microsoft word 3-王剑飞.doc 14-18 advances in systems science and applications (2010), vol.10, no.1 issn 1078-6236 international institute for general systems studies, inc. extensions of liouville’s theorem jianfei wang applied science college, harbin university of science and technology, harbin 150080, china email: jianfei1965@sohu.com abstract in this paper, we first generalize liouville’s theorem into the general forms based on power series representations for analytic functions. second, in simply connected domains harmonic functions can be identified as real parts of analytic functions. observing the relations between analytic functions and harmonic functions, we extend liouville’s theorem to harmonic functions by the harnack’s inequality. the generalized liouville’s theorems obtained in this paper will help us to further study the properties of entire functions and harmonic functions. keywords liouville’s theorem entire function extension harmonic function harnack’s inequality analytic 1. introduction if f (z) is analytic on the whole complex plane, then it is said to be an entire function. for entire functions, there exists a beautiful theorem, known as liouville’s theorem. it gives many important properties of entire functions, these properties have been widely applied to complex analysis. and it enables us to prove the fundamental theorem of algebra. therefore, it is necessary for us to further discuss liouville’s theorem. in this paper, we first derive two extension forms of liouville’s theorem by simplifying some conditions of classical liouville’s theorem. meanwhile we get two results for analytic functions. second, in simply connected domains harmonic functions can be identified as real parts of analytic functions. and there are many consequences for analytic functions. some of these are the infinite differentiability of analytic functions, liouville’s theorem, and the maximum modulus theorem. hence we think that these results have analogues for harmonic functions. in terms of the ideas, we extend liouville’s theorem to harmonic functions. the generalized liouville’s theorems provide theoretical basis for us to further study the properties of entire functions and harmonic functions. 2. inequalities and lemmas we list some useful inequalities and lemmas before giving to the generalized liouville’s theorems. theorem 2.1 (liouville’s theorem) [1] the only bounded entire functions are the constant functions. lemma 2.1 [2] let ϕ be harmonic on a simply connected domain d . then there is an analytic function f such that re fϕ = on d . lemma 2.2 (mean-value theorem for harmonic function) [3] let ϕ be harmonic in a domain containing the dish z r≤ . then 2 0 1(0) ( ) 2 i tre dt π ϕ ϕ π = ∫ (1) lemma 2.3 (poisson integral formula) [2] let ϕ be harmonic in a domain containing the dish z r≤ . then for 0 r r≤ < , we have advances in systems science and applications (2010), vol.10, no.1 15 ( ) ( ) ( ) 2 2 2 2 202 2 cos i t i rer rre dt r r rr t πθ ϕ ϕ π θ − = + − −∫ (2) proof: using lemma 2.1, we have re fϕ = here f is analytic on a simply connected domain d . (assuming the domain d includes the circle :rc z r= as well as its interior) applying cauchy integral formula, we obtain ( ) ( ) ( )1 2 rc f f z d z r i z ζ ζ π ζ = < −∫ (3) for fixed z , with z r< , the function ( ) 2 f z r z ζ ζ− is an analytic function of ζ inside and on rc . hence by cauchy theorem ( ) 2 1 0 2 rc f z d i r z ζ ζ π ζ = −∫ (4) we add it to equation (eq.) (3): ( ) ( )2 1 1 2 rc zf z f d i z r z ζ ζ π ζ ζ ⎛ ⎞ = +⎜ ⎟⎜ ⎟− −⎝ ⎠ ∫ ( )( ) ( ) 22 2 1 2 rc r z f d i z r z ζ ζ π ζ ζ − = − −∫ (5) if we parameterize rc by , 0 2 ,i tre tζ π= ≤ ≤ eq. (5) becomes ( ) ( )( ) ( ) 22 2 20 1 2 i t i t i t i t r z f z f re r e dt re z r re z π π − = − −∫ (6) ( ) ( )( ) 22 2 02 i t i t i t f rer z dt re z re z π π − − = − −∫ (7) ( )22 2 202 i t i t f rer z dt re z π π − = − ∫ (8) writing z in the polar form iz re θ= , we have ( ) ( )2 2 2 202 i t i i t i f rer rf re dt re re πθ θπ − = − ∫ (9) ( ) ( ) 2 2 2 2 202 2 cos i tf rer r dt r r rr t π π θ − = + − −∫ (10) finally, by taking the real part of this equation, we arrive at poisson integral formula ( ) ( ) ( ) 2 2 2 2 20 , , 2 2 cos r tr rr dt r r rr t π ϕ ϕ θ π θ − = + − −∫ (11) or, equivalently, ( ) ( ) ( ) 2 2 2 2 202 2 cos i t i rer rre dt r r rr t πθ ϕ ϕ π θ − = + − −∫ (12) poisson integral formula expresses the values of a harmonic function in a region is wang: extensions of liouville’s theorem 16 completely determined by its values on the boundary. using poisson integral formula and mean value theorem for harmonic function, we may derive the following important inequality. theorem 2.2 (harnack’s inequality) [2] let ϕ be harmonic and nonnegative in a domain containing the dish z r≤ . then for 0 r r≤ < , we have ( ) ( ) ( )0 0ir r r rre r r r r θϕ ϕ ϕ− + ≤ ≤ + − (13) proof: applying lemma 2.3 to harmonic functionϕ . then for 0 r r≤ < , ( ) ( ) ( ) 2 2 2 2 202 2 cos i t i rer rre dt r r rr t πθ ϕ ϕ π θ − = + − −∫ (14) observing that ( ) ( ) ( )2 22 2 2 cosr r r r rr t r rθ− ≤ + − − ≤ + (15) since ϕ is nonnegative, we have ( ) ( ) ( ) ( ) ( ) ( ) 2 2 2 2 2 2 20 0 02 cos i t i t i tre re re dt dt dt r r rr tr r r r π π πϕ ϕ ϕ θ ≤ ≤ + − −+ −∫ ∫ ∫ (16) hence ( ) ( ) ( )2 2 0 0 1 1 2 2 i t i i tr r r rre dt re re dt r r r r π πθϕ ϕ ϕ π π − + ⋅ ≤ ≤ ⋅ + −∫ ∫ (17) finally, by lemma 2.2, the proof is completed. 3. the generalized liouville’s theorems in this section, we consider the extension problem of classical liouville’s theorem. this is the main work of this paper. firstly, we obtain two extension forms of liouville’s theorem by reducing the condition of classical liouville’s theorem. these theorems are stated below. theorem 3.1 i f f is an entire function, and suppose there are a nonnegative integer n , and two positive constants ,r m such that ( ) nf z m z≤ when z r≥ , then f is a polynomial with ( )deg f n≤ or constant. proof: the case 0n = will be treated first. in this case, ( )f z m≤ when z r≥ . since f is an entire function, it must be continuous on the closed domain z r≤ .under such circumstances it is known from calculus that the function must be bounded there. in other words, there exists a positive constant g such that ( )f z g≤ when z r≤ . by taking { }max , 0n m g= > , we have ( )f z n≤ (18) when z < +∞ . thus from theorem 2.1 there follows f is constant. this completes the proof. then the general case 1n ≥ will be dealt with. since f is an entire function, it must have a maclaurin series representation, namely ( )2 0 1 2( ) n nf z c c z c z c z z= + + + + + < +∞ (19) where ( ) ( ) ( ) ( ) 1 0 1 0; 0,1, 2, ! 2 n n nz r f f c d r n n i ζ ζ π ζ += = = > =∫ (20) for an arbitrary integer 1p ≥ and a large enough 0r r> , we consider that advances in systems science and applications (2010), vol.10, no.1 17 ( ) 0 1 1 2n p n pz r f c d i ζ ζ π ζ+ + += = ∫ (21) notice that ( ) 0 1 1 2n p n pz r f c ds ζ π ζ + + += ≤ ∫ 0 1 1 2 n n pz r m d s ζ π ζ + += ≤ ∫ 01 0 0 1 2 2 p p m mr r r π π += ⋅ ⋅ = (22) we can get ( )0 1,2,n pc p+ = = . thus ( )2 0 1 2( ) n nf z c c z c z c z z= + + + + < +∞ (23) it says that f is a polynomial with ( )deg f n≤ . by modifying the condition of theorem 3.1, we easily obtain the following results. corollary 3.1 i f f is an entire function, and suppose there are a nonnegative integer n , and three positive constants ,r m and n such that ( ) nf z n m z≤ + when z r≥ , then f is a polynomial with ( )deg f n≤ or constant. corollary 3.2 i f f is an entire function, and there is a positive integer n such that ( )lim 0nz f z k z→∞ = > , then f is a polynomial with ( )deg f n≤ . theorem 3.2 let f be analytic in the extended complex plane. then f is constant. proof: since f is analytic in the extended complex plane, it must have a maclaurin series representation, namely ( ) 0 ( ) n n n f z c z z ∞ = = < +∞∑ (24) meanwhile, z = ∞ is a removable singularity of f . hence ( )f z has a finite limit as z approaches 0z . recalling (24), the conclusion is obtained. secondly, in simply connected domains harmonic functions can be identified as real parts of analytic functions. based on the relations between analytic functions and harmonic functions, harmonic function in 2r is considered [4] in generalizing liouville’s theorem. we state liouville’s theorem for harmonic functions as follows. theorem 3.3 (liouville’s theorem for harmonic functions) [5] let ϕ be harmonic in the whole real plane 2r and bounded from above or below there. thenϕ is constant. proof: assume that ϕ is bounded from above there, namely, there exists a constant m such that mϕ ≤ for any 2z r∈ . clearly, mψ ϕ= − is harmonic and nonnegative in the whole real plane 2r . using theorem 2.2, for 0 r r≤ < < +∞ , we have ( ) ( ) ( )0 0ir r r rre r r r r θψ ψ ψ− + ≤ ≤ + − (25) letting r →+∞ , deduce that ( ) ( )0ire θψ ψ= (26) where [ )0,r∈ +∞ . wang: extensions of liouville’s theorem 18 notice that r is an arbitrary nonnegative real number hence mψ ϕ= − is constant. it follows immediately that ϕ is constant. 4. conclusions summing up, we first transform some useful inequalities and lemmas to generalize liouville’s theorem. then using power series representations for analytic functions [6] , we derive two extension forms of liouville’s theorem by simplifying some conditions of classical liouville’s theorem. finally, observing the relations between analytic functions and harmonic functions, we extend liouville’s theorem to harmonic functions by harnack’s inequality. in this course, we know that the generalized liouville’s theorems will help us to further study the properties of entire functions. meanwhile, we can get some important conclusions for harmonic functions [7] based on the consequences for analytic functions in the future. references [1] yuquan zhong. functions of complex variable. higher education press, beijing, 2004: 127-130. [2] saff, e. b. fundamental of complex analysis with applications to engineering and science. china machine press, beijing, 2004: 221-226. [3] jiarong yu. functions of complex variable. higher education press, beijing, 2000: 156-157. [4] weixing dai, shaobo zhou. attraction and stability for neutral stochastic differential delay equations. advances in systems science and applications, 2008, 8(2): 212-219. [5] zhenhua jiao. a note on liouville’s theorem. journal of hangzhou dianzi university, 2006, 26(2): 96-98. [6] jiancong chen, zhihong guan and hu chen. fuzzy association analysis of complex system in correlation. advances in systems science and applications, 2008, 8(2): 251-257. [7] elias zafiris. categorical modeling of natural complex systems. part i: functorial process of representation. advances in systems science and applications, 2008, 8(2): 187-200. advances in systems science and application (2016) vol.16 no.2 94-102 driver drowsiness detection system meriem boumehed1, belal alshaqaqi2, abdullah salem baquhaizel2 and mohamed el amine ouis2 1preparatory school in science and technology 2university of sciences and technology of oran mohamed boudiaf (usto-mb) oran, algeria abstract drowsiness and fatigue of drivers are amongst the significant causes of road accidents. every year, they increase the amounts of deaths and fatalities injuries globally. in this paper, a module for advanced driver as-sistance system (adas) is presented to reduce the number of accidents due to drivers fatigue and hence increase the transportation safety; this system deals with automatic driver drowsiness detection based on visual information and artificial intelligence. we proposed an algorithm to locate, track, and analyze both the drivers face and eyes to measure percentage of eye dosure, a scientifically supported measure of drowsiness associated with slow eye closure. keywords drowsiness detection, adas, face detection and tracking, eyes detection and tracking 1 introduction currently, transport systems are an essential part of human activities. we all can be victim of drowsiness while driving, simply after a night’s sleep too short, altered physical condition or during long journeys. the sensation of sleep reduces the driver’s level of vigilance producing dangerous situations and increases the probability of an occurrence of accidents. driver drowsiness and fatigue are among the important causes of road accidents. every year, they increase the number of deaths and fatalities injuries globally. in this context, it is important to use new technologies to design and build systems that are able to monitor drivers and to measure their level of attention during the entire process of driving. in this paper, a module for adas (advanced driver assistance system) is presented in order to reduce the number of accidents caused by driver fatigue and thus improve road safety. this system treats the automatic detection of driver drowsiness based on visual information and artificial intelligence. we proposed an algorithm to locate, track and analyze both the driver face and eyes to measure percentage of eye closure. the remainder of this paper is organized as follows, section 2 presents the related works, section 3 presents the proposed system and the implementation of each block of the system, the experimental results are shown in section 4 and in the last section conclusions and perspectives are presented. advances in systems science and application (2016) vol.6 no.2 95 2 related works some efforts have been reported in the literature on the development of the notintrusive monitoring drowsiness systems based on the vision. malla et al.[1] develop a light-insensitive system. they used the haar algorithm to detect objects[2] and face classifier implemented by[3] in opencv[4] libraries. eye regions are derived from the facial region with anthropometric factors. then, they detect the eyelid to measure the level of eye closure. vitabile et al. implement a system to detect symptoms of driver drowsiness based on an infrared camera[5]. by exploiting the phenomenon of bright pupils, an algorithm for detecting and tracking the driver’s eyes has been developed. when drowsiness is detected, the system warns the driver with an alarm message. bhowmick et kumar use the otsu thresholding to extract face region[6,7]. the localization of the eye is done by locating facial landmarks such as eyebrow and possible face center. morphological operation and k-means is used for accurate eye segmentation. then a set of shape features are calculated and trained using non-linear svm to get the status of the eye. hong et al. define a system for detecting the eye states in real time to identify the driver drowsiness state[8]. the face region is detected based on the optimized jones and viola method[2]. the eye area is obtained by an horizontal projection. finally, a new complexity function with a dynamic threshold to identify the eye state. tian et qin build a system that checks the driver eye states. their system uses the cb and cr components of the ycbcr color space[9]. this system locates the face with a vertical projection function, and the eyes with a horizontal projection function. once the eyes are located the system calculates the eyes states using a function of complexity. under the light of what has been mentioned above, the identification of the driver drowsy state given by the perclos is generally passed by the following stages: 1) face detection, 2) eyes location, 3) face and eyes tracking, 4) identification of the eyes states, 5) calculation of perclos and identification the driver state. 3 the proposed system in this section, we discuss our presented system which detects driver drowsiness. the overall flowchart of our system is shown in fig 1. 96 meriem b., belal a., abdullah salem b, etc: driver drowsiness detection system a. face detection the symmetry is one of the most important facial features.we modeled the symmetry in a digital image by a one-dimensional signal (vector accumulator) with a size equal the width of the image, which gives us the value corresponding to the position of the vertical axis of symmetry of objects in the image. the traditional principle to calculate the signal of symmetry is for each two white pixels which are on the same line we increment the value in the medium between these two pixels in the accumulating vector. (the algorithm is applied on an edge image, we called a white pixel: the pixel with value 1). we introduce improvements on the calculation algorithm of symmetry into an image to adapt it to the detection of face, by applying a set of rules to provide a better calculation of symmetry of the face. instead of computing the symmetry between two white pixels in the image, it is calculated between two windows (z1 and z2) (fig.2). fig. 1 flowchart of the proposed system for each window z1, we sweep the window z2 in the area determined by the parameters s min, s max, and h. we increment the signal of symmetry between advances in systems science and application (2016) vol.6 no.2 97 these two windows if the sum of white pixels is located between two thresholds s1 (maximum) and s2 (minimum). fig. 2 the new method to improve the calculation of the symmetry in the image (a) original image (b) edge detection (c) symmetry signal (d) localization of the maximum of symmetry (e) region of interest roi (f) result fig. 3 face detection using symmetry then we extract the vertical region of the image contours (region of interest roi) corresponding to the maximum index of the obtained signal of symmetry. next, we take a rectangle with an estimated size of face (because the camera is fixed and the driver moves in a limited zone so we can estimate the size of the face using the camera focal length after the step of camera calibration) and we 98 meriem b., belal a., abdullah salem b, etc: driver drowsiness detection system scan the roi by searching the region that contains the maximum energy corresponding to the face (fig.3). we propose a checking on two axes: the position variance of the face detected according to time; i.e., in several successive images, it is necessary that the variance of the positions of the detected face is limited; because the speed of movement of the face is limited of some pixels from a frame to another frame which follows. b. eyes localization since the eyes are always in a defined area in the face (facial anthropometric properties), we limit our research in the area between the forehead and the mouth (eye region of interest ’eroi’) (fig.4.a). we benefit from the symmetrical characteristic of the eyes to detect them in the face. first, we sweep vertically the eroi by a rectangular mask with an estimated height of height of the eye and a width equal to the width of the face, and we calculate the symmetry. the eye area corresponds to the position which has a high measurement of symmetry. then, in this obtained region, we calculate the symmetry again in both left and right sides. the highest value corresponds to the center of the eye. the result is shown in fig.4.b. (a) eroi (b) result fig. 4 eyes localization using symmetry c. tracking the tracking is done by template matching using the sad algorithm (sum of absolute differences). sad(x, y) = n∑ i=1 m∑ i=1 |i(x+ i, y + j)−m(i, j)| (1) we proposed to make a regular update of the reference model m to adjust it every time when light conditions changes while driving, by making a tracking advances in systems science and application (2016) vol.6 no.2 99 test: tracking = { good if sad ≤ th bad if sad > th (2) d. eyes states the determination of the eye state is to classify the eye into two states: open or closed. we use the hough transform for circles [10] (htc) on the image of the eye to detect the iris. for that, we apply the htc to the edge image of the eye to detect the circles with defined rays, and we take at the end the circle which has the highest value in the accumulator of hough for all the rays. then, we apply the logical ’and’ logic between edges image and complete circle obtained by the htc by measuring the intersection level between them “s”. finally, the eye state “stateeye” is defined by testing the value “s” by a threshold: stateeye { open if s ≤ th closed if s ≥ th (3) the results are shown in fig.5. (a) and (b) edge detection (c) and (d) eyes states results fig. 5 eyes states using htc e. driver state we determine the driver state by measuring perclos. if the driver closed his eyes in at least 5 successive frames several times over a period of up to 5 seconds, 100 meriem b., belal a., abdullah salem b, etc: driver drowsiness detection system it is considered drowsy. 4 experimental results to validate our system (fig.6), we test on several drivers in the car with real driving conditions. we use an ir camera with infrared lighting system operates automatically under the conditions of reduced luminosity and night even in total darkness. fig. 6 our system installed in the car based on ir camera the results of the eye states are illustrated in table 1. table 1 results obtained from the system driver frames number false eyes sates false rate open closed d1/day 420 17 0 4% d2/day 430 15 0 3.50% d3/day 245 7 1 3.20% d1/night 200 3 1 2% d2/night 200 1 0 0.50% d3/night 200 6 3 4.50% according to the obtained results, our system can determine the eye states advances in systems science and application (2016) vol.6 no.2 101 with a high rate of correct decision. 5 conclusion and perspectives in this paper, we presented the conception and implementation of a system for detecting driver drowsiness based on vision that aims to warn the driver if he is in drowsy state. this system is able to determine the driver state under real day and night conditions using ir camera. face and eyes detection are implemented based on symmetry. hough transform for circles is used for the decision of the eyes states. the results are satisfactory with an opportunity for improvement in face detection using other techniques concerning the calculation of symmetry. moreover, we will implement our algorithm on a dsp (digital signal processor) to create an autonomous system working in real time. references [1] a. malla, p. davidson, p. bones, r. green and r. jones. (2010), “automated video-based measurement of eye closure for detecting behavioral microsleep”, in 32nd annual international conference of the ieee, buenos aires, argentina. [2] p. viola and m. jones. (2001), “rapid object detection using a boosted cascade of simple features”, in proceedings of the ieee computer society conference on computer vision and pattern recognition. [3] r. lienhart and j. maydt. (2002), “an extended set of haar-like features for rapid object detection”, in proceedings of the ieee international conference on image processing. [4] open source computer vision library: reference manual, available at: http://opencvlibrary.sourceforge.net [5] s. vitabile, a. paola and f. sorbello. (2010), “bright pupil detection in an embedded, real-time drowsiness monitoring system”, in 24th ieee international conference on advanced information networking and applications. [6] b. bhowmick and c. kumar. (2009), “detection and classification of eye state in ir camera for driver drowsiness identification”, in proceeding of the ieee international conference on signal and image processing applications. [7] n. otsu. (1979), “a threshold selection method from gray-level histograms”, ieee transactions on systems,man and cybernatics, pp.62-66. 102 meriem b., belal a., abdullah salem b, etc: driver drowsiness detection system [8] t. hong, h. qin and q. sun. (2007), “an improved real time eye state identification system in driver drowsiness detection”, in proceeding of the ieee international conference on control and automation, guangzhou, china. [9] z. tian et h. qin. (2005), “real-time driver’s eye state detection”, in proceedings of the ieee international conference on vehicular electronics and safety. [10] m. s. nixon and a. s. aguado. (2008), feature extraction and image processing(2nded.), jordan hill, oxford ox2 8dp, uk. corresponding author belal alshaqaqi can be contacted at: alshaqaqi belal@hotmail.fr microsoft word 7 s. x. mei j. l. xie--3d numerical simulation in combustion space of an oxy-fuel glass furnace.doc advances in systems science and applications (2011), vol.11, no.3-4 257-263 issn 1078-6236 international institute for general systems studies, inc 3d numerical simulation in combustion space of an oxy-fuel glass furnace s. x. mei and j. l. xie school of materials science and engineering, wuhan university of technology, hubei, wuhan 430070, china abstract in glass manufacturing, oxy-fuel combustion is a new technology. in design of oxy-fuel glass furnace, the numerical simulation and virtual reality techniques have become necessary tools. in this paper, the numerical simulation in the combustion space of a designed oxy-fuel glass furnace was carried out with the results displayed by the virtual reality technology. the gas phase was expressed with k-ε two-equation model; the combustion was described with non-premixed model; the radiation was expressed with discrete ordinates radiation model. the results of simulation agree well with the related reference, showing that the flame from each burner is individual and similar, and most fuel streams combust adequately with high temperature except that from the burner 1, which is suggested to be moved to other place. what we learned from the simulation results can give direct insight in the combustion space of the furnace, and can be used to guide the design of the oxy-fuel glass furnace. keywords glass furnace; oxy-fuel; combustion space; virtual reality; numerical simulation 1.introduction in glass manufacturing, the oxy-fuel furnaces, using pure oxygen instead of air as the primary oxidant, began to rapidly instead of the conventional air-fuel furnaces since 1990 because of many benefits, such as glass quality improvement, fuel reduction, productivity increase, emissions reduction (nox, so2, particulates), and expansion of the existing furnace [1, 2, 3]. the oxy-fuel technology is different from the air-fuel technology, to meet higher quality of the produced glass, reduce energy consumption, and prolong furnace life, it is important to obtain a rational flow field in the combustion space of the furnace. in the past,most designs of the glass furnaces were based on empirical rules or traditional methods, being difficult to obtain information such as temperatures and pressures in the combustion space of the furnace. and now with the computational fluid dynamics (cfd) modelling has been widely used to predict and estimate the flow field of the glass furnace [4, 5, 6], it is more scientific to design and optimize the oxy-fuel glass furnaces using mathematical simulation. in this paper, an oxy-fuel glass furnace was designed, and the numerical simulation in the combustion space of the furnace was carried out. by using the virtual reality technology for the simulation results, the fields of temperature, velocity and species concentration were displayed clearly. what we learned from the simulation results can give direct “insight” in the furnace, and can be used to guide the design of the oxy-fuel glass manufacturing technology. 2.geometrical model fig. 1 shows schematically the configurations of the combustion space of the glass furnace. there are seven pairs of staggered burners, and two exhaust ports which are between the first pairs of burners. fig. 2 shows the mesh. structural hexahedral grid was used in the whole combustion space with mesh refined around the burners. 258 mei:3d numerical simulation in combustion space of an oxy-fuel glass furnace fig.1. configurations of the combustion space of the glass furnace fig.2. meshes 3. mathematical model there occur fluid flow, heat transfer, and combustion phenomena inside the combustion space. to deal with turbulence, thermal radiation, and combustion, some submodels are as follows. 3.1 turbulence model in eularian system we solve the fluid phase continuity and momentum equations using the k-ε model, which is widely used in engineering [7, 8, 9]. the general form of the governing equations for the gas phase is given as follow: ( ) ,/ [ ( / )] /j j j j pv x x x s sϕ ϕ ϕρ ϕ ϕ∂ ∂ = ∂ γ ⋅ ∂ ∂ ∂ + + (1) where ρ is the fluid density, φ is the general different variable, γφ is the effective viscosity, sφ is the source term of the gas phase, sp,φ, the sourse term from the interaction with the discrete phase. 3.2 combustion model combustion is modeled by the non-premixed modeling, which involves the solution of transport equations for one or two conserved scalars (the mixture fractions f). equations for individual species are not solved. instead, species concentrations are derived from the predicted mixture fraction fields. interaction of turbulence and chemistry is accounted for with an assumed-shape probability density function (pdf). the mean (density-averaged) mixture fraction equation is: ( ) / ( ) [( / ) ]t t mf t vf f s sρ ρ μ σ ′∂ ∂ +∇⋅ = ∇ ⋅ ⋅∇ + + (2) where the source term sm is due solely to transfer of mass into the gas phase from liquid fuel droplets or reacting particles (e.g., coal), and s ′ is any other source term. 3.3 radiation model radiation heat transfer from surface to surface and radiation absorption and emission by h2o and co2 in the combustion atmosphere is modeled with the discrete ordinates (do) radiation model [10], which solves the radiative transfer equation for a finite number of discrete solid angles, each associated with a vector direction s fixed in the global cartesian system (x,y,z). the do model transforms equation (7) into a transport equation for radiation intensity in advances in systems science and applications (2011), vol.11, no.3-4 259 the spatial coordinates (x,y,z). the do model solves for as many transport equations as there are directions s . the solution method is identical to that used for the fluid flow and energy equations. ( ) ( ) ( ) ( ) ( ) 42 4 0 , / , / [ , , ] / 4s sdi r s ds a i r s an t i r s s s d π σ σ π σ π′ ′ ′+ + = ⋅ + ⋅ φ ω∫ (3) where r is the position vector, s is the direction vector, s ′ is the scattering direction vector, s is the path length, a is the absorption coefficient, n is the refractive index, σs is the scattering coefficient, σ is the stefan-boltzmann constant (5.672 * 10-8 w/m2-k4), i is the radiation intensity, which depends on position ( r ) and direction ( s ), t is the local temperature,φis the phase function, and ω′ is the solid angle. 4. boundary conditions and numerical solution method the numerical procedure was based on a well-known finite volume method. to solve the elliptic form of differential equations, appropriate boundary conditions are required. (i) the fuel is natural gas with the distribution listed in table 1, the oxidizer is pure oxy. at the inlet, all velocities and temperature were specified. (ii) at the outlet, the pressure was at an ambient atmosphere. (iii) at an impenetrable wall of the furnace, the wall temperatures on different region were specified, and the usual non-slip conditions were applied. to account for the wall effect in the nearby regions, the wall-function was introduced to link velocities in the near-wall region. the flow field equations with boundary conditions were solved numerically. the process was repeated until convergence was achieved table 1 fuel distribution ratio (%) no. 1# 2# 3# 4# 5# 6# 7# distribution ratio 6 13 9.5 7.5 6 4.5 3.5 no. 1’# 2’# 3’# 4’# 5’# 6’# 7’# distribution ratio 6 13 9.5 7.5 6 4.5 3.5 5. results and discussions to visualize the flame envelope, concentration of co (carbon monoxide) can be used as an indicator since the flame is rich in soot, which is formed in fuel-rich regions. fig. 3 displays the iso-surface of co concentration of 0.04. we can see the “flame” shape from each burner. fig. 4 shows the temperature contours through the burner plane, and fig. 5 shows the temperature contours at the vertical middle slice of burner 2. fig.3. iso-surface of co concentration of 0.04 figs. 3-5 show that not only the “flame” shape, but also the temperature field agree well with the related reference [11], indicating the reliability of the simulation results. from figs. 3 and 4 we can see that the length and width of each “flame” from burners 2-7 correspond with the fuel 260 mei:3d numerical simulation in combustion space of an oxy-fuel glass furnace distribution ratio listed in table 1, while the length of the “flame” from burner 1 is the shortest, indicating the incomplete combustion of the fuel. fig.4. temperature contours through the burner plane fig.5. temperature contour at the vertical middle slice of burner 2 fig. 6 shows the velocity vector on the vertical middle slice of burner 2, fig.7 displays the streamlines from burners 1-7. from fig. 6 we can see that the gas stream runs straightly from the inlet to the opposed wall, and then rebounds when meets the wall, resulting in a wider sized reversed flows under the crown. from fig.7 it can be observed that there is large backflow under the crown in the whole space, being beneficial for decreasing the crown heat duty. fig.6. velocity vector at the vertical middle slice of burner 2 fig.7. streamlines from burners 1-7 fig.8. temperature contours at the bottom surface figs. 8-10 show the temperature iso-lines at the bottom surface, the cross middle slice of burners and the crown, respectively. from fig 9 it can be observed that the crest “flame” advances in systems science and applications (2011), vol.11, no.3-4 261 temperature is around the burner 2, resulting in the corresponding local high temperature region not only on the bottom (see fig. 8) but also on the crown (see fig. 10). except the high temperature region, in other regions the temperature on the bottom and on the crown is uniform. fig.9. temperature iso-lines at the cross middle slices of burners fig.10. temperature iso-lines at the crown to know the temperature field along the length direction (x direction) in the whole combustion space, we created a series of vertical slices, of which the average temperature values are shown in the curve of fig. 11. it can be observed that there are 13 temperature crests corresponding with 13 burners, except the burner 1, indicating the incomplete combustion process there, which should be improved. the maximum of the temperature crest is near the burner 2, around which the expected hot-spot located, indicating the rationality of the results. fig.11. average temperatures at the vertical slices along x direction the above analyses results show that the combustion condition from the burner 1 is quite bad. it is necessary to find the reason. fig. 12 displays the velocity vector through the burner plane, fig. 13 shows the velocity vector at the vertical middle slice of exhaust port, and fig. 14 displays the co2 concentration through the burner plane. from figs. 12-14 it can be known that since the burner 1 is near the outlet 1, and the fuel distribution ratio of burner 1 is much less than that of the neighbor burners (2# and 2’# ), when large flue gas pass by the burner 1 (see fig. 14), and run to the outlets finally, the oxy is insufficient around the burner 1, resulting in the incomplete combustion. to improve this condition, it is suggested to move the burner 1 to other place. 262 mei:3d numerical simulation in combustion space of an oxy-fuel glass furnace fig.12. velocity vector through the burner plane fig.13. velocity vector at the vertical middle slice of exhaust ports fig.14. co2 concentration through the burner plane 6. conclusion in this paper, the numerical simulation in the combustion space of a designed oxy-fuel glass furnace was carried out. the virtual reality techniques was used to display the simulation results which are agree well with the related reference. the results show that the “flame” from each burner is individual and similar on the whole, and most fuel streams combust adequately with high temperature except that from the burner 1. the maximum flame temperature in the oxy-flames is higher than that in the air-fired furnace, and the maximum average temperature along the length direction is near the second pairs of burners, corresponding to the predicted region of hot-spot. there are back flows occurring near the crown, being beneficial for decreasing the crown heat duty. to obtain more rational flow field the placement of burner 1 will be discussed in future works. acknowledgements the authors are grateful for the supports provided by the national key technology r&d program (2006baf02a26). references [1] simpson, neil g., wilcox richard, et al., oxy-fuel technologies for boosting and 100% conversions of cross fired furnaces. glass technology: european journal of glass science and technology part a. 48(4) (2007) 168-175. [2] habel michael, lievre kevin, inskip julian, et al., advanced cleanfire® hri oxy-fuel boosting application lowers emissions and reduces fuel consumption. ceramic engineering and science proceedings: 68th conference on glass problems a collection of papers presented at the 68th conference on glass problems. 29(1) (2008) 203-211. advances in systems science and applications (2011), vol.11, no.3-4 263 [3] kobayashi h, evenson e., and xue y., development of an advanced batch/cullet preheater for oxy-fuel fired glass furnaces. ceramic engineering and science proceedings: 68th conference on glass problems a collection of papers presented at the 68th conference on glass problems. 29(1) (2008) 137-148. [4] s.l. chang, c.q. zhou, and b. golchert, eulerian approach for multiphase flow simulation in a glass melter. applied thermal engineering. 25 (2005) 3083–3103. [5] vishal sardeshpande, u.n. gaitonde, and rangan banerjee, energy conversion and management, 48(10) (2007) 2718-2738. [6] a. abbassi, and kh. khoshmanesh, numerical simulation and experimental analysis of an industrial glass melting furnace. applied thermal engineering. 28(5-6) (2008) 450-459. [7] u. shah, c. zhang, j. zhu, et al., validation of a numerical model for the simulation of an electrostatic powder coating process. international journal of multiphase flow. 33(5) (2007) 557-573. [8] d.m. hargreaves, and n.g. wright, on the use of the k–ε model in commercial cfd software to model the neutral atmospheric boundary layer. journal of wind engineering and industrial aerodynamics. 95(5) (2007) 355-369. [9] fatih üneş, investigation of density flow in dam reservoirs using a three-dimensional mathematical model including coriolis effect. computers & fluids. 37(9) (2008) 1170-1192. [10] jian c q, dutta a, mukhopadhyay a, et al., explicit coulpling between combustion space and glass tank simulation for complete furnace analysis. presented at the 1st balkan conference on glass science and technology , volos , greece ,2000. [11] lankhorst a. m., bauer r. a., coupled combustion modeling and glass tank modeling in oxy and air-fired glass melting furnaces. the 4th int sem on mathematical simulation in glass melting ,1997. advances in systems science and applications (2013) vol.13 no.3 227-232 thoughts on the general systems theory michael v. tokarev corporation axis ltd. miami, fl, s.-petersburg, russia abstract the most common definition of systems is formulated as “a combination of elements”, or even as “any object is a system”. many authors have presented the formal definitions of an aggregation of two multitudes: elements and relationships. but a single, universally accepted systems definition still does not exist. in this article, why this happens, what matters in definitions do not have the answers yet, how “general systems theory” metascience and “systems engineering” particularistic science relate, the paradoxes of general systems theory, and much more will be discussed. keywords systems’ definition, general systems theory 1 introduction known classic of general systems theory v.n. sadowski brings dozens of existing systems’ definitions [1]. other authors add new definitions, but a single, universally accepted systems definition still does not exist. why this happens, what matters in definitions do not have the answers yet, how “general systems theory” metascience and “systems engineering” particularistic science relate, the paradoxes of general systems theory [1], and much more will be discussed in this article. these issues arise primarily in the course of the analysis of existing definitions and their research by v.n. sadowski in the book “foundations of general systems theory”. from our point of view as long as it is the best and most comprehensive study to date of the existing state of things in terms of the general system theory (gst). the purpose of this paper is to formulate the most pressing issues in the context of system definitions and attempt to identify ways to address them. 2 existing definitions and issues arising from them. entropy the most common definition of systems is formulated as “a combination of elements”, or even as “any object is a system”. many authors have presented the formal definitions of an aggregation of two multitudes: elements and relationships [2]. our task is not to refute or criticize existing definitions. we’re just going to try to formulate questions to the definitions and try to find answers to those for which it is possible within this article. in the analysis of any definition there is the first important question, the answer to which is connected to all systems definitions and classifications. this michael v. tokarev: thoughts on the general systems theory 228 question will be worded as follows: “does a system really exist?” not many researchers ask this question, but those who are wondering, as a rule, respond to it positively. it seems to us that the question can’t be answered at all. why? let’s think constructively. were there systems before the mankind? if we answer in the affirmative (systems have always existed), what is the point in trying to determine them? they already exist, someone created them, and therefore defined? and what about those systems that the man himself creates? and what about those systems, which include both elements: those existed prior humans and artificial elements? this question is based on the idea that if we say that the system is there, so we know what we say, and this system has already been described by someone and you can explore it as a system. obviously, prior to mankind, to be exact, even to the end of xix century, no one knew what the system in the modern sense were, did not describe them. so, we can only speak of the objects or entities existence, but not the system. if the answer to the existence of systems prior humans is negative ( system concept was invented by a man, and prior him the concept did not exist), then we are just coming to terms of kant transcendental idealism, according to which the person is not doing anything, but only “reflects”. it seems to us that the answer to this fundamental question must be sought in the following. a man created the concept of a system. he “adjusts” the existing reality to his concept. in fact we find the objects, try to assign them the system properties (which we have invented ourselves) and state that systems exist. it’s fine if scientists understand that examine just a model, not the system. the researchers come to paradoxical conclusions: for example, the type (genus, family) of animals-is also a system. obviously, from the point of view of nature of it is not so, because the concept of genus, species, family also came up with people. certainly, nature was not busy with splitting all living beings into the genera and species. in fact, these systems are completely virtual. otherwise we’ll have to admit a creator, who invented not only flora and fauna, but also classified them. thus, if we wave away the non-scientific idea of a universal creator, we are forced to admit that the system did not exist before humans and does not exist now. here we see the fundamental paradox of the system science: everybody talks about a non-existing, and not only talks, but also uses in their practice and research. of course, aside from all of this there are artificial systems that people originally created as such. but in this case, and the conceptual apparatus is immediately objective. in particular, if we build a car as a system of interacting elements that has to deliver goods from point a to point b, we define the system on the basis of interacting elements, goals and functions. 229 advances in systems science and applications (2013) vol.13 no.3 thus, solving this major issue, we split at least (and fundamentally) natural and artificial systems, knowing that the latter are systems a priori. now we continue the argument about natural systems (which exist only in our minds?). usually, natural systems are classified as animate (living) and inanimate (nonliving) ones. if every living individual can be somehow dragged into the system concept, the inanimate objects can be appealed to only in the form of models, as we do not know enough about them, about the hierarchy of these systems, and possibly even the dimension in which they exist. for example, a stone lying in the mountains. is it a system? from the point of view of the most general definition (“any object is a system”) a stone is undoubtedly the system. from the point of view of the extension of the concept (for example, an observer or function), the stone can’t be considered a system, until we started to study it and / or until it is comes into operation. of course, we can think of the functions of internal heat, molecules and atoms inside it, the crystal structure, etc. but is it a system while it is lying, and it is not being studied? in that case, why bother to talk about such an object as a system? only because of a common definition of the system itself? no less obvious the application of the notion “system” to objects of fauna. is a particular bird a system? from a biological point of view (only because it is more convenient to be studied by a man)-of course. and from the point of view of other fauna? possible-it is also a system. a bird feeds, feed others, in addition it gets a lot more besides input actions. we may say that here the interpretations of definitions of the hierarchy, morphology and interaction of elements, the micro and macro levels are performed. one point confuses us: the bird itself does not know it and, moreover, does not take any concerted action to be in the system. just feeds, breeds and dies. from this point of view, it is like the inanimate nature, stone, which also does not “know” that it is the system. there is only one answer to all these questions: a human called it a system. he himself invented system. and each person (the researcher), has its own system. in fact, this system is some kind of a model for our understanding of a particular entity. the model itself, in turn, is also a system. and when the researchers begins to study the model, he/she has to disengage himself/herself from the “system” (more precisely, from the subject). no researcher argues with this. do we do with an object, which exists, not realizing that it a system? but the most paradoxical is that the person giving the definition of the system, tries to simplify the study of all the systems to unify their properties, suggesting that in the future it may be possible by studying one (the one that he understands) system to project (extrapolate) the system properties on the other system he understands less. if this does not work directly, it makes changes to the definimichael v. tokarev: thoughts on the general systems theory 230 tions, complicating them by adding new elements and features. such “fitting” of definitions for a research specific needs leads to the greater separation of new definitions from classical ones, and, as a consequence, to greater entropy in the system definitions themselves. discuss the entropy and the second law of thermodynamics, which, from our point of view, is applied indiscriminately to almost all known systems. interpretation of second law of thermodynamics is unambiguous: “entropy of an isolated system may increase or remain unchanged. reduction of entropy in an isolated system is impossible” [3]. what do many systems researchers do? they apply the second law of thermodynamics and the concept of entropy itself to any of the studied systems. in this case, few people pay attention to the paradox that, for example, self-organizing system does not increase the entropy inside themselves, but reduce, not desintegrate, but are being created, and organize themselves. this error has been long known to physicists and many other scientists. the concept of an isolated system in the definition of thermodynamics means no system interchange with the environment, not only by means of substance (closed system), but also by energy. the gross error is also the application of the concept of systems entropy to the systems which are not in equilibrium (initially entropy is a measure of the thermodynamic system in equilibrium). another mistake made by researchers who follow the fashion-application of the laws of thermodynamics to any system, including ones of not physical, and certainly not of a thermodynamic nature. for example, how does a perfect system of geometric axioms relates to thermodynamics and how we can apply the principles of entropy change to it? however, despite these considerations, all systems eventually die (entropy increases). why? is there a common cause of death of all the systems? and may the root cause of this common cause be the determination, which is man-made? meaning that the man gives the definition of the system in which it ends its existence (and this definition may include this course latently, possible only in the context of the investigator). most researchers believe that the system is, by definition, hierarchical, that the system itself may enter into other systems as an element (subsystem), and each of the elements of the system is the system on its own level of consideration. however, if at the intuitive level, this can be imagined, at the level of universal definition systems this property may seem controversial. why? if we assume that the system as a term was invented by people (not nature and not the creator), then considering a particular system in a large number of cases, the researcher does not appeal to the macro or micro systems, even not assuming that there is a hierarchy. as an example, study the alphabet. with this study, it is possible to ignore the fact that a particular alphabet is a subset of 231 advances in systems science and applications (2013) vol.13 no.3 all human alphabets. but at the micro level, in some context, we are absolutely not interested in each letter as a system (for example, lettering or placing them side by side in constructing words, etc.). one can argue that in this case a systematic approach does not apply. but in course of the alphabet model study we can look at it from the point of view of other system properties, in particular the main-integrity (agregation of letters has properties not possessed by each letter separately). perhaps we have no questions only to this particular property system-a property of integrity (the interaction of elements of the system leads to system properties / functions that each element separately unable to perform or “the object properties can’t be reduced to the sum of the properties of its constituent elements and non-deducibility of the last properties of the whole”. 3 system paradoxes the urgency of our questions is confirmed by systemic paradoxes in v.n. sadowski works [1]. and in particular, the first paradox of hierarchy, which is the following: for a complete description of the element as a “system element” there must be a full description of the system that can’t be described fully until each element is described. therefore, the question of the legality of a systematic approach analysis of micro and macro-level systems remains a question, the answer to which is possible in our view only in specific applications, with significant restrictions in the context, or the use of artificial systems with known properties, as-built for specific purposes. let’s consider a simple example. a man, as a biological system, consists of elements (subsystems): circulatory, digestive system, musculoskeletal, respiratory, etc. if we try to describe the circulatory system, as a separate, outside of the body, we can miss important features of nutrients carried by the blood, oxygen, and various chemical elements necessary for metabolism, etc. for a complete description of the circulatory system, as part of the body, we need to have a complete description of the body (including the blood system). thus, for the study of any element of a complex system, we have to build a model that is different from the object due to substantial simplification. considering the provision of oxygen, we can abstract away from the digestive system, which significantly affects feedback and the circulatory system, which in some cases may be affected so much that the work of the respiratory system will be blocked. even more interesting example is when a person studies a community of people, to which he /she belongs, and his/her decisions can affect the outcome of the michael v. tokarev: thoughts on the general systems theory 232 system. in this case, self-knowledge is not just difficult, it is impossible. on the other hand, if we consider a vehicle as the system, studying any of its units (subsystem) is simplified by the fact that creating a car we base its structure on system properties. the rest of the paradoxes is formulated similarly based on the fact that it is impossible to explore a part, not knowing to the whole and vice versa. here we do not just agree with v.n. sadowski, but get confirmation of these issues. namely: “the attempt to interpret paradoxes considered static, applied to the system knowledge, taken out of its development, inevitably leads to the conclusion that the system thinking is impossible.” 4 how to answer these questions in what direction should one look for the answers to these questions? the first way is, of course, a sharp decline in the number of systems, consideration of only artificial systems, which were originally created as a system, not disseminating research results to other systems. the second method, proposed by v.n. sadowski, is to explore the hierarchical systems by fixing some elements/connections/properties. as is done in the systems of equations, where the number of unknowns by more than 1 greater than the number of equations. solving such systems of equations, the researcher captures part of the unknown variables and then solves the solvable system of equations. unfortunately, this method, as well as the method of successive approximations, proposed v.n. sadovski, does not allow to solve systemic paradoxes efficiently, in real time, with the right qualities in the study of systems. references [1] sadowski v.n. (1974), foundations of general systems theory, moscow, nauka. [2] mesarovic m.d. and yasuhiko takahara. (1975), general system theory: mathematical foundations, system research center, cleveland, ohio. [3] osipov a.i. and uvarov a.v. (2004), “entropy and its role in science”, journal.issep.rssi.ru, vol.8, no.1. corresponding author author can be contacted at: mtokarev@axisconsulting.ru advances in systems science and applications (2012), vol.12, no.2 133-140 research on the 3-dimensional grab design system based on case-based reasoning yuantao sun1 and ran li2 1college of mechanical engineering, tongji university, shanghai, 200439,china 2changjiang three gorges navigation administratioiyichang,443133, china abstract with the increasing intense competition of mechanical manufacturing market, companies must shorten products design cycles, improve products’ quality, and reduce the costs of products. this paper takes into account the characteristics of grab product design and put forward the method the integrate three-dimensional design and case-based reasoning technology for the grab design, and propose the procedure of grab bucket design based on case-based reasoning and apply the method to grab bucket design field,. in the paper, the 3-d grab case library and explore intelligent reasoning module are established. and case-modified technology through solidworks and its secondary development are achieved. at last, a set of three-dimensional grab design system is developed, which can realize grab virtual assembly, kinematics analysis, interference check, finite element analysis and creating engineering drawings so that it makes the applications of intelligence design possibility. keywords digital manufacturing, case-based reasoning, virtual design, intelligent design, grab 1 introduction the grab is a main handling device which is used for loading and unloading staple scattered material. but all countries have not a uniform design rules for grab bucket because of the diversity of the material shape, the type of material or load capacity and working environment. for example in germany, various manufacturers have their own reference data and design rules. and in china, there is no similar national standard. the study on the grab design also is less[1]. in the actual design process, the experience occupies a large proportion. in general, designers use the method such as a simple analogy, size, and zoom in or out according to the reference drawings to revise and design new grab products based on the similar types of grab products drawings. but the method is very difficult to avoid interference between the parts of grab and guarantee the products have reasonable dredging rate and so on. with the increasing intense competition of mechanical manufacturing market, companies must shorten products design cycles, improve products’ quality, and reduce the costs of products. to improve the efficiency and quality of grab design, 134 yuantao sun:research on the 3-dimensional grab design system based on case-based· · · the paper focus on the method how to integrate both of intelligent design and three-dimensional grab modeling . in common, the intelligence design method is the process that choosing the design satisfied solution through the rbr (rule-base reasoning) way base on record the product design knowledge and discipline in rule form, through the rbr (rule-base reasoning) way to choose the design satisfied solution. intelligent design is method which record product design knowledge and law based on the usual rule form, and make the design plan which meet the requirements through the rule-based reasoning (rule-based reasoning, rbr). however in the design field, a wealth of experience and fragmented knowledge are difficult to be summarized in the form of rules. those problems limit such system application for the mechanical design field in a narrow range. these problems restrict the application of this kinds system in mechanical design field. recent years, with the intelligence technology studying, case-based design (case-based design, cbd) method has shown haven a good ability to resolve these issues. along with the studying of the intelligence technology, cbd has show its ability on these problem solving. case-based reasoning (case-base reasoning, cbr) is a form of artificial intelligence from machine imitation to machine thinking[2-3]. the characteristics of grab product design are took account in the paper. three-dimensional design and case-based reasoning technology are integrated in the grab bucket design field. the procedure of grab bucket design based on case-based reasoning is put forward in the paper. the design system establishes the 3-d grab case library and explore intelligent reasoning module. and achieve case-modified technology through solidworks and its secondary development. at last, a set of three-dimensional grab buckets design system is developed, that it makes grab intelligence design come true. 2 the principal of grab design based on case reasoning because the foundation of cbr is that similar problems have similar solutions, the case-based design method is very close to the actual design process[4]. faced to new design requirements, the system first select the case which is most close to design requirements from the past design cases library ,and then simulate to get design program which meet current requirements. the results case also can be modified again as a reference case that means the system has a self-learning ability. the advantages are: (1) no need to build the rule or model, but only collect the past cases to establish the cbr system; (2) only need determine the related case characters. cbr system’s case can be improved and enlarged during the using, as long as some examples are added in the system; (3) no need to reasoning from the beginning, only through a completed program which can generate the solution quick; (4) easy to maintain. the increasing new cases advances in systems science and applications (2012), vol.12, no.2 135 not only achieve the purpose of the study, but also reflect the customer demand character. 2.1 design workflow of grab design based on cbr design workflow is shown as fig.1. designer input the parameter of grab first. then the system search the cases library through the key words and find out a number of similar cases and identify and select the most similar one as design reference .the user can get the guidance information to modify the case and establish the three-dimensional geometric model for the product through human-computer interface after a series of steps such as virtual assembly model, kinematic analysis, interference checking, structural analysis, mechanical steps finally, to generate two-dimensional drawings, at the same time, new product design is completed as a new three-dimensional cases which is stored in the cases library. fig.1 design flow of grab based on cbr so that the grab intelligence design depend on the following key technologies: case expression retrieval cases modify case modification. 136 yuantao sun:research on the 3-dimensional grab design system based on case-based· · · 2.2 case expression because the number of cases and the degree of perfection of case base is the cornerstone of using cbr technology for compute-aided design .the establishment of case library should take the content of the case, the expression of the case and the organization of the case base etc. into consideration. it is considered that the software solidworks has the functions of three-dimensional solid modeling, so that the software and its secondary development are used to the established geometric model can display three-dimensional appearance of entities in actual time[5]. through the operation of the model the entitle shapes modification of the model can be concisely completed and the same times. the physical changes slice product model, interference detects, motion simulate, finite elements analyze and so on can be concisely completed. in view of the characteristics of grab, in order to facilitate search and matching cases, the paper put forward the method to define cases that the cases =cases three-dimensional graphic + case feature the data. feature data refers to the factors that has a decisive influence on the product design, and can be extracted from the cases. the following is workflow. at first, the three-dimensional graphic of the cases are set up in solidworks, then feature data of the cases are input into the cases library. in practical applications, regarding to different storage requirements of case’s three-dimensional graphic library and feature data, different storage management functions are established in system respectively so that they can be integrate into the same interface to be operated. so that a cases consist of a set of three-dimensional graphics include the related sub-components hierarchy and a number of feature date of the case. the method makes the case query and retrieval more easier. fig.2 shows the three-dimensional graphics, the structure and the expression of feature data in the design system respectively. fig.2 case expression of grab in order to make the model more practical, both hands of modeling methods and modeling sequence should be take into; on the one hand, it should be easily modified ,that means the main size of upon -bearing beam , under-bearing beam, advances in systems science and applications (2012), vol.12, no.2 137 stay bar, bucket body can be stretched and reduced in the corresponding parts through changing the numerical size, and they are easily be matched each other also ); on the other hand, it is easy to establish model, and the model’s level and structure coincide with the actual as far as possible. the system use the software solidworks, through top-down design methods to establish the model of grab’s every part. case features are summarized from the case. the description of the case feature is the core of case representation. in the system, the case features are classified three types to be expressed according to different case objects. 1. every complete set of grab product is a case object. the type of grab as one of case feature will be placed as the first important. in general, there are several types of grab: double-cables two-jaw grab, four-cable double-jaw grab, four-cable grab and four cables multiple jaws and so on. there are also several special type grab such as hydraulic grab .and then features such as lifting capacity, material characteristics and work occasions will be took into account as the design requirements. 2 .the main part of grab such as upon -bearing beam, under-bearing beam, stay bar, bucket body also are deal as cases .the feature are expressed according to the range of lifting capacity, shaping feature, mating dimension. 3. the case feature of grab accessory such as pin is geometrical dimensions and processing technology. these features are input into product lib as database. then the features will be modified by the product data management module and matched with the corresponding three-dimensional drawings .so that the corresponding three-dimensional drawings will be easily got through the feature database retrieval. 2.3 case retrieval the key to case-based reasoning achieved successfully is the case retrieval and matching .the function of the case retrieval and matching is that the most similar cases will be retrieved and used as the template of new design program. to cases which be retrieved, there are two requirements: 1) the cases should be less as possible. 2) the cases are most similar with new product. so that the retrieval algorithm is the core of cases retrieval. at present there are three kind of retrieval algorithm: nearest neighbor method, induction indexing method, knowledge inducting method. because the feature of grab can be classed according to grab parameter such as dimension, type of grab. the system uses a knowledge-guided strategy and the recent adjacent strategy. the nearest neighbor method and the knowledge inducting method are used in the system .those grab parameter are define as feature key word and are given weight value .after calculating case similarity according to the weight, the most similarity case is selected and modified. 138 yuantao sun:research on the 3-dimensional grab design system based on case-based· · · the following equation is similarity formula sim (l, k) = 1− √√√√ n∑ i=1 wi × ( fki − fli fli )2 where sim means the similarity between objective design and the case, the more value of sim is large ,the more similarity is close. fli means design requirement feature item l includes attribute variable i. fli means the case feature item k includes attribute variable i . if fki and fli is same, then fki − fli value is 0, otherwise is 1).wi is weight of the attribute variable i. for example, when the case is searched , considering the design requirement and the using condition for the grab ,its feature array is (the grab type, lifting capacity ,material ,bucket cubage),the weight array is (0.4,0.2,0.3,0.1). 2.4 the case modification base the knowledge it is a key problem to propose the advice about the modification through case reasoning. after the system concluded the experts’ design experience on grab, it establishes knowledge base. the knowledge base has concluded the main parts’ design experience and data on different types of grabs, modify suggest on overall geometric parameters and the related reasoning calculate formula. if there is a difference between the case base is the grab capacity when a new grab is designed., you can propose the suggestion that modify the grab’s width in a certain range firstly according to the ratio between the two grab capacity in the case base and then the modified modules are called to the grab three-dimensional model can be modified. because the application of the interface with solidworks and the top down design theory, the function can be applied that the grab three-dimensional model could automatically update as the parameter changes. topdown design means the ability to design the relevant sub-components in the assembly environment, not only the relevance between the size parameters, but also realize the automatic and total relevance between the geometry appearances and parts. the user can design some other parts on the case that the assembly layout diagram is already, and make sure that the assembly layout diagram and the parts’ size are totally automatic relevance, so to make the system’s modified modules in three-dimensional model more convenient. as long as the key sizes are modified on the matching position, the relevant parts can be automatically modified and updated. 3 application to design a double -cable double-jaw grab, whose load capacity is 12t, grab material is sand, bucket capacity is 1.6t/ m3, after retrieval the similar grab advances in systems science and applications (2012), vol.12, no.2 139 bucket case from the three dimensional case library. in the case library, one which loads capacity is 10t, grab material is sand, bucket capacity is 6m3 can be got. considering the difference with the objective product, keeping the same bucket area, the bucket wide is modified from 2.05m to 2.46m, and then it can satisfy the capacity requirement. the new capacity is 7.2m3. after the accordingly adjustment for other components, the 12t three dimensional model and drawing can be got shown as fig.3 and fig.4 respectively. fig.3 grab model fig.4 grab part drawing 4 conclusion the intelligent t design process of the grab based on the cbr and three-dimensional entity’s model technology is introduced in this paper. the basic structure of implementation, the expression s of cases and their retrieval method are also described. an intelligent t design system is developed using solidworks and its secondary development. the application of the system has shortened the design cycle and improved the design efficiency. at present improving the knowledge 140 yuantao sun:research on the 3-dimensional grab design system based on case-based· · · library is still in progress. acknowledgements the paper is sponsored by “the open research projects supported by the project fund of the hubei province key laboratory of mechanical transmission and manufacturing engineering wuhan university of science and technology”. the sponsor project series number is 2007a22. references [1] li ran. (2005), research and development of 3-dimensional grab bucket design system based on cbr, degree description of wuhan university of technology, pp.2-5. [2] dubois.d, hullermeier, e and prade.h. (2006), “fuzzy methods for casebased recommendation and decision support”, journal of intelligent information systems, vol.27, no.2, pp.95-115. [3] hamza.h, belaid.yand, belaid.a. (2007), “case-based reasoning research and development”, proceedings 7th international conference on case-based reasoning, pp.404-418. [4] lee sangjae, kim kyoung-jae. (2009), “using case-based reasoning for the design of controls for internet-based information systems”, expert systems with applications, vol.36, no.3, pp.5582-5591. [5] ding yufeng, wei zhongling. (2006), “research on parametric process planning technology based on three-dimensional part model”, journal of wuhan university of technology, vol.28, no.s1, pp.502-506. microsoft word 4-李国东.doc advances in systems science and applications (2010), vol.10, no.1 19-24 issn 1078-6236 international institute for general systems studies, inc. analysis b-scan image by cnn and polyfit* guodong li1,2 and wenxia xu3,4 1school of mathematics and physics, north china electric power university, beijing 102206, china 2industrial systems engineering, university of regina, wascana parkway, regina, sask. s4s 0a2, canada 3dep of electric engineering, chengdu university of information technology, chengdu 610225, china 4xinjiang weather modification office, urumqi 830002, china email: lgdzhy@ncepu.edu.cn, xwxqiuye@live.cn abstract in this paper, we have be processed ultrasound with cell neural network(cnn). first we detected the edge from b-scan images, and then we analysis the data that from the edge of image. some interesting result has been finding. there are some regular between the number and the patient’s b-scan image. keywords b-scan image edge detection liver damages 10-degree polynomials fitting 1. introduction in the last two decades, medical image processing technology has been developed rapidly. it can help doctor to diagnose patients' disease more accurately. medical image processing with computer has attracted much attention recently [1]. a lot of methods for dealing with medical images have appeared [2],[3]. in chua's articles (see [4] or [5]), many important and interesting cnns are described. one of them is the edge detection cnn, which can detect the edge in gray-scale images. in a recent conference paper, a robustness theorem for designing edge detection cnns is set up [6][7][8]. detecting edges of gray-scale images may be required for a variety of purposes, such as machine vision, image analysis, and image pick-up and so on. edge detection is one of the most important steps for image recognition since there is a direct relationship between edge and object recognition. a lot of scene information can be interpreted from the edges. in natural images, most edges are associated with abrupt changes in intensity distribution and can be approximately modelled as step edges. pathological changes of chronic hepatitis patient include liver fibrosis, heterotrichosis and liver cancer, etc; however, the three states of liver disease are difficult to distinguish from b-scan image by naked eyes. digital analysis of the b-scan images will help doctors to diagnose the stage of the damages of patients' livers. 2. edge detection cnn in a chua’s exposition [1], many important and interesting cnns are described. one of them is edgegray cnn, which can detect the edges in gray-scale images. the template of the standard edge gray cnn has the form。 the standard m×n cnn architecture is composed of cells jic , . the dynamics of each cell is given via the following equation [1]: jiljki r rk r rk lkljki r rk r rk lkjiji zubyaxx ,,,,,,, +++−= ++−= −=++−= −= ∑ ∑∑ ∑ (1) * this work is supported by meteorology bureau of xinjiang uighur autonomy, science and technical item (project no. 201012). li: analysis b-scan image by cnn and polyfit 20 ( )11 2 1 ,,, −−+= ++++++ ljkiljkiljki xxy njmi nmljki ,2,1;,2,1 ],1[],1[),( == ×∈++ (2) where jijijiji zuyx ,,,, ,,, represent state, output, input, and threshold respectively; lklk ba ., , are the elements of the a-template and the b-template respectively. figure 1 the dynamic routes of the cnn the standard edge cnn's local rules and template are listed as follows: local ruler jiu , → )(, ∞jiy (1) white → white, independent of neighbours. (2) black → white, if all nearest neighbours are black. (3) black →black, if at least one nearest neighbours is white. (4) grey →black, if the laplacian operator zu ji >∇ , 2 . (5) grey →white, if the laplacian operator zu ji <∇ , 2 . (6) grey →0, if the laplacian operator 0, 2 =∇ jiu . the standard cnn template has been generalized by the following theorem. theorem 1[6]: let the positions of cnn template parameters be described by (3).then the cnn can perform the local ruler of the detection of edges in gray-scale images, if the following parameter inequalities hold: cbzzcz 6 .2 8 .1 −<<− 000 00 000 1,1 zz ccc cbc ccc baa −= ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ −−− −− −−− = ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ = (3) where a > 1, 1,1 >> cb ∑ ∑ ≠ −∈ ++−δ∇ )0,0(),( }1,0,1{),( ,,, 2 lk lk ljkijiji ucbuu where: advances in systems science and applications (2010), vol.10, no.1 21 ⎪⎩ ⎪ ⎨ ⎧ − >δ =δ otherwise u if c ub ji jilk 1 g )( , ,, (4) 3. segmentation of the b scan image figures 3 show the ultrasound b-scan images (ubsis) of six people's livers. each original ubsi is an rgb image with 720×576 pixels. we chose the same eath part in the ultrasound b-scan images as that shown in fig.2, which is on the left of the cholecyst and up of the portal vein. expert doctors analyze this region when they diagnose sufferer according to b-scan image. so we cut this region and then processed it with edge gray detected cnn. clinical diagnoses for the six ubsis are listed as follows. figure 3(1) is obtained from a healthy hepatitis b virus (hbv) carrier. figure 3(2) is the ubsi of a normal liver without hbv infection. the ubsi of a young (under 20 years old) chronic hepatitis b patient without liver cirrhosis is shown in figure 3 (3) and that of an old one (over 50 years old) is given in figure 3 (4). figure 3 (5) and figure 3 (6) are the ubsis of a patient with ascites and a patient with marked liver cirrhosis, respectively. 4. application of the edge detection cnn in this section, we shall use edge detection cnn to process the part of b-scan image. the template parameters of the cnns are given in table 1. table 1 template parameters of the cd cnn a b c z g 4 16 2 0.4 0.4 the edge detection cnn can be used to process rgb images. an rgb image is usually represented by an m×n×3 (for the above ubsis, m = 576, n = 720) data matrix p where m, n stand for the rows and the columns of the pixels in the image, and 3 represents 3 color planes ---red, green and blue, each color plane with 256 levels of intensity denoted as (r, g, b) = ( ) ( ) ( )( ),3:,:,~,2:,:,~,1:,:,~ ppp . such an image is called an rgb 24-bit map. in order to use cnn to process rgb images, we use a transform *pp → (5) to change the 256 levels of intensity of each color plane into the levels of intensity in [-1, 1]. consequently, the lighten pixels in the original rgb image p correspond to smaller values in the transformed image *p and vice versa. in the following discussions, we always assume that the levels of intensity of input rgb images 1p and 2p have been transformed via formula (5); the color planes of processed images have been changed into the rgb forms when the processed images are shown as color pictures. firstly the edge detection cnn given in table 1 is used to process the part of the ubsis shown in figure 3. the processing results are demonstrated in figure 4. it is difficult to give a correct judgement for non medical researchers. however, we arrange the pixels of the three color planes p(:,:,1), p(:,:,2), p(:,:,3), of each processed image (the output ( )"" , ∞jix , not the output ( )"" , ∞jiy ) in row-wise packing scheme [5]. then wavelet transform is applied to process the images in figure 4. li: analysis b-scan image by cnn and polyfit 22 figure 2 the selected part of the b scan image the part in the white rectangle is chosen for edge cnn analysis. patient liver no1(a) patient liver no2(a) patient liver no3(a) patient liver no4(a) patient liver no5(a) patient liver no6(a) figure 3 segmentation of patient liver patient liver no1(b) patient liver no2(b) patient liver no3(b) patient liver no4(b) patient liver no5(b) patient liver no6(b) figure 4 the result of the b scan image processed by edge detect cnn. 5. introduction of polynomial fitting given data ( )ii yx , (i=0,1,…,m),we shall find a polynomial p(x) such that the error’s square sum is the least, i.e. ( )[ ] min 0 2 =−∑ = m i ii yxp suppose that φ is the set of polynomials of degree ( )mnn ≤ . let ( ) ∑ = φ∈= n k k kn xaxp 0 and advances in systems science and applications (2010), vol.10, no.1 23 ( )[ ] min 0 2 00 2 =⎟ ⎠ ⎞ ⎜ ⎝ ⎛ −=−= ∑ ∑∑ = == m i n k i k ik m i ii yxayxpi (6) it is obvious that ∑ ∑ = = ⎟ ⎠ ⎞ ⎜ ⎝ ⎛ −= m i n k i k ik yxai 0 2 0 is a function of ,,,, 10 naaa . so the problem is to find the extreme value of ( )naaaii ,,, 10= . by the necessary condition of extreme points, we have: 02 0 0 =⎟ ⎠ ⎞ ⎜ ⎝ ⎛ −= ∂ ∂ ∑ ∑ = = m i j i n k i k ik j xyxa a i nj ,1,0= (7) i.e. njyxax m i i j i n k k m i kj i ,1,0 00 0 ==⎟ ⎠ ⎞ ⎜ ⎝ ⎛ ∑∑ ∑ == = + (8) in matrix form, (8) becomes ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ = ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ + ∑ ∑ ∑ ∑∑∑ ∑∑∑ ∑∑ = = = == + = = + == == m i i n i m i ii m i i nm i n i m i n i m i n i m i n i m i i m i i m i n i m i i yx yx y a a a xxx xxx xxm 0 0 0 1 0 0 2 0 1 0 0 1 0 2 0 00 1 (9) it is easy to show that the coefficient matrix of (9) is positive definite, therefore (9) has a unique solution. solving (9), we obtain ka (k=0, 1,…,n),and thus the polynomial ( ) ∑ = = n k k kn xaxp 0   (10) it is easy to show that (10) satisfies, i.e. (10) is the fitting polynomial. figure 3 and figure 4 are 6 livers’ original b-scan images and resulting images processed by cnn, respectively. each original b-scan image is an rgb image with 720×576 pixels. in this paper we deal with these b-scan images by edcnn. then by polynomial fitting (10-degree), we find that the coefficients of the polynomials are related to the damage of the patients’ livers. 6. concluding remarks we have detected 6 b-scan images by cnn. then using polynomial fitting, we obtain the following data as shown in table 2. data of normal liver is no. 1 in table 2, from 2 to 5 are b-hepatitis in table 2. from the tables 2, it is easy to see that the value of a1 is closely related to the condition of the liver. most values of a1 are smaller than 7, and few are smaller than 12. therefore, according statistic regular[9], we can conclude that if a1 is larger than 7, the liver is normal; and if a1 is smaller than 7, the liver is chb. we can use this way to diagnose the chb. li: analysis b-scan image by cnn and polyfit 24 table 2 data of edge b-scan analysis by cnn a10 e-47 a9 e-41 a8 e-35 a7 e-29 a6 e-24 a5 e-19 a4 e-13 a3 e-09 a2 e-04 a1 e-00 a0 e+5 1 -2.1535 4.9041 -4.7809 2.6041 -8.6754 18.181 -2.3715 18.313 -7.4913 12.252 1.8006 2 1.1509 -2.2140 1.7396 -0.7036 1.4633 -1.0387 -0.1564 3.6640 -2.5314 5.3752 2.0242 3 -0.3927 0.9139 -0.9437 0.5649 -2.1400 5.2406 -0.8123 7.4907 -3.6249 6.5730 1.9970 4 -0.8249 1.7516 -1.6135 0.8469 -2.7937 5.9910 -0.8262 6.8948 -3.0492 5.1082 2.0678 5 1.3496 -2.4470 1.7724 -0.6260 0.9187 0.6274 -0.4292 6.0397 -3.5273 6.9943 1.9646 6 0.5490 -0.8967 0.5185 -0.7841 -0.4708 2.6723 -0.5914 6.5870 -3.5138 6.6962 1.9830 references [1] b. li. evole and application of medicine image in technology of computer image processed. journal of medicine engineer, 1993, 10(4): 360-363. [2] s. r. xia. medicine image searchers technology base on charpter fuse and feedback correlative. spaceflight medicine and medicine engineer, 2004, vol. 12, 429-433. [3] z. l. tian. two dimension wave transform and application in processed of medicine image. journal of medicine treatment appliance in china, 2004, 28(6): 409-411. [4] l.o.chua. cnn: a version of complexity. int. j. bifurcation and chaos, 1997, 7(10): 2219-2425. [5] l.o.chua, t.roska. cellular neural networks and visual computing. cambridge: cambridge university press, 2000, 89-94 [6] g.li, l.min, h.zang. design for robustness edgegrey detection cnn, international conference on communication, circuits and systems and chaos, 2004, (2): 1161-1165. [7] zhao, x., liao, x., jia, x "the neural networks of regression analysis and its application", advances in systems science and applications, 2006, 6(1): 7-15. [8] jian, j.g., zhang, f.k., chen, d.y. analysis of absolute stability for a class of neural networks. advance in systems science and application, 2005, 5(1): 59-64. [9] fanling kong, zhiguo zhang, guizhi wang. a martingale method of optimization design in investment and effectiveness model. advances in systems science and applications, 2007, 7(1): 1-6. advances in systems science and applications (2014) vol.14 no.3 278-285 the new form of mixed economy with rationing: agent based approach valery l. makarov and albert r. bakhtizin central economics and mathematics institute of russian academy of sciences, moscow abstract in nowdays, there is a number of mixed economy types in different country. however, any one type has its advantages and insufficient. in this paper, a new form of mixed economy with rationing based on agent is proposed, and the results of computational simulation are investigated. keywords mixed economy types, agent based approach 1 introduction. there is sizable number of mixed economy’ types. generally speaking every existing economy is mixed one in a sense that it uses various economic mechanisms. every country tries to find optimal combination of the mechanisms. the most famous mechanisms are market, rationing, planning, direct distribution, gifts, some others. in the paper we explore the rationing as a way to implement a variety of incentives for people’ activity. we assume that the population is divided to the six social clusters with different incentives to work. the first social cluster (business oriented people) has standard incentive: getting profit. other clusters’ members (state and military servants, scientists, art workers, priests, etc.) try to lift on the social stairs, getting higher rank. having a position of a certain rank a person obtains the possibility to consume goods according to fixed norms. so, we have mixture of the two mechanisms: market and rationing. mathematical economists start to study this kind of models long ago. see, tobin james [1]. the literature was growing since that. we indicate a few. in howard david h. one can find some results about dependence of a market equilibrium on quantity constraints for a consumer [2]. in makarov v. l, vasil’ev v. a. there are number of proofs of the equilibrium’ existence for various mixed economies [3]. it is important to mentioned that in the all papers quantities norms (standards. rates, quotas) are given. the only task related to changing the norms is comparative static’s problem. for example, what happened to equilibrium prices, when some norm is increasing. no studies of mechanisms, seeking the norms. in a reality the norms are generating by a society. much depends on a type of the society. under totalitarian regime the norms are given by a dictator. other regimes produce a number of way how to generate the norms, starting from a decision of relatively small groups (elite) to taking into account the opinion of everybody (referendum). advances in systems science and applications (2014) vol.14 no.3 279 a number of papers attempt to reveal the mechanism for formation of social standards [4]. it is clear, that mathematical modeling is not an effective method to study the problem. the agent based approach looks promising. there are number of papers where the agent based models are successfully used for generating of so called social norms. see, for example, epstein joshua m. [5]. epstein proposes a simple agent-based model where agents learn behavioral rules (i.e., accept certain standards formed in the society). agents are placed in order in a certain loop. agents interact with each other. each agent has a fixed position in the loop and is endowed with two characteristics. the first is “quota”, a binary variable which depends on the values of the quotas of the neighboring agents. the neighbors taken into consideration are situated within the radius of vision, which is the second characteristic of the agent. the simulations demonstrated that after certain rules have been established in the society, the majority of agents subconsciously keep following them. fent, thomas described an agent-based model developed for analyzing emergence, stability, and change of social standards within a certain population of artificial agents [6]. the information on social standards is kept within a social network, which unites the agents of the model. each agent has an inner group (i.e., a set of agents with the same social standards) and an outer group. accordingly, the agents of the model receive utility from following social rules of their inner group and from refusing the rules of outer group. the model was applied to explaining the emergence of temporal social rules, prevalent in certain groups. the quantities’ norms are the special type of social ones. the social norms play role of constraints in behavior. for example, taboo is a typical social norm. the quantities’ norms look as prices. one can buy or not by the price, one can use the quantities’ norm or not. prices can be set by market (people) or by a state agency. so with norms. its can be generated by a society (people) or by an agency. essential difference between prices and norms lies on the speed of changes. in the described below model we consider that the norms change once in a time unit period. 2 mixed economy with rationing based on agent we can mentioned the recent paper with the some rules of generating the quantities’ norms (tips). savarimuthu et al. describe an agent-based model with a micro-level mechanism for the emergence of social standards [7]. the approach differs from the common macro-level mechanism, which is implemented in most papers on social standards. agents have memory of events and radius of vision, i.e. a parameter which determines the square of the surface where the agent can observe events, related to violating or following social standards. the pro280 valery l. makarov:the new form of mixed economy with rationing: agent based ... posed algorithm for the emergence of social standards is applied to a visit to a restaurant where agents make decisions about tips. computational simulations revealed that the time for emergence of social standards directly depends on the length of memory (the number of recent events memorized by the agent) and on the value of the radius of vision. we constructed the agent based model, where agents are people of the six social clusters mentioned above. moreover people in the clusters are divided to three ranks; high, middle and low ones with exclusion of the first social cluster. the first cluster (business oriented people) has objective to achieve maximal profit. the other people destinations are to get high rank. the quantities’ norms associated with the clusters and the ranks in. an agent consume goods from two sources: from market and rationing spheres. in a simple case we consider that an agent used only one source: market, if he/she belongs to the first cluster, or rationing, otherwise. now we describe the first version of the agent based model, where each social cluster produces one “good”, and the norms relate to the first good only. other “goods” we interpret as development’s levels of the clusters. so the production function of the first cluster looks as a1 (t) = n1(t) α1(t) · k1(t− 1)β 1(t) · a2(t)γ2 · a3(t)γ3 · a4(t)γ4 · a5(t)γ5 · a6(t)γ6 (1) where a1 (t) is a production of the first cluster in the period t; aj (t) is the level of the cluster’ j development in the period t; n1 (t) is the quantity of agents, populated cluster 1; k1 (t− 1) is accumulated capital (production funds) to the beginning of the period t. other clusters’ production functions are about the same type. so there is a strong dependence of all clusters on each other. it is impossible to develop one cluster with no development of others. the consumption of an agent ci is defined by his/her budget, if he belongs to the first cluster and is equal to the existing at the given period norm of a rank & cluster for others. so a macro-path {a (t) , c (t)} of the agent-based model, which is the outcome of simulations, depends on the norms, generated by the mechanism, we try to design. the comparison of the simulated trajectories can be done by standard way, using well know criteria. the simple way to generate the norms we used is the following. the initial norms for all clusters and ranks are given. further agents make influence on the advances in systems science and applications (2014) vol.14 no.3 281 norms according to the formula. csr(t) = { csr(t− 1) · ex(n) · fx(n)(r) · ox(n)(t− 1); s = x csr(t− 1)/(ex(n) · fx(n)(r) · ox(n)(t− 1)); s ̸= x (2) where ex(n) the level of n-th agent’s influence on the total level of consumption f social cluster x is calculated as: ex(n) = 1 + µ · (1/n) (3) with n the number of agents in the society and µ a coefficient, which determines the power of the agent. by default µ equals 10 for agents of the first (high) rank, 3 for the second and 1 for the third rank. this implies that the influence is an order higher for the agents of the first (high) rank if compared to the third rank. it is important to mentioned that the state of agents is changing over time. the age of the agent increases and the life expectancy decreases in each moment of time. with a certain probability, estimated on the basis of statistics, the agent can have a child. upon reaching the age 18, the child with a high probability enters the cluster of parents. agents can migrate from one cluster to another according to some rules. we produced a number of scenarios, how agents influence to the norms, divided to the two blocks. the first block deals with assessing the change in the socio-cluster society due to the change in the mechanism for the formation of the norms. the second estimates the impact of various factors on the processes for the formation of the norms. the first block contains three scenarios to make calculations. the first one is the baseline scenario for the development of the socio-economic system in our model. the second calculation deals with a slight modification in the work of the model. while the first calculation studied the totality of agents, in which each agent was capable to directly influence the consumption quotas for all social clusters, in the second calculation agents elect representatives of their clusters proportional to the number of agents in the cluster. in other words, we supplement the model with a special (new or additional) cluster. the only function of the agents in this cluster is to regulate social quotas. the agents of the supplementary cluster change quotas according to the formula analogous to presented above. the only difference is that the agents of the supplementary cluster have higher level of influence in changing quotas. moreover, the very change of the quota in each cluster happens stochastically: the quota is changed with probability 1 for the cluster represented by the agent. for other social clusters the quotas are changed with probability 0.5. this implies that the government elected by the agents fully 282 valery l. makarov:the new form of mixed economy with rationing: agent based ... fig.1 the results of computational simulations (x coordinates is years, y coordinates is the volume of production) supports its electorate but does not hurt others strongly. the third scenario is similar to the second. the difference is higher values of probability for the decrease of the quotas for the social clusters, which agents are not represented in the regulating body. fig.1 demonstrates the results of calculations with respect to the estimated values of production in all the three scenarios. we see a justification of the assumption about the need for equal rights of all clusters for successful development of the society. when the regulating body strengthens the policy aimed at discriminating certain social clusters, it is immediately reflected on the volume of production in the whole economic system. note that a more democratic principle of change of social quotas (when each agent has a direct influence on parameters) is slightly more efficient. the second block of calculations quantitatively assesses the impact of various factors on the processes for formation of social quotas. namely, it distinguishes between two mechanisms of formation: by each agent or by a selected group. first, we studied the dependency of the values for social quotas on the initial number of agents in each social cluster. in this case the size of the first social cluster was 50% of the society, and the sizes of other clusters were 10%. in the second case the sizes of all the 6 clusters were equal (16(.6)%). note that the initial level of social quotas was the same in both cases. the results for the 8th year of the model are presented in table 1 (all values are normalized relative to the initial values, i.e., all quotas equal to 1 in the zero year). the results demonstrate that in the first case (when the first social cluster initially includes 50% of the society), quotas of the first social cluster increased advances in systems science and applications (2014) vol.14 no.3 283 table 1 the impact of the initial number of agents in clusters social cluster experiment-1.1 (each agent can change quota) experiment-1.2 (quotas are changed only by representatives of a special cluster) experiment-2.1 (each agent can change quota) experiment-2.2 (quotas are changed only by representatives of a special cluster) first 1.053 1.042 1.001 0.996 second 0.962 0.944 1.011 1.002 third 0.930 0.981 0.983 1.010 fourth 0.954 0.956 1.035 0.987 fifth 0.936 0.948 0.996 1.003 sixth 0.933 0.967 1.022 0.989 standard deviation (for all clusters) 0.03060 0.02568 0.01479 0.00713 table 2 the impact of the number of ranks social cluster ranks are present ranks are absent first 0.981 0.995 second 0.943 0.981 third 1.034 0.975 fourth 0.967 1.036 fifth 1.045 0.983 sixth 0.954 1.055 standard deviation (for all clusters) 0.03478 0.02759 more than quotas of other clusters. note that the standard deviation of the values for all clusters is larger than in case of change only by the representatives of a special cluster. we conducted experiments no.2.1 and 2.2 several times. each time we obtained a certain (different) set of values for social quotas. consequently, even making a society more balanced (i.e., with the equal number of agents in social clusters), in the end we inevitably have a differentiation in consumption. we assessed the dependency of the values of social quotas on the presence/absence of ranks in the society. (with the absence of ranks the agents have equal abilities to change social quotas). similarly to previous estimations we conducted calculations for 8 model years (see results in table 2; all values are normalized relative to the initial values, i.e., all quotas equal to 1 in the zero year). we see that quotas change faster in the presence of ranks. in other words the society with the differentiation s more inclined to an additional stratification in income. 284 valery l. makarov:the new form of mixed economy with rationing: agent based ... the final experiment deals with estimating the dependency of social quotas on their initial values. in initializing the model we increased the quotas for the first cluster, decreased for the sixth cluster, and left unchanged for other clusters. our purpose was to study the impact of initial misbalance in social quotas on the values of social quotas in the long run (15 years). the results (table 3) demonstrate a clear long run tendency: the discrepancy in social quotas increases. this implies that no movement towards balance can happen without an exogenous disturbance. table 3 the impact initial values of social quotas social initial values of each agent only representatives cluster social quotas changes quotas of a special cluster change quotas first 1.500 1.544 1.525 second 1.000 0.820 0.833 third 1.000 0.811 0.824 fourth 1.000 0.810 0.829 fifth 1.000 0.809 0.826 sixth 0.500 0.397 0.406 3 conclusion to sum up, the results of computational simulations revealed that: (1) successful development of a society with social clusters requires equal rights of social clusters; (2) democratic principle of change in quotas (when each individual has a direct influence on the values of quotas) is a more efficient economic mechanism; (3) social quotas change faster for the dominant social cluster than for other clusters; (4) even trying to make a society more balanced (i.e, with equal number of individuals in social clusters) at the end we inevitably face differentiation in production and consumption; (5) a society with differentiation is more inclined to additional stratification in the level of consumption. references [1] tobin james. (1952), “a survey of the theory of rationing”, econometrica, vol.20 (1952), pp.521-553. [2] howard david h. (1977), “rationing, quantity constraints, and consumption theory”, econometrica, vol.45, pp.399-412. advances in systems science and applications (2014) vol.14 no.3 285 [3] makarov v. l, vasil’ev v. a. (1989), “equilibrium, rationing an stability”, matecon, vol.25 no.4, pp.4-95. [4] neumann martin. (2010), “norm internalisation in human and artificial intelligence”, journal of artificial societies and social simulation, vol.13 no.1, 2010. publishers. [5] epstein joshua m. (2000), “learning to be thoughtless: social norms and individual computation, the brookings institution and santa fe institute, center on social and economic dynamics”, working paper, no.6, january 2000. [6] fent, thomas. (2006), “collective social dynamics and social norms”, mpra paper, no.2841. [7] bastin tony roy savarimuthu, stephen cranefield, maryam a. purvis and martin k. purvis. (2010), “obligation norm identification in agent societies”, journal of artificial societies and social simulation, vol.13 no.4. this work is supported by russian humanitarian scientific fund (grants no.1402-00431 and no.15-02-00276). corresponding author albert r. bakhtizin can be contacted at: albert.bakhtizin@gmail.com advances in systems science and applications (2013) vol. 13 no. 3 218-226 integrated mechanisms of organizational behavior control v.n. burkov, m.v. goubko, n.a. korgin and d.a. novikov institute of control sciences, russian academy of sciences, moscow, russia abstract problems of control mechanisms integration are formulated and discussed in the framework of mechanism design for organizational behavior control. unified schemes for control mechanisms description and design are proposed. an example of the integrated production cycle optimization mechanism is considered. keywords mechanism design, management theory, organizational behavior, integrated control mechanisms. 1 introduction according to general control methodology [1], control is the activity of subjects controlling other subjects or objects. formal models of control in organizational systems are studied in the framework of mechanism design (see surveys in [2] and [3]), theory of contracts [4, 5] and collective choice theory [6, 7], involving results of game theory [10, 11] and operations research (see textbooks [8, 9]), and are applied in microeconomics and management theory. a set of control mechanisms may be considered as a “kit”, which contains elementary blocks, intended for solving of typical problems of planning, organizing, motivating and controlling. but in practice one faces not typical but real complex problems of organizational control (according to [12] controlled complex systems nowadays are usually decentralized, hierarchical and networked, hence, heterogeneity approach in control models would surely be in demand). hence the technique of complexing and integrating different control mechanisms is required. below an attempt of integrated control mechanisms construction is taken for the problem of production cycle optimization. 2 control mechanisms suppose that a man is engaged in a control loop. then there is a need to control him or her – a controlled object which • is subjectified and acts according to individual interests and preferences (the principle of rationality); • appears not completely known to a corresponding control subject (the principle of information asymmetry); • may not reveal true information and not perform the expected actions. in the sequel, such controlled objects are said to be economic agents or, shortly, agents; the corresponding control subject is referred to as a principal. 219 advances in systems science and applications (2013) vol. 13 no. 3 the “principal – agent” terminology seems convenient and is traditionally used to describe control mechanisms. a control mechanism represents a certain set of rules and procedures involved by the principal for making decisions that influence on the behavior of active economic agents (in particular, on information revealed and actions chosen by them), hence the principal has to solve the problems of control mechanism design. in fig. 1 the typical aggregated scheme of a control mechanism structure is provided in a operational form. fig. 1 typical aggregated scheme of a control mechanism subjects of control are: norms and restrictions of agents’ activity (institutional control), preferences of the agents (motivational control, including incentive problems) and awareness of the agents (informational control). sequence of moves in an organizational system (being common for all stages of a control cycle) under a fixed control mechanism is shown in fig. 1: stage i: the principal makes the first move, i.e., informs the agent of a control mechanism (“rules of play”) in a general form. for instance, this could be the relation between the amount of resource allocated and the reported need in the resource; alternatively, rules of play may be specified by the relationship between the reward and results achieved (the state of the agent). stage ii: the agent reports information on uncertain parameters to the principal (e.g., submits a claim for a resource or reports information on his or her preferences). stage iii: the principal informs the agent of control mechanism parameters. for instance, he or she assigns a plan as the result of the agent’s activity expected by the principal. stage iv: the agent chooses an action, and the result of activity is formed. stage v: the principal receives information on the agent’s action. stage vi: the principal informs the agent of his or her own action according to the control mechanism (e.g., the amount of resource allocated, the reward and so v.n. burkov: integrated mechanisms of organizational behavior control 220 on). stages i-iii (see fig. 1) correspond to the planning cycle, while stages iii-vi are related to the implementation cycle. the lower rectangle in fig. 1 represents the model of an economic agent. when solving design problems for sophisticated (complex) control mechanisms, it appears reasonable to describe the mechanisms using the notation similar to that of idef0 standard [13] and the input-output schemes in control science (see fig. 2). fig. 2 the general input-output scheme of a control mechanism within the framework of such representation, the six general stages discussed above (see fig. 1) could be formulated in the following way: a. the external conditions and constraints imposed on the principal during the process of making management decisions. b. the parameters of a control mechanism, being defined by the principal (stages i and iii). c. the input information from (on) agent (agents), being necessary to make a management decision (stages ii and iv). d. the result of the mechanism applied, i.e., the management decision made by the principal (stages iii and vi). now let’s consider the technique of integrating several mechanisms, which follow the unified description (fig. 1 and fig. 2), in order to solve certain control problem (the production cycle optimization problem is used as an example). 221 advances in systems science and applications (2013) vol. 13 no. 3 3 integrated mechanism of production cycle optimization production cycle optimization (pco) appears an important factor in improving the efficiency of a manufacturing process and in reducing the demand for circulating assets. pco represents an integrated mechanism, since the underlying process includes four primary stages as follows. stage 1. data acquisition regarding capabilities of pco (such information is supplied by units, i.e., shop floors and offices). stage 2. pco planning to ensure the required rate of optimization (what operations should be optimized and at what rates). stage 3. incentive scheme design (what rewards should be assigned to the units for reducing operation time of a manufacturing process). stage 4. implementation of pco plan and providing the rewards to the units. a simplified description of the pco-mechanism is shown in fig. 3. fig. 3 stages of production cycle optimization according to the stages mentioned, one would observe that (at least) four basic mechanisms are necessary: • stage 1 (data acquisition) – the mechanism for counter plans, which motivates the units to report higher capability estimates for pco; • stage 2 (pco planning) – the coordinated planning mechanism, which guarantees the required rate of pco in the sense of the minimum total costs to motivate such optimization; • stage 3 (incentive scheme design) – the incentive mechanism for a unit as the result of a certain optimization rate achieved; the mechanism must ensure truth-telling of all units; • stage 4 (implementation of pco plan) – the mechanism of predictive selfcontrol, which serves for well-timed informing of possible frustration of the plans. the mechanisms for counter plans ensure better estimates of feasible work time optimization (reported by the units) via coordinating the rewards for tense plans, the penalties caused by plan non-fulfillment and the rewards for plan overfulfillment. when an agent is paid merely for fulfillment (or overfulfillment) of the plan assigned by a principal, the agent is not interested in having a high (“tight”) plan. the reason is performing it would require additional efforts (costs) of the agent. for instance, an agent may inform the principal of his or her preferences, eo ipso reporting his or her estimate of the plan (referred to as a “counter plan”). v.n. burkov: integrated mechanisms of organizational behavior control 222 within the framework of incentive scheme for counter plans, an agent is given rewards for reporting counter plans that better meet the principal’s interests (yet, are “tighter” for the agent). the planning mechanism enables solving an optimization problem of defining the planned rates of operation time optimization to minimize the total optimization rate; this mechanism is also used to evaluate a reward norm for a unit optimization rate. for example, resource allocation mechanism is intended to distribute a specific resource based on claims of control objects (agents) for a desired quantity of the resource. note this is done under the conditions of deficiency (the resource is limited), while a control subject (a principal) has no information on optimal quantity for every agent. the mechanism ensures truth-telling of the claims submitted by the agents. modifying the mechanism parameters (priorities of the agents, prices, rates, etc) allows the principal to minimize his or her losses caused by the gap between actual distribution and its optimal counterpart (the latter could be achieved if the principal knew exact quantities of the resource needed by the agents). another example is a transfer price mechanism, which is involved to perform mutual payments within a company (in the system of internal accounting) or between companies that enter the same corporation. in the case of vertical integration, a transfer price is a tool to distribute profits between participants of a technological chain. horizontal integration being considered, a transfer price provides a tool of coordinating interests between the enterprises (units – agents) and a principal. the latter acquires from the former information on the price and quantity of a product they are ready to supply; in other words, the principal finds the relationship between an optimal output (or a plan) and a transfer price. based on the acquired information, the principal establishes a transfer price of the product such that the total output of the agents equals the required one. under a sufficiently great number of the agents, the mechanism of transfer prices ensures, first, truth-telling (since the agents benefit from reporting the actual dependence between the optimal output and the transfer price) and, second, minimum manufacturing costs for the products. the corresponding incentive mechanism could be represented by a proportional unified incentive scheme; here the income (bonus) of a unit equals the product of the reward norm and the planned rate of optimization. collective incentive scheme is designed for situations when a principal turns out unable to separately observe the action of every agent (the principal merely knows a certain aggregated rate, e.g., the result of collective activity). imagine that the principal can evaluate the minimum costs to-be-incurred by the agents to achieve the required result of collective activity. in this case, the efficient 223 advances in systems science and applications (2013) vol. 13 no. 3 incentive scheme takes the following form. the minimum costs of each agent are compensated (provided that the result of collective activity agrees with the requirements of the principal). moreover, sometimes the principal has no costs related to observing individual actions of every agent; thus, the principal’s workload to acquire and process information is substantially reduced. unified incentives mechanism is employed in situations when a principal has to motivate large groups of agents, to involve “democratic” management methods and to decrease the amount of processed data. under unified incentives, the relationship between the reward and labor intensity of the agents (alternatively, the results attained by them) is identical for all agents. in several cases, the described unification leads to no loss of efficiency, while the wages fund is spent in an optimal way. yet, unified control may be inefficient, when non-consideration of individual features of the agents results in inefficient spending of financial resources. the mechanism of predictive self-control allows a principal to obtain well-timed information on possible deviations from the planned rates of pco; this is done by coordinated assignment of the penalties for plan correction (they depend on the moment a certain unit reports of such correction) and the penalties for plan non-fulfillment. mechanism of predictive self-control is intended for well-timed informing a principal of possible deviations (from a plan) in the agents’ activity. the earlier the principal gets aware of possible deviations from the plan (e.g., in due dates, financial investments, etc), the more efficient and well-timed would be his or her decision (e.g., additional measures to eliminate deviations and reduce losses, or plan correction); note the agents report deviations. the matter is that the penalties of the agents (in the case of plan correction) depend on the moment the agents report of the correction (they are smaller if the report is early); moreover, these penalties are less than in the case of plan non-fulfillment. detailed description of these mechanisms may be found in [3]. fig. 4 shows a block diagram of the integrated mechanism of pco. the algorithm of pco-optimization mechanism operation: (1). each unit informs the principal of an optimization rate of the corresponding operations in a manufacturing process (depending on the reward norm). (2). the principal determines a minimum reward norm such that the duration of the manufacturing process is reduced to a required rate. the principal defines an optimization rate of the corresponding operations, using a rank-order tournament in the case of multiple-valued solution. (3). the units implement the tasks regarding production cycle optimization. (4). the principal computes the rewards of the units and pays them to the latter. v.n. burkov: integrated mechanisms of organizational behavior control 224 fig. 4 the block diagram of pco-optimization mechanism 4 conclusions the theory of control in organizations traditionally studies a certain system of nested control problems (solutions to “special” problems are widely used to solve more “general” ones) [3]. today, there are two common ways to describe an organizational system model (as well as to formulate and solve the corresponding control problems). they are referred to as the “bottom-top” approach and the “top-bottom” approach. according to the first (“bottom-top”) approach, particular problems are solved first; using the obtained solutions to the particular problems, general ones are treated then. for instance, a particular problem could be that of incentive scheme design. suppose this problem has been solved for any possible staff of the organization. next, one may pose the problem of staff optimization, i.e., choosing a certain staff to maximize the efficiency (under a proper optimal incentive scheme). an advantage of this approach lies in its constructivity. a shortcoming lies in high complexity due to the large number of possible solutions of the upper-level problem, each requiring the solution of the corresponding set of particular subproblems. the second (“top-bottom”) approach eliminates this shortcoming; this approach states that the upper-level problems must be solved first, while their solutions serve as constraints for the particular lower-level problems. in fact, we doubt whether (e.g., when creating a new department) a manager of a large-scale company would first think over the details of regulations describing the lowestlevel employees’ interactions. quite the contrary, he would delegate this task to 225 advances in systems science and applications (2013) vol. 13 no. 3 the head of the department (providing him with necessary resources and authorities). construction of an efficient control system for an organization requires combining both approaches in theory and in applications. as well, both approaches may be applied when designing certain integrated control mechanism. the challenge is to develop a simple and efficient technique of integration. references [1] novikov, d. and rusjaeva, e. (2012), “foundations of control methodology”, advances in systems science and application, vol. 12, no. 3, pp. 33-52. [2] burkov, v., goubko, m., korgin, n., and novikov, d. (2013), “mechanisms of organizational behavior control”, advances in systems science and application, vol. 13, no. 1, pp. 1-21. [3] novikov, d. (2013), theory of control in organizations, new york. nova science publishers. [4] bolton, p. and dewatripont, m. (2005), contract theory, cambridge, mit press. [5] salanie, b.m. (2005), the economics of contracts. 2nd edition, massachusetts, mit press. [6] moulin, h. (1995), cooperative microeconomics: a game-theoretic introduction, princeton, princeton university press. [7] nitzan, s. (2010), collective preference and choice, cambridge, cambridge university press. [8] camerer, c. (2003), behavioral game theory: experiments in strategic interactions, princeton, princeton university press. [9] myerson, r.b. (1991), game theory: analysis of conflict, london, harvard university press. [10] hillier, f. and lieberman, g. (2005), introduction to operations research (8th ed.), boston, mcgraw-hill. [11] taha, h. (2011), operations research: an introduction (9th ed.), new york, prentice hall. v.n. burkov: integrated mechanisms of organizational behavior control 226 [12] forrest, j., novikov, d. (2012), “modern trends in control theory: networks, hierarchies and interdisciplinarity”, advances in systems science and application, vol.12, no.3, pp.1-13. [13] idef0 overview at idef.com, http://www.idef.com/idef0.htm, last reviewed on july 7, 2011. corresponding author author can be contacted at: novikov@ipu.ru advances in systems science and applications (2013) vol.13 no.4 392-400 stability results for an inverse parabolic problem j. damirchi1 and a. shidfar2 1faculty of mathematics, statistics and computer science, department of mathematics, semnan university, semnan, iran. 2department of mathematics, iran university of science and technology, narmak, tehran 16844, iran. abstract in this paper, a one dimensional inverse parabolic problem in a quarter plane will be considered. the unknown function in a boundary is estimated from an over specified condition at a fixed location inside the region by solving an illposed integral equation. the tikhonov regularization method of the 1st order is applied in order to stabilize the solution of the ill-posed problem. the solution of the inverse problem is defined by minimization of the tikhonov functional. some analytical results for regularization parameter determination and stability of solution of inverse problem are derived. keywords tikhonov regularization method, inverse parabolic problem, regularization parameter, stable solution, ill-posed problem 1 introduction inverse parabolic problems play a crucial role in applied mathematics, physics and engineering science. they arise for example, in the study of heat conduction processes, diffusion, control theory [1-10]. in recent years, a lot of attention has been devoted to the study of inverse parabolic problems. hence, the last 20 years have seen growing attention paid in the literature to the development, analysis, and implementation of accurate methods for the solution of inverse parabolic problems, i.e., the determination of unknown boundary condition g(t) in the parabolic partial differential equation. in this paper, we investigate an inverse parabolic problem in quarter plane to obtain an unknown function g(t) from over specified data p(t) at a fixed location inside the body. this inverse problem can be reduced to the operator integral equation ag = g. it is well known that the operator equation ag = g is an ill-posed problem, when a solution is unstable with respect to small variations in input data. since in the usual case where only a measured or computation approximation gδ is available, some kind of regularization methods is required in order to obtain a reasonable stable approximation gδ to g [6, 11-15]. in order to obtain stable solution for this ill-posed problem the tikhonov regularization method is applied for operator equation ag = g for retrieving solution in a stable manner. applying general result in the theory of tikhonov regularization method for ill-posed inverse problem which consists in solving the unconstrained minimization problem, we can find stable solution. since the regularization paadvances in systems science and applications (2013) vol.13 no.4 393 rameter play an important role in applying the tikhonov regularization method to the operator equation, we obtain this parameter based on the error in input data directly, which can be one of the advantage of our method for selecting regularization parameter with respect to other methods [16-19]. the organization of this paper is as follows, in forcecoming section, mathematical formulation for this inverse parabolic problem in a quarter plane is introduced. in section 3, thikhonov regularization method is stated and we use this method to construct a stable solution for this ill-posed problem. some theoretical results will be proved about the solution of this ill-posed problem and the choice of regularization parameter. it will be shown that the solution of thikhonov regularization method is stable under small errors in input data and the existence of this stable solution is shown. we conclude this article with a brief conclusive discussion in section 4. finally, some references are given at the end of this paper. 2 mathematical formulation in this section, we consider the following inverse parabolic problem in a quarter plane: ut = uxx, 0 < x, 0 < t < t (1) u(x, 0) = f(x), 0 < x (2) u(0, t) = g(t), 0 < t < t, (3) |u(x, t)| ≤ m, 0 < x, 0 < t < t (4) where f(x) and g(t) are piecewise known continuous functions and m is a positive number.the problem consist of using an overspecified data p(t), which is given by p(t) = u(1, t), 0 < t < t (5) to determine the unknown function g(t). for the forward problem (1)-(4), the unique bounded solution u(x, t) is given by [20], u(x, t) = x√ 4π ∫ t 0 g(τ)√ (t− τ)3 e − x2 4(t−τ)dτ + 1√ 4πt ∫ ∞ 0 (e− (x−ξ)2 4t − e− (x+ξ)2 4t )f(ξ)dξ (6) due to the setting x = 1, and using (5) we obtain ∫ t 0 e − 1 4(t−τ)√ (t− τ)3 g(τ)dτ = 2 √ πp(t)− 1√ t ∫ ∞ 0 (e− (x−ξ)2 4t − e− (x+ξ)2 4t )f(ξ)dξ (7) 394 j. damirchi: stability results for an inverse parabolic problem which is written in the form∫ t 0 h(t− τ)g(τ)dτ = 2 √ πp(t)− 1√ t ∫ ∞ 0 (e− (x−ξ)2 4t − e− (x+ξ)2 4t )f(ξ)dξ where the kernel h(t, τ) = 1√ (t−τ)3 e − 1 4(t−τ) is a continuous function on [0, t ] × [0, t ]. now, we formulate (7) in term of an operator integral equation : (ag)(t) = g(t), 0 < t < t (8) where the operator a is defined by (ag)(t) = ∫ t 0 e − 1 4(t−τ)√ (t− τ)3 g(τ)dτ, 0 < t < t and g(t) = 2 √ πp(t)− 1√ t ∫ ∞ 0 (e− (x−ξ)2 4t − e− (x+ξ)2 4t )f(ξ)dξ, 0 < t < t the above integral equation of the first kind (7) cannot be reduce into an integral equation of the second kind by differentiation and the problem is inherently ill-posed. for 1 ≤ p < ∞, a is a compact linear operator in lp[0, t ]. zero is not an eigenvalue of a and is the only point in the spectrum of a, thus a−1 exist and is unbounded, so if g(t) on 0 ≤ t ≤ t is in the range of a, g(t) is uniquely determined from g = a−1g. in practice, with g(t) obtained from measurement, small error in g(t) lead to enormous errors in g(t) because a−1 is unbounded. in order to regularize the problem, we use tikhonov regularization method for finding approximate stable solution for ill-posed integral equation (8). in the next section we describe tikhonov regularization method. 3 tikhonov regularization method for ill-posed integral equation it is well known that the tikhonov regularization method is one of the useful tools for solving an ill-posed problem of the form (7)[12-13, 19]. since in practical purpose, the input data are non smooth, we apply tikhonov regularization method to construct stable solution for solving ill-posed equation (8). based on the main concept of tikhonov regularization method, we introduce smoothing functional mα[g,g] which is called also thikhonov functional and stabilizing functional ω(g) as follows: mα[g,g] = ∥ag −g∥2l2[0,t ] + αω(g) advances in systems science and applications (2013) vol.13 no.4 395 α is regularization parameter and ω(g) is a stabilizer of the 1th order constant coefficient which is defined as: ω(g) = ∫ t 0 (g2(τ) + g′2(τ))dτ we construct a regularize solution for the integral equation (8) by using the following minimization problem which is state by the following theorem: theorem 1. for every function g ∈ l2[0, t ], and any positive α, there exists an element gα ∈ w 1 2 , such that smoothing functional mα[g,g] attain it greatest lower bound. infmα[g,g] = mα[gα, g] proof. we have mα[g,g] = ∫ t 0 ( ∫ t 0 e − 1 4(t−τ)√ (t− τ)3 g(τ)dτ −g(t))2dt+ α ∫ t 0 (g2(τ) + g′2(τ))dτ a condition for a minimum of this functional is vanishing of its first variation, this is written in the form 1 2 d dε mα[g + εη,g]|ε=0 = ∫ t 0 [ ∫ t τ ∫ t 0 h(t− s)h(t− τ)g(s)dsdt− ∫ t τ h(t− τ)dt − α(g′′(τ)− g(τ))]η(τ)dτ + g′(τ)η(τ)|t0 =0 (9) here, η(τ) is an arbitrary variation of the function g(τ) such that both g(τ) and g(τ) + εη(τ) belong to the class of admissible function. condition (9) will be satisfied if∫ t τ ∫ t 0 h(t− s)h(t− τ)g(s)dsdt− ∫ t τ h(t− τ)dt = α(g′′(τ)− g(τ)) (10) and g′(0) = g′(t ) = 0 (11) the equation (10) is called euler-lagrange equation, therefore the minimizer gα(t) for tikhonov functional is determined by the solution of euler-lagrange equation corresponding to functional mα[g,g]. the solution of (10) with boundary condition (11) is unique by classical theorems in ode. now, we can assume that the regularized solution of the above minimization problem, gα as an regularizing operator r(g,α) such that gα = r(g,α) where 396 j. damirchi: stability results for an inverse parabolic problem α = α(δ,gα) in accordance with the error in the initial data g and δ, measuring the error in data, we select α in a suitable way such that r(g,α) is a regularizing operator for the equation (8), and gα = r(gα, α(δ)) can be taken as an approximate stable solution for ill-posed problem (8). theorem 2. if gt (t) ∈ c[0, t ] be the exact solution of the original problem (8) associated with the exact right hand member g = gt ; that, (agt )(t) = gt (t). then, for every positive number ε, there exists a positive number δ(ε) such that for every gδ ∈ l2[0, t ] the inequality ∥gt (t)−gδ(t)∥l2[0,t ] ≤ δ < δ(ε) implies the inequality ∥gα(δ)(t)− gt (t)∥c[0,t ] ≤ ε where gα = r(gα(δ), α(δ)) be the solution of (8) associated with perturbed data gδ for all α satisfying α(δ) = δγ , 0 ≤ γ < 2. proof. since gα is a minimizer of functional mα, we have mα(δ)[gα, gδ] ≤ mα(δ)[gt , gδ] therefore, ∥agα −gδ∥2 ≤ mα(δ)[gα, gδ] ≤ ∫ t 0 (agt (t)−gδ(t)) 2dt+ α(δ) ∫ t 0 (g2t (τ) + g′2t (τ))dτ = ∫ t 0 (gt (t)−gδ(t)) 2dt+ δγ ∫ t 0 (g2t (τ) + g′2t (τ))dτ ≤ δ2 + δγ ∫ t 0 (g2t (τ) + g′2t (τ))dτ ≤ δγ(1 + ∫ t 0 (g2t (τ) + g′2t (τ))dτ) = δγn with n = 1 + ∫ t 0 (g2t (τ) + g′2t (τ))dτ . consequently, the elements gt (t) and gα(δ)(t) belong to the compact subset e of element g of c[0, t ] such that e = {g(t)| ∥g∥2w 1 2 ≤ n} since e is compact in c[0, t ], and the operator a is continuous, the mapping a : e → ae is continuous and one to one, therefore the inverse mapping a−1 : ae → e is also continuous. this means that, ∀ε > 0, ∃η(ε), ∥gt −gα∥ ≤ η(ε), agt = gt , agα = gα advances in systems science and applications (2013) vol.13 no.4 397 then ∥gt − gα∥c[0,t ] ≤ ε on the other hand we have ∥gt −gα∥2l2 = ∫ t 0 (agα −gt (t)) 2dt < δ2 an so, ∥gt (t)− gα(δ)(t)∥c[0,t ] = ∥a−1agt −a−1agα(δ)∥ ≤ ∥a−1∥∥agt −agα(δ)∥ on the other hand ∥agt −agα(δ)∥l2 ≤ ∥agt −gδ∥+ ∥agα(δ) −gδ∥ ≤ ∥gt −gδ∥+ ∥agα(δ) −gδ∥ ≤ δ + δ γ 2 √ n ≤ δ γ 2 (1 + √ n) therefore, ∥gt − gα(δ)∥ ≤ ∥a−1∥δ γ 2 (1 + √ n) the above results show that δ(ε) should be chosen in the form δ(ε) ≤ [ ε ∥a−1∥(1 + √ n) ] 2 γ such that the theorem is satisfied. the above theorem shows that when we construct regularizing operator by minimizing the smoothing functional mα[g,g], the regularization parameter α can be obtain according to the error in the right hand member. by the next theorem we will show that g depends continuously on the initial data f, p. theorem 3. if the exact data ft (x) and pt (t) satisfy (8) and the approximate data fδ(x) and pδ(t) also satisfy (8), then inequalities ∥ft (x)− fδ(x)∥l2[0,∞] < δ and ∥pt (t)− pδ(t)∥l2[0,t ] < δ imply that ∥gt (t)−gδ(t)∥l2[0,t ] < δd, d = (9 √ 2πt + 12π) 1 2 398 j. damirchi: stability results for an inverse parabolic problem proof. by using of cauchy-schwartz inequality, we have ∥gt (t)−gδ(t)∥2 =∥2 √ π(pt (t)− pδ(t))− 1√ t ∫ ∞ 0 (e− (1−x)2 4t − e− (1+x)2 4t )(ft (x)− fδ(x))dx∥2 ≤3[ ∫ t 0 4π(pt (t)− pδ(t)) 2dt+ ∫ t 0 1 t ( ∫ ∞ 0 e− (1−x)2 4t (ft (x)− fδ(x))dx) 2dt + ∫ t 0 1 t ( ∫ ∞ 0 e− (1+x)2 4t (ft (x)− fδ(x))dx) 2dt] ≤3(4δ2π + ∫ t 0 1 t ( ∫ ∞ 0 (e− (1−x)2 4t )2dx)( ∫ ∞ 0 (ft (x)− fδ(x)) 2dx)dt + ∫ t 0 1 t ( ∫ ∞ 0 (e− (1+x)2 4t )2dx)( ∫ ∞ 0 (ft (x)− fδ(x)) 2dx)dt) ≤3(4δ2π + δ2 ∫ t 0 1 t ∫ ∞ 0 e− (1−x)2 2t dxdt+ δ2 ∫ t 0 1 t ∫ ∞ 0 e− x2 2t dxdt) ≤3(4πδ2 + 3 √ 2πtδ2) =δ2(12π + 9 √ 2πt ) theorem (3) and (4) show that by choosing α in such a way that the choice for α is consistent with the accuracy δ of the initial data, then the element gα = r(gα, α) obtained with the aid of the regularizing operator r(g,α), can be taken as approximate stable solution of equation (8). we summarize the above results by the following theorem: theorem 4. if gt (t) is the exact solution of equation (8) with exact data ft (x) and pt (t), then for every ε > 0 and the approximate data fδ(x) and pδ(t) also satisfy (8), there exist δ(ε) and α(δ) such that inequalities ∥ft (x)−fδ(x)∥l2[0,∞] < δ and ∥pt (t)− pδ(t)∥l2[0,∞] < δ imply that inequality ∥gt (t)−gα(t)∥ < ε where gα(δ) = r(gδ, α(δ)). 4 conclusion in this paper, we have introduced tikhonov regularization method and we have shown why it is important to use it in order to solve ill-posed problems. we have shown that the tikhonov regularization technique is introduced to treat the instability of obtaining a stable solution. the choice of regularization parameter and stability of solution are proved. it is very interesting to extend these results for higher dimensional problem and for nonstandard heat equation. advances in systems science and applications (2013) vol.13 no.4 399 references [1] ramm a g. (2005), inverse problems, springer, new york. [2] ramm a g. (2001), “an inverse problem for the heat equation”, j. math. anal. appl, vol.264, pp.691-697. [3] cannon j r, lin y, wang s. (1992), “determination of source parameter in parabolic equations”, meccanica, vol.27, pp.85-94. [4] chen q, liu j j. (2006), “solving an inverse parabolic problem by optimization from final measurement data”, j. comput. appl. math, vol.193, pp.183-203. [5] dehghan m, tatari m. (2006), “determination of a control parameter in a one-dimensional parabolic equation using the method of radial basis functions”, math. comput. modelling, vol.44, pp.1160-1168. [6] zheng g h, wei t. (2011), “spectral regularization method for solving a time-fractional inverse diffusion problem”, appl. math. comp, vol.38, pp.317-336. [7] isakov v. (1998), inverse problems for partial differential equations, springer, new york. [8] alifanov o m. (1994), inverse heat transfer problem, springer, berlin heidelberg. [9] shidfar a, azary h. (1997), “an inverse problem for a nonlinear diffusion equation”, nonlinear analysis theory method, application, vol.28, no.4, pp.589-593. [10] shidfar a, karamali g r, damirchi j. (2006), “an inverse heat conduction problem with a nonlinear source term”, nonliner analysis theory method, application, vol.65, pp.615-621. [11] liu j j. (2005), regularization method and application for the ill-posed problem, science press, beijing. [12] engl h w, hanke m, neubauer a. (1996), regularization of inverse problems, kluwer academic publishers, dordrecht. [13] egger h, engl h w. (2005), “tikhonov regularization applied to the inverse problem of option pricing: convergence analysis and rates”, inverse problems, vol.21, pp.1027-1045. 400 j. damirchi: stability results for an inverse parabolic problem [14] fu cl, xiong xt, fu p. (2005), “fourier regularization method for solving the surface heat flux from interior observations”, math. comput. model, vol.42, pp.48998. [15] cheng j, yamamoto m. (2000), “one new strategy for a-priori choice of regularizing parameters in tikhonov’s regularization”, inverse problems, vol.16, pp.31-36. [16] tikhonov a n, arsenin v y. (1997), solution of ill-posed problems, washington, d.c: v.h. winston & sons. [17] bakushinski a b. (1984), “remarks on choosing a regularization parameter using the quasioptimality and ratio criterion”, comput. math. math. phys, vol.24, pp.1812. [18] wang z, liu j. (2009), “new model function methods for determining regularization parameters in linear inverse problems”, appl. numer. math, vol.59, pp.2489-2506. [19] liu w, sun x, shen j. (2012), “a v-curve criterion for the parameter optimization of the tikhonov regularization inversion algorithm for particle sizing”, optics, laser technology, vol.44, pp.1-5. [20] cannon j r. (1984), the one-dimensional heat equation, addison-wesley, california. corresponding author j. damirchi can be contacted at: damirchi.javad@gmail.com adv syst sci appl 2018; 4; 12-19 published online at http://ijassa.ipu.ru/ severity of breast masses prediction in mammograms based on optimized naive bayes diagnostic system abeer s. desuky faculty of science, al-azhar university, cairo, egypt e-mail: abeerdesuky@azhar.edu.eg abstract. mammography is the most effective tool for breast masses screening. it is a special ct scan technique used only to detect breast tumors early and accurately. detecting tumors in its early stage has improved the survival rate for breast cancer patients. computer aided diagnostic systems help the physicians to detect breast cells abnormalities earlier than other traditional procedures. in this paper, an improved naive bayes classifier based on chicken swarm optimization algorithm (cso-nbc) is analyzed on mammographic mass dataset. the main aim of this research is to increase physician's ability to determine the severity of a mammographic mass lesion from the bi-rads features and the patient's age using the bio inspired chicken swarm optimization (cso) algorithm for naive bayes classifier (nbc) improvement. the dataset is preprocessed and divided to train the cso-nbc system and test it by 5-folds cross validation technique. the performance of our proposed classification system is compared with papers' results of other researchers to show the efficiency of our system in predicting severity of breast tumors with the highest accuracy. keywords: breast masses mammography, naive bayes, chicken swarm optimization, computer aided diagnosis. 1. introduction cancer is one of the top leading causes of death. it is caused by uncontrolled growth of cells which invade and spread around the body, often resulting in death. cancer is the second leading cause of death globally, and was responsible for 8.8 million deaths in 2015 [1]. globally, nearly 1 in 6 deaths is due to cancer according to the who (world health organization). per the us national cancer institute, 60 percent of the world’s new cancer cases happen in asia, africa, central and south america, and 70 percent of global cancer deaths occur in those same regions as well. moreover, breast cancer death-rate was 571 000 equivalently 6.5 percent of all deaths in 2015 [2]. because of this fact, early detection and true diagnosis is an important issue and plays a key role to reduce mortality of this disease. mammography is an efficient imaging appliance for early breast cells abnormality detection. mammography exists in two types: first type, screening mammography which is breast x ray used to test the changes in breast area for early detection of breast tumors in women with no symptoms of cancer. it can also detect tiny calcium deposits (micro calcification) that is one of cancer manifestations. the second type, diagnostic mammography which is a breast x ray to check the signs of breast cancer after detecting a mass or other symptom. symptoms include breast size/shape change, skin thickening, pain or nipple discharge [3, 4]. significant improvements can be made in the lives of breast cancer patients by detecting cancer early and avoiding delays in care, computer aided diagnosis (cad) can help the 13 a. s. desuky copyright ©2018 assa. adv. in systems science and appl. (2018) physicians to do this task faster and more accurate than traditional procedures. during the last decade with development of machine learning approaches in diagnostic systems, breast cancer detection has improved. the aim of using machine learning approaches is to minimize mistakes that may occur by specialists in diagnosis [5]. different methods have been proposed to detect and classify masses in mammogram images. ramani and vanitha [4] used weighted histogram algorithms to select features from mammogram data then classified the selected features using random forest, naive bayes and ann algorithms. another method proposed in [6] based on missing values imputation and three models svm with polynomial kernel, ann with pruning parameters and dt with chi-squared interaction detection were derived for prediction. bio-inspired swarm algorithms can be used to enhance effectiveness of cad systems. this paper presents the bio inspired chicken swarm optimization (cso) algorithm to improve naive bayesian classifier (nbc) for the severity of breast masses prediction in mammograms. 2. the proposed method 2.1. chicken swarm optimization chicken swarm optimization (cso) algorithm proposed by meng et al. [7], is a bio-inspired swarm algorithm which mimics the behavior of chicken swarms in nature. each chicken swarm contains several groups of one rooster with two or more hens and many chicks move in a hierarchy groups and searching for food which has an effect on the swarm movements. dividing the chicken swarm into groups and identifying the chickens into roosters (r), hens (h) and chicks (c) mainly depends on the fitness values of the chickens. the chickens with best fitness values would be the roosters. chicks would be the chickens with worst fitness values while the others would be the hens (mothers). these statuses are updated according to the change in the fitness values only every several (g) time steps. there are three basic movements in the swarm as proposed in [7]: first movement is rooster’s movement which has the better fitness values since it can search for food in a wider range than the rest of chicks with less fitness values, the movement of the rooster is defined in eq. 1 and eq. 2: (1) where x is the selected rooster, is a gaussian distribution with mean 0 and standard deviation , and f is the fitness value of the corresponding rooster x. (2) where k is rooster index, i is the related position and ε is the smallest integer number in the computer used to avoid zero-division-error. second movement is hen’s; hens follow the roosters to search for food and their movement is defined in eq. 3 – eq. 5: (3) (4) (5) severity of breast masses prediction in mammograms 24 copyright ©2018 assa. adv. in systems science and appl. (2018) where r1 and r2 are the indices of the rooster and the randomly chosen chicken (rooster or hen), r1 ≠ r2, and rand is a uniform random value between [0, 1]. third movement is the movements of the chicks which follow their mothers, defined in eq. 6: (6) where xm,j is the position of the chick’s mother, and l is a randomly chosen parameter between 0 and 2. 2.2. naive bayesian classifier the bayesian classifier represents a statistical classification algorithm as well as a supervised learning method for classification based on bayes theorem [8]. assuming a probabilistic model, in bayes theorem, posterior probability of a data sample x with unknown class label is calculated as in eq. 7. (7) where each data sample is represented by an n-dimensional feature vector, x = (x1, x2.... xn) and h is some hypothesis, such that the data sample x belongs to a specified class label l1, l2.... lm. the nb classifier will predict that the given data sample x belongs to the class label li that have the higher posterior probability, i.e. for p (li | x) is the maximum [9]. for 1 ≤ j ≤ m (8) 2.3. the dataset researchers proposed several (cad) computer aided diagnostic systems in the last few years. these systems help doctors in deciding to perform short term follow-up examination or proceed a breast biopsy on a suspicious lesion seen in a mammogram instead. table 1. mammographic mass dataset features feature type values and labels missing values bi-rads ordinal (nonpredictive) 0 assessment incomplete 1 negative (non-predictive) 2 benign findings 3 probably benign 4 suspicious abnormality 5 highly suggestive of malignancy 2 age integer 18 96 (patient's age in years) 5 shape nominal 1 round 2 oval 3 lobular 4 irregular 31 margin nominal 1 circumscribed 2 microlobulated 3 obscured 4 ill-defined 5 spiculated 48 density ordinal 1 high 2 iso 3 low 76 15 a. s. desuky copyright ©2018 assa. adv. in systems science and appl. (2018) 4 fat-containing severity binominal (goal field) 0 benign 1 malignant 0 bi-rads (breast imaging reporting and data system) is developed to standardize nomenclature between radiologist and doctors. bi-rads assessment is done based on shape, age, margin density. bi-rads features and the patient's age in mammographic data set can be used to predict the severity of a mammographic mass lesion. this data set contains the patient's age, a bi-rads assessment, and three bi-rads features together with the class feature (severity) for 445 malignant and 516 benign masses that have been collected and identified on full field digital mammograms at the institute of radiology of the university erlangen-nuremberg between 2003 and 2006. table 1 shows that each instance has bi-rads features, age, shape, margin, density and severity features. mammographic data set includes 162 missing values [6, 10]. 2.4. naive bayes based on chicken swarm optimization traditional nbc is a statistical classifier with mutual independency among the features [4]. in the medical domain, all the symptoms do not contribute equally in diagnosing a specific disease. for example, in heart disease domain, the bmi feature is having less impact than the prior-stroke feature in predicting the probability of being heart patient [9]. in our proposed work called csonbc we assign weights to each feature according to their efficiency in prediction. we use cso technique to learn the features weight in nbc. no information or assumptions are needed about the weight in nbc, the mechanism of cso can help us get automatically the optimal weights. fig. 1 cso-nb classification technique we use the weight to increase the classification performance of the nbc. a high weight will be assigned to features that have more impact in prediction and low weight are assigned to features that have less impact. so, our main aim is to get optimum weight configuration that gives the highest accuracy for cso-nbc. the major steps of the working technique are shown in figure 1 and described as follows: severity of breast masses prediction in mammograms 26 copyright ©2018 assa. adv. in systems science and appl. (2018) first step: the mammographic mass data is preprocessed divided into training and testing sets then, the nbc applied on it firstly to measure the nb classification performance. second step: an initial generation for weight individuals (chickens) is formed in a random manner. each chicken swarm is a vector of all weight variables ranging from -10 to 10. for the weight individuals, the generation size n should be determined first. third step: the total generation cost of each chicken is calculated. fourth step: generate new rooster around the global best swarm by searching for the best fitness index and the worst fitness index of chicken swarm in same iteration and store their indexes; then, cso-nbc replaces the worst index with the best index. fifth step: third and fourth steps are repeated till get the best individual weight which is the one getting the highest classification accuracy. last step: the test instances can be classified under the learned weight by cso-nbc and the performance measures can be determined. 3. experimental results the effectiveness of any diagnostic system is evaluated by its ability to give the maximum accurate classification result [11]. per the real nature for any given case and the prediction from diagnostic system, there are only four possible outcomes, true negatives (tn) as well as true positives (tp) correspond to a correct diagnosis; that is, cases are successfully labeled as uninfected and infected patients, respectively; false positives (fp) refer to uninfected patient being classified as infected; false negatives (fn) are infected patients incorrectly classified as uninfected. the most popular performance metrics is accuracy which can be calculated as: n n n (9) to evaluate our system (cso-nbc), besides the classical accuracy, the standard metrics of sensitivity, g-mean (geometric mean), and f-measure have been used. n (10) (11) 2 2 n (12) a diagnosis system should have a high accuracy, sensitivity, g-mean and f-measure [4,12]. he experiments were implemented on a computer with intel core™ i7 processor and 6gb ram running microsoft windows 7 professional and the algorithm is coded in matlab 15. first, we replaced all the missing features values substitute it with the average in the mammographic mass dataset then, applied information gain (ig) feature selection technique [13], which is used to evaluate the importance of a feature by measuring the information gain (entropy) with regard to the class to remove useless features and improve the classification performance. the instances are randomly divided into training and testing sets in a 5-folds cross validation manner. the used parameters in evolution were: 100 for population size, 1000 generations (iterations) and time steps (g=10). tables 2, 3, 4 and 5 show the performance evaluation: classification accuracy, sensitivity, gmean and f-measure respectively of the nb classifier and the proposed technique using the full 17 a. s. desuky copyright ©2018 assa. adv. in systems science and appl. (2018) mammographic mass dataset features and the subset of selected features after applying information gain technique and removing “density” feature. the results show a perceivable improvement in classification performance for cso-nbc and show that with ig feature subset selection our proposed technique enhanced the classification performance against the enhancement with nb classifier. table 2 classification accuracy table 3 sensitivity nb cso-nbc full data 70.16 89.52 reduced data 87.57 82.94 table 4 g-mean nb cso-nbc full data 78.81 82.63 reduced data 85.11 87.37 table 5 f-measure nb cso-nbc full data 77.93 85.80 reduced data 86.48 87.36 table 6 comparison of algorithms applied on mammographic mass data method accuracy proposed cso-nbc 87.20 nb [5] 82.49 svm [6] 81.25 bagging svm-smo [9] 82.00 bagging dt [9] 83.40 the experiments also showed table 6 -that the proposed algorithm cso-nbc outperforms algorithms proposed by other researchers applied on the mammographic mass data. 4. conclusion mammography is used to help in the early detection (diagnosis) of breast diseases. the computer aided diagnostic systems can help physicians in detection (diagnosis) of nb cso-nbc full data 78.87 83.88 reduced data 85.42 87.20 severity of breast masses prediction in mammograms 28 copyright ©2018 assa. adv. in systems science and appl. (2018) abnormalities early than traditional procedures. many algorithms and diagnostic systems have been developed recently. the objective is not to replace medical researchers and professionals, but to increase their abilities to take decisions about the disease. this paper proposed a new automated mammogram classification system cso-nbc based on the improved naive bayes classifier using chicken swarm optimization algorithm. the mammographic mass data is preprocessed by replacing all the missing features values and applying information gain (ig) feature selection technique to remove the useless features. the dataset then divided to training and testing sets using 5-folds cross validation technique. the experiments showed the efficiency of our proposed classification system in predicting severity of breast tumors since it performs better than the traditional nb classifier also, better than algorithms proposed by other researchers for the same mammographic mass data. references [1]world health organization (2018, may) cancer [online]. available: http://www.who.int/cancer/en/ [2] world atlas (2018, may) explore the world [online]. available: http://www.worldatlas.com [3] sickles, e. a., wolverton, d. e., & dee, k. e. (2002). performance parameters for screening and diagnostic mammography: specialist and general radiologists. radiology, 224(3), 861869. [4] ramani, r. & suthanthira vanitha, n. (2014). computer aided detection of tumors in mammograms, int.j. image, graphics and signal processing, 4, 54-59. [5] güzel, c., kaya, m. & yıldız, o. (2013). breast cancer diagnosis based on naïve bayes machine learning classifier with knn missing data imputation, awerprocedia information technology & computer science, 4, 401-407, [online]. available: www.awer-center.org/pitcs [6] mokhtar, s. a., elsayad, a. m. (2013). predicting the severity of breast masses with data mining methods, international journal of computer science issues, 10(2). [7] meng, x., liu, y., gao, x. & zhang, h. (2014). a new bio-inspired algorithm: chicken swarm optimization. in: tan y., shi y., coello c.a.c. (eds.) advances in swarm intelligence. icsi 2014. lecture notes in computer science, vol 8794. springer, cham, https://doi.org/10.1007/978-3-319-11857-4_10 [8] ozer, p. (2008). data mining algorithms for classification, b.sc thesis, redbound university nijimegan. [9] jeyarani, d. s., anushya, g., raja, r. & pethalakshmi, a. (2013). a comparative study of decision tree and naive bayesian classifiers on medical datasets, international journal of computer applications (0975 – 8887), international conference on computing and information technology (ic2it-2013), 5-7. [10] luo, s.-t., cheng, b.-w. (2012). diagnosing breast masses in digital mammography using feature selection and ensemble methods, journal of medical systems, 36(2), 569-77. https://doi.org/10.1007/s10916-010-9518-8. [11] el bakrawy, l. m. & desuky, a. s. (2015). a hybrid classification algorithm and its application on four real-world data sets, international journal of computer science and information security (ijcsis), 13(10): 93-97. [12] han, j. & kamber, m. (2001). data mining concepts & techniques. san francisco, ca, usa: morgan kaufmann publishers inc. 19 a. s. desuky copyright ©2018 assa. adv. in systems science and appl. (2018) [13] kumar, c. s. & sree, r. j. (2014). application of ranking based attribute selection filters to perform automated evaluation of descriptive answers through sequential minimal optimization models, ictact journal on soft computing, 5(1), 860-868. advances in systems science and application (2015) vol.15 no.3 267-278 modern changes of the climatic conditions and rhythmicity of the long-term oscillations of the mode parameters of water bodies as the integral index of the climate fluctuation(based on the urals example) nadezhda s. rasskazova1, aleksandr v. bobylev1, michail n. bubin2 and alexander v. malaev3 1south ural state university russian federation 454080, chelyabinsk lenin v.76. 2yurginski technological institute (branch) of the national research tomski polytechnic university 652055, russian federation, kemerovo region, yurga, st. leningradski, 26. 3 chelyabinsk state pedagogical university, russian federation, 454074, chelyabinsk, st.bazhova 48 abstract the mechanism of how global factors of terrestrial and extraterrestrial origins influence the natural objects is very complicated and has not been fully studied yet. the majority of scientists think that the lunar and solar tidal forces produce an effect on the water surface of the sea, making the ocean currents pulse under certain rhythms. their pulsation causes changes in the global moisture and heat transfer on the earth. as a result, there appear complex rhythmic oscillations of the planets climate and inland waters. this article represents an analysis of the reasons for the modern climate fluctuations, including those at the local level (the urals region). keywords rhythmicity of the long-term river flow oscillations; dependence on cosmo and geophysical factors; prevailing atmospheric circulation type c (longitudinal); rossby waves 1 introduction determining the rhythms of natural phenomena and their reasons represents one of quite important tasks of the modern science. rhythmicity is the major parameter of the long-term oscillations of the natural processes in time and space. the major rhythms that determine the development character of the natural phenomena are the cosmic cycles. the cosmic rhythms of various durations and different origins permeate through and regulate every development process of the earth. the authors have found out that the reasons for the modern climate fluctuations in the urals, along with the major rhythms, are the longitudinal atmospheric circulation type c, which prevails in the zauralie area, and the rossby waves. to determine the major rhythms, the authors used fourier analysis and the difference integral curve method. to reveal the connection between the long-term oscillations of the hydrological characteristics and the cosmoand 268 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... geophysical factors which represent a reflection of the climate fluctuations, cluster analysis was used. the accuracy of the research results was estimated using the mathematical statistics methods. 2 material & methods there are two essentially different rhythm categories in nature: the limited set of the cosmic rhythms and the unlimited number of the derivative interaction rhythms or environment rhythms. unlike the interaction rhythms which appear in particular environments and in limited territories, the cosmic rhythms reveal themselves, with amazing stability, in every geocomponent and geosphere. according to their duration, rhythms are divided into centuries-old and interdecadal ones. among the centuries-old rhythms, those with durations of 120, 300, 600, and 1,200 to 2,000 years are the most often found in nature. but it is not only natural phenomena that obey the rhythmicity laws. l.n. gumilev [1] noticed that even ethno-social processes taking place on earth are also subject to rhythms, and e.v. maximov [2] found that the passionary shocks described by gumilev [1] , which lead to disappearance of some nations and birth of some others, are connected with the cosmic rhythms, and in particular with one of the most important cosmic rhythms with the duration of about 1,800 years. on the shores of the picturesque lake zurich there are two ancient terraces: high cliffs where, in the rock strata, once can clearly distinguish the strata of different eras. in these stratified rocks, scientists found some very clear evidence of the 1,800-year rhythm. the same rhythm was also determined in the sequence of the oozy deposits, the glacier movements, humidification oscillations, and after all in the climate fluctuations. of course, we cannot observe any rhythms of such durations in that short life we live. however, there are a whole series of space-generated interdecadal rhythms the effects of which can be felt during our lifetime: 80 to 90, 28 to 35, 21 to 22, 15, 11 to 13, 5 to 8, and 2 to 4 years. these are the interdecadal rhythms. the interconnection between the terrestrial and cosmic phenomena becomes the most evidential if based on the example of the rhythms occurring in the physical and geographical processes recorded by means of instrumental observations which are introduced in table 1. today, the overwhelming majority of researchers have no doubt that rhythmicity is an integral natural law of the earths geosphere, and that its rhythmical fluctuations are caused by the cosmic rhythms. the cosmic rhythms give rise to rhythms in the geosphere, similar to themselves, while derivative rhythms produced by their interference cause accidental fluctuations. rhythmical oscillations have been found in the solar activity, the earths magnetic field, atmospheric precipitations, air and water temperature, tides, and advances in systems science and application (2015) vol.15 no.3 269 many other natural phenomena. they are conditioned by three groups of factors: astronomical, geophysical and circulating. table 1 average annual indices reflecting the connection between the long-term oscillations of the annual river flow and the cosmoand geophysical factors dendro-gram number index factors represented by the index symbol unit of measure 175 relative numbers of sunspots (wolf numbers) solar activity w n/a 195 troposphere-effective indices of solar activity (according to loginov) t 186,187 19-year constituent of the lunar and solar tide-generating force potential at 56◦ and 60◦ north polar tide f19 cm2/sec2 188 average annual values of the day length fluctuations earth axis nutation △t of 0.00005 sec 176,177 longitudinal and latitudinal coordinates of the icelandic low activity of pressure formations φ,λ degree 178 soi (southern oscillation index) intensity of ocean currents isoi mb 179 nao (north atlantic oscillation index) inao mb 185 longitudinal circulation indices in the northern atlantic (january) rn mb 189,190 water temperature anomalies in the northern atlantic in smed squares d and e △t◦ degree 180,181 annual number of days with cyclones in the 4th and 8th vitels zones global atmospheric circulation z4, z8 days192 recurrence of the western type a.c. w 193 recurrence of the eastern type a.c. e 194 recurrence of the longitudinal type a.c. c 184 annual anomalies of the surface air temperature in the arctic underlying surface humidification long-term river flow oscillations in the urals t degree 182,183 aridity index dm (etp,atp) dm n/a 197-199,301-317 humidification coefficient chumid n/a q m3/sec 270 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... the oscillations of the hydrological parameters of water bodies (river flows and levels) are taken as the integral index of the climate fluctuations. the mechanism of how global factors of terrestrial and extraterrestrial origins influence the natural objects is very complicated and has not been fully studied yet. the majority of scientists think that the lunar and solar tidal forces produce an effect on the water surface of the sea, making the ocean currents pulse under certain rhythms. their pulsation causes changes in the global moisture and heat transfer on the earth. as a result, there appear complex rhythmic oscillations of the planets climate and inland waters. the fluctuations of the solar activity, in its turn, is connected with the gravitational effect of the planets in the solar system and with the effect produced on it by other space bodies: comets, asteroids and meteorites coming to us from outer space and breaking the abovementioned rhythms. in particular, the average rhythm of meteoric impacts on the earth is about 80 years[3]. as for the nature of various cycles, scientific literature still introduces quite contradictory hypotheses. the coincidence in the durations of the cycles of specific phenomena is normally considered to be the major argument for the existence of some connections between the external factors and the oscillations of the hydrometeorological elements. but this kind of approach is one-sided and insufficient to make statements on the existence of those connections and to reveal their nature. 3 results so we spread our research both in the time and space aspects. as it was said above, we used the fourier analysis and difference integral curves to determine the rhythms, and cluster analysis was used to reveal the connection between the natural phenomena and cosmic rhythms. the research resulted in the map (fig. 1) that the author made using the cartographic method. according to the studies that have been performed [4] the full 80-year solar rhythm can be observed in the series with execution lengths of over 100 years, such the wolf numbers, the nao (north atlantic oscillations) indices, the soi (el niño southern oscillation indices), and others. it means that all the rhythms are the reflection of the solar rhythms and have the cosmic origin. but unfortunately, their accuracy is so low that determining such a rhythm requires series of 250 years (3 x 80) and more. the map (fig. 1) shows that the rhythm of 80-90 years (the full solar rhythm), the rhythm of 28-35 years (the brückner cycles), as well as others, are actually cosmic rhythms which should be considered as inherent features, not only of the entire planet but also of the urals in particular. the duration of every rhythm gets longer southwards; the rhythms of 28-35 advances in systems science and application (2015) vol.15 no.3 271 and 80-90 years are no exception. in the north, the duration of the brckner cycle in the flow of the rivers in the mountain taiga areas is 28 years (7%-14% of the total dispersion of the characteristic), while in the south it increases up to 34-35 years. fig. 1 schematic map of the major rhythms in the long-term oscillations of the river flow in the basin of the tobol river. as for the 80-to-90-year rhythm, it increases up 90 (3 x 30) years southwards. the difference integral curves also prove the presence of the abovementioned rhythms in the flow of the rivers in the urals. the difference integral curves, an example of which is represented on fig. 2 below, show that there are rhythm 272 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... elements with durations of about 35 and 80 years. as it was said above, those rhythms are of global character and are typical of the entire geosphere. fig. 2 difference integral curves of the long-term oscillations of the river flow in the basin of the tobol river. rhythms of 30 to 35 years have been found in the geological deposits, in the growth of the annual rings, which is related to the humidification dynamics, and in the spectrum of the air haziness caused by volcanic eruptions. in the hydrosphere, the 30-year periods have been found in the oscillations of the water level of the caspian sea. at the same time, we have to mention that the long-term river flow osciallations are inert to the cosmoand geophysical factors. the delays of the long-term flow oscillations are caused by the peaks of the solar activity and the large-scale processes taking place in the atmosphere and hydrosphere. as for the modern climate fluctuations, the authors have determined three reasons for the climate fluctuations occurring in the southern urals this years: (1)the prevailing longitudinal atmospheric circulation type in the zauralie area is c; (2)the 35-year rhythmic oscillations (the brückner cycles) that depend on the cosmic rhythms; (3)rossby waves. 4 discussion 1. the longitudinal air mass transport in the zauralie area is the one that prevails, for the ural mountains are located longitudinally and the west siberian plain is open to the invention of the arctic air masses, and it takes them about 6 hours to reach the southern urals. if the process is directed southwards, it causes sharp temperature falls, and if it goes northwards, then in wintertime it gives rise advances in systems science and application (2015) vol.15 no.3 273 to warm air masses with large amount of precipitations. in the flat land it leads to thaws and snowdrifts and to heavy snowfalls, avalanches and mudflows in the mountains. lately, the frequency of the longitudinal processes has been growing, which leads to weather anomalies: extremely high and low air temperatures, heavy rains and snowfalls, and longer droughty periods. as for the territory of the region, in summertime low air pressure prevails there. arctic air masses come there from the barents sea and kara sea, while tropical air masses travel from the south, from kazakhstan and central asia. when continental tropical air comes, hot and dry weather sets in. western winds from the atlantic ocean bring humid and changeable weather. representative for this region is the period comprising several a.c. epochs. to explore the effect produced by the cosmogeophysical factors on the climate fluctuations in the region, we used the values of the water discharge and river level variations. a period when the longitudinal circulation prevails (c epoch) was chosen as the period of the study. the research was carried out using the cluster analysis. the analysis results were represented in the form of a dendrogram (fig. 3). the optimal connectivity level was taken as 0.3 (according to the connectivity criterion [4]), where 8 clusters are determined. the first cluster at the level of 0.3 is formed by the rivers of the forest and forest-steppe zones of the territory that was explored, with the probability of the connection between the objects (rivers) of 85% to 99%. the second cluster consists of two sub-clusters combined at the level of 0.3. the first sub-cluster is represented by a group of rivers of the urals forest-steppe and mountain taiga zones situated at approximately the same latitude. they become combined at the level of 0.7, with the connection probability of 94% to 97%. the second sub-cluster is represented by the rivers of the zauralie and preduralie areas (forest and forest-steppe zones; r=0.8, p=99,9%), which proves the synchronism of the long-term oscillations of the river flows in the basins of the rivers kama and tobol during the c epoch. at the level of 0.6, this sub-cluster is joined by the index of the water temperature anomalies in the north atlantic (smed square d, the southwestern part of the north atlantic), which proves the interrelation of the processes occurring in the hydrosphere and atmosphere. and with the probability of the connection between the objects of 88% to 93%, we can suppose that there are also some other factors producing their effects on the long-term river flow oscillations. the next cluster also consists of two sub-clusters. the first sub-cluster comprises three factors: tide-generating force index (at φ=60◦). wolf numbers and the index of the numbers of cyclones in the zauralie area (z8) (r=0.7, p=94% 95%). their combination proves the close interconnection between the cosmic 274 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... factors and the processes occurring in the atmosphere and hydrosphere. in the second sub-cluster, we observe a correlation between the tide-generating force index (at φ=56◦) and the air aridity index at (r=0.7, p=99%), which proves that the abovementioned conclusion is right. the sub-clusters get combined at the level of 0.3. the connection probabilities between the objects in the sub-clusters exceed 95%, except for the connection between the indices z8 and dm (p=83%) and between z8 and f19 (at φ=56◦) (p=85%). fig. 3 dendrogram of the connection between the river flow and the cosmoand geophysical factors during the c epoch the c epoch also reveals a considerable connection between the intensity indices of the ocean currents, the soi and nao indices (r=0.6, p=97%). at the level of 0.6, we see the indices of the number of cyclones (z4) at etp (the preduralie area) combine with those of the process recurrence of the w form (p=85%). at the level of 0.3, they get reunited with the soi and nao indices. the connection probability between the soi and nao indices and the recurrence index of the western form of the atmospheric circulation is quite considerable (94% 95%), which proves the fact of their close interconnection. but the accuracy of the connection between the soi and nao indices and z4 is not high (p<80%). the recurrence value of the eastern form processes of the atmospheric circulation (e) correlates with the index of the water temperature anomalies in the northeastern part of the north atlantic, which proves yet again that the there is advances in systems science and application (2015) vol.15 no.3 275 a connection between the processes in the atmosphere and hydrosphere. in the c epoch, the recurrence of the processes of this form depends on the position of the longitudinal center of the icelandic low (λicl.low), which is proved by the combination of those indices at the level of 0.5 with the connection probability of 94%. the latitudinal coordinate of the icelandic low (φicl.low) correlates with the longitudinal circulation index in the north atlantic and the troposphereeffective index of the solar activity (the connection probabilities of 93% and 97% relatively). this allows representing the interrelation chain as follows: solar activity →...→ sea level oscillations →...→ large-scle atmospheric processes →...→ annual river flow moreover, in the dendrograms of the other periods under study, we found stable and reliable connections within the cosmoand geophysical factors themselves and also between them and the long-term oscillations of the annual river flow: (1)the index of the number of cyclones in the zauralie area (z8) correlates with the solar activity (w); (2)the latitudinal value of the center of the icelandic low (φicl.low) correlates with the nao index and the recurrence of the w form processes; (3)the recurrence of the c form processes correlates with the long-term changeability of the annual river flow of all the natural zones, except for the rivers situated in the area of the stovas critical parallels (59◦ 62◦), where the major influence is produced by the w form. hence, according to the results of the studies, the chain of the interrelations between the cosmoand geophysical factors and the long-term river flow oscillations, represented in the most detailed way, looks as follows: solar activity → large-scale atmospheric processes → ocean currents → number of cyclones (z) and anticyclones (az) → air humidification (aridity indices) → annual river flow. 2. along with the average climate conditions and the abovementioned climatic cycles, the peculiar thing about the climatic regime is its variations. the modern period is characterized by a global temperature trend. it is subject to interannual and decadal (with characteristic rhythmicity of about a decade) variations. if we consider the period of instrumental observations only, we can state that finding the planetary average appears more reasonable. but if we excluded, from our consideration, the 19th century instrumental observations as the least reliable, the modern 100-year trend would totally disappear. instead, we could only speak of the temperature growth that started in the 1980s and in the first decade of the third millennium and consider it an example of the positive temperature 276 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... anomaly of the 1940s warming type. at the same time, it would be reasonable to distinguish changeability from changes. changeability is understood as variations around some average condition, but changes is what characterizes a trend, a transition to a brand new state, with the changes occurring both to the average regime itself and the changeability condition. this differentiation is methodologically convenient, but the physical value of this approach is doubtful, for a lot of things depend on the time scale taken into consideration: the same phenomenon may act both as changeability and changes if viewed from different time scales. as appears on fig. 1, the major rhythms in the zauralie area are those of about 13.8, 19 and 35 years. in 2014, we could observe the 35-year brckner cycle (it may vary from 28 to 35 years; it is closer to 35 years in the zauralie area). the authors have determined this rhythm as one of the major ones. a similar abnormally cold summer was observed in the southern urals in 1980. according to the scientists forecast, the next anomalous summer is expected in 2050, and climatologists even promise us a minor ice age then. hence, the rhythms of nature represent a result of the joint effect produced on it by the space and geophysical factors. 3. not so long ago, scientists found a correlation between the anomalous weather phenomena[5]: the hardest frosts[6] in the past 20 years in the usa (50 below zero c◦) and the thaws at the beginning of the year in russia. the reason for the climate changeability on the opposite sides of the earth is the rossby waves. they may occur in different seasons (k. nyullis, press secretary of the world meteorological organization). last summer, during a two-month period, there was a blocking anticyclone “hanging” over russia. such blocking anticyclones are typical representatives of a natural phenomenon called rossby solitons (spatial solitary waves). rossby waves are air waves which are generated in the atmosphere at moderate altitudes and are moving from the east towards the west. they participate in the formation of cyclones and anticyclones. the atmospheric rossby waves appear in the atmospheres of planets as a result of the shift in the vortices due to the effect of the coriolis forces during latitudinal variations of the area. such waves were first determined in the earths atmosphere in 1939 by carl-gustaf arvid rossby. such formations are not unique; they are found at different levels of the organization of matter, and in particular, the rossby soliton is exactly what determines the spiral structure of the galaxies. the most striking and best-known rossby soliton example is jupiters great red spot which has been observed for more than 300 years [7]. in theory, these waves had been known since the late 19th century. back in 1893, the academy of sciences of vienna published an article by max margules, where the author stated that in a rotating liquid or gaseous environment there advances in systems science and application (2015) vol.15 no.3 277 appear drift waves caused by the global rotation of the environment where they are developing. a little later, in the works by the german geophysicist and oceanologist bernhard haurwitz, those waves were considered as an additional element of the earths ocean. but it was in 1939 when the swedish geophysicist carl-gustaf rossby became the first to draw special attention to the extremely important role played by those waves in the global atmospheric circulation. since then those drift waves, as well as the global vortices they cause, have been named after him. the oceanic rossby waves cause fluctuations of the height of the sea surface. it had been difficult to determine the wave lengths before satellite altimetry was invented. the observations made with the nasa/snes topex/poseidon satellite proved the existence of the oceanic rossby waves. the satellite observations showed the majestic progress of rossby waves in every ocean basin, at low and middle latitudes in particular. the waves may live for months or even years, so that they cross the pacific ocean. rossby waves have also been considered as an important mechanism to monitor the warming of europas ocean. most probably, rossby waves also participate in the generation of magnetic fields in nature and are capable of causing local fluctuations of the condition of the lower troposphere. 5 conclusions according to our studies, the following phenomena accumulated in the zauralie area in summer 2014: a) recurrence of the brückner cycle with the duration of 34 years (1980 2014); b) intensification of the longitudinal shift that prevails in the zauralie area and is caused by el niño (rossby waves are the “conductor” of el niño); c) the shift being blocked in the zauralie area by a rossby wave; d) “hanging” of the weather conditions caused by el niño [8]. thus, the rhythms observed in nature are the results of the effect produced on it by the cosmogeophysical factors. in particular, “interference” in this interaction between the geophysical factors (rossby wave) and celestial bodies in the solar system and outside it is able to “disturb” that interaction and lead both to local (zauralie) and global climate fluctuation, which has been observed more than once in the history of our planet. references [1] gumilev l.n. (1992), ethnogenesis and the biosphere of earth, oscow. [2] maksimov e.v.. (1995) rhythms on earth and in space, publishing house of saint petersburg state university. [3] rasskazova n.s. (2013), “rhythmic changeability of the regime characteristics of water bodies as the integral index of the climate fluctuation (based on 278 nadezhda s. rasskazova, aleksandr v. bobylev, michail n. bubin and alexander... the urals example)”, climate change and industrial city ecology, pp. 49-51. http://www.forum-ecology.ru/content/forum/catalog 2013/194 catalog ef 2013.pdf [4] bubin m.n. and rasskazova n.s. (2013), “rhythmicity of the long-term river flow oscillations as the integral index of the climate fluctuation (based on the urals example) ”(monograph by m.n. bubin, n.s. rasskazova of the institute of technologies in yurga), tomsk nrpu publishing house, 278 pages. [5] (2013), access mode : “the el nino phenomenon” [electronic resource]. http://nature.web.ru/db/msg.html?mid=1158162 [6] (2014), access mode: “anomalous frosts” [electronic resource]. http://okoplanet.su/pogoda/pogodaday/225659-ssha-anomalnye-holoda.html [7] rossby soliton, (2010), access mode:“jupiters great red spot and anomalous summer in russia” [electronic resource]. http://okoplanet.su/pogoda/listpogoda/48596-soliton-rossbi-bolshoe-krasnoe-pyatnoyupitera-i.html [8] rasskazova n.s., (2009), “el nino-southern oscillation signal depression and its possible reasons”,science & technology. theses of the reports of the 29th russian school dedicated to the academician v.p. makeevs 85th anniversary, miass, pp. 83. corresponding author nadezhda s. rasskazova can be contacted at: yal05@mail.ru advances in systems science and application (2016) vol.16 no.2 81-93 an artificial volition architecture for autonomous robotics t. e. raptis1 and k. karamanos2 1division of applied technologies, national centre for science and research “demokritos”, athens, 15341 greece. 2physics department, university of athens, zografou, gr-15784 abstract we introduce a new computational architecture capable of exhibiting an archetypal type of volition in a simplified modularized version of m. minskys hive-mind. in this model, three relatively independent computational cores which themselves can also be whole multi-agent systems are engaged in an endless interaction each one representing the internal “imaginative” world, the external world interface and the arbitrator or internal observer. volition then is expected to occur as the result of an endless antagonism for control between the internal and the external world models. keywords a.i., robotics, volition, free will. 1 introduction history of a. i. is marked with a division between two main abstractions, the one of “connectionists” that try to imitate directly the real brain functions closer to the spirit of old wieners cybernetics paradigm and “symbolists” who advocate the old belief that mind is algorithmic, first introduced by alan turing. both schools have had various successes in diverse fields but when it comes to the inner psychological or subjective experiences they face a major obstacle which is the correct objective definition and experimental verification of such abstract concepts as volition (exercise of will), consciousness and/or self-awareness. these problems also arose in the early phase of development of cognitive sciences and still constitute a major set of debatable issues at the heart of this research. questions concerning the possibility of externally verifying ones own awareness of a self and a personal identity have given rise to the argument of “philosophical zombies” against functionalist interpretations of the identity and awareness problem.although the particular architecture proposed seems to be unique, at least to our knowledge, there are several parts that may be comprised into the proposed structure and have already been proposed as separate items and implemented into various robotic platforms. a crucial property for self-aware agents is that of self-observation which requires a constant scan of both the history of the external agent behavior and the internal loop activity and their correlation. recent work from the cornell group[1-3] has revealed the possibility of an advanced reverse engineering of 82 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics complex environmental signals and accurate modeling of external reality. such models can then be contrasted with the internal personality model and affect future planning. furthermore, the agent can effectively randomize its own behavior to avoid stereotypical pre-programmed behavior by an internal mechanism which optimizes the antagonistic need to preserve the personality model and the environmental pressures by an adaptation protocol which affects the probabilities/weights of certain actions. attempts to introduce a “self-model” compatible with the idea of an “inner world” have been superficially carried out based on a 3 -dimensional mental space together with the so called “big five model” of psychology. previous work by the waseda group[4-6] on a 3-d mental space has shown promising results in control of a robotic face (emotional agent). this way we can build “idiosyncratic agents” with a certain predefined set of preferences. a second important step in the construction of intelligent agents is the previous work of luc steels on “fluid construction grammars” which led to the development of the talkingheads project. in this, a number of bots were sustained inside the internet moving in a network of host computers which were equipped with steerable cameras. they were then able to analyze optical signals and classify objects while developing a primitive linguistic structure through which they were communicating their experience into one another. the same model can also be utilized in an advanced volitional agent who would become capable of classifying both internal and external signals and thus reaching at level 3.1 or 3.3 of tables 1(a) (b) of the next section. it could also be used for an agent applying an artificial linguistic structure to purely internal variables. the present proposal attempts to provide an example of an advanced architecture that could encompass all previous developments and unify their separate approaches in a unique frame thus extending their capabilities towards an advanced agent design with truly inherent volitional attitude that would arise as a result of the internal dynamics of its major components. the presentation is based on a high level description of the whole architecture ignoring the details of the software implementation which in principle could be numerous. it is intended in giving a correct understanding of the foundations of an alternative volitional theory that could possibly be applicable in other fields as cognitive sciences and human psychology. in section 2, we lay the foundations of our model while in section 3 we describe their possible implementation in more detail using a top-down approach. in section 4 we conclude and also discuss the significance of learning as an additional concept of a higher level that was not absolutely essential in the previous development. advances in systems science and application (2016) vol.16 no.2 83 2 foundations of artificial volition adopting a practical approach, we choose to concentrate in the necessary and sufficient conditions for a behavioral evaluation of the existence of volitional attributes of an agent. present state-of-the -art in cognitive sciences allows one to write a generic test that can be applied to an arbitrary set of agents in order to discriminate between several levels of awareness as proposed in [1]. this is summarized in the table 1(a) below. table 1 (a) the awareness level awareness level discriminating question categorization 0 is it animate? alive/dead 0.1 “does the system move or act on its own, i.e., without obvious prompting by external forces?” autonomous/nonautonomous 0.2 “is the systems spontaneous behavior modified by events/conditions in the environment?” 1 ”does the system appear to be trying to approach or avoid any object or occurrence of an event in its environment” 1.1 ”does the system have different sets of goals active during different environmental or bodily conditions?” modal value-driven automaton/“pacman ghost” 2 “does the system develop new adaptive approach or avoidance patterns over time?” 2.1 “can the system engage in a task that requires working memory (e.g. delayed non-match-to-sample)?” 2.2 “can the system engage in a task that requires longterm memory?” 2.3 “can the system engage in a behavior(e.g. gameplaying, navigation) that requires evaluation of multiple possibilities without action?” most animals 3 “does the organism send and selectively respond to social cues? ” 3.1 “can the agent pick up and move around objects in its environment?” 3.2 “does the system communicate using language that has syntax as well as semantics?” chimpanzee some observations are due at this point. nowadays, robotic arms and even autonomous robots exist that are capable of fulfilling 3.1 although they are no more than modal automata thus we should move this question into level 1. secondly, the overall description of level 3 seems incomplete in that it does not explicitly shows true volitional acts in the absence of competition with a truly intelligent environment which includes also other intelligent agents either artificial or human. we thus propose the following modified table 1(b) for level 3.0. 84 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics table 2 (b) the modified awareness level 3 “does the organism send and selectively respond to social cues?” 3.1 “does the agent attempts to escape captivity enforced by an external agent or situation?” 3.2 “is the agent capable of enforcing a certain task to other less intelligent agents?” 3.3 “does the system communicate using language that has syntax as well as semantics?” we can now proceed to a separate examination of the three different levels. in fact we propose that the above hierarchy is not strictly necessary and that certain properties may be intermixed at least in artificial agents. that is to say that we can in principle separate between different architectures or implementations of systems that could emulate several of the characteristics pertaining at different levels without strictly obeying the above hierarchy. in particular, we would like to make a crucial separation between consciousness as a purely subjective state and volition as a more primordial level necessary for its existence. in fact, it is not in principle possible to directly assert the existence of an internal “feeling” by the agent.it is only possible to assert the resistive actions taken by the agent against external obstacles or environmental “laws” that are not preprogrammed. next we concentrate on a mechanism capable of exhibiting volitional effects at the level of 3.1 without necessarily reproducing all of the characteristics of the previous levels. specifically, we seek for the minimal behavioral test that could verify the ability of an agent to exhibit an element of “free will” in the sense of a) either randomizing its own behavior in order to cope with a contradiction or a conflict between its own self model and its obtained world model, or b) undertaking evasive actions against an enforced captivity or restrain by another agent. we thus concentrate on the fact that the presence of any kind of “free will” is definitely asserted only in an antagonistic situation where the agents will is exerted against another agent as a resistive force. this assumption does not exclude a cooperative behavior which would only occur under a state of agreement or symbiosis between different agents. to explain the significance of the levels 3.1 and 3.2 we may also add a quote from a dialogue between c. zvosil and d. greenberger[8] where they criticize the incomplete nature of turings test with respect to free will. “· · · assume an artificial world and/or an artificial creature. assume further a super selection rule which should not be broken by any circumstances; eg eating from the tree of genesis. this creature should be termed intelligent if it breaks its super selection rule. in this approach it is evident that uncontrollability is a trade-off for intelligence and that it might be impossible to create a machine which is both intelligent and a advances in systems science and application (2016) vol.16 no.2 85 reliable server.” we choose to call this principle, a “non serviam” principle from the latin expression for denial of service. we will next show that it is possible to extract from the above general argumentation a full computing architecture that can fulfill the above principle. the significance of such a construct is that a) if it is possible than it will probably be realized in the future yielding new problems for security and reliability of services (friendly ai problem) and b) it seems to be more closely associated with human behavior than other mechanistic approaches. at first we will use a metaphorical example to lay down the description of the principles of the model. the basic obstacle in present state-of-the art is that the mind is still represented much like an operating system which waits for input from the external environment and then reacts to it according to a prescribed set of functions. what is clearly missing from all these approaches is the notion of a purely internal world with autonomous life independent of any external activity. such is the case of imagination but also of dreaming. the only way we can achieve a machine with a “dreaming state” is to endow such a construct with an internal autonomous and endless loop of which the variables are equally sensed by an “internal observer” (io) as well as the external response signals albeit not always on an equal footing. we thus present this general idea with the following schematic. in fig.1, the three modules should be interpreted as relatively independent computational engines of different characteristics and purpose. while the inner world (iw) is represented by an autonomous activity, like a huge dynamical system, the io module is more like an operating system which needs to be fed with inputs from both subsystems having the environmental variables fed by sensory systems and also form the interface of the machine with the outside world. each module may act as an independent agent in a continuous dialogue with the other two thus forming a tri-dialogical system. the dynamics envisioned behind this scheme can be described with the aid of fig.2. the io module acts like an independent agent “trapped” between the activities of the iw core which is fed to it through appropriate internal sensory inputs and those provided by the external interface. the io module then attempts an endless evaluation and arbitration task between the conflicting activities in order to control the appropriate responses that are also imposed by the third survival task (st) module which plays the role of an autonomic nervous system. in what follows we give precise meaning and some possible technical methods to implement the above logic. we of course assume the existence of an appropriate environmental interface for both input and response signals (eg. a full robotic body). 86 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics fig. 1 a tri-dialogical system fig. 2 a dynamic response process 1. the iw-module is a specially constructed closed computational loop which is to have access to some or all the external inputs not for processing but in the form of direct perturbations of its own internal variables. it is also prewired in a certain way with the io module in a way that allows for alterations of its internal structure -not just its variables!-imposed by the io module. (practically this means that the io module may alter even the form of the “equations of motion” of the internal dynamics). this is supposedly an endless internal activity. at advances in systems science and application (2016) vol.16 no.2 87 this point one may ask how we can guarantee that this will indeed be an endless computation without solving the halting problem! this will be further clarified into the next paragraph that fully describes the dynamics with the environment mediated by the io activity. at the moment it is important to accept the iw module as an isolated “dreaming” machine or a kind of “subconscious” in the overall architecture. 2. the st-interface module contains all such attributes that represent predetermined protocols for the survival of the agent in an arbitrary environment including input signal processing and automatic response circuitry. it is thus close to a model of the autonomic nervous system. 3. the io is by necessity the part of the machine responsible for the categorization and evaluation of the activities of both the internal and external environmental variables. it thus holds and updates a primitive self-model made by the subsequent observations of each ones activity. on the other hand, it also plays the role of an arbitrator between possible conflicting demands of the two separate iw and st agents in case of conflicting demands posed to the external response modules. in order for this scheme to exhibit full volitional attributes there is one more crucial step in the definition of the io module. we thus propose to introduce an “egotistic” attitude to the agent in the following sense. the fundamental property of “ego” is not to be just a self-model in the form of a data structure but also to be “demanding”. this is also evident from early childhood psychology. the only way we can give a precise technical meaning to this proposition is to introduce a certain kind of internal hyper-tasks attributed to the io module beyond the simple self-identification task. we thus attempt to derive such a hyper-task from a generalized functional control problem. we assume that the simplest hyper-task is the attempt by the machine to create a “higher self” model into which all or as many as possible of the external variables have been “internalized” or in fact “enslaved”. this is to be understood here in a somewhat more broad manner than a “master-slave” configuration is often realized in engineering. this is quite a broad concept that can extend even to systems with continuous variables. we interpret this internalization process as follows. assume a composite system, e.g. a neural network, which has several processing nodes and several other input nodes (fig.3). the processing nodes are to be considered as “immediately active” in the sense that they carry a certain processing capacity while the sensory nodes form an interface for some in principle unknown external nodes that represent the activity of the environment. thus all nodes f1, f2, · · · , fn represented by circles are “known” to the system while the nodes e1, e2, · · · , em represented by red squares may not even have a functional expression due to their extreme complexity. in a sense, although the agent may 88 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics possess a primitive notion of identity it still considers the rest of its world as an extension that can be manipulated. the systems dynamics executes the hypertask of trying to minimize a generalized distance between the functional forms of the external world source nodes and the internal nodes. thus the system should start to transform the external world continuously until it would be able to bring all or most of the external variables to a state in which they conform to the internal world dynamics. while this in actual reality is an open and endless task (considering the rest of the universe as the “environment”!) it is still an accurate representation of the origin of the agents volition. in fact, we propose that this is the only mathematically acceptable expression of an agents “free will”. an interesting consequence of the above type of dynamics that we can predict is that cooperation with such an agent would be possible through “enticement” which presupposes a degree of access or knowledge to its own internal variables that make up its “character” or self-model without sacrificing an amount of fluidity in it. conclusively, we may end this section with the following propositions. fig. 3 the internalization process of artificial volitional attributes proposition 1. a sufficient condition for the appearance of volitional attributes in a network of parallel turing machines is the existence of conflicting dynamics between computational tasks of different purpose that result in the formation of a master-slave configuration inside the network. proposition 2. a necessary condition for ascertaining the volitional attribute of an agent is the testability of its capacity to interpret restrictions as such even in the absence of any knowledge about their origin (“suspicious agent”) and seek more freedom in the transformation of a given environment. 3 possible implementation and dynamics a more concise and detailed description of the system units, their interconnection and dynamics follows. the design philosophy explained in the previous section advances in systems science and application (2016) vol.16 no.2 89 is further elaborated on the technical details of a possible implementation. design of the “inner world” model can be better described with the aid of the analog computing paradigm which is here simplified on a system of differential equations. these can be seen as an average over a multitude of microscopic degrees of freedom that could be realized by neuromorphic circuitry. we believe this to be true for natural biological circuits although we do not make strict use of neural implementations as this is unimportant to our purpose. let then x1,· · · ,xnin and x1,· · · ,xnout be a set of internal and external variables defined as follows : the internal variables form the basis of the inner loop dynamics while they are coupled to the set of external variables. the external variables are coupled to environmental signals and they originate at the st module where a first processing of the system inputs takes place. system response is due to both the internal dynamics of the st module as well as the supervising inputs from the io module which mediates between the iw module and the st module. the inner loop attempts to bring the set of external variables under its own control. this of course is an always incomplete task in the sense that the external variables obey an in principle unknown, non-stationary dynamics which may also have different underlying laws from a previous instant to the next. the philosophy behind this control scenario is that the inner loop attempts not just to identify the external variables dynamics but it tries to “enslave” them in order to follow its own internal differential equations. that is to say, the inner loop attempts to internalize the external world. it is possible to borrow from nature the additional well known biological principle of antagonistic signals. for this, we would have to define pairs of antagonistic variables. one can utilize for example the well known model of competitive lotka-volterra equations to simulate the inner loop dynamics. the choice of the system is somewhat arbitrary and has been chosen for having stable attractors. one could in principle try other choices or a direct neural implementation. even a spiking network could have been used but at the cost of an increased complexity. in fact, we may assume the continuous version of a pde system or equivalently the corresponding pseudo-spectral kernel by which the iw module builds the subsequent configurations of its own internal state. assuming an appropriate sampler interface between the io and the iw modules, the io “sees” snapshots of the iw activity the same way it sees the external world through the additional interface of the st module. moreover, the inner loop module should be coupled to a set of supervising signals from the io module that affect a matrix of coefficients for the ode/pde system that defines the iw loop dynamics. the inner loop dynamics may be made to obey a 1st order hyper-task. this 90 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics is described by an internal update process of the system coefficient matrix or of the corresponding kernel parameters towards the production of more coherent and symmetric patterns of activity (artistic agent). this is also influenced by the perturbations from the external variables so that an antagonistic dynamics between the inner and the outer world models is established. the criteria of what constitutes a like pattern are prewired into the particular agent and form an integral part of its “character”. (martian spiders for example may have a very peculiar idea of what constitutes a beautiful face!) the io module is primarily responsible for a 2nd order hyper-task of reducing the functional distance between the iw and the outer world as it appears to the interface variables provided by the st module. this 2nd order hyper-task is eventually linked to the 1st order one through the attempt to bring the environmental dynamics close to fulfilling the directive of the iw module of producing more symmetric and coherent patterns. in principle there is no limit in the hierarchy of hyper-tasks that could be implemented and could also be said to stand for the agents “talents” in a metaphorical sense. the io module should contain a reverse engineering engine which separately examines time-series of both the internal and the external variables and their history which are recorded into a dedicated part of memory (long term memory). it should then extract mathematical models of both and attempt a possible match by proposing certain changes into both the internal parameters and the external assumed environmental parameters. the supervising signals towards the iw module can then be directly implemented into the underlying system of differential equations while the supervising signals towards the st module must be appropriately translated into actions through an interpreter that extracts from them the discrete signals towards the drivers of the external devices in a manner that can best allow the modification of the assumed external dynamical functions. such a modification of the external world may also obey certain directives like the overall entropy reduction in the surrounding space. other scenarios can also be tested with more complicated directives. in fact, the io module can be updated to include a more general planner under certain learning strategies. for the moment we may ignore this advanced capability as it is not inside our objectives to test learning strategies of which are many in the trade. the io module is in addition responsible for continuously building and refining a self model by projecting the present state of match or distance between the iw and the external world model into the mental space representation which may affect the decision making and subsequent planning. in total, the io module is responsible for all the abstract representations of both the self and the external world. the io module can also incorporate a fcg engine for artificial linguistics in order to fulfill the final level of self-expression. advances in systems science and application (2016) vol.16 no.2 91 the st module must contain certain instructions for preprogrammed tasks like step generators for locomotion, or other necessary movements, auto-charging in case of a power source present, “food” hunting if abscent etc. in general it may be characterized as the equivalent of the autonomic nervous system. on the other hand, the st module must interpret the supervising signals from the io module in order to feed appropriately the drivers of the external peripheral devices (legs, arms, sensors) so that the combined motion will affect the environment in a way fit to the higher task of enforcing the desired structure to the external world dynamics. building the appropriate interpreter for the drivers is a non-trivial task as it requires also an amount of computational geometry in order to take into account the exact actual details of the environment and the robotic devices and the possible ways of their interaction. at this level, it is necessary to include a system of somatosensory perception which entails the robot with the capability of describing itself in both shape and movement inside a certain environment. 4 discussion and conclusions the above general scheme is here presented as having certain advantages over other existing approaches due to its holistic nature that attempts to incorporate the most fundamental elements of what human beings intuitively know about themselves. to our opinion these advantages include. • deeper understanding of the foundational principles behind awareness in artificial and natural systems. • enhanced capabilities of operation of autonomous systems in real time. • practical approach to the problem of measuring various forms of awareness in situ. • prediction of possible malignant applications of true self-aware software or hardware and experimentation with counter-measures (“asimov laws” onr report[9]). • practical investigation of the problem of friendly ai (hal9000 problem). in principle there seems to be no restriction in the kind of hyper-tasks that could be implemented. what has not been examined in detail here due to space restrictions is the role of learning processes in a direct interaction with the io layer. it seems reasonable to assume that a learning machine with the structure described above could also discover its own set of hyper-tasks or modify previously programmed ones unless this is strictly forbidden via some dedicated censorship circuitry (freudian viewpoint). with respect to the last two points above it deserves to mention that there 92 t. e. raptis and k. karamanos:an artificial volition architecture for autonomous robotics exists indeed a possibility that a learning machine could also have been imprinted or even develop on its own! a “seek and destroy” attitude. in fact, it is quite possible that a “skynet” scenario is in principle technically feasible and could become reality in the next 20 years or so as already predicted by de garis[10] and kurzweil[11]. the architecture presented above to our opinion justifies such fears and shows the necessity towards more research in the direction of friendly ai as well as the need to ask for demilitarization of robotics. acknowledgements: the first author would like to express his gratitude to the members of the computational applications group of d.a.t.-ncsrd for helpful discussions and support in the preparation of this document. references [1] bongard j. and lipson h. (2005) “active coevolutionary learning of deterministic finite automata”, journal of machine learning research, 6(10):1651-1678. [2] bongard j., zykov v. and lipson h. (2006), “resilient machines through continuous self-modeling”, science, 314(5802): 1118-1121. [3] bongard j. and lipson h. (2007), “automated reverse engineering of nonlinear dynamical systems”, proceedings of the national academy of science, vol. 104, no. 24, pp. 9943-9948. [4] miwa, h., umetsu t., takanishi a. and takanobu h. (2001), “robot personality based on the equations of emotion defined in the 3d mental space”, proceedings of the 2001 ieee international conference on robotics and automation, pp.2602-2607. [5] miwa, h., umetsu t., takanishi a. and takanobu h. (2001), “human-like robot head that has personality based on equations of emotion”, preprints of the sixth symposium on theory of machines and mechanisms,pp.1-8. [6] miwa, h., umetsu t., takanishi a. and takanobu h. (2000), “robot personalization based on the mental dynamics”, proceedings of the 2000 ieee/rsj international conference on intelligent robots and systems,pp.8-14. [7] chadderdon and g. l. (2008), “assessing machine volition: an ordinal scale for rating artificial and natural systems”, adaptive behavior 16, pp.246-263. [8] svozil, k. (1993), randomness and undecidability in physics, world scientific. advances in systems science and application (2016) vol.16 no.2 93 [9] lin p., bekey g. and abney k. (2008), autonomous military robotics: risk, ethics, and design, ethics & emerging technologies group, california polytechnic state university. [10] de garis h. (2005), the artilect war, etc publication. [11] kurzweil r. (1990), the age of intelligent machines, mit press. corresponding author t. e. raptis can be contacted at: rtheo@dat.demokritos.gr advances in systems science and applications (2014) vol.14 no.2 170-182 air materiel supply intelligent collaborative decision-making mode based on ontology and multi-agent xiong li1, chang-xin liu1, fang bing1 and xiongyi li2 1school of management, shanghai university, china 2graziadio school of business and management, pepperdine university, usa abstract to solve the complexity problem of each node and the diversity problem of decision-influencing factors in air materiel supply chain, an air materiel supply intelligent collaborative decision-making mode based on ontology theory and multi-agent has been proposed. to better organize the collaborative knowledge utilized by agents and facilitate agents’ adaptive collaborative decision-making ability, an ontology-based approach is presented in this paper. the knowledge is separated into shared ontology and private ontology to ensure both the agent communicative interoperability and the privacy of strategic knowledge. then, the collaborative mode and action planning of agents are analyzed, and the architecture of intelligent collaborative decision-making system of two-echelon air materiel supply chain has been designed. thus a platform for the consultations and coordination of each agent has been provided, and an effective decisionmaking method has been proposed to decision-makers of air materiel supply. keywords intelligent collaborative decision, ontology, multi-agent, action planning 1 introduction the decision-making mode of the air materiel supplying is actually one of the most important subjects in the field of air materiel management and engineering. in the decision-making process, the airline and air force mainly use the best analytical model available for the aircraft spares provisioning problem, which includes the order, transportation, storage and consumption of air materiel[1-3].in supply chain network, the complexity of each node and the diversity of decisioninfluencing factors are indispensable attributes, however, when making decisions, decision makers only consider a certain part or some key factors in the decision process. usually, the actual decision authority and processes are distributed among the members in the supply chain, who are primarily concerned with optimizing their own objectives. as a result, making appropriate decisions to attain optimal performance in air materiel supply chain is a very challenging problem. traditionally, contractual agreements and complex accounting schemes are used to ensure that the supply-chain works effectively during daily operations. and the centralized or hierarchical decision-making process in these supply chains readvances in systems science and applications (2014) vol.14 no.2 171 sults in losses of efficiency in a competitive and dynamic market environment. members in a supply chain have to coordinate their individual decision-making closely and cooperatively to achieve optimal supply chain performance with balanced services level and inventory level, the each node of supply systems can no longer be viewed in isolation, they must be managed in the context of the total business and the associated key linkages of the business, and the collaboration is an attractive strategy in air materiel supply chain network. due to the complexity of each node and the diversity of decision-influencing factors in the field of air materiel supply, it is worth to research the air materiel supply intelligent collaborative decision-making system to aid managers to make decisions. decision support systems (dss) is the area of the information systems (is) discipline that is focused on supporting and improving managerial decision-making[4]. the history of dss reveals the evolution of a number of sub-groupings of research and practice. the major dss sub-fields include personal decision support systems (pdss), group support systems (gss), organization decision support systems(odss), negotiation support systems (nss), intelligent decision support systems (idss),distributed intelligent decision support systems (didss),knowledge management-based dss (kmdss), distributed intelligent cooperative decision support systems(dicdss) [5-8],etc. with the development of decision support systems, it has been applied in many areas, such as the application to support the temporal and spatial distributed decision-making process in supply chain collaborative planning[9]. a software agent is a program that performs a specific task on behalf of a user, independently or with little guidance, it is characterized with environment awareness, ongoing execution, autonomy, adaptiveness, mobility, intelligence, anthropomorphism, reproduction, independence, collaboration, distribution, etc. the software agent is crucial factor as dss components to build intelligent collaborative decision support systems characterized by cooperating agents, either human or non-human actors[9-11]. ontology is a knowledge representation method with a philosophical concept as the branch of metaphysics which has been widely used in science and technology. from computer specialists perspective, ontology means a vocabulary and a set of terms and relations that define, with the needed accuracy, a set of entities enabling the definition of classes, hierarchies, and other relations among them [12], it has been established as a powerful paradigm to enable knowledge sharing, and became the foundation for many multi-agent system applications as a means to achieve semantic interoperability among heterogeneous agents systems. many intelligent cooperative decision support systems based on ontology and multi-agent has been proposed by researchers[13-15]. for example, chang-shing lee, etc. present an ontology-based intelligent decision support agent to apply to project monitoring and control to reduce the human efforts and 172 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... the costs of the project [13]. in the study, the natural language processing agent, the fuzzy inference agent, the performance decision support agent, the capability maturity model integration ontology and project personal ontology are designed. through the collaborative work of the agents and ontologies, the ontology-based intelligent decision support agent can work effectively for project monitoring and control of capability maturity model integration. compared with the above articles, this paper has the following characteristics. at the first, the collaborative mode of agents based on ontologies has been designed, and according to the collaborative mode, the collaborative decision-making problem in agents located on the different nodes of supply chain network or between managements and agents can be resolved. secondly, the approach of dynamic description logics(ddl) has been applied to planning the action sequences of agents in the instance of air material supply decisions. at last, the structure of air materiel supply intelligent collaborative decision support system based on the collaborative mode of ontologies and agents is designed. the details of this article are as follows. the paper proposes a framework for building decision support systems using software agent technology and ontology to support organizations characterized by physically distributed, enterprise-wide, heterogeneous information systems. in section 2, the structure and interoperability mode of ontology is presented, the multi-agents collaborative mode based on ontology is also analyzed, and a platform for the consultations and coordination of each agent has been provided. section 3 analyzes the action planning of agent. in section 4, the architecture of intelligent collaborative decision-making system of two-echelon air materiel supply chain is designed, and an effective decision-making method is presented to decision-makers of air materiel supply. section5 presents conclusions of this paper. 2 the multi-agents’ collaborative mode based on ontology 2.1 ontology based intelligent agent applications an agent is defined as a software entity that is situated in some environment, and is capable of autonomous action in that environment in order to meet its design objectives[16]. a generic agent has a set of goals, certain capabilities to perform tasks and some knowledge about its environment. to achieve its goals, an agent needs to use its knowledge to reason about its environment and the behaviors of other agents, to generate plans and to execute these plans. a multiagent system(mas) consists of a group of agents, interacting with one another to collectively achieve their goals. by absorbing other agents’ knowledge and capabilities, agents can overcome their inherent bounds of intelligence. air materiel supply intelligent collaborative decision-making system is the intelligent system based on air materiel supply knowledge. in the air materiel supply chain network, advances in systems science and applications (2014) vol.14 no.2 173 managers get consensus about some decision problems through the collaboration of intelligent agents or the intelligent agents and managers. agent technology can provide flexible, distributed, and intelligent solutions for air materiel supply management. however, each agent is an independent entity, one agent needs to exchange the domain knowledge, related concepts with other agents in the internet, and the software agents need achieve the collaborative protocols regulating the set of rules that govern the interaction of agent. ontology as an explicit specification of a conceptualization, and the agreements about the conceptual frameworks for modeling domain knowledge, it is more than just a vocabulary and taxonomy of terms. it provides a set of well-founded constructs that can be leveraged to build meaningful higher level knowledge and relationships between terms[17]. by describing a set of concepts and the relationships between them, ontology can construct both the hierarchical architecture of the knowledge and the descriptive logics of regulations and activities. according to the ontology structure, inference rules can be defined to guide agents collaborative behaviors to adapt to various collaborative environments. as a novel knowledge organization concept, ontology is widely applied in the domain of information sharing, supply collaborative management, multi-agent systems, dsss(decision support systems) and rule-based reasoning systems to enable interoperable decision knowledge structures for knowledge sharing and utilization[18-20]. to effect the intelligent agent applications in the collaborative environment, ontology can help agents in representing and storing domain knowledge, enabling a semantic interoperable environment, reaching mutual understanding, reasoning and querying the knowledge repositories , maintaining a secure system access, etc. in air materiel supply chain network, the members embrace mainly air materiel manufacturers, maintenance contractors, airlines, pooling providers of air materiel, etc. the pooling provider of air materiel is the air materiel warehouse or airline, manufacturer, maintenance contractor in essence which provides service of air materiel supply for the airlines, and the pooling provider charges the airlines in some calculation for the air materiel supply service in accordance with the agreement. in the decision-making analysis process of production, order and storage, generally the air materiel manufacturers, maintenance contractors and pooling providers are considered as the analysis object. due to supply members needing to exchange the concept, domain knowledge and the agents’ related activity specification in the collaborative decision process of air materiel manufacturers, maintenance contractors and pooling providers, two aspects have to be considered. the first is to solve the ontology interoperability problem in agent communication. the second is to build sophisticated private ontology and shared ontology to facilitate collaborative behavior deployment[15]. the private ontology abstracts the knowledge in supply chain collaborative decision which config174 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... ures software agents computing and inference parameters. the shared ontology is developed in accordance with some criteria to ensure mutual understanding between agents developed by different air materiel supply chain members, and it is structured as the types of schemas: concept and agentaction, etc. the structure and interoperability mode of ontology are displayed by fig.1. the contents of shared ontology developed by pooling providers embrace the air materiel category, air materiel name, payment method, order quantity, ordering price, shared price, unit cost of transportation , unit cost of inventory, unit cost of missing parts, agent name, agent address, aid (agent identifier), etc. the contents of private ontology developed by pooling providers comprise order cycle, manufacturer’s credibility, transportation methods, ordering price concessions degree, the way of ordering price concessions, agent negotiation strategies, etc. the contents of shared ontology developed by manufacturers embrace the air materiel category, air materiel name, sale price, delivery time, terms of service, unit operating costs, agent name, agent address, aid, etc. the contents of private ontology developed by manufacturers comprise pooling providers credibility, ordering price concessions degree, the way of ordering price concessions, agent negotiation strategies, etc. fig.1 the structure and interoperability mode of ontology 2.2 the multi-agent collaborative mode in the collaborative decision process of air materiel supply chain members, the members exchange the information and define the software agents collaborative mode through the ontology interoperability. the software agents solve the decision problem by interacting together. there is a set of intelligent agents defined advances in systems science and applications (2014) vol.14 no.2 175 as ordering decision agent, available inventory agent, manufacturer select agent, manufacturer production decisions agent, transport planning agent, etc. in the air materiel supply chain network. when the pooling providers make ordering decision, the ordering decision agent needs to choose adaptive agents from defined agents according to the task, and carry out manufacturer select agent to select possible manufacturers. then the pooling provider sends information of order request to the possible manufacturers, exchanges shared ontology with them. according to the shared ontology exchanged and respective private ontology, the related parameters are calculated collaboratively by the ordering decision agent, available inventory agent and manufacturer production decision agent. at last, the order quantity and order price, etc. are negotiated repeatedly to determine the final order scheme by the ordering decision agent, manufacturer production decision agent and transport planning agent. the multi-agents collaborative mode is displayed by fig.2. fig.2 the multi-agent collaborative mode 3 agents action plan an agent is a tuple, and concrete agents have the same architecture. the components of task agent is defined as follows: w=, act is the act ability component, and represents available behaviours, as well as the situations in which these plans are applicable. its action derive from collaborative mode of agent. beliefs comprise information known by the agent, and regularly updated as a result of perception. desire represents situations that the agent reacts to by adopting plans, corresponding to desired states. intention structures comprise the set of partially instantiated plans currently adopted by the agent[21-23]. when the members of air materiel supply chain make ordering decision, the 176 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... ordering decision agent embrace the order information of a certain air materiel. the content of its belief as follows: belief = {av i(avi), equal(avi, “agent5“), prc(prc), equal(prc, “agent7“), mpd(mpd), equal(mpd, “agent8“), m(m), equal(m, “materiel6“), p (pj), equal(pj , “producer − 1, p roducer − 2, ..., p roducer − n“), ma1(m, a1), equal(a1, v1), ..., man(m, an), equal(an, vn), ..., pb1(p1, b11), equal(b11, x1), ..., pbn(pj , bjn), equal(bjn, xn), ...} concept avi represents the collection of available inventory agent. concept prc represents the collection of manufacturer select agent. concept mpd represents the collection of manufacturer production decisions agent. concept tp represents the collection of transport planning agent. concept m represents the collection of air materiel category. concept p represents the collection of manufacturer.mak(m,ak) represents the kth attribute of an item air materiel named m is ak . relation of equal(a,vk) represents that the value of a is vk.pbk(pj ,bjk) represents the kth attribute of a manufacturer named pj is bjk. relation of equal (bjk, xk) represents that the value of bjk is xk. the collection of belief represents that it knows an available inventory agent named agent5, an manufacturer select agent named agent7, an manufacturer production decision agent named agent8, an item air materiel named materiel6 and n attributes of the air materiel, a set of manufacturers named producer-j and n attributes of every manufacturer. the attributes of the air materiel may be the index of price, performance specifications and transport conditions, etc. in the decision process, ordering decision agent receives the ordering instructions from upper layer agent to formulate an order scheme of an air materiel. according to collaborative mode of agent, the ordering decision agent sets the desire as follows: {ai(m, q0) ∧hr (q0) ∧ pc (m, p) ∧hr (p) ∧oq (m, q) ∧hr (q) ∧op (m,ω)∧ hr (ω) ∧ap (m,ϕ) ∧hr(ϕ) ∧ ep (m, e) ∧hr (x)} . relation of ai(m,q0) represents the available inventory of air materiel named m is q0. relation of pc(m, p) represents manufacturer of air materiel named m is p. relation of oq(m, q) represents order quantity of air materiel named m is q. relation op(m, ω) represents the collaboration price of air materiel named m advances in systems science and applications (2014) vol.14 no.2 177 is ω. relation ap(m, ϕ) represents profit distribution parameters of air materiel named m is ϕ.relation ep(m, e) represents optimal expected profit is e in the collaborative ordering decision process. concept hr (x) represents returning to the upper layer agent. ordering decision agent determines the order quantity, collaboration price, profit distribution parameters and optimal expected profit, then the order scheme is returned to upper layer agent. after the objective of agents determined, it searches the act ability base according to the collaborative mode, and plan the sequence action for achieving the goals. the act of ordering decision agent comprises the action as follows: requestavailableinventory(availableinventory agent5,materiel, attribute1, ..., attributen, availablequantity)= < { ai(availableinventory agent5), k(availableinventory agent5), m (materiel), k(materiel), ma1(materiel, attribute1) , k(attribute1),..., man (materiel, aattributen) , k(attributen), ai(materiel, availablequantity)} , { k(availablequantity), ai(materie1, availablequantity)} > requestproducerschoose(producerschoose agent7, producer-j, attributej1, ..., attributejn, producer)= < { pc(producerschoose agent7), k(producerschoose agent7), p (producer-j), k(producer-j), pb1(producer-1, attribute11) , k(attribute11),..., pbn (materiel, aattributejn) , k(attributejn),..., pc(materiel, producer-j)} , { k(producer-j), ai(materie1, producer-j)} > requestdecisionsplan(manufacturerproductiondecision agent8, quantity, price, assigning parameters, optimalexpectedprofit)= < {mpd(manufacturerproductiondecision), k(manufacturerproductiondecision), m (materiel), k(materiel), p (producer-j), k(producer-j), ai (materiel, availablequantity) , k(availablequantity), pc (materiel, producer-j) , k(producer-j), (oq(materiel,quantity)∧op(materiel, price)∧ap(materiel, assigning parameters) ∧ep(materiel, optimalexpectedprofit))} , { k(quantity), oq(materiel,quantity), k(price), op(materiel, price), k(assigning parameters), ap(materiel, assigning parameters), 178 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... k(optimalexpectedprofit), ep(materiel, optimalexpectedprofit) } > returnordersplan(producer, quantity, price, assigning parameters, optimalexpectedprofit)= < {mpd(manufacturerproductiondecision), k(manufacturerproductiondecision), m (materiel), k(materiel), pc(materiel, producer-j), k(producer-j), oq(materiel,quantity), k(quantity), op(materiel, price), k(price), ap(materiel, assigning parameters), k(assigning parameters), ep(materiel, optimalexpectedprofit), k(optimalexpectedprofit), -(hr(producer-j)∧hr(quantity)∧hr(price)∧hr(assigning parameters)∧ hr(optimalexpectedprofit))}, { hr (producer-j), hr(quantity), hr (price), hr (assigning parameters), hr(optimalexpectedprofit) } > the concept k represents the value of an element has been determined. the requestavailableinventory indicates if the attribute of surplus stock, repairing parts and parts waiting for repair of an air materiel and an available inventory agent are known, then the number of available stock can be obtained through the available inventory agent. the requestproducerschoose indicates if the attribute of quality, order price, reliability, maintainability of an air materiel produced by different manufacturers and a manufacturer select agent are known, then the manufacturers can be obtained through the manufacturer select agent. the requestdecisionsplan indicates if the number of available inventory and manufacturer of an air materiel and a manufacturer production decisions agent are known, then order quantity, collaborative price, profit distribution parameters and optimal expected profit can be determined through the manufacturer production decisions agent. the appropriate ordering scheme of an air materiel can be formulate. the returnordersplan indicates if manufacturer, the number of available inventory, order quantity, collaborative price, profit distribution parameters and optimal expected profit are known, then it can be considered as an order scheme returned to upper layer agent. the action achieving decision objective as follows: requestavailableinventory (avi,m, a1, . . . , an, q0) requestproducerschoose (prc, pj , aj1, . . . , ajn, p) requestdecisionsplan (mpd, q, ω, ϕ, e) returnordersplan (p, q, ω, ϕ, e) the sequence action represents firstly the available stock and manufacturer can be obtained through available inventory agent and manufacturer selection agent, advances in systems science and applications (2014) vol.14 no.2 179 then the request of ordering is sent to manufacturer production decision agent for obtaining the collaborative price, profit distribution parameters and optimal expected profit of an air materiel, at last the order scheme is returned to upper layer agent. 4 the structure of the intelligence collaborative decision-making support system owing to the distribution of the geographical location of the air materiel supply chain members, the heterogeneity of the network and internal business data resources, and the dynamic of decision support activities, the problem of mistakes and misunderstandings among the supply chain members are easy caused. to facilitate the collaborative decision support activities in the air materiel supply chain members, it is essential to integrate and filter the information of inventory, distribution, production and logistics, and to provide a shared domain knowledge structure to enable the members to reach mutual understanding with each other. at the same time, collaborative mechanism of the members and collaborative mode of agent need to be designed, and the interaction between man and computer is realized based on web. fig.3 the system architecture there are many types of decision agents owned by the members of air materiel supply chain. in the every decisions process, each decision agent calls on the corresponding agent in supply chain network according to the decision tasks, and they complete the decision task through the collaborative mode. the system 180 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... architecture shows as fig.3. it mainly comprises the interface, task management module, internal database and management system of each node in air materiel supply chain, data warehouse and management system, the agents integrated base, model base, knowledge base, method base, problem base, case base, shared ontology base, private ontology base, the management systems, etc. when the members of air materiel supply chain submit decision problems to the system through interface, the task management module calls on the corresponding agent to resolve and deal with the decision problems, and the shared ontologies are exchanged to guarantee the mutual understanding of agents through the interface agent and ontology agent. then the collaborative decision-making model is established, and decision problems are solved by calculation and reasoning through calling on the database, model base, knowledge base, method base and case base. the subsequent collaborative processes between the members involve iterative exchanges of proposals and counter-proposals until the final result is achieved. 5 conclusions in this paper, an intelligence collaborative decision-making mode based on the ontology and multi-agent technology has been proposed. a generic structure and interoperability mode of ontology has been developed. it involves the knowledge representation method and collaborative protocol. the collaborative scheme of multi-agent has also been presented. then, the collaborative mode and action planning of agent are analyzed, and the architecture of intelligent collaborative decision-making system of two-echelon air materiel supply chain has been designed. a platform for the consultations and coordination of each agent has been provided, and an effective decision-making method has been proposed to decisionmakers of air materiel supply. the proposed decision-making mode is still in the development stage, in the future, the model can be extended to encompass more collaboration considerations in air materiel supply chains. references [1] jeffery k. cochran, theodore p. lewis.(2002) , “computing small-fleet aircraft availabilities including redundancy and spares”, computers & operations research, vol.29, no.5, pp.529-540. [2] jhugo simao, warren powell.(2009) , “approximate dynamic programming for management of high-value spare parts”, journal of manufacturing technology management, vol.20, no.), pp.147-160. [3] lee loo hay, chew ek peng, teng suyan, chen yankai.(2008) , “multiobjective simulation-based evolutionary algorithm for an aircraft spare parts advances in systems science and applications (2014) vol.14 no.2 181 allocation problem.”, european journal of operational research, vol.189, no.2, pp.476-491. [4] d.arnott, g. pervan.(2005) , “a critical analysis of decision support systems research”, journal of information technology, vol.20, no.2, pp.67-87. [5] g. phillips-wren, m. mora, g.a. forgionne, j.n.d. gupta.(2009), “an integrative evaluation framework for intelligent decision support systems”, european journal of operational research, vol.195, no.3, pp.642-652. [6] t. padma, p. balasubramanie.(2011) , “domain experts knowledge-based intelligent decision support system in occupational shoulder and neck pain therapy”, applied soft computing, vol.11, no.2, pp.1762-1769. [7] m.n. nguyen, d. shi, c. quek.(2008) , “a nature inspired yingcyang approach for intelligent decision support in bank solvency analysis.”, expert systems with applications , vol.34, no.4, pp.2576-2587. [8] xuan f. zha, ram d. sriram, marco g. fernandez, et al.(2008), “knowledge-intensive collaborative decision support for design processes: a hybrid decision support model and agent.”, computers in industry, vol.59, no.9, pp.905-922. [9] minhong wang, huaiqing wang, doug vogel, et al.(2009), “agent-based negotiation and decision making for dynamic supply chain formation”, engineering applications of artificial intelligence, vol.22, no.7, pp.1046-1055. [10] soheil boroushaki, jacek malczewski.(2010), “measuring consensus for collaborative decision-making: a gis-based approach”, computers, environment and urban systems, vol.34, no.4, pp.322-332. [11] babak khosravifar,jamal bentahar,rabeb mizouni, et al.(2013), “agentbased game-theoretic model for collaborative web services: decision making analysis”, expert systems with applications, vol.40, no.8, pp.3207-3219. [12] t.r. gruber.1993), “a translation approach to portable ontologies”, knowledge acquisition, vol.5, no.2, pp.199-220. [13] chang-shing lee, mei-hui wang, jui-jen chen.(2008), “ontology-based intelligent decision support agent for cmmi project monitoring and control”, international journal of approximate reasoning, vol.48, no.1, pp.62-76. [14] p. wongthongtham, e. chang, t.s. dillon, et al.(2006), “ontology-based multi-site software development methodology and tools”, journal of systems architecture, vol.52, no.11, pp.640-653. 182 xiong li:air materiel supply intelligent collaborative decision-making mode based on ... [15] m.j. dibley, h. li, j.c. miles, et al.(2011), “towards intelligent agent based software for building related decision support”, advanced engineering informatics, vol.25, no.2, pp.311-329. [16] mauricio paletta,pilar herrero.(2011), “simulating collaborative systems by means of awareness of interaction among intelligent agents”, simulation modelling practice and theory, vol.19, no.1, pp.17-29. [17] rudi studer, v.richard benjamins, dieter fensel.(1998), “knowledge engineering: principles and methods”, data & knowledge engineering, vol.25, no.12, pp.61-197. [18] ghassan beydoun, graham low, numi tran, et al.(2011), “development of a peer-to-peer information sharing system using ontologies”, expert systems with applications, vol.38, no.8, pp.9352-9364. [19] gong wang, t.n. wong, xiaohuan wang.(2013), “an ontology based approach to organize multi-agent assisted supply chain negotiations”, computers & industrial engineering, vol.65, no.1, pp.2-15. [20] alejandro rodrguez-gonzalez, javier torres-nino, gandhi hernandezchan, et al.(2012), “using agents to parallelize a medical reasoning system based on ontologies and description logics as an application case”, expert systems with applications, vol.39, no.18, pp.13085-13092. [21] yang yue-fu,sun li-quan.(2010), “the study of multi-agent cooperative framework based on pervasive computing”, journal of harbin university of science and technology, vol.15, no.1, pp.19-23. [22] egon ostrosi,alain-jerome fougeres, michel ferney.(2012), “fuzzy agents for product configuration in collaborative and distributed design process”, applied soft computing, vol.12, no.8, pp.2091-2105. [23] guoyin jiang, bin hu, youtian wang.(2010), “agent-based simulation of competitive and collaborative mechanisms for mobile service chains”, information sciences, vol.180, no.2, pp.225-240. corresponding author author can be contacted at: liuchangxin5128@msn.com. мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 43-52 concordance of private and public interests: dynamic graph representation, identification and simulation modeling a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky department of applied mathematics and computer science, southern federal university, rostov-on-don, russia annotation a universal approach to the description of dynamics of complex systems, identification of their models, and finding the solutions by means of computer simulation is proposed (on the example of cppi-models). the dynamic graph balance models permit to reflect in a convenient visual form the main matter-energetic processes of real-world systems dynamics. computer simulation is proposed to use both for solving of complex game theoretic problems and for their identification. an idea of representative scenarios is developed in this frame. a special computer software should be developed for the implementation of the method. key words: computer simulation, concordance of private and public interests, dynamic graph balance models, identification. 1 introduction in the seminal paper [7] a static model of joint consideration of private and public interests was proposed. they proved that if each agent's payoff function is a convolution by minimum of the private and public parts then a pareto optimal nash equilibrium exists in the agents' game in normal form. an investigation of the models of concordance of private and public interests (cppi-models) was continued by the authors [8,9]. the conditions of system compatibility in cppi-models based on the notion of price of anarchy [1] were studied, and different mechanisms of control providing the system compatibility were analyzed. dynamic versions of cppi-models were also built [3,4]. this paper makes a contribution to dynamic graph representation, identification and simulation of cppi-models as instruments of the applied systems analysis. first, there is a number of mathematical models which permit to describe the state of complex systems including explicit or implicit consideration of their dynamics. some examples are markov chains, finite automates, petri nets, queuing systems. in this paper we develop a technique of dynamic graph balanced cppi-models [15]. second, the standard methods of econometrics [6] and theory of identification [14] are based on long time series of reliable data which are often absent in real applications. we propose to solve the problems of structural and numerical identification by means of building a special computer software. third, computer simulation [13] is an appropriate method of solving complex dynamic problems. we specify this method for cppi-models with different information structure and introduce the idea of representative scenarios. the rest of the paper is organized as follows. in section 2 dynamic graph balanced cppimodels are described. section 3 gives an idea of computer simulation with cppi-models with different information structure based on a small number of representative scenarios. section 4 is concerned with computer simulation support of the identification of cppi-models. section 5 concludes. 44 a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky: concordance of private and public interests: dynamic graph … 2 dynamic graph balanced cppi-models the description of the state of a complex system in the moment of time t taking into consideration its structure by means of dynamic multidigraphs [15] includes the following elements. 1. a set of vertices y(t)=(y1(t),...,yn(t)) where n(t) is a number of vertices in the moment t. decompose the set y(t) onto two non-intersecting subsets: t y(t) = y1(t)  y2(t),y1(t)  y2(t) =  (it is possible that y2(t) = ). let’s name the vertices from the subset y1(t) compartments and denote them by squares and the vertices from the subset y2(t) transformers and denote by circles. 2. a set of arcs z(t)={zij k (t)}, 1 i, j  n(t), 1  k  n, where zij k (t) is the arc from the vertex yi to the vertex yj (in particular the loop if i=j) on which a resource k can move in the moment t; n is a total number of resources within the system. 3. a set of state variables of the compartments x(t) = {xi k (t)}, 1 i n(t), 1  k  n, where xi k (t) is a value of the resource k in the compartment yiy1 in the moment t. then xi(t) is a state vector of the compartment yi in the moment t (the collection of all its resources). 4. a set of state variables of the arcs f(t) = {fij k (t)}, 1 i, j  n(t), 1  k  n, where fij k (t) is a weight of the arc zij k (t), i.e. a number of the resource k moved during the time [t,t+1] from the vertex yi to the vertex yj, i j, or a quantity of increase (decrease) of the resource k in the compartment yi during the same time, i=j (∆t=1). in each considered situation (problem) the set f(t) can be split into two non-intersecting subsets: t f(t)=f1(t)  f2(t), f1(t)  f2(t) =  (in particular it is possible that f2(t) = ). variables from the set f1(t) are called regulated (they change in the strength of given rules) and variables from the subset f2(t) are called regulators (they can change arbitrarily in the admissible set). 5. a set of limitations on the compartments capacity x = {xi k }, 1 in(t), 1  k  n, where xi k is a maximal number of the resource k which can be stored in the compartment yi. 6. a set of limitations on the carrying capacity of the arcs f = {fij k }, 1 i, j n(t), 1  k  n, where fij k is a maximal number of the resource k which can be moved from the vertex yi to the vertex yj, ij, or produced (destructed) in the compartment yi, i=j, during the time unit. thus, the extended state of a complex system is a set s(t) = < y(t), z(t), x(t), f(t), x, f>. to avoid a consideration of digraphs with multiple arcs let’s map to each vertex yiy1 the only value xi(t) and to each arc zijz the only weight aij(t). then a dynamic structure of the system consists of the separate “scalar” structures each of which represent a certain aspect of matter and energy interactions within the system. the partition of a set of vertices of the dynamic digraph onto two parts permits to describe the principal matter-energetic processes in the realworld systems such as 1) movement (transfer, exchange) of resource between the compartments; 2) production/destruction of the resource in the compartments; 3) transformation of the resource. the respective models can be called dynamic graph balanced ones. describe the processes by such models. 1. a movement of the resource k between two compartments yi and yj in the moment t can be performed if the arc zij k (t) exists (fig.1). fig.1 a movement of the resource k between compartments yi and yj advances in systems science and application(2016) vol.16 no.4 45 assume that in the moment t the stocks of the resource k in the compartments yi and yj are equal to xi k (t) and xj k (t) respectively and the state variable of the arc zij k (t) is fij k (t). then ).()()1(),()()1( tftxtxtftxtx k ij k j k j k ij k i k i  (1) 2. production/destruction of the resource k in the compartment yi in the moment t is possible if the loop zii k (t) exists (fig. 2). fig.2 production/destruction of the resource k in the compartment yi the case fii k (t) > 0 corresponds to the production and the case fii k (t) < 0 to the destruction of the resource k. the result is ).()()1( tftxtx k ii k i k i  (2) 3. a transformation of some resource into other one is possible if a vertex-transformer from the set y2 exists. it is the most complicated class of processes which contains a number of subclasses. the subclasses can be classified by different criterions such as а) a simple transformation (resource k into resource l); b) a composite transformation (one resource into several ones, many resources into one or many to many); or a) a unit transformation (within one compartment); b) a binary transformation (between two compartments); c) a multiple transformation (between several compartments). consider the case bb as an example. assume that initial stocks of the resource are xi k (t), xi l (t), xj l (t). the transformation satisfies the equations ),()()1(),()()1( ),()()1( tftxtxtftxtx tftxtx l pj l j l j l ip l i l i k ip k i k i   (3) where yi, yj are compartments, yp is a transformer. now consider as a more detailed example a known predator-prey model[12] ,, 21222 2 21111 1 xxx dt dx xxx dt dx   (4) where x1(t), x2(t) are biomasses of the prey and predator respectively in the moment t; ε1, ε2 are coefficients of the natural increase of the populations; γ1, γ2 are coefficients of the predatorprey interaction. a representation of the model (4) by means of the dynamical digraph is shown in fig. 3. fig.3 a representation of the predator-prey model by means of the dynamical hierarchical digraph 46 a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky: concordance of private and public interests: dynamic graph … the loops z11 1 and z22 2 describe an increase of the prey biomass (resource 1) and a decrease of the predator biomass (resource 2) in the compartments y1 and y2 respectively, and the transformation y3 describes a simple binary transformation of the prey biomass into the predator biomass. in the general case a natural dynamics of the system resources is represented by a balance equation for each compartment and each resource: ,1),(,,1 ),()()()1( )()( nktnlji tftftxtx tsy k jl k tsy ij k j k j jlji      (5) the relation (5) added by initial data represents in fact a simulation model describing the system dynamics with consideration of its structure. the equation (5) considers both a passive regulation of the system (due to change of fjl k , fmj k  f1 in the strength of given relations) and an active one (due to choice of fjp k , fqj k  f2). the active regulation could additionally change the sets y and z. it is natural to name the changes of the sets x and f resource ones and the changes of the sets y and z structural ones. the totality of resource and structural changes determines the dynamics of the system s. an extended state of the system s(t) is also changed by external impacts. the system environment can be represented by a vertex y0 with the state vector x0(t) = (x0 1 (t), ... , x0 n (t)). respectively the set of arcs z(t) is added by elements of the type z0i k (t), zi0 k (t) and the set of state variables of the arcs f(t) by elements f0i k (t), fi0 k (t), 1 ≤ i ≤ n(t), 1 ≤ k ≤ n. an influence of the environment is considered in (5) without loss of generality with the condition that y0 can belong to the sets sj + , sj . besides, an external impact can change the sets y(t), z(t). if there are several sources of impact then it is necessary to introduce several external vertices y01 ,..., y0m , respective arcs and state variables. consider as an example the predator-prey model with man-made impact ,, 221222 2 121111 1 xxxx dt dx xxxx dt dx   (6) where in comparison with the model (4) the characteristics of man-made exploitation of the community are added, namely an intensity of use λ and methods of use α, β. a representation of the model (6) by means of the dynamical hierarchical digraph is given in fig. 4. in comparison with the fig.3 to the compartments y1 ("preys") and y2 ("predators") the compartment y0 reflecting the community environment (a source of exploitation) and arcs z10 1 , z20 2 with state variables αλx1, βλx2 are added [15]. fig.4 modeling of exploitation in the predator-prey system by means of the dynamical digraph advances in systems science and application(2016) vol.16 no.4 47 now consider cppi-model written in a discrete form:    ni ijj max; (7)    ni t i t i ss ;1;0 (8) max;)]()([ 0   t t tt i t i t iii xcsurpj (9) ;0 t i t i ru  (10)     ni t ii tt ubxx ;1 (11) .,...,1,0;);,(1 ttniuxgrr t i t i t i t i  (12) in this dynamic model each agent shares his resource ir between a production of a public good )( iu and a private activity )( ii ur  . respectively, his current payoff is a sum of the private gain )( iii urp  and the share in the consumption of the public good )(xcsi . his integral payoff is given by the formula (9), where ipc, are continuous increasing concave functions, .0)0()0(  ipc the utilitarian social welfare function (7) is also introduced. if it is associated with a principal then a choice of the variables is s. t. (8) is considered as an economic control of the principal. the equations of dynamics are given by formulas (11)-(12), where x is a state vector, and the function of its dynamics is linear for simplicity. the implementation of the model dynamics (11)-(12) can be given by the following algorithm: given ;,,,, 0000 niryzx ii  ttotfor 0: );(:{ 11 t i t ii t i t i urpyy       ni t ii tt ubxx ;: 1 );(: 1 ttt xczz   })()(: 1 tt i t i tt i t i t ii t i zsyxcsurpr   . this algorithm can be represented by a dynamic graph balanced model as follows (fig. 5). notice that the equation (12) is specified by means of this model. the relations for j and ij are omitted for simplicity, they can be represented similarly. in fact, in fig. 5 a general relation )(:1 t ijk t jk t jk xx  is presented graphically as fig.5 a balance relation 48 a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky: concordance of private and public interests: dynamic graph … by default, it is supposed that t ijk t ijk  )( (a linear transformation), and t ik t ijk x (a transference). fig.6 cppi-model presented by the dynamic graph balanced model 3 computer simulation with cppi-models schematically, a process of simulation modeling with a cppi-model can be represented as follo ws. each scenario mj ,...,1 includes a set of the principal's control variables t t n i t jis 11)( }{  and a s et of the agents' control variables t t n i t jiu 11)( }{  . these two sequences generate the respective syste m trajectory t t t jx 1)( }{  and payoffs )(1)( ,}{ j n iji jj  . however, it is important to consider an information structure of the hierarchical game (7)(12). to develop a classification of information structures in the hierarchical differential games with many followers three attributes characterizing the principal's strategy can be used: 1) absence/presence of a feedback of the principal's strategy on the state of a controlled dynamic system. this attribute has two basic values: open-loop strategies (ol) which depend only on the moment of time t , and closed-loop strategies (cl) which depend on the game position ))(,( txt [2]; 2) absence/presence of a feedback of the leader's strategy on the followers' strategies. in the first case we deal with a stackelberg game, and games of the second type we propose to call germeier games [10,11]; 3) methods of hierarchical control. here we differentiate compulsion, when the principal influences the followers' sets of feasible strategies, and impulsion, when the principal influences the followers' payoff functionals [15]. advances in systems science and application(2016) vol.16 no.4 49 in turn, the followers can choose one of the three modes of behavior: (a) isolation, when the followers act independently and come to a nash equilibrium; (b) cooperation, when they pool resources and combine efforts to maximize the summarized payoff functional; (c) collaboration, when the followers voluntarily maximize the principal's payoff functional. notice that in the case of cppi-models cooperation and collaboration coincide because    ni ijj . to explain the proposed classification we use the following two tables. table 1 basic information structures in the hierarchical games [11] without a feedback on the followers' controls with a feedback on the followers' controls without a feedback on the system state t1 t2 with a feedback on the system state x1 x2 table 2 maximal guaranteed payoffs of the principal for different information structures principal followers inaction (s-const) impulsion t1 , x1 (st) t2 , x2 (ger) isolation (ne) 0 nej stimp nej  gerimp nej  cooperation (c) 0 cj stimp cj  gerimp cj  in table 1 the types of leader's strategies using the denotations proposed in [11] are shown. the table 2 should be explained in more details. in hierarchical differential games the principle of optimality is a maximal guaranteed strategy of the principal with consideration of an optimal reaction of the followers. the respective maximal guaranteed payoffs of the principal for the enumerated information structures are collected in the table 2. in the case of isolation it is supposed that the optimal reaction of the followers is their nash equilibria set ne. in the case of cooperation the optimal reaction of the followers is the set c of points of maximum of their summary payoff functional. in this paper we consider only a case of impulsion, when the principal chooses a vector of strategies ),...,( 1 nsss  in the modes t1 , x1 (stackelberg games) or t2 , x2 (germeier games). the strategies can be ol ( t1 , t2 ) or cl ( x1 , x2 ).in the degenerate case of inaction s is constant (no control). thus, in the case of inaction ),(inf0 ujj neu ne   ),(inf0 ujj cu c   )(inf0 max ujj uu  . in the case of impulsion for stackelberg games we have ),(infsup )( usjj sneuss stimp ne    , ),(infsup )( usjj scuss stimp c    , and for germeier games )),~((infsup )~(~~ usjj sneuss gerimp ne     , )),~((infsup )(~~ usjj scuss gerimp c     , where usussuss  ~ :},:~{ ~  . now we can describe an approach to the implementation of the characterized information structures in the simulation mode. in the case t1 strategies have the form )(),( tuts . a scenario mj ,...,1 represents a pair of discrete control trajectories t t n i t jis 11)( }{  , t t n i t jiu 11)( }{  for which a discrete phase trajectory t t t jx 1)( }{  and the payoffs )(1)( ,}{ j n iji jj  are calculated. 50 a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky: concordance of private and public interests: dynamic graph … in the case of information structure t2 strategies have the form )()),(,( tututs . asarule, impulsion is implemented by a mechanism of reward and punishment of the type       ,),( ,)(),( ))(,( otherwisets ututs tuts p r r (13) where ru is a set of the agents' strategies encouraged by the principal, )(),( tsts pr strategies of reward and punishment respectively ),(),(( i r iii p ii usjusj  . in many situations it is possible to think that pprr stssts  )(,)( . for each agent's control trajectory t t n i t jiu 11)( }{  belonging to the scenario mj ,...,1 , at each step tt ,..,1 the condition (13) is checked, and the respective value t jx )( and the respective summands of payoffs are calculated. in the case of information structure x1 strategies have the form ))(,()),(,( txtutxts , and rules of the type            ,)(, ... ,)(, ,)(, ))(,( )( 2 )( 2 1 )( 1 )( l m j m lj lj j xtxs xtxs xtxs txts            ,)(, ... ,)(, ,)(, ))(,( )( 2 )( 2 1 )( 1 )( f l j l fj fj j xtxu xtxu xtxu txtu mj ,...,1 , (14) where  f r f k f l fl r l k l m l xxxxxxxxxx ,...,,... 11 , should be applied step by step for all scenarios t t t j t t t j us 1)(1)( }{,}{  , mj ,...,1 . the implementation of the game x2 is more complicated and is omitted here. it is important to note the following thing. in the majority of organizational and socioeconomic systems a number of scenarios reflecting qualitatively different control strategies is quite small. in fact, for a qualitative differentiation of control strategies it is sufficient to study "strong", "moderate", and "weak" types of them. for example, if an effort is measured on a scale [0,1] then the values 0, 1/2, 1 representatively reflect the mentioned strategies. if it is not sufficient due to a reason then the values 1/4, 3/4 can be considered additionally, and so on. such scenarios can be called representative ones. therefore, in the set of representative scenarios the complete enumeration becomes a practically implementable procedure. 4 computer simulation support of the identification of cppi-models the problems of identification of mathematical models are solved by econometrics [6] and theory of identification [14]. but these theories have at least two essential shortages. first, their implementation requires to use long time series of reliable data of observations. but such time series are often hardly available and even absent (for example, in social processes such as corruption). second, the standard methods solve only the problem of numerical identification, i.e. determination of the numerical values of model parameters. in the same time, in applied systems analysis a problem of structural identification is more important (which classes of functions should be used in the model). we propose a unified computer-based approach to solving the problems of structural and numerical identification (cppi-models are used as an example). in mathematical modeling it is well known that to provide adequacy one must build hierarchical gradually complicated sequences of models. first, the simplest model is built which reflects only the most essential features of a modeled object. as a rule, the model contains only a few parameters and allows for an analytic investigation. after that investigation is made and some results are received, it becomes clear in which direction the model is bounded and which properties of the real system it does not describe. then the model is perfected in the required advances in systems science and application(2016) vol.16 no.4 51 direction, and the procedure repeats. in the limit, the respective sequence of gradually complicated models can describe the system's behavior with any required accuracy. this methodology is used for both structural and numerical investigation of mathematical models. as for the structural identification, we always start from linear functions. specific features of the models should always be considered. in cppi-models the functions c and ip are increasing, thus we use linear functions )0(  abaxy . after that, it is natural to explore power functions. in cppi-models the functions c and ip are concave, thus functions )1(  baxy b are used. again, it is natural to start from 2/1b ; if it is necessary then other values 1b can be analyzed. at last, other classes of continuous increasing concave functions can be taken if necessary. in the numerical identification a feasible range of values is determined for each parameter (for example, maxmin aaa  ). using the idea of representative scenarios, the values 2/)(,, maxminmaxmin aaaa  are examined first. if it is not sufficient, a dichotomy of the segments ],2/)[(],2/)(,[ maxmaxminmaxminmin aaaaaa  is made, and so on. the idea is implemented by developing a computer software providing to a user the possibility of choice of different classes of functions and values of their parameters. the essence of the software consists in a modification of genetic algorithms for solving problems with several input parameters. the genetic algorithm decreases a number of required calculations on some orders in comparison with known numerical methods. the numerical methods give the value of a result with almost any given accuracy, but in the case of more than four input parameters too many calculations are required. in turn, genetic algorithms don't guarantee such an accuracy but practically don't depend on the number of model parameters. in the same time, the accuracy of the order 10 -6 can be achieved which is more than sufficient for cppi-models. a possible development of the software includes usage of parallel calculations. though genetic algorithms can decreases the number of calculations on some orders there are still hundreds of thousands and even millions of them, therefore paralleling seems a logical way. each stream will contain its own genetic algorithm and input data set. the solutions of each stream are cumulated for further processing. another promising direction of the software development is to seek a solution not in the domain of parameters values but in the domain of types of the solved problem. a modification of the genetic algorithm of another type is required for this. at last, the software can be developed by adding a possibility of the formation of arbitrary kinds of problems. the respective tool is parsing, i.e. a decomposition of a given term on more fine formulas which form a new term available for processing. thus, the described approach assumes an active usage of the methodology of applied systems analysis based on simulation modeling. if a precise analytical solution of an optimization problem is troubled then it can be investigated qualitatively on the base of scenario method using the idea of representative scenarios determined by means of systems analysis. therefore, simulation modeling serves as a universal numerical method of solving complex mathematical problems. 5 conclusion in this paper a universal approach to the description of dynamics of complex systems, identification of their models, and finding the solutions by means of computer simulation is proposed (on the example of cppi-models). the dynamic graph balance models permit to reflect in a convenient visual form the main matter-energetic processes of real-world systems dynamics. this technique is applied to cppi-models. computer simulation is proposed to use both for 52 a.v. antonenko, o.i. gorbaneva, g.a. ougolnitsky: concordance of private and public interests: dynamic graph … solving of complex game theoretic problems and for their identification. an idea of representative scenarios is developed in this frame. a special computer software should be developed for the implementation of the method. references [1] n. nisan, t. roughgarden, e. tardos and v. vazirani(2007), algorithmic game theory, cambridge university press. [2] basar t. and olsder g.j. (1999), dynamic noncooperative game theory, siam. [3] chistyakov a.e., nikitina a.v., ougolnitsky g.a. et al. (2015), "a differential game model of preventing fish kills in shallow waterbodies", game theory and applications. vol. 17. [4] dyachenko v.k., ougolnitsky g.a. and tarasenko l.v. (2014), "computer simulation of social partnership in the system of continuing professional education", advances in systems science and applications, vol.14, no.4, pp.27-37. [5] dockner e., jorgensen s., long n.v. and sorger g. (2000). differential games in economics and management science. cambridge university press. [6] dougherty c. (2011) introduction to econometrics. oxford university press. [7] germeier yu. b., vatel i.a. (1975). "equilibrium situations in games with a hierarchical structure of the vector of criteria", lecture notes in computer science, 27, 460-465. [8] gorbaneva o.i. and ougolnitsky g.a. (2013), "purpose and non-purpose resource use models in two-level control systems", advances in systems science and applications, vol.14, no.4, pp.378-390. [9] gorbaneva o.i. and ougolnitsky g.a. (2015) , "system compatibility, price of anarchy and control mechanisms in the models of concordance of private and public interests", advances in systems science and applications, vol.15, no.1, pp. 45-59. [10] gorelov m.a. and kononenko a.f. (2015). "dynamic models of conflicts. iii. hierarchical games", automation and remote control, vol.76, no.2, pp. 264-277. [11] kononenko a.f. (1977), " game-theory analysis of a two-level hierarchical control system", ussr computational mathematics and mathematical physics, vol.14, no.5, pp. 72-81. [12] kot m. (2001), elements of mathematical ecology. cambridge university press. [13] law a.m. and kelton w.d. (2000) simulation modeling and analysis. mcgraw-hill. [14] ljung l. (1999) system identification: theory for the user. prentice hall. [15] ougolnitsky g.a. (2011) sustainable management. nova science publishers. corresponding author guennady a. ougolnitsky can be contacted at: ougoln@gmail.com. microsoft word 3-hu rong.doc issn 1078-6236 international institute for general systems studies, inc. existence, uniqueness and asymptotic properties of neutrastochastic functional differential equations with markovian switching hu rong 1,2, hu shigeng 1 1department of mathematics, huazhong university of science and technology, wuhan 430074. china 2department of mathematics in science institute, wuhan university of technology, wuhan 430070. china abstract this paper considers the existence and uniqueness of solution to neutral stochastic functional differential equation with markovian switching with local lipschitz condition but neither the linear growth condition. and we discuss the asymptotic properties of this solution including moment boundedness and moment average boundedness in time. a one-dimension nonlinear example is discussed to illustrate the theory. keywords moment boundedness lyapunov function stochastic functional di fferential equations markovian switching generalized ito formula 1. introduction and preliminaries many practical systems may experience abrupt changes in their structure and parameters caused by phenomena such as component failures or repairs, changing subsystem interconnections, and abrupt environmental disturbances. the hybrid systems driven by continuous-time markov chains have recently been developed to cope with such situation, which have therefore received a great deal of attention, and have played a more and more important role in recent years. stochastic functional differential equations with markovian switching have been studied by many authors, and we here mention [8-12], in which they mainly discuss the asymptotic property of the solution, including the stability and moment boundedness and so on with the linear growth condition. kolmanovskii [1] studied the neutral stochastic differential delay equations with markovian switching, and discussed the existence and uniqueness of the solution of the equation and the moment asymptotic boundedness and moment exponential stability. mao [2] discussed the almost surely asymptotic stability of nsdde. in this paper, we will mainly consider neutral stochastic functional differential equations with markovian switching and discuss the existence and uniqueness of a global solution without the linear growth condition, and asymptotic properties including moment boundedness and moment average boundedness in time of the this global solution. 42-54 advances in systems science and applications (2011), vol. 11, no. 1-2 consider the neutral stochastic functional differential equations with markovian switching of the form: [ ( ) ( , ( ))] ( ( ), , ( )) ( ( ), , ( )) ( ),t t td x t u x r t f x t x r t dt g x t x r t dw t   (1) where ( ) ( ), [ ,0]tx x t       which is regarded as in ([ ,0]; ), ( )( 0)nc r r t t  is a right-continuous markovian chain on the probability space taking values in a finite state space  1, 2, ,s n  . moreover : ([ ,0]; ) ,n n nf r c r s r    : ([ ,0]; ) ,n n n mg r c r s r     : ([ ,0]; ) .n n mu c r s r    let  },}{,, 0 pff tt  be a complete probability space with a filtration 0}{ ttf satisfying the usual conditions (i.e. it is right continuous and 0f contains all p-null sets). let ( )( 0)w t t  be an m-dimensional brownian motion defined on this space. let 0 and )];0,([ nrc  denote the family of continuous functions  from ]0,[  to nr with the norm |)(|sup|||| 0    . if a is a vector or matrix, its transpose is denoted by ta . let ( ) 0r t t  , be a right-continuous markovian chain on the probability space taking values in a finite state space },,2,1{ ns  with generator nnij  )( given by jio jio ij ii itrjtrp   ),( ),(1{})(|)({   where 0 . here 0ij is the transition rate from i to j if ji  while    ij ijii  . we assume that the markovian chain )(r is independent of the brownian motion )(w . for any 2( , ) ( ; )nv x i c r s r  , define an operator lv from ([ ,0]; )n nr c r s   to r by advances in systems science and applications (2011), vol.11, no.1-2 43 1 ( , , ) ( ( , ), ) ( , , ) [ ( , , ) ( ( , ), ) ( , , )] 2 t x xxlv x i v x u i i f x i trace g x i v x u i i g x i         1 ( ( , ), ), n ij j v x u i j     (2) where 2 1 ( , ) ( , ) ( , ) ( , ) , , , ( , ) .x xx n i j n n v x i v x i v x i v x i v x i x x x x                    if x(t) is a solution to eq.(1) and let ( ) ( ) ( , ( ))tz t x t u x r t  (as is the following), then by the generalized ito  formula, we have 0 ( ( ), ( )) ( (0), (0)) ( ( ), ( )) t ev z t r t ev z r e lv z s r s ds   , where ( ( ), ( )) ( ( ), , ( ))tlv z t r t lv x t x r t . in this paper the following assumptions are imposed as standing hypothesis. assumption 1.1 both f and g are locally lipschitz continuous. assumption 1.2 for each ,i s there is constant (0,1)i  such that 0 ( , ) ( , ) ( ) ( ) ( ),iu i u i d               (3) where  is a probability measure and those , ([ ,0]; ).nc r    assume moreover that u(0,i)=0, f(0,0,i)=0, g(0,0,i)=0. in general, these assumptions will only guarantee a unique maximal local solution to eq.(1) for any given initial data ([ ,0]; )nc r   and 0(0)r i s  . however, the additional conditions imposed in it, we will guarantee that this maximal local solution is in fact a unique global solution, which is denoted by 0( , , ),x t i and this solution has properties 0limsup ( , , ) , p p x e x t i k   * 00 1 limsup ( , , ) , t p p x e x t i ds k t      (4) where 0  and 0p  are proper parameters, pk and * pk  are positive constants independent of  and 0i . 44 rong: existence, uniqueness and asymptotic properties of neutral…… for the convenience of reference, several elementary inequalities are given in the following which will be used frequently. for any , nx y r , , , 0. x y x y                  (5) 1 1( ) (1 ) , 1,0 1.p p p p px y x y p          (6) ( ) ,0 1.p p px y x y p     (7) 2 2 2 ( ) ,0 1. 1 x y x y          (8) before we state our main results, let us cite several useful lemmas. lemma 1.3 for any ( ) ( ; ), , 0,nh x c r r b  when ,x  ( ) ( ),h x o x  then sup[ ( ) ] . nx r h x b x      in this paper, when we use the notation ( )o x  , it is always under the condition x  . in addition, throughout this paper, const represents a positive constant, whose precise value or expression is not important. ( )i x const always implies that ( )( )ni x x r is bounded above. note that the notation ( )o x  includes the continuity. hence lemma 1.3 can be rewritten as ( ) .b x o x const     in this paper, let 2( , ) ( ) ( ) p t n iv x i x q x x r  . ( )n n iq r i s  are positive definite matrices and 0p  . clearly, we have 2 2( , ) , p pp p i iq x v x i q x  (9) here min ( )i iq q . by (2), we have 12( , , ) ( ) [2 ( , , ) ( , , ) ( , , )] 2 p t t t i i i p lv x i z q z z q f x i g x i q g x i     advances in systems science and applications (2011), vol.11, no.1-2 45 2 22 2 ( 2) ( ) [ ( , , )] ( ) , 2 p p t t t i i ij i j p p z q z z q g x i z q z    (10) where ( , )z x u i  . lemma 1.4 let i be the last term of (10), then we have , p pi m z (11) where 2 2max ( ) 0. p p p i ii i j i ij jm q q    proof clearly, (11) is obtained directly. we only need to prove that 0pm  . we may suppose 1 2 .nq q q   noting that 1 1q q , then 2 2 22 2 11 1 1 11 1 1 1 1 1 1 1 1 0. p p pp p p j j j j j j j m q q q q q                lemma 1.5 assume 1p  , let x(t) be a solution of eq.(1) with 0x  , we have limsup ( ) (1 ) limsup ( ) , p pp t t e x t e z t       (12) where  1( ) ( ) ( , ( )), max .t i n iz t x t u x r t      proof by (3) and (6), we have 01( ) (1 ) ( ) ( ) p p pe x t e x t d              1 0 (1 ) sup ( ) sup ( ) . p pp s t s t e z s e x s             this implies 1 0 0 sup ( ) (1 ) sup ( ) sup ( ) . p p p pp s t s t s t e x s e z s e x s                   so, we can get sup ( ) p s t e x s     . then limsup ( ) (1 ) limsup ( ) p pp t t e x t e z t       . 2. a basic lemma the following lemma plays a key role in this paper. 46 rong: existence, uniqueness and asymptotic properties of neutral…… lemma 2.1 under assumptions 1.1 and 1.2, let 1,p  if there exist constants 00, , , , , 0( ,1 ),i j i ija k k i s j m       positive definite matrices iq and probability measures j , such that ( , , ) ( ( , ), )lv x i v x u i i     0 0 ( ( ) ( ) ),j jp i i ij j j a x k k d e x               (13) then for any initial data ([ ,0], )nc r   and 0(0)r i s  there exists a unique global solution 0( , , )x t i to eq.(1) and this solution satisfies (4). proof for any given initial data ([ ,0]; )nc r   and 0i s , write 0( , , ) ( )x t i x t  , we will divide the whole proof into three steps. step 1 let us first show the existence of the global solution x(t). under assumption 1.1 and 1.2, eq.(1) admits a unique maximal local solution ( )( )x t t    , where  is the explosion time. let ( ) ( ) ( , ( ))tz t x t u x r t  , define the stopping time  inf 0 : ( ( ), ( )) , ( )k t v z t r t k k n      . since  is bounded, when k is large enough, ( ( ), ( ))v z r k   for      , thus, 0k  . if    , when t  , z(t) may explode. hence,  : ( ( ), ( )) , ( )t v z t r t k k n        shows that k  . thus, we may assume 0 ( )k k n    . obviously, k is increasing and ( ) . .k k a s     . if we could show , . .a s   , then . .a s   . thus it need only, for any 0t  , ( ) 0kp t   as k  . fix 0t  . now we prove that ( ) 0kp t   as k  . first note that if k   , then by the continuity of x(t) and the right continuity of r(t), ( ( ), ( ))k kv z r k   . hence, by (13), we have advances in systems science and applications (2011), vol.11, no.1-2 47 ( ) ( ( ), ( )) ( ) ( ( ), ( ))k k k k k kkp t v z r p t ev z t r t           0 0 ( (0), ) ( ( ), ( )) kt ev z i lv z s r s ds     0( (0), )ev z i + 0 00 [ ( ) ( ) ( ) ] k j j t r rj j j e k k x s d x s ds                     0 0 0 0 ( (0), ) ( ) ( ) ( ) k kj j t t j j j ev z i k t k e d x s ds x s ds                     0 0 0 0( (0) ( (0), ), ) ( ) : ,j j t j v u i i k t k d k                 where the index r represents r(t), max (0 )j i ijk k j m    , and tk is a positive constant independent of k. so we can get 1( ) 0( ).k tp t k k k     that shows that x(t) is a global solution to eq.(1). step 2 let us now show inequality (4). by (13), we obtain that 0 ( ( ), ( )) ( (0), (0)) [ ( ( ), ( ))] t se ev z t r t ev z r e l e v z s r s ds    0 00 ( (0), (0)) [ ( ) ( ) ( ) ]j j t s r rj j j ev z r e e k k x s d e x s ds                     01 ( ) 0 0 0 1( (0) ( (0), ), ) ( 1) ( ) : ,jt t j j v u i i k e k e d c ke                         where 1c is a positive constant independent of t and 1 0k k   is a positive constant independent of  and 0i . hence, we have limsup ( ( ), ( )) . t ev z t r t k   then the required assertion (4) follows from (9) and (12). step 3 finally, using (13), we obtain that 0 ( ) t p a e x s ds    0 00 ( ( ), ( )) [ ( ) ( ) ( ) ]j j r rj j j t lv z s r s k k x s d x s de s                     0 0 0 0 2 0( (0) ( (0), ), ) ( ) : ,j j j v u i i k t k d c k t                   48 rong: existence, uniqueness and asymptotic properties of neutral…… where min i ia a and 2c is a positive constant independent of t. the assertion (4) follows directly. the proof is therefore complete. denote the left side of (13) by  and establish the inequality 0 ( ( ) ( ) ) ,j j ij j j k d e x i              (14) where ( ). p p ii a x o x      (15) by lemma 2.1, we have ( ) . 2 p pia x o x const      this together with (15) yields . 2 pia i x const    substituting this into (14) shows that the condition (13) are required. to get (14) and (15), some conditions imposing on the coefficients f and g. these conditions are considered in the next section. 3. main results recall  to denote the left hand of (13). if p>2, by (9) (10) and (11) 1 12 2( ) ( , , ) ( ) ( , ) ( , , ) p p t t t i i i ip z q z x q f x i p z q z u i q f x i      2 22 2 1 2 3 4 ( 1) ( , , ) [ ] : . 2 p pp p i p i p p q z g x i m q z i i i i         (16) we firstly list the following conditions that we will need: (h1) there exist , , 0,i ia   positive-definite matrices iq and a probability measure  , such that 02 2 2 ( , , ) ( ) ( ) ( ).t i i ix q f x i a x d o x                  (h2) there exist 0, , 0i ir r   and a probability measure  , such that 01 1 1 ( , , ) ( ) ( ) ( ).i if x i r x r d o x                (h3) there exist 0, , 0,i i    positive-definite matrices iq and a probability measure  , such that 01 1 1 ( , , ) ( ) ( ) ( ).i ig x i x d o x                  advances in systems science and applications (2011), vol.11, no.1-2 49 we can now state our main result in this paper. theorem 4.1 under assumptions 1.1 and 1.2, if the conditions (h1)-(h3) hold, 2 , 2 3p    and 2 2 2 2 1 ( )( 2) [ 1 ( 1)(1 ) p p i i i i i i ip p i i r r p p a q p                  2( 1)( ) (1 sgn( 2 ))], 2 i i i p q         (17) then for any initial data ([ ,0], )nc r   and 0(0)r i s  there exists a unique global solution 0( , , )x t i to eq.(1) and this solution satisfies (4). proof let 0( ) ( , , )x t x t i and  be sufficiently small. now we estimate 1 4i i respectively. first, by the condition (h1), the inequalities (5) and (7), we can have 01 2 2 222 1 [ ] p p pp i i ii a p q x x d              0 01 2 2 2 222 [ ][ ( )] p p pp i i iq x d d op x                     01 22 ( 2) ( 2) [ p p p p p i i i i p x q a x ap d p                        0 0( 2) ( 2) ( ) ( ) p p p p i p x d o d o x p                           0 02 ( 2) ( ) ( 2) ( ) ( ) ( )].i p p p i p s d s d p                          (18) next, by the condition (h2), the inequalities (5) and (7), we obtain 01 112 2 [( 2) ] 1 p p pp i i p i q p x p d p            01 1 1 [ ( ) ( ) ( )]i ir x r d o x              0 2 ( 1) ( 1) [( 2) ( 2) 1 p p p p i i i p xp q p r x p r d p p                      50 rong: existence, uniqueness and asymptotic properties of neutral…… 0 01 ( 1) ( 1) ( ) ( ) p p p pp i i p x p r d o d o x p                            0 01 ( 1) ( ) ( 1) ( ) ( ) ( )]. p p p i i p s p r d s d p                         (19) then by the condition (h3) and the inequalities (5), (7) and (8), we can get 02 222 3 ( 1) [ ] 2 p p pp i i p p i q x d             0 2 222 22 2 2 [ ( )] 1 ii i i dx o x             2 22 2 2 022 ( 2) (2 2)( 1) [ 2 2 p ppp pi i i i i i p xp p q x d v p                       2 22 0 0 2 2(2 2) ( 2) ( ) ( ) 1 2 p p p pi i p x d o d o x p                             2 22 2 0 0 ( 2) ( ) (2 2) ( ) ( ) ( )], 1 2 p pp i i i p s d s d p                          (20) where , (0,1)i v  are constants. it is easy to see that 012 4 [ ][(1 ) ]. p p pp p i i ii m q x d             (21) then substituting (18)-(21) into (16), we can get  whose form is similar to (14), where 1 22 ( 2) ( 2) ( 2) ( 2) { p p i i i i i e p p e q a ai p p p                    2 2 1( 1) ( 1) ( 1) ( 1) [( 2) 1 p ip p i ii i i q p e e p e p r p r p p p                       2 2 2 2 2 1 2 ( 1) ( 2) ]} [ 2 1 p pppp i i i i i i i i i i i i p p p r p re x q e v                     advances in systems science and applications (2011), vol.11, no.1-2 51 2 2 2( 2) (2 2) (2 2) ( 2) ] ( ) ( ). 2 1 2 p p pi i e p e p x o x o x p p                        if 2  , then we have 12 ( ), p p p i ii p q a x o x       where 2 2( 2) ( 2) ( 2) ( 2 ] { ) [1 p p i i i ii i e p p e a a e p p                      1( 1) ( 1) ( 1) ( 1) [( 2) 1 i p i i i q p e e p p r p r p p p                  1( 2) ]} : ( ).p i i i ip r p re a     by (17), we have (0) 0ia  . since  is sufficiently small, we get 0ia  . therefore, the form of i is similar to (15). if 2  , then we can get 12 ( ), p p p i ii p q a x o x       where 2 2 2 2 21 ( 2) ( 2) [ 2 1 p p i i i i i i i i i i i p e p a a q e v p                        2 ( 2) ( 2) ] : ( , ). 1 i i i e p a v p              choosing that ( )i i i i     and by (17), we get ( , ) 0ia v  . then we also have 0ia  , and the form of i is similar to (15). thus, by lemma 2.1, we can get that for any initial data  and 0i , there exists a unique global solution 0( , , )x t i to eq.(1) and this solution satisfies (4). theorem 4.2 under assumptions 1.1 and 1.2, if the conditions (h1)-(h3) hold, 2 , 2 3p    and 2 1 2 21 ( )( 2) [ 1 ( 1)(1 ) p p i i i i i i i i i i i r r p p a q p                     52 rong: existence, uniqueness and asymptotic properties of neutral…… 2( 1)( ) (1 sgn( 2 ))], 2 i i i p q         (22) then for any initial data ([ ,0], )nc r   and 0(0)r i s  there exists a unique global solution 0( , , )x t i to eq.(1) and this solution satisfies (4). the proof is mostly the same as the one we provided previously, only when we estimate the 1 3i i , we use the inequality (6) not (7). 4. one-dimension nonlinear example let us discuss a one-dimension nonlinear neutral stochastic functional differential with markovian switching to illustrate our theory. 04 2 4 2 1 [ ( ) 0.1 ( 1)] [ ( ) ( ) ( ) ] ( ) ( ).rd x t x t b x t cx t d x t d dt x t dw t             let 3, , 3, 1, 0.1, 0, 0i i ip q e b          , then 0 05 5 55 3 4 1 1 1 4 ( , , ) ( ) ( ), 5 5 t i i ix q f x i b x cx dx x d b x d o x                  04 4 4 1 ( , , ) ( ),if x i b x d o x       2 ( , , ) .g x i x  now, 1 4 , , , 1, , 0. 5 5i i i i i i i i ia b r b r          by theorem 4.1, when 166 25ib  , we can conclude that there exists a unique global solution to eq.(1), and the solution has properties (4) . references [1] kolmanovskii, v.b., neutral stochastic differential delay equations with markovian switching. stoch. anal. appl., 2003, 21 : 819-847. [2] x. mao, y. shen and c. yuan, almost surely asymptotic stability of neutral stochastic differential delay equations with markovian switching. stochastic processes and their applications, 2008, 118: 1385-1406. [3] x. ouyang, s. hu, stochastic optimization of firms investment decision under indeterminate conditions. advances in systems science and applications.2008, 1: 46-52. advances in systems science and applications (2011), vol.11, no.1-2 53 [4] s. qu, m. gong, a sliding mode control strategy for uncertain systems with time delays. advances in systems science and applications.2008, 3: 506-512. [5] j. luo, comparison principle and stability criteria for stochastic delay differential equations with poisson jump and markovian switching. nonlinear anal., 2006, 64: 253-262. [6] x. mao, stability of stochastic differential equations with markovian switching. stochastic processes appl., 1999, 79: 45-67. [7] l. wang, s. hu, comparison principle and stability of stochastic functional differential equations with markovian switching. advances in systems science and applications.2008, 2: 220-227. [8] x. mao, exponential stability of stochastic delay interval systems with markovian switching. ieee trans. automat. control., 2002, 47(10):1604-1612. [9] x. mao, robustness of stability of stochastic differential delay equations with markovian switching. stability control:theory appl., 2000, 3(1): 48-61. [10] a. v. svishchuk, yu. i. kazmerchuk, stability of stochastic delay equations of it ô form with jumps and markovian switchings, and their applications in finance. theor. probab. math. stat., 2002, 64: 167-178. [11] x. mao, asymptotic stability for stochastic differential delay equations with markovian switching. functional differential equations, 2002, 9: 201-220. [12] c. yuan, x. mao, asymptotic stability in distribution of stochastic differential equations with markovian switching. stochastic processes appl., 2003, 103: 277-291. 54 rong: existence, uniqueness and asymptotic properties of neutral…… advances in systems science and application (2015) vol.15 no.1 45-59 system compatibility: price of anarchy and control mechanisms in the models of concordance of private and public interests olga i. gorbaneva and guennady a. ougolnitsky southern federal university, russia abstract the problem of system compatibility is considered. its solution ensures maximization of the social welfare by consideration of individual interests of the agents. the conditions of system compatibility and respective control mechanisms are analyzed for the models of concordance of common and private interests in the agents’ resource allocation. keywords concordance of interests; control mechanisms; hierarchical game theory; system compatibility 1 introduction a problem of concordance of interests in the active systems may be considered in two aspects. first, it is well known that an egoistic behavior of independent active agents often implies a less social welfare than in the case of their coordinated actions. the quantitative side of this problem is named “inefficiency of equilibria”[1] and can be characterized by the price of anarchy index introduced by papadimitriou [2].second, the agents can allocate their resources between private and public interests. in the seminal paper by germeier and vatel [3],it is shown that if payoff functions of all agents are convolutions by minimum of the functions of public and private interests then in the respective game there is a pareto-optimal nash equilibrium (i.e. the price of anarchy is equal to the ideal value of one). we continue to investigate models of that type in literature [4,5] and in the present paper. mathematical methods of solution of the static problems of concordance of interests of active agents are developed in the theory of incentives [6], the information theory of hierarchical systems [7,8], the theory of control in organizations [9,10], mechanism design [1]. it should be noticed that the mechanism design investigates another setting of the problem: how to motivate agents to report the true information about their type (the problem of strategy-proofness, or incentive compatibility). in the present paper a notion of system compatibility is introduced. the system compatibility means that individually optimal controls of agents form the globally optimal vector of controls for a social welfare function. this setting is close to the problem of meta-game synthesis in the theory of active systems [11]. conditions of the system compatibility are studied for the models of concordance of private and public interests (cppi-models). as the conditions are quite re46 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... strictive, the control mechanisms directed to ensure the system compatibility are proposed. in the framework of the authors’ concept [12], the administrative and economic mechanisms with/without a feedback are constructed. the rest of the paper is organized as follows. in the section 2 the principal notions such as price of anarchy, system compatibility, control mechanisms, and cppi-models are introduced. economic mechanisms without feedback and with it as well as algorithms of their implementation are considered in the sections 3 and 4 respectively. administrative control mechanisms are discussed in the section 5. section 6 concludes. 2 principal notions. let’s consider a set n = {1, 2, ..., n} of active agents maximizing their payoff functions gi(u1, ..., un) → max (1) s.t. ui ∈ ui, i ∈ n (2) let the solution of the game (1)-(2) be a nash equilibrium une ∈ ne,u = (u1, ..., un). introduce a utilitarian social welfare function. g0(u) = ∑ j∈n gj(u). let umax be a solution of the maximization problem g0(u) → max, u ∈ u = u1 × ...× un (3) denote gmax 0 = g0(u max) . definition 1. the model(1)-(3)is system compatible if ∀une ∈ ne g0(u ne) = gmax 0 . a quantitative measure of the system compatibility is the price of anarchy [1-2] pa = min une∈ne g0(u ne) gmax 0 (4) it is evident that a model is system compatible if pa = 1 . the system compatibility is a rare phenomenon, and it is worthwhile to use control mechanisms for its achievement[9]. suppose that maximization of the social welfare (3) is the objective of a specific agent (center, principal, mechanism designer and so on) who has an ability of impact on the sets of feasible controls and/or payoff functions of the other agents to provide this objective. denote the first possibility as ui = ui(qi) , and the second one gi = gi(pi, ui), where q, p are vectors of the principal’s administrative and economic controls respectively. in the context of literature[12] we can differentiate the following control mechanisms (methods of control). advances in systems science and application (2015) vol.15 no.1 47 table 1 control mechanisms principal’s impact without a feedback (γ1) with a feedback (γ2) on the sets of feasible controls of the agents(administrative one,or compulsion) qi = const qi = qi(u) on the agents’ payoff functions(economic one,or impulsion) pi = const pi = pi(u) thus, the principal can exert influence on the sets of feasible controls of the agents (administrative control mechanism, or compulsion) or on the agents’ payoff functions (economic control mechanism, or impulsion). both mechanisms can include not or include a feedback on control. in the first case a hierarchical game of the type γ1(stackelberg game) holds, meanwhile the second case generates a hierarchical game of the type γ2(germeier game). so, four types of control mechanisms are possible (table 1). notice that now or in dependence of the used control mechanism. definition 2. a control mechanism q(p) in the model (1)-(3) is system compatible if gmax 0 = g0(u ne(q)) or gmax 0 = g0(u ne(p)) respectively. for definiteness let’s specify the model (1)-(3) in the form gi(u) = pi(ri − ui) + sic(u) → max, 0 ≤ ui ≤ ri, i ∈ n (5) g0(u) = ∑ j∈n pj(rj − uj) + c(u) → max, ∑ j∈n sj = { 1, ∃i : si > 0, 0, otherwise. (6) here, ri > 0 is a resource of the i-th agent; ui is a part of the resource assigned for production of the public payoff described by a function c(u); si is a share of the i-th agent in the public payoff; pi(ri − ui) is a function of the i-th agent’s private interest. the functions pi, c are supposed to be continuously differentiable and concave in all arguments. so, each agent shares his resource between public and private interests according to the ratio ui and ri − ui respectively. thus, the model (5)-(6) describes a concordance of the private and public interests in resource allocation. our investigation of models of the type (5)-(6) (cppimodels) develops the approach by germeier and vatel and burkov and opoitsev [3,11]. economic control mechanisms in the model (5)-(6) are implemented by the principal’s choice of the values si. to use administrative mechanisms one should suppose additionally that the principal can bound feasible controls of the agents: q̃i ≤ ui ≤ q̄i, i ∈ n (7) 48 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... the control mechanisms from the table 1 can be specified for the cppi-model (5)-(7) as follows (table 2). table 2 mechanisms of system compatibility in cppi-models principal’s impact without a feedback (γ1) with a feedback (γ2) on the sets of feasible controls of the agents(administrative one,or compulsion) q̃i ≤ ui ≤ q̄i, q̃i, q̄i = const q̃i(u) ≤ ui ≤ q̄i(u) on the agents’ payoff functions(economic one,or impulsion) si = const si = si(u) 3 economic control mechanisms without a feedback. assume that in the model (5)-(6) ∀i si = const .using the first order conditions we get that the internal system compatibility in the model (5)-(6) holds only if ∂c ∂ui = 0, i ∈ n (8) let’s notice that it is true also for models with partly coincident interests in more general form [5] gi(u) = pi(ui) + sic(u) → max, ui ∈ ui, i ∈ n (9) so, the following statement is true. theorem 1. suppose that ∃i ∈ n : ∂c/∂ui ̸= 0. then for system compatibility in the model (5)-(6) it is necessary that ∀i ∈ n ui = 0 ∨ ui = ri . in other words, the system compatibility in the model (5)-(6) is possible only if all agents are pure individualists (ui = 0) or pure collectivists (ui = ri) . example 1 (linear cppi-model). consider a linear specification of the model (5)-(6): gi(u) = ki(ri − ui) + sik ∑ j∈n uj → max, 0 ≤ ui ≤ ri, i ∈ n (10) g0(u) = ∑ j∈n kj(rj − uj) +k ∑ j∈n uj → max; s : 0 ≤ si ≤ 1, ∑ j∈n sj = { 1, ∃i : s > 0, 0, otherwise. (11) advances in systems science and application (2015) vol.15 no.1 49 here ki > 0,k > 0 are given constants which characterize efficiencies of the functions of private and public payoffs respectively. the first order conditions for the agents and the principal give respectively: ∂gi ∂ui = ksi − ki { ≥ 0, si ≥ ki k ⇒ une i = ri; < 0, si < ki k ⇒ une i = 0; ∂g0 ∂ui = k − ki { ≥ 0, ki ≤ k ⇒ umax i = ri; < 0, ki > k ⇒ umax i = 0. thus, the system compatibility is possible and holds on the bounds of the segments of feasible controls. two partitions of the set n can be defined: n = i0 ∪ c0, i0 ∩ c0 = ∅ : i0 = {i ∈ n : ki > k ⇒ umax i = 0}(immanent individualists); c0 = {i ∈ n : ki ≤ k ⇒ umax i = ri}(immanent collectivists). n = i(s) ∪ c(s), i(s) ∩ c(s) = ∅ : i(s) = {i ∈ n : si < ki k ⇒ une i = 0}(controlled individualists); c(s) = {i ∈ n : si ≥ ki k ⇒ une i = ri}(controlled collectivists). the immanent partition is determined by the objective properties of the model (10)-(11), meanwhile the controlled partition results from the optimal reaction of the agents on the choice of a vector s = (s1, ..., sn) by the principal. notice that i ∈ i0 ⇒ ki/k > 1 ⇒ ∀s ∈ s si ≤ ki/k ⇒ i ∈ i(s), i.e. ∀s ∈ s i0 ⊂ i(s) . the inverse statement is wrong: let ki ≤ k ⇒ i ∈ c0 but si = 0 ⇒ i ∈ i(s). similarly, it is simple to show that ∀s ∈ s c(s) ⊂ c0. it is also clear that if ∃s ∈ s : i(s) = i0, c(s) = c0 then the model is system compatible. to find her optimal control s∗ the principal should solve a discrete optimization problem g0(s) = ∑ j∈i(s) kjrj +k ∑ j∈c(s) rj → max (12) s : 0 ≤ si ≤ 1, i ∈ n, ∑ j∈n sj = { 1, ∃i : si > 0, 0, otherwise. (13) if c0 = ∅ then the model is system compatible for ∀s ∈ s .in this case n = i(s) = i0, g max 0 = gi0 = ∑ j∈n kjrj (individualistic society). 50 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... suppose that c0 ̸= ∅ if ∑ i∈c0 ki ≤ k then let for i ∈ c0, s ∗ i = ki k + εi so that∑ i∈c0 s∗i = 1. then (s∗) = c0 , and the model is system compatible. particularly, if i0 = ∅ then n = c0 = c(s∗), gmax 0 = gc0 = k ∑ j∈c0 rj (collectivistic society). in general case i0 ̸= ∅, c0 ̸= ∅, ∑ j∈c0 kj > k. here |c0| > 1. subject to ∀s ∈ s , i0 ⊂ i(s) it is impossible to put any i ∈ i0 into c(s), thus the problem (12)-(13) is reduced to the construction of a set c(s) ⊂ c0 such that g0(s) = ∑ j∈i(s)=c0\c(s) kjrj +k ∑ j∈c(s) rj → max, ∑ j∈c(s) kj ≤ k, s ∈ s this problem is still being solved. example 2 (power-linear cppi-model). consider the following specification of (5)-(6): gi(u) = ki √ ri − ui + sik ∑ j∈n uj → max, 0 ≤ ui ≤ ri, i ∈ n (14) g0(u) = ∑ j∈n kj √ rj − uj +k ∑ j∈n uj → max, 0 ≤ si ≤ 1, ∑ j∈n sj = { 1,∃i : s > 0, 0, otherwise. (15) the first order conditions for the agents and the principal give respectively: 0 = ∂gi ∂ui = ksi − ki 2 √ ri − ui ⇒ une i = { ri − k2i 4k2s2i , si ≥ ki 2k √ ri ; 0, otherwise; 0 = ∂g0 ∂ui = k − ki 2 √ ri − ui ⇒ umax i = { ri − ki 4k2 , ki ≤ 2k √ ri; 0, otherwise. thus, the system compatibility in the model (14)-(15) holds only if ∀i ∈ n , ki ≥ 2k √ ri (16) if this condition holds, then n = i(s) = i0, gmax 0 = gi0 = ∑ j∈n kj √ rj for any s ∈ s. if (16) doesn’t hold, then the problem of system compatibility can be formulated in a weaker form of maximization the price of anarchy (4) by a mechanism of economic control. the following problem of discrete optimization arises g0(s) = ∑ j∈i(s) kj √ rj + ∑ j∈c′(s) [krj + kj 2 2ksj − k2j 4ks2j ] → max (17) advances in systems science and application (2015) vol.15 no.1 51 s.t. (13), where i(s) = {i ∈ n : si ≤ ki 2k √ ri ⇒ une i = 0}, c ′(s) = {i ∈ n : si > ki 2k √ ri ⇒ une i = ri − k2i 4k2s2i } this problem is reduced to the following one: ∑ j∈c′(s) ( 2k2j sj − k2j s2j ) + λ( ∑ j∈c′(s) sj − 1) → max (18) the foc gives: −2k2j 2s2j + k2j 2s3j + λ = 0, s.t. s ∈ s. multiplication by sj gives the system s3j − 2k2j sj λ + 2k2j λ = 0∑ i∈c′(s) si = 1 to find sj and λ the following should be done: 1) to express sj by λ from the first equation; 2) to substitute the expression into the second equation and find λ; 3) to substitute the value of λ back to the expression from 1) and find the respective si. given λ the first equation can be solved analytically by cartan method (the solution is omitted due to its awkwardness) or numerically. as only real solutions such as 0 ≤ si ≤ 1 are feasible, the following conclusions can be received: (1) a solution exists only if λ < 0. rewriting the equation as s3j = 2k2j λ (sj − 1), we get λ < 0 due to positive left part and negative terms in the numerator and in the brackets; (2) when λ < 0 the only real solution si of the equation exists, and 0 ≤ si ≤ 1. actually, lets find the derivative of the expression f (sj) = s3j − 2k2j sj λ + 2k2j λ : f ′ (sj) = 3s2j − 2k2j λ > 0. therefore, the function f(si) increases in r , and if a root of the equation f(si) = 0 exists then it is unique.now let’s prove that the root exists and satisfies 0 ≤ si ≤ 1. f (0) = 2k2j λ < 0, f (1) = 1 > 0. subject to continuousness of f(si) the property is proved. 52 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... (3) sj increases in λ. let’s find the derivative ∂sj ∂λ of an implicit function: 3s2j · s′j − 2k2jλ·s′j−2k2j sj λ2 − 2k2j λ2 = 0, or s′j = 2k2j (1−sj) λ(3s2jλ−2kj) < 0 that implies. (4) ∑ i∈c′ si increases in λ. therefore, it is possible to choose such λ that ∑ i∈c′(s) si = 1. thus, a problem with two variables is reduced to a problem with one variable. the problem can be solved by the method of dichotomy in which the left bound of a segment is equal to λ1 : ∑ i∈c′(s) si > 1 (for example,λ1 = −ε √ b2 − 4ac , and the right bound is a big enough λ2 : ∑ i∈c′(s) si < 1. 4 economic control mechanisms with a feedback. now assume that in the model (5)-(6) si = si(ui) or even si = si(u). according to foc, the internal system compatibility is possible only if ∂si(u) ∂ui c(u) = [1− si(u)] ∂c(u) ∂ui , i ∈ n (19) this condition is less restrictive than (8) when si = const. both empirical and theoretical approaches can be used in the following analysis. in the context of empirical approach widely spread in practical activity methods of resource allocated are investigated. for example, consider a natural method of proportional allocation si(u) =  ui∑ j∈n uj , ∃m : um > 0, 0, otherwise. (20) in this case (19) takes the form ∑ j ̸=i uj [ ∂c(u) ∂ui ∑ j∈n uj − c(u)] = 0, i ∈ n and therefore the following statement is evident. theorem 2. the mechanism of proportional allocation (20) is system compatible in the cppi-models with linear function of public payoff c(u) and any functions of private payoffs. example 3. suppose that gi(u) = ki √ ri − ui+sik ∑ j∈n uj , where is determined advances in systems science and application (2015) vol.15 no.1 53 by (20). then ∂gi ∂ui = ∂g0 ∂ui = k− ki 2 √ ri−ui , une i = umax i = { ri − k2i 4k2 , ki ≤ 2k √ ri, 0, otherwise; gmax 0 = ∑ j∈i kj √ rj+k ∑ j∈c′ rji = {i ∈ n : ki > 2k √ ri}, c ′ = {i ∈ n : ki ≤ 2k √ ri}. theoretical approach to building of system compatible economic impulsion mechanisms is based on germeier’s theorem for games of the type γ2 (see appendix). let’s apply this theorem to the linear model (5)-(6). we obtain sdi is arbitrary (g0 does not depend on s ), spi ≡ 0; li = kiri; ei = {ui = 0}; di = {(si, ui) : si > kiui k ∑ j∈n uj , n∑ i=1 si = 1}; k2 = gi0 = ∑ j∈n kjrj to find k1 we must solve an optimization problem g0(u) = ∑ j∈i(s) kj(rj − uj) +k ∑ j∈c(s) uj → max s.t. kiui k ∑ j∈n uj < si ≤ 1, ∑ j∈n sj = { 1, ∃i : si > 0, 0, otherwise, , 0 ≤ ui ≤ ri, i ∈ n from the first order condition ∂g0 ∂ui = k − ki ⇒ u∗i = { ri, ki ≤ k (set c0), 0, otherwise (set i0). it is proved that it is possible to find si from the set di: si∈c0 = kiri k ∑ j /∈i0 rj + εi, n∑ i=1 εi = 1− ∑ i/∈i0 kiri k ∑ j /∈i0 rj ; si∈i0 = 0. therefore, like in γ1 formulation (example 1), we obtain g0(u ∗) = ∑ j∈i0 kjrj +k ∑ j∈c0 rj 54 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... three cases are possible: 1. ∀i ∈ n , ki > k ⇒ u∗i = 0, k1 = gmax 0 = gi0 = ∑ j∈n kjrj , in this case k2 = k1, and gi = kiri = li . 2. ∀i ∈ n , ki ≤ k ⇒ u∗i = ri, k1 = gmax 0 = gc0 = k ∑ j∈n rj , that is k2 ≥ k1, and gi = kri > li, hence, condition si > kiui k ∑ j∈n uj is satisfied. in this case n = c0 = c(s), i = ∅, gmax 0 = gc0 = k ∑ j∈n rj , s∗, is any allocation satisfying (13) (“collectivistic” society). 3. ∃l,m : kl > k, km ≤ k. here the solution of the problem (12)-(13) is given by the following algorithm: to assign s∗i =  kiri k ∑ j /∈i0 rj + εi, ki ≤ k; 0, ki > k. then u∗i = { ri, ki ≤ k, 0, ki > k, and gmax 0 = k ∑ j∈i0 rj + ∑ j /∈i0 kjrj . in this case c0 = c(s) ̸= ∅, i0 = i(s) ̸= ∅ (≪mixed≪ society). in all cases 1-3 there is a system compatibility of the model (12)-(13) and economic impulsion mechanism s∗. gi(u) = ki √ ri − ui + sik ∑ j∈n uj → max, 0 ≤ ui ≤ ri, i ∈ n g0(u) = ∑ j∈n kj √ rj − uj +k ∑ j∈n uj → max, 0 ≤ si ≤ 1, ∑ j∈n sj = { 1,∃i : s > 0, 0, otherwise. we obtain sdi is arbitrary (g0 does not depend on s ),spi ≡ 0; li = ki √ ri; ei = {ui = 0}; di = {(si, ui) : si > kiui k( √ ri + √ ri − ui) ∑ j∈n uj }; k2 = gi0 = ∑ j∈n kj √ rj . to find k1 we must solve an optimization problem g0(u) = ∑ j∈n kj √ rj − uj +k ∑ j∈n uj → max ki( √ ri − √ ri − ui) k ∑ j∈n uj < si ≤ 1, ∑ j∈n sj = { 1, ∃i : si > 0, 0, otherwise, , 0 ≤ ui ≤ ri, i ∈ n advances in systems science and application (2015) vol.15 no.1 55 from the first order condition u∗i = { ri − ki 2 4k2 , ri ≤ 2k √ ri; 0, otherwise. it is proved that it is possible to find si from the set di: s∗i =  ki( √ ri− ki 2k ) k ∑ j /∈i0 ( rj− kj 2 4k2 ) + εi, ui > 0; 0, ui = 0. where n∑ i=1 εi = 1− ∑ i/∈i0 ki( √ ri− ki 2k ) k ∑ j /∈i0 ( rj− kj 2 4k2 ) ; therefore k1 = ∑ j∈i0 kj √ rj + ∑ j /∈i0 ( kri + k2i 4k ) > k2. the following cases are possible. 1. ∀i ∈ n , ki > 2k √ ri ⇒ u∗i = 0, k1 = gmax 0 = gi0 = ∑ j∈n kj √ rj , in this case k2 = k1, and gi = ki √ ri = li. 2. ∀i ∈ n , ki ≤ 2k √ ri ⇒ u∗i = ri, k1 = gmax 0 = ∑ j∈n ( kri + k2i 4k ) , i.e. k2 ≤ k1, li = ki √ ri; ei = {ui = 0}; di = { (si, ui) : si > kiui k( √ ri+ √ ri−ui) ∑ j∈n uj } ; and gi = kri + k2i 4k > li, hence, condition si > kiui k( √ ri+ √ ri−ui) ∑ j∈n uj is satisfied. in this case n = c ′, i0 = ∅, ∑ j∈n ( kri + k2i 4k ) , s∗ is any allocation satisfying (13) (“collectivistic” society). 3. ∃l,m : kl > 2k √ rl, km ≤ ki > 2k √ rm. here the solution of the problem (12)-(13) is given by the following algorithm: to assign s∗i =  ki( √ ri− ki 2k ) k ∑ j /∈i0 ( rj− kj 2 4k2 ) i , ki ≤ 2k √ ri; 0, ki > k. then u∗i = { ri − ki 2 4k2 , ri ≤ 2k √ ri; 0, otherwise. and gmax 0 = ∑ j∈i0 kj √ rj + ∑ j /∈i0 ( kri + k2i 4k ) . in this case c0 = c(s) ̸= ∅, i0 = i(s) ̸= ∅ (≪mixed≫ society). in all cases 1-3 there is a system compatibility of the model (12)-(13) and economical impulsion mechanism s∗. 5 administrative control mechanisms. suppose that the principal can bound the agents’ sets of feasible controls. consider the case of administrative control without a feedback. then the model (5)-(6) takes the form gi(q̃i, q̄i, u) = pi(ri − ui) + sic(u) → max, q̃i ≤ ui ≤ q̄i, si ∈ [0, 1]; (21) 56 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... g0(q̃, q̄, u) = ∑ j∈n pj(rj − uj) + c(u) → max, 0 ≤ q̃i ≤ q̄i ≤ ri, i ∈ n. (22) it is clear that if there are no restrictions then the problem (22) has a trivial solution q̃i = q̄i = umax i , i ∈ n . therefore the principals control cost should be considered. then (22) takes the form g0(q̃, q̄, u) = ∑ j∈n pj(rj − uj) + c(u)− c(q̃, q̄) → max, 0 ≤ q̃i ≤ q̄i ≤ ri, i ∈ n, where c(q̃, q̄) is a continuously differentiable and convex in all arguments compulsion cost function. let’s give a simple example. example 4. assume that gi(q̃i, q̄i, u) = ki(ri − ui) + sik ∑ j∈n uj → max, q̃i ≤ ui ≤ q̄i, si ∈ [0, 1]; g0(q̃, q̄, u) = ∑ j∈n kj(rj − uj) +k ∑ j∈n uj − ∑ j∈n (m̃j q̃j + m̄j q̄j) → max, 0 ≤ q̃i ≤ q̄i ≤ ri, i ∈ n where k, ki, m̃i, m̄i are known positive constants. the foc give ∂gi ∂ui = sik − ki ⇒ une i = { ri, ki ≤ sik, 0, ki > sik; ∂g0 ∂ui = k − ki ⇒ umax i = { ri, ki ≤ k, 0, ki > k; ∂g0 ∂q̃i = −m̃i, ∂g0 ∂q̄i = −m̄i ⇒ q̃max i = q̄max i = 0 notice that if ki > k then ki > sik, therefore umax i = une i = 0, and compulsion is not required (the model is system compatible). otherwise two cases should be differentiated: (a) ki ≤ sik ⇒ umax i = une i = ri, and compulsion is not required again; (b) sik < ki ≤ k ⇒ une i = 0, umax i = ri . in this case the principal’s payoffs with and without compulsion should be compared. denote by m the set of agents for whom sik < ki ≤ k. if compulsion holds we have g0(r, 0, r) = k ∑ j∈m rj − ∑ j∈m m̃jrj , and if not then g0(0, 0, 0) = ∑ j∈m kjrj . thus, compulsion is rational if ∑ j∈m (k − m̃j − kj)rj > 0 holds. advances in systems science and application (2015) vol.15 no.1 57 6 conclusion. in the present paper the problem of system compatibility was analyzed. its solution ensures maximization of the social welfare by considering individual interests of the agents. this setting is close to the problem of mechanism design but distinct because in the latter problem it is required to motivate agents to report the true information about their types instead of the direct maximization of the social welfare. the conditions of system compatibility are quite restrictive, therefore to provide them it is worthwhile to construct control mechanisms. in the framework of authors’ concept, the mechanisms are classified by two attributes: direction of impact (sets of feasible controls of the agents or their payoff functions) and presence or absence of feedback in the control system. the first attribute differentiates administrative and economic control mechanisms (methods of compulsion and impulsion respectively), while the second one leads to hierarchical games of the types γ1 and γ2. the conditions and mechanisms of system compatibility are analyzed in the class of models of concordance of the private and public interests in resource allocation (cppi-models). some preliminary results about system compatibility of the cppi-models are obtained. thus, system compatibility in cppi-models when si = const is reachable only if all agents are pure individualists (all resources are assigned for private interests) or pure collectivists (all resources are assigned for public interest). the exact dichotomous partition is built by specific algorithms of discrete optimization. economic mechanisms with a feedback simplify the achievement of system compatibility. administrative mechanisms of system compatibility are under development. the research perspectives include: investigation of the system compatibility for more general classes of models; considering of corruption (an additional feedback on bribe); analysis of dynamic settings, including phase constraints (requirements of sustainable development), investigation of the conditions of time consistence. appendix (germeier theorem). assume that payoff functions of both players m1(x1, x2),m2(x1, x2) are continuous on compact sets of feasible controls x1, x2. introduce the punishment function xp1 (x2) such that m2(x p 1 , x2) = min x1∈x1 m2(x1, x2), and the dominant strategy of the player 1 xd1 (x2), which satisfies the condition m1(x d 1 , x2) = max x1∈x1 m(x1, x2). introduce also the following values and sets: l2 = max x2∈x2 m2(x p 1 (x2), x2); e2 = {x2 ∈ x2 : m2(x p 1 (x2), x2) = l2}; d2 = {(x1, x2) ∈ x1 ×x2 : m2(x1, x2) > l2}; 58 olga i. gorbaneva and guennady a. ougolnitsky: system compatibility,price of anarchy and... k1 = sup (x1,x2)∈d2 m1(x1, x2) ≤ m1(x ε 1, x ε 2) + ε (d2 = ∅ ⇒ k1 = −∞); k2 = min x2∈e2 max x1∈x1 m1(x1, x2). then the guaranteed payoff of the player 1 (leader) in the game γ2 (in which the first player knows the choice of the second player) is equal to w1 = max(k1,k2), and the respective -optimal guaranteeing strategy has the form x̃ε1(x2) =  xε1, x2 = xε2, k1 > k2, xd1 (x2), x2 ∈ d2, k1 ≤ k2, xp1 (x2), otherwise, where xε1 and xε2 are described above. acknowledgements the work is supported by the russian foundation for basic research, project # 15-01-00432. references [1] n. nisan, t. roughgarden, e. tardos and v. vazirani. (2007), algorithmic game theory, cambridge university press. [2] papadimitriou c.h. (2001), “algorithms, games, and the internet”, proc. 33th symposium theory of computing, pp.749-753. [3] germeier yu.b. and vatel i.a. (1975), “equilibrium situations in games with a hierarchical structure of the vector of criteria” ,lecture notes in computer science, vol.27, pp.460-465. [4] gorbaneva o.i. and ougolnitsky g.a. (2013), “purpose and non-purpose resource use models in two-level control systems”, advances in systems science and application ,vol.13,no.4,pp.378-390. [5] ougolnitsky g.a. (2011), “games with differently directed interests”, contributions to game theory and management, vol.4,pp.327-338.[collected papers presented on the fourth international conference game theory and management / editors l. petrosyan, n. zenkevich.] [6] laffont j.-j. and martimort d. (2002), the theory of incentives: the principal-agent model, princeton university press. [7] kononenko a.f. (1974), “game-theory analysis of a two-level hierarchical control system”, ussr computational mathematics and mathematical physics,vol.14,no.5,pp.72-81. advances in systems science and application (2015) vol.15 no.1 59 [8] kukushkin n.s. (1994), “a condition for existence of nash equilibrium in games with public and private objectives”, games and economic behavior,vol.7,pp.177-192. [9] d. novikov. (2013), mechanism design and management: mathematical methods for smart organizations, nova science publishers. [10] novikov d. (2013), theory of control in organizations, nova science publishers. [11] burkov v.n. and opoitsev v.i. (1974), “metagame approach to the control in hierarchical systems”, automation and remote control, vol 35, no.1, pp.93-103. [12] ougolnitsky g. (2011), sustainable management, nova science publishers. corresponding author olga i. gorbaneva, ph.d. can be contacted at: gorbaneva@mail.ru guennady a. ougolnitsky, ph.d. can be contacted at:ougoln@gmail.com advances in systems science and applications (2012) vol.12 no.4 373-387 predictive inferences for a future number of failures coming from underlying models under parametric uncertainty konstantin n. nechval1, nicholas a. nechval2, maris purgailis2, uldis rozevskis2, vladimir f. strelchonok3 and max moldovan4 1applied mathematics department, transport and telecommunication institute lomonosov street 1, lv-1019, riga, latvia 2statistics department, evf research institute, university of latvia raina blvd 19, lv-1050, riga, latvia 3informatics department, baltic international academy lomonosov street 4, lv-1019, riga, latvia 4australian institute of health innovation, university of new south wales, level 1 agsm building, sydney nsw 2052, australia abstract in this paper, we present an accurate procedure to obtain prediction limits for the number of failures that will be observed in a future inspection of a sample of units, based only on the results of the first in-service inspection of the same sample. the failure-time of such units is modeled with a two-parameter weibull distribution indexed by scale and shape parameters β and δ, respectively. it will be noted that in the literature only the case is considered when the scale parameter β is unknown, but the shape parameter δ is known. as a rule, in practice the weibull shape parameter δ is not known. instead it is estimated subjectively or from relevant data. thus its value is uncertain. this δ uncertainty may contribute greater uncertainty to the construction of prediction limits for a future number of failures. in this paper, we consider the case when both parameters β and δ,are unknown. in literature, for this situation, usually a bayesian approach is used. bayesian methods are not considered here. we note, however, that although subjective bayesian prediction has a clear personal probability interpretation, it is not generally clear how this should be applied to non-personal prediction or decisions. objective bayesian methods, on the other hand, do not have clear probability interpretations in finite samples. the technique proposed here for constructing prediction limits emphasizes pivotal quantities relevant for obtaining ancillary statistics. and represents a special case of the method of invariant embedding of sample statistics into a performance index. two versions of prediction limits for a future number of failures are given. keywords weibull distribution, parametric uncertainty, future number of failures, prediction limits 374 konstantin n. nechval:predictive inferences for a future number of failures coming from... 1 introduction this paper extends the results of nelson [1]. nelson’s prediction limits were motivated by the following application. nuclear power plants contain large heat exchangers that transfer energy from the reactor to steam turbines. such exchangers typically have 10,000 to 20,000 stainless steel tubes that conduct the flow of steam. due to stress and corrosion, the tubes develop cracks over time. cracks are detected during planned inspections. the cracked tubes are subsequently plugged to remove them from service. to develop efficient inspection and plugging strategies, plant management can use a prediction of the added number of tubes that will need plugging by a specified future time. nelson presents simple prediction limits for the number of failures that will be observed in a future inspection of a sample of units. the past data consist of the cumulative number of failures in a previous inspection of the same sample of units. life of such units is modeled with a weibull distribution with a given shape parameter value. prediction of an unobserved random variable is a fundamental problem in statistics. hahn and nelson [2], patel [3], and hahn and meeker [4] provided surveys of methods for statistical prediction for a variety of situations on this topic. in the areas of reliability and life-testing, this problem translates to obtaining prediction intervals for lifetime distributions. nordman and meeker [5] compared probability ratio, simplified probability ratio and likelihood ratio methods proposed by nelson [1], assuming known the weibull shape parameter δ. in this paper, we use a frequentist procedure, which is called ‘within-sample prediction of future order statistics’, when the time-to-failure follows the twoparameter weibull distribution indexed by scale and shape parameters β and δ. we consider the case when both parameters β and δ are unknown. the technique proposed here for constructing prediction limits emphasizes pivotal quantities relevant for obtaining ancillary statistics and represent a special case of the method of invariant embedding of sample statistics into a performance index applicable whenever the statistical problem is invariant under a group of transformations, which acts transitively on the parameter space (nechval et al. [6-7]). conceptually, it is useful to distinguish between “new-sample” prediction, “within-sample” prediction, and “new-within-sample” prediction. some mathematical preliminaries for the within-sample prediction are given below. 2 mathematical preliminaries for within-sample prediction theorem 1 let x1 ≤ . . . ≤ xk be the first k ordered observations (order statistics) in a sample of size m from a continuous distribution with some probability density functio fθ(x) and distribution function fθ(x), where θ is a parameter (in general, vector). then the joint probability density function of x1 ≤ . . . ≤ xk advances in systems science and applications (2012) vol.12 no.4 375 and the lth order statistics xl(1 < k < l < m) is given by gθ(x1, . . . , xk, xl) = gθ(x1, . . . , xk)gθ(xl|xk), (1) where gθ(x1, . . . , xk) = m! (m− k)! πk i=1fθ(xi)[1− fθ(xk)] m−k, (2) gθ(xl|xk) = (m− k)! (l − k − 1)!(m− l)! [ fθ(xl)− fθ(xk) 1− fθ(xk) ]l−k−1[1− fθ(xl)− fθ(xk) 1− fθ(xk) ]m−l fθ(xl) 1− fθ(xk) = (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j [ 1− fθ(xl) 1− fθ(xk) ]m−l+j fθ(xl) 1− fθ(xk) = (m− k)! (l − k − 1)!(m− l)! m−lx j=0 � m− l j � (−1)j [ fθ(xl)− fθ(xk) 1− fθ(xk) ]l−k−1+j fθ(xl) 1− fθ(xk) (3) represents the conditional probability density function of xl given xk = xk. proof.the joint density of x1 ≤ . . . ≤ xk and xl is given by gθ(x1, . . . , xk, xl) = (m)! (l − k − 1)!(m− l)! ky i=1 fθ(xi)[fθ(xl)− fθ(xk)] l−k−1fθ(xl) [1− fθ(xl)] m−l = gθ(x1, . . . , xk)gθ(xl|xk). (4) it follows from (4) that gθ(xl|x1, . . . , xk) = gθ(x1, . . . , xk, xl) gθ(x1, . . . , xk) = gθ(xl|xk), (5) i.e., the conditional distribution of xl l, given xi = xi for all i = 1, . . . , k , is the same as the conditional distribution of xl , given only xk = xk,which is given by (3). this ends the proof. corollary 1.1.the conditional probability distribution function of xl given xk = xk is pθ(xl ≤ xl|xk = xk) =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h 1− fθ(xl) 1− fθ(xk) im−l+1+j = (m− k)! (l − k − 1)!(m− l)! m−lx j=0 � m− l j � (−1)j l − k + j hfθ(xl)− fθ(xk) 1− fθ(xk) il−k+j . (6) 376 konstantin n. nechval:predictive inferences for a future number of failures coming from... corollary 1.2. let x1 ≤ . . . ≤ xk be the first k order statistics in a sample of size m from the two-parameter weibull distribution with the probability density function fθ(x) = δ β ( x β )δ−1 exp[−( x β )δ] (x > 0), (7) where θ = (β, σ),β > 0 and σ > 0 are the scale and shape parameters, respectively. then the conditional probability distribution function of xl given xk = xk is pθ{xl ≤ xl|xk = xk} = 1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h exp � − xδl − xδk βδ �im−l+1+j . (8) theorem 2 if in (8) the scale parameter is unknown, then the predictive probability distribution function of xl based on (xk, δ) is given by pδ n�xl xk �δ ≤ � xl xk �δ } = 1− m! (l − k − 1)!(m− l)! × � l − k − 1 j � (−1)j m− l + 1 + j � πk−1 s=0 h� xl xk �δ − 1)(m− l + 1 + j) + (m− k + 1 + s) i�−1 . (9) proof.we reduce (8) to pθ n�xl xk �δ ≤ �xl xk �δ | �xk β �δ = �xk β �δo =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + jh exp(−ω[νδ − 1]) im−l+1+j =pδ{v δ ≤ νδ|w = ω}, (10) where v = xl/xk is the ancillary statistic whose distribution does not depend on the parameter β. since xk does not depend on v , w = (xk/β) δ is the pivotal quantity, whose distribution is known and does not depend on the parameters β and δ, we eliminate the parameter from the problem as pδ{xl ≤ xl} = z ∞ 0 pθ{xl ≤ xl|xk = xk}gθ(xk)dxk, (11) where gθ(xk) = m! (k − 1)!(m− k)! f k−1 θ (xk) h 1− fθ(xk) im−k fθ(xk), xk ∈ (0,∞), (12) advances in systems science and applications (2012) vol.12 no.4 377 represents the probability density function of the kth order statisticxk k. indeed, it follows from (12) that gθ(xk)dxk = m! (k − 1)!(m− k)! h 1− exp � − �xk β �δ�ik−1 exp � − �xk β �δ(m−k) � exp � − �x β �δ� d �x β �δ = m! (k − 1)!(m− k)! [1− e−ω]k−1e−ω(m−k+1)dω = g(ω)dω. (13) it follows from (10) and (13) that pδ{v δ ≤ νδ} = z ∞ 0 pδ{v δ ≤ νδ|w = ω}g(ω)dω =1− (m)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j � πk−1 s=0 � (νδ − 1)(m− l + 1 + j) + (m− k + 1 + s) ��−1 . (14) now (9) follows from (14). this ends the proof. corollary 2.1.if the parameter δ = 1 , i.e. we deal with the exponential distribution, then the predictive probability distribution function of xl based on xk is given by p n�xl xk � ≤ � xl xk �o = 1− m! (l − k − 1)!(m− l)! × l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j � πk−1 s=0 h� xl xk − 1 � (m− l + 1 + j) + (m− k + 1 + s) i�−1 . (15) theorem 3 let x1 ≤ . . . ,≤ xk be the first k ordered observations from a sample of size m from the two-parameter weibull distribution (7). then the joint probability density function of the pivotal quantities w2 = δóδ , w3 = � óβ β �δ̂ , (16) conditional on fixed zk = (zi, . . . , zk),where zi = (xi/óβ)δ̂, i = 1, . . . , k,are ancillary statistics, any k − 2 of which form a functionally independent set,óβ and óδ are the estimators of β and δ,based on the first k ordered observations (x1 ≤ . . . ≤ xk) from a sample of size m from the two-parameter weibull distribution (7), such that w2 and w3 are the pivotal quantities (in particular, the 378 konstantin n. nechval:predictive inferences for a future number of failures coming from... maximum likelihood estimators of β and δ, óβ = �h kx i=1 xδ̂i + (m− k)xδ̂k i /k �1/δ̂ (17) and óδ = �� kx i=1 xδ̂i lnxi+(m−k)xδ̂k lnxk �� kx i=1 xδ̂i +(m−k)xδ̂k �−1 − 1 k kx i=1 lnxi �−1 (18) respectively, lead to the pivotal quantities w2 and w3 )is given by f(ω2, ω3|z(k)) =ϑ•(z(k))ωk−1 2 ky i=1 zω2 i ωkω2−1 3 exp � − ωω2 3 � kx i=1 zω2 i + (m− k)zω2 k �� =ϑ•(z(k))ωk−2 2 ky i=1 zω2 i ω ω2(k−1) 3 exp � − ωω2 3 � kx i=1 zω2 i + (m− k)zω2 k �� ω2ω ω2−1 3 =f(ω2|z(k))f(ω3|ω2, z (k)), ω3 ∈ (0,∞), (19) where ϑ•(z(k)) = h z ∞ 0 γ(k)ωk−2 2 ky i=1 zω2 i � kx i=1 zω2 i + (m− k)zω2 k �−k dω2 i−1 (20) is the normalizing constant, f(ω2|z(k)) = ϑ(z(k))ωk−2 2 ky i=1 zω2 i � kx i=1 zω2 i + (m− k)zω2 k �−k , ω2 ∈ (0,∞), (21) ϑ(z(k)) = h z ∞ 0 ωk−2 2 ky i=1 zω2 i � kx i=1 zω2 i + (m− k)zω2 k �−k dω2 i−1 , (22) f(ω3, ω2|z(k)) = hpk i=1 z ω2 i + (m− k)zω2 k ik γ(k) ω w2(k−1) 3 × exp � − ωw2 3 h ( kx i=1 zω2 i + (m− k)zω2 k i� ω2ω ω2−1 3 , ω3 ∈ (0,∞). (23) proof.the joint density x1 ≤ . . . ≤ xk is given by fθ(x1, . . . , xk) = m! (m− k)! ky i=1 δ β ( xi β )δ−1 exp(−( xi β )δ) exp(−(m− k)(( xk β )δ). (24) advances in systems science and applications (2012) vol.12 no.4 379 using óβ and óδ (the maximum likelihood estimators of β and δ obtained from solution of (17) and (18)) and the invariant embedding technique [8-14], we transform (24) as follows: fθ(x1, . . . , xk)dóβdóδ = m! (m− k)! ky i=1 x−1 i δk ky i=1 �xi β �δ exp � − kx i=1 �xi β �δ − (m− k) �xk β �δ� dóβdóδ =− m! (m− k)! óβóδk ky i=1 x−1 i �δóδ �k−2 ky i=1 �xióβ �δ̂( δ δ̂ )� óβ β �δ̂( δ δ̂ )(k−1) × exp � − � óβ β �δ̂( δ δ̂ ) � kx i=1 �xióβ �δ̂( δ δ̂ ) + (m− k) �xkóβ �δ̂( δ δ̂ )���óδ( δóδ ) β � óβ β �δ̂( δ δ̂ )−1 dóβ��− δóδ2dóδ � =− m! (m− k)! óβóδk ky i=1 x−1 i ωk−2 2 ky i=1 zω2 i ω ω2(r−1) 3 exp � − ωω2 3 h kx i=1 zω2 i + (m− k)zω2 k i� d(ωω2 3 )dω2 =− m! (m− k)! óβóδk ky i=1 x−1 i ωk−2 2 ky i=1 zω2 i ω ω2(k−1) 3 exp � − ωω2 3 h kx i=1 zω2 i + (m− k)zω2 k i� ω2ω ω2−1 3 dω2dω3. (25) normalizing (25), we obtain (19). this ends the proof. it will be noted that more general case of distributions indexed by location and scale parameters has been considered in [15]. theorem 4 if in (8) both parameters β and δ are unknown, then the predictive probability distribution function of xl based on (xk, óδ) and conditional on fixed z(k) is given by p n�xl xk �δ̂ ≤ � xl xk �δ̂|z(k)o =1− m! (l − k − 1)!(m− l)! × z ∞ 0 l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j� k−1y s=0 h��� xl xk �δ̂�ω2 − 1 � (m− l + 1 + j) + (m− k + 1 + s) i�−1 f � ω2|z(k) � dω2. (26) 380 konstantin n. nechval:predictive inferences for a future number of failures coming from... proof. we reduce (9) to pδ n�xl xk �δ̂( δ δ̂ ) ≤ � xl xk �δ̂( δ δ̂ ) o = 1− m! (l − k − 1)!(m− l)! × l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j � k−1y s=0 h�� xl xk �δ̂( δ δ̂ ) − 1 � (m− l + 1 + j) + (m− k + 1 + s) i�−1 =1− m! (l − k − 1)!(m− l)! × l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j � k−1y s=0 h (νω2 2 − 1)(m− l + 1 + j) + (m− k + 1 + s) i�−1 =p ¦ v w2 2 ≤ νω2 2 © , (27) where v2 = (xl/xk) δ̂ is the ancillary statistic whose distribution does not depend on the parameters β and δ. since the pivotal quantity w2, whose distribution is given by (21), does not depend on v2, it follows from (21) and (27) that p ¦ v2 ≤ ν2|z(k) © = z ∞ 0 p ¦ v w2 2 ≤ νω2 2 © f � ω2|z(k) � dω2, (28) where the unknown parameters β and δ are eliminated from the problem. now (26) follows from (28). this ends the proof. 3 prediction limits for a future number of failures consider the situation in which m units start service at time 0 and are observed until a time tc when the available weibull failure data are to be analyzed. failure times are recorded for the k units that fail in the interval [0, tc]. then the data consist of the k smallest-order statistics x1 ≤ . . . ≤ xk ≤ tc and the information that the other m?k units will have failed after tc.with time (or type i) censored data, tc is prescribed and k is random. with failure (or type ii) censored data, k is prescribed and tc = xk is random. the problem of interest is to use the information obtained up to tc to construct the weibull within-sample prediction limits (lower and upper) for the number of units that will fail in the time interval [tc, tω].for example, this tω could be the end of a warranty period. consider the situation when tc = xk. under conditions of theorem 4, the lower prediction limit for the number of units that will fail in the time interval [tc, tω] is given by llower = lmax − k, (29) advances in systems science and applications (2012) vol.12 no.4 381 where lmax = max k tω|zk © ≤ α � (30) the upper prediction limit for the number of units that will fail in the time interval [tc, tω] is given by lupper = lmin − k − 1, (31) where lmin = min k tω|zk © ≥ 1− α � (32) in the above case, where both parameters β and δ are unknown, the prediction limits (lower and upper) for the number of units that will fail in the time interval [tc, tω] are based on (xk, óδ) and conditional on fixed z(k). if l, which satisfies (30), does not exist then lmax = k and the lower prediction limit for the number of units that will fail in the time interval [tc, tω] is given by llower = lmax − k = 0. (33) if l, which satisfies (32), does not exist then lmin = m+1 and upper prediction limit for the number of units that will fail in the time interval [tc, tω] is given by lupper = lmin − k − 1 = m− k, (34) 4 second version of prediction limits for a future number of failures in this section, we wish to show how to obtain the second version of prediction limits for a future number of failures. the methodology is based on the following results. theorem 5 let x1 ≤ . . . ≤ xk be the first k ordered observations from a sample of size m from the two-parameter weibull distribution (7). then the joint probability density function of the pivotal quantities w1 = � óβ β �δ , w3 = δóδ , (35) conditional on fixed z(k) = (zi, . . . , zk), where zi = (xi/óβ)δ̂, i = 1, . . . , k are ancillary statistics, any k−2 of which form a functionally independent set,óβ andóδ are, for instance, the maximum likelihood estimators for β and δ based on the first k ordered observations (x1 ≤ . . . ≤ xk) from a sample of size m from the 382 konstantin n. nechval:predictive inferences for a future number of failures coming from... two-parameter weibull distribution (7), which can be found from solution of (17) and (18), is given by f(ω1, ω2|z(k)) = ϑ•(z(k))ωk−2 2 ky i=1 zω2 i ωk−1 1 exp(−ω1[ kx i=1 zω2 i + (m− k)zω2 k ]) = f(ω2|z(k))f(ω1|ω2, z (k)), ω1 ∈ (0,∞), ω2 ∈ (0,∞), (36) where ϑ•(z(k)) = � z ∞ 0 γ(k)ωk−2 2 ky i=1 zω2 i � kx i=1 zω2 i + (m− k)zω2 k �−k dω2 �−1 (37) is the normalizing constant, f(ω2|z(k)) is given by (21), f(ω1, ω2|z(k)) = �pk i=1 z ω2 i + (m− k)zω2 k �k γ(k) ωk−1 1 exp � − ω1 � kx i=1 zω2 i + (m− k)zω2 k ) �� ω1 ∈ (0,∞), (38) proof. the joint density of x1 ≤ . . . ≤ xk is given by fθ(x1, . . . , xk) = m! (m− k)! ky i=1 δ β ( xi β )δ−1 exp(−( xi β )δ) exp(−(m− k)( xk β )δ). (39) using the invariant embedding technique [8-14], we transform (39) to fθ(x1, . . . , xk)dóβdóδ = m! (m− k)! ky i=1 x−1 i δk ky i=1 �xi β �δ exp � − kx i=1 �xi β �δ − (m− k) �xk β �δ� dóβdóδ =− m! (m− k)! óβóδk ky i=1 x−1 i �δóδ �k−2 ky i=1 �xióδ �δ̂( δ δ̂ )� óβ β �δ(k−1) × exp � − � óβ β �δ h kx i=1 ( xióδ ) δ̂( δ δ̂ ) + (m− k) �xkóδ �δ̂( δ δ̂ ) i�� δ β � óβ β �δ(k−1) dóβ�(− δóδ2 )dóδ2 =− m! (m− k)! óβóδk ky i=1 x−1 i ωk−2 2 ky i=1 x−1 i zω2 i ωk−1 1 exp � − ω1 h kx i=1 zω2 i + (m− k)zω2 k i� dω1ω2. (40) normalizing (40), we obtain (36). this ends the proof. corollary 5.1. if the parameter δ is known then w1 ∼ f(ω1) = kk γ(k) ωk−1 1 exp(−ω1k), ω1 ∈ (0,∞). (41) advances in systems science and applications (2012) vol.12 no.4 383 theorem 6 if in (8) the scale parameter β is unknown, then the predictive probability distribution function of xl based on (óβ, δ) and conditional on fixed xk is given by pδ{xl ≤ xl|xk = xk} =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h 1− (m− l + 1 + j) xδl − xδk k óβδ �−k (42) proof.we reduce (8) to pθ{xl ≤ xl|xk = xk} =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h exp � − � óβ β �δ xδl − xδkóβδ �im−l+1+j =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h exp � − ω1 xδl − xδkóβδ �im−l+1+j . (43) now, we eliminate the unknown parameter β from the problem and find (42) as pδ{xl ≤ xl|xk = xk} = z ∞ 0 pθ{xl ≤ xl|xk = xk}f(ω1)dω1. (44) this ends the proof. corollary 6.1.if the parameter δ = 1, i.e. we deal with the exponential distribution, then the predictive probability distribution function of xl based on óβ and conditional on fixed xk is given by xk p{xl ≤ xl|xk = xk} = 1− 1 b(l − k,m− l + 1) l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h 1 + (m− l + 1 + j) xl − xk k óβ i−k , (45) where k óβ = kx i=1 xi + (m− k)xk. (46) theorem 7 if in (8) both parameters β and δ are unknown, then the predictive probability distribution function of xl based on (wideparenβ,wideparenδ) and 384 konstantin n. nechval:predictive inferences for a future number of failures coming from... conditional on fixed xk and z(k) is given by pθ{xl ≤ xl|xk = xk; z (k)} = 1− (m− k)! (l − k − 1)!(m− l)! × z ∞ 0 l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j h 1 + (m− l + 1 + j) ��xlóβ �δ̂ω2 − �xkóβ �δ̂ω2 �� kx i=1 zω2 i + (m− k) zω2 i �−1i−k × f(ω2|z(k))dω2. (47) proof. we reduce (8) to pθ{xl ≤ xl|xk =xk} = 1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j � exp � − � óβ β �δh�xlóβ �δ̂( δ δ̄ ) − �xkóβ �δ̂( δ δ̄ ) i��m−l+1+j =1− (m− k)! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j exp h� − ω1 ��xlóβ �δ̂ω2 − �xkóβ �δ̂ω2 ��im−l+1+j . (48) now, we eliminate the unknown parameters β and δ from the problem and find (47) as pθ{xl ≤ xl|xk = xk; z (k)} = z ∞ 0 z ∞ 0 pθ{xl ≤ xl|xk = xk}f(ω1, ω2|z(k))dω1dω2 = z ∞ 0 z ∞ 0 pθ{xl ≤ xl|xk = xk}f(ω1|ω2, z (k))f(ω2|z(k))dω1dω2. (49) this ends the proof. under conditions of theorem 7, the lower prediction limit for the number of units that will fail in the time interval [tc, tω] is given by llower = lmax − k, (50) where lmax = max k tω|xk = xk; z (k))} ≤ α � , (51) advances in systems science and applications (2012) vol.12 no.4 385 the upper prediction limit for the number of units that will fail in the time interval [tc, tω] is given by lupper = lmin − k − 1, (52) lmin = min k tω|xk = xk; z (k))} ≥ 1− α � , (53) in the above case, when both parameters β and δ are unknown, the prediction limits (lower and upper) for the number of units that will fail in the time interval [tc, tω] are based on (óβ, óδ) and conditional on fixed xk,z (k). if l, which satisfies (51), does not exist then lmax = k and the lower prediction limit for the number of units that will fail in the time interval [tc, tω] is given by llower = 0. if l, which satisfies (53), does not exist then lmin = m+ 1 and upper prediction limit for the number of units that will fail in the time interval [tc, tω] is given by lupper = m− k. 5 numerical example for the sake of simplicity, but without loss of generality, we consider (for illustration) the special case of theorem 2 where m = 40 items simultaneously tested have life times, which follow the weibull distribution with δ = 1 . in other words, we deal with the exponential distribution. two items have failed by the inspection at times, x1 = 45 and x2 = 100 hours. let us assume that the situation takes place when tc = xk = 100 hours, where k = 2. suppose, say, tω = 450 hours. taking into account (15), we find the lower prediction limit for the number of units that will fail in the time interval [tc, tω] as llower = lmax − k = 3− 2 = 1, (54) lmax = max k tω} ≤ α � = 3, α = 0.05 (55) p{xl > tω} = m! (l − k − 1)!(m− l)! l−k−1x j=0 � l − k − 1 j � (−1)j m− l + 1 + j� k−1y s=0 h� tω xk − 1 � (m− l + 1 + j) + (m− k + 1 + s) i�−1 , (56) the upper prediction limit for the number of units that will fail in the time interval [tc, tω] is given by lupper = lmin − k − 1 = 17− 2− 1 = 14, (57) lmin = min k tω|xk = xk; z (k))} ≥ 1− α � = 17. (58) 386 konstantin n. nechval:predictive inferences for a future number of failures coming from... it will be noted that when both parameters β and δ are unknown, the lower and upper prediction limits for the number of units that will fail in the time interval [tc, tω] can be found either from (29) and (31), which are based on (xk, óδ),or from (50) and (52), which are based on (óβ, óδ). conclusion and future work the methodology described here can be extended in several different directions to handle various problems that arise in practice. we have illustrated the prediction method for log-location-scale distributions (such as the weibull or exponential distributions). application to other distributions could follow directly. acknowledgements this research was supported in part by grant no. 06.1936, grant no. 07.2036, grant no. 09.1014, and grant no. 09.1544 from the latvian council of science and the national institute of mathematics and informatics of latvia. references [1] nelson w. (2000), “weibull prediction of a future number of failures”, quality reliability engineering international, vol.16, pp.23-26. [2] hahn g.j and nelson w. (1973), “a survey of prediction intervals and their applications”, journal of quality technology, vol.5, pp.178-188. [3] patel j.k. (1989), “prediction intervals-a review”, communications in statistics. theory and methods, vol.18, pp.2393-2465. [4] hahn g.j and meeker w.q. (1991), statistical intervals: a guide for practitioners, new york: wiley. [5] nordman d.j and meeker w.q. (2002), “weibull prediction for a future number of failures”, technometrics, vol.44, pp.15-23. [6] nechval n.a, nechval k.n and vasermanis e.k. (2003), “effective state estimation of stochastic systems”, kybernetes (an international journal of systems & cybernetics), vol.32, pp.666-678 1218-1224. [7] nechval n.a, purgailis m, berzins g, cikste k, krasts j, and nechval k.n. (2010), “invariant embedding technique and its applications for improvement or optimization of statistical decisions”, in: k. al-begain, d. fiems and w. knottenbelt (eds.), analytical and stochastic modeling techniques and applications, lncs, vol.6148, berlin, heidelberg: springer-verlag, pp.306320. advances in systems science and applications (2012) vol.12 no.4 387 [8] nechval n.a, nechval k.n and purgailis m. (2011), ‘prediction of future values of random quantities based on previously observed data”, engineering letters, vol.9, pp.346-359. [9] nechval n.a and purgailis m. (2010), “improved state estimation of stochastic systems via a new technique of invariant embedding”, in: chris myers (ed.) stochastic control, publisher: sciyo, croatia, india, pp.167-193. [10] nechval n. a, nechval k. n, purgailis m, strelchonok v. f. (2011), “planning inspections in the case of damage tolerance approach to service of fatigued aircraft structures”, international journal of performability engineering, vol.7, pp.279-290. [11] nechval n.a, purgailis m, nechval k.n, and strelchonok v.f. (2012), “optimal predictive inferences for future order statistics via a specific loss function”, iaeng international journal of applied mathematics, math.1180, springer, berlin, vol.42, pp.40-51. [12] nechval n.a, purgailis m, nechval k.n, and bruna i. (2012), “optimal inventory control under parametric uncertainty via cumulative customer demand”, in:lecture notes in engineering and computer science: proceedings of the world congress on engineering, wce vol.i, pp.4-6 july, london, u.k, pp.6-11. [13] nechval n.a, purgailis m, nechval k.n and bruna i. (2012), “optimal prediction intervals for future order statistics from extreme value distributions”, in: lecture notes in engineering and computer science: proceedings of the world congress on engineering 2012, wce 2012, vol.iii, pp.4-6 july, london, u.k, pp.1340-1345. [14] nechval n.a and purgailis m. (2012), “stochastic control and improvement of statistical decisions in revenue optimization systems”, in:stochastic control, ivan ganchev ivanov (ed.), croatia, india, publisher: sciyo, pp.151-176. [15] paramonov yu.m (1992), methods of mathematical statistics in problems on the estimation and maintenance of fatigue life of aircraft structures, (in russian), riga: riiga. corresponding author konstantin n. nechval can be contacted at:e-mail: konstan@tsi.lv мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 1-12 analysis method of structural equation modeling a. v. burkov 1 , e. a. murzina 2 1 applied statistics and informatics department, mari state university, lenina 1, yoshkar-ola, 424002, mari el, russia 2 department of management and law, volga state university of technology, lenina 3, yoshkar-ola, 424002, mari el, russia annotation the goal of this article is to define a set of factors which most strongly influence a probability of finding (opening) self-employment by graduates of high schools of the united states of america. different methods for reducing the number of independent variables are reviewed and compared. first, it is the correlation analysis. this method was used for the assessment of these factors’ influence on self-employment of graduates of high schools of the united states of america. secondly, it is the classical factor analysis, and finally it is the confirmatory factor analysis with application of structural equations. these methods were used to decrease the quantity of the variables influencing self-employment. confirmatory factor analysis with application of structural equations was used for the confirmation of the received results. as the data base the information of the national science foundation of the united states of america was used. moreover in this article the factors which have the most significant influence on selfemployment of american graduates are described. keywords: correlation analysis, factor analysis, structural equations, higher education, labor market. 1 introduction in the modern world education is one of the most important prerequisites of successful development of society. sustained economic growth of any country and its competitiveness in the times of globalization of world economy are impossible without highly educated and highly skilled laborers. when considering the tasks facing higher education and statistics of education, the development of a methodological basis of statistical analysis of higher education expert labor market and the employment of such specialists for jobs in their degree field seems potentially productive and extremely important. it is important for national and international studies. the experience of the usa is interesting from the point of view that if in the russian federation the introduction of a two-stage system of higher education is still developing, in the usa the similar experience of training specialists is already obtained. to our mind, special attention should be paid to the criteria of education quality which is used in the usa. the usa is a recognized leader among other countries with market economy, it is a country with a highly developed private sector. a considerable share of concentration of medium-sized and small business is in many respects caused by the mentality: in the usa the choice of speciality is influenced by possibilities of employment, but the prestige of a business also influences the education one’s gets. 2 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. the organization of business by graduates from american higher education institutions is influenced by the whole set of factors. taking into consideration everything said before, at the first stage of our research it is necessary to define factors influencing the basis of a new business organized by american university graduates. 2 research methodology the methods used in the research to find out factors which most strongly influence a probability of finding (opening) self-employment by graduates of high schools in the usa are the following: correlation analysis, eigenvalue method, factor analysis. to prove the results the method of structural equations was used. the data base was the information of the national science foundation of the usa (http://www.nsf.org). 3 findings and discussion our goal is to define a set of factors which most strongly influence a probability of finding (opening) self-employment by graduates of high schools in the usa. we used the database on experts with higher education of the national science foundation of the usa. there are 447 parameters for each graduate. for the construction of a regression model we cannot use all of 447 parameters from the database, therefore, it is necessary to make a reduction of data which is carried out in several stages. on the basis of expert estimation, the parameters which have no influence on the process of self-employment have been removed from the database. for example: ”taking courses during last (reference) week“, ”living in the usa during last (reference) week“ or ”reason for working less than 35 hours a week“. at the given stage the number of parameters was reduced from 447 to 224. using the correlation analysis (pair correlation), from 224 parameters those which most strongly influence the probability of self-employment have been selected. as a result, 36 parameters have been selected with the module of pair correlation coefficients of 0,1 or more (table 1). the parameters with high autocorrelation have been removed from the selected 36 parameters: for example, parameters ”age group [5 year intervals]” and ”year, date of birth [recoded for public use]” as they correlate with the “age” parameter (the value of correlation is 0,99). moreover, some variables can be grouped together: for example, ”employer size“ or ”type of educational institution [employer]” (table 1). hence, such parameters cannot be used as independent variables and should be removed. after the removal of highly correlated parameters and parameters that can be grouped together only 18 parameters were left. it is noteworthy that at first there were 100 000 cases in the database, but after the removal of cases with missing values, there were 569 cases left in the database. to check the received representative sampling we compared histograms of distribution of the variable "age", both for the general set (fig.1) and for sampling (fig.2). so, the histogram of distribution of the variable “age” for the general set does not considerably differ from the same histogram for sampling, it is possible to make a conclusion about the representative sampling. on the basis of the remained 18 parameters the 3 and 2 factor models were received using factor analysis. the kaiser criterion (kaiser, 1960) based on eigenvalues was used to find out the number of factors (table2) and the cattell’s scree test (cattell, 1966). advances in systems science and application(2016) vol.16 no.4 3 table 1 correlation ot the parameters name descriptions included group correlations age age x 0,11 agegr age group [5 year intervals] x 0,11 biryrp year, date of birth [recoded for public use] -0,11 emsec sm employer sector [summary code] x 0,31 emsec pb employer sector [recoded for public use] x 0,29 hdacy r academic year of highest degree -0,14 dgryr year of highest degree x -0,14 hdacy 3 year of highest degree [3 year intervals] x -0,14 hday5 year of highest degree [5 year intervals] x -0,14 hsyr year of receiving high school diploma x -0,11 acdrg type of degree valid during the week of oct. 1 x -0,12 actrd t activity, research, development, and teaching x -0,14 acttc h activity, teaching x -0,16 emed employer is an educational institution x -0,25 emsize employer size x -0,49 edtp type of educational institution [employer] x 0,33 fptind full-time/part-time status including all jobs during the reference week x -0,13 nedtp type of a non-educational institution [employer] x -0,57 nrfam reason for working outside the highest degree field: family-related reasons x 0,10 nrocn a reason for working outside the highest degree field: a desired job is not available x -0,10 waacc work activities on principal job: accounting, finance, contracts x 0,16 waprs m summarized primary work activity x 0,15 wasvc work activities in the principal job: professional services x 0,13 wasal e work activities in the principal job: sales, purchasing, marketing x 0,15 watea work activities in the principal job: teaching x -0,14 newbu s new business x 0,25 baacy r academic year of first bachelor degree -0,11 bayr year of first bachelor’s degree x -0,11 mrdac yr academic year of most recent degree -0,14 mryr year of most recent degree x -0,14 mr3yr year of most recent degree [3-year intervals] x -0,14 mr5yr year of most recent degree [5-year intervals] x -0,14 d2ayr degree award date based on academic year -0,12 d2yr year of second highest degree x -0,12 4 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. d3ayr degree award date based on academic year -0,10 d3yr year of third highest degree x -0,10 histogram (spreadsheet в workbook1.stw 225v*100402c) age = 100402*5*normal(x; 46,9476; 11,8435) 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 age 0 2000 4000 6000 8000 10000 12000 14000 16000 18000 n o o f o b s fig. 1 histogram of distribution of the variable "age" for the general set histogram (spreadsheet в workbook1.stw 23v*569c) age = 569*5*normal(x; 49,7856; 10,4602) 20 25 30 35 40 45 50 55 60 65 70 75 80 age 0 20 40 60 80 100 120 n o o f o b s fig.2 histogram of distribution of the variable "age" for sampling table 2 eigenvalues of factors eigenvalues extraction: principal components eigen value % total variance cumulative eigenvalue cumulati ve % 1 6,3724 88 35,40271 6,37249 35,40271 2 2,9252 16 16,25120 9,29770 51,65391 3 1,3341 97 7,41221 10,63190 59,06611 4 1,1181 76 6,21209 11,75008 65,27820 5 1,0623 68 5,90204 12,81244 71,18025 the scree plot of eigenvalues is given in fig.3. from the table of eigenvalues it is clear, that the greatest share of variance is being described by the first two factors, the same fact is confirmed by the plot, it has excesses on the second and third points (fig.3). hence, it is most logical to consider the 2 and 3 factor models. of the factor analysis methods we selected the principal component analysis with various variants of rotation of axes for the 3 factor and 2 factor models. for rotation of axes the following methods were used: unrotated, varimax raw, varimax normalized, biquartimax raw, advances in systems science and application(2016) vol.16 no.4 5 biquartimax normalized, quartimax raw, quartimax normalized, equamax raw and equamax normalized. plot of eigenvalues number of eigenvalues 0 1 2 3 4 5 6 7 8 v a lu e fig.3 the scree plot of eigenvalues as the criterion of of the rotation of axes method, the values of factor loadings were used. the method with the greatest values of factor loadings is considered the best. factor loading is considered to be high if its value is 0,5 or more. the variable joined in the factor which has the greatest factor loading. moreover, the best model should have a minimal correlation between factors. after analyzing models with various variants of rotation of axes, we came to the conclusion, that the optimal method is varimax normalized, both for the 3 factor and 2 factor models. factor loadings for the given models are given in table 3 and table 4 (the most significant loadings are in bold type). relying on factor loadings, it is possible to assign variables to those factors. let’s have a closer look at the received models and describe them in detail. the received 3 factor model describes dependence of probability of self-employment on the following 3 factors: ”experience“, ”the attitude to education and science“ and ”business characteristics“. the factor «experience» describes work experience of a graduate and is linear approximation of the following characteristics of a high school graduate: ”age“, ”year of highest degree“, ”year of receiving a high school diploma“, ”year of first bachelor’s degree“, ”year of most recent degree“, ”year of second highest degree“, ”year of third highest degree“. it is noteworthy that factor loadings of variables for this factor are higher than 0,9, that shows good approximation of variables by the given factor. the factor ”the attitude to education and science“ shows the attitude to education and science, how much the activity of a graduate is connected with education and science, it is linear approximation of the following characteristics: ”activity, research, development, and teaching“, ”activity, teaching“, ”employer is an educational institution“, ”work activities in the principal job: teaching“. factor loadings for this factor are not so unequivocal if compared with the previous ones, but all their values are high and not lower than 0,6. table 3 factor loadings for the 3 factor model factor loadings (varimax normalized) clusters of loadings are marked; they determine the oblique factors for hierarchical analysis factor 1 factor 2 factor 3 age 0,95299 6 0,06146 8 0,02065 4 dgryr 0,92170 1 0,082739 0,06004 9 table 4 factor loadings for the 2 factor model factor loadings (varimax normalized) clusters of loadings are marked; they determine the oblique factors for hierarchical analysis factor 1 factor 2 age 0,954366 0,012021 dgryr 0,916927 0,135735 6 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. hsyr 0,95689 6 0,06113 0 0,02092 9 actrd t 0,044594 0,65320 6 0,280236 acttc h 0,04986 3 0,90217 4 0,01144 5 emed 0,02353 2 0,80138 2 0,049644 fptind 0,11708 2 0,16849 9 0,36340 0 nrfam 0,057464 0,16323 3 0,52586 1 nrocn a 0,09272 4 0,06485 9 0,28800 0 waac c 0,01964 4 0,346718 0,52497 4 wasv c 0,03913 7 0,049291 0,27536 7 wasal e 0,003421 0,279992 0,58086 5 wate a 0,03008 7 0,82129 8 0,14496 0 newb us 0,077575 0,102705 0,36437 0 bayr 0,97749 3 0,05961 8 0,03345 3 mryr 0,92306 0 0,084538 0,05954 2 d2yr 0,95271 9 0,02518 6 0,04798 4 d3yr 0,95504 6 0,07871 4 0,07030 6 hsyr 0,958248 0,011453 actrdt 0,022482 0,700040 acttch 0,105162 0,874567 emed 0,068952 0,790894 fptind 0,149831 0,079367 nrfam 0,014094 0,046620 nrocn a 0,114510 0,004039 waacc 0,031704 0,453740 wasvc 0,053400 0,110096 wasal e 0,016311 0,399886 watea 0,088993 0,767415 newbu s 0,060479 0,176327 bayr 0,979465 0,006301 mryr 0,918140 0,137439 d2yr 0,953615 0,029312 d3yr 0,960593 0,017869 finally, the third factor «business characteristics» describes a new business where a graduate is employed, it is linearization of characteristics: ”full-time/part-time status including all jobs during the reference week“, ”reason for working outside the highest degree field: family-related reasons“, ”reason for working outside the highest degree field: a desired job is not available“, ”work activities in the principal job: accounting, finance, contracts“, ”work activities in the principal job: professional services“, ”work activities in the principal job: sales, purchasing, marketing“, ”new business“. factor loadings for the variables of this factor are not so significant, they all exceed two times the values of factor loadings for other factors and for some variables do not exceed 0,15. the received 2 factor model describes the dependence of probability of self-employment on the following 2 factors: «experience and environment conditions» and «business characteristics». the first factor «experience and environment conditions» describes work experience of a graduate and work conditions. it is linear approximation of the following characteristics: “age“, “year of highest degree“, “year of receiving a high school diploma“, ”full-time/part-time status including all jobs during the reference week“, ”reason for working outside the highest degree field: a desired job is not available“, ”year of first bachelor’s degree“, ”year of most recent degree“, ”year of second highest degree“ and ”year of third highest degree“. practically all factor loadings of variables for the given factor have high value of more than 0,9, except for factor loadings for parameters ”full-time/part-time status including all jobs during the reference week“ and ”reason for working outside the highest degree field: a desired job ia not available“, the value of their factor loadings does not exceed 0,15. this fact shows the low influence of these parameters on this factor. the second factor ”business characteristics“, as well as the third factor in the 3 factor models characterizes the business in which a graduate is employed. the given factor is linearization of the following parameters: ”activity, research, development, and teaching“, ”activity, advances in systems science and application(2016) vol.16 no.4 7 teaching“, ”employer is an educational institution”, ”reason for working outside the highest degree field: family-related reasons“, ”work activities in the principal job: accounting, finance, contracts“, ”work activities in the principal job: professional services“, ”work activities in the principal job: sales, purchasing, marketing“, ”work activities in the principal job: teaching“, ”new business“. the situation with factor loadings for this factor is similar to the first factor of the given model. factor loadings change within 0,8-0,04, and the parameter ”reason for working outside the highest degree field: family-related reasons” has the least influence on the given factor. below there are given tables of correlations between factors (table 5 and table 6), both for the 3 factor and the 2 factor models. table 5 correlations between factors for the 3 factor model correlations between oblique factors (clusters of variables with unique loadings) fact or 1 facto r 2 facto r 3 fac tor 1 1,00 0000 0,033 756 0,088 356 fac tor 2 0,03 3756 1,000 000 0,182457 fac tor 3 0,08 8356 0,182457 1,000 000 table 6 correlations between factors for the 2 factor model correlations between oblique factors (clusters of variables with unique loadings) factor 1 factor 2 factor 1 1,000000 0,022367 factor 2 0,022367 1,000000 the level of correlation between factors in the both models is low enough, this confirms their importance. in the 3 factor model the correlation between factors is 0,03, 0,08 and -0,18 accordingly, and in the 2 factor model the correlation between factors is -0,02. if to consider the correlation between probability of self-employment and the received factors (table 7 and table 8), we can come to the following conclusions. first, taking into consideration the 3 factor model it is obvious that the third factor has the greatest influence on probability of self-employment, it is followed by the first factor and then the second one. secondly, with the growth of the first or third factors, the value of probability decreases, and with the growth of the second factor, the probability increases. thirdly, in the 2 factor model the both factors have equally negative influence on probability of self-employment. table 7 correlation between probability of self-employment and the received factors for the 3 factor model correlations marked correlations are significant at p <, 05000 n=569 (casewise deletion of missing data) factor 1 factor 2 factor 3 selfemp l -0,20 0,11 -0,37 table 8 correlation between probability of self-employment and the received factors for the 2 factor model correlations marked correlations are significant at p <, 05000 n=569 (casewise deletion of missing data) factor 1 factor 2 selfempl -0,22 -0,22 if to compare the 3 factor and the 2 factor models, it is possible to draw the following conclusion: though the correlation between factors in the 2 factor model is less than in the 3 factor model, from the point of view of the explanation of factor value, the 3 factor model is better. we will verify the results of the factor analysis using structural equations. reliability of the received factors was verified by the confirming factor analysis using structural equations. using the structural equations for both the 3 factor and the 2 factor models, we employed the method of maximum likelihood estimation together with the method of least squares, with absence of correlation between factors and the residuals since low correlation between factors was given above (table 5 and table 6). the data for the analysis was the matrix 8 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. of correlations between parameters. the results of the analysis of the constructed structural equations were the following. table 9 the confirming factor analysis using structural equations for the 3 factor model model estimates para meter estim ate stan dard erro r t statistic p rob. l evel (experience)-1> [age] 0,980 0,00 7 1,409422e+02 0 ,000 (experience)-2> [dgryr] 1,000 0,00 0 8,086041e +03 0 ,000 (experience)-3> [hsyr] 0,980 0,00 7 1,416687e +02 0 ,000 (experience)-4> [bayr] 0,985 0,00 5 1,880929e +02 0 ,000 (experience)-5> [mryr] 1,000 0,00 0 8,033911e +03 0 ,000 (experience)-6> [d2yr] 0,988 0,00 4 2,397669e +02 0 ,000 (experience)-7> [d3yr] 0,977 0,00 8 1,200892e +02 0 ,000 (teaching)-8> [actrdt] 0,814 0,08 2 9,930843e +00 0 ,000 (teaching)-9> [acttch] 1,000 0,00 0 1,475717e +09 0 ,000 (teaching)-10> [emed] 0,930 0,03 3 2,829592e +01 0 ,000 (teaching)-11> [watea] 0,951 0,02 3 4,060509e +01 0 ,000 (business)-12> [fptind] 0,013 0,27 1 4,958209e -02 0 ,960 (business)-13> [nrfam] 0,085 0,27 0 3,138192e -01 0 ,754 (business)-14> [nrocna] 0,193 0,26 3 -7,332709e-01 0 ,463 (business)-15> [waacc] 0,813 0,27 5 2,961953e +00 0 ,003 (business)-16> [wasvc] 0,169 0,26 5 6,393056e -01 0 ,523 (business)-17> [wasale] 0,806 0,27 3 2,948677e +00 0 ,003 (business)-18> [newbus] 0,272 0,25 5 1,066804e +00 0 ,286 (delta1)-19 (delta1) 0,040 0,01 4 2,915270e +00 0 ,004 (delta2)-20 (delta2) 0,000 0,00 0 2,041707e -01 0 ,838 (delta3)-21 (delta3) 0,040 0,01 4 2,915235e +00 0 ,004 (delta4)-22 (delta4) 0,030 0,01 0 2,913469e +00 0 ,004 (delta5)-23 (delta5) 0,000 0,00 0 3,872706e -01 0 ,699 (delta6)-24 (delta6) 0,024 0,00 8 2,912069e +00 0 ,004 (delta7)-25 (delta7) 0,046 0,01 6 2,916457e +00 0 ,004 (delta8)-26 (delta8) 0,338 0,13 3 2,533464e +00 0 ,011 (delta9)-27 (delta9) 0,000 0,00 0 (delta10)-28 (delta10) 0,135 0,06 1 2,217215e +00 0 ,027 (delta11)-29 0,097 0,04 2,168875e 0 advances in systems science and application(2016) vol.16 no.4 9 (delta11) 5 +00 ,030 (delta12)-30 (delta12) 1,000 0,00 7 1,368791e +02 0 ,000 (delta13)-31 (delta13) 0,993 0,04 6 2,173207e +01 0 ,000 (delta14)-32 (delta14) 0,963 0,10 1 9,498401e +00 0 ,000 (delta15)-33 (delta15) 0,338 0,44 7 7,571400e -01 0 ,449 (delta16)-34 (delta16) 0,971 0,09 0 1,082963e +01 0 ,000 (delta17)-35 (delta17) 0,351 0,44 0 7,979269e -01 0 ,425 (delta18)-36 (delta18) 0,926 0,13 8 6,699364e +00 0 ,000 for the 3 factor model the results are described in table 9 (the significant facts are in bold type), and the normal probability plot of residuals is given in fig.4. from table 9 it is obvious that the ways for the factors “experience” and “the attitude to education and science” are significant as they have high value of t-statistics and low probability. hence, it is possible to conclude that the considered model precisely describes the set forth above factors. however, we can observe that the factor “business characteristics” is not so well described by the given model, not all ways for this factor have high t-statistics and low probability. for additional verification of the importance of the model the normal probability plot of residuals (fig.4) has been constructed. from the plot it is obvious that the residuals of model are situated close to the straight line of normal distribution that confirms the importance of the constructed model. on the whole we can conclude that the results received, using the diagram of ways for the 3 factor model, coincide with the results of the factor analysis for this model which proves their correctness. table 10 the confirming factor analysis using structural equations for the 2 factor model model estimates param eter estima te stan dard erro r t statistic p rob. l evel (experience)-1> [age] -1,000 0,00 0 2,083354e+05 0 ,000 (experience)-2> [dgryr] 0,980 0,01 0 1,020890e +02 0 ,000 (experience)-3> [hsyr] 1,000 0,00 0 3,205600e +12 0 ,000 (experience)-4> [fptind] 0,254 0,22 7 1,117756e +00 0 ,264 (experience)-5> [nrocna] 0,136 0,23 8 5,715618e -01 0 ,568 (experience)-6> [bayr] 0,997 0,00 2 6,361259e +02 0 ,000 (experience)-7> [mryr] 0,980 0,01 0 1,000551e +02 0 ,000 (experience)-8> [d2yr] 0,995 0,00 3 3,952639e +02 0 ,000 (experience)-9> [d3yr] 0,997 0,00 1 6,970720e +02 0 ,000 (business)-10> [actrdt] 0,814 0,08 2 9,930843e +00 0 ,000 (business)-11> [acttch] 1,000 0,00 0 3,545235e +02 0 ,000 (business)-12> [emed] 0,930 0,03 3 2,829592e +01 0 ,000 (business)-13> [nrfam] 0,027 0,24 2 1,124848e -01 0 ,910 (business)-14> -0,588 0,15 0 10 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. [waacc] 9 3,700352e+00 ,000 (business)-15> [wasvc] -0,164 0,23 6 -6,967788e-01 0 ,486 (business)-16> [wasale] -0,530 0,17 4 3,040265e+00 0 ,002 (business)-17> [watea] 0,951 0,02 3 4,060509e +01 0 ,000 (business)-18> [newbus] -0,346 0,21 4 1,617683e+00 0 ,106 (delta1)-19 (delta1) 0,000 0,00 0 2,061573e +00 0 ,039 (delta2)-20 (delta2) 0,040 0,01 9 2,103603e +00 0 ,035 (delta3)-21 (delta3) 0,000 0,00 0 (delta4)-22 (delta4) 0,936 0,11 5 8,127443e +00 0 ,000 (delta5)-23 (delta5) 0,981 0,06 5 1,515202e +01 0 ,000 (delta6)-24 (delta6) 0,006 0,00 3 2,068245e +00 0 ,039 (delta7)-25 (delta7) 0,040 0,01 9 2,104467e +00 0 ,035 (delta8)-26 (delta8) 0,010 0,00 5 2,072333e +00 0 ,038 (delta9)-27 (delta9) 0,006 0,00 3 2,067659e +00 0 ,039 (delta10)-28 (delta10) 0,338 0,13 3 2,533464e +00 0 ,011 (delta11)-29 (delta11) 0,000 0,00 0 (delta12)-30 (delta12) 0,135 0,06 1 2,217215e +00 0 ,027 (delta13)-31 (delta13) 0,999 0,01 3 7,562196e +01 0 ,000 (delta14)-32 (delta14) 0,655 0,18 7 3,508443e +00 0 ,000 (delta15)-33 (delta15) 0,973 0,07 8 1,253796e +01 0 ,000 (delta16)-34 (delta16) 0,719 0,18 5 3,888715e +00 0 ,000 (delta17)-35 (delta17) 0,097 0,04 5 2,168875e +00 0 ,030 (delta18)-36 (delta18) 0,881 0,14 8 5,966712e +00 0 ,000 let’s now consider the 2 factor model. the parameters of the 2 factor model are given in table 10 and the normal probability plot of residuals is given in fig.5. as well as in the factor analysis, the factor “experience and environment conditions” is well enough described by the diagram of ways, practically all the ways have high t-statistics and low probability. the exceptions are only the ways for the parameters “full-time/part-time status including all jobs during the reference week” and “reason for working outside the highest degree field: a desired job is not available”. the second factor is not so well described by the given model. the ways for the parameters “reason for working outside the highest degree field: family-related reasons”, “work activities in the principal job: professional services” and “new business” have low tstatistics and high probability. the normal probability plot of residuals for this model is practically similar to the normal probability plot of residuals for the 3 factor model this proves that the model conforms to real data. on the whole, the model confirms the results of the factor analysis. considering all the results described above we can draw the conclusion that the received structure of correlation between the parameters and the factors which include them coincides with the structure of correlation of the factor analysis for the 3 factor and the 2 factor models. hence, advances in systems science and application(2016) vol.16 no.4 11 the results of the confirming factor analysis using structural equations, on the whole, confirm the results of the factor analysis. normal probability plot normalized residuals -3,0 -2,5 -2,0 -1,5 -1,0 -0,5 0,0 0,5 1,0 1,5 2,0 value -3 -2 -1 0 1 2 3 e x p e c te d n o rm a l v a lu e fig. 4 the normal probability plot of residuals for the 3 factor model normal probability plot normalized residuals -2,0 -1,5 -1,0 -0,5 0,0 0,5 1,0 1,5 2,0 value -3 -2 -1 0 1 2 3 e x p e c te d n o rm a l v a lu e fig. 5 the normal probability plot of residuals for the 2 factor model further on, on the basis of the received factors the regression models showing dependence of probability of self-employment on given factors will be constructed. for it various methods are to be used to find out the best regression model for binary responses. references [1] brown m.b., and forsythe a.b(1974). "robust tests for the equality of variances". journal of the american statistical association, no.69, pp. 364-367. [2] dallal g.e. and wilkinson l(1986). "an analytic approximation to the distribution of lilliefor’s test statistic for normality". the american statistician, vol.40, no.4, pp. 294-296 [3] de leeuw j. and van rijckevorsel j(1980). homals and princals some generalizations of principal components analysis. in: data analysis and informatics, e. diday et al, eds. amsterdam: north-holland. [4] dziuban c.d. and shirkey e.c(1974). "when is a correlation matrix appropriate for factor analysis? ", psychological bulletin, no. 81, pp. 358-361. [5] encyclopedia of statistical sciences 4. ny: wiley, pp. 608-610. [6] frigge m., hoaglin, d.c., and iglewicz b(1987). "some implementations for the boxplot". in: computer science and statistics proceedings of the 19th symposium on the interface, r. m. heiberger and m. martin, eds. alexandria, va.: american statistical association. 12 a. v. burkov, e. a. murzina : analysis method of structural equation modeling. [7] glaser, r. e(1983). levene’s robust test of homogeneity of variances. [8] harman h.h(1976). modern factor analysis, 3rd ed., chicago: university of chicago press. [9] hendrickson a.e. and and white p.o (1964). "promax: a quick method for rotation to oblique simple structure". british journal of statistical psychology, no. 17, pp. 65-70. [10] hoaglin d.c., mosteller, f., and tukey j.w. (1983). understanding robust and exploratory data analysis. new york: john wiley & sons, inc. [11] hoaglin d.c., mosteller, f., and tukey j.w.(1985). exploring data tables, trends, and shapes. new york: john wiley & sons, inc. [12] hoyle, r.h. (ed.). (1995). structural equation modeling. concepts, issues, and applications. thousand oaks, ca: sage. [13] israлls a. (1987). eigenvalue techniques for qualitative data. leiden: dswo press. [14] jennrich r.i. and sampson p.f. (1966). rotation for simple loading. psychometrika, vol.31, no.3,pp. 313-323. [15] joreskog k. g. (1977). factor analysis by least-square and maximum likelihood methods. in: statistical methods for digital computers, volume 3, k. enslein, a. ralston, and r.s. wilf, eds. new york: john wiley & sons, inc. [16] kaiser h.f. (1963). image analysis. in: problems in measuring change, c.w.harris, ed. madison, wisc.: university of wisconsin press. [17] lilliefors h.w. (1967). "on the kolmogorov-smirnov tests for normality with mean and variance unknown". journal of the american statistical association, no.62, pp. 399-402. [18] loh w.y. (1987). "some modifications of levene’s test of variance homogeneity". journal of statistical computation and simulation, no.28, pp. 213-226. [19] rao c.r. (1955). "estimation and test of significance in factor analysis". psychometrika, vol.20, no.2, pp. 93-111. [20] rummel r.j. (1970). applied factor analysis. evanston: ill.: northwestern university press. [21] theil h. (1953b). estimation and simultaneous correlation in complete equation systems. the hague: central planning bureau. [22] tukey j.w. (1977). exploratory data analysis. reading, mass.: addison-wesley. [23] velleman p.f. and hoaglin d.c. (1981). applications, basics, and computing of exploratory data analysis. boston: duxbury press. [24] wilkinson j h. (1965). the algebraic eigenvalue problem. oxford: clarendon press. [25] yalyalieva, t.v. and murzina, e.a. (2015) , "the system of parameters efficiency of financial supervision". advances in systems science and application vol.15 no.4, pp.384-391. corresponding author elena a. murzina can be contacted at: elena.murzina@gmail.com. advances in systems science and applications (2011), vol. 11, no. 1-2 149-162 monotone hybrid methods for a finite family of nonexpasive multi-valued maps and equilibrium problems ∗ suthep suantai and watcharaporn cholamjiak department of mathematics, faculty of science, chiang mai university, chiang mai 50200, thailand centre of excellence in mathematics, che, si ayutthaya rd., bangkok 10400, thailand email: scmti005@chiangmai.ac.th, c-wchp007@hotmail.com abstract in this paper, we introduce a new monotone hybrid iterative scheme for finding a common element of the set of common fixed points of a finite family of nonexpansive multivalued maps and the set of the solutions of the equilibrium problem in a hilbert space. moreover, we also introduce a new iterative scheme for finding a common fixed point of a finite family of nonexpansive multi-valued maps in a banach space. strong convergence theorem of the proposed iteration is established. keywords nonexpansive multi-valued map monotone hybrid method, 1. introduction letd be a nonempty convex subset of a banach spacee. let f be a bifunction fromd×d to r, where r is the set of all real number. the equilibrium problem for f is to find x ∈ d such that f(x, y) ≥ 0 for all y ∈ d. the set of such solutions is denoted by ep (f). the set d is called proximinal if for each x ∈ e, there exists an element y ∈ d such that ‖x−y‖ = d(x,d), where d(x,d) = inf{‖x − z‖ : z ∈ d}. let cb(d),k(d) and p (d) denote the families of nonempty closed bounded subsets, nonempty compact subsets, and nonempty proximinal bounded subsets of d, respectively. the hausdorff metric on cb(d) is defined by h(a,b) = max { sup x∈a d(x,b), sup y∈b d(y,a) } for a,b ∈ cb(d). a single-valued map t : d → d is called nonexpansive if ‖tx− ty‖ ≤ ‖x − y‖ for all x, y ∈ d. a multi-valued map t : d → cb(d) is said to be nonexpansive if h(tx, ty) ≤ ‖x−y‖ for all x, y ∈ d. an element p ∈ d is called a fixed point of t : d → d (respectively, t : d → cb(d)) if p = tp (respectively, p ∈ tp). the set of fixed points of t is denoted by f (t ). the mapping t : d → cb(d) is called quasi-nonexpansive[18] if f (t ) 6= ∅ and h(tx, tp) ≤ ‖x − p‖ for all x ∈ d and all p ∈ f (t ). it is clear that every nonexpansive multi-valued map t with f (t ) 6= ∅ is quasi-nonexpansive. but there exist quasi-nonexpansive mappings that are not nonexpansive, see [17]. the mapping t : d → cb(d) is called hemicompact if, for any sequence {xn} in d such that d(xn, txn) → 0 as n → ∞, there exists a subsequence {xnk } of {xn} such that xnk → ∗this research is supported by the centre of excellence in mathematics, the commission on higher education, thailand, the thailand research fund and the graduate school of chiang mai university for the financial support. issn 1078-6236 international institute for general systems studies, inc. 1 2 1 2 150 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . p ∈ d. we note that if d is compact, then every multi-valued mapping t : d → cb(d) is hemicompact. a mapping t : d → cb(d) is said to satisfy condition (i) if there is a nondecreasing function f : [0,∞)→ [0,∞) with f(0) = 0, f(r) > 0 for r ∈ (0,∞) such that d(x, tx) ≥ f(d(x, f (t ))) for all x ∈ d. a family {ti : d → cb(d), i = 1, 2, ..., n} is said to satisfy condition (ii) if there is a nondecreasing function f : [0,∞)→ [0,∞) with f(0) = 0, f(r) > 0 for r ∈ (0,∞) such that d(x, tix) ≥ f(d(x, n⋂ i=1 f (ti))) for all i = 1, 2, ..., n and x ∈ d. in 1953, mann [10] introduced the following iterative procedure to approximate a fixed point of a nonexpansive mapping t in a hilbert space h: xn+1 = αnxn + (1− αn)txn, ∀n ∈ n, (1) where the initial point x0 is taken in c arbitrarily and {αn} is a sequence in [0,1]. however, we note that mann’s iteration process (1) has only weak convergence, in general; for instance, see [1, 7, 15]. in 2003, nakajo and takahashi [12] introduced the method which is the so-called cq method to modify the process (1) so that strong convergence is guaranteed have recently been made. they also proved a strong convergence theorem for a nonexpansive mapping in a hilbert space. recently, tada and takahashi [20] proposed a new iteration for finding a common element of the set of solutions of an equilibrium problem and the set of fixed points of a nonexpansive mapping t in a hilbert space h . in 2005, sastry and babu [16] proved that the mann and ishikawa iteration schemes for multi-valued map t with a fixed point p converge to a fixed point q of t under certain conditions. they also claimed that the fixed point q may be different from p. more precisely, they proved the following result for nonexpansive multi-valued map with compact domain. in 2007, panyanak [13] extended the above result of sastry and babu [16] to uniformly convex banach spaces but the domain of t remains compact. later, song and wang [19] noted that there was a gap in the proofs of theorem 3.1(see [13]) and theorem 5 (see [16]). they further solved/revised the gap and also gave the affirmative answer to panyanak [13] question using the following ishikawa iteration scheme. in the main results, domain of t is still compact, which is a strong condition (see [19], theorem 1) and t satisfies condition(i) (see [19], theorem 1). in 2009, shahzad and zegeye [17] extended and improved the results of panyanak [13], sastry and babu [16] and song and wang [19] to quasi-nonexpansive multi-valued maps. they also relaxed compactness of the domain of t and constructed an iteration scheme which removes advances in systems science and applications (2011), vol. 11, no. 1-2 151 the restriction of t namely tp = {p} for any p ∈ f (t ). the results provided an affirmative answer to panyanak [13] question in a more general setting. in the main results, t satisfies condition(i)(see [17], theorem 2.3) and t is hemicompact and continuous (see [17], theorem 2.5). question: how can we modify iteration process for a nonexpansive multi-valued map t which the domain of t is not necessary to be compact to obtain strong convergence theorems for finding a common element of the set of solutions of an equilibrium problem and the set of fixed points of t ? in the recent years, the problem of finding a common element of the set of solutions of equilibrium problems and the set of fixed points in the framework of hilbert spaces and banach spaces have been intensively studied by many authors, for instance, see [2, 3, 4, 5, 6, 8, 14, 20] and the references cited theorem. in this paper, we introduce a monotone hybrid iterative scheme for finding a common element of the set of a common fixed points of a finite family of nonexpasive multi-valued maps and the set of solutions of an equilibrium problem in a hilbert space. let d be nonempty, closed and convex subset of a hilbert space h and αi n ∈ (0, 1) for all i = 0, 1, ...,m with∑m i=0 α i n = 1, ∀n ≥ 0 and rn ∈ (0,∞). for an initial point x0 ∈ d = c0, compute the sequence {xn} by the iterative process f(un, y) + 1 rn 〈y − un, un − xn〉 ≥ 0, ∀y ∈ d, yn = ∑m i=0 α i nz i n, z i n ∈ tiun, ∀i = 1, 2, ...,m, z0n = un, cn+1 = {z ∈ cn : ‖yn − z‖ ≤ ‖xn − z‖}, xn+1 = pcn+1x0, n ≥ 0, (2) where ti is a nonexpansive multi-valued map for all i = 1, 2, ...,m. 2. preliminaries the following lemmas give some characterizations and a useful property of the metric projection pd in a hilbert space. let h be a real hilbert space with inner product 〈·, ·〉 and norm ‖ · ‖. let d be a closed and convex subset of h . for every point x ∈ h , there exists a unique nearest point in d, denoted by pdx, such that ‖x− pdx‖ ≤ ‖x− y‖, ∀y ∈ d. pd is called the metric projection of h onto d. we know that pd is a nonexpansive mapping of h onto d. lemma 2.1. [11] let d be a closed and convex subset of a real hilbert space h and let pd be the metric projection from h onto d. given x ∈ h and z ∈ d. then z = pdx if and only if the following holds: 〈x− z, y − z〉 ≤ 0, ∀y ∈ d. lemma 2.2. [12] let d be a nonempty, closed and convex subset of a real hilbert space h and pd : h → d be the metric projection from h onto d. then the following inequality holds: ‖y − pdx‖2 + ‖x− pdx‖2 ≤ ‖x− y‖2, ∀x ∈ h, ∀y ∈ d. 152 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . lemma 2.3. [11] let h be a real hilbert space. then the following equations hold: (i) ‖x− y‖2 = ‖x‖2 − ‖y‖2 − 2〈x− y, y〉, ∀x, y ∈ h; (ii) ‖tx+ (1− t)y‖2 = t‖x‖2 + (1− t)‖y‖2 − t(1− t)‖x− y‖2, ∀t ∈ [0, 1] and x, y ∈ h . by using lemma 2.3, we obtain the following lemma. lemma 2.4. let h be a real hilbert space. then for each m ∈ n ‖ m∑ i=1 tixi‖2 = m∑ i=1 ti‖xi‖2 − m∑ i=1,i 6=j titj‖xi − xj‖2, xi ∈ h and ti, tj ∈ [0, 1] for all i, j = 1, 2, ...,m with ∑m i=1 ti = 1. lemma 2.5. [9]letd be a nonempty, closed and convex subset of a real hilbert spaceh . given x, y, z ∈ h and also given a ∈ r, the set {v ∈ d : ‖y − v‖2 ≤ ‖x− v‖2 + 〈z, v〉+ a} is convex and closed. for solving the equilibrium problem, we assume the bifunction f : d × d → r satisfies the following conditions: (a1) f(x, x) = 0 for all x ∈ d; (a2) f is monotone, i.e., f(x, y) + f(y, x) ≤ 0 for all x, y ∈ d; (a3) for each x, y, z ∈ d, lim supt↓0 f(tz + (1− t)x, y) ≤ f(x, y); (a4) f(x, ·) is convex and lower semicontinuous for each x ∈ d. lemma 2.6. [2] let d be a nonempty, closed and convex subset of a real hilbert space h . let f be a bifunction from d ×d to r satisfying (a1)-(a4) and let r > 0 and x ∈ h . then, there exists z ∈ d such that f(z, y) + 1 r 〈y − z, z − x〉 ≥ 0, for all y ∈ d. lemma 2.7. [6] for r > 0, x ∈ h , defined a mapping tr : h → 2d as follows: tr(x) = { z ∈ d : f(z, y) + 1 r 〈y − z, z − x〉 ≥ 0, for all y ∈ d } . then the followings hold: (1) tr is single-value; (2) tr is firmly nonexpansive, i.e., for any x, y ∈ h , ‖trx− try‖2 ≤ 〈trx− try, x− y〉; (3) f (tr) = ep (f); (4) ep (f) is closed and convex. advances in systems science and applications (2011), vol. 11, no. 1-2 153 lemma 2.8. let d be a closed and convex subset of a real hilbert space h . let t : d → cb(d) be a nonexpansive multi-valued map with f (t ) 6= ∅ and tp = {p} for each p ∈ f (t ). then f (t ) is a closed and convex subset of d. proof. first, we will show that f (t ) is closed. let {xn} be a sequence in f (t ) such that xn → x as n→∞. we have d(x, tx) ≤ d(x, xn) + d(xn, tx) ≤ d(x, xn) +h(txn, tx) ≤ 2d(x, xn). it follows that d(x, tx) = 0, so x ∈ f (t ). next, we show that f (t ) is convex. let p = tp1 + (1− t)p2 where p1, p2 ∈ f (t ) and t ∈ (0, 1) . let z ∈ tp, by lemma 2.3, we have ‖p− z‖2 = ‖t(z − p1) + (1− t)(z − p2)‖2 = t‖z − p1‖2 + (1− t)‖z − p2‖2 − t(1− t)‖p1 − p2‖2 = td(z, tp1) 2 + (1− t)d(z, tp2) 2 − t(1− t)‖p1 − p2‖2 ≤ th(tp, tp1) 2 + (1− t)h(tp, tp2) 2 − t(1− t)‖p1 − p2‖2 ≤ t‖p− p1‖2 + (1− t)‖p− p2‖2 − t(1− t)‖p1 − p2‖2 = t(1− t)2‖p1 − p2‖2 + (1− t)t2‖p1 − p2‖2 − t(1− t)‖p1 − p2‖2 = 0, hence p = z. therefore p ∈ f (t ). lemma 2.9. [21] let p > 1, r > 0 be two fixed numbers. then a banach space e is uniformly convex if and only if there exists a continuous, strictly increasing, and convex function g : [0,∞)→ [0,∞) with g(0) = 0 such that ‖λx+ (1− λ)y‖p ≤ λ‖x‖p + (1− λ)‖y‖p − ωp(λ)g(‖x− y‖), for all x, y ∈ br(0) = {x ∈ e : ‖x‖ ≤ r} and λ ∈ [0, 1] where ωp(λ) = λ(1−λ)p+λp(1−λ). by using lemma 2.9, we can prove the following lemma by induction. lemma 2.10. let e be a uniformly convex banach space and br(0) = { x ∈ e : ‖x‖ ≤ r} be a closed ball of e. then there exists a continuous strictly increasing convex function g : [0,∞)→ [0,∞) with g(0) = 0 such that ‖ m∑ i=1 αixi‖2 ≤ m∑ i=1 αi‖xi‖2 − α1α2g(‖x1 − x2‖), for all m ∈ n, xi ∈ br(0) and αi ∈ [0, 1], i = 1, 2, ...,m with ∑m i=1 αi = 1. by interchanging the roles of vectors xi in lemma 2.10 and summing the inequalities together we obtain the following lemma. 154 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . lemma 2.11. let e be a uniformly convex banach space and br(0) = {x ∈ e : ‖x‖ ≤ r} be a closed ball of e. then there exists a continuous strictly increasing convex function g : [0,∞)→ [0,∞) with g(0) = 0 such that for each j ∈ {1, 2, ...,m}, ‖ m∑ i=1 αixi‖2 ≤ m∑ i=1 αi‖xi‖2 − αj m− 1 ( m∑ i=1 αig(‖xj − xi‖) ) , for all m ∈ n, xi ∈ br(0) and αi ∈ [0, 1] for all i = 1, 2, ...,m with ∑m i=1 αi = 1. proof. let j ∈ {1, 2, ...,m} be fixed. by lemma 2.10, there is a continuous strictly increasing convex function g : [0,∞)→ [0,∞) with g(0) = 0 such that ‖α1x1 + α2x2 + α3x3 + ...+ αmxm‖2 ≤ m∑ i=1 αi‖xi‖2 − αjα1g(‖xj − x1‖) ‖α1x1 + α3x3 + α4x4 + ...+ α2x2‖2 ≤ m∑ i=1 αi‖xi‖2 − αjα2g(‖xj − x2‖) ... ‖α1x1 + α4x4 + α5x5 + ...+ α3x3‖2 ≤ m∑ i=1 αi‖xi‖2 − αjαj−1g(‖xj − xj−1‖) ‖α1x1 + α4x4 + α5x5 + ...+ α3x3‖2 ≤ m∑ i=1 αi‖xi‖2 − αjαj+1g(‖xj − xj+1‖) ... ‖α1x1 + αmxm + α2x2 + ...+ αm−1xm−1‖2 ≤ m∑ i=1 αi‖xi‖2 − αjαmg(‖xj − xm‖). by summing up above inequalities, we obtain ‖ m∑ i=1 αixi‖2 ≤ m∑ i=1 αi‖xi‖2 − αj m− 1 ( m∑ i=1 αig(‖xj − xi‖) ) . 3. main result first, we prove a strong convergence theorem for a finite family of nonexpansive multivalued mappings which satisfies the condition (ii) in a uniformly convex banach space. theorem 3.1. let d be a nonempty, closed and convex subset of a uniformly convex banach spacee. let ti : d → cb(d) be a nonexpansive multi-valued map for all i = 1, 2, ...,m with⋂m i=1 f (ti) 6= ∅ and tip = {p} for each p ∈ ⋂m i=1 f (ti). assume that {ti : i = 1, 2, ...,m} satisfies the condition (ii) for all i = 1, 2, ...,m and αi n ∈ (0, 1) with 0 < lim infn→∞ α i n ≤ advances in systems science and applications (2011), vol. 11, no. 1-2 155 lim supn→∞ α i n < 1 for all i = 0, 1, 2, ...,m. let x0 ∈ d and let {xn} be the sequence in d generated by iteration process: xn+1 = m∑ i=0 αi nz i n, (3) where z0n = xn, zin ∈ tixn for all i = 1, 2, ...,m and ∑m i=0 α i n = 1. then {xn} converges strongly to a common fixed point of ti, i = 1, 2, ...,m. proof. let p ∈ ⋂m i=1 f (ti). by the nonexpansiveness of ti, we have ‖xn+1 − p‖ ≤ m∑ i=0 αi n‖zin − p‖ = α0 n‖xn − p‖+ m∑ i=1 αi nd(zin, tip) ≤ α0 n‖xn − p‖+ m∑ i=1 αi nh(tixn, tip) ≤ ‖xn − p‖, (4) which implies that limn→∞ ‖xn − p‖ exists. for each i = 1, 2, ...,m, we have ‖zin − p‖ = d(zin, tip) ≤ h(tixn, tip) ≤ ‖xn − p‖. it follows that {‖zin − p‖} is bounded for all i = 1, 2, ...,m. put r = max1≤i≤m{supn ‖zin− p‖}. by lemma 2.11, there is a continuous strictly increasing convex function g : [0,∞)→ [0,∞) with g(0) = 0 such that ‖xn+1 − p‖2 ≤ m∑ i=0 αi n‖zin − p‖2 − α0 n m m∑ i=1 αi ng ( ‖zin − xn‖ ) = α0 n‖xn − p‖2 + m∑ i=1 αi nd(zin, tip) 2 − α0 n m m∑ i=1 αi ng ( ‖zin − xn‖ ) ≤ α0 n‖xn − p‖2 + m∑ i=1 αi nh(tixn, tip) 2 − α0 n m m∑ i=1 αi ng ( ‖zin − xn‖ ) ≤ ‖xn − p‖2 − α0 n m m∑ i=1 αi ng ( ‖zin − xn‖ ) . it follows that α0 n m m∑ i=1 αi ng ( ‖zin − xn‖ ) ≤ ‖xn − p‖2 − ‖xn+1 − p‖2. this implies that g ( ‖zin − xn‖ ) → 0 as n → ∞ for all i = 1, 2, ...,m. since g is continuous strictly increasing with g(0) = 0, we can conclude that ‖zin − xn‖ → 0 as n → ∞ for all i = 1, 2, ...,m. also d(xn, tixn) ≤ ‖zin − xn‖ → 0 as n → ∞ for all i = 1, 2, ...,m. since that {ti}mi=1 satisfies the condition (ii), we have d(xn, ⋂m i=1 f (ti))→ 0. thus there is a 156 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . subsequence {xnk } of {xn} such that ‖xnk − pk‖ < 1 2k for some {pk} ⊂ ⋂m i=1 f (ti) and all k. from (4), we obtain ‖xnk+1 − pk‖ ≤ ‖xnk − pk‖ < 1 2k . next, we shall show that {pk} is cauchy sequence in d. notice that ‖pk+1 − pk‖ ≤ ‖pk+1 − xnk+1 ‖+ ‖xnk+1 − pk‖ < 1 2k+1 + 1 2k < 1 2k−1 . this implies that {pk} is cauchy sequence in d and thus converges to q ∈ d. since d(pk, tiq) ≤ h(tiq, tipk) ≤ ‖q − pk‖ for all i = 1, 2, ...,m and pk → q as n→∞, it follows that d(q, tiq) = 0 for all i = 1, 2, ...,m and thus q ∈ ⋂m i=1 f (ti) and {xnk } converges strongly to q. since limn→∞ ‖xn − q‖ exists, it follows that {xn} converges strongly to q. this completes the proof. note that in theorem 3.1 in order to have strong convergence of the iterative sequence {xn} defined by (3), we need to assume that {ti}mi=1 satisfy the condition (ii). in the following theorem, we introduce a new monotone hybrid iterative scheme (2) for finding a common element of the set of a common fixed points of a family of nonexpasive multi-valued maps and the set of solutions of an equilibrium problem in a hilbert space, and we prove strong convergence of the sequence {xn} defined by (2) without the condition (ii). theorem 3.2. let d be a nonempty, closed and convex subset of a real hilbert space h . let f be a bifunction from d × d to r satisfying (a1)-(a4) and let ti : d → cb(d) be nonexpansive multi-valued maps for all i = 1, 2, ...,m with ⋂m i=1 f (ti) ∩ ep (f) 6= ∅ and tip = {p} for each p ∈ ⋂m i=1 f (ti). assume that αi n ∈ (0, 1) with 0 < lim infn→∞ α i n ≤ lim supn→∞ α i n < 1 for all i = 0, 1, 2, ...,m and rn ∈ (0,∞) with lim infn→∞ rn > 0. then the sequence {xn} generated by (2) converges strongly to p⋂m i=1 f (ti)∩ep (f)x0. proof. we split the proof into six steps. step 1. show that pcn+1x0 is well defined for every x0 ∈ d. by lemma 2.8, we obtain that ⋂m i=1 f (ti) is a closed and convex subset of d. since ep (f) is also closed and convex, then ⋂m i=1 f (ti) ∩ ep (f) is a closed and convex subset of d. from the definition of cn+1, it follows from lemma 2.5 that cn+1 is closed and convex for each n ≥ 0. let v ∈ ⋂m i=1 f (ti) ∩ ep (f). from un = trnxn, we have ‖un − v‖ = ‖trnxn − trnv‖ ≤ ‖xn − v‖, (5) for every n ≥ 0. from this, we have ‖yn − v‖ = ‖ m∑ i=0 αi nz i n − v‖ ≤ m∑ i=0 αi n‖zin − v‖ = α0 n‖un − v‖+ m∑ i=1 αi nd(zin, tiv) ≤ α0 n‖un − v‖+ m∑ i=1 αi nh(tiun, tiv) ≤ ‖un − v‖ ≤ ‖xn − v‖. (6) advances in systems science and applications (2011), vol. 11, no. 1-2 157 so, we have v ∈ cn+1, thus ⋂m i=1 f (ti)∩ep (f) ⊂ cn+1. therefore pcn+1x0 is well defined. step 2. show that limn→∞ ‖xn − x0‖ exists. since ⋂m i=1 f (ti) ∩ ep (f) is a nonempty, closed and convex subset of h , there exists a unique v ∈ ⋂m i=1 f (ti) ∩ ep (f) such that v = p⋂m i=1 f (ti)∩ep (f)x0. from xn = pcnx0, cn+1 ⊂ cn and xn+1 ∈ cn, ∀n ≥ 0, we get ‖xn − x0‖ ≤ ‖xn+1 − x0‖, ∀n ≥ 0. on the other hand, as ⋂m i=1 f (ti) ∩ ep (f) ⊂ cn, we obtain ‖xn − x0‖ ≤ ‖v − x0‖, ∀n ≥ 0. it follows that the sequence {xn} is bounded and nondecreasing. therefore limn→∞ ‖xn−x0‖ exists. step 3. show that xn → w ∈ d as n→∞. for m > n, by the definition of cn, we see that xm = pcmx0 ∈ cm ⊂ cn. by lemma 2.2, we get ‖xm − xn‖2 ≤ ‖xm − x0‖2 − ‖xn − x0‖2. from step 2, we obtain that {xn} is cauchy. hence, there exists w ∈ d such that xn → w as n→∞. step 4. show that ‖zin − xn‖ → 0 as n→∞ for every i = 1, 2, ..,m. from xn+1 ∈ cn+1, we have ‖xn − yn‖ ≤ ‖xn − xn+1‖+ ‖xn+1 − yn‖ ≤ 2‖xn − xn+1‖ → 0 (7) 158 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . as n→∞. for v ∈ ⋂m i=1 f (ti) ∩ ep (f), by lemma 2.4 we have ‖yn − v‖2 = ‖ m∑ i=0 αi nz i n − v‖2 ≤ m∑ i=0 αi n‖zin − v‖2 − m∑ i=1 α0 nα i n‖zin − un‖2 ≤ α0 n‖un − v‖2 + m∑ i=1 αi nd(zin, tiv)2 − m∑ i=1 α0 nα i n‖zin − un‖2 ≤ α0 n‖un − v‖2 + m∑ i=1 αi nh(tiun, tiv)2 − m∑ i=1 α0 nα i n‖zin − un‖2 ≤ ‖un − v‖2 − m∑ i=1 α0 nα i n‖zin − un‖2 ≤ ‖xn − v‖2 − m∑ i=1 α0 nα i n‖zin − un‖2. this implies that m∑ i=1 α0 nα i n‖zin − un‖2 ≤ ‖xn − v‖2 − ‖yn − v‖2 ≤ m‖xn − yn‖, where m = supn≥0{‖xn − v‖+ ‖yn − v‖}. by our assumptions and (7), we obtain ‖zin − un‖ → 0 as n→∞, ∀i = 1, 2, ...,m. (8) from lemma 2.7, we obtain ‖un − v‖2 = ‖trnxn − trnv‖2 ≤ 〈trnxn − trnv, xn − v〉 = 〈un − v, xn − v〉 = 1 2 { ‖un − v‖2 + ‖xn − v‖2 − ‖xn − un‖2 } , hence ‖un − v‖2 ≤ ‖xn − v‖2 − ‖xn − un‖2. advances in systems science and applications (2011), vol. 11, no. 1-2 159 therefore, by lemma 2.4, we get ‖yn − v‖2 = ‖ m∑ i=0 αi nz i n − v‖2 ≤ m∑ i=0 αi n‖zin − v‖2 ≤ α0 n‖un − v‖2 + m∑ i=1 αi nd(zin, tiv)2 ≤ α0 n‖un − v‖2 + m∑ i=1 αi nh(tiun, tiv)2 ≤ ‖un − v‖2 ≤ ‖xn − v‖2 − ‖xn − un‖2. it follows that ‖xn − un‖2 ≤ ‖xn − v‖2 − ‖yn − v‖2 ≤ m‖xn − yn‖, where m = supn≥0{‖xn − v‖+ ‖yn − v‖}. from (7), we obtain ‖xn − un‖ → 0 as n→∞. (9) from (8) and (9), we have ‖xn − zin‖ ≤ ‖xn − un‖+ ‖un − zin‖ → 0 as n→∞. (10) step 5. show that w ∈ ⋂m i=1 f (ti) ∩ ep (f). from (9) and lim infn→∞ rn > 0, we get ‖xn − un rn ‖ = 1 rn ‖xn − un‖ → 0 , n→∞. (11) from xn → w as n → ∞ and (9), we obtain also that un → w. we shall show that w ∈ ep (f). by un = trnxn, we get f(un, y) + 1 rn 〈y − un, un − xn〉 ≥ 0, ∀y ∈ d. from the monotonicity of f , we have 1 rn 〈y − un, un − xn〉 ≥ f(y, un), ∀y ∈ d, hence 〈y − un, un − xn rn 〉 ≥ f(y, un), ∀y ∈ d. from (11) and condition (a4), we have 0 ≥ f(y, w), ∀y ∈ d. 160 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . for t with 0 < t ≤ 1 and y ∈ d, let yt = ty+ (1− t)w. since y, w ∈ d and d is convex, then yt ∈ d and hence f(yt, w) ≤ 0. so, we have 0 = f(yt, yt) ≤ tf(yt, y) + (1− t)f(yt, w) ≤ tf(yt, y). dividing by t, we obtain f(yt, y) ≥ 0, ∀y ∈ d. letting t ↓ 0 and from (a3), we get f(w, y) ≥ 0, ∀y ∈ d. therefore, we obtain w ∈ ep (f). next, we will show that w ∈ ⋂m i=1 f (ti). for each i = 1, 2, ...,m, we have d(w, tiw) ≤ ‖w − xn‖+ ‖xn − zin‖+ d(zin, tiw) ≤ ‖w − xn‖+ ‖xn − zin‖+h(tiun, tiw) ≤ ‖w − xn‖+ ‖xn − zin‖+ ‖un − w‖. it follows from step 4 that d(w, tiw) = 0 and thus w ∈ f (ti) for all i = 1, 2, ...,m. step 6. show that w = p⋂m i=1 f (ti)∩ep (f)x0. since xn = pcnx0, by lemma 2.1, we have 〈z − xn, x0 − xn〉 ≤ 0 for all z ∈ cn. since w ∈ ⋂m i=1 f (ti) ∩ ep (f) ⊂ cn, we get 〈z − w, x0 − w〉 ≤ 0 for all z ∈ ⋂m i=1 f (ti)∩ep (f). again by lemma 2.1, we obtain thatw = p⋂m i=1 f (ti)∩ep (f)x0. this completes the proof. corollary 3.3. let d be a nonempty, closed and convex subset of a real hilbert space h . let f be a bifunction from d × d to r satisfying (a1)-(a4) and let t : d → cb(d) be a nonexpansive multi-valued map for all i = 1, 2, ...,m with f (t ) ∩ ep (f) 6= ∅ and tp = {p} for each p ∈ f (t ). assume that βn ∈ (0, 1) with 0 < lim infn→∞ βn ≤ lim supn→∞ βn < 1 and rn ∈ (0,∞) with lim infn→∞ rn > 0. for an initial point x0 ∈ d = c0, compute the sequence {xn} by the iterative process f(un, y) + 1 rn 〈y − un, un − xn〉 ≥ 0, ∀y ∈ d, yn = βnun + (1− βn)zn, zn ∈ tun, cn+1 = {z ∈ cn : ‖yn − z‖ ≤ ‖xn − z‖}, xn+1 = pcn+1x0, n ≥ 0, then the sequence {xn} converges strongly to pf (t )∩ep (f)x0. proof. putting t1 = t and ti = i for i = 2, 3, ...,m, where i : d → cb(d) such that id = {d} for all d ∈ d in theorem 3.2, we obtain the desired result directly from theorem 3.2. advances in systems science and applications (2011), vol. 11, no. 1-2 161 corollary 3.4. let d be a nonempty, closed and convex subset of a real hilbert space h . let t : d → cb(d) be a nonexpansive multi-valued map for all i = 1, 2, ...,m with f (t ) ∩ ep (f) 6= ∅ and tp = {p} for each p ∈ f (t ). assume that βn ∈ (0, 1) with 0 < lim infn→∞ βn ≤ lim supn→∞ βn < 1. for an initial point x0 ∈ d = c0, compute the sequence {xn} by the iterative process yn = βnxn + (1− βn)zn, zn ∈ txn, cn+1 = {z ∈ cn : ‖yn − z‖ ≤ ‖xn − z‖}, xn+1 = pcn+1x0, n ≥ 0, then the sequence {xn} converges strongly to pf (t )x0. proof. putting f(x, y) = 0 for all x, y ∈ d in corollary 3.3, we obtain the desired result directly from corollary 3.3. the main result of this paper holds true under the assumption that tp = {p} for all p ∈ f (t ). this condition was introduced by shahzad and zegeye [17]. the following example gives an example of a nonexpansive multi-valued map t which satisfies the property that tp = {p} for all p ∈ f (t ) and tx is not a singleton for all x /∈ f (t ). example. consider d = [0, 1]× [0, 1] with the usual norm. define t : d → cb(d) by t (x, y) =  {(x, 0)}, x 6= 0, y = 0 {(0, y)}, x = 0, y 6= 0 {(x, 0), (0, y)}, x, y 6= 0 {(0, 0)}, x, y = 0. open problem: can we drop the condition that tp = {p} for all p ∈ f (t ) in the main result of this paper ? references [1] bauschke,h. h., matoušková,e., reich,s. (2004). projection and proximal point methods: convergence results and counterexamples. nonlinear anal. 56 (2004) 715-738. [2] blum,e., & oettli,w. (1994). from optimization and variational inequalities to equilibrium problems. math. student 63 (1994) 123-145. [3] ceng,l.-c., & yao,j.-c. (2008). a hybrid iterative scheme for mixed equilibrium problems and fixed point problems. j. comput. appl. math. 214 (2008) 186-201 [4] cholamjiak,p. (2009). a hybrid iterative scheme for equilibrium problems, variational inequality problems and fixed point problems in banach spaces. fixed point theory appl. volume 2009 (2009). article id 719360. 18 pages. [5] cholamjiak,p. & suantai,s. (2009). a new hybrid algorithm for variational inclusions, generalized equilibrium problems and a finite family of quasi-nonexpansive mappings. fixed point theory appl. volume 2009 (2009). article id 350979. 20 pages. 162 suantai:monotone hybrid methods for a finite family of nonexpasive. . . . . . [6] combettes,p.l. & hirstoaga,s.a. (2005). equilibrium programming in hilbert spaces. j. nonlinear convex anal. 6 (2005) 117-136. [7] genal,a. & lindenstrass,j. (1975). an example concerning fixed points. israel j. math. 22 (1975) 81-86. [8] kangtunyakarn,a. & suantai,s. (2009). hybrid iterative scheme for generalized equilibrium problems and fixed point problems of finite family of nonexpansive mappings, nonlinear analysis: hybrid systems. 3 (2009) 296-309. [9] kim,t.h. & xu,h.k. (2006). strongly convergence of modified mann iterations for with asymptotically nonexpansive mappings and semigroups, nonlinear anal. 64 (2006) 11401152. [10] mann,w.r. (1953). mean value methods in iteration, proc. amer. math. soc. 4 (1953) 506-510. [11] marino,g. & xu,h.k. (2007). weak and strong convergence theorems for strict pseudocontractions in hilbert spaces. j. math. anal. appl. 329 (2007) 336-346. [12] nakajo,k. & takahashi,w. (2003). strongly convergence theorems for nonexpansive mappings and nonexpansive semigroups. j. math. anal. appl. 279 (2003) 372-379. [13] panyanak,b. (2007). mann and ishikawa iterative processes for multivalued mappings in banach spaces. comput. math. appl. 54 (2007) 872-877. [14] peng,j.-w. , liou,y.-c. , yao,j.-c. (2009). an iterative algorithm combining viscosity method with parallel method for a generalized equilibrium problem and strict pseudocontractions. fixed point theory appl. volume 2009 (2009). article id 794178. 21 pages. [15] reich,s. (1979). weak convergence theorems for nonexpansive mappings in banach spaces. j. math. anal. appl. 67 (1979) 274-276. [16] sastry,k.p.r. & babu,g.v.r. (2005). convergence of ishikawa iterates for a multivalued mappings with a fixed point. czechoslovak math. j. 55 (2005) 817-826. [17] shahzad,n. & zegeye,h. (2009). on mann and ishikawa iteration schemes for multivalued maps in banach spaces. nonlinear analysis 71 (2009) 838-844. [18] shiau,c. , tan,k.k. , wong,c.s. (1975). quasi-nonexpansive multi-valued maps and selection. fund. math. 87 (1975) 109-119. [19] song,y. & wang,h. (2008). erratum to ”mann and ishikawa iterative processes for multivalued mappings in banach spaces”[comput. math. appl. 54 (2007) 872-877]. comput. math. appl. 55 (2008) 2999-3002. [20] tada,a. & takahashi,w. (2007). weak and strong convergence theorems for a nonexpansive mapping and an equilibrium problem. j. optim theory appl. 133 (2007) 359-370. [21] xu,h.k. (1991). inequality in banach spacees with applications. nonlinear. anal. 16 (1991) 1127-1138. advances in systems science and application (2015) vol.15 no.3 279-298 simd acceleration of spmv kernel on multi-core cpu architecture j.saira banu and dr m.rajasekhara babu school of computing science and engineering, vit university, vellore, india. abstract as the field of sparse matrix vector multiplication (spmv) matures and its breadth of application increases, the need for parallel implementation becomes necessary. spmv is proved to be a bottleneck due to its irregular access patterns. various storage formats for sparse matrices have been proposed to solve this issue. the quad tree-compressed row storage (qcsr) format shows good performance improvement over other formats such as compressed row storage (csr) and blocked csr (bcsr) for spmv. this paper extends qcsr format to exploit the single instruction multiple data (simd) registers which are available in current processors. programming with simd registers in a single core processor achieves parallelism with reduced power and without any additional hardware requirement, as in the case of graphics processing unit (gpu) computing. to program effectively in simd units intels streaming simd extension (sse) instructions are used. in this paper computational performance of qcsr-spmv is determined for over a collection of 10 benchmark matrices on simd units of x86 architecture. experimental results demonstrate qcsr-simd achieves significant average speedup of 2.0 x compared to csr-simd. keywords openmp; gpu; sparse matrix; spmv; csr; qcsr; simd; sse; performance; programming model; optimization; parallel computing. 1 introduction there are multiple ways of accomplishing parallel computing in single machine to cloud environment. in single machine, parallel computing is achieved through thread level parallelism with open multiprocessing (openmp), message passing method with message passing interface (mpi) and through simd technique with sse instruction. increasing clock speed has been ceased recently which gave a way for multi or many cores cpu. the other alternative method to achieve many core computing is through gpu. in gpu, parallelization is achieved with the expense of new hardware. parallel computing in cloud environment is realized by distributing the computing task to various computing resources and gathering the result. in cloud computing allocating task is a major issue. guiyi wei, et al addressed this issue of allocating the task to various computing resources in the cloud[1]. it involves allocation matrix and expense matrix. if a matrix used in an algorithm is dense then there is no issue but algorithm involving sparse matrix is a concern one. 280 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... a lot of work is being done in the field of sparse matrices and spmv, to improve the efficiency. emphasis has been laid on the parallel implementation of spmv in the past decade. sparse matrix structures are present in various applications, and efficient method for optimizing the performance of these applications is a critical issue. spmv is an operation found in many computational science applications. improving the speedup of spmv boosts the performance of these applications. extensive research has been done in improving the performance of spmv using thread level and distributed memory parallelism. samuel williams et al. examined spmv on various multicore design namely amd quad core, intel quad core[2], etc. shengfei liu, et al. worked on spmv with compressed sparse row (csr) and blocked csr (bcsr) format[3]. they have implemented multithreaded spmv using openmp. partitioning the matrices in to sub matrices and hybrid programming model with mpi and openmp are the future methods proposed in the paper to increase the performance of the kernel. dakuan discussed spmv with mpi[4]. gerald schubert et al. examined the performance of spmv kernel with hybrid programming (openmp and mpi) model[5]. this paper does not explore the load balancing issues. m.krotiewski and m.dabrowski presented a parallel implementation of spmv on multicores with wellknown optimization such as matrix reordering, blocking and prefetching[6]. with the rise of high performance computing, most of the focus is now on gpu computing. in literature, different formats of storage have been developed to either improve the space efficiency, or the time of access of non-zero elements for various operations. some formats are specific to the execution environment such as cpu or gpu. tomas oberhuber et al. proposed a new format such as new row grouped csr which performed well in gpu devices[7]. jilin zhang et al. have implemented spmv using a new type of storage format, quad tree csr (qcsr) format using gpus[8]. this format outperforms the blocked csr (bcsr) format. recently vectorization capability using sse instructions of the current processor proved to be better technology to improve the performance of many applications. susana ladra, et al. exploited simd instructions in current processors to improve the classical string algorithms[9]. kai zeng et al. proved the usage of simd technique in computed tomography ct reconstruction is efficient rather than gpu implementation[10]. image processing applications such as feature detection, stereo vision class model estimation, and object detection achieved better performance on the compressed sparse extended ( csx ) simd architecture compared with gpu[11]. s.j.pennycook et al. explored the use of simd registers for molecular dynamics problem set[12]. to obtain better performance on numerical kernels, martin kong et al. proposed a 3-step framework such as data locality, multicore parallelism and simd advances in systems science and application (2015) vol.15 no.3 281 execution of programs[13]. libo huang et al. presents a dynamic vectorization method to address the constraints in current simd engines such as register variation in the processors requires changing of vector operands and aligned memory accesses[14]. nasersedaghati proposed an extension to the instruction available in instruction set architecture (isa) to overcome the disadvantage of simd for stencil computation application[15]. neil g. dickson et al. found that explicit vectorization on the cpu gave 9x-12x speedup over the original cpu version, and was also found to be 2x faster than the fully optimized gpu version that uses explicit memory coalescing. these papers signifies the importance of single core optimization on cpu[16]. ji-lin zhang et al. have implemented spmv using csr format with simd methodology and thread-level parallelizing[17]. they found out that these methods gave a speed up of around 2.11 over that of the non-optimized method. kai zhang et al. improved upon this and broke the performance bottlenecks of the simd processors like the low utilization of simd processors and the memory bandwidth[18]. their method showed a good speedup over the normal csr vector kernel. nazligoharian et al. demonstrated that csr is the best storage format for information retrieval and query processing among various storage format of sparse matrix[19].these paper explored simd in spmv and its importance. in this paper, we investigate the impact of simd acceleration on spmv kernel. for this, spmv kernel with csr and qcsr format has been accelerated, i.e., simd is applied to all suitable operations and implemented in haswell processor using sse instructions. the main contribution of this paper are as follows. 1. made an extensive study on strorage formats of sparse matrix structure and analyzed the space complexity for benchmark matrices from different applications. 2. simd acceleration is presented for all the data-parallel kernnels of spmv with csr and qcsr format. 3. implementation and optimization is performed with sse instructions. 4. simd implementation of csr outperforms the na?ve and thread level parallelism. 5. with simd optimization,spmv with qcsr format gives 2 fold of speedup compared to csr-simd. the rest of the paper is organized as follows: section 2 deals with architecture specifications, section 3 describes the spmv algorithm with its storage format analysis; section 4 explores about simd optimization of spmv kernel with csr and qcsr storage formats; section 5 deals with the results and observations; section 6 talks about the conclusion. 282 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... 2 architecture specification since the beginning of the processor, its performance has been steadily increasing. developments in semiconductor capability are the main reason for this remarkable change. advancement in architectural features like instructionlevel parallelism, data-level parallelism (dlp) thread-level parallelism (tlp) and memory-level parallelism (mlp) brought remarkable improvement in performance. here, in this paper dlp is achieved through simd, and tlp is done through openmp. 2.1 dlp-simd simd is one of the classifications of the flynns taxonomy of computer architec− ture[20]. it exploits dlp by performing the same operation on multiple data points simultaneously. simd architecture consists of multiple processing elements that perform identical task concurrently on the data elements as shown in fig.1. fig. 1 simd architecture. sse instructions are now available in current generation processors to program effectively with simd registers. these instructions are used effectively to vectorise the code and to modify the data layout. automatic vectorization has certain limitations such as handling loops of irregular access patterns, pointers, etc. since automatic vectorization has restricted applicability, programmers write vectorised code using intrinsic which is directly expanded to machine instructions. programming using intrinsic are tied to a specific instruction set. with the help of these sse instructions, simd programming is performed in high level programming languages such as c. compared to gpu, there is no overhead incurred, like moving data from device to device or thread processing. advances in systems science and application (2015) vol.15 no.3 283 it can also be combined with thread level parallelism technique to increase the speed up further. sse is included from pentium iii processors by adding 128 bit registers and the instructions that can operate on them. advanced intel processors supports avx which includes 256 bit register with extra simd instructions in addition to sse instruction set. sse instruction set includes integer and floating point arithmetic instructions, comparison, shuffling, data type conversion, bitwise operations, minimum, maximum, conditional copies, crc 32 and population count. currently in many applications such as gaming, graphics, physics, and mathematics, simd instructions are used to perform shuffling, scalar product, checksum calculations and complex operations. there are various versions of sse such as ssesse4.2 which are capable of processing 2 double precision floating point numbers or 4 single precision 32 bit floating point numbers. sse intrinsic is used to improve the performance of an application which has fine grain parallelism in it. when dealing with sparse matrices, one of the most important restrictions is that the average width of non-zero elements per row must be greater than the simd width of the computer device. else, the simd registers are only partially utilized, and do not give a considerable increase in the efficiency. 2.2 tlp openmp thread level parallelism is achieved when a single process is divided into various subparts, and each part is carried out by a different thread, so at the end, it could be combined to give the result. openmp is an api that provides shared memory multi-processing programming using high level languages such as c, c++ and fortran. there are various methods to perform thread level parallelism for spmv-csr kernel. parallelism can be done row wise or column wise or block wise. here, we have performed row wise parallelism, where each row is assigned to a thread. the accesses to the data are independent and hence can easily be parallelized. one advantage of using the row wise partitioning is that individual threads operate on different parts of the final resultant array unlike that of the column wise partitioning, where all threads have to write to all parts of the array. 3 spmv algorithm spmv forms a basic computational kernel in many scientific and industrial applications. modern high performance computer systems (hpc) such as multicores, gpu,simd computation using coprocessor and special xmm registers rely on spmv computation for numerous task. spmv is usually represented as y = a ∗x where a represents a sparse matrix and y and x represents a vector. sparse matrix is a matrix which has more number of zero elements than the non-zero values. storing a sparse matrix as it is in a memory is a space overhead and in some applications it doesnt fit in to the available memory. to effectively reduce the storage space many data structures have been proposed in the literature. 284 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... these data structure stores only the non-zero elements of the matrix and thereby reduces the computation of spmv algorithm. this is important in hpc community since all modern computer system is equipped with small storage devices. the compression of sparse matrix using efficient storage format influence the performance of spmv operation.for the remainder of the paper, we use the following notation: n : total number of non-zero values n : order of the matrix s( ): storage space occupied nr : number of rows 3.1 storage formats sparse matrices are stored in computer memory with specific formats that give high performance of spmv operation. few of the popular formats of storing sparse matrices with space and time complexity are discussed here.equations 1-5 represents the space complexity of various storage formats. 3.1.1 the coordinate format (coo) this is one of the most general purpose formats. it consists of three 1-dimensional arrays one to store the non-zero values, one to store the corresponding column index, and the other to store the corresponding row index [21]. fig. 2 example matrix. for example, the coo format for the matrix shown in fig.2 is represented as : data = [1 1 1 1 1] col = [1 1 2 2 4] row = [1 2 2 3 3]. a time complexity : the time complexity to convert this format to/from the compressed row storage (csr) format is o (n+n) b space complexity : the space complexity of this format excluding the data array is given by s(coo) = 2 ·n · s(n) (1) advances in systems science and application (2015) vol.15 no.3 285 3.1.2 the compressed row storage format (csr) this is one of the most efficient and hence most popular storage formats. it consists of three 1-dimensional arrays c one to store the non-zero value, one to store the corresponding column index, and the other to store the row pointer value, that gives the total non-zero values in each row [21] for the example in figure 2: csr format is given as: data = [1 1 1 1 1] col = [1 1 2 2 4] row-ptr = [0 1 3 5 5] a. time complexity : the time complexity to convert this format to/from the coordinate (coo) format is o (n+n). b. space complexity : the space complexity of this format excluding the data array is given by s(csr) = n · s(n) + n · s(n) (2) 3.1.3 the quadtree storage format quadtree is a recursive tree data structure given by ivan simecek [22]. the matrix is recursively divided into four quadrants until each block is equal to a predefined size (called density). these blocks or nodes can be of three types: a. empty: the entire node is made up of zeroes. b. mixed: the node consist of a mixture of zeroes and non-zero elements. c. full: the entire node id made up of non-zero elements. empty nodes are ignored, and the mixed and full nodes are stored in any format that is found to be appropriate for the application.the matrix in fig.1, is divided in to four quadrants and each quadrant is checked for the node types. here the second quadrant is empty and it is ignored and all other quadrants are further retrieved through the indices for its operation as shown in fig.3. the advantages of this format are: a. easy conversion from popular formats. b. easy modifications to the data. c. the recursive style of programming leads to better performance due to better cache memory utilization. a. time complexity: the entire algorithm can be broken down into smaller functions. the time complexity of each such function is as given below: 286 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... (1) to find the total non-zero values in each quadrant: o (average per row. (y2 y1 + 1)). (2) to check the given quadrant is empty or not:o (log2 average per row. (y2 y1 + 1)). (3) for converting the matrix into quadtree format, the complexity depends on the density. b. space complexity: the space complexity completely depends on the density chosen and the creation of leaves. 3.1.4 the qcsr (quadtree compressed row storage) format the qcsr format is a combination of the quadtree format and the csr format(jilin zhang, etal, 2013)[8]. after dividing the matrix into various nodes using the quadtree logic, the mixed and full nodes are stored in the csr format. though it is found to have a space overhead, the implementation of spmv is much faster in this format. a. time complexity : the time complexity of this format is the same as that of the quadtree format, in addition to which each quadrant is converted into the csr format. b. space complexity : this format gives an overhead in space.the maximum overhead over the csr format is given by: soh(qcsr) = (2sr + sl)xo(4d− 1) (3) where, sr is the space occupied by the index pointer to the region, sl is the region length and d is the maximum depth of the tree. fig. 3 splitting of matrix. 3.1.5 minimal quadtree format: this format is developed by i.simecek et al.[23]. it is a derivative of the quadtree format and consists of a bit stream of 1s and 0s. the matrix is recursively divided advances in systems science and application (2015) vol.15 no.3 287 into blocks as in quadtree format, but here, each block is represented by a single bit0 if all the elements in the block are zeroes, 1 otherwise. the example given in fig. 1, can be divided into 4 blocks if the density is 2, and can be represented as 4 bits1011 for only the second block is an empty block, and the others have non-zero elements. a. time complexity : the total time complexity for this conversion is o(n(n+ √ n)) · log2avg per row b. space complexity : the minimal size of the mqt format is given by: s(mqtmin) = 4 · (n 3 + log4( n2 n )) (4) the maximal size of the mqt format is given by: s(mqtmax) = 4 · (1 3 + log4( n2 n )) (5) as minimization of memory is an important criterion for spmv operation, we compared thespace complexity of various formats for the benchmark matrices obtained from university of florida database [24]. table 1 shows the properties of benchmark matrices used. table 2 depicts the space occupied for these matrices in coo, csr, qcsr and minqcsr format. table 1 overview of the benchmark matrices benchmark matrices number of rows number of columns number of nonzero elements bcsstk07 420 420 4140 adder dcop 08 1813 1813 11242 mhd4800b 4800 4800 16160 meg4 5860 5860 26324 gemat11 4929 4929 33185 cell1 b 7055 7055 34855 ex12 3973 3973 42092 sina 5743 5743 102265 na5 5832 5832 155731 from table 2, it is evident that minimal qcsr, which is just a series of bits is well suited for storing the structure of matrices and hence not suitable for spmv operation. csr is the next format that has the least space complexity and proved to be well utilized in spmv operation. qcsr format has the space overhead but that is not accounted for spmv operation. hence, among all the 288 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... formats that have been developed so far, the csr format is the most efficient and effective format. a lot of improvements have been done on this format to improve its efficiency. in this paper, we analyse the csr format along with the qcsr format for spmv computation in simd architecture. table 2 the space complexity of various formats for benchmark matrices benchmark matrices original coo csr qcsr minqcsr bcsstk07 391kb 97kb 76kb 118kb 394b adder dcop 08 6.35mb 326kb 176kb 265kb 999b mhd4800b 44mb 436kb 263kb 416kb 810b meg4 65.7mb 425kb 386kb 854kb 739b gemat11 46.5mb 648kb 506kb 1.2mb 367b cell1 b 95.2mb 587kb 386kb 854kb 739b ex12 30.4mb 1.13mb 635kb 845kb 349b sina 63.6mb 2.95mb 1.52mb 3.19mb 46b na5 65.9mb 4.54mb 2.3mb 4.25mb 384b 3.2 spmv csr spmv kernel acts as a core kernel of many iterative algorithms and scientific applications. it is one of the time consuming kernel in these methods. optimizing this kernel plays a vital role in improving the performance of these applications. the performance of spmv kernel is based on the storage format used. csr format has been proved to be the best format with space efficiency. algorithm 3.1 and 3.2 shows spmv with csr format in naive and thread level parallelism. algorithm 3.1 spmv with csrnaive inputs ptr pointer array data data array column column index array x vector for multiplication output: z result array 1. begin 2. for each item i from 0 nr do 3. z[i] = 0 4. for each item j in ptr[i] to ptr[i+1] -1 do 5. temp = column [j] 6. z[i] = z[i]+ (data[j] * x[temp]) 7. end for 8. end for advances in systems science and application (2015) vol.15 no.3 289 9. end algorithm 3.2 spmv with csr-thread level parallelism inputs ptr pointer array data data array column column index array x vector for multiplication output: z result array 1. begin 2. ♯ pragma omp parallel num threads(4) 3. ♯ pragma omp parallel for private (k, j, i) schedule(static, 10) 4. for i = 0 to n 1 do 5. z[i] = 0 6. for j = ptr[i] to ptr[i+1]-1 do 7. temp = colj 8. z[i] = z[i] + (val[j] * x[temp]) 9. end for 10. end for 11. end 3.3 spmv qcsr as elucidated in section 3.1, the qcsr format is a combination of the quadtree format and the csr format. the given matrix is first divided into various quadrants using the quadtree format, and the mixed and full nodes are converted into the csr format as shown in algorithm 3.3 and 3.4 [8]. here, sr stands for start row, er for end row, sc for start column, ec for end column, a[][] is the input matrix, and density is the size of the quadrant. qcsr format compared to csr format has space overhead. this overhead is due to storage of block and intermediate node information. this overhead accounts for transformation and representation of matrix and not in spmv kernel. this implies that qcsr is suitable for spmv kernel with all general cases. algorithm 3.3 matrix conversion to qcsr format inputs sr start row sc start column 290 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... er end row ec end column 1. begin 2. int mr=(sr+er)/2; 3. int mc=(sc+ec)/2; 4. int r=er-sr; 5. int c=ec-sc; 6. if( r≥ density && c ≥ density ) 7. { 8. if( !isempty(sr,mr,sc,mc) ) 9. trans(sr,mr,sc,mc); 10. if( !isempty(sr,mr,(mc+1),ec) ) 11. trans(sr,mr,(mc+1),ec); 12. if( !isempty((mr+1),er,sc,mc) ) 13. trans((mr+1),er,sc,mc); 14. if( !isempty((mr+1),er,(mc+1),ec) ) 15. trans((mr+1),er,(mc+1),ec); 16. } 17. else if(c>density) 18. { 19. trans(sr,er,sc,mc); 20. trans(sr,er,(mc+1),ec); 21. } 22. else if(r>density) 23. { 24. trans(sr,mr,sc,ec); 25. trans((mr+1),er,sc,ec); 26. } 27. else 28. { 29. if( !isempty(sr, er, sc, ec)) 30. convert the quadrant into csr format 31. } 32. end algorithm 3.4 isempty(sr,er,sc,ec) inputs sr start row sc start column advances in systems science and application (2015) vol.15 no.3 291 er end row ec end column 1. begin 2. int isempty() 3. for all elements i from sr to er 4. { 5. for all elements j from sc to ec 6. { 7. if( a[i][j] != 0) 8. return 0 9. } 10. } 11. return 1 12. end 4 simd optimization 4.1 spmv ccsr simd optimization is performed with the help of sse 4.2 instructions. here, nr is the total number of rows in the matrix, val is the array used to store the nonzero elements, col is the array to store the column indices, row is the row pointer array, and x is the vector for multiplication. algorithm 3.5 gives simd optimization of spmv kernel with csr format. in vectorization using simd technique, four values are operated simultaneously, as the size of the simd register is 4. the non-zero elements were loaded into the data register, and the corresponding vector elements in the x register. here data and x are the special xmm registers available in the current generation processors. these values are multiplied simultaneously, hence decreasing the total number of operations performed. algorithm 3.5 spmv with csr-simd inputs nr c number of rows val c data array row c pointer array x vector for multiplication output: z result array 1. begin 292 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... 2. for every element i from 0 to nr-1 do 3. for every element j in row[i] to row[i+1]-1 do 4. data = {val[j] , val[j+1], val[j+2], val[j+3]} 5. x = {x[col[j]] , x[col[j+1]], xcol[j+2]], x[col[j+3]]} 6. y = mm mul ps(data, x) 7. zi = sum{y} 8. j = j + 4 9. end for 10. end for 11. end 4.2 spmv -qcsr qcsr format is an improvisation on the quad tree format, in which each of the blocks is stored in the csr format. we have performed sparse matrix vector multiplication using this format. one condition that has to be met to maximize efficiency is that, the total non-zero values in each block should be greater than or equal to the size of the simd register. if not, the remaining bits in the simd register must be padded with zeros which results in time overhead. qcsr with sse instructions is specified in algorithm 3.6. the following instructions are used to vectorize the spmv implementation of qcsr format. 1. mm loadu ps(array) : this is used for loading the registers with the elements in the array. 2. mm mul ps(reg1, reg2) : this is used to perform vectorized multiplication of the elements in the two mentioned registers. the result is stored in reg1. 3. mm hadd ps(reg1, reg2) : this is used to perform vectorized addition of the elements in the two mentioned registers. the result is stored in reg1. 4. mm storeu ps(ptr, reg) : this is used to store the results in the register to the pointer. algorithm 3.6 qcsr with simd optimization each block in qcsr format consists of the following components: startrow: the row index of the first row in the block. startcolumn: the column index of the first column in the block. endrow: the row index of the last row in the block. endcolumn: the column index of the last column in the block. structure csr: csr representation of the block, which contains advances in systems science and application (2015) vol.15 no.3 293 nz: the total non-zeroes values in the block data[n]: data array for the block. ptr[n]: row-pointer array for the block col[n]: column indices array for the block multiply (each block) i: to keep track of the number of rows in the block. k: to keep track of the number of non-zero elements in each row. t: to keep track of the ptr arrays index. temparr: to store the padded data values. data1, vect1, temp, res1: simd registers. res: to store the multiplication result. 1. begin 2. for i from startrow to endrow, do 3. for k from csr.ptr[t] to csr.ptr[t+1], do 4. q= csr.ptr[t+1] c csr.ptr[t]; 5. a = minimum ( q-k, 4) 6. for j from 0 to a, do 7. temparr[j]= csr.data[k+j]; 8. if j < 4 9. pad the remaining places with 0. 10. end if 11. data1= mm loadu ps ( temparr ) 12. vect1= mm loadu ps ( v[col[k]], v[col[k+1]], v[col[k+2]], v[col[k+3]] ) 13. temp= mm mul ps (data1 , vect1) 14. res1= mm hadd ps (temp , temp ) 15. res1= mm hadd ps (res1 , res1) 16. mm storeu ps ( c , res1) 17. res[i] = res[i] + c 18. end for 19. t++ 20. k=k+4 21. end for 22. end 294 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... 5 results and discussion in this section we analyse the performance improvement of csr-spmv and qcsrspmv format using instructions included in sse 4.1/4.2 extension. to carry out experiments, we have used intel(r) coretm i5-4200u cpu@ 1.60 ghz (2 cores) with 4 gb ram. visual studio 2010 compiler was used for compilation. the width of the simd register in the test system is 4. in this experiment, the density size has been chosen to be 16 such that the nonzero elements in each block are equal to or greater than the width of simd register. speed-up is a performance index, which is a ratio of the time taken to implement a program sequentially to the time taken to do the same in parallel. we first performed spmv using the csr format both sequentially and in parallel using openmp. in parallel implementation, we have considered the block size to be 10. each thread is given a single row and multiplication of these rows is performed in parallel. one disadvantage of this method is, if the total non-zero elements in one of the rows are much greater than that in the other rows, then stalling of the other threads occurs. the remaining threads have to be idle until the thread with that particular row finishes execution. next, we implemented the same using simd vectorization and compared this with that of sequential execution. the simd registers are of size 128 bits, and hence 4 32-bit integers can be loaded in it at a given time. each row is considered separately, and the nonzero elements in it are loaded into the register, 4 at a time. if the number of non-zero elements left is less than 4, it is padded with zeroes. this overcomes the disadvantage of stalling that occurs in thread level parallelism. another overhead that this method avoids is the time taken for the creation of threads. the efficiency of the execution of spmv using simd completely depends on the arrangement of the non-zero elements in the matrix. fig. 4 shows results obtained when implementing spmv using csr format with simd and openmp techniques. from the graph shown in fig. 4, we determine that the simd vectorization technique gives better results when compared to that of sequential and parallel execution. hence, we implemented simd version of spmv for various benchmark matrices using the qcsr format using this technique of simd vectorization. in this method of implementation, each quadrant is considered separately and is treated as a separate matrix. csr spmv using vectorization, as discussed above, is then carried out for each quadrant. here, the size of the quadrant depends on the density chosen. since the density is usually in powers of two and much smaller than the original matrix size, the probability that the total non-zero elements in each row in the quadrant is less than 4 is quite high. hence the number of times the simd registers have to be loaded and padded is reduced. these results were compared with the advances in systems science and application (2015) vol.15 no.3 295 results obtained from implementing spmv with regular csr format with simd vectorization. fig. 5 shows the graph of the speed-up obtained by the qcsr format. the graph clearly shows that the qcsr format when implemented with simd is more efficient than that of the csr format with simd. fig. 4 computing time for optimization schemes using csr. fig. 5 computing time for optimization schemes using csr. 6 conclusions this paper explores the use of the simd technique provided in current processors, to accelerate the performance of spmv using qcsr format. as simd units are readily available with the current generation processors, and do not incur ad296 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... ditional processing cost unlike gpu programming and threading, it is necessary to avail the feature of this hardware to improve the performance of various applications with reduced power . results show that the speedup of spmv qcsr with simd is an average of 2, compared to that of spmv csr with simd, for the benchmark matrices chosen. it is also apparent that the efficiency of this technique completely depends on the arrangement of the non-zero elements per row. references [1] guiyi wei, athanasios v. vasilakos , yao zheng and naixue xiong, (2010), “a game-theoretic method of fair resource allocation for cloud computing services”, the journal of supercomputing, springer, vol. 54, no. 2, pp. 252269. [2] samuel williams, leonid oliker, richard vuduc, john shalf, katherine yelick, james demmel (2009), “optimization of sparse matrixcvector multiplication on emerging multicore platforms”, journal parallel computing, vol. 35, no.3, pp. 178-194. [3] shengfei liu, yunquan zhang, xiangzheng sun and rongrong qiu (2009), “performance evaluation of multithreaded sparse matrix-vector multiplication using openmp,” 11 th international conference on high performance computing and communications. [4] dakuan (2009), “cui parallel sparse matrix algorithms for numerical computing matrixvector multiplication”, iims postgraduate seminar. [5] schubert, g., hager, g., fehske, h. and wellein, g. (2011), ”parallel sparse matrix-vector multiplication as a test case for hybrid mpi+ openmp programming”. in parallel and distributed processing workshops and phd forum (ipdpsw), 2011 ieee international symposium, pp. 1751-1758 . ieee. [6] m.krotiewski, m.dabrowski (2010), “parallel symmetric sparse matrixvector product on scalar multi-core cpus”, journal parallel computing, vol. 36, no. 4, pp. 181-198. [7] tomas oberhuber, atsushi suzuki and jan vacata, (2011), “new rowgrouped csr format for storing sparse matrices on gpu with implementation in cuda”, corr journal, vol abs/1012.2270, 2010. [8] jilin zhang, enyi liu, jian wan, yongjian ren, miao yue and jue wang,(2013), “implementing sparse matrix-vector multiplication with qcadvances in systems science and application (2015) vol.15 no.3 297 sr on gpu”, international journal on applied mathematics & information sciences, vol.7, pp. 473-482. [9] susana ladra, oscar pedeira, jose duato and nieves r.brisaboa, (2012), “exploiting simd instructions in current processors to improve classical string algorithms”, advances in databases and information systems, lecture notes in computer science, vol.7503, pp. 254-267. [10] kai zeng, erwei bai and ge wang,(2007), “a fast ct reconstruction scheme for a general multi-core pc”, international journal of biomedical imaging, vol. 2007 no. 1, pp. 1-1. [11] amir fijani, fouzhan hosseini,(2011), “image processing applications on a low power highly parallel simd architecture”, aero ’11 proceedings of the 2011 ieee aerospace conference, pp. 1-12. [12] s. j. pennycook, c. j. hughes, m. smelyanskiy and s. a. jarvis, (2013), “exploring simd for molecular dynamics, using intel r xeonrprocessors and intelr xeon phi tm coprocessors”, ipdps ’13 proceedings of the 2013 ieee 27th international symposium on parallel and distributed processing, pp. 1085-1097. [13] martin kong, richard veras and kevin stock (2013), “when polyhedral transformations meet simd code generation”, proceedings of the 34th acm sigplan conference on programming language design and implementation, pp. 127-138. [14] libo huang, zhiying wang,nong xiao, yongweng wang and qiang dou, (2013), “dynamic streamization model execution for simd engines on multicore architectures”, ieee transactions on computer -aided design of integrated circuits and systems, vol. 32, no. 11. [15] nasersedaghati, renji thomas, louis pouchet, radu teodorescu and p. sadayappan,(2011), “stvec: a vector instruction extension for high performance stencil computation”, pact ’11 proceedings of the 2011 international conference on parallel architectures and compilation techniques, pp. 276-287. [16] neil g. dickson, kamran karimi, firashamze (2011), “importance of explicit vectorization for cpu and gpu software performance”, journal of computational physics, vol. 230, no. 13, pp.5383-5398. 298 j.saira banu and m.rajasekhara babu:simd acceleration of spmv kernel on ... [17] ji-lin zhang, li zhuang, jian wan, xiang-hua xu, cong-feng jiang and yong-jian ren, (2011), “combine optimized sparse matrix-vector multiplication for csr format”, proceeding chinagrid ’11 proceedings of the 2011 sixth annual chinagrid conference, pp. 124-129. [18] kai zhang, shuming chen, yaohuawang and jianghuawan,(2013), “breaking the performance bottleneck of sparse matrix vector multiplication on simd processors”, ieice electronics express, vol. 10, no.9 pp. 1-7. [19] nazligoharian, ankit jain and qian sun (2012), “comparative analysis of sparse matrix algorithms for information retrieval”, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.14.8914. [20] roland leiba, sebastian hack and ingowala,(2012), “extending a c-like language for portable simd programming”, proceedings of the 17th acm sigplan symposium on principles and practice of parallel programming, pp. 65-74. [21] nathan bell and michael garland, (2008), efficient sparse matrix-vector multiplication on cuda, nvidia technical report nvr-2008-004. [22] ivan simecek, (2009). “sparse matrix computations using the quadtree storage format”, synasc ’09 proceedings of the 2009 11th international symposium on symbolic and numeric algorithms for scientific computing, pp. 168-173. [23] i.simecek, d.langer, p.tvrdik, (2012), “minimal quadtree format for compression of sparse matrices storage”, synasc ’12 proceedings of the 2012 14th international symposium on symbolic and numeric algorithms for scientific computing, pp. 359-364. [24] t.davis (1997), “the university of florida sparse matrix collection” [eb/ol], http://www.cise.ufl.edu/research/sparse/matrices. corresponding author j.saira banu can be contacted at: jsairabanu@vit.ac.in microsoft word 13-杨冰陈喜等.doc issn 1078-6236 international institute for general systems studies, inc. real-time 3d smoking simulation based on computational fluid dynamics and gpu computation for virtual surgery� bing yang, xi chen, cheng yang, jiaxiang guo, jie han, jiayu tang and zhiyong yuan school of computer, wuhan university, wuhan 430072, china email: zhiyongyuan@whu.edu.cn (* corresponding author) abstract in this paper a smoking simulation model based on computational fluid dynamics (cfd) for virtual surgery is proposed. navier-stokes equations are used to model smoke as a constant and incompressible fluid. in order to realize the visual effects of real-time 3d smoking, we utilize gpu (graphics processing unit) and cuda (compute unified device architecture) computing technology to perform the massive parallel computation by executing multithreading concurrently. the rendering of the smoke is implemented by the ray-tracing algorithm based on hardware. the visual effects of the real-time dynamic smoking simulation are given by the paper. the results show that the methods applied to the smoking simulation are elucidating and practical. keywords virtual surgery cfd gpu 3d smoking simulation 1. introduction virtual surgery training system based on virtual reality can be used to improve the skill level of the surgeons and to reduce the incidence of complications in surgical practice. the ultimate advantage of this simulation system is to bridge the gap between the basic training of surgeon and the practical operation on patients, which means a real configurable training environment without the limitation of repeated training. however, the current surgical training systems have not yet reached the necessary level of sufficient reality, and the application of the simulation technology to the field of surgery has not yet been widely accepted [1]. the real simulation of smoke is an important research which has not yet been soundly solved. by far, there are few researches referring to smoke simulation in the virtual surgical simulation. considering that smoke is a kind of fluid, a fluid dynamics method is adopted to solve the problems. usually, navier-stokes equations are used to represent the accurate mathematic equations of fluid dynamics [2], but this method is only applicable under simple conditions. with the * the research work described in this paper was supported by the national natural science foundation of china (grant no. 61070079) and the fundamental research funds for the central universities of china (grant no. 6081006). ∗ ∗ advances in systems science and applications (2011), vol.11, no.1-2 173-179 development of computer technology, it has been possible to compute the equations by numerical calculation and as a result, a new subject, computational fluid dynamics, has been born. in 1997, foster and metaxas [3, 4] adopted incompressible navier-stokes equations to a coarse mesh and solved nonlinearity partial differential equations by explicit differential equations, but it was too slow to show real-time animations. in 1999, stam proposed a fluid solution with unconditionally stable and simulated the flow of the fluid with a random time step [5]. because the simulation process of virtual surgery requires great visual effect and excellent quality of real-time, on the premise of great visual effect, the solution to the smoke model built by navier-stokes equations should meet the need of high real-time. therefore, this paper implements a rapid solution to the three-dimensional smoke simulation model by cuda+gpu technology. cuda is a development environment for gpu computation. it is a new software and hardware structure based on c programming language. the code which runs in the display core can be wrote directly by c language and it is unnecessary to learn the special commands or structures of the display core [6]. it can perform large-scale parallel computation by executing multithreading concurrently, which offers a significant opportunity to meet the need of real-time simulation. therefore, this paper accomplishes a rapid solution to navier-stokes equations by cuda technology so as to meet the challenge of real-time. compared with cpu, the result of this paper represents great advantages. to meet the need of real-time quality and real visual effect of the smoke simulation during the surgical training process, this paper adopts a real-time smoke synthesis and simulation model based on computational fluid dynamics. navier-stokes equations are adopted to model smoke as a constant, incompressible fluid with stable temperature. in order to simulate the smoke with a real-time effect, the gauss-seidel iterative method is employed to obtain diffusive items of the density field governing equations. successive over relaxation method is used to derive the amended velocity field which meets the mass conservation condition. and then, ray tracing algorithm is used to implement smoke rendering [7]. for the purpose of getting real-time smoke, the solutions to navier-stokes equations based on gpu are performed by cuda. 2. modeling of smoking simulation 2.1 smoking model smoke state is determined by its density, temperature and velocity. supposing that smoke is incompressible fluids with a stable temperature, then the smoking model can be expressed by incompressible navier-stokes governing equations: continuity equation and momentum equation [8]. 0=⋅∇ v (1) 174 yang: real-time 3d smoking simulation based on computational fluid dynamics…… fvvvv t v +∇+∇⋅−= ∂ ∂ 2)( (2) skv t +∇+∇⋅−= ∂ ∂ ρρρ 2)( (3) equation (1) is the conservative differential form of incompressible smoke navier-stokes continuity equation. where v⋅∇ is called the divergence of velocity (divergence). in the 3d space, divergence is defined as ),,()( z w y v x uvdivv ∂ ∂ ∂ ∂ ∂ ∂ ==⋅∇ . equation (2) is a simplified incompressible navier-stokes smoking velocity momentum equation, while ignoring the pressure gradient item. v is velocity vector, ν represents the kinematic viscosity coefficient and f corresponds to the vector of external body forces. equation (3) is the smoking density momentum equation, in which ρ is the density of the smoke, k is the diffusion constant and s is the scalar field of the amount of substance injected. 2.2 solving the smoking model in 3d space fluid model, we consider smoke fluid to be stored in a discrete grid of a cube. the density field is divided into a n×n×n grid. each cell in it has a number which stands for the density of the cell. if the number is 0, it means that there is no fluid particle in this cell. the larger the number is, the greater the density is. the density field is shown in fig. 1, in which a) shows the three-dimensional fluid mesh and b) is the density distribution of each layer. a) 3d fluid grid b) distribution of density field in each layer of fluid gird fig. 1 density field in each layer of the 3d fluid gird, the outermost ring is the boundary of grid. there are five state types for the boundary: no slip, free slip, inflow, outflow and periodic. in this paper, we employ free slip for the smoking simulation, which means that the density values on boundary are set to 0. the structure of velocity field is similar to the density field. each cell in the grid is a velocity vector, which includes 3 velocity component values at position (x, y, z) of a cell. fdm advances in systems science and applications (2011), vol.11, no.1-2 175 method is used to solve the governing equations. for convenience, density field is solved first, and then the velocity field. solution by parts is used to resolve governing equation of the density field. the detailed steps are as follows: when solving the equation, two of the three items of the right side of equation are set to zero. therefore, the solving process can be divided into 3 steps: solving the adding source item, solving the diffusion item by g-s relaxation method; solving the advection item using tracing backwards method to get the solution. at this point, the calculation of the density field is completed. governing equation of smoke simulation velocity field is exactly the same format with the density field equation. therefore we only need to modify the density field equation and apply the solution by parts to solve the velocity field. for the time being, we complete the calculation of velocity field and velocity field to be amended. the sequences of solving the smoking model governing equation are solving velocity field equations, solving to be amended velocity field which meets mass conservation conditions, solving density field governing equation, repeating till the end of calculation. 3. 3d smoking simulation based on gpu computing considering that smoke is a constant and incompressible fluid with stable temperature, the smoking model can be expressed by incompressible navier-stokes governing equations. the algorithms for the smoking model are highly data-parallel and therefore a good match for the cuda based on gpu programming model to improve the reality of the simulation. 3.1 cuda+gpu model in cuda+gpu architecture, the main function is executed in cpu host and the parallel running data is transferred to the gpu device to implement the parallel computation. all of the tasks executed by cuda are finished by a series of threads in the device allocated by the user. an array of threads is specified as a block. all the blocks are distributed in a parallel running gird. the programming frame based on cuda+gpu is shown in fig. 2 fig. 2 cuda+gpu programming architecture 3.2 3d smoking simulation model based on cuda+gpu 176 yang: real-time 3d smoking simulation based on computational fluid dynamics…… in 3d fluid model, the smoke is considered to be saved in a discrete mesh of a square. at first, the cuda api functions transfer the data of velocity field and density field into the dram of the device and then the corresponding computation of resolving the velocity field and density field is executed in the device. in the end, the result obtained in the device is transferred to the host so as to implement the rendering. the whole frame is shown in fig. 3. fig. 3 simulation model of 3d smoking based on cuda+gpu considering that the bandwidth between the device and device memory is much higher than that between device memory and host memory, it is highly necessary to minimize the data transmission between the host and the device. in our test, the velocity field and the density field are initialized in the host firstly, and then they are sent to the device. at time t, the simulation model is resolved in device and as a result the velocity field and the density field of time t+1 can be got. for the reason that the final rendering is implemented in the host, the computation result in the device should be written back into the host. considering that only the data of density filed is used when rendering, it just need to write the density filed to the host. meanwhile, the velocity field and the density filed at time t+1 exchange with the data of time t to prepare for the next iteration. in the detailed running process, the 3d mesh is corresponding to the parallel running grid in the device, the slice of each layer in the fluid mesh is corresponding to the thread block in the device and the gird in the slice to the thread in the device. the final solving of the smoking model is implemented by the allocated thread. 4. smoking simulation results the rendering of the smoke is implemented by the ray-tracing algorithm based on hardware [7]. g-s iterative method is applied to solve the diffusive item of smoking density governing equation. sor method is utilized to solve possion equation and the velocity field to be amended which meets the mass conservation condition [8]. the ray tracing method is used to render the advances in systems science and applications (2011), vol.11, no.1-2 177 smoke. cuda+gpu technology is employed to solve the navier-stokes equations based on gpu to obtain a visual effect of real-time smoking. in the following, the smoking simulation based on cuda+gpu is implemented, and the results are displayed in opengl. the hardware configuration of our experimental platform is as follows: dell precision t3500, intel(r) xeon(r) w3503 2.40ghz cpu, 4gb memory, nvidia tesla c1060. the software configuration is: windows vista, visual c++.net 2005. the size of the fluid mesh in the investigation is 323232 ×× . we forced our codes to execute 1000 times in order to get an objective running time of the smoking model. a frame of dynamic smoking simulation results is shown in fig. 4. besides , the efficiency of the whole smoking simulation for virtual surgery (including the spatial positioning of the surgical instruments) is tested and the real-time smoking simulation result for virtual surgery is shown in fig. 5. fig. 4 a frame of dynamical smoking simulation fig. 5 real-time smoking s imulation for virtual surgery the performance of the smoking simulation system implemented on the cpu and cuda+gpu are compared in table 1. we get 36.0 speedup for the mesh. the result listed in table 1 shows that significant increase of the performance is achieved and the frame per second of the whole system reaches to 56.2, which meets the requirement of real-time simulation for virtual surgery. 178 yang: real-time 3d smoking simulation based on computational fluid dynamics…… table 1 performance comparison on the cpu and the cuda+gpu data processing location total running time (ms) average running time (ms) speedup ratio frame per second (fps) cpu 193384 193.384 4.7 cuda+gpu 5370 5.370 36.0 56.2 5. conclusion this paper investigates the methods for building the real-time dynamical smoking model based on cfd and employs the incompressible navier-stokes equations as the governing model to the smoke to be constant and incompressible. in order to develop a realistic simulation of smoking, we apply g-s iterative method to solve the diffusive item of smoking density governing equation. sor method is utilized to solve possion equation and the velocity field to be amended which meets the mass conservation condition. the ray tracing method is used to render the smoke and cuda technology is employed to solve the navier-stokes equations based on gpu to obtain a real-time smoking visual effect. the results show that the smoking simulation model proposed is practical and elucidating. references [1] montgomery k., bruyns c., wildermuth s. et al. surgical simulator for operative hysteroscopy. ieee visualization 2001, washington dc, pp.14-17, 2001. [2] jiyuan tu, guan heng yeoh, chaoqun liu. computational fluid dynamics: a practical approach. 1st ed., oxford: butterworth-heinemann, 2007. [3] foster n., metaxas d. controlling fluid animation. proceedings of the cgi_97, belgium: ieee computer society, pp.178-188, 1997. [4] foster n., metaxas d. modeling the motion of a hot, turbulent gas. proceedings of the 24th annual conference on computer graphics and interactive techniques. new york, pp. 181-188, 1997. [5] stam j. stable fluids. siggraph 99 conference proceedings, annual conference series, los angeles, pp. 121-128, august 1999. [6] nvidia: cuda programming guide ver. 2.0. [7] ronald fedkiw, jos stam, henrikwann jensen. visual simulation of smoke. in siggraph 2001 conference proceedings, annual conference series, pp.15-22, august 2001. [8] zhiyong yuan, qian yin, shikun feng, jun hu. real-time dynamic bleeding synthesis and simulation of endoscopic images. geomatics and information science of wuhan university, 34(3): 317-320, 2009. (in chinese) advances in systems science and applications (2011), vol.11, no.1-2 179 advances in systems science and application (2016) vol.16 no.1 1-18 cybernetics 2.0 dmitry a. novikov institute of control sciences, moscow, russia moscow institute of physics and technology, moscow, russia abstract the evolution of cybernetics (from n. wiener to the present day) is briefly considered. a new development stage of cybernetics (the so-called cybernetics 2.0) is discussed as a science on general regularities of systems organization and control. the author substantiates the topicality of elaborating a new branch of cybernetics, i.e., organization theory (o3) which studies an organization as a property, process and system. keywords cybernetics; organization; system; control. 1 cybernetics of n. wiener this section is intended to consider in brief the history of cybernetics and describe “classical” cybernetics. let us call it “cybernetics 1.0”. cybernetics is the science of general regularities of control and information transmission processes in different systems, whether machines, animals or society. cybernetics studies the concepts of control and communication in living organisms, machines and organizations including self-organization. it focuses on how a (digital, mechanical or biological) system processes information, responds to it and changes or being changed for better functioning (including control and communication). cybernetics is an interdisciplinary science. it originated “at the junction” of mathematics, logic, semiotics, physiology, biology and sociology. among its inherent features, we mention analysis and revelation of general principles and approaches in scientific cognition. control theory, communication theory, operations research and others represent most weighty theories within cybernetics 1.0. in ancient greece, the term “cybernetics” denoted the art of a municipal governor (e.g., in platos laws). a. ampere (1834) related cybernetics to political sciences: the book [2] defined cybernetics (“the science of civil government”) as a science of current policy and practical governance in a state or society. b. trentowsky (1843, see [3]) viewed cybernetics as “the art of how to govern a nation.” in its tektology (1925, see [4]), a. bogdanov examined common organizational principles for all types of systems. in fact, he anticipated many results of n. wiener and l. bertalanffy, as the both were not familiar with bogdanovs works. 2 dmitry a. novikov: cybernetics 2.0 the modern (and classical!) interpretation of the term “cybernetics” as “the scientific study of control and communication in the animal and the machine” was pioneered by norbert wiener in 1948, see the monograph [5]. two years later, wiener also added society as the third object of cybernetics[6]. among other classics, we mention william ashby [7, 8] (1956) and stafford beer [9] (1959), who made their emphasis on the biological and “economic” aspects of cybernetics, respectively. therefore, cybernetics 1.0 (or simply cybernetics) can be defined as “the science ofcontrol and data processing in animals, machines and society.” an alternative is the definition of cybernetics (with capital c, to distinguish it from cybernetics whenever confuse may occur) as “the science of general regularities of control and data processing in animals, machines and society.” the second definition differs from its first counterpart in the words “general regularities,” which is crucial and will be repeatedly underlined and used below. in the former case, the matter concerns “the umbrella brand,” i.e., the “integrated” results of all sciences dealing with problems of control and data processing in animals, machines and society. the latter case covers partial “intersection” of these results (see fig.1 figuratively speaking, the central rode of the “umbrella.”), i.e., usage of common results for all component sciences. furthermore, we will adhere to this approach over and over again for discrimination between the corresponding umbrella brand and the common results of all component sciences in the context of different categories such as interdisciplinarity, systems analysis, organization theory, etc. cybernetics today (disciplines included in cybernetics in the descending order of their “grades” of membership, see fig.1, with year of birth if available): –control theory (1868-the papers [10, 11] published by j. maxwell and i. vyshnegradsky); –mathematical theory of communication and information (1948-k. shannons works[12, 13]); –general systems theory, systems engineering and systems analysis (1968-the book [14] and 1956-the book [15]); –optimization (including linear and nonlinear programming; dynamic programming; optimal control; fuzzy optimization; discrete optimization, genetic algorithms, and so on); –operations research (graph theory, game theory and statistical decisions, etc.); –artificial intelligence (1956-the dartmouth summer research project on artificial intelligence); –data analysis and decision-making ; –robotics and others (purely mathematical and applied sciences and scientifadvances in systems science and application (2016) vol.16 no.1 3 ic directions, in an arbitrary order) including systems engineering, recognition, artificial neural networks and neural computers, ergatic systems, fuzzy systems (rough sets, grey systems, etc), mathematical logic, identification theory, algorithm theory, scheduling theory and queuing theory, mathematical linguistics, programming theory, synergetics and all that jazz. fig. 1 the composition and structure of cybernetics in its components, cybernetics intersects considerably with many other sciences, in the first place, with such metasciences as general systems theory and systems analysis and informatics (see below and [16]). there exist a few classical monographs and textbooks on cybernetics with its “own” results; here we refer to [1, 6, 7, 9, 14, 17–23]. on the other hand, textbooks on cybernetics (mostly published in the former ussr) include many of the above-mentioned directions (par excellence, control in technical systems and informatics)-see [24–28]. the prefix “cyber” induces new terms on a regular basis, viz., cybersystem, cyberspace, cyberthreat, cybersecurity, etc. in a broader view of things, this prefix embraces all connected with automation, computers, virtual reality, internet and so on. alongside with general cybernetics, there exist special (“sectoral”) types of cybernetics[27]. a most natural approach (which follows from wieners extended 4 dmitry a. novikov: cybernetics 2.0 definition) is to separate out technical cybernetics, biological cybernetics and socioeconomic cybernetics besides theoretical cybernetics (i.e., cybernetics). it is possible to compile a more complete list of special types of cybernetics (see references in [16]): physical cybernetics (to be more precise, “cybernetical physics”, see [29, 30]), social cybernetics, educational cybernetics, quantum cybernetics (quantum systems control, quantum computing), etc. as standing apart, we mention a branch of biological cybernetics known as cybernetic brain modeling integrated with artificial intelligence, neural and cognitive sciences. a romantic idea to create a cybernetic (computer-aided) brain at least partially resembling a natural brain stimulated the founding fathers of cybernetics (see the works of w. ashby[8], g. walter[31], m. arbib[32], f. george[33], k. steinbuch[34] and others) and their followers (for a modern overview, we refer to [35]). 2 cybernetics of cybernetics and other types of cybernetics in addition to wieners classical cybernetics, the last 50+ years yielded other types of cybernetics declaring their connection with the former and endeavoring to develop it further. no doubt, the most striking phenomenon was the appearance of second-order cybernetics (cybernetics of cybernetics, metacybernetics, new cybernetics; here “order” corresponds to “reflexion rank”). cybernetics of cybernetic systems is associated with the names of m. mead, g. bateson and h. foerster and puts its emphasis on the role of subject/observer performing control[36–40]. the central concept of second-order cybernetics is an observer as a subject refining the subject from the object (indeed, any system is a “model” generated from reality for a certain cognitive purpose and from some point of view/abstraction). h. foerster noted that “a brain is required to write a theory of a brain. from this follows that a theory of the brain, that has any aspirations for completeness, has to account for the writing of this theory. and even more fascinating, the writer of this theory has to account for her or himself. translated into the domain of cybernetics; the cybernetician, by entering his own domain, has to account for his or her own activity. cybernetics then becomes cybernetics of cybernetics, or second-order cybernetics.”[38]. in contrast to wieners cybernetics, second-order cybernetics possesses the conceptual-philosophical character (for a mathematician or engineer, it is demonstrative that all publications on second-order cybernetics contain no formal models, algorithms, etc.). in fact, this type of cybernetics “transmits” the complementarity principle (with insufficient grounds) from physics to all other sciences, phenomena and processes. moreover, a series of works postulated that any system must have positive feedback loops amplifying positive control actions (e.g., advances in systems science and application (2016) vol.16 no.1 5 see [41]). but any expert in control theory knows the potential danger of such loops for system stability! the “biological” stage in second-order cybernetics is associated with the names of h. maturana and f. varela[42–44] and their notion of autopoiesis (self-generation and self-development of systems). f. varela underlined that “first-order cybernetics is the cybernetics of observed systems; second-order cybernetics is the cybernetics of observing systems.” the latter focuses on feedback of a controlled system and an observer. therefore, the key terms of second-order cybernetics are recursiveness, selfregulation, reflexion, autopoeisis. for a good survey of this direction, we refer to [45, 46]. however, the historical picture has appeared much more colorful and diverse, not confining to the second order. some authors adopt the terms “third-order cybernetics” (social autopoeisis; second-order cybernetics considering autoreflexion) and “fourth-order cybernetics” (third-order cybernetics considering observers system of values), but they are conceptual and still have no generally accepted meanings (e.g., see a discussion in [47–53]). for instance, v. lepsky wrote: “third-order cybernetics can be formed basing on the thesis “from observing systems to self-developing systems.” in this case, control is gradually transformed into a wide spectrum of support processes for system self-development, namely, social control, stimulation, maintenance, modeling, organization, “assembly/disassembly” of subjects and others.”[54]. we point out other directions (see table 1): homeostatics (yu. gorsky and his scientific school), a science studying contradictions control for the sake of maintaining the permanency of processes, functions, development trajectories, etc.[55]; neo-cybernetics (b. sokolov and r. yusupov), an interdisciplinary science which elaborates a methodology of stating and solving analysis and synthesis problems of intelligent control processes and systems for complex arbitrary-nature objects[56, 57]; neo-cybernetics (s. krylov)[58]; control methodology (d. novikov)[59]; new cybernetics, post-cybernetics (g. tesler), a fundamental science about general laws and models of informational interaction and influence in processes and phenomena running in animate, inanimate and artificial nature[60]. interestingly, k. kolin had proposed almost a same definition to informatics 20 years before g. tesler, see [61]; evergetics (v. vittikh), a value-oriented science about control processes in a society, which focuses on problem situations for a group of heterogeneous actors 6 dmitry a. novikov: cybernetics 2.0 with different viewpoints, interests and value preferences[62]. in other words, evergetics can be defined as third-order cybernetics for interacting control subjects. according to vittikhs fair remark, in everyday social life control processes will be realized by the “tandem” of common and professional control experts (theoreticians): the former face concrete problem situations in daily routine and acquire conventional knowledge (in the sense of h. poincare) on the situation and define directions of its control, whereas the latter create necessary methods and means for their activity. involvement of “common” people into social control processes is an important development trend of control science. subject-oriented control in noosphere, the so-called hi-hume cybernetics (v. kharitonov and a. alekseev), a science mostly considering subjectness and subjectivity of control[63]. table 1 different types of cybernetics type main authors period cybernetics n. wiener, w. ashby, s. beer the 1948-1950s second-order cybernetics m. mead, g. bateson, h.foerster the 1960-1970s autopoiesis h. maturana, f. varela the 1970s homeostatics yu. gorsky the 1980s conceptual cybernetics of third and fourth orders v. kenny, r. mancilla, s.umpleby the 1990-2010s neo-cybernetics b. sokolov, r. yusupov the 2000s neo-cybernetics s. krylov the 2000s third-order cybernetics v. lepsky the 2000s new cybernetics, post-cybernetics g. tesler the 2000s control methodology d. novikov the 2000s evergetics v. vittikh the 2010s subject-oriented control in noosphere (hi-hume cybernetics) v. kharitonov, a. alekseev the 2010s it is possible to introduce the notion of “fifth-order cybernetics”[16] as fourthorder cybernetics considering the mutual reflexion of control subjects[46] making coordinated decisions, etc. note that all types of cybernetics in table 1 are conceptual, i.e., absorbed by cybernetics. the observed variety of the approaches claiming (explicitly or implicitly) to be a new mainstream in classical cybernetics development seems natural, as reflecting the evolution of cybernetics. with the lapse of time, certain approaches will be further developed, others will stop growing. of course, it is extremely desirable to obtain a general picture with integration, generalization and joint positioning of all existing approaches or most of them. advances in systems science and application (2016) vol.16 no.1 7 3 cybernetics 2.0 the history of cybernetics and its state-of-the-art, as well as the development trends and prospects of several components of cybernetics (mainly, control theory see also [64]) is briefly considered in [16]. what are the prospects of cybernetics? to answer this question, let us address the primary source-the initial definition of cybernetics as the science of control and communication. its interrelation with control seems more or less clear. at the first glance, this is also the case for communication: by the joint effort of scientists (including n. wiener), the mathematical theory of communication and information appeared in the 1940s (quantitative models of information and communication channels capacity, coding theory, etc.). but take a broader view of communication. both in the paper [65] and in the original book [5], n. wiener explicitly or implicitly mentioned interrelation or intercommunication or interaction-reasonability and causality (cause-effect relations). really, in feedback control systems, control-effect is defined by its cause, i.e., the state of a controlled system (plant); conversely, control supplied to the input of a plant is induced by its cause, i.e., the state of a controller, and so on. no doubt, the channels and methods of communication are important but secondary whenever the matter concerns universal regularities for animals, machines and society. a much broader view of communication implies interpreting communication as intercommunication, e.g., between elements of a plant, between a controller and a plant, etc. including different types of impacts and interactions (material, informational and other ones). “intercommunication” is a more general category than “communication.” in the general systems context, intercommunication corresponds to the category of organization (see its definition and discussion below). therefore, a simple correction (replacing “communication” with “organization” in wieners definition of cybernetics) yields a more general and modern definition of cybernetics: “the science of systems organization and their control.” we call it cybernetics 2.0. making such substitution, we get distanced from informatics. consider the soundness and consequences of this distancing. cybernetics and informatics.nowadays, cybernetics and informatics form independent interdisciplinary fundamental sciences[61]. according to a figurative expression of b. sokolov and r. yusupov[56], informatics and cybernetics are “siamese twins.” yet, in nature siamese twins represent pathology for instance, the definition of informatics as the “union” of general laws of informatics and control would induce a megascience without concrete content, subsisting at conceptual level exclusively. 8 dmitry a. novikov: cybernetics 2.0 cybernetics and informatics have a strong intersection (including the level of common scientific base-statistical information theory). their accents much differ. the fundamental ideas of cybernetics are wieners “control and communication in the animal and the machine,” whereas the fundamental ideas of informatics are formalization (theory) and computerization (practice). accordingly, in the mathematical sense cybernetics bases on control theory and information theory, whereas informatics proceeds from theory of algorithms and formal systems. note, that this distinction partly elucidates why some sciences often related to informatics or computer sciences have not been mentioned: theory of formal languages and grammars, “true” artificial intelligence (knowledge engineering, reasoning formalization, behavior planning, etc. instead of artificial neural networks as a modern empirical engineering science), automata theory, computational complexity theory, and so on. the subject of modern informatics (or even the “umbrella brands” of informational sciences) covering information science, computer science and computational science[66] are informational processes. indeed, on the one hand, information processing arises everywhere (!), not only in control and/or organizing. on the other hand, informational processes and corresponding information and communication technology are integrated into control processes so that their discrimination seems almost impossible. a close cooperation of informatics and cybernetics at partial operational level will be continued and even extended in future. organization and organization theory.according to the definition provided by merriam-webster dictionary, an organization is: 1. the condition or manner of being organized; 2. the act or process of organizing or of being organized; 3. an administrative and functional structure (as a business or a political party); also, the personnel of such a structure. well use the notion “organization” mostly in its second and first meanings, i.e., as a process and a result of this process. the third meaning (an organizational system) as a class of controlled objects appears in theory of control in organizational systems[67, 68]. at descriptive (phenomenological) and explanatory levels[69], “system organization” reflects how and why exactly so, respectively, a system is organized (organization as a property). at normative level, “system organization” reflects how it must be organized (requirements to the property of organization) and how it should be organized (requirements to the process of organization). note that nowadays also exists “theory of organizations” (“organizational theory”) a branch of management science, both in its subject (organizational systems) and methods used. unfortunately, numerous textbooks (and just a few advances in systems science and application (2016) vol.16 no.1 9 monographs!) give only descriptive generalizations on the property and process of organization in their introductions, with most attention then switched to organizational systems, viz., management of organizations (for instance, see the classical textbook [70]). a scientific branch responsible for the posed questions (organization theory, or o3 (organization as a property, process and system, by analogy to c3 control, computation, communication[16, 64]) has almost not been developed to-date. yet, this branch obviously has a close connection and partial intersection with general systems theory and systems analysis (mostly focused on descriptive level problems and a little bit dealing with normative level ones), as well as with methodology (as the general science of activity organization[59, 69]). creating a full-fledged organization theory is a topical problem of cybernetics! consider the correlation of the two basic categories in the definition of cybernetics 2.0 (“organization” and “control”). control is “an element, function of different organized systems (biological, social, technical ones) preserving their definite structure, maintaining activity mode, implementing a program, a goal of activity.” control is “an impact on a controlled system, intended for ensuring its necessary behavior”[68]. consequently, the categories of organization and control do intersect, but do not coincide. the former fits system design and the latter fits system functioning (a conditional analogy: organization corresponds to deism (the creator of a system does not interfere in its functioning), while control corresponds to teism (the opposite picture)); they are jointly realized during system implementation and adaptation, see fig.2. in other words, organization (strategic loop) “foregoes” control (tactical loop). fig. 2 organization and control the domains in fig.2 have the following content (as examples): i. design (construction) of systems (including their stuff, structure and functions)organization but not control (despite that theory of control in organizational systems suggests stuff control and structure control). 10 dmitry a. novikov: cybernetics 2.0 ii. joint design of a system and a controlled object. adaptation. control mechanisms adjustment. iii. functioning of controllers in technical systems-control but not organization. on the one part, control process calls for organization (organization as a stage in fayols management cycle and a function of organizational control, see [67]). on the other part, organization process (e.g., system life cycle) might and should be controlled. organization and control can have a “hierarchical” correlation. generally speaking, the correlation of organization and control is far from trivial and requires further perception. for instance, in multi-agent systems decentralized control (choosing the laws and rules of autonomous agents interaction) can be treated as organization. another example is the bible as a tool of organization[71] (a system of norms making common knowledge and implementing institutional control of a society). following the complication of systems created by mankind, the process and property of organization will attract more and more attention. indeed, control of standard objects (e.g., controller design for technical and/or production systems) gradually becomes a handicraft rather than a science; modern challenges highlight standardization of activity organization technologies, creation of new activity technologies, etc. (activity systems engineering). a fruitful combination of organization and control within cybernetics 2.0 would give a substantiated and efficient answer to the primary question of activity systems engineering: how should control systems for them be constructed? actually, this is a “reflexive” question related to second-order and even higher-order cybernetics. mankind has to learn to design and implement control systems for complex systems (high-technology manufacturing, product life cycle, organizations, regions, etc.), similarly to the existing achievements in technical systems engineering. cybernetics is important from general educational viewpoint, since it forms the integral modern scientific world outlook. cybernetics 2.0.we have defined cybernetics 2.0 as the science of (general regularities in) systems organization and their control. a close connection between cybernetics and general systems theory and systems analysis[16], as well as the growing role of technologies leads to a worthy hypothesis. cybernetics 2.0 includes cybernetics (wieners cybernetics and higher-order cybernetics), cybernetics, and general systems theory and systems analysis with results in the following forms: general laws, regularities and principles studied within metasciences-cybernetics and systems analysis; advances in systems science and application (2016) vol.16 no.1 11 a set of results obtained by sciences-components (“umbrella brands”-cybernetics and systems studies uniting appropriate sciences); design principles of corresponding technologies. keywords for cybernetics 2.0 are control, organization and system. similarly to cybernetics in its common sense, cybernetics 2.0 has a conceptual core (cybernetics 2.0 with capital c). at conceptual level, cybernetics 2.0 is composed of control philosophy (including general laws, regularities and principles of control), control methodology, organization theory (including general laws, regularities and principles of (a) complex systems functioning and (b) development and choice of general technologies), as illustrated by fig.3. basic sciences for cybernetics 2.0 are control theory, general systems theory and systems analysis, as well as systems engineering-see fig.3. complementary sciences for cybernetics 2.0 are informatics, optimization, operations research and artificial intelligence-see fig.3. fig. 3 the composition and structure of cybernetics 2.0 12 dmitry a. novikov: cybernetics 2.0 the general architecture of cybernetics 2.0 (see fig.3) admits projection to different application domains and branches of subject-oriented sciences depending on a class of posed problems (technical, biological, social, etc.). 4 the prospects of cybernetics 2.0 further development of cybernetics has several alternative scenarios as follows: –the negativistic scenario (the prevailing opinion is that “cybernetics does not exist” and it gradually falls into oblivion); –the “umbrella” scenario (owing to past endeavors, cybernetics is considered as a “mechanistic” (non-emergent) union, and its further development is forecasted using the aggregate of trends displayed by the basic and complementary sciences under the “umbrella brand” of cybernetics); –the “philosophical” scenario (the framework of new results in cybernetics 2.0 includes conceptual considerations only-the development of conceptual level); –the subject-oriented (sectoral) scenario (the basic results of cybernetics are obtained at the junction of sectoral applications); –the constructive-optimistic (desired) scenario (the balanced development of the basic, complementary and “conceptual” sciences is the case, accompanied by the convergence and interdisciplinary translation of their common results, with subsequent generation of conceptual level generalizations (realization of wieners dream “to understand the region as a whole”). the development of cybernetics 2.0 in the conditions of intensified sciences differentiation provides the following: –for scientists specialized in cybernetics proper and the representatives of adjacent sciences: the general picture of a wide subject domain (and a common language of its description), the positioning of their results and promotion in new theoretical and applied fields; –for potential users of applied results (authorities, business structures): (1) confidence in the uniform positions of researchers; (2) more efficient solution of control problems for different objects based on new fundamental results and associated applied results. main challenges are control in social and living systems. several classes of control problems seem topical, namely: –network-centric systems (including military applications, networked and cloud production); –informational control and cybersafety; –life cycle control of complex organization-technical systems; –activity systems engineering. among promising application domains, we mention living systems, social systems, microsystems, energetics and transport. advances in systems science and application (2016) vol.16 no.1 13 there exists a series of global challenges to cybernetics 2.0 (i.e., observed phenomena going beyond cybernetics 1.0), see [16]: 1)the scientific tower of babel (interdisciplinarity, differentiation of sciences; in the first place, in the context of cybernetics-sciences of control and adjacent sciences); 2)centralization collapse (decentralization and networkism, including systems of systems, distributed optimization, emergent intelligence, multi-agent systems, and so on); 3) strategic behavior (in all manifestations, including interests inconsistency, goal-setting, reflexion and so on); 4) complexity damnation (including all aspects of complexity and nonlinearity (figuratively, in this sense cybernetics 2.0 has to include nonlinear automatic control theory studying nonlinear decentralized objects with nonlinear observers, etc.) of modern systems, as well as dimensionality damnation-big data and big control[72]). thus, the main tasks of cybernetics 2.0 are developing the basic and complementary sciences, responding to the stated global challenges, as well as advancing in appropriate application domains. and here are the main tasks of cybernetics 2.0: 1) ensuring the interdisciplinarity of investigations (with respect to the basic and complementary sciences, as illustrated by fig.3); 2) revealing, systematizing and analyzing the general laws, regularities and principles of control for different-nature systems within control philosophy; this would require new and new generalizations; 3) elaborating and refining organization theory (o3). we have described the phylogenesis of a new stage of cybernetics-cybernetics 2.0. further development of cybernetics would call for considerable joint effort of mathematicians, philosophers, experts in control theory, systems engineering and many others involved. references [1] ackoff r. and emery f. (2005), on purposeful systems: an interdisciplinary analysis of individual and social behavior as a system of purposeful events. 2nd ed., new york: aldine transaction, pp.303. [2] ampère a.-m. (1843),essai sur la philosophie des sciences, paris: chez bachelier, pp.140-142. [3] trentowski b. (1843), stosunek filozofii do cybernetyki, czyli sztuki rza̧dzenia narodem, warsawa, pp.195. 14 dmitry a. novikov: cybernetics 2.0 [4] bogdanov a. (1926), algemeine organisationslehre (tektologie),berlin: hirzel, no.i. [5] wiener, n. (1961), cybernetics, or, control and communication in the animal and the machine; m.i.t. press, pp.194. [6] wiener n. (1950), the human use of human beings. cybernetics and society, boston: houghton mifflin company, pp.200. [7] ashby w. (1956), an introduction to cybernetics, london: chapman and hall, pp.295. [8] ashby w. (1952), design for a brain: the origin of adaptive behavior., new york: john wiley & sons, pp.298. [9] beer s. (1959), cybernetics and management, london: the english university press, pp.214. [10] maxwell j.c. (1968), “on governors”, proceedings of the royal society of london, vol.16, pp. 270-283. [11] vyshnegradsky i. (1877), “on direct-action controllers”, izvestiya st. petersburg practical technological institute, vol.1, pp.21-62. [12] shannon c(1974), “a mathematical theory of communication”, bell system technical journal, vol.27, pp.379-423,623-656. [13] shannon c. and weaver w. (1948), the mathematical theory of communication, illinois: university of illinois press, pp.144. [14] bertalanffy l. (1968), general system theory: foundations, development, applications, new york: george braziller, pp.296. [15] kahn h. and mann i. (1956), techniques of systems analysis, santa monica: rand corporation, pp.168. [16] novikov d. (2016), cybernetics: from past to future, berlin, springer, pp.107. [17] beer s. (1972), brain of the firm: a development in management cybernetics, london: herder and herder, pp.319. [18] george f. (1977), the foundations of cybernetics, london: gordon and breach science publisher, pp.286. advances in systems science and application (2016) vol.16 no.1 15 [19] george f.h. (1979), philosophical foundations of cybernetics, kent: abacus press, pp.157. [20] mesarović m., mako d. and takahara y. (1970), theory of hierarchical multilevel systems, new york: academic, pp.294. [21] wiener n. (1953), ex-prodigy: my childhood and youth, mathematics & statistics, pp.317. [22] wiener n. (1964). “god & golem, inc.: a comment on certain points where cybernetics impinges on religion”, philosophy & phenomenological research, vol.28, no.1, pp.99. [23] wiener n. (1964). i am mathematician, cambridge: the mit press, pp.380. [24] druzhinin v and kontorov d.s. (1989), introduction to conflict theory, moscow: radio i svyaz, pp.288. [25] glushkov v. (1964), introduction to cybernetics, kiev: ukr. ssr academy of sciences, pp.324. [26] korshunov yu. (1987), mathematical foundations of cybernetics, moscow: energoatomizdat, pp.496. [27] kuzin l. (1979), the foundations of cybernetics, moscow: energiya, vol. 1. pp.504 and vol. 2. pp.584. [28] lerner a. (1972), fundamentals of cybernetics, berlin: springer, pp.294. [29] fradkov a. (2006), cybernetical physics: from control of chaos to quantum control (understanding complex systems),berlin: springer, pp.236. [30] turchin v. (1977), the phenomenon of science, new york: columbia university press,1977. pp.348. [31] walter w. (1963). “the living brain”, journal of nervous & mental disease, vol.120, no.1, pp.255. [32] arbib m. (1972), the metaphorical brain: an introduction to cybernetics as artificial intelligence and brain theory, new york: wiley, pp.384. [33] george f. (1962), the brain as a computer., new york: pergamon press, pp.437. [34] steinbuch k. (1963), automat und mensch. kybernetische tatsachen und hypothesen, berlin: springer-verlag, pp.392. 16 dmitry a. novikov: cybernetics 2.0 [35] pickering a. (2010), the cybernetic brain, chicago: the university of chicago press, pp.537. [36] bateson g. (1972), steps to an ecology of mind, san francisco: chandler pub. co, pp.542. [37] foerster h. (1995), the cybernetics of cybernetics, minneapolis: future systems, pp.228. [38] foerster h. (2003), understanding understanding: essays on cybernetics and cognition, new york: springer-verlag, pp.362. [39] heylighen f. and joslyn c. (2001), cybernetics and second-order cybernetics / encyclopedia of physical science & technology (3thed.), new york: academic press, pp.155-170. [40] mead m. (1968), the cybernetics of cybernetics / purposive systems. h. von foerster et al.(ed.), new york: spartan books, pp.1-11. [41] maruyama m. (1963), “the second cybernetics: deviation-amplifying mutual causal processes”, american scientist, vol. 5, no. 2, pp.164-179. [42] maturana h. (1980), varela f. autopoiesis and cognition, dordrecht: d. reidel publishing company, pp.143. [43] maturana h. (1987), varela f. the tree of knowledge, boston: shambhala publications, pp.231. [44] varela f. (1975), “a calculus for self-reference”, international journal of general systems, vol.2, pp.5-24. [45] lefevbre v. (2002), “second-order cybernetics in the soviet union and western countries”, reflexive processes and control, vol.2, no. 1, pp.96-103. (in russian). [46] novikov d. (2014), chkhartishvili a. reflexion and control: mathematical models, london: crc press, pp.298. [47] boxer p and kenny v. (1992), “lacan and maturana: constructivist origins for a 30 cybernetics”, communication and cognition, vol.25, no.1, pp.73100. [48] kenny v. (2009), “theres nothing like the real thing. revisiting the need for a third-order cybernetics”, constructivist foundations, vol.4, no.2, pp.100111. advances in systems science and application (2016) vol.16 no.1 17 [49] mancilla r. (2011), “introduction to sociocybernetics (part 1): third-order cybernetics and a basic framework for society”, journal of sociocybernetics, vol.42, no.9, pp.35-56. [50] mancilla r. (2013), “introduction to sociocybernetics (part 3): fourthorder cybernetics”, journal of sociocybernetics, vol.44, no.11, pp.47-73. [51] müller k. (2013), “the new science of cybernetics: a primer”, journal of systemics, cybernetics and informatics, vol.11, no.9, pp.32-46. [52] umpleby s. (2008), “a brief history of cybernetics in the united states”, austrian journal of contemporary history,2008. vol. 19. no. 4. p. 28-40. [53] umpleby s. (1990), “the science of cybernetics and the cybernetics of science”, cybernetics and systems, vol.21, no.1, pp.109-121. [54] lepsky v. (2014), “the philosophy and methodology of control in the context of scientific rationality development”, xii all-russian meeting on control problems. -moscow: trapeznikov institute of control sciences, pp.77857796. [55] gorsky yu. (1988), a system-informational analysis of control processes, novosibirsk: nauka, pp.327. [56] sokolov b. and yusupov r.m. (2014), “analysis of interdisciplinary interaction between modern informatics and cybernetics: theoretical and practical aspects”, xii all-russian meeting on control problems. moscow: trapeznikov institute of control sciences ras, pp.8625-8636.. [57] sokolov b. and yusupov r.m. (2014), “neocybernetics in the modern structure of system knowledge”, robototekhnika i tekhnicheskaya kibernetika,2014. no. 2(3). p. 3-10. (in russian). [58] krylov s. (2008), neocybernetics: algorithms, evolution mathematics and future technologies, moscow: lki, pp.288. [59] novikov d. (2013), control methodology, new york: nova science publishers, pp.76. [60] tesler g. (2004), new cybernetics, kiev: logos, pp.404. [61] kolin k. (2010), philosophical problems of informatics, moscow: binom, pp.270. 18 dmitry a. novikov: cybernetics 2.0 [62] vittikh v.a. (2014), “evolution of ideas on management processes in the society: from cybernetics to evergetics”, group decision and negotiation. available at: http://link.springer.com/article/10.1007/s10726-0149414-6/fulltext.html. [63] kharitonov v. and alekseev a.o. (2015), “the concept of subjectoriented control in social and economic systems”, polythematic electronic journal of kuban state agricultural university, (electronic source), kuban state agricultural university, vol.05, no.109. available at http://ej.kubagro.ru/2015/05/pdf/43.pdf. [64] forrest j. and novikov d. (2012),“modern trends in control theory: networks, hierarchies and interdisciplinarity ”, advances in systems science and application, vol.12, no.3, pp.1-13. [65] rosenblueth a., ewiener n. and bigelow j. (1943), “behavior purpose and teleology”, philosophy of science, no.10, pp.18-24. [66] kolin k k. (2006), “becoming of informatics as fundamental science and the complex scientific problem”, sistemy i sredstva inform, pp.7-58. [67] burkov v, goubko m, kondratev v, korgin n and novikov d. (2013), mechanism design and management: mathematical methods for smart organizations. prof(ed.), new york: nova science publishers, pp.163. [68] novikov d. (2013), theory of control in organizations, new york: nova science publishers, pp.341. [69] novikov a. and novikov d. (2013), research methodology: from philosophy of science to research design.,amsterdam, crc press, pp.130. [70] daft r. (2012), organization theory and design(11thed), new york: cengage learning, pp.688. [71] prangishvili i. (2000), systems approach and system-wide regularities, moscow: sinteg, pp.258. [72] novikov d. (2015), “big data and big control”, advances in systems studies and applications, vol.15, no.1, pp.21-36. corresponding author dmitry a. novikov can be contacted at: novikov@ipu.ru advances in systems science and applications (2012) vol.12 no.4 347-352 a universal nonlinear control law for the synchronization of arbitrary 3-d continuous-time quadratic systems zeraoulia elhadj1 and j.c.sprott2 1department of mathematics, university of tébessa, 12002, algeria. 2department of physics, university of wisconsin, madison, wi 53706, usa abstract in this letter we present a universal nonlinear control law for the synchronization of arbitrary 3-d continuous-time quadratic systems. this control law does not require any type of conditions on the considered systems. keywords synchronization, universal nonlinear control law, chaos. pacs numbers: 05.45.-a, 05.45.gg 1 introduction several methods have been successfully applied to chaos synchronization. for example, in [1] a method is introduced to synchronize two identical chaotic systems with different initial conditions. an adaptive control approach is presented in [2], a backstepping design was presented in [3], an active control method is presented in [4-6], and a nonlinear control scheme was given in [7-9]. consequently, there are many applications of chaos synchronization in physical, chemical, and ecological systems, and in secure communications as shown in [1-2,10-13]. in this letter, we apply nonlinear control theory to synchronize two arbitrary 3-d continuous-time quadratic systems. the proposed control law does not need any conditions on the considered systems, and hence it is a universal synchronization approach for general 3-d continuous-time quadratic systems. in other words, the present letter is concerned with synchronization of nonlinear systems in the framework of nonlinear observers. the investigation is restricted to a pair of quadratic three dimensional systems, for which a control feedback can be chosen in such a way that global asymptotic stability of the error system can be established in the framework of classical lyapunov theory. this restriction is justified by the importance of this type of systems in real applications [14] which is certainly a useful result. 2 synchronization using a universal nonlinear control law in this section, we consider two arbitrary 3-d continuous-time quadratic systems. the one with variables x1, y1, and z1 will be controlled to be the new system given by 348 zeraoulia elhadj:a universal nonlinear control law for the synchronization of arbitrary...  x′1 = a0 + a1x1 + a2y1 + a3z1 + f1(x1, y1, z1) y′1 = b0 + b1x1 + b2y1 + b3z1 + f2(x1, y1, z1) z′1 = c0 + c1x1 + c2y1 + c3z1 + f3(x1, y1, z1) (1) where  f1(x1, y1, z1) = a4x 2 1 + a5y 2 1 + a6z 2 1 + a7x1y1 + a8x1z1 + a9y1z1 f2(x1, y1, z1) = b4x 2 1 + b5y 2 1 + b6z 2 1 + b7x1y1 + b8x1z1 + b9y1z1 f3(x1, y1, z1) = c4x 2 1 + c5y 2 1 + c6z 2 1 + c7x1y1 + c8x1z1 + c9y1z1 (2) and the one with variables x2, y2, and z2 as the response system  x′2 = d0 + d1x2 + d2y2 + d3z2 + g1(x2, y2, z2) + u1(t) y′2 = r0 + r1x2 + r2y2 + r3z2 + g2(x2, y2, z2) + u2(t) z′2 = s0 + s1x2 + s2y2 + s3z2 + g3(x2, y2, z2) + u3(t) (3) where  g1(x2, y2, z2) = d4x 2 2 + d5y 2 2 + d6z 2 2 + d7x2y2 + d8x2z2 + d9y2z2 g2(x2, y2, z2) = r4x 2 2 + r5y 2 2 + r6z 2 2 + r7x2y2 + r8x2z2 + r9y2z2 g3(x2, y2, z2) = s4x 2 2 + s5y 2 2 + s6z 2 2 + s7x2y2 + s8x2z2 + s9y2z2 (4) here (ai, bi, ci)0≤i≤9 ⊂ r30 and (di, ri, si)0≤i≤9 ⊂ r30 are bifurcation parameters, and u1(t), u2(t), u3(t) are the unknown (to be determined) nonlinear controller such that two systems (1) and (3) can be synchronized. first, let us define the following quantities depending on the above two systems in which we can proceed with our proposed method:  ξ1 = a1 + d1 + a4(x1 + x2) + d4(x1 + x2) + a7y1 + a8z1 + d7y2 + d8z2 ξ2 = a2 + d2 + a5(y1 + y2) + d5(y1 + y2) + a9z1 + d9z2 ξ3 = a3 + d3 + a6(z1 + z2) + d6(z1 + z2) ξ4 = η1 + η2 + η3 ξ5 = b1 + r1 + b4(x1 + x2) + r4(x1 + x2) + b7y1 + b8z1 + r7y2 + r8z2 (5) and advances in systems science and applications (2012) vol.12 no.4 349  ξ6 = b2 + r2 + b5(y1 + y2) + r5(y1 + y2) + b9z1 + r9z2 ξ7 = b3 + r3 + b6(z1 + z2) ξ8 = η4 + η5 + η6 ξ9 = c1 + s1 + c4(x1 + x2) + s4(x1 + x2) + c7y1 + c8z1 + s7y2 + s8z2 ξ10 = c2 + s2 + c5(y1 + y2) + s5(y1 + y2) + c9z1 + s9z2 ξ11 = c3 + s3 + c6(z1 + z2) + s6(z1 + z2) ξ12 = η7 + η8 + η9 (6) where  η1 = d4x 2 1 + d7x1y2 + d8x1z2 + d1x1 − a4x2 − a7x2y1 − a8x2z1 η2 = −a1x2 + d5y 2 1 + d9y1z2 + d2y1 − a5y 2 2 − a9y2z1 − a2y2 η3 = d6z 2 1 + d3z1 − a6z 2 2 − a3z2 − a0 + d0 η4 = r4x 2 1 + r7x1y2 + r8x1z2 + r1x1 − b4x 2 2 − b7x2y1 − b8x2z1 η5 = −b1x2 + r5y 2 2 + r9y1z2 + r2y1 − b5y 2 2 − b9y2z1 − b2y2 η6 = r6z 2 1 + r3z1 − b6z 2 2 − b3z2 − b0 + r0 η7 = s4x 2 1 + s7x1y2 + s8x1z2 + s1x1 − c4x 2 2 − c7x2y1 − c8x2z1 η8 = −c1x2 + s5y 2 1 + s9y1z2 + s2y1 − c5y 2 2 − c9y2z1 − c2y2 η9 = s6z 2 1 + s3z1 − c6z 2 2 − c3z2 − c0 + s0 (7) the above quantities comes from the formulation of the problem as the system in (8) below. now let the error states be e1 = x2−x1, e2 = y2−y1, and e3 = z2−z1. then the error system is given by e′1 = ξ1e1 + ξ2e2 + ξ3e3 + ξ4 + u1(t) e′2 = ξ5e1 + ξ6e2 + ξ7e3 + ξ8 + u2(t) e′3 = ξ9e1 + ξ10e2 + ξ11e3 + ξ12 + u3(t) (8) we propose the following universal control law for the system (3): u1 = −(ξ1 + 1)e1 − (ξ2 + ξ5)e2 − ξ4 u2 = −(ξ6 + 1)e2 − (ξ7 + ξ10)e3 − ξ8 u3 = −(ξ3 + ξ9)e1 − (ξ11 + 1)e3 − ξ12 (9) then the two 3-d continuous-time quadratic systems (1) and (3) approach synchronization for any initial condition. indeed, the error system (8) becomes 350 zeraoulia elhadj:a universal nonlinear control law for the synchronization of arbitrary...  e′1 = −e1 − ξ5e2 + ξ3e3 e′2 = ξ5e1 − e2 − ξ10e3 e′3 = −ξ3e1 + ξ10e2 − e3 (10) and if we consider the lyapunov function v = e21+e22+e23 2 , then it is easy to verify the asymptotic stability of the error system (10) by lyapunov stability theory since we have dv dt = −e21 − e22 − e23 < 0 for all (ai, bi, ci)0≤i≤9 ⊂ r30, (di, ri, si)0≤i≤9 ⊂ r30 and for all initial conditions. in particular, if the two systems (1) and (3) are chaotic, then the control law (9) guarantees also their synchronization for any initial condition. a practical example of this situation can be found in [8]. on the other hand, any 3-d continuoustime quadratic chaotic system can be stabilized (controlled) to a stable 3-d continuous-time quadratic system that converges to an equilibrium point (to a 3-d continuous-time quadratic system that converges to a periodic solution). furthermore, any 3-d continuous-time quadratic system can be chaotified to a chaotic 3-d continuous-time quadratic system. 3 example the most known example of 3-d quadratic systems, is the original lorenz system given by:  x′1 = a1x1 − a1y1 y′1 = b1x1 − y1 − x1z1 z′1 = −c3z1 + x1y1 (11) to apply the above method, we consider the one with variables x2, y2, and z2 as the response system  x′2 = d1x2 − d1y2 y′2 = r1x1 − y2 − x2z2 + u2(t) z′2 = −s3z2 + x1y2 + u3(t) (12) thus, the universal control law for the system (12) is given by: u1(t) = −(ξ1 + 1)e1 − (ξ2 + ξ5)e2 − ξ4 u2(t) = e2 − ξ8 u3(t) = −ξ9e1 − (ξ11 + 1)e3 − ξ12 (13) advances in systems science and applications (2012) vol.12 no.4 351 where  ξ1 = a1 + d1, ξ2 = −a1 − d1, ξ4 = d1x1 − a1x2 − d1y1 + a1y2 ξ5 = −z1 − z2, ξ6 = −2, ξ8 = −x1z2 + z1x2 − y1 + y2 ξ9 = y1 + y2, ξ11 = −c3 − s3, ξ12 = x1y2 − x2y1 − s3z1 + c3z2 (14) in particular, if the two systems (11) and (12) are chaotic, then the control law (13) guarantees their synchronization for any initial condition. 4 conclusion we have presented a universal nonlinear control law (without any conditions) for the synchronization of arbitrary 3-d continuous-time quadratic systems. this universal law (9) can be considered either as a stabilization, or a control, or as a chaotification approach for the system under consideration. references [1] l.m pecora and t.l carroll. (1990), “synchronization in chaotic systems”, phys. rev. lett, vol.64, pp.821-824. [2] l kocarev and u parlitz. (1995), “general approach for chaotic synchronization with application to communication”, phys. rev. lett, vol.74, pp.50285031. [3] x. tan, j. zhang and y. yang. (2003), “synchronizing chaotic systems using backstepping design”, chaos, solitons & fractals, vol.16, pp.37-45. [4] h. k chen, t. n lin and j. h chen. (2003), “the stability of chaos synchronization of the japanese attractors and its application”, jpn. j. appl. phys, vol.42, pp.7603-7610. [5] m. c ho and y. c hung. (2002), “synchronization two different systems by using generalized active control”, phys. lett. a, vol.301, pp.424-428. [6] m.t.yassen, (2005), “chaos synchronization between two different chaotic systems using active control”, chaos, solitons & fractals, vol.23, pp.131-140. [7] h.k. chen. (2005), “global chaos synchronization of new chaotic systems via nonlinear control”, chaos, solitons & fractals, vol.23, pp.1245-1251. [8] j. h park. (2005), “chaos synchronization of a chaotic system via nonlinear control”, chaos, solitons & fractals, vol.25, pp.579-584. [9] l. huang, r. feng and m. wang. (2004), “synchronization of chaotic systems via nonlinear control”, phys. lett. a, vol.320, pp.271-275. 352 zeraoulia elhadj:a universal nonlinear control law for the synchronization of arbitrary... [10] a. pikovsky, m. rosenblum and j. kurths. (2001), synchronization, a universal concept in nonlinear science, cambridge university press, cambridge. [11] e. mosekilde, y. mastrenko and d. postnov. (2002), chaotic synchronization: applications for living systems, world scienctific, singapore. [12] h. k. chen. (2005), “synchronization of two different chaotic systems: a new system and each of the dynamical systems lorenz”, chen and lu, chaos, solitons & fractals, vol.25, pp.1049-1056. [13] j. lu, x. wu and j. l. (2002), “synchronization of a unified chaotic system and the application in secure communication”, phys. lett. a, vol.305, pp.365370. [14] g. chen. (1999), controlling chaos and bifurcations in engineering systems, crc press, boca raton, fl. corresponding author zeraoulia elhadj can be contacted at:zeraoulia@mail.univ-tebessa.dz, and zelhadj12@yahoo.fr. advances in systems science and application (2016) vol.16 no.1 85-94 design of full adder using subthreshold dtpt logic kishore sanapala and r.sakthivel school of electronics engineering, vit university, vellore, tamilnadu, india. abstract as technology scaling has enabled small, design of digital circuits optimal for subthreshold region is becoming an active area for ultra low power applications.this paper presents the design of full adder using subthreshold dynamic threshold pass transistor (sub-dtpt) logic to achieve low power with acceptable performance. the simulations are carried out in cadence 90nm technology for vdd=0.2v and 1.2v. from the simulations the power delay product (pdp) of the proposed design is found to be extremely low in subthreshold region and is reduced by more than 80% when compared with the earlier reports. in strong inversion region the proposed adder is achieved more than 50% savings in delay when compared with the existing designs. keywords full adder; pdp; subthreshold region; sub-dtpt logic; ultra low power. 1 introduction increase in the usage of modern portable battery operated devices such as wearable electronics, cellular phones, ipods and remote sensors are leading to the demand of ultra low power applications. from [? ? ? ? ? ? ? ], it is noted that digital sub-threshold circuit design became the assured method for achieving the ultra-low power applications with acceptable performance. the sub-threshold voltage is basically the supply voltage (vdd) less than the threshold voltage (vth) of the transistor, circuits operating in this region consider the sub-threshold leakage current of the device for the necessary computations. exponential reduction of power at the cost of reduced performance is the impact with the subthreshold region of operation. this impacts being positive, there has been interest for the digital computations which uses the subthreshold leakage current achieving in ultra low power consumptions in portable computing devices. the formerly parasitic sub-threshold leakage current is exploited by the operation in the sub-threshold or weak-inversion region, and this exploited leakage current is considered as the primary operation current[? ]. these primary currents are considered to be much weaker comparatively to the standard strong-inversion currents; this will increase the time for charging or discharging the capacitive nodes which will limit the operation frequency of the circuit. in order to achieve ultra low power benefits, these leakage currents expected to drive the logic should be minimized in the device off state. for an mos transistor, when the gate to source voltage (vgs) of the transistor 86 kishore sanapala and r.sakthivel: design of full adder using subthreshold dtpt logic is biased under the threshold voltage (vth), the subthreshold or weak inversion region of operation occurs. vth is independent of the drain bias and the channel in case of a long channel device. the case differs when comes to submicron channel lengths and results in the effect of drain induced barrier lowering (dibl)[? ]. the drain current in the subthreshold region of operation is [? ] ids = i0e (vgs − ηvds − vth)/ηvt (1− e −v ds vt ) (1) where i0 = µ0c0x w l (n− 1)vth 2 (2) and the parameters vt thermal voltage (kt/q = 26mv at 3000k) η dibl coefficient n subthreshold swing coefficient (n= 1 + cdep/cox) µ0 zero bias mobility cdep depletion capacitance cox oxide capacitance w effective width of the channel l length of the channel vds drain to source voltage as device behavior depends on many parameters, in this study of dtpt logic, bulk terminal voltage is the key variable choice. varying the bulk terminal potential yields to further perceive into the device behavior to changes at the drain/ source and gate terminals. in the standard cmos configuration the bulk of the nmos is tied to the ground and the bulk of the pmos is tied to the vdd for an inverter. it is found that the drain current increases, when the bulk potential raise above the ground and the vdd to below threshold for the nmos and pmos devices respectively[? ]. one solution to increase/decrease subthreshold currents in on/off states is to use dtmos configuration[? ], where the bulk terminal is tied to its gate as shown in fig.1 (a). using dtmos pass gate along with the augmented device tends to a new configuration as shown in fig.1 (b) called dtpt[? ]. this paper presents the full adder design using sub-dtpt logic for the first time in the literature. the remaining of this paper is organized as follows. section 2 presents the brief description about existing full adder designs reviewed in comparison with the proposed design. the proposed sub-dtpt full adder design is described in section 3. simulation results and performance parameters advances in systems science and application (2016) vol.16 no.1 87 of the comparisons made with respect to different adders are presented in section 4. finally some conclusions are summarized in section 5. 2 existing full adder designs the adder is the basic building block in many of the vlsi systems such as digital signal processors (dsp) and microprocessors. since adder is the core module in many of the arithmetic operations, enhancing the performance of this module would lead to enhance the overall system performance. thus the design of full adder with lower pdp becomes the engineers interest for implementing the modern vlsi systems. number of adder circuits was designed using different logic styles and circuit techniques to reduce the power consumption and delay. the functionality of the adders compared in table1 is similar but differ in the design methodologies. each logic style tends to favor one of performance aspects like power, delay, and area but at the expense of the other. the different full adders are considered for performance comparison in conventional strong inversion region is discussed briefly in the following. the standard complementary metal oxide semiconductor (cmos) adder is the more robust and most conventional design[? ? ? ]. it is designed using the regular cmos structure with pull up and pull down networks. another conventional complementary pass transistor logic (cpl) with swing restoration logic is proposed in[? ? ]. it produces the complementary output of the many intermediate switching nodes. voltage degradation is the main issue in the performance of cpl design. this was improved with the transmission gate full adder (tgfa) designs[? ? ]. transmission gate (tg) is constructed by connecting the pmos and nmos transistors in parallel. but the driving capability of the tgfa design is less when cascaded. this results in performance degradation. later many hybrid logic styles which use more than one logic style in implementation are proposed. a 14 transistor hybrid adder was proposed in [? ? ] and zhang et al. proposed hybrid pass logic with static cmos output drive (hspc) adder [? ]. in this logic the xor and xnor functions are generated concurrently by using pass transistors and implemented in cmos module to produce full swing outputs. chiou-kou tung et al. proposed another hybrid fa core which uses the mirror type static cmos logic to improve output driving capability[? ]. sumeer goel et al. proposed a new hybrid adder which also targets high speed applications and is noise immune[? ]. mariano aguirre et al. proposed a hybrid double pass transistor logic (dpl) and swing restored cpl (srcpl) full adders[? ]. multiplexing of the boolean functions xnor/xor and or/and is used as internal logic style for implementing this full adder architecture. the different full adders considered for performance comparison in subthresh88 kishore sanapala and r.sakthivel: design of full adder using subthreshold dtpt logic old region (for vdd= 200mv) is discussed briefly in the following. direct synthesis fa [? ] is implemented based on karnaugh map driven digital design and is realized using standard nand and xor gates. the other three different adders proposed in [? ] are min3 stacked fa, min3 mirrored fa, and min3 ijcnn fa. these three adders are based on the stacked minority 3 element[? ], mirror minority 3 element[? ], cmos inverter and dynamically reconfigurable ijcnn element[? ? ]. min3 mirrored adder is suitable only for low performance applications because of its large power dissipation. min3 stacked and ijcnn adder architectures are applicable for low noise margin systems. 3 proposed sub-dtpt full adder many circuit techniques are introduced to operate the digital circuit in subthreshold region to meet the ultra low power requirement. in all the subthreshold digital circuits, power supply less than the threshold voltage of the transistors is used to power the circuit. in the subthreshold region, there is no conducting inversion channels, so therefore, the transistors behave in different manner as compared to when they are operated in a strong inversion region. therefore, there will be a change in the circuit properties, such as noise margin, tolerance to temperature and process variations. thus there are noted favorable changes, such as increased transconductance gain, near-ideal static noise margin, sensitivity of subthreshold circuit to power supply, temperature and process variations [? ? ]. this increase in the sensitivity may lead to improper functioning of the subthreshold circuits. thus to make sure proper functioning of the subthreshold circuits, moalemi et al.[? ] and lindert et al.[? ] proposed a logic family called sub-dtpt logic which gives lesser pdp when compared other subthreshold logic families reviewed in [? ]. (a) (b) fig. 1 (a) standard dtmos (b) dtmos with pass gate sub-dtpt logic uses the dynamic threshold transistors whose gates are tied to the substrates. to mitigate the drop problem, restoration has been commonly used; this will assist the pass-gate pull-up at the output. a new idea in pass-gate advances in systems science and application (2016) vol.16 no.1 89 logic is to use dynamic threshold mos (dtmos)[? ] for restoration, but the combination of the two will help scale the voltage even further. to provide a new symmetric design to the circuit, augmented circuit styles need to be customized for pass-gate logic, which makes it to add a secondary auxiliary device. the dtmos pass gate with the augmented device is known to be sub-dtpt and is shown in fig.1(b) and the standard dtmos is shown in fig.1(a). fig. 2 subthreshold dtpt nand gate fig. 3 full adder using sub-dtpt nand gates the design of nand gate using the sub-dtpt logic is as shown in the fig.2. 90 kishore sanapala and r.sakthivel: design of full adder using subthreshold dtpt logic the combinational circuit full adder is designed in sub-dtpt logic, since the full adders are the basic digital blocks in many of the arithmetic units. because of the universal characteristics, of nand/nor gates, engineers prefer these universal gates in the design of combinational digital circuits. in the proposed design, the full adder is realized using only nand gates. since nor gate has more delay because of the series stack of pmos transistors in series. the full adder using subthreshold dtpt nand gate is shown in fig.3. 4 results and discussion the proposed sub-dtpt full adder is simulated using cadence virtuoso tool in 90nm technology for the supply voltages of 1.2v (strong inversion region) and 200mv (subthreshold region). the obtained simulation waveforms are shown in fig.4. the performance parameters power, delay and power delay product (pdp) are measured for the proposed design and also compared with the existing designs. the comparisons for the proposed adder with different existing full adders in conventional strong inversion region and subthreshold region are shown in table 1 and table 2 respectively. the longest delay is considered as the cell delay which is obtained by considering the 50% of the input voltage swing to the 50% of output voltage swing. the average power is measured using the predefined calculator functions available in the cadence virtuoso tool. from the comparisons in the table 1, it is noted that in strong inversion region, the proposed design achieved major s avings in terms of delay and is 36.25% less than the fa-srpl, 58.26% less than the fa-dpl designs proposed in [? ], approximate 25% less than the sumeer goel design[? ] and chiou-kou tung table 1 results of the simulations for full adders in 90nm technology with vdd=1.2v design reference power(µw) delay(ns) pdp(fj) c-cmos [? ? ] 1.5799 0.1274 0.20127 cpl [? ? ? ] 1.7683 0.0791 0.13987 tgfa(16t) [? ? ] 1.7459 0.3258 0.5688 tgfa(20t) [? ? ] 1.7796 0.2348 0.41785 14t hybrid [? ? ] 3.3328 0.3377 1.1254 hspc-hybrid [? ] 1.576 0.2301 0.3626 mirror type-hybrid [? ? ] 7.707 0.1406 1.0836 fa-hyb [? ] 6.21 0.143 0.888 fa-dpl [? ] 7.34 0.254 1.864 fa-srpl [? ] 7.4 0.167 1.235 proposed [present] 12.3704 0.10646 1.31695 advances in systems science and application (2016) vol.16 no.1 91 (a) (b) fig. 4 (a) simulation response of the proposed full adder design for vdd=1.2v; (b) simulation response of the proposed full adder design for vdd=0.2v. et al.[? ], more than 53% less than the hybrid hspc adder design and 20t tgfa design, more than 65% savings than the vesterback et al.[? ] adder and 16t tgfa and also 16.4% less than the conventional cmos adder. it is also noted that the power delay product (pdp) of the proposed adder is 29.2% less than the fa-dpl design. table 2 results of the simulations for full adders in 90nm technology with vdd=0.2v design reference power(pw) delay(ns) pdp(aj) direct synthesis [? ? ] 854.1 263.5 225.3 c-cmos [? ? ] 931.6 162.4 151.3 min3-stacked [? ? ] 191.8 4767.7 914.5 min3-mirrored [? ? ] 1160 159.1 184.6 min3-ijcnn [? ? ] 5206 173.7 904.1 proposed [present] 2293 69.25 158.79 from the comparisons in table2 it is noted that proposed adder has achieved major savings in terms of pdp and is more than 82% lesser than the min3ijcnn and min3 stacked adder designs, 29.5% lesser than the direct synthesis adder design and also 13.98% lesser than the min3 mirrored adder design. from the above result analysis, it is observed that the delay of the sub-dtpt design is extremely low but the power is found to be very large. this increase in power is due to the excessive gate current caused by forward biasing the source body junctions. 5 conclusions in this paper, the design of full adder using dynamic threshold pass transistor logic is presented. the simulations are done using cadence virtuoso 90nm 92 kishore sanapala and r.sakthivel: design of full adder using subthreshold dtpt logic technology for supply voltages of 1.2v and 0.2v. at 1.2v, the results of the proposed design are compared with many logic styles like conventional cmos, cpl, tgfa, hybrid, srpl and dpl. at 0.2v, the results of the proposed logic is compared with different logic designs proposed earlier like standard cmos and min-3 based adders. from the simulations, it is found that in subthreshold region of operation the proposed design achieves major savings in the pdp compared with the previous reports. the proposed logic is useful in the implementation of the chip which overcomes the issues like power, delay, dibl, and process sensitivity. further improvements can be done for the proposed design in reducing the area overhead. references [1] k.roy and h.soeleman (1999), “ultra low power subthreshold digital logic circuits”, proc. ieee conf. low power electronics and design, pp. 94-96. [2] a.chandrakasan and r.broderson (1995), low power digital cmos design, springer, boston. [3] a.wang, b.h.calhoun and a.chandrakasan (2006), subthreshold design for ultra low power systems, springer. [4] j.m.rabaey and m.pedran (1996), low power design methodologies. springer, boston. [5] nyathi and jabulani (2006), “logic circuits operating in subthreshold voltages”, proc. ieee conf. low power electronics and design, pp. 131-134. [6] j.kao, s.narendra and a.chandrakasan (2006), “subthreshold leakage modeling and reduction techniques”, proc. ieee conf. low power electronics and design, pp. 131-134. [7] k.roy and h.soeleman (2001), “robust subthreshold logic for ultra low power operation”, ieee trans. circuits and systems, vol. 9, pp. 90-99. [8] j.m.rabaey, a.chandrakasan and b.nikolic (2003), digital integrated circuits-a design perspective (2nded.), pearson eduction. [9] n.lindert, t.sugii and s.tang et al. (1999), “dynamic threshold pass transistor logic for improved delay at lower supply voltages”, ieee journal of solid state circuits, vol.34, pp. 85-89. [10] zhang.m, j.gu and ch.chang (2003), “a novel hybrid pass logic with static cmos output drive full adder cell”, proc. int.symp circuits and systems, pp. 317-320. advances in systems science and application (2016) vol.16 no.1 93 [11] n.h.weste and david harris (2004), cmos vlsi designa circuits and systems perspective (3rded.), addison wesley. [12] issam.s, khater.a and bellaouar.a et al.(1996), “circuit techniques for cmos low power, high performance multipliers”, ieee journal of solid state circuits, vol.31, pp. 1535-1544. [13] a.m. shams, darwish.t and m.bayourmi (2002), “performance analysis of 1-bit low power full adder cells”, ieee trans. very large scale integration systems, vol. 10, pp. 20-29. [14] n.jhuang and h.wu. (1992), “a new design of the cmos full adder”, ieee journal of solid state circuits, vol.27, pp. 840-844. [15] k.navi, r.f.mirzaee and m.h.moyaeri at al. (2007), “a novel low power full adder cell with new technique in designing logical gates based on static cmos inverter”, micro electronics.j., elsevier, vol. 38, no. 1, pp. 130-139. [16] k.navi, v.foroutan and m.r.azghadi at al. (2009), “a novel low power full adder cell with new technique in designing logical gates based on static cmos inverter”, micro electronics.j., elsevier, vol. 40, pp. 1441-1448, aug. 2009. [17] c.k.tung, s.h.sheih and y.c.hung. (2007), “a low power hybrid cmos full adder for embedded system”, proc. ieee conf. design diagnostic. electron. circuits syst.,vol.13, pp. 1-4. [18] s.goel, a.kumar and m.a.bayoumi (2006), “design of robust, energy efficient, full adders for deep sub-micrometer design using hybrid cmos logic style”, ieee trans. very large scale integration systems, vol. 14, pp. 13091321. [19] mariano agurrie-hernandez and m.l.aranda (2011), “cmos full adders for energy efficient arithmetic applications” , ieee trans. very large scale integration systems, vol. 19, pp. 718-721. [20] kristian.aunet.s. (2006), “six subthreshold full adder cells characterized in 90nm cmos technology”, proc. ieee conf. design diagnostic. electron. circuits syst., vol.13, pp. 25-30. [21] aunet.s and y.berg (2005), “three sub-fj power delay product subthreshold cmos gates”, ifip vlsi soc, perth, australia. [22] aunet.s, oelmann.b and abdalla.s et al. (2004), “reconfigurable subthreshold cmos perceptron”, international joint conference on neural networks, ijcnn, budapest, hungary, vol.13., pp. 25-29. 94 kishore sanapala and r.sakthivel: design of full adder using subthreshold dtpt logic [23] aunet.s, norweigan and kretselement (2003), patent application, no.20035537, leiv eiriksson nyskapning, trondheim, norway. [24] b.h.calhoun, al.wang and a.chandrakasan (2005), “modeling and sizing for minimum energy operation in subthreshold circuits”, ieee journal of solid state circuits, vol.40, pp.1778-1786. [25] vahid.m and ali a.k. (2007), “subthrelod pass transistor logic for ultra low power operation”, ieee computer society annual symposim on vlsi., pp.490-491. [26] ramesh.v, s.dasgupta and r.p.agarwal (2009), “device and circuit design challenges in the digital subthreshold region for ultra low power applications”, hindawi publishing corporation vlsi design. [27] jiangmin.gu and ch.chang (2003), “ultra low voltage, low power 4-2 compressor for high speed multiplications”, proc. ieee international symposium on circuits and systems, pp. 321-324. [28] alioto.m, di cataldo.g and palumbo.g. (2007), “mixed full adder topologies for high performance low power arithmetic circuits”, micro electronics.j., elsevier, vol. 38, no. 1, pp. 130-139. [29] c.h.chang, j.m.gu and zhang.m. (2005), “a review of 0.18µm full adder performances for tree structured arithmetic circuits”, ieee trans. very large scale integration systems, vol. 13, pp. 686-695. [30] p.bhattacharyya, bijoy.k and s.ghosh et al (2014), “performance analysis of a low power high speed hybrid 1-bit full addert circuit”, ieee trans. very large scale integration systems, vol. 23, pp. 2001-2008. corresponding author kishore sanapala can be contacted at: kishorelendi@gmail.com study of external and internal factors affecting enterprise’s stability catherine g. zinovieva margarita v. kuznetsova tatyana v. dorfman pavel v. limarev juliya a. limareva nosov magnitogorsk state technical university, magnitogorsk, russia abstract. the paper analyzes the factors of enterprise’s stability meaning system’s ability to maintain its qualitative certainty via self-organization targeted to overcome the environment. the authors prove that enterprise’s stability is characterized as system’s activeness and adaptiveness and is formed under the influence of a set of external/internal environmental factors. the first directly depend on enterprise’s operations arrangement, the second are external to that arrangement and are outside of enterprise’s influence zone. so, the most significant direct influence factors are consumers, competitors, suppliers, contact audiences. indirect influence factors do not act on an enterprise directly but cause the environment to change which may affect an enterprise. this article covers on such factors like economic, state and political, scientific and technological, legal and socio-demographic, etc. finally the authors come to the strong conclusion: one of the most important ways to enterprise’s stability is the training of organizational culture of staff, featuring its philosophy, its basic operational principles (company’s relationships with suppliers, consumers and competitors), management style, attitude to staff, etc. keywords: enterprise’s stability, external environment factors, internal environment factors, direct and indirect influence factors, consumers, competitors, suppliers, contact audiences, activeness, system’s adaptiveness, innovational activity, pre-adaptive elements, human potential. 1. introduction an enterprise is one of complex systems functioning as the unity of stability and unsteadiness. synergetic approach understands system’s complexity as the ability for self-organization, typical, first of all, for nonequilibrium systems. stability of an enterprise may (temporarily) act as equilibrium. in equilibrium state, each element exercises its function, stably and independently from other elements. meantime the stability of an enterprise is ensured through management, i.e., the ability of entrepreneur/manager to maintain steadiness of functions. but, as an enterprise is an open system, equilibrium may not be a moment of its existence. own existence of an enterprise is an open nonequilibrium system capable for self-organization. in that case, the stability mechanism acts otherwise: the system behaves as a whole but the level of elements freedom decreases and coherent interaction of elements, forces, goals, motives, wishes occurs. in such a system, fluctuations are not suppressed but intensified, thus causing the growing role of actions coordination and efforts cooperation inside the system. the objective of this paper is to prove that stability of an enterprise is characterized as system’s activeness and adaptiveness and is formed under the influence of a set of external/internal environmental factors. in the economic literature, self-organization is equivalent to self-adjustment borrowed from engineering. but j. m. keynes in his day showed that self-adjustment idea does not reflect the specifics of the contemporary economy and causes fatal consequences in practice. keynes offered to add to self-organization the relative organization on the macroeconomic level. the stability mechanism of a separate enterprise is the same. so, stability of an enterprise is system’s ability to maintain its qualitative certainty via selforganization targeted to overcome the environment [8]. 2. choosing methodology. to assess the level of stability of an enterprise and build its stable development strategy, the factors affecting operational stability should be analyzed. in the study of those factors we are based on the following methodological provisions: – stability as a complex dynamic feature is not just subject to the influence of a great number of factors but is formed and maintained by their whole interacting aggregate; – the influence of any factor depends on the development of other factors and the aggregate of subjective conditions created in the economic system. being closely interrelated, stability factors often are oppositely directed affecting enterprise’s operation: some are positive, some – negative. the negative effect of some factors is able to reduce or even eliminate the positive effect of others, and vice versa. 3. main part. the availability of a great number of various factors makes it necessary to group them. various characteristics may become the classification basis: – by place of existence – external and internal; – by nature of effect – objective and subjective; – by level of influence – basic and secondary; – by structure – simple and complex; – by time of effect – permanent and temporary, etc. we hold the opinion that enterprise’s stability is formed under the influence of a set of factors of internal and external environment. the first directly depend on enterprise’s operations arrangement, the second are external to that arrangement and are outside of enterprise’s influence zone. external factors of enterprise’s stability are the conditions which may not as a rule be changed but should be accounted for by the subject of an enterprise as they affect the state of its affairs, i.e., the factors outside an enterprise. external factors are interrelated: change of a one may cause change of others and therefore their effect on enterprise’s stability is correlated. among the total aggregate of external factors of enterprise’s stability, foreign economic trade, production, market, information and innovational infrastructure factors may be specified. external factors affecting the level of enterprise’s stability are logically divided into two groups: – direct influence factors; – indirect influence factors. direct influence means the straight impact of environment on an enterprise. indirect influence is the environmental impact on an enterprise causing changes of its operating conditions. external factors of direct influence include consumers, competitors, suppliers, contact audiences – those are the factors manifesting through the activity of enterprise’s subjects, although having objective reasoning. an enterprise in the course of its activities continuously interacts with various elements and processes, so a rather significant change of activities always causes the change of an enterprise itself. the reasons like growth of raw materials prices and loan interests, loss of buyers, etc. may change enterprise’s structure, its actual form and size. the most important factor is consumers forming the sales market. there is a widespread point of view that the only true purpose of a business is to create consumers. that means the following: enterprise’s survival and justification of existence depends on its ability to find consumers of its activity and meet their demands. in the scientific literature, there are various approaches to consumers classification. to study and analyze the consumers, their aggregate may be grouped as follows: – direct consumers; – manufacturers receiving products/services for production purposes; – intermediaries buying products/services for resale; – governmental organizations buying products/services for own use. regarding the relationships between an enterprise and its products buyer, in the course of business an entrepreneur does not have to enter into relationships with investors on calling for capital (having own capital), does not have to enter into relationships with employees (doing all production activities independently), but there has to be a relationship with the consumer/buyer of products/services. in any case, an enterprise has to cooperate with the consumer as the need for production process arrangement and its length depend on the consumer: it is the consumer which determines whether the production process of a business will further exist. the consumer, from producer’s point of view, is its main partner. 4. discussion however, in the contemporary world producer’s dependency on the consumer changed greatly. as noted by j. baudrillard, the idea of primary needs is a myth [3, p. 54]. the things are not exhausted with what they serve for, but are imposed prestige and signs of social position. political economy traditionally explains the reasoning of production by the existing needs. as opined by j. baudrillard, it camouflages the internal feasibility of production order [3, p. 64]. in fact, producers manage the needs via advertising. the real power lies in the real spheres of decision making, management of needs, manipulation of signs and people [3, pp. 99-100]. thus, that external factor of stability is gradually getting the internal nature. the main stability factors include the competition of business entities. in many cases, not consumers but competitors determine which kind of products may be sold and at what price. consumers are not the only object of competition between businesses. the latter may compete for labor resources, materials, capital and the right to use some technical novelties. internal factors like working conditions, labor compensation, nature of relationships between leaders and subordinates are depending on competitive response. competitive success is ensured to those which can find new needs, arrange production of new goods and introduce new technologies, thus making competition a tool of economic play, making businesses review their strategies. competition affects the quantity and quality of products made, removes inefficient enterprises from production, assists in rational use of resources, prevents from producer’s (monopolistic) dictatorship over consumers. for growth and prosperity, enterprise’s subject needs capital suppliers. those may be banks, individuals, governmental institutions engaged in loans granting. the wish of investors to invest their capital may be a certain criterion of enterprise’s success as the higher its performance statements, the higher the probability to get the required funds on good terms and conditions. the significance of factor suppliers of material resources is characterized by the fact that in an enterprise the portion of material resources in production cost-price is 60-80% and higher. contact audiences are the external forces directly affecting enterprise’s decision making due to various kinds of interests related to its business. contact audiences may be characterized as follows: – governmental authorities on supervision and regulation of business; – mass media (advertising agencies, newspapers, magazines, radio and tv); – public organizations, trade unions, civil public opinion groups, etc.; – local contact audiences: communities, religious organizations, etc. governmental institutions play an important role in regulating the activities of enterprises. enterprise’s legal status reasons a certain taxation procedure. various governmental regulators are authorized to set forth the procedure of financial accounting keeping, issue licenses, set the standards and operating conditions, etc. indirect influence factors do not affect an enterprise straightly, but cause the changes in the external environment which may affect an enterprise. they are economic, governmental and political, scientific and technological, legal, socio-demographic and others. indirect influence environment is usually more complex than that of direct influence. top management, taking incomplete information as the basis and trying to prognosticate possible consequences for an enterprise often has to rely upon the suggestions about such environment. for efficient interaction with the external environment, the leader has to continuously analyze the dynamics of such environment (changing structure of factors) taking into account the following: – structure of factors is rather complex and there is plenty of factors; – the level of each factor’s impact on business structure is different; – some factors are continuous, others – short-term; – external environment changes are mobile, chaotic and rather violent, making hard to monitor them. to ensure survival and achievement of the goals to be sought, enterprise’s subject should be able to efficiently response and adapt to the changing environment. enterprise’s stability is system’s activeness and adaptiveness. adaptiveness is the representation in the reflecting system of the specifics of the reflected one, or system’s response to environment’s influence. adaptiveness in the general sense is characterized as passive adaptation of a system to the external environment [2]. it is oriented at the compensation of negative consequences caused by external factors. enterprise’s adaptiveness is characterized by its ability to change the functioning method in compliance with the environment’s changes and may not be considered absolutely passive as it is exercised through conscious activity of individuals – employees and owners of an enterprise. nevertheless, it is relative activeness as it is forced by external factors. activeness suggests feasible response of a system directed to the external environment and related to the elimination of negative external factors by acting on the elements of the environment producing them [7]. enterprise’s activeness is the implementation of measures aimed at decrease of external environment factors’ dependency. however, if during the adaptation the dependency is decreased by elimination of the influence of factors, active behavior suggests preventive measures of an enterprise aimed to eliminate the external environment factors. w.r. ashby [14] worded the basic principle of systems stability – law of requisite variety, according to which only a variety may destroy a variety so the variety of system’s states should not be lower than external environment’s variety. n.v. chepurnykh notes: if an enterprise is capable to affect the external environment’s factors, it may be used for reducing environment’s variety [12, p. 76]. we opine that through adaptiveness an enterprise reduces the variety of the external environment by eliminating its factors. enterprise’s active behavior contributes to the growth of the variety of states of an enterprise itself. enterprise’s adaptiveness as a way of functioning in compliance with the external environment changes regarding direct factors of the environment is characterized by the level of control of enterprise’s top management over its liabilities. in economic literature on business matters some four groups of enterprise’s stability factors are specified: 1) innovations and investments; 2) resources and their use (material/financial, human, technical/technological facilities, information); 3) marketing quality and level of use; 4) organizational culture of a business. it should be noted that the above groups include dozens of specific factors selectively acting in each business structure. an enterprise may increase the variety of its states and keep ahead of the environment’s variety via production optimization, introduction of new technologies and kinds of products. enterprise’s effect on the environment is achieved by establishment and development of competitive advantages causing the consumer market structure to change ensuring reliability of raw materials and semi-finished products supply and sales. according to m. porter’s theory of five competitive forces [18], achievement of competitive advantages by an enterprise is ensured by the developed technology, innovative equipment, intellectual resources and other elements. enterprise’s stability may be ensured either via keeping traditional business forms or via innovations. innovation notion relates to a new product or service, their production method, novelties in organizational, financial, scientific and research and other spheres. any improvement ensuring costs saving or creating any conditions for such saving may be deemed an innovation. the innovational process combines science, technology and management covering the aggregate of production, exchange, consumption relationships. under institutional evolutionary theory, a number of models was developed explaining company’s development as the result of its innovative activity. j. schumpeter specified seven directions of innovative changes on the microlevel: products, technologies, sale markets, raw materials, semi-finished products, organizational structure principles. through innovations, an enterprise increases the variety of products range, resources used, production and management methods. it was found that between innovative growth and relative quality of products there is a certain correlation, while enterprises offering a great number of new products compared to its competitors the probability to offer more useful products from consumer’s point of view is higher. intensive development of production implies opportunity to increase production output without extra resources as distinct from extensive development which is oriented to call for extra production resources [13]. intensive use of resources means the introduction of new technologies and production optimization. an essential consequence of the intensification process is growing products quality while cutting cost-price, which makes them more attractive for consumers. therefore, an enterprise becomes less dependent on certain buyers’ groups. the use of new saving technologies enables to increase production output using less resources and makes production more flexible and not dependent on a particular kind of resource. meantime, cutting cost-price makes price increase on raw materials and semi-finished goods a less substantial factor. increasing innovations quantity and quality, intensifying and diversifying production, an enterprise may affect the environment’s variety via increasing the variety of own states: product differentiation and product line expansion, opportunities for consumption of larger volume of interchangeable resources, products quality improvement, increasing number of buyers and markets. staff is the central factor of any enterprise. this is because an individual, his/her norms of behavior and values greatly determine the achievement of enterprise’s goals. to ensure staff’s normal working conditions, it is required to take into consideration individual skills of employees, develop their creative potential and apply it, accounting for socio-psychological and physiological needs of people, improve professional training thus ensuring social and psychological stability of the staff team. the account for those aspects should be the cornerstone of enterprise’s potential. the model of keeping enterprise’s stability suggests the availability of two parameters at least: business efficiency and free acts of the leader and the staff. freedom means growing role of choice of each element functioning as acts of certain subjects. systems methodology worded so-called law of varieties exchange: in social systems, dropping variety on the macrolevel (economic life of the society) is accompanied by growing variety on the microlevel (in our case, we may speak of an enterprise) [11, p. 125]. 5. result. this, the ability of top management to manage human resources rapidly taking into consideration the growing role of employee’s personality, knowing his/her motivations, ability to create them and align with enterprise’s goals may become a reserve of enterprise’s stability. transferring the center of gravity to an individual, a manager creates the conditions for a new level of system’s existence where an individual acts not as a cog in the management system machine in general but as an independent element. from the point of view of system management methodology, organizational activity should be aimed at the development of active, constructive consciousness, accounting for free choice alternatives of production parties. as opined by i. novik, systems analysis of alternatives suggests three main methods: 1) analysis of compliance level of living program and its implementation tool; 2) pros and contras analysis enabling to assess objective/subjective conditions in the state of instability; 3) estimating consequences of alternatives [11, pp. 122-123]. to solve the strategic tasks, the information on the environment is required. information flows as a means of control ensure enterprise’s links with markets, consumers, scientific and engineering novelties, thus achieving enterprise’s work optimization. meantime, the information collected on the organizational environment should possess a number of characteristics: reliability (assuredness in non-distortion of the information received), timeliness (no delays in getting the information required), trustworthiness, exactness (quality characteristic describing object’s condition), sufficiency (information about an object should cover all of object’s state), necessity (only necessary knowledge about an object is permitted), usefulness (effect of its use should exceed the costs on its receipt), regularity, etc. information’s impact on sustainable development may open new opportunities by solving the following tasks: – growing requirements to information quality and content; – defining enterprise’s need for information by kinds and forms; – defining basic directions on collecting, processing and storing initial data; – planning actions on collecting and exchanging information between enterprise’s divisions; ensuring feedback; – creating database used in the course of development and implementation of marketing programs, planning and control at an enterprise; – creating system for meeting the technical needs in connection with collection, processing and storage of information; automation and computerizing of administrative and management work; – development of protection and security of information. the information system’s efficiency which determines stable operation of an enterprise is determined by flexibility, i.e., the ability to timely describe emerging problems and give possible solution variants, thus checking for potential opportunities of an enterprise, tasks and goals, alternative solutions, possible risks and effect. thus, information forms manager’s ability to analyze and choose alternatives, i.e., contributing to the creation of internal variety for keeping enterprise’s stability. it also should be noted that the decisive role in enterprise’s dynamics is played by the subject – the manager. the abilities ensuring system’s self-organization in crisis conditions depend on so-called pre-adaptive elements. (pre-adaptation is the emergence of some or other useful characteristics in changing systems before they become actually useful [4, p. 79]). on the personal level it means cognitive complexity of thinking, high variation of human behavior and motivation. t.i. zaslavskaya [7] opines that the most important and decisive feature ensuring social integrity is human potential. a subject (a manager) should be correctly motivated, possess professional knowledge, morals, ability to think progressively. 6. conclusion. thus, one of the most important ways to enterprise’s stability is the training or organizational culture of employees [5], characterizing enterprise’s philosophy, its basic principles (interrelation of enterprise’s subject with suppliers, consumers and competitors), management style, of enterprise’s subject with suppliers, consumers and competitors), management style, attitude to employees, etc. stability is one of the sides of dynamic existence of an enterprise, the other side is its development. references 1. abalyan a.s. building the system of enterprise’s dynamic stability in modernization conditions // scientific review. – 2013. – no. 6. – p. 178-183. 2. agafonov v.a. principles of building structural model of strategy // strategic planning and development of enterprises. moscow: central economic and mathematical institute of ras, 2001. p. 21-22. 3. baudrillard j. for a critique of the political economy of the sign. moscow: biblion – russkaya kniga, 2003. 272 p. 4. vasilyeva l.n. theory of elites (synergetic approach) // public sciences and contemporaneity. 2005. no. 4. p. 75-85. 5. votchel l.m. philosophical analysis of ontological reasons of business: phd thesis, author’s abstract; magnitogorsk state university. magnitogorsk, 2000. 156 p. 6. golovanov p.v. improving financial and economic stability of industrial enterprises // scientific review. – 2013. – no. 7. – p. 158-160. 7. zaslavskaya t.i. business stratum of the russian society: notion, structure, identification // economic and social changes: public opinion monitoring. 1994. no. 5. p. 7-15. 8. zinovieva ye.g. enterprise’s existence: dialectics of stability and unsteadiness: phd thesis, author’s abstract; magnitogorsk state university. magnitogorsk, 2006. 138 p. 9. zinovieva ye.g., usmanova ye.g. enterprise’s stability factors // scientific review. – 2014. – no. 6. – p. 402-409. 10. kaminskiy m.a. economic stability of construction sector enterprises // scientific review. – 2013. – no. 9. – p. 692-694. 11. novik i. systematicity of optimal choice // public sciences and contemporaneity. 1995. no. 1. p. 118-126. 12. chepurnykh n.v. economy and ecology: development, catastrophes. moscow: nauka, 1996. 271 p. 13. sheremet a.d. methodology of financial analysis. moscow: infra-m, 1995. 176 p. 14. ashby w.r. principles of self-organization // principles of self-organization. moscow: mir, 1966. p. 314-343. 15. aldrich h., zimmer c. entrepreneurship through social networks / sexton p., smilor r. (eds) the art and science of entrepreneurship. – cambridge (ms), 1986. 16. carland j.w. diffierentioning entrepreneurs and small business owners: a conceptualization // academy of management rev. – 1984. – no. 9. – p. 358. 17. kirtzner i. competition and entrepreneurship. chicago: the university of chicago press, 1973. 18. porter m. competitive advantage. n.y.: macmillan, 1985. advances in systems science and applications (2012) vol.12 no.1 89-95 rapid manufacturing metallic parts via selective laser melting ruidi li1, yusheng shi1, jinhui liu2, mingzhang du3 and zhan xie3 1state key laboratory of material processing and die & mould technology, huazhong university of science and technology, wuhan 430074, china 2heilongjiang institute of science and technology, harbin, 150027, china 3shichuan petroleum perforating materials ltd, longchang, 642177, china abstract 316l stainless steel parts were manufactured via selective laser melting in this work. the surface morphology, microstructure, density and mechanical property were characterized. it is found that the surface contains a little amount of oxide and splash with balling effect. microstructure in low magnification shows features of scan molten tracks and molten pools; microstructure in high magnification depicts very fine crystal under rapid cooling. the tensile strength of as-received samples is 652.12 mpa, with a density of 95.6%. therefore the gas atomized 316l stainless steel powders could be used in manufacturing high quality parts with complex shapes via selective laser melting method. keywords selective laser melting, metallic parts, 316l stainless steel, powder 1 introduction selective laser melting (slm) is a newly developed rapid manufacturing technology, which can directly manufacturing intricate metal parts according to a three dimensional model from metal powders[1-4]. this forming technique is based on means of adding layers to build net shape components. each layer represents a slice geometry graph of the objective part in two-dimensional pattern. owing to its unlimited flexibility of geometry and complexity, slm is capable of manufacturing short run components which are not easily made through other forming method. moreover, metal parts made by slm need no or very little postprocessing procedure, and it is differ from selective laser sintering (sls). during slm process, metallic powders are fully melted because of a higher laser energy input. while in sls process, the forming mechanism is melting high polymer powder to binder un-melted metal powder. thus sls needs very trivial procedures such as post treatment to degrease to remove the binder. so it is urgent to develop and optimize slm technique with an aim to improve its practicability and universality in advanced manufacturing field. at present, slm technology is faced with many problems such low density, balling phenomenon, delaminating, warp and crack etc. that is due to the complex heat transfer mode under rapid moving gauss heat source, inducing acute temperature variation[5-6]. metal powders are melted in an extremely short time 90 ruidi li:rapid manufacturing metallic parts via selective laser melting coupled with multi-lines and multi-layers process, resulting in a very complex physicochemical process and multi modes heat and mass transfer. under this circumstance, the densification is especially important for creating high quality parts. in addition, the as prepared microstructure is particular comparing with other microstructure obtained from conventional forming method. therefore, investigation into a series of special phenomenon during slm is essential. in this paper, the densification, microstructure, balling effect, mechanical property of 316l stainless steel part was presented and the corresponding mechanisms were also addressed. 2 experimental procedures 2.1 powder materials 316l stainless steel powders (99% purity) were used in this experiment. these powders were prepared through water atomization in institute of powder metallurgy, central south university (csu). the morphology of the starting powder was examined by a quanta 200 scanning electron microscope (sem). it can be seen that these powders show sphere shape and bimodal-size feature, which can induce a high loose density and it is favorable for slm experiment. fig.1 sem figure showing morphology of 316l stainless steel powders 2.2 slm forming the forming process was carried on in the hrpm-iislm system which was developed by huazhong univ. of sci.&tech. (hust). the shaping space of the equipment was 250mm(l)×250mm(w)×200mm(h). this slm equipment contained a 100w continuous wave fiber laser. the building chamber of this system could also be vacuumed and protected by inert gas. forming process of slm was followed the listed procedures. (i) build 3d-cad model and transfer it into stl file. slice this 3d-cad model into horizontal layers according layer thickness and input it into slm system (ii) a quality of powders is dropped, then a roller spread a powder layer (iii)high energy laser scan the powder bed, thus the powders in scanned zone advances in systems science and applications (2012) vol.12 no.1 91 are completely melted (iv) the working platform descend a layer thickness (v) repeat procedure (b)-(d) until an integrated part is formed scan speed of 50∼100mm/s, laser power of 90∼100w, scan interval of 0.05∼0.2mm and layer thickness of 0.02∼0.1mm were selected for slm process. using above procedures and processing parameters, samples of slm-parts can be made as shown in fig.2. (a) (b) fig.2 samples of as fabricated slm parts 2.3 characterization at laser, densities of slm parts were measured through archimedes laser. surface morphologies and microstructure of as-received samples were analyzed by a quanta 200 sem. the element composition of the laser processed material was measured by energy dispersive x-ray spectrometer analysis. 3 experimental procedures 3.1 surface morphology surface characteristic could represent melting and solidification feature of slm process. fig.3 shows the surface morphologies. from the low magnification of fig.3a, it can be seen that the surface is relatively smooth in addition to some white flakes and splash. edx analysis resulted that the white flakes on the surface were oxygen rich, which is formed because of oxidation. it should be noted that the oxidation is disadvantageous for slm technique, due to a worsened wetting ability caused by oxide. in next layer forming process, the melted liquid could not wet the previous layer easily. therefore, it should be cautious to control oxygen content in atmosphere to prevent oxidation. moreover, it also should control oxygen contents in powder materials. fig.3b is the high magnification of splash form fig.3a, showing the detailed characteristic of splash. it exhibits spherical feature, with a large number of spheres on the top surface. the forma92 ruidi li:rapid manufacturing metallic parts via selective laser melting tion mechanism of balling effect during solidification process can be explained as follow. in liquid metal solidifying process, the surface energy tends to reach a lower value which is according to the lowest energy principle. thus under surface tension action, the many spheres are generated. (a) (b) fig.3 sem figure showing top surface morphologies of as prepared part (a) (b) fig.4 metallograph showing low magnification of microstructure. a: top view; b: side view 3.2 microstructure slm-produced microstructure is very different from microstructure obtained by conventional forming method. hence investigation on its microstructure is requisite. fig.4 shows the metallograph of top view and side view respectively, overall it can be seen that the microstructure in low magnification is dense with few pores. it is widely accepted that pores in a metal part are detrimental to mechanical property, so a high density is a significant index of a metal part. in this work the final density of 316l stainless is 95.6%, which is measured by displacement method. consequently this dense structure facilitates mechanical property. fig.4a describes the microstructure from top view. it can be seen that scan tracks are overlapped, and formed into a dense layer. fig.4b shows the microstructure advances in systems science and applications (2012) vol.12 no.1 93 from side view. it indicates that molten pools are accumulated tightly with a small amount of pores. fig.5 illustrates the microstructure in high magnifica(a) (b) fig.5 sem figures showing microstructure in high magnification of slmproduced material tion of slm-produced material. it is obvious that the grains are very fine and the grains grow along multiple directions. the extremely fine size is caused by the significantly rapid cooling rate during solidification process. a rapid solidification could inhibit grain growth, accordingly forming a large number of fine grains. the grains during molten pool solidification tend to grow along the easy growth direction. for cubic crystal, the growth directions are < 110 >. when one of the easy growth directions coincides with the heat flow direction, growth conditions are optimum. for this reason, among the randomly oriented grains, those grains have one of their < 100 > crystallographic axis most aligned with the heat flow direction will be favored. however, slm is a multi-line based technology in every cad slice layer, the moving heat resource controlled by a complex scan path can easily induce a very complicated heat diffusion process, thereby a varied heat transfer direction. this reason leads to a variable growth direction in order to accommodate heat flow direction. thus the as-received microstructures are formed. here, it should be pointed out that the slm produced parts contain fine microstructure, which is in favor of mechanical property. 3.3 tensile strength evaluation mechanical property is an important factor which ultimately determines its practical application. thus the tensile strength experiment was conducted with an eye to evaluate its mechanical property. a tensile test specimen was made by wire-electrode cutting from bulk slm fabricated samples. fig.6 expresses the stress-strain curves of extension test of slm-prepared 316l stainless steel materials. by calculation from above data, the tensile strength value of 652.12 mpa is obtained. to an extent, this mechanical property is equivalent to forgeable piece. therefore, this slm-produced parts can be used as many practical requirement in metal components. 94 ruidi li:rapid manufacturing metallic parts via selective laser melting fig.6 stress-strain curves from extension test 4 conclusions based on the investigation on selective laser melting 316l stainless steel powders, the following conclusion can be drawn. the gas atomized 316l stainless steel powders with spherical shape could be used as slm material for rapid manufacturing metallic parts. the corresponding properties and characteristic for evaluation slm-produced material are addressed. the surface morphology shows multi-lined feature, with a little amount of oxide and splash with balling effect. moreover, the microstructure in low magnification is composted of scan molten tracks and molten pools, which reflect the accumulative characteristic of rapid prototype technique. the microstructure in high magnification indicates the intrinsic nature of rapid cooling rate, resulting in extraordinary fine crystal. finally, the density of 95.6% is measured by drainage, possessing of tensile strength of 652.12 mpa. overall, slm technology in manufacturing 316l stainless steel parts shows its excellent feature in both flexible formability and high mechanical property, which can meet demand for practical applications. acknowledgements the authors would like to give thanks to the national high-tech program (863) of china (2007aa03z115), independent fund of state key lab. of material processing and die & mould technology of huazhong university of sci & technol and open fund of state key lab. of powder metallurgy of central south university of china (2008112022) the authors also thank for hua yan, bin hua and the analytic and testing center of huazhong university of science & technology for their assistance. advances in systems science and applications (2012) vol.12 no.1 95 references [1] yevko v, park cb, zak g,benhabib b. (1998), “cladding formation in laserbeam fusion of metal powder”, rapid prototyping, vol.4, pp.168-184. [2] mumtaz k.a, erasenthiran p, hopkinson n. (2008), “high density selective laser melting of waspaloy”, j. mater. process. technol, vol.195, pp.77-87. [3] asgharzadeh h, simchi a. (2005), “effect of sintering atmosphere and carbon content on the densification and microstructure of laser-sintered m2 high-speed steel powder”, mater. sci. eng. a, vol.403, pp.290-298. [4] simchi a, pohl h. (2004), “direct laser sintering of iron-graphite powder mixture”, mater. sci. eng. a, vol.383, pp.191-200. [5] simchi a. (2006), “direct laser sintering of metal powders: mechanism, kinetics and microstructural features”, mater. sci. eng. a, vol.428, pp.148158. [6] shen yf, gu dd, wu p. (2008), “development of porous 316l stainless steel with controllable microcellular features using selective laser melting”, mater. sci. technol, vol.24, pp.1501-1505. corresponding author author can be contacted at: shiyusheng@263.net advances in systems science and application (2016) vol.16 no.2 15-38 intrusion detection model using pca and ensemble of classifiers sumaiya thaseen1 and ch.aswani kumar2 1school of computing science and engineering, vit university, chennai, tamil nadu. 2school of information technology and engineering, vit university, vellore. abstract most of the intrusion detection systems examine all network features to identify intrusions with different classification approaches. the major challenges for any intrusion detection model is to achieve maximum accuracy with minimal false alarms. while many ensemble techniques are present to improve the accuracy of intrusion detection models, building an ensemble that can be generically applied for any network traffic is still a difficult task. in this paper, we propose a hybrid model for intrusion detection integrating base classifiers such as svm, linear discriminant and quadratic discriminant analysis.the aim of this paper is to identify the class label by constructing an individual classifier for each of the attack type and merging the results of every classifier. the resultant decision of the class label is obtained using weighted majority voting approach. we analyzed the performance of the model on two different data sets such as nsl-kdd and unsw-nb datasets. the experimental results indicate that the ensemble produces high accuracy in comparison to the base classifiers. as there is a huge class imbalance problem in network traffic, it is also observed that rather than relying on a single classifier, predicting the class label by weighted majority voting of svm, linear and quadratic discriminant classifier is an optimal solution which is proposed in this paper. keywords accuracy; intrusion; linear discriminant; quadratic discriminant; support vector machine. 1 introduction malicious intruders in the network are increasing day by day due to the rapid development of internet. the intruders can access, manipulate and disable the systems connected on the internet. intrusion detection systems (ids) are designed to discover the unauthorized access to computers in the network. intrusion detection systems are classified in two categories signature detection and anomaly detection. signature detection is used to identify attacks based on the known pattern of attacks. anomaly detection compares unknown profiles with known profiles and then identifies the unknown traffic profile as an attack. anomaly detection techniques have high false positive rates. many machine learning techniques have been used by researchers to overcome the disadvantages of anomaly detection models. several intelligent approaches such as svms [1], 16 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... anns [2], petri nets and data mining approaches[3, 4] have been used to build an ids. there are also many ensemble methods of machine learning which are more efficient than individual techniques that can reduce the false alarm rate and increase the classification accuracy. the different ensemble methods are bagging, boosting and stacking. bagging and boosting are mostly used to implement intrusion detection models as the stacking technique requires more time. there are two major limitations of existing approaches. the first limitation is even though there are many sophisticated detection techniques only few focus on feature representation for normal traffic and attack traffic which is a major issue to enhance performance of the classifier. the second issue is the computation time involved in integrating multiple techniques which may degrade the efficiency of on-line detection. the contribution of this research is to to build a desirable ids model with high accuracy using machine learning ensemble techniques with a feature reduction technique to identify the suitable features in ids. ensemble is preferred because the aggregation of multiple classifier predictions improves the accuracy of ids. in this paper we specify a manager as a combination of classifiers that will be generated depending on the class labels in each dataset. for instance, the manager will have 5 different classifiers in the svm, mnb and ldc ensemble for nsl-kdd data set as there are five class labels and the manager will have 9 different classifiers in the svm, ldc and qdc for unswnb dataset. thus the manager has a variable number of classifiers generated dynamically according to the number of class labels available in the dataset. thus the ensemble model utilizes a individual classifier for every class type and is an integration of base classifiers svm, linear discriminant and quadratic discriminant with the resultant class label predicted by weighted majority voting ensemble which is deployed to classify the different kinds of network attacks. a similar ensemble of classifiers was already developed [5] using four dif-ferent base classifiers namely linear discriminant classifier (ldc), quadratic discriminant classifier (qdc), k-nearest neighbor (knn) and back propagation and tested on four datasets namely hearth, diabetes, iris and transfusion. the uniqueness of the proposed model over earlier developed ensemble techniques are 1) the model can be evaluated on any real time intrusion detection datasets and the managers can be extended based on the number of class labels in the samples. 2) any combination of base classifiers can replace the existing techniques to improve the model. the rest of this paper is summarized as follows. section 2 provides a discussion on various developed intrusion detection models. section 3 provides the background of various techniques utilized in the model such as svm, linear disadvances in systems science and application (2016) vol.16 no.2 17 criminant , quadratic discriminant classifier and weighted majority approach. section 4 gives the overview of proposed intrusion detection model. experimental results and discussions are discussed in section 5. section 6 finally concludes the paper. 2 related work in this section we will discuss the various intrusion detection models developed using machine learning approaches, models developed by integrating classifiers and the ensemble techniques developed for intrusion detection. various artificial intelligence methods have been developed for intrusion detection models such as fuzzy logic [6], k-nearest neighbors [7], support vector machines [1], artificial neural networks [2], naïve bayes networks [8], decision trees [9], and genetic algorithms [10]. sumaiya et al. developed an intrusion detection model by using pca as the dimensionality reduction technique and svm as the classifier [11, 12]. the kernel parameters of svm are optimized by considering the variance of samples available in the same and different classes. this model provided a better classification accuracy. ajith et al. built a light weight ids using genetic programming approaches [13]. the experimental results proved that the accuracy was better in comparison to traditional intrusion detection models. kuanga et al. developed a hybrid kpca svm with ga model for intrusion detection [14]. the authors used kpca to extract the primary features of intrusion detection dataset. svm multilayer model is employed as the classifier to identify the attack. chebrolu et al. evaluated the performance of two feature selection techniques such as bayesian networks (bn) and classification and regression trees (cart) and also an integration of cart and bn [15]. results illustrate that feature selection is very effective in the development of real world intrusion detection models. sandhya et al. developed a hybrid intrusion detection model by combining decision trees and svm and also an ensemble of other base classifiers [3]. this hybrid model maximized accuracy and minimized complexity and results illustrated that the proposed model provided an accurate ids. perin et al. built a three layer multi classifier intrusion detection model to increase the overall accuracy [16]. the performances were analyzed from a vari-ety of combination techniques such as fuzzy k-nn classifier, na¨ıve bayes classifier and back propagation neural network classifier and the decision obtained from multiple classifiers are combined into a single result. the results proved that the detection performance is better than deploying a single classifier when using a full feature set or partial feature set. chandra and yao developed an ensemble based neural network wherein the outputs are combined in a form that resulted in a sig-nificant improvement in the generalization performance [17]. srinivas et al. built 18 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... an ensemble model of svm, artificial neural network (ann) and multivariate adaptive regression splines (mars) and analyzed the performance which was superior to individual approaches with respect to classification accuracy [18]. syarif et al. improved the accuracy and false positive rate of intrusion detec-tion by constructing an ensemble of bagging, boosting and stacking [19]. the base classifiers for these ensemble models were na¨ıve bayes, decision tree, rule induction and nearest neighbor. their results indicated that the accuracy for known intrusions was more than 99% but novel intrusions were identified with accuracy levels of 60%. thus bagging and boosting did not improve the accuracy significantly whereas stacking decreased the false positive rate by 46%. these ensembles increase the execution time and hence are not practical to be implemented in an ids. bahri et al built a ensemble method called greedyboost and compared with adaboost and c4.5 [20]. however the base classifier details are not specified in the paper but the results indicated that greedyboost scores a higher precision and recall in comparison to probe, u2r and r2l attacks present in kdd’99 dataset. bukhtoyarov et al designed neural network classifiers by applying a probabilistic approach to the network intrusion detection. genetic programming based ensembling (gpen) was deployed to design neural network ensembles [21]. they also analyzed with the kddcup 1999 dataset and classified the attacks. cordeiro and pappa utilized the particle swarm optimization (pso) by weighing the classifications obtained from different classifiers [22]. the four classification algorithms used were knn, näıve bayes, rocchio and svm. they used datasets of users of video social network for classification. their results outperformed the single classifiers. the motivation for selecting algorithms in the ensemble is due to the fact that an ensemble based on the four expert algorithms: linear discriminant classifier (ldc), quadratic discriminant classifier (qdc), k-nearest neighbor (knn) and back propagation. the ensemble is obtained by integrating the experts opinion with a weight coefficient assigned by weighted majority voting is already tested on four widely used datasets of hearth, diabetes, iris and transfusion. this ensemble resulted in better accuracy in comparison to simple majority voting approach, mean, maximum, minimum and median combiner [23]. thus many hybrid models using ensemble of techniques for intrusion detection have been developed but the major issue in ensemble approaches is the models were built and tested only on kdd datasets and thus a generic model that can deploy and test for any real time datasets is the necessity for the current scenario. the proposed model aims to overcome the issues in the existing ensemble approach and also with the advantage of developing modular structures that can have interchangeable positions. another advantage of our proposed ensemble deadvances in systems science and application (2016) vol.16 no.2 19 signs is that the algorithms can be replaced anytime with a more precise one. hence the approach aims to improve intrusion detection accuracy using simple techniques in ensemble learning integrated with pca as a dimensionality reduction technique. 3 background 3.1 preprocessing data preprocessing is very essential for huge data such as network traffic. reduction of redundant data and normalization are essential to be performed in preprocessing to build a balanced set of data. normalization is the process of transforming the data within a small specified range. the different normalization techniques are min-max normalization, z-score normalization and normalization by decimal scaling. we select z-score technique because it considers the mean and standard deviation of the attribute. d1 = b−mean (f) std (f) (1) where , mean(f)= sum of all attribute values of f std(f) = standard deviation of all values of f. 3.2 dimensionality reduction dimensionality reduction transforms the data in the high dimensional space to a lower dimension. pca performs a linear mapping of the data to a lower dimension such that the maximum variance for the data is obtained. the procedure begins with the construction of correlation matrix and computation of eigen vectors. the eigen vectors that correspond to the highest eigen values will be deployed to reconstruct the variance of the original data. transformation thus the original space is reduced to the space obtained by a few eigen vectors. the advantages of using dimensionality reduction are as follows: • time and storage space is reduced. • improves the performance of the machine learning model. • visualization of data is easier as the dimensions are reduced to 2d or 3d. 3.3 datasets we have analyzed the knowledge discovery and data mining 1999 standard dataset such as nsl-kdd data set which is widely used as intrusion detection benchmark datasets. the nsl-kdd data set contains roughly 33,300 samples. this dataset is chosen because of the following benefits [24]: 1) no redundant records in the training set. 2) due to the reduction in the number of records, the complete data 20 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... can be used for both training and testing. each packet in the dataset can be classified in any one of the classes namely normal, dos, u2r, r2l and probe. all the class labels except normal indicate the different attacks in the dataset. the other dataset analyzed in our model is unsw-nb dataset obtained from the cyber range lab of the australian center for cyber security (accs). this dataset contains nearly 1,56,000 samples falling in one of the nine classifica-tion categories namely normal, analysis, backdoor, reconnaissance, exploits, fuzzers, generic, dos and shellcode. this dataset has the advantages of con-taining current attacks in the network domain. 3.4 support vector machine support vector machines are widely used for classification and regression problems. svm is preferred over other techniques due to the low generalization error and less over fitting issues that arise from the training data set. there exist a possibility of high generalization error or overfitting if the model doesnt scale well on instances not available in the training set. svm is very effective on data samples that are separable in a linear fashion. the objective is to identify the hyperplane h that can split the instances into two categories such that samples in one class fall entirely on one side of h. as we can determine unlimited number of candidate hyperplanes, svm selects only the hyperplane that maximizes distance to the closest data samples in either class. this is known as margin maximization. the major features of svm are: • deals with very large data sets efficiently. • multiclass classification can be done with any number of class labels. • high dimensional data in both sparse and dense formats are supported. • expensive computing not required. • used in many applications like e-commerce, text classification, bioinformatics, banking and other areas. there are many real time applications where such a hyperplane does not exist. in such cases, svm utilizes a function to transform the data into a different feature space such that there is a possibility of separation. the function that performs such a transformation is called as kernel function. kernels play a major role in svm. the different kernel functions widely used along with svm are [25] as given below: i) linear kernel: k (xi, xj) = xixj j) polynomial kernel : k ( x, x ′ ) = ( xx ′ + 1 )d k) rbf kernel: k ( x, x ′ ) = exp ( −γ∥ x− x ′ ∥2 ) l) sigmoid kernel: k (xi, xj) = tanh ( yxi txj + r ) k (xi, xj) = tanh ( yxi txj + r ) advances in systems science and application (2016) vol.16 no.2 21 svm can be extended to multi-class classification, a set of binary classifiers are trained one for each class depending on the data set and its respective class labels.ie. if we train the nsl-kdd data set, then let i=15 be a index in the set s=(normal, probe, dos, u2r and r2l) and let bi denote the matching binary classifier for the target set s. similarly if we train the unsw-nb dataset, then let i= 1· · · 9 be a index in the set s=(normal, analysis, backdoor, reconnaissance, exploits, fuzzers, generic, dos and shellcode) and let bi denote the matching binary classifier for the target set s. thus the observations are classified using one-versus-all approach in both the datasets. to distinguish among the binary classifiers, we deploy manager to denote one set of classification. svm produces the best results when the rbf kernel function is utilized. experimental results show that the performance of svm classifiers will differ with the selection of rbf function. therefore in this paper we train the svm manager with five different rbf values= [ 5, 2, 1, 0.5 ,0.1] for both datasets to ensure that the svm algorithm is utilized maximally. this approach will ensure greater diversity of managers in ensemble classifier as the accuracy will vary for each binary classifier according to the selected rbf values in the vector. construct a set of binary classifiers f 1 , f 2· · · f n for 1· · ·n classes each trained to differentiate one class from the rest. a multi class categorization can be obtained by combining them according to the maximal output before applying the sgn function. argmax gk(x) where gk(x) = n∑ i=1 yiai kk(x, xi) + bk where k = 1 · · · n. (2) wherein gk (x) returns a signed real value which is the distance from the hyper plane to the point x. this value is referred as the confidence value. the higher the value, the more is the confident that the point x belongs to positive class. hence we need to assign x to the class having highest confidence value. given normal data χ = {x1, x2, ...xm} ∈ rd and let r be the radius of the hypersphere and c ∈ rd which is the center. the optimization problem can be solved by determining the minimum enclosing hypersphere. minimize r2 subject to ∥ ϕ (xj)− c ∥2 ≤ r2, j = 1, ...m (3) l (c, r, α) = r2 + m∑ j=1 αj{∥ ϕ (xi − c) ∥2 − r2} (4) 22 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... setting the derivatives δl (c, r, α) δc = c m∑ j=1 αj (ϕ (xj)− c) = 0 (5) we can obtain the following equation,∑m j=1 αj = 1 and c = ∑m j=1 αjϕ (xj) hence the equation (4) becomes, l (c, r, α) = m∑ j=1 αjk (xj , xj)− m∑ i,j=1 αiαjk (xi, xj) (6) which is the dual form of equation (4). the dual form of α can be obtained by solving the optimization problem, maximizing, w (α) = m∑ i=1 αik (xi, xi)− m∑ i,j=1 αiαjk (xi, xj) (7) subject to∑m i=1 = αi = 1 and αi ≥ 0 , i = 1 to m it should be noted that lagrange multiplier can be non-zero only if the inequality constraint is an equality for the solution. the complementarity conditions are satisfied by the optimal solutions α,(c, γ) αi{∥ ϕ (xi)− c ∥2 − r2}, i = 1...m (8) hence it implies that the training samples x i lie on the surface of the optimal hypersphere corresponding to αi > 0. the decision function becomes, f (x) = sgn ( r2 − ∥ ϕ (x)− c ∥2 ) this implies, = sgn(r2 − ϕ(x).ϕ(x)− 2 m∑ i=1 αiϕ(x).ϕ(xi) + m∑ i,j=1 αiαj(ϕ(xi).ϕ(xj))) = sgn(r2 − k(x, x)− 2 m∑ i=1 ϕ(xi)k(x, xi) + m∑ i,j=1 ϕiϕjk(xi, xj)))) (9) thus the aim of obtaining minimum enclosing hypersphere containing all training samples is satisfied. advances in systems science and application (2016) vol.16 no.2 23 3.5 linear discriminant classifier discriminant analysis is a classification problem where more than two groups of populations are known a priori and one or more samples from the population are classified according to the characteristics measured. the assumption is that the population πi has a probability density function of x which has a mean vector ui and variance-covariance matrix σ (similar for all populations). it is specified as f (x | πi) = 1 | σ |1/2(2π)p/2 exp [ −1 2 (x− ui) ′ σ−1 (x− ui) ] (10) we classify to the population for which p i f(x— i ) is the highest. lda is used when the variance-covariance matrix is not dependent on the population from the available data. in such cases the decision rule is based on the linear score function which is a function of the means for each of our g population ui and also the variance ccovariance matrix. the linear score function is as follows: si l (x) = −1 2 ui ′ σ−1ui ′ + ui ′ σ−1x+ logpi = d̂i0 + p∑ j=1 d̂ijxj + logpi (11) where di0 = −1 2ui ′ σ−1uiui ′ σ−1 dij = jth element of ui ′ σ−1 the far left hand expression represents a linear regression with intercept di0 and regression coefficients dij . di l (x) = −1 2ui ′ σ−1ui ′ + ui ′ σ−1x = d̂i0 + ∑p j=1 d̂ijxj given a sample unit with measurements x1,x2 · xp, the sample unit is classified into the population that has the highest linear score. this is comparable to the population that has the highest membership of posterior probability. linear score has to be calculated for each class of population and then the assignment of the sample to the population with highest score. but as this function utilizes unknown parameters ui and σ these parameters have to be determined from the data. hence discriminant analysis requires estimation of the following prior probabilities: pi = pr (πi) ; i = 1, 2, ...g population means: these can be determined by sample vectors. ui = e (x | πi) ; i = 1, 2, ...g variance-covariance matrix: this can be determined using the pooled variancecovariance matrix. 24 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... σ = var (x | πi) ; i = 1, 2, ...g (12) generally these parameters are determined from training set for which the population membership is already calculated. conditional density function parameters population means: i can be determined by replacing in the sample means xi variance-covariance matrix : let s i represent the sample variance-covariance matrix for the population i. thus the variance-covariance matrix σ can be determined by substituting the pooled variance-covariance into the linear score as given below: sp = ∑g i=1 (ni − 1)si∑g i=1 (ni − 1) (13) to obtain the linear score, ˆsil (x) = −1 2xisp −1xi + ui ′ σ−1x+ logpi = d̂i0 + ∑p j=1 d̂ijxj + logpi where, d̂i0 = −1 2xisp −1xi and dij = jth element of xisp −1 this is a function of the sample mean vectors, the variance-covariance matrix and prior probabilities for different populations. thus the expression looks similar to linear regression formula with a term for intercept and a linear combination of response variables with the natural log of the probabilities. thus the decision rule is to classify the sample item into the population that has the highest calculated linear score. the steps to identify the class label is as follows step 1: delete one observation from the sample. step 2: compute the discriminant function using the remaining observations. step 3: calculate the discriminant function from step 2 to identify the class label of the observation removed from sample in step 1. steps 1-3 are repeated for all the samples. calculate the misclassified observations. 3.6 quadratic discriminant analysis qda is very similar to lda where in the assumption is that the each class measurements are distributed normally. but in qda there is no such assumption that the covariance of each of the class is identical. if the normality assumption holds true, then the best possible test for a hypothesis that the given measurement from a given class is named as the likelihood test. assume that there are only 2 groups (yϵ{0, 1}) and the means of each class are specified as uy = 0 ,uy = 1 and the covariances are defined as σy=0 and σy=1 . then the likelihood ratio will be specified as, advances in systems science and application (2016) vol.16 no.2 25 likelihoodratio = exp ( −1 2 (x− uy=1) t ∑−1 y=1 (x− uy=1) )√ 2π | σy=1 | −1 exp ( −1 2 (x− uy=0) )t ∑−1 y=0 √ 2π | σy=0 ∥ −1 < t (14) for a specific threshold ’t’. the sample estimates of the mean vector and variance-covariance matrices will substitute the population quantities in the formula. 3.7 ensemble approach using wma the basic idea of majority voting is that the votes are initialized to each managers opinion. the opinion with the highest votes is selected as the final result. littlestone and warmuth [37] have specified that the number of errors can be reduced in an ensemble model by introducing weights to the majority voting technique. we utilize voting approach in this paper as given in [37]. each manager is initialized with a weight obtained from managers accuracy in classifying the sample. as each manager based on the dataset contains different number of binary classifiers bi, we need to consider each manager’s opinion for every class i separately. thus we can divide the manager’s opinion into two categories: • managers which classify given sample as an object of class i (output value 1) • managers which assert that the given observation fits to some other class than i. (output value 0 ). the voting approach is repeated for each sample x and for each binary classifier inside the manager. this results in an ensemble manager one for each class. we define a set of weight coefficients w as a 15 element vector for nsl-kdd dataset where each element j represents weight for jth manager in ensemble ie. w = (w1, w2...w15) and similarly w as a 27 element vector for unsw-nb dataset. ie. w = (w1, w2...w27). to obtain the final decision function, we consider the weights used in the voting approach. for every single observation x, we obtain fifteen output values (y1, y2...y15) for nsl-kdd dataset and twenty seven output values (y1, y2...y27) for unsw-nb dataset, one output value per manager. each value can be a positive or negative,i.e yj = {−1, 1} where value 1 correspond to managers output 1 and negative value -1 represents managers output 0. the final decision is evaluated by the equation given below y = sgn ( n∑ i=1 wiyi ) (15) 26 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... where n= 1· · · 15 for nsl-kdd dataset and n=1· · · 27 for unsw-nb dataset. each coefficient wi is multiplied with output from ith manager yi and the final decision is determined by the sign of the sum of weight coefficients for all managers. 4 proposed model the proposed model is an integrated intrusion detection system combining pca as the dimensionality reduction technique and a hybrid model of base and ensemble classifiers. in stage 1, the raw data is sent to a preprocessing unit for performing normalization by z-score technique and the noisy attributes are removed using dimensionality reduction. the resultant subset is then fed to the stage 2 which is the ensemble layer for classification. three classifiers are deployed in stage 2 namely svm. linear and quadratic discriminant classifiers. thus the ensembled approach is a merger of different supervised classifiers. a 5-fold cross validation is performed to split the data into training and testing sets. the class label is obtained by majority voting from the three classifier results. fig 1 depicts the proposed intrusion detection model. in the next subsection we discuss the approach for obtaining optimal subset using pca and ensemble approach for classification of network traffic label. algorithm for obtaining the optimal feature subset using pca: input(training set, test set) output(optimal training set, optimal test set) step 1: determine the size of training and test data step 2: scale the training and test data step 3: subtract the mean for each row m = ∑n k=1 xk n (16) wherein x specifies the individual elements and ‘n’ denotes the no. of samples. step 4: determine the covariance matrix c = xixit n (17) where x represents the matrix after subtracting the mean and xt is the transpose matrix and n is the total number of elements. step 5: determine the eigenvectors and eigenvalues of the covariance matrix. σv = λv (18) step 6: obtain a feature vector = ( eig1, eig2...eigp ) where eig1 is principal component and p ≤ n. select ’m’ such eigen vectors that match to the largest advances in systems science and application (2016) vol.16 no.2 27 ‘m’ eigenvalues in the set. algorithm for obtaining the class label using hybrid model given: classifier m1,(svm)m2,(ldc),m3(qdc), ensemble(wma) input: optimal attribute dataset d output: class label step1: initialize all the weights in d. wi = 1/n, where n is the total number of elements. step2: for every sample data di fit the svm classifier to (xt, yt) using weights wi for each class label k = 1...k obtain the hypothesis xs ← argmin (λ | f (xj) | +(1− λ)) max k(xi,xj)√ k(xi,xj)k(xi,xj) d ← d ∪ {xs} label(d) l← l ∪ {s} step 3: for every sample data di i) compute the sample estimates π̂m,ûm,σ̂ ii) make two transformations: sphere the data points based on factoring σ̂ and project to the subspace by the centroids. thus a transformation of aϵrp(k−1) iii) given any data, xϵrp transform to x̄ = axϵrk−1 x̄ = axϵrk−1 and classify according to class m = 1...k for which 1 2∥ x̃− ũm ∥2 − log π̂m is lowest where ũj = aûj step 4: for every sample data di, perform qda by repeating the process in step 3 but with different and for each class and compute the class label. step 5: the class label is predicted by weighted majority voting from results of steps (2), (3)and (4) given in equation (15) step 6: determine the performance of the model and analyze the accuracy and time complexity of the different algorithms. 4.1 class label prediction using weighted majority approach (wma) in this paper, we first deploy svm, ldc and qdc classifiers individually to result in a good generalization performance. after passing through each of the individual classifiers, a majority voting technique is adopted to predict the resultant class label from the different classifier labels. the major advantage of using ensemble approach is that the performance is improved because the approach selects only the class label which are correctly identified by all the classification techniques. the efficiency of existing approaches is compared by deploying a validation set to obtain weights for each classifier. all managers assign initial weight with value 1 and all the managers weights 28 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... are combined using wmv. the predicted class label is then compared with the target label present in the validation set. if the manager has made a mistake in identifying the class label then the weights will be subtracted by a learning factor β. learning factor is user defined and the values range between 0 and 1. this process is repeated for every sample in the dataset. after each test in which a mistake arises the sum of the weights is at most ‘u’ times the sum of the weights before the test wstart for specific u < 1. if the initial eight is wbegin and the final weight is wfin , then wbeginu f ≥ wfin must be true where ‘f’ is the number of faults. f ≤ log (wbegin −wfin) log (1/u) (19) then semble approach overcomes the difference in misclassification. the overall complexity of the model is o(n4 *m) where n is the number of attributes and m is the number of training instances. ldc performs computation in o(n), qdc performs computation in o(n2 ) and svm performs computation in o(nm) time. fig. 1 proposed intrusion detection model advances in systems science and application (2016) vol.16 no.2 29 5 experiment setup and results the experimental study was conducted with nsl-kdd and unsw-nb datasets as discussed in the section 3.2. the experiments utilized matlab-2013b installed on windows 7 ultimate 64-bit machine. a 5-fold cross validation is performed on all the classifiers for splitting the dataset into 5 non repeated subsets for training, validation and testing. in this paper we compare the efficiency of the proposed algorithms with the metrics namely accuracy rate and elapsed time. • accuracy: number of test instances correctly classified by the model. • elapsed time : time to complete the detection by a classification approach. an important advantage of intergrating complementary classifiers is to improve accuracy and generalization performance. we assume the accuracy of each classifier separately. the ensemble approach is compared by considering the average score from each classifier. we analyzed the efficiency of each classifier in the base and ensemble. thus the managers accuracy is defined by ei = ∑n i=1ai n (20) where ‘n’ is the number of classes. 5.1 study 1: nsl-kdd dataset the initial step is to perform preprocessing on the data and dimensionality reduction using pca to remove the noisy attributes in the traffic that do not contribute for classification. table 1 shows the features retrieved after dimensionality reduction on the nsl-kdd dataset. the symbolic features and the features with less variance are removed from the dataset. experimental results for each manager separately for svm, ldc, qdc and ensemble managers on the nslkdd datasets are specified in tables 2-5 respectively. the results show that by deploying different classifiers for each class, higher accuracy is achieved for all classes namely normal, dos, probe, u2r and r2l which is highlighted in each table. it is also shown that in comparison to base classifiers and ensemble managers, the latter results in high accuracy above 99% for all the class labels. time required for classification for each of the managers is presented in table 6. the time consumption for the ensemble wma is relatively higher in comparison to base classifiers but it can be neglected as the accuracy is high. figs 2-5 represent the accuracies of svm, ldc, qdc and ensemble wma managers respectively. it is evident that wma produces highest accuracy in comparison to all base classifiers though individually svm, ldc and qdc produces optimal results for a single or two class label only. 30 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... table 1 features selected by pca in nsl-kdd dataset 22 features selected after dimensionality reduction in nslkdd data set. service, dst bytes, dst host diff srv rate, flag, dst host serror rate, dst host srv count, same srv rate , dst host same srv rate, serror rate, src bytes, dst host srv diff host rate, host, dst host rerror rate duration, srv diff host rate, dst host srv rerror rate er ror rateprotocol type, srv rerror rate, is guest login, srv count, num compromised. table 2 accuracy obtained using svm manager with different rbf kernel in nsl-kdd dataset manager normal probe dos u2r r2l svm 1 (rbf:5) 92.9 98.08 74.41 90.08 87.5 svm 2 (rbf:2) 91.93 97.79 51.42 84.1 100 svm 3 (rbf: 1) 90.39 99.8 85.71 85.2 88.88 svm 4 (rbf:0.5) 89.73 97.78 71.82 100 83.3 svm5 (rbf:0.2) 99.54 97.94 56.25 99.2 89.47 table 3 accuracy obtained by linear discriminant analysis in nsl-kdd dataset manager normal dos probe u2r r2l ldc 1 98.87 99.83 82.05 96.99 99.12 ldc 2 93.22 97.79 94.12 86.77 99.65 ldc 3 96.58 98.56 100 99.45 98.56 ldc 4 99.58 99.6 86.48 100 81.81 ldc 5 70.4 93.92 71.84 55.86 97.42 table 4 accuracy obtained by quadratic discriminant analysis in nsl-kdd dataset manager normal probe dos u2r r2l qdc 1 92.898 97.94 82.75 83.81 89.83 qdc 2 91.73 98.28 86.29 86.81 93.32 qdc 3 87.17 97.55 67.92 91.71 92.21 qdc 4 90.59 97.92 94.05 87.44 92.45 qdc 5 89.83 97.8 78.98 86.02 95.31 advances in systems science and application (2016) vol.16 no.2 31 table 5 accuracy obtained by wma managers in nsl-kdd dataset manager normal probe dos u2r r2l wma 1 98.93 99.47 94.61 99.45 99.11 wma 2 78.77 94.47 94.59 82.78 99.6 wma 3 78.89 94.31 88.1 82.36 84.85 fig. 2 accuracy of svm managers in kdd dataset fig. 3 accuracy of ldc managers in nsl-kdd dataset fig. 4 accuracy of qdc managers in nsl-kdd dataset fig. 5 accuracy of wma managers in nsl-kdd dataset 32 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... table 6 elapsed time for managers in nsl-kdd managers elapsed time (secs) svm 31.96 s ldc 63.95 s qdc 66.09 s wma 53.36 s 5.2 study ii: unsw-nb dataset after preprocessing on the dataset, dimensionality reduction using pca is performed and the results are shown in table 7. experimental results for each manager separately for svm, ldc, qdc and ensemble manager on the unsw-nb datasets are specified in tables 811 respectively. the results show that by deploying unique classifiers for each class higher accuracies are obtained for all the class labels namely normal, analysis, backdoor, exploits, reconnaissance, fuzzers, exploits, dos and shellcode respectively. the highest accuracy for all the class labels are obtained using ensemble manager higher than the base classifiers as weighted majority approach reduces the misclassification rate. time required for classification for each of the managers is presented in table 12. the time consumption for the ensemble wma is relatively higher in comparison to base classifiers but as the dataset is huge the time span is relatively huge in comparison to the nsl-kdd dataset. figs 6-9 represent the graphical accuracies using svm, ldc, qdc and wma managers for the unsw-nb dataset respectively. it is inferred that svm performs poorly for attacks namely reconnaissance, dos and analysis with accuracy of 35%, 29% and 40% respectively. ldc performs poorly for attacks such as shellcode, reconnaissance, dos and fuzzers with accuracy of 40%,36%,38% and 28% respectively. qdc performs poorly for backdoor, analysis, exploits, shellcode and reconnaissance with accuracy of 34%,33%,30% and 33% respectively whereas wma results in an average accuracy for reconnaissance,exploits, shellcode and fuzzers more than 50% which is a considerable increase in comparison to base classifiers and other class labels also produce higher accuracy. table 7 attributes retrieved after dimensionality reduction using pca in unswnb dataset. 21 attributes selected after pca id, dur, service, state, spkts, dpkts,sbytes,dbytes, rate, sttl, dttl, sload, dload, sloss, dloss, sinpkt, dinpkt, sjit, djit, swin, stcpb, dtcpb, dwin, tcprtt, synack, ackdat, smean, dmean, trans depth, response body len, ct srv src, ct state ttl, ct dst ltm, ct src dport ltm, ct dst sport ltm, is-2 login, ct 2 cmd, ctflw 4 mthd, ct src ltm, ct srv dst, is sm ips ports. advances in systems science and application (2016) vol.16 no.2 33 table 8 experimental results using svm manager in unsw-nb dataset manager normal analysis backdoor reconna issance exploits fuzzers generic dos shellcode svm 1 99.94 66.21 59.14 90.28 86.76 91.35 99.61 91.35 91.26 svm 2 51.66 60.94 50.81 71.89 64.86 98.33 84.06 76.66 86.66 svm 3 69.14 93.79 45.81 48.21 56.17 99.72 86.89 56 78.87 svm 4 99.03 39.28 45.95 60.79 87.76 53.05 98.4 70.2 27.83 svm 5 93.76 43.33 65.51 25.97 47.81 57.88 72.64 58.62 43.13 svm 6 82.34 34.76 62 30.39 52.22 63.57 87.37 80.76 73.22 svm 7 99.55 31.33 47.36 33.55 75.49 42.67 99.26 47.42 42 svm 8 96.19 32.15 47.07 33.4 52.54 33.84 98.83 28.95 34.18 svm 9 99.65 79.63 63.73 70.78 87.38 96.28 99.65 93.97 87.51 table 9 experimental results using ldc manager in unsw-nb dataset manager normal analysis backdoor reconnaissance exploits fuzzers generic dos shellcode ldc 1 87.91 53.49 54.15 75.07 67.88 63.91 99.09 71.95 39.61 ldc 2 86.41 50.22 52.95 36.07 84.34 55.66 94.71 54.16 84.94 ldc 3 98.67 75.14 62.59 36.84 65.45 56.08 92.36 38.67 50 ldc 4 97.59 80.11 92.3 32.98 65.79 42.87 99.26 46.25 41.37 ldc 5 96.59 89.2 90.44 50.89 84.26 47.55 63.43 35.26 36.74 ldc 6 94.22 76.56 85.66 36 75.44 94.59 80.82 48.21 66.99 ldc 7 87.56 44.7 59.53 37.39 55.4 28.41 99.57 38.39 43.27 table 10 experimental results using qdc manager in unsw-nb dataset manager normal analysis backdoor reconnaissance exploits fuzzers generic dos shellcode qdc 1 82.11 46.41 77.41 26.38 48.33 50 65.21 67.1 67.16 qdc 2 97.86 28.67 25.03 35.8 60.34 57.06 80.68 75.49 58.76 qdc 3 86.85 32.98 34.03 33.42 83.7 60.34 59.06 85.68 50 qdc 4 91.73 43.43 52.76 36.34 46.27 51.37 94.03 77.36 30.17 qdc 5 84.21 83.33 63.37 42.19 51.76 45.83 80.94 74.15 42.23 qdc 6 81.24 35.82 81.51 24.4 52.8 53.95 87.59 81.92 41.95 qdc 7 88.39 40.52 78.94 49.49 38.58 54.98 87.58 68.85 64.19 qdc 8 87.48 44.45 58.95 35.43 54.54 66.31 81.54 91.66 50.63 qdc 9 88.25 46.7 63.85 36.72 62.33 62.83 96.43 82.48 53.56 table 11 experimental results using wma manager in unsw-nb dataset manager normal analysis backdoor reconnaissance exploits fuzzers generic dos shellcode wma1 99.99 78.41 86.71 65.97 76.66 91.35 91.85 58.94 86.66 wma1 98.51 51.95 64.71 71.89 68.84 98.33 90.88 50.46 91.26 wma 3 88.93 52 62.09 48.21 71.14 99.72 81.45 90.75 78.87 wma 4 99.41 50.8 100 60.79 73.97 53.05 97.65 ‘100 57.83 wma 5 93.2 67.32 84.61 90.28 59.9 57.88 90.19 67.67 53.13 wma 6 54 100 50 60.39 46.84 63.57 84.32 68.8 73.22 wma 7 95.25 55.89 57.09 63.55 62.67 42.67 99.15 58.8 62 wma 8 81.71 74.72 50.22 63.4 88.21 63.84 91.4 69.34 64.18 wma 9 99.21 58.67 65.62 70.78 87.92 96.28 84.13 64.96 87.51 table 12 elapsed time for managers in unsw-nb data set managers elapsed time (secs) svm 464.82 s ldc 485.56 s qdc 622.50 s wma 751.61 s 34 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... fig. 6 proposed intrusion detection model fig. 7 proposed intrusion detection model fig. 8 proposed intrusion detection model advances in systems science and application (2016) vol.16 no.2 35 fig. 9 proposed intrusion detection model 5.3 discussion in this paper, we have analyzed the accuracy of the base and ensemble classifier managers. it is to be noted that the ensemble managers accuracy is relatively high in comparison to the base classifiers and managers accuracies are above 99% for nsl-kdd dataset and the range is from 88-100% for unsw-nb dataset. the variation is due to the imbalance of samples in the dataset for the classes exploits, shellcode and reconnaissance. the reason for the high accuracy in both the datasets are 1) proper selection of training data 2) well-selected parameters in wma technique namely the learning factor and weight assignment wi for each sample. 3) misclassification error is reduced due to weight assignment to each sample. there were poor results obtained for ldc and qdc manager only in the unsw-nb dataset. but svm performs well for nsl-kdd dataset and in specific for class labels normal, fuzzers, generic and dos in the unsw-nb dataset. svm produced poor accuracy of 85% for dos and qdc produced poor accuracy of 91% and 95% for u2r and r2l attacks in the nsl-kdd dataset. the accuracy was as low as 49.49% for reconnaissance class label because the test data contains samples that were not utilized while training the classifiers. thus in comparison to wma, the base classifiers svm, ldc and qdc perform poorly and also with increase in elapsed time. the constraint on the weights generated by wma also is a major factor in improving the accuracy of the intrusion detection model. the weights can lie between 0 and 1, the higher the value the more is the possibility of correctly predicted. thus this model can be extended to include other ensemble techniques but it should consider the time complexity as many ensemble approaches complete the task in extremely long time. 36 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... 6 conclusions the classification accuracy can be improved with minimal elapsed time by integrating opinions from multiple managers into single using an ensemble approach. we have deployed weighted majority voting to integrate results from different managers. the three base classifiers namely svm, ldc and qdc were experimentally compared using two different datasets namely nsl-kdd and unswnb. thus wma results in good accuracy for both the datasets. the accuracy improvement is of 1% in comparison to base classifiers. thus the success of the model is due to the generated weights in wma which were tuned by the learning factor. thus the integration of base classifiers in to an ensemble manager with wma will be very well utilized in intrusion detection as it has been tested on recent datasets with modern day attacks. the other improvement in our approach is instead of relying on binary classification techniques, we have utilized multi class classification for the base managers. in future we can integrate optimization techniques with wma to tune the parameters for generating weights. references [1] khan l., awad m. and thruaisingham b. (2007), “a new intrusion detection system using support vector machines and hierarchical clustering”, the vldb journal, vol. 16, pp. 507-521. [2] wang g., hao j., ma j. and huang l. (2010), “a new approach to intrusion detection using artificial neural networks and fuzzy clustering, expert systems with applications”, vol. 37, pp. 6225-6232. [3] sandhya peddabachigiri, ajith abraham, crina grosan and johnson thomas. (2007), “modeling intrusion detection system using hybrid intelligent systems”, journal of network and computer applications, vol. 30, pp. 114-132. [4] lee w, stolfo s and mok k. (1999), “a data mining framework for building intrusion detection model”, in: proc. of ieee symposium on security and privacy, pp. 120-32. [5] a. kausar, m. ishtiaq, m.a.jaffar and a.m.mirza. (2010), “optimization of ensemble based decision using pso”, in: proceedings of the world congress on engineering, vol. 10. [6] tsang c.h., kwong s. and wang h. (2007), “genetic fuzzy rule mining approach and evaluation of feature selection techniques for anomaly intrusion detection”, pattern recognition, vol. 40, pp. 2373-2391. advances in systems science and application (2016) vol.16 no.2 37 [7] li y. and guo l. (2007), “an active learning based tcm-knn algorithm for supervised network intrusion detection”, computer and security, vol. 26, pp. 459-467. [8] amor n.b., benferhat s., elouedi z.and zhang x.t. (2004). näıve bayes vs decision trees in intrusion detection systems, sac04:proceedings of the 2004 acm symposium on applied computing, new york,ny,usa: acm press. [9] xiang c., yong p.c. and meng l.s. (2008),“design of multiple-level hybrid classifier for intrusion detection system using bayesian clustering and decision trees” , pattern recognition letters, vol. 29, no. 7, pp. 918-924. [10] shafi k. and abbass h.a. (2009), “an adaptive genetic based signature learning system for intrusion detection”, expert systems with applications, vol. 36, no. 10, pp. 12036-12043. [11] sumaiya thaseen and ch.aswani kumar. (2014), “ intrusion detection model using fusion of pca and optimized svm”, 2014 international conference on computing and informatics (ic3i) 27-29 nov 2014, pp. 879-884. [12] i.sumaiya thaseen and ch.aswani kumar. (2016), “improving accuracy of intrusion detection model using pca and optimized svm”, cit journal of computing and information technology, vol.24, no.2 ,pp. 133-148 [13] ajith abraham, crina grosan and carlos martin-vide. (2007), “evolutionary design of intrusion detection programs”, international journal of network security, vol. 4, no. 3, pp. 328-339. [14] fangjun kuanga, weihong xua and siyang zhang.(2014). “a novel hybrid kpca and svm with ga model for intrusion detection”, in applied soft computing , vol. 18, pp. 178-184. [15] srilatha chebrolu, ajith abraham and johnson p.thomas. (2005), “ feature deduction and ensemble design of intrusion detection systems”, computers and security, vol. 24, pp. 295-307. [16] alexandre balon-perin. (2012), “ensemble-based methods for intrusion detection”, norwegian university of science and technology. [17] chandra a. and yao x.(2006), “evolving hybrid ensembles of learning machines for better generalization”, neurocomputing, vol. 69, no. 7, pp. 686700. 38 sumaiya thaseen and ch.aswani kumar: intrusion detection model using pca and ... [18] srinivas mukkamala, andrew h. sung, ajith abraham. (2007). “intrusion detection using and ensemble of intelligent paradigms”, journal of network and computer applications, vol. 28, no. 2, pp. 167-182. [19] i.syarif, e.zaluska, a.prugel-bennett and g.wills. (2012), “application of bagging, boosting and stacking to intrusion detection”, in: machine learning and data mining in pattern recognition, pp. 593-602. [20] e. bahri, n.harbi and h.n.huu. (2011), “approach based ensemble methods for better and faster intrusion detection”, in: computational intelligence in security for information systems, pp. 17-24. [21] v.bukhtoyarov and v.zhukov. (2014), “ensemble-distributed approach in classification problem solution for intrusion detection systems”, in: intelligent data engineering and automated learning-ideal 2014, pp. 662-668. [22] cordeiro junior, z. and g.l.pappa. (2011), “a pso algorithm for improving multi-view classification”, in: 2011 ieee congress on evolutionary computation ( cec ), pp. 925-932. [23] a.kausar, m.ishtiaq, m.a.jaffar and a.m.mirza. (2010), “optimization of ensemble based decision using pco”, in proceedings of the world congress on engineering, wce, vol. 10. [24] http://nsl.cs.unb.ca/nsl-kdd/ [25] s.j.horng, m.y.su, y.h.chen, t.w.kao, r.j.chen, j.l.lai and c.d.perkasa. (2011), “a novel intrusion detection system based on hierarchical clustering and support vector machines”, expert systems and applications, vol. 38, no. 1, pp. 306-313. corresponding author sumaiya thaseen can be contacted at: sumaiyathaseen@gmail.com advances in systems science and applications (2012) vol.12 no.4 327-346 what are natural systems, actually? alan rayner bath bio art, 81 dovers park, bathford, bath ba1 7ud, uk abstract to this day, the applicability of abstract, definitive logic and mathematics to natural systems is rarely challenged or even questioned. consequently we find ourselves predominantly living, working and researching in a way that contradicts how we naturally are in the world as it naturally is. this seems unwise, to put it mildly. in such circumstances can it be any surprise when we are drawn into needless conflict and misunderstanding, unable to work out what it means to live in an ecologically sustainable way, and prone to inflict profound psychological, social and environmental harm on ourselves and our natural neighbourhood? a way out from this predicament is offered by what has been called ‘natural inclusionality’, which has the potential radically to transform human understanding of natural systems and our place within them. keywords boundary perceptions, evolutionary process, flow-geometry, intangibility, natural inclusion, place-time, receptivity, space perceptions 1 introduction: abstract and natural perceptions of space and boundaries by imposing non-existent rigid structure onto naturally continuous space and dynamic boundaries, and enshrining this within three dimensional planes set at right-angles to one another, abstract mathematical logic engenders profound paradox[1-3]. natural space cannot be confined within a box-frame, nor can it be cut into discrete segments that can be moved around independently and/or relative to one another. natural space is infinite at all scales i.e. dynamically distinguishable into different localities, but not divisible into separable, discretely packaged units or quantities. by the same token, natural boundaries cannot cut, they can only dynamically configure space. to cut space, as abstract models require, they would have to be reduced to zero thickness, and hence be nowhere. this paradox of treating space and boundaries as rigidly definable fabric or structure is laid bare by the following statements: “when a smaller box s is situated, relatively at rest, inside the hollow space of a larger box s, then the hollow space of s is a part of the hollow space of s, and the same “space”, which contains both of them, belongs to each of the boxes. when s is in motion with respect to s, however, the concept is less simple. one is then inclined to think that s encloses always the same space, but a variable part of the space s. it then becomes necessary to apportion to each box its particular space, not thought of as bounded, and to assume that these two spaces are in motion with respect to each other”[4]. 328 alan rayner:what are natural systems, actually “space is another framework we impose upon the world . . . here the mind may affirm because it lays down its own laws; but let us clearly understand that while these laws are imposed on our science, which otherwise could not exist, they are not imposed on nature. . . .euclidian geometry is . . . the simplest, . . . just as the polynomial of the first degree is simpler than a polynomial of the second degree. . . . the space revealed to us by our senses is absolutely different from the space of geometry” poincaré[5]. a moment’s contemplation reveals how utterly unrepresentative of the space and boundaries of our natural experience this treatment is. all we have to do is ask ‘what needs to be present for natural form to be distinguishable?’ or, more concretely, ‘why is our direct experience of walking into a brick wall different from that of walking through an open doorway?’ or, ‘what makes it possible to paint a picture?’ it then becomes apparent that the only way of answering these questions is to acknowledge the occurrence of at least two kinds of natural presence: a receptive context or medium which provides freedom for local movement and/or expression, and local formative content, which informs or configures that context. the former is necessarily spacious, the latter necessarily cohesive. moreover, for form to be and become distinguishable, each of these presences must naturally include the other. spacious presence alone would be formless void, and formative presence alone would have no shape or size. they are necessarily distinct, but mutually inclusive presences. they can neither be abstracted from one another as independent entities, nor be homogenised into ‘oneness’. the only way in which this necessity can be fulfilled is for one of these presences, natural space, ultimately to be everywhere, continuous, intangible (i.e. frictionless) and immobile, and for the other ultimately to be somewhere, distinctive, tangible and continually in motion. natural space and figural boundaries are hence, respectively, continuous and dynamically distinct energetic interfacings between the insides and outsides of all natural forms as flow-forms[1-3]. in summary, natural systems are hence actually radically different in their dynamic, evolutionary organization from the self-contained abstract scientific models that are widely used to represent and simulate their behaviour, both conceptually and practically. this difference arises from the preand/or postimposition of rigidly definitive structure in abstract models onto the spatial and dynamic continuity of natural flow-geometry. such imposition is embedded in the foundations of propositional and dialectic logics and both classical and modern mathematics[1-3]. whereas it has predictive utility in unchanging or repetitive systems, it cannot be expected to account adequately for the dynamics of evolutionary systems, where it has the potential to give rise to serious and damaging misunderstanding and miscalculation. in the latter systems a radically different, more natural approach is needed in which conventional discontinuous treatments advances in systems science and applications (2012) vol.12 no.4 329 of space and boundaries are replaced by continuous ones. this more natural approach is available through the fluid boundary logic of what has been called ‘natural inclusionality’[2-3]. here, i discuss the benefits of ‘naturalizing’ systems science using this approach to understand all forms as flow-forms energetic configurations of space in space, displaying varying degrees of deformability, permeability and connectivity depending on context. 2 natural inclusionality conceptually, natural inclusionality is simply a way of understanding all natural evolutionary form or organization as flow-form, an energetic configuration of space in space. implicit within this simple description are, however, two radical innovations in thought: (1). the recognition that natural boundaries are intrinsically energetic ‘dynamic interfacings’ between distinct localities, not the ‘inert limits’ of discrete objects[1,6]. (2). the recognition that natural boundaries can only be dynamic through the inclusion of space as infinite, intangible, frictionless presence[1-3]. correspondingly, the dynamic origin of all natural form as an energetic configuration of space can be understood in terms of an evolutionary process of natural inclusion: the co-creative, fluid-dynamic transformation of all through all in receptive spatial context[2-3,7]. this understanding differs radically from the conception of the evolutionary origin of biological species by ‘natural selection, or the preservation of favoured races in the struggle for life’[8]. it similarly challenges any notion of cosmological origins and expansion from an isolated locality. such notions of ‘something out of nowhere’ are an abstract product of theoretically dividing or unifying space and form[9]. natural inclusionality therefore calls for what amounts to a paradigmatic transformation a radical de-framing and reframing of abstract perceptions of space and boundaries that have imposed non-existent definition onto naturally continuous systems for millennia. such a transformation has profound implications for every aspect of our philosophical understanding of and relationship with nature, including human nature: logical, mathematical, scientific, artistic, theological, linguistic, educational, social, psychological and political. the need for such a transformation was recognised by polanyi[10] when he stated that: “for once men have been made to realize the crippling mutilations imposed by an objectivist framework once the veil of ambiguities covering up these mutilations has been definitely dissolved many fresh minds will turn to the task of reinterpreting the world as it is, and as it then once more will be seen to be”. to summarize in general terms, natural inclusionality is a kind of awareness that helps us to appreciate our selves and other tangible forms as dynamic in330 alan rayner:what are natural systems, actually habitants of nature, not discrete subjects and objects rigidly set apart from one another. this awareness comes with recognizing that natural space is a limitless intangible presence everywhere, which permeates throughout and beyond all tangible expressions of energy, whether in the form of radiation or massy bodies. natural space cannot be cut and can neither resist nor be resisted by nor be removed from the presence and movement of tangible forms. far from being just empty distance between, outside or occupied by discrete material objects or structures as is assumed by abstract logic natural space is a receptive presence, vital for movement and communication. as natural dynamic inclusions of space, all forms are variably fluid flow-forms. their boundaries are energetic configurations of space, not exclusions from space. when they move, they do not move through space; instead space permeates through them. with this awareness comes an appreciation of self-identity as an inclusion of neighbourhood a fluid inclusion, not a rigid exclusion of others’ identities. our understanding of physical reality is such as to bring profound compassion for ourselves and other life forms, and is a source of deep inspiration and creativity. it calls for an expansion of conventional theoretical reasoning to include more fluid, artistic and poetic forms of expression. in more technical and philosophical terms, natural inclusionality is a new philosophy and fluid boundary logic of self-identity and ecological and evolutionary diversity and sustainability. it is intended to supersede the abstract rationality that has dominated human thought for millennia, based on definitive logic that can only apply to inert material systems that are unknown to exist anywhere in nature. whereas abstract rationality treats space as empty distance between, occupied by or outside completely definable tangible material structures or objects with discrete boundary limits, natural inclusionality recognizes space as a limitless, indivisible, receptive (non-resistive) ‘intangible presence’ vital for movement and communication. this allows all form to be understood as flow-form, distinctive but dynamically continuous, not singularly discrete. the simple move from regarding intangible space and tangible boundaries as mutually exclusive sources of discontinuity and discrete definition to mutually inclusive sources of continuity and dynamic distinction enables self-identity to be understood as a dynamic inclusion of neighborhood. intangible space is included throughout and beyond all tangible figural forms as configurations of energy, whether as massy bodies or mass-less electromagnetic radiation. 3 the relationship between natural inclusionality and ecological sustainability the natural inclusional perception of living systems as flow-forms that receive, retain and pass on energy in the process of growing, living and dying is inconsistent advances in systems science and applications (2012) vol.12 no.4 331 with any model of them as independent entities. to be entirely self-contained is to be an inert, hermetically closed structure with no capacity for take up or loss of energy between inner world and outer world. the nearest any life forms actually get to this condition is when they form survival capsules such as spores, seeds, pupae and cysts that carry them through periods of scarcity. this, by contrast with the darwinian perception of survival by competitive exclusion of others, is what real biological ‘survival’ or ‘preservation’ entails. in such a dormant condition they are incapable of any active growth or relationship with others. but no sooner is any activity resumed that can support growth, so too is any life form’s capacity to lose as well as take up energy through its necessarily permeable bodily boundaries and those of others in its vicinity. it is therefore clear that the availability of sources of energy is the principal influence that governs the growth, organization and function of all natural forms of organic life as variably open systems. any activity or pattern of development in which energy loss through permeable boundaries persistently exceeds energy acquisition will result in unsustainable deficit. on the other hand, any pattern of development that permanently prevents energy loss also prevents energy gain. for any living system to sustain itself, its primary need is therefore to be able to attune its activities and development to correspond with energy availability and hence with the local conditions of its habitat. this availability varies, both in amount and rate of supply, due to seasonal and climatic fluctuations, and where and in what form it is located. it also changes due to the growth, death and decomposition of the systems themselves, which respectively deplete and replenish supplies as they come under one another’s simultaneous mutual influence. real life does not, therefore, inhabit an even playing field of energy, space and time. instead it continually both changes and responds to changes in the contextual circumstances of its natural neighbourhood in an improvisational process of autocatalytic flow, which gives rise to evolutionary and ecological complexity and succession[1,6]. through this process of ‘natural inclusion’ an opening is made dynamically for an extraordinary diversity and complexity of interdependent forms and patterns of life to co-evolve over myriad nested temporal and spatial scales. the breathtaking variety that we can find in a crumb of soil, a patch of chalk grassland, a coral reef and a tropical forest comes into being under the guidance of no more and no less than the responses and contributions of its membership to natural energy flow in a natural ‘sustainability of the fitting’[11-13]. fig.1 illustrates the general principles arising from observations of how living systems (except modern human cultures) attune their patterns of growth and development to variable availabilities of energy sources. as natural inclusional energetic inner-outer interfacings of continuous space, the boundaries of real organisms, populations and communities do not remain constant throughout their 332 alan rayner:what are natural systems, actually life span, but fluidly vary in permeability, deformability and contiguity (connectivity)[6,14]. they change in dynamic relationship with the availability of energy predominantly assimilated from sunlight into organic compounds via the process of photosynthesis, and rendered into chemical form (adenosine triphosphate) via the oxidative-reductive reactions of respiration as a form of combustion. moreover, these changes themselves entail alterations in boundary chemistry induced by and involving shifts in availability and production of oxidizing and reducing power[6,15]. fig.1 the interplay between boundary-proliferating (‘differentiation’) and boundary-condensing (‘integration’) processes in energy-rich (stippled) and energy-restricted circumstances. this interplay enables energy to be assimilated (allowing regeneration and proliferation of boundaries), conserved (by conversion of boundaries into relatively impermeable form), explored for (through internal distribution of energy) and recycled (via redistribution/reconfiguration of boundaries) in spatial capsules, channels, branches and networks of life forms in dynamic attunement with their natural neighbourhood. thin lines indicate relatively more permeable boundaries, thick lines relatively impermeable boundaries and dotted lines degenerating boundaries[6]. the ecological and evolutionary sustainability of natural life forms, from the cells and tissues in a human body to the trees in a forest correspondingly depend upon close harmonization with (as distinct from unilateral adaptation to) advances in systems science and applications (2012) vol.12 no.4 333 fig.2 ‘fungal foraging’[6,16]. the mycelium of the wood-decaying fungus, hypholoma fasciculare, finds an ‘oasis in a desert’, by fluid-dynamically spreading and narrowing its energetic focus. the fungus has been inoculated into a tray full of soil on a block of wood (‘starter’ food source), with an uncolonized wood block (‘bait’ food source) placed some distance away from it. distinct stages are shown in the radial spreading of the fungal colony from the inoculated wood block, followed by the redistribution and directional focusing of its energy following upon contact with the bait. as indicated in fig.2, similar fluid dynamic patterns of gathering in, conservation of, exploration for and redistribution of energy supplies within variably connective channels and capsules of receptive space are found throughout the living world, from subcellular to ecosystem scales of organization. 334 alan rayner:what are natural systems, actually the diversity, complementary nature and changeability of all within their neighbourhood, to which they themselves contribute. when energy supplies become scarce, sustainable living systems pool and redistribute internal resources within integrated structures and survival capsules they do not compete to proliferate faster on the dwindling supplies than their neighbours. when supplies are abundant they proliferate and differentiate. moreover, as is beautifully illustrated by the exploratory patterns of some kinds of fungi, this ability to attune their capacity to differentiate and integrate activity in dynamic relationship with energy availability allows life forms to locate and sustain supplies in heterogeneous habitats with extraordinary efficiency. as illustrated in fig.2, they do this through a combination of all round exploration and directional focus. sustainability, not supremacy, is therefore the path of evolutionary and ecological continuity. natural energy flow is variably fluid, circulatory and redistributive along pressure gradients from higher concentration (relative ‘abundance’) to lower concentration (relative ‘scarcity’), as illustrated, for example by atmospheric and ocean currents. the primary need for all life forms is not to seek competitive advantage through the unilateral accumulation of energy ‘wealth’ at the expense of their neighbourhood, but to sustain themselves and their offspring as variable channels for natural energy flow. they are more like members of a relay team than a set of autonomous individuals striving to be first past the post. to succeed in this they have to be open to the energetic influence of their neighbourhood at the same time as sustaining the distinctiveness but not discreteness (or separateness) of their inner worlds from their outer worlds through their dynamic boundaries. any ecological or evolutionary or management model that treats an individual or group as a discrete, autonomous object or subject with the set objective of promulgating and preserving its self at all costs as sole survivor of a war of attrition is therefore partial and unsustainable in a changeable world of natural energy flow. unfortunately, just such models are implicit in the objectivistic framing of natural energy flow that continues to underpin our strategic planning for a desirable future and perceptions of what it means to be sustainable. we confuse sustainability with self-preservation, just as darwin[8] did when describing ‘natural selection’ as the preservation of favoured races in the struggle for life. why? 4 unsustainable logic: self-dislocation from natural neighbourhood notions of adversarial ‘competition’ and coercive ‘co-operation’, which respectively underlie individualistic ‘capitalism’ and collectivistic ‘socialism’, are predicated upon definitive logic that is incompatible with the cumulative energetic transformation of an evolving system[2]. it is presupposed that individual or group advances in systems science and applications (2012) vol.12 no.4 335 entities can be defined independently from their spatial context and correspondingly that their ‘future’ can be fully defined by present or ‘initial conditions’. as recognized by bateson[17], this narrows the focus of perception and purpose at the outset of enquiry into nature instead of in the process of discovery (cf. fig.2) and can give rise to the familiar idea that undesirable present ‘means’ can be justified by desirable future ‘ends’. human beings may be cognitively and culturally predisposed to make this presupposition through a combination of our inter-related capacities for categorization, sociality, abstract thought, tool and language use and awareness of mortality[12-14,18]. on the other hand, the imagination that comes alongside these capacities offers the creative potential to escape the restrictions imposed by purposive abstract objectivity through what is actually the more comprehensive worldview of natural inclusionality[2-3,12]. as terrestrial, omnivorous, bipedal primates unable to digest cellulose but equipped with binocular vision and opposable thumbs that enable us to catch and grasp, we are predisposed to view the geometry of our natural neighborhood in an overly definitive way. we are prone to see the world in terms of what it can do for us and to us as detached observers or abstracted ‘exhabitants’, not how we are inextricably involved in it as natural inhabitants. we perceive ‘boundaries’ as the limits of definable ‘objects’ and ‘space’ as ‘nothing’ a gap or absence outside and between these objects[1]. as discussed earlier, this perception of space and boundaries as definitively discontinuous is incompatible with the comprehension of continuity and change[23,19]. if two adjacent locations in space and/or time are distinguished by a boundary, which one does the boundary belong to? if it belongs to both of them, how can the mutual exclusivity of definitive logic be satisfied, and where do both cease to be both and become either one or the other? if it belongs to neither, then where does one location end and the other begin and what really comes between them? in the case of a curved boundary, does it belong to whatever lies within it or to whatever lies without it? if two distinct locations are both contained within a larger location, are they mutually exclusive or coexistent? upon such dilemmas rests the whole gamut of alternative propositional (either/or) and dialectical/transcendental logics (both/and in mutual opposition) that have been in conflict for millennia and continue to be so[20]. so too do the ‘holons’ as ‘janus-faced’ entities combining individual and collective aspects, and ‘holarchies’ as nested arrays of holons, of koestler[21] in his ‘open hierarchical systems theory’[22-23]. that it is nonetheless possible to avoid this perception is, however, evident from the indigenous cultures that sustain a much stronger sense of inclusion in nature, aided by the preservation of oral, aural and nomadic traditions [24-25]. 336 alan rayner:what are natural systems, actually according to walker[26], “cross-cultural views of the self define individuality in terms of boundaries, locus of control and inclusiveness versus exclusiveness, or that which is intrinsic versus that which is extrinsic to the self [27-28]. cultures that emphasize firm boundaries and high personal control tend to view the self as exclusionary or ‘self contained’. fluid boundary, strong field control cultures, view the self as “ensembled,” meaning that the self is inclusive of other individuals. while ‘self contained’ individualism is indigenous to the united states and to the european countries from which its dominant ethnic groups draw their roots, ‘ensembled’ individualism is far more prevalent as a percentage of all known cultures[29]. ensembled individualism is also indigenous to aboriginal, native american, senoi and other cultures that are widely known to use dreams for social purposes.” the perception of completely definable objects separated by intervals of space as ‘gaps of nothingness’ sets the scene for the hard line logic of abstract rationality to become established in the foundations of our mathematical, scientific, theological, linguistic, governmental and economic endeavors. it also profoundly affects our perceptions of ‘self’ and ‘self-interest’. the definitive supposition that one thing is not another thing, and, specifically, that ‘one self cannot be another self’ leads to what c.s. lewis[30] called ‘the philosophy of hell’, in which ‘to be means to be in competition’. it is easy to see that this detached perception of nature and human nature in unnatural opposition could lead to profound human conflict and jealous possessiveness. with the continuous presence of space throughout and beyond all form erased from consideration, ‘subjective self’ and ‘objective other’ are brought into fear-full confrontation. priorities are inverted from seeking sustainable relationship with others in a natural ‘communion of diversity’, to seeking cancerous dominion over other as the only certain route to ‘self-preservation’[25]. sustaining ‘ego’ becomes the focus of attention at the expense of the natural neighbourhood upon which individual self-identity actually depends to sustain itself. love and trust of others break down into xenophobia and avarice. can this abstraction actually be intellectually justified as a means of representation consistent with sensory experience (i.e. evidence) and that makes consistent sense? in a word, no, it cannot, because energy/matter cannot physically be cut away from space[2,18,31-32]. nonetheless, the dissociation of matter from space is embedded in the numerical and geometrical foundations of classical and modern mathematics. here it may be recalled that euclidean geometry is the abstract geometry of zerodimensional (size-less) numerical points, one-dimensional (breadth-less) lines, two-dimensional (depthless) planes and three-dimensional solids (self-contained volumes). its figures are used to represent definitive tangible structure and yet advances in systems science and applications (2012) vol.12 no.4 337 can only actually represent the intangible presence in the core of tangible form because it is impossible to reach zero without removing the tangible presence. the same applies to the so-called ‘non-euclidean’, riemannian and lobachevskian geometries of curved surfaces. the scientifically inconvenient truth is hence that abstract euclidian and noneuclidean points, lines and planes/curved surfaces can consist only of intangible presence, not tangible presence! by the same token, it is impossible to drive or rotate a solid body from or around a solid fixed centre. the central ‘still’ point, axis or plane of symmetry of any bodily form can only consist of intangible presence, with correspondingly zero pressure. in effect, conventional mathematics and its discontinuous underpinning logic thereby treat ‘1’, as a ‘unit of tangible presence’, as if it is ‘0’, a vanishing point of intangible presence. they literally attempt to construct ‘one thing from nothing’ and then to sum an infinite number of these one things up into an infinite ‘whole’ as a ‘one’ that is also ‘many’, whilst discounting the very presence that truly is infinite, at all scales. this difficulty can only be resolved realistically by accepting that in nature, tangible and intangible presences are distinct but mutually inclusive. this is the point recognized by the fluid geometry of natural inclusionality. here, space and boundaries are regarded as mutually inclusive sources of continuity and dynamic distinction with variable connectivity, not mutually exclusive sources of discontinuity and discrete definition, as in euclidean and non-euclidean geometries. so far, the only mathematical formulation explicitly to accept and incorporate this natural inclusion of non-local space in and throughout local figural form is the ‘transfigural mathematics’ introduced in 1985 by lere shakunle[32-34]. natural inclusionality effectively transforms the fixed frameworks of euclidean and non-euclidean geometries into fluid framings of omnipresent, non-local intangible space everywhere, within (intra-), throughout (trans-), between (inter-) and beyond (extra-) local tangible energetic form[32]. this opens the possibility of a dynamic, co-creative, mutually inclusive relationship between internally and externally situated non-resistive (and hence receptive) intangible spatial presence and locally situated, tangible energetic presence. 5 variable connectivity: the sustainable self-cultivation of life all that may therefore be needed to unlock our imagination and the world of real, live organisms and communities from the unnatural confinement imposed by abstract rationality is the simple understanding that space cannot be cut, occupied, confined or excluded. space is a continuous presence throughout and beyond the boundaries of natural figures. by the same token, these boundaries are energetic interfacings between inner and outer realms, not fixed limits. this simple move from regarding space and boundaries as sources of discontinuity and 338 alan rayner:what are natural systems, actually discrete definition to sources of continuity and dynamic distinction is the ecological and evolutionary point of departure of ‘natural inclusionality’ from objective rationality. the underlying logic of natural inclusionality can be described as ‘the understanding of all form as flow-form, an energetic configuration of space throughout figure and figure in space’, such that space, as a receptive (non-resistive) presence, is not assumed to be discontinuous (i.e. to stop at discrete boundary limits)[12,26]. correspondingly, we can recognize the impossibility of defining or measuring anything in absolute numerical terms anywhere, because all form has both a ‘figural’, energetic inner-outer interfacing or dynamic boundary, which makes it distinct, and a ‘transfigural’ (this term was first conceived by lere shakunle in 1985) ‘through the figure’ spatial reach that cannot be sliced or limited. the continuous space throughout and beyond the figure pools it within the co-creative, influential neighbourhood of all others: local ‘self’ as an ‘including middle’ finds identity in its non-local neighbourhood as neighbourhood finds identity through its local ‘self’. without spatial continuity, figures are rendered into lifeless bodies, integral or fractional numbers and idealized geometric points, lines and solids. with space included, we can escape the confinement and inconsistencies of the ‘excluded middle’, discrete boundary logic of ‘one opposed to other’ that has held human imagination to ransom for millennia. this enables us to move on to a more natural and comprehensive form of reasoning in the fluid boundary logic of each in the other’s mutual influence. the real meanings of ‘zero’ and ‘infinity’ as qualities of space and sources of creativity, not abstract quantities of material, are brought into our natural accounting systems, not excluded by abstract definition. the following simple exercise might help illustrate the difference between the hard-line, space-cutting view of discontinuous models and fluid-line understanding of natural inclusionality. draw an outline of two figures using a dotted line on a plain sheet of paper. the ‘paper’ infinitely stretched would represent what in the transfigural geometry developed by lere shakunle is called ‘omni-space’[32,34]. the space within each figure represents ‘intra-space’, the space between figures ‘inter-space’, the space beyond the figures ‘extra-space’ and the space transcending the figures’ permeable and dynamic boundaries ‘trans-space’. you can see how the continuous non-local space everywhere (omni-space’) is locally configured into distinctive, but not discrete regions. in the way that you have drawn them, the figures are not contiguous (connected), and so their ‘intra-spaces’ can only communicate through the ‘inter-space’ and ‘trans-space’ between and permeating their boundaries as energetic interfacings and restraining influences (not restrictive material definitions or external forces see later). nonetheless, they advances in systems science and applications (2012) vol.12 no.4 339 inhabit the same limitless pool of omni-space everywhere. if you were now to draw the figures closer together, so that their boundaries first connect and then coalesce at one or more points, their intra-space now becomes continuous (cf. fig.3). on the other hand, if you were to take a pair of scissors and cut around the dotted lines, the figures will drop out of their spatial context as discontinuous individual entities. this ‘dropping out’ of context is what discontinuous models of reality effectively do they treat boundaries as cut-out zones between discrete inner realms and outer realms, instead of dynamic relational interfacings through which these realms remain continuous through trans-space. fig.3 distinct but not discrete figures of space in space (redrawn by philip tattersall from original pencil sketch by alan rayner, 2010). fig.3 illustrates the dynamic relationships between figural flow-forms as energetic configurations of space throughout figure and figure in space. it also serves to distinguish the natural inclusional dynamic relationship between distinct but not discrete flow-forms both from reductive schemas that cut off inner from outer spatial realms and from connective and holistic schemas where individual dynamic locality is eschewed from a seamless, purely figural whole or ‘unity’. since the cartoons can only represent an instantaneous ‘slice’ through the figures, the dotted lines shouldn’t be taken to represent ‘sieves’ but more the seething ‘fluid mosaic’ that constitutes real biological membranes. a very simple example of what is represented in the cartoon can also be seen between surface-tense droplets 340 alan rayner:what are natural systems, actually of water condensing on a surface. as they expand and come into proximity their tensely curved inner-outer interfacings first touch and then coalesce in a visible rush as each flows reciprocally into the other and the tension of their boundaries is released. fig.4 stages (from top left clockwise) in fusion between the protoplasm-filled cellular tubes (hyphae) within the mycelium of the basidiomycete fungus, phanerochaete velutina. the tubes are internally partitioned into distinct compartments by septa, which have a door-like pore in their middle. as fusion occurs (third picture in the sequence) the cell walls and membranes around initially distinct tubes coalesce, so that their intracellular cytoplasm, which in its turn contains membrane bound organelles (nuclei and mitochondria) becomes continuous. a visible recoil can occur in the receptive hypha when the tubes coalesce. (photographed by dr a.m. ainsworth). a living illustration of the process of figural boundaries coming into proximity, contiguity and conjugation occurs during the process of hyphal fusion that is found in many fungi[16] and is shown in fig.4. here some fundamental differences between rationalistic and natural inclusional perceptions of connectivity and continuity emerge: (1). in rationalistic thought, continuity is equated with ‘connectedness’ because space is regarded as void, a source of discontinuity or disruptive gap between and around ‘things’ as discrete objects. hence the only way of deriving continuity in this ‘whole way of thinking’, is either by totally excluding space and boundaries from form as a continuous line or network of width-less threads, or by totally conflating space with form in a seamless [distinction-less] whole. such exclusion or conflation is neither consistent with evidence/experience nor does it advances in systems science and applications (2012) vol.12 no.4 341 make consistent sense. (2). in natural inclusional thought, space is a continuous omnipresence that cannot be cut, occupied, confined or excluded, and form is dynamically continuous through its energetic inclusion of space throughout figure and figure in space. distinction and difference are hence accommodated in a natural fluid continuum, without contradiction. local identity is recognised as a dynamic inclusion of non-local space in which all forms are pooled together (but not merged into complete unity) in natural communion as flow-forms. (3). correspondingly, the treatment of continuity by objective rationality as the same as connectedness as exemplified in conventional calculus, where continuity is approximated by connecting infinitesimal discontinuous units is an idealized abstraction that is physically impossible. the very idea of complete ‘whole units’ existing anywhere, at any scale in nature as an energetically open, fluid system does not make sense. the fluidly variable connectivity of natural inclusionality arises from the coming together (contiguity/interconnectivity), fusion (confluence/intra-connectivity) and dissociation (individuation/differentiation) of energetic paths, corridors or channels of included space in labyrinthine branching systems and networks (i.e. as shown in fig. 2), not the ‘ties that bind all into a web of one’[1,31,35]. 6 flow-networking the space-including processes of regeneration, degeneration, differentiation and integration illustrated in figs 1-4 are very different from the purely tangible connectedness of modern network theory. they inform us about how energy is assimilated, located, conserved and redistributed in real-world sustainable systems, as distinct from abstract mathematical models. this is the understanding that i suggest we need to incorporate into systems science. as we do this, there are a number of principles that we need to remember. rather than being formed by stringing together a given set of initially independent entities, flow-networks grow into place through a combination of selfdifferentiating (boundary-maximizing) and self-integrating (boundary-minimizing) processes that configure and reconfigure space in dynamic correspondence with energy availability. for example, fungal mycelia form when a spore germinates by first swelling symmetrically as it takes in water and nutrients across its bounding cell wall and membrane. the resulting structure then becomes polarized, hence breaking spherical symmetry and increasing surface area to volume ratio, through the emergence of a germ-tube or ‘hypha’ with a parabolic growing tip. as this tube elongates, its growth accelerates exponentially, as the absorptive surface increases, before attaining a more or less constant rate of extension, whence branches begin to emerge, each with their own parabolic growing tips. a den342 alan rayner:what are natural systems, actually dritic (tree-like) system of hyphal branches develops, which radiates out in all directions. eventually, in many fungi, as resources are depleted by the growing system, some of the branches begin to fuse or anastomose with one another, so converting the inner part of the system into a network of labyrinthine channels (as in fig.4). during this process, the branches open up their external boundaries to one another, so that the inter-space initially between them becomes the continuum of intra-space within them. in other words, they ‘let go’ of their individuated self-identity ‘agenda’ in the process of coalescing by self-integration (cf. fig.2). within the integrated system, the branches do not disappear, but retain their form as connective channels of intra-space. the nodes in this system are the places from which the branches originally arose, rather than the loci of initially discrete entities. the branch-identities are the connective channels in the system, not the ‘knots’ or local centres through which network transactions are administratively controlled. at no stage in the evolution of the system have these identities been fully dislocated from one another or the limitless pool of common space in which they are immersed and of which they are dynamic inclusions. by growing into place, these dynamic systems exhibit indeterminacy, the potential for indefinite expansion and transformation within boundaries that vary in their deformability, permeability and connectivity depending on contextual circumstances. this contrasts with the determinacy assumed by many to apply to creatures like our individual selves, sentenced to death within a fixed frame of bodily space and time and so bustling through life as if there were no place else to care for, notwithstanding the continuum of our social space. such indeterminacy brings scope for continual improvisation, discovery and learning through co-creative evolutionary play that is not fixed on a pre-determined course, but eases its own passage through a process of autocatalytic flow in which the flow of current lowers resistance to subsequent flow: sheep, wildebeest, ants and humans all exhibit this phenomenon as they create paths by following in one another’s wake[6]. some fungal mycelia making their way through ancient forest in this fashion are thought to cover up to square kilometres of ground and to be thousands of years old. by connecting their internal space in parallel rather than purely in series (as applies to dendritic systems, lacking anastomoses/cross links), flow-form networks greatly increase their conductivity and consequent capacity to store (i.e. ‘memorize’) and supply power at or to localities on their boundaries (cf figs 1, 3). in fungi, this increased capacity is what allows mycelial systems literally to ‘mushroom’ as well as to produce survival structures such as sclerotia (of which ‘ergots’ are a well known example) and rapidly extending cable-like aggregations known as ‘rhizomorphs’ because of their root-like appearance and growth. mycelial systems that lack or lose the ability to form anastomoses are prone to advances in systems science and applications (2012) vol.12 no.4 343 become dysfunctional and degenerate, proliferating numerous branches from local nodal sites in a way that looks very similar to some unrealistic ‘maps’ that have been made of the internet using purely abstractive analytical techniques. local, well connected centers in flow-form networks drain resources from the system, and inhibit its expansion. in fungi, fruit bodies and storage structures may form at such centers. in human organizations they have the potential to develop into exploitative growths and megalithic power structures. degenerative processes in flow-form networks are vital as a means of preventing retention of power by core components of the system. for example, ‘fairy rings’, consisting of an annulus of spreading mycelium, result from the degeneration of the colony centre and release of its resources to supply the growing margin. in the absence of such degeneration, expansion of the system stalls. death is vital to the possibility of continuing life: it feeds life and opens up new possibilities for reconfiguration it does not annihilate life in the way that the rationalistic view of space as an absence of presence may lead us to believe. the ability of flow-form networks to differentiate, integrate and degenerate, by varying the dynamic properties of their boundaries in tune with their circumstances and avoiding the wastage implicit in rationalistic ‘cost-cutting’, allows them to produce extraordinarily efficient organizations in highly heterogeneous situations. in fungi inhabiting the forest floor, for example, this ability allows them to make connections between local sources of nutrients in decaying wood, leaf litter and roots, to form an underground communicative infrastructure, which brings the lives and deaths of the trees into a common circulation . so, altogether, these living flow networks are far more sensitively attuned to the ever-reconfiguring space that their channels embody, than the inflexible meshwork entrapments our current abstractions represent. how do these principles translate into management praxis? i have just two general suggestions: (1). being alive to any life-form’s unique situation, the way it attunes with its neighbourhood, the complex relationships that such attunement entails, and the ease with which these relationships can be destroyed by insensitive intervention. (2). value one’s own learning experience, be prepared to share this with others and value others’ unique experience, rather than simply following or desiring some ‘one size fits all’ doctrine, fad or short term ‘fix’. i have the feeling that these suggestions might sound rather obvious and lacking any absolute, clear, fixed, authoritarian direction. they might seem like not much more than we might gather about life’s patterns and uncertainties from our everyday experience as relational human beings good neighbours using all our sentient faculties. i do hope so! 344 alan rayner:what are natural systems, actually references [1] rayner a.d.m. (2004), “inclusionality and the role of place, space and dynamic boundaries in evolutionary processes”, philosophica, vol.73, pp.51-70. [2] rayner a.d. (2011), “space cannot be cut: why self-identity naturally includes neighbourhood”, integrative psychological and behavioural science, vol.45, pp.161-184. [3] rayner a.d.m. (2011), naturesscope: unlocking our natural empathy and creativity an inspiring new way of relating to our natural origins and one another through natural inclusion, o books. [4] einstein a. (1954), relativity, university paper back, london: methuen & co. [5] poincaré h. (1905), science and hypothesis, dover publications, walter scott publishing company ltd. [6] rayner a.d.m. (1997), degrees of freedom living in dynamic boundaries, london: imperial college press. [7] rayner a.d.m. (2006), “natural inclusion: how to evolve good neighbourhood”, available from http://www.inclusional-research.org/naturalinclusion.php. [8] darwin c. (1859), on the origin of species by means of natural selection, or the preservation of favoured races in the struggle for life, down,bromley, kent. [9] rayner a, sidebottom b, peleshok d and tattersall p. (2012), “place-time: the flow-geometry of space”, available from http://www.bestthinking.c om/article/permalink/1798?tab=article&title=place-time-the-flow-geometr y-of-space. [10] polanyi m. (1958), personal knowledge: towards a post-critical philosophy, london, routledge and kegan paul, pp.381. [11] rayner a.d.m. (2008), “natural communion: poems and paintings about our human inclusion in the evolutionary flow of place-time”, available from http://www.inclusional-research.org/furtherreading/naturalcommunion.pdf. [12] rayner a.d.m. (2010), “inclusionality and sustainability attuning with the currency of natural energy flow and how this contrasts with abstract economic rationality”, environmental economics, vol.1, pp.98-108. advances in systems science and applications (2012) vol.12 no.4 345 [13] elstrup o. (2009), “the ways of humans: modelling the fundamentals of psychology and social relations”, integrative psychological and behavioural science, vol.43, pp.267-300. [14] elstrup o. (2010), “the ways of humans: the emergence of sense and common sense through language production”, integrative psychological and behavioural science, vol.44, pp.82-95. [15] rayner a.d.m, z.r. watkins, j.r. beeching. (1999), “self-integration an emerging concept from the fungal mycelium”, in “the fungal colony” eds nar gow and gm gadd, cambridge university press, pp.1-24. [16] dowson c.g, rayner, a.d.m & boddy l. (1986), “outgrowth patterns of mycelial cord-forming basidiomycetes from and between woody resource units in soil”, journal of general microbiology, vol.132, pp.203-211. [17] bateson g. (1972), steps to an ecology of mind: collected essays in anthropology, psychiatry, evolution, and epistemology, university of chicago press. [18] rayner a.d.m & jarvilehto t. (2008), “from dichotomy to inclusionality: a transformational understanding of organism-environment relationships and the evolution of human consciousness”, transfigural mathematics, vol.1, no.2, pp.67-82. [19] smith b. (1997), “boundaries: an essay in mereotopology”, in l. hahn (ed.), la salle: open court, the philosophy of roderick chisholm, pp.534-561. [20] valsiner j. (2009), “baldwin’s quest: a universal logic of development”, in j.w. clegg (ed), new brunswick, london: transaction publishers, the observation of human systems lessons from the history of anti-reductionist empirical psychology, pp.45-82. [21] koestler a. (1976), the ghost in the machine, london: hutchinson. [22] rayner a.d.m, coates d, ainsworth a.m, adams t.j.h, williams e.n.d & todd, n.k. (1984), “the biological consequences of the individualistic mycelium”, in d.h. jennings & a.d.m rayner, (eds),cambridge university press, the ecology and physiology of the fungal mycelium, pp.509-540. [23] wilber k. (1996), a brief history of everything, boston: shambhala publications. [24] cairns h.c & harney b.y. (2004), dark sparklers, h.c.cairns. 346 alan rayner:what are natural systems, actually [25] taylor s. (2005), the fall. winchester, new york: o books. [26] walker e.m. (2003), “the confusion of dreams between selves and the other: non-linear continuities in the social dreaming experience”, in w.g. lawrence (ed.), london: karnac books, experiences in social dreaming, pp.215-227. [27] heelas p & lock a. (1981), indigenous psychologies: the anthropology of the self, london: academic press. [28] sampson e. (1988), “indigenous psychologies of the individual and their role in personal and societal functioning”, american psychologist, vol.43, pp.15-22. [29] sampson e. (2000), “reinterpreting individualism and collectivism: their religious roots and monologic versus dialogic person-other relationship”, american psychologist, vol.55, pp.1425-1432. [30] lewis c.s. (1942), the screwtape letters, geoffrey bles. [31] tesson k.j.a. (2006), dynamic networks: an interdisciplinary study of network organization in biological and human organizations, phd thesis, university of bath. [32] shakunle l.o. & rayner a.d.m. (2009), “transfigural foundations for a new physics of natural diversity the variable inclusion of gravitational space in electromagnetic flow-form”, journal of transfigural mathematics, vol.1, no.2, pp.109-122. [33] shakunle l.o. (1994), spiral geometry. the principles (with discourse), hitit verlag, berlin, germany. [34] shakunle l.o & rayner a.d.m. (2008), “superchannel inside and beyond superstring: the natural inclusion of one in all iii”, transfigural mathematics, vol.1, no.3, pp.9-55, pp.59-69. [35] barabási a-l. (2002), linked: the new science of networks, perseus publishing. corresponding author alan rayner can be contacted at:alan@admrayner.plus.com advances in systems science and applications (2011), vol. 11, no. 1-2 27-41 existence results for semilinear fractional functional differential equations with state-dependent delay∗ yong-kui chang1, m.mallika arjunan2, g. m. n’guér ékata3 and v. kavitha4 1department of mathematics, lanzhou jiaotong university, lanzhou, gansu 7300070, p.r. china 2department of mathematics, karunya university, karunya nagar, coimbatore641 114, tamil nadu, india 3 department of mathematics, morgan state university, 1700 e. cold spring lane, baltimore, m.d. 21251, usa 4 department of mathematics, karunya university, karunya nagar, coimbatore641 114, tamil nadu, india email: lzchangyk@163.com, arjunphd07@yahoo.co.in, gaston.n’guerekata@morgan.edu, kavi velubagyam@yahoo.co.in abstract according to theories for α-resolvent family (sα(t))t≥0 and fixed point methods, this paper is mainly concerned with existence of mild solutions to a semilinear fractional functional differential equation with state-dependent delay in a complex banach spacex. some sufficient conditions are established without the compactness of (sα(t))t≥0. keywords fractional differential equations α-resolvent family state-dependent delay fixed point 1. introduction in this paper, we establish the existence of mild solutions to the following fractional func∗this research is supported by nnsf of china (10901075), the key project of chinese ministry of education (210226), the scientific research fund of gansu provincial education department (0804-08) and “qing lan” talent engineering funds (ql-05-16a) by lanzhou jiaotong university. issn 1078-6236 international institute for general systems studies, inc. 28 chang:existence results for semilinear fractional functional . . . . . . tional differential equation with state-dependent delay dαx(t) = ax(t) + f(t, xρ(t,xt)), t ∈ j = [0, b], (1) x0 = ϕ ∈ b, (2) where b > 0, 0 < α < 1 and a : d(a) ⊂ x → x is the infinitesimal generator of an α-resolvent family (sα(t))t≥0 defined on a complex banach space x. the function xs : (−∞, 0] → x, xs(θ) = x(s + θ), belongs to some phase space b that will be defined later ( see section 2), f : j ×b → x, ρ : j ×b → (−∞, b] are appropriate functions and ϕ belongs to the phase space b with ϕ(0) = 0. the fractional derivative dα is understood here in the riemann-liouville sense. the theory of functional differential equations has emerged as an important branch in nonlinear analysis. it is worth mentioning that several important practical problems have lead to investigations of functional differential equations of various types ( see the books of hale et al. [11], wu [31], and the references therein). on the other hand, functional differential equations with state-dependent delay appear frequently in applications as model of equations and for this reason, the study of this type of equation has gained great attention in the last decades, we refer to [4, 5, 9, 14, 15, 22] and the references therein. differential equations of fractional order play a very important role in describing some real world problems. for example some problems in physics, mechanics and other fields can be described with the help of fractional differential equations, see [2, 7, 17, 24, 25] and references therein. the theory of differential equations of fractional order has recently received much attention and now constitutes a significant branch in differential equations. lots of research papers and monographs have appeared devoted to fractional differential equations, for example see [1, 3, 6, 8, 18, 19, 20, 21, 26, 27, 28, 30, 33, 34] and the references therein. motivated by the above mentioned works, the purpose of this paper is to investigate the existence results of mild solutions to a semilinear fractional functional differential equation with state-dependent delay described in the general abstract form (1)-(2). the main technique is based upon the α-resolvent family (sα(t))t≥0 combined with suitable fixed point theorems. this paper is organized as follows. in section 2, we introduce notations, definitions and some lemmas which are used in the sequel. in section 3, we prove the existence of mild solutions advances in systems science and applications (2011), vol. 11, no. 1-2 29 for the problem (1)-(2). 2. preliminaries from now on, we set j = [0, b]. we denote by x a complex banach space with norm ‖ · ‖,c(j,x) the space of all x-valued continuous functions on j , endowed with the topology of uniform convergence with norm ‖x‖∞ := sup t∈j ‖x(t)‖. and l(x) the banach space of all linear and bounded operators on x. moreover, br(z0,z) denotes the closed ball with center at z0 and radius r > 0 in z. definition 2.1 [25] assume that f ∈ cm(r+,x). if α ∈ (m− 1,m), where m ∈ n, then the riemann-liouville fractional derivative of order α ∈ (m− 1,m) is the expression dα t f(t) = dm dtm ∫ t 0 gm−α(t− s)f(s)ds, where for β > 0 gβ(t) =  tβ−1 γ(β) for t > 0, 0 for t ≥ 0. definition 2.2 [25] let α > 0 and f : r+ → r be in l1(r+,x). then the riemann-liouville integral is given by: iαf(t) = 1 γ(α) ∫ t 0 (t− s)α−1f(s)ds. recall that the laplace transform of a function f ∈ l1(r+,x) is defined by: f̂(λ) = ∫ ∞ 0 e−λtf(t)dt, re(λ) > ω, if the integral is absolutely convergent for re(λ) > ω. 30 chang:existence results for semilinear fractional functional . . . . . . definition 2.3 [26] let a be a closed and linear operator with domain d(a) defined on a banach space x and α > 0. let ρ(a) be the resolvent set of a. we call a the generator of an α-resolvent family if there exists ω ≥ 0 and a strongly continuous function sα : r+ → l(x) satisfying sα(0) = i such that {λα : re(λ) > ω} ⊂ ρ(a) and (λα −a)−1x = ∫ ∞ 0 e−λtsα(t)xdt, re(λ) > ω, x ∈ x. in this case, sα(t) is called the α-resolvent family generated by a. for construction of solution by using α-resolvent family, we refer to [26] and the references therein. we also refer to [23, 29] for more information about resolvent or solution operator. remark 2.1 [26] note that if a is the generator of an α-resolvent family (sα(t))t≥0 then the laplace transform of sα(t) is ŝα(λ) = (λα −a)−1. in this paper, we will employ the axiomatic definition of the phase space b introduced by hale and kato in [12] and follow the terminology used in [16]. thus, (b, ‖ · ‖b) will be a seminormed linear space of functions mapping (−∞, 0] to x, and satisfying the following axioms: (a1) if x : (−∞, b) → x with b > 0, is continuous on [0, b] and x0 ∈ b, then for every t ∈ [0, b) the following conditions hold: (i) xt ∈ b; (ii) there exists a positive constant h such that ||x(t)|| ≤ h‖xt‖b ; (iii) there exist two functions k(·),m(·) : r+ → [1,+∞) independent of x(t) with k continuous and m locally bounded such that ‖xt‖b ≤ k(t) sup{||x(s)|| : 0 ≤ s ≤ t}+m(t)‖x0‖b. denote kb = sup{k(t) : t ∈ [0, b]} and mb = sup{m(t) : t ∈ [0, b]}. (a2) for the function x(·) in (a1), xt is a b-valued continuous function on [0, b]. (a3) the space b is complete. advances in systems science and applications (2011), vol. 11, no. 1-2 31 an example of phase space b satisfying (a1) − (a3) is the following space c0 g (see [31], pp.44), where g : [−∞, 0]→ [0,∞) is a given continuous nondecreasing function such that: (i) g(0) = 1 and g(−∞) =∞. (ii) the function g(t) = sup { g(t+s) g(t) : −∞ < s ≤ −t } is locally bounded for t ≥ 0. let c0 g = {φ : (−∞, 0]→ x;φ is continuous and lim s→−∞ |φ(s)| g(s) = 0} then c0 g , together with the following norm: ‖φ‖c0 g = sup |φ(s)| g(s) , satisfies axioms (a1)− (a3). the next lemma is a consequence of the phase space axioms and is proved in [14]. lemma 2.1 ([14]) let ϕ ∈ b and i = (γ, 0] be such that ϕt ∈ b for every t ∈ i . assume that there exists a locally bounded function jϕ : i → [0,∞) such that ‖ϕt‖b ≤ jϕ(t)‖ϕ‖b for every t ∈ i . if x : (−∞, b]→ r is continuous on j and x0 = ϕ, then ‖xs‖b ≤ (mb + jϕ(max{γ,−|s|})‖ϕ‖b +kb‖x‖max{0,s}, for s ∈ (γ, b], where we denotedkb = sup t∈j k(t) andmb = sup t∈j m(t), ‖x‖max{0,s} = sup {‖x(θ)‖, θ ∈ [0,max{0, s}]}. to conclude the current section, we recall the following well-known results. theorem 2.2 [10, theorem 6.5.4]. let d be a closed convex subset of a banach space x and assume that 0 ∈ d. let γ : d → d be a completely continuous map. then, either the set {x ∈ d : x = λγ(x), 0 < λ < 1} is unbounded or the map γ has a fixed point in d. lemma 2.3 ([13, 32]) suppose b ≥ 0, α > 0 and a(t) is a nonnegative function locally integrable on 0 ≤ t < t ( for some t ≤ +∞), and suppose u(t) is nonnegative and locally integrable on 0 ≤ t < t with u(t) ≤ a(t) + b ∫ t 0 (t− s)α−1u(s)ds 32 chang:existence results for semilinear fractional functional . . . . . . on this interval; then u(t) ≤ a(t) + ∫ t 0 [ ∞∑ n=1 (bγ(α))n γ(nα) (t− s)nα−1a(s) ] ds. 3. existence results in this section, we present and prove the existence results for the fractional differential problem (1)-(2). first, we present its mild solution. definition 3.1 a function x : (−∞, b] → x is called a mild solution of (1)-(2) if x0 = φ, xρ(s,xs) ∈ b for each s ∈ j and x(t) = ∫ t 0 sα(t− s)f(s, xρ(s,xs))ds, for each t ∈ j. (3) we are now in a position to state and prove our existence result for the problem (1)-(2). for the study of this, we first list the following hypotheses: (h1) there exist m > 0 and δ > 0 such that ‖sα(t)‖l(x) ≤meδt, t ∈ j . (h2) the function f : j → b → x is completely continuous and there exists a continuous function µ : j → (0,+∞) such that ‖f(t, ψ)‖ ≤ µ(t)‖ψ‖b, (t, ψ) ∈ j × b. (h3) the function t → ϕt is well defined and continuous from the set r(ρ−) = {ρ(s, ψ) : (s, ψ) ∈ j × b, ρ(s, ψ) ≤ 0} into b. moreover, there exists a continuous and bounded function jϕ : r(ρ−)→ (0,∞) such that ‖ϕt‖b ≤ jϕ(t)‖ϕ‖b for every t ∈ r(ρ−). remark 3.1 for more details on the hypothesis (h3), we refer to [14]. theorem 3.1 assume that the hypotheses (h1)-(h3) hold, then the problem (1)-(2) has at least one mild solution on (−∞, b]. advances in systems science and applications (2011), vol. 11, no. 1-2 33 proof. let y = {u ∈ c(j,x) : u(0) = ϕ(0) = 0} endowed with the uniform convergence topology and n : y → y be the operator defined by nx(t) = ∫ t 0 sα(t− s)f(s, xρ(s,xs))ds, for each t ∈ j. (4) where x : (−∞, b] → x is such that x0 = ϕ and x = x on j . from axiom (a1) and our assumption on ϕ, we infer that nx(·) is well defined and continuous. let ϕ : (−∞, b] → x be the extension of ϕ to (−∞, b] such that ϕ(0) = ϕ(0) = 0 on j and jϕ = sup{jϕ : s ∈ r(ρ−)}. we will prove that n(·) is completely continuous from br(0, y ) to br(0, y ). step 1: n is continuous on br(0, y ). let (xn)n∈n be a sequence in br(0, y ) and x ∈ br(0, y ) such that xn → x in y . from the axiom (a1), it is easy to see that (xn)s → xs uniformly for s ∈ (−∞, b] as n → ∞. by (h2) we have ‖f(s, xn ρ(s,(xn)s) )− f(s, x ρ(s,(x)s) )‖ ≤ ‖f(s, xn ρ(s,(xn)s) )− f(s, x ρ(s,(xn)s) )‖+ ‖f(s, x ρ(s,(xn)s) )− f(s, x ρ(s,(x)s) )‖, which implies that f(s, xn ρ(s,(xn)s) ) → f(s, x ρ(s,(x)s) ) as n → ∞ for each x ∈ j . by axiom (a1), lemma 2.1 and dominated convergence theorem, we obtain ‖n(xn)−n(x)‖ = sup t∈j ∥∥∥∥∫ t 0 sα(t− s) [ f(s, xn ρ(s,(xn)s) )− f(s, x ρ(s,(x)s) ) ] ds ∥∥∥∥ → 0 as n→∞. thus, n(·) is continuous. step 2: n maps bounded sets into bounded sets. let µ∗ = sup 0≤τ≤b µ(τ). if x ∈ br(0, y ), from lemma 2.1, follows that ‖xρ(t,xt)‖b ≤ r ∗ = (mb + j ϕ )‖ϕ‖b +kbr. (5) and so ‖n(x)(t)‖ = ∥∥∥∥∫ t 0 sα(t− s)f(s, xρ(s,xs))ds ∥∥∥∥ 34 chang:existence results for semilinear fractional functional . . . . . . ≤ ∫ t 0 ‖sα(t− s)‖l(x)‖f(s, xρ(s,xs))‖ds ≤m ∫ t 0 eδ(t−s)µ(s)‖xρ(s,xs)‖bds ≤mµ∗r∗ ∫ t 0 eδ(t−s)ds ≤mµ∗r∗ eδb δ . this implies that ‖n(x)‖∞ ≤mµ∗r∗ eδb δ = l. step 3: n maps bounded sets into equicontinuous sets. let t1, t2 ∈ j with t1 > t2 and x ∈ br(0, y ). then ‖n(x)(t1)−n(x)(t2)‖ = ∥∥∥∥∫ t1 t2 sα(t1 − s)f(s, xρ(s,xs))ds + ∫ t2 0 [ sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs))ds ∥∥∥∥ ≤ i1 + i2, where i1 = ∫ t1 t2 ‖sα(t1 − s)f(s, xρ(s,xs))‖ds, i2 = ∫ t2 0 ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ ds. here i1 and i2 tend to 0 as t1 → t2 independently of x ∈ br(0, y ). in fact, i1 = ∫ t1 t2 ‖sα(t1 − s)f(s, xρ(s,xs))‖ds ≤ ∫ t1 t2 ‖sα(t1 − s)‖l(x)‖f(s, xρ(s,xs))‖ds ≤m ∫ t1 t2 eδ(t1−s)µ(s)‖xρ(s,xs)‖bds ≤mµ∗r∗ ∫ t1 t2 eδ(t1−s)ds = mµ∗r∗ [eδ(t1−t2) − 1 δ ] . advances in systems science and applications (2011), vol. 11, no. 1-2 35 hence lim t1→t2 i1 = 0. and i2 = ∫ t2 0 ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ ds → 0, since f is compact and sα is strongly continuous, ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ → 0 as t1 → t2 uniformly for x ∈ br(0, y ). we conclude that lim t1→t2 i2 = 0. step 4: the operator n maps br(0, y ) into a relatively compact set in x. indeed from the strong continuity of sα(·) and (h2), the set { sα(t− s)f(s, xρ(s,xs)), t, s ∈ [0, b], x ∈ br(0, y )} is relatively compact in x. moreover, for x ∈ br(0, y ), using the mean value theorem for the bochner integral, we obtain nx(t) ∈ tconv{sα(t− s)f(s, xρ(s,xs)) : s ∈ [0, b], x ∈ br(0, y )}, for all t ∈ [0, b]. consequently the set {nx(t) : x ∈ br(ϕ|j , y )} is relatively compact in x, for every t ∈ [0, b]. step 5: a priori bounds. set λ = {x ∈ y such that x = λn(x) for some 0 < λ < 1}. let x ∈ λ. then for each t ∈ [0, b], we have ‖x(t)‖ ≤ λ ∫ t 0 ‖sα(t− s)‖l(x)‖f(s, xρ(s,xs))‖ds ≤m ∫ t 0 eδ(t−s)µ(s)‖xρ(s,xs)‖bds ≤mµ∗ ∫ t 0 eδ(t−s) [ (mb + j ϕ )‖ϕ‖b +kb‖x‖max{0,s} ] ds ≤mµ∗ ∫ t 0 eδ(t−s) [ (mb + j ϕ )‖ϕ‖b +kb‖x‖s ] ds ≤mµ∗(mb + j ϕ )‖ϕ‖b eδb δ +mµ∗kb ∫ t 0 eδ(t−s)(t− s)α−1(t− s)1−α‖x‖sds ≤ θ1 + θ2 ∫ t 0 (t− s)α−1‖x‖sds, where θ1 = mµ∗(mb + j ϕ )‖ϕ‖b eδb δ θ2 = mµ∗kbe δbb1−α. 36 chang:existence results for semilinear fractional functional . . . . . . in view of lemma 2.3, we have for all t ∈ j , ‖x(t)‖ ≤ θ1 [ 1 + ∫ t 0 ∞∑ n=1 (θ2γ(α))n γ(nα) (t− s)nα−1 ] ds ≤ θ1 [ 1 + ∞∑ n=1 (θ2γ(α))nbnα nαγ(nα) ] = θ1 [ 1 + ∞∑ n=1 [θ2γ(α)bα]n γ(nα+ 1) ] ≤ θ1λα[θ2γ(α)bα], where λα[θ2γ(α)bα] = ∞∑ n=0 [θ2γ(α)bα]n γ(nα+ 1) is the mittag-leffer function. this implies that ‖x‖∞ ≤ θ1λα[θ2γ(α)bα]. hence combining step 1–step 5 and using the theorem 2.2, we obtain that n has a fixed point which is a mild solution of (1)-(2) on (−∞, b]. next, we give an existence result when the nonlinearity f has a sublinear growth with its state variable. let us list the following condition: (h2∗) the function f : j → b → x is completely continuous such that there exist a continuous function µ : j → (0,+∞) and a continuous nondecreasing function w : [0,+∞) → (0,+∞) satisfying ‖f(t, ψ)‖ ≤ µ(t)w (‖ψ‖b), (t, ψ) ∈ j × b, lim inf ξ→+∞ w (ξ) ξ = γ < +∞. theorem 3.2 assume that the hypotheses (h1), (h2∗) and (h3) are satisfied. then the problem ( 1)-(2) admits at least one mild solution on (−∞, b] provided that meδbµ∗kb δ γ < 1,where µ∗ = sup 0≤τ≤b µ(τ). (6) proof. let n be the operator defined by (4). we shall complete the proof by schauder’s fixed point theorem. let r∗ be defined as (5). we claim that there exists a positive number r such that nbr(0, y ) ⊆ br(0, y ). if it is not true, then for each r > 0, there exists xr(·) ∈ br(0, y ), but nxr /∈ br(0, y ), that is, advances in systems science and applications (2011), vol. 11, no. 1-2 37 ‖n(xr)(t)‖ > r for some t(r) ∈ j , where t(r) denotes t depending on r. however, on the other hand, we have from (h1), (h2∗) that r < ‖n(xr)(t)‖ = ∥∥∥∥∫ t 0 sα (t− s) f ( s, xρ(s,xrs) ) ds ∥∥∥∥ ≤ m ∫ t 0 eδ(t−s)µ (s)w (r∗) ds ≤ m eδb δ µ∗w (r∗) . dividing both sides by r and taking the lower limit, we get meδbµ∗kb δ γ ≥ 1, where contradicts (6). hence for some positive r, nbr(0, y ) ⊆ br(0, y ). just the same as the proof in theorem 3.1, we can show that n is continuous on br(0, y ) and n maps br(0, y ) into a relatively compact set in x. next we prove that the family {nx : x ∈ br(0, y )} is an equicontinuous family of functions. let t1, t2 ∈ j with t1 > t2 and x ∈ br(0, y ). then ‖n(x)(t1)−n(x)(t2)‖ = ∥∥∥∥∫ t1 t2 sα(t1 − s)f(s, xρ(s,xs))ds + ∫ t2 0 [ sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs))ds ∥∥∥∥ ≤ i1 + i2. where i1 = ∫ t1 t2 ‖sα(t1 − s)f(s, xρ(s,xs))‖ds, i2 = ∫ t2 0 ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ ds. here i1 and i2 tend to 0 as t1 → t2 independently of x ∈ br(0, y ). indeed, i1 = ∫ t1 t2 ‖sα(t1 − s)f(s, xρ(s,xs))‖ds 38 chang:existence results for semilinear fractional functional . . . . . . ≤ ∫ t1 t2 ‖sα(t1 − s)‖l(x)‖f(s, xρ(s,xs))‖ds ≤m ∫ t1 t2 eδ(t1−s)µ(s)w (‖xρ(s,xs)‖b)ds ≤mµ∗w (r∗) ∫ t1 t2 eδ(t1−s)ds = mµ∗w (r∗) [ eδ(t1−t2) − 1 δ ] . hence lim t1→t2 i1 = 0. and i2 = ∫ t2 0 ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ ds → 0, since f is compact and sα is strongly continuous, ∥∥∥[sα(t1 − s)− sα(t2 − s) ] f(s, xρ(s,xs)) ∥∥∥ → 0 as t1 → t2 uniformly for x ∈ br(0, y ). we deduce that lim t1→t2 i2 = 0. thus, by the arzela-ascoli theorem n is a completely continuous operator. in view of schauder’s fixed point theorem, we deduce that n has a fixed point which is a mild solution of (1)-(2) on (−∞, b]. this finishes the proof. according to theorem 3.2, we can easily obtain the following consequence for the sublinear growth case. (h2∗∗) the function f : j → b → x is completely continuous such that there exist a continuous function µ : j → (0,+∞) and a contant ϑ ∈ (0, 1) satisfying ‖f(t, ψ)‖ ≤ µ(t) [ 1 + (‖ψ‖b)ϑ ] , (t, ψ) ∈ j × b. corollary 3.1 suppose that (h1),(h2∗∗) and (h3)hold. then the problem (1)-(2) has at least one mild solution on (−∞, b]. references [1] r.p. agarwal, y. zhou and y. he, existence of fractional neutral functional differential equations, comp. math. appl., 59(2010), 1095-1100. advances in systems science and applications (2011), vol. 11, no. 1-2 39 [2] r.p. agarwal, m. belmekki and m. benchohra, a survey on semilinear differential equations and inclusions involving riemann-liouville fractional derivative, advances in difference equations, 2009(2009), article id 981728, 47 pages. [3] b. ahmad, j. j. nieto, existence results for a coupled system of nonlinear fractional differential equations with three-point boundary conditions, comp. math. appl, 58(2009), 1838-1843. [4] a. anguraj, m. mallika arjunan, hernández, e, existence results for an impulsive neutral functional differential equation with state-dependent delay, appl. anal., 86(2007), 861– 972. [5] w. aiello, h.i.freedman, j. wu, analysis of a model representing stage-structured population growth with state-dependent time delay, siam j. appl. math., 52(1992), 855–869. [6] z. bai, h. lü, positive solutions for boundary value problem of nonlinear fractional differential equations, j. math. anal. appl., 311(2005), 495-505. [7] b. bonila, m. rivero, l. rodriquez-germa and j.j. trujilio, fractional differential equations as alternative models to nonlinear differential equaitons, appl. math. comput., 187(2007), 79-88. [8] y. k. chang, j. j. nieto, some new existence results for fractional differential inclusions with boundary conditions, math. comput. modelling, 49(2009), 605-609. [9] c. cuevas, g. m. n’guérékata, m. rabelo, mild solutions for impulsive neutral functional differential equations with state-dependent delay, semigroup forum, 80(2010), 375–390. [10] a. granas, j. dugundji, fixed point theory. springer-verlag, new york, 2003. [11] j.k. hale and s. verduyn lunel, introduction to functional differential equations, springer-verlag, new york, 1993. [12] j. hale and j. kato, phase space for retarded equations with infinite dealy, funkcial. ekvac., 21(1978), 11-41. 40 chang:existence results for semilinear fractional functional . . . . . . [13] d. henry, geometric theory of semilinear parabolic equations lecture notes in mathematics, vol. 8, springer-verlag, new york, 1981. [14] e. hernández, a. prokopczyk, luiz. ladeira, a note on partial functional differential equations with state-dependent delay, nonlinear anal. rwa, 7 (2006), 510–519. [15] e. hernández, mark a. mckibben., on statedependent delay partial neutral functionaldifferential equations, appl. math. comput., 186(2007), 294-301. [16] y. hino, s. murukami and t. naito, functional differential equations with unbounded delay, lecture notes in mathematics, 1473, springer-verlag, berlin, 1991. [17] g. jumarie, an approach via fractional analysis to non-linearity induced by coarsegraining in space, nonlinear anal. rwa, 11 (2010), 535-546. [18] a. a. kilbas, hari m. srivastava, and juan j. trujillo, theory and applications of fractional differential equations. north-holland mathematics studies, 204. elsevier science b.v., amsterdam, 2006. [19] v. lakshmikantham and j. v. devi, theory of fractional differential equations in a banach space, eur. j. pure appl. math., 1(2008), 38-45. [20] v. lakshmikantham and a. s. vatsala, basic theory of fractional differential equations, nonlinear anal., 69(2008), 2677-2682. [21] v. lakshmikantham and a. s. vatsala, general uniqueness and monotone iteration technique in fractional differential equations, appl. math. lett., 21(2008), 828-834. [22] w. s. li, y. k. chang, j. j. nieto, solvability of impulsive neutral evolution differential inclusions with state-dependent delay math. comput. modelling, 49(2009), 605–609. [23] c. lizama, regularized solutions for abstract volterra equations, j. math. anal. appl., 243(2000), 278-292. [24] y. f. luchko, m. rivero, j. j. trujillo and m. p. velasco, fractional models, nonlocality and complex systems, comp. math. appl., 59(2010), 1048-1056. advances in systems science and applications (2011), vol. 11, no. 1-2 41 [25] k.s. miller and b. ross, an introduction to the fractional calculus and fractional differential equations, john wiley and sons, inc., new york, 1993. [26] g.m. mophou and g. m. n’guérékata, existence of mild solutions of some semilinear neutral fractional functional evolution equations with infinite delay, appl. math. comput., 216(2010), 61-69. [27] g.m. n’guérékata, a cauchy problem for some fractional abstract differential equations with non local conditions, nonlinear anal., 70(2009), 1873-1876. [28] i. podlubny, fractional differential equations, academic press, new york, 1999. [29] j. pruss, evolutionary integral equations and applications, monographs mathematics volume 87, birkhuser-verlag, 1993. [30] s. g. samko, a. a. kilbas and o. i. marichev, fractional integrals and derivatives. theory and applications, gordon and breach, yverdon, 1993. [31] j. wu, theory and applications of partial functional differential equations, applied mathematical sciences 119, springer-verlag, new york, 1996. [32] h. ye, j. gao and y. ding, a generalized gronwall inequality and its application to a fractional differential equations, j. math. anal. appl., 338(2007), 1075-1081. [33] s. zhang, monotone iterative method for initial value problem involving riemannliouville fractional derivatives , nonlinear anal., 71 (2009), 2087-2093. [34] y. zhong, j. feng and j. li, existence and uniqueness for fractional neutral differential equations with infinite delay, nonlinear anal., 71 (2009), 3249-3256. introduction preliminaries existence results advances in systems science and application (2015) vol.15 no.2 99-118 towards a quantum model for meditation françois dubois1 and christian miquel2 1 dept. of mathematics, conservatoire national des arts et métiers, paris, france. 2 association for the development of mindfulness and the center for mindfulness in medicine, health care, and society. abstract we study the meditative states of human beings from the conceptual framework provided by the fractaquantum hypothesis : analogously to an atom, man can from his “quiet” base state explores various states of higher energy as loving or mystical state. we then look what energy states are explored during meditation: is it the “hyperfine” structure of his base state? is there a love ecstatic state? a very high energy structure mystical state? on one hand we illustrate these hypothesis from the experience of a large part of mystical traditions such as hinduism or buddhism and on another hand from contemporary cognitive sciences. in addition, quantum mechanics indicates that any interaction between energy levels is mediated by a boson of exchange. so we aim to identify the nature of this boson linking the various human being energy levels. keywords fractaquantum hypothesis. 1 fundamental state and fractaquantum hypothesis this work is placed in the context of the fractaquantum hypothesis: the quantum approach is applicable to all indivisible scales in nature, regardless of their size [1]. in addition, the underlying question is that of a quantum model for humans, knowing that a quantum model is not interested in specific individuals, but rather to what they have in common, and makes them “indistinguishable” through this approach. we are guided by the atomic analogy. it is known that an atom is, from the quantum point of view, correctly described by its energy levels. these can be calculated as the eigenvalues of the hamiltonian operator describing the interaction between the atom and the electromagnetic environment outside. thus, the atom has energy levels as proposed by niels bohr [2] and exactly calculable (at least for the hydrogen atom) with the quantum mechanics methods; we refer e.g. the book of cohen-tannoudgi et al [3]. the fundamental level | 0 >, a first excited level | 1 >, a second noted as | 2 >, etc. when the atom passes from the level | 0 > to the level | 1 >, it absorbs the energy gap corresponding to an absorbing wave of frequency ν according to the relation h ν = ε1 − ε0 , (1) where h is planck’s constant. when the atom moves from level | 1 > to the basic level | 0 >, it forwards a ν frequency wave according to the (1) equation. 100 françois dubois and christian miquel: towards a quantum model for meditation in this quantum description, the initial analogy recalled here is located in the human being excited levels. we admit that everyday life, with joys, sorrows, desires and frustrations, is in fact a sum of a fundamental state modulations. this base state could be defined in a first approach, implicitly compared to all states of greater excitement. it is the “current state” of the profane or everyday life, made of low excitations, unstable discontinuous changes according to the every moment exchange of energy made with the outside world, without the usual subject being necessarily aware. these can be closer to what christophe andré [4] calls “states of mind” these mixtures of feelings, thoughts and emotions match each time to a state of temporary power of my relationship to the world, changing during each new configuration, thus being replaced by another state, with both fluctuations and matching micro energy jumps. we call this fundamental state of weakness and disorganized excitement of everyday life, the “daily agitated state”, noted | 0 >, as we will later see. this basic structure is referred to an atom (known technically as molecular orbital) and physics tells us that a “fine structure” (see e.g. the book [3]) exists for each level. the difference of energy between sublevels that forms this fine structure is much smaller than the difference ε1 − ε0 of the relation (1). still in the same analogy, we propose that a succession of energy states | 0, j > and levels ε̃j actually compose the fundamental level | 0 >. as for the atom, it is a “fine structure” even though we cannot yet offer any explicit representation. but as evoked earlier, different feelings, anger, sadness, desires, etc. are all “understate of excitation” of the base state. we know that anger requires energy then falls, such as desires and their acomplishments, or sadness and fear. a real spectroscopist work starts here, classifying energy gaps of these states. following descartes, spinoza was the first to go this route, since his whole ethics is to show how ethics invites to deviate from the resentmentfull sad passions towards joy thus corresponding to the maximum intensity of being, as presented by damasio [5]. the wheel of emotions of plutchik shows how the eight basic emotions (joy, fear, disgust, anger, sadness, surprise, trust, anticipation) can each pass through three degrees or intensity[6]. moreover, descartes opened the way with his treatise on the passions, where he quoted six basic emotions: admiration, hate, love, desire, joy, sadness. from irritationton anger then rage just as aprehension, from right to fear, or in a positive way from serenity to joy then extasy there is allways a fluctuation going along with possible macro-jumps of energy. the lauri nummenmaa et al. finnish study of the aalto university recently deepened and nuanced this classification of emotions, showing in spinoza’s way, that they affect more or less subjective and energy body according to their nature[7]. more than seven hundred volunteers from several countries were asked advances in systems science and application (2015) vol.15 no.2 101 to show on a human silhouette the parts of their bodies that were superactivated, or otherwise impaired, while feeling one of the seven emotions, including the five emotions currently recognized as primary by researchers: joy, sadness, anger, fear, disgust. the results show that the body activation is “very low” to “low” when passing from depression to sadness, contempt or shame, every time with a lower limbs sub-activation and a slight activation of the upper body. conversely when going from disgust to fear and anger, the body activation becomes stronger, mainly in the upper areas of the body. both emotions of happiness and especially love match the strongest activations, this time of the entire body. let’observe how energy, atomic levels energy differences, and by extension man’s excitation levels are of the same nature as an engine’s mechanical energy. this physical concept is well understood since sadi carnot and josiah gibbs pioneering works during the 19th century[8, 9]. 2 from base state to love and mystical states if we now focus on excited states above the base state, we assume that the first excited state | 1 > is love. like the equation (1) it “gives energy”, places the individual in a very specific state. we all have experienced it and know that after a while, passion time, one falls from the state of love to the natural state. it has been described by all cultures over the course of time. as we write these lines in trouville, we believe the christian bobin books (we refer e. g. to the two fundamental la part manquante or le très-bas [10]), to alain de botton, swiss-born english writer who develops a philosophy of love [11], to pierre de ronsart : “mignonne allons voir si la rose?” [12]. to the disappointment in love, staying in the love state | 1 >, while the other is no longer in the relationship, we refer to the works of poets, popular singers and e.g. to “l’écharpe” of de maurice fanon [13] or to the famous french song “mon amant de saint jean” [13]. our idea is not to stop the quantum description to the two previous states, the fundamental state and the being in love. doesn’t plato, in the σνµπóσιoν (the sumpósion), invite us to see the desire, or eros to see the being in love as a springboard in the quest for the beauty, the good and the true, higher states of contemplation, involving most unconscient couples? we propose that the mystical state would be considered “second excited state” of man and noted | 2 >. the mystical state experience is more rare than the being in love. a very poetic description is given by teresa of avila [14] or john of the cross [15] in the christian culture, among the sufi poets, or closer to us by karlfried graf dürkheim [16] or jiddu krishnamurti [17]. we believe that the mystic state is as natural as the being in love. a more “excited” state: the mystics often speak about being “delighted” transported out of them, the body sometimes trembling, approaching the trance phenomena, due to such an excitement. most of the time 102 françois dubois and christian miquel: towards a quantum model for meditation these experiences are fleeting and fragile as these state are difficult to reconcile with the agitation and violence of everyday life. a priori, this state is limited in time and john of the cross speaks of “la noche oscura” to describe the shock of returning to the “fundamental” state. we think possible to exceed the mystical state, beyond the level | 2 > to state | 3 >? what happens then to the over-excitement we were talking about? will it dilute to an appeased base state no longer crossed by any micro-excitation, or will it immerse in a higher energy state? is it only an excitement modulation experience, not such as a fine structure, but rather as “hyperfine” structure as in quantum physics. in addition to their fine structure, atoms have a hyperfine structure involving the coupling the proton magnetic moment with the electron orbital moment and spin of the electron with other quantum effects and an additional coupling being the contact between the electron and the “frontier” of proton, see cohentannoudji et al. op cit.. these famous couplings are widely studied transitions by atomic physicists. should we now speak of a hyperfine structure of mystical excitement, or a new comparable quantum jump towards a state of excitement or a peaceful energy stage higher level | 3 >? to go forward, we propose no longer to focus on ecstatic mystics experiences, but to consider the already mapped meditation experiences, for over two thousand years, with the multiple mindmaps that a regularly trained spirit can experience and more with the recent development of neuroscience, meditation is itself becomes the subject of numerous scientific researches. 3 the different stages and experiences of meditation what is meditation? according to francisco varela [18], who identified the common structures between buddhist meditations, sufi, and orthodox prayers, we always find the same movement of mobilization and attention refocusing, done in three steps: first suspending the attention focus to the world watched as an external object that constantly comes, thus leading to focus (samatha in buddhism) calmly, back to oneself and especially on breath. second, this suspension allows to change the sight and flow of attention, re-directing it towards internal feelings, to all the body or mental micro-excitations that occur at any moment while observing (vipassana) according to michel bitbol [19], their variations as small as unpermanents without being caught by them. by maintaining this self re-directed attention a new state is created marking a real break from the usual excitement state level. according to varela and bitbol, it is then characterized by the third step as being inherent to any meditation practice: attention becomes just open and available as a single “host state”. this is, according to them and to the experiences of people engaged in meditation, a open standby state, obadvances in systems science and application (2015) vol.15 no.2 103 ject and knowledge free and independent of any excitation trace. this relative vacuum time can be very brief, lasting a few seconds or hundredths of seconds, several minutes if not more, time for new objects or stimuli to inevitably come in the field of consciousness. according to bitbol who incorporates the terms of romano [20], the most confusing and sometimes scaring for beginners meditation is to approach this terra incognita of the internal events without knowing in advance what will happen, to accept this short vacuum time that will leave emerge, to getin a “immersed in the seem”, rather than in a “taking reflective distance” (see [19], op. cit.). what is happening now, during meditation? everything? and nothing! thoughts, emotions, pleasant or painful sensations arise, drowsiness or mental agitation threatens. the main thing is to learn not to get caught by them, gradually discovering that there are several possible stages in meditation, each one with interesting correlations to brain. we can see these thoughts as “virtual bosons”, as a classical concept for quantum field theory calculus. if there is communication or exchange between two partners two entities, a interaction boson is transmitted and received (see e.g. [21]). in the case of non-communication, such as the thought of a subject without explicit communication, the boson is emitted and reabsorbed by the same entity, which is characteristic of the virtual bosons, as described e.g. in literature [22]. 4 a first type of meditation experience, exploration of the base states in the first stage, meditation opens to the discovery of internal micro-sensations which are not usually under attention. still, away from the stress of the outside world, indeed attention drops unconsciously from the level of ordinary consciousness to a succession of moods and multiple excitations, a state of mindfulness to inner events. this is not an external posture of the observer, analytically decomposing internal sensations as many objects to observe, but it is using intuitive ability to self-presence, falling below the subject-object cutoff as written by michel bitbol [23]. we learn to see in the field of consciousness, without getting caught by them, internal micro-physical sensations, emotions and partial impulses, that come and go, fleeting as bosons, without leaving trace. this ability to carefully observe the emergence, deployment and loss of multiple emotions or sensations that pass through our bodies every moment without being drawn by them corresponds to a change in the normal work of the brain. current neuroscientists studies show that during full conscience mindfulness meditation, rational areas of thought (broca’s and wernicke’s) are less active, while internal perceptions areas of our own body are instead stimulated and overactive: the fronto-limbic cortex and fronto-parietal cortex that promotes introceptive percep104 françois dubois and christian miquel: towards a quantum model for meditation tion of internal sensations, the anterior cingulate cortex, which helps strengthen body sensations and pain responses as presented in summary of various neuroscience studies about the effects of meditation on the brain including katya rubia’s, of the university of london [24], the left prefrontal cortex that promotes stress management and emotional well-being, particularly in response to negative emotions such as sadness. the more you meditate and become “expert”, the more changes occur in the brain, which reshapes, with, among others, a decrease in the amygdala which is the place of the automatic stress reactions, but also a sensible thickening of the gray matter for a better control of the cortex, an elongation of telomeres responsible for maintaining cell life (stress instead shortening the telomeres and the matching duration of cells life), and a better capacity for sustained attention focus. when the meditator is able to install a certain stability in his mind, we make the assumption that this very fine meditative observation of his internal energy states corresponds, according to the quantum physics model, to a lowering of the daily agitated state, to what we call a “quiet state”, in which it is possible to become aware of the fine and even hyperfine structures man’s base state and micro-sensations and excitations which allways succeed in him. similar observations occur in quantum physics and common buddhist meditation: there is no streaming, but discrete and ephemeral appearances, followed by disappearances and new energy emergencies, whenever discontinued. from the restless state to the exploration of micro-movements and fine or hyperfine structures of the base state, involves attention gymnastics, which gradually through the meditations, instills a four steps new cycle and cognitive mechanism leading different areas of the brain power to work together (see antoine lutz [25]), gradually calming and fixing attention as shown in a irm study of wendy hasenkamp et al. from atlanta university [26]. mental vagrancy which usually occurs by “default” brain operation in automatic mode at the sensory and motor cortices, the meditation, awareness of this mental wandering is followed by a conscious awarness in the cingulate cortex, somatosensory cortex and anterior insula, which allows to pause for a moment the always prevailing elsewhere mind flows. this awareness allows attention to detach the object vagrancy, mobilizing this time the dorsolateral prefrontal cortex and anterior parietal regions. it can then refocus, sustained and quiet, in the prefrontal cortex, with the setting of a broad and open attentive presence to what is happening. 5 a second type of meditative experience, parallel with the being in love if learning meditation is a deepening of man’s finest base state, certain types of meditation help in many traditions, exploring new, more excited states, which would probably refer to love states previously presented, offering all the features advances in systems science and application (2015) vol.15 no.2 105 of the | 1 > step of love or | 2 > step of mystic experience. to support our hypothesis we take examples drawing on two often opposed major eastern traditions: hinduism and buddhism. in hinduism, the bhakti yoga meditation, or path of devotion, is famous since it was themed for the first time in the bhagavad gita [27], between the 5th and 1st century bc. primarily oriented towards vishnu or krishna this bhakti allows to develop and grow, in everyday life and in meditation, a state of love transport with one’s heart divinity. the bhakti implies in fact a complete abandonment to the divinity, installing a person to person relationship to the god, the devotee receiving or discovering an energy and intensity of love far from with the usual feelings of the secular world. as one might expect, in the most advanced stages, such a loving devotion naturally extends towards a mystical state, since it is then to unite and merge into the infinite love of the god, devotees of vishnu or krishna sometimes being remarked in the streets by real love and energy transport while they chant and meditate, seeking to perpetuate this state of ecstasy in love before it fatally falls as this dark night described by st john of the cross. in pathological cases, and specially in the bengali bhakti as noted by esnoul [28], these second fervor states may exceed the meditation framework and overflow crowds of jubilant devotees phenomenons with nervous or functional disorders caused by extreme emotional stimuli. in buddhism, the compassion and kindness loving meditations highlighted by matthieu ricard [29] for tibetan buddhism, are to develop a feeling of infinite love, without limit, extending the feeling of kindness, empathy and compassionate love, even to people we do not like or comply. during these meditations cultivating compassion, meditator goes through strong and sometimes unpleasant thrills leading through into mental intensity and higher energy states, with strong activations in certain areas of brain, measured by neurociences. during the kindness love meditation, richard davidson and antoine lutz [30] (from cern) have registered a higher activation in brain regions that manage empathy for the suffering of others: the anterior insula and the anterior cingulate cortex, but also the three areas that allow us to put ourselves in the place of others (medial prefrontal cortex, superior temporal sulcus, temporoparietal junction). experienced meditators feel even more compassion and suffering of others, with additional activation of the somatosensory cortex that makes them part of their own body. according to matthieu ricard [29], other studies about this loving kindness meditation testify logically from the decreased activity of the amygdala, which manages aggression and fear, while areas related to empathy as the insula were more active, increase in size, creating more neural connections and thus reshaping the brain. the feeling of kindness love goes even further than a simple 106 françois dubois and christian miquel: towards a quantum model for meditation increasing empathy, according to studies reported by matthieu ricard, particularly by promoting and developing oxytocin, this maternal love peptide produced by hypothalamus. 6 a third type of meditative experience, approaching the mystical state if these meditations based on love seem to scale up from the base state to the love state level | 1 >, even to the mystical state | 2 >, are there non devotional meditative practices which would specifically lift to the so called mystic level | 2 >? it seems that this is the case of advanced dhyana samadhi states of concentration and contemplation, which have been cataloged in all meditation traditions of the two great hindu and buddhist currents. for over two thousand years, these states of consciousness were indeed widely listed and presented, sometimes in hinduism as the accession to new energy strata and higher conscience, sometimes in buddhism as stages of consciousness, obtaining advanced powers to keep beware of without stopping. in hinduism: the yoga sutras, between 200 and +500, codified and compiled for probably several centuries experiences and stages developed by yogis during their yoga exercises and meditation. in buddhism: the abhidharma, between first and 5th century, is still an impressive collection of texts and reviews about pyschological and philosophical experiences experienced by the buddha and buddhist meditators. michel bitbol, in his research on consciousness, builds among others on this corpus. let’s present now the first advanced states of consciousness found in these various traditions, stopping for the convenience on the terms and stages described by buddhism (see e.g.[31]), and more specifically in the dhyanas or jhanas, more or less ultimate absorbing states which can reach after long mediations and samadhi concentrations. the first dhyana state that seems to be the very definition of the mystical state given at the beginning of this article is first the experimentation of a pure inner energy or virya excitement state, excitement feeling arising in the depths of oneself. this burst of excitement and energy can be experienced with weak signs as heat, tingling, creepy, or cause, as with the mystics, most violent movements that literally carry the meditator, almost out of himself, giving the impression of being transported untill almost levitating or feel of emerging from his own body in a state of priti passadhi rapture and pure joy which then shows access to the second state. this state of joy can almost naturally lead to a third pure happiness state free of body sensation or excitement, again with a mental experience level break, and the access to a fourth prasrabdhbi equanimity state, of serenity and immutable peace. which can then lead to be completely absorbed into new dhyana or jhana more subtle states according to the texts: the fifth stage, dilution or absorption in an unlimited space; sixth, absorption in a state of unlimited consciousness; seven advances in systems science and application (2015) vol.15 no.2 107 and eight, the absorption in other areas be empty or all of any form, perfectly silent, beyond perceptions and non-perceptions. hindu texts and the different stages of yoga present variations and terms inversions, distinguishing such as mircea eliade [32] after concentration dhyana, and the more and more refined samadhi: first, an enstasis or standard limits liberation in samadhi with support, still maintaining the support of the meditation object. thus leading to miraculous yogic powers [33], and finally when exceeded to the ultimate samadhi without any support, which is a perfect and sudden enlightenment, a most deep inside enstasis, in a higher state of consciousness, “a total and saturated by direct intuition be”. we refer to mircea eliade [34] and to zimmer’s description [35] about these states of samadhi, according to the great vedanta philosopher cankara at 8th century. it is not surprising that these absorption states reached through meditation find the states evidenced by the mystics. in her spiritual journey, st. teresa and reached in the fifth home of her soul a pure attention succeeding an infinite bliss, which is according to bitbol [19] , the third dhyana buddhist, removing the awareness of body sensations to become a pure attention, just after a feeling of joy and bliss. hatching such states requires, ability to focus attention, then let go with abandon, are essential as observed too by jeanne guyon, 13th century mystical: “surrender is the key to the the inner”[36]. regardless the difference stages and terms in yoga or buddhism, favoring samadhi or dhyana. the main thing is to realize that these two traditions mapped these mystical level “two” states quantum of our assumptions for a long while, each time with the same degree of change, a break level, an extension and intensification of energy awareness in both aspects physiological and cognitive. physiological studies in neuroscience allow, again, to aproach indirectly and by comparaison the energy transformations that probably occur in the brain, when such statements are reached in meditation for a short or long period. measures that were conducted on trained meditators [37] show indeed both increased alpha and theta waves of deep relaxation and rem (rapid eye movement) sleep, but also the increase of gamma waves thus reflecting the activation and mobilization of neural resources allowing mental effort and the emergence of a global intuitive flash. as they allow, according chaskalson [37], to daily accumulate sound, visual and cognitive data, to suddenly understand by example in a flash: it is a train, the same waves that cause these flashes of intuitive understanding known by the great scientists, artists and meditators. this is the “aha!”, the “eureka” that last a few milliseconds, which suddenly allows to see the reality otherwise. according chaskalson, even without speaking of the states of samadhi or dhyana we have presented, the mental meditation training allows monks to have the same type of illumination and intuitive flash, with a broad and comprehensive vision that can 108 françois dubois and christian miquel: towards a quantum model for meditation last up to five minutes. how to explain or account for neuroscience? studies by lutz in his research laboratory at the university of wisconsin [25], on expert meditators who have accumulated more than 10,000 hours of meditation, show that stress areas located in the amygdala are much less mobilized, even with an objective stress (sending aggressive sounds), then providing them an attentional and emotional stability. in other studies on more experienced meditators, totaling more than 50,000 hours of meditation, antoine lutz also shows that synchronization of gamma waves is actually much sharper than in average individuals[38], which indicates levels of “increased” consciousness with more integration of the different brain areas begining to work in synergy. 7 paradoxes of mystical states: third state, or return to the base states? if the existence of mystical states corresponding to a “jump” into a higher energy level seems to be demonstrated that these become mystical states: they can exceed that in a later stage we would call a quantum state level | 3 >, or they eventually return to the base state? hinduism and yoga seem to militate in favor of the possibility of reaching a final status level | 3 >. in hinduism, mystical peak condition naturally to contemplation, to the union or merger in a higher power or divine, which goes beyond feelings states absorption described above. this was already revealed in the most ancient hindu texts, bitbol reminder that the chandogya upanishad already stated that six centuries bc “dhyana meditation is more than consciousness” (see [19]). for hindus, the ultimate experience is the effect of the fusion and resorption in the primordial ocean, conceived on the model of the wave or the water drop (atman or ultimate principle in itself), which returns blend into the primordial ocean the absolute (brahman). cankara, large non-dualistic philosopher who belongs to the hindu tradition vedanta to 8th century ad, will be part of the same line when his distinguished according zimmer commentator different samadhi with and without support, to finally end with a oceanic fusion in the absolute: “in the first type of samadhi savikalpa with mindfulness the subject-object duality, the oscillating vitality of consciousness assumes the form of brahman, but remains conscious of itself, has a beneficial ecstasy, nirvikalpa samadhi, absorption without consciousness is immersed in the self, without distinction about objects, such as waves vanish in the water” (see zimmer [35]). but this merger in the infinite seems quantum level | 2 > or | 3 >, is it not the same resorption time in a base state, naturally peaceful and informal, which refers to the paradox of a new level | 0 > ? this is what seems to indicate for which buddhism different stages of dhyana absorption are not an end in itself: they simply allow to abandon the game categories and normal mental projectionadvances in systems science and application (2015) vol.15 no.2 109 s, to evade excitations which continue to occur without one is now assigned or forced to react by simply return to a peaceful state originally, where we show the world just leaves appear here and now, in a perpetual present. nirvana simply means extinction, blow and quench the thirst and the projection of his desires, to let appear the phenomena as they are, in their “like this being” (just as well, tathagata, the epithet enlightened buddha, literally meaning: and went, and came, or more exactly according bitbol: standing in the well). this is what allows bitbol to write: “the state of absorption once pushed into its ends, leading to a territory which is not a stranger, and that seems unusual because he has been stripped of its grid cadastral: what is, as it is” (see [19]). in this sense, it is just an awakening to what is presented to each moment, as if pondering managed to reverse in the original and fundamental quantum field which emerge discrete events or appearances, which immediately effaced, as strange bosons enigmatic. to speak of this final state, the tibetan tradition of talk “open presence” and the zen tradition back to an originally awake and calm the conscience, conscience hishiryo of non-thought, before any thought of where the boundary between the self and the world continues to find itself in what is sometimes mistakenly called the “nature of buddha” (nature of what is just, well). our thinking leads to a strange paradox: how to reconcile the hindu metaphor the drop of water that melts into an infinite ocean of energy, with the buddhist image of the empty mirror dust, which has nothing to reflect? maybe he should accept the aporia as the place and the reverse of the same coin, and that the ultimate mystical state is ultimately nothing more than a return to the base state “appeased” rights, ie the base state within the structure hyperfine state quiet day. that is the common point between the “semi-silent buddhist and theological verb: they look like each other to prepare the mind to meet the unprecedented and the unnamed themselves, evoking in a case like a proto-type, and the painting in the other case under the guise of an over-kind” (see [19]). compared to our quantum assumptions, this would imply that the quantum leap inducing a change of consciousness is thinner in what is called metaphorically an ocean of energy, would be correlative or followed by a return of consciousness ordinary that we named the base state appeased by letting go and extinction apprehensions and tensions of the ego. 8 transition from one state to another and quantum physics at the end of this presentation, meditative states seem to respond well to different intensity scales, and to a lesser extent the affective and emotional states of the human spirit can, within the scope of increased excitement, split from a low state to an average state, then exacerbated such as frustration, anger and rage. but if, for the emotional states of plutchik, we understand that the external stimulus 110 françois dubois and christian miquel: towards a quantum model for meditation makes suddenly switch from one state to another, how will the transition occur with a much stronger meditative state | 0 > to | 1 >, | 2 > or even | 3 > while the meditator seeks to abstract from any external stimulus? we will try to approach this issue from three aspects: the relationships dynamics, bosons, and finally heat exchanges, which each help to determine whether the changes produced by meditation fully meet the requirements of quantum physics. 9 what dynamics of relationships, what breaking levels? in this second part, it is therefore to ask how the transition from one stage to another can be explained. we know that during meditation, thoughts come and go, such as interaction bosons, light, photon to the electromagnetic field. are thoughts bosons that occasionally emerge from the quantum field of a selfrefocused and soothed consciousness? this hypothesis probably deserves further developments! one can also imagine that they are the result of a measurement made by our conscience, the trace of a reduction of the wave packet, the trace of the interaction between the macroscopic observer which is our conscience and microscopic phenomena of our body, specially our brain. let’s recall that we have identified three different scales in energy levels structuring the human psyche. we start from the everyday agitated state, the ordinary mental state of everyday social exchanges. with a high energy input, one goes to the love state then beyond to the mystical state that allows to shift to a contemplation state. the “daily hectic” state itself is made of myriad of sub-states, fine structure of the energy system. of these, joy, anger, etc. as many levels of excitations in the daily regular. in this new spectral system fine structure, the “quiet” state is fundamental for us. the“quiet” state is the reference state of the meditator who gradually descends towards “hyperfine” sub-levels to lead to a new appeased base state, we may call “vacuum state”, in analogy with the vacuum of a photon free electromagnetic field. so we have three interlocked structures for the psychic structure, with levels of energy exchanges inside each similar subsystem, but at very different levels between subsystems. as always in quantum mechanics, the relationship between energy levels is carried by discrete jumps, transfers with a definite energy, such as the photo-electric effect introduced with the relation (1) at the beginning of this contribution. in this case, the interaction mediator, the “boson” as called by the physicists is the photon, unbreakable grain light provided in 1905 by albert einstein [39]. to extend the analogy between the atomic system and human psyche, to develop our attempt of spectroscopy of the human psyche, we must now search this boson of interaction between energy levels of the psyche. it is known that looking for a boson is still a “big deal” in physics. the understanding of the weak interaction with the weinberg-salam model [40] allowed to hypothesize the “intermediate advances in systems science and application (2015) vol.15 no.2 111 boson”, highlighted at cern in 1984 [41]. more recently, the unification of the electro-weak interaction with the strong interaction that binds protons and neutrons in the atomic nucleus [42] led to the discovery of a boson planned in 1964 by robert brout, françois englert, peter higgs [43] and probably many others! we here suggest the book of gilles cohen tannoudji and michel spiro [44] to the reader. 10 a boson for meditation in quantum physics, the change in energy level for the system is possible if a quantum of energy is exchanged with the outside world. a photon, elementary particle of light, allows the transition between energy e0 and energy e1 if (and only if!) its frequency ν is exactly compatibe with the two previous energies through the equation (1). reciprocally, when the atom descends from e1 level to e0 level, it emits a ν frequency photon. the photon is the interaction boson, which mediates the exchange between the atom and the outside world. the question now, as part of a quantum model for humans, is to understand what these energy grains are, these interaction bosons that allow for example to move from the ordinary state to the being in love. this transition generally occurs abruptly; this is the famous “lightning strike”, whose name recalls the exchange of light within the atomic system. butthe matter is also to understand what are these micro-energy exchanges that allow the mediator to explore the hyperfine structure of the human spectrum, and to gradually descend towards the base level, called “vacuum” in the case of the system obtained by quantifying the entire electromagnetic field. in the case of meditation, reversing the outside world look towards the inner being allows the progressive exploration. but it’s probably a simple “quiet state” overall support framework avoiding the strongest disturbances of everyday life. the precise shift from an energy level to the following one should happen (if the quantum framework of this model is correct) throught a very low energy expense of the subject. the outside world disturbance (better: a disturbance from the outside world) results in a still possible energy contribution moving the meditator from an energy level to a more agitated level. what is the nature of this energy input? let’s first recall that it must exacly be on a tuned frequency, as the photon, through a compatibility relationship similar to the (1) relation. if the outside world sends detuned energy grains, even on much larger frequencies, the system state does not change. this is a great discovery of albert einstein in 1905 while he made the assumption of photon. considering the very complex structure of human being, we believe that we should not look for this energy exchange in pure electromagnetism as for the atoms, but probably rather in a still to highlight assembly of biomolecular structures ultimately organizing trade within the human psyche. 112 françois dubois and christian miquel: towards a quantum model for meditation we can hypothesize in line with théodule ribot [45] that a “lightning strike” results of the resonance between an unconscious rememberance buried in the memory and a glimpse of present, with a previous incubation and updating work on the buried memory. so this exchange of information, whom the subject is generally unaware, causes the “lightning strike”, the shift from of the “daily status” to the being in love. is the thought making this exchange of information? it seems clear that thinking requires an energy expenditure (see e.g. the work of giuseppe vitiello [46]) and results in the exchange of information within the bio-physical body system. but it is likely that the resonant equation harmony (1) is not reached and that thought does not allow such a transition. it seems instead that, if one refers to the assumption made for the “lightning strike” during the appearance of the being in love, that the transition between two states escapes the consciousness of the subject. no one chooses to fall in love and neither not to be in love. we should probably seek the exchange of energy between the states of consciousness, hyperfine energy levels similar to those the atom’s in a non conscious process. of course, this process is stimulated by the thought that controls the aware mental state; but the transition itself is beyond thought. we can assume it is produced by an exchange boson. boson which needs to be highlighted. boson allowing the transition between two states of consciousness. boson unconsciously emitted by the body bio-physical system. universal boson which would not depend on the individual: everything could then be explained by reference to a biological level to a neurotransmitter being an interaction boson, not that far from what happens when you’re in love, which cause and match the change of states we are interested in. the lighting of what is happening in the state changes during meditation helps to bring other complementary lightings, surprisingly finding what happens in the love state. changing of state to leave the daily agitated state, requires at least three conditions: first, as seen with varela, a suspension of the link engaged in to the world and a redirection of sight to the pre-reflective life then the setting of that we have called a quiet state, especially with the techniques of concentration on the breath that allow you to install a minimum of stability within discontinuous appearance / disappearance of various excitations and events; finally, a necessary control of attention, which tirelessly leads intentional thought ready to escape to outside, and inside, to turn into a simple and intuitive attentive presence, open and available at what is happening at every moment. this important and even decisive role of attention as a third factor is confirmed by studies in neuroscience that highlight the importance of the prefrontal cortex mobilization, being, among others, the seat of the thought control by the attention. but what is interesting is that these three factors are necessary conditions but not sufficient for a significant change of state in meditation. neither introspective sight advances in systems science and application (2015) vol.15 no.2 113 redirection nor installing a state of tranquility or opening a wide attention as a wellcoming and 360 degrees acceptance as stated bitbol, is sufficient to explain or cause a change of state, the daily agitated state to a mystical state, or emptyness or wellcoming: that change will never be determined and causally provoqued, most of the time it appears suddenly and discontinuously, even brutally, as in the state of love. in soto zen, it is said that you can just promote the enlightenment emergence, but cause it in no case: it will happen spontaneously, naturally, automatically and unconsciously as liked to say deshimaru [47], without any intervention of man. under a neurotransmitter biological influence, a sudden compatibility ofrelationship or, to resume zen terminology, when mysteriously body, mind and breath finally get synchronized and reunited in a “fair” posture, in synch, then, along with the “letting go”, a sudden illumination may occur. rinzai zen is even more radical, by not advocating a progressive illumination that can be prepared as in the soto zen, but a sudden illumination. pai-chang huai-hai, t’chan eighth century chinese master (the chinese t’chan is the original form that will create the japanese zen), was thus responding to a disciple who asked him how to reach the issue: “it can only be reached by the sudden illumination” [48]. hence the further development of koans, these seemingly absurd riddles that teachers offer their disciple. when the student finally drops his intellectual efforts to understand, a sudden illumination, satory, may occur, which leads him to another level of reality or consciousness. “satori is a spiritual experience; it describes the sudden click of buddhist enlightenment. if it is difficult to describe to the uninitiated, the intellectual approach is easily accessible. just imagine archimedes in his bathtub discovering the same name famous thrust: “eureka!”. he experiences a cognitive process in which violence is like an illumination. suddenly the incomprehensible lights in front of the mind flash. this is a sudden process whose instantaneous contrasts with the heaviness of a verbal explanation”, as suggested in the abc-book “ombres nippones” on the website kichigai.com. in the hindu and vedanta tradition, access to samadhi and ultimate levels of consciousness can only, in the same way, occur suddenly, with a sharp break of ontological level of consciousness, as noted by mircea eliade [32] (page 70): “the level break india aims to achieve is in the samaddhi. this enstasis is in fact a rapture, since it is experienced without being provoqued. the enstasis is equivalent to a reintegration of different modes towards the real mode: primary non-duality, before the bipartition of the real into subject-object, undifferentiated fullness (with feeling of) unit and bliss. there is a return to the origin, but enriched by dimensions of freedom and trans-awareness (or consciousness).” 114 françois dubois and christian miquel: towards a quantum model for meditation 11 state changes and thermal changes during meditation, the subject navigates among the hyperfine structures of the quiet state. we want to establish statistics relating to the occupation of each of the energy levels. but we can not talk about temperature because a priori we study a single subject. however, one can imagine averaging up the time the subject spends in each state. this approach is a little daring because there is a real state change dynamic over time. some tried to describe it with thermostatic statistical tools, which implies invariance in time throughout the experiment. however, let’s try continue this analogy. the nj population of the εj level of statistical physics now represents the time θj spent in the ej state of energy. then we can through the partition function “z” (see e.g. the book of bernard diu et al. [49]) defining a “temperature” of the meditator, modulo this transformation by considering the population of nj states during a time θj : z = ∑ j exp ( − εj kt ) , nj = 1 z exp ( − εj kt ) . (2) we then observe on the equation (2) that the psychic temperature is negative (!) if the subject spends more time in the excited state than in the base state. in meditations, it would also implies that when going from daily agitated state to quiet state, there should be a decrease in temperature and a regulation of vital functions. or this is exactly what we see as neuroscience studies on the subject show that meditation causes almost mechanically a decrease in the release of stress cortisol, a regulation of blood pressure and heart rythm, an elevation of immune defenses [38], and for some types of meditation, a decrease in skin temperature as presented by manocha et al. [50]. other tibetan meditation or yoga can also focus on increasing the temperature of the skin to fight against the cold. 12 conclusion in conclusion, if we take the fractaquantum hypothesis in its very maximum, the analogy observed between the atomic system and the psychic system can continue: just as the photon allows moving from one atomic system state to another, there should be a boson (still to identify!) which would be the direct cause of the transition between two states of mind. this boson could be chemical or electromagnetic or purely physical or biophysical. without going to such a prediction, the fractaquantum hypothesis can minimally and by analogy notice the existence of a same mechanism explaining the process and state changes to quantum level and psychological level particularly in meditation. meditation, when it reaches a stable and peaceful state of attention would thus reach or experience a contentfree underlying quantum field, which advances in systems science and application (2015) vol.15 no.2 115 manifests itself locally and occasionally by small bursts or quanta of energy, fragments of thoughts, emotions or fleeting sensations appearing then disappearing immediately, which may ultimately lead to a state of joy and soothed presence. as if those changes and modulations tiny lived during meditation, were musical notes reflecting a fundamental quantum field, quantum song. acknowledgments the english translation is a collaboration with claire couratier. the authors are happy to thank her! references [1] f. dubois. (2002), “fractaquantum hypothesis”, res-systemica, volume 2, 5th european congress of system science, heraklion, greece. [2] n. bohr. (1913), “on the constitution of atoms and molecules, part 1”, philosophical magazine, vol. 26, pp.1-24. [3] c. cohen-tannoudji, b. diu, f. laloë. (1977), quantum mechanics, hermann, paris. [4] c. andré. (2009), states of mind: learning serenity,, odile jacob, paris. [5] a. damasio. (2003),looking for spinoza: joy and sadness, the brains of emotions, odile jacob, paris. [6] r. plutchik, h.r. conte. (1997), circumplex models of personality and emotions, herausgeber. [7] l. nummenmaa, e. glerean, r. hari and j.k. hietanen. (2013), “bodily maps of emotions”, proceedings of the national acadeny of science of the usa, vol. 111, issue 2, p. 646-651, new york. [8] s. carnot. (1824), reflections on the motive power of heat, bachelier, paris. [9] j. gibbs. (1902), elementary principles in statistical mechanics, charles scribner’s sons, new york. [10] c. bobin. (1989), the missing part, the very low, gallimard, paris. [11] a. de botton. (1994), essays in love, macmillan, london. [12] p. de ronsart. (1555), “sweety, let’s go see if the rose”, (the original is in french: “ode xvii à cassandre”), les quatre premiers livres des odes de p. de ronsard vandomois, dediés au roy, a paris. chez la veufve maurice 116 françois dubois and christian miquel: towards a quantum model for meditation de la porte, au clos bruneau, à l’enseigne sainct claude. 1555. avec privilege du roy ; achevé d’imprimer le xxv. de janvier 1555. [13] m. fanon. (1963), with fanon, cbs, paris. [14] teresa de avila. (1921), the interior castle, or the mansions, the benedictines of stanbrook, thomas baker, london. [15] juan de yepes alvarez. (1959), the dark night, translated and edited by e. allison peers, doubleday. [16] k.g. dürkheim. (1967), hara: the vital center of man, translated by sylvia-monica von kospoth, barnes and noble. [17] j. krishnamurti. (1976), the first and last freedom, foreword by aldous huxley, gollancz, london. [18] n. depraz, f. varela, p. vermersch. (2003), on becoming aware. a pragmatics of experiencing, advances in consciousness research, john benjamins publishing company, amsterdam. [19] m. bitbol. (2014), does consciousness has an origin? , flammarion, paris. [20] c. romano. (2005), the life song: faulkner phenomenology, gallimard, paris. [21] c. itzykson, j.b. zuber. (1980), quantum field theory, mcgraw hill, new york. [22] f. dubois. (2004), “fractaquantum tracks”, res-systemica,, volume 4, issue 2. [23] m. bitbol. (2010), from the inside of the world: towards a philosophy and a science of relations, flammarion, paris. [24] k. rubia, a.b. smith, r. halari, f. matsukura, m. mohammad, e. taylor, m.j. brammer. (2009), “disorder-specific dissociation of orbitofrontal dysfunction in boys with pure conduct disorder during reward and ventrolateral prefrontal dysfunction in boys with pure adhd during sustained attention”, the american journal of psychiatry, vol. 166, issue 1, p. 83-94. [25] a. lutz, l.l. greischar, n.b. rawlings, m. ricard, r.j. davidson. (2004), “long term meditators self-induce high amplitude gamma synchrony during mental practice”, proceedings of the national acadeny of science of the usa, vol. 101, issue 46, p. 16369-16373. advances in systems science and application (2015) vol.15 no.2 117 [26] w. hasenkamp, c.d. wilson-mendenhall, e. duncan, l.w. barsalou. (2012), “mind wandering and attention during focused meditation: a finegrained temporal analysis of fluctuating cognitive states”, neuroimage, vol. 59, p. 750-760. [27] a.m. esnoul, o. lacombe (translation). (1972), bhagavad gita, translation and comments, fayard, paris. [28] a.m. esnoul. (1956), “emotional current inside the ancient brahmanical current”, bulletin de l’ecole française de l’extrême orient, volume 48-1, p. 141-207. [29] m. ricard. (2014), advocacy for altruism, benevolence force, édition nil, paris. [30] r.j. davidson, a. lutz. (2008), “buddha’s brain: neuroplasticity and meditation”, ieee signal process mag., vol. 25, issue 1, p. 172-176. [31] j. kornfield. (1994), buddha’s little instruction book, bantam books, new york. [32] m. eliade. (1954), yoga: immortality and freedom, translated by willard r. trask, princeton university press, princeton, 2009. [33] patanjali. (1952), yoga-sutra, chapter iii, p. 16. [34] m. eliade. (1978), a history of religious ideas, translated by willard r. trask, university of chicago press, chicago. [35] h. zimmer. (1953), philosophies of india, joseph campbell. [36] madame guyon. (1995), the middle short stories and other spiritual; subversive simplicity, edited by m.-l. gondal, grenoble, eds jérome millon, grenoble. [37] m. chaskalson. (2011), the mindful workplace: developing resilient individuals and resonant organizations with mbsr, wiley-blackwell, hoboken. [38] a. lutz. (2012), “meditative brain”, cerveau et psycho, number 52, p. 27-33. [39] a. einstein. (1905), “on the electrodynamics of moving bodies”, translated by meghnad saha (1920), annalen der physik, vol. 17, p. 891-921. [40] s.l. glashow. (1961), “partial symmetries of weak interactions”, nuclear physics, vol. 22, p. 579-588. 118 françois dubois and christian miquel: towards a quantum model for meditation [41] s. van der meer. (1981), “stochastic cooling in the cern antiproton accumulator”, ieee transactions on nuclear science, vol. 28, p. 1994-1998. [42] cern. (2012), announcement of the discovery of the boson “beh” at a press conference on the experiences atlas and cms, 04 july 2012. [43] f. englert, r. brout. (1964), “broken symmetry and the mass of gauge vector mesons”, physical review letters, vol. 13, issue 9, p. 321-323. [44] g. cohen tannoudji, m. spiro. (2012), the higgs boson and the mexican hat; a new big narrative of the universe, gallimard, paris. [45] t. ribot. (1910), essay on the passions, alcan, paris. [46] g. vitiello. (1995), “dissipation and memory capacity in the quantum brain model”, international journal of modern physics b, vol. 09, issue 08, p. 973990. [47] t. deshimaru. (1992), the zen way to martial arts, translation by nancy amphoux, penguin books. [48] hoai-häı. (1985), “po-chang kouang-lou”, in tch’an zen; roots and blooms, hermes number 4, fayard. [49] b. diu, c. guthmann, d. lederer, b. roulet. (1997), statistical physics, hermann, paris. [50] r. manocha, d. black, d. spiro, j. ryan, c. stough. (2010), skin temperature changes of a mental silence orientated form of meditation compared to rest”, journal of the international society of life information sciences, vol. 28, issue 1, p. 23-31. corresponding author françois dubois can be contacted at: francois.dubois@cnam.fr advances in systems science and applications (2012) vol.12 no.1 38-45 the vibration response analysis on a 6-ups pkm based on fem jinquan li and zhenxing luan school of automation, beijing university of posts and telecommunications, beijing 100876, china abstract in this paper, a finite element full-scale entity model for 6-ups pkm and its accessories were established based on substructure method and were further confirmed through a modal experiment and a finite element modal analysis. the displacements on response curves of the nodes along the virtual x, y and z axes of pkm were obtained through a random vibration response analysis on pkm performed with a finite element method. based on these resultant curves, the law of vibration response on the pkm was investigated, and the resultant parameters clearly demonstrated the vibration response characteristic of pkm and its response scope magnitude, which can provide an important theoretical base for the optimal structure design of pkm. keywords pkm, random vibration response, stewart platform, fem, dynamics 1 introduction the bkx–i parallel kinematic machine (pkm)[1] is a typical using of stewart platform in machine. it is a 6 -ups pkm, shown as fig.1. fig.1 bkx-i pkm the pkm’s rigid body dynamics model can be built by nearly all classical mechanics methods, such as newton-euler method, lagrange method, virtual displacement principal, kane equation[2-3], etc. so, these methods based on dynamic can be chosen proper and simply according to different type machines advances in systems science and applications (2012), vol.12, no.1 39 which are extensively used in control system. but, the telescopic shafts in pkm shows the flexible characters that the rigid parallel kinematic structures don’t have during the high speed working. the shaft’s elastic deformation, dynamic stress and elastic vibration can cause the whole machine’s impact, noise and fatigue[4]. the pkm’s dynamic system is actual a multi-elastic system, having the characters of mechanism-structure couple, time-varying, nonlinearity, etc. so, during the period of the pkm’s design and optimization, the finite element method is usually used in the modal analysis and the response analysis to reflect the machines’ dynamic characters and the convenience of solution accurately[5]. the machine structure’s random vibration response analysis is one of spectral analysis in the dynamic analysis. it is mainly used in the probability statistics of structure to random vibration response analysis. it is also a relationship curve of power spectral destiny and frequency which reflects the intensity of load and the frequency in time changing. 2 the machine’s finite element solid modeling method the bkx-i pkm’s geometric structure is a complex spatial structure. directly establish its finite element model by using the software of ansys would be complex. in this paper, the geometry solid model including 14 substructures was established by 3d software pro/e. the substructures included the machine frame, motional platform, 6 telescopic shafts and 6 crosses of universal joint which the universal joint’s other accessories belonged to the machine frame and telescopic shaft. the main idea for making substructures was that it could make every moving structure as the substructure while the machine was working. the motional platform is a substructure because it is the pkm’s actuator and also an individual moving part. the machine’s frame is an entirety and holds stationary relative the earth when the machine is working. the universal joint’s upper part was fixed the frame firmly by all kinds of connection, so they could be ranged as a substructure. inside the telescopic shaft assembly, although it’s lower slide bar and upper swing bar had the tendency of motion relatively, both of them and the universal joint’s lower part which connected with them were supposed as a substructure in order to simplify the structure because their connections and power transmission were obtained by the ball screw. the universal joint’s cross as the main member of connecting the machine frame and telescopic shaft is always an individual active structure in any time, so it also was a substructure. after establishing the geometry solid model, it could be transferred to finite element model by the interface between the pro/e and ansys. during the transferring, the relationship between the original pro/e substructures would disappear. the couple configuration of every joint face needs to be made again in ansys to simulate spherical hinge, universal joint and other structures. the 40 jinquan li:the vibration response analysis on a 6-ups pkm based on fem couple configuration was that it equalized the entire contact node’s displacement dof which between the two contact face, and each other’s rotational dof need not to be constrained to make the contact node have the tendency of rotation relatively. after coupling, it could use the solid45 unit to meshing every solid model freely. then the pkm’s finite element solid model including 170 thousand solid units was established, shown as fig.2. fig.2 bkx-i pkm fem model 3 the confirmation and the modal analysis on the finite element model of pkm the modal experiment and finite element theoretic modal analysis on the machine were carried out respectively. then, the comparison of results was made to confirm the validity of finite element model. during the experiment modal analysis the method of single point exciting and multi-point receiving[6] was used. in the experiment, the machine was fixed on the foundation firmly first, which its position and posture were adjusted to as the same as finite element model’s. then the excitation with the hammer were made, while the 4 acceleration sensors picked up and recorded the system’s single exciting signal and multi-response signals at the same time. the experiment statistic’s acquisition and disposal were made by the software of dasp. after a series of signals disposing like analog digital conversion and faster fourier transformation, the system’s transfer function, amplitude-frequency were obtained which can reflect the functional relationship of machine’s dynamic characteristics and then the machine’s natural frequency could be got by normalizing the modal parameter. the boundary condition of finite element theoretic modal analysis was that the machine basement’s dof of displacement was zero. then the block-lanczos method was used during the finite element modal analysis. after that, all of the machine’s 10 order natural frequency, vibration shape and its animation were advances in systems science and applications (2012), vol.12, no.1 41 obtained in ansys. the comparable result between the machine’s natural frequency got from the finite element modal analysis and the machine’s natural frequency got from the experiment was shown in the table 1. from the table we could get that the proper finite element model could be build with the error less than 10% after many times of simplifying the structures and choosing proper finite element unit. table 1 the results comparison between fem’s analysis and experimental modal analysis modal fem actually order calculation value error 1 25.836 23.357 9.6% 2 49.436 45.472 8.0% 3 67.988 63.434 6.7% 4 76.475 72.365 5.4% 4 the theoretic basis of response analysis on pkm when n order natural frequencies of pkm are ω1, ω2, . . . ωn and the corresponding vibration shapes are x1,x2, . . .xn , which natural frequencies are arranged according to sort ascending and the ω1 and x1 are the pkm’s basic frequency and basic vibration shape, the relationship of them can be expressed as follows: [k]xi = ωi 2[m ]xi (i = 1, 2, . . . n) (1) because the integer stiffness matrix [k] and the integer mass matrix [m] are real symmetric positive definite matrix which the vibration shape vectorsx1,x2, . . .xn are a group of basis of n-dimensional space vector, the node displacement vector {ϕ (t)} can be expressed as follows: {ϕ (t)} = q1x1 + q2x2 + . . .+ qnxn = [x]{q} (2) here, q1, q2, . . . qn is the coordinate of {ϕ (t)} in the coordinate system which the vibration shape vectors are basis and they are the time function. so, {q} = {q1 q2 . . . qn}t · [x] is the vibration shape matrix of system which [x] = [x1 x2 . . . xn]. the dynamic equation of machine in the entire coordinate system can be obtained as follows: [m ]{ϕ̈}+ [c]{ϕ̇}+ [k]{ϕ} = {f} (3) 42 jinquan li:the vibration response analysis on a 6-ups pkm based on fem the eq.(2) is substituted into eq.(3) and the equation can be obtained as follows: [m][x]{q̈}+ [c][x]{q̇}+ [k][x]{q} = {f (t)} (4) the equation is left multiplicity by [x]t and the equation can be obtained as follows after collate: mq̈+cq̇+kq = f(t) (5) here, m–main mass matrix of system. m = [x]t [m][x] = diag(m1,m2, . . . ,mn) k–main stiffness matrix of system. k = [x]t [k][x] = diag(k1, k2, . . . , kn) here, ωi 2 = ki/mi c–main damp matrix of system c = [x]t [c][x] = diag(c1, c2, . . . , cn) and ci = 2ξiωimi. here, the ξi is the damp ratio corresponding to the i order vibration shape. f (t) –main active load vector. f (t) = [x]t {f (t)}. expand the eq.(5) and every equation is divided by m̄i , the equation can be obtained as follows after collate: q̈i + 2ξiωiq̇i + ωi 2qi = f i (t) /mi (i = 1, 2, . . . , n) (6) this is a typical single freedom vibration equation belong to linear constant coefficient quadratic differential equation. the chief response of structure system in modal space can be got after solving the equation. qi = eξiωit ( qi0 cosωdit+ q̇i0 + ξiωiqi0 ωdi sinωdit ) + 1 ωdi ∫ t 0 fi (τ) mi e−ξiωi(t−τ) sinωdi (t− τ) dτ (7) here,ωdi = ωi √ 1− ξi 2 {q}0 = [x]t [m ]{ϕ0} {q̇}0 = [x]t [m ]{ϕ̇0} last, when qi are replaced by eq.(7) in eq.(2), the relationship equation between the displacement response in the original physics coordinate system and the excitation could be obtained. advances in systems science and applications (2012), vol.12, no.1 43 5 the random vibration response analysis based on fem the random vibration analysis based on the fem was solved on the base of modal analysis after extending the modal. in ansys, the nodes’ dof of machine base plane were fixed. after extending the modal, the random displacement vibration loads of 1 × 10−5 m amplitude were imposed on servo motors which located on the 6 telescopic shafts and the main axis motors which located on the center of motional platform along the machine’s virtual x and y axes. according to the machine structure damper rate which got from the modal experiment, it could definite the damper rate as 2%. finally, random vibration analysis was implemented to the pkm’s general structure to get the curve of the machine structure node’s displacement response when the random vibration frequency was between zero and 200hz. fig.3 the random vibration response curve of motional platform centre node along the x axis the displacement response curve of the node centre on the motional platform along the virtual x axis was shown as fig3, on the y axis curve shown as fig 4, on the z axis curve shown as fig5. for saving the words, other nodes’ curves were not shown. in fig3, fig4 and fig5, the horizontal axis was frequency which its unit was hz, the longitudinal axis was displacement which its unit was meter. 6 the analysis and conclusions of calculation results some conclusions were got from the pkm’s random vibration response curves: first, all of the virtual x, y, z axes had a great response around the 26hz of the machine’s first order natural frequency. it was the resonance response which didn’t happen in the second and third order natural frequency. at the same time, 44 jinquan li:the vibration response analysis on a 6-ups pkm based on fem fig.4 the random vibration response curve of motional platform centre node along the y axis fig.5 the random vibration response curve of motional platform centre node along the z axis a smaller resonance happened in the fourth order natural frequency. from the modal shape of machine and its animation, the result could be made that the pkm’s first order modal shape was the machine’s whole modal shape, the second and third order modal shape were machine’s local modal shape. the vibration response excitation which loaded on the centre of motional platform didn’t excite the machine’s second and third order response. the result of machine’s vibration analysis and the result of machine’s modal analysis were closely fit each other. second, the motional platform centre’s maximum resonance amplitude was 4.5× 10−6 m which was half of the loaded maximum resonance amplitude. comparing to the vibration excitation amplitude, the vibration response value was advances in systems science and applications (2012), vol.12, no.1 45 less than 2 orders of magnitude when the pkm didn’t resonate, so its influence could be neglected during the actually manufacturing progress. third, the motional platform centre resonated in some low-order natural frequency along the virtual z axis, but compared with the biggest random vibration amplitude, its values were less than 4 orders of magnitude in vibration magnitude. so it could be nearly neglected. fourth, the result, which the random response along the x and y axis’ was greater than that along the z axis’, was just like the rigidity which along the machine x and y axis’ was far lower than that along the z axis’. the result of machine’s vibration analysis and the result of machine’s rigidity analysis[6] were closely fit each other. acknowledgements the project is supported by national postdoctoral fund.(china 2005038083). references [1] jinquan lin, hongsheng ding, tie fu, siqin pang. (2003), “the parallel machine tool’s history, status quo and prospect”, machine tool & hydraulics, no.3, pp.3-8. [2] zhao-cai du, yue-qing yu, li-ying su. (2006), “effects of inertia parameters of moving platform on dynamic characteristic of flexible parallel mechanism and optimal design”, optic and precision engineer, vol.14, pp.1009-1016. [3] shan-zeng liu, yue-qing yu, zhao-cai du, jian-xin yang. (2007), “recent development and current status of parallel manipulators”, modular machine tool & automatic manufacturing technique, no.7, pp.4-10. [4] ju f, choo y s, cui f s. (2006), “dynamic response of tower crane induced by the pendulum motion of the payload”, int.j.solids structures, vol.43, no.2, pp.376-389. [5] jin-quan li, tie fu, zhen-xing luan. (2008), “the dynamic analysis on a 6ups pkm based on fem. optic and precision engineer”, journal of beijing university of posts and telecommunications, vol.31, no.5, pp.40-43. [6] jin-quan li, hong-sheng ding, tie fu, si-qin pang, xie dian-huang. (2004), “the experimental modal analysis of the bkx-i model pkm”, manufacturing technology and machine tool, no.10, pp.45-48. advances in systems science and applications (2014) vol.14 no.3 244-253 symbolic dynamics applied to velocity time-series in wind farms k. karamanos1, i.s. mistakidis2, s.i. mistakidis3 and i.e. sarris4, 1 complex systems group, institute of nuclear and particle physics, ncsr demokritos, gr-15310, aghia paraskevi, attiki, greece. 2 department of physical sciences and applications, hellenic army academy, vari, attiki, greece. 3 zentrum für optische quantentechnologien, universität hamburg, luruper chaussee 149, 22761 hamburg, germany. 4 department of energy technology, technological & educational institute of athens, ag. spyridona 17, 12210 athens, greece. abstract the development of standards for wind farms, presupposes the correct description of wind potential and this can be done with the field measurements of wind flow by cup anemometers. the utilization of new concepts, coming from the world of cybernetics of nonlinear science and complex systems could open the road to uncover information hidden in both the mean polar velocity and the mean angle time-series. in particular, with the use of block entropies, it is shown that we can achieve a better and deeper understanding of the phenomenon of filtered turbulence, produced by time-series of the average wind velocity logged every ten minutes. the present analysis allows in principle a characterization of the experimental time-series in terms of the complexity for selected stationary windows of the signal, as well as the underlying mechanisms of the filtered turbulence. keywords wind farm, anemometer, wind velocity measurements, filtered turbulence, symbolic dynamics, block-entropies. 1 introduction. in the context of energetic problems of modern societies, a proposed solution has been the use of alternative energy sources. one particular realization of this idea is for instance the installation and usage of wind farms for electricity. a cluster of wind turbines in the same site used to produce energy is called wind farm. investment in a wind farm consults temporal measurements of the wind potential on the prospective site by using suitably located towers. the wind velocity, the pressure and temperature, are collected in the measurement tower by cup anemometers and meteorological instruments [1]. the height of the tower is up to the hub height of the planned wind turbines, and the data are recorded frequently, i.e. contains the averaged quantities every ten minutes, and for at several months. these data, in principle, may allow the developer to decide if the investment in a wind farm is economically feasible for the selected site. understanding and quantify aspects of their behaviour and nature is one of advances in systems science and applications (2014) vol.14 no.3 245 the most important and challenging problem of modern engineering [2]. the long recordings of the averaged wind velocity time-series from the anemometer, are a kind of filtered turbulent temporal data set with intermediate properties that may be analysed. the most important problem in this framework is the prediction of the production of electric energy in wind farms [2]. the problem of prediction is exactly, what connects these studies with cybernetics and nonlinear science. 0 1 2 3 4 5 6 7 8 9 10 0 20 months m /s 0 1 2 3 4 5 6 7 8 9 10 0 5 months m /s 0 1 2 3 4 5 6 7 8 9 10 0 300 months d eg re es 0 1 2 3 4 5 6 7 8 9 10 0 150 months d eg re es (c) (a) (d) (b) fig.1 time-series of the velocity field between 23 march 2005 and 16 december 2005. the collected data have been obtained from the mountains of peloponnesus, greece. shown are the distributions over time (obtained from the anemometer) of (a) the mean polar velocity in m/s, (b) the standard deviation of the mean velocity, the distribution of (c) the mean angles (degrees) and (d) the standard deviation of the angles. 0 1 2 3 4 5 6 7 8 9 10 0 10 20 30 months m /s 0 1 2 3 4 5 6 7 8 9 10 0 200 400 months d eg re es w 6 w 12 w 10 w 1 w 5 (a) (b) w 11 w 8w 7 w 3 w 2 w 4 w 9 fig.2 segments of experimental time-series (a) for the mean value of the velocity and (b) the mean angle of the anemometer depicted in fig.1. the vertical lines indicate the months during the measurements. 246 k. karamanos : symbolic dynamics applied to velocity time-series in wind farms the development of standards for wind farms presupposes the correct description of wind potential and this can be done only with the field measurements of wind force by anemometers. herein, we utilize new concepts coming from the world of nonlinear physics, to quantify and understand the relative complexity of wind signals of filtered or mild turbulence. although field measurements of filtered turbulence by mechanical purpose anemometers presuppose enormous and violent averagings on the molecular nature of turbulence, also averagings in time (months, epochs), not to mention the inertial phenomena in the measuring apparatus. we intend to show that a coherent and self-consistent mathematical description of turbulence within the arsenal of nonlinear physics is not only possible, but also beneficial for physicists and engineers. in this paper, one set of wind speed data from one measurement tower situated in the mountains of the region achaia, peloponnesus, greece, in the form of polar velocity and angle was analyzed. the data set consists of some thousands of wind speed values, recorded over every ten minutes by a cup anemometer, covers the period march 2005 to december 2005. more specifically, the wind polar speed and angle was measured using cup anemometers and the 38500 recordings logged into a digital anemograph logging equipment system. the complexity and variability of the data is analyzed here by the use of nonlinear techniques, and the information hidden in both the mean polar velocity and the mean polar angle time series are discussed. more precisely, the temporal evolution of nonlinear characteristics is studied by applying a recently proposed technique [4-5]. the original continuous time, of the mean polar velocity/angle data are projected to symbolic sequence and a block entropy analysis by the novelty of lumping follows [4-5]. the paper is articulated as follows : in sec 2, we recall basic facts about symbolic sequences, and the block entropy analysis by lumping. sec 3, will be devoted to the application of the entropy analysis by lumping to anemometer recordings. finally, in the last sec 4, we draw the main conclusions and discuss future plans. 2 symbolic dynamics a way to examine transient phenomena is to analyze the original time series (anemometer recordings) into a sequence of distinct time windows (epochs). the basic aim is to discover a clear difference of dynamical characteristics as time evolves by employing techniques from the toolbox of symbolic dynamics. in particular here we employ the notion of block entropy analysis by lumping [4-10]. towards this direction, within a stationary time window, the block entropy serves as a measure of "complexity" of the signal. the lower the value of entropy, the more "ordered" it is. in the following, in order to proceed with the analysis of the experimental data we will briefly review the concepts of symbolic dynamics and advances in systems science and applications (2014) vol.14 no.3 247 0 2 4 6 8 10 12 0 0.2 0.4 0.6 0.8 1 n h (n ) / n 0 2 4 6 8 10 12 0 0.2 0.4 0.6 0.8 1 n h (n ) / n w 1 w 2 w 3 w 4 w 5 w 6 w 7 w 8 w 9 w 10 w 11 w 12 (a) (b) fig.3 block-entropy per letter as a function of the word length for the various stationary windows shown in figs.2 (a) and (b). we observe a reduction of the block-entropy per letter which can be interpreted as a sign of complexity reduction of the respective time-window of the signal. the notion of block entropies [3-18]. we restrict ourselves to the simplest possible coarse graining of the recording. this is given by choosing a threshold c and assigning the symbols "1" and "0" to the signal, depending on whether it is above or below the threshold (binary partition or bipartition). in this way, each stationary time window of the original aiolic time-series for a given threshold is transformed into symbolic sequences, which contains "linguistic" or "symbolic dynamics" characteristics. more specifically, the block entropies, depending on the word-frequency distribution, are of special interest, extending shannon’s classical definition of the entropy of a single state to the entropy of a succession of states. thus, each entropy takes a large (small) value if there are many (few) kinds of patterns, i.e it decreases while the organization of patterns is increasing. in this manner we can argue that the block entropy constitutes a measure the complexity of a stationary signal. in particular, we estimate the block entropy by lumping [4-5]. lumping is the reading of the symbolic sequence by "taking portions", as opposed to gliding [1017] where one has essentially a "moving frame". in general, the basic novelty of the analysis by lumping is that, unlike the fourier transform or the conventional entropy by gliding, it gives results that can be related to algorithmic aspects of the sequences. it is useful to transform the initial raw data of the anemometer recording into symbolic sequences taking values in the alphabet {0, 1}, according to the rules ai = 1 if a(ti) > e[a(ti)] and ai = 0, if a(ti) < e[a(ti)]. here the quantities a(ti) denote the values of the measured mean velocity/polar angle at 248 k. karamanos : symbolic dynamics applied to velocity time-series in wind farms 0 2 4 6 8 10 12 0.5 1 1.5 2 2.5 3 3.5 n h (n ) w 2 w 6 w 12 fig.4 the observed scaling of the block entropy h(n) (eq. 6) as a function of the word length n. the experimental values show a linear best fit for small word lengths with a very good precision. the slope of the line gives the kolmogorovsinai entropy which in 1d coincides with the lyapunov exponent. this scaling is consistent with the corresponding theoretical predictions, see nicolis and gaspard [13]. time ti and e[a(ti)] =< a(ti) > is the mean value in the particular time windows. so, let us consider a subsequence of length n , selected out of a very long (theoretically infinite) symbolic sequence. we stipulate that this subsequence is to be read in terms of distinct "blocks" of length n, i.e : · · ·a1...an︸ ︷︷ ︸ b1 an+1...a2n︸ ︷︷ ︸ b2 ... ajn+1...a(j+1)n︸ ︷︷ ︸ bj+1 · · · . (1) we call this reading procedure "lumping" and we shall implement it in the following to the experimental time-series of the filtered turbulence. it is also useful to mention the following quantities which give the information content of the sequence : – the dynamical (shannon-like) block-entropy for blocks of length n has the form : h(n) = ∑ (a1,...,an) p(n)(a1, ..., an) ln p (n)(a1, ..., an), (2) where the probability of occurrence of a block a1, ..., an, denoted p(n)(a1, ..., an), is defined by the fraction (when it exists) in the statistical limit as follows : no of blocksa1, ..., anencounteredwhen lumping total no of blockswhen lumping , (3) advances in systems science and applications (2014) vol.14 no.3 249 table 1 the kolmogorov-sinai (ks) entropy h, the estimated error δh and the corresponding percentage h/ln(2) (see fig.5) in respect to the maximum value of the ks entropy for the different windows (see figs. 2(a),(b)) window no. h δh h/ ln 2(%) w1 0.2192 0.0350 31.6 w2 0.2784 0.0163 40.2 w3 0.2588 0.0407 37.3 w4 0.2040 0.0128 29.4 w5 0.2301 0.0152 33.2 w6 0.2065 0.0194 29.8 w7 0.2611 0.0576 37.7 w8 0.1879 0.0414 27.1 w9 0.1781 0.0611 25.7 w10 0.1245 0.0398 18.0 w11 0.3140 0.0535 45.3 w12 0.1277 0.0201 18.4 starting from the beginning of the sequence. however, the associate entropy per letter reads : h(n) = h(n) n . (4) – on the other hand, the entropy of the source (a topological invariant), defined in the limit (if it exists) reads : h = lim n→∞ h(n), (5) which is the discrete analog of metric or the kolmogorov-sinai entropy. therefore, in order to determine the abundance of long blocks one is led to examine the scaling properties of h(n) as a function of n (see also fig.4). 3 results in terms of symbolic dynamics. to begin with, in fig. 3 we depict the block entropy by lumping per letter as a function of the word length for the selected time windows that we present in fig. 2. we note that a complete absence of structure in the signal, would lead to an horizontal line in the block entropy diagram. as one can observe this is not the present case. however, one important conjecture, due essentially to ebeling and nicolis [13-14] states that the most general (asymptotic) scaling of the block entropies takes the form h(n) = y + nh+ gnµ0(lnn)µ1 , (6) 250 k. karamanos : symbolic dynamics applied to velocity time-series in wind farms 1 2 3 4 5 6 0 10 20 30 40 50 # of window h / ln (2 ) (% ) 1 2 3 4 5 6 0 10 20 30 40 50 # of window h / ln (2 ) (% ) (a) (b) fig.5 the normalized kolmogorov-sinai entropy (taken as the normalized slope of the linear part of the block entropy h(n)), for the respective time-windows of (a) the mean velocities (w1-w6), and (b) the polar angles (w7-w12). where y, h and g are constants and µ0 and µ1 are the corresponding non-scaling constant exponents. in the following, we attempt to examine the behavior of eq.(6) for each of the twelve stationary windows under study, which are depicted in fig. 2. hence, in fig. 4 we present the typical variation of the block entropy by lumping h(n) as a function of the word length n for three representative windows. this study reveals that if we restrict ourselves to the first five values of h(n), a linear scaling is observed with a great precision. in this manner, we next perform a least square method for this region and we estimate the slope h. note, that the associated correlation coefficients r, are close to 1 with a precision better than 10−4. working similarly for the rest of the 9 time windows, we conclude that for n < 6 the same behavior is observed, i.e. the equation for the scaling of the block entropy by lumping, is transformed to the remarkably simple linear relation h(n) = y + nh. (7) this means that g = 0 , for n < 6. in fig 2(a),2(b) we isolate 6 time windows for the mean velocity and 6 time windows for the polar angle respectively, which present a good overall stationary behavior according to our tests. their corresponding ks entropy is given in table 1. the ks entropy (complexity) of the whole window w6 (mean velocity) is of the order of 30%. this shows that the underlying mechanism of filtered turbulence has an underlying organized molecular basis, as it does not correspond to a completely random process (bernoulli shift). the maximum value of ks entropy for the mean velocity is about 41% for the window (w2), while the minimum value is about 30% (see also fig.5). a close inspection of the w2 window, shows that it has a structure more reminiscent of advances in systems science and applications (2014) vol.14 no.3 251 noise and this is confirmed by its high value of ks entropy. the ks entropy of the whole window for the mean angle (w12) is about 19% and it is quite low. the highest value is achieved in w11 which is of the order of 46%. we note that when g = 0 and h > 0, long words are penalized exponentially. we focus on the quantity h, namely the kolmogorov-sinai entropy defined as the slope of eq. (7). we notice that for a one-dimensional process the kolmogorov-sinai entropy coincides with its lyapunov exponent. the lyapunov exponent under these conditions gives a measure of the chaoticity (or dynamical randomness) of the signal. for a two-letter alphabet, the kolmogorov-sinai entropy h takes values from zero to ln 2 (see the discussion in sec 2), so that one can normalize dividing by ln 2 and obtain the respective percentage. hence, it is important to note that the linear part of the scaling helps for a classification and categorization of the recordings. the question which arises naturally, is whether this is an independent algorithmic law of nature. this seems to be an open problem for the moment. however, our results strongly support this hypothesis. remark : we restrict ourselves to the region n < 6, because the maximum statistical accuracy for the block entropies by lumping is of the order of lnl, where l is the total number of points (the size of the window). in our case l, is of the order of 2800, so that n < 8 and due to the underestimation of the higher entropies, we have enough statistical precision for n < 6. 4 conclusions. in this paper, we made an attempt to apply techniques and methods from cybernetics, nonlinear physics and general systems to the temporal unfolding of the phenomenon of filtered turbulence in wind farms as measured by cup anemometers. the tremendous molecular averaging implied by this simplistic procedure, has been equilibrated by the high specialization and the novelty of the techniques used in general systems. block entropy analysis by lumping, as introduced by karamanos et al. [4-5], is for the first time used for the understanding and categorization of time series of filtered turbulence in anemometer recording. a first linear region has been revealed, as it has been already happened for em preseismic precursors [6], cardiac signals of coronary patients [9], and dna strands in oligonucleotide basis acgt [8]. hence, it is important to note that the linear part of the scaling helps us for a complete classification and categorization of the recordings. as it has been already pointed out, the question which arises naturally, is whether this is an independent algorithmic law of nature. this seems to be an open problem for the moment. however, our results strongly support this hypothesis. future projects include the enhancement of data and recordings, the crosschecking of many complexity measures, the month-to-month monitoring of the 252 k. karamanos: symbolic dynamics applied to velocity time-series in wind farms dynamics and the application of prediction techniques, with special use to wind parks. references [1] probst, o. and cárdenas, d. (2010), “state of the art and trends in wind resource assessment”, energies, 3, 1087-1141. [2] huang, z. and chalabi, z.s. (1995), “use of time-series analysis to model and forecast wind speed”, journal of wind engineering and industrial aerodynamics, vol.56, no.2-3, pp.311-322. [3] nicolis, g., rao, g., rao, j. and nicolis, c. (1988), generation of spatially asymmetric, information-rich structures in far from equilibrium systems, in: christiansen. p. l., and parmentier, r. d. (eds.). (1989), “structure, coherence, and chaos in dynamical systems”, r. d., manchester university press. [4] karamanos, k., and nicolis, g. (1999), “symbolic dynamics and entropy analysis of feigenbaum limit sets”, chaos, solitons and fractals, vol.10, no.7, pp.1135-1150. [5] karamanos, k. (2001), “entropy analysis of substitutive sequences revisited”, journal of physics a: mathematical and general, vol.34, no.43, pp.9231. [6] karamanos, k., peratzakis a., kapiris p., nikolopoulos s., kopanas j. and eftaxias k. (2005), “extracting preseismic electromagnetic signatures in terms of symbolic dynamics”, nonlinear processes in geophysics (npg), vol.12, pp.835-848. [7] karamanos, k., dakopoulos, d., aloupis., k., peratzakis, a., athanasopoulou, l., nikolopoulos, s., kapiris, p. and eftaxias, k. (2006), “preseismic electromagnetic signals in terms of complexity”, physical review e, vol.74, no.1, 016104, pp.21. [8] karamanos, k., kotsireas, i.s., peratzakis, a., and eftaxias, k. (2006), “statistical compressibility analysis of dna sequences by generalized entropy-like quantities: towards algorithmic laws for biology?”, wseas transactions on systems, vol.5, no.11, pp.2503-2509. [9] karamanos, k., nikolopoulos, s., hizanidis, k., manis, g., alexandridi, a. and nikolakeas, s. (2006), “block entropy analysis of heart rate variability signals”, int. j. bif. chaos, vol.16, no.7, pp.2093-2101. [10] khinchin, a.i. (1949), mathematical foundations of statistical mechanics, dover publications. [11] mistakidis, i., karamanos, k., and mistakidis, s. (2013), “statistical versus optimal partitioning for block entropies”, kybernetes, vol.42, no.1, pp.35-54. advances in systems science and applications (2014) vol.14 no.3 253 [12] nicolis, g. (1995), introduction to nonlinear science, cambridge university press. [13] nicolis, g., and gaspard, p. (1994), “toward a probabilistic approach to complex systems”, chaos, solitons and fractals, vol.4, no.1, pp.41-57. [14] ebeling, w., and nicolis, g. (1992), “word frequency and entropy of symbolic sequences: a dynamical perspective”, chaos solitons and fractals, vol.2, pp.635-650. [15] steuer, r., molgedey, l., ebeling, w. and jimenez-montano, m.a. (2001), “entropy and optimal partition for data analysis”, eur. phys. j. b, vol.19, pp.265-269. [16] nicolis, g., nicolis, c., and nicolis, j. s. (1989), “chaotic dynamics, markov partitions, and zipf’s law”, journal of statistical physics, vol.54, no.3, pp.915-924. [17] nicolis, j. (1991), “chaos and information processing: a heuristic outline”, world scientific. [18] nikolopoulos, s., kapiris, p., karamanos, k., and eftaxias, k. (2004), “a unified approach of catastrophic events”, natural hazards and earth system science, vol.4, no.5/6, pp.615-631. [19] sammis, c. g., and sornette, d. (2002), “positive feedback, memory, and the predictability of earthquakes”, proceedings of the national academy of sciences of the united states of america, vol.99(suppl 1), pp.2501-2508. corresponding author author can be contacted at: sarris@teiath.gr. advances in systems science and applications (2017) vol.17 no.1 40 control cost in adaptive systems with an identifier under disturbance description uncertainties alexander bunich v.a. trapeznikov institute of control sciences of the russian academy of sciences, 117997, profsoyuznaya street, 65, moscow, russia; e-mail: bunfone@ipu.ru abstract. a problem of the criterion optimization of a linear-quadratic system with a stationary disturbance and set transfer function on control is considered. to solve the problem, a regularized method of substitution of estimates of the spectral density of the disturbance into the control cost functional is proposed. using the non-regularized method of substitution is shown to be able to lead to deflate the control cost. key words: linear-quadratic system, control cost, spectral smoothing of disturbances, substitution method. 1. substitution method and control cost a natural way of synthesis automatic systems with incomplete initial description is a reduction to a completely defined analog with substitution in the control law of corresponding empirical estimates instead of unknown parameters, calculated by an identifier. the identifier, as a sensor of parametric plant disturbances, may also be used to solve auxiliary problems, for instance, to predict slow faults, what demonstrates a universality of systems with an identifier. however, a justification of the substitution method, whose name was taken from the statistical estimation theory [1], requires considerable efforts. in the first turn, the optimization abilities are limited by two-block structures with separation of processes of forming the controls (identifier and controller with tuned parameters). generically, restricting the class of strategies increases the control cost in comparison to the optimal (dual [2]) strategy. besides that, the bayesian scheme of recalculating a posteriori parameters distributions over past observations requires setting a priori distributions of these parameters. finally, it is necessary to modify this scheme for problems with infinite control horizon [3]. the forced substitution of the initial problem statement of the synthesis with an asymptotic version implies considerable losses (a detailing of the term of ''asymptotic'' is needed: under increasing a sample, normalized risks may converge almost surely, in probability, in mean). formally, the synthesis problem (of dual control) for a plant with a parametric uncertainty relates to the stochastic optimal control of a completely defined plant with partially observed state ),( tt xx  , extended by partial inclusion of unknown plant parameters  , and involves calculating a priori distributions of the parameter )/( t u ydp  with respect to the observations t t yyy ,...,( 1 ), where the sub-index u underlines the dependence of these distributions on selecting a control strategy )},({ 1tt t uyu (in problems with a scalar criterion, the class of selectors is enough). if the strategy selection does not influence the rate of studying (passive feed-back), then the filtering algorithm admits 41 a. bunich: control cost in adaptive systems with an identifier under disturbance description uncertainties decomposition with emphasizing the parameters estimation in a functionally autonomous block (the identifier). computational problems of the dual control are supplemented with inadequate assumption on known a priori distribution of the parameters, characterizing the so called ''a priori hardness''. a natural way to overcome this hardness is the asymptotic statement of the synthesis problem. losses of the substitution are considerable. firstly, there are not distinguished asymptotically optimal systems with different transient quality, so the problem statement needs detailing. secondly, since a problem with infinite control horizon is considered, then a new constraint on the admissible strategies appears, which enables the closed-loop system stability, having no finite analogs. in systems with an identifier, investigating the stability is implemented within the method of ''frozen coefficients''. correct using the heuristic principle of ''frozen coefficients'', excluding effects of the parametric resonance, is possible under artificial decreasing the rate of controller retune with respect to the identification rate). asymptotically optimal strategies may be constructed without an identification by use of the direct method, as well as in systems with an identifier, when for a closed-loop plant conditions of the consistent estimation are violated (for instance, the gauss-markov conditions for the least squares method (lsm) [4]). let us be distracted from a more detailed discussion of the place and role of the identification in control systems. we will be restricted by more weak purpose of the criterion optimization [5] (the criterion optimization problem is always solvable, involving cases, when the lower bound of the control cost in the class of admissible controllers is not achieved, and the problem of the argument optimization is unsolvable). let us adopt the following assumptions. the plant is linear time-invariant, all variables are scalar and measured without disturbances; the transfer function of the plant on control is known; the disturbance 0}{  ttvv is centered stationary gaussian random process with the spectral density of an unknown order ,ks for which a consistent estimate ,...}2,1,{*  nkss n is known by use of the input/output observations  0},{ ttt yu , where k is a cone of spectral densities of stationary random processes of the autoregression/moving average with the 1l -metrics. the discrete-time ,...1,0t plant is described by the equation ttt vubya  )()( , 1)0( a with known coefficients of co-prime operator polynomials in the backward one time step shift operator . the initial data are disturbance independent random values with finite second moments. the controller (output feed-back) tt yu )()(   , 1)0(  is admissible, if the characteristic polynomial of the closed-loop system bag   is stable (the schur one). the control cost is the output variance in the steady-state mode. it turns out, that under accounting even the simplifications listed the criterion optimization is rather complex and the ''natural'' substitution method may decrease the control cost. let n be the completion of .k the cone n includes spectral densities of any stationary disturbances (indeed, k contains polynomials, in particular, convolutions of any density from n with the fejer kernels of any orders, forming the approximative unit in 1l ). thus, k may be advances in systems science and applications (2017) vol.17 no.1 42 considered as a class of parametric approximations of gaussian stationary random processes with absolutely continuous spectrum. on n , a concave functional ,||inf)( 20    dbfwi y f     is defined, where 0 yw is the transfer function from the disturbance to the output of the system with some fixed admissible controller,  is the ring of rational functions without poles in the closed disk }.1|{| z due to the completeness of the affine parameterization of control systems by the functional parameter f [6] kssi ),( is the control cost, that is the lower bound of the control cost over the class of all admissible controllers. the substitution method to determine the control cost )(si recommends the estimate *).(si the recommendation is justified, if the cost functional is uniformly continuous: in this case, to small errors of estimating the spectral density, small errors of the estimation by use of the substitution method correspond. however, the cost functional meets only the more weak condition of the semi-continuity from below [7]. indeed, let us continue uniformly the continuous cost functional from k on .n from the definition of the functional cost and the runge theorem it follows that 0)( si for any densities taking zero values on some arcs of the circumference 1|| z , and the set of such densities is dens in ,n so .0)( si however, in contrast to the last said, the cost is strictly positive for any density ks and the assumption on the uniform continuity is a delusion. in a number of cases, the designer possesses information on the spectral make-up of the disturbance in the form of the inclusion s with an a priori given class k . this information may be set by the inclusions ns . the paper is organized as follows. in section 2, a class (of degenerated problems) is emphasized. since the degeneracy is equivalent to the disturbance singularity, the control cost is zero. in section 3, an example of decreasing the control cost under using the substitution method (a consistent estimate of the spectral density of a regular disturbance) is presented. in section 4, a formula for the control cost is presented, and a procedure of the regularization of calculating the cost with using substitution of the metrics in the space of spectral densities of disturbances is proposed. required reference information was presented in section 1. 2. degenerate problems synthesis problems with zero cost will be referred as degenerate. let us resemble that the factorization method assumes the process to be regular, that is control system performance under the conditions of stationary inflow of new (renovating the disturbance) information from outside. a stationary process is regular if and only if its spectral density exists and has no ''deep zeros'' in the sense of convergence of the logarithmic integral: .ln 1ls this a rather weak constraint does not even guarantee the fourier series convergence for sln on a set of a positive measure (an example of a.n. kolmogorov), in computational practice it is conventionally substituted by considerably more strict one of existing the spectrum (absolute convergence of this series). correspondingly, the complex problem of the density s factorization is reduced to the separation for sln . a technical way of reducing the density factorization to separating its logarithm is complicated by the fact that the separation operator or ''natural projecting'' p ,  kj n k k kj n nk k eceсp            1 || , 43 a. bunich: control cost in adaptive systems with an identifier under disturbance description uncertainties is not bounded in the 1l -norm. the high prefilter sensitivity to variations of the spectral make-up of the disturbances indicates that 1l -norm is not an adequate closeness measure; and possible way of regularization of synthesis of wiener controllers is substitution of the metrics. from the category point of view, the regularity property is not typical: spectral densities meeting the regularity criterion forms in the space of all spectral densities the set of the first baire category. the regularity (singularity) property is preserved under transforming the process by a stable filter. the factorization solves the inverse problem: determining the transfer function of the prefilter via the spectral density of the disturbance. these well known results admit a generalization to multivariate processes of a constant rank (with substitution in the criterion ss det ). for more general classes of stationary processes, effectively verified regularity criteria are not known, while determining the prefilter via the spectral density is a complex problem. a general theory of singular processes perhaps does not exist, however enough simple way is available to transfer from singular processes to regular ones by adding (mixing) the white noise of a small power (what is interpreted by ''measurement noise''). to pursue the purpose, one determines a prefilter that factorizes the density s with small 0 , reducing the case of singular disturbance to regular one (non-formal question of selecting  is solved experimentally). under a fixed spectral density s , the control cost does not increase monotonically under as ,0 while solving the argument optimization problem is extremely non-regular and for the limit system (as )0 does not exist (let us resemble that for a singular disturbance the criterion optimization problem has zero solution). forming control has a dual function: directing (minimizing the action risk) and studying (minimizing future risks due to an uncertainty non-removed) [6]. the adaptive systems with an identifier (asi) are emphasized with their ability to parallelize these functions, what is seen from the structure asi scheme with a functionally autonomous block, the identifier, calculating in the real time plant parameters estimates and controlling another block, the tuned controller, constructed at the stage of the synthesis of the main control system loop. justification of the identification (indirect) approach to the synthesis problem is reduced to the fact, that if the control cost does not depend on transients in the identifier, then substitution of unknown plant parameters with their consistent estimates will not enlarge the control cost. 3.decreasing the control cost let us show, by use of an example of a minimum-phase plant, that the substitution method may decrease the control cost [7]. in this example, the control cost coincides with the error variance of the optimal one step ahead prediction. the example idea is to approximate a density with the ''deep zero'' by spectral densities of regular processes. for a positive density ks and a small number 0 , we will set s ( )()  s under ),(   and )exp()( 2 s under ).,(   the functions 1)( ls  are bounded from zero and so ,)( rs  where r is the class of spectral densities of regular processes. in accordance to the construction, ss l1  as ,0 however advances in systems science and applications (2017) vol.17 no.1 44     dsln as 0 , and by virtue of the szegő formula [8] ,0)( 0 2     where )(2  is the error variance of the optimal one step ahead prediction for the process with the spectral density s , coinciding, due to the minimum-phase plant property, with the control cost. thus, for some family of regular processes the following conditions hold: , 1 ss l  rs  (1a) ),(0)(lim 0 sisi     (1b) hence, the control cost takes as small as needed values in any neighborhood of the point .s one may prove that the latter affirmation is valid also without the constraints of minimum-phase and under any delays in the measurement. lemma on decreasing the control cost. for any spectral density ks , there exists a family of densities from k , meeting conditions (1). the lemma name is explained by the inequality 0)(inflim   ikv < ),(si (2) where v is any neighborhood of the point .ks due to this inequality, 1l -metrics is not an adequate measure of closeness of disturbances in the spectral make-up, what is expressed by the incorrectness of the criterion optimization problem, possible decreasing the control cost, defined by the substitution method, in comparison to its magnitude under exactly known spectral distribution density. the problem regularization is possible by using additional information, or by introducing a new metrics in the class of spectral densities of regular disturbances. 4. calculating the cost and solution regularization the control cost )(si has a simple geometric interpretation: it is the square of the distance from the point 201 lhwb yi  to 2h subspace. its calculation enables one to regularize the criterion optimization problem by the metrics substitution: in the new metrics, the cost functional is slowly changed under small variations of the spectral make-up of disturbances. to calculate the control cost, let us introduce the canonical (interior-exterior factorization of the polynomial eibbb  , factorization *)2( 1hh of the spectral density ,s the distance )( from the point 2l to 2h subspace. let us fix some admissible (support) controller and designate 00, uy ww the transfer functions of the control system with the support controller. lemma. the control cost is defined by the formula )()( 012 hwbsi yi  . (3) indeed, the distance d from the point hwy 0 to the 2hbi subspace is equal to 2 0 2 | || |inf fbhw iyhf   , and due to the unimodular property of ib , the equality holds )( 0122 hwbd yi  . (4) but hbe is an exterior function, so the linear manifold hbe is dense in 2h , bh is dense in в 2bh , and 2d ,| || |inf 2 2whmw (5) 45 a. bunich: control cost in adaptive systems with an identifier under disturbance description uncertainties where m 0 yw .ib the value 2 2| || |wh is the control cost of the control system with the spectral density s of the disturbance, and due to the completeness of the class m , 2d is the control cost, what was to be proven. remark. as a support regulator, one may select the h -optimal regulator. on the code r of spectral densities of regular processes, let us introduce a new metrics ,| || |),( 22121 hhss  where 2,1h are solutions of the problem of the factorization of the densities rs 2,1 with norming 0)0( h . it is easily to verify, that the mapping introduced as is a metric: ),0[: rr . let ssn  r , for solutions hhn , of the problem of the factorization of corresponding densities , 2 hh h n  but  )(|||| ssssss nnn ,)(|| sshh nn  (6) where from due to the cauchy inequality, we will obtain: chhss nn 2 2 2 1 | || || || |  (7) with a large enough constant c . the limit transfer as n proves the following implication: ssn  r . 1 ss l n however, the inverse affirmation is mistaken. let, for instance, the density s is a singular process density, while ns is its polynomial approximation, then .rs the new metrics regularizes the criterion optimization problem, and in the new metrics on the class of spectral densities of regular disturbances the functional )(2/1 si is a lipschitz one. let us underline, that the metrics introduced does not permit a continuation to more wide classes of densities, in contrast to the 1l -metrics, under definition of which the regularity condition is not pre-assumed. for rather poor classes of  , the regularization is not required. example. let  admit embedding in the simplex interior (parametric uncertainty). the continuity follows from the concave property. references [1] borovkov a.a., mathematical statistics, nauka publ., moscow, 1984. (in russian) [2] feldbaum a.a., foundations of the theory of optimal automatic systems, nauka publ., moscow, 1966. (in russian) [3] fomin v.n., methods of control of linear discrete-time plants, lgu publ., leningrad, 1985. (in russian) [4] kogan m.m. and neymark yu.i., “the identifiability of locally-optimal adaptive control laws under indirect observations”, automation and remote control, vol. 51, no. 1, 1990, pp. 53-62. [5] tsypkin ya.z., “the optimality in problems and methods of the advanced control theory”, vestnik an sssr, no. 9, 1982, pp. 116-121. [6] fomin v.n., fradkov a.l., and yakubovich v.a., adaptive control of dynamic plants, nauka publ., moscow, 1981. (in russian) [7] bunich a.l., “degenerate linear-quadratic problem of discrete plant control under uncertainty”, automation and remote control, vol. 72, no. 11, 2011, pp. 2351-2363. [8] hoffman k., banach spaces of analytical functions, prentice hall, englewood cliffs, n.j., 1962. advances in systems science and application (2015) vol.15 no.1 90-98 the birth triangle: a new approach to study the birth of systems j.-f. vautier french society of systems science (afscet), paris, france abstract this paper focus on the birth of systems. it starts with a physical model which provides some necessary conditions to the birth of a specific system: the fire. next, a more general model: the birth triangle is presented. similarly to the fire triangle, in the birth triangle two objects interact together when there is an activator. the notion of environment, in which the two objects are included, is also introduced. different examples of application of this model from physics to human organizations are examined: conglomerates of matters, molecules, cells and markets. the first main interest of the birth triangle is to propose ways to induce the birth of systems and also to stop their development. the second main interest of this model is to propose a new framework for a triggering condition and thus a way to revise the question of causality of different factors. indeed, with this model, it is the triangle which is causal (the set of the necessary factors and their conjunction). this new vision permits to avoid some quarrels to know which factor is the most responsible for something and some debates on which is the first inducer of an effect. keywords fire triangle, birth, system, field 1 introduction a lot of studies deal with the behavior of systems, the occurrence of some specific properties like for example emergences (properties of systems which cannot be observed in the elements [1]) or the collapse of systems [2]. considering the systemic methods, the same remarks can be done: a lot of methods aim to represent the behavior and/or the collapse of systems. the system dynamics, for example, belongs to this kind of methods [3]. but there are few studies about the birth of systems. then, this article will be focused on this topic. moreover, a lot of studies deal with causes [4-6]. some methods try in particular to find variables that most influence the others [7-9]: the control variables. indeed, the question is often: what are the variables which induce variations of the other ones? then, this article will propose a specific answer to this question related to the factors which cause the birth of systems. finally, this work starts with a physical model which provides some necessary conditions to the birth of fire (which is considered as a system). this model is the fire triangle. it was adapted to get a more general one: the birth triangle. then, the birth triangle will be firstly described. next, examples of application of this model from physics to human organizations will be provided. finally, a advances in systems science and application (2015) vol.15 no.1 91 discussion will be presented. 2 the birth triangle the fire triangle it is proposed to study the question of the birth of systems with a metaphor: the birth of fire. the common representation used to explain this birth is called the fire triangle. in a few words, this representation indicates that the birth of a fire depends on the conjunction and the amount of three factors which have to be in the same place at the same time: something to burn: a fuel, another factor which is necessary for the chemical reaction with the fuel: an oxidizing agent (usually oxygen), an activating energy which is necessary to induce the reaction between the two previous factors. the different components of the birth triangle similarly to the fire triangle, in the birth triangle two objects interact together when there is an activator. the notion of environment, in which the two objects are included, is also introduced here. fig.1 the birth triangle in this model, objects are generally “physical” and may have different sizes: two molecules (of fuel and oxygen), two people, two towns, two countries· · · there may be also virtual objects like the cells of conway’s game of life [10]. but objects are not qualities or emotions like gladness, sadness· · · . the activator induces a movement of one object closer to the other one in such a way that a system (i.e. two objects grouped and interacting together) may exist after (this movement). the activator is a field which may be a: 92 j.-f. vautier: the birth triangle: a new approach to study the birth of systems ... pushing field. object 2 moves closer to object 1 due to an internal energy (in object 2) or the two objects are pushed toward each other due to an external energy. this field is related to: · intrinsic pushing energies: if the objects are living beings, it means that only object 2 moves closer to object 1 (a prey tries to run away when the predator attacks). then, when a predator attacks a prey, a system of pursuit may be created; and/or · extrinsic pushing energies: heat (for example the birth of fire), mechanical energy (for example to make a mayonnaise) or pulling field (in the meaning of attraction). the two objects are pulled toward each other due to an internal energy (in the two objects) or due to an external energy. this field is related to: · intrinsic pulling energies of the two objects (leading to a mutual attraction). for example, two boxers move closer to constitute a pair after the ringing of bell; and or · extrinsic pulling energies: introduction of new entities in the environment like a catalyst. in a few words, energies are what permit the existence of the fields (which induce a movement of one object closer to another one). according to this general definition, energies are not only physical ones. they may correspond also for example to wishes of people. moreover, for living beings or human organizations, this latest kind of energy may appear to create a movement after an event, an accident, an information, a chance meeting. for example, the murder of archduke franz ferdinand of austria in 1914 was an event which leads to the start of the first world war. the wishes of fighting appeared (even if they were “fed” by grievances existing previously). the european peoples constituted in 1914 a system of belligerents; the environment supports the movement of moving one object closer to the other one. it means that there are also some characteristics of the environment that maintain together the two objects (to be able to interact to build a system afterwards): a small space inside a bottle or a bowl a place, a town which “maintain” people together. it means also that there are not obstacles in the environment between the objects. this may result from the removal of these obstacles, previously present in the environment, removal which may be caused by action of an entity (a human for example) or not (the natural removal of the link between a fruit and its tree). in the birth triangle, only the interactions between object 1 and object 2 are considered. object 1 is the entity on which the focus is. object 2 may consist of advances in systems science and application (2015) vol.15 no.1 93 several other entities which interact with object 1. moreover the question of the shape or the form of the set of these entities which interact in order to build the system is not considered in this article. 3 examples of birth of systems: from the physics to the human organizations conglomerates of matter in this case, the physical forces of gravitation are only taken into account. they attract the objects toward each other. when an apple falls down on the ground, this phenomenon can be described with the birth triangle: object 1: the earth; object 2: an apple; environment: the space around these two objects; activator: a pulling field related to gravitational forces in the same way, two magnets attract (pulling field) or repel (pushing field) each other according to their nature, orientation and distance from one to each other i.e. according to the field related to electromagnetic forces. “life bricks” let us consider now the chemical domain (the fire belongs to this domain). s. miller succeeded to produce some molecules: the amino acids which are often considered as the “life bricks”[11]. object 1: methane (ch4); object 2: ammonia (nh3) and hydrogen (h2); environment: aqueous solution in a bowl; activator: a pushing field (electric shock). cells concerning the process of fecundation: object 1: an ovule; object 2: a spermatozoid; environment: liquid solution in a small space; activator: for in vitro fecundation, there is pushing field resulting from an electric stimulation or mechanical push (a sting with a pipette). for natural fecundation, there is a pulling field. ovule and spermatozoids move closer toward each other. markets let us consider a market as a set of economic actors which sell and buy different kind of products: object 1: a company; 94 j.-f. vautier: the birth triangle: a new approach to study the birth of systems ... object 2: its clients; environment: a place to support the interactions between these two objects i.e. the transfer of products between a selling company and a buying actor; activator: a pulling field due to, for example, the occurrence of new offers of products from the company and/or new demands from the clients. 4 discussion several points are presented. first of all, this discussion deals with some ways to induce the birth of a system or to stop its development. next, the focus is on one characteristic of the objects: the fertile soil. finally, the notion of causality is questioned. ways to induce the birth of a system or to stop its development according to the birth triangle, if there is a birth of a system then a pushing or a pulling field exists. it means that if these kinds of fields are identified somewhere, then a system may exist if there are an environment and sufficient fertile soil in the objects. it is a way to look for the existence of some potential new systems. according to this model, what should be done, in practice, to induce the birth of systems? the way is often to: change the environment, for example, by introducing new entities (catalysts), or energy (e.g. heat), or removing distances, obstacles· · · between the objects in the environment and/or increase in object 1 and 2 the potentiality of interaction with the other one. for example, let us consider: object 1: a baker who also sells some kinds of cakes; object 2: the clients for cakes; environment: the location of the bakery; activator: a pulling field. then to induce the purchase of the cakes, it could be important to work, respectively, on: the quality of the cakes, their colors, their shapes· · · ; the needs of cakes from the clients (a way to increase them may be an advertising campaign for example); all information to locate the bakery· · · in order to increase the foot traffic to the storefront; if there are not enough clients, a special event may be organized with promotions and low costs or prepared like during the days of christmas or epiphany. this example shows that this birth triangle is in fact already used in marketing· · · even if people did not know it. on the other hands, the birth triangle may be also advances in systems science and application (2015) vol.15 no.1 95 a way to observe a marketing mix strategy with a new framework of examination. moreover, a hypothesis is proposed: it is possible to use also the birth triangle to stop the development of a system. it means that the different factors of the birth triangle could be pointed out during the development of the system. if we consider, for example, the problem of irregular migrants, a lot of countries try to prevent these migrants from crossings their borders. according to a birth triangle view, these actions deal with the environment in putting obstacles between the objects. nevertheless, some people propose also other approaches, for example trying to decrease the pushing field by improving the standard of living of people in their country with investments· · · then, in this way, the birth triangle, as a generalization of the fire triangle, may propose some ways to stop the development of systems, in particular, in decreasing the level of the different factors to remove the conjunction. importance of the “fertile soil” in the objects to create a system from interactions between object 1 and object 2, there must be already some “fertile soil” in the objects. for physical and chemical domains the fertile soil is in fact a basic property of the objects (the mass, the atomic composition· · · ) which will not change when the system will be made for human organizations like for example a market, it is different. its birth needs: sufficient “fertile soil” in the selling company i.e. some people who can answer the demands and understand the wishes, the needs expressed by the potential clients, who can provide some specific products for clients and this fertile soil may exist, in the company, in a domain but not in another one. in particular, it consists of top managers who think that their company has to answer the demands, has to go in this direction and who can provide resources and decide the development of teams in specific domains otherwise the response of the company will be certainly: “it is not my business”; sufficient “fertile soil” in the clients i.e. some people who can provide “demands” to the potential selling company, top managers of a buying company who think they may send these demands to this selling company. it concerns also the capacity of changing its buying habits by purchasing new products. an example: a seed may fall down from the tree to the ground but it will grow only if there is sufficient fertile soil in the ground (object 2: water). if the seed is too young (and then not ready to sprout) or if there is a stone (environment) or no possibility for water to enter the seed (the activator: a pulling field), there will be not sufficient conditions to the birth of the plant. besides, the seed may stay sometimes a long time in this state if there is not a chance meeting with a suitable environment 96 j.-f. vautier: the birth triangle: a new approach to study the birth of systems ... another example: if a problem occurs, for example a shortage of a product (like during the oil crisis of 1973), it may induce a pushing field from the clients of a company which sells this kind of products. but a system will appear only if there are sufficient fertile soils in object 1 (enough products to sell) and object 2 (enough money) and an environment suitable for development (without obstacles). then, this view, from the birth triangle, contrasts with another view which indicates that a problem, an accident, an event would be the triggering factor of modifications of the organizations. according to the birth triangle, an event, like the murder of archduke franz ferdinand of austria in 1914, may induce some changes in the objects (concerning for example the wishes of people to fight together) which may create, in this case, a pulling field but the birth of a system needs also some fertile soil in the objects (for example the resources necessary for the fighting). finally, the question may be even more complex for certain kinds of human organizations. indeed, the necessary level of fertile soil in object 1 may depend on the environment. for example, to own a research department or not in its company may be a weakness or not. that depends often on the other companies of this market or those which might enter this market. then, let us note that this birth triangle proposes some conditions to the birth of systems (factors and their conjunction). for physical process, the birth of a system is more predictable than for human organizations since, in this latest case, we cannot know often precisely the fertile soil of object 1 or 2 which is necessary. a causality based on the conjunction of necessary conditions the birth triangle proposes a new framework for a triggering condition and thus a way to revise the question of causality of different factors. with this model, it is the triangle which is causal (the set of the necessary factors and their conjunction). furthermore, we are not in a cumulative vision in which the value of a factor can compensate the value of another one (it is, for instance, the propriety which underlines the possibility to calculate a mean value). we enter in a more complex reality in which no component can compensate the lack of another one. for example, in order to create a market, the needs for the product are important but the capacity of the clients to pay for the product is naturally important too· · · this new vision permits often to avoid some quarrels to know which factor is the most responsible for something and some debates on which is the first inducer of an effect. 5 conclusion let us note that, in the birth triangle, the efficiency of each factor depends on the other factors and a suitable conjunction between the three factors is necessary to advances in systems science and application (2015) vol.15 no.1 97 induce the birth of a system. and to end this article, here are some perspectives of works for the future: the size of systems does not seem to be linked to a pushing or pulling field.is it true the more the systems are large the more the objects seem to be modified to induce the birth of systems and also after when the systems work. does a relationship exist really between the size of systems and the amount and the nature of the modifications? is there a possibility to combine some birth triangles together? for example, may an object be represented as a triangle? in a causal analysis, could it be possible to use the birth triangle to help to identify and to go back to the root causes of an event, a problem? if we consider a fire, there is an auto-activation of the reaction until sufficient fuel or oxygen are present in the environment of this system i.e. until new elements can support the combustion regularly. in other words, it could be represented as a continuous birth of the system. is this kind of functioning may be observed in a few or in a lot of systems? finally, it could be certainly interesting to better define the limits of this formalization. for example, it seems to be important that the two objects are different to build a system, like in a musical chord, the necessary difference between two notes. but is there a minimal difference? references [1] l. von bertalanffy. (1968),general system theory: foundations, development, applications, george braziller, new york, usa. [2] j. diamond. (2005),how societies choose to fail or succeed, collapse, penguin books, newyork, usa. [3] p. senge. (1990),the fifth discipline: the art and practice of the learning organization. [4] j.-f. vautier. (2014), “a causal contextualization based on the four causes of aristotle”, 9th congress of the european union for systemics (eus-ues), valencia, spain. [5] s. anderson. (2009), “root cause analysis: addressing some limitations of the 5 whys”, http://www.qualitydigest.com/inside/fda-compliancenews/root-cause-analysis-addressing-some-limitations-5-whys.html. [6] t.minoura. (2007), “talks about problems with 5-whys”, http://www.taproot.com/archives/710. 98 j.-f. vautier: the birth triangle: a new approach to study the birth of systems ... [7] j.-f. vautier. (1994), “structural analysis of a man-machine system (samms): presentation of the method”, xiith congress of the international ergonomics association (iea) toronto vol. 4, pp. 59-60. [8] a. dassens and r. launay. (2008), “systemic study of the risks analysis: presentation of a general approach”, systemic study of the risks analysis, ag 1585, editions t.i. a. dassens and r. launay. (2008), “etude systémique de l’analyse de risques :the original is in french.) [9] m. godet. (1994), from anticipation to action: a handbook of strategic prospective , unesco publishing. [10] m. gardner. (1970), “mathematical games. the fantastic combinations of john conway’s new solitaire game ‘life’ ”, scientific american,no.223,pp.120-123. [11] s. miller. (1953), “a production of amino acids under possible primitive earth conditions”, science, vol.117, pp.528-529. corresponding author j.-f. vautier can be contacted at:jean-francois.vautier@cegetel.net advances in systems science and applications (2014) vol.14 no.2 129-143 adaptive cfar tests for detection and recognition of target signals in radar clutter konstantin n. nechval1 and nicholas a. nechval2 1applied mathematics department, transport and telecommunication institute lomonosov street 1, lv-1019, riga, latvia 2statistics department, evf research institute, university of latvia raina blvd 19, lv-1050, riga, latvia abstract in this paper, adaptive cfar tests are described which allow one to classify radar clutter into one of several major categories, including bird, weather, and target classes. these tests do not require the arbitrary selection of priors as in the bayesian classifier. the decision rule of the recognition techniques is in the form of associating the p-dimensional vector of observations on the object with one of the m specific classes. when there is the possibility that the object does not belong to any of the m classes, then this object is to be classified as belonging to one of the m classes or to class m+ 1 whose distribution is unspecified. the tests are invariant to intensity changes in the clutter background and achieve a fixed probability of a false alarm. the results obtained in this paper agree with the simulation results, which confirm the validity of the theoretical predictions of performance of the suggested adaptive cfar tests. keywords radar clutter, target signal, detection, recognition, adaptive cfar tests 1 introduction modern air traffic control radar systems rely heavily on automatic target detection and tracking to maximize air traffic safety. moving target indicator and moving target detector algorithms achieve good target detection performance through the suppression of most or all forms of radar clutter. unfortunately, real-time information on airborne hazards to aircraft, such as birds and storm systems, is also suppressed. the ability to classify clutter and hence identify these hazards can thus contribute significantly to air traffic safety. the process of classification can be formalized as follows. the unprocessed radar data are passed through a feature extractor, which transforms the available data samples into a set of separable features. these features are derived from the reflection coefficients computed using the multisegment version of burgs formula[1]. the aforementioned coefficients (that contain all spectral information, including the mean doppler shift) are then transformed and grouped to satisfy the requirements for multivariate gaussian behaviour. only information that is different from class to class is maintained, and in such a form that a reliable decision, based on a discriminant function derived from the above features, may 130 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... be made. the classification problem consists in the following. there are m classes (populations), the elements (objects) of which are characterized by p measurements (features). next, suppose that we are investigating a certain object on the basis of the corresponding p measurements. we postulate that this object can be regarded as a random drawing from one of the m populations but we do not know from which one. we suppose that m samples are available, each sample being drawn from a different class. the elements of these samples are realizations of p-dimensional normal random variables with unknown parameters. after a sample of p-dimensional vectors of observations on the object is drawn from a class known a priori to be one of the above set of m classes, the problem is to infer from which class the sample has been drawn. the decision rule should be in the form of associating the sample of observations on the object with one of the m samples and declaring that the object has come from the same class as the sample with which it is associated. when there is the possibility that the object does not belong to any of the m classes, then this object is to be classified as belonging to one of the m classes or to class m+ 1 whose distribution is unspecified. stehwien and haykin [2] solved the problem of statistical classification of radar clutter in a bayesian framework. in this paper, the problem is treated in a nonbayesian setting. a classification technique is described which allows one to classify radar clutter into one of several major categories, including bird, weather, and target classes. this technique is based on applying the theory of generalized maximum likelihood ratio testing for composite hypotheses. the unknown parameters are then estimated using maximum likelihood estimators. this approach does not require the arbitrary selection of priors as in the bayesian classifier. yet the generalized likelihood ratio test (glrt) is widely preferred because of its nice asymptotic (large sample size) properties such as consistency, unbiasedness, and constant false alarm rate (cfar). it is also called the uniformly most powerful invariant (umpi) test since it exhibits the ump property among the class of tests that are invariant to a natural set of transformations. the asymptotic performance of the glrt becomes equivalent to the test with perfectly known parameters. the main feature of the proposed classification technique is the class elimination rule. when certain conditions are met, the decision is taken to eliminate specific class from further considerations, and the classification process is continued with a reduced number of classes. the class elimination rule is based on the generalized likelihood ratio. the outline of the paper is as follows. a problem of signal detection in clutter is considered in section 2. section 3 is devoted to a problem of target signal recognition. advances in systems science and applications (2014) vol.14 no.2 131 2 signal detection in clutter the problem of detecting the unknown deterministic signal s in the presence of a clutter process, which is incompletely specified, can be viewed as a binary hypothesis-testing problem. the decision is based on a sample of observation vectors xi = (xi1, ..., xip) ′, i = 1(1)n, each of which is composed of clutter wi = (wi1, ..., wip) ′ under the hypothesish0 and a signal s = (s1, ..., sp) ′ added to clutter wi under the alternativeh1, where n > p. the two hypotheses that the detector must distinguish are given by h0 : x = w(clutteralone) (1) h1 : x = w + cs ′ (signalpresent) (2) where x = (x1, ...,xn) ′ (3) w = (w1, ...,wn) ′ (4) are n > p random matrices, and c = (1, ..., 1) ′ (5) is a column vector of n units. it is assumed that wi, i = 1(1)n, are independent and normally distributed with common mean 0 and covariance matrix (positive definite) q, i.e. wi ∼ np(0,q), ∀i = 1(1)n. (6) thus, for fixed n, the problem is to construct a test, which consists of testing the null hypothesis h0 : xi ∼ np (0,q) , ∀i = 1(1)n. (7) versus the alternative h1 : xi ∼ np (s,q) , ∀i = 1(1)n. (8) where the parameters q and s are unknown. remark 1. characterization of the multivariate normality is given by the following theorem. theorem 1 (characterization of the multivariate normality). let xi, i = 1(1)n, be n independent p-multivariate random variables (n ≥ p+2) with common mean bfa and covariance matrix (positive definite) bfq. let zk, k = p+2, , n, be defined by zk = k − (p+ 1) p k − 1 k (xk − x̄k−1) ′s−1 k−1 (xk − x̄k−1) 132 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... = k − (p+ 1) p ( |sk| |sk−1| − 1 ) , k = p+ 2,...,n, (9) where x̄k−1 = k−1∑ i=1 xi/(k − 1), (10) sk−1 = k−1∑ i=1 (xi − x̄k−1)(xi − x̄k−1) ′, (11) then the xi (i = 1, . . . , n) arenp (a,q) if and only if zp+2, ..., zn are independently distributed according to the centralf distribution with p and 1, 2, ...n− (p+ 1) degrees of freedom, respectively. proof. the proof is similar to that of the characterization theorems [3-4] and so it is omitted here. 2.1 goodness-of-fit testing for the multivariate normality the results of theorem 1 can be used to obtain test for the hypothesis of the form h0 : xi follows np (a,q) versus ha : xi does not follow np (a,q) , ∀i = 1(1)n the general strategy is to apply the probability integral transforms [5] of zk, ∀k = p + 2 (1)n, to obtain a set of i.i.d. u(0, 1) random variables under h0. under ha this set of random variables will, in general, not be i.i.d. u(0, 1). any statistic, which measures a distance from uniformity in the transformed sample (say, a kolmogorov-smirnov statistic), can be used as a test statistic. 2.2 gmlr statistic and its distribution one of the possible statistics for testing h0 versus h1 is given by the generalized maximum likelihood ratio (gmlr) gmlr=max θ∈θ1 lh1(x; θ) / max θ∈θ0 lh0(x; θ), (12) where q = (s,q) ,q0 = {(s,q) : s = 0,q ∈ qp},q1 = q − q0,q = {(s,q) : s ∈ rp,q ∈ qp}, qp denotes the set of p × p positive definite matrices. under h0, the joint likelihood for x based on (7) is lh0(x; θ)= (2π)−np/2|q|−n/2 exp ( − n∑ i=1 x′iq −1xi/2 ) ., (13) under h1, the joint likelihood for x based on (8) is lh1(x; θ)= (2π)−np/2|q|−n/2 exp ( − n∑ i=1 (xi − s)′q−1(xi − s)/2 ) . (14) advances in systems science and applications (2014) vol.14 no.2 133 it can be shown that gmlr= ∣∣∣q̂0 ∣∣∣n/2∣∣∣q̂1 ∣∣∣−n/2 , (15) where q̂0=x ′x/n, (16) q̂1= (x ′ − ŝc′)(x ′ − ŝc′)′/n, (17) and ŝ=x ′c/n (18) are the well-known maximum likelihood estimators of the unknown parameters q and s under the hypotheses h0 and h1, respectively. it can be shown, after some algebra, that (15) is equivalent finally to the statistic y = t ′ 1t −1 2 t1/n, (19) where t1 = x′c,t2 = x′x.it is known that(t1,t2) is a complete sufficient statistic for the parameter q = (s,q) thus, the problem has been reduced to consideration of the sufficient statistic (t1,t2). it can be shown that under h0 , the result (19) is a q-free statistic y which has the property that its distribution does not depend on the actual covariance matrixq. this is given by the following theorem. theorem 2 (pdf of the gmlr statistic y). under h0, the statistic y is subject to a noncentral beta-distribution with the probability density function (pdf) fh1(y;n, q) = [ b (p 2 , n−p 2 )]−1 y (p 2 ) −1 (1− y) (n−p 2 ) −1 ×e−q/2 1f1 (n 2 ; p 2 ; qy 2 ) , 0 h, then h1 (signalpresent), ≤ h, then h0 (clutteralone), (27) and can be written in the form of a decision rule u(v) over {v : v ∈ (0,∞)} , u(v) = { 1, v > h (h1), 0, v ≤ h (h0), (28) where h > 0 is a threshold of the test which is uniquely determined for a prescribed level of significance α so that sup θ∈θ0 eθ {u(v)} = α. (29) advances in systems science and applications (2014) vol.14 no.2 135 for fixed n, in terms of the probability density function (26), tables of the central f -distribution permit one to choose h to achieve the desired test size (false alarm probability pfa ), pfa = α = ∞∫ h fh0(v;n)dv. (30) furthermore, once h is chosen, tables of the noncentral f -distribution permit one to evaluate, in terms of the probability density function (25), the power (detection probability pd) of the test, pd = γ= ∞∫ h fh1(v;n, q)dv. (31) the probability of a miss is given by β = 1− γ. (32) it follows from (30) that the gmlr test is invariant to intensity changes in the clutter background and achieves a fixed probability of a false alarm, i.e. the resulting analyses indicate that the test has the property of a constant false alarm rate (cfar). also, no learning process is necessary in order to achieve the cfar. thus, operating in accordance to the local clutter situation, the test is adaptive. when the parameter q = (s,q) is unknown, it is well known that no the uniformly most powerful (ump) test exists for testing h0 versus h1 [8]. however, some hypothesis testing problems that do not admit ump decision rules (tests) nevertheless exhibit certain natural invariance properties [8-9]. these properties suggest restricting attention to a limited class of decision rules, viz., the invariant decision rules. it is then sometimes possible to derive decision rules that are ump within this limited class. in this sense, invariance is a concept of fundamental importance in hypothesis testing. the following theorem shows that the test (27) is umpi for a natural group of transformations on the space of observations. theorem 4 (umpi test). for testing the hypothesis h0(1) versus the alternative h1(2), the cfar test given by (27) is uniformly most powerful invariant (umpi). proof. the proof is similar to that of nechval [10] and so it is omitted here. a robustness property of the v-test can be studied in the following set-up. let x = (x1, ...,xn) ′ be an n × p random matrix with a pdf φ,let cnp be the class of pdf s on rnp with respect to lebesque measure dx, and let h be the set of nonincreasing convex functions from [0,∞) into [0,∞). we assume n ≥ p + 1. 136 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... for s ∈ rp and q ∈ qp define a broad class of pdfs on rnp as follows cnp(s,q) =  f ∈ cnp : f(x; s,q) = |q|−n/2η ( n∑ i=1 (xi − s)′q−1(xi − s) ) , η ∈ h  (33) in this model, it can be considered the testing problem h0 : ϕ ∈ cnp(0, q), q ∈ qp (34) versus h1 : ϕ ∈ cnp(s,q), s ̸= 0, q ∈ qp (35) and shown that v-test is umpi. clearly if (x1, ...,xn) is a random sample of xi ∼ np(s,q), i = 1(1)n, or x ∼ nnp(cs ′, in ⊗q), the pdf φ of x belongs to cnp (s,q) . further if f (x; s,q)) belongs to cnp (s,q), then g∗(x; s,q) = ∞∫ 0 f(x; s, rq)dg∗(r) (36) also belongs to cnp (s,q) where g∗ is a distribution function on (0,∞), and so cnp (s,q) contains the (np-dimensional) multivariate t-distribution, the multivariate cauchy distribution, the contaminated normal distribution, etc. [10-12]. here the following theorem holds. theorem 5 (robustness property). for problem (34)-(35), the cfar v-test is umpi and the null distribution of v is f -distribution with d.f.s p and n− p, i.e., the cfar test is still umpi in a broad class of distributions given by (33), and the null distribution under any member of the class is the same as that under normality. proof. the proof is similar to that of nechval [10] and so it is omitted here. 2.4 risk minimization for fixed n, in terms of the above probability density functions in (25) and (26), the probability of making the first type of wrong decision (false alarm probability) is found by α(h;n) = ∞∫ h fh0(v;n)dv (37) and the probability of making the second type of wrong decision (the probability of a miss) by β(h;n, q)= h∫ 0 fh1(v;n, q)dv. (38) advances in systems science and applications (2014) vol.14 no.2 137 any value of s will result in a value for q that is greater than zero. as the value of s increases, the value of q will also increase. a good detector is certainty expected to minimize α and β in some manner. for example, the neyman-pearson criterion defines optimality to be that of maximizing 1− β subject to the constraint that α ≤ α0, where α0 is a fixed constant between zero and unity. for this criterion, the optimum threshold can be found from (30). in general the structure of an optimum detector depends on the signal (or the signal-to-noise ratio). let us assume that a noncentrality parameter q representing the generalized signal-to-noise ratio (gsnr) is given. if we let wα and wβ be the unit weight (cost) of the probability of making the first type of wrong decision (α) and the probability of making the second type of wrong decision (β), respectively, then the optimal threshold of test, h∗, can be found by solving the following optimization problem (with respect to h): minimize r(h;n, q) = wαα(h;n) + wββ(h;n, q) (39) subject to h ∈ (0, 1), (40) where r (h;n, q) is a risk representing the weighted sum of the false alarm risk and the miss risk. it can be shown that h∗ satisfies the equation wαfh0(h ∗;n) = wβfh1(h ∗;n, q). (41) generally, the miss risk is more important that the false alarm risk, so that wa ≤ wb. if the sample size of observations, n is not bounded above, then the optimal value n∗ of n can be found as n∗ = inf n : ( α(h∗;n)+β(h∗;n,q) ≤ ϑ, h∗= arg min h∈(0,1) r(h;n, q) ) , (42) where ϑ is a preassigned value of the sum of the false alarm risk and the miss risk. 3 target signal recongnition suppose that the hypothesis h0: (clutter alone) is rejected. then the target (signal in clutter) classification problem using the target identity information consists in the following. let the target signal belong to one of m classes and each class has equal a priori probability. there is available a sample of radar measurements of size n from each class. the elements of the sample from the jth class are realizations of p-dimensional random variables si (j) ∼ np (s (j) ,q (j)) , i = 1 (1)n, 138 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... with unknown parameters s (j) and q (j) for each j ∈ {1, . . . ,m}. we are investigating a detected target signal on the basis of the corresponding sample of size n of p-dimensional radar measurements ri = (ri1, ..., rip) ′ , i = 1 (1)n, where ri ∼ np (s,q) . we postulate that this target signal can be regarded as a random drawing from one of the m classes but we do not know from which one. the problem is to classify a detected target signal as belonging to one of the m specified classes. when there is the possibility that a target signal does not belong to any of the m above classes, it is desirable to recognize this case. let ri and si (j) be the ith observation of the target and jth class variable, respectively. it is assumed that all observation vectors,ri = (ri1, ..., rip) ′ , si (j) = (si1 (j) , ..., sip (j)) ′ , i = 1 (1)n, are independent of each other, where n is a number of paired observations. let xi (j) = ri − si (j) , i = 1 (1)n, be paired comparisons leading to a series of vector differences. thus, for classification of a detected target signal as belonging to the jth class, it can be obtained and used a sample of n independent observation vectors x (j) = (x1 (j) , ...,xn (j)) , j ∈ {1, . . . ,m} . it is assumed that under h0 (j),xi (j) ∼ np (0,q+q (j)) , ∀i = 1 (1)n, where q+q (j) is a positive definite covariance matrix. under h1 (j), xi (j) ∼ np (a (j) ,q+q (j)),∀i = 1 (1)n, where a (j) = (a1 (j) , ..., ap (j)) ′ ̸= (0, ..., 0)′ is a mean vector. for fixed n, the problem is to construct a test which consists of testing the null hypothesis h0 (j) : xi (j) ∼ np (0,q+q (j)) , ∀i = 1 (1)n, versus the alternative h1 (j) : xi (j) ∼ np (a (j) ,q+q (j)) ,∀i = 1 (1)n, where the parameters a (j) ,q and q (j) are unknown. the cfar test of h0(j) versus h1 (j) (j ∈ {1, ...,m}) is based on the statistic given by (23), v(j) = [n(n− p)/p] ( â′(j) [ ĝ1(j) ]−1 â(j) ) , (43) where ĝ1(j) = (x ′(j)− â(j)c′)(x ′(j)− â(j)c′)′ = n∑ i=1 (xi(j)− â(j))(xi(j)− â(j))′. (44) the test of h0 versus h1, based on the gmlr statistic v(j), is given by v(j) { > h(j), then h1(j) (targetdoesnotbelongtoclassj), ≤ h(j), then h0(j) (targetbelongstoclassj), (45) where h (j) > 0 is a threshold of the test which is uniquely determined for a prescribed level of significance α(j) so that sup θ(j)∈θ0(j) eθ(j) {u(v(j))} = α(j), (46) advances in systems science and applications (2014) vol.14 no.2 139 where θ (j) = (a (j) ,q+q (j)) ,θ0 (j) = {(a (j) ,q+q (j)) : a (j) = 0, (q + q (j)) ∈ qp} , u(v(j)) = { 1, v(j) > h(j) (h1(j)), 0, v(j) ≤ h(j) (h0(j)). (47) thus, if v(j) > h(j) then the jth target class is eliminated from further consideration. if (m− 1) target classes are so eliminated, then the remaining class (say, kth) is the one to which a detected target signal being classified belongs. if all the target classes are eliminated from further consideration, we decide that a detected target signal belongs to the (m+1)th class whose distribution is unspecified. if the set of target classes not yet eliminated has more than one element, then we declare that a detected target signal belongs to the class j∗ if j∗ = arg max j∈d (h(j)− v(j)), (48) where d is the set of target classes not yet eliminated by the above test. now consider the situation in which a detected target signal s is related to the true target signal of the jth class, s(j), by s = us(j) = u(s1(j), ...sp(j)) ′, j ∈ {1, ...,m} (49) where ν is a scalar amplitude parameter. it is assumed that the target signal vectors s(j), j = 1(1)m, are known. the generalized maximum likelihood ratio statistics for this recognition problem are given by max υ {max q lh1(j)(x; υ,q)} / max q lh0(j)(x;q), (50) where lh0(j)(x;q)= (2π)−np/2|q|−n/2 exp ( − n∑ i=1 x′iq −1xi/2 ) , (51) lh1(j)(x; υ,q)= (2π)−np/2|q|−n/2 exp ( − n∑ i=1 (xi − υs(j))′q−1(xi − υs(j))/2 ) (52) are the likelihood functions under h0(j) and h1(j), j ∈ {1, . . . ,m} , respectively, and max υ {max q lh1(j)(x; υ,q)} = max υ 1 (2π)np/2 ∣∣∣⌢q1(j) ∣∣∣n/2 exp ( −np 2 ) , (53) 140 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... max q lh0(j)(x;q) = 1 (2π)np/2 ∣∣∣⌢q0 ∣∣∣n/2 exp ( −np 2 ) . (54) the well-known maximum likelihood estimates (mles) of the unknown covariance matrix q under the respective hypotheses, h0(j) and h1(j), are given by ⌢ q0 = 1 n n∑ i=1 xix ′ i = 1 n xx ′, (55) ⌢ q1(j) = 1 n n∑ i=1 (xi − υs(j))(xi − υs(j))′ = 1 n (x − υs(j)c′)(x − υs(j)c′)′. (56) after several algebraic manipulations, (28) reduces to the following clutter-adaptive test of detection of the jth target signal, j ∈ {1, . . . ,m} :{ > h(j), then h1(j), ≤ h(j), then h0(j), (57) where z(j) = [s′(j)(xx ′)−1xc] 2 [s′(j)(xx ′)−1s(j)][1− c′x ′(xx ′)−1xc] , (58) h(j) > 0 is a threshold of the test which is uniquely determined for a prescribed level of significance α(j) so that the probability of a false alarm is equal to α(j). theorem 6 (pdf of the gmlr statistic v(j)). the probability density function of v(j) under hypothesis h1(j) is given as follows: fh1(j)(v(j);n, q(j)) = 1∫ 0 f(v(j);g, n, q(j))f(g;n)dg, (59) where f(g;n) = γ ( n−1 2 ) γ (n−p 2 ) γ ( p−1 2 )(1− g) p−3 2 g n−p−2 2 , (60) for 0 ≤ g ≤ 1, and f(v(j); g, n, q(j)) = γ ( n−p+1 2 ) exp ( − q(j)g 2 ) γ (n−p 2 ) γ ( 1 2 ) advances in systems science and applications (2014) vol.14 no.2 141 ×[1− v(j)] n−p−2 2 [v(j)]− 1 2 1f1 ( n− p+ 1 2 ; 1 2 ; q(j)gv(j) 2 ) (61) for 0 < v(j) < 1. in (61) 1f1 (a; b;x) is the confluent hypergeometric function, and q(j) is the generalized signal-to-noise ratio (gsnr) defined by q(j) = gsnr = nυ2s′(j)q−1s(j). (62) under hypothesis h0(j), no signal is present. thus, if one sets q(j) = 0 in (61), fh0(j)(v(j);n) = γ ( n−p+1 2 ) γ (n−p 2 ) γ ( 1 2 ) [1− v(j)] n−p−2 2 [v(j)]− 1 2 , 0 < v(j) < 1. (63) proof. the proof is similar to that of theorem 2 and so it is omitted here. finally, in terms of the above probability density functions in (59) and (63) the probability of false alarm is given by pfa(j) = 1∫ h(j) fh0(j)(v(j);n)dv(j) (64) and the probability of detection of the jth target signal is pd(j) = 1∫ h(j) fh1(j)(v(j);n, q(j))dv(j). (65) thus, if v(j) < h(j) then the jth target class is eliminated from further consideration. if (m− 1) target classes are so eliminated, then the remaining class (say, kth) is the one to which a detected target being classified belongs. if all the target classes are eliminated from further consideration, we decide that we deal with a clutter alone. if the set of target classes not yet eliminated has more than one element, then we declare that a detected target belongs to the class j∗ if j∗ = arg max j∈d (v(j)− h(j)), (66) where d is the set of target classes not yet eliminated by the above test. 142 konstantin n. nechval: adaptive cfar tests for detection and recognition of target... 4 conclusion the main idea of this paper is to find a test statistic whose distribution, under the null hypothesis, does not depend on unknown (nuisance) parameters. this allows one to eliminate the unknown parameters of noisy processes, which are changing with time and position, from the problem. the authors hope that this work will stimulate further investigation using the approach on specific applications to see whether obtained results with it are feasible for realistic applications. acknowledgments this research was supported in part by grant no.06.1936 and grant no.01.0031 from the latvian council of science and the national institute of mathematics and informatics of latvia. this support is gratefully acknowledged. references [1] kay, s.m. and makhoul, j. (1983), “on the statistics of the estimated reflection coefficients of an autoregressive process”, ieee trans. on assp, vol. 31, pp.1447-1455. [2] stehwien, w. and haykin, s. (1989), “a statistical radar clutter classifier”, proceedings of the ieee national radar conference, dallas, texas, pp.164169. [3] nechval, n.a. and nechval, k.n. (1998), “characterization theorems for selecting the type of underlying distribution”, abstracts of communications of the 7th vilnius conference on probability theory and mathematical statistics & the 22nd european meeting of statisticians, tev, vilnius, lithuania, pp.352-353. [4] nechval, n.a., nechval, k.n., and vasermanis, e.k. (2000), “technique of testing for two-phase regressions”, proceedings of the second international conference on simulation, gaming, training and business process reengineering in operations, riga, latvia, pp.129-133. [5] nechval, n.a. (1988), “a general method for constructing automated procedures for testing quickest detection of a change in quality control”, computers in industry, vol. 10, pp.177-183. [6] abramowitz, m. and stegun, i.a. (1964), handbook of mathematical functions, national bureau of standards, new york. advances in systems science and applications (2014) vol.14 no.2 143 [7] nechval, n.a. (1992), “radar cfar thresholding in clutter under detection of airborne birds”, proceedings of the 21st meeting of bird strike committee europe, bsce, jerusalem, israel, pp.127-140. [8] lehmann, e.l. (1959), testing statistical hypotheses, john wiley, new york. [9] nechval, n.a. (1997), “adaptive cfar tests for detection of a signal in noise and deflection criterion”, digital signal processing for communication systems (edited by t. wysocki, h. razavi, b. honary), kluwer academic publishers, boston·dordrecht·london, pp.177-186. [10] nechval, n.a. (1997), “umpi test for adaptive signal detection”, signal processing, sensor fusion, and target recognition vi (edited by i. kadar), proc. spie, vol. 3068. orlando, florida usa, paper no.3068-73, 12 pages. [11] nechval, n.a. and nechval, k.n. (1999), “cfar test for moving window detection of a signal in noise”, proceedings of the 5th international symposium on dsp for communication systems, curtin university of technology, perth-scarborough, australia, pp.134-141. [12] nechval, n.a., nechval, k.n., and purgailis, m. (2011), “statistical pattern recognition principles”, international encyclopedia of statistical science (edited by miodrag lovric), springer-verlag, berlin, heidelberg, part 19, pp.1453-1457. corresponding author konstantin n. nechval can be contacted at: konstan@tsi.lv. microsoft word 1-robert vallee.doc advances in systems science and applications (2010), vol.10, no.1 1-5 issn 1078-6236 international institute for general systems studies, inc. internal time of a dynamical system robert vallée professor emeritus université paris-nord, president of wosc email: r.vallee@afscet.asso.fr abstract we have a dynamical system. the evolution of its state, x(t) at instant t, is given by a differential equation dx(t)/dt = f(x(t),t), independent of the environment. we propose to introduce a time s, or internal time, different from time t, or reference time. for this purpose we consider duration, or time elapsed between two instants. reference duration, between instants t1 and t2 is obviously given by dr(t1,t2) = t2-t1. any duration, for example internal duration di(t1,t2), must satisfy certain conditions. once we have an internal duration di(t1,t2), we can generate an internal time s = di(t0,t). the choice of di(t1,t2) depends upon the “weight” of reference duration t2-t1, seen from the internal point of view, or equivalently that of infinitesimal reference duration dt between t and t+dt. we propose that the internal duration corresponding to reference duration dt is equal to (d(x(t)/dt)2 dt. in a way (dx(t)/dt)2 is an index of the “importance” of instant t. as an example, we consider an “explosive-implosive” dynamical system described by a certain evolution equation. the corresponding internal time varies from -∞ to +∞ while reference time varies from 0 to +∞. interpretations (physiology, cosmology) are given. keywords explosion-implosion internal duration cosmological time 1. introduction we want to give a definition of the internal time, or intrinsic time of a dynamical system evolving independently of its environment. we proposed this definition for the first time in 1996. we developed it mainly in an article (vallée, 2005) which the present text reproduces partly. the notion of internal timeis opposed to that of external time, or reference time, taken for granted and which is used in the evolution equation. the basic idea is that the internal time does not elapse if the state of the system does not change, a conception close to that of aristotle for whom time ceases to be known when the “soul” does not vary. so if x(t), belonging to a finite dimensional linear space, is the state of the system at reference instant t, any real positive and increasing function, null for argument 0, of a norm of dx(t)/dt, is a measure of the intensity of change of the system at instant t .we make the most simple choice, that of the square of the euclidian norm (or scalar square) (dx(t)/dt)2 which may be seen as an index of “importance” of instant t. so we consider that the internal duration corresponding to reference duration dt (between t and t+dt) is equal to (dx(t)/dt)2 dt and we define (vallée, 1996, 2001) the internal duration di(tl , t2) of interval (t1,t2), whose reference duration is t2 –t1, by di(t1,t2) = ∫t1,t2 (dx(t)/dt)2 dt (1) so if (dx(t)/dt)2 is equal to 0 on the interval, the internal duration is 0, and if (dx(t)/dt)2 is equal to1, the internal duration is equal to the reference duration. in short, the higher the values of (dx(t)/dt)2 on the interval, the greater the internal duration. we can now, define the internal time s(t) by s(t) = di(t0,t) = ∫t0, t (dx(s)/ds)2 ds (2) where t0 is any reference instant. so s(t) if determined up to an arbitrary additive constant. of course we have di(t1,t2)=s(t2)-s(t1). (3) robert: internal time of a dynamical system 2 it is interesting to verify if equation (3) is consistent with the axioms that a « time » must satisfy. if f(a,b) is a duration attached to interval (a,b), we must have, with a≤b≤c, f(a,b) + f(b,c) ≡ f(a,c), f(a,b) >0 for b> a, f(a,a) ≡ 0 (4) f (a,b) increasing with b and decreasing with a. with the hypothesis that f is differentiable, it is easy to solve functional equation (4). we have f(a+da,b) + f(b,c+dc) ≡ f(a+da, c+dc) replacing f(a+da,b) by f(a,b) + ∂f(a,b)/∂a da and the like for f(b,c+dc) and f(a+da, c+dc), we obtain ∂f(a,b)/∂a ≡ ∂f(a,c)/∂a so ∂f(a,b)/∂a is independent of b. consequently, by integration with respect to a, we have f(a,b) = f(a) +g(b) but since f(a,a) ≡ 0 ≡ f(a) + g(a) we have g = -f so f(a,b) = g(b) – g(a) which is consistent with (3). 2. examples we call explosion the evolution of a system whose state vector has a modulus starting with value 0 at t = 0 then increasing with t, and such that the modulus of its speed vector starts with value +∞ at t = 0. the first instants of the evolution of the system have an exceptional importance since (dx(t)/dt)2 tends to +∞ when t tends to 0. we have here an idealisation as well as in the case of what we call implosion where x(t) decreases with t and attains value 0 at the final instant while (dx(t)/dt)2 tends to +∞ system may be explosive at the beginning and implosive at the end, then we say that we have an explosion-implosion. for the sake of simplicity we shall suppose now that x(t) is a mere scalar. we start with an explosion-implosion (vallée, 1996, 2001) defined by the differential equation x (t)/dt = q/p sgn(p-t) (q2 x2 (t))1/2 / x(t), x(0) = 0, p et q > 0, 0≤ t ≤ p (5) where sgn(p-t) is the sign of p-t. the solution of this equation is given by function x(t) = q/p (p2 (p-t)2)1/2 (6) whose graph is an half-ellipse (great axis 2p, small axis 2q). we say that we have an elliptic explosion-implosion. when t varies from 0 to 2p, x(t) increases from 0 to q, then decreases from q to 0,with a speed of infinite absolute value at t = 0 and t = 2p.the square of the speed of evolution is given by (dx(t)/dt)2 = q2/p2 (p-t)2/ t(2p-t) = q2/2p (1/t 2/p + 1/2p-t) which shows that the “importance” of instant t is infinite at the beginning (t = 0) and at the end (t=2p). if we integrate (dx(t)/dt)2 from t1 to t2 we obtain the internal duration of the reference time interval (t1,t2) di(t1,t2) = q2/2p (log (t2/t1) 2(t2-t1) /p -log (2p-t2 /2p-t1)) (7) now, choosing t1 = p and t2 = t we have an internal time s(t) defined by s(t) = di(p,t) = q2/2p (log t 2t/p +2 -log (2p-t)) (8) of course constant q2/2p may be supressed, then we have another internal time σ(t) = q2/2p (log t – 2t/p – log(2p-t) we see that when the reference time t varies from 0 to 2p, generating a finite reference duration of the evolution equal to 2p, the internal time svaries from -∞ to +∞generatingan infinite internal duration of the evolution. the initial instant 0 is pushed back to -∞ and the final instant 2p is pushed forward to +∞.the internal duration of any interval (0, t) is infinite as well as the internal duration of any interval (t, 2p). we consider now the differential equation advances in systems science and applications (2010), vol.10, no.1 3 dx(t)/dt = q/p (q2 + x2(t))1/ 2 / x(t) , x(0) = 0, p et q >0, 0≤τ< + ∞ (9) its solution is given by function x(t) = q/p ((p+t)2 – p2)1/2 (10) whose graph is the right part of an half-hyperbola (great axis 2p,“small axis” 2q).we say that we have an hyperbolic explosion : x(t) increases from 0 to +∞ with an infinite speed at instant 0 and, for the great values of t, x(t) behaves like (t-p) a/p. the square of the speed is given by (dx(t)/dt)2 = q2/p2 (p+t)2 / t(t+2p) = q2/2p (1/t + 2/p -1/2p+t) it shows that the « importance » of instant t is infinite at t = 0. calculations, similar to those for the elliptic case, give the internal duration of interval (t1,t2) and consequently an internal time such as s(t) = q2 /2p (log t + 2t/p – log (2p+t) (11) when t varies from 0 to +∞, s(t) varies from -∞ to +∞, the initial instant 0 being pushed back to −∞. any interval (0, t), of reference duration t, has an infinite internal duration. moreover s(t) behaves as q2 /2p log t for small values of t and as q2/p2 t for great values. an intermediary case, which we call parabolic explosion (vallée, 1996, 2001), is obtained when p and q tend to infinity while q2/p keeps a constant value 2h. starting indifferently from equation (6) or (9), we obtain dx(t)/dt = 2h / x(t), x(0) = 0, h>0 (12) the solution is given by function x(t) = 2 (ht)1/2 (13) whose graph is an half-parabola of parameter h. we see that xt) increases from 0 to +∞ with an infinite initial speed. then we have (dx(t)/dt)2 = h/t the “importance” of instant t is infinite at t = 0 and tends to 0 when t tends to +∞. the calculation of internal duration generates an internal time s(t) = h log t, (14) the initial instant 0 is pushed back to -∞ and any interval (0,t) has an infinite internal duration. 3. interpretations in the first interpretation we consider an elliptic explosion-implosion as the evolution of a living being whose birth may be compared to a kind of explosion and the end of life as an involution, or a kind of implosion, more or less quick. in our model the implosive part is symmetrical with the explosive one. this is not very realistic, since it seems that the implosive part must be shorter. nevertheless if we consider only the qualitative aspect of the conclusions we can say that, from the internal time point of view, the initial instant (conception) is pushed back to ∞and the final instant (death) is pushed forward to +∞ (vallée,1991,1996). the first part of this qualitative conclusion is in accordance with the natural feeling of a human being (if we consider this case) of having no beginning. we can also consider the case of a parabolic explosion limited at the instant of death. the initial instant is pushed back to −∞ and the internal time is proportional to the logarithm of the elapsed reference time. it elapses slowly at the beginning and more and more quickly near the end. this is close to the ideas of lecomte du noüy (lecomte du noüy, 1936). for him the physiological duration of an astronomical time interval (reference time interval in our terminology), of given length, is proportional to the speed of healing of wounds. this speed varying roughly as the inverse of age, a logarithmic physiological time is generated. but a remark seems necessary in order to avoid apparent paradoxes: the internal time of a conscious being may be different from the perceived internal time. the second interpretation concerns internal duration in cosmology. the cases of elliptic explosion-implosion, parabolic or hyperbolic explosion have common traits with certain robert: internal time of a dynamical system 4 cosmological models with primordial explosion followed by final implosion or with primordial explosion only. generally speaking, the differential equation giving the evolution of the universe, whose state at instant t is described by the cosmological scale factor r(t), is according to lemaître, friedmann, robertson (berry, 1989) (dr(t)/dt)2 = 8πg/3ρ(t) r2(t) kc2 + λ/3 r2(t), r(0) = 0 (15) where g is the gravitational constant, c the speed of light, k the index of curvature (k = -1, we have a space with negative curvature; k = 0, we have a flat space; k = +1, the space has a positive curvature (then r(t) may be considered as the radius of the universe), λ the cosmic constant or cosmic repulsion term, ρ(t) the density of matter or its material equivalent in case of pure radiation. in the material case ρ(t) = a/r3(t) and in the case of pure radiation it is equal to b/r4(t), a and b being constants. equation (5), corresponding to an elliptic explosion-implosion, gives if we consider the square of its two members (dx(t)/dt)2 = q4/p2 / x2(t) q2/p2 if we substitute r(t) to x(t), the above equation takes one of the possible forms of (15) if ρ(t) = b/r4(t), k = +1, λ= 0 . we have a case of pure radiation with positive curvature and null cosmic constant. more precisely q = c p and p = 1/c2 (b 8πg/3)1/ 2. then (dr(t)/dt)2 = 8πg/3 b/r2(t) – c2 (16) the internal time of this system, which we propose to call generalized cosmological time (vallée, 1995, 1996, 2005) is then given, according to (8), by s(t)= c2p/2 (log t 2t/p – log (2p-t) (17) the initial reference instant t = 0 (“big bang”) is pushed back to -∞ and the final referenceinstant t = 2p (“big crunch”) is pushed forward to +∞. but classically a cosmological model with pure radiation is accepted mainly as an approximation valid when the density of matter (a/r3(t)) is negligible compared to the (equivalent) density of matter of pure radiation (b/r4(t)). this happens when r(t) is small, so when t is close to 0. in that case –kc2 is negligible as well as λ/3r2 and it is not even necessary to suppose that λ = 0 . we then have (dr(t)/dt)2 = 8πg/3 b/r2(t) (18) this equation is that of the radiation-dominated era at the beginning of the universe, or that of an evolution with pure radiation, null cosmic constant and flat universe. it corresponds to a parabolic explosion whose differential equation, after taking the square of its two members, gives (dx(t)/dt) = 4h2 / x2(t) we have just to substitute r(t) to x(t) and (b 2πg/3)1/ 2 to h. according to (14), the internal time of this universe is s(t) = (b 2πg/3)1/2 log t (19) the initial instant t = 0 (“big bang”) is pushed back to -∞. we recognize here what milne has called cosmological time (milne, 1948). references [1] berry, m. principles of cosmology and gravitation. institute of physics publishing, bristol & philadelphia, 1989. [2] lecomte du noüy. p. le temps et la vie. gallimard, paris, 1936. [3] milne, e. a. kinematic relativiy. clarendon press, oxford, 1948. [4] vallée, r. perception, memorisation and multidimensional time, kybernetes, 1991, vol.20, 15-28. [5] vallée, r. cognition et système. essai d’épistémo-praxéologie. l’interdisciplinaire, limonest (france), 1995. advances in systems science and applications (2010), vol.10, no.1 5 [6] vallée, r. temps propre d’un système dynamique, cas d’un système explosif-implosif, in actes du 3ème congrès international de systémique, (e. pessa, m.p. penna, dirs), edizioni kappa, rome, pp.967-970, 1996. [7] vallée, r. time and dynamical systems. systems science vol.27, 97-100, 2001. [8] vallée, r. time and systems, kybernetes, vol.34, 9-10, 1563-1569, 2005. adv syst sci appl 2016; 16(2); 70-80 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/351 copyright ©2016 assa. adv. in systems science and appl. (2016) improved prediction of post-operative life expectancy after thoracic surgery abeer s. desuky, lamiaa m. el bakrawy al-azhar university, cairo, egypt e-mail: abeerdesuky@yahoo.com, lamiaabak@yahoo.com abstract: monitoring health outcomes is essential to enhance quality initiatives, healthcare management and consumer education. thoracic surgery is the data collected for patients who underwent major lung resections for primary lung cancer. the application of machine learning techniques for predicting post-operative life expectancy in the lung cancer patients is an area with little research and few concrete recommendations. in order to use machine learning techniques effectively, attribute ranking and selection is an integral component to successful health outcome prediction. in this paper, we present three attribute ranking and selection methods to improve algorithms performance for health outcomes research. two papers results for other researchers are used in comparison to show the efficiency of our proposed attribute ranking and selection methods. keywords: attribute ranking, machine learning, prediction, thoracic surgery. 1. introduction integrating computer applications into the medical field have directly affected the productivity and accuracy of doctors nowadays. measuring health outcomes is one of these applications. clearly, there is a growing role for health outcomes in the purchasing and management of healthcare. these days cancer is one of the major causes of death in the most countries. currently, lung cancer is the most frequent augury for thoracic surgery [1]. researchers applied different strategies, such as examination in early stage, to identify the type of cancer before the emergence of symptoms. furthermore, new methods for the early prediction of cancer therapy outcome have been developed [2].with the raise of new techniques in the field of medicine, massive datasets of cancer have been collected and now available to researchers in the medical field. however, the most challenging task is predicting a disease outcome accurately. so, the current research efforts examine the use of machine learning techniques for discover and identify models and relationships between them, from large datasets, the data is analyzed to extract useful information that supports disease augury, and to improve models that predict patient’s health more accurately [3,4]. huge datasets usually lead to crumble the performance and accuracy of the machine learning systems. datasets with high dimensional attributes have more processing complexity with longer computational time for prediction. attribute ranking and selection is a solution to complex datasets [5]. several attribute and selection methods have been presented in the machine learning domain. the main aim of these methods is to remove attributes that can be irrelevant, misleading, or redundant which increase search space size resulting in difficulty to process data further thus not contributing to the learning process. attribute and ranking selection is the process of choosing best attributes from all the attributes that are useful to discriminate classes [6, 7]. mailto:abeerdesuky@yahoo.com mailto:lamiaabak@yahoo.com 71 a. s. desuky, l.m. el bakrawy copyright ©2016 assa. adv. in systems science and appl. (2016) the rest of this paper is organized as follows: brief introduction on machine learning algorithms and attribute ranking and selection methods have been applied to disease prognosis and prediction are introduced in sections (2 and 3). the details of the proposed methods and the data set are presented in section (4). section (5) shows experimental results. conclusions are discussed in section (6). 2. related work the most major operation that performed on lung cancer patients is the thoracic surgery. survival rate is very critical factor for sawbones to decide on which patient surgery would be performed. selection of the appropriate patient for surgery is one of the common clinical decision challenges in thoracic surgery, bearing in mind risk and benefits for a patient, both in short-term (e.g. post-operative complications, including death-rate in the first month) and long-term perspective (e.g. survival for 1-5 years) [8]. a variety of different machine learning algorithms and attribute ranking and selection methods have been applied to disease prognosis and prediction in the last decades. a comprehensive search was performed relevant to the use of machine learning algorithms in cancer receptivity, recurrence and survival prediction [2]. k. kourou et al. [2] presented predictive models based on various supervised machine learning techniques including support vector machines, bayesian networks, artificial neural networks, and decision trees as an aim to model cancer risk or patient outcomes. in their work, maciej zieba et al.[9], used boosted svm for predicting post-operative life expectancy. in their research, they applied oracle-based approach for extracting decision rules from the boosted svm in order to solve imbalanced data problems. sindhu et al. [1] used six classification approaches-naive bayes, j48, part, oner, decision stump and random forest-toanalyse thoracic surgery data and they found that random forest gives the best classification accuracy with all split percentages. another paper [10] have analyzed and compared the performance of four machine learning techniques (naïve bayes, simple logistic regression, multilayer perceptron and j48) with their boosted versions by different metrics. their results indicate that boosted simple logistic regression technique is generally better or at least competitive against the rest of four machine learning techniques with 84.53% prediction accuracy. 3. machine learning machine learning is a branch of artificial intelligence which utilizes statistical, optimization and probabilistic techniques that allows computers to “learn” from past examples and to detect hard-to-discern patterns from large, noisy or complex data sets. these techniques have become a popular tool in medical diagnosis, which can find and identify models and relationships between them from large, noisy or complex datasets[3]. the inputs are the information about the patient's age, gender, past medical history, past medical procedures, family medical history and current symptoms , while labels are the illnesses. in some cases, these inputs are missed because some tests haven't been applied to the patient, so we do not apply machine learning techniques unless we confirm that the patient will give us valuable information. if the medical diagnosis is wrong, decision may lead to a wrong or no treatment, so machine learning is extremely used to diagnose and detect cancer [4]. more recently, it has been widely applied in the field of cancer prediction and prognosis which are differ from cancer detection and diagnosis. there are three types of cancer prediction and prognosis: one of them is prediction of cancer receptivity. in this type, one is trying to predict the probability of cancer progression before occurrence of the disease. second type is the prediction of cancer recurrence by trying to predict the probability of redeveloping cancer after treatment and after a period of time during which the cancer cannot improved prediction of post-operative life expectancy after thoracic surgery 72 copyright ©2016 assa. adv. in systems science and appl. (2016) be detected. third type is the prediction of cancer survivability by trying to predict an outcome which usually refers to life expectancy, survivability, progression and tumor-drug sensitivity. these days, different types of cancer such as prostate, brain, cervical, esophageal, leukemia, head, neck, breast, and thoracic are appear to be compatible with machine learning prediction. the thoracic datasets is concerned with classification problem related to the post-operative life expectancy in the lung cancer patients[2,3,4] in order to improve machine learning techniques when the datasets have a large number of features or attributes, attribute ranking and selection is used to identify the most relevant attributes and remove the redundant and irrelevant attributes from the dataset. attribute ranking and selection algorithms can be divided into wrapper and filter methods. the wrapper methods select attributes based on an estimation of the accuracy according to target learning algorithm. after applying the learning algorithm, wrapper searches the feature space by removing some attributes and testing the effectiveness of attribute removing on the prediction metrics. the attribute which make important difference in learning process should be selected as high quality attribute, while filters methods estimate the quality of selected attributes independently from the learning algorithm. it depends on the statistical correlation between the set of attributes and the target attribute, since the value of correlation identify the importance of target attribute [6,11]. by using filtering methods attributes can be ranked independently, then according to the ranking result optimal subset of attributes can be selected [12]. 4. the proposed method 4.1 dataset description table 1. characteristic of dataset features. name description characteristics dgn diagnosis specific combination of icd-10 codes for primary and secondary as well multiple tumors if any nominal pre4 forced vital capacity fvc numeric pre5 volume that has been exhaled at the end of the first second of forced expiration fev1 numeric pre6 performance status zubrod scale nominal pre7 pain before surgery binary pre8 haemoptysis before surgery binary pre9 dyspnoea before surgery binary pre10 cough before surgery binary pre11 weakness before surgery binary pre14 t in clinical tnm size of the original tumor, from oc11 (smallest) to oc14 (largest) nominal pre17 type 2 dm diabetes mellitus binary pre19 mi up to 6 months binary pre25 pad peripheral arterial diseases binary pre30 smoking binary pre32 asthma binary age age at surgery numeric risk1y 1 year survival period t value if died binary 73 a. s. desuky, l.m. el bakrawy copyright ©2016 assa. adv. in systems science and appl. (2016) thoracic surgery data is dedicated mainly to elicit surgical risk for real-life clinical lung cancer patients. the data was collected retrospectively by mareklubicz et al. [13] at wroclaw thoracic surgery centre for consecutive patients –ages from 21 to 87 years old who underwent major lung resections for primary lung cancer in the years 2007–2011. the centre is associated with the department of thoracic surgery of the medical university of wroclaw and lower-silesian centre for pulmonary diseases, poland, while the research database constitutes a part of the national lung cancer registry, administered by the institute of tuberculosis and pulmonary diseases in warsaw, poland. the dataset includes 470 instances (70 true and 400 false) and 16 attributes with no missing values and binary valued class (death within one year after surgery – survival). 4.2 research methodology in this work, version 3.7.12 of weka (waikato environment for knowledge analysis) toolkit [14] has been used for analysis. it is the product of the university of waikato (new zealand) and it is licensed under the gnu general public license. weka is a popular suite of machine learning software written in java, also it provides access to sql database and process the result retrieved by a database query. we have run our experiments on a system with a 2.30 ghz intel(r) coretmi5 processor and 512 mb of ram running microsoft windows 7 professional (sp2). cross-validation (10 folds) has been used in this study to validate the results. in this model, the dataset is partitioned into complementary 10 equal sized subsets. the analysis is performed on 9 subsets (training) and validating the analysis on one subset (testing). ten rounds of cross-validation are performed and in each round another subset 2 through 10 used as testing dataset. the validation results are averaged over the ten rounds in the final phase. researchers in the machine learning field have proposed numerous attribute ranking and attribute selection methods. the main aim of these methods is to eliminate redundant or irrelevant attributes from the original set of attributes. in our work, we use the attribute ranking methods (information gain (ig) attribute evaluation, symmetrical uncertainty (su) attribute evaluation and relief-f (rf) attribute evaluation) information gain (ig) attribute evaluation [15] is used to evaluate the importance of an attribute by measuring the information gain with regard to the class. the bases of ig depend on entropy which measure the randomness of the system. information gain can be calculated by the following equation: attribute)| h(class-h(class) = attribute) ig(class, (1) where h is the entropy which stands for the greek alphabet eta. symmetrical uncertainty (su) attribute evaluation [16]is used to evaluate the importance of an attribute by measuring the symmetrical uncertainty with respect to the class. symmetrical uncertainty compensates for the inherent bias in information gain. symmetrical uncertainty is given by the following equation: e))h(attribut-h(class) /(attribute) ig(class,*2 = attribute) su(class, (2) relief-f (rf) attribute evaluation is used to rank the quality of features depending on how well their values differ from the cases that are close to each other. it is sensible to predict that a valuable feature should have different values between cases belong to different classes and have the same value for cases from the same class [17]. the aim of this paper is to analyze the effect of number of attributes on accuracy of machine learning techniques to solve the problem for prediction of the post-operative life https://en.wikipedia.org/wiki/gnu_general_public_license https://en.wikipedia.org/wiki/machine_learning https://en.wikipedia.org/wiki/java_(programming_language) https://en.wikipedia.org/wiki/complement_(set_theory) improved prediction of post-operative life expectancy after thoracic surgery 74 copyright ©2016 assa. adv. in systems science and appl. (2016) expectancy in the lung cancer patients. reducing the number of attributes and increasing the accuracy is required to minimize the computational time of prediction techniques. in this study, we used information gain, symmetrical uncertainty and relief-f as attribute ranking methods to reduce the number of attributes (from 16 to 13 attributes), then we examined the quality of techniques naïve bayes, simple logistic regression, j48, multilayer perceptron, and svm after applying the three ranking methods for prediction of post-operative life expectancy after thoracic surgery. the quality of the proposed methods is evaluated by comparing the performance of naïve bayes, simple logistic regression, j48 and multilayer perceptron techniques with and without using attribute ranking methods as first step. also, our proposed` methods is compared to boosted naïve bayes , boosted simple logistic regression, boosted j48, boosted multilayer perceptron and boosted svm. 5. experimental results performances of the methods were analyzed by using six metricsaccuracy, f measure, roc curve, gmean, tnr and tpr [9, 10]. accuracy is the percentage of observations that were correctly predicted by the method. it was used to evaluate the performance of each algorithm. n)+tn/(p+tp =accuracy (3) table 2. shows the confusion matrix which clarifies the prediction tendencies tp (true positive), tn (true negative), fp (false positive) and fn (false negative) of considered machine learning technique. table 2. confusion matrix predicted outcome p n actual value p tp fn n fp tn accuracy is not a reliable metric for the real performance of a machine learning technique, because it will yield misleading results if the data set is imbalanced (i.e. when the number of samples in different classes vary greatly). since thoracic surgery data is imbalanced data with 70 true and 400 false instances we used f measure (f1 score), roc curve, gmean, tnr and tpr. where, f measure was used to test the accuracy depending on harmonic mean of precision & recall. fn)+fp+2tp/(2tp = measure f (4) while, roc curve was also used as an effective method to evaluate the performance of predicted models by plotting the true positives against the false positives and area under the roc curve is used for predicting accuracy of models. the gmean (geometric mean) is a widely used quality rate and is defined as equation: tnr tpr=gmean  (5) where tnr (specificity or true negative rate) is described by: fp) + tn/(tn = tnr (6) and tpr (sensitivity or true positive rate) and described by the equation: fn) + tp tp/( = tpr (7) 75 a. s. desuky, l.m. el bakrawy copyright ©2016 assa. adv. in systems science and appl. (2016) table 3. shows the accuracy of naïve bayes, simple logistic regression, j48 and multilayer perceptron techniques with and without using attribute ranking methods. also, it shows the accuracy of boosted naïve bayes, boosted simple logistic regression, boosted j48, boosted multilayer perceptron for prediction of post-operative life expectancy after thoracic surgery. results show that using ig and su as ranking methods before applying naïve bayes gives the better accuracy than applying naïve bayes without using ranking methods and with boosted naïve bayes. also, in the case of applying simple logistic, using the three ranking methods gives better accuracy than applying simple logistic without using ranking methods and with boosted simple logistic. similarity, in the case of applying multilayer perceptron, using the three ranking methods gives better accuracy than applying multilayer perceptron without using ranking methods and with boosted multilayer perceptron. but in the case of applying j48 without using ranking methods gives better accuracy than applying j48 with using ranking methods and applying j48 with using ranking methods gives better accuracy than boosted j48. table 3. also shows that the simple logistic technique applied with the three ranking methods gives the best accuracy. table 3. prediction accuracy comparison of machine learning techniques using thoracic surgery data set ml techniques method accuracy naïve bayes original [10] 77.74 boosted [10] 78.32 su 82.12 rf 77.74 ig 82.13 simple logistic original [10] 84.55 boosted [10] 84.53 su 84.68 rf 84.68 ig 84.68 multilayer perceptron original [10] 80.91 boosted [10] 80.70 su 81.27 rf 81.28 ig 81.28 j48 original [10] 84.64 boosted [10] 79.34 su 84.46 rf 84.47 ig 84.47 table 4. shows the f measure and roc curve of naïve bayes, simple logistic regression, j48 and multilayer perceptron techniques with and without using attribute ranking methods. also, it shows the f measure, roc curve of boosted naïve bayes, boosted simple logistic regression, boosted j48, boosted multilayer perceptron for prediction of post-operative life expectancy after thoracic surgery. results show that applying naïve bayes without using ranking methods gives the better f measure than using the three ranking methods before applying naïve bayes and with boosted naïve bayes, but, boosted naïve bayes gives the best roc curve. in the case of applying simple logistic, it gives the same results for the f measure in all methods, but using the three ranking methods gives the best roc curve. in the case of applying multilayer perceptron, using su and ig ranking methods gives the best f measure and the best roc curve. in the case of applying j48, boosted j48 gives the best f measure but applying j48 with and without using ranking methods gives better roc curve than boosted j48. improved prediction of post-operative life expectancy after thoracic surgery 76 copyright ©2016 assa. adv. in systems science and appl. (2016) table 4. prediction measures comparison of machine learning techniques using thoracic surgery data set ml techniques method f measure roc naïve bayes original [10] 0.13 0.68 boosted [10] 0.12 0.60 rf 0.06 0.66 su 0.06 0.66 ig 0.07 0.67 simple logistic original [10] 0.00 0.53 boosted [10] 0.00 0.61 rf 0.00 0.50 su 0.00 0.50 ig 0.00 0.50 multilayer perceptron original [10] 0.22 0.60 boosted [10] 0.18 0.56 rf 0.20 0.58 su 0.24 0.55 ig 0.24 0.55 j48 original [10] 0.00 0.50 boosted [10] 0.18 0.61 rf 0.02 0.50 su 0.00 0.50 ig 0.00 0.51 table 5. shows the tpr, tnr, and gmean for support vector machine after applying the three attribute ranking and selection methods and boosted support vector machine. the rf gives the best prediction quality where it has the higher gmean value. it shows also that the proposed methods give better gmean and tnr than boosted svm but boosted svm gives better tpr. table 5. performance evaluation of boosted svm vs. svm with ranking methods method tpr tnr gmean boosted svm(bsi) [9] 60.00 72.00 65.73 svm (ig) 44.30 99.80 66.49 svm (su) 44.30 99.80 66.49 svm (rf) 51.40 99.80 71.62 6. conclusion in this study, the quality of three attribute ranking and selection methods has been evaluated to improve the prediction for life expectancy of lung cancer patients after thoracic surgery. five machine learning techniques before and after applying the attribute ranking and selection methods have been compared with their boosted versions. the results show that boosting is not always the better choice where attribute ranking and selection can perform better in improving prediction accuracy. other attribute selection and machine learning techniques can be introduced in the future work to gain a better prediction model performance of the dataset. references [1] v. sindhu, s. a. s. prabha, s. veni , and m. hemalatha, “thoracic surgery analysis using data mining techniques” , international journal of computer technology & applications , vol. 5, pp 578-586, may, 2014 [2] konstantina kourou, themis p. exarchos, konstantinos p. exarchos, michalis v. karamouzis, dimitrios i. fotiadisa, “machine learning applications in cancer prognosis and prediction”, computational and structural biotechnology journal, vol 13, pp 8-17, 2015. http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/article/pii/s2001037014000464 http://www.sciencedirect.com/science/journal/20010370 77 a. s. desuky, l.m. el bakrawy copyright ©2016 assa. adv. in systems science and appl. (2016) [3] kwetishe joro danjumal, “performance evaluation of machine learning algorithms in post-operative life expectancy in the lung cancer patients”, international journal of computer science issues, vol. 12, no. 2, pp. 189-199, 2015 [4] joseph a. cruz, david s. wishart, “applications of machine learning in cancer prediction and prognosis”, cancer informatics, vol. 2, pp. 59–77, 2006 [5] mehdi naseriparsa, amir-masoudbidgoli,tourajvaraee, “a hybrid feature selection method to improve performance of a group of classification algorithms”, international journal of computer applications, vol. 69, no 17, pp 28-35, 2013 [6] pinar yildirim, “filter based feature selection methods for prediction of risks in hepatitis disease”, international journal of machine learning and computing, vol. 5, no. 4, pp. 258-263, august 2015 [7] samina khalid, tehmina khalil, shamila nasreen, “a survey of feature selection and feature extraction techniques in machine learning”, science and information conference (sai), pp. 372-378, 2014 [8] koklu murat, humar kahramanli, and novruz allahverdi. “applications of rule based classification techniques for thoracic surgery”, tiim international conference, 27-29 may 2015, italy, pp 1991-1998. [9] m. zięba, j .m. tomczak, m. lubicz, and j. świątek, “boosted svm for extracting rules from imbalanced data in application to prediction of the post-operative life expectancy in the lung cancer patients” , applied soft computing, vol. 14, pp 99-108, jan. 2014. [10] md. ahasan uddin harun and md. nure alam, “predicting outcome of thoracic surgery by data mining techniques”, ijarcsse, vol. 5, no. 1, pp 7-10, 2015 [11] mark a. hall, “correlation-based feature selection for discrete and numeric class machine learning”, international conference on machine learning, pp 359-366, 2000 [12] vishal gupta, et al., “performance of various feature selection techniques under loaded networks”, international journal of computer applications, vol. 78, no. 2, september 2013 [13] uci machine learning repository. url: https://archive.ics.uci.edu [14] machine learning group at the university of waikato. url: http://www.cs.waikato.ac.nz/ml/ [15] c. sunil kumar and r.j. rama sree, “application of ranking based attribute selection filters to perform automated evaluation of descriptive answers through sequential minimal optimization models”, ictact journal on soft computing: special issue on distributed intelligent systems and applications, vol. 5, no. 1, october 2014 [16] jasmina novaković, perica strbac, dusan bulatović , “toward optimal feature selection using ranking methods and classification algorithms”, yugoslav journal of operations research, vol. 21, no. 1, pp. 119-135, 2011 [17] niket kumar choudhary, yogita shinde, rajeswari kannan, vaithiyanathan venkatraman, “impact of attribute selection on the accuracy of multilayer perceptron”, ijitkmi, vol. 7, no. 2, pp. 32-36, 2014 http://ieeexplore.ieee.org/search/searchresult.jsp?searchwithin=%22authors%22:.qt.khalil,%20t..qt.&newsearch=true http://ieeexplore.ieee.org/search/searchresult.jsp?searchwithin=%22authors%22:.qt.nasreen,%20s..qt.&newsearch=true https://archive.ics.uci.edu/ http://www.cs.waikato.ac.nz/ml/ advances in systems science and applications (2012) vol.12 no.1 27-37 study on the horizontal subgrade reaction of expressway protective guard pillar lu yang1, daiheng chen2 and bo xiao1 1shenyang university of technology school of civil engineering and architecture ,shenyang,110023 2tokyo university of science engineering department , tokyo ,102-0073 abstract it is quite necessary to describe the relationship between the horizontal subgrade reaction and the horizontal displacement with a simple form. this study provides the analysis of the horizontal subgrade reaction of protective guard pillar in the conflict between automobiles and protective guard pillars. making use of the relationship between soil load and horizontal load, we can build the theory of assessment of the coefficient of horizontal subgrade reaction, and then simulate the relationship between soil and column with spring in the help of the theory. in this paper, we use mohr-coulomb yield criterion and integrate abaqus finite element numerical simulations, to find that it is the difference between the greatest horizontal stress components σx and the smallest horizontal stress components σy that causes the fact that horizontal load leads to the plastic yielding of protective guard pillar and earth coupling model soil. and also, the paper provides the theoretical formula of coefficient of horizontal subgrade reaction and the ultimate bearing capacity of horizontal subgrade reaction, then build the theoretical model of coefficient of horizontal subgrade reaction when soil is under the action of horizontal force, to offer design and analysis reference for practical projects. keywords protective guard pillar, level of load, abaqus, level reaction site, mohr-coulomb yield criterion, numerical simulation 1 introduction the accident that automobiles collide guardrails is one of the main forms of road traffic accidents, and their collision has become main research direction of traffic safety[1-3]. in the analysis of building structures, spring is usually exerted in the foundation to simulate the relationship between soil and structures, so as to predict the coefficient of soil subgrade reaction in the system of building structures, and then we can ascertain spring coefficient. usually, the prediction method of coefficient of soil subgrade reaction refers only to semi-infinite space soil or beam on elastic foundation[4-5], so it’s necessary to revise the existing prediction method of coefficient of soil subgrade reaction, and ascertain its failure mechanism. this research makes use of the relationship between soil load and horizontal load, taking their elastic plastic behavior into consideration[6-7], and builds the prediction method of coefficient of soil subgrade reaction which is applied to buiding structure system, expecting that the research result is helpful 28 lu yang:study on the horizontal subgrade reaction of expressway protective guard pillar for the selection of coefficient of engineering soil subgrade reaction. take the columns of semi-rigid guardrail system as object of study, and research each component’s characteristics of deformation and energy-absorbing in the impact of shock. because of the fact that the column is rammed and buried into soil,and the nonlinear characteristics of soil’s height, in-depth study of the relationship between surrounding soil and column and also the characteristics of deformation and energy-absorbing of column is of great importance. this paper integrates finite element numerical simulations, providing the analysis of horizontal subgrade reaction and limiting uniform force,and finally simulates the relationship between soil and column with a series of horizontal springs, to build the theoretical model of coefficient of horizontal subgrade reaction when the soil is under the action of horizontal force;offers design and analysis references for practical projects. 2 previous research for coefficient of subgrade reaction ks piles’ horizontal resistance is represented by coefficient of subgrade reaction ks. coefficient of subgrade reaction ks is proposed by kinds of proposals. usually, the derived relational expression kh0 = α · e0 · d−3/4 acts as the coefficient of subgrade reaction when the benchmark displacement is 10mm,and we can represent the coefficient of subgrade reaction with the nonlinear curve of second order kh = kh0/y 1/2 of the relative displacement y between piles and the surrounding subgrade.among them,:determinate number=80 (for cohesive soil ,when e0 is derived by n,is 60). gold proposed that relation curve of horizontal load pmax and displacement could be represented approximately by three crease line, among them, y is the relative displacement of piles and surrounding subgrade, 1(y < 0.6mm),1/4(0.6mm < y < 10mm), 1/12(y > 10mm) acts as the slope of each line[8]. as can be seen from the above, kh’s character is that its value reduces gradually with deformation. seed and others analyzed many horizontal subgrade response results, and proposed that stress reduction factor should be decreased with depth, and show bandwidth viriation with depth[9]. the deeper, the wider. also, they suggested that when analyzing and designing, we could use average curve. the most important data of pile detail design includes coefficient of subgrade reaction ks that responses deformation behavior. in existing d sees the formula (2). esign, ks’s precise value is not as precise as other soil parameters (such as: undrained cohesion, internal friction angle, etc.). usually, we can estimate the relationship of horizontal subgrade reaction with young’s modulus es: non-sticky soil kh = 3×es/d, sticky soil kh = 1.6×es/d, d is pile diameter. in evaluating pile’s coefficient of horizontal subgrade reaction, its horizontal advances in systems science and applications (2012) vol.12 no.1 29 load’s horizontal displacement and the effect of horizontal subgrade reaction are pretty important, only in reasonable way to build appropriate design value, and it’s essential to operate the experiment of loading pile load horizontally or conduct the simulation of correct value. the traditional method is to conduct the experiment of pile’s destructive load, and it’s rather difficult in loading condition or test technology[10-11]. finite element method is the most appropriate method at this stage, through the material parameters obtained from laboratory testing, the mechanical behavior of making use of simulation analysis to determine the pile-soil interaction is worthy of wide concern. 3 numerical simulation analysis analyzing and calculating the factors that affect pile’s horizontal subgrade reaction and displacement correctly is of the utmost importance.this study simulates the contact problem of pile-soil interaction with abaqus finite element software, and makes use of the one by one contact algorithm in abaqus, and then build contact couple between the side of pile and the soil; for pile shaft, we can use elastic plastic model, and for soil, we can simulate with mohr-coulomb model; taking the effect of initial earth stress into consideration, we introduce soil lateral pressure coefficient to achive balance in the stage of ‘geostatic’, and this is very effective for simulation of geotechnical issues. 3.1 the establishment of finite element model of column soil in the geometric model, use large body to simulate semi-infinite space, and in the simulate calculation , the radius of the soil is much lager than the radius of the pile’s cross-section. as the picture 1 shown. in this model, the soil is built by eight-node entity reduced element, and others are four node shell element. the whole mode system includes 89158 nodes and 76656 elements, and material models are all elastic plastic materials that consider isotropic hardening. the main parameters that can be revised in desigh are as follows(see table 1). 3.2 interpretation of result mohr-coulomb damage and strength criteria is widely applied to geotechnical engineering. the constitutive model that is used in abaqus is the classical expansion of mohr-coulomb yield criterion[12]. at some point the role of the shear stress equal to the shear strength, the damage occurred, and the role of shear strength in the face of the linear relationship between normal stress; allow the material isotropic hardening or softening, however, the shape of the flow of the model potential function in the meridional plane is hyperbola, and there is no cusp in the π plane, potential function is therefore completely smooth, ensuring the uniqueness of the plastic flow direction. fig.2 is the curve of path and displcement for the soil in front of column at 30 lu yang:study on the horizontal subgrade reaction of expressway protective guard pillar fig.1 mesh model different loading time on the depth of the soil. as can be seen from figure, the displacement of the soil in front of the column is gradually increasing with the load; and is of inverse relationship with the buried depth. because of the role of soil bound, the region that column and soil doesn’t saperate, has the characteristic that displacement reduces with buried depth; at about 1.05 meters in depth accurs the phenomenon that column saperate from the soil. kh is the ratio of the distribution force destiny q and relative displacement y,that is q/y; q is obtained form the distribution stress in the surface of the pile, and here we should firstly study the stress σx of the point in front of pillar. fig.3 is the curve of the relationship between the path in depth and σx (σx = s11) , as can be seen, according to the horizontal load, s11 is negative value, and its absolute value is roughly linear and gradually become larger. fig.4 is the distribution force of the soil in front of pillar in depth, as can be table 1 model of material parameters pile length pile length 2.2 m buried in soil 1.5m elastic modulus 2.1e12 pile diameter 0.14m of column pa elastic coefficient of modulus of 5.0e+6pa soil lateral 0.538 soil pressure soil poisson’s ratio 0.35m friction angle 20 soil cohesion 3000pa expansion angle 0.1 advances in systems science and applications (2012) vol.12 no.1 31 seen, with the development of deformation, the latter phenomenon of spin-off accurs, so we can predict the theoretical starting point of the spin-off, and the following spin-off broadens gradually. 4 limiting distribution force qcr based on the study on limiting distribution force qcr, we can draw the qualitative conclusions as follows: fig.2 the depth way-deflection curve (1) the limiting distribution force qcr that varies with the the soil density γ, this function as shown in fig.5, and the performance is not obvious. especially in a lot of literatures, qcr is in proportional relationship with proportion γ, here the relationship qcr∞rgx isn’t performed very well. the relationship is as stated above, σx is the middle stress, and yield stress is related to σy and σz. that is to say, usually, when we consider coulomb earth pressure, we take it as plane strain issue, σz is middle stress, and coulomb limiting stress is determined by the difference of σy and earth pressure σx∞rgx and believe that σy∞rgx. the performance of the relationship is: the movement of the soil pile, is in threedimensional stress state in soil that in front of the pile. coulomb limiting stress is determined by the difference of σy and earth pressure σx. consequently, the function qcr∞rgx can’t exist. (2) about the relationship with cohesion c, according to the research and analysis into numerical simulations, we can gain a deeper knowledge: no matter big or small the depth x is, qcr will become bigger with the cohesion c. consequently, the expression of qcr is: 32 lu yang:study on the horizontal subgrade reaction of expressway protective guard pillar fig.3 depth way-load direction curve of pressure and resistance fig.4 depth way-distributed force relational graph qcr = qcro + cc × x (1) qcr can be ascertained from the following expression qcro ∼= 2 cosϕ 1− sinϕ cb (2) meanwhile, the relationship between coefficient cc and cohesion c is ascertained.for previous research on the construction of pile, coefficient cc is not dependent on cohesion, but has founction relationship with friction angle ϕ in 3kprg form and density γ. advances in systems science and applications (2012) vol.12 no.1 33 fig.5 r= 1800 and r= 900situation q/e0b and v/b relations fig.6 in logarithm graph q/e0b and v/b relations based on the research above, the approximate expression of limiting force qcr is: qcr ∼= 2 cosϕ 1− sinϕ b(c+ 0.24k1k2rg 1 + sinϕ 1− sinϕ x b ) (3) 34 lu yang:study on the horizontal subgrade reaction of expressway protective guard pillar k1 and k2 are the modified coefficient of soil proportion γ and cohesion c. k1 = 0.3 + 0.8× r0 r , r0 = 1800kg/m3 (4) k2 = 0.5 + 0.5× c c0 , c0 = 3000pa (5) 5 the proposal of the relationship of horizontal subgrade distribution force q and horizontal displacement v the purpose of this study is to propose the analysis of the simple model of horizontal subgrade reaction in the conflict between automotives and protective guard pillars. therefore, based on the statement above, it’s quite essential to describe the relationship between the horizontal subgrade distribution force q and the horizontal displacement v with a simple form. this paper researched the logarithmic chart that is shown in fig.6, and express the relationship between q and q approximately in the chart with three lines. in the logarithmic chart, firstly, for initial elastic deformation, can be expressed approximately as the line whose slope λ = 1; secondly, the latter non-elastic deformation can be expressed approximately as the line whose slope λ = 1/2; at last, if horizontal subgrade distribution force q reaches limiting force qcr, then it can be expressed approximately as the line whose slope λ = 0 . from the fig.6 we can see: in the early stages of non-elastic deformation, horizontal subgrade distribution force q’s limiting resistance is qcro in the place where x = 0, sees the formula (2). however, we can see from the chart, when the depth x is big, the starting point transit form the line whose slop λ = 1 to the line whose slope λ = 1/2, when we calculate horizontal subgrade reaction qcr, we can use the following expression : qre = p2 × [qcro + 0.1(qcr − qcro)] (6) qcr can be ascertained by it. therefore, q can be ascertained by the following expression : q =  p1 × c1e0v (v ≤ vre) cxe0 √ vb (vre ≤ v ≤ vcr) p2 × qcr (vcr ≤ v) (7) p1 and p2 in equation (6) and (7) modified coefficients when considering the friction between soil and pillar. when the friction coefficients µ = 0, p1 = p2 = 1; when µ ̸= 0, the calculation of p1 and p2 will be stated later. the unknown quantities vre, cx and vcr in equation (7) can be ascertained as follows: firstly, when v reaches vre, q reaches qre from the first expression in equation advances in systems science and applications (2012) vol.12 no.1 35 (7) we can obtain: vre = qre p1c1e0 (8) and also, when v = vre from the first and second expressions in equation (7)we can obtain: cx = qre e0 √ bvre (9) so when v = vcr, from the second and third expression in equation (7)we can obtain: vcr = ( p2qcrcxe0 )2 b (10) the predicting value according to the expression (3) is quite closed to the abaqus numerical simulations, therefore, the theoretical model of the horizontal resistance of the protective guard pillar in subgrade is feasible. 6 conclusion as is shown in the research result, with the help of the abaqus software, we built the coupling model of pillar and soil under the action of horizontal load, and analyzed soil’s mechanical behavior when pillar is under the action of horizontal load, and also, proposed theoretical expression about the coefficient of horizontal subgrade reaction. the simulation results fit the trend of the literature[13] experimental data very well, indicating that abaqus has a great ability in dealing with highly nonlinear issues of the interaction between pillar and soil. (1) along the depth of the pillar, the coefficient of horizontal subgrade reaction is not certain, in the region more nearer to the surface, the value is relatively smaller, and this is the result that the horizontal resistance is small for the region near the surface . (2) in pillar’s slewing deformation, the coefficient of horizontal subgrade reaction is not certain too. it will become smaller with the development of the deformation. this is because the effects of plastic deformation become greater gradually. (3) according to the argument above, using mohr-coulomb yield criterion, we can find that soil plastic yielding damage is caused by the difference between σx and σy. because of the action of horizontal load, σx that is s11) is negative value, and its absolute value is roughly linear and gradually become larger, and σy (that is s22)’s absolute value gradually become smaller. the greatest stress component is σx, and the smallest stress component is not σz (that is s33) which in gravity direction ,but is σy nd the cause of the entire yielding is the difference between σx and σy. (4) proposed the theoretical expressions of the limiting distribution force qcr 36 lu yang:study on the horizontal subgrade reaction of expressway protective guard pillar and horizontal subgrade reaction q and horizontal displacement v . acknowledgements supported by key laboratory of geological hazards in three gorges reservoir of ministry of education area, china three gorges universityunder grant no.2008kdz10. especially thanks to professor chen from tokyo university of science for his instruction and help when the author worked at tokyo university of science engineering department during mar.2008-mar.2009. references [1] nathaniel r. seckinger, a.m.asce, akram abu-odeh. (2005), “performance of guardrail systems encased in pavement mow strips”, journal of transportation engineering , asce / november, pp.851-860. [2] atahan a.o, and ross h. e, jr. (2004), “computer simulation of recycled content guardrail post impacts”, transp. eng., pp.733-741. [3] tabiei a and jin w. (2000), “roadmap for crashworthiness finite element simulation of roadside safety structures”, finite elem. anal. design, pp.145157. [4] yuanshun bi huanren zhang, “tilted bored pile the foundation system of economy and safe”, the user convention of germany bao e mechanic equipment ltd. [5] xunshan zhan, jianzhi wang, guanxiong wang. (2001), “mechanics parameter of gao xiong jie yun red-line soil layer”, proceeding of 9th geotechnical engineering proseminar (paper no. a031), august 30-31 and september 1. [6] mizuue oozutsuchiya maneka. (2002), “the analysis of horizontal subgrade reaction”, proceeding of 37th subgrade proseminar, pp.1477-1478. [7] kaimi seiyanakai teruokido hiraki. (2002), “the actual locale experiment of the horizontal load”, model proceeding of 37th subgrade proseminar, pp.1467-1468. [8] seed h.b and peacock w.h. (1971), “test procedures for measuring soil liqufaction characteristics”, journal of the soil mechanics and foundations division, asce, vol.97 no.sm8, proc., paper 8330, pp.1099-1119. [9] torinami shinsuke, tomi kougi, “the non-linear analysis of horizontal subgrade reaction and pile foundation”, proceeding of japan civil engineering proseminar(11), issn, pp.61-64. advances in systems science and applications (2012) vol.12 no.1 37 [10] kanaako naoru, tomi kougi, kokuhu taiyou. (2004), “experiment of lateral resistance of pile foundation and horizontal subgrade reaction”, proceeding of subgrade proseminar, vol.jgs39, pp.1515-1516. [11] cishun chen, weixhao xu, xinjie lv. (2000), “subgrade reaction coefficient of layer soil system”, journal of china mechanic academy. vol.16, no.1, pp.69-80. [12] jinchang wang, yekai chen. (2006), the usage of abaqus in civil engineering, zhengjiang university press. [13] suzuki yasuji. (2002), “the horizontal load experiment of subgrade reaction model”, proceeding of 37th subgrade proseminar, pp.1469-1470. advances in systems science and application (2015) vol.15 no.3 220-232 labour market model — focused on final products a.oinarov1, s. baizakov2, n.baizakov3 and jeffrey forrest4 1 the board of jsc “kazakhstan public-private partnership” 2 economic research institute, str. temirkazyk 65, astana, 0100000, kazakhstan; 3 association for “leadership in education development”, 36 saraishyk street, apt 105, astana, 0100000, kazakhstan; 4 school of business, slippery rock university, slippery rock, pa 16057, usa. abstract this paper studied history of development of commodity and cash flows balancing models. keynesian model , monetary model and mundell-fleming model are investigated, the issues of mutual consistency of indicators of economic growth is studied. at the same time , the mathematical formulation of market equilibrium of levels of production, employment, incomes and prices is gives. a new model of labour market focusing economy on the final product is established. keywords labour market; gross output; labour-intensity of products; interindustry balance; final product 1 market imbalance theory is an instrument to identify macroeconomic imbalances this section studied history of development of commodity and cash flows balancing models. there are assumptions that adequately reflected development of market forces of economic development. relationship between indicators of economic growth in both sectors of the economy was established. if keynesian model developed a balance between real economy and financial economics by using only one indicator, then monetary model used two, and mundell-fleming model used three indicators of economic growth. it also studied the issues of mutual consistency of indicators of economic growth. historically, analysis of production repeatability since 40-ies of the last century was conducted based on keynesian model, which assesses the market equilibrium between commodity and cash flows. using this model as an anti-crisis tool to eliminate the consequences of the great depression of the us economy in the 1930s is associated with the name of the us president f. roosevelt. circuit diagram of the replenishment cycle of this model is shown in figure 1. according to keynesian model we may admit the constancy of prices for goods and services, and consider investment growth and increase in other consumption costs as an impetus for economic development by the state. as prices for goods and services are constant, commodity and cash flows balancing based on this model is carried out due to one indicator nominal gdp. in other words, economic growth indicator is nominal gdp, which serves as an indicator of real advances in systems science and application (2015) vol.15 no.3 221 economic growth, as prices are constant, and volumes of real gdp and nominal gdp are the same because of constant prices. fig.1 principle circuit diagram of the replenishment cycle of keynesian model fig.2 principle circuit diagram of the replenishment cycle of monetary model in the mid-80s of the last century a model of economic management has 222 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... changed. keynesian model was replaced by more progressive model a model of monetary policy. the reason of progressivity of monetary model was that it considered variability of prices of goods and services that was clear in real life and which is a key advantage of the market economy.conditions of goods and services are rapidly renewed because of the influence of market competition of goods and services, and there will be their new prices, and prices of previous goods and services will change due to scientific and technological progress. monetary model in production repeatability analysis started to be actively used since the 1980s. its appearance on practice of economic management of the country is associated with the name of the president of r. reagan. monetary model means constant money velocity and it uses two gdp indicators: nominal and real. a real gdp is defined in prices of basic year and characterizes the dynamics of market forces of real economy. a nominal gdp is defined in prices of the year and characterizes the dynamics of market forces of monetary and financial system. gdp deflator connects their balancing between each other. but later it is turned out that a weak point of the model of monetary policy is assuming the constancy of money velocity, which is considered as a basis for monetary model. fig.3 principle circuit diagram of the replenishment cycle of monetary model advances in systems science and application (2015) vol.15 no.3 223 multinational companies and companies with developed foreign trade relations started to see the narrowness of this model during their economic analysis. in this regard, since the 1990s of the last century fleming mundell model was used, which overcame limitations of the monetary model. the advantage of fleming mundell model is that it admits that money velocity is dynamic. however, this model again assumes the constancy of prices for goods and services. production repeatability analysis according to fleming mundell model (since 1990) considers dynamics of exchange rate, which serves as a reliable development instrument of international relations of the world countries. it uses three gdp indicators: nominal, real and gdp based on purchasing power parity; figure 3 shows a principle circuit diagram based on fleming mundell model. firstly, mundell-fleming model assumes that prices are constant, and secondly, it links three gdps to each other. to ensure conceptual relationships of three gdps and stop assuming that prices for goods and services are constant, market equilibrium of levels of production, employment, incomes and prices was established. fig.4 principle circuit diagram of the replenishment cycle of kazakhstan model of double regulation 224 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... a new model of market equilibrium of levels of production, employment, incomes and prices in employment terms summarizes keynesian, monetary and fleming mundell models, and do not contradict their structure principles, but develops them. 2 economic content of kazakhstan model of double regulation this section gives mathematical formulation of market equilibrium of levels of production, employment, incomes and prices.. a new model of labour market focusing economy on the final product is established, which is shown in the example of the republic of kazakhstan. it explains why the market forces of labour and capital are developing according to the laws of free competition despite wishes of entrepreneurs to increase competitiveness of their companies. an article of american scientists was published in the column ”new world order” in magazine ”foreign affairs” about ”labour, capital and ideas in economy of power laws”, which states that the future belongs not to those who provide cheap labour or owns capital (they will be driven out by automation), but to those who are willing to introduce innovations and create products, services and business models based on them. new technologies do not only integrate existing sources of labour and capital, but also generate products that replace both labour and capital. according to the authors of the article, innovations in the nearest future, will lead to great changes, associated with distribution of incomes, as profit sources will be represented by high technology and globalization, in which this process is usually demonstrated as power laws. in this case, an employee in the usa will receive the same salary as employees working in the industrial sector in china or india. this is good news for developing countries, and problem for the united states and europe, where there are no possibilities of reducing production costs and monetary growth.can kazakhstan take advantage of the situation in the new world order? economic research institute (kazakhstan) applies a system of techniques to do forecasting and analysis. innovative model of the labour market plays an important role in it, which is focused on the final result. top modules of balance labour model are coefficients of direct labour intensity of products. to measure labour intensity we can use the unit of man-hours, man-years or jobs, and for their calculation we can use the table ”input output”, which are produced in kazakhstan every year: ti = li xi , i = 1, n (1) where ti coefficients of direct labour intensity of products; li number of people employed in the economy; xi gross output. advances in systems science and application (2015) vol.15 no.3 225 to measure labour intensity we can use the unit of man-hours, man-years or jobs. full coefficients of labour intensity per unit of production are calculated by coefficients of direct labour intensity: tj = n∑ i=1 aijti + tj , i, j = 1, n (2) where aij input-output coefficient i type of economic activity for issue j type of economic activity; ti, tj full labour intensity coefficients for production of yi or yj component of final product y ; if we consider that coefficients of direct t = (t1, t2, . . . , tn) and full t = (t1, t2, . . . , tn) labour intensities are raw vectors, then the ratio (2) can be written in matrix form: t = ta+ t (3) or t = t(e −a)−1 (4) where e is identity matrix, a is matrix coefficient, representing production process technology of this year. (e −a)−1 input coefficient matrix of production processes represented by full material costs in the table of interindustry balance. labour intensity turns into full material costs in accordance with duality principle of koopmans and kantorovich. if we introduce a notation = (e −a)−1 then formula (4) can be rewritten as follows: t = tb (5) as it is seen from the formula, the only link between direct and full labour intensity indicators of products is still input coefficient matrix of production processes. it means that coefficients of of full labour intensity of each goods and services provided are very important, which reflect the influence of production processes. if we consider l as a number of jobs in all sectors of national economy and apply module (1), we will get the following formula to calculate a number of people employed in the economy: l = n∑ j=1 lj = n∑ j=1 tjxj = tx (6) where x is gross output of goods and services. if in formula (1) we marked a number of people employed in the economy by the components of raw vector li, then we can have the record shown in formula 226 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... (6) by representing the same number of people employed in the economy as that of components of raw vector j. if we multiply both sides of (5) by y-final vector of production, we can have the following model of the labour market:= tby . and if we take into account the formula (6), then balancing model will have a complete form, where the basic law of labour market and capital is reflected: tx = ty (7) on the other hand, a similar assessment on the level of national economy of the country represents the average value formed by the ratio of nominal gdp to gross output: c = t t = y x left side (7) is equal to ty where t according to formula (5) is full labour intensity of products, nd y is the final product. and the cost of the final product is considered to be equal to nominal gdp. right side (7) is equal to tx where t according to the module (1) is direct labour intensity of products, nd is gross output of goods and services provided. based on this simple formula the yield of gdp (cdp (t)) is estimated as per turnover unit x(t). in fact, the gross output value x(t), which was correctly identified by famous american scientist harrington emerson in his main work “twelve principles of productivity” (1911), at business level it is ”total expenditure = material costs + labour costs + capital costs. or: total expenditure =qp+tw+tr; where compensation for intermediate materials, and other current material resources, used during production of goods is qp; remuneration of labour costs of employees is tw; remuneration of fixed capital is tr. thus, there is relationship between the final product and overall costs to create it x. so, by putting in place of t in the formula (7) its value from the formula (1) we will have t = l/x. then we will have l = ty . it means that a number of people employed in the economy is equal to multiplication of full labour intensity of products by nominal gdp. and full labour intensity of products is t , which in turn, is determined by multiplication of direct labour intensity of products t to full costs matrix.b = (e −a)−1. multiplicationt x according to (6) alsomeans number of working time fund used or number of jobs on economic activity, which are required to produce gross output (x), necessarry for production of the final product (y ). on the one hand, development model of the labour market (7) defines the advances in systems science and application (2015) vol.15 no.3 227 marginal valuation of the ratio of direct and full labour intensity of products for every type of economic activity, on the other a similar assessment at the level of national economy of the country represents the average value, formed by the ratio of nominal gdp to gross output: t/t = y/x (8) a number of people employed in the economy of state is similar to the total number of employees in all economic activities. therefore, new interpretation of the main situation of the market in labour terms is calculated by the model (7), which is completed by equation (8). in this case, the share of direct labour intensity of goods and services in their full labour intensity is equal to the level of production of the final product (y ) in the gross output (x). as a rule, destructive forces of laws of market equilibrium appear in macroeconomic imbalances, which are easily detected by analytical tools built on the basis of the table of interindustry balance. thus, according to the above formula (5), full material costs matrix allows to turn actual labour used during production of gross output (x) into abstract labour. however, this abstract labour, expressed in full labour intensity of products is needed to evaluate the effectiveness of production of a particular amount of the final product (y ). as defined in formula (1), direct labour intensity is expressed by the ratio of a number of people employed in every economic activity to the total volume of gross output, issued in this sector. multiplication of raw vector of direct labour intensity of a particular type of human activity by inverse matrix of material costs expresses a column vector of full labour intensity of production of the final product (y)). figure 5 shows a piece of values of direct and full labour intensity of products for certain types of economic activity of kazakhstan for 2012. as we can see from the diagram, the first indicator reflects the specific amount of jobs (man-years, man-hours, man-days) required to produce 1 million tenge worth of gross output. for example, in agriculture, hunting and forestry, it amounted for 965 jobs, and in production area of crude oil and natural gas only 10. figure 5 also shows that full labour intensity is greater in number than direct labour intensity. this is due to the fact that the first indicator reflects abstract labour that expresses the social form of unit cost of the final product produced, and the second represents economy of a particular type of human activity. here, the “black box”, which allows you to transfer direct labour intensity (t) into full (t ) is represented by inverse matrix of interindustry balance (b), t = t ∗ b. it also makes it possible to determine rapid development of technological improvement of production, since it is its dynamics (τ) that is a factor defining the rapidness of this process. that is why the advantages of new kazakh model, focused on 228 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... the final result, are provided on the basis of analytical calculations, conducted using reverse (τ)-matrix. the results are shown in table 1. fig.5 diagrams of the values of direct and full labour intensity, defined on the basis of the table of interindustry balance of kazakhstan for 2012. the first three columns show the components of final demand, translated into jobs, and the fourth column reflects total results. thus, we see that to cover final consumption, to have gross savings, and to attract export of the national economy we will require 6903.22, 1907.75 and 2227.89 thousand jobs, respectively. in general, the final demand (fourth column) can be satisfied if kazakhstan will have more than 11 million jobs. since supply of the country’s economy, according to the data of interindustry balance of 2012, is amounted to just tx = 8507000 jobs in total, then to cover the final demand of the economy additional capacity will be fully required, which must be repaid by import financing of foreign goods (the fifth column). in general, as we can see from the sixth column, kazakhstan’s jobs defined on the basis of labour intensity of product in calculation equal to 1 million tenge worth gross output ( by economic activity) significantly deviated from the jobs defined on the basis of full labour intensity (y ). however, their ratio (direct labour intensity 0,177 jobs, full 0.298) in terms of the national economy is equal to the ratio of scientific and technological excellence of real sector of the country, because we have the equation: c = t t = l x + l y = y x (9) advances in systems science and application (2015) vol.15 no.3 229 index value of the level of scientific and technological excellence of real sector of the country for kazakhstan is 0.587. this equation (9) of market equilibrium of levels of production, employment, incomes and prices in man terms is based on the balance of aggregate demand and aggregate supply. final demand in the economy of the country is generally determined by total costs of final consumption of households (y 1), general government costs (y 2), non-profit organizations (y 3), changes in inventories (y 4), purchase minus retirement of valuables (y 5), gross saving of fixed capital (y 6) and export (y 7). final demand = y 1 + y 2 + y 3 + y 4 + y 5 + y 6 + y 7 an aggregate supply, which is used to cover it, is defined as total of final demand and import with a minus sign (−y 8): final product = ∑ y i− y 8 thus, market equilibrium of levels of production, employment, incomes and prices is carried out by all agents of production, exchange, distribution and consumption. for example, to measure performance of the economy in working time we have the results of final demand in kazakhstan: 11,005,000 people. (8.505 million of which is covered by human resources within the country, and 2.5 million beyond the country). volumes of aggregate demand and aggregate supply are balanced, and they are equal to the work of 11,005,000 people, while direct labour intensity of products for production of 1 million tenge worth gross products in kazakhstan in 2012 amounted to 177 people, and full labour intensity 298 people. this equilibrium in formula (9) is defined in current prices and it expresses a market equilibrium. to assess the level of economic development of the country comparison base in comparable prices or quantum indexes should be determined based on this level. since nominal gdp (ngdp) is monetary terms of the final product and is defined in prices of this year, then, multiplying both sides of a well-known equation of gdp deflator: ngdp = b∗rgdp by purchasing power (pp), real value of the final product (q) can be defined, which corresponds to the balance of supply and demand in the country: q = pp ∗ngdp = c ∗rgdp, where c = pp ∗ pb = rgdp x ∗ ngdp rgdp = rgdp x (10) the main indicator of equation (10) is an indicator of scientific and technological progress (stp) c. its multiplication by real gdp is identical to multiplication of nominal gdp with purchasing power. it turns out that stp’s contribution helped to divide gdp deflator into two indicators. one of them 230 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... is purchasing power, the other is coefficient of scientific technological progress itself. as a result, the formula (10) balances the final product both in terms of money and terms of commodity. it is developed by generalizing and developing the equation of exchange of monetarists, which do not contradict its structure principles, but there are variables that are also represented by purchasing power and an stp index at real gdp. it means that the cost of the final product is equal to the nominal value of the product, i.e.q = pp*ngdp. therefore, multiplying both sides of equation ngdp = pb ∗ rgdp by purchasing power, a balance between nominal and real gdp can be set with weighing coefficient pp and c, respectively: pp ∗ngdp = c ∗ rgdp (11) it explains the viability of creative formula of scientific and technological progress: ċ c = ẏ y − ẋ x (12) rapid change of stp coefficient in equation (12) depends on difference in speed of the final product and gross output. let us assume that principle of using direct and full labour intensity to determine technical and technological excellence of the national economy is applied in the countries all over the world, which have developed foreign economic relations. then value of the dollar, euro, kazakhstani tenge and other currencies will be established according to the level of economic development within the state itself and its foreign relations. in this case, assessment principle of stp level deviations between the countries can be preserved, and real exchange rates of the national currencies will be set in accordance with rules of sdr design. such an approach can serve as an effective tool for building a methodology for sustainable development of the global economy: measured economy is managed economy. in conclusion, it can be summarized that it is possible to construct an analytical model for managing a market economy. it can be built in labour terms using the principle of double regulation and the table “input-output”. thus, balances of commodity and cash flows on economic activity in labour terms are formed with imbalances: ti ∗xi − ti ∗ yi = ±θi and at the level of the national economy of the state commodity and cash flows are always balanced: ∑ ti ∗xi − ∑ ti ∗ yi = 0 where components of final demand are yi defined, and their total value is the final demand, which corresponds to the final product of the country. advances in systems science and application (2015) vol.15 no.3 231 t a b le 1 b al an ci n g m o d el of th e la b o u r m a rk et o f k a za k h st a n a cc o rd in g to d a ta o f th e in te ri n d u st ry b al a n ce of 20 12 (p ie ce ). 232 a.oinarov, s. baizakov, n.baizakov and jeffrey forrest: labour market model — focused... these balance equations fully meet the conditions of globalization of the world economy and can be used to ensure market equilibrium of the levels of production, employment, incomes and prices anywhere in the world. the construction of an analytical model of market forces for development of real sector and financial system based on parity of their autonomous and parallel operation is the final result of this study. corresponding author s.baizakov can be contacted at: baizakov37@mail.ru advances in systems science and applications (2014) vol.14 no.3 294-301 estimates of the relative deviation for the equity home bias yirong ying1, ye ying1 and jeffrey forrest2 1 college of economics, shanghai university, shanghai, 200444, china 2 department of mathematics, slippery rock university, slippery rock, pa16057, usa, abstract there are a variety of explanations of equity home bias, the underlying difference being that different factors are emphasized, such as trade costs that impede international diversification, and the possibility that terms of trade responses to supply shocks provide international risk sharing. we found that with the ratio of world consumption of home vs. foreign goods and the ratio of world demand for home vs. foreign goods used for physical investment, we have three different estimates of the relative deviation about terms of trade, and then figured out three expressions representing these estimates. what we concluded helps to uncover the puzzle of the equity home bias. keywords international equity and bond portfolios, capital flows, current account 1 introduction though international capital flows have rallied with the liberalization of the global capital market two decades ago, equity home bias still exits in all industrial economies. abundant researches have been contributed to solve this puzzle. there are two major explanations of the continuously sizable equity home bias. the first one emphasizes on transaction costs and information barriers in cross-border financial transactions and suggests that international risk sharing is insufficient. the second one emphasized on the possibility that terms of trade changes in response to supply shocks may provide international insurance against these shocks, so that even a portfolio with home bias delivers efficient international risk sharing. seen from these very different opinions, where we start the work can make a big difference to what we conclude about the equity home bias. much published research made use of statistical methods and empirical analysis to interpret the home bias under capital account liberalization by transaction costs, information barriers and financial openness (recently from behavioural finance, cultural factors and legal system perspectives). ferreira and miguel, bekaert and wang and mondria and wu attributes much of home bias to information, familiarity and capital market openness [1-3]. another class of literature tried to decide home bias determinants through characterization of steady-state equilibrium portfolios in a general equilibrium model. engel and matsumoto analysed international equity portfolio choices in a advances in systems science and applications (2014) vol.14 no.3 295 model with money, sticky prices and trade in bonds [4-5]. under price stickiness, the short-run output is fixed, so that a positive productivity shock leads to a fall in employment and labor income, but an increase in profits. ownership of local equity is thus an effective hedge against labor income risk. heathcote and perri investigated the importance of physical investment for equity portfolios for the first time [6]. the hp model only generates realistic equity home bias if the terms of trade respond strongly to total factor productivity shocks. since the empirical evidence concerning the response of the terms of trade to technology shocks is mixed, it is important that our model does not require strong terms of trade effects of productivity shocksnevertheless, there is sizable equity home bias. the main contribution of this paper is that with the ratio of world consumption of home vs. foreign goods and the ratio of world demand for home vs. foreign goods used for physical investment, we have three different estimates of the relative deviation about terms of trade, and then figured out three expressions representing these estimates. 2 model we consider two symmetric countries, home (h) and foreign (f ), each with a representative household. country i = h,f produces one good using labor and capital. goods and financial assets (stocks and bonds) are traded in perfectly competitive markets. country i is inhabited by a representative household who lives in periods t = 0, 1, 2.... the household has the following life-time utility function: e0 ∞∑ t=0 βt ( c1−σ i,t 1− σ − l1+ω i,t 1 + ω ) (1) with w > 0. ci,t is country i’s aggregate consumption in period t and li,t is labor input. like much of the macroeconomics and finance literature, we take the coefficient of relative risk aversion to be greater than one: σ > 1. ci,t is a composite good given by: ci,t = [ a1/φ ( cii,t )(φ−1)/φ + (1− a)1/φ ( cij,t )(φ−1)/φ ]φ/(φ−1) , j ̸= i (2) where ci j,t is country i’s consumption of the good produced by country j at time t. ϕ > 0 is the elasticity of substitution between the two goods. in the (symmetric) deterministic steady state, a is the share of consumption spending devoted to the local good. we assume a preference bias for local goods, 1/2 < a < 1. the welfare-based consumer price index that corresponds to these preferences is: pi,t = [ a(pi,t) 1−φ + (1− a)(pj,t) 1−φ ]1/(1−φ) , j ̸= i (3) 296 yirong ying: estimates of the relative deviation for the equity home bias where pi,t is the price of good i. likewise, the associated investment price index is: p i i,t = [ ai(pi,t) 1−φi + (1− ai)(pj,t) 1−φi ]1/(1−φi) , j ̸= i (4) 3 relative deviation estimation coeurdacier, kollman and martin (2010) introduced a concept in their research on equity home bias, relative deviation. ∧ zt = (zt − z)/z denotes the relative deviation of a variable zt from its steady state value z. here, variables without a time subscript refer to the steady state. chh,t + cfh,t = p−φ h,t [ ach,tp φ h,t + (1− a)cf,tp φ f,t ] , cff,t + chf,t = p−φ f,t [ acf,tp φ f,t + (1− a)ch,tp φ h,t ] coeurdacier et al. (2010) derived from the above first-oder conditions that: yc,t ≡ chh,t + cfh,t cff,t + chf,t = q−ϕ t ωa [( pf,t ph,t )ϕ cf,t ch,t ] , with ωz(x) ≡ 1 + x ( 1−z z ) x+ ( 1−z z ) (5) where yc,t is the ratio of world consumption of home goods over world consumption of foreign goods, while qt ≡ ph,t/pf,t denotes the country h terms of trade. the ratio of world demand for home vs. foreign goods used for physical investment yi,t ≡ ihh,t+ifh,t iff,t+ihf,t can similarly be expressed as: yi,t ≡ q−ϕi t ωai (p i f,t p i h,t )ϕi if,t ih,t  (6) coeurdacier et al. (2010) found a zero-order portfolio such that the ratio of home to foreign marginal utilities of aggregate consumption (c−σ h,t/c −σ f,t ) is equated to the consumption-based real exchange rate (rert ≡ ph,t/pf,t), up to the following first-order condition: −σ ( ∧ ch,t− ∧ cf,t ) = ∧ rert (7) which is a linearized version of a risk sharing condition that holds under complete markets. it follows from the definition of home and foreign cpi indices (see equation 3) that: ∧ rert = ∧ ph,t− ∧ pf,t = (2a− 1) ∧ qt (8) advances in systems science and applications (2014) vol.14 no.3 297 due to consumption home bias (a > 1/2), an improvement of the home terms of trade leads to an appreciation of the home real exchange rate. lemma 1 if α, β are non-zero constants and α+β ̸= 0, then we have for continues functions f1 and f2 that ∧ (αf1(qt) + βf2(qt)) = αf1(q) αf1(q) + βf2(q) ∧ f1(qt) + βf2(q) αf1(q) + βf2(q) ∧ f2(qt) (9) proof : ∧ (αf1(qt) + βf2(qt)) = αf1(qt) + βf2(qt)− (αf1(q) + βf2(q)) αf1(q) + βf2(q) = α (f1(qt)− f1(q)) + β (f2(qt)− f2(q)) αf1(q) + βf2(q) = αf1(q) · ∧ f1(qt)+βf2(q) · ∧ f2(qt) αf1(q) + βf2(q) lemma 2 for continues functions f1 and f1, we have that ∧ (f1(qt) · f2(qt)) = ∧ f1(qt) + ∧ f2(qt) + ∧ f1(qt) · ∧ f2(qt) (10) proof : ∧ (f1(qt) · f2(qt)) = f1(qt)f2(qt)− f1(q)f2(q) f1(q)f2(q) = (f1(qt)− f1(q)) f2(qt) + f1(q)f2(qt)− f1(q)f2(q) f1(q)f2(q) = f1(qt)− f1(q) f1(q) · f2(qt) f2(q) + f2(qt)− f2(q) f2(q) = ∧ f1(qt) ( 1 + ∧ f2(qt) ) + ∧ f2(qt) = ∧ f1(qt) + ∧ f2(qt) + ∧ f1(qt) · ∧ f2(qt) lemma 3 for continues functions f1 and f1, we have that ∧( f1(qt) f2(qt) ) = ∧ f1(qt)− ∧ f2(qt) 1 + ∧ f2(qt) (11) 298 yirong ying: estimates of the relative deviation for the equity home bias proof : ∧( f1(qt) f2(qt) ) = f1(qt) f2(qt) − f1(q) f2(q) f1(q) f2(q) = f1(qt)f2(q)− f1(q)f2(qt) f1(q)f2(qt) = (f1(qt)− f1(q)) f2(q) + f1(q)f2(q)− f1(q)f2(qt) f1(q)f2(qt) = ∧ f1(qt) · f2(q) f2(qt) − f2(qt)− f2(q) f2(q) · f2(q) f2(qt) = ∧ f1(qt)− ∧ f2(qt) 1 + ∧ f2(qt) in order to improve the expression of equation 8, we need to modify equation 9, 10 and 11 and present as follows: lemma 4 for continues functions f1 and f1, we have that ∧ (αf1(qt) + βf2(qt)) = α ∧ f1(qt) + β ∧ f2(qt) α+ β (12) lemma 5 for continues functions f1 and f2, we have that ∧ (f1(qt) · f2(qt)) = ∧ f1(qt) + ∧ f2(qt) (13) lemma 6 for continues functions f1 and f1, we have that ∧( f1(qt) f2(qt) ) = ∧ f1(qt)− ∧ f2(qt) (14) obviously, lemma 4 is unique when f1(q) equals to f2(q) as is implied by ex-ante symmetry of two countries in our model, and lemma 5 and lemma 6 is one special case when ignoring second-order relative deviations and ∧ f2(qt) in the dominator, respectively. theorem 1 when equation 5 holds, then the relative world consumption demand for the home good obeys ∧ yc,t = − [ (1− (2a− 1)2)φ+ (2a− 1)2 1 σ ] ∧ qt (15.a) advances in systems science and applications (2014) vol.14 no.3 299 proof:since ∧ yc,t = ∧ q−ϕ t ωa [( pf,t ph,t )ϕ cf,t ch,t ] = ∧ q−ϕ t ωa(x), we derive from lemma 4-6 that: ∧ ωa(x) = ∧( 1 + x ( 1−a a ) x+ ( 1−a a ) ) = ∧( 1 + x ( 1− a a )) − ∧( x+ ( 1− a a )) = (1− a) ∧ x−a ∧ x = (1− 2a) ∧ x then, ∧ yc,t = ∧ q−ϕ t + ∧ ωa(x) = −ϕ ∧ qt+(1− 2a) ( −ϕ+ 1 σ ) (2a− 1) ∧ qt = − [ (1− (2a− 1)2)φ+ (2a− 1)2 1 σ ] ∧ qt and we can similarly prove for expression 2 and 3. theorem 2 when equation 5 holds, then the relative world consumption demand for the home good obeys ∧ yc,t = − [ (1− (2a− 1)2 a )φ+ (2a− 1)2 a 1 σ ] ∧ qt (15.b) theorem 3 when equation 5 holds, then the relative world consumption demand for the home good obeys ∧ yc,t = − [ (1− (2a− 1)2 1− a )φ+ (2a− 1)2 1− a 1 σ ] ∧ qt (15.c) note that λ > 0(as 1/2 < a < 1 implies 0 < 1 − (2a− 1)2). thus, an improvement in the home terms of trade lowers worldwide relative consumption of the home good. introduce the following signs, whose expressions are illustrated in fig.1: λ11 = 1− (2a− 1)2, λ12 = (2a− 1)2, λ21 = 1− (2a− 1)2 a λ22 = (2a− 1)2 a , λ31 = 1− (2a− 1)2 1− a , λ32 = (2a− 1)2 1− a 300 yirong ying: estimates of the relative deviation for the equity home bias fig.1 home terms of trade lowers worldwide relative consumption of the home good. acknowledges this research was supported by national natural science foundation of china (71171128) and research fund of program foundation of education ministry of advances in systems science and applications (2014) vol.14 no.3 301 china (10yja790233). references [1] ferreira, m.a. and a.f. migue. (2007). “the determinants of domestic and foreign bond bias”. working paper serious, may. [2] bekaert g. and x. s. wang. (2009), “home bias revisited”. nber working paper series, february. [3] mondria j. and t. wu. (2010), “the puzzling evolution of the home bias, information processing and financial openness”. journal of economic dynamics and control, vol.34, pp.875-896. [4] engel, c. and a. matsumoto. (2006), “portfolio choice in a monetary openeconomy dsge model”. nber working papers, may, pp.12214. [5] coeurdacier n., r. kollman and p. martin.(2010), “international portfolios, capital accumulation and foreign assets dynamics”. journal of international economics, vol.80, pp.100-112. [6] heathcote j. and f. perri. (2007), “the international diversification puzzle is not as bad as you think”. nber working papers, october, pp. 13483. corresponding author yirong ying and wenjun lv can be contacted at: yrying@staff.shu.edu.cn advances in systems science and application (2016) vol.16 no.1 47-61 development of a biodegradable additive from brown algae of the russian far east for the production of plastic packaging tatyanav. chadova, lyudmila o. korshenko, elena s. smertina, alexey e. nekrasov and natalya v. berlova far eastern federal university. vladivostok, russian; vladivostok branch of the russian customs academy, vladivostok. russian abstract the article presents the development of a biodegradable additive from brown algae of the russian far east for the packaging polymer industry. we consider a possibility of using cellulose derived from algae of the primorsky krai, as an additive to polymeric packaging products, which impartsthe properties of biodegradability to the polymer. a way to produce the biodegradable additive for plastic packaging is to reduce the duration of the technological process, increase the yield and quality of the end product and expand the resource base through the use of cheap and non-traditional local raw materials. keywords biodecomposition; biodegradation, brown algae; biodegradable packaging, polymeric materials. 1 introduction russia annually produces about 70 million tonnes of municipal solid waste (msw), the major share of which (more than 50%) refers to packages. up to 10 thousand hectares of land is annually alienated for msw landfills and dumps. this land also includes fertile lands, withdrawn from the agricultural use. currently, in the russian federation measures to implement new investment projects in the sphere of solid waste management are developed and being taken[1]. only 3% of this amount are recycled, and the rest are incinerated or disposed of in landfills. however, incineration is an expensive process that also leads to the formation of the highly toxic (such as furans and dioxins) compounds. time needed for degradation of packaging materials in natural conditions may range from a few decades to hundreds of years, the use of biodegradable materials in the packaging industry will lead to a significant reduction of this period. therefore, the topic of the development of biodegradable materials is relevant and is of practical and scientific interest. polymeric materials have become part of our lives and have replaced wood, metal, glass, ceramics. a large number of scientific papers and reviews[2-20] are dedicated to the studies on the degradation of polymers. according to the international organization for standardization, biodegradable plastics are polymers the degradation of which takes place under the influence of microorganisms, bac48 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... teria and fungi. the use of such plastics minimizes the adverse effects on the environment ecology. the degradation rate depends on a number of factors such as the type of a polymer, the type and concentration of degrading materials, moisture, temperature and a series of other factors. the corresponding public opinion and the legislative ways of control and regulation contribute to the accelerated proliferation of the production technologies of these packaging materials. in many countries, the government prohibits the use of plastic bags and disposable plastic utensils made of traditional types of plastic, and one has to pay for plastic bags in supermarkets, while paper bags are free. in recent years, the production of packaging materials is the leading sector of the economy and is growing worldwide, including in russia. based on this, one of the relevant directions is the production of environmentally friendly biodegradable packaging. the development of new highly efficient technologies, the creation of products based on plant raw materials, their further assimilation and standardization become more popular due to the wide spectrum of biological effects. particular attention is paid to the need to use local plant raw materials with sufficiently large natural reserves of natural or cultivated natural sources that are promising in terms of their introduction into production. one of the ways to produce biodegradable materials is to create composite materials based on natural biodegradable polymers such as cellulose, chitosan, gelatin, polypeptides, casein, and others. this research work aims to study the possibility of using local plant raw material brown algae of the primorsky krai in order to create a new biodegradable polymeric composition, which along with the high molecular basis comprises organic fillers (starch, cellulose) that should serve as a nutrient medium for microorganisms. the advantage of biogenic packaging made using local algae lies precisely in the high speed of growth some species can grow up to a meter during twelve hours. 2 main part at the department of commodity research and examination of goods of the far eastern federal university, the work was undertaken to develop a method for producing cellulose from brown algae of the primorsky krai. the main objective of the development was to provide an efficient method for producing cellulose from sea algae, mostly of the fucus genus, the use of which as a raw material makes it possible to clean coastal areas littered with the cast ashore algae and improve their ecological condition. today, modern biotechnology is aimed at the production of both traditional and new types of substances and materials. in addition, at this point, renewable plant raw materials can become a solution. such technology can be successfully advances in systems science and application (2016) vol.16 no.1 49 used for the further development of the industry in the far east. today, many physical and technical characteristics of biodegradable polymers are not inferior to the characteristics of conventional plastics. biodegradable polymers are degradable in natural conditions under the influence of such environmental factors as light, temperature, humidity and with the participation of living microorganisms as well. simultaneously, high-molecular substances degrade into low molecular ones, which are safe for the environment. another positive characteristic of biopolymers is that they decompose in a short time from several months to several years. thus, the creation of compositions that along with the high molecular basis include organic fillers (starch, cellulose, amylose, amylopectin, dextrin, etc.), which are a nutrient medium for microorganisms is a promising trend in the packaging industry. by the general structure, the majority of algae cells are similar to the cells of green plant such as corn or tomatoes; they have rigid cell wall, which consists primarily of cellulose, hemicellulose and pectins. cellulose pulp, alkyl ketones and compounds containing carbonyl groups are the catalysts of photo and biodegradation of the films based on polyethylene, polypropylene, or polyethylene terephthalate. photo and biodegradation of such films begins after 8-12 weeks, the remnants of the film completely disappear during harrowing and ploughing, thus loosening the soil. the main commercial algae is brown algae, so this paper puts more emphasis on the representatives of this division. brown algae (phaeophycophyta) are spread exclusively in the seas; their appearance resembles the one of higher plants. representatives of brown algae have a yellowish-brown colour of thalli, which is due to the presence of chlorophyll and carotenoids, as well as the brown pigment from the group of xanthophylls-fucoxanthine. their vegetative body (thallus) is of complex internal and external structure. the thallus is composed of several types of tissues with different functions. the chemical composition of brown algae is highly dependent on the species, season and habitat. according to the literature, brown algae contains from 2.5 to 17% of algal cellulose, which, due to certain differences from the usual, is called eucellulose[21]. extraction of cellulose from natural materials is based on the action of reagents that solve or destruct non-cellulosic components contained in plant tissues (proteins, fats, waxes, resins, lignin and polysaccharides cellulose companions). extraction methods depend on the type of plant material and designation of cellulose. the main method is alkaline pulping processing of plant materials with diluted sodium hydroxide solution. in the course of the research, we obtained the optimal technological parameters of the cellulose extraction process, thus allowing us to increase the yield of the final product. 50 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... the authors propose a new method for producing cellulose from non-wood cellulose-containing raw materials, which comprises the milling of dried raw material, delignification with dilute acid by heating followed by separation of the solid phase, which is treated with the aqueous solution of sodium hydroxide under heating, and then the desired product is filtered off, washed with water and dried. an important advantage of the method proposed by the authors is the possibility to improve the sanitary condition of the coastline sections cluttered with deposits of cast ashore dry algae and improve the environmental situation in general. the study found that raw algal cellulose fibres are very small, in the shape of irregular flakes and dissolve in the schweitzer reagent. thus, sea algae is an ideal raw material for bio-packaging material. 3 methodology the method consists in the purification of large extraneous impurities of dried brown algae, mostly of the fucus genus, followed by refinement through grinding and sifting. 4 to 6% solution of sulfuric acid (2so4) is added to the sifted fine fraction with the 1: 45-55 hydromodule, and is heated at the boiling point of the mixture (100 to 105◦c) for 5-10 minutes under reflux with continuous mechanical stirring. the reacted and cooled down mixture is filtered, and the mass precipitated on the filter is washed with hot water (60 to 70◦c) until neutral and dried to constant weight. then, 5-10% solution of the koh alkali is added to the resulting solid residue, followed by hot distilled water (from 60 to 70◦c) in equal amounts calculated to provide a hydro module of 1/50-60 at an alkali solution concentration of 2.5-5.0%, and it is heated at the boiling point of the mixture for 5-10 minutes with continuous mechanical stirring. the cooled mixture is filtered with it being rinsed with water as well. after complete filtration, the solid residue in the filter, which is the desired product, is washed and dried in natural conditions. after it is dried, it is in the form of brown powder, which goes into the reactions characteristics for cellulose: it is dissolved in schweizer’s reagent and forms a gelatinous cellulose mass if diluted acid is added to the resulting solution. lignin-like residual content and other organic substances in the resulting product amounts to no more than 5-8 wt.%. the proposed method differs from the known methods[22-25], as local sea brown algae is used as the non-wood cellulose raw material; at the same time, the delignification is carried out by heating in a 4-6% solution of sulfuric acid (s/l = 1/45-55) at the boiling point temperature of the mixture for 5-10 min; the alkaline treatment of the obtained solid residue is carried out after washing with water by heating in a 2.5-5.0% solution of potassium hydroxide koh (s/l = 1/50-60) at the boiling point for 5-10 min. advances in systems science and application (2016) vol.16 no.1 51 technological parameters of the proposed method ensure efficient processing of specific algal raw material with the product being based on cellulose, containing lignin residues (lignin-like substances) and trace content of other organic substances. the conducted field tests showed that the resulting cellulosic semi-finished product acts as an effective filler of composite polymeric materials that serves as a nutrient medium for microorganisms, destroying high molecular basis, and can be successfully used in the production of various types of packages and containers with limited lifetime in particular. for a more complete utilization of the algal mass and higher profitability the author suggest using acidic and alkaline filtrates for obtaining a valuable nutrient of the fertilizer, for this purpose, it is necessary to mix and separate them for the formation of the flocculated settling fraction, which is washed with water until neutral and dried. discussion.there are certain developments that envisage the use of different plant fillers as natural bio additives: flax waste, wheat stalks, corn stalks, wastes of sawmill and wood chemical industries, food production waste, etc. today polymer blends are widely used (most often polyethylene) with corn and starch, biodegradable polymers with natural additives cellulose, soy flour, spent grains. the combination of a synthetic polymer and a natural material can give a new set of properties. the authors studied the technical level in the sphere of development of the biodegradable additives into polymeric materials of a wide application range; a patent search is conducted, which revealed that many scientists, both in russia and europe mainly work in the field of biodegradable materials that are used as packaging materials[26-46]. from the perspective of rational environmental management as a natural supplement, it is feasible to consider the resources that one or another region is rich with. the fauna of the far east is a rich source of unique plants. algae make up the bulk of plant organisms in bodies of water and are the most productive plants. large amounts of sea weed and algae are cast ashore in the promorsky krai; these plants contain fiber plant cellulose, which, as the authors suggest, can be used as a natural filler for biodegradable composition. it is the study on the degradation of algae that a large number of scientific papers and reviews are dedicated to [47-57]. by its general structure, most algae cells have similar cells to the cells of green plants such as corn or tomatoes. the studies have found that algae (cyanobacteria) are capable of synthesizing cellulose, which can be used as a biodegradable additive. the far east is rich with brown algae, so in our work we focused on representatives of this division. the chemical composition of brown algae is highly dependent on the species, 52 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... season and habitat. it is found that brown algae contain up to 17% of algal cellulose, eu-cellulose, which is structurally similar to cellulose of higher plants, and thus it is a potential source[57]. the use of dry brown algae, which can be considered as non-wood cellulose waste, allows expanding the raw material base for producing cellulose, reduce the consumption of more valuable wood and get cheaper products due to the use of affordable renewable non-conventional raw materials. each plant contains cellulose, but not every plant is suitable for industrial extraction of cellulose. when choosing raw materials, the content of fibre is important, as well as structural features of the constituent cellular tissues, the possibility of using industrial processing methods, the quality of the product obtained as a result of this treatment, the prevalence of plant raw materials, cost of collection, storage. despite the proven similarity of the algal cellulose with the cellulose of higher plants in structure and properties, with the use of dry brown algae as cellulosecontaining raw materials, the known cellulose production methods are ineffective and cannot exhibit sufficiently convincing results. the reason for this is that algal fibre contains lignin-like substances that are difficult to remove as they have strong bonds with it, as well as due to the residual amount of alginic acid, prone to certain conditions, to gel formation, and in some cases chlorophyll formation. there is a known method for manufacturing cellulose powder from various kinds of lignocellulosic materials derived from wood semi-finished products in the course of their processing on the cellulose and paper plants, straw of herbaceous plants, waste paper (rf patent. no. 2478664, published on 2013.04.10) which includes the destruction of said materials via exposure to a lewis acid solutions of low concentration and an organic solvent with stirring, followed by washing and drying of the desired product[24]. if there is a common idea consisting in the use of lewis acids, this method includes unspecified number of variants, each of which requires an independent selection of optimum conditions for obtaining powdered cellulose: concentration values and the type of lewis acid in various organic solvents, liquid module, temperature and a destructive treatment of the lignocellulosic material, the intensity of stirring and the conditions of drying of the end product, depending on the used raw material. furthermore, cellulose production in the known manner on an industrial scale involves the use of substantial quantities of organic solvent, requiring certain safety precautions during handling and disposal. when this method is used for dry brown algae as a cellulose-containing raw material, it is ineffective and does not provide the production of the necessary target product. there is a method of cellulose production from paddy straw (rf patent no. 2418122, published on 2011.05.10) comprising two pulping stages, the first of advances in systems science and application (2016) vol.16 no.1 53 which is carried out in an alkaline medium with subsequent separation of the cellulose product, and the second-in an acidic environment, with a mixture of peracetic acid, acetic acid and hydrogen peroxide in the presence of a stabilizer, in this case, mixtures of organ ophosphonates, wherein second pulping stage is carried out in the presence of ozone at a rate of 2-4 g/hr[25]. the disadvantage of this method is the difficulty and cost of technology that are due to the need to conduct the process in the presence of ozone and due to the use of hydrogen peroxide, which is subject to rapid decomposition, and organic stabilizers. furthermore, the known method is also ineffective when dry brown algae is used as raw cellulose-containing material. there is a method of cellulose production from straw (rf patent no. 2423570, published on 2011.07.10) comprising impregnation in the reactor and maceration of chopped straw with the aqueous solution of sodium hydroxide with a concentration of 20-30g/l in na2o units at a temperature of 30-80◦c at a ratio of the solution weight to the weight of dry chopped straw of 7:1[24]. the impregnated chopped straw is maintained at a predetermined temperature for 30 minutes, and then the liquid phase is withdrawn. heated water is added to the mass, the mass temperature is raised to 96◦c and the pulping is carried out at this temperature for 2 h 30 min. when used as a dry cellulosic raw material of brown algae containing non-hydrolysable substances strongly bonded with fibre and similar in elemental composition to lignins of higher plants, the known method is inefficient and does not provide a marketable product. the closest to the claimed method is a methodof producing cellulose from nonwood plant raw material with the content of the native cellulose being not more than 50%, described in the rf patent no. 2448118, publ. on 2012.04.20[27]. according to the conventional method, the raw material is washed with water at 40-70◦c and atmospheric pressure for 0.5-4 hours, and processing is carried out with an aqueous solution of nitric acid with a concentration of 2-8% at 90-95◦c for 4-20h followed by separation of the solid phase, which is treated with the aqueous solution of sodium hydroxide with a concentration of 1-4% at 60-95◦c for 1-6 hours. a disadvantage of this method is the long duration of the technological process, in addition, with the cellulose extraction from dry brown algae, the fibre of which contains lignin-like components and alginic acids that are hard to separate, it is possible to use the method with positive results and it does not provide an acceptable yield and quality of the end product due to the complexity of removing the tightly bound substances contained in algal cellulose and, on the one hand, and the partial hydrolysis of the fibre itself after prolonged treatment, on the other hand. result.algae are a non-food biomass, the use of which does not pose a threat to the production of packaging and food safety. algae grow 20-30 times faster 54 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... than land plants (some species can double their mass several times a day). the lack of a hard shell and almost no lignin makes the processing technology easier and more efficient than biomass processing of any ground material. it should be noted that the intensification of research in the field of polymers is important not only for further successful development of the market of biodegradable plastic packaging, but is one of the most promising ways to solve the global ecological problem related to the environmental pollution with polymeric materials wastes. the proposal to use local raw materials as raw materials base can solve the problem of non-waste production in the industry. among the possible priority areas we can note the development of ocean resources and the ecology, new technologies and materials. today, modern biotechnology is aimed at the production of both traditional and new types of substances and materials. moreover, at this point, renewable plant raw materials can become a solution to this problem, and the proposed technology for producing biodegradable additives for plastic packaging can be successfully used for the further industrial development in the far east region of russia. the development presented in the article is an invention (patent no. 2556115 russian federation, ipc s08b15/02 (2006.01), appl. dated 15.04.14, published on 07.10.15), which relates to algae processing methods, in particular, of the sea brown algae cast ashore; this method can be used to produce fillers for synthetic polymersthat provide biodegradation of polymer compositions and are required for manufacturing materials with a controlled lifetime and to producecellulose semi-finished products used as raw materials for the chemical industry, in the manufacture of paper and cardboard. cellulose is not a product intended for direct consumption: cellulose obtained after a single technological cycle serves as a raw material for other processing cycle; varying degrees of cellulose purification determine its various applications. an important advantage of the proposed method is the possibility to improve the sanitary condition of the coastline sectionscluttered with deposits of cast ashore dry algae and improve the environmental situation in general. 4 conclusions 1. the technical result of the proposed method for producing the biodegradable additive from the far east brown sea algae for plastic packaging is to reduce the duration of the technological process, increase the yield and quality of the end product and expand the resource base through the use of cheap and nontraditional local raw materials. 2. in the course of the research, we obtained the optimal technological parameters of the cellulose extraction process, thus allowing us to increase the yield of advances in systems science and application (2016) vol.16 no.1 55 the final product. 3. technological parameters of the proposed method ensure efficient processing of specific algal raw material with the end product being based on cellulose, containing lignin residues (lignin-like substances) and trace content of other organic substances. 4. the resulting cellulosic semi-finished product acts as an effective filler for composite polymeric materials that serves as a nutrient medium for microorganisms, destroying high molecular basis, and can be successfully used in the production of various types of packages and containers with limited lifetime in particular. 5. during the study, the authors found that raw algal cellulose fibres are very small, in shape of irregular flakes and dissolve in the schweitzer reagent. thus, sea algae is an ideal raw material for bio-packaging material. 6. for a more complete utilization of the algal mass and higher profitability, we propose to use acidic and alkaline filtrates for obtaining a valuable nutrient of the fertilizer, for this purpose, it is necessary to mix and separate them for the formation of the flocculated settling fraction, which is washed with water until neutral and dried. 7. the novelty of the technological solution of cellulose production from brown sea algae is confirmed by rf patent no. 2556115 “the process for producing cellulose from brown sea algae”. acknowledgement the authors are deeply grateful to their colleagues victor e. vaskov,olesya a. zdor, margarita d. boyarova, who are not the authors, but with whose assistance the study was conducted. references [1] the committee on environment and natural resources. (2014), available at: www.magas. bezformata.ru/listnews/tonn-tverdih-bitovihothodov/23761174/. [2] perrin f.x., merlatti c., aragon e. and margaillan a. (2009), “degradation study of polymer coating: improvement in coating weatherability testing and coating failure prediction”, progress in organic coatings, vol.64, no.4, pp.466-473. [3] eglin d., mortisen d. and alini m. (2009), “degradation of synthetic polymeric scaffolds for bone and cartilage tissue repairs”, soft matter , vol.5, no.5, pp.938-947. 56 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... [4] ito m and nagai k. (2008), “degradation issues of polymer materials used in railway field”, polymer degradation and stability, vol.93, no.10, pp.1723-35. [5] lucas n, bienaime c, belloy c, queneudec m, silvestre f and nava-saucedo je. (2008), “polymer biodegradation: mechanisms and estimation techniques”, chemosphere , vol.73, no.4, pp.429-442. [6] hakkarainen m. (2007), “solid phase microextraction for analysis of polymer degradation products and additives ”,chromatography for sustainable polymeric materials, pp.23-50. [7] jorgensen m, norrman k and krebs fc. (2008), “stability/degradation of polymer solar cells”, solar energy materials and solar cells, vol.97, no.7, pp.686-714. [8] borup r, meyers j, pivovar b, kim ys, mukundan r, garland n et al. “scientific aspects of polymer electrolyte fuel cell durability and degradation”, chemical reviews, vol.107, no.10, pp.3904-51. [9] zeynalov e. and friedrich j.f. (2007), “anti-radical activity of fullerenes and carbon nanotubes in reactions of radical polymerization and polymer thermal/thermo-oxidative degradation”, materialprufung, vol.49, no.5, pp.265-270. [10] jayasekara r., harding i., bowater i. and lonergan g. (2005), “biodegradability of a selected range of polymers and polymer blends and standard methods for assessment of biodegradation”,journal of polymers and the environment , vol.13, no.3, pp.231-251. [11] gu jg and gu jd. (2005), “methods currently used in testing microbiological degradation and deterioration of a wide range of polymeric materials with various degree of degradability: a review”, journal of polymers and the environment, vol.13, no.1, pp.65-74. [12] kikkawa y. and doi y. (2003) “surface morphology and enzymatic degradation of biodegradable polymers”, sen-i gakkaishi, vol.59, no.10, pp.313-318. [13] santovena a, alvarez-lorenzo c, concheiro a, llabres m and farina jb. (2004), “rheological properties of plga film-based implants: correlation with polymer degradation and spf66 antimalaric synthetic peptide release”, biomaterials, vol.25, no.5, pp.925-931. [14] gu j.d. (2003), “microbiological deterioration and degradation of synthetic polymeric materials: recent research advances”, international biodeterioration& biodegradation, vol.52, no.2, pp.69-91. advances in systems science and application (2016) vol.16 no.1 57 [15] lucarini m., pedulli g.f. and motyakin m.v. (2003), “schlick s. electron spin resonance imaging of polymer degradation and stabilization”, progress in polymer science , vol.28, no.2, pp.331-340. [16] santerre jp, shajii l. and leung bw (2000), “relation of dental composite formulations to their degradation and the release of hydrolyzed polymericresin-derived products”, critical reviews in oral biology & medicine,vol. 12, pp.136-51. [17] wilkie ca(1999), “tga/ftir: an extremely useful technique for studying polymer degradation”, polymer degradation and stability,66:301-06,1999. [18] “degradation of polyethylene terephthalate (pet) by lignolytic fungi”(2000), international biodeterioration & biodegradation, vol. 55, pp.30809. [19] reiner cn. (2004), “fluid catalytic cracking degradation of polyethylene terephthalate by various catalysts”, abstracts of papers of the american chemical society ,vol. 227,643-ched, pp.119-135. [20] gijsman p, meijers g and vitarelli g. (1999), “comparison of the uvdegradation chemistry of polypropylene, polyethylene, polyamide 6 and polybutylene terephthalate”, polymer degradation and stability, vol 65, pp.43341. [21] cellulose and its derivatives (1974), bayklz n. and l. segal.(ed.), mir. [22] pat. ru no2478664. (2013), “a method for producing cellulose powde”, kuvshinov la, frolova sv, av kuchin, published . [23] pat. ru no2418122(2011), a method for producing pulp from rice straw: vurasko av driker bn, mertin ev, galimov ar, chistyakov kn, publ. [24] pat. ru no2423570(2011), a method for producing pulp from straw: pazukhin ga, moncef sh.r., publ. [25] pat. ru no2448118(2012), a method for producing pulp from non-wood plant raw material with the content of the native cellulose is not more than 50% and a method for producing therefrom carboxymethylcellulose, publ. [26] pat. ru no 2352597(2008) biodegradable granular polyolefin composition and method of preparation. alexander ponomarev, t-ka no 2008125461/04, priority 2008.06.25, publ. 2009.04.20.32.; pat. ru no 2114865 biodegradable crosslinked polimery.nikomed imaging as not (no). 3. no94045151 / 04, priority 1992.03.06, publ. 1998.07.10.. 58 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... [27] pat. ru no 2114865(1992) biodegradable crosslinked polimery.nikomed imaging as not (no). 3. no94045151 / 04, priority 1992.03.06, publ. 1998.07.10. [28] pat. ru no 2095379(1992) polymer composition for producing biodegradable moldings and biodegradable moldings novamonts.p.a. (it). 3. no 92016573/04, priority 04.05.1992, publ. 10.11.1997. [29] pat. ru no ru 2174132(2000) production of plastics based on natural polymers used in the production of thermoformed articles of various configurations. moscow state university of applied biotechnology, w-ka no 2000116003/04, priority 23.06.2000, publ. 27.09.2001. [30] pat. ru no 2180670(2000) biodegradable thermoplastic composition based krahmala.vserossiysky research institute starch, t-ka no 2000100058/04, priority 06.01.2000, publ. 20.03.2002. [31] pat. number ru 2349612(2007) a biodegradable thermoplastic composition using the waste confectionery. state educational institution of higher professional education “moscow state university of food production”, ministry of education of the russian federation, the w-ka no2007141897 / 04, priority 14.11.2007, publ. 20.03.2009. [32] pat. s no es2315533(2002) water-solublecontainerobristclosuresswitzerland, z-ka es20030762682t, submitted 2003.06.26 priority 2002.07.03, published. 2009-04-01. [33] pat. britain no gb 2452227(2007) packagingmaterialsportabrandsltd. zka no gb20080009345, submitted 2008.05.23 priority 2007.08.28, published. 2009-03-04. [34] pat. korea no kr20000005282(2006) water-soluble bag packaging system and packaging method novartis ag [ch]; chris chraftind products inc [us]. z-ka no kr19987007990 submitted 1998.10.08, priority 1996.04.08, published. 2000-01-25. [35] pat. france no fr2691467(1993) hot water soluble saccharide based sheet material for thermo-forming rigidpackaging containers comprises starch based saccharide, alkali salt of proteinaceous material, organic plasticiser and water gervaisdanone sa. z-ka no fr19920006279, submitted 1992.05.22, published. 1993-11-26. [36] pat. united states no us7402618(2006) biodegradable composition for the preparation of tableware, drink container, mulching film and package and advances in systems science and application (2016) vol.16 no.1 59 method for preparing the samexu; hao z-ka no 11/418,628, submitted may 6, 2006, published. july 22, 2008. [37] pat. united states no us7368160(2005) packaging film biax international inc.ca z-ka no 11/142,296, submitted june 2, 2005, published. may 6, 2008. [38] pat. united states no 7288196(2005) packagingmaterialandinformation recording by packaging informationrecording media packaged by packaging material.science applications international corporation ca. z-ka no 11/101,441, submitted april 8, 2005, published. october 30, 2007/. [39] pat. united states no 71324637(2003) biodegradable composition having improved water resistance and process for producing sameyoul chon chemical co., ltd.kr. z-ka no 10/490,671 submitted june 9, 2003, pct filed june 09, 2003 pct no.: pct/kr03/01120, published. november 7, 2006. [40] pat. united states no 7083673(2004)biodegradable or compostable containers new ice limitedgbno10/977,082, submitted october 29, 2004, published august 1, 2006. [41] pat. united states no 706765(2002) non-synthetic biodegradablestarch-based composition for production of shaped bodieskasetsart university, bangkok, th z-ka no 10/282,205, submitted october 29, 2002, published. june 27, 2006. [42] pat. united states no 6878199(2003) biodegradable orcompostablecontainersnew ice limited (gb) z-ka no 10/341,288, submitted january 13, 2003, published. april 12, 2005. [43] pat. united states no 6723264(1997) method of making biodegradablepackaging material bussey, jr.; harry (fl), bussey, iii; buddy harry (nj) z-ka no 08/867,771, submitted june 2, 1997, published april 20, 2004. [44] pat. united states no 6706684(2001) method for preparing a collagen material with controlled in vivo degradation imedexbiomateriaux,trevoux, fr. z-ka no09/856,858, submitted may 29, 2001, pct filed november 09, 1999 , published. march 16, 2004. [45] pat. united states no 6573340(2003) biodegradable polymer films and sheets suitable for use as laminate coatings as well as wraps and otherpackagingmaterials.biotecbiologischenaturverpackungen gmbh & co. kg (de) z-ka no 09/648,471.submitted09/648,471, published. june 3, 2003. 60 tatyanav. chadova, lyudmila o. korshenko, etc.: development of a biodegradable... [46] pat. united states no 6337097(2002) biodegradable and edible feed packaging materialskansas state university research foundation (manhattan, ks). z-ka no09/408,373 submittedseptember 29, 1999, published. january 8, 2002. [47] caceres tp, megharaj m and naidu r. (2008), “biodegradation of the pesticide fenamiphos by ten different species of green algae and cyanobacteria”, current microbiology, vol. 57, pp.643-46. [48] warshawsky d, ladow k and schneider j. (2007), “enhanced degradation of benzo[a]pyrene by mycobacterium sp in conjunction with green alga”, chemosphere, vol. 69, pp.500-506. [49] alderkamp ac, van rijssel m and bolhuis h. (2007), “characterization of marine bacteria and the activity of their enzyme systems involved in degradation of the algal storage glucanlaminarin”, fems microbiology ecology, vol. 59, pp.108-117. [50] zhou c.s. and ma h.l. (2006), “ultrasonic degradation of polysaccharide from a red algae (porphyrayezoensis)”, journal of agricultural and food chemistry, vol. 54, pp.2223-2228. [51] meneses c.g.r, saraiva l.b., mello hnd., de melo jls and pearson hw. (2005), “variations in bod, algal biomass and organic matter biodegradation constants in a wind-mixed tropical facultative waste stabilization pond”, water science and technology, vol. 51, pp.183-190. [52] hirooka t., nagase h., uchida k., hiroshige y., ehara y., nishikawa j. et al. (2005), “biodegradation of bisphenol a and disappearance of its estrogenic activity by the green alga chlorella fusca var”, vacuolata. environmental toxicology and chemistry, vol. 24, pp.1896-1901. [53] michel g., helbert w., kahn r., dideberg o. and kloareg b. (2003), “the structural bases of the processive degradation of iota-carrageenan, a main cell wall polysaccharide of red algae”, journal of molecular biology, vol. 334, pp.421-433. [54] regel r.h., ferris j.m., ganf g.g. and brookes j.d. (2002), “algal esterase activity as a biomeasure of environmental degradation in a freshwater creek”, aquatic toxicology, vol. 59, pp.209-223. [55] ding hb and sun my. (2002), “biodegradation of algal fatty acids in oxic and anoxic systems: effect of structural association and relative roles of aerobic and anaerobic bacteria”, abstracts of papers of the american chemical society, vol. 223, 067-fuel. advances in systems science and application (2016) vol.16 no.1 61 [56] ivanova e.p., bakunina i.y., sawabe t., hayashi k., alexeeva y.v., zhukova n.v. et al .(2002), “two species of culturable bacteria associated with degradation of brown algae fucusevanescens”, microbialecology, vol. 43, pp.242-249. [57] avakova o.g. and bogolitsyn k.g. (2004), “vegetable fiber: structure, proper-ties, applications”. publishing house: university, forestry journal, no. 4.c, pp.122-129. corresponding author tatiana can be contacted at: yal05@mail.ru microsoft word 17 y. g. jiang--the application of norm functional multi-dimensional space theory in dynamic design.doc advances in systems science and applications (2011), vol.11, no.3-4 331-338 issn 1078-6236 international institute for general systems studies, inc the application of norm functional multi-dimensional space theory in dynamic design y. g. jiang machine-electrical engineering department, chengdu electromechanical college, chengdu, 610031, china abstract the dynamic design for space over three-dimension is a tough problem in mechanical design, because there are no proper cognition and analysis methods for multi-dimensional space. so, more abstract and comprehensive methods are in need. normed function space theory is a fully-fledged mathematics theory for multi-dimensional space, which is a powerful aid in analysis and representation. by using arbitrary function, dynamic design problem can be solved, such as the description of multi-variable system, relation modeling and optimization. variable systems with totally different physical meanings may have an identical norm function structure, which makes the abstraction and analysis of complex problem easier. the design is the modeling based on physical essence and the analysis and calculation of model, instead of traditional geometry analysis. the method is not analogy but precise calculation, which can produce an optimum solution in norm function. centrifugal governing system is a typical mechanical power system. according to dynamic requirement, the energy arbitrary function is set up and solved. that completes the mechanical structure design. the example shows, the arbitrary function design is workable and digitized. keywords multi-dimensional space route planning dynamic design norm function space 1.introduction along with the development of science and technology, the major task in mechanical design is the dynamic design of multi-variable system. the dimensions of multi-variable dynamic space are over three, at least four including time, without visual geometry. the major task of dynamic design is the representation and analysis of ultra three-dimension space. nowadays, the mechanical design is basically in analogy and static design stage [1-2]. it’s difficult to represent and analyze dynamic space without scientific method to express multi-dimension space. in mechanical structure design, the following problems cannot be solved at the same time: the representation of motion position and performance in mechanical drive system, the elastodynamic characteristic of mechanism motion, the time-sharing occupation of space, the spatial encounter and capture of moving object, the influence of different motion parameter on mechanical and motion characteristics, the structure of high-speed mechanism and dynamic strength and so on. these involve position, tangent line, normal line, time, velocity, accelerated velocity, which can be concluded to study function and its derivative. norm function space theory is a mathematic theory to describe and analyze multi-dimension space, which establish the relation between three-dimension and infinite space. the normal three-dimension space is a fully-fledged mathematic method, with direct geometry view, which helps the comprehension of the concepts in high-dimension, even infinite dimension space [3-4]. dynamic design problem, such as the description, the representation and the overall optimization of multi-dimension space in mechanical design, can be analyzed with norm function space theory. so, mechanical design can be transferred from geometry analogy to precise calculation, from three-dimension to multi-dimension. the content is expanded by the project is funded by the sichuan educational commission in the area of nature science (contract no. 2004a163). 332 jiang: the application of norm functional multi-dimensional space theory in dynamic design changing the method of design, thus complete the fullness and scope in mechanical design. 2.the arbitrary function problem of n-dimension design space 2.1 the expansion and composition of design space mechanical design has developed from static to dynamic design, from three-dimension to multidimension, from mechanism to inter-discipline of mechanism, electricity, magnetism, liquid and light, from geometry analogy to precise calculation, from sequential to parallel design. the emphasis is transferred from structure design to life process of product. the continual development of mechanical design will expand dimension and content in design space further. the structure parameter and the function or differential equations set which expresses physical process are foundation to form overall performance. the multi-dimension space of mechanical design consists of structure parameter or function sequential set. normally, the spatial variable is represented by vector: [ ]tn21 x,,x,x=x the sub-spaces are collections of parameters with identical characteristics. the design space can be divided into sub-space according to following functions: geometry shape, shape restrict, geometry factor; requirements on tolerance and surface technology; statics performance; motion performance; dynamics performance; process performance; and other performances concerned with geometry topology. all these performances are represented with functions and derivatives, such as the design of contour curve which need to consider its function and derivative, with nth continual derivatives to ensure smoothness of contour. the mechanical motion design hopes for motion equation with continual derivatives over 3 to ensure steady movement. 2.1.2 the arbitrary function problem of design space arbitrary function is more abstract and comprehensive than higher mathematics and linear algebra. to study space composed of different functions, the property and trend can be analyzed with abstract operator. normally, the problems of arbitrary function are the representation and analysis of multi-dimension function space and to choose one set optimized function which meets the requirement from functions group. the design process is to choose representation parameter according to preset dynamic characteristics from the system, to derive the solution which meets the design requirement, to choose one set optimized function, such as contour, motion, strength function and so on, in order to meet the standard. for example, a design with, in certain range, least deflection in vibrating frequency and lightest mass is a reverse problem of dynamics. that is, to preset requirements of the system on frequency, stress, strain and displacement, which determine the material and geometry characteristics in mechanical system. the problem is to solve the extremum of design arbitrary function. so, dynamic design problem is also arbitrary function problem: to represent and analyze multi-dimension space in mechanical design and to solve the extremum of arbitrary function. the first arbitrary function problem of dynamic design is to study n-dimensional variable and establish a space which can express n-dimensional variable. the second arbitrary function problem of dynamic design is to study the variable and its change rate at the same time. the key problem is to establish a norm function space which can express the function and its n th derivative at the same time. the third arbitrary function problem of dynamic design is to choose optimized solution which meets dynamic requirements and proper operator to solve and design the extremum of arbitrary function. 3.the representation and simplification of n-dimension design space with arbitrary function 3.1.1 the representation of design space advances in systems science and applications (2011), vol.11, no.3-4 333 traditional mechanical design space is three-dimensional euclid space, which is represented by distance and angle, and can be analyzed continuously and with extremes. for multi-dimension, such as linear space, normally there is no need to represent but only introduce norm into it, and make it a normed space. with norm, the concepts of open set, closed set, convergence and continuation can be introduced correspondingly. in order to solve the first arbitrary function problem of dynamic design, a normed space should be defined. definition: supposing x is the vector space in number field k (real number field r or complex number field c), if for every x∈x , appointed one real number ║x║ which called norm of x, it meets following norm axioms. (1) homogeneity: xx αα = ; (2) triangle inequality: yxyx ++ ≤ ; (3) positive definiteness: 0≥x 00 =⇔= xx as above, x, xy∈ ,if k∈α , x is called a nomed vector space in k, normed space for short. when norm must be expressed definitely, it is marked as (x,║•║). if k=r (or c),the normed space in k is real (or complex) . the way to define norm diversifies. norm can be defined according to its physical meaning. furthermore, the distance of three-dimensional space, the displacement, velocity and accelerated velocity of motion can also be defined into norm separately, which can also be combined together for norm. the same function can be bestowed different norm. for example, c[a,b] is defined all continuous derivable functions in [a,b], which forms linear space x. for every x∈x , we define norm: )( bta txmaxx ≤≤ = or )()( txmaxtxmaxx b≤t≤ab≤t≤a += or dttxx a∫= b )( they can all be normed space. n-dimensional euclid space set distance as norm: so it is also a normed space. norm is a representation of multi-dimension. the definition of norm is artificial, which makes its physical meaning diversified. some norm has such a complicated physical meaning that it is unfathomed by using three-dimension. multi-dimension can express much richer physical meaning than three. norm can express n variables at the same time, which can represent and analyze variables from several design spaces systematically and comprehensively at the same time. 3.1.2 the simplification of n-dimension design space with norm norm analysis is highly comprehensive and abstract. it is not focused on specific structure and physical meaning of function, but on the function system and its commonality. mechanical design space analyzed with arbitrary function can unify geometry, statics, kinematics and dynamics design into one norm structure. the diversification of norm definitions makes the study comprehensive and flexible. the normed space with distance as its norm is a traditional geometry space. and the one with the linear combination of displacement, velocity and accelerated velocity can show the difference among different moving modes easily and clearly. normed space with linear combination of 2 1 1 2 ⎟ ⎠ ⎞ ⎜ ⎝ ⎛ = ∑ = n i ix ξ 334 jiang: the application of norm functional multi-dimensional space theory in dynamic design contour function and n th derivative as norm can also express the smoothness of curve and the continuity of processing easily. define for any )(tu mc∈ ,set norm: p ma p p a pm uu 1 , ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎝ ⎛ ∂= ∑ ≤ { }∞<∈= pm, m u:)t(cus is a normed space. its completion is a banach space, called sobolev space. the second arbitrary function problem of dynamic design can be solved by sobolev space definition, that is, to study variable and its change rate at the same time. though two functions with equal norm║u║m,p may not be the same, the norm of identical functions must be equal. so, accompanied by other conditions, norm can be used to study function. distance, motion and energy space all involves function and its derivatives. sobolev space is concerned with motion: suppose the motion space is composed of displacement, velocity, acceleration velocity and accelerating acceleration velocity, represented with u ,u ′,u ′′ ,u ′′′ respectively, )(tu mc∈ , set norm: ( )2 2 2 2 1 1 1 1 1 1 22 2 2 2 , u dt u dt u dt u dt p pa m p p a m u u ξ ξ ξ ξ ξ ξ ξ ξ ≤ ⎛ ⎞ ′ ′′ ′′′= ∂ = + + +⎜ ⎟⎜ ⎟ ⎝ ⎠ ∑ ∫ ∫ ∫ ∫ ║u║m,p norm can express for variables u ,u ′,u ′′ ,u ′′′ at the same time. space variables in dynamics are: u ,u′,u ′′ , um ′′ , in which m is mass. space variables of contour curve are ( )tc0 , ( )tc1 , ( )tc 2 ,…, ( )tc n , which can also be defined into norm with similar method. they share the same form, all belonging to sobolev space. sobolev space, which combines parameter and its derivative to define norm, includes the main content in dynamic design by unified structure of norm function. this highly abstract and comprehensive structure makes the analysis of mechanical design space more simple and comprehensive. 4.example for dynamic design analyzed with arbitrary function 4.1.1 flexible and digital characteristics to define norm with several parameters or only one, no matter which way, the analysis of arbitrary function directly with norm is discrete and fussy. if to analyze directly with calculus, it’s also discrete and fussy for function space because calculus can only study one function. operators in arbitrary function are all effective methods to analyze arbitrary function. the optimization of arbitrary function, effective theoretical method, is an expansion of differential calculus. gateaux differential is equivalent to vatiation in variational calculus, or direction derivative in differential calculus. fréchet differential is equivalent to gradient in differential calculus or complete differential or gradient. by using gateaux and fréchet differential, the extreme value of arbitrary function in linear space can be solved easily, and so can the necessary condition for local extreme. if x is a linear space and f is real value arbitrary function in x, with gateaux or fréchet differentiable, the necessary condition for extreme in x∈0x is, for every x∈h , ( ) 0; =hxfδ . that is, the extremum of arbitrary function can be solved by vatiation, which is the commonest and simplest method in analysis and calculation. at present, the main task for dynamic design is to solve kinematics and dynamics problems, such as vibration frequency, periodic and inertial force, which involves mainly in analytical mechanics. the major theory of analytical mechanics is deduced according to the necessary condition for arbitrary function extreme, such as hamilton’s principle and lagrange equation. advances in systems science and applications (2011), vol.11, no.3-4 335 a x y z d b c m ω θ φ m m hamilton’s principle focuses on energy arbitrary function composed of general displacement and its derivative, time and so on. if the constraint is ideal and complete, and the conservative force is drive force, the real motion is the motion which makes get exrtreme value. ( ) ( ) ( ) ( )[ ] ttqv-t,qq,ttt,qqlqs t t t t d,d, 2 1 2 1 ∫∫ ′=′= if the function which makes arbitrary function s(q) get extremum is existing and exclusive, according to arbitrary function differential theory, the necessary condition equivalent to arbitrary function extremum is: 0 q l q l t aa = ∂ ∂ −⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ ′∂ ∂ d d this is called lagrange equations set [5-6]. the solution can be very easy and simple only by partial differential or differential. for example: to design a centrifugal governor shown in fig.1, a slide c with mass m can slide along a vertical axis, and the whole system can also rotate on this axis. point a is fixed, point b and d hinge two particles with mass m. the requirement: when the system acquires stable angular velocity ω, the vibration periodic is t, and determine the length of lever a. solution: (1) list laplace equation: the angle between plane abcd and axis x isφ , θ = dac∠ , degree of freedom f=2. to choose θ , φ is generalized coordinates, constraint is ideal and stable. the cartesian coordinate of particle m: φcosθ sinax = θθ cossinay = θcosaz = lagrange function: ( ) ( ) ( ) ( )2 22 2m ml θ,φ,θ ,φ 2 aθ aφ sinθ 4a θ sinθ 2mga cosθ 2mga cosθ 2 2 ⎡ ⎤ ⎡ ⎤′ ′ ′ ′ ′= + + + +⎣ ⎦ ⎣ ⎦ to solve partial derivative: fig.1 the vibration periodic design of centrifugal 336 jiang: the application of norm functional multi-dimensional space theory in dynamic design to solve partial derivative and simplification: θsinθma4θma2)φθφ,θ,(l θ 222 ′+′=′′ ′∂ ∂ lagrange equation: 2 2 2 2 2 2(2ma θ 4ma θ sin θ) (2mφ 4mθ )a cosθ (2m 2m)ga sin θ t d d ′ ′ ′ ′⎡ ⎤+ = + +⎣ ⎦ ( ) ( ) ( )2 2 2 gm 2m sin θ θ mφ 2mθ cosθm+m sinθ a ⎡ ⎤′′ ′ ′+ = −⎢ ⎥⎣ ⎦ (2) to solve relatively balanced condition and micro-vibration periodic, the system rotates on vertical axis at even angular velocity ω, ωφ =′ , on balanced position: ( ) 0 a gmmθcosmω2 =⎥⎦ ⎤ ⎢⎣ ⎡ +− to simplify, then: 2 2 gm 2mθ sin θ mω cosθ (m m) sinθ a ⎡ ⎤′+ = − +⎢ ⎥⎣ ⎦ at balanced position: a gmmθmω s 2 )(cos += solve angle sθ , then put it into lagrange equation: [ ] θsin)θcosθ(cosmωθ)θsinm2m( s 22 −=′′+ by using formula: ) 2 θθsin() 2 θθsin(2θcosθcos ss s −+ −=− define sθθφ −= , take micro-vibration approximation: s s θ 2 θθ = + sss θsinφ) 2 φsin(θsin2θcosθcos −=−=− define and set the micro-vibration approximation: [ ] ss 2 s 2 θsin)φθsin(mωφ)θsinm2m( −=′′+ φ θsinm2m θsinmωφ s 2 s 22 + − =′′ s 2 s 22 2 θsinm2m θsinmωω + = micro-vibration frequency: s 2 s 2 θsin m m21 θsinωω + = the micro-vibration periodic near the balanced position is: advances in systems science and applications (2011), vol.11, no.3-4 337 s 2 s 2 θsin θsin m m21 ω 2π ω π2t + == solve: 2 1 2 s m m 2 ωtθsin − ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎣ ⎡ −⎟ ⎠ ⎞ ⎜ ⎝ ⎛= π set: 0.6=t s, 3 m m = , sω π10= solve: by the angle sθ at balanced position: 0435.0 θcosω g m m1 a s 2 = ⎟ ⎠ ⎞ ⎜ ⎝ ⎛ + = (m) by simple dynamic design example above, we know that according to the expression of micro-vibration periodic, the structure of a centrifugal governor which meets dynamics requirements can be produced. the step is that based on the physical characteristics of system, to describe and establish arbitrary function rule by choosing motion parameters, and to deduce function solution that is in accordance with the design condition. it is different from tradition. this design is not confined to size or three-dimension any more, but dealing with the physical essence of system directly, choosing parameters flexibly, establishing arbitrary function rule and solving it. compared with traditional die forming and hard jig assemblies, mpf/mpp can perform in dieless and jigless manufacturing. the flexibility characteristic can be completely shown by using active elements in the eg of the mpf/mpp tooling. for example, using mppf tooling, the deformation path can be changed to gain the best loading for raising the formability limit through increases in the contact area between the punches and the sheet, which changes the loading condition. various outputs from a 3d surface model can be digitally used to adjust the eg using a computer control system to configure the surfaces of tools. by means of the flexible and digital characteristics of mpf/mpp, larger and more complex panel parts can be formed and assembled incrementally and continuously with the dt/jt system. 5.conclusion to describe multi-dimension in mechanical design with arbitrary function theory, the selection of parameter is not geometry size, the object is not geometrical body, the representation is not limited to three-dimension figure any more, but to build model just according to the system’s physical essence, and with diverse expression, to express different kinds of parameter at the same time. the design is not limited to the arrangement of three-dimensional figure, but to analyze and compare parameters form different design field with space theory. by using normed function, the convergence and extreme problems of function in banach space can be discussed and designed. by utilizing hilbert space, the optimization problem can be discussed. and by using differential operator, dynamic space function, derived function and optimized solution of parameter can be solved. to sum up, arbitrary function theory can solve the difficulty in dynamic design, and facilitate dynamic design. acknowledgements =sθsin 0.4082, 094924.θs = a gmmθmω s 2 )()cos( += 338 jiang: the application of norm functional multi-dimensional space theory in dynamic design the project is funded by the sichuan educational commission in the area of nature science (contract no. 2004a163). references [1] a. illarramendi1, j. l. azpeitia1, r. bueno1,new control techniques based on state space observers for improving the precision and dynamic behaviour of machine tools, annals of cirp,54/(1) (2005): 632635. [2] y.altintas.,k. eykorkmaz. feedrate optimization for spline interpolation in high speed machine tools[j]. cirp, 52(1)(2003):288-305. [3] francois isnard , gordon dodds. dynamic positioning ofclosed chain robot mechanisms in virtual reality environments. international conference on intelligent robots and systems , victoria , b. c. , canada ,1998. [4] luokang,xiongnan the structure identification of link gear based on connected matrix aggregtion operation, sichuan industry technology college scientific journal, 19 (2) (2000) 35∶ ~37. [5] wang shuang. functional analysis and optimization theory [m]. beijing: beijing aerospace university press, 2004. [6] isnard f , dodds g ,claude vallée. efficient multi arm closed chain dynamics computation for visualisation . international conference on intelligent robots and sys2 tems , grenoble , france ,1997. мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 81-88 the asymmetry of the present-day social and demographic conditions of the north vitali s.belozerov, natalia a. shchitova ,vasiliyv.chichikhin north-caucasus federal university, stavropol, russia annotation the integral analysis of the current socio-demographic conditions in the north caucasus, the specific features of natural population development, migration and their influence on the dynamics of the ethnic structure are presented in the article. substantial regional disparities in the demographic development of the territories have been found. the regional differences in the reproduction and the ethnic composition of the population of the mountain and plain land regions have been shown. the four types of regions with their specific way of population generation have been identified. special attention is paid to the processes connected with the transformation of the ethnic structure of the population. the increasing number of the areas rapidly changing their ethnic composition, the growth of multi-ethnicity in the plain land regions, derussification and ethnic homogenization in the mountain areas have been noted. a significant asymmetry of the north caucasus regions in terms of their social and economic development proving the territorial disparities between the republics in the south-east and the relatively developed plain regions in the north-west has been revealed. key words: the north caucasus, demographic conditions, migration of the population, ethnic structure, interregional disparities, social and economic processes.. 1 introduction the north caucasus is one of the most peculiar regions in russia in terms of social and demographic issues. here, quite developed, self-sustaining provincial and falling behind areas adjoin each other on a relatively small territory. some of them are rather politically stable and characterized by a rapid economic growth, while other regions are inert and inactive and still others are seen to have growing social inequality and deteriorating economic conditions. among the factors, influencing the social and economic development of the north caucasus region in the first place is its geopolitical and cross-border position to have gained pace after the collapse of the soviet union. in fact, it is the southern outpost of russia, providing access to the south caucasus, the black sea and the caspian sea. it is located between two highly differing in terms of socio-cultural, and ethno-demographic processes territories slavic, primarily orthodox (central russia, belarus and ukraine) and the caucasus, being intervened in ethnic and religious terms (north caucasian republics and south caucasus states). the north caucasus is one of the most important and at the same time complex and unstable ethnic and political elements in the structure of the russian nationhood. the main causes of the political precariousness arise from the ethnic heterogeneity of the geo-space, high multiethnic, overlapping and intersection of ethnic settlement areas, the contradictions in the political and economic purposes of different ethnic, cultural and social groups [5]. hence, the conditions there are quite loaded and controversial and resolved not only through negotiations but sometimes 82 vitali s.belozerov, natalia a. shchitova ,vasiliyv.chichikhin: the asymmetry of the present-day social and … with force applied. often ill-conceived administrative and territorial structure itself proved to be a detonator of the conflict. being observed in the last decades growth in the ethnic identification takes place along with the boost in religious consciousness. for many centuries, the north caucasus is the contact zone of the peoples belonging to two major world religions christianity and islam, the representatives of other religions being present as well. however, one cannot ignore the fact of non-traditional religious movements in the area and their destructive influence, in particular, the islamic fundamentalist movement (e.g. wahhabism). in general, it is no secret that the socio-economic and political situation in the north caucasus is deteriorated by a number of factors preventing strengthening of the unity of the russian nation and harm onisation of interethnic relations. it is impossible to understand the nature of these problems without a thorough analysis of the social and demographic processes, which largely influence not only the quality of human potential, but also the quality of life of the population. 2 sources of information and research methods in our study the main accent was made on finding answers to the following questions: 1. what is the demographic and migration situation in the north caucasus as a whole and in its certain regions? 2. how inhomogeneous is the social and demographic extent of the north caucasus, what are the reasons for this heterogeneity? 3. how is the ethnic structure of the population of the north caucasus transformed under the influence of the demographic and migration processes? 4. how do ethnic and demographic processes affect the population’s standard of living and socio-economic situation in the north caucasus? we used the following sources of information in our study: population current record data (mainly demographic and migration records) available on the official website of the federal state statistics service rosstat and its regional offices in the north caucasus; data collected during the census of 1989 (the all-union 1989 census), 2002 and 2010 (all-russia national population census of 2002, national population census of 2010 ) on the ethnic structure of the north caucasus’ regions’ inhabitants; design and analysis data of the rating agencies ria-rating and expert ra on the level of social and economic development of the north caucasus regions; results of the previous complex and industry (demographic, migration, ethnic, social and economic) researches of the north caucasus conducted, including the authors of the article [1][8]. the main research methods used are quantitative ones. they are conventionally applied for this kind of work. the focus is made on the multilateral statistical analysis of demographic, migration, ethnic and socio-economic processes, ratios and indicators of birth rate, death rate, natural population increase, net migration, the share of the various ethnic groups in the general population being calculated. the main qualitative methods used in this research are aimed at identifying the territorial disparities of the social and demographic processes in the north caucasus combined with the analysis of the temporal trends over the last 15 20 years. those methods are widely applied in social and economic geography. the method of multi-immensity is based on the classic geographical papers and it is understood as the manipulation of scale ranks of the territory within the particular object of the research to find the spatial patterns of events and processes. multi-immensity of this study includes the following territorial “steps”: federal (comparison of advances in systems science and application(2016) vol.16 no.4 83 the north caucasus with the processes of the russian federation as a whole), regional (the subjects of the russian federation in the north caucasus), intra-regional (municipal districts). the typological method deals with the development of such groupings of geographical features used to identify the qualitative differences. we have proposed a typology of the north caucasus regions according to the nature of the population. this method allowed us to conduct the zoning of the north caucasus territory in terms of socio-economic development and living standards. in general, finding the territorial disparities of the processes was carried out through the intensity vectors in the directions of north south and west east. the main spatial heterogeneity is the asymmetry of the mountain land (south or more precisely south-east) and plainl and (north-west) parts of the north caucasus. the mountain part of the republic includes dagestan, chechnya, ingushetia, north ossetia, kabardino-balkaria, karachay-cherkessia and the plain part includes rostov, krasnodar and stavropol regions, and the republic of adygea. 3 main results and their discussion specifics features of the processes in the mountainous and plain land parts of the north caucasus. demographic processes. in the 60 – 80s of the 20 th century, the plain land and mountainous parts of the north caucasus “changed” their features of the demographic conditions. in the plain land part, the high level of natural growth inherent in it fell to the population replacement level since the second half of the 19 th century and currently to the restricted population replacement level. in the mountainous area, on the contrary, the restricted population replacement level of 1960s turned into the expanded one and even population explosion” was observed there. at the present time, the north caucasus differs greatly for the better in terms of birth and death rate from the rest of russia and the region can be quite clearly divided into “relatively successful south” (mountainous area) and “disadvantaged north” (the plain land).the range of the crude birth rate (cbr) is from 25 29 to 10 11 ‰. in all the republics, the cbr is quite high. for example, in chechnya it is 25 29 ‰, in ingushetia 21 27 ‰, in dagestan about 19 ‰.when moving to the north-west, the birth rate falls and reaches its minimum in rostov region (10 11 ‰.). at the same time, since the 2000s, the slow growth of birth rate is quite obvious in all regions of the north caucasian plains (from 9 10 and 11 12 ‰) being a consequence (as will be proved below) of change in the ethnic and demographic structure and increase in the share of non-slavic, more demographically active population. however, it should be noted that in most regions of the north caucasus, as well as in the whole country, in the 1990s early 2000s the birth rate was significantly lower than it is today. the particular negative peak at the border of the centuries was due to the widespread protracted demographic crisis caused by the serious political and socio-economic disturbances. the second most important demographic process is population mortality. in rostov region, its indicators are the highest and even slightly higher than the national average ones, they are slightly lower in krasnodar and stavropol regions. in the republics, the crude death rate (cdr) is not high. as is known, the value of cdr is highly dependent on the age structure of the population and life expectancy. in mountainous republics, the age structure of the population is much younger than in that in the plain land regions. for example, in stavropol region, the share of young people is 18%, older people is 22%, and in chechnya, these figures are 35 and 9% respectively, in ingushetia they are 31% and 10%, in dagestan they are 27% and 11% respectively. at the same time, the republics show the highest average life expectancy: in ingushetia it is almost 81 years old, in north ossetia 79 years, in dagestan, kabardino-balkaria and karachay-cherkessia it is 78 years (in plain land regions it is lower than 70 years). as a result, the north caucasus republics show the highest natural population growth in russia: in chechnya it is 19.9 ‰, in ingushetia – 17.7 ‰, in dagestan – 13.3 ‰. in the plain land regions, the natural population growth is slightly positive or even negative. 84 vitali s.belozerov, natalia a. shchitova ,vasiliyv.chichikhin: the asymmetry of the present-day social and … thus, the modern period in the north caucasus is characterrised by the increase in the interregional disparities in demographic development the south shows an impressive multi-ethnic human potential increase, and the north the mono-ethnic slavic community of nations is aging, weakened and eroded. migration processes. it is obvious that the demographic processes govern the vector of the interregional migration. from the 1960s the population of the mountain are as accumulating a powerful demographic potential was actively drawing upon the labour deficient plain territories. plain land areas are, on the one hand, a powerful accumulator of the north caucasian migration flows, and on the other hand, they are a kind of “a corridor of the migration winds” between the north caucasian republics and the rest of the country. being at the forefront of the caucasus problems, krasnodar, rostov and especially stavropol region appeared to be the russian base in the caucasus and at the same time the buffer against the acuteness of the ethno-political and socio-economic crisis. the location next to the hot spots of chechnya, dagestan, and the zone of the osset-ingush conflict, etc. and what is more as acute hearths of inter-ethnic tensions in the countries of the south caucasus (georgia-south ossetia, azerbaijani-armenian, georgian-abkhazian) resulted in a mass influx of the forced migrants to stavropol, don and kuban regions. gradually the stressful migration factors have lost their significance. the current quiet geopolitical conditions have exacerbated the impact of economic and demographic factors on the migration process. the inter-regional migration growth in the plain land regions will remain positive due to the high demographic potential in the republics. the north caucasian peoples have high demographic potential growth rates and are actively advancing on the plain land, adjacent to the caucasus territories. the population and young people in the first place, will be leaving to choose new places of residence in the economically developed regions. at the same time, the ukrainian instability hearth is heading toward the renewed stress migration flow in the regions of the plain land part of the north caucasus. territorial heterogeneity of demographic and migration processes in the north caucasus can be divided into four types of the regions with their specific features of population formation: – regions, where natural population growth exceeds the total one (chechnya) and has a steady migration outflow; – regions, where the excess of the natural and total growth alternate (dagestan, ingushetia, north ossetia, kabardino-balkaria, karachay-cherkessia) and at certain periods they either accept or give a large number of migrants; – regions, where natural growth is negative or very low, and the total one is almost always positive or higher than the natural growth (stavropol and krasnodar regions), the population growth is achieved due to the migrants. here we can also observe the growth of the intraregional disparities in demographic development in the form of the formation of areas of high positive and high negative population growth being the consequence of the transformation of the population ethnic structure. – regions, where the total population growth is negative or slightly positive (rostov region, adygea). migration there does not overlap the natural population decline. ethnic processes. the ethnic structure of the north caucasus population is largely dependent upon the correlation of the three groups of ethnic groups to be the russians, the title nations of the national entities (the adygei, karachai, cherkesses, kabardin, balkar, ossetian, ingush, chechen and the peoples of dagestan), and a number of non-indigenous peoples (the greeks, armenians, azerbaijani, germans, meskhetian turks, etc.). the dynamics of the ethnic structure is dependent upon the balance of the natural and mechanical population growth. from the beginning of colonisation to the present time, two trends are observed: at an early stage, it is “slavyanization” and at present, it is strengthening of the caucasian characteristics. advances in systems science and application(2016) vol.16 no.4 85 for nearly two centuries, the two main ethnic groups – the slavic people (mostly russians) and the title nations of the republics had been changing roles in the course of their interaction. initially, the main ethnic group to ensure the inclusion of the mountain peoples in the russian economy and social sphere were russians. currently, the north caucasian peoples having high demographic growth rates are actively “attacking” the plain land part of the north caucasus, which is seen not only in expanding the settlement territory but also in active participation in the economic activity [6]. at the same time, the expansion of the title ethnic groups of the republics originally was seen in the rural districts and the agricultural sector, and now the expansion takes place in the cities and in the new sectors of the economy (including business, management, etc.) and more young people are found to be engaged in the educational sphere. simultaneously, the title ethnic groups are concentrated in “their” original regions, whereas the “aliens” are pushed to leave. the main factor of ethnic homogenisation of there publics’ population is not an ethnic incompatibility and hatred, but growing competition for becoming scarce job and study placements, comfortable living conditions, water and land. these processes often turn into covertly or overtly confrontational and they are accompanied by the change in the settlement of the ethnic groups [5]. these changes are particularly well seen in the geography of the displacement of the russian ethnos, having gone through several stages in its evolution in the republics of the region. at the initial stage, they chose the administrative, metropolitan and industrial centers as their place of residence, as a rule. at the next stage, there occurs a slowdown in the growth of the russian population and the outflow of the russians from the rural areas. later the inclusion of the title ethnic groups of the republics in the urbanisation process in the conditions of expanded reproduction was accompanied by their active resettlement in the rural areas by means of “pushing out” and displacement of the russian population living there. the third stage is characterised by the sustained reduction in the absolute and relative indicators of the russians in both rural and urban areas. the outflow of the russians was seen throughout all the settlements and especially the cities, including the capital ones in the conditions of the deep economic crisis and the lack of effective national policy, high ethnic tensions, actively propagating xenophobia and actual civil insecurity of the russian and russianspeaking population. the structure of the non-indigenous population of the north caucasus is changing rapidly. the active outflow of some nations (particularly the germans emigrating to germany) combined with the increasing inflow of other nations (the armenians) and the emergence of the new nations (the azerbaijanis, meskhetian turks). some of them, first of all, the armenians and the greeks, and in recent years meskhetian turks have formed the compact settlement areas. especially noticeable is the increase in the proportion of the armenian ethnos. if in the national republics the armenian population was declining (except adygea republic), in the steppe areas of the north caucasus the influx of the armenians has increased dramatically, primarily in stavropol territory and kuban [3][4]. new in their resettlement was the settlement, along with the cities and suburbs, in the rural areas. the indicator of the concentration of the armenian population in areas traditionally inhabited regions of the steppe ciscaucasia dropped significantly. a new feature about their resettlement was the settlement in the cities, suburbs, and rural areas. the indicator of the armenian population concentration in the traditionally inhabited areas has significantly reduced. contrasting zones of socio-demographic area of the north caucasus. the course of the modern demographic, migration and ethnic processes in the north caucasus depends on both political and geographical changes and socio-economic conditions in the caucasus as a whole in the post-soviet years. the active migration of migrants from the crisis regions resulted in a more stable transformation of the ethnic structure of both receiving and giving territories has changed the trends of their social and economic development. the north caucasus regions occupying the peripheral position in the russian sociogeographical space, form a rather complex discrete conglomerate, its parts being contrasted with 86 vitali s.belozerov, natalia a. shchitova ,vasiliyv.chichikhin: the asymmetry of the present-day social and … each other on the main socio-economic characteristics [10], the inter-regional disparities in terms of welfare, development of social and cultural infrastructure and other indicators have reached here enormous values. the comparative analysis of the socio-economic situation in the north caucasus regions show a close correlation between the parameters of ethno-demographic and socio-economic processes [16]. the discrepancy between the level of the socio-economic well-being and demographic prosperity of the population is becoming a source of the key contradictions adding to the instability of the regional development. the crisis of the early 1990s led to the considerable deterioration of the population’s living conditions in all areas of the north caucasus, but the economic boom of the 2000s manifested itself in many ways and led to a significant socioeconomic differentiation of the regions and to the formation of depression areas and areas of socio-economic growth. currently, the north caucasus can be divided into three contrasting zones, differing in the nature and pace of socio-economic development and the living standards of the population [15]. the western zone is represented by krasnodar territory and rostov region being stable leaders, developing in the most dynamic way in economic and social sphere. however, the falling migration growth there does not compensate for the natural population decline. the central zone includes one region – stavropol territory, where a living standard is going up due to the economic growth, unemployment reduction, income increase and decrease in the proportion of the poor people. demographic and migration conditions are relatively favorable there. the western and central zones are characterised by multi-ethnicity strengthened through active, sometimes a point settling of the title peoples of the north caucasus republics of the south caucasus states, “pushing out” the russians from the old russian regions. many traditionally “rural” ethnic groups (dargin, chechens) having settled in the eastern agricultural areas are gradually shifting to the west and to the cities [2]. the influx of dargin, chechen, karachay, cherkess in large cities and towns is growing due to educational migration, which is becoming a real channel of social mobility and transformation of the ethnic villagers into townspeople [1][14]. all this increases the inter-ethnic tensions, forms ahostile attitude of the local population to the ethnic migrants and raises domestic aggression. the southern zone–the north caucasus republics can be divided into two subtypes. the first one the least affected in the 1990s early 2000s karachay-cherkessia, north ossetia and kabardino-balkaria, which currently are having deep economic and social problems, but having several high indicators in the social sphere (e.g. housing, the number of students, health indicators). the second subtype of the southern zone includes chechnya, ingushetia characterised by the low levels of socio-economic development and a large number of unresolved issues negatively affecting the quality of life in general. these republics are the outsiders as regards the difficulties in the labour market, observed for a long time, despite the efforts of the federal and regional authorities. the highest unemployment rate has been recorded in ingushetia: almost half of its inhabitants do not have a permanent place of employment. in two regions chechnya and dagestan the unemployment rate is over 10% [16]. it is believed that these figures are not reliable as the population of these republics have shadow employment [9], existing at least, in two forms. the first form of employment is non-incorporated entrepreneurs engaged in farm households and producing products for sale. the second form is work for private individuals [13]. in this subtype, the russian population in chechnya is declining fast: from 1989 till 2010 it has fallen 11 times, in ingushetia the figure is almost 8 times, in dagestan it is 1.6 times. there has been a rapid reduction in long-time territorial expansion of the russians in the caucasus. the most dramatic was the fate of nearly 300-thousand russian population of the chechen republic from which the forced exodus has exceeded 90%. the most striking change in the ethnic structure of the population is observed in the “old russian” regions of chechnya, ingushetia and advances in systems science and application(2016) vol.16 no.4 87 dagestan. over the last 50 years, not only the dominance the dominance of the russian population but also the prevalence has been lost and in the 1990s, these trends increased significantly and in a number of “old russian” areas almost complete exodus of the russian population is observed [6]. 4 conclusions 1. the north caucasus having a high population size is gradually losing its demographic advantages. the territories with both natural population decline and a negative migration balance are expanding. in the plain land regions, the falling migration gain does not compensate for the natural population decline. 2. multi-ethnicity and derussification of the migration growth has considerably increased. the area of the positive migratory growth of the russians and armenians is falling rapidly. the migration of the rural ethnic groups into the cities is growing and the ethnic structure of those groups is being transformed. 3. the regions of the north caucasus, being at the crossroads of inter-regional and crosscountry migration flows fulfil the function of integration of different nations. here, numerous ethnic cultures of numerous language groups interact (slavic, armenian, iranian, greek, german, nakh-dagestani, abkhazian-circassian, turkic, etc.). at the same time, russian population in all regions of the north caucasus is rapidly going down. 4. strengthening of multi-ethnicity in the plain land regions of the north caucasus owing to the active settlement of the title peoples of the north caucasus republics of the south caucasus states, “pushing out” of the russians from the old russian regions, enhances inter-ethnic tensions, forms the hostile attitude of the local population to the ethnic migrants and growth of domestic aggression. 5. the russians outflow from the republics prevents from the rapid revival of the industrial economy sectors. areas of the russians’ mass exodus are a risk zone for the integrity of the russian state and indicators of ethnic tensions proving the propagation of nationalist sentiments and actions, extremism and other extreme forms of ethnic tensions manifestations. the most negative result of this situation is the russian center influence weakening and the spread of political, economic and religious expansion of the muslim countries. 6. the impact of migration and ethno-demographic processes on the development and socioeconomic stability in the north caucasus regions is evident. the discrepancy between demographic and socio-economic well-being becomes a source of political instability in the republics. however, per capita income is falling behind the population growth, despite some economic recovery. high birth and life expectancy rates raise dependency burden on the economically active population, whose structure is now dominated by the unemployed. the location of regions with high birth rate and younger age structure of the population in one place requires the adequate response from the federal authorities. the socio-economic policy is to be set taking into account the regional conditions and specific measures aimed at eliminating contrasts in socio-economic development of the population are to be proposed taking into account the ethnodemographic characteristics. acknowledgment this researchers was supported by the russian foundation for basic research № 16-06-00179 "development and approbation of system of geo information monitoring of ethno-demographic processes (on the example of regions of the north caucasus)". references 88 vitali s.belozerov, natalia a. shchitova ,vasiliyv.chichikhin: the asymmetry of the present-day social and … [1] belozerov v.s., panin a.n., prikhodko r.a. and chihichin v.v. (2014), "migration processes in stavropol territory: trends and current situation". science. innovation. technologies. no.4. p. 96 -108. [2] belozerov v.s., panin a.n., prikhodko r.a.,chihichin v.v. and cherkasov a.a(2014). "ethnic atlas of the stavropol territory". stavropol. p.304. [3] belozerov v.s. and panin a.n.cherkasov a.a(2014). "geo-information monitoring of settlement and migration of the armenians. migration processes: problems of migrants adaptation and integrations". proceedings of the international scientific conference. stavropol. p. 266 273. [4] belozerov v.s., panina.n. and chihichinv.v (2008). "ethnic atlas of stavropol territory". stavropol, p.208. [5] belozerov v.s. and polyanp.m (2010). "demographic processes in the north caucasus. permyakov’s collection of papers". p. 478 493. [6] belozerov v.s(2005). "ethnic map of the north caucasus". p .204. [7] belozerov v.s(2000). "demographic processes in the north caucasus". stavropol. p.166. [8] zolnikovayu.f. and belozerov v.s(2014.). "settling and migration of the armenians in the northern caucasus. the role of migration in the socio-economic and demographic development of the sending and receiving countries of europe and asia: the regions of eastern europe and central asia". proceedings of the conference. pp. 347-350. [9] zubarevich n.v(2010).russia’s regions: inequality,crisis, modernisation. moscow, p.160. [10] rating of socio-economic status of the subjects of the russian federation. 2014 year results[electronic resource] http://vid1.rian.ru/ig/ratings/rating_regions_2015.pdf. [11] ryazantsev s.v(2003). "modern demographic and migration portrait of the north caucasus". stavropol, p.376. [12] i.v.starodubrovskaya, n.v.zubarevich, d.v.sokolov, t.p. intigrinova, n.i.mironov, h.g.magomedov(2011).the north caucasus: a modernized challenge. p.328. [13] solovyov i.a (2009). "the regional features of the modern migration in the south of russia". regional studies. no. 4-5, pp. 54-60. [14] chihichin v.v(2014). "academic migration in the north caucasus: causes, geography and potential problems". science. innovation. technologies. no. 2, pp.161-179. [15] shchitova n.a.,polushkovsky b.v. and a.i.belousov(2011). "spatial analysis of the quality of life in the south of the european part of russia". bulletin of stavropol state university. vol.76, no. 5, pp.236-241. [16] shchitova n.a.chichikhin v.v.( 2014). "comparative analysis of the socio-economic development of the north caucasus regions". science. innovation. technologies. no.1, pp.161 174. corresponding author vitali s.belozerov can be contacted at: yal05@mail.ru. advances in systems science and applications (2013) vol.13 no.3 233-248 ways of fusing different types of information and how systemic yoyo model is applied in complex systems evaluation and estimation xiaojun duan1 and yi lin2 1department of mathematics and systems science national university of defense technology changsha 410073, pr china 2department of mathematics slippery rock university slippery rock, pa 16057 usa abstract continuing the works in literature[1], we show in this paper how different types of information can be fused together consistently in order to produce accurate evaluations and estimations for complex systems. the theoretical part of this presentation is based on the standard statistical reasoning, while the ending part constructs three case studies in order to validate the main thinking logic and results obtained in literature[1] and in this paper. it is shown that (1) for linear systems, when fusing data of different types, the weights placed on the data have profound effects on the outcomes and the achieved precisions, meaning that in this case, the unique optimal weight matrix is determined by the precisions of the data (gauss-markov theorem of linear models); (2) for nonlinear models, when fusing heterogeneous sets of data with varied scales of precision, the structure of the weight matrix is no longer uniquely determined by the precisionsbut also related to the degree of model nonlinearity, indicating that the classical gauss-markov theorem of linear models no longer holds true. at the same time, a specific method of determining the optimal weighting factor and the relevant computational method for estimating the parameters are established. combined with the process of conserved information applied in systems evaluation, we provide three case studies, including (1) how to quantitatively measure prior knowledge and observational data so that prior knowledge can be considered in obtaining much improved optimal systems evaluations; (2) how to excavate new sources of observational data of processes so that the established models can be validated jointly using process data collected under different test environments and the directly measured information of the specific indices of concern in order to improve the quality of systems evaluation and estimation and to obtain model validation results of better accuracy; and (3) how to more effectively fuse prior knowledge and heterogeneous sets of data. all of these case studies further witness the epistemological validity of the information conservation existing in the systemic recognition process beneath the systems model description, prior knowledge, and observational data, and their transformational relationship, as obtained in literature[1]. xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 234 because other than establishing the theory, particular procedures are provided, conclusions of this work can be directly employed in system evaluations and estimations and related works. this work shows how systemic thinking can be practically applied to benefit the efforts of system evaluation and model estimations involved in various engineering projects. keywords systemic yoyo model, data fusion, linear/nonlinear model, parameter estimation, gauss-markov theorem 1 introduction in literature[1], after clarifying the relationship between observational quantities and the target indices to be measured, the systemic yoyo model[2] is established for system evaluations and estimations (fig.1), where the form of the model, observational data, and prior knowledge are the main sources of information useful for the evaluation, estimation, and prediction of the performance index of the system to be measured. fig.1 the systemic yoyo model for complex system evaluation from analyzing the characteristics of and connections between the three main sources of information, model descriptions, prior knowledge, and observational data, for system evaluations and estimations, the following conservation law of information for system analysis is obtained. aeim ×beid × ceip = a (1) where a, b, and c are constants, im stands for the information content described by the model, id the information content of the autoptic test data, and ip the information content of the prior knowledge. the constant a should somehow depict the minimum amount of information required to satisfy the given precision (in the estimate of the model or parameters), where the precision is 235 advances in systems science and applications (2013) vol.13 no.3 given in terms of the model accuracy and parameter estimation precision. for the detailed expressions of im , id, and ip , see [1]. with this law of conservation is established, duan and lin use it to investigate the evolution direction of the process of a system evaluation and estimation. continuing what is obtained in literature[1], in this paper, we show that for linear systems, when different types of data are available for systems evaluation and estimation, then the analysis outcomes and precisions achieved are greatly determined by the weights placed on the data. more specifically, the gauss-markov theorem, established on the method of least squares method, holds true. however, when nonlinear systems are involved, the classical gauss-markov theorem of linear models no longer holds. that is, the structure of the weight matrix is not uniquely determined by the precisions of the available data. to this end, we provide a specific method of determining the optimal weighting factor and the relevant computational method for estimating the parameters. after this theoretical exploration, combined with the conservation law of information of system evaluations and estimations, we construct three case studies to show (1)how to quantitatively measure prior knowledge and observational data so that better evaluation and estimation results can be obtained using prior knowledge; (2)how to excavate the available observational data of processes so that the process information collected under different test environments and directly measured data of the specific indices can be employed jointly to fine-tune the model, leading to improved system evaluations, estimations, and more accurate test results, and (3)how to make prior knowledge and heterogeneous sets of data work effectively together. these case studies further verify the validity of the conservation of information of the model information, prior knowledge, and observational data in the recognition process of systems and the evolutionary relationship between three main sources of information. this paper is organized as follows: section 2 looks at various ways one can fuse information in his analysis of complex systems. section 3 focuses on three specific cases studies. and, the paper is concluded by section 4. 2 ways information fusion takes place in processes of system evaluation with the requirement of precision given, we can optimize the process of a system evaluation. that is such a problem as how to obtain the optimal evaluation results when multiple types of models and multiple kinds of data are available. this end can be analyzed by placing various weights on the multiple kinds of data and prior information. xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 236 when dealing with information by combining heterogeneous data, the most typical case is the fusion of such information that are of different types and various precisions. when observational information is expressed by using a parametric model, the problem of how to fuse the information together can be transformed into that of estimating the parameters of some regression models. here, by different types of information, we mean such information that is of different functional relationships with the parameters to be estimated so that their various orders of derivatives are also different. if in our treatment we have to deal with different types of information of varying precision, different weightings placed on these information will have direct effect on the estimations of the parameters. sometimes, the effects can also be quite significant. hence, how to place weights on the information that are of heterogeneous types and varied precisions becomes a key technique for obtaining high accuracies in the parameter estimations. as for the parameter estimations of linear regression models, gauss-markov theorem provides the most optimal weighting method for observational data of different precisions [3]. as for nonlinear regression models, current publications have assumed that all observational data have the same scale of precision. that is, the random errors in these data are identically independently distributed [3]. to this end, our work in this section theoretically shows that when jointly dealing with heterogeneous sets of data of varied scales of precision, the structure of the weight matrix is no longer uniquely determined by the precision of the observational data; instead, it also has something to do with the degree of nonlinearity of the model, which can be measured by different orders of derivative functions. that is, the classical gauss-markov theorem of linear models does not hold true anymore. in the following, we will specifically address the problem of how to determine the most optimal weighting factor, while providing the relevant computational method for estimating the parameters. 2.1 the optimally weighted information fusion of linear system evaluation according to the research on the mean square errors of parameter estimations of linear systems, it is readily to show that the estimate corresponding to the optimal weighting factor is the bayesian estimation, which is gauss-markov theorem[1,6]. for the problem of estimating parameters βp×1, assume that there are the following two types of observational information:{ ym×1 = xm×pβp×1 + εm×1 eε = 0;cov(ε, ε) = σ2 1im×m (2) and { β̃k×1 = zk×pβp×1 + ηk×1 η ∼ n(0, σ2 2ik×k) (3) 237 advances in systems science and applications (2013) vol.13 no.3 satisfying eεηt = 0. if we treat model (2) as the direct observational data, while (3) the prior information on the parameters, then the bayesian estimation of the parameters is the solution of the following extremum problem: min β∈rp σ−2 1 ∥ y −xβ ∥22 +σ−2 2 ∥ β̃ − zβ ∥22 (4) that is given as follows: β̂b = (σ−2 1 xtx + σ−2 2 ztz)−1(σ−2 1 xty + σ−2 2 zt β̃) (5) therefore, the following conclusions can be shown readily[3-5]: (1)eβ̂b = β; that is the bayesian estimate is unbiased; and (2)mse(β̂b) = tr(σ−2 1 xtx + σ−2 2 ztz)−1 < mse(β̂ls) = σ2 1tr(x tx)−1 that is, by making use of appropriate prior knowledge, one can always improve the estimation precision of the parameters, where tr(ak×k) = ∑k i=1 ai,i is the trace of the matrix. when fusing these data, the weights placed on the data have profound effects on the outcomes and the achieved precisions. these conclusions indicate that when fusing observations of different precisions, the unique optimal weight matrix is determined by the precisions of the data, which in essence is still the gaussmarkov theorem of linear models established on the least squares method. 2.2 the optimally weighted information fusion of nonlinear system evaluation for nonlinear models, we show in theory that when fusing heterogeneous sets of data with varied scales of precision, the structure of the weight matrix is no longer uniquely determined by the precisions. that is, the classical gaussmarkov theorem of linear models no longer holds true. at the same time, we establish a method for determining the optimal weighting factor and the relevant computational method for estimating the parameters. for the sake of convenience, for the parameter θ that is to be estimated, we assume that we have the prior information (3) for a linear model, where the error satisfies eεηt = 0. so, the weighting problem of fusing these two types of observational data can be reduced to the following minimization problem: min ρ∈r1,ρ>0 min β∈rp ∥ y − f(x,β) ∥22 +ρ ∥ β̃ − zβ ∥22 (6) then, we have the following result: theorem1. denote s(β) =∥ y −f(β) ∥22 +ρ ∥ β̃−zβ ∥22, s(β̂) = minβ s(β), c =∑m i=1 ḟ 2 i +ρ ∑k i=1 z 2 i , d = ∑m i=1 ḟif̈i, ξ = ∑m i=1 ḟiεi+ρ ∑k i=1 ziηi, a = σ2 1 ∑m i=1 ḟ 2 i + xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 238 ρ2σ2 2 ∑k i=1 z 2 i , then under the assumed conditions (i) and (ii), the following estimation holds true: β̂ − β =c−1ξ + c−2 m∑ i=1 f̈iεiξ − 3 2 c−3dξ2 + c−3 m∑ i,j=1 f̈if̈jεiεjξ − 9 2 c−4d m∑ i=1 f̈iεiξ 2 + 9 2 c−5d2ξ3 (7) with the bias and mean square error approximated as follows: e(β̂ − β) = −1 2σ 2 1c −2d− 3 2c −3dρ(ρσ2 2 − σ2 1) ∑k i=1 z 2 i mse(β̂) =c−2a+ 6c−4d2σ4 1 + 3c−4aσ2 1 m∑ i=1 f̈2 i + 135 4 c−6d2a2 − 36c−5ad2σ2 1 (8) where the assumed conditions (i) and (ii) are given below: (i) the derivative of f(t, β) with respect to the parameter β exists and is continuous, and lim m→+∞ 1 m m∑ i=1 = ( df(ti, β) dβ )2 = ω1(β) > 0 (9) (ii) the second order derivative of f(t, β) with respect to β exists and is continuous, and lim m→+∞ 1 m m∑ i=1 = ( d2f(ti, β) dβ2 )2 = ω2(β) (10) proof. taking the series expansion of ṡ(β̂) about the true value of the parameter β produces ṡ(β̂) = ṡ(β) + s̈(β)(β̂ − β) + 2−1... s (β)(β̂ − β)2 notice that ṡ(β̂) = 0, ṡ(β) = −2( ∑m i=1 ḟiεi + ρ ∑m i=1 ziηi), s̈(β) = 2c − 2 ∑m i=1 f̈iεi, ... s (β) = 6d − 2 ∑m i=1 ... f iεi, so, ignoring all terms of third or higher order derivatives produces β̂ − β = c−1 { ξ + m∑ i=1 f̈iεi(β̂ − β)− 3 2 d(β̂ − β)2 + 1 2 m∑ i=1 ... f iεi(β̂ − β)2 } = c−1ξ + c−2 m∑ i=1 f̈iεiξ − 3 2 c−3dξ2 + c−3 m∑ i,j=1 f̈if̈jεiεjξ − 9 2 c−4d m∑ i=1 f̈iεiξ 2 + 9 2 c−5d2ξ3 239 advances in systems science and applications (2013) vol.13 no.3 that is equ.(7) holds true. because the expected values of normal variables to odd powers are zero and ε and η are independent, we obtain the first equation in equ. (8) by ignoring the error terms of the fourth and higher orders and then calculating the expected value such that e(β̂−β)2 = c−2eξ2+3c−4e( m∑ i=1 f̈iεi) 2ξ2+ 45 4 c−6d2eξ4−12c−5d m∑ i=1 f̈iεiξ 3 (11) because eξ2 = σ2 1 m∑ i=1 ḟ2i + ρ2σ2 2 k∑ i=1 z2i =̂a, eξ4 = 3σ4 1( m∑ i=1 ḟ2i ) 2 + 6ρ2σ2 1σ 2 2 m∑ i=1 ḟ2i k∑ i=1 z2i + 3ρ4σ4 2( m∑ i=1 z2i ) 2 = 3a2, e( m∑ i=1 f̈iεi) 2ξ2 = 2d2σ4 1 + σ4 1 m∑ i=1 ḟ2 i m∑ i=1 f̈2 i + ρ2σ2 1σ 2 2 m∑ i=1 f̈2 i k∑ i=1 z2i = 2d2σ4 1 + σ2 1a m∑ i=1 f̈2 i , e m∑ i=1 f̈iεiξ 3 = 3dσ4 1 m∑ i=1 ḟ2i + 3ρ2σ2 1σ 2 2d k∑ i=1 z2i = 3daσ2 1, substituting these equations into equ.(11) leads to e(β̂−β)2 = c−2a+6c−4d2σ4 1 +3c−4aσ2 1 m∑ i=1 f̈2i + 135 4 c−6d2a2−36c−5ad2σ2 1 that is the second equation in equ.(8). qed. theorem2. for mse(β̂)(ρ) in theorem 1, the solution to the following minimization problem min ρ mse(β̂)(ρ) (12) exists, satisfying min ρ mse(β̂)(ρ) < minρmse(β̂)( σ2 1 σ2 2 ). proof. because lim ρ→+∞ mse(β̂)(ρ) = σ2 2( ∑k i=1 z 2 i ) −1, mse(β̂)(ρ) is infinitely difxiaojun duan, yi lin:ways of fusing different types of information and how systemic... 240 ferentiable on [0,+∞), mse(β̂)(ρ) has its minimum value on [0,+∞). because d dρ mse(β̂)(ρ) = ( k∑ i=1 z2i )(−2c−3a+ 2ρc−2σ2 2 − 24c−5d2σ4 1 − 12c−5σ2 1 m∑ i=1 f̈2 i + 6ρc−4σ2 1σ 2 2 m∑ i=1 f̈2 i − 405 2 c−7d2a2 + 135ρc−6d2aσ2 2 + 180c−6ad2σ2 1 − 72ρc−5d2σ2 1σ 2 2) (13) we have d dρ mse(β̂)(0) = −( k∑ i=1 z2i )(2( m∑ i=1 ḟ2 i ) −2σ2 1 + 93 2 ( m∑ i=1 ḟ2 i ) −5d2σ4 1+ 12( m∑ i=1 ḟ2 i ) −5 m∑ i=1 ḟ2 i m∑ i=1 f̈2 i σ 4 1) < 0 also, because lim ρ→+∞ d dρmse(β̂)(ρ) = 0, and when ρ → +∞, each term starting from the third in equ.(13) is a higher order infinitesimal when compared to the previous two terms; and when ρ > σ2 1σ −2 2 , the sum of the previous two terms is greater than zero there is ρ0 > 0 so that d dρmse(β̂)(ρ) > 0 when ρ ∈ [ρ0,+∞). therefore, the solution of min ρ mse(β̂)(ρ) satisfies ρ̂ ∈ (0, ρ0). and because d dρ mse(β̂)( σ2 1 σ2 2 ) = σ4 1 2c5 k∑ i=1 z2i (33d 2 − 12( m∑ i=1 ḟ2i + σ2 1 σ2 2 k∑ i=1 z2i ) m∑ i=1 f̈2i ) ̸= 0, ρ = σ2 1σ −2 2 is not a solution of min ρ mse(β̂)(ρ). hence, theorem 2 is proven. qed. remarks: (1) theorem 2 indicates that for nonlinear models, because their least squares estimates are generally biased, when fusing observational data of varied scales of precision, the weight matrix obtained by using gauss-markov theorem, which is derived out of the least squares estimation for linear models, is no longer optimal. the optimal weights can be obtained by solving the minimization problem in equ.(12). (2) if the prior knowledge equ. (3) is a nonlinear model and the first and second order derivatives of the model are the same as those of nonlinear function f(x,β), then the optimal weight can be approximated by ρ = σ2 1σ −2 2 . 241 advances in systems science and applications (2013) vol.13 no.3 2.3 the parameter estimation of heterogeneous data fusion when solving problem (6), one can simply follow the following iterative method: step 1: for an initial weight value ρ0 = σ2 1σ −2 2 , solve the following minimization problem min β∈r1 ∥ y − f(x,β) ∥2 +ρ0 ∥ β̃ − zβ ∥2 (14) to obtain its solution β̂(1); step 2: calculate the mean square error mse(β̂(1))(β̂(1), β) of the estimated parameter at β̂(1); step 3: solve the minimization problem min ρ>0 mse(β̂(1))(β̂(1), β) to obtain ρ1; and step 4: repeat steps 1-4 with the initial value ρ0 replaced by ρ1 until the estimated value of the parameter becomes stable. fig.2 the three samples and their relationships to the information flow of the yoyo model 3 case studies in this section, we will use case studies to illustrate the systemic yoyo model and the evolutionary process naturally existing in system evaluations and estimations. fig.2 shows the connection of these examples and their individual relationships with the information flow of the systemic yoyo model. we will mainly consider the scenario of supplementing prior information. because of the shortage of observational data, the convergence of the system evaluation and estimation process is very slow or becomes stagnant after converging to a certain degree. however, after additional prior information becomes available, the stagnated process will continue to converge; and the speed of the resumed convergence is dependent on the quality of the newly supplied prior information and how consistent it is with the observational data. we will look at three examples to respectively illustrate the following: (1) how to quantitatively measure both the prior and data information; (2) how to xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 242 excavate a new source of observational process data so that the quality of the ultimate system evaluation and estimation is improved; and (3) how to fuse prior information with heterogeneous sets of data effectively. fig.3 the relationship between the accuracy of post fusion parameter estimation and sample size (the left represents posterior fusion estimation accuracy (measured by information); the right the posterior fusion accuracy improvement (measured by gain in information)) table 1 the relation between the fisher information gain and the increase in sample size increase of sample size 1 2 3 4 5 6 7 informa(1)without prior 29.2893 12.9757 7.7350 5.2786 3.8965 3.0284 2.4411 (2)fuse prior,w = 0.4 9.3127 6.0794 4.3675 3.3328 2.6511 2.1741 1.8249 tion gain (3)fuse prior,w = 1 3.8965 3.0284 2.4411 2.0220 1.7106 1.4716 1.2836 increase of sample size 8 9 10 11 12 13 14 informa(1)without prior 2.0220 1.7106 1.4716 1.2836 1.1325 1.0089 0.9062 (2)fuse prior,w = 0.4 1.5601 1.3537 1.1892 1.0555 0.9451 0.8527 0.7744 tion gain (3)fuse prior,w = 1 1.1325 1.0089 0.9062 0.8199 0.7464 0.6833 0.6287 increase of sample size 15 16 17 18 19 informa(1)without prior 0.8199 0.7464 0.6833 0.6287 0.5809 (2)fuse prior,w = 0.4 0.7075 0.6496 0.5992 0.5551 0.5161 tion gain (3)fuse prior,w = 1 0.5809 0.5389 0.5017 0.4686 0.4390 example1. let us look at the information measurement of the prior and observational data. take the prior parameter variance to be sigma0 = 50 and the sample variance sigma1 = 100. let us vary the sample size from 1 to 20 and consider three scenarios: no prior information is fused, and prior information is fused with the 243 advances in systems science and applications (2013) vol.13 no.3 consistency weights w = 1 and w = 0.4, respectively. figure 3 shows the relationship between the accuracy of the post-fusion parameter estimation and the sample size. the fisher information gain is shown in table 1. evidently, when the threshold of fusion accuracy is fixed at 1.5, one needs to repeat his test ten times if he does not fuse any additional prior information; if he fuses additional prior information and sets its weight at w = 0.4 (that means the consistency between the additional prior information and the test data is measured by weight 0.4), then he needs to repeat the test 9 times; and if he fuses additional prior information with weight set at w = 1 (that means the additional prior information has very good consistency with the test data), then he only needs to repeat the test 6 times. this result indicates that with correct fusion of prior information, the same requirement of precision can be met with a fewer number of times the test is repeated. example2. in this case study, we will see how we can speed up the convergence of our recognition of the underlying system by making use of process information and data. the precision evaluation of active homing radar [6] is a complex recognition process of systems. the impact error of clustered warhead missiles with active homing radar is mainly composed of the measurement error of the radar navigation system, the instrumental error of the inertial navigation system (ins), the method error of the terminal guidance, the error of the distribution, and some random error. let us take the impact error of an active homing radar system [6] as our system evaluation performance index and compare the outcomes of the following two methods: one is to fuse the different observational information, directly observed index data, and some indirectly observed data; and the other the point estimate method [7] for the impact error. the main difference here lies in that the later, more traditional method uses only the final impact point error information so that the evaluation outcomes are strongly influenced by the random error of the impact points, while the former method validates the procedure error model with systematic error and the characteristic of the random error by making using of all the observational information. that is how the former method provides more robust evaluation results. in order to avoid analyzing a heterogeneous population created by different testing states, we will transform these different testing states to standard full range process testing states so that the consequent analysis will become manageable. the impact errors of different testing states are denoted as follows: △lradar stands for the assessed impact deviation caused by the error of the radar measurement in the standard overall operational test; ∆l̃radar the measured impact xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 244 deviation caused by the error of radar measurement in a substitute test. for a missile with active homing radar, its warhead-target relative positional x-, y-, and z-errors are mainly caused by its radar’s measurement error. what an actual radar measures includes the range r, azimuth angle a, and elevation angle e. these measurements satisfy the following transformational connection with the measured warhead-target relative x-, y-, and z-positions: x = r cose cosa y = r sine z = r cose sina (15) that is, the performance index (that is to be evaluated), the relative positional error of warhead and target, is a function of the observational quantities (range r, azimuth angle a, and elevation angle e). for details, see fig.4. fig.4 the relationship between radar measurements r, a, and e, and the warhead-target relative x-, y-, and z-positions according to the gaussian law of error propagation, combined with equ. (15), we can obtain the transfer relation from the x-, y-, and z-errors to the r-, a-, and e-errors below: △x △y △z  =  ∂x ∂r ∂x ∂a ∂x ∂e ∂y ∂r ∂y ∂a ∂y ∂e ∂z ∂r ∂z ∂a ∂z ∂e  △r △a △e  =  cose cosa −r cose sina −r sine cosa sine 0 r cose cose sina r cose cosa −r sine sina  △r △a △e  (16) where △x, △y, and △z stand for the warhead-target relative positional errors of the standard testing state, and △r, △a, and △e the errors in the radar measurements r, a, and e of the whole testing state. 245 advances in systems science and applications (2013) vol.13 no.3 firstly, we use the observational data of actual tests to obtain the errors △̃r, △̃a, and △̃e of radar measurements. secondly, we analyze the iterative process of the errors in the measured range and angles. in the following, we use a monopulse radar system as our example to specifically analyze the influencing factors on the errors in radar measured range and angles. based on the analysis on the sources of errors in radar measurements r, a, and e [6], and the law of error synthesis, we can obtain the main errors in radar measured ranges and angles as follows: u2 angle = △2 rader +△2 target +△2 enviroment + ... = △2 thermanoise +△2 phaseunbalance +△2 angularglint+ △2 dynamiclag +△2 clutterinterference + ... u2 distance = △2 rader +△2 target +△2 enviroment + ... = △2 thermanoise +△2 angularglint +△2 dynamiclag+ △2 clutterinterference + ... if we look at radar measured ranges, we see that the systematic error is mainly the dynamic lag error. the time-dependent random error mainly includes the error of clutter interference, thermal noise, and distance glint. now, let us consider the methods of computation for the impact errors of active homing radar systems, as mentioned earlier, under two different testing states: one is to fuse different observational information, including the direct observational index data and indirect observational data(the concrete fusion model refers to subsection 2.2 and 2.3), and the other the point estimate method of impact errors [7]. our simulated radar measurements are the range r, azimuth angle a, and elevation angle e. other than analyzing the single point measurements at the impact moments [7], we also provide a method on how to combine observed process quantities. we respectively model the errors in radar measured range and angle signals, and the speed and acceleration of the actual measured ranges and angles for the two testing states: the simulated substitute tests and the whole process tests. assume that the random error term includes the independent errors of clutter interference, thermal noise, and distance glint. for the systematic deviation, we mainly consider the error caused by dynamic lag. by comparing the point estimate method of impact points and that of combining process information that is used to calculate the impact point error caused by radar errors, the outcomes are listed in table 2. the point estimate method does not employ the process data of the active homing radar. instead, it only applies the observational data of the last moments to calculate the impact point deviation caused by radar errors. when the sample size is small, the outcome xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 246 table 2 the impact error estimate comparisons between the two methods considered (significance level α= 0.01) unit (meter) cross impact cross impact longitudinal longitudinal error point error confidence impact error impact error estimation interval point estimation confidence interval real impact error -12.93 [-15.72,-10.13] 18.78 [18.09,19.47]for standard overall test substitute test -13.96 [-49.43,21.51] 20.23 [-29.56,70.02]point conversion (traditional method) substitute test -12.56 [-15.33,-9.80] 18.33 [17.61,19.05] conversion method fusing with indirect observational information of this method is greatly affected by random factors. on the other hand, the method developed in this research righteously employs the physical background information and observed process data so that interval estimates can be directly produced from the estimation of the parameters. comparing to the point estimate method, our method reduces the effect of random errors and makes the system evaluation and estimation process converge more quickly. example3. let us now look at how to place optimal weights for nonlinear models. assume f(t, β) = 1 + (5 + tβ)0.1, y(t) = f(t, β + ε(t), ε(t) iid∼n(0, 0.012), β̃ = β + η, η ∼ n(0, 0.052) let t = 0.01 ∗ j, j = 1, ..., 100, and the true value of β is 8. let us generate 50 observational data {yi(t)}1001 , β̃i, i=1,...,50. then, when ρ = 0.012/0.052, let us solve the minimization problem (6) 50 times, producing the mean square error 0.0394. and when ρ = 2.33∗0.012/0.052, the mean square error reaches the minimum value 0.0372. however, in theory, the mean square error is 0.0361. this end indicates that for a nonlinear system, the optimal choice of weights for fusing information not only depends on observational precisions but also the model curvature involved. all the three examples above verify the transformational relationship of system model information, prior knowledge, and observational data in the recognition process of the underlying system, where the model information can be 247 advances in systems science and applications (2013) vol.13 no.3 strengthened through the usage of process data, and when the observational data is insufficient, the shortage in information can be made up by supplementing additional prior knowledge. 4 summary continuing literature[1], in this paper we studied how to fuse heterogeneous sets of data together consistently so that better results can be obtained for system evaluations and estimations. it is shown that if the types of data considered are few and the available observational data is insufficient, one can consider obtaining process information and additional prior knowledge. it is because the limited amount of observational data could make the process of system evaluation and estimation converge extremely slowly or stop converging completely after reaching a certain degree. with process information or additional prior knowledge added, the convergence will continue. the speed of the resumed convergence is dependent on how the process information is applied, how good quality the prior knowledge is and how consistent the newly adopted prior information is with the available observational data. this end has been well illustrated by the case studies considered in section 3. when there is only a small sample available for a specific system evaluation and estimation, this work provides the theoretical guideline for how to excavate other sources of information and how newly adopted information should be fused with what is available. acknowledgements this work is supported by the natural science foundation of china (60974124), the program for new century excellent talents in university and the projectsponsored by srf for rocs, sem in china, the key lab open foundation for space flight dynamics technique (sfdlxz-2010-004). references [1] xj duan, y. lin. (2011), “conservation law of information and its application in evaluation and estimation of complex systems”, kybernetes: the international journal of cybernetics, systems, and management science, vol.40 no.1/2, pp.262-274 [2] y. lin. (2008), systemic yoyos: some impacts of the second dimension, taylor and francis, new york. [3] s. s. mao. (1999), bayesian statistics. beijing: china statistics publishing house. xiaojun duan, yi lin:ways of fusing different types of information and how systemic... 248 [4] d. m. bates and d. g. watts. (1997), nonlinear regressive analysis and its application, translator: bocheng wei. beijing: china statistics publishing house. [5] x. p. zhang, j. h. zhang, and h. w. xie (2003), “a few discussion of samples, a prior information and bayesian statistical decision”, acta electronica sinica, vol.31 no.4, pp.536-538. [6] d. c. wang, j. h. ding, w. d. chen (2006), radar measurement technique in precise tracking, beijing: publishing house of electronics industry. [7] g. wang, x. j. duan, z. m. wang (2009), “conversion method of impact dispersion in substitute equivalent tests vased on error propagation”, defence science journal, vol.59 no.1, pp.15-21. corresponding author xiaojun duan can be contracted at: xjduan@nudt.edu.cn advances in systems science and application (2016) vol.16 no.2 1-14 evaluation of a germ stability of a differentiable mapping defined by a mathematical model abdykappar ashimov1 and yuriy borovskiy2 1academician of national academy of science of republic of kazakhstan. 2candidate of physical and mathematical sciences. abstract the paper presents a set of algorithms to assess: the set of singular points of a differentiable mapping defined by some mathematical model; the stability of the germ of this mapping in its singular point, and (in the case of its stability) the form of such a germ in cases of corank 1 and all the possible relations of dimensions of the domain and mapping image. the stability of these germs in all singular points of the mapping under study is a necessary condition for the stability of the mapping in its domain. the implementation of the developed algorithms can be used to verify the model under study by investigating mappings defined by it. keywords germ of smooth mapping; germ stability; singular point; algorithm; mathematical model. 1 introduction as is known, in natural science and economics to study properties of the objects under investigation are widely used mathematical models related to different classes. important stage of accepting a model to use it in practice is verification and validation v&v[1, 2]. successful testing of the model using v&v methods gives reason to carry on the practice the conclusions drawn on the basis of calculations of the model. note that the currently used v&v methods contain no estimates of conditions for preserving qualitative properties of the considered models under arbitrarily small changes in these models. that is, these methods permit transfer of initial model (under its arbitrarily small changes) to the model with totally different properties[3]. necessary conditions for preserving qualitative properties of model under its small changes can be conditions of the weak structural stability of the model, being dynamical system and stability conditions of mappings, defined by the model under study[4–6]. the study of authors ashimov et al on the basis of the robinson theorem on sufficient conditions for the weak structural stability and an algorithm for constructing a symbolic set provide convenient for practice numerical algorithms for the evaluation of weak structural stability of dynamical systems[4, 7–9]. such system stability is required for small sensitivity of its phase portrait to small disturbances in model. these papers also provide the application examples of these algorithms to macroeconomic models. 2 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... in literature [8] authors propose v&v method, based on stability estimation of the defined by the model differentiable mappings in the sense of definition used in the theory of singularities of differentiable mappings[5, 6]. literature [5] provide sufficient theoretical conditions for the stability of differentiable mapping when the mapping is an immersion, submersion or submersion with a fold. proposed in [8] on the basis of these theoretical propositions the numerical algorithms permit to evaluate the set of singular points of the studied mapping, and also its stability (and instability) when the mapping belongs to the specified class of mappings. if using these algorithms the defined by model mappings are evaluated as stable, the appropriate model can be considered as more adequate to describe corresponding phenomena of reality. however, the proposed in [8] algorithms are insufficient to assess the stability of these differentiable mappings in general case (if they do not belong to classes of immersions, submersions or submersions with a fold). based on the fact that a required condition for the stability of a differentiable mapping is the stability of its germs in all the singular points of the mapping, the present work provides (based on theoretical propositions of arnold et al[6]) an algorithm to estimate the stability of such germs. also this paper proposes three algorithms to evaluate the form of stable germ with corank 1 for the cases of all possible ratios of dimensions of the domain and the image of the studied mapping. the use of the developed algorithms in the framework of v&v mathematical model would enable to evaluate the conditions of the low sensitivity of qualitative properties of the models to its small perturbations. 2 algorithms to evaluate a set of singular points of differentiable mapping and germ stability of the mapping in its singular point as noted above, the evaluation of stability of the given by model mapping f can be replaced by the evaluation of stability of germs of f in all singular points of this mapping. ashimov et al[8] proposed an algorithm to estimate the set of singular points of the differentiable mapping in the parallelepiped d f : d0 → e (1) (where dim(d) = m, d0 the set of internal points d) using the set of parallelepipeds d̃ with arbitrarily small size, covering the estimated set. for the convenience of the reader, this algorithm (as well as an illustration of its application), we present below. designate the vector of arguments of the mapping (1) through p = (p1, ..., pm) ∈ d, and respective image of the point p vector of the model solutions designate through y = y(p) = (y1, ..., yn) ∈ e, ( d ⊂ rm and e ⊂ rn some domains). in this case the jacobian matrix with dimension v ∗ n for the mapping (1) in the advances in systems science and application (2016) vol.16 no.2 3 point p would be written as follows: j(p) = ( ∂yi ∂pj (p) ) i=1,...,n; j=1,...,m (2) evaluation of the jacobian matrix (2) in some point p ∈ d, derived using the numerical differentiation, we also designate through j(p). designate total quantity of maximal order minors in j(p) through l. estimate of determinant value of such a minor of order min(v, n) in j(p) for p ∈ d designate through |mi(p)|, i = 1, ..., l. algorithm 1 to estimate a set of singular points of mapping (1). 1) parallelepiped d is divided into sufficiently great number of (elementary) parallelepipeds dk with the same size, and a grid p composed of n points is determined, which are vertices of chosen parallelepipeds: p = {pj : j = 1, ..., n}. 2)values of all j(pj) matrix elements are computed for j = 1, ..., n . 3) for i = 1, ..., l =; are computed determinants |mi(pj)|. 4) for every i = 1, ..., j = 1, ..., l the set d(i) is determined in the following way. d(i) is a union of all (closed) parallelepipeds dk with the property: not all values of |mi(pj)| in vertices dk have the same sign. 5) the set d̃ = ∩l i=1d(i) is found. 6) if set d̃ is empty, then stop. 7) if not, the steps 1 5 of this algorithm are performed with replacement of domain d by d̃ and decrease of parallelepipeds size, participating in subdividing d̃ until the diameter dk is not less than some given forward number . the abovementioned algorithms (as well as provided in [8] algorithms for evaluating points of a fold) were tested based on the whitney mapping, defined by relations y1 = x31 + x1x2, y2 = x2,where (x1, x2) ∈ r2. it is known that singular points of this mapping form parabola 3x21 + x2 = 0 and all points of this parabola, except the origin of coordinates, are fold points; origin of coordinates the point of cusp[6]. fig.1 presents the derived by the mentioned algorithms points estimates of the whitney mappings fold as the set of vertices of parallelepipeds p̃ covering the fold. note that the distance from every point in the set p̃ to the fold does not exceed the diameter of elementary square from the set , namely, the number 0.05 √ 2. here are further required to construct an algorithm to evaluate the stability of the germ of a differentiable mapping notations and the theoretical propositions from [6]: definitions of a stable and infinitesimally v stable germ, as well as theorems, enabling to obtain sufficient conditions for the stability of the germ by conditions of its infinitesimally v stability. then the conditions from the definition of infinitesimally v stable germ we have rewritten equivalently as condition 6. the proposed further algorithm 2 permits to evaluate condition 4 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... 6, that is, the infinitesimal v stability, and therefore the stability of the germ under study. fig. 1 estimate of the whitney mappings fold. here are further required to construct an algorithm to evaluate the stability of the germ of a differentiable mapping notations and the theoretical propositions from [6]: definitions of a stable and infinitesimally v -stable germ, as well as theorems, enabling to obtain sufficient conditions for the stability of the germ by conditions of its infinitesimally v -stability. then the conditions from the definition of infinitesimally v -stable germ we have rewritten equivalently as condition 6. the proposed further algorithm 2 permits to evaluate condition 6, that is, the infinitesimal v -stability, and therefore the stability of the germ under study. consider a germ of a differentiable mapping f : rm → rn in some of its singular point x0 ∈ rm . we denote also the mapping f : d → rn, specifying the germ, determined in a certain parallelepiped d ⊂ rm centered at point x0 . such a germ is called stable if for any other arbitrarily close (in the corresponding topology) to it mapping f there exist diffeomorphisms in some neighborhoods of points of the domain x0 and image f (x0) , during application of which the mapping f converted into f . the formal definition of the (left − right, differentiable) advances in systems science and application (2016) vol.16 no.2 5 stability of germ f is provided in the monograph[4]. definition 1. a germ f is called stable, if for arbitrarily small neighborhood u of point x0 exists such a neighborhood e of the mapping f , that for any mapping f from e is found the point x ∈ u such that the germ f in x (left-right, differentiable) equivalent to the germ f in x0 . topology in the set of all differentiable germs in point x0 is given by the set of neighborhoods of the form[6]: e = {f : sup |α|≤k,x∈u,i=1,...,n | ∂ |α| ∂αx f i(x)− ∂|α| ∂αx f i 0(x)|} (3) here f i coordinate function of the germ f = (f 1, ..., fn), (i = 1, ..., n); f0 fixed germ the center of neighborhood e; u some neighborhood of point x0 ; x = (x1, ..., xn), α = (α1, ..., αm); αj integral non-negative number; |α| = |α1 + ...+ αm| ; k arbitrarily large number; ∂|α| ∂αx = ∂|α| ∂α1x1...∂αmxm . further we introduce the following notations. let ∂f / ∂xj a germ of the mapping, defined by partial derivatives of coordinate functions of the mapping f : ∂f / ∂xj = ( ∂f 1/ ∂xj , ..., ∂fn/ ∂xj ) ; j = 1, ...,m. basis germs of mappings rm → rn we call constant germs er, (r = 1, ..., n) which have r-th coordinate function identically equal to 1, and the rest coordinate functions zero: e1 = (1, 0, ..., 0), ..., en = (0, 0, ..., 1). let ax0 algebra of all germs of differentiable functions rm → r in point x0, (ax0) n ax0 module of all germs of differentiable mappings rm → rn in point x0. designate trough submodule in (ax0) n generated by the following m+ n2 germs of mappings: ∂f / ∂xj ,where j = 1, ...,m and f ier,where i, r = 1, ..., n. definition 2. [5]a germ of the mapping f in point x0 is called infinitesimally v −stable, if factor module (ax0) n/b is generated above r images of basis germs e1, e2, ..., en. note that from infinitesimal v -stability of a germ follows its stability. this fact follows from the following theorems 3 and 4, provided in arnold et al[6]. theorem 3. infinitesimal v -stability of a germ is equivalent to its infinitesimal stability. theorem 4. (mather theorem (local version)) infinitesimally stable germ is stable. using facts: submodule b is a set of all linear combinations of germs of mappings ∂f / ∂xj and f ier with coefficients from ax0 and fulfillment of condition of definition 2 is equal to the fact that any element of the factor module (ax0) n/b of the form g+b where g some germ from (ax0) n can be written as the sum of some linear combination of germs e1, e2, ..., en with numerical coefficients and submodule b reformulate the condition of infinitesimal 6 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... v -stability of the germ f from definition 2 by the following equivalent way. condition 5. any germ g from (ax0) n can be presented as the sum of some linear combination with numerical coefficients of constant germs e1, e2, ..., en and linear combination with coefficients from ax0 of the mentioned germs ∂f / ∂xj , f ier. since for the case x0 = 0 any germ g from (ax0) n can be written as the sum of constant germ (resulting from linear combination with numerical coefficients of germs e1, e2, ..., en) and the germ taking zero value in point x0, and the last germ is written as the sum of linear combinations of functions of the form xjer (xj-j-th coordinate of vector x) with coefficients from ax0 , then in condition 5, the term “any germ g”, can be replaced by “any germ of the form xjer where j = 1, ...,m; r = 1, ..., n”. moreover, for sufficiently small neighborhood x0=0 coefficients of germs, generating submodule can be considered in the rough (with precision till infinitely small of high-order infinitesimality) polynomial germs qk(x) from ax0 , of order not exceeding k, where k sufficiently large fixed number: qk(x) = ∑ α:|α|≤k cαx α (4) here xα = (x1)α1 ...(xm)αm , cα coefficient of the polynomial. therefore, for the case x0=0 condition 5 infinitesimally v -stable germ f can equivalently be rewritten in the following way. condition 6. any germ of the mapping of the form xjer where j = 1, ...,m; r = 1, ..., n (with precision till infinitely small of high-order infinitesimality) can be written as linear combination of germs ∂f / ∂xj , f ier, with polynomial coefficients as (4). from theorem 4 results that condition 6 is sufficient for the stability of germ f in its singular point the origin of coordinates. here are based on an assessment of condition 6 enlarged algorithm to evaluate the stability of the germ in a singular point x0 of the mapping f (defined in dε parallelepiped centered at x0 , where ε is a diameter of the parallelepiped) defined by some model. estimate of singular point x0 can be obtained by algorithm 1 and selecting one point from the obtained set p̃ . algorithm 2 to estimate the stability of germ of the differentiable mapping 1) appropriate parallel shift and normalizing coordinate system in rm yields that singular point x0 corresponds with the origin of coordinates, and parallelepiped dε is close to cube. here ε sufficiently small given number. 2) dividing every edge dε into sufficiently large even integer n of equal segments, yields a grid pε composed of nm points. 3) determine grid functions f i, i = 1, ..., n, corresponding to coordinate functions of the mapping f and determined in grid junctions p ∈ pε. advances in systems science and application (2016) vol.16 no.2 7 4) using numerical differentiation (with step less than step of the grid pε) mine grid functions f i j , i = 1, ..., n, j = 1, ...,m,determined in grid functions pε and being estimates of partial derivatives ∂f / ∂xj in points p ∈ pε. 5) setting grid mapping linear combination of germs by its coordinate grid functions determined in pε (here i = 1, ..., n ): yi = ri(x, {αj : j = 1, ...,m; |αj | ≤ k}, {αi,r : r = 1, ..., n; |αi,r| ≤ k}) = m∑ j=1 qj(x)f i j + n∑ r=1 qi,r(x)f r (5) here qj and qi,r some polynomials of degrees, not exceeding k with arbitrary coefficients αj and αi,r respectively. k sufficiently large fixed number. 6) using numerical differentiation (with step less than step of the grid pε ) determine grid functions f i β, i = 1, ..., n, β = (β1, ..., βm), |β| ≤ k, determined in grid junctions pε and being estimates of partial derivatives ∂|β| ∂βx f i.(for the cases |β|=0 and |β|=1 these functions were determined above, correspondingly, in steps 3 and 4 of the algoritm). 7) setting grid functions ri β, i = 1, ..., n, β = (β1, ..., βm), |β| ≤ k , determined in grid functions pε and being estimates of partial derivatives ∂|β| ∂βx ri using the derived f i β values and values of partial derivatives of polynomials: ∂|β| ∂βx qj , ∂|β| ∂βx qi,r in the mentioned grid functions. 8) for every r = 1, ..., n and j = 1, ...,m the steps 9, 10, 11 are performed. 9) for i = 1, ..., n determine functions m i r,j ({αj : j = 1, ...,m; |αj | ≤ k}, {αi,r : r = 1, ..., n; |αi,r| ≤ k}) = sup x∈pε;β:|β|≤k ∣∣∣∣∣ri β(x)− ∂|β| ∂βx ( xjeir )∣∣∣∣∣ (6) here for β = 0, ∂|β| ∂βx ( xjeir ) = xj , if i = r and j = j, if not, ∂ ∂βx ( xjei i ) = 0. for cases |β| > 1 ∂|β| ∂βx ( xjei i ) ≡ 0. 10) finding estimate of deviation of the germ xjer from approximating it linear combination. m = 1 ε inf {αj :j=1,...,m;|αj |≤k},{αi,r:r,i=1,...,n;|αi,r|≤k} ( sup i∈{1,...,n} m i r,j ) (7) 11) if m > δ, where δ sufficiently small given forward number, then go to step 13. 8 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... 12) the germ f is assessed as stable. stop. 13) two-fold decrease of ε and all of the edges of parallelepiped dε . 14) if ε > ε0 , where ε0 sufficiently small given forward number, then go to step 2. 15) question about f germ stability left unanswered. stop. 3 algorithms to estimate the form of stable germ of corank 1 of the differentiable mapping after evaluation by the algorithm 2 of the studied germ f in its singular point x0 as a stable, there is a question on the evaluation of its form in accordance with [6] classification of genotypes of stable germs. in [6] there are three morin theorems determining the form of stable germ in the case where the germ has corank 1; the question on classification of stable germs with greater corank is still open. here corank of germ f its singular point x0 is determined as the difference between the number min(m,n) and the rank of the jacobian matrix of germ f in x0 . this section presents these morin theorems and developed algorithms enabling for respective estimation of the form of the germ f , and to evaluate some of the other characteristics of the germ under study. let there be a differentiable mapping f : rm → rn defined in some neighborhood of its singular point 0 ∈ rm, f (0) = 0 ∈ rn. next, consider the options of relations of dimensions values m and n. 2.1. let the dimension of the image and the full domain of the germ f match: n = m. the morin theorem[6], which describes these stable germs is as follows. theorem 7. a stable germ f : (rn, 0) → (rn, 0), of corank 1 left-right equivalent to the germ (i.e., driven by a diffeomorphic changes of coordinates in the spaces of domain and the image to the form) ỹ1 = (x̃1)k+x̃2(x̃1)k−2+...+x̃k−1x̃1), ỹ2 = x̃2, ... ỹn = x̃n, (8) here ỹ = (ỹ1, ỹ2, ..., ỹn) and x̃ = (x̃1, x̃2, ..., x̃n) new coordinates correspondingly in spaces of image and domain of the germ f . k integer (indicator of germ), 2 ≤ k ≤ n+ 1. we present an enlarged algorithm for assessing the axis oỹ for mapping (8), and the k order value of genotype specified in this theorem of the mapping f . it is based on the obvious remark that in the space of images rn there is the only direction (defined by the axis oỹ), satisfying the following property: the derivative of the mapping coordinate f corresponding to ỹ1 at the origin of coordinates advances in systems science and application (2016) vol.16 no.2 9 o in any direction in domain of f is zero. for any other axis in space of images the corresponding derivative at point o in some direction is different from zero. algorithm 3 1) finding by algorithm 1 estimate p of the set of singular points of the mapping f : d → rn, d ⊂ rn. 2) if p is empty, then stop. there are no singular points of the mapping f . 3) choosing the singular point x0 ∈ p and estimating the stability of the germ f in x0 by algorithm 2. 4) if the germ f is not assessed as stable, then stop. 5) computing the jacobian estimate of the mapping f in point x0 and its rank r = rank(j(x0)). 6) if r ̸= n− 1 then stop. corank f in singular point x0 is greater than one. 7) setting the shifted mapping y = f (x−x0)+f (x0) = f (x), for which x = 0 is the investigated singular point and f (0) = 0 . 8) finding the estimate of axis oỹ1. 8.1) for arbitrary unit vector ỹ ( |ỹ| = 1 ) setting coordinate function of n variables a = f̄ỹ(x) = prỹf̄ (x) projection of vector f̄ (x) to vector ỹ. 8.2) setting function of n variables s = g(ỹ) = ∑n i=1 ( ∂f̄ỹ(x) ∂xi )2 x=0 , where ∂f̄ỹ(x) ∂xi appropriate difference derivative. 8.3) finding the estimate of constrained minimum s0 of function s = g(ỹ) under constraint |ỹ| = 1. 8.4) if s0 ≈ 0 then the detected ỹ is an estimate of axis unit vector ỹ1, if not, then stop. in this case recalculation of steps 1-8 of the algorithm is possible with decreased values of grid steps and increments. 9) finding the estimate of value of indicator k of the germ f . 9.1) for arbitrary unit vector x ( |x| = 1 ) setting the constraint of coordinate function f̄ỹ1(x) on axis, determined by this vector: c = fx(t) = f̄ỹ1(tx). 9.2) assignment k = 2. 9.3) finding the estimate of maximum value of module of k-th derivative of function c = fx(t) at zero: m = max x,|x|=1 |f (k) x (t)|t=0|, where the estimate in the form of corresponding difference derivative is used as a derivative. 9.4) if m > ε, where ε sufficiently small number, then output of k value and stop. 9.5) increasing k value by one: k := k + 1. 9.6) if k ≤ n+ 1, then go to step 9.3. 9.7) indicator k of point x0 is not defined. stop. 2.2 consider now the case, when the germ image dimension f̄ is strongly greater than the domain dimension: n > m. take the morin theorem[5], describ10 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... ing the mentioned stable germs for this case. theorem 8. in the case n > m stable germ f̄ : (rm, 0) → (rn, 0), of corank 1 is right-left equivalent the germ (that is, reduced by diffeomorphic substitution of coordinates in spaces of domain and image to the form) ỹ1 = (x̃1)k + x̃2(x̃1)k−2 + · · ·+ x̃k−1x̃1, ỹ2 = x̃k(x̃1)k−1 + x̃k+1(x̃1)k−2 + · · ·+ x̃2k−2x̃1, · · · ỹt = x̃(t−1)k−t(x̃1)k−1 + x̃k+1(x̃1)k−2 + · · ·+ x̃tk−tx̃1, ỹt+1 = x̃2, · · · ỹn = x̃m. (9) here k ≥ 2 − integer; t = n −m + 1; t(k − 1) ≤ m; ỹ = (ỹ1, ỹ2, . . . , ỹn) and x̃ = (x̃1, x̃2, . . . , x̃n) − new coordinates correspondingly in spaces of image and domain of the germ f . we present an enlarged algorithm for assessing the plane ∏ spanned by the axes ỹ1, ỹ2, . . . , ỹt for mapping (9), as well as the value of k genotype order of the mapping f specified in this theorem. in the construction of this algorithm is used the following remark. for any axis of the plane ∏ the coordinate derivative of the mapping f corresponding to the axis at the origin o in any direction in space of the inverse images is zero. for any axis that is not in from the space of images the derivative at point o in some direction is different from zero. algorithm 4. 1) finding by algorithm 1 estimate p of the set of singular points of the mapping f : d → rn, d ⊂ rm. 2) if p is empty, then stop. there are no singular points of the mapping f . 3) choosing the singular point x0 ∈ p and estimating the stability of the germ f in x0 by algorithm 2. 4) if the germ f is not assessed as stable, then stop. 5) computing the jacobian estimate of the mapping f in point x0 and its rank r = rank(j(x0)). 6) if r ̸= m − 1 then stop. corank f in inverse image in singular point x0 is greater than one. 7) setting the shifted mapping y = f (x−x0)+f (x0) = f̄ (x), for which x = 0 is the investigated singular point and f̄ (0) = 0. 8) finding the estimate of plane ∏ , spanned by the axes ỹ1, ỹ2, ...,ỹt (where t = n−m+ 1) for the mapping (9). 8.1) for arbitrary unit vector ỹ ( |ỹ| = 1 ) setting coordinate function n of variables a = f̄ỹ(x) = prỹf̄ (x) projection of vector f̄ (x) to vector ỹ. 8.2) setting function of n variables s = g(ỹ) = ∑n i=1 ( ∂f̄ỹ(x) ∂xi )2 x=0 , where ∂f̄ỹ(x) ∂xi advances in systems science and application (2016) vol.16 no.2 11 appropriate difference derivative. 8.3) assignment k = 1, where k enumerator of coordinate axes of the required plane ∏ . 8.4) finding the estimate of constrained minimum s0 of function s = g(ỹ) under constraints:|ỹ| = 1, ỹ · ỹi = 0,where i = 1, . . . , k − 1 numbers of the derived at the previous steps coordinate axes of the plane ∏ , · sign of scalar product (when k = 1 this constraint is not used). 8.5) if s0 ≈ 0 then the detected value ỹ is an estimate of axis unit vector ỹk, if not, then stop. in this case recalculation of steps 1-8 of the algorithm is possible with decreased values of grid steps and increments. 8.6) increasing k by one: k := k + 1. 8.7) if k ≤ t, then go to step 8.4. 9) finding the estimate of value of indicator k of the germ f . 9.1) for arbitrary unit vector x̃ ∈ rm (|x̃| = 1) and arbitrary unit vector ỹ ∈ rn (|ỹ| = 1) setting the constraint of coordinate function f̃ỹ(x) on axis, determined by this vector: c = fx,y(t) = f̃y(tx). 9.2) assignment k = 2. 9.3) finding the estimate of maximum value of module of k-th derivative of function c = fx,y(t) at zero: m = max y∈π,|y|=1,x,|x|=1 ∣∣∣f (k) x,y (t)|t=0 ∣∣∣, where the estimate in the form of corresponding difference derivative is used as a derivative. 9.4) if m > ε, where ε sufficiently small number, then output of k value and stop. 9.5) increasing k value by one: k := k + 1. 9.6) if t(k − 1) ≤ m, then go to step 9.3. 9.7) indicator k of point x0 is not defined. stop. 2.3 now consider the rest case, when the dimension of f̄ germ image is less than the dimension of domain: n < m. we present the morin theorem[6] for this case, describing the mentioned stable germs of corank 1. theorem 9. in the case n < m the stable germ f̄ : (rm, 0) → (rn, 0) of corank 1, the corank of the second differential the genotype of which at zero does not exceed 1, right-left is equivalent to the germ (i.e., by diffeomorphic substitutions of coordinates in spaces of domain and image reduced to the form). ỹ1 = (x̃1)k + x̃2(x̃1)k−2 + · · ·+ x̃k−1x̃1 ± (x̃n+1)2 ± (x̃n+2)2 ± · · · ± (x̃m)2, ỹ2 = x̃2, · · · ỹn = x̃n. (10) here ỹ = (ỹ1, ỹ2, . . . , ỹn) and x̃ = (x̃1, x̃2, . . . , x̃n) − new coordinates correspondingly in spaces of image and domain of the germ f . k − integer (genotype order),2 ≤ k ≤ n+ 1. 12 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... below is presented an enlarged algorithm for assessing the axis oỹ1 for the mapping (10), and k order value of the genotype of the specified in this theorem mapping f . the first part of the algorithm (steps 1-8) is similar to the corresponding steps of algorithm 3. algorithm 5. 1) finding the estimate p of the set of singular points of the mapping f : d → rm, d ⊂ rn using algorithm 1. 2) if p is empty, then stop. there are no singular points of the mapping f . 3) choosing the singular point x0 ∈ p and estimating the stability of the germ f in x0 by algorithm 2. 4) if the germ f is not assessed as stable, then stop. 5) computing the jacobian estimate of the mapping f in point x0 and its rank r = rank(j(x0)). 6) if r ̸= m− 1 then stop. corank f in singular point x0 is greater than one. 7) setting the shifted mapping y = f (x−x0)+f (x0) = f̄ (x) , for which x = 0 is the investigated singular point and f̄ (0) = 0 . 8) finding the estimate of axis oỹ1 . 8.1) for arbitrary unit vector ỹ ∈ rm(|ỹ| = 1) setting coordinate function n of variables a = f̄ỹ(x) = prỹf̄ (x) − projection of vector f̄ (x) to vector ỹ. 8.2) setting function of m variables s = g(ỹ) = ∑n i=1 ( ∂f̄ỹ(x) ∂xi )2 x=0 , where ∂f̄ỹ(x) ∂xi − appropriate difference derivative. 8.3) finding the estimate of constrained minimum s0 of function s = g(ỹ) under constraint |ỹ| = 1 . 8.4) if s0 ≈ 0 then the derived value is an estimate of axis unit vector oỹ1 , if not, then stop. in this case recalculation of steps 1-8 of the algorithm is possible with decreased values of grid steps and increments. 9) finding the estimate of value of indicator k of the germ f . 9.1) for arbitrary unit vector x(|x| = 1) setting the constraint of coordinate function f̄ỹ1(x) on axis, determined by this vector: c = fx(t) = f̄ỹ1(tx) . 9.2) assignment i = 1. i enumerator of axes, corresponding to morse singularity of function f̄ỹ1(x). 9.3) finding the estimate of maximum value of module of 2nd derivative of function c = fx(t) at zero: m = max x,|x|=1;xx̃j=0,j ε , where ε sufficiently small number, then output of the required k value and stop. 9.10) increasing k value by one: k := k + 1. 9.11) if k ≤ n+ 1 , then go to step 9.8. 9.12) characteristics k of point x0 is not defined. stop. the use of algorithms 2, 3 or 4 enables to evaluate the axis oỹ1 (or the plane∏ for the case n > m ) in the space of the mapping image under study, with the following property. the coordinate derivative of the mapping f corresponding to this axis (or any axis of the mentioned plane for the case n > m ) at the origin (corresponding to a singular point x0 ) in any direction in space of the inverse images is zero. detection of the specified axis (or plane) may be important to study the behavior of the germ (in the singular point) of the mapping defined by the studied model. coordinate function corresponding to this axis gives an example of the local invariant of model, in the sense that in some neighborhood of the singular point under study the derived function is constant (with accuracy up to infinitesimals of highest order) with respect to all exogenous parameters of the model involved in the construction of the initial mapping. if the studied point is estimated (using algorithm 1) as a non-singular (regular) for this mapping, then space of images does not contain the axis with the above properties and the mapping has no local invariant in the neighborhood of this regular point. 4 conclusion this paper presents a set of algorithms that allow: to estimate the set of singular points of a differentiable mapping created by mathematical model. to estimate the stability of germ of this mapping in its singular point. in the case of such a stability of germ of corank 1 to estimate k indicator of its genotype and the corresponding axis in the domain of the germ, which determines the local invariant of the mapping under study. the developed algorithms can be used to evaluate the possibility of transferring obtained on the basis of model results into practice. 14 abdykappar ashimov and yuriy borovskiy: evaluation of a germ stability ... references [1] o. balci. (1998), “verification, validation and testing”, in handbook of simulation: principles, methodology, advances, applications, and practice, john wiley & sons, new york, pp. 335-393. [2] r.c. kennedy, x. xiang, g.r. madey and t.f. cosimano. (2005), “verification and validation of scientific and economic models”, agent 2005 conference proceedings, chikago, pp. 177-192. [3] v.i. arnold (1988), geometrical methods in the theory of ordinary differential equations, springer-verlag, new york. [4] c. robinson (1980), “structural stability on manifolds with boundary”, journal of differential equations,no. 37. pp. 1-11. [5] m.golubitsky and v. gueillemin. (1973),stable mappings and their singularities,springer-verlag, new york, heidelberg, berlin. [6] v.i. arnold, s.m. gusein-zade and a.n. varchenko. (1985), “singularities of differential maps”, birkhauser, boston, basel, stuttgard, vol. 1. [7] a. ashimov et al (2013), macroeconomic analysis and parametrical control of a national economy, springer, new york. [8] a. ashimov et al. (2014), “the theory of parametric control of macroeconomic systems and its applications (i)”, advances in systems science and application, vol. 14, no. 1, pp. 1-21. [9] e.i. petrenko (2006), “development and realization of the algorithms for constructing the symbolic set”, differential equations and control processes , no. 3, pp. 55-96. corresponding author abdykappar ashimov can be contacted at: ashimov37@mail.ru мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 22-28 system of criteria and indicators for the development of resource-based multiclusters dmitry l. napolskikh, tatyana v., yalyalieva, nina i. larionova, elena a. murzina volga state university of technology annotation the paper presents the system of criteria and indicators of development of resource-based multiclusters able to be used for the purposes of government control of performance efficiency of the territory on which they are located. the limitations of the paper are general adaptation mechanisms of a region economic system within the frames of innovation development that would significantly enrich the theoretical aspect of the study. an indisputable advantage of the suggested system of the parameters for the purposes of governmental control of the resource-based multiclusters efficiency is its versatility, as it is adapted to any kind of regional economic system. key words: governmental control, economy clustering, natural resource-based multicluster, efficiency of governmental cluster policy. 1 introduction in the modern conditions of global economy, the competitiveness of the cluster residents is interpreted as a concentrated expression of the production, scientific, educational and technical advantages implemented in innovation technology, goods and services [1]. the multi-subject composition of the multicluster can be classified according to the subject composition of the “triple spiral” model, i.e., it appears possible to single out «state and municipalities», «science and education» and «business [2]. the aggregate social and economic effect produced through implementation of governmental cluster policy, which is achieved by applying structural modifications to the economy, can be used as the efficiency indicator [3]. the formation of resource-based multiclusters at the regional level is based not only on building an industry infrastructure but also on the formation of a network interaction structure between innovation businesses and the government. in this case, the methods used to assess the natural resource-based multiclusters can be applied to control the efficiency of the use of the state resources [4]. the practical importance of this approach lies mainly in the opportunities to formulate and implement large national-and regional-scale investment and innovation projects. issues of government regulation of the cluster have been studied in works by [5, 6]. in this regard, studying theoretical approaches to improving government control of the formation of natural resource-based multiclusters at the regional level becomes an important scientific task. 2 data and method advances in systems science and application(2016) vol.16 no.4 23 most international studies on the problem of efficiency of the resource-based multiclusters are devoted to comparative analysis of transaction costs of economic activity at the regional level. in the modern economy system the structure of transaction costs are determined by a wide range of technological and ecological factors [7, 8]. minimization of transaction costs as a main function of government control allows us to consider costs of interactions between economic agents in the multicluster. an multicluster performance efficiency for economic agents can be defined as the ratio of the benefits of transaction costs reduction in multicluster and the institution’s maintenance costs and/or institutional constraints loss: where: me = multicluster efficiency tc = transaction costs imc = institution maintenance costs icl = institutional constraints loss efficiency of informal institutions of control of the formation of resource-based multiclusters can be evaluated using the following formula: where: uie = informal institution efficiency lgc = losses associated with government control iic= informal interactions costs icl = informal constraints loss if the informal institution is of a “grey” (illegal) nature, its efficiency can be evaluated using the following formula: where: gie = grey institution efficiency cla = losses associated with activity legitimization lgc = losses associated with government control iic = informal interaction costs icl = losses associated with informal constraints rial = losses associated with risks of illegal activities in the context of the network analysis of the of the formation of resource-based multiclusters at the regional level, it is expedient to use the instruments of the mathematical graph theory. considering the development criteria of the resource-based multiclusters, we can single out such parameters as the density of the institutional environment and its innovation conductivity, which produce effect on the diffusion speed of the management and production innovations within the cluster. let us make an assumption that any economical agent of the multicluster may be connected with other economical subjects through identifiable bilateral relations, both direct connections and indirect ones operating through intermediaries. the multicluster density can be characterized as the ratio of the stable formal connections between the multicluster organizations and businesses to their total number. the multicluster density can be represented with the formula below: where: d = multicluster density sxi = number of the stable formal connections between the ith muticluster agent with the other ones 24 dmitry l. napolskikh, tatyana v., yalyalieva, nina i. larionova, et al.: system of criteria and indicators… n = total number of the multicluster agents considering the multicluster density, it is required to single out the institutional environment integrity parameters related to the cluster. the multicluster integrity can be characterized as the ratio of the sum of the number of the indirect connections of each multicluster agent with the other ones to their total number: where: i = multicluster integrity uxi = number of the multicluster agents connected with indirect non-formal connections with the ith cluster agent n = total number of the multicluster agents the multicluster complementarity can be characterized as the average number of the primary communication channels between each multicluster agent and the other ones: where: c= multicluster complementarity axi = number of the primary communication channels between each multi cluster agent and the other ones n = total number of the multicluster agents the multicluster conductivity is considered as the average length of the bilateral interaction chains between the cluster organizations and businesses which are not directly interconnected but build communication channels for innovation diffusion, technology transfer, as well as for informational and sociocultural interactions between the cluster agents. the multicluster conductivity can be represented with the formula below. where: φ = multicluster conductivity xixj = communication channel between the ith and the jth multicluster agents which are not directly interconnected lxixj = aggregate length of the bilateral interaction chains between the multicluster agents building the communication channel. 3 analysis and results to assess the impact of innovation on economic development of the regions of the russian authors applied regression analysis. as a final indicator of innovative development of regional economic systems used the "share of innovative goods, works and services in their total volume". the analysis revealed that in the russian regions is no direct correlation between the share of innovative products and factors, which in accordance with the theory of cluster should define an innovative vector of territory development. the results of the correlation analysis are presented in table 1. using the territorial method to identify and assess the development of the resource-based multiclusters at the regional level allows assessing the synergetic effect produced by the interaction between primary sector and innovation businesses. a local resource sphere is created in the territory of this cluster, which is intended for providing innovation technologies, outside the region as well. advances in systems science and application(2016) vol.16 no.4 25 table 1 the results of correlation analysis of innovative development of the russian regions indicators of innovative development of the russian regions № 1.1 № 1.2 № 1.3 № 1.4 № 2.1 № 2.2 № 2.3 № 2.4 № 3 1.1. the proportion of organizations implementing technological innovation 1,00 1.2. the proportion of organizations implementing organizational innovation 0,69 1,00 1.3. the proportion of organizations implementing marketing innovation 0,35 0,48 1,00 1.4. the proportion of small innovative businesses, implementing technological innovation 0,22 0,12 0,26 1,00 2.1. number of advanced production technologies 0,27 0,27 0,26 0,21 1,00 2.2. the number of used advanced production technologies 0,28 0,28 0,29 0,25 0,74 1,00 2.3. the number of organizations engaged in research and development 0,28 0,22 0,24 0,19 0,88 0,74 1,00 2.4. the number of issued patents for inventions and utility models 0,26 0,18 0,20 0,14 0,74 0,67 0,96 1,00 3. the proportion of innovative goods, works and services in the total volume 0,00 0,03 0,07 0,05 0,06 0,15 0,03 0,01 1,00 competitive advantage of multicluster are the result of the synthesis of competitiveness factors, formed at previous stages of regional economic development. the main factors of competitiveness at different evolution stages of territorial economic systems are presented in table 2. table 2 the main factors of competitiveness at the evolution stages of territorial economic systems the evolution stages of territorial economic systems the main factors of competitiveness industrial agglomeration of industry transportation and logistics advantages, reducing uncertainty and transaction costs on the basis of geographical concentration, rapid response to competitors' innovations innovative industrial zones staffing and infrastructural benefits of innovative development, reducing uncertainty and transaction costs with the use of formal institutions and on the basis of an explicit contract with the participants of cooperation local innovation networks information benefits, reducing uncertainty and transaction costs with informal institutions, social capital formation, diffusion of management innovations multiclusters innovative advantages of joint activities in the network of scientific and technical cooperation, formation of the institutional environment of innovative development, a partnership with the government and the local community the current russian model for multicluster relations between the state, business and society is characterized under the three main segments of the institutional socio-economic interactions: «white sphere». it combines the formal institutions of the legislative and administrative regulation within the framework of multicluster, which include: registration, licensing, arbitration proceedings, auctions, etc. 26 dmitry l. napolskikh, tatyana v., yalyalieva, nina i. larionova, et al.: system of criteria and indicators…  «gray sphere». it includes informal institutions, shadow rent from business, political "bargaining" with the regional management, etc.  «black sphere». it includes informal practices of corrupt interactions, interaction with the criminal world, raider grabs businesses, etc. a variety of business and government in the formation and development of clusters is given space with two axes shown in figure 1. fig. 1. models and institutional spheres of interaction between business, government and society in the framework of multicluster it was found out in the course of the research that l resource-based multiclusters do not only broaden and develop but, in time, they can also become narrower and disintegrate. this kind of dynamics and flexibility of resource-based multiclusters represents their distinctive feature and requires permanent governmental control. in time, efficient active clusters become objects of big governmental investments. the economic institutions uniting into a multicluster on the basis of vertical and horizontal integration build a unified distribution sphere innovation technologies and the essential condition for the efficient transformation from innovations to competitive advantages is establishing a network of stable connections between all the cluster agents. formation of resource-based multiclusters at the regional level may, on the whole, be regarded as a response to excessive transaction costs [9,10]. at this point, multicluster integration processes are characterized by aggregation and consolidation of enterprises. this implicitly proves the desire for diversification and economic control of enterprise risks associated with imperfect institutional environment and excessive transaction costs. comprehensive mechanism for efficiency control of transaction costs which is used in economy to reduce them involves development of specification and protection of property rights, standardization of measurements, accounting and reporting, maintenance of monetary system, improvement of law enforcement effectiveness and efficiency, as well as implementation of measures aimed at eliminating unnecessary administrative burdens and infrastructure markets of various transactions. the basic requirements of institutional changes to improving effectiveness of natural resource-based multiclusters include recognition of the critical role of the government control of gray sphere white sphere the degree of involvement of residents in the process of territorial development the level of institutional development of the business environment contradi ctions model affiliates model policymaking model corporate model advances in systems science and application(2016) vol.16 no.4 27 economic development; government’s commitment to economic development; accounting of institutional transformation costs; review of efficiency of the current control. 4 discussion from the position of the institutional theory, the functioning sphere of the multicluster institutions represents an environment that is commonly called institutional. the study of the institutional environment of a territory as an evolving endogenous factor was initiated by the representatives of the school of economics of washington university in the 1970s [11]. douglass north uses the institutional environment term to define the institutional limits which exist at the macro level and determine the possible conditions of contractual agreements between individuals [12]. oliver williamson defines the institutional environment as an established system of the informal “rules of the game” which build the sociocultural context of economic activity [13]. in the context of this research, the institutional environment of natural resource-based multiclusters is interpreted as the aggregate of the institutional connections which surround and fill the regional economic system and produce their effect on it. on the other hand, the development degree of the governmental control mechanisms applied to the cluster policy efficiency at the regional level is determined by the activities of various groups of interests within the multicluster, while the effect they produce on the environment depends on the level of the institutional control of the interrelations inside those groups. thus, we can speak about a network aspect of functioning of the natural resource-based multiclusters which sets trajectories of interactions between economic entities. network analysis for the purposes of government control of the formation of resource-based multiclusters at the regional level allows to: · identify the influence of informal relationships between economic agents within a multicluster on competitive ability and efficiency of a multicluster as a whole. · evaluate structural consequences for the economic system of a multicluster due to change in equilibrium of institutional environment of a territorial unit. · identify the optimal organizational structure of communication channels in the resourcebased multiclusters at the regional level. 5 conclusion studying the formation of resource-based multiclusters in the modern practices of the economic development allows determining the formation trends of an efficient governmental control system at the regional level. the research shows that the formation of the governmental control mechanisms is based on a general study of the transaction costs of economic activity. the resource-based multicluster have their potential both for generating fundamental and applied knowledge and for managing innovation projects. the efficiency parameters suggested by the author for the modern cluster policy implementation for the governmental control purposes, such as density, integrity, complementarity and conductivity of multicluster will allow controlling the development efficiency of the resource-based multiclusters. the paper offers the method for calculating performance efficiency of an cluster institution for economic agents, as well as efficiency of informal institutions. the article formulates recommendations for implementing control principles to improve government regulation of multiclusters. the results proposed in the study can find application in the sphere of government control under conditions of the natural resource-driven economy. thus, the proposed efficiency indicators of multicluster effectiveness were used by government authorities of the republic of mari el (russia) while developing the programs for development of the region’s economy. the 28 dmitry l. napolskikh, tatyana v., yalyalieva, nina i. larionova, et al.: system of criteria and indicators… mechanisms for government regulation considered in this article were used by the city administration of yoshkar-ola in the process of formation of a local natural resource-based innovation cluster. it is planned to implement future research findings into the practice of public administration in russia through long-term programs of collaboration between the volga state university of technology and government agencies of the mari el republic. acknowledgment this researchers was supported by the russian foundation for basic research. project № 16-3600126 mol_a "development of mathematical methods to assess the efficiency of formation of innovative multiclasters at the sub-national level." references [1] porter, m.e., (2009). "clusters and economic policy: aligning public policy with the new economics of competition". institute for strategy and ompetitiveness. [2] enright, m., (1996). "regional clusters and economic development: a research agenda. in: business networks: prospects for regional development",walter, d.g. (ed.)., berlin, isbn10: 3110151073, pp: 190-213. [3] feldman, m.p., (1994). "the geography of innovation". springer science and business media,dordrecht, isbn-10: 0792326989, pp: 154. [4] dritsaki, c. and a. adamopoulos, (2005). "a causal relationship and macroeconomic activity: empirical results from european union". am. j. applied sci, no.2, pp.504-507. [5] liu, c.c. and c.y. chen, (2004). "a computable general equilibrium model of southern region in taiwan: the impact of the tainan science-based industrial park". am. j. applied sci. [6] kim, y.d., s. yoon and h.g. kim, (2014). "an economic perspective and policy implication for social enterprise ". am. j. applied sci., no.11, pp. 406-413. [7] larionova, n. i., napolskikh, d. l. and. yalyalieva, t.v. (2015) "institutional aspects of state control the effectiveness of innovation clusters development", actual problems of economics. no.5, pp. 20-24. [8] larionova, n. i., napolskikh, d. l. and. yalyalieva, t.v. (2014) "ensuring efficiency control of institutional environment of the cluster", american journal of applied sciences, vol.11, no.9, pp. 1594-1597. [9] larionova, n. i., napolskikh, d. l., yalyalieva, t.v., and shebashev, v.r. e. (2014) "governmental control of the formation efficiency of educational clusters at the regional level", american journal of applied sciences, vol.11, no.9, pp. 1594-1597. [10] larionova, n. i.,.napolskikh, d. l and. yalyalieva, t.v. (2015) " theoretical approaches to improving government control systems for educational clusters development", actual problems of economics. no. 4, pp. 285-288. [11] coase, r.h., (1992). "the institutional structure of production". am. economic rev., no.82, pp.713-719. [12] north, d., (1994). "economic performance through time". am. economic rev., no.84, pp. 360-361. [13] williamson, o., (2000). "the new institutional economics: taking stock, looking ahead". j. economic literature, vol.38, pp.595-613. corresponding author dmitry l. napolskikh can be contacted at: yal05@mail.ru. advances in systems science and applications (2012) vol.12 no.1 46-53 a method for determining importance degree of customer requirements in software quality function deployment lixiong gong1, mingzhong yang1, shunshenga guo and qi wang2 1school of mechanical and electrical engineering, wuhan university of technology,wuhan 430070,china 2vic, monash university, victoria 3800 ,australia abstract a method based on fuzzy analytic hierarchy process was provided to determine customer requirements in this study, and the importance degree was analyzed by using trapezoid fuzzy function. the method overcame the impact on subjective judgments & preference, and caused decision-making more reasonable. finally, the effectiveness and practicality of the method was verified by the example of fetching customer requirement in the process of software project development. keywords analytic hierarchical process, customer requirements, quality function deployment, house of quality 1 introduction software is a product that integrating knowledge and procedure. software engineers must focus on how to consider the maximum requirements of users in programming. therefore, the software must be designed to collect and analyze customer requirements, and be oriented to maximize the value of customers in the whole software development. however, what are customer requirements for a complex problem must be investigated and analyzed. data existing software programming showed that more than 50% unsuccessful projects of software development occurred in the wrong requirement analysis. quality function deployment (qfd) is a widely used customer-driven quality, design and manufacturing management tool. it is becoming a methodology of modern design theory applied for new products design and old products improvement[1-3]. customer requirements were mapped to the corresponding technical characteristics of products design that will be around customer requirements by the house of quality (hoq) in qfd[4]. but the success rate of traditional qfd is subject to restrictions because of various shortcomings, such as: fussy process, too large matrix, complicated manual calculation, unreasonable score mechanism. for the determination of the house of quality problems, the non-linear programming method that determined customer demands and technical requirements of the relationship was provided by large number of experimental data in quotation[5], and the method of multi-feature map based on taguchi theory was used to obtain the above-mentioned relationship in quotation[6]. however, advances in systems science and applications (2012), vol.12, no.1 47 these methods must rely on a large number of relevant experimental data. consequently, the experimental time and cost will be fundamental obstacles. with the development of decision-making and information science, fuzzy analytic hierarchy process (fahp) is more and more widely applied in various fields[7-8]. it decreased subjective judgmental errors account of the fuzzy factor, so the final set of index weight are more realistic. based on this, the algorithm model based on fahp was provided to determine the importance degree in this study, and applied to get customer requirements in software programming. 2 structure of software quality function deployment methodology 2.1 house of quality the core of qfd is the demand for conversion, the house of quality, that is qfd matrix, is a quality function deployment plans that associate with customer requirements and technical characteristics of products. software qfd run through the whole process of software development, which is derived from manufacture qfd and originated in japan[9]. the house of quality of software qfd is similar to the traditional qfd, as is shown in fig.1: fig.1 structure of hoq the house of quality can converted customer requirements to the quality characteristics. it is composed of six matrices, as is shown in fig.1: (1) whats matrix, said customer demand; (2) plan matrix, said the evaluation of whats matrix. (3) hows matrix, the demands for what to do; (4) the relationship matrix, the relationship between whats and the hows; (5) the matrix of hows inner relationship; (6) technology matrix, said that the evaluation of the technical cost: comparison of competitiveness or feasibility. the conversion of “what needs” to “how to do” was completed after building house of quality. 2.2 process of qfd implementation in general, the implementation of software qfd consists of two basic processes: the extraction of customer requirements and waterfall decomposition of customer requirements. information of customer requirements through face-to-face, telephone, e-mail, network, on-site investigation, after-sales service, were collected, 48 lixiong gong:a method for determining importance degree of customer requirements . . . classified, organized and analyzed to form well-organized user requirements and weighted importance degree. then, software engineers disassembled customer requirements and constructed hoq according to technical feasibility, practicability, economy and other aspects. the importance degree of qfd indexes were determined after completed above process. 3 model of determining importance degree of customer requirements in software qfd based on fahp 3.1 algorithm of fahp it is reasonable decision-making for fahp because of overcoming shortcomings of human subjective judgments, choices and preferences. this study was used trapezoidal fuzzy number to score the weight because it is more accordant with actual states and more extensive application, although trigonometric number, logarithmic trigonometric number, normal distribution function can be used while scoring. fahp algorithm steps are as follows: step 1: hierarchical structure construction. put the goal of the desired problem on the top layer of the hierarchical structure, and then put the evaluation criteria on the second layer of the hierarchical structure. further, the third layer is the sub-indexes of evaluation criteria. finally, the candidate alternatives lay in the bottom layer. hierarchical structure is shown in fig.2. fig.2 hierarchical structure step 2: constructing the fuzzy judgment matrix in this study, the fuzzy judgment matrix x is the matrix of the combination of each candidate alternative and evaluation criteria, and the fuzzy judgment matrix x is represented by trapezoid fuzzy numbers such as 1, 3, 5, 7 and 9; the advances in systems science and applications (2012), vol.12, no.1 49 values are shown in tab.1. the judgment matrix is as follows: x =  x11 x12 . . . x1n x21 x22 . . . x2n ... ... . . . ... xn1 xn2 . . . xnn  where xij = (aij , bij , cij , dij , ) , and aij , bij , cij , dij denote four values of trapezoid fuzzy numbers respectively. table 1 membership of trapezoid fuzzy numbers meaning of scale of value of relative importance relative importance fuzzy numbers equal importance 1 ( 1, 1, 1, 1 ) weak importance 3 ( 2, 2.5, 3.5, 4 ) strong importance 5 ( 4, 4.5, 5.5, 6 ) demonstrated importance 7 ( 6, 6.5, 7.5, 8 ) absolute importance 9 ( 9, 9, 9, 9 ) step 3: testing the consistency of the matrix and calculating fuzzy weights testing the consistency of fuzzy matrix after clarifying above fuzzy matrix, algorithm of testing the consistency of ahp is described in quotation[10]. return to step 2 to re-structure trapezoidal fuzzy judgment matrix and calculate it until the consistency is passed, and then, calculating fuzzy weights, fuzzy weighting formula is defined as follows: dx̄ = n∑ j=1 xij ⊗  n∑ i=1 n∑ j=1 xij −1 (1) where x̄ is judgment matrix, xij is the element of judgment matrix. step 4: single sorting weight of same hierarchy for any of two trapezoidal fuzzy numbers m and n, the possibility of m ≥ n is following results. theorem: suppose m = (r1, r2, r3, r4), n = (s1, s2, s3, s4) are two trapezoidal fuzzy numbers, then the possibility degree is v (m ≥ n) =  1 r3 ≥ s2 r4−s1 (r4−r3)+(s2−s1) r3 ≤ s2, r4 ≥ s1 0 r4 ≤ s1 (2) 50 lixiong gong:a method for determining importance degree of customer requirements . . . where r1 > 0, r2 > 0, r3 > 0, r4 > 0, s1 > 0, s2 > 0, s3 > 0, s4 > 0 approach to sorting the weights of hierarchical indexes is as follows: (1) calculating possibility degree of dx̄1 ≥ dx̄2 , . . . , dx̄1 ≥ dx̄n .calculating possibility degree of dx̄1 ≥ dx̄2 , . . . , dx̄1 ≥ dx̄n according to formula (2), that is v ( dx̄1 ≥ dx̄2 ) , v ( dx̄1 ≥ dx̄3 ) , . . . , v ( dx̄1 ≥ dx̄n ) . (2) calculating the possibility value that x̄1 is greater than other matrices. d ( x̄1 = minv ( dx̄1 ≥ dx̄2 , dx̄3 , . . . , dx̄n )) and analog,d ( x̄2 ) ,d ( x̄3 ) ,. . . d ( x̄n ) can also be calculated. (3) normalizing the matrix and getting the weight vectors ( wx̄1 ,wx̄2 . . .wx̄n ) . step 5: sorting the general hierarchy the sorting general hierarchy means that weights of all schemes are ordered in the layer of candidate alternative. the general sorting value is the weight of the scheme multiply the value of single sorting of the same hierarchy. 3.2 application in accordance with steps of fahp, calculated as follows: step 1: constructing trapezoidal fuzzy judgment matrix, testing consistency and completing the single-sort in the same hierarchy. (1) constructing trapezoidal fuzzy judgment matrix c̄ in c layer and calculating weight. the trapezoidal fuzzy judgment matrix c̄ is as follows: c̄ = [ c1 c2 ] = [ (1, 1, 1, 1) (1, 1.5, 2.5, 3) (0.333, 0.4, 0.6667, 1) (1, 1, 1, 1) ] (2) calculating synthetically fuzzy values. dc1 = 2∑ j=1 x1j ⊗  2∑ i=1 2∑ j xij  −1 = (0.3334, 0.4838, 0.8974, 1.2) dc2 = (0.2223, 0.2709, 0.4237, 0.6) (3) single sorting in c layer. single sorting in c layer using step 4 of section 3.1, then d (c1) = minv (dc1 ≥ (dc2) = 1 d (c2) = minv (dc2 ≥ (dc1) = 0.8251. the weight can be concluded while normalizing the matrix, it is shown as follows: (wc1 ,wc2) = (0.5479, 0.5421) , and then advances in systems science and applications (2012), vol.12, no.1 51 (wr1 ,wr2 ,wr3) = (0.5270, 0.4730, 0.0007) , (ws1 ,ws2) = (0.5479, 0.4521) . fig.3 hierarchical structure of software qfd development step 2: sorting the general hierarchy. vr1 = wr1 ×wc1 = 0.5479× 0.5270 = 0.2883, vr2 = 0.2592, vr3 = 0.0004, vs1 = 0.2477, vs2 = 0.2044. therefore, the general sorting and respective weight is as follows: r1(0.2883), r2(0.2592), s1(0.2477), s2(0.2044), r3(0.0004). as can be seen from the result the most important customer requirements are robust procedure & low faults, user-friendly & easy to operation and low prices. 4 conclusion it is very important to determine the importance degree, relationship degree and competitive forces in hoq. in this study, fahp method is used to determine the importance degree of customer requirements of the effective factors in the model of software qfd. humans are often uncertain in assigning the evaluation scores in conventional ahp. fahp can capture the vagueness of human thinking style 52 lixiong gong:a method for determining importance degree of customer requirements . . . and effectively solve multi-criteria decision making problems. by using fahp and appropriate calculations, there are extremely accurate for fetching customer requirements and the quality characteristics. the application shows that fahp can access to design new products, software development and evaluation of customer satisfaction that provide a new thinking for exactly getting customer requirements. acknowledgements the project is supported by the hubei international cooperation key projects, china (no. 2007ca008). references [1] hauser j r, clausing d. (1988), “the house of quality”, harvard business review, vol.66, no.3, pp.63-73. [2] sireliy y, kauffmann p, ozan e. (2007), “integration of kano’s model into qfd for multiple product design”, ieee transactions on engineering management, vol.54, no.2, pp.380-390. [3] zheng l y, chin k s. (2007), “qfd based optimal process quality planning”, international journal of advanced manufacturing technology, vol.26, no.7, pp.831-841. [4] karsak e e. (2004), “fuzzy multiple objective programming framework to prioritize design requirements in quality function deployment”, computers & industrial engineering, vol.47, no.2-3, pp.149-163. [5] dawson d, asking g. (1999), “optimal new product design using quality function deployment with empirical value functions”, quality and reliability engineering international. vol.15, no.1, pp.17-32. [6] kumar p, barua p b, gaindhar j l. (2000), “quality optimization (multicharacteristics) through taguchis technique and utility concept”, quality and reliability engineering international, vol.16, no.2, pp.475-485. [7] han shilian, li xuhong, liu xinwang. (2004), “stmulti-person and multicriteria evaluation and selection of logistic centers with fuzzy analytic hierarchical process method”, system engineering theory & practice,vol.24, no.7, pp.128-134. [8] hu yaoguang, fan yushun. (2006), “decision model based on fahp for selection of enterprise core business systems”, computer integrated manufacturing systems, vol.12, no.2, pp.215-219. advances in systems science and applications (2012), vol.12, no.1 53 [9] haag s. (1996), “quality function deployment: usage in software development”, in communications of the acm, vol.39, no.1, pp.41-49. [10] sang song, lin yan, ji zhuoshang. (2002), “an improved ahp method for mcdm in ship type’s demonstration”, journal of dalian university of technology, vol.42, no.2, pp.204-207. advances in systems science and application (2015) vol.15 no.3 254-266 optimal estimation and precision analysis of measuring data fusion model xuanying zhou, jiongqi wang, zhengming wang and zhangming he college of science, national university of defense technology, changsha, hunan, china. abstract data fusion is an effective method to improve the data processing accuracy, and the fusion weight has a great influence to the accuracy of the fusion estimation. in this paper, we study the problem of optimal weight and parameter estimation for the linear fusion model with unequal precision measuring data. the properties of the unbiased estimation are discussed for linear model, and then the optimal estimation for the fusion model with unequal precision measuring data is given. it is proved that there exists the optimal fusion weight and it is unique. besides, the accuracy of the optimal estimation for the multivariate linear model is analyzed and some conclusions suitable for practical application are obtained, which can provide the theory foundation for the experiment design and the data selection. finally, two simulations are offered to validate the theories conclusions in this paper. keywords data fusion; regression model; fusion weight; parameter estimation; precision analysis. 1 introduction the main purpose of data fusion is to improve the accuracy of measuring data, to establish a proper processing model, and to give an effective and reliable fusion algorithm [1, 2]. fusion with different types and unequal precision data is the most typical situation in the data fusion processing [3-5]. after the data are modeled in a parametric model, the data fusion problem can be transferred into the parameter estimation problem [6,7]. to evaluate the performance of the parameter estimation result, we need an evaluation standard, i.e., evaluation criterion or optimal criterion. such criteria include minimum mean-square error (mse) criterion, maximum likelihood (ml) criterion, maximum a posteriori (map) criterion, best linear unbiased estimation (blue) criterion and least squares estimation (lse) criterion, etc. whether the estimated parameter satisfies the need of the application depends on the estimation criteria as well as the data accuracy and the model properties [8]. obviously, the selection of the evaluation criterion is affected by the characteristic of the estimation parameter, the demand of the estimation accuracy and the complexity of the estimation algorithm. specifically, the parameters, estimated according to lse criterion, will lead to the minimal norm of the obseradvances in systems science and application (2015) vol.15 no.3 255 vation residual, i.e. the differencebetween the observed value and the calculated value. usually, lse does not involve the dynamic and statistical information of the parameter to be estimated. therefore, lse is easy implemented but with low estimation accuracy. nevertheless, when we are short of the error information about the measuring data, lse also can provide us with an acceptable solution. mse criterion is the best in terms of that mse has the minimal mean square error. however, this method needs some statistical prior information, such as the first and the second moment of the data and parameter. the map and the ml estimation are both related to the conditional probability density functions, and the estimation is hard to be obtained except for some special cases. therefore, in the actual application, the efficient and reliable data fusion algorithm for data fusion should be selected according to the specific situation. although most of the fusion systems are nonlinear, they can be linearized into some linear regression model when proper base functions are selected or the nonlinear iterative means are adopted. that is to say, nonlinear fusion problem can be approximated to process with the linear fusion problem. furthermore, when the parameters and the measuring data are with the normal distribution, some optimal estimation methods, such as blue, mse and lse, are equivalence to each other [10]. therefore, the minimal linear variance criterion is usually applied for the actual application. following the introduction in section 1, the structure of the paper is organized as follows. the form of unbiased estimation for linear model is given and the estimation characters are discussed in section 2. in section 3, the optimal weight and parameter estimation of unequal-precision data fusion are researched. besides, it is proved that there exists the optimal fusion weight and it is unique, and the accuracy of multivariable optimal estimation for linear fusion model is analyzed. section 4 provides two numerical examples to validate the proposed theory and method. finally, the paper is concluded in section 5. 2 unbiased estimate of linear model consider the linear measuring regression model as follow: y = xβ + ε, ε ∼ (0, σ2i) (1) where y = ym×1 is the measuring data, x = xm×n is the design matrix and rank(x) = n, β = βn×1 is the estimated parameter vector, and ε = εm×1 is the measuring random error vector with the zero expectation and diagonal covariance, i.e. ε ∼ (0, σ2i). for the parameter estimation problem in model (1), the lse β̂ls = ( xtx )−1 xty has some good properties as follows: property 1: β̂ls is the umvue (uniform minimum variance unbiased esti256 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... mation) of the parameter β, and moreover, ∀c ∈ rn, ctβ̂ls is the linear unbiased estimation of the parameter ctβ. property 2: assuming e ∥∥∥xβ̂ls −xβ ∥∥∥2 = (xβ̂ls −xβ)t(xβ̂ls−xβ) and then e ∥∥∥xβ̂ls −xβ ∥∥∥2 = nσ2. property 3: assuming cov(β̂ls) = e(β̂ls − β)(β̂ls − β)t, mse(β̂ls) = e(β̂ls − β)t(β̂ls − β) and then cov(β̂ls) = σ2s−1,mse(β̂ls) = σ2 n∑ i=1 λ−1 i , where s = xtx and λi, i = 1, · · · , n are all eigenvalues of the matrix s; remark 1: property 1 shows that in the actual engineering application, different constant vector c can be chosen for estimating some components or the linear combination of the parameter β. remark 2: property 2 shows that the estimation precision of xβ̂ls is directly proportional to n. that is to say, the more the number of the estimated parameters, the lower estimation precision will be obtained. therefore, the proper base function and parameter model should be chosen when modeling the measuring data to let the parameter number as few as possible. the sparse parameter modeling methods are often used in the real application [11, 12]. remark 3:property 3 shows that when the model (1) is multi-collinearity, i.e. the matrix xtx has some extremely small eigenvalues λi, the large mse(β̂ls) or mse(ctβ̂ls) will lead to the bad estimation accuracy. the regularizing methods, a series of biased estimation methods, were proposed to handle the multi-collinear problem in application [13-15]. by choosing the proper regularizing factor µ and the regularizing matrix d with full column rank to solve the following optimization problem min β∈rn,µ>0 ∥y −xβ∥22 + µ ∥dβ∥22 (2) the solution of (2) can be easily calculated as follows β̂r = (xtx + µdtd)−1xty (3) compared to lse, the regularizing estimation has properties as follows: property 4: eβ̂r = (xtx + µdtd)−1xtxβ. i.e. the regularizing parameter estimation is biased; property 5: there exists µ,d, and make mse(β̂r) < mse(β̂ls), i.e. the regularizing estimation can better than lse by choosing some proper regularization parameter and matrix. in the linear measuring regression model, ε ∼ (0, σ2i) means the measures are irrelevant and the precision are equal. in actual, if the measures are relevant and have unequal precision, the model (1) can be transferred to y = xβ + ε, ε ∼ (0, σ2g) (4) advances in systems science and application (2015) vol.15 no.3 257 where σ2 is known or unknown, and g is a known positive definite matrix. for model (4), its umvue is the weighted least squares estimation (wlse)β̃wls =( xtg−1x )−1 xtg−1y , and mse(β̃wls) = tr ( xtg−1x )−1 . in real application, in order to get the lse for the linear fusion model, the measuring data should be parametric modeling to make it satisfy the model (1) or (4). actually, a typical application of model (4) is the unequal precision data fusion processing problem with several kinds of measuring equipment. although the measuring equations of different equipment are non-linear, the proper basis function can be chosen or the nonlinear iterative means can be adopted to linearize the measuring equations. therefore, the lse or wlse can be an important theory foundation for the measuring data fusion. no matter the parameters estimated by model (1) or (4), the statistic properties of the measuring random error, including the correlation, meaning, variance, and covariance and so on, need to be estimated at first. σ2 in the model (1) or (4) reflects the accuracy of the measuring data. therefore, as the base of unequal precision data fusion, the estimation of the parameter σ2 is very important. besides, when the estimation performance of the lse (wlse) is worse, the information of the parameter σ2 is also needed to be used in order to build the biased estimation of the parameter β. there is the property about the estimation of σ2: property 6:assume the observation error in model (1) satisfy the normal distribution, i.e. ε ∼ n(0, σ2i), then σ̂2 = rss/(m− n),where rss = m∑ i=1 µ2 i =∥∥∥y −xβ̂ls ∥∥∥2 is the measuring residual square sum, µi = yi−xiβ̂ls , i = 1, · · · ,m is the ith residual between of the measuring data and the calculated value and eσ̂2 = σ2, mse(σ̂2) = 2σ4 / (m− n). remark 4:property 6 shows that the parameters σ2 can be estimated by the residual if the precision of the actual measuring data is unknown. the estimation value and its accuracy of the parameter σ2 are related to the number of the measuring data as well as the dimensional of the estimated parameter. as a result, in the actual application, the estimation variance can be decreased by increasing the sampling number of the measuring data. 3 optimal fusion estimation of linear model with unequal precision data in many measuring processing problem, like trajectory tracking, the unequal precision data fusion processing often need to be considered. obviously, the weighting methods for different measuring data have a great influence to the accuracy of the fusion estimation. 258 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... 3.1 optimal estimation in unequal precision linear fusion model considering s kinds of unequal precision linear data fusion model: y1 = x1β + ε1, ε1 ∼ (0, σ2 1im1), · · · · · · · · · ys = xsβ + εs, εs ∼ (0, σ2 s ims), eεiεj t = omi×mj , i, j = 1, · · · , s, i ̸= j (5) the definitions of the parameter in (5) are same to that in model (1), and the optimal estimation can be given by the follow theorem. theorem 1: for model (5), ∀c ∈ rn, the uniformly minimum variance estimation of ctβ is ctβ̃f where β̃f = ( s∑ i=1 σ−2 i s∑ i=1 σ−2 i xt i xi) −1( s∑ i=1 σ−2 i s∑ i=1 σ−2 i xt i yi) (6) proof : assuming t = 1 /√ s∑ i=1 σ−2 i , then model (6) can be rewritten as:  tσ−1 1 y1 = tσ−1 1 x1β + tσ−1 1 ε1, tσ −1 1 ε1 ∼ (0, t2im1), · · · · · · · · · tσ−1 s ys = tσ−1 s xsβ + tσ−1 s εs, tσ −1 s εs ∼ (0, t2ims), (7) assuming y = [tσ−1 1 y t 1 , · · · , tσ−1 s y t s ]t, x = [tσ−1 1 xt 1 , · · · , tσ−1 s xt s ] t, ε = [tσ−1 1 εt1 , · · · , tσ−1 s εts ] t, combining with (6) and (7), the fusion model (5) can be written as follows: y = xβ + ε, ε ∼ (0, t2im),m = s∑ i=1 mi (8) using the lse, the theorem can be proved that ∀c ∈ rn, the uniformly minimum variance estimation of ctβ is ctβ̃f , where β̃f = (xtx)−1xty = ( s∑ i=1 t2σ−2 i xt i xi) −1( s∑ i=1 t2σ−2 i xt i yi) (9) the proof is completed. 3.2 optimal weight for the linear fusion model the purpose of data fusion is to find the optimal weight ρi and then to optimize the fusion problem and obtain the optimal parameter estimation. theorem 1 advances in systems science and application (2015) vol.15 no.3 259 above shows that the optimal weight of the data fusion model with unequal precision measuring data is ρi = σ−1 i /√ s∑ i=1 σ−2 i = tσ−1 i , and moreover, it satisfies s∑ i=1 ρ2i = 1. this indicates that the optimal only related to the data accuracy σ−1 i . for convenient, the weight method for two types of unequal-precision linear observed data is discussed firstly: y1 = x1β + ε1, ε1 ∼ (0, σ2 1im1), y2 = x2β + ε2, ε2 ∼ (0, σ2 2im2), eε1ε t 2 = om1×m2 (10) considering the following optimization problem: arg β min 2∑ i=1 ρ2i ∥yi −xiβ∥2 2∑ i=1 ρ2i = 1 (11) and the solution can be easily obtained as follow: β̂(ρ) = ( 2∑ i=1 ρ2ix t i xi) −1( 2∑ i=1 ρ2ix t i yi) (12) theorem 2: under the assumption of model (10), the solution of the optimization problem arg ρ mine ∥∥∥β̂(ρ)− β ∥∥∥2 = arg ρ minmse(β̂(ρ)) is: ρi = σ−1 i /√ σ−2 1 + σ−2 2 , i = 1, 2 (13) proof : by calculating e ∥∥∥β̂(ρ)− β ∥∥∥2 = tr( 2∑ i=1 ρ2ix t i xi) −2( 2∑ i=1 ρ4iσ 2 ix t i xi) assuming a = xt 1 x1, b=xt 2 x2, and then f(ρ1, ρ2) = (ρ21a+ ρ22b)−1(ρ41σ 2 1a+ ρ42σ 2 2b)(ρ21a+ ρ22b)−1 = σ2 2 ( ρ21 ρ22 a+b )−1 (ρ41 ρ42 σ2 1 σ2 2 a+b )( ρ21 ρ22 a+b )−1 (14) as both of a and b are real symmetric positive definite matrices, and can be similarity diagonalized simultaneously, that is to say existing an invertible matrix 260 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... p , let a = ptλp, b=ptp , where λ = diag(λ1, · · · , λn), therefore f(ρ1, ρ2) = σ2 2p t ( ρ21 ρ22 λ + i )−1 (ρ41 ρ42 σ2 1 σ2 2 λ + i )( ρ21 ρ22 λ + i )−1 (p−1)t = σ2 2p tdiag [( ρ41 ρ42 σ2 1 σ2 2 λi + 1 )( ρ21 ρ22 λi + 1 )−2 ] (p−1)t (15) assuming that t = ρ21 ρ22 , σ = σ1 σ2 , λi = a, g(t) = (t2σ2a+ 1) / (ta+ 1)2, let dg(t) dt = 2a(tσ2−1) (ta+1)3 = 0 and then t = σ−2. besides, d2g(t) dt2 |t=σ−2 = 2a(σ2 + a) (ta+ 1)4 > 0 (16) noticed that ρ21 + ρ22 = 1, and the function g(t) gets the minimum value at ρi = σ−1 i /√ σ−2 1 + σ−2 2 , i = 1, 2, then e ∥∥∥β̂(ρ)− β ∥∥∥2 = tr( 2∑ i=1 ρ2ix t i xi) −2( 2∑ i=1 ρ4iσ 2 ix t i xi) = tr [ σ2 2p tdiag [( ρ41 ρ42 σ2 1 σ2 2 λi + 1 )( ρ21 ρ22 λi + 1 )−2 ] (p−1) t ] = σ2 2 n∑ i=1 t2σ2λi+1 (tλi+1)2 (17) since the each item in equation (17) gets the minimum at ρi = σ−1 i /√ σ−2 1 + σ−2 2 , i = 1, 2, therefore, the solution of arg ρ mine ∥∥∥β̂(ρ)− β ∥∥∥2 is ρi = σ−1 i /√ σ−2 1 + σ−2 2 , i = 1, 2. the proof is complete. remark 5: the conclusion of theorem 2 has important application value. in the actual problem, unequal precision data are usually measured, so the data weight has an important effect to the data fusion accuracy. theorem 2 shows that the unique optimal fusion weigh depends on the data precision in the linear fusion model with unequal precision measuring data. actually, this is the gauss-markov theorem applies to the lse for the linear model. however, the gauss-markov theorem shows the optimal estimation only can be given when the data accuracy is known, and while theorem 2 shows the optimal estimation can be obtained by solving the optimal problem arg ρ mine ∥∥∥β̂(ρ)− β ∥∥∥2 when the data precision is unknown. that is to say, for model (10), ρi = σ−1 i /√ σ−2 1 + σ−2 2 , i = 1, 2 is the necessary and sufficient condition if β̂(ρ) = ( 2∑ i=1 ρ2ix t i xi) −1( 2∑ i=1 ρ2ix t i yi) is the advances in systems science and application (2015) vol.15 no.3 261 uniformly minimum variance solution (optimal solution) for the parameter β. from theorem 2, for the linear fusion model (5), assuming ρ = [ρ1, · · · , ρs], the parameter estimation and optimal weight determination can be handled by the following two-step minimal problem (1)  arg β min s∑ i=1 ρ2i ∥yi −xiβ∥2 s∑ i=1 ρ2i = 1 (18) (2) arg ρi mine ∥∥∥β̃(ρ)− β ∥∥∥2 (19) 3.3 accuracy analysis of optimal fusion estimation for convenient, consider the parameter estimation accuracy problem of two kinds of unequal precision data fusion model. suppose that β̂(i), i = 1, 2 are the estimation by measuring yi, i = 1, 2, separately, β̂(1, 2) is the traditional joint estimation of these two kinds of measuring data, and β̂f is the optimal fusion estimation, i.e. β̂(1) = (xt 1 x1) −1xt 1 y1 β̂(2) = (xt 2 x2) −1xt 2 y2 β̂(1, 2) = (xt 1 x1 +xt 2 x2) −1(xt 1 y1 +xt 2 y2) β̂f = (σ−2 1 xt 1 x1 + σ−2 2 xt 2 x2) −1(σ−2 1 xt 1 y1 + σ−2 2 xt 2 y2) (20) then the follow conclusions can be drawn. theorem 3: for the different estimation for the parameter β there are: (1)e ∥∥∥β̂f − β ∥∥∥2 ≤ min{e ∥∥∥β̂(1)− β ∥∥∥2, e ∥∥∥β̂(2)− β ∥∥∥2, e ∥∥∥β̂(1, 2)− β ∥∥∥2} (21) (2) e ∥∥∥β̂(1, 2)− β ∥∥∥2 < max{e ∥∥∥β̂(1)− β ∥∥∥2, e ∥∥∥β̂(2)− β ∥∥∥2} (22) (3)if σ2 2/σ 2 1 ≤ 2, and then e ∥∥∥β̂(1, 2)− β ∥∥∥2 ≤ min{e ∥∥∥β̂(1)− β ∥∥∥2, e ∥∥∥β̂(2)− β ∥∥∥2} (23) proof : (1) from the lse properties: e ∥∥∥β̂(1)− β ∥∥∥2=σ2 1tr(x t 1 x1) −1,e ∥∥∥β̂(2)− β ∥∥∥2 = σ2 2tr(x t 2 x2) −1 e ∥∥∥β̂(1, 2)− β ∥∥∥2= tr(σ2 1x t 1 x1+σ2 2x t 2 x2)(x t 1 x1+xt 2 x2) −2 e ∥∥∥β̂f − β ∥∥∥2 = tr(σ−2 1 xt 1 x1+σ−2 2 xt 2 x2) −1 (24) 262 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... obviously, e ∥∥∥β̂f − β ∥∥∥2 ≤ min{e ∥∥∥β̂(1)− β ∥∥∥2,e∥∥∥β̂(2)− β ∥∥∥2} and furthermore, from theorem 2, e ∥∥∥β̂f − β ∥∥∥2 ≤ e ∥∥∥β̂(1, 2)− β ∥∥∥2. (2) assuming e ∥∥∥β̂(1)− β ∥∥∥2 ≤ e ∥∥∥β̂(2)− β ∥∥∥2 and a = xt 1 x1, b = xt 2 x2, e ∥∥∥β̂(1, 2)− β ∥∥∥2 ≤ e ∥∥∥β̂(2)− β ∥∥∥2 need to be proved, the follow inequality need to be proved first: tr(σ2 1a+σ2 2b)(a+b)−2 ≤ σ2 2tr(b −1) (25) that is tr[(σ2 1a+σ2 2b)− (a+b)σ2 2b −1(a+b)] ≤ 0 (26) noticed that σ2 1a+σ2 2b−(a+b)σ2 2b −1(a+b) =a(σ2 1a −1−2σ2 2a −1−σ2 2b −1)a < 0. then equation (26) is right, and thus equation (22) is proved. (3) following the symbol and assumption in (2), if want to prove e ∥∥∥β̂(1, 2)− β ∥∥∥2 ≤ e ∥∥∥β̂(1)− β ∥∥∥2, i.e. to prove tr(σ2 1a+σ2 2b)(a+b)−2 ≤ σ2 1tr(a −1) that is tr[(σ2 1a+σ2 2b)− (a+b)σ2 1a −1(a+b)] ≤ 0 (27) noticed that when σ2 2/σ 2 1 ≤ 2, and then σ2 1a+σ2 2b − (a+b)σ2 1a −1(a+b) =(σ2 2 − 2σ2 1)b −1 − σ2 1ba−1b < 0, then equation (27) as well as equation (23) is proved. the proof is complete. obviously, the conclusion in theorem 3 can also be adapted to s kinds of unequal precision data fusion processing problem. it has the great effect to experiment designed and data selection scheme optimization problem in the actual application. equation (21) shows that the accuracy of several sensors optimal fusion estimation is the best comparing to the any single or any combination sensors joint estimation. and equation (22) shows that the precision of several sensors joint (traditional weighted scheme) estimation is better than the worst single sensors estimation, but the estimation precision of several sensors joint can better than the best single sensors if each sensors measuring accuracy reaches some certain conditions. 4 numerical examples in this section, two calculation examples are given to validate the proposed theory and algorithm for optimal weight and parameter estimation of unequal-precision data fusion. example 1: fusion processing of static measuring data advances in systems science and application (2015) vol.15 no.3 263 assuming two unequal-precision equipment measure the physical signal β. suppose the real value of the signal β is 10, and randomly create 100 highprecision measuring data (the root mean square error, rms, is 3), 100 mediumprecision measuring data (the rms is 4) and 100 low-precision measuring data (the rms is 6), and simulate 100 times. the estimation of β and its variance is get. seen in the table 1 as follow (the root variance is come from the 100 simulate data statistic) the true value of the physical quantity is β = 10. we have 100 groups of data. in each group, there are 100 high-precision data samples (the standard deviation of is 3), 100 medium-precision data samples (the standard deviation of is 4) and 100 low-precision data samples (the standard deviation of is 6). the parameter estimated and its mse is shown in table 1 below. (the estimated variance is obtained based on statistics of 100 groups of observed data.) table 1 parameter estimate result in different weighted methods hhhhhh result method highprecision only mediumprecision only lowprecision only traditional weighted jointestimation with highand medium-precision data traditional weighted joint estimation with highand low-precision data optimal fusion estimation truth parameter 10 10 10 10 10 10 estimated value 9.975 10.087 10.114 9.983 10.052 9.995 mean square error 0.087 0.131 0.154 0.071 0.116 0.045 in the linear measuring data fusion processing, the unique fusion weight depends on the measuring data accuracy. solve the minimum optimization problem (18) and (19), and get the mse of the estimated parameter is smallest, 0.045. for the highand medium-precision data fusion, the precision of two data satisfies σ2 2/σ 2 1 < 2, therefore, the mse of the parameter with the traditional joint weighted method is 0.071, which is better than that only with the high-precision data, 0.087. for the highand low-precision data fusion, the precision of two data does not dissatisfy σ2 3/σ 2 1 < 2, the mse of the estimated parameter with the traditional weight joint method is 0.116, which is worse than that only with the high-precision data, 0.087, while better than the mse, 0.154, which only use the low-precision data. example 2: fusion processing of dynamic tracking data assuming that gps and bds are tracking and measuring a dynamic target, simultaneity, and the measuring data are the single point positioning data. suppose (x(t), y(t), z(t))t is the position of the target orbit at time t, the positioning accuracy of gps in every direction is 1m, and that of bds is 3m. simulate 80 groups of measuring data by the theoretical orbit, including t = 0.05× j, j = 1, · · · , 600 gps and bds positioning data in each group. in the tracking period, the orbit data is model by the cubic spline function of the optimal node according 264 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... to the reference [6]. the spline coefficient is estimated first and then the orbit (x̂(k)(tj), ŷ (k)(tj), ẑ (k)(tj),̂ ẋ (k)(tj),̂ ẏ (k)(tj),̂ ż (k)(tj))(ρ), k = 1, 2, · · · , 80, j = 1, 2, · · · 600 can be calculated sequentially. assume that, mser(ρ21, ρ 2 2) = 1 80 1 600 80∑ k=1 600∑ j=1 √ (x(tj)− x̂(k)(tj)) 2 + (y(tj)− ŷ(k)(tj)) 2 + (z(tj)− ẑ(k)(tj)) 2 (ρ21, ρ 2 2) (28) where (x(tj), y(tj), z(tj)) t is the theoretical orbit and table 2 gives the estimation precision for the orbit position with different weighted methods. table 2 parameter estimate result in different weighted methods xxxxxxxxxxx method estimation result weighted only use gps data only use bds data traditional weighted estimation optimal fusion estimation weighted method (1, 0) (0, 1) (1/2, 1/2) (9/10, 1/10) mse 1.512 3.247 1.243 0.716 5 conclusion data fusion is one effective method to improve the precision of data processing. this paper researches the optimal weight and parameter estimation of unequal precision linear data fusion. for the linear fusion model, the optimal weight depends on the precision of the measuring data only, and it is consistent with the classical gauss-markov theory. furthermore, when the data precision in the fusion model is unknown, the parameter estimation and the optimal weight can be obtained by minimizing the mean square error. the parameter estimation precision of multivariate linear fusion model is given in this paper as well as some conclusions, which can be used in the practical engineering application. the accuracy of the optimal fusion estimation with multi-measuring data is better than that of only using any single measuring data and the traditional weighted estimation. the estimation precision with the traditional weighted method is better than that only with the low-precision data, and if the data precision of each measuring data satisfies some certain condition, the precision with the traditional weighted method can better than that only with the high-precision data. and these conclusions can be the support to experiment design and data selected scheme. note that the optimal fusion weight of the unequal precision discussed in this paper is the linear fusion model. the nonlinear model needs to be considered further, although the theory and method in this paper is the basic of nonlinear advances in systems science and application (2015) vol.15 no.3 265 problem. besides, the optimal fusion weight of the multi-structure linear regression model is obtained under certain evaluation criterion. different evaluation criterion leads to different optimal fusion weight. the evaluation criterion corresponding to the minimal mse of parameter estimation is used in the paper, i.e. arg ρ mine ∥∥∥β̂(ρ)− β ∥∥∥2. nevertheless, this criterion has some limitations, e.g. the contribution of each parameter to the problem is not distinguished. certainly, other evaluation criteria should be considered for specific issues, which will be studied in the future. references [1] hall d l.(1992). mathematical techniques in multi-sensor data fusion, artech house, boston, london. [2] jacqueline le moigne and james smith. (2000). “image registration and fusion in remote sensing for nasa”, proceedings of 2000 international conference on information fusion, paris, france. [3] y. bar-shalom, h. chen and m. mallick. (2004). “one-step solution for the multi-step out-of-sequence-measurement problem in tracking”, ieee transactions on aerospace and electronics systems. [4] x. r li and vesselin p. (2003), “a survey of maneuvering target trackingpart v: multiple-model methods ”, proceeding of spie conference on signal and data proceeding of small targets, san diego, ca, usa. [5] prieto, j., mazuelas, s., bahillo, a., fernandez, p., lorenzo, r. m., and abril, e. j. (2012), “adaptive data fusion for wireless localization in harsh environments”, signal processing, ieee transactions. [6] z. m. wang and d. y. yi. (2011), measurement data modeling and parameter estimation, crc press, china. [7] y. barshalom, x. r. li and t. kirubarajan. (2001), esitmationwith applications to tracking and navigation, wiley, new york. [8] simon d. (2006). optimal state estimation, john wiley & sons, inc. new york. [9] haiyin zhou, jiongqi wang and xiaogang pan. (2013), fusion theory and methods for satellite state estimation, science press, china. 266 xuanying zhou, jiongqi wang and zhengming wang, zhangming he: optimal... [10] jiongqi wang, haiyin zhou and yi wu. (2007), “the theory of data fusion based on state optimal estimation”, acta mathematicae applicatae sinica, china. [11] cetin m., malioutov d. m. and willsky a. s. (2002), “a variational technique for source localization based on a sparse signal reconstruction perspective”, proceedings of the 2002 ieee international conference on acoustics, speech, and signal processing, orlando. [12] tenorio l. (2001), “statistical regularization of inverse problems”, siam review. [13] tikhonov, a. n., and arsenin, v. y. (1977), solutions of ill-posed problems, john wiley & sons, new york. [14] voutilainen, a., stratmann, f., and kaipio, j. p. (2000), “a nonhomogeneous regularization method for the estimation of narrow aerosol size distributions”, journal of aerosol science. [15] aster, r. c., borchers, b. and thurber, c. h. (2013), parameter estimation and inverse problems , academic press. corresponding author xuanying zhou can be contacted at: julia chow07@163.com advances in systems science and application (2015) vol.15 no.3 202-219 a novel multi-attribute decision making methodology and application yong liu1 and yi lin 2 1 school of business, jiangnan university, wuxi, 214122, china; 2 mathematics department, slippery rock university of usa, pennsylvania, usa. abstract in the multi-attribute decision making problems, how to effectively extract the decision making rules and rank the schemes are much more important research contents. however, the acquisitions of the decision making rules often are ignored. in view of this, with respect to the problems that there exist a lot of preference information and fuzzy information in the real decision making information system, a novel decision making methodology based on dominance intuitionistic fuzzy rough set is constructed in the paper, and then it is applied to audit risk assessment and risk judgment. based on the analysis of model and example, the result shows that the proposed model can well realize the extraction of the decision making rules and the ranking of the schemes, and effectively deal with the intuitionistic fuzzy information system with preference information. keywords preference information; decision making rule; dominance distance index 1 introduction as a useful mathematical tool to deal with knowledge with inaccuracy, uncertainty and fuzziness, rough set theory was initially proposed by pawlak [1]. it has been widely applied in many fields, such as knowledge discovery, data mining, decision analysis, and pattern recognition[2–4]. the classical rough set theory conducts data reasoning on the basis of equivalence relationship, while it is difficult to satisfy the harsh conditions of equivalence relationship in the practical applications. at the same time, the binary relationship existing on the field of discourse is often a fuzzy relationship and a similarity relationship instead of an equivalence relationship. in view of this, based on the idea and method of the fuzzy set theory[5] put forward the fuzzy rough sets theory. because this new theory can well describe the uncertainty of various types of knowledge and more objectively reflect the physical world, it has been rapidly becoming a research focus of rough set theory, leading to its rapid development. as a result of simultaneously considering the positive, negative and hesitancy degrees for an object to belong to a set, intuitionistic fuzzy sets possess stronger ability of information expression and well describe and portray delicate ambiguities of the nature of the objective world when compared with the traditional fuzzy sets[6, 7]. therefore, intuitionistic fuzzy sets and rough sets are first proadvances in systems science and application (2015) vol.15 no.3 203 posed to hybrid, leading to the construction of the intuitionistic fuzzy rough set model [8]. due to the important theoretical value and application implications, intuitionistic fuzzy rough set theory has soon become a hot academic research area. currently, most of the related research on intuitionistic fuzzy rough set lies in the aspects of constructing different models and exploring their relevant properties. for the related researches on constructing different models and exploring their properties, the relationship between intuitionistic fuzzy set theory and rough set theory firstly is revealed, and then they employ intuitionistic fuzzy set to define approximation operators in the intuitionistic fuzzy approximation space. by making use of the cut set of intuitionistic fuzzy sets, the upper and lower approximation operators of intuitionistic fuzzy rough set and the axiomatic method of the approximation operators based on general binary intuitionistic fuzzy relationship are respectively constructed [8–10], and then it is well known that the upper and lower approximation sets of intuitionistic fuzzy rough sets are intuitionistic fuzzy sets by making the proof [11]. based on intuitionistic fuzzy residual implication and intuitionistic fuzzy relationship, an intuitionistic fuzzy rough set model is established[12]. however, it is difficult to apply this model to deal with an information system with noise data. with the intuitionistic fuzzy triangle model t = min, intuitionistic fuzzy t-conorms s= max, and intuitionistic fuzzy inverse operator n, the approximation operators of the intuitionistic fuzzy rough sets is defined, and the intuitionistic fuzzy rough set models based on the general intuitionistic fuzzy logic operators are developed[13–15]. by using the thought of intuitionistic fuzzy set and rough set, an improved intuitionistic fuzzy rough set model based on hamming distance and establish the models such properties as interval, symmetry, complete similarity and complete dissimilarity are proposed[16], while the novel intuitionistic fuzzy rough set based on general intuitionistic fuzzy information systems is constructed in order to expand the model and its application[17]. the interval-valued intuitionistic fuzzy rough set based on the thought of implication is established, and then the related properties of the models are developed[18, 19]. by combining interval intuitionistic fuzzy set and rough set, the interval intuitionistic fuzzy rough set models based on interval intuitionistic fuzzy relationship are constructed[19–21]. by using interval-valued intuitionistic fuzzy compatibility relationship, the interval-valued intuitionistic fuzzy rough set model based on the concept of double universes and relevant properties is constructed, and then it is applied into the decision making[22, 23]. for the related literature on attribute reduction, a genetic algorithm is proposed to reduce attributes by making use of the characteristics of intuitionistic fuzzy information systems[24], while the kind of attribute reduction algorithm of intuitionistic fuzzy rough set based on mutual information by combining in204 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... formation entropy theory and intuitionistic fuzzy rough set is constructed and designed[25]. by using the intuitionistic fuzzy distance formula, an intuitionistic fuzzy rough model and an attribute reduction algorithm is constructed[26–28]. as discussed above, although there exist many related literatures on the intuitionistic fuzzy rough set model to construct models and discuss their theories from different levels, there are few intuitionistic fuzzy rough set models in the existing literatures those can be applied into the multi-attribute decision making and effectively deal with the real problems. in the multi-attribute group decision making problems, the dominated and dominating relationship for the attributes should be considered, and the subjective preferences of attribute value from decision-makers are based on some methods to aggregate, so that the acquisition of the decision rules and the ranking of the schemes can be made. however, it is difficult for the group decision making methods based on rough set[29–38] to deal with multi-attribute decision making problems with noise data, preference information and fuzzy information. in view of this, with respect to the shortcoming of the size comparison of the intuitionistic fuzzy numbers, the concept of the intuitionistic fuzzy dominance distance index is defined, and then it is used to construct a novel intuitionistic fuzzy rough set model, finally an example illustrates the effectiveness and applicability of the proposed model. 2 dominance distance index intuitionistic fuzzy sets proposed by atanassov[6, 7] are an expansion and further development of the traditional fuzzy sets. as a result of taking into account to the membership and non-membership information by adding a new attribute parameter: non-membership function in intuitionistic fuzzy sets, it provides additional options for describing the properties of things and possesses a stronger capability of dealing with uncertainty information. that is, intuitionistic fuzzy sets can well describe and portray delicate ambiguities of the nature of the objective world. 2.1 intuitionistic fuzzy dominance distance index definition 1. (intuitionistic fuzzy sets[6, 7]. suppose thatx = {x1, x2, ..., xn} is a nonempty, finite set of objects with xi(i = 1, 2, ..., n) being the ith object. then the set a = {< x, µa(x), υa(x) > |x ∈ x} of triplets is called an intuitionistic fuzzy set, where µa(x) and υa(x) are respectively known as the membership and non-membership for the object to belong to , that is, µa(x) : x → [0, 1], x ∈ x → µa(x) ∈ [0, 1] (1) υa(x) : x → [0, 1], x ∈ x → υa(x) ∈ [0, 1] (2) satisfying 0 ≤ µa(x) + υa(x) ≤ 1 , for any x ∈ x. and πa(x) = 1 − µa(x) − υa(x), x ∈ x stands for the degree of hesitation or uncertainty for the object to advances in systems science and application (2015) vol.15 no.3 205 belong to . so the intuitionistic fuzzy number is denoted as α =< µa(x), υa(x) >. definition 2. (intuitionistic fuzzy sets[6, 7]. for any intuitionistic fuzzy number α =< µ, ν >, the score function s(α) of this number is defined as follows: s(α) = µ− ν, s(α) ∈ [−1, 1] (3) the larger s(α) is, the greater the intuitionistic fuzzy number α =< µ, ν > is. for example, assume that both intuitionistic fuzzy numbers are respectively α1 =< 0.8, 0.1 > and α1 =< 0.9, 0.1 >, because s(α1) = 0.7 and s(α2) = 0.8, then s(α1) < s(α2). so we can regard the intuitionistic fuzzy number α2 as being greater than α2. definition 3. (intuitionistic fuzzy sets[6, 7]. for any intuitionistic fuzzy number α =< µ, ν >, the accuracy function h(α) of this number is defined as follows: h(α) = µ+ ν 2 (4) the larger h(α) is, the greater the intuitionistic fuzzy number α =< µ, ν > is. for example, for intuitionistic fuzzy numbers α3 =< 0.7, 0.2 > and α4 =< 0.4, 0.2 >, according to definition 4, the accuracy functions of these intuitionistic fuzzy numbers are respectively h(α3) = 0.45 and h(α3) = 0.3. therefore,α3 > α4. according to the definition 2 and 3, based on the score function s(α) and precision function h(α), the intuitionistic fuzzy numbers are compared. for any both intuitionistic fuzzy numbers αi =< µi, υi > and αk =< µk, υk >, if s(αi) ≥ s(αk), then αi ≥ αk; if s(αi) = s(αk) and h(αi) ≥ h(αk), then αi ≥ αk. for example, for intuitionistic fuzzy numbers α5 =< 0.8, 0.1 > and α6 =< 0.6, 0.3 >, due to s(α5) = 0.7, s(α6) = 0.3, therefore α5 > α6; while for the intuitionistic fuzzy numbers α7 =< 0.7, 0.3 > and α8 =< 0.5, 0.1 >, due to s(α7) = 0.4, s(α8) = 0.4, h(α7) = 0.5, h(α8) = 0.3, therefore α7 > α8. however, there exist the shortcomings that how much their uncertainty allotted to the membership and non-membership for two intuitionistic fuzzy numbers, so that the size of the intuitionistic fuzzy numbers cannot be exactly determined, for example, the intuitionistic fuzzy numbers α7, α8. in order to determine the size relationship of the intuitionistic fuzzy numbers, a novel method should be proposed. definition 4. for any given two intuitionistic fuzzy numbers αi =< µi, υi > and αk =< µk, υk >, then the dominance distance index ifdd(xi, xk) of the intuitionistic fuzzy numbers αi and αk can be defined as follows: ifdd(xi, xk) =  1 µi ≥ µk, υi ≤ υk 0 µi < µk, υi > υk 1 2 + 1 4 µi−υi−µk+υk µ(ω) other (5) 206 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... where, µ(ω) stands for the background measure, and it is generally taken as µ(ω) = max {1− υ1, 1− υ2, ..., 1− υn}−min {µ1, µ2, ..., µn}, while ifdd(xi, xk) expresses the degree of intuitionistic fuzzy numbers vi better than the intuitionistic fuzzy number vk. based on the definition 4, intuitionistic fuzzy numbers can be compared. for example, suppose that the background measure µ(ω) is 1, for the intuitionistic fuzzy numbers α9 =< 0.8, 0.1 > and α10 =< 0.5, 0.3 >, such that ifdd(xi, xk) = 0.55, therefore, α11is greater than α12 by the possibility of 55%. 2.2 dominance relationship based on dominance distance index definition 5. assume that an ordered 4-tuple s = (u,a, v, f) is known as an preference information system, if u = {u1, u2, ..., un} is a finite and nonempty set, known as the universe; a = c ∪ d is a finite, nonempty attribute set, while c = {a1, a2, ..., am} and d are respectively the condition attribute set and the decision attribute set; v = ∪va stands for the value domain of the information system s, where va is the value of u with respect to the attribute a ∈ a; f : u ×a → v is an information function. if va is an intuitionistic fuzzy number αa =< µa, υa >, then the information system s is called as the intuitionistic fuzzy preference information system, and denoted as ifs. definition 6. suppose that there exist an intuitionistic fuzzy preference information system ifs = (u,a, v, f), for ∀a ∈ p ⊆ a, xi, xk ∈ u, f(xi, a) = αia =< µia, υia >∈ v , f(xk, a) = αka =< µka, υka >∈ v , λ ∈ (0.5, 1], if ifdd(xi, xk) ≥ λ, then the dominance relationship between the objects xi and xk with respect to the α attribute or the attribute set p can be called as a dominance relationship based on the dominance distance index with threshold value λ, written as (xi, xk) ∈ r≥λ a , (xi, xk) ∈ r≥λ p . property 1.suppose that there exist an intuitionistic fuzzy preference information system ifs = (u,a, v, f), ∀a ∈ p ⊆ a, ∀xi, xk ∈ u, f(xi, a) = αia =< µia, υia >∈ v, f(xk, a) = αka =< µka, υka >∈ v , it holds true: 0 ≤ ifdda(xi, xk) ≤ 1 (6) 0 ≤ ifddp (xi, xk) ≤ 1 (7) proof. according to the definition 4, (1)if µia ≥ µka, υia ≤ υka, then ifdda(xi, xk) = 1; (2)if µia < µka, υia > υka, then ifdda(xi, xk) = 0; (3)if µia ≥ µka, υia ≥ υka or µia ≤ µka, υia ≤ υka, because of the background value measure µ(ω) = max {1− υ1, 1− υ2, ..., 1− υn}−min {µ1, µ2, ..., µn}, then 2(min i (µia)−max i (1−υia)) ≤ µia−υia−µka+υka ≤ 2(max i (1−υia)−min i (µia)), so that −2 ≤ µia−υia−µka+υka max i (1−υia)−min i (µia) ≤ 2, and then there exists advances in systems science and application (2015) vol.15 no.3 207 0 ≤ 1 2 + 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) ≤ 1, therefore, 0 ≤ ifdda(xi, xk) ≤ 1. in conclusion, 0 ≤ ifdda(xi, xk) ≤ 1 can be obtained. let the weight vector of attributes is w = (w1, w2, ..., wp), and it satisfies wt > 0 and p∑ t=1 wt = 1, and then there exists ifddp (xi, xk) = p∑ t=1 wtifddat(xi, xk), therefore, 0 ≤ ifddp (xi, xk) ≤ 1. property 2. for ∀a ∈ p ⊆ a, xi, xk ∈ u , f(xi, a) = αia =< µia, υia >∈ v and f(xk, a) = αka =< µka, υka >∈ v , λ ∈ (0.5, 1], if (xk, xs) ∈ r≥λ a , the following holds true: ifdda(xi, xk) ≤ ifdda(xi, xs). proof. it suffices to show ifdda(xi, xk)− ifdda(xi, xs) ≤ 0. for ifdda(xi, xk), there exist three cases:ifdda(xi, xk) = 1. ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) and ifdda(xi, xk) = 0. for ifdda(xi, xs), there also are three cases. for (xk, xs) ∈ r≥ a , according to the dominance relationship r≥λ a , that is, ifdda(xk, xs) ≥ λ by analyzing the situation, there exit the following 6 cases. (1) when the object xi is definitely better than the objects xk and xs, that is αi > αk and αi > αs, thus ifdda(xi, xs) = 1, ifdda(xi, xk) = 1, therefore, ifdda(xi, xs)− ifdda(xi, xk) = 0. (2)when the objectxi is definitely better than the object xs, and not necessarily better than the object xk, ifdda(xi, xs) = 1 and ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) can be acquired. therefore, the following holds true: ifdda(xi, xs) − ifdda(xi, xk) = 1 2 − 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) . according to definition 4, there exits −2 ≤ µia−υia−µka+υka max i (1−υia)−min i (µia) ≤ 2, that is, ifdda(xi, xk) − ifdda(xi, xs) ≤ 0. (3) when the object xi is definitely better than the objects xs and must be inferior to the object xk, ifdda(xi, xs) = 1 and ifdda(xi, xk) = 0 can be obtained. thus, there exists ifdda(xi, xs) − ifdda(xi, xk) = 1 that is, ifdda(xi, xk)− ifdda(xi, xs) ≤ 0. (4) when the object xi may be superior to the object xs and may be superior to the object xk, ifdda(xi, xs) = 1 2+ 1 4 µia−υia−µsa+υsa max i (1−υia)−min i (µia) and ifdda(xi, xk) = 1 2+ 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) can be acquired. thus, ifdda(xi, xs)−ifdda(xi, xk) = 1 4 µka−υka−µsa+υsa max i (1−υia)−min i (µia) due to ifdda(xk, xs) ≥ 1 2 , thus µka − υka − µsa + υsa ≥ 0, therefore, it then follows that ifdda(xi, xs)− ifdda(xi, xk) ≥ 0. when the object xi may be superior to the object xs and must be inferior to the object xk, ifdda(xi, xs) = 1 2+ 1 4 µia−υia−µsa+υsa max i (1−υia)−min i (µia) and ifdda(xi, xk) = 0 can be obtained. thus ifdda(xi, xs)− ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µsa+υsa max i (1−υia)−min i (µia) . 208 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... from the fact that −2 ≤ µia−υia−µsa+υsa max i (1−υia)−min i (µia) ≤ 2, it follows that ifdda(xi, xs)− ifdda(xi, xk) ≥ 0. when the object xi must be inferior to the objects xs and xk, there exist ifdda(xi, xs) = 0 and ifdda(xi, xk) = 0, and then ifdda(xi, xs) − ifdda(xi, xk) = 0 can be concluded. in conclusion, ifdda(xi, xs)−ifdda(xi, xk) ≥ 0 therefore, ifdda(xi, xk) ≤ ifdda(xi, xs). qed property 3. the dominance relationship based on dominance distance index satisfies the condition of transitivity. proof. it suffices to show that for ∀a ∈ p ⊆ a, xi, xk ∈ u, if ifdda(xi, xk) ≥ 1 2 and ifdda(xk, xs) ≥ 1 2 , then ifdda(xi, xs) ≥ 1 2 . when ifdda(xi, xk) ≥ 1 2 , there exists µia − υia − µka + υka ≥ 0 . accordingly there are two cases which are respectively certainly better and perhaps superior for the object xi than the object xk. that is, if µi ≥ µk, υi ≤ υk, then ifdda(xi, xk) = 1. if ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) and µia − υia − µka + υka ≥ 0 then ifdda(xi, xk) ≥ 1 2 . when ifdda(xk, xs) ≥ 1 2 similarly there are also two cases which are respectively and certainly better and perhaps superior for the object xk than the object xs. thus there exist the following 4 cases where the object xi is better than the object xs. (1)if ifdda(xi, xk) = 1 and ifdda(xk, xs) = 1, then µi ≥ µk, υi ≤ υk and µk ≥ µs, υk ≤ υs, and then µi ≥ µs, υi ≤ υs. therefore ifdda(xi, xs) = 1. (2)if ifdda(xi, xk) = 1, ifdda(xk, xs) = 1 2 + 1 4 µka−υka−µsa+υsa max i (1−υia)−min i (µia) and µka−υka−µsa+υsa ≥ 0, then µia−υia−µsa+υsa ≥ 0. according to definition 4, ifdda(xk, xs) ≥ 1 2 can be obtained. (3)ifdda(xk, xs) = 1 2+ 1 4 µka−υka−µsa+υsa max i (1−υia)−min i (µia) and µia−υia−µka+υka ≥ 0 and ifdda(xk, xs) = 1 2 + 1 4 µka−υka−µsa+υsa max i (1−υia)−min i (µia) and µka − υka − µsa + υsa ≥ 0, such that µia − υia − µsa + υsa ≥ 0 and according to the definition of the dominance distance index, ifdda(xk, xs) ≥ 1 2 can be obtained. (4) if ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µka+υka max i (1−υia)−min i (µia) , µia−υia−µka+υka ≥ 0 and ifdda(xk, xs) = 1, then µia − υia − µsa + υsa ≥ 0. according to the definition of the dominance distance index, therefore, ifdda(xk, xs) ≥ 1 2 . in conclusion, ifdda(xk, xs) ≥ 1 2 can be always obtained. qed. property 4. for ∀a ∈ p ⊆ a, xi, xk ∈ u , αia =< µia, υia >∈ v , αka =< µka, υka >∈ v ,then the following hold true: ifdda(xi, xk)+ifdda(xk, xi) = 1. proof. according to equation (5), (1)if µia ≥ µka, υia ≤ υka, then ifdda(xi, xk) = 1 and ifdda(xk, xi) = 0, therefore ifdda(xi, xk) + ifdda(xk, xi) = 1 can be obtained. advances in systems science and application (2015) vol.15 no.3 209 (2) when µia < µka, υia > υka, it follows that ifdda(xi, xk) = 0 and gdda(xk, xi) = 1. thus, we have gdda(xi, xk) +gdda(xk, xi) = 1. (3) when the other conditions, ifdda(xi, xk) = 1 2 + 1 4 µia−υia−µka+υka µa(ω) and ifdda(xk, xi) = 1 2 + 1 4 µka−υka−µia+υia µa(ω) can be acquired. thus, ifdda(xi, xk) + ifdda(xk, xi) = 1 can be obtained. qed. 3 dominance intuitionistic fuzzy variable precision rough set model in order to deal with preference attributes, with respect to the information system with preference value[39? –41], by introducing dominance relationship into the rough set mode and taking advantage of it to substitute for the indistinguishable relationship constructed the dominance rough set model, and then acquired the decision making rules. based on the thought of dominance rough set model, the dominance relationship based on intuitionistic fuzzy dominance distance index can be used to substitute the equivalence relationship of rough set, and then the dominance intuitionistic fuzzy rough set model is constructed. 3.1 the construction of model definition 7. suppose that there exist an intuitionistic fuzzy preference information system ifs = (u,a, v, f), for p ⊆ a, cl≥t ⊆ d, a threshold value λ ∈ (0.5, 1] and β ∈ (0.5, 1], if d+ p (x) = {y ∈ u |ifdda(x) ≥ λ, ∀a ∈ p } and d− p (x) = {y ∈ u |ifdda(x) < λ, ∀a ∈ p } respectively stand for the λ−p dominating set and dominated set with respect to x, and then the lower approximation and upper approximation of the decision making class cl≥t are respectively defined as follows: aprλ p (cl≥t ) = ∪{x ∈ u : d+ p (x) ⊆ cl≥t } (8) aprλp (cl≥t ) = ∪{x ∈ u : d− p (x) ∩ cl≥t ̸= ∅} (9) and thus, [ aprλ p (cl≥t ), apr λ p (cl≥t ) ] is called as the dominance intuitionistic fuzzy rough set. the λ−p -lower approximation aprλ p (cl≥t ) of (cl≥t ) can be called as the positive domain of the dominance intuitionistic fuzzy rough set and interpreted as the set of the union of all condition classes with confidence threshold value , where the classified objects definitely belong to the upward union (cl≥t ). accordingly, the λ − p -upper approximation aprλ p (cl≥t ) of (cl≥t ) can be interpreted as the union of all classes with confidence threshold value λ, where the classified objects possibly belong to the upward union (cl≥t ). accordingly, based on the lower approximation and upper approximation of the intuitionistic fuzzy rough set, the λ boundary domain and classification quality of the set (cl≥t ) can be defined as follows: bndλp (cl≥t ) = aprλp (cl≥t )− aprλ p (cl≥t ) (10) 210 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... αλ p (cl≥t ) = |aprλ p (cl≥t )| |aprλp (cl≥t )| (11) analogously, the λ -lower and λ -upper approximations of can be respectively defined as follows: aprλ p (cl≤t ) = ∪{x ∈ u : d− p (x) ⊆ cl≤t } (12) aprλp (cl≤t ) = ∪{x ∈ u : d+ p (x) ∩ cl≤t ̸= ∅} (13) and thus, [ aprλ p (cl≤t ), apr λ p (cl≤t ) ] stands for the dominance intuitionistic fuzzy rough set. accordingly, based on the lower approximation and upper approximation of the intuitionistic fuzzy rough set, the λ boundary domain and classification quality of the set (cl≥t ) can be defined as follows: bndλp (cl≤t ) = aprλp (cl≤t )− aprλ p (cl≤t ) (14) αλ p (cl≤t ) = |aprλ p (cl≤t )| |aprλp (cl≤t )| (15) 3.2 attribute reduction and preference rules definition 8. suppose that that there exist an intuitionistic fuzzy preference information system ifs = (u,a, v, f), for p ⊆ a and the given threshold value λ ∈ (0.5, 1], the classification quality of cl can be defined as follows: γλp (cl) = |u − ((∪bnd(cl≥t )) ∪ (∪bnd(cl≤t )))| |u | (16) the classification quality γλp (cl) of cl stands for the ratio of the relation between all the correctly classified objects and all the objects with respect to the attribute set p in the information system. for every minimal subset p ⊆ c, the attribute set p satisfying γλp (cl) = γλc(cl) s called as the reduction of c with respect to cl and denoted byredcl(p ). the preferential decision rule is one kind of dependence form between condition preference attribute and decision preference attribute. based on dominance relationship with the dominance distance index, the rough approximation is acquired, and then the preferential decision rule can be induced and shown as follows: for the given threshold value λ, d≥ -decision rules can take on the following form: if f(x, q1) ≥ rq1 ∧ f(x, q2) ≥ rq2... ∧ f(x, qp) ≥ rqpthenx ∈ cl≥t ; for the given threshold value λ, d≤ -decision rules can take on the following advances in systems science and application (2015) vol.15 no.3 211 form: if f(x, q1) ≤ rq1 ∧ f(x, q2) ≤ rq2... ∧ f(x, qp) ≤ rqpthenx ∈ cl≥t ; where {q1, q2, ...qp} ⊆ c, iff(x, q1) ≤ rq1 ∧ f(x, q2) ≤ rq2... ∧ f(x, qp) ≤ rqpthenx ∈ cl≥t , t ∈ {1, 2, ..., l}. 3.3 comprehensive dominance degree definition 9. suppose that the intuitionistic fuzzy preference information system ifs = (u,a, v, f), for ∀aj ∈ p ⊆ a, xi, xk ∈ u , f(xi, aj) = αij =< µij , υij >∈ v , f(xk, aj) = αkj =< µkj , υkj >∈ v the following is called as the dominance degree with equal weight with respect to the attribute set p for the object xi over the object xk based on the dominance distance index: ifddp (xi, xk) = 1 |p | ∑ ∀a∈p ifdda(xi, xk) (17) where |p | is the cardinality of the attribute set p . in the situation of real-life decision making, due to the fact that different decision makers may very well weigh the attributes differently, the results, obtained on the basis of this uniform treatment of equal weights, can be obviously expected to be inconsistent with the reality. therefore, there is a need to study the situation that the attributes are given different weights. definition 10. suppose that the intuitionistic fuzzy preference information system ifs = (u,a, v, f), for ∀aj ∈ p ⊆ a, xi, xk ∈ u , f(xi, aj) = αij =< µij , υij >∈ v , f(xk, aj) = αkj =< µkj , υkj >∈ v , let the weight vector of the attribute be w = (w1, ..., wt, ..., w|p |), satisfying wt > 0 and |p |∑ t=1 wt = 1. then ifddp (xi, xk) = |p |∑ t=1 wtifdda(xi, xk) (18) is referred to as the different weight dominance degree with respect to the attribute set p for the object xi over the object xk based on the dominance distance index. definition 11. suppose that the intuitionistic fuzzy preference information system is ifs = (u,a, v, f), for ∀xi ∈ u , p ⊆ c, the comprehensive dominance degree of the object xi in all the objects based on the dominance distance index is defined as follows: ifddp (xi) = 1 |u | − 1 ∑ i ̸=k ifddp (xi, xk) = 1 |u | − 1 ∑ i̸=k |p |∑ t=1 wtifdda(xi, xk) (19) 212 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... 3.4 the decision steps based on dominance intuitionistic fuzzy rough set to sum up, the decision steps are following: step1: by adjusting the given threshold value , based on the condition attribute set, the different classifications can be determined. step2: based on the decision attribute set, all classifications can be determined. step3: according to all classifications based on the condition attribute set and the decision attribute set, the lower approximation and the upper approximation of the decision classifications can be acquired, and then the attribute reduction can be obtained, so that t the decision rules can be extracted. step4: by the classification results based on the condition attribute set or the decision attribute set, based on the expressions (18) and (19), the comprehensive dominance degrees of the objects are calculated, and then the ranking of the objects can be determine. 4 case analysis information system security design is that by testing the legal compliance, the confidentiality, the usability, and the reliability of the audiles, the security risk of information system is assessed, and then some audit opinions and suggestions of information system security can be given. while the risk guidance information system security audit based on the risk identification and evaluation is that through the comprehensive analysis of the audiles operation environment and information system running environment, the factors affected the system security are extracted, and then according to the risk assessment, the implementation audit scope and key target are determined, so that the substantive tests are implemented. therefore, the evaluation and judgment of audit risk possess are much more important in the information system security audit process. currently, few scholars take advantage of rough set to make the audit risk judgment of the information system security, and the evaluation is mainly based on the knowledge and experience of audit personnel in the practice process, however, it lacks the scientific and rationality, which has effect on the implementation of information system security audit, and ultimately affects the audit results. in recent years, the audit risk assessment and risk judgment expert system is established to effectively decrease audit judgment deviation and reduce the audit risk. however, due to the complexity and uncertainty of the objective world, as well as the limitation of human ability to understand, there always exist a variety of preference information and fuzzy information in the audit risk assessment and risk judgment expert system, while it is difficult for the existing representation and processing methods of the information system security audit based on the expert experience to acquire exactly knowledge. in view of this, based on the proposed model in advances in systems science and application (2015) vol.15 no.3 213 this paper, the attribute reduction is used to extract the decision rule of the audit risk assessment of the information system security and acquire the key factors and bottleneck factors of affecting the audit risk assessment of the information system security, and then the ranking of the different information system security can be determined, so that the proposed model can provide for audit personals a much reasonable and effective audit risk assessment method and tool. according to the actual audit cases, the related data can be collected and shown in the table 1. in the table 1, there exist 10 audited objects denoted as u = {x1, x2, ..., x10}, and the five attributes c = {a1, a2, a3, a4, a5} which respectively good system environment, good system control, reliable financial data, reliable audit software and standard operation. for each condition attribute, based on the audit results and their own professional quality, its values can be obtained by the comprehensive judgment from the information system audit experts, for example, f(x2, a1) =< 0.5, 0.4 > stands for the fact that 50% experts think that the system environment of the audited object x2 is good, while 40% experts think that it is bad and 10% experts hesitate. the decision attribute set is denoted as d = {d}, and it indicates whether the audit risk of information system security is acceptable. for example, f(x2, d) = 1 expresses that experts regard that the audit risk of the audited object x2 is acceptable. table 1 the audit risk assessment intuitionistic fuzzy information system of the information system security u a1 a2 a3 a4 a5 d u1 < 0.2, 0.6 > < 0.1, 0.7 > < 0.4, 0.4 > < 0.5, 0.4 > < 0.4, 0.4 > 0 u2 < 0.5, 0.4 > < 0.3, 0.6 > < 0.3, 0.6 > < 0.5, 0.2 > < 0.5, 0.4 > 0 u3 < 0.2, 0.6 > < 0.1, 0.8 > < 0.2, 0.7 > < 0.3, 0.6 > < 0.3, 0.6 > 0 u4 < 0.7, 0.1 > < 0.6, 0.4 > < 0.7, 0.2 > < 0.8, 0.2 > < 0.7, 0.2 > 1 u5 < 0.3, 0.6 > < 0.2, 0.7 > < 0.2, 0.7 > < 0.2, 0.6 > < 0.3, 0.7 > 0 u6 < 0.6, 0.3 > < 0.6, 0.4 > < 0.6, 0.3 > < 0.7, 0.2 > < 0.5, 0.4 > 1 u7 < 0.2, 0.6 > < 0.2, 0.6 > < 0.5, 0.4 > < 0.5, 0.4 > < 0.2, 0.6 > 1 u8 < 0.1, 0.6 > < 0.2, 0.6 > < 0.4, 0.5 > < 0.3, 0.6 > < 0.2, 0.7 > 0 u9 < 0.7, 0.2 > < 0.6, 0.4 > < 0.8, 0.1 > < 0.6, 0.3 > < 0.8, 0.2 > 1 u10 < 0.6, 0.2 > < 0.6, 0.2 > < 0.8, 0.2 > < 0.4, 0.5 > < 0.4, 0.5 > 1 step1: according to the condition attribute set c, when λ = 0.55, based on the dominance distance index, the universe can be divided into the following classifications: u/c = {x1, x2, x3, x4} where x1 = {x1, x3, x5} , x2 = {x7, x8} , x3 = {x2, x6, x9, x10} , x4 = {x4} and 214 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... x1 < x2 < x3 < x4. step2:according to the decision attribute set d, the universe can be divided into the following classifications: u/d = {cl1, cl2}. where cl1 = {x1, x2, x3, x5, x7, x8}, u/d = {cl1, cl2}. step3: according to the classifications based on the condition attribute set and the decision attribute set, the lower approximation aprλ p (cl≥1 ) of (cl≥1 ) and the upper approximation aprλ p (cl≥1 ).the lower approximation aprλ p (cl≤2 ) of (cl≤2 ) and the upper approximation aprλ p (cl≤2 ) respectively are: aprλ p (cl≤1 ) = {x1, x3, x5, x7, x8} , aprλp (cl≤1 ) = {x1, x2, x3, x5, x6, x7, x8, x9, x10} aprλ p (cl≥2 ) = {x4} , aprλp (cl≥2 ) = {x2, x4, x6, x9, x10} γλ,βp (cl) = | {x1, x3, x5, x7, x8} ∪ {x4} | | {x1, x2, x3, x4, x5, x6, x7, x8, x9, x10} | = 6 10 = 0.6 when λ = 0.55, the reduction {a1, a2} based on the genetic algorithm can be acquired, therefore, it is well known that the factors of good system environment and good system control should be considered in the audit risk assessment of information system. according to the reduction {a1, a2}, the probabilistic decision rules can be generated and shown in the table 2. table 2 the probabilistic decision rules based on the reduction {a1, a2} with λ = 0.55. rules support number confidence a1 6< 0.3, 0.6 > and a2 6< 0.2, 0.7 > 100%−−−→ d = 0 2 100% a1 6< 0.2, 0.6 > and a2 6< 0.2, 0.6 > 100%−−−→ d = 0 3 100% a1 >< 0.6, 0.3 > and a2 >< 0.6, 0.4 > 100%−−−→ d = 1 4 100% step4: because there are four classifications based on the condition attribute set, and it satisfies x1 < x2 < x3 < x4, that is, {x1, x3, x5} < {x7, x8} < {x2, x6, x9, x10} < x4. for the object x4, due to the fact that there only exists the object x4 in the x4, so its comprehensive dominance degree no longer needs to calculate; for the classes {x1, x3, x5} , {x7, x8} and {x2, x6, x9, x10}, their comprehensive dominance degrees should be computed. according to the attribute dependency based on the proposed the model, the attribute dependency degree of each attribute can be acquired and then they are standardized, thus the weight vector is obtained as follows w = (w1, w2, w3, w4, w5) = (0.2125, 0.2024, 0.1964, 0.1998, 0.1889). for {x1, x3, x5}, based on the formula (18) and (19), ifdda(x1) = 0.5204, advances in systems science and application (2015) vol.15 no.3 215 ifdda(x3) = 0.1745 and ifdda(x5) = 0.3035 can be calculated, therefore, x1 > x5 > x3. for {x7, x8}, because the both objects are in the classification, it is necessary to only calculate the dominance degree with weights, and then it will exist ifdda(x7) = ifdda(x7, x8) = 0.8167 and ifdda(x8) = ifdda(x8, x7) = 0.1833, therefore, x7 > x8. for {x2, x6, x9, x10}, based on the formula (18) and (19), ifdda(x2) = 0.1469, ifdda(x6) = 0.4039, ifdda(x9) = 0.5743 and ifdda(x10) = 0.3563 can be calculated, therefore, x2 < x10 < x6 < x9. to sum up, the ranking of the audited objects is x3 < x5 < x1 < x7 < x8 < x2 < x10 < x6 < x9 < x4. based on the above calculation and analysis, the proposed model possesses certain fault-tolerant ability by adjusting parameter λ, and it can well do with the intuitionistic fuzzy information system with preference information and realize the extraction of group decision rules and the ranking of schemes, so that it can well deal with the real multi-attribute decision making problems with preference information and fuzzy information. 5 conclusion in order to achieve the law of mining the real decision information system and extract decision rules, with respect to the intuitionistic fuzzy preference information system, we construct the dominance intuitionistic fuzzy rough set based on dominance distance index, and then exploit it to extract decision-making rules and determine the ranking of decision schemes. the results show that the hybrid model can well treat the decision making problems with fuzzy information and preference information and realize the mining for the law of the decision information system, the extraction of group decision rules and the ranking of schemes, meanwhile, the proposed model can be applied into the project evaluation, military system decision and other fields. however, the proposed model cannot well deal with the intuitionistic fuzzy preference information system consisting of noise data. for how to solve the multi-attribute decision making problems with noise data, the further study will involve in it. references [1] z. pawlak. (1982),”rough set”. int j of computer and information science,vol.11, no.5, pp.341-356. [2] z. pawlak, and a. skowron. (2007a),”rudiments of rough sets”. information sciences, vol. 1, no. 1, pp. 3-27. [3] z. pawlak, and a. skowron. (2007b),”rough sets: some extensions”. information sciences, vol. 1, no. 1, pp. 28-40. 216 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... [4] z. pawlak, and a. skowron. (2007c),”rough sets and boolean reasoning”. information sciences, vol. 1, no. 1, pp.41-73. [5] d. dubois, and h. prade. (1999), the oriental aspect of reasoning about data, academic publishers, london. [6] k. atanassove. (1986), “intuitionistic fuzzy sets”, fuzzy sets and systems, vol. 20, pp. 87-96. [7] k. atanassove. (1989),“ more on intuitionistic fuzzy sets”, fuzzy sets and systems, vol. 33, pp. 37-46. [8] d. coker. (1998),“fuzzy rough sets are intuitionistic l-fuzzy sets”,fuzzy sets and systems, vol. 96, no. 3, pp. 381-383. [9] c. cornelis, m.d. cock, and e. e. kerre. (2003), “intuitionistic fuzzy rough sets: at the crossroads of imperfect knowledge”, expert systems, vol. 20, no. 5, pp. 260-270. [10] c. cornelis, g. deschrijver, and e. e. kerre. (2004), “implication in intuitionistic fuzzy and interval-valued fuzzy set theory: construction, classification, application”, int j of approximate reasoning, vol. 35, no. 1, pp. 55-95. [11] p. jena, and s.k. ghosh. (2002), “intuitionistic fuzzy rough sets”, noteson intuitionistic fuzzy sets, vol. 8, pp. 1-18. [12] x.l.xu, y.j. lei, and q.y. tan. (2008), “intuitionistic fuzzy rough sets based on triangle norm”, control and decision, vol. 23, no. 8, pp. 900-904. [13] l.zhou, and w.z. wu. (2008), “on generalized intuitionistic fuzzy approximation operators”, information sciences, vol. 178, no. 11, pp. 2448-2465. [14] l. zhou, w.z. wu, and w.x. zhang. (2009), “on characterization of intuitionistic fuzzy rough sets based on intuitionistic fuzzy implicators”, information sciences, vol. 179, no. 7, pp. 883-898. [15] w.h. xu, y.f. liu, w.x. sun. (2014), “uncertainty measure of atanassov’s intuitionistic fuzzy t equivalence information systems”, journal of intelligent & fuzzy system, vol. 26, no. 4, pp. 1799-1811. [16] h.y. yin, y.j. lei, and y. lei. (2008),“similarity measures of intuitionistic fuzzy rough sets based on the hamming distance”, computer applications and software, vol. 35, no. 12, pp. 15-16,75. advances in systems science and application (2015) vol.15 no.3 217 [17] t. feng, j.s. mi, s.p. zhang. (2014), “belief functions on general intuitionistic fuzzy information systems”,information sciences, vol. 271, no.1, pp. 143-158. [18] m.l. lin and w.p. yang. (2011), “properties of interval-valued intuitionistic fuzzy rough sets with implicators”, j of shandong university(natural science), vol. 46, no. 8, pp. 104-109. [19] z.m. zhang, c. wang, d.z. tian. (2014), “a novel approach to intervalvalued intuitionistic fuzzy soft set based decision making”, applied mathmatical moddeling, vol. 38, no. 4, pp.1255-1270. [20] z.m. zhang, and j.f. tian. (2011), “interval-valued intuitionistic fuzzy rough sets based on implicators”, control and decision, vol. 25, no. 4, pp. 614-618. [21] z.m. zhang, y.c. bai, and j.f. tian. (2010), “intuitionistic fuzzy rough sets based on intuitionistic fuzzy coverings”, control and decision, vol. 25, no. 9, pp. 1369-1373. [22] h.l. yang. (2012), “interval valued fuzzy rough set model on two different universes and its application”, fuzzy systems and mathematics, 26, no. 4, pp. 168-173. [23] s.q. luo, w.h. xu. (2014), “rough atanassov’s intuitionistic fuzzy sets model over two universes and its applications”, scientific world journal, vol. 2014, no. 5, pp.348-383. [24] y.l. lu, y.j. lei, and j.x. hua. (2009), “attribute reduction based on intuitionistic fuzzy set”, control and decision, vol. 24, no. 3, pp. 335-341. [25] h. chen, h.c. yang. (2011), “one new algorithm for intuitiontistic fuzzyrough attribute reduction”, journal of chinese computer systems, vol. 32, no. 3, pp. 506-510. [26] b. huang and d.k. wei. (2011), “distance-based rough set model in intuitionistic fuzzy information systems and application”, system engineering theory and practice, vol. 3, no. 17, pp. 1356-1363. [27] b. huang and h.x. li. (2011), “evaluation rules acquisition of performance audit for it projects in china based on dominance intuitionistic fuzzy rough set model”, computer science, vol. 38, no. 10, pp. 223-227. [28] h. esmail, j. maryam, l. habibolla. (2013), “rough set theory for the intuitionistic fuzzy information”, systems international journal of modern mathematical sciences, vol. 6, no. 3, pp. 132-143. 218 s. baizakov, n.baizakov and jeffrey forrest: a novel multi-attribute decision... [29] j. wang, s.y. liu, j. zhang. (2006), “rough set approach to group decision making based on linguistic information processing”, journal of systems engineering, vol. 21, no. 1, pp. 18-23. [30] g. xie, j.l. zhang, k.k, lai. (2008), “variable precision rough set for group decision making: an application”, international journal of approximate reasoning, no. 49, pp. 331-343. [31] w.j. bi, c.h. chen. (2008), “approach to multiple decision tables analysis based on variable precision rough set”, systems engineering and electronics, vol. 30, no. 6, pp. 1074-1078. [32] f. tian, l.. liu, w.j. you. (2008), “multi form preference information group decision making approach based on rough set theory”, computer integrated manufacturing system, vol. 14, no. 12, pp. 2408-2413. [33] y.p. jiang, a.m. liang. (2011). “a method based on rough sets for multi attribute group decision making with incomplete interval linguistic information”, journal of systems and management, vol. 20, no. 4, pp .485-489. [34] g. wei, s.y. wang, k.k, lai. (2011), “optimal stable interval in vprs based group decision making :a further application”, expert systems with applications, vol. 38, no. 11, pp. 13757-13763. [35] l. zhao, z. xue. (2009), “multi attribute group decision making information system security assessment based on vprs”, journal of shanghai jiaotong university, vol. 43, no. 7, pp. 1161-1166. [36] c.y. lee, h. lee, h. seol. (2012), “evaluation of new service concepts using rough set theory and group analytic hierarchy process”, expert systems with applications, vol. 39, no. 3, pp. 3404-3412. [37] j.j. zhu, j.g. zheng, j.b. li. (2012), “rough classification algorithm for uncertain extension group decision making”, control and decision, vol. 27, no. 6, pp. 851-856. [38] w. xiong, q.y. su, j.l. li. (2012), “the group decision making rules based on rough sets on large scale engineering emergency”, systems engineering proscenia, vol. 4, pp. 331-337. [39] s. greco, b. matarazzo, r. slowinski. (1998), “a new rough set approach to multi-criteria and multiattribute classification”, lecture notes in artificial intelligence, vol. 1424, pp. 60-67. advances in systems science and application (2015) vol.15 no.3 219 [40] s. greco, b. matarazzo, r. slowinski. (1999), “rough approximation of a preference relation by dominance relations”, european journal of operational research, vol. 117, no. 1, pp. 63-83. [41] s. greco, b. matarazzo, r. slowinski. (2001), “rough sets theory for multicriteria decision analysis”, european journal of operational research, vol. 129, no. 1, pp. 1-47. [42] s. greco, b. matarazzo, r. slowinski, j. stefanowski. (2001), “an algorithm for induction of decision rules consistent with dominance principle”, lecture notes in artificial intelligence, vol. 2005, pp. 304-313. corresponding author yong liu can be contacted at: clly1985528@163.com adv sist sci appl 2017; 17(2); 43-51 published online in http://ijassa.ipu.ru/ojs/ijassa/article/view/258 copyright ©0000 assa. adv. in systems science and appl. (0000) reconfigurable architecture for image feature detection rajesh nandalike 1 , saroja devi hande 2 1) department of electronics and communication engineering e-mail: nrajesh7@gmail.com 2) department of computer science and engineering nitte meenakshi institute of technology, bengaluru, india e-mail: hsarojadevi@gmail.com abstract: hardware-based developments are useful for variety of image based applications such as highly challenging video surveillance. fpga comprises of combination of the hardware attributes of an asic, supporting reconfigurability, reduced time-to-market and real-time performance. hardware-based feature detection is a promising solution that exploits inherent parallelism in algorithms to accomplish significant improvement in speed. the efficient usage of resources still remains a challenge that specifically determines the cost of hardware, which can be addressed using our approach. in this paper, we have presented the hardware architecture for feature detection part of the scale invariant feature transform (sift) algorithm for frames from an hd-720p video. the proposed architecture is designed using xilinx system generator (sysgen) tool and implemented on genesys2 kintex-7 fpga development board. keywords: field programmable gate array (fpga), scale-invariant feature transform (sift), feature detection, keypoint detection, feature point, image feature. 1. introduction feature detection and matching are fundamental functionalities in many image applications, video surveillance and also computer vision. they are computationally intensive, requiring considerable amount of resources in terms of time, power, memory and silicon. scale invariant feature transform (sift) algorithm proposed by lowe is one of the potent algorithms in image matching and object recognition [1,2]. high resource utilization and computational complexity make sift algorithm challenging to cope up with the realtime performance for software implementation. feature-based identification is the prevailing object identification strategy. it employs one of the feature extraction algorithms to extract the important features of the image. the development of numerous feature extraction algorithms has been taking place during last few years. canny and sobel’s edge detectors, binary robust independent elementary feature (brief) and harris corner detectors are local feature extraction algorithms that are employed to extract features from an image in object recognition systems [3,4]. every algorithm attempted to enhance the uniqueness and robustness across image transformation operations. this paper proposes hardware architecture for feature detection on fpga, so that the chip acts as a standalone system for feature detection. based on the analysis of feature detection algorithms, five steps must be performed for feature detection and matching, visualization, pre-processing, feature detection, feature descriptor building, feature matching and postprocessing. sift algorithm is used for feature detection in this paper. the approach aims at suitable algorithmic selection and testing using matlab. different algorithms are tested and the corresponding hardware implementation architectures 44 rajesh n., saroja devi h. copyright ©2017 assa. adv. in systems science and appl. (2017) are investigated. feature detection techniques are implemented on fpga, which are directly adaptable by robotic units and air-borne vehicles. 2. scale invariant feature transform (sift) sift is the most competent approach to identify and describe invariant features of an image [5]. the features obtained are invariant to image scaling, rotation and partially invariant to change in illumination. it is an algorithm where image data is converted to scaleinvariant coordinates corresponding to its local features. sift consists of four major stages:  scale-space peak selection: in the first stage, the identification of feasible interest points in an image with their position and scale is accomplished. this is utilized effectively to find out the interest points that are stable, by constructing the gaussian pyramid and searching for the interest points in a series of difference-of-gaussian (dog) images.  key point localization: in the second stage, the feature points are restricted to sub-pixel precision and are waived if they are ambiguous.  orientation assignment: in the third stage, orientations are allotted to each and every feature point position depending on the image gradient directions. the orientation, scale and location to each feature point empowers sift to construct an accepted perspective for the feature point that is invariable to identical transformations.  key point descriptor: in the fourth stage, image feature descriptor is built for each feature point. the sift algorithm forms a depiction for each feature point depending on a patch of pixels in its neighborhood. the patch that has been formerly centered about the feature point's location will be rotated depending on the dominant orientation and scaled to suitable size. the sift descriptor produces 128 feature vectors for each feature point [69]. 3. feature detection the properties of features found in an image make them relevant for identifying different images of a similar scene. detecting good features is a complex problem. the following section gives a discussion on the properties of the ideal local feature. the desired properties from a feature depend on the actual application. a feature is a point-of-interest in an image. it is a slice of information which is suitable for determining the computational assignment identified with respect to a particular application. the properties of a good feature are: they are consistent over several images of the same scene, insensitive to noise, invariant towards certain transformations [10]. sift uses scale-space extrema as candidate features. this work concentrates on the scale-space extrema detection with focus on dedicated hardware implementation. the gaussian kernel is well-defined in 2d as, (1) in equation (1), σ determines the width of kernel and is often referred as the inner scale. it is the standard deviation while the σ 2 is the variance. the term 1/(2πσ 2 ) is the normalization constant and it makes the integral over the exponential function unity. this constant in the equation makes it a normalized kernel with the integral unity for every σ. this means that increasing the σ value effectively decreases the height of kernel while increasing its width. reconfigurable architecture for image feature detection 45 copyright ©2017 assa. adv. in systems science and appl. (2017) the procedure of sift feature point detection comprises of developing difference-ofgaussian (dog) pyramid of the image. from the developed dog pyramid, maxima and minima known as scale-space extrema (also known as feature points or interest points) are recognized. the procedure of sift feature identification is given in fig. 1. fig. 1. computational flow of feature detection 3.1. dog pyramid construction the sift algorithm focuses on the image locations that display immense neighborhood transforms in their visual appearances. the dog pyramid is built with the objective of identifying the feature points. the input image i(x, y) is convolved with a gaussian kernel k(x, y; σ), where σ is the size of gaussian kernel. the product is gaussian-filtered image symbolized by equation (2). (2) where conv2(•) symbolizes the 2-d convolution procedure and gaussian kernel is represented by equation (3). (3) the dog image is the difference between the two gaussian filtered images over successive scales, as depicted by equation (4). (4) where k is a multiplicative factor. the convolution of image with the variable size gaussian kernel produces a mass of blurred images with the quantity of blur dependent on scale aspect. the size of gaussian kernel is varied using a constant multiplicative factor k. 3.2. stable key-point detection after the creation of dog image pyramid [11], the local maxima and minima [12,15] is found by comparing every pixel with its 26 neighborhood in the 3x3 regions. these local maxima and minima feature points are shown in fig. 7. when feature point candidates have been created, the low contrast and solid edge response points must be eliminated to make it sturdy against disturbance. this approach gives rise to a lot of feature points depending on the image size. quality of feature points is more important than the quantity for reliable object recognition. the features that have greater probability to be found in the other version of the image exhibit better quality. to check the stability of feature points, the local extrema are compared with a minimum threshold value. the extrema that have relatively higher minima or maxima value pass the threshold test and are considered, while the weak feature points having low contrast are discarded, resulting in less but more stable candidates. 4. sysgen model for feature detection the architecture developed for each of the stage feature detection and matching is first implemented in software, then feature detection based on sift algorithm is implemented in 46 rajesh n., saroja devi h. copyright ©2017 assa. adv. in systems science and appl. (2017) sysgen model based environment [13]. this was carried on to deal with the issues concerned with that of synchronization and data type matching. upon obtaining a simulated design, hardware co-simulation of the same carried on. this was intended to perform functional verification also taking into account the issues concerned with routing delays, latency and timing constraints. implementation of feature detection on fpga is followed on a ‘design module’ basis [14]. initially, the input image which is a 2-d matrix is converted into a 1-d vector. this 1-d vector from matlab workspace environment is passed as an input to the gaussian filter blocks of sysgen to perform 2-d convolution operation between input image and gaussian kernels. the add/sub blocks are used to find the difference between two gaussian images will be taken. the results then passed through a 3x3 window generator subsystem and maxima and minima blocks. finally using and and or gates strong features of the image will be detected. then finally the resultant of the sysgen model will be a 1-d vector taken onto the matlab workspace. this 1-d vector is then converted back to its 2d form using a matlab program to obtain the strong key-points. the overall sysgen model of feature detection is as shown in fig. 2. the model shown can be categorized into the following major subsystems: subsystem 1: 3x3 window generator. subsystem 2: maxima and minima block. subsystem 3: 5x5 gaussian filter. fig. 2. sysgen model for feature detection 4.1. subsystem 1: sysgen model for 3x3 window generator this subsystem implements the 3x3 window on every dog images. the design comprises of six delays and two virtex line buffers to form a 3x3 window on every dog images. the virtex line buffers have a depth equal to the number of columns in an input image.the sysgen model of 3x3 window generator is as shown in the fig. 3. reconfigurable architecture for image feature detection 47 copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 3. sysgen model for window generator 4.2. subsystem 2: sysgen model for maxima and minima blocks the implementation of the maxima and minima blocks is carried out in this subsystem. these blocks are used for finding the maximum value among the 26 neighbors and employs ‘compare and select’ process. the maxima and minima are found by comparing every pixel with its 26 neighborhood in the 3x3 regions. these local maxima and minima points are treated as candidate feature points. the design comprises of 26 relational blocks with comparison is set to a > b for maxima block and a < b for minima block, to find the maximum and minimum value among the 26 neighbors. the sysgen model of maxima and minima block is as shown in the fig. 4. fig. 4. sysgen model for maxima and minima block 4.3. subsystem 3: sysgen model for 5x5 gaussian filter the 1-d vector from matlab workspace environment is passed as an input to the convolution blocks of sysgen to perform 2-d convolution operation between input image and gaussian kernels. the resultant of the sysgen model will be a 1-d vector taken onto the matlab workspace. this 1-d vector is to be then converted back to its 2-d form using a matlab program to obtain the gaussian filtered image. the overall sysgen model of 5x5 gaussian filter is as shown in fig. 5. convolution block subsystem gives 2-d convolution of input image with the corresponding gaussian kernels. this design comprises of five multiplier blocks, concatenated using delay blocks and virtex line buffers with a depth equal to the number of columns in an input image. the design of a 2-d convolution block is shown in fig. 5. 48 rajesh n., saroja devi h. copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 5. sysgen model for gaussian filter 5. results and discussions the experimental result of feature detection for the above proposed architecture is presented. the performance can be considered based on three parameters: accuracy, speed and hardware utilization. there is a trade-off present between these parameters; higher accuracy consumes more hardware resources and processing time. in this work, the preference is given to accomplish high accuracy in real-time restraint. processing time: the processing time is anticipated based on the number of clock cycles necessary to accomplish every task and the operational frequency for that particular module [16]. the processing time is anticipated as number of clock cycles per task divided by operational frequency in hertz. sift module processing time: the maximum operating frequency of the anticipated design is 85.390mhz. the sift feature extraction module takes 1280x720 clock cycles to examine the input image and identify the feature points. feature detection: a high definition image is selected in the process of analysing the performance of architecture. the hd image shown in fig. 6 is selected. at this stage the input image was convolved with the gaussian kernel. (a) (b) fig. 6. (a) input image (b) gaussian smoothened image reconfigurable architecture for image feature detection 49 copyright ©2017 assa. adv. in systems science and appl. (2017) the next step is to take difference of gaussian to get the minima and maxima points. it gives the co-ordinates of the pixel. the feature points of the image are obtained. the feature points for different values of size of the gaussian kernel or standard deviation (σ) are shown in the fig. 7. increase in the σ value decreases the height of the kernel while the width of the kernel increases. the above analysis is carried out for different 200 images. the results of synthesis show that the proposed architecture requires less resources and achieves the real-time processing requirements. speed can be further increased by using higher frequency. the system is able to achieve maximum frequency of 100mhz. furthermore it is also possible to increase the frequency beyond 100mhz by introducing more pipeline stages at the expense of more resources. however increasing the frequency directly impacts the power requirements. in applications that require lower frame rate a lower clock can be used to get more efficiency in terms of power and resources. (a) (b) (c) (d) fig. 7. feature points using gaussian kernels of different variances: (a) σ = 5 (b) σ = 5.62 (c) σ = 5.93 (d) σ = 6.25 the hardware utilized the feature detection module in the proposed architecture is given in table 1. the fpga essentially consists of hardware resources such as memory, slice registers, slice luts, lut flip flop pairs and dsp blocks [11,12]. the results are reported from the synthesis reports generated by xilinx ise environment. table 1. utilized hardware for the feature detection device utilization summary (estimated values) logic utilization used available utilization number of slices 19487 25350 7.68% number of slices lut 5345 101400 5.27% number of slices registers 8427 202800 4.15% lut as flip flop pairs 6959 101400 6.86% 50 rajesh n., saroja devi h. copyright ©2017 assa. adv. in systems science and appl. (2017) number of brams 34 350 10.46% number of iobs 33 400 8.25% the input image size of 1280x720 is considered with five gaussian levels and a kernel size of 5. the effect of changing number of octaves is negligible since the resources are time shared among the octaves. the maximum operating frequency of hardware is 100 mhz and minimum period is 10ns. the total time to extract the number feature points does not depend on the number of feature points. for this design, we have achieved 108 frames per second for hd-720p video which is much above the real-time processing requirement. this gives an overview of synthesis results for the target fpga xilinx genesys 2 kintex – 7 xc7k325t (package: fbg676, speed grade: -1). xilinx vivado (version 2014.4) has been used for the synthesis of design. table 1 recapitulates the hardware resources used in the implementation the of the feature detection module. conclusion the hardware architecture for feature detection based on scale invariant feature transform (sift) is proposed in this paper. the design of computationally effective hardware architecture for feature recognition and coordinating system on a solitary fpga chip is projected. this system has the capability to detect sift features, extract the descriptors for the detected features and complete features matching for two images taken at separate standpoints, rotation, scaling and change in illumination. the real time feature identification and matching for a series of images are accomplished. the architecture is implemented on a xilinx genesys 2 kintex 7 fpga. the results presented shows that the hardware is reliable and supports real-time applications on an hd image size of up to 1280x720. the maximum operating frequency of hardware is 100 mhz and minimum period is 10ns. it is possible to use this feature extraction hardware in realtime high definition video applications. references [1] lowe, d. g. (2004). distinctive image features from scale-invariant keypoints, int. j. computer vision, 60(2), 91-110. [2] lowe, d. g. (1999). object recognition from local scale-invariant features, proc. of the seventh ieee int. conf. on computer vision, kerkyra, greece, 1150-1157. [3] jian wu, j. cui, z., sheng, v. s., zhao, p., su, d. & gong, s. (2013). a comparative study of sift and its variants, measurement science review, 13(3). https://doi.org/10.2478/msr-2013-0021 [4] harris, c. (1988). a combined corner and edge detector, proc. 4th alvey vision conf., manchester, uk, 147-152. [5] wang, j., zhong, s., yan, l. & cao, z. (2014). an embedded system-on-chip architecture for real-time visual detection and matching, ieee transactions on circuits and systems for video technology, 24(3), 525-538. [6] mishra, p., nidhi a.i., kishore, j.k., nandini, s. & iffat, u. (2014) embedded hardware architectures for scale and rotation invariant feature detection, proc. of ieee int. conf. electronics, computing and communication technologies (ieee conecct), bangalore, india. https://doi.org/10.2478/msr-2013-0021 reconfigurable architecture for image feature detection 51 copyright ©2017 assa. adv. in systems science and appl. (2017) [7] alhwarin, f., wang, c., risti-durrant, d. & graser, a. (2008). improved siftfeatures matching for object recognition, proc. of bcs int. academic conf., london, uk, 179-190. [8] wang, z., xiao, h., he, w., wen, f. & yuan, k. (2013) real-time sift-based object recognition system, in proc. ieee int. conf. on mechatronics and automation (icma), takamatsu, japan, 1361-1366. [9] cheung, w. & hamarneh, g. (2009). n-sift: n-dimensional scale invariant feature transform, ieee trans. on image processing, 18(9), 2012-2021. [10] tuytelaars, t. & mikolajczyk, k. (2007). local invariant feature detectors: a survey, foundations and trends in computer graphics and vision, 3(3), 177-280. [11] raut, n.p. & gokhale, a.v. (2013). fpga implementation for image processing algorithms using xilinx system generator, iosr j. of vlsi and signal processing (iosr-jvsp), 2(4), 26-36. [12] swaraj, d. & madhumati, g.l. (2014). fpga implementation of sift algorithm using xilinx system generator, int. j. emerging trends in electrical and electronics, 10(10), 80-85. [13] qasaimeh, m., sagahyroon, a. & shanableh, t. (2014). a parallel hardware architecture for scale invariant feature transform (sift), int. conf. multimedia computing and systems (icmcs), marrakech, morocco, https://doi.org/10.1109/icmcs.2014.6911251 . [14] bonato, v., marques, e. & constantinides, g.a. (2008). a parallel hardware architecture for scale and rotation invariant feature detection, ieee trans. circuits and systems for video technology, 18(12), 1703-1712. [15] rajesh, n., kulkarni, r.r., sarojadevi, h. (2014). hardware architecture for scale and rotation invariant feature detection for image registration, int. conf. emerging research in computing, information, communications and applications, 239-244. [16] zhong, s., wang, j., yan, l., kang, l., cao, z. (2013). a real-time embedded architecture for sift, j. systems architecture, 59, 16–29. https://doi.org/10.1109/icmcs.2014.6911251 microsoft word 9 gang xu, yutang dai, jianlei cui--micro-manufacturing technology using duv laser and nc.doc 270-279 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc micro-manufacturing technology using duv laser and nc* gang xu, yutang dai and jianlei cui key laboratory of fiber optic sensing technology, wuhan university of technology, wuhan 430070, china abstract combining deep uv laser with nc technology, complex 3d structures can be generated. for fused silica, the ablation depth and surface roughness are investigated after 157nm laser processing. under laser fluence of ~4j/cm2 and spot size of ~35μm, the ablation rate of fused silica is about 75nm/pulse for 157nm laser. a better ablation quality could be achieved under the repetition rate of ~10hz. assisted by numerical control, three-dimensional micro-structures are fabricated using 157nm laser technique. keywords duv laser, micro-manufacturing, numerical control, fused silica 1.introduction moems (micro-optical-electro-mechanical system) is a very dynamic research direction developed from mems technology in recent years, which is a new type of micro-optical system resulting from the combination of micro-mechanical, micro-electronics and micro-optics. moems device has smaller size, lighter weight, so has lower power consumption, smaller inertia, shorter response times, etc[1,2,3,4]. but the forecast advantages are not fully reflected, in large part due to the poor level of micro-mechanical processing[5]. currently, the majority of such photonics chips or mems parts are planar devices based on silicon or silica-on-silicon substrates but the manufacturing procedures are still largely reliant on combinations of conventional exposure and etching steps. in particular, three-dimensional structuring of most mems parts is usually reliant on liga process[6]. there are so many constraints using abovementioned process that developments of more cost-effective process is needed. among various new manufacturing methods, laser micro-processing is a promising choice. one of the main attractions of laser micro-processing is that lasers offer great flexibility in the rapid prototyping and evaluation of different designs. many different process routes can also be tried out with the same laser tool in a relatively short time so the developmental cycle is also much faster than with many conventional techniques. sio2 has low thermal expansion coefficient, permeability to gases and to ionic contaminants, excellent chemical inertness, good electric insulation, and has more attractive mechanical properties than ordinary glass [7,8,9], it has become one of the most important materials for fabricating optical lens and moems devices, such as micro fluidic chip, soc (system on chip) and mcm(multi-chip module). however, fused silica is relatively transparent for visible light and ultraviolet radiation, which can not be effectively absorbed, so the issue of how to process it efficiently is of crucial importance. due to development of special optical component, three dimension micromachining with deep-ultraviolet laser (duv, mainly 157nm laser) has become feasible technically. the photon energy of 157 nm laser is up to 7.9-ev [10], can directly hit the ionic bond of sio2 by photo-chemical effect. besides, 157nm laser have the benefit of being compatible with mask projection systems, so that a precision laser-ablation processing can be performed on glass surfaces. the main process steps of laser direct ablation with the sample are shown in fig. 1. * national natural science foundation of china (grant no.50775169 and no.60537050) advances in systems science and applications (2011), vol.11, no.3-4 271 1 2 3 4 laser beam is irradiated on substrate through a mask photonic energy is absorbed and ionic bonding is destroyed material particles spray out and redeposit around micro structure is shaped fig.1. the ablation steps of 157nm laser process in this study, preliminary research on micro-ablation of 157nm duv laser was performed. the main aim of those researches is to investigate the micro-processing performance of 157nm laser.influence of laser process parameters on surface qulity was investigated. 3d micro structures produced by duv laser and nc process, are demonstrated. 2.duv laser micromachining system a dual-laser system (m2000: dual laser processing tool) manufactured by exitech ltd of england was used. the system is adaptive to wide variety of laser processing requirements, based around uv laser materials processing. two laser sources are fitted to the m2000. in one case the beam from a high repetition rate uv dpss laser operating at 355nm illuminates an aperture. this is then imaged onto the workpiece using a high resolution projection lens. in the other case an f2 laser is fitted which operates at a wavelength of 157nm. the beam from this is imaged onto the workpiece using a high resolution reflective lens. fig. 2 illustrates schematic of dual-laser precision micro-ablation system.. the 157nm laser source is the model m-100 manufactured by tui laser of germany. the pulse duration is 20ns, and the pulse repetition rate is 100hz. the output of the 157nm laser source is 1.5w, and the laser pulse energy is 25mj. due to the high absorption of 157nm radiation in the atmosphere, at atmospheric pressure the photon penetration of o2 is only ~50μm, the entire beam path from the laser output coupler to the sample was purged with high purity nitrogen. the purge gas escaped from the bottom of the schwarzschild lens onto the sample. the purging also helps prevent contamination of the optics by deposited films from airborne organic contaminants that are photo-dissociated by 157nm radiation. 272 xu: micro-manufacturing technology using duv laser and nc the workpiece mounted on the worktable can be moved by 4-axis (x, y, z and r) nc stages. motions of the stages and the firing of the laser are coordinated by the aerotech unidex 500 motion controller. providing up to four axes of synchronized servo control and four axes of stepper control, the unidex 500 supports the motion requirements of today’s most demanding machines. the parameters of the motion control stages are shown in table 1, and fig. 3 shows the schematic of the motion system. fig.2. schematic of optical system of duel laser process tool fig.3. schematic of the motion control stages table 1 the parameters of the aerotech motion control stages axis description travel drive resolution maximum feedrate typical feedrate x x workpiece 300mm ac linear motor 0.5um 60000mm/min 10000mm/min y y workpiece 300mm ac linear motor 0.5um 60000mm/min 10000mm/min z focus 10mm ac brushless motor 0.1um 500mm/min 250mm/min r workpiece rotation ±45° ac brushless motor 0.8mdeg 5rpm 1rpm a attenuator 0-1 stepper motor 0.01 1/sec 0.25/sec advances in systems science and applications (2011), vol.11, no.3-4 273 3.laser micro-ablation experiments 3.1 ablation rates fig. 4 shows vuv–uv transmission spectra of fused silica at room temperature. fused silica is relatively transparent when laser wavelength is larger than 180nm. but the transparency decrease dramatically as the wavelength is lower than 175nm. this means that 157nm laser beam can be better absorbed by fused silica. not only is the range of materials which can be machined enlarged by using 157nm radiation, but the higher machining quality can be achieved. in general, the ablation rate is largely affected by absorption coefficients α of materials and laser fluences f. in the 157nm laser ablation, the material removal is predominantly photon-chemical process. wavelength(nm) tr a n sm is si o n ( % ) 100 80 40 60 20 0 140 150 170160 200190180 fig.4. uv transmission of fused silica for incident laser fluences just above an ablative threshold value ft, the etch depth per pulse t is approximately: ⎥ ⎦ ⎤ ⎢ ⎣ ⎡ = tf ft ln1 α (1) for a laser pulse duration τ giving a temperature rise δt to the vaporization point of a material with density ρ, specific heat cp and reflectivity r of the surface at the laser wavelength, the ablation threshold is given by: ( )r ct f p t − δ = 1 1 ρ α (2) for photon absorption, the penetration depth of 1/α is dependent on laser wavelength, bandgap energy of materials and so on. the higher the photon absorption, the smaller the ablation threshold and etch depth per pulse (eqs. 1 and 2). as for fused silica ablated by 157nm laser, the penetration depth is about 59nm, and the ablation threshold of ft is about 1.0j/cm2. through calculation, the ablation depth is about 81nm per pulse as fluence of 157nm laser was 4j/cm2. in this study, we also measured laser ablation depth under conditions with different laser pulses and constant energy (f=4j/cm2), as shown in fig. 5. the 157nm laser spot size was about 35μm×35μm. as the number of pulses is smaller than 100, the ablation rates is 73-77nm per pulse, generally close to the calculation result. but the ablation rate reduces gradually to about 50nm per pulse as the number of pulses increase to 600[11]. that is result from energy reduction when the ablated material far from the focusing plane, even though the energy is still constant on the focusing plane. 274 xu: micro-manufacturing technology using duv laser and nc fig.5. relationship between ablation depth and number of laser pulses (f=4j/cm 2 ) 3.2 micromachining of silica glasses 157nm laser beam can be largely absorbed (α=170,000 cm-1) by fused silica, it can thus be used for 3d micro-structuring tool for silica materials. fig. 6 shows channels formed by a 20μm×20μm f2 laser beam scanning along the horizontal direction in fused silica. the size of each channel is ~120μm×20μm. the ablating conditions are: repetition rate for each column is 8hz, 12hz and 16hz, respectively; scanning velocity for each row is 0.18mm/min, 0.23mm/min and 0.28mm/min, respectively; laser fluence is ~4j/cm2, the scanning cycles are three loops. the plume dynamics in laser ablation of silica glasses generates non-volatile products, so after ablation, some of re-deposited particles scattered onto and over nearby the ablated surface, as shown in fig. 6(a). (a) (b) fig.6. the images of the channels ablated on silica glasses using 157nm laser (a) ablation debris deposited around the channels, immediately after ablation (b) the sem image after ultrasonic cleaning in hf solution advances in systems science and applications (2011), vol.11, no.3-4 275 submicron order glass particulates will collect and partially sinter on surfaces to form coatings around small structures, and growing thicker and broader with the volume of material removed by the laser. since there is no strong bonding between the re-deposited particles, the debris tends to be rather weak in structure and can be easily removed by etching process using vol. <10% hf acid. the sem image shown in fig. 6(b) is the result after 5 minutes ultrasonic cleaning in 5% hf solution. it is seen that, the channels produced by 157nm laser is relatively regular and precise. using a profiler, the depth of each channels were measured. fig. 7 shows the depth of the channels under different ablating conditions. the measured results is very close to the theoretical value. the depth of the channel can be calculated by: l f t nh v ⋅ ⋅ ⋅ = (3) l: length of laser spot along scanning direction v: shift velocity of the workpiece f: pulse repetition frequency t: etch depth per pulse n: scanning loops fig.7. the depth of the channels under different frequency and velocity fig.8. the relationship between roughness and laser parameters (frequency, velocity) 276 xu: micro-manufacturing technology using duv laser and nc we also investigated ablated surface roughness of fused silica under different process parameters. like shown in fig. 6, the short channels (4×4 array) were produced using a quadrate beam, the spot size ~35μm×25μm, the laser fluence ~4j/cm2. we changed laser pulse repetition rate from 5hz to 30hz, and the shift velocity of worktable from 0.05mm/min to 0.30mm/min. fig. 8 shows the ablated bottom roughness ra of channels for the different scanning velocity and repletion rate. it is obvious that, there exists a strong dependence of the surface quality on the laser process parameters. on the whole, when laser repletion rate increases, the averaged roughness of the channel bottom would become larger, i.e. the surface quality become worse. and, slower laser scanning velocity such as v=0.05mm/min is not recommended, because the averaged roughness shows higher value. using high velocity (v=0.2 mm/min or 0.30mm/min) and low pulse repetition rate (f=5hz), the ablated surface roughness is very low. but take into account the machining efficiency, the repletion rate of 10hz is better. from the trend of ablated roughness, the influence of the heat produced during the laser process and the cleaning process with hf solution can not be ignored. in fact, when the laser ablation was performed under higher repetition rate and lower scanning speed, the accumulation of heat become bigger at same point, so that thermo-stress would formed at bottom of the channel. after ultrasonic cleaning with hf solution, cracks resulting from the thermo-stress would be produced, like the upper-right channel shown in fig. 6. this causes the increase of bottom roughness. 3.3 micromachining assisted by numerical control nc programming is an important part of laser processing, which affects the processing quality and production efficiency[12]. in this paper, the nc programming system is alphacam which programmed by licom system ltd of england. based upon geometric graphics information, this system can automatically generate and optimize the tool-path. as the user input the process parameters on the basis of the tool-path, the system then generate laser processing nc code automatically, includes the laser fire on/off control. the flow chart of laser micro-manufacturing is shown in fig. 9. trajectory optimization is aimed at generating the reasonable laser processing tool-path, getting the shortest processing path and the least number of laser switching, and then improving the efficiency of laser processing. typically, the processing parameters include laser fluence, beam size, pulse repetition frequency, scanning velocity (for processing with scanning mode) and total pulse numbers (usually for machining at a constant position). creating toolpath determining the optimal processing parameters step and repeat mode direct writing mode generating nc code actual machining drawing 3d structure fig.9. the flow chart of laser processing advances in systems science and applications (2011), vol.11, no.3-4 277 as examples of laser nc micromachining, fig. 10 shows two microstructures fabricated by 157nm laser. fig. 10(a) is a chinese character “田”. the process tool-path was decided as from line 1 to line 6. the process parameters were: a quadrate beam with size of ~35μm×25μm, the laser fluence ~4j/cm2; scanning velocity 0.15mm/min, pulse repetition rate 12hz. fig. 10(b) is a special shaped channel ablated into a smf-28 fiber. laser scanning was performed along a path as line1 to 2, and so on. laser spot is about 10μm×10μm, and scanning velocity is 0.12mm/min. on the whole, the micromachining quality of those microstructures is considerably good. (a) (b) fig.10. the nc process tool-path and ablated images of 157nm laser micro-fabrication (a) on a silica chip (b) on a silica fiber 4. discussion fused silica has a ~9.3ev bandgap energy, equivalent to 133nm wavelength. practical lasers are not yet available at the bandgap energy (9.3ev). at 157 nm wavelength, fused silica is only weakly absorbing, α=10cm-1[10], but from our experiments and literature [10,11,13,14,15], the 157nm lasers nevertheless provide sufficiently strong interaction that smooth etched surfaces can be precisely excised, free of micro-cracks and with little debris. the inverse of the slope provides an effective single-photon absorption coefficient of αeff =170,000cm-1 — a 17,000-fold 278 xu: micro-manufacturing technology using duv laser and nc increase that attests to the importance of defect formation during such single-pulse interactions. this is due to, the 157nm output of the f2 laser, which strongly couples energy into glass via defects, dopants or near-bandedge states, provides one approach for micromachining glass. in general, germanium and other elements are doped into silica fibers or glasses. fused silica has ~8-ev bandgap energy, and bonding energy of si-o is 4.5ev [10]. the dopants lower the bandgap and provide copious new defect centers. and then, induced by irradiation of 157nm laser, many defects such as nbohcs, odcs and e’ centers were also created. the possible processes for photoinduced defect formation in the fused silica upon f2 laser irradiation can be given by: ( ) ( ) ( ) hvsi o si si o si e nbohc ≡ − − ≡ ⎯⎯→ ≡ • + • − ≡ ′ ⅰ 2 1 2 ( )( ) 2 (e ) hv si o si si si o odcs si si si ≡ − − ≡ ⎯⎯→≡ − ≡ + ≡ − ≡ ⎯⎯→ ≡ • ′ ⅱ the photoexcited states of si–o–si rings would result in dissociation into pairs of nbohcs and e’ centers [process (i)]or formation of odcs accompanied by interstitial oxygen molecules [process (ii)]. the photoinduced odc defects were also excited by f2 laser, which led to the formation of e’ centers. all of them would cause the ablation threshold was lowered, sometimes to 70%. due to the large photon energy of 7.9-ev, it is easy to break the bonding of silica glasses for 157nm laser, only by one-photon-absorption processes. in fact, when fused silica exposed to 157nm laser irradiation, ion (si+, o+) emission phenomenon and changes of ion intensities have been observed [16]. that is also evidence of photo-chemical interaction. unlike material removal by 193nm and 248nm excimer laser, which are mainly by photo-thermal effect [17], few of heat generated during ablation of 157nm laser [18,19]. however, there is still evidence of weak thermal effect in 157nm laser ablation. for example, slight color-difference resulted from heat on glass surface can be observed, as shown in fig. 6 and fig. 10. after all, the 157nm laser ablation is one-photon-absorption process for silica glasses or fibers, accompanying with weak thermal effect only. 5. conclusions a new micro-manufacturing tool -157nm duv laser is demonstrated in this paper. for 157nm laser ablation of fused silica, the material removal is predominantly photo-chemical reaction induced by one-photon absorption process. experiments and analysis show that: (1) the ablation rate of fused silica is about 75nm/pulse under laser fluence of ~4j/cm2 and spot size smaller than 35μm. but when the number of pulses increases, the ablation rate would drop down gradually. (2) take into account the machining efficiency, a better ablation quality can be achieved under the repetition rate of ~10hz. (3) 157nm duv laser processing has great potential for precision microfabrication of hard materials for practical applications. acknowledgements this work was financially supported by the national natural science foundation of china (grant no.50775169 and no.60537050). references [1] m. edward motamedi, ming c.wu and kristorfer s.j.pister. micro-opto-electro-mechanical devices and on-chip optical processing. opt. eng., vol. advances in systems science and applications (2011), vol.11, no.3-4 279 36 (1997) 1282-1297. [2] w. zhang and j. xiong. micro-optical-electromechanical-system,moems. china machine press, (2006). [3] ming c. wu. micromachining for optical and optoelectronic systems.vol.85, no.11(1997) 18333-1856. [4] fei-fan chen, ling yin and yun-long li. review and prospects of research on micro-opto-electromechanical system. microfabrication technology. no.3(2002) 1-7. [5] hiroaki misawa and saulius juodkazis. 3d laser microfabrication. wiley-vch verlag gmbh&co.kgaa, (2006). [6] j.f. li, microfabrication technology of threedimensional microdevices and their mems application. inorganic materials, 17(4) (2003) 657-664. [7] ivan fanderilik. silica glass and its application. amsterdam: elsevier(1991). [8] levy d.h. and gleason k.k. reactions of hydrogenated defects in fuled silica caused by thermal treament and deep ultraviolet irradiation. appl.phys.lett.,60(14)(1992) 1667-1669. [9] y. wang and l. liu. silica glass. chemical industry press, (2006). [10] p.r. herman, r.s. marjoribanks, et al. laser shaping of photonic materials: deep-ultraviolet and ultrafast lasers. applied surface science, 154-155(8)(2000)577-586. [11] y. dai, d. jiang and g. xu. microstructuring of photonic materials by deep-ultraviolet laser. proceedings of spie, 7284(2009)72840p1-6. [12] ryoichi kuwano, tsuyoshi tokunaga, yukitoshi otani and norihiro umeda. beam shaping optics for yag laser processing fabricated by computerized numerical control lathe. optical review vol. 12, no. 6 (2005) 476–479. [13] h. deng, y. rao, z. ran, x. liao, w. liu. photonic crystal fiber based fabry-perot sensor fabricated by using 157nm laser micromachining. acta optica sinica. 28(2)(2008)255-258. [14] y. dai, w. li and d. jiang. micro ablation of optical fibers using vacuum ultraviolet laser. proc. int. conf. integr. commer. micro nanosyst, sanya china. (2007)1353-1356. [15] j. greuters and n.h. rizvi. laser micromachining of optical materials with a 157-nm fluorine laser, proc. spie, belgium, 4941(2003) 77-83. [16] s.r. john, j.a. leraas, s.c. langford and j.t. dickinson. laser induced ion emission from wide bandgap materials. applied surface science. 253(2007) 6283–6288. [17] y. liao, y. chen, c. chao and y. liu. surface morphology and sub-surface damaged layer of various glasses machined by 193nm arf excimer laser. proceedings of spie, 5717(2005)110-117. [18] a. pusel, p. hess, et al. photochemical processing (157nm) of semiconductor surfaces without heating. surface science, 74(11)(1999)433-435. [19] vaidynathan a, walker t, et al. comparison of keldysh and perturbation formulas for one-photon absorption. j. physics review, 20b(5)(1979)3526–3527. advances in systems science and applications (2012) vol.12 no.1 54-66 the application of semantic web technology in manufacturing grid h.j.zhang1,2 and y.f.hu1,2 1the school of mechanical and electronic engineering, wuhan university of technology, wuhan 430070,china 2hubei digital manufacturing key laboratory, wuhan university of technology, wuhan 430070, china abstract concerning the knowledge-intensive environment, the paper introduces semantic web(sw) technology to manufacturing grid (mgrid). resources and services in mgrid are described in the well-defined meaning, so that computers can understand resource and service information to interoperate seamlessly. in the paper, a semantic-aware mgrid architecture (samga) based on ws-resource framework (wsrf) is first presented, in which the stateless web services are semantically described using owl-s and the stateful resource descriptions are semantically enhanced using owl. then a semantic matching algorithm is proposed for samga. finally, samga has been applied to the magnetic bearing resource and service sharing platform (mbrssp). keywords manufacturing grid, semantic web technology, semantic matching algorithm, architecture 1 introduction through the network, manufacturing grid (mgrid)[1] makes all kinds of enterprises and resources geographically distributed connected and form a virtual organization (vo), which is centralized in logic but distributed in physics. in this vo, all resources can be shared and all enterprises can be employed to work collaboratively towards to a common target of dealing with a distributed manufacturing task or problem. the kernel of mgrid is to share manufacturing resources and use services provided by manufacturing resources pellucidly. however, mgrid resources are far more diverse and complex than those of grid computing and there isn’t a unified format for describing manufacturing resources and services. there exist lack of relationships and semantics. moreover, manufacturing service composition is notoriously complex and challenging. these problems seriously hinder the development and application of mgrid. even though web services description language (wsdl) was used extensively in web service and grid at present, regrettably, it couldn’t address these problems due to lack of semantic parts. tangmuarunkit argued that existing resource description and resource selection in the grid was highly constrained[2]. due to the complexity of some manufacturing processes, some manufacturadvances in systems science and applications (2012) vol.12 no.1 55 ing tasks in mgrid should be decomposed into several subtasks, which cannot be further decomposed and can be executed by a single resource service. extensive studies have been conducted related to web service composition problem in service-oriented system and distributed system. existing research efforts for web service composition have been undertaken in two orthogonal directions: 1) manual composition, and 2) automated composition. manual composition is a time-consuming task, and it is supported by major it enterprises such as ibm and microsoft. in this approach, the web service developer selects the outsourced web services that are relevant to their composition intents and programs the interaction logic of the component services with a low level programming language such as bpel[3]. for the automated composition, the information must be understandable by computers, so that they can perform more of the tedious work involved in finding, sharing, and finally combining information in mgrid. the mutual understanding between computers is also important for the optimal composition of mgrid resources. the appearance of sw[4] brings a new hope of solving the above issues. sw is an evolving extension of the world wide web in which the semantics of information and services on the web is defined, making it possible for the web to understand and satisfy the requests of people and machines to use the web content. many researches have applied sw technologies to grid. all resources and services in grid are adequately described in the well-defined meaning that is machine-processable, which is favor of enabling computers and people to work in cooperation. in the grid system, “semantics” plays an important role in the following aspects: • more effectively discovering and managing dynamic resources in virtual organizations, such as the description and processing of semantic service, resource classification, event notification, origin tracking. the problem of heterogeneity of information in the distributed environment would be resolved at the bottom of the grid architecture; • linking and cooperating with the information stored and available within a grid automatically, so that it can support the knowledge-intensive applications at the top of the architecture. the rest of the paper is organized as follows. in section 2, some related work is summarized. in section 3, samga is presented. section 4 introduces a semantic matching algorithm which is especially designed for samga. section 5 provides the resource and service of magnetic bearing resource and service sharing platform based on samga. the conclusion is given in section 6. 56 h.j.zhang:the application of semantic web technology in manufacturing grid 2 related work early in 2001, a number of researchers were increasingly conscious of the necessity of combination of semantic web and grid. this was firstly captured by david de roure in the uk e-science program[5]. global grid forum (ggf) also established semantic grid research group (sem-grg) to realize the added value of emerging web technologies and approaches, in particular semantic web and web 2.0, for grid users and developers. there are some representative research and projects abroad, such as mygrid, combechem, geodise etc; in china, there is the national grand fundamental research 973 program: study on basic theory, model and method of semantic grid, which would mainly resolve three issues: standard resource organization, semantic interconnection and intelligent aggregation[6]. above researches are mostly applied in the environment of computing grid or data grid. there are few researches of semantic web technology for manufacturing. for instance, yang et al. presented a semantic web services approach for automated integration of manufacturing systems and services[7]. lemaignan et al. presented a proposal for a manufacturing upper ontology in order to draft a common semantic net in manufacturing domain, and applied the ontology in the automatic cost estimation and semantic-aware multi-agent system for manufacturing[8]. he et al. proposed an ontology-based manufacturing resource discovery architecture, and gave the discovery algorithm of manufacturing resources based on semantic extending and qos[9]. li et al. used ontology to describe resources and services in manufacturing grid, and extend uddi to support service publishing and retrieval[10]. it has been realized that sw have an extremely broad development prospect, but the application in manufacturing just begin, especially in mgrid. most researches have focused on the description of resource and service of mgrid, not on the semantic-enabled mgrid architecture and semantic matching algorithm. therefore, this paper proposes a new semantic-aware architecture and a semantic matching algorithm for mgrid. 3 semantic-aware mgrid architecture based on concepts and technologies of wsrf and sw, the semantic-aware mgrid architecture (samga) is designed as shown in fig.1. for the purpose of the pellucid interoperation, sw technology should be applied upon grid middleware. it is important to note that the “semantics” permeates the full vertical extent of architecture and is not just a semantic (or knowledge) layer on top: it is semantics in, on and for the grid. the function of every layer as follows: advances in systems science and applications (2012) vol.12 no.1 57 fig.1 the semaritic-aware mgrid architecture (1) fabric layer: its basic function is that local resources would be encapsulated into global resources which could be shared for the mgrid application layer. at the endpoint of physical resources node, the interfaces of lifetime management, state management and notification would be provided; web services standardize various functional operations provided by resources node and offer the standard 58 h.j.zhang:the application of semantic web technology in manufacturing grid interfaces of access services. web services would be described by using service upper ontology owl-s and manufacturing domain ontologies which are written in standard ontology language owl. a flexible and high-quality “ws-resource” would be constructed by combining standard web services with dynamic manufacturing resources. ws-resource is the extended grid service according with wsrf specification, hiding the heterogeneity of manufacturing resources and being rich in “semantics”. (2) mgrid core middleware layer: mgrid core middleware is essentially a semantic enhanced web service container, because the sw standards provide the web service container with the capabilities of processing semantic information, such as services composition monitoring, batch job processing and service level agreement (sla) management. it is the basic running environment of grid service, which will be deployed in each grid node in advance. it is recommended that mgrid core middleware is developed on globus toolkit 4.0.5 core. (3) mgrid public service layer: this is a service set which provides the common operations for mgrid applications, including information service, data service, task management and semantic service. the semantic service includes three components: • to provide the domain ontology library (knowledge workers who are a little familiar with the knowledge of manufacturing can easily use protg[11] to establish and maintain ontology library) and the general ontology library (e.g., wordnet [12], cyc[13], sumo[14]); • to provide the api of ontology libraries (e.g., jena framework is open source and grown out of work with the hp labs semantic web programme, which provides an owl api and sparql query engine[15]); • to provide the ontology reasoner (e.g., racer[16], pellet[17], fact++[18]), owl-s based matching engine[19] and owl & owl-s editor[20]. (4) manufacturing service layer: on the basis of the mgrid public service layer, the intelligent toolkits for the manufacturing applications are developed, for example, capp expert system based on artificial intelligence. (5) mgrid application layer: in virtue of the domain dependent programming model (e.g. corba, com, javabean etc.) and man-machine interaction mechanism, the man-machine interfaces (e.g. visual component portlet, servlet, etc.)are provided to mgrid users for mgrid applications. both owl-s and the web service modeling ontology (wsmo) [21] can describe services semantically. owl-s uses owl in combination with wsdl. wsmo uses f-logic to perform inferences with services, and its xml-based format gives external agents access to the service features. the paper adopts owl-s as mgrid services description language (see fig. 5). advances in systems science and applications (2012) vol.12 no.1 59 4 procedures of mgrid resource & service publication and discovery based on sw 4.1 mgrid resource & service publication elenius et al. developed an owl-s plug-in for protégé as owl-s editor [20]. the owl-s editor provides a graphical user interface to create and modify an owl-s description including all three parts: serviceprofile, servicemodel and servicegrounding. publishing sw services is more complex than web services. however, it can contribute to automated discovery, negotiation, composition, execution, and monitoring of web services. first of all, rsp develop wsdl documents of web services. then with the help of owl-s editor, rsp fill in serviceprofile information and use service–model to define service executable process. at last, rsp import the above wsdl documents into owl-s files. in order to describe mgrid services formally as much as possible, rsp can enter some keywords and browse ontolgoies that include the keywords, as well as synonyms, superclass or subclass of these ontologies, when entering iopes. information. thus rsp select the right ontology names as the iopes terms. 4.2 mgrid resource & service discover the procedures of semantic-aware service discovery are shown in fig.3. rsc just need to enter input and (or) output term(s) via the man-machine interaction interface. in the same way of resource & service publication, after that ontology reasoner performed reasoning and expansion of ontology, rsc select the right ontology names. the selected ontology names and those related terms (e.g., synonyms, superclass or subclass) are sent to mds4 as keywords. mds4 connect to all mgiis and return the service profile information to owl-s based matching engine. matching engine would extract i/o information from mds4 and then start to perform the matching algorithm (detailed in section 5) to compare i/o information between returned services and rsc request. service query interface will show the matching results (including service name and address). thus rsc connect the selected mgris to schedule this service node. 5 semantic matching algorithm the matching algorithm is the key to the implementation of the sw technology in mgrid. the algorithm is the main process of the owl-s based matching engine in the core middleware layer. the principles of the semantic matching algorithm, which are based on the semantic ontology library-wordnet, are as follows: 60 h.j.zhang:the application of semantic web technology in manufacturing grid fig.2 semaritic-aware mgrid service dicovery model (1) the smaller the semantic distance, the greater the similarity. the semantic distance between one concept and itself is 0, and their similarity is 1. when the semantic distance between concepts is infinite, their similarity is 0. in the hierarchical structure tree (see fig.3), the semantic distance refers to the length of the shortest path between two concepts. as shown in fig.3, the semantic distance of o31 to o36 is denoted by l1 +l2, where l1 refers to the distance of o31 to o11, l2 refers to the distance of o36 to o11 , and o11 is the nearest common concept of o31 and o36. (2)the more the semantic overlap, the greater the similarity. the semantic overlap refers to the level of the same meaning that both concepts involve. as shown in fig.3, the semantic overlap of o31 to o36 is denoted by l, where l refers to the distance of o11 to o0, and o0 is the top concept in the hierarchical structure tree. generally speaking, the larger the semantic overlap , the smaller the difference between the meanings of both concepts. for example, the similarity between o31 to o32 is larger than that between o11 to o12. (3) the more detailed the classification, the lower the similarity. there are fine and coarse classifications of concepts in the semantic dictionary. hence, the concept density isn’t a fixed value. there must consider the concept density in the algorithm . (4) the similarity is asymmetric. generally speaking, the similarity from o1 to o2 is unequal to that from o2 to o1. for example, the similarity from “lathe” to “machine tool” is larger than that from “machine tool” to “lathe”. because machine tool includes lathe, grinder, shaper etc. therefore, we introduce the vector curves, such as the broken blue curves with arrow in the fig.3. according to the above principles, the similarity of resource functionality from concept o1 to concept o2 is defined sim(o1 → o2) as follows: advances in systems science and applications (2012) vol.12 no.1 61 sim(o1 → o2) = l l+ α(o1, o2) l1 ρ(o) + [1− α(o1, o2)] l2 ρ(o) where α(o1, o2) = { l+l1 2l+l1+l2 , l1 ≤ l2 l+l2 2l+l1+l2 , l1 > l2 is the adjusting parameter of the asymmetry of similarity. the concept density [22] ρ(o) = m−1∑ i=0 nhypi 0.20 descendants0 ,where nhyp refers to the mean number of hyponyms per node, h refers to the height of the subhierarchy, and m refers to the number of senses in the hierarchical structure tree. there is an example of how to cacluate the similarity in the field of machine tool service. the goal of calculating the similarity is to match two resource names. the resource name stands for the resource functionality. the hierarchial structure tree in the wordnet is shown in fig.4. from the fig.4, we can obtain the data in the table 1. table 1 the similarity calculation o1 → o2 l l1 l2 m desc− α(o1, o2) nhpy ρ(o) sim(o1 endanto → o2) lathe→ 9 2 0 2 5 0.45000 1.15091 0.43018 0.81138 machine tool grinder→ 9 1 0 2 5 0.47368 1.15091 0.43018 0.89099 machine tool shaper→ 9 1 0 2 5 0.47368 1.15091 0.43018 0.89099 machine tool machine tool 9 0 2 2 5 0.45000 1.15091 0.43018 0.77874→lathe machine tool 9 0 1 2 5 0.47368 1.15091 0.43018 0.88033→grinder machine tool 9 0 1 2 5 0.47368 1.15091 0.43018 0.88033→shaper lathe→ 10 1 1 2 5 0.45000 1.00000 0.66667 0.86957 milling machine lathe→ 9 2 1 2 5 0.47619 1.15091 0.43018 0.72396 grinder in the table 1, it’s easy to see that sim(lathe → tool) > sim(lathe → grinder) and sim(lathe → machine tool)> sim(machine tool → lathe). the first inequation shows when a lathe is required, the resource node of machine tool is much better for the searching result than that of grinder. the second inequation verifies the asymmetry of similarity. we set a threshold for the similarity in the mgrid 62 h.j.zhang:the application of semantic web technology in manufacturing grid sysytem. if sim(o1 → o2) is no less than the value of threshold, we argue that o1 is similar enough to o2. fig.3 the hierarchical structure tree fig.4 an example of hierarchical structure tree in wordnet 6 application a magnetic bearing is a new type of high performance bearing. however, each type of magnetic bearings must be designed and manufactured according to the concrete objects. a large number of resources are needed in the development of magnetic bearing. in order to realize the sharing and collaborative work of all needed resources, according to samga, the magnetic bearing resource and service sharing platform (mbrssp) is developed. resource & service providers (rsps) can publish the own resources through the manufacturing resources publication center, as shown in fig.4. the publication center provides the semantic template for resource and service publication and automatically generates the semantic documents. resource & service demanders (rsds) can search the needed resources at the resource optimal-allocation center as shown in fig.5; rsds can advances in systems science and applications (2012) vol.12 no.1 63 also reserve all kinds of resources for mgrid tasks through the co-reservation system, as shown in fig.6. the semantic matching can provide the better searching result for both the optimal-allocation and the co-reservation. fig.5manufacturing resource & servicepublication center based on owl&owls fig.6 mgrid resource & service optimal-allocation center 64 h.j.zhang:the application of semantic web technology in manufacturing grid 7 conclusion this paper presented a semantic-aware mgrid architecture that exploits semantic web technologies to solve manufacturing problems. the semantic environment at the bottom of samga can adequately support the knowledge-intensive application of mgrid on the top layer. with the support of samga, a new information model suitable for the semantic and grid environment is designed, which can realize the fuzzy matching. furthermore, the information model is also called the pull model, which avoids a large number of redundant information. fig.7 mgrid resource & service co-reservation system acknowledgements this paper is supported by the national natural science foundation key project of china: digit manufacturing basic theories and key techniques under network environment (no.50335020), and the hubei digital manufacturing key laboratory opening fund project: research on resource service search and optimalselection theories and experiments in manufacturing grid system (no.sz0621). references [1] fei tao, yefa hu and zude zhou. (2008), “study on manufacturing grid and its executing platform”, international journal of manufacturing technology and management, vol.14, no.1-2, pp.35-51. [2] hongsuda tangmuarunkit, stefan decker and carl kesselman. (2003), “ontology based resource matching in the grid-the grid meets the semantic web”, proceedings of the 2nd inter. semantic web conference, pp.706-721. advances in systems science and applications (2012) vol.12 no.1 65 [3] matjaz b. juric, benny mathew, poornachandra sarang. (2004), business process execution language for web services, birmingham.u.k.: packt. [4] http://www.semanticweb.org [5] http://www.intsci.ac.cn/ai/sg.html [6] zhonghhua yang, robert gay, chunyan miao, et al. (2005), “automating integration of manufacturing systems and services: a semantic web services approach”, 31st annual conference of ieee industrial electronics society, pp.2255-2260. [7] severin lemaignan, ali siadat, jean yves dantan, et al. (2006), “mason: a proposal for an ontology of manufacturing domain”, ieee workshop on distributed intelligent systems: collective intelligence and its applications, pp.195-200. [8] yu’an he, tao yu, lilan liu, et al. (2006), “research on manufacturing resource discovery based on ontology and qos in manufacturing grid”, 2006 international conference on cyberworlds, pp.209–215. [9] jing li, xiangxu meng, shijun liu, ran yuan. (2007), “ontology-based resource description in manufacturing grid”, computer supported cooperative work in design, pp.646–650. [10] http://protege.stanford.edu [11] http://www.cyc.com [12] http://wordnet.princeton.edu [13] http://www.ontologyportal.org [14] http://jena.sourceforge.net [15] http://www.racer-systems.com [16] http://pellet.owldl.com [17] http://owl.man.ac.uk/factplusplus [18] naveen srinivasan, massimo paolucci, katia svcara. (2004), “adding owls to uddi. implementation and throughput”, proceedings of the 1st international workshop on semantic web services and web process composition.san diego, california. 66 h.j.zhang:the application of semantic web technology in manufacturing grid [19] daniel elenius, grit denker, david martin, et al. (2005), “the owl-s editor-a development tool for semantic web services”, in proceedings of the second european semantic web conference. [20] http://www.wsmo.org [21] zhizhong liu, huaimin wang, bing zhou. (2007), “a two layered p2p model for semantic service discovery”, journal of software, vol.18, no.8, pp.1922-1932 [22] eneko agirre, german rigau. (1995), “a proposal for word sense disambiguation using conceptual distance” (submited), international conference of recent advances in natural language processing. corresponding author y. f. hu can be contacted at: haijun@whut.edu.cn advances in systems science and application (2016) vol.16 no.1 72-84 the development system of independent (commercial) sector of the cultural complex of the region (on the example of the khanty-mansiysk autonomous okrug-ugra) yulia s. rod1, lina s. khromtsova1, svetlana a. yesipova2 and anna i. panenko2 1institute of management and economics, yugra state university, khanty-mansiysk, russia. 2yugra state university, khanty-mansiysk, russia. abstract the article presents the results of a study of the current state of independent (commercial) cultural complex of khanty-mansiysk autonomous okrug ugra, including analysis of statistical indicators of development of independent (commercial) cultural complex, analysis of russian and international experience of development, the sociological study of satisfaction of residents, visitors and entrepreneurs the terms of service, business practices, existing in the region in the field of culture. the aim of the study was to develop proposals for the development of independent (commercial) cultural complex of the region and its interaction with the public sector. as a result of study a system of statistical indicators that reflects the status of independent (commercial) sector of the cultural complex of ugra; the technique of “mapping” of the independent (commercial) sector of the cultural complex of ugrahave been developed. a set of activities aimed at the development of the cultural environment of the region through creation of a favorable investment climate and cooperation between the two sectors of the cultural complex of yugra state (municipal) and independent (commercial); support of innovative projects in the field of culture and the creation of “creative clusters”; the development of public-private partnership in the sphere of culture of ugra was formulated. keywords independent (commercial) sector of culture; the technique of “mapping”; systems of creative clusters; public-private partnership 1 introduction the composition of the commercial (or independent) sector of the cultural complex of khanty-mansiysk autonomous okrug-ugra (hereinafter khmao-ugra) includes cinema circuits, art galleries, organizations of show business, publishing and bookselling network and other cultural institutions, the main purpose of which is to obtain profits. state and independent sectors represent two parts of a consistent regional cultural situation. there are no insurmountable walls between them. the same advances in systems science and application (2016) vol.16 no.1 73 creative professionals sometimes work in both sectors. there are examples of cooperation between state and independent culture organizations. many culture organizations have recently had independent status. finally, new, experimental directions, which then become the constituency of the regional cultural situation and work out also in the state organizations of culture, often develop particularly in the independent sector. thus, the independent sector is an innovative resource and creative reserve of the regional culture. but its value is not limited. independent cultural institutions create their own jobs and independent cultural products and services. they make a significant contribution to the development of the regional environment, enable raising of cultural diversity in the region. they are often more flexible than public sector organization, perceive new trends more quickly, develop new technologies, respond to social needs and therefore occupy niches which are not occupied by public organizations, due to some reasons. although region authorities are not able to influence this sector directly in the same way that they affect the state organizations however, they can conduct district policy in relation to the independent sector, aimed at stimulating its growth, the maintenance of diversity and the use of its resources for further cultural development of khmaougra. a serious obstacle to the development of region policy in relation to the independent sector in the field of culture is the lack of reliable statistical data on the number, composition and activities of independent culture organizations. the role and importance of entrepreneurship development in the field of culture do not cause doubts, as evidenced by the analysis of methods and forms of state support in various countries. for instance, the policy of support to small and medium-sized enterprises in austria and germany[1,2] aims at support young entrepreneurs, venture enterprises. sectorial projects in the creative industries including the sphere of culture have a priority. in china, the government directs most of its efforts at reducing administrative barriers, while providing substantial fiscal benefits to micro and small enterprises. large-scale industrial development policy of spatial concentration in the form of free trade zones develops also[3,4]. direct support of art and culture from the public sector takes the form of subsidies, grants and awards. the allocation of funds differs among european countries in accordance with their cultural priorities. in addition, autonomous regions and municipalities make significant contributions to the culture at the local level in some countries, such as germany and poland. public financial support for culture is distributed through foundations, arts councils in some countries. lottery funds have a special role for culture in some countries (e.g. italy). 74 yulia s. rod et al: the development system of independent (commercial) sector of ... indirect support for culture is carried out in reduction of incomes associated with payment of national or local taxes. there is a general trend in many european countries towards the introduction of legal measures in the field of taxation for donations or sponsorship in the field of culture (transnational public-privatepartnership concep). there are government programs aimed at supporting organizations working in the field of culture in the russian federation and some of its constituents. thus,vgrants to support innovative projects in the field of contemporary art; grants to music organizations established by the russian federation subjects and municipal entities, as well as independent music collectives; grants to theatres under the jurisdiction of subjects of the russian federation and municipal entities and independent theatre ensembles; and other forms of support are provided in the framework of implementation of the federal target program “culture of russia (2012-2018)”. 2 materials and methods this research was carried out in the framework of the state contract for a comprehensive study of independent (commercial) cultural complex of khmao-ugra. analysis of statistical data in the field of development of (independent) commercial cultural sector in the region, analysis of russian and international experience of state support of subjects of the cultural sector, analysis of “mapping” technique and creation of “creative clusters”, sociological research have been used as the main methods of research.a large-scale survey of residents and guests of khmao-ugra (using questionnaires), as well as a survey of representatives of independent (commercial) sector of the cultural complex of the region (interviews) to identify issues and trends of its development included the sociological study held during the work. 2099 people out of 22 municipalities of khmao-ugra have been interviewed in the survey. on average, 95 people (56% of women and 44% of men) have been interviewed in each municipality.60% of people are employed and 40% do not work out of musters of respondents, almost two-thirds have children (62%). 3 the main part the state and municipal sector culture of ugra was represented by a multidisciplinary network of cultural institutions on all types of cultural activities, consisting of 466 cultural institutions in 2014. 231 public library, 112 organizations of cultural and leisure type 3 crafts establishments, 8 theatres, 34 museums, 3 parks of culture and leisure (urban gardens), 5 concert organization and 1 independent company, 1 institution of cinema and film distribution, 2 institutions for the protection of monuments of history and advances in systems science and application (2016) vol.16 no.1 75 culture and 5 other institutions act on the territory of autonomous district. 3 educational organization conducting educational activity on educational programs of secondary vocational education, 58 children’s music, art, choreographic school and art schools, which operate on the territory of municipal formations conduct educational activities in the field of culture on the territory of ugra. we conducted a grouping of small and medium enterprises of non-state (commercial) sector of culture of khanty-mansiysk autonomous okrug-ugra on activities according to statistics provided by the territorial body of federal state statistics service for khmao-ugra, budget institution of khmao-ugra “museum of geology of oil and gas”, budget institution “state library of ugra”(table 1). table 1 the number of small and medium enterprises of private (commercial) sector of culture in khmao-ugra line of activity as of rate of growth %01.01.2010 01.01.2015 manufacture of products of national art crafts 25 26 104 additional education of children (in culture) 18 22 122 film production 7 12 171 distribution of a motion picture 8 6 75 screening 9 11 122 activity in the field of art 10 12 120 activity in the field of creation work of art 2 3 150 activity in the field of art literary executive works 6 7 117 the activities of the organization and the production of theatrical and opera performances, concerts and other stage performances 15 18 120 the activities of actors, directors, composers, artists and other representatives of creative professions, acting on an individual basis 4 5 125 the activity of concert and theatre halls 1 10 1000 the activity of fairs and amusement parks 12 17 142 the activity of dance halls, discos, schools of dances 11 11 100 other entertainment activities 7 8 114 activities of libraries, archives, institutions of club type 18 20 112 the activities of museums and protection of historical sites and buildings 4 5 125 total: 157 192 122,3 thus, there are 192 subjects of small and average business of private (commercial) sector of culture on 01.01.2015 in the region. over the past 5 years the rate of growth of such institutions was 22.3%. the share of small and medium 76 yulia s. rod et al: the development system of independent (commercial) sector of ... enterprises of private (commercial) sector of culture in khmao-ugra is less than 1% of all business entities in the region. participants of the mass survey were asked to answer a series of questions, with the aim to study satisfaction level of the residents of khmao-ugra by level, quantity and range of services in the field of culture: 1.“what cultural facilities do you visit?” the most popular, according to respondents, are the palaces of culture, clubs (fig. 1) fig. 1 the level of demand for culture facilities. 2.“how often do you visit the culture facilities?” fig. 2 the frequency of culture facilities visits. half of the respondents visit cultural institutions at least 1-2 times per month (fig. 2). on answer option “other” respondents were asked for their own answers, the main ones are almost every day (7%) and do not attend at all (5%). 3.“please, specify the reasons why do not you attend or visit the culture faciladvances in systems science and application (2016) vol.16 no.1 77 ities rarely”. fig. 3 the reasons for rear visits of culture facilities. respondents cited excessive employment and “high cost” of services as the main reasons for rear visits (fig. 3). 4.“please rate the level of satisfaction with the quality of services provided by the region cultural institutions” fig. 4 satisfaction with a quality of services. two-thirds of respondents rated the level of satisfaction with the activities of cultural institutions as positive (fig. 4) 5. “please rate the level of professionalism of the region cultural institutions” more than half of respondents rated the level of professionalism as high (52,3%), a quarter of respondents rated the level as average (24.6 percent) and 14% as low. 6. “what the new culture facilities would you like to see in your community?” the respondents were offered the following answer options: amusement park; entertainment center with a skating rink; theme parks; entertainment nightclub; -imax cinema, planetarium, aquarium, dolphinarium to the last question. thus the summary of questionnaire survey of residents of the region the following conclusions have been made: 78 yulia s. rod et al: the development system of independent (commercial) sector of ... virtually all types of cultural objects are represented in the region clubs (house of culture) (49%) are the most popular among inhabitants of the region; half of the respondents visit cultural institutions at least 1-2 times a month; the range of services offered satisfied about 80% of the population; two-thirds of respondents satisfied with the level of satisfaction with the quality of services provided by the region cultural institutions (67%); more than half of respondents rated the level of professionalism of employees of cultural institutions of khmao-ugra as high (52.3%); 70% of citizens rated the level of satisfaction with the activities of cultural institutions positively. a survey of hotel and other tourist cities in khanty-mansiysk, surgut, nizhnevartovsk has been conducted to identify the degree of satisfaction the guests of ugra of the level, quality and range of services in the sphere of culture. 78 people have been interviewed just in 13 hotels in these cities. among these cultural facilities that attract visitors of our region were named as follows: archeopark, temple complexes, center of national cultures, biathlon center, the ethnographic museum under the open sky torum maa and the other. more than 23% of respondents indicated “cinemas”, 16% “temple complexes” and 11% “cultural-leisure centers” as cultural objects that attract them in the cities of ugra. preferences of region visitors in terms of cultural events were as follows: about one third of respondents are attracted to concerts and performances (concerts, cfi), 27% of respondents named sport events (biathlon) and 19.2% are attracted to exhibitions and nearly 18% festivals (the spirit of fire, rescue and save). 87.2% of all respondents answered positively on the survey question “are you satisfied with the range of services offered in the field of culture”, while 12.8% indicated that they were not satisfied by reason of the fact that there are not enough facilities incities ”where you can go with children”, there are no circuses, zoos, leisure centers. guests are most satisfied with the variety of services available in khanty-mansiysk (93.5% of the respondents), because of the large number of cultural events and the presence of a more diverse culture. guests were also asked to rate the quality of services offered in the field of culture in khmao-ugra on a five point scale. over half of respondents rated the quality of services offered in the field of culture in our region as “excellent”, indicating that all activities (especially in khanty-mansiysk) are always organized at “the highest level”. respondents who rated the quality of services as “good” and “satisfactory”, outlined the following reasons: -there are not enough places for young people; -there are not enough advertising activities and culture facilities; advances in systems science and application (2016) vol.16 no.1 79 -the lack of information in the hotel; -there are not enough facilities (museums, galleries), working in the evening. guests of our region would like to see an amusement park, new cinemas, theater, virtual museum, the circus, the museum of modern art, leisure centers, zoo as new culture facilities. meetings were held with entrepreneurs in khanty-mansiysk, surgut, nizhnevartovsk and 2 municipalities: surgut and nizhnevartovsk in the study of the independent (commercial) culture sector. analysis of the results of the interviews showed in general, underdeveloped commercial sector of culture in khmao ugra. major obstacles to business development in the field of culture, according to leaders of organizations in the industry, were as follows: the problem of underreporting of income; significant rents; the high cost of advertising; the complex of problems connected with competition and the “overflow” of clients with low incomes from one organization to another; the problem of lack of qualified personnel. 4 conclusions and suggestions 1.we developed a map, which gives an idea about the distribution of institutions and organizations independent (commercial) sector of the cultural complex of ugra based on the russian[5] and abroad[6] mapping data obtained in sociological research of management and administrative bodies of municipalities of the region, as well as information obtained from the territorial body of federal state statistics service of the khanty-mansiysk autonomous okrug ugra; budgetary institution of the khanty-mansiysk autonomous okrug ugra “museum of geology of oil and gas”; budgetary institution “state library of ugra” (fig.5). according to the results of the mapping of the organizations of small and medium-sized business of the cultural sector of khmao ugra, you can make a general conclusion, according to which commercial (independent) enterprises are not the leaders in the field of region culture. the sector is most complete and diverse developed in surgut, nizhnevartovsk, khanty-mansiysk, nefteyugansk, beloyarsk, megion and raduzhny. almost all activities are widely represented here. outsiders are berezovsky, sovetsky and kondinsky areas, as there are nocommercial enterprises in the field of culture even in region centers. 2.on the basis of domestic experience of creation of special institutions for the development of culture is currently possible to develop the following infrastructure requirements as the basis for investment development of independent (commercial) cultural sector in the region: 1) the creation of national centres of cultural development, whose activities should be focused on the disparities levelling in the quality of provision and diversity of the spectrum of cultural services for the population as a major, and 80 yulia s. rod et al: the development system of independent (commercial) sector of ... small towns of the region, as well as ensuring maximum involvement of local people in joint cultural and creative activity: fig. 5 the distribution of institutions and organizations of independent (commercial) sector of the cultural complex of ugra. 2) the creation of centres of international cultural exchange, whose activity should be aimed at community involvement in the process of intercultural integration and orientation processes of local cultural services (including support of private cultural initiatives) in the system of cultural “subregions; 3) the creation of centres for monitoring the state of activity of the organizations of the cultural sector, engaged in the maintenance and development of infrastructure to ensure the safety of these values and guaranteeing access to citizens, and replication of successful projects in the field of culture, the creation and promotion of network projects, maintaining information networks (including mapping), the promotion of small businesses. the activities of all the above centres must respond to the needs of regions, and the characteristics and priorities of the centres should be developed by the authorities of subjects of the russian federation in the field of culture. 3.a model of interaction between these sectors, including the expansion of the organizational structure of the department of culture of khmao-ugra has been developed with the aim of developing the cultural environment of the region through creation of a favorable investment climate and cooperation between the two sectors of the cultural complex of ugra state and independent (commercial) advances in systems science and application (2016) vol.16 no.1 81 sectors (fig. 6). the basic premise of creating additional structural unit (department for cooperation with the independent (commercial) sector of culture)is the lack of reliable statistical data on the number, composition and activities of independent organizations culture. the tasks of the structural unit will include: -the collection and analysis of information about independent (commercial) organizations; -consultation of creative communities; -providing information about the presence and functioning of an independent (commercial) organizations. -providing information about the forms of state support for the subjects of independent (commercial) culture sector. 4.there is a sector of activity in the independent sector that lies on the boundary of the sphere of culture and sometimes goes beyond this boundary. however, creative and cultural resources are actively used in many cases, and their products are becoming an integral part of ugra cultural environment. therefore, the following proposal relates to attract investors to the region for the construction of entertainment (shopping) centers, as well as increase the level of satisfaction of residents and guests of the region cultural sphere. these sites offer visitors not only a wide selection of products, but also services of cultural complex associated with providing of innovative resources: theme parks, amusement parks for adults and children, indoor aquarium, terrarium, zoo, cinemas, ice rinks. in addition to the shopping center you can create such cultural facilities for residents and visitors to the region as museum of local celebrities, the museum of entertaining science,museum for children, a miniature park and other. 5.the introduction of additional forms of support: subsidies from the budget of the khanty-mansiysk autonomous okrug ugra on the implementation of innovative projects in cultural sphere; the competition for the award of the government of khmao-ugra for the implementation of innovative educational projects in the field of culture; the competition for the grant of the government of kmao-ugra; on the implementation of innovative projects in the field of theater and concerts; the competition for the grant of the government of khmao-ugra on the development of innovative museum technology are suggested with the aim of further development of the cultural environment of the region and the increase of innovative activity in the sphere of culture and art. the use of regional tax concessions for entities engaged in entrepreneurial activity in the sphere of culture and art should be also provided for in the legislation of khmao-ugra. 6.it is possible to acknowledge the tremendous role of creative clusters in the development of the commercial sector of culture: the activities of the centres are wholly aimed at providing services in the field of culture, organization of leisure 82 yulia s. rod et al: the development system of independent (commercial) sector of ... and creative development of the areas population of all ages, and it is conducted on a commercial basis. we have proposed a model of the circuit of the creative cluster on the area of municipal formation of surgut, because it is a major hub of the region (table 2). table 2 “galleries”+g2:j16 educational projects industrial interactive platform oil and gas lecture studios debating club wood and paper ethnic cinema halls mass cinema creative workshop fine art art films dance documentary film the design and simulation of techniques handicraft experimental theatre platform radio journalism “virtual museums of the world” animated cartoonsphotostudios interactive platform of street culture project “wall” (graffiti) libraries branch “booksurfing” the club of sport achievements clubs of city-quests and sport orienteering board games club rooms for forums, festivals, presentations, private exhibitions related services hire of sports equipment centres of coworking timeclub hippodrome amusement park the formation of creative clusters should occur with the mandatory support of state or regional authorities, which may create favorable conditions for interaction between representatives of business and creative environment. 7.propositions for the development of public-private partnership in the sphere of culture of the khanty-mansiysk autonomous okrug ugra is based on the study of russian[7] and international experience implementing projects in the field of culture on the basis of the mechanism of state-private partnership[? ], the analysis of existing region programs and activities in the field of culture and legislation of khmao-ugra on participation in public-private partnerships, the study of satisfaction of residents and visitors to the region level, quality and range of services in the field of culture, the study views of the representatives of independent (commercial) sector of the cultural complex of ugra. in order to improve the security of the population and guests of municipalities cultural institutions (museums, libraries, cinemas, circuses, zoos), the construction of new objects of culture, reconstruction and equipment of existing cultural institutions and cultural heritage the following activities are offered: -conclusion of contracts of rent of objects of cultural significance of khmaoadvances in systems science and application (2016) vol.16 no.1 83 ugra with commercial organizations providing funding for the maintenance and repair of buildings; equipped with modern technology. so renting for long term rent movie theaters will allow to implement their modernization with modern equipment. renting museums to specialized business entities will reduce the financial costs of the region budget associated with these facilities. renting individual objects of cultural heritage of khmao-ugra to business entities will allow the financial cost of their reconstruction at the expense of investors, providing for the latter a significant reduction of the rent. -conclusion of service contracts (outsourcing) of separate functions for maintenance of buildings and premises of cultural institutions (museums, theatres, libraries); -the conclusion of concession agreements in creation (construction, reconstruction) of objects of infrastructure of cultural heritage. thus, the use of concession agreements is possible to perform repair and restoration work on cultural heritage sites of regional importance within the framework of the implementation of the state program “development of culture and tourism in the khanty-mansi autonomous district-ugra on 2014 2020”. -conclusion of investment agreements with business entities for the construction of new cultural facilities that are absent in the municipalities of khmaougra at the time of the study (the circus; zoo; folk arts and crafts enterprise, producing handicrafts in industrial scale). the use of the above forms of public-private partnership in the sphere of culture of khanty-mansiysk autonomous okrug-ugra will attract private investment in solving problems of infrastructural support of the cultural sphere of the region. references [1] 2014 sba fact sheet. (2014), austria-european comission, 2. [2] 2014 sba fact sheet. (2014), germany-european comission, 2. [3] liu, x. and lim, h. (ed.) (2008), “sme development in china: a policy perspective on sme industrial clustering”, sme in asia and globalization, eria research project report. [4] tang. (2013), facilitating sme: the chinese perspective, mekong forum, khom kaen, thailand. [5] kuzovnikova l. (2010), “historical and cultural resources mapping as a stage creative industries development”, in l.e. vostryakov (ed.) ecology of culture, moscow: dashkovik. 84 yulia s. rod et al: the development system of independent (commercial) sector of ... [6] brown j. (2003), presentation at “creative industries mapping” seminar, november 10-13, petrozavodsk. available at: http://www.cpolicy.ru/doc.plx?id=64 [7] evmenov a.d. and golubev g.m. (2012), “introduction prospects of private-public partnership concept in cultural sphere of the russian federation”, russian entrepreneurship, vol.202, no.4, pp.17-22. corresponding author yulia s. rod can be contacted at:sci.publ@gmail.com microsoft word 9-shang shaoqiang.doc issn 1078-6236 international institute for general systems studies, inc. dentability and convexity shang shaoqiang1 , suyalatu wulede2 1department of mathematics, harbin institute of technology , harbin 150001, china 2college of mathematics science, inner mongolia normal university, huhhot 010022, china email: yizhitu@163.com, suyila@imnu.edu.cn abstract in this paper, we introduced the notion of weak denting point of ( )u x (respectively, uniformly dentable) and described the characterization of reflexive very convex (respectively, uniformly convex) spaces by using the notion of weak denting point (respectively, uniformly dentable), and studied the properties of them. keywords dentability convexity denting point banach space 1. introduction and preliminaries throughout this paper, x will denote a real banach space and x will denote its conjugate space . set ( ) { : , 1}, ( ) { : , 1}, { : ( ), ( ) 1 }.xs x x x x x u x x x x x s f f s x f x x          for a convex set ,c x ext c will denote the extreme point of c . for ( )f s x and 0 ,  set ( , )f f  will denote the slice{ : ( ) : ( ) 1 }.x x u x f x    the weak topology of x is denoted by ( , )x x  and the weak  topology of x is denoted by ( , ).x x  let d be a subset of .x d is said to be dentable if for any 0  there is a x d  such that ( \ ( )) ,x co d b x   where ( ) { : , }.b x x x x x x       a point x d is said to be denting point of d if for any 0 ,  we have ( \ ( )) .x co d b x m.a.rieffel [1] first introduced the notion of dentability and proved that x has the radon-nikodym property whenever every bounded subset of x is dentable. h.b.maynard[2] improved the result of m.a.rieffel and proved that x has the radon-nikodym property if and only if x is dentable. it is known that there is a close connection between the extreme point and the denting point. for example, if m is a compact convex set in banach space ,x then x is extreme point of m if and only if x is denting point of .m in 1993, congxin wu and yongjin li [3] introduced the notion of strongly convex banach spaces and proved that if x is reflexive banach space, then x is dentable if and only if every point of ( )s x is denting point of ( ).u x this is only a result about describing the straight relations between dentability and convexity. * this research has been financially supported by the national natural science foundation of china (no.11061022). ∗ 110-122 advances in systems science and applications (2011), vol. 11,no.1-2 definition 1 let b be a bounded subset of .x a point 0x b is said to be weakly exposed point of b if there is a function ( )x s x  such that 0( ) sup{ ( ): }x x x x x b   and for any sequence 0{ } , ( ) ( ) , ( )n nx b x x x x n    imply that 0 , ( ).w nx x n  definition 2 a point 0 ( )x u x is said to be very extreme point of ( )u x if for 0  there exists a weak neighborhood v of 0 in weak topology ( , )x x  such that for any ( )f s x does not exist , ( )a b u x v  satisfying that ( ) ,{ (1 ) } ( ) ,f a b ta t b u x v      0 1 ( ) , 2 x a b  where [ 0 ,1 ].t  definition 3 let d be a subset of .x d is said to be weak dentable if for any weak neig -hborhood v of 0 in weak topology ( , ),x x  there is a vx d such that ( \( )).w v vx co d x v  a point x d is said to be weak denting point of d if for any weak neighborhood v of 0 in weak topology ( , ),x x  we have ( \ ( )) .wx co d x v  definition 4 let d be a subset of .x a point x d  is said to be weak  denting point of d if for any weak  neighborhood v of 0 in weak  topology ( , ),x x  we have ( \ ( )) .wx co d x v    definition5 ( )u x is said to be uniformly dentable. if for any 0, ( )x s x   there exists 0  such that inf{ ( ) ( ( ( ) \ ( ))) , } .xx x x co u x b x x s       definition 6 [4] a space x is said to be uniformly convex if and only if for any 0  there is a 0  such that for , ( ) ,x y s x if 1 , 2 x y   then .x y   definition 7 [5] a space x is said to be very convex if and only if for any ( ) ,x s x { } ( )nx s x and for some xx s  there holds ( ) 1, ( ),nx x n   then 0,( ).w nx x n  definition 8 [3] a space x is said to be strongly convex if and only if for any ( ) ,x s x { } ( )nx s x and for some xx s  there holds ( ) 1, ( ),nx x n   then 0 , ( ).nx x n  lemma 1 [6] x is uniformly convex if and only if for any 0, ( ),f s x   there is a 0  and some compact set c with dim 1c  such that ( , ) { : ( , ) }.f f y x d y c    lemma 2 let x be a strictly convex space. if 01 0{ } ( ), ( ) , ,n n xx u x x s x x s        0 0( ) ( ) , ( ),nx x x x n   then there exists a net 1{ } { }n nx x     such that .wx x   proof if 01 0 0 0{ } ( ), ( ) , , ( ) ( ) , ( ),n n x nx u x x s x x s x x x x n            we may assume that n mx x  for all .n m ( )u x is weak  compact set, so there exists 0 ( )x u x  such that 0x  is accumulation point of 1{ }n nx   about weak  topology ( , ).x x  we construct a family of sets 0 0 { : x x u u  is weak  neighborhood of 0x  in weak  topology ( , )}x x  and advances in systems science and applications (2011), vol.11, no.1-2 111 define a order by inclusive relation, i.e., 0 0x x u u  if and only if 0 0 , x x u u  thus we obtain a ordered set . we also construct another family of sets 0 0 1{ { } :n nx x u x u      is weak  neighborhood of 0x  in weak  topology ( , )},x x  then by zermelo axiom, there is a mapping f such that 0 0 1 1( { } ) { } .n n n nx x f u x u x         put 0 1( { } ) ,n nx x f u x       then 1{ } { }n nx x       is a net in x and .wx x   denoting 0 0| ( ) | ,r x x then 1.r otherwise, 0 1 ,r  we consider a neighborhood 0 0 0 1 { :| ( ) ( )| (1 )} 2 z z x x x r     of point 0 .x because 0| ( )| 1,nx x  there is an integer n such that for all n n there holds inequality 0 1 | ( )| , 2n r x x   hence 0 1 : | ( ) | 2n n r x x x        is finite set, furthermore, there exists a weak  neighborhood v in weak  topology ( , )x x  such that 0 1 :| ( ) | 2n n r x x x v          because of ( , )x x  is hausdorff topology. by 0 ,wx x   we know that there is a 0 such that for 0  there holds 0 0 0 1 { :| ( ) ( )| (1 )} . 2 x z z x x x r v         one hand, by 0 0 0 1 { :| ( ) ( )| (1 )}, 2 x z z x x x r        we have 0 1 | ( )| . 2 r x x    on the other hand, by 0 0 0 1 { } { }, { :| ( ) ( )| (1 )} 2nx x x z z x x x r v             and 0 1 :| ( ) | , 2n n r x x x v          we have 0 1 | ( )| , 2 r x x    which leads to a contradiction, this shows that 0 0( ) 1,x x  noticing that 0( ) 1x x  and the hypothesis that xis strictly convex, we have 0 ,x x  but 0x x  is impossible because of 0 0 0, ( ) ( ), ( )w nx x x x x x n       and 1{ } { } ,n nx x       which leads to 0 .wx x   2. main results and proof theorem 1 let x be a reflexive banach space. then x is very convex if and only if every point of ( )s x is weak denting point of ( ).u x proof we divide the proof into three parts. firstly, we will prove that if every point of ( )s x is weak denting point of ( ),u x then x is strictly convex. suppose that x is not strictly convex, then there exist three different point 0 1 2, ,x x x in )(xs such that 0 1 2 1 1 . 2 2 x x x  we choose a scalar 0l  and define an linear functional 0f on a subspace 0 1 2{ : , }x x x r      of x such that 0 1 2( ) ( ) ( )f x x l l          112 shang: dentability and convexity for some 0.  since 0x is finite dimensional subspace of ,x then there exists a real number 0m  such that 0|| || .f m by hahn-banach theorem, there exists ( )f s x such that || ||f m and 0( ) ( )f y f y whenever 0 .y x therefore we have 1 0 1 0 1 2( ) ( ) ( 0 ) ( ) 0( ) ,f x f x f x x l l l           2 0 2 0 1 2( ) ( ) (0 ) 0( ) ( ) ,f x f x f x x l l l           0 1 2 0 1 2 1 1 1 1 1 1 ( ) ( ) ( ) . 2 2 2 2 2 2 f x f x x f x x l l l                    we consider a weak neighborhood 0 0: ( ) ( ) | 2 x v x f x f x         of 0x in weak topology ( , ),x x  where v denotes the weak neighborhood of 0 in weak topology ( , ),x x  then 1 2 0, ,x x x v  hence 1 2 0, ( ) \ ( ) ,x x u x x v  it follows that  0 0( \ ( )).wx co u x x v  this contradicts that 0x is weak denting point of ( ).u x secondly, we will prove the necessity of theorem 1. by hahn-banach theorem, we know that for ( )x s x  there exists a ( )f s x such that ( ) 1.f x  if ( ) , ( ) 1, ( ) ,n nx u x f x n   then || || 1, ( ) .nx n  let , || || n n n x y x  then ( ) 1.nf y  because x is very convex (which implies strictly convex), we ( ) ( ( ) \{ })f x f u x x and , ( ).w ny x n  on the other hand, | ( ) | ,n n n nf x x f x y f y x       hence | ( ) ( )| 0, ( ). nf x f x n   this shows that x is weakly exposed point of ( )u x and f is corresponding weakly exposing function. if ( ) \ ( )y u x x v  (where v is the weak neighborhood of 0 in weak topology ( , ),x x  then there exists scalar 0m  such that ( ) ( ) ,f x f y m  hence    ( ) sup ( ): ( )\( ) sup ( ): ( ( )\( )) ,wf x m f y y u x x v f y y co u x x v       this shows that ( ( ) \ ( )) ,wx co u x x v  hence x is weak denting point of ( ).u x thirdly, we will prove the sufficiency of theorem 1. suppose that 1( ),{ } ( )n nx s x x s x    and ( ) 1, ( )nf x n  for some .xf s by the reflexivity of ,x we know that ( )u x is weak sequential compact, hence there exists a subsequence 1 1{ } { } kn k n nx x    such that , ( ) . k w nx x k  because ( )u x is closed convex set, we have ( ) ( ) ( ) , w u x u x u x  hence 1.x  on the other hand, ( ) 1.x f x   this shows that 1.x  by ( ) ( ) 1f x f x  and the fact that x is strictly convex, we have .x x furthermore, we can deduce that , ( ) ,w nx x n  this shows that x is very convex. assume the contrary,i.e., there exist weak neighborhood x v of x in weak topology ( , ),x x  and a subsequence 1 1{ } { } in i n nx x    such that . inx x v  by the assumption that every point of ( )s x is weak denting point of ( ) ,u x we have ( ( ) \ ( )).wx co u x x v  hence there is a function ( , )g x x x    which separates x and ( ( ) \ ( )) ,wco u x x v i.e., there is advances in systems science and applications (2011), vol.11, no.1-2 113 a scalar 0r  such that ( ) sup ( ( ( ) \ ( ))).wg x r g co u x x v   evidently, ( ( )\( )) , i w nx co u x x v  thus ( ) ( ) . ing x g x r  on the other hand, by the reflexivity of ,x we know that there exist a subsequence 1 1{ } { } i il n l n ix x    such that 1{ } il n lx   converges weakly to x . which contradicts that ( ) ( ) . ing x g x r  . theorem 2 x is uniformly convex if and only if ( )u x is uniformly dentable and reflexive. proof we divide the proof into three parts. firstly, we will prove that x is uniformly convex if and only if for any 0 , ( ) ,x s x    there exists 00 , ( )x s x   such that 0( ) 1,x x  then  0( , ) : , .f x x x x x x      suppose that x is uniformly convex. by lemma 1 we know that for any 0, ( ),x s x    there exist a scalar 0 and a compact setc with dim 1c  such that ( , ) :f x x x   ( , ) .d x c  let ( , ) ,x f x  then ( ) 1 .x x    we select a 00 , ( )x s x   such that 0( ) 1x x  because of uniform convexity implies reflexivity, then 0( ) ( ) ,x x x x    it follows that 0 0( ) 1 . 2 2 2 x x x x       by the assumption that x is uniformly convex, we have 0 .x x   this shows that  0( , ) : , .f x x x x x x      conversely, suppose that for any 0 , ( ) ,x s x    there exists 00, ( )x s x   such that 0( ) 1,x x  then  0( , ) : , .f x x x x x x      by lemma 1 we know that x is uniformly convex. secondly, we will prove the necessity of theorem 2. suppose that x is uniformly convex. evidently, x is reflexive. let ( ) , .xx s x x s  by we have proved above, there exist 0  such that  ( , ) : , ,f x y y x y x      i.e., if ( ) ( ) ,x x x y    then .y x   hence, for ( ) \ ( ) ,y u x b x we have .y x   (where ( ) { : , }b x y y x y x     ). we take , 2   then ( ) ( ) .x x x y      therefore, inf{ ( ) ( ( ( ) \ ( )))} .x x x co u x b x      this shows that ( )u x is uniformly dentable. thirdly, we will prove the sufficiency of theorem 2. by the reflexivity of ,x we know that for ( )x s x  there exists 0 ( )x s x such that 0( ) 1.x x  hence, there exists 0 such that 0 0{ ( ) ( ( ( ) \ ( )))}x x x co u x b x    because of ( )u x is uniformly dentable. we take   and let ( , ) ,x f x  then ( ) 1 ,x x    114 shang: dentability and convexity i.e., 0( ) ( ) .x x x x      by the inequality 0 0{ ( ) ( ( ( ) \ ( )))} ,x x x co u x b x    we have 0 0( ) { : , }.x b x x x x x x      this shows that  0( , ) : , .f x y y x y x      by we have proved above, we know that x is uniformly convex. theorem 3 if x  is strictly convex space, then weak  denting points of ( )u x are dense in ( ).s x proof firstly, we will prove that if for ( )x s x  there exists 0 ( )x s x such that 0( ) 1,x x  then x is weak  denting points of ( ).u x for any weak  neighborhood v of 0 in weak  topology ( , ),x x  there exists a scalar 0r such that for any ( ) \ ( )y u x x v    there holds inequality 0 0( ) ( ) .x x y x r   assume the contrary,i.e., 0 0( ) sup ( ( )\( )),x x x u x x v    it is obvious that there exists ( ) \ ( )ny u x x v    such that 0 0( ) ( ) 1,( ).ny x x x n    by lemma 2, we know that there is a 1{ } { }n nx y       such that ,wx x   which contradicts to that ( ) \ ( ).ny u x x v    thus  0 0( ) sup ( ) : ( ) \ ( )x x r y x y u x x v         0sup ( ) : ( ( ) \ ( ) )y x y co u x x v       0sup ( ) : ( ( ) \ ( ) ) ,wy x y co u x x v       hence 0 ( ( ) \ ( )).wx co u x x v     this shows that x is weak denting points of ( ).u x secondly, we will prove that for any ( )x s x  and 0  there is a ( )y s x  which attains its norm on ( )s x such that .y x    let ( ).x s x  for 0,  take a  0,1 such that     25 1 0 . 4 5 1        we consider the norm neighborhood 1 : 4 z z x          of x  in norm topology ( , ) ,x   then 5 1 3 1 . 4 4 z      thus z z x x x x x x z z z z z z                        1 4 4 4 1 1 4 5 1 5 1 5 1                      25 1 . 4 5 1        by bishp-phelp theorem, we know that there exist 0 0, ( )z x z s x   with  0 0 0z z z  such that 0 1 : . 4 z z z x            taking 0 0 , z y z     then  0 1y z  and ,y x    this shows that weak  denting points of ( )u x are dense in ( ).s x advances in systems science and applications (2011), vol.11, no.1-2 115 theorem4 let d be a bounded closed convex set of x .then d is weak dentable if and only if for any weak neighborhood v of 0 in weak topology ( , ),x x  there is a slice  ( , , ) : , ( ) sup ( )s x d y y d x y x d       and 0 ( , , )x s x d such that 0( , , ) .s x d x v   where 0  . proof sufficiency. for any 1 2, , , nf f f x  and 0 ,  let   1 : ( ) | . n i i v x f x     then, by the conditions given here, we know that there exist a slice 0 ( , , )x s x d and 0 ( , , )x s x d such that 0( , , ) ,s x d x v   it follows that 0( ) sup ( ) .x x x d    noticing that  0\ ( ) \ ( , , ). : , ( ) sup ( )d v x d s x d y y d x y x d         and  : , ( ) sup ( )y y d x y x d     is weak closed convex set of weak topology ( , ),x x  we know that  0( \ ( )) : , ( ) sup ( ) ,wco d x v y y d x y x d       it follows that 0 0( \ ( )) ,wx co d x v  this shows that d is weak dentable. necessity. suppose that d is weak dentable. for any 1 2, , , nf f f x  and 0 ,  let 1 : ( ) | . 2 n i i v x f x          then, there a is 0x x such that 0 0( \ ( )) .wx co d x v  hence there is a function x x  which separates 0x and 0( \ ( )) ,wco d x v i.e., there is a scalar 0r  such that 0 0( ) sup ( ( \ ( ))).wx x r x co d x v    let 0sup ( ) ( ) ,x d x x r     then 0( ) sup ( ) sup ( ) ,x x x d r x d        this shows that 0 ( , , ) .x s x d furthermore, for any ( , , ) ,y s x d we can deduce 0 0( ) sup ( ) sup ( ) sup ( ) ( ) ( ) ,x y x d x d x d x x r x x r             which leads to 0 .y x v  otherwise, 0 ,y x v  furthermore 0 0\( ) ( \( )).wy d x v co d x v    by the inequality 0 0( ) sup ( ( \ ( ))),wx x r x co d x v    we have 0( ) ( ).x x r x y   a contradiction. theorem5 let x be a reflexive banach space. if every point of ( )s x is weak denting point of ( ) ,u x then every point of ( )s x is weakly exposed point of ( ).u x proof suppose that 0 ( )x s x is weak denting point of ( ) ,u x by hahn-banach theorem we know that there is a function 0 ( )x s x  such that 0 0( ) 1.x x  we will prove that if 0x attains its norm at another point 0 ( ) ,y s x then 0 0 .x y otherwise, 0 0 ,x y it is obvious that    0 0 0 0 0 1 1 1 1, 2 2 x y x y x     thus  0 0 1 1. 2 y x  by hahn-banach theorem we know that there is a function 0 ( )y s x  such that    0 0 0 0 .y y y x  we consider a weak neighborhood 0 0 0 0 0 0 0 | ( ) ( ) |1 : ( ( )) 2 4 y y y x v x y x y x            of point 0 0 1 ( ) , 2 y x 116 shang: dentability and convexity it is clear that 0 0, ,x y v thus  0 0, \ ,x y u x v furthermore  0 0 1 ( ) ( \ ), 2 wy x co u x v  this shows that  0 0 1 2 y x is not weak denting point of ( ) ,u x which contradicts to the hypothesis that every point of ( )s x is weak denting point of ( ).u x if 1 0 0 0{ } ( ) , ( ) ( ) 1, ( ) ,n n nx u x x x x x n        then, by the hypothesis that x is reflexive space we know that there exist 1 1{ } { } kn k n nx x    and ( )x u x such that ,( ). k w nx x k  hence 0 0 0 0( ) ( ) ( ) 1, ( ) , knx x x x x x k      it follows that 1,x  this shows that ( ).x s x by the proof above, we have 0 ,x x thus 0 , ( ) k w nx x k   and we can deduce that 0 , ( ).w nx x n  if 1{ }n nx   does not converges weakly to 0 ,x then there exist a weak neighborhood 1v of 0 in weak topology ( , )x x  and a subsequence 1{ } in ix   of 1{ }n nx   such that 1 0 1{ } ( ) . in ix x v     noticing that 0x is weak denting point of ( )u x , we know that 0 0 1( ( ) \ ( )).wx co u x x v  hence there is a function x x  which separates 0 1( ( ) \ ( ))wco u x x v and 0 ,x i.e., there is a scalar 0r  such that  0 0 1sup ( ( ( ) \ ( ))).wx x r x co u x x v    noticing that 1 0 1{ } ( ) , in ix x v    we have 0 1 0 1( ) \ ( ) ( ( ) \ ( )). i w nx u x x v co u x x v    by the inequality 0 0 1( ) sup ( ( ( ) \ ( ))),wx x r x co u x x v    we have 0( ) ( ). inx x r x x   on the other hand, by we have proved above, there exists a subsequence 1{ } jn jx   of 1{ } in ix   such that 1{ } jn jx   converges weakly to 0 ,x which leads to    0 0 .x x r x x   a contradiction. hence 0 ( )x s x is weakly exposed point of ( ).u x theorem6 let x be a separable reflexive banach space. if 0 ( )x u x is extreme point of ( ) ,u x then 0 ( )x s x is weak denting point of ( ).u x proof suppose that 0 ( )x u x is extreme point of ( ).u x if 0 ( )x s x is not weak denting point of ( ),u x then there exists a weak neighborhood v of 0 in weak topology ( , )x x  such that 0 0( ( ) \ ( )).wx co u x x v  noticing that 0( ( )\( )) ( )wco u x x v u x  and the assumption that x is reflexive banach space, we know that 0( ( ) \ ( ))wco u x x v is weak compact set. by krein-milman theorem, we have 0 0( ( ( ) \ ( ))) ( ( ) \ ( )),w w wco ext co u x x v co u x x v   hence 0 0( ( ) \ ( )) ( ) \ ( ) . wwext co u x x v u x x v   in fact, by the assumption that x is separable banach space, we know that weak topology ( , )x x  is metrizable space, hence, for any 0( ( ) \ ( )) ,wx ext co u x x v  there exists a sequence 1{ (1 ) }n n n n nt x t y    such that  (1 ) , .w n n n nt x t y x n    where   00,1 , , ( ) \ ( ).n n nt x y u x x v   by the reflexivity of ,x we know that there exists a sequence    in n such that 0 0, , i i w w n nx x y y  0, ( ). i w nt t i  it is obvious that  0 0 0 00,1 , , ( )\( ) w t x y u x x v   and (1 ) , i i i i w n n n nt x t y x   advances in systems science and applications (2011), vol.11, no.1-2 117 thus  0 0 0 01 .x t x t y   case (i): if 0 0,t  then 0 0( ) \ ( ) ; w x y u x x v   case (ii): if 0 1,t  then 0 0( ) \ ( ) ; w x x u x x v   case (iii): if  0 0,1 ,t  then, by 0 0 0, ( ) \ ( ) w x y u x x u  0( ( ) \ ( ))wco u x x v  and 0( ( ) \ ( )) ,wx ext co u x x v  we have 0 0 0( ) \ ( ) . w x x y u x x v    combining case (i),(ii) and (iii), we have 0 0( ( ) \ ( )) ( ) \ ( ) . wwext co u x x v u x x v   because   0 0( ( ( ) \ ( ))) ( ( ) \ ( )) ,w wext u x co u x x v ext co u x x v   we have 0 0 0( ( ) \ ( )) ( ) \ ( ) . wwx ext co u x x v u x x v    this is a contradiction. theorem7 if 0 ( )x s x is weak denting point of ( )u x , then 0x is very extreme point of ( )u x . proof suppose that 0 ( )x s x is not very extreme point of ( ),u x then there exist 0 0 ,  0 ( )f s x  such that for any neighborhood v of 0 in weak topology ( , ),x x  there exists , ( )a b u x v  satisfying that     0 0 1 ( ),| | , 1 ( ) , 2 x a b f a b ta t b u x v        where  0,1 .t  we consider a weak neighborhood 0 0: | ( ) | 6 v x f x       of 0 in weak topology ( , )x x  and a family of sets  v (where v is balanced convex neighborhood of 0 in weak topology ( , ),x x  then weak topology ( , )x x  has a neighborhood base  v v  of 0. for any ,v v  there exist , ( ) ( )a b u x u v     satisfying that 0 0 1 ( ), ( ) , 2 x a b f a b        { (1 ) } ( ) ,ta t b u x v v       where [0,1].t let ,a a a b b b            (where , ( ) , ,a b u x a b v v          ), then 0 0 0 0 0 1 1 1 ( ) ( ) , ( ) | ( ( )) | , 2 2 2 2 x a b a b f a x f a b                  0 0 0 0 0 0 0 0( ) ( ) | ( ) | . 2 6 3 f a x f a x f a a               noticing that ( )u x is convex set and v v  is balanced convex set, we know that 1 1 1 ( ) ( ) , ( ) , ( ) . 2 2 2 a b u x a b v v a b v v                    let 0 0 0: ( ) . 3 v x f x       it is obvious that 0 0 .a x v   similarly, 0 0 .b x v   hence 118 shang: dentability and convexity 0, ( ) \ ( ) ,a b u x x v     furthermore, 0 0 1 ( ) ( ( ) \ ( )). 2 a b co u x x v     on the other hand, 0 1 1 ( ) ( ) 2 2 a b x a b          and 1 ( 2 a  ) ,b v v    so we have 0 1 ( ) . 2 a b x v v       because  v v  is neighborhood base of 0 in weak topology ( , ),x x  from the definition of weak closure about weak topology ( , ),x x  we know that 0 0 0( ( ) \ ( )).wx co u x x v  which contradicts that 0 ( )x s x is weak denting point of ( ).u x theorem8 let x be a reflexive separable banach space. x x is a banach space with the norm  ( , ) max , .x y x y if 1 2,x x are weak denting points of ( ) ,u x then 1 2( , )x x is weak denting point of ( ) ( ) ( ).u x u x u x x   . proof suppose that 1 2,x x are weak denting points of ( ) .u x by theorem 7 we know that 1 2,x x are very extreme points of ( ).u x it is obvious that 1 2,x x are extreme points of ( ).u x we will prove that 1 2( , )x x is extreme point of ( ) ( ).u x u x if 1 2 1 2( , ), ( , ) ( ) ( )y y z z u x u x  and 1 2 1 2 1 2 1 1 2 2 1 1 1 1 1 1 ( , ) ( , ) ( , ) ( , ), 2 2 2 2 2 2 x x y y z z y z y z     then  1 1 1 2 2 2 1 1 ( ), . 2 2 x y z x y z    hence 1 1 1 2 2 2, .x y z x y z    this shows that 1 2 1 2 1 2( , ) ( , ) ( , ).x x y y z z  therefore 1 2( , )x x is extreme point of ( ) ( )u x u x . by theorem 6, 1 2( , )x x is weak denting point of ( ) ( ) ( ).u x u x u x x   theorem9 let x be a very convex space and x x is a banach space with the norm  ( , ) max , .x y x y if 1 1, ( ) ,x y s x then 1 1( , )x y is weak denting point of ( ).u x x if 1 11, 1,x y  then 1 1( , )x y is not weak denting point of ( ).u x x proof firstly, we will prove that ( , ( ) ) ( , ) ( , ).x x x x x x x x        for 0, ( ) ,f x x     we consider a eighborhood{( , ) : ( , ) }x y f x y  of (0,0) in weak topology ( ,( ) )x x x x   and define 1 2( ) ( ,0), ( ) (0, )f x f x f x f x  for any ,x x then ( )if x f x  and ( 1,2)if i  are linear functions. this shows that *, ( 1,2).if x i  we construct a set 1 2: ( ) : ( ) , 2 2 g x f x x f x                then g is open set in ( , ) ( , ).x x x x   for any ( , ) ,x y g we have 1 2( , ) ( ,0) (0, ) ( ) ( ) , 2 2 f x y f x f y f x f y          hence {( , ) : | ( , ) | }.g x y f x y   this shows that topology ( , ) ( , )x x x x   is smaller than topology ( , ( ) ).x x x x   for * 1 20, , ,f f x   we also consider the above open setg in topology and define advances in systems science and applications (2011), vol.11, no.1-2 119 1 2( , ) ( ) , ( , ) ( )f x y f x g x y f x  for ( , ) ,x y x x  then 1 1( , ) . ( , ) ,f x y f x f x y  2 2( , ) . ( , )g x y f x f x y  and ( , ) , ( , )f x y g x y are linear functions. this shows that , ( ) ,f g x x   we consider a neighborhood ( , ): ( , ) ( , ): ( , ) 2 2 x y f x y x y g x y               of (0,0) in weak topology ( ,( ) ),x x x x   then ( , ) : ( , ) ( , ) : ( , ) . 2 2 x y f x y x y g x y g                this shows that topology ( ,( ) )x x x x   is smaller than topology ( , ) ( , ).x x x x   this completes the proof that ( ,( ) )x x x x   ( , ) ( , )x x x x    . secondly, we will prove that if 1 1, ( ),x y s x then 1 1( , )x y is weak denting point of ( ).u x x if 1 1, ( ) ,x y s x then, by hahn banach theorem we know that there exists 1 ( )f s x such that 1 1( ) 1.f x  we choose ( )nx u x with 1,nx  such that  1 1 1( ) ( ) 1, .nf x f x n   by the assumption that x is very convex space, we have 1, ( ) .w nx x n  it follows that for any weak neighborhood v of 0 in weak topology ( , ),x x  there exists 1 0r  such that for any 1( ) \ ( )x u x x v  there holds 1 1 1 1( ) 1 ( ) .f x f x r   otherwise, there exists 1( ) \ ( )nx u x x v  such that 1( ) 1,( ),nf x n  which leads to 1,( ).w nx x n  this contradicts that 1( )\( )nx u x x v  .similarly, there exists 2 ( )f s x such that 2 1( ) 1f y  and for any weak neighborhood v of 0 in weak topology ( , ),x x  there exists 2 0r  such that for any   1\ ( )y u x y v  there holds 2 1 2 2( ) 1 ( ) .f y f y r   for any ( , ) ,x y x x  let 1 2( , ) ( ) ( ) ,f x y f x f y  then f is linear function on ,x x and 1 2 1 2( , ) ( ) ( ) . .f x y f x f y f x f y    2 2 1 22 f f  2 2 x y  2 2 2 2 1 2 1 24 max , 4 ( , )f f x y f f x y    . this shows that ( ) .f x x   we consider any weak neighborhood w of (0,0) in weak topology ( ,( ) )x x x x   and notice that 1 1, ( ) ,x y s x then, by ( ,( ) )x x x x   ( , ) ( , ),x x x x    we know that there exists a weak neighborhood v of 0 in weak topology ( , )x x  such that 1 1 1 1( ) ( ) ( , ) .x v y v x y w     for 1 1( , ) ( ( ) ( ) ) \ (( ) ( )),x y u x u x x v y v     we have 1( ) \ ( )x u x x v  or 1( ) \ ( ).x u x y v  without loss of generality, we may assume 120 shang: dentability and convexity that  1( ) \ ,x u x x v  then 1 1 1 1 1 2 1 2 1( , ) ( , ) ( ( ) ( )) ( ( ) ( )) .f x y f x y f x f x f y f y k      thus 1 1 1( , )f x y k   1 1sup ( , ) : ( , ) ( ) ( ) \ ( ) ( )f x y x y u x u x x v y v      1 1sup ( , ) : ( , ) ( ) ( ) \ (( , ) )f x y x y u x u x x y w     1 1sup ( , ) : ( , ) ( ( ) ( ) \ (( , ) ))f x y x y co u x u x x y w     1 1sup ( , ) : ( , ) ( ( ) ( ) \ (( , ) )) .wf x y x y co u x u x x y w    hence 1 1 1 1( , ) ( ( ) ( ) \ (( , ) )).wx y co u x u x x y w   obviously, ( ) ( ) ( ),u x x u x u x   which leads to 1 1 1 1( , ) ( ( ) \ (( , ) )) ,wx y co u x x x y w   this shows that 1 1( , )x y is weak denting point of ( )u x x . thirdly, we will prove that if 1 11, 1,x y  then 1 1( , )x y is not weak denting point of ( ).u x x case (i). if 1 1( ), , (0,1),x s x y a a   then, there exist 1 1 2 2 2 2 1 , 1 3 3 a a y y y y a a                   such that 1 1 1 1 2 2 (1 ) 1, (1 ) 1. 3 3 y y y y a a y y y y a a                 by hahn-banach theorem, there exists *f x such that 1 2 2 2 2 ( ) 1, ( ) 1 , ( ) 1 , 3 3 a a f y f y f y a a        hence 1 1 2 2 2 2 ( ) , ( ) 3 3 a a f y y f y y a a       and 0 2 2 . 6 a a    we consider a weak neighborhood 1 0{ : | ( )| }x f x y   of 1y in weak topology ( , ) ,x x  then 1 0{ : | ( )| }x f x y   can be written by 1 0 1{ : | ( )| } ,x f x y y v     where v  is a weak neighborhood of 0 in weak topology ( , ).x x  evidently, 1, ,y y y v    it follows that 1 1 1 1( , ),( , ) ( ( ) ( ) \ ( )( )).x y x y u x u x x v y v       by ( ,( ) ) ( , ) ( , ),x x x x x x x x        we know that there exists a weak neighborhood v of 0 in weak topology ( , )x x  such that 1 1 1 1( , ) ( ) ( ) .x y v x v y v      noticing that 1 1 1 1 1 1 ( , ) ( , ) ( , ) 2 2 x y x y x y   and the equality ( ) ( ) ( ) ,u x x u x u x   we have 1 1 1 1 1 1( , ) ( ( ) ( ) \ ( ) ( )) ( ( ) \ (( , ) )) ,x y u x u x x v y v co u x x x y v         furthermore, 1 1 1 1( , ) ( ( ) \ (( , ) )).wx y co u x x x y v   hence, 1 1( , )x y is not weak denting point o ( ).u x x case (ii). if 1 10, ( ),y x s x  let 1 ( ) , . 2 z u x z  by hahn-banach theorem, there exists *f x such that ( ) 1.f z  we consider a neighborhood 1 : ( ) 2 v x f x        of 0 in weak advances in systems science and applications (2011), vol.11, no.1-2 121 topology ( , ),x x  then , .z z v  which leads to 1 1 1( , ) , ( , ) ( ( ) ( ) \ ( ) (0 )).x z x z u x u x x v v       similar to case (i), there exists a v such that 1 1( ,0) ( ) (0 ).x x v v     noticing that 1( ,0)x  1 1 1 1 ( , ) ( , ), 2 2 x z x z  we have 1 1 1( ,0) ( ( )\( ) (0 )) ( ( )\(( ,0) )).w wx co u x x x v v co u x x x v         this shows that 1( ,0)x is not weak denting point of ( ).u x x . theorem10 let x be a strongly convex space. x x is a banach space with the norm  ( , ) max , .x y x y if 1 1, ( ) ,x y s x then 1 1( , )x y is denting point of ( ).u x x if 1 11, 1,x y  then 1 1( , )x y is not denting point of ( ).u x x proof case (i). 1 1, ( ).x y s x using a similar method to that used in the proof of theorem 9, we can prove that 1 1( , )x y is denting point of ( ).u x x case (ii). 1 11, 1.x y  by the assumption that x is strongly convex space, we know that x is very convex space. by theorem 9, we know that 1 1( , )x y is not weak denting point of ( ) ,u x x so there exists a weak neighborhood v of (0,0) in weak topology ( , ),x x  such that 1 1 1 1( , ) ( ( ) \ ( , ) ).wx y co u x x x y v   it is obvious that there exists a scalar 0  such that 1 1 1 1( , ) ( , ) .b x y x y v   hence 1 1 1 1 1 1( , ) ( ( )\ ( , )) ( ( )\ ( , )).wx y co u x x b x y co u x x b x y     this shows that 1 1( , )x y is not denting point of ( ).u x x references [1] m.a.rieffel. dentable subsets of banach spaces with applications to a radon-nikodym theorem, funct.anal.proc.(conf. irvine, calif., 1966), acad.press.london. thompson.wash -ington, d.c. 71-77. [2] h.b.maynard. a geometrical characterization of banach spaces having the radon-nikodym property. trans.amer.math.soc, 1973, 185: 493-500. [3] c.x.wu, y.j.li. strong convexity in banachspaces. chin.j.math, 1993, 13(1): 105-108. [4] j.a.clarkson. uniformly convex spaces. trans.amer.math.soc, 1936, 40: 396-414. [5] tegusi, suyalatu, y.j.li. very convex banach spaces. northeast.math.j, 1997, 13(1): 1-4. [6] x.n.fang. slice and convexity, smoothness of banach spaces. chin.j.math, 1999, 19(3): 293-298. [7]xu xu,tai kuang. q-operator for multi-attribute decision-making with incomplete informati-on. advances in systems science and applications, 2006, 1: 148-154. [8] long zhou, mianyun chen. investigation on fussy degree criterion of electric power equipm-ent. advances in systems science and applications, 2004, 2: 319-324. [9] jinghong pan, xuemou wu. a pansystem topology approach grayscale image processing. advances in systems science and applications, 2003, 2: 151-156. 122 shang: dentability and convexity мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 89-99 modeling the outbreak of an infectious disease on a heterogeneous network m. hajizadeh 1 , f. r. vishkaie 2 , f. bakouie 3 , s. gharibzadeh 3 1 faculty of biomedical engineering, amirkabir university of technology (tehran polytechnic) tehran, iran 2 faculty of electrical engineering, shahid beheshti university, tehran, iran 3 institute for cognitive and brain sciences, shahid beheshti university, tehran, iran annotation the outbreak of infectious diseases is a global public health threat for the international community. modelling the propagation of epidemics in a society is one of the important fields in epidemiology science. it is essential to know the number of infected cases for estimating and controlling the spread of disease in the affected countries. in this study, we used complex network theory to model the spread mechanism of epidemic disease in social networks. we modeled a social complex network by graph theory. individuals are considered as nodes and acquaintances between them are considered as links. disease virus can transmit along the links between nodes (people) according to different situations. in this work, we proposed a dynamic model for simulating the outbreak of infectious disease on a social network based on the susceptible, exposed, infected and recovered (seir) dynamical categories. it has been tried to study the heterogeneity on the network by considering two key factors in the epidemic propagation: 1) the communications weights between individuals in the network 2) different body resistances of people based on age. the proposed dynamic model was applied on a real social network which was constructed in our previous research. we compared the proposed model with two different dynamic models. finally, the simulations were compared with the reported data of infected cases of sars outbreak in hong kong in 2003. the results indicated some similarity between our proposed model and the real reported data. based on the results, it could be concluded that considering communications weights and body resistances of people captures the dynamic of disease spread in a proper way. key words: social network, infectious disease, heterogeneity, body resistance, communications weights. 1 introduction complex network theory could help to understand the spread mechanism of epidemic disease in social networks [1, 2]. graph theory is a powerful tool for modeling complex networks [3-6]. in graph theory, individuals are considered as nodes and acquaintances between individuals are considered as links [7]. virus transmission can occur along the existing links according to different situations. some studies indicated that human interactions are important for the spread of infectious disease [8]. in the real-word, social networks links strengths aren’t equal and some interactions have a high probability of infection transmission than others, i.e. the weight of links shows the amount of communications between individuals [7]. prevalence of infectious diseases such as influenza, meningitis, pertussis, yellow fever and etc., since the beginning of history has been a major concern of humanity. some of them are 90 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… widely spread in the last decade: severe acute respiratory syndrome (sars) in 2002, methicillin-resistant staphylococcus aureus (mrsa) in 2005 and ebola in 2014. several types of research have studied the spread of infectious disease in social networks [913]. some of them, considered the complex network approach in their researches. christian l. althaus presented mathematical modeling that was appropriate for estimate the outbreak of ebola in west africa in 2014. he applied seir model in the simulations [14]. abdulrahman et al. presented and analyzed a model for controlling the spread of ebola virus disease (evd) in a population by slir dynamic model. they defined 12 parameters in their model such as rate of public enlightenment, enhanced personal hygiene due to public enlightenment, number of quarantined individuals and availability of isolation centers. their simulations showed that improved personal hygiene and quarantining of infectious individuals are enough to control the spread of evd [15]. small and tse modeled transmission of sars in hong kong with complex small world network. they used seir dynamic model. they supposed different probabilities of transition between the states of disease in the network. in their model, random nodes are isolated in order to control the spread of infection [16]. colizza et al. used real airline networks data for their study because they believed that people’s transportations are causes of sars disease propagation. they considered the population of each city is classified into seven different compartments, and hospitalized as well as infectious individuals are able to transmit the infection. they concluded that their models are fit for the forecast and analysis of emerging disease spreading at the global level [17]. some scientists have attempted to produce their database through questionnaires [18]. previously, we recorded the communications between 100 persons during one week and constructed a sample of a social network consisting human communications. it was shown that the constructed social encounters network has a small-world topology. in the next step, the common cold outbreak was modeled with sir dynamic model [1]. in this study, we used our previous social network for modeling sars propagation by seir model. in a disease spread network, there are some sources of heterogeneity. for example, individuals’ characteristics against the disease and the level of interaction between people are some bases of inhomogeneity. in this paper, we tried to apply this heterogeneity factors in the simulations. for this purpose, we proposed a dynamic model with two important features: 1) communications levels between individuals which is reflected in the connections’ weights of the network 2) different body resistances based on ages of people which is used as different thresholds for each node (people). our suggested model considers cumulative weights of all links between each node and its infected neighbors. we named this model as cumulative weight model (cw). this model was compared with two different dynamics. the first one is abramson and kuperman’s model (ak) that it is based on the number of infected neighbors for each node [19]. this model doesn’t consider the amount of communication between persons. the second model is rajabi vishkaie et al. model (r) that it’s based on the largest connection weight between one node and its infected neighbors [1]. at the end, the simulation results of these models were compared with reported data of infected cases of sars outbreak in hong kong in 2003 [20]. 2 methods dataset: previously, for data gathering, we used the questionnaire distribution method in a small social network. in order to construct the network structure, 15 participants were asked in the questionnaires to record their encounters in a week. moreover, they were asked to record their contact time weekly. based on the level of contact between the persons, we assigned different amounts for edge weight in integer values from 0 to 9: more time of weekly contacts receives more scores. advances in systems science and application(2016) vol.16 no.4 91 the questionnaires were distributed among the author’s acquaintances. these participants were asked to distribute the questionnaires between their own acquaintances (exactly those people that the participants mentioned in their own questionnaires). in this way, the network was expanded. all the participants were asked to record their encounters with their acquaintances; also, we requested participants to record the encounters between their acquaintances. there were 15 participants. the number of all people that they recorded was 100. the produced network (database network) had 326 links between persons [1]. in this study, since we used our previous database, the run time must be 7 days but it isn’t long enough to see the propagation of disease. hence, the run time must be a multiple of 7. therefore, the run time was chosen 98 days. (we considered the run time long enough to see the changes.) 2.1. basic definitions in disease transmission the dynamic of the disease spread governed by seir model (that ‘s’ stands for susceptible, ’e’ shows exposed persons, ‘i’ shows infected person, and ‘r’ indicates recovered or deceased person). it has reported that sars has incubation period between 2 to 7 days (or longer) [21]. in our model, we supposed the incubation period is 7 days ( ). this means that when the sars’s virus goes inside a person body (s)he goes to exposed state (e) while (s)he doesn’t have disease symptoms in this period (first week). after this incubation period, the person goes to infected state (i) which has disease symptoms such as fever, cough etc. here, we assumed the symptoms period is about 7 days. since cdc (centers for disease control and prevention) recommends that persons with sars limit their interactions outside the home until 10 days after their respiratory symptoms have gotten better [22], we considered infection period ( ) as 17 (10+7) days which each infected person could carry the virus. after the infected period (s)he goes to recovered state (r) in which (s)he can’t transmit the infection anymore (see fig.1). fig.1 schematic of state transition by seir dynamic model for sars disease, this dynamic model has four state for a society that has infectious disease: s (susceptible), e (exposed), i (infected) and r (recovered) in our network, we assumed that at the start time of the simulations, sars viruses infected 1% of the population and before this time, nobody is sick. we defined two time-based vectors: where state(t) is a vector that indicates the conditions of each node (individuals) and n is the number of network’s nodes. we also defined probability vector as: where (t) is a vector that i ndicates the likelihood of getting illness for the node i at time t. generally, if the interaction between two people increases, the probability of getting illness ( (t)) will increase, too. the states of nodes in time step t+1 (state (t+1)) will change according to: (t+1) = 0 (s) ; 1 (e) ; 2 (i) 3 (r) http://www.cdc.gov/ 92 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… one of the important factors in the outbreak of epidemic disease is the probability of virus transmission from each infected nodes to its neighbor nodes (q). we assumed that q is between 0 and 1 with a normal distribution. another factor in propagation is the threshold (th). we considered it as body resistance against the disease. if a person has more body resistance, (s)he will have a higher threshold. in this study, at first, in some simulations we considered th as a fixed parameter between 0 and 1 to all the individuals for obtaining results in different situations. since the disease mortality depends on age [23], so body resistance against the virus may be different according to the people’ age. hence, in some simulations, we assigned different th to the network nodes based on the ages. 2.2. transmission dynamics we applied three different models for the dynamic of the disease spread: 1) ak model. this model is proposed by abramson et al. and it was used for modeling infectious disease [19]. this model has expressed based on the number of the infected neighbors. where is the number of the infected neighbors for the node i at time t, and q is the probability of transmission from each infected neighbor to node i and it is a random quantity between 0 and 1. (t) is the probability of becoming ill for node i at time t. in simulations if (t) is more than th, node i will go to the e state. (th is the body resistance against the disease for each node.) 2) r model. this model is proposed by rajabi et al. and it was used for modeling common cold outbreak [1]. where is the weight of the link between the node i and one of its infected neighbors. if the node i has several infected neighbors, simulations will check all of the links and if it get a connection which has a (t) more than th, the contagion transmits to the node i and it will go to the e state. (th is the body resistance against the disease for each node.) in fact, one dangerous connection (a link that its weight is more than probability *10) is sufficient to transmit the infection and the other connections will be ignored. 3) cw model. the amount of interactions between persons (w) and the number of infected neighbors have the main roles in disease spread. as the number of infected neighbors increase, the probability of getting illness will be increased too. also, the probability of transmission from each infected neighbor to a susceptible node (q) is significant. in order to consider these factors and the heterogeneities on the network, we changed eq. (4) (ak model) to eq. (6). this equation comprised cumulative weights of links to each node and its infected neighbors: where (t) is the probability of becoming ill for i node at time t. is the sum of the edges weights between the node i and its infected neighbors at time t. if (t) is more than th, node i will go to the e state. (th is the body resistance against the disease for each node.) 3 results the diagrams of the susceptible, exposed, infected and recovered persons for three models is illustrated in fig.2. in these three diagrams, the total number of susceptible, exposed, infected and recovered persons is equal to 100 (number of network’s vertices) at each time. for all 3 advances in systems science and application(2016) vol.16 no.4 93 dynamics, it was indicated that the number of exposed people is 1 from the first day until 7 th day. during this period, the number of infected people is 0. fig.2 number of susceptible, exposed, infected and recovered persons for (a) ak model (eq. (4)) with mean of the q=0.4 and th=0.5 (b) r model (eq. (5)) with th=0.1 (c) cw model (eq. (6)) with mean of the q=0.4 and th=0.5 the results of simulation for ak model with the mean of q=0.4 and th=0.5 are shown in fig.2.a. as it is shown in this figure, the number of infected persons is increased from 7 th day to 40 days from the beginning of the simulation, and then it will be decreased till the 46 th day. after that, it will be increased till 48 th day and then it will be decreased till the 77 th day. the maximum number of infected persons is 55 which is reached in the 48 th day. the number of susceptible persons becomes 0 at the 52 nd day and it will be remained 0 until the end of run time. therefore, all the people (100 persons in the social network) became sick in ak model. simulation results for r model with th=0.1 are illustrated in fig.2.b. as it is shown in this diagram, the number of infected persons is increased from 7 th day till 32 nd day of simulation cycle and then it will be decreased till the 49th day. the maximum of the infected subject is 24 in the 32 nd day. the number of susceptible individuals remained in 75 from 25 th day till the end of simulation cycle. in this model, 25 persons in the social network became infected. the results of implementation for cw model with the mean of q=0.4 and th=0.5 are indicated in fig.2.c. as it is shown in this figure, the number of infected people is increased from 7 th day till 40 th day, and then it will be decreased till 41 st day quickly. again it will increase till 42 nd day, after that it will be decreased till the 62 th day. the maximum number of infected people is 81 which is reached in the 40 th day of simulation cycle. the number of susceptible people became 2 in the 38 th day and it will remain in 2 until the end of the simulation, i.e.in cw model 98 persons became sick. with giving different quantities to the q and th for each dynamic model, the number of susceptible people will change. the results are shown in fig.3. 94 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… fig.3 (a) number of susceptible persons for (a) ak model (eq. (4)) when thq : th=0.8 & q=0.4 (b) r model (eq. (5)) when th<0.5 : th=0.2 and when th>0.5 : th=0.8 (c) cw model (eq. (6)) when thq : th=0.8 & q=0.4 in this figure when th q in ak model any node won’t be infected although in cw model 26 persons became ill. also, when th q in ak and cw model after 34 days from beginning of simulation number of susceptible people became 0. it means, in this case, all the susceptible individuals will be sick (see fig.3.a and fig.3.c). in fig.3.b for r model when th=0 all the people in the network will be infected, since the number of susceptible people becomes 0 at 33 rd day. in addition, when 0 th 0.5 only 25 persons became infected and it remained 25 from 25 th day till the end of the simulation. moreover, when th 0.5 any person won’t be infected. in the real world, old people and kids have lower body resistance against the disease. in order to have a more plausible model, we consider different body resistances (th) in our simulations. we assign different quantities to the th based on the persons’ ages in the network. old people and kids were given the lower th and younger people were given higher th. the simulation results for three models in which th quantities are assigned based on persons’ ages is shown in fig4. the diagram of the number of susceptible, exposed, infected and recovered people for ak model in which th quantities are based on individuals ages is shown in fig.4.a. as you see the number of infected individuals is increased from 7 th day till 40 th day, and then it will be decreased till the 69 th day. the maximum number of infected individuals is 75 which is reached in the 40 th day of simulation cycle. the number of susceptible individuals became 0 in the 45 th day and it will remain in 0 until the end of the simulation. therefore, all the people in this social network became sick. for r model as it is shown in fig.4.b, the number of infected individuals is increased from 7 th day till 56 th day of simulation cycle and then it will be decreased till the 65 th day. the maximum of the number of infected individuals is 8 from the 40 th day till 48 th day. the number of susceptible individuals remained in 89 from 41 st day till the end of run time. thus, in this model, just 11 persons in the social network became ill. the simulation results for cw model in which th quantities are assigned based on ages and with the mean of q=0.7 is shown in fig.4.c. as it is shown in this figure, the number of infected people is increased from 7 th day till 40 th day, and then it will be decreased till the 44 th day. again it is increased till 48 th day and after that, it will be decreased till the 70 th day of the simulation period. the maximum number of infected people is 73 which is reached in the 40 th day of simulation cycle. the number of susceptible individuals became 1 in the 46 th day and it will remain in 1 until the end of the simulation. so, 99 persons became infected. advances in systems science and application(2016) vol.16 no.4 95 fig.4 number of susceptible, exposed, infected and recovered persons giving th quantity based on age for (a) ak model (eq. (4)) with mean of the q=0.7 (b) r model (eq. (5)) (c) cw model (eq.(6)) with mean of the q=0.7 in order to have a better comparison with real data from sars infection data for hong kong since 15 february 2003 [20], the number of infected people in fig.4 for each model is depicted in fig.5. fig.5 number of infected persons (a) ak model (eq. (4)) with mean of q=0.7 and giving th quantity based on age (b) r model (eq. (5)) and giving th quantity based on age (c) cw mode l (eq. (6)) with mean of q=0.7 giving th quantity based on age (d) daily reported sars infection data for hong kong since 15 february 2003 all these diagrams have the same trend of propagation of disease. ak and cw model have higher amplitude for the number of infected people. in addition cw model has a higher slope than ak and r model. 4 discussion 96 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… in this study, we examined three different dynamic models for simulating the outbreak of an infectious disease (here the sars was modeled) on a social network. in all the simulations, at the beginning of the disease spread, 1% of people will receive disease virus, (in this study the number of nodes is 100 persons) therefore just one node will go to exposed status. according to fig.2, in the first 7 days, there are no changes in the number of infected people. in fact, the number of infected people depends on the changing status of the first exposed person in the network. the incubation period is supposed 7 days, and the contagion couldn’t transmit from the first exposed person to others in this period. the ak model (eq. (4)) considers these factors: the number of infected neighbors and the probability of infection transmission from an infected person to its susceptible neighbors (q). this probability depends on the transmission rate. in the transmission of epidemic diseases, another important factor is the weights of communications between persons [24, 25]. there are different strategies to assign a weight to a link based on the amount of communication between two nodes [26, 27]. previously, we allocated different amounts for links weights in integer values from 0 to 10 based on the level of communication between the people and we applied this factor (w) in r model (eq. (5)) [1]. according to r model, having an infected neighbor with strong communication weight is enough to become sick. in our data, most of the links have the weight equal to 1. for these links, the probability of becoming sick is equal to 0.1 (eq. (5)) and it is not more than th (th= 0.1). so this dynamic doesn’t let the communication links with weight 1 to transmit the infection. this is the reason for the r model’s result wherein a few people have been ill (see fig.2.b). in the real-word social networks, the communications between individuals aren’t equal. in order to consider heterogeneity on the network, we changed eq. (4) (ak model) to eq. (6) (cw model). in the proposed model the communication weights of all the infected neighbors are considered. the diagram trend of infected people in the cw model is similar to the ak because both of them consider q. in both of these models, almost all the people became infected. besides, the cw model shows more fluctuations in disease propagation than the ak and r model. it may be as a result of considering the summation of communication weights. with giving different amounts to the q and th the below results were achieved from simulations (see fig.3). according to fig.3.a in ak model we observe: if th ≥ q: any node can’t be infected. if th < q: all nodes will be infected. these former statements can be obtained from the following equations: at the beginning of the disease spread just one node will be sick. so some nodes have just one infected neighbor and the remaining ones does not have any infected neighbors. so we have or → =1, 1>1-q > 1-th →th> 1 →th > → =1-q > 1-th →th> 1 →th > therefore, node i doesn’t become infected. it means that if th ≥ q at the beginning of the spread of disease, the infection does not transmit from node i to any node. in r model, we observe that (see fig.3.b): if th ≥ 0.5: no node will be infected. if th < 0.5: some persons will be infected. if th = 0: all the nodes will be infected. in this model as th decreases, the probability of becoming infected will increase. in our data, the weights of links are between 0 and 10. in the beginning, there is only one infected node which its maximum weight of connected links is 5. according to eq. (5) advances in systems science and application(2016) vol.16 no.4 97 ( ), if th ≥ 0.5, the infection can’t transmit from first infected node to its neighbors. in cw model, we observe that (see fig.3.c): if th ≥ q: some persons will be infected. if th < q: all the nodes will be infected. in an overview, when th≥ q in ak model no one became sick, but in cw model some persons became sick. as cw model considers the weight of links and so it is more impressible than ak model. in addition to aforementioned factors, the characteristics of individuals are not the same and people may have different body resistance against the disease, i.e. each person may have a particular threshold of becoming sick (th). for example, infection risks may be related to age. previously hethcote suggested that in realistic infectious disease models it would be beneficial to include the age of individuals [28]. here we allocated different quantities for the resistance body from 0 to 1 based on the persons’ age (see fig.4 and fig.5). for the r model, the number of infected individuals in fig.4 are fewer than it in fig.2, since in our network data most of the people are young and they have high th (compare fig.2 with fig.4). and as mentioned before, the r model depends on th very much. comparing the number of infected persons in our simulated models with real data suggests that the ak and the cw showed more similarity to real result (see fig.5). however, the cw results have more fluctuations similar to real data which is as a result of considering both q and w. 5 conclusions and future prospect modeling transmission of infectious diseases in a society is one the most important field in epidemiology science. here, we have attempted to model the transmission of sars disease in a small population with three different models. based on our results, weights of communications between individuals and the thresholds of body resistances of people are important factors for disease spread. the cw model which has all these factors is more similar to the reported data of infected cases of sars outbreak in hong kong in 2003 [20]. based on these results it may be concluded that the cw captures the dynamic of disease spread in a proper way. as suggestions for the future prospect, it will be useful to model the vaccination effect of the person who has more connections with other people in a society, since in other studies it has expressed that the prevalence of the disease depends on the degree distribution [2]. also, it has shown that quarantining infected persons can reduce the infectious in a society [16]. considering this state in simulated models will be interesting. we believe the model can be applied to another contagious diseases such as ebola, influenza, aids and etc., because these diseases have some similar features. for example, ebola and sars have the analogous incubation period. acknowledgment we want to thank dr.sajad jafari for useful guidance and improvements. references [1] f. r. vishkaie, f. bakouie, and s. gharibzadeh(2014), "common cold outbreaks: a network theory approach," communications in nonlinear science and numerical simulation, vol. 19, pp. 3994-4002. [2] j. saramäki and k. kaski(2005), "modelling development of epidemics with dynamic small-world networks," journal of theoretical biology, vol. 234, pp. 413-421. [3] m. newman(2010), networks: an introduction: oxford university press. 98 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… [4] m. kaiser, r. martin, p. andras, and m. p. young(2007), "simulation of robustness against lesions of cortical networks," european journal of neuroscience, vol. 25, pp. 3185-3192. [5] r. albert(2005), "scale-free networks in cell biology," journal of cell science, vol. 118, pp. 4947-4957. [6] s. r. proulx, d. e. promislow, and p. c. phillips(2005), "network thinking in ecology and evolution," trends in ecology & evolution, vol. 20, pp. 345-353. [7] l. pellis, f. ball, s. bansal, k. eames, t. house, v. isham, et al.(2015), "eight challenges for network epidemic models," epidemics, vol. 10, pp. 58-62. [8] j. m. read, k. t. eames, and w. j. edmunds(2008), "dynamic social networks and the implications for the spread of infectious disease," journal of the royal society interface, vol. 5, pp. 1001-1007. [9] s. eubank, h. guclu, v. a. kumar, m. v. marathe, a. srinivasan, z. toroczkai, et al(2004)., "modelling disease outbreaks in realistic urban social networks," nature, vol. 429, pp. 180-184. [10] f. fasina, a. shittu, d. lazarus, o. tomori, l. simonsen, c. viboud, et al.(2014), "transmission dynamics and control of ebola virus disease outbreak in nigeria, july to september 2014," euro surveill, vol. 19, p. 20920. [11] l. danon, t. a. house, j. m. read, and m. j. keeling(2012), "social encounter networks: collective properties and disease transmission," journal of the royal society interface, p. rsif20120357. [12] v. colizza, a. barrat, m. barthélemy, and a. vespignani(2007), "predictability and epidemic pathways in global outbreaks of infectious diseases: the sars case study," bmc medicine, vol. 5, p. 34. [13] t. yoneyama, s. das, and m. krishnamoorthy(2010), "a hybrid model for disease spread and an application to the sars pandemic," arxiv preprint arxiv:1007.4523. [14] c. l. althaus(2014), "estimating the reproduction number of ebola virus (ebov) during the 2014 outbreak in west africa," plos currents, vol. 6. [15] n. abdulrahman, a. sirajo, and a. abdulrazaq(2015), "a mathematical model for controlling the spread of ebola virus disease in nigeria," international journal of humanities and management sciences, vol. 3. [16] m. small and c. k. tse(2005), "small world and scale free model of transmission of sars," international journal of bifurcation and chaos, vol. 15, pp. 1745-1755. [17] v. colizza, a. barrat, m. barthélemy, and a. vespignani(2007), "predictability and epidemic pathways in global outbreaks of infectious diseases: the sars case study," bmc medicine, vol. 5, p. 1. [18] j. mossong, n. hens, m. jit, p. beutels, k. auranen, r. mikolajczyk, et al.(2008), "social contacts and mixing patterns relevant to the spread of infectious diseases," plos med, vol. 5, p. e74. [19] m. kuperman and g. abramson(2001), "small world effect in an epidemiological model," physical review letters, vol. 86, p. 2909. [20] m. small and c. k. tse(2005), "clustering model for transmission of the sars virus: application to epidemic control and risk assessment," physica a: statistical mechanics and its applications, vol. 351, pp. 499-511. [21] g. chowell, p. w. fenimore, m. a. castillo-garsow, and c. castillo-chavez(2003), "sars outbreaks in ontario, hong kong and singapore: the role of diagnosis and isolation as a control mechanism," journal of theoretical biology, vol. 224, pp. 1-8. [22] c. f. d. c. a. prevention(2004), "frequently asked questions about sars". [23] j. gjorgjieva, k. smith, g. chowell, f. sánchez, j. snyder, and c. castillo-chavez(2005), "the role of vaccination in the control of sars," math biosci eng, vol. 2, pp. 753-769. [24] y. sun, c. liu, c.-x. zhang, and z.-k. zhang(2014), "epidemic spreading on weighted complex networks," physics letters a, vol. 378, pp. 635-640. advances in systems science and application(2016) vol.16 no.4 99 [25] k. van kerckhove, n. hens, w. j. edmunds, and k. t. eames(2013), "the impact of illness on social networks: implications for transmission and control of influenza," american journal of epidemiology, vol. 178, pp. 1655-1662. [26] c. kamp, m. moslonka-lefebvre, and s. alizon(2013), "epidemic spread on weighted networks," plos comput biol, vol. 9. [27] j. stehlé, n. voirin, a. barrat, c. cattuto, l. isella, j.-f. pinton, et al.(2011), "highresolution measurements of face-to-face contact patterns in a primary school," plos one, vol. 6. [28] h. w. hethcote(2000), "the mathematics of infectious diseases," siam review, vol. 42, pp. 599-653. corresponding author f. bakouie can be contacted at: f_bakouie@sbu.ac.ir. mailto:f_bakouie@sbu.ac.ir advances in systems science and applications (2014) vol.14 no.1 22-35 optimizing investment decisions for a set of projects under uncetainty of future returns konstantin n. nechval1 and nicholas a. nechval2 1applied mathematics department, transport and telecommunication institute lomonosov street 1, lv-1019, riga, latvia 2statistics department, evf research institute, university of latvia raina blvd 19, lv-1050, riga, latvia abstract project portfolio selection is a crucial decision in many organizations, which must make informed decisions on investment, where the appropriate distribution of investment is complex, due to varying levels of risk, resource requirements, and interaction among the proposed projects. in this paper, we present a mathematical model of investment for a set of projects under uncertainty of future returns. the model addresses the problem of an investor with access to a limited pool of capital, who makes decisions on investments. the problem is to decide how much to invest in each project so as to maximize the total expected return by the end of the horizon in relation to a given utility function. we discuss optimal investment decisions for the cases where the return from investment is a random variable. the single-period and multi period cases of investment decisions are considered. the paper presents closed form solutions for commonly adopted utility functions. keywords projects, future return, uncertainty, investment decisions, optimization 1 introduction project portfolio investment is the periodic activity involved in investing a portfolio of projects, that meets an organizations stated objectives without exceeding available resources or violating other constraints. some of the issues that have to be addressed in this process are the organization’s objectives and priorities, financial benefits, intangible benefits, availability of resources, and risk level of the project portfolio [1]. difficulties associated with project portfolio investment result from several factors: (i) there are multiple and often-conflicting objectives, (ii) some of the objectives may be qualitative, (iii) uncertainty and risk can affect projects, (iv) the project portfolio may need to be balanced in terms of important factors, such as risk and time to completion, (v) some projects may be interdependent, and (vi) the number of feasible portfolios is often enormous. in addition to these difficulties, due to resource limitations there are usually constraints such as finance, work force, and facilities or equipment, to be considered. as some researchers have noted [2], the major reason why some projects are selected but not completed is that resource limitations are not always formally advances in systems science and applications (2014) vol.14 no.1 23 included in the project selection process. in cases where resource limitations are at fault for a failed project, a selection model that incorporated resource limitations could have aided the decision maker in avoiding such mistakes [1]. portfolio selection becomes more complex when resource availability and consumption are not uniform over time. there are many different techniques that can be used to estimate, evaluate, and choose project portfolios [3-4]. some of these techniques are not widely used because they address only some of the above issues, they are too complex and require too much input data, they may be too difficult for decision makers to understand and use, or they may not be used in the form of an organized process [5]. among all of the techniques that are available, optimization techniques are the most fundamental quantitative tool for project portfolio selection and address most of the important issues [6]. however, they have largely failed to gain user acceptance [7], and few modeling approaches, from a variety of optimization approaches that have been developed, are being utilized as aids to decision making in this area [8]. according to hess “management science has failed altogether to implement project selection models; we have proposed more and more sophistication with less and less practical impact” [9]. one of the major reasons for the failure of traditional optimization techniques is that they prescribe solutions to portfolio selection problems without allowing for the judgment, experience and insight of the decision-maker [7]. a literature review we conducted in this field clearly showed that, although there are many different methods for project evaluation and portfolio selection that have their own advantages, no single technique addresses all of the issues that should be considered in project portfolio selection [10]. among published methodologies for project portfolio selection, there has been little progress towards achieving an integrated framework that: (a) simultaneously considers all the different criteria in determining the most suitable project portfolio, (b) takes advantage of the best characteristics of existing methods by decomposing the process into a flexible and logical series of activities and applying the most appropriate technique(s) at each stage. well-known pragmatic difficulties that make project selection challenging include the following: (1) success uncertainty (unknown rewards): whether (or how well) a project will succeed, both technically and in the market, may be uncertain. (2) changing opportunities (randomly arriving opportunities): the opportunity of funding a project may be uncertain. new projects are continually being proposed throughout the year, stimulated by changes in technology and market opportunities that might help the company to achieve its goal. (3) need for quick funding decisions (on-line decision): the organizations proposing projects lack a comprehensive view of all projects activities company-wide. it is essential 24 konstantin n. nechval: optimizing investment decisions for a set of projects... for top management to give the organization prompt accept/revise/reject decisions, since such feedback helps to coordinate their activities. (4) combinatorial complexity: the costs and benefits from different projects may interact. 2 problem statement the problem we will examine can be formulated as follows. let w0 denote the initial wealth (measured in monetary units) of the investor and assume that there are m projects (opportunities) for investments which belong to a set of m risk categories, with corresponding random rates of return r1, r2, ..., rm, among which the investor can allocate his wealth. the investor can also invest in a project of a riskless category offering a sure rate of return s. if we denote by u1, ..., um the corresponding amounts to invest in the projects belonging to the set of m risk categories and by (w0 − u1 −λ− um) the amount of investment in the project of the riskless category, the final wealth is given by w1 = s(w0 − u1 − · · · − um) + m∑ j=1 rjuj , (1) or equivalently w1 = sw0 + m∑ j=1 (rj − s)uj . (2) the objective is to maximize over u1, ..., um, e{g(ω1)} (3) where g is a known utility function for the investor, w1 is a random variable. we assume that the given expected value is well defined and finite for all w0, ui, and that g is concave and twice continuously differentiable. we will not impose constraints on u1, ..., um. this is necessary in order to obtain the results in convenient form. a few additional assumptions will be made later. if a given amount w0 of the initial wealth is available for investment within n time periods, we are interested in determining decision rules on how much to invest in each project so as to maximize the total expected return by the end of the horizon in relation to a given utility function. 3 investment decisions for the single case let us consider the preceding problem for every value of initial wealth and denote by µ∗ j = µ∗ j (ω0), j = 1, ...,m, the optimal amounts to invest in projects belonging to the set of m risk categories when the initial wealth is w0. we say that the investment portfolio {µ1 j (ω0), ..., µ m j (ω0)} is partially separated if u∗j (w0) = cjh(w0), j = 1,...,m (4) advances in systems science and applications (2014) vol.14 no.1 25 where cj , j = 1, ...,m, are fixed constants and h(w0) is a function of w0 (which is the same for all j). when partial separation holds, the ratios of amounts invested in projects belonging to the set of m risk categories are fixed and independent of the initial wealth; that is, u∗j (w0) u∗k(w0) = cj ck , forj, k ∈ {1, ...,m}, ck ̸= 0 (5) actually, in the cases we will examine, when partial separation holds, the investment portfolio {µ1 j (ω0), ..., µ m j (ω0)} will be shown to consist of affine (linear plus constant) functions of w0 that have the form u∗j (w0) = cj [a+ bsw0], j = 1,...,m (6) where a and b are constants characterizing the utility function g. in the special case where a=0 in (6), we say that the optimal investment portfolio is completely separated in the sense that the ratios of the amounts invested in projects belonging to both the set of m risk categories and the riskless category are fixed and independent of initial wealth. here the following theorem holds. theorem 1. if the utility function satisfies − g′(w1) g′′(w1) = a+ bw1, forallw1 (7) where g′ and g′′ denote the first and second derivatives of g, respectively, and a and b are some scalars, then the optimal investment portfolio is given by (6). furthermore, if g(w0) is the optimal value of the problem, i.e., g(w0) = max u1,...,um e{g(w1)} (8) then we have −g′(w0) g′′(w0) = a s + bw0, forallw0 (9) proof. let us assume that an optimal investment portfolio exists and is of the form u∗j (w0) = cj(w0)[a+ bsw0], j = 1,...,m (10) where cj(w0), j = 1, ...,m, are some differentiable functions. we will prove that dcj(w0)/dw0 = 0 for all w0 and hence the functions cj must be constant. we have for every w0, by the optimality of µ∗ j (ω0) for j = 1, ...,m, de{g(w1)} duj = e { g′ [ sw0 + m∑ k=1 (rk − s)ck(w0)(a+ bsw0) ] (rj − s) } = 0 26 konstantin n. nechval: optimizing investment decisions for a set of projects... j = 1(1)m (11) differentiating the m equations in (11) with respect to w0 yields e   (r1 − s)2 · · · (r1 − s)(rm − s) · · · · · · · · · · · · (rm − s)(r1 − s) · · · (rm − s)2  g′′(w1)(a+ bsw0) ×  dc1(w0)/dw0 ... dcm(w0)/dw0  = −  e { g′′(w1)(r1 − s)s [ 1 + m∑ k=1 (rk − s)ck(w0)b ]} ... e { g′′(w1)(rm − s)s [ 1 + m∑ k=1 (rk − s)ck(w0)b ]}  = −  e { g′′(w1)(r1 − s)s [ 1 + m∑ j=1 (rj − s)cj(w0)b ]} ... e { g′′(w1)(rm − s)s [ 1 + m∑ j=1 (rj − s)cj(w0)b ]}  . (12) using relation (7), we have g′′(w1) = − g′(w1) a+ b [ sw0 + m∑ j=1 (rj − s)cj(w0)(a+ bsw0) ] = − g′(w1) (a+ bsw0) [ 1 + m∑ j=1 (rj − s)cj(w0)b ] . (13) substituting in (12) and using (11), we have that the right side of (12) is the zero vector. the matrix on the left in (12), except for degenerate cases, can be shown to be nonsingular. assuming that it is indeed nonsingular, we obtain dcj(w0) dw0 = 0, j = 1,...,m (14) advances in systems science and applications (2014) vol.14 no.1 27 and cj(w0) = cj , where cj are some constants, thus proving (6). we now turn our attention to proving relation (9). we have g(w0) = e{g(w1)} = e g s 1 + m∑ j=1 (rj − s)cjb w0 + m∑ j=1 (rj − s)cja  (15) and hence g′(w0) = e g′(w1)s 1 + m∑ j=1 (rj − s)cjb  (16) g′′(w0) = e g′′(w1)s 2 1 + m∑ j=1 (rj − s)cjb 2 (17) the last relation after some calculation and using (13) yields g′′(w0) = − e { g′(w1)s [ 1 + m∑ j=1 (rj − s)cjb ]} s a+ bsw0 (18) by combining (16) and (18), we obtain the desired result: −g′(w0) g′′(w0) = a s + bw0 (19) this ends the proof. it can be shown that the following utility functions satisfy (19): exponential: −e−w/a, forb = 0 (20) logarithmic ln(ω + a), forb = 1 (21) power [1/(b− 1)](a+ bω)1−(1/b), otherwise (22) naturally in our problem only concave utility functions from this class are admissible. furthermore, if a utility function that is not defined over the whole real line is used, the problem should be formulated in a way that ensures that all possible values of the resulting final wealth are within the domain of definition of the utility function. 28 konstantin n. nechval: optimizing investment decisions for a set of projects... 4 investment decisions for the multiperiod case it is now easy to extend the one-period result of the preceding analysis to the multiperiod case. we will assume that the current wealth can be used to invest amounts in projects at the beginning of each of n consecutive time periods. if we denote: wυ the wealth of the investor at the beginning of the υth period, uj(υ) the amount to invest in project belonging to the jth risk category at the beginning of the υth period, rj(υ) the rate of return of project which belongs to the jth risk category at the beginning of the υth period, sυ the rate of return of project which belongs to the riskless category at the beginning of the υth period then we have (in accordance with the single-period model) the system equation wν+1 = sνwν + m∑ j=1 (rj(ν) − sν)uj(ν),v = 0, 1, ..., n − 1 (23) we assume that the vectors rυ = (r1(υ), ..., rm(υ)), υ = 0, ..., n − 1, are independent with given probability distributions that result in finite expected values throughout the following analysis. the objective is to maximize e{g(wn )}, the expected utility of the terminal wealth wn , where we assume that g satisfies for all w − g′(w) g′′(w) = a+ bw (24) applying the dynamic programming algorithm to this problem [11-12] , we have gn (ωn ) = g(ωn ) (25) gν(wν) = max u1(ν),...,um(ν) e gν+1 sνwν + m∑ j=1 (rj(ν) − sν)uj(ν)  , ν = 0, 1, ..., n−1 (26) from the solution of the one-period problem we have that the optimal policy at the beginning of period n − 1 is of the form u∗n−1(wn−1) = cn−1[a+ bsn−1wn−1], (27) where cn−1 is an appropriate m-dimensional vector, µ∗ n−1 = [µ∗ 1(n−1), ..., µ ∗ m(n−1)] ′. furthermore, we have −g′ n−1(w) g′′ n−1(w) = a sn−1 + bw. (28) advances in systems science and applications (2014) vol.14 no.1 29 hence, applying the result of this section in (26) for the next to the last period, we obtain the optimal policy u∗n−2(wn−2) = cn−2 ( a sn−1 + bsn−2wn−2 ) (29) where cn−2 is again an appropriate m-dimensional vector. proceedings similarly, we have for the υth period u∗ν(wν) = cν ( a sn−1 · · · sν+1 + bsνwν ) (30) where cυ, υ = 0, 1, ..., n − 1, are m-dimensional vectors that depend on the probability distributions of the rates of return rj(υ) of projects belonging to the risk categories and are determined by optimization of the expected value of the optimal cost-to-go functions gυ. these functions satisfy −g′ ν(w) g′′ ν(w) = a sn−1 · · · sν + bw, ν = 0, 1, ..., n − 1 (31) thus one can see that the investor, when faced with the opportunity to reuse sequentially his wealth, uses a policy similar to that of the single-period case. 5 multiperiod project portfolio investment with markowitz mean-variance optimization to our knowledge, no analytical or efficient numerical method for finding the optimal multiperiod portfolio policy for the constrained mean-variance formulation of markowitz has been reported in the literature. this section presents an optimal solution to the constrained mean-variance formulation of the multiperiod project portfolio investment problem. we consider a portfolio with (m + 1) risky projects, with random rates of returns. let w0 be an initial wealth of an investor at time 0. the investor can allocate his wealth among the (m + 1) projects. the wealth can be reallocated among the (m + 1) projects at the beginning of each of the following (t − 1) consecutive time periods. the rates of return of the risky projects at time period τ within the planning horizon are denoted by a vector rτ = [rτ(0), rτ(1), ..., rτ(m)], where rτ(j) is the random return for project j at time period τ . it is assumed in this paper that vectors rτ , = 0, 1, ..., t − 1, are statistically independent and return rτ has a mean e{rτ} = [e{rτ(0)}, e{rτ(1)}, ..., e{rτ(m)}]′ and a covariance cov{rτ} =  στ(00) · · · στ(0m) ... . . . ... στ(0m) · · · στ(mm)  (32) 30 konstantin n. nechval: optimizing investment decisions for a set of projects... which can be found from the data of observations. let wτ be the wealth of the investor at the beginning of the τth period, and let uτ(j), j ∈ {1, ...,m}, be the amount invested in the jth risky project at the beginning of the τth time period. the amount investigated in the 0th risky project at the beginning of the τth time period is equal to uτ(0) = wτ − m∑ j=1 uτ(j). (33) an investor is seeking a best multiperiod investment strategy, (uτ(0), uτ(1), ..., uτ(m)) for τ = 0, 1, 2, ..., t − 1, such that either (i) the expected value of the terminal wealth wt , e{wt }, is maximized if the variance of the terminal wealth, v ar{wt }, is not greater than a preassigned risk level v∗, or (ii) the variance of the terminal wealth, v ar{wt }, is minimized if the expected terminal wealth, v ar{wt }, is not smaller than a preassigned level e∗. mathematically, a meanvariance formulation for multiperiod project portfolio investment can be posed as one of the following two forms: (i) maximize v ar{wt } (34) subject to v ar{wt } ≤ ν∗ (35) m∑ j=1 uτ(j) ≤ e{wτ} (36) µτ(j) ≥ 0, j = 1, ...,m, τ = 0, 1, 2, .., t − 1 (37) with wτ+1 = m∑ j=1 rτ(j)uτ(j) + wτ − m∑ j=1 uτ(j)  rτ(0) = rτ(0)wτ +∆′ τuτ (38) and (ii) minimize v ar{wt } (39) subject to e{ωt } ≥ e∗ (40) m∑ j=1 uτ(j) ≤ e{wτ} (41) µτ(i) ≥ 0, j = 1, ...,m, τ = 0, 1, 2, .., t − 1 (42) advances in systems science and applications (2014) vol.14 no.1 31 with wτ+1 = m∑ j=1 rτ(j)uτ(j) + wτ − m∑ j=1 uτ(j)  rτ(0) = rτ(0)wτ +∆′ τuτ , (43) where ∆τ = [∆τ(1),∆τ(2), ...,∆τ(m)] ′ = [rτ(1) − rτ(0), rτ(2) − rτ(0), ...,rτ(m) − rτ(0)] ′ (44) formulation (i) or (ii) enables an investor to specify a risk level he can afford when he is seeking to maximize his expected terminal wealth or specify an expected terminal wealth he would like to achieve when he is seeking to minimize the corresponding risk. a strategy of multiperiod project portfolio investment is an investment sequence, ut = [u0, u1, u2, ..., un−1] (45) where uτ = [uτ1 , uτ2 , ..., uτm ] ′, ∀τ = 0(1)t − 1 (46) more specifically, ut is a feedback strategy and uτ maps the wealth at the beginning of the τth period, wτ , into a project portfolio decision in the τth period, i.e., uτ = uτ (wτ ). a multiperiod project portfolio investment strategy, u∗ t , is said to be efficient if there exists no other portfolio one, ut , such that e{wt ;ut } ≥ e{wt ;u ∗ t } and v ar{wt ;ut } ≤ v ar{wt ;u ∗ t } with at least one equality strictly. by varying the value of v∗ in (i) or the value of e∗ in (ii), the set of efficient multiperiod project portfolio investment strategies can be generated. an equivalent formulation to either (i) or (ii) in generating efficient multiperiod portfolio strategies is (iii) maximize e{wt } − ϑv ar{wt } (47) subject to m∑ j=1 uτ(j) ≤ e{wτ} (48) uτj ≥ 0, j = 1, ...,m, τ = 0, 1, 2, .., t − 1 (49) with wτ+1 = m∑ j=1 rτ(j)uτ(j) + wτ − m∑ j=1 uτ(j)  rτ(0) = rτ(0)wτ +∆′ τuτ (50) 32 konstantin n. nechval: optimizing investment decisions for a set of projects... table 1 the observations of returns for three risky projects. year(t) project 0 1 2 1 1.3000 1.2250 1.1490 2 1.1030 1.2900 1.2600 3 1.2160 1.2160 1.4190 4 0.9540 0.7280 0.9220 5 0.9290 1.1440 1.1690 6 1.0560 1.1070 0.9650 7 1.0380 1.3210 1.1330 8 1.0890 1.3050 1.7320 9 1.0900 1.1950 1.0210 10 1.0830 1.3900 1.1310 11 1.0350 0.9280 1.0060 12 1.1760 1.7150 1.9080 where ϑ ∈ [0,∞). it will be noted that if u∗ t } solves (iii), then u∗ t } solves (i) with v∗ = v arwt ;u∗ t , and solves (ii) with e∗ = ewt ;u∗ t . note that ϑ is equal to ∂e{wt }/∂v ar{wt } at the optimal solution of (iii). problem formulation (iii) is preferable to be adopted in investment situations where an investor is able to specify his desirable trade-off between the expected terminal wealth and the associated risk. 6 numerical example consider the case of a stationary multiperiod process with t = 2. an investor has one unit of wealth at the very beginning of the planning horizon, i.e., w0 = 1. the investor is trying to find the best allocation of his wealth among three risky projects, 0, 1, and 2 in order to maximize e{w2} while keeping his risk not exceeding 0.07; that is, v∗ = 0.07. the observations of returns for risky projects, 0, 1, and 2 are given in table 1. at first consider the case of a stationary multiperiod process with t = 1. the problem is to maximize e{ω1} (51) subject to v ar{ω1} ≤ ν∗ = 0.07 (52) 2∑ j=1 u0(j) ≤ e{w0} = w0 = 1 (53) advances in systems science and applications (2014) vol.14 no.1 33 u0(j) ≥ 0, j = 1, 2, ... (54) with w1 = 2∑ j=1 r0(j)u0(j) + w0 − 2∑ j=1 u0(j)  r0(0) = r0(0)w0 +∆′ 0u0. (55) using (55) and table 1, it can be shown that e{w1} = e{r0(0)}w0 + e{∆′ 0}u0, (56) where e{r0(0)} = 1.0891, e{∆′ 0} = [0.1246, 0.1455] (57) v ar{ω1} = [ω0, u ′ 0]cov{[r0(0),△′ 0] ′}[ω0, u ′ 0] ′ (58) with cov{[r0(0),∆′ 0] ′} =  0.0099 0.0015 0.0021 0.0015 0.0407 0.0374 0.0021 0.0374 0.0723  (59) using solver software (ms excel), we obtain from maximization of (51) a percent investment (on w0 = 1) in three risky projects, 0, 1, and 2, at period τ = 0 as follows: µ∗ 0(0) = 0%, µ∗ 0(1) = 26.92%, and µ∗ 0(2) = 73.08%. the corresponding expected terminal wealth and the risk level are given by e{w1} = 1.229 and v ar{w1} = 0.07, respectively. now consider the case of a stationary multiperiod process with t = 2. the problem is to maximize e{w2} (60) subject to v ar{ω2} ≤ ν∗ = 0.07 (61) 2∑ j=1 u0(j) ≤ e{w0} = w0 = 1 (62) 2∑ j=1 u1(j) ≤ e{w1} (63) uτ(j) ≥ 0, j = 1, 2, τ = 0, 1 (64) with wτ+1 = m∑ j=1 rτ(j)uτ(j) + wτ − m∑ j=1 uτ(j)  rτ(0) = rτ(0)wτ +∆′ τuτ (65) 34 konstantin n. nechval: optimizing investment decisions for a set of projects... using (65) and table 1, it can be shown that e{w2} = e{r1(0)r0(0)}w0 + e{r1(0)∆′ 0}u0 + e{∆′ 1}u1 (66) where e{r1(0)r0(0)} = 1.1960 (67) e{r1(0)∆′ 0} = [0.1371, 0.1605] (68) e{∆′ 1} = [0.1246, 0.1455] (69) v ar{w2} = [w0, u ′ 0, u ′ 1]× cov{[r1(0)r0(0), r1(0)∆′ 0,∆ ′ 1] ′}[w0, u ′ 0, u ′ 1] (70) where cov{[r1(0)r0(0), r1(0)∆′ 0,∆ ′ 1] ′} =  0.0491 0.0028 0.0053 0.0022 0.0037 0.0028 0.0496 0.0490 0.0447 0.0427 0.0053 0.0490 0.0942 0.0426 0.0823 0.0022 0.0447 0.0426 0.0407 0.0374 0.0037 0.0427 0.0823 0.0374 0.0723  (71) using solver software (ms excel), we obtain from maximization of (60) a percent investment (on w0 = 1) in three risky projects, 0, 1, and 2, at period τ = 0 as follows: µ∗ 0(0) = 100%, µ∗ 0(1) = 0%, and µ∗ 0(2) = 0%; and a percent investment (on e{w1} = 1.0891) in three risky projects, 0, 1, and 2, at period τ = 1 as follows: µ∗ 1(0) = 39.80%, µ∗ 1(1) = 47.24% and µ∗ 1(2) = 12.96%. the corresponding expected terminal wealth and the risk level are given by e{w2} = 1.2807 and v ar{w2} = 0.07, respectively. 7 conclusion the problem considered in this paper is to decide how much of our available resources to invest in each project so as to maximize the total expected return by the end of the horizon in relation to a given utility function. we formulate the problem in terms of dynamic programming, which allows one to obtain optimal investment decisions for a set of projects under uncertainty of future returns in a simple form. the derived optimal multiperiod project portfolio investment strategy provides investors with the best strategy to follow in a dynamic investment environment. acknowledgments this research was supported, in part, by the latvian council of science and the national institute of mathematics and informatics of latvia under grant no. 06.1936 and grant no. 01.0031. this support is gratefully acknowledged. advances in systems science and applications (2014) vol.14 no.1 35 references [1] schniederjans, m. and santhanam, r. (1993), “a multi-objective constrained resource information system project selection. method”, european journal of operational research, vol.70, pp.244-253. [2] lucas, h.c. (1973), “computer-based information systems in organizations”, science research associates, chicago. [3] cooper, r.g., edgett, s.j. and kleinschmidt, e.j. (1997), “portfolio management in new products: lessons from the leaders-i”, research technology management, vol.40, pp.16-28. [4] dos santos, b.l. (1989), “selecting information system projects: problems, solutions and challenges”, proceedings of the hawaii conference on system sciences, pp.1131-1140. [5] cooper, r.g. (1993), winning at new products, ma: addison-wesley, reading. [6] jackson, b. (1983), “decision methods for selecting a portfolio of r&d projects”, research management, vol.10, pp.21-26. [7] mathieu, r.g.and gibson, j.e. (1993), “a methodology for large scale r&d planning based on cluster analysis”, ieee transactions on engineering management, vol.30, pp.283-291. [8] liberatore, m.l. and titus, g.j. (1983), “the practice of management science in r&d project selection”, management science,, vol.29, pp.962-974. [9] hess, s.w. (1993), “swinging on the branch of a tree: project selection applications”, interfaces, vol.23, pp.5212 [10] archer, n.p. and ghasemzadeh, f. (1996), “portfolio selection techniques: a review and a suggested integrated approach”, innovation research center, working paper 46, school of business, mcmaster university, hamilton. corresponding author konstantin n. nechval can be contacted at: konstan@tsi.lv advances in systems science and application (2015) vol.15 no.1 72-89 trying to evaluate human dignity in a social group antonio caselles spanish society of general systems abstract this study attempts to make progress in the way an instrument can be created, by consensus, to monitor the development of human beings quality of life in all aspects. based on various recent studies into human values, quality of life and subjective well-being, and on the universal declaration of human rights, this study takes human dignity as the supreme value, and development, freedom and equality (with solidarity, justice and peace as the subsidiary values) as the subsidiary values. all these values were disaggregated hierarchically by considering the literature on this matter to obtain measurable variables. they were converted into indices and geometric averages were used to aggregate the considered variables into each level. obviously, this is an initial attempt, and the definitive instrument will need more in-depth studies and ampler consensuses. keywords human values; subjective well-being; quality of life; universal declaration of human rights. 1 introduction from immemorial times, and without going into much detail, philosophers in particular, and human beings in general, have attempted to determine and explain what they need to survive and to obtain satisfactory conditions of life, both individually and also in their social group. in the present-day, and taking advantage of new technologies, we can aspire to find a procedure to give a numerical value to these conditions in a given social group. that is, to devise a system that evaluates human organizations, and not just individuals (or perhaps it does). to this end, it would be necessary to first determine the partial objectives to be reached by asking: what is well? do objectively right social conduct norms exist? if so, how can their degree of implantation and effectiveness be evaluated? evidently this it is the field of ethics. the aim is to find what is good. yet “good” seems a relative thing (“one man’s meat is another man’s poison”, “it never rains to everyone’s taste”, are well-known sayings). for this reason, what the common good actually is must be determined by consensus, and must represent the common desires or preferences in a social group. the larger social group is humanity as a whole (at the moment), and we have the united nations (un), which issued the universal declaration of human rights (udhr) in 1948. this is the first manifestation of the existence of a global consensus on “what is well” that we know. in the udhr, “human dignity” is determined implicitly as the upper rank value, and its subordinated values are also determined [1]. psychology studies the so-called “subjective well-being”, which we could consider advances in systems science and application (2015) vol.15 no.1 73 the equivalent to happiness. some authors, e.g., khaneman and krueger, have attempted to evaluate a group’s happiness by averaging the happiness of their component individuals (determined through questionnaires)[2]. other authors, e.g., parra-luna, have proposed their own scale of values and suggested a way to evaluate it[3]. regarding the philosophical treatment of udhr, the book of moncho is most interesting[1]. from the abundant literature on the matter we have selected, more or less rightly, these two traditional approaches based on (a) human dignity, and (b) subjective happiness, since economic well-being is considered in the udhr to be a subsidiary objective of human dignity. therefore, first we set out to analyze the human dignity concept from the udhr perspective by hierarchically disaggregating its components until we obtain the directly measurable components; second, we do this to suggest a way to evaluate all these lower level components; third, we do this to confer a value to the degree of respect to human dignity in a given social group (city, country, region, world, etc.) from these values by means of mathematical formulas and with a staggered aggregation process. finally, we analyze the happiness or subjective well-being concept, and attempt to find aspects or factors not contemplated in the udhr that can affect human dignity; e.g., environmental caring and sustainability. we attempt to integrate them into this hierarchy as a suggestion for future consensus. 2 human dignity from the udhr analysis and following moncho [1], we deduce that human dignity is considered here to be the supreme value or the upper rank value, and development, freedom and equality to be immediate the subordinated values. solidarity, justice and peace are subordinated to equality (which acts as justification for them). in order to establish common criteria, we attempt to summarize definitions for these seven basic concepts: dignity: the equivalent to being a “person”; that is to say, subject of operations, and not a “thing”; that is to say, an object or usable instrument. this definition assumes that self-conscience and reason exist in a person. development: survival and self-fulfillment options, which include: life/ health, social progress (education, culture, etc.) and standard of life (economic resources, comforts, etc.). freedom: no restrictions to self-fulfillment would be the total freedom which, obviously in a group, must be limited by the dignity of the other group members. equality: non discrimination to face opportunities and rights, and obviously with the limits determined by the social group’s resources. 74 antonio caselles: trying to evaluate human dignity in a social group solidarity: considered synonymous of brotherhood; that is to say, mutual aid. justice: mechanisms of prevention, protection and compensation for individuals or groups to face possible damage or benefits. peace: absence of violence, coercion and fear. next we attempt to not only identify articles on the udhr which explicitly mention the diverse components or ingredients of human dignity existing, but to also locate other possible components in the literature that are not explicitly mentioned in the present udhr, but could perhaps be included in the future according to our criterion. 2.1 development for this aspect, the udhr includes: health, education, sufficient rent and free time. health: includes feeding, dress, house and health care (article 25-1). education: refers to knowledge and aptitudes, and emphasizes the following subsidiary values: ◦ in the state (article 26-1) • free-of-charge and obligatory nature of elementary and fundamental education. • generalization for technical and professional education. • free access through merits to higher education. ◦ like objective (article 26-2) • respect human rights and fundamental liberties. • understanding, tolerance and friendship. • peace-keeping. • values that do not explicitly appear in the udhr, like education objectives (perhaps they are implicit in the three previous ones), but frequently appear in the literature, like desirable: courage, love, joy, calm, prudence, respect of opinions and other people’s customs, communication and cooperation, self-control, self-knowledge, self-acceptance, self-esteem, flexibility, loyalty, integrity/honesty (word), self-discipline, honor the elderly and parents, empathy, generosity, efficiency (yield rate, rapidity, success, social recognition), exploration, reliability, environmental respect, pleasure, responsibility (in case of failure), and to have objectives that go beyond ones own person (see for instance: [4-11]). advances in systems science and application (2015) vol.15 no.1 75 sufficient income: enough to ensure health and family well-being (article 25-1). free time: includes rest, leisure and paid vacations (article 24). sustainability and environment care do not appear explicitly in the udhr. 2.2 freedom it includes the following as subsidiary values: opinion and publication by any means (article 19). choice of partner (in marriage) (article 16). pacific meeting (article 20). pacific association (article 20). choice of work (article 23). choice of asylum (not for common or anti-un crimes) (article 14). displacement (in the territory of a state) (article 13-1). to leave any country and to return to ones own country (article 13-2). trade unions (to endow and/or to affiliate) (article 23-4). choice of childrens type of education (article 26-3). thought, conscience and religion (aims and values) (article 18). ◦ to change ones own religion or beliefs. ◦ to show ones own religion or beliefs (individual and collectively, publicly and privately, education, cult and observance). access to public functions (article 21). legislation: indirectly by means of electing legislators and governors, or directly (article 21-1). vote (article 21-3). the following do not appear explicitly: right to strike, freedom to hire (market), sexual freedom (with consent and fidelity). 2.3 equality the udhr understands equality to be equality in rights and liberties (article 2) and facing the law (article 7). subsidiary values of equality are solidarity, justice and peace because equality justifies solidarity as both justice and peace [1]. 76 antonio caselles: trying to evaluate human dignity in a social group solidarity ◦ non discrimination: by race, color, sex, language, opinion, national or social origin, economic position, birth, etc. (article 2). ◦ same rights: economic, social, cultural and social security (article 22). it clarifies by stating that it refers to rights that are indispensable to people’s dignity, them free developing their personality, and always limited by the resources of the state itself and international cooperation. • right to work: under equitable, satisfactory conditions (article 23-1). • the same wage for the same work (article 23-2). • security and social welfare: insurances for disease, unemployment, widowhood, disability, old age and involuntary loss of means of subsistence (articles 23-1, 25-1 and 26). ◦ the solidarity mechanisms that are not explicit in the udhr: • urgent palliative for ensuring survival. • reintegration mechanisms: for workers and delinquents. • welfare aid for people with difficulties. justice ◦ protection by law (article 7). ◦ appeal to courts (article 8). ◦ presumption of innocence and guarantees of defense (article 11-1). ◦ non retroactivity of laws (article 11-2). ◦ right of property (article 17). ◦ protection of author rights (article 27-2). ◦ protection of human rights (article 28). ◦ duties to the community (article 29-1). ◦ respect of others rights and liberties (article 29-2). ◦ no opposition to the general un principles (article 29-3). ◦ no contradiction (article 30). ◦ effectiveness and efficiency of justice are not explicitly considered. peace: no coercion, no violence, no fear, no misery. ◦ prohibition of slavery and servitude (article 4). advances in systems science and application (2015) vol.15 no.1 77 ◦ prohibition of torture and cruel and/or degrading treatment (article 5). ◦ right to legal personality (article 6). ◦ prohibition of arbitrariness: in cases of detention, prison and exile (article 9). ◦ right to being heard publicly and with justice by an impartial court (article 10). ◦ right to having a nationality and being able to change it (article 15). ◦ right to participating in the scientific progress and in artistic and cultural activities (article 27-1). 3 subjective well-being, happiness or satisfaction with one’s own life for our purpose, these three concepts are considered synonymous. as our objective in this section is to find the factors that influence subjective well-being, and this objective agrees with that in the recent publication of vijayamohanan and asalatha[5], for details we recommend reading this publication. from the literature review that these authors did, we emphasize the following factors that determine human happiness: (a) material conditions and consumption (income, unemployment, inequality, inflation, free time, etc.). (b) satisfactory family life (partner, children, relatives, etc.). (c) personal and family health. (d) satisfaction in the workplace. (e) ones own character or personality. (f) environmental, socio-demographic or institutional factors (community life, friends, liberties, activities, social control, religion/values, etc.). according to these authors, or those mentioned by them, the main factor that influences individual happiness is one’s own character or personality (determined by genetic and environmental factors), followed by health. money seems to have less influence than what people think, mainly from a minimum. sex and age seem to have little influence and depend on certain aspects. influence of marriage is different for men than it is for women. the united nations development program distinguishes between well-being and happiness. the well-being components are health, work and standard of life. the components of happiness are: life with a purpose, receiving respectful 78 antonio caselles: trying to evaluate human dignity in a social group treatment and having a network of social support[12]. the evaluation of these components is made by means of interviews, which determine the percentage of people who state having these components. in order to measure the degree of personal happiness kahneman and krueger propose a scale of adjectives (happy, enjoying, pleasant, depressed, angry and frustrated) and a questionnaire to determine the distribution of the time a person spends between pleasant and disagreeable situations[2]. it distinguishes 19 possible situations in daily life; e.g. intimate relationships, meeting people after work, eating supper, relaxing, eating, exercising, etc.. later it proposes a formula to average the happiness of the individuals in a social group (sum of the degree-of-satisfaction and time-in-the-situation products of different individuals in various situations). as observed, this approach does not attempt to analyze the causes of happiness, but to evaluate happiness as the result. perhaps it would be necessary to ask with questionnaires because of each answer, and to thus attempt to reach the corresponding cause. the literature includes several questionnaires that have been devised to determine an individual’s level of satisfaction with his/her life. a very popular one is the oxford questionnaire[9], although it has received severe critics[13]. to conclude our literature review of happiness causes, we state that we have not found anything new that is objectively measurable to be incorporated into calculations at the human dignity level. therefore, if what is more influential seems to indicate respect to a person’s happiness is his/her own character, and the other conditions have already been considered between the factors that most influence his/her human dignity, we agree with most of the opinions voiced in the literature that the best way to determine a person’s degree of happiness is to directly ask him/her, and with more or less disaggregation according to the questionnaire used. 4 other scales to measure the value of a society or group of people among the more recent studies on this subject, we selected the following for them possibly leading, more or less quickly, to a numerical global evaluation of a society or a group of people which are, at the same time, a source of ideas or suggestions for our purpose. an approach to the subject in more detail is found in the work of parra-luna[3]. this author distinguishes nine groups of values: health, wealth, security, knowledge, freedom, justice, conservation of the environment, quality of activities, prestige. each group is disaggregated into its respective components to obtain 84 measurable lower level ones. the proposed averaging formula is arithmetic. thus, for example, the health value is disaggregated as so: (a) life expectancy, which includes that of 1-year-old children, mortality at the age of 1 year and advances in systems science and application (2015) vol.15 no.1 79 mortality at the age of 50; (b) quality of life, which includes days not worked due to disease or an accident; (c) sanitary means available, which includes the people for 10000 inhabitants and hospital beds for 10000 inhabitants. all the lowest level components are assumed to be measurable and registered in official statistics. dolan et al. propose a detailed hierarchy of the factors that influence subjective well-being or happiness based on the literature review they did[7]. they distinguish seven groups of factors: income, personal characteristics, socially developed characteristics, distribution of available time, attitudes to and beliefs in others, relations with others, and characteristics of an ample social environment. these factors are composed of sub factors; for example, in the factor distribution of available time, they distinguish: hours worked, hours spent traveling from home to work, taking care of others, voluntary community service, physical exercise and religious activities. all the sub factors are assumed to be measurable and are recorded in statistics or can be determined by surveys. they propose aggregation by means of a linear model with an uncertainty term (y = a + b1 · x1 + b2 · x2 + . . . + ε). as mentioned earlier, kahneman and krueger propose a procedure to measure the happiness of a country based on individual surveys, distributed by what each individual does in his/her own time and how he/she feels like in all the considered situations[2]. schwartz puts forward a hierarchy of basic values to motivate action and interrelations (some are hardly compatible) that are common to all cultures (justified), between which there would be differences only for priorities and each value’s relative importance[6]. these values are aggregated into 10 higher ranking values: self-direction, stimulation, hedonism, success, power/prestige, security, conformity (with social norms), tradition (acceptance of one’s own culture and customs), benevolence (group aid) and universalism (well-being for everyone and for nature). the proposed measurement method is based on surveys conducted by means of validated questionnaires. the priorities of the diverse values in each individual and each culture are calculated. the evaluation at the group level is assumed based on averages. no higher level value to group these 10 basic values is considered musek proposes a hierarchy that begins with two macro categories: dionysian values and apollonian values[10]. dionysian values are classified into hedonistic (sensual and heath related) and values of power (profit, success, etc.). apollonian values are classified into moral values (traditional, social, etc.) and satisfaction values (cognitive, cultural, etc.). singular values are fitted to this scheme. nevertheless, the author observes that this hierarchy can change with country and with each individual’s age and time. 80 antonio caselles: trying to evaluate human dignity in a social group maslow proposes a hierarchy of values (motivating priorities), which is not altogether justified and receives considerable feedback[11]. for him, level 1 (physiological) is occupied by: breathing, food, water, sex, dreaming, homeostasis and excretion; level 2 (security) is occupied by: corporal security, and those of work, resources, morality, family, health and property; level 3 (love/belonging) by: friendship, family and sexual intimacy; level 4 (esteem) by: self-esteem, confidence, success, respect of others and to be respected by others; level 5 (self-fulfillment) by: morality, creativity, spontaneity, problem solving, lack of prejudices and acceptance of facts. 5 the development concept considered by the undp the development concept that the un uses attempts to include all the aspects implied in: (a) a prolonged, healthy, creative life; (b) knowledge acquisition; (c) a decent standard of life; (d) political freedom; (e) human rights; (f) to interact freely with others[12]. in our opinion, its amplitude is total. however, a constructive process is followed there, which tends to produce an adequate index for measuring development. the un began by defining the human development index (the hdi, in 1990), whose calculation was specified in detail by anand and sen[14]. for this calculation, the following variables were used: life expectancy at birth (years), adult literacy rate (%), combined gross enrolment ratio (%), and gross domestic product (gdp ) per capita (ppp us$). the hdi measures health by life expectancy when born, wealth by gdp and education by the percentage of people with a degree or registered in regulated studies. in order to certainly complete the hdi, other indices were created, among them the hpi2, which is a poverty index applied to developed countries. this index is based on four variables: probability of not surviving to the age of 60; the long-term unemployment rate; the proportion of adults who lack functional aptitudes; the proportion of the population that lives below the poverty threshold. later, hdi − d (that corrects the hdi by considering inequality)[12], hybrid hdi (that uses a geometric average instead of an arithmetic one in order to be more sensitive to minor differences in lower valued factors), the multidimensional poverty index (mpi) (that analyzes deprivations in hdi components), and the gender inequality index (gii) appeared. in the undp , today the idea seems to develop and perfect these indices in order to obtain a suitable, efficient specification and measurement to monitor human development in diverse countries. we suggest restricting the development concept to the components that are deduced from the udhr, that is, health, education, sufficient rent and free time, to which we would add what concerns sustainability and the environment. thus development, along with freedom and equality (which includes solidarity, justice and peace), as described by the udhr, would be the components of the supreme advances in systems science and application (2015) vol.15 no.1 81 value: human dignity. 6 attempting to suggest a proposal for future consensus let us now consider the scale that the udhr implicitly developed and which we have attempted to make explicit based on the philosophical approach of moncho [1] (see heading 2.). it is now necessary to associate all the values specified in udhr articles with a measurable variable that is registered, or can be registered, in the official statistics of a given country or region. this scale also needs the incorporation of new values which, in 1948, the date when the udhr came into being, were not as important as they are now. attempting to find the best way to measure the concepts specified in the various udhr articles and incorporating new values into it can imply remarkable work, which also requires a consensus. however we dare, as an attempt, to suggest a preliminary proposal which, if it contains something interesting, will have to be improved in subsequent approaches. following the implicit scheme in the udhr, considering the statistics that the undp uses and stressing that our suggestion is merely a preliminary attempt[12], we propose the following procedure to evaluate the degree of human dignity in a social group (world, country, region, municipality, company, etc.). 1. specify the implicit hierarchy of values in the udhr and complement it by considering the statistics and procedures that the undp habitually uses. 2. evaluate all the first-level variables (input or basic) as they appear in the statistics or by means of scores provided by experts. 3. transform all these values into indices with the formula habitually used in the undp ; that is: (present value minimum value)/(maximum value minimum value). 4. calculate the indices of the second-level variables (those that depend solely on first-level variables) by using the geometric average (as the undp suggests). indices with a negative sense will be entered as (1-index) in the corresponding formula. 5. calculate the indices of the third-level variables (those depending on the secondor firstlevel variables) by also using the geometric average. 6. continue in this way until the human dignity index is calculated. this procedure is simple and quite feasible in a computer. in order to complicate it, a weight or importance measure may be assigned to each index. for geometric averages, each weight will appear as an exponent of its corresponding index. for instance, idignity = (idevelopment) a · (ifreedom)b · (iequality)c 82 antonio caselles: trying to evaluate human dignity in a social group where a, b, c are higher than zero, and a+ b+ c = 1. this calculation can be facilitated using logarithms. as a first approach, the list of variables (turned into indices) can be the following: they are numbered between 0 and 142 (both inclusive). when their codified name begins with a y , it is understood that they have a negative sense and that their respective aggregation formula will consider (1− y · ··). i000 human dignity i001 development i002 health (disaggregated according to [4]) i003 resources i004 cost in health per capita (ppa in us$) i005 doctors (for every 10000 inhabitants) i006 hospital beds (for every 10000 inhabitants) y007 risk factors i008 children not immunized against: i009 diphtheria, pertussis and tetanus (% of children aged 1 year) i010 measles (% of children aged 1 year) i011 incidence of hiv i012 young people (% aged 15-24 years) (women) (men) i013 adults (% aged 15-49 years) (total) y014 mortality i015 infantile (for every 1000 live born babies) i016 children aged under 5 (for every 1000 live born babies) i017 adults (for every 1000 inhabitants) (women) (men) i018 rates of death by non transmissible diseases, standardized for age (for every 100000 inhabitants) i019 education (disaggregated according to [4]) i020 achievements in education i021 literacy rates for adults (% of 15-year-olds or older) i022 population who at least finished secondary education (% of 25-year-olds old and older) i023 access to education i024 rate of enrollment in primary education (% of the population of primary education ages) (gross) (net) i025 rate of enrollment in secondary education (% of the population of secondary education ages) (gross) (net) i026 rate of enrollment in tertiary education (% of the population of tertiary education ages) (gross) i027 efficiency of primary education y028 dropout rate, all levels (% of the cohort in primary education) y029 repetition rate, all levels (% of the total enrollment in primary education in advances in systems science and application (2015) vol.15 no.1 83 the previous year) i030 quality of primary education i031 student-teacher ratio (number of students per teacher) i032 teachers trained in primary education (%) i033 income/work/standard of living (disaggregated according to literature[4]) i034 jobs-population ratio (% of the population aged 15-64 years) i035 formal jobs i036 (% of all jobs) i037 womens rate/mens rate ratio y038 vulnerable jobs i039 (% of all jobs) i040 womens rate/mens rate ratio y041 people who work and live with less than us$ 1.25 per day (% of all jobs) y042 unemployment rate per levels of education (% of the labor force with the indicated level of education) i043 primary education or less i044 secondary education or better y045 infantile work (% of children aged 5-14 years) i046 obligatory and paid maternity (in days) i047 free time y048 working hours per year (average per worker) i049 days of paid vacations per year (average per worker) i050 stainability and non vulnerability (disaggregated according to [4]) i051 fitted net saving (% of gross net income) y052 ecological footprint of consumption (hectares per capita) i053 proportion of the total provision of primary energy y054 fossil fuels (%) i055 renewable sources (%) y056 emissions of carbon dioxide per capita (tons) i057 protected area (% of the terrestrial area) y058 population that lives on degraded terrain (%) y059 population with no access to improved services i060 water (%) i061 sewerage (%) y062 deaths caused by intra-domiciliary, atmospheric and water contamination (per 1 million people) y063 population affected by natural disasters (annual average, per 1 million people) i064 freedom i065 of choice i066 of partner (score) 84 antonio caselles: trying to evaluate human dignity in a social group i067 of work (score) i068 of asylum (score) i069 of displacement in a country (score) i070 to leave and return to a country (score) i071 of childrens type of education (score) i072 of religion or beliefs (score) i073 of access to public functions (score) i074 of opinion and publication i075 freedom of press index y076 number of jailed journalists i077 of pacific meetings (score) i078 of pacific associations (to create and to become a member, including trade unions) (score) i079 of pacific manifestations (including religion or beliefs) (score) i080 of legislation (directly or by choosing legislators) i081 to vote (score) y082 % corruption/bribe victims i083 degree of democratic decentralization (score) i084 % of political participation i085 equality i086 solidarity i087 non discrimination (race, color, sex, language, opinion, national or social origin, economic position, birth, etc.) i088 seats in parliament (score) i089 population who have at least completed secondary education (score) i090 rate of participation in the labor force (score) i091 same rights (economic, social, cultural, equal wage for the same work, fair wage)(everything limited by the resources of the state and international cooperation) (score) i092 social security i093 disease (score) i094 unemployment (score) i095 widowhood (score) i096 disability (score) i097 old age (score) i098 non voluntary loss of means of subsistence (score) i099 attending to people who have difficulties surviving (not in udhr) (score) i100 justice i101 protection by law i102 physical (score) i103 of property (score) advances in systems science and application (2015) vol.15 no.1 85 i104 of author rights (score) i105 of human rights (score) i106 obligations to others and the community (score) i107 resources to courts (score) i108 presumption of innocence (score) i109 guarantees of defense (score) i110 non retroactive laws (score) i111 non contradictory laws (score) i112 effectiveness and efficiency in application of laws i113 hearing (score) i114 publicity (score) i115 impartiality (score) i116 non abuse (score) i117 no errors (score) i118 celerity (score) i119 peace i120 no coercion i121 no slavery (score) i122 no servitude (score) i123 right to legal personality (score) i124 right to nationality and its change (score) i125 right to participate in scientific progress (score) i126 right to participate in artistic and/or cultural activities (score) i127 non violence i128 no torture (score) i129 no cruel and/or degrading treatment (score) y130 rate of homicides y131 rate of robberies y132 rate of assaults i133 no fear (disaggregated according to [4]) y134 selling/purchasing arms y135 refugees y136 displaced internally y137 civil war i138 victims i139 intensity i140 no misery y141 % undernourished y142 insufficiency average 86 antonio caselles: trying to evaluate human dignity in a social group 7 conclusions and discussion different approaches have been observed in the literature as to not only the identity of the superior values in a human group, but also its hierarchy and composition. the values detected there have been checked against the subjective well-being components and the result is that well-being depends mostly on genetic factors and personality, while the other factors are assumed in other previously considered values. it is observed that values and their priorities can change according to country and time. accordingly, if the intention is to design a scale of values that is useful for any human group, it is necessary to reach a global consensus and to update it whenever necessary. the undp is immersed in constructing an index that measures the general value of both a human group and partial indices that measure different aspects of this value. this process comes across several difficulties, such as those that derive from the nonexistence or inaccuracy of necessary data and the nonexistence of a consensus about the definition and structure of diverse partial values. in order to advance in determining a definition and a structure between the partial values and the total value of a human group, we based our work on the udhr (a first universal consensus) and on the philosophical studies, which are also based on the udhr. with this information and the suggestions that we found in the undp reports and in studies by different authors, we drew a list of indices, structured by levels, and a general formula to aggregate lower level indices into higher level ones. in this way, the upper rank/higher level index would be the human dignity index (idig) (level 6). like the indices of level 5 (with whose aggregation the idig would be calculated), we would obtain the human development index (idev ), the human freedom index (ifre) and the human equality index (iequ). from this stage onward, we would suppress the adjective “human” for being redundant and obvious. level 4 indices, subsidiaries of the idev , would include the index of health (ihea), the index of education (iedu), the index of income (iinc), the index of free time (ifrt ) and the index of sustainability (isus). we would continue in the same way until we reach the indices of level 1, which would be calculated from the statistical data or from the scores provided by experts. as formulas to calculate the indices of level 1 and to aggregate indices, we agree with the habitual ones employed in the undp reports (present-minimum)/(maximum-minimum) and the geometric average (perhaps weighed), respectively. all the indices must be aggregated in a positive sense. the indices in a negative sense (in the list in heading 6, those that begin with a y) would enter the respective aggregation formulas as the 1-index. despite being obvious, we insist that this proposal (based on the consulted literature) has to be considered a germ or an attempt, must be the object of a more advances in systems science and application (2015) vol.15 no.1 87 detailed or refined study, and has to be submitted to consensus; for instance, the definition of each index, deduced from its assigned components, ingredients or dimensions, and also the weight assigned to each component when averaging. another question to consider in future studies is that which refers to crosssectional indices. we assumed that each component, ingredient or dimension appears once in the structure; that is to say, the considered structure is hierarchical or tree-shaped. however, it is possible that a given component can be considered as a starting point to pertain to more than one branch. in this case, we chose the branch that was more concordant with the concept being dealt with; for example, free medical aid could be considered within the health concept and also on the development branch. however, we considered that it was more likely a solidarity subject than (although also) a health subject. there are also some concepts, such as poverty or gender inequality, that do not appear in the list of heading 6, but include components that either appear on several branches of the tree or do not appear on any. thus in the multidimensional poverty index (mpi)[12], 10 components enter: 2 in health (nutrition and infantile mortality), 2 in education (school enrollment and training years) and 6 in standard of life (goods, floor, electricity, water, sewerage and fuel to cook). both the health ones are on different branches (y141 and i015, respectively). in education, training years do not appear on the list and school enrollment is disaggregated into three levels (i024, i025, and i026). we see that the six standard of life ones do not appear on the list. since many of the items on the list were obtained from tables in the undp report[12] and were assumed to be a first approach, we observed that perhaps it would be advisable to better select the components of the list so that the cross-sectional indices, such as the mpi, are based solely on the components on the list. a similar situation occurs with the gender inequality index (gii); the ideal situation would be that all the components of the cross-sectional indices are included on the list, but this ideal situation is perhaps not attainable if simplicity is preferred. in this last case, the cross-sectional indices would be independent indices of the human dignity index, which we are attempting to design herein. we sincerely hope that somebody finds new ideas in this paper which contribute to or are useful for better monitoring the progress of humanity on our planet, and we encourage anyone who feels motivated by this matter to continue with this type of work. references [1] j.r. moncho. (2003), theory of superior values (in spainish:teoŕia de los valores superiors), campgrafic, valencia, spain. [2] d. kahneman and a.b. krueger. (2006), “developments in the measuremen88 antonio caselles: trying to evaluate human dignity in a social group t of subjective well-being”, the journal of economic perspectives, no.20, pp.3-24. [3] f. parra-luna. (2013), “axiological systems theory: a general model of society”, triplec, no.6-1, pp.1-23. [4] s. roth. (2013), “common values? fifty-two cases of value semantics copying on corporate websites”, human systems management, no.32, pp.249265. [5] p.n. vijayamohanan and b.p. asalatha. (2013), “objectivizing the subjective: measuring subjective wellbeing”, munich personal repec archive, no.45005, http://mpra.ub.uni-muenchen.de/45005/. [6] s.h. schwartz. (2012), “an overview of the schwartz theory of basic values”, online readings in psychology and culture, no.2-1, http://dx.doi.org/10.9707/2307-0919.1116. [7] p. dolan, t. peasgood and m. white. (2008), “do we really know what makes us happy? a review of the economic literature on the factors associated with subjective well-being”, journal of economic psychology, no.29, pp.94-122. [8] w.a. kritsonis. (2007), “ways of knowing through the realms of meaning”, national forum press, houston, tx. [9] p. hills and m. argyle. (2002), “the oxford happiness questionnaire: a compact scale for the measurement of psychological well-being”, personality and individual differences, no.33, pp.1073-1082. [10] j. musek. (1994), “values and value orientations in the background of european cultural traditions”, anthropos (international issue), ljubljana. [11] m. ferguson and abraham maslow. (1980), millennium: glimpses into the 21st century, k. dychtwald and a. villoldo.(eds.), j.p. tarcher, los angeles, ca. [12] united nations development program. (2010), human development report 2010, new york, usa. [13] t.b. kashdan. (2004), ”the assessment of subjective well-being (issues raised by the oxford happiness questionnaire)”, personality and individual differences, no.36, pp.1225-1232. advances in systems science and application (2015) vol.15 no.1 89 [14] s. anand and a. sen. (1994), “human development index: methodology and measurement”, human development report office, new york.vol.5,no.2,pp.1433-1434. corresponding author antonio caselles can be contacted at: antonio.caselles@uv.es advances in systems science and applications (2012) vol.12 no.2 113-121 sup-pixel edge detection technology and its application in non-contact measurement of precise parts bo lei, hong lu and chenghuo shang school of mechanical and electronic engineering, wuhan university of technology, wuhan, china abstract the precision of traditional methods of edge detection of precise parts with linear ccd is determined by the size of ccd cells. the precision can only reach µm at high cost. a sup-pixel edge detection method of ccd based on least square method and derivative operator method is proposed to improve the measurement precision. the image gradient obtained using derivative operator method. point a with the max gradient is then regarded as the point of the edge, and then the pixels nearby edge are interpolated linearly. for example, a high-speed data acquisition system is designed using high accuracy linear ccd tcd1501d and high speed a/d converter tlc5510. fifo memory cy7c460a is used to store the converted data. using the sample pulse of ccd as the system control clock, a high speed data acquisition and storing system with simple circuit is designed in this paper. experimental results show that the resolution µ of the system design is 0.036 pixels and the measurement precision is improved by one order of magnitude over traditional methods of edge detection. it reaches sub-pixel accuracy. keywords edge detection, sub-pixel, linear ccd 1 introduction with the growth of heavy industries and the demand for high quality production, many industries require some form of technology for controlling their production quality. one common quality control (qc) technique employs linear ccd cameras for precise parts measurement. the edge detection and its localization in the image are, in particular, used for dimensional inspection and for localizing objects in industrial applications[1]. the precision is determined by the size of ccd cells used for measurement by linear ccd. people try to improve the precision of the measurement system in this way, for example, improving the manufacturing process to decrease the size of ccd cells, and improving the precision of system by amplifying measured parts with optical system. but the precision can only reach µm at high cost. in order to obtain accurate edge measurements, it is necessary to determine the location of an edge to a greater resolution than the spacing between the pixels of the image sensor, that is to say, at sub-pixel resolution[2]. ohtani[3] summarizes the common sub-pixel edge detection techniques and divide them into first derivative algorithm, second derivative algorithm, template 114 bo lei:sup-pixel edge detection technology and its application in non-contact... matching, edge fitting, and statistical approaches. linearity interpolation, which is a mature theory, is widely used in many fields, such as image and signal processing. but the interpolation result is too smooth, and it would loose some edge of images information[4]. therefore, a sub-pixel measurement method of ccd is proposed to so that the edge of static images is detected using derivative operator method in digital image processing, and the pixels nearby edge are then interpolated linearly. 2 design of measurement system 2.1 design of optical system the optical system is as shown (see fig.1). after the amplified by the optical system, the shade of part is imaged on the photosensitive cells of ccd. fig.1 optical system 1-laser 2-test part 3-objective lens 4-imaging objective lens 5-linear ccd 2.2 system composition the system is composed of illumination unit, test part, optical imaging unit, linear ccd, signal acquisition unit, fifo memory and pc computer (see fig.2). we choose laser as illumination unit. and linear ccd is a toshiba ccd linear image sensor (tcd1501d). the parallel beam irradiates the edge of the test part, it passes through the optical amplification unit and images on the advances in systems science and applications (2012) vol.12 no.2 115 photosensitive surface of ccd. after being sampled and a/d conversion by the signal acquisition unit, the output signals of ccd are stored at fifo memory (cy7c460a). then the digital signals are transmitted to pc computer. 2.3 composition of hardware 2.3.1 linear ccd tcd1501d the tcd1501d which includes sample-and-hold circuits is a high sensitive and low dark current 5000 elements ccd image sensor. the sensor is designed for facsimile, image scanner and ocr. the device is operated by 5v (pulse), and 12v power supply[5]. the features of tcd1501d are shown as table 1 below. table 1 the features of tcd1501d image sensor elements size 7µm by 7µm on 7µm centers photo sensor region high sensitive and low voltage dark signal pn photodiode clock 2 phase (5v) internal circuit s/h circuit package 22pin 2.3.2 analog-to-digital converters tlc5510 the tlc5510 is cmos, 8-bit, 20msps analog-to-digital converters (adcs) that utilize a semiflash architecture. the tlc5510 operates with a single 5-v supply and typically consume only 130 mw of power. included is an internal sample-and-hold circuit, parallel outputs with highimpedance mode, and internal reference resistors. the semiflash architecture reduces power consumption and die size compared to flash converters. by implementing the conversion in a 2-step process, the number of comparators is significantly reduced. the latency of the data output valid is 2.5 clocks. the tlc5510 uses the three internal reference resistors to create a standard, 2v, full-scale conversion range using vdda[6]. so its peripheral circuit is simplified. 2.3.3 fifo memory cy7c460a the cy7c460a is 8k words by 9-bit wide first-in first-out (fifo) memory. each fifo memory is organized such that the data is read in the same sequential order that it was written. full and empty flags are provided to prevent overrun and underrun. three additional pins are also provided to facilitate unlimited expansion in width, depth, or both. the depth expansion technique steers the control signals from one device to another by passing tokens. the read and write operations may be asynchronous; each can occur at a rate of up to 50 mhz. the write operation occurs when the write (w) signal is 116 bo lei:sup-pixel edge detection technology and its application in non-contact... fig.2 system composition low. read occurs when read (r) goes low. the nine data outputs go to the high-impedance state when r is high[7]. 2.3.4 single-chip computer at89s52 the at89s52 is a low-power, high-performance cmos 8-bit microcontroller with 8k bytes of in-system programmable flash memory. the at89s52 provides the following standard features: 8k bytes of flash, 256 bytes of ram, 32 i/o lines, watchdog timer, two data pointers, three 16-bit timer/counters, a six-vector twolevel interrupt architecture, a full duplex serial port, on-chip oscillator, and clock circuitry[8]. 2.4 design of hardware system after being sampled and a/d conversion the output video signals of ccd are transmitted to pc computer. and it is transformed into digital images by pc computer. the key parts of the signal acquisition unit are high-speed analog-todigital converter (tlc5510), fifo memory (cy7c460a) and simple chip computer (at89s52). its principle is shown (see fig.3). advances in systems science and applications (2012) vol.12 no.2 117 fig.3 block diagram of hardware system through differential amplifier and low-pass filter, the output signals of ccd (opt) are sent into a/d converts tlc5510. under the control of the a/d convert controlling clock (clk) that generated by the sample pulse of ccd (sp), the simulated output signal of ccd (opt) are converted into 8-bit digital signals that is corresponding to its simulated amplitude. if enabling port (oe) captures the low level signal, the converted signals are sent to the 8-bit input data ports (a0a0) of fifo memory (cy7c460a). and under the control of write clock (w), which is generated by the sample pulse of ccd (sp), the converted signals are stored in fifo in turn. relative address methods are used in fifo memory, and the address coding generator can be omitted, so the circuits are simplified. upon power-up, the fifo must be reset with a master reset (mr) cycle. this causes the fifo to enter the empty condition signified by the empty flag (ef) being low, and both the half full (hf), and full flags (ff) being high. so while enabling port (oe) of tlc5510 captures the low level signal, the reset 118 bo lei:sup-pixel edge detection technology and its application in non-contact... port (mr) of fifo memory should capture low level to reset the fifo .the shift pulse (sh) of the ccd is sent to the trunk line (into) of the simple chip computer (at89s52). we can get the beginning time and end time of the output signals of ccd by (into). so the signals of 5000 elements of ccd are stored into fifo memory in turn. then the at89s52 reads data which stored into the fifo with bus mode and sends them to pc computer one by the serial ports. the 74ls373 is used for improving the stability of reading data. 3 analysis of the measurement signal of ccd when the test part is imaged on the photosensitive cells of ccd, the ccd converts the optical signals of the test part into discrete voltage signals. the value of each discrete voltage signal corresponds to the light intensity of the photosensitive cells, and the order of output signal is corresponding to the position of the cell. after preprocessing the discrete output voltage signal of ccd, such as high pass filter, differential amplifier and low-pass filter. the edge of one-dimensional image signal of ccd is corresponding to the physical boundary of the test part. the low voltage signals of the ccd output correspond to the shadow of the test part and the high voltage signals is corresponding to the bright area where the test part cant shade the light. the ideal edge signal is a step signal (see fig.4(a)). but the edge signal actually is a gradually-changed signal which decreases gradually(see fig.4(b)), and the actual edge point is located at this area. fig.4 edge signal of one dimension image 4 sub-pixel edge detection algorithm we define that the resolution equals the numbers of cells x when the ccd sensor gets valid signal, and the signal of ccd is valid if the signal to noise ratio (snr) s>1. the image snr is signal square to noise square ratio. and if we take s=1, the valid signal f(x) is the least local variance of all pixel h(x). f (x) = h(x) (1) advances in systems science and applications (2012) vol.12 no.2 119 we can get the image gradient by derivative operator method. then the point a with the max gradient is regarded as the point of the edge. and we get several points nearby point a. the number of points is m. we can get the polynomial curve equation (2) and the max residue by polynomial curve fitting. according to error theory and data processing method [9], and take confidence probability p= 95.44%, can get (3): g(x) = k ∗ x+ c (2) and h(x) = 2δ (3) we take ∆g(x) = 2δ, therefore obtain from (2) and (3): ∆x = 2δ/|k| (4) which leads to: µ = 2δ/|k| (5) the magnification of optical system is β, and it leads to µ = 2δ/(|k| ∗ β) (6) but the output of ccd is discrete signals, so the actual edge point p is between the point p and the point p’. the point p’ is nearest point of the point p. we take h(xa) = ξ ·h(xp ) + (1− ξ) ·h(xp ) (7) and the linear factor ξ ∈ [0, 1],then obtain the position of actual edge point xa is: xa = (h(xa)− c)/k (8) 5 application fig.5 one of the pictures of the experiment data according to the algorithm mentioned above, we measured the location a specific edge with the system. the number of measure times is 5. the magnification of optical system β=20, and the polynomial curve equation under the condition of 120 bo lei:sup-pixel edge detection technology and its application in non-contact... the number of points m=7. one of data of the experiment is shown (see fig.5), we get the results of its polynomial curve fitting (see fig.6).the experimental results show that the performance of this algorithm is better than the traditional method in precision of measurement. fig.6 results of polynomial curve fitting table 2 the features of tcd1501d order ξ=0.3(pixel) ξ=0.5(pixel) ξ=0.7(pixel) µ(pixel) 1 1048.6 1048.3 1048.1 0.021 2 1048.7 1048.3 1048.0 0.034 3 1049.2 1048.9 1048.6 0.031 4 1049.2 1048.7 1048.2 0.047 5 1049.1 1048.7 1048.2 0.048 statistical results average 1049.0 1048.6 1048.2 0.036 standard deviation 0.288 0.268 0.228 0.011 then we get the results under the condition of linear factor ξ=0.3, ξ=0.5 and ξ=0.7 . the results are shown as table 2 blow. the data of the table means the location of the pixel. there are no abnormal data in this table. with the comparison analysis on actual location of the specific edge, the result is best under the condition of ξ=0.7. the repetitive error (2δ) is 0.45, which is lest than a pixel, and the resolution µ of the system is 0.036pixel in theory. thus it reaches sub-pixel accuracy. advances in systems science and applications (2012) vol.12 no.2 121 6 conclusions in this paper, a sub-pixel measurement method of ccd was presented, which detect edge of static images with derivative operator method in digital image processing, and then the pixels nearby edge are interpolated linearly. the technique is to improve the precision of measurement at low cost. at the same time the resolution of the measurement system is calculated in theory. finally the results of experiment verify that the system can extract the accurate sub-pixel edge position. acknowledgements this work was supported by hubei province science foundation (no.2008cdb274) and wuhan high-tech development project foundation (no.200812121559). references [1] o. faugeras. (1993), three dimensional computer vision, cambridge, the mit press:cambridge. [2] a.j. tabatai and o.r. mitchell. (1984), “edge location to subpixel values in digital imagery”, ieee trans. pattern analysis machine intell, pp.188-201. [3] k. ohtani, m. baba. (2001), “a fast edge location measurement with subpixel accuracy”, ieee proceedings instrumentation and measurement technology, pp.2087-2092. [4] zhang xiong, bi duyan, yang baoqiang. (2007), “an edge preserved image interpolation method”, journal of air force engineer university, vol.8, no.3, pp.78-80. [5] toshiba. (1996), the toshiba ccd linear image sensor tcd1501d. [6] li cai, wang an. (2003), “application of 8 bit high speed a/d converter tlc5510”, foreign electronic devices and components, pp.59-61. [7] cypress. (2002), datasheet of cy7c460a. [8] atmel company. (2002), datasheet of at89s52. [9] fei yetai. (2000), error theory and data processing, china machine press. beijing, pp.72-82. advances in systems science and applications (2014) vol.14 no.1 1-21 the theory of parametric control of macroeconomic systems and its applications(i) a. ashimov, zh. adilov, r. alshanov, yu. borovskiy and b. sultanov kazakh national technical university named after k. satpayev, 22 satpaev street, almaty, 05x0013, kazakhstan abstract this work consists of three parts and presents the recent results of development of the theory of parametric control of macroeconomic systems and some its applications for solving a number of concrete problems. keywords mathematical model, structural stability, parametrical identification, parametric control introduction development of adequate methods on the basis of mathematical models for macroeconomic analysis and of evaluating optimal values of parameters (economic policy tools) for macroeconomic systems control on the level of national economies, regional economic unions and world economic system is urgent problem, sharply necessity in solving which was emphasized by the latest global crisis. nowadays, mathematical models of corresponding macroeconomic systems without comprehensive testing for possibility of their application are widely used for macroeconomic analysis (including scenario analysis) and evaluating optimal parameter values of economic policy for controlof macroeconomic systems evolution [1-15]. this paper is devoted to the development of the theory of macroeconomic analysis and evaluating optimal parameter values of economic policy for control of macroeconomic systems on the basis of the corresponding mathematical models, tested for possibility of their application, and it consists of three parts. the first part describes the components of the parametric control theory and its algorithmic foundations. the second part describes mathematical foundations of the parametric control theory, and the third part-the developed theory applications for solving a number of applied problems on the basis of some mathematical models of macroeconomic systems. part 1. components of the parametric control theory and its algorithmic foundations 1.1 components of the parametric control theory of macroeconomic systems given the following facts: solution of either continuous or discrete dynamical system [which can include both controllable parameter vectors-state policy tools (µ), and uncontrollable 2 a. ashimov: the theory of parametric control of macroeconomic systems and ... parameter vectors (a)] depends on initial condition vectors and parameters (coefficients) of this system; solution of static system (for instance, static model of small open economy) depends on parameters (coefficients) of this system; for judging by the study results of dynamical system about an object, described by it, an existence of structural stability (or robustness) of this system is required [16]; for judging by the study results of (static or dynamic) model about an object, described by it, an existence of stability of mapping, defined by this model, is required [17]; and also a condition of macroeconomic model (presented by one of dynamic or static system) stability is required at small perturbations of the initial statistical data for parametric identification of the model (input parameters) and the following components of the parametric control theory are proposed [18-19]. 1. the methods for forming the set (library) of macroeconomic mathematical models. these methods are oriented towards the description of various specific socio-economic situations. 2. the methods for estimating the conditions for robustness (structural stability) of the dynamical mathematical models, the methods for estimating the stability indicators and the methods for estimating stability of mappings, set by models of national economic system from the library (without parametric control). 3. the methods for adjusting the structural instable dynamical mathematical model to obtain its structural stability (methods for attenuation of structural instability). choosing (or synthesizing) the algorithms for attenuation of structural instability for the mathematical model of macroeconomic system. 4. the methods for choosing and synthesizing the laws of parametric control of macroeconomic system based on its dynamical mathematical models. the methods for setting and solving the parametric control problems in terms of corresponding mathematical programming problems on the basis of static mathematical models of macroeconomic systems. 5. the methods for estimating the robustness (structural stability) of dynamical mathematical model. the methods for estimating the stability and the methods for estimating the stability of mappings, set by models of macroeconomic systems (with parametric control). 6. the methods for adjusting the constraints on parametric control of macroeconomic system in the case of structural instability of its mathematical model with parametric control. specification of constraints on the parametric control of macroeconomic system. 7. the methods for studying the effects of uncontrollable parameters and advances in systems science and applications (2014) vol.14 no.1 3 functions (uncontrollable factors) on the results ofsolving of variational calculus problems of synthesis and choice (among given finite algorithms set) of parametric control laws. study of bifurcation points of extremals of variational calculus problems of choosing optimal laws of parametric control. the methods for studying the effects of uncontrollable factors variance on the solution results of mathematical programming problems based on static mathematical models. 8. approach for choosing recommendations on evaluating political rules in the frame of implementing the laws of parametric control of macroeconomic system on the base of the analysis of dependences of optimal criteria values of corresponding parametric control problems on uncontrollable factor values. this paper presents the general results of component-specific development of the parametric control theory (its mathematical and algorithmic foundations). within the framework of the methods for forming the set (library) of macroeconomic mathematical models, it is proposed an algorithm for parametric identification of large-scale macroeconomic models, which uses jointly two identification criteria. within the framework of the methods for examining mathematical models stability, it is proposed numerical algorithms for stability indicators estimation and numerical algorithms for estimation stability of mappings, set by model (in terms of the theory of differentiated mappings singularities); within the framework of the methods for examining the weak structural stability, it is described the proposed numerical algorithm based on the robinson theorem about sufficient conditions of weak structural stability of dynamical mathematical models. within the framework of the methods for choosing and synthesizing parametric control of national economy, based on continuous and discrete non-autonomous dynamical systems, as well as discrete dynamical systems with additive noise, there are formulated and proved the corresponding theorems about conditions for the existence of solutions of variational calculus problems on synthesis and choice (among given finite algorithms set) of optimal parametric control laws. within the framework of the methods for studying the effects of uncontrollable factors variance on the solution results of variational calculus problems on synthesis and choice of optimal parametric control laws, there are formulated and proved the theorems about conditions for the continuous dependence of optimal criteria values of variational calculus problems on uncontrollable parameters (uncontrollable function values). within the framework of studying the bifurcations of extremals of the variational calculus problem of choosing the optimal parametric control laws, there are formulated and proved the theorems about sufficient conditions for the existence of appropriately defined bifurcational point of extremals of the variational 4 a. ashimov: the theory of parametric control of macroeconomic systems and ... calculus problem; there is proposed an approach for choosing the recommendations on evaluating political rules within the framework of implementation of appropriate economic tools for adjusting national economy on the base of analyzing dependences of optimal criteria values of corresponding parametric control problems on uncontrollable factors values. 1.2 algorithm for the parametric identification of large-scale macroeconomic models the following algorithm for the parametric identification of large-scale macroeconomic models is proposed within the framework of elaborating the 1st component of the parametric control theory [19]. the parametric identification problem for discrete dynamical macroeconomic model is finding the estimates of unknown values of its parameters (to which belong unknown values of exogenous functions of model and unknown initial values of its dynamical equations), at those one can obtain the minimum of the objective, characterizing the deviations of output variables values of the model from corresponding observed values (of known statistical data for the period t = t1, t1 + 1, ..., t2). this problem comes to finding the minimum of the multi variable function (parameters) in some closed domain d of euclidean space with constraints, overlaying both on endogenous variables values of the model (e constraints)and on initial parameter values (f constraints). in the case of large number of dimensions n of this domain, the standard methods for finding function extrema are often ineffective because of presence of several local minima of the objective. below we present the algorithm, allowing for features of the parametric identification problems for macroeconomic models and allowing passing over the mentioned problem of “local extrema”. constraints e are formed by economic meaning of endogenous variables of model (for example, by their non-negativity). domain of type d = ∏ni i−1[a i, bi], where [ai, bi] is segments of possible values of the parameter pi; i = 1, ..., n was considered as a range, defined by f constraints for evaluating possible values of exogenous parameters.herewith, parameter estimates, for which had observed values, were searched either within [ai, bi] segments with centers at corresponding observed values (in the case of one such value) or within some segments, covering observed values (in the case of several such values). other [ai, bi] segments for search of the parameters were chosen using indirect estimates of their possible values.nedler-mead algorithm of directed search was used in calculating experiments for finding the minimal values for continuous function k: d → r of several variables [20]. use of this algorithm for initial point p1 ∈ d can be interpreted in terms of (converged to local minimum p0 = argmink of criterion k) sequencep1, p2, ..., where k(pj + 1) ≤ k(pj), pj ∈ d, j = 1, 2,... we will assume that point can be found accurately enough, when we describe the following algorithm. advances in systems science and applications (2014) vol.14 no.1 5 for solving the parametric identification problem of model in question on the base of obvious assumption about divergence (in general case) of minimum points of two different functions, two criteria of the following type were proposed: ka(p ) = √√√√ 1 nµ(t2 − t1 + 1) t2∑ t=t1 na∑ i=1 αi ( yi(t)− yi∗(t) yi∗(t) )2 kb(p ) = √√√√ 1 nβ(t2 − t1 + 1) t2∑ t=t1 nβ∑ i=1 αi ( yi(t)− yi∗(t) yi∗(t) )2 here {t1, ..., t2} is identification period;yi(t), yi ∗ (t)-correspondingly computed and observed values of model output variables, ka(p )-subsidiary criterion, kb(p )-basic criterion; nb > na; αi > 0 and βi > 0 are some weight coefficients, values of which are defined during solution of parametric identification problem for dynamical system; ∑na i=1 αi = nα, ∑nb i=1 βi = nβ. the minimization problems based on the model of corresponding criterion (ka and kb) in the domain d, we will call the problem a and the problem b.theaggregate algorithm for solving the parametric identification problem of model was chosen in terms of the following steps: 1. for some vector of initial values of parameter p1 ∈ d, solve problems a and b simultaneously. then, find the minimum points pa0 and pb0 of criteria ka and kb, respectively. 2. if kb(pb0) < ε for some sufficiently small number ε, then the model parametric identification problem is solved. 3. otherwise, choose the point pb0 as the initial point p1, solve problem a and, choosing the point pa0 as the initial point p1, solve problem b. go to step 2. after sufficiently large number of iterations of stages 2 and 3, initial values of the parameters might leave neighborhoods of the non-global minima in one criterion with help of the other and thereby solve the parametric identification problem. the following methods for evaluating the stability indicators and the structural stability of mathematical models are proposed within the framework of elaborating the 2 component of the parametric control theory. 1.3 methods for evaluating the stability of mathematical models of macroeconomic systems 1.3.1 methods for evaluating the weak structural stability of dynamical models the methods of analysis of the robustness (structural stability) of mathematical model of national economic system are based on: fundamental results on dynamical systems theory in the plane; 6 a. ashimov: the theory of parametric control of macroeconomic systems and ... methods of verification of mathematical models belonging to certain classes of structurally stable systems (classes of morse-smale systems, ω-robust systems, -systems, systems with weak structural stability). at present, the theory of parametric control of market economic development has available a number of theorems about structural stability of specific mathematical models (the model of the neoclassical theory of optimal growth; model of national economic system taking into consideration the influence of the share of public expenses and of the interest rate of governmental loans on economic growth; model of national economic systems taking into consideration the influence of international trade and currency exchanges on economic growth; and others) formulated and proved on the basis of the aforementioned fundamental results. along with analysis of the structural stability of specific mathematical models (both with and without parametric control), based on results of the theory of dynamical systems, one can consider approaches to the analysis of structural stability of mathematical models of national economic system by means of computer simulations. we shall consider below the construction of a computational algorithm for estimating the structural stability of mathematical models of national economic system on the basis of robinson theorem (theorem a) on weak structural stability [21]. theorem.let n ′ be some manifold, and n a compact subset in n ′ such that the closure of the interior of n is n. let some vector field be given in a neighborhood of the set n in n ′. this field defines the c1-flux f in this neighborhood. let r(f,n) denote the chain-recurrent set of the flux f on n. let r(f,n) be contained in the interior of n. let it have a hyperbolic structure. moreover, let the flux f upon r(f,n) also satisfy the transversability conditions of stable and unstable manifolds. then the flux f on n is weakly structurally stable. in particular if r(f,n)an empty set, then the flux f is weakly structurally stable on n. a similar result is also correct for the discrete-time dynamical system (cascade) specified by the homeomorphism (with image) f : n → n ′. therefore, one can estimate the weak structural stability of the flux (or cascade) f via numerical algorithms based on this theorem by means of numerical estimation of the chain-recurrent set r(f,n) for some compact region n of the phase space of the considered dynamical system. let us further propose an algorithm of localization of the chain-recurrent set for a compact subset of the phase space of the dynamical system described by a system of ordinary differential (or difference) equations and algebraic system. the proposed algorithm is based on the algorithm of construction of the symbolic image [22]. a directed graph (symbolic image), being a discretization of the advances in systems science and applications (2014) vol.14 no.1 7 shift mapping along the trajectories defined by this dynamical system, is used for computer simulation of the chain-recurrent subset. suppose an estimate of the chain-recurrent set r(f,n) of some dynamical system in the compact set n of its phase space has been found. for a specific mathematical model of the economic system, one can consider, for instance, some parallelepiped of its phase space including all possible trajectories of the economic system evolution for the considered time interval as the compact set n. the localization algorithm for the chain-recurrent set consists of the following: 1. define the mapping f defined on n and given by the shift along the trajectories of the dynamical system for the fixed time interval. 2. construct the partition c of the compact set n into cells ni. assign the directed graph g with graph nodes corresponding to the cells and branches between the cells ni and nj corresponding to the conditions of the intersection of the image of one cell f(ni) with another cell nj . 3. find all recurrent nodes (nodes belonging to cycles) of the graph g. if the set of such nodes is empty, then r(f,n) is empty, and the process of its localization ceases. one can draw a conclusion about the weak structural stability of the dynamical system. 4. the cells corresponding to the recurrent nodes of the graph g are partitioned into cells of lower size, from which a new directed graph g is constructed (see item 2 of the algorithm). 5. go to item 3. items 3, 4, 5 must be repeated until the diameters of the partition cells become less than some given number ε. the last set of cells is the estimate of the chain-recurrent set r(f,n). the method of estimating the chain-recurrent set for a compact subset of the phase space of a dynamical system developed here allows one, in the case in which the obtained chain-recurrent set r(f,n) is empty, to draw a conclusion about the weak structural stability of the dynamical system. in the case that the considered discrete-time dynamical system is a priori the semi-cascade f, one should verify the invertibility of the mapping f defined on n (since in this case, the semi-cascade defined by f is the cascade) before applying robinson’s theorem a for estimating its weak structural stability. let us give a numerical algorithm for estimating the invertibility of the differentiable mapping f : n → n ′, where some closed neighborhood of the discrete-time trajectory {f ′(x0), t = 0, ..., t} in the phase space of the dynamical system is used as n. suppose that n contains a continuous curve l,which sequentially connects the points {f ′(x0), t = 0, ..., t}. one can choose as such curve a piecewise linear curve with nodes at the points of the above mentioned discrete-time trajectory of the semi-cascade. 8 a. ashimov: the theory of parametric control of macroeconomic systems and ... an invertibility test for the mapping f : n → n ′ can be implemented in the following two stages: 1. an invertibility test for the restriction of the mapping f : n → n ′ to the curve l, namely, f : l → f(l). this test reduces to the ascertainment of the fact that the curve f(l) does not have points of self-crossing, that is, (x1 ̸= x2) ⇒ (f(x1) ̸= (f(x2)). for instance, one can determine the absence of self-crossing points by means of testing monotonicity of the limitation of the mapping f onto l along any coordinate of the phase space of the semi-cascade f. let us choose sufficiently large set of points like xi = (x1i , x 2 i , ..., x n i ) ∈ l, yi = f(xi), yi = (y1i , y 2 i , ..., y n i ) and coordinate number of these points (j). if for all xji , i = 1, ..., n at xji1 < xji2 theine quality yji1 < yji2 is met (or at xji1 < xji2 the inequality yji1 > yji2 is met), then f : l → f(l) mapping is estimatedas invertible. 2. an invertibility test for the mapping f in neighborhoods of the points of curve l (local invertibility). based on the inverse function theorem, such a test can be carried out as follows: for a sufficiently large number of chosen points x ∈ l one can estimate the jacobians of the mapping f using the difference derivations: j(x) = det( ∂fi∂xj (x)), x, j = 1, ..., n. here i, j are the coordinates of the vectors, and n is the dimension of the phase space of the dynamical system. if all the obtained estimates of jacobians are nonzero and have the same sign, one can conclude that j(x) = 0 for all x ∈ l and, hence, that the mapping f is invertible in some neighborhood of each point x ∈ l. an aggregate algorithm for estimating the weak structural stability of the discrete-time dynamical system (semi-cascade defined by the mapping f) with phase space n ′ ∈ r′′ defined by the continuously differentiable mapping f can be formulated as follows: 1. find the discrete-time trajectory {f ′(x0), t = 0, ..., t} and curve l in a closed neighborhood n which is required to estimate the weak structural stability of the dynamical system. 2. test the invertibility of the mapping f in a neighborhood of the curve l using the algorithm described above. 3. estimate (localize) the chain-recurrent set r(f,n). by virtue of the evident inclusion r(f,n1) ⊆ r(f,n2) for n1 ⊂ n2 ⊂ n ′, one can use any parallelepiped belonging to and containing l as the compact set n. 4. if r(f,n) = φ, draw a conclusion about the weak structural stability of the considered dynamical system in n. this aggregate algorithm can be also applied for estimating the weak structural stability of a continuous-time dynamical system (the flux f), if the trajectory l = {f ′(x0), 0 ≤ t ≤ t} of the dynamical system is considered as the curve l. in this case, item 2 of the aggregate algorithm is omitted. the mapping f t for some fixed t(t > 0) can be accepted as the mapping f in item 3. advances in systems science and applications (2014) vol.14 no.1 9 1.3.2 methods for evaluating the weak structural stability of dynamical models by definition of orlov [18], the mathematical model of an economic system in general view is some mapping f : a → b transferring values of initial (exogenous) data p ∈ a to solutions (values of endogenous variables) y ∈ b. after constructing a mathematical model of some real-life phenomena or process and defining some actual values of the point p by known measured data or solving the parametric identification problem, the question about adequacy of the analyzed model arises. the condition of model stability relative to admissible perturbations of the initial data [16] is a one of the conditions of the model adequacy. in case of such stability, small perturbations of the model’s initial data results in small changes of its solution. in the mentioned monograph, the definitions of the basic stability indicators are introduced (these definitions are presented below). monograph, however, does not propose any algorithm for computing the considered indicators of the mathematical model stability. below we present the developed algorithms for evaluating the mathematical model stability indicators which characterize stability of solutions of the mathematical model relative to initial data perturbations. at that, all of the model parameters and variables must be made dimensionless beforehand. let x = (x1, x2, ..., xk) be some vector of values of the model exogenous parameters for the time interval t ∈ {0, ..., t}. let x0 = (x1 0 , x 2 0 , ..., x k 0 ) denote the respective vector of base values for the same time interval. the vector that incorporates the values of parameters and initial values of the variables of differential (or difference) equations is considered as the vector x. the vector of measured statistical data used for finding the model equation coefficients is considered as vector x for the econometric models. let p = (p1, p2, ..., pk) be a vector of the normalized input data of the mathematical model where pi = xi xi 0 , i = 0, ..., k. the vector p0 = (1, 1, ..., 1). let be a space of the normalized input data vectors which includes all admissible sets p, a ⊂ rk is a metric space with the euclidean metric defined by the space rk, p0 ∈ a. let y = y (p) = (y 1, y 2, ..., y k) be a selected vector of the values of endogenous variables for some chosen interval (or moment) of time obtained for the selected values of p. the vector that incorporates the values of some selected set of the model endogenous variables for the aforesaid interval (or moment) of time is considered as vector y for the dynamical models. the vector of coefficients of the model equations or vector of values of some selected set of the model en10 a. ashimov: the theory of parametric control of macroeconomic systems and ... dogenous variables for the aforesaid interval (or moment) of time is considered as vector y for the econometric models. in particular, with p = p0, introduce the notation y0 = y0(p) = (y 1 0 , y 2 0 , ..., y k 0 ). the normalized vector of values of the endogenous variables for the moment of time t1 is denoted by y = y(p) = (y 1 y 1 0 , y 2 y 2 0 , ..., y n y n 0 ); y0 = y(p0) = (1, 1, ..., 1). let b ⊂ rn be a region which contains all possible output values y for p ∈ a with the euclidean metric of space rn, y0 ∈ b. the considered model defines the mapping f of set a into set b. for the selected point p ∈ a and number α > 0, let uα(p) denote the intersection of a neighborhood of the point p with radius α with set a: uα(p) = {p1 ∈ a : ρ(p1, p) ≤ α} here and below, ρ(., .) denotes the euclidean distance between two points of the euclidean space. for some subset b1 ⊂ b, let d(b1) denote the diameter of set v1, that is d(b1) = sup(ρ(y1, y2) : y1, y2 ∈ b1) definition1.1. the number β(p, α) = d(f(uα(p))) is defined as the stability indicator of the econometric model at the point for ¿0. algorithm1.1 for evaluating the model stability indicator β(p, α) by the monte carlo method is as follows: 2. define the vector of normalized input data p = (p1, p2, ..., pk), number α > 0, and set uα(p). 3. generate a set of sufficiently large number m of pseudo-random points (p1, p2, ..., pm ) uniformly distributed in β(p, α). for this purpose, consecutively generate the coordinates pij(i = 1, ..., k; j = 1, ...,m) of the point pj in numerical segments [pi − α, pi + α] covering uα(p) using a generator of uniformly distributed pseudo-random numbers. if the inequality k∑ i=1 (pij − pi)2 ≤ α2 holds (i.e. xj ∈ uα(p)), this point is added to the created set. 4. for each point pj of the set, define point yj = f(pj), j = 1, ...,m , by simulation. 5. evaluate β = max(ρ(yi, yj) : i, j = 1, ...,m). 6. stop. with α = 0.01, the obtained number β/2 characterizes the (maximum) percentage change of values of the model output variables under the perturbed input data by 1%. definition1.2. the number β(x) = inf 0≤α≤α0 β(p, α) is called the absolute stability indicator of the econometric model at point x ∈ a. here, α0 is the maximal advances in systems science and applications (2014) vol.14 no.1 11 admissible relative deviation of values of the model input data. algorithm 1.2 for evaluating the absolute stability indicator β(p) of the econometric model is as follows: for the selected value α0 and numbers j = 0, 1, 2, consecutively find (by algorithm 1.1) numbers βj = β(p, α0/2 j), and then evaluate the number β(p) = inf j=0,1,2,... βj if β(p) turns out to be less than some a priori given small number (i.e. β(p) is considered to be approximately zero), then the mapping f defined by the analyzed model is evaluated at point p continuously depending on the input values. definition1.3. the number γ = sup p∈a β(p) is called the maximal absolute stability indicator of the model for region a. algorithm1.3 for evaluating the maximal absolute stability indicator of the model by the monte carlo method is as follows: 1. generate the set of sufficiently large number m of pseudo-random points (p1, p2, ..., pm ) uniformly distributed in . 2. for each point pj in the set and chosen ,α0 > 0 find the numbers β(pj) by algorithm 1.2. 3. determine the number γ = max j=1,...,m β(pj). 4. stop. if the number turns out to be less than some a priori given small number ε(i.e. is considered to be approximately zero), then the mapping f defined by the analyzed model is evaluated in set a continuously depending on the input values. the developed algorithms were applied for evaluating econometric model of a small open economy and computable general equilibrium model of economic branches. 1.3.3 methods for evaluating the stability of mappings, defined by models, in terms of the theory of differentiated mappings singularities this section describes the methods for evaluating in sense of definition the stability of smooth mappings f : d → e, defined by statical or dynamical model [17]. as domains (d) of the mappings in question are used corresponding domains of possible values of uncontrollable, controllable parameters, coefficients of the model econometric equations in question, and also domain of possible values of observed (statistical) data, used for functions building, which set models econometric equations of the model. range of values (e) of the mapping f contains 12 a. ashimov: the theory of parametric control of macroeconomic systems and ... set of possible values of endogenous variables of the model. existence of such stability property indicates preservation of qualitative properties of mapping, using which the model is described, at small variances of this mapping. when real economic phenomena are described adequately using the mathematical model, stability (or instability) of the mapping, presented by the model, may indicate stability (or instability) of corresponding dependencies of possible values of economic indicators on external (controllable or uncontrollable) factors at small variances of these dependencies. instability of mapping, set by the model, may also indicate inadequacy of the model in question [17]. an algorithm for estimating critical point set of the mappings in question is presented within the framework of this study of the given mappings stability. this algorithm, in particular, allow sestimating the maximality of jacobian matrix rank of mapping at all points of its domain, i.e. to check whether either the investigated map is immersion or submersion. for the immersion case, an algorithm for estimating injectivity of the mapping is proposed. there are formulated and proved the statements, which allow to estimate the stability (and in some case ratios of the image dimension to the counter image dimension-instability) of the mapping in question, when conditions of immersion and injectivity or condition of submersion are satisfied for the mapping in question. there are presented the statements about stability conditions for the mapping in question in the case if this mapping is submersion with fold [23]. there is presented the algorithm, which allows to estimate the mapping as submersion with fold and estimate the stability of the mapping in question in this case. 1.3.3.1 algorithm for estimating the critical point set of the mappings, set by the model hereinafter, we will imply the mapping, defined by the mathematical model when we use the mapping f : d → e (1) let’s denote the mapping arguments vector (1) through p = (p1, ..., pn) ∈ d, and corresponding p point image the model solutions vector denote through y = y(p) = y(p1, ..., pυ) ∈ e, (d ⊂ rn and e ⊂ rυ are some regions). in this case, jacobian matrix with size vn for the mapping (1) at the point p will be written in the following form: j(p) = ( ∂yi ∂pj (p) ) i=1,...,υ;j=1,2,...,n (2) we will also denote the jacobian matrix estimate (2) at some point p ∈ d, obtained by numerical differentiation from (1), through j(p). advances in systems science and applications (2014) vol.14 no.1 13 within the framework of the solution of f mapping stability studying problem, we will present an algorithm, which allows to estimate the j(p) matrix rank maximality for p ∈ d, that is the algorithm for condition estimate rank((j(p))) = min(υ, n), p ∈ d (3) when this condition is satisfied, the mapping f will noth ave any critical point in the domain d [17]. for (3) condition estimate to find any nonsingular minor (of the matrix j(p) of order min(υ, n)) for each p ∈ d is enough, taking into account that total amount of maximal-order minors in j is l = cn υ = υ! n!(υ−n)! if n < υ and l = cυ n = n! υ!(n−υ)! , if n ≥ υ. we will denote the determinant value estimate of such minor of order min(υ, n)) in j(p) for p ∈ d through |mi(p)|, i = 1, ..., l. algorithm1.4. the aggregate algorithm for estimating condition(1.3)for estimation of mapping critical point set. 1) domain d divides to sufficiently large amount of (elementary) parallelepipeds dk of the same size, and define the net p from n points, those are vertices of chosen parallelepipeds: p = {pj : j = 1, ..., n}. 2) compute the values of all j(pj) matrix elements for j = 1, ..., n . 3) for i = 1, ..., l, j = 1, ..., n , compute the determinants |mi(pj)|. 4) for each i = 1, ..., l the set d(i) is defined in the following way. d(i) isa sumof all (closed) parallelepipeds dk, which have property that not all values of |mi(pj)| at dk vertices have the same sign. 5) find the set d̃ = ∩l i=1d(i). 6) if the set d̃ is empty, then condition (3) is evaluated as satisfied. stop. 7) otherwise,the steps 1) 6) of present algorithm are performed substituting domain d by d̃ and diminishing sizes of parallelepipeds, participating in partitioning of d̃. sufficiently large amount iteration of the steps of the presented above algorithm allow either to estimate satisfaction of the condition (3) or to obtain the estimate (using set d̃)the set of critical points of f mapping. 1.3.3.2 algorithm for estimating the nonlocal injectivity of mapping,set by model in this section it assumes that n < υ (d domain dimension of mapping (1) is less than erange of values dimension). this section presents an algorithm for estimating conditions of nonlocal injectivity (absence of non-close points in d, having equal images at mapping f) of mapping (1), set by model in domain d. satisfaction of mentioned nonlocal injectivity condition and condition (3), ensuring local injectivity in neighborhood of each point means an existence of inverseto fmapping, determined in set f (d). fix sufficiently small number ε > 0. for each point pj ∈ p through d(pj) 14 a. ashimov: the theory of parametric control of macroeconomic systems and ... denote set of all points pk ∈ p which have |pj − pk| > ε. here | · | is magnitude of vector. algorithm1.5 for estimating the condition of the nonlocal injectivity is presented in terms of the following steps. 1) compute the numbers mj = min pk∈d(pj) |f (pk)− f (pj)| for each point pj ∈ p . 2) compute m == minpj∈pmj and determine the points pj , pk ∈ p such that |f (pk)− f (pj)| = m. 3) repeat steps 1) and 2) of this algorithm substituting the net p by net p1 being the vertices of less size parallelepipeds and containing all points of the net p, being distant from one of pj , pk points by the distance not exceeding 2ε. there are two possible cases given sufficiently large amount of iterations of steps of the presented above algorithm. a) the sequence of obtained values are diminishing about proportional to the p1 net step. in this case,f mapping is estimated as non-injective (in other words, f (d) set is estimated as self-crossing). b) condition m > ε is met for all nets in question with sufficiently small step. in this case, f mapping is estimated as injective (in other words, f (d) set is estimated as non-self-crossing). 1.3.3.3 estimating the stability of mappings set by the model this section presents propositions about sufficient conditions for f mapping stability in open domain d0 = d \ γ(d), where γ(d) is boundary of set d, within the framework of determining stable mapping [17]. it also presents an aggregate algorithm for estimating mapping f : d → e, as stable submersion with fold. according to stability theorem of mazer [17], mapping f is stable in manifold d0, if it is infinitesimally stable in d0. condition of infinitesimal stability of mapping h(p) in k(y) is formulated in terms of solubility in relative to mappings and of the following homologous equation [17]. µ(p) = −j(p)h(p) + k(f (p)) (4) here µ(p) is an arbitrary infinitesimal deformation of f mapping, which is presented in terms of smooth correspondence to each point p ∈ d0 of tangent vector to manifold e atpoint f (p); h(p) is the smooth vector field in d0; k(y) is the smooth vector field in e. consider two possible cases. 1. let n < υ let condition (3) be satisfied, that is f mapping is the immersion advances in systems science and applications (2014) vol.14 no.1 15 in d. according to the theorem about inverse function (when f mapping injectivity condition is satisfied, which checks by the algorithm described in item 2) in set f (d0) the smooth mapping f−1 : f (d0) → d0, inverse to f is defined. if in equation (4) input h(p) = 0, and k(y) = µ(f−1(y)), then this equation become an identity µ(p) = 0+µ(f−1(f (p))). this means that the mentioned mappings, h(p) and k(y), are solutions of homologous equation (4), that is this equation is soluble. consequently, the first statement of the following proposition is true. proposition 1.3.1. let given n < υ for all points of chosen set d,condition (3)for injective mapping(1.1)is satisfied.then mapping(1) is stable in domain d0 = d \ γ(d). if condition(3) is not satisfied for any point p ∈ d0, then mapping(1) is not stable in its domain. the second statement of the proposition 1.3.1 results from the sentences 2.4 and 3.12 [23].formulate them. statement 2.4. let x is compact and dimy ≥ 2dimx+1. mapping f : x → y is stable, if and only if f isone-to-one immersion. statement3.12. let x is compact manifold and dimy = 2dimx. mapping f : x → y is stable, if and only if it is an immersion with normal crossings. from these sentences results that when condition dimy ≥ 2dimx is satisfied, the immersion property is a requirement for stability, that is when (3)is not satisfied, mapping f cannot be stable. note that although in formulations of mentioned sentences x manifold compactness condition is used, proofs of immersion properties for stable mappings given dimy ≥ 2dimx + 1(or dimy = 2dimx) are based on the theorems 5.6 and 5.7 [23], which do not require x manifold compactness. proposition 1.3.1 is fully proved. 2. let n ≥ µ let condition (3) be satisfied, that is f mapping is a submersion in d0. test, in this case, the solubility of homologous equation (4). consider first the case, when some v-order minor determinant mi(p) of jacobian matrix has constant sign for p ∈ d0, that is, |mi(p)| ≥ ε ≥ 0 in d0. let, for determinacy, such anon singular minor mi(p) consists of first six columns of matrix j(p). for arbitrary deformation µ(p), written in terms of v-dimensional column vector in equation (4) determine h(p) in the following way. first v coordinates of n-dimensional column vector h(p) set using column −(mi(p)) −1µ(p), and all other coordinates of vector h(p) assume to be zero. if put k(y) = 0, then equation (4) becomes the identity: µ(p) = −j(p)h(p) + 0 = mi(p)(mi(p)) −1µ(p) (5) consider the general case of satisfaction of condition (3). let for each point p ∈ d (and some its neighborhood) be found its nonsingular minormi(p). choose from such neighborhoods uj finite domain cover: d ⊂ u s j=1uj . make subject to this cover the partition of unity into domains d in terms of s smooth functions 16 a. ashimov: the theory of parametric control of macroeconomic systems and ... φj(p) ≥ 0 [23], where φj(p) = 0 at p ∈ rn\uj(j = 1, ..., s) and ∑s j=1 φj(p) = 1 for p ∈ d. consider arbitrary deformation µ(p) in equation (4). then, for each neighborhood uj and its corresponding nonsingular minor mi(p), make (according to the method proposed in the previous paragraph) the vector field hj(p) in uj , which is the solution of equation (4) in uj at k(y) = 0. determine the vector field h(p) in d0 using the formula: h(p) = s∑ j=1 φj(p)hj(p) thus, from (4) (where h(p) should be substitute hj(p)) results that the vector fields h(p) and k(y) = 0 are the solutions of (4): −j(p)h(p) + k(f (p)) = −j(p) s∑ j=1 φj(p)hj(p) + 0 = s∑ j=1 φj(p)[−j(p)hj(p)] = s∑ j=1 φj(p)µ(p) = µ(p) s∑ j=1 φj(p) = µ(p) this means that the equation (4) is soluble. therefore, the following proposition is true. proposition 1.3.2. let given n ≥ υ for chosen set d condition(3)be satisfied for mapping (1).then mapping(1) is stable in domain d0 = d \ γ(d) 3. now consider the case, when given n ≥ υ condition (3) is not satisfied for some points of domain d0 (that is, the case, when domain d0 contains the critical points of mapping f). denote by s1(f ) set of points of domain d0, in which jacobian matrix of mapping frank is less by unit than the maximal one, that is s1(f ) = {p ∈ d : rank(j(p) = υ − 1} it is known that when the additional condition (j1f ◃▹ s1, where j1f is 1stream of mapping f, s1 is submanifold in the space of 1-streams j1(d0, e) consisting of streams with 1 co-rank, ◃▹ is the sign of transversality), is satisfied, the set s1(f ) is a submanifold in d0 with dimension υ − 1 [23]. definition. let the mapping f : d0 → e satisfy condition j1f ◃▹ s1. point p ∈ s1(f ) is called as fold point, if sum of tangent space to s1(f ) and kernel of tangent mapping df at this point has dimension n, that is, if tps1(f ) +ker(df )p = tpd (6) since sum of dimensions of summands of lhs is equal to the dimension of rhs (υ − 1) + (n − υ + 1) = n, then condition (6) is equivalent to that these advances in systems science and applications (2014) vol.14 no.1 17 summands have the only common point origin of coordinates and cos∠(tps1(f ),ker(df )p) ̸= 1 (7) definition. mapping f : d0 → e is called a submersion with folds, if eachits singular pointis a fold point. in this case submanyfold s1(f ) is called as a fold. it is known that if f : d0 → e is submersion with fold, then f mapping constraint on s1(f ) fold is immersion. the following theorem is valid [23]. consequence 1.3.4. if f : d0 → e is submersion with fold and f |s1(f ) is injective, then mapping f : d0 → e is stable. we present an algorithm for estimating the satisfaction of these conditions of this consequence. algorithm 1.6. aggregate algorithm for estimating mapping f : d0 → e as stable submersion with fold. let given n ≥ υ, after use of algorithm 1.4, the estimate of singular points set of the mapping fin terms of non-empty set d̃ ⊂ d and set of p̃ vertices of elementary parallelepipeds, containing d̃, be found. algorithm steps for testing transversality condition (j1f ◃▹ s1) is not presented here because of unhandiness. we will assume that condition j1f ◃▹ s1 is estimated as satisfied. 1. for each primary parallelepipeddk ⊂ d̃, satisfaction of condition rank(j(p)) = υ − 1 is estimated in the following way. let {pkj }2 n j=1 be set of parallelepiped vertices dk; {|mi(p k j )|}li=1 is set of values of υ − 1-order minor determinants of jacobian matrix j(pkj ) at point pkj ; l = cυ−1 n n = n!n (υ−1)!(n−υ+1)! . if for chosen dk one can find such number i, that all determinants |mi(p k j )| for j = 1, ..., 2n have the same sign, then rank(j(p)) is estimated by number υ − 1 in parallelepiped dk. if rank(j(p)) is estimated by number υ − 1 for all dk ⊂ d̃, then set d̃ is considered to be the estimate of submanifold s1(f ). otherwise, if one can find such parallelepiped dk ⊂ d̃, that for each chosen i = 1, ..., l numbers |mi(p k j )| for j = 1, ..., 2n have different signs, then subset d̃ is not estimated as a fold. stop. 2 the next steps 3, 4, 5 are performed for each point of the net p̃ . 3. for p ∈ p̃ basis vectors (e1, ..., eυ−1) of tangent space tps1(f ) are estimated in the following way. choose m (where m ≫ υ−1) close (except this point itself) to p points of the net p̃ : {p1, ..., pm}. identify the following set of m vectors {fi = pi − p : i = 1, ...,m}. here the points pi and p are consider edasradiusvectors. linear envelope of arbitrary set of n-dimensional vectors (e1, ..., eυ−1) denote by t = t (e1, ..., eυ−1). the distance from point fi up to the plane t 18 a. ashimov: the theory of parametric control of macroeconomic systems and ... denote by d(fi, t ). sum of squares of distances from points fi to t denote by s(e1, ..., eυ−1) = m∑ i=1 (d(fi, t )) 2 (8) coordinates of required e1, ..., eυ−1 vectors are determined by the least squares method from the condition of s(e1, ..., eυ−1) function minimum. 4. (g1, ..., gυ−n+1) basic vectors of ker(df )p kernel for p ∈ p̃ are estimated in the following way. 4.1. since mapping matrix (df )p with theoretical rank υ − 1 is estimated by numerically found jacobian matrix j(p) , all υ − 1-order minors of which have determinants close to zero, we firstly determine the row close to linear combination of other rows of the matrix j(p). in case, if matrix rank is less by unit than the number of its rows, according to the theorem about principal minor, one of matrix rows is a linear combination of its other rows and its elimination does not change the kernel of the linear operator, appropriate to this matrix. let {j1, j2, ..., jυ} be the set of all normalized (the elements of each row divide to the magnitudeof this row, if magnitude of any row is zero, then the problem in item 4.1 is solved) row of j(p) matrix. let p i, (i = 1, ..., υ) be linear envelope of all rows of mentioned set, except the row j i, which is considered as plane in the space rn. let mi be the distance from the point j i ∈ rn to plane p i : mi = d(j i, p i). mi value can be found by finding minimum of function of υ − 1 variable (α1, ..., αi−1, αi+1, αυ): di(α1, ..., αi−1, ..., αi+1, αυ) = |j2 − υ∑ j=1,j ̸=i αij j |2 choose number i relevant to minimal number mi. denote the matrix j(p) of size (υ − 1)× n with removed i-row by j̃(p). 4.2. solve the linear homogeneous system from (υ − 1) equation with n unknowns and with matrix of system j̃(p) by gauss method. herewith, basis of its decision space (g1, ..., gυ−n+1) is found. 5. for p ∈ p̃ estimate angle cosine (cosφp) between the planes tps1(f ) and ker(df )p in rn in the following way. let (e1, ..., eυ−1) and (g1, ..., gυ−n+1) be the found above estimates of bases of these planes; (α1, ..., αυ−1) and (β1, ..., βυ−n+1) be variables sets. let the vector e = ∑υ−1 i=1 αiei be arbitrary vector of plane tps1(f ), vector g = ∑υ−1 i=1 βigi be arbitrary vector of plane ker(df )p. define function f (expressing angle cosine between the vectors e and g) on n variables in terms of (α1, ..., αi−1, β1, ..., βn−υ+1) y = f(α1, ..., αi−1, β1, ..., βn−υ+1) = fg |f ||g| advances in systems science and applications (2014) vol.14 no.1 19 its maximal value takes as the value of required φp. 6. choose small number ε > 0. if for all p ∈ p̃ condition cosφp < 1− ε is met, then based on the definition, d̃ set is the fold estimate and mapping f : d0 → e is estimated as the submersion with fold. otherwise, if p ∈ p̃ is found, for which cosφp ≥ 1 − ε, then d̃ set is not estimated as fold. stop. 7. estimate the injectivity of the mapping f on fold d̃. since f mapping constraint on fold is immersion then local injectivity of this constraint is guaranteed. nonlocal injectivity is estimated by algorithm 1.5, in which one should substitute the net p by the net p̃ . number ε, used in this algorithm, should exceed doubled diameter of elementary parallelepipeds of d̃ set. two cases are possible given sufficiently large amount of step iterations of algorithm 1.5. a) the sequence of obtained m values are diminishing about proportional to the p̃1 net step. in this case, f mapping constraint on the fold d̃ is estimated as non-injective (in other words, f (d̃) set is estimated as self-crossing). additional study of f (d̃) self-crossing points for normality is required. b) condition m > ε is met for all nets in question with sufficiently small step. in this case, f mapping constraint on the fold d̃ is estimated as injective (in other words, f (d̃) set is estimated as non-self-crossing). based on the consequence 1.3.4, the mapping f : d0 → e, in this case, is estimated as stable. references [1] r.a. mundell. (2000), “a reconsideration of the twentieth century”, the american economic review, vol.90, no.3. [2] j.a. frenkel, a. razin. (1987), “the mundell-fleming model: a quarter century later”, national bureau of economic research, working paper , cambridge, massachusetts, no.2321. [3] c. erceg, l. guerrieri, c. gust. (2005), “sigma: a new open economy model for policy analysis”, board of governors of the federal reserve system, international finance discussion papers, july, no.835. [4] j. andrés, p. burriel, á estrada. (2006), “bemod: a dsge model for the spanish economy and the rest of the euro area”, documentos de trabajo banco de españa, no.0631. [5] n. stähler, c. thomas. (2011), “fimod-a dsge model for fiscal policy simulations”, deutsche bundesbank discussion paper. series 1: economic studies, no.06. [6] s. paltsev, j.m. reilly, h.d. jacoby, r.s. eckaus, j. mcfarland, m. sarofim, m. asadoorian, m.h. babiker. (2005), “the mit e20 a. ashimov: the theory of parametric control of macroeconomic systems and ... missions prediction and policy analysis (eppa) model: version 4”, http://globalchange.mit.edu/research/publications/697, report no.125. [7] p.d. dixon, m.t. rimmer. (1998), “forecasting and policy analysis with a dynamic cge model of australia”, preliminary working paper,no.op-90, http://www.monash.edu.au/policy/elecpapr/op-90.htm. [8] v.l. makarov, a.r. bahtizin, and s.s. sulakshin. (2007), “application of computable models in state administration”, moscow: nauchniy expert, russian. [9] g. gandolfo. (2014), “international trade theory and policy”, springer heidelberg new york dordrecht london. [10] k. farmer, m. schelnast. (2013), “growth and international trade. an introduction to the overlapping generations approach”, springer heidelberg new york dordrecht london. [11] r.c. (2004), “fair, estimating how the macroeconomy works”, cambridge, massachusetts: harvard university press, london, england. [12] s. murchison, a. rennison. (2006), “to tem: the bank of canada’s new quarterly projection model”, bank of canada technical report, no.97. [13] m. adolfson, s. laséen, j. lindé and l.svensson. (2011), “optimal monetary policy in an operational medium sized dsge model”, journal of money, credit and banking, blackwell publishing, vol.43(7), pp.1287-1331. [14] d. acemoğlu. (2008), “introduction to modern economic growth”, princeton university press. [15] f.j. andré, m.a. cardenete, and c. romero. (2010), “designing public policies. an approach based on multi-criteria analysis and computable general equilibrium modeling”, lecture notes in economics and mathematical systems, 1st ed, vol. 642, springer. [16] v.i. arnold. (1988), “geometrical methods in the theory of ordinary differential equations”, springer-verlag, new york. [17] v.i. arnold, s.m. gusein-zade, a.n. varchenko. (1985), “singularities of differential maps. volume i”, basel, stuttgard: birkhauser, boston. [18] a.i. orlov, econometrics. (2002), “a textbook”, ekzamen, moskow, russian. advances in systems science and applications (2014) vol.14 no.1 21 [19] a.a. ashimov, b.t. sultanov, zh.m. adilov, yu.v. borovskiy, d.a. novikov, r.a. alshanov, as.a. ashimov. (2013), “macroeconomic analysis and parametrical control of a national economy”, springer, new york. [20] j.a. nelder, r. mead. (1965), “a simplex method for function minimization”, the computer journal, no.7, pp.308-313. [21] robinson c. (1980), “structural stability on manifolds with boundary”, journal of differential equations, no.37, pp.1-11. [22] e.i. petrenko. (2006), “development and realization of the algorithms for constructing the symbolic set”, differential equations and control processes (electronic journal), no.3, pp.55-96. [23] . golubitsky, gueilleminv. (1973), “stable mappings and their singularities”, new york, heidelberg, berlin: springer-verlag. corresponding author a. ashimov can be contacted at: ashimov37@mail.ru advances in systems science and applications (2014) vol.14 no.2 183-189 competitive equilibrium of a sequence of incomplete markets with a continuum of agents guo-sheng zhang the school of economics, trade and event management, beijing international studies university, beijing, china, 100024 abstract by introducing the large economical ideas to multi-period financial market, we have constructed the multi-period economy with incomplete market and a continuum of agents. the competitive equilibrium has been proposed and the existence has been claimed. our equilibrium definition is a development compared to that described by radner, and our conclusion for equilibrium existence has generalized the related results obtained by aumann and zhang, if only the future contracts and goods are traded on security-spot markets. keywords perfect competition, equilibrium, incomplete markets, correspondence integral 1 introduction consider a sequence of markets at successive dates, no one of which is complete in the arrow debreu sense, i.e., at every date and for every commodity there will be some future dates and some events at those date for which the spot goods and future contracts contingent on those events are traded. for such economy, radner had first proposed the concept of common expectations that require traders to associate the same future prices to same future exogenous events[1]. an equilibrium is a set of prices at the first date, a set of common price expectations for the future, and a consistent set of individual plans for agents such that, given the current prices and price expectations, each individual agents plan is optimal for him, subject to an appropriate sequence of budget constraints. radner’s common expectation is a foundation for modern incomplete market theory. but a basic assumption of such model is that the current prices and price expectations for future contracts be not affected by a single agents action, which needs the market be perfect competition. otherwise a change in an individuals offer to buy or sell can easily upset the prevailing prices, so that the equilibrium will never be achieved, and price system is meaningless. as early as 1964, aumann had suggested that the most natural mathematical model for a commodity market with such perfect competition is one in which there are a continuum of traders (like the continuum of points on a line)[2]. for decades, aumann’s large economy has always been a vigorous field of economics. a deficiency of modern incomplete market theory is the lack of introduction of aumann’s large economy ideas. as a novelty, zhang first discussed security-spot markets with a measurable space of agents[3-4], where aumann’s ideas have been 184 guo-sheng zhang: competitive equilibrium of a sequence of incomplete markets with ... applied to security-spot market, particularly, to financial markets. but these research works are primary: the model is two-periods, and the existing discussion needs to proceed technically. in this paper, by combining aumann’s ideas with radner’s model, we discuss the economy with a sequence of incomplete markets and a continuum of agents. first, the concept of a competitive equilibrium for such economy is proposed, which has developed radner’s definition. then, we claim the existence of equilibrium. the sufficient conditions are all used by aumann and radner[5-6]. for the security-spot market with goods and future contracts, our conclusion is the generalization of related results proved by aumann and zhang[4,5]. 2 a multiperiod model for a sequence of incomplete markets with a continuum of agents consider an economy extending through a finite sequence of elementary dates 1, 2, ..., t , in an environment with a finite set s of alternative states. the set of events observable at date t will be represented by partition φt of s . it is assumed the sequence of partition, φt is monotonous, no decreasing in fineness, that is, φt+1 is as fine as φt 1 . also, take φ1 = {s}. for each date, there is a finite set of commodities, numbered 1, 2, ..., l. trade contract (such as future contract) at date t in event a, denoted by θhtu(a,b), specifies the number of units of commodity h that the trader will receive from the market at date u ≥ t in event b(θhtu(a,b)) < 0 means the delivery to market; u > t means a future trade and u = t means a spot trade). for each pair of dates t and u such that u > t , and each commodity h, there is a given family f h tu of events, which is either empty or is a partition of s. in the latter case, φu must be as fine as f h tu. assume that . assume further that if f h tu is not empty and t ≤ v ≤ u, f h vu is as fine as f h tu. in other word, if at date t, one can buy a contract for receipt at date u contingent on event b, then at a later date v , one can do the same. a portfolio plan, which was described as a trade plan by radner, is an array (θhtu(a,b)), one for each combination(h,t,u,a,b) such that for a in φt , b in f h tu, b ⊆ a, t ≤ u. the security price paid at date t in event a for receipt of commodity h at date u in event b will be denoted by phtu(a,b) .when t = u, phuu(a) is spot price of commodity h. an array p = {phtu(a,b)} will be called a commodity price system. in the situation just described, there is for each event pair (t,a), with a ∈ φt, a market in contracts for current and future receipt, with cost to be made currently in units of account. to simplify the notation, let denote the set of all pair ( t,a) such that, t = 1, 2, ..., t , a ∈ φt. endow order for m = (t,a) and n = (u,b) in 1φ is said to be as fine as partition φ′ if, for every a’ in φ ,either a ⊂ a′ or a’ ∩ a = ϕ advances in systems science and applications (2014) vol.14 no.2 185 m : m ≤ n if and only if t ≤ u, a ⊇ b. we assume short sales have low bound l, −l ∈ r++. then all of the portfolio plans can be denoted as z = ×m∈mzm with zm denoting the vector space of all arrays of number θhtu(a,b) ≥ l; h = 1, 2, ..., l; u = t, t+ 1, ..., t ; b ⊆ a,b ∈ f h tu for each m = (t, a) ∈ m . thus a portfolio θ = (θm)m∈m,θm ∈ zm, is a point in z. for m = (t, a), n = (u,b), if u ≥ t and b ⊆ a, a ∈ φt, b ∈ f h tu, we denote θhmn = θhtu(a,b); or θmn = (θ1mn, θ 2 mn, · · · θlmn) = 0. the payment at m for portfolio plan θ = (θm)m∈m , θm = (θmn)n≥m ∈ zm,given the commodity price system p = (pm)m∈m , is the inner product pm · θm. let {a,f, u} be the measure space of agents. for any agent a ∈ a, we denote his preference as ≺a, which is a binary relation on (×m∈mrl +)× (×m∈mrl +). we assume the ≺a is complete, transitive, reflexive and satisfies a) closeness: the set {(x, y) ∈ (×m∈mrl +) × (×m∈mrl +)|x≺ay} is closed in (×m∈mrl +)× (×m∈mrl +) and b) monotony: if x,y are two points in ×m∈mrl + , such that x < y, then x≺ay. the set of all preference relations on ×m∈mrl + satisfying all of these assumptions is denoted by β. we endow β with the topology of closed convergence on (×m∈mrl +) × (×m∈mrl +). for each a ∈ a, we require ε = (≺a, ea) : a → β × (×m∈mrl +) is measurable and ea is integrable, where ea ∈ ×m∈mrl + is the real endowment of agent α. we call ε large sequence of security-spot market. each agent must select a consumption plan xa = (xam)m∈m ∈ ×m∈mrl +, and a portfolio plan θa = (θam)m∈m ∈ z within his budget set. an pair (xa, θa) is called an assign if (xa, θa) : a → (×m∈mrl +)× z is integrable (we endow z with bore -field). give p = (pm)m∈m , agent a′s budget set is defined as xa(p) = {(x, θ) : x = (xm)m∈m ∈ ×m∈mrl +, θ = (θm)m∈m ∈ z, such that for each m ∈ m,pm · θm ≤ 0, xm ≤ eam + ∑ j≤m θjm} the agent a′s demand correspondence is then defined as ξa(p) = {(x, θ) ∈ xa(p)|∀(x′, θ′) ∈ xa(p), x≻ax ′} definition 2.1 an equilibrium of the economy ε is an assign (xa, θa) of consumption-portfolio plan and a price system p such that a · e, a ∈ a, (xa, θa) ∈ ξa(p) (1)∫ a θadu = 0, ∫ a xadu = ∫ a eadu (2) where ∫ a θadu = 0 means ∫ a θamdu = 0 for each m ∈ m . this definition has obviously generalized radner’s definition for pure exchange 186 guo-sheng zhang: competitive equilibrium of a sequence of incomplete markets with ... economy, that is, the finite market participants is replaced by infinite market participants. note that (1) implies a.e., a ∈ a, (xa, θa) is the best consumptionportfolio in his budget set and (2) implies the market is clear at each dataevent pair. our definition is also a development for the equilibrium described by debreu[7], which only considered spot-market with one-period. but compared with general incomplete market as that geankoplos proposed[8], we only consider future contracts as securities here. equilibrium existence our main result can be described as the following. theorem 3.1 if {a,f, u} is atomless2 and endowment satisfies ∫ a eam ≫ 03, for any m ∈ m , then economy ε has equilibrium. for each m ∈ m , let p = ×m∈mpm, pm be the set of all nonnegative vectors in zm whose coordinates sum to 1.denote the excess demand correspondence as ξ(p) = ∫ a ξa(p)du− ∫ a (ea, 0)du 4, here (ea, 0) ∈ (×m∈m )rl +×z. we will give some properties of ξ(p) restricted on . for the reason of shortening this paper, some simple proof processes similar to that used by debreu are omitted here[9]. proposition 3.1 under conditions of theorem 3.1, ξ(p) is non-empty, compact, lower bounded and upper hemicontinuous at every p ≫ 0 in p . proof. we only prove ξ(p) is upper hemicontinuous. we first prove xa(p) is continuous at each p ≫ 0 in p .the graph of correspondence xa(p) is obviously closed in p× (×m∈mrl +)×z. thus, xa(p) is upper hemicontinuous on p . to show that xa(p) is lower hemicontinuous at any point po ≫ 0 in p , we consider a pt sequence in p converging to po(t → ∞) and a point (xo, θo) ∈ xa(p o).denote xo = (xom)m∈m , θo = (θom)m∈m , θom ∈ zm . we will discuss the problem under two cases. i) pom · θom < 0 for any m ∈ m . because of ptm → pom, we have ptm · θom < 0 for large enough t. so (xo, θo) ∈ xa(p t), which satisfies the condition appearing in the definition of lower hemicontinuity. ii) ∃m1,m2 · · · ,mg ∈ m , such that pomj · θomj = 0, j = 1, 2, . . . , gand pom · θom < 0, for m ̸= m1,m2 . . . ,mg. for any j, we select a point θ ′ mj ∈ zmj satisfying pomj · θ′mj < 0. this is possible because −l > 0.thus pomj ·θ′mj < pomj ·θomj , and for large enough t, the hyperplane {θmj ∈ zmj |ptmj · θmj = 0} intersects the straight line through θ′mj and θomj in a unique point θ t mj . 2{a,f, u} is called atomless if, for any b ∈ f , u(b) > 0, there is c ∈ f such that 0 < u(c) < u(b). 3for x ∈ l, x ≫ 0 means all of its coordinates are strictly positive. 4for integrals of correspondences and their properties, see definition of part , d. of [7]. advances in systems science and applications (2014) vol.14 no.2 187 define θt as θt = (θtm)m∈m , θtm =  θom, m ̸= m1,m2, · · · ,mg θ t m, m = mj and θ t mj is between θ′mj and θomj 0, others. it is easily checked that θtm → θom(t → ∞) for any m ∈ m , and ptm · θtm ≤ 0. let xom = (xohm ), h = 1, 2, . . . , l, eam = (eahm ), h = 1, 2, . . . , l for each m ∈ m . for any h, if xohm < eahm + ∑ j≤m θohjm, we can take xthm such that xth < eahm + ∑ j≤m θthjm, and xthm → xohm (t → ∞). this is possible because xohm < eahm + ∑ j≤m θthjm for large enough t. if xohm = eahm + ∑ j≤m θohjm, then we take xthm = eahm + ∑ j≤m θthjm, which also implies xthm → xohm . therefore, we can find xt = (xtm)m∈m → xo such that (xt, θtm) satisfies the condition for lower hemicontinuity of xa(p). note that xa(p) is nonempty ((0, 0) ∈ xa(p)) and compact, according to debreu[10],ξ(p) is upper hemicontinuous at each p ∈ p , p ≫ 0 . # proposition 3.2 under the conditions of theorem 3.1, ξ(p) satisfies boundary condition, that is, if pt ≫ 0 in converges to p0 in ∂p, then d(0, ξ(pt)) → (t → ∞). proof. by the similar discussion to debreu (1982, p. 729), the proposition holds if only for a.e.,a ∈ a d(0, ξ(pt)) → ∞. suppose that the conclusion does not hold. then there is a subsequence (pt ′ ) such that d(0, ξa(p t′)) is bounded. for each t′, one can select (ct ′ , θt ′ ) ∈ ξa(p t′) in such a way that sequence (ct ′ , θt ′ ) is bounded. therefore, one can extract from (pt ′ , ct ′ , θt ′ ) a sequence (pt ′′ , ct ′′ , θt ′′ ) converging to (po, co, θo). by proposition 3.1 we have (co, θo) ∈ ξa(p o). since po ∈ ∂p , there is mo ∈ m such that some coordinates of pomo is zero. without loss of generality, we suppose the first coordinate of pomo is zero. then by replacing the first coordinate of θomo with large one, we can get θ′ = (θ′m)m∈m ∈ z, satisfy in pom · θ′m ≤ 0 and eau + ∑ j≤u θ′ju > eau + ∑ j≤u θoju. thus, we can obtain x′ = (x′m)m∈m , such that (x′, θ′) ∈ xa(p o) and x′ > x0. this contradicts the monotony of preference ≺a.# proposition 3.3 under conditions of theorem 3.1, walras law holds, that is, for any p ∈ p , p ≫ 0 and (x, θ) ∈ ∫ a ξa(p)du, we have (pm · θm)m∈m = 0 and xm = ∫ a eamdu+ ∑ j≤m ∫ a θajmdu. proof. for any a ∈ a, (x, θ) ∈ ξa(p), by the monotony of preference ≺a, it is easily proven that (pm · θam)m∈m = 0 and xam = eam+ ∑ j≤m θajm. by integrating the two sides of (pm · θam)m∈m = 0 and xam = eam + ∑ j≤m θajm, the conclusion holds. # 188 guo-sheng zhang: competitive equilibrium of a sequence of incomplete markets with ... proposition 3.4 under conditions of theorem 3.1, is convex-valued. noticing the fact that ξa(p) is convex-valued for any a ∈ a, the proposition is the direct corollary of theorem 3 of part i, d.ii of [9]. proof of theorem 3.1. for the real number b > 0, let b·pm = {b·p : p ∈ pm}. notice that if we replace p by p′ : p′ = ×m∈m (bm · pm), all of the conclusions of proposition 3.1-proposition 3.4 are true. assume the number of dimensions of pm is vm, and ∑ m∈m νm = ν. let p ′ m = νm ν · pm, p′ = ×m∈mp ′ m, and e = {p ∈ p′|p ≫ 0, there is(x, θ) ∈ ξ(p), such that ∑ n≥m,n,m∈m l∑ h=1 θhmn ≤ 0}. since ξ(p) is bounded below, when p varies in e, the θ with (x, θ) ∈ ξ(p) remains bounded. so does χ. so, by the boundary conditions, there cant be in e a sequence pt converging to p0 in ∂p ′. consequently, the distance from p ∈ e to ∂p ′ is bounded below by a strictly positive real number. thus, there is a closed convex cone c with vertex 0 in ×m∈mrm, such that e ⊂ intc and c\0 ⊂ int(×m∈mr+ m), where rm denotes the vector space of all arrays of real number ghmn ∈ r for any n = (u,b) ≥ m,b ∈ f h mn and h = 1, 2, ..., l, and r+ m is the positive cone of rm. let ξ′(p) = {θ ∈ z : thereis(x, θ) ∈ ξ(p)}. if we restrict ξ′(p) on p ′, it is easily proved that ξ′(p) is convex-valued, bounded below and satisfies walras law. we further claim ξ′(p) is upper hemicontinuous at each p ≫ 0, p ∈ p′. suppose pt in p′ converge to po ∈ p′, po ≫ 0 and θt ∈ ξ′(pt) converge to θo, which is a portfolio plan. by the definition of ξ′(p) , there is (xt, θt) ∈ ξ(pt). let u be a compact neighborhood of p0 contained in the relative interior of p′. then when t large enough, (xt, θt) ∈ ξ(pt) is uniformly bounded. thus we can take a subsequence (xt ′ , θt ′ ) → (xo, θo) with xo ∈ ×m∈mr+ m. by proposition 3.1, we have (xo, θo) ∈ ξ(po), which yields θo ∈ ξ′(po). this shows that ξ′(p) is upper hemicontinuous at any po ∈ p′, po ≫ 0. according to debreu[9], there is a point p∗ ∈ c ∩p′ such that ξ′(p∗)∩co ̸= ϕ. let θ∗ ∈ ξ′(p∗) ∩ co. by walras’ law, the point ( 1 v , 1 v , · · · , 1 v ) in p′ belongs to e, hence to c. therefore, ∑ n≤m,n,m∈m l∑ h=1 θ∗hmn ≤ 0. consequently,p∗ ∈ e, hence p∗ ∈ intc. moreover, by another application of the walras’ law, we has p∗ · θ∗ = 0. this equality together with p∗ ∈ intc and θ∗ ∈ co implies θ∗ ∈ 0. then (x∗, θ∗) is a equilibrium for price p ∗.# references [1] aumann, r. j. (1964), “space cannot be cut: why self-identity naturally includes neighbourhood”, integrated psychological behaviour, vol.32, pp.3950. advances in systems science and applications (2014) vol.14 no.2 189 [2] aumann, r. j. (1966), “existence of competitive equilibria in market with a continuum of traders”, econometrica, vol.34, pp.1-17. [3] yun, x. (2007), “optimal portfolio on optimal control”, advances in system science and applications, vol.7, no.1, pp.66-71. [4] arrow, k., debreu, g. (1954), “existence of an equilibrium for a competitive economy”, econometrica, vol.22, pp.265-290. [5] debreu, g. (1982), “existence of competitive equilibrium”, handbook of mathematical economics, vol. ii, chapter 15, pp.697-743, north holland, new york.http://www.calphysics.org/articles/chown2007.html [6] radner, r. (1972), “existence of equilibrium of plans, prices, and price expectations in a sequence of market”, econometrica, vol.40, no.2, pp.289-302. [7] hildenbrand, w. (1974), core and equilibria of a large economy, princeton university press. [8] geankoplos, j. (1990), “an introduction to general equilibrium with incomplete asset market”, journal of mathematical economics, vol.19, pp.1-38. [9] jifeng, z. and xuemou, wu. (2002), “pansystems or principles to societyeconomy systems”, advances in system science and applications, vol.2, no.1, pp.24-31. [10] hart, o. d. (1975), “on the optimality of equilibrium when the market structure is incomplete”, journal of economic theory, vol.11, pp.418-443. [11] bingan, j. etc.(2007), “equilibrium of manufactures’ r&d decision-making in defense procurement”, advances in system science and applications, vol.7, no.1, pp.117-120. [12] guosheng zhang (1998), on large finacial market, mathematics in economy (in chinese), vol.15, no.3, pp.1-6. [13] guosheng zhang (2000), “spot-financial market with large characteristics”, systems engineering-theory & paractice (in chinese), vol.20, no.3, pp.3945. [14] guosheng zhang (2007), “competitive equilibrium of large security-spot market with incomplete asset structure”, journal of systems science and complexity, vol.20, no.3, pp.386-396. corresponding author author can be contacted at: zhangguosheng@bisu.edu.cn. advances in systems science and application (2015) vol.15 no.3 233-241 computational study of the particles interaction distance under the influence of steady magnetic field n. k. lampropoulos1, e. g. karvelas 1,2 and i. e. sarris1 1 department of energy technology, technological & educational institute of athens, ag. spyridona 17, 12210 athens, greece. 2 department of civil engineering, university of thessaly, volos, greece. abstract a computational method for the estimation of the particles’ maximum interaction distance, when these are under the influence of a steady magnetic field, is presented. the computational model was developed for the simulation of the particles and the forces that are exerted on them. the chains of particles are formed under the influence of these forces. also, the mean value of cogglomeration is estimated with simulations, which have different number of particles under a steady concentration and a uniform magnetic field. keywords interaction distance, magnetic driving, aggregations 1 introduction for the treatment of serious diseases, drug is inserted into the human body through a blood vessel and through the blood circulation reaches the infected area. as a result of the above mentioned injection, the drug circulates all over the human body and causes damages in areas that are not infected. in order to overcome this difficulty, we attach drug to particles and drive them to the infected area. for the particles’ guidance to the targeted area, a magnetic resonance imaging (mri) device is needed. this method was presented for the first time in the 1970’s [1, 2]. it minimizes the side effects, as the drug is driven only to the infected area. the above mentioned method is applied with even better results on small blood vessels with low blood flow rates [3]. other factors that affect the efficiency of the proposed method is the size, the material of the particles and the magnetic intensity. the smaller a particle is, the smaller the magnetic response will be. as a result, the magnetic driving to the infected area is non-feasible [4,5]. to overcome this problem, magnetic particles are used in order to form cogglomerates [6]. it is proven that the total magnetic moment of cogglomerates is higher than this of the isolated particles, and therefore, clusters are more magnetically responsive [7]. consequently, the magnetic driving is feasible. when the aggregates reach the targeted area, break into isolated particles. this is feasible with the usage of superparamagnetic particles, which lose their magnetism when there is no magnetic field [6]. it is verified that the computational methods can simulate with accuracy or with a small discrepancy the experimental data [8]. interaction distance of par234 n. k. lampropoulos, e. g. karvelas and i. e. sarri computational study of the particles... ticles plays an important role in the accuracy of the results, when these are compared with the experimental data. the minimum interaction distance is chosen, in order to export accurate results in minimum time, because the magnetic moments that are exerted on a particle are calculated from particles that are located inside the minimum interaction distance. the aim of this study is to define the minimum particle’s interaction distance under a constant magnetic field. the forces acting on a particle are described in section 2 and the results in section 3. finally, discussion is presented in section 4. 2 numerical model for the propulsion model of the particles, four major forces are considered, i.e. the magnetic force from mris main magnet static field, as well as the magnetic field gradient force from the special propulsion gradient coils. the static field caters for the aggregation of nanoparticles, while the magnetic gradient navigates the agglomerations. moreover, the contact forces among the aggregated nanoparticles and the wall is used. the stokes drag force for each particle is considered, while only spherical particles are used in the calculation process. finally, gravitational forces due to gravity and the force due to buoyancy are added. the motion of particles is given by the newton equations. mi ∂ui ∂t = fmag i + f nc i + f tc i + f hydro i + f boy i +w i (1) ii ∂ωi ∂t = mdrag i +m con i + tmag i (2) where the index i stands for the particle i. all quantities indicated in (1) and (2) are presented in table 1 and depicted in figure 1. the bold variables represent vector quantities. the numerical model for the forces is given in [8]. fig. 1 forces and moments of a particle. advances in systems science and application (2015) vol.15 no.3 235 fig. 2 domain of particle close magnetic field inside the mri bore. 2.1 magnetic field in the mri bore the magnetic field b in the mri bore is given by b = b0 + g̃+b1 (3) and is depicted in figure 2. b0 is the mri superconducting magnet field that is constant and uniform, g̃ is the gradient field and b1 is the time dependent radio frequency field [9]. the steady magnetic field b0 is used for the cogglomeration forming and the gradient magnetic field g̃ for the navigation of the cogglomerates into the desired area. 2.2 particle’s interaction distance the magnetic interaction force, f , in the parallel direction between two spheres is inversely proportional to the fourth power of the separation distance [10,11]: f ∝ mimj (h+ ai + aj)4 (4) where, mi and mj are the magnetic moments of the center of each sphere and h is the nearest distance between the surfaces of two spheres with radii ai and aj . although, a weak interaction is being observed for particles that are more than five radii apart, as time goes by, the interaction force is getting stronger, because the particles are getting faster closer to each other [8]. this study is trying to define the minimum interaction distance di of the particle, as is depicted in figure ??, in which the mean length of aggregations remains stable. to estimate the minimum particle’s interaction distance, five series of simulations were performed with 100, 200, 300, 400, 500 particles, respectively. for each series, fifteen simulations have been performed. each time, the particle’s interaction distance varied from 5 to 20 radius. 236 n. k. lampropoulos, e. g. karvelas and i. e. sarri computational study of the particles... table 1 forces and moments simulation quantities quantity description vi velocity ωi rotational velocity t time ii mass moment of inertia matrix mi mass fmag i total applied magnetic force f nc i normal contact force f tc i tangential contact force f boy i buoyancy force w i weight force f drag i drag force mdrag i drag moments m con i contact moments tmag i torque due to the magnetic field 3 numerical method the openfoam platform was used in order to calculate the flow field and the uncoupled equations of particles’ motion [12]. the simulation process reads as follows: firstly, the fluid flow is found using the pressure correction method. upon finding the flow field (pressure, velocity) the motion of particles is evaluated by the lagrangian method by solving eq. (1) and (2) along the trajectory of each particle. the equations are solved in time using the euler time marching method. the stability of the algorithm is guaranteed with a time step of 10−6s. table 2 simulation parameters flow domain and computational grid of case 1 (3d) no concentr. (mg/ml) particles volume (m3) dimension in x-dir (m) dimension in y-dir (m) dimension in z-dir (m) 1 1.125 100 7.303× 10−11 4.18× 10−4 4.18× 10−4 4.18× 10−4 2 1.125 200 1.381× 10−10 5.17× 10−4 5.17× 10−4 5.17× 10−4 3 1.125 300 2.018× 10−10 5.94× 10−4 5.83× 10−4 5.83× 10−4 4 1.125 400 2.686× 10−10 6.6× 10−3 6.38× 10−3 6.38× 10−5 5 1.125 500 3.489× 10−10 7.04× 10−4 7.04× 10−4 7.04× 10−4 the spacing of the computational grid is equal to 2 ∗ diameter of the particles that were simulated. the summary of the domain parameters for the simulated cases is tabulated in table 2. advances in systems science and application (2015) vol.15 no.3 237 4 results (a) 0.1ms (b) 1ms (c) 3ms (d) 5ms fig. 3 snapshots of aggregation process of 500 particles under a uniform magnetic field of b0 = 0.4t . simulations with 100, 200, 300, 400, 500 particles were performed. for each number of particles, fifteen simulations were tested. each time the particle’s interaction distance varied from 5 to 20 radius with an increment of 1. the concentration and the magnetic field of the domain were 1.125 kg/m3 and 0.4t , respectively. the diameter of each particle was 11 um and the relative magnetic permeability (µr) of the fe3o4 particles was 12.3. the young modulus (y ) and the poisson ratio (ν) of the material was 109 m−2pa and 0.5, respectively. the tangential stiffness (γ) was 10 nsm−1 and the coefficient of friction (µ) was 0.5 . the density of the fluidic environment (ρf ) and particle (ρp) was 1000 and 1087 kg/m3. from figure 4 is depicted that the mean length of the chains remains steady, 238 n. k. lampropoulos, e. g. karvelas and i. e. sarri computational study of the particles... when the particle’s interaction distance is more than 11 radius. the mean length and std deviation of aggregations are steady above 11 radii. on the other hand, when the particle’s interaction distance is between 5 and 10 radii, the results show discrepancies from the experimental data. these discrepancies appeared due to truncation error, which is minimized when the particle’s interaction distance is more than 11 radii. apparently, the forces that are exerted on each particle are eliminated more than 11 radii when a magnetic field is applied. the computational platform does not estimate all the forces that are exerted on the particle from each neighbour, due to the existence of a certain interaction distance each time. for the simulations of 200 to 500 particles, the mean length of aggregations that is estimated from the computational platform is in a good agreement with the experimental data, as is depicted in figure 5. the simulations that were performed with 100 particles show a small discrepancy from the experimental data. this difference occurs due to the small numbers of particles that were used in the simulation, in comparison with the measurements that were conducted using a number of particles in the order of 106 . fig. 4 mean length of aggregations with different particle’s interaction distance. 5 discussion after all the above mentioned simulations, it is observed that under the influence of a steady magnetic field of 0.4 t and concentration of 1.125 kg/m3 the results are steady, when the particles interaction distance is above 11 diameters, as is depicted in figure 4. in figure 5, is observed that the computational platform estimates accurately the experimental data [6]. although, in the simulations of 100 particles the computational platform shows discrepancies with the experimental data, in these with more particles the proposed method is in agreement advances in systems science and application (2015) vol.15 no.3 239 with the experimental data. the small number of particles is the cause of these discrepancies. in figure 3, we present some snapshots of the aggregation process. in time of 0.1ms, the particles are isolated, as they have not form chains. from 1ms to 3ms, chains are observed, but some particles are still isolated. they have, thus, the tendency to form aggregations either with other particles that are isolated or with the chains that are already formed. in time of 5ms, the aggregation process has ended. as a result, the chains are the longest possible under the current magnetic field. small number of particles is still isolated and they will not form aggregations as time goes by, because there is no force exerted on them. fig. 5 mean length of aggregations with different number of particles. fig. 6 mean length of aggregations in each ms in figure 6, the aggregation process is presented in accordance with time. the 240 n. k. lampropoulos, e. g. karvelas and i. e. sarri computational study of the particles... mean length of the aggregation is estimated for each ms of the process till its completion. the aggregation process is not linear, because the particles react with each other. as a result, the size of the chains is constantly changing. for tests of 100 particles, the aggregation process shows discrepancies with the other simulations, due to the small number of particles that has been simulated. factors, such as size, the material of the particles and the magnetic intensity of the field are important for the interaction distance. higher intensity of the magnetic field creates stronger forces between particles. as a result, the particles interact in greater distances. therefore, the mean length of aggregates tends to be larger. bigger particles with the same magnetic intensity become into stronger magnets, due to bigger magnetic volume. as a result, smaller particles are attracted by bigger ones. the material of the particles plays an important role in the particles’ interaction distance. particles with high material magnetic susceptibility attract particles that are located in greater distances than those which have low. as a result of all the above, the interaction distance becomes greater. 6 conclusion in this work, a parametric study for the particles’ interaction distance is presented. after the simulations, it was observed that there is no interaction between the particles, which have a distance above 11 diameters. the computational method simulated the experimental results with a minimal discrepancy. acknowledgments the work is funded by the nanother program (magnetic nanoparticles for targeted mri therapy) through the operational program cooperation 2011 of gsrt, greece. discussions with dr klinakis from brfaa, greece, prof. zergioti from ntua, greece and the people from future intelligence ltd. are also acknowledged. references [1] a. senyei, k. widder, and c. czerlinski, (1978), “magnetic guidance of drug carrying microspheres,” appl. phys., vol. 49, pp. 3578–3583. [2] k. widder, a. senyei, and g. scarpelli, (1978), “magnetic microspheres : a model system of site specific drug delivery in vivo.” proc. soc. exp. biol. med., vol. 158, pp. 141–146. [3] b. yellen, z. forbes, and d. halverson, (2005), “targeted drug delivery to magnetic implants for therapeutic applications,” magn. magn. mater, vol. 293, pp. 647–54. advances in systems science and application (2015) vol.15 no.3 241 [4] q. pankhurst, j. connolly, s. jones, and j. dobson, (2003), “applications of magnetic nanoparticles in biomedicine,” physiscs d:applied physics, vol. 36, pp. 166–181. [5] o. petracic, (2010), “superparamagnetic nanoparticle ensembles,” superlattices and microstructures, vol. 47, p. 569. [6] j.-b. mathieu and s. mantel, (2009), “aggregation of magnetic microparticles in the context of targeted therapies actuated by a magnetic resonance imaging system,” j. appl. phys., vol. 106, p. 044904. [7] p. babinec, a. krafcik, m. babincova, and j. rosenecker, (2010), “dynamics of magnetic particles in cylindrical halbach array: implicationsfor magnetic cell separation and drug targeting,” med. biol. eng. comput., vol. 48, pp. 745–753. [8] n. k. lampropoulos, e. g. karvelas, and i. e. sarris, (2014), “computational modeling of an mri guided drug delivery system based on magnetic nanoparticle aggregations for the navigation of paramagnetic nanocapsules”, under review. [9] c. l. epstein and f. w. wehrli. (2005), “magnetic resonance imaging,” elsevier encyclopedia on mathematical physics, [online]. available: http://www.math.upenn.edu. [10] t. fujita and m. mamiya, (1987), “interaction forces between nonmagnetic particels in the magnetized magnetic fluid,” j. of magnetism and magnetic materials, vol. 65, pp. 207–210. [11] a. mehdizadeh, r. mei, j. f. klausner, and n. rahmatian, (2010), “interaction forces between soft magnetic particles in uniform and non-uniform magnetic fields,” acta mechanica sin., vol. 26, pp. 921–929. [12] h. g. weller, g. tabor, h. jasak, and c. fureby, (1998) “a tensorial approach to computational continuum mechanics using object-oriented techniques, computers in physics,” vol. 12, no. 6, pp. 620–631. corresponding author e. g. karvelas can be contacted at: karvelas@uth.gr microsoft word 2-yirong ying.doc 6-13 advances in systems science and applications (2010), vol.10, no.1 issn 1078-6236 international institute for general systems studies, inc. dynamics of price model with nonlinear demand function∗ yirong ying1, ke chen1 and jeffrey forrest2 1college of international business and management, shanghai university, shanghai 200444, china 2department of mathematics, slippery rock university, slippery rock, pa16057, usa abstract in this paper we derive the dynamic price model by introducing a general nonlinear form of the demand function into the traditional cobweb model, and establish three propositions about the existence and structure stability of limit cycles by applying the qualitative theory of ordinary differential equations. the dynamic characteristics under the specific nonlinear demand function in six individual situations are discussed and illustrated by the petroleum oil daily price data from january 2 to october 27, 2009. keywords demand function limit cycle dynamic characteristics of price petroleum oil price 1. introduction the cobweb theory proposed by ragnar frisch and jan tinbergen belongs to the category dynamic equilibrium analysis. it investigates the change of commodity prices in the supply-demand equilibrium and the inherent stability problems. the early cobweb model assumes that producers have the same expectations, only to produce one commodity or participate in one market, and to adjust the supply by the price of the last period. besides, the demand and supply functions are of a linear form of the prices, and the market is clear at every period (namely, the overall balance of supply and demand), etc. recently, the researchers continue to relax these assumptions or add the new ones in order to make the situation studied closer to the real market, and apply the new methods to analyze the complex performances and equilibrium conditions of the price in the nonlinear dynamic cobweb model. for instant, chiarella (1988) introduces a general nonlinear supply function into the traditional cobweb model with adaptive expectations, and finds that the dynamics of the model is driven by a chaotic type of single-hump map; hommes (1994) studies the dynamics of the price-quantities model derived from the cobweb model with adaptive expectations and nonlinear supply and demand curves, also find the chaotic dynamic behavior even if both the supply and demand functions are monotone; brock & hommes (1997) analyze the nonlinear dynamic equilibrium process by allowing the suppliers with heterogeneous expectations to switch between the naïve and rational ones freely, but restricting the linear function formation of supply and demand in the cobweb model, and dominate that a bifurcation route will change to the chaos and strange attractors when the intensity of choice to switch prediction strategies increases; goeree & hommes (2000) develop the brock & hommes’s evolutionary cobweb model by expanding the linear supply and demand to the nonlinear ones, while reaching similar conclusions. particularly, with adding other more realistic hypotheses into the traditional cobweb model, some recent literatures show that the improved models also generate chaotic dynamic behaviors. chiarella et al. (2006) discuss the cobweb model with boundedly rational heterogeneous producers (i.e., risk-averse producers), and conclude that each dimension of heterogeneity enriches the cobweb dynamics with respect to the case of homogeneous producers; choudhary & orszag (2008) study the traditional cobweb model in one market with local externalities which imply that the ∗ this work has been supported partly by the research fund of subject construction for reading material of financial economics. advances in systems science and applications (2010), vol.10, no.1 7 firms must forecast both price and quantities, and obtain the evidence of clusters of firms whose output behavior is correlated as equilibrium is reached; dieci & westerhoff (2009) consider that the producers face two kinds of commodity markets and tend to enter the more profitable one in the recent past, and find such a switching process is a further source of nonlinearity for the dynamics of prices. in addition, the more widely the new research methods are applied, the more deeply the chaotic behaviors of dynamic cobweb model are researched, including the model of stability and convergence, unstable conditions and the possibility of internal changes, and the evolutionary steps of the prediction strategy choice. for instance, ying yi-rong (1996) applies the ordinary differential equation qualitative theory, especially the “limit cycle” theory, to explore the existence and uniqueness of stable limit cycles, and the magnitude and period of price shocks in the nonlinear cobweb model. onozaki et al. (2000) use the classical homoclinic point theorem to discuss the nonlinear cobweb model with adaptive production adjustment, and find the observable chaos (strange attractors) as well as topological chaos (saddle point) associated with homoclinic points. (for the mathematical and physical mechanisms underlying the appearance of chaos, please consult with (lin and ouyang, 1998; lin, 2008)). chiarella et al. (2006) construct the geometric decay processes (gdp) of supply quantities to study the dynamic features of the nonlinear cobweb model while the supply function contains the limited memory or infinite memory. li et al. (2008) introduce the gaussian white noise into the demand function, and transform the equilibrium price model into a quasi-integrable hamiltonian system, then take the random averaging method to discuss the first-pass damage of this system and give the conditional reliability function and the probability distribution function. in short, there is a diversification trend in the study of dynamic cobweb model due to the complexity of economic and social activities. in fact, chaos is everywhere so that different assumptions and research methods are introduced and developed to explore the different sides of the complex phenomena. for enriching the studies of nonlinear cobweb models, this paper focuses on the dynamic equilibrium of price in a single commodity market with a general nonlinear formation of demand function basing on ying yi-rong’s earlier study, and derives three propositions about the existence and structure stability of limit cycles from this nonlinear cobweb model. furthermore, we discuss the dynamic characteristics under the specific nonlinear demand function in six individual situations, and find that there is the specific demand function in a shorter term illustrated by the petroleum oil daily price data. but there exist many other nonlinear forms of the demand function which are hard to determinate because of the chaotic behaviors existing in the real market. in the next section we will construct the general nonlinear demand function and the nonlinear dynamic cobweb model. the third section focuses on the existence and stability of the limit cycles under different conditions; and the fourth section expands the conclusions developed in section 3 by discussing the dynamic price equilibrium with the special nonlinear demand function form. the final section concludes the presentation of this work. 2. modeling 2.1 nonlinear demand function ying yi-rong (1996) assumes that the current demand is a nonlinear function of the price and its growth rate. that is dt dppcpccaptd )()( 2 210 ++++=α (1) where ( )d t is demand of the goods at time t; p is price of the goods at time t; 210 ,,,, cccaα are parameters. denote 2( )f p ap bp c= + + ,then ying: dynamics of price model with nonlinear demand function 8 ( ) ( ) dpd t ap f p dt α= + + (2) in eq. (2), ( )d t is made up of the simple linear impacts of price ( apα + ) and the nonlinear impact ( ( ) dpf p dt ) that explains the consuming attitude respect to the growth rate of price. this assumption makes the model closer to the real market than the simple linear demand function in the traditional cobweb model. in accordance with microeconomics views, besides the price of the commodity itself, the other important factors, generally including income, consumption preferences, and other commodity prices, can directly generate either positive or negative impact on the commodity demands. furthermore, these factors also have some unascertained impact on the consumer price sensitivity and pricing behavior of suppliers. the uncertainty involved eventually leads to the proposed nonlinear demand function. for example, with the growth in income, the consumer increases the demand but reduces the degree of consuming price sensitivity at the same time. so rational price hikes by the supplier might not necessarily lead to declining demand. in addition, in some commodity markets, the short-term change in consumption preferences may greatly improve or reduce the degree of consuming price sensitivity with respect to the price increases, which directly cause volatile fluctuations in demand. therefore, given the consuming attitude that is mainly influenced by the sensitivity of consumers to price changes, ( )f p in eq. (2) may take any function formation. if still denoted by ( )f p , then eq. (2) can be considered as the general form of any nonlinear demand function. 2.2 the nonlinear cobweb equilibrium model basing on ying yi-rong’s study (1996), in general, one can assume that the supply at time t is still a linear function of the price. the general form is as follows: ( )s t bpβ= + , 0b > (3) where ( )s t stands for the supply of the goods at time t with ,bβ being the parameters. assumed that the suppliers do pricing all the time to make the rate of change in price proportional to the amount of shortage, which is caused by the stock fallen below a certain threshold level. that is, ( ( ) )dp q t q dt λ= − − , 0λ > (4) where ( )q t is stock of the goods at time t; q is the threshold level of stock set by suppliers;λ the proportion of the stock shortage. let 0 ( ) (0) ( ( ) ( )) t q t q s t d t dt= + −∫ with (0)q being the initial level of ( )q t . then 0 [ (0) ( ( ) ( )) ] tdp q q s t d t dt dt λ= − − + −∫ , 0λ > (5) substituting eqs. (2), (3) into eq. (5), one can easily obtain the following second order differential equation 2 2 ( ) ( ) ( )d p dpf p b a p dt dt λ λ λ α β− + − = − , 0λ > (6) when p p= , then ( ) ( )s t d t= , 0dp dt = . so one can get p b a α β− = − (b a≠ ), where p is the equilibrium price. let ( )p p t p= − and substitute it into eq. (6), we have advances in systems science and applications (2010), vol.10, no.1 9 2 2 ( ) ( ) 0d p dpf p p b a p dt dt λ λ− + + − = (7) without loss of generality, we still denote ( )f p p+ by ( )f p , meanwhile let t τ λα = , 0 b a λμ = − > − , we can get 2 2 ( ) 0d p dpf p p d d μ τ τ + + = (8) let p x= , dy x dτ = − . then eq. (8) can be rewritten as the rayleigh equations by using lienard transformation. that is ( )dx y f x d dy x d τ τ ⎧ = −⎪⎪ ⎨ ⎪ = − ⎪⎩ (9) where 0 ( ) ( ) x f x f p dpμ= ∫ . 3. analysis according to the qualitative theory of ordinary differential equations, there are three conclusions in terms of eq. (9). proposition 1 if (0) 0f > , the singular point (0,0) is an unstable focus or unstable node of eq. (9); if (0) 0f < , the singular point (0,0) is a stable focus or stable node of eq. (9). proof. the characteristic equation of the linear part in eq. (9) at point (0,0) is 2 (0) 1 0fλ μ λ+ + = therefore the proposition 1 must be true according to the qualitative theory of ordinary differential equations. proposition 2 if the divergence of eq. (9) is constant within a simply connected domain area g, then there does not exist any limit cycle in eq.(9) so that all closed orbits are located within the area g. proof. let us compute the divergence of eq. (9) as follows: (9)| ( ) ( ( )) 2 0div x y f x x y ∂ ∂ = + − = > ∂ ∂ hence proposition 2 can be shown to hold true by applying theorem 1.10 derived by ye yan-qian (1984). proposition 3 if eq. (9) satisfies the following two conditions: (a) ( ) 0xf x < , when 0x ≠ and x is sufficiently small; (b) there exist a positive constant m and constants k and k ′ with k k ′> such that ( )f x k≥ ,when x m> ; ( )f x k ′≤ ,when x m< − .then eq. (9) has the stable limit cycle. proof. the function in eq. (9) shows that there exist ( )g x x≡ . then the three conditions in theorem 5.1 in the monograph of ye yan-qian (1984) can be satisfied. thus considering the conclusions of proposition 1, proposition 3 must be true. ying: dynamics of price model with nonlinear demand function 10 4. applications 4.1 case 1: constpf =)( let 0( )f p c= , 0c is a none-zero constant. substituting ( )f p into eq. (8) leads to 2 02 0d p dpc p d d μ τ τ + + = (10) (1) if 2 0( ) 4 0cμδ = − > , then 0 2( )c b a> − , and 1 2 1 2 p t p tp a e a e= + , where 1 2,a a are any constants, 1 2,p p the two distinct real roots of the characteristic equation of eq. (10). when 1 0p > and 1 2p p> , the price keeps on moving further and further away from the equilibrium price p (for the convenience of drawing the figure, assume 0p = ); when 1 0p < and 2 0p < , the price gradually moves toward the equilibrium price. figure 1 and 2 show the two kinds of price movements at some specific values 0 15c = and 0.2μ=− . in the same fashion, we can produce figure 3 – 6. 1 1.1 1.2 1.3 1.4 1.5 1.6 1.7 1.8 1.9 2 -200 -150 -100 -50 0 50 100 150 200 0 1 2 3 4 5 6 7 8 9 10 -20 -15 -10 -5 0 5 10 15 20 figure 1 (0) 20, (0) 30p p′= ± = ± figure 2 (0) 20, (0) 30p p′= ± = ± (2) if 2 0( ) 4 0cμδ = − = , then 0 2 1 2( ) c tp a a t e μ−= + , where 1 2p p= . when 0 0c > , after reaching the equilibrium price, the price continues to move further and further away from it, see figure 3 ( 0 10c = , 0.2μ =− ) for more details. when 0 0c < , the price firstly experiences larger fluctuations in the vicinity of the equilibrium price, then moves closer to it, figure 4 ( 0 10c = , 0.2μ =− ). 3 3.5 4 4.5 5 5.5 -30 -20 -10 0 10 20 30 40 50 60 70 0 1 2 3 4 5 6 -5 0 5 10 15 20 25 figure 3 (0) 20, (0) 10p p′= − = − figure 4 (0) 20, (0) 10p p′= = (3) if 2 0( ) 4 0cμδ = − < , then 0 2( )c b a< − , and 0 2 cos( ) c p ae t μ ω ψ−= + , where 0 1 2, 2 cp p iμ ω= − ± , ψ stands for the initial phase. when 0 0c > , the price contains a period of oscillations with an increasing amplitude around the equilibrium price, and then moves away advances in systems science and applications (2010), vol.10, no.1 11 from it, figure 5. however, when 0 0c < , the price contains a period of oscillations with a decreasing amplitude around the equilibrium price, and finally close in to it, figure 6. 0 50 100 150 -100 -80 -60 -40 -20 0 20 40 60 80 0 50 100 150 -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 figure 5 0 2.5c = , 0.2μ = − , (0) 2, (0) 0p p′= = figure 6 0 2.5c = − , 0.2μ = − , (0) 2, (0) 0p p′= = 4.2 case 2: the short-term fluctuations of petroleum oil price according to the u.s. government's official energy statistics, the daily data of petroleum oil prices (fob, except weekends and holidays) obviously has a growing uptrend during january 2 to october 27 in 2009 (see figure 7). what’s shown is that petroleum oil equilibrium price points were continuously broken within the interval of the 207 said days due to changes in supply and demand, and gradually went upward from the initial equilibrium price point with the short-term irregular fluctuations or a random walk. however, in a smaller time interval, the price fluctuations show the similar characteristics of dynamic changes as in figure 1– 6. for example, the data set is simply divided into two stages at the price point of june 11, 2009. in stage 1, although the petroleum oil price irregularly moves without a clear direction in the first 40 days, it goes upward and moves far away from the initial price after that; eventually it reaches the highest price on june 11. it has the similar characteristics as described in figure 3. in the stage 2, before october 10, the petroleum oil price seems to keep moving around a certain equilibrium price with decreasing amplitude. it is similar in principle to what is shown in figure 6. but it is not easy to find whether its movement within the rest of the time period would have the similar performance as those in the previous figures. 0 20 40 60 80 100 120 140 160 180 200 35 40 45 50 55 60 65 70 75 80 stage 1 stage 2 data source: http://tonto.eia.doe.gov/dnav/pet/hist/rwtcd.htm; figure 7 01/02/2009-10/27/2009 petroleum oil price trend. some evidences can be found from this real market experience. generally speaking, if these factors, such as price, income, preferences, and the prices of other commodities, change regularly without a significant fluctuation, or no war, no pestilence, no earthquakes, etc. to occur, both suppliers and consumers’ price sensitivity may vary little within a shorter term or a sufficiently shorter interval. that is, the price sensitivity can be assumed to be constant theoretically so that ( )f p in eq. (9) can be roughly seen as a constant in a short time period. ying: dynamics of price model with nonlinear demand function 12 however, if those factors change chaotically, or one or some of the events mentioned above suddenly occur, then the price sensitivity is easily changed drastically even within a sufficiently shorter interval. so ( )f p may be a linear or even nonlinear function instead of a constant. therefore, the price movements of the petroleum oil always stand for a combined effect of two individual performances even in a sufficiently small interval. one is similar to what are described in the cases of figure 1– 6, another represents a complex process. given 10 days as a sufficiently short interval for the petroleum oil price data, then there are 198 samples by sampling randomly 10 consecutive days. all the price movement may be roughly divided into two categories with ( )f p being constant or non-constant. so, let us simply choose two special samples for easy comparative analysis. additionally, assume the lag one of the initial price in the sample is the equilibrium price under the adaptive expectations, which is estimated by smoothing spline interpolation method (with the smoothing coefficient being 0.5). figure 8 shows that during april 27 to may 8, the fitting curve by fourier interpolation method goes up to cross the equilibrium price line on april 28, and keeps on moving further upward. if excluding the irregular change, and contrasting with figure 3, this curve implies that there exists a constant ( )f p greater than zero in the petroleum oil price dynamics of this sample. whereas the fitting curve obtained in the same way in figure 9 is not similar to any of those in figure 1– 6 during the 10th to 24th of february. it has a period of almost six trading days of oscillation with the same amplitude underneath the equilibrium price line, and hovers over the line in the last four days. it can be inferred that ( )f p in figure 9 is no longer a constant but a nonlinear function. unfortunately, it is not easy to find its specific functional expression. because it is more difficult to determine the existence and stability conditions of limit cycles according to proposition 3, more in depth study will be needed. 0 1 2 3 4 5 6 7 8 9 49 50 51 52 53 54 55 56 57 58 59 0 1 2 3 4 5 6 7 8 9 34 35 36 37 38 39 40 figure 8 petroleum oil price trend in 04/27-05/08/2009 figure 9 petroleum oil price trend in 02/10-02/24/2009 5. conclusions based on the studies of ying yi-rong (1996), the hypothesis of a linear demand function in the traditional cobweb model is extended to a nonlinear form. then for the given nonlinear demand function of the general form ( ) ( ) dpd t ap f p dt α= + + , of which ( )f p could be any function, the nonlinear dynamic price model is derived and transformed into the rayleigh equations, and three propositions on the limit cycles are proved to hold true by applying the qualitative theory of ordinary differential equations. on this basis, two specific cases are discussed. one is the study of the dynamics of the nonlinear price model with a constant ( )f p , where six situations of price fluctuations are graphed. another is the analysis of the petroleum oil price movement from january 2 to october 27, 2009; it is found that there may appear two kinds of situations where ( )f p is either a constant or non-constant in a sufficiently short term by contrasting to what are shown in figure 1– 6. for non-constant ( )f p , it might be caused by significant changes in price, income, preferences, and prices of other commodities, or by the advances in systems science and applications (2010), vol.10, no.1 13 appearance of one or more uncontrollable events. so, ( )f p might be a nonlinear function (to this end, please consult with (lin, 2008), where it is shown that the study of economics should be mainly about interactions of nonlinear forces). but its specific functional expression is difficult to find, constituting an open problem for future studies. references [1] carl chiarella. the cobweb model: its instability and the onset of chaos. economic modelling. 1988, 24(4): 377-384. [2] cars h. hommes. dynamics of the cobweb model with adaptive expectations and nonlinear supply and demand. journal of economic behavior & organization. 1994, 24 (3):315-335. [3] brock, w.r., hommes, c. rational route to randomness. econometrica. 1997, 65(5):1059 –1096. [4] jacob k. goeree, cars h. hommes. heterogeneous beliefs and the non-linear cobweb model. journal of economic dynamics and control, 2000, 24(5-7):761-798. [5] carl chiarella, xue-zhong he, hing hung, peiyuan zhu. an analysis of the cobweb model with boundedly rational heterogeneous producers. journal of economic behavior & organization, 2006, 61(4):750-768. [6] m. ali choudhary, j. michael orszag. a cobweb model with local externalities. journal of economic dynamics and control, 2008, 32(3):821-847. [7] roberto dieci, frank westerhoff. stability analysis of a cobweb model with market interactions. applied mathematics and computation, 2009, 215(6): 2011-2023. [8] lin lin, ying yi-rong, dang xin-yi. the significance, methodology and application about infinite. xi’an: press of northwestern university, 1996. (in chinese) [9] yi lin. systemic yoyos: some impacts of the second dimension. auerbach publications, an imprint of taylor and francis, new york, 2008. [10] yi lin and s. c. ouyang. invisible tao and realistic nonlinearity proposition. kybernetes: the international journal of cybernetics, systems and management science, vol. 27, nos. 6 – 7 (1998), pp. 809 – 822. [11] tamotsu onozaki, gernot sieg, masanori yokoo. complex dynamics in a cobweb model with adaptive production adjustment. journal of economic behavior & organization, 2000, 41(2):101-115. [12] jiaorui li, wei xu, wenxian xie, zhengzheng ren. research on nonlinear stochastic dynamical price model. chaos, solitons and fractal, 2008 (37):1391-1396. [13] yan-qian ye. limit cycle. shanghai: press of science and technology, 1984. (in chinese) microsoft word 5 peng zhao, buyun sheng--process planning service mechanism based on saas mode.doc 240-248 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc. process planning service mechanism based on saas mode peng zhao and buyun sheng hubei digital manufacturing key laboratory, wuhan university of technology, wuhan 430070, p.r.china abstract software-as-a-service (saas) is a new mode that the software is deployed as a series of services on the internet. the customers could choose the necessary function services or the better services to meet their normal and individual requirements. for process planning is a highly individual work, process planning service mechanism based on saas mode was proposed and researched. the service-oriented architecture of capp was established. process planning services were designed to service components, which could communicate each other with messages. the thought of process service bus (psb) was put forward and applied to manage the process planning services needed in real time. on this basis, a case was analyzed to show how to apply the service mechanism to gears process planning, the feasibility and validity of the process planning service mechanism were demonstrated by the case. keywords process planning software-as-a-service service-oriented architecture process service bus 1. introduction with the development of china's manufacturing informatization technology in the last decade, the technology informatization ( cad/cam ) and management informatization (pdm/erp) have made significant achievements and brought certain economic and social benefits to the country and enterprises. however, in actual production circumstance, process planning is considering as the connecting bridge between product design and manufacture, whose informatization degree lags far behind the actual application requirements[1]. the development and commercialization of capp are difficult, even become a bottleneck of the application of cims, which can be ascribed to the complicated and changeable characters of process planning project. not only does process planning exist in almost all production activities including product design, production schema, shop fabrication, stock management, quality management, sales management, cost management and other business activities, but also be related to parts’ types, product structures, technology and equipments and so on. furthermore, it is influenced by personnel practical experiences and the constraints of production management system, any change of these factors may lead to revise the process planning project. for the actual state and conditions of different enterprises are different, there are various different process planning requirements. further speaking, the individual process planning service of the capp system can become a decisive factor for the development of capp systems. a flexible multi-user service mechanism is needed to meet the enterprises' individual requirements. in detail,the system architecture, the service management and the serving model of capp must be unified, complete and agile. ref.[2-5] gave some researches on capp system architecture using components technology. the related capp system architectures are classified as traditional system architecture, which is need to be revise wholly when system updating or modifying are happened because the connections among components are tight coupling and complex and the communication standards are different. ref.[6-9] constructed capp systems with web services. although these studies were still based on traditional system architecture, they provided us one new thinking way: forming system with services, just like web services. with the emergence and application of service-oriented architecture (soa), it became a hot point doing research on capp based on soa. because the interface between the modules in this advances in systems science and applications (2011), vol.11, no.3-4 241 architecture is unified and standard, different manufacturers provide integrated solutions for better compatibility between each other, and the architecture is easy to be extended while the business expanding. the attendant problem is about service management, namely,how do the service "components" run in the service-oriented architecture? ref.[10,11] adopted the thought of enterprise service bus(esb) when building system architectures and which can be used for reference. esb shows that the services are distributed and are connected via bus-like method, which supports message communication and data dynamical interaction based on a widely accepted and open standard. process service bus (psb) is proposed to express process planning service management clearly. so far, it is still not enough to meet the multi-user individual service requirements, given that the service thought claims to choose and configure services according to the users' practical requirements. traditional capp systems are set up in the users' computers as a whole software, and most functions of the software are not be needed or are not suitable for the different applying environments. so a better serving model is needed. ref.[12-16] described one serving model named as application service provider(asp) and ref.[17,18] discussed the applications of asp. however, asp is just a primary serving model. ref.[19-22] used an advanced serving model, software-as-a-service(saas), which is more comprehensive and sophisticated. it is summarized that soa,psb and saas are coincident, trying to provide customers with flexible, diverse and personalized services, rather than a fixed, simple software. in the paper, process planning service mechanism based on saas mode is proposed and be researched detailedly. 2. saas mode saas is a mode of software service by renting software or on-line services [23]. the core of saas is deploying software into services and accessing services via network. software vendors provide services such as maintenance, daily operation and technical support service, which conveniences users to focus on their own business, regardless of the purchasing related hardware and software, and the cumbersome system maintenance and upgrade. this service mode has much difference from the traditional software services mode, especially on the ways of paying, deploying, serving, using and upgrading [24]. saas maturity model is assorted into four grades according to the data structure and deployment for customers[25] (see fig.1). fig.1. saas maturity model on the first grade, application service vendors offer individuals a customization sample, which includes independent database, virtual content, website, etc. what the software provides is not a service but an application to each customer. software vendors haven’t consider the various 242 zhao: process planning service mechanism based on saas mode needs of customers as a whole, thus the personal demands wouldn’t be better satisfied. the ideal saas model should be the fourth grade: the clients connect into the server cluster, and a layer used to balance the loads of clients is added between clients and servers, which also can deal with the customization, expansibility and the efficiency of multi-tenant. so it is regarded to meet the demands of customers excellently and solve the load problem by coordinating the multi-tenants. on the one hand, the processes of domestic commercial capp have fixed in the procedure, which results that the system can not respond to the part’s process changes to meet individual customer requirements opportunely. on the other hand, most of systems, designed by one common process planning management model, couldn’t meet individual customer requirements, so the workload of software development increases while the software life cycle shortens. capp lags behind the use of modern information-based manufacturing requirements due to the particularity and complexity of process planning. saas model provides a good solution for the development of the capp system. in this model, customers will consider that each service is provided as their needs. 3. process planning service mechanism in the saas model, to provide customers with individual services, the capp system must be studied in the aspect of process planning service mechanism, including system architecture, design & realization of process planning services and service management mechanism and other relative problems. 3.1 system architecture capp should not only be general but also adaptable to different requirements of individual users, which determines the system itself should be flexible. according to the saas model, software vendors provide process planning services to users, and users access to services after registering and searching. it should be emphasized that the provided software must be in the form of services. when service-oriented architecture (soa) is introduced to capp system, independent, standard and loosely coupled services could be provided because these services in soa architecture can be configurable, reconfigurable, scalability. in accordance with the thinking of soa[26], capp system is divided into multi-level services, business could be completed by calling services. users can invoke services dynamically and configure on-demand system and choose the right interface style, business processes and management through customizing services. the architecture is divided into four layers: the user layer, business layer, services layer and data layer (see fig.2). (1) user layer. system users are classified into enterprise users and individual users. users login in system with web browser or other network equipments directly after identity authentication. (2) business layer. in accordance with process planning task, process management modules invoke appropriate services through psb. for process planning is multi-threaded task, a number of process management modules may be needed.,but there is only one process management center that harmonizes and controls their work process. (3) services layer. this layer contains four categories of basic services: process planning service, process service, process rules service and basic system service (such as data service, logical service, agency service). each type of services contains many fine granularity services. these services can be searched and used after services registration by uddi[27], short for universal description discovery and integration. the whole system is based on the soa architecture and the services are independent, standards, loosely coupled and call or visit each other by process controlling , information mapping. (4) database layer. it supports the system data provision and contains process knowledge national natural science fund project of china (contract no.50620130441) scientific&technological project of wuhan city (contract no.200810321153) advances in systems science and applications (2011), vol.11, no.3-4 243 database, process files database, process resources database and process cases database. users can customize the type and structure of the database and add, modify or delete them without having any impact on other users’. fig.2. the framework of capp system 3.2 design & realization of process planning services generally, a service describes the input and output parameters, right, accessing control lists, security and the quality of service (priority levels, delivery, transaction characteristics, recovery semantics, etc.) and services levels, including response time, availability rate. under an ideal circumstance, process planning services could be designed in accordance with the existing definition of the domain. if the existing capp systems are expected to be reused, process planning services could be got by packaging components interface, traditional message format or apis. the common service design method is that classifying functional services wholly with top-down manner, these coarse granularity services could be further classified fine granularity services or atomic services. according to the analysis of the process planning field, seven categories services are abstracted: cad parts information service, process decision service, process edited service, manufacturing resources service, process data service, process knowledge service and output of process service. messages transfer among the seven services (see fig.3). fig.3. message passing among services 244 zhao: process planning service mechanism based on saas mode fig.4. the interfaces of process decision service in the service-oriented architecture, it is necessary to ensure that services are loosely coupled and can be dynamically discovered and binded to other services in order to deal with changes in business processes or interactive partners flexibly. so process planning services are packaged into service-oriented component models, which could afford one kind uniform invoking way for different interfaces[7].the component declaration should be took to regard service as component like as following: <component name=" processordercomponent "> <implementation.java class="processdecision.processorder "/> <reference name="addservice">processdivivdecomponent </reference> </component> 3.3 process service management all functions provided as services, only is there a reasonable mechanism for service management, users’ requirements would be met. psb is proposed to manage the services in the process planning system. it is an architecture mode, supporting virtual communication, interaction among services and services management. it affords the connection between the service vendors and requesters, even the services are not perfectly matched. the services participants don’t interact directly, but through psb, which supplies management functions and virtualizations, including the location and identity, interactive protocol, quality of service (qos), and many other virtualizations, to achieve and expand the core services definition. fig.5. the basic psb mode advances in systems science and applications (2011), vol.11, no.3-4 245 in the basic psb mode, the message flows link the various communication participants into the bus (see fig.5). some participants invoke the provided services while others just are interested in the published services information. the endpoint that services and psb interact is named service interactive point (sip), which could be web services endpoint, webspere message queue, or remote method invocation (rmi) agents. the service registering table captures the following metadata: the requirements and functions of sip, the hoped interaction ways (synchronous or asynchronous, http or jms, etc.),the requirements of qos (the first safe choice, etc.) and other information about supporting other sip. process planning services put messages on the bus or take away from the bus. sip will deal with these dynamic messages and ensure the qos of the hosting interactions. 4. application the case is about the application in a gear enterprise, which has an annual output of 1.5 million middle-module gears and 50 thousand spline shafts and other gear boxes. in accordance with the requirements of the function of the system,the coarsegranularity process planning services were: gear aided design service,gear intensity checking service,process edit service, typical process management service, users management service,gear routing service,etc. fine-granularity services were defined based on detail function requirements. interfaces and function of each service were designed. take gears routing service as an example, the correlation between gears routing service and other services was described and the interfaces could be defined (see fig.6) fig.6 the interfaces of gears routing service it was designed with java language and packaged into service component,the following java code described the interface definition and function realization of gear routing service: package services.gearrouting; public interface gearroutingservice { string getgearinfo(string info); string getgearrouting( string info ); string sendgearprocess( string routing ); } package services. gearrouting; import org.osoa.sca.annotations.*;@service(gearproecess;.class) public class gearrouting implements gearroutingservice { public string getgearinfo(string info){…} 246 zhao: process planning service mechanism based on saas mode public string getgearrouting( string info ){…} public string sendgearprocess( string outing ){..} } psb was realized by model driven architecture (mda). in practical application, the core services were retrieved and matched from the local service or network database. the so-called process planning system was in essence composed dynamically with scalability, agility and availability. fig.7 shows the dynamical correlation of services on the psb. fig.7 the dynamical correlation of services on the psb more efficient and robust process planning services were needed to be searched to meet the process planning requirements in a high extent. the user interface was different with different practical requirements. the interface of gear process planning system was designed as web pages(see fig.8). fig.8. interface of gear process planning system 5. conclusion this paper researches on process planning service mechanism, including how to compose advances in systems science and applications (2011), vol.11, no.3-4 247 process planning service-oriented architecture based saas model, design & realization of process planning services packaged as service-oriented component and process service management with psb. capp system based saas model represented the developing direction on the network, personalized, flexible, modular, and integrated. it does not only show its advantage in information technology costs, efficiency to small and medium enterprises, but also meets the needs of different business’ personalization process planning requirements. it is still the primary stage that saas mode is applied in process planning,although some achievements and progress have been made,the feasibility and practicality of process planning service mechanism based on saas mode remains to be proven in actual production. acknowledgements this paper is supported by national natural science fund project of china (contract no.50620130441), scientific and technological project of wuhan city (contact no. 200810321153) and youth science and technology chen guang project of wuhan city (contact no. 200750731289). references [1] wang j, li tj. development and future trend of capp, journal of liaoning institute of technology, 27 (1) ,2007,54-57 (in chinese) [2] sheng buyun,wang tianhu, luo dan, et al. study and practice on capp software component base based on 3d platform, machinery,45(512),2007,20-22 (in chinese) [3] zhao f l, wu s y. a cooperative framework for process planning, int. j. computer integrated manufacturing,12(2):168-178. [4] zhu jia cheng, zhang lian ting, song hui. study of capp domain component and component-base, journal of hefei university of technology(natural science),26(3),2003, 427-431(in chinese) [5] liu changyi, zhang gewei, qiu feivan, et al. development of component-based integratable capp system, china mechanical engineering,14(24),2003,2120-2124 (in chinese). [6] w.d. li.a web-based service for distributed process planning optimization,computers in industry, 56(3), 2005, 272-288. [7] w.d. li, s.k. ong, a.y.c. nee.a web-based process planning optimization system for distributed design,computer-aided design, 37(9), 2005, 921-930 [8] k.-d. bouzakis, g. andreadis, a. vakali, m. sarigiannidou. automating the manufacturing process under a web based framework ,advances in engineering software, 40(9), 2009, 956-964 [9] jong myoung ko, chang ouk kim, ick-hyun kwon.quality-of-service oriented web service composition algorithm and planning architecture ,journal of systems and software, 81( 11), 2008, 2079-2090 [10] qing li, jian zhou, qi-rui peng,et al. business processes oriented heterogeneous systems integration platform for networked enterprises ,computers in industry, 61(2), 2010, 127-144 [11] richard olejnik, teodor-florin fortiş, bernard toursel. web services oriented data mining in knowledge architecture .future generation computer systems, 25( 4), 2009, 436-443. [12] michael alan smith, ram l. kumara theory of application service provider (asp) use from a client perspective, information & management, 41( 8),2004, 977-1002 [13] heeseok lee, jeoungkun kim, jonguk kim. determinants of success for application service provider: an empirical test in small businesses, international journal of human-computer studies, 65( 9), 2007, 796-815 248 zhao: process planning service mechanism based on saas mode [14] marco l. bittencourt, edilson g. borges,et al. an application service provider for finite element analysis,advances in engineering software, 39( 11), 2008, 899-910 [15] bong-keun jeong, antonis c. stylianou. market reaction to application service provider (asp) adoption: an empirical investigation information & management, 47(3), 2010, 176-187 [16] liz mason, william lefkovics, melissa craft, et al. application service providers, configuring exchange server, 2000, 2001, 397-431 [17] william w. cato, r. keith mobley. the application service provider internet-based solution, computer-managed maintenance systems (second edition), 2002, 141-148 [18] choon seong leem, hong joo lee. development of certification and audit processes of application service provider for it outsourcing, technovation, 24( 1), 2004, 63-71 [19] the e-hub evolution: from a custom software architecture to a software-as-a-service implementation ,computers in industry, volume 61, issue 2, february 2010, pages 145-151 david concha, javier espadas, david romero, arturo molina [20] vânia gonçalves, pieter ballon.adding value to the network: mobile operators’ experiments with software-as-a-service and platform-as-a-service models, telematics and informatics, 28( 1), 2011, 12-21 [21] ye li-na, lin lan-fen. research of capp service based on saas, computer engineering, 36(22), 2010,268-271(in chinese) [22] chen p, xue hx. research of saas for small medium-sized enterprises informationization, manufacture information engineering of china, 37 (1),2008,10-13(in chinese) [23] minyifei. the difference analyze between saas with the traditional models. http://minyifei.cn/myf/post/169.html (2009) [24] wu h. software as a service in the first series of courses say: seize the strategic framework for the long tail market. (2006) [25] she w, duan zm, liu yp. research on e-commerce platform based on soa, china computer & network,11,2008, 38-41 [26] wang xh, yao sj, jiao zy, et al. research on uddi-based web service discovery, computer and modernization,65,2009,31-34 [27] rao ph, zhang yh, liu diet. soa design and implementation based-on service component architecture, computer systems & applications,.8,2008 advances in systems science and applications (2012) vol.12 no.2 103-112 the stability of milling of thin-walled workpiece tongyue wang1,2, ning he1, liang li1 and dahu liu1 1nanjing university of aeronautics and astronautics, nanjing, 210016 2huaiyin institute of technology, huaian, 223003 abstract in order to control the cutting chatter in machining of thin-walled workpieces, the dynamic milling model of thin-walled workpieces is analyzed and built based on the analysis of degrees in two perpendicular directions of tool-workpiece system. in high speed milling of 2a12 aluminum alloy, the compensation method based on the modification of inertia effect was proposed and accurate cutting force coefficients were obtained. modal parameters of tool-workpiece system were acquired via modal analysis tests. the stable lobe for high speed milling of 2a12 aluminum alloy thin-walled workpieces and limit cutting axial-depth at different cutting radial-depth were obtained. the results were verified with cutting tests. the method can be used in milling of thin-walled workpieces to select cutting parameters properly. all these work lay a reliable foundation to the further studies on the cutting chatter rules of thin-walled workpieces. keywords thin-walled workpiece, cutting chatter, dynamic milling model, inertia modification, stable lobe, limit cutting depth 1 introduction with the structural properties of light weight, high strength et al, thin-walled workpiece has been used widely in many fields such as aeronautics & astronautics, mold & die manufacturing. because of the inherent poor rigidity, complicated structures, large metal allowance and bad processing properties, the milling of thin-walled components is difficult for cutting deformation and vibration. a noted previous research on end milling of thin-walled structures was carried out in ref.[1]. a flexible thin-walled rectangular plate was clamped on three edges (cccf) was assumed, the effect of the deflection on the chip load and cutting geometry was not considered in the force calculation. sutherland and devor[2] presented an improved model to take into account the effect of the deflection on the chip load. a dynamic model for milling of a very flexible cantilever plate with rigid end mill by neglecting the time varying structural properties and the changes in the immersion boundaries was built by altintas et al.[3]. budak and altintas went forward one-step by considering the milling of a flexible cantilever plate with slender end mills that incorporate with the mechanistic force model and finite element methods[4]. lim et al. developed a mechanistic force model for predicting the machining errors caused by tool deflection[5]. budak considered the dynamic model of milling of thin-walled workpiece as a mdof(multi-degree 104 tongyue wang: the stability of milling of thin-walled workpiece of freedom) system in[6]. duncan et al. proved the frf (frequency response function) of the tip and error for measuring force coefficients had influenced the prediction accuracy of stability limit[7]. although these models are very useful to analyze the milling of thin-walled structures, the machining of thin-walled workpiece is limited due to the non-liner dependency between the forces and the continuously changing tool immersion angle and chip thickness. in this paper, considering the degrees in two perpendicular directions of toolworkpiece is analyzed and built. the authors proposed the compensation method based on the modification of inertia effect. theoretical analysis to high speed milling stability of thin-walled workpiece is carried out via high speed milling tests and modal analysis tests. the stable lobe and limit cutting axial-depth at different cutting radial-depth for high speed milling of 2a12 aluminum alloy thin-walled workpieces are obtained. 2 dynamic milling model of thin-walled workpiece milling of thin-walled workpiece is usually expressed by the dynamic milling model as shown in figure 1, the cutter and the workpiece are modeled as twodegree-of freedom structures, respectively. after the first revolution, the cutter starts leaving a wavy surface behind because of the bending vibration of the cutter in the normal direction, which is the direction of radial cutting force. when the second revolution starts, the wavy surface causes the tool-workpiece system fluctuating. hence, the resulting dynamic chip thickness is no longer constant. the general dynamic chip thickness can be divided into two parts, one is the intended static chip thickness and the other is the dynamic chip thickness produced owing to vibrations at the present time and one spindle revolution period before. it can be expressed as follows: h(ϕj) = [fz sinϕj + (υj,0 − υj)]g(ϕj) (1) where fz is the feed rate per tooth (mm/rev-tooth),ϕj is the instantaneous angular immersion of tooth j, and (υj,0 − υj) are dynamic offsets of the cutter at the previous and present tooth periods, respectively. the function g(ϕj) is a unit step function, which is used to decide the cutting status of the cutter tooth. if g(ϕj) = 1, the cutter is in cutting, and g(ϕj) = 0, the cutter is out of cutting. henceforth, the static component of the chip thickness fz sinϕj can be dropped from the above equation because it does not contribute to the regeneration dynamic chip thickness. then, the chip thickness can be written in terms of the fixed coordinate system x and y (see fig.1) as follows: h(ϕj) = [∆x sinϕj +∆y cosϕj ]g(ϕj) (2) advances in systems science and applications (2012) vol.12 no.2 105 where ∆x = x− x0 and ∆y = y − y0. here (x, y) and (x0, y0) represent the dynamic offsets of the cutter at the present and previous tooth periods, respectively. fig.1 dynamic milling model of thin-walled workpieces the tangential and radial cutting forces acting on the tooth are proportional to the chip thickness and the axial depth of cut ftj = ktah(ϕj), frj = krftj (3) where kt is the tangential milling force coefficient which is experimentally determined for a tool-workpiece material pair,kr is the ratio of radial force coefficient to tangential force coefficient, a is the cutting width or axial depth of cut. resolving the cutting forces in the x and y directions and summing the cutting forces contributed by all teeth, the milling forces formulate can be built. the dynamic milling forces can be expressed in matrix form: ( fx fy ) = 1 2 akt ( αxx αxy αyx αyy )( ∆x ∆y ) (4) where time-varying directional force coefficients are given by: 106 tongyue wang: the stability of milling of thin-walled workpiece αxx = n−1∑ j=1 −gj [sin 2ϕj + kr(1− cos 2ϕj)] αxy = n−1∑ j=1 −gj [(1 + cos 2ϕj) + kr sin 2ϕj ] αxy = n−1∑ j=1 −gj [(1− cos 2ϕj)− kr sin 2ϕj ] αyy = n−1∑ j=1 −gj [sin 2ϕj − kr(1 + cos 2ϕj)] (5) as the tool rotates, the directional force coefficients vary with time, they can be expanded into fourier series. take the average component of the fourier series expansion, the directional force coefficients are written as: αxx = 1 2 [cos 2ϕ− 2krϕ+ kr sin 2ϕ] ϕex ϕst αxy = 1 2 [− sin 2ϕ− 2krϕ+ kr cos 2ϕ] ϕex ϕst αyx = 1 2 [− sin 2ϕ+ 2krϕ+ kr cos 2ϕ] ϕex ϕst αyy = 1 2 [cos 2ϕ− 2krϕ− kr sin 2ϕ] ϕex ϕst (6) where n is the cutter tooth number, ϕst and ϕex are the entry and exit angles of a tooth, respectively. the transfer function matrix [φ(iw)] of the milling system is: [φ(iw)] = ( φxx(iw) φxy(iw) φyx(iw) φyy(iw) ) (7) where φxx(iw) and φyy(iw) are the direct transfer functions in the x and y directions, and φxy(iw) and φyx(iw) are the cross transfer functions. considering the vibration at the present and previous tooth period, system equation at the chatter frequency wc in the frequency domain can be written as: [f ]eiwct = 1 2 akt[1− eiwct][a0][φ(iwc)][f ]eiwct (8) where a0 is the directional milling matrix. advances in systems science and applications (2012) vol.12 no.2 107 3 principle of dynamic force coefficient measurement and inertia modification a force measuring chain is functionally illustrated in fig.2. in milling, the cutting force produced between the cutter and workpiece acts on the workpiece and is transmitted to the dynamometer. the cutting force leads to a deformation of the piezozlectric sensors of the dynamometer, which produce electric charges in correspondence with the deformation of the sensors. the electric charges are processed into a data file which reflects the size of cutting force following several steps of amplifier and filter. the validity of the acquired force data depends not only on the precision of the hardware devices used and their parameter settings, but also the dynamic effect of the dynamometer itself, particularly when measuring forces in high speed milling. in high speed milling, the tooth passing frequency and/or the high harmonic frequency components induced from the impact effect of the milling cutter are often in the neighborhood of the natural frequencies of the dynamometer. so the vibration of the structure including the dynamometer is exaggerated in the output signals and the measured force signals are damaged. fig.2 diagram of force measuring principle the structure including the dynamometer and workpiece can be treated as typical spring-damper-mass model. to obtain the accurate cutting force data in high speed milling, compensation of the inertia forces is necessary to improve the measurement results. it can be modified as follows: → f= → f0 − → fi (9) where → f is the modified cutting force (n), → f0 is the directly measured cutting force (n), → fi is the inertia force (n). the inertia force → fi can be written as → fi= (me +mw)a , me is the equivalent mass of the dynamometer (kg), mw is the mass of workpiece (kg), and a is the measured acceleration(m/s2). following the quick mechanistic method of calibrating and the above inertia modification approach, a set of high speed milling experiments are conducted at following cutting conditions: workpiece: 2a12 aluminum alloy 108 tongyue wang: the stability of milling of thin-walled workpiece equipment: mikron ucp 710 high speed machining center, kistler 9265b dynamometer, kd1001a acceleration sensor tool: yg813 carbide tipped tool, two flutes, diameter=20mm test parameters: slotting, axial depth of cut=1mm, radial depth of cut=3mm, spindle revolution=10000 rev/min, feed rate per tooth changes from 0.01mm/z to 0.055mm/z at 0.005mm/z increment after treating the experimental results, the values of force coefficients can be worked out: kt=4252mpa,kr=0.858. 4 stability analysis of milling of thin-walled workpiece[3,4,6] 4.1 the limit axial depth of cut after solving the characteristic equation (8) of dynamic milling system, the eigenvalue is given as: λ = − 1 2a0 (a1 ± √ a12 − 4a0) (10) where { a0 = φxx(iwc)φyy(iwc)(αxxαyy − αxyαyx) a1 = αxxφxx(iwc) + αyyφyy(iwc) (11) in [3,4,6] , altintas and budak presented the solution of stability is in detail and provided the limit depth of cut and spindle speed, i.e. stability lobes as: alim = −2πλr nkt (1 +k2) (12) n = − 60wc n(ε+ 2kπ) (13) where κ = sinwct 1−coswct , λr is the real part of the eigenvalue, ψ = arctanκ is the phase shift of the eigenvalue, ε = π − 2ψ is the phase shift between the current chatter mark and the previous chatter mark. 4.2 end milling with a flexible cutter a above mentioned yg813 carbide tipped tool with 2 flutes, 20mm diameter is used in end milling of aluminum alloy 2a12. the gage distance is 150mm from the collet. the transfer function of the cutter attached to the spindle is measured in both feed and normal directions with an impact hammer instrumented with a pcd208c02 piezoelectric force transducer and a b&w22100 acceleration sensor. the modal parameters such as modal mass, modal dampness etc are identified from modal analysis software uteklma and are given in table 1. later the modal parameters are used to simulate the stability lobes for milling advances in systems science and applications (2012) vol.12 no.2 109 of 2a12 alloy. four kinds of radial depth of cut are adopted. the results are given in fig.3. the region above the curve is unstable cutting region and under the curve is stable cutting region. table 1 identified modal parameters direction modal mass m(kg) dampness c(n ∗ s/m) stiffness k(n/m) nature frequency ωc(hz) tangential 1.37 533.1 1.67e+7 556 radial 1.32 510.0 1.57e+7 550 (a) ae = 1mm (b) ae = 2mm (c) ae = 3mm (d) ae = 4mm fig.3 stability lobe 4.3 analysis of experimental results of varying axial depth of cut two sets of varying axial depth of cut experiments are carried out to investigate the influence of axial depth of cut to cutting stability. the principle scheme is given in fig.4. the 2a12 aluminum alloy workpiece with 30 inclined angles is 110 tongyue wang: the stability of milling of thin-walled workpiece adopted in order to increase or decrease the axial depth of cut during the milling period. the above mentioned yg813 carbide tipped tool with 2 flutes, 12mm diameter is used with the 150mm gage distance measured from the collet. the milling stability is studied by analyzing the radial cutting force, which influences the stability most. during the milling period, the 3mm radial depth of cut, 10000 rev/min spindle revolution, 0.01mm/z feed rate per tooth are kept unchanged. the axial depth of cut changes from 0.1mm to 3.98mm and from 3.98mm to 0.1mm, respectively. the radial cutting force results measured via the kistler 9265b dynamometer are shown in fig.5. although the varying way is different, the radial cutting force fluctuated severely at about 2mm axial depth of cut together. the limit axial depth of cut seems to be about 2mm at 3mm radial depth of cut, 10000 rev/min spindle revolution. obviously, the measured results are in very close agreement with the results of stability lobe (see fig.3(c)). the stability lobes presented a reasonable range for selecting the cutting parameters. fig.4 diagram of varying axial depth of cut test principlee (a) axial depth of cut, 0.1mm to 3.98mm (b) axial depth of cut, 3.98mm to 0.1mm fig 5 radial cutting force of varying cutting-depth test advances in systems science and applications (2012) vol.12 no.2 111 5 conclusions on the basis of analysis of degrees in two perpendicular directions of tool-workpiece system, the dynamic milling model of thin-walled workpieces is analyzed and built. the compensation method based on the modification of inertia effect is proposed and accurate cutting force coefficients are obtained through high speed milling test. modal parameters of tool-workpiece system are acquired via modal analysis tests. the modal parameters and cutting force coefficients are used to simulate the stability lobe at 4 cutting radial depth of cut for high speed milling of 2a12 aluminum alloy thin-walled workpieces. the results are verified with varying axial depth of cut tests. all these work lay a reliable foundation to the further studies on the high speed milling stability of thin-walled workpiece. acknowledgements the authors greatly appreciate the financial support provided by the national natural science foundation of china (no.10477008), the foundation of education department of jiangsu province (no.07kjb460008) and the research foundation of dml-hyit(hgdml-0801). references [1] w.a. kline, r.e. devor, j.r. lindberg. (1982), the prediction of cutting forces in end milling with application to cornering cuts, international journal of machine tool design and research, vol.22, no.1, pp.7-22. [2] j.w. sutherland, r.e. devor. (1986), an improved method for cutting force and surface error prediction in flexible end milling systems, asme journal of engineering for industry,vol. vol.108, no.b-4, pp.269-279. [3] y. altintas, e. budak. (1995), analytical prediction of stability lobes in milling, annals of the cirp, vol.44, no.1, pp.357-362. [4] e. budak and y. altintas. (1998), analytical prediction of chatter stability in milling. part ii:application of the general formulation to common milling systems. trans, asme journal of dynamci systems, measurement and control, vol.120, no.1, pp.22-36. [5] e.m. lim, h. feng, c. menq, z. lin. (1995), the prediction of dimensional error for sculptured surface productions using the ball-end milling process. part 1: chip geometry analysis and cutting force prediction, international journal of machine tools manufacture, vol.35, no.8, pp.1149-1169. 112 tongyue wang: the stability of milling of thin-walled workpiece [6] e. budak. (2006), analytical models for high performance milling. part i: cutting forces, structural deformations and tolerance integrity, international journal of machine tools and manufacture, vol.46, no.12-13, pp.1478-1488. [7] g.s. duncan, m. kurd, t.l. schmitz. (2006), uncertainty propagation for selected analytical milling stability limit analyses, transactions of namri/sme, vol.34, no.1, pp.17-24. advances in systems science and applications (2012) vol.12 no.1 96-102 dmu-oriented design based on cases and the research on kbe system yiming qian, xu chen, bin zhang and jinjia zhou school of mechanical engineering, chongqing university, chongqing 400044, p.r.china abstract by analyzing the technology composition of digital mock-up (dmu) and knowledge based engineering (kbe) system, the dmu-oriented design methods based on cases is proposed, which is significant for the development of model design methodology. the design flow used in the kbe system is developed based on the product cases and neural network technology, and the kbe system for motorcycle design is described in detail. keywords knowledge acquisition, dmu, kbe, intelligent design 1 dmu&kbe dmu (digital mock-up) technology is the technology of product development and design growing with the development of computer, computer graphics knowledge and other emerging technologies[1-3]. there is no consistent definition so far. the main model of dmu is a dendriform relation model based on the hierarchical structure of products. it describes the assembly information, functional information, movement relation information and cooperative relation information of the whole product, and also expresses design parameter and engineering semantic constraint of products parts and describes the design information about each stage of the life circle of the whole product. fig.1 the framework of kbe advances in systems science and applications (2012) vol.12 no.1 97 kbe (knowledge-based engineering) is a computer integrated disposal technology, which provides the best solutions to engineering problems and tasks through the driving and reproduction of knowledge[4]. the four cores of kbe technology are knowledge system, knowledge acquisition, product modeling and analysis technique. knowledge system is mainly used to denote and process engineering design knowledge, facing to the engineering design personnel and embodying the intelligent level of system. knowledge acquisition technique is primarily applied in engineering knowledge acquisition, including automatic acquisition and manual acquisition, embodying the knowledge of experts in all fields, making the whole design system improve engineering design and analysis ability step by step, thereby making the system achieve the goal of kbe system of engineering design. the framework of kbe is presented in fig.1. 2 case-based design case-based design (cbd) method is a reasoning method based on case applying in the design filed. its design idea originates from human thinking modes. confronted with the new design requirements, the similar design conditions, which happened before, usually emerge into the designers mind firstly, and according to them, the designer identifies a new design scheme associated with the standardized design rules[5-6]. the contents of case consist of the description of design problems, the design rationales, the evaluation of design scheme and the final design scheme. at present, the application of artificial neural network in reusing case experience and knowledge of parts is the relative advanced technology. the design process of case-based and nn-based (neural network-based) kbe system is summarized in fig.2. the whole process involves four key technologies: case knowledge reuse, case filtering, case supplement and modification and case base maintenance. the whole process involves four key technologies: case knowledge reuse, case filtering, case supplement and modification and case base maintenance. nn(neural network) indicates the given conception or knowledge by interlinking a great deal of nerve cells and connecting weight distribution. in the engineering that uses artificial network to acquire case knowledge, these cases which are systematically filtered and have appropriate similarity and corresponding results should be provided by the special adaptive algorithm of nn to learn the samples and constant modifying connecting weight distribution to meet the demands in order to make nn acquire the same exporting cases as much as possible under the condition of the identical input. this is the ideal process for nn to learn automatically. on this occasion, the case experience and knowledge will be transformed into the joint strength of each nerve cell in nn. when the error between output case and filtration one comes to a certain precision, nn finishes 98 yiming qian:dmu-oriented design based on cases and the research on kbe system the process of learning, and stores the acquired nn into case base simultaneously and establishes the corresponding relation with the relevant case. fig.2 the design flow of case-based and nn-based kbe the structure-function model of motorcycle parts is the base of establishing artificial nn, and the artificial nn of cases can be set up according to the structurefunction model of parts. we should analyze the cases structure after choosing the motorcycle parts cases. firstly, the designer should analyze the relationship between the structure parameters of the parts and its function, and confirm what main structure parameter affects its function. secondly, the designer presents the exact numerical range of every key parameter and function parameter, and the two parameters have different function relation in different count range, which shows the relationship between structure and function. we can adjust the function relation to make the output function and the case function similar as much as possible. for instance, the modal parameter of frame structure is the primary parameter that affects the comfort of four-wheel motorcycle. when designing the frame, the main structure parameters that have influence on modal parameter are the height and width between the front and the back upright pole of the frame, the length and width between the bottom and the top horizontal pipe of frame, the whole height of frame, the height of mid upright pole, the length of frame tail, the assembly angle and the materials of frame, displayed in fig. 3. according to the design experience data that enterprise accumulates step by step, the fitting function relation of module parameter of multiple parameters can be advances in systems science and applications (2012) vol.12 no.1 99 established to make the structure parameter based on cases acquire the function parameters that are close to the cases. this is the artificial nn for system to establish the motorcycle parts and the process of finish studying. then we describe the design requirements, inputting the description into the learned nn, a integrated case experience scheme is output, and the final design scheme will be formed through modifying it according to the cad system; or, we directly call the cases combined with cad system to modify, at last, an outcome that meets the design requirements will be got. the application of nn and the directly call of cases is a side-by-side route, proceeding simultaneously. this is the design flow case-based and nn-based kbe system. fig.3 frame of four-wheel motorcycle searching cases includes the following procedures: distributing index, looking up the relative cases and choosing the best cases. the index of cases is to determine when we can use the cases in future and it indicates the parts deserved to learn in cases. when establishing the index, its relativity, the generality and the feasibility should be taken into account. the filtration of cases can be chosen with regard to the similarity of cases, and the calculation of the similarity of cases will be listed later. 2.1 the calculation of similarity of property when two properties are identical, the similarity is 1. the calculation formula of two properties is as follows: sim(ti, si) = 1− △d ti (1) in the formula, sim is the similarity function; ti is to show the certain property in design requirements of the object case; si is to show the corresponding property in the source case;△d = |ti − si|. 2.2 the definition of attribute weights the setting of property can be designed according to the requirements of nn, and the user can configure the corresponding weights to the every property according to the different focus on design requirements. 100 yiming qian:dmu-oriented design based on cases and the research on kbe system 2.3 the calculation of the whole similarity of product sim(t, s) = n∑ i=1 wi × sim(ti, si) (2) n∑ i=1 wi = 1 in the formula, sim is the whole similarity function, t is to show the object case; s is to show the case in the case base; n is to show the number of property of each case; wi is to show the weight of the property of i. the modification and supplement of case is the difficulty of cbd method, and it includes identifying the difference between case and problem, finding out the parts that require modification and reservation. according to the characteristics of case, the types of modifications are as follows: (1) direct change. when the cases fit the design requirements completely, it only requires altering the relative parameter of the previous case. (2) change with modification. do partial correction and make use of the relative knowledge in this field. (3) scheme change based on frame. store the new cases into case base after finishing designing new product 3 the development of kbe system and its application example in the motorcycle design with the powerful secondary development function of ug, we can successfully develop kbe of knowledge-oriented motorcycle based on ug, integrating it with pdm software (team center), multi-cae soft wares and database system to realize the intellectualization and being knowledgeable. the system consist mainly of four modules: the assembly-oriented design guide, the management of parts base, the analysis of man-machine engineering and the interface with pdm system. 3.1 the assembly-oriented design guide the process of this kbe module is, first, entering the assembly-oriented design guide, configuring the basic performance parameter of the designed vehicle model. kbe program forces the designer to design and assemble parts from system-class parts. because the system-class basic part of the whole motorcycle is frame, so the frame is the first part that is to be designed and assembled. fig.5 shows the main interface of assembly design guide of frame of a four-wheeled atv motorcycle. after the design parameters are completely defined in the design platformthe management system of parts base makes use of cases inquiry system to inquire advances in systems science and applications (2012) vol.12 no.1 101 the storage position of the parts model in accordance with the requirements in the parts storeroom according to the design knowledge in the knowledge storeroom and these parameters, and it will be showed on the interface for the choice of designer according to similarity. after choosing the frame, the designers can establish the new frame according to the procedure based on cases and nn. after designing the new frame, we store it named by the number of file into the parts storeroom, and at the same time register the storage position of the file into the registration table of parts. the designer can enter the design of the rest parts fixed on the frame after finishing designing frame, and the process is similar to the design of frame. fig.4 interface for frame design of oriented-assembly 3.2 the management of parts storeroom and the analysis of man-machine engineering the cases base of kbe system is to establish the parts base and the management system of parts base under the condition of three-dimension cad circumstance. these new established parts are stored into the right position of parts storeroom by the management system and registered its position information into registration table. the designers must simulate and analyze every function of the vehicle after accomplishing the assembly design of the vehicle. the following are kbe system and the integration interfaces of cae analysis software: adams, msc. patran, msc. nastran, anasys and so on. 3.3 the interface to the pdm system the main function of this module is to achieve the integration between kbe system and pdm system. the knowledge of design procedure of extracting dmu mainly include the scheme management procedure of dmu and the procedure of development assembly of the motorcycle dmu, storing the relative information with procedure into pdm system and achieving the reuse of the procedure 102 yiming qian:dmu-oriented design based on cases and the research on kbe system knowledge and experience. 4 conclusion this paper has discussed the composition of dmu and kbe technology, presented the products design procedure of kbe system based on cases and nn. then, the kbe of motorcycle design based on ug that integrates with pdm and cae is developed. the integration of dmu, pdm, cae and kbe makes the analytical technology of product, the experiences of experts and the analysis results reused in the engineering design, improving the design efficiency and quality of products greatly, and it is an important direction of intelligent development of product design. acknowledgements project of nature sustentation fund for chongqing science committee (cstc, 2007bb0116). references [1] guo gang. (2004), digital design and management of new products, chongqing: chongqing university press. [2] wu meiping, liao wenhe. (2008), “application and research of development management knowledge of virtual products in dmu development of helicopter”, science and technology of mechanical, vol.27, no.5, pp.633-639. [3] chen xu, huang zehao, guo gang. (2005), “technology research of dmu of motorcycle products development”, modern manufacture engineering, no.3, pp.69-71. [4] li zhi, jing xianlong, jia huaiyu. (2006), “expression and reuse technology of knowledge for products design”, shanghai transportation university journal, vol.40, no.7, pp.1183-1186. [5] gu jianguang, zhang weihua, xie hongyu. (2008), “knowledge acquisition technology of product design by integrating cases with experiences”, computer integration manufacture system, vol.14, no.3, pp.418-424. [6] dai rong, he yuling, he xiansong. (2009), “integration application research of cases reasoning and regulation reasoning of motorcycle by intelligent designing”, computer integration manufacture system, vol.15, no.3, pp.411423. analysis and studying of cascading failures in gene networks advances in systems science and applications (2011), vol. 11, no. 1-2 163-172 issn 1078-6236 international institute for general systems studies, inc. analysis and study of cascading failures in gene network * wang shudong 1 , shi songtao 2 , ge yanru 3 , sun longxiao 1 , xu dashun 4 , meng dazhi 1,2 1 college of information science and engineering, shandong university of science and technology, qingdao, shandong 266510, china 2 school of software engineering; college of applied science, beijing university of technology, beijing 100124, china 3 department of control science and engineering, huazhong university of science and technology, wuhan, hubei 430074, china 4 department of mathematics, southern illinois university carbondale, makanida 62958, usa email: wangshd2008@yahoo.com.cn, dzhmeng07@yahoo.com.cn abstract genome forms gene networks in terms of complicated interactions to realize its functions. further research for gene networks can help to comprehend and predict many unknown functions of genome. in this work, cascading failure models of weighted gene networks are built and the robustness of the models is also analyzed and discussed. based on the data of normal and lung adenocarcinoma stages, by simulating and analysing the cascading failures of the two gene network models, that cascading failures occur more likely in the networks for adenocarcinoma experimental groups than the one of normal control group is discovered. in the numerical experiments, we notice that nine genes of experimental group and eight genes of control group are of very strong destructibility for the robustness of experimental and control networks respectively. the failures of these genes can lead to the collapse or paralysis of the whole network. therefore, we conclude that these genes might play important roles in keeping normal level or developing lung adenocarcinoma of organisms. when applying the methods of modeling and analysing cascading failures in gene network to other diseases’ data, biomedical scientists can be enlightened for understanding the mechanism of diseases and predicting the functions of significant genes. keywords systems biology gene network cascading failure network statistics 1. introduction research on the biological functions of genome is a major issue in life science [1-3] . the study for gene networks describing the complicated genetic interactions in genome is an important way to understand biological functions [4-9] . so far, many methods have been proposed to build and analyze gene networks [4] . amy hin yan tong et al. [5] investigated the correspondence between the dense local neighborhoods in gene regulatory network and biological functions of yeasts. mark kittisopikul and gürol m. süel [6] studied the biological significance of feed-forward loop motifs in gene network of escherichia coli and discovered that most * this work was supported by the national natural science foundation of china (grant nos. 60874036, 60503002) and sdust research fund. mailto:wangshd2008@yahoo.com.cn http://www.pnas.org/search?author1=mark+kittisopikul&sortspec=date&submit=submit http://www.pnas.org/search?author1=g%c3%bcrol+m.+s%c3%bcel&sortspec=date&submit=submit 164 wang: analysis and study of cascading failures in gene network feed-forward loops have two kinds of regulatory functions. hallinan j. s. et al. [7] analyzed the relationships among the motifs, feedback loops and their dynamical features in gene regulatory network. bowers et al. [8] proposed a computational approach-logic analysis of phylogenetic profiles to identify detailed relationships among genes or proteins on the basis of genomic data, which was applied into 4873 distinct orthologous protein families of 67 fully sequenced organisms, and identified 750, 000 triplets previously unknown logic regulatory relationships. shudong wang et al. [9] constructed a logical network with 16 active genes of shoot in different external stimuli, and analyzed the dynamics of the logical network. now, the theoretical models of cascading failures and their mechanisms, prevention and control for various actual complex networks have been relatively deeply studied [10-17] . for instance, r. kinney et al. [16] analyzed the cascading failures in the north american power grid. the results show that deliberate attacks can lead to a substantial decline in the transmission efficiencies of power grid, while random failures have nearly no influence. ashley g. smart et al. [17] investigated the relationships between structure and robustness in the metabolic networks of escherichia coli, methanosarcina barkeri, staphylococcus aureus, and saccharomyces cerevisiae using a cascading failure model based on a topological flux balance criterion and found that metabolic networks are exceptionally robust compared to appropriate null models. but the reports are rare about cascading failures in gene network. this work investigates the influences of cascading failures on gene networks. based on the documental data, by comparing cascading failures of gene networks for normal and lung adenocarcinoma groups, we discover control networks are quite robuster than lung adenocarcinoma experimental ones. this indicates the change from normal organisms into lung adenocarcinoma ones may result from dysfunctions or gene mutations (considered as failures) of some genes. through numerical experiments, we notice that failures of some genes in experimental and control groups can lead to collapse or paralysis of the whole network. these genes might play important roles in keeping normal level or developing lung adenocarcinoma in organisms. this paper is organized as follows: the work background and existing methods for studying gene networks are introduced in the first part. the cascading failure model and its algorithm used in this work are presented in detail in the second part. the methods and main results of numerical experiments of cascading failure model are described based on two data sets of (normal) control group and (lung adenocarcinoma) experimental group in the third part. the obtained results are analyzed and discussed, and the possible corresponding biological significances are also pointed out in the fourth part. 2. the model and algorithm of cascading failure in this research, we consider cascading failures of complex gene networks, so we treat genes no different from nodes of complex networks. we use the capacity-load in cascading failure model. let  wevg ,, be a complex (directed or undirected) gene network with node-set  1,2, ,v n  , edge-set e and weight-set w . suppose ijw is the weight from node i to j in complex gene network g . then the edge-length from node i to j is defined as the reciprocal value ijw 1 of ijw . if 0ijw , then the edge-length from node i to j is  . the greater the weight between two nodes is, the lesser the edge-length is; the lesser the weight is, the greater the edge-length is, and vice versa. the shortest paths from node i to j are these paths corresponding to the smallest sum of edge-length in all the paths from node i to j . obviously, the shortest paths from node i to j are not always unique. suppose there exist app:ds:have app:ds:describe app:ds:always advances in systems science and applications (2011), vol. 11, no. 1-2 165 p shortest paths from node i to j . then the load of any shortest path r is defined as p 1 of the product of all the weights in r , i.e.   p w rji ij , . the load jl of node j is defined as the sum of the loads of all the shortest paths passing through node j . the capacity jc of node j is proportional to its initial load 0 jl , i.e.   01 jj lc  , 1,2, ,j n  , where constant 0  is a tolerance factor. if the load of a node is greater than its capacity, then it is called a failure node. after deleting node i , and causing is failure nodes (including node i ), then is is defined as the size of cascading failure of node i and n s d i i  as the size-ratio of cascading failure. if cfi td  , then the network breaks down, otherwise, the network doesn’t have failure. this is a criterion of network failure, where cft is the threshold of network failure. let       cfi cfi td td isign ,0 ,1 )(1 . then the percentage of failure nodes of the network n isign p n i   1 )(1 ; the largest size-ratio of cascading failure  max max , 1,2, ,ir d i n   ; the average size-ratio of cascading failure 1 n i i d r n   . let 1, 2( ) 0, i i d d sign i d d     ( d is a variable parameter). then the cumulative probability of size-ratio of cascading failure n isign ddp n i   1' )(2 )( , which indicates the probability of size-ratio id of cascading failure greater than d . obviously, maxr , r and '( )p d d are the important parameters measuring the robustness or fragility of network. based on the above mentioned definitions and symbols, we present the algorithm (cfa)of cascading failure model as follows: ① input the weight matrix of complex gene network  , ,g v e w . ② calculate initial load 0 jl of node j and its capacity   01 jj lc  , nj ,,2,1  . 1i . ③ delete node i and its incident edges in the network. ④ calculate the load of every node in the present network and compare the capacity with the load of every node. if the load is lesser than the capacity for every node in the present network, then go to ⑤, otherwise, delete every node and its incident edges whose load is greater than its capacity, go to ④. ⑤ if the size-ratio of cascading failure after deleting node i is greater than or equal to the threshold cft of network failure, then the network breaks down. ⑥ 1 ii . if ni  , then go to ③. ⑦ calculate the largest size-ratio of cascading failure maxr , average size-ratio of cascading 166 wang: analysis and study of cascading failures in gene network failure r , and the cumulative probability of size-ratio of cascading failure '( )p d d . 3. the methods and r esults of numerical experiments 3.1 data sources data used in this work are from the results of lung adenocarcinoma network studied by yuanyuan zhang et al. (the detailed data sources can be seen in). we use the mutual information network and directed (or 1-order logic) network for control group (abbreviated as n) and lung adenocarcinoma experimental group (abbreviated as ac). in the mutual information network (directed weighted network), mutual information value (u value) is the weight of networks. we denote the weight matrices of mutual information (directed-weighted) network for control and experimental groups by nm ( nl ) and acm ( acl ), respectively. table 1 the list of  ddp ' along with the change of the network threshold nett in control and experimental networks respectively. when 34.0d ,  ddp ' is equal to zero nett 0.52 0.5 0.45 0.4 d ac n ac n ac n ac n 0.01 1 1 1 1 1 1 1 1 0.02 1 1 1 1 1 0.0794 1 0.0714 0.03 1 0.1053 1 0.0870 0.1351 0.0794 0.1667 0.0714 0.04 0.2069 0.1053 0.1875 0.0870 0.1351 0.0794 0.1667 0.0714 0.05 0.2069 0.1053 0.1875 0.0870 0.1351 0.0794 0.1667 0.0714 0.06 0.2069 0.1053 0.1875 0.0870 0.1351 0.0794 0.1667 0.0714 0.07 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.08 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.09 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.1 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.11 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.12 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.13 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.14 0.1724 0.1053 0.1563 0.0870 0.1351 0.0794 0.1667 0.0714 0.15 0.1724 0.1053 0.1563 0.0870 0.1351 0 0.1667 0.0714 0.16 0.1724 0.1053 0.1563 0.0870 0.1351 0 0.1667 0.0286 0.17 0.1724 0.1053 0.1563 0.0870 0.1351 0 0.1667 0.0286 0.18 0.1724 0.1053 0.1563 0 0.1351 0 0.1667 0 0.19 0.1724 0.1053 0.1563 0 0.1351 0 0.1667 0 0.2 0.1724 0.1053 0.1563 0 0.1351 0 0.1667 0 0.21 0.1724 0.1053 0.1563 0 0.1351 0 0.1667 0 0.22 0.1724 0.0790 0.1563 0 0.1351 0 0.1667 0 0.23 0.1724 0.0790 0.1563 0 0.1351 0 0.1667 0 0.24 0.1724 0 0.1563 0 0.1351 0 0.1667 0 0.25 0.1724 0 0.1563 0 0.1351 0 0.1667 0 0.26 0.1724 0 0.1563 0 0.1351 0 0.1667 0 0.27 0.1724 0 0.1563 0 0.1351 0 0.1667 0 advances in systems science and applications (2011), vol. 11, no. 1-2 167 0.28 0.1724 0 0.1563 0 0.1351 0 0.1667 0 0.29 0.1724 0 0.0625 0 0.1351 0 0.1667 0 0.3 0.1724 0 0.0625 0 0 0 0.1667 0 0.31 0.1724 0 0.0625 0 0 0 0 0 0.32 0.069 0 0 0 0 0 0 0 0.33 0.069 0 0 0 0 0 0 0 0.34 0.069 0 0 0 0 0 0 0 3.2 the methods and results 3.2.1 the results of mutual information gene network in order to highlight the characteristics of the network structure, we analyze the changes of '( )p d d along with the network threshold nett . when taking different network thresholds, we can obtain the mutual information networks with different coarse granularities. the corresponding weight matrices nm and acm are the inputs in the above cfa algorithm. obviously, the greater the network threshold nett is, the coarser the granularity is, the more the lost information is and the computational complexity is relatively low; on the contrary, the lesser the network threshold nett is, the finer the granularity is, the less the lost information is, but the computational complexity is relatively high. the detailed data and the changing curves are in table 1 and fig. 1. from table 1 and fig. 1, it is obvious that '( )p d d of experimental network is clearly higher than that of control one under any network threshold. this indicates the ratio of nodes in experimental network which can result in cascading failures is much greater than the one of control group. with the increasing of d ,  ddp ' in control network reduces to zero earlier than in experimental one. in other words, taking certain appropriate d , control network has no failure while experimental network has more failures. moreover, with the increasing of network threshold nett , the platform value of  ddp ' of control network is 0.0714, 0.0794, 0.0870 and 0.1053 respectively, showing gradually increasing tendency. this indicates that with the decreasing of the numbers of nodes and edges, cascading failures are more likely to occur in the gene networks, namely: the robustness goes worse. the genes resulting in the cascading failures of control and experimental groups under all four network thresholds are nras, pik3ca, mapk9, top2a and fgf1, ret, wt1, tcl1a, hrk, respectively. the detailed situations can be seen in table 2 and 3. fig. 1 taking the network threshold nett as 0.4, 0.45, 0.5, 0.52 respectively, the changing curves of )( ' ddp  along with d in control and experimental networks. table 2 the list of the relative greater size-ratio *d of cascading failure of control network under different network thresholds nett , where * denotes the gene. the following presentation is similar. 0.52 0.5 0.45 0.4 app:ds:clearly app:ds:appropriate app:ds:situation app:ds:relative 168 wang: analysis and study of cascading failures in gene network gene *d (%) gene *d (%) gene *d (%) gene *d (%) nras 23.68 nras 17.39 nras 14.2857 nras 15.71 pik3ca 23.68 pik3ca 17.39 pik3ca 14.2857 pik3ca 17.14 mapk9 21.05 mapk9 17.39 mapk9 14.2857 mapk9 17.14 top2a 23.68 top2a 17.39 rbl1 14.2857 rbl1 15.71 top2a 14.2857 top2a 15.71 table 3 the list of the relative greater size-ratio *d of cascading failure of experimental network under different network thresholds nett . 0.52 0.5 0.45 0.4 gene *d (%) gene *d (%) gene *d (%) gene *d (%) fgf1 34.48 fgf1 31.25 fgf1 29.73 fgf1 30.95 ret 31.03 ret 28.125 ret 29.73 fgf2 30.95 wt1 31.03 wt1 28.125 wt1 29.73 hspb2 30.95 tcl1a 34.48 tcl1a 31.25 tcl1a 29.73 ret 30.95 hrk 31.03 hrk 28.125 hrk 29.73 wt1 30.95 tcl1a 30.95 hrk 30.95 3.2.2 the results of directed gene network to comprehensively measure the robustness and fragility of directed weighted gene network, we analyze the situations of r , maxr and '( )p d d with the changes of network thresholds nett (table 4 and table 5). from table 4 and 5, we discover that r , maxr and '( )p d d of experimental network are clearly greater than the ones of control network under any network threshold. this shows that cascading failures occur in experimental network more easily than in control one. the genes resulting in cascading failures of control and experimental networks under five network thresholds are bad, ing1, raf1, traf3 and esr2, hspb2, nov, tal1 respectively. the detailed situations can be seen in table 6 and 7. app:ds:relative app:ds:situation app:ds:clearly app:ds:situation advances in systems science and applications (2011), vol. 11, no. 1-2 169 fig. 2 taking the network threshold nett as 0.1, 0.125, 0.15, 0.175, 0.2 respectively, the changing curves of )( ' ddp  along with d in control and experimental networks. table 4 the list of r , maxr along with the change of network threshold nett in control and experimental networks respectively. nett stage no. of nodes no. of edges r maxr 0.100 ac 60 487 0.1355 0.2167 n 98 1124 0.0858 0.1531 0.125 ac 60 392 0.1385 0.2000 n 95 887 0.0756 0.1158 0.150 ac 59 338 0.1390 0.2034 n 90 700 0.0635 0.1000 0.175 ac 58 285 0.1281 0.1897 n 86 560 0.0686 0.1163 0.200 ac 58 240 0.0888 0.1552 n 77 446 0.0727 0.1169 table 5 the list of '( )p d d along with the change of network threshold nett in control and experimental networks respectively. when 22.0d ,  ddp ' is equal to zero. nett 0.1 0.125 0.15 0.175 0.2 d ac n ac n ac n ac n ac n 0 1 1 1 1 1 1 1 1 1 1 0.01 1 1 1 1 1 1 1 1 1 1 0.02 0.5167 0.3469 0.4833 0.2947 0.4237 0.2333 0.3621 0.2326 0.3448 0.2597 0.03 0.5167 0.2959 0.4833 0.2947 0.4237 0.2000 0.3621 0.1977 0.3448 0.2208 0.04 0.5167 0.2449 0.4833 0.2316 0.4237 0.1889 0.3621 0.1744 0.2931 0.1818 0.05 0.5167 0.2449 0.4833 0.2211 0.4237 0.1556 0.3621 0.1512 0.2931 0.1818 170 wang: analysis and study of cascading failures in gene network 0.06 0.5167 0.2347 0.4667 0.2211 0.4068 0.1222 0.3448 0.1512 0.2759 0.1558 0.07 0.4833 0.2143 0.4667 0.2000 0.3898 0.0889 0.3103 0.1163 0.2069 0.1299 0.08 0.4833 0.2143 0.4667 0.1684 0.3898 0.0778 0.3103 0.1163 0.2069 0.1169 0.09 0.4500 0.2143 0.4167 0.1158 0.3390 0.0333 0.2414 0.0814 0.1552 0.1169 0.10 0.4500 0.1939 0.4167 0.0632 0.3390 0.0333 0.2414 0.0233 0.1552 0.0909 0.11 0.3667 0.1837 0.3833 0.0105 0.3051 0 0.2241 0.0116 0.0517 0.0390 0.12 0.3000 0.1122 0.3000 0 0.2712 0 0.2241 0 0.0517 0 0.13 0.3000 0.0510 0.3000 0 0.2712 0 0.2069 0 0.0517 0 0.14 0.2333 0.0102 0.2333 0 0.2542 0 0.1552 0 0.0517 0 0.15 0.2333 0.0102 0.2333 0 0.2542 0 0.1552 0 0.0517 0 0.16 0.1333 0 0.1833 0 0.1864 0 0.0690 0 0 0 0.17 0.0833 0 0.0833 0 0.0339 0 0.0690 0 0 0 0.18 0.0833 0 0.0833 0 0.0339 0 0.0517 0 0 0 0.19 0.0500 0 0.0333 0 0.0169 0 0 0 0 0 0.20 0.0500 0 0.0333 0 0.0169 0 0 0 0 0 0.21 0.0333 0 0 0 0 0 0 0 0 0 0.22 0 0 0 0 0 0 0 0 0 0 table 6 the list of the relative greater size-ratio *d of cascading failure of control network under different network thresholds nett , where * denotes the gene. 0.1 0.125 0.15 0.175 0.2 gene *d (%) gene *d (%) gene *d (%) gene *d (%) gene *d (%) apc 15.31 ing1 11.58 elk1 10 traf3 11.63 bad 11.69 akt1 13.27 apc 10.53 raf1 10 bad 10.47 atf2 11.69 axl 13.27 bad 10.53 traf3 10 fas 9.30 ing1 11.69 fosl2 13.27 mll 10.53 bad 8.89 hck 9.30 apc 10.39 grb2 13.27 pml 10.53 hck 8.89 ing1 9.30 hck 10.39 bad 12.24 raf1 10.53 ing1 8.89 nras 9.30 mll 10.39 ing1 12.24 akt1 9.47 mll 8.89 raf1 9.30 traf3 10.39 nras 12.24 axl 9.47 akt1 7.78 akt1 8.14 cxcl2 9.09 sell 12.24 elk1 9.47 axl 8.14 raf1 9.09 tp53 12.24 grb2 9.47 cxcl2 8.14 sfrs3 7.79 traf3 12.24 nras 9.47 mcc 11.22 bcl2 8.42 mll 11.22 cxcl2 8.42 notch1 11.22 mapk9 8.42 mapk3 11.22 traf3 8.42 raf1 11.22 aven 8.42 rara 11.22 supt4h1 11.22 elk1 10.2 hck 9.18 app:ds:relative advances in systems science and applications (2011), vol. 11, no. 1-2 171 mapk9 9.18 table 7 the list of the relative greater size-ratio *d of cascading failure of the experimental network under different network thresholds nett , where * denotes the gene. 0.1 0.125 0.15 0.175 0.2 gene *d (%) gene *d (%) gene *d (%) gene *d (%) gene *d (%) gli2 21.67 fes 20.00 hspb2 20.34 nov 18.97 extl3 15.52 tp63 21.67 ros1 20.00 nov 18.64 wnt3 18.97 nov 15.52 ret 20.00 hspb2 18.33 erg 16.95 tcl1a 18.97 ros1 15.52 fes 18.33 wnt3 18.33 esr2 16.95 extl3 17.24 esr2 10.34 tal1 18.33 tcl1a 18.33 extl3 16.95 esr2 15.52 gli2 10.34 esr2 16.67 esr2 16.67 fes 16.95 hspb2 15.52 hspb2 10.34 extl3 16.67 gli2 16.67 cxcl3 16.95 tal1 15.52 il1a 10.34 nov 16.67 cxcl3 16.67 tal1 16.95 tp63 15.52 tal1 10.34 e2f1 15.00 il1a 16.67 wnt3 16.95 hrk 15.52 tcl1a 10.34 cxcl3 15.00 nov 16.67 tp63 16.95 hspb2 15.00 tal1 16.67 hrk 16.95 ros1 15.00 wnt3 15.00 tcl1a 15.00 4. conclusion and analysis in this research, we analyze and investigate the cascading failures in control and experimental networks. through numerical experiments, we discover: under all the network thresholds, cascading failures occur in experimental networks more easily than in control ones for undirected and directed weighted gene networks. this indicates that the normal organisms are quite robust while diseased organisms are more fragile. in table1, we notice that: with the increasing of network thresholds, the platform values of  ddp ' of control network are gradually increasing. this shows that with the decreasing of the numbers of nodes and edges, the robustness goes worse. in other words, with the increasing of the numbers of nodes and edges, the robustness goes better. this indicates the intrinsic reason of organisms functioning normally and stably maybe is that most genes play their own roles in organisms. in the process of numerical experiments, we notice that failures of genes bad, ing1, raf1, traf3, nras, pik3ca, mapk9, top2a of control group and esr2, hspb2, nov, tal1, fgf1, ret, wt1, tcl1a, hrk of experimental group under all the network thresholds result in collapse or paralysis of the whole network. this provides some useful reference informations for the normal or dieased organisms. for example, activation of gene bad may induce apoptosis in human lung adenocarcinoma cells [18] . in other words, the failure of gene bad leads to the defunctionalization of inducing the apoptosis of human lung adenocarcinoma cells and the organism might suffer from lung adenocarcinoma. gene mapk9 may enhance the stability of tumor suppressor p53 and its failure can reduce the stability of p53. thus the organism might develop into cancer. references app:ds:relative app:ds:normally app:ds:enhance app:ds:stability app:ds:stability 172 wang: analysis and study of cascading failures in gene network [1] motoki shiga, ichigaku takigawa, hiroshi mamitsuka. annotating gene function by combining expression data with a modular gene network. bioinformatics, 2007, 23(13):i468-i478. [2] marco sardiello, michela palmieri, alberto di ronza, et al.. a gene network regulating lysosomal biogenesis and function. science, 2009, 325(5939):473-477. [3] roded sharan1, trey ideker. modeling cellular machinery through biological network comparison. nature biotechnology, 2006, 24: 427-433. [4] thomas schlitt, alvis brazma. present approaches to gene regulatory network modelling. bmc bioinformatics, 2007, 8(suppl 6) s9:1-22. [5] amy hin yan tong, guillaume lesage, gary d. bader, et al.. global mapping of the yeast genetic interaction network. science, 2004, 303(5659):808-813. [6] mark kittisopikul, gürol m. süel. biological role of noise encoded in a genetic network motif. pnas, 2010, 107 (30): 13300-13305. [7] hallinan, j.s., jackway, p.t. network motifs, feedback loops and the dynamics of genetic regulatory networks. computational intelligence in bioinformatics and computational biology, 2005. cibcb '05, 1-7. [8] bowers p m, cokus s j, eisenberg d, et a1.. use of logic relationships to decipher protein network organization. science, 2004, 306(5706):2246-2249. [9] shudong wang,yan chen, qingyun wang, eryan li, yansen su, dazhi meng. analysis for gene networks based on logical relationships. journal of systems science and complexity, 2010, 23(5): 999-1011. [10] asavathiratham c. the influence model: a tractable representation for the dynamics of networked markov chains. elec. eng. and comp. sci. dept., mit, 2000. [11] dobson i, chen j, thorp j s, et al.. examining criticality of blackouts in power system models with cascading events. proceedings of 35th hawaii international conference on system sciences, 2002, 63-72. [12] moreno y, gómez j b, pacheco a f. instability of scale-free networks under node-breaking avalanches. europhys. lett., 2002, 58(4):630-636. [13] moreno y, pastor-satorras r,vázquez a, vespignani a. critical load and congestion instabilities in scale-free networks. europhys. lett., 2003, 62:630-636. [14] crucitti p, latora v, marchiori m. model for cascading failures in complex networks. phys. rev. e, 2004, 69, 045104:1-4. [15] xiaofan wang, xiang li, guanrong chen. the theory and application of complex networks. tsinghua press, beijing, 2006.4. [16] r.kinney, p.crucitti, r.albert, v.latora. modeling cascading failures in the north american power grid. eur. phys. j. b, 2005, 46:101-107. [17] ashley g. smart, luis a. n. amaral, julio m. ottino. cascading failure and robustness in metabolic networks. pnas, 2008, 105(36):13223-13228. [18] chen-tzu kuo, ming-jen hsu, bing-chang chen, et al.. denbinobin induces apoptosis in human lung adenocarcinoma cells via akt inactivation, bad activation, and mitochondrial dysfunction. toxicol lett, 2008, 177(1):48-58. http://www.nature.com/nbt/journal/v24/n4/full/nbt1196.html#a1#a1 http://www.pnas.org/search?author1=mark+kittisopikul&sortspec=date&submit=submit http://www.pnas.org/search?author1=g%c3%bcrol+m.+s%c3%bcel&sortspec=date&submit=submit http://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=10629 http://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=10629 http://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=10629 advances in systems science and applications (2012) vol.12 no.3 272-282 improved airline seat inventory control policies under parametric uncertainty of customer demand models nicholas a. nechval1, konstantin n. nechval2 and maris purgailis1 1university of latvia, evf research institute, statistics department, raina blvd 19, lv-1050 riga, latvia 2transport and telecommunication institute, applied mathematics department, lomonosov street 1, lv-1019 riga, latvia abstract most models, which are used for solving airline seat inventory control problems, are developed in the literature under the assumptions that the parameter values of the models are known with certainty. when these models are applied to solve real-world problems, the parameters are estimated and then treated as if they were the true values. the risk associated with using estimates rather than the true parameters is called estimation risk and is often ignored. when data are limited and/or unreliable, estimation risk may be significant, and failure to incorporate it into the model design may lead to serious errors. in this paper, we consider the static and dynamic problems of airline seat inventory control under parametric uncertainty, which are invariant with respect to a certain group of transformations. since common practice for airlines is to charge several different fares for a common pool of seats, this paper presents the policies that have been used to address the problem of when to refuse booking requests for a given fare level to save the seat for a potential request at a higher fare level. in this paper, we present the innovative technologies for constructing the static and dynamic policies of the airline seat inventory control.on the basis of the ‘unbiasedness performance index’. the idea of prediction of a future cumulative customer demand for the seats on a flight via the order statistics from the underlying distribution, introduced in the paper, allows one to use the invariant embedding technique in order to eliminate the unknown parameters from the problem and to use the previous and current sample data as completely as possible. the proposed unbiased static and dynamic policies are more efficient as compared with the policies, where the unknown parameters of the airline customer demand models are estimated and then treated as if they were the true values. an illustrative example is given. keywords airlines, demand, uncertainty, airline booking, optimization 1 introduction passenger reservations systems have evolved from low level inventory control processes to major strategic information systems. today, airlines and other transportation companies view revenue management systems and related information advances in systems science and applications (2012) vol.12 no.3 273 technologies as critical determinants of future success. indeed, expectations of revenue gains that are possible with expanded revenue management capabilities are now driving the acquisition of new information technology. each advance in information technology creates an opportunity for more comprehensive reservations control and greater integration with other important transportation planning functions. the airline seat inventory control problem lies at the heart of airline revenue management. it is common practice for airlines to sell a pool of identical seats at different prices according to different booking classes to improve revenues in a very competitive market. in other words, airlines sell the same seat at different prices according to different types of travelers (first class, business and economy) and other conditions. the question then arises whether to offer seats at a relatively low price at a given time with a given number of seats remaining or to wait for the possible arrival of a higher paying customer. assigning seats in the same compartment to different fare classes of passengers in order to improve revenues is a major problem of airline seat inventory control. this problem has been considered in numerous papers. for details, the reader is referred to a review of yield management, as well as perishable asset revenue management, by weatherford et al. [1], and a review of relevant mathematical models by belobaba [2]. this paper deals with the airline seat inventory control problem when customers for different fare levels are booked into a common seating pool in the aircraft. the following assumptions are made: (1) single-leg flight: bookings are made on the basis of a single departure and landing; no allowance is made for the possibility that bookings may be part of larger trip itineraries, (2) independent demands: the demands for different fare classes are stochastically independent, (3) low before high demands: the lowest fare reservations requests arrive first, followed by the next lowest, etc., (4) no cancellations: cancellations, no-shows and overbooking are not considered, (5) nested classes: any fare class can be booked into seats not taken by bookings in lower fare classes, (6) fare classes: the business and economy fare classes are considered. the first purpose of this paper is to present the innovative information technologies for constructing the static and dynamic policies of the airline seat inventory control on the basis of the ‘unbiasedness performance index’. the static and dynamic policies (unbiased) are more efficient (from the point of view of airline revenue management) as compared with the policies, where the unknown parameters of the airline customer demand models are estimated and then treated as if they were the true values. the second purpose of this paper is to introduce the idea of prediction of a future cumulative customer demand for the seats on a flight via the order statistics from the underlying distribution, where only the functional form of the 274 nicholas a. nechval:improved airline seat inventory control policies under parametric... distribution is specified, but some or all of its parameters are unspecified. this idea allows one to use the technique of invariant embedding of sample statistics in a performance index in order to eliminate the unknown parameters from the problem [3-4]. the technique represents a simple and computationally attractive statistical method based on the constructive use of the invariance principle in mathematical statistics. unlike the bayesian approach, an invariant embedding technique is independent of the choice of priors, i.e., subjectivity of investigator is eliminated from the problem. it allows one to find the improved invariant statistical decision rules, which have smaller risk than any of the well-known traditional statistical decision rules, and to use the previous and current sample data as completely as possible. 2 state-of-the-art and progress beyond airline seat inventory control is a very profitable tool in the airline industry. a major problem of airline seat inventory control is to sell the same seat at different prices according to different types of travelers (first class, business and economy) and other conditions in order to improve revenues. this problem has been considered in numerous papers. littlewood [5] was the first to propose a solution method of the seat inventory control problem for a single leg flight with two fare classes. the idea of his scheme is to equate the marginal revenues in each of the two fare classes. he suggests closing down the low fare class when the certain revenue from selling low fare seat is exceeded by the expected revenue of selling the same seat at the higher fare. that is, low fare booking requests should be accepted as long as c2 ≥ c1pr{y1 > µ1}, (1) where c1 and c2 are the high and low fare levels respectively, y1 denotes the demand for the high fare (or business) class,µ1 is the number of seats to protect for the high fare class and pr{y1 > µ1} is the probability of selling more than µ1 protected seats to high fare class customers. the smallest value of µ1 that satisfies the above condition is the number of seats to protect for the high fare class, and is known as the protection level of the high fare class customers. the concept of determining a protection level for the high fare class can also be seen as setting a booking limit, a maximum number of bookings, for the low fare class. both concepts restrict the number of bookings for the low fare class in order to accept bookings for the high fare class. it should be remarked that there is no protection level for the low fare (or economy) class;µ2 is the booking limit, or number of seats available, for the low fare class; the low fare class is open as long as the number of bookings in this class remains less than this limit. thus,is the booking limit, or number of seats available, for the low fare class; the low fare class is open as long as the number advances in systems science and applications (2012) vol.12 no.3 275 of bookings in this class remains less than this limit. thus, µ1+µ2 is the booking limit or number of seats available, for the high fare class at time. the high fare class is open as long as the number of bookings in this and low classes remain less than this limit. richter [6] gave a marginal analysis, which proved that (1) gives an optimal allocation (assuming certain continuity conditions). optimal policies for more than two classes have been presented independently by curry [7], wollmer [8], and brumelle & mcgill [9]. 3 airline booking policies which are used in practice 3.1 static airline booking policy under complete information it will be noted that (1) represents the static policy of airline seat inventory control (or airline booking) under complete information. if fθ , the probability distribution function of y1 with the parameter θ (in general, vector), is continuous and strictly increasing, the definition (1) of µ1 is equivalent to µ1 = arg ( f̄θ(µ1) = γ ) (2) where γ = c1/c2, (3) f̄θ(τj) = 1− fθ(τj). (4) 3.2 static airline booking policy under parametric uncertainty in practice, under parametric uncertainty, i.e. when the parameter θ is unknown, the performance index, f̄θ(µ1) = γ, (5) is usually used to construct the static policy given by µ1 = arg ( f̄θ̂(µ1) = γ ) , (6) where θ̂ represents the maximum likelihood estimator of θ . the performance index (5) is named as ‘maximum likelihood performance index’. the static policy (6) based on (5) is named as ‘static maximum likelihood airline booking policy’. 3.3 dynamic airline booking policy under parametric uncertainty the static policy of airline booking is optimal as long as no change in the probability distributions of the customer demand is foreseen. however, information on the actual customer demand process can reduce the uncertainty associated with the estimates of demand. hence, repetitive use of a static policy over the booking period, based on the most recent demand and capacity information, is the general way to proceed. 276 nicholas a. nechval:improved airline seat inventory control policies under parametric... 4 improved airline booking policies proposed in the paper 4.1 static unbiased airline booking policy under parametric uncertainty this policy is based on the performance index, eθ{f̄θ(µ1)} = γ, (7) which takes into account (2) and the previous data of cumulative customer demand y1 for the seats on a flight. it allows one to construct the static airline booking policy given by µ (µnb) 1 = arg ( eθ{f̄θ(µ1)} = γ ) , (8) where µ1 ≡ µ1(θ̂),θ̂ represents either the maximum likelihood estimator of θ or sufficient statistic s for θ, i.e., µ1 ≡ µ1(s) the performance index (7) is named as ‘unbiasedness performance index’.the static policy (8), which is based on (7), is named as ‘static unbiased airline booking policy’. the relative bias of the static airline booking policy is given by γ(µ1) = ∣∣eθ{f̄θ(µ1)} − γ ∣∣ γ 100% (9) 4.2 dynamic airline booking policy under complete information in this section, we consider a flight for a single departure date with m predefined reading dates at which the dynamic policy is to be updated, i.e., the booking period before departure is divided intom readings periods:(0, τ1], (τ1, τ2], . . . , (τm−1, τm] determined by them reading dates:τ1, τ2, . . . , τm. these reading dates are indexed in increasing order:0 < τ1 < τ2 < . . . < τm, where (τm−1, τm] denotes the reading period immediately preceding departure, and τm] is at departure. typically, the reading periods that are closer to departure cover much shorter periods of time than those further from departure. for example, the reading period immediately preceding departure may cover 1 day whereas the reading period 1-month from departure may cover 1 week. let us suppose that the cumulative passenger demand for the high fare class at the kth reading date (time τk, 1 ≤ k ≤ m) is y1k representing the kth order statistic from the underlying distribution with the probability distribution function gθ(y1k), where θ is a parameter (in general, vector). in other words,y1k represents the number of seats sold for the customers of the high fare class at the kth reading date. we assume that the cumulative passenger demands for the high and low fare classes are stochastically independent. each booking of a seat of the high fare class generates average revenue of c1. each booking of a seat of the low fare class generates average revenue of c2 , where c2 < c1.let µ1k be an individual protection level for the high fare class at time τk (the kth advances in systems science and applications (2012) vol.12 no.3 277 reading date).this many seats are protected for the high fare class from the low fare class. there is no protection level for the low fare class;µ2k is the booking limit for the low fare class at time τk; the low fare class is open as long as the number of bookings in this class remains less than this limit. thus,µ1k + µ2k is the booking limit for the high fare class at time τk the high fare class is open as long as the number of bookings in this and low classes remain less than this limit. the maximum number of seats that may be booked by fare classes in the next at time τk prior to flight departure is the number of unsold seats µ◦ k. under the complete information, the dynamic airline booking policy is given by µ1k = arg ( ḡθ(µ1k|y1k) = γ ) , k = 1, 2, . . . ,m− 1, (10) where ḡθ(µ1k|y1k) = 1−gθ(µ1k|y1k), (11) gθ(µ1k|y1k) represents the conditional probability distribution function of the mth order statistic y1m . the number of unsold seats protected for the high fare class from the low fare class in the next at time τk prior to flight departure is the number of unsold seats,µ◦ 1k,which is given by µ◦ 1k = min(µ◦ k, µ1k − y1k). (12) 4.3 dynamic unbiased airline booking policy under parametric uncertainty under the parametric uncertainty, the dynamic unbiased airline booking policy is given by µunb 1k = arg ( eθ{ḡθ(µ1k|y1k)} = γ ) , k = 1, 2, . . . ,m− 1, (13) where µ1k ≡ µ1k(θ̂),θ̂ represents either the maximum likelihood estimator of θ or sufficient statistic s for θ, i.e.,µ1k ≡ µ1k(s).the number of unsold seats protected for the high fare class from the low fare class in the next at time τk prior to flight departure is the number of unsold seats µ◦ 1k,which is given by µ ◦(unb) 1k = min(µ◦ k, µ unb 1k − y1k). (14) 5 mathematical preliminaries theorem 1 let x1 ≤ . . . ≤ xk be the first k ordered observations (order statistics) in a sample of size m from a continuous distribution with some probability density function fθ(x) and distribution function fθ(x) where θ is a parameter (in general, vector). then the joint probability density function of x1 ≤ . . . ≤ xk and the lth order statistics xl(1 ≤ k ≤ l ≤ m) is given by gθ(x1, . . . , xk, xl) = gθ(x1, . . . , xk)gθ(xl|xk), (15) 278 nicholas a. nechval:improved airline seat inventory control policies under parametric... where gθ(x1, . . . , xk) = m! (m− k)! k∏ i=1 fθ(xi) [ 1− fθ(xk) ]m−k , (16) gθ(xl|xk) = (m− k)! (l − k − 1)!(m− l)! [fθ(xl)− fθ(xk) 1− fθ(xk) ]l−k−1[ 1− fθ(xl)− fθ(xk) 1− fθ(xk) ]m−l fθ(xl) 1− fθ(xk) = (m− k)! (l − k − 1)!(m− l)! k−l−1∑ j=0 ( l − k − 1 j ) (−1)j [ 1− fθ(xl) 1− fθ(xk) ]m−l+j fθ(xl) 1− fθ(xk) = (m− k)! (l − k − 1)!(m− l)! m−l∑ j=0 ( m− l j ) (−1)j [fθ(xl)− fθ(xk) 1− fθ(xk) ]l−k−1+j fθ(xl) 1− fθ(xk) (17) represents the conditional probability density function of xl given xk = xk. proof. the joint density of x1 ≤ . . . ≤ xk and xl is given by gθ(x1, . . . , xk, xl) = m! (l − k − 1)!(m− l)! k∏ i=1 fθ(xi) [ fθ(xl)− fθ(xk) ]l−k−1 fθ(xl) [ 1− fθ(xl) ]m−l =gθ(x1, . . . , xk)gθ(xl|xk). (18) it follows from (4) that gθ(xl|x1, . . . , xk) = gθ(x1, . . . , xk, xl) gθ(x1, . . . , xk) = gθ(xl|xk), (19) i.e., the conditional distribution of xl, given xi = xi for all i = 1, . . . , k, is the same as the conditional distribution of xl, given only xk = xk, which is given by (17). this ends the proof. corollary 1.1. the conditional probability distribution function of xl given xk = xk is pθ{xl ≤ xl|xk = xk} =1− (m− k)! (l − k − 1)!(m− l)! l−k−1∑ j=0 ( l − k − 1 j ) (−1)j m− l + 1 + j [ 1− fθ(xl) 1− fθ(xk) ]m−l+1+j = (m− k)! (l − k − 1)!(m− l)! m−k∑ j=0 ( m− l j )[fθ(xl)− fθ(xk) 1− fθ(xk) ]l−k+j . (20) advances in systems science and applications (2012) vol.12 no.3 279 corollary 1.2.let x1 ≤ . . . ≤ xk be the first k order statistics in a sample of size m from the two-parameter weibull distribution with the probability density function fθ(x) = δ β (x β )δ−1 exp [ −( x β )δ ] (x > 0), (21) where θ = (β, δ), β > 0 and δ > 0 are the scale and shape parameters, respectively. then the conditional probability distribution function of xl given xk = xk is pθ{xl ≤ xl|xk = xk} = 1− (m− k)! (l − k − 1)!(m− l)! l−k−1∑ j=0 ( l − k − 1 j ) (−1)j m− l + 1 + j [ exp(− xδl − xδk βδ ) ]m−l+1+j . (22) theorem 2 if in (22) the scale parameter β is unknown, then the predictive probability distribution function of xl based on (xk, δ) is given by pδ {(xl xk )δ ≤ ( xl xk )δ ) } = 1− m! (l − k − 1)!(m− l)! × l−k−1∑ j=0 ( l − k − 1 j ) (−1)j m− l + 1 + j ( πk−1 s=0 [(( xl xk )δ − 1 ) (m− j +1+ j) + (m− k+1+ s) ])−1 . (23) proof.we reduce (22) to pθ {(xl xk )δ ≤ ( xl xk )δ | (xk β )δ = (xk β )δ} = 1− (m− k)! (l − k − 1)!(m− l)! l−k−1∑ j=0 ( l − k − 1 j ) × (−1)j m− l + 1 + j [ exp ( − ω[νδ − 1] )]m−l+1+j =pδ{v δ ≤ νδ|w = ω} (24) where v = xl/xk is the ancillary statistic whose distribution does not depend on the parameter β. since xk does not depend on v ,w = (xk/β) δ is the pivotal quantity, whose distribution is known and does not depend on the parameters β and δ, we eliminate the parameter β from the problem as pδ{xl ≤ xl} = ∫ ∞ 0 pθ{xl ≤ xl|xk = xk}gθ(xk)dxk, (25) where gθ(xk) = m! (k − 1)!(m− k)! f k−1 θ (xk) [ 1− fθ(xk) ]m−k fθ(xk), xk ∈ (0,∞), (26) 280 nicholas a. nechval:improved airline seat inventory control policies under parametric... represents the probability density function of the kth order statistic xk.indeed, it follows from (26) that gθ(xk)dxk = m! (k − 1)!(m− k)! [ 1− exp ( − (xk β )δ)]k−1 exp ( − (xk β )δ(m−k) ) exp ( − (x β )δ) d (x β )δ = m! (k − 1)!(m− k)! [ 1− e−ω ]k−1 e−ω(m−k+1)dω = g(ω)dω. (27) it follows from (24) and (27) that pδ{v δ ≤ νδ} = ∫ ∞ 0 pδ{v δ ≤ νδ|w = ω}g(ω)d(ω) =1− m! (l − k − 1)!(m− l)! l−k−1∑ j=0 ( l − k − 1 j ) (−1)j m− l + 1 + j( πk−1 s=0 [ (νδ − 1)(m− l + 1 + j) + (m− k + 1 + s) ])−1 . (28) now (23) follows from (28). this ends the proof. corollary 2.1.if the parameter δ = 1 , i.e., we deal with the exponential distribution, then the predictive probability distribution function of xl based on xk is given by p {(xl xk ≤ xlxk )} =1− m! (l − k − 1)!(m− l)! × l−k−1∑ j=0 ( l − k − 1 j ) (−1)j m− l + 1 + j( πk−1 s=0 [( xl xk − 1 ) (m− l + 1 + j) + (m− k + 1 + s) ])−1 . (29) 6 illustrative example of airline booking policies let x1, . . . , xn be the random sample of the previous independent observations of the cumulative customer demand for the high fare class, which follow the exponential distribution with the probability density function (21)(δ = 1) , where the parameter β is unknown. then the static policies of airline booking under parametric uncertainty are given as follows. advances in systems science and applications (2012) vol.12 no.3 281 the static maximum likelihood airline booking policy follows from (6): µml 1 = ln γ−s/n, (30) where s = ∑n i=1xi is the sufficient statistic for β,with v = s/β ∼ f(ν) = 1 γ(n) νn−1 exp(−ν), ν ≥ 0, (31) and the relative bias, r(µml 1 ) = ∣∣eθ{f̄θ(µ ml 1 )} − γ ∣∣ γ 100% = 1 + (1 + ln γ−1/n)−1 − γ γ 100%. (32) if, say,n = 1 and γ = 0.4, then rrb(µ ml 1 ) = 30%.thus, in this example the static maximum likelihood airline booking policy has the relative bias equal to 30%.it follows that the protection level for customers of the high fare class will be determined incorrectly. this may lead to serious loss. the static unbiased airline booking policyfollows from (8): µunb 1 = [γ−1/n − 1]s, (33) where the relative bias r(µunb 1 ) = 0. the dynamic unbiased airline booking policyfollows from (13) and (29): µunb 1k = arg ( m! (m− k − 1)! m−k−1∑ j=0 ( m− k − 1 j ) (−1)j 1 + j ( πk−1 s=0 [(µ1k y1k − 1 ) (1 + j) + (m− k + 1 + s) ])−1 = γ ) , k = 1, 2 . . . ,m− 1, (34) µ ◦(unb) 1k = min ( µ◦ k, µ unb 1k ,−y1k ) . (35) 7 conclusion the methodology, which is developed in this paper for the use in the airline industry under parametric uncertainty of airline customer demand models, may be found to be useful in other industries such as hotels, car rental companies, shipping companies, etc. while the details of problems considered in this project can change significantly from one industry to the next, the focus is always on making better demand decisions and not manually with guess work and intuition but rather scientifically with models and technology, all implemented with disciplined processes and systems. 282 nicholas a. nechval:improved airline seat inventory control policies under parametric... acknowledgements this research was supported in part by grant no. 06.1936, grant no. 09.1014, and grant no. 09.1544 from the latvian council of science and the national institute of mathematics and informatics of latvia. references [1] weatherford l.r., bodily s.e. and pfeifer p.e. (1993), “modeling the customer arrival process and comparing decision rules in perishable asset revenue management situation”, transportation science, vol.27, pp.239-251. [2] belobaba p.p. (1987), “airline yield management: an overview of seat inventory control”, transportation science, vol.21, pp.66-73. [3] nechval n.a., berzins g., purgailis m., nechval k.n. (2008), “improved estimation of state of stochastic systems via invariant embedding technique”, wseas transactions on mathematics, vol.7, pp.141-159. [4] nechval n.a., nechval k.n., purgailis m., rozevskis u. (2011), “improvement of inventory control under parametric uncertainty and constraints”, in: dobnikar, a., lotric, u., ster, b. (eds.), adaptive and natural computing algorithms, lncs, vol.6594, part ii, pp.136-146, springer, heidelberg. [5] littlewood k. (1972), “forecasting and control of passenger bookings”, in: proceedings of the 12th agifors symposium, pp.95-117, american airlines, new york. [6] i richter h. (1982), “the differential revenue method to determine optimal seat allotments by fare type”, in: proceedings of the xxii agifors symposium, pp.339-362, american airlines, new york. [7] curry, r.e. (1990), “optimal airline beat allocation with fare classes nested by origins and destinations”, transportation science, vol.24, pp.193-203. [8] wollmer, r.d. (1992), “an airline seat management model for a single leg route when lower fare classes book first”, operations research, vol.40, pp.2637. [9] brumelle, s.l. and mcgill, j.i. (1993), “airline seat allocation with multiple nested fare classes”, operations research, vol.41, pp.127-137. corresponding author nicholas nechval can be contacted at:jnechval@junik.lv мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 29-42 a new decision method for multi-criteria decision making with numerical values based on criteria reduction zhaobin li 1 , jian liu 2 , zhuo zhang 1 , si-feng liu 1 1 school of economics and management, nanjing university of aeronautics and astronautics, nanjing, china, 210016 2 school of economics and management, nanjing university of science and technology, nanjing, china, 210094 annotation the work is contribution a new decision method to address the challenge (large number of criteria) in multi-criteria decision making (mcdm) problems with numerical values. this new method involves criteria reduction based on the rough set theory and the relation of criteria values (tolerance and advantage relations). using this method and building a discrenibility matrix for numerical value mcdm problems, find useful criteria and avoid useless criteria. then, we find a new way to obtain the weights based on the discernibility matrix when criteria weights of alternatives are completely unknown. later, we also propose a new method to rank the alternatives according to weighted combinatorial advantage values (wcav). finally, we use a realistic voting example to demonstrate the proposed method. key words: multi-criteria decision making, criteria reduction, discernibility matrix, obtain weight, relation. 1 introduction with the development of information technology, most decision makers (dms) face the problem of how to make a wise decision when there are massive data in the decision table. how can we filter information is a potential application and development area for mcdm/maut in an internet or mobile environment that was proposed by wallenius et al. on management science, 2008 [1]. undoubtedly, we need to find out the useful data that really affect the decision making also take to human subjective. the related problems have been one of the most popular research topics in decision making science since 2008 [2-4]. massive data of mcdm problems contain three situations that are large number of criteria, large 4umber of alternatives or both. in this paper, we focus on the problems which are large number of criteria in the decision table. extracting useful information from large quantity of uncertain problems has become an important research field in computer science-attribute reduction [5-6]. rough set theory has been recognized as one of the most powerful techniques to deal with uncertainty problems since its appearance in 1982 [7-8]. the original rough set approach validated to be very useful in dealing with discrete problems. rough set theory [9-10] is based on equivalence relation and captures useful information from a great deal of information through attribute reduction is a basic research method. attribute (criteria) reduction find useful information by using smaller criteria set b a which is to make the criteria set a describes replaced and described by b. in this paper, we want to find out the useful criteria set b from a, and use b to make a decision will address the problem proposed by wallenius et al. on management science, 2008. 30 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… obtaining criteria weights is also an important research topic of mcdm problems. for uncertain of criteria weights problems, there are several methods to obtain them, such as obtaining criteria weights approach of subject [11-12], objective obtaining criteria weights approach[13-14], subjective and objective obtaining criteria weights approach [15], the owa operator weights [16] and feedback model [17], etc. as for the uncertain criteria weights problems, most approaches obtain the criteria weights according to their deviation of criteria values to facilitate the ranking of alternatives. generally speaking, the bigger the deviation of criteria values are, the larger the weight will be [18]. in reality, we find that some criteria values change lager than others but these criteria only have a little influence on the result, while some tiny changes of few criteria would lead to different consequences for some mcdm problems. according to the traditional methods of mcdm problems, we need several procedures, such as unifying criteria and obtaining criteria weights and information fusion as well as ranking and selecting the most desirable alternative(s) [18-19]. in order to rank alternatives, we need to compare expectations of combinatorial criteria values [20-21]. but we can not filter off the absolute disparity through unifying criteria, so different criteria are incomparable even if we unify them into the same meaning. in this paper, we think the criteria can be compared only on the same criteria. some errors would be produce in the processing of uniform criteria. in order to filter off these errors, we propose through comparing the weighted combinatorial advantage values (wcav) [22] to rank alternatives in this paper. this paper is organized as follows: in section 2, for numerical value mcdm problems, there are a large number of criteria in the decision table. we propose a criteria reduction technique based on tolerance and advantage relation and rough set theory to find out useful criteria, respectively. in section 3, we propose a new way of obtaining criteria weights according to the discernibility matrix and using weighted combinatorial advantage value (wcav) instead of traditional methods to rank those alternatives. in section 4, we validate the method a useful and effective tool for mcdm problems. finally, section 5 discusses the conclusion and future work. 2 a large number criteria mcdm problems and criteria reduction in this section, we will give one type of numerical value mcdm problems that involve a large number of criteria in decision tables. the primary goal of mcdm problems is to rank the alternatives and select the most desirable one(s). to achieve this goal, several processing steps are needed to compare the alternatives based on multiple criteria. in each processing step, specific algorithms and operations are involved. the problems concerned in this paper are how to find out the useful criteria and obtain weights of useful criteria and to rank alternatives or select the most desirable alternative(s). 2.1 addressing numerical value mcdm problems example mcdm problem (supplier choice): the commercial aircraft corporation of china, ltd. (cacc) builds huge commercial aircrafts to serve commercial airlines in china. to build aircrafts the company needs to buy and use some key parts from international or domestic suppliers. therefore, the cacc must make a scientific decision to choose the most desirable supplier that relate to the success of commercial aircraft program. there are lots of complicated factors that affect decision makers (dms) to make decision and they need to combine all information for every supplier and analyze them as well as select the most desirable supplier(s) [22]. suppose that there are five international suppliers in the first round competing for the cacc demand of some key parts of the huge commercial aircrafts, and these five suppliers are represented by 1 2 3 4 5{ , , , , }a a a a aa . suppose we invite 100 experts to make judgments for the sake of obtaining the degrees to which alternative ai satisfies and does not satisfy criteria cj (i=1, app:ds:comparability advances in systems science and application(2016) vol.16 no.4 31 2, 3, 4, 5; j=1, 2, ……, m). there are two kinds of poll results “yes” or “no” to the question whether alternative ai satisfies criteria cj. then we should choose the most desirable choice based on the results in table 1. table 1 the result of “yes” answers from 100 experts u c1 c2 …… cm …… a1 a2 a3 a4 a5 55 61 55 65 89 50 62 55 65 85 …… …… …… …… …… 58 60 70 65 55 …… …… …… …… …… 2.2.1 the challenge of numerical value mcdm problems obviously, there are a large number of criteria in decision table 1. the first problem is how to make a wise decision within limited time when the dms face a large number of criteria? what kind of decision support do dms want under large number of criteria environment? to address the problem, we need to find out useful criteria that really affect the decision results. this is a significant potential research area for mcdm problems. 2.2.2 the principles and methods of criteria reduction for numerical value mcdm problems we often face a question whether we can remove some criteria from an information table while preserving its basic properties, that is, whether a table contains some superfluous criteria. through criteria reduction we want to find out useful information that really affects the processing of making decision. this work is to our best knowledge the first one that applies rough set theory [7-10] to do criteria reduction. in addition, the rough set theory contains many attributes reduction [5-6] techniques, such as the reduction of attributes based on similarity relation [23], advantage relation or disadvantage relation [22] and automatic threshold estimation [24] etc. in this paper, we apply rough set theory to find out useful information from the decision table by using criteria reduction based on tolerance relation and advantage relation of criteria value. now, let’s use the concrete example shown in table 1 to illustrate our idea in using tolerance relation or advantage relation and rough set theory to find out useful information. generally speaking, there are three different type criteria  the criteria of cost type (the smaller of criteria values, the better)  the criteria of benefit type (the larger of criteria values, the better)  the criteria of middle type (at a special point of criteria values, the better) obviously, all the criteria belong to benefit type in table 1. there are two different criteria values of alternatives a1 and a2 on criterion c1 are f(a1, c1) and f(a2, c1), such as f(a1, c1)=55, f(a2, c1)=61. obviously, both of them have different criteria values on c1, i.e., f(a1, c1)< f(a2, c1). for benefit criteria, if f(a1, c1)is smaller than f(a2, c1), it means that the alternative a2 locates at a more advantage position than a1 on criterion c1 when we want to compare them. in other words, alternative a2 is better than a1 on criteria c1. in this paper, we use 2 1 1a a c to indicate this advantage relation of them. at this situation, c1 is a useful criterion when we compare these two alternatives a1 and a2. two criteria values of alternatives a2 and a4 on criterion c1 are f(a2, c1) and f(a4, c1), i.e., f(a2, c1)=61 and f(a4, c1)=65. obviously, both of them have different criteria values on c1, i.e., f(a2, c1)< f(a4, c1). for benefit criteria, if f(a2, c1)is smaller than f(a4, c1), it means that the alternative a2 locates at a more disadvantage position than a4 on criterion c1 when we compare them. at the same time, the alternative a4 locates at a more advantage position than a2 on criterion c1. in this paper, we use 2 4 1a a c to indicate this disadvantage relation of them. thus, c1 is still a useful criterion when we compare these two alternatives a2 and a4. 32 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… two criteria values of alternatives a1 and a3 on criterion c1 are f(a1, c1) and f(a3, c1), obvious f(a1, c1)=55 and f(a3, c1)=55. both of them have the same criteria value on c1, i.e., f(a1, c1)= f(a3, c1). for benefit or cost criteria, if two alternatives have the same criteria value on a criterion, it means that these two alternatives locate at the same position on this criterion when we compare them. that means the alternative a1 and a3 locate at the same position on criterion c1. according to attribute reduction based on tolerance relation and rough set theory, if two alternatives have the same values on the same criteria, it means that these two criteria values have a tolerance relation for numerical value mcdm problems. when two alternatives have different values on the same criterion, it means that a tolerance relation does not exist. in this paper, we use 1 3 1a a c to indicate this tolerance relation of them. thus, c1 is a useless criterion when we compare these two alternatives a1 and a3, so we can remove it. as the previous discussion, there are three different relations of criteria values for numerical value mcdm problems. these three relations are advantage relation, disadvantage relation and equivalence relation. definition 1. suppose that {a1, a2, … , an} indicates a set of n alternatives, and{c1, c2, … , cm} is a set of m criteria for numerical value mcdm problems. f(ai, cj) and f(ak, cj) indicate the possible outcome of alternatives ai and ak on criterion cj, three relations of the criteria values are defined as follows: ( , ) ( , ) ( , ) ( , ) ( , ) ( , ) i k j i j k j i k j i j k j i k j i j k j a a c f a c f a c a a c f a c f a c a a c f a c f a c        (1) our idea regarding how to find out useful information: as the previous discussion, when we want to compare two alternatives, we do not need to consider those criteria with the same value on the same criteria. that means we do not need to consider those criteria that two alternatives locate at the same position on these criteria. however, we need to consider those criteria that two alternative locate at different positions on these criteria. from the previous discussion, criterion c1 is useless to compare alternatives a1 and a3, but it is needed when we want to know which is better between alternatives a1 and a2. that means the same criteria may play different roles for different alternatives in the decision table. thus, if we want to find out the useful criteria that we need, we need to make a comparison on every pair of alternatives in the decision table. to check every pair of alternatives separately, we need to construct a discernibility matrix [25] and find out useful criteria for all the alternatives in the decision table. the data in the discernibility matrix indicate a set of criteria that must be considered when we want to compare two corresponding alternatives. we think this criteria reduction method reflect and convey the below information.  from the perspective of the alternative: every alternative is chosen as the most desirable alternative(s) in decision making. thus, it is necessary to find out these kinds of desirable criteria that locate in a comparative advantage position as the useful criteria in decision making for every alternative and the larger these criteria’s weights are the better. thus, the alternative could have its weighted combinatorial advantage in decision making.  from the perspective of the competitors: those criteria that locate at a less advantage position could be chosen as the useful criteria in decision making for their opponents and the heavier of those kinds of criteria are, the better. thus, the competitors could have their own weighted combinatorial advantage in decision making and they will have more chance to be selected as the best choice in decision making. from the perspective of the decision makers: they want to make a scientific and reasonable decision within limited time. so they need to remove those useless criteria that locate in the same position when they compare two different alternatives. 2.2 building the discernibility matrix advances in systems science and application(2016) vol.16 no.4 33 as we proposed in section 2.1.2, need to construct three discernibility matrixes as criteria reduction based on tolerance relation, advantage relation and disadvantage relation and rough set theory in order to find out useful criteria from the decision table. the discernibility matrix for numerical values mcdm problems as follows: definition 2. suppose that {a1, a2, … , an} indicates a set of n alternatives, and{c1, c2, … , cm} is a set of m criteria for numerical value mcdm problems. f(ai, cj) and f(ak, cj) indicate the possible outcome of alternatives ai and ak on criterion cj, m is a discernibility matrix for numerical value mcdm problems based on tolerance relation defined as follows: 1 1 11 1 1 n n n n nn a a a m m m a m m           , where { : ( , ) ( , )}j i j k j ik c f a c f a c m else      c (2) and ikm denotes those criteria of two alternatives ai and ak have the different criteria values on the same criteria in set c. in other words, ikm is a set of criteria that contains all the criteria of two alternatives ai and ak have different criteria values in set c. obviously, if there is i k ja a c , there will be k i ja a c when ( , ) ( , )i j k jf a c f a c . so there is mik=mki in the discernibility matrix. thus, the discernibility matrix is a symmetric matrix. definition 3. suppose that {a1, a2, … , an} indicates a set of n alternatives for numerical value or interval number mcdm problems, {c1, c2, … , cm} is a set of m criteria. for the numerical value mcdm problems, m is a discernibility matrix for numerical value mcdm problems based on advantage relation defined as follows: 1 1 11 1 1 n n n n nn a a a m m m a m m            , where, { :j i k j ik c a a c m else     c (3) and ikm is a set of criteria which contains those criteria that the alternative ai locates at a more advantage position than ak on those criteria. definition 4. suppose that {a1, a2, … , an} indicates a set of n alternatives for numerical value or interval number mcdm problems, {c1, c2, … , cm} is a set of m criteria. for the numerical value mcdm problems, m is a discernibility matrix for numerical value mcdm problems based on disadvantage relation defined as follows: 1 1 11 1 1 n n n n nn a a a m m m a m m            , where, { :j i k j ik c a a c m else     c (4) and ikm is a set of criteria which contains those criteria that the alternative ai locates at a more disadvantage position than ak on those criteria. there is tm m between these two discernibility matrixes. so we will get the same useful criteria by using the discernibility matrix based on advantage relation, the discernibility matrix based on disadvantage relation or both. thus, we will just use the discernibility matrixes based on advantage and tolerance relations to find out the useful criteria avoid the useless criteria in this paper. we use the relative discernibility function of discernibility matrix by using boolean reasoning techniques [6-7, 25-26]. we can get the useful criteria for all the alternatives. 3 the method for a large number of criteria mcdm problems 34 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… according to the traditional methods of mcdm problems, we need several procedures, such as unifying criteria and obtaining criteria weights and information fusion as well as ranking and selecting the most desirable alternative(s)[18]. in this paper, we apply the method of criteria reduction based on tolerance and advantage relations and rough set theory to a large number criteria mcdm problems. so, we will change the traditional procedure of mcdm problems. 3.1 resolution procedure for the a large number of criteria mcdm problems to solve the above problems, a new resolution procedure is proposed, as shown in fig. 1. a brief description of the resolution procedure is given below. fig.1 the resolution procedure for a large number of criteria mcdm problems first, by using eq. (2)-(4), we construct a discernibility matrix to find out the useful criteria for numerical value mcdm problems. second, by using eq. (5) and (6), we find a way to obtain criteria weights for useful criteria and all criteria in the decision table, respectively. third, by using eq. (7)-(11), we find a new algorithm to rank alternatives of numerical value of mcdm problems. finally, we validate this new resolution procedure a feasibility and validity method for mcdm problems. all equations (eq. (5)-(11)) and details (sections 4.1 and 4.2) will be provided in the next sections. 3.2 obtaining criteria weights of the useful criteria the challenge of obtaining criteria weights: there are several methods to obtain the criteria weights. according to the traditional opinion, the most popular one in the existing methods is to obtain the criteria weights based on the deviation of criteria values [18]. but we think this method has two places that deserve further consideration.  first, for some mcdm problems, some criteria values change lager than others but these criteria only have a little influence on the result, while some tiny changes of few criteria would lead to different consequences.  second, building the discernibility matrix and finding out useful criteria depend on whether the criteria values are the same or not, but do not depend on whether the deviation of criteria values larger or smaller. large decision table find out useful criteria rank and select alternatives validate the method using eq. (2)-(4) using eq. (5)-(6) using eq. (7)-(10) sections 4.1 and 4.2 information aggregation using eq. (11) obtain criteria weights for the useful criteria advances in systems science and application(2016) vol.16 no.4 35 how to find a scientific and reasonable method and obtain the criteria weights is becoming a very important research area in mcmd fields. based on the previous representing, we propose a new way to obtain the criteria weights of mcmd problems. our idea about how to obtain criteria weights: in the discernibility matrix, ( , 1,2, , )ikm i k n indicates a useful set which contains those criteria that two alternatives ai and ak have different criteria values. for the dms, ( , 1,2, , )ikm i k n contains all the criteria that we must compare if we want to know which is the better between the alternatives ai and ak in table 2. the times of the criteria appear in the discernibility matrix mean how many times we need to consider it. the times of criteria appear in the advantage discernibility matrix means how many alternatives locate at a more advantage position than others on these criteria. every alternative wants to become the best one in decision making. thus, these alternatives want those criteria with big weights in advantage matrix. at the same time, the times of these criteria appear in the disadvantage discernibility matrix means how many alternatives locate at a disadvantage position on these criteria. every alternative hopes its competitors having big weights in disadvantage matrix. thus, the criteria weights should have a proportional relation with the times it appears in the discernibility matrix. definition 5. suppose that 1 2{ , , , }na a a is a set of n alternatives of mcdm problems, 1 2 '{ , , , }mc c c indicates a set of m’ useful criteria,and 1 2 '{ , , , }m   is a set corresponding to all useful criteria weights. ( 1,2, ')j j m  for useful criteria are listed as follows: ' 1 | | | | j j m j j c c     (5) where, | |jc indicates the times of criteria cj appearing in the discernibility matrix, ' 1 | | m j j c   indicates the total times of all useful criteria appearing in the discernibility matrix, and 'm indicates how many useful criteria in table 2. definition 6. suppose that 1 2{ , , , }na a a is a set of n alternatives of mcdm problems, 1 2{ , , , }mc c c indicates a set of m criteria o,and 1 2{ , , , }m   is a set corresponding to all criteria weights. ( 1,2, , )j j m  for all criteria are listed as follows: 1 | | | | j j m j j c c     (6) where, | |jc indicates the times of criteria cj appearing in the discernibility matrix, 1 | | m j j c   indicates the total times of all criteria appearing in the discernibility matrix, and m indicates how many criteria in the decision table. as previous representing in section 2.2.2, we also use the same reasons to explain through eq. (5) and (6) to obtain the criteria weights is reasonable and scientific. as the previous discussion of the risk preference assumptions, we get different advantage orders for the same two interval numbers under a special situation. different dms will get different discernibility matrixes for the interval number mcdm problems. thus, we will get different criteria weights for different types of dms in the same decision table. 3.3 ranking and selecting the most desirable alternative(s) for numerical value mcdm problems the challenge of ranking alternatives: there are several existing techniques to rank alternatives in mcdm, such as gower plots and decision balls method [21], theseus method [20], information fusion [18], weight restrictions (reza farzipoor saen 2009) and rational research method [27] etc. we must make all criteria the same meaning if we want to compare them. but it would produce some errors in this step. furthermore, uniting criteria and comparing the 36 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… weighted combinatorial expectation of alternatives still has another problem. we can not filter off the absolute disparity through unifying criteria, so different criteria are incomparable even if we unify them into the same meaning. our idea of how to rank and select alternatives: we think the criteria only on the same criteria can be compared. through criteria reduction and obtaining criteria weights we don’t need to consider the deviation between two criteria values, because we only consider whether they are the same or not and find out the useful criteria as well as obtain the criteria weights. for the dms, they just need to know the criteria that locate at an advantage or a disadvantage in decision making. they do not need to consider these useless criteria. in this paper, we propose through comparing the weighted combinatorial advantage values (wcav) of alternatives to rank alternatives and select the most desirable alternative(s) [22]. for benefit type criteria, criteria values of two alternatives ai and ak on criterion cj are f(ai, cj) and f(ak, cj). if f(ai, cj)is larger than f(ak, cj) it means that the alternative ai locates at a more advantage position than a2 on criterion cj. in other words, alternative ai is better than ak on criteria cj. undoubtedly, if alternative ai is better than ak on criteria cj, it also means that alternative ak locates at a more disadvantage position than ai on criterion cj. in this paper, we use i k ja a c to indicate the relation. if there is i k ja a c , it means that alternative ak locates at a more advantage position than ai on criterion cj, respectively, there has f(ai, cj)< f(ak, cj). two criteria values of two alternatives on the same criterion. if two alternatives have the same criteria values, it means that these locate at the same position on this criterion in decision making. we use i k ja a cav to express an advantage value (av) between decision alternative ai and ak on criterion cj. so we get the advantage value as follows: 1 0 1 i k j i k j a a c i k j i k j a a c av a a c a a c       (7) 1 2a awav represents a weighted advantage value (wav) of criteria between ai and ak for useful criteria as follows: 1 2 '/ 1 / 2 / 'i k i k i k i k ma a a a c a a c a a c mwav av av av         (8) 1 2a awav represents a weighted advantage value (wav) criteria between ai and ak for all criteria as follows: 1 2/ 1 / 2 /i k i k i k i k ma a a a c a a c a a c mwav av av av         (9) comparing every pair of all alternatives, we construct the weighted advantage relation matrix (warm), so we get the wadm for all the alternatives of mcdm problems as follows: 1 1 1 2 1 2 1 2 2 2 1 2 n n n n n n a a a a a a a a a a a a a a a a a a wav wav wav wav wav wav warm wav wav wav               (10) for the wav there are the following characteristics: 1. 0 i k k ia a a awav wav  2. 0 i ka a i kwav a a   3. 0 i ka a i kwav a a  4. i k k ia a a awav wav  kawcav represents a weighted combinatorial advantage value (wcav) of alternatives for alternative ak in decision tables, we can get the wcav as follows: 1 1k k ia a a i k wcav wav n     (11) app:ds:comparability advances in systems science and application(2016) vol.16 no.4 37 the larger the wcav of ak, the better alternative ak. therefore, all the alternatives can be ranked according to the wcav, respectively. thus, the best alternative can be selected. as the previous discussion of the risk preference assumptions, we get different advantage values for the same two intuitionistic fuzzy sets under a special situation. thus, we will get different weighted combinatorial advantage values for different types of dms in the same decision table. 4 applying the method to numerical value mcdm problems in this section, an example for numerical value mcdm problem is used to illustrate the feasibility and validity of the proposed method. in order to demonstrate the method of criteria reduction based on tolerance relation and rough set theory,we proposed an effective tool for mcdm problems. suppose that the dms have seven criteria of 1 2 3 4 5 6 7{ , , , , , , }c c c c c c cc to make a decision, such as quality c1, competitive c2, price c3, design plan c4, delivery time c5, safety index c6, and sale service c7. suppose we invite 100 experts to make their judgement and voting for all the competitors based on seven criteria, there are two kinds of poll results “yes” or “no”. which is the best one in table 2? table 2 the result of “yes” from 100 experts u c1 c2 c3 c4 c5 c6 c7 a1 45 50 75 20 50 40 48 a2 61 62 65 54 45 50 50 a3 45 55 30 54 45 45 70 a4 65 65 30 65 70 45 65 a5 89 85 65 65 65 50 50 as denoted in fig. 1, we need several steps (find useful criteria, obtain criteria weights, rank alternatives, etc.) to select the most desirable alternative(s). we can build two different discernibility matrixes based on tolerance and advantage relations to find out the useful criteria, respectively. first, we will build the discernibility matrix based on tolerance relation and rough set and validate the method in section 4.1. second, we will build the discernibility matrix based on advantage relation and rough set and validate it in section 4.3. as we have denoted in fig. 1 that we need. 4.1 the procedures for numerical value mcdm problems based on tolerance relation step 1 find out the useful criteria. according to criteria of table 2 by using eq. (2), we construct the discernibility matrix based on tolerance relation and find out the useful criteria as follows: 2 3 4 5 6 7 1 2 3 6 7 1 2 4 5 1 2 3 5 7 1 2 3 5 6 7 c c c c c c c c c c c c c c c m c c c c c c c c c c c                     c c c c c as the previous representing, we know discernibility matrix is a symmetric matrix. in this paper we just give the upper triangular matrix of it. if we want to compare alternatives a1 and a2, we must compare those criteria that m12 contains. as we proposed in section 2, c is a criterion set which contains all the criteria in table 2. thus, we need to compare all criteria in table 2 if we want to know which is better between alternatives a1 and a2. 38 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… we use a relative discernibility function of discernibility matrix by using boolean reasoning techniques[27-30]. we get 1 2 4 5{ , , , }c c c c a group of useful criteria. if we want to compare alternatives these four criteria are needed. step 2 obtain criteria weights of the useful criteria. by using eq. (5) and the useful criteria of 1 2 4 5{ , , , }c c c c , we can get the criteria weights as follows: 1 2 4 5 1 5 2 1 , , , . 4 18 9 4        step 3 rank and select the most desirable alternative(s). by using eq. (7), (8) and (10) as well as the useful criteria of 1 2 4 5{ , , , }c c c c , we construct the warm as follows: 0 0.5 0.25 1 1 0.5 0 19 36 1 1 0.75 19 36 0 1 1 1 1 1 0 5 18 1 1 1 5 18 0 warm                        by using eq. (11), we get the wcav for all alternatives as follows: 1 2 3 4 5 0.6875, 0.2431, 0.4444, 0.6806, 0.8194.a a a a awcav wcav wcav wcav wcav        therefore, the ranking order of all the alternatives is 5 4 2 3 1 0.2778 1 0.5278 0.75 a a a a a . thus, the alternative a5 is the best choice in table 2. 4.1.2 validating the method for numerical value mcdm problems based on tolerance relation in section 4.1.1, we got the alternatives order for numerical value mcdm problems of cacc by using the new method based on the rough set and tolerance relation. several algorithms are needed in it, such as finding out useful criteria through criteria reduction and making a decision using useful criteria. in this section, we will validate this method that we proposed an effective and useful tool of mcdm problems. to address this problem, we use all the criteria in table 2 to make a decision. if we get the same ranking order and the most desirable alternative as we have got in section 4.1 by using the useful criteria, it means that our method is correct. using all criteria in table 2 of mcdm problems to make a decision, by using the eq. (6) to obtain the criterion weights as follows: 1 2 3 4 5 6 7 9 61, 10 61, 8 61, 8 61, 9 61, 8 61, 9 61.c c c c c c c             by using eq. (7), (9), (10) and all the criteria,we get the warm as follows: 0 27 61 18 61 45 61 45 61 27 61 0 26 61 29 61 36 61 18 61 26 61 0 36 61 43 61 45 61 29 61 36 61 0 17 61 45 61 36 61 43 61 17 61 0 warm                        by using eq. (11), we get the wcav for all alternatives as follows: 1 2 3 4 5 0.5533, 0.0492, 0.3566, 0.3811, 0.5779.a a a a awcav wcav wcav wcav wcav        therefore, the ranking order of all the alternatives is 5 4 2 3 1 0.2787 0.4754 0.4262 0.2951 a a a a a . obviously, we get the same ranking order and the most desirable choice in two different situations. so, finding out and using the useful criteria to make decision is an effective and useful tool for numerical value mcdm problems, especially, when there are large number of criteria in decision tables. in section 4.1, we validate the new method for criteria reduction based on rough set and tolerance relation. through this method, we can find out the useful criteria and avoid the useless criteria. just use these useful criteria to make decision. thus, we find a useful method for advances in systems science and application(2016) vol.16 no.4 39 criteria reduction based on tolerance relation to make decision when there have large number of criteria of numerical value mcdm problems. 4.2 the procedures for numerical value mcdm problems based on advantage relation step 1 find out the useful criteria. according to criteria of table 2 by using eq.(3), we construct the discernibility matrix based on advantage relation and find out the useful criteria as follows: 3 5 3 5 3 3 1 2 4 6 7 1 2 3 6 3 6 2 4 6 7 7 7 7 1 2 4 5 6 7 1 2 4 5 7 1 2 4 5 5 7 1 2 4 5 6 7 1 2 4 5 1 2 3 4 5 6 1 2 3 6 c c c c c c c c c c c c c c c c c m c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c c                       as the previous representing, we know that the discernibility matrixes based on advantage and disadvantage relations have a relation of transposed, it means that they are transposed matrix each other. in this paper we just use the advantage discernibilty matrix to find out the useful criteria and make a decision. as the discussion in section 2.2, we know that if we want to compare alternatives a1 and a2, alternatives a1 locates at a more advantage position than a2 on criteria c3 and c5 in table 2. at the same time, alternatives a2 locates at a more advantage position than a1 on criteria c1, c2, c4, c6, and c7 in table 2. we use a relative discernibility function of discernibility matrix by using boolean reasoning techniques [24-26, 27-30]. we get 1 2 3 4 5 7{ , , , , , }c c c c c c a group of useful criteria. if we want to compare alternatives these four criteria are needed. step 2 obtain criteria weights of the useful criteria. by using eq. (5) and the useful criteria of 1 2 3 4 5 7{ , , , , , }c c c c c c , we can get the criteria weights as follows: 1 2 3 4 5 7 9 10 8 8 9 9 , , , , , . 53 53 53 53 53 53            step 3 rank and select the most desirable alternative(s). by using (7), (8) and (10) as well as the useful criteria of 1 2 3 4 5 7{ , , , , , }c c c c c c , we construct the warm as follows: 0 19 53 10 53 37 53 37 53 19 53 0 18 53 37 53 36 53 10 53 18 53 0 27 53 35 53 37 53 37 53 27 53 0 9 53 37 53 36 53 35 53 9 53 0 warm                        by using eq. (11), we get the wcav for all alternatives as follows: 1 2 3 4 5 0.4858, 0.1698, 0.3302, 0.2480, 0.5519.a a a a awcav wcav wcav wcav wcav        therefore, the ranking order of all the alternatives is 5 4 2 3 1 0.1698 0.6981 0.3396 0.1887 a a a a a . thus, the alternative a5 is the best choice. 4.2.2 validating the model for numerical value mcdm problems in this section, we will validate this method based on the rough set and advantage relation an effective and useful tool of mcdm problems. to address this problem, we still use all the criteria in table 2 to make a decision. if we get the same ranking order and the most desirable alternative as we have got in section 4.1 by using the useful criteria, it means that our model is effectively. using all criteria in table 2 of mcdm problems to make a decision, by using the eq. (6) to obtain the criterion weights as follows: 1 2 3 4 5 6 7 9 10 8 8 9 8 9 , , , , , . 61 61 61 61 61 61 61              40 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… by using eq. (7), (9), (10) and all criteria,we get the warm as follows: 0 27 61 18 61 45 61 45 61 27 61 0 26 61 29 61 36 61 18 61 26 61 0 27 61 43 61 45 61 29 61 27 61 0 17 61 45 61 36 61 43 61 17 61 0 warm                        by using eq. (11), we get the wcav for all alternatives as follows: 1 2 3 4 5 0.5533, 0.0492, 0.3197, 0.3343, 0.5779.a a a a awcav wcav wcav wcav wcav        therefore, the ranking order of all the alternatives is 5 4 2 3 1 0.2787 0.4754 0.4262 0.2951 a a a a a . obviously, we get the same ranking order and the most desirable choice in two different situations. so, finding out and using the useful criteria to make a decision is an effective and useful tool for numerical value mcdm problems, especially, when there are large number of criteria in decision tables. thus, we find a useful model to make a decision when dms face large number of criteria of numerical value mcdm problems. if we can get the same ranking order, it means that this model is a useful tool. in fact, if we get the same the most desirable alternative(s), we can say this model is still a useful tool for mcdm problems. because the decision makers are always concerned with finding the best alternative(s), so getting the most desirable alternative(s) seems to be more important than the other ranks. in this section, we use two different methods to address the “large decision table” (e.g. a large number of criteria) challenge in multiple criteria decision making. this new method involves criteria reduction based on rough set and criteria value relation (tolerance and advantage relations). we also validate the new method. thus, we can say that we find an effective method to address the problem “there may be not one of having insufficient information, but rather one of having too much or an unknown quality of information how we filter information for mcdm in an internet or mobile environment.” 4.3 validating the model using maximizing deviation method in this paragraph, we will compare the results from our method with the one based on the maximum deviation method to obtain the criteria weights. by using the maximizing deviation method for interval numbers proposed in [18], we need to take all the criteria into consideration in table 2. if we use all the criteria in table 4 to make a decision, the ranking order of all the alternatives is also 5 4 2 3 1a a a a a . alternative a5 is the best choice. we get the same best choice and ranking order of all alternatives. 5 conclusion this paper mainly focuses upon how to make a wise and reasonable decision within limited time when dms face a large number of criteria in mcdm problems. in this paper, our solution mainly focuses on four aspects. first, according to numerical value mcdm problems, we proposed a method of criteria reduction based on tolerance relation and then build a discernibility matrix to find out useful criteria. just use useful criteria to make a decision. second, through the idea of building the discernibility matrix, we proposed a new method to obtain criteria weights. third, a new method based on weighted combinatorial advantage criteria value and criteria reduction, we give the model of advantage matrix of criteria. finally, we compare the ranking result by using useful criteria and all criteria in the table to make a decision. we validated the method of finding out the useful criteria through criteria reduction a useful and advances in systems science and application(2016) vol.16 no.4 41 scientific method for mcdm problems. our work is an underline research field of mcdm problems with large number of criteria. in future work, we will study the criteria reduction algorithms to find the useful criteria when the criteria are dependent. acknowledgment this research is supported by the national natural science foundation of china (no. 71671092; 71301075), national natural science foundation of jiangsu province, china (no. bk20130770), international postdoctoral exchange fellowship program of china (no. 20140072), postdoctoral science foundation funded project of jiangsu province, china (no. 1501040a). we thank the department editor and anonymous reviewers for their helpful comments. references [1] j. wallenius, j. s. dyer, p. c. fishburn, r. e. steuer, s. zionts and k. deb(2008), "multiple criteria decision making, multicriteria utility theory: recent accomplishments and what lies ahead", management science, vol.54, no. 7, pp. 1336-1349. [2] e. m. feit, m. a. beltramo, f. m. feinberg and reality check(2010), "combining choice experiments with market criteria to estimate the importance of product criteria", management science, vol.56, no. 2, pp. 785-800. [3] p. ghemawat and d. levinthal(2008), "choice interactions and business strategy", management science, vol. 54, no. 9, pp. 1638-1651. [4] v. v. podinovski (2010), "set choice problems with incomplete information about the preferences of the decision maker", european journal of operational research, vol. 207, no. 1, pp. 371-379. [5] w. ziarko(1993), "variable precision rough set method", journal of computer and system sciences, vol.46, no. 1, pp. 39-59. [6] z. plawlak(1997), "rough set approach to knowledge-based decision support", european journal of operational research, vol. 99, no. 1, pp. 48-57. [7] z. plawlak(1982), "rough sets", international journal of computer and information sciences, vol. 11, no. 5, pp. 341-356. [8] w. c. lee and c. c. wu(2009), "some single-machine and m-machine flowshop scheduling problems with learning considerations", information sciences, vol. 179, no. 22, pp. 38853892. [9] z. pawlak and a. skowron(2007), "rough sets: some extensions", information sciences, vol. 177, no. 1, pp. 28-40. [10] d.q. miao, y. zhao, y.y. yao, h.x. li and f.f. xu(2009), "relative reducts in consistent and inconsistent decision tables of the pawlak rough set method", information sciences, vol. 179, no. 24, pp. 4140-4150. [11] t. l. saaty and l. vargas (1987), "uncertain and rank order in the analytic hierarchy process", european journal of operational research, vol. 32, no. 1, pp. 107-117. [12] r. f. saen(2009), "a decision method for ranking suppliers in the presence cardinal and ordinal criteria, weight restrictions, and nondiscretionary factors", annals of operations, vol. 172, no. 1, pp. 177-192. [13] g. r. jahanshahloo, l. f hosseinzadeh and m izadikha(2006), "an algorithmic method to extend topsis for decision-making problems with interval criteria", applied mathematics and computation, vol. 175, no. 2, pp. 1375-1384. [14] j. j. zhang, d. s. wu and d. l. olson(2005), "the method of grey related analysis to multiple criteria decision making problems with intuitionistic fuzzy set", mathematical and compute modeling , vol. 42, no. 9-10, pp. 991-998. [15] j. ma, z. p. fan and l. h. huang(1999), "a alternative and objective integrated approach to 42 zhaobin li, jian liu, zhuo zhang, et al.: a new decision method for multi-criteria decision making with numerical values… obtain criteria weights", european journal of operational research, vol. 112, no. 2, pp. 397-404. [16] y. m. wang and k. s. chin(2011), "the use of owa operator weights for cross-efficiency aggregation", omega, vol. 39, no. 5, pp. 493-503. [17] c. fu and s. l. yang(2011), "an attribute weight based feedback method for multiple attributive group decision analysis problems with group consensus requirements in evidential reasoning context", european journal of operational research, vol. 212, no. 1, pp. 179-189. [18] z. s. xu(2004), uncertain multiple criteria decision making methods and applications, tsinghua university press, beijing, vol. 11, no. 5, pp. 341-356. [19] y. h. qian, j. y. liang and c. y. dang(2008), "interval ordered information systems", computers and mathematics with applications, vol. 58, no. 8, pp. 1994-2009. [20] e. fernazdez and j. navarro(2011), "a new approach to multi-criteria sorting based on fuzzy outranking relations: the theseus method", european journal of operational research, vol. 213, no. 2, pp. 405-413. [21] l. c. ma and h.l. li(2011), "using gower plots and decision balls to rank alternatives involving inconsistent preferences", decision support systems, vol. 51, no. 3, pp. 712-719. [22] liu j., liu p., liu s. f., zhou x. z. and zhang t(2015). "a study of decision process in mcdm problems with large numbers of criteria", international transactions in operational research, vol. 22, no. 2, pp. 237-264. [23] e.a. abo-tabl(2011), " a comparison of two kinds of definitions of rough approximations based on a similarity relation", information sciences, vol. 181, no. 12, pp. 2587-2596. [24] j. b. dos santos, c. a. heuser, v. p. moreira and l.k. wives(2011), "automatic threshold estimation for data matching applications", information sciences,vol.181, no. 13, pp. 26852699. [25] y.y. guan and h.k. wang(2006), "set-valued information systems", information sciences, vol. 176, no. 17, pp. 2507-2525. [26] s. greco, b. matarazzo and r. slowinski(2001), "rough sets theory for multi-criteria decision analysis", european journal of operational research, vol. 129, no. 1, pp. 1-47. [27] w. wang, b. payam and b. andrzej (2011), "rational research method for ranking semantic entities", information sciences, vol. 181, no. 13, pp. 2823-2840. [28] a. skowron and c. rauser(1992), the discernibility matrices and functions in information system, in: r. slowinski (eds.), intelligent decision support: handbook of application and advances of rough sets theory, kluwer academic publisher, dprdrecht, pp. 331-362. [29] a. skowron(1995), "extracting laws from decision tables: a rough set", computational intelligence, vol. 11, no. 2, pp. 371-388. [30] r. w. swiniarskia and a. skowronb(2006), "rough set methods in feature selection and recognition", pattern recognition letters, vol. 24, no. 6, pp. 833-849. corresponding author jian liu can be contacted at: jianlau@njust.edu.cn. http://www.sciencedirect.com/science/journal/00200255 http://www.sciencedirect.com/science/article/pii/s0167865502001964 http://www.sciencedirect.com/science/article/pii/s0167865502001964 http://www.sciencedirect.com/science/article/pii/s0167865502001964 http://www.sciencedirect.com/science/article/pii/s0167865502001964 http://www.sciencedirect.com/science/journal/01678655 microsoft word 4 jiang jianhua, sheng buyun, yang mingzhong--research on owa based multi-source heterogeneous data fusion.do 232-239 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc. research on owa based multi-source heterogeneous data fusion jianhua jiang, buyun sheng and mingzhong yang school of mechanical and electronic engineering, wuhan university of technology, wuhan 430070, china abstract for the problem of multi-source heterogeneous data fusion, an architecture model of multi-source heterogeneous data fusion was designed. in the process of data fusion, triangular fuzzy number (tfn) was used to give uniform express of multi-data description in quantity, the ordered weight average (owa) was used to deal with the preference of decision-maker and design the algorithm of data fusion. at last, the feasibility was validated by an example. keywords data fusion; triangular fuzzy number; ordered weight average 1. introduction data fusion is a process of multi-source data cooperating as to reduce data redundancy and capture comprehensive information from data. it has become a research focus in data processing, object identification, situation assessment and intelligent decision-making domains. in general view, the data for data fusion come from multi-sensors[1], the data type is almost numerical and the methods are mainly from statistics and artificial intelligent, and some literatures have researched the multi-source data fusion[2,3,4]. in fact, except for numerical value, there are many other expression ways of data such as language and symbols, and the multi-ways of data expression will led to the ambiguity, difference and heterogeneity of data. on the other hand, in order to make an ultimate decision, a decision-maker must integrate a wide range of heterogeneous data and information. therefore, in allusion to the features of heterogeneous data, this paper set focus on the method of multi-source heterogeneous data fusion and its application. 2. multi-source heterogeneous data fusion methods and its architecture 2.1 multi-source heterogeneous data fusion methods the methods for data fusion in decision level mainly include weight average[5], d-s evidence theory[6] and voting[7]. (1) weight average method the method uses formula ∑ iji tw to compute decision support value, where iw denotes weight of data source i and ijt denotes support value of data source i to decision j . it evaluates every decision according the support value, it is easy to operate and gives full consideration to the importance of data sources. but it is difficult to eliminate the impact of subjective factors in determining the weights. (2) d-s evidence theory it defines the space including all possible results of the object to be identified as a framework set d , its subset marked d2 and for da ⊆∀ : ]1,0[2: →dm where 0)( =φm , 1)(2 =∑ ⊆ da am , φ is empty set, and m is called basic probability assignment function (bpaf) which actually refers to assign the trust value to subsets of d on advances in systems science and applications (2011), vol.11, no.3-4 233 issn 1078-6236 international institute for general systems studies, inc basis of the available evidence. in practice, many bpafs should be combined for different evidences of a problem may led to different im , it can be realized by the following formal: )()( 1 i aa i amkam i ∑∏×= =∩ − ( ni ≤≤1 ) where )( i a i amk i ∑∏= ≠∩ φ and dai ⊆ d-s evidence theory is based on bpaf and able to compute uncertainty caused by unknown factors. but it requests all the items in d mutually exclusive and is difficult to calculate when there are too many bpafs. (3) voting it treats every data source as a voter and chooses a decision by comparing the votes obtained, the votes is defined as follow: ))(()( iji asupfasup = where ia denotes decision i ( ni ≤≤1 ), )( iasup denotes the total votes it obtained, )( ij asup denotes the support value of data source j to decision i and its value is 0 or 1, and f may be defined as sum. it is difficulty to find bpaf for d-s evidence theory and distinguish two decisions have same votes, this paper uses owa to fusion multi-source data in full considering the preference of decision-makers. 2.2 multi-source heterogeneous data fusion architecture literature [2] proposes an architecture for multi-source data fusion as shown in figure 1, it takes into account user requirements and source credibility, and uses proximity knowledge base, knowledge of reasonableness and voting method to resolve data conflicts. figure 1 schematic of data fusion under the guidance of the above model, an architecture model of multi-source heterogeneous data fusion in decision level was designed as shown in figure 2. the fusion engine is made up of four parts of data warehouse, decision support value computing, owa operator weight vector computing and data conversion and sorting. figure 2 multi-source data fusion architecture (1) data warehouse is used to integrate multi-source data and eliminate data heterogeneity and differences by operation of data selection, feature extraction and statistical analysis. 234 jiang:research on owa based multi-source heterogeneous data fusion (2) decision support computing module obtains related data from data warehouse according to decision-making attributes and computes the support value ijs of data source i to decision j . (3) owa operator weight vector computing module computes the weight of data source according to the fuzzy semantic parameters provided by decision-maker, and these parameters reflect the preferences of decision-makers to the data source. (4) data conversion and sorting calculates ' ijs in combining with data source credibility (or importance degree) and owa weight, and sorts ' ijs according its value. at last, the final decision value can be calculated by ' ijs and iw . 3. algorithm of data fusion 3.1. data type and its features data can be expressed in quantity and quality, and numbers for quantity while language variables for quality[8]. according to difference of data description, data type can be divided into quantitative and qualitative two types, this paper focus on 4 types of data description, as shown in table 1. table 1 data type type data description notes random variable random variable submits to certain distribution quantitative two value type data value either 1 or 0, or true or false degree type the degree general described by 7 or 9 standard qualitative terminology based description method depends on vocabulary space in case of large samples, random variable submits to normal distribution marked ),(~ 2σμx , where μ denotes expectation and σ denotes standard variance, and 9974.0)33( =+<<− σμσμ xp . two-value type data is used to give answer for affirming or denying, if agreed the value is 1, otherwise 0, it can also be described by true or false. degree type data is generally described by degree adverbs such as good, a little good and so on, the grading standard widely used is 7 or 9. terminology based data is described by terms from a pre-defined vocabulary space and the space depends on specific circumstance. 3.2 computing of tfn based support value in allusion to the ambiguity in data description, tfn is used to compute the decision support value. (1) transform method for random data supposed that: σ30 −= ux σ6 0' xxx − = if support value increases with the increased of random variable x and ]3,3[ σμσμ +− is divided into n intervals, support value can be defined as: ⎪ ⎪ ⎩ ⎪⎪ ⎨ ⎧ +> + + ≤<+ + −≤ = σ σσ σμ 3,)1,1,1( )1(66,)1,,( 3,)0,0,0( )( 00 ' ux x n ixx n i n ix n i x xs (1) advances in systems science and applications (2011), vol.11, no.3-4 235 issn 1078-6236 international institute for general systems studies, inc if support value decreases with the increased of random variable x , support value defined as: )()1,1,1(' )( xss x −= (2) transform method for two value type data supposed that two-value type data described by set {1, 0}, the number of answer 1 and 0 is n and m respectively. if support value is defined by number of value 1, support value can be defined as: )/,/,/()( mnnmnnmnnxs +++= (2) (3) transform method for degree type data if data described by the 7 grading standard, the support value can be defined according to table 2. table 2 fuzzy quantification of degree type data (4) transform method for terminology based data if the vocabulary space w includes n terminologies, and they are ordered by their contribution value for decision from low to high as },...,,{ 110 −= nwwww , the support value defined as: ))1/(),1/(),1/(()( −−−= nininiws i (3) 3.3 computing of owa operator[9] weight vector supposed that rrf n →: and a n dimension vector ),...,,( 21 nwwww = related to f makes: ∑ == n i iin bwaaaf 121 ),...,,( (4) where ]1,0[∈iw , ni ≤≤1 and 11 =∑ = n i iw . ib is the ith largest element of ia . then f is called n dimension owa operator and ),...,,( 21 nwwww = can be represented by a function as: )/)1(()/( nifnifwi −−= (5) where ni ,...,2,1= , f is named fuzzy semantic operator (fso) and defined as: ⎪ ⎩ ⎪ ⎨ ⎧ < ≤≤−− < = xb bxaabax ax xf ,1 ,)/()( ,0 )( (6) where ]1,0[,, ∈bax . in addition, owa also defines a measure operator for reflecting preference of decision-makers named )(worness and it is defined as: inverse proportion direct proportion )(xs very high very low (0.00,0.00,0.17) high low (0.00,0.17,0.34) a little high a little low (0.17,0.34,0.50) common common (0.34,0.50,0.67) a little low a little high (0.50,0.67,0.84) low high (0.67,0.84,1.00) very low very high (0.84,1.00,1.00) 236 jiang:research on owa based multi-source heterogeneous data fusion ∑ = −== n i iwin n wornessc 1 )(1)( (7) 3.4 algorithm of data fusion supposing there are n decisions ),...,,( 21 naaaa = , m data sources ),...,,( 21 mssss = and ip is the credibility or importance of is , the algorithm of data fusion is described as: step1: compute support value of data sources to decisions access data from data warehouse and transform it to support value of every decision according to the method of section 3.1. the support value marked as: ),,( ijijijij cbas = where ijs is support value of data source i to decision j , ),,( jiijij cba is tfn expression of ijs and 10 ≤≤≤≤ ijijij cba . step2: compute owa operator weight vector considering the preference of decision-makers, choose fuzzy semantic parameters for formula 6. in most cases, fuzzy semantic parameters is defined as “the majority”, “at least half” and “as much as possible”, their parameters are )8.0,3.0( , )5.0,0( and )1,5.0( respectively. according to the parameters, the fso )(xf can be determined. with fso, formula 5 and 7, owa weight vector ),...,,( 21 nwwww = and )(worness can be calculated. step3: transform ijs according to ip and ijs in order to use owa weight vector, ijs must be transformed and sorted, the transform method as follow: supposed that: ijiij sps ×=min_ ijiijiij spsps ×−+=max_ (8) ijin i i averageij sp p ns ∑ = = 1 _ define: if 5.0≤c 0)( =ch , cm c 2)( = , cl c 21)( −= if 5.0≥c 12)( −= ch c , cm c 22)( −= , 0)( =cl ijs can be transformed into ' ijs by following formula: min_)(_)(max_)( ' ijcaverageijcijcij slsmshs ++= (9) step 4: compute final support value for every decision the final support value of decisions can be computed by the following formula: ∑ == m i ijij bws 1 nj ,...,2,1= (10) where ijb denotes the ith largesse element of ),...,,( '' 2 ' 1 njjj sss . step5: make a decision according to the final support value of decisions. 4. an example this paper takes steam turbine product development decision-making as example. supposing there are 5 kinds of product named 1a , 2a , 3a , 4a and 5a , the data can be collected include market demand, product feedback, product parameters, historical data, failure statistics, advances in systems science and applications (2011), vol.11, no.3-4 237 issn 1078-6236 international institute for general systems studies, inc expert advice and so on. based on these data sources, products can be compared from 6 aspects, they are assessment of market demand 1a , average failures per year 2a ( 8.0,5.3 == σμ ), longest average no failure time 3a (month as unit and 53.2,28.12 == σμ ), economy evaluation 4a , customer evaluation 5a and expert advice 6a , and the preliminary results are shown in table 3. table 3 support value of product and data source credibility (1) in table 3, 1a and 4a are degree-type data and transformed by table 2, 2a and 3a are random data and transformed by formula 1 (supposing n=15), 5a is two-value type data(the value in table 3 is the ratio of positive evaluation) and transformed by formula 2, 6a is terminology based data and transformed by formula 3. the transformed results are shown in table 4. (2) if “the majority” is selected as fso, the parameters a and b in formula 6 are 0.3 and 0.8 respectively. according to formula 5 and 6, the owa weight vector can be calculated as: )0,27.0,33.0,33.0,067.0,0(=w put every iw into formula 7, the value of c can be calculated and 37.0=c . (3) according to formula 8 and 9, the data in table in table 4 can be transformed as shown in table 5. (4) for every column data in table 5, sorted them from high to low according the second value and computed final support value by formula 10, the result is shown in table 6. (5) from table 6, product 3a gets the highest support value, so it should be developed full table 4 disposed attribute values ia ip a1 a2 a3 a4 a5 1a 0.5 (0.67,0.84,1.00) (0.34,0.5,0.67) (0.84,1.00,1.00) (0.17,0.34,0.50) (0.00,0.17,0.34) 2a 0.9 (0.13,0.14,0.20) (0.00,0.00,0.00) (0.53,0.55,0.60) (0.80,0.85,0.86) (0.02,0.23,0.27) 3a 0.9 (0.20,0.24,0.26) (0.53,0.56,0.6) (0.66,0.68,0.73) (0.13,0.16,0.20) (0.80,0.83,0.86) 4a 0.3 (0.34,0.50,0.67) (0.84,1.00,1.00) (0.17,0.34,0.50) (0.50,0.67,0.84) (0.67,0.84,1.00) 5a 0.6 (0.65,0.65,0.65) (0.73,0.73,0.73) (0.48,0.48,0.48) (0.54,0.54,0.54) (0.67,0.67,0.67) 6a 0.4 (0.33,0.33,0.33) (1.00,1.00,1.00) (0.66,0.66,0.66) (0.00,0.00,0.00) (0.33,0.33,0.33) ia ip a1 a2 a3 a4 a5 1a 0.5 large common very large a little small very small 2a 0.9 5.25 8.64 3.28 1.81 4.76 3a 0.9 8.36 13.28 15.14 7.08 17.31 4a 0.3 commo n very good a little bad a little good good 5a 0.6 0.65 0.73 0.48 0.54 0.67 6a 0.4 improve fully recommen d as usual out of market improve 238 jiang:research on owa based multi-source heterogeneous data fusion table 5 results of data transformed ia a1 a2 a3 a4 a5 1a (0.50,0.62,0.75) (0.25,0.37,0.50) (0.63,0.75,0.75) (0.13,0.25,0.37) (0.00,0.13,0.25) 2a (0.18,0.19,0.27) (0.00,0.00,0.00) (0.71,0.74,0.81) (1.07,1.14,1.16) (0.27,0.31,0.36) 3a (0.27,0.32,0.35) (0.71,0.75,0.81) (0.89,0.92,0.98) (0.18,0.22,0.27) (1.07.1.12,1.16) 4a (0.15,0.22,0.30) (0.38,0.45,0.45) (0.07,0.15,0.22) (0.22,0.30,0.38) (0.30,0.38,0.45) 5a (0.58,0.58,0.58) (0.65,0.65,0.65) (0.43,0.43,0.43) (0.48,0.48,0.48) (0.60,0.60,0.60) 6a (0.20,0.20,0.20) (0.60,0.60,0.60) (0.39,0.39,0.39) (0.00,0.00,0.00) (0.20,0.20,0.20) table 6 final support values a1 a2 a3 a4 a5 value (0.23,0.27,0.31) (0.43,0.49,0.53) (0.52,0.54,0.56) (0.20,0.27,0.35) (0.28,0.32,0.36) 5. conclusion in this paper, an architecture model for multi-source heterogeneous data fusion was constructed, tfn based data processing for multi-data description in quantity was researched and owa operator based data fusion algorithm was designed, the practical application indicates that the algorithm designed is effective. the study of this paper presents a feasible option for constructing intelligent decision support system and has certain reference value for similar data processing and fusion. acknowledgements this paper is supported by national natural science fund project of china (contract no.50620130441), scientific and technological project of wuhanc city (contact no. 200810321153) and youth science and technology chenguang project of wuhan city (contact no. 200750731289). references [1] d. l. hall, j. llinas. hand book of multi-sensor data fusion. beijing: electronic industry press, 2008. [2] r. r.yager. a framework for multi-source data fusion. journal of information sciences vol.163 (2004) 175~200. [3] zhou haiyin, li donghui, jiang yueping. high precision method for determining the position of aerocraft based on mult-sensors information fusion. journal of hunan university(natural sciences), 34(4) (2007) 37~40. [4] r.r.yager. modeling intelligence information: multi-source fusion. ieee international conference on computational intelligence for homeland security and personal safety,orlando, fl, mar 31-apr 01, 2005. [5] wang guangyun, li weihua, hua wenjian, etal. a method for heterogeneous uncertain information fusion and its application. international conference on signal processing proceedings, vol.3 (2004) 2253~2256. [6] ma linru, yang lin, wang jianxin. research on security information fusion from multiple heterogeneous sensors. journal of system simulation, 20(4) (2008) 981~985. [7] s. kang. two-phase identification algorithm based on fuzzy set and voting for intelligent multi-sensor data fusion, lecture notes in computer science, vol. 4252 (2006) 769~776. [8] f.herrera, l.martinez. an approach for combing linguistic and numerical information based on the 2-tuple fuzzy linguistic representation model in decision-making. international advances in systems science and applications (2011), vol.11, no.3-4 239 issn 1078-6236 international institute for general systems studies, inc journal of uncertainty, fuzziness and knowledge-based systems, 8(5) (2000)539~562. [9] z.zhang, c.zhang. result fusion in multi-expert systems based on owa operator. proceedings of the 23rd computer science conference, (2000) 234~240. advances in systems science and applications (2012) vol.12 no.3 283-296 central limit theorems for the single point catalytic super-brownian motion with immigration xu yang and mei zhang laboratory of mathematics and complex systems, school of mathematical sciences, beijing normal university, beijing 100875, people’s republic of china abstract we establish the central limit theorems for the single point catalytic superbrownian motion with deterministic immigration and single point catalytic superbrownian motion immigration on the schwartz space and a weighted sobolev space. for the catalytic immigration case, the weak convergence depends on the branching rates of both immigration part and non-immigration part. keywords super-brownian motion, catalytic, conditional log-laplace functional, immigration, central limit theorem 1 introduction and main results catalytic super-brownian motion is the superprocess with brownian motion as underlying spatial motion,whose branching occurs only in the presence of some catalysts. dawson and fleischmann [1] considered the case of single point catalytic super-brownian motion, namely, the underlying particles move as independent brownian motions in r. the life time of each particle is exponentially distributed. when it dies, it splits according to critical branching only if they pass 0, which is called the (single) catalyst point. fleischmann and xiong [2], yang and zhang [3] and li and wang [4] proved the large deviation, moderate deviation and central limit theorem for the single point catalytic super-brownian motion, respectively. if we suppose the situation where there are additional particles added, we need to consider the process with immigration. superprocesses with immigration are studied by many authors; see e.g. [5-10]. in the present paper,first we will prove the central limit theorem for the single point catalytic super-brownian motion with deterministic immigration on schwartz space, and then extend the result to a weighted sobolev space. we shall also investigate the processes with catalytic immigration. in this case, the weak convergence depends not only on the branching rate ϱ of non-immigration part, but also on the branching rate ϱ0 of immigration part. 1.1 notations and preliminaries first we introduce some notations. let p ≥ 2, hp(x) = (1 + x2)− p 2 and cp(r) denote the set of all real-valued continuous functions φ on r such that φ(x)/hp(x) has a finite limit as |x| → ∞. equipped with the norm 9φ9p := sup{|φ(x)|/hp(x) : x ∈ r}, cp(r) is a banach space. let c+ p (r) denote all the positive functions 284 xu yang:central limit theorems for the single point catalytic... of cp(r) and mp(r) be the set of all measures µ on r such that ⟨µ, hp⟩ < ∞. suppose that 9µ9p = ⟨µ, hp⟩. let c∞(r) be the set of bounded infinitely differentiable functions on r with bounded derivatives. let s (r) ⊂ c∞(r) denote the schwartz space of rapidly decreasing functions on r and s+(r) be the collection of non-negative elements of s (r). that is, each f ∈ s (r) is infinitely differentiable and for each non-negative integer k and each non-negative integer α we have lim |x|→∞ xk dα dxα f(x) = 0. now we introduce some basic results in the following subsection, which can be found in [11]. we define the hilbertian norms {q0, q1, q2, · · · } on s (r) by qn(f) 2 = n∑ k=0 ∫ r (1 + x2)n(f (k)(x))2dx. the hermite polynomials on r are given by gk(x) = (−1)kex 2 dk dxk e−x2 , k = 0, 1, 2, · · · . based on those we define the hermite functions hk(x) = 1 4 √ π √ 2kk! e−x2/2gk(x), k = 0, 1, 2 · · · . then hk ∈ s (r) and {hk : k ≥ 0} is a complete orthonormal system in l2(r). let ⟨·, ·⟩ denote the inner product of l2(r). for f ∈ s (r) we write f = ∞∑ k=0 ⟨f, hk⟩hk and define ∥f∥2n = ∞∑ k=0 (2k + 1)2n⟨f, hk⟩2 (1) for n = 0,±1,±2, · · · . let hn(r) be the completion of s (r) with respect to ∥ · ∥n. by approximation we can extend ⟨·, ·⟩ to a bilinear form between h−n(r) and hn(r). let ⟨·, ·⟩n denote the inner product of hn(r). for g, f ∈ hn(r) we have ⟨g, f⟩n = ∞∑ k=0 (2k + 1)2n⟨g, hk⟩⟨f, hk⟩ = ⟨πng, f⟩, where πng = ∞∑ k=0 (2k + 1)2n⟨g, hk⟩hk ∈ hn(r). then h−n(r) and hn(r) are dual spaces with the duality ⟨·, ·⟩. advances in systems science and applications (2012) vol.12 no.3 285 lemma 1.1 for every n ≥ 0 there is a constant c(n) > 0 such that qn(f) ≤ c(n)∥f∥n and ∥f∥n ≤ c(n)q2n(f), f ∈ s (r). the sequence of norms defined by (1) induces a topology on the set h∞ := ∞∩ k=0 hn(r), which is compatible with the metric ρ defined by ρ(f, g) = ∞∑ k=0 ∥f − g∥k 2k(1 + ∥f − g∥k) . then (s (r), ρ) (written as s (r) for simplicity) is a nuclear space. this implies s ′(r) = ∞∪ n=0 h−n(r) ⊃ · · · ⊃ h−2(r) ⊃ h−1(r) ⊃ h0(r) ⊃ h1(r) ⊃ h2(r) ⊃ · · · ⊃ ∞∩ n=0 hn(r) = s (r). a subset b of the nuclear space s (r) is said to be bounded if it is bounded in each norm ∥ · ∥n, that is, sup x∈b ∥x∥n < ∞ for each n ≥ 0. for each bounded set b ⊂ s (r) we define the semi-norm pb on s ′(r) by pb(f) = sup{|f(x)| : x ∈ b}, f ∈ s ′(r). we endow s ′(r) with the topology generated by the collection of semi-norms {pb : b ⊂ s (r) is bounded}, which is called the strong topology. then s ′(r) is a nuclear space. 1.2 models for a process x taking its value in mp(r), let pr,ν denote its conditional law given xr = ν. suppose that p is the heat kernel in r with constant ς > 0: 1√ 2πςt exp { − a2 2ςt } , t > 0, a ∈ r · · · . (2) for φ ∈ c+ p (r) fixed, u(t, z; ϱ) denotes the unique non-negative solution to the log-laplace equation u(t, z; ϱ) = ptφ(z)− ϱ ∫ t 0 pt−r(z)u 2(r, 0; ϱ)dr, t ≥ 0, z ∈ r (3) 286 xu yang:central limit theorems for the single point catalytic... and ω(t, z; ϱ, ϱ0) denotes the unique non-negative solution to the log-laplace equation ω(t, z; ϱ, ϱ0) = ∫ t 0 pt−r [u(r, ·; ϱ)] (x)dr − ϱ0 ∫ t 0 pt−r(z)ω 2(r, 0; ϱ, ϱ0)dr, t ≥ 0, z ∈ r · · · . (4) where ϱ and ϱ0 are positive constants and {pt : t ≥ 0} denotes the brownian semigroup corresponding to (2). let a be the generator of {pt : t ≥ 0}. in the following we suppose µ, η ∈ mp(r). we say ξµ,ϱ = {ξµ,ϱt : t ≥ 0} is a single point catalytic super-brownian motion, if ξµ,ϱ0 = µ, and for r ≥ 0 and ν ∈ mp(r), − logpr,ν exp {−⟨ξt, φ⟩} = ⟨ν, u(t− r, ·; ϱ)⟩, 0 ≤ r ≤ t, φ ∈ c+ p (r) · · · . (5) suppose z = {zt : t ≥ 0} is the single point catalytic super-brownian motion with z0 = µ, deterministic immigration controlled by η and log-laplace functional given by − logpr,νe −⟨zt,φ⟩ =⟨ν, u(t− r, ·; ϱ)⟩+ ∫ t r ⟨η, u(t− s, ·; ϱ)⟩ds, 0 ≤ r ≤ t, φ ∈ c+ p (r) · · · . (6) now we suppose that xξ = {xξ t : t ≥ 0} is the single point catalytic superbrownian motion with single point catalytic immigration determined by ξη,ϱ0 . let p0,ν,η denote its conditional law given xξ 0 = ν and ξ0 = η. by theorem 3.2 of [12] we have the log-laplace functional of xξ: − logp0,ν,η exp { −⟨xξ t , φ⟩ } =− logp0,ν,η [ p0,ν,η exp { −⟨xξ t , φ⟩ } ∣∣∣∣ {σ(ξs : 0 ≤ s ≤ t)} ] =− logp0,η exp { −⟨ν, u(t, ·; ϱ)⟩ − ∫ t 0 ⟨ξs, u(t− s, ·; ϱ)⟩ds } =⟨ν, u(t, ·)⟩+ ⟨η, ω(t, ·; ϱ, ϱ0)⟩, µ, η ∈ mp(r), t ≥ 0, φ ∈ c+ p (r) · · · . (7) in the following, we breviate u(t, ·; ϱ) and ω(t, ·; ϱ, ϱ0) by u(t, ·) and ω(t, ·), respectively. we construct the superprocesses z, ξη,ϱ0 and xξ on the same probability space (ω,f ,p). advances in systems science and applications (2012) vol.12 no.3 287 1.3 main results let {z(k) t : t ≥ 0} be the single point catalyst super-brownian motion with deterministic immigration characterized by(3) and (6), with ϱ replaced by 1/k, let w (k) t = k 1 2 ( z (k) t − µpt − ∫ t 0 ηpsds ) . theorem 1.2 as k → ∞, the sequence {w (k) t : t ≥ 0} converges weakly to the gaussian process {wt : t ≥ 0} in c([0,∞),s ′(r)) with w0 = 0 and laplace functional given by p exp { − ⟨wt, f⟩ } =exp {∫ t 0 ⟨µ, pt−r(·)⟩p 2 r f(0)dr + ∫ t 0 ds ∫ s 0 ⟨η, ps−r(·)⟩p 2 r f(0)dr } , where f ∈ s+(r). remark 1.3 if η ≡ 0 in theorem ??, the result can be found in [4]. theorem 1.4 for any n ≥ 3, c([0,∞),s ′(r)) can be replaced by c([0,∞), h−n(r)) in theorem ??. similarly, we can establish the central limit theorem for the single point catalytic super-brownian motion with immigration controlled by another single point catalytic super-brownian motion. the weak convergence depends on ρ and ρ0, which are the branching rates of non-immigration and immigration parts. to specify the effects of ϱ and ϱ0 on the convergence of xξ, we suppose that ϱ = γ1k −1, ϱ0 = γ2k −β(β > 0), where k ∈ n, γ1, γ2 are positive constants. let {ξ(k)t : t ≥ 0} be defined by (3) and (5) with ϱ0 replaced by γ2k −β(β > 0). let {x(k) t : t ≥ 0} be defined accordingly by (7), (3) and(4) with ϱ, ϱ0 replaced by γ1k −1, γ2k −β(β > 0), respectively. define y k t = kα ( x (k) t − µpt − tηpt ) (8) theorem 1.5 as k → ∞, the sequence {y (k) t : t ≥ 0} converges weakly to the gaussian process {yt : t ≥ 0} in c([0,∞),s ′(r)) with y0 = 0 and laplace functional given by p exp { − ⟨yt, f⟩ } = exp {∫ t 0 fs(f)ds } , f ∈ s+(r) (9) 288 xu yang:central limit theorems for the single point catalytic... where fs(f) is determined by the following table: β α fs(f) (0, 1) 1 2β γ2⟨η, pt−s(·)⟩ (spsf(0)) 2 1 1 2 γ1⟨µ, pt−s(·)⟩ (psf(0)) 2 + γ1 ∫ s 0 ⟨η, ps−r(·)⟩ (prf(0)) 2 dr +γ2⟨η, pt−s(·)⟩ (spsf(0)) 2 (1,∞) 1 2 γ1⟨µ, pt−s(·)⟩ (psf(0)) 2 + γ1 ∫ s 0 ⟨η, ps−r(·)⟩ (prf(0)) 2 dr theorem 1.6 for any n ≥ 3, c([0,∞),s ′(r)) can be replaced by c([0,∞), h−n(r)) in theorem ??. since the proofs are similar, we only show the case β = 1 of theorems ??–??. 2 proofs of of theorem ?? and theorem ?? (case β = 1) in the proof of theorem ??, we need the following proposition. proposition 2.1 as k → ∞, the finite dimensional distributions of {y (k) t : t ≥ 0} converge weakly to a s ′(r)-valued gaussian process {yt : t ≥ 0} with y0 = 0 and laplace functional determined by (9). proof. to simplify the notations, we consider the two dimensional distributions of y (k). first we calculate the laplace transform of (⟨y (k) t1 , f⟩, ⟨y (k) t2 , f⟩). by the markov property, (7), (3), (8), [10] and [12], for f ∈ c+ p (r), t2 > t1 > 0, we get the laplace transform of (⟨y (k) t1 , f⟩, ⟨y (k) t2 , f⟩): logp0,µ,η exp { − θ1⟨y (k) t1 , f⟩ − θ2⟨y (k) t2 , f⟩ } = logp0,µ,η [ p0,µ,η ( exp { − ⟨x(k) t1 , θ1k 1 2 f + v(k)(t2 − t1, ·; θ2)⟩ − ∫ t2 t1 ⟨ξ(k)s , v(k)(t2 − s, ·; θ2)⟩ds }∣∣∣∣ {σ(ξ(k)s : s ≤ t2) })] + k 1 2 ( ⟨µpt1 , θ1f⟩+ ⟨µpt2 , θ2f⟩+ t1⟨ηpt1 , θ1f⟩+ t2⟨ηpt2 , θ2f⟩ ) = logp0,η exp { − ⟨µ, u(k)(t1, ·; θ1, θ2)− ∫ t1 0 ⟨ξ(k)s , u(k)(t1 − s, ·; θ1, θ2)⟩ds − ∫ t2 t1 ⟨ξ(k)s , v(k)(t2 − s, ·; θ2)⟩ds } + ⟨µpt1 , k 1 2 θ1f⟩ + ⟨µpt2 , k 1 2 θ2f⟩+ t1⟨ηpt1 , k 1 2 θ1f⟩+ t2⟨ηpt2 , k 1 2 θ2f⟩. advances in systems science and applications (2012) vol.12 no.3 289 using markov property, (3) and (8), we get logp0,ν,η exp { − θ1⟨y (k) t1 , f⟩ − θ2⟨y (k) t2 , f⟩ } = logp0,η exp { − ⟨µ, u(k)(t1, ·; θ1, θ2)− ∫ t1 0 ⟨ξ(k)s , u(k)(t1 − s, ·; θ1, θ2)⟩ds − ⟨ξ(k)t1 , ω(k)(t2 − t1, ·; θ2)⟩ } + ⟨µpt1 , k 1 2 θ1f⟩+ ⟨µpt2 , k 1 2 θ2f⟩+ t1⟨ηpt1 , k 1 2 θ1f⟩+ t2⟨ηpt2 , k 1 2 θ2f⟩ =k−1γ1 ∫ t2−t1 0 ⟨µ, pt2−r(·)⟩ [ v(k)(r, 0; θ2) ]2 dr + k−1γ1 ∫ t1 0 ⟨µ, pt1−r(·)⟩ [ u(k)(r, 0; θ1, θ2) ]2 dr + k−1γ2 ∫ t2−t1 0 ⟨η, pt2−l(·)⟩[ω(k)(l, 0; θ1, θ2)] 2dl + k−1γ2 ∫ t1 0 ⟨η, pt1−r(·)⟩[s(k)(r, 0; θ1, θ2)]2dr + k−1γ1 ∫ t2−t1 0 dl ∫ l 0 ⟨η, pt2−r(·)⟩[v(k)(r, 0; θ2)]2dr + k−1γ1t1 ∫ t2−t1 0 ⟨η, pt2−r(·)⟩[v(k)(r, 0; θ2)]2dr + k−1γ1 ∫ t1 0 dl ∫ l 0 ⟨η, pt1−l(·)⟩[u(k)(r, 0; θ1, θ2)]2dr, where θ1, θ2 ≥ 0, v(k)(·, ·; θ2) is the non-negative solution to v(k)(r, x; θ2) = k 1 2 θ2prf(x)− γ1k −1 ∫ r 0 pr−l(x) [ v(k)(l, 0; θ2) ]2 dl, u(k)(·, ·; θ1, θ2) is the non-negative solution to u(k)(r, x; θ1, θ2) = pr [ k 1 2 θ1f + v(k)(t2 − t1, ·; θ2) ] (x) − γ1k −1 ∫ r 0 pr−l(x) [ u(k)(l, 0; θ1, θ2) ]2 dl, ω(k)(·, ·; θ2) is the non-negative solution to ω(k)(r, x; θ2) = ∫ r 0 pr−l [ v(k)(l, ·; θ2) ] (x)dl − γ2k −1 ∫ r 0 pr−l(x) [ ω(k)(l, 0; θ2) ]2 dl, 290 xu yang:central limit theorems for the single point catalytic... and s(k)(·, ·; θ1, θ2) is the non-negative solution to s(k)(r, x; θ1, θ2) = pr [ ω(k)(t2 − t1, ·; θ2) ] (x) + ∫ r 0 pr−l [ u(k)(l, ·; θ1, θ2) ] (x)dl −γ2k−1 ∫ r 0 pr−l(x) [ s(k)(l, 0; θ1, θ2) ]2 dl. it is easy to show that k−1/2v(k), k−1/2u(k), k−1/2ω(k) and k−1/2s(k) are all convergent as k → ∞. then lim k→∞ logp0,µ,η exp { − θ1⟨y (k) t1 , f⟩ − θ2⟨y (k) t2 , f⟩ } =γ1θ 2 1 ∫ t1 0 ⟨µ, pt1−r(·)⟩ [ prf(0) ]2 dr + γ1θ 2 2 ∫ t2 0 ⟨µ, pt2−r(·)⟩ [ prf(0) ]2 dr + 2γ1θ1θ2 ∫ t1 0 ⟨µ, pt1−r(·)⟩prf(0)pt2−t1+rf(0)dr + γ1θ 2 1 ∫ t1 0 dl ∫ l 0 ⟨η, pt1−r(·)⟩ [ prf(0) ]2 dr + γ1θ 2 2 ∫ t2 0 dl ∫ l 0 ⟨η, pt2−r(·)⟩ [ prf(0) ]2 dr + 2γ1θ1θ2 ∫ t1 0 dl ∫ l 0 ⟨η, pt1−r(·)⟩prf(0)pt2−t1+rf(0)dr + γ2θ 2 1 ∫ t1 0 ⟨η, pt1−r(·)⟩ [ rprf(0) ]2 dr + γ2θ 2 2 ∫ t2 0 ⟨η, pt2−r(·)⟩ [ rprf(0) ]2 dr + 2γ2θ1θ2 ∫ t1 0 dl ∫ l 0 ⟨η, pt1−r(·)⟩r(t2 − t1 + r)prf(0)pt2−t1+rf(0)dr. recalling (9), we can obtain the result by the method of [12] page 110. 2 in the following, for fixed interval i := [0, t ], t > 0, we introduce the banach space ci p (r) of all continuous maps u of i into cp(r) equipped with the norm ∥u∥ip := sup{9u(t)9p : t ∈ i}. the proofs of the following lemmas ??, ?? and ?? are essentially similar to those of [1, lemma 2.5.2], [1, lemma 2.6.2] and [1, lemma 3.2.1], so we only present the results and omit the proofs here: lemma 2.2 there are two positive constants ε1, ε2, such that for |θ| < ε1, there is a unique solution u = uθ ∈ ci p (r) to the following equation u(t, x) = θptf(x)− ∫ t 0 pt−r(x) [u(r, 0)] 2 dr, 0 ≤ t ≤ t, x ∈ r satisfying ∥u∥ip < ε2. advances in systems science and applications (2012) vol.12 no.3 291 set v(t, x) := θptf(x)− u(t, x), 0 ≤ t ≤ t, x ∈ r, |θ| < ε1. let v(n) denote the nth derivative of v with respect to θ, taken at θ = 0. put ∥sf∥t := sup { |ptf(0)| : t ∈ (0, t ] } , f ∈ cp(r), t > 0. lemma 2.3 for each f ∈ cp(r), 0 ≤ t ≤ t, n ≥ 2, v(2)(t, x) = 2 ∫ t 0 pt−r(x) [prf(0)] 2 dr, x ∈ r and there is a constant cn such that 9v(n)(t)9p ≤ n!cn∥sf∥ntα(t)n−1, where α(t) = t 1 2 + t p+1 2 . lemma 2.4 for each k ≥ 1, n ≥ 2, f ∈ cp(r), p0,µ,η [∣∣∣⟨y (k) t , f⟩ ∣∣∣2] = 2 ∫ t 0 ⟨µ, pt−s(·)⟩(s2 + 1) [psf(0)] 2 ds +2 ∫ t 0 ds ∫ s 0 ⟨η, pt−r(·)⟩ [prf(0)] 2 dr, and there exists a constant cn such that∣∣∣p0,µ,η [ ⟨y (k) t , f⟩ ]n∣∣∣ ≤ cnt n 4 ∥sf∥n1 n−1∑ i=1 ( 9 µ 9p + 9 η 9p )i . by the proofs of proposition ?? and [1, lemma 3.2.1], for all t2 > t1 > 0 and φ,ψ ∈ cp(r) we have p0,µ,η [ ⟨y (k) t1 , φ⟩+ ⟨y (k) t2 , ψ⟩ ] = 0 (10) and (−1)np0,µ,η [ ⟨y (k) t1 , φ⟩+ ⟨y (k) t2 , ψ⟩ ]n =⟨kµ, ū(n)k (t1, ·, φ, ψ)⟩+ ⟨kη, s̄(n)k (t1, ·, φ, ψ)⟩+ ∑ 2≤j≤n−2 ( n− 1 j )[ ⟨kµ, ū(n−j) k (t1, ·, φ, ψ)⟩+ ⟨kη, s̄(n−j) k (t1, ·, φ, ψ)⟩ ] (−1)jp0,µ,η [ ⟨y (k) t1 , φ⟩+ ⟨y (k) t2 , ψ⟩ ]j , (11) 292 xu yang:central limit theorems for the single point catalytic... where n ≥ 2, v̄k(·, ·, ψ) is the non-negative solution to v̄k(r, x, ψ) = k− 1 2prψ(x)− ∫ r 0 pr−l(x) [ v̄k(l, 0, ψ) ]2 dl, ūk(·, ·, φ, ψ) is the non-negative solution to ūk(r, x, φ, ψ) = pr [ k− 1 2φ+ v̄k(t2 − t1, ·, ψ) ] (x)− ∫ r 0 pr−l(x) [ ūk(l, 0, φ, ψ) ]2 dl, ω̄k(·, ·, ψ) is the non-negative solution to ω̄k(r, x, ψ) = ∫ r 0 pr−l [ v̄k(l, ·, ψ) ] (x)dl − ∫ r 0 pr−l(x) [ ω̄k(l, 0, ψ) ]2 dl, and s̄k(·, ·, φ, ψ) is the non-negative solution to s̄k(r, x, φ, ψ) = pr [ ω̄k(t2 − t1, ·, ψ) ] (x) + ∫ r 0 pr−l [ ūk(l, ·, φ, ψ) ] (x)dl − ∫ r 0 pr−l(x) [ s̄k(l, 0, φ, ψ) ]2 dl. then by (10), (11) and calculations, we get the following estimate. lemma 2.5 there exists a positive constant c0 such that p0,µ,η [ ⟨y (k) t , phf⟩ − ⟨y (k) t+h, f⟩ ]6 ≤ c0h 3 2 ∥sf∥61, 0 ≤ t < t+h ≤ 1, k ≥ 1, f ∈ s (r). similarly to [1, lemma 3.2.2], we have the sixth moment of the increments of the process {y (k) t : t ≥ 0}: lemma 2.6 there is a positive constant c0 such that p0,µ,η [ ⟨y (k) t − y (k) s , f⟩ ]6 ≤ c0(t− s) 3 2 , 0 ≤ s < t ≤ 1, k ≥ 1, f ∈ s+(r). proof. applying the elementary inequality |x+ y|n ≤ 2n−1(|x|n + |y|n), x, y ∈ r, n ≥ 0, there are positive constants d1, d2, d3 such that p0,µ,η [ ⟨y (k) t − y (k) s , f⟩ ]6 ≤ d1 ( p0,µ,η [ ⟨y (k) s , pt−sf − f⟩ ]6 +p0,µ,η [ ⟨y (k) s pt−s − y (k) t , f⟩ ]6) ≤ d2 ( ∥s(pt−sf − f)∥61 + (t− s) 3 2 ∥sf∥61 ) ≤ d3 ( ∥pt−sf − f∥6∞ + (t− s) 3 2 ∥f∥6∞ ) advances in systems science and applications (2012) vol.12 no.3 293 by lemmas ?? and ??. since phf−f h −→ ς 2f ′′ = af as h → 0, h 7→ phf−f h is strongly continuous on h ∈ [0, 1]. this implies that there exists a positive constant d4 such that ∥phf − f∥∞ ≤ hd4 for h ∈ [0, 1]. let c0 = d3(d 6 4 + ∥f∥6∞). we finish the proof. 2 lemma ?? and kolmogorov’s criterion lead to lemma 2.7 for any f ∈ s (r), the sequence {⟨y (k) t , f⟩ : t ≥ 0; k ≥ 1} is tight in c([0,∞),r). proof of theorem ??. by lemma ?? and proposition ??, the result follows from [13, theorem 6.15]. 2 proposition 2.8 for each k ≥ 1 and f ∈ cp(r), m (k) t (f) := ⟨y (k) t , f⟩− ∫ t 0 ⟨y (k) s , af⟩ds− ∫ t 0 ( ⟨ξ(k)r , k 1 2 f⟩ − ⟨η, prk 1 2 f⟩ ) dr, t ≥ 0 is a continuous martingale. proof. fix k ≥ 1. by the markov property of {x(k) t : t ≥ 0}, for t > s ≥ 0, p0,µ,η ( ⟨x(k) t , f⟩ − ∫ t s ⟨x(k) r , af⟩dr ∣∣fs ) =⟨x(k) s , pt−sf⟩+ (t− s)⟨ξ(k)s , pt−sf⟩ − ∫ t s ( ⟨x(k) s , pr−saf⟩+ (r − s)⟨ξ(k)s , pr−saf⟩ ) dr. then by [14, proposition 1.5], we have p0,µ,η ( ⟨x(k) t , f⟩ − ∫ t 0 ⟨x(k) r , af⟩dr ∣∣fs ) =⟨x(k) s , f⟩ − ∫ s 0 ⟨x(k) r , af⟩dr + ∫ t−s 0 ⟨ξ(k)s , prf⟩dr and ⟨µ, ptf⟩+ t⟨η, ptf⟩ − ∫ t 0 ( ⟨µ, praf⟩+ r⟨η, praf⟩ ) dr = ⟨µ, f⟩+ ∫ t 0 ⟨η, prf⟩dr. recalling y (k) t = k 1 2 (x (k) t − µpt − tηpt), we have p0,µ,η [ m (k) t (f) ∣∣fs ] =m (k) s (f). 2 294 xu yang:central limit theorems for the single point catalytic... lemma 2.9 there is a locally bounded function t→ c(t) on [0,∞) such that sup k≥1 p0,µ,η { sup 0≤t≤t ∣∣∣⟨y (k) t , f⟩ ∣∣∣2} ≤ c(t )∥f∥22, t ≥ 0, f ∈ h2(r). proof. we only consider f ∈ s (r). for each k ≥ 1, p0,µ,η { sup 0≤t≤t ∣∣∣⟨y (k) t , f⟩ ∣∣∣2} ≤ 3 p0,µ,η { sup 0≤t≤t m (k) t (f)2 } + 3 p0,µ,η {[∫ t 0 ∣∣∣⟨y (k) s , af⟩ ∣∣∣ ds]2} +3 p0,η {[∫ t 0 ∣∣∣⟨ξ(k)s , k 1 2 f⟩ − ⟨η, psk 1 2 f⟩ ∣∣∣ ds]2} . using hölder inequality, fubini theorem and (5), we have p0,η {[∫ t 0 ∣∣∣⟨ξ(k)s , k 1 2 f⟩ − ⟨η, psk 1 2 f⟩ ∣∣∣ ds]2} ≤2 t ∫ t 0 ds ∫ s 0 ⟨η, ps−r(·)⟩ (prf(0)) 2 dr. by proposition ??, m (k) t (f) is a continuous martingale. then by doob’s martingale inequality, hölder inequality and fubini theorem, we have p0,µ,η { sup 0≤t≤t m (k) t (f)2 } ≤ 12 p0,µ,η {∣∣∣⟨y (k) t , f⟩ ∣∣∣2}+ 12 t {∫ t 0 p0,µ,η [∣∣∣⟨y (k) s , af⟩ ∣∣∣2] ds} +24 t ∫ t 0 ds ∫ s 0 ⟨η, ps−r(·)⟩ (prf(0)) 2 dr. by lemma ??, there is a constant c such that [prf(0)] 2 ≤ 1√ 2πςr q0(f) 2 ≤ c√ 2πςr ∥f∥20. it is easy to show that ⟨µ, ps−r(·)⟩ ≤ 1√ 2πς(s−r) (9µ 9p +µ([− √ 2pπςs, √ 2pπςs]) ) . then by lemmas ?? and ??, there is a locally bounded function t → c1(t) on advances in systems science and applications (2012) vol.12 no.3 295 [0,∞) such that p0,µ,η [∣∣∣⟨y (k) s , f⟩ ∣∣∣2] ≤ c2(s)∥f∥20 and p0,µ,η [∣∣∣⟨y (k) s , af⟩ ∣∣∣2] ≤ c2(s)∥f∥22 for any all s ≥ 0. this finishes the proof. 2 proof of theorem ??. using [13, corollary 6.16], the result follows from proposition ??, lemmas ?? and ??. 2 acknowledgements mathematics subject classifications (2000):supported by nsfc (no.11071021),primary 60j80; secondary 60g20, 60j68. references [1] dawson, d.a, fleischmann, k. (1994), “a super-brownian motion with a single point catalyst”, stoch. proc. appl, vol.49, pp.3-40. [2] fleischmann k, xiong j. (2005), “large deviation principle for the single point catalytic super-brownian motion”, markov processes related fields, vol.11, pp.519-533. [3] yang, x, zhang, m. (2010),“ moderate deviation for the single point catalytic super-brownian motion”, acta mathematica sinica, english series, vol.28, pp.1799–1808. [4] li, z.h, wang, l. (2010), “fluctuation limits of the super-brownian motion with a single point catalyst” (submitted). [5] hong w.m. (2002), “longtime behavior for the occupation time processes of a super-brownian motion withrandom immigration”, stoch. proc. appl, vol.102, pp.43-62. [6] hong w.m. (2003), “large deviations for the super-brownian motion with super-brownian immigration”, theoret. probab, vol.16, pp.899-922. [7] hong w.m, li z.h. (1999), “a central limit theorem for super-brownian immigraion”, appl. prob, vol.36, pp.1218-1224. [8] li z. h, shiga t. (1995), “measure-valued branching diffusions: immigraiton, excursions and limit theorems”, math. kyoto univ, vol.35, pp.233274. [9] zhang, m. (2004), “large deviation for super-brownian motion with immigration”, j. appl. prob, vol.41, pp.187-201. [10] zhang, m. (2005), “functional central limit theorem for the super-brownian motion with super-brownian immigration”, j. theoret. probab, vol.18, pp.665-685. 296 xu yang:central limit theorems for the single point catalytic... [11] li, z.h. (2011), measure-valued branching markov process, springer, berlin. [12] iscoe i. (1986), “a weighted occupation time for a class of measure-valued critical branching brownian motion”, probab.th. rel. fields, vol.71, pp.85116. [13] walsh, j.b. (1986), “an introduction to stochastic partial differential equations”, in: ecole d’eté de probabilités de saint-flour xiv -1984, pp.265-439, lecture notes math, pp.1180, springer, berlin. [14] ethier, s.n, kurtz, t.g. (1986), markov processes: characterization and convergenc, wiley, new york. corresponding author xu yang can be contacted at:xuyang@mail.bnu.edu.cn мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 62-70 method of parametric structural synthesis of the machine-assembly departments anatoliy serduk, aleksandr sergeev, aleksandr rusyaev, sergey kamenev orenburg state university, orenburg, russia annotation reconstruction and technical re-equipment of enterprises is preferable to the creation of a similar capacity as a result of new construction. stringent requirements apply to projects of modernization of production, design methods used, duration and quality of the project works. traditional methods of designing high-tech assembly areas, based on average calculations are inefficient and weakly correlated with the automated production process. important research task is to develop a method of automated design of assembly areas, based on computer modeling of technological processes and automated synthesis of structure and parameters of the production equipment. the article describes the formalized representation of the production process machining industries, including operations of separation and assembly. developed algorithmic, mathematical and information support of the process simulation of the machining industry, including operations of separation and assembly. implemented the lexicographic algorithm generates the best options for production jobs. these options are processed statistically. developed a method of structural-parametric synthesis of machining industries. the method takes into account changes in production item due to the simulation of a sample optimal replacement jobs that have the highest probability. the use of the developed method will ensure the optimality of the design solutions according to the criteria of efficiency of functioning of the machining industries. keywords: machine-assembly plant, structural-parametric synthesis, lexicographic algorithm, simulation plant 1 introduction a reorganization of existing enterprises is most efficient way to adjust for challenging market environment. return of the reorganization investment is estimated to be 3 times faster comparing to establishment of similar facilities. in the conditions of the formation of a new technological order the tasks of the reconstruction and re-equipment are important for vast majority of machine-building enterprises in the country especially in adopting of new product ramping. highly competitive product markets require very efficient organization of manufacturing facilities saving material and human resources. level of manufacturing automation is most crucial to achieve production efficiency and, therefore, production of the competitive products. in this regard, modernization of manufacturing projects, design styles, timing and quality of a design work are subjected to very demanding requirements. the traditional averaged calculations methods that used for design of the high-tech machineassembly plant (map) are ineffective with low correlation to the automated manufacturing process. for this reason computer modeling of manufacturing processes is necessary to improve the methods of a computer-aided design of the map. on the basis of this, the improvement of computer-aided design of the map based approaches is important scientific challenge that would provide great benefit to the country advances in systems science and application(2016) vol.16 no.4 63 economy. solution of the specified problem is possible with development of computer modeling of manufacturing processes that provides the automated structural-parametric synthesis of the map. the developed method will provide the optimality of accepted design solutions according to criteria of operational efficiency of the machine-assembly lines. 2 review of works on the subject of research a great contribution to the theory of design and optimization of manufacturing processes in the machining production made the researches of [1]-[8] and other authors. the successful solutions for design and optimization of manufacturing processes are offered by the companies who develops the special software such as the dassault systems, siemens plm software, autodesk, anylogic company, mstu «stankin», sprut-technology. in accordance to numerous analyses in the field of a machining production design, systems of operating scheduling, systems of simulation techniques which implemented on the machinebuilding enterprises the design-automation systems cannot be developed without connection with the operating scheduling regardless of production seriality. therefore the leading world companies provide the set of solutions for production design including the package to compose production schedules which are used in the maintenance process of created system. the domestic developments in the field of operating scheduling such as «fobos», «sprut-okp» aren’t the tools for manufacturing processes modeling that makes difficulties for getting of statistics on the production assortment. the solutions offered by foreign vendors which are leaders in the field of modeling (tecnomatix, delmia) has a wide set of tools but their manufacturing application for domestic enterprises involves the additional costs on adaptation to production standards, purchasing of unused modules and the development of specific subsystems. therefore, the development of automated system for design of the machine-assembly departments that include the module for composition of production schedules has a great importance. the mathematical model was based on the method for automated construction of sequence diagrams which was developed in orenburg state university and described at the works of [5] [9-12]. 3 simulation algorithm for the working process of the map that takes into account the separation and assembly operations the technological process of manufacturing of the product named «drill pipe» was used as model object for development of the product structure. this technological process contains the following operations: sawcut the pipe in a predetermined size, cutoff the workpieces for manufacturing of the nipple and coupling, machining of the pipe, nipple and coupling, welding the nipple and coupling with the pipe, machining of the drill pipe. the product structure in the form of a bi-directional graph (fig. 1) is proposed to account partitioning and assemblage operations. on this figure left to right shown the direction of workpieces flow and right to left – the information about workpieces required for pipe assembly. in the block specified the nodes which are connected with graph according to the technological process of manufacturing and composition of the product. at the first blocks the nodes wherein specified the composition of the product are equal to zero, at the last blocks – on the place of node that indicate the following manufacturing stage, assigned the value «–1», which fix the machining completion. the operations of partitioning and assemblage of the products introduce significant changes to the work of modeling subsystem. thus, for example, without the data of operations the user by itself calculates and sets the number of workpieces that start up to manufacturing. in our 64 anatoliy serduk, aleksandr sergeev, aleksandr rusyaev, et al: method of parametric structural synthesis of the machine... modeling case it is necessary to automatically provide the compliance between the number of assembling products, number of assembly units and amount of a raw stock for workpieces production. number of machined workpieces in a batch start up depends on whether the product is unit or workpiece and is described by the expression:     .. .. , , prodprod prodprod vfromtakenzifzdevedv zincludevifzv z (1) where vprod. – number of products to be produced, pcs.; z – number of assembly units with the same name belonging to the product or number of workpieces produced from raw stock, pcs. fig. 1 bi-directional graph of the product development with account the blanking and assembling operations the amount of a raw stock (pcs.) required for manufacturing of given number of workpieces can be determined by expression:      residuewithoutzondividedvifzv residuewithoutzondividedtdonvifzv z prodprod prod .. . ,div ',1div (2) thus, this expression takes into account the situation where, for example, the one bar can be sawn to 8 workpieces for nipple. but if required to make ten such workpieces then it is necessary to put the two bars on sawing operation. for the selection of an acceptable variant of the work order (wo) it is necessary to analyze the product composition and distribute the productive time of assembly units (au) between the equipment groups: , 100 ∑, iji i i ft fdf f    at i=1.. g, j=1..n, (3) where fi – fund of productive time of the equipment group i, min; тi,j∑ – total productive time of the workpieces machining in composition of product j on the equipment group i, min; g – number of the equipment groups, pcs; n – assortment of the delivered products, pcs. then the total productive time of the workpieces machining in composition of product j on advances in systems science and application(2016) vol.16 no.4 65 the equipment group i determines by expression: jji prod n j prodji vтt . 1 .∑, ∑ ,   , (4) where тprod.i,j – productive time of product j on the equipment group i with account of product composition (min), determined by equation: ∑ 1 ,,, kol k jauprod zтт kjiji   , (5) where тau i,j,k – productive time of assembly unit k of product j on the equipment group i, min; kol – number of au in product j, pcs. the developed structure for describing of the products allows keeping and using the information about assembly units and raw stock in the modeling process. based on equations above the algorithm for modeling of manufacturing process was developed (fig. 2), the corresponding software was accomplished and estimation of adequacy was performed. the input data are the capabilities of primary and auxiliary equipment and geometrical position of process modules. the output data are the actual duration of work cycle, coefficient of equipment utilization, productivity, utilization by the equipment groups and machine tools, sequence diagram for work of map. fig.2 simulation algorithm for the working process of the map that takes into account the separation and assembly operations 66 anatoliy serduk, aleksandr sergeev, aleksandr rusyaev, et al: method of parametric structural synthesis of the machine... 4 synthesis algorithm for operating schedule based on the use of the truncated lexicographic algorithm in the conditions of multiproduct manufacture the multiple passes of manufacturing process on the different variants of wo for collection of statistical data about the productive efficiency, in the dependence of content of the batch run are required. as it is well known, the brute force approach of all possible variants of the batch run has a factorial dependency and cannot be used for real life problems even on the modern highly productive computers. the formation of work orders relates to problems in combinatorial optimization and may be reduced to statement of knapsack problem. let exists the set of items consisting of the m names and l parameters. then the characteristics of items will defined the by vector ][ia  , which determines the values of each parameter of the item: ],...,,[][ ,2,1, liii aaaia   , where  mi ,...,2,1∈ . the number of items of the each name n[i] may be the any integer in the range  maxm0,1,...,∈ n[i] , where mmax – maximum number of the identical items. within a predetermined range of constraints  ][],...,2[],1[ minminminmin lssss   and  ][],...,2[],1[ maxmaxmaxmax lssss   it is required to find the combinations of packing which meets the specified constraints. in the mathematical statement the conditions of possibility of the packing for given combination of items are described by the following expressions:              ][≤≤][ ... ]2[≤≤]2[ ]1[≤≤]1[ max 1 min max2 1 min max1 1 min ∑ ∑ ∑ lsnals snas snas mlm m i mm m i mm m i (6) when the condition (6) meets the constraints for combination k then number of items in the knapsack ][∑ kn will be ∑ 1 ∑ ][][ m i inkn   . subject to the introduced mathematical support the algorithm for generation of the work orders variants based on the branch and bound method was developed if in a production flow of the processing of the parts from different groups without their interleaving is found a situation when the some machine tools are not active due to the specific of a manufacturing cycle then it may be concluded that the given situation involves the abrupt reduction of the productive efficiency. therefore, attempt to optimize the operating schedule through interleaving the parts of various groups in a sequence can be engaged. solution of this problem was realized with truncated lexicographic algorithm (fig. 3). the number of steps in this algorithm was reduced at the expense of the selection from the sum of all rearrangements of the sequence run the rearrangements with the predefined interval only. for statistical treatment of the modeling results for sampling of the wo the advanced algorithm was developed that allows to receive the frequency histograms and cumulative probability plots for the various factors of map efficiency. comparison of modeling results of the work orders with and without optimization reveals higher efficiency for optimization with the sequence run. to account the sensibility of the map to alterations of the manufactured product assortment the technique for selection of representative batches based on the estimated probability of higher efficiency of the manufacturing system was realized (fig. 4). selection of the representative batches that allowed to use in the developed computer model of the map in the algorithm of advances in systems science and application(2016) vol.16 no.4 67 structural-parametric synthesis of the production equipment is subject to the wide assortment of produced objects. fig.3 synthesis algorithm for operating schedule based on the use of the truncated lexicographic algorithm fig.4 statistical analysis of the work orders 68 anatoliy serduk, aleksandr sergeev, aleksandr rusyaev, et al: method of parametric structural synthesis of the machine... 5 the development of the method of structural-parametric synthesis for machine-assembly plants the next stage is the development of the method of structural-parametric synthesis for machineassembly plants. the introduced method uses the three strategies based on using of genetic algorithms, hooke-jeeves method and alternating-variable descent method. in the process of computing experiments the structural-parametric synthesis method is found to be most effective if based on the genetic algorithms. it may be explained by the fact that the two other methods use the same principle of optimization, i.e. here it matches the optimal value of the one parameter when the other parameters are constant. the advantage of the genetic algorithm is the variations of all synthesized values on every iteration step that as the result allows to get the uniform set of parameters that provide the maximum efficiency of the machineassembly departments. the other distinctive feature of the algorithm is that algorithm during the execution takes into account all used efficiency indexes: usage and duration of the actual equipment work. this approach allows to avoid the procedure of criteria convolution so as to proceed for the complex efficiency index. meanwhile, the parameters selection is performed with regard to extreme value from all efficiency indices. for solution integrity in the problem of structural-parametric synthesis of the machineassembly departments the algorithm for synthesis of coordinate arrangement of equipment was introduced. in the capacity of synthesis method the genetic algorithm based on the chromosome conception wherein the values of genes not modified but their locations are changing is implemented. this algorithm was realized in the software «komponovka 1.3». on the basis of computer model of the manufacturing system the method of structuralparametric synthesis of the machine-assembly departments was developed (fig. 5) that take into account the partitioning and assemblage operations of the produced objects. this method takes into account the change of manufacturing assortment at the expense of modeling the sampling of optimal wo which has a maximum probability of production starting. fig. 5 method of structural-parametric synthesis of the map on the framework of introduced algorithms the program «plan3d» was developed which is possible to transfer the generated model of map to cad-system kompas and automatically advances in systems science and application(2016) vol.16 no.4 69 construct the layouts of designed machine-assemble department in 2d and 3d formats. on the layout the table with the selected parameters of manufacturing equipment is also displayed (fig. 6). fig. 6 3d model of equipment arrangement for machine-assembly department the design data that formed as a result of automated structural-parametric synthesis are easy integrated into uniform information space of enterprise and allow to make changes via tools of cad-system kompas. 5 conclusions in this work it was determined that the modern approach to design the machine-assemble departments is a detailed analysis of the design solutions related to selection of structure and parameters of manufacturing equipment via computer modeling of map with the running on multiple manufacturing processes. the formalized presentation for functioning process of the machine-assemble department was performed. this presentation which describes the dynamics of assembly units manufacturing is based on the diagrammatic work method and may be used for most categories of the machinebuilding products. the algorithm for synthesis of optimal operating schedules which allows reducing a turnaround time by 10-15% was developed. the method for designing of machine-assembly departments which distinctive feature is the usage of operating scheduling procedures and computer modeling of the map functioning was developed. it was proved that the inclusion the procedures of an operating scheduling into method for design allows to increase the map efficiency on 23%. the program realization of a developed algorithms and procedures was designed, tested and approved. the developed software allows to form the complex of design solutions realized as 3d model of the equipment arrangement and nomenclature of the equipment parameters on the base of computer simulation of functioning and analysis of structure and parameters of the map. 70 anatoliy serduk, aleksandr sergeev, aleksandr rusyaev, et al: method of parametric structural synthesis of the machine... references [1] g. n. melnikov and v. p. voronenko(1990), designing of the machine-assembly departments, machine building, moscow,. [2] l. u. lishinsky(1990), structural and parametric synthesis of the flexible manufacturing systems, machine building, moscow,. [3] r. r. zagidullin(2010), "control the processes of the machine building enterprise", stin, no. 11, pp. 34-39. [4] i. serduk, a. i. sergeev (2005), "designing of the flexible manufacturing systems with the predefined payback period", stin, no. 11, pp. 20-26. [5] i. serduk, r. r rakhmatullin and a. p. zelenin(2010), "preproject analysis of the flexible manufacturing cells and decision support tools", herald of machine building, no.10, pp. 86-91 [6] u. m. solomencev, r. r. zagidullin and e. b. frolov(2010), "planning in the modern production control systems", information technologies and computational systems, no. 4, pp. 77-87. [7] r. m. usupov(2012), "national society of simulation modeling of the russia begin of the way", cad/cam/cae observer, no 2, pp. 10-18. [8] y. koren, u. heisel, f. jovane, t. moriwaki, g. pritschow, g. ulsoy and h. van brussel (1999), reconfigurable manufacturing systems. university of michigan. [9] s. rusyaev and t. o. rusyaeva(2012), "model for calculation of capacity of an automated product stock", software products and systems, no. 2, pp. 82-86. [10] i. sergeev, m. a. kornipaev, a. a. kornipaeva and a. s. rusyaev(2010), "genetic algorithms in the structural-parametric synthesis of flexible production systems", russian engineering research, vol. 30, no. 4, pp. 404-407. [11] i. sergeev and a. s. rusyaev(2013), "mathematical description of the shift", eastern european scientific journal, no. 4, pp. 64-68. [12] i. sergeev, m. a. kornipaev, r. r rakhmatullin and a. s. rusyaev (2008). "structuralparametric synthesis of the flexible manufacturing systems with application of the genetic algorithms", the study, orenburg,. corresponding author anatoliy serduk can be contacted at: yal05@mail.ru. advances in systems science and applications (2014), vol.14, no.2 158-169 teaching chaos with a pendulum to greek secondary school students constantine skordoulis, vasilis tolias, dimitris stavrou, kostas karamanos and aristotelis gkiolmas department of education, national and kapodistrian university of athens, athens, greece abstract in this paper we report on how a commercial “chaos pendulum system” was used to teach aspects of chaos theory to greek upper secondary school students. the principal aim of this project was to investigate to what extent students can develop an understanding of the chaotic behaviour using representations in the phase space. the didactical methodology used was that of the “physics workshop” where students follow a discovery type laboratory course with detailed lab worksheets and are guided to the study of chaotic motion starting from observations of the oscillation of a simple physical pendulum. interesting conclusions concerning students’ ability in using the technology associated with the microcomputer based laboratory (mbl), the development of a qualitative understanding of the chaotic behaviour using the representation in the phase space and in interpreting graphs can be drawn from the analysis of their recorded answers. 1 why teaching chaos? the study of non-linear systems has become a major area of research in science as well as in philosophy. this research has led to a shift in the emphasis concerning fundamental scientific c philosophical concepts such as ‘determinism’, ‘predictability’, ‘causality’, and ‘chance’. there are systems, the chaotic pendulum is one of them, which are very ‘sensitive’ to small changes in the initial conditions when the process is running. subtle changes of the initial parameters lead to a large deviation in the path followed so that detailed prediction of the system’s behavior is not possible. despite the irregular behavior of the system in the representation of the movement in the phase space a characteristic structure emerges indicating a kind of order. a number of non-linear systems develop complex self-organizing structures. self-organizing structures as well as structures that appears in the phase space could be described by the model of fractals. in this sense, chaos, self-organisation and fractals are intimately linked. from an educational point of view this new domain has an important contribution to scientific literacy by developing a more adequate worldview and a more advances in systems science and applications (2014), vol.14, no.2 159 relevant view concerning the nature of science[1] . more specifically, teaching nonlinearity to students makes them understand that the mechanistic worldview has been challenged, the physical world is not governed by simple deterministic laws leading to complete predictability, that causality and chance are an integral part of physical reality. a considerable amount of research work in science education has been done on issues of non-linear systems in the last years[1-4] . nevertheless there is a deficiency in empirical investigations concerning the possibility to teach and learn of aspects of the chaos theory using representations in the phase space. the principal aim of this research is to investigate to what extent students can develop an understanding of the chaotic behavior using representations in the phase space. for this task our approach adopts the mbl methodology, which allows a direct connection between the real experiment and the abstract graphical representation, based on the ‘workshop physics’ project using the pasco chaotic pendulum system. in the following paragraphs a short description of the mbl methodology, the ‘workshop physics’ project and the pasco chaotic pendulum system is given. 2 microcomputer based laboratory c (mbl) methodology the development of the laboratories using mbl-tools has been inspired by the pedagogical approaches applied in ‘realtime physics’[5] and ‘tools for scientific thinking’[6] and by research in physics education[7-9] . for some decades sensors interfaced with a computer have been used in most physics research laboratories. an attachment of a sensor to a computer creates a very powerful system for collection, analysis and display of experimental data. today several systems, specially developed for schools and undergraduate courses, are commercially available. in a microcomputer based laboratory, students do real experiments using different sensors connected to a computer via an interface. experimental data from real-world experiment are immediately presented on the computer screen in a graphical format and can immediately be analysed. one of the main educational advantages of using mbl is the real-time display of experimental results and graphs thus facilitating direct connection between the real experiment and the abstract representation. because data are quickly taken and displayed, students can easily examine the consequences of a large number of changes in experimental conditions during a short period of time. brasell (1987)[10] has shown that even a very short delay in the presentation of experimental data is detrimental for student comprehension. 160 constantine skordoulis:teaching chaos with a pendulum to greek secondary school students 3 the ‘workshop physics’ course the ‘workshop physics’ course has developed since 1986 at dickinson college by p. laws (2004) and is based on mbl methodology. the major objective of workshop physics courses is to help students understand the basis of knowledge in physics as an interplay between observations, experiments, definitions and mathematical description. the main idea of the course is to engage students in laboratory activities without any previous lecturing. in this discovery type laboratory course the students gain real-world experience using hands-on experiments. students following “workshop physics” courses learn to use computer tools to collect data and to develop mathematical models of phenomena. the workshop physics curriculum is centred on a workbook-style activity guide, that consists of twenty-eight units[11] . these units are based on activities that include predictions, qualitative observations, explanations, equation derivations, mathematical model building, quantitative experiments and problem solving. students in our sample class did activities selected from unit 15 on “oscillations, determinism and chaos”. this unit includes a series of observations and experiments on a physical pendulum system that becomes increasingly complex and eventually appears to undergo chaotic behaviour. the major advantage of the workshop physics for us is the ‘step by step’ approach used in unit 15 where, starting from observations of the motion of the physical pendulum, students can gradually built up their knowledge up to a point of understanding chaotic behaviour. 4 the chaotic pendulum system the chaotic pendulum system by pasco is one of the chaotic pendulums that are commercially available. a detailed comparison of the chaotic pendulums available has been published[12]. the pasco chaos pendulum can be used in two different configurations: a horizontal and a vertical. both configurations have been analysed and modelled in the relevant literature. we have used the horizontal configuration, shown in fig.1, a brief description of which follows in the next paragraphs. the oscillator consists of an aluminum disk connected to two springs. a mass attached on the edge of the aluminum disk makes the oscillator non-linear. the frequency of the sinusoidal driver can be varied to investigate the progression from predictable motion to chaotic motion. magnetic damping can be also adjusted to change the character of the chaotic motion. the angular position and velocity of the disk are recorded as a function of time using a rotary motion sensor connected to a computer via an interface. data are recorded and displayed advances in systems science and applications (2014), vol.14, no.2 161 fig.1 the pasco chaos pendulum with the horizontal configuration using the data studio software. the chaotic behaviour of the driven non-linear pendulum is explored by graphing its motion in phase space and by making a poincare plot. these plots are compared to the motion of the pendulum when it is not chaotic. a real time phase space diagram is made by graphing the angular velocity versus the displacement angle of the oscillation, while the poincare plot is also graphed in real time and can be superimposed on the phase plot. this is achieved by recording the point on the phase plot once every cycle of the driver arm as the driver arm blocks a photogate. the variables of the experiment (driving frequency, displacement amplitude and magnetic damping) can be varied so as to cause the regular motion to become chaotic.students, using the data studio software, can view the following graphs: 1. angular displacement (θ) vs. time (t) 2. angular velocity (ω) vs. time (t) 3. phase space: angular velocity (ω) vs. angular displacement (θ) 4. poincare plot: angular velocity (ω) vs. angular displacement (θ) plotted only once per period of the driving force. the phase space and the poincare plot are particularly useful for recognizing chaotic oscillations. when the motion is chaotic, the graphs do not repeat. 5 didactical activities our sample class consists of six gifted final year secondary school students, with high scores in maths and physics, selected from a private school based in athens and worked in two teams of three persons during three lab sessions of three hours each (a total of nine hours of lab instruction) in asel where the apparatus is based. 162 constantine skordoulis:teaching chaos with a pendulum to greek secondary school students the students had not any previous knowledge either of the concepts or the graphical representations (eg phase space, poincare mapping) associated with chaotic motion and they were not lectured on them before commencing on the lab course. in accordance with the greek upper secondary school science curriculum the students had a fairly good knowledge of simple harmonic motion, rotational dynamics and electromagnetism and so it was rather easy for them to understand the operation of the dc motor driver, the rotational sensor and the interface during a demonstration lab period conducted by the teacher. the guidelines of the activities selected from unit 15 on “oscillations, determinism and chaos” of ‘workshop physics’ were modified to meet the tasks and the aim of the present study. the students had to make observations, discuss the phenomena, give explanations and draw graphics, following the instructions in a worksheet designed by our group to meet the needs of the present research. the performance of the students, both individually and as a team, has been evaluated from the questions completed in the laboratory worksheets and from the discussions among them and with the teacher. during the lab period the teacher was giving the necessary information concerning the operation of the devices, analyzing the instructions for doing the experiment when necessary, and intervening only when the discussion among students had reached a dead end. below a detailed description of the procedure followed is given highlighting the ‘step by step’ approach and the students’ performance in every step. activity one: observation of the oscillation of the physical pendulum. in the first activity students made a series of observations, at different oscillation angles, of the motion of the aluminium disk mounted on the low friction bearing with and without various quantities of mass attached to its edge. this disk was part of the physical pendulum system. the students keep a record of their observations and discuss about the oscillation of the physical pendulum (fig.2). activity two: time series and phase space plot of the physical pendulum. next the rotary motion sensor was attached to a universal laboratory interface. the students using the data studio software could create a plot of the angular displacement of the pendulum as a function of time in real time (time series plot) as well as a plot of the phase space (angular velocity as a function of angular displacement phase plot). the students observe the time series graph, the phase plot graph and the experiment in real time. it has to be emphasised that the students had no difficulty in understanding advances in systems science and applications (2014), vol.14, no.2 163 fig.2 observation of the oscillation of the physical pendulum the phase space graph. they started from the graph of the time series, which is very familiar to them and they can easily connect it with the motion of the disk (experiment in the real world). then they correspond specific points of the time series graph to the graph of the phase space (see fig.3) and connect these with the oscillation in the real experiment. fig. 3 time series and phase space plot of the physical pendulum activity three: determination of the natural frequencies of the system. students were asked to modify the physical pendulum so that a string with 164 constantine skordoulis:teaching chaos with a pendulum to greek secondary school students the two springs attached and the driver motor are coupled to the aluminium disk with the edge mass in place (see fig.4). fig.4 determination of the natural frequencies of the system then with the motor turned off the students explored the natural frequencies of the system with and without the spring coupling and with and without the edge mass in place. the students using data studio software create the time series graph ω(t) and take a direct reading of the natural frequency of the system in the various cases. then students activate the driver motor and they set the driver frequency very close to the natural frequencies of the physical pendulum system thus creating a situation “near resonance”. the shape of the phase space graph is elliptic (see fig.5). fig.5 the shape of the phase space graph “at resonance” the amplitude in the time series graph remains constant (within experimental error) while the shape of the graph in the phase space is a circle. the students had some difficulty in understanding the shape of the phase space graph and the teacher had to talk about the phase difference between θ(t) and advances in systems science and applications (2014), vol.14, no.2 165 ω(t) in the corresponding time series graphs at resonance. at this point the poincare plot was introduced. a photogate was placed in the driver arm so as the plot can be produced. the students could not understand the physical significance and could not make any connection with the actual experiment despite the efforts of the teacher. activity four: exploring conditions for chaotic motion-sensitivity to initial conditions students were asked to place the damping magnet fairly close to the metal pendulum disk, activate the driver motion and set the driver frequency in a position that is different than any of the natural frequencies of the system and see if they can achieve a situation in which there is an irregular pattern in the time series graph (angular position vs. time). a typical pattern is shown in fig.6 for the time series graph and the corresponding phase space plot. fig.6 the time series graph and the corresponding phase space an important observation in this sample activity was to have each participant attempt to repeat a pattern like that shown in fig.6 by setting up the same initial conditions for the pendulum. all attempts failed after only a few seconds of run time. the students realize that this system is so sensitive to the initial values of angular position and angular velocity that it is impossible to recreate the initial conditions accurately enough to repeat a pattern on a graph. in this activity students initially show serious difficulties to explain the final form of the graph in the phase space and connect it directly with the motion of the device. nevertheless, they managed to give a qualitative explanation of the 166 constantine skordoulis:teaching chaos with a pendulum to greek secondary school students shape of the phase space graph by connecting it to the time series graph and then connect the motion of the pendulum with the time series graph. more specifically, students connect the random number of peaks on the waveform in the time series graph with the number of oscillations of the pendulum. in the next step, they relate the random number of peaks on the waveform in the time series graph with the irregular way of shaping the graph in the phase space and by extension to the motion of the pendulum. concerning the poincar plot, students can make use of the experimental apparatus and produce the plot, they can use the shape of the plot as a criterion to identify chaotic motion but they cannot relate it to phase space plot and/or understand the physical meaning of it despite the intervention of the teacher. in order to clarify if and to what extent students had develop an understanding of the phenomena discussed during this laboratory period, the same students have been asked six months later to repeat the course. in this second phase, the main emphasis was on students’ predictions about the process of the experiment and the kind of graphs expected. after the experiment was carried out students had to compare their predictions and their graphs with the graphs produced by the software based on the date collected from the apparatus and discuss the possible agreement or differences. the results of this second phase showed that students were able to make predictions about the behaviour of the system which are very successful, and their graphs in most cases are similar to the graphs drawn by the software. their explanations were also very satisfactory. in the following fig.7 and 8, the graphs drawn by the two groups of students are presented concerning their prediction about phase space plots for the oscillation of the physical pendulum and chaotic motion. fig.7 phase space plots for the oscillation of the physical pendulum and chaoticmotion by group a advances in systems science and applications (2014), vol.14, no.2 167 fig.8 phase space plots for the oscillation of the physical pendulum and chaoticmotion by group b 6 discussion and conclusions the results presented here are based mainly on students’ worksheets and the discussion, the teacher had with the students during and after the lab period. in sum, we can say, that despite the difficulties students have, they can go through these activities and develop up to a point a qualitative understanding of the chaotic behavior using the representation in the phase space. the ‘step by step’ approach, ie starting from observing the motion of the simple physical pendulum and ending with the study of the complex motion of the chaotic pendulum system, in combination with the simultaneous presence of the real experiment and the corresponding graphs facilitate students’ understanding. the students had the opportunity to discuss the graphs of simple motions, which are up to a point familiar to them, and connect them with the real experiments. they connect the time series graph with the phase space graph and by extension with the motion of the pendulum. based on the gained knowledge they could make also connections between the real experiment and the graphs in the case of the chaotic motion and give up to a point explanations of the observed complex behavior and the corresponding graphs. interventions of the teacher are necessary and concern mainly the clarification of the graphs in the phase space. this study gave insights about difficulties and opportunities of teaching a new field of physics using the mbl methodology. the results presented here are stem from a preliminary explorative study. thus they need further elaboration and should be tested in a broader sample. a further investigation based on the results presented here is intended. this work also contributes in the field of pendulum 168 constantine skordoulis:teaching chaos with a pendulum to greek secondary school students studies[13] by revealing aspects of the educational dynamic of the pendulum. apart from using the simple pendulum in cross disciplinary teaching giving students an understanding of the way that science has developed in relation to other forms of cultural life, the physical pendulum can be utilised to promote knowledge about the nature of science where causality and chance are characteristics of an integrated scientific worldview. references [1] komorek, m., stavrou, d., & duit, r. (2003), “non-linear physics in upper physics classes: educational reconstruction as a frame for development and research in a study of teaching and learning basic ideas of nonlinearity. in: d. psillos, p. kariotoglou, v. tselfes, e. hatzikraniotis, g. fassoulopoulos, & m. kallery (eds.)”, science education research in the knowledge based society, pp.269-276. [2] adams, h.m. & russ, j.c. (1992), “chaos in the classroom: exposing gifted elementary school children to chaos and fractals”, journal of science education and technology 1, pp.91-209. [3] chacn, r., batres, y. & cuadros, f.(1992), “teaching deterministic chaos through music”, physics education, vol.27, pp.151-154. [4] duit, r., komorek, m. & wilbers j. (1997), “studies on educational reconstruction of chaos theory”, research in science education. vol.27, pp.339357. [5] sokoloff, d., thornton, r. and laws, p. (1998), “realtime physics”, active learning laboratories. [6] thornton, r. (1989), “tools for scientific thinking: learning physical concepts with real-time laboratory measurements tools in redish, e. and risley, j. (eds), proc. conf”, computers in physics instruction, pp.177-189. [7] thornton, r. (1997), “learning physics concepts in the introductory course: microcomputer-based labs and interactive lecture demonstrations in wilson, j. (ed.)”, conference on the introductory physics course, pp.6986. [8] laws, p.w., (1997a), “a new order for mechanics”, conference on the introductory physics course, pp.125-136. [9] mcdermott, l. (1997), “how research can guide us in improving the introductory course”, proceedings of the conference on introductory physics course, pp.33-45. advances in systems science and applications (2014), vol.14, no.2 169 [10] brasell, h., (1987), “the effect of real-time laboratory graphing on learning graphic representation of distance and velocity”, journal of research in science teaching, vol.24, pp.385. [11] laws, p.w., (1997b), “workshop physics activity guide, new york: wiley laws, p.w., (2004), a unit on oscillations, determinism and chaos”, american journal of physics, vol.72, no.4, pp.446. [12] blackburn, j. a. & baker g. l. (1998), “a comparison of commercial chaotic pendulums”, american journal of physics , vol.66, no.9, pp.821. [13] matthews, m. r., (2001), “how pendulum studies can promote knowledge of the nature of science”, journal of science education and technology, vol.10, no.4, pp.359. [14] komorek, m. & duit, r., (2004), “the teaching experiment as a powerful method to develop and evaluate teaching and learning sequences in the domain of non-linear systems”, international journal of science education, vol.26, pp.619-633. corresponding author constantine skordoulis can be contacted at: kkaraman@ulb.ac.be. advances in systems science and applications (2017) vol.17 no.1 1 the concept of intelligent tutoring for enterprise staff as a component of integrated manufacturing control system development rustem a. sabitov, gulnara s. smirnova 1 , bulat. r. sirazetdinov, natalya. u. elizarova, alexander. v. eponeshnikov kazan national research technical university named after a. n. tupolev — kai, kazan, russian federation abstract. modern industrial manufacturing has high requirements for employees. workers must continually improve their competence. you must constantly test, train and retrain specialists of enterprises. in the light of manufacturing control problems and cognitive issues of identification, verification and other problems for processes of identification, modeling, and control, intelligent tutoring system for specialists training can be considered as an important part of manufacturing intelligent control and planning system. the development and a pilot implementation of adaptive approaches to the universal computer system of individual training, functioning on the basis of subject and meta knowledge bases and meta knowledge bases with the latest advances in information and communication technologies, and computing and artificial intelligence is a very important part of intelligent control and planning for manufacturing system. staff must continuously learn novel the techniques and use of the learning system in the educational process for solving practical problems in individual courses within their existing production competence and future qualification requirements. to improve the system functionality and its compliance with the modern educational process needs it is necessary to take into account the comments and suggestions from users. this system also can be used for more productive generation of ideas to improve the system and eliminate bottlenecks. actually an intelligent tutoring system helps to solve the broader problems, i.e. a learner understands and sees not only one option for solving the a problem task and working on a predetermined pattern, but begins to perceive the subject areas of the enterprise as a part of the whole system. key words: learning management systems, intelligent tutoring system, model identification, manufacturing control, cognitive issues of identification, positive-formed formulas. 1. introduction modern industrial manufacturing has high requirements for employees. workers must continually improve their competence. you must constantly test, train and retrain specialists of enterprises. an adaptive program of development of new competencies is needed to perform advanced enterprise plans. they should be formed depending on the program and production plans, as well as based on the analysis of skills of learners themselves. in the light of manufacturing control problems and cognitive issues of identification, verification and other problems for processes of identification, modeling, and control, intelligent tutoring system for specialists training can be considered as an important part of manufacturing intelligent control and planning system [1]. the proposed approach can be considered as a component of the global problem of forecasting and management of the development of productive forces and production relations of society depends on the individual capabilities of each worker, and as a component of the inverse problem of advanced training of specialists, who is "needed" for the economy. training of highly skilled professionals that are competitive in the labor market, as well as competent and able to work effectively in the specialty, is a prerequisite for the successful development of any country. the recently formed lack of trained engineers in high-tech industries reflects the prevailing tension between the modern engineering complexity and the actual training level of technical specialists with higher education. 1 corresponding author. e-mail: seyl@mail.ru 2 r. s. sabitov et.al.: the concept of intelligent tutoring for enterprise staff as a component of integrated manufacturing control system development today it is a large and growing in scale problem of many businesses. there is a lack of education: instead of bringing innovative ideas to the production, investments are needed for further training and retraining of young professionals. universities are not always able to train a sufficient quantity of decent level specialists in information technologies, data mining, machinebuilding design and technological specialties, integrated logistics, production planning and modeling, analysis and design of business processes, etc. [2]. the essence of the proposed approach in this paper is quite simple and can use well-known scientific and technical solutions. as necessary conditions for the successful introduction of the results and further development of a new training system are the rational use of modern and advanced information technologies, taking into account recent advances of cognitive science and developing of original high-tech methods for solving problems in the modern theory of intelligent control, the theory of “c cubed” (control, computation and communication), etc. 2. statement of the problem and the main results teaching many technical and mathematical subjects is impossible without learning problem solving. it is necessary not only to demonstrate the examples of reasoning in solving various problems, but also to offer the student a tool for problem solving and self-check of the results. comparison of the numerical value of the response task obtained by a student with a reference value is not sufficient to control the level of the student’s assimilation of the problem-solving technique. especially if he or she is asked to choose an answer from a fixed set of options. first, the student should find a solution in the general form (to obtain a formula), and only then perform the calculations. thus, we need to be able to compare two formulas up to equivalence transformations (in the general case is algorithmically undecidable), but first it is necessary to get the formula from the user (in a machine-readable format). second, to trace the student’s reasoning, you need to know at least what formulas and in what sequence he or she used (because we cannot count on the analysis of natural reasoning). that is, the task of getting the formula from the user and comparing it with the existing needs to be repeated many times. modern information technologies training mainly provides students with electronic versions of educational hypertext material, possibly using multimedia tools, as well as remote access to these information resources and remote interaction with the teacher and other students, including real time. the achievements of artificial intelligences are still used in a much lesser extent. this conclusion applies to the learning management systems and learning content (lcms), both commercial and open source systems (e.g., moodle). content of this work is the development and pilot implementation of adaptive approaches to the universal computer system of individual training, functioning on the basis of subject knowledge and meta knowledge bases with the latest advances in information and communication technologies, computing and artificial intelligence as an important part of intelligent control and planning manufacturing system. the work includes the presentation of the basic theoretical foundations, detailed solutions of typical examples; explanations of practical applications of the mathematical results, as well as challenges for independent solutions. it is not simple tests that require calculate a numeric value, or select one of several options, but it must provide control of the correct course of solving the problem, review and updating of the results. we are talking about the formation of the individual learning program material based on real opportunities to learn using the level of training dynamic selection and feedback. the entire process can be analyzed and logged, while training the next lesson; you will have more efficiency using the accumulated experience of the system. 3. brief description of the algorithm of the module "universal tasks solver on the basis of positive-formed formulas" the main components of the pilot software package [3] is the module "universal tasks solver on the basis of positiveformed formulas". the basis of knowledge representation here is a advances in systems science and applications (2017) vol.17 no.1 3 classic first-order language l of positively constructed formulas. positively formed formula the formula is derived from n concluding statements (n ≥ 1) using standard quantifiers and logical connectives &, ⸀, and which do not limit (in its syntax) representability of various properties. moreover, the presentation of knowledge in the form of language l using only standard existential and universal quantifiers, logical connectives conjunction and disjunction is not explicitly recorded and assumed in conjunctive and disjunctive branching-formulas [4]. output is to focus on the method of contradiction. inference rule in the calculus j positively constructed formulas denoted as ω. thus, an algorithm for finding solutions using calculus j, based on the language l can be represented as follows:  have a starting base b (which is given), i.e. many well-known facts;  have the relations of computability in the knowledge base that can be used to form the basis of questions to b;  by the operation ω find answers to some questions and expand the facts base b;  if the answer is made by the target issue, i.e. refutation of contradictions, we come to a successful conclusion;  if the answer to the question is made with disjunctive subformula branching, the fact base is split by the number of bases corresponding to the number of elements in the disjunctive formula;  in case of splitting database operations using ω again is searched for the new database response and so on, to successfully complete location solutions or not. the algorithm of the universal problem solvers in this case can be summarized as follows:  the raw data (given, goal) enter the solver from the graphical user interface (interactive equation editor);  the raw initial data base is formed by the facts;  loop solver associated with the knowledge base (which contains the ratio of computability) and selects the required ratio of computability;  on the basis of computability relations the solver generates a design according to the structure formulas, applies the rule ω and given the potential ramifications of the facts base (raw data), through all the branches in the original conjunctive branching;  solver adds facts to the base;  the process continues as long as the option is not found/solutions in the form of a refutation of the contradictions, otherwise the message about the absence of the possibility of solving the problem with this facts and knowledge database. schematically, the process of finding a solution to the problem can be represented by the following relationships between the graphical user interface, solver of problems, the original data and knowledge base (fig. 1-2) 4 r. s. sabitov et.al.: the concept of intelligent tutoring for enterprise staff as a component of integrated manufacturing control system development fig. 1. process of finding a solution to the problem 4. principles of formation of the formula editor for interactive intelligent tutoring system computer software formula editors are of great importance in teaching technical subjects. they are needed when you create training material, as well as directly in the learning process. it offers ideas for the construction of the formula editor you would really want to use. on the basis of these ideas a working formula editor has already been created. this equation editor is used in the interactive learning system of intellectual training of engineers [7]. these ideas are as follows. form formula by typing the name of the keyboard characters, operators, etc., that you want to insert. moreover, it made directly in the text of the formula. the system then replaces the typed text to the desired mark, operator or member. enough variants of the name that you use it (integral matrix, matrix, etc.). also, the system tries to guess what you want, with the first letter typed, and offers options for items to choose from. as you type, a list of options varies. this approach works well in integrated development environments used by programmers. here the system guesses what the programmer is going to write on, and offers options for continuing or replacement text. this feature is known as code completion. our experiments showed that the use of this approach for formulas reduces the formula typing time compared with the traditional interface to enter a formula using the mouse. revealed the following features. beginners do not need to study the menu and the editor of formulas to find the right icon. they already know names of the needed elements. the student, typing formula, can continue to think about the formula itself. he or she doesn’t need to transfer attention to finding the right button, the window or the icon. inserting the name of the element is made directly in the text of the formula, and not in a separate window. interface to enter a formula should contain the necessary elements to enter and control the formula using the mouse as in traditional systems, since it also has certain advantages. advances in systems science and applications (2017) vol.17 no.1 5 fig. 2. algorithm for finding the solution of the problem allow you to specify the semantics of typed formulas. that is, allow the user to specify exactly what he got – a superscript or degree opening bracket arithmetic expression or a list of function parameters, vector, matrix, etc. automatic guessing what is typed by the user, based on the context surrounding the formula. that is, if a function is assumed in calculations of v, but there is a variable with the same name, then the bracket for the symbol v rather marks the beginning of the list of its parameters, than anything else. also in the opposite direction, it is possible to use semantic information from a typed formula to automatically generate a list of user-entered or used variables, functions, and other objects. this is true for both research and training systems. for example, to request the user to provide the meaning of an introduction of a variable already present in the context of the problem being solved. a formula may be managed fully with the keyboard, waiting for the user’s next action. for example, adding new parameters of a function by pressing the comma. inserting new rows and columns by pressing the enter key and the space bar, respectively. add to a user the possibility to undo and redo actions and also the opportunity to return to any of the formulas in the timeline, copy from an old version of the desired piece and insert it into the current version. further some of the implementation details are provided. 6 r. s. sabitov et.al.: the concept of intelligent tutoring for enterprise staff as a component of integrated manufacturing control system development it is proposed to move away from the standard approach abstraction and unification of structures and methods. for a software implementation does not use the formula editor arithmetic data types (adt) and / or parse trees formats for mathematical formulas (mathml, etc.) and reflects the perception of the structure of the formula written by the user. for example, a formula does not seem tree operators /, operations on arguments (adt), and an ordered sequence of its constituent elements (variables, operators, etc.), see fig. 3. fig. 3. representation of formulas in the computer memory fig. 4. presentation exponentiation operation in the computer memory exponentiation is represented not as a transaction with two arguments (mathml, etc.). it is a characteristic of an element (variable, function, etc.) the extent to which he was elevated (fig. 4). division operation will then be presented in two ways (fig. 5). this approach simplifies the system. system behavior when navigating and editing a formula corresponds to the user's expectations, and not abstract structures in computer memory. it then becomes easier to predict what the user wants to do next, but has not yet made, and to help the user to perform these actions. the program realizes the possibility of undo and redo actions. changes are made through the design pattern (command). in this case a formula itself does not change (fig. 5). if you want to make changes to the formula, the program creates a special object (command object) with two methods – "apply changes" and "undo the changes." this object allows you to make changes to the formula these objects are stored in a special stack. if necessary, they in turn are popped from the stack, used for undo and placed in another stack to repeat actions. the command object, which inserts a new element into the formula, when recalling the method "apply changes", it inserts the same object as the first call. objects that removes an element after canceling of the removal return to the formula the object that has been deleted, and do not generate a new identical object. this allows you to insert, delete, and modify the formula to implement elements directly via pointers to the object in memory (since reached the immutability of pointers), and not only use the location element in the formula with respect to its root / start. application of direct pointers to mutable object simplifies writing and empowering online and accelerates the work itself online. advances in systems science and applications (2017) vol.17 no.1 7 fig. 5. substitution the division operation in the computer memory implementation of these ideas has already shown tremendous speed gain in typing a set of formulas and invaluable assistance to concentrate on the essence of the case that the user is working on. actually, the same formula editor can be used for scientific research and for writing articles. together with our colleagues, we continue to develop it. thus, the role of an editor of mathematical formulas is twofold: on the one hand, it should allow a student to be comfortable enough to express their thoughts, and on the other – should allow effective handling of user’s input. in doing so, it would be desirable to reduce the possibility of typing errors. moreover, the editor’s usage should not assume that a user has special knowledge, representation of formulas should be familiar and intuitive. these requirements are met by the flexible editor. formulas in the editor are presented in a familiar graphical form, as they would be printed in a textbook. the editor uses the concept of a template formula – a sort of a "canonical" representation in which this type of formula is written (which makes them easy to compare). a user selects a pattern corresponding to the desired formula, and then fills it with slots of necessary constants, variables and functions. symbols of constants, variables and functions are presented by graphic primitives, roaming through the mechanism of drag-n-drop. in addition, the editor allows you to further adjust the difficulty of a task, varying the amount of "designer details" available to the user. templates in the system are generated automatically from text formulas. in a formula (generated by parsing the text representation) variables, functions, constants are indicated. they are not inserted in the template automatically, instead of that cells for substitution variables, functions and constants are put in their places. other elements (operators, etc.) are inserted into the template unchanged. 6. conclusions the development and a pilot implementation of adaptive approaches to the universal computer system of individual training, functioning on the basis of subject and meta knowledge bases with the latest advances in information and communication technologies, computing and artificial intelligence is a very important part of intelligent control and planning for 8 r. s. sabitov et.al.: the concept of intelligent tutoring for enterprise staff as a component of integrated manufacturing control system development manufacturing system. staff must continuously learn novel techniques and use a learning system in the educational process for solving practical problems in individual courses within their existing production competence and future qualification requirements. to improve the system functionality and its compliance with the modern educational process needs it is necessary to take into account comments and suggestions from users [5-6]. staff can also use a complex learning system to have a holistic picture for more productive generation of ideas which can be used for improving the production system and eliminating bottlenecks. actually, an intelligent tutoring system helps to solve the broader problems, i.e. a learner understands and sees not only one option for solving a task and working on a predetermined pattern, but begins to perceive the subject areas of the enterprise as a part of the whole system. this happens largely because using the latest achievements of artificial intelligence, telecommunications and information technologies it is possible to study fast and find efficient solutions of a various options with the learner’s active participation. this also significantly simplifies the process of domain knowledge modeling. a similar effect can also be achieved with conventional training methods, i.e. using a constant direct contact between a teacher and a learner. obviously, to achieve a similar effect in the case of traditional teaching methods in average much more time and money is required than with the learning process of the developed software system. here the learning system actually allows you to simulate the process of individual learning. the effect of the individual approach in this case is very significant and convenient. you can continue with your work at any time that is not practicable at the standard training scheme. in the case of a group of traditional forms of education it is extremely difficult for a teacher to give sufficient time for an individual approach for each group member. the system can be used very effectively in groups where there are different levels of competence of learners. weaker learners will not hinder the work of strong learners. all motivated learners have an opportunity to improve their performance through their individual work. references [1] vassilyev, s.n., sabitov r.a., degtyarev, g.l., kozlov, v.v, malivanov, n.n., sabitov, sh.r. and sirazetdinov, r.t. adaptive approach to developing advanced distributed elearning management system for manufacturing. preprints of the 13th ifac symposium on information control problems in manufacturing, moscow, russia, pp. 2198-2203, 2009 [2] vasiliev s.n. and sabitov r.a. the knowledge economy and the intelligent management. proceedings of the x international conference chetaev, kazan, 2012. [3] vassilyev, s.n., sabitov sh.r., sirazetdinov b.r., smirnova, g.s., sukonnova a.a. architecture and functions tracking intelligent tutoring system "volga" №4, knrtu-kai named after a.n. tupolev, russia, kazan, 2012 [4] vasiliev s. n., zherlov a. k., fedosov e. a., b. e. fedunov (2000) intelligent control of dynamic systems. – moscow: physical and mathematical literature [5] smirnova, g.s., sabitov, r.a., sabitov, sh.r., korobkova, e.a, kislov, a.s. programmethodical complex of operations management №4, knrtu-kai named after a.n. tupolev, russia, kazan, 2012 [6] smirnova, g.s., sabitov, r.a., elizarova, n.y., sh.r. sabitov operational management system of enterprise production processes "1c: mes: cloudy production management. journal “automation in industry”, 2014, № 8, russia, moscow [7] sirazetdinov b.r., smirnova, g.s., korobkova e.a. principles of formation equation editor for interactive intelligent tutoring system for training engineers in aerospace disciplines. international scientific-practical conference "modern technologies and materials – the key point in revival of domestic aircraft building", kazan, 2010 advances in systems science and applications (2011), vol. 11, no. 1-2 1-26 time-optimal control of infinite variables parabolic systems with time lags given in integral form g. m. bahaa1 and m. m. tharwat 2 1 department of mathematics, faculty of science, taibah university, al-madinah al-munawarah, saudi arabia 2 departmentofmathematics,universitycollege,ummal-qurauniversity,makkah,saudiarabia emall: bahaa−gm@hotmail.com, zahraa26@yahoo.com abstract in this paper, the time-optimal control problem for second order parabolic system and also for (n× n)– parabolic systems with infinite number of variables involving constant time lags appearing in integral form in both the state equation and in the boundary condition is presented. some specific properties of the optimal control are discussed. keywords time-optimal control (n × n) parabolic systems operator with an infinite number of 1. introduction distributed parameters systems with delays can be used to describe many phenomena in the real world. as is well known, heat conduction, properties of elastic-plastic material, fluid dynamics, diffusion-reaction processes, the transmission of the signals at a certain distance by using electric long lines, etc., all lie within this area. the object that we are studying (temperature, displacement, concentration, velocity, etc.) is usually referred to as the state. the time-optimal control problems of distributed second order parabolic systems with finite number of variables involving time lags appearing in the boundary condition have been widely discussed in many papers and monographs. a fundamental study of such problems is given by (wang, 1975) and was next developed by (knowles, 1978) and (wong, 1987). it was also intensively investigated by (kowalewski, 1988; 1990a; 1990b; 1993; 1998; 1999; 2009), (kowalewski and duda, 1992 ), (kowalewski and krakowiak, 1994; 2000; 2006; 2008), (kotarski, 1997), (kotarski & el-saify and bahaa, 2002b), (kotarski and bahaa, 2007) and (elsaify, 2005; 2006) in which linear quadratic problem for parabolic systems with time delays given in the different form (constant time delays, time-varying delays, time delays given in the integral form, etc.) were presented. the necessary and sufficient conditions of optimality for systems consists of only one equation and for (n × n) systems governed by different types of partial differential equations defined on spaces of functions of infinitely many variables and also for infinite order systems are discussed for example in ( gali, i. m. & el-saify, h. a. 1982; 1983), (el-saify & bahaa, 2001; 2003), (el-saify, h. a., serag, h. m, & bahaa, g. m. 2000), (el-saify, 2005; 2006), (kowalewski, 2009) and (kowalewski and krakowiak, 2008) in which the argument of (lions, 1971 and lions & magenes, 1972) were used. variables time lag.variables time lag issn 1078-6236 international institute for general systems studies, inc. 2 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . making use of the dubovitskii-milyutin theorem in (kotarski, el-saify & bahaa, 2002a,b), (bahaa, 2003; 2005a,b; 2008) and (bahaa and kotarski, 2008), the necessary and sufficient conditions of optimality for similar systems governed by second order operator with an infinite number of variables and also for infinite order systems were investigated. the interest in the study of this class of operators is stimulated by problems in quantum field theory. in particular, the papers of (kowalewski & krakowiak, 2006, 2008), the time-optimal boundary control problem for a second order distributed parabolic systems with finite number of variables in which constant time lags appear in integral form in both the state equation and the boundary condition is presented. some particular properties of optimal control are discussed. in this paper we recall the problem in a more general formulation. we consider the timeoptimal distributed and boundary control problem for second order parabolic system and also for (n×n) –second order parabolic systems with infinite number of variables involving constant time lags appearing in integral form in both the state equation and in the boundary condition simultaneously. such an infinite variables parabolic systems can be treated as a generalization of the mathematical model for a plasma control process. the quadratic performance functional defined over a fixed time horizon are taken and some constraints are imposed on the boundary control. following a line of the lions scheme (lions, 1971) and (lions & magenes, 1972), necessary and sufficient optimality conditions for the neumann and dirichlet problem applied to the above systems were derived. the optimal control is characterized by the adjoint equations. this paper is organized as follows. in section 2, we introduce sobolev spaces with infinite number of variables. in section 3, we formulate the mixed neumann problem for infinite variables parabolic systems involving time lags. in section 4, the time-distributed control problem for this case is formulated, then we give the necessary and sufficient conditions for the time control to be an optimal. in section 5, we concluded and generalized our results. 2. sobolev spaces with infinite number of variables this section covers the basic notations, definitions and properties, which are necessary to present this work (berezanskii, 1975), ( gali & el-saify 1982; 1983), (el-saify & serag & bahaa, 2000) and (el-saify & bahaa, 2001). let (pk(t)) ∞ k=1 be a sequence of weights, fixed in all that follows, such that; 0 < pk(t) ∈ c∞(r1), ∫ r1 pk(t)dt = 1, with respect to it we introduce on the region r∞ = r1 × r1 × . . . , the measure dρ(x) by setting, dρ(x) = p1(x1)dx1 ⊗ p2(x2)dx2 ⊗ . . . , (r∞ 3 x = (xk) ∞ k=1, xk ∈ r1). on r∞ we construct the space l2(r∞, dρ(x)) with respect to this measure i.e., l2(r∞, dρ(x)) is the space of quadratic integrable functions on r∞. we shall often set l2(r∞, dρ(x)) = l2(r∞). advances in systems science and applications (2011), vol. 11, no. 1-2 3 it is classical result that l2(r∞) is a hilbert space for the scalar product (φ, ψ)l2(r∞) = ∫ r∞ φ(x)ψ(x)dρ(x). we next consider a sobolev space in the case of an unbounded region. for functions which are ` = 1, 2, . . . times continuously differentiable up to the boundary γ of r∞ ( γ is meant to be the boundary of the support of the measure dρ(x)) and which vanish in a neighborhood of ∞, we introduce the scalar product (φ, ψ)w `(r∞) = ∑ |α|≤` (dαφ,dαψ)l2(r∞), where dα is defined by dα = ∂|α| (∂x1)α1(∂x2)α2 · · · , |α| = ∞∑ i=1 αi, and the differentiation is taken in the sense of generalized functions on r∞, and after the completion, we obtain the sobolev space w `(r∞). so in short, sobolev space w 1(r∞) is defined by : w 1(r∞) = {φ|φ,dφ ∈ l2(r∞)}. as in the case of a bounded region, the space w 1(r∞) form the space with positive norm ||.||w 1(r∞). we can construct the space w−1(r∞) = (w 1(r∞))∗ with negative norm ||.||w−1(r∞) with respect to the space w 0(r∞) = l2(r∞) with zero norm ||.||l2(r∞), then we have the following equipped, w 1(r∞) ⊆ l2(r∞) ⊆w−1(r∞), ||φ||w 1(r∞) ≥ ||φ||l2(r∞) ≥ ||φ||w -1(r∞). letl2(0, t ;w 1(r∞)) be the space of square integrable measurable functions t→ φ(t) of ]0, t [→ w 1(r∞), where the variable t denotes the “ time ”; t ∈]0, t [, t <∞. this space is a hilbert space with respect to the scalar product (φ, ψ)l2(0,t ;w 1(r∞)) = ∫ t 0 (φ(t), ψ(t))w 1(r∞)dt, and its dual is the spacel2(0, t ;w−1(r∞)), analogously, we can define the spacesl2(0, t ;l2(r∞)) which we shall denote by l2(q). let ω ⊂ r∞ is a bounded, open set with boundary γ, which is ac∞ manifold of dimension (n − 1). locally, ω is totally on one side of γ and denote by w 1(ω,r∞, dρ(x)) (briefly w 1(ω,r∞)) the sobolev space of vector function y(x) defined on ω. the construction of the cartesian product of n-times to the above hilbert spaces can be construct, for example (w 1(ω,r∞))n = w 1(ω,r∞)×w 1(ω,r∞)× · · · ×w 1(ω,r∞)︸ ︷︷ ︸ n−times = n∏ i=1 (w 1(ω,r∞))i, 4 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . with norm defined by: ||φ||(w 1(ω,r∞))n = n∑ i=1 ||φi||w 1(ω,r∞), where φ = (φ1, φ2, ..., φn) = (φi) n i=1 is a vector function and φi ∈w 1(ω,r∞). finally, we have the following chain: (l2(0, t ;w 1(ω,r∞)))n ⊆ (l2(q))n ⊆ (l2(0, t ;w−1(ω,r∞)))n, where (l2(0, t ;w−1(ω,r∞)))n are the dual spaces of (l2(0, t ;w 1(ω,r∞)))n. the spaces considered in this paper are assumed to be real. 3. existence and uniqueness of solutions consider now the distributed-parameter system described by the following parabolic delay equation: ∂y ∂t +a(t)y + ∫ b a c(x, t)y(x, t− h) dh = u, x ∈ ω, t ∈ (0, t ), h ∈ (a, b), (1) y(x, t′) = φ0(x, t′), x ∈ ω, t′ ∈ [−b, 0), (2) y(x, 0) = y0(x), x ∈ ω, (3) ∂y(x, t) ∂ηa = ∫ b a d(x, t)y(x, t− h) dh+ v, x ∈ γ, t ∈ (0, t ), h ∈ (a, b), (4) y(x, t′) = ψ0(x, t′), x ∈ γ, t′ ∈ [−b, 0), (5) where ω and γ have the same properties as in section 2. we have y ≡ y(x, t;u), u ≡ u(x, t), v ≡ v(x, t), q ≡ ω× (0, t ), q ≡ ω× [0, t ], q0 ≡ ω× [−b, 0) σ ≡ γ× (0, t ), σ0 ≡ γ× [−b, 0), t is a specified positive number representing a time horizon, c is a given real c∞ function defined on q, d is a given real c∞ function defined on σ, h is a time lag such that h ∈ (a, b) and a > 0, φ0 and ψ0 are initial functions defined on q0 and σ0, respectively. the parabolic operator ∂ ∂t + a(t) in the state equation (1) is a second order parabolic operator with infinite number of variables anda(t) (berezanskii, 1975), (gali & el-saify, 1982; 1983) and (kotarski & el-saify & bahaa, 2002b ) is given by: a(t)y(x) = ( − ∞∑ k=1 1√ pk(xk, t) ∂2 ∂x2 k √ pk(xk, t) + q(x, t) ) y(x) = − ∞∑ k=1 d2 ky(x) + q(x, t)y(x), (6) advances in systems science and applications (2011), vol. 11, no. 1-2 5 where dky(x) = 1√ pk(xk, t) ∂ ∂xk √ pk(xk, t)y(x), (7) and q(x, t) is a real-valued function in x which is a bounded and measurable on ω ⊂ r∞, such that q(x, t) ≥ ξ0 > 1, ξ0 is a constant. the operatora(t) is a bounded second order self-adjoint elliptic partial differential operator with an infinite number of variables maps w 1(ω,r∞) onto w−1(ω,r∞). for this operator we define the bilinear form as follows: definition 3.1. for each t ∈ (0, t ), we define a family of bilinear forms on w 1(ω,r∞) by: π(t; y, φ) = (a(t)y, φ)l2(ω,r∞), y, φ ∈w 1(ω,r∞), (8) where a(t) maps w 1(ω,r∞) onto w−1(ω,r∞) and takes the above form. then π(t; y, φ) = ( a(t)y, φ ) l2(ω,r∞) = ( − ∞∑ k=1 d2 ky(x) + q(x, t)y(x), φ(x) ) l2(ω,r∞) = ∫ ω ∞∑ k=1 dky(x)dkφ(x) dρ(x) + ∫ ω q(x, t)y(x)φ(x) dρ(x). lemma 3.1. the bilinear form π(t; y, φ) is coercive on w 1(ω,r∞), that is π(t; y, y) ≥ λ ||y||2w 1(ω,r∞), λ > 0. (9) proof. it is well known that the ellipticity ofa(t) is sufficient for the coerciveness of π(t; y, φ) on w 1(ω,r∞). π(t;φ, ψ) = ∫ ω ∞∑ k=1 dkφ(x)dkψ(x) dρ+ ∫ ω q(x, t)φ(x)ψ(x) dρ. then π(t; y, y) = ∫ ω ∞∑ k=1 |dky(x)|2 dρ(x) + ∫ ω q(x, t)|y(x)|2 dρ(x) ≥ ∞∑ k=1 ||dky(x)||2l2(ω,r∞) + ξ0||y(x)||2l2(ω,r∞) = ||y(x)||2w 1(ω,r∞) + ξ0||y(x)||2l2(ω,r∞) ≥ ||y(x)||2w 1(ω,r∞) = λ||y||2w 1(ω,r∞), λ > 0. 6 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . also we have: ∀y, φ ∈w 1(ω,r∞) the function t→ π(t; y, φ) is continuously differentiable in (0, t ) and π(t; y, φ) = π(t;φ, y) } (10) equations (1)–(5) constitute a neumann problem. then the left-hand side of the boundary condition (4) may be written in the following form: ∂y(u) ∂ηa = ∞∑ k=1 (dky(u)) cos(n, xk) = g(x, t), (11) where ∂ ∂ηa is a normal derivative at γ, directed towards the exterior of ω, cos(n, xk) is the k − th direction cosine of n, with n being the normal at γ exterior to ω, and g(x, t) = ∫ b a d(x, t)y(x, t− h) dh+ v(x, t), x ∈ γ, t ∈ (0, t ), h ∈ (a, b). (12) first we shall prove sufficient conditions for the existence of a unique solution of the mixed initial boundary value problem (1)–(5) for the cases where the control u or v belong to l2(q) or l2(σ) respectively. to this purpose, for any pair of real numbers r, s ≥ 0, we introduction the sobolev space w r,s(q) (lions and magenes, 1972, vol. 2, p. 6) defined by w r,s(q) = l2 (0, t ;w r(ω,r∞)) ∩w s ( 0, t ;l2(ω,r∞) ) (13) which is a hilbert space normed by(∫ t 0 ||y(t)||2w r(ω,r∞)dt+ ||y||2w s(0,t ;l2(ω,r∞)) )1/2 , (14) where w s ( 0, t ;l2(ω,r∞) ) denotes the sobolev space of order s of functions defined on (0, t ) and taking values in l2(ω,r∞). the existence of a unique solution for the mixed initial-boundary value problem (1)–(5) on the cylinder q can be proved using a constructive method, i.e., first, solving (1)–(5) on the sub-cylinder q1 and in turn on q2, and so on, until the procedure covers the whole cylinder q. in this way, the solution in the previous step determines the next one. for simplicity, we introduce the following notation: ej , ((j − 1)a, ja), qj = ω× ej , σj = γ× ej , j = 1, 2, . . . . (15) case 1: u ∈ l2(q) using theorem 6.1 of lions & magenes (1972, vol. 2, p. 33), we can prove the following lemma. advances in systems science and applications (2011), vol. 11, no. 1-2 7 lemma 3.2. let u ∈ l2(q), (16) fj(x, t) ∈ l2(qj), (17) where fj = u(x, t)− ∫ b a c(x, t)yj−1(x, t− h) dh, yj−1(·, (j − 1)a) ∈w 1(ω,r∞), (18) gj ∈w 1 2 , 1 4 (σj), (19) where gj(x, t) = ∫ b a d(x, t)yj−1(x, t− h) dh+ v(x, t). then, there exists a unique solution yj ∈ w 2,1(qj) for the mixed initial-boundary value problem (1), (4) and (18). proof. we observe that for j = 1, y0|q0(x, t−h) = φ0(x, t−h) and y0|σ0(x, t−h) = ψ0(x, t− h). then the assumptions (17)–(19) are fulfilled if we assume that φ0 ∈ w 2,1(q0), y0 ∈ w 1(ω,r∞), v ∈ w 1 2 , 1 4 (σ) and ψ0 ∈ w 1 2 , 1 4 (σ0). these assumptions are sufficient to ensure the existence of a unique solution y1 ∈ w 2,1(q1). in order to extend the result to q2, we have to prove that y1(·, a) ∈w 1(ω,r∞), g2 ∈w 1 2 , 1 4 (σ2) and f2 ∈ l2(q2). really, from theorem 3.1, p.19 of lions & magenes vol.1, y1 ∈ w 2,1(q1) implies that the mapping t → y1(·, t) is continuous from [0, a] → w 1(ω,r∞). thus y1(·, a) ∈ w 1(ω,r∞). then using the trace theorem of lions & magenes (1972, vol. 2, p. 9) we can verify that y1 ∈ w 2,1((q1) implies that y1 → y1|σ1 is a linear, continuous mapping of w 2,1(q1)→ w 1 2 , 1 4 (σ1). assuming that d is a c∞ function and v ∈ w 1 2 , 1 4 (σ), the condition g2 ∈ w 1 2 , 1 4 (σ2) is fulfilled. also it is easy to notice that the assumption (17) follows from the fact that y1 ∈ w 2,1(q1) and u ∈ l2(q). then, there exists a unique solution y2 ∈ w 2,1(q2). finally, we can extend our result to any qj , j = 3, 4 . . .. theorem 3.3. let y0, φ0, ψ0, v and u be given with y0 ∈ w 1(ω,r∞), φ0 ∈ w 2,1(q0), v ∈ w 1 2 , 1 4 (σ), ψ0 ∈w 1 2 , 1 4 (σ0) and u ∈ l2(q). then, there exists a unique solution y ∈w 2,1(q) for the mixed initial-boundary value problem (1)–(5). moreover, y(·, ja) ∈w 1(ω,r∞) for j = 1, 2, . . .. case 2: v ∈ l2(σ) using theorem 15.2 of lions & magenes (1972, vol. 2, p. 81), we can prove the following lemma. lemma 3.4. let u ∈w− 1 2 ,− 1 4 (q), v ∈ l2(σ) (20) fj ∈w− 1 2 ,− 1 4 (qj), (21) yj−1(·, (j − 1)a) ∈w 1 2 (ω,r∞), (22) 8 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . gj ∈ l2(σj). (23) then, there exists a unique solution yj ∈ w 3 2 , 3 4 (qj) for the mixed initial-boundary value problem (1), (4) and (22). proof. for j = 1, the assumptions (21)–(23) are fulfilled if we assume that φ0 ∈w 3 2 , 3 4 (q0), y0 ∈ w 1 2 (ω,r∞) and ψ0 ∈ l2(σ0). these assumptions are sufficient to ensure the existence of a unique solution y1 ∈ w 3 2 , 3 4 (q1). in order to extend the result to q2, we have to prove that y1(·, a) ∈ w 1 2 (ω,r∞), y1|σ1 ∈ l2(σ1) and f2 ∈ w− 1 2 ,− 1 4 (q2). first using theorem 3.1 of lions & magenes (1972, vol. 1, p. 19) we can prove that y1 ∈ w 3 2 , 3 4 (q1) implies that the mapping t → y1(·, t) is continuous from [0, a] → w 3 4 (ω,r∞) ⊂ w 1 2 (ω,r∞). hence y1(·, a) ∈ w 1 2 (ω,r∞). again, from trace theorem of lions & magenes (1972, vol. 2, p. 9), we can verify that y1 ∈ w 3 2 , 3 4 (q1) implies that y1 → y1|σ1 is a linear, continuous mapping of w 3 2 , 3 4 (q1) → w 1, 1 2 (σ1). thus y1|σ1 ∈ l2(σ1). moreover, it is worth mentioning that the assumption (21) follows from the fact that y1 ∈ w 3 2 , 3 4 (q1) and u ∈ w− 1 2 ,− 1 4 (q). then, there exists a unique solution y2 ∈ w 3 2 , 3 4 (q2). finally, we can extend our result to any qj , j = 3, 4 . . .. theorem 3.5. let y0, φ0, ψ0, v and u be given with y0 ∈w 1 2 (ω,r∞), φ0 ∈w 3 2 , 3 4 (q0), ψ0 ∈ l2(σ0), v ∈ l2(σ) and u ∈ w− 1 2 ,− 1 4 (q). then, there exists a unique solution y ∈ w 3 2 , 3 4 (q) for the mixed initial-boundary value problem (1)–(5). moreover, y(·, ja) ∈ w 1 2 (ω,r∞) for j = 1, 2, . . .. now we shall verify the existence of a unique solution for the problem (1), (2), (3) and (5) with the dirichlet boundary condition involving a time lag y(x, t) = g(x, t) (24) where g is given by the formula (12). making use of the results of lions & magenes (1972, vol. 2, p. 33 and p. 81) we can prove the following lemmas and theorems. case 3: u ∈ l2(q) lemma 3.6. let u ∈ l2(q), (25) fj ∈ l2(qj), (26) yj−1(·, (j − 1)a) ∈w 1(ω,r∞), (27) gj ∈w 3 2 , 3 4 (σj), (28) and the following compatibility relation is fulfilled yj−1(x, (j − 1)a) = gj(x, (j − 1)a), on γ. (29) then, there exists a unique solution yj ∈ w 2,1(qj) for the mixed initial-boundary value problem (1), (24) and (27). advances in systems science and applications (2011), vol. 11, no. 1-2 9 proof. for j = 1, the assumptions (26)–(28) can be satisfied if we assume that φ0 ∈ w 2,1(q0), v ∈ w 3 2 , 3 4 (σ) and ψ0 ∈ w 3 2 , 3 4 (σ0). these assumptions are sufficient to ensure the existence of a unique solution y1 ∈ w 2,1(q1) if y0 ∈ w 1(ω,r∞) and the following compatibility relation is satisfied y0(x, 0) = g1(x, 0), on γ. (30) in order to extend the result to q2, we have to prove that y1 ∈ w 2,1(q1) and it is necessary to impose the compatibility relation y1(x, a) = g2(x, a), on γ (31) and it is sufficient to verify that f2 ∈ l2(q2), (32) y1(·, a) ∈w 1(ω,r∞), (33) g2 ∈w 3 2 , 3 4 (σ2). (34) first using the solution in the previous step and the condition (25) we can prove immediately the condition (32). to verify (33), we use the fact that y1 ∈ w 2,1(q1) implies that the mapping t → y1(·, t) is continuous from [0, a] → w 1(ω,r∞) (by theorem 3.1 of lions & magenes (1972, vol. 1, p. 19)), hence y1(·, a) ∈w 1(ω,r∞). from the trace theorem of lions & magenes (1972, vol. 2, p. 9) y1 ∈w 2,1((q1) implies that y1 → y1|σ1 is a linear, continuous mapping of w 2,1(q1) → w 3 2 , 3 4 (σ1). assuming that d is a c∞ function and v ∈ w 3 2 , 3 4 (σ), the condition (34) is fulfilled. then, there exists a unique solution y2 ∈ w 2,1(q2). finally, we can extend our result to any qj , j = 3, 4 . . .. theorem 3.7. let y0, φ0, ψ0, v and u be given with y0 ∈ w 1(ω,r∞), φ0 ∈ w 2,1(q0), v ∈ w 3 2 , 3 4 (σ), ψ0 ∈ w 3 2 , 3 4 (σ0), u ∈ l2(q) and the compatibility relation (29) is fulfilled. then, there exists a unique solution y ∈w 2,1(q) for the mixed initial-boundary value problem (1), (2), (3) (5) and (24) with y(·, ja) ∈w 1(ω,r∞) for j = 1, 2, . . .. case 4: v ∈ l2(σ) lemma 3.8. let u ∈w− 3 2 ,− 3 4 (q), v ∈ l2(σ) (35) fj ∈w− 3 2 ,− 3 4 (qj), (36) yj−1(·, (j − 1)a) ∈w− 1 2 (ω,r∞), (37) gj ∈ l2(σj). (38) then, there exists a unique solution yj ∈ w 1 2 , 1 4 (qj) for the mixed initial-boundary value problem (1), (24) and (37). 10 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . proof. we observe that for j = 1, the assumptions (36)–(38) are satisfied if we assume that φ0 ∈ w 1 2 , 1 4 (q0), y0 ∈ w− 1 2 (ω,r∞) and ψ0 ∈ l2(σ0). these assumptions are sufficient to ensure the existence of a unique solution y1 ∈ w 1 2 , 1 4 (q1). next for j = 2, using the solution in the first step, it is sufficient to verify that f2 ∈ w− 3 2 ,− 3 4 (q2), y1(·, a) ∈ w− 1 2 (ω,r∞) and y1|σ1 ∈ l2(σ1). then it is worth mentioning that the condition f2 ∈ w− 3 2 ,− 3 4 (q2) follows from the fact that y1 ∈ w 1 2 , 1 4 (q1) and u ∈ w− 3 2 ,− 3 4 (q). since y1 ∈ w 1 2 , 1 4 (q1) implies that the mapping t → y1(·, t) is continuous from [0, a] → w 1 4 (ω,r∞) (by theorem 3.1 of lions & magenes (1972, vol. 1, p. 19)), hence y1(·, a) ∈ w 1 4 (ω,r∞) ⊂ l2(ω,r∞) ⊂ w− 1 4 (ω,r∞) ⊂ w− 1 2 (ω,r∞). we shall prove that y1|σ1 ∈ l2(σ1). we must notice that for proving y1|σ1 ∈ l2(σ1) we cannot use the trace theorem of lions & magenes (1972, vol. 2, p. 9), since y1 ∈w 1 2 , 1 4 (q1). it is worth mentioning that this difficulty can be avoided by using the condition y1|σ1 = g1 ∈ l2(σ1). this implies that y1|σ1 ∈ l2(σ1). then, there exists a unique solution y2 ∈w 1 2 , 1 4 (q2). finally, we can extend our result to any qj , j = 3, 4 . . .. theorem 3.9. let y0, φ0, ψ0, v and u be given with y0 ∈ w− 1 2 (ω,r∞), φ0 ∈ w 1 2 , 1 4 (q0), ψ0 ∈ l2(σ0), v ∈ l2(σ) and u ∈ w− 3 2 ,− 3 4 (q). then, there exists a unique solution y ∈ w 1 2 , 1 4 (q) for the mixed initial-boundary value problem (1), (2), (3), (5) and (24). moreover, y(·, ja) ∈w− 1 2 (ω,r∞) for j = 1, 2, . . .. 4. optimal distributed control now, we shall restrict our considerations to the case of the distribute control for the neumann problem. therefore, we shall formulate the minimum-time problem for (1)–(5) in the context of the theorem 3.3, i.e., u ∈ u = {u ∈ l2(q) : |u(x, t)| ≤ 1}. (39) we shall define the reachable set h such that h = {y ∈ l2(ω,r∞) : ||y − zd||l2(ω,r∞) ≤ ε} (40) where zd ∈ l2(ω,r∞) and ε > 0. solving the stated minimum-time problem is equivalent to hitting the target set h in minimum time, that is, minimizing the time t, for which y(t;u) ∈ h and u ∈ u . moreover, we assume that there exists a t > 0 and u ∈ u with y(t ;u) ∈ h (41) then we have the following theorem theorem 4.1. if the assumption (41) holds, then the seth is reached in minimum time t∗ by an admissible control u∗ ∈ u . moreover∫ ω [zd − y(t∗;u∗)] [y(t∗;u)− y(t∗;u∗)] dρ ≤ 0, ∀u ∈ u. (42) advances in systems science and applications (2011), vol. 11, no. 1-2 11 proof. let us define the following set t∗ := inf{t : y(t;u) ∈ h for some u ∈ u} (43) the minimum is well defined, as (41) guarantees that this set is nonempty. by definition, we can choose tn ↓ t∗ and admissible controls {un} such that y(tn;un) ∈ h, n = 1, 2, 3, . . . . (44) each un is defined on ω × (0, tn) ⊃ ω × (0, t∗). to simplify the notation, we denote the restriction of un to ω× (0, t∗) again by un. the set of admissible controls then forms a weakly compact, convex set in l2(ω×(0, t∗)), and so we can extract a weakly convergent subset {um}, which converges weakly to some admissible control u∗. consequently, theorem 3.3 implies that y(t;u) ∈ w 1(ω,r∞) ⊂ l2(ω,r∞) for each u ∈ l2(q) and t > 0. then using theorem 1.2 of (lions, 1971, p. 102) and theorem 3.3 it is easy to verify that the mapping u → y(t∗;u) from l2(ω × (0, t∗)) into l2(ω,r∞), is continuous. since any continuous linear mapping between banach spaces is also weakly continuous (dunford and schwartz, 1958), theorem v. 3.15, the affine mapping u → y(t∗;u) must also be weakly continuous. hence, y(t∗;um)→ y(t∗;u∗) weakly in l2(ω,r∞). (45) moreover, dy(u) dt ∈ l2 ( [0, t∗];l2(ω,r∞) ) , (46) for each u ∈ u , by definition of w 2,1(ω× (0, t∗)) and ||y(tm;um)− y(t∗;um)||l2(ω,r∞) = ∣∣∣∣∣∣∣∣∫ tm t∗ ẏ(σ;um) dσ ∣∣∣∣∣∣∣∣ l2(ω,r∞) (47) ≤ √ tm − t∗ (∫ tm t∗ ||ẏ(σ;um)||2l2(ω,r∞) dσ )1/2 .(48) applying theorem 1.2 of (lions, 1971) and theorem 3.3 again, the set {ẏ(um)} must be bounded in l2(0, t∗;l2(ω,r∞)), and so ||y(tm;um)− y(t∗;um)||l2(ω,r∞) ≤m √ tm − t∗. (49) combining (45) and (49) shows that y(tm;um)− y(t∗;u∗) = (y(tm;um)− y(t∗;um)) + (y(t∗;um)− y(t∗;u∗)), (50) converges weakly to zero in l2(ω,r∞), and therefore y(t∗;u∗) ∈ h as h is closed and convex, hence weakly closed. this shows that h is reached in time t∗ by an admissible control accordingly, t∗ must be the minimum time and u∗ an optimal control. we shall now prove the second part of our theorem. indeed, from theorem 3.1 (lions and magenes, 1972, vol. 1, p. 19) y(u) ∈ w 2,1(q) implies that the mapping t → y(t;u) is 12 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . continuous from [0, t ] → w 1(ω,r∞) ⊂ l2(ω,r∞) is continuous for each fixed u, and so y(t∗;u) 6∈ inth , for any u ∈ u , by the minimality of t∗. from our earlier remarks, the set a(t∗) = {y(t∗;ux) : ux ∈ u}, (51) is weakly compact and convex in l2(ω,r∞). applying theorem 21.11 of (choquet, 1969) to the setsa(t∗) and h shows that there exists a nontrivial hyperplane z ∈ l2(ω,r∞) separating these sets, that is, ∫ ω zy(t∗;u) dρ ≤ ∫ ω zy(t∗;u∗) dρ ≤ ∫ ω zy dρ (52) for all u ∈ u and y ∈ l2(ω,r∞) with ||y − zd||l2(ω,r∞) ≤ ε. from the second inequality in (52), z must support the seth at y(t∗;u∗). sincel2(ω,r∞) is a hilbert space, z must be of the form z = µ(zd − y(t∗;u∗)) for some µ > 0. (53) subsequently, dividing (52) by µ gives the desired result (42). we shall apply theorem 4.1 to the control problem of (1)–(5). to simplify (42), we introduce the adjoint equation, and for every u ∈ u we define the adjoint variable p = p(u) = p(x, t;u) as the solution of the following system −∂p(u) ∂t +a∗(t)p(u) + ∫ b a c(x, t+h)p(x, t+h;u) dh = 0, x ∈ ω, t ∈ (0, t∗− b), (54) −∂p(u) ∂t +a∗(t)p(u) + ∫ t∗−t a c(x, t+ h)p(x, t+ h;u) dh = 0, x ∈ ω, t ∈ (t∗− b, t∗− a), (55) −∂p(u) ∂t +a∗(t)p(u) = 0, x ∈ ω, t ∈ (t∗ − a, t∗), (56) p(x, t∗;u) = zd(x)− y(x, t∗;u), x ∈ ω, (57) ∂p(u) ∂ηa∗ (x, t) = ∫ b a d(x, t+ h)p(x, t+ h;u) dh, x ∈ γ, t ∈ (0, t∗ − b), (58) ∂p(u) ∂ηa∗ (x, t) = ∫ t∗−t a d(x, t+ h)p(x, t+ h;u) dh, x ∈ γ, t ∈ (t∗ − b, t∗ − a), (59) ∂p(u) ∂ηa∗ (x, t) = 0, x ∈ γ, t ∈ (t∗ − a, t∗), (60) where ∂p(u) ∂ηa∗ (x, t) = ∞∑ k=1 (dkp(u)) cos(n, xk), (61) a∗(t)p(u) = ( − ∞∑ k=1 d2 k + q(x, t) ) p(u). (62) advances in systems science and applications (2011), vol. 11, no. 1-2 13 remark 4.2. if t∗ < b, then we consider (55) and (59) on ω×(0, t∗−a) and γ×(0, t∗−a), respectively. the existence of a unique solution to the problem (54)–(60) on the cylinder ω× (0, t∗) can be proved using a constructive method. it is easy to notice that for given zd and u, the problem (54)–(60) can be solved backwards in time starting from t = t∗, i.e., first, solving (54)–(60) on the sub-cylinder qk and in turn onqk−1, and so on, until the procedure covers the whole cylinder ω × (0, t∗). for this purpose, we may apply theorem 3.3 (with an obvious change of variables). hence, using theorem 3.3, the following result can be proved. theorem 4.3. let the hypothesis of theorem 3.3 be satisfied. then for given zd ∈ l2(ω,r∞) and any u ∈ l2(q), there exists a unique solution p(u) ∈ w 2,1(ω × (0, t∗)) for the adjoint problem (54)–(60). now, we have the main result. theorem 4.4. if the assumptions concerning system (1)–(5) and controllability condition (41) are satisfied, then the time-optimal control u∗ exists and is characterized by the following condition ∫ t∗ 0 ∫ ω p(u∗)(u− u∗) dρ dt ≤ 0, ∀u ∈ u, (63) where p(u∗) is the solution of the adjoint system (54)–(60). proof. we simplify the left-hand side of the inequality (42) using the adjoint equation (54)– (60). for this purpose, setting u = u∗ in (54)–(60), multiplying both sides of (54), (55) and (56) by y(u) − y(u∗), then integrating over ω × (0, t∗ − b), ω × (t∗ − b, t∗ − a) and ω× (t∗− a, t∗) respectively and then adding both sides of (54), (55) and (56), we get∫ t∗ 0 ∫ ω ( −∂p(u ∗) ∂t +a∗(t)p(u∗) ) (y(u)− y(u∗)) dρ dt + ∫ t∗−b 0 ∫ ω (∫ b a c(x, t+ h)p(x, t+ h;u∗) dh ) × (y(x, t;u)− y(x, t;u∗)) dρ dt + ∫ t∗−a t∗−b ∫ ω (∫ t∗−t a c(x, t+ h)p(x, t+ h;u∗) dh ) (y(x, t;u)− y(x, t;u∗)) dρ dt =− ∫ ω p(x, t∗;u∗)(y(x, t∗;u)− y(x, t∗;u∗)) dρ+ ∫ t∗ 0 ∫ ω p(u∗) ∂ ∂t (y(u)− y(u∗)) dρ dt + ∫ t∗ 0 ∫ ω a∗p(u∗)(y(u)− y(u∗)) dρ dt + ∫ t∗−b 0 ∫ ω ∫ b a c(x, t+ h)p(x, t+ h;u∗) (y(x, t;u)− y(x, t;u∗)) dh dρ dt + ∫ t∗−a t∗−b ∫ ω ∫ t∗−t a c(x, t+ h)p(x, t+ h;u∗) (y(x, t;u)− y(x, t;u∗)) dh dρ dt = 0. (64) 14 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . then, applying (57), the equation (64) can be expressed as ∫ ω (zd − y(t∗;u∗))(y(x, t∗;u)− y(x, t∗;u∗)) dρ = ∫ t∗ 0 ∫ ω p(u∗) ∂ ∂t (y(u)− y(u∗)) dρ dt+ ∫ t∗ 0 ∫ ω a∗p(u∗)(y(u)− y(u∗)) dρ dt + ∫ b a ∫ ω ∫ t∗−b 0 c(x, t+ h)p(x, t+ h;u∗) (y(x, t;u)− y(x, t;u∗)) dt dρ dh + ∫ t∗−t a ∫ ω ∫ t∗−a t∗−b c(x, t+ h)p(x, t+ h;u∗) (y(x, t;u)− y(x, t;u∗)) dt dρ dh. (65) using (1), the first integral on the right-hand side of (65) can be rewritten as ∫ t∗ 0 ∫ ω p(u∗) ∂ ∂t (y(u)− y(u∗)) dρ dt =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ t∗ 0 ∫ ω p(x, t;u∗) (∫ b a c(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dh ) dρ dt + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt advances in systems science and applications (2011), vol. 11, no. 1-2 15 =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ t∗ 0 ∫ ω ∫ b a p(x, t;u∗)c(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dh dρ dt + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ b a ∫ ω ∫ t∗ 0 p(x, t;u∗)c(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dt dρ dh + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ b a ∫ ω ∫ t∗−h −h p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ b a ∫ ω ∫ 0 −h p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh − ∫ b a ∫ ω ∫ t∗−b 0 p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh − ∫ b a ∫ ω ∫ t∗−h t∗−b p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ b a ∫ ω ∫ 0 −h p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh − ∫ b a ∫ ω ∫ t∗−b 0 p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh − ∫ t∗−t a ∫ ω ∫ t∗−a t∗−b p(x, t′ + h;u∗)c(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dρ dh + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt. (66) the second integral on the right-hand side of (65), in view of green formula, can be 16 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . expressed as ∫ t∗ 0 ∫ ω a∗p(u∗)(y(u)− y(u∗)) dρ dt = ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt + ∫ t∗ 0 ∫ γ p(u∗) ( ∂y(u) ∂ηa − ∂y(u∗) ∂ηa ) dγ dt− ∫ t∗ 0 ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt. (67) using the boundary condition (4), the second component on the right-hand side of (67) can be written as ∫ t∗ 0 ∫ γ p(u∗) ( ∂y(u) ∂ηa − ∂y(u∗) ∂ηa ) dγ dt = ∫ t∗ 0 ∫ γ p(x, t;u∗) (∫ b a d(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dh ) dγ dt = ∫ t∗ 0 ∫ γ ∫ b a p(x, t;u∗)d(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dh dγ dt = ∫ b a ∫ γ ∫ t∗ 0 p(x, t;u∗)d(x, t)(y(x, t− h;u)− y(x, t− h;u∗)) dt dγ dh = ∫ b a ∫ γ ∫ t∗−h −h p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh = ∫ b a ∫ γ ∫ 0 −h p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh + ∫ b a ∫ γ ∫ t∗−b 0 p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh + ∫ b a ∫ γ ∫ t∗−h t∗−b p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh = ∫ b a ∫ γ ∫ 0 −h p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh + ∫ b a ∫ γ ∫ t∗−b 0 p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh + ∫ t∗−t a ∫ γ ∫ t∗−a t∗−b p(x, t′ + h;u∗)d(x, t′ + h)(y(x, t′;u)− y(x, t′;u∗)) dt′ dγ dh. (68) the last component in (67) can be rewritten as ∫ t∗ 0 ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt = ∫ t∗−b 0 ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt + ∫ t∗−a t∗−b ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt+ ∫ t∗ t∗−a ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt. (69) advances in systems science and applications (2011), vol. 11, no. 1-2 17 substituting (68) and (69) into (67) and then (66) and (67) into (65), we obtain∫ ω (zd − y(t∗;u∗))(y(x, t∗;u)− y(x, t∗;u∗)) dρ =− ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt − ∫ b a ∫ ω ∫ 0 −h c(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dρ dh − ∫ b a ∫ ω ∫ t∗−b 0 c(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dρ dh − ∫ t∗−t a ∫ ω ∫ t∗−a t∗−b c(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dρ dh + ∫ t∗ 0 ∫ ω p(u∗)a(y(u)− y(u∗)) dρ dt + ∫ b a ∫ γ ∫ 0 −h d(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dγ dh + ∫ b a ∫ γ ∫ t∗−b 0 d(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dγ dh + ∫ t∗−t a ∫ γ ∫ t∗−a t∗−b d(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dγ dh + ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt − ∫ t∗−b 0 ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt − ∫ t∗−a t∗−b ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt− ∫ t∗ t∗−a ∫ γ ∂p(u∗) ∂ηa∗ (y(u)− y(u∗)) dγ dt + ∫ b a ∫ ω ∫ t∗−b 0 c(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dρ dh + ∫ t∗−t a ∫ ω ∫ t∗−a t∗−b c(x, t+ h)p(x, t+ h;u∗)(y(x, t;u)− y(x, t;u∗)) dt dρ dh. (70) then, using the fact that y(x, t;u) = y(x, t;u∗) = φ0(x, t) for x ∈ ω and t ∈ [−b, 0), and y(x, t;u) = y(x, t;u∗) = ψ0(x, t) for x ∈ γ and t ∈ [−b, 0), we obtain∫ ω (zd − y(t∗;u∗))(y(x, t∗;u)− y(x, t∗;u∗)) dρ = ∫ t∗ 0 ∫ ω p(x, t;u∗)(u− u∗) dρ dt, (71) then, substituting (71) into (42), we get (63) and this finishes proof of the theorem. 5. generalization time-optimal control problem presented her can be extended to certain different two cases. case 1: time-optimal control problem for (2 × 2) coupled system of parabolic equations with infinite number of variables, in which time lags appear in integral form in both the state equation 18 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . and the boundary condition. case 2: time-optimal control problem for (n×n) coupled system of parabolic equations with infinite number of variables, in which time lags appear in integral form in both the state equation and the boundary condition. 5.1 time-optimal control problem for (2 × 2) coupled system of parabolic equations with infinite number of variables. we can extend the discussions to study the time-optimal control problem for 2×2 coupled system of parabolic equations with infinite number of variables, in which time lags appear in integral form in both the state equation and the boundary condition. consider now the distributed-parameter system described by the following (2× 2) coupled system of parabolic equations with infinite number of variables, for i = 1, 2, ∂yi ∂t + a(t)yi + ∫ b a ci(x, t)y(x, t− h) dh = ui, x ∈ ω, t ∈ (0, t ), h ∈ (a, b), (72) yi(x, t ′) = φi,0(x, t′), x ∈ ω, t′ ∈ [−b, 0), (73) yi(x, 0) = yi,0(x), x ∈ ω, (74) ∂yi(x, t) ∂ηa = ∫ b a di(x, t)yi(x, t− h) dh+ vi, x ∈ γ, t ∈ (0, t ), h ∈ (a, b), (75) yi(x, t ′) = ψi,0(x, t′), x ∈ γ, t′ ∈ [−b, 0), (76) where a(t)yi(x) = ( − ∞∑ k=1 d2 k + q(x, t) ) yi(x) + 2∑ j=1 aijyj(x) ∀ i = 1, 2, (77) aij = { 1, i ≥ j; −1, i < j. (78) it is easy to see that a(t) is (2× 2) matrix which takes the form a(t) =  − ∞∑ k=1 d2 k + q + 1 −1 1 − ∞∑ k=1 d2 k + q + 1  2×2 . (79) also we have yi ≡ yi(x, t;u), ui ≡ ui(x, t), vi ≡ vi(x, t), u ≡ (u1, u2), ci and di, i = 1, 2, are real c∞ functions defined on q and σ, respectively, φi,0 and ψi,0, i = 1, 2, are initial functions defined on q0 and σ0, respectively. now we discuss the case of the distributed control for the neumann problem. then, as section 3, for u = (u1, u2) ∈ (l2(q))2, we can obtain the following results. advances in systems science and applications (2011), vol. 11, no. 1-2 19 lemma 5.1. let u ∈ ( l2(q) )2 , (80) fj = (f1,j , f2,j) ∈ ( l2(qj) )2 , (81) where fi,j(x, t) = ui(x, t)− ∫ b a ci(x, t)yi,j−1(x, t− h) dh, i = 1, 2, yj−1(·, (j − 1)a) = (y1,j−1(·, (j − 1)a), y2,j−1(·, (j − 1)a)) ∈ ( w 1(ω,r∞) )2 , (82) gj = (g1,j , g2,j) ∈ ( w 1 2 , 1 4 (σj) )2 , (83) where gi,j(x, t) = ∫ b a di(x, t)yi,j−1(x, t− h) dh+ vi(x, t), i = 1, 2. then, there exists a unique solution yj ∈ ( w 2,1(qj) )2 for the mixed initial-boundary value problem (72), (75) and (82). theorem 5.2. let yi,0, φi,0, ψi,0, v and u be given with yi,0 ∈w 1(ω,r∞), φi,0 ∈w 2,1(q0), vi ∈ w 1 2 , 1 4 (σ), ψi,0 ∈ w 1 2 , 1 4 (σ0) and ui ∈ l2(q), i = 1, 2. then, there exists a unique solution y ∈ ( w 2,1(q) )2 for the mixed initial-boundary value problem (72)–(76). moreover, y(·, ja) ∈ ( w 1(ω,r∞) )2 for j = 1, 2, . . .. now, we shall formulate the minimum-time problem for (72)–(76) in the context of the theorem 5.2, i.e., u ∈ u = {u ∈ ( l2(q) )2 : |ui(x, t)| ≤ 1, i = 1, 2}. (84) we shall define the reachable set h such that h = {y ∈ ( l2(ω,r∞) )2 : 2∑ i=1 ||yi − zi,d||l2(ω,r∞) ≤ ε} (85) where zi,d ∈ l2(ω,r∞), i = 1, 2, and ε > 0. solving the stated minimum-time problem is equivalent to hitting the target set h in minimum time, that is, minimizing the time t, for which y(t;u) ∈ h and u ∈ u. moreover, we assume that there exists a t > 0 and u ∈ u with y(t ;u) ∈ h (86) then, as section 4, we can prove the following theorem theorem 5.3. if the assumption (86) holds, then the set h is reached in minimum time t∗ by an admissible control u∗ ∈ u. moreover 2∑ i=1 ∫ ω [zi,d − yi(t∗;u∗)] [yi(t ∗;u)− yi(t∗;u∗)] dρ ≤ 0, ∀u ∈ u. (87) 20 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . now, we introduce the adjoint equation, and for every u ∈ u we define the adjoint variable p = p(u) = p(x, t;u) as the solution of the following (2× 2) system −∂pi(u) ∂t +a∗(t)pi(u)+ ∫ b a ci(x, t+h)pi(x, t+h;u) dh = 0, x ∈ ω, t ∈ (0, t∗−b), (88) −∂pi(u) ∂t +a∗(t)pi(u)+ ∫ t∗−t a ci(x, t+h)pi(x, t+h;u) dh = 0, x ∈ ω, t ∈ (t∗−b, t∗−a), (89) −∂pi(u) ∂t + a∗(t)pi(u) = 0, x ∈ ω, t ∈ (t∗ − a, t∗), (90) pi(x, t ∗;u) = zi,d(x)− yi(x, t∗;u), x ∈ ω, (91) ∂pi(u) ∂ηa∗ (x, t) = ∫ b a di(x, t+ h)pi(x, t+ h;u) dh, x ∈ γ, t ∈ (0, t∗ − b), (92) ∂pi(u) ∂ηa∗ (x, t) = ∫ t∗−t a di(x, t+ h)pi(x, t+ h;u) dh, x ∈ γ, t ∈ (t∗ − b, t∗ − a), (93) ∂pi(u) ∂ηa∗ (x, t) = 0, x ∈ γ, t ∈ (t∗ − a, t∗), (94) where ∂pi(u) ∂ηa∗ (x, t) = ∞∑ k=1 (dkpi(u)) cos(n, xk), (95) a∗(t)pi(u) = ( − ∞∑ k=1 d2 k + q(x, t) ) pi(u) + 2∑ j=1 ajipj(u), (96) and aji is the transpose of aij . hence, as the proof of theorem 4.4 in section 4, we can prove the following theorem. theorem 5.4. if the assumptions concerning system (72)–(76) and controllability condition (86) are satisfied, then the time-optimal control u∗ = (u∗1, u ∗ 2) exists and is characterized by the following condition 2∑ i=1 ∫ t∗ 0 ∫ ω pi(u ∗)(ui − u∗i ) dρ dt ≤ 0, ∀u ∈ u, (97) where p(u∗) is the solution of the adjoint system (88)–(94). 5.2 time-optimal control problem for (n × n )coupled system of parabolic equations with infinite number of variables. we can extend the discussions to study the time-optimal control problem for n×n coupled system of parabolic equations with infinite number of variables, in which time lags appear in integral form in both the state equation and the boundary condition. advances in systems science and applications (2011), vol. 11, no. 1-2 21 consider now the distributed-parameter system described by the following (n×n) coupled system of parabolic equations with infinite number of variables, for i = 1, 2, . . . , n, ∂yi ∂t +a(t)yi + ∫ b a ci(x, t)y(x, t− h) dh = ui, x ∈ ω, t ∈ (0, t ), h ∈ (a, b), (98) yi(x, t ′) = φi,0(x, t′), x ∈ ω, t′ ∈ [−b, 0), (99) yi(x, 0) = yi,0(x), x ∈ ω, (100) ∂yi(x, t) ∂ηa = ∫ b a di(x, t)yi(x, t− h) dh+ vi, x ∈ γ, t ∈ (0, t ), h ∈ (a, b), (101) yi(x, t ′) = ψi,0(x, t′), x ∈ γ, t′ ∈ [−b, 0), (102) where a(t)yi(x) = ( − ∞∑ k=1 d2 k + q(x, t) ) yi(x) + n∑ j=1 aijyj(x) ∀ i = 1, 2, . . . , n, (103) aij = { 1, i ≥ j; −1, i < j. (104) it is easy to see thata(t) is (n× n) matrix which takes the form a(t) =  − ∞∑ k=1 d2 k + q + 1 −1 · · · −1 1 − ∞∑ k=1 d2 k + q + 1 · · · −1 ... ... ... ... 1 1 · · · − ∞∑ k=1 d2 k + q + 1  n×n . (105) also we have yi ≡ yi(x, t;u), ui ≡ ui(x, t), vi ≡ vi(x, t), u ≡ (u1, u2, . . . , un), ci and di, i = 1, 2, . . . , n, are real c∞ functions defined on q and σ, respectively, φi,0 and ψi,0, i = 1, 2, . . . , n, are initial functions defined on q0 and σ0, respectively. now we discuss the case of the distributed control for the neumann problem. then, as section 3, for u = (u1, u2, . . . , un) ∈ ( l2(q) )n, we can obtain the following results. lemma 5.5. let u ∈ ( l2(q) )n , (106) fj = (f1,j , f2,j , . . . , fn,j) ∈ ( l2(qj) )n , (107) where fi,j(x, t) = ui(x, t)− ∫ b a ci(x, t)yi,j−1(x, t− h) dh, i = 1, 2, . . . , n, 22 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . yj−1(·, (j − 1)a) = (y1,j−1(·, (j − 1)a), y2,j−1(·, (j − 1)a), . . . , yn,j−1(·, (j − 1)a)) ∈ ( w 1(ω,r∞) )n , (108) gj = (g1,j , g2,j , . . . , gn,j) ∈ ( w 1 2 , 1 4 (σj) )n , (109) where gi,j(x, t) = ∫ b a di(x, t)yi,j−1(x, t− h) dh+ vi(x, t), i = 1, 2, . . . n. then, there exists a unique solution yj ∈ ( w 2,1(qj) )n for the mixed initial-boundary value problem (98), (101) and (108). theorem 5.6. let yi,0, φi,0, ψi,0, v and u be given with yi,0 ∈w 1(ω,r∞), φi,0 ∈w 2,1(q0), vi ∈ w 1 2 , 1 4 (σ), ψi,0 ∈ w 1 2 , 1 4 (σ0) and ui ∈ l2(q), i = 1, 2, . . . , n. then, there exists a unique solution y ∈ ( w 2,1(q) )n for the mixed initial-boundary value problem (98)–(102). moreover, y(·, ja) ∈ ( w 1(ω,r∞) )n for j = 1, 2, . . .. now, we shall formulate the minimum-time problem for (98)–(102) in the context of the theorem 5.6, i.e., u ∈ u = {u ∈ ( l2(q) )n : |ui(x, t)| ≤ 1, i = 1, 2, . . . , n}. (110) we shall define the reachable seth such that h = {y ∈ ( l2(ω,r∞) )n : n∑ i=1 ||yi − zi,d||l2(ω,r∞) ≤ ε} (111) where zi,d ∈ l2(ω,r∞), i = 1, 2, . . . , n, and ε > 0. solving the stated minimum-time problem is equivalent to hitting the target seth in minimum time, that is, minimizing the time t, for which y(t;u) ∈ h and u ∈ u . moreover, we assume that there exists a t > 0 and u ∈ u with y(t ;u) ∈h (112) then, as section 4, we can prove the following theorem theorem 5.7. if the assumption (112) holds, then the seth is reached in minimum time t∗ by an admissible control u∗ ∈ u . moreover n∑ i=1 ∫ ω [zi,d − yi(t∗;u∗)] [yi(t ∗;u)− yi(t∗;u∗)] dρ ≤ 0, ∀u ∈ u . (113) now, we introduce the adjoint equation, and for every u ∈ u we define the adjoint variable p = p(u) = p(x, t;u) as the solution of the following (n× n) system, i = 1, 2, . . . , n, −∂pi(u) ∂t +a∗(t)pi(u) + ∫ b a ci(x, t+ h)pi(x, t+ h;u) dh = 0, x ∈ ω, t ∈ (0, t∗ − b), (114) advances in systems science and applications (2011), vol. 11, no. 1-2 23 −∂pi(u) ∂t +a∗(t)pi(u)+ ∫ t∗−t a ci(x, t+h)pi(x, t+h;u) dh = 0, x ∈ ω, t ∈ (t∗−b, t∗−a), (115) −∂pi(u) ∂t +a∗(t)pi(u) = 0, x ∈ ω, t ∈ (t∗ − a, t∗), (116) pi(x, t ∗;u) = zi,d(x)− yi(x, t∗;u), x ∈ ω, (117) ∂pi(u) ∂ηa∗ (x, t) = ∫ b a di(x, t+ h)pi(x, t+ h;u) dh, x ∈ γ, t ∈ (0, t∗ − b), (118) ∂pi(u) ∂ηa∗ (x, t) = ∫ t∗−t a di(x, t+ h)pi(x, t+ h;u) dh, x ∈ γ, t ∈ (t∗ − b, t∗ − a), (119) ∂pi(u) ∂ηa∗ (x, t) = 0, x ∈ γ, t ∈ (t∗ − a, t∗), (120) where ∂pi(u) ∂ηa∗ (x, t) = ∞∑ k=1 (dkpi(u)) cos(n, xk), (121) a∗(t)pi(u) = ( − ∞∑ k=1 d2 k + q(x, t) ) pi(u) + n∑ j=1 ajipj(u), (122) and aji is the transpose of aij . hence, as the proof of theorem 4.4 in section 4, we can prove the following theorem. theorem 5.8. if the assumptions concerning system (98)–(102) and controllability condition (112) are satisfied, then the time-optimal control u∗ = (u∗1, u ∗ 2, . . . , u ∗ n) exists and is characterized by the following condition n∑ i=1 ∫ t∗ 0 ∫ ω pi(u ∗)(ui − u∗i ) dρ dt ≤ 0, ∀u ∈ u , (123) where p(u∗) is the solution of the adjoint system (114)–(120). 6. conclusions and perspectives the results presented in the paper can be treated as a generalization of the results obtained by knowles (1978), and kowalewski and krakowiak (2006; 2008) onto the case of time optimal distributed and boundary control of second order infinite variables parabolic systems with deviating arguments appearing in the integral form both in state equations and in boundary conditions. we considered a different type of control, namely, the control function defined in the distributed and boundary of the spatial domain. sufficient conditions for the existence of a unique solution of such parabolic equations with neumann boundary conditions are proved (lemmas 3.2; 3.4; 3.6; 3.8; 5.1 and 5.5) and (theorems 3.3; 3.5; 3.7; 3.9; 5.2 and 5.6). the optimal control is characterized by using the adjoint equations (theorems 4.3; 4.4; 5.4 and 5.8). the conditions (41; 86 and; 112) plays a fundamental role in controllability problems 24 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . for time-delay parabolic systems. with regard to the controllability assumption (41; 86 and 112), we can investigate the exact controllability problem for the parabolic system (1)-(5). in this paper, we considered the time-optimal distributed and boundary control problem for infinite variables parabolic systems with non-homogeneous neumann and dirichlet boundary conditions. we can also consider an analogous minimum time problem for hyperbolic systems with non-homogeneous neumann and dirichlet boundary conditions. finally, we can consider the time-optimal control problem for discrete time delay distributed and boundary parameter systems. the ideas mentioned above will be developed in forthcoming papers. references [1] bahaa, g. m. (2003). quadratic pareto optimal control of parabolic equation with statecontrol constraints and an infinite number of variables. ima j. math. control and inform., 20, 167-178. [2] bahaa, g. m. (2005a). time-optimal control problem for parabolic equations with control constraints and infinite number of variables. ima j. math. control and inform., 22, 364375. [3] bahaa, g. m. (2005b). time-optimal control problem for infinite order parabolic equation with control constraints. differ. equ. control process. electron. j., 4, 64-81. available at http://www.neva.ru/journal. [4] bahaa, g. m. (2008). optimal control problems of parabolic equations with an infinite number of variables and with equality constraints. ima j. math. control and inform. 25, 37-48. [5] bahaa, g. m. and kotarski, w. (2008). optimality conditions for (n × n) infinite order parabolic coupled systems with control constraints and general performance index. ima j. math. control and inform. 25, 49-57. [6] berezanskii, ju. m. (1975). self-adjointness of elliptic operator with an infinite number of variables. ukrain. math. z. 27, 729-742. [7] choquet, g. (1969). lectures on analysis, vol.2, w.a. benjamin, new york. [8] dunford, n. and schwartz, j. (1958). linear operators, vol. 1, john wiley and sons, new york. [9] el-saify h. a. (2005). optimal control for (n × n) parabolic system involving time lag. ima j. math. control and inform. 22(3), 240-250. [10] el-saify h. a. (2006). optimal boundary control problem for (n × n) infinite order parabolic lag system. ima j. math. control and inform. 23(4),433-445. [11] el-saify h. a.& bahaa, g. m. (2001). optimal control for (n× n) systems of hyperbolic types. revista de matemáticas aplicadas, 22, 41-58. advances in systems science and applications (2011), vol. 11, no. 1-2 25 [12] el-saify, h. a.& bahaa, g. m. (2003). optimal control for (n × n) coupled systems of petrowsky type with an infinite number of variables. mathematica slovaca, 53, 291-311. [13] el-saify, h. a., serag, h. m, & bahaa, g. m. (2000). on optimal control for (n×n) elliptic system involving operators with an infinite number of variables. advances in modelling & analysis, 37, 47-61. [14] gali, i. m. & el-saify, h. a. (1982). optimal control of a system governed by hyperbolic operator with an infinite number of variables. j. math. anal. appl., 85, 24-30. [15] gali, i. m. & el-saify, h. a. (1983). distributed control of a system governed by dirichlet and neumann problems for a self-adjoint elliptic operator with an infinite number of variables. j. optim. theory appl., 39, 293-298. [16] knowles, g. (1978). time-optimal control of parabolic systems with boundary conditions involving time delays. j. optim. theor. appl., 25(4), 563-574. [17] kotarski, w., el-saify, h. a.& bahaa, g. m. (2002a). optimal control of parabolic equation with an infinite number of variables for non-standard functional and time delay. ima j. math. control and inform., 19, 461-476. [18] kotarski, w., el-saify, h. a. and bahaa, g.m. (2002b). optimal control problem for a hyperbolic system with mixed control-state constraints involving operator of infinite order. int. j. pure and appl. math., 1, 241-254. [19] kowalewski, a. (1988). boundary control of distributed parabolic system with boundary condition involving a time varying lag. international journal of control , 48(6), 2233-2248. [20] kowalewski, a. (1990a). feedback control for distributed parabolic system with boundary condition involving a time varying lag. ima j. math. control inform., 7(2), 143-157. [21] kowalewski, a. (1990b). optimal control of distributed parabolic system involving time lags. ima j. math. control inform., 7(4), 375-393. [22] kowalewski, a. (1993). optimal control of parabolic systems with time-varying lags. ima j. math. control inform., 10(2), 113-129. [23] kowalewski, a. (1998). optimal control of a distributed parabolic systems with multiple time varying lags. int. j. control, 69(3), 361-381. [24] kowalewski, a. (1999). optimization of parabolic systems with deviating arguments. int. j. control, 72(11), 947-959. [25] kowalewski, a . ( 2009). time optimal control of infinite order hyperbolic systems with time delays. int. j. appl. math. comput. sci,19, ( 4), 597-608 [26] kowalewski, a. and duda, j. (1992). on some optimal control problem for a parabolic system with boundary condition involving a time-varying lag. ima j. math. control inform., 9(2), 131-146. 26 bahaa:time-optimal control of infinite variables parabolic systems. . . . . . [27] kowalewski, a. and krakowiak, a. (1994). time-optimal control of parabolic time lag system, appl. math. and comp. sci., 4(1), 19-28. [28] kowalewski, a. and krakowiak, a. (2000). time-optimal control of parabolic system with time lags given in the integral form. ima j. math. control inform., 17(3), 209-225. [29] kowalewski, a. and krakowiak, a. (2006). time-optimal boundary control of parabolic system with time lags given in the integral form. inter. j. of appl. math. and comp. sci., 16(3), 287-295. [30] kowalewski, a. and krakowiak, a. (2008). time-optimal boundary control of infinite order parabolic system with time lags. inter. j. of appl. math. and comp. sci., 18(2), 189198. [31] li, x., & yong, j. (1995). optimal control theory for infinite dimensional systems. systems & control: foundations & applications. birkhäuser, boston.basel.berlin 1-448. [32] lions, j. l. (1971). optimal control of systems governed by partial differential equations. springer-verlag, band 170. [33] lions, j. l. & magenes, e. (1972). non-homogeneous boundary value problem and applications. i, ii, springer-verlag, new york. [34] tanabe, h. (1965). on differentiability and analyticity of weighted elliptic boundary-value problems. osaka mathematical j., 2, 163-190. [35] wang, p.k.c. (1975). optimal control of parabolic systems with boundary conditions involving time delays. siam j. control, 13(2), 274-293. [36] wong, k.h. (1987). optimal control computation for parabolic systems with boundary conditions involving time delays. j. optim. theory appl., 53, 475-57. advances in systems science and applications (2013) vol.13 no.4 379-391 purpose and non-purpose resource use models in two-level control systems olga i. gorbaneva and guennady a. ougolnitsky southern federal university, russia abstract in the paper a two-level control system consisting of one element in top level and one element in bottom level is considered. both levels have purpose use and non-purpose use interests. the model of resource allocation between purpose and non-purpose use interests for different classes of payoff functions is investigated. the model is built as a two-person game where the stackelberg equilibrium is found. analytical and numerical results are presented. keywords resource allocation, purpose use interests, non-purpose use interests, stackelberg equilibrium 1 introduction many problems of the social-economic development are solved due to the federal financing. the financing has different forms (grants, subsidies, assignments, credits) and is always strictly purpose oriented, i.e. the allocated resources should be spent for the pre-scribed needs only. there are legislative sanctions for the non-purpose use of the federal financing. nevertheless, the non-purpose use of federal resources is widely spread and may be considered as a variety of the opportunistic behavior meeting the private interests of active agents[1]. the non-purpose resource use is closely connected with corruption, especially with so-called “returns” when federal resources in some programs are assigned in exchange of a bribe and only partly satisfy the prescribed social destination being mostly used in the private interests of bribe-givers. it is natural to consider the problem of non-purpose resource use from the point of view of the concordance of interests in hierarchical control systems. this permits to use the mathematical formalism of the hierarchical game theory [2], the theory of incentives [3] and the theory of organizational systems [4-5]. in the same time, namely the models of resource allocation in the hierarchical systems with respect of their non-purpose use are scantily known and are analyzed by the authors’ methodology [6]. in this paper the emphasis is on the dependence of the distribution of resources between the purpose and non-purpose use on different classes of payoff functions characterizing the common (social) interest and the private interests of the resource distributor and resource recipient. in section 2 the structure of investigation is presented. sections 3 and 4 include analytical and numerical results respectively. section 5 concludes. 380 olga i. gorbaneva: purpose and non-purpose resource use models in ... 2 structure of investigation let us consider a two-level control system which consists of one element on the top level a1 (resource distributor-she) and one element on the bottom level a2 (resource recipient-he). without loss of generality we can equate to one the number of resources on the top level. the distributor delegates a part of her resources to the recipient for the purpose use, and the other part she keeps for the non-purpose use. in turn, the bottom level divides the received resource between the purpose use and his private non-purpose use. both levels have their shares in the purpose-use income and have their private payoff functions (fig.1). fig.1 the structure of the modeled system the model is built as a hierarchical two-person game in which a stackelberg equilibrium is sought [2]. the payoff functions of both players include two terms: a non-purpose income and the respective share of the purpose-use income. so, the payoff functions are: g1(u1, u2) = a1(1− u1) + b1(u1, u2)c(u1, u2) → max u1 , g2(u1, u2) = a2(u1, 1− u2) + b2(u1, u2)c(u1, u2) → max u2 subject to 0 ≤ ui ≤ 1. and the following conditions for the functions a, b and c: ai ≥ 0, ∂ai ∂ui ≤ 0, ∂ai ∂uj ̸=i ≥ 0,bi ≥ 0, ∂c ∂ui ≥ 0, i = 1, 2. here the subscript 1 relates to the parameters of the top level (the leader), and the subscript 2 relates to the parameters of the bottom level (the follower); ui a part of resources assigned by the i-th level for the purpose use (respectively the part 1ui leaves for the private non-purpose use); advances in systems science and applications (2013) vol.13 no.4 381 gi the i-th level payoff function; ai the i-th level function of his/her private interest; bi a share of the purpose-use income received by the i-th level; c a function of the purpose-use income of the whole system (society, organization). as functions a and c power, exponential and logarithmic functions of the variables u1 and u2 are considered which are cumulative ones, i.e. a1=a1(1-u1), a2=a2(u1 (1-u2)), c=c(u1u2). in this case the share u1u2 of resources is assigned for the purpose use. the relations a1=a1(1-u1), a2=a2(u1 (1-u2)) reflect the system’s hierarchical structure. the non-purpose income of the top level does not depend on the part of resources assigned by the bottom level for the purpose use, but the non-purpose income of the bottom level depends on the share of resources received by him from the top level. the following distributions of the purpose-use income b are considered: (1) uniform one, in particular, for n = 2 bi = 1 2 , i = 1, 2. (2) proportional one b1 = u1 u1 + u2 , b2 = u2 u1 + u2 . in this paper it is assumed that b1 + b2 = 1. the strategy of a player i is a part ui of his/her resources as-signed for the purpose use. the top-level player moves first, i.e. she chooses a value u1 and informs about it the bottom-level player who chooses his optimal reaction u2. the aim of investigation is study of the influence of the relations between functions a1, a2, b1, b2, c to the solution of the game (stackelberg equilibrium). the following types of non-purpose use functions were used: power function with an exponent smaller than one (a(x) = axα, 0 < α < 1, a > 0), linear function (a(x) = ax), power function with an exponent greater than one (a(x) = axk, k > 1, a > 0); exponential function (a(x) = a(1− e−λx), λ > 0, a > 0); logarithmic function (a(x) = a log2(1 + x), a > 0). almost all of the functions satisfy the conditions ∂a/∂x ≥ 0, ∂2a/∂x2≤0 (except the second condition for the function a (x) = axk, k > 1). similarly, the following types of purpose use functions were used: power function with an exponent smaller than one (c (x) = cxα, 0 < α < 1, c > 0); 382 olga i. gorbaneva: purpose and non-purpose resource use models in ... linear function (c(x) = cx); power function with an exponent greater than one (c(x) = cxk, k > 1, c > 0); exponential function (c(x) = c(1− e−λx), λ > 0, c > 0); logarithmic function (c(x) = c log2(1 + x), c > 0). thirteen of the possible twenty five combinations of the functions a and c are studied analytically, namely: (1) combinations of the one-type functions (both functions a and c are power, exponential, or logarithmic ones); (2) combinations of any non-purpose use function with linear purpose-use function; (3) combinations of any purpose-use function with non-purpose use linear function. six of other twelve cases are investigated numerically. 3 analytical investigation of the different classes of models first, let’s consider the following parameterization: a1(u1, u2) = a1(1− u1), a2(u1, u2) = a2u1(1−u2), c(u1, u2) = clog2(1 + u1u2), b1 = b, b2 = 1− b. in this case the payoff functions are g1(u1, u2) = a1(1− u1) + bclog2(1 + u1u2) (1) g2(u1, u2) = a2u1(1−u2) + (1− b)clog2(1 + u1u2) (2) subject to 0 ≤ ui ≤ 1, i = 1, 2. omitting the calculations, consider each branch of the stackelberg equilibrium separately: i. u = (0; 0) , if a2 > (1− b) c/ln2 or a1 > bc/ln2 (fig.2), i.e. for one of the players the non-purpose resource use is much more profitable than the purposeuse activity; therefore, it is disadvantageous for him/her to assign resources for the purpose-use activity, and in this case it is also disadvantageous for the other player. the payoffs in this case are equal to: g1 = a1, g2 = 0. ii. u = (1; 1) , if a2 < (1− b) c/ln2 and a1 < bc/ln2 (fig.2), i.e. for both players the purpose-use activity is much more profitable, and each of them assigns all resources for it. the payoffs are equal to: g1 = bc, g2 = (1− b) c. iii. u = (bc/ (a1ln2)− 1; 1), if the conditions bc/2ln2 < a1 < bc/ln2 and a2 < a1 (1− b) /b are satisfied (fig.2), i.e. for the top-level player it is profitable to assign only a part of her resources for the purpose use because her incomes from both activities are comparable, meanwhile for the bottom-level player it is profitable to assign all his resources to the purpose-use activity. the payoffs are advances in systems science and applications (2013) vol.13 no.4 383 fig.2 equilibrium outcomes in the game(1) (2) equal to: g1 = 2a1 − bc ln 2 + bclog2 ( bc a1 ln 2 ) , g2 = (1− b)clog2 ( bc a1 ln 2 ) iv. u = ((1− b) c/ (a2ln2)− 1; 1), if (1− b) c/2ln2 < a2 < (1− b) c/ln2 and a2 > a1 (1− b) /b (fig.2), i.e. for both players it is profitable to assign only a part of their resources for the purpose-use activity because their incomes from both types of activities are comparable. but the leader gives to the follower exactly the fixed number of resources which he planned to allocate for the purpose use, therefore compel-ling him to assign all his resources for the purpose use. the players’ payoffs are equal to g1 = 2a1 − a1(1− b)c a2 ln 2 + bclog2 ( (1− b)c a2 ln 2 ) , g2 = (1− b)clog2 ( (1− b)c a2 ln 2 ) . second, consider the following parameterization: a1(u1, u2) = a1(1− u1) k, a2(u1, u2) = a2(u1(1−u2)) k, c(u1, u2) = c(u1u2), b1 = b, b2 = 1− b. 384 olga i. gorbaneva: purpose and non-purpose resource use models in ... then the payoff functions have the form g1(u1, u2) = a1(1− u1) k + bc(u1u2) → max u1 (3) g2(u1, u2) = a2(u1(1−u2)) k + (1− b)c(u1u2) → max u2 (4) subject to 0 ≤ ui ≤ 1, i = 1, 2. the stackelberg equilibrium outcomes are the following : ū = { (1; 1), (a1 < bc)& (a2 < (1− b)c) (0; 0), (a1 > bc) ∨ (a2 > (1− b)c) consider the cases separately (fig.3): fig.3 equilibrium outcomes in the game(3) (4) i. u = (0; 0), if a2 > (1− b) c or a1 > bc, i.e. for one of the players the nonpurpose resource use is much more profitable than the purpose-use one, therefore it is disadvantageous for her/him to finance the purpose-use activity. the payoffs are: g1 = a1, g2 = 0. ii. u = (1; 1), if a2 < (1− b) c and a1 < bc, i.e. for both players the pur-poseuse activity is much more profitable, and each of them assigns all resources for it. the payoffs are equal to: g1 = bc, g2 = (1− b) c. when even one of the functions of purpose or non-purpose re-source use is power with an exponent greater than one (and the other function is the same or linear) then it is advantageous for both players to allocate resources or only to the purpose use (altruistic strategy), or only to the non-purpose use (egoistic strategy). advances in systems science and applications (2013) vol.13 no.4 385 third, consider the following case: a1(u1, u2) = a1(1− e−λ(1−u1)), a2(u1, u2) = a2(1− e−λu1(1−u2)), c(u1, u2) = c(1− e−λu1u2), b1 = b, b2 = 1− b. then the payoff functions have the form g1(u1, u2) = a1(1− e−λ(1−u1)) + bc(1− e−λu1u2) (5) g2(u1, u2) = a2(1− e−λu1(1−u2)) + (1− b)c(1− e−λu1u2) (6) subject to 0 ≤ ui ≤ 1, i = 1, 2. let’s consider all stackelberg outcomes separately (fig.4): fig.4 equilibrium outcomes in the game(5) (6) i. u = (0; 0), if ( a1 > bceλ ) &(a1 > (1− b) c) or a1 > bceλ √ a2 (1−b)c , i.e. for one of the players the non-purpose resource use is much more profitable than the purpose-use one, therefore it is disadvantageous for her/him to finance the purpose-use activity. the payoffs are: g1 = a1(1− e−λ), g2 = 0. ii. u = (1; 1), if a1 < bce−λ and a2 < (1− b) ce−λ, i.e. for both players the purpose-use activity is much more profitable, and each of them assigns all resources for it. the payoffs are equal to:g1 = bc(1− e−λ), g2 = (1− b)c(1− e−λ). iii. u = ( 1; 12 − 1 2λ ln a2 (1−b)c ) , if a1 < b 2 √ a2c 1−be −λ 2 and (1− b) ce−λ < a2 < (1− b) ceλ , i.e. for the top-level player it is profitable to assign all her resources to the purpose-use activity, meanwhile for the bottom-level player it is advantageous to divide his resources. payoffs are the following: 386 olga i. gorbaneva: purpose and non-purpose resource use models in ... g1 = bc ( 1− e −λ ( 1 2 − 1 2λ ln a2 (1−b)c )) = b ( 1− e− λ 2 )√ a2c (1− b) , g2 = (1− b) ( 1− e− λ 2 )√ a2c (1− b) . iv. u = ( 1 2 − 1 2λ ln a1 bc ; 1 ) , if bce−λ < a1 < bceλ and a2 < (1 − b) √ bc a1 ce λ 2 , i.e. the situation is opposite to the previous one. the payoffs are: g1 = a1 − a1e −λ 2 √ bc a1 + bc− e− λ 2 √ bca1, g2 = (1− b)c ( 1− e− λ 2 √ a1 bc ) . v. u = ( 2 3 − 2 3λ ln 2a1 b √ 1−b a2c ; λ−ln 2a1a2 b(1−b)c 2 ( λ−ln 2a1 b √ 1−b a2c ) ) , if the conditions b 2 √ a2c 1−be −λ 2 < a1 < b 2 √ a2c 1−be λ and (1− b) √ 2a1 bc e −λ 2 < a2 < bc2 2a1(1−b)e λ are satisfied, i.e. for both players it is profitable to divide their resources. the payoffs are omitted due to their tediousness. at last, consider the following case: a1(u1, u2) = a1(1− u1), a2(u1, u2) = a2 (u1(1−u2)) , c(u1, u2) = c(u1u2) α, b1 = b, b2 = 1− b. then the payoff functions have the form g1(u1, u2) = a1(1− u1) + bc(u1u2) α → max u1 (7) g2(u1, u2) = a2 (u1(1−u2)) + (1− b)c(u1u2) α → max u2 (8) the stackelberg equilibrium has the form (fig.5): ū =  (1; 1), (a1 < bcα)&(a2 < (1− b)cα),( 1−α √ αbc a1 ; 1 ) , (a1 > bcα)&(a2 < a1),( 1−α √ (1−b)cα a2 ; 1 ) , (a2 > (1− b)cα)&(a2 > a1). in this case the egoistic strategy is disadvantageous for both players. besides, the leader is always able to compel the follower to assign all his resources to the advances in systems science and applications (2013) vol.13 no.4 387 fig.5 equilibrium outcomes in the game(7) (8) purpose use. the payoffs are: g1 = a1 − a1 ( α(1− b) a2 ) 1 1−α + b ( αα(1− b)αc a2α ) 1 1−α , g2 = ( αα(1− b)c a2α ) 1 1−α . when the purpose-use function is linear, and the function of non-purpose use is power with exponent smaller than one, the altruistic strategy is profitable for the bottom-level player. the egoistic strategy is disadvantageous for both players. the thirteen analyzed cases are grouped by the structure of equilibrium outcomes of the game: i. one outcome when both functions of purpose and non-purpose use are power with exponent smaller than one. in this case it is profitable for both players to divide their resources between purpose and non-purpose use. ii. two outcomes (0; 0) and (1; 1) (fig.3) when: • the function of non-purpose use is power with exponent smaller than one, and the function of purpose use is linear; • both functions are linear or power with exponent greater than one, in any combination. iii. three outcomes (fig.5) when the non-purpose resource use function is linear, and the purpose-use function is power with exponent smaller than one. in this the altruistic strategy is profitable even for one player. iv. four outcomes (fig.2) in cases when one of the functions is linear, and the 388 olga i. gorbaneva: purpose and non-purpose resource use models in ... other is logarithmic. v. five outcomes (fig.4) when • both functions are linear or exponential in any combination except the case when they are both linear. • both functions are logarithmic. 4 numerical analysis let’s consider an example of the numerical analysis for the following parameterization: a1(u1, u2) = a1(1− u1) α, a2(u1, u2) = a2(u1 (1− u2)) α, c(u1, u2) = c(1− e−λu1u2), b1 = b, b2 = 1− b. in this case the game has the form g1(u1, u2) = a1(1− u1) α + bc(1− e−λu1u2) → max u1 (9) g2(u1, u2) = a2(u1 (1− u2)) α + (1− b)c(1− e−λu1u2) → max u2 (10) to find an optimal strategy of the bottom-level player let’s calculate the derivative of the function g2 with respect to the variable u2 and equate it to zero: ∂g2 ∂u2 (u1, u2) = − a2αu1 α (1− u2) 1−α + λu1(1− b)ce−λu1u2 = 0 (11) let’s prove that the method of bisection is applicable for the solution of (11). note that the second derivative of the function g2 with respect to the variable u2 is negative ∂2g2 ∂u22 (u1, u2) = a2α(1− α)u1 α (1− u2) 2−α − λ2u1 2(1− b)ce−λu1u2 < 0, and therefore the function ∂g2/∂u2 is monotone. now let’s calculate the signs of ∂g2/∂u2 in the ends of the segment [0; 1]. ∂g2 ∂u2 (u1, 0) = −a2αu1 α + λu1(1− b)c (12) ∂g2 ∂u2 (u1, u2) → u2→1− −a2αu1 α 0+ + λu1(1− b)ce−λu1u2 → u2→1− −∞ (13) if (12) is positive then the equation can be solved by bisection, and the solution will be the point of maximum due to the negativity of the second derivative. if (12) is negative then the method of bisection is not applicable but the left side of the equation is monotone and therefore it is negative in the segment [0; 1], so the function g2 de-creases and the point of maximum is u2 = 0. thus, advances in systems science and applications (2013) vol.13 no.4 389 u2 ∗ = { 0, −a2αu1 α + λu1(1− b)c < 0, ∈ (0; 1), −a2αu1 α + λu1(1− b)c > 0. the top-level player can use the information to impel the bottom-level player to choose a non-zero strategy u2 > 0. it is necessary for this to assure the condition −a2αu1 α+λu1 (1− b) c > 0. solving the inequality with respect to u1 we rewrite the condition as u1 > 1−α √ a2α λ(1−b)c . the top-level player can satisfy the condition only if 1−α √ a2α λ(1−b)c < 1 , or a2 < ( λ(1−b)c α )1−α . if the top-level player cannot choose the strategy then the bottom-level player chooses u2 = 0. in this case g1(u1, 0) = a1(1− u1) α. as far as the function g1 decreases with respect to u1 we receive u1 = 0. let’s summarize: i. if a2 > ( λ(1−b)c α )1−α then the leader cannot influence to the follower and u2 = 0, therefore u1 = 0. this case takes place when the effect from non-purpose activity on the bottom level is essentially greater than the effect from his purposeuse activity. ii. if a2 < ( λ(1−b)c α )1−α then the leader can impel the follower to assign a part of his resources for the purpose-use activity by choosing u1 > 1−α √ a2α λ(1−b)c . this case takes place when the effect from purpose-use activity on the bottom level is essentially greater than the effect from his non-purpose use activity. the case of parameterization a1(u1, u2) = a1log2(2− u1), a2(u1, u2) = a2log2(1 + u1(1− u2)), c(u1, u2) = c(u1u2) α, b1 = b, b2 = 1− b. is considered similarly. the payoff functions have the form g1(u1, u2) = a1log2(2− u1) + bc(u1u2) α, (14) g2(u1, u2) = a2log2(1 + u1(1− u2)) + (1− b)c(u1u2) α (15) the results of analysis are presented in fig.6. if a1 < αbcln2 and a2 > α (1− b) cln2 then u1 = 1−α √ (1−b)cα ln 2 a2 , u2 = 1. the payoffs of the players are: g1 ( 1−α √ (1−b)cα ln 2 a2 , 1 ) = a1log2 ( 2− 1−α √ (1−b)cα ln 2 a2 ) + bc ( 1−α √ (1−b)cα ln 2 a2 )α g2  1−α √ (1− b)cα ln 2 a2 , 1  = (1− b)c ( (1− b)cα ln 2 a2 ) α 1−α . 390 olga i. gorbaneva: purpose and non-purpose resource use models in ... fig.6 equilibrium outcomes in the game(14) (15) if a1 < αbcln2 and a2 < α (1− b) cln2 then the altruistic strategy is advantageous for both players: u1 = 1, u2 = 1. the payoffs are:g1 (1, 1) = bc, g2 (1, 1) = (1− b) c. if a1 > αbcln2 then it is profitable for the top-level player to allocate for the purpose use a part of her resource u1 ∈ ( 0;min { 1−α √ (1−b)cα ln 2 a2 ; 1 }) . 5 conclusion in this paper the problem of non-purpose resource use is considered from the point of view of analysis and design of the control mechanisms providing the concordance of interests in the hierarchical (two-level) systems. interests of the agents are described by their payoff functions including two terms: profit from the purpose and non-purpose resource use respectively. different classes of the payoff functions are studied. the top-level control agent (resource distributor) is treated as leader, and the bottom-level agent (resource recipient) as follower what results in the concept of stackelberg equilibrium. the analytical and numerical investigation allows for the following conclusions. • when both purpose and non-purpose interests functions are power (k < 1) then it is advantageous for both players to invest a part of their resources to the purpose use and the other part to the non-purpose use; • when even one of the payoff functions is power (k > 1), and the other is also power (k > 1) or linear then it is advantageous for both players or assign the resources only for the purpose use (the altruistic strategy), or only for non-purpose use (the egoistic strategy); • in other cases the following situations are possible: advances in systems science and applications (2013) vol.13 no.4 391 a) when the payoff from the non-purpose activity of a player is much greater than the payoff from the purpose activity then the egoistic strategy is advantageous; b) when the payoff of non-purpose activity of both players is much smaller than the payoff from the purpose activity then the strategy of pure altruism is advantageous; c) when the payoffs of the purpose and non-purpose activities are comparable then it is profitable to divide the resources between the purpose and non-purpose interests. acknowledgements the work is supported by rfbr (projects 12-01-00017, 12-01-31287) and by southern federal university. references [1] williamson o.e. (1985), the economic institutions of capitalism, free press. [2] basar t, olsder g.y. (1999), dynamic noncooperative game theory, siam, philadelphia, pp.536. [3] laffont j.j, martimort d. (2002), the theory of incentives: the principalagent model, princeton university press, pp.421. [4] prof. d. novikov (2013),mechanism design and management: mathematical meth-ods for smart organizations, n.y: nova science publishers, pp.163. [5] novikov d. (2013), theory of control in organizations, n.y: nova science publishers, pp.341. [6] ougolnitsky g. (2011), sustainable management, n.y: nova science publishers, pp.285. corresponding author guennady a. ougolnitsky can be contacted at: ougoln@gmail.com microsoft word 14 haoming dong, dingfang chen shaojun zou--development framework for gantry crane training system based on v advances in systems science and applications (2011), vol.11, no.3-4 309-314 issn 1078-6236 international institute for general systems studies, inc development framework for gantry crane training system based on virtual reality haoming dong1,2 , dingfang chen1 and shaojun zou2 1wuhan university of technology , wuhan 430063, china; 2wuhan supervise and test institute of especial equipment, , wuhan 430018, china abstract with the rapid development of science and technology, and the improvement of production safety and management standards, the use of virtual reality technology to gantry crane operator teaching simulation to train high quality personnel has become a pressing need. this article points out that virtual reality technology will be applied to gantry crane operator teaching simulation, a virtual simulation on the gantry crane operator teaching system and the composition of achieving. this article focuses on the virtual simulation platform for the establishment and adopted by key technologies, introduces the 3d modeling, kinetic analysis of crane, and the control hardware system, and analyzes and forecasts the application of the simulation system. keywords virtual reality, gantry crane, simulation training system 1.introduction in recent years, the national production safety situation is not optimistic. the gantry crane accidents have occurred occasionally. the main reason is that some gantry crane operators are lack of expertise and operating skills, poor quality of the safety operation, sometimes even wrong operation and so on. at present, the gantry crane as the category of special equipment in china has established a very sound safety training management system for special operations, but the teaching process and method is much lagged. in the teaching process, as a result of outdated teaching materials and methods, the actual operation conditions is limited, leading to the problem that gantry crane operators did not really acquire advanced theoretical knowledge. the practice of operating capacity is not strong, resulting in ineffective training and mining a large number of potential causes of accidents. therefore, it is a pressing task to change the status quo of operating personnel training of the gantry crane, and to improve the quality of personnel in professional and operational skills to ensure the security. implementation of the system can fundamentally change the gantry crane operation training status, and improve the training effectiveness and economize resources, and improve the operating personnel quality and safety skills, and improve the existing examination procedures, and eliminate safety incidents occurring due to human factors, and provide technical support for scientific management to special equipment for operating personnel. 2.virtual reality virtual reality (vr) is the modern hi-tech which can generate realistic visual, hearing, touch the specific scope of the integration of virtual environments with the computer technology. with the necessary equipment, users can interact with the virtual environmental object and exert the impact. in this way, the users can get the immersive experience of the real environment [1]. vr technology is that users can immerse in an artificial virtual environment and then design and complete the mission by interaction with the computers fully through the virtual reality software and its external devices. see its concept in figure 1. 310 dong: development framework for gantry crane training system based on virtual reality it combines advanced information technologies (such as network computing, graphics and image processing, multimedia technology, the new sensors, simulation, etc.), and even psychology, bionics, arts and other fields of research. the purpose is to enable people to build an immersion combined with the actual situation, the efficient multi-dimensional interactions and more harmonious human-computer environment. fig. 1 i3 of vr 3.research of simulated system on gantry crane the system is developed on osg system. it is easy to develop all types of vr system. the simulated system of gantry crane is composed of several parts, including visual system, human-computer system, hardware system. see the function of these parts in figure 2. fig.2 module of function 3.1 creation of visual system the system creates the visual system of gantry crane for the trained personnel. the visual scene includes the exterior and the crane. the factory, material, device and operators constitute the visual scene. and the scene is created by the technology of image processing, such as romance, texture, shade, lighting and so on. these technologies can immerse the operator in the scene. [3] the main body of crane provides suppositional operation platform, and with organic anastomosis of true driver's cage together. the trainee can see operation action's looking at scenery effect in the true cab. see the scenes in figure 3~ 4. advances in systems science and applications (2011), vol.11, no.3-4 311 fig.3 workshop scene fig.4 harbor scene 3.2 the research of human-computer interaction system in order to obtain the real interactive training effect, the sway simulation testing based on physical property and virtual interactive study are to be carried out for the training and assessment system of gantry crane. they mainly includes: (1) the research of elasto-dynamics system is to simulate the carrying sway and simulated failures and accidents of gantry crane under the circumstances of different carrying capacity, different braking time, and different running speed and so on. [2] (2) interactive simulation technology based on the collision detection technology is to achieve the real-time simulation of the gantry crane carrying objects and the collision avoidance of obstacles. and, the stress analysis of wirerope is the key to solve the problem of carrying sway and collision detection during the carrying process. operating condition of gantry machine can be divided into the following parts: the start (i.e. to accelerate movement), the uniform speed movement, the brake (i.e. to decelerate movement). the system does not obtain acceleration during the uniform speed movement, that is, it is easy to carry out real-time simulation since there is no horizontal force attached to a crane and non-sway problem. as for the stage of startup or braking, it will not affect the forces of each other by taking into account the fact that the car and the cart are rigidly connected with the interactive movement. in order to facilitate the analysis, we only take into account the forces on the wire rope during the running of the car when we analyze. when the car starts running or parking brake, the goods will make a sway which will foist additional horizontal forces on the crane structure. the emergence of such a dynamic load has a certain impact on the operation of the crane and will 312 dong: development framework for gantry crane training system based on virtual reality produce sway problems. when the dynamic load calculation of the car braking operating is similar to the operation condition of the startup, the additional horizontal force generated is opposite to it. therefore, the mechanical condition of different car models and dynamic load calculations can be simplified for the car to start, namely the force analysis and the calculation of dynamic load during the acceleration. see the fig.5. figure 5 force analysis during the movement of the car the dynamics analysis of large cart can be gradually worked out through the establishment of the above model and we can conduct calculating simulation. this system enables operator's to interact with the three-dimensional scene well. the operator can be trained to move cars or carts, high or low lift hook, and obtain all kinds of safety device skill in various scenes so that the operator acquires preliminary understanding of the basic principles and fundamental of the gantry crane, and experiences essentials to operate the gantry crane and taste the moving mode of all kinds of safety devices,. 3.3 the research of hardware control system based on the present operating room of the gantry crane, the research of the system utilize the sensor, the data communication and acquisition techniques to design and develop a set of controlling and demanding system in accordance to a general-purpose gantry crane control mode (controller mode), through which each sensor can collect control signal, and connect the main correspondence line through the interface i/o. the hardware control system collect various types of control lever, button signal and control signal to put on processing through the transferring relevant dynamic link storehouse. the transferring procedure is as follows: hplx = pciclose(hplx) public hplx as long public addr as boolean public dwvendorid as integer public dwdeviceid as integer public fuseint as boolean apart from the signal acquisition system, the hardware control system also includes the driver's cab structural design, reasonable layout of the control box, and display devices, sensors, power supply, as well as the location of signal alignment, and so on. on the basis of the above, we focus on the timeliness and anti-jamming of the operation signal. through the research and integration of the above systems, we can enable the trainees to interact directly with the control system. the operating movement signal is gathered by the control system and carried on the real-time processing, then the control visual system responses and gives feedback to the trainee via display device of the visual system. meanwhile the advances in systems science and applications (2011), vol.11, no.3-4 313 operating movement signal is simultaneously delivered to the system of training and assessment, which enables the experts system to make judgments on the operating movements which provide the fundamental basis for the final appraisal training. these four modules are organically linked together to provide users with realistic training environment. 4.the application and merits of virtual reality in training the system is a set of simulation operating system used in the hands-on training, skills training, safety education and practice of the gantry crane. it mainly promotes or improved the effects of training and assessment of the gantry crane operator in following aspects. (1) the use of the immersion of virtual reality to improve the understanding of the structure and working principle of the hoisting machinery. 2)the use of the interactivity of virtual reality to enhance the actual ability of operating personnel. 3)accident reappearance and case analysis the trainees may observe and analyze their own hands-on training process through on-site reappearance during the operation training process, especially when there occur glaring errors, such as a collision, hitting obstacles, gnawing rails, and safety devices movements. the use of the reappearance effect can help to determine and analyze the reasons for the error. it is suitable for students without any experience of operating gantry crane to carry out hands-on training, and some green hands who have qualified to be operators to conduct the actual practice. compare with the previous methods of training and assessment, the simulation operating system of the gantry crane based on the virtual reality technology will have a powerful functions. the advantages are as follows in terms of conducting training, assessment, education of security techniques: 1)be able to set a variety of conditions conveniently, efficiency, economy and high security compared by calculating, the training in terms of the present system can consumes only about one kilowatt in energy and occupies an area of within 10 square. while the previous real vehicle training consumes over 20 kilowatts in terms of energy and occupies over 200 square. if 3,000 persons will be trained each year in our city and 40 hours will be spent on each person per training hours, the present system only in terms of energy consumption will be able to save: (20-1)×3000×40=2280000 kilowatts hour in terms of the current electric charge of 0.5 yuan/ kwh, the total savings on electricity amount to 1.14 million yuan. 2)good reproducibility the factors concerning the working and operating conditions of the gantry crane is hard to manage, so the reproducibility of the real vehicle test is poor. while gantry crane simulator based on virtual reality technology can conveniently carry out the data acquisition, model selection and the environment settings of simulation model with a good reproducibility. 3)high security high-speeding and maximal driving, as well as very dangerous safety experiment can be safely carried out with the utility of the simulation operating system of gantry crane, which cannot achieve in the real vehicle experiment. 5.conclusion the virtual reality technology is an integrated multi-disciplinary technology. its basic philosophy is to use the method of modern science and technology to create artificially a virtual space in which people can achieve interaction such as watching, listening, and moving and so on, just like in the real environment. the users may enter this environment through the computer and can control and interact with the object in the system with timeliness and interactiveness in the three-dimensional environments as its chief features. 314 dong: development framework for gantry crane training system based on virtual reality the application of the virtual reality to the gantry crane operating personnel training is to design the idea of virtual scenes and conceive it to be viewable and feasible; to achieve the vivid spot effects; to provide simultaneously all kinds of direct-viewing and natural sensation interactive method such as the sensations of hearing, seeing, touching, etc. through the computer user’s connection; to utilize the computer simulation trainer to train the operating personnel to the full while shorten the training cycle, and reduce the training expense, thus enhance training quality and efficiency of gantry crane operator. developed gantry crane virtual training system which based on virtual reality technology achieves good results in the process of actual use. acknowledgements the project is funded by general administration quality supervision, inspection and quarantine of people’s public of china (project no. 20070k0229), and wuhan science and technology bureau (project no. 200711021381) references [1] shao-jun zou. on management of hoisting and conveying machinery operators. hoisting and conveying machinery. 2008 vol 03 [2] zhonghua lu, chenlin shen, zhe liu. research on distributed multi-screen overhead crane simulator. 2008 3rd international conference on pervasive computing and applications. icpca08 [3] guoqian wei, zhengyan zhang. research on the scene displaying scheme of the training simulator for crane. proceedings of 2008 ieee 9th international conference on computer aided industrial design & conceptual design vol.1. 2008-11-01 [4] jian li. the development and application of container crane training simulator. master dissertation of shanghai maritime university [5] de segura, j.d.g., peral, r. using virtual reality for gesture and vocal interface validation in industrial environments. 17th international conference on artificial reality and telexistence. icat 2007 , page: 294-5 [6] guoqian wei, qiuhua tang.research on the framework of the crane's virtual design system and its key technologies[c].2006 ieee international conference on robio: 1293-1298 advances in systems science and applications (2012) vol.12 no.3 238-257 foundations of control methodology dmitry novikov1,2 and elena rusjaeva1 1institute of control sciences ,russian academy of science, moscow, russia 2moscow institute of physics and technology, moscow, russia abstract control methodology is defined as the theory of control activity organization. methodology of control activity, its characteristics, logical and temporal structures are described. philosophical foundations of control methodology are introduced. keywords control, methodology, control activity, control philosophy 1 introduction methodology is the theory of organization of an activity[1-2]. such definition uniquely determinates the subject of methodology which is organization of an activity (an activity is a purposeful human action). methodology being treated as the theory of organization of an activity, one should naturally consider the notion of an “organization”. according to the definition provided by merriamwebster dictionary and[3], an organization is: 1) the condition or manner of being organized; 2) the act or process of organizing or of being organized; 3) an administrative and functional structure (as a business or a political party); also, the personnel of such a structure. thus, an organization may be considered as the property of being organized (the first meaning) and the process of organizing including the result of this process (the second meaning). the third meaning is an organizational system[3]. let us classify an activity based on its ultimate goal (play-learning-labor[2]). in this case, one distinguishes among: -the methodology of play activity (in the first place, play of children); -the methodology of learning activity; -the methodology of labor (professional) activity. next, professional (or practical) activity can be subdivided into: -practical activity (in the fields of material and immaterial production). in the above sense, most of people are engaged in practical professional activity; -specific forms of professional activity such as philosophy, science, art, and religion. accordingly, we separate out philosophic activity, scientific activity, art activity, and religious activity. in scientific literature, one can find relatively full coverage of the methodologies of scientific activity (research methodology), practical activity, educational activity, as well as the basics of the methodologies of art activity and play activity[1-2]. advances in systems science and applications (2012) vol.12 no.3 239 control activity is the primary subject of this paper. control activity represents a type of practical activity, see fig.1. control methodology is the theory of organization of control activity1[4]. within the framework of general methodology approaches[2], it is possible to construct methodologies of other types of practical activity (learning activity, medical activity, etc.) by analogy with control methodology. methodology considers organization of an activity. organizing an activity means arranging it as an integral system with clearly defined characteristics, a logical structure and the accompanying process of its realization, the temporal structure. the corresponding reasoning lies in the pair of the dialectic categories “historical (temporal)” and “logical”. fig.1 types of activity: a classification the logical structure includes the following components of control activity: subject, object, topic, forms, means, methods, and result. the following characteristics of activity are external with respect to this structure: features, principles, conditions, and norms. the process of activity implementation is usually considered within the framework of a project realized in a time sequence by phases, stages and steps. furthermore, this sequence is common for all kinds of activity[2]. the completeness 1alternative approaches to the definition of methodology take place, as well. for instance, methodology is considered as the theory of methods (generally, methods of scientific research). following such idea, one can understand control “methodology” as epistemological foundations of control science. another approach lies in treating methodology as the theory of methods of practical activity. as a branch of cybernetics, the corresponding control “methodology” studies general methods (ways of implementation) of control, e.g., disturbance-based control or deviation-based control. in addition, there exist narrower “branchwise” interpretations, e.g., “project management methodology” as the whole set of general rules of efficient project management. another example consists in “quality management methodology”. and so on. and these are different control “methodologies”! 240 dmitry novikov:foundations of control methodology of an activity cycle (a project) is defined by the following three phases: -design phase,which yields the models of activity of a control subject and a controlled system, as well as and the plan of their implementation; -technological phase,which yields implementation of control actions; -reflexive phase,which yields an estimate of the results of control activity and indicates the necessity of its further correction or “launching” of a new project (i.e., designing a new control system). therefore, it is possible to suggest the following “textitscheme of control methodology”[4]: 1.the characteristics of control activity (features and principles); 2.the logical structure of control activity including subject, object, topic, forms, means, methods, and result of control activity; 3.the temporal structure of control activity (phases, stages, and steps). let us outline the structure and logic of the paper. first the general scheme of any activity is discussed (see section 2). in section 3 characteristics, logical and temporal structures of control activity are presented. philosophical foundations of control methodology are given in section 4. 2 control activity let us consider the basic structural (procedural[1,5]) components of any activity of some subject, see fig.2 (for convenience, the margins of a subject (individual or collective one) are marked by the dotted rectangle. the chain “need → motive→ goal→ tasks→ technology→ action→ result,” highlighted by thick arrows in fig.2, corresponds to a single “cycle” of activity. the goal is decomposed with respect to conditions, norms and principles of activity into a set of tasks. next, taking into account the chosen technology (that is, a system of conditions, forms, methods and means to solve tasks), a certain action is chosen; note that technology includes content and forms, methods and means. the above-mentioned action leads (under the influence of an environment) to a certain result of activity). a particular position within the activity structure is occupied by those components referred to as either self-regulation (in the case of an individual subject) or control (in the case of a collective subject), see fig.4. self-regulation represents a closed control loop. during the process of self-regulation the subject modifies the components of his activity based on the assessment of the achieved results (see the thin arrow in fig.2). thus, we have discussed primary characteristics of activity and the corresponding structural components. now, let us proceed to control issues. starting the discussion about control, one should precisely formulate what control is; therefore, below we give a series of common definitions, paying special attention to the purposefullness of control: advances in systems science and applications (2012) vol.12 no.3 241 fig.2 structural components of activity control is “the process of checking to make certain that rules or standards being applied” (macmillan dictionary). control is “the act or activity of looking after and making decisions about something” (merriam-webster dictionary). control is “an influence on a controlled system with the aim of providing the required behavior of the latter[3]”. there exist numerous alternative definitions, which consider control as a certain element, a function, an action, a process, a result, an alternative, and so on. we would not intend to state another definition; instead, let us merely emphasize that control being implemented by a subject2 should be considered as an activity. such approach, when control is meant as a type of practical activity3 (control activity, management activity-see above) puts many things into place; in fact, it explains “versatile character” of control and balances different approaches to this 2this eliminates from consideration situations when control is implemented by a technical system (activity is inherent to human beings only). hence, control methodology, as the theory of organization of control activity, studies exclusively (!) situations when control is performed by a human being or by a group of people. furthermore, choosing between two remaining alternatives (a controlled system comprises people or represents a techniqal system), we will be mainly focused on the first alternative as the most complicated one. in addition, we emphasize that the activity of a researcher designing a control system is not control activity but scientific activity. similarly, the activity of an engineer designing a technical system is not control activity but practical (engineering) activity. 3at first glance, interpreting control as a sort of practical activity seems a bit surprising. the reader knows that control is traditionally seen as something lofty and very general; however, activity of any manager is organized similarly (satisfies the same general laws) to that of any practitioner, e.g., a teacher, a doctor, an engineer. moreover, sometimes “control” (or management activity) and “organization” are considered together. 242 dmitry novikov:foundations of control methodology notion. let us clarify the last statement. if control is considered as activity of a control subject (principal), then implementing this activity turns out to be a function of a control system; moreover, the control process corresponds to the process of activity, a control action corresponds to its result, etc.[3]. in other words, if a principal and controlled system both represent subjects (see fig.3), then control is activity (of principals) regarding organization of activity (of controlled subjects). therefore, control methodology is the theory of organization of control activity, i.e., the activity of subjects controlling other subjects or objects. one can further increase the level of reflexion (who organizers whose activity). on the one hand, in a multilevel control system the activity of a top manager may be considered as activity regarding organization of activities of his subordinates; in turn, their activity consists in organization of activity of their subordinates, and so on. on the other hand, an army of consultants represent experts in organization of management activity (first of all, the matter applies to management consulting). such consultants regularly operate the term “control methodology” (sometimes, incorrectly and inappropriately). let us take the general formulation of a control problem for a certain system. assume there exist a control subject (a principal) and a controlled system or control object (in terminology of automatic control theory). the state of the controlled system depends on external disturbances, actions of the control subject (control actions) and, probably, on actions of the system, see fig.3 (if the control object appears active). a problem of the control subject consists in performing control actions (see the thick line in fig.3) to ensure a required state of the control object. this is done using information on external disturbances (see the dashed line in fig.3). the so-called subject-object (input-output) structure of a control system is illustrated by fig.3. this is the basic structure used in control theory to study control problems for systems of different nature. fig.3 structure of a control system the primary input-output structure of a control system illustrated by fig.3 advances in systems science and applications (2012) vol.12 no.3 243 bases on the scheme of activity presented by fig.2. the point is that both the control subject and control object carry out the corresponding activity. combining the structure of both sorts of activity according to fig.2, one obtains the structure of control activity illustrated by fig.4. note the following. from agents point of view, the principal is a part of an external environment (numbers of actions in fig.2 and fig.4 coincide), which exerts an influence for a definite purpose (double arrows (1)-(4) and (6) in fig.2), see fig.4. some components of environmental influence may even have a random (nondeterministic) character, and be beyond the principals control. along with actions of the controlled system, these actions exert an impact on the outcome (the state) of the con-trolled system (double arrow (5) in fig.2); see also external disturbances in fig.4. the structure given by fig.4 may be augmented by adding new hierarchical levels. the principles used to describe control in multilevel systems remain unchanged. however, multilevel systems have specifics distinguishing them from a serial combination of two-level “blocks”. fig.4 structural components of control activity for a controlled system, a criterion of operating efficiency depends on its state and (in some cases) on the control actions. the most important feature is whose viewpoint serves a reference in efficiency analysis. suppose that one knows the relationship between the state of a controlled system and control actions applied. hence, it seems possible to consider operating efficiency of a controlled system as a certain function of control actions. such function is referred to as a control 244 dmitry novikov:foundations of control methodology efficiency criterion. consequently, any control problem4 could be formally stated as follows. find feasible control actions ensuring the maximal efficiency (such controls are said optimal). to succeed, one should solve an optimization problem, notably, choose an optimal control (optimal controls). we have provided the general scheme of any human activity. a cycle of activity terminates with achieving a certain result. thus, the efficiency of activity is assessed in the sense of estimating the corresponding result. the presence of a measurable result (otherwise, control makes no sense!) allows for estimating the level of goal attainment as an anticipated, expected result of activity. the efficiency of activity is a degree of conformity between the result and goals of the subject performing the activity. exerting an impact on the components of activity (controlling them), one may influence on the result and efficiency of the activity. control is the activity of a control subject with respect to a controlled system. when a control subject coincides with a controlled system, the matter concerns self-regulation. the result of activity performed by a control subject is defined by his state and the state (the result of activity) of a controlled system. hence, the efficiency of control activity (control efficiency) represents a degree of conformity between the result of operation of a controlled system and goals of a control subject. evaluation of controls ensuring maximal efficiency is the scope of optimization. in fact, optimization consists in finding the best (optimal) alternatives in a set of feasible alternatives under given conditions. let us emphasize the relevance of every word in this statement. using the term “the best alternatives”, we assume there is a certain criterion (several criteria) and a way (ways) to compare alternatives. it is of crucial importance to account for the existing conditions and constraints; varying them leads to optimality of other alternatives under the same criterion (criteria). we have earlier discussed control efficiency. the efficiency being measured, control aims for efficiency optimization (in fact, maximization) under given constraints and conditions. in fact, optimization consists in finding the best (optimal) alternatives in a set of feasible alternatives under given conditions. given the general structure of control activity, lets procedd to its characteristics, logical and temporary structures. 4a problem is something requiring execution or solution; a goal of activity specified in certain conditions. in this book, the term control problem has two meanings. the first (wide) one is searching for an optimal control within the framework of the general model (efficiency maximization as the goal of control activity). the second (narrow) meaning consists in searching for an optimal control of a certain type (e.g., resource allocation problems, operative control problems, etc.). advances in systems science and applications (2012) vol.12 no.3 245 3 characteristics, logical and temporal structures of control activity characteristics, logical and temporal structures of control activity are presented in a summarised form in tables 1-3 correpondingly (see details in[4]). table 1 characteristics of control activity characteristics organization of control activity 1.the personalized nature of control activity; 2.independent goal-setting by a control subject (principal); 3.the mediated outcome of control activity; 4.the creative character of control activity; features of activity 5.the necessity of modeling (predicting, forecasting the behavior of a controlled system under specific control actions); 6.the responsibility of a control subject for the process and result of his her activity and activity of subjects and/or objects controlled by him/her; 7.development and adaptation. principles of activity principles of hierarchy; unification; purposefulness; openness; efficiency; responsibility; non-interference;social and state control; development; completeness and prediction; regulation and resource provision; feedback; adequacy; well-timed control; predictive reflection; adaptivity; rational centralization; democratic control; coordination; ethics . conditions of activity motivational, personnel-related, material and technical, methodical, organizational, financial, regulatory and legal, and informational conditions. norms: 1)general; universal ethical,legal and other norms. 2) specific norms of managerial ethics, organizational culture. table 2 the logical structure of control activity structural components organization of control activity active subject control subject (individual or collective). the object of activity control object and/or controlled subject (individual or collective). 246 dmitry novikov:foundations of control methodology structural components organization of control activity the subject of activity elements of a controlled system, components of activity of a controlled subject. for instance, in the case of organizational systems: staff of the system; structure of the system; constraints and norms of activity of participants; goals and preferences of participants; awareness of participants; the sequence of function-ing. the result of activity state of a control object, result of activity of a controlled subject; consumed resources. the forms of activity organization individual and collective control; unified and personalized control. projectand process-based management; reflectory (situational) and forward-looking control. hierarchical control, distributed control, and network control. the functions of activity in the case of organizational control: planning, organizing, motivating, and controlling. tasks in the case of organizational control: monitoring and analysis of the actual state of a controlled system, forecasting the evolution of the system, goal-setting, planning and distributing the resources, motivation (incentives), control and operative management, analysis and improvement of activity. the methods of activity in the case of organizational systems: staff control; structure control; institutional control (normative control, i.e., control of constraints and norms of activity); motivational control (economic control, i.e., control of preferences); informational control (socio-psychological control, i.e., control of information being available to controlled subjects at the moment of decision making). the means of activity in the case of organizational control: orders, directives, instructions, plans, strategies, policies, norms, standards, procedures, regulations concerning activity organization. advances in systems science and applications (2012) vol.12 no.3 247 structural components organization of control activity mechanisms in the case of organizational control: mechanisms of active expertise, mechanisms of active expertise, transfer pricing mechanisms, mechanisms of contract renegotiation, mechanisms of cost-benefit analysis, mechanisms of institutional control, mechanisms of informational control, integrated rating mechanisms (mechanisms of data aggregation), rankorder tournaments (tenders), multi-channel mechanisms, mechanisms of assignment, mechanisms of exchange, mechanisms of predictive self-control, mechanisms of production cycle optimization, incentive mechanisms for cost reduction, resource allocation mechanisms (including costs and incomes), mechanisms of self-financing, mechanisms of structure choice, mechanisms of staff choice, mechanisms of joint financing, mechanisms of consent, incentive mechanisms, insurance mechanisms, etc. table 3 organizing the process (temporal structure) of control activity temporal structure a control activity cycle phases stages steps 1. design phase 1.1. conceptual stage 1.1.1. identifying contradictions a contradiction between the actual (or forecasted) state of a controlled system and its desired state. 1.1.2.stating a problem a control problem as the need for exerting an impact on activity (state) of a controlled system; such need must be recognized by a control subject. 1.1.3.defining the goal of control. identifying contradictions defining the goals of control as a desired state (result of activity) of a controlled system (in the narrow sense, as a way of organizing of controlled subjects activity). 1.1.4.choosing criteria criteria for describing/assessing the state (result of activity) of a controlled system. control efficiency criteria. 248 dmitry novikov:foundations of control methodology temporal structure a control activity cycle phases stages steps 1.2. modeling stage 1.2.1. constructing a model constructing a model of a controlled system (taking into account its active property if necessary). studying the dependence of the controlled systems state (result of controlled subjects activity) on control actions and the state of an external environment. 1.2.2. optimization solving the problem of optimal control synthesis (for the constructed model of a controlled system). analyzing stability and adequacy of solutions. 1.3. the stage of control planning 1.3.1. decomposing formulating control problems as the goals for specific subproblems ensuring a definite overall goal of control (within the framework of existing constraints). 1.3.2.aggregation coordinating the results of solution of specific control problems, assessing the feasibility of joint application of different methods, means, forms and mechanisms of control. 1.3.3. analyzing the conditions (available resources) analyzing the influence of conditions (resource constraints) on the efficiency of control activity, including resources decomposition by methods, forms, means of control, etc. 1.3.4. making up the program of control identifying the controlled system. choosing conditions, methods, means, forms and mechanisms of control. solving the problem of optimal control synthesis. advances in systems science and applications (2012) vol.12 no.3 249 temporal structure a control activity cycle phases stages steps 1.4.the stage of technological preparations for control 1.4.1. technological preparations detailed elaboration and preparation of necessary conditions, methods, means and forms of control. 2. implem entation phase 2.1. organizing stage implementing conditions, methods, means, forms and mechanisms of control. resources allocation. distributing functions and tasks among elements of a controlled system. 2.2. motivating stage implementing the mechanisms of non-financial and financial incentives of controlled subjects. 2.3. monitoring stage organizing the system of permanent assessment of the activity performed by a controlled subject and/or an external environment. 2.4. the stage of operational management well-timed correction of conditions and mechanisms of control based on monitoring results. 3. reflexiv -e phase the stage of accounting and controlling acquiring information on the results of activity performed by a control subject and a controlled system, results assessment (comparison with posed goals). the stage of activity analysis (results analysis) reflexion as a way of control subjects recognition of his/her activity integrity, as well as of the goals, content, forms, and means of such activity. analyzing the obtained results (taking into consideration resources consumed). the stage of decisions correction in the case of cyclic (repetitive) activity, “local” modification of its content and parameters based on analysis of achieved results. 250 dmitry novikov:foundations of control methodology temporal structure a control activity cycle phases stages steps the stage of activity improvement systematic reviewing of the whole organizational structure of control activity (in particular, efficiency criteria adopted, as well as methods, forms, means and mechanisms of control). 4 philosophical foundations of control methodology a foundation is a sufficient condition of something (entity, cognition, an idea or activity). from the control theory point of view cybernetics and systems analysis are remarkable for occupying the interdisciplinary or overdisciplinary position and may be treated as applied dialectics. within the framework of these approaches, control activity is a complex system intended for preparing, substantiating and implementing solutions to complex problems of different character (e.g., political, social, economic, technical problems, etc.)[6-11]. by comparing the conceptions adopted by different scientific disciplines (viz., philosophy, psychology, sociology and systems analysis or systems engineering), one would easily choose the general structure of activity (see fig.2). but the fundamental foundations for control methodology are given by philosophy. philosophy studies activity as a universal way of human existence. accordingly, humans represent active creatures. human activity covers material-practical, intelligent and spiritual operations, external and internal processes. activity is the behavior of mind just exactly as the behavior of arms, whereas human activity makes up cognition process similarly to human behavior. activity enables an individual to reveal his/her particular place in the world and to assert himself/herself as a social being. having reached a certain level of epistemological maturity, scientists perform “reflexion” by formulating general laws in corresponding scientific fields, i.e., create metasciences. on the other part, any “mature” science becomes the subject of philosophical research. for instance, the philosophy of physics appeared at the junction of the 19th century and the 20th century as the result of such processes[12]. originated in the 1850s, research in the field of control theory5 led to the appearance of other metasciences, i.e., cybernetics[6-7,11] (in the 1950s) and systems analysis[8-10] (later). moreover, cybernetics quickly became the subject of philosophical investigations (e.g., see[11,13]) conducted by “fathers” of cybernetics 5following the established tradition, we will occasionally call control science by control theory (yet, keeping in mind that the name is narrower than the subject). advances in systems science and applications (2012) vol.12 no.3 251 and professional philosophers. the 20th century was accompanied with rapid progress of management science[14-16] as a branch of control theory studying practical control in organizational systems. by the beginning of the 2000s, management science engendered management philosophy. books and papers entitled “management philosophy” appeared exactly at that times (for instance, see[13,15,17]); as a rule, their authors represented professional philosophers. generally speaking, one may acknowledge the long-felt need for more precise mutual positioning of philosophy and control. consider fig.5 illustrating different connections between the categories of philosophy and control; they are treated in the maximal possible interpretation (philosophy includes ontology, epistemology, logic, axiology, ethics, aesthetics, etc.; control is viewed as a science and a type of practical activity). we believe that the three domains shaded in fig.5 are the major ones. control philosophy (as a branch of philosophy). historically (and similarly to the subjects of most modern sciences), control problems analysis was first the prerogative of philosophy. r. descartes was used to say, “philosophy is like a tree whose roots are metaphysics and then the trunk is physics. the branches coming out of the trunk are all the other sciences”. historical and philosophical analysis implies that first control theorists were exactly philosophers. confucius, lao-tzu, socrates, platon, aristotle, n. machiavelli, t. hobbes, i. kant, g. hegel, k. marx, m. weber, a. bogdanovcthis is a short list of philosophers that laid down the foundations of modern control theory for the development and perfection of managerial practice. presently, concrete control problems are no more the subject of philosophical analysis. philosophy (as a form of social consciousness, the theory of general principles of entity and cognition, human attitude to the reality, as the science of universe laws of natural development) studies general prob-lems and laws separated out by experts in certain sciences. by analogy to the notions of “historical philosophy”, “cultural philosophy”, “legal philosophy”, etc. (see philosophical encyclopedias), one can define control philosophy as a branch of philosophy connected with comprehension and interpretation of control processes and control cognition, studying the essence and role of control. such meaning of the term “control philosophy” (see the dashed-line contour in fig.5) has rich internal structure and covers epistemological research of control science, the analysis of logical, ontological, ethical and other foundations (both for control science and management science). the basic goals of research in control philosophy are as follows: 1.identifying the content of control as a science and practical activity, analyzing their subject and place in the system of scientific knowledge; 2.performing the ideological, methodological and logical-epistemological anal252 dmitry novikov:foundations of control methodology fig.5 philosophy and control ysis of primary notions, results, techniques, functions and theories in control science; 3.translating philosophical laws to enrich the content of control laws; 4.involving the achievements of control theory and practice to enrich the content of philosophical categories and laws; 5.substantiating the feasibility and conditions of using common approaches to control problems in systems of interdisciplinary nature, constructing uniform control theory; 6.performing methodological analysis of control with application to different areas of human activity and different classes of control objects; 7.substantiating philosophically the key directions in control theory and practice; 8.systematizing and classifying theories of control; 9.identifying and systematizing axiological dominants in control theory and practice; 10.developing the integrated conceptual framework of control science (including the terminology of all embedded theories). let us formulate a series of “questions” determining perspective directions of research in control philosophy (according to experts in control theory, these issues lie “in the plane” of control philosophy). • what would general laws and regularities studied by philosophy gain for control theory and practice? which modern directions of philosophical research advances in systems science and applications (2012) vol.12 no.3 253 can find (alternatively, have already found) applications in control science (structuralism, post-structuralism, hermeneutics, etc.)? what are the manifestation and influence of general scientific meaningfulness and interdependency of adopted terminology? • what are the epistemological specifics of control science? are there general approaches to the statement and solution of control problems? how does control science position itself in the general system of sciences? what is the epistemological status of a researcher in control theory and practice? • how are basic categories of philosophy (a language, ordinary consciousness, ethics, a law, philosophy, a science, art, a religion, a political ideology, etc.) correlated with that of control science (control, an activity, an organization, decision making)? how is the latter group of categories correlated with other categories (such as a human being, nature, a society, production)? • which laws (features) of control science formation as a metascience can be identified in historical retrospective and at the modern stage of its development? what is the connection between control theory and practice (again, in historical retrospective and in future perspective)? • how does philosophy (as the “quintessence of culture”) affect the formation of “organizational culture” in control theory and practice? what is the interrelation between universal principles, laws and features of development of particular organizational, social and cultural formations in control theory and practice? cybernetics (as a branch of control science, studying its most general theoretical laws). for many scientific disciplines, there exists a range of problems related to their foundations and traditionally referred to as the philosophy of a corresponding science. control science follows this tradition, as well. foundations of control science also include general laws of efficient control (representing the subject of cybernetics). nowadays, one often faces the opinion that cybernetics has become old-fashioned as a scientific discipline and no more pretends to the role of certain universal control science. this is true, but only in part. as a matter of fact, in the middle of the 1940s cybernetics appeared the theory of “control and communication in the animal and the machine” (see the pioneering monograph[11]). furthermore, it originated even as the theory of general laws of control. triumphal advancements of cybernetics during the 1950-1960s (e.g., technical cybernetics, economic cybernetics, biological cybernetics, etc., and well as their close connections to operations research, mathematical theory of control; plus intensive implementation of results in designing new and upgrading existing technical and information systems) created the illusion of the universal character of cybernetics and inevitability of its rapid development in future. however, the evolvement of cybernetics slowed down in the early 1970s. this “integral” science branched out into a set of 254 dmitry novikov:foundations of control methodology partial directions and “mingled with details”; indeed, the number of subbranches grew and all of them showed independent development (almost without identification and systematization of general laws). curiously enough, the only bearers of canonical cybernetic traditions were philosophers, whereas experts in control theory lost their confidence in ample opportunities of cybernetics. things can’t carry on as they are. on the one hand, philosophers vitally need knowledge of the subject (actually, the generalized knowledge). in this context, v. ilin mentioned that “philosophy represents second-rank reflexion; it provides theoretical grounds to other ways of spiritual production. the empirical base of philosophy consists in specific reflections of different types of cognition; philosophy covers not the reality itself, but the treatment of reality in figurative and category-logical forms” (see references in[18]). on the other hand, experts in control theory need “to see the wood for the trees”. hence, one can hypothesize that cybernetics must and would play the role of control philosophy in its second meaning (as a branch of control theory, studying its most general laws). here the emphasis should be made on constructive development of control philosophy, i.e., on formation of its content through obtaining concrete results (probably, first partial results and then general ones). management “philosophy”. a detailed analysis of modern textbooks on management science, sociology and psychology of management separates out the following categories6 used to describe managerial practice (see fig.6). management “philosophy” tops the pyramid demonstrated in fig.6. it reflects the maximally abstracted level of description and consideration of solving the problems of managerial practice. there are intensive discussions regarding the comprehension of management “philosophy”, its subject and main content. for instance, the following opinions are quoted in[18]: -“possibly, management philosophy is the pragmatism, where an essential characteristic of a human being lies in actions, purposeful activity. cognizing exactly the laws of human activity must form the object of management philosophy” (l. bessonova); -“management philosophy considers axiological, epistemological, and methodological foundations of human activity in control processes” (v. diev), and so on. the examples of inhomogeneous definitions could be continued. many authors of textbooks on management science adopt the term “personal management philosophy” (similarly to the existence of numerous opinions regarding necessary qualities of a good leader, there are many different management philosophies). in other words, sometimes management “philosophy” is commonly treated as an6note that the corresponding terms are generally not defined explicitly and addressed somewhat inadvertently (in management science). advances in systems science and applications (2012) vol.12 no.3 255 alyzing the set of qualities of an efficient manager and his/her decisions leading to a success. fig.6 levels and categories of managerial practice description fig.7 control philosophy, cybernetics and management “philosophy” almost all authors agree with the following. management “philosophy” is a system of ideas, views and beliefs of managers about human nature and society, control problems and ethical principles of their behavior (this system forms mostly empirically). yet, we believe such definition appears eclectic and not operational. our approach is to understand management “philosophy” (“the top of 256 dmitry novikov:foundations of control methodology management”) as a branch of control science dealing with generalization of laws of successful managerial practice. we have briefly analyzed the correlation of control philosophy (as a branch of philosophy studying general problems of control theory and practice), cybernetics (as a branch of control science generalizing the methods and results of solving theoretical problems of control) and management (as a branch of control science generalizing the experience of successful managerial practice), see fig.7. 5 conclusion the paper has endeavored to systematize control methodology (as the theory of organizing of control activity). philosophical foundations of the methodology of control activity, its characteristics, as well as the logical and temporal structures were described. a series of “questions” determining perspective directions of research in control philosophy were posed. authors hope that a unified approach (based on control methodology) to the consideration and research of control activity will help to interconnect and develop parallely mathematical control theory and control philosophy. references [1] novikov, a., novikov, d. (2013), research methodology: from philosophy of science to research design, leiden: crc press. [2] novikov, a., novikov, d. (2007), methodology, moscow: sinteg. (in russian). [3] novikov, d. (2013), theory of control in organizations, n.y.: nova science publishers. [4] novikov, d. (2013), control methodology, n.y.: nova science publishers. [5] leontjev, a. (1978), activity, consciousness and personality, prentice: prentice-hall. [6] ashby, w. (1956), an introduction to cybernetics, london: chapman and hall. [7] beer, s. (1995), brain of the firm, 2nd ed. ny: wiley. [8] bertalanffy, l. (1968), general system theory: foundations, development, applications, n.y.: george braziller. [9] mesarovic, m., mako, d. and takahara, y. (1970), theory of hierarchical multilevel systems, new york: academic. advances in systems science and applications (2012) vol.12 no.3 257 [10] peregudov, f. and tarasenko, f. (1993), introduction to systems analysis, oh: columbus, glencoe/mcgraw-hill. [11] wiener, n. (1965), cybernetics: or the control and communication in the animal and the machine, 2nd ed, massachusetts: the mit press. [12] heisenberg, w. (1962), physics and philosophy: the revolution in modern science, new york: harper & row publishers, inc. [13] rosenberg, a. (2000), a philosophy of science: a contemporary introduction, london: routledge. [14] ackoff, r. (1999), classic writings on management, 2nd ed, ny: wiley. [15] drucker, p. (2006), the effective ecutive: the definitive guide to getting the right things done, n.y.: collins business. [16] mintzberg, h. (1983), structure in fives: designing effective organizations, englewood cliffs, nj: prentice-hall. [17] kirkeby, o. (2000), management philosophy: a radical-normative perspective, heidelberg: springer. [18] novikov, d. and rusjaeva, e. (2012), “philosophy and conbtrol”, philosophy problems, no.5, pp.19-26. (in russian). corresponding author author can be contacted at: novikov@ipu.ru. advances in systems science and applications (2014) vol.14 no.3 254-278 mathematical models of informational and strategic reflexion: a survey novikov d.a. and chkhartishvili a.g. institute of control sciences, moscow abstract the paper is dedicated to a survey (in the framework of game theory and theory of collective behavior) of modern approaches to mathematical modeling of reflexive games and reflexive processes in control. keywords mathematical models, informational and strategic reflexion 1 introduction. 1.1 reflexion a fundamental property of human entity lies in the following. in addition to natural (“objective”) reality, there exists its image in human minds. furthermore, an inevitable gap (mismatch) takes place between the latter and the former. in the sequel, the described image will be called a part of reflexive reality. traditionally, purposeful study of this phenomenon relates to the term “reflexion”. the term reflexion (from latin reflex ‘bent back’; was first suggested by j. locke) means [1]: • principle of human thinking, guiding humans towards comprehension and perception of ones own forms and premises; • subjective consideration of a knowledge, critical analysis of its content and cognition methods; • the activity of self-actualization, revealing the internal structure and specifics of spiritual world of a human. to elucidate the whole essence of reflexion, let us consider the case of a single subject. he/she possesses certain beliefs about natural reality; however, a subject may perform reflexion (construct images) with respect to these beliefs (thus, generating new beliefs). generally, this process is infinite and results in formation of reflexive reality. the reflexion of a subject with respect to his/her own beliefs of reality, principles of his/her activity, etc., is said to be self-reflexion or reflexion of the first kind. we emphasize that most social research works concentrate on self-reflexion. in philosophy, self-reflexion represents the process of individuals thinking about beliefs in his/her own mind [2]. reflexion of the second kind takes place with respect to other subjects (includes beliefs of a subject about possible beliefs, decision principles and self-reflexion of other subjects). 1.2 reflexion and control control is an element, a function of organized systems of different nature (biological, social, technical, etc.), preserving their definite structure, sustaining their advances in systems science and applications (2014) vol.14 no.3 255 mode of activity and implementing the program or goal of their activity; control is a purposeful impact exerted on a controlled system to ensure its required behavior [3]. assume there is a control subject (a principal) and a controlled system (control object-in terminology of technical systems-or a controlled subject). the state of a controlled system depends on external disturbances, control actions applied by a principal and possibly on actions performed by the controlled system (if the latter represents an active subject), see fig. 1. the principals problem lies in choosing control actions (see the thick line in fig.1) to ensure the required behavior of a controlled system taking into account information on external disturbances (see the dashed line in fig.1). the so-called input-output structure of a control system (fig.1) is typical for control theory dealing with control problems in systems of different nature. the presence of feedback (see the double line in fig.1) which provides a principal with information on the state of a controlled system is the key (but not compulsory!) property of a control system. some researchers interpret feedback as reflexion (as an image of the controlled systems state in the “mind” of a control subject). this forms the first aspect of interrelation between control and reflexion. fig.1 the structure of a control system a series of scientific directions investigate the interaction and activity of a control subject and controlled system. control science (or control theory in the terminology of corresponding experts) mostly focuses on the interaction between a control subject and controlled system. control methodology [4] is the theory of organizing of control activity, i.e., the activity performed by a control subject. we emphasize that activity can be mentioned only with respect to active subjects (e.g., a human being, a group, a collective). in the case of passive (e.g., technical) systems, the term “functioning” is used instead. in the sequel, we believe that a control subject and controlled system appear active (otherwise, there is a clear provision for the opposite). hence, each of them may perform (at least) 256 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... self-reflexion, constructing “images” of the process, organization principles and results of his/her own activity. this is the second aspect of interrelation between control and reflexion. searching for optimal control (i.e., the most efficient admissible control) requires control subjects ability of predicting controlled systems response to certain control actions. one of prerequisites is a model of a controlled system. generally speaking, a model is an image of a certain system; an analog (a scheme, a structure or a sign system) of a certain fragment of the natural or social reality, a “substitute” for the original in cognition process and practice. a model can be considered as an image of a controlled system in the mind of a control subject. modeling (as a process of “reflecting”, i.e., constructing this image) can be viewed as reflexion. furthermore, a controlled system may predict and assess the activity performed by a control subject. and so, we obtain the third aspect of interrelation between control and reflexion. the fourth aspect lies in the following. a control subject or controlled system performs reflexion with respect to external subjects and objects, phenomena or processes, their properties and laws of activity/functioning. for instance, the matter concerns an external environment (for a control subject), an external environment and/or other elements of a controlled system (for a fixed element of a controlled system). indeed, suppose that a controlled system includes several active agents; each of them may perform reflexion with respect to the others. exactly this aspect-mutual reflexion of controlled subjects-is discussed in game-theoretical models. of crucial importance here is that the process and/or result of reflexion can be controlled, i.e., can represent a component of controlled systems activity, being modified by a control subject for a definite goal. precisely this relationship between control and reflexion enables informational control and reflexive control, considered below. 1.3 game theory formal (mathematical) models of human behavior have been constructed and studied for over last 150 years. gradually, these models find wider application in control theory, economics, psychology, sociology, etc., as well as in practical problems. in the sequel, we will understand a game as the interaction of subjects with noncoinciding interests. still, an alternative interpretation treats a game as a type of unproductive activity whose motive consists not in the corresponding results, but in the process of activity itself (see [2, 5], where the notion of a game is assigned a broader sense). game theory represents a branch of applied mathematics, which analyzes models of decision making in the conditions of noncoinciding interests of opponents (players); each player strives for influencing the situation in his/her favor [6-7]. advances in systems science and applications (2014) vol.14 no.3 257 in what follows, a decision-maker (a player) is called an agent. the major task of game theory is describing the interaction among several agents with noncoinciding interests, where the results of agents activity (payoff, utility, etc.) generally depend on actions of all agents. such description yields a forecast of a rational and “stable” outcome of the game-the so-called game solution (equilibrium). describing a game means specifying the following parameters: a set of agents; preferences of agents (relationships between payoffs and actions). each agent is supposed to strive for maxi-mizing his/her payoff (and so, the behavior of each agent appears purposeful); a set of feasible actions of agents; awareness of agents (information on essential parameters, being available to agents at the moment of their choice); sequence of moves (the sequence of obtaining information and choosing actions). the above parameters define a game; unfortunately, they are insufficient for forecasting its outcome, i.e., a solution (or an equilibrium) of the game-the set of rational and stable actions of agents. nowadays, game theory suggests no universal concept of equilibria. by adopting different assumptions regarding principles of agents decision making, one can construct different solutions. thus, designing an equilibrium concept forms a basic problem for any game-theoretic research; this book does not represent an exception, as well. reflexive games are defined as a direct interaction among agents, where they make decisions based on hierarchies of their beliefs. in other words, awareness of agents is extremely important. 1.4 the role of awareness. common knowledge in game theory, psychology, distributed systems and other fields of science (see the overviews in [8-9]), one should consider not only agents beliefs about essential parameters, but also their beliefs about the beliefs of other agents, etc. the set of such beliefs is called the hierarchy of beliefs. we will model it using the tree of awareness structure of a reflexive game (see below). in other words, situations of interactive decision making (modeled in game theory) require that each agent “forecasts” opponents behavior prior to his/her choice. and so, each agent should possess definite beliefs about the view of the game by his/her opponents. on the other hand, opponents should do the same. consequently, the uncertainty regarding the game to-be-played generates an infinite hierarchy of beliefs of game participants. a special case of awareness concerns common knowledge when beliefs of all orders coincide. a rigorous definition of common knowledge was introduced in [10]. notably, common knowledge is a fact with the following properties: 1 ) all agents know it; 258 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... 2 ) all agents know 1; 3 ) all agents know 2 and so on-ad infinitum. the formal model of common knowledge was originally proposed in [11]. later on, many investigators refined and redeveloped it see surveys and references in [1,12-18]. the present paper is almost completely dedicated to models of agents awareness in game theory (viz., hierarchies of beliefs and common knowledge). thus, we give several references demonstrating the role of common knowledge in different fields of science-philosophy, psychology, etc. (see also the overview in [19]). in philosophy, common knowledge has been studied in convention analysis [10, 20]. in psychology, one would face the notion of discourse (from latin discursus ‘argument’). it means human thinking in words, being mediated by past experience; discourse acts as the process of connected logical reasoning, where a next idea stems from the previous one. the importance of common knowledge in discourse comprehension has been explored in [19, 21]. mutual awareness of agents turns out significant in distributed computer systems [13, 15, 22], artificial intelligence [23-24] and other fields. game theory often assumes that all1 parameters of a game are a common knowledge. such assumption corresponds to the objective description of a game and enables addressing the nash equilibrium2 concept [25] as a forecasted outcome of a noncooperative game (a game, where agents do not agree about coalitions, data exchange, joint actions, redistribution of payoffs, etc.). thus, the assumption regarding common knowledge allows claiming that all agents know which game they play and that their beliefs about the game coincide. generally, each agent may possess individual beliefs about parameters of a game. and so, each belief corresponds to a subjective description of the game [6] (see also modern models of awareness in [26-29]). consequently, agents participate in the game, having no objective views of it or interpreting this game in different ways (rules, goals, the roles and awareness of opponents, etc.). unfortunately, still no universal approaches have been proposed for equilibria design under insufficient common knowledge. on the other part, within the “reflexive tradition” of the humanities, the surrounding world of each agent includes the rest agents; moreover, beliefs about other agents get reflected during the process of reflexion (in particular, variations of beliefs may result from nonidentical awareness). however, researchers have not succeeded in deriving constructive formal outcomes in this field to date. 1if the initial model incorporates uncertain factors, specific procedures of uncertainty elimination are involved to obtain a deterministic model. 2an agents action vector is a nash equilibrium if none of them benefits by unilateral deviation from it (provided that the rest agents choose the corresponding components of the nash equilibrium). a more rigorous definition could be found below. advances in systems science and applications (2014) vol.14 no.3 259 hence, an urgent problem lies in designing and analyzing mathematical models of games, where agents awareness is not a common knowledge and agents make decisions based on hierarchies of their beliefs. such class of games is called reflexive games [30-32]. we will provide a formal definition later. the term “reflexive games” was introduced by v. lefebvre in 1965, see [33]. however, the cited work and his other publications [34-37] represented qualitative discussions of reflexion effects in interaction among subjects (actually, no general concept of solution was suggested for this class of games). similar remarks apply to [38-41], where a series of special cases of players awareness was studied. the monograph [32] concentrated on systematical treatment of reflexive games and an endeavor of constructing a uniform equilibrium concept for these games. according to game theory and reflexive models of decision making, it seems reasonable to distinguish between strategic reflexion and informational reflexion. informational reflexion is the process and result of agents thinking about (a) the values of uncertain parameters and (b) what his/her opponents (other agents) know about these values. here the “game” component actually disappears-an agent makes no decisions. strategic reflexion is the process and result of agents thinking about which decision making principles his/her opponents (other agents) employ under the awareness assigned by him/her via informational reflexion. therefore, informational reflexion often relates to insufficient mutual awareness, and its result serves for decision making (including informational reflexion). strategic reflexion takes place even in the case of complete awareness, precessing agents choice of an action. in other words, informational and strategic reflexion can be studied independently, but the both occur in the case of incomplete or insufficient awareness. 1.5 general approaches to the description of information and strategic reflexion according to [12, 42-43], there are two different approaches to the description of awareness structures, viz., syntactic and semantic ones. recall that syntactics means syntax of sign systems, i.e., the structure of sign combinations and rules of their formation, “translation”and interpretation irrespective of their values and functions of sign systems. semantics studies sign systems as tools of meaning expression; here the basic subject lies in interpretations of signs and sign combinations. foundations of these approaches were laid in mathematical logic [44-45]. within the framework of syntactic approach, an hierarchy of beliefs is described explicitly. suppose that beliefs are defined by a probability distribution. then hierarchies of beliefs (at a certain level) correspond to distributions on the product of the set of states of nature and distributions reflecting beliefs of preceding levels [46]. an alternative is using “logic formulas”-rules of transforming elements of an initial set based on logic operations and operators such as “player i believes the probability of event . . . is not smaller than α” [43, 47]. a knowledge is modeled 260 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... by propositions (formulas) constructed according to certain syntactic rules. according to semantic approach, beliefs of agents are defined by probability distributions on the set of states of nature. hierarchies of beliefs get generated only by virtue of these distributions. in the elementary (deterministic) case, a knowledge represents the set θ of feasible values of an uncertain parameter and different partitions {pi}i∈n of this set. an element of the partition pi containing θ ∈ θ forms the knowledge of agent i, namely, the set of values of the uncertain parameter, being indistinguishable for this agent under a known fact θ [11-12]. the correspondence (or “equivalence”) between syntactic and semantic approaches was established in [18, 42] and other works. we also cite experimental research on hierarchies of beliefs [48-50]; see the surveys in [51-52]. the above overview points at two existing “extremes”. the first one lies in common knowledge. here j. harsanyis merits [53] are (a) reducing all information on an agent (determining the latters behavior) to a single characteristic-agent‘s type-and (b) constructing a bayes-nash equilibrium by hypothesizing that the probability distribution of types is a common knowledge. the second “extreme” relates to infinite hierarchy of compatible or incompatible beliefs. (for an example, see the structure discussed in [46]. on the one hand, it describes all possible bayesian games and all possible hierarchies of beliefs. on the other hand, it appears very general and, consequently, very cumbersome, thus interfering with constructive statement and solution of specific problems). most research on awareness seeks to answer the following question. when does an hierarchy of agents beliefs describe a common knowledge and/or reflect adequately their awareness? [19, 54]. the dependence of game solutions on a finite hierarchy of compatible or incompatible beliefs of agents (the whole range between the above “extremes”) has been studied in [1, 32]. 1.6 theory of collective behavior traditionally, game-theoretic models and/or models of collective decision making utilize one of two assumptions regarding mutual awareness of agents [51]. the first one implies that all essential information and decision principles adopted by agents are known to all agents, all agents know this fact and so on (such reasoning could be infinite). actually, this is the concept of a common knowledge, which serves, e.g., in constructing a nash equilibrium. the second assumption claims that each agent (according to his/her awareness) follows a certain procedure of individual decision making and has “almost no idea” of the knowledge and behavior of the rest agents. the first approach appears canonical in game theory, while the second approach has become popular in models of collective behavior. yet, a variety of intermediate situations exists between these “extreme cases”. imagine that informational reflexion takes no place-a common knowledge on essential external parameters is observed. let an agent have performed an advances in systems science and applications (2014) vol.14 no.3 261 act of strategic reflexion, i.e., an attempt to predict the behavior of other agents (not their awareness but decision principles). this agent chooses his/her actions using the forecast (we believe he/she possesses reflexion rank 1). another agent (with reflexion rank 2) possibly knows about the existence of agents having reflexion rank 1. consequently, such agent endeavors to predict their behavior, as well. again, this line of reasoning could be infinite. a series of questions arises immediately. how does the behavior of a collective of agents depend on their distribution by reflexion rank (the number of agents with a specific rank in a collective)? suppose that the shares of reflexing agents can be controlled. what are the optimal values of these shares? here optimality is “measured” in terms of some criterion defined on the set of agents actions. classic game-theoretic models proceed from the following. in a normal form game, agents choose nash equilibrium actions. however, investigations in the field of experimental economics indicate this not always the case (e.g., see [55] and the overview [56]). the divergence between actual behavior and theoretical expectations has several explanations: limited cognitive capabilities of agents [57] (decentralized evaluation of a nash equilibrium represents a cumbersome computational problem [58]). furthermore, sometimes nash equilibria provide no adequate description to the real behavior of agents in experimental single stage games (agents have not enough time for “correcting” their wrong beliefs about essential parameters of a game [59]). for instance, d. bernheims concept of rationalizable strategies requires unlimited rationality from agents (their high cognitive capabilities); agents full confidence in that all the opponents would evaluate a nash equilibrium; incomplete awareness; the presence of several equilibria. therefore, there exist at least two foundations (“theoretical” and “experimental” ones) for considering models of collective behavior of agents with different reflexion ranks. in contrast to game theory, the theory of collective behavior analyzes the behavior dynamics of rational agents under rather weak assumptions regarding their awareness. for instance, far from always agents need a common knowledge about the set of agents, sets of feasible actions and goal functions of opponents. alternatively, agents may not predict the behavior of their opponents (as in game theory). moreover, making decisions, agents may “know nothing about the existence of” specific agents or possess aggregated information about them. the most widespread model of collective behavior dynamics is the model of indicator behavior (see references in [51]). the essence of the model consists in the following. suppose that at instant t each agent observes the actions of all 262 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... agents {xt−1 i }i∈n that have been chosen at the preceding instant t− 1, t = 1, 2, . . . the initial action vector x0 = (x01, . . . , x 0 n) is assumed known. each agent can evaluate his/her current goal -an action maximizing his/her goal function provided that at a current instant all agents choose the same actions as at the previous instant: wi({xt−1 −i }) = argmax y∈ℜ1 fi(y, x t−1 −1 ), t = 1, 2, . . . , i ∈ n (1) according to the hypothesis of indicator behavior, at each instant an agent makes a “step” from his/her previous action to the current goal: xti = xt−1 i + γti [wi(x t−1 i )− xt−1 i ], i ∈ n, t = 1, 2, . . . (2) where γti ∈ [0; 1] designate “the values of steps”. for convenience, such collective behavior can be called “optimization behavior” (thus, we emphasize its difference from play behavior). the approaches adopted by the theory of collective behavior and game theory agree in the following sense. the both study the behavior of rational agents, while game equilibria generally represent equilibria for dynamic procedures of collective behavior. for instance, the nash equilibrium specifies an equilibrium for the dynamics (2) of collective behavior. to make the picture complete, note one more aspect, as well. the theory of collective behavior proposes another approach (going beyond the scope of this book), namely, evolutionary game theory [60]. this science studies the behavior of large homogeneous groups (populations) of individuals in typical repeated conflicts; each strategy is applied by a set of players, whereas a corresponding goal function characterizes the success of specific strategies (instead of specific participants of such interaction). thus, game theory often employs maximal assumptions regarding agents awareness (e.g., the hypothesis of existing common knowledge), while the theory of collective behavior involves the minimal assumptions. the intermediate position belongs to reflexive models. and so, let us discuss the role of (informational and strategic) reflexion in decision making by agents. 2 mathematical models of informational and strategic reflexion 2.1 reflexion in game theory and models of collective behavior: the structure of problem domain game theory and the theory of collective behavior analyze interaction models for rational agents. approaches and results of these theories can be considered at three interconnected epistemological levels (that correspond to different functions of modeling [2])-see fig.2 [1]: -phenomenological level, where a model aims at describing and/or explaining advances in systems science and applications (2014) vol.14 no.3 263 fig.2 descriptive and normative models of informational and strategic reflexion table 1 modeling of informational and strategic reflexion: a comparison of approaches parameter informational reflexion strategic reflexion parameter informational reflexion strategic reflexion model of a “game” awareness structure reflexive structure equilibrium information equilibrium reflexive equilibrium control information control reflexive control the behavior of a system (a collective of agents); predictive level (the aim is forecasting the system behavior); normative level (the aim is ensuring a required system behavior). in game theory, a common scheme consists in (1) describing the “model of a game” (phenomenological level), (2) choosing an equilibrium concept defining the stable outcome of a game (predictive level) and (3) stating a certain control problem-find values of controlled “game parameters” implementing a required equilibrium (normative level). an interested reader would find the corresponding illustration in fig.2. taking into account informational reflexion leads to the necessity of constructing and analyzing awareness structures. this enables defining an informational equilibrium, as well as posing and solving informational control problems-see fig.2. taking into account strategic reflexion generates a similar chain marked by heavy lines in fig.2: “models of strategic reflexion” “reflexive structure” 264 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... “reflexive equilibrium” “reflexive control”. a comparison of approaches to modeling of informational and strategic reflexion is given by table 1. 2.2 awareness structure and informational equilibrium consider the set of agents: n = {1, 2, ..., n}. denote by θ ∈ θ the uncertain parameter (we believe that the set θ is a common knowledge for all agents). the awareness structure ii of agent i includes the following elements. first, the belief of agent i about the parameter θ ; denote it by θi, θi ∈ θ. second, the beliefs of agent i about the beliefs of the other agents about the parameter θ; denote them by θij , θij ∈ θ, j ∈ n . third, the beliefs of agent i about the beliefs of agent j about the belief of agent k; denote them by θijk, θijk ∈ θ, j, k ∈ n . and so on (evidently, this reasoning is generally infinite). in the sequel, we employ the term “awareness structure”, which is a synonym of “informational structure” and “hierarchy of beliefs”. therefore, the awareness structure ii of agent i is specified by the set of values θij1...jl , where l runs over the set of nonnegative integer numbers, j1, ..., jl ∈ n , while θi1...il ∈ θ. the awareness structure i of the whole game is defined in a similar manner; in particular, the set of the values θi1...il is employed, with l running over the set of nonnegative integer numbers, j1, . . . , jl ∈ n , and θij1...jl ∈ θ. we emphasize that the agents are not aware of the whole structure i; each of them knows only a substructure ii. thus, an awareness structure is an infinite n-tree; the corresponding nodes of the tree describe specific awareness of real agents from the set n , and also phantom agents (complex reflexions of real agents in the mind of their opponents). a reflexive game γi is a game defined by the following tuple: γi = {n, (xi)i∈n , fi(·)i∈n , i} (3) where n stands for a set of real agents, xi means a set of feasible actions of agent i, fi(·) : θ × x ′ → ℜ1 is his/her goal function (i ∈ n); θ indicates a set of feasible values of the uncertain parameter and i designates the awareness structure. therefore, a reflexive game generalizes the notion of a normal-form game (det -ermined by the tuple {n, (xi)i∈n , fi(·)i∈n , i} ) to the case when agents’ awaren -ess is reflected by an hierarchy of their beliefs (i.e., the awareness structure i). within the framework of the accepted definition, a “classical” normal-form game is a special case of a reflexive game (a game under a common knowledge among the agents). consider the “extreme” case when the state of nature appears a common knowledge; for a reflexive game, the solution concept (proposed in this book based on an informational equilibrium, see below) turns out equivalent to advances in systems science and applications (2014) vol.14 no.3 265 the nash equilibrium concept. to proceed and formulate a series of definitions and properties, we introduce the following notation:∑ + stands for a set of finite sequences of indexes belonging to n ;∑ is the sum of ∑ + and the empty sequence; |σ| indicates the number of indexes in the sequence σ ∈ ∑ (for the empty sequence, it equals zero); this parameter is known as the length of an index sequence. imagine θi represents the belief of agent i about the uncertain parameter, while θii means the belief of agent i about his/her own belief. it seems then natural that θii = θi. in other words, agent i is well-informed on his/her own beliefs. moreover, he/she assumes that the rest agents possess the same property. formally, this means that the axiom of self-awareness is accepted: ∀i ∈ n,∀τ, σ ∈ n : θτiiσ = θtiσ. in particular, being aware of θτ for all τ ∈ ∑ + such that |τ | = γ, an agent may explicitly evaluate θτ for all τ ∈ ∑ + with |τ | < γ. in addition to the awareness structures ii(i ∈ n), one may also analyze the awareness structures iij (i.e., the awareness of agent j according to the belief of agent i), iijk, and so on. let us identify the awareness structure with the agent being characterized by it. in this case, one may claim that n real agents (i − agents, where i ∈ n) having the awareness structures ii also play with phantom agents (τ − agents, where τ ∈ ∑ + , |τ | ≥ 2) having the awareness structures iτ = {θτσ}, σ ∈ ∑ . it should be emphasized that phantom agents exist merely in the minds of real agents; still, they have an impact on their actions; these aspects will be discussed below. assume that the awareness structure i of a game is given; this means that the awareness structures are also defined for all (real and phantom) agents. within the framework of the hypothesis of rational behavior, the choice of an action xτ performed by a τ − agent is described by his/her awareness structure iτ . hence, the mentioned structure being available, one may model agents reasoning and evaluate his/her action. on the other hand, while choosing his/her action, the agent models actions of the rest agents (i.e., performs reflexion). therefore, estimating the game outcome, we should account for the actions of real and phantom agents. a set of actions x∗τ , τ ∈ ∑ + , is called an informational equilibrium, if the following conditions are met: 1. the awareness structure i possesses finite complexity v [30]; 2. ∀λ, µ ∈ σ : iλi = iµi ⇒ xλi ∗ = xµi ∗; 3. ∀i ∈ n, ∀σ ∈ σ: x∗σi ∈ arg max xi∈xi fi(θσi, x ∗ σi1, ..., x ∗ σi,i−1, xi, x ∗ σi,i+1..., x ∗ σi,n) (4) 266 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... here condition 1 claims that a reflexive game involves a finite number of real and phantom agents (what happens when this assumption is rejected, is discussed in [61]). condition 2 expresses the requirement that the agents with an identical awareness choose identical actions. finally, condition 3 reflects rational behavior of agentsceach agent strives for maximizing the individual goal function via a proper choice of his/her action. for this, an agent substitutes actions of the opponents into his/her goal function; the actions are rational in the view of the considered agent (according to the available beliefs of the rest agents). the “classical” concept of a nash equilibrium is remarkable for its self-sustained nature. notably, assume that a repeated game takes place and all agents (except agent i) choose the same equilibrium actions. then agent i benefits nothing by deviating from his/her equilibrium action; evidently, this feature is directly related to the following. beliefs of all agents about reality are adequate, i.e., the state of nature appears a common knowledge. generally speaking, the situation may change in the case of an informational equilibrium. indeed, after a single play of the game some agents (or even all of them) may observe an unexpected outcome due to an inconsistent belief about the state of nature (or due to an inadequate awareness of opponents beliefs). anyway, the self-sustained nature of the equilibrium is violated; actions of agents may change as the game is repeated. informational equilibrium is stable [62], if each agent (real or phantom) observes exactly the expected result (in this case agents awareness does not change). some models of awareness dynamics are considered in [63]. 2.3 informational control the model of informational control (purposeful impact on agents awareness to form informational structure which leads to the desired informational equilibrium) includes an agent (or several agents) and a principal. each agent is characterized by the cycle “awareness of the agent → action of the agent → result observed by the agent → awareness of the agent”. generally speaking, these components vary for different agents. at the same time, the cycle could be viewed common for the whole controlled subsystem (i.e., for the complete set of agents). this feature is indicated by the word “agent(s)” in fig.3. the interaction between an agent (agents) and the principal is characterized by the following elements: an informational impact of the principal, which forms a certain awareness of an agent (agents). it seems possible to study the principals influence on the outcome observed by an agent (agents), see the chain “principal → observed outcome” in fig.3; an actual outcome of the agents action (or agents actions), which has an impact on the preferences of the principal. advances in systems science and applications (2014) vol.14 no.3 267 fig.3 the model of informational control implementing informational control, the principal (as usual) strives to maximize his/her utility. assume the principal can form any awareness structure from a certain feasible set. the problem of informational control may be posed as follows. find an awareness structure from the set of feasible structures, which maximizes the principals utility in a corresponding informational equilibrium (perhaps, taking into account the principals costs to form such an awareness structure). define the following objects: the set ψx(i) ⊆ x ′ of the action vectors of real agents, representing equilibria under the awareness structure i; and the set ψi(x) of awareness structures, making the action vector x of real agents an equilibrium (solution to the inverse problem). let us give a formal statement to the control problem. assume that the goal function of the principal, φ (x, i), is defined on a set of real agents actions and awareness structures. next, suppose that the principal can form any awareness structure from a certain set ℑ′. under the awareness structure i ∈ ℑ′, the action vector of real agents is an element of the set of equilibrium vectors φx (i). we emphasize that the set ψx(i) may be empty; in the case of a missed equilibrium, the principal cannot predict the outcome of a game. to avoid this problem, introduce the set of feasible structures leading to the non-empty set of equilibria: ℑ = {i ∈ ℑ′|ψx (i) ̸= ∅}. imagine that, under the specified awareness structure i ∈ ℑ, the set of equilibrium vectors ψx(i) includes (at least) two elements. as a rule, one of the following assumptions is then adopted [3]: 1) the hypothesis of benevolence(hb), which implies that agents always choose 268 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... the equilibrium desired by the principal; 2) the principle of maximal guaranteed result(pmgr), i.e., the principal expects the worst-case equilibrium of the game. using either the hb or the pmgr, one has the problem of informational control in two settings as follows: max x∈ψx(i) φ(x, i) −−→ i∈ℑ max; (5) min x∈ψx(i) φ(x, i) −−→ i∈ℑ max; (6) naturally, if for any i ∈ ℑ the set ψx(i) consists of a single element, formulas (1) and (2) coincide. in the sequel, the problem (5) (alternatively, (6)) is called the informational control problem in the form of the goal function. now, provide an alternative formulation to the problem of informational control (being independent from the goal function of the principal). assume that the principal wants agents choose an action vector x ∈ x ′. the question arises, “for which vectors and by which awareness structure i would the principal achieve this?” in other words, the second possible formulation of the informational control problem is to find the following components. first, the attainability set, viz, the one composed of the vectors x ∈ x ′ such that for each of them the set of awareness structures ψi(x) ∩ ℑ is nonempty (7) or consists of a single element. (8) second, the corresponding feasible awareness structures i ∈ ψi(x)∩ℑ, meeting the above property for each vector x. note that the condition (7) “corresponds” to the hb, while the one of (8) “corresponds” to the pmgr. the problem (7) (alternatively, (8)) will be referred to as the problem of informational control in the form of the attainability set. once again, we underline that the second formulation of the problem does not depend on the goal function of the principal. it merely reflects the possibility of bringing the system to a certain state by informational control. methods and examples of problems (5)-(8) solution are described in [1, 64]. particular case is the concordant informational control ; here agents are informed about the fact of control implementation by a principal, and still they trust messages of the principal. evidently, implementing such control requires specific conditions [65]. advances in systems science and applications (2014) vol.14 no.3 269 2.4 reflexion structure and reflexion equilibrium publications on strategic reflexion models, the so-called level k models, appeared in the mid-1990s [66-68]. in 2004, they were generalized by the cognitive hierarchies model(chm) [48]. the survey [56] identified four basic approaches to the construction and study of strategic reflexion theoretical models within the framework of game theory and experimental economics (see also experimental results in [48, 69-72] ). we cite fundamental works only (references to later research can be found in [56]). notably, the four basic approaches are: the level k approach [73]; the approach of quantal best response equilibria [74]; the quantal level k approach [50]; the approach of cognitive hierarchies [48]. all of this approaches are mainly generalized by the following model. the hypothesis of indicator behavior implies that choosing his/her actions by the procedure (2), an agent does not ponder over that the rest agents act similarly. otherwise, an agent would perform reflexion and (making decisions at a time instant ) seek for the best response to the actions of the rest agents, forecasted according to (2). in this case, the state of goal is no more defined by formula (1). instead, we obtain wi(x t −i) = argmax y∈ℜ1 fi(y, x t −i) (9) here xt−i satisfies (1). we will believe that a reflexing agent of rank 1 considers the rest agents as non-reflexing. similarly, it is possible to consider agents with higher reflexion ranks (the term “an agent of reflexion rank k” possesses many synonyms (a step k player, a level k player, a k-level player, a smart k-player, etc c see [75-76] and the survey [51]). for this, define ℵ = {n0, n1, . . . , nm} as a partition of the agents set n , where ni is the set of agents with reflexion rank i, i = 0, m , and m specifies the maximal reflexion rank, ni = |ni| , i ∈ n, m∑ i=0 ni = n. we will call ℵ a reflexive partition [77]. suppose that an agent with reflexion rank k exactly knows the sets (shares) of the agents with ranks k′ < k − 1. moreover, assume that he/she considers all agents as having reflexion rank k − 1. in other words, this agent does not concede the existence of agents with the same (or even higher) reflexion rank than his/her rank. in addition, the agent in question may incorrectly estimate the sets of agents possessing reflexion ranks k − 1, k, .... consider a given initial action vector of the agents. let us study the following dynamic reflexive model of their decision making. the corresponding expressions for the one-step “game” model represent a special case, when decisions are made 270 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... one-time under γ1i ≡ 1, i ∈ n . reflexion rank 0 take agents with reflexion rank 0 (belonging to the set n0. assume that they choose actions, thinking that the rest agents act similarly to the previous period. formula (1) yields xti = xt−1 i + γti [wi(x t−1 −i )− xt−1 i ], i ∈ n0, t = 1, 2, . . . (10) n the case n0 = n (no reflexing agents), all agents observe the real trajectory (x0, . . . , xt, . . . ) of the agents action vectors, see (10). reflexion rank 1 x1tj = x1t−1 j + γtj [wj(x t −j)− x1t−1 j ], j ∈ n1 (11) for agent j ∈ n1, the forecasted trajectory is defined by (x0, . . . , (x1tj , x t −j), . . .); however, actually the trajectory (x0, . . . , (x1tj∈n1 , xti∈n0 ), . . .) is realized. this means that the real trajectory may differ from the forecasted trajectories of agents with reflexion ranks 0 and 1 [1]. reflexion rank 2 suppose that each agent j with reflexion rank 2 (j ∈ n2) exactly knows the set n0; moreover, he/she considers all agents from the set n1 ∪ n2\{j} as having reflexion rank 1. in the general case of several agents with reflexion rank 2, this agent wrongly assigns rank 1 to them. consequently, he/she can “forecast” the behavior of the opponents. therefore, his/her choice is the best response to the expected outcome: x2tj = x2t−1 j + γtj [wj(x t−1 i∈n0 , x1tl∈n1∪n2\{j})− x2t−1 j ], j ∈ n2 (12) for agent j ∈ n2, the forecasted trajectory is given by (x0, . . . , (x2tj , x1 t l∈n1∪n2\{j}, xti∈n0 ), . . .), while actually the trajectory (x0, . . . , (x2tj∈n2 , x1tl∈n1 , xtl∈n0 ), . . .) is realized. reflexion rank k(k ≤ m) the behavior of agents with reflexion rank k is described by analogy to the three cases above (reflexion ranks 0, 1 and 2). this is done on the basis of the following awareness structure of the agents. for agent j with reflexion rank k, denote by ℵjk the subjective reflexive partition (the beliefs of the agent about the partitions of all agents): ℵjk = (n0, n1, . . . , nk−2, nk−1 ∪nk ∪ . . . ∪nm\{j}︸ ︷︷ ︸ k , {j}, ∅, . . . , ∅︸ ︷︷ ︸ m−k−1 ), j ∈ nk (13) an agent with reflexion rank chooses actions by the procedure xktj = xkt−1 j +γtj [wj(x t l∈n0 , x1tl∈n1 , . . . , x[k−1]tl∈nk−1∪nk∪...∪nm\{j})−xkt−1 j ], j ∈ nk (14) advances in systems science and applications (2014) vol.14 no.3 271 in the “static” case, this agent selects the action xk∗j (ℵjk) = argmax y∈ℜ1 fj(y, x 1 l∈n0 , x11l∈n1 , . . . , x[k−1]1l∈nk−1∪nk∪...∪nm\{j}), j ∈ nk. (15) therefore, a reflexive structure represents the set of subjective reflexive partitions of all agents. assume that agents beliefs about the reflexion ranks of each other satisfy (13). then the awareness structure is uniquely defined by the reflexive partition ℵ. the vector of agents actions x∗(ℵ) = {xk∗j (ℵjk)}j∈nk, k=0,m (16) is said to be a reflexive equilibrium of the game γℵ = {n,fi(·)i∈n ,ℵ} [51, 77]. in other words, a reflexive equilibrium forms the set of agents actions being the best responses to opponents actions (according to an existing reflexive structure). by virtue of the assumptions regarding the existence and uniqueness of best responses, a reflexive equilibrium always exists. furthermore, a reflexive equilibrium seems rather exotic. generally, the actions of agents are not the best responses to opponents actions. detailed classification of strategic reflexion models is given in [51, 77]. the described general model of reflexive collective behavior would hardly lead to general analytical derivations. nevertheless, it may provide a basis for developing particular analytical models or general simulation models (e.g., according to the classification suggested in [78]). such models serve for describing and forecasting collective behavior (human beings, mobile robots, program agents) in various situations. for instance, we refer an interested reader to [1] for reflexive simulation models of evacuation, reflexive models of transport flows and other numerous examples from different applications. by proper variation of reflexive partitions, one can change the actions of agents, i.e., perform reflexive control [1, 77]. consider reflexive partition as a control parameter. it is possible to formulate controllability problem, as follows. under a given set ℑ of feasible reflexive partitions, find the set of agents action vectors x(ℑ) = ∪ ℵ∈ℑ x(ℵ) that can be realized by reflexive control. the inverse problem lies in obtaining the “minimal” set of feasible reflexive partitions (in a certain sense), allowing to realize a given agents action vector. now, let us address the control problem. suppose that the preferences of a control subject (a principal) are described by his/her real-valued goal function f0(q(x∗)) defined on the set of aggregated outcomes (q : ℜn → ℜ1), i.e., f0(·) : ℜ1 → ℜ1. using the expression (16), the efficiency of the reflexive partition ℵ can be characterized by k(ℵ) = f0(q(x∗(ℵ))). consequently, the problem of reflexive control (in terms of reflexive partitions) 272 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... can be formally stated as [1, 77] k(ℵ) → max ℵ∈ℑ . (17) let km be the maximal value of the efficiency criterion in the problem (17) under a fixed maximal reflexion rank m. the problem of the maximal rational rank of reflexion (a rank being pointless to exceed for the principal in the sense of controllability or/and efficiency of reflexive control) is to find: m∗ = min{m|m ∈ arg max w=0, 1, 2, ... kw}. to proceed, we discuss conformity of subjective reflexive partitions of the agents. suppose that each agent observes merely the aggregated outcome. trajectories forecasted by the agents may differ from the real trajectory (see the general reflexive model of collective behavior). this motivates the agents to doubt the correctness of their subjective reflexive partitions. imagine that the agents observe just the aggregated outcome of the game (in addition to their own actions). by analogy to the condition of stable informational control (see above), one can introduce the condition of stable reflexive partition. notably, require that the aggregated outcome for the real trajectory coincides with the forecasted aggregated outcomes for all agents. stability of reflexive partitions is closely associated with learning in games. observing the behavior of opponents (which differs from the forecasted behavior), agents may modify their beliefs about the reflexion ranks of the opponents or pass to higher levels of reflexion. under a fixed reflexive partition ℵ ∈ ℑ, we have realization of the action vector (16). and the aggregated outcome q(x∗(ℵ)) is realized. according to agent with reflexion rank , the following vector is realized x̃jk(ℵjk) =(xl∈n0 , x1l∈n1 , x2l∈n2 , . . . , x[k − 1]l∈nk−1∪nk∪...∪nm\{j}, xkj), j ∈ nk, k = 0, m the condition of stable reflexive partition ℵ ∈ ℑ takes the form q(x̃jk(ℵjk)) = q(x ∗ (ℵ)), j ∈ nk, k = 0, m. the problem of reflexive control (ℵ) can be stated on the set of stable reflexive controls (if nonempty). in practice, this means that the principal forms an optimal partition of the agents into reflexion ranks. in such partition, the agents do not doubt the correctness of their beliefs about reflexion ranks of the opponents (based on observing the results of the “game”). 3 conclusion mathematical models of informational and/or reflexive structures and equilibria (and models of corresponding control problems) allow the following: • from the decision theory viewpoint, extending the class of collective behavior advances in systems science and applications (2014) vol.14 no.3 273 models for intelligent agents performing a joint activity under incomplete awareness and missed common knowledge; • from the descriptive viewpoint, enlarging the set of outcomes that can be “explained” (within the framework of the model) as stable results of agents interaction; accordingly, extending the controllability domain (for control problems); • from the normative viewpoint, posing/solving the problems of collective behavior by choosing a proper structure of agents awareness. numerous applied models of informational or/and reflexive control in economic, social and organizational systems, military problems and other fields are described in [1, 32, 64]. as strategic objectives of future investigations, we mention integration of informational reflexion models with strategic reflexion ones. in other words, it seems promising to construct a language for uniform joint description of informational and reflexive structures. references [1] novikov d.a., ckhartishvili a.g. (2014), reflexion and control: mathematical models, leiden: crc press. [2] novikov a.m., novikov d.a. (2013), research methodology, c leiden, crc press. [3] novikov d.a. (2013), theory of control in organizations, n.y.: nova scientific publishing. [4] novikov d.a. (2013), control methodology, c n.y.: nova scientific publishing. [5] huizinga j. (2008), homo ludens, london. roughtledge. [6] germeier yu. non-antagonistic games, 1976. dordrecht: d. reidel publishing company, 1986. [7] myerson r. (1991), game theory: analysis of conflict, c london: harvard univ. press. [8] geanakoplos j. (1994), common knowledge / handbook of game theory, vol.2, amsterdam: elseiver, pp.1438-1496. [9] morris s., shin s.s. (1997), “approximate common knowledge and coordination: recent lessons from game theory”, journal of logic, language and information, vol.6, pp.171-190. [10] lewis d. (1969), convention: a philosophical study, cambridge: harvard university press. [11] aumann r.j. (1976), “agreeing to disagree”, the annals of statistics, vol.4, no.6, pp.1236-1239. [12] aumann r.j. (1999), “interactive epistemology i: knowledge”, international journal of game theory, no.28, pp.263-300. 274 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... [13] fagin r., halpern j., moses y., vardi m. (1999), “common knowledge revisited”, annals of pure and applied logic, vol.96, pp.89-105. [14] fagin r., halpern j., moses y., vardi m. (1995), reasoning about knowledge, cambridge: mit press. [15] fagin r., halpern j., vardi m. (1991), “a model-theoretic analysis of knowledge”, journal of assoc. comput. mach., vol.38, no.2, pp.382-428. [16] hill b. (2010), “awareness dynamics”, journal of philosophical logic, no.39, pp.113-137. [17] li j. (2009), “information structures with unawareness”, journal of economic theory, no.144(3), pp.977-993. [18] simon r. (1999), “the difference of common knowledge of formulas as sets”, international journal of game theory, vol.28, pp.367-384. [19] fagin r., geanakoplos j., halpern j., vardi m. (1999), “the hierarchical approach to modeling knowledge and common knowledge”, international jounal of game theory, vol.28, pp.331-365. [20] van huyck j., cook j., battalio r. (1997), “adaptive behavior and coordination failure”, j. of economic behavior and or-ganization, vol.32, pp.483-503. [21] clark h.h., marshall c.r. (1981), “definite reference and mutual knowledge”, elements of dicourse understanding, (ed. by a.k. joshi, b.l. webber, i.a. sag), cambridge: cambridge university press, pp.10-63. [22] halpern j., moses y. (1990), “knowledge and common knowledge in a distributed environment”, journal of assoc. comput. mach., vol.37, no.3, pp.549-587. [23] gray j. (1978), “notes on database operating system / operating systems: an advanced course”, lecture notes in computer science, vol.66. berlin: springer, pp.393-481. [24] mccarthy j., sato m., hayashi t., igarishi s. (1979), on the model theory of knowledge, technical report stan-cs-78-657, stanford: stanford university. [25] nash j.f. (1951), “non-cooperative games”, ann. math., vol.54, pp.286-295. [26] copic j., galeotti a. (2007), “awareness equilibrium. mimeo”, essex: university of essex. [27] feinberg y. (2004), subjective reasoning games with unawareness. research paper no.1875. stanford: graduate school of business. [28] heifetz a., meier m., schipper b. (2008), “a canonical model of interactive unawareness”, games and economic behavior, no.62, pp.304-324. [29] rego l., halpern j. (2012), “generalized solution concepts in games with possibly unaware players”, international journal of game theory, no.41, pp.131-155. [30] chkhartishvili a.g., novikov d.a. (2004), “models of reflexive decision-making”, systems science, vol.30, no.2, pp.45-59. advances in systems science and applications (2014) vol.14 no.3 275 [31] novikov d.a., chkhartishvili a.g. (2003), “information equilibrium: punctual structures of information distribution”, auto-mation and remote control, vol.64, no.10, pp.1609-1619. [32] novikov d.a., ckhartishvili a.g. (2003), reflexive games, moscow: sinteg (in russian). [33] lefebvre v.a. (1965), basic ideas of the reflexive games logic / proc, ≪problems of systems and structures researches≫, moscow: ussr academy of science (in russian). [34] lefebvre v.a. (2010), algebra of conscience. 2nd ed. berlin: springer. [35] lefebvre v.a. (1973), conflicting structures, moscow: soviet radio (in russian). [36] lefebvre v.a. (2010), lectures on reflexive game theory, los angeles: leaf & oaks. [37] lefebvre v.a. (1998), sketch of reflexive game theory / proc. of workshop on multi-reflexive models of agent behavior, los alamos, new mexico, usa, pp.144. [38] ereshko f.i. (2001), modelling of reflexive strategies in control systems moskow: cc ras (in russian). [39] gorelik v.a., kononenko a.f. (1982), game–theoretical models of decision making in ecological and economic systems. moscow: radio and communication (in russian). [40] soros g. (1994), the alchemy of finance: reading the mind of the market, new york: wiley. [41] taran t.a., shemaev v.n. (2004), “boolean reflexive control models and their application to describe the information struggle in socio-economic systems”, automation and remote control, vol.65, no.11, pp.1834-1846. [42] aumann r.j., brandenbunger a. (1995), “epistemic conditions for nash equilibrium”, econometrica, vol.63, no.5, pp.1161-1180. [43] heifetz a. (1999), “iterative and fixed point belief”, journal of philosophical logic, vol.28, pp.61-79. [44] hintikka j. (1962), knowledge and belief, ithaca: cornell university press. [45] kripke s. (1959), “a completeness theorem in modal logic”, journal of symbolic logic, no.24, pp.1-14. [46] mertens j.f., zamir s. (1985), “formulation of bayesian analysis for games with incomplete information”, international journal of game theory, no.14, pp.1-29. [47] wolter f. (2000), “first order common knowledge logics”, studia logica, vol.65, pp.249-271. [48] camerer c., ho t., chong j. (2004), “a cognitive hierarchy model of games”, the quarterly j. of economics, no.8, pp.861-898. 276 novikov d.a. and chkhartishvili a.g.:mathematical models of informational and ... [49] nagel r. (1995), “experimental results on interactive competitive guessing”, american economic review, vol.85, no.6, pp.1313-1326. [50] stahl d., wilson p. (1994), “experimental evidence on players’ models of other players”, journal of economic behavior and organization, vol.25, pp.309-327. [51] novikov d.a. (2012), “models of strategic behavior”, automation and remote control, vol.73, no.1, pp.1-19. [52] weber r. (2001), “behavior and learning in the ≪dirty face≫ game”, experimental economics, vol.4, pp.229-242. [53] harsanyi j. ”games with incomplete information played by “bayesian” players”, management science. part i: 1967, vol.14, no.3, pp.159-182; part ii: 1968, vol.14, no.5, pp.320-334; part iii: 1968, vol.14, no.7, pp.486-502. [54] brandenburger a., dekel e. (1993), “hierarchies of beliefs and common knowledge”, journal of economic theory, vol.59, pp.189-198. [55] the handbook of experimental economics / ed. by j. kagel and a. roth. c princeton: princeton university press, 1995. [56] wright j., leyton-brown k. (2010), beyond equilibrium: predicting human behavior in normal form games / proc. of conf. associat, advancement of artificial intelligence (aaai-10), pp.461-473. [57] kahneman d., slovic o., tversky a. (1982), judgment under uncertainty: heuristics and biases, cambridge: cambridge university press. [58] nisan n., roughgarden t., tardos e. et al. (2007), algorithmic game theory, c cambridge: cambridge university press. [59] bernheim d. (1984), “rationalizable strategic behavior”, econometrica, no.5, pp.1007-1028. [60] weibull j. (1995), evolutionary game theory, c cambridge: mit press. [61] chkhartishvili a.g. (2003), “bayes-nash equilibrium: infinite-depth point information structures”, automation and remote control, vol.64, no.12, pp.1922-1927. [62] novikov d.a., chkhartishvili a.g. (2005), “stability of information equilibrium in reflexive games”, automation and re-mote control, vol.66, no.3, pp.441-448. [63] chkhartishvili a.g. (2010), “reflexive games: transformation of awareness structure”, automation and remote control, vol.71, no.6. pp.1208-1216. [64] chkhartishvili a.g. (2004), “game-theoretical models of informational control”, moscow: pmsoft (in russian). [65] chkhartishvili a.g. (2012), “concordant informational control”, automation and remote control, vol.73, no.8, pp.1401-1409. [66] costa-gomes m., crawford v. (2006), “cognition and behavior in two-person guessing games: an experimental study”, aer, vol.96, pp.1737-1768. advances in systems science and applications (2014) vol.14 no.3 277 [67] nagel r. (1995), “unraveling in guessing games: an experimental study”, aer., vol.85, pp.1313-1326. [68] stahl d., wilson p. (1995), “on players models of other players: theory and experimental evidence”, games and economic behavior, vol.10, pp.213-254. [69] burchardi k., penczynski s. (2010), “out of your mind: estimating the level-k model”, london: london school of econom-ics, working paper. [70] camerer c., ho t., chong j. (2003), “models of thinking, learning and teaching in games”, aea paper proc, vol.92, no.2, pp.192-195. [71] choi s., gale d., kariv s. (2005), “behavioral aspects of learning in social networks: an experimental study”, advances in behavioral and experimental economics, ed. by john morgan, jai press. [72] crawford v., costa-gomes m., iriberri n. (2012), “structural models of nonequilibrium strategic thinking: theory, evidence and applications”, journal of economic literature. [73] costa-gomes m., broseta b. (2001), “cognition and behavior in normal-form games: an experimental study”, economet-rica, vol.69, no.5, pp.1193-1235. [74] mckelvey r., palfrey t. (1995), “quantal response equilibria for normal form games”, games and economic behavior, vol.10(1), pp.6-38. [75] mccain r. (2010), learning level-k play in noncooperative games, philadelphia: drexel university, working paper. [76] stahl d. (1993), “evolution of smartn players”, games and economic behavior, no., pp.604-617. [77] korepanov v.o., novikov d.a. (2012), “the reflexive partitions method in models of collective behavior and control”, automation and remote control, vol.73, no.8, pp.1424-1441. [78] novikov d.a. (2010), “≪cognitive games≫: a linear impulse model”, automation and remote control, vol.71, no.4, pp.718-730. corresponding author d.a. novikov can be contacted at: novikov@ipu.ru. advances in systems science and applications (2012) vol.12 no.1 1-11 simulation based design for refrigerator development k. fukuyo graduate school of innovation and technology management, yamaguchi university, tokiwadai 2-16-1, ube, yamaguchi,7558611,japan abstract the simulation based design (sbd) is one of the new approaches to shorten development periods, reduce cost of development and avoid failures. this approach emphasizes an in-depth analytical understanding of target products or services in the upstream of a design process by using the simulation technology. this short communication presents some cases of simulation in the development of refrigerator as examples of bsd approach. keywords function s-rough sets, structure of function s-rough sets, relation theorem, function transfer rough law mining 1 introduction manufactures of refrigerators are forced to shorten development periods, reduce cost of development and avoid the failures and incidents. for those purposes, they are now trying to introduce new technologies and processes such as cad/cam/cae, digital engineering, or concurrent engineering instead of traditional ones. the simulation based design (sbd) is one of the new approaches to meet those purposes. according to bossak[1], its concept is “simulation of the entire life cycle of the product, from concept development to detailed design, prototyping, testing, manufacturing operations, maintenance and disposal” and it consists of “modeling methods and computational tools, virtual reality environment, an infrastructure for collaborative engineering and integration technologies and tools.” the concept of sbd is almost the same as those of the analysis-leads or analysis led design (ald, the terminology of isi and yamaguchi univ.) and the simulation driven product developmenttm (sdpd, the terminology of fluent). either word emphasizes a radical front-loading of design, i.e., an in-depth consideration in the upstream of a design process. an ordinary design process consists of the requirement definition, conceptual design, project design, and detail design phases. manufacturing and testing start almost as soon as the detailed design starts. some researchers argue that a sufficient analytical understanding of a target product/service is required in the upstream of design, i.e., the requirement definition and conceptual design phases and it results in decrease in the cost of development[2]. in the other words, a poor understanding of target products/services will bring some serious problems in the downstream of the design process, cause many trials and errors in the manufac2 k. fukuyo:the application of semantic web technology in ... turing and testing process, and significantly increase cost and time to market. therefore, the necessity for the analytical understanding in the upstream of a problem-solving process has been mentioned since a long time ago[3-4]. nowadays, a progress of information technology allows manufacturers to use computer simulations as a tool for the analytical understanding of target products/services; hence new approaches of design such as sbd, ald, and sdpd tm are proposed. 2 simulation methods in refrigerator development concept of sbd has become important among the refrigerator manufacturers, although sbd is not perfectly deployed in the design and manufacturing process. in this industry two types of simulation are mainly used for design process. one is cycle simulation, which is the system simulation of a refrigerating cycle, and the other is computational fluid dynamics (cfd). this short communication treats the latter. (a) dynamic and changing. process knowledge should be updated along with the emergence of new techniques and new methods. cfd simulation applied for designing refrigerators are as follows: cortella et al. applied a fem (finite element method) code to carry out unsteady 2-dimensional simulation of a refrigerated display cabinet[5]. chang et al. applied a cfd code with a k−ε turbulent flow model to refrigerators equipped with freezing cabinets at the top[6]. afonso and matos used a commercial cfd code, fluent tm to study the effect of radiation shields around a condenser and compressor[7]. gupta et al. used the cfd model to study effects of various operating and design parameters on the refrigerator performance[8]. haldar et al. applied their developed code to simulate free convection inside a freezer[9]. the author of this short communication has also carried out cfd analyses in the upstream of the refrigerator development. some cfd analyses carried out by the author are presented in the following sections as examples of bsd. 3 realization of thermal uniformity inside refrigerators fig.1 shows a household refrigerator with a conventional air-supply system. the fan supplies the air cooled by an evaporator to the fresh food cabinet, the vegetable cabinet and the freezing cabinet. in the fresh food cabinet, the cooled air is supplied from openings in the rear panel. temperature of the supplied air is set at a sufficiently low temperature to cool the freezing cabinet. thus the flow rate of the supplied air to the fresh food cabinet is suppressed to prevent freezing and drying foods, and increasing the heat gain from outside the refrigerator. however, the suppression of the flow rate causes rises in temperature in the door pockets and the upper shelf. besides this problem, the conventional air-supply system can not increase cooling rate because of the suppression of the flow rate. advances in systems science and applications (2012) vol.12 no.1 3 to meet the requirements for keeping food fresh, a new air-supply system for improving the thermal uniformity and the cooling rate inside the fresh food cabinet was developed by adding a blower and jet slots to a conventional cooled air supply system[10]. in the design process of this new air-supply system, cfd is applied to study air distribution inside refrigerators. fig.1 cross section of the refrigerator 3.1 design of the air-supply system fig.2 shows the schematics of the fresh food cabinets with a conventional air supply-system and our developed system. the volume of each cabinet is 0.24 m3. each air-supply system has ten openings and six grills. the openings supply cooled air at 0.17 m3 min−1 and -10◦c to keep the cabinet temperature below 5◦c, and the grills returns the air to an evaporator. the flow rate of the supplied air to the cabinet is suppressed to prevent freezing and drying foods, and increasing the heat gain. a thermal sensor is set on the rear panel of each cabinet. when the sensor detects that the surrounding temperature is lower than the preset one, air-supply is stopped. system a is the conventional air-supply system. the temperatures in the door pockets and the upper shelf will rise, because of the distances between those places and the supply openings. system b is the new air-supply system. one blower, three jet slots and two return grills are added to system a to improve the thermal uniformity and the cooling rate. the flow rate of the air jetted by the additional blower is 0.3 m3 min−1. the total opening area of the jet slots is one sixth of that of the supply openings. the velocity of the jets is thus ten times higher than that of the cooled air. the jet stirs the air inside the cabinet and improves thermal uniformity. the jet also improves the heat transfer on the surface of foods and cools it rapidly. 4 k. fukuyo:the application of semantic web technology in ... the blower collects the air inside the cabinet through the grills and jets air from the slots. the air does not pass through the evaporator. the air is thus at about the same temperature as the average temperature inside a cabinet and does not freeze foods. (a) (a)system a (b) (b)system b fig.2 layouts of air supply systems it takes much time to repeat the measurements for studying the air distribution inside the cabinet with the new system. in order to cut down the development period, we applied cfd simulations to the study on the air distribution inside the cabinet. 3.2 simulation results fig.3 shows the simulated thermal distribution for each air-supply system. regarding system a, there are high temperature regions above the top shelf and advances in systems science and applications (2012) vol.12 no.1 5 the top door pocket, and there is also a low temperature region in the center of each shelf. in system b, the air is kept at about 2◦c in most of the area, because the jetted air sufficiently stirs the air inside the cabinet. the heat gains of system b increases by 5% compared to system a because of the decrease in temperature of the regions that are not sufficiently cooled by system a. however, the standard deviation of the temperature for the fresh-food cabinet with system b is about half of that of system a. simulation results are validated by comparing with the measurement after the (a) (a)system a (b) (b)system b fig.3 thermal distributions in the central cross section simulation. the difference between the temperatures at the points indicated by circles in fig. 3 is 0.5◦c in the simulation and it is 0.3◦c in the measurement. it is concluded from these simulation results that system b improves thermal uniformity. after these simulations, the project design and detail design were carried out and the refrigerators with this new air-supply system were released 6 k. fukuyo:the application of semantic web technology in ... to the market. 4 thermal-bridge problem in order to decrease heat exchange between the inside and outside of refrigerators, insulating walls are applied. however, supporting elements of the insulating walls (frames and reinforcements) often act as thermal bridges and increase heat gain or loss. boughton et al. have shown that heat loss through thermal bridges accounted for about 30% of a refrigerators overall heat loss[11]. applying cfd and a heat-flow visualization technique to the upstream stage of the refrigerator design will help to know and solve the thermal-bridge problems. below is an example of the thermal-bridge problems[12]. 4.1 thermal-bridge problem around gaskets fig.5 shows a cross section of the region where the insulating wall and door of a refrigerator contact through a gasket. the wall and door consist of polyurethane foam insulators and reinforcements. the foam of the insulators is filled with isobutane. the wall and door are 100.0-mm-thick, and are separated by a 4.0mm-wide gasket. the left side of the insulation is the inside of the refrigerator and the right side is the outside. the inside and outside air are at temperatures of 273 and 303k respectively. the heat transfer coefficients on both surfaces are set to 10.0 w m−2 k. the upper side of the wall and lower side of the door are set to be adiabatic boundaries. two cases of steady-state simulations were carried out. the region shown in fig.4 is divided into ten thousand of 1.0 mm2 control volumes. the inner reinforcements were made of steel sheets in case (a) and of abs resin in case (b). the thermal properties of the material used in this region are listed in table 1. 4.2 simulation results fig.5 (a) and (b) shows the distributions of heat-flow intensity in cases (a) and (b) respectively. the respective heat flow intensities are 1.8 mw m−3 and 64 kw m−3 at the marked points in fig.5 (a) and (b); this indicates that the steel inner reinforcements conduct heat about 28 times as large as abs reinforcements. 5 radical change in the layout of cabinets the design for environment (dfe), i.e., environmentally conscious design is an important issue for manufacturers. they pay attention to the energy efficiency of their products to decrease co2 emission resulting from their energy consumption. in the field of the household refrigerators, manufacturers develop high performance insulation walls, compressors, heat exchangers, and control devices to improve the energy efficiency of their products. moreover, they try to change advances in systems science and applications (2012) vol.12 no.1 7 fig.4 cross-sectional view of the region where an insulating wall is connected with an insulating door. unit: mm the refrigerants to natural ones to keep the ozone layer. however, the improvement of energy efficiency will reach the limit if the shapes of the refrigerators are fixed. manufacturers should try to change the shape of the refrigerators to improve the energy efficiency. the current mainstream refrigerators shown in fig.1 are called “bottom freezer type” because of the location of the freezing cabinet. the layout of each cabinet is decided based on the usability or ergonomics. this layout has not changed for many years. however, there are two following problems in this type of refrigerator from the view point of thermal engineering. • nearness of the compressor and freezing cabinet causes temperature rise in the freezing cabinet because of heat gain from the compressor. • placing the compressor in the limited space decreases the heat transfer from the compressor. these problems decrease refrigerators’ energy efficiency. one of the solutions of these problems is change in the shape of the refrigerators. the effects of change in the layout of cabinets were studied by cfd[13]. table 1 conditions of the heat gain calculation conductivity density specific heat [w m1 k] [kg m−3] [j kg−1 k−1 ] air 0.026 1.205 1006 gasket 0.16 1230 1.42 inner reinforcements (steel) 43 7880 473 (abs resin) 0.22 1040 1000 insulator 0.03 100 836 outer reinforcement (steel) 43 7800 473 8 k. fukuyo:the application of semantic web technology in ... fig.5 distributions of heat-flow intensities for the region shown in fig.4 fig.6 layouts of cabinets and compressors 5.1 alternative layouts of cabinets some alternative shapes of refrigerators have been already proposed as shown in fig.6. the “top compressor type” refrigerators are the ones that have the compressor at the top. they increase the heat transfer from the compressor. the “separate type” refrigerators are the ones that the compressors are separated from the refrigerators’ cabinets. this type of refrigerator not only increases the heat transfer from the compressor but also prevent the freezing cabinet from gaining the compressor’s heat. the heat gains of the cabinets of the ordinary and alternative types from the ambient air were calculated by cfd on the conditions listed in table 2. 5.2 simulation results the result is shown in fig.7. the heat gains of the alternative types are smaller than that of the current one. this tells that the energy efficiency of the alternative advances in systems science and applications (2012) vol.12 no.1 9 types will be improved. fig.7 heat gain of cabinets from the ambient air table 2 thermal property of materials fresh food & vegetable cabinets freezing cabinet size width [m] 0.65 0.65 depth [m] 0.65 0.65 height [m] 1.35 0.45 insulation wall thickness [m] 0.04 0.06 thickness [m] 0.04 0.06 conductivity[w m1 k−1] 0.02 0.02 heat transfer inside[w m−2 k−1] 10.0 10.0 outside[w m−2 k−1] 5.0 temperature cabinet [◦c] 5.0 -18.0 ambient [◦c] 30.0 compressor [◦c] 50.0 however, there are some problems in the serviceability or usability dimension. the top compressor type doesn’t allow users to put something on the ceiling and the location of the compressor makes it difficult to maintain the compressor. moreover, the separate type needs an excess area and time to set it up. this indicates that it is not only the engineering performance that decides the value of a product/service. a design of a product/service should be carried out in view of performance, features, reliability, conformance, durability, serviceability, aesthetics, and perceived quality [14]. this is why an in-depth consideration by 10 k. fukuyo:the application of semantic web technology in ... using simulation is required in the upstream of a design process. 6 conclusion three examples of sbd in the refrigerator development are presented in this short communication. as shown in each example, by using cfd simulations, we can study the performance of target products analytically in the upstream of the design process, avoid to repeat the measurements and test in the downstream, and cut down the development period and cost. acknowledgements the research is supported by national natural science foundation of china(no. 11171065). references [1] m.a. bossak. (1998), “simulation based design”, technology, pp.8-11. [2] l.r. jenkinson, j.f. marchman iii. (2003), aircraft design project, butterworth-heinemann . [3] d.w. ver planck, b.r. teare, jr. (1954), “engineering analysis”, an introduction to professional method, john wiley & sons. [4] g. polya. (1957), how to solve it, princeton university press. [5] g. cortella, m. manzan, c. comini. (2005), “cfd simulation of refrigerated display cabinets”, international journal of refrigeration, no.24, pp.250-260. [6] w.r. chang, j.y. lin, h.c. hsu. (2001), “air flow simulation and energy estimation for household refrigerators/freezers”, proceedings of 52th annual international appliance technical conference, columbus, vol.17, no.6, pp.750-771. [7] c. afonso, j. matos. (2006), “the effect of radiation shields around the air condenser and compressor of a refrigerator on the temperature distribution inside it”, international journal of refrigeration, no.29, pp.1144-1151. [8] j.k. gupta, m.r. gopal, s. chakraborty. (2007), “modeling of a domestic frost-free refrigerator”, international journal of refrigeration, no.30, pp. 311-322. [9] s.c. haldar, k. manohar, g.s. kochhar. (2008), “conjugate conductioncconvection analysis of empty freezers”, industrial management and data systems, no.49, pp.783-790. advances in systems science and applications (2012) vol.12 no.1 11 [10] fukuyo k. et al. (2003), thermal uniformity and rapid cooling inside refrigerators, international journal of refrigeration, vol.26, no.2, pp.249-255. [11] b.e. boughton, a.m. clausing, t.a. newell. (1996), “an investigation of household refrigerator cabinet thermal loads”, hvac&r research, no.2, pp.135-148. [12] k. fukuyo. (2003), “heat flow visualization for thermal bridge problems”, international journal of refrigeration, no.26, pp.614-617. [13] k. fukuyo, et al. (2005), “conflicts between eco-design and usability of refrigerators”, 4th international symposium on environmentally conscious design and inverse manufacturing ecodesign, vol.12-14, no.12, pp.3d-11s. [14] d.a. garvin. (1987), “competing on the eight dimensions of quality”, harvard business review, vol.65, pp. 101-109. advances in systems science and applications (2012) vol.12 no.4 353-372 ecosystem food-webs as dynamic systems: educating undergraduate teachers in conceptualizing aspects of food-webs’ systemic nature and comportment kostas karamanos1, aristotelis gkiolmas1, anthimos chalkidis1, constantine skordoulis1, maria papaconstantinou2 and dimitrios stavrou3 1athens science & education laboratory, faculty of primary education, university of athens, greece. 2faculty of informatics, ionian university, corfu, greece. 3faculty of primary education, university of crete, greece. abstract the current research aimed at attempting to familiarize greek prospective primary school teachers with notions of the dynamic systems nature of ecosystem food-webs and evaluate the results. the sample consisted of 85 undergraduate students of the faculty of primary education, of the university of athens, greece, who had chosen the optional topic of the autumn semester: “environmental science: the laboratory approach”. students were initially given specifically designed pre-test worksheets, to find out their initial knowledge and pre-instructional ideas about the structure and the properties of food-webs as systems. following this, instructional-teaching took place in four stages. in the first stage of the instruction process, students received real data concerning pellets of barn owls (tyto alba) gathered from two areas of the united states, and were asked to complete given worksheets. at this stage, students’ discussions were recorded, hence trying to extract answers to issues such as: the food-webs’ ability to reflect biodiversity; their mental representation of a food-web as a system; and what would happen should a “node” of the food-web-network disappear. the second stage involved students interacting with computer software in two phases. in the first, students were asked to reconstruct, on a computer screen, the barn owl food-web, given the arrows and the food-web elements. in the second stage, students used the system dynamics modeler of the netlogo (version 4.0.4) simulation environment, as an interactive tool for learning. in the third stage, after familiarizing themselves with the modeler, using the simple predator-prey ecosystem “wolf sheep predation (docked)” model, students interacted with a variation of the modeler, created in order to learn things about the stability and instability of food-webs as systems and, as a result, of ecosystems. in the fourth-final-stage, following the instruction, post-test worksheets were handed out to students and stratified samples of them were interviewed on further ideas, predictions and possible extensions of their knowledge. 354 kostas karamanos:ecosystem food-webs as dynamic systems... the results gathered helped to reach encouraging conclusions related to the fact that students seemed to have conceptualized food-webs as dynamic systems. keywords food-webs, dynamical systems, stability, teaching, netlogo 1 introduction in the recent scientific literature, the study of ecological food-webs is important in understanding the systemic behavior of ecosystem populations, in revealing aspects of their comportment as complex systems, as well as in making predictions about their future states. typical questions in this latter field of research are: whether an ecological system will become more stable or unstable in the case that one population becomes extinct, or whether the interconnectedness in the food-web is increased or decreased. other problems concerning the systemic nature of food-webs are: (i) the degree of relation if any between the biodiversity and the connectivity in a food-web, (ii) the relative importance of the abundance of a population in the food-web as compared to its biomass content. the relationship between the systemic structure of food-webs and the properties of ecological systems as “complex systems” has been thoroughly researched. it is stated, for instance, that when a hypothetical food-web is formed by adding edges at random to a set of n “nodes”, a connectivity avalanche occurs when the number of edges is approximately n/2, which means that critical phenomena arise[1]. the critical phenomena have the form of a phase-change, where a connected sub-web is clearly formed, and the food-web looks more “connected”. such relations between the systemic structure, food-web behavior and the complexity of ecological systems, are the basis of this research. the broader research project, carried out in the faculty of primary education of the university of athens, greece, aimed at familiarizing prospective primary school teachers with the concepts of complex systems, specifically ecological systems. similar researches have been carried out in many european and north american educational systems[2], due to its outmost importance in developing complexity-thinking abilities in future educators, especially with respect to non-deterministic thinking, holistic attitudes, absence of pre-existing ideas about a central-control in all ecological phenomena and non mechanistic, non-clockwise treatment of all procedures in nature[3-4]. the inherent relation between systemic-thinking and “complexity-thinking” is among the core ideas in this research[5-6]. a person who has developed the ability to conceive a system as whole and is additionally able to conceptualize the relations between these parts, is also capable of understanding complex system properties such as emergence and feedback loops[6]. the well known property of complex systems “the whole is greater than the sum of the parts” is implicit in the systemic treatment of natural entities. advances in systems science and applications (2012) vol.12 no.4 355 food-webs are a clear case of the direct relation between aspects of ecological complexity and the systemic nature of ecosystems. in a wide known definition given by asher[7]: “. . . ‘complexity’ is the multiplicity of interconnected relationships and levels.” the existence of this property of multiplicity in ecological food-webs, when systemically treated, is obvious. asher declares that all the characteristics so often attributed to complex systems (and, therefore, complex ecological systems), such as emergence and non-linearity, are consequences of the fundamental properties identified in his definition[8]. another famous statement about complexity is that: “. . . complex interactions result from multiple pathways linking organisms with abiotic resources”[9]. once more these pathways become revealed in detail, within the ontological, systemic representations of food-webs in various ecosystems. a broader link between ecological complexity and the systemic nature of foodwebs is the notion of “connectivity”, i.e. the number, form, complicatedness and scale of connections between organisms (but also between abiotic components) in an ecosystem[8]. the greater and broader the connectivity of a food-web, the more properties of complex systems is found in it. in fig.1, connectivity is depicted as one of the aspects of ecological complexity[8]. fig.1 connectivity as one of the aspects of ecological complexity. connectivity and complexity increases from top to bottom. lastly, there is the concept of ‘stability’. the majority of ecologists believe that as food-web complexity increases, the stability of the ecosystem as a whole increases[10-12], in the sense that no populations or species tend to become extinct or to have an unrestrained increase. 356 kostas karamanos:ecosystem food-webs as dynamic systems... 2 method 2.1 the objectives the main scopes of the current research were: (i) to test and improve the ability of prospective primary school teachers (undergraduate students of the faculty of primary education) to create representations and then construct (on paper and on screen) relatively simple food-webs, given the species that participate in these webs. (ii) to help students conceptualize the main properties of food-webs as systems. for example, the property that not all populations have the same importance in the food-web, or that biodiversity and multiplicity of food-webs depend on geographic latitude (they increase as one moves from the polar to the equatorial zone). (iii) to increase students’ abstractive and comparative thinking, by providing them with a scientific/theoretical food-web, and then asking them to construct one similar to it based on the data of food-web-findings from certain areas. (iv) enticing students to examine the relationship between food-web interconnectedness and ecosystem stability. lateral objectives of this research were also to: help students understand and derive information from systemic diagrams; depict graphical representations of physical concepts, such as stability and instability, and to verbally express opinions about systemic properties of ecological systems. 2.2 research instruments in addition to pre-test and post-test worksheets, with each one lasting one and a half teaching hour1, undergraduate students were given data in form of tables, diagrams of food-webs, and worksheets. during each of the four teaching sessions (each lasting two hours) students had to complete a worksheet containing both open and closed questions. recorded classroom discussions were instigated so as to reach unanimous conclusions. computer interaction was a composite part of each teaching session. specifically: (i) “microsoft word”, with regards to its “drawing” tools, and (ii) “netlogo version 4.0.4.”[13] as well as the system dynamics’ modeler, built-in in netlogo. 2.3 sample and the settings the sample consisted of 85 undergraduate students (66 women and 19 men) from the faculty of primary education of the university of athens, greece, who 1pre-test and post-test sheets, as well as the teaching sequences worksheets, are available, upon request of the reviewers. advances in systems science and applications (2012) vol.12 no.4 357 had chosen the optional subject “science and environment: the laboratory approach”. students were grouped with respect to the following parameters: a. year of studies (a1freshmen to a4senior) b. selected biology as an exam subject for their university entry national exams (b1 implies ‘yes’ and b2 implies ‘no’). c. students’ orientation in the final year of high-school, namely: (i) science, (ii) humanities and law and (iii) technology, economics and informatics (denoted as: c1, c2 and c3, respectively). d. at high-school or university, did the student select optional topics relevant to computers? this parameter implies a pre-existing familiarity with computers (d1 implies ‘yes’ and d2 implies ‘no’). the sample of 85 students was divided into five groups; four of 18 students and one of 13 students. the first author of the paper met up with each group once a week, and the students worked either alone (e.g. when completing worksheets) or in groups of two or three, when working on a computer. in the interviewing phase, each student underwent a recorded interview, lasting about an hour. 3 the instruction process and the delivered material the overall teaching sequence started with a very short interview with each one of the 85 undergraduate students, in an effort to group them according to the parameter values mentioned above. their basic knowledge of ecosystems and food-webs, if any, were also questioned. each group of students was then asked to complete a pre-test worksheet2, with questions, tasks, and interactions with the computer. each group received four weekly teaching sessions (i to iv), each lasting two teaching hours (two 50 minute-sessions). teaching session i: students are given real data from a table[14], depicting findings regarding the pellets of the barn owl (tyto alba) from two different regions of the united states of america: northwestern and southwestern. the table with the pellet findings is given in fig.2. the barn owl (tyto alba) is a very important bird species for the study of food-webs and biodiversity, as it is a “high-level” predator and it has the property of emitting ‘pellets’, where the spinal cord of its preys remains intact. therefore, by studying the pellets of the barn owl, one can more or less study the biodiversity and the abundance of the species living in a given area. 2both the pre-test and the post-test sheets, as well as the teaching sequences worksheets, are available, upon request of the reviewers. 358 kostas karamanos:ecosystem food-webs as dynamic systems... fig.2 the pellet findings for tyto alba in northwest and southwest u.s.a.[14] given the above table, students were asked to answer certain questions (more details follow in the results section). teaching session ii: in the second teaching session, which took place one week later, students were given a worksheet with an actual representation of the food-web of the barn owl (fig.3), as derived from scientific literature[15]. the worksheet contained questions asking to correlate the theoretical foodweb of the barn owl with the actual two food-webs that stem from the table in fig.2. advances in systems science and applications (2012) vol.12 no.4 359 fig.3 the food-web of the barn owl, as given in the scientific literature[15]. students attempt to represent, on paper, two food-webs (one for each region), where the names of prey and predators must correspond to the given pelletfindings table. this involved having to exclude or add species to fig.3. following this, students answered questions regarding the complexity of the food-web as a system, in relation to geographical latitude, the relative importance of species and their biomasses in the food-web system. teaching session iii: in the third teaching session, students were given a simpler table of the barnowl-pellet findings, containing only those species which have a significant change in their abundance between the northwestern and the southwestern states. students were asked to work in groups of two or three and to open an existing word (*.doc) file named “exercise in the food-web of the barn owl.” a screenshot of this word file is depicted in fig.4. students were asked to cooperate in moving the pre-existing arrows of the word file and the pre-existing shapes, to form, on the screen, a simplified foodweb of the owl, once for northwest and once for southwest. each of the two files was uniquely named. they were then asked to answer a series of questions on the worksheet, the most significant being: ‘question number 9. a) which one of the two food-webs that you have represented on screen is more 360 kostas karamanos:ecosystem food-webs as dynamic systems... fig.4 a screenshot of the microsoft word file exercise in the food-web of the barn owl. the objects (animals, plants and arrows) are movable. complex? the one for the northern or the one for the southern united states? justify our answer. b) in which of the two food-webs does the species of the barn owl (tyto alba) have more possibilities of becoming extinct? why?’ teaching session iv in the final teaching session, the students were asked once again to work in groups. this time, they interacted with the netlogo version 4.0.4 programming and the simulation environment[13]. specifically, they worked with the “wolf sheep predation (docked)” model[16]. this model is accompanied with a “system dynamics modeler” file, created in java platform[16]. both these files in the model are depicted in fig.5 and 6 respectively. students, after trying to interpret the terms in the “system dynamics modeler” screen, focused on interacting with the sliders in the “aggregate model” section (right part of the interface screen) of the “wolf sheep predation (docked)” model. this section is nothing more than the graphical representation of a simple one-predator-and-one-prey population system, obeying to two simple lotkavolterra equations. students were guided by questions in their worksheet, to try and “destabilize” this system, by finding certain positions of the three sliders (‘predatorefficiency’, ‘predationrate’ and ‘wolves-death-rate’) where the wolves or the sheep become extinct, and hence, the other population becomes extinct too or deviates exponentially. the effort was to guide them to the conclusion that this is relatively easy, when there is such a simple food chain. advances in systems science and applications (2012) vol.12 no.4 361 fig.5 a screenshot of the netlogo ”wolf sheep predation (docked)” model[16], with both sections of the interface screen (the ”agent model” and the ”aggregate model”) activated. fig.6 a screenshot of the system dynamics modeler screen, belonging to the netlogo wolf sheep predation (docked) model[16] in the next step of this session, students were confronted by a variation of this simple prey-and-predator system, where, instead of one there are two species of prey. the idea behind this concept is depicted in fig.7 and in fig.8. in fig.7, the simple food chain of the initial “wolf sheep predation (docked)” model, is explained. in fig.8, the new simple food-web is shown. this food-web was constructed, by creating a very simple variation of the “system dynamics modeler” 362 kostas karamanos:ecosystem food-webs as dynamic systems... that was named: “system modeler food-web without grass”. in fig.9, below the new “system modeler food-web without grass” screen, fig.7 the food chain corresponding to the initial wolf sheep predation (docked) model, in the aggregate model part. fig.8 the simple food-web corresponding to the new “system modeler food-web without grass” model. which the students had in front of them in the computers, is depicted. and in fig.10, the interaction / interface screen of the same model (“sysfig.9 a screenshot of the “system modeler food-web without grass” model (“diagram” window from “system dynamics modeler”) tem modeler food-web without grass”) is shown, where the students were again advances in systems science and applications (2012) vol.12 no.4 363 asked to “destabilize” the system, by making one or two of the three populations extinct, with the use of the sliders. the concept behind this variation is that, in this case, it was more difficult fig.10 a screenshot of the interface of the “system modeler food-web without grass”. to “destabilize” the eco-system. twenty days following the final teaching session, a post-test was delivered to all groups of students3, containing identical questions to the initial pre-test worksheet. following the teaching session, a phase of one-hour interviews with a “stratified” sample of 17 students, begun. 4 results and discussion initially, students were handed a diagram, similar to fig.3, containing only the names and the images of the species involved in the barn owl food-web, without the predation connections between them. they were then asked to hand-draw the food-web, as they thought it should be and provide an explanation behind their food-web diagram. a. 72 out of the 85 students (percentage 84.7% of the sample) drew the web with arrows pointing, as they should, from the prey to the predator. b. 57 out of the 85 (percentage 67.1%) used in their argumentation on why they drew lines like that, the rationale “who is eaten by whom”, instead of “who eats who”.both these findings imply that the students have conceptualized a “bottom to top” perspective of ecosystems as systems, instead of the “top to bottom” perspective, the former being considered as the appropriate systemic treatment[17]. at a later stage, the students were asked to create on paper, a barn-owl-foodweb for the northern and sothern states of america. here, the systemic-thinking 3both the pre-test and the post-test sheets, as well as the teaching sequences worksheets, are available, upon request of the reviewers. 364 kostas karamanos:ecosystem food-webs as dynamic systems... skill examined was to correlate the theoretically existing elements (“stocks”) and relations (“flows”) in a system, with the actual ones, even if the latter have sometimes different names. the drawing of a student, alex, for the northern and for the southern united states, are depicted below, in fig.11. the drawings of alex show good correlation between given data and theory, fig.11 the food-web of the barn owl, as drawn by a student, alex. northwest u.s.a. (left) and southwest u.s.a. (right) even though this specific student reveals a “misconception”, by creating a more complex food-web in the north instead of in the south. due to the discrepancies between the theoretical food-web and actual barn-owlpellet table data, with regards to the type of species present, a “scoring system” was established for the drawings in the worksheets: whenever the student includes one of the eight species in his drawing, one (1) point is given. whenever the student misses one species, or includes a wrong one, zero (0) points are given. the actual scores and the corresponding numbers of students, are depicted in table 1 below. by calculating the average and the standard deviation of the scores of the students, we find: mean value = 4.36 sd = 1.85 the resulting mean value and standard deviation for the sample reveal a satisfactory ability to correlate to systems, achieved by the teaching sequence, which is an encouraging outcome. advances in systems science and applications (2012) vol.12 no.4 365 table 1 the scores of the n=85 students, in correlating the parts (“stocks”) of the theoretical and of the actual (real-data) barn-owl-food-web. number of students score 4 0 3 1 5 2 12 3 21 4 13 5 19 6 6 7 2 8 in concluding the results concerning objective (i), some statistical analysis was performed. frank draper [18], based on earlier work of barry richmond[19] proposes seven skills of systems thinking of students and the associated levels of activity[18]. each skill has also levels of acquainting it within it. for the current work, only the following skills and the specific sub-levels were considered relevant, ordered in increasing achievement difficulty: table 2 some of the proposed skills of systems’ thinking of students and the associated levels of activity[18] as they were selected for the current work. skill 1 (structural thinking), level ii identifying stocks and flows in phenomena skill 3 (generic thinking), level ii tabincellcrecognizing and using stockand-flow generic structures skill 4 (operational thinking), level ii building paper-and-pencil stock-and-flow diagrams skill 5 (scientific thinking), level i manipulating and modifying preconstructed computer models skill 6 (closed-loop thinking), level i identifying simple internal causal relations in the pre-test and post-test work sheets, with questions and interactions with the computer, there was effort to measure all these skills, based on a marine food-web, as retrieved by[20]. 366 kostas karamanos:ecosystem food-webs as dynamic systems... fig.12 the marine food-web delivered to the students in the pre-test and posttest worksheets[20]. by semantically-analyzing and content-analyzing students’ answers in both pre and post test, each was graded from 0 to 5 according to the skills demonstrated by their answers. following this, the non-parametric wilcoxon signed-rank test was carried out. there were statistically significant shifts of systemic-thinking skills in all levels of the scale 0-5. the results are briefly depicted in table 3 below. table 3 results of the non-parametric wilcoxon signed-rank test for the sample of n=85 students. grade significance level 0 p= .018 (¡0.05) 1 p= .014 (¡0.05) 2 p= .021 (¡0.05) 3 p= .023 (¡0.05) 4 p= .019 (¡0.05) 5 p= .011 (¡0.05) the findings are again considered encouraging, for the impact of the chosen teaching method. objective (ii) referring to the “relative importance” of each population (“stock”) in the foodwebas-a-system, the students had to use the table of food-web-pellets, together with simple calculators, to answer specific questions : q1: which species of prey contributes more to the barn owl’s food in each area of u.s.a. (northern and southern)? advances in systems science and applications (2012) vol.12 no.4 367 q2: is this the same species for both areas? q3: if all the species of shrew (blarina, cryptotis and sorex) become extinct, which one of the two areas’ food-webs will be affected more? or will they be equally affected? justify your answer. q4: if the only species of vole (microtus) becomes extinct, which one of the two areas’ food-webs will be affected more? or will they be equally affected? justify your answer. we briefly present the “correct” answers to q1-q4 and the corresponding ratio of students that found them: q1: the mice. (correct answers: 76 out of 85. percentage 89.4%) q2: yes, it was the same. (correct answers: 71 out of 85. percentage 83.5%) q3: the south will be affected much more. (correct answers: 40 out of 85.percentage 47.1%) q4: the two areas will be almost equally affected. (correct answers: 32 out of 85.percentage 37.6%) in the following fig.13, the percentages of correct answers, with respect to each question, are presented as a bar chart. fig.13 the percentages of students’ correct answers, with respect to each question. what can be observed is that within the same food-web as a system, the students are adequately capable of finding the prevailing “stock” (population), but this capability is reduced, when having to think comparatively between two food-webs/ systems. this has also been noticed by other researchers[21]. the answers obtained, regarding the question outlined below are shown in table 4. ‘the diversity of species and the complexity of the food-webs increases, de368 kostas karamanos:ecosystem food-webs as dynamic systems... creases or does something else, as one moves from the poles to the equator? explain your answer.’ from the n=85 undergraduate students, the results were: table 4 results of the non-parametric wilcoxon signed-rank test for the sample of n=85 students. “increases” “decreases” “do not know/ other answers i cannot answer” pre-test 50 14 5 16 post-test 73 6 r.2 4 in fig.14, the findings of table 4 are expressed in the form of pie charts. it is easily concluded, that following the teaching sequence, students seem to have conceptualized that as the geographical latitude decreases in absolute values, the diversity and the connectedness of food-webs as systems increases. fig.14 comparative pie-charts, depicting the relative (%) answers of the students to the question referring to biodiversity and the geographical latitude, in the preand post-test. objective (iv). after interacting with the word file “exercise in the food-web of the barn owl” (fig.4), students were confronted with two questions in their worksheets: q1: which one of the food-webs is more complex? the one of the north or the one of the south? justify your answers. q2: in which one of the two webs, the barn owl faces greater dangers of becoming extinct? it was considered as encouraging that in both questions the students selected the correct statement, which means: in q1 : 63 of the 85 students (74.11%) answered “the food-web of the south”. in q1 : 55 of the 85 students (64.71%) answered “in the food-web of the north”. advances in systems science and applications (2012) vol.12 no.4 369 table 5 the answers of the students referring to stability as related to food-web complexity. “i agree” “i disagree” “do not know/ other answers i cannot answer” pre-test 46 21 11 7 post-test 68 11 5 1 furthermore, both in the pre-test and the post-test, there was the question:‘do you agree or disagree with the following statement: “the more complex the foodweb is in an ecosystem, within an area, the more stable the ecosystem is, i.e. the more difficult it is for species populations to become extinct.” justify your selection.’ from the n=85 undergraduate students, the results were: it is apparent that the opinion of students, after interacting with the software and with the barn-owl-pellet data, lies closer to that which prevails in science nowadays. additionally, students’ ability to correlate food-web interconnectedness with ecological stability improved through their interaction with the “aggregate model” of the netlogo “wolf sheep predation (docked)” model. following attempts to ‘destabilize’ the system by moving the sliders in fig.5, students were asked to fill-in the cloze question in the worksheet: the system becomes unstable, i.e. one population is extinct, in n1 =. . . cases or in n1=. . . combinations of the three slider values [here we expect the same integer number]. these values are . . . . . . . . . the same was repeated for the systemmodeler food-web without grass (fig.10) this new system becomes unstable, i.e. one population is extinct, in n2 =. . . cases or in n2=. . . . combinations of the two slider values [here we expect the same integer number]. these values are. . . . . . . in the first case n1 varied from 1 to 6, and in the second case n2 varied from 0 to 2. the exact results are depicted in table 6. table 6 the findings for the n=85 students, as regards the (integer) number of cases they succeeded in destabilizing the system of populations. integer number mean value standard deviation n1 3,418 1,082 n2 0,760 0,112 the drawn here is that the more “complex” the food-web the more difficult it is to “bring it to instability”. 370 kostas karamanos:ecosystem food-webs as dynamic systems... 5 conclusions the overall teaching sequence used, including drawing of shapes, interaction with simple computer software, oral instruction, as well as discussion and personal interviews with the undergraduate students, seemed to have had adequate results in their conceptualizing the systemic properties and comportment of food-webs. the students within the evaluation phase in each worksheet, as well as during the post-test stage, gave answers and descriptions regarding: food-web structure, relation of food-web complexity with ecosystem stability and the dynamic system properties of the food-web, which were close enough to the scientifically accepted ones.the research is still continuing with interviews and new samples of students are being used for instruction. references [1] green d.g, sadedin s. & leishman t. g. (2009), “self-organisation”, in: s. e jorgensen, (editor), ecosystem ecology. amsterdam: academic press, pp.98-106. [2] dresner m. (2008), “using research projects and qualitative conceptual modeling to increase novice scientists’ understanding of ecological complexity”, ecological complexity, vol.5, no.3, pp.216-221. [3] jacobson m, j. &wilensky u. (2006), “complex systems in education: scientific and educational importance and implications for the learning sciences”, journal of the learning sciences. vol.15, no.1, pp.11-34. [4] jacobson m, j. (2001), “problem solving, cognition and complex systems: differences between experts and novices”, complexity, vol.6, no.3, pp.41-49. [5] d’ apollonia s, charles e, s, & boyd g. m. (2004), “acquisition of complex systemic thinking: mental models of evolution”, educational research and evaluation, vol.10, no.4-6, pp.499-521. [6] charles e.s, & d’ apollonia s. (2004), “developing a conceptual framework to explain emergent causality: overcoming ontological beliefs to achieve conceptual change”, in: k. forbus, d. gentner & t. reiger (editors), proceedings of the 26th annual cognitive science society. mahwah, n.j: lawrence erlbaum associates. [7] asher w. (2001), “coping with complexity and organizational interests in natural resource management”, ecosystems, vol.4, pp.742-757. advances in systems science and applications (2012) vol.12 no.4 371 [8] cadenasso m. l, pickett s.t, a. & grove j m. (2006), “dimensions of ecosystem complexity: heterogeneity, connectivity and history”, ecological complexity, vol.3, no.1, pp.1-12. [9] carpenter s.r. & kitchell j. f. (1988), “introduction”, in: s.r. carpenter & j. f. kitchell (editors), complexity in lake communities, new york: springerverlag, pp.1-8. [10] macarthur r. (1955), “fluctuations of animal populations and a measure of community stability”, ecology, vol.36, pp.533-536. [11] margalef r. (1968), perspectives in ecological theory, chicago: university of chicago press. [12] mccann k s. (2000), “the diversity-stability debate”, nature, no.405, pp.228-233. [13] wilensky u. (1999), “netlogo”, available at: http://ccl.northwestern.edu/ net logo. center for connected learning and computer-based modeling, northwestern university evanston, il. [14] annenberg media. (2007), “ecosystems, unit 4”, in “the habitable planet: a systems approach to environmental science-professional development guide”, a course produced by the harvard-smithsonian center for astrophysics & the harvard university center for the environment, [online]. available at: http://www.learner.org/courses/envsci/support/guide unit4.p df. [accessed 13 august 2010] [15] wells kristi. (2007), “the barn owl pellet, science 3450, science methods 5-8. solo unit project 12-04-07”, bemidji state university. [online] available at: http://www.bemidjistate.edu/academics/departments/science/k12-scie nce-units/barn-owl-project-grade7-biology.pdf. [accessed 13 august 2010] [16] wilensky u. (2005), netlogo wolf sheep predation (docked) model. available at: http://ccl.northwestern.edu/netlogo/models/wolfsheeppredation (docked). center for connected learning and computer-based modeling, northwestern university, evanston, il. [17] magntorn o, & helldén, g. (2007), “reading nature from a ‘bottom-up’ perspective”, journal of biological education, vol.41, no.2, pp.68-75. [18] draper f. (1993), “a proposed sequence for developing systems thinking in a grades 4-12 curriculum”, system dynamics review, vol.9, no.2, pp.207-214. 372 kostas karamanos:ecosystem food-webs as dynamic systems... [19] richmond b.( 1993), “systems thinking: critical thinking skills for the 1990s and beyond”, system dynamics review, vol.9, no.2, pp.113-133. [20] pauly d. (1998) “fishing down marine food-webs’ as an integrative concept”, proceedings of the expo’98 conference on ocean food-webs and economic productivity, lisbon, portugal, 1-3 july 1998.[online]. available at: http://cordis.europa.eu/inco/fp5/icons/pauly1.gif. [accessed 13august 2010]. [21] gallegos l, jerezano m e, & flores f. (1994), “preconceptions and relations used by children in the construction of food chains”, journal of research in science teaching, vol.31, no.3, pp.259-272. corresponding author aristotelis gkiolmas can be contacted at: agkiolm@primedu.uoa.gr advances in systems science and applications (2014) vol.14 no.3 286-293 empirical study on location choice of foreign banks in china tong xue school of international economics and trade, bisu, beijing, china abstract using the panel data from 2006 to 2010, this paper analyses the determinants of location choice of foreign banks in china with the amount of assets of foreign banks in different regions as dependant variables. the empirical results show that market opportunity and following clients factors are both important determinants for foreign banks location choice. meanwhile the presence of foreign banks tends to be more in financial centers of china, holding everything else being equal. however, the non-performing loan ratio of commercial banks of one region is not important factor for foreign banks to consider for now. keywords foreign banks, china, location choice. 1 introduction with the complete liberalization of banking industry at the end of 2006, foreign banks accelerated the pace of entering into china. by 2010, 185 banks from 45 countries and regions set up 216 representative offices in china; 37 banks from 14 countries and regions were locally incorporated which maintained 223 branches. in addition, there were 90 foreign bank branches established by 74 banks from 25 countries and regions. and the total assets of foreign banking institutions in china increased 29.13 percent year-on-year to rmb1.74 trillion (china banking regulatory commission (cbrc), 2010). according to the authors statistics, these foreign banks maintained business branches in 34 cities in china. shanghai, beijing, shenzhen, guangzhou and tianjin have the largest number of business branches respectively, in which shanghai is the top one and has 102 business branches, accounting for 30% of the number of business branches of foreign banks in china. it can be seen that there is an agglomeration effect in the location choice made by foreign banks in china, i.e. the offices of foreign banks concentrate in those more developed and first-deregulated cities, such as shanghai, beijing, etc. therefore, what factors determine the location choice of foreign banks in china? and which regions can attract foreign banks to do more business? these are interesting problems worth concerning. among foreign literatures, in addition to the studies on the location choices among various host countries, there are some researches on the internal location choices in a host country, such as literature[1] and [2]. in china, some researchers made some descriptive analysis on the features of location choices of foreign banks in china (for example [3-5]) using panel data, zhang and yang made the empirical analysis on the determinants of location of foreign banks in 16 cities in advances in systems science and applications (2014) vol.14 no.3 287 china for the first time[6]. li examined the determinants in 24 cities on the basis of zhang and yang[6-7]. he and yeung[8] examined the locational distribution in 32 cities with conditional logit models. existing studies, however, all took the number of institutions but not the amount of assets of foreign banks as the dependent variables, while the number of institutions cannot accurately reflect the business development status of foreign banks. this paper aims to contribute to the existing literature by using the amount of assets of foreign banks as the dependent variable to investigate the determinants of location choice of foreign banks in china. 2 hypothesis development on the basis of previous studies and taking into account the availability of data, this paper empirically examines four factors that affect the location distribution of foreign banks in china. 2.1 market opportunity existing literatures prove that market opportunity is one of the important factors that affect the location choice of foreign banks. the market opportunity is generally measured from two aspects, economic development level and size of financial sector. the higher the level of economic development, (which is usually measured by gdp, gdp per capita or the growth rate of gdp), and the larger the financial markets, the demand of households and businesses for financial services is greater, which can bring more profit opportunity for foreign banks. many studies have proved it. for example, buch and yamori proved that from the perspective of national level, the entry of foreign banks is positively related with the gdp or gdp per capita of host countries[9-10]. the further study by focarelli and pozzolo showed that profit opportunities resulting from a high expected economic growth and the prospect of competing with relatively less efficient banks appear to be a key factor affecting the expansion abroad[11]. goldberg and grosse found that foreign bank presence in various states in the u.s., as measured by assets or offices, is positively related with the size of the states banking market[2]. studies on location distribution of foreign banks in china have the identical results. zhang and yang proved that disposable income per capita is positively related with the office number of that city[6]. li proved that the office number of foreign banks is significantly positively related with both disposable income per capita and credit aggregates[7]. therefore, based on the above discussion, the following are inferred: h1: foreign bank presence is drawn to the economically developed areas. h2: foreign bank presence is drawn to the areas with large banking market sizes. 288 anne talkington: an extension of a logistic model for microbial kinetics 2.2 “follow-the-customer” factor theories on multinational banking indicated that one of the motives of multinational banks to go abroad is to follow the customers. numerous empirical studies also proved that foreign banks presence is significantly correlated with the foreign trade or fdi volume between the home and host countries. meanwhile, some studies showed the foreign trade or fdi volumes of one area play important role in location choice of foreign banks within one country. for example, goldberg and grosse proved that the amount of fdi in the state is a significant determinant for attracting foreign bank assets, while the foreign trade volume is not[2]. studies on china made by zhang and yang and li also proved that number of foreign banks’ offices in one city is positively related with the foreign trade volume, while the amount of fdi is not a significant determinant[6-7]. therefore, the following is speculated: h3: foreign bank presence is drawn to the areas with more foreign trade or fdi volume. 2.3 institutional factors institutional factors are also important for location distribution of foreign banks. numerous studies found that foreign banks tend to enter the financial centers of one country, which result in agglomeration effect. setting up offices in financial center makes it convenient for foreign banks to establish good relationships with regulatory authorities and other financial institutions. it is also convenient to gain funds and information with low cost and to receive specialized services from other companies. studies made by bagchi-sen and o huallachin all showed foreign banks in the u.s. concentrated in the financial centers[1,12]. he and yeung showed that foreign banks agglomerate in beijing, shanghai, and cities hosting the regional branches of the people’s bank of china (central bank of china)[8]. based on the above discussion, it is reasonable to deduce the following: h4: foreign bank presence agglomerates in financial centers in china. 2.4 risk factor in consideration of risk control, foreign banks usually make the regional risk assessments for location choice. therefore, financial risk status of one region is also one of factors affecting the location distribution. empirical study made by li used weighted average nonperforming loan ratios of 24 cities as independent variable. but the coefficient was not significant[7]. this paper tries to study whether the asset distribution of foreign banks is related with the risk status of one region and thus the following hypothesis is proposed: h5: foreign bank presence is negatively related with the financial risk level of one region. advances in systems science and applications (2014) vol.14 no.3 289 3 methodology and data sources 3.1 dependent variable as mentioned before, different from the previous studies on china, this paper uses the amount of assets of foreign banks but not the number of offices in different regions as the dependent variable. however, due to the availability of data, the amount of assets of foreign banks in different provinces, autonomous regions, or municipalities but not in different cities is used, which is denoted as fbasset (in natural logarithmic form). the data is from financial performance report of each region from 2006 to 2010. since foreign banks did not enter some provinces until recently, this paper selects 20 among 31 provinces, autonomous regions, or municipalities which have foreign bank presence at least after 20081. 3.2 explanatory variables based on the previous studies and the availability of data, definitions of explanatory variables are as follows: (1) the natural logarithmic form of gdp and the growth rate of gdp are introduced to measure the level of economic development, which are denoted as gdp and gdpgrowth respectively and the data are from china statistical yearbook. (2) the ratio of total banking loan or total amount of banking assets to gdp are introduced to measure the banking market sizes of one region. the reason is that compared with provinces and autonomous regions, municipalities are much smaller with respect to area and population. then it is biased to use the absolute value of total banking loan or banking assets to measure the banking market sizes. two variables are denoted as loan and asset respectively and the data are from financial performance report. (3) to test the “follow-the-customer” hypothesis (h3), the natural logarithmic form of volume of foreign trade (trade) and the natural logarithmic form of realized amount of foreign direct investment (fdi) are included as explanatory variables. the data are from china statistical yearbook. (4) dummy variable (center) is introduced to test whether foreign bank presence agglomerates in financial centers in china (h4), which represents the locations of regional branches of the central bank, assigning a value of “1” for beijing, shanghai, tianjin, chongqin, guangdong, liaoning, shandong, sichuan, jiangsu, hubei and shanxi, and zero for other provinces. (5) the nonperforming loan ratios of commercial banks of each region (npl) are introduced to measure the financial risk status. the data are from distribution of npls of commercial banks by region , annual report of china banking reg1twenty provinces, autonomous regions, or municipalities include beijing, shanghai, tianjin, chongqin, guangdong, fujian, liaoning, shandong, sichuan, jiangsu, hubei, hunan, jiangxi, anhui, shanxi, guangxi, yunnan, heilongjiang, zhejiang, hainan. 290 anne talkington: an extension of a logistic model for microbial kinetics ulatory commission of each year. in summary, regression model based on panel data is as follows: fbassetit =cit + β1oppotunityit + β2followit+ β3centeri + β4nplit + εit where i =1, ...n, t = 1, ..., t 4 regression results first, the correlation test of all variables except dummy variable is made and the result is shown in table 1. table 1 correlation coefficients among variables fbasset gdp gdpgrowth loan asset trade fdi npl fbasset 1.000 gdp 0.564 1.000 (0.000) gdpgrowth -0.060 -0.050 1.000 (0.552) (0.615) loan 0.529 -0.006 -0.335 1.000 (0.000) (0.952) (0.001) asset 0.540 0.010 -0.326 0.926 1.000 (0.000) (0.919) (0.001) (0.000) trade 0.831 0.789 -0.158 0.374 0.372 1.000 (0.000) (0.000) (0.114) (0.000) (0.000) fdi 0.713 0.790 0.072 0.121 0.127 0.866 1.000 (0.000) (0.000) (0.471) (0.228) (0.204) (0.000) npl -0.217 -0.427 0.234 -0.476 -0.479 -0.335 -0.354 1.000 (0.043) (0.000) (0.029) (0.000) (0.000) (0.002) (0.001) notes: p values are in parentheses. as the table 1 shows, among the explanatory variables, loan and asset are highly correlated and the correlation coefficient is 0.926. trade and fdi are also highly correlated with the coefficient being 0.866. to mitigate the multicollinearity issue of estimates, the interaction of loan and asset, and trade and fdi are introduced in the model, to show the effect of banking market sizes and “follow-the-customer” factor on the location choice respectively. besides, gdp is also highly correlated with trade and fdi. so the significance of gdp and trade*fdi is tested separately. advances in systems science and applications (2014) vol.14 no.3 291 table 2 regression results on location choice of foreign banks in china (20062010) variable (1) (2) gdpgrowth 0.0630* (1.66) gdp 1.5105*** (5.37) asset*loan 0.1550*** 0.1112** (2.97) (2.13) fdi*trade 0.1054*** (5.45) center 1.7155*** 2.0640*** (3.57) (3.26) npl -0.0143 0.0099 (-0.76) (0.51) adjusted r2 0.5830 0.4990 number of observation 82 89 hausman test(p value) 2.05 6.18 (0.73) (0.10) estimation metho random effects random effects notes: (1) the sample period is from 2006 to 2010 and the number of cross-section is 20. the actual sample is less than 100 because some data are default or wrong. (2) the constants are left out in the results. (3) *significant at the 10% level; ** significant at the 5% level; *** significant at the 1% level, t values are in parentheses. (4) hausman-test statistic shows whether the random effects method is suitable. the empirical results of table 2 show that the asset distribution of foreign banks in china is significantly positively related with gdp or gdp growth rate of that region, demonstrating that the level of economic development is one of important determinants for foreign banks presence. thus, h1 is supported. in addition, the disposable income per capita is also introduced into the model instead of gdp and the result is similar. this is slightly different from the results of zhang and yang and li which showed the number of offices of foreign banks is positively related with the disposable income per capita but not gdp[6-7]. both of results (1) and (2) verify that the foreign bank presence is significantly positively related with the total banking loan or total amount of banking assets, proving foreign bank presence is drawn to the areas with large banking market sizes (h2). this is consistent with li[7]. the coefficient of trade*fdi is significantly positive, proving that foreign bank presence is drawn to the areas with more foreign trade or fdi volume (h3). foreign banks may follow their customers and enter the area with more international trade or fdi, providing financial services to those customers. this result is consistent with other studies on china. 292 anne talkington: an extension of a logistic model for microbial kinetics the coefficient of center is significantly positive and thus h4 is supported very well, showing that holding everything else being equal, foreign banks tend to operate in financial centers to obtain the advantages in funds, information and others. this result is consistent with he and yeung[8]. h5 is not proved since the coefficient of npl is positive or negative but not significant, showing that the foreign bank presence is not significantly related with the nonperforming loan ratios of commercial banks, which is consistent with li[7]. this indicates that foreign banks do not take the nonperforming loan ratios as an important factor when making location choice. one of reasons may be that foreign banks do not consider the nonperforming loan ratios of commercial banks equivalently with the financial risk status of that region. 5 conclusion based on the previous studies and using the panel data from 2006 to 2010, this paper analyses the determinants of location choice of foreign banks in china with the amount of assets of foreign banks in different regions as dependant variables. market opportunity, “follow-the-customer”, institution factor and risk factor are all included. the empirical results show that when foreign banks make the location choice, market opportunity is one of the important factors so that the economically developed regions with large banking sizes are more attractive to foreign banks. on the other hand, foreign bank also consider “follow-the-customer” strategy so that foreign bank presence is drawn to the areas with more foreign trade or fdi volume. meanwhile the presence of foreign banks tends to be more in financial centers of china, holding everything else being equal. however, the non-performing loan ratio of commercial banks of one region is not important factor for foreign banks to consider for now. acknowledgements this research is supported by funding project for academic human resources development in institutions of higher learning under the jurisdiction of beijing municipality (phr(ihlb)). references [1] bagchi-sen, s. (1991), “the location of foreign direct investment in finance, insurance, and real estate in the united states”, geografiska annaler, vol.73b, pp.187-197. [2] goldberg, l.g., and r. grosse. (1994), “location choice of foreign banks in the united states”, journal of economic business, no.46, pp.367-379. advances in systems science and applications (2014) vol.14 no.3 293 [3] zheng, bohong., jianzhong tang. (2001), “study on location of multinational banks in china”, world regional studies, vol.10, no.4, pp.21-28. [4] xie, shouhong, mingfeng wang. (2004), “a study on the location of foreign financial institutions in china”, human geography, vol.19, no.3, pp.50-55. [5] bai,yongping, fajun ji. (2010), “dynamic development and spatial distribution of foreign-funded banks in china”, journal of northwest normal university (natural science) , no.1, pp.108-113. [6] zhang, hongjun, chaojun yang. (2007), “empirical study on location choice and motivation of foreign banks in china”, journal of financial research, no.9, pp.160-172. [7] li, aixi. (2009), “behavior and empirical study on location choice of foreign banks in china”, journal of international trade, no.3, pp.118-124. [8] he, canfei, and godfrey yeung. (2010), “the locational distribution of foreign banks in china: a disaggregated analysis”, regional studies, vol.45, no.6, pp.733-754. [9] buch, c.m. (2000), “why do banks go abroad? evidence from german data”, financial markets, institutions & instruments, vol.9, no.1, pp.33-67. [10] nobuyoshi yamori. (1998), “a note on the location choice of multinational banks: the case of japanese financial institutions”, journal of banking & finance, vol.22, no.1, pp.109-120. [11] focarelli d. and a.f. pozzolo. (2005), “where do banks expand abroad? an empirical analysis”, journal of business, no.78, pp.2435-2463. [12] o huallachin b. (1994), “foreign banking in the american urban system of financial organization”, economic geography, vol.70, no.3, pp.206-228. corresponding author tong xue can be contacted at: xuetong@bisu.edu.cn. a new modified block iterative algorithm for a system of equilibrium problems and a fixed point set of uniformly φ-asymptotically nonexpansive mappings∗ siwaporn saewan and poom kumam department of mathematics, faculty of science, king mongkut’s university of technology thonburi (kmutt),bangmod, bangkok 10140, thailand email: 52501406@st.kmutt.ac.th, poom.kum@kmutt.ac.th abstract in this paper, we construct a new modified block hybrid projection algorithm for finding a common element of the set of common fixed points of an infinite family of closed and uniformly quasi -φasymptotically nonexpansive mappings, the set of the variational inequality for an α-inverse-strongly monotone operator, the set of solutions of a system of equilibrium problems. moreover, we obtain a strong convergence theorem for the sequences generated by this process in the framework banach spaces. the results presented in this paper improve and generalize some well-known results in the literature. keywords modified block iterative algorithm inverse-strongly monotone operator variational inequality a system of equilibrium problem uniformly quasi-φ-asympto tically nonexpansive mapping. 1. introduction let c be a nonempty closed convex subset of a real banach space e with ‖ · ‖ and e∗ the dual space of e and a : c → e∗ be an operator. the classical variational inequality problem for an operator a is to find x∗ ∈ c such that 〈ax∗, y − x∗〉 ≥ 0, ∀y ∈ c. (1) the set of solution of (1) is denote by v i(a,c). recall that let a : c → e∗ be a mapping. then a is called (i) monotone if 〈ax−ay, x− y〉 ≥ 0, ∀x, y ∈ c, (ii) α−inverse-strongly monotone if there exists a constant α > 0 such that 〈ax−ay, x− y〉 ≥ α‖x− y‖2, ∀x, y ∈ c. such a problem is connected with the convex minimization problem, the complementary problem, the problem of finding a point x∗ ∈ e satisfying ax∗ = 0. ∗this work was completed with the support by the national research university project of thailand’s office of the higher education commission (under nru-csec project no.54000267). issn 1078-6236 international institute for general systems studies, inc. advances in systems science and applications (2011), vol. 11, no. 1-2 123-148 quasilet {fi}i∈γ : c × c −→ r be a bifunction, {ϕi}i∈γ : c −→ r be a real-valued function, where γ is an arbitrary index set. the system of equilibrium problems, is to find x ∈ c such that fi(x, y) ≥ 0, i ∈ γ, ∀y ∈ c. (2) if γ is a singleton, then problem (2) reduces to the equilibrium problem, is to find x ∈ c such that f(x, y) ≥ 0, ∀y ∈ c. (3) the above formulation (3) was shown in [5] to cover monotone inclusion problems, saddle point problems, variational inequality problems, minimization problems, optimization problems, variational inequality problems, vector equilibrium problems, nash equilibria in noncooperative games. in addition, there are several other problems, for example, the complementarity problem, fixed point problem and optimization problem, which can also be written in the form of an ep (f). in other words, the ep (f) is an unifying model for several problems arising in physics, engineering, science, optimization, economics, etc. in the last two decades, many papers have appeared in the literature on the existence of solutions of ep (f); see, for example [5, 13] and references therein. some solution methods have been proposed to solve the ep (f); see, for example, [5, 13, 15, 16, 21, 24, 30, 29, 28, 35, 46] and references therein. for each p > 1, the generalized duality mapping jp : e → 2e ∗ is defined by jp(x) = {x∗ ∈ e∗ : 〈x, x∗〉 = ‖x‖p, ‖x∗‖ = ‖x‖p−1} for all x ∈ e. in particular, j = j2 is called the normalized duality mapping. if e is a hilbert space, then j = i , where i is the identity mapping. consider the functional defined by φ(x, y) = ‖x‖2 − 2〈x, jy〉+ ‖y‖2, ∀x, y ∈ e. (4) as well know that if c is a nonempty closed convex subset of a hilbert space h and pc : h → c is the metric projection of h onto c, then pc is nonexpansive. this fact actually characterizes hilbert spaces and consequently, it is not available in more general banach spaces. it is obvious from the definition of function φ that (‖x‖ − ‖y‖)2 ≤ φ(x, y) ≤ (‖x‖+ ‖y‖)2, ∀x, y ∈ e. (5) if e is a hilbert space, then φ(x, y) = ‖x − y‖2, for all x, y ∈ e. on the author hand, the generalized projection (alber [2]) πc : e → c is a map that assigns to an arbitrary point x ∈ e the minimum point of the functional φ(x, y), that is, πcx = x̄, where x̄ is the solution to the minimization problem φ(x̄, x) = inf y∈c φ(y, x), (6) existence and uniqueness of the operator πc follows from the properties of the functional φ(x, y) and strict monotonicity of the mapping j (see, for example, [1, 2, 12, 17, 37]). remark 1.1. if e is a reflexive, strictly convex and smooth banach space, then for x, y ∈ e, φ(x, y) = 0 if and only if x = y. it is sufficient to show that if φ(x, y) = 0 then x = y. from ( 4 ), we have ‖x‖ = ‖y‖. this implies that 〈x, jy〉 = ‖x‖2 = ‖jy‖2. from the definition of j, one has jx = jy. therefore, we have x = y; see [12, 37] for more details. 124 saewan:a new modified block iterative algorithm for a system of. . . . . . let c be a closed convex subset of e, a mapping t : c → c is said to be l-lipschitz continuous if ‖tx− ty‖ ≤ l‖x− y‖,∀x, y ∈ c and a mapping t is said to be nonexpansive if ‖tx − ty‖ ≤ ‖x − y‖, ∀x, y ∈ c. a point x ∈ c is a fixed point of t provided tx = x. denote by f (t ) the set of fixed points of t ; that is, f (t ) = {x ∈ c : tx = x}. recall that a point p in c is said to be an asymptotic fixed point of t [31] if c contains a sequence {xn} which converges weakly to p such that limn→∞ ‖xn − txn‖ = 0. the set of asymptotic fixed points of t will be denoted by f̃ (t ). a mapping t from c into itself is said to be relatively nonexpansive [25, 36, 45] if f̃ (t ) = f (t ) and φ(p, tx) ≤ φ(p, x) for all x ∈ c and p ∈ f (t ). the asymptotic behavior of a relatively nonexpansive mapping was studied in [6, 7, 8]. t is said to be φ-nonexpansive, if φ(tx, ty) ≤ φ(x, y) for x, y ∈ c. t is said to be relatively quasi-nonexpansive if f (t ) 6= ∅ and φ(p, tx) ≤ φ(p, x) for all x ∈ c and p ∈ f (t ). t is said to be quasi-φ-asymptotically nonexpansive if f (t ) 6= ∅ and there exists a real sequence {kn} ⊂ [1,∞) with kn → 1 such that φ(p, tnx) ≤ knφ(p, x) for all n ≥ 1 x ∈ c and p ∈ f (t ). we note that the class of relatively quasi-nonexpansive mappings is more general than the class of relatively nonexpansive mappings [6, 7, 8, 23, 33] which requires the strong restriction: f (t ) = f̃ (t ). a mapping t is said to be closed if for any sequence {xn} ⊂ c with xn → x and txn → y, then tx = y. it is easy to know that each relatively nonexpansive mapping is closed. definition 1.2. ([9]) (1) let {ti}∞i=1 : c → c be a sequence of mapping. {ti}∞i=1 is said to be a family of uniformly quasi-φ-asymptotically nonexpansive mappings, if ∩∞i=1f (ti) 6= ∅, and there exists a sequence {kn} ⊂ [1,∞) with kn → 1 such that for each i ≥ 1 φ(p, tni x) ≤ knφ(p, x), ∀p ∈ ∩∞i=1f (ti), x ∈ c, ∀n ≥ 1. (7) (2) a mapping t : c → c is said to be uniformly l-lipschitz continuous, if there exists a constant l > 0 such that ‖tnx− tny‖ ≤ l‖x− y‖, ∀x, y ∈ c. (8) remark 1.3. it is easy to see that anα−inverse-strongly monotone is monotone and 1 α -lipschitz continuous. in 2004, matsushita and takahashi [22] introduced the following iteration: a sequence {xn} defined by xn+1 = πcj −1(αnjxn + (1− αn)jtxn), (9) where the initial guess element x0 ∈ c is arbitrary, {αn} is a real sequence in [0, 1], t is a relatively nonexpansive mapping and πc denotes the generalized projection from e onto a closed convex subset c of e. they proved that the sequence {xn} converges weakly to a fixed point of t . in 2005, matsushita and takahashi [23] proposed the following hybrid iteration method (it is also called the cq method) with generalized projection for relatively nonexpansive mapping advances in systems science and applications (2011), vol. 11, no. 1-2 125 t in a banach space e: x0 ∈ c chosen arbitrarily, yn = j−1(αnjxn + (1− αn)jtxn), cn = {z ∈ c : φ(z, yn) ≤ φ(z, xn)}, qn = {z ∈ c : 〈xn − z, jx0 − jxn〉 ≥ 0}, xn+1 = πcn∩qnx0. (10) they proved that {xn} converges strongly to πf (t )x0, where πf (t ) is the generalized projection from c onto f (t ). in 2008, iiduka and takahashi [14] introduced the following iterative scheme for finding a solution of the variational inequality problem for an inversestrongly monotone operator a in a 2-uniformly convex and uniformly smooth banach space e : x1 = x ∈ c and xn+1 = πcj −1(jxn − λnaxn), (11) for every n = 1, 2, 3, . . ., where πc is the generalized metric projection frome ontoc, j is the duality mapping from e into e∗ and {λn} is a sequence of positive real numbers. they proved that the sequence {xn} generated by (11) converges weakly to some element of v i(a,c). takahashi and zembayashi [39, 40], studied the problem of finding a common element of the set of fixed points of a nonexpansive mapping and the set of solutions of an equilibrium problem in the framework of banach spaces. in 2009, wattanawitoon and kumam [41] using the idea of takahashi and zembayashi [39] extend the notion from relatively nonexpansive mappings or φ-nonexpansive mappings to two relatively quasi-nonexpansive mappings and also proved some strong convergence theorems to approximate a common fixed point of relatively quasi-nonexpansive mappings and the set of solutions of an equilibrium problen in the framework of banach spaces. cholamjiak [10], proved the following iteration: zn = πcj −1(jxn − λnaxn), yn = j−1(αnjxn + βnjtxn + γnjszn), un ∈ c such that f(un, y) + 1 rn 〈y − un, jun − jyn〉 ≥ 0, ∀y ∈ c, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn), xn+1 = πcn+1x0, (12) where j is the duality mapping on e. assume that αn, βn and γn are sequence in [0, 1]. then {xn} converges strongly to q = πfx0, where f := f (t ) ∩ f (s) ∩ ep (f) ∩ v i(a,c). in 2010, saewan et al. [33] introduced a new hybrid projection iterative scheme which is difference from the algorithm (12) of cholamjiak in [10, theorem 3.1] for two relatively quasinonexpansive mappings in a banach space. motivated by the results of takahashi and zembayashi [40], cholumjiak and suantai [11] proved the following strong convergence theorem by the hybrid iterative scheme for approximation of common fixed point of countable families of relatively quasi-nonexpansive mappings in a uniformly convex and uniformly smooth banach space: x0 ∈ e, x1 = πc1x0, c1 = c yn,i = j−1(αnjxn + (1− αn)jtxn, ) un,i = t fmrm,nt fm−1 rm−1,n · · ·t f1 r1,nyn,i cn+1 = {z ∈ cn : supi>1 φ(z, jun,i) ≤ φ(w, jxn)}, xn+1 = πcn+1x0, n ≥ 1. (13) 126 saewan:a new modified block iterative algorithm for a system of. . . . . . then, they proved that under certain appropriate conditions imposed on {αn}, and {rn,i}, the sequence {xn} converges strongly to πcn+1x0. we note that the block iterative method is a method which often used by many authors to solve the convex feasibility problem (see, [18, 20], etc.). in 2008, plubtieng and ungchittrakool [27] established strong convergence theorems of block iterative methods for a finite family of relatively nonexpansive mappings in a banach space by using the hybrid method in mathematical programming. chang et al. [9] proposed the modified block iterative algorithm for solving the convex feasibility problems for an infinite family of closed and uniformly quasi-φasymptotically nonexpansive mapping, they obtain the strong convergence theorems in a banach space. in 2010, saewan and kumam [34] obtain the following result for the set of solutions of the generalized equilibrium problems and the set of common fixed points of an infinite family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings in a uniformly smooth and strictly convex banach space e with kadec-klee property. very recently, qin, cho and kang [28] purposed the problem of approximating a common fixed point of two asymptotically quasi-φ-nonexpansive mappings based on hybrid projection methods. strong convergence theorems are established in a real banach space. h. zegeye, e. u. ofoedu and n. shahzad [46] introduced an iterative process which converges strongly to a common element of set of common fixed points of countably infinite family of closed relatively quasinonexpansive mappings, the solution set of generalized equilibrium problem and the solution set of the variational inequality problem for a α-inverse strongly monotone mapping in banach spaces. motivated and inspired by the work of chang et al. [9], qin et al. [30], takahashi and zembayashi [39], wattanawitoon and kumam [41], zegeye [44] and saewan and kumam [34], we introduce a new modified block hybrid projection algorithm for finding a common element of the set of the variational inequality for an α-inverse-strongly monotone operator, the set of solutions of the system of equilibrium problems and the set of common fixed points of an infinite family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings in a 2-uniformly convex and uniformly smooth banach space. the results presented in this paper improve and generalize some well-known results in the literature. 2. preliminaries a banach space e is said to be strictly convex if ‖x+y 2 ‖ < 1 for all x, y ∈ e with ‖x‖ = ‖y‖ = 1 and x 6= y. let u = {x ∈ e : ‖x‖ = 1} be the unit sphere of e. then a banach space e is said to be smooth if the limit lim t→0 ‖x+ ty‖ − ‖x‖ t exists for each x, y ∈ u. it is also said to be uniformly smooth if the limit is attained uniformly for x, y ∈ u . let e be a banach space. the modulus of convexity of e is the function δ : [0, 2]→ [0, 1] defined by δ(ε) = inf{1− ‖x+ y 2 ‖ : x, y ∈ e, ‖x‖ = ‖y‖ = 1, ‖x− y‖ ≥ ε}. advances in systems science and applications (2011), vol. 11, no. 1-2 127 a banach space e is uniformly convex if and only if δ(ε) > 0 for all ε ∈ (0, 2]. let p be a fixed real number with p ≥ 2. a banach space e is said to be p-uniformly convex if there exists a constant c > 0 such that δ(ε) ≥ cεp for all ε ∈ [0, 2]; see [3, 38] for more details. observe that every p-uniformly convex is uniformly convex. one should note that no a banach space is p-uniformly convex for 1 < p < 2. it is well known that a hilbert space is 2-uniformly convex, uniformly smooth. it is also known that if e is uniformly smooth, then j is uniformly norm-to-norm continuous on each bounded subset of e. remark 2.1. the following basic properties can be found in cioranescu [12]. (i) if e is a uniformly smooth banach space, then j is uniformly continuous on each bounded subset of e. (ii) if e is a reflexive and strictly convex banach space, then j−1 is norm-weak∗-continuous. (iii) if e is a smooth, strictly convex, and reflexive banach space, then the normalized duality mapping j : e → 2e ∗ is single-valued, one-to-one, and onto. (iv) a banach space e is uniformly smooth if and only if e∗ is uniformly convex. (v) each uniformly convex banach space e has the kadec-klee property, that is, for any sequence {xn} ⊂ e, if xn ⇀ x ∈ e and ‖xn‖ → ‖x‖, then xn → x. we also need the following lemmas for the proof of our main results. lemma 2.2. (beauzamy [4] and xu[42]). if e be a 2-uniformly convex banach space. then for all x, y ∈ e we have ‖x− y‖ ≤ 2 c2 ‖jx− jy‖, where j is the normalized duality mapping of e and 0 < c ≤ 1. the best constant 1 c in lemma is called the p-uniformly convex constant of e. lemma 2.3. (beauzamy [4] and zalinescu [43]). ife be a p-uniformly convex banach space and let p be a given real number with p ≥ 2. then for all x, y ∈ e, jx ∈ jp(x) and jy ∈ jp(y) 〈x− y, jx − jy〉 ≥ cp 2p−2p ‖x− y‖p, where jp is the generalized duality mapping of e and 1 c is the p-uniformly convexity constant of e. lemma 2.4. (kamimura and takahashi [17]). let e be a uniformly convex and smooth banach space and let {xn} and {yn} be two sequences of e. if φ(xn, yn)→ 0 and either {xn} or {yn} is bounded, then ‖xn − yn‖ → 0. lemma 2.5. (alber [2]). let c be a nonempty closed convex subset of a smooth banach space e and x ∈ e. then x0 = πcx if and only if 〈x0 − y, jx− jx0〉 ≥ 0, ∀y ∈ c. lemma 2.6. (alber [2, lemma 2.4]). let e be a reflexive, strictly convex and smooth banach space, let c be a nonempty closed convex subset of e and let x ∈ e. then φ(y,πcx) + φ(πcx, x) ≤ φ(y, x), ∀y ∈ c. 128 saewan:a new modified block iterative algorithm for a system of. . . . . . let e be a reflexive, strictly convex, smooth banach space and j is the duality mapping from e into e∗. then j−1 is also single value, one-to-one, surjective, and it is the duality mapping from e∗ into e. we make use of the following mapping v studied in alber [2] v (x, x∗) = ‖x‖2 − 2〈x, x∗〉+ ‖x∗‖2, (14) for all x ∈ e and x∗ ∈ e∗, that is, v (x, x∗) = φ(x, j−1(x∗)). lemma 2.7. (alber [2]). let e be a reflexive, strictly convex smooth banach space and let v be as in (14) . then v (x, x∗) + 2〈j−1(x∗)− x, y∗〉 ≤ v (x, x∗ + y∗), for all x ∈ e and x∗, y∗ ∈ e∗. a set valued mapping b : e ⇒ e∗ with graph g(b) = {(x, x∗) : x∗ ∈ bx}, domain d(b) = {x ∈ e : bx 6= ∅}, and rang r(b) = ∪{bx : x ∈ d(b)}. b is said to be monotone if 〈x − y, x∗ − y∗〉 ≥ 0 whenever (x, x∗) ∈ g(b), (y, y∗) ∈ g(b). we denote a set valued operator b form e to e∗ by b ⊂ e × e∗. a monotone b is said to be maximal if its graph is not property contained in the graph of any other monotone operator. ifb is maximal monotone, then the solution set b−10 is closed and convex. let e be a reflexive, strictly convex and smooth banach space, it is knows that b is a maximal monotone if and only if r(j + rb) = e∗ for all r > 0. define the resolvent of b by jrx = xr. in other words, jr = (j + rb)−1j for all r > 0. jr is a single-valued mapping from e to d(b). also, b−1(0) = f (jr) for all r > 0, where f (jr) is the set of all fixed points of jr. define, for r > 0, the yosida approximation of b by trx = (jx − jjrx)/r for all x ∈ c. we know that trx ∈ b(jrx) for all r > 0 and x ∈ e. let a be an inverse-strongly monotone mapping of c into e∗ which is said to be hemicontinuous if for all x, y ∈ c, the mapping f of [0, 1] intoe∗, defined by f (t) = a(tx+(1−t)y), is continuous with respect to the weak∗ topology of e∗. we define by nc(v) the normal cone for c at a point v ∈ c, that is, nc(v) = {x∗ ∈ e∗ : 〈v − y, x∗〉 ≥ 0, ∀y ∈ c}. (15) lemma 2.8. (rockafellar [32]). letc be a nonempty, closed convex subset of a banach spacee and a is a monotone, hemicontinuous operator of c into e∗. let b ⊂ e × e∗ be an operator defined as follows: bv = { av +nc(v), v ∈ c; ∅, otherwise. (16) then b is maximal monotone and b−10 = v i(a,c). lemma 2.9. (chang et al.[9]) . let e be a uniformly convex banach space, r > 0 be a positive number andbr(0) be a closed ball ofe. then, for any given sequence {xi}∞i=1 ⊂ br(0) and for any given sequence {λi}∞i=1 of positive number with ∑∞ n=1 λn = 1, there exists a continuous, strictly increasing, and convex function g : [0, 2r) → [0,∞) with g(0) = 0 such that, for any positive integer i, j with i < j, ‖ ∞∑ n=1 λnxn‖2 ≤ ∞∑ n=1 λn‖xn‖2 − λiλjg(‖xi − xj‖). (17) advances in systems science and applications (2011), vol. 11, no. 1-2 129 lemma 2.10. (chang et al.[9]) . let e be a real uniformly smooth and strictly convex banach space, and c be a nonempty closed convex subset of e. let t : c → c be a closed and quasi-φ-asymptotically nonexpansive mapping with a sequence {kn} ⊂ [1,∞), kn → 1. then f (t ) is a closed convex subset of c. for solving the equilibrium problem for a bifunction f : c × c → r, let us assume that f satisfies the following conditions: (a1) f(x, x) = 0 for all x ∈ c; (a2) f is monotone, i.e., f(x, y) + f(y, x) ≤ 0 for all x, y ∈ c; (a3) for each x, y, z ∈ c, lim t↓0 f(tz + (1− t)x, y) ≤ f(x, y); (a4) for each x ∈ c, y 7→ f(x, y) is convex and lower semi-continuous. for example, let a be a continuous and monotone operator of c into e∗ and define f(x, y) = 〈ax, y − x〉,∀x, y ∈ c. then, f satisfies (a1)-(a4). the following result is in blum and oettli [5]. lemma 2.11. (blum and oettli [5]). letc be a closed convex subset of a smooth, strictly convex and reflexive banach space e, let f be a bifunction from c×c to r satisfying (a1)-(a4), and let r > 0 and x ∈ e. then, there exists z ∈ c such that f(z, y) + 1 r 〈y − z, jz − jx〉 ≥ 0, ∀y ∈ c. lemma 2.12. (takahashi and zembayashi [39]). let c be a closed convex subset of a uniformly smooth, strictly convex and reflexive banach space e and let f be a bifunction from c × c to r satisfying conditions (a1)-(a4). for all r > 0 and x ∈ e, define a mapping t fr : e → c as follows: t fr x = {z ∈ c : f(z, y) + 1 r 〈y − z, jz − jx〉 ≥ 0, ∀y ∈ c}. then the following hold: (1) t fr is single-valued; (2) t fr is a firmly nonexpansive-type mapping [19], that is, for all x, y ∈ e, 〈t fr x− t fr y, jt fr x− jt fr y〉 ≤ 〈t fr x− t fr y, jx− jy〉; (3) f (t fr ) = ep (f); (4) ep (f) is closed and convex. lemma 2.13. (takahashi and zembayashi [39]). let c be a closed convex subset of a smooth, strictly convex, and reflexive banach space e, let f be a bifunction from c ×c to r satisfying (a1)-(a4) and let r > 0. then, for x ∈ e and q ∈ f (t fr ), φ(q, t fr x) + φ(t fr x, x) ≤ φ(q, x). 130 saewan:a new modified block iterative algorithm for a system of. . . . . . 3. main results in this section, we prove the new convergence theorems for finding the set of solutions of system of generalized mixed equilibrium problems, the common fixed point set of a family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings, and the solution set of variational inequalities for an α-inverse strongly monotone mapping in a 2-uniformly convex and uniformly smooth banach space. theorem 3.1. let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach space e. for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)-(a4). let a be an α-inverse-strongly monotone mapping of c into e∗ satisfying ‖ay‖ ≤ ‖ay −au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let {si}∞i=1 : c → c be an infinite family of closed uniformly li-lipschitz continuous and uniformly quasiφ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 such that f := (∩∞i=1f (si))∩ (∩mj=1ep (fj))(∩v i(a,c)) is a nonempty and bounded subset in c. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows:  vn = πcj −1(jxn − λnaxn), zn = j−1(αn,0jxn + ∑∞ i=1 αn,ijs n i vn), yn = j−1(βnjxn + (1− βn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πcn+1x0, ∀n ≥ 1, (18) where j is the duality mapping on e, θn = supq∈f (kn − 1)φ(q, xn), for each i ≥ 0, {αn,i} and {βn} are sequences in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant of e. if∑∞ i=0 αn,i = 1 for all n ≥ 0, lim infn−→∞ βn(1− βn) > 0 and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. proof . we first show that cn+1 is closed and convex for each n ≥ 0. clearly c1 = c is closed and convex. suppose that cn is closed and convex for each n ∈ n. since for any z ∈ cn, we known that φ(z, un) ≤ φ(z, xn) + θn is equivalent to 2〈z, jxn − jun〉 ≤ ‖xn‖2 − ‖un‖2 + θn. hence, cn+1 is closed and convex. next, we show that f ⊂ cn for all n ≥ 0. since by the convexity of ‖ · ‖2, property of φ, lemma 2.9 and by uniformly quasi-φ-asymptotically nonexpansive of sn for each q ∈ f ⊂ cn, advances in systems science and applications (2011), vol. 11, no. 1-2 131 we have φ(q, un) = φ(q, t fmrm,n t fm−1 rm−1,n ...t f2r2,nt f1 r1,nyn) ≤ φ(q, yn) = φ(q, j−1(βnjxn + (1− βn)jzn) = ‖q‖2 − 2〈q, βnjxn + (1− βn)jzn〉+ ‖βnjxn + (1− βn)jzn‖2 ≤ ‖q‖2 − 2βn〈q, jxn〉 − 2(1− βn)〈q, jzn〉+ βn‖xn‖2 + (1− βn)‖zn‖2 = βnφ(q, xn) + (1− βn)φ(q, zn), (19) and φ(q, zn) = φ(q, j−1(αn,0jxn + ∑∞ i=1 αn,ijs n i vn)) = ‖q‖2 − 2〈q, αn,0jxn + ∑∞ i=1 αn,ijs n i vn〉+ ‖αn,0jxn + ∑∞ i=1 αn,ijs n i vn‖2 = ‖q‖2 − 2αn,0〈q, jxn〉 − 2 ∑∞ i=1 αn,i〈q, jsni vn〉+ ‖αn,0jxn + ∑∞ i=1 αn,ijs n i vn‖2 ≤ ‖q‖2 − 2αn,0〈q, jxn〉 − 2 ∑∞ i=1 αn,i〈q, jsni vn〉+ αn,0‖jxn‖2 + ∑∞ i=1 αn,i‖jsni vn‖2 −αn,0αn,jg‖jvn − jsnj vn‖ = ‖q‖2 − 2αn,0〈q, jxn〉+ αn,0‖jxn‖2 − 2 ∑∞ i=1 αn,i〈q, jsni vn〉 + ∑∞ i=1 αn,i‖jsni vn‖2 − αn,0αn,jg‖jvn − jsnj vn‖ = αn,0φ(q, xn) + ∑∞ i=1 αn,iφ(q, sni vn)− αn,0αn,jg‖jvn − jsnj vn‖ ≤ αn,0φ(q, xn) + ∑∞ i=1 αn,iknφ(q, vn)− αn,0αn,jg‖jvn − jsnj vn‖. (20) it follows from lemma 2.7, that φ(q, vn) = φ(q,πcj −1(jxn − λnaxn)) ≤ φ(q, j−1(jxn − λnaxn)) = v (q, jxn − λnaxn) ≤ v (q, (jxn − λnaxn) + λnaxn)− 2〈j−1(jxn − λnaxn)− q, λnaxn〉 = v (q, jxn)− 2λn〈j−1(jxn − λnaxn)− q, axn〉 = φ(q, xn)− 2λn〈xn − q,axn〉+ 2〈j−1(jxn − λnaxn)− xn,−λnaxn〉. (21) since q ∈ v i(a,c) and a is an α-inverse-strongly monotone mapping, we have −2λn〈xn − q, axn〉 = −2λn〈xn − q,axn −aq〉 − 2λn〈xn − q, aq〉 ≤ −2λn〈xn − q,axn −aq〉 ≤ −2αλn‖axn −aq‖2. (22) by lemma 2.2 and ‖axn‖ ≤ ‖axn −aq‖, ∀q ∈ v i(a,c), we also have 2〈j−1(jxn − λnaxn)− xn,−λnaxn〉 = 2〈j−1(jxn − λnaxn)− j−1(jxn),−λnaxn〉 ≤ 2‖j−1(jxn − λnaxn)− j−1(jxn)‖‖λnaxn‖ ≤ 4 c2 ‖jj−1(jxn − λnaxn)− jj−1(jxn)‖‖λnaxn‖ = 4 c2 ‖jxn − λnaxn − jxn‖‖λnaxn‖ = 4 c2 ‖λnaxn‖2 = 4 c2 λ2 n‖axn‖2 ≤ 4 c2 λ2 n‖axn −aq‖2. (23) 132 saewan:a new modified block iterative algorithm for a system of. . . . . . substituting (22) and (23) into (21), we have φ(q, vn) ≤ φ(q, xn)− 2αλn‖axn −aq‖2 + 4 c2 λ2 n‖axn −aq‖2 = φ(q, xn) + 2λn( 2 c2 λn − α)‖axn −aq‖2 ≤ φ(q, xn). (24) substituting (24) into (20), we also have φ(q, zn) ≤ αn,0φ(q, xn) + ∑∞ i=1 αn,iknφ(q, xn)− αn,0αn,jg‖jvn − jsnj vn‖ ≤ αn,0knφ(q, xn) + ∑∞ i=1 αn,iknφ(q, xn)− αn,0αn,jg‖jvn − jsnj vn‖ = knφ(q, xn)− αn,0αn,jg‖jvn − jsnj vn‖ ≤ φ(q, xn) + supq∈f (kn − 1)φ(q, xn)− αn,0αn,jg‖jvn − jsnj vn‖ = φ(q, xn) + θn − αn,0αn,jg‖jvn − jsnj vn‖ ≤ φ(q, xn) + θn. (25) and substituting (25) into (19), we obtain φ(q, un) ≤ φ(q, xn) + θn. (26) thus, this show that q ∈ cn+1 implies that f ⊂ cn+1 and hence, f ⊂ cn for all n ≥ 0. this implies that the sequence {xn} is well defined. from definition of cn+1 that xn = πcnx0 and xn+1 = πcn+1x0,∈ cn+1 ⊂ cn we have φ(xn, x0) ≤ φ(xn+1, x0), ∀n ≥ 0. (27) form lemma 2.6, it follows that φ(xn, x0) = φ(πcnx0, x0) ≤ φ(q, x0)− φ(q, xn) ≤ φ(q, x0), ∀q ∈ f. (28) by (27) and (28), then {φ(xn, x0)} are nondecreasing and bounded. so, we obtain that lim n→∞ φ(xn, x0) exists. in particular, by (5), the sequence {(‖xn‖− ‖x0‖)2} is bounded. this implies {xn} is also bounded. we denote m := sup n≥0 {‖xn‖} <∞. (29) moreover, by the definition of θn and (29), it follows that θn −→ 0 as n −→∞. (30) next, we show that {xn} is a cauchy sequence in c. since xm = πcmx0 ∈ cm ⊂ cn, for m > n, by lemma 2.6, we have φ(xm, xn) = φ(xm,πcnx0) ≤ φ(xm, x0)− φ(πcnx0, x0) = φ(xm, x0)− φ(xn, x0). advances in systems science and applications (2011), vol. 11, no. 1-2 133 since limn−→∞ φ(xn, x0) exists and we taking m,n→∞ then, we get φ(xm, xn)→ 0. from lemma 2.4, we have limn→∞ ‖xm − xn‖ = 0. thus {xn} is a cauchy sequence and by the completeness of e and there exist a point p ∈ c such that xn → p as n→∞. now, we claim that ‖jun−jxn‖ → 0, as n→∞. by definition of xn = πcnx0, we have φ(xn+1, xn) = φ(xn+1,πcnx0) ≤ φ(xn+1, x0)− φ(πcnx0, x0) = φ(xn+1, x0)− φ(xn, x0). since lim n→∞ φ(xn, x0) exists, we also have lim n→∞ φ(xn+1, xn) = 0. (31) again form lemma 2.4, that lim n→∞ ‖xn+1 − xn‖ = 0. (32) from j is uniformly norm-to-norm continuous on bounded subsets of e, we obtain lim n→∞ ‖jxn+1 − jxn‖ = 0. (33) since xn+1 = πcn+1x0 ∈ cn+1 ⊂ cn and the definition of cn+1, we have φ(xn+1, un) ≤ φ(xn+1, xn) + θn. by (30) and (31), that lim n→∞ φ(xn+1, un) = 0. (34) applying lemma 2.4, we have lim n→∞ ‖xn+1 − un‖ = 0. (35) since ‖un − xn‖ = ‖un − xn+1 + xn+1 − xn‖ ≤ ‖un − xn+1‖+ ‖xn+1 − xn‖ it follows from (32) and (35), that lim n→∞ ‖un − xn‖ = 0. (36) since j is uniformly norm-to-norm continuous on bounded subsets of e, we also have lim n→∞ ‖jun − jxn‖ = 0. (37) next, we will show that xn → p ∈ f := ∩mj=1ep (fj) ∩ (∩∞i=1f (si)) ∩ v i(a,c). (i) we show that xn → p ∈ ∩∞i=1f (si). since xn+1 = πcn+1x0 ∈ cn+1 ⊂ cn, it follow from (25), we have φ(xn+1, zn) ≤ φ(xn+1, xn) + θn, by (30) and (31), we get lim n→∞ φ(xn+1, zn) = 0 (38) 134 saewan:a new modified block iterative algorithm for a system of. . . . . . again form lemma 2.4, that lim n→∞ ‖xn+1 − zn‖ = 0. (39) since ‖zn − xn‖ ≤ ‖zn − xn+1‖+ ‖xn+1 − xn‖ from (32) and (39), we have lim n→∞ ‖zn − xn‖ = 0. (40) by using the triangle inequality, we obtain ‖xn+1 − zn‖ ≤ ‖xn+1 − xn‖+ ‖xn − zn‖. (41) by (32) and (40), we get lim n→∞ ‖xn+1 − zn‖ = 0. (42) since j is uniformly norm-to-norm continuous, we obtain lim n→∞ ‖jxn+1 − jzn‖ = 0. (43) from (67), we note that ‖jxn+1 − jzn‖ = ‖jxn+1 − (αn,0jxn + ∑∞ i=1 αn,ijs n i vn)‖ = ‖αn,0jxn+1 − αn,0jxn + ∑∞ i=1 αn,ijxn+1 − ∑∞ i=1 αn,ijs n i vn‖ = ‖αn,0(jxn+1 − jxn) + ∑∞ i=1 αn,i(jxn+1 − jsni vn)‖ = ‖ ∑∞ i=1 αn,i(jxn+1 − jsni vn)− αn,0(jxn − jxn+1)‖ ≥ ∑∞ i=1 αn,i‖jxn+1 − jsni vn‖ − αn,0‖jxn − jxn+1‖, and hence ‖jxn+1 − jsni vn‖ ≤ 1∑∞ i=1 αn,i (‖jxn+1 − jzn‖+ αn,0‖jxn − jxn+1‖). (44) from (33), (43) and lim inf n→∞ ∑∞ i=1 αn,i > 0, we get lim n→∞ ‖jxn+1 − jsni vn‖ = 0. (45) since j−1 is uniformly norm-to-norm continuous on bounded sets, we have lim n→∞ ‖xn+1 − sni vn‖ = 0. (46) using the triangle inequality, that ‖xn − sni vn‖ = ‖xn − xn+1 + xn+1 − sni vn‖ ≤ ‖xn − xn+1‖+ ‖xn+1 − sni vn‖. from (32) and (46), we have lim n→∞ ‖xn − sni vn‖ = 0. (47) advances in systems science and applications (2011), vol. 11, no. 1-2 135 on the other hand, we observe that φ(q, xn)− φ(q, un) + θn = ‖xn‖2 − ‖un‖2 − 2〈q, jxn − jun〉+ θn. it follows from θn −→ 0, ‖xn − un‖ −→ 0 and ‖jxn − jun‖ −→ 0, that φ(q, xn)− φ(q, un) + θn −→ 0 as n −→∞. (48) from (19), (20) and (24), we compute φ(q, un) ≤ φ(q, yn) ≤ βnφ(q, xn) + (1− βn)φ(q, zn) ≤ βnφ(q, xn) + (1− βn)[αn,0φ(q, xn) + ∑∞ i=1 αn,iknφ(q, vn) −αn,0αn,jg‖jvn − jsnj vn‖] = βnφ(q, xn) + (1− βn)αn,0φ(q, xn) + (1− βn) ∑∞ i=1 αn,iknφ(q, vn) −(1− βn)αn,0αn,jg‖jvn − jsnj vn‖ ≤ βnφ(q, xn) + (1− βn)αn,0φ(q, xn) + (1− βn) ∑∞ i=1 αn,iknφ(q, vn) ≤ βnφ(q, xn) + (1− βn)αn,0φ(q, xn) + (1− βn) ∑∞ i=1 αn,ikn[φ(q, xn)− 2λn(α− 2 c2 λn)‖axn −aq‖2] ≤ βnφ(q, xn) + (1− βn)αn,0knφ(q, xn) + (1− βn) ∑∞ i=1 αn,iknφ(q, xn) −(1− βn) ∑∞ i=1 αn,ikn2λn(α− 2 c2 λn)‖axn −aq‖2 = βnknφ(q, xn) + (1− βn)knφ(q, xn)− (1− βn) ∑∞ i=1 αn,ikn2λn(α− 2 c2 λn)‖axn −aq‖2 ≤ knφ(q, xn)− (1− βn) ∑∞ i=1 αn,ikn2λn(α− 2 c2 λn)‖axn −aq‖2] ≤ φ(q, xn) + supq∈f (kn − 1)φ(q, xn)− (1− βn) ∑∞ i=1 αn,ikn2λn(α− 2 c2 λn)‖axn −aq‖2 ≤ φ(q, xn) + θn − (1− βn) ∑∞ i=1 αn,ikn2λn(α− 2 c2 λn)‖axn −aq‖2 and hence 2a(α− 2b c2 )‖axn −aq‖2 ≤ 2λn(α− 2 c2 λn)‖axn −aq‖2 ≤ 1 (1−βn) ∑∞ i=1 αn,ikn (φ(q, xn)− φ(q, un) + θn). (49) from (48), {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, lim infn−→∞(1− βn) > 0 and lim infn−→∞ αn,0αn,i > 0, for i ≥ 0 and kn −→ 1 as n −→∞, we obtain that lim n→∞ ‖axn −aq‖ = 0. (50) from lemma 2.6, lemma 2.7 and (23), we compute φ(xn, vn) = φ(xn,πcj −1(jxn − λnaxn)) ≤ φ(xn, j −1(jxn − λnaxn)) = v (xn, jxn − λnaxn) ≤ v (xn, (jxn − λnaxn) + λnaxn)− 2〈j−1(jxn − λnaxn)− xn, λnaxn〉 = φ(xn, xn) + 2〈j−1(jxn − λnaxn)− xn,−λnaxn〉 = 2〈j−1(jxn − λnaxn)− xn,−λnaxn〉 ≤ 4λ2n c2 ‖axn −aq‖2 ≤ 4b2 c2 ‖axn −aq‖2. 136 saewan:a new modified block iterative algorithm for a system of. . . . . . applying lemma 2.4 and (50) that lim n→∞ ‖xn − vn‖ = 0 (51) and we also obtain lim n→∞ ‖jxn − jvn‖ = 0 (52) from sni is continuous, for any i ≥ 1 lim n→∞ ‖sni xn − sni vn‖ = 0. (53) again by the triangle inequality, we get ‖xn − sni xn‖ ≤ ‖xn − sni vn‖+ ‖sni vn − sni xn‖. from (47) and (53), we have lim n→∞ ‖xn − sni xn‖ = 0, ∀i ≥ 1. (54) since j is uniformly continuous on any bounded subset of e, we obtain limn→∞ ‖jxn − jsni xn‖ = 0, ∀i ≥ 1. (55) since xn −→ p and j is uniformly continuous, it yields jxn −→ jp. hence, from (55), we get jsni xn −→ jp, ∀i ≥ 1. (56) since j−1 : e∗ −→ e is norm-weake*-continuous, we have sni xn ⇀ p, for each i ≥ 1. (57) on the other hand, for each i ≥ 1, we have |‖sni xn‖ − ‖p‖| = |‖j(sni xn)‖ − ‖jp‖| ≤ ||j(sni xn)− jp||. by (56), we obtain ‖sni xn‖ −→ ‖p‖ for each i ≥ 1. since e is uniformly convex banach spaces then e has the kadec-klee property, we get sni xn −→ p for each i ≥ 1. by the assumption that ∀i ≥ 1, si is uniformly li-lipschitz continuous, hence we have. ‖sn+1 i xn − sni xn‖ ≤ ‖sn+1 i xn − sn+1 i xn+1‖+ ‖sn+1 i xn+1 − xn+1‖+ ‖xn+1 − xn‖+ ‖xn − sni xn‖ ≤ (li + 1)‖xn+1 − xn‖+ ‖sn+1 i xn+1 − xn+1‖+ ‖xn − sni xn‖. (58) by (32) and (54), it follows that ‖sn+1 i xn − sni xn‖ → 0. from sni xn −→ p, we have sn+1 i xn → p, that is sisni xn → p. in view of closeness of si, we have sip = p, for all i ≥ 1. this imply that p ∈ ∩∞i=1f (si). advances in systems science and applications (2011), vol. 11, no. 1-2 137 (ii) we show that xn → p ∈ ∩mj=1ep (fj). applying (19) and (25), we get φ(p, yn) ≤ φ(p, xn)+θn. from lemma 2.13 and un = ωm n yn, when ωj n = t qj rj,nt qj−1 rj−1,n ...t q2 r2,nt q1 r1,n , j = 1, 2, 3, ...,m, ω0 n = i , for p ∈ f , we observe that φ(p, un) = φ(p,ωm n yn) ≤ φ(p,ωm−1 n yn) ... ≤ φ(p,ωj nyn) ... ≤ φ(p, yn) ≤ φ(p, xn) + θn ∀j = 1, 2, 3, ...,m. (59) follows from lemma 2.13, that φ(un,ω j nyn) ≤ φ(p,ωj nyn)− φ(p, un) ≤ φ(p, xn)− φ(p, un) + θn = ‖p‖2 − 2〈p, jxn〉+ ‖xn‖2 − (‖p‖2 − 2〈p, jun〉+ ‖un‖2) + θn = ‖xn‖2 − ‖un‖2 − 2〈p, jxn − jun〉+ θn ≤ ‖xn − un‖(‖xn + un‖) + 2‖p‖‖jxn − jun‖+ θn. (60) from (36), (37), θn → 0 as n→∞ and lemma 2.4, we get lim n→∞ ‖un − ωj nyn‖ = 0∀j = 1, 2, 3, ...,m. (61) by using triangle inequality, we have ‖xn − ωj nyn‖ ≤ ‖xn − un‖+ ‖un − ωj nyn‖. from (36) and (61), we have lim n→∞ ‖xn − ωj nyn‖ = 0 ∀j = 1, 2, 3, ...,m. (62) again by using triangle inequality, we have ‖ωj nyn − ωj−1 n yn‖ ≤ ‖ωj nyn − xn‖+ ‖xn − ωj−1 n yn‖. from (62),we also have lim n→∞ ‖ωj nyn − ωj−1 n yn‖ = 0 ∀j = 1, 2, 3, ...,m. (63) since j is uniformly norm-to-norm continuous, we obtain lim n→∞ ‖jωj nyn − jωj−1 n yn‖ = 0 ∀j = 1, 2, 3, ...,m. from rj,n > 0 we have ‖jωj nyn−jωj−1 n yn‖ rj,n → 0 as n→∞ ∀j = 1, 2, 3, ...,m, and fj(ω j nyn, y) + 1 rj,n 〈y − ωj nyn, jωj nyn − jωj−1 n yn〉 ≥ 0, ∀y ∈ c. 138 saewan:a new modified block iterative algorithm for a system of. . . . . . by (a2), that ‖y − ωj nyn‖‖jωj nyn−jωj−1 n yn‖ rn ≥ 1 rj,n 〈y − ωj nyn, jωj nyn − jωj−1 n yn〉 ≥ −fj(ωj nyn, y) ≥ fj(y,ω j nyn), ∀y ∈ c, and ωj nyn → p we get f(y, p) ≤ 0 for all y ∈ c. for 0 < t < 1, define yt = ty + (1 − t)p. then yt ∈ c which imply that fj(yt, p) ≤ 0. from (a1), we obtain that 0 = fj(yt, yt) ≤ tfj(yt, y) + (1− t)fj(yt, p) ≤ tfj(yt, y). thus fj(yt, y) ≥ 0. from (a3), we have fj(p, y) ≥ 0 for all y ∈ c and j = 1, 2, 3, ...,m. hence p ∈ ep (fj) ∀j = 1, 2, 3, ...,m. this imply that p ∈ ∩mj=1ep (fj). (iii) we show that xn → p ∈ v i(a,c). indeed, define b ⊂ e × e∗ by bv = { av +nc(v), v ∈ c; ∅, v /∈ c. (64) by lemma 2.8, b is maximal monotone and b−10 = v i(a,c). let (v, w) ∈ g(b). since w ∈ bv = av +nc(v), we get w −av ∈ nc(v). from vn ∈ c, we have 〈v − vn, w −av〉 ≥ 0. (65) on the other hand, since vn = πcj −1(jxn − λnaxn). then by lemma 2.5, we have 〈v − vn, jvn − (jxn − λnaxn)〉 ≥ 0, and thus 〈v − vn, jxn−jvnλn −axn〉 ≤ 0. (66) it follows from (65), (66) and a is monotone and 1 α -lipschitz continuous, that 〈v − vn, w〉 ≥ 〈v − vn, av〉 ≥ 〈v − vn, av〉+ 〈v − vn, jxn−jvnλn −axn〉 = 〈v − vn, av −axn〉+ 〈v − zvn, jxn−jvnλn 〉 = 〈v − vn, av −avn〉+ 〈v − vn, avn −axn〉+ 〈v − vn, jxn−jvnλn 〉 ≥ −‖v − vn‖‖vn−xn‖α − ‖v − vn‖‖jxn−jvn‖a ≥ −h(‖vn−xn‖α + ‖jxn−jvn‖ a ), where h = supn≥1 ‖v − vn‖. take the limit as n → ∞, (51) and (52), we obtain 〈v − p, w〉 ≥ 0. by the maximality of b we have p ∈ b−10, that is p ∈ v i(a,c). finally, we show that p = πfx0. from xn = πcnx0, we have 〈jx0 − jxn, xn − z〉 ≥ 0, ∀z ∈ cn. since f ⊂ cn, we also have 〈jx0 − jxn, xn − y〉 ≥ 0, ∀y ∈ f. advances in systems science and applications (2011), vol. 11, no. 1-2 139 taking limit n→∞, we obtain 〈jx0 − jp, p− y〉 ≥ 0, ∀y ∈ f. by lemma 2.5, we can conclude that p = πfx0 and xn → p as n → ∞. this completes the proof. if si = s for each i ∈ n, then theorem 3.1 is reduced to the following corollary. corollary 3.2. let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach space e. for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)-(a4). let a be an α-inverse-strongly monotone mapping of c into e∗ satisfying ‖ay‖ ≤ ‖ay − au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let s : c → c be a closed l-lipschitz continuous and quasi-φ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 such that f := (f (s)) ∩ (∩mj=1ep (fj)) ∩ (v i(a,c)) is a nonempty and bounded subset in c. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows: vn = πcj −1(jxn − λnaxn), zn = j−1(αnjxn + (1− αn)jsnvn), yn = j−1(βnjxn + (1− βn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πcn+1x0, ∀n ≥ 1, (67) where j is the duality mapping one, θn = supq∈f (kn−1)φ(q, xn), {αn}, {βn} are sequences in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant of e. if lim infn−→∞(1 − βn) > 0 and lim infn−→∞ αn(1− αn) > 0, then {xn} converges strongly to p ∈ f , where p = πfx0. for a special case that i = 1, 2, we can obtain the following results on a pair of quasi-φasymptotically nonexpansive mappings immediately from theorem 3.1. corollary 3.3. let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach space e. for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)-(a4). let a be an α-inverse-strongly monotone mapping of c into e∗ satisfying ‖ay‖ ≤ ‖ay − au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let s, t : c → c be two closed quasi-φ-asymptotically nonexpansive mappings and ls , lt -lipschitz continuous, respectively with a sequence {kn} ⊂ [1,∞), kn → 1 such that f := f (s) ∩ f (t ) ∩ (∩mj=1ep (fj)) ∩ v i(a,c) is a nonempty and bounded subset in c. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows: vn = πcj −1(jxn − λnaxn), zn = j−1(αnjxn + βnjs nvn + γnjt nvn), yn = j−1(δnjxn + (1− δn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πcn+1x0, ∀n ≥ 0, (68) 140 saewan:a new modified block iterative algorithm for a system of. . . . . . where j is the duality mapping on e, θn = supq∈f (kn − 1)φ(q, xn), {αn}, {βn}, {γn} and {δn} are sequences in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant ofe. if αn+βn+γn = 1 for all n ≥ 0 and lim infn−→∞ αnβn > 0, lim infn−→∞ αnγn > 0, lim infn−→∞ βnγn > 0 and lim infn−→∞ δn(1− δn) > 0, then {xn} converges strongly to p ∈ f , where p = πfx0. corollary 3.4. let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach spacee. for each j = 1, 2, ...,m let fj be a bifunction fromc×c to r which satisfies conditions (a1)-(a4). let a be an α-inverse-strongly monotone mapping of c into e∗ satisfying ‖ay‖ ≤ ‖ay−au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let {si}∞i=1 : c → c be an infinite family of closed quasi-φnonexpansive mappings such that f := ∩∞i=1f (si)∩ (∩mj=1ep (fj)) ∩ v i(a,c) 6= ∅. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows: vn = πcj −1(jxn − λnaxn), zn = j−1(αn,0jxn + ∑∞ i=1 αn,ijsivn), yn = j−1(βnjxn + (1− βn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn), xn+1 = πcn+1x0, ∀n ≥ 0, (69) where j is the duality mapping on e, {αn,i} and {βn} is sequence in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2uniformly convexity constant of e. if ∑∞ i=0 αn,i = 1 for all n ≥ 0, lim infn−→∞(1− βn) > 0 and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. proof since {si}∞i=1 : c −→ c is an infinite family of closed quasi-φ-nonexpansive mappings, it is an infinite family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings with sequence kn = 1. hence the conditions appearing in theorem 3.1 f is a bounded subset in c and for each i ≥ 1, si is uniformly li-lipschitz continuous are of no use here. by virtue of the closeness of mapping si for each i ≥ 1, it yields that p ∈ f (si) for each i ≥ 1, that is, p ∈ ∩∞i=1f (si). therefore all conditions in theorem 3.1 are satisfied. the conclusion of corollary 3.4 is obtained from theorem 3.1 immediately. corollary 3.5. [44, theorem 3.2] let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach spacee. let f be a bifunction fromc×c to r satisfying (a1)-(a4). leta be an α-inverse-strongly monotone mapping of c into e∗ satisfying ‖ay‖ ≤ ‖ay − au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let {si}ni=1 : c → c be a finite family of closed quasi-φnonexpansive mappings such that f := ∩ni=1f (si)∩ep (f)∩v i(a,c) 6= ∅. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows:  zn = πcj −1(jxn − λnaxn), yn = j−1(α0jxn + ∑n i=1 αijsizn), f(un, y) + 1 rn 〈y − un, jun − jyn〉 ≥ 0, ∀y ∈ c, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn), xn+1 = πcn+1x0, ∀n ≥ 0, (70) advances in systems science and applications (2011), vol. 11, no. 1-2 141 where j is the duality mapping on e, {αn,i} is sequence in [0, 1], {rn} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant of e. if αi ∈ (0, 1) such that ∑n i=0 αi = 1, then {xn} converges strongly to p ∈ f , where p = πfx0. corollary 3.6. let c be a nonempty closed and convex subset of a uniformly convex and uniformly smooth banach space e. let f be a bifunction from c × c to r satisfying (a1)-(a4). let {si}∞i=1 : c → c be an infinite family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 and uniformly li-lipschitz continuous such that f := ∩∞i=1f (si) ∩ ep (f) is a nonempty and bounded subset in c. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows: yn = j−1(αn,0jxn + ∑∞ i=1 αn,ijs n i xn), f(un, y) + 1 rn 〈y − un, jun − jyn〉 ≥ 0, ∀y ∈ c, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πcn+1x0, ∀n ≥ 0, (71) where j is the duality mapping one, θn = supq∈f (kn−1)φ(q, xn), {αn,i} is sequence in [0, 1], {rn} ⊂ [a,∞) for some a > 0. if ∑∞ i=0 αn,i = 1 for all n ≥ 0 and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. 4. deduced to hilbert spaces if e = h , a hilbert space, then e is 2-uniformly convex (we can choose c = 1) and uniformly smooth real banach space and closed relatively quasi-nonexpansive map reduces to closed quasi-nonexpansive map. moreover, j = i , identity operator on h and πc = pc , projection mapping from h into c. thus, the following corollaries hold. theorem 4.1. let c be a nonempty closed and convex subset of a hilbert space h . for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)(a4), bj : c −→ e∗ be a continuous and monotone mapping and ϕj : c → r be a lower semicontinuous and convex function. let a be an α-inverse-strongly monotone mapping of c into h satisfying ‖ay‖ ≤ ‖ay −au‖, ∀y ∈ c and u ∈ v i(a,c) 6= ∅. let {si}∞i=1 : c → c be an infinite family of closed and uniformly quasi-φ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 and uniformly li-lipschitz continuous such that f := ∩∞i=1f (si)∩ (∩mj=1gmep (fj , bj , ϕj))∩ v i(a,c) is a nonempty and bounded subset in c. for an initial point x0 ∈ h with x1 = pc1x0 and c1 = c, we define the sequence {xn} as follows:  zn = pc(xn − λnaxn), yn = αn,0xn + ∑∞ i=1 αn,is n i zn, un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : ‖z − un‖ ≤ ‖z − xn‖+ θn}, xn+1 = pcn+1x0, ∀n ≥ 0, (72) where θn = supq∈f (kn − 1)‖q − xn‖, {αn,i} is sequence in [0, 1], {rj,n} ⊂ [a,∞) for some a > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < α/2. if ∑∞ i=0 αn,i = 1 for all n ≥ 0 142 saewan:a new modified block iterative algorithm for a system of. . . . . . and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. remark 4.2. theorem 4.1 improve and extend the corollary 3.7 in zegeye [44] in the aspect for the mappings, we extend the mappings from a finite family of closed relatively quasinonexpansive mappings to more general an infinite family of closed and uniformly quasi-φasymptotically nonexpansive mappings. 5. applications 5.1 zero points of an inverse-strongly monotone operator next, we consider the problem of finding a zero point of an inverse-strongly monotone operator of e into e∗. assume that a satisfies the conditions: (c1) a is α-inverse-strongly monotone, (c2) a−10 = {u ∈ e : au = 0} 6= ∅. theorem 5.1. let c be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach space e. for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)-(a4), bj : c −→ e∗ be a continuous and monotone mapping and ϕj : c → r be a lower semicontinuous and convex function. let a be an operator of e into e∗ satisfying (c1) and (c2). let {si}∞i=1 : c → c be an infinite family of closed uniformly li-lipschitz continuous and uniformly quasi-φ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 such that f := ∩∞i=1f (si) ∩ (∩mj=1gmep (fj , bj , ϕj)) ∩a−10 is a nonempty and bounded subset in c. for an initial point x0 ∈ e with x1 = πc1x0 and c1 = c, we define the sequence {xn} as follows: zn = j−1(αn,0jxn + ∑∞ i=1 αn,ijs n i vn), yn = j−1(βnjxn + (1− βn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, cn+1 = {z ∈ cn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πcn+1x0, ∀n ≥ 0, (73) where j is the duality mapping on e, θn = supq∈f (kn − 1)φ(q, xn), for each i ≥ 0, {αn,i} and {βn} are sequences in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant of e. if∑∞ i=0 αn,i = 1 for all n ≥ 0, lim infn−→∞(1− βn) > 0 and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. proof. setting c = e in corollary 3.4, we also get πe = i. we also have v i(a,c) = v i(a,e) = {x ∈ e : ax = 0} 6= ∅ and then the condition ‖ay‖ ≤ ‖ay − au‖ holds for all y ∈ e and u ∈ a−10. so, we obtain the result. advances in systems science and applications (2011), vol. 11, no. 1-2 143 5.2 complementarity problems let k be a nonempty, closed convex cone in e. we define the polar k∗ of k as follows: k∗ = {y∗ ∈ e∗ : 〈x, y∗〉 ≥ 0, ∀x ∈ k}. (74) if a : k −→ e∗ is an operator, then an element u ∈ k is called a solution of the complementarity problem ([37]) if au ∈ k∗, and 〈u,au〉 = 0. (75) the set of solutions of the complementarity problem is denoted by cp (a,k). theorem 5.2. let k be a nonempty closed and convex subset of a 2-uniformly convex and uniformly smooth banach space e. for each j = 1, 2, ...,m let fj be a bifunction from c × c to r which satisfies conditions (a1)-(a4), bj : c −→ e∗ be a continuous and monotone mapping and ϕj : c → r be a lower semicontinuous and convex function. let a be an αinverse-strongly monotone mapping of k into e∗ satisfying ‖ay‖ ≤ ‖ay − au‖, ∀y ∈ k and u ∈ cp (a,k) 6= ∅. let {si}∞i=1 : k → k be an infinite family of closed uniformly li-lipschitz continuous and uniformly quasi-φ-asymptotically nonexpansive mappings with a sequence {kn} ⊂ [1,∞), kn → 1 such that f := ∩∞i=1f (si) ∩ (∩mj=1gmep (fj , bj , ϕj)) ∩ cp (a,k) is a nonempty and bounded subset in k. for an initial point x0 ∈ e with x1 = πk1x0 and k1 = k, we define the sequence {xn} as follows: vn = πkj −1(jxn − λnaxn), zn = j−1(αn,0jxn + ∑∞ i=1 αn,ijs n i vn), yn = j−1(βnjxn + (1− βn)jzn), un = t fmrm,nt fm−1 rm−1,n ...t f2 r2,nt f1 r1,nyn, kn+1 = {z ∈ kn : φ(z, un) ≤ φ(z, xn) + θn}, xn+1 = πkn+1x0, ∀n ≥ 0, (76) where j is the duality mapping on e, θn = supq∈f (kn − 1)φ(q, xn), for each i ≥ 0, {αn,i} and {βn} are sequences in [0, 1], {rj,n} ⊂ [d,∞) for some d > 0 and {λn} ⊂ [a, b] for some a, b with 0 < a < b < c2α/2, where 1 c is the 2-uniformly convexity constant of e. if∑∞ i=0 αn,i = 1 for all n ≥ 0, lim infn−→∞(1− βn) > 0 and lim infn−→∞ αn,0αn,i > 0 for all i ≥ 1, then {xn} converges strongly to p ∈ f , where p = πfx0. proof . as in the proof of takahashi in [37, lemma 7.11], we get that v i(a,k) = cp (a,k). so, we obtain the result. references [1] y. i. alber, s. reich, an iterative method for solving a class of nonlinear operator equations in banach spaces, panamer. math. j. 4 (1994) 39–54. [2] y. i. alber, metric and generalized projection operators in banach spaces: properties and applications, in: a.g. kartsatos (ed.), theory and applications of nonlinear operators of accretive and monotone type, marcel dekker, new york, 1996, pp. 15–50. 144 saewan:a new modified block iterative algorithm for a system of. . . . . . [3] k. ball, e. a. carlen, e. h. lieb, sharp uniform convexity and smoothness inequalities for trace norm, invent. math. 26 (1994) 137–150. [4] b. beauzamy, introduction to banach spaces and their geometry, 2nd ed., noth holland, 1985. [5] e. blum, w. oettli, from optimization and variational inequalities to equilibrium problems, math. student 63 (1994) 123–145. [6] d. butnariu, s. reich, a. j. zaslavski, asymptotic behavior of relatively nonexpansive operators in banach spaces, j. appl. anal. 7 (2001) 151–174. [7] d. butnariu, s. reich, a.j. zaslavski, weak convergence of orbits of nonlinear operators in reflexive banach spaces, numer. funct. anal. optim. 24 (2003) 489–508. [8] y. censor, s. reich, iterations of paracontractions and firmly nonexpansive operators with applications to feasibility and optimization, optimization 37 (1996) 323–339. [9] s. s. chang, j. k. kim, x. r. wang, modified block iterative algorithm for solving convex feasibility problems in banach spaces, journal of inequalities and applications, vol. 2010, article id 869684, 14 pages. [10] p. cholamjiak, a hybrid iterative scheme for equilibrium problems, variational inequality problems and fixed point problems in banach spaces, fixed point theory and applications volume 2009 (2009), article id 312454, 19 pages. [11] w. cholamjiak and s. suantai, convergence analysis for a system of equilibrium problems and a countable family of relatively quasi-nonexpansive mappings in banach spaces, abstract and applied analysis, volume 2010 (2010), article id 141376, 17 pages. [12] i. cioranescu, geometry of banach spaces, duality mappings and nonlinear problems, kluwer, dordrecht, 1990. [13] p. l. combettes, s. a. hirstoaga, equilibrium programming in hilbert spaces, j. nonlinear convex anal. 6 (2005) 117–136. [14] h. iiduka, w takahashi,weak convergence of a projection algorithm for variational inequalities in a banach space, j. math. anal. appl. 339 (2008) 668–679. [15] c. jaiboon and p. kumam, a general iterative method for solving equilibrium problems, variational inequality problems and fixed point problems of an infinite family of nonexpansive mappings, j. appl. math. comput. (2010) volume 34, numbers 1-2, 407–439. [16] p. katchang and p. kumam, a new iterative algorithm of solution for equilibrium problems, variational inequalities and fixed point problems in a hilbert space, j. appl. math. comput. (2010) 32: 19–38. [17] s. kamimura, w. takahashi, strong convergence of a proximal-type algorithm in a banach space, siam j. optim. 13 (2002) 938–945. advances in systems science and applications (2011), vol. 11, no. 1-2 145 [18] f. kohsaka and w. takahashi, block iterative methods for a finite family of relatively nonexpansive mappings in banach spaces, fixed point theory and applications, vol. 2007, article id 21972, 18 pages, 2007. [19] f. kohsaka and w. takahashi, existence and approximation of fixed points of firmly nonexpansive-type mappings in banach spaced siam journal on optimization, vol. 19,no. 2, pp. 824-835, 2008. [20] m. kikkawa and w. takahashi, approximating fixed points of nonexpansive mappings by the block iterative method in banach spaces, international journal of computational and numerical analysis and applications, vol. 5, no. 1, pp. 5966, 2004. [21] p. kumam, a new hybrid iterative method for solution of equilibrium problems and fixed point problems for an inverse strongly monotone operator and a nonexpansive mapping, j. appl. math. comput. (2009) 29: 263–280. [22] s. matsushita, w. takahashi, weak and strong convergence theorems for relatively nonexpansive mappings in banach spaces, fixed point theory and appl. 2004 (2004) 37–47. [23] s. matsushita, w. takahashi, a strong convergence theorem for relatively nonexpansive mappings in a banach space, j. approx. theory 134 (2005) 257-266. [24] a. moudafi, second-order differential proximal methods for equilibrium problems, j. inequal. pure appl. math. 4 (2003) (art.18). [25] w. nilsrakoo, s. saejung, strong convergence to common fixed points of countable relatively quasi-nonexpansive mappings, fixed point theory and appl. volume 2008 (2008), article id 312454, 19 pages. [26] s. park, fixed points, intersection theorems, variational inequalities, and equilibrium theorems, international journal of mathematics and mathematical sciences volume 24 (2000), issue 2, pages 73-93. [27] s. plubtieng, and k. ungchittrakool, hybrid iterative methods for convex feasibility problems and fixed point problems of relatively nonexpansive mappings in banach spaces, fixed point theory and applications, volume 2008, article id 583082, 19 pages. [28] x. qin, s. y. cho and s. m. kang, on hybrid projection methods for asymptotically quasiφ-nonexpansive mappings, applied mathematics and computation, volume 215, 3874– 3883. [29] x. qin, s. y. cho and s. m. kang, strong convergence of shrinking projection methods for quasi-φ-nonexpansive mappings and equilibrium problems, j. comput. appl. math. volume 234, 750–760. [30] x. qin, y. j. cho and s. m. kang, convergence theorems of common elements for equilibrium problems and fixed point problems in banach spaces, j. comput. appl. math. 225 (2009) 20–30. 146 saewan:a new modified block iterative algorithm for a system of. . . . . . [31] s. reich, a weak convergence theorem for the alternating method with bregman distance, in: a.g. kartsatos (ed.), theory and applications of nonlinear operators of accretive and monotone type, marcel dekker, new york, (1996) 313–318. [32] r. t. rockafellar, on the maximality of sums of nonlinear monotone operators, trans. amer. math. soc 149(1970), 75–88. [33] s. saewan, p. kumam and k. wattanawitoon, convergence theorem based on a new hybrid projection method for finding a common solution of generalized equilibrium and variational inequality problems in banach spaces, abstract and applied analysis, volum 2010, article id 734126, 26 pages. [34] s. saewan and p. kumam, modified hybrid block iterative algorithm for convex feasibility problems and generalized equilibrium problems for uniformly quasi-φ-asymptotically nonexpansive mappings, abstract and applied analysis volume 2010, article id 357120, 22 pages. [35] s. saewan and p. kumam, a hybrid iterative scheme for a maximal monotone operator and two countable families of relatively quasi-nonexpansive mappings for generalized mixed equilibrium and variational inequality problems, abstract and applied analysis volume 2010, article id 123027, 31 pages. [36] y. su, d. wang, m. shang, strong convergence of monotone hybrid algorithm for hemirelatively nonexpansive mappings, fixed point theory and appl, volume 2008 (2008), article id 284613, 8 pages. [37] w. takahashi, nonlinear functional analysis, yokohama-publishers, 2000. [38] y. takahashi, k. hashimoto, m. kato, on sharp uniform convexity, smoothness, and strong type, cotype inequalities, j. nonlinear convex anal. 3(2002), 267–281. [39] w. takahashi, k. zembayashi, strong and weak convergence theorems for equilibrium problems and relatively nonexpansive mappings in banach spaces, nonlinear anal. 70 (2009) 45–57. [40] w. takahashi, k. zembayashi, strong convergence theorem by a new hybrid method for equilibrium problems and relatively nonexpansive mappings, fixed point theory and appl. volume 2008 (2008), article id 528476, 11 pages. [41] k. wattanawitoon, p. kumam, a strong convergence theorem by a new hybrid projection algorithm for fixed point problems and equilibrium problems of two relatively quasinonexpansive mappings, nonlinear analysis: hybrid systems, vol. 3, no. 1,(2009) 11-20 [42] h. k. xu, inequalities in banach spaces with applications, nonlinear anal. 16 (1991) 1127–1138. [43] c. zalinescu, on uniformly convex functions, j. math. anal. appl. 95(1983) 344–374. [44] h. zegeye, a hybrid iteration scheme for equilibrium problems, variational inequality problems and common fixed point problems in banach spaces, nonlinear anal, 72 (2010) 2136–2146. advances in systems science and applications (2011), vol. 11, no. 1-2 147 [45] h. zegeye, n. shahzad, strong convergence for monotone mappings and relatively weak nonexpansive mappings, nonlinear anal, 70 (2009) 2707–2716. [46] h. zegeye, e. u. ofoedu and n. shahzad, convergence theorems for equilibrium problem, variational inequality problem and countably infinite relatively quasi-nonexpansive mappings, applied mathematics and computation, volume 216, 3439–3449. [47] s. zhang, generalized mixed equilibrium problem in banach spaces, appl. math. mech. -engl. ed., 30 (2009) 1105–1112. 148 saewan:a new modified block iterative algorithm for a system of. . . . . . microsoft word 6-liu wei.doc issn 1078-6236 international institute for general systems studies, inc. structural analysis of leukemia related gene network liu wei1, zhang yuanyuan2, xu xiuzhu2 , yang meixi2, xu dashun 3 and wang shudong 2 1department of mathematics and information, ludong university, yantai, shandong 264025, china 2. college of information science and engineering, shandong university of science and technology, qingdao, shandong 266510, china 3. department of mathematics, southern illinois university carbondale, makanida 62958, usa email: wangshd2008@yahoo.com.cn ,liuyiwei1030@yahoo.com.cn abstract genome forms gene networks in the complicated interactive ways and further study on the cancer-related gene networks can help to understand and predict many unknown biological functions of cancer-related genome, and then obtain the useful informations about molecular mechanism of the formation of cancer. in this research, using the method of mutual information, we construct gene networks corresponding to normal control group and diseased experimental group of acute lymphoblastic leukemia (abbreviated as all). through contrasting the structure difference of two kinds of networks, we find 23 structural key genes of all with significant degree-difference, where 21 genes are confirmed to be closely related to the formation of all. the match ratio between the prediction by model and literature is up to 91.3%. according to the effectiveness of the method in this work, we can predict that the remaining two genes cxcl1 and tacstd2 are closely related to all. the significant differences of the network structures between control and experimental groups will enlighten the biomedical scientists that the reason of normal organism suffering from leukemia maybe is the great changes of some genes in vivo. the finding of structural key genes of all will help the biomedical scientist to further research the pathogenesis of all. keywords systems biology gene network mutual information leukemia 1. introduction leukemia is one of the top 10 human cancers. now many biomedical scientists are engaging in discovering the oncogenes and tumor-suppressor genes of leukemia, hope to know well the pathogenesis of leukemia, and find out the effective methods to treat leukemia. at present, besides doing clinical experiments directly, the main research ways to discover the pathogenesis of leukemia are based on the gene expression profiles data of dna chip. for instance, the method of cluster analysis was used to analyze the pattern of genes expressed in leukemic blasts from 360 pediatric all patients. distinct expression profiles identified each of the prognostically important leukemia subtypes, including t-all, e2a-pbx1, bcr-abl, *supported by national natural science foundation of china (grant nos. 60874036, 60503002) and sdust research fund. ∗ advances in systems science and applications (2011), vol.11, no.1-2 71-81 tel-aml1, mll rearrangement (yeoh et al., 2002)[1]. the method of pattern recognition was utilized to analyze the gene expression profiles data of patients with leukemia(yakovlev et al., 2002) [2]. the method of multivariate statistics was used to identify genes modulated by adhesion of human precursor b leukemia cells that regulate proliferation and apoptosis, highlighting new pathways that might provide insights into future therapy aiming at targeting apoptosis of leukemia cells (astier et al., 2003) [3]. the analysis of neural network internal structure allowed the identification of specific phenotype markers and the extraction of peculiar associations among genes and physiological states. at the same time, the neural network outputs provided assignment to multiple classes, such as different pathological conditions or tissue samples, for previously unseen instances (silvio et al., 2003)[4]. the method of graphical gaussian models (ggms) was used to describe gene association networks and to detect conditionally dependent genes (korbinian strimmer et al., 2005)[5]. a robust gene selection approach based on a hybrid between genetic algorithm and support vector machine was formalized and the major goal of this hybridization was to exploit fully their respective merits for identification of key feature genes (or molecular signatures) for a complex biological phenotype (shaoqi rao et al., 2005)[6]. the data-mining methods were applied to biomedical research. the results and how these results would affect the diagnosis and treatment of all in the future were discussed (morton, geoffrey, 2010)[7]. these above-mentioned methods can compensate for the deficiency of clinical trials, and they are of important significance to the prediction, diagnosis and treatment of cancer. systems biology is a newly interdiscipline after genomics, proteomics. now, forward modeling and reverse modeling are two main methods to study systems biology. modeling and analysis method of complex network is one of the important ways of reverse network modeling and has been widely applied to the study of complex biological system. the complex network models of biological system mainly include genetic networks, protein interaction networks, metabolic networks, signal networks and cellular networks, etc. harald lahm et al. researched the complex network of family of endogenous agglutinin in 2004[8], which would help to understand certain types of malignant phenotypes. adriano v. werhli et al. constructed raf signal transduction network using gaussian and bayesian model respectively based on the data of system expression profiles in 2006[9]. theodore j. perkins et al. analyzed the gene regulatory network of drosophila using the method of reverse network modeling, which can explain the activation of genes[10]. shudong wang et al. constructed the logical network of arabidopsis genes under different external stimuli using reverse network modeling of information entropy, simulated and analyzed the dynamical behaviours of the obtained logic network [11]. in this work, based on the gene expression profiles of control group and experimental groups of all, we construct the mutual information network of healthy marrow, b_all and t_all respectively. through contrasting the changes of network structures under different groups, we discover the significant difference in the networks for control and experimental groups. furthermore, we find 23 structural key genes of all. the paper is organised as follows. we introduce the related background in the first part, discuss the analysis method of complex network in the second one, construct the mutual information gene network and give the main results in the third part, and analyze the biological significances of the experimental results combining with the gene functions in the fourth part. 2. analysis method of complex network complex network can show the complex relationships among large numbers of elements more clearly, and better explain the mutual influence relationships between the structures and functions. therefore, further study on complex network of the genes of diseases can help to understand pathogenesis of diseases. biological systems are composed of interactions and mutual regulations of many elements, such as genes, proteins, protein complexes and 72 liu: structural analysis of leukemia related gene network transcription factors, etc. if these elements and the interactions and mutual regulations of elements can be simplified as nodes and edges respectively, then the complicated biological systems can be abstracted to complex networks, for example, undirected network, directed network, and weighted network etc. now, the analysis method of complex network is mainly studying the overall (average path length, clustering coefficient, assortativity coefficient, degree distribution, etc.) and local (community structure, motif etc.) properties of network. in this work, we will analyze the structure of complex life systems to obtain further understanding its function, which is of important reference values for life scientists to predict and cure disease. let  ,g v e be a complex network with node-set  1,2, ,v n  and edge-set e . the statistics of complex network used in the work are as follows: (1) average degree ( k ) the degree of a node is defined as the number of nodes adjacent to it. the average degree is the average of the degrees of all nodes in the network, denoted by k . (2) average path length ( l ) the shortest path length is defined as the distance between any two nodes in the network. the average path length is the average of all of the shortest path lengths in the network, denoted by l . (3) average clustering coefficient (c ) the clustering coefficient of node i in the network is defined as the number of triangles dividing by the number of triples connected with node i. it reflects the tightness of the network connection. the average clustering coefficient is the average of the clustering coefficients of all nodes in the network, denoted by c . (4) modularity (q ) modularity is defined as 2( )i ij i i i q q e a    where ije denotes the proportion of the edges connecting two different communities to all the edges of network, and ia is the sum of each element of thi row which denotes the proportion of the edges connecting with thi community to all the edges of network. modularity is the probability measure of community structure of network, and is mainly used to judge the quality of community structure. (5) average coreness ( k ) k -core of a graph is defined as the remaining subgraph after repeatedly removing those nodes whose degrees are less than or equal to k . the coreness of a node denotes the depth of the node in the core. the average coreness is the average of the corenesses of the graph, denoted by k . (6) average betweenness centrality (b ) the betweenness centrality of a node is defined as the number of the shortest paths through the node in a network. it is an important index which can reflect the topological properties of the network. the average betweenness centrality is the average of the betweenness centralities of all nodes in the network, denoted byb . (7) non-isolated node proportion ( i ) the isolated node of a network is defined as the node that is not adjacent to any other nodes. it usually doesn't work in the network. the non-isolated node proportion is the number of non-isolated node dividing the number of all nodes. it can reflect the intensive extent of the nodes associated in the network. advances in systems science and applications (2011), vol.11, no.1-2 73 3. construction and analysis of gene network of mutual information 3.1 establishment of working database all the data in this research are from e13159 sample library in gpl570 of ncbi. e13159 sample library is composed of more than 2000 samples. we extract 40 samples from normal samples, b_all samples and t_all samples respectively. for description convenience, we refer to normal samples as the control group, b_all samples and t_all samples as two experimental groups to establish three original databases which contain more than 20,000 genes sample data of human genome. it is impossible to establish a gene network of mutual information with more than 20,000 genes using the present computer from the view of computational complexity. therefore, we should process the original database firstly. according to 301 cancer-related genes we have mastered [12], we select 801 probes corresponding to these genes. if there are several probes corresponding to a gene, then we choose the probe that has the highest expression profiles corresponding to the gene. through the above processing, we get 286 genes as the research object. because we focus on gene network structure, the genes whose expression levels are almost 0 or 1 in all of the samples have no contribution to the differences of network structure. hence, we delete the genes whose the proportions of data expression profiles of samples being 1 are greater than 90% or less than 15% in each library. at last we get 58, 73 and 85 genes from the normal group, the b_all group and the t_all group respectively to form the corresponding working databases. next we will construct the gene networks of mutual information based on the 3 working databases. 3.2 selection of threshold 3.2.1 comparison of network statistics firstly, we should discretize p_value of the working database appropriately. specific methods are as follows: divide the range of p_value [0, 1] into 20 parts, and mark with 1, 2, …, 20, respectively. this kind of discretization has finer granularity than 0-1 discretization, and it lose less information. secondly, we can get a complete connected weighted network of all genes in each library by use of the mutual information formula, and the mutual information value is denoted by the weight. in order to highlight the specificity of network structure and obtain useful biological informations, the mutual information values of the complete connected network are need to be coarse-grained to seek threshold. as the distribution interval of mutual information value in each database is different, all the mutual information values of the 3 working databases are normalized, so that different databases can be comparable. finally, we analyze 7 statistics (average degree, average path length, average clustering coefficient, modularity, average coreness, average betweenness centrality, non-isolated node proportion) with the variation of thresholds between 0.1 and 0.9 in steps of 0.01, as shown in fig. 1. 74 liu: structural analysis of leukemia related gene network fig. 1. the change situation of 7 network statistics with the increasing of threshold, where the abscissa and vertical axis denotes the threshold and the corresponding statistics respectively. red, green, blue c urve denotes the change curve of normal group, b_all group and t_all group respectively. 3.2.2 construction of mutual information network we discover that when the threshold is taken as 0.53, 7 statistics of the networks for control and two experimental groups will be separated obviously by studying the statistics curves changing with the increasing of threshold in fig. 1. so we take 0.53 as the threshold to construct the mutual information gene networks of 3 working databases, as shown in fig. 2: (a) normal (b) b-all (c) t-all fig. 2. when the threshold is taken as 0.53, the mutual information gene networks of 68 coordinate positioning genes of 3 working databases, where (a),(b),(c) denotes normal, b_all and t_all network, respectively. obviously, the relationships of many nodes in the three gene networks change in quantity, such as most of isolated nodes in the normal mutual information gene network are no longer isolated in b-all or t-all network, especially the changes of t-all network are significantly. at the same time, many non-isolated nodes in the normal network become isolated in b-all or t-all network, the connectivity of nodes in the network also changes significantly. 3.3 selection of structural key genes a: some genes are isolated in the normal and t-all network, but high connectivity in b-all network. because there are 41 non-isolated nodes in the b_all network, take the genes whose degree changes in quantity are more than 20 as the significant genes called the structural key genes. at last, we discover that 8 genes ar, erbb4, esr2, evi-1, pdgfrb, fat1, cxcl1 and cxcl2 satisfy the condition. their degrees in 3 gene networks are showed as table 1. table 1 the gene list of type a advances in systems science and applications (2011), vol.11, no.1-2 75 b: some genes are isolated nodes in normal network and b-all network, but high connectivity in the t-all network. because there are 38 non-isolated nodes in the b_all network, take the genes whose degree changes in quantity are more than 19 as the structural key genes. at last, we find 8 genes fer, fes, lta, wnt3, tcl1a, btk, msh4 and ptk2 genes satisfy the condition, their degrees in 3 gene networks are showed as table 2. table 2 the gene list of type b gene normal b_all t_all fer* 0 0 34 fes* 0 0 27 lta* 0 0 32 wnt3* 0 0 26 tcl1a* 0 0 31 msh4* 0 0 25 ptk2* 0 0 24 btk* 0 0 23 c: the degrees of some genes have changed more significantly in the networks of experimental groups than in the one of normal control group. take the genes whose degree changes in quantity are more than 18 as the structural key genes. at last, we find 7 genes cdh1, kdr, igf-1, fosl1, cttn, mos and tacstd2 satisfy the condition, their degrees in 3 gene networks are showed as table 3. table 3 the gene list of type c gene normal b_all t_all cdh1* 0 31 27 tacstd2 0 30 27 fosl1* 0 21 21 cttn* 0 20 20 kdr* 20 0 0 igf1* 20 0 0 mos* 19 0 0 4. conclusion and analysis we discover some genes whose degrees change significantly between the networks for control group and experimental groups. for example, some genes are very active (with greater degree) in the normal control group, and abnormal silence (isolated nodes) in the experimental groups, while some others are very active in the experimental groups, and abnormal silence in the gene normal b_all t_all ar* 0 26 0 erbb4* 0 27 0 esr2* 0 26 0 evi1* 0 33 0 fat1* 0 28 0 pdgfrb* 0 35 0 cxcl1 0 21 0 cxcl2* 0 25 0 76 liu: structural analysis of leukemia related gene network normal control group. that is to say, these genes have great contribute to structure changes of mutual information gene networks for control and experimental groups, so we define those genes whose degrees change significantly between control group and experimental groups as “the structural key gene”. now, we discuss these genes in three conditions as follows: a: considering genes ar, erbb4, esr2, evi-1, pdgfrb, fat1, cxcl1 and cxcl2, they are isolated nodes in normal and t-all gene network, but with greater degree in the b-all network. because their degrees in the b-all network significantly increase, we can speculate that their functions or the life processes involved in should be to promote cell proliferation and division, and they should promote the development and deterioration of cancer forward. the specific functions of the 8 genes as follows: yun cai et al. investigated the methylation status of the androgen receptor gene (ar) in leukemia cell lines. results showed the presence of both methylated and unmethylated cpg islands of the ar promotor in leukemia cell lines. in the normal blood samples, only unmethylated bands were observed[13]. the over-expression of erbb4 gene (leukemia virus oncogene, homologous chromosome 4) is related with the occurrence and development of tumor. in leukemias, the downstream pathways of erbb family signaling are frequently activated, and introduction of activated egfr into hematopoietic cell lines induced proliferation, survival and abrogation of cytokine dependency of the cells. however, the activation mechanisms of the erbb family signaling are largely unknown[14]. otabek imamov et al. demonstrated the novel role for esr2 in regulating the differentiation of pluripotent hematopoietic progenitor cells, and suggested that the esr2(er beta)-mouse should be a potential model for myeloid and lymphoid leukemia, and that esr2 agonists might have clinical value in the treatment of leukemia if the esr2 is not itself mutated in this disease. the human esr2 gene has been mapped to chromosome 14q22[15]. evi-1 gene is frequently overexpressed in leukemias having 3q26 abnormalities such as t(3;3)(q21;q26) and inv(3)(q21 q26), and subjects to structural alteration in t(3;21)(q26;q22). ogawa s et al. presented another case of structural alteration of evi-1 gene in a case of inv(3)(q21 q26), in which evi-1 is truncated and a shorter form of evi-1 protein is expressed upon rearrangement of the gene. their result also supports an idea that evi-1 is a relevant oncogene whose overexpression or structural changes might play a crucial role in development of human leukemias[16]. pdgfrb is constitutively activated by gene fusion with different partners in myeloproliferative disorders with peculiar clinical characteristics. translocation t(5;12)(q33;p13), resulting in an etv6/pdgfrb gene fusion, is a recurrent chromosomal abnormality associated with chronic myelomonocytic leukemia (cmml). an analogous translocation was also found in four cell lines with features of pre-b acute lymphoblastic leukemia (all) [17]. thomas dunwell et al. demonstrated fat1 gene is significantly more methylated in b-all compared to t-all[18]. gill d. et al. demonstrated the prolonged survival of b-cll in vitro without additional stromal cell support. furthermore, novel cytokines ccl2 and cxcl2 appear to prolong the survival of b-cll cells in vitro. this culture system and these chemokines may allow us to gain insight into factors modulating b-cll survival and potentially could lead to targeted therapy as well as serve as an appropriate model to test new therapies[19]. cxcl1 is expressed by macrophages, neutrophils and epithelial cells, and has neutrophil chemoattractant activity. cxcl1 plays a role in spinal cord development by inhibiting the migration of oligodendrocyte precursors and is involved in the processes of angiogenesis, inflammation, wound healing, and tumorigenesis. based on the specific functions of the above genes, we discover that they are closely related to b_all, maybe they are the oncogenes of b_all. b: considering genes fer, fes, lta, wnt3, tcl1a, btk, msh4 and ptk2, they are isolated nodes in normal and b-all network, but with greater degree in the t-all network. because their degrees in the t-all network significantly increase, we can speculate that their functions or the life process involved in should be to promote cell proliferation and division, and advances in systems science and applications (2011), vol.11, no.1-2 77 they should promote the development and deterioration of cancer forward. the specific functions of the 8 genes as follows: expression of fer gene in a wide range of cell types indicates a general role in intracellular signalling or differentiation processes. j. groffen et al. found that the negative regulation to be the main cause for dysfunctioning of the fer promoter in a t-cell leukemia cell line[20]. the human fes proto-oncogene is expressed as a transcript of about 3.0 kb in both normal and leukemic myeloid cells. tesch h. et al. detected truncated fes transcripts of about 0.9 kb in a panel of human lymphoma and lymphoid leukemia cell lines, but not in normal untransformed hematopoietic cells[21]. zhou mx, et al. discovered that the partially purified lta (tnf-alpha) obtained from the eu-1 cell line also suppressed the proliferation of tnf-sensitive primary leukemic cells, and this inhibitory activity was abolished by an anti-tnf-alpha specific antibody. the results demonstrated that tnf-alpha is an inhibitor of in vitro proliferation of bcp-all cells from most patients[22]. transfection with a wnt3 plasmid resulted in a small increase in reporter gene activity, which was augmented by the transfection with the lrp6 coreceptor. recent microarray analyses have demonstrated that the wnt3 gene is overexpressed in cll, compared with normal b and t cells[23]. in t-cell prolymphocytic leukemia (t-pll), chromosomal imbalances affecting the long arm of chromosome 22 are regarded as typical chromosomal aberrations secondary to a tcrad-tcl1a fusion due to inv(14) or t(14;14) [24]. bruton's tyrosine kinase (btk) deficiency results in a differentiation block at the pre-b cell stage. likewise, acute lymphoblastic leukemia cells are typically arrested at early stages of b cell development. janet d. rowley et al. identified kinase-deficient splice variants of btk throughout all leukemia subtypes[25]. msh4 gene belongs to the human dna mismatch repair system (mmr).the principal function of mmr is to edit replication and to reverse dna polymerase errors. inactivation of mmr greatly increases spontaneous mutation rates. h hirai et al. suggested that disruption of mmr may play an important role in the development of human lymphoid leukemias[26]. focal adhesion kinase (fak) is constitutively activated and tyrosine phosphorylated in bcr/abl-transformed hematopoietic cells. yi le et al. suggested that fak is critical for leukemogenesis and might be a potential target for leukemia therapy[27]. based on the specific functions of the above genes, we discover that they are closely related to t_all, maybe they are the oncogenes of t_all. c: considering genes cdh1, kdr, igf-1, fosl1, cttn, mos and tacstd2, their degrees change more significantly in the networks for experimental groups than those for control group. the specific functions of the 7 genes as follows: e-cadherin, the gene product of cdh1, plays a key role in cell–cell adhesion. decrease or loss of e-cadherin expression accompanied by cdh1 promoter methylation has been reported in many human cancers. john r. et al. found that all normal donor samples expressed e-cadherin mrna, whereas both samples of acute myelogenous leukemia and chronic lymphocytic leukemia had a significant reduction or absence of expression. however, normal blast counterparts expressed only a low level of e-cadherin surface protein[28]. kdr (flt) gene plays an important role in new angiogenesis,and it is related with tumor grade, growth and prognosis. kai neben et al. showed that unique gene expression patterns can be correlated with flt3-itd and flt3-tkd. this might lead to the identification of further pathogenetic relevant candidate genes particularly in aml with normal karyotype[29]. igf-1 is important in blood formation and regulation and has been shown to stimulate the growth of both myeloid and lymphoid cells in culture. since infants who develop leukemia are likely to have had at least one transforming event occur in utero, julie a. ross et al. hypothesized that high levels of igf-1 may both produce a larger baby and contribute to leukemogenesis[30]. fosl1(fra-1) gene, which is a tax1-inducible fos-related gene, was isolated and tax1 or serum-responsive cis elements were analyzed to obtain further insight into the mechanism of tax1 action. tax1 of human t-cell leukemia virus type 1 stimulates the expression of several cellular immediate-early genes[31]. patients with lymph node metastasis apt to emerge the amplification of cttn (ems1) gene, 78 liu: structural analysis of leukemia related gene network accompanied the phenomenon of ems1 over-expression, tumor cell invasion and metastasis increases simultaneously. the mutation of ems1 maybe one of the reasons of all[32]. cytogenetic analysis of an infant with down syndrome with concomitant acute myelogenous leukemia revealed a unique t(8;16)(q22;q24). in situ chromosomal hybridization was used to demonstrate that the protooncogene mos was translocated from chromosome 8 to chromosome 16. mark j. et al. reported that the transposition of mos in association with acute leukemia[33]. tacstd2 is considered as a cancer-associated antigen, and the antibody of its extracellular domain can reduce the invasiveness of tumor cell. based on the specific functions of the above genes, we discover that they are closely related to all, maybe they are the oncogenes of all. therefore, comparing the 23 structural key genes of all predicted by model with the functions of these genes showed in literatures, we discover that 21 genes (marked with “*” in table 1, 2 and 3) are confirmed to be closely related to the formation of all. the match ratio between the prediction by model and literature is up to 91.3%. hence, we can predict that the remaining two genes cxcl1 and tacstd2 are also closely related to all based on the effectiveness of the method in this method. the authenticity and reliability of the above speculations, as well as their pathogenic mechanism of all are needed to be further validated by biological experiments. in this research, we use 7 statistics parameters to characterize the network structure from different angles. from the results of the numerical experiments, we can see that the 7 parameters can not measure the network structure fully. whether or not there is a proper structural parameter, which can give the network structure a more comprehensive characterization? in addition, the number of samples in the databases has a certain degree impact on the result, so how to choose a reasonable number of samples need us to further explore as well. references [1] eng-juh yeoh et al.. classification, subtype discovery, and prediction of outcome in pediatric acute lymphoblastic leukemia by gene expression profiling. cancer cell, 2002, 1(2): 133-143. [2] yakovlev et al.. variable selection and pattern recognition with gene expression data generated by the microarray technology. mathematical biosciences, 2002, 176(1): 71-98. [3] anne laurence astier et al.. temporal gene expression profile of human precursor b leukemia cells induced by adhesion receptor: identification of pathways regulating b-cell survival. blood, 2003, 101(3): 1118-1127. [4] silvio bicciato et al.. pattern identification and classification in gene expression data using an autoassociative neural network model. biotechnology and bioengineering, 2003, 81(5): 594-606. [5] juliane schäfer and korbinian strimmer. an empirical bayes approach to inferring large-scale gene association networks. bioinformatics, 2005, 21(6): 754-764. [6] shaoqi rao et al.. a robust hybrid between genetic algorithm and support vector machine for extracting an optimal feature gene subset. genomics, 2005, 85(1): 16-23. [7] morton, geoffrey. data mining the genetics of leukemia. qspace at queen's research & learning repository, 2010. [8] harald lahm et al.. tumor galectinology: insights into the complex network of a family of endogenous lectins. glycoconjugate journal, 2004, 20(4): 227-238. [9] adriano v. werhli, marco grzegorczyk, dirk husmeier. comparative evaluation of reverse engineering gene regulatory networks with relevance networks, graphical gaussian modelsand bayesian networks. bioinformatics, 2006, 22(20): 2523-2531. [10] theodore j. perkins, johannes jaeger, john reinitz, leon glass. reverse engineering the gap gene network of drosophila melanogaster. plos comput biol., 2006, 2(5): 0417-0428. advances in systems science and applications (2011), vol.11, no.1-2 79 [11] shudong wang,yan chen, qingyun wang, eryan li, yansen su, dazhi meng. analysis for gene networks based on logical relationships. journal of systems science and complexity, 2010, 23(5): 999-1011. [12] lǚsheng si, xu li. oncogene, tumor suppressor gene, tumor-related genes. shanxi science and technology press, xi’an, 2002.12. [13] yun cai et al.. androgen receptor cpg island methylation status in human leukemia cancer cells. leukemia, 2009, 27(2): 156-162. [14] jong woo lee et al.. kinase domain mutation of erbb family genes is uncommon in acute leukemias. leukemia research, 2006, 30(2): 241-242. [15] otabek imamov et al.. estrogen receptor beta in health and disease. biology of reproduction, 2005, 73(5): 866-871. [16] ogawa s et al.. abnormal expression of evi-1 gene in human leukemias. hum cell, 1996, 9(4): 323-332. [17] j cools et al.. chronic myeloproliferative disorders: a tyrosine kinase tale. leukemia, 2006, 20: 200-205. [18] thomas dunwell et al.. a genome-wide screen identifies frequently methylated genes in haematological and epithelial cancers. molecular cancer, 2010, 9(44): 1-12. [19] gill d., burgess m., knop l. et al.. identification of two novel chemokines (ccl2 and cxcl2) in b-chronic lymphocytic leukemia (b-cll) and prolonged survival of primary b-cll in vitro. blood, 2008, 112(11): abstr. 3157. [20] carmel m, shpungin s, nir u.. role of positive and negative regulation in modulation of the fer promoter activity. gene, 2000, 241(1): 87-99. [21] tesch h. et al.. expression of truncated transcripts of the proto-oncogene c-fps/fes in human lymphoma and lymphoid leukemia cell lines. oncogene, 1992, 7(5): 943-952. [22] zhou mx, et al.. effect of tumor necrosis factor-alpha on the proliferation of leukemic cells from children with b-cell precursor-acute lymphoblastic leukemia (bcp-all): studies of primary leukemic cells and bcp-all cell lines. blood, 1991, 77(9): 2002-2007. [23] dennis a. carson et al.. activation of the wnt signaling pathway in chronic lymphocytic leukemia. pnas, 2004, 101(9): 3118-3123. [24] stefanie bug et al.. recurrent loss, but lack of mutations, of the smarcb1 tumor suppressor gene in t-cell prolymphocytic leukemia with tcl1a–tcrad juxtaposition. cancer genetics and cytogenetics, 2009, 192(1): 44-47. [25] janet d. rowley et al.. deficiency of bruton's tyrosine kinase in b cell precursor leukemia cells. pnas, 2005, 102(37): 13266-13271. [26] h hirai et al.. mutations and loss of expression of a mismatch repair gene, hmlh1, in leukemia and lymphoma cell lines. blood, 1997, 89(5): 1740-1747. [27] yi le et al.. fak silencing inhibits leukemogenesis in bcr/abl-transformed hematopoietic cells. american journal of hematology, 2009, 84(5): 273–278. [28] john r. et al.. hypermethylation of e-cadherin in leukemia. blood, 2000, 95(10): 3208-3213. [29] kai neben et al.. distinct gene expression patterns associated with flt3and nras-activating mutations in acute myeloid leukemia with normal karyotype. oncogene, 2005, 24: 1580-1588. [30] julie a. ross et al.. big babies and infant leukemia: a role for insulin-like growth factor-1. cancer causes and control, 1996, 7(5): 553-559. [31] h tsuchiya et al.. human t-cell leukemia virus type 1 tax activates transcription of the human fra-1 gene through multiple cis elements responsive to transmembrane signals. j virol., 1993, 67(12): 7001-7007. [32] donald pasqualea et al.. clonal evolution with +11q 13, t(1;7) and t(1;4) at relapse in a patient with ph positive acute lymphocytic leukemia (all) treated with single agent front line imatinib followed by dasatinib. hematology, 2007, 12(6): 505 509. 80 liu: structural analysis of leukemia related gene network [33] mark j. et al.. translocation of the mos gene in a rare t(8;16) associated with acute myeloblastic leukemia and down syndrome. cancer genetics and cytogenetics, 1989, 37(2): 221-227. advances in systems science and applications (2011), vol.11, no.1-2 81 advances in systems science and applications (2013) vol.13 no.3 276-298 using simulations of netlogo as a tool for introducing greek high-school students to eco-systemic thinking aristotelis gkiolmas1, kostas karamanos1, anthimos chalkidis1, constantine skordoulis1, maria papaconstantinou2 and dimitrios stavrou3 1department of primary education, university of athens, greece. 2department of informatics, ionian university, corfu, greece. 3department of primary education, university of crete, greece. abstract in this paper, the effectiveness of the netlogo programming environment is investigated, as regards assisting students (of various levels of achievement) of the greek higher secondary education, to understand how some simple ecosystems are structured and to model the systemic behaviour of such ecosystems by conceptualising their complexity features. this paper is part of a wider research on teaching ecosystem complexity to high-school students, with the use of information and communication technologies (ict’s). specific models from the netlogo models’ library were used and students of the 2nd class of the greek lyceum (aged between 16 and 17) participated in the investigation. apart from oral instruction, the students were asked to run the netlogo simulations and do specific things with the models, answering simultaneously questions on worksheets provided to them. the studying and the evaluation of the worksheets by the researchers, as well as the postinstructional evaluation of the students, both oral (by means of cassettes) and written, through the use of an evaluation sheet, gave research findings that proved to be encouraging in that: a) students developed a greater understanding of the complex/systemic behaviour of ecosystems and b) they were capable, to a certain extent, of analyzing the systemic relations within simple ecosystems and built analogous relations in other, also simple, ecosystems. keywords eco-systemic thinking, education, netlogo, classroom, complexity. 1 introduction the aim of the current research was to investigate whether high-school students’ ability to achieve systemic thinking, and in particular eco-systemic thinking, could be slightly enforced and enhanced with the use of a specific software. systems-thinking is defined in many ways, and one way of defining it, is as: “the ability to understand and interpret complex systems” [1]. ecosystems are considered complex systems with respect to three essential properties of them [2]: (i) sustained diversity and individuality of components, (ii) localized interactions among these components and (iii) an autonomous process that selects from these components based on the results of local interactions a subset for repli277 advances in systems science and applications (2013) vol.13 no.3 cation or enhancement. taking these properties into consideration, it becomes apparent that eco-systems’ thinking is a perfectly valid expression of systemic thinking. many attempts, on the other hand, have been made so far in the literature to define what “modeling” of systems and especially of complex systems [4]. even in the case that the famous alternative definition of the (so called) complex “adaptive” systems (cas’s) (i.e. complex systems that “learn” as time evolves and modify their comportment) given by holland (1995)[5] is used, attributing to such systems four basic properties: aggregation, nonlinearity, diversity and flows, it becomes again obvious that ecosystems are indeed complex (adaptive) systems, which “learn” as time evolves, and thus, eco-systemic thinking is a form of systemic thinking and eco-systemic modeling an aspect of systems’ modeling. ecosystems are also, in a variety of other aspects “complex systems” [6] and can, therefore, be treated through their complexity characteristics, which also reveal systemic behavior, since the dynamics within complex systems have all the characteristics of system dynamics [7]. in this respect, the current research examines whether students, of greek upper secondary education, can understand the systemic features of a simple modelecosystem, consisting of only one prey and one predator. the definition of what “learning about the systemic nature of ecosystems” actually means, is given as: “the main goal of having students learn about systems is not to have them talk about systems in abstract terms, but to enhance their ability (and inclination) to attend to various aspects of particular systems, in attempting to understand or deal with the whole” [8], and is employed in this research. in his classic paper comparing the way of thinking of “novices” and an “experts” in complex systems, jacobson provides a table, namely table 1 in which he juxtaposes the mental models and beliefs of complex systems experts and novices [3]. some of these mental models and beliefs are directly pertinent to system-thinking. these mental models and beliefs, taken from table 1 of jacobson, are presented in the table below, together with the system-thinking ability that was treated as related to each one: (see tab.1 ) the current research emphasizes on the fourth column of the above table and attributes a score for each one of the system-thinking abilities achieved by the students, as it progresses. this is a measure of the students’ improvement in systems’ thinking as will be shown in the “results” section of the paper. a separate model of treating knowledge and learning in complex systems, very popular in the field of science education, is the model of “sbf structure, behavior and function”, developed mainly by cindy hmelo-silver and her associates [9-11]. in the sbf treatment and evaluation of knowledge about complex systems, the learning subject is considered to have adequate knowledge about a (usually complex) system if he/she can describe in exact terms the system’s: aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 278 table 1 juxtaposition of the complex systems novices’ and experts’ mental models and beliefs, those related to systems’ thinking, as taken by jacobson [3] category of component belief component beliefs associated with clockwork mental model component beliefs associated with complex systems mental model associated system-thinking ability 1. understanding phenomena reductive (e.g., step-wise sequences, isolated parts) nonreductive: whole-is-greaterthan-the-parts seeing only the stocks vs. seeing all the stocks and the flows and the arrows and the faucet controls 4. action effects small actions → small effects small action → big effect seeing a proportional increase/decrease in the stock and the flow vs. seeing a nonproportional increase/decrease in the stock and the flow 6. complex actions from complex rules from simple rules seeing a stock as affecting only very close stocks by the neighbouring flows vs. seeing a stock as affecting even very distant stocks by arrays of flows 7. final causes or purposefulness of natural phenomena teleological nonteleological or stochastic not observing the arrows, i.e. the “closed” feedback loops, seeing only “open” procedures vs. observing the arrows, i.e. the “closed” feedback loops, and not seeing only “open” procedures 8. ontology static structures events equilibration processes seeing the stocks as reaching a final, constant size and the flows as stopping vs. seeing the stocks as reaching a final constant size but the flows as never stopping table 2 the s-f elements of the modeled wolves-sheep-grass ecosystem and their respective system dynamics’ depictions properties and aspects of the modeled ecosystem as system, related to the sbf conceptual model. elements of this property and aspect system dynamics expression of this element structure s1) the wolves s2) the sheep s3) the grass s1) a box (“stock”) called “wolves” or “wolves’ population” or equivalent s2) a box (“stock”) called “sheep” or “sheep’s population” or equivalent s3) a box (“stock”) called “grass” or “amount of grass” or equivalent function f1) wolves eat sheep f2) sheep eat or do not eat grass f3) wolves die if they do not find something to eat for some time f4) sheep die if they do not find something to eat for some time (in the case: “grass plays a role”) f5) sheep die due to predation by wolves f6) grass is reduced when eaten by sheep (in the case: “grass plays a role”) f7) new wolves are added due to wolves’ births f8) new sheep are added, due to sheep’s births f9) grass is replaced with a certain rate (in the case: “grass plays a role”) f1) an arrow showing the predation or something equivalent f2) an arrow showing the eating of grass (loss of grass) or something equivalent f3) something depicting the deaths of wolves f4) something depicting the deaths of sheep f5) something depicting the deaths of sheep due to predation f6) an arrow depicting the loss of grass f7) something depicting the addition of wolves f8) something depicting the addition of sheep f9) something describing the replacement of grass 279 advances in systems science and applications (2013) vol.13 no.3 (i) structure (i.e. the parts that constitute the system), (ii) behavior (i.e. the algorithms that underlie the system and make it perform the way it performs and (iii) function (i.e. what actions the system really carries out, in the observable level). for the scope and the needs of the current research, the aspect of “behavior” of the system and the relative characteristics of systems dynamics were considered highly difficult to achieve for high-school students, so emphasis was given only to the aspects of “structure” (“s”) and “function” (“f”), as well as their corresponding system dynamics’ expressions. in the table 2 [9], the s and f characteristics of the modeled simple ecosystem of wolf predating sheep (and, optionally, sheep eating grass) are presented, together with the corresponding system dynamic representation(s) for each characteristic, that the student is expected to find and express. table 2 will also be used in the analysis of the results section of the research. 2 method 2.1 the objectives a number of methods and tools have been developed, in order to reveal students’ and teachers’ conceptions of simple dynamic systems and ecosystems [12]. a variety among these methods of revealing and improving the understanding of “systems”, have used agent-based systems and especially netlogo as a tool [1316], and furthermore, some of them used the environment of netlogo to see how a student changes from “expert” to “novice” in complex systems [10, 17]. there is clear evidence that even laypeople that have completed secondary education cannot easily define the structure or understand even the simplest simulations of ecosystem dynamics [18]. on the other hand, netlogo educational researchers have demonstrated that this modeling environment helps students, to a certain extent, to create mathematical models of a system, in the form of equations, or create dynamical/graphical models of a system [19-21]. therefore, the strategy of the current research, which makes it slightly different than all the aforementioned researches, focused on six didactical objectives, which stem mainly from authentic system-thinking abilities related to finding relations in the system and representing it, whereas the other researches seek for learners’ mental models regarding the systems’ nature or the learners abilities to model the system. in this context with the help of the teaching sequence carried out in this research, there are as already said six objectives. these objectives are that the students should: • understand, through computer simulation, the meaning of the entities of a modeled ecosystem. for example: a “stock” (a box) called “wolves” corresponds to the wolves’ population, and a “faucet control” corresponds to “wolves’ birth rate”. aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 280 • be able to conceptualise the effects that parameter changes on the gui (graphical user interface) of netlogo have on the modeled system. • be in a position to reproduce, on paper, a similar model of an ecosystem, using the same set of symbols as the ones they were faced with, originally. the only things altered are the names of the constituents of the model. • develop the ability of constructing a similar ecosystem dynamics’ model, given only the parts of it. • in their oral and graphical description of the dynamics of the system, be a little more closer on the side of the “expert”, compared to the one of the “novice”, after interacting with the netlogo models, than before interacting with them. the comparison between “expert” and “novice” descriptions, refers to table 1 of jacobson and especially its last column on the right (“associated system thinking abilities”) [3]. • in their oral and graphical description of the dynamics of the system they should get a higher score (as explained in the results) regarding the sbf description of the system [9-10], after the interaction with the netlogo models. a secondary objective, apart from these six, was to be able to realise that while intervening on an (eco)system, the parts of it that are affected are much more than the ones that one initially sees, a skill that relates even to their active citizenship virtues. 2.2 the instrument as a working environment for teaching students, the programming language/environment netlogo was used (version 4.0.4). netlogo is a programming tool suitable for simulating [22], studying and understanding of complex systems. it is a modern variation of the logo programming language, simulating the function of multiagentbased systems. many researchers so far have used netlogo in educating people about understanding the nature of complex systems [23-25]. also specific use of this programming and modeling environment has been done to enhance students modeling abilities and skills, in domains such as graphical modeling or mathematical modeling [18, 20]. each agent of netlogo (typically called a “turtle”) follows a simple set of rules, defined by the writer of the code. this set of rules “guide” the agent in its motion among certain pixels of the screen, named now “patches”. the code additionally defines the action which the turtle should perform on the patch when it meets it, such as changing its colour. the agents act, to a certain extent, independently from each other and yet the “turtle’s” choice of which “patch” to move to and what task to perform on it, is also determined by the “status” of the patch it arrives at the degree of influence being usually determined by the programmer. netlogo provides an extensive “models’ library”, for various complex systems. the two simulations used for our teaching sequence come directly from this 281 advances in systems science and applications (2013) vol.13 no.3 models library and are: the wolf sheep predation model and the “wolf sheep predation (docked)” model [26-27], which is a variation of the first, concerning more the dynamics and the differential equations (mathematics) of the system. both models describe a simple predator (wolf) and prey (sheep) ecosystem, the first of the two giving also the option of adding a third population to the system (the grass). in both these simulations, there are two kinds of agents: the “wolves” and the “sheep”. [the “grass”, which can be also optionally activated, only in the first of the two models, as mentioned above, is not an agent. it is a netlogo “patch” (pixel) property]. the agents are interconnected by relationships of a simple prey-and-predator nature. both kinds of agent follow also a rate of reproduction, controlled by the user of the simulation, and they also die when they run out of energy, apart for the sheep deaths due to predation. the first model, “wolf sheep predation” is used mainly to familiarize the student with the multi-agent simulation, in the sense that the students interact with it, change its parameters and see the results both on the simulation screen but also on the graph screen which, in turn, depicts the time-evolution of the populations. the focus of interest for the teaching sequence is the second simulation, called the “wolf sheep predation (docked)” [27]. when opening the corresponding file, the student is faced with two screens. the one is the typical netlogo console, depicted in fig.1. in the left part of this screen the typical agent-model (simulation) of netlogo for one-prey and one-predator simulation exists. it is called “agent model” and the students may interact with it, as in the previous model, to visualise the time-evolution of the populations when altering parameters of the system. the right part of the screen is called the aggregate model and its evolution is only graphically depicted. this part of the screen describes the evolution of the two populations, based strictly on the mathematical treatment, and provided, in turn, by a form of the famous set of the two lotka-volterra equations [28-29]. epistemologically and didactically, the difference between the two parts of the screen is that on the left part (the simulation/agent screen) randomness plays an important role (“sheep die when and if they encounter wolves”) where as on the right part of the screen (the aggregate model screen) randomness and stochastic behaviour play no role, since the rate of dying of both populations are strictly determined by “rates” (i.e. differential equations, mathematics). apart from this two-sided screen, another, very important for this research, part of the second simulation is the screen called “system dynamics modeler”, created in a java environment, which is shown in fig.2. this screen is exactly the one that focus is given on, being directly related to the “aggregate model”. it represents the dynamic system-model of the simple prey-predator system, depicting additionally its interrelations, aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 282 fig.1 a screenshot of the netlogo model: “wolf sheep predation (docked)” feedback loops and flows/stocks. 2.3 the sample and the settings this teaching sequence was part of a wider research, concerning the teaching and learning of ecosystem complexity at the upper-secondary-education student level, with the use of computers. for this research, a sample of 10 voluntarily participating students was used. the students belonged to two different senior secondary schools, and were at the 2nd class of the greek lyceum (ages between 16 and 17). their orientation was either the technical or science education, therefore guaranteeing a satisfactory background in mathematics, physics and biology. the two schools were of neighbouring areas of athens, greece, and the two groups had similar socioeconomic status, gender mix and school-grade achievement. the students were grouped according to their school achievement in the previous year of the high-school to three categories: –low achievement students (la). their overall grade of graduation from the first class of the senior high school was less than 14 (in the greek scale of 0 to 20, where 20 is the best possible grade). –medium achievement students (ma). their overall grade of graduation from the first class of the senior high school was between 14 and 18.5. (from 18.5 283 advances in systems science and applications (2013) vol.13 no.3 fig.2 a screenshot of the netlogo system dynamics modeler of the model: “wolf sheep predation (docked)”. up to 20, the grades characterization is “excellent” in greece). –high achievement students (ha). their overall grade of graduation from the first class of the senior high school exceeded 18.5. there are two things to notice: a. the choice of the grade score of 14 as a limit is not arbitrary. in the greek high school system very few students fall below 10 in the scale of 0 to 20 and also if one gets a grade of 14, he/she is considered as adequately having followed the class in the respective year of studies, even if he has lost a severe amount of teaching hours. b. in this particular research the parameterisation with respect to the school achievement (la, ma and ha) was not taken into consideration in detail, but only as an indicative index. in further researches of this research group, samples of students are examined with respect to their systemic thinking and systemic modeling abilities and in these researches, the school achievement plays an important role as a parameter. each one of the two groups of the ten students, was taught separately by the first of the authors, for sixteen teaching hours, and he provided each student with 4 worksheets one for each quartet of teaching hours which they aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 284 completed during the process of the instruction. the students worked with computers, in groups of two or three, depending on computer availability. there was never a student sitting alone in a computer screen. in addition to the printed worksheets, students answers and the classroom discussion were tape-recorded. 3 the instruction process and the delivered material first teaching session (4 didactic hours) the first two teaching hours (of the first quartet of hours) were dedicated in familiarizing students with the netlogo environment, particularly with prey-andpredator simulations. for this purpose, the “wolf sheep predation” model was used. within this part of the first teaching session (the 2 teaching hours), the students were asked to handle each button and “slider” in the netlogo console making trials with it and watching the model evolving on the screen with time and try to learn the button’s or slider’s function. after each question on the worksheet, there followed a class discussion, mediated by the instructor. the following are characteristic questions from the first-2-hour worksheet: • question number 3: run the simulation at a low speed (“lower”) and try to find out what the sliders “initial-number-sheep”, “initial-number-wolves”, “wolf gain from food”, “sheep-reproduce” and “wolf-reproduce” actually do. write your answers down. • question number 4: ‘now let us discuss and reach a common conclusion about the role of each of the sliders’. in the rest of the first teaching session (two more teaching hours), through some other questions, the students are introduced visually and verbally to the concept of “(eco)system instability”, since some population may occasionally become extinct. second teaching session (4 didactic hours) during the second teaching session, the simulation “wolf sheep predation (docked)” was introduced. the participants were asked first to investigate the “agent model” section of the netlogo screen, which resembles the previous simulation. they interacted with this section, choosing various values and combinations of values for the sliders, and noticing the model’s evolution with time, as well as its final outcome, both in the simulation area and on the graph. their attention was then driven to the right part of the screen, the “aggregate model”, being now asked to find the meaning and the role of the sliders, noticing that the outcome of the user’s interaction with the system is shown only graphically. simultaneously, the students were given on their computers’ screens a typed presentation about the lotka-volterra equations. the students were then encouraged to discuss their findings about the sliders’ role with the class. next, the 285 advances in systems science and applications (2013) vol.13 no.3 students were prompted, by the worksheet, to write down or discuss verbally the relation and differences between the agent and the aggregate model, observing the effects live, by pushing the “compare” and “step compare” buttons. third teaching session (4 didactic hours) in the third teaching session (4-hour), the class was introduced to the system dynamics modeler. the aim was to conceptualize the direct connection of this model with the aggregate models’ parameters and the way in which this relation is established. the students were asked what the flows, stocks, arrows and faucet controls actually mean in the modeler, and how they can interfere with them (through the mouse of their pc). towards the end of the worksheet, a first low-level direct objective was achieved, to see what the hypothetical effect would be on the system dynamics modeler, when they altered the position of a slider in the aggregate model worksheet. question number 6: fill in the gaps below and create similar sentences: “when i increase ...... in the simulation, in essence i make the box (stock) ...... in size”. “when i reduce ...... in the simulation, in essence i make the arrow ...... in width.”, or “let less flow pass through the faucet control ......”. we also asked for verbal answers, which we tape-recorded. later in the same session, the students were given the opportunity to think and reconsider question 6. the issue is whether only one thing would be affected in the modeler if one changed something in the netlogo screen, or more things would be affected. afterwards, the notion of “systemic thinking” and especially “ecosystemic thinking” was introduced, and its importance to the conservation of the earth, as a whole. at the end of this third 4-hour session, a paper was given to the students, reading: “we see the parts of an ecosystem as interrelated and not as separate entities (this is called “reductionist thinking”). it is easy to see that human interference constitutes one of these parts and can affect many more parts than those that we initially estimated!!” fourth teaching session (4 didactic hours) in the last session, further skills of systemic thinking and modeling were developed, through the system dynamics modeler. mainly, it consisted of two stages: at the first stage, students were given a piece of paper other than the worksheet, which was identical to the model depicted in the modeler, but the terms the words had been removed (refer to fig.3). here the students were asked to fill in the following words or phrases: water deposits in an area / average rainfall rate in the area / birth rate in the area / water returning to the sea, to the lakes or to the under-surface horizon / consumption of water in cubic meters per person / population increase in the area / water increase in the area / peoples’ deaths in the area / population of the area aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 286 fig.3 the system dynamics modeler of the netlogo model: “wolf sheep predation (docked)”, without the terms in it.. / rate of deaths in the area. they did this filling of gaps in the way they think proper. at the second stage of this fourth session, students were given the elements of the system dynamics modeler, as shown at fig.4. they were allowed to use each part/element of the figure freely, in order to fig.4 the elements (stocks, flows etc) of the dynamic model of the system. create a description of the following (eco)-system: “the amount of solar energy entering the baikal lake in russia, in relation to the seaweeds growing and dying in this lake, and the photosynthesis they perform, engaging this energy.” a crucial question was posed to the students after the second and before the third teaching session. it was related to the skills of “novice’s” vs. “expert’s” description of the system as given by jacobson [3], as well as to the sbf understanding of the system as given by hmelo-silver et. al. [9-10]). the question was this: 287 advances in systems science and applications (2013) vol.13 no.3 overall question: “you are asked now to write an essay (around one page or one page and a half) about what you do see in that system of the populations. describe, among other things, what this ecosystem consists of, what each part of the ecosystem does, and what the ecosystem does as a whole. when is a population extinct? when is it endangered? when does a population fall and when does it augment tremendously? what parameters affect each population and how does a population affects (or is affected by) another or others? you are also encouraged a lot to make drawings about the relations of the populations, using “boxes”, various kinds of “arrows”, “faucets”. write things or comments, or titles (legends) on or in every part of your figure that you thing it is appropriate to write. write and draw a lot!” the same question was also handled to the students after the completion of the fourth and final teaching session of the whole sequence, and the results were compared. 4 results and discussion despite the limitation of the small size of the sample, which does not allow to draw general conclusions, and makes this more of a case study, this research depicted that the use of the netlogo as a teaching environment, together with the system dynamics modeler, could potentially help high-school students in understanding simple (eco)system structures, and to acquire or slightly improve their skills on representing and even building models. triangulating the results with the ones of using other modeling tools than the system dynamics modeler of netlogo, such as the stagecast creator (sc) [30], used for primary school students, the oral answers of the students were used in the present survey as an encouraging feedback with respect to the use of the software as a means of understanding the model. at first the results of this research are presented in an indicative form. the answer of a girl, christine, of low school achievement, is quoted, to the tape-recorded discussion, following the aforementioned question number 6, in the third teaching session: ‘yes, by varying the sliders positions in the aggregate model, i can better see what every shape in the system dynamics modeler means.’ in question number 7, of the same worksheet, where they were asked to reconsider their answer in number 6, they are prompted to see the wider interconnectedness of the factors in this simple prey-and-predator population system. this is not at all easy for high-school students, as was shown in the research of barman, griffiths and okebukola, recorded as “misconception no. 2”[31]. there, students were found rather to believe that the change in one population affects only its closest pray and itself. in the answer to question 7, a student, john, is aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 288 quoted, who is of a medium school achievement: ‘it seems that one motion of the slider can change the whole system. so if we touch the “predator-efficiency”, we increase the number of births of wolves, thus we increase predation-rate of them which, in turn, increases the number of dying sheep. as a result, the number of wolves is increased in one hand, but the amount of sheep (food) is decreased, so the wolves will have problems. many such changes are done.’ the concept investigated and aimed at improving here is a concept of what white calls “naive ecology”[32]. the relations in this simple prey-and-predator system are strictly causal (‘the wolf eats the sheep thus sheep die’, or ‘the wolves have no sheep to eat, thus they gradually become extinct’) and the learners are asked to follow such causal relations, as far as they can, relatively to the initial point of focus, which could be everything (e.g. ‘wolves die’). white’s research also depicted that the persons cannot follow such a causal string of facts in ecology too far from the point of origin. the students’ answers are also close enough to the eco-systemic, holistic conclusion, before this was finally distributed to them, written on paper, therefore their stands and attitudes seem to have been adequately investigated and, possibly, changed. katherine, a student with school achievement belonging to the lower part of the sample, wrote: ‘this indicates that whatever we touch in an ecosystem, this also affects the other factors too, and, therefore, we must never change something that needs not be touched, because this would affect many more factors than the ones we initially estimated!’ the one of the two higher system-thinking objectives, which is to represent a very similar system but with differently named parts or terms, was satisfactorily achieved, as is represented by the fig.5 given by denis, who belonged to the medium area of the sample as regards school grades achievement. this skill here refers mainly to the systems-based inquiry (s-bi) protocol used by sweeney and sterman, in ecosystems, and especially its first two parts: part i: systemic scenarios: participants consider system dynamics in [six] simple scenarios. these scenarios emphasize feedback dynamics. part ii: homology challenges: participants imagine [six] related but different systemic scenarios. this requires participants to use homological reasoning.’ [12]. homological reasoning, as well as analogical system-thinking, both in their very simple expressions, was cultivated, to a certain extend, to all of the students, as denis’s drawing (who reflects an average of the sample) reveals. passing to an ecological phenomenon other than “predation”, which is “water consumption”, the students appear to be 100% correct in adopting the terminology and conceptualising the analogies. 289 advances in systems science and applications (2013) vol.13 no.3 fig.5 the denis’ first sheet. finally the other, even higher, system-thinking and system-constructing objective, aimed at, was to construct a model again similar to the initial having only the elements of it. such modeling tasks are not easy, as researchers have shown [17, 33]. hogan used only textbook-based teaching and asked the learners to construct simple food webs, in the sense of only putting the arrows in them (‘who eats what’), giving the reasoning for their choices. the results were poor. jensen and brehmen asked a sample of post-secondary-education individuals to study a predator-and-prey model very close to the one we chose. it is depicted in fig.6. the task here was to establish population equilibrium in this system. the learners tried this on the computer (changing a numeric parameter, i.e. the “foxes”) and what they relied on, was either mathematical approach or negative feedback-loops. the results were again poor (less than 50% did it) and as one of the main reasons for this low performance, the authors give: ‘one possible reason why the rabbit-fox problem is difficult may be that it is not easily decomposed into subsystems or homomorphs. since they are so closely intertwined, it is not feasible to split the system into a rabbit system and a fox system. therefore it is not possible to simplify it. this may contribute to the difficulties in forming a clear-cut model of the system.’ [18]. aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 290 fig.6 the figure of jensen and brehmer fig.7 the sheet as completed by george (at left) and christine (at right). therefore, one important competency in understanding an (eco)system, is decomposing it to its parts, which is what the half of our last activity is about. the 291 advances in systems science and applications (2013) vol.13 no.3 other half of the activity is attempting, with these same parts, to build a similar ecosystem. as a measure of the results of this last activity on the sample, we present two drawings: figure 7 depicts the model made by the highest-school-achievement student of the whole sample (george), as well as the one made by the lowestschool-achievement student of the whole sample (christine). it can be noticed that, even though the “better student” has made a more flexible shape, and he adds even subtle-unneeded details, whereas the “weaker student” sticks more to the original and omits certain terms, the basic representation of this alternative ecosystem (“lake baikal and the seaweeds”) is essentially achieved in both drawings. as regards the “numerical” representations of the results, each one concerning the six didactic objectives mentioned in section 2.1 of the current paper, the results were: first objective. the 7 out of the 10 students corresponded totally correctly the elements of the system dynamics model with the entities it describes. it seems to have been an easy task, especially with the aid of the netlogo model. second objective. 8 out of 10 students gave all the correct answers when writing what is affected in the modeled system by each slider’s and each button’s variation. again the proportion of successful completion of the task is high, combined with the use of the netlogo models. third objective. once again 7 out of the 10 learners constructed on paper a quite similar model of an ecosystem, with only the names of the terms altered. high performance also appeared here, with the help of the netlogo screen. fourth objective. here as was expected there was relative failure. only 3 out of 10 students succeeded in building correctly on paper, the dynamics’ model of a system “from the scratch”, given only the elements of it. it is a high systemic-thinking ability and, even with netlogo, it cannot be easily enhanced. fifth objective. as regards the table 1 of the classic paper of jacobson a score was attributed to each of the 10 students, as regards their performance of the last right column of table 1. the question of reference is of course the “overall question”, mentioned before in the paper, in which the students’ answers are scrutinized with respect to their content. each time a student acts as a “novice” on the modeled (complex) ecosystem, in his/her writings, descriptions or in his/her shapes of the system’s dynamics he/she gets a “0”, for the corresponding aspect mentioned in aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 292 the last column on the right of table 1. on the contrary, each time the student acts as an “expert” on the modeled ecosystem he/she gets an 1 for this aspect. since there are five aspects of systemic thinking, it is obvious that the maximum score each student could take is “5” and the minimum score is “0”. for the whole sample of 10 students the respective maximum score is “50” and the minimum score is “0”. the results and the scores of this attributing system are shown in table 3, both before and after the study of the netlogo’s system dynamics’ modeler (the “overall question” was delivered to them twice). interpreting the table 3 the (preand post-interaction with the model) scores achieved by the 10 students, as regards the jacobson [3] model of mental modes, component beliefs and the associated with them system-thinking abilities. category of component belief and associated system-thinking ability overall score achieved by the 10 students before interacting with the two netlogo models overall score achieved by the 10 students after interacting with the two netlogo models 1. understanding phenomena 42 44 4. action effects 14 17 6. complex actions 20 39 7. final causes or purposefulness of natural phenomena 19 38 8. ontology 12 15 findings of table 3, one can see that: – the study of the system dynamic modeler of netlogo and the interaction with netlogo model “wolf sheep predation (docked)” was very helpful for the students , as regards “complex actions” and “final causes or purposefulness” of systems, which in systemic terms means it helped them to “see very distant interactions within elements of the system” and to “see closed loops within the system and not only open processes”. especially the first of these two aspects, is in close proximity, with what sharona t. levy and uri wilensky define as a “mid-level” which means that the students understand better a system if they can describe it by means of the actions of a group of neighbouring and interacting parts [34], not by the actions of one single part alone and not by the action of all the parts as a whole. – in the domain of “understanding phenomena”, which in systemic terms means: “seeing most the compartments and the flows in the system”, the students performance was significantly good from the beginning, and the interaction with the system dynamic modeler of netlogo and the model itself, failed to improve it significantly. – in the domains of “action effects”, and “ontology” which in systemic terms mean respectively: “seeing beyond proportionality between stocks’ size and flows’ rates in the system”, and “seeing the size of stocks as dynamic, even when it looks 293 advances in systems science and applications (2013) vol.13 no.3 constant”, the students performance was poor in the beginning, and the interaction with the system dynamic modeler of netlogo and the model itself, seems to have failed also to make it better. sixth objective. once more, the question that was scrutinized is the “overall question” and the drawings or the writings of students in it, about the system, both before and after their study of the system dynamics modeler and the “wolf-sheep predation (docked)” model of netlogo. analyzing again the content of the writings and the drawings of the students in this question, a scoring system was established: – for each one of the three “structure” elements of the system that the student managed to find, he got a score “1”. consequently the score for each student would range from “0” to “3”, and the score for the whole sample of 10 students would range from “0” to “30”. – for each one of the nine “function” elements of the system that the student managed to find, he got a score “1”. consequently the score for each student would range from “0” to “9”, and the score for the whole sample of 10 students would range from “0” to “90”. – additionally, it was decided that for each “faucet control” or “rate” that the student would describe, for the corresponding “function” element, he would get “0.5” point more, since this reveals a slightly deeper understanding of the relative “function” aspect of the system. thus, the score for each student here would range from “0” to “4.5”, and the score for the whole sample of 10 students would range from “0” to “45”. the cumulative results for sbf treatment of system dynamics are shown in table 4. the results show that the interaction with the system dynamics’ modeler of table 4 the (preand post-interaction with the model) scores achieved by the 10 students, as regards the sbf model [9] of knowledge of , learning about and description of the system and the associated with them system dynamics’ expressions (“behavior” is excluded). aspect of the systems description according to the sbf model of knowledge and learning (“behavior” is excluded, as said) “structure” overall score achieved by the 10 students before interacting with the two netlogo models overall score achieved by the 10 students after interacting with the two netlogo models “structure” 21 28 “function” 33 68 “function” with “faucet controls” 10 41 netlogo, as well as with the model “wolf-sheep predation (docked)”: – helped the students conceptualise better to a certain extent the elements that constitute the system (the “structure”) . it should be stressed here that aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 294 the element mostly missed before interaction was “grass”, which is an expected result. – helped the students much more to conceptualise the “functional” aspects of the system, which are mainly: (i) the “flows” (inflows, outflows) and (ii) the “interaction arrows” between system elements. this is a gain for system dynamics’ knowledge and students’ systemic thinking. – proved extremely helpful in making students conceptualise and describe the “faucet controls” which are actually the “rates”, within the system. this is an aspect of the system description and understanding of its “function”, that was very poor before the interaction with the netlogo models and the dynamics’ modeler, and improved significantly afterwards. 5 conclusions taking into consideration the restrictions of the current research (small sample size, absence of pre-test and post-test evaluations) and the imposed settings on the teaching sequence (few hours of teaching, students working out of their classroom schedule) , it is, nevertheless, argued, that there are significant conclusions from the research. at first, it is concluded that the combined use of a netlogo model and the system dynamics modeler associated with it can make students improve their understanding and their system-analysing abilities of a simple modeled ecosystem. secondly, it is concluded that upon interaction with netlogo models and system dynamics’ modeler, a student who is a “novice” in conceiving and describing systems and relations within systems, can slightly move towards the direction of being an “expert” on it, or at least on some of the system’s aspects. a third conclusion is that low-level system thinking abilities exist at large even before working with netlogo model, and remain there after interacting with it. the same is valid, on the opposite view, for high-level thinking abilities. they are mostly absent and remain absent after the interaction with the netlogo models and the system dynamics modeler. the abilities that are mostly enhanced by netlogo, are the medium-level system-thinking, system-analysing and systemrepresenting abilities. finally, a core concept of this research and a conclusion of it, is that a learner who is weak in: (i) seeing detailed structures in a system, (ii) seeing loops, (iii) predicting time-evolutions in it, (iv) observing large-reaching interactions among the elements of the system and (v) making generalizations about the system behavior will definitely benefit somehow, if he/she is taught about the system, not only with the graphic representation of the system dynamics but also with the study of (and interaction with) an agent-based simulation (such as a netlogo 295 advances in systems science and applications (2013) vol.13 no.3 model) of the system. questions regarding the time-evolution of the system, questions regarding extreme events (“what if a stock becomes too large or too small”? “what if a stock vanishes”? “what if a flow diminishes in thickness”?) and questions regarding random events (“stochasticity”) (“what if wolves find no sheep to eat”?) may not be easily answered for the novice, by the graphic or mathematical representation of the system dynamics, but are clarified by the netlogo simulations. references [1] evagorou m., korfiatis k., nicolaou c. and constantinou, c. (2009), “an investigation of the potential of interactive simulations for developing system thinking skills in elementary school: a case study with fifth-graders and sixth-graders”, international journal of science education, vol.31, no.5, 15 march 2009, pp.655-674. [2] levin s., a. (1998), “ecosystems and the biosphere as complex adaptive systems”, ecosystems, vol.1, pp.431-436. [3] jacobson m. j. (2001), “problem solving, cognition, and complex systems: differences between experts and novices”, complexity, vol.6, pp.41-49. [4] lesh r. (2006), “modeling students’ modeling abilities: the teaching and learning of complex systems in education”, journal of the learning sciences, vol.15, no.1, pp.45-52. [5] holland j. (1995), hidden order: how adaptation builds complexity, addison-wesley, reading, ma. [6] jorgensen s., e. (editor). (2009), ecosystem ecology, editions: elsevier b.v. amsterdam, the netherlands. [7] d’appolonia s.t., charles e.s., & boyd g.m. (2004), “acquisition of complex systemic thinking: mental models of evolution”, educational research and evaluation, vol.10, pp.499-521. [8] american association for the advancement of science, aaas. (1993), benchmarks for scientific literacy, oxford university press, new york. [9] hmelo c., e. holton, d. l. & kolodner, j., l. (2000), “designing to learn about complex systems”, the journal of the learning sciences, vol.9, no.3, pp.247-298. [10] hmelo-silver c. e. & pfeffer m., g. (2004), “comparing expert and novice understanding of a complex system from the perspective of structures, behaviors and functions”, cognitive science, vol.28, pp.127c138. aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 296 [11] hmelo-silver c. e., marathe s. & liu l. (2007), “fish swim, rocks sit and lungs breathe: expert-novice understanding of complex systems”, journal of the learning sciences, vol.16, no.3, pp.307-331. [12] sweeney l., b. and sterman j. d. (2007), “thinking about systems: student and teacher conceptions of natural and social systems”, system dynamics review, vol.23 no.2/3, (summer/fall 2007), pp.285-311. [13] schieritz n. (2002), “integrating system dynamics and agent-based modelling”, proceedings of the 20th international conference of the system dynamics society, palermo, italy. july 28 august 1, 2002. retrieved 20th march 2013, from: http://www.systemdynamics.org /conferences/2002/proceed/papers/schieri1.pdf [14] schieritz n., & grossler a. (2003), “emergent structures in supply chains a study integrating agent-based and system dynamics modelling”, proceedings of the 36th hawaii annual international conference on system sciences, hawaii, usa, 6-9 january 2003. [15] bousquet f. & le page c. (2004), “multi-agent simulations and ecosystem management: a review”, ecological modelling, vol.176, no.3-4, pp.313-332. [16] meij, jan van der & jong ton de (2006), “supporting students’ learning with multiple representations in a dynamic simulation-based learning environment”, learning and instruction, vol.16, no.3, pp.199-212. [17] hashem k. & mioduser d. (2011), “the contribution of learning by modeling (lbm) to students’ understanding of complexity concepts”, international journal of e-education, e-business, e-management and e-learning, vol.1, no.2, pp.151-155. [18] jensen e. and brehmer b. (2003), “understanding and control of a simple dynamic system”, system dynamics review, vol.19, no.2, pp. 119-137. [19] thompson k. (2007), models as mindtools for environmental education: how do students use models to learn about a complex socio-environmental system? phd thesis, the centre for research on computer-supported learning and cognition, faculty of education and social work, the university of sydney, australia. [20] thompson k. (2008), “the value of multiple representations for learning about complex systems”, proceedings of the 8th international conference for the learning sciences (icls), utrecht, the netherlands. vol.2, pp.398406. 297 advances in systems science and applications (2013) vol.13 no.3 [21] levy s. t. & wilensky u. (2011), “mining students’ inquiry actions for understanding of complex systems”, computers & education, vol.56, issue 3, pp.556c573. [22] wilensky u. (1999), netlogo. [online] “center for connected learning and computer-based modeling”, northwestern university, evanston, il. available at: http:// ccl.northwestern.edu/ netlogo. [accessed 2 june 2009]. [23] levy s. t. & wilensky u. (2009), “students’ learning with the connected chemistry (cc1) curriculum: navigating the complexities of the particulate world”, journal of science education and technology, vol.18, no.3, pp.243254. [24] goldstone r. l. & wilensky u. (2008), “promoting transfer by grounding complex systems’ principles”, journal of the learning sciences, vol.17, no.4, pp.465-516. [25] vattam s. s., goel a. k., rugaber s., hmelo-silver, c. e., jordan, r., gray, s., & sinha, s.(2011), “understanding complex natural systems by articulating structure-behavior-function models”, journal of educational technology & society, vol.14, no.1, pp.66c81. [26] wilensky u. (1997), “netlogo wolf sheep predation model”, [online] center for connected learning and computer-based modeling, northwestern university, evanston, il. available at: http://ccl.northwestern.edu/netlogo /models/wolf sheep predation.[accessed 2 june 2009]. [27] wilensky u. (2005), “netlogo wolf sheep predation (docked) model”, [online] center for connected learning and computer-based modeling, northwestern university, evanston, il. available at: http://ccl.northwestern.edu/netlogo/models/wolfsheeppredation(docked). [accessed 2 june 2009]. [28] lotka a.j. (1956), elements of mathematical biology, dover, new york. [29] wilensky u. & reisman k. (1999), “connected science: learning biology through constructing and testing computational theories an embodied modeling approach”, inter journal complex systems(cx)[online], manuscript. no.234. availableat: http://www.interjournal.org/. [30] papaevripidou m., constantinou c.p. and zacharia z.c. (2007), “modeling complex marine ecosystems: an investigation of two teaching approaches with fifth graders”, journal of computer assisted learning, vol.23, pp.145157. aristotelis gkiolmas: using simulations of netlogo as a tool for introducing greek ... 298 [31] barman c. r., griffiths a. k. and okebukola p. a. o. (1995), “high school students’ concepts regarding food chains and food webs: a multinational study”, international journal of science education, vol.17, no.6, pp.775782. [32] white p. a. (1997), “naive ecology: causal judgments about a simple ecosystem”, the british journal of psychology, vol.88, may ’97, pp.219-233. [33] hogan k. (2000), “assessing students’ systems reasoning in ecology”, journal of biological education, vol.35, no.1, pp.22-28. [34] levy s. t. & wilensky u. (2008), “inventing a ‘mid level’ to make ends meet: reasoning between the levels of complexity”, cognition and instruction, vol.26, no.1, pp.1-47. [35] wilensky u. and reisman k. (2006), “thinking like a wolf, a sheep or a firefly: learning biology through constructing and testing computational theoriescan embodied modeling approach”, cognition and instruction, vol.24, no.2, pp.171-209. corresponding author author can be contacted at: agkiolm@primedu.uoa.gr 82-92 advances in systems science and applications (2011), vol. 11, no. 1-2 existence and nonexistence of positive solutions for a kirchhoff-type equation with inhomogeneous strong allee effect ruyun ma and guowei dai department of mathematics, northwest normal university, lanzhou 730070, p.r. china email: mary@nwnu.edu.cn, daiguowei@uwnu.edu.cn abstract in this paper, we deal with the nonlocal semilinear elliptic equation with inhomogeneous strong allee effect { −m (∫ ω 1 2 |∇u| 2 dx ) ∆u = λf(x, u) in ω, u = 0 on ∂ω, where the nonlocal coefficient m (∫ ω 1 2 |∇u| 2 dx ) is a continuous function of ∫ ω 1 2 |∇u| 2 dx. by means of variational approach, we prove that the equation has at least two positive solutions for largeλ under suitable hypotheses about nonlinearity. we also prove some nonexistence results. in particular, we shall give a positive answer to the conjecture by liu, wang and shi’s of [1]. keywords nonlocal differential equation variational method positive solutions inhomogeneous strong allee effect. 1. introduction in this paper we study the following problem{ −m (∫ ω 1 2 |∇u| 2 dx ) ∆u = λf(x, u) in ω, u = 0 on ∂ω, (1) where ω is a smooth bounded domain in rn for n ≥ 1, the nonlocal coefficient m(t) is a continuous function of t = ∫ ω 1 2 |∇u| 2 dx. we shall give a positive answer to a conjecture by liu, wang and shi’s of [1]. the problem (1) is a generalization of a model introduced by kirchhoff[2]. more precisely, kirchhoff proposed a model given by the equation ρ ∂2u ∂t2 − ( ρ0 h + e 2l ∫ l 0 ∣∣∣∣∂u∂x ∣∣∣∣2 dx ) ∂2u ∂x2 = 0, (2) where ρ, ρ0, h, e, l are constants, which extends the classical d’alembert’s wave equation, by considering the effect of the changing in the length of the string during the vibration. a distinguishing feature of equation (2) is that the equation contains a nonlocal coefficient ρ0 h + e 2l ∫ l 0 ∣∣∂u ∂x ∣∣2 dx which depends on the average 1 2l ∫ l 0 ∣∣∂u ∂x ∣∣2 dx of the kinetic energy 1 2 ∣∣∂u ∂x ∣∣2 on [0, l], and hence the equation is no longer a pointwise identity. the equation{ − ( a+ b ∫ ω |∇u| 2 dx ) ∆u = f(x, u) in ω, u = 0 on ∂ω, (3) issn 1078-6236 international institute for general systems studies, inc. advances in systems science and applications (2011), vol. 11, no. 1-2 83 is related to the stationary analogue of the equation (2). equation (3) received much attention only after lions [3] proposed an abstract framework to the problem. some important and interesting results can be found, for example, in [4-15]. in the context of population biology, the nonlinear function f(x, u) ≡ ug(x, u) represents a density dependent growth if g(x, u) is a function depending on the population density u. while traditionally g(x, u) is assumed to be declining to reflect the crowding effect of the increasing population, allee suggested that physiological and demographic precesses often possess an optimal density, with the response decreasing as either higher or lower densities. such growth pattern is called an allee effect. if the growth rate per capita is negative when u is small, we call it a strong allee effect; if the growth rate per capita is small than the maximum but still positive for small u, we call it a weak allee effect (for detail, see [16] or [17]). under the special case of equation (3) with a = 1, b = 0 and f(x, u) satisfies inhomogeneous strong allee effect growth pattern, liu, wang and shi[16] prove that the equation{ −∆u = λf(x, u) in ω, u = 0 on ∂ω (4) has at least two positive solutions for large λ if ∫ c(x) 0 f(x, s) ds > 0 for x in an open subset of ω, where c(x) ∈ c1(ω) such that f(x, c(x)) = 0(see the assumption of (f2)). they also prove some nonexistence results. in particular, they conjecture that the nonexistence holds if∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω (see remark 1.7 of [16]). motivated by above, we generalize existence and nonexistence results for the semilinear elliptic equation (4) to the case of nonlocal semilinear elliptic equation (1). more precisely, if f(x, u) satisfies inhomogeneous strong allee effect growth pattern and the nonlocal coefficient m(t) satisfies some suitable conditions, we establish the existence of at least two positive solutions for the nonlocal problem (1) withλ large enough. we also prove some nonexistence results for the nonlocal problem (1). in particular, we shall give a positive answer to the conjecture by liu, wang and shi. to the best of our knowledge, this is the first paper that discusses the nonlocal semilinear elliptic equation with inhomogeneous strong allee effect via variational method. we point out the nonlocal coefficient m(t) raises some of the essential difficulties. for example, the way of proving the geometry condition of mountain pass theorem in [16] can not be used here because the functional of (1) is notc2 function under our assumptions. in order to overcome this difficulty, we divided ω into b1 and b2 by comparing the value of c(x) with b, then use poincaré inequality to prove it(see lemma 3.3). this paper is organized as follows. in section 2, we present our main resuts and some necessary preliminary lemmas. in sections 3, we use variational method and sub-supersolution method to prove the main results. in section 4, we prove the conjecture of liu, wang and shi’s and give some examples which satisfy our hypotheses. 2. main resuts and preliminaries in this section, we give our main results and some necessary preliminary lemmas which will be used in the following proof. for simplicity we write x = h1 0 (ω) with the norm ‖u‖ = 84 ma:existence and nonexistence of positive solutions for a kirchhoff-type . . . . . .(∫ 1 0 |∇u| 2 dx ) 1 2 . hereafter, f(x, t) and m(t) are always supposed to verify the following assumptions: (f1) f(x, u) ∈ c(ω× r+) and f(x, ·) ∈ c1(r+) for any x ∈ ω; (f2) there exist b(x) ∈ c(ω), c(x) ∈ c1(ω) such that 0 < b(x) < c(x) and f(x, 0) = f(x, b(x)) = f(x, c(x)) = 0 for any x ∈ ω; (f3) for almost all x ∈ ω, f(x, s) < 0 for any s ∈ (0, b(x)) ∪ (c(x),+∞) and f(x, s) > 0 for any s ∈ (b(x), c(x)). remark 2.1. note that the weak maximum principle (theorem 8.1 of [18]) and strong maximum principle (theorem 8.1 of [18]) also hold for the nonlocal problem (1) because m(t) satisfies the assumption (m ). (m ) ∃m0 > 0 such that m(t) ≥ m0. definition 2.1. we say that u ∈ x is a weak solution of (1), if m (∫ ω 1 2 |∇u|2 dx )∫ ω ∇u∇ϕdx = λ ∫ ω f(x, u)ϕdx for any ϕ ∈ x. define φ(u) = m̂ (∫ ω 1 2 |∇u|2 dx ) ,ψ(u) = ∫ ω f (x, u) dx, where m̂(t) = ∫ t 0 m(s) ds, f (x, u) = ∫ u 0 f(x, t) dt. we redefine f(x, u), such that f(x, u) ≡ 0 when u ∈ (−∞, 0) ∪ (c(x),∞), but it does not change the solution set of (1) by the weak maximum principle, since all the solution of (1) satisfies 0 ≤ u(x) ≤ c(x). then the energy functional iλ(u) = φ(u) − λψ(u) : x → r associated with problem (1) is well defined. then it is easy to see that iλ ∈ c1 (x,r) is weakly lower semi-continuous and u ∈ x is a weak solution of (1) if and only if u is a critical point of iλ. from the regularity assumptions on f(x, u), any critical point u of iλ(·) is a classical solution of (1) (see [19, 20]), and from the strong maximum principle and the definition of modified f(x, u) above, u is either zero or satisfies 0 < u(x) < c(x) for any x ∈ ω. moreover, we have i ′λ(u)v = m (∫ ω 1 2 |∇u|2 dx )∫ ω ∇u∇v dx− λ ∫ ω f(x, u)v dx = φ′(u)v − λψ′(u), for any v ∈ x. from (m ) and lemma 4.1 of [21] we can easily see that φ′ is of (s+) type, i.e. if un ⇀ u in x and lim n→+∞ (φ′(un)− φ′(u), un − u) ≤ 0, then un → u in x . it is clear that ψ′ is weak-strong continuous (or see lemma 1.2 of [1]). so i ′λ is of (s+) type. advances in systems science and applications (2011), vol. 11, no. 1-2 85 our main existence result is as follows: theorem 2.1. if m(t) satisfies (m ) and f(x, u) satisfies (f1)–(f3), and ω1 is an open subset of ω such that ∫ c(x) 0 f(x, s) ds > 0 (5) for x ∈ ω1, then for λ large enough, (1) has at least two positive solutions, and (1) has no solution for small λ. in order to prove our main existence result we need the following lemma: lemma 2.1 (see [1].) suppose that f satisfies (f1)–(f3). if u(x) is an integrable function in ω, and there is a measurable subset ω0 of ω with positive measure, such that∫ c(x) 0 f(x, s) ds > 0 in ω0 and ∫ c(x) 0 f(x, s) ds ≤ 0 in ω \ ω0, then ∫ u(x) 0 f(x, s) ds ≤ ∫ c(x) 0 f(x, s) ds in ω0 and ∫ u(x) 0 f(x, s) ds ≤ 0 in ω \ ω0, now we turn to the nonexistence of the positive solutions of (1) when (5) does not hold for any x ∈ ω. we define c = maxx∈ω c(x), f(u) = maxx∈ω f(x, u). our main nonexistence result is theorem 2.2. if ∫ c 0 f(u) du ≤ 0, then (1) has no positive solution for any λ > 0. in order to prove our main nonexistence result, we recall a theorem in [22] for (1) with the special case of m(t) ≡ 1 and f(x, u) ≡ f(u). in fact, the theorem also holds for the nonlocal problem (1) with f(x, u) ≡ f(u). because the proof is similar to the proof of[22], we omit it here (for detail, see the proof of theorem 1 in [22]). let us assume that f : r→ r is a c1 function and let the following conditions hold: there exist 0 ≤ s0 < s1 < s2, such that f(si) = 0, i = 1, 2, f(s0) ≤ 0, f(s) < 0, s0 < s < s1, f(s) > 0, s1 < s < s2 (6) and let ∫ s2 s0 f(s) ds ≤ 0. (7) we have the following lemma. lemma 2.2. assume that f satisfies (6) and (7). let ω be a bounded domain with smooth boundary. if (1) with f(x, u) ≡ f(u) has a positive solution u, then u can not satisfy{ umax = maxx∈ω u(x) ∈ (s1, s2), u(x) > 0, x ∈ ω. (8) 86 ma:existence and nonexistence of positive solutions for a kirchhoff-type . . . . . . remark 2.2. note that our assumptions (f1)–(f3) are weaker than (f1)–(f4) of[1] even in the case ofm(t) ≡ 1. in fact, from (f1)–(f3), we can easily see that there exist a positive constant β such that f(x, s) ≤ βs for any s ≥ 0 and a.e. x ∈ ω, i.e., the condition (f4) of [1]. we do not need the conditions of b(x) ∈ c1,α(ω)(0 < α < 1) and f(·, u) ∈ c1,α(ω) for any u ≥ 0 because we do not need energy functional of (1) is ac2 function in x in our proof. remark 2.3. the condition of f(x, ·) ∈ c1(r+) for any x ∈ ω can be relaxed to f(x, ·) is locally lipschitz in r+ for any x ∈ ω. in fact, lemma 2.2 also holds when f : r → r is a locally lipschitz function because the symmetry results of [23] holds under this weaker condition. 3. proof of theorem 2.1 and 2.2 in this section we will prove theorem 2.1 and 2.2. lemma 3.1. if m(t) satisfies (m ), and f(x, u) satisfies (f1)–(f3) and (5), then for λ large enough, iλ(·) has a global minimum point u1 such that iλ(u1) < 0. proof. since ∫ c(x) 0 f(x, s) ds > 0 in ω1, then there exists a measurable set ω0 ⊂ ω with positive measure, such that ∫ c(x) 0 f(x, s) ds > 0 in ω0, and ∫ c(x) 0 f(x, s) ds ≤ 0 in ω \ ω0. from (m ) and the definition of m̂(t), we have m̂(t) ≥ m0t. in view of lemma 2.1, we have iλ(u) = m̂ (∫ ω 1 2 |∇u|2 dx ) − λ ∫ ω f (x, u) dx ≥ m0 ∫ ω 1 2 |∇u|2 dx− λ ∫ ω (∫ u(x) 0 f(x, s) ds ) dx ≥ m0 ∫ ω 1 2 |∇u|2 dx− λ ∫ ω0 (∫ u(x) 0 f(x, s) ds ) dx− λ ∫ ω\ω0 (∫ u(x) 0 f(x, s) ds ) dx ≥ m0 ∫ ω 1 2 |∇u|2 dx− λ ∫ ω0 (∫ c(x) 0 f(x, s) ds ) dx ≥ m0 ∫ ω 1 2 |∇u|2 dx− λ ∫ ω0 a1 dx = m0 2 ‖u‖2 − λ|ω0|a1 → +∞, as ‖u‖ → +∞, (9) where a1 = maxx∈ω0 |f (x, c(x))|. since iλ is weakly lower semi-continuous, iλ has a minimum point u1 in x . next we shall prove iλ(u1) < 0, thus u1 is a positive solution of (1). in fact, we only need to verify that when λ is large there exists a u0 ∈ x , such that iλ(u0) < 0 = iλ(0). we define u0(x) = 0 in ω \ ω1ε, and u0(x) = c(x) in ω1 and properly in ω1ε \ ω1 such that u0 ∈ x , where ω1ε = {x ∈ ω : dist(x,ω1) ≤ ε}. using the similar method with[1], we have iλ(u0) ≤ m̂ (∫ ω 1 2 |∇u0|2 dx ) − λ ∫ ω1 f (x, c(x)) dx− λ [−a1 (|ω1ε| − |ω1|)] dx. (10) advances in systems science and applications (2011), vol. 11, no. 1-2 87 since ∫ c(cx) 0 f(x, s) ds > 0 when x ∈ ω1 and ∫ c(x) 0 f(x, s) ds is continuous, then there must exists an open subset ω2 with ω2 ⊂ ω1 and δ > 0, such that |ω2| > 0 and ∫ c(x) 0 f(x, s) ds ≥ δ for x ∈ ω2. choose ε small enough, such that δ|ω2| + a(|ω1| − |ω1ε|) > 0. again using the similar method with [1], we have iλ(u0) ≤ m̂ (∫ ω 1 2 |∇u0|2 dx ) − λ [δ|ω2|+a1(|ω1| − |ω1ε|)] . therefore when λ large enough, iλ(u0) < 0, and consequently when λ is large enough, (1) has a positive solution u1(x) satisfying iλ(u1) = infu∈x iλ(u) < 0. next we shall use mountain pass theorem to prove that (1) has another positive solution u2. first we prove iλ(u) satisfies palais-smale condition. definition 3.1. we say that iλ satisfies (p.s.) condition in x , if any sequence {un} ⊂ x such that {iλ(un)} is bounded and i ′λ(un) → 0 as n → +∞, has a convergent subsequence, where (p.s.) means palais-smale. lemma 3.2. if m(t) satisfies (m ), f satisfies (f1)–(f3) and (5), then iλ satisfies (p.s.) condition. proof. suppose that {un} ⊂ x , |iλ(un)| ≤ c0 and i ′λ(un) → 0 as n → +∞. in view of (9), we have c0 ≥ iλ(un) ≥ m0 2 ‖un‖2 − λ|ω0|a. hence, {‖un‖} is bounded. without loss of generality, we assume that un ⇀ u, then i ′(un)(un− u)→ 0. therefore, we have un → u by the (s+) property of i ′λ. lemma 3.3. ifm(t) satisfies (m ), f satisfies (f1)–(f3), then there exist ρ > 0 and γ > 0 such that iλ(u) ≥ γ for every u ∈ x with ‖u‖ = ρ. proof. we define b = minx∈ω. for any u(x) ∈ x , we also define b1 = {x ∈ ω : u(x) < b}, b2 = {x ∈ ω : u(x) ≥ b}. it is well known that the embedding of x ↪→ lp(ω) is continuous when 2 < p ≤ 2∗, where 2∗ is the critical exponent. from poincaré inequality, we have b|b2| 1 p ≤ (∫ b2 up dx ) 1 p ≤ c1 (∫ b2 |∇u|2 dx ) 1 2 ≤ c1 (∫ ω |∇u|2 dx ) 1 2 = c1‖u‖, 88 ma:existence and nonexistence of positive solutions for a kirchhoff-type . . . . . . where c1 is the embedding constant of x ↪→ lp(ω). thus, we have iλ(u) = m̂ (∫ ω 1 2 |∇u|2 dx ) − λ ∫ ω f (x, u) dx ≥ m0 2 ‖u‖2 − λ ∫ b1 f (x, u) dx− λ ∫ b2 f (x, u) dx ≥ m0 2 ‖u‖2 − λ ∫ b2 f (x, u) dx ≥ m0 2 ‖u‖2 − λa2|b2| ≥ m0 2 ‖u‖2 − λa2 ( c1 b )p ‖u‖p = ‖u‖2 ( m0 2 − λa2 ( c1 b )p ‖u‖p−2 ) , where a2 = max(x,s)∈b2×[b,c] |f (x, s)|. therefore, there exist m0b p 2λa2c p 1 > ρ > 0 such that iλ(u) ≥ ρ2 ( m0 2 − λa2 ( c1 b )p ρp−2 ) := γ > 0 for every ‖u‖ = ρ and fixed λ. proof of theorem 2.1 concluded. first let us show that iλ satisfies the conditions of mountain pass theorem (see theorem 2.10 of [24]). by lemma 3.2, iλ satisfies (p.s.) condition in x . by lemma 3.3, for fixed λ > 0, there exist min { ‖u0‖, m0b p 2λa2c p 1 } > ρ > 0, γ > 0 such that iλ(u) ≥ γ > 0 for every ‖u‖ = ρ, where u0 comes from (10). on the other hand, since iλ(0) = 0 and from the proof lemma 3.1, there exists u0 ∈ x such that iλ(u0) < 0 and ‖u0‖ > ρ. so from mountain pass theorem, iλ has another critical point u2 such that iλ(u2) ≥ γ > 0 > iλ(u1). therefore, u2 is another positive solution of (1). finally we show that (1) has no positive solution when λ is small. we assume (1) has a positive solution u, let (λ1, ϕ1(x)) be the principal eigen-pair of the problem{ −∆φ = λφ in ω, u = 0 on ∂ω, (11) such that ϕ1(x) > 0 in ω. we rewrite (1) as the following form{ −∆u = λ f(x,u) m( ∫ ω 1 2 |∇u|2 dx) in ω, u = 0 on ∂ω. (12) multiplying (11) by u, multiplying (12) by ϕ1, subtracting and integrating in ω, we obtain 0 = ∫ ω [ λ1uϕ1 − λϕ1 f(x, u) m (t) ] dx = ∫ ω uϕ1 m (t) [ m (t) λ1 − λ f(x, u) u ] dx, (13) where t = ∫ ω 1 2 |∇u| 2 dx. if λ < m0λ1 β , then by remark 2.2, we have m (t) λ1 − λ f(x, u) u ≥ m0λ1 − λ f(x, u) u > m0λ1 − λβ > 0. advances in systems science and applications (2011), vol. 11, no. 1-2 89 that is contrary to (13). so for small λ, (1) has no positive solution. proof of theorem 2.2. the proof is similar to the proof of [1]. for the sake of completeness, we include it here. if there exists a positive solution (λ, u∗) for (1), then u∗ is a subsolution of { m (∫ ω 1 2 |∇u| 2 dx ) ∆u+ λf(u) = 0 in ω, u = 0 on ∂ω, (14) since m (∫ ω 1 2 |∇u∗| 2 dx ) ∆u∗ + λf(u∗) ≥ m (∫ ω 1 2 |∇u∗| 2 dx ) ∆u∗ + λf(x, u∗). and c is supersolution of (13). so by the standard comparison arguments, (13) has a positive solution u such that u∗ ≤ u ≤ c. but if we let s0 = 0, s1 = b and s2 = c, f satisfies (6) and (7), then by lemma 2.2, (13) has no positive solution. this is a contradiction. so (1) has no positive solution if ∫ c 0 f(u) du ≤ 0. 4. proof of a conjecture and some examples in this section we will prove the conjecture of liu, wang and shi’s and give some typical consequences of theorem 2.1 to theorem 2.2. in [1], liu, wang and shi conjecture that the nonexistence holds with a weaker condition:∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω. (15) in fact, as we will see in the following proposition, the condition (15) is more strong than∫ c 0 f(s) ds ≤ 0. therefore, by theorem 2.2, the conjecture is right. proposition 4.1. if f(x, u) satisfies (f1)–(f3) and ∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω, we have ∫ c 0 f(s) ds ≤ 0. proof. from (f1)–(f3), we can easily see that f(x, s) ≤ 0 when s ∈ [c(x), c]. thus, we have ∫ c c(x) f(x, s) ds ≤ 0. then, for any x ∈ ω, we have 0 ≥ ∫ c(x) 0 f(x, s) ds = ∫ c 0 f(x, s) ds− ∫ c c(x) f(x, s) ds ≥ ∫ c 0 f(x, s) ds. in particular, ∫ c 0 f(s) ds ≤ 0. now, we give some examples which satisfy our hypotheses. example 4.1. let m(t) = a + bt with t = ∫ ω 1 2 |∇u| 2 dx, here a, b are two positive constants and f(x, u) = u(u − b(x))(c(x) − u) with b(x) ∈ c(ω), c(x) ∈ c1(ω) such that 0 < b(x) < c(x) for any x ∈ ω. it is clear that m(t) and f(x, u) verify our assumptions (m) and (f1)–(f3). example 4.2. we consider a special case of example 4.1:{ ∆u+ λu(u− b(x))(c(x)− u) = 0 in ω, u = 0 on ∂ω, (16) 90 ma:existence and nonexistence of positive solutions for a kirchhoff-type . . . . . . where b(x) ∈ c(ω), c(x) ∈ c1(ω) such that 0 < b(x) < c(x) for any x ∈ ω. we have known that f(x, u) satisfies (f1)–(f3) from example 4.1. moreover, we have∫ c(x) 0 f(x, s) ds = ∫ c(x) 0 s(s− b(x))(c(x)− s) ds = 1 12 [c(x)]3(c(x)− 2b(x)). then by theorem 2.1, if there exists an open subset ω1 ⊂ ω, such that c(x) > 2b(x) in ω1, then (16) has at least two positive solutions for large λ. if c(x) ≡ 1 for all x ∈ ω, we obtain∫ 1 0 f(s) ds = ∫ 1 0 max x∈ω s(s− b(x))(1− s) ds = ∫ 1 0 max x∈ω [s2 − s3 + b(x)(s2 − s)] ds = 1 12 − b 6 , since s2 − s ≤ 0 for s ∈ [0, 1]. then by theorem 2.2, if b = minx∈ω b(x) ≥ 1 2 , then (16) has no positive solution for any λ > 0. example 4.3. let m(t) ≡ 1 and f(x, s) = s(s − 1)(c(x) − s) with 3 2 ≤ c(x) for any x ∈ ω. we can easily obtain∫ c 0 f(s) ds = ∫ c 0 max x∈ω s(s− 1)(c(x)− s) ds = ∫ c 0 (c(x)s2 − s3 + s2 − c(x)s) ds = c3 3 − c4 4 + ∫ c 0 max x∈ω c(x)(s2 − s) ds = c3 3 − c4 4 + max x∈ω c(x) ( c3 3 − c2 2 ) ds = c3 3 − c4 4 + c ( c3 3 − c2 2 ) ds = c3 12 [c− 2]. so ∫ c 0 f(s) ds ≤ 0 if and only if c ≤ 2. on the other hand,∫ c(x) 0 f(x, s) ds = ∫ c 0 s(s− 1)(c(x)− s) ds− ∫ c c(x) s(s− 1)(c(x)− s) ds ≥ ∫ c 0 s(s− 1)(c(x)− s) ds = −c 4 4 + 1 + c(x) 3 c3 − c(x) 2 c2. advances in systems science and applications (2011), vol. 11, no. 1-2 91 if ∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω, we have 0 ≥ −c 4 4 + 1 + c(x) 3 c3 − c(x) 2 c2 ⇒ 4(1 + c(x))c− 6c(x) ≤ 3c2. in particular, we have 4(1 + c)c− 6c ≤ 3c2 ⇒ c ≤ 2. however, it is clear that∫ c 0 f(s) ds ≤ 0 ; ∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω. therefore, the condition “ ∫ c(x) 0 f(x, s) ds ≤ 0 for any x ∈ ω” is more strong than the condition “ ∫ c 0 f(s) ds ≤ 0” in this example, which verifies proposition 4.1 by a concrete example. remark 4.1. in [25], dancer and yan proved when c(x) ≡ 1 and {x ∈ ω : b(x) < 1/2} is of positive measure, then (16) may have many positive solutions of local minimum type. the results of example 4.2 shows that the condition ∫ 1 0 f(s) ds ≤ 0 is optimal for the nonexistence of positive solution of (16). however, we do not know whether ∫ c 0 f(s) ds ≤ 0 is optimal for the nonexistence of positive solution of (1). references [1] g.liu, ,y. wang, & j. shi, existence and nonexistence of positive solutions of semilinear elliptic equation with inhomogeneous strong allee effect, appl. math. mech. -engl. ed. 30(11), 1461-1468 (2009) [2] g.kirchhoff, mechanik, teubner, leipzig. (1883). [3] j. l. lions, on some questions in boundary value problems of mathematical physics, in: proceedings of international symposium on continuum mechanics and partial differential equations. rio de janeiro, 1977, in: de la penha, medeiros (eds.), math. stud., vol. 30, north-holland, 1978, pp. 284-346. [4] a. arosio, & s.pannizi, on the well-posedness of the kirchhoff string. trans. amer. math. soc. 348 (1996) 305-330. [5] cavalcante, cavalcante m. m., , v.n., & j.a. soriano, global existence and uniform decay rates for the kirchhoff-carrier equation with nonlinear dissipation. adv. differential equations 6 (2001) 701-730. [6] f. j. s. a. a corrê, s.d.b.menezes, &j. ferreira, on a class of problems involving a nonlocal operator. appl. math. comput. 147 (2004) 475-489. [7] f. j. s. a. corrêa, & g.m. figueiredo, on a elliptic equation of p-kirchhoff type via variational methods. bull. austral. math. soc. vol. 74 (2006) 263-277. 92 ma:existence and nonexistence of positive solutions for a kirchhoff-type . . . . . . [8] p.d’ancona, , & s. spagnolo, global solvability for the degenerate kirchhoff equation with real analytic data. invent. math. 108 (1992) 247-262. [9] m.chipot,, & b. lovat, some remarks on nonlocal elliptic and parabolic problems. nonlinear anal. 30 (1997) 4619-4627. [10] m.dreher, the kirchhoff equation for the p-laplacian. rend. semin. mat. univ. politec. torino 64 (2006) 217-238. [11] m.dreher, the ware equation for the p-laplacian. hokkaido math. j. 36 (2007) 21-52. [12] g.ai, & r. hao, existence of solutions for a p(x)-kirchhoff-type equation. j. math. anal. appl. 359 (2009) 275-284. [13] dai, g., & liu, d. infinitely many positive solutions for a p(x)-kirchhoff-type equation. j. math. anal. appl. 359 (2009) 704-710. [14] fan, x. l. on nonlocal p(x)-laplacian dirichlet problems. nonlinear anal. 72 (2010) 3314-3323. [15] he, x., & zou, w. infinitely many positive solutions for kirchhoff-type problems. nonlinear anal. 70 (2009) 1407-1414. [16] allee, w. c. the social life of animals. w.w norton, new york(1938). [17] cantrell, r. s., & cosner, c. spatial ecology via reaction-diffusion equation. wiley series in mathematical and computational biology, john wiley sonsltd. [18] gilbarg, d., & trudinger, s. elliptic partial differential equations of second order. springer-verlag, berlin. [19] lieberman, g. m. boundary regularity for solutions of degenerate elliptic equations. nonlinear anal. tma 12 (1988) 1203-1219. [20] lieberman, g. m. the natural generalization of the natural conditions of ladyzenskaja and ural’tzeva for elliptic equations. comm. partial differential equations 16 (1991) 311361. [21] g.dai, nonsmooth version of fountain theorem and its application to a dirichlet-type differential inclusion problem. nonlinear analysis 72 (2010) 1454-1461 [22] e. n.dancer, & k. schmitt, on positive solution of semilinear elliptic equations. proc. amer. math.soc. 101 (1987), no.3, 445-452. [23] b. gidas,, w. ni, w. & nirenberg, l. symmetry and related properties via the maximum principle. comm. math. phys. 68 (1979), 209-243. [24] m. willem, minimax theorems. birkhäuser, boston. [25] e. n.dancer,, & s. yan, construction of various types of solutions for an elliptic problem. calculus variations and partial differential equations 20(1)(2004) , 93-118. microsoft word 4-li xiaogen.doc issn 1078-6236 international institute for general systems studies, inc. the key techniques study of constructing the roaming modelin virtual environment li xiaogen 1,2, wang zong-min1, wang an-ming2 ,yang xiang-bing3 1 department of water conservancy and environment , zhengzhou university, zheng zhou 450001, p.r china 2 north china institute of water conservancy and hydroelectric power,zhengzhou, 450011 p.r.china 3programme and construction commission of linzhou city,linzhou henan 456550, p.r.china email: lixiaogen@ncwu.edu.cn abstract taking the digitization of zhengzhou university campus as an example, the 3-d model of every buildings of zhengzhou university were constructed with vr technology and 3d image processing technology. the paper has studied the questions that need to be solved in the construction process of digitizing campus and in the model roaming process, this solution has been successfully used in the scientific research item of digitizing zhengzhou university. the study results can be summarized as follows:⑴the texture mapping technology make the 3d model roaming more smoother;⑵the construction technologies of many trees models make the data volume worse;⑶the 3d digitization of zhengzhou university campus provides a new idea to construct 3d virtual large scene model. keywords virtual reality 3d module the roaming virtual reality-geographical information system (vrgis) the texture mapping technology 1. introduction nowadays, virtual reality-geographical information system which was the aim of establishing digital globe in the future, was a hot spot but difficult issue in the field of geographical information system and it was a high level to implement immersed roaming in the established virtual environment. digital campus project, which was the demonstration project carried out by education ministry of henan province, has realized the virtualization of campus. in the process of constructing the campus model a few difficult questions were solved. the used ideas which were less in the references about virtual reality were a few new thoughts to construct the 3d virtual large scene which was the bigger data volume model in the future[1][2].  this study was supported by innovative prominent talents project foundation for henan universities in 2005,henan innovation project for university prominent research talents in 2005(haipurt)(2005kycx015), important science & technology foundation of henan province. this support is greatly appreciated. ∗ advances in systems science and applications (2011), vol.11, no.1-2 55-60 2. the basic principles of establishing virtual campus according to the basic principle that entity positioning in actual world is decided by both eyes of a person, virtual scenes which are composed of transverse wave and longitudinal wave are projected onto a screen with divided projectors. the transverse wave and longitudinal wave are separately projected by two groups of projectors. at this moment, beholders should wear special glasses which is made up of two layers of membrane. the two layers of glasses could filtrate transverse waves and longitudinal waves respectively[3][4]. because the structure of both glasses are contrary, they could receive a set of transverse wave and a set of longitudinal wave respectively. that is to say, people’s eyes could respectively receive a set of transverse wave and a set of longitudinal wave. so, two groups of projector suit both eyes of a person respectively. in this way, beholders could see distinct virtual geographical environment in the screen[5][6]. in this virtual environment, we adopt the creator of multigen corporation to construct the model of campus. the scenes are rendered by vega presented by multigen corporation. finally, steering wheel and remote controlling pole are used as interactive equipments to simulate roaming of driving vehicles or overlook the whole campus from the sky by air. the established virtual environment are presented as follows: fig. 1, fig. 2, fig. 3, fig. 4. fig. 1 comprehensive administration building fig. 2 students’ dormitory building on the second phase 56 li: the key techniques study of constructing the roaming modelin virtual environment fig. 3 cafeteria and rest house in the living areas fig. 4 flank greenbelt of electric engineering department on the second phase 3. a few key techniques 3.1 texture patching pictures digital camera will be used to collect surface texture pictures to show actual scenes. rgb or rgba format graphics file are recommended in vega. when exporting texture files, attention should be focused on length and width of these graphics. the pixel of length and width of the graphics texture files should be limited within power of two. the strict limitation guarantees the normal display of patching picture texture. as for large area of lawn and roof, multi-surface patching picture should be adopted. mipmap bilinear and mipmap trilinear should be adopted as the ways of filtration so as to avoid twinkle of graphics texture effectively[7]. 3.2 processing of trees as for trees along the roads, it is a difficult technique problem to display them quickly and effectively while the speed of roaming is not decreased. this problem is settled effectively with the adoption of billboard and interactive patching pictures. in addition, large quantities of trees could be produced by applying the instance technique provided by creator. trees of the same kind could share the same texture and memory is used only once, which greatly economizes resources of the system and speed up displaying on the screen[8]. 3.3 combination of building units if the number of building is small (which means the number of files which need to be combined is small),we could produce the final document by applying the methods of duplicating and paste with one of the files as dominance. meanwhile, we could integrate the advances in systems science and applications (2011), vol.11, no.1-2 57 actual scenes by copying other building models into the present documents in the way of paste, placing them in the proper position and adjusting the levels of documents in the catalogue frame work. if single building units at different stages are processed, it is not proper to apply the above stated methods, while single document is dominant. in this way, single building units are introduced with the method of external reference so that graphics texture files of building models and independence of three-dimension model file could be maintained and the actual scenes could be integrated[9]. 3.4 roaming of sight to make the users interact with the virtual environment naturally and simultaneously, we could manipulate joystick like steering wheel to roam in the campus. we could realize this aim with the application of c++ procedure and dynamic model which is based on vega. on the basis of analysis of vega structure and operational principle, we could upgrade the immersed sense of roaming system by applying directinput and controlling poll of force feedback. users could sense that they are immersed in roaming by controlling the automobiles in the virtual scenes with steering wheel manipulated. there are ten bull type trigger buttons on the input equipment which separately symbolize “forward” “backward” “turn on” “turn off”. the other two buttons are used to simulate steering wheel and accelerator. ( as shown in fig.5 and fig. 6). fig. 5 virtual roaming environment fig. 6 steering wheel 3.5 accomplishment of collision test the phenomenon of perforation should be avoided in virtual environment. tools for collision joints are not directly provided in creator, while perforation could be solved by collision test. collision test could be carried out among entities in the virtual scene where collisions probably happen, which is main characteristic of video rendering software. it could be described 58 li: the key techniques study of constructing the roaming modelin virtual environment as evaluation of intersection point in mathematics. in collision test and intersection test, vector which is composed of certain number of directed line segments will be considered. there are generally seven methods to define vector entity. these methods confine the position of the entities and determine the minimum test results. they are as follows: z, hat, tripod, los, bump, xyzpr, volume. each method has its own applicatory field. for example: the method of “z” is fit to be used in calculation of elevation. the method of “los” is fit to be used in tests of eyesight direction. if collisions happen, collision responses will be made according to the information which is stored in gathered intersection test results[10]. to carry out collision test, we intercalate anti-collision for every building by applying c++ procedure in vega environment, that is to set up a category which prevents perforation. effect of mountain climbing could be achieved by applying the method of “z” to calculate the collision between dynamic objects and ground. the method of “bump” could be applied to calculate collision between dynamic objects and buildings and trees in the virtual scene. 4. conclusion through the above study as follow the conclusion can be summarized. the texture mapping technology make the 3d model roaming more smoother, which can make the observers’ eyes more comfy. the construction technologies of many trees models make the data volume worse, which make the computer run faster. the 3d digitization of zhengzhou university campus provides a new thought to construct 3d virtual large scene model which has bigger data volume in the future. 5. future work it is not the best choice to apply multigen creator to building modeling in large-area real scene. its superiority lies in modeling of geometric objects. while the best way is to patch picture by modeling with autocad or 3dsmax software and then switch them into creator to revise. in addition, authenticity, applicability and interest could be promoted by further improvement of dynamic modeling to support coordinated roaming of individuals. references [1] li xiaode, peng bin, virtual reality technique development summarize[j], innovation forum,,10-12, 2004. advances in systems science and applications (2011), vol.11, no.1-2 59 [2] yu li, wang cheng. digital campus three dimension simulation system based on the virtual reality [j], computer simulation newspaper. 2004:99-102. [3] wang yujiang, the virtual reality discuss based on multigen creator and vega[j], mapping information and engineering newspaper, 2003:14-16. [4] zhenjun wu, jianhui deng, hong min, “ gis-based management and analysis system for landslide monitoring information” (in chinese), rock and soil mechanics, 11, p1739-1743, 2004. [5] ashis k. saha, ravi p. gupta, irene sarkar, manoj k. arora, elmar csaplovics, an approach for gis-based statistical landslide susceptibility zonation-with a case study in the himalayas, landslides 2, p61–69, 2005. [6] guirong zhang , kunlong yin, “ web database construction of geo-hazard information system in zhejiang province based on webgis“(in chinese), tile chinese journal of geological hazard and control,3, p114~118, 2005. [7] gui-rong zhang, kunlong yin , liu li-ling etc, “a real time regional geological hazard warning system in terms of webgis and rainfall “(in chinese), rock and soil mechanics,3,p1312~1317, 2005. [8] yiping wu, kunlong yin, “the landslide data management information system” (in chinese), hydrogeology and engineering geology, 1, p 14-16, 2007. [9] daan liu, xiaojia liu, “the development of geologic engineering monitoring information system” (in chinese), journal of engineering geology, 4, p351-356, 2007. [10] xinli hu, huiming tang, “ frame design of space prediction and eevaluation system for the slope geological hazard based on gis” (in chinese), geological science and technology information, 21, p99-103, 2002. 60 li: the key techniques study of constructing the roaming modelin virtual environment advances in systems science and applications (2013) vol.13 no.2 144-166 fashion erp systems and supply chain coordination wing-yan li, tsan-ming choi and pui-sze chow business division, institute of textiles and clothing, the hong kong polytechnic university, hung hom, kowloon, hong kong abstract the enterprise resource planning (erp) system is an enterprise information system designed to integrate and optimize the business processes and transactions in a corporation. an erp system enables an organization to integrate all the primary business processes in order to enhance efficiency and achieve a competitive edge. it also facilitates the smooth flow of functional information and practices across the supply chain. tal apparel limited is chosen to be the case study target and its cooperation with j.c. penney under vendor-managed inventory (vmi) scheme is also being investigated. we also investigate the two latest recently implemented modules in tal called fashion planning workbench (fpw) and supply chain order (sco), which relate closely to the company’s production planning and order arrangement. based on the case study results, and the literature review, we identify the critical successful factors in coordinating supply chains with erp systems which include: (i) assessing the fitness between organization and the target erp system, (ii) managing the cost and time of erp implementation (iii) arranging extensive education, training and rewards to employees, (iv) ensuring the internal processes are working properly and fully integrated before it is implemented in managing the supply chain, (v) building up a trusting and collaborative relationship and (vi) ensuring information visibility. finally, future research opportunities are explored. keywords erp systems, coordination, fashion planning workbench, supply chain order, fashion industry 1 introduction in fashion industry, requirements to reduce the product cost while meeting quality requirements and improving the time-to-market lead to many operational challenges and magnify competition among companies. fashion companies are hence encouraged to use advanced technologies to improve operational efficiency. the enterprise resource planning (erp) system is an enterprise information system designed to integrate and optimize the business processes and transactions in a corporation. an erp system enables an organization to integrate all the primary business processes in order to enhance efficiency and maintain a competitive position. it also facilitates the smooth flow of functional information and practices across the supply chain (sc). erp platform is specifically necessary in the fashion industry. this is because the fashion industry and its supply chains face a demand-driven market and it becomes upmost important to obtain the latest 145 advances in systems science and applications (2013) vol.13 no.2 market information and share the information among the channel members in order to facilitate the operation [1]. there is a considerable amount of literature reviewing the erp system implementation [2-21]. in this paper, we will focus on erp system implementation and its modular functionalities in a manufacturing company and how it can help achieve supply chain coordination. owing to data availability, tal apparel limited is chosen to be the case study target. the company has recently adopted two modules, namely: fashion planning workbench (fpw) and supply chain order (sco). these two latest applied modules concern about production planning and order arrangement, which are critical for fashion manufacturers to schedule their production, even the production cycle, and to be responsive to customer orders. with the case study and literature, we aim at investigating the following four research perspectives: 1. how do erp systems benefit fashion supply chain coordination and internal operations? 2. what are the possible issues that fashion companies may encounter in supply chain coordination and erp system implementation? 3. what are the critical successful factors for achieving supply chain coordination with erp system implementation? 4. what are the trends of erp system implementation and coordination? the organization of the rest of the paper is as follows. we conduct in section 2 a comprehensive review of the literature on the apparel supply chain characteristics, supply chain coordination, erp system in fashion industry and the challenges and successful factors in erp system implementation. then, we introduce the case study research methodology, review and discuss the system implementation project of tal in section 3. afterwards, we explore the above perspectives and discuss the finding from the literature and the case in section 4. finally, we have a conclusion and further research direction in section 5. 2 literature review 2.1 apparel supply chain characteristic in apparel supply chain, there are several characteristics. (i) as it is marketoriented, the brand-owners are usually the controller and have the power to influent other units in the supply chain [22]. (ii) it is a subcontracting supply chain in which the apparel retailers are usually big companies that place orders with manufacturers whereas those manufacturers make the product under a maketo-order production structure in order to lower the risk of building up obsolete inventories [23]. (iii) demand is uncertain and unpredictable. most products of the apparel industry are seasonal with short product life cycle while the demands are driven by ever-changing market trends, so it requires quick response and shorter lead-time [24] (iv) information flow is complicated as a buyer may connect wingyan li:fashion erp systems and supply chain coordination 146 to a number of manufacturers or the other way round [25], as work process in fashion industries are often highly interdependent and time sensitive. accordingly, the business value is co-created by all partners who provide different services [26]. 2.2 supply chain coordination there is a vast amount of literature about supply chain coordination. according to [27], coordination is a central lever of supply chain management. and it plays a key role in focusing on the innovation, flexibility, and speed that serve as the sources of competitive advantage necessary for survival in global competition [24,28]. simatupang and sridharan (2004)[29] state that two or more independent companies work jointly to plan to execute sc operations is with greater success than when acting in isolation. but in reality, the objectives and interests of different members along a supply chain are usually diversified and conflicting. therefore supply chain coordination becomes vitally important to achieve the all-level consensus, through which different members along a supply chain can react to market requirements in highly congruous ways [30]. under the coordination scenario in a two-echelon supply chain, achieving a win-win improvement situation can provide improved business success for both parties [31] and becomes a strategic response to the challenges that arise in the market [32]. the inter-firm coordination processes are characterized by effective communication, information exchange, partnering and performance monitoring [33]. and structures of coordination are related to the levels of information sharing and integration of physical flow from the perspective of operations management [34]. supply chain partners are working for joint planning, joint product development, mutual information exchange and integrated information systems, cross coordination on several levels in the companies on the network, long-term cooperation and fair sharing of risks and benefits [35]. lack of coordination occurs when decision makers have incomplete information or have incentives that are not compatible with system-wide objectives [36]; and it may result in poor performance of sc. the consequences of lack of coordination are: inaccurate forecasts, low capacity utilization, excessive inventory, inadequate customer service, low inventory turns, high inventory costs, slow time to market, as well as poor performance in order fulfillment response, quality, customer focus and customer satisfaction [37]. 2.3 erp systems in fashion industry in fashion industry, the product life-cycle is relatively short whilst the demand is uncertain. the emergence of fast fashion further leads to a more dynamic market with more frequently changing merchandising and purchasing decisions, and hence erp systems can play a crucial role [38]. many textiles and clothing 147 advances in systems science and applications (2013) vol.13 no.2 companies have employed the company-wide centralized erp system to improve information sharing and operations [39]. in the literature, the importance of erp system for fashion companies has been explored. au and ho (2002)[40] conduct a review study on how erp system enhances supply chain management and comment that erp system can support the fashion supply chains by allowing the implementation of e-commerce business models. erp system can also provide companies with flexibility and help solve many operational problems such as achieving more coordinated decisions on material resource planning (mrp) and capacity scheduling [41]. it allows effective planning of all resources in an organization. fin (2006)[42] investigates the relationship between electronic data interchange (edi) in apparel industry and the three performance levels: operational, financial and strategic, and finds that edi can help in reduction of lead time from several weeks to 3days. according to hodge (2002)[43], an erp system can be defined as an information system for identifying and planning the enterprise-wide resources needed to take, make, ship, and account for customer order. in general, erp systems used in textile and clothing industries are modular in structure and they can be tailored to the needs of individual companies, both large and small, with variable numbers and features of the modules. functionalities commonly covered in an erp system include finance, sales and distribution, production planning and scheduling, manufacturing and quality management, inventory management, global enterprise management, and customer relationship management[44]. for example, choi et al.(2012) [45] examine the implementation of the order-to-cash module in levi strauss & co. in chinahong kong, showing that a modular design allows erp implementation in different processes to fit the companys requirement with flexibility. 3 case study tal apparel ltd & j.c. penney 3.1 methodology in this paper, we follow the relevant literature and employ a case study methodology, supplemented by literature review, to investigate four perspectives and propose future research directions. this research focuses on a single, in-depth case study, based on the publicly available data from the company’s website, and several published articles on the target company. according to benbasat et al. (1987)[46], case study methodology examines a phenomenon in its natural setting, employing multiple methods of data collection to gather information. it provides a systematic way of looking at events, collecting data, analyzing information, and reporting the results. as a result, a sharpened understanding of why the instance happened as it did and what might become important to look at more extensively in future research is investigated and discussed [47]. this approach is deemed as wingyan li:fashion erp systems and supply chain coordination 148 appropriate for this piece of research because the topic is exploratory in nature and a number of insights and critical issues can be identified by examining the target case. 3.2 overview of target company founded in 1947 with headquarter in hong kong, tal apparel ltd (tal) is one of the world leaders in production of innovative clothes that combine style, comfort and functionality. it specializes in the manufacture of quality men’s and women’s garment for the world’s leading brands and creates one in every six dress shirts sold in the united states. (tal official website)[48] j. c. penney co.ltd. was founded in 1902 in kemmerer, wyoming by james cash penney and william henry mcmanus. today j.c. penney offers a range of family apparel, jewelry, shoes, accessories, and home furnishing products through a chain of department stores and its company website. headquartered in plano, tx, the company operates in the united states and puerto rico, with a total of 1102 department stores in 49 states. it receives merchandise from over 2,500 suppliers, both domestic and foreign. the company also operates buying and quality assurance inspection offices in 15 countries. (j.c. penney annual report 2011)[49]. 3.3 erp implementation in tal in 1990s, tal was one of those first companies that use a shop-floor system to manage apparel bundles on the factory floor. each bundle has a barcode and whenever a sewer finishes an operation, she scans the bundle and it’s tracked with the shop-floor systems. the system, named the apparel manufacturing management information systems (ammis), is still in use, though its database has been upgraded [50]. tal is known for using information technology to improve shop floor efficiency and process flow to differentiate itself from the competitors. lawson m3 enterprise management systems implementation in early 2002, tal implemented lawson m3 enterprise management systems (m3 stands for make, move and maintain) to manage its internal business process. tal invested us$ 12 to 15 million to develop and perfect the system in around 5 years starting from 2000 [51]. this system is provided by lawson software which is a global provider of enterprise software, services and support to customers primarily in three sectors: services, trade and manufacturing/distribution. lawson’s solutions include enterprise performance management, human capital management, supply chain management, enterprise resource planning, customer relationship management, manufacturing resource planning, enterprise asset management and industry-tailored application [52]. tal implemented m3 enterprise management systems with the following five modules: customer sales & service, enterprise asset management, supply chain management, manufacturing operation and finance management. table 1 summarizes their corre149 advances in systems science and applications (2013) vol.13 no.2 sponding functions and benefits (adapted from the infor official website). tal has been using the lawson m3 enterprise software for supply chain management, order management, capacity planning, production planning, inventory management, procurement, vendor-managed inventory, material requirements planning, and finance. then tal spent two years with lawson’s research and development team to co-develop two more modules in 2006, which are: fashion planning workbench (fpw) and supply chain order (sco). fashion planning workbench (fpw) in the past, tal used excel spreadsheets to plan the mediumand long-term capacity for individual factories, which took hours to download data from the erp system and is susceptible to human error. now the workbench tool allows executives in the company to see order capacity across the group and handle multi site planning in a more efficient manner. it takes only 30 minutes and that’s with 200,000 manufacturing orders in the system the company has 10 factories located in seven countries, and it can swap short orders from one factory to another one according to individual factory loads. with the workbench tool, tal can arrange six-month production planning [53]. supply chain order (sco) in traditional material requirements planning (mrp), demands from multiple customer orders were satisfied at an aggregate level. it did not tell which order was specifying the demand, and it lacked visibility. when tal needed to make changes in a customer order like size breakdown, it had to go though three more steps to make sure the changes were done. however in tal’s business where most things are made-to-order, materials are highly specific to particular orders, and often need to be uniquely linked, rather than grouped together as a whole, so tracking and having visibility to customers level is crucial. the sco module improves the way tal plans and manages order chains by linking orders to supply chain data. specifically, if tal changes the quantity of a customer’s order, all other related material requirement orders are taken care of. tal implemented sco at 10 of tal’s production sites, which help to save 80 percent of the time spent for revising related orders already in the system, while reducing inventory levels by approximately 30 percent. in short, fpw and sco help tal to smooth the production workflow and even out the production cycle, and a flat production cycle results in $3 millions saving a year. [54-55]. wingyan li:fashion erp systems and supply chain coordination 150 table 1 m3 enterprise management systems (adapted from infor official website)[56] modules functions and benefits customer sales & service • provides support for market development, sales and after-sales service processes • integrates web, mobile and traditional order management tools • supports sales channels using the same availability, pricing, and fulfillment controls enterprise asset management • integrates with other operational, planning and execution applications • synchronizes maintenance with production plan • increases asset reliability and improves operational capacity and output • enhances scheduling • reduces the need for emergency repair and maintenance cost supply chain management • controls and optimizes information, material and financial flows • drives efficiencies from planning to procurement, from warehouse to work-in-progress • provides functionality based on open standards • helps to capitalize on new mobile and data collection manufacturing operations • simplifies complex challenges with the capability to manage traceability, change control, and enhance manufacturing environments • provides a single source of product data and improves information availability • brings control and transparency to the manufacturing processes finance management • records operational transactions from other modules • handles accounting, budgeting, and helps consolidate all required reports and documents • coordinates multiple units in terms of the financial requirements • detects critical trends effectively and in a timely manner continuous upgrade of the erp system in tal afterwards, tal continued to launch its system upgrade in hong kong, mainland china, thailand, malaysia, indonesia and vietnam in 2009 and the upgrade project was complete in january 2010, on time and on budget. the upgraded system helps tal improve productivity across its global operations, by helping users personalize their information workspace and bring together enterprise applications, business intelligence, desktop tools, and group collaboration. thus it helps tal to work in a more productive and dynamic manner [52]. 151 advances in systems science and applications (2013) vol.13 no.2 3.4 vmi scheme implementation between tal and j.c. penney traditionally, information systems are designed such that they are not accessible to exterior companies. there are organizational boundaries isolated a company’s system from the vendors and customers. useful information like sales data, inventory level and order status cannot reach it partners along the supply chain effectively. this would result in longer lead time and lower flexibility. under vmi scheme with tal, j.c. penney uses edi over a third-party network. there is both centralized and non-centralized ordering, the stores place orders directly to supplier, and many items are handled through automated replenishment system. edi can also be used to transmit pos-data. tal collects point-of-sales data of j.c. penney’s 1,102 stores. the data is fed into tal’s demand forecasting model yielding model stock of stock keeping units (skus) for individual stores. tal then decides how many shirts to be made in which color and size[57]. this entire program is designed and operated by tal. the model stock and sku information is now dynamically linked to tal’s factory floor, so the factories promptly know the detailed orders and then deliver the shirts directly to individual stores, using the technology of x-docking. although forecasts are never completely accurate, tal’s collaborative information and fast manufacturing capabilities enable the company to react quickly to changes. if necessary, tal can yield shirts ready for shipment in just four hours. according to [58], j.c. jenney is doing something beyond vmi called supplier managed availability (sma), which is an extension of vmi. there are no minimum or maximum inventory level requirements placed on the supplier whereas the supplier evaluates solely on product availability in the store when the product is needed. depending on the urgency of the need at a particular store, the supplier may choose to ship the product much more rapidly by air (express).tal’s coordinated power makes the chain operating in a more efficient mode. under vmi scheme, both the manufacturer and the retailer are benefited. j.c. penney can respond more quickly to consumer demand and it has shown sales increases by 20-25% and improvement in inventory turnover by 30% [59]. this form of collaboration completely eliminates warehouse inventory thus bypassing warehouse handling, improved fill rates, and increased customer satisfaction. the direct-ship processes have helped j.c. penney save roughly $30 million of the average monthly inventory expenses [60]. on the other hand, the main advantage of the vmi scheme for the manufacturer is the mitigation of demand uncertainty; it monitors the pos data and determines the most cost-effective way to meet the projected availability requirements, enhancing tal’s scheduling and planning, allowing a specified service level to be maintained at minimal inventory and production cost. tal claims that the vmi program can result in costs savings of up to 15% through reduced wingyan li:fashion erp systems and supply chain coordination 152 inventory and operational costs. (tal official website)[48]. close collaboration with j.c. jenney allows tal to design products for it. when tal’s design teams in new york and dallas come up with a new style, within a month its factories can churn out 100,000 new shirts and offer for sale in 50 penney stores [36]. as tal participates in the whole process including design, production, marketing and distribution, it can be more flexible and accurate to response to the consumer market, resulting in greater sales and fewer closeouts. 4 discussions 4.1 how do erp systems benefit fashion supply chain coordination and internal operation? allowing information accessibility among supply chain partners erp system allows integration and coordination, supports information flow and activities both within a company and along supply chain, provides reliable services to make information flow smoothly amongst all partners and ensures that any partner can get information timely and accurately when in need [25]. under the vendor managed inventory scheme between tal and j.c. jenney, tal (the manufacturer) can access to j.c. penney (the retailer)’s system to obtain point-of-sale data so that tal can have a clearer picture about the real market demand. this successfully reduces the forrester’s bullwhip effect and results in just-in-time inventory and delivery, quick response, collaborative planning and forecasting in their supply chain. also the two members maintain the updated single forecasting process in the system by sharing information related to future demand, resulting in substantial cost reduction. developing a successful synchronization strategy in the case of tal and j.c. penney, sophisticated skills in the deployment and use of technologies such as edi, advanced planning and scheduling (aps), forecasting software, and enterprise resource planning (erp) help the manufacturer in synchronizing with its retailers for information sharing, inventory and logistics management. information sharing helps supply chain partners build up trust and responsibility interdependence. the stronger the responsibility interdependence between supply chain partners, the more they need to become involved in sharing private information, learning how to open another’s viewpoint, as well as resolving shared problems [29]. thus, it develops a differential compatible advantage and creates higher switching costs for retailers and substantial entry barriers for competitors. in addition, most commodity manufacturers do not have the sophistication in a complex partnership between supplier and customer [61]. identifying mediumto long-term production planning tal has been using the lawson m3 enterprise management system in its factories in seven countries to help minimize production peaks and valleys. to achieve 153 advances in systems science and applications (2013) vol.13 no.2 that, it needs to match production capacities across all factories with orders that it expects six to nine months in advance. tal balances production load across its multiple sites and assists with planning at the style/color level. this gives tal the ability to better plan capacity across the company’s various production sites and simulate load based on possible operational and business changes [53]. thus, tal is better positioned to manage seasonal peaks and valleys for production. achieving a higher visibility in material requirements planning erp system helps tal plan and manage order chains by linking orders with supply chain data. this helps the company managing issues related to changes along the supply chain as demand and supply information are all in one single view. this allows improvement of delivery accuracy, reduction in inventories, and manual work elimination. with more immediate access to comprehensive integrated business information, tal has cut its reporting time by 50 percent. besides, improved demand management with relevant information consolidated to a single database also allows the company for better customer service using real-time information, and increase customer responsiveness [50]. 4.2 what are the possible issues that fashion companies may encounter in supply chain coordination and erp system implementation? a trade-off between autonomy and control collaboration implies visibility of internal activities and metrics by external parties. the adoption of an integrated approach throughout the supply chain requires a trade-off between autonomy and control between each supply partner relationship [62]. in the case of tal and j.c. penney, j.c. penney shares packing, shipping, inventory and product movement with tal and let tal control its inventory level. it is actually giving an important function to supply chain partners when outsourcing the inventory management. many companies in virtual integration may not be willing to allow partners to view their systems and processes as their ability to perform will become more transparent, which may create pressure on themselves, especially for small companies. therefore, this becomes an organizational challenge to reach an acceptable balance among autonomy and control between supply chain partners. conflict in inter-organization relationships with effective coordination, partners should work together to reach a mutual goal. however, such approach will be difficult for many firms to take on board as they have focused on squeezing margins with their suppliers rather than cooperating with them [63]. conflict in inter-organization relationships refers to the disagreements that occur in cooperation relationship or the incompatibility of activities, shared resources and goals between partners [64]. and this is impediment to effective inter-organization information sharing. if the supply chain partners do not cooperate with each others, this would affect the accuracy of wingyan li:fashion erp systems and supply chain coordination 154 information being shared, for example, the required lead time provided by the manufacturers and the expected due date by the retailers. power asymmetry according to hingley (2005)[65], inter-organizational relationships emphasize the necessity for symmetry and mutuality and that symmetric dependence structures foster longer-term relationships while asymmetric relationships are associated with less stability and more conflict. under vmi, we can observe a significant shift in power from the retailer to the manufacturer since it is now the manufacturer who manages the inventory. in the case of tal and j.c. penney, they are giant manufacturer and retailer respectively with high bargaining power; this leads them to become inter-dependant and work together under fair terms. tal is willing to design and operate entire program for vmi and j.c. penney is willing to share pos data and design work with tal. but for other manufacturers in the apparel industry that are small or medium in scale and lack resources and bargaining power in the supply chain, the cost involved in implementing new it systems is high. those small manufacturers will be worried about an unequal distribution of costs and benefits between the various members of the supply chain [26]. these will obviously hinder smaller firms to share their information. operation reengineering shifts the management power implementing an erp system may force the reengineering of key business processes or developing new business processes to support the organizations goals [66]. and redesigned processes require corresponding realignment in organizational control to sustain the effectiveness of the reengineering efforts. and this kind of alignment may change the functional areas and many social systems within the organization. the resulting changes may significantly affect organizational structures, policies, processes, and employees. according to burca (2005)[63], implementation of erp system changes the career prospects and aspirations of many people in the company. power was perceived to move strongly in the direction of people who had embraced the it agenda. as a result, management may feel that there is a change, and want to resist it. this would definitely affect erp implementation. 4.3 what are the critical successful factors for achieving supply chain coordination with erp system implementation? assessing the fit between organization and the target erp system according to swan et al. (1999)[67], organizational misfits of erp exist due to the conflicting interests of the user organization and the erp vendor. in order to have successful erp implementation, erp implementation managers as well as top management of the organization should be able to assess the fit between their organization and the target erp system before its adoption, identify gaps between the erp generic functionality and the specific organizational require155 advances in systems science and applications (2013) vol.13 no.2 ment, and then decide how these gaps will be handled [68]. after erp system implementation, the impact of erp and process adaptations should be well measured and managed so as to minimize the potential business disruptions and user resistance. in the case of tal’s erp implementation, tal spent two years with lawson’s research and development team co-developing new erp modules (the fashion planning workbench and supply chain order systems) to ensure the modules fit tal and work efficiently. managing the cost and time of erp implementation erp systems come in modular fashion and do not have to be implemented entirely at once as the length and cost of implementation is affected to a great extent by the number of modules being implemented and the scope of the implementation [69]. the adopting company is suggested to follow a phase-in approach in which few modules are implemented at a time. besides, the implementation costs increase with the degree of customization. therefore, the company should have a thoughtful plan in order to balance customization and cost, keep the project plan aggressive, but achievable and flexible. in the case of tal’s erp system implementation, tal first adopted 5 modules in 2002; after the system had been run smoothly for few years, it added two more new modules in 2006 and eventually upgraded the erp systems in 2009 on time and on budget. gradual implementation and modification of erp system can also lower user resistance and the risk that may encounter. arranging extensive education, training and rewards to employees education and training is probably the most widely recognized critical success factor for erp system implementation, because user understanding is essential. erp implementation requires a critical mass of knowledge to enable people to solve problems within the framework of the system. if employees do not understand how the system works, they will invent their own processes using those parts of the system they are able to manipulate [70]. top management must allocate adequate expense on education and end-user training and incorporate such expense as part of the erp budget. according to volwer (1999)[71], it is suggested that reserving 10-15% of the total erp implementation budget for training will give an organization an 80% chance of implementation success. once the selected employees are trained after investing a huge sum of money, it is a challenge to retain them; especially the market is seeking for skilled erp system consultants. employees could double or triple their salaries by accepting other positions [69]. it is recommended the company should provide attractive retention strategies such as bonus programs, company perks, salary increases, continual training and education, and appeals to company loyalty. ensuring internal processes work properly under vmi, the manufacturer receives sales data from the retailer via edi and wingyan li:fashion erp systems and supply chain coordination 156 the information will be transmitted to the manufacturer’s erp system for inventory planning, scheduling and production. the manufacturer should ensure its internal processes integrated properly so that the orders can link up with other supply chain data. norris et al. (2001)[72] have stressed that before an organization can participate in e-supply chain management (e-scm) it needs to ensure that its internal processes are fully integrated. the assertion is that processes that span the supply chain cannot work if the internal processes do not work correctly in the first place. building up trusting and collaborative relationships all parties involved must view collaboration as a strategic asset and an operational priority in order to foster trust among trading partners. organizations wishing to extend their processes will have to develop more trusting and collaborative relationships with their business partners. all parties need to recognize that success for one part of the supply chain means success for all and all improvements made to the operation of the supply chain will ultimately benefit all member firms [73]. smooth communication is the foundation for supply chain management. frequent communication contributes to faster problem resolution, trust, and relationship building as well as quicker decision-making and those are resulted from having access to up-to-date information. moreover, willingness to share information enhances the quality and relevance of the information that is shared. for example, sharing actual sales data combined with rolling forecasts helps improving mutual supply chain decision making. in addition, willingness to share future product strategies and technology plans rather than simply sharing forecasting data allows more cooperation and integration. ensuring information visibility good supply chain management is essential for a successful company. supply chain management aims at going beyond the boundary of a single company. this also applies the corresponding use of information. in order to have efficient and effective supply chain operations, information visibility, such as those on inventory level and transportation schedule, is critical. regarding information visibility across the supply chain, it should be managed with strict policies, disciplines and monitoring [72]. allowing all partners in the supply chain to view and dynamically manage both demand and capacity data raises opportunities for the simultaneous improvement in customer service levels and reduction in overall inventory levels and associated costs [74]. partners in virtual integration need to be willing to share information with each other in order for the end-to-end process to work correctly and organizations also need to understand the implications of integration across the entire supply chain [75]. according to kelle and akbulut (2005)[76], if one of the supply chain partners forces its optimal policy on the 157 advances in systems science and applications (2013) vol.13 no.2 other partner, the total operating cost of the system can be much higher than the case with a coordinated ordering/setup and shipment policy. so the potential for cost reduction in policy coordination motivates sharing of information, too. 4.4 what are the trends of erp system implementation and coordination? intensive cooperation between the organization and its erp vendor competing in a dynamic environment and meeting global challenges requires agility. successful companies must be able to respond quickly and cost-effectively to change. the change could be of any type, for example shifting in customer demands and supply chain partners, modifications of a business model, business process or business expansion [77]. organizations need to convert their business into responsive, demand-driven, profit-making enterprises by optimizing their operations, so a company with erp system would be more active in cooperating with its erp vendor to create new modules that can increase operational effectiveness and market responsiveness. this phenomenon also appears in the case of tal, where tal actually teams up with its erp vendor to co-create erp modules for effective operation and coordination. the need of extended information system applications and technology from the case of tal we can see the erp system should be monitored and upgraded continually to enhance its functions and efficiency, so as to increase the company’s competitiveness. according to sleeper (2004)[78], only 35% of the organizations are satisfied with the erp they use at the moment, and the main reason for dissatisfaction is that the software does not map well with the business goals. a major problem with existing erps is the “misfit” between delivered functionality and needed functionality, described as a gap between the processes the erp supports and the processes the organizations work by. thus, there is an increasing interest among vendors to improve future erp-systems [77], and further look into how to identify and present business requirements for the future erp. this creates the need for extended information system applications and technology (e.g. erp ii, soa, web 2.0 or software as a service saas, etc.). staff (2006)[79] describes mysap as the next generation of erps, and says that this software is fast, flexible, and an efficient foundation offering organizations new functions, greater productivity, and integrated analytical insights into business processes which increase the flexibility to design the software so that it meets customer requirements. this is described as role-specific development aiming at increasing employee productivity by supporting them with a tool tailored to their specific work tasks. amongst companies that are relatively satisfied with their existing erp operations, many of them are now considering extension of the functionalities provided by the original erp systems (i.e. erp ii) toward e-business, supply chain management (scm), customer relationship management, supplier relationship wingyan li:fashion erp systems and supply chain coordination 158 management, business intelligence, and manufacturing execution systems, etc. according to tenkorang et al. (2011)[80], erp ii aggregates and manages the data surrounding all the transactions of an enterprise as accurately as possible in real time. it also facilitates opening up of the system to make information available to trading partners in the supply chain. 5 conclusion and future research opportunities through the study of the talj.c. penney case and the reviewed literature, we have demonstrated how the erp system can improve business performance and organization in tal and how the system helps coordination between the retailer and the manufacturer. erp allows useful information to be integrated and flowed effectively within the organization and across the supply chain, which is important to uncertain and demand-driven fashion industry. the latest implemented modules – fashion planning workbench (fpw) and supply chain order (sco) – help tal to smooth the production workflow and even out the production cycle. besides, information sharing between tal and j.c. penney benefit in (i) maintaining optimal stock level, (ii) reducing inventory and operational cost, (iii) enhancing manufacturers’ scheduling and planning and (iv) building up close relationship between manufacturer and retailer. combining the findings from the case study and the literature review, below is the summary of what we have discussed in this paper. 5.1 the benefit of erp systems in fashion supply chain coordination and internal operations four major benefits in fashion supply chain coordination and internal operation are identified: (i) it allows information accessible among supply chain partners, providing reliable services to make information flow smoothly between all partners and get information from partners timely and accurately when partner needs. (ii) it develops a successful synchronization strategy, creating a differential compatible advantage and creates higher switching costs for supply chain partners. (iii) it helps the manufacturers to identify mediumto long-term production planning. (iv) it helps achieving a higher visibility in material requirements planning by linking orders with supply chain data. 5.2 the possible issues that fashion companies may encounter in supply chain coordination and erp system implementation from the case of tal and j.c. penney, we reveal the possible issues in supply chain coordination and erp system implementation, namely: (i) the trade-off between autonomy and control, (ii) conflict in inter-organization relationships, (iii) power asymmetry between the supply chain partners, and (iv) the management power shifted though operation reengineering. 159 advances in systems science and applications (2013) vol.13 no.2 5.3 the critical successful factors for achieving supply chain coordination with erp system implementation in order to have a successful erp systems implementation and make sure it can work effectively in supply chain coordination, the following six factors should be considered: the company should (i) assess the fitness between organization and the target erp system, (ii) manage the cost and time of erp implementation, (iii) arrange extensive education, training and rewards to employees, (iv) ensure the internal processes are working properly and fully integrated before it spans the supply chain. (v) the supply chain partners should build up a trusting and collaborative relationship and have better communication. (vi) supply chain participants should ensure information visibility, allowing all partners in the supply chain to view and dynamically manage both demand and capacity data. 5.4 the trends of erp system implementation and coordination according to the case study and the literature review, we have observed two trends of erp system implementation and coordination. (i) intense cooperation between organization and its erp vendors: companies with erp system installed would be more active in cooperating with their erp vendors to create new modules that can increase operational effectiveness and market responsiveness. (ii) the need for extended information system applications and technology: there are many companies which have implemented erp systems and are relatively satisfied with their operations are now considering the extension of the functionalities provided by the original erp systems. 5.5 the challenges and further research opportunity in erp system implementation mismatch between existing and new erp system as discussed in section 4, there are some new versions of erp systems developed in order to handle the problem of mis-fit between erp system and organizational functionality. however according to kumar and van hillegerberg (2000)[12], this creates new problems in the migration between different versions, and either it could be that the new version is not backward compatible or it could be that the customer organization has made modifications of the erp, and these modifications do not automatically adjust to the new version of the erp. thus, there is a need to examine further how companies should extend their erp systems and the factors that should be considered when extending erp systems. small and medium-sized firms hardly develop a costly erp system companies that invest in technology can improve overall efficiency and effectiveness and have a competitive advantage. but the implementation cost of those technology systems is high, only large companies are able to invest in these big projects to improve their performance whilst small and medium-sized firms hardly develop such a system on their own. besides, not many small companies are taking a proactive approach to implementing erp or even sharing information. wingyan li:fashion erp systems and supply chain coordination 160 the usual reason is they tend to rely on the industrial giants to offer help (e.g. teaming up with big suppliers such as tal can help enhance their own supply chains). then it may come to another problem mentioned by burca (2005)[63]: those small companies may have a distrust of passing confidential information to customers and suppliers as this may indicate how well they are performing. further study in this area may indicate how small and medium-sized firms can put practices in place to ensure that collaboration takes place. references [1] fernie, j. (1994), “quick response-an international perspective”, international journal of physical distribution and logistics management, vol.24, no.6, pp.38-46. [2] ehie, i.c, madsen, m. (2005), “identifying critical issues in enterprise resource planning (erp) implementation”, computers in industry, vol.56, no.6, pp.545-57. [3] moon, y.b. (2007), “enterprise resource planning (erp): a review of the literature”, international journal of management and enterprise development, vol.4, no.3, pp.235-264. [4] bhatti, t.r. (2005), “critical success factors for the implementation of enterprise resource planning (erp): empirical validation”, the 2nd international conference on innovation in information technology. decision sciences, vol.33, no.4, pp.505-536. [5] escalle, c.x, cotteleer, m.j, austin, r.d. (1999), enterprise resource planning (erp): technology note, harvard business school publishing, boston, ma. [6] trunick, p.a. (1999), “erp: promise or pipe dream?”, transportation & distribution, vol.40, no.1, pp.23-26. [7] zhang, l, lee, m, zhang, z, banerjee, p. (2003), “critical success factors of enterprise resource planning aystems implementation success in china”, proceedings of the 36th hawaii international conference on system sciences, pp.236-245. [8] kennerley, s, neely, b. (2001), “enterprise resource planning: analysing the impact”, integrated manufacturing systems, vol.12, no.2, pp.103-13. [9] kohli, a. k, jaworski, b.j. (1990), “market orientation: the construct, research propositions, and managerial implications”, journal of marketing, vol.54, pp.1-18. 161 advances in systems science and applications (2013) vol.13 no.2 [10] kraemmerand, p, moller, c, boer, h. (2003), “erp implementation: an integrated process of radical change and continuous learning”, production planning & control, vol.14, pp.338-348. [11] kumar ,v, maheshwari, b, kumar, u. (2002), “enterprise resource planning systems adopting process: a survey of canadian organizations”, international journal of production research, vol.40, no.3, pp.509-523. [12] kumar, k, van hillegersberg, j. (2000), “erp experiences and evolution”, communications of the acm, vol.43, no.4, pp.22-26. [13] livermore c, ragowsky a. (2002), “erp systems selection and implementation: a cross cultural approach”, proceedings of the eighth americas conference on information systems, pp.1333-1339. [14] luo, w, strong, d. m. (2004), “a framework for evaluating erp implementation choices”, ieee transactions on engineering management, vol.51, no.3, pp.322-333. [15] menon, a, bharadwaj, s. g, howell, r. (1996), “the quality and effectiveness of marketing strategy: effects of functional and dysfunctional conflict in intra-organizational relationships”, journal of academic marketing science, vol.24, pp.299-313. [16] nah, f, lau, j, kuang, j. (2002), “critical factors for successful implementation of enterprise systems”, business process management journal, vol.7, pp.285-296. [17] narver, j. c, slater, s. f. (1990), “the effect of a market orientation on business profitability”, journal of marketing, vol.54, no.4, pp.20-35 [18] ngai, e.w.t. law, c.c.h, wat, f.k.t. (2008), “examining the critical success factors in the adoption of enterprise resource planning”, computer in industry, vol.59, pp.548-564. [19] rolland, c, prakash, n. (2000), “bridging the gap between organizational needs and erp functionality”, requirements engineering, vol.5, no.3, pp.180-193 [20] scott, j, vessey, i. (2002), implementing enterprise resource planning systems: the role of learning from failure, information systems frontier, vol.2, pp.213-232. [21] sheu, c, chae, b, yang, c. (2004), national differences and erp implementation: issues and challenges, omega, vol.32, pp.361-371. wingyan li:fashion erp systems and supply chain coordination 162 [22] abecassis-moedas, c. (2006), “integrating design and retail in the clothing value chain: an empirical study of the organization of design”, international journal of operations & production management, vol.26, no.4, pp.412-28. [23] mather, c. (2004), garment industry supply chain, women working worldwide: manchester. [24] fisher, m.l. (1997), “what is the right supply chain for your product?”, harvard business review. harvard business school publication corp, vol.75, pp.105-116. [25] ma, b, zhang, k. j. (2009), “research of apparel supply chain management service platform”, in management and service science, mass’09, ieee international conference on, pp.1-4. [26] cherbakov, l. (2005), “impact of service orientation at the business level”, ibm systems journal, vol.44, no.4, pp.653-68. [27] ballou, r.h, gilbert, s.m, mukherjee, a. (2000), “new managerial challenges from supply chain opportunities”, industrial marketing management, vol.29, no.1, pp.7-18. [28] lee, h. and whang, s. (2002), “the impact of the secondary market on the supply chain”, management science, vol.48, no.6, pp.719-731. [29] simatupang, t.m. and sridharan, r. (2004), “the collaborative supply chain”, international journal of logistics management, vol.3, no.1, pp.1530. [30] collins. (1995), collins cobuild english dictionary, harpercollins publishers, london. [31] mcclellan, m. (2003), collaborative manufacturing, st lucie press, delray beach, fl. [32] xu, l, beamon, b. (2006), “supply chain coordination and cooperation mechanisms: an attribute-based approach”, the journal of supply chain management, vol.42, no.1, pp.4-12. [33] stank, t.p, crum, m.r, arango, m. (1999), “benefits of interfirm coordination in food industryin supply chains”, journal of business logistics, vol.20, no.2, pp.21-41. [34] sahin, f, robinson, p. (2002), “flow coordination and information sharing in supply chains: review, implications and directions for future research”, decision science, vol.33, no.4, pp.505-536. 163 advances in systems science and applications (2013) vol.13 no.2 [35] larsen, s.t. (2000), “european logistics beyond 2000”, international journal of physical distribution and logistics management, vol.30, no.6, pp.377387. [36] cao, n, zhang, z, to, k.m, ng, k.p. (2008), “how are supply chains coordinated?: an empirical observation in textile-apparel businesses”, journal of fashion marketing and management, vol.12, no.3, pp.384-97. [37] ramdas, k, spekman, r.e. (2000), chain or shackles: understanding what drives supplycchain performance, interfaces, vol.30, no.4, pp.3-21. [38] bruce, m, daly, l. (2006), “buyer behavior for fast fashion”, j fashion market manaegment, vol.10, pp.329-344. [39] callaway, e. (1999), “enterprise resource planning: integrating applications and business processes across the enterprise”, computer technology research corporation charleston, cs. [40] au, k.f, ho, c.k. (2002), “electronic commerce and supply chain management: value-adding service for clothing manufacturers”, j integrated manufacture system, vol.13, pp.247-254. [41] doshi, g. (2006), information technology and textile industry, available at http://www.fiber2fashion.com, retrieved on 5 nov 2012, [42] fin, b. (2006), “performance implications of information technology implementation in an apparel supply chain”, supply chain management: an international journal, vol.11, no.4, pp.309-316. [43] hodge, g. l. (2002), “enterprise resource planning in textiles”, journal of textile apparel tech management, vol.2, no.3, pp.1-8. [44] hui, c. l, tse, k. choi, t. m, liu, n. (2010), “enterprise resource planning systems for the textiles and clothing ”, innovative quick response programs in logistics and supply chain management, (cheng, t.c.e., choi, t.m. eds.), springer, pp.279-296. [45] choi, t. m, chow, p. s, liu, s. c. (2012), “implementation of fashion erp systems in china: case study, review and future challenges”, international journal of production economics, doi: 10.1016/j.ijpe.2012.12.004 , in print. [46] benbasat, i, goldstein, d.k, mead, m. (1987), “the case research strategy in studies of information systems”, mis quarterly, vol.11, no.3, pp.369-386 wingyan li:fashion erp systems and supply chain coordination 164 [47] flyvbjerg, b. (2006), “five misunderstandings about case study research”, qualitative inquiry, vol.12, no.2, pp.219-245. [48] tal group official website: http://www.talgroup.com/en/index.html, retrieved on 11 nov 2012. [49] j.c.penney annual report. (2011), j.c. penney official website: http://www.jcpenney.com/dotcom/index.jsp, retrieved on 11 nov 2012. [50] thilmany, j. (2007), tal at 60: faster than ever, apparel magazine. [51] hirsch s. (mar 2005), e-management-suppliers become closer partners, available at http://www.intracen.org, retrieved on 12 nov 2012. [52] hanson, b. (2010), apparel manufacturing leader live on lawson solutions to improve production, available at http://www.whichplm.com/news/, retrieved on 12 nov 2012. [53] lawson software. (2007), “tal’s quest for manufacturing nirvana: adding the newest m3 modules”, lawson fashion newsletter. [54] lawson software. (2007), lawson m3 fashion helps tal balance production load, available at http://www.fibre2fashion.com, retrieved on 23 nov 2012. [55] lawson official website: http://www.lawson.com/solutions/, retrieved on 11 nov 2012. [56] infor official website: http://www.infor.com, retrieved on 10 dec 2012. [57] umble, e. j, haft, r. r, umble, m. m. (2003), “enterprise resource planning: implementation procedures and critical success factors”, european journal of operational research, vol.146, no.2, pp.241-57. [58] hausman, w. (2003), “supplier managed availability”, wall street journal. [59] buzzell, r.d, ortmeyer, g. (1995), “channel partnerships streamline distribution”, sloan management review, vol.36, no.3, pp.85-96. [60] atkinson, w. , j.c. penney. (2006): “pioneer of supply chain efficiency”, apparel magazine,(april 2006). [61] koudal, p, long, w.t. (2005), “the power of synchronization: the case of tal apparel group”, deloitte research. 165 advances in systems science and applications (2013) vol.13 no.2 [62] graham, g, hardier, g. (2000), “supply chain management across the internet”, international journal of physical distribution & logistics management, vol.30, pp.286-95. [63] burca, s, fynes, b.marshall, d. (2005), “strategic technology adoption: extending erp across the supply chain”, journal of enterprise information management, vol.18, no.40, pp.427-40. [64] anderson, j.c, narus, j.a. (1990), “a model of distributor firm and manufacturer firm working partnerships”, journal of marketing, vol.54 no.1, pp.42-58. [65] hingley, m. k. (2005), “power to all our friends? living with imbalance in suppliercretailer relationships”, industrial marketing management, vol.34, no.8, pp.848-58. [66] minahan, t. (1998), enterprise resource planning, purchasing, vol.16, pp.112-117 [67] swan, j, newell, s, robertson, m. (1999), “the illusion of ‘best practice’ in information systems for operations management”, european journal of information systems, vol.8, no.4, pp.284-293. [68] soh, c, sia, s.k, tay-uap, j. (2002), “cultural fits and misfits: is erp a universal solution”, communications of the acm, vol.43, no.4. pp.47-51. [69] bingi, p, sharma, m.k, godla, j.k. (2006), “critical issues affecting an erp implementation”, information systems management, vol.16, no.3, pp.7-14. [70] ptak, c, schragenheim, e. (2000), erp: tools, techniques and applications for integrating the supply chain, the st, lucie press series on resource management. [71] volkoff, o, sterling, b, nelson, p. (1999), “getting your money worth from an enterprise system”, ivey business journal, vol.64, no.1, pp.54-57. [72] norris, g, hurley, j.r, hartley, k.m. dunleavy, j.r. (2001), e-business and erp: transforming the enterprise, published by john wiley & sons.inc. [73] scalet, s.d. (2001), “the cost of secrecy”, cio magazine; july 2001, available at: www.cio.com. [74] kehoe, d, boughton, n. (2001), “paradigms in planning and control across manufacturing supply chains-the utilization of internet technologies”, international journal of operations & production management, vol.21, no.4, pp.516-525. wingyan li:fashion erp systems and supply chain coordination 166 [75] venkatraman, n, henderson, j.c. (1998),“ real strategies for virtual organizing”, sloan management review, vol.44, no.1, pp.33-48. [76] kelle, p, akbulut, a. (2005), “the role of erp tools in the supply chain information sharing, cooperation, and cost optimization”, international journal of production economics, vol.93, pp.41-52. [77] johansson, b. (2009), “why focus on roles when developing future erp systems?”, in barry, c, conboy, k, lang m, wojtkowski, g and wojtkowski, w. (eds.), information systems development: challenges in practice, theory, and education, springer, vol.1, pp.547-60. [78] sleeper, s.z. (2004). “amr analysts discuss role-based erp interfaces the user-friendly enterprise”, sap design guild. available at: http://www.sapdesignguild.org/editions/edition8/print amr.asp. [79] staff, c. (2006), sap’s next generation erp software enhances and improves business processes, caribbean business. [80] addo-tenkorang, r., helo, p. (2011). “enterprise resource planning (erp): a review literature report”, proceedings of the world congress on engineering and computer science, vol.2, pp.19-21. corresponding author author can be contracted at jason.choi@polyu.edu.hk. microsoft word 11 x.h.yang,c.h. li--design and simulation for conjugate cam mechanism.doc 286-293 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc design and simulation for conjugate cam mechanism x.h. yang and c.h. li school of mechanical engineering, shandong university of technology, zibo , 255049, china abstract fuzzy synthetically appraising theory is put forward and to be used in the optimal selection of the law-of-cam motion. and its model is also established. an example has been given to illustrate the validity of this method. by compiling program about the optimal design and the simulation system, the motion and process simulation of the conjugate cams mechanism are realized. so the system can improve the efficiency, reduce the cost and solve the precision problem of high speed and over-loading about cam segmentation. keywords conjugate cam mechanism fuzzy evaluation simulation software 1.introduction the conjugate cam mechanism has many excellences, including high speeding and overloading, exact and effective control, steady capability and so on. it is a part of intermission and displacement mechanism, atc equipment and walking equipment, which is widely used in wrapping machine, agricultural machine, printed machine and wiring assemble line and so on. meanwhile, there exist a lot of faults. for example, there are complicated design processes, uneasy to ensure design quality, difficult to manufacture, to depend on import product for domestic need of high speeding and over-loading conjugate cam [1-4]. the cam’s dynamic capability can be effectively improved by fuzzy synthetically appraising theory and optimal selection the cam’s contour line. some problems appeared during the process of design can be earlier solved by combining with advanced virtual design and manufacture technology which can shorten the cycle of design, improve the precision of design and effectively enhance the conjugate cam’s design and the manufacture level. 2.fuzzy synthetically selection of the-law-cam mation it is important to select the-law-follower in the process of conjugate cam design. according to some fuzzy factors including the cam’s working condition, economy, precision, and kinematics and dynamics characteristics of the motion law, meanwhile, synthetically considering a few relevant factors influenced the law-of-follower, the motion capability of conjugate cam can be effectively improved by fuzzy synthetically appraising theory. during the process of confirming the motion law, both loading and speed must be satisfied. so, we regard the two factors as sub-factors of the basic requirements[5-8]. other factors include working conditions, economics, and precision and so on. its secondary appraising rational tree, such as fig. 1: the process to choose the optimal motion law of the cam’s followers by fuzzy synthetically appraising as follows: a. the motion law-of-follower deduced by positive sequence is known as spare set. it can be denoted by the following set fundatin of shan dong and sdut ( no y2008f02、2005kjm04) { }1 11 12 1, , , mv v v v= ⋅⋅⋅ advances in systems science and applications (2011), vol.11, no.3-4 287 fig.1. secondary appraising tree of connection b. the necessarily considered factors such as the maximum speed mv , the maximum acceleration ma , extracted the root-mean-square of speed rmsa , and move-loading torque characteristic value ( )m av and so on, are known as key factors, denoted by iu { }1 11 12 1, , , nu u u u= ⋅⋅⋅ c. as a result of different design requirement and the use spot, the importance of key factors is different. so, a fuzzy subset is established by analysing every key factor. { } ( )1 11 12 13 1 1, , , ,0 1, 1,2, ,n iw w w w w w i n= ⋅⋅⋅ ≤ ≤ = ⋅⋅⋅ in the set, iw1 is important degree coefficient of iu , which is known as weight coefficient. d. according to ( )1,2, ,iu i n= ⋅⋅⋅ , satisfactory subject degree of every element in the spare set is confirmed., forming the following matrix: 111 112 121 122 1 1 ~ 1 1 1 2 1 ... ... ,0 1 ... ... ... ij n n nm r r r r r r r r r ⎡ ⎤ ⎢ ⎥ ⎢ ⎥= ≤ ≤ ⎢ ⎥ ⎢ ⎥ ⎢ ⎥⎣ ⎦ it denotes the subject degree from element jv to factor iju in the spare set. e. the choice of motion law ( ) ( )1m1211 1nm1n21n1 12m122121 11m112111 11211 11 b,...,b,b r...rr ............ r...rr r...rr ,...,, = ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ ⋅= ⋅= nwww rwb hereinto, ( )1 1 1 1 1, 2,..., , n k j jk j b w r k m = = ⋅ =∑ jv corresponding to 1max kb , jv is the best choice.the flow chart of synthetically choice is denoted as fig. 2. whole requirement precision economyworking condition basic requirement speedload 288 yang: design and simulation for conjugate cam mechanism f. example ①requirement of use: middle speed, light-load, low noise. condition choice condition analysis, pick-up the spare set pick-up the weight a>− compute ija aa/a ij >− pick-up the parameters of characteristic curves compute ijr ( ) rd/f,f0.1 ijimax >− rab1 ×= whether to choose the other requirements 1bb = pick-up the weight c>− compute c cc/cij >− rcb2 ×= pick-up the secondary weight ( )21, www =>− 2211 b*wb*b +=×= wbw for i=0 to 5 extract ( )ibmax , output the relevant motion law end y n fig.2 the flow chart of the cam’s followers’ motion law confirmed by fuzzy synthetically theory advances in systems science and applications (2011), vol.11, no.3-4 289 ②spare set: { }1 11 12 13 14 15, , , ,v v v v v v= = {sine, amending constant speed, amending trapezoid, amending sine, equal addition and minus} ③characteristic values of every curve are regarded as key factors: ( ){ }1 , , , , ,m m rms m mm u v a av a jτ= because of the higher speed, ma and mj should be lower. so, its weight coefficient should be a little bigger; because of the light-load, weight coefficient of mv and ( )m av can be a little lower; because of the lower noise, weight coefficient of mτ should be proper bigger. synthetically considering various requirements, the right weight modulus can be taken as: { } { } 1 11 12 13 14 15 16, , , , , 0.2,1,0.1,0.5,0.6,0.8 w w w w w w w= = after computing every percentage of their weights: { }0.0625,0.3125,0.03125,0.15625,0.1875,0.25w = ④characteristic values of various motion law, such as tab.1: tab.1 characteristic values of various motion law subject degree can be formed with ( )max0.1 /ij i ijr f f d= + − ( ) ( )max min / 1 0.1i id f f= − − . the matrix can be formed: ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ = 082.011.057.0 083.088.01.01 85.0145.084.01.0 15.0113.092.01.0 166.08.01.05.0 1.04.01.011.0 r1 ⑤adopting synthetic appraise model of weighted average, namely: 1 , 1, 2,..., m j j ij j b w r i m = = ⋅ =∑ that: ( )1 0.511,0.298,0.746,0.779,0.456b w r= ⋅ = hereinto, the amending sine curve is the best choice which is corresponding to max 0.779,jb = the following curve is amending trapezoid curve (corresponding to 0.746jb = ), sequenc e number curve’s name mv ma ( )m am rmsa mj mτ 1 2 3 4 5 amending equal speed equal addition and minus amending trapezoid sine the mending sine 1.275 2.0 2.0 2.0 1.76 8.013 4.0 4.888 6.238 5.528 5.671 8.0 8.048 8.126 5.435 4.001 4.0 4.232 4.443 3.908 201.4 ∞ 61.43 39.48 69.47 63.3 — 26.71 44.14 34.17 290 yang: design and simulation for conjugate cam mechanism sine, equal addition and minus, amending constant speed. known by experience, amending sine curve among several curves has the strongest use, the better capability/performance. the method’s accuracy and feasibility are validated. the amending trapezoid is in the second place. 3.motion simulation for conjugate cam because contour curve of cam is formed by multiple curves, the correct contour curve of cam will not be gained if the input cam’s parameters and motion law aren’t proper (fig. 3). in order to avoid manufacturing the unqualified cam, the cam’s rationality and accuracy must be detected by motion simulation after gaining the relevant datum of designing cam model. we pick up the above data of the design model and change them into date format which can be recognized by graph software autocad and ug. the cam’s accurate contour graph and threedimensional dynamic simulation model can be automatically built in the graph environment. in the whole process of design, no only the relevant datum can be disposed of; but the visualization of designing results can be achieved by motion simulation. the designer can amend design parameters at any moment. so, the optimal design scheme can be gained, the ratio of waster during the practical process can be reduced, meanwhile, the cost can be reduced too. fig.3. the error cam’s contour curve the flow of simulation analysis and numerical controllable programming module is mainly divided into two parts: first, implement graph simulation in autocad2006 by compiling data interface files (camshape.scr) on the platform of vc++6.0, we can extremely intuitively judge whether contour is distortion or comes to a point. meanwhile, we can compile the interface files that numerical controllable line incision machine tool needs by autolisp language. fig.4. the flow chart of simulation and program model second, implement three-dimensional graph simulation and numerical controllable programming in ug nx3.0 by compiling data interface files (camshape.scr) on the platform of advances in systems science and applications (2011), vol.11, no.3-4 291 vc++6.0, simulate the processing movement of the cutting tool and cutting situation of the workpieces by using the ug’s formidable processing module. examine whether there occurs over-cutting or interference collision in order to validate the accuracy and the rationality of the procedure in front of the post-processing and make the prompt revision, thus guarantee accuracy of numerical controllable programming of cam’s contour. its flow chart of simulation and program model is denoted as fig. 4. the motion simulation is denoted as fig. 5. fig.5 the three-dimensional motion simulation of conjugate cam 4.machining simulation for conjugate cam for the regular cam tested by motion simulation model, we pick up the data which program needs from the graphics software and send data files to the automatic program model, and meanwhile, we choose the type of program, then, the numerical control line incision 3b or machining g programming can be realized according to practical needs. and the accuracy can be detected by machining simulation. if there were some problems, the program would be compiled over again. by the procedure of drawing data in autocad, the cam’s counter data can be gained and changed into 3 b line cutting codes such as fig. 6. and the knife track codes such as fig. 7; the machining simulation is denoted as fig. 8. fig.6. the 3 b line cutting codes 292 yang: design and simulation for conjugate cam mechanism fig.7. the knife track codes at last, by the program’s transferring model, send the correct program to the cam machining tool and complete the cam’s machining. to do so, the design, simulation and manufacture are integrated; it can share the datum and enhance the work efficiency. fig.8. the machining simulation of conjugate cam 5.conclusions according to some fuzzy factors including the cam’s working condition, economy, precision, and kinematics and dynamics characteristics of the motion law, the fuzzy optimal design for the conjugate cam is realised by applying the fuzzy synthetically appraising theory. and its model is also established. meanwhile, program the procedure of the fuzzy optimal design and simulation system. by combining existing advanced graphic simulation software, the design, simulation and manufacture are integrated. therefore, integration work of design and manufacture is completed. so, the system can share the datum, improve the efficiency, reduce the cost and solve the design precision problem of the high-speed, over-loading conjugate cam. references [1] yin ming-fu,zhao zhen-hong. study on one-side machining principle and tool path control method of the globoidal cam. china mechanical engineering. 02(2005)127-130. [2] he hong-mei,zhang mei. cad on the compound law of motion of follower in cam mechanism design. mechanical science and technology. 2(2001)218-219. [3] zhang ming-hong,mu an-le. development of planar disc cam mechanism cad system and three-dimension entity sculpt and movement simulation. modern electronic technique. 19(2004) 4-6. advances in systems science and applications (2011), vol.11, no.3-4 293 [4] zhang.yuhua,shin, joong-ho. a computational approach to profile generation of planar cam mechanisms. journal of mechanical design. 1(2004)183-188. [5] yao yan-an, zhang ce, yan hong-sen. motion control of cam mechanisms. mechanism and machine theory. 4(2000) 593-607. [6] zhang yu-hua,xin chong-gao. design of disk cams with reciprocating roller follower. journal of tianjin university (science and technology). 34(6) (2001)766-770. [7] mosier, ronald g , modern cam design , international journal of vehicle design, automotive design process. 23(1) (2000)38-55. [8] long-long wu, shiau-huei wu, hong-sen yan. simplified graphical determination of diskcam curvature. mechanism and machine theory. 34(1999)1023-1036. advances in systems science and application (2015) vol.15 no.4 384-391 the system of parameters efficiency of financial supervision tatiana v. yalyalieva and elena a. murzina volga state university of technology, department of management and law, yoshkar-ola, russia abstract the solution for the problem of monitoring the effectiveness of the process depends on how timely and adequate state control bodies will react to changes in the external environment and the needs of society. the modern system of government performance management control is another step for russia on its way to enter the european and world space. the presented results of the study can be applied to the government’s management and control. parameters of the effectiveness of controls offered in this article have been used by state bodies of the republic of mari el, russia, while working on the program of innovative development of the region. it is planned to implement the results of future research in the russian practice of governance based on a long-term program of cooperation volga state university of technology and government bodies of the republic of mari el, russia. the article suggests methods for assessing the efficiency of management aimed at public economic supervision, both in a municipality and in a region as a whole. based on the data obtained from the method, a mechanism has been developed to improve the management efficiency. keywords governmental control; economic institutions; legal regulation; effectiveness of institutions. 1 introduction governmental control is performed by officials representing the authorized bodies, within their competences, by collecting explanations, checking the accounting and reporting information, inspecting the premises and territories used to gain income (profit), and by using other instruments provided by the law. one of the most significant areas of governmental control is financial control. financial control is traditionally considered as a part of the governmental control system which covers the sphere of formation, distribution and use of the monetary funds. financial control is control of the legality and appropriateness of the activities performed in the sphere of formation, distribution and use of the monetary funds of the state and municipalities, which is aimed at efficient social and economic development of the country and its regions. representing a constituent part of the economic management, financial control ensures the interconnection between companies and the governmental bodies that have the corresponding rights and authorities. advances in systems science and application (2015) vol.15 no.4 385 calculation of the payments additionally charged as resulting from the controlling activity: components represented in the standardized manner. based on the calculation results, we build a trend line to find the expected amount of the payments in the future. as a result of this work, we have carried out a comprehensive study of the theoretical bases of the governmental control. formation of a certain type of market structures, may, on the whole, be regarded as a response to excessive transaction costs. at this point, integration processes in russia are characterized by aggregation and consolidation of enterprises. this implicitly proves the desire for diversification and economic control of enterprise risks associated with imperfect institutional environment and excessive transaction costs. it should be noted that control over efficiency of costs or any other category may be an extremely complicated and expensive procedure. an example would be the control which the government exercises over taxpayers since it has to maintain the costly tax-collection apparatus. 2 research methodology the calculation method is suitable for the remaining efficiency types, as it allows calculating the total efficiency indicator for the structural components, which are multidirectional indicators expressed in different measures, so the total indicator allows bringing all of them to a unified scale and making their convolution. however, the normalization procedure requires either a data array for several municipalities or information available for a specific period. some issues related to the improvement of the governmental control efficiency were also examined in the works by shemyakina [1]. however, we find it possible to continue studying the subject. certain theses were already mentioned in some of the works by the authors of this article [2–4]. we are going to proceed with the analysis here. the suggested system contains three major blocks of criteria: 1) general efficiency assessment of the controlling activity; 2) assessment of the parameters related to organization and implementation of field audits; 3) assessment of the parameters related to organization and implementation of office audits. to make the analysis and develop appropriate efficiency assessment methods for the work done by governmental bodies, we used the information provided by the governmental bodies of the mari el republic, and by the taxation authorities in particular. the efficiency was analyzed by the major parameters of their activities. here are the most popular coefficients and criteria, which can be divided into three group of parameters: a) general assessment of the controlling activity performed by a taxation authority; b) assessment of the organization and implementation of office tax audits; c) assessment of the organization and implementation of field tax audits. this system of parameters includes: 386tatiana v. yalyalieva and elena a. murzina:the system of parameters efficiency of financial ... a. dynamics of the payments additionally charged as resulting from the controlling activity: ∆a = at + 1−at, where a stands for the amount of the payments additionally charged as resulting from the controlling activity, and t is one year. this value represents the comparison of the amounts of the payments additionally charged for several reported periods, and any positive dynamics would mean improvement of the controlling activity efficiency. b. specific share of the payments additionally charged as resulting from the controlling activity in the total amount of the incoming payments: sadd = a p × 100, where a stands for the amount of the payments additionally charged as resulting from the controlling activity, and p is the total amount of the incoming payments. this value is calculated as the ratio of the amount of the additionally charged payments to the total amount of the payments charged. c. specific share of the payments recovered as resulting from the controlling activity in the total amount of the additionally charged payments: srec = r a × 100, where a stands for the amount of the payments additionally charged as resulting from the controlling activity, and r is the total amount of the payments recovered to the budget system of russia as resulting from the tax audit. this value is calculated as the ratio of the amount of the payments recovered as resulting from the controlling activity to the amount of the payments additionally charged as resulting from the tax audits. it shows what share of the payments additionally charged came to the budget system. d. general effectiveness coefficient: ceff = r c × 100, where r stands for the amount of the payments recovered to the budget system of russia as resulting from the tax audit, and c is the costs of the activities performed by the taxation authority. this value is calculated as the ratio of the amount recovered from the payments additionally charged as resulting from the tax audit to the amount of the costs advances in systems science and application (2015) vol.15 no.4 387 of the activities performed by the taxation authority. it shows the effectiveness of the costs required for the recovery. e. legality coefficient of the extra charges resulting from tax audits: clegal = a−d a × 100, where d stands for the deduction amount of the payments additionally charged as resulting from the decisions of judicial and superior bodies, and a is the amount of the payments additionally charged as resulting from the controlling activity. this value represents the judicial acknowledgement of the legality of the fines and penalties imposed. the ideal value is 100%. 3 findings and discussion in scientific and journalistic literature, governmental control is defined in the broad sense as a special method of ensuring legality. its key tasks are ensuring that all the taxes and duties stipulated by the law are duly paid to the budgets of various levels and preventing tax evasion. the study of the institutional environment of a territory as an evolving endogenous factor was initiated by the representatives of the school of economics of washington university in the 1970s coase [5]. douglass north uses the institutional environment term to define the institutional limits which exist at the macrolevel and determine the possible conditions of contractual agreements between individuals north [6]. the aggregate social and economic effect produced through implementation of governmental programs, which is achieved by applying structural modifications to the economy, can be used as the efficiency indicator feldman [7]. the multi-subject composition of the parties of the contract relations within a cluster can be classified according to the subject composition of the triple spiral model, i.e. it appears possible to single out the state (state and self-government bodies), business and universities enright [8]. oliver williamson defines the institutional environment as an established system of the informal rules of the game which build the sociocultural context of economic activity williamson [9]. when analyzing the efficiency for the control purposes, it is required to consider the main principles that the financing of the state expenses is based on (including those when allotting the funds intended implementation of the target programs). in this case, the methods used to assess the educational clusters can be applied to control the efficiency of the use of the state resources dritsaki and adamopoulos [10]. in the modern conditions, the competitiveness of the subfederal cluster territory is interpreted as a concentrated expression of the production, scientific and technical, and organizational institutional advantages implemented in high technology goods and services porter [11]. 388tatiana v. yalyalieva and elena a. murzina:the system of parameters efficiency of financial ... in the context of this research, the institutional environment is interpreted as the aggregate of the institutions and institutional connections which surround and fill the subfederal economic system and produce their effect on it. the institutional environment creates favorable conditions for the formation of an interaction network between the agents and governmental bodies e.a.murzina[2] , t.v. yalyalieva [3]. at the same time, while changing in the course of time, the institutional aspects of the governmental control mechanisms form the limits of the economic behavior of innovation companies, thus establishing the sociocultural norms taking effect on the behavior of the agents of economy napolskikh [12]. in other words, governmental control means checking on how the laws are observed and elimination of errors and violations [13]. for instance, nikolaos, petros, eleni, evangelos and dimitrios, in their article, looks at governmental control from the position of the necessity to observe the financial discipline as a precondition for individuals and legal entities to properly fulfill their obligations to the state and considers fiscal supervision to be the major function of the taxation control system [14]. larionova, yalyalieva, napolskikh, shebashev [15] representatives of another school of economists and financiers introduce their grounds as follows: “governmental control is a system of measures aimed at supervising the legality, appropriateness and effectiveness of the activities related to the formation of the monetary funds of the state at all the levels of management” [15]. some issues related to the improvement of the governmental control efficiency were also examined in the works by shemyakina [1]. however, we find it possible to continue studying the subject. certain theses were already mentioned in some of the works by the authors of this article [2–4]. we are going to proceed with the analysis here. integration trends in the russian economy are mainly associated with transition processes and enterprises’ adaptation to the imperfect market conditions. the natural reaction of a rational economic agent to excessive transaction costs would be trying to independently reduce them. but as a result, the reproduction system may fall into an institutional trap. since it is impossible to rely on natural tendencies of transaction costs to decrease, the need arises for the government active interference in the process of their ‘natural’ reduction. as a subject of government control, dual nature of transaction costs, especially transformation processes with a high share of informal sector, as well as significant difference in conditions of functioning of various types of markets require a differentiated approach to transaction costs reduction in the national economy. comprehensive mechanism for efficiency control of transaction costs which is used in economy to reduce them involves development of ideology, specification and protection of property rights, standardization of measurements, accounting and reporting, maintenance of monetary system, improvement of law enforceadvances in systems science and application (2015) vol.15 no.4 389 ment effectiveness and efficiency, as well as implementation of measures aimed at eliminating unnecessary administrative burdens and infrastructure markets of various transactions. 4 concluding remarks an indisputable advantage of the suggested system of the parameters for the purposes of governmental control of the efficiency is its versatility, as it is adapted to any kind of system and is valid in the conditions of establishing an uninterrupted regional and national system resources. hence, we suggest assessment methods for efficiency of governmental control of regional economy aimed at economic supervision over the regional activities. based on these methods, we develop a mechanism to improve the resources management efficiency and find the essential ways to improve it. this mechanism is expected to represent a complex of measures aimed at changing the factors producing the biggest effect on the efficiency of governmental control of regional economy. for every factor and efficiency type, we have to suggest ways for efficiency improvement, and those ways have to consist in specific measures. as seen from the data we obtained, the recovery level of the payments resulting from the controlling activity is extremely low; it did not exceed 43% in the period of the analysis. in 2012 it was even less than 30%.it means that only one third of the amount additionally charged reached the budget system of the country. at the regional level, this value is several points lower in the reported period than the average value for the country. however, within the period we can observe a growing trend for this value, which represents a positive result of the work done by the taxation authorities. the recovery value, as estimated per governmental body official, grew by 35% within the period. totally for the entire country, it amounted to 978,000 rubles in 2012, 952,000 rubles in 2013, and 1,106,000 rubles in 2014. the main reasons for the decrease of the value of the payments additionally charged as resulting from the field audits, if compared to the same value of the previous year, are reduction of the number of the audits performed, which, in turn, resulted from the reduction of the number of the officials who actually performed the field audits, and selection of the subjects to be included in the audit plan, basing on the recovery possibility factor applied to the payments additionally charged as resulting from the audits. thanks. this researchers was supported by the ministry of economic development and trade of the mari el republic (agreement no. 126/8 of october 8, 2013) and volga state university of technology, state-financed scientific-research work no 20 “ efficiency of state control in the sphere of economic activity of the 390tatiana v. yalyalieva and elena a. murzina:the system of parameters efficiency of financial ... region ”, no 9 “development institutions of local markets and the formation of an innovative cluster”. references [1] blumenthal, o. i., shemyakina, m.s. (2013), “improvement of system of tax deductions and regulation of a tax rate as ways of increase of efficiency of tax administration of receipts of the personal income tax”. new university: economics law, vol.9, pp.10-16. [2] murzina, e. a. (2014), “fiscal nihilism of subjects of small and medium business as threat of economic security of russia”. new university: economics & law, pp.5-6 and pp.83-87. http://dx.doi.org/10.15350/2221-7347.2014.56.00063 [3] yalyalieva, t. v. (2014), “problems and strategies of quality control of public services at the regional level”,new university: economics & law, pp.7-8 and pp.50-53. http://dx.doi.org/doi: 10.15350/2221-7347.2014.7-8.0008 [4] yalyalieva, t. v. (2014), “economic state control effective land management”, actual problems of economics, vol.9, pp.126-128/ [5] coase, r.(1992),“the institutional structure of production”. the american economic review, vol.82, no.4, pp.713-719. [6] north, d.(1994), “economic performance through time.” the american economic review, vol. 84, no.3, pp.360-361. [7] feldman, m.p.,(1994), the geography of innovation. 1st edn, springer science and business media, dordrecht, pp.154. [8] enright, m., (1996),“regional clusters and economic development: a research agenda”. business networks: prospects for regional development, walter, d.g. (ed.), berlin, pp.190-213. [9] williamson, o. (2000), ”the new institutional economics: taking stock, looking ahead” journal of economic literature vol.38, no.3, pp.595-613. [10] dritsaki, c. and a. adamopoulos,(2005), “a causal relationship and macroeconomic activity: empirical results from european union”. am. j. applied sci.vol.2, pp.504-507. [11] porter, m.e., (2009), “clusters and economic policy: aligning public policy with the new economics of competition”, nstitute for strategy and competitiveness. advances in systems science and application (2015) vol.15 no.4 391 [12] larionova.n.i., napolskikh d. l, yalyalieva , and shebashev v.r e.(2014), “governmental control of the formation efficiency of educational clusters at the regional level”, american journal of applied sciences, vol.11, no.9, pp.1594-1597. [13] koteeswaran, s., p. visu and j. janet (2012), “a review on clustering and outlier analysis techniques in datamining”. am. j. applied sci, vol.9, pp.254258. [14] nikolaos, e., k. petros, p. eleni, p. evangelos and v. dimitrios, (2013), “the impact of the greek economic crisis on the greek construction companies”. back to basics: the statistical cost accounting model. am. j. econom. bus. admin, vol.5, pp.168-174. [15] larionova n.i., yalyalieva t.v.and napolskikh d.l. (2014), “ensuring efficiency control of institutional environment of the cluster”. published online vol.11, no.9, http://www.thescipub.com/ajas.toc [16] kim, y.d., s. yoon and h.g. kim (2014),“ an economic perspective and policy implication for social enterprise”. am. j. applied sci, vol.11, pp.406413. [17] liu, c.c. and c.y. chen (2004), “a computable general equilibrium model of southern region in taiwan: the impact of the tainan science-based industrial park”. am. j. applied sci, vol.1, pp.220-224. [18] murzina, e.a., larionova, n, i. and yalyalieva t.v. (2015), “supervising the efficiency of governmental control of regional economy: legal aspects and economic essence”. mediterranean journal of social sciences, vol 6, pp.370-375 corresponding author tatiana can be contacted at:yal05@mail.ru advances in systems science and applications (2012) vol.12 no.3 297-304 an improvement on kim’s inequality for pricing asian option yirong ying1, sheng xu1 and jeffrey forrest2 1college of management, shanghai university, shanghai, 200444, china 2department of mathematics, slippery rock university, slippery rock, pa16057, usa abstract in his study of a simple one-dimensional parabolic pde, kim[11] obtained an inequality that the price of asian options satisfies. this paper aims to improve this particular kim’s inequality. our inequality has an advantage over kim’s, for our estimation has an upper bound as r vanishes. therefore, our newly established estimation can help us to determine the solution error of a second order partial differential equation with variable coefficients. keywords asian option, parabolic pde, error estimation 1 introduction as a new financial product, asian options, also known as average options, can be seen as innovative european options. the commonality with the european options is that investors are only allowed to exercise their option contracts on the maturity dates, while the difference is that investors of asian options decide whether or not to exercise their option contracts based on the price level of the average share price during the contract term. because the value of european options on the maturity date has nothing to do with the price path and depends only on the maturity date of the stock price, it is difficult to prevent speculators from manipulating the maturity price and consequently from arbitrage. on the other hand, because asian options are associated with the price path, they can be applied to ease market behaviors. they represent the exotic options most actively traded in financial derivatives market today. their difference with the usual-sense stock options is the implementation of price limits so that the exercise price is the average price of the stock price of the secondary market over a period of six months prior to exercise. their difference with the standard options is that the options’ maturity-date benefits are determined not by the prevailing market price of the underlying asset, but by the average price of the underlying asset over some time period during the option contract period. this time period is known as an average period, over which either the arithmetic or geometric mean is applied. asian options can be well designed to avoid stock price manipulation and damage of company interests caused by insider trading. so, they are welcomed by both investors and issuers. additionally, compared to the standard options, asian options also possess such advantages as lower prices and can be applied to hedge the risk of a specified time period. along with baskets, spreads, as well as 298 yirong ying:an improvement on kim’s inequality for pricing asian option other strategies, asian options have been extensively studied by many scholars. even so, few of published works deal with the problem of pricing general asian options. in terms of the pricing of asian options, [1-4] represent some of the most recent results, while dewynne et al. provide a simplified means of pricing asian options using parabolic pde[5]. a progressive solution is established by using laplace transforms and power series expansions. equations of kolmogorov type have also turned out to be relevant in option pricing in the setting of certain models for stochastic volatility and in the pricing of asian options. frentz et al. numerically solve the cauchy problem of a general class of second order degenerate parabolic differential operators of kolmogorov type with variable coefficients by using posteriori error estimates and an algorithm developed for adaptive weak approximation of stochastic differential equations[6]. on the basis of these works, we show how to apply the relevant results in the context of mathematical finance and option pricing. the approach outlined in this paper circumvents many of the difficulties confronted by any deterministic approach based on, for example, a finite-difference discretization of the partial differential equation. meanwhile, frentz[7] also analyze the second order pde operators arising in the pricing of asian options, while proving the optimal interior regularity. min dai presents a lattice algorithm for pricing both europeanand american-style moving average barrier options (mabos)[8]. they develop a finite-dimensional partial differential equation (pde) model for discretely monitoring mabos and solve it numerically by using a forward shooting grid method. however, their modeling pde for continuously monitored mabos is of infinite dimensions and cannot be solved directly by using existing numerical methods. recently, bayraktar et al. construct a sequence of functions based on the value of asian option[9]. as a result, each term of the sequence is the unique classical solution of a parabolic pde so that they provide the relevant numerical approximation. deelstra et al. obtain an approximation formula by using comonotonic bounds[10], leading to four different approximations: the upper, the improved upper, the lower, and the intermediary bounds. in this way, they improve the traditional hybrid moment matching method. their methods have the advantage that they can be applied in other frameworks, e.g., in lévy settings, as well. these results can be adapted to deal with options written in a foreign currency, such as compo and quanto options. kim studies a simple one-dimensional parabolic pde that the price of asian options satisfies[11]. the result indicates that the generalized solution is a classical solution and cannot be obtained. additionally, kim’s inequality is shown to be unbounded as the parameter r vanishes. by improving kim’s inequality, our estimation has an upper bound as r vanishes. advances in systems science and applications (2012) vol.12 no.3 299 2 main results kim (2009)[11] considers the following parabolic type equation ut + 1 2 [x− e− ∫ t 0 dv(s)q(t)]2σ2uxx = 0 (1) along with the boundary condition u(t, x) = max(x−k1, 0) (2) where v(t) stands for dividend yield, σ the volatility of the underlying asset, and q(t) the trading strategy. first let us introduce a lemma about gauss estimation. suppose that g(x) is a continuous function in the interval [−r,r] satisfying 1 2 ≤ g(x) ≤ 3 2 , x ∈ [−r,r] denote q := {(t, x) ∈ r2 : 0 < t < 2, |x| < r}, ω := {(t, x) ∈ q : t > g(x)},σ := {(t, x) ∈ q : t = g(x)}. lemma1[11] let ω and σ be defined as above and a(t, x) a function satisfying{ lu := ut − a(t, x)uxx = 0, (t, x) ∈ ω u = 0, (t, x) ∈ σ assume that u ∈ c1,2 loc (ω) ∩ c(ω̄) and a(t, x) satisfy 0 ≤ a(t, x) ≤ 1, ∀(t, x) ∈ ω then, we have the following estimate: |u|0;ω′ ≤ ( 16√ 2π )r−1e− r2 32 |u|0;ω ,ω′ := {(t, x) ∈ ω : |x| < r 2 } (3) proposition 1 if |x| < r, then∫ e ϕ(2, x− y)dy < 2 ∫ 2r 0 ϕ(2, y)dy + 1√ 2πr e−2r2 (4) where e = ∪ j∈z ((4j + 1)r, (4j + 3)r). proof: first we prove that the following equality holds true:∫ e ϕ(2, x− y)dy = ∞∑ j=1 ∫ (4j−1)r+x (4j−3)r+x ϕ(2, y)dy + ∞∑ j=0 ∫ (4j+3)r−x (4j+1)r−x ϕ(2, y)dy 300 yirong ying:an improvement on kim’s inequality for pricing asian option in fact,∫ e ϕ(2, x− y)dy = ∫ ∞∪ j=1 ((−4j+1)r,(−4j+3)r) ϕ(2, x− y)dy + ∫ ∞∪ j=0 ((4j+1)r,(4j+3)r) ϕ(2, x− y)dy = ∞∑ j=1 ∫ (−4j+3)r (−4j+1)r ϕ(2, x− y)dy + ∞∑ j=0 ∫ (4j+3)r (4j+1)r ϕ(2, x− y)dy = ∞∑ j=1 ∫ x−(−4j+3)r x−(−4j+1)r −ϕ(2, y)dy + ∞∑ j=0 ∫ x−(4j+3)r x−(4j+1)r ϕ(2, y)dy ∵ ϕ(2, y) = 1√ 8π e− y2 8 is an even function, ∴ ∫ e ϕ(2, x− y)dy = ∞∑ j=1 ∫ (−4j+3)r−x (−4j+1)r−x ϕ(2, y)dy + ∞∑ j=0 ∫ (4j+3)r−x (4j+1)r−x ϕ(2, y)dy = ∞∑ j=1 ∫ (4j−1)r+x (4j−3)r+x −ϕ(2, y)dy + ∞∑ j=0 ∫ (4j+3)r−x (4j+1)r−x ϕ(2, y)dy secondly, we prove that when −r < x < 0, the following inequality holds true ∞∑ j=1 ∫ (4j−1)r+x (4j−3)r+x ϕ(2, y)dy ≤ ∫ 2r 0 ϕ(2, y)dy + 1 2 √ 2πr e−2r2 (5) in fact, when −r < x < 0, ∞∑ j=1 ∫ (4j−1)r+x (4j−3)r+x ϕ(2, y)dy < ∞∑ j=0 ∫ (4j+2)r 4jr ϕ(2, y)dy = ∫ 2r 0 ϕ(2, y)dy + ∞∑ j=1 ∫ (4j+2)r 4jr ϕ(2, y)dy ∵ ∞∑ j=1 ∫ (4j+2)r 4jr ϕ(2, y)dy < ∫ ∞ 4r ϕ(2, y)dy = ∫ ∞ 4r 1√ 8π e− y2 8 dy < 1√ 8π ∫ ∞ 4r y 4r e− y2 8 dy = 1 2 √ 2πr e−2r2 advances in systems science and applications (2012) vol.12 no.3 301 ∴ ∞∑ j=1 ∫ (4j−1)r+x (4j−3)r+x ϕ(2, y)dy ≤ ∫ 2r 0 ϕ(2, y)dy + 1 2 √ 2πr e−2r2 thirdly, we prove that as 0 < x < r, we have ∞∑ j=0 ∫ (4j+3)r−x (4j+1)r−x ϕ(2, y)dy ≤ ∫ 2r 0 ϕ(2, y)dy + 1 2 √ 2πr e−2r2 (6) in practice, as 0 < x < r, ∞∑ j=0 ∫ (4j+3)r−x (4j+1)r−x ϕ(2, y)dy ≤ ∞∑ j=0 ∫ (4j+2)r 4jr ϕ(2, y)dy = ∫ 2r 0 ϕ(2, y)dy + ∞∑ j=1 ∫ (4j+2)r 4jr ϕ(2, y)dy ≤ ∫ 2r 0 ϕ(2, y)dy + 1 2 √ 2πr e−2r2 put (5) and (6) together, the proof is completed. qed. proposition 2, if |x| < r, then∫ e ϕ(2, x− y)dy < √ 1− e−r2 + 1√ 2πr e−2r2 (7) proof: for arbitrary r > 0, we have inequality ∫ 2r 0 ϕ(2, y)dy < 1 2 √ 1− e−r2 . in fact, (∫ 2r 0 ϕ(2, y)dy )2 = ∫ 2r 0 ∫ 2r 0 1 8π e− x2+y2 8 dxdy ≤ ∫∫ d 1 8π e− x2+y2 8 dxdy (8) = 1 4 (1− e−r2 ) where d = {(x, y)|x2 + y2 ≤ 8r2, x ≥ 0, y ≥ 0}. hence ∫ 2r 0 ϕ(2, y)dy < 1 2 √ 1− e−r2 . by substituting the above inequality in proposition 1, inequality (7) is obtained at once. qed. proposition 3 if k > 1, r2 < 2ln2 16k−9 , then∫ (4j+2)r 4jr ϕ(2, y)dy ≤ ∫ 8jr (8j−4)r ϕ(2, √ ky)dy (9) 302 yirong ying:an improvement on kim’s inequality for pricing asian option proof: after changing the variable t = y+4r 2 , we have∫ 8jr (8j−4)r ϕ(2, √ ky)dy = 1√ 8π ∫ 8jr (8j−4)r e− ky2 8 dy = 1√ 8π ∫ (4j+2)r 4jr eln2− k(y−2r)2 2 dy as r2 < 2ln2 16k−9 , 4jr < y < (4j + 2)r, j = 1, 2, · · · , the following inequality holds true:ln2− k(y−2r)2 2 + y2 8 ≥ 0. hence 1√ 8π ∫ (4j+2)r 4jr e− y2 8 dy ≤ 1√ 8π ∫ (4j+2)r 4jr eln2− k(y−2r)2 2 dy. that is ∫ (4j+2)r 4jr ϕ(2, y)dy ≤ ∫ 8jr (8j−4)r ϕ(2, √ ky)dy. qed. proposition 4 if k > 1, r2 < 2ln2 16k−9 , then∫ e ϕ(2, x− y)dy < 2 ∫ 2r 0 ϕ(2, y)dy + 1√ 2πkr e−2kr2 (10) proof: combining the results of proposition 2 and proposition 3 leads to∫ e ϕ(2, x− y)dy ≤ 2 ∫ 2r 0 ϕ(2, y)dy + 2 ∞∑ j=1 ∫ (4j+2)r 4jr ϕ(2, y)dy ≤ 2 ∫ 2r 0 ϕ(2, y)dy + 2 ∞∑ j=1 ∫ 8jr (8j−4)r ϕ(2, √ ky)dy ≤ 2 ∫ 2r 0 ϕ(2, y)dy + 2 ∫ ∞ 4r ϕ(2, √ ky)dy ≤ 2 ∫ 2r 0 ϕ(2, y)dy + 1√ 2π ∫ ∞ 4r y 4r e− ky2 8 dy = 2 ∫ 2r 0 ϕ(2, y)dy + 1√ 2πkr e−2kr2 .qed. proposition 5 if k > 1, r2 < 2ln2 16k−9 , then ∫ e ϕ(2, x − y)dy < √ 1− e−r2 + 1√ 2πkr e−2kr2 . proof: by substituting inequality (8) in proposition 4, one knows that inequality (10) is true. qed. 3 conclusions with the increasing globalization of investment in recent years, a variety of asian options has obtained a wide range of applications and development. for example, quanto option represents a contingent claim that the option’s benefit depends on the prices of financial derivatives in a foreign currency, while the actual transaction of the option is paid in local currency. quanto options, as a kind of option advances in systems science and applications (2012) vol.12 no.3 303 designed for avoiding different types of risks involved in international trades and investments, represent a financial derivative instrument. they can be used as a hedge in international trades against fluctuations in exchange rates. therefore, with the rapid development of international trades and the globalization of the financial industry, this option will definitely attract more attention. the revenue function of quanto option stands for a joint concern of foreign asset prices and exchange rates, creating many alternative investing and hedging opportunities. additionally, other than asian options, power options represent yet another new class of options. they alter the pricing structure of assets and greatly increase the sensitivity of price over time. power-asian options are a unification of the two, making it possible for investors to better avoid risks. an example is the moving average call option. it has been often used to design a poison pill, a business strategy used to increase the likelihood of negative results over positive ones against a party that attempts a takeover. the moving average calls, issued to existing shareholders, would be triggered by the event of a hostile takeover. the french investment bank compagnie financiere indosuez and the french construction company bouygues have successfully issued such options/warrants to protect themselves against potentially unfriendly investors. kim’s inequality is very useful for our work[12]. however, kim’s estimation is possibly unbounded as r vanishes. our inequality, developed in this paper, has an advantage over kims inequality, because our estimation is bounded on the top as r vanishes. therefore, our work in this paper may help us to look for the solution error of a second order partial differential equation with variable coefficients. acknowledgements this research is supported by the research fund of program foundation of ministry of education of china (10yja790233). references [1] xu chenglong, zhou jing, ren xuemin. (2007), “arbitrage analysis of a class of deposit product with option style”, journal of tong ji university (natural science), vol.35, no.7, pp.994-997. [2] wu zhen, wang guangchen. (2007), “a black-scholes formula for option pricing with dividends and optimal investment problems under partial information”, journal of system science and mathematic sciences, vol.27, no.5, pp.676-683. [3] wang xiaotian. (2010), “scaling and long-range dependence in option pricing i: pricing european option with transaction costs under the fractional 304 yirong ying:an improvement on kim’s inequality for pricing asian option black-scholes model”, physica a: statistical mechanics and its applications, vol.389, no.3, pp.438-444. [4] chen caisheng, li gang, zhou jidong, wang wenchu. (2008), mathematical physics equations, science press. [5] dewynne, shaw. (2008), “differential equations and asymptotic solutions for arithmetic asian options: ‘black-scholes formulae’ for asian rate calls”, european journal of applied mathematics, vol.19, pp.353-391. [6] marie frentz, kaj nyström. (2010), “adaptive stochastic weak approximation of degenerate parabolic equations of kolmogorov type”, journal of computational and applied mathematics, no.234, pp.146-164. [7] marie frentz, kaj nyström, andrea pascucci, sergio polidoro. (2010), “optimal regularity in the obstacle problem for kolmogorov operators related to american asian options”, math. ann, vol.347, pp.805-838. [8] min dai, peifanli, jine.zhang. (2010), “a lattice algorithm for pricing moving average barrier options”, journal of economic dynamics & control, no.34, pp.542-554. [9] erhan bayraktar, hao xing. (2011), “pricing asian options for jump diffusion”, mathematical finance, pp.117-143. [10] griselda deelstra, alexandre petkovic, michèle vanmaele. (2010), “pricing and hedging asian basket spread options”, journal of computational and applied mathematics, no.233, pp.2814-2830. [11] seick kim. (2009), “on a degenerate parabolic equation arising in pricing of asian options”, journal of mathematical analysis and applications, vol.351, pp.326-333. [12] yi-rong ying, yu-yuan tong, jeffrey forrest. (2010), “pricing asian options: approach of decomposition and estimation”, advances in systems science and applications, vol.10, no.2, pp.241-247. corresponding author author can be contacted at:jeffrey.forrest@sru.edu microsoft word 15 li shuiping, li xiaotian,shi lugang--simulation of the flow field of cement mixer based on numerical metho advances in systems science and applications (2011), vol.11, no.3-4 315-321 issn 1078-6236 international institute for general systems studies, inc simulation of the flow field of cement mixer based on numerical methods li shuiping, li xiaotian, shi lugang school of mechanical and electronical engineering, wuhan university of technology, wuhan 430070, p.r.china abstract different kinds of research methods of the flow field in the cement mixer were introduced and compared, theory bases of the numerical simulation method of the flow field in the cement mixer were described, whole of thoughts of realization this method were analyzed. and then, the geometrical model and mathematical mode of the cement mixer were created by fluent, which was based on cfd theory; and the characteristic of flow field in the cement mixer was simulated by mixture model, simple algorithm and standard k-ε turbulence model. the calculation results showed that the distribution of flow field in cement mixer could be simulated preferably by the numerical technology, which supplied the valuable theoretical guidance for designing of the cement mixer. keywords cement mixer numerical simulation flow field 1.introduction cement slurry is frequently needed to grout the cracks of the buildings when reinforcing the constructions. currently, most cement slurry is produced by the common mechanical mixer which makes grout by vane, scraper or spiral device. and the mixer is composed by gear, mixing shaft and impeller. it will easily produce the dead angle and precipitation at the bottom of mixing tank in the mixing process due to the blade is fixed on the shaft. apart from that, it has the problems which are big energy waste, low efficiency and can not prepare for different water-cement ratio cement slurry fast. therefore, the development of a new type of high-performance cement slurry mixing equipment has become rather important. currently,cement mixer commonly can be divided into three categories:colloidal mills mixer, blades mixer, water-jet mixer. although these devices are generally able to achieve the desired technical effect, the mixing efficiency and cost is another matter. colloidal mills mixer is an advanced equipment, mixing, high efficiency, high concentration slurry can be stirred, but the structure is more complex and costly. the structure of blades mixer is simple, the cost is inexpensive, and widely used in engineering, but the efficiency is low. water-jet mixer is a low-cost equipment, have the characteristics of high speed stirring, but because the mixing uniformity is not high, to some extent these limit its wider use. so how to provide advanced cement mixing equipment is an important issue that engineering and technical personnel in the country have to faced. for the design of cement mixer, traditional theory and experimental measurements usually used in analysis of the flow field. study found that in the mixer there are various of slurry flow forms at the same time, combined with the interaction of cement powder and water, make the two-phase flow field very complex and difficult to carry out a detailed theoretical analysis. also the complex flow characteristics and opacity of the slurry limit the application of experimental methods. in process of today's research about fluid mechanical design, the cfd method to simulate the flow fluid inside the machine has become an important technical means when the software of cfd have emerged. in engineering, numerical simulation has been more widely 316 li: simulation of the flow field of cement mixer based on numerical methods used, it can be replaced experiment to cut down the cost in some extent, in addition, numerical simulation can provide a lot of information about flow field, it can give reliable basis for the optimal design of fluid machinery for designers. comprehending of the cement slurry flow characteristics and the mechanism of mixing is a basic prerequisite for the structural design [1]. the main purpose of this paper is to simulate and analyze the flow field of the equipment based on the numerical simulation, and then study the mixed performance. 2.research methods of the mixer flow 2.1 introduction of research methods of the flow field 2.1.1 theoretical analysis method on the basis of the relevant experimental data and the law of conservation of energy and so on, it will establish the velocity field distribution of fluid flow, the stress field distribution of fluid flow , mass conservation law and energy conversion law by the mathematical analysis, thus construct the mathematical model of fluid flow[2] .as a result, it can forecast the intensity and characteristics of the fluid flow of mixer as well as its function and influence on mass transfer and energy exchange. 2.1.2 experimental methods (1) visualization of fluid method visualization of fluid method which can monitor the trajectory of fluid is a basic method of studying flow field. there are anisotropic thin (particulate) method, laser-induced fluorescence method and dye tracer method [3]. anisotropic thin (particulate) act refers to adding the anisotropic particles (such as titanium dioxide, mica, aluminum stearate, etc.) which have the reflective capacity to the fluid so that the anisotropic particles flow with the fluid, and the strong reflective capacity of particulate makes the flow trajectory of fluid can be visible. in addition, we are able to study the movement law of two-phase fluid and the migration law of material according to the selective dispersion of some particles in fluid (for example, water phase is chosen for mica, while hydrophobic organic phase are chosen for aluminum and mica). (2) video measurement using advanced camera equipments (such as magnetic resonance fluid imager, laser doppler veloeimetry, particle image vefocimetry etc.) to film the track of fluid flow and then obtain it’s boundary conditions as well as the related parameters [3]. piv technique is non-contact measurement technology which has many advantages such as high measurement accuracy, wide range of speed measurement speed, having simple principle and small influence from outside .consequently, it is able to depict the whole velocity vector of the total velocity field, so the flow field information of one profile can be analyzed at once. 2.1.3 numerical simulation method numerical simulation is that adopting numerical method and logical way to reflect changes prototype process and the rules of movement in the computer. mathematical model, different from the physical model, is often made from a single equation, a group of equation or groups formulas. we can get the various locations of the basic physical quantities (such as speed, pressure, temperature, concentration, etc.) and their changes in the very complex flow by numerical simulation. in addition, it can optimize the design of structural by combining with cad. 2.2comparison of the research methods of the mixer’s internal flow field advances in systems science and applications (2011), vol.11, no.3-4 317 table 1 comparison of the research methods numerical simulation method combined with traditional theoretical analysis method and experimental measurement method form a perfect system of studying the problems of the mixing equipment’s internal fluid flow. from table 1, we can we see clearly that these methods have their own advantages and disadvantages. both traditional theoretical analysis method and experimental test method can not obtain the characteristics of flow field mainly due to the following two points: first, the visualization methods such as staining piv and tracer method can not be used to measure the flow field which has high concentration and poor visibility; second, at present, all of the contact velocimeter can not test the velocity vector of cementing slurry flow field accurately. therefore, the slurry flow field can only be studied by numerical simulation. 3. numerical simulation method of the internal flow field of mixer 3.1 theoretical basis of numerical simulation the theoretical basis of numerical simulation is computational fluid dynamics. fluid flow is dominated by physical conservation laws which relates to the principal basic laws including mass conservation law, momentum conservation law, energy conversion and conservation law, components conversion and balance law,and so on[2,5]. these laws of the basic equations can be expressed as follows: ( ) ( ) ( )div u div grad s t φ ρφ ρ φ φ∂ + = γ + ∂ (1) expand the form of: ( ) ( ) ( ) ( )u v w t x y z ρφ ρ φ ρ φ ρ φ∂ ∂ ∂ ∂ + + + ∂ ∂ ∂ ∂ ( ) ( ) ( ) s x x y y z z φ φ φ∂ ∂ ∂ ∂ ∂ ∂ = γ + γ + γ + ∂ ∂ ∂ ∂ ∂ ∂ (2) whereφ is the general dependent variable, and so can represent the solution of variables of u v w t、 、 、 ; γ is the generalized diffusion coefficient; s is the generalized source term. fheoretical nnalysis method experimental method numerical simulation m ain a dvantage the results whose all kinds of influencing factors are clearly visible have universality. it is the theoretical basis which is used to guide the experimental research and verify the new numerical calculation method. the experimental results which are true and believable,are the basis of theoretical analysis and numerical method. it is easy to choose different physical parameters to make the various effective and sensitive tests because you are not restricted by physical model and mathematical model. m ain d isadvantages we are often required to abstract and simplify the object of calculation. for the non-linear situation, only the analytic results of simple flows can be given. experiments are often limited by model size, flow disturbances and measuring accuracy, in addition ,they will also encountered lots of problem such as the enormous cost of funding and manpower,material resources and the long cycle. it strongly rely on the mathematical model,and numerical processing method will lead to false results. 318 li: simulation of the flow field of cement mixer based on numerical methods in the macro category of macro continuous mechanics, two-phase flow analysis can be divided into two types: (1) local homogeneous flow analysis; (2) two-phase slip flow analysis. cementing slurry flow in stirred tank belongs to two-phase slip flow which mainly has three types of analysis model: euler-lagrange model, eulerian-eulerian model that is a mixture of the two-fluid model and mixture model. euler-lagrange model centralizes the advantages of macroscopic pseudo-fluid model and micro-dynamics model, that is, take into account not only the solid phase and liquid-phase interactions, and can describe the collision between particles, but it needs to record the spatial location and movement of every particle in the flow field , therefore a larger computer memory is necessary. currently, there are certain difficulties in the simulation of high concentration of solid-liquid two-phase flow [4, 5]. at present, multi-fluid model is given a wide range of applications. in a multi-fluid model, the continuous (liquid) phase movements is described by euler equations, assuming that the dispersed (particle) phase is the mutual penetration and continuous quasi-fluid, its movement is also described by euler equation, and then solve the momentum, energy and quality equations of each phase. however, the computational complexity of multi-fluid model results in poor stability, so it confronts many difficulties in the practical application. mixture model is a simplified multiphase flow model, which is used to simulate each phase of the multiphase flow at different speeds, however, the local equilibrium under short spatial scale is assumed. for mixture model, there is a strong coupling between different phases. when there exist a wide range of particle distribution, mixture model can achieve good results as perfect multiphase flow model. due to space limitations, the control equations of two-phase flow will not be listed here. 3.2 the process of numerical simulation of the internal flow field at present, the numerical simulation of mixing process is on the basis of macro flow field. first of all, it obtains the velocity field distribution law of macro-flow field and then joins tracer when the flow field is stable. through the calculation of tracer concentration changes with space and time to simulate the mixing process. this method applies to the flow field of relatively low concentration. in spite of that, there are great difficulties in the simulation of high concentration slurry flow field. in this paper, cementing slurry two-phase flow field was simulated by mixture model, and then the mixing effects were assessed according to the results of simulation. the aim of numerical simulation is to obtain the useful data used to instruct the optimal design of mixer, such as mixing time and mixing power number. the basic flow of numerical simulation is shown in figure1. in the process of numerical simulation, we can create the geometric model according to the structure of mixer. generally, we should deal with the physical mode reasonably during the establishment of geometric model since because it is usually rather complex [6, 8]. in order to reduce the unnecessary computation time, the physical model need to be simplified as much as possible, such as the cover of the mixer which has little impact on the effect of stirring, therefore, it ought to be neglected during the geometric model establishment; for those structures which have greater impact on the effect of stirring, not only to retain but also to establish them accurately. the establishment and dispersion of the mathematical model which directly affects the accuracy and precision of simulation are the core parts of numerical analysis. the key of numerical simulation is to establish the mathematical model which can reflect the essence of engineering problems. for different fluids, we can choose different suitable mathematical models, such as rng k-ε two-equation turbulence model or two-realizable k-ε model may be the best choice in a flow field with a cyclone or strong bending wall [9]. we discretize the regional of calculation before numerical calculation, that is, divide the space of calculation into many sub-regions, and to identify nodes in each region in order to generate grid. and then, the control equations are discretize in the discrete grids, as a result, partial differential equations turn into the algebraic equations of each node for numerical solution. algebraic equations have a advances in systems science and applications (2011), vol.11, no.3-4 319 variety of discrete formats including quick format which can provide higher calculation accuracy when we study swirling flow by structured grid. so the quick format may be the best choice during the numerical simulation of the internal flow field in mixer. the physical model of mixer establishing geometric model mesh generation establishing mathematical model algorithm optimization numerical simulation visual data mining analyze, compare,optim ize judgment of convergence yes no figure 1 the basic flow of numerical simulation apart from a few simple questions, the discrete equations can not be directly solved, because certain adjustments of the dispersion equations are necessary, and we ought to deal with the order and way of unknown variables (speed, pressure, temperature, etc.) especially in order to solve them easily. simple algorithm is widely used in engineering. before carrying out the computer simulation, setting the model boundary conditions is required because the boundary conditions which include the inlet velocity, boundary conditions, outlet pressure, the concentration of slurry, are necessary for solving the discretized equations. setting appropriate initial value and sub-relaxation factor appropriate before the iterative solution in favor of iterative calculations converge faster and more easily. once convergence of iterative solution achieves prescriptive convergence precision, the data we need can be obtained through the user interface. 4. process of numerical simulation of flow field of mixer by fluent 4.1 tools of numerical simulation of flow field on the basis of the theory and method above, we can achieve the numerical simulation of flow field by gambit and fluent soft wares [10]. fluent is specific cfd software which is used to simulate and analyze a complex geometry in the field of fluid flow and heat exchange while gambit is professional pre-processing software for computational fluid dynamics, its main features include three aspects: geometric modeling, mesh generation and assignment of boundaries. among them, mesh generation is the most important function. it generates grid file containing boundary information. fluent can realize numerical simulation of flow field of mixer effectively in connection with its pre-processing software gambit. 320 li: simulation of the flow field of cement mixer based on numerical methods 4.2 the establishment of geometric model mesh generation is the necessary condition to establish discrete control equations, because the quality of the mesh has a direct impact on the accuracy of simulation results. through analyzing the physical model of mixer, this paper used gambit to establish the mesh model on the basis of geometric model of mixer, as shown in figure 2. figure 2 the mesh model of mixer to reduce quantity of mesh cells and to improve precision and convergence, we used structured mesh in the nozzles. for the internal structure with complex geometric shape, the flow field changes rapidly, it is inconvenient to adopt structured mesh, so a structured-unstructured mixed type mesh was employed in this computation. 4.3 solve the discrete mathematical model mesh model file will be optimized after being imported in fluent. we discreted control equations of mesh nodes by numerical discretization of quick format in this paper. numerical algorithm is one of the methods for solving the discrete control equations of grid model, and simple algorithm is quite applicable. 4.4 preliminary analysis of simulation results to a large extent, the effect of mixing depends on the internal structure of flow field, so the existence of turbulence which has complex flow field is an important prerequisite for enhancing the effect of mixing. figure 3 velocity vectors colored by radial velocity (m/s) after analyzing the velocity vector distribution chart, it can be seen that the mixer produced a strong vortex under the action of the spiral blade combined the nozzles. the radial velocity is advances in systems science and applications (2011), vol.11, no.3-4 321 shown in figure 3. 5. future research at present, the literatures about the numerical simulation of cement mixer are rare. although the numerical simulation method has its unique advantages, there are many aspects need to be further improved, which mainly include the numerical simulation accuracy and convergence rate. numerical simulation accuracy of multiphase flow mainly depends on the discrete mathematical model, therefore how to establish mathematical model in accordance with the mechanical structure of mixer is one of the difficulties of numerical simulation. low–order dispersion and high-order dispersion has its own excellent shortcomings, so the way to solve the discrete control equations is a very key technology. the purpose of numerical simulation of flow field is obtaining the useful data to optimize the mechanical structure, thus making further analysis of the simulation results is one of the future research directions. acknowledgements this research is funded by the national key technology r&d program of china (no. 2006baj03a09-08). references [1] zhao jianjun,yuan shouqi,liu houlin, huang zhongfu and tan mingao. simulation of solid-liquid two-phase turbulent flow in double-channel pump based on mixture model. journal of transactions of the csae, 24 (1)( 2008) 7-10. [2] wang fujun. analysis of computational fluid dynamics—principle and application of cfd software. beijing: tsinghua university press , 2006. [3] zhan hanhui, chenghao, liu jianwen and zhan xuehui. principle of secondary flow. changsha:central south university press, 2005. [4] anders darelius, anders rasmuson, berend vanwachem, ingela niklasson and staffan folestad. cfd simulation of the high shear mixing process using kinetic theory of granular flowand frictional stress models. journal of chemical engineering science, 63(8)( 2008) 2188-2197. [5] zhou lixing. dynamics of multiphase turbulent reacting fluid flows, beijing: national defense industry press, 2002. [6] wu bo,yan hongzhi,duan yiqun. study on 3-d turbulence numerical simulation and wear characteristics of slurry pump. journal of china mechanical engineering. 20(6)( 2009) 719-780. [7] huang si, wang guo-yu. a 3d numerical simulation of solid-liquid turbulent flow with high solid-concentration in a centrifugal pump. journal of coal mining machinery. vol.11(2005) 53-57. [8] k.h. javed, t. mahmud and j.m. zhu. numerical simulation of turbulent batch mixing in a vessel agitated by a rushton turbine. journal of chemical engineering and processing, 45(2)( 2006) 99–112. [9] benjamin coesnon, mourad heniche, christophe devals, franc¸ois bertrand and philippe a. tanguy. a fast and robust fictitious domain method for modeling viscous flows in complex mixers: the example of propellant make-down. international journal for numerical methods in fluids, 58(4)( 2008) 427–449. [10] fluent inc. fluent user guide.2003. advances in systems science and application (2016) vol.16 no.3 1-10 on mutual coincidence of return on capital measuring principles by adam smith and karl marx (a critique of “capital in the twenty-first century” by thomas piketty) s.baizakov1,y.utembayev2,a.r. oinarov3 and j.forrest4 1 scientific supervisor, jsc "economic research institute"; astana, kazakhstan; 2 independent expert, astana, kazakhstan; 3 oinarovazamat, chairman of the board of the jsc "kazakhstan center for public-private partnership", astana, kazakhstan; 4 jeffrey forrest, school of business, slippery rock university, philadelphia, usa annotation in his book “capital in the xxi century” by thomas piketty justified the two basic laws of capitalism. on the basis of these economic laws thomas piketty concluded destabilizing role of accumulated national capital. based on his critical analysis this paper studied alternative ways of identification and knowledge of the objective laws’ system that would become the tools of detection imbalances in the economy and assessing the impact of the regulatory impacts on the development of a market economy. develop an appropriate model for the analysis of regulatory impacts. keywords capital, profitability, scientific and technological potential, law, macroeconomics, balance, effect methods thomas piketty, author of the bestselling book “capital in the twenty-first century”, has paid serious attention to the methods of measuring economic growth, including the method of measurement of nominal gdp, defining its by multiplication of return on capital by its volume. he correctly noted that “the concepts of inflation and growth are not always very well defined. the decomposition of the nominal growth (the only kind that can be observed with the naked eye, as it were) into a real component and inflation component is in part arbitrary and has been the source of numerous controversies” [1]. speaking about the effects of accumulation and return on capital’s decrease, he notes, “...on the basis of historical experience, the most likely outcome is that the volume effect will outweigh the price effect, which means that the accumulation effect will outweigh the decrease in the return on capital.”[1]. at the same time, he pays attention to that in agriculture there is a reverse process: “... capital (such as farmland in the case in point), it is inevitable that beyond a certain point, the price effect will outweigh the volume effect... there is no better illustration of the maxim “too much capital kills the return on capital” than the relative value of land and land rents in the new world and the world” [1]. in general in the piketty’s book, the principle of two-dimensional measurement 2 s.baizakov, y.utembayev, a.r.oinarov, j.forrest:on mutual coincidence ... of economic growth was successfully used to analyze the price effect, as a form of the capital’s cost in its form of goods, and volume effect, as a form of the capital’s cost of in its form of money. the same principle is used for two-dimensional measurement of the two fundamental economic laws of capitalism (in the terminology of the author). however, deeper analysis shows that these laws have no any relation to the fact that the “the volume effect is will outweigh the price effect, which means that the accumulation effect will outweigh the decrease in the return on capital”. moreover it hasn’t related to“...farmland in the case in point it is inevitable that beyond a certain point, the price effect will outweigh the volume effect”. pikettydoes his findings on the basis of historical retrospective in some combination of it with the dynamics of economic laws. the reason for these inconsistencies and inaccuracies has often derived from an outdated theoretical basis of establishment of analytical toolsthat are being used now in the practice of market economies institutions and management agents’ regulatory impact assessment. in our opinion, the main factor that has a destabilizing effect on the sustainability of the market economy is the outdated persistence of a one-dimensional measure of return on capital, which is recommended by adam smith as a tool to assess the true value of goods and money[2]. so, up to this day the principle of a one-dimensional measure of return on capital by adam smith has existed in the theoretical basis of current models of balanced economic growth. as an example, the traditional formula of the gdp deflator (inflation) pb, can be taken for, which has the known form as follows: pb = ngdp rgdp (a) where ngdp-is nominal gdp, which by its content is the final product in money terms. the theoretical basis of the formula (a) is the adam smith’s principle of a one-dimensional measure of return on capital. according to this formula, if nominal gdp (ngdp) grows faster than real gdp (rgdp) in the long term, what is happening in some developing countries nowadays, then the gdp deflator (pb) as an indicator of inflation, fueled by the national currencies devaluation may tend to infinity (pb −→ ∞). this is not the only example, which defines the adam smith’s principle of one-sidedness of return on capital’s measurement and its limitation as a balanced growth’s theoretical base models. thus, in the equation of monetarism, by replacing pb*rgdp onto ngdp, we have got: ngdp = v ∗m where m money supply, v velocity of money. advances in systems science and application (2016) vol.16 no.3 3 since v (velocity of circulation) according to the monetarism model is a constant, and taken m (money supply) tending to infinity, then nominal gdp (ngdp) may also tend to infinity. even though thomas piketty properly finds out the formula (a)’s limit to assess the balanced economic growth, he himself falls into the formula’s trap. for instance, in his analysis of the real growth rate’s measurement tools thomaspikettyhas failed to go beyond the limits of the adam smith’s principle of a one-dimensional measure of return on capital. his artistic success in the establishment of two economic laws of national income and the national capital, could not spread further towards the correct breakdown of “nominal growth into real and inflationary components”, and towards the elimination of errors in evaluation methods for growth rateand inflation. in short, thomas piketty diagnosedproperly the market economy disease, but was not able to determine the treatment methods. identification and knowledge of the market economy development laws is needed to make key management decisions and adjustments of previously taken ones. otherwise, inaccurate directions and contradictory readings of analytical tools result in the destabilizing and social inequalities that are listed in the thomas piketty’sbook. as already mentioned, the theoretical basis of all actual current market equilibrium and balanced growth’s models, including piketty’s economic equilibrium model, up to this day is founded on the adam smith’s one-sided principle of return on capital’s measurement. adam smith’s one-sided approach consists in reducing the economic growth’stwo-dimensional,by working time and national currency, measurement system into the one-dimensional system measured by money. for adam smith, “the annual cost of the product” is identified with “the cost that is newly established during the year[3]. more precisely, he was correct that the value of the annual product can be reduced to the same components of the newly created during the year value. adam smith and his modern followers are right in establishing the value of the final product by deducting material resources spent costs from the “cost of the ready product”. from a mathematical point of view these transactions by a.smith and his followers are perfect. however, economy development indicators, both by the adam smith’s theory and by the karl marx’s theory, are measured in the first place by working time. by both theories the price of the working time is the base measure for productive forces of labour and capital . thus, adam smith specifically states that “labour measures the value, not only of that part of price which resolves itself into labour, but of that which resolves itself into rent, and of that which resolves itself into profit”[2]. then he gives a specific example:“in the price of corn, for example, one part pays the rent of the landlord, another pays the wages or maintenance of the labourers and labouring cattle employed in producing it, and the third pays the 4 s.baizakov, y.utembayev, a.r.oinarov, j.forrest:on mutual coincidence ... profit of the farmer. these three parts seem either immediately or ultimately to make up the whole price of corn. a fourth part, it may perhaps be thought is necessary for replacing the stock of the farmer, or for compensating the wear and tear of his labouring cattle, and other instruments of husbandry. but it must be considered, that the price of any instrument of husbandry, such as a labouring horse, is itself made up of the same time parts; the rent of the land upon which he is reared, the labourof tending and rearing him, and the profits of the farmer, who advances both the rent of this land, and the wages of this labour. though the price of the corn, therefore, may pay the price as well as the maintenance of the horse, the whole price still resolves itself, either immediately or ultimately, into the same three parts of rent, labour, and profit.”[ibid, p.105]. this approach of adam smith is static, and focused on the short term and for a momentary effect when the scientific and technological potential of the country has no time to be updated, or when its impact on the real economy can be ignored. karl marx’s approach is focused on long-term development and the market economy sustainable development. but his approach does not replace the adam smith’s principle of return on capital, but summarizes and complements it. a commodity, according to marx, is in development, its dividing onto “the goods and money is the law of the expression of the product as a commodity” [3]. it is this law thomas piketty did not consider by partly remaining within the framework of the adam smith’s principle of a one-dimensional measurement. he failed to take into account that goods and money were developing and transforming into capital, which became a form of its product or its form of money. in the third volume of “capital” which explores the development and selfdevelopment of capital in general karl marx introduces two-dimensional measurement of capital in its form of money and capital in the form of its goods. these measures or economic growth measurement “rulers” are being presented by working time in form of man-hours and money in form of the national currency of each country. marx’s summary with respect to the adam smith’s one-dimensional approach is the following: “he does not distinguish the dual nature of labour itself, i.e.does not distinguish between labour,as a cost of wage creating value, and labour, as a specific and useful labour creating commodities (use-value). the total amount of goods produced in a year i.e. an entire annual product is a product of useful labour active during the past year; all these products exist only due to the fact that publicly employedlabourwas consumed in diversified extensive system of various kinds of usefullabour. that is the only reason the cost of the means of production has been consumed during manufacturing but kept in the total cost of manufactured commodities, preserved as reappearing in a new kind. consequently, the entire annual product is a result of the useful labour expended during advances in systems science and application (2016) vol.16 no.3 5 the year. but only a part of the annual cost of the product has being created newly; this part is the newly created cost for the year, which embodies the sum of labour,spent during the year” [3]. our studies done based oneconomic science latest achievements, in particular the nobel prize leonid kantorovich and tjallingkoopmans’ duality principle, show that adam smith is partly right in saying that “the cost of the annual product”,according to that principle, can be reduced to “the newly created cost for the year.” the adam smith’s principle of return on capital allows to estimate the cost of capital in its form of money; while karl marx is right in saying that the total costs of producing a particular product, as capital, in the form of its goods are consumed in the production of the final product “capital plus income” [4]. both of these indicators are not just theoretical, but also have a crucial practical significance. in accordance with the principle of duality, they form the basis for the twodimensional measurement of the balanced economic growth. thus, the newly created value for the year by adam smith is an exact expression of the cost of capital in its form of money. the practical implementation of this concept is found in the nominal gdp, which is the cost of the final product. and the cost of the annual product by karl marx is an exact expression of the cost of capital in its form of goods. the practical implementation of this concept is in real gdp, which is to represent the volume of actually produced final product. the conclusion is that to eliminate errors in the measurement of inflation and real growth,which is possible by adding the adam smith’s one-dimensional approach with the karl marx’s approach based on assessmentof the labour’ total costs, that are objectively necessary for the production of the given structure’s final product amount. therefore, the rationale of the new market equilibrium equation can start from that point where thomaspikettystopped his theoretical research of two-dimensional measurement: of the national income ĺcby money, and of the national capital ĺc by working time, man-years. in this case, the duality theory allows you to take advantage of the twodimensional measurement criterion. firstly, the cost of the working time is measured by estimating the product’s direct and full labour-intensity of economic activities. second, the cost of the final product (nominal gdp) and resources total costs for the production of nominal gdp are set in monetary terms. according to the duality principle, solutions for interlinked problems identified on the basis of the industries intersectional balance report will meet following criteria [5-7]: l = t ∗x = t ∗ y (b) where t is direct line and tproduct fulllabour-intensity, l the entire fund of working time, in man-years, yfinal product value (nominal gdp) and x6 s.baizakov, y.utembayev, a.r.oinarov, j.forrest:on mutual coincidence ... resources full costs for the nominal gdp production. fourthly, we introduce the extension transformation below: the tool for solving the contradictory problem is extension transformation. by using certain transformations, an unfeasible problem can be transformed a feasible one. we only introduce the general concept and the basic types of transformation. in the formula (b), each component of the full labour-intensity tof the final product value y is determined by the scalar multiplication of the product’s direct labour-intensity components t by components of each column of the full costs technology matrixb = (e a)-1, which serves as a carrier of scientific and technological progress and, therefore, the true value of goods and services. at the national level we have the equalities t = l / x and t = l / y. from here we denote the level of scientific and technological potential (stp) as c and get as follows: c = t t = y x (c) stp ratio of the country c with its growth (+,-) is determined either by the magnitude of the margin, which is correlated with changes in the prices of goods and services, or by the amount of rent which changingrate is correlated with the intensity of productive resources utilization. in any case, the effect of volume, which represents the level of scientific and technological potential of the country, is measured by the difference between the growth in labour productivity, defined by value of the final product and used in its resources production, as they have both utilized the very same equal fund of working time: c c = y/l y/l − x/l x/l . the new equation of market equilibrium derives from the basis of adding the adam smith’sprinciple of return on capital to the principle based on the ratio of direct and fulllabour-intensity, as determined by the karl marx’s theory of capital’s labour value. thus, on one hand, the new equation ofthe balanced economic growth of nominal value of the final product (ngdp = y) defined by the formula: ngdp = c ∗x. in this formula the level of national income is measured by the value of nominal gdp. it is determined by multiplying the level of scientific and technological potential (stp) of the country c by a volume of domestic capital x. here, the scientific and technological potential (stp) of the country c represents coefficient of national capital efficiency. this equation is the analog of the first fundamental advances in systems science and application (2016) vol.16 no.3 7 law of market economy development, defined on the basis of the principle of duality. on the other hand, multiplying both sides of the gdp deflator equation by the purchasing power of the national currency (not to be mixed up with purchasing power parity) pp, we have got as follows: pp ∗ngdp = pp ∗ pb ∗rgdp hence we have a qualitatively new equation of balanced economic growth, which determines the actual volume of the final product rq: rq = pp ∗ngdp = pp ∗ pb ∗rgdp (d) where ngdp, as well as previously, represents nominal gdp, which determines the cost of the final product, and its multiplication by the level of the true value of money represents the real final product. since the pp * pb is equal to c = ngdp / xas by definition the value of money pp = rgdp / x, while the gdp deflator pb = ngdp / rgdp, then the purchasing power of money (not to be mixed up with the parity of purchasing power of money) is defined by formula: pp = c pb , where the value of the ccoefficient, which represents the level of scientific and technological potential of the country, is determined by the ratio of direct labourintensity to its full labour-intensity (t / t). this ratio is determined due to the principle of duality solution for interlinked problems, based on the data of given reports on the industries intersectional balance of the country; and it is equal to the ratio of the value of the final product cost used in the country to the total consumption of resources used for its production (y / x). to change economic policy, aimed at establishing a real growth index of prices of goods and services, and the level of depreciation of the national currency, is doable due to the system of economic laws and regulations, which are derived from combining the adam smith’s principle of one-dimensional measure of return on capital to the karl marx’sprinciple of two-dimensional measurement of return on labourand capital. the main ones, which are caused by the definition of the true cost of price indices of goods and services by the purchasing power of money, are as follows: 1) the law on determining overall impact of the takenincentives on innovative investmentsinto the economy and on scientific and technological improvement of the production process, including the strength of competition c(t): c(t) = ngdp/x(t) ̸= 1, 8 s.baizakov, y.utembayev, a.r.oinarov, j.forrest:on mutual coincidence ... though on the basis of return on equity ratio by adam smith an stp coefficient: c (t) = pp ∗ pb = 1,because by definition, the value of money pp = rgdp/ngdp , and the gdp deflator pb = ngdp/rgdp . 2) the basic law on determiningreal purchasing power of money, as the ratio of the scientific and technological potential of the country to the gdp deflator pp(t): pp (t) = c (t) /pb (t) . 3) the law on determining prices of goods and services pc(t) = 1/pp(t): pc (t) = 1/pp (t) = ngdp (t) /(c (t) ∗rgdp (t)). 4) leading law on determining real growth of the final product on the basis of nominal gdp growthrq(t): rq (t) = pp (t) ∗ngdp (t) . 5) the control law on determining real growth of the economy, on the basis of real gdp growth rq(t): rq (t) = c (t) ∗rgdp (t) . 6) the law on defininga gdp deflator, the general price deflation of goods and services pb(t): pb (t) = c (t) /pp (t) = ngdp (t) /rgdp (t) . 7) the law on determining net benefits of promoting scientific and technological advances and competition (±△c (t)): ±△c (t)% = (c (t) /c (t− 1))%− 100%. since any incentives on scientific and technological improvements have certain expenses of costs of labour, capital and materials then their return is associated with risks, and return on these funds can be a value greater than 100 and less than 100. however, where the competitive environment is developed there the contribution of science and technology potential into the development of the real economy, as a rule, would be positive. the exceptions are those developed countries where sources of reduction of material resources through scientific and technological advances have been exhausted; and those countries trying to maintainthe efficiency of their economies at the expense of reducing their national currenciespurchasing value. however, these attempts are easily detected by a tool developed with the help of a system of laws to promote competition in the market economy. advances in systems science and application (2016) vol.16 no.3 9 the path to sustainable economic growth goes through the open labour market, focused on the final product production. thus, last five six years the actual “restoration” of the old world monetary and financial system by transferring private risk to the global (system-wide) level, and by the public debt exponential increase in a number of leading countries and the money supply, makes the issue of transition to a new world economic order more difficult. the problem of transition from the current crisis and unstable state of the world economy to a sustainable growth mode should be investigated in the system, as shown above, in the unity of the macroeconomic, sectoral and microeconomic aspects of market economy development based on feedbacks that exist between the various regulatedagents, and of patterns of the long-term economic development countries of the world. this approach significantly expands the understanding of the causes of the global financial crisis, allowing to realise mechanisms of modern economy reproduction and to develop reliable tools for its sustainable development. thus, the history of economic thought tells that objective laws of market economy can justify new, adequate to the new world economic order, models of economic governance, by which it will be able to ensure the unity of the three product types: nominal, real and final. by and large, the core to ensure the unity of three gdp growth rates is the final product, the real value of which is determined by the equation (d) under the influence of the real purchasing power of the national currency at nominal gdp and the rate of scientific and technological potentialat the real gdp. according to this model of balanced economic growth the procedure of targeting nominal gdp is implemented instead of inflation targeting’monetary policy.this is the tate-shafarevich group: xe/k . here it is as a subscript: ax . references [1] piketty and thomas(2014),capital in the twenty-first century,the belknap press of harvard university press. [2] smith and adam(2007),an inquiry intothe nature and causes of the wealth of nations,moscow eksmo. [3] marx and karl(1978),capital,critique of political economy:the process of circulation of capital,vol.iv, 648 p. [4] marx and karl(1975),capital,critique of political economy:the process of capitalist production as a whole, vol.iii, part 1. [5] y. hasanli, s. bayzakov, v. valiyev and g. sarsembaeva. (2011), “modeling 10 s.baizakov, y.utembayev, a.r.oinarov, j.forrest:on mutual coincidence ... of the multiplicative effects of opening of the work places on the bases of ‘intersectoral labor balance’ (on example of azerbaijan and kazakhstan)”, ecomodazkz_full pape eng. [6] s. baizakov(2014), “model of the labour market, based on the final product”, science and innovation, no.12, pp.31-34. [7] s. baizakov(2015), “declaration on market laws of economic development of world countries”, veo of russia, no.195, pp.143-156. corresponding author sailau baizakov can be contacted at: baizakov37@mail.ru advances in systems science and applications (2013) vol.13 no.3 198-217 currency wars and a possible self-defense (i): how currency wars take place jeffrey forrest1, zack hopkins2 and sifeng liu3 1department of mathematics, slippery rock university, slippery rock, pa 16127, usa 211934 route 89, wattsburg, pa 16442, usa 3school of economics and management, nanjing university of aeronautics and astronautics, nanjing 210016, pr china abstract going along with the common knowledge that money can easily destroy a person, a family, and a healthy business enterprise, this work investigates the nature of currency wars in two parts. the first part shows, by making use of the systemic yoyo model and bernanke-gertler model of fundamental value of capital, how money can be purposefully and strategically employed as a weapon of mass destruction. based on how a currency war could be potentially raged against a nation, the second part of this work makes use of the results of feedback systems to develop a self-defense mechanism that could conceivably protect the nation under siege. keywords systemic yoyo, foreign investment, financial crisis, financial liberalization, speculative attack 1 introduction the main focus of this work is on financial crises that are caused and/or created purposefully by certain group(s) of people in order to acquire economic and political gains. here, financial crises include all those crises with respect to either respectively or jointly currency, credit, bank, debt, and the markets of stocks, bond, and financial derivatives, etc. they generally mean severe fluctuations and chaos that appear in the financial area of a nation, interfering very negatively with the operation of the real economy. each of these crises is accompanied by sudden deterioration of all or most financial indexes, stock market crash, capital flight, credit destruction, extremely tight money supply, rising interest rate, bank runs, bankruptcy of a large number of financial institutions, major decrease in the official reserves, inability to repay the interest and principal of maturing debt, currency devaluation both internally and externally, etc. because the modern-day national economies have been closely intertwined with each other to a high degree of globalization, when a nation or geographic region suffers from economic and financial crises, in terms of their damaging effects and adverse impacts, the crises tend to possess international and global characteristics. since the 1920s, there had appeared many large scale financial crises. the 199 advances in systems science and applications (2013) vol.13 no.3 most noteworthy of these crises include 1) the global stock market crisis that started on october 28, 1929, in the new york stock exchange and spread to other nations quickly, leading to a worldwide financial and economic crisis until 1933; 2) six us dollar crises one after another during the time period from 1960 to 1973; 3) a bank bankruptcy wave that started with the failure of the united states national bank in san diego and quickly spread to many other western countries during 1973-1975; 4) the debt crisis that broke out in august 1982 and quickly involved over 50 developing countries from around the world; 5) the international stock market crisis that started with the plummet of the dow jones industrial average on october 19, 1987; the consequent chaos of the wall street momentarily spread throughout all the major stock exchanges of the western world; 6) the british pound crisis of 1992; 7) mexican financial crisis of 1994; 8) the southeast asian financial crisis of 1997; 9) the russian currency, financial crisis of 1998; 10) the brazilian currency, financial crisis of 1999; 11) the argentinean currency, financial crisis of 2001. this work makes use of the systemic yoyo model and bernanke-gertler model of fundamental value of capital to show that when a large amount of foreign investments gathers in one place over either a long period or a short period of time and then leaves suddenly and massively, that local economy has to suffer through a positive bubble, caused by the increased money supply as a consequence of the foreign investments, and then a following negative, disastrous bubble, caused by the sudden dry-out of the money supply. because a large number of economic activities are either unexpectedly delayed or totally impossible to complete, the local investors are actually unable to continue to collect their originally expected dividends for many time periods to come. in other words, foreign investments can be employed as a weapon of mass destruction, if they leave strategically and suddenly, no matter whether they come quickly in a short period of time or slowly over a relatively longer period of time. with the conclusion that money can be employed as a weapon of mass destruction is established, this work then is turned to design a self-defense mechanism for the purpose of eliminating or greatly reducing the disastrous aftermath of such currency warfare. this work is organized as follows: in the first part, with relevant concepts systematically gathered in section 2, recent speculative attacks and currency crises jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 200 are compared in section 3. by focusing on the analysis of the fundamental value of capitals, section 4 shows that how potentially currency wars can be launched purposefully. in the second part we present a possible way to build up a national defense against possible currency wars. then a few final remarks conclude this presentation. 2 the basic concepts and systemic intuition the general concept of financial crises includes the following four classes of crises: currency, bank, foreign debt, and systems. here, by currency crisis it represents such a situation that due to purposeful and targeted speculative activitiesagainsta nation, the nation’s currency suffers from drastic devaluation, or the nation’s government is forced to drastically increase its interest rate or spend a large amount of the foreign reserves to defend its currency. by bank crisis, it means actual and/or potential bank runs or such a scenario that a number of banks stop repaying their debts because of their falling into bankruptcy or the government is forced to interfere by providing large amounts of support. by foreign debt crisis it implies such a case that a nation can no longer repay its foreign debts on time, no matter whether the debtors are governments or private individuals. by systematic financial crisis, it stands for the destructive effects on the real economy due to severe damages of the financial infrastructure so that the efficiency of the financial markets is greatly affected. the connotation of systematic financial crises might overlap with those of other kinds of crises, while currency and bank crises might not lead to severe damages to a nations payment system. so, neither currency crises nor bank crises can be identified with systematic financial crises. in many circumstances, the specific definition of financial crises simply means currency crises. the early investigation of currency crises can be traced back to at least [1]. paul krugman treats a currency crisis as an internationalbalance-ofpayments crisis. he believes that in order to prevent their currencies to devaluate, countries with either a fixed exchange rate system or pegged exchange rate system would pay the price of either spending their international reserves or increasing inflation due to their raising domestic interest rates; when the governments give up on the fixed exchange rate system or the pegged exchange rate system, their currencies experience drastic drops in value, leading to currency crises. currency crises are generally indicated by the collapse of the fixed exchange rate system or forced adjustment to the system, such as official devaluation of the local currency, expanded floating range of the exchange rate, drastic decrease in international reserves, noticeable rise in the interest rate of local currency, etc. according to the literature, there are four main criteria for judging currency crises: (1) sudden and large scale changes in exchange rate; (2) the weighted average of exchange rate and foreign reserves fluctuates widely; (3) the weighted 201 advances in systems science and applications (2013) vol.13 no.3 average of exchange rate, foreign reserves, and interest rate vibrates wildly; (4) import drops drastically. the second criterion was established by kaminsky, lizondo and reinhard who believe that a currency crisis represents such a scenario that is caused by either devaluation of the nation’s currency or drastic drop in the nation’s international reserve or both as a consequence of an attack on a nation’s currency [2]. because of this reason, currency crises can be verified afterward by using the index named exchange market pressure (emp). this index stands for a weighted average between the monthly percentage change in the exchange rate of the local currency and that of the international reserves.because this index increases with the devaluation of the local currency and the loss of international reserve, major increases in this index indicate a strong pressure to sell off the local currency. because this second criterion possesses very practical operationality, it has been employed most widely among these four criteria. most of the financial crises that occurred since the 1990s have brought along with them clear characteristics of currency crises; and another outstanding feature of recent currency crises is that they have been accompanied by bank crises. such scenarios are referred to as twin crises. since the 1980s, along with the gradual strengthening of globalization of the capital markets, international flows of capital have been developed unprecedentedly with ever increasing magnitude, velocity, and accompanying dangers and risks. large amounts of short-term capitals that are not under the watch of any national government and international financial organization have been moving freely in the international financial markets for pursuing profit opportunities by making use of various new financial tools, trading platforms,and advanced trading technologies. these new characteristics of the international capital have led to frequent occurrence of turmoil in the international financial markets; speculative attacks happen often with ever increasing strength and duration of impact. the asian currency crises that started in 1997 further indicate that the currency crises caused by speculative attacks can also possibly develop into full scale financial crises and deepening social crises of large magnitude. the so-called international speculative capital or hot money represents such capital that is frequently moved within and between various markets in pursuit of short-term, high levels of profits without any particular fields of investment focus. speculative capital tends to be short term even though there are exceptions to this rule of thumb. one of the modern characteristics of international speculative capitals is their camouflage. at the same time, these capitals can also go along with the market cycles by pursuing midand long term investments. additionally, not all short-term capitals are speculative. for example, the short-term capitals involved in the financial intermediation and settlement of international trades, short-term interbank funds, banks’ short-term positions for allocation, etc., are jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 202 not speculative in nature. along with the expansion in the size, circulation speed, and coverage of the international capital markets have international speculative capitals grown. based on their predictions on the changes in exchange rate, interest rate, security prices, gold price, or the prices of certain commodities, speculative capitals could be suddenly involved in large scale in both long and short trades in a short period of time. by substantially altering the composites of their portfolios and by affecting the confidence of the holders of other assets, these speculative capitals cause severe instability in the market prices so that short-term high profit opportunities could be created. such market behaviors that disturb the market prices and appear suddenly are known as speculative attacks. as limited by the constraint of pursuing after quick gains, international speculative capitals generally choose to attack such economic sectors and geographic areas that can hold large amount of capitals and allow fast movements of funds with expected high returns, and few financial regulations. these economic fields include (foreign) currencies, futures, options, precious metals, real estates, etc. from what has happened empirically in the past, it seems that international speculative capitals have been fond of attacking the currency of one nation or the currencies of several nations at the same time. most often seen speculative attacks are those that assault either fixed exchange rate systems or regulated exchange rate systems. speaking generally, when a nation either employs a fixed exchange rate system or pegs a target exchange rate, the speculators often make their judgment that as long as the official parity or the target exchange rate does not conflict with the fundamental conditions and states of the economy, the official exchange rate will be maintained. however, if the speculators believe that the fundamental states of the current economy could not sustain the prevalent level of exchange rate for long, they would launch a speculative attack in order to speed up the dissolution of the fixed exchange rate system so that opportunities of quick profits would be generated. under the conditions of either fixed or pegged exchange rate system, as long as there appears either a domestic inflation or recession accompanied with sustained current account deficits, the governmental promise on the fixed exchange rate would lose its reliability. it is because in these situations there is a heavy pressure to devaluate the local currency; in order to maintain the promised exchange rate, the government will be forced to mobilize and spend its international reserve. even with the support of the international financial markets, the fundamental imbalances existing in the economic states still cannot be corrected, which can most likely delay the occurrence of the devaluation of the local currency, although the devaluation will sooner or later happen inevitably. if speculators predicted this forthcoming event, they would mobilize their capitals ahead of 203 advances in systems science and applications (2013) vol.13 no.3 time and launch their speculative attack in order to position themselves for quick profits. by making use of the spot and forward transactions, futures contracts, options trades, and swaps of various financial tools, speculators carry out their multi-dimensional speculative strategy by positioning their capitals at the same time on the markets of foreign exchanges, securities, and all different forms of financial derivatives. because of the fixed exchange rate or the promise that the government would maintain the rate fixed, the risk to the speculators is actually quite low, because the direction along which the exchange rate would move is clear. to say the least, even if the prediction is incorrect, the worst is that the exchange rate parity did not change so that the most the speculators could lose is their minimal amounts of trade costs. hence, once a speculative wave is started, the magnitude in general is large, leading to the expected consequences, as a self-fulfilling prophecy of a humongous scale. 3 recent speculative attacks and currency crises during the time period of bretton woods system after world war ii, the strength and power of private capitals grew drastically; relevant speculative activities evolved with increasing levels of energy. their attacks on various national currencies from around the world were mostly successful and amplified with evergrowing vigor and intensity. the most typical are the british pound crisis of the late 1967, french franc crisis of august 1969, and the u.s. dollar crisis of 1971-1973. if we say that the root problem for bretton woods system to eventually collapse were the defects of the system itself, then the direct triggering factor for the system’s collapse would be the speculative attack of the international short-term hot money on the u.s. dollar-the base currency. when bretton woods system was over, the world was in a wave of deregulation, strengthening the market mechanism, promoting economic and financial liberalization. correspondingly, the international financial markets become further liberalized and global. along with the application of modern technology of communication and computer networks, financial derivatives and methods of trading are developed in abundance. all these political, societal, and technological advances provided the space for international capitals to grow and to be mobilized unprecedentedly. with their greatly increased speed of mobility, international capitals have launched frequent speculative attacks. among the most typical are the attacks on the pegged exchange rate system employed by some countries of latin america in the early 1980s, the mexican peso crisis of 1994, and financial crises of eastern asia during 1997-1998. the following provides a list of recent speculative attacks. for more details please consult with wang and hu [3]. case 1: the attacks on the exchange rate system of latin american countries in the early 1980s jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 204 in 1978, chile, uruguay, and argentina decided to employ a crawling peg exchange rate system. each of these national governments established its plan to gradually depreciate its local currency against u.s. dollar. however, in their implementations, their rates of inflation were much higher than that in the u.s.a., while their degrees of depreciation were much smaller than the difference of the u.s. inflation rates. therefore, the over-evaluations of their local currencies made the deficits of their current accounts rise. during 1981-1982, the interest rate in the international financial markets reached an historical high, making the burdens of foreign debts and the deficits of the current accounts of these three countries difficult to sustain, so that it became inevitable for the local currencies to devaluate while departing from the targeted exchange rates. under this background, these countries respectively experienced speculative attacks, corresponding currency devaluations and the consequent capital flights, and crises of domestic financial institutions’ runs. case 2: mexican peso crisis of 1994 in 1982 after having suffered from its debt crisis, under the supervision of the imf mexico implemented a comprehensive policy for economic adjustment and reform, while tightening its economy and dramatically reducing its fiscal deficit. in 1987, mexico re-fixed its exchange rate between mexican peso and u.s. dollar. in january 1989, mexico started to employ a crawling peg exchange rate system, which was changed to a moving target regional exchange rate system in december 1992, while gradually expanding the floating range for peso. this series of measures of economic reform achieved a certain degree of success; the national economy steadily recovered. however, in 1994, mexican economy once again stalled while accompanied by political instability. therefore, the expectation and rumor for peso to depreciate grew; capitals fled one after another. interventions of the central bank made the market interest rate rise drastically, while the national foreign reserves were depleted quickly. on december 30, mexican government eventually had to allow peso to depreciate. however, the new exchange rate established after the depreciation immediately suffered from speculative attacks so that mexican government had to implement a floating exchange rate system. after then, the domestic economic conditions and the political situation made foreign investors extremely nervous, causing continued capital flight, banks subjected to runs, and the economy falling into crises. in the newly adopted floating exchange rate system, peso continued to depreciate; until the end of 1995, peso had reached consecutive historical lows one after another. case 3: the speculative attack on the joint floating mechanism of european monetary system (1992-1993) the national currencies within the european monetary system of the nations that were members of the european community had followed a joint floating ex205 advances in systems science and applications (2013) vol.13 no.3 change rate mechanism. this mechanismhad led to the creation of the european currency unit (ecu) and established the statutory central exchange rates between the individual national currencies and the european currency unit. therefore, between these member nations a system of fixed exchange rates was employed, while externally a joint floating rate system was implemented. in the early 1990s, the member nations of the european community experienced varied degrees of economic turmoil. uniformities in their respective macroeconomic states, such as inflation rates, unemployment rates, fiscal deficits, economic growth rates, etc., started to be broken. over time, it became clear that some member nations could no longer maintain their statutory exchange rates with ecu, providing opportunities for the international speculative capitals. in the late 1992, an initial round of speculative attacks appeared. among the first group of victims of the attacks were finnish marks and swedish krona. at the time, neither finland nor sweden was a member of the european community. however, they all hoped to join so that they voluntarily pegged their own currencies with ecu. under the speculative attacks, finland quickly gave up its fixed exchange rate and drastically depreciated its currency on september 8. on the contrary, swedish government decided to protect its krona by raising its short-term interest rate to 500% annually. that eventually defeated the speculative attacks. at the same time, british pound and italian lira also suffered from continued attacks. on september 11, european monetary system agreed for lira to depreciate 7%. although german central bank spent around 24 billion marks to support lira, three days later lira was still forced out of the european monetary system. by this time, the bank of england had lost several billions of u.s. dollars to protect its pound. even so, on september 16, the british pound was still forced to float freely. although french franc also suffered from speculative attacks, through joint interventions with germany and by greatly increasing the interest rate, the value of franc recovered. this crisis of european monetary system started in 1992 and lasted until 1993, during which speculative attacks often occurred. at the end of 1992, portuguese currency escudo depreciated; spanish currency pesetas was devalued once again, while swedish krona and norwegian krone started to float. in the early part of 1993, ireland’s pound depreciated, portuguese escudo depreciated another time, and spanish pesetas experienced its third round of depreciation. on the other side, french franc and danish kroner successfully stood against the sporadic speculative attacks. case 4: 1997 currency crises of east asia (1997-1998) in the 1980s and early 1990s, nations in southeast asia sped up their steps toward financial liberalization by completely opening up their domestic financial markets to attract maximal scales of foreign investments. such fast-speed jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 206 liberalization led to drastic economic growth, which was known as “southeast asia miracle”. however, after entering the mid-1990s, the rising labor costs were translated into the decreased international competitiveness in their products so that deficits began to appear in the current accounts of some southeast asian countries. because these countries did not in a timely basis upgrade their industrial structures in order to keep pace with the increasing competitiveness of their products, the continued influx of the foreign capitals together with domestic investments led to the formation of economic bubbles and overheated sector of real estates. for example, in 1996, thailand’s balance of foreign debts had reached over 90 billion u.s. dollars with more than 40 billion dollars of shortand mid-term foreign debts, both of which surpassed the corresponding levels of foreign reserves at the start of 1997. additionally, because of the overheated investments, particularly the overheated investments in real estates, the bad debts of thai financial institutions had amounted to more than 30 billion u.s. dollars in early 1997. therefore, the public and foreign investors started to worry about the economic conditions and financial order in thailand, which inevitably helped to consolidate the expectation for baht to depreciate. at the same time, international speculators were also building up their monetary energy and preparing to launch their large scale attacks. on february 14, thai currency baht depreciated 5% against u.s. dollars, making the covered speculative attacks public. after then, baht suffered from ever increasing pressure to depreciation further; and interventions of bank of thailand were quickly exhausting the national foreign reserves. after mid-may, speculative capitals launched their new rounds of attacks, creating an eleven-year new low for the exchange rate of baht. several nations in southeast asia jointly intervened in the foreign exchange markets by buying in baht, while bank of thailand sacrificed 5 billion u.s. dollars of foreign reserves and once again raised its short-term interest rate. however, all of these still could notrebuild the public confidence and drive off the speculative attacks. eventually on july 2, bank of thailand was forced to allow baht to float freely in the exchange markets, causing baht to depreciate 20% against u.s. dollar on that single day. subsequently, the speculative attacks quickly spread over to the neighboring nations and regions so that philippines, malaysia, indonesia, singapore, hong kong, and taiwan were all affected. other than hong kong, the local currencies of these countries and regions all depreciated against u.s. dollar with different scales. at the same time, all these nations and regions except singapore and taiwan fell into deep financial and economic crises. after october of the same year, the crises spread over to south korea, causing south korean currency won to depreciate deeply against u.s. dollar and making the economy of south korea fall deeply into an economic crisis. case 5: russian currency crisis (1998) 207 advances in systems science and applications (2013) vol.13 no.3 at the early stage of russian economic transition, large amounts of international capitals entered russia. as of july 1, 1997, the accumulated foreign investments totaled 18 billion u.s. dollars, about 10 billion dollars of which were short-term capitals and invested in the securities markets. although the economic growth of the real economy was nearly zero, the stock market rallied rapidly in the first half of 1997; bubble expansion appeared in the prices of financial assets. when the currency crises of southeast asia broke out and started to spread toward the regions of northeast asia, the probability for russia to experience a similar currency crisis was very big, considering the market long-term expectation of instability for russian economy. starting in november 1997, speculators launched their attack on ruble with many others followed. during the end of 1997 and early part of 1998, russian government and central bank sacrificed foreign reserves to purchase ruble, while expanding rubles floating range and increasing the interest rate from 21% to 35%. although the situation was temporarily stabilized, large amounts of foreign reserves were lost. however, after may, another wave of speculation started, causing rubles exchange rate to fluctuate severely again. though russian government raised the interest rate to 150% and sought for international assistance in order to counter the attack, the situation was never successfully under control. on august 17, russian government had to expand ruble floating exchange corridor; that is, allowing ruble to depreciate. however, unexpectedly, the situation continued to worsen. eventually on september 2 without a choice the exchange corridor was abandoned. that meant that russian government gave up its over-two-year old managed target floating mechanism of ruble. case 6: brazilian currency crisis (1999) the currency crises that appeared in east asia and russia increased the market expectation for brazilian currency to depreciate, causing large amounts of capital leaving brazil since september 1998. at the end of 1998, brazilian congress did not pass the bills to increase the benefit taxes of civil servants and those of retired civil servants as contained in the fiscal adjustment plan, while the 1998 deficits of the fiscal balance, trades, and current accounts exceeded the expected levels. so, along with the decreasing market confidence, speculative attacks were triggered, forcing brazilian central bank to allow its real to float freely against u.s. dollar. all the speculative attacks that were launched since the 1980s have shown a good number of new characteristics, including amplifying scales and forces, enhancing dimensionality in terms of their strategies of attack, covering larger geographic areas, increasing publicity of planned attacks, etc. in particular, the main reasons for the amplifying scales and forces of the speculative attacks include: 1.the size of the international speculative capital has been increasing conjeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 208 stantly. since the 1980s, the economic growth of the main industrialized nations has been slowed, while the reform for financial liberalization of various countries has been strengthened. that caused large amounts of capitals toenter into the international financial markets to look for new opportunities. such continuous accumulation of international speculative capitals, together with the multiplying effect of international credit, makes the available speculative capitals increase multiple times. based on the relevant estimates of the international monetary fund, there were at least 7,200 billion u.s. dollars of speculative capitals currently floating around the international financial markets. that amount is equivalent to 20% of the world gdp; and each day, over 1,200 billion u.s. dollars of speculative capitals are looking for various profit opportunities. that is nearly one hundred times more than the amount of trades of real consumable goods. 2.the capitals available for speculation purposes have shown the tendency to collaborate and to take actions jointly. along with the rapid development of communication technology and the expanding availability of internet, a 24/7 operational system for the globalized exchange markets has appeared so that capitals can be transferred from one exchange market to another instantly. that makes the once stranglers and disbanded international speculators develop into powerful speculation assemblages. as a third force different from various national currency authorities and international financial organizations, they have constituted a major thread to the stability of each nation’s exchange system and the normal operation of the international currency system. 3.the introduction of financial derivatives provided a leveraged platform of trading for the speculators. since the 1970s, financial innovations have led to the profligate development of new financial derivatives and their trades. because of the characteristic of high leverages of the financial derivatives, speculators can trade such financial products that are valued tens or even hundreds of times of the little amounts of capital they mobilized. that enables a hedge fund to embark on trades worthseveral hundreds of billions of u.s. dollars with a small amount of capital, affecting the entire international financial markets. since the 1990s, the strategies of speculative attacks have been further developed. the traditional speculator would simply use the price differences between spot and future trades, while the current strategies of speculative attacks have been quite complicated. by utilizing all kinds of available financial tools, the speculator gets involved in various markets of the traditional tools and derivatives. it can be said that current speculative attacks make use of the inherent linkages among the prices of various financial products, traded in different markets, to make comprehensive and profitable arrangement of capitals. traditional speculative attacks have been isolated and scattered in different geographic regions. however, along with the recent deepening globalization 209 advances in systems science and applications (2013) vol.13 no.3 of the international financial markets, and along with the further unification of regional economies, speculative attacks have also showed clear regional characteristics. no matter whether it is the latin american currency crisis of the early 1980s, the mexican crisis of 1994, the european monetary system crisis of 1992, or the southeast sian crisis of 1997, each started with the initial shocks in such a market that contained the most concentrated imbalances. then, the initial shocks spread over to other neighboring markets, displaying a dynamic process of mutual influences. traditional speculations used to be hidden or semi-public arbitrage activities. however, since the 1990s, along with the deregulation and financial liberalization of the international financial markets and the widening use of information and network technologies, originally speculative activities have been gradually evolved into open and purposeful attacks on specified currencies. such openness and publicity could be and have been strengthening the market expectation of depreciation of the target currency. for example, at the end of 1997, george soros published an article to openly declare that he and associates would launch another attack on hong kong dollars. after that, he announced in newspapers that the next currency crisis would appear in russia, which was later shown to be accurate. soon after then, soros commented that brazilian currency was evaluated too high so that brazilian real would be his next target of attack, which was indeed what happened next. before we start to develop and present our main results, let us first look at the systemic intuition that lies underneath this work. when von bertalanffy pointed out that the fundamental character of living things is its organization [4], the customary investigation of individual parts and processes cannot provide a complete explanation of the phenomenon of life, this holistic view of nature and social events has spread over all corners of science and technology [5]. accompanying this realization of the holistic nature, in the past 80 some years, studies in systems science and systems thinking have brought forward brand new understandings and discoveries to some of the major unsettled problems in the conventional science [6-7]. similar to how numbers are theoretically abstracted, systems can also be proposed out of any and every object, event, and process of concern. for instance, behind collections of objects, say, apples, there is a set of numbers such as 0 (apples), 1 (apple), 2 (apples), 3 (apples), ...; and behind each organization, such a regional economy, there is an abstract, theoretical system within which the relevant whole, component parts, and the related interconnectedness are emphasized. and, it is because of these interconnected whole and parts that the totality is known as an economy. in other words, when internal structures can be ignored, numbers can be very useful; otherwise the world consists of dominantly systems jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 210 (or structures or organizations). historically speaking, on top of numbers and quantities has traditional science been developed; and along with systemhood comes the systems science. that jointly gives rise of a 2-dimensional spectrum of knowledge, where the classical science, which is classified by the thinghood it studies, constitutes the first dimension, and the systems science, which investigates structures and organizations, forms the genuine second dimension [8]. the importance of the systems science, the second dimension of knowledge, cannot be in any way over-emphasized. for example, when there are difficulties in studying dynamics in an n-dimensional space, one can conveniently get help from a higher-dimensional space. in particular, when a one-dimensional flow is stopped by a blockage located over a fixed interval, the movement of the flow has to cease. however, if the flow is located in a two-dimensional space, instead of being completely stopped, the 1-dimensional blockage would only create a local (minor) irregularity in the otherwise linear movement of the flow (that is how nonlinearity appears [9]. additionally, if one desires to peek into the internal structure of the 1-dimensional blockage, he can simply take advantage of the second dimension by looking into the blockage from either above or below the blockage. that is, when an extra dimension is available, science will gain additional strength in terms of solving more problems that have been challenging the very survival of the mankind since the start of history. additional to the afore-described strong promise of systems science, on the basis of the blown-up theory [10], the systemic yoyo model, fig.1, is formally introduced by lin for each and every system [11], be they tangible or intangible, physical or intellectual. (a) (b) (c) fig.1 the eddy motion model of the general system in particular, on the basis of the blown-up theory and the discussion on whether or not the world can be seen from the viewpoint of systems [12-13], the concepts of black holes, big bangs, and converging and diverging eddy motions are coined together in the model shown in fig.1. in other words, each system or object 211 advances in systems science and applications (2013) vol.13 no.3 considered in a study is a multi-dimensional entity that spins about its either visible or invisible axis. if we fathom such a spinning entity in our 3-dimensional space, we will have a structure as shown in fig.1(a). the side of black hole sucks in all things, such as materials, information, energy, profit, etc. after funneling through the short narrow neck, all things are spit out in the form of a big bang. some of the materials, spit out from the end of big bang, never return to the other side and some will (fig.1(b)). for the sake of convenience of communication, such a structure as shown in fig.1(a), is referred to as a (chinese) yoyo due to its general shape. more specifically, what this systemic model says is that each physical or intellectual entity in the universe, be it a tangible or intangible object, a living being, an organization, a culture, a civilization, etc., can all be seen as a kind of realization of a certain multi-dimensional spinning yoyo with either an invisible or visible spin field around it. it stays in a constant spinning motion as depicted in fig.1(a). if it does stop its spinning, it will no longer exist as an identifiable system. what fig.1(c) shows is that due to the interactions between the eddy field, which spins perpendicularly to the axis of spin, of the model, and the meridian field, which rotates parallel to axis of spin, all the materials that actually return to the black-hole side travel along a spiral trajectory. to show this yoyo model can indeed, as expected, play the role of intuition and playground for systems researchers, literature [5, 9] have successfully applied it to investigate newtonian physics of motion, the concept of energy, economics, finance, history, foundations of mathematics, small-probability disastrous weather forecasting, civilization, business organizations, the mind, among others. now, what is important to our work in hand here, which constitutes the necessary intuition for our reasoning, is that the systemic yoyo model implies that each economy is a system so that it can be investigated as a pool of rotational fluids of information, knowledge and money. and when the world economy is concerned with, we really have a theoretical ocean of rotational pools (of fluids) that interact with other. with time, some regional pools are destroyed while some stronger, more powerful spin fields are formed. in other words, the so-called speculative attacks, as described earlier, are natural phenomena that appear along with the globalization of the world economy. so, one natural question is how such attacks appear and how their damaging effect can be maintained at the theoretical height first and then at the level of real life practice. in the rest of this work, we will address this question by using some of the recent results of systems research. 4 one possible form of currency wars according to literature[14], the fundamental value of a particular capital is equal to the present value of the dividends the capital is expected to generate throughout the indefinite future. symbolically, the fundamental value qt of a depreciable jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 212 capital in period t is given by qt =et( ∞∑ i=0 [ (1− δ)id(t+1+i)∏i j=0r q t+1+i ]) =et( dt+1 rq t+1 + (1− δ)1dt+2 rq t+1r q t+2 + (1− δ)2dt+3 rq t+1r q t+2r q t+3 + ...) (1) where et stands for the mathematical expectation as of period t ,δ the rate of physical depreciation of the capital, dt+i the dividends and rq t+1 the relevant stochastic gross discount rate at t for dividends received in period t+ 1. then, we can rewrite equ. (1) as follows: qt = et( dt+1 + (1− δ)qt+1 rq t+1 ) (2) because of various reasons, such as fads, the market price st of the capital differs persistently from the capitals fundamental value qt. when st ̸= qt, we say, as in literature [14], there is a bubble. however, to be more specific, when st > qt, we say that there is a positive bubble in the market place; and a negative bubble, when st < qt. in the realistic market place, asset prices mostly like deviate from the fundamental values due to various reasons, such as liquidity trading or to waves of alternating optimism or pessimism. if a bubble exists at period t with probability p to persist into the next period, then by using the mathematical expectations the difference between the market price and the fundamental value of the capital in period (t+ 1) satisfies the following: p(st+1 −qt+1) + (1− p) · 0 = a[(st −qt)r q t+1] + (1− a) · 0 (3) it means that the mathematically expected (st+1 −qt+1) value with probability p for st+1 ̸= qt+1 to happen is equal to the expected growth of the t-period difference (st −qt)r q t+1 with probability a(> p) for st −qt ̸= 0 . that is, what is expected is a more severe “bubbl”, since a/p > 1 . so, if we assume a/p < 1 , it means that the bubble in period (t+ 1) is expected to be less severe than in period t. from equ.(3), it follows that we know the following expression is true: p · st+1 −qt+1 rq t+1 = a(st −qt) (4) now by taking the mathematical expected value for in period t, we have the following: et( st+1 −qt+1 rq t+1 ) = p( st+1 −qt+1 rq t+1 ) + (1− p)· = a(st −qt) (5) 213 advances in systems science and applications (2013) vol.13 no.3 equ.(2) implies that qt =et[ dt+1 + (1− δ)st+1 − (1− δ)st+1 + (1− δ)qt rq t+1 ] =et[ dt+1 + (1− δ)st+1 rq t+1 − (1− δ) st+1 −qt+1 rq t+1 ] =et[ dt+1 + (1− δ)st+1 rq t+1 ]− (1− δ)et st+1 −qt+1 rq t+1 =et[ dt+1 + (1− δ)st+1 rq t+1 ]− (1− δ)a(st −qt) therefore, we have qt + (1− δ)a(st −qt) =et[ dt+1 + (1− δ)st+1 rq t+1 ] which is the equivalent to: st[qt + (1− δ)a(st −qt)] st =et[ dt+1 + (1− δ)st+1 rq t+1 ] by isolating the factor in the numerator and cross multiplying the rest from the left hand side onto the right hand side produce the following: st =et [dt+1 + (1− δ)st+1]st rq t+1[qt + (1− δ)a(st −qt)] =et [dt+1 + (1− δ)st+1]st rq t+1[a(1− δ)st − a(1− δ)qt] +qt =et [dt+1 + (1− δ)st+1]st rq t+1[bst + (1− b)qt] , where b = a(1− δ) =et dt+1 + (1− δ)st+1 rq t+1[b+ (1− b)qt st ] = et dt+1 + (1− δ)st+1 rs t+1 therefore, we have derived at st = et dt+1 + (1− δ)st+1 rs t+1 (6) where the return on stocks rs t+1 is related to the fundamental return of the capital rq t+1 as follows: rs t+1 = rq t+1[b+ (1− b) qt st ] (7) jeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 214 if st > qt , then rs t+1 < rq t+1 , meaning that the expected stock return st = et dt+1+(1−δ)st+1 rs t+1 is less than the fundamental return of st = et dt+1+(1−δ)st+1 rq t+1 . from equ.(3), it follows that st −qt = p a × st+1 −qt+1 rq t+1 (8) which implies that at period t+ 1, there is a positive bubble if and only if there is a positive bubble at period t. because 0 < p < a < 1, if st+1 > qt+1 (assuming a positive bubble), then from 0 < p a < 1 , equ.(8) implies that the current bubble st −qt is smaller than the fundamental return of the bubble in period t+ 1. in other words, the bubble in period t+ 1 gets more severe than when in period t. now, if st+1 < qt+1 with 0 < p < a < 1, then equ.(8) implies that the underpricing st+1 of the asset qt+1 in period t+ 1 is more severe than that in period t. moreover, this equation also indicates that at period t+ 1 there is a negative bubble if and only if there is a negative bubble at period t. these two conclusions evidently contradict the efficient market hypothesis, because these analogies indicate that if in period t the market over prices the asset, then the overpricing will continue forever; and if in period t the market underprices the asset, then the underpricing will also continue forever. that is, neither positive nor negative bubble will ever crash. so there are two possibilities: (1) the model in equ.(8) does not hold true in general, or (2) the efficient market hypothesis is not ever true. evidence appears to show that the efficient market hypothesis holds true at least occasionally and also bubbles, both positive and negative, do burst [15-16]. so, the model in equ.(8) needs to be modified in order to describe the more realistic market situation better. if in equ.(8) we assume 0 < p < a < 1, then a similar analysis as above indicates that the phenomenon of overpricing or under pricing disappears over time. and, if 0 < p = a < 1 is assumed, then the model in equ.(8) implies that the existing underpricing or overpricing stays fundamentally stable. if 0 < p < a < 1 and a ≈ 0 are assumed, then equ.(8) implies that p a ≈ +∞ so that the fundamental return st+1−qt+1 rq t+1 of the asset approaches 0, meaning that the existing bubble gradually disappears with time. next, let us focus on the analysis of equs.(1) and (2). in particular, assume that in the (t+ i)-th period foreign investments are suddenly increased drastically due to expected appearance of activities from the current weak economic state and weak local currency. so for this period, δ(= the physical depreciation rate of capital) would increase due to the increased amount of money supply, which in turn pushes up the inflation. so (1− δ)i would decrease drastically if the influx of foreign investments is large. at the same time, rq t+i (= the stochas215 advances in systems science and applications (2013) vol.13 no.3 tic gross discount rate of the (t+ i)-th period) would also increase, because of the increased inflation while the dividends dt+i would generally decrease due to the reason that everybody would like to reinvest much of the available capital back into the maket in order to capture the rising book values, including stocks, real estate, and others with of impressive increasing prices. that is, the present value of the return of the (t+ i)-th period (1−δ)idt+i rq t+1...r q t+i would drop from the level of expectation. in reality, the investor would hold onto the increased book value by receiving less tangible returns, hoping that the book value would continue to rise drastically. due to the wide and conveniently availability of capital, caused by the increased money supply, the local economic activities pick up too in large quantities, while the interest rate also goes up due to the fact that the central bank, in order to control the inflation, revises the interest rate and attempts to limit the money supply. at this very moment of financially prosperity, assume that a huge amount of foreign investments suddenly leave in the (t+ j)-th period, where j > i, because of the much higher prices in assets, in capital investments, etc., for them to take profit in order to move their capitals to other regions to capture new economic opportunities. so, in the (t+ j)-th period, when a huge amount of foreign investment leaves, most of the economic activities that got started because of the foreign investments become stalled and/or negatively impacted. therefore, a good portion of the local investments is forced to be retained with the interrupted economic activities. that is, the investors of the local investments can no longer receive their expected dividends, and at the same time a large amount of the investments evaporates totally if many of the stalled activities can no longer be continued until their expected completion, or are indefinitely delayed. that of course costs additional local capital to do the cleanup of what is left behind, unfinished, and unusable ruins. in particular, in equs. (1) and (2), the dividends dt+j of the (t+ i)-th period decreases drastically and the remaining future dividends, if they still come fortunately as expected, would have to be used to bail out (to finish up) some of the other potentially possibly profitable projects. that is, right before the large amount of foreign investments leaves suddenly and strategically, the local economy is more active than ever before. hence, the stochastic gross discount rate rq t+i of the (t+ i)-th period would be much lower than rq t+j of the (t+ j)-th period, because much higher returns on earlier investments are optimistically expected. so, the present ratio: (1−δ)idt+i rq t+1...r q t+i...r q t+j would in reality be very close to zero. summarizing what is just analyzed above, one can see that when a large amount of foreign investments gathers in one place over either a long period or a short period of time and then leaves suddenly and massively, that local ejeffrey forrest: currency wars and a possible self-defense (i): how currency wars ... 216 conomy has to suffer through a positive bubble, caused by the increased money supply as a consequence of the foreign investments, and then a following negative, disastrous bubble, caused by the sudden dry-out of the money supply. and due to a large number of economic activities that are either unexpectedly delayed or totally impossible to complete, the local investors are actually unable to continue to collect their originally expected dividends for many periods to come. that is, foreign investments can be employed as a weapon of mass destruction, if they leave strategically and suddenly, no matter whether they come quickly in a short period of time or slowly over a relatively longer period of time. acknowledges the authors would like to give thanks to the national high-tech program (863) of china (2007aa03z115), independent fund of state key lab. of material processing and die & mould technology of huazhong university of sci & technol and open fund of state key lab. of powder metallurgy of central south university of china (2008112022). the authors also thank for hua yan, bin hua and the analytic and testing center of huazhong university of science & technology for their assistance. references [1] krugman p. (1979), “a model of balance-of-payments crises”, journal of money. credit and banking, vol.11, pp.311-25. [2] kaminsky g, lizondo s. and reinhart c. m. (1998), “leading indicators of currency crises”, imf staff papers. palgrave macmillan, vol.45, no.1, pp.1-48, march. [3] wang r. x, hu g. h. (2005), international finance, wuhan. hubei: press of wuhan university of science and technology. [4] von bertalanffy l. (1924), einfuhrung in spengler’s werk, literaturblatt kolnische zeitung. [5] lin y. and forrest b. (2011), systemic structure behind human organizations: from civilizations to individuals, new york: springer. [6] lin y. (1999), general systems theory: a mathematical approach, new york: plenum and kluwer academic publishers. [7] klir g. (1985), architecture of systems problem solving, new york. ny: plenum press. 217 advances in systems science and applications (2013) vol.13 no.3 [8] klir g. (2001), facets of systems science, new york: springer. [9] lin y. (2008), systemic yoyos: some impacts of the second dimension, new york: auerbach publications, am imprint of taylor and francis. [10] wu y. and lin y. (2002), beyond nonstructural quantitative analysis: blown-ups, spinning currents and modern science, river edge nj: world scientific. [11] lin y. (2007), “systemic yoyo model and applications in newton’s, kepler’s laws, etc”, kybernetes: the international journal of cybernetics, systems and management science, vol.36 no.3-4 pp.484-516. [12] lin y. (1988), “can the world be studied in the viewpoint of systems”, mathl. comput. modeling, vol.11, pp.738-742. [13] lin y, ma y. and port. r. (1990), “several epistemological problems related to the concept of systems”, math. comput. modeling, vol.14, pp.52-57. [14] bernanke b., gertler m. (1999), “monetary policy and asset price volatility”, economic review. federal reserve bank of kansas city, fourth quarter, pp.1751. [15] beechey m, gruen d. and vickrey j. (2000), “the efficient markets hypothesis: a survey”, reserve bank of australia in its series rba research discussion papers numbered rdp. [16] smith v. l, suchanek g. l. and williams a. w. (1988), “bubbles. crashes. and endogenous expectations in experimental spot asset markets”, econometrica, vol.56, no.5 pp.1119-1151. corresponding author author can be contracted at jeffrey.forrest@sru.edu microsoft word 16 y. sun, j. luo, g.f. mi, x.j.wang,x.lin--numerical simulation and defect analysis in the casting of the nod advances in systems science and applications (2011), vol.11, no.3-4 322-330 issn 1078-6236 international institute for general systems studies, inc numerical simulation and defect analysis in the casting of the nodular cast iron truck rear axle y. sun1, j. luo1, g.f. mi2, x.j. wang1 and x. lin3 1state key laboratory of mechanical transmission, chongqing university, chongqing 400030, china 2 college of material science and engineering,henan polytechnic university, jiaozuo 454000, china 3 state key laboratory of solidification processing, northwestern polytechnical university, xi’an 710072, china abstract the paper analyzes the reasons of some typical defects such as shrinkage porosity (macro or micro), shrinkage cavity, cold shut, segregation and hot crack which occurring the truck rear axle casing using nodular cast iron during the casting process. a three-dimensional (3d) cad engineering model is created by the pro/e software, one 3d fdm (finite difference methods) numerical solidification model and z-cast simulation software are used to study the casting solidification of the rear axle casing using nodular cast iron. based on the simulation research results, the structure design of the truck rear axle casing is improved and the casting technique is developed, these simulations and experiments show that these optimization methods are contribute to reduce the casting defect and improve the product quality. keywords pouring system, mold filling, solidification, simulation, casting technique 1.introduction as an important component of truck, the rear axle casing need to withstand the bending stress, torque stress and fatigue during truck transport (takeda 1994), so the rear axle casing will have a good comprehensive mechanical properties and behaviors (such as, high strength and good toughness) by a good casting (pinkaew 2007). the general manufacture technique of a truck rear axle is casting technology, and the general material for manufacturing the truck rear axle is nodular cast iron which is of some good casting characteristics and behaviors (wang, 2003). but as a result of graphite expansion action and effect, which can lead to the fluid feeding channel is blocked and lose the feeding ability in later period of casting solidification, so many kinds of defects will present in the casting, such as cavity, shrinkage and crack. on the other hand, the structure of pouring system and casting technology can affect the stability on casting mold filling and solidification process, when the casting mold filling is lack of physical stability, it can lead to splashing, spatter, gas evolved and molten metal oxidized etc., the formation of casting defects includes sand burning, pore, oxidation and slag (wei, 1996). these above defects reduce casting strength, toughness and quality of the truck rear axle, which is very dangerous to a truck transport. many experimental studies on casting technology are carried out in order to improve casting quality of a truck rear axle and to guarantee the safety of truck transport (pinkaew 2007; huang 2004; gui 2004). in previous research works, these issues are usually dealt with according to personal intuition and experience( wei,1996; liu, 2005), the cost of this traditional research method is very high and the time is more. now the cad/cam/cae techniques are employed to cope with optimization design in casting process (wu, 2005; luo, 2006), which is a more efficiently and low-cost method and can save a plenty of time (luo, 2007; mi, 2007). the paper is supported by the natural science foundation project of cq cstc2008bb3303 and the program for new century excellent talents in university (ncet-08-0607). 323 sun: numerical simulation and defect analysis in the casting of the nodular cast iron truck rear axle the paper will study the casting process of the truck rear axle casing by the numerical simulation method, and then improve the casting technology and casting structure in order to improve the defects of he truck rear axle casing. 2.analysis of casting defects 2.1 structure and material of the truck rear axle casing fig.1 shows that structure of the truck rear axle casing, which is a symmetrical structure, the middle part of the truck rear axle casing is a rotator with an internal flange, the two waist parts are a smoothing transitional surface from the rotator, and there has a internal cavity structure with some small sidesteps. the truck rear axle casing has some groove, sidesteps and flanges, the internal mould cavity structure of the rear axle casing is a main structure forms, which is made up of complex geometrical structure. the basic size of the rear axle casing is 1559mm×448mm×196mm, the thinnest wall of the rear axle casing is 10mm, the thickest part of the rear axle casing is 20mm, and the average thickness of the rear axle casing is 15mm. fig.1. structure of the rear axle casing the nodular cast iron is used to manufacture the truck rear axle casing. the type of material is qt450-10 (trademark in china.), which is ferrite type nodular cast iron. this kind of nodular cast iron includes a high c and si element contents, the content of p and s element are in a low level, the nodular cast iron has a good fluidity and casting behaviors. so the casting defects such as misrun, cold-shut is difficult to appear in the casting workpiece at an appropriate temperature and reasonable pouring system condition, the casting workpiece has a good mechanical properties (strength and toughness ). 2.2 casting defects of the truck rear axle casing fig.2 shows an old pouring system of the truck rear axle casing, which is used in the actual casting process at some factory. the pouring system locates at the legs of the truck rear axle casing, and has two casting ladles and two risers. this structure design of pouring system is of compact and simple characteristics, it is advantageous to a good mold filling and save molten metal on the casting process. the shortcoming of this pouring system in fig.2 is as following by the experimental studies and theory analysis: (1) there are two casting ladles. because the two casting ladles must work simultaneously, so this pouring system makes the casting technology complex and casting quality controlling difficult. (2) there are some differences of casting temperature and beginning time between the two casting ladles, so the casting temperature field is not uniformly at the two sides of the truck rear axle casing, which disturbs the direction of heat transfer, temperature gradient and solidification process, so the casting workpiece is likely to present many defects (such as shrinkage) at the later period of casting solidification. (3) this pouring system increases the instability of mould filling when the molten metal liquid is poured into two independent casting ladles at same time. it is very possible that the sand advances in systems science and applications (2011), vol.11, no.3-4 324 burning, splashing, spatter, gas evolved, slag and molten metal oxidized take place in the casting process, which has a bad effect on the casting quality. fig.2. the old pouring system (a) slag (b) sand burning fig.3. the casting defects fig.3 shows an actual typical appearance casting defects of the rear axle casing, such as slag and sand burning, by adopting this pouring system. 3.new pouring system design and technology optimization 3.1 new pouring system design based on actual experiments and above analysis, the pro/e software is used to establish the three-dimensional (3d) cad engineering model for the truck rear axle casing with a new pouring system, the cad model is one precondition for the structure design and numerical simulation of new truck rear axle casing. by improving the pouring system and considering many important parameters (such as casting shrinkage rate, machining allowance etc.), a 3d engineering model of the rear axle casing with a new pouring system is designed as shown in fig.4. the cad model is saved as stl format by pro/e software, in order that the 3d model can be input fdm numerical simulation software for meshing and simulation. 325 sun: numerical simulation and defect analysis in the casting of the nodular cast iron truck rear axle fig.4. the new pouring system and 3d solid model 3.2 new structure analysis 3.2.1 model of casting process simulation the z-cast fdm simulation software is used to simulate casting mold filling and solidification process in order to check the pouring system reliability, optimize the structure design and guide the casting technology. the simulation model is shown in fig.4 which is input from the pro/e model. the continuity equation, momentum conservation (navier-stokes) equation and energy (enthalpy) conservation equation are respectively as following, 0= ∂ ∂ + ∂ ∂ + ∂ ∂ z w y v x u (1) ug x p z uw y uv x uu t u x 2)( ∇++ ∂ ∂ −= ∂ ∂ + ∂ ∂ + ∂ ∂ + ∂ ∂ μρρ vg y p z vw y vv x vu t v y 2)( ∇++ ∂ ∂ −= ∂ ∂ + ∂ ∂ + ∂ ∂ + ∂ ∂ μρρ (2) wg z p z ww y wv x wu t w z 2)( ∇++ ∂ ∂ −= ∂ ∂ + ∂ ∂ + ∂ ∂ + ∂ ∂ μρρ s z tk zy tk yx tk x z tcw y tcv x tcu t tc + ∂ ∂ ∂ ∂ + ∂ ∂ ∂ ∂ + ∂ ∂ ∂ ∂ =⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ ∂ ∂ + ∂ ∂ + ∂ ∂ + ∂ ∂ )()()( ρρρρ (3) where , ,u v w is the velocity on three direction. , ,x y z is direction factors. g is the gravity, ρ is density, f is the rate of volume, k is the heat conduction ratio, s is the inner heat source. the details of mathematical model please refer to the user guide manual of z-cast. the 3d simulation model is meshed using one uniform method (about 5mm per grid), the meshing numbers of casting solid model is 353×168×90. 3.2.2 results analysis of mold filling simulation fig.5, fig.6 and fig.7 are the temperature field when the casting mold filling states is 30%, 60% and 100% respectively. the 30% filling state (in fig.5) shows that the internal temperature gradient of casting workpiece is small, the temperature at the inlet of ingate is little higher than that of at the bottom of flange. the 60% filling state (in fig.6) shows the casting temperature of the flange body ( in middle of the rear axle casing) is 1300 . ℃ fig.5. the temperature field about the 30% casting mold filling states advances in systems science and applications (2011), vol.11, no.3-4 326 the temperature of two risers and two edges of the rear axle casing are higher than that of the flange body. the molten metal flows into the casting mould steadily through the ingate and long pouring channel, which lies on one side of the casting workpiece. the liquid level of the casting mould rises also steadily in the mold filling process, so the bottom of casting workpiece has little defects about cold shut and feeding insufficiency in this way. fig.6. the temperature field about the 60% casting mold filling states it shows obviously that the casting workpiece establishes a good temperature gradient when the mold filling percents is 100% in fig.7. the temperature of the risers is higher than that of the casting body. this kind of temperature distribution is helpful to bring into full filling and overflow, and reduce the existences of internal shrinkages defect in the casting workpiece. fig.7. the temperature field about the 100%casting mold filling states 3.2.3 results analysis of casting solidification simulation fig.8. the temperature field about solidification time is 900 second 327 sun: numerical simulation and defect analysis in the casting of the nodular cast iron truck rear axle fig.9. the temperature field about solidification time is 1200 second at the 1200s solidification final stage (in fig.9), it shows obviously that solidification speed of the ingate is faster than that of casting body, the temperature of the ingate is lower that that of casting body, so the pouring channel is closed. it makes the risers can not play the role of feeding and overflow at the end period of solidification, and then the expanded solid graphite make metal liquid attempt to flows back into pouring channel from the casting body, so many defects ( such as hot crack and shrinkage cavity) take place at this solidification stage. in addition, there has one interface between the ingate and casting mould, the temperature at combining site where the fluid enter casting mould by the ingate is higher than other parts, it causes the solidification time is long at the combining site, many defects also present at this site, the reason is that the structure of inner channel near this combining site is very complex. so the pouring system , riser and structure of the ingate need to improve. 3.3 optimization a new improvement pouring system is designed, the position of the casting ladles and risers are optimized, and one new structure of the ingate at the combining site (between pouring channel and casting workpiece) is proposed, at the aid of the cad technique( pro/e software) and many numerical simulations (such as z-cast, procast and fluent softwares ) for mold filling and solidification process. for avoiding early solidification of molten metal in the ingate, the height of the ingate increases from 12mm to 24mm, and the length of the ingate improve from 75mm to 50mm considering the feeding ability of the risers and sand burning ratio. correspondingly the width of the horizontal runner changes from 22mm to 24mm, the diameter of the sprue bottom part which is located in the middle of the runner adds from 41mm to 43mm, and the position of the riser is also change. the temperature field of the new improvement system is shown in fig.10 and fig.11 respectively when the solidification time is 900 second and 1200 second after the mold filling process is finished. fig.10. the temperature field about solidification time is 900 second (improved) advances in systems science and applications (2011), vol.11, no.3-4 328 it shows obviously that the temperatures at the ingate and the bottom of riser are higher than that of casting body, so the pouring channel is unblocked and feeding at the final stage of casting solidification, this good temperature distribution in fig.11 is also helpful to reduce many internal defects in casting workpiece. compared fig.9 with fig.11, the numerical simulation results show that the improvement system is of availability. fig.11. the temperature field about solidification time is 1200 second (improved) on the other hand, the truck rear axle casing has a complicated internal structure. in order to improve the casting quality, the internal structure of the truck rear axle casing is also optimized as shown in fig.12. cooperated with the optimization of pouring system, the casting technology is improved to suit for the new optimization structure simultaneously. the results of numerical simulation and experiences show that the truck rear axle casing has a good casting profile and high mechanical properties, and the casting production meet the requirement of design and quality. fig.12. the optimized internal structure and pouring system of casting workpiece 4. discusses compared to fig.2, the runner of the new pouring system (in fig. 4) employ a new smooth circle transition structure, which can ensure that molten metal liquid flow steadily from the ingate into the casting workpiece with an uniform speed and adequate quantity, this kind of mold filling method can prevent the fluid form a turbulence state, and avoid the air mixture enter the casting, reduce the defects such as metal oxide, shrinkage porosity (macro or micro). the new pouring system is also of great benefit to the casting behaviors and improve gases release from molten metal, it can make molten metal liquid feeding and reduce the appearance of slag, shrinkage and hot cracks. at the same time, it is very convenient that workers operate the pouring system, in one word, the integrated methods including the new structure, casting 329 sun: numerical simulation and defect analysis in the casting of the nodular cast iron truck rear axle technology and pouring system are useful to establish effectively an sequential solidification procedure from mid part to two edge part of the truck rear axle casing. 5. conclusion the mold filling and casting solidification process of the truck rear axle casing using nodular cast iron are studied. the pouring system and structure of rear axle casing are improved, and the casting technique is developed. the integrated methods including the new casting structure, casting technology and pouring system are useful to establish a great directional and sequential solidification procedure from middle part to two edge part of the truck rear axle casing. the optimization pouring system reduces the defects in casting workpiece. the simulation research results and experiments show that these optimization methods are contributed to reduce the casting defect and improve the product quality. acknowledgements the authors gratefully acknowledge the support of a grant from the natural science foundation project of cq cstc 2008bb3303 and cq cstc 2009ba3026, the ph.d. programs foundation of ministry of education of china (no.20070611030 ), the outstanding young scholar program of natural science foundation of hubei province (no. 2006abb027), the program for new century excellent talents in university (ncet-08-0607), changjiang scholars and innovative research team in university (irt 0763), the project of the state key laboratory of mechanical transmission in chongqing university and the fund of the state key laboratory of solidification processing in northwestern polytechnical university. the paper is also supported by the exchange study foundation program of chongqing university of the project 211 tertiary system. references [1] takeda, nobuyuki et al. stress analysis of rear axle case for heavy-duty truck. (optimum design of rear cover fixed area), transactions of the japan society of mechanical engineers, part a, 60(579) (1994) 2612-2617. [2] pinkaew, tospol, et al. experimental study on the identification of dynamic axle loads of moving vehicles from the bending moments of bridges. engineering structures, 29(9) (2007) 2282-2293. [3] wang h.l. the casting techniques: typical automobile workpieces. polytechnical university press, beijing (2003). [4] wei qt. casting technology. xi'an: northwestern polytechnical university press (1996). [5] liu, b.c. development trend of casting technique and computer simulation, foundry technology, 26(7) (2005) 611-617. [6] lin q.a. pro/engineering design guide beijing: beijing univeristy press (2000). [7] g.f.mi, et al. development and application of numerical simulation for the mold filling process of casting. journal of henan polytechnical university (natural science), 26(3) (2007) 334-339. [8] huang x, et al. development & research of 13-ton automobile axle casing casting, automobile science and technology, (4) ( 2004) 27-30. [9] wu c.g, et al, 2007. the foundry technology design and numerical simulation of automotive rear axle by proportional solidification and macroporous outflow, china foundry machinery & technology, (5) (2007) 25-28. [10] wu m, et al. numerical study of the thermal-solutal convection and grain sedimentation during globular equiaxed solidification, material science forum, 475-479(5) (2005) 2725-2730. [11] luo j , et al.the 3d simulation of liquid core change of cylinder steel rolling forming on soft-reduction continuous casting process. isdm2006 international conference, oct.15-17, advances in systems science and applications (2011), vol.11, no.3-4 330 2006, wuhan, p.r.china. journal of wuhan university of technology, 28 (si.) (2006) 637-639. [12] luo j, et al.numerical study of liquid core solidification in influence of soft reduction deformation on steel slab continuous casting process, proceedings of the 5th international conference on physical and numerical simulation of materials processing (icpns2007), october 23~27, zhengzhou, p.r.china. material science forum, 575-578 (2008) 80-85. [13] luo j, luo q, lin yh, xue j. a new approach for fluid flow model in gas tungsten arc weld pool using longitudinal electromagnetic control. welding journal, 82 (8) (2003) 202s-206s. [14] gui q.s. casting technology of back bridge housing, foundry technology, 25(5) (2004) 337-339. microsoft word 6 bing zhang, ying wang, dejun chen--a research on technology project credit evaluation model based on ahp and advances in systems science and applications (2011), vol.11, no.3-4 249-256 issn 1078-6236 international institute for general systems studies, inc a research on technology project credit evaluation model based on ahp and fcem bing zhang, ying wang and dejun chen school of information engineering, wuhan university of technology, wuhan 430070, p.r.china abstract this paper analyzes the basic requirements of the technology project credit evaluation, and presents a technology project credit evaluation model based on analytic hierarchy process (ahp) and fuzzy comprehensive evaluation method (fcem). combining with the actual situation of one scientific research center, a technology project credit evaluation index system is established, and its weight of each evaluation index is determined by ahp, and then an evaluation results is analyzed and evaluated through fcem, also it’s valuable theoretical foundation for the management of the technology project. keywords ahp; fcem; project credit evaluation 1. introduction the technology project credit evaluation is an important part of technology project process management. in the technology project concluding, it is necessary to evaluate the technology project implementation process, which is not only the summary of project implementation process, but also the archive of credit of undertakers, which provide important historical basis for the future project application approval procedures. currently, the technology project credit evaluation is based on subjective qualitative assessment, there is no reasonable technology project credit evaluation index system, or the evaluation indexes are lack of scientific weight distribution. to solve this problem, this paper presents a comprehensive evaluation method based on ahp and fcem. combining with the actual situation of one scientific research center, this paper proposes a technology project credit evaluation index system, which evaluate the credit of project stakeholders from the contract compliance, reporting significant matters, and implementation within the stipulated time those three aspects, and determines the weight of each index by ahp and checks the consistency, then taking one credit evaluation results in a project as example, calculates the technology project credit situation through fuzzy analysis and quantitative assessment to validate this model. 2. technology project credit evaluation index system technology project credit evaluation system should be operated from the multi-level, multi-angle, which should be able to fully reflect the technology project's credit rating commitment, combined with an actual situation of r & d center, based on ahp, a three-level technology project credit evaluation index system is proposed , as shown in figure 1. figure 1 show that, this system is made up of 3 respects of contract compliance, reporting significant events and implementation within the stipulated time, and has 8 indexes; of course, it can be adjusted according to actual situation. explain the specific content of each index as follows: (1) completion condition of assessment indicators: refers to the completion condition of the content stipulated in the contract. (2) rate of progress is the completion situation of the project progress according to the contract rules. (3) reporting significant events is to account for the significant issues to the virtual 250 zhang: a research on technology project credit evaluation model based on ahp and fcem coordination center faithfully and timely. u 1u 2u 3u 11u 12u 31u 32u 32u 34u 35u figure 1 technology project credit evaluation index system (4) submit research plan: refers to submit their work outline within the specified time, such as it is finished after being noticed in two months. (5) submit contract: refers to hand over contract within the specified time, such as it is finished after being noticed in three months. (6) submit the sheet about the execution situation: refers to submit it in scheduled time, such as finishing it on 15 of the first month of each quarter. (7) submit the acceptance of applications: refers to submit it before the deadline specified in the contract. (8) submit archive data: refers to submit it within the specified time after project acceptance, such as 1 month after the inspection (30 days) for submission. 3 the establishment of technology project credit evaluation model based on the ahp and fcem ahp is used first to define the weight of each level index, and then use fcem for project credit evaluation. the model is established in following steps: 1.tablishment of factors set from the project credit evaluation index system in figure 1 can be seen, there are two levels of evaluation indices, and now the first level is defined },,{ 321 uuuu = ; the second level is },{ 12111 uuu = , },,,,{ 35343332313 uuuuuu = 。 2. establishment of reviews set according to the actual situation of the r&d center mentioned before, reviews are set to 4 levels, namely, n=4, },,,{ 4321 vvvvv = represent {excellent, good, medium, poor}. reviews level can be adjusted and determined according to the specific situations. 3. determine weights set of evaluation indexes by ahp using the ahp to determine weights set of evaluation indexes can be divided into the following steps: (1) construct judgment matrix judgment matrix represents the relative importance between two elements in the same level to some element in the upper level, the evaluation about the importance of indexes at all levels are a subjective process, based on expert evaluation results or the result of the questionnaire. use advances in systems science and applications (2011), vol.11, no.3-4 251 issn 1078-6236 international institute for general systems studies, inc ),...2,1,(, njibb ji = to represent the indexes. ijb expressed the value of the importance that ib relative to jb , construct the judgment matrix p through 1-9 ratio scaling, and the matrix have reciprocity and basic consistency, that is, 0>ijb , 1=iib , 1=∗ jiij bb .   ⎥ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎢ ⎣ ⎡ = nnnn n n bbb bbb bbb p 21 22221 11211                                       (1)                      table 1 1-9 description of proportion quotients ijb meaning explanation 1 equal importance both have the same importance 3 weak importance ib important than jb slightly 5 strong importance ib important than jb obvious 7 very strong importance ib more important than jb 9 absolute importance ib absolute important than jb obvious 2、4、6、8 between the various levels above the importance is between the adjacent levels 1、1/2、… 、1/9 reverse comparison the importance of jb relative to ib with experience and knowledge related to credit evaluation and project management, construct judgment matrix of each level as 1p 2p 3p . in order to facilitate analysis and more intuitive, we graphically shows the matrix, as shown in table 2, 3, 4. table 2 judgment matrix 1p table 3 judgment matrix 2p table 4 judgment matrix 3p 3p 31u 32u 33u 34u 35u 31u 1 1 1 1/2 2 2p 11u 12u 11u 1 2 12u 1/2 1 1p 1u 2u 3u 1u 1 2 1/2 2u 1/2 1 1/3 3u 2 3 1 252 zhang: a research on technology project credit evaluation model based on ahp and fcem 32u 1 1 1 1/2 2 33u 1 1 1 1/2 2 34u 2 2 2 1 3 35u 1/2 1/2 1/2 1/3 1 (2) calculate the weight vector of each level and make consistency check first, calculate the maximized eigenvalue and eigenvector of judgment matrix. generally, it is to use geometric averaging (root method) or normative column average (sum method) to calculate the approximate eigenvectors [2], and then calculate the maximized eigenvalue. geometric average method: calculate the product of each element of each row, then calculate the nth root of each product; and then normalized the obtained vector. vector obtained above is the approximate eigenvectors, if the consistency check is passed, the vector is the relative weight vector of each index. calculating the maximized eigenvalue and eigenvector using geometric averaging mean is as follows: ① calculate the geometric average of all elements of each row of the judgment matrix. based on n n j iji bw ∏ = = 1 , ( )ni ,,2,1= ,so ( )tnwwww ,,, 21= 。 ② normalize w , ∑ = = n i iii www 1 , ni ,,2,1= , then t ni wwww ),,,( 2= which is approximate eigenvectors, and its value of each element is the weight value of each index. ③ calculate the largest eigenvalue maxλ ( )∑ = = n i i i nw pw 1 maxλ (2) in the formula 2, vector ( )ipw is the first i component of pw . then check on the consistency of judging matrix. matrix consistency test as follows: ① calculate the inconsistent level (ci) of judgment matrix. 1 max − − = n n ci λ (3) in the formula 3, maxλ is the maximized eigenvalue of )1( >nn order matrix. ② calculate the random consistency level (ri) of judgment matrix, which only determined by the order of the judgment matrix. note that, when 20 ≤< n , there is no inconsistency issue, matrix does not need be tested. standards of random consistency level shown in table 5: ③ calculate the consistency ratio of judgment matrix (cr). ri cicr = (4) advances in systems science and applications (2011), vol.11, no.3-4 253 issn 1078-6236 international institute for general systems studies, inc table 5 random consistency level ri n 1 2 3 4 5 6 7 8 9 10 ri 0 0 0.58 0.9 1.12 1.24 1.32 1.41 1.45 1.49 the method to determine the consistency of judgment matrix is [3]: when .10 0 (2) when outliers are present, {yt, t = 0, 1, 2, · · ·} is unobservable. instead the time series {zt, t = 0, 1, 2, · · ·} is observed which follows the model: zt = yt + dt, where dt is a deterministic perturbation. let ψj = 0 if j < 0, then we have dt = ψt−t0w0, if time is io; or dt = w0δt,t0 , if time t0 is ao , where w0 denotes the outlier’s magnitude at t = t0. in the present setting, a chromosome ξ is a string of characters of assigned length n that can be evaluated in terms of the ff, where n is the number of observations of the time series, as each locus ξj is corresponding to an observation zj where an outlier may occur. so, ξ = (ξ1, ξ2, · · · , ξn). then a gene ξj = 0, if the locus is an outlier-free time point, ξj = 1 if the observation at this time point is an additive outlier, and ξj = 2 if it is an innovational outlier. if k outliers are located at t1, t2, · · · , tk and denoting by z = (z1, z2, · · · , zn)′ the observed time series and by y = (y1, y2, · · · , yn)′ the unobserved realization of yt, then z = ψw + y , where w = (w1, w2, · · · , wk) ′ is the vector of the outliers’ magnitudes at t1, t2, · · · , tk and ψ is the matrix of the n × k elements ψjh defined as follows: if th is the timing of an io, then ψjh = ψj−th if j > th and 0, otherwise. if an ao is occurring in th, then ψjh = 1 if j = th and 0, otherwise. the idea is to seek for the matrix ψ which maximizes the likelihood function. when n is large, γ−1 may be replaced by the matrix of inverse autocovariances γi . then we have the likelihood function of the time series z l = p(z | ξ,w ) = 2π−n/2(detγi)1/2 exp{−1 2 (z −ψw )′γi(z −ψw )} (3) the joint maximum likelihood estimate w ∗ of w , given the pattern of the outlying observations and γi, is w ∗ = (ψ′γiψ)−1ψ′γiz and y ∗ = z −ψw ∗ (4) advances in systems science and applications (2012) vol.12 no.4 403 in practice, however, γi is seldom known, so we have to estimate it from the data. similar to baragona et al.[1], using the interpolator estimate of yt. we may obtain inverse autocorrelations interpolation estimates ρ̂ij , j = 1, · · · ,m and the estimate γ̂i0 of inverse variance and the estimates of inverse autocovariance γ̂ij = γ̂i0 · ρ̂ij , j = 1, · · · ,m. once the inverse autocovariances have been estimated, the ψj , j = 1, · · · , q may be computed using equation (2). nevertheless, these estimates are biased because of the presence of the outliers. so we have to resort to an iterative scheme for (4), which is repeated until convergences. if convergence is attained, and since the likelihood is maximized with respect to the inverse autocorrelations, the argument of the exponential in likelihood’s formula (3) is simply −n/2 , which is shown by the following proposition(the proof is omited). proposition 1 using the former representation and estimations, then we have −1 2 (z −ψw )′γ(z −ψw ) = −1 2 n therefore, we have logl = −n 2 log(2π) + 1 2 log(detγi)− 1 2 n thus, the likelihood function basically depends on detγi. because that the fitness function(ff) has to be positive, we let f (ξ) = a · blog(detγi)−ck, where b is a real constant such that b > 1, and the constant terms in the exponent were dropped. we use a = 10, b = 1.0001 and c = 14 in this paper. we adopted (1) population size s = n + 1; (2) probability of crossover pc = 0.8; (3) probability of mutation pm = 0.05; (4) the maximum number of outliers within a chromosome g = 10; (5) the number of iterations of the series’ adjustment/parameters’ estimation procedure was 3;(6) we perform as many iterations of the genetic algorithm as possible within a reasonable time period. 5 simulation studies example a in the following example, we consider the model (1− 0.8b + 0.3b2)xt = εt zt = (1− 0.5b)xt − 4δt,30 + 5δt,31 − 5δt,80 − 4δt,90+ 6× 1− 0.36b + 0.85b2 1− 0.6b δt,40 + 1− 0.36b + 0.85b2 1− 0.6b et where {εt} and {et} are all normal white noise, their means are zero and variance σ2 = 1. we create 101 observations x0, x1, · · · , x100 of xt and 100 observations z1,...,z100 of zt by simulation. obviously, it is ao at t = 30, 31, 80, 90 singly and io at 404 ping chen:the identification of outliers in armax models via genetic algorithm t = 40 , and outlier magnitudes are w30 = −4, w31 = 5, w80 = −5, w90 = −4 and w40 = 6, respectively. applying our method to the above data and prewhitening the input series. making {xt} follows an arma model: (1− 0.88088b + 0.32738b2)xt = εt. then we take the same manipulation to prewhiten {zt}. by analyzing filtered cross correlation coefficient of {zt} and {xt}, we obtain the transfer function 1.03416−0.73967b for {xt}. delete the influence of input process {xt} in response process {zt}, and let z∗t = zt − (1.03416 − 0.73967b)xt. we have that {z∗t } is an arma series include outliers. we detect the outliers in {z∗t } by applying the above method. let m = 2, q = 9, g = 10, s = 101. because the 30th point is very close to 31th point and they influence each other, it is difficult to identify the outliers. in this case, one needs larger number of iterations. we take 1000 iterations by standard genetic algorithm. and obtain the best individual at 612th iterations: t = 30(ao), w∗ 30 = −4.7172; t = 31(ao), w∗ 31 = 4.1010; t = 40(io), w∗ 40 = 4.0487; t = 80(ao), w∗ 80 = −5.6440; t = 90(ao), w∗ 90 = −4.6228 the outcome is consistent with our prearrangement. the outliers in {zt} process are detected successfully, and there is no misjudgement. 6 conclusions there are ao and io in our model, and also the outliers(ao) present consecutively. it is quite difficult to detect outliers, however, all of the outliers in the above model have been detected successfully, which shows our method is also efficient for the ao and io problem in armax model. some other case studies also show that our method is effective in detecting outliers’ location and type and in estimating their size for armax model. acknowledgements the research is supported by national natural science foundation of china (no.11171065) references [1] baragona r, battaglia f, calzini c. (2001), “genetic algorithms for the identification of additive and innovation outliers in time series”, computational statistics & data analysis, vol.37, pp.1-12. advances in systems science and applications (2012) vol.12 no.4 405 [2] peña d, sánchez i. (2005), “multifold predictive validation in armax time series models”, journal of the american statistical association, vol.100, pp.135-146. [3] chen p, li l, liu y and lin j.g. (2010), “detection of outliers and patches in bilinear time series models”, mathematical problems in engineering, vol.2010, pp.1-10. [4] chen p, yang j, li l.y. (2013), “synthetic detection of change point and outliers in bilinear time series models”, international journal of systems science, http://dx.doi.org/10.1080/00207721.2013.777983. (in press) [5] huang l, pang w, wang k.p, zhou c.g, xiao y. (2004), “improved genetic algorithm for vehicle routing problem with time windows”, advances in systems science and applications, vol.4, pp.118-124. [6] box g.e.p, jenkins g.m, reinsel g.c. (1994), time series analysis: forecasting and control, third edition, prentice-hall, englewood cliffs, nj. corresponding author author can be contacted at cp18@263.net.cn advances in systems science and applications (2014) vol.14 no.2 190-198 design & simulation of smitha: a structured and scalable architecture for 3d network on chip systems sanju v1, niranjan chiplunkar2, m khalid3 and p venkata krishna3 1nitte meenakshi institute of technology, yelahanka, bangalore, india 2nmamit, nitte p.o, karkala, india 3vit university,vellore, india abstract network on chip is a new paradigm in asic design, and new topologies are emerging from this domain of work, scalable modular interconnect for three dimensional high performance applications is a typical 3d architecture(smitha) exactly. in this paper, smitha will be analyzed, and be designed in verilog hdl. keywords smitha, network on chip, design and simulation 1 introduction early vlsi designs were implemented using common bus architecture. as time passed by the density of the ic’s rose dramatically.bus architecture failed to give the required amount of performance. this gave rise to a new architecture called network on chip[1-3]. the idea behind this is to use concepts of data networking on silicon. however, network on chip is becoming the backbone of all high performance chips, many topologies have been introduced in this novel field. mesh and torus[2,4] being the most commonly used topology. in this paper we like to introduce smitha, a new three dimensional topology for network on chip based systems. in this paper, a new three dimensional topology for network on chip based systems is introduced, and then, the internals of the topology and the implementation of the topology are explained. 2 a new 3d topology : smitha smitha: scalable modular interconnect for three dimensional high performance application (smitha) is a new three dimension topology for network on chip based systems. the construction of the topology can be explained as follows. the base topology is obtained by removing the root node and interconnecting the nodes of a complete binary tree along the same layer as shown in fig.1. the nodes are addressed based on the layer in which they belong and relative position within the layer. the three dimension topology is obtained by placing the above discussed base topology one over the other and by interconnecting the adjacent levels as follows. advances in systems science and applications (2014) vol.14 no.2 191 fig.1 smitha base topology with 3 layers fig.2 smitha base topology with three levels having three layers / level 192 sanju v: design & simulation of smitha: a structured and scalable architecture ... • the connections between odd level and an even are done by connecting the odd layers though the left and even layers through the right. for example interconnection between level one and two. • similarly the connections between even level and an odd are done by connecting the odd layers though the right and even layers through the left. for example interconnection between level two and three. this is depicted in fig.2. 3 design and sumulation of smitha the architecture discussed above was designed and implemented using verilog hdl. the code was tested and verified using modelsim and synthesised for cyclone ii fpga on altera de2 70 board. the design starts from the smallest unit in the topology which is the node which interconnect to form the topology. the node is designed with five interfaces which connect to the external world. fig.3 gives a representation of a node. fig.3 internals of a smitha node fig.4 internals of a smitha node interface advances in systems science and applications (2014) vol.14 no.2 193 each node has five interfaces which are named left interface (li), right interface(ri), top left interface (tli), top right interface (tri), bottom interface (bi). these connect to the neighbouring node to do their specified operation. fig.4 shows the internals of each of the interface. it consists of the connecting pins namely req, ack , data, clk which are request pin,acknowledgement pin, data pin and the clock pin. it consists of the routing algorithm in the routing logic and the whole control of the node in control logic. it also consists of temporary storage purpose namely the registers – send and receive and the buffers – send and receive. 3.1 packet format the packet has three components namely the destination address, source address and the data bits. fig.5 smitha packet format fig.5 depicts the same. the source and destination address consists of the level number to which the node belongs, the ring in the level and the position of the node in that ring. the level numbering starts from 0 to n 1 where n is the largest number of level the architecture holds. similar arguments holds good for ring number and node number which are within each level. similar argument hold good for source addressing also. data bits can be of any size depending on the problem. for example a node addressing (1, 0, 0, 2, 0, 0, 25) represents a data 25 from level two ring zero node zero to level one ring zero node zero. the total size of the packet is 2 ∗ ( log2 (n) + log2 (m) + ∑m k=1 2 k ) + t data bits. 3.2 communication between nodes fig. 6 shows how to nodes in smitha topology communicates. the nodes in smitha communicate with each other with first raising the req pin. the receiving node checks whether the receive buffer is full or not. if the receive is not full then busy bit and the ack bit is set. the code illustrates the same. pueduo code on recieve req // pseudo code when request signal is received { if(!full(recievebuffer) // checking receive buffer is full or not 194 sanju v: design & simulation of smitha: a structured and scalable architecture ... { busybit = 1; // if receive buffer is not full then make busy bit ack = 1; // and ack bit high } } fig.6 two nodes communicate in smitha once receiving the request signal, the data sending module will clock and send the data one bit at a time. after sending the data the busy bit of the sending module is kept low. on the receiving side, the data is received on with the help of receive register. after the reception of data, the busy pin is kept low. now the packet is checked by the routing algorithm to decide the next path. if the send buffer of the next interface is not full then the packet is send to the next interface else it is kept in the receive buffer of the current interface. the pseudo code below illustrates the same. pseudo code onack sendingside { for(i = 0 ; i< length(packet); i++) { clk = 1; data = sendregister[i]; clk = 0; } busybit = 0; advances in systems science and applications (2014) vol.14 no.2 195 } pseudo code onack recieveside { for(i = 0 ; i< length(packet); i++) { recieveregister[i] = data; } busybit = 0; nextinterface = route(recieveregister); if(!full(sendbuffer(nextinterface))) { store(nextinterfacebuffer); else store(recievebuffercurrentinterface); } } 3.3 control logic it is a integral part of the system. it basically checks and starts the operation of sending the packet. furthermore the packets which in the receive buffer are forwarded to their respective location by the logic. the pseudo code below depicts the same. pseduo code controllogic { if(!empty(sendbuffer)&&!busybit) { busybit = 1; req = 1 ; } if(!empty(receivebuffer)) { nextinterface = route alg(packet) ; if(!full(sendbuffer(nextinterface))) enqueue(packet) } } 3.4 output verification the above circuit was coded in verilog and tested using modelsim and was downloaded to de2 – 115 board to test the code. the code was tested for absolute 196 sanju v: design & simulation of smitha: a structured and scalable architecture ... fig.7 output waveform advances in systems science and applications (2014) vol.14 no.2 197 accuracy with different conditions. one of such output is as below (fig.7)it describes the packet going from (1,2,1) to (1,2,0 ) and then it moves to the second level to (2,2,0) and ends at ( 2,2,1). 4 conclusion the paper briefly draws a light into the new three dimensional architecture and discusses the fpga implementation using verilog hdl by means of altera’s quartus and modelsim for simulation purpose. the paper starts with the node implementation to the topology. the implementation discusses the internals of the node with every part of the node in detail. it clearly brings out the working of the module by the waveform shown. references [1] william j. dally , brian towles. (2001), “route packets not wires : on chip interconnection network”, dac 2001. [2] tobias bjerregaard , shankar mahadevan. (2006), “a survey of research and practices of network on chip”, acm computing survey. [3] ahmed hemani, axel jantsch, shashi kumar, adam postula, johnny berg, mikael millberg,dan lindqvist, “network on a chip: an architecture for billion transistor era”, dac. [4] ville rantala, teijo lehtonen, juha plosila. (2006), “network on chip routing algorithms”, tucs technical report, no.779. [5] radu marculescu, ieee, natalie enright jerger, yatin hoskote, umit y. ogras and li shiuan peh, “outstanding research problems in noc design: system, micro architecture and circuit perspectives”, ieee transactions on computer aided design of integrated circuits and systems,vol.28,no.1. [6] j. nurmi. (2005), “network on chip: a new paradigm for system on chip design”, proceedings 2005 international symposium on system on chip, no.17. [7] l.benini, g.de micheli, “networks on chips: a new soc paradigm”, ieee computer, vol.35. [8] s. kumar, a. jantsch, j. soininen, m. forsell, m.millberg, j. oberg, k. tiensyrja, a. hemani. (2002), “a network on chip architecture and design methodology”, proceedings international symposium vlsi (isvlsi), pp.117-124. 198 sanju v: design & simulation of smitha: a structured and scalable architecture ... [9] dan marconett, “a survey of architectural design and implementation tradeoffs in network on chip systems”, university of california, davis. [10] nilanjan banerjee, praveen vellanki, karam s.chatha. (2005), “power and performance model for network-on-chip architectures and test in europe conference and exhibition (date)”, transactions on computers, vol.54, no.8, pp.1025c1040. corresponding author sanju v can be contacted at:sanjuv21@gmail.com. advances in systems science and application (2015) vol.15 no.1 21-36 big data and big control d.a.novikov abstract this paper analyzes and structures modern challenges in the field of big data handling, arising before hardware developers, experts in applied mathematics and artificial intelligence, as well as researchers focused on different application domains. the author outlines further development of the theory and applications of big control, i.e., control based on big data. keywords big data, data analysis, big control, big data management 1 introduction i am a great believer in the simplicity of things and as you probably know i am inclined to hang on to broad & simple ideas like grim death until evidence is too strong for my tenacity. c sir ernest rutherford from letter to irving langmuir (june 10, 1919). the term “big data” means unstructured data whose volume exceeds the available handling capabilities in required time. this term appeared just a few years ago [1], but became very popular among it experts, scientists, business analysts and other specialists. for instance, a retrieval request to google returns millions of links on the subject. what are the capabilities and risks carried by big data? what challenges and problems do they formulate before scientists, experts in different application domains (particularly, specialists in systems science and control theory-big data usage in control problems and big data management), the education system as a whole? 2 big data. big analytics. big visualization in information technology, big data (perhaps, the term was first mentioned in the special issue of nature[1]) represents a direction of theoretical and practical investigations on the development and application of handling methods and means for the big volumes of unstructured data. big data handling comprises their1: acquisition; transmission; storage (including recording and extraction); processing (transformation, modeling, computations and analysis); 1in certain classifications, big data handling is associated with 4d (data discovery, discrimination, distillation and delivery/dissemination). 22 d.a.novikov: big data and big control usage (including visualization) in practical, scientific, educational and other types of human activity. in the narrow interpretation, the term “big data” sometimes covers only the technologies of their acquisition, transmission and storage. in this case, big data processing (including construction and analysis of corresponding models) is called big analytics (including big computations), whereas visualization of the corresponding results (depending on users cognitive capabilities) is called big visualization (see fig. 1). the universal cycle of big (generally, any) data handling is illustrated by fig. 2. here the key role belongs to an object and a subject (a “customer”); the latter requires knowledge on the state and dynamics of the former. however, sometimes there exists a chasm between the data acquired on an object and the knowledge necessary for a subject. primary data must be preprocessed, i.e., transformed into more or less structured information. subsequently, necessary knowledge is extracted from this information depending on a specific task solved by a subject. fig.1“the big triad”2: data, analytics, visualization particularly, a subject may adopt this knowledge for object control , viz., exerting purposeful impacts on an object to ensure its required behavior. control can be automatic in a special case (an inanimate subject). perhaps, the term “big control”3 will become common soon for indicating control based on big data, big 2we will not discuss another fashionable triad (big data, high-performance computations, cloud technologies 3for justice’ sake, note that in the recent fifteen years experts in control theory have tended to consider the problems of control, computations and communication jointly (the so-called c3 problem (control, computation, communication)). according to this viewpoint, control actions are synthesized in real time taking into account the existing delays in communication channels and information processing time (including computations). there is another generally accepted term (large-scale systems control), but big data can be generated by “small” systems.] advances in systems science and application (2015) vol.15 no.1 23 analytics and, possibly, big visualization.4 fig.2 the universal cycle of big data handling the qualitative analysis of numerous publications on big data leads to the author’s subjective expert appraisal for the current distribution of attention paid by researchers and developers (but not users!) to big data handling problems. this expert appraisal is demonstrated in fig.3. fig.3 the current distribution of attention paid by researchers and developers to big data handling problems 4an alternative interpretation of “big control” concerns control of big data handling processes. actually, this represents an independent and nontrivial problem. 24 d.a.novikov: big data and big control in other words, the overwhelming majority of big data investigations create the technologies of big data acquisition, transmission, storage and preprocessing, whereas big analytics and visualization receive by far less consideration. big control problems are almost not studied, despite that control objects become more and more complicated and multi-scale, while “networkism” represents a leading trend of modern control theory [2]! 3 civilization-scale problems however, is the current state of affairs (see the distribution in fig.3) reasonable? on the one hand, the answer is affirmative. indeed, technologies followed exactly this path of development; moreover, data analysis and visualization requires data acquisition and storage (no doubt, with the feasibility of rapid accessing and processing). on the other hand, the existing “disbalance” is the result of the following. nowadays, mankind realizes the potential utility of any data, but does not completely understand what should be done with the growing avalanche of data. this problem seems not novel, since a class of similar civilization-scale problems has appeared recently, formally called the problems of anticipatory technologies development. to elucidate this idea, consider the correlation among science, technologies and practice (see fig.4). during different periods of mankind development, science often initiated creation and adoption of certain technologies; sometimes, the chain was “inverse” (presently, we observe exactly this picture!). fig.4 science, technologies and practice really, appeal to history. starting from the middle of the 19th century, for over 100 years mankind had experienced the triumphal development of science and the technological boom (science anticipated and predetermined technologies, the latter were widely adopted in practice; generally speaking, before science was advances in systems science and application (2015) vol.15 no.1 25 stimulated by practice).in this context, we mention electricity, communications, nuclear power engineering, electronics, etc. most technologies were clear and available to average men.5 later on (approximately, since the 1970s) the situation gradually changedcthe accumulated fundamental scientific results were enough for technologies growth (i.e., some technologies even “anticipated” science). consequently, their “demand for science” was reduced (perhaps, living systems made an exception). nevertheless, technologies demonstrate further development, and the pace of their development even increases. for instance, in the recent decade advances in information and communication technologies (ict) have been even anticipating practice6, including human recognition of new technological capabilities, development prospects and associated threats.7 in other words, the threshold of the third millennium was remarkable for a turning point: mankind had earlier mastered science and technologies to satisfy its needs, whereas afterwards technologies rather dictated the directions, restrictions and conditions of scientific development, at the level of individuals and at the level of states and mankind as a whole.8 exactly this effect is called “anticipatory technologies development”. mankind will have to realize the corresponding civilization-scale problems and learn to respond to them. similar things happen with big data: human beings have mastered technologies for accumulating giant volumes of data, but are still unable to process and utilize them. the major problem concerns the comprehension of why should big data processed (rather than how should it be done). to have a clear view of the challenges faced by scientists and engineers, we 5amusingly, the heroes of fantasy books on technological breakthrough, owing to isolation (e.g., cyrus smith from j. verne’s the mysterious island) or a trip to the past (e.g., the hero from m. twain’s a connecticut yankee in king arthur’s court), are often engineers or average men but not scientists. 6as a positive feature, we acknowledge that advances in ict were stimulated by the development of applied mathematics, particularly “network mathematics” (random graphs, largedimensional graphs, network games and others, see the surveys [3-6]). 7scattered manifestations of this phenomenon took place earlier; e.g., a. noble and the participants of the manhattan project thought about the ethical responsibility of scientists. 8the problems of information security have become common. today is the right time for thinking not just about technological security (cybersecurity), but about the social, economic and other security of information technologies, i.e., the security of ict users, their groups and society as a whole against informational impacts (here a striking example concerns social media including online social networks [7]). recall that all important decisions (at the level of an enterprise and at the level of a government) are made on the basis of information coming from certain sources and processed by certain methods (far from always, a decision-maker knows these sources and methods). therefore, we have to acknowledge the relevance of the socioeconomic security of ict, i.e., the security of an individual, an economy, a society and a country against the consequences of decisions made by modern ics (including decision support systems applied in economics, military science, politics, etc.). 26 d.a.novikov: big data and big control discuss a series of questions. which data are actually big? where do big data arise? what are the modern applications of big data? how can one improve the efficiency of their usage in future? what is the role of science in big data handling? 4 sources and customers of big data they include large groups, namely: science (astronomy and astrophysics, meteorology, nuclear physics, highenergy physics, geoinformation systems and navigation systems, distant earth probing, geology and geophysics, aerodynamics and hydrodynamics, genetics, biochemistry and biology, etc.); internet (in the wide sense) and other telecommunication systems; business, commerce and finances, as well as marketing and advertising (including trading, targeting and adviser systems, crm-systems, rfidcradiofrequency identifiers used in sales, transportation, logistics and so on); monitoring (geo-, bio-, eco-; space, air, etc.); security (military systems, antiterrorist activity, etc.); power engineering (including nuclear power engineering), smartgrid; medicine; governmental services and public administration; production and transport (objects, units and assemblies, control systems, etc.). numerous applications9 of big data in these fields can be found in popular science literature (or even “glossy” journals) available at public internet sources. we will not describe these applications here to avoid embarrassing “zettabytes” and “yottabytes”. in almost all fields cited, the modern level of automation is such that big data have automatic generation. therefore, the following question gains growing importance. what is the volume of “lost” data flows (due to insufficient capabilities or time for their storage or processing)? this question seems correct for an engineer in ict, but not for a scientist or a user of big data processing results. rather, the former and the latter would ask “what are essential losses in this case?” and “what are the changes if we successfully acquired and processed all data?”, respectively. 9the principal idea of using big data is revealing “implicit regularities”, i.e., answering nontrivial questions: epidemic prediction based on information from social networks and sales in drugstores; medical and technical diagnostics; retention of clients by analyzing sellers’ behavior in stores (the spatial movements of rfid-tags of products); and others. advances in systems science and application (2015) vol.15 no.1 27 5 which data are big? scientific challenges traditionally, big data are unstructured data whose volume exceeds the available handling capabilities in required time. however, this definition appears somewhat“cunning”: data considered big today cease to be such tomorrow owing to the progress of data handling methods and means. data that looked big several hundreds or even thousands of years ago (in the absence of automatic treatment) are easily processed today by home computers. the competition between the (hypothetic) computational demands of mankind and corresponding technical capabilities has been known very long ago. of course, the capabilities have been always chasing the needs. and the gap between them represents a monumental stimulus for science development. researchers have to suggest simpler (yet, adequate) models, design more efficient algorithms, etc. sometimes, the definition of big data includes the so-called 5v properties (volume, velocity, variety, veracity, validity). alternatively, the difference between the big volume of conventional data and big data proper is that the latter form the big flow of unstructured10 data (in the sense of volume and velocity as the volume per unit time). in the wide comprehension, the unstructuredness of big data (text, video, audio, communications structures, etc.) is actually their characteristic feature and a challenge for applied mathematics, linguistics, cognitive sciences and artificial intelligence. creation of real-time processing technologies11, including the feasibility of implicit information revelation, for large flows of text, audio, video and other information forms the mainstream of applications of the above sciences12 to ict. therefore, we observe a direct (and explicit) query from technologies to science. the second explicit query concerns adaptation of traditional statistical analysis, optimization and other methods to big data analysis. furthermore, it is necessary to develop new methods with due consideration of big data specifics. a modern fashionable trend is boosting analytics tools (generally, business analytics) for big data. but their list almost coincides with the classical kit of statistical tools (or is even narrower, since some methods are inapplicable to big data). this is also the case for: machine learning methods (neural networks, bayesian networks, fuzzy inference and other logical inference, etc.); 10data unstructuredness can be the result of their omissions and/or different scales of studied phenomena and processes (in space and time, see the so-called multi-scale systems). 11in the first place, these technologies must perform data aggregation (e.g., detecting changes in technological data or storing aggregated indices). really, one does not need all data (especially, “homogeneous” data). 12mathematics rather easily operates structured data; and so, data structuring makes an important problem. 28 d.a.novikov: big data and big control high-dimensional optimization problems (in addition to traditional parallel computing, intensive research focuses on distributed optimization, l1-optimization; e.g., see [8-13]); discrete optimization methods (here an “alternative” lies in application of multiagent program systems [14]) and others. the common feature in the stated queries of technologies to science is adaptation or small modification of well-known tried-and-true methods. we have to be aware of the following. generally, automatic modeling (by traditional tools13) based on raw data represents just a fashionable delusion.14 we expect to suggest algorithms and apply them to bulky volumes of unstructured (often irrelevant) information, thereby improving the efficiency of decision-making. such delusions occurred in the history of science at the early development stages of cybernetics and artificial intelligence.15 they resulted in numerous disappointments and considerably impeded the development of these scientific directions. there exist no miracles in science: generally, new conclusions require new models and new paradigms (e.g., see the books on science methodology [15-16]). 6 general challenge the complexity of the surrounding world grows at a smaller rate than the capabilities of data detection (“measurement”) and storage. perhaps, these capabilities have exceeded the ability of mankind to realize the feasibility and reasonability of their usage. in other words, we “choke” with data, trying to find what to do with them. however, there exists an alternative viewpoint of this situation as follows. obtaining big data (having an arbitrary large volume) is possible and easy enough (obvious examples arise in combinatorial optimization, nonlinear dynamics or thermodynamics, see below). but we have to understand how to manage big data (and ask the nature correct questions). furthermore, it is possible to construct an arbitrary complex model using big data and then try to reach a higher accuracy within the model. but the associated dilemma is whether we obtain new results or not (in addition to very many new problems16). long ago math13an additional encumbrance is the accumulated experience of a researcher/developer and the traditions of its scientific school. successful solution of a certain problem leads to the conviction that same methods (only!) are applicable to the rest open problems. 14in some cases, additional information can be obtained by increasing the volume of data (under correct processing). 15the behavior demonstrated by a cybernetic system is always the result of embedded algorithms (stochastic, nondeterministic, etc.), despite the seeming generation of new knowledge or revelation of new (“unexpected”) behavior. this is especially the case for interaction of very many elements. 16we recognize the importance of model’s adequacy and stability of modeling results, but omit these problems. advances in systems science and application (2015) vol.15 no.1 29 ematicians and physics knew that increasing the dimensionality and complexity of a model (aspiration for considering more factors and relations among them) does not necessarily improves the quality of modeling results; sometimes, it even carries to the point of absurdity.17 let us study a series of examples. example 1. the book [17] by nobel prize winner h. simon considered the following example. imagine an ant walking along a beach. the ant may try to minimize the efforts required for moving from one point to another; and so, it escapes rocks, sometimes turns back, etc. if we observe just the horizontal projection of the ant’s path (without the knowledge of relief), explaining its behavior (a very winding, complex path) seems difficult. h. simon arrived to an important conclusion. the existing variety and complexity of human behavior are explained not by their complex principles of decision-making (actually, these principles appear simple), but by the diversity of related situations. one would hardly disagree with this opinion. really, nontrivial results can be provided by a complex model based on simple input data and by a simple model based on complex input data. ideally, nontrivial results should be generated by simple models owing to the correct choice of relevant and simple input data (mathematicians are used to say that“simplicity is the sign of truth”). example 2. suppose that, as if by magic, scientists of the 18th century receive a modern laptop and a tomographic scanner of laptops (with user manuals). the scientists make the tomographic image of the laptop and save it in the memory. by analyzing these bulky and very detailed data on the physical structure of the laptop, they would hardly understand the operation principle of the laptop. a correct paradigm, a correct conceptual model forms a necessary (yet, not sufficient) condition of a success. example 3. take any np-hard problem of combinatorial optimization [18], e.g., the travelling salesman problem. there are about n! variants to-be-analyzed for exact solution search; in the case of n∼100, the number of variants exceeds the computational capabilities of mankind. the network excitation control problem for a real social network possesses the dimensionality of n∼ 106[19]. this property of np-hard problems has been known for decades, motivating researchers to develop (generally, heuristic) methods of approximate solution search with estimated guaranteed accuracy at reasonable time. a similar example concerns nonlinear dynamics models: the“observations” of dynamic chaos demonstrated by a rather simple (medium-dimensional) nonlinear dynamic system may require the memory of all computers in the world, but generate no new knowledge. example 4. a classical example from the history of physics is the discovery of 17not to mention situations, when existing scientific paradigms make it impossible in principle to model system behavior on a large time horizon (e.g., accurate weather forecasting). 30 d.a.novikov: big data and big control the law of universal gravitation. for two decades, t. brahe (1546−1601) observed planetary motion in the solar system. for those days, his records represented big data. based on brahe’s records, i. kepler (1571−1630) formulated his empirical (!) laws of planetary motion: the orbit of a planet is an ellipse with the sun at one of the two foci; the square of the orbital period of a planet is proportional to the cube of the semi-major axis of its orbit. keplers laws of planetary motion aggregated brahes data, and the motion of any planet could be calculated by them (instead of brahe’s many-volumed recordings) with a high accuracy. in other words, brahe learned to describe18 planetary motion, whereas kepler learned to describe and predict it. however, kepler’s laws did not explain why planets move so. the answer was later given using the law of universal gravitation established by i.newton (1643-1727). perhaps, kepler’s laws would be derived by any modern computer (with adequate algorithms embedded and initial data entered). at the same time, modern computers would be unable to obtain the law of universal gravitation without the corresponding model of mass interaction. kepler’s laws are the “corollaries” of the law of universal gravitation (can be deduced from it), just like brahe’s results follow from kepler’s laws. therefore, the law of universal gravitation made useless (epistemologically superfluous) both kepler’s laws and brahe’s big data. today’s experience in big data handling testifies that, in most cases, we are at the level of brahe and apply titanic efforts to reach the level of kepler. but a qualitative jump occurs only under the appearance of generalizations (at the level of newton) that radically simplify the situation. once again, note that the whole secret is“putting correct questions to the nature.” example 5. the second classical example from the history of physics relates to the development of molecular-kinetic theory. it shows that generation of a big data flow appears easy. the whole point is what to do with this flow and what questions to answer. consider the following mental experiment (the problem of ideal gas behavior description). one cubic meter of air is in the normal conditions. it contains approximately 1025 molecules. the motion of molecules and their collisions are completely described by kinematics and dynamics (within the framework of the ideal gas model). in other words, there exist no fundamental obstacles in such description. during a second, each molecule suffers from 109 collisions with other molecules. the real-time description of such a system (the coordinates and velocities of all molecules) requires the minimum data flow of 1035 bytes per second. 18recall the basic functions of scientific cognition (particularly, modeling) [16]: description (the phenomenological function)explanationprediction (the prognostic function) control (the normative function). advances in systems science and application (2015) vol.15 no.1 31 this data flow exceeds even the modern processing capabilities of mankind (to say nothing of the processing capabilities in the 1850s when the theory appeared)! perhaps, physicians recognized the senselessness of such detailed analysis (even despite its possibility in principle) and passed to macrodescription in terms of aggregated characteristics (temperature, volume, pressure) and, subsequently, to the description of probabilistic distributions within the framework of statistical physics. but if physicians of that period were able to perform necessary calculations, science would never have statistical physics! today a similar effect arises, e.g., in informational control problems for online social networks. transition from microdescription using graphs with tens of millions of nodes and billions of connections (no doubt, this is a big control problem) to macrodescription in terms of probabilistic distributions [20] leads to realistic theoretical study, still preserving the key properties of an object. and finally, note the following. depending on their source, big data can be natural or artificial. in the former case, data are generated by some independent object and we (“investigators”) decide what should be “measured” (examples 1, 2 and 4). in the latter case, the source of data is a model (examples 3 and 5); complexity (data flow) is partially controlled and defined during simulation. “recipes.” there exist four large groups of subjects (see fig.5) operating (explicitly or implicitly) big data in their professional (scientific and/or practical) activity: manufacturers of big data handling tools (software/hardware developers, suppliers, consultants, integrators, etc.); designers of big data handling methods (experts in applied mathematics and computer science); specialists in application domains (scientists focused on real objects or their models) that represent big data sources; customers utilizing or planning to utilize the results of big data analysis in their activity. representatives of the mentioned groups interact with each other (see the dashed lines in fig.5). the normative (“ideal”) division of “responsibility areas” is illustrated by fig.6; here the thickness of arrows corresponds to the level of involvement. proceeding from sensus communis and not claiming to be constructive, we formulate the following general “recipes” for the listed groups of subjects. for manufacturers of big data handling tools: with the course of time, it will be difficult to sell big data solutions (including analytical ones) without suggesting new adequate mathematical methods and stipulating for the feasibility of close cooperation between customers, the developers of appropriate methods and specialists in application domains. for mathematicians (the author’s “brothers-in-arms”): a topical query con32 d.a.novikov: big data and big control fig.5 subjects operating big data fig.6 division of “responsibility areas” cerns adapting well-known methods and developing new processing methods (in the first place, with nonlinear complexity!) for the large flows of unstructured data representing a good testing area for new models, methods and algorithms (to the extent possible, at the expense of manufacturers and/or customers). for specialists in application domains: big data technologies lead to new capabilities for acquiring and storing the bulky arrays of “experimental” information, conducting the so-called computing experiments; the associated methods of applied mathematics enable systems generation and rapid verification of hypotheses (revelation of implicit regularities). for customers: the expensive technologies of big data acquisition and storage would hardly be economically sound without involving specialists in appropriate methods and subject areas (only if it is absolutely clear which questions a customer would like to answer using big data19). some threats. in addition to the emphasized necessity of searching for ade19though, it is possible to store data de bene esse (e.g., to verify a certain hypothesis in future based on them). advances in systems science and application (2015) vol.15 no.1 33 quate simple models and the alerting trend of anticipatory technologies development, we expect the future relevance of the following problems (the list below is unstructured and incomplete). • the informational security of big data. this requires adaptation of wellknown methods and tools, as well as development of fundamentally new ones. • the energy efficiency of big data. even today, data processing centers represent a considerable class of power consumers. the bigger are data to-beprocessed, the higher is energy needed. • the principle of complementarity was established in physics long ago; it declares that measurements modify the state of a system. however, does it apply to social systems whose elements (people) are active, i.e., possess their own interests and preferences, choose their actions independently, etc. [21]? a demonstration of this principle lies in the so-called information manipulation (strategic behavior). according to theory of choice [4,22], an active subject reports information by forecasting the results of its usage; generally speaking, an active subject does not adhere to truth-telling. another example concerns the so-called active forecasting: a system changes its behavior based on new knowledge about itself [23]. are these and similar problems (e.g., crowdsourcing [4,24], conformity behavior [25], etc.) eliminated or aggravated in the case of big data? • as far as we have mentioned the principle of complementarity, it is necessary to recall the principle of uncertainty in the following (epistemological) statement [16]: the current level of science development is characterized by certain mutual constraints imposed on results “validity” and results applicability. in the context of big data, this principle means the existence of a rational balance between the level of detail in the description of a studied system and the validity of results and conclusions to-be-made on the basis of this description. • a traditional assumption in design and operation of information systems (corporate systems, decision support systems of governmental services, inter-agency circulation of documents, etc.) is that all information in such systems must be complete, unified and publicly available (under existing access rights). but it is possible to show the “distorting-mirror” reality to each person, i.e., to create an individual informational picture20, thereby performing informational control [4, 7, 21, 23]. should we strive for or struggle against these effects in the field of big data. 20at the very least, a fragment of the “objective” picture (hushing up the whole truth); at the most, an arbitrary inconsistent system of beliefs about the reality. 34 d.a.novikov: big data and big control 7 instead of conclusion thus and so, data have always been big. new, more and more perfected tools of big data acquisition, storage and processing appear intensively. we wish we managed to perform these operations in real time. this calls for developing appropriate directions of applied mathematics and computer science (a topical query from technologies and practice to modern science). there also exists a pressing need for mass training of specialists in big data, big analytics and big visualization (with focus on concrete application domains). however, this is not enough: we have to accumulate knowledge (in appropriate branches of science) and create models for compact and adequate description of studied phenomena and processes (with due consideration of a solved problem). in other words, it is desired to advance from the level of brahe to the level of newton in each possible area of big data applications. otherwise, we are doomed to handle particulars, not seeing the wood for the trees (see e. rutherford’s citation above). furthermore, the anticipatory development of technologies has become the civilization-scale problem to-be-accounted by scientists and engineers (including the field of big data and big control), as well as by the consumers of appropriate methods and tools created by them. references [1] special issue. (2008), nature. [2] forrest j. and novikov d. (2012), “modern trends in control theory: networks, hierarchies and interdisciplinarity”, advances in systems science and application, vol.12, no.3. [3] dorogovtsev s. (2010), lectures on complex networks, oxford university press, oxford. [4] gubanov d., korgin n., novikov d. and raikov a. (2014), e-expertise: modern collective intelligence, springer [5] jackson m. (2010), social and economic networks, princeton univ. press, princeton. [6] novikov d. (2014), “games and networks”,automation and remote control, vol.75, no.6, pp. 1145-1154. [7] gubanov d., chkhartishvili a. and novikov d. (2010), social networks: models of informational influence, control, and confrontation, ed. by d.novikov, fizmatlit, moscow, russian. advances in systems science and application (2015) vol.15 no.1 35 [8] noam nisan and tim roughgarden. (2009), algorithmic game theory, ed. by n. nisan, t. roughgarden, e. tardos, and v. vazirani. n.y, cambridge university press. [9] boyd s., parikh n. and chu e. et al. (2011), “distributed optimization and statistical learning via the alternating direction method of multipliers”, foundations and trends in machine learning, no.3, pp.1-122. [10] granichin o.n. and pavlenko d.v. (2010), “randomization of data acquisition and ℓ1-optimization (recognition with compression)”, automation and remote control, vol.71, no.11, pp. 2259-2282. [11] nesterov y. (2012), “efficiency of coordinate descent methods on hugescale optimization problems”,siam journal on optimization, vol.22, no.2, pp.341-362. [12] polyak b.t. and tremba a.a. (2012), “regularization-based solution of the pagerank problem for large matrices”, automation and remote control, vol.73, no.11, pp.1877-1894. [13] shoham y. and leyton-brown k. (2009), multiagent systems: algorithmic, game-theoretical and logical foundations, cambridge university press, cambridge. [14] wooldridge m. (2002), an introduction to multiagent systems, john wiley & sons, n.y. [15] kuhn t. (1962), the structure of scientific revolutions, university of chicago press, chicago. [16] novikov a and novikov d. (2013), research methodology: from philosophy of science to research design, crc press, leiden. [17] simon h. (1996), the sciences of the artificial,3rd edition, the mit press. [18] garey m. and johnson d. (1979), computers and intractability: a guide to the theory of np-completeness, w. h. freeman and company, san francisco. [19] novikov d. (2014), “models of network excitation control”,procedia computer science, vol.31, pp.184-192. [20] breer v., novikov d. and rogatkin a. (2014), “microand macro-models of social networks”, control sciences, no.5, pp.28-33. 36 d.a.novikov: big data and big control [21] novikov d. (2013), theory of control in organizations, nova science publishers, n.y. [22] aizerman m. and aleskerov f. (1995), theory of choice, elsevier, amsterdam. [23] novikov d. and chkhartishvili a. (2014), reflexion and control: mathematical models, crc press, leiden. [24] surowiecki j. (2005), the wisdom of crowds, anchor, n.y. [25] breer v. and novikov d. (2013), “models of mob control”,automation and remote control, vol.74, no.12, pp.2143-2154. corresponding author d.a.novikov can be contacted at:novikov@ipu.ru advances in systems science and application (2015) vol.15 no.1 60-71 tackling foundations of systemic complexity on reality and decision systems hermı́nio duarteramos emeritus professor, new university of lisbon, lisbon abstract what is reality? that’s an old question, a big question raised by so many thinkers, since very ancient times. many responses has been given, so many answers are purposed in these days. here we advance on quantum territories and we also approach the human decision theory to interpret the complex meaning of reality under new systemic complexity viewpoints. keywords reality, systemic complexity, decision system, qbism, bayesian probability. 1 introduction traditionally, thoughts on reality were put forward among philosophers and scientists, around force lines of idealism and materialism dipoles, as plato (427-347 bc) and lucretius (99-55 bc) traced on matter, or over generated currents between theism and deism bipoles, as wilhelm leibniz (1646-1716) and albert einstein (1879-1955) did conduct about the universe [1]. today we find some other ways of thinking, based on new scientific interpretations, using abstract fields of mathematics, as john von neumann (1903-1957) supposing a consciousness creation for object attributes, and quantum physics, as niels bohr (1885-1962) negating deep reality down phenomenal facts [2]. nowadays, what can we say about the meaning of reality? first of all, it is a general topic, which appears inevitably at the front panel of globalization’s science, thinking on worldwide vertiginous trends of scientific research without national frontiers. the reality meaning is a complex problem to understand systems we meet on nature, focusing our consciousness on things, because we aren’t aware of essentials on material structures and human-nature interactions.nevertheless, today we know much better what reality can be than some centuries ago, and even decades behind or from late years. here i join three basic ideas, integrating them with the general system theory to interpret what reality is, namely the wavefunction for the quantum formulation of the ultimate matter, the complexity of an observed system and the subjective probability on decision making.in such way, i briefly tackle how to set up general foundations of the so called systemic complexity, hopping to contribute with a new and powerful vision on system sciences. it is inevitable to approach physical and mathematical concepts, at least evocating some descriptive principles, because physics and mathematics are sciences advances in systems science and application (2015) vol.15 no.1 61 of the reality modelling, that humans use to understand the world around. 2 what is reality? outside my mind there are things on the objective (or concrete) world, which are only matter and signals, and inside the mind it runs a subjective (or abstract) world with forms of ideas and signs1. the human detects signals that are radiated from matter and he sees or hears the impressive reality, and then he judges on what he is seeing or hearing by decision making on some important cues, which are detected signals contributing to the awareness of reality. on one hand, the physics (the natural philosophy in past times) measures and relates features of the objective reality, since electromagnetic waves as signals of light until quantum waves as signals of the ultimate matter, i.e., the physics objectivity interprets the external reality by its fundamental properties, and so the objectivity in physics doesn’t describe the world “as it really is” [3], but it is an epistemic knowledge devoted to measurable features of the world, on which an open worldwide community of independent observers do agree. on the other hand, cognition sciences (encompassing the brain, the mind and the soul) study features of the subjective reality, since sensing signals and their perception until the ultimate decision on inner appearances, i.e., the cognitive subjectivity also represents interpretations on the outer reality. when plato’s prisoners emerged from the cave, they were only able to make out shadows and reflexions as reality, not just because the sunlight burned their eyes but also due to the subjective probability of decision making in the brain’s mind. within this frame, i wish to participate freely on the open discussion of some fundamental thoughts on complexity, to better understand systems both of objective and subjective realities, trying to discover an all-embracing model for a single scientific paradigm of reality systems. each one of us makes decisions taking systems from the natural matter and observing their behaviour, and then we interpret what it is going on actually in our mind as a real representation. how can we do that?basically, a system is one concrete structure of several interconnected parts within a limited space and interacting in concrete modes (doesn’t a matter what they are at the moment), emerging to us inherent black box outputs2, as reactions to eventual stimulating inputs or owing to its own 1we don’t discuss here the mind’s cognition nature of brain’s imagerial and soul’s imaginal information processes. 2the system science uses efficiently the concept of “black box” and its “outputs”, being equivalent to the whole functional behaviour and ignoring the real system structure. 62 hermı́nio duarteramos: tackling foundations of systemic complexity on reality... autonomous existence, besides the necessary vital energy feeding things to exist with an own intentionality. 3 systemtic complexity we call systemics the science on the function integration of whatever composite system in order to process its behaviour by a functional model.in open systems the outputs emerge from the internal operation to the outward environment, and otherwise inward outputs define closed systems. generally speaking, the functional organization of a given system is characterized by the following four operational systemic essentials: the acrony or composition, the axony or interactivity, the aquadry or boundary, the adaptacy or optimum workability.however, the most notable feature is that the real system always denotes an existential intentionality that must be truly expressed by their outputs into the environment, which may be detected by instruments or human observers. these external information signals establish a quintessential systemic principle, the telonomy or aim outputs representing the intrinsic intentionality of the system in action [4].that is to say, a system accomplishes its existential purpose if its bulk operation obeys to those five systemic principles, and we need to know them clearly in order to acknowledge its real behaviour. in complex systems (distinct from complicated ones) we don’t know quite well one or more systemic essentials, otherwise the correct knowledge of all these characteristics defines a simplex system (distinct from simple ones).in the case of a simplex system all systemic essentials are well known, and so we have a definite acrony and complete axony and certain aquadry and determinate adaptacy and an accurate telonomy. the four operational systemic essentials of a simplex system may be expressed mathematically onto a transfer function3 to calculate its right behaviour, denoting the systemic telonomy at each time, but complex systems incorporate additional barriers for obtaining reasonable outcomes, which are variable with the correctness of systemic essentials. the degree of complexity for the system’s observer is related to the imperfect knowledge he has on systemic essentials, for instance, the complexity of a simplex system is zero degree, and a complex system of 2nd degree has two systemic essentials that we don’t known entirely, either indefinite acrony or incomplete axony or uncertain aquadry or indeterminate adaptacy or inaccurate telonomy. 3the transfer function of a system requires initial conditions to be defined in the complex frequency domain s = σ + iω, where σ is a real frequency and iω is the imaginary angular frequency. advances in systems science and application (2015) vol.15 no.1 63 the concept of systemic complexity4 was introduced to approach complexities encountered in material or concrete (objective) systems, and the same functional methodology was extended to mind or abstract (subjective) systems, both trying to reduce complexity into simplexity. here they are some foundations of the systemic complexity of general and singular systems to interpret the observed natural world, fusing accessible conscious information on imagerial and imaginal mind processes.we use it to acknowledge both objective and subjective reality, for the near contact world till the ultimate mater in the void real space. the world scientific community in the twentieth century did interpret the reality using thoughts derived from the quantum theory, since the famous “copenhagen interpretation” by the physicist niels bohr supposing that reality is created by observation, and the wholeness of david bohm (1917-1992), until parallel universes purposed by hugh everett (1930-1982), and also the world consisting of potentials and actualities under the uncertainty principle of werner heisenberg (1901-1976) considering somewhat as probability waves in the middle between possibility and reality, after louis de broglie (1892-1987) having translated the reality into a “statistical interpretation” [2]. what is the quantum interpretation of the ultimate reality? it’s known that the wavefunction ψ(r, t) was formulated by the physicist erwin schrödinger (1887-1961) describing the behaviour of elementary particles through an equation5, which is interpreted as representing probability values of particle states on its existential trajectory (traced by the space-vector r at each time t). this mode of thinking do gives us a way to predict the electron behaviour on the move, calculating its own evolution in space, and pointed out a potential way to interpret the real face of things. the wavefunction corresponds to the frequency probability6 of the state occurrence, which is an objective probability, for instance, it’s value is 50 % in case 4the adjective systemic clearly qualifies the complexity we are talking about, distinguishing it from other complexity types as we find in mathematics (kolmogorov complexity) or in computer sciences (algorithm complexity) and also in physics (pattern complexity) [5]. 5the time dependent schrödinger equation for a single non-relativistic particle moving in an electric field is i~ ∂ ∂t ψ(r, t) = [ − ~2 2m ∇2 + v (r, t) ] ψ(r, t) where ~ is the constant’s planck, m is the particle’s mass, v is its potential energy, ∇ is the laplacian, and ψ (r, t) is the position-space wavefunction. 6frequency (or objective) probability po is defined by the ratio of the number of favourable cases (nfav) and the number of possible cases (npos), or mathematically po = nfav / npos. 64 hermı́nio duarteramos: tackling foundations of systemic complexity on reality... of an electron7, because this particle assumes at each time just one natural state between two possible (up/down) spin states. the logic consequence is that the quantum reality itself has a probabilistic character, while our common world reality seems to exclude a statistical temper. from that quantum narrative, the so called “schrödinger’s cat paradox” shows an intriguing quantum ambiguity of the cosmic nature, having been very much discussed among thinkers on the quantum world [6].it can be analysed considering a closed system formed by a cat inside a sealed room, where there is a glass flash full of poison, besides a radioactive source emitting particles into the inner space and also a particle decay detector which will trigger automatically the fall of a heavy hammer on the flash immediately after detecting a preset particle decay threshold, shattering it and pouring the poison out.thus, the cat will be killed, but in the meantime the cat was simultaneously alive and dead, according to the wavefunction of the particle represented by state probabilities.however our direct observation certificates correctly if the cat is alive or not, since it cannot be in both states at the same time. the paradox lies in the fact that the animal is alive and dead inasmuch we don’t observe what is going on inside the hermetic room (the cat doesn’t look at the emitted particles for sure), and when an observer looks at the particles he can say rigorously the cat is either alive or dead. this happens after the superposition of quantum states ends by observation, producing the collapse of the wavefunction, and then it doesn’t represent any more an objective probability, being the observed reality inevitably either one or another possible concrete real state. that was the reason why albert einstein felt serious doubts on quantum theory implications, and he did write to his friend max bohr (1882-1804) “i cannot believe that god plays dice with the universe” [7].during decades many scientists endeavoured to explain such logic contradiction of the strange quantum paradox.nevertheless, they struggled without success.let us now analyse the schrödinger’s cat system under the systemic complexity point of view. the acrony is definite because we are aware of all components inside the closed room; the axony is complete once interactions will follow exactly some serial causations of the automation mechanism breaking the glass with poison; the aquadry is certain due to rigid boundaries build by room walls; the adaptacy is determined supposing the structure will work perfectly as it was planned. so it seems to be a simplex system, but we can’t attribute a single state to the cat, the quantum ambiguity remains and we continue ignoring the true reality of 7the electron is an elementary particle of fixed mass and electric charge (corpuscular reality, discovered in 1897) and also found to be a wave (waving reality, discovered in 1927). advances in systems science and application (2015) vol.15 no.1 65 the system before its observation. why? because the system is closed, and the systemic theory can be effectively applied only to open systems, measuring their outputs or system telonomy, i.e., the observer belongs to the global system of reality acknowledgement, and he cannot be ignored.notice that the so called observer can be a human or an instrument, and for this the observation may correspond to a measurement whose result is a real measure. in the closed system we can say that the non-observed (or not measured) reality is represented by the state objective probability ascribed to the wavefunction, according to the quantum theory, and the observed reality on the consequent open system is described by a physical state convincing the observer about the reality character.before observation, in spite of the observer awareness on the inner operating systemic essentials, he didn’t get any external system telonomy and the reality was complex to him. when we observe the reality of the open system we verify the cat is alive or dead without any doubt, and before observation the reality of the closed system was ambiguous (due to the quantum wavefunction), being the cat alive and dead, although we were conscious that just one of two possibilities could arrive, saying that the cat is either alive or dead. that’s the real problem, a restless and/or question to the reality in those closed/open systems, revealing a fundamental difference between the quantum world, which is an ocean of waves, and the macroscopic world on the human scale, where we live.at the existential limit of the nature, the quantum theory makes out the reality with all possible objective probability states, and at the human existential nature, the reality to observers means something concrete far from objective probability of real states. each part of the reality on our newtonian world exists each time apparently with only one form, because multiple own states may not superpose at the same time in any physical system. what is such potential concreteness of our observation? maybe it is the old concept potentia of the aristotle (384-322 bc) philosophy, maybe it isn’t.recently that contradiction of the quantum theory has been overshoot by the so called qbism 8, applying the subjective probability of bayes to the wavefunction. 8the term qbism is derived from the quantum bayesian view, using the bayesian probability of elementary particle states in the wavefunction. 66 hermı́nio duarteramos: tackling foundations of systemic complexity on reality... 4 tackling foundations of systemic complexity in the eighteenth century, the english presbyterian and statistician thomas bayes (1701-1716) had a different idea on the probability notion, and in 1812 the french mathematician and astronomer pierre-simon, marquis de laplace (1749-1827), disseminated it among other scientists of the natural philosophy, gaining credibility last century in several statistical applications, as in economics and also engineering decision and in cognition sciences. in fact, the bayesian probability 9 measures the human conviction on estimated cues or contents of signals that we detect from the reality to make decisions.for instance, it happens when a physician regards his patient and evaluates the symptoms to diagnose a disease, and he would be self believed on the most credible hypothesis, guessing it by means of subjective probabilities, even he uses a practical method giving weights to the collected data.such subjective probability quantifies the human conviction level on the observed reality, being entirely different from the objective probability based on illness occurrences, which results from the disease frequency and it is not related to the human convincement about the reality itself. the modern decision theory has applied successfully the bayesian probability, conveying an important tool for engineers, economists, sociologists and other users, and even to solve trivial situations we have in our everyday life ourselves.undoubtedly, we are observers of the reality and making decisions always with subjective probabilities on all things we experience around us. now we pay attention to quantum reality supposing that wavefunction superposes subjective (or bayes) probability states instead objective (or frequency) probability states.reasoning like that we don’t see any paradox. indeed, a human can guess the schrödinger’s cat state in the closed system, assigning it a state either alive or dead, according to the subjective probability of his creed, and because the quantum reality to him is what the wavefunction assigns to be in reality. 9bayesian (or subjective) probability is defined by the bayes’ rule p (h|e) = p(e|h) p(e) .p (h) where | denotes a condition probability; h stands for competing hypothesis (competitors) from which one chooses the most probable; the evidence e (cue) corresponds to new data that were not used in computing the prior probability; the prior probability p(h ) is the probability of h before e is observed; the model evidence p(e) or the marginal likelihood (diagnosticy weight) is the same for all possible hypothesis being considered; the likelihood function p(e |h ) is the probability of observing e given h (impact weight); the posterior probability p(h |e) is the probability of h given e, i.e., after e is observed (credibility weight); the impact factor p(e |h ) / p(e) represents the impact of e on the probability of h. advances in systems science and application (2015) vol.15 no.1 67 such reasoning affirms that a quantum state is just a subjective probability assignment, and the reality of the quantum world tosses the observer’s beliefs to elect that one most evident.this is hard to accept for the traditional scientific thought, “it is really hard to believe something you don’t actually believe” [8], yet for a qbist it is a consistency question, considering the subjective probability a logic extension of the truth table.that extended logic encourages us to generalize the quantum reasoning to the macroscopic world, as we find it explicitly in decision system methodologies. we concern the collapsed wavefunction after observation (or measurement) which is characterized by subjective probabilities, and so the observers’ personal information gives us expectations or degrees of belief.when i look at the world nature my mind interprets detected cues by sensing and perception of the natural world as appearances in correspondence to reality. whether i don’t get information signals from the reality, this reality exists but i can’t say anything about it, and whether i detect some real signals i will build a creed on reality, assembling captured cues in my mind and weighting them to obtain the most evident subjective probability to make decisions.this is the apparent reality or the reality we think it is. consequently, the matter on macroscopic world relies on the subjective probability, as well as it must occur on the quantum level when it is observed (or measured).that’s a question of believing, but to believe grounding on a probability induced by the own reality.the behaviour of the observed quantum world is the same as that of our macroscopic world, taking in account the reality reveals the same physical principles. all this means that the natural interpretation of the reality is the reality itself, even when it exists inside a closed quantum system or whenever we are not observing it, and so the material substance is nothing more than a human judgement, reasoning on detected signals radiated from matter to process the subjective probability according to the intrinsic character of things.the reality remains “the thing-in-itself”, as immanuel kant (1724-1804) said, however to me it is the appearance i judge to be the reality in extremely rapid decision processes. we arrive to the conclusion that the reality exists as it is, but our consciousness decides its appearance using a very pragmatic heuristic, estimating some cues by perception in very short time, and seeking the accumulated evidence to adopt cue weights for diagnosticy and owing to their credibility proportions for each competitive hypothesis, integrating both weights in products as evidence factors, and 68 hermı́nio duarteramos: tackling foundations of systemic complexity on reality... choosing the evident hypothesis with the maximum sum of those factors 10.this optimization heuristic translates the bayesian probability into our everyday mental process, for instance buying a shirt or choosing a new car to buy it. whether the decision time is very short we feel acting under intuitive pulses, judging without any elaborate reasoning, and we can say the intuition feeling relies upon a very rapid decision making.this is an interesting by-product of the present search.the truth is that when we are observing our world we are always reasoning with subjective probability pulses. i see the green forest because i detect light cues from the trees using my eyes (and the integrated biological vision system), and they convince me that the tree colour is green, matching the wave frequency 11 of the light detected from trees with the learned language stored in my neural memory. the consciousness answers to the subjective probability of the real wavefunction whenever the human observes the reality, because it appears as an open system, whose telonomy is the appearance of the reality that affects him.therefore, the observer is not separated from the observed [9], they both interact intimately. it comes now the opportunity to ask if all systems are real or not.taking systems from the reality we observe them without any doubt, each concrete system belongs to the reality, doesn’t matter it is a macroscopic or a quantum system, and its systemic model is also real, because i design it in a representative abstract nature but its functional operation corresponds to a concrete system. a system model makes decisions exactly as the natural system itself, supposing systemic essentials are perfect in the case of simplex systems, or it approximates more or less right outputs in complex model systems with imperfect systemic es10the decision theory defines a heuristic based on the cue diagnostic [p(e)], that ranks detected cues with normalized weights according to their importance for the decision making, and also on the cue credibility [p(e |h )] of each hypothesis representing cue confidences by their impact weights, as we can see in the following table with three cues, and choosing the best hypothesis between two competitors (hypothesis a and b), defining evidence factors [p(e).p(e |h )] for each cue and each competitor by the product of diagnosticity weights (wdi) and credibility weights (wci), and electing the maximum sum of evidence factors: 11here frequency means the number of wave cycles per second (not the frequency of occurrences in an arbitrary gap of time, which is used to define objective probabilities). advances in systems science and application (2015) vol.15 no.1 69 sentials.all this is what an observer can infer observing both natural and model realities and using detected cues from outside of both black boxes (of natural and model systems), because results are exactly the same for human consciousness if system outputs are equal, or we approach them as much as they seem to be similar. for instance, a refrigerator’s user doesn’t need to know how the refrigerator system works, being only concerned on their cold outputs, but in the case of a refrigerator’s designer he must be aware of its five systemic essentials to produce good equipments. concluding, a concrete (natural or model) system means a black box to the observer, and we don’t need any especial knowledge about what is going on inside, supposing it operates as it was designed, because we reduce the reality to telonomy measurements and excluding the side-by-side information that doesn’t matter to the intentionality target. beyond that, decision systems are also real?as we know, in such situations we organize data sets which inform us on reality features, and we decide electing one of them through subjective probabilities, and the choice represents the most valuable or convincing hypothesis.only later real facts will confirm whether our choices are acceptable or unacceptable judgements. in any case, the methodology is universal, based on wavefunction collapses by observation and using the bayesian probability tool, i.e., a decision system is an abstract system but the systemic complexity governs their behaviour as well as in concrete natural systems.this important corollary is proved by systemic essentials of decision systems analysing the character of detected cues. the acrony is definite because in general we can fairly estimate all diagnosticity weights before comparing hypothesis; the axony is complete when we can clearly guess credibility weights of cues for competitive hypothesis; the aquadry is uncertain because it is possible missing some relevant cues, owing to the proper finiteness of cue boundaries; the adaptacy is determined in the case of impartial or not biased choices, otherwise it will come some errors by overconfidence bias and psychological anchoring heuristics, provoking undetermined optimum working points. following these main lines, normally a decision system is an open complex system, the telonomy being the best possible output from global weighted cues and their mutual relationships, although we cannot always assume the optimum decision. what we must do is trying to reduce the degree of complexity, recollecting better information on systemic essentials, endeavouring to attain simplexities or at least approximating as much as possible to simplex systems. 70 hermı́nio duarteramos: tackling foundations of systemic complexity on reality... here it is the recommended strategy. 5 conclusion in conclusion, this approach lengthens the copenhagen interpretation, which judges that reality is created by observation, defending the thesis it is much more than that by reasoning on the reality as a mind formulation founded on subjective probabilities. indeed, i can summarize the previous discussion in few basic items, raising important issues to be debated in general and in particular cases, as it follows: 1. the reality means my belief on decision making by perception of signals, which are radiated from things as system acrony (measurement problem). 2. i use a subjective probability scheme in my mind to decide what reality is, choosing the most convincing accumulated axony (subjective problem). 3. i can’t know the total meaning of the observed reality, because the biological structure of my body limits every systemic aquadry (boundary problem). 4. the correlation among collected data and psychological bias and heuristics may deviate from the best choice, and i cannot assure attaining the optimum system adaptacy (optimum control problem). 5. common decisions in our current life are intuitive, we make them in extremely short time and so the reality seems apparently to be evident from the system telonomy (consciousness problem). we are now seeing why i am convinced that kant’s appearance of reality is complex under the systemic standpoint, and so i hope we will be able to handle complex systems of worldwide realities developing the systemic complexity theory by scientific methodologies. actually, science research must get down to essential fundamentals, and science theory consists of human rational attempts to represent both appearance and reality of things and ideas. are reality and appearance a unique thing?the science and technology will develop detecting tools for supra-human sensing contacts with non-observable reality by humans, and we research on human-machine interfaces [10] to integrate simplexicity into complexity of nature systems inserting trends to observe the reality itself in a new systemic supra-appearance. that is my conviction, which emerges from decision making in my mind based on some detected signals and memory ones.in reality, that’s the reality. advances in systems science and application (2015) vol.15 no.1 71 references [1] a. north whitehead. (1933), adventures of ideas, cambridge at the university press, london. [2] nick herbert. (1985), quantum reality. beyond the new physics, archor books, new york. [3] john ziman. (2000), real science, what it is, and what it means, cambridge university press, cambridge. [4] hermı́nio duarteramos. (2006), “hard and soft systems intentionality”, 7th congress of ues, lisbon; also in res-systemica, no.7, afscet. [5] g. t. arecchi. (1997), “pattern generation and competition in nonequilibrium optical media: a case study for complexity”, the physics of complex systems, ios press, oxford. [6] david lindley. (1996), where does the weirdness go?, basic books, new york. [7] max born, albert einstein. (1971), the born-einstein letters, walter and company, new york. [8] christopher fuchs. (2012), “interview with a quantum bayesian”, maximilian schlosshauer, (2011). elegance and enigma: the quantum interviews, springer, frontiers collection. [9] david bohm. (1996), on dialogue, routledge, london. [10] r. antunes, f. v. coito, h. duarteramos. (2013), “skill evaluation in point-to-point human-machine operation”, applied mechanics and materials, vol.394, trans tech publications, zurich-durten. corresponding author hermı́nio duarteramos can be contacted at: hduarteramos@gmail.com. microsoft word 3.dan luo, jianhua jiang, buyun sheng--design and implementation of network based collaborative design system. 226-231 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc design and implementation of network based collaborative design system dan luo, jianhua jiang and buyun sheng school of mechanical and electronic engineering, wuhan university of technology, wuhan 430070, china abstract in order to support collaborative design in network environment, according to the aim and requirements of network based collaborative design system (ncds), the functional architecture of ncds was designed and its key technology includes improving of collaborative application sharing, multimedia communication in real-time, management of collaborative interaction and network environment of collaborative work were researched. finally, a ncds was developed and its correctness and feasibility was validated by applied in practice. keywords collaborative design; ncds; product designing; audio & video communication; member management 1. introduction with the popularity of computer network and the implementation of parallel engineering, collaborative design technology[1] has become the research focus in advanced manufacturing domain. to construct an effective, convenient and practical ncds is the precondition of realizing internet based different design for users dispersed geographically. ncds supports product designing, information sharing and exchanging, cax accessing and invocation, product design projects discussing and modifying among product designing engineers in different locations by the communication of audio & video, word and images. in this way, the process of product designing can be carried out across time and space and its cycles be shortened, product development costs be reduced and personalized product development capacity be enhanced substantially. 2. system design 2.1 aims and requirements of ncds the purpose of ncds is to make each participant watch and participate in the real-time collaborative design process, and provide a management tool for design organizer to organize and control the whole process of produce designing. the aim and requirement of ncds is to: (1) provide a platform for information exchanging. it should support real-time and mutual communication in audio & video, text interactive and usb interface based camera images, more than 256 people online simultaneous, audio and video information exchanging with 4 users for each user at the same time. (2) tracking the entire process of remote designing. it achieves screen broadcasting of remote designing and other participants receiving in real-time smoothly. (3) support remote operation of collaborative design. organizer initiates a remote designing and designates the participants at the principle of one designer any time and the designer can only do the operation allowed by the organizer. (4) organize and manage the process of collaborative design. it manages the identity data, password and permissions of designers, and designates the jobs to be finished, allocates permissions to participants or expel a designated staff and manages the process of designing. (5) other requirements. the system should be scalable, interface attractively and easy to advances in systems science and applications (2011), vol.11, no.3-4 227 issn 1078-6236 international institute for general systems studies, inc use. 2.2 functional architecture of ncds according to the aims and requirements of the system, a ncds was designed, it includes audio & video meeting and remote screen broadcasting, remote operation control and console management four parts, as shown in figure 1. n c d s audio and video meeting remote screen broadcast remote operation control console management video & audio communication text charting network communication screen transmit screen receive control client controlled client user management role management figure 1 functional structure of ncds (1) audio and video meeting. it realizes information sharing by audio & video, text and network communication modules. audio communication process includes audio signal acquisition, compression, send and playback while video communication collects data from video card and send it to network after compressed, at same time it extract the received data from other members and playback it on the display, text charting provides real-time text messages interactive and supports text sending for a group, network communication transmits and receives data and choose transmission way according the type of message. (2) remote screen broadcast. it includes screen transmit and receive two modules. the screen transmit captures the screen of server, calculates images difference and sends it after compressed. while screen receive incepts the screen of server and extracts it in real-time, synthesizes images and displays it. (3) remote operation control. the control client simulates the images in local machine according the screen received from server client (controlled client), encodes the operation commands into characters and sent them back to server client. the server client receives control commands from control client and simulates the scene. (4) console management. the user management is responsible for constructing workgroup and managing user data, the role management defines and manages every role in designing process, and manager can change the designer and speaker according the need. 3 .key technology 3.1 improvements of collaborative application sharing collaborative application sharing refers to dispatch application view of server to other clients and collect all other input command to the server. it includes centralized sharing and replication sharing. in allusion to the inherent defects of the two types of sharing technologies, this paper improves the sharing method according to perceived degree of collaboration, collaborative function of collaboration tools and fast transmission of cad models and images. (1) improve the perceived degree of collaboration. for perceiving client information, perceptual information and perceptual result table was used to store online users’ status information both in server and clients as shown in figure 2. node 1 sends the perceived results received by the event sensing module from clients to the server and updates the perception table. in receiving the results, the server updates perceptual 228 luo: design and implementation of network based collaborative design system table and sends the updated information to other nodes, those other nodes then update their perceptual list according received information from the server. in this way, all the nodes can perceive the changing status of every node. figure 2 event sensing process of clients in addition, multi-cursor technology also supports the parallel perceiving of collaboration. all users’ cursors displayed in a screen at the same time and distinguished by different colors for different group or by showing designers’ name under the cursors[3]. (2) improve the function of collaboration tools in order to make common application software supports displaying of multi-cursors, revocation and redoing of operation, decomposition and analysis of input event stream, the interface of common software must be extended and it can be solved by the locking mechanism in application-level. the process of the improving is shown in figure 3 where solid line denotes information flow of collaborative operation and dotted line denotes information flow of collaborative tool. the collaborative tool can realize many control and cooperation mechanism such as certification and conflict coordinate among members, direct interaction, multi-input processing, locking in operation level. figure 3 improvement of collaborative tool to software (3) fast transmission of cad models and images for the problem of low transmission speed caused by large quantity of images data, incremental transfer technology and images compressing from the aspects of time-dependent, image compression and color transformation were used in this paper. time-dependent of screen images refers to changes in the adjacent images confined to a few places and the changes of image color limited in a few tint. so, only the data reflecting changed parts of the screen images should be compressed and transmitted. in this process, some technologies including the message mechanism used to capture the changes of screen images, comparing these changes in a relatively large time interval, inter-frame prediction and compensation were used to solve the inconsistencies between screen changes and the messages. advances in systems science and applications (2011), vol.11, no.3-4 229 issn 1078-6236 international institute for general systems studies, inc small piece based encoding was used to compress images of the screen and its encoded form of data packets as shown in figure 4. any changes in a rectangular area of screen images will be divided into a 16*16 sub-block named a piece, and then appropriate encoding form is used to each piece and a byte is adopted to as sign to show the encoding form of the piece. figure 4 data format of small piece based encoding color transformation refers to change screen data into the transition color (16-bit color) and divide the screen images into small pieces size of 16*16, use small piece comparative law to capture changes of the screen images, and adopt small piece based encoding algorithm to compress changes and re-assemble them into a packet. 3.2 multimedia based communication in real-time the technology of multimedia based collaborative design supported communication includes audio & video capture, its compression and special treatment. as to audio & video capture, video for window (vfw) sdk is used to capture, play and edit audio & video information into avi format, and mpeg4/divx is used to compress, store and transmit audio & video. special treatment of audio system includes voice consistency and delay, mute detection and interactive approach in full-duplex. data buffer technology is used to solve the voice consistency and delay[4]. by coding the sample value of voice and using a serial number to indicate the sample time, the missed number was discarded as to avoid sending mute information. limited input & output scope of audio is used to avoid the echo in full-duplex interactive process. special treatment of video ensures the video receiving quickly by throwing away some video frames. the video frame includes i, p and b three types. i frame is coded frame and it is the most important, while b frame is the least important. so, the principle of abandoning frame is: (1) abandon b frame firstly, then p frame and i frame lastly. (2) all the abandoned frames should distribute in the media flow as uniformly as possible. (3) adjust the rate of abandoning frames according to the network status. 3.3 management of collaborative interaction (1) member management member management in ncds refers to permission management and access control of members, permission altering, authorization and cancellation dynamically. role based access control can be used to achieve the management of permissions. roles including administrator, head designer, general member and spokesman. the administrator is responsible for assigning roles to members, designer has the operation right, general member and spokesman can send messages to other members. a member can be assigned multiple roles at the same time and there is only one head designer and one designer in a collaboration group any time. (2) speaking management in process of collaboration, each user can obtain the speaking right or give up it in preventing some speaker occupied the right too long or refused to give up speaking right. so, four types of speaking management mechanism is used[5]: 1) centralized control. an administrator is in charge of assigning speaking right to members. 230 luo: design and implementation of network based collaborative design system 2) grab initiatively. users sent application to the control machine, it then deprives the speaking right of current speaker and assigns the right to the user according to agreed strategy. 3) give up initiatively. the current speaker gives up its right initiatively. 4) speaking freely. everyone can speak and receive messages at freedom while obtain the operation right by competing freely. in the process of speaking, each participant can receive any other audio & video and only 8-way audio & video is allowed in order to ensure smooth communication. (3) token control. token is used to avoid inconsistency caused by multiple designers in operating on a same model at the same time. a head designer manages a token can be applied by any designer. during the designing process, only the token owner can operate the model and only his operation commands can be transmitted to the client of sharing the application, while members without token only can see the operation process or communicating each other by the multimedia environment provided by the ncds. 3.4 network environment for ncds for the ip is allocated by isp dynamically, it’s impossible to establish a permanent association between distributed members by ip address. moreover, there is a communication problem of how to penetrate nat. in these cases, a management server may be used to resolve them by storing user information and their current status of collaboration. in addition, all the data stored in the server transmitted to the recipient as to establish stable communication channels between recipient and server and solve the issue of nat. in order to ensure the security of enterprise information, port mapping technology is used to build a link between internal lan ip and the real ip. port mapping can also map a virtual ip of a host to a real ip or map a number of ports in a host with permanent ip address into different port of different machines in lan. 4.system implement and application example 4.1 system implement a ncds was developed based on tcp/ip protocol and can be run on windows platform, the hardware including camera, headphones, microphone and pc, the system developed on borland delphi and socket technology was used. the system including user management, main application interface, audio & video service, design service, screen broadcast and receive, text communication and port mapping. the user management is used to add, delete or change workgroup and its members, define roles and configuration network. main application provides an interface for users and audio & video service developed on vfw technology. design service uses vnc as kernel and called in main application interface. screen broadcast and receive adopts vncx.dll and port mapping realized by port mapping technology. 4.2 application example figure 5 running example of ncds advances in systems science and applications (2011), vol.11, no.3-4 231 issn 1078-6236 international institute for general systems studies, inc in the process of a collaborative design example as shown in figure 5, multiple members watching on the model and only one designer has the permission of operating on the model while other members can only watch or communicate each other by multimedia communication. in figure 5, the design service region was showing the shared screen and only one designer can operate on it at one time, the text communication region was showing the content of communication information between members, all the online users were listed in user region, there are four video images can be seen in the figure and the lower right corner is the audio & video service region. 5 conclusions collaborative design is an important technology of advanced manufacturing model while ncds is an effective tool for realizing it. from the viewpoint of developing a ncds, several key technologies were analyzed and a ncds was developed successfully in this paper. the current version is only provided an environment for collaboration design and our future work is to integrate it with pdm software. acknowledgements this paper is supported by national natural science fund project of china (contract no.50620130441), scientific and technological project of wuhan city (contact no. 200810321153) and youth science and technology chen guang project of wuhan city (contact no. 200750731289). references [1] gao shuming, he fazhi. survey of distributed and collaborative design. journal of computer-aided design & computer graphics, 16(2) (2004) 149~157. [2] tian ling, chen jizhong, zhao huishe. net-based collaborative design tools. journal of china mechanical engineering, 15(19) (2004) 1774~1777. [3] xu baomin,wanglixin. concept and implementation of improved application shared. journal of computer engineering, 29(11) (2003) 161~162. [4] ci jianwei,zhang huazhong. research on real-time video dealing methods of multimedia network. journal of computer engineering, 31(2) (2005) 193~194. [5] he fazhi. research on collaboration support technology and tool for cscw based cad system. wuhan: dissertation of wuhan university of technology,2000. advances in systems science and application (2015) vol.15 no.4 366-383 deduce user search progression with feedback session b.bazeer ahamed1 and t.ramkumar2 1faculty in computer science & engineering, sathyabama university, chennai,tamil nadu,india 2department of computer applications, avc college of engineering, mayiladuthurai,tamil nadu,india abstract the web is a medium for accessing a great variety of information stored in various locations. as data on the web grows rapidly it leads to several problems such as increased difficulty of finding relevant information. when a user submits a query to the search engine, it must be able to retrieve information according to the user’s intention. but search engine retrieves the list of pages ranked based on that similarity to the query. sometimes the results are not according to users interests, because many relevant terms may be absent from queries and words may be ambiguous. therefore, the results produced by the search engine are not satisfactory to fulfill the user query request. in order to solve this ambiguity, the proposed work is to discover the number of diverse user search goals for a query and represent each goal with some keywords automatically. keywords clustering; hits; restructuring search results; classified typical meticulousness; feedback session. 1 introduction with the fast growth of the web, a user can obtain abundant information easily by submitting a query to a search engine. many existing search engines use keyword matching as the search mechanism, which usually causes the situation that a large number of non-relevant documents containing query terms are founded out, and the user will make strenuous efforts to browse these non-relevant documents. thus, it is not simple to find out the real user goal from such short queries. data mining involves the use of sophisticated data analysis tools to discover previously unknown, valid patterns and relationships in large data sets. these tools can include statistical models, mathematical algorithms and machine learning methods (algorithms that improve their performance automatically through experience such as neural networks or decision trees). consequently, data mining consists of more than collecting and managing data, it also includes analysis and prediction. data mining can be performed on data represented in quantitative, textual or multimedia forms. data mining applications can use a variety of parameters to examine the data. they include association (patterns where one advances in systems science and application (2015) vol.15 no.4 367 event is connected to another event such as purchasing a pen and purchasing paper), sequence or path analysis (patterns where one event leads to another event such as birth of a child and purchasing dress), classification (identification of new patterns such as coincidences between duct tape purchases and plastic sheeting purchases), clustering (finding and visually documenting groups of previously unknown facts, such as geographic location and brand preferences) and forecasting (discovering patterns from which one can make reasonable predictions regarding future activities, such as prediction that people who join an athletic club may take exercise classes). information retrieval (ir) is essentially a matter of deciding which documents in a collection should be retrieved to satisfy a user’s need for information. the user’s information need is represented by a query or profile, and contains one or more search terms, plus some additional informations. hence, the retrieval decision is made by comparing the terms of the query with the index terms appearing in the document itself. the decision may be binary (retrieve/reject), or it may involve estimating the degree of relevance that the document has to the query. a stemming algorithm is a process of linguistic normalization, in which the variant forms of a word are reduced to a common form. it is important to appreciate that we use stemming with the intention of improving the performance of information retrieval systems [1]. unfortunately, the words that appear in documents and in queries often have many morphological variants. thus, pairs of terms such as “computing” and “computation” will not be recognized as equivalent without some form of natural language processing (nlp). improved hypertext-induced topic selection (hits) algorithm is a very popular and effective algorithm to rank documents based on the link information among a set of documents. the algorithm presumes that a good hub is a document that points to many others, and a good authority is a document that many documents point to. hubs and authorities exhibit a mutually reinforcing relationship: a better hub points at many good authorities, and a better authority is pointed to by many good hubs. to run the algorithm, we need to collect a base set, including a root set and its neighborhood, the inand out-links of a document in the root set [2]. because the hits algorithm ranks documents only depending on the in-degree and out-degree of links, it will cause problems in some cases. for example, a) mutually reinforcing relationships between hosts and b) topic drift. both problems can be solved or alleviated by adding weights to documents. the first problem can be solved by giving the documents from the same host much less weight, and the second problem can be alleviated by adding weights to edges based on text in the documents or their anchors. the simple modification to the hits algorithm for the first problem achieves a remarkable better precision, while further precision can be obtained by adding content analysis [3]. clustering is a 368 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... common descriptive task where one seeks to identify a finite set of categories or clusters to describe the data. clustering is the process of identification of classes, also called clusters or groups, for a set of objects whose classes are unknown. the objects are so clustered that the intra class similarities are nearly maximized and the interclass similarities are minimized based on some criteria defined on the attributes of objects. once the clusters are decided, the objects are labeled with their corresponding clusters, and common features of the objects in a cluster are summarized to form the class description. agglomerative hierarchical clustering techniques is used to cluster the pseudo documents since agglomerative hierarchical clustering is a classical clustering algorithm, originating analogically like k-means from the statistics domain. the main advantage of ahc is to create descriptions of clusters, removes redundant descriptions and attaching cluster to another one whose description is a subset of its description [4]. improved hits algorithm is used to rank cluster results relevant for a particular topic. the web results are restructured. finally, we introduce user editable browser that allow the user to perform editing operations such as deletion and emphasis while browsing the search results. the rest of the paper is organized as follows. section 2 reviews various techniques for effective inferring user search goals and generalized feedback session techniques. deduce user search progression is presented in section 3. experiments measures in section 4.section 5 concludes the paper and shows our future directions on this topic. 2 literature review 2.1 identifying user goals from web search results yao-sheng chang et al propose a novel probabilistic inference model which effectively employs syntactic features to discover a variety of confined user goals by utilizing web search results.[5] on the basis of analyzing the user goals in the viewpoint of natural language processing (nlp) process. assume that the user goal should be expressed with the form of a hidden sentence in his/her mind. in general, a typical sentence includes a subject (s), a verb (v), and an object (o). also, assume that the subject of the hidden sentence in user mind is the user himself/herself and the combined pair of the verb and object is called vo pair. on the basis of vo pairs, potential user goal can be represented. for example when users submit a query “michael jackson”, predict that the hidden sentence in the user mind is “i want to download michael jacksons music,” and the potential user goal is the vo pairs “download music” (verb + object). the object can be regarded as the noun after the verb [5]. the user senses his own sentence according to his own mindset but awareness to the mechanism is much different. advances in systems science and application (2015) vol.15 no.4 369 2.2 an overview of personalization in web search the above mentioned technique has some drawbacks as follows: more challenges to adopt vo-pair classes to certain languages and time consuming to identify vo-pair. in order to overcome these drawbacks indu chawla introduces web search personalization algorithms to improve the web search experience by using an individuals data e.g. user’s domain of interest, preferences, query history, browser history etc . using these factors they extract the results that are the most relevant to that individual. personalization can be broadly categorized in two types: context oriented and individual oriented [6]. context oriented personalization include factors like the nature of information available, the information currently being examined, the applications in use, when, and so on. individual oriented personalization uses user interests, query history, browser history, pages visited etc [7]. even though the web search personalization algorithm improve the web search experience by using an individual’s data, but has some drawbacks as follows [8]: in context oriented personalization expecting searchers to provide context information explicitly as part of their search is not ideal. many users are simply unwilling to provide this type of additional information and even asking for it can lead to frustration certainly asking the user for anything close to personal information is liable to alienate many users because of privacy concerns [9]. and in the second option where the user can choose from among the categories provided, the problem is to know about the users’ interest for displaying the categories to the user. moreover, searchers often do not have enough knowledge available to them to explicitly express such context information even if they were inclined to do so [10]; in individual oriented personalization a single user profile or model can contain a too large variety of different topics so that new queries can be incorrectly biased and also users are becoming more concerned about threats to privacy in the online environment. 2.3 query recommendation using query logs in search engines baeza et al [11] present an algorithm to recommend related queries to a query submitted to a search engine. the related queries are based in previously issued queries, and can be issued by the user to the search engine to tune or redirect the search process. the method proposed is based on a query clustering process in which groups of semantically similar queries are identified. the clustering process uses the content of historical preferences of users registered in the query log of the search engine [12]. the method not only discovers the related queries, but also ranks them according to a relevance criterion. ranking the queries according to two criteria: first one by the similarity of the queries to the input query (query submitted to the search engine).trail by the next one by the measures, how much the answers of the query have attracted the attention of users.the combination 370 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... of measures (a) and (b) defines the interest of a recommended query. 2.4 automatic identification of user goals in web search uichin lee et al [13] study whether and how to identify the user goal automatically without any explicit feedback from the user. two types of features for the goal identification task,one is past user-click behavior and another one is anchor-link distribution. first feature is based on the intuition that the user’s goal for a given query may be learned from how users in the past have interacted with the returned results for this query. if the goal of a query is navigational, then in the past users should have mostly clicked on a single website corresponding to the one they have in mind. on the other hand, if the goal is informational, in the past users should have clicked on many results related to the query. thus by observing how the results for a particular query have been clicked so far, and easily tell whether the current user who issues that query has a navigational or an informational goal. the anchor link distribution is used to find the destination of the links with the same anchor text as the query. for example, for a navigational query pub med, a single authoritative website exists (which is www.ncbi.nlm.nih.gov)[9]. as a result, if extract all the html links with the anchor text pub med, to find that a dominating portion of these links point to that single website. on the other hand, for an informational query hidden markov model, because of lack of a single authoritative site, can expect that the links with the anchor text hidden markov model point to a number of different destinations. 2.5 learn from web search logs to organize search results xuanhui wang et al propose a different strategy for partitioning search results, which addresses these two deficiencies through imposing a user-oriented partitioning of the search results.[10] learn “interesting aspects” of similar topics from search logs and organize search results based on these “interesting aspects”. finally generate more meaningful cluster labels using past query words entered by users. to know what the users are really interested in given this query, first retrieve its past similar queries in preprocessed history data collection. based on the similarity scores, we rank all the documents in history data set. the top ranked documents provide us a working set to learn the aspects that users are usually interested in. each document in history data set corresponds to a past query, and thus the top ranked documents correspond to q’s related past queries. 2.6 query-sets: using implicit feedback and query patterns to organize web documents barbara poblete et al present a new document representation model based on implicit user feedback obtained from search engine queries.[14] the main objective advances in systems science and application (2015) vol.15 no.4 371 of this model is to achieve better results in non-supervised tasks, such as clustering and labeling, through the incorporation of usage data obtained from search engine queries. this type of model allows user to discover the motivations of users when visiting a certain document. the query document model reduces the feature space dimensions considerably, because the number of terms in the query vocabulary is smaller than that of the entire website collection. this model is very similar to the vector model, with the only difference that instead of using a weighted set of keywords as vector features, we will use a weighted set of query terms. fig. 1 example of the query document representation. 2.7 user intent based searching haibo yu et al propose a user intent based searching mechanism (uibs )in order to enable precise discovery of a web site (for navigational searching) as well as aggregating of their information (for informational searching) and performing further activities (for transactional searching).[15] this mechanism includes three main components: a web site capability description interface for the explicit description of a web site’s capability, a search interface that enables explicitly describing user’s search intent and requirements and their relevant responses, and a query engine which provides relevant search results based on user’s intent and requirements. in a uibs enabled user, a uibs client should be installed which can issue uibs search requests based on user inputs or other sources of user intents, preferences and query [9]. the client should also process the received uibs results, and presents the processed result to the user or other applications. it explicitly describes the general information of the web site and the links to published content as well as web services that the web site can provide. the 372 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... wscd works like a site map and presents all information that the web site wants to publish. 2.8 generalized feedback session techniques zheng lu et al extended the previously discussed algorithm in an effective manner. he focused on user search goals by organizing search results by aspect learned from user click through logs.[16, 17] in the feedback sessions related to the given query will be extracted from user clickthrough logs. the feedback session is defined as the series of both clicked and unclicked urls and ends with the last url that was clicked in a session from user click-through logs. the clicked urls tell what users require and the unclicked urls reflect what users do not care about [18]. then, map each feedback sessions to pseudo-documents using keywords which can efficiently reflect user information needs. by using kmeans clustering techniques to cluster the pseudo-documents to infer user search goals [19]. finally restructure the web search results inferring user search goals. fig.2 represents the data flow diagram for the existing system. but there is some confines that we are focused in obtainable system i.e. generate noisy and redundant search results problem and cluster labels generated are not informative enough to allow the user to identify the right search result. do not generate the search results if the query has been entered for the first time. fig. 2 dataflow diagram for the generalized feedback session techniques. 3 deduce user search progession with the fast growth of the web, a user can obtain abundant information easily by submitting a query to a search engine. many existing search engines use keyword matching as the search mechanism, which usually causes the situation that advances in systems science and application (2015) vol.15 no.4 373 a large number of non-relevant documents containing query terms are founded out, and the user will make strenuous efforts to browse these non-relevant documents. thus, it is not simple to find out the real user goal from such short queries. a variety of ranking algorithms have been proposed and used in many web search engines. people search for information using search engines. search engines, however, cannot always return good ranked search results which satisfy user’s search intentions adequately. hence, it is difficult to recognize users’ search intentions just by analyzing their input queries. the search results sometimes do not correspond to the user’s search intentions because of this diversity. in this case, the user must check the search results sequentially until he obtains sufficient information from the linked pages from the page containing the search results. the query contains parts of user general verbal communication and special characters which are not required for analysis as they do not truly reflect the relevance of a search result. if this query is used for analysis, it may give inconsistent and inaccurate results. therefore the user query will be pre-processed to identify the root words. the feedback sessions has been introduced to infer user search goals for a query. then, map each feedback sessions to pseudo-documents using keywords which can efficiently reflect user information needs and relate to the given query will be extracted from user clickthrough logs. the feedback session is defined as the series of both clicked and unclicked urls and ends with the last url that was clicked in a session from user click-through logs. the clicked urls tell what users require and the unclicked urls reflect what users do not care about. in order to apply the evaluation method to large-scale data, the single sessions in user click-through logs are used to minimize manual work. because from user click-through logs, we can get implicit relevance feedbacks, namely “clicked” means relevant and “unclicked” means irrelevant. a possible evaluation criterion is the typical meticulousness (tm) which evaluates according to user implicit feedbacks. tm is the average of precisions computed at the point of each relevant document in the ranked sequence however, tm is not suitable for evaluating the restructured or clustered searching results. therefore we introduce classified typical meticulousness (ctm) system has been initiated to evaluate the performance of the restructured web search results. where the ctm of the class including more clicks namely confer is calculated, ctm selects the tm of the class that user is interested in (i.e., with the most clicks/confer). then, map each feedback sessions to pseudo-documents using keywords which can efficiently reflect user information needs. k-mean clustering is too simple and it performs poorly for large set of data. and also it requires prior knowledge of number of clusters to be generated. therefore, instead of using k-mean clustering we had used agglomerative hierarchical clustering to cluster the pseudo documents. 374 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... 4 experiments and measures the query submitted by the user contains parts of speech and special characters which are not required for analysis as they do not truly reflect the relevance of a search result. if this query is used for analysis, it may give inconsistent and inaccurate results. therefore, the user query will be pre-processed to identify the root words. fig. 3 data preprocessing fig. 4 general feedback session progression. feedback session consists of both clicked and unclicked urls and ends with the last url that was clicked in a single session. the clicked urls tell what users require and the unclicked urls reflect what users do not care about. table advances in systems science and application (2015) vol.15 no.4 375 1 shows a feedback session with lists of 10 search results of the given query “computer”, where “0” in the click sequence represents the “unclicked url” and remaining represents “clicked url”. the pseudo-document can be used to infer user search goals. the mapping of feedback sessions into a pseudo-document includes two steps. they are as follows: first enrich the urls with additional textual contents by extracting the titles and snippets of the returned urls appearing in the feedback session. in this way, each url in a feedback session is represented by a small text paragraph that consists of its title and snippet. then, some textual processes are implemented to those text paragraphs, such as transforming all the letters to lowercases, stemming and removing stop words. fig. 5 map feedback sessions to pseudo-documents. improved hits (hyperlink induced topic search) algorithm is applied to the clusters generated by agglomerative hierarchical clustering to rank cluster results relevant for a particular topic. ranking is done by assigning relevance weight to the cluster results. restructure the web search results based on result generated by improved hits algorithm. each webpage will be treated as a node, with hyperlinks treated as directed links from one node to another. each node i is assigned an authority score a(i) and hub score h(i). given a directed graph, the authority and hub score is defined as follows: a(i) = the sum of the hub scores of the nodes pointing to node i. h(i) = the sum of the authority scores of the nodes that node i is pointing to. the higher a node’s authority/hub score is, the better authority/hub is. this makes intuitive sense; a node is a good authority if good hubs are pointing to it, and a node is a good hub if it is pointing to good authorities. these authority and hub updates can be described through matrix notation. suppose that we define matrices a, u and v as follows:a = the adjacency matrix of the graph. that is a(ij) = 1 if node i has a directed link to node j, and 0 otherwise. u = a column matrix containing the hub score of all of the nodes. 376 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... v = a column matrix containing the authority score of all of the nodes. with these definitions, it follows that for any iteration k: u(k) = a ∗ v (k)andv (k) = transpose(a) ∗ u(k) fig. 6 hubs are pages that link to authorities. search engines always return millions of search results, it is necessary to organize them to make it easier for users to find out what they want. restructuring web search results is an application of inferring user search goals. the inferred user search goals are represented by the vectors and the feature representation of each url in the search results can be computed. then, categorize each url into a cluster centered by the inferred search goals. perform categorization by choosing the smallest distance between the url vector and user-search-goal vectors. by this way, the search results can be restructured according to the inferred user search goals. the adjacency matrix of the graph a =  0 0 1 0 0 1 0 0 0 at =  0 0 0 0 0 0 1 1 0  assume the initial hub weight vector u =  1 1 0  advances in systems science and application (2015) vol.15 no.4 377 compute the authority weight vector by: v = at · u =  0 0 0 0 0 0 1 1 0  •  1 1 1  =  0 0 2  then, the updated hub weight v = a · u =  0 0 1 0 0 1 0 0 0  •  0 0 2  =  2 2 0  editable browser for selecting the results & re ranking the result based on user editing strategies. we enhanced an editing operation such as deletion and emphasis that can be employed while users are browsing web search results. this system enables users to edit any portion of the page of web search results at any time while searching. our system detects user’s search intentions from the editing operation. then our work propagates user’s search intentions based on their editing operation to all of the web search results. this system guesses the user’s search intentions, for example assuming, “this user does not want this kind of the result”, if the user deletes a part of the search results, or “this user wants this kind of the results more”, if the user emphasizes a part of the search results. after guessing the user’s search intentions, our system re-ranks the search results according to the intention and shows the re-ranked results to the user. in this way, the user can easily obtain optimized search results. fig. 7 editing browser processing. 378 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... fig. 8 restructured outputs with editable options. fig. 9 restructured result after deletion operation. advances in systems science and application (2015) vol.15 no.4 379 fig. 10 feedback session for the query “the sun”. user editable browser has been introduced to perform editing operations such as deletion and emphasis while browsing the search results. when the user deletes a part of the search result, the system degrades search results which include the deleted term or sentence. when the user emphasizes a part of the search result, the system upgrades the search results which include the emphasized term or sentence. the system re-ranks the search results according to the user intention and shows the re-ranked results to the user. deletion operation: deletion is an operation that indicates what types of search results the user does not want to obtain from the system. consider the sample class shown in the fig.7 , if the user does not want the 36rd link the user uses the deletion operation to remove the link from the generated search result. performance evaluation based on restructured web search results each url in the click session is categorized into one class, we introduce the typical meticulousness(tm) process for calculating the user click through log. tm will always be the highest value namely 1 no matter whether users have so many search goals or not. therefore, there should be a risk to avoid classifying search results into too many classes by error. so, we can further extend tm by introducing the above risk and propose a new criterion called “classified tm” it is calculated by ctm = ( 1 y r=1∑ y rr r )× (1− dij/c 2 y ) x where y is the number of relevant (or clicked) documents in the retrieved ones, r is the rank, rr is the number of relevant retrieved documents of rank r and x is used to adjust the influence of risk on ctm which can be learned from training data. we select 10 queries and empirically decide the number of user search goals of these queries. then, we cluster the feedback sessions and restructure the search results with inferred user search goals. we tune the parameter x to make 380 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... ctm the height. based on the above process, the optimal x is from 0.6 to 0.8 for the 10 queries the mean and the variances of the optimal x are 0.697 and 0.005, respectively. thus, we set x to be 0.7. table 1 user editable browser comparisons with other methods for 100 ambiguous queries method probabilistic inference model personalizati on algorithms query recommender algorithm identification of user goals generalized feedback session user editable browser i 0.7124 0.787 0.755 0.562 0.7173 0.8055 ii 0.7911 0.625 0.584 0.742 0.8031 0.8875 iii 0.801 0.749 0.611 0.632 0.9895 1 iv 0.5391 0.6245 0.512 0.654 0.6646 0.6787 v 0.632 0.745 0.741 0.524 0.7845 0.833 sample manual calculation for single class result: 1.ctm = 1/3(1/1 + 2/3 + 3/4) ∗ (1− 0)0.7 = 0.8055 2.ctm = 1/4(1/1 + 2/2 + 3/4 + 4/5) ∗ (1− 0)0.7 = 0.8875 3.ctm = 1/2(1/1 + 2/2) ∗ (1− 0)0.7 = 1 4.ctm = 1/5(1/1 + 2/3 + 3/5 + 4/7 + 5/9) ∗ (1− 0)0.7 = 0.6787 5.ctm = 1/2(1/1 + 2/3) ∗ (1− 0)0.7 = 0.833 in order to demonstrate that when inferring user search goals, clustering our proposed feedback sessions are more efficient than other clustering search results and clicked urls directly in order to further compare our method with the existing method, we test the 100 most ambiguous queries such as “apple”, “the sun”, “car”, “mobile”, “college”, “earth” and so on. ctm has the highest mean average which is significantly higher than existing method. fig. 11 ctm vs. mean average of other methods. 5 conclusion this paper aims to discover the number of diverse user search goals for a given query and predict with some keywords automatically. first, the given user query is pre-processed to find the root word. then, a feedback session has been introadvances in systems science and application (2015) vol.15 no.4 381 duced to infer user search goals for a query. the feedback session is defined as the series of both clicked and unclicked urls and ends with the last url that was clicked in a session from user click-through logs. each feedback sessions are mapped to pseudo-documents using keywords which can efficiently reflect user information needs. by using agglomerative hierarchical clustering techniques to cluster the pseudo documents. improved hits (hyperlink induced topic search) algorithm is used to rank cluster results relevant for a particular topic. we introduced classified typical meticulousness(tm) process for calculating the user click through log, the web results are restructured. finally, user editable browser has been introduced allow the user to perform editing operations such as deletion and emphasis while browsing the search results. this paper can be extended to support query recommendation there by suggesting queries that can helps the user to form queries more precisely. acknowledgement the authors would like to express their sincere thanks to the editorial board members and anonymous reviewers for providing their valuable comments. the first author would like to express their thanks to management and other authorities for providing an feasible environment to carry out the research work successfully. references [1] beeferman. d and berger. a (2000), “agglomerative clustering of a search engine query log”, acm sigkdd proc. sixth int’l conf. knowledge discovery and data mining(sigkdd ’00), pp.407-416. [2] joachims .t text mining, j. franke, g. nakhaeizadeh and renz, eds.(2003), “evaluating retrieval performance using click-through data”, physica/springer verlag, pp.79-96. [3] beitzel. s, jensen .e., chowdhury. a, and frieder. o (2007), “varying approaches to topical web query classification,” proc. 30th ann.int’l acm sigir conf. research and development (sigir ’07) , pp.783-784. [4] joachims .t,(2002), “optimizing search engines using click through data,”,proc. eighth acm sigkdd int’l conf. knowledge discovery and data mining (sigkdd ’02), pp.133-142. [5] yao-sheng chang, kuan-yu he, scott yu and wen-hsiang lu (2006), “identifying user goals from web search results”,in ieee. 382 b. bazeer ahamed and t.ramkumar:deduce user search progression with ... [6] indu chawla,(2010), “an overview of personalization in web search”, in ieee. [7] cao.h, jiang.d, pei.j, he.q, liao.z, chen.e and li.h (2008), “contextaware query suggestion by mining click-through”, proc.14th acm sigkdd int’l conf. knowledge discovery and data mining (sigkdd’08), pp.875-883. [8] chen .hand dumais. s, (2000), “bringing order to the web: automatically categorizing search results”, proc. sigchi conf. human factors in computing systems (sigchi ’00), pp.145-152. [9] zeng .h.-j, he. q.-c, chen .z, ma. w.-y and ma. j.(2004), “learning to cluster web search results”,proc. 27th ann. int’l acm sigir conf. research and development in information retrieval (sigir ’04), pp.210-217. [10] shital c. patil, prof. r.r. keole, (2013), “web usage mining and webcontent mining c a combine approach for enhancing search result delivery”, vol.3, no.10. [11] baeza-yates .r, hurtado. c and mendoza. m, (2004), “query recommendation using query logs in search engines”, the journal of supercomputing(edbt ’04), pp.588-596. [12] joachims .t,“optimizing search engines using click through data”. [13] lee.u, liu.z and cho. j, (2005),“automatic identification of user goals in web search”,proc. 14th int’l conf. world wide web (www ’05), pp.391400. [14] poblete.b and ricardo .b.-y (2008), “query-sets: using implicit feedback and query patterns to organize web documents,”,proc. 17th int’l conf.world wide web (www ’08), pp.41-50. [15] haibo yu , tsunenori mine and makoto amamiya(2011), “towards user intent based searching”, 2011 international joint conference of ieee trustcom-11/ieee icess-11/fcst-11. [16] wang .x and zhai. c.-x (2007), “learn from web search logs to organize search results,”,proc. 30th ann. int’l acm sigir conf. research and development in information retrieval (sigir ’07),pp.87-94. [17] zheng lu, hongyuan zha, xiaokang yang, weiyao lin and zhaohui zhengr, (2013), “a new algorithm for inferring user search goals with feedback sessions,”, ieee transactions on knowledge and data engineering,vol.25, no.3. advances in systems science and application (2015) vol.15 no.4 383 [18] takehiro yamamoto, satoshi nakmura and kutsumi tanaka, “an editable browser for reranking web search results”. [19] huang c.-k, chien l.-f, and oyang y.-j, (2003), “relevant term suggestion in interactive web search based on contextual information in query session logs,”,j. am. soc. or information science and technology, vol.54, no.7, pp. 638-649. corresponding author b.bazeer ahamed can be contacted at:bazeerahamed@gmail.com advances in systems science and application (2015) vol.15 no.4 392-399 the market laws declaration of economic development of the countries worldwide s.baizakov1, a. oinarov2, j. yi-lin forrest3 and n.baizakov4 1jsc, economic research institute, astana, kazakhstan; 2jsc, kazakh center of public and private partnership, astana, kazakhstan; 3pennsylvania state system of higher education, astana, kazakhstan; 4jsc stana-project, astana, kazakhstan abstract the mathematical formulation of a problem of market balance of levels of production, employment, the income and the prices is considered. the new market model of work which focuses economy on the end result is constructed. it becomes clear why development of market forces of work and the capital in economy branches, contrary to desires of businessmen on accumulation of competitiveness of the enterprises, happens according to laws of development of market economy. keywords equlibrium; balanced growth; table input-output; model; laws 1 introduction the analytical work to assess the costs and benefits of the regulatory impact on the development of the national economy was successfully carried out in europe and in other developed countries. appropriate tools in russian-speaking countries are known under the following names: regulatory impact analysis (ria) and regulatory impact assessment (ria). the ability to analyze the influence of the regulatory impacts on the economy, and to use them in the evaluation of scientific and technological changes in the economy, is very useful for developing countries. if we study the regulatory impact analysis tools on the development of the national economy, it is possible to distinguish two types of them. the first type of regulatory impact influences on macroeconomic indicators,and the second type of regulatory impact influences on the performance of enterprises at the microeconomic level. the first one mainly influences on the cost of capital, in the form of money, while the second one has an impact on the cost of capital, in the form of the product. it should be noted that changes in the cost of capital in the form of money, and the cost of capital in the form of the product are measured with units of national or global currency banknotes. the president of the republic of kazakhstan nursultannazarbayev in his article “the fifth way”, which has been written during the crisis of 2007-2009, marked the presence of the third type of regulatory impact influencing on the level of development of the national economy[1]. the “fifth way” demonstrates the necessity of an objective assessment of the social and economic subsequences advances in systems science and application (2015) vol.15 no.4 393 of those three types of regulatory impacts. the conclusion that appears from the logic of the “fifth way” is to ensure the mutual consistency of the indicators of all types of regulatory impact that have a crucial importance in determining the true value of money capital in macroeconomics and the true cost of the commodity-capital in microeconomics. this principle of mutual consistency of regulatory impacts defined in the “fifth way” has become the foundation of the concept of building cross-sectoral assessment models of levels of production market balance, employment, incomes and prices developed by the initiative group of kazakhstan in early 2010. thus, the initiative group of kazakhstan in 2010, sparing no effort and energy performed the implementation of the ideas proposed by the president of kazakhstan on the assessment of the influence of economic governance policies not only on the true cost of capital in the form of money, and on the true cost of capital in the form of the product. this initiative group reviewed in detail the basic principles of market economy management models associated with the name of keynes, monetarists and mundell-fleming[2]. as we know, those models are now extensively used in the analysis of the market economy development. firstly, it appeared that the main indicator of market forces analysis of capital in the form of money in all those three models is a nominal gdp, and the main indicator of market forces analysis of capital in the form of goods – a real gdp in the keynesian and monetarist models, gdp, and according to purchasing – power parity in the mundell-fleming model. secondly, the kazakhstani initiative group has proved that the principles of building models of the monetarists and mundell-fleming have a single macroeconomic root which was defined by a known keynesian theory of the equation of demand and supply for the goods and services. since the principles of all those models of economy analysis and management have been developed with the use of the system of the same macroeconomics indicators which had been created by keynes. thirdly, it was discovered that in the system of equations of market balance, created by keynes and his followers, there is no microeconomics indicator: all models of the economy analysis are based on macroeconomic theory and are not related to microeconomic theory and microeconomic indicators. fourthly, it became known that keynes is the author of separation of a unified indicators system of the national economy, which was defined in the classical economic science, in the part of macroand microeconomics. fifthly, it is determined that all of those models are based on the one assumption between the capital in the form of money and the capital in the form of the product. keynes has based his principle of construction of his model on the macroeconomic stability of the prices for goods and services, and the monetarists 394 s.baizakov, a. oinarov, j. yi-lin forrest and n.baizakov: the market laws declaration... – on macroeconomic stability of the money turnover velocity, but fleming and mundell – on the macroeconomic stability of the national currency exchange rate. consequently, representatives of the initiative group during the formation of the kazakh model of economic management had a reason to accept the hypothesis of a possible omission of keynes made in the course of a single economic system breakdown on macroeconomics and microeconomics, and to mark the uncritical attitude of monetarists as well as of mundell and fleming towards this omission. 2 the market laws declaration of economic development of the countries worldwide there are no words that all of those models were popular management tools of market economies of that time, this can be proved by the sustainable development of the world economy until 2007. however, keynes in due time determining the indicators of macroeconomics used the method to minus the material costs from sales of goods and services. and he created a new macroeconomic indicator called the gross domestic product. this arithmetic operation allowed the authors of macroeconomic theory to get rid from microeconomics indicators. as a result, in the study of the cost of the financial capital are mainly engaged macroeconomists, and with cost of commodity capital microeconomists. as it became known much later, after keynes, the determination of the solve equations of mutual dependence of the final result of the production from the gross sales of goods and services does not require arithmetic but deeper calculations using matrix algebra. the searches of the kazakh initiative group showed that the foundation of this dependence is determined by the system of equations describing the dependence between macroeconomic and microeconomic indicators as follows[3]: by economic sectors– t(i)x(i)− t (i)y (i) = ±θ(i), i = 1, 2, ..., n by national economy– c ≡ t t = y x , where t(i)= l(i) x(i)– time spent by the people themselves, in relation to one tenge, received from the sale of goods and services by economic activity, expressed in hours, days, months, years; t (i) = t(i)b– the total time spent by people on the job, per one tenge, received from the sale of goods and services by economic activity, expressed in hours, days, months, years; tx = l– national fund of working time, calculated by gross domestic product (x) is determined by the balance of the time model of inter-sectoral working advances in systems science and application (2015) vol.15 no.4 395 people, expressed in man-hours per year, man-days a year, man-year, people a year; ty = l– national fund of working time, calculated by the volume of the end product (y) is determined by the balance of the time model of inter-sectoral working people, expressed in man-hours per year, man-days a year, man-year, people a year; tiyi − tixi = θij– the difference between the cost of working time, defined in terms of the final product (y) and the working time defined in terms of gross product (x), expressed in man-hours per year, man-days a year, man-year, people a year; b = (e−a)(−1)– inverse matrix of the full costs of the kazakhstan economy, as determined by inter-branch model, where a the technological matrix of the economy of kazakhstan by types of economic activity, and e special algebraic matrix units. the final result of the analysis of the regulatory impact on the development of the national economy of kazakhstan is the difference between the growth rate of capital in the form of money in macroeconomics and of capital in the form of goods in microeconomics: ċ c = ẏ y = ẋ x , this difference arises from the opposite direction of capital flows in the form of money and of capital in the form of the product which have different flow rate. it can be called the coefficient of scientific and technological change, as it has appeared, as a manna from heaven, from the ceiling, with strict equality of the working time spent on the creation of the end product (y = c + g + i + nx) in macroeconomics and gross domestic product (x) in microeconomics. so that, this difference is a net contribution of scientific and technological progress determined by the difference between productivity of the end product and productivity of gdp.in other words this success comes from within the country’s economic system determined by its interdisciplinary development model. now we can explain any innovative initiative which is linked to scientific and technological changes in the national economy not only by an abstract definition of “scale” production or solow residue, but by the cost spent on the production of specific working hours, the specific productivity of labor and capital, the ratio of scientific and technological progress. in addition, every entrepreneur can perform the same calculation for each type of their economic activity. now, the true value of the national currency can be estimated by dividing the rate of scientific and technological changes on the gdp deflator, and the index of inflation will be determined by a formula which is different from the formula of the gdp deflator. the system of laws of the market economy determined due to the coefficient of scientific and technological change, includes the following 396 s.baizakov, a. oinarov, j. yi-lin forrest and n.baizakov: the market laws declaration... related equations: • equation of the law determining the overall impact of stimulating scientific and technological improvements – c(t): c(t) = y (t)/x(t) • equation of the law determining the purchasing power of money – pp: pp(t) = (c(t) ∗ i2(t))/i1(t); • equation of the law determining the prices of goods and services – pc(t)=1/pp(t): pc(t) = 1/pp(t) = i1(t)/(c(t) ∗ i2(t)); • basic law equation determining the real volume of the end product, measured by the real cost of capital in the form of money, as an indicator of real economic growth – i3(t): i3(t) = pp(t) ∗ i1(t); • basic law equation determining the real volume of the end product, measured by the real cost of capital in the form of goods as the benchmark of real economic growth – i3(t): i3(t) = c(t) ∗ i2(t); • equation of the law of general price deflation – b(t): b(t) = c(t)/pp(t) ≡ i1(t)/i2(t); • equation of the law determining the net benefits of stimulating scientific and technological improvements – dc(t)%: dc(t)% = ċ c = ẏ y = ẋ x . it should be noted that the index i1(t), is the rate of growth of gross domestic product for the price of the current year, while the index i2(t) the price of the previous year. progress of the us economy and china for 2002-2011 according to the inputoutput balance of these countries. as the table shows, in the period from 2002-2008, both countries were characterized by trends to reduce the cost of capital in the form of money (y), which falls per unit of capital in the form of goods (x) .this means that both countries, return on equity, in the form of money decreased, rates of stp were declining trend. this situation lasted until 2009 flesh. in spite of the substantial difference in the level of indicators of scientific and technological progress, the two countries after the crisis of 2007-2008 has a regulating effect on the level of the coefficient advances in systems science and application (2015) vol.15 no.4 397 of ntp so that stop further depreciation of the national currency. but they have different approaches. for example, in the united states since 2009 have been lowered growth rate of the cost of capital in the form of money and the cost of capital in the form of the product, so as to achieve better rate of scientific and technological progress. in china, by contrast, produced faster growth rate of growth capital in the form of money, compared to the growth rate of capital in the form of the product. table 1 comparative assessment of indicators of scientific and technological progress of the us economy and china for 2002-2011 according to their inputoutput balance (mm in usd). years usa china end product (y) gross product (x) ntp coefficient=y/x end product (y) gross product(x) ntp coefficient =y/x 2002 10416078.53 18873334.93 55.2 1324753 3794147 34.9 2003 10938122.88 19828498.07 55.2 1485275 4457621 33.3 2004 11736535.23 21264711.1 55.2 1760018 5372764 32.8 2005 12549008.41 23072266.04 54.4 2001206 6527490 30.7 2006 13280785.08 24479922.11 54.3 2354040 8160175 28.8 2007 13844412.73 25795266.08 53.7 3019764 10740915 28.1 2008 14214545.69 26565031.96 53.5 3951859 13913136 28.4 2009 13775598.08 24802899.21 55.5 4538041 15149965 30 2010 14230107.43 25810105.94 55.1 5401801 18070490 29.9 2011 14770666.86 26918120.33 54.9 6777332 22271025 30.4 source: developed by the author based on the table “input-output” of these countries. 3 discussion and conclusion now when we already know the trends of scientific and technological progress, which are determined by the correlation between the two main indicators of the market economy, namely, between the cost of the end product, which expresses the cost of capital in the form of money and the value of gross domestic product, which expresses the cost of capital in the form of the product, you can do strong statement: the law of scientific and technological progress and related laws of the market economy are a reflection of objective processes occurring in it, and they are independent of the will of the people. no one in the world can not change or cancel them. therefore, people can only identify, understand and learn these laws. do not consider their actions in the practice of economic management equivalent to curb the development of a market economy, its inhibition of further liberalization. this is to ensure that the laws of market economy, equivalent to the laws of natural science and nature. i can even say so that a natural disaster happens in one place, in one country, and the collapse in the developed economies as a 398 s.baizakov, a. oinarov, j. yi-lin forrest and n.baizakov: the market laws declaration... contagious disease spreads rapidly on the economies of developing countries. therefore, science and technology indicator changes justified by the kazakhstani initiative group, as the law of the market economy, is suitable for use in the analysis of the economies of all countries of the world. and it will be best if we call these economic laws system tools capable for economic analysis and regulatory impact assessment models such as the keynesian model, the monetarists and the model of mundell and fleming. this is due to the fact that any of these models is reduced to the particular case of the following generalized equation, kazakhstan proposed an initiative group to assess the actual final product anywhere in the world, including even such large countries like the us and china: pp ∗ngdp = c ∗rgdp (a) this indicator measuring the end product is significantly different from the nominal gross domestic product (gdp), real gdp and the gdp determined by purchasing-power parity, which are included in the current model of keynesianism deystvuyushie, monetarists, mundell and fleming. model (a) determines the power of the final product, used in the country and consists not only of the domestic product of the country, but also its product to come from external trade turnover. the new indicator measuring economic growth is the development of money equilibrium assignment (pp * ngdp) capital level of trade (with * rgdp) capital. it is the final product, and not the nominal (ngdp) and real gdp (rgdp), let alone the gdp determined by purchasing power parity, is a tool for sustainable development of the national economy and the real measure of wellbeing of its people. the measure defined by the equation of market equilibrium (a), is common for the economy and the economy of consumption of goods and services, as well as for the economy of the monetary and financial system. it is a new measure, adapted to the level of development of the productive forces of labor and capital in a globalized world economy. it can be used to assess the regulatory impact on the economic development of any country, any region of the world, since the value is the true purchasing power of the national currency, and the value is a measure of realindicator of scientific and technological change. this measure determining indicators of scientific and technological changes in the past never been used and is a new phenomenon in the development of world economics. and because the core of this declaration on market laws of economic development of the countries of the world is the progress in economic science, especially in the science of economic management, which are hidden multibillioneffects from the use of reliable tools of the regulatory impact. advances in systems science and application (2015) vol.15 no.4 399 references [1] the fifth way, (news), september 22, 2009. [2] y. hasanli, s. bayzakov, v. valiyev, g. sarsembaeva. (2011), “modeling of the multiplicative effects of opening of the work places on the bases of ‘intersectoral labor balance’ (on example of azerbaijan and kazakhstan)”, ecomodazkz full pape eng. [3] ba�zakov s. (2014), “model of the labour market, based on the final product”, science and innovation, no.12, pp.31-34. corresponding author sailau baizakov can be contacted at: baizakov37@mail.ru delay-dependent stability criteria of stochastic uncertain hopfield neural networks with unbounded distributed delays and impulses r. raja1, r.sakthivel 2 and s.marshal anthoni 3 1department of mathematics, periyar university, salem-636 011, india 2department of mathematics, sungkyunkwan university, suwon 440-746, south korea 3department of mathematics, anna university of technology, coimbatore-641 047, india email: antony.raja67@yahoo.com abstract this paper is concerned with the stability analysis problem for a class of delayed stochastic uncertain hopfield neural networks with unbounded distributed delays and impulses. a new lyapunov-krasovskii functional is constructed for the addressed system and several freeweighting matrices combined with the s-procedure are employed to derive the delay-dependent stability criterion. the criterion is derived and formulated in terms of linear matrix inequality (lmi). in addition to that, two illustrated examples with simulation results are given to show the effectiveness of the obtained theoretical results. keywords delay-dependent stochastic hopfield neural networks distributed delays lyapunov krasovskii functional linear matrix inequality impulses. 1. introduction during the past several years, the stability of a unique equilibrium point of hopfield neural networks [2] with delays have received especially considerable attention due to their extensive applications in solving optimization problem, traveling salesman problem and many other subjects in recent years [4, 5, 6, 14, 15, 16, 17, 18, 23, 24]. basically, the stability results of delayed hopfield neural networks can be classified into two categories: delay dependent stability and delay independent stability. delay-dependent stability results are generally less conservative than delay-independent stability when the delays are small. on the other hand, time-delays occurring in the interaction between neurons will affect the stability of a network by creating instability, oscillation and chaos phenomena. recently, a number of global stability criteria of hopfield networks with time-delays have been proposed (see [13, 15, 16, 17, 18]). the dynamical systems are often classified into two categories of either continuous-time or discrete-time systems. apart from this two systems, yet there is a somewhat new category of dynamical systems, which is neither continuous-time nor purely discrete-time, these are called dynamical systems with impulses. a basic theory of impulsive differential equations has been developed in [11]. the stability conditions in [19, 20, 22] were established by using the impulsive condition. in the real world, there are two common disturbances that affects the issn 1078-6236 international institute for general systems studies, inc. advances in systems science and applications (2011), vol. 11, no. 1-2 93-109 network process, one is stochastic perturbations and the other one is uncertain parameters. recently, there are some research papers about stochastic neural networks, has been investigated, see for example [3, 5, 6, 7, 9, 10, 16, 17, 18]. in [12] the stability problem for both discrete and distributed delays were discussed. yang [21] has been investigated the stability of neural networks with distributed delays. in practical, uncertainties often exist in most engineering and communication systems and may cause undesirable dynamic network behaviors such as oscillation, instability and chaos. more specifically, the connection weights of the neurons are inherent dependent on certain resistance and capacitance values that inevitably bring in uncertainties during the parameter identification process. in the literature, uncertainties can possibly be described by norm bounded, polytopic or linear fractional uncertainties characterizations and have been widely employed in the field of robust control and performance analysis. based on the above descriptions, this paper aims to develop the problem of asymptotic stability for delayed stochastic uncertain hopfield neural networks with unbounded distributed delays and impulses. by constructing an appropriate lyapunov-krasovskii functional, employing several free-weighting matrices and s-procedure, we obtain a delay-dependent stability criterion in terms of lmis. finally, two numerical examples with simulation results are provided to demonstrate the usefulness of the main results in this paper. 2. network model and preliminaries the delayed stochastic hopfield neural network model with unbounded distributed delays and impulses is defined by the following state equations: dxi(t) = [ − aixi(t) + n∑ j=1 bijfj(xj(t− τ)) + n∑ j=1 cij ∫ t −∞ kj(t− s)fj(xj(s))ds+ ji ] dt + n∑ j=1 σij(t, xj(t))dwj(t), t 6= tk (1) xi(tk) = ikx(t−k ), t = tk, k = 1, 2, ... where xi(t) is the state of the ith neuron at time t; ai > 0 denotes the passive decay rate; bij and cij are the synaptic connection strengths; fj denotes the neuron activation functions; ji is the constant input from outside the system; τ represents the continuous delay and the delay kernel kj is a real valued continuous function defined on [0,+∞] and satisfies, for each i, ∫∞ 0 kj(s)ds = 1. the stochastic disturbance w(t) = (w1(t), w2(t), ..., wm(t))t is an mdimensional brownian motion; σij(·, ·) is locally lipschitz continuous and satisfies the linear growth condition as well; xi(tk) = ikx(t−k ) is the impulse at moment tk, the fixed moment of time tk satisfy t1 < t2 <, ..., limk→+∞ tk = +∞ and x(t−) = lims→t− x(s); ik is a constant real matrix at the moments of time tk. let pc([−τ, 0],rn) denotes the set of piecewise right continuous functions φ : [−τ, 0]→ rn with the sup-norm |φ| = sup−τ≤s≤0‖φ(s)‖. for given t0, and φ ∈ (pc[−τ, 0],rn), the initial condition of system (1) is described as x(t0 + t) = φ(t), for t ∈ [−τ, 0], φ ∈ 94 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . pc([−τ, 0],rn). the following assumptions is utilized throughout this paper: (c1) the activation function f(x) is boundless and satisfies 0 ≤ fi(ξ1)− fi(ξ2) ξ1 − ξ2 ≤ li, for any ξ1, ξ2 ∈ r, ξ1 6= ξ2, i = 1, 2, ..., n (c2) the function fi(xi(·)) = 0, i = 1, 2, ..., n satisfy 0 ≤ fi(xi(t)) xi(t) ≤ li, fi(0) = 0, ∀xi(t) 6= 0, i = 1, 2, ..., n where li, i = 1, 2, ..., n are positive constants. remark 2.1 the above conditions ensures that the nonlinear resulting neuron activation functions should be non-monotonic and be more general than the usual sigmoid function as well as the commonly used lipschitz condition. the equilibrium point y∗ = [y∗1, y ∗ 2, ..., y ∗ n] of system (1) will be shifted to the origin by the transformation y(·) = x(·)− x∗, transforms system (1) into the following form dy(t) = [−ay(t) +bg(y(t− τ)) + c ∫ t −∞ k(t− s)g(y(s))ds]dt+ σ(t, y(t))dw(t), t 6= tk (2) y(tk) = iky(t−k ), t = tk, k = 1, 2, ... y(t0 + t) = ψ(t), t ∈ [−τ, 0] where y = [y1, y2, ..., yn]t , a = diag[a1, a2, ..., an], b = [bij ], c = [cij ], k(t − s) = diag[k1(t − s), k2(t − s), ..., kn(t − s)], g(y) = [g1(y1), g2(y2), ..., gn(yn)] with gj(yj(t)) = fj(yj(t) + x∗j )− fj(x∗j ). note that since each function fj(·) satisfies the assumptions (c1) and (c2), hence each gj(·) satisfies g2j (ξj) ≤ l2 jξ 2, ξjgj(ξj) ≥ g2j (ξj) lj ∀ξj ∈ r, gj(0) = 0 (c3) there exist a constant matrix d0 such that trace[σt (t, y(t))σ(t, y(t))] ≤ yt (t) d0 y(t) lemma 2.2 [1] (s-procedure) let ti ∈ rn×n (i = 0, 1, ..., p) be symmetric matrices. the conditions on ti, (i = 0, 1, ..., p) αtt0α > 0, ∀α 6= 0 s.t. αttiα ≥ 0 (i = 0, 1, ..., p) hold, if there exist τi ≥ 0 (i = 0, 1, ..., p) such that t0 − p∑ i=1 τiti > 0 advances in systems science and applications (2011), vol. 11, no. 1-2 95 lemma 2.3 let u, v,w andm be real matrices of appropriate dimensions withm satisfying m = mt , then m + uvw +w tv tut < 0, ∀ v tv ≤ i if and only if there exist a scalar ε > 0 such that m + ε−1uut + εw tw < 0 definition 2.4 [25] the function v : [t0,∞)× rn → r+ belongs to class v0 if (1) the function v is continuous on each of the sets [tk−1, tk) × rn and for all t ≥ t0, v (0, t) ≡ 0; (2) v (x, t) is locally lipschitzian in x ∈ rn; (3) for each k = 1, 2, ..., there exist finite limits lim (q,t)→(x,t−k ) v (q, t) = v (x, t−k ) lim (q,t)→(x,t+k ) v (q, t) = v (x, t+k ) with v (x, t+k ) = v (x, tk) satisfied. in the following section, we will develop delay-dependent condition for the given system such that the origin of the delayed stochastic hopfield neural network (2) is asymptotically stable. 3. asymptotic stability criterion before discussing the stability analysis of the problem, we firstly introduce the ito’s formula for a general stochastic system. let v (y(t), t) : c([−τ, 0],rn × r+ → r+) be a positive function which is continuously twice differentiable in y and once differentiable in t. thus, an operator l acting on v (y(t), t), is defined by lv (y(t), t) = vt(y(t), t) + vy(y(t), t)[−ay(t) +bg(y(t− τ)) + c ∫ t −∞ k(t− s)g(y(s))ds] + 1 2 trace[σt (t, y(t))vyy(y(t), t)σ(t, y(t))] (3) where vt(y(t), t) = ∂v (y(t), t) ∂t , vy(y(t), t) = (∂v (y(t), t) ∂y1 , ∂v (y(t), t) ∂y2 , ..., ∂v (y(t), t) ∂yn ) vyy(y(t), t) = (∂2v (y(t), t) ∂yi∂yj ) n×n now, the following theorem gives a new stability criterion for system (2) without uncertain parameters. 96 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . theorem 3.1 suppose that the assumption (c1)− (c3) is satisfied. if there exist matrices p = [ p11 p12 p t12 p22 ] ≥ 0 with p11 > 0, r = [ r11 r12 rt12 r22 ] ≥ 0 q1 > 0, q2 > 0, h1, h2, h3, h4, diagonal matrices e > 0, s > 0 and a positive scalar ρ > 0 such that the following inequalities hold: p < ρi (4) itk p11ik + 2itk p12ik + itk p22ik − p11 − 2p12 − p22 < 0 (5) ω =  ξ11 ξ12 ξ13 ξ14 −τatp t12 + τp t22 −τh1 −τatr22 ∗ ξ22 ξ23 −ht 4 −τp t22 −τh2 0 ∗ ∗ ξ33 0 τbtp12 −τh3 −τbtr22 ∗ ∗ ∗ −e τctp12 −τh4 −τctr22 ∗ ∗ ∗ ∗ −τr11 −τr12 0 ∗ ∗ ∗ ∗ ∗ −τr22 0 ∗ ∗ ∗ ∗ ∗ ∗ −τr22  < 0 (6) where ξ11 = −p11a−atp t11 + p12 + p t12 + ρd0 +q1 + lel+ τr11 +h1 +ht 1 − τr12a ξ12 = p12 −h1 +ht 2 , ξ13 = p11b +ht 3 + τr12b ξ14 = p11c +ht 4 + τr12c, ξ22 = −q1 −h2 −ht 2 , ξ23 = ls −ht 3 , ξ33 = −q2 − 2s, l = {l1, l2, ..., ln}. then the origin of system (2) is the unique equilibrium point and it is globally asymptotically stable. proof. define new state variables g1(t) = −ay(t) +bg(y(t− τ)) + c ∫ t −∞ k(t− s)g(y(s))ds (7) g2(t) = σ(t, y(t)) (8) to prove the asymptotic stability result, let us consider the following lyapunov functional candidate for system (2) as v1 = δt1 (t)pδ1(t), v2 = ∫ t t−τ yt (s)q1y(s)ds, v3 = ∫ t t−τ gt (y(s))q2g(y(s))ds v4 = ∫ 0 −τ ∫ t t−τ δt2 (s)rδ2(s)dsdσ, v5 = n∑ j=1 ej ∫ ∞ 0 kj(ξ) ∫ t t−ξ g2j (yj(γ))dγdξ (9) advances in systems science and applications (2011), vol. 11, no. 1-2 97 where p = [ p11 p12 p t12 p22 ] ≥ 0 with p11 > 0, r = [ r11 r12 rt12 r22 ] ≥ 0 and δ1(t) = [ y(t)∫ t t−τ y(s)ds ] , δ2(t) = [ y(t) g1(t) ] , by newton-leibnitz formula, the following equation is true for any matrices hi (i = 1, 2, 3, 4) with appropriate dimensions: 2 [ yt (t)h1 + yt (t− τ)h2 + gt (y(t− τ))h3 + (∫ t −∞ k(t− s)g(y(s))ds ) h4 ] × [ y(t)− y(t− τ)− ∫ t t−τ g1(s)ds ] = 0 (10) when t 6= tk the derivative of v can be calculated by using ito’s differential formula. then the trajectories of the system (2) is given as: lv1 = 2δt1 (t)p δ̇1(t) = 2  y(t)∫ t t−τ y(s)ds t p11 p12 p t12 p22 −ay(t) +bg(y(t− τ)) + c ∫ t −∞k(t− s)g(y(s))ds y(t)− y(t− τ)  = −2yt (t)p11ay(t) + 2yt (t)p11bg(y(t− τ)) + 2yt (t)p11c (∫ t −∞ k(t− s)g(y(s))ds ) + 2yt (t)p12y(t)− 2yt (t)p12y(t− τ)− 2 ∫ t t−τ yt (s)p t12ay(t)ds+ ∫ t t−τ yt (s)p t12b × g(y(t− τ))ds+ 2 ∫ t t−τ yt (s)p22y(t)ds− 2 ∫ t t−τ yt (s)p22y(t− τ)ds+ 2 (∫ t t−τ yt (s)ds ) × p t12c (∫ t −∞ k(t− s)g(y(s))ds ) + trace[σt (t, y(t))pσ(t, y(t))] (11) lv2 = yt (t)q1y(t)− yt (t− τ)q1y(t− τ) (12) lv3 = gt (y(t))q2g(y(t))− gt (y(t− τ))q2g(y(t− τ)) (13) lv4 = τδt2 (t)rδ2(t)− ∫ t t−τ δt2 (s)rδ2(s)ds (14) lv5 = n∑ j=1 ej ∫ ∞ 0 kj(ξ)g 2 j (yj(t))dξ − n∑ j=1 ej ∫ ∞ 0 kj(ξ)g 2 j (yj(t− ξ))dξ = gt (y(t))eg(y(t))− n∑ j=1 ej ∫ ∞ 0 kj(ξ)dξ ∫ ∞ 0 kj(ξ)g 2 j (yj(t− ξ))dξ ≤ yt (t)lely(t)− n∑ j=1 ej (∫ ∞ 0 kj(ξ)g 2 j (yj(t− ξ))dξ )2 = yt (t)lely(t)− (∫ t −∞ k(t− s)g(y(s))ds ) e (∫ t −∞ k(t− s)g(y(s))ds ) (15) 98 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . it is noted from (c2) that, gi(yi(t− τ))[gi(yi(t− τ))− liyi(t− τ)] ≤ 0, i = 1, 2, ..., n (16) now, by applying the s-procedure, we find that system (2) is asymptotically stable, if there exist s = diag{s1, s2, ..., sn} such that lv = lv1 + lv2 + lv3 + lv4 + lv5 + 2 [ yt (t)h1 + yt (t− τ)h2 + gt (y(t− τ))h3 + (∫ t −∞ k(t− s)g(y(s))ds ) h4 ][ y(t)− y(t− τ)− ∫ t t−τ g1(s)ds ] ≤ lv1 + lv2 + lv3 + lv4 + lv5 + 2 [ yt (t)h1 + yt (t− τ)h2 + gt (y(t− τ))h3 + (∫ t −∞ k(t− s)g(y(s))ds ) h4 ][ y(t)− y(t− τ)− ∫ t t−τ g1(s)ds ] −2 n∑ i=1 sigi(yi(t− τ))(gi(yi(t− τ))− liyi(t− τ)) ≤ lv1 + lv2 + lv3 + lv4 + lv5 + 2 [ yt (t)h1 + yt (t− τ)h2 + gt (y(t− τ))h3 + (∫ t −∞ k(t− s)g(y(s))ds ) h4 ][ y(t)− y(t− τ)− ∫ t t−τ g1(s)ds ] −2gt (y(t− τ(t)))sg(y(t− τ(t))) + 2gt (y(t− τ(t)))sg(y(t− τ(t))) ≤ yt (t)[−p11a−atp t11 + p12 + p t12 + ρd0 +q1 + lel+ τr11 +h1 +ht 1 − τr12a] ×y(t) + yt (t)[p12 −h1 +ht 2 ]y(t− τ(t)) + yt (t)[p11b +ht 3 + τr12b]g(y(t− τ)) +yt (t)[p11c +ht 4 + τr12c] (∫ t −∞ k(t− s)g(y(s))ds ) + yt (t)[−τatp t12b + τp t22 × (∫ t t−τ y(s)ds ) + yt (t)[−τh1] (∫ t t−τ g1(s)ds ) + yt (t− τ)[−q1 −h2 −ht 2 ] ×y(t− τ) + yt (t− τ)[−ht 3 + ls]g(y(t− τ)) + yt (t− τ)[−τp t22] (∫ t t−τ y(s)ds ) +yt (t− τ)[−τh2] (∫ t t−τ g1(s)ds ) + gt (y(t− τ))[−q2 − 2s]g(y(t− τ)) + yt (t− τ)) ×[−ht 4 ] (∫ t −∞ k(t− s)g(y(s))ds ) + gt (y(t− τ)[τbtp12] (∫ t t−τ y(s)ds ) +gt (y(t− τ)[−τh3] (∫ t t−τ g1(s)ds ) + (∫ t −∞ k(t− s)g(y(s))ds )t [−e] advances in systems science and applications (2011), vol. 11, no. 1-2 99 × (∫ t −∞ k(t− s)g(y(s))ds ) + (∫ t −∞ k(t− s)g(y(s))ds )t [τctp12] (∫ t t−τ y(s)ds ) + (∫ t −∞ k(t− s)g(y(s))ds )t [−τh4] (∫ t t−τ g1(s)ds ) + (∫ t t−τ y(s)ds )t [−τr11] × (∫ t t−τ y(s)ds ) + (∫ t t−τ y(s)ds )t [−τr12] (∫ t t−τ g1(s)ds ) + (∫ t t−τ g1(s)ds )t [−τr22] × (∫ t t−τ g1(s)ds ) (17) it is easy to see y(t)− y(t− τ)− ∫ t t−τ g1(s)ds = ∫ t t−τ(t) g1(s)ds = 1 τ ∫ t t−τ(t) τ(τ−1τ(t)g1(s))ds (18) this together with (16), implies lv = 1 τ ∫ t t−τ(t) βt (t, s)ωβ(t, s)ds where βt (t, s) = [yt (t) yt (t− τ) gt (y(t− τ)) (∫ t −∞ k(t− s)g(y(s))ds )t yt (s) gt1 (s)] this implies that lv (y(t), t) < 0. when t = tk, we obtain the following result: v (y(tk), (tk))− v (y(tk), t − k ) = δt1 (tk) p11 p12 p t12 p22  δ1(tk)− δt1 (t−k ) p11 p12 p t12 p22  δ1(t−k ) = yt (t−k ) { itk p11 p12 p t12 p22  ik − p11 p12 p t12 p22 }y(t−k ) = yt (t−k )itk p11iky(t−k ) + yt (t−k )itk p12iky(t−k ) + yt (t−k )itk p t 12iky(t−k ) + yt (t−k )itk p22iky(t−k )− yt (t−k )p11y(t−k )− yt (t−k )p12y(t−k ) − yt (t−k )p t12y(t−k )− yt (t−k )p22y(t−k ) = yt (t−k ) [ itk p11ik + 2itk p12ik + itk p22ik − p11 − 2p12 − p22 ] y(t−k ) based on the lyapunov stability theorem, it follows that the delayed stochastic hopfield neural network (2) is globally asymptotically stable in the mean square. the proof of the theorem is completed. � 100 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . 4. robust asymptotic stability criterion consider the system (2) with norm bounded parameter uncertainties that is dy(t) = [−(a+ ∆a(t))y(t) + (b + ∆b(t))g(y(t− τ)) + (c + ∆c(t)) ∫ t −∞ k(t− s) g(y(s))ds]dt+σ(t, y(t))dw(t), t 6= tk (19) y(tk) = iky(t−k ), t = tk, k = 1, 2, ... where a+ ∆a(t), b + ∆b(t) and c + ∆c(t) are of the following structure: [∆a(t) ∆b(t) ∆c(t)] = mf (t)[n1 n2 n3] where m,n1, n2, n3 are known constant matrices with appropriate dimensions and bounded which satisfies f t (t)f (t) ≤ i, t ≥ 0 theorem 3.2 suppose that the assumption (c1)− (c3) is satisfied. if there exist matrices p = [ p11 p12 p t12 p22 ] ≥ 0 with p11 > 0, r = [ r11 r12 rt12 r22 ] ≥ 0 q1 > 0, q2 > 0, h1, h2, h3, h4, diagonal matrices e > 0, s > 0 and four positive scalars ρ > 0, ε1 > 0, ε2 > 0, ε3 > 0 such that the following inequalities hold: p < ρi (20) itk p11ik + 2itk p12ik + itk p22ik − p11 − 2p12 − p22 < 0 (21) advances in systems science and applications (2011), vol. 11, no. 1-2 101 ω1 = θ11 θ12 θ13 θ14 −τatp t12 + τp t22 −τh1 p11m 0 −τatr22 0 ∗ ξ22 ξ23 −ht 4 −τp t22 −τh2 0 0 0 0 ∗ ∗ θ33 0 τbtp12 −τh3 0 0 −τbtr22 0 ∗ ∗ ∗ θ44 τctp12 −τh4 0 0 −τctr22 0 ∗ ∗ ∗ ∗ −τr11 −τr12 0 τr12m 0 0 ∗ ∗ ∗ ∗ ∗ −τr22 p22m 0 0 0 ∗ ∗ ∗ ∗ ∗ ∗ −ε1i 0 0 0 ∗ ∗ ∗ ∗ ∗ ∗ ∗ −ε2i 0 0 ∗ ∗ ∗ ∗ ∗ ∗ ∗ ∗ −τr22 τr22m ∗ ∗ ∗ ∗ ∗ ∗ ∗ ∗ ∗ −ε3i  < 0 (22) where θ11 = −p11a−atp t11 + p12 + p t12 + ρd0 +q1 + lel+ τr11 +h1 +ht 1 − τr12a + ε1n t 1 n1 + τε2n t 1 n1 + τε3n t 1 n1 θ13 = p11b +ht 3 + τr12b − ε1nt 1 n2 − τε2nt 1 n2 − τε3nt 1 n2, θ14 = p11c +ht 4 + τr12c − ε1nt 1 n3 − τε2nt 1 n3 − τε3nt 1 n3, θ33 = −q2 − 2s + ε1n t 2 n2 + τε2n t 2 n2 + τε3n t 2 n2, θ44 = −e + ε1n t 3 n3 + τε2n t 3 n3 + τε3n t 3 n3 and ξ22,ξ23, l are stated as in theorem 3.1. then the origin of system (2) is the unique equilibrium point and it is globally asymptotically stable. proof. in order to prove the robust asymptotic stability, we use the same lyapunov-krasovskii functional as defined in (9). by replacing a,b and c in (2) with a+ ∆a(t), b + ∆b(t) and c + ∆c(t), respectively and by ito’s differential formula, we can calculate the trajectories of 102 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . the system (19), then we have ω1 = ω + ε−11  p11m 0 0 0 0 p22m  × [ mtp11 0 0 0 0 mtp22 ] + ε1  −n1 0 n2 n3 0 0  × [ −nt 1 0 nt 2 nt 3 0 0 ] +τε−12  0 0 0 0 r12m 0  × [ 0 0 0 0 mtr12 0 ] + τε2  −n1 0 n2 n3 0 0  × [ −nt 1 0 nt 2 nt 3 0 0 ] +τε−13  0 0 0 0 0 r22m  × [ 0 0 0 0 0 mtr22 ] + τε3  −n1 0 n2 n3 0 0  × [ −nt 1 0 nt 2 nt 3 0 0 ] therefore, under condition (20)-(22), system (19) is robustly globally asymptotically stable with respect to the uncertain parameters ∆a(t), ∆b(t) and ∆c(t). this completes the proof of the theorem. � advances in systems science and applications (2011), vol. 11, no. 1-2 103 if we neglect the impulsive term and stochastic perturbations in (2), then it reduces to ẏ(t) = −ay(t) +bg(y(t− τ)) + c ∫ t −∞ k(t− s)g(y(s))ds (23) corollary 3.3 suppose that the assumption (c1)− (c3) is satisfied. if there exist matrices p = [ p11 p12 p t12 p22 ] ≥ 0 with l11 > 0, r = [ r11 r12 rt12 r22 ] ≥ 0 q1 > 0, q2 > 0, h1, h2, h3, h4 and diagonal matrices e > 0, s > 0 such that the following inequalities hold: ω =  ξ11 ξ12 ξ13 ξ14 −τatp t12 + τp t22 −τh1 −τatr22 ∗ ξ22 ξ23 −ht 4 −τp t22 −τh2 0 ∗ ∗ ξ33 0 τbtp12 −τh3 −τbtr22 ∗ ∗ ∗ −e τctp12 −τh4 −τctr22 ∗ ∗ ∗ ∗ −τr11 −τr12 0 ∗ ∗ ∗ ∗ ∗ −τr22 0 ∗ ∗ ∗ ∗ ∗ ∗ −τr22  < 0 (24) where ξ11 = −p11a−atp t11 + p12 + p t12 +q1 + lel+ τr11 +h1 +ht 1 − τr12a ξ12 = p12 −h1 +ht 2 , ξ13 = p11b +ht 3 + τr12b ξ14 = p11c +ht 4 + τr12c, ξ22 = −q1 −h2 −ht 2 , ξ23 = ls −ht 3 , ξ33 = −q2 − 2s, l = diag{l1, l2, ..., ln}. then the origin of system (2) is the unique equilibrium point and it is globally asymptotically stable. proof. by arguing similar to the proof of theorem 1, we can show that the equilibrium point of system (2) is globally asymptotically stable in the mean square. this completes the proof of the theorem. � remark 3.4 the authors chen and cao, [2] discussed the global asymptotic stability of delayed hopfield neural networks. wan et al. investigated the mean square exponential stability of stochastic delayed hopfield neural networks. in [18], wang et al. proposed the robust stability for stochastic hopfield neural networks with time delays and zhang et al., [24, 26] obtained the global stability results for delayed hopfield neural network. as a result, in all the above mentioned references, impulsive effect has not been taken into account. however, in our paper, we 104 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . derived the delay-dependent stability results for stochastic hopfield neural networks with impulsive effects. therefore, our main result is new, quite effective and leads to less conservative results when compared with some existing works [2, 8, 19, 20]. remark 3.5 it is noteworthy that in our paper, we employed several free weighting matrices and s-procedure to derive the delay-dependent stability criterion. the derived criterion is obtained in lmi forms whose feasibility can be readily checked by using the matlab lmi toolbox. different from the conventional stability criteria that depend on the m-matrix computation, no tuning of parameters will be needed when employing our lmi-based stability criteria. moreover, two numerical examples with simulation results will show the effectiveness of the stability conditions in this paper. 5. illustrated examples in this section, we provide two numerical examples to demonstrate the effectiveness of the main results presented in this paper. example 4.1 consider a delayed stochastic hopfield neural network (2) with parameters as: a = [ 1.7679 0 0 1.8860 ] , b = [ −0.2376 −0.4769 −0.6707 −0.7654 ] , c = [ −0.1052 −0.5069 −0.0257 −0.2808 ] , l = [ 0.5219 0 0 1.8993 ] , d0 = [ 0.33 0 0 0.25 ] , ik = i = [ 0.2 0 0 0.2 ] it can be checked that system (2) satisfies the assumptions (c1) − (c3). for the delay bound τ = 0.957, we have obtained the following feasible solutions to the lmis (4) (6) in theorem 1 p11 = [ 25.7039 −1.6632 −1.6632 39.0949 ] , p12 = [ 4.3286 −0.9483 −0.9483 3.8977 ] , p22 = [ 8.8576 −1.6157 −1.6157 7.3501 ] q1 = [ 18.9787 −6.6658 −6.6658 16.4166 ] , q2 = [ 29.9330 20.6811 20.6811 51.6907 ] , r11 = [ 21.3444 −6.6192 −6.6192 19.6772 ] r12 = [ 6.4801 −1.9881 −1.9881 4.5859 ] , r22 = [ 7.2495 −2.6422 −2.6422 3.5791 ] , h1 = [ −1.5463 0.9452 −1.2486 1.3717 ] h2 = [ 1.6552 0.8485 −0.6090 1.9241 ] , h3 = [ 1.8527 1.2408 2.4549 1.7661 ] , h4 = [ −1.0116 2.4402 −1.5775 1.8661 ] advances in systems science and applications (2011), vol. 11, no. 1-2 105 s = [ 18.3098 0 0 8.9915 ] , e = [ 22.2513 0 0 16.6959 ] , ρ = 1.8663× 103 in order to show the significant improvement of our results, we summerize the comparisons between the previous works and the obtained result. for this example, the delay-dependent stability analysis in [3, 26, 27], cannot be satisfied for any τ > 0. table 1 shows the maximum upper bound of the previous works [8, 19, 20] as 0.4121, 1.7484 and 1.7644, respectively. however, by theorem 1, we have that the origin of delayed stochastic hopfield neural networks with impulsive effect is globally asymptotically stable for any constant allowable upper bound τ > 0. hence, it is clear that the proposed method shows the less conservativeness than the existing works [2, 8, 19, 20]. fig. 1 state trajectories of y1, y2 for example 1 example 4.2 consider a delayed stochastic uncertain hopfield neural network (2) with the following parameters: a = [ 0.7679 0 0 0.8860 ] , b = [ −0.1746 −0.8642 −0.2892 −0.7300 ] , c = [ −0.8252 −0.4912 −0.4732 −0.8858 ] , m = [ 0.07051 0 0 0.0342 ] , n1 = [ 0.3526 −0.1904 0.3322 −0.1564 ] , n2 = [ 0.2446 0.3674 −0.1753 0.2956 ] , 106 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . table 1: maximum allowable bound of the delay method maximum upper bound of τ in ref [2, 26, 27] in ref [8] 0.4121 in ref [19] 1.7484 in ref [20] 1.7644 in this paper for any large finite τ > 0 n3 = [ 0.1981 −0.1313 0.1185 0.1645 ] , l = [ 0.07051 0 0 0.0342 ] , d0 = [ 0.85 0 0 0.85 ] , for τ = 0.957 and by solving the lmis (20)-(22) in theorem 3.2, we get the following feasible solution as follows: p11 = [ 133.1390 −39.2261 −39.2261 103.1396 ] , p12 = [ 5.2825 −5.7032 −5.7032 10.4098 ] , p22 = [ 18.1026 −10.6945 −10.6945 25.6053 ] q1 = [ 30.8319 −27.8887 −27.8887 47.6789 ] , q2 = [ 90.1163 28.7709 28.7709 147.2412 ] , r11 = [ 39.5332 −37.1370 −37.1370 66.5761 ] r12 = [ 21.2050 −14.1850 −14.1850 21.1791 ] , r22 = [ 31.5030 −17.8344 −17.8344 21.7477 ] , h1 = [ −13.2142 4.3068 −3.1365 −1.7307 ] h2 = [ 13.9993 −1.7353 −3.6838 8.4480 ] , h3 = [ −9.5150 11.0562 −15.8051 17.3187 ] , h4 = [ 12.1911 −1.3902 −7.3453 12.0990 ] s = [ 72.1101 0 0 263.5610 ] , e = [ 491.5915 0 0 339.6219 ] , ρ = 7.1793× 103 ε1 = 26.1631, ε2 = 15.5658, ε3 = 16.3759. thus, all the conditions in theorem 3.2 are satisfied. therefore, the model (19) with above given parameters is globally asymptotically stable in the mean square. 6. conclusion this paper has studied the problem of stability analysis for a class of delayed stochastic uncertain hopfield neural networks with distributed time-varying delays and impulses. a delay-dependent asymptotic stability condition is developed in terms of an lmi, which can be easily checked by using recently developed algorithms in solving lmis. finally, two numerical examples has been provided to demonstrate the usefulness and the reduced conservatism of the proposed results. advances in systems science and applications (2011), vol. 11, no. 1-2 107 references [1] s. boyd, l. e. ghaoui, e. feron and v. balakrishnan, linear matrix inequalities in system and control theory, philadelphia, pa:siam, 1994. [2] a. chen, j. cao and l. huang, an estimation of upper bound of delays for global asymptotic stability of delayed hopfield neural networks, ieee transactions on circuits and systems i: fundamental theory and applications, 49 (2002) 1028-1032. [3] c. j. chen, t. l. liao and c. c. hwang, exponential synchronizations of a class of chaotic neural networks, chaos, solitons and fractals, 24 (2005) 197-206. [4] l. chua and l. yang, celluar neural networks: theory and applications, ieee transactions on circuits and systems i, 35 (1998) 1257-90. [5] h. huang and j. cao, exponential stability analysis of uncertain stochastic neural networks with multiple delays, nonlinear analysis: real world applications., in press. [6] h. huang, d. w. c. ho and j. lam, stochastic stability analysis of fuzzy hopfield neural networks with time-varying delays, ieee transactions on circuits and systems ii, 52 (2005) 251-255. [7] j. j. hopfield, neural networks and physical systems with emergent collective computational abilities, proceeding of the national academy of sciences, 79 (1982) 2554-2558. [8] h. he, a. n. michel and k. wang, global stability and local stability of hopfield neural networks with delays, physics review e, 50 (1994) 4206-4213. [9] y. r. liu, z. d. wang and x. h. liu, on global exponential stability of generalized stochastic neural networks with mixed time-delays, neurocomputing, 70 (2006) 314-326. [10] x. x. liao and x. mao, exponential stability and instability of stochastic neural networks, stochastic analysis and its applications, 14 (1996) 165-185. [11] v. lakshikantham, d. bainov and p. s. simenov, theory of impulsive differential equations, world scientific, singapore. [12] ju. h. park, on global stability criterion for neural networks with discrete and distributed delays, chaos, solitons and fractals, 30 (2006) 897-902. [13] h. qiao, j. peng and z. xu, nonlinear measures: a new approach to exponential stability analysis for hopfield-type neural networks, ieee transactions on neural networks, 12 (2001) 360-370. [14] v. singh, simplified lmi condition for global asymptotic stability of delayed neural networks, chaos, solitons and fractals, 29 (2006) 470-473. [15] p. van den driessche and x. zou, global attractivity in delayed hopfield neural networks model, siam journal on applied mathematics, 58 (1998) 1878-1890. 108 raja: delay-dependent stability criteria of stochastic uncertain hopfield . . . . . . [16] l. wan and j. sun, mean square exponential stability of stochastic delayed hopfield neural networks, physics letters a, 343 (2005) 306-318. [17] z. wang, y. liu, k. fraser and x. liu, stochastic stability of uncertain hopfield neural networks with discrete and distributed delays, physics letters a, 354 (2006) 288-297. [18] z. wang, h. shu, j. fang and x. liu, robust stability for stochastic hopfield neural networks with time delays, nonlinear analysis: real world applications, 7 (2006) 11191128. [19] d. y. xu and z. c. yang, impulsive delay differential inequality and stability of neural networks, journal of mathematical analysis and its applications, 305 (2005) 107-120. [20] d. y. xu, w. zhu and s. j. long, global exponential stability of impulsive integrodifferential equation, nonlinear analysis, 64 (2006) 2805-2816. [21] h. yang and t. chu, lmi conditions for stability of neural networks with distributed delays, chaos solitons fractals, 34 (2007) 557-563. [22] t. yang, impulsive control, ieee transactions on automatic control, 44 (1999) 10811083. [23] q. zhang, x. g. wei and j. xu, delay-dependent global stability results for delayed hopfield neural network, chaos solitons and fractals, 34 (2007) 662-668. [24] j. zhang and x. jin, global stability analysis in delayed hopfield neural network models, neural networks, 13 (2000) 745-753. [25] y. zhang and j. t. sun, stability of impulsive neural networks with time delays, physics letters a, 348 (1-2) (2005) 44-50. [26] q. zhang, x. wei and j. xu, global asymptotic stability of hopfield neural networks with transmission delays, phys lett a, 318 (2003) 399-405. [27] q. zhang, x. wei and j. xu, delay-dependent exponential stability of cellular neural networks with time-varying delays, chaos, solitons and fractals, 23 (2005) 1363-1369. advances in systems science and applications (2011), vol. 11, no. 1-2 109 advances in systems science and applications (2017) vol.17 no.1 46 dynamic models of mob excitation control1 ivan n. barabanov a and dmitry a. novikov a,b a) v.a. trapeznikov institute of control sciences of russian academy of sciences, 117997, profsoyuznaya street, 65, moscow, russia; b) moscow institute of physics and technology, 141700, institutskiy per., 9, dolgoprudny, moscow region, russia e-mail: ivbar@ipu.ru, novikov@ipu.ru abstract. this paper formulates and solves the mob excitation control problem in the continuous-time setting by introducing an appropriate number of “provokers” at each moment of control. key words: collective behavior, granovetter’s model, stochastic models of mob control. 1. introduction in the case of conformity decision-making of agents that perform binary choice between “passivity” and “activity,” the associated control problems involve the models of collective behavior based on classical granovetter’s model [1] (see the surveys in [2, 3]). within this model, each agent is characterized by his threshold, a number from the interval [0; 1], in the following way. he decides to be active if the fraction of active agents in his neighborhood exceeds the threshold; otherwise, the agent prefers passivity. the dynamics of the fraction of active agents depends on its initial value and the distribution function of the agents’ thresholds. hence, a goal-directed (exogenous) change in the number of agents at the initial and/or subsequent moments affects the behavioral dynamics of the whole group. consider the collective behavior of agents forming a mob [4]. in this case, control consists in choosing a fraction of always active agents to-be-introduced to the mob (the so-called “provokers”). the problems belonging to this class can be classified by several bases such as discrete-time setting or continuous-time setting, single or multiple application of the control actions (constant controls and time-varying controls, respectively), and open-loop or feedback control. the discrete-time optimal choice problem for a single constant control applied by a control subject (principal) to a mob was formulated and solved in [4]. in [5], this problem was generalized to the case of multiple open-loop controls in the discrete-time setting. the present paper is dedicated to the continuous-time mob control models. further exposition follows the paper [6], in particular, the proofs of propositions 2-6 can be taken there. the mob model proper is imported from [7, 8], actually representing a generalization of granovetter’s model to the continuous-time case as follows. suppose that we know the initial fraction x0  [0; 1] of active agents at the zero moment. then the evolution of this fraction x(t) in the continuous time t ≥ 0 is governed by the equation ( )x f x x  , (1) where f(∙) is a known continuous function possessing the properties of a distribution function, f(0) = 0 and f(1) = 1. actually, this is the distribution function of the agents’ thresholds [2, 4]. similar to [4, 5], by applying a control action u(t)  [0; 1] (introducing provokers), we obtain the controlled dynamic system ( ) (1 ( )) ( )x u t u t f x x    . (2) this paper is organized in the following way. section 2 studies the reachability set and the monotonicity of the system trajectories in the control action. next, section 3 is focused on the case of constant controls according to the above classification. section 4 considers the models 1 this work was partially supported by rsf grant 16-19-10609. 47 i.n. barabanov and d.a. novikov: dynamic models of mob excitation control where control excites the whole mob. and finally, section 5 deals with the case of feedback control. 2. reachability set and monotonicity first, we formulate a lemma required for further analysis. consider functions 1( , )g x t and 2 ( , )g x t : 0[ , )r t r   that are continuously differentiable with respect to x and continuous in t . by assumption, the functions 1g and 2g are such that the solutions to the cauchy problems for the differential equations ( , )ix g x t , 1,2i  , with initial conditions 0 0 0( , ),t x x r , admit infinite extension in t . denote by 0 0( ,( , )), 1,2ix t t x i  , the solutions of the corresponding cauchy problems. lemma [6]. let 0 1 2, ( , ) ( , )x r t t g x t g x t      . then 0 1 0 0( ,( , ))t t x t t x   2 0 0( , ( , ))x t t x . note that, for validity of this lemma, one should not consider the inequality 1 2( , ) ( , )g x t g x t for all x r . it suffices to take the union of the reachability sets of the equations ( , )ix g x t , 1,2i  , with the chosen initial conditions 0 0( , )t x . denote by xt(u) the fraction of active agents at the moment t under the control action u(∙). the right-hand side of the expression (1) increases monotonically in u for each t and ∀ x ∈ [0; 1]: f(x) ≤ 1. hence, we have the following result. proposition 1. let the function f(x) be such that f(x) < 1 for x ∈ [0; 1). if 0 1 2( ) ( )t t u t u t    and x0(u1) = x0(u2) ( 0 1x  ), then 0t t  : xt(u1) > xt(u2). indeed, by the premises, for all t and x<1 we have the inequality 1 1 2 2( ) (1 ( )) ( ) ( ) (1 ( )) ( )u t u t f x x u t u t f x x       , as the convex combination of different numbers (1 and f(x)) is strictly monotonic. the point x=1 forms the equilibrium of the system (1) under any control actions u(t). and so, it is unreachable for any finite t. using the above lemma, we find that xt(u1) > xt(u2) under same initial conditions. suppose that the control actions are subjected to the constraint   0    ,     ,u t t t  (3) where δ ∈ [0; 1] means some constant. we believe that t0=0, x(t0)=x(0)=0, i.e., initially the mob is in the nonexcited state. if the efficiency criterion is defined as the fraction of active agents at a given moment t > 0, then the corresponding terminal control problem takes the form ( ) ( ) max, (2), (3). t u x u     (4) here is a series of results (propositions 2-4) representing the analogs of the corresponding propositions from [5]. proposition 2. the solution of the problem (4) is given by u(t) = δ, t ∈ [0; t]. denote by ˆ ˆ( , ) min { 0 | ( ) }tx u t x u x    the first moment when the fraction of active agents achieves a required value x̂ (if the set ˆ{ 0 | ( ) }tt x u x  is empty, just specify ˆ( , )x u = +∞). within the current model, one can pose the following time-optimal problem: ( ) ˆ( , ) min, (2), (3). u x u     (5) proposition 3. the solution of the problem (5) is given by u(t) = δ, t ∈ [0; τ]. by analogy with the discrete-time models [5], the problem (4) or (5) has the following practical interpretation. the principal benefits most from introducing the maximum admissible number of provokers in the mob at the initial moment, doing nothing after that (e.g., instead of first decreasing and then again increasing the number of introduced provokers). this structure of advances in systems science and applications (2017) vol.17 no.1 48 the optimal solution can be easily explained, as in the models (4) and (5) the principal incurs no costs to introduce and/or keep the provokers. what are the properties of the reachability set d = ( ) [0; ] ( )t u t x u   ? clearly, [0;1]d , since the right-hand side of the dynamic system (2) vanishes for x = 1. in the sense of potential applications, a major interest is attracted by the case of constant controls (u(t) = v, t ≥ 0). here the principal chooses the same fraction v  [0; ∆] of provokers at all moments. let xt(∆) = xt(u(t) ≡ δ), t ∈ [0; t], and denote by d0 = [0; ] ( ) [0;1]t v x v    the reachability set under constant control actions. according to proposition 1, ( )tx v represents a monotonic continuous mapping of [0; ∆] into [0; 1] such that (0) 0tx  . this leads to the following. proposition 4. d0 = [0; xt(∆)]. consider models taking into account the principal’s control costs. given a fixed “price” λ ≥ 0 of one provoker per unit time, the principal’s costs over a period τ ≥ 0 are defined by   0 (      .)u t dtc u     (6) suppose that we know a pair of monotonic functions characterizing the principal’s terminal payoff h(∙) from the fraction of active agents and his current payoff h(∙). then the problem (4) can be “generalized” to 0 ( ( )) ( ( )) ( ) max, (2), (3). t t t u h x u h x t dt c u         (7) under existing constraints on the principal’s “total” costs c, the problem (7) acquires the form 0 ( ( )) ( ( )) max, (2), ( ) . t t u t h x u h x t dt c u c        (8) a possible modification of the problems (4), (5), (7), (8) is the one where the principal achieves a required fraction x̂ of active agents by the moment t (the cost minimization problem): ( ) min, ˆ( ) , (2). t u t c u x u x      (9) the problems of the form (7)-(9) can be easily reduced to standard optimal control problems. example 1. consider the problem (9), where ( )f x x , x0 = 0 and the principal’s costs defined by (6) with 0 1  . this yields the following optimal open-loop control problem with fixed bounds: [0, ] 0 (1 ), ˆ(0) 0, ( ) , 0 , ( ) min . t u x u x x x t x u u t dt           (10) 49 i.n. barabanov and d.a. novikov: dynamic models of mob excitation control for the problem (10), construct the hamilton-pontryagin function  (1 )h u x u   . by the maximum principle, this function takes the maximum values in u . as h is linear in u , its maximum is achieved at an end of the interval [0, ] depending on the sign of the factor at u , i.e.,    sign 1 1 1 . 2 u x      (11) the fact that the hamilton-pontryagin function is linear in control actions actually follows from the same property of the right-hand side of the dynamic system (2) and the functional (6). in other words, we have the result below. proposition 5. if the constraints in the optimal control problems (7)-(9) are linear in control actions, then the optimal open-loop control possesses the structure described by (11). that is, at each moment the control action takes either the maximum or the minimum admissible value. the hamilton equations acquire the form (1 ), . h x u x h u x              the boundary conditions are imposed on the first equation only. for 0u  , its solution is a constant; for u   , the solution is    0 0( ) 1 1 ( ) . t t x t x t e      the last expression restricts the maximum number of provokers required for mob transfer from the zero state to x̂ : 1 1 log . ˆ1t x    and there exists the minimum time min 1 1 log , ˆ1 t x    during which control actions take the maximum value  , being 0 at the rest moment. particularly, a solution of the problem (10) has the form min min , , 0, t t u t t t       (12) when the principal introduces the maximum number of provokers from the very beginning, maintaining it during the time mint . the structure of the optimal solution to this problem (a piecewise constant function taking the value of 0 or  ) possibly requires minimizing the number of control switchovers (discontinuity points). such an additional constraint reflects situations when the principal incurs extra costs to introduce or withdraw provokers. if this constraint appears in the problem, the best control actions in the optimal control set are either (12) or min min , [ , ], 0, . t t t t u t t t        3. constant control in the class of the constant control actions, we obtain cτ(v) = λ v τ from formula (6). under given functions f(∙), i.e., a known relationship xt(v), the problems (7)-(9) are reduced to standard scalar optimization problems. example 2. choose f(x) = x, t = 1, x0 = 0, h(x) = x, and h(x) = γ x, where γ ≥ 0 is a known constant. it follows from (2) that   0   1   – ex ( .)p  t t u yx dyu         (13) for the constant control actions, xt(v) = 1 – vte . advances in systems science and applications (2017) vol.17 no.1 50 the problem (7) becomes the scalar optimization problem [0; ] 1 maxv v e v v v               . (14) next, the problem (8) becomes the scalar optimization problem [0; ] 1 maxv v e v v             . (15) and finally, the problem (9) acquires the form [0;1] min, ˆ1 . v v v e x       its solution is described by v = 1 log ˆ1 x       . 4. excitation of whole mob consider the “asymptote” of the problems as t → +∞. similarly to the corresponding model in [5], suppose that (a) the function f(∙) has a unique inflection point and f(0) = 0, (b) the equation f(x) = x has a unique solution q > 0 on the interval (0; 1) so that (0; ) ( ) , ( ;1) ( )x q f x x x q f x x      . several examples of the functions f(∙) satisfying these assumptions are provided in [5]. the principal seeks to excite all agents with the minimum costs. by the above assumptions on f(∙), if for some moment τ we have x(τ) > q, then the trajectory xt(u) is nonincreasing and lim ( ) 1t t x u   even under u(t) ≡ 0 t   . as mentioned in [5], this property of the mob admits the following interpretation. the domain of attraction of the zero equilibrium without control (without introduced provokers) is the half-interval [0; q). in other words, it takes the principal only to excite more than the fraction q of the agents; subsequently, the mob itself surely converges to the unit equilibrium even without control. denote by u τ the solution of the problem : ( ) [0; ], ( ) 0 ( ) min u u t x u q u t dt       (16) calculate qτ = 0 ( )u t dt    and find τ * = arg 0 min   qτ. the solution to the problem (16) exists under the condition [0, ] * ( ) max 1 ( )   x q x f x f x      (17) for practical interpretations, we refer to [5]. owing to the above assumptions on the properties of the distribution function, the optimal solution to the problem is characterized as follows. proposition 6. if the condition (17) holds, then u τ (t) ≡ 0 for t > τ. example 3. the paper [9] constructed the two-parameter function f(∙) describing in the best way the evolvement of active users in the russian-language segments of online social networks livejournal, facebook and twitter. the role of the parameters is player by a и b. this function has the form      , arctan( ( )) arctan( ) , arctan( (1 )) arctan( )   7;1  5 ,   0;1  .a b a x b ab f x a b a b ab        choose a = 13 that corresponds to facebook and b = 0.4. in this case, q ≈ 0.375 and * ≈ 0.169; the details can be found in [5]. 51 i.n. barabanov and d.a. novikov: dynamic models of mob excitation control 5. feedback control in the previous sections, we have considered the optimal open-loop control problem arising in mob excitation. an alternative approach is to use feedback control. consider two possible statements having transparent practical interpretations. within the first statement, the problem is to find a feedback control law ( ) :[0;1] [0;1]u x  ensuring maximum mob excitation (in the sense of (4) or (5)) under certain constraints imposed on the system trajectory and/or control actions. by analogy with the expression (3), suppose that the control actions are bounded:  ( ) , 0;1 ,xu x   (19) and there exists an additional constraint on the system trajectory in the form ( ) ,    0,tx t   (20) where δ > 0 is a known constant. the condition (20) means that, e.g., a very fast growth of the fraction of excited agents (increment per unit time) is detected by appropriate authorities banning further control. hence, trying to control mob excitation, the principal has to maximize the fraction of excited agents subject to the conditions (19) and (20). the corresponding problem possesses the simple solution * ( ) ( ) min ; max 0; 1 ( ) x f x u x f x            , (21) owing to the properties of the dynamic system (2), see the lemma. the fraction in (21) results from making the right-hand side of (1) equal to the constant δ. note that, under small values of δ, the nonnegative control action satisfying (20) may cease to exist. the second statement of feedback control relates to the so-called network immunization problem [4]; here the principal seeks to reduce the fraction of active agents by introducing an appropriate number (or fraction) of immunizers–agents that always prefer passivity. denote by w ∈ [0; 1] the fraction of immunizers. as shown in the paper [4], the fraction of active agents evolves according to the equation  (1 ) ( ) ;1  ., 0x w f x x x   (22) let ( ) :[0;1] [0;1]w x  be a feedback control. if the principal is interested in reducing the fraction of active agents, i.e., ( ) 0, 0,x t t  (23) then the control actions must satisfy the inequality ( ) 1 ( ) x w x f x   . (24) the quantity min [0;1] max 1 ( )x x f x         characterizes the bottom restrictions on the control actions at each moment when the system (22) is “controllable” in the sense of (23). 6. conclusion this paper has described the continuous-time problems of mob excitation control using introduction of provokers or immunizers. a promising line of future research is analysis of a differential game describing informational opposition of two control subjects (principals) that choose in continuous time the fractions (or numbers) of introduced provokers u and immunizers w, respectively. the corresponding static problem [7] can be a “reference model” here. the controlled object is defined by the dynamic system [4, 10] (1 ) (1 2 ) ( )x u w u w uw f x x       . another line of interesting investigations concerns the mob excitation problems with dynamic (open-loop and/or feedback) control, where the mob dynamics is modeled by the transfer advances in systems science and applications (2017) vol.17 no.1 52 equation [7] of the form       , (1 ) , 0p x t u u f x x p x t t x           . in this model, the mob state at each moment is described by a probability distribution function p(x, t), instead of the scalar fraction of active agents. references [1] granovetter, m., “threshold models of collective behavior”, ajs, vol. 83, no 6, pp. 1420-1443, 1978. [2] breer, v. v., “models of conformity behavior: a survey”, control sciences, vol. 1, pp. 2-13, 2014 [in russian]. [3] slovokhotov, yu. l., “physics and sociophysics. part 1”, control sciences, vol. 1, pp. 220, 2012 [in russian]. [4] breer, v. v., novikov, d. a., and rogatkin, a. d., “stochastic models of mob control”, automation and remote control, vol. 77, no. 5, pp. 895-913, 2016 [5] barabanov, i. n., and novikov, d. a., “dynamic models of mob excitation control in discrete time”, automation and remote control, vol. 77, no. 10, pp. 1792-1804, 2016. [6] barabanov, i. n., and novikov, d. a., “dynamic models of mob excitation control in continuous time”, upravlenie bol'simi sistemami (large scale systems control), vol. 63, pp. 71-68, 2016 [in russian]. [7] rogatkin, a. d., “continuous-time granovetter’s model”, upravlenie bol'simi sistemami (large scale systems control), vol. 60, pp. 139-160, 2016 [in russian]. [8] akhmetzhanov, a.r., worden, l., and dushoff, j., “effects of mixing in threshold models of social behavior”, phys. rev. e, vol. 88, no. 1, 2013, article number 012816. [9] batov, a. v., breer, v. v., novikov, d. a., and rogatkin, a. d., “microand macromodels of social networks. ii. identification and simulation experiments”, automation and remote control, vol. 77, no. 2, pp. 321-331, 2016. [10] novikov, d. a., “models of informational confrontation in mob control”, automation and remote control, vol.77, no.7, pp. 1259-1274, 2016. microsoft word 1 d.h. chen, k. ushijima--estimation of maximum compressive load for circular tubes under axial impact.doc advances in systems science and applications (2011), vol.11, no.3-4 203-213 issn 1078-6236 international institute for general systems studies, inc. estimation of maximum compressive load for circular tubes under axial impact d.h. chen1 and k. ushijima2 1department of mechanical engineering, faculty of engineering, tokyo university of science, tokyo, 1628601, japan 2department of mechanical engineering, faculty of engineering, kyushu sangyo university, fukuoka, 8138503,japan abstract the influence of impact velocity on the crushing behaviour of cylindrical shells subjected to an axial impact was investigated using a finite element analysis. the effects of the material properties, tube geometries and impact velocity v0 on the initial peak stressσ1 are explored. in this study, the applied material is assumed to be insensitive to the strain rate, and the effect of impact velocity is discussed as an inertia effect. it is shown that the initial peak stressσ1 during dynamic loading increases with increase of the impact velocity v0, which is due to the fact that the displacement in radial direction is delayed as the velocity v0 increases. also, based on our numerical simulations, the peak stressσ1 can be regarded as a function of the ratio of tube thickness to radius t/r, hardening modulus to young's modulus eh/e and impact velocity to elastic stress wave speed v0/c. moreover, an approximate equation to evaluate the peak stress is proposed and in good agreement with the fem results and other researcher's results under a relatively low impact velocity (v0<40m/s). keywords elastic-plastic cylindrical shells, axial impact, energy absorption, fem 1.introduction thin-walled structures such as circular and square tubes have been widely used in automotive and aerospace engineering as impact energy absorbing devices. a large number of studies concerning the static and dynamic response of thin-walled tubes subjected to axial load have been conducted by many researchers, and investigated some important parameters such as crushing distance, the peak load and the buckling shape in crashworthiness design [1-13]. figure 1 shows a schematic of compressive axial stressσx and displacement ux for a cylindrical tube subjected to axial loading. the purpose of this study is to explore the effect of impact velocity on the peak load for circular tubes, and to propose an empirical equation to estimate the peak load by numerical simulation. 2.method of analysis in this study, the dynamic numerical simulation of the impact crush test was carried out using the non-linear fe commercial code, msc.dytran. the geometry of the fe model and its boundary condition is shown in fig.2. the model is struck from the upper edge by a rigid mass m having an initial kinetic energy t0=mv0 2/2=99 kj. in the fe model, 4-node key-hoff shell elements(quad4, pshell) with three integration points are used to evaluate domain integrals, and the whole model is divided into 1440 elements. here, parameters l, t and r are the tube length, thickness and mean radius, respectively. also, the lower end of a tube is fixed to another rigid body, and a contact condition between the tube and the striker, and a self-contact condition at the inner and the outer surface of the tube are defined with the dynamic and static frictional coefficients of 0.2 and 0.3, respectively. 204 chen: prediction of maximum moment of rectangular tubes subjected to pure bending σ1 axial displacement , ux ax ia l s tre ss , σ x fig.1. schematic of axial compressive stress-displacement relationship for a tube subjected to axial impact fig.2. shell geometry and loading condition the analyzed fe model with a densityρ=2685 kg/m3 is assumed to be isotropic, and to obey the mises yield criterion with strain hardening, and a strain rate insensitive bilinear relationship between the uniaxial stress and strain as: ( ) ( ) ( )⎩ ⎨ ⎧ >−+ ≤ = eee ee yyhy y σεσεσ σεε σ (1) here, the young's modulus, e, yielding stress, σy and hardening coefficient, eh are assumed to be 72.4 gpa, 72.4 mpa and 3.62 gpa, respectively. all models in our calculation have the same tube length l=150 mm, mean radius r=25 mm and thickness t=1 mm, unless otherwise mentioned. 3.results and discussion 3.1 effect of impact velocity v0 on the initial peak stressσ1 0 0.01 0.02 0.03 0 50 100 :v0=5 (km/h) ax ia l f or ce (k n ) axial displacement (m) b a c d :v0=180 (km/h) :v0=360 (km/h) fig.3. comparison of axial force and displacement behaviour for a tube under different impact velocity v0 figure 3 shows comparisons of axial compressive force and displacement diagram for a tube under the impact velocity v0=5, 180 and 360 km/h. advances in systems science and applications (2011), vol.11, no.3-4 205 issn 1078-6236 international institute for general systems studies, inc. also, the deformed shapes at the initiation of the initial peak stress and soon after the stress for the case of v0=5 km/h and 360 km/h are shown in fig.4. it is evident from fig.4 that the initial peak stress is associated with the initiation of local buckling deformation which occurs near the upper and lower ends of a tube. from the relationship between the axial load and displacement for each v0 which is shown in fig.3, the initial peak stress can be calculated and summarized in fig.5. also in fig.5, quasi-static buckling stress for the tube obtained by implicit finite element code, msc.marc, is shown by a dashed line. it is found from this figure that the initial peak stressσ1 for a lower impact velocity is almost equal to the value of the quasi-static result, and the stress value becomes higher as the impact velocity v0 increases. fig.4. comparisons of deformed shape at points 'a', 'b', 'c' and 'd' in fig.3. 'a': buckling point for v0=5(km/h); 'b': point just after buckling for v0=5(km/h); 'c': buckling point for v0=360(km/h); 'd': point just after buckling for v0=360 (km/h) 0 100 200 300 400 0 250 500 in iti al p ea k s tre ss σ 1 (m p a) impact velocity v0 (km/h) quasi−static buckling stress fig.5. variation of peak stress σ1 with impact velocity v0 0 0.08 0.16 0 200 400 a xi al s tre ss σ x (m pa ) position x (m) impacted endv0=5 (km/h) t=6.2 (μsec) 4.16 3.16 2.16 1.16 0.16 0.08 0.05 0.03 0.02 0.01 :yield stress :quasi−static buckling stress (a) v0=5(km/h) figure 6 illustrates propagations of the stress wave in the axial direction for v0=5 km/h (fig.6(a)) and 360 km/h (fig.6(b)) until the initial peak stress occurs. also in figure 6, values of the yield stressσy and the quasi-static buckling stress are shown by dotted and dashed lines, respectively. it is evident from fig.6(a) that for the case of a lower impact velocity(v0=5 km/h), 206 chen: prediction of maximum moment of rectangular tubes subjected to pure bending the amplitude of the stress wave in axial direction at the initiation of impact is relatively small, and the stress wave travels and reflects many times along the tube until the local buckling deformation arises near the fixed end. in the end, the axial stress distribution developing over the tube becomes almost uniformly at t=0.16μs, and its amplitude increases stably to the quasi-static buckling stress. on the other hand, for the case of a higher impact velocity (fig.6(b)), the development of axial stress distribution apparently differs from that for a lower impact velocity(fig.6(a)). 0 0.08 0.16 0 400 800 a xi al s tre ss σ x (m p a) position x (m) impacted endv0=360 (km/h) t=0.14 (μsec) 0.12 0.1 0.03 0.02 0.01 :yield stress :quasi−static buckling stress (b) v0=360(km/h) fig.6. axial stress distribution until the initiation of the peak stressσ1 in fig.6(b), the stress concentration can be found at x=0.156 m. such a stress concentration occurs by the existence of the flange at the impacted end, and does not affect the overall buckling behaviour of a circular tube. it is evident from fig.6(b) that the amplitude of the stress wave bigger than the value of quasi-static buckling stress develops near the impacted end at the beginning. however, such the high stress field is relatively narrow, and no buckling behaviours seem to be observed even if the amplitude of stress wave is bigger than that of the quasi-static result. moreover, even though the axial stress developing over the tube becomes larger than the quasi-static result at t=0.12μs, no buckling can be observed. finally, the local buckling can be observed at t=0.14μs when the amplitude of the axial stress is almost twice larger than the quasi-static buckling stress. the mechanism of the buckling for circular tubes under higher impact velocities will be discussed as follows. fig.7. comparisons of deformed shape near the fixed end at the initiation of the peak stressσ1 based on the previous work concerning the quasi-static axial compressive behaviour for circular tube, it is evident that the initial peak stress is associated with the sufficient local bending deformation which is observed near the tube end. that is, during the axial compression, the tube seems to expand in radial direction, but near the tube end, such a movement is restricted by the existence of the fixed boundary condition. as a result, the local bending deformation can be observed near the tube end, and the local bending deformation is necessary for initiating the buckling behaviour. figure 7 shows comparisons of deformed shape near the fixed end when the initial peak stress is observed for some cases of impact velocity v0=5 km/h, 180 km/h and 360 advances in systems science and applications (2011), vol.11, no.3-4 207 issn 1078-6236 international institute for general systems studies, inc. km/h. it is evident that for the case of a higher impact velocity, the amount of the local bending deformation is smaller than that under a slower impact velocity. in order to initiate the buckling behaviour where the smaller amount of the local bending deformation occurs under higher impact velocity, that is, in order to increase the radial displacement, a large amount of axial stress should be needed. therefore, it could be understood that the reason why the initial peak stress increases as increase of the impact velocity v0 is due to the fact that the impact for a higher impact velocity causes a smaller radial displacement than that for a lower impact velocity. moreover, the reason why the bending deformation decreases as v0 increases can be explained by the radial inertia effect. that is, the more faster the impact velocity becomes, the more rapidly the axial stress increases, but the expansion in radial direction would be delayed by the radial inertia effect. such a mechanism can be observed in fig.8. 0 200 400 0 0.8 1.6 0 400 800 ra di al d is pl ac em en t u r ( m m ) impact velocity v0 (km/h) :ur :σx a xi al s tre ss σ x ( m p a) ur σx fig.8.variation of radial displacement ur and axial stress σx at the apex of wrinkle with impact velocity v0 0 1 2 0 200 400 600 a xi al s tre ss σ x ( m p a) radial displacement ur (mm) :peak stress :σx−ur curve v0=300 (km/h) v0=360 (km/h) v0=180 (km/h) v0=5 (km/h) fig.9. variation of axial stressσx with radial displacement ur at the apex of wrinkle for v0=5, 180, 300 and 360 (km/h) figure 8 presents the relationship between the radial displacement ur and the impact velocity v0 under almost the same amount of the axial stressσx. it is evident that the displacement ur decreases as increase of v0 even if the same amount ofσx occurs. moreover, the relationship between the axial stressσx and the radial displacement ur for four cases of v0= 5, 180, 300 and 360 km/h are chased and the initial peak stress for every v0 is summarized in fig.9 by marks ●. from the figure, the reason why the higher velocity causes the larger initial peak stress can be explained as follows. while the axial stress increases by propagating the stress waves over the tube, the radial displacement ur develops by the axial compression, but the rate of ur relatively decreases as increase of v0. as a result, the axial stress increases by the stress wave reflecting many times until the sufficient radial displacement can be reached for initiating the local buckling. in order to discuss the influence of v0 on the peak stress quantitatively, an effective parameter considering the effect of mechanical properties on the inertia effect is needed. figure 10 shows the relationship between the normalized axial stress σx/e and displacement ux/l for 208 chen: prediction of maximum moment of rectangular tubes subjected to pure bending three cases of a tube problem having different elastic modulus e, tube density ρ and the impact velocity v0. in the figure, parameter c represents the elastic wave speed for a one-dimensional rod, and can be shown as (e/ρ)1/2. here, all models have different values of e, ρ and v0, but keep the same ratio v0/c. it is evident that all models behave the same response of the axial stressσx/e and ux/l diagram. strictly speaking, the elastic and the plastic wave speed for a circular tube are distinct from the elastic wave speed c for a one-dimensional rod. however, based on the fact that the relationship between the normalized axial stressσx/e and displacement ux/l is the same under the same v0/c, the non-dimensional parameter v0/c can be used for evaluating the effect of impact velocity on the initial peak stressσ1. fig.10. normalized axial compressive stress and displacement behaviour for tubes having the same ratio of v0/c 3.2 effects of tube geometries and material properties on the initial peak stress as the other effective parameters for the dynamic initial peak stressσ1, tube geometries (such as mean radius, r and thickness, t) and material properties (such as young's modulus, e, hardening coefficient, eh and yield stress,σy) can be given, and the effects of these parameters onσ1 are discussed as follows. 0 0.1 0.2 0 100 200 300 σ x (m pa ) ux / l :t=0.5(mm) , r=12.5(mm) v0=40 (km/h) l=160(mm) :t=1.0(mm) , r=25.0(mm) :t=1.5(mm) , r=37.5(mm) fig.11. comparison of the compressive stress and displacement behaviour for tubes having the same ratio of tube thickness to radius t/r figure 11 summarizes effects of mean radius r and thickness t on the axial stress and displacement behaviour for tubes having different tube radius r and thickness t, but the same ratio of t/r. it is evident from fig.11 that if the ratio of t/r has the same value, the initial peak stressσ1 has the same value. advances in systems science and applications (2011), vol.11, no.3-4 209 issn 1078-6236 international institute for general systems studies, inc. 0 0.1 0.2 0 0.002 0.004 σ x / e ux / l e ρ v0=40 (km/h) σy=e/1000 e0 2e0 ρ0 2ρ0 c e0=72.4 (gpa) ρ0=2685 (kg/m3) eh e0/20 e0/10 e0 /ρ0 fig.12. normalized axial compressive stressσx/e and displacement ux/l behaviour for tubes having the same ratios of eh/e and v0/c figure 12 shows the comparison of the relationship between two tubes having the different e and eh but the same ratio of eh/e. in the figure, the stress value is normalized by e. it can be observed from fig.12 that the normalized initial peak stressσ1/e can be arranged as a function of the ratio, eh/e. moreover, the yield stressσy for a tube also affects the initial peak stress, but if the valueσy is quite smaller than the peak stressσ1, the effect seems to be negligible. from the above results and discussions, the normalized initial peak stressσ1/e for a tube obeying a bilinear stress and strain relationship can be expressed by a function composed of four parameters, v0/c, t/r, eh/e andσy /e as: ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ = ee e r t c vf e yh σσ ,,,0 1 1 (2) figure 13 shows the relationship between the normalized impact velocity v0/c and the peak stressσ1/e for some combinations of ratios, t/r, eh/e andσ1/e. it is clear from the figure that the stressσ1/e increases linearly with the parameter (v0/c)2. also, the slope and intersect of the relationship depend on these ratios, eh/e and t/r. 0 20 40 0 0.004 0.008 0.012 σ 1 / e (v0 / c)2 10−5 t/r σy=e/1000 σy=3e/1000 eh/e 0.01 0.05 0.04 0.08 0.01 0.05 fig.13. normalized impact velocity v0/c and the peak stressσ1/e for some combinations of t/r, eh/e andσy/e based on these charateristics, the normalized initial peak stressσ1/e can be written by the following type of equation as: 2 2 0 1 1 c c vc e +⎟ ⎠ ⎞ ⎜ ⎝ ⎛= σ (3) where, 210 chen: prediction of maximum moment of rectangular tubes subjected to pure bending .,,,,,,, 3221 ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ =⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ = ee e r tfc ee e r tfc yhyh σσ equation (3) means that the initial stress can be expressed by the sum of the inertial term including the parameter v0/c and the quasi-static term for v0=0. here, the quasi-static term c2 has already been discussed in the previous study[11], and proposed an approximate equation as follows: ( ) ( ) ( ) ( ) ( ) . 1 2 13 0 17.0 02 1 2 ⎪ ⎪ ⎩ ⎪⎪ ⎨ ⎧ >−⎟ ⎠ ⎞ ⎜ ⎝ ⎛⋅+ ≤ − = = − xrte e e r t e xrt r t e c k ee hy static hσ ν σ (4) where, ( ) ⎟ ⎠ ⎞ ⎜ ⎝ ⎛ − − = − = 0 2 0 22, 13 x r t e eek e x y hy σ νσ figure 14 shows comparisons of the parameter c2 obtained by fem and its approximation obtained by eq.(4). in the figure, results forσy/e=1/1000 and 3/1000 correspond to solid and dotted lines, respectively. it is clear that both results coincide with each other, and the approximate equation which is shown as eq.(4) can be used to estimate the quasi-static peak stress with a good accuracy. 0 0.05 0.1 0 0.005 0.01 0.015 c 2 t / r σy/e=0.001 σy/e=0.003 eh/e 0.01 0.05 0.1 line :equation (4) :σy/e=0.001 :σy/e=0.003 fig.14. comparisons of quasi-static term c2 in eq.(3) for tubes having some combinations of t/r,σy/e and eh/e in figure 15, results of the parameter c1 for some ratios, eh/e,σy/e and t/r are summarized. it is found from the figure that the parameter c1 increases with increase of eh/e, and its value becomes larger as the ratio of tube thickness to radius t/r decreases. also, the value c1 depends on the yield stress σy. for example, c1 for t/r=0.02 and 0.08 under the same ratios, eh/e=0.1 andσy/e=0.001 are almost equal to 13.0 and 8.0, respectively, so that the difference between them is almost 5.0. however, by concerning the term (v0/c)2, for example, if the impact velocity v0 is 100km/h, the order of the parameter (v0/c)2 is about 10-5, and the scale of the difference is one digit smaller than the parameter c2. based on the fact, the parameter c1 can be written by the following equation as: 7.0 1 50 ⎟ ⎠ ⎞ ⎜ ⎝ ⎛= e ec h (5) and shown in fig.15 as a solid line. advances in systems science and applications (2011), vol.11, no.3-4 211 issn 1078-6236 international institute for general systems studies, inc. 0 0.04 0.08 0.12 0 5 10 15 c 1 eh / e t/r 0.04 :equation (5) σy/e=0.001 0.08 σy/e=0.003 fig.15. comparisons of quasi-static term c1 in eq.(3) for tubes having some combinations of t/r,σy/e and eh/e consequently, an approximate equation for the non-dimensional initial peak stress for a cylindrical tube subjected to axial impact load is proposed in this paper as follows: static h ec v e e e 1 2 0 7.0 1 50 σσ +⎟ ⎠ ⎞ ⎜ ⎝ ⎛ ⎟ ⎠ ⎞ ⎜ ⎝ ⎛= (6) 0 200 400 0 500 1000 σ 1 (m p a) v0 (km/h) :eh/e=0.01 :eh/e=0.05 :eh/e=0.1 t/r=0.04 :equation (6) (a) for the case of t/r=0.04 0 200 400 0 500 1000 1500 σ 1 (m pa ) v0 (km/h) :eh/e=0.01 :eh/e=0.05 :eh/e=0.1 t/r=0.08 :equation (6) (b) for the case of t/r=0.08 fig.16. estimation of the peak stress σ1 under some cases of impact velocity v0 figure 16 shows the relationship between the initial peak stress and the impact velocity for some cases of eh/e for the ratio of t/r=0.04 (fig.16(a)) and 0.08 (fig.16(b)). in these figures, solid lines correspond to the approximate results obtained by eq.(6), and the symbols show the numerical results obtained by fem. it is clear from these figures that the predictedσ1/e agrees well with the numerical results for a wide range of impact velocity v0. 4. validation of the proposed prediction for σ1 212 chen: prediction of maximum moment of rectangular tubes subjected to pure bending in order to check the validity of the proposed prediction forσ1 as shown in eq.(6), the same impact problem studied by karagiozova and jones[5] is examined and compared the prediction with their results in fig. 17. in the figure, the solid line shows the approximation by eq.(6), and dotted line and solid circles are the predictions and fem results which can be found in karagiozova and jones' paper[5]. here, the model is intended for a circular tube made of aluminium alloy, and the material and geometrical parameters of the model are shown in fig. 17. it is evident that for relatively low impact velocity (v0<40m/s), the proposed approximation gives a good agreement with numerical results obtained by karagiozova and jones[5], which means that the proposed approximation can be effectively used for estimating the initial peak stressσ1 under a relatively low impact velocity. 0 40 80 200 400 600 v0 (m/sec) σ 1 (m pa ) : fem : prediction karagiozova and jones(2001) : eq.(6) e =72.4(gpa), eh =542.6(mpa), σy=295(mpa) ρ =2685(kg/m3), t =1.65(mm), r =11.875(mm) fig.17. comparison of estimating the peak stressσ1 under some cases of impact velocity v0 5. conclusion in this paper, large displacement numerical simulation based on fem is undertaken to explore the relationship between the impact velocity v0 and the initial peak stressσ1 for circular tubes when subjected to an axial impact. based on our numerical results, the following points have been revealed. (1) the initial peak stressσ1 becomes higher with increases of the impact velocity v0. that is because the local displacement ur in radial direction decreases as v0 increases, under the same axial stressσx. (2) the initial peak stressσ1 can be expressed by the sum of a term including v0/c and a term obtained by quasi-static numerical simulation. (3) the proposed approximate equation forσ1 can be effectively used under low impact velocity (v0<40m/s). references [1] karagiozova d. and alves m., jones n., inertia effects in axisymmetrically deformed cylindrical shells under axial impact. int. j. impact engng. 24(10)(2000) 1083-1115. [2] karagiozova d. and alves m., transition from progressive buckling to global bending of circular shells under axial impact – part i: experimental and numerical observations. int. j. solids struct. 41(5-6)(2004) 1565-1580. [3] karagiozova d. and alves m., transition from progressive buckling to global bending of circular shells under axial impact – part ii: theoretical analysis. int. j. solids struct. 41(5-6)(2004) 1581-1604. [4] karagiozova d. and jones n., dynamic effects on buckling and energy absorption of cylindrical shells under axial impact. thin-walled struct. 39(7)(2001) 583-610. [5] karagiozova d. and jones, n., influence of stress waves on the dynamic progressive and dynamic plastic buckling of cylindrical shells. int. j. solids struct. 38(38-39)( 2001) 6723-6749. advances in systems science and applications (2011), vol.11, no.3-4 213 issn 1078-6236 international institute for general systems studies, inc. [6] karagiozova d. and jones, n., on dynamic buckling phenomena in axially loaded elastic-plastic cylindrical shells. int. j. of non-linear mech. 37(7)(2002) 1223-1238. [7] karagiozova d. and jones, n., on the mechanics of the global bending collapse of circular tubes under dynamic axial load dynamic buckling transition. int. j. impact engng. 35(5)(2008) 397-424. [8] karagiozova d., nurick g.n. and yuen s.c.k., energy absorption of aluminium alloy circular and square tubes under an axial explosive load. thin-walled struct. 43(6)(2005) 956-982. [9] lu g. and mao r., a study of the plastic buckling of axially compressed cylindrical shells with a thick-shell theory. int. j. mech. sci. 43(10)(2001) 2319-2330. [10] rusinek a., zaera r., forquin p., and klepazko, j.r., effect of plastic deformation and boundary conditions combined with elastic wave propagation on the collapse site of a crash box. thin-walled struct. 46(10)(2008) 1143-1163. [11] ushijima k., haruyama s. and chen d-h. evaluation of first peak stress in axial collapse of circular cylindrical shell. trans. jsme. series a 70(700)(2004) 1695-1702. [12] wei z.g., yu j.l. and batra r.c. dynamic buckling of thin cylindrical shells under axial impact. int. j. impact engng. 32(1-4)(2005) 575-592. zhao h. and abdennadher s. on the strength enhancement under impact loading of square tubes made from rate insensitive metals. int. j. solids struct. 41(24-25)(2004) 6677-669. advances in systems science and applications (2013) vol.13 no.3 249-275 agent-based models simulations for high frequency trading d. dezsi, e.scarlat and i. măries department of economic cybernetics, bucharest university of economics, romania abstract high frequency computer-based trading (hft) represents a challenging topic nowadays, mainly due to the controversy it creates among investors on the financial market. the hereto paper compares two types of agent-based models, one with zero-intelligence traders and the other with intelligent traders in order to simulate the tick-by-tick high frequency trades on the stock market for the selected u.s. stocks. the simulations of the agent-based models are done with the help of adaptive modeler software application which uses the interaction of 2,000 heterogeneous agents to create a virtual stock market for the selected stock with the scope of forecasting the price. within the intelligent agent-based model the population of agents is continuously adapting and evolving by using genetic programming in forming new agents by using the trading strategies of the best performing agents and replacing the worst performing agents in a process called breeding, while the zero-intelligence agent based model does not evolve, agents do not breed, and they trade in a random manner. after comparing the fitting of the two models with the real data, the results show that in almost all the cases the intelligent agent-based model performed better when compared to the zero-intelligence agent-based model, which could be interpreted as lower market efficiency, allowing for predictions of the stock market price, or even stock market manipulation. also, the zero-intelligent agent-based model generates more trades and lower wealth for the population, compared to the intelligent agentbased model. the high-frequency data turns out to be very hard to simulate and analyse due to its particularities which differentiate them from daily data, as price changes are discrete, being multiples of the minimum price increment, the price changes not being independent. keywords high frequency trading, agent-based modeling, zero-intelligence traders, double auction, financial market. 1 introduction the changes of the stock market structure due to the technological improvement which led to the high speed computer-based trading have switched the market from an investor-focused mechanism to a trader-focused mechanism, where the investors’ trust and concerns are ignored. according to a study conducted by beddington in 2012 on the impact of the computer trading on the financial markets based on which the report entitled the future of computer trading in d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 250 financial markets-an international perspective [1] was released, although the investors are worried regarding the possible abuse of the market generated by the high-frequency trading (hft) and claims on market manipulation using hft techniques are reported by institutional investors all over the world, it seems like so far economic research has provided no direct evidence that hft has increased market abuse, but the authors of the report also underlined that the research regarding the measurement of market abuse is at an early stage and incomplete, not being the main focus of the report. furthermore, brogaard (2010) [2] examines the impact of hft on the u.s. equity markets and the results show that hft activities are not detrimental to non-hft activities and that hft tends to improve market quality. despite the lack of research on this subject, and even though the abuse is present or not in the market, regulators and policy makers should take in consideration the perception of the investors which show worries in this regard, because this is what determines trading behavior, portfolio management and investment decisions. thus regulators should increase their ability to detect abuse and obtain significant empirical proof on either to confirm or deny the manipulation related to hft in order to restore market confidence. market abuse through hft is studied by the economists due to its implications on market liquidity, volume, pricing efficiency and even social welfare. according to the studies conducted by cumming et al. (2012) [3] on the possible risks of market abuse generated by hft, the authors find that hft lead to lower incidence of manipulation, as also another study conducted by aitken et al. (2012) [4] reports that hft improves market efficiency without harming the market integrity. the papers mentioned above analyze market abuse at the end of the trading day, while empirical studies for the hft market abuse during the continuous trading period have not been conducted yet. these results are in contrast with the surveys regarding the perception of the investors over the degree of manipulation in the market generated by hft, which were conducted on the stock markets all over the world, and which underline the considerable concerns showed by the large investors regarding the market abuse and the lack of action and detection of manipulation by the regulators. the statistical modeling facts, models and challenges for high frequency financial data have been studied by cont [5] who outlined the empirical characteristics of high frequency financial time series, providing an overview of stochastic models for the continuous-time dynamics of the limit order book described as queuing systems, pointing that the gap between microstructure models and stochastic models should be filled in by inputting theoretical issues from microstructure models to design new stochastic models with a better economic interpretation. furthermore, ponta et al. [6] propose a non-homogeneous normal compound 251 advances in systems science and applications (2013) vol.13 no.3 poisson process for describing non-stationary returns for high-frequency financial time series for the italian stock exchange, also testing if the model can reproduce some stylized facts of high-frequency financial time series. in order to model the high-frequency foreign exchange market, aloud et al. [7] constructed an agentbased model which is able to reproduce the stylized facts of the trading activity on the foreign exchange market, by using zero-intelligence directional-change event trading strategy. according to the results obtained by li and krause (2009) [8] by comparing the market structures with near-zero-intelligence traders with the use of agentbased model simulations, they have observed that the properties of returns arising in double auction markets are not very sensitive to the trading rules employed. also, sunder (2004) [9] underlines that artificial and computer intelligence are very important in understanding the distinction between individual behaviour and market outcomes, outlining that computer simulations helped to discover that allocative efficiency is largely independent of variations in individual behaviour under classical conditions, while herbert simon was convinced that “the possibility of building a mathematical theory of a system or of simulating that system does not depend on having an adequate microtheory of the natural laws that govern the system components. such a microtheory might indeed be simply irrelevant” [9]. the aim of our research is to identify the main distinctions between the highfrequency financial data simulation results of two agent-based models, one with intelligent agents and the other with zero-intelligence agents. in order to achieve our research aim, we use the adaptive modeler software to simulate the two types of agent-based market models-the zero-intelligence agent-based model and the intelligent agent-based model-for price forecasting of real world market-traded securities such as stocks. thus, heterogeneous agents trade a stock floated on the stock exchange market, placing orders depending on their budget constraints and trading rules, where the virtual market is simulated as a double auction market. for the intelligent agent-based model, the software uses evolutionary computing such as the strongly typed genetic programming [10] in order to create adaptive, evolving and self-learning market modeling and forecasting solutions, while for the zero-intelligence agent-based model the agents do not breed nor evolve, as they trade in a random manner. the results obtained from the simulations are compared revealing that the high frequency tick-by-tick data is best modeled by the intelligent agent-based model, although the high-frequency data turns out to be very hard to simulate and analyse due to its particularities which differentiate them from daily data, such as discrete price changes, which are multiples of the minimum price increment, thus the price changes are not independent. to the best of our knowledge, none of the works in the literature have attemptd. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 252 ed to compare the results obtained from simulating zero-intelligence agent-based models with the results obtained from simulating intelligent agent-based models for high frequency financial data. the organization of this paper is as follows. section 2 provides an overview of the high frequency financial data and its statistical properties, section 3 describes the datasets used in this study, section 4 presents the agent-based models used in the simulations, while the results of the simulations that have been performed are presented in section 5, the paper ending with the conclusions and avenues for future work. 2 high frequency financial data and their statistical properties the high frequency financial data or high frequency trading (hft) represents a subset of algorithmic trading or a trading strategy where a large number of small size orders are sent into the market at high speed, therefore the securities are bought and sold by a computer algorithm and held for a very short period, usually seconds or milliseconds. the automated trading applies on the order-driven electronic platforms which aggregate all outstanding orders in an order book and market orders are executed in a mechanical manner, usually a double auction mechanism on the stock market. for a better comprehension, table 1 illustrates the hierarchy of time scales for financial data, as pointed out by cont in [5], while thousand of buying or selling orders are submitted in a 10-second interval. the hft is one of the most significant market structure developments in table 1 financial data series time scales regime time scale issues ultra-high frequency (uhf) 10−3 − 0.1s microstructure high frequency (hf) 1-100s trade execution daily 103 − 104s trading strategies recent years, as stated by the securities and exchange commission (sec), but is also controversial due to its divers impact on investors’ perception, which are worried about the degree of manipulation in the market generated by hft, according to surveys conducted on the stock markets. despite the investors worries described above, research studies on the impact of hft over the market show that hft tends to improve market quality [2], while vuorenmaa (2012) [11] reached the conclusion that the benefits of the hft have the most weight and dominate the negative aspects of hft. recently, eu parliament voted in favour of mifid ii, a draft law for strengthening the regulation of hft and other financial instruments. the new rules will apply starting with 2014 and aims to reduce speculation without harming 253 advances in systems science and applications (2013) vol.13 no.3 the real economy, by introducing a synchronised clock for trading shares, bonds, commodities and other instruments across the eu so regulators can spot abuses more easily in a market where many exchanges and platforms trade the same shares, and by extending the minimum time in which high-frequency orders have to remain in the market from 3 milliseconds currently to at least 500 milliseconds. that way, purely speculative business with high-frequency transactions will be discouraged. recall that high-frequency and algorithmic trading impact on the financial market has raised worries from behalf of the regulators and the investors since may 6, 2010, when about usd 862 billion was erased from stock values in 20 minutes before share prices recovered from the plunge, while on august 1, 2012 a trading malfunction at knight capital group which poured in the market unintended orders lead to high traded volume and high price volatility, which finally led to a usd 440 million trading loss for the company. according to a study conducted by beddington in 2012 [1] on the impact of the computer trading on the financial markets based on which the report entitled the future of computer trading in financial markets-an international perspective was released, although the investors are very worried regarding the possible abuse of the market generated by the hft, so far economic research has provided no direct evidence that hft has increased market abuse. furthermore, according to the studies conducted by cumming et al. (2012) [3] on the possible risks of market abuse generated by hft, the authors find that hft lead to lower incidence of manipulation, while also another study conducted by aitken et al. (2012) [4] reports that hft improves market efficiency without harming the market integrity. the papers mentioned above analyze market abuse at the end of the trading day, while empirical studies for the hft market abuse during the continuous trading period have not been conducted yet. on the other hand, the surveys conducted regarding the perception of the investors over the degree of manipulation in the market generated by hft underline the considerable concerns showed by the large investors regarding the market abuse and the lack of action and detection of manipulation by the regulators. also, hft allows the program trader to “see” ahead of others the major incoming orders and to act benefitting from its speed advantage, in order to gain profits, as hft has been shown to have a potential sharpe ratio, which is a measure of reward per unit of risk, thousands of times higher than the classical buy-and-hold strategies [12]. a major disadvantage is that it can generate trade crashes, such as the one on may 6, 2010, and the one on august 1, 2012, thus eroding the trust of investors in the market. as regards to statistical properties of the tick-by-tick high frequency financial data compared to daily frequency data, the main differences are the fact that d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 254 price changes are discrete, being multiples of the minimum price increment as it can be observed in the histograms from fig.1 below. (a) (b) (c) (d) (e) (f) (g) (h) (i) (j) fig.1 histogram of log returns for adbe(a), aa(b), aapl(c), bac(d), c(e), xom(f), ge(g), ibm(h), pfe(i), wfc(j) the autocorrelation function of high frequency price returns is significantly negative at the first lag and for higher lags it rapidly decreases to zero, meaning 255 advances in systems science and applications (2013) vol.13 no.3 that price changes are not independent [13]. another particularity is that trades occur at irregular intervals, they are random and endogenous, linked to the behaviour of the price and to previous trading history, and also seasonality effects is observed, as the trading activity is more intense at the market opening and closing [5]. 3 data the high frequency financial data in this paper was retrieved from the website http://www.finam.ru/analysis/profile041ca00007/default.asp and contains tickby-tick data from bats stock exchange. the tick-by-tick data was retrieved from this website, for the trades that occurred on the bats (better alternative trading system) stock exchange during regular trading hours on december 3rd, 4th, 5th and 6th of 2012 limited to 20,000 quotes for each stock, and includes the date, time and closing price information for each stock. the unique dataset used in this paper, namely the tick-by-tick high-frequency financial dataset contains real trading data for the following stocks floated on the bats, a computerized u.s. stock exchange market: addobe systems inc. (symbol: adbe), alcoa inc. (symbol: aa), apple inc. (symbol: aapl), bank of america (symbol: bac), citigroup inc. (symbol: c) exxon mobil (symbol: xom), general electric (symbol: ge), ibm (symbol: ibm), pfizer inc. (symbol: pfe), wells fargo (symbol: wfc). the bats stock exchange has been chosen for the data retrieval because it is one of the leading u.s. venues for hft, covering 12-13% of all u.s. equity trading on a daily basis according to the stock exchange’s website http://www.batstrading.com/about/. the distinction between the hft and non-hft was not defined in this paper, although taking in consideration the nature of the bats stock market, we consider the trades to be mainly hfts. table 2 symbols and numbers of observations for the 10 analysed stocks stock name symbol sector no. of obs. (bars) period of time addobe systems adbe technology 20,000 3rd − 6th of december 2012 alcoa inc. aa basic materials 12,612 3rd − 6th of december 2012 apple inc. aapl technology 19,999 3rd − 5th of december 2012 bank of america bac financial 20,000 3rd − 4th of december 2012 citigroup inc. c financial 20,000 3rd − 4th of december 2012 exxon mobil xom basic materials 20,000 3rd − 6th of december 2012 general electric ge industrial goods 20,000 3rd − 5th of december 2012 ibm ibm technology 6,144 3rd − 6th of december 2012 pfizer inc. pfe healthcare 20,000 3rd − 5th of december 2012 wells fargo wfc financial 20,0000 3rd − 5th of december 2012 the time period of trades is expressed in milliseconds, the number of observations and period of time for each stock is summarized in table 2, while the descriptive statistics for the tick-by-tick log returns are presented in table 3, in d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 256 which the excess kurtosis can be observed. also, fig.1 shows the histogram for the log returns of the analysed stocks, in order to show how the returns are distributed, being possible to see the discrete character of returns even after the logarithmic transformation. table 3 descriptive statistics for the tick-by-tick log returns stock name symbol mean×106 std.dev.×104 skewness kurtosis addobe systems adbe 0.486 1.46 -0.360791 43.17925 alcoa inc. aa 1.120 2.64 0.086025 19.32359 apple inc. aapl -3.700 1.90 -16.53734 979.9840 bank of america bac -5.060 1.37 -0.334374 23.19635 citigroup inc. c -0.909 1.10 -0.400060 32.14271 exxon mobil xom -0.443 1.10 -1.226893 41.91198 general electric ge 0.0472 1.24 0.436346 18.27284 ibm ibm -0.890 1.94 0.449297 32.84855 pfizer inc. pfe 1.230 1.05 0.267775 16.65013 wells fargo wfc -0.653 1.21 3.666686 156.9951 4 the agent-based models specifications an agent-based model is a computational model for simulating the actions and interactions of multiple agents in order to analyze the effects on a complex system as a whole, and represents a powerful tool in the understanding of markets and trading behavior. an agent-based model of a financial market consists of a population of agents (representing investors) and a price discovery and clearing mechanism (representing a virtual market). as regards to the financial markets, agent-based models can successfully replicate time series features like fat-tailed distributions and volatility clustering, on which standard financial models offer few explanations. conventionally, financial markets have been studied using analytical mathematics based on a generalization of market participants and other simplifications and idealizations. however, the behavior of financial markets as observed in reality can’t be fully described by such mathematical models. in reality, market prices are established by a large diversity of investors with different decision making methods and different investment goals. the complex dynamics of these heterogeneous investors and the resulting price formation process require a simulation model of multiple heterogeneous agents and a virtual market [14]. research has shown that complex behavior can emerge from simulations of agents with relatively simple decision rules. furthermore, commonly observed stylized facts of financial time series (price and order flow) have been reproduced by several agent-based market models [7,15-17]. although agent-based research has been introduced over 30 years ago in economics, it is still considered a niche field because of its detailed and complex 257 advances in systems science and applications (2013) vol.13 no.3 platforms. another impeachment in using computational agent-based modeling is represented by the fact that it uses source codes which usually are not disseminated along with the paper, thus being hard to be evaluated, tested and developed for further research by the rest of the community, as underlined by barr et all (2008) [12]. furthermore, agent-based modeling can generate important facts when used for institutional design. at a large scale they are already used for computer simulation by government agencies, an important example being the traffic simulations, while as a small to intermediate size model it can be used by the government agencies as an advisory and prediction instrument. fig.2 agent-based model cycle the agent-based models referred to in this paper are simulated in adaptive modeler software, which supports up to 2,000 agents and 20,000 observations for each simulation. to the best of our knowledge, none of the work in the literature has attempted to compare the simulation results generated from the simulations of zero-intelligence agent-based models with the results obtained from simulating intelligent agent-based models for high frequency financial data. each model consists of a population of agents and a virtual market on which the agents trade the envisaged security. the agents are autonomous entities representing the traders of the stock market, each having their own wealth (cash and shares) and their own trading strategy. the general cycle of the agent-based models used in this paper, which repeats at each of the 20,000 analyzed quotes, is illustrated in fig.2 and fig.3 in more details. it can be described as follows: 1. import real stock market data: real stock market data is imported for the simulated stock. 2. initialization of the model: the cycle starts after the model initialization and after the parameters are set. at the initialization, the agents receive an d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 258 initial wealth of 100,000 and no shares. each agent receives a trading rule which is called the genome, which is randomly created by taking in account the selected genes (which represent functions) using genetic programming described in [10]. fig.3 agent-based model flowchart there are no broker fees in these simulations, and no market maker. all the parameters for each of the two models and their values are described in table 4. 259 advances in systems science and applications (2013) vol.13 no.3 table 4 general settings of the models. market and agents’ parameters configuration in the simulations model parameter type parameter name parameter value no. of trading periods (bars) max. 20,000; see table 1. market no. of agents 2,000 parameters spread 20.005% variable broker fee 0% zi agentwealth distribution equal for all agents: 100,000 based position distribution equal for all agents% model : initial position 0% min. position unit 20% agent min. initial genome depth 2 parameters max. initial genome depth 5 genes rndpos, rndlimit, advice breeding cycle frequency 1,000,000 bars (so that the agents never breed, don’t evolve) no. of trading periods max. 20,000; see table 1. market no. of agents 2,000 parameters spread 20.005% variable broker fee 0% wealth distribution equal for all agents: 100,000 position distribution equal for all agents% : initial position 0 min. position unit 20% max. genome size 1000 max. genome depth 20 min. initial genome depth 2 max. initial genome depth 5 intelligent genes curpos, levunit, rmarket, agentvmarket, long, short, cash, bar, based invpos, rndpos, close, bid, ask, model agent average, min, max, volume, >, parameters change, +, dir, isupbar, upbars, pos, breeding cycle frequency 1 bar (so that the agents breed at each new bar, conditioned of the fact that they are of minimum breeding age) minimum breeding age 80 initial selection: 100% of agents of minimum randomly select breeding age or older parent selection 5% agents of initial selection will breed mutation probability 10% per offspring 3. receive new quote bar: a new quote bar from the imported real stock market data series is received by the virtual market. 4. agents evaluate trading rules and place orders: agents get access to historical prices and evaluate their trading rules according to the genomes allocated in the initialization process, resulting in a desired position as a percentage of wealth limited by the budget constraints, and a limit price. for the zero-intelligence agent-based model, a desired position and a limit price order are generated in a random manner. for the intelligent agent-based model, the position is also gend. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 260 erated in a random manner, while the limit price is generated after a technical analysis has been performed, according to the genome structure which represent trading functions. 5. virtual market clearing and forecast generation: the virtual market determines the clearing price, executes all executable orders, and forecasts the price for the next bar, using a double auction mechanism. 6. breeding (only for the intelligent agent-based model): during the breeding process, new agents are created from best performing agents in order to replace the worst performing agents, creating new genomes by recombining the parent genomes through a crossover operation. the breeding process repeats at each bar, with the condition that the agents must have a minimum breeding age of 80 bars, in order to be able to assess the agents performance. 7. the model waits for a new quote: if the model receives new quotes, it will repeat the process described at points 4-6. if there are no more quotes to be processed, the simulation ends. in order to obtain random seed, the adaptive modeler software uses the mersenne twister algorithm [18] to generate pseudo random number sequences for the initial creation of trading rules or genomes and for the crossover and mutation operators of the breeding process. more information regarding the adaptive modeler software and how it works can be found at http://altreva.com/adaptive modeler users guide.htm. 4.1 zero-intelligence traders model the concept of zero intelligence (zi) traders has been introduced by gode and sunder (1993) [19] in order to study the lower limit of rationality required to participate in a double auction market, which implies that agents have no strategy and they behave in a random manner subject to a budget constraint, so that the market mechanism can be observed. in case the zi traders model generate good results, this means that the market mechanism used is incentive-compatible and most probably is not influenced by the trading strategies used by the investors. this represents a very important result for the market mechanism design, as a market mechanism which performs well despite the trader’s irrationality is preferred over the mechanism that performs well only with perfectly rational traders, as underlined in walia (2003) [20]. in ladley’s (2012) [21] point of view, any effects that are not observed in the zi model cannot solely be due to the market mechanism and requires the interaction of the investors strategies, therefore it can be possible to separate the effects of the market mechanism and trader strategy and to determine the driving forces within markets. in this context, a model that assumes stock market traders have zero intelligence has been found to mimic the behavior of the london stock exchange very closely, the research being conducted by farmer et al. (2004) [22] at 261 advances in systems science and applications (2013) vol.13 no.3 the santa fe institute, which suggest that the movement of the markets depend less on the strategic behavior of the traders and more on the design of the trading system. a double auction is a trading mechanism in which buyers and sellers can enter bid or ask limit orders and accept asks or bids entered by other traders [23]. this trading mechanism was chosen to be used for the virtual market simulation in the adaptive modeler models because most of the stock markets are organized as double auctions, which is considered to be a continuous-time game of incomplete information [24]. double auction stock markets converge to the equilibrium derived by assuming that the traders are profit-maximizing bayesian, being an example of a microeconomic system, as described in hurwicz (1986) [25] and smith (1982) [26]. in the double auction markets, agents introduce bid or ask orders, each order consisting of a price and quantity. the bids and asks orders received are put in the order book and an attempt is made to match them. the price of the trades arranged must lie in the bid-ask spread (interval between bid price and ask price). furthermore, the use of double auction trading mechanism generates stochastic waiting times between two trades, as stated by scalas [27]. an example of the order book resulted from buy and sell orders collected by the virtual market can be viewed in the fig.4, 5, 6 and 7, where blue bars represent buy orders, red bars represent sell orders and yellow bars represent buy and sell orders at the same price before the market clearing. fig.4 and 6 illustrate the order book before the clearing, while fig.5 and 7 illustrate the order book after clearing, so only the unexecuted buy (blue) and sell (red) orders remain visible in these two figures. the clearing price can be observed in the top of the figures. in the hereto paper, the zi traders model is designed within the context of the double auction mechanism, which sets the order placement, clearing price, the supply and demand. the model is implemented in the adaptive modeler software and will be used in this paper to simulate tick-by-tick high-frequency data. the adaptive modeler software in the zi agent-based model uses only three types of genes or trading strategies, as mentioned in table 4, namely rndpos, rndlim and advice. the rndpos gene returns a position value ranging from -100% to 100% which is randomly generated from a uniform distribution, the rndlim gene returns a random limit price that is generated as follows: the last closing price is taken from either the virtual market or the real market which is chosen randomly, then the price is multiplied with a normally distributed random value with µ = 1 and σ = 3.5 · σm where σm is the standard deviation of the log returns of the last 20 bars of the chosen market. the advice gene combines the position generated by rndpos function and the limit price value generated by the rndlim function into a buy or sell order advice. the buy or sell order is introduced in the market after comparing the desired position with the agent’s d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 262 fig.4 order book example from the zi-abm simulation for aapl stock. blue bars represent buy volume, red bars represent sell volume, yellow bars represent buy and sell volume at the same price before clearing; clearing price can be seen above the chart fig.5 order book example from the zi-abm simulation for aapl stock. blue bars represent buy volume, red bars represent sell volume after clearing; clearing price can be seen above the chart 263 advances in systems science and applications (2013) vol.13 no.3 fig.6 order book example from the zi-abm simulation for aapl stock. blue bars represent cumulative buy volume, red bars represent cumulative sell volume, yellow bars represent buy and sell volume at the same price before clearing; clearing price can be seen above the chart fig.7 order book example from the zi-abm simulation for aapl stock. blue bars represent cumulative buy volume, red bars represent cumulative sell volume after clearing; clearing price can be seen above the chart d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 264 current position and calculating the number of shares that need to be bought or sold, taking also in consideration the available cash. this order generation process is explained in detail in the paper agent-based simulation of a financial market by raberto (2001) [17], together with the ability of this model to exhibit the stylized facts of financial time series, such as fat tails and volatility clustering, using simple trading rules, budget restriction of the agents, order limit prices, and the creation and matching of demand and supply curves. the zero-intelligence agent-based model always skips the breeding step by setting the breeding cycle frequency parameter at a very high value (1,000,000) as mentioned in table 4, therefore agents are not replaced with better ones, the population does not evolve, nor adapt to the market conditions, agents trading in a random manner. 4.2 the intelligent agent-based model the intelligent agent-based model cycle starts by receiving a new quote bar, so that agents can place a new order or remain inactive according to their trading strategy. after all agents have evaluated their trading strategy, the virtual market determines the clearing price, executes all executable orders and releases the price forecast for the next bar. after that, breeding of new agents and replacement by evolutionary operations such as crossover and mutation can take place, a process which repeats itself for each bar. the trading rules of the model uses historical price data as input, either from the virtual market either from the real market, and return an advice consisting of a desired position, as a percentage of wealth, and an order limit price for buying or selling the security. the trading rules are implemented by genetic programming technology explained bellow. through evolution the trading rules are set to use the input data and functions (trading strategies) that have the most predictive value. the agents’ trading rules development is implemented in the software by using the strongly typed genetic programming (stgp) approach, and use the input data and functions that have the most predictive value in order for the agents with poor performance to be replaced by new agents whose trading rules are created by recombining and mutating the trading rules of the agents with good performance. the stgp was introduced by montana (2002) [10], with the scope of improving the genetic programming technique by introducing data types constraints for all the procedures, functions, variables and constants, thus decreasing the search time and improving the generalization performance of the solution found. therefore, the genomes (programs) represent the agents’ trading rules and they contain genes (functions), thus agents trade the security on the virtual market based on their analysis of historical quotes. in order to obtain random sequences for the initial creation of trading rules or genomes and for the crossover and mutation operators of the breeding process, the adaptive modeler software 265 advances in systems science and applications (2013) vol.13 no.3 uses the mersenne twister algorithm [18] to generate pseudo random number sequences. more information regarding the adaptive modeler software and how it works can be found at http://altreva.com/adaptive modeler users guide.htm. also, all the genes as functions are described in the guide, along with the types of the arguments used. during the breeding process which is specific only to the intelligent agent-based model, new offspring agents are created from some of the best performing agents to replace some of the worst performing agents. in order to achieve this, at every bar, agents with the highest breeding fitness return are selected as parents, and the genomes (trading rules) of pairs of these parents are then recombined through genetic crossover to create new genomes that are given to new offspring agents. these new agents replace agents with the lowest replacement fitness return. the fitness functions are a measurement of the agents investment return over a certain period, therefore the breeding fitness return is computed as a short term trailing return measure of the wealth over the last 80 analyzed quotes and represents the selection criterion for breeding (best agents), while the replacement fitness return is computed as the average return per bar and represents the selection criterion for replacement (worst agents). fig.8 oaddobe stock simulation-intelligent agent-based model 5 simulation results the results of the simulations are illustrated in fig.8-25, and were conducted over each of the ten analyzed stocks, for both types of agent-based models, with intelligent agents and zero intelligence agents. the results of the simulations mainly illustrate a better performance of the intelligent agent-based model in terms of forecast accuracy, except for the stocks addobe and citigroup. the results must be further tested for how stylized facts are reproduced with these models, compared to real stock quotes. d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 266 fig.9 addobe stock simulation-zero-intelligence agent-based model fig.10 apple stock simulation-intelligent agent-based model fig.11 apple stock simulation-zero-intelligence agent-based model 267 advances in systems science and applications (2013) vol.13 no.3 fig.12 bank of america stock simulation-intelligent agent-based model fig.13 bank of america stock simulation-zero-intelligence agent-based model fig.14 citigroup stock simulation-intelligent agent-based model d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 268 fig.15 citigroup stock simulation-zero-intelligence agent-based model fig.16 exxon mobil stock simulation-intelligent agent-based model fig.17 exxon mobil stock simulation-zero-intelligence agent-based model 269 advances in systems science and applications (2013) vol.13 no.3 fig.18 general electric stock simulation-intelligent agent-based model fig.19 general electric stock simulation-zero-intelligence agent-based model fig.20 ibm stock simulation-intelligent agent-based model d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 270 fig.21 general electric stock simulation-zero-intelligence agent-based model fig.22 pfizer stock simulation-intelligent agent-based model fig.23 pfizer stock simulation-zero-intelligence agent-based model 271 advances in systems science and applications (2013) vol.13 no.3 fig.24 wells fargo stock simulation-intelligent agent-based model fig.25 wells fargo stock simulation-zero-intelligence agent-based model it is worth mentioning that wealth is much better and uniformly distributed among agents in the case of zero-intelligence agent-based model, as it can be observed in the fig.8-25, which is of great importance for market design, while intelligent agent-based models generate discrepancies among agents’ wealth distribution. on the other hand, the wealth in the intelligent agent-based model simulations is much better preserved compared to the zero intelligence agentbased models, due to the fact that intelligent agent-based model is not a closed economy due to the replacement of the agents which bring new capital in the virtual market, while the zero-intelligence agent-based model uses the same traders during all the simulation period. this generates a divergence between the buy orders and sell orders due to lack of cash, thus buy orders. as illustrated in the average genome size window in the fig.9-25 for the zero-intelligence agent-based models the value remains constant at 3, as there are only 3 genes (functions) that generate the trades of the population, while for the intelligent agent-based model the value ranges between 30 and 80, showing how the agents strategies constantly adapt and evolve during the simulations. d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 272 in order to compare the fit of the models to the real data and to compare the efficiency of the models, the root mean squared error and mean absolute error have been computed and listed in table 5. the results show that the intelligent agent-based model performed better compared to the zero-intelligent agent-based model in terms of price forecast. table 5 root mean squared error and mean absolute error for the simulations stocks model simulated root mean squared error mean absolute error addobe (adbe) zi-abm 0.010 0.006 i-abm 0.032 0.003 alcoa (aa) zi-abm 0.004 0.003 i-abm 0.002 0.001 apple (aapl) zi-abm 0.279 0.124 i-abm 0.110 0.056 bank of america (bac) zi-abm 0.007 0.003 i-abm 0.002 0.000 citigroup inc. (c) zi-abm 0.027 0.011 i-abm 0.110 0.003 exxon mobile (xom) zi-abm 0.040 0.020 i-abm 0.010 0.005 general electric (ge) zi-abm 0.013 0.006 i-abm 0.003 0.001 ibm (ibm) zi-abm 0.062 0.041 i-abm 0.037 0.020 pfitzer inc. (pfe) zi-abm 0.035 0.011 i-abm 0.003 0.001 wells fargo (wfc) zi-abm 0.016 0.007 i-abm 0.009 0.002 6 conclusions taking into account the previous studies on the impact of high frequency trading on the financial markets, the results have shown that this algorithmic trading has led to higher liquidity, improved market efficiency without harming market integrity and lower incidence of market manipulation. although the results of the research in this field is rather optimist thus far, regarding the impact of hft over the financial market, the studies are few and do not comprise all the intra-day price impact, nor the impact of the unexecuted orders which are introduced into the market by these trading algorithms. thus, investors are worried regarding the possible manipulation of the market following these high frequency trades, and currently regulators focus more and more on a better policy for regaining the trust of the investors in the market by limiting the hft in the markets by imposing that the trading systems are configured such as not to cause market disturbances, a higher control of the market strategies is intended to be implemented, and a fee is also to be charged in the case of excessive use of the system and a limit introduced on the ratio of unexecuted orders. the results show that in almost all the cases the intelligent agent-based mod273 advances in systems science and applications (2013) vol.13 no.3 el performed better when compared to the zero-intelligence agent-based model, which could be interpreted as lower market efficiency, allowing for predictions of the stock market price, or even stock market manipulation. this has to be further studied and efficient market hypothesis should be tested. the hereto paper brings to light the importance and accuracy of the agent-based models when it comes to stock market price forecast and modeling, and more important it applies these types of models on high frequency trading data, obtaining high accuracy in forecasting. further research should focus on improving the fit of these models with the real stock market by adding brokerage fee and allowing zero-intelligence agents to breed (but not to adapt), thus renewing the population to obtain a zerointelligence agent-based model closer to reality. acknowledgements this work was co-financed from the european social fund through sectorial operational programme human resources development 2007-2013, project number posdru/107/1.5/s/77213 “ph.d. for a career in interdisciplinary economic research at the european standards”. references [1] j. beddington. (2012), “the future of computer trading in financial markets”, the govornment office for science, london. [2] j. brogaard. (2010), “high frequency trading and its impact on market quality”, social science research network. [3] d. cumming, f. zhan and m. aitken. (2012), “high frequency trading and end-of-day manipulation”, government office for science, london. [4] m. aitken, f. harris, t. mclnish and a. aspris. (2012), “high frequency trading assessing the impact on market efficiency and integrity”, government office for science, london. [5] r. cont. (2011), “statistical modeling of high frequency financial data: facts, models and challenges”, ieee signal processing magazine. [6] l. ponta, m. raberto, s. cincotti and e. scalas. (2012), “modeling nonstationarities in high-frequency financial time series”, in cord conference proceedings. d. dezsi, e.scarlat and i. măries :agent-based models simulations for high... 274 [7] m. aloud, m. fasli, e. tsang, a. dupuis and r. olsen. (2013), “modelling the high-hrequency fx market: an agent-based approach”, university of essex, technical reports, colchester. [8] x. li and a. krause. (2009), “a comparison of market structures with near-zero-intelligence traders”, intelligent data engineering and automated learning-ideal 2009, lecture notes in computer science, vol.5788, pp.703-710. [9] s. sunder. (2004), “markets as artifacts: aggregate efficiency form zerointelligence traders”, in m. e. augier and j. g. march, editors, models of a man: essays in memory of herbert a. simon, the mit press, cambridge. [10] d. montana. (2002), “strongly typed genetic programming”, evolutionary computation, vol.3, no.2, pp.199-230, . [11] t. vuorenmaa. (2012), “the good, the bad, and the ugly of automated high-frequency trading”, social science research network. [12] j. barr, t. tassier and l. ussher. (2008), “a future of agent-based models in economics: a panel discussion, eastern economic association annual meeting, boston”, eastern economic journal, vol.34, pp.550-565. [13] r. cont. (2001), “empirical properties of asset returns: stylized facts and statistical issues”, quantitative finance, vol.1, pp.223-236. [14] “www.altreva.com”, altreva. (2003), available: http://www.altreva.com /technology.html. [15] v. alfi, m. cristelli, l. pietronero and a. zaccaria. (2009), “minimal agent based model for financial markets i: origin and self-organization of stulized facts”, the european physical journal, vol.67, no.3, pp.385-397. [16] g. daniel. (2006), asynchronous simulations of a limit order book, university of manchester, ph.d. thesis. [17] m. raberto, s. cincotti, s. focardi and m. marchesi. (2001), “agent-based simulation of a financial market”, physica, vol.299, pp.319-327. [18] m. matsumoto and t. nishimura. (1998), “mersenne twister: a 623dimensionally equidistributed uniform pseudorandom number generator”, acm trans. on modeling and computer simulation, vol.8, no.1, pp.3-30. [19] d. gode and s. sunder. (1993), “lower bounds for efficiency of surplus extraction in double auctions, in the double auction market: institutions, 275 advances in systems science and applications (2013) vol.13 no.3 theories, and evidence”, in the double auction market, eds. d. friedman and j. rust, sfi studies in the science of complexity, proc.vol. xiv. [20] v. walia, a. byde and d. cliff. (2003), “evolving market design in zero-intelligence trader markets”, in ieee international conference on ecommerce (ieee-cec03), newport beach, ca., usa. [21] d. ladley. (2012), “zero intelligence in economics and finance”, the knowledge engineering review, cambridge university press, vol.27, pp.273-286. [22] d. farmer, p. patelli and i. zovko. (2004), “the predictive power of zero intelligence in financial markets”, proceedings of the national academy of sciences of the united states of america, vol.102, no.6, pp.2254-2259. [23] d. gode and s. sunder. (1993), “allocative efficiency of markets with zerointelligence traders: market as a partial substitute for individual rationality”, journal of political economy, vol.101, pp.119-137. [24] j. rust, r. palmer and j. miller. (1989), “a double auction market for computerized traders”, santa fe institute, new mexico. [25] l. hurwitz. (1986), “on informationally decentralized systems”, in decision and organization, minnesota, univ. of minnesota press, pp.297-336. [26] v. l. smith. (1982), “microeconomic systems as an experimental science”, the american economic review, vol.72, no.5, pp.923-955. [27] e. scalas. (2007), “mixtures of compound poisson processes as models of tick-by-tick financial data”, chaos, solitons & fractals, vol.34, pp.33-40. corresponding author d. dezsi can be contracted at: dianadezsi@yahoo.com microsoft word 2. k. masuda and d. h. chen--prediction of maximum moment of rectangular tubes subjected to pure bending.doc 214-225 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc. prediction of maximum moment of rectangular tubes subjected to pure bending k. masuda and d. h. chen department of mechanical engineering, tokyo university of science, 1-3 kagurazaka, shinjuku-ku, tokyo, 162-8601, japan abstract in this paper, the collapse behaviors of rectangular tube subjected to pure bending are studied by using the finite element method. such bending collapse has been studied for a long time, including the landmark study by kecman. according to these studies, there are two types of collapses. the first type is a collapse due to buckling at the compression flange. the second type is a collapse due to plastic yielding at the flanges. however, there may be another collapse. for a rectangular tube in which the web is wider than the flange, it is found that collapse due to buckling at the compression web may occur. further, an approximation prediction method is proposed for estimating the maximum bending moment of rectangular tubes in which the web buckling is also taken into account. its validity is verified by comparing with the numerical results by fem under various conditions. keywords fem, pure bending, rectangular tube, buckling, effective width 1.introduction evaluation of a car’s crush behavior when the car is subjected to an oblique load, such as in an offset crash, is becoming increasingly important for general car design. it is thus vital to understand the crushing characteristic of rectangular tubes that are used as general components in car. because the oblique load can be decomposed into axial load and pure bending, the pure bending of the rectangular tubes has been widely studied for a long time [1, 2, 3], including the landmark study by kecman [1]. according to these studies, there are two types of collapses. the first type is a collapse due to buckling at the compression flange. the second type is a collapse due to plastic yielding at the flanges. however, when the web is wider than the flange, it is considered that a collapse due to buckling at the compression web may occur because it was already reported in bending of open section beams [4]. in the present study, the effects of the material and geometrical properties of rectangular tubes on their bending collapse are studied by using the finite element method. further, based on the numerical results obtained, a method for estimating the maximum bending moment of rectangular tubes subjected to pure bending is proposed. in addition, a validity of fe analysis result under bending collapse has been already verified by comparing the experimental results by kyriakides [5] with the previous numerical results by authors [6] under pure bending with cylindrical tubes. 2.analytical method the commercial fem analysis package msc. marc[7] is used in this study to analyze large elastoplastic bending of the rectangular tubes shown in fig. 1. in the present calculation, one end of the rectangular tube was completely fixed to a rigid wall. pure bending was applied from the other end by modeling a lid rotating about the axis of z under rotary control. the effects of various geometric parameters, such as tube thickness t, tube flange width c1, and tube web width c2, on the bending collapse were investigated. the value of lid thickness tf was set to five times t as referred to guarracino [8] because the lid must be stiff enough to prevent distortion. advances in systems science and applications (2011), vol.11, no.3-4 215 issn 1078-6236 international institute for general systems studies, inc. the tube material used in the analysis was assumed to be homogeneous and isotropic elastic perfectly plastic material that conforms to von mises yield conditions. in this study, it was assumed that young’s modulus e = 72.4 gpa, and poisson’s ratio v = 0.3. the influence of the material properties on the bending collapse of the rectangular tube was investigated in terms of the yield stress yσ . in this study, the updated lagrange method was used to formulate the geometric nonlinear behavior, and the algorithm based on the newton-raphson method and the return-mapping method were used to solve the nonlinear equation. the rectangular tubes were modeled using four-node quadrilateral thickness shell elements (element type 75). the elements were divided the flange and wed width into 20 sub lengths, and divided the axial length in a way that the elements become almost square. in addition, the rectangular length used in the analysis was assumed to be long enough in order to exclude the influence of the boundary conditions. the ratio of the length and flange width l/c1 was set to 1/ 6l c > . l t t x y z2c 1c z y x m output zθ input tf =5t a aa-a fig.1.tube geometry and loading condition 3.results and discussion 3.1 comparison between proposal method by kecman and results of present numerical analyses first, we show kecman’s method for estimating the maximum bending moment of rectangular tubes subjected to pure bending. for a rectangular tube subjected to pure bending, the buckling stress bucσ of the compression flange was derived in the following equation. ( ) ( ) 22 2 5.23 0.16 1 12 1buc e a t b a πσ ν ⎛ ⎞⎛ ⎞= +⎜ ⎟⎜ ⎟− ⎝ ⎠⎝ ⎠       b a 2 eayσ 1y 1 1 y b y y σ − (a) b yσ (b) yσ 2 b a fig.2. schematic representation of axial stress distribution proposed by kecman: (a) buc yσ σ< ; (b) 2buc yσ σ≥ where e, v, a, b, and t are respectively, young’s modulus, poisson’s ratio, flange width, web width and tube thickness. in addition, a is 1c t+ and b is 2c t+ . kecman presented a proposal method in which was decided by relations of the buckling stress bucσ and yield stress yσ . (1) in the case of buc yσ σ< if the buckling stress bucσ is less than the yield stress yσ , the compression flange buckles 216 masuda: prediction of maximum moment of rectangular tubes subjected to pure bending and the edges stress come up to yield stress yσ . in order to consider this phenomenon, an effective width eσ is introduced in the following simplified equation. ( )0.7 0.3 2buc ea a σ σ ⎛ ⎞ = +⎜ ⎟⎜ ⎟ ⎝ ⎠y       as a result, stress distribution in the maximum moment is shown in fig. 2(a). in the figure, y1 in which the distances from a compression flange to the neutral axis is derived from the condition of zero axial loads. therefore, y1 is given by the following equation. ( )1 3 2e y a b b a a b + = + +       by summing moments through the cross-section, the maximum bending moment is derived in the following equation. ( ) ( )2 max 2 3 2 4 3 e aa b a bm t b a b σ ⎛ ⎞+ + ⋅ +⎜ ⎟ ⎝ ⎠⋅ ⋅ ⋅ +y=   (2) in the case of 2buc yσ σ≥ in this case, stress distribution in the maximum moment is shown in fig. 2(b). namely, it is assumed that the maximum moment is equal to a fully plastic moment mp. the maximum bending moment is derived in the following equation. ( ) ( ) ( )2 max 0.5 2 5pm m t a b t b tσ ⎡ ⎤⋅ − + −⎣ ⎦y= =   (3) in the case of 2y buc yσ σ σ≤ < first, if the buckling stress bucσ is equal to the yield stress yσ , it is assumed that the maximum moment is equal to a elastic moment me in which the stress of flanges are equal to the yield stress yσ . this elastic moment me is derived in the following equation. ( )6 3e bm t b aσ ⎛ ⎞⋅ ⋅ ⋅ +⎜ ⎟ ⎝ ⎠ y=   and in the case of 2y buc yσ σ σ≤ < , the maximum bending moment is derived from linear interpolation: ( ) ( )max 7buc y e p em m m m σ σ σ − − y = +   0 0.005 0.01 0.015 0.02 0.025 0 1 2 t / c1 m m ax / (σ y c 1 c 2 t ) σbuc<σy 2σy≤σbuc l = 300 mm c1 = 50 mm σy/ e = 1/1000 kecman [1] c2/c1=1 c2/c1=2 c2/c1=1 c2/c1=2 2σy≤σbuc σbuc<σy fig.3. comparison between kecman’s proposal and results of fem with relation of t/c1 and ( )max 1 2/ ym c c tσ ⋅ ⋅ ⋅ advances in systems science and applications (2011), vol.11, no.3-4 217 issn 1078-6236 international institute for general systems studies, inc. figure 3 compares these proposal methods by kecman and results of present numerical analyses for two levels of aspect ratio c2 /c1 with c1 =50mm, l=300mm, / 0.001y eσ = . as can be seen from this figure, in the case of high-aspect ratio in which the web is wider than the flange, the results of maximum moment under various t/c1 between the kecman’s proposal and results of fem have a margin of error. in particular, the error increases with decreasing t/c1. therefore, it is found that a region which does not apply to kecman’s proposal exists. in order to estimate the maximum moment, it is vital to reveal the bending collapse mechanism of rectangular tubes. 3.2 two types of collapse mechanism pointed out by kecman an investigation of two types of collapse mechanism pointed out by kecman was presented by using square tubes in which the aspect ratio c2 /c1 was set to 1. figure 4 shows the relation of a tube curvature / lκ θ= and moment m for a square tube with t=0.9mm, c1=50mm, c2=50mm, / 0.001y eσ = ( 1.52buc yσ σ= ). and figure 4 also shows the relations of the tube curvature / lκ θ= and axial stress /x yσ σ at point b and c (refer to schematic representation of cross-section in figure 4). as can be seen from the figure, the maximum moment is in good agreement with the value of kecman’s proposal equation (7). the axial compression stress /x yσ σ at point b in the middle of compression flange increases until the moment becomes maximum moment, and the value /x yσ σ comes up to 1. in addition, the axial compression stress /x yσ σ at point c in the quarter of web width increases until the moment becomes maximum moment. figure 5 shows the axial stress distribution of cross-section at phase ( )α and ( )β corresponding to 1/ 0.025l mθ −= and 10.065m− in fig. 4. as can be seen from the figure, the absolute value of the axial stress when the maximum moment occurs is greater than the value at phase ( )α in all cross-section positions. in addition, the axial stress distribution when the maximum moment occurs is in good agreement with kecman’s proposal. it is confirmed from the above investigation that in the case of 2 1/ 1c c = and y bucσ σ≤ , the collapse type is not due to buckling at the compression flange and web, but due to plastic yielding at the flanges. therefore, to estimate the maximum moment by kecman’s theory is possible in this case. 0 0.05 0.1 0 50 100 150 200 250 0 0.5 1 1.5 θ / l [m−1] m [ n m ] σ x /σ y moment σ x (point b) σ x (point c) eq.(7) l = 300 mm t = 0.9 mm c1 = 50 mm c2 = 50 mm extension side compression side b cc2/4 σy / e = 1/1000 σx /σy =1(α ) (β ) fig.4. relations of / lθ and , /x ym σ σ for square tube with t=0.9mm, c1=50mm, c2=50mm 218 masuda: prediction of maximum moment of rectangular tubes subjected to pure bending 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 1.5 σ x / σ y s / (2c1+2c2 ) s extension side compression side a b a b l = 300 mm t = 0.9 mm c1 = 50 mm c2 = 50 mm σy / ε =1/1000 kecman (α ) (β ) fig.5. axial stress distribution of cross-section for the square tube shown in fig.4 0 0.05 0.1 0 40 80 0 0.5 1 θ / l [m−1] m [ n m ] σ x /σ y moment σ x (point b) σ x (point c) eq.(4) eq.(1) l = 300 mm t = 0.4 mm c1 = 50 mm c2 = 50 mm extension side compression side b cc2/4 σy /e= 1/1000 (α ) (β ) fig.6. relations of / lθ and , /x ym σ σ for square tube with t=0.4mm, c1=50mm, c2=50mm figure 6 shows the relation of a tube curvature / lκ θ= and moment m for a square tube with t=0.4mm, c1=50mm, c2=50mm, / 0.001y eσ = ( 0.31buc yσ σ= ). and figure 6 also shows the relations of the tube curvature / lκ θ= and axial stress /x yσ σ at point b and c (refer to schematic representation of cross-section in figure 6). as can be seen from the figure, the maximum moment is in good agreement with the value of kecman’s proposal equation (4). the axial compression stress /x yσ σ at point b in the middle of compression flange decreases before the moment becomes maximum moment, and the maximum value /x yσ σ is in good agreement with the equation (1) of elastic buckling stress. in addition, the axial compression stress /x yσ σ at point c in the quarter of web width increases until the moment becomes maximum moment. figure 7 shows the axial stress distribution of cross-section at phase ( )α and ( )β corresponding to 1/ 0.012l mθ −= and 10.038m− in fig. 6. as can be seen from the figure, although the axial compression stress in the middle of compression flange decreases due to buckling in the middle of compression flange, the axial compression stress in both edges advances in systems science and applications (2011), vol.11, no.3-4 219 issn 1078-6236 international institute for general systems studies, inc. of compression flange increase because buckling doesn’t occur in both edges. immediately after buckling, stress increment in the both edges is greater than stress decrement in the middle of compression flange. therefore, total force of compression side and moment increase. in addition, the stress distribution of the web changes linearly because buckling doesn’t occur in the web. therefore, the axial stress distribution when the maximum moment occurs is in good agreement with kecman’s proposal by using an effective width in the compression flange. it is confirmed from the above investigation that in the case of 2 1/ 1c c = and y bucσ σ> , the collapse type is due to buckling at the compression flange. therefore, to estimate the maximum moment by kecman’s theory is possible in this case. 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 kecman σ x / σ y s / (2c1+2c2 ) s extension side compression side a b a b l = 300 mm t = 0.4 mm c1 = 50 mm c2 = 50 mm (α ) (β ) fig.7. axial stress distribution of cross-section for the square tube shown in fig.6 3.3 collapse mechanism which is different from kecman’s indication in the case of high-aspect ratio in which the web is wider than the flange, it was confirmed that collapse due to buckling at the compression web occur. 0 0.005 0.01 0.015 0 100 200 300 0 1 2 θ / l [m−1] m [ n m ] σ x /σ y moment σ x (point b) σ x (point c) eq.(5) l = 300 mm t = 0.5 mm c1 = 20 mm c2 = 100 mm extension side compression side b cc2/4 σy / e = 1/1000 σx /σy =1 (α ) (β ) fig.8. relations of / lθ and , /x ym σ σ for rectangular tube with t=0.5mm, c1=20mm, c2=100mm 220 masuda: prediction of maximum moment of rectangular tubes subjected to pure bending figure 8 shows the relation of a tube curvature / lκ θ= and moment m for a rectangular tube with t=0.5mm, c1=20mm, c2=100mm, / 0.001y eσ = ( 2 12.83 , / 5buc y c cσ σ= = ). and figure 8 also shows the relations of the tube curvature / lκ θ= and axial stress /x yσ σ at point b and c (refer to schematic representation of cross-section in figure 8). as can be seen from the figure, the maximum moment is less than the value of kecman’s proposal equation (5). in addition, the axial compression stress /x yσ σ at point b in the middle of compression flange increases until the moment becomes maximum moment, and the value /x yσ σ comes up to 1. and the axial compression stress /x yσ σ at point c in the quarter of web width decreases before the moment becomes maximum moment. figure 9 shows the axial stress distribution of cross-section at phase ( )α and ( )β corresponding to 1/ 0.036l mθ −= and 10.048m− in fig. 8. as can be seen from the figure, the axial stress distribution in the compression flange is constant value and the absolute value is almost 1 when the maximum moment occurs. and the axial stress distribution in the compression web doesn’t increase linearly. therefore, the sum of axial stress when the maximum moment occurs is less than the kecman’s proposal as much as it is shown by arrows of figure 9. 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 1.5 σ x / σ y s / (2c1+2c2 ) s extension side compression side a b a b l = 300 mm t = 0.5 mm c1 = 20 mm c2 = 100 mm σy / ε =1/1000 kecman (α ) (β ) fig.9. axial stress distribution of cross-section for the rectangular tube shown in fig.8 it is found from the above investigation that in the case of high-aspect ratio and y bucσ σ< , the collapse type is not due to buckling at the compression flange but due to buckling at the compression web. therefore, to estimate the maximum moment by kecman’s theory is impossible in this case. figure 10 shows the relation of a tube curvature / lκ θ= and moment m for a rectangular tube with t=0.4mm, c1=50mm, c2=100mm, / 0.001y eσ = ,( 0.30buc yσ σ= ), ( 2 1/ 2c c = ). and figure 10 also shows the relations of the tube curvature / lκ θ= and axial stress /x yσ σ at point b and c (refer to schematic representation of cross-section in figure 10). as can be seen from the figure, the maximum moment is less than the value of kecman’s proposal equation (4). the axial compression stress /x yσ σ at point b in the middle of compression flange decreases before the moment becomes maximum moment, and the maximum value /x yσ σ is in good agreement with the equation (1) of elastic buckling stress. in addition, the axial compression stress /x yσ σ at point c in the quarter of web width decreases before the moment becomes advances in systems science and applications (2011), vol.11, no.3-4 221 issn 1078-6236 international institute for general systems studies, inc. maximum moment. figure 11 shows the axial stress distribution of cross-section at phase ( )α and ( )β corresponding to 1/ 0.007l mθ −= and 10.016m− in fig. 10. as can be seen from the figure, the axial stress in the compression flange is concentrated in the edges when the maximum moment occurs. and the axial stress distribution in the compression web doesn’t increase linearly. therefore, the sum of axial stress in the maximum moment is less than the kecman’s proposal as much as it is shown by arrows of figure 11 because the equation (2) applies to the axial stress distribution of compression flange, and linearly approximation doesn’t apply to the axial stress distribution of compression web. it is found from the above investigation that in the case of high-aspect ratio and y bucσ σ≥ , the collapse type is not only due to buckling at the compression flange but also due to buckling at the compression web. therefore, to estimate the maximum moment by kecman’s theory is impossible in this case. 0 0.02 0.04 0.06 0 100 200 0 0.5 1 θ / l [m−1] m [ n m ] σ x /σ y moment σ x (point b) σ x (point c) eq.(4) eq.(1) l = 300 mm t = 0.4 mm c1 = 50 mm c2 = 100 mm c b extension side compression side c2/4 σx /σy = 1/1000 (α ) (β ) fig.10. relations of / lθ and , /x ym σ σ for rectangular tube with t=0.5mm, c1=20mm, c2=100mm 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 kecman σ x / σ y s / (2c1+2c2 ) s extension side compression side l = 300 mm c1 = 50 mm c2 = 100 mm t = 0.4 mm a b a b (α ) (β ) fig.11. axial stress distribution of cross-section for the rectangular tube shown in fig.10 222 masuda: prediction of maximum moment of rectangular tubes subjected to pure bending 3.4 proposal method of maximum moment considering the web buckling 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 fem kecman’s proposal σ x / σ y s / (2c1+2c2 ) s extension side compression side l = 300 mm c1 = 50 mm c2 = 30 mm t = 0.4 mm present proposal (b) 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 fem kecman’s proposal σ x / σ y s / (2c1+2c2 ) s extension side compression side l = 300 mm c1 = 50 mm c2 = 100 mm t = 0.4 mm present proposal (a) fig.12. axial stress distribution in the range of buc yσ σ< : (a) by kecman’s method; (b) by present method figure 12 shows a schematic representation of axial stress distribution in the maximum moment after buckling of the compression flange ( buc yσ σ< ). figure 12(a) shows kecman’s proposal in which doesn’t consider the web buckling, and figure 12(b) shows present proposal in which considers the web buckling. as can be seen in the figure(b), an effective width ae applies to the compression web as well as the compression flange. it is assumed that the effective width ae is independent of initial web width as referring to karman’s theory [9]. a coefficient α in which represents the axial tension stress is derived in the following equation. ( ) ( ) 1 2 8 2 ea t a b y t α − = + − −       figures 13(a) and (b) show comparison between results of fem and proposals with the axial stress distribution. as can be seen in the figures, in the case of high-aspect ratio 2 1/ 2c c = , present proposal in which considers the web buckling is in good agreement with the result of fem. and in the case of low-aspect ratio 2 1/ 0.6c c = , kecman’s proposal in which doesn’t consider the web buckling is in good agreement with the result of fem. b a 2 eayσ 1y 1 1 y b y y σ − (a) b a 2 ea yσ 1y yασ b a 2 ea yσ 1y yασ (b) fig.13. axial stress distribution by results of fem, kecman’s proposal and present proposal: with (a) 2 1/ 2c c = ; (b) 2 1/ 0.6c c = figure 14 shows a schematic representation of axial stress distribution in the maximum moment when the compression flange doesn’t buckle ( buc yσ σ≥ ). figure 14(a) shows kecman’s proposal in which doesn’t consider the web buckling, and figure 14(b) shows present proposal in which considers the web buckling. as can be seen in the figure, an effective width ae applies to only the compression web. a coefficient β in which represents the axial tension stress is derived in the following equation. advances in systems science and applications (2011), vol.11, no.3-4 223 issn 1078-6236 international institute for general systems studies, inc. ( ) 1 2 9 2 ea a t a b y t β + = + − −       figures 15(a) and (b) show comparison between results of fem and proposals with the axial stress distribution. as can be seen in the figures, in the case of high-aspect ratio 2 1/ 5c c = , present proposal in which considers the web buckling is in good agreement with the result of fem. and in the case of low-aspect ratio 2 1/ 2c c = , kecman’s proposal in which doesn’t consider the web buckling is in good agreement with the result of fem. we show present method for estimating the maximum bending moment of rectangular tubes subjected to pure bending. in the case of buc yσ σ< , a position of the center of gravity in the tension web g is derived in the following equation. b ayσ / 2b yσ (a) b / 2ea yσ 1y yβσ (b) a fig.14. axial stress distribution in the range of buc yσ σ≥ : (a) by kecman’s method; (b) by present method 0 0.5 1 −1.5 −1 −0.5 0 0.5 1 fem kecman’s proposal σ x / σ y s / (2c1+2c2 ) s extension side compression side l = 300 mm c1 = 20 mm c2 = 100 mm t = 0.5 mm present proposal (a) 0 0.5 1 −2 −1 0 1 2 fem kecman’s proposal σ x / σ y s / (2c1+2c2 ) s extension side compression side l = 200 mm c1 = 20 mm c2 = 40 mm t = 0.5 mm present proposal (b) fig.15. axial stress distribution by results of fem, kecman’s proposal and present proposal: with (a) 2 1/ 5c c = ; (b) 2 1/ 2c c = ( )1 1 1 10 3 2 g b y⎛ ⎞= +⎜ ⎟ ⎝ ⎠     therefore, in the case of buc yσ σ< , the maximum moment in which the compression web buckles is derived in the following equation. ( )( ) ( ){ ( )( )max 1 1 2 2 2 2 (11) 2 2 4 e y e e abm t a t b t b y g a t b t aσ α α ⎫⎛ ⎞= − − + − + − − + − ⎬⎜ ⎟ ⎝ ⎠⎭         where 1,y α and g are respectively, the value of equation (3), (8) and (10). it is found from the above investigation that in the case of buc yσ σ< , the maximum moment is derived in the following equation. ( ) ( )( ) ( )max . 4 , . 11 12m min eq eq=     224 masuda: prediction of maximum moment of rectangular tubes subjected to pure bending { }1 ( 2 )( ) 2 ( ) ( 2 )( ) 2 ( ) (13)max 12 2 4 ab em t a t b t b y g a t b t ay eσ β β= − − + − + − − + − moreover, in the case of buc yσ σ≥ , the maximum moment in which the compression web buckles is derived in the following equation. where 1,y β and g are respectively, the value of equation (3), (9) and (10). it is found from the above investigation that in the case of buc yσ σ≥ , the maximum moment is derived in the following equation. ( ) ( ) ( )( ) ( )max . 5 , . 7 , 13 14m min eq eq eq=   figure 16 shows comparison between results of fem and kecman’s proposal and present proposal with the maximum moment for two levels of /y eσ . as can be seen in the figure, the lower values of kecman’s proposal and present proposal is in good agreement with the results of fem. 0 0.005 0.01 0.015 0.02 0.025 0 0.5 1 1.5 2 2.5 t / c1 m m ax / (σ y c 1 c 2 t ) l = 300 mm c1 = 50 mm kecman’s proposal 2σy≤σbuc σbuc<σy σbuc<σy σy/ e = 1/1000 σy/ e = 1/500 present proposal c2 = 100 mm σy/ e = 1/1000 σy/ e = 1/500 comparison between results of fem and kecman’s proposal and present proposal with relation of 1/t c and ( )max 1 2/ ym c c tσ ⋅ ⋅ ⋅ for two levels of /y eσ . 4. conclusion in this paper, the investigation of the bending collapse for rectangular tubes by using numerical analysis of the finite element method was presented. for a rectangular tube in which the web is wider than the flange, it is found that collapse due to buckling at the compression web may occur. it is possible to estimate the maximum moment under various the material and geometrical properties by using the present proposal in which applies an effective width to the web, and the kecman’s proposal. references [1] kecman d. bending collapse of rectangular and square section tubes, international journal of mechanical sciences, vol.25 (1983) 623-636. [2] kim t. h. and reid s. r. bending collapse of thin-walled rectangular section columns, computers and structures, vol.79 (2001) 1897-1911. [3] lu g. and yu t. x. energy absorption of structures and materials, (2003) section5 crc pr i llc. [4] parh m. s. and lee b. c. prediction of bending collapse behaviours of thin-walled open section beams, thin-walled structures, vol.25(3) (1996) 185-206. [5] kyriakides s. and ju g. t. bifurcation and localization instabilities in cylindrical shells advances in systems science and applications (2011), vol.11, no.3-4 225 issn 1078-6236 international institute for general systems studies, inc. under bending, international journal of solids and structures, vol.29 (1992) 1117-1171. [6] chen d. h. et al. study on elastoplastic pure bending collapse of cylindrical tubes, transactions of the japan society of mechanical engineers, series a, vol.74(740) (2008) 520-527. [7] msc. marc manual, 2003. [8] guarracino f. on the analysis of cylindrical tubes under flexure: theoretical formulations, experimental data and finite element analyses, thin-walled structures, vol.41 (2003) 127-147. [9] karman v. t. et al. strength of thin plates in compression, trans asme, vol.54 (1932) мягкая вероятность ошибки при передаче информации advances in systems science and application(2016) vol.16 no.4 71-80 specifics of functional planning and architectural organization of religious educational complexes as a new type of educational institutions olga i. zhovkva kyev national university of construction and architecture, ukraine annotation the article identifies the specifics of shaping the architecture of religious educational complexes as a new type of religious educational institutions. on the basis of the consolidation of international scientific and practical experience in the design and construction of religious educational institutions, the author substantiated the functional structure and functional relationships of religious educational complexes, provided the typology and characteristics of main functional groups of premises. the article specifies the methods of functional and planning organization of religious educational complexes, including modular, pavilion, combined and compact, wherein the latter is regarded as optimal in case of designing religious educational complexes in an urban area. the planning schemes and a composition structure of religious educational complexes are also discussed. alongside with the development of complete compositional solutions of religious educational complexes, it is proposed to create dynamic evolving compositions, making it possible to carry out further reconstruction or construction of new functional units. the article also discusses the types of volumetric and spatial composition, identifies the main approaches to creating architectural expression and harmonizing the architectural space of religious educational complexes. key words: architecture of religious educational institutions, sacral architecture, synamic architecture, functional and planning organization, volumetric and spatial composition, universal space, transformability, harmonization. 1 introduction the experience of the development of civilization shows that education is the foundation of spiritual, social and economic development of society. in recent years, religious life in ukraine has significantly intensified. according to statistics, a steady increase in the number of religious organizations has been observed in ukraine since 1997. new religious buildings and complexes are constructed, which determines the need for clergymen, superiors and counsellors – graduates of religious educational institutions. the relevance of the research issue for ukraine is also demonstrated by the amendments recently introduced to some laws of ukraine by the supreme council of ukraine, related to the possibility of the foundation of educational institutions by religious organizations. for instance, official religious organizations are allowed to found higher, vocational, general education, preschool and out-of-school institutions that will give children from religious families an opportunity to receive education in a religious environment and strengthen moral values of the younger generation. currently, although ukraine is a multi-religious society, the architecture of educational institutions of some christian (catholic, protestant) denominations, as well as muslim and 72 olga i. zhovkva: specifics of functional planning and architectural organization of religious educational complexes… jewish institutions, remains understudied, despite the presence of a fairly large number of representatives of these denominations in the territory of ukraine. the issue of designing religious educational complexes as a new type of religious educational institutions has not been studied in both domestic and foreign science. identification of the specifics of shaping the architecture of religious educational complexes will make it possible to improve design solutions, optimize and improve their functional and planning structure, harmonize the relationships between the architecture and architectural environment. 2 methodology the study is based on the approaches that regard religious educational complexes as a complex system phenomenon, which requires a comprehensive study by related sciences such as philosophy, art history, religious studies, history. the following methods were used in the course of the study: field survey of religious educational institutions of ukraine, cis countries and countries outside the former soviet union; questionnaire survey of representatives and management of educational institutions; comparative analysis of foreign and domestic experience of design; system approach; environmental approach; graphic-analytical method and experimental modeling method. the method of field survey was used to study the current state and the experience of design and construction of christian, jewish and muslim religious educational institutions in the territory of ukraine and abroad (in russia, belarus, uzbekistan, poland, italy, turkey). the review of the current state of the design and usage of religious educational institutions in ukraine provided an opportunity to identify both disadvantages and positive aspects, while comparative analysis made it possible to determine the ways of further improvement and development of these institutions. the disadvantages of religious educational institutions in ukraine are the following: facilities and resources of 50-60% of the surveyed objects don’t meet contemporary requirements to modern educational process. the majority of religious educational institutions do not have their own premises, designed for academic process, and are housed in adapted buildings not having all necessary functional groups of premises. the applied method made it possible to identify the need for the modernization of existing religious educational institutions, as well as the need for the introduction of a new type of building – a religious educational complex designed with regard to the basic modern educational requirements. the method of questionnaire survey involved interviews and a survey among teachers and management of religious educational institutions. respondents had to answer, what the spiritual education will be like in the near future, in which direction it is necessary to improve exiting institutions, what a religious educational institution of the 21st century should be like. clergymen were also asked to answer the questions about the desired optimal functional composition of modern religious educational institutions and their spatial organization. on the basis of the analysis of a large number of answers and interviews it was assumed that there is a need for the introduction of a new type of integrated religious educational institution with a sufficiently large amount of functions – a religious educational complex. the introduction of the concept of a religious educational complex to the architectural vocabulary made it possible to offer new trends of shaping the architecture of religious educational institutions, aimed at improving the conditions for obtaining spiritual education. the method of comparative analysis was used to study the typology of religious educational institutions of ukraine, cis countries and countries outside the former soviet union. this method made it possible to identify similarity and differences in the architecture of religious advances in systems science and application(2016) vol.16 no.4 73 educational institutions of different countries, different denominations and types, to propose a universal functional structure of institutions that can be used by christian, muslim and jewish denominations when designing religious educational complexes. 3 the main part the creation of religious educational complexes (campuses) is one of perspective lines of the development of religious educational institutions. for the development of the concept of the creation of religious educational complexes it is necessary to recall the experience of medieval universities. towns-universities, such as oxford or cambridge, are the examples of modern cooperated educational establishments, comprising several educational institutions [20]. the contemporary experience of design and construction of religious educational institutions of the catholic church should also be taken into account [21]. therefore, alongside with the modernization, reconstruction and expansion of the resource base of existing religious educational institutions, an attempt should be made to organize campuses – complexes comprising several religious educational institutions of different levels. however, in order to provide the conditions for educational process in an educational establishment, it is necessary to carry out a series of studies to substantiate the functional composition and zoning of the complex with regard to the technological process, power supply and the specifics of operation. an educational unit comprises: a group of classrooms for training specialists of primary, secondary (including vocational) and hither level, as well as teachers’ rooms and rooms of departments. these rooms can be housed either in one unit or in several separate educational units. functional longevity of an educational unit will provide the flexibility and dynamism of its planning organization and structure [4]. if territorial resources permit, it is expedient to use a separate unit of a religious educational complex for the organization of production and training workshops (for icon-painting, goldembroidery, carpentry -specialized in manufacturing kiots and iconostases), as well as classrooms for choral singing and teaching choirmasters. a sacred unit comprises sacred rooms and a group of auxiliary premises (baptistery, acolyte’s room, sacristy). in the general composition of a religious educational complex, the sacred unit is usually placed separately, but if the territorial resources are limited it can have a form of an extension to educational, refectory or housing units of the complex, or have the form of a house church. the architecture of a sacred unit brings uniqueness and specificity into the composition of religious educational complexes of different denominations [15]. the architectural and planning composition of this unit has a number of traits that are common for the studied five major religions [17, 18,]. for example, the sacred core of an orthodox educational institution (an orthodox church) is formed by the following three parts: the altar, the place of prayer and the narthex [19]. the altar in the orthodoxy is regarded as a sacred spot; the middle part is a place of prayer for the congregation; the narthex is a scope of earthly existence. an orthodox church is considered to be a symbol of the universe, and its entire architecture symbolically embodies the christian idea of the connection between earthly and heavenly [3]. the altar of an orthodox church is always oriented to the east and separated from the prayer hall by the iconostasis, which is regarded as a baffle separating the two worlds [12]. a catholic church (the sacred core of a catholic educational complex) has, in general, the same structure as an orthodox church. an essential difference of its internal structure is the absence of the iconostasis. a catholic church also consists of three parts: the central part of the church; altar, or presbytery, where the holy gifts are stored; and the narthex [11]. the planning structure of a catholic religious building, in a similar way to the configuration of an orthodox 74 olga i. zhovkva: specifics of functional planning and architectural organization of religious educational complexes… church, can include special places for chorus and separate premises for church ministers and for the storage of vestments [9]. the altar of a catholic church is also oriented to the east. protestant houses of prayer are free from exuberant decorations, icons and sculptures. such decoration is unnecessary for protestants, because in this faith there is no worship of icons in the form in which it can be observed in catholic and orthodox churches [1]. a church can be housed in any building, which is rented or purchased; there are no strict prescriptions regarding the orientation of certain groups of premises. in general, however, the protestant theology does not contradict the theological decisions of oecumenical councils. protestant houses of prayer usually also have a three-part structure: the entrance area (a counterpart of narthex), the prayer hall itself and the sacred area with a place for a sermon. in the prayer hall of a mosque, next to the mihrab, a minbar (a pulpit for a preacher) is constructed, which is a counterpart of the pulpit of protestant religious buildings, the ambon in orthodox and catholic churches, the bimah in synagogues. as in the orthodox christianity, the rectangular foundation of mosques symbolized the earth, while the spherical dome – the heaven [13]. the internal arrangement of a synagogue is also based on the construction of a temple, which in turn recreates the internal design of the tabernacle (framed rectangular space with the sanctuary in the middle). that is why a synagogue usually has a rectangular shape. in the part where in a christian church there is a sanctuary, in synagogue there is a receptacle called aron kodesh, containing torah scrolls. in the center of a synagogue there is an elevated platform from which the torah is read – the bimah [8]. synagogues must be placed in such a way that the wall, near which the aron kodesh is placed, is directed towards jerusalem. for jewish synagogues this means orientation to the east, as well as for christian and muslim religious buildings. in the planning structure of a settlement a synagogue is usually placed at the highest point, which is also characteristic for religious buildings of the above-mentioned denominations. for this very reason the sacred core of the general composition of a religious educational complex should play a dominant part in the general architectural and spatial composition. thus, the above-discussed sacred units of religious educational complexes of five major religions as an integral component of such complexes, while definitely having their distinctive features, have also much in common. an administrative unit of religious educational complexes comprises the following service rooms: an office of a superior (a rector), offices of deputies, accounting department, administration, etc.; a group of health care premises; a group of auxiliary and subsidiary premises (premises for the production of candles); a refectory group. multifunctional religious educational complexes may also comprise a missionary and charity unit, in which groups of premises for missionary-volunteer activities should be provided; a group of premises for psychological and physical rehabilitation of children and adults who find themselves in difficult circumstances; rooms for counselling with priests, psychologists, lawyers; rooms for temporary stay of people who need shelter; a medical aid station; a refectory; sacred rooms. the analysis of the structures of existing religious educational institutions showed that in order to ensure their functional longevity or the possibility of technical modernization with regard to modern requirements of constantly changing educational process it is reasonable to use universal units with flexible planning, on the basis of which dynamically developing systems can be organized [4]. due to the principle of universality, a unit can perform an educational, administrative or recreational function. such a unit is a single flowing space, divided in a way determined by functional reasons. only basic structural elements of this unit are fixed: vertical structures and utility systems, as well as main load-bearing structural elements [5]. advances in systems science and application(2016) vol.16 no.4 75 a general functional structure of major religious educational complexes (intended for 700 and more students) contains the maximum number of functional units, which are united into architectural and planning schemes by means of functional and compositional relationships. the units of of religious educational institutions form various volumetric and spatial compositions, among which the following types can be specified: closed, semiclosed and open (table 1). table 1. types of volumetric and spatial composition of religious educational complexes closed open now let us focus on the types and the analysis of the basic groups of premises, which form the above-discussed units of religious educational complexes, and propose a typology of premises (table 2). the entrance group of premises of various functional units usually consists of an entrance hall, a cloakroom and sanitary facilities. the primary purpose of the entrance hall is to rationally organize the flows of students, teachers and guests of the complex. the entrance hall of a religious educational complex can be used not only for serving visitors, but also for educational, missionary and club work; it can also serve as a reception room. such an entrance hall/reception room can be designed as open space, forming a part of back rooms or a corridor. the optimal solution is the one that makes it possible in case of necessity to separate the entrance hall as an independent space using sliding partitions. the main groups of premises in modern religious educational complexes are training and applied training groups of premises comprising classrooms, lecture rooms, music and rehearsal classrooms, choral singing classrooms, icon-painting, gold-embroidery, cabinetry and carpentry workshops. classrooms intended for 10-25 students are the main premises of the complex, because students spend there most of their time. the most optimal form of a school room may be a rectangular shape with a longitudinal external wall and the dimensions in the axes 3 * 6 m, 6 * 9 m. such proportions make it possible to arrange students’ desks in three rows, thus ensuring the optimum illumination of all work areas. using another constructive decision makes it possible to obtain a classroom of the third type, intended for 25-50 students. when designing religious educational complexes, more complex forms of classrooms can also be used: trapezoidal, pentagonal, hexagonal, etc. classrooms having a trapezoidal shape make it possible to conveniently group desktops with the possibility of allotting a separate area for a personal prayer. in the premises having a pentagonal or square shape it is easier to organize workplace illumination using natural light [6]. however, nowadays the requirements to the illumination of premises cannot be considered to be most important ones, because students often use a pc for capturing information received from a teacher; in this case the presence of natural lightning is not a mandatory requirement. nonstandard forms of classrooms open new possibilities for architectural and artistic solution of the interior of a religious educational complex, providing more freedom in arranging equipment and making it possible to organize a sacred area in classrooms. 76 olga i. zhovkva: specifics of functional planning and architectural organization of religious educational complexes… table 2. typology of premises multipurpose room classroom for 20-25 students, audience workshop reading room assembly hall exhibition room by virtue of the latest technologies, currently it is not necessary to organize equipment of a classroom in such a way that all the attention is focused on a teacher’s work area. the possibility of installing several monitors in a classroom to broadcast the lecture material eliminates the need to focus attention on one area, which also makes it possible to arrange workplaces not in a straightline scheme, but radially or in segments. the survey showed that the process of training in higher religious educational is effective in classrooms with a capacity of no more than 25-30 students for seminaries and 20-25 students for academies. advances in systems science and application(2016) vol.16 no.4 77 lecture rooms should be intended for 50-100 students. a progressive and cost-effective direction of the universalization of educational units and creating audiences of large capacity is the possibility of the transformation of medium sized lecture rooms into larger ones. handicraft rooms (icon-painting, gold-embroidery, cabinetry workshops) should be placed on the ground floor; it is also possible to place them in basements with mandatory natural lighting. the cutural and recreational group of premises comprises library premises, a publishing center, internet rooms; a group of premises for meetings and discussions; club and entertainment premises, exhibition rooms, halls, lounges. a library is the embodiment of an educational institution, its intellectual center, where the knowledge of previous generations is stored. it is formed on the basis of the premises of a reading room with an area for working with books and periodicals, reading newspapers and viewing microfilms, listening to audio records, the premises of circulation department, catalogue, storeroom, open-access fund. the size of a reading room depends on the number of students in the religious complex. the transition to the electronic storage of information eliminates the need for the organization of reading rooms in the classic sense with fixed workplaces. students can receive required materials, books or publications in an electronic form and process them in any place of the educational complex, which is most comfortable for the work. the most important features of a library are flexibility and multi-functionality – the possibility of using the working rooms also for classes, meetings, exhibitions. a group of entertainment premises is very important in the structure of a religious educational complex. the assembly hall of a religious educational complex is generally used to hold solemn meetings, deliver lectures, etc. when designing an assembly hall, it is necessary to provide for the possibility of its multiple use, for example, as a conference hall, a concert or an exhibition hall. one of the options is using the assembly hall as a large lecture room. besides, an assembly hall should have the capability to be transformed into a number of smaller audiences intended for 50-100 places, which will increase the efficiency of its use. modular planning should be used when designing medium-capacity religious educational complexes. efficiently connected functional units will ensure the optimal arrangement of the basic functions of religious educational complexes in individual units, provide a communication link between units through covered walkways or galleries. the advantage of this method of designing religious educational complexes is the possibility to isolate units and simultaneously provide a direct connection with the entrance area. pavilion planning is most widely used when designing medium-sized and large religious educational complexes. on the basis of this planning method it is possible to extend religious educational institutions through the construction of additional buildings and link units. this type of planning structure makes it possible to use plots with a complex terrain [10]. a disadvantage of this composition is the elongation of the functional links between units and groups of premises. combined planning is used when designing large-capacity religious educational complexes. it makes it possible to create expressive architectural compositions, optimize functional links, ensure optimal orientation of premises. according to the planning scheme, several types of compositions of religious educational complexes can be identified: t-shaped, h-shaped, e-shaped, o-shaped, perimetral, rectilinear, curvilinear, etc. all existing compositions are the result of the improvement of functional and technological schemes and the aspiration to improve the architectural space of religious educational complexes. according to the communication structure, the following compositional schemes of religious educational complexes can be identified: corridor-type, non-corridor, hall, enfilade and combined. 78 olga i. zhovkva: specifics of functional planning and architectural organization of religious educational complexes… corridor schemes are characteristic of linear buildings of religious educational complexes. a corridor system consists of cells that are linked by a common linear communication – a corridor. a non-corridor system integrates cells containing a part of the cycle of the general process in a religious educational complex. the link between the cells is represented by premises – a hall or a lounge, rather than a linear communication. such premises usually have their own particular purpose – a place of meetings, recreation, discussions and debates. the schemes described above can be used when designing religious educational complexes of different denominations (orthodox, catholic, protestant, muslim, jewish). when designing religious educational complexes, it is reasonable to combine the above-listed schemes, thus obtaining complex mixed and dynamic compositions. the prevalent compositional techniques of shaping and harmonizing the internal space of religious educational complexes are coloristics and ergodesign techniques that help create a truly comfortable artistic environment for effective learning and spiritual development. 4 discussion the study is based on the works by both domestic and foreign scientists in different areas of science: scientific works on religion by p. florensky; scientific works on the history of sacral architecture and the history of the church by n. budur, i. bondar'; scientific works, devoted to the study of architecture of religious educational institutions, by o. zhovkva, r. stots'ko and v. proskuryakov. the study is also based on the works, devoted to the general theoretical problems of designing educational institutions, by such scientists as v. ezhov and p. solobay. the study also relies on the works, devoted to the general theoretical and practical problems of architecture and the development of sacral architecture, by a. gutnov, a. gayduchenya, e. voznyak, sh. shukurov, o.boyko, e. kotlyar, v.kutsevich et al. these authors address the issues of shaping an architectural and artistic image, compositional solutions of architecture objects, including sacred ones. the architecture of orthodox and catholic religious educational institutions is also discussed in the monograph by r. stots'ko and v. proskuryakov “architecture of religious educational institutions of the ukrainian greek catholic church” and the monograph by o. zhovkva “architecture of orthodox religious educational institutions of ukraine", in which the authors conducted a study of the architecture of religious educational institutions of two influential christian denominations and analyzed the typology of these institutions. "selected works on art" by the priest p. florensky provide a deeper understanding of the symbolic essence of an orthodox temple and sacred art. in the work “iconostasis” florensky paid much attention to the canonical interpretation of the configuration of an orthodox temple and its altar. according to florensky, “... the altar means a human soul, and the temple – a body" [15]. this work made it possible to deeper understand the essence of the configuration of a temple as an integral part of an orthodox educational complex. the issues of the universality and transformability of space were discussed in the works by a number of scientists and architects, such as a. gutnov, a. gayduchenya. for example, a. gutnov in the book "world of architecture" discussed the issue of universal and free space of buildings and structures, relying on the idea of free planning declared by le corbusier [14]. using universal spaces and units in designing religious educational complexes makes it possible to improve the rationality of functional planning organization of such complexes. the above research and writings offer an opportunity to understand the state of knowledge on the issue of designing religious educational complexes and sacred buildings as their integral part. advances in systems science and application(2016) vol.16 no.4 79 these works cover the issue of temple construction and shaping the architecture of religious educational institutions. however, the issue of the complex shaping of the architecture of religious educational complexes of the major world religions remains understudied. that is why the issue of the development of architectural and planning solutions of religious educational complexes requires a comprehensive study and the preparation of relevant scientific guidelines and design recommendations. 5 conclusion a modern trend in designing religious educational complexes is the pursuit of universality and ensuring the possibility of the versatile use of premises and units, which would eliminate the problems of fast “moral” depreciation and obsolescence. the internal segmentation of such universal structure is carried out using facilities not associated with basic design structures: mobile partitions, walls, screens, etc. the basic requirement to the functional planning organization of premises of religious educational complexes is the possibility of flexible transformation, free planning. free planning will not only ensure the freedom in the search for optimal planning solutions and flexible transformation during the operation of a complex, but also brings into the internal space something new through the use of transformable partitions and bay window partitions that will make the interior more expressive. currently, a religious educational complex is understood as an environment that corresponds to the aesthetic and philosophical views of our time and provides favorable conditions for study, spiritual and physical development, self-improvement, teaching and scientific activity, accommodation and recreation. a religious educational complex should ensure the proper functioning of the educational process, provide an opportunity to hold meetings, exchange thoughts and have a rest. a religious educational complex is not only a temple of science, but also an environment designed for the education of an intelligent and humane personality. therefore, the architecture of a religious educational complex as a temple of science should keep up with the times [16]. the proposals, obtained as a result of the study, may be used in the process of designing religious educational institutions and complexes. the study determined the prerequisites for the development of a modern religious educational institution, as well as the need for the systematization of the scientific knowledge on architectural design and forecasting the development of religious educational complexes. on the basis of the consolidation of international scientific and practical experience in the study, design and construction of religious educational institutions the author substantiated the need for the development of a new type of a religious institution – a religious educational complex, its functional composition, structure and functional links. the study proposed a typology of the premises of religious educational complexes and a list of functional groups of premises of a modern religious educational complex including: educational, applied training (workshops), sacred, cultrural, recreational, sports, administrative, medical, refectory, utility and housing groups of premises. the paper specifies the methods of functional and planning organization of religious educational complexes – modular, pavilion, combined and compact, wherein the latter is regarded as optimal in case of designing religious educational complexes in an urban area. acknowledgment the study was conducted within the framework of the plans of research works of the kiev national university of building and architecture. 80 olga i. zhovkva: specifics of functional planning and architectural organization of religious educational complexes… the study was carried out with scientific and methodical support of the saint petersburg state university of architecture and civil engineering. some of the study results have been more than once reported at international scientific-practical conferences. the author expresses special gratitude to the staff of the design institute “mosproekt-2" named after m.v. posokhin for design and engineering support in the preparation of this study. references [1] bondar' and i.a.(2011). "lutheranism in the context of common european spiritual development: specifics of religious and cultural relations", thesis for the degree of candidate of philosophical sciences, kiev. p. 13 [2] boyko, o.g.(2010). "architecture of stone synagogues of right-bank ukraine of the 16th20th centuries", thesis for the degree of candidate of architecture, lviv. [3] budur, n.(2009). orthodox temple. moscow: rossa. p. 17 [4] gayduchenya, a.a.,(1983). dynamic architecture: main directions, development, methods. kiev: budivelnik. pp. 10-11 [5] gutnov, a.e. (1985). world of architecture: language of architecture. moscow: molodaya gvardiya. pp. 74-77 [6] yezhov, v.i.(1977). technology of designing school complexes. moscow: centre of scientific and technical information on civil engineering and architecture. p. 11,14. [7] zhovkva, o.i.(2009). architecture of orthodox religious educational institutions of ukraine. chernivtsi: llc druk art. [8] kotlyar, e.a.(2001). "synagogues of ukraine in the second half of the 16th century – early 20th century as a historical and cultural phenomenon", thesis for a degree of a candidate of art criticism, kharkiv. p. 9 [9] kutsevich, v.v. (ed.)(2002). religious buildings and structures of various denominations: design guidebook. kiev: kievzniiep. p. 70 [10] solobay, p.a.(2012). "typological foundations of shaping the architecture of higher education complexes", thesis for a degree of doctor of architecture, kiev. p. 15 [11] stots'ko, r.z. and v.i. proskuryakov(2009).architecture of religious educational institutions of the ukrainian greek catholic church: monograph. lviv: publishing house of the national university “lviv polytechnic”. p. 60 [12] florensky, p.a.(1996). selected works on art. moscow: mysl'. p. 97 [13] shukurov, sh.m.(2014). architecture of a modern mosque. origins. moscow: progress – tradition. pp. 46-47 [14] hasan uddin khan, p.j.(1998). international style: modernist architecture from 1925 to 1965. cologne: taschen. pp. 27-36 [15] jodidio, ph.(1997). new forms: architecture in the 1990s. cologne: taschen. pp. 113-114. [16] radoslav zuk’s seventh in a series of ukrainian catholic churches, built on geometric systems in holy trinity church, kerhonkson ny(1978). progressive architecture, no.10, pp.64. [17] religious architecture(1992-1993). techniques et architecture, no.405, pp. 34-66. [18] two churches(1986). progressive architecture, no.2, pp.81-82. [20] schöber, u.(2007). 1000 kirchen und kloster. cologne: naumann & göbel. pp. 49 [21] architektura. warszawa (1980). "ukrainian catholic clergy in diaspora (1751-1988): annotated list of priests who served outside of ukraine". rome. no.5, pp. 19-34 [23] contemporary materials recreate heritage(1988). building & construction, no.5, pp.47. [24] dom suits both sacred and secular events(1988). building & construction, no.5, pp.4445. corresponding author olga i. zhovkva can be contacted at: yal05@mail.ru. advances in systems science and applications (2012) vol.12 no.1 18-26 empirical analysis of credit risk measure using fuzzy adaptive neural-networks:for small and medium-sized enterprises zhichun xu and zongjun wang school of management, huazhong university of science technology,wuhan 430074 p.r.china abstract measuring credit risks of small and medium-sized enterprises should take a series of unique characteristics into consideration. firstly, non-financial conditions should be combined to produce more accurate evaluation results. secondly, to overcome the vague of non-financial and sharp boundaries of financial conditions, fuzzy approaches and a nonlinear classifier based on fuzzy adaptive neuralnetwork(hereinafter referred to as “fan”) were introduced. a sample family of 65 firms was collected to test the performance of the above classifier, among which 29 firms were default and 36 firms were good. the result was shown as roc curves. to compare the performance of the model, the empirical results of another 4 classifiers were shown at the same figures and tables. the results showed that fan has the best performance and can effectively distinguish the good and default firms. keywords fuzzy adaptive neural-network, small and medium-sized enterprises, credit risk, non-financial conditions 1 introduction it is well known that small and medium-sized enterprises (smes) have a series of unique characteristics differing from large enterprises. for instance, many small and medium-sized and businesses are lack of audited financial statements, and the reliability of their financial statements is doubtful to lenders. also, non-financial factors of the borrowers have deep impacts on the probability to meet their future obligations. these non-financial factors include the owners’ personal finances, as well as other attributes like their morality and beliefs. at last, macroeconomic environments and industry policies have important influence to operating of the firms. based on the above reasons, when we construct model to measure credit risk of smes, the unique characters should been taken serious consideration. in this paper, we will concentrate on these characters and put forward to a nonlinear classifier based on fuzzy adaptive neural-network to solve these problems.(einstein, 1954). the paper is structured as follows. the next section provides a survey of the technologies about credit risk modeling for smes and related work. section 3 describes our data set used in the paper. section 4 presents the feature extraction and the procedure of factors formation. section 5 discusses the empirical results advances in systems science and applications (2012) vol.12 no.1 19 and analyzes the results together with a comparison between the two methods. we conclude the paper in section 6. 2 technologies about credit risk measure for smes unlike many literatures emphasizing financial conditions, edmister (1988) claimed that numerical financial ratios and human credit analysis could be combined to produce more accurate evaluation results for credit risk measure. neglecting the information provided by these qualitative factors might result in undesirable consequences[1]. based on the judgments, many non-financial conditions were introduced to model the credit risk of smes. but the non-financial conditions always were linguistic and vague, they could not be easily treated like financial ratios in linear discriminant models such as mda and logistic regression model. even for financial information, the sharp boundaries also need seriously treated[2].to overcome this problem, fuzzy approaches were introduced. rommelfanger (1999) proposed the use of fuzzy logic to check the credit solvency of small business firms [3]. weber (1999) used fuzzy logic for credit worthiness evaluation[4]. syau et al. (2001) used fuzzy membership function to model the credit worthiness of an enterprise[5]. chen (1999) presented a fuzzy credit-rating approach to deal with the problem arisen from the credit rating table used in taiwan. the evaluation criteria were modeled as a hierarchical decision structure, and fuzzy integral was employed for aggregating credit information[6]. resent research has been done by jiao et al (2007). in their paper, fuzzy adaptive network was used to model the credit rating of small financial enterprises. the data of the credit rating problem was first represented by the use of fuzzy numbers. the network based on inference rules was constructed and was trained or learned by using the fuzzy number training data. because of the learning and the linguistic representation ability of the fuzzy adaptive network, they showed that it was ideally suited for the modeling of credit rating[2]. compared to jian’s work (2007), we consider the non-financial information when we classify smes by using fuzzy approaches together with neural-network. furthermore, we conduct factor analysis as a pretreatment process to reduce the complexity and detect structure in the relationships between variables.in addition, we apply analytic hierarchy process to get the weights of non-financial conditions. the input vectors have 5 dimensions, the output has only one dimension. the result will show on roc curves. 3 data acquisition statistics on smes are problematic almost by nature: the majority of them aren’t public companies and the financial information together with non-financial infor20 zhichun xu:empirical analysis of credit risk measure using fuzzy adaptive neural-networks. . . mation can’t be obtained publicly. in contrast to most other western countries, there is no long official default history that can be used. only by combining different sources, and expensive and labor-intensive process, can a usable database of default be constructed. as far as default companies are concerned, we adopt the definition in the basel accord ii. the basel ii (2004) suggest a conservative definition of default for a bank, that a default is considered to have occurred with regard to a particular obligor when the obligor is past due more than 90 days on any material credit obligation to the banking group, or the bank considers that the obligor is unlikely to pay its credit obligations to the banking group in full, without recourse by the bank to actions such as realizing security. the accounting data used in this study are collected by a state-owned bank’s credit management system. because the systems have been run since 2002, we get the data in the period between 2002 and 2008. when a firm is granted a credit rating by the bank, the bank’s loan managers are obliged to register their customers’ fixed information at first time together with accounting report each quarter of every year. during the period of 2002 to 2008, a total of 65 loans are selected. the data set comprises of 36 good loans (gl) and 29 bad (or default) loans (bl). all loans are under the normal loan scheme. all of the 65 loans correspond to 65 firms. the financial conditions are collected from the credit management system. the ratios are calculated based on the data in the balance sheets and income statements. the ratios of the default firms are calculated by the data of the years before default. the ratios of the good firms are computed by the newest data we can get. to obtain the objective evidence according to a company’s non-financial conditions, a reliable way is to collect the views of experienced bank loan officers who are asked to assess the degree of each information source using five linguistic terms, namely“very good”,“good”,“fair”,“bar”, and“very bad”. we collect the 65 firms’ basic information, and select 12 factors representing 4 kinds of abilities. the 12 non-financial factors are selected referring to the practices of loan risk classification of chinese commercial banks. we produce consulting tables and distribute them to 30 bank’s account managers and loan officers, and receive 24 answers. the answers are expressed by the above five linguistic. all the financial conditions and non-financial conditions are listed in table 1. 4 feature selection and factor formation described as above section, the financial conditions listed as table 1 have 10 items and non-financial conditions have 12 items. if all of them are used as inputs when we train fuzzy adaptive neural network, the nodes in hidden layer will become too large to properly work. in order to reduce the number of variables and to advances in systems science and applications (2012) vol.12 no.1 21 detect structure in the relationships between variables, an effective method to deal with financial conditions is the application of factor analytic techniques. we can combine the variables into several factors and reduce the complexity. table 1 variables used in the model financial conditions(f) non-financial conditions −−debt payment ability −−debt payment ability(nf1) debt-to-asset ratio(f1) period matching of liabilities and assets(nf11) quick ratio(f2) ability of financing from outside in time(nf12) current ratio(f3) collateral(nf3) −− earning ability(nf2) −− earning ability sustained profitability(nf21) return on assets(f4) core business(nf22) return on net sales(f5) quality of profit(nf23) −− management ability −− management ability(nf3) total assets turnover(f6) innovation ability(nf31) inventory turnover(f7) customer degree of satisfaction(nf32) equipment and technologies(nf33) −− growth ability −−characteristics of company and owner(nf4) sale growth ratio(f8) education degree of the owner(nf41) net asset growth ratio(f9) moral behavior of the owner(nf42) net profit growth ratio(f10) specifications of management(nf43) differ from those financial conditions can be expressed by numerical ratios, non-financial conditions are the problem we must seriously consider. as a result, we firstly put our focuses on the fuzzy processing of non-financial information. similar to the work of jiao(2007), our purpose is to get an aggregative score to express the information. in order to get the total score, we should calculate the objective evidence of each information source. to obtain the objective evidence according to a company’s non-financial conditions, we get five linguistic expressions by the above section’s methods. for the subsequent fuzzy operations, these linguistic terms should be translated into fuzzy numbers. we use five triangular fuzzy numbers to describe these linguistic terms according to a kind of conversion scale. the definitions of these fuzzy numbers are (80, 90, 90), (70, 80, 90), (60, 70, 80), (50, 60, 70), and (50, 50, 70), respectively, as shown in fig.1. by the fuzzy numbers, we can compute the average numbers of the non-financial conditions of each firm. it should be noted that another problem is to decide the weighting of each nonfinancial criterion. since this weighting depends on the particular problem at a particular time, the decision-maker usually determines it. an effective method to 22 zhichun xu:empirical analysis of credit risk measure using fuzzy adaptive neural-networks. . . fig.1 membership functions of the five rating levels for non-financial conditions build up the weighting structure is analytic hierarchy process. since the focus of this paper is to illustrate the application of fuzzy logic and neural network, the process of analytic hierarchy won’t be described in detail. we simplely give the result in table 2. table 2 normalized weights for the non-financial conditions category variables nf11 nf12 nf13 nf21 nf2 nf23 weights 0.23 0.45 0.32 0.25 0.35 0.30 variables nf31 nf32 nf33 nf41 nf42 nf43 weights 0.19 0.23 0.42 0.31 0.53 0.16 based on fuzzy arithmetic the aggregation of n fuzzy triangular numbers xi =( xi l, xi c , xi r ) ,, i=1, 2, 3, .n, with normalized weights w = (w1,w2, . . . ,wn) can be carried out using the following equations: ac = n∑ i=1 wixi c al = n∑ i=1 wixi l ar = n∑ i=1 wixi r (1) the final aggregated triangular fuzzy number is assumed as a = ( al, ac , ar ) . using the normalized weights listed in table 2, the triangular fuzzy numbers for non-financial conditions and eq. (1), the aggregated triangular fuzzy numbers for the four non-financial conditions categories of each sample can be obtained, respectively. in order to compare or rank the credit ratings of the fuzzy numbers obtained above, the overall existence ranking index (oeri) will be used. for triangular fuzzy numbers, the overall existence measure om for fuzzy set a can be written advances in systems science and applications (2012) vol.12 no.1 23 as: om (a) = 4al − ac + ar 4 (2) using eq.(2), we can obtain the oeri ratings of each sample company. we describe the process of feature selection as follows. step1. use factor analytic approach to treat the financial conditions. as a result, we obtain four main components. step2. build up the membership functions of the five rating levels for each non-financial criterion and the result criterion. step3. decide the weights of each criterion using analytic hierarchy process. step4. calculate the objective evidence of the four non-financial conditions categories by aggregating the triangular fuzzy numbers and rank the fuzzy numbers using the overall existence ranking index approach. finally, a ranking numerical value for each sample can be obtained. using factor analysis technique, we can combine 10 variables into several main components. by analyzing the data described above, we get four main components representing the four categories respectively. together with the aggregative score representing the non-financial conditions, we get five factors which represent the characters of a company and construct input vectors with five dimensions. 5 classification procedure and results based on the above discussions, every sample company is represented by a five feature vector. the classifier will be constructed by fuzzy adaptive neural network. in order to check the performance of the classifier, a generalized dynamic fuzzy neural network (hereinafter referred to as “gdfnn”) introduce by wu et al.(2007), a common bp neural network and logistical regression classifiers are used to process the data respectively. a method name round-robin will be used to extract training samples and checking samples. during every round, n-1 samples are used as training samples and the remaining one is used as checking sample. the process will be repeated until all samples are checked once. the performance of classification will be evaluated by roc methodology. the experiments were run on matlab7.0. the results are shown as fig.2,fig.3 and table 3. by using fan classifier, the roc curve line area(az) reached 0.959, and corresponding true default rate and false default rate were 89.6% and 11.2%,respectively, and the overall rate of accuracy was 89.2%. another classifier is constructed by gdfnn, a general dynamic fuzzy neural network. the network is constructed based on rbf network and had four layers. the first layer is input layer. the second layer is membership function layer. the third layer is if-then rule layer and the last layer is normalizing layer. the gdfnn is equivalent to tsk fuzzy logic system. the classifier is trained 24 zhichun xu:empirical analysis of credit risk measure using fuzzy adaptive neural-networks. . . like the above fan classifier. and the result is shown as fig.2 either. the roc curve line area(az) reach 0.893, and the overall rate of accuracy is 81.5%. the classifier bp haven’t carried out fuzzy computation, and az=0.714, the rate of accuracy is 64.5%. the result of logistic regression classifier is listed as fig.2 roc curve created by fan,gdfnn and bp classifier fig.3 roc curve created by fan table 3. one can find that the rate of accuracy was lower than fan and gdfnn, but higher than bp. the above four classifiers are tested by financial information and non-financial information. in order to test the effect of non-financial information, we build up a contrast experiment. the classifier is fan, and the input vectors are conditions advances in systems science and applications (2012) vol.12 no.1 25 including non-financial and not including non-financial condition. the result is shown as roc curves. the higher curve is for the input with non-financial conditions just as fig.1, and the lower curve line is for the input without non-financial conditions. the az is equal to 0.959 and 0.757 respectively. the results show that non-financial conditions can raise the rate of accuracy of default. table 3 classification table predicted classifier observed group percentage correct .00 1.00 group 0 28 8 77.7 logistic regression group 1 4 25 86.2 overall percentage 81.5 group 0 32 4 88.8 fan group 1 3 26 89.2 overall percentage 89.2 group 0 30 6 83.3 gdff group 1 6 23 79.3 overall percentage 81.5 group 0 25 11 69.4 bp group 1 12 17 58.6 overall percentage 64.5 6 conclusion small and medium-sized enterprises have uniquely financial features, models and methods to measure their credit risk need special treatments. in this paper, we have presented that non-financial conditions are helpful to produce more accurate evaluation results. in order to overcome the vague of non-financial and sharp boundaries of financial conditions, fuzzy approaches and a nonlinear classifier fuzzy adaptive neural network are conducted. our numerical results demonstrated that fuzzy adaptive neural network is ideally suited for the modeling of problems such as credit rating or other financial systems when the variables or even the problem are vaguely defined. in addition, we showed the learning ability can improve the initially represented approximate model as more data become available. references [1] r.o. edmister. (1988), “combining human credit analysis and numerical credit scoring for business failure prediction”, akron business economic re26 zhichun xu:empirical analysis of credit risk measure using fuzzy adaptive neural-networks. . . view, vol.19, no.3, pp.6-14. [2] y. jiao, y.r. syau, and s.e lee. (2007), “modeling credit rating by fuzzy adaptive network”, mathematical and computer modeling, vol.45, pp.717731. [3] h.j.rommelfanger. (1999), “fuzzy logic based systems fro checking credit solvency of small business firms”, soft computing in financial engineering, pp.371-387. [4] r.weber. (1999), “applications of fuzzy logic for credit worthiness evaluation”, soft computing in financial engineering, pp.388-401. [5] y.r. syau, h.t. hsieh, e.s. lee. (2001), “fuzzy numbers in the credit rating of enterprise financial conditions”, review of quatitative finance and accounting, vol.45, pp.351-360. [6] l.h. chen, t.w.chiou. (1999), “a fuzzy credit rating approach for commercial loans, a taiwan case, omega international”, journal of management science, vol.27, pp.407-419. corresponding author zhichun xu can be contacted at: spring0267@yeah.net advances in systems science and application(2015) vol.15 no.1 37-44 a subspace multi-relation model for the expression of system complexity dongyun yi1,2, yangyang liu2, xiaojie wang2, zheng xie2, xianggui qu3, chengli zhao2 and chengping hou2 1 national key laboratory of high performance computing, changsha, hunan, 410073, china; 2 department of mathematics and system science, national university of defense technology, changsha, hunan, 410073, china; 3 department of mathematics and statistics, oakland university,rochester, michigan, 48309, usa. abstract modeling expression of complexity is one of the frontier research fields in system science. classical graphical models characterize two-relations but lack of the ability of depicting multi-relations among objects in a system. in this paper, a multi-relation expression model is proposed based the subspaces of the attribute space of objects in a system. properties of the multi-relation model are analyzed and an algorithm is provided to find the multi-relation subspaces in a data set. the relationship between the multi-relation model and a hyper graphical model is also discussed. keywords complexity; relation; subspace; hyper-graph. 1 introduction there are two types of research in studying the complexity of a system [1-3], one considers how to characterize and compute the overall complexity of a system [4]. boltzmann entropy in statistical physics is such a representative which gives an overall measure for the system status. the other type of research describes the interactive relationships among various objects in a system. graphical models from graph theory or complex network are typical representatives [5]. a graphical model regards the objects in a system as vertices and relationships among objects as edges that can be undirected, directed or weighted. graphical models have a lot of applications in the analysis of e-mail data, communication data, web data, biological data, financial data, social networks, actor networks, etc. [6-8]. however, graphical models only depict two-relation between objects, lack of the description of multi-relation in a system in practice. recently, some research articles discuss the multi-relation using hyper-graph models [9-13], but no rigorous definitions and generating methods of hyper-graphs from data are given. in this paper, we propose a subspace multi-relation model to characterize the this research is supported by the natural science foundation of china(no.61473302), the basic research project of nudt and the open fund from hpcl. 38 dongyun yi : a subspace multi-relation model for the expression of system complexity different types of relationships in the attribute space of objects. in this model, each relationship exists in its own subspace, which is a part of the attribute space. the traditional graph model is a special case of our model, which can be called one-dimension relation model.based on the subspace clustering principle [14-16], we give the algorithm to find subspace multi-relation from the object attribute space, which is helpful to understand the complexity of data from multiple perspectives. real data are used to verify the validity of our model. we discuss the vector hyper-graph models at the end. 2 methods 2.1 definitions of subspace multi-relation in this section, we define the subspace multi-relation model. definition 2.1 if m objects have a common property of p , we call the m objects as m-relation. in the examples provided by tables 1 and 2, there are three events and three individuals. property p means that there are individuals appearing in the same event. in table 1, any two individuals appear simultaneously in one event, and therefore there exists a two-relation between any two individuals, which can be expressed by the fully connected network model. in table 2, we notice that three individuals simultaneously appear in the event a. individuals 1, 3 appear in the event b and individuals 2, 3 appear in the event c. thus there are one threerelation and two two-relations by definition 2.1. the difference between table 2 and table 1 shows that: firstly, three-relation cannot be obtained from tworelation; secondly, the definition of the multi-relation needs consider the subspace arising relation. thus we should make some modification to definition 2.1. table 1 any two individuals appear simultaneously in one event, and hence, two-relations between any two of them exist. event a event b event c individual 1 1 1 0 individual 2 1 0 1 individual 3 0 1 1 table 2 there are three individuals appearing in the event a at the same time, which indicates that a three-relation exists. event a event b event c individual 1 1 1 0 individual 2 1 0 1 individual 3 1 1 1 establishing relations among objects needs to consider the object attribute space. for example, a person can be described by an attribute space including gender, age, education, hobby, strength, event, and so on; an email attribute space includes sender, recipient, delivery time, subject, body, attachments, and people names that appear in the mail, and so on. hence each object can be expressed by a numerical vector in its attribute space. assuming there are m objects denoted by xi, i = 1, 2, · · · ,m, of which attribute space is n-dimensional. then each object can be represented as a n-dimensional vector xi = (xi1, · · · , xin)t ∈ rn, i = advances in systems science and application(2015) vol.15 no.1 39 1, 2, · · · ,m. in order to be able to obtain the multi-relations among objects from the attribute space, we give the definition of subspace multi-relation. definition 2.2 suppose that rn = x1 × x2 × · · · × xn,and if there exists a sub-sequence denoted by i1 < i2 < · · · < ik, aij ⊂ xij is a subset of xij , j = 1, 2, · · · , k,which makes the subset b = ai1 × ai2 × · · · × aik in the k-dimension subspace xi1 ×xi2 × · · · ×xik has a common property denoted by p . then we call set b as the k-dimension subspace m-relation, where m is number of element in set b. definition 2.2 shows that the relation not only means that the objects in set b have a common property, but also tells us that the relation occurs in which attribute space. as defined, the multi-relation in table 1 can be written as (x1, x2;a), (x1, x3;b), (x2, x3;c), and the multi-relation in table 2 can be written as (x1, x2, x3;a), (x1, x3;b), (x2, x3;c). therefore there are three one-dimension subspace two-relations in table 1, and two one-dimension subspace two-relations and one one-dimension subspace three-relation in table 2. we have defined the kdimension subspace multi-relation. how can we compute subspace multi-relations in real data? the following definition 2.3 provides a solution. definition 2.3 given ϵ > 0, let rn = x1 × x2 × · · · × xn, if there is a subsequence of 1, 2, · · · , n denoted by i1, i2 < · · · < ik, aij ⊂ xij is a point set of xij , which makes the subset b = ai1 × ai2 × · · · × aik in xi1 ×xi2 × · · · ×xik be covered by the hypercube u(ϵ), ϵ > 0 in rk, then the subset b is called the k-dimension subspace m-object relation, m is the number of element in subset b. note: in definition 2.3, u(ϵ) refers to the k-dimensional hypercube of side length 2ϵ in rk, where ϵ > 0 depends on the specific situation. for example, if we define u(ϵ) = (−ϵ, ϵ), ∀ϵ > 0 in the space of r1, then, there are one one-dimension subspace three-relation and two one-dimension subspace two-relations in table 2 by simple calculation. 2.2 properties of subspace multi-relation theorem 2.1 subspace multi-relation is of downward compatibility. that is to say, if b is a k-dimension subspace m-relation, then for any k1,m1, 0 < k1 < k, 0 < m1 < m, b is also k1-dimension subspace m1-relation. proof. let b = {xi|i = 1, 2, · · · ,m} has the k-dimension m-relation. given ϵ > 0, then there exists a positive integer k,rk ⊂ rn, which makes the projection xi|rk = (xi1 , xi2 , · · · , xik)t of xi covered by the hypercube of u(ϵ), ϵ > 0. thus for any subspace rk 1 with rk 1 ⊂ rk, and any subset {xi|i = i1, · · · , im1}, ∥ xi − xj ∥rk1≤∥ xi − xj ∥rk< ϵ. so b is k1-dimension subspace m1-relation. 40 dongyun yi : a subspace multi-relation model for the expression of system complexity theorem 2.2 subspace multi-relation is of upward aggregation. assuming that all one-dimension subspace multi-relation in the set of objects s denoted by si, i = 1, 2, · · · ,m . if si1i2···ik = si1 ∩si2 ∩· · ·∩sik ̸= ∅, 1 ≤ i1 < i2 < · · · < ik ≤ m , then si1i2···ik is k-dimension subspace mi1i2···ik-relation, where mi1i2···ik is the number of element in set (si1i2···ik . proof. for any xi, xj ∈ si1i2···ik , xi = (xi1 , · · · , xik), xj = (xj1 , · · · , xjk), to eliminate the influence of the different subspace calculation, we use the manhattan distance which is normalized on dimension as following: drk(xi, xj) = k∑ s=1 |xis − xjs |/k as xi, xj ∈ si1i2···ik ,then |xis − xjs | < ϵ,thus drk(xi, xj) = ∑k s=1 |xis − xjs |/k < ϵ note: here we assume that the ϵ > 0 is uniform on different subspace in the proof, but the ϵ > 0 can be different according to the actual problems. 3 algorithm we have given the definitions of subspace multi-relation and discussed its properties. in this section we will discuss the searching algorithm of subspace multirelation and give some practical examples. the subspace multi-relation is downward compatibility and upward aggregation, which allows us to adopt a bottom-up searching algorithm to obtain subspace multi-relation from low-dimension to high-dimension. firstly, multirelation is searched in one-dimension subspace. secondly any two of one-dimension subspace multi-relation are used to compute intersection, and which being not empty is to form two-dimension subspace multi-relation. thirdly, three-dimension subspace multi-relation could be found by computing the intersection of any two of two-dimension subspace relations. finally, based on this process all the existing subspace multi-relations in attribute space will be found. this method is capable of avoiding the effect of the high-dimension disaster problem which is encountered during the direct search in the high-dimension space, and hence can greatly improve the computational efficiency. assume that there are m objects xi, i = 1, 2, · · · ,m, which are expressed as xi = (xi1 , · · · , xin)t ∈ rn, i = 1, 2, · · · ,m in the attribute space. here we want to find the complex relations among the objects in attribute space based on subspace multi-relation model. the main steps of subspace multi-relation searching algorithm are given as follow: step 1: search one-dimension subspace multi-relation. let j = 1, according to some clustering methods, we cluster the projection {xij , i = 1, 2, · · · ,m} in advances in systems science and application(2015) vol.15 no.1 41 one-dimension subspace rj with each cluster covered by an interval of length 2ϵ. the result is denoted by cj(i), i = 1, 2, · · · ,mj , and let sj(i) = {x : xij ∈ cj(i)}, i = 1, 2, · · · ,mj , which is belong to rn. j = j + 1 until j = n. step 2: search two-dimension subspace multi-relation. if sj1j2(i1i2) = sj1i1 ∩ sj2i2 ̸= ∅ let mj1j2 (i1i2) = num(sj1i1 ∩ sj2i2 ) then there exists a two dimension mj1j2 (i1i2) -relation,1 ≤ j1 < j2 ≤ n, 1 ≤ i1 ≤ mj1 , 1 ≤ i2 ≤ mj2 , i1 < i2. step 3: similarly, we can search for higher dimension subspace multi-relations. k=3; if sj1j2...jki1i2...ik = sj1i1 ∩ sj2i2 ∩ ... ∩ sjkik ̸= ∅ 1 < i1 < i2 < · · · < ik ≤ m, 1 ≤ it ≤ mi, t = 1, 2 · ··, k then we say sj1j2···jk(i1i2···ik) = sj1i1 ∩ sj2i2 ∩ · · · ∩ sjkik has k-dimension mi1i2···ik -relation, where mi1i2···ik = num(sj1j2···jk(i1i2···ik)) k=k+1,if all of the intersection is empty, then stop searching. note 1: the clustering method of one-dimension space above can be suitably selected according to actual data such as k-means method or hierarchical clustering method; note 2: when dealing with big data, in order to control the number of subspace multi-relation, the number n of objects in which is asked to be greater than a given positive integer that is used to characterize how many objects are interesting, and parameter ϵ refers to the tightness of relation. 4 examples 4.1 experiment 1 in the department of mathematics of some college, the data set consists of 84 staff vertices, where each vertex has a total of 52 attributes including gender, education, job title, hobby, physical health, etc. the searching algorithm given above is used to find subspace multi-relation in data. the parameters ϵ is set to 0.5 and n (the number of staffs in subspace multi-relation) is 5. the total 125 subspace multi-relations are found. by use of traditional graph model we can only obtain fully connective graph which lost a lot of information 42 dongyun yi : a subspace multi-relation model for the expression of system complexity in data. that is because classical two-relation just indicate whether two objects have relation, but do pay attention to neither the simultaneously relation of how many objects nor in which subspace the relation happened. the more complex relation that exists in actual data can be expressed by the subspace multi-relation model, which can not only tell us how many objects have relation but also point out the attribute space which the relation exists in. in example given above, staffs link together differently because of schoolfellows, countrymen, or the same research field and so on. 4.2 experiment 2 enron was one of the most important companies in the u.s. energy industry. in 2001, the accounting fraud scandal of the enron was exposed. then the federal energy planning board took e-mail communications between enron employees posted online [17] for publicity and academic research. here, we want to discover some interesting groups based on these email. from the body content of those emails which happened before the fraud scandal we extract 2062 different people as vertices and set 144 staff’s mailbox as the attributes of vertices. let ϵ be 0.5 and n be 60, we finally find 275 subspace multi-relations by our searching algorithm. it is very interesting that the highest dimension of subspace in all subspace multi-relation is 8 with 86 mailboxes, referring to 8 people including chief operating officer, director of risk management, vice president and chief, online president of enron, enron ceo in north america, and other two important figures, who all are suspected of fraud, which tell us maybe the event come from people’s talking. 5 subspace multi-relation and vector hyper-graph as two-dimension relation corresponds to the graph model, subspace multirelation corresponds to the hyper-graph model. recently studies on hyper-graph model have mainly focused on how to expand the existed properties in graph theory into hyper-graph, such as hyper-path, hyper-chain, graph partitioning and so on [9-13]. furthermore, hyper-graph models do not indicate that the multi-relation exists in different subspaces. here we propose the concept of vector hyper-graph corresponding to subspace multi-relation. definition 5.1 let x = {x1, x2, · · · , xs}, xi ∈ rn be a finite set. a vector hypergraph of x is denoted by g = (e1,rk1), · · · , (ep,rkp) where ei is the finite collection of subset of x, such that: (1)ei ̸= θ, i = 1, 2 · ··, p advances in systems science and application(2015) vol.15 no.1 43 (2) p ∪ i=1 ei = x,rki < rn, 1 ≤ ki ≤ n, i = 1, 2 · · · p in the vector hyper-graph g, xi is called a vertex, (ei,rk i ) is called vector hyperedge. there exists a natural connection between the vector hyper-graph and subspace multi-relation. assume that a vector hyper-graph x has a vector hyperedge (ei,rki) = (xi1 , xi2 , · · · , ximi , rki), then the objects xi1 , xi2 , · · · , ximi has multi-relation on subspace rki , which means ki-dimension mi-relation, and vice versa. hence we can use the vector hyper-graph to express subspace multi-relation and study the complex relation in system. 6 conclusions we discuss the method to express the complexity of a system by a subspace multi-relation model. based on the attribute space of objects, we establish the rigorous definition of subspace multi-relation. the traditional graph model or complex network is a special case which is equivalent to the one-dimension tworelation. moreover, an algorithm is given to search for the subspace multi-relation and its effectiveness is verified by real data. finally, vector hyper-graphs related to multi-relation models are discussed. the equivalence of the subspace multirelation and the vector hyper-graph is pointed out, which provides a new method to study system complexity through vector hyper-graphs. references [1] nicholas w. watkins and mervyn p. freeman. (2008), “natural complexity”, science, vol.320, no.5874, pp.323-324. [2] yi lin. (2008), systemic yoyos:some impacts of the second dimension, crc press. [3] pavel kabat. (2012), “systems science for policy evaluation”. science, vol.336, no.6087, pp.1398. [4] george sugihara, robert may, hao ye, chih-hao hsieh, ethan deyle, michael fogarty, and stephan munch. (2012), “detecting causality in complex ecosystems”, science, vol.338, no.6106, pp.496-500. [5] j. leskovec, j. kleinberg and c and faloutsos. (2005), “graphs over time: densification laws, shrinking diameters and possible explanations”, acm sigkdd international conference on knowledge discovery and data mining (kdd). [6] m.e.j. newman. (2003), “the structure and function of complex networks”, siam rev, vol.45, no.2, pp.167-256. 44 dongyun yi : a subspace multi-relation model for the expression of system complexity [7] s. boccaletti, v. latora, y. moreno, m. chavez and d.-u. huang. (2006), “complex networks: structure and dynamics”, phys. rep, vol.424, no.4, pp.175-308. [8] l. da, f. costa, f.a. rodrigues, g. travieso and p.r.u. boas. (2006), “characterization of complex networks: a survey of measurements”, adv. phys, vol.56, no.1, pp.167-242. [9] claude berge. (1989), hypergraphs: combinatorics of finite sets, northholland. [10] d.s.johnson and h.o.pollak. (2006), “hypergraph planarity and the complexity of drawing venn diagrams”, journal of graph theory, vol.11, no.3, pp.309-325. [11] vitaly i.voloshin. (2009), introduction to graph and hypergraph theory, nova science publishers, inc. [12] xie zheng, yi dong-yun, ouyang zhen-zheng and li dong. (2012), “hyperedge communities and modularity reveal structure for documents”, china. phys. lett., vol.29, no.3. [13] dominique lasalle and george karypis. (2013), “multi-threaded graph partitioning”, 27th ieee international parallel & distributed processing symposium. [14] charu c.aggarwal, joel l.wolf, philip s.yu, cecilia m. procopiuc and jong soo park. (1999), “fast algorithms for projected clustering”, acm sigmod record, new york, vol.28, no.2, pp.61-72. [15] rakesh agrawal, johannes gehrke, dimitrios gunopulos and prabhakar raghavan. (2005), “automatic subspace clustering of high dimensional data”, data mining and knowledge discovery, vol.11, no.1, pp.5-33. [16] kriegel hans-peter, kroger peer and zimek arthur. (2009), “clustering high-dimensional data: a survey on subspace clustering, pattern-based clustering, and correlation clustering”, acm transactions on knowledge discovery from data, nwe york, vol.3, no.1, pp.1-58. [17] enron email dataset, http://www.cs.cmu.edu/ enron/ corresponding author dongyun yi: dongyunyi@sina.com advances in systems science and application (2016) vol.16 no.3 21-32 survey on secured password authentication for iot vanitha m school of information and technology, vit university, vellore 632014, tamilnadu,india abstract password is easy and mostly used method to provide security and authentication. now in iot environment many devices like rfid, smartcard, and wireless sensor devices use passwords for security. but passwords are susceptible to many attacks like loss of password, eavesdropping, forgetting passwords, etc., by saving passwords in a single server it is more prone to loss of password if the server is compromised. in order to avoid this it is better to divide and save passwords in multiple server and this helps us to overcome loss of password on servers. this paper has used light weight cryptographic algorithm to secure these passwords in the network because this type of algorithm occupies less space and the execution speed will be high. some of the recent light weight cryptographic algorithms are humming bird, present, tea, xxtea and humming bird which was surveyed by eisenbarth et al.[1]. these are specially designed for smart cards, rfid etc. this paper has implemented an algorithm that divides and saves password on two servers and use present algorithm on one server and use humming bird algorithm on another server for providing the efficient password security. so our model will be more secured than any other existing works. keywords rfid; smart card; iot; present; humming bird 1 introduction currently, internet of things (iot) is a communication paradigm that predicts a near future, in which the objects of normal life will be prepared with microcontrollers for digital communication that will make them able to communicate with one another and with the users, becoming an essential part of the internet discussed by l. atzori et al.[2]. the iot aims at making the internet even more pervasive by enabling easy access and providing good security. this model finds application in different domains, such as home automation, medical aids ,industrial automation, elderly assistance, mobile healthcare, traffic management, intelligent energy management and smart grids, automotive, and many others. in order to provide security for those devices we require some sort of light weight cryptographic algorithms for the two reasons such as 1. for the low resource-devices, e.g. battery-powered devices, the cryptographic operation with a limited amount of energy. 2. lower resource devices that are the footprint of the lightweight cryptographic primitives are smaller than the conventional cryptographic ones. 22 vanitha m: survey on secured password authentication for iot previously password based authentication model transmitted the hash value of the password in the public channel which is easily accessible by the attacker. burande et al surveyed the data and said that 55 percent of the password can be cracked in 8hours using the password recovery toolkit which is commercially available[3]. the first protocol for the password based authentication is pki(public key infrastructure) model. the purpose of a pki is to validate the information being transferred in a network and to confirm the identity of the clients using passwords. this model is used by gong et al[4], the client can send the password to the server by using public key encryption. bellovin et al have proved that the password is unaffected to offline dictionary attack[5]. halevi et al given the proofs for pki model with the formal definitions[6]. the second protocol for the password authentication is password-only model, bellare et al[7], boyko et al[8] and yang et al[9] have introduced a method called encrypted key exchange. the password is used as a secret key for key exchange. they were the first to give password only protocol which is more secure under cryptographic assumption. but consider the scenario that all the passwords are stored in the single server; if that server is compromised due to attacks the password will be disclosed. the overcome this issue , the third protocol called two server password only model were introduced ,where two servers cooperate , even if one server is compromised the attacker cannot act like a client with the information from the compromised server. katz et al protocol called koy protocol can run in parallel and generates secret session key between client and the two servers[10]. goldreich et al discussed about the session key generation[11]. due to denial of service attack if one of the servers shuts down another server can continue to provide services to the clients. all the above mentioned protocols two servers share the random password pw1 and pw2 = pw. yi et al proposed a new symmetric solution for the password only authenticated key exchange (pake)[12], one server s1 with an encryption of the password pk1 and another server s2 with the encryption of the password pk2, where pk1 and pk2 are the encryption keys of s1 and s2. they have concluded that their protocol is secure against passive and active attacks even if one of the servers is compromised. this paper has propose the new solution for the two servers pake. current works considered diffie-hellman key exchange protocol for the clients to establish a shared key over a communication channel. another protocol they have used is elgamal encryption which consists of encryption, decryption and key generation algorithms. we use light weight cryptographic algorithm which is specifically designed for smart cards, rfid tags, and wireless sensor devices. password is advances in systems science and application (2016) vol.16 no.3 23 divided in to two halves, one in server1 and second in server2. both the servers are not aware of each others encryption keys. security analysis has shown that our protocol is secure against the attacks like smart card loss attack, forgery attack and dos attack. moreover proposed work have improved the efficiency of the system by reducing the execution speed as well as the memory occupied by the algorithm is also reduced. the paper is organized as follows, section 2 introduces related work, and light weight cryptographic algorithm named humming bird algorithm in section 3 and section 4 discusses the proposed work, section 5 shows the security analysis and last section deliberates the conclusion. 2 related work password-authenticated key exchange is an authentication method where a client and a server who share a password and verify each other with that password and both will agree on a cryptographic key discussed by katz et al[13]. passwords which are required to verify the clients are stored on a particular server. if the server is cooperated, due to some horrible operations like hacking or fixing a trojan horse, passwords which are stored in the server gets discovered. to overcome this problem, ford et al proposed the threshold password authenticated key exchange with n servers[14]. di raimando et al proposed a protocol[15], which requires 1/3 of the servers to be compromised. brainard et al developed the first two server protocol with pki based setting and implemented using the public key techniques such as ssl[16]. katz et al proposed a koy protocol with the proof of security[10]. here two servers cooperate to authenticate a client and if one server is cooperated, the hacker still cannot act as a client with the evidence from the granted server. this protocol was based on katz-ostrovsky-yung protocol called koy protocol proposed by katz[17]. in this protocol the client randomly choose a password pwc, and two servers a and b are delivered with random password parts pwc1 and pwc2 where pwc1+pwc2=pwc. at higher level this protocol can be seen as implementation of two koy protocol, one between client and server a where server b helps for confirmation and one between client and server b where server a helps for confirmation. koy protocol is symmetric in the sense two servers correspondingly contribute to the authentication and key exchange. this protocol is secure under passive adversaries but everyone roughly performs twice work as koy protocol. in order to be secure from active attacks the user work remains same but the servers work increases by 2-4 times. the advantage of koy protocol is its structure and its disadvantage is ineffectiveness of practical use. subsequently in 2005, yang et al[18] system built on brainard et al.[16] work, yang et al suggested an asymmetric setting, where a service server, interacts with the client and a back-end server called control server. control 24 vanitha m: survey on secured password authentication for iot server helps service server with the authentication, and only service server and the client decide on a secret session key in the end. they suggested a pki based asymmetric two-server pake protocol in 2005 and several pake protocols has been proposed. recently in 2013, yi et al proposed a two server password-authenticated key exchange (pake) will share a password between the client and two servers for authentication[12]. so it is difficult to hack both the servers because if one server is compromised , the attacker wont pretend to be the client on other server. they proposed a solution for two-server pake, where the client found a different cryptographic keys with the two servers and it runs in parallel and is more efficient. in their protocol they have provedh(k1, 1) ⊕ b1 = h(k1, 1) ⊕ (h(k1, 0) ⊕ h1) = h1, the server1 accepts the message m6 and computes the secret session key sk1 = sk1 because k1 = k1. also they proved that their protocol is secured against the active and passive attacks in case that one of the two servers is compromised. 3 preliminaries 3.1 present algorithm present is a lightweight block cipher designed by bogdanov et al[19], the algorithm is notable for its compact size i.e., 2.5 times smaller than aes algorithm. light weight algorithms are suitable for constrained environments such as rfid tags, smart cards and sensor networks. following pseudocode shows that each of the 31 rounds consists of an xor operation to familiarize a round key ki for 1 ≤ i ≤ 32, where k32 is used for post-whitening, a linear bit wise permutation and substitution layer. it uses a single 4-bit s-box , which is applied 16 times in parallel in each round. genroundkeys() for i = 1 to 31 do arkey(state,ki) sblayer(state) player(state) end for arkey(state,k32) in applications that claim efficient use of space, the block cipher will be implemented as encryption-only. in this way it can be used for challenge-response authentication protocols, it could be used for both encryption and decryption by using the counter mode. taking such considerations into account so decided to make present a 64-bit block cipher. it supports both encryption and decryption is smaller than an encryption-only aes and produces an ultra-lightweight solution. the encryption sub keys can be calculated on-the-fly . advances in systems science and application (2016) vol.16 no.3 25 differential cryptanalysis.the case of differential cryptanalysis is captured by the following theorem discussed by biham[20]. theorem 1.characteristics that involve 10 sboxes over 5 rounds. the following two-rounds involves two s-boxes per round and holds with probability 2−25 over five rounds. ∆ = 0000000000000011 → 0000000000030003 → 0000000000000011 = ∆ linear cryptanalysis.the case of the linear cryptanalysis of present algorithm is handled by wang[21] in the following theorem. theorem 2.let ε4r be the maximal bias of a linear approximation of four rounds of present.then ε4r ≤ 1 27 . the theorem is formally proved in a. bogdanov1 et al[19] paper, and it is used directly to bound the maximal bias of a 28-round linear approximation by 26 × ε4r = 26 × (2−7)7 = 2−43 therefore a cryptanalyst need only approximate 28 rounds in present to mount a key retrieval attack, linear cryptanalysis of the cipher would require of the order of 284 known plaintext/cipher texts. such data requirements exceed the available text. 3.2 humming bird algorithm hummingbird, which is recently designed by smith et al[2] is an ultra-lightweight cryptographic basic for encryption and authentication in severely resource-constrained devices like passive rfid tags. it is an elegant combination of a block cipher and stream cipher with a 16-bit block size, 256-bit key size, and 80-bit internal state. the size of the key and the internal state of hummingbird provide a security level which is adequate for many rfid applications. moreover, hummingbird is resistant to the most common attacks to block ciphers and stream ciphers including differential and linear cryptanalysis, structure attacks, algebraic attacks, birthday attacks and cube attacks. hummingbird encryption process differential cryptanalysis: here ek(x) denote the encryption function of hummingbird with 256-bit key k. recall that ek(x), denotes the 16-bit block cipher encryption . then ek(x) is the composition of four ek(x). for a function f (x) from fm 2 to fm 2 , the differential between f (x) and f (x + a), where + is the bit-wise addition by df (a, b), is given by df (a, b) = |x|f (x) + f (x+ a) = b, xεfm 2 engels et al proved that, the differential of ek(x) has the same upper bound as ek(x), the block cipher component in ek[22]. they have tested the reduced version of hummingbird for more instances of different pairs. from those experimental results, the standard differential cryptanalysis method is not valid for 26 vanitha m: survey on secured password authentication for iot table 1 add caption algorithm : encryption process input: a 16-bit plaintext pti and four rotors ria(i = 1, 2, 3, 4) output: a 16-bit cipher text cti [block encryption] 1: v 12a = esr1(pti + nr1a) 2: v 23a = esr2(v 12a + nr2a) 3: v 34a = esr3(v 23a + nr3a) 4: cti = esr4(v 34a + nr4a) [internal state updating] 5: lfsra+1 ← lfsra 6: r1a+1 = r1a + nv 34a 7: r3a+1 = r3a + nv 23a + nlfsra+1 8: r4a+1 = r4a + nv 12t+ nr1a+1 9: r2a+1 = r2a + nv 12t+ nr4a+1 10:return cti hummingbird with time complexity. linear cryptanalysis: for the linear cryptanalysis of ek(x), consider |ek(a, b)|, the absolute value of the walsh transform of ek(x), where ek(a, b) = σx∈f 16 2 − 1, a, b ∈ f 16 2 , a ̸= 0 where < x, y > is the inner product of two binary vectors x and y. the absolute value of the walsh transform of encryption function could be limited by the square root of 216. so, hummingbird is resistant to linear cryptanalysis attack in iot applications. 4 novel protocol for two-server password only authentication in this work , used two different light weight cryptographic algorithms for encrypting the passwords in two servers. algorithms such as present and humming bird help to encrypt the password and store it in the servers. two servers jointly work to authenticate the client and offer services proposed by yang et al[23]. client broadcast the message to both the servers , the servers cooperate to authenticate the password. our protocol uses two different algorithms so it will be very hard for the attacker to break because even if one of the servers is compromised its very difficult to hack the other one. our model used more efficient algorithm which is specially designed for the resource constraint devices when compared to the other. the protocol runs mainly in three stages initialization, registration and authentication. fig.1 explains the working process of password authentication. advances in systems science and application (2016) vol.16 no.3 27 fig. 1 working process of proposed work 1. server1 initialises runs present and sends its secret key of server1 to client. 2. server2 initialises runs humming bird and sends its secret key of server2 to client. 3. server1s secret key and sever2s secret key are received by client. 4. client registers to server1 and sends information after encrypting using secret key of respective server to respective server. the password is separated and sent to server1 and server2. 5. id and pwd1 is sent to server1. 6. id and pwd2 is sent to server2. 7. user logins to server1 with his registered id and password. 8. server1 and server2 communicates each other and server2 receives the full password and compares with user entered password. returns success if password matches else returns failure. 4.1 initialization the server then performs the following operations. 1. the server1 s1 generates the secret key according to present algorithm. 2. the server2 s2 generates secret key according to humming bird algorithm. 4.2 registration before authentication, each client c is essential to register to both s1 and s2 through different channels. first of all, the client c accesses the secret key of both servers s1 and s2. next, the client c picks a password pwc and encrypts the password using the secret key. after that, the client c recalls the password pwc. the two secure channels are essential for all two server pake protocols, where a password is encrypted by means of two different secret keys, which are safely broadcasted to the two servers, during registration. although, the idea of private key cryptosystem, the encryption key of one server should be unfamiliar 28 vanitha m: survey on secured password authentication for iot to another server and the client needs to memorize the secret code or password just behind registration. the two servers s1 and s2 have settled on the password confirmation information of the client c during registration. the id of each client has to be unique. 4.3 authentication/login 1. the client c broadcast a request message m1 to the s1 server. the message includes the authentication information of the client and a nonce. 2. the two servers exchange messages m2 and m3 based on the authentication information gathered during the registration phase. 3. now if the client is genuine (based on password check) then the client is allowed to have access. else the request for login is denied. advantage is to achieve better performance because parallelism is used and it prevents replay attack because of nonce. 4.4 encryption 1.here the plain text is the password which is divided as per the two-server password only authentication and according to novel architecture each half of the passwords are encrypted by client using present and humming bird. 2.but the change here is the password are not sent to same receiver but are sent to different servers and they use different algorithms namely present and humming bird. the detailed process of encryption is illustrated in fig.2. fig. 2 encryption process advances in systems science and application (2016) vol.16 no.3 29 4.5 decryption here the two parts of plain text are sent to different servers which use present and humming bird algorithms and these receive the respective plain text and decrypt them using present and humming bird. the two parts of plain text are merged only during login phase. the detailed process of decryption is illustrated in fig.3. fig. 3 decryption process 5 security analysis 5.1 smart card loss attack suppose if the customer lost his/ her smart card, by accessing the storage medium the attacker can try to access the data. then the attacker may try to use this information to login to the server. in out model its very hard to hack the password because passwords are divided and stored in two different servers. our protocol is against to smart card loss attack. 30 vanitha m: survey on secured password authentication for iot 5.2 dos attack suppose an attacker has stolen the smartcard, and the attacker try to update the information. in that case our model is not vulnerable to the attacker because two different algorithms have been used. even if one server is compromised the other server wont respond. smart card will be logged if it exceeds the number of login attempt limit. therefore out protocol can withstand dos attack. 5.3 forgery attack the attacker cannot create legal login request and cannot act as an authorized user, because without knowing the users password the attacker has no way to get the exact value. 6 conclusion this project presents a novel protocol for two server password only authentication. this protocol is secure against many attacks in case if one of the servers is compromised still the intruder will not be able to gain access to the password because only half of the password will be revealed. the performance is also efficient because of the parallel computation. therefore, our scheme is more suitable, particularly for applications which need enhanced security in the iot environment since making use of efficient light weight cryptographic algorithm. references [1] t. eisenbarth, s. kumar, c. paar, a. poschmann, a. and uhsadel, l. (2007), “a survey of lightweight-cryptography implementations”, ieee design & test of computers, (6), 522-533. [2] atzori, luigi, antonio iera and giacomo morabito. (2010),“the internet of things: a survey”, computer networks, 54.15, pp. 2787-2805. [3] burande, n. and gumaste, s. v. (2013), “survey on public key encryption for two server password only authenticated key exchange”, available at http://www.schneier.com/blog/archives/2006/12/ realworld passw.html. [4] gong, l., lomas, m., needham, r. m. and saltzer, j. h. (1993), “protecting poorly chosen secrets from guessing attacks”, selected areas in communications,ieee journal on, 11(5), 648-656. [5] bellovin, s. m. and merritt, m. (1992), “encrypted key exchange: passwordbased protocols secure against dictionary attacks”, in research in security and privacy, proceedings, 1992 ieee computer society symposium , pp. 72-84. advances in systems science and application (2016) vol.16 no.3 31 [6] halevi, shai and hugo krawczyk. (1999), “public-key cryptography and password protocols”, acm transactions on information and system security (tissec) 2.3, pp. 230-268. [7] m. bellare, d. pointcheval and p. rogaway. (2000), “authenticated key exchange secure against dictionary attacks”, proc. 19th intl conf.theory and application of cryptographic techniques (eurocrypt 00), pp. 139-155. [8] v. boyko, p. mackenzie and s. patel. (2000), “provably secure passwordauthenticated key exchange using diffie-hellman”, proc. 19thintl conf. theory and application of cryptographic techniques(eurocrypt 00), pp. 156-171. [9] yang, y. and bao, f. (2010), “enabling use of single password over multiple servers in two-server model”, ieee 10th international conference on in computer and information technology (cit), 2010, vol.34, pp. 846-850. [10] katz, p. mackenzie, g. taban and v. gligor,(2005), “two-server passwordonly authenticated key exchange”, roc. applied cryptography and network security (acns 05), pp. 1-16. [11] goldreich, o. and lindell, y. (2001), “ession-key generation using human passwords only”, in advances in cryptology-crypto 2001, pp. 408-432. [12] yi, xun, san ling and huaxiong wang. (2013), “efficient two-server password-only authenticated key exchange.”, ieee transactions on parallel and distributed systems, 24.9,pp. 1773-1782. [13] katz, j. and yung, m. (2007), “scalable protocols for authenticated group key exchange”, journal of cryptology, 20(1), vol.18, pp. 85-113. [14] ford, w. and kaliski jr, b. s. (2000), “ server-assisted generation of a strong secret from a password”, in enabling technologies: infrastructure for collaborative enterprises, 2000.(wet ice 2000). procedings, ieee 9th international workshops, pp. 176-180. [15] m. di raimondo and r. gennaro. (2003), “provably secure threshold password authenticated key exchange”, proc. 22nd intl conf.theory and applications of cryptographic techniques (eurocrypt 03), pp. 507-523. [16] brainard, j. g., juels, a., kaliski, b. and szydlo, m. (2003), “ultralightweight cryptography for low-cost rfid tags: hummingbird algorithm and protocol, a new two-server approach for authentication with short secrets”,in usenix security, vol.3, pp. 201-214. 32 vanitha m: survey on secured password authentication for iot [17] jonath katz, philip mackenzie, gelareh taban and virgil gligor. (2005), two-server password only authenticated key exchange, springer, pp. 1-16. [18] y. yang, f. bao and r.h. deng,(2005), “a new architecture for authentication and key exchange using password for federated enterprise”, proc. 20th ifip intl information security conf. (sec05), pp. 95-111. [19] a. bogdanov, l. r. knudsen, g. leander, c. paar, a. poschmann, m. j. b. robshaw, y. seurin and c. vikkelsoe (2007), “present an ultra lightweight block cipher”, proc. ches, springer, pp. 450-453. [20] biham, e. (1994),“new types of cryptanalytic attacks using related keys”, journal of cryptology, 7(4), 229-246. [21] m.wang(2007), “differential cryptanalysis of present”, proc .ches, pp. 1-4. [22] engels, x. fan, g. gong, h. hu and e. m. smith (2009), “ultra-lightweight cryptography for low-cost rfid tags: hummingbird algorithm and protocol”, centre for applied cryptographic research (cacr) technical reports cacr. [23] y. yang, r.h. deng and f. bao (2006), “a practical password-based twoserver authentication and key exchange system”, ieee trans. dependable and secure computing, vol. 3, no. 2, pp. 105-114. corresponding author vanitha m can be contacted at: mvanitha@vit.ac.in. investigation on fuzzy degree criterion of electric power equipment 180-187 advances in systems science and applications (2011), vol.11, no.1-2 issn 1078-6236 international institute for general systems studies, inc. research on energy management of hev lithium ion battery wu tiezhou 1,2 , chen xueguang 1 , zhang yongfei 2 and xiao qing 2 1 department of control science & engineering, huazhong university of science & technology,wuhan 430074 2 hubei university of technology school of electrical & electronic engineering,wuhan 430068 abstract battery system is a more complicate and fragile link in the hev vehicle. the performance of battery system directly affects the whole energy and the function and cost of hev system. therefore, energy management of battery system can optimize the working state of battery and make it match better with other system. the paper mainly discusses the charging management of lithium ion battery and energy feedback of braking. results of experiment demonstrate the proposed energy management strategy has an obvious effect of saving energy and can reduce the charging time and improve efficiency. keywords energy management hev lithium ion battery energy feedback of braking 1. introduction with the rapid development of automobile industry, energy and environmental issues become even more acute. in order to solve the energy consumption and pollution caused by these two problems, hybrid electric vehicle (hev) has been rapid development [1]. storage battery can store and provide clean power, and has been widely used in hybrid electric vehicles. compared to other batteries, lithium ion battery [2] has many advantages, such as the high voltage, energy density, no memory effect and the long lifetime. although battery technology has been great progress, traditional methods such as constant voltage battery charging method, constant current method, phase-wise constant current method and the constant current constant voltage method [3] have some problems, such as charging too long, and seriousness of polarization effects, reduced cycle life of the battery, taking noting of energy feedback during the period of braking into account, so it is necessary for new battery fast charging technology [4]and new method of energy feedback during the period of braking. according to maas laws [5], to charge the battery fast, the charge current should comply with the optimal acceptable current curve, which is a negative exponential curve. in practice, it is very difficult to meet the requirement. common methods such as fast charging pulse charge method [6], intermittent charging method [7], achieve fast-charging by eliminating the polarization effects [8]. pulse charging method is an effective method of fast charging. but, it is difficult to determine frequency of pulse [9]. in this paper, the pulse charging method is adapted, advances in systems science and applications (2011), vol. 11, no. 1-2 181 using variable frequency pulse charge method to reduce the charge time and improve charging efficiency. there is a transient and frequent braking in the process of hev vehicle’s driving. it will offer much motive to vehicle if the inertia energy be accumulated. the paper will research how to recycle the maximum possible feedback braking energy without destroying the battery. 2. principles of variable frequency pulse charge method and description of algorithm 2.1 principles of variable frequency pulse charge method in accordance with the battery emf equivalent model [10], the resistance of the battery has three parts: ohmic resistence r , electrochemical polarization resistance dz and concentration polarization resistance zk , the total resistance is as follows: kdb zzrz   (1) the less the zb is, the less chemical energy converted from the electrical energy is. so in order to reduce the energy loss during charging, it is necessary to determine the optimal frequency fop of charging to reduce the battery impedance zb. the battery resistance changes as the battery soc, so the charge frequency varies with the soc. to find the optimal pulse frequency conveniently, we need to charge the battery with different frequencies of pulses, acquire the corresponding average currents, and then select the pulse frequency with maximum average current as the best pulse frequency. 2.2 algorithm description the charging of lithium-ion battery continues with a constant period tsum . a cycle of tsum can be divided into three time sub-periods: tfull for fullness detection, tseek for determining of the optimal frequency and tch for charging. according to experimental results, tsum is set as 5min. the charging process is repeated until the charging completed. tch f should be much higher than tfull + tseek. work model of variable frequency pulse charging as shown in fig. 1. initialization, t = 0 start timer whether t≥tsum judge soc>95% or vb>vsup estimate soc , measure battery voltage vb search optimal frequency f beginning charging, the frequency is f y n n end y fig.1 work model of variable frequency pulse charging the optimal frequency is determined as follows: 1) calculate the average charging current with frequency fn. if the m-th sampling value of current is dib,n(m), and the sum of the sampled currents id dacc(n),and sample value, sample design and for the currents, setting the initial = 0: ;)()()( , mdndnd nibaccacc  m=m+1 (2) after calculating the average current of m sub-samples: m nd nd acc avg )( )(  (3) 2) calculating the average current of sampling value of current at varied frequencies, n=n+1. nnmd m nd m m nibavg ...3,2,1)( 1 )( 1 ,    (4) 3) determining optimal frequency according to (5):  nnndmaxff avgnop ,...2,1)),((|  (5) 182 wu: research on energy management of hev lithium ion battery182 wu: research on energy management of hev lithium ion battery182 wu: research on energy management of hev lithium ion battery182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery wu: research on energy management of hev lithium ion batterywu: research on energy management of hev lithium ion batterywu: research on energy management of hev lithium ion batterywu: research on energy management of hev lithium ion batterywu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery ? ? advances in systems science and applications (2011), vol. 11, no. 1-2 183 3. battery group charge management with energy braking energy 3.1 determination feedback current and braking time according to mass’ law, in the charging process, the maximum acceptable charging current could be expressed as equation(6): t tt eii  (6) charge the lithium-ion battery which had released capacity oc with maxci under the constant-current charging mode, t later, the battery’s voltage reaches the charging final voltage sv , then maxci is the maximum acceptable charging current, and sv could be found in the battery instruction manual of manufacturer. t is determined as following: there are two factors: one is whether it influences the battery’s cycle life or not when charging battery with a high current in t , the other is considering the durable time of peak power rate. we choose t =2min, considering the braking time of 510 bus in wuhan are mostly within 2minutes. 3.2 braking energy feedback charging strategy first of all, we should judge battery’s charge final voltage. if battery’s voltage does not reach it, then detection feedback current ri , and the maximum acceptable charging current was 0maxci in the brake moment. under the circumstance of 0maxcr ii  , the braking energy could be feedback. when clock the braking time rt , if ttr  and 0maxcr ii  , start energy feedback; with the continuous braking, energy feedback stop if ttr  or 0maxcr ii  . strategy implementation process shown in fig. 2 braking begin sb vv  read ri 0maxcr ii  flag=0 flag=1 0maxcr ii  ttr  output charge signal,then begin charge output uncharged signal braking over n y y y y n n n y n ,0 r t 0flag calculate 0maxc i beginning time output charge signal,then begin charge fig. 2 flowchart of braking feedback charging current 4. experiments and conclusions 4.1 lithium-ion batteries variable frequency pulse charging experiment batteries are selected: the gbs-lfmp90ah square lithium-ion batteries are selected for our variable frequency pulse charging experiments, which are produced by zhejiang jia-beth green energy co., ltd. experimental method: the number of sampling value of current at varied frequencies is selected and the tsum is determined for our experiments. 182 wu: research on energy management of hev lithium ion battery 182 wu: research on energy management of hev lithium ion battery 184 wu: research on energy management of hev lithium ion battery advances in systems science and applications (2011), vol. 11, no. 1-2 185 there are 40 frequencies for sampling, which start with f1=100hz,…fn=fn-1+500hz(n=2,…49). when the cycle tsum is too short or too long, the charging time will be extended. in our experiments, the tsum is 6min. the results are shown in fig. 3 and fig. 4. fig 3 shows a charging curve of the optimal frequency. 0 2000 4000 6000 8000 10000 12000 14000 12024 48 72 96 time(min) fr eq u en cy (h z ) fig.3 optimal charging frequency as shown in fig. 4, the optimal charge frequency is effected by factors such as soc(state of charge). we have charged the batteries using constant current constant voltage charging method and have charged the same batteries at 1khz, 100 khz pulse, and then we have charged the same batteries using the method given in this paper. the curves are as follows: fig.4 flowchart of different charging mode as can be seen from fig4, the voltage/time curve for variable frequency pulse charging method is steeper than those for other methods, implies that it requires less time for charging. 4.2 comparative experiment of energy feedback we have a experiment using dongfeng’s 1500kg hev, the friction coefficient: f = 0.02, frontal area: a = 2.872m * m, drag coefficient: cd = 0.4. the experiment data show as table 1: table 1 comparative table of energy feedback parameter some feedback all feedback no feedback initial soc 0.9 0.7 0.8 end soc 0.87 0.68 0.77 energy consumer(kwh) 4.52 4.68 4.49 feedback energy(kwh) 0.31 0.87 0 energy conservation (%) 6.85 18.58 0 according to table 1, when in average circulation, while battery’s soc under 0.7, hev could feedback all of the braking energe, and could saving many energy. 5. conclusions at present, the battery management system takes insufficiently the battery energy management into account, or lacks the energy management strategy, thus causes drop of battery's state of health, affects the using and the service life of the battery. in this paper, a fast charging technology for charging pulse is proposed, whose key idea is to accurately determine the pulse frequency and to charge with variable optimal frequency, and a kind of energy feedback strategy is proposed too. experimental results show that the fast charging technology for charging pulse can significantly shorten the charging time, and can be used in hev battery management system, the energy feedback strategy can reduce energy’s consumption, they can be used in hev. references [1] li jianping. the controlling & suggestion of vehicle’s emission [j]. research & application for car.1999,5:18-23. [ch] [2] becker h., schuh h., brunarie j.. high power lithium-ion battery: first field trial results of an advanced backup system for advanced cellular networks. intelec, twenty-seventh international, 2005, 257 – 262. 186 wu: research on energy management of hev lithium ion battery advances in systems science and applications (2011), vol. 11, no. 1-2 187 [3] r. c. cope and y. podrazhansky, “the art of battery charging,” in 14thannu. battery conf. applications and advances, 1999, pp. 233–235. [4] p. m. hunter and a. h. anbuky, “vrla battery rapid charging under stress management,” ieee trans. ind. electron., vol. 50, no. 6, pp.1229–1237, dec. 2003. [5] blake dickinson, jay gill. issues and benefits with fast charging industrial batteries[j]. battery conference on applications and advances, 2000, 223-229. [6] j. zhang, j. yu, c. cha, and h. yang, “the effects of pulse charging on inner pressure and cycling characteristics of sealed ni/mh batteries,”j. power sources, vol. 136, no. 1, pp. 180–185, sep. 2004. [7] lopez, j. ,gonzalez, m. ,viera, j.c. ,blanco, c. fast-charge in lithium-ion batteries for portable applications[j]. telecommunications energy conference, 2004,19-24. [8] p.-h. cheng and c.-l. chen. high efficiency and nondissipative fast charging strategy[j]. electric power applications, 2003,539-545. [9] tristan juergens, robert f. nelson, michael a. ruderman. a new high rate, fast charge, sealed, lead acid battery[c].wescon, 1994, 230-235. [10] martin coleman, chi kwan lee, chunbo zhu, william gerard hurley. state-of-charge determination from emf voltage estimation: using impedance, terminal voltage, and current for lead-acid and lithium-ion batteries. ieee transactions on industrial electronics, 2007,54(5):2550-2557. advances in systems science and applications (2012) vol.12 no.3 258-271 a geometric interpretation of gravity theory n.n. popov dorodnitsyn computer centre of russian academy of science abstract a system of postulates which give a geometric meaning to the fundamental notions of gravity theory is introduced. the fundamental equation of geometric gravity theory is derived. the consistency of the system of postulates is shown. a series of examples adequately describing the gravitational field are considered. mathematically rigorous definitions of a black hole and dark energy are given. keywords pseudo-riemannian space, scalar curvature, metric gravity equation, spherically symmetric spatial black hole, dark energy 1 introduction hilbert’s sixth problem of axiomatizing those branches of physics in which mathematics is prevalent, posed by hilbert in 1900 among other problems, has the status of being too vague. however, in the context of some particular area of physics, such as gravity theory, this problem can be stated rigorously. the foundation of physical gravity theory is two fundamental physical concepts, of the gravitational field and of the mass of a body. in general relativity theory (grt), which is essentially relativistic gravity theory, a substantial progress in the mathematical interpretation of one of the basic concepts of the theory, namely, that of the gravitational field, has been made. from the geometrical point of view, the gravitational field is interpreted as a metric pseudo-riemannian 4space. however, grt does not provide such a precise mathematical definition of mass, or, to be more precise, the distribution density of mass. thus, the grt fundamental equation contains both a purely mathematical left-hand side, which is generated by a metric space, and a purely physical right-hand side, which is the energy-momentum tensor of the physical system under consideration. note that grt imposes no constraints on the choice of the energy-momentum tensor; this leads to the possibility of constructing unrealistic models and, thereby, provides evidence for the insufficiency of the principles on which grt is founded. our purpose in this paper is to present a complete translation of gravity theory into the language of differential geometry, in which the physical notion of the mass distribution density has a purely geometric interpretation. 2 the first postulates of mathematical gravity theory the notion of a metric space is the basis on which mathematical gravity theory is constructed. to construct the theory, it suffices to choose a pseudo-riemannian 4-space of signature (−−−+), on which a twice covariant symmetric nondegenerate tensor field gij(x1,k, x4) is defined; we have advances in systems science and applications (2012) vol.12 no.3 259 det |gij | ̸= 0, gij = gji, i, j = 1, . . . , 4 we refer to the tensor as the metric of the pseudo-riemannian space. general requirements to a metric space are as follows: (1) smoothness, i.e., the continuity of all components of the metric on the entire space and the continuous differentiability of the components up to the second order with respect to all variables almost everywhere except, possibly, on singular sets; (2) the preservation of the metric signature at each point of the space. note that the metric smoothness condition is stated in a somewhat relaxed form in order to make it possible to consider pseudo-riemannian spaces with discontinuous scalar curvature. note that the space structure and dimension chosen above are not regarded to be final. they are only sufficient for constructing geometric gravity theory; thus, we introduce them as sufficient conditions for constructing an adequate theory rather than as fundamental postulates. we proceed to state two postulates of geometric gravity theory. as mentioned above, the main physical objects of gravity theory are the gravitational field and the distribution density of the matter mass. the objective of gravity theory is determining laws governing the interaction of the gravitational field with the distribution density of gravitational mass. the first postulate of geometric gravity theory can be stated as follows. postulate 1. the gravitational field is a metric of a pseudo-riemannian space. this assertion is the basis of grt and does not need any comments. before stating the second postulate, we introduce the following notation[1-2]: γk ij is the pseudo-riemannian connection, which is defined by γk ij = −1 2 gkl( ∂glj ∂xi + ∂gli ∂xj + ∂gij ∂xl ); (1) rij is the ricci tensor, that is, rij = γk ij ∂xk − γk ik ∂xj + γk pkγ p ij − γk pjγ p ik; (2) r is the scalar curvature, that is, r = gijrij ; (3) where the summation over repeated indices is implied; and is the contravariant metric tensor. postulate 2′ (weak statement). if the mass density of the matter is nonzero at some point of the space, then so is the scalar curvature of the space at this point, 260 n.n. popov:a geometric interpretation of gravity theory and vice versa. this postulate says nothing about the form of the dependence between these scalar functions. it only indicates the existence of a relation between them. to determine this relation, we use two principles; namely, the principle of minimal action and the correspondence principle. these principles make it possible to derive the fundamental equation of geometric gravity theory and refine the statement of postulate 2. 3 the third postulate, the fundamental equation of gravity theory, and the strengthened statement of postulate 2′ according to postulate 2′, there is a dependence between the scalar function of the matter mass density and the function of the space scalar curvature. to find this dependence, we introduce a composite scalar function χ (ρ) , which depends on the matter mass density ρ in an unknown way, and a scalar curvature function r(gij), which is a composite function depending on the metric gij . the general form of the gravitational field equations can be obtained by applying the principle of minimal action. the field equations are obtained as the eulerclagrange equations under the variation of the field action. for the field action functional we take the quantity sg = ∫ ω (r + χ) √ |g|d4x, (4) where dω = √ |g|d4x is the standard volume 4-form on the pseudo-riemannian space,g = det(gij), and the integral is the whole 3-space (x1, x2, x3) and over the interval x41 ≤ x4 ≤ x42 of the time coordinate x4. postulate 3. in geometric gravity theory, the following relation holds: δsg δgij = 0. postulate 3 means that the sought dependence between the composite functions r(g) and χ(ρ) minimizes the action functional (4) with respect to the contravariance metric of the pseudo-riemannian space gij . theorem 1. if postulate 3 holds for functional (4), then rij − 1 2 rgij = 1 2 χgij . (5) proof. relation (4) implies δsg δgij = δ ∫ ωr(g) √ |g|d4x δgij + δ ∫ ω χ(ρ) √ |g|d4x δgij . advances in systems science and applications (2012) vol.12 no.3 261 the first term is the hilbert variational derivative[3] δ ∫ ωr(g) √ |g|d4x δgij = rij − r 2 gij . the second term can be represented in the form δ ∫ ω χ(ρ) √ |g|d4x δgij = ∫ ω(δχ(ρ) √ |g|+ χ(ρ)δ √ |g|)d4x δgij . taking into account the relation δ √ |g| = −1 2 √ |g|gijδgij and the fact that the variation of the scalar. function χ(ρ), which does not depend on the metric gij , vanishes, we obtain δ ∫ ω χ(ρ) √ |g|d4x δgij = −χ(ρ) 2 gij . it follows that δsg δgij = rij − r 2 gij − χ(ρ) 2 gij = 0, which completes the proof of the theorem. equation (5) implies the presence of a linear dependence between the functions r and χ. indeed, multiplying the rightand left-hand sides of equation (5) by the contravariant metric tensor gij and convolving both sides over the indices i and j, we obtain the relation χ = −r 2 . (6) according to (6), the scalar curvature r turns out to depend on the matter mass density ρ. however, we do not know the particular form of this dependence so far. to determine it, we apply the correspondence principle, which can be stated as follows: if the metric gij weakly converges to the minkowski metric ηij and, moreover, g44 = 1+ aφ c2 + o( 1 c3 ) and gij = −δij + o( 1 c3 ) for i, j = 1, . . . , 4, i ̸= j , where c−1 is a small parameter and φ is a twice differentiable function, then equation (5) must degenerate into the poisson equation for newtonian gravity theory, which is ∆φ = 4πgρ, (7) where ∆ is the laplace operator, φ is the newtonian gravitational potential, g is the gravitational constant, and c is the speed of light. simple calculations show that equation (5) does degenerate into the poisson equation (7) under the condition ρ = c2 32πg r. (8) relation (8) can be regarded as a stronger statement of postulate 2’. postulate 2 (strong statement). the matter mass distribution density in space 262 n.n. popov:a geometric interpretation of gravity theory is directly proportional to the scalar curvature of the pseudo-riemannian space: ρ = ær , where æ = c2 32πg .it follows from postulate 2 and equation (5) , based on postulate 3, that the fundamental equation of gravity theory can be represented in the form rij − 1 2 rgij = −8πg c2 ρgij , i, j = 1, . . . , 4, (9) or in the form of the system of two relations rij = r 4 gij , i, j = 1, . . . , 4, (10) r = 32πg c2 ρ the first of which is a direct consequence of postulate 3 and the second, of postulate 2. thus, physical theory of gravitational fields can be translated into the language of differential geometry. according to postulates 1 and 2, the fundamental physical concepts of gravitational theory, such as gravitational field and matter mass density, are interpreted in the geometric as the metric and the scalar curvature (up to proportionality), respectively, of a pseudo-riemannian space. 4 consistency and physical adequacy of postulates 1-3 to verify the consistency of the system of postulates stated above and the physical adequacy of these postulates, consider a model of a spherically symmetric space with a ball of radius r1 at the center of symmetry. the ball is uniformly filled with a matter of mass ρ and constant density; outside the ball, there is no matter. from the geometric point of view, we have a spherically symmetric pseudo-riemannian space with constant scalar curvature r = 32πg c2 ρ inside a ball of radius r1 and vanishing scalar curvature outside the ball. the problem is to determine the metric of such a space. hereafter, we always assume that the system units used for measuring physical quantities is chosen so that g, c = 1 . the general form of a stationary spherically symmetric metric of a pseudoriemannian 4-space in spherical coordinates t, r, θ, φ is[4] ds2 = g44(r)dt 2 + g11(r)dr 2 + g22(r)(dθ 2 + sin2θdφ2), (11) where g11 and g22 are negative unknown functions and g44 is a positive function, which depend on the variable r, and r, t ∈ (0,∞), r ∈ (0,∞), θ ∈ [0, π], φ ∈ [0, 2π). the components of metric (11) must satisfy system (10) . using relations (1) advances in systems science and applications (2012) vol.12 no.3 263 and (2), we can reduce system (10) for metric (11) to the form ( g′22 g22 )′ + 1 2 ( g′44 g44 )′ − 1 2 ( g′11 g11 )( g′22 g22 + g′44 2g44 ) + 1 2 ( g′22 g22 ) 2 + 1 4 ( g′44 g44 ) 2 = r 4 g11, 1 2 ( g′22 g11 )′ + g′22 2g11 ( g′11 2g11 + g′22 g22 + g′44 2g44 )− 1 2g11g22 (g′22) 2 − 1 = r 4 g22, (12) −1 2 ( g′44 g11 )′ − g′44 2g11 ( g′11 2g11 + g′22 g22 + g′44 2g44 ) + 1 2g11g44 (g′44) 2 = r 4 g44, where g′ij = dgij dr , r > 0 ,at r ≤ r1, and r = 0 at r > r1 . for the sake of generality, we assume that the unknown scalar function r in system (12) depends on the r coordinate. the following theorem is valid. theorem 2. if g22(0) = 0 and g44(0) < ∞, then the scalar curvature r(r) does not depend on r in the domain where it is nonzero, and system (12) has a unique solution depending on one free parameter. if, in addition,g44(0) = 1, then the solution of system (12) is unique; moreover, at r ≤ r1, where r1 < √ 12 r , it coincides with the de sitter metric[5] ds2 = (1− r 12 r2)dt2 − dr2 1− r 12r 2 − r2(dθ2 + sin2θdφ2), (13) and at r > r1, it coincides with the schwarzschild metric[6] ds2 = (1− r 12 r31 r )dt2 − dr2 1− r 12 r31 r − r2(dθ2 + sin θ2dφ2). (14) proof. system of equations (12) can be simplified by passing to the generalized spherical coordinates x1, . . . , x4,where x1 = r3 3 ,x2 = − cos θ,x3 = φ,and x4 = t. let us introduce the following new notation for the components of the metric tensor: g11 = −(3x1) 4 3 f1(x1), g22 = −f2(x1), g44 = f4(x1). in this notation, metric (11) takes the following form in the generalized spherical coordinates: ds2 = f4dx4 2 − f1dx1 2 − f2( dx2 2 1− x22 + (1− x2)dx3 2). (15) using the arbitrariness of the scaling multiplier of the coordinate x1 , we can achieve f1f2 2f4 = 1. (16) 264 n.n. popov:a geometric interpretation of gravity theory system (10) for metric (15) can be represented as −1 2 ( f ′ 1 f1 )′ + 1 2 ( f ′ 2 f2 ) 2 + 1 4 ( f ′ 1 f1 ) 2 + 1 4 ( f ′ 4 f4 ) 2 = −r 4 f1, 1 2 ( f ′ 2 f1 )′ − 1 2f1f2 (f ′ 2) 2 − 1 = −r 4 f2, (17) −1 2 ( f ′ 4 f1 )′ + 1 2f1f4 (f ′ 4) 2 = r 4 f4. condition (16) gives the additional equation f ′ 1 f1 + 2f ′ 2 f2 + f ′ 4 f4 = 0, (18) where f ′ i = dfi dxi . system (17), (18) contains four unknown functions f1 ,f2 ,f4 , and r ; only two of these functions, say f2 and r , can be regarded to be independent. this follows from relations (3) and (16). let us express the unknown functions f1 and f4 in terms of f2 and r . for this purpose, note that the third equation in system (17) can be represented as 1 2f1f4 f ′ 1f ′ 4 − 1 2 ( f ′ 4 f4 )′ = r 4 f1. (19) adding the first equation in system (17) to equation (19) and using (18) ,we obtain the relation ( f ′ 2 f2 )′ + 3 2 ( f ′ 2 f2 ) 2 = 0. twice integrating it, we obtain f2 = λ(3x1 + α) 2 3 , where λ and α are arbitrary constants of integration. by the assumption of the theorem, we have f2(0) = 0 , which implies α = 0 and f2 = λ(3x1) 2 3 . (20) we seek f ′ 4 f4 in (19) in the form f ′ 4 f4 = c(x1)f1 . substituting the last relation into equation (19) , we obtain f ′ 4 = c − 1 2 ∫ rdx1 λ2(3x1) 4 3 , where c is a constant. integrating both sides of this relation, we arrive at the formula f4 = β − c λ2(3x1) 1 3 + 1 2λ2 ∫ rdx1 (3x1) 1 3 − 1 2λ2 ∫ r (3x1) 1 3 dx1, (21) advances in systems science and applications (2012) vol.12 no.3 265 where β is constant of integration. by the assumption of the theorem, the function f4(0) is bounded; therefore, considering a solution in some neighborhood of zero, we must set c = 0 , and the formula for f4 takes the form f4 = β + ∫ r(x1)dx1 2λ2(3x1) 1 3 − 1 2λ2 ∫ r(x1) (3x1) 1 3 dx1. (22) condition (16) implies f1 = 1 λ2(3x1) 4 3 1 f4 . (23) the functions f1 ,f2, and f4 specified by (20), (22), and (23) satisfy system (17) of differential equations only if βλ3 = 1,∫ r(x1) (3x1) 1 3 dx1 = r(x1) 2 (3x1) 2 3 . (24) it follows from relation (24) that the function r(x1) must be constant in the domain where it is nonzero. thus, inside the ball of radius r1 , the components of metric (25) have the form f2 = λ(3x1) 2 3 , f4 = 1 λ3 − r 12 (3x1) 2 3 λ2 , f1 = 1 λ2(3x1) 4 3 1 f4 . (25) outside the ball, the scalar curvature vanishes, and relation (21) implies that the components of metric (15) have the form f2 = λ(3x1) 2 3 , f4 = 1 λ3 − c λ2(3x1) 1 3 . (26) these components satisfy also system (17),(18). the constant c in (26) is found from the condition that the metric components must be continuous on the entire space, which implies c = r 12r1 3. system (17), (18) has the unique solution (25), (26), which contains one free parameter λ. taking into account the additional condition g44(0) = 1 in the theorem, which implies λ = 1, and passing to the usual spherical coordinates, we see that relations (13) and (14) hold. this completes the proof of the theorem. thus, in the framework of the proposed axiomatics, we have constructed a mathematically rigorous model of a spherically symmetric space adequately describing the spherically symmetric gravitational field generated by a spherical gravitational source with constant mass density. theorem 2 are easy to extend to a more general case. 266 n.n. popov:a geometric interpretation of gravity theory 5 stationary spherically symmetric spaces with scalar curvature having finitely or countably many discontinuities of the first kind and a mathematically rigorous definition of spherically symmetric black holes consider a stationary spherically symmetric space endowed with a metric of the general form (11). suppose that the scalar curvature of this space is a piecewise smooth function r(r) with at most countably many discontinuities of the first kind at points r1, r2, . . . , rn, . . . numbered in increasing order. for such a space, the following theorem is valid. theorem 3. if the components of metric (11) satisfy the conditions g22(0) = 0 and g44(0) = 1 , then the scalar curvature of the spherically symmetric space under consideration is a piecewise constant function taking the constant values r(r) = rk at rk−1 < r ≤ rk for k = 1, . . . , n, . . ., where r0 = 0. the metric of such a space is everywhere continuous, provided that 2m(r) r < 1 for r ∈ (0,∞), and has the form ds2 = (1− 2m(r) r )dt2 − dr2 1− 2m(r) r − r2(dθ2 + sin2θdφ), (27) where m(r) = r∫ 0 ∑ k>0 rk 4 [θ(x− rk−1)− θ(x− rk)]x 2dx. the components of metric (27) satisfy system (12) of differential equations everywhere except at a finite or countable set of points r1, . . . , rn, . . . . proof. this theorem is proved by the same method as theorem 2. as in theorem 2, we show that, under the assumptions of theorem 3, the components of metric (11) in the spherical coordinate system have the form g44 = 1− 1 4r r∫ 0 r(x)x2dx, g22 = −r2, g11 = −g44 −1, (28) where r is a piecewise smooth function with at most countably many discontinuities of the first kind at points r1, r2, . . . . the functions in (28) satisfy system (12) only under the condition dr dr = 0, (29) which is an analogue of condition (24) in the proof of theorem 2. relation (29) implies that r(r) must take constant values in its domains of continuity. suppose that r(r) takes a value rk at rk−1 < r ≤ rk ; then r(r) = ∑ k=1 rk(θ(r − rk−1)− θ(r − rk)), advances in systems science and applications (2012) vol.12 no.3 267 which implies the assertion theorem 3. the components of metric (27) are everywhere continuous if 2m(r) r < 1 for r ∈ (0,∞). if 2m(r) r = 1 at some r = rg. in this case, there are two possibilities: either the spherically symmetric space is bounded by a hypersphere with radial parameter rg (if 2m(r) r > 1 at r > rg ), or this space can be extended (if 2m(r) r < 1 at r > rg ). in the latter case, the space contains a spherically symmetric body of radius rg with generally nonuniform mass distribution density, and on the boundary of this body, the metric exhibits an irregular behavior. we refer to such bodies as spherically symmetric stationary black holes. below we give the definition of the simplest stationary black hole. definition 1. a globular body of radius rg with constant mass density satisfying the condition r = 12 rg2 is called a stationary spherical black hole. the sphere of radius rg being the surface of a black hole is called its horizon level, or the schwarzschild sphere. it follows from the definition that rg = √ 12 r . note that the signature of the space inside a black hole and outside it remains invariable, which agrees with condition (2) on the metric of a pseudo-riemannian space. all components of metric (27) are continuously differentiable on the entire space except on the surface of the black hole, on which the behavior of metric (27) is irregular, namely, g11(rg) = −∞. the mathematical properties of a stationary spherically symmetric black hole in the formalism suggested here substantially differ from the properties of black holes investigated in the framework of grt[7]. if the spherically symmetric space has the same scalar curvature r at all points, then the space is the de sitter closed elliptic space[5] determined by metric (27) of the form ds2 = (1− r 12 r2)dt2 − dr2 1− r 12r 2 − r2(dθ2 + sin2θdφ) with r < √ 12 r . this space is bounded by a sphere with radial parameter r < √ 12 r . such a model corresponds to a homogeneous closed space filled with a matter with constant mass density. consider yet another parameter of the spherically symmetric space generated by a ball of radius r1 with constant mass density ρ1 and a matter with constant mass density ρ2 filling the whole space outside the ball. the scalar curvature of the space inside the ball is calculated by r1 = 32πρ1 and outside the ball, by 268 n.n. popov:a geometric interpretation of gravity theory r2 = 32πρ2. the components of metric (27) for each space have the form g44 = 1− r1 12 r2 at r < r1, g44 = 1− r1 −r2 12 r1 3 r − r2 12 r2 at r ≥ r1, and g22 = −r2, g11 = −g44 −1(r) if r1 12 r1 2 < 1. the spherically symmetric space is bounded by a sphere of radius r2, which is the least positive root of the cubic equation r3 − 12 r2 r + r1 −r2 r2 r1 3 = 0. a criterion for the transformation of the ball into a black hole is r1 12 r1 2 = 1. 6 a model of a nonstationary spherically symmetric pseudo-riemannian space with constant scalar curvature and the definition of dark energy in the preceding section, we described a stationary spherically symmetric space with discontinuous scalar curvature. here, we consider a mathematical model of a spherically symmetric pseudo-riemannian space with nonstationary fridmantype metric[8] ds2 = dx4 2 − r2(x4)(dx1 2 + sin2x1(x2 2 + sin2x2dx3 2)), (30) where r is the curvature radius of the three-dimensional hypersphere, which depends on the parameter x4; x1 ∈ (0, 2π); x2 ∈ (0, π); x3 ∈ (0, 2π); and x4 ∈ (0,∞). the following theorem is valid. theorem 4. system (10) for metric (30) reduces to the single second-order nonlinear differential equation r d2r dx42 − ( dr dx4 ) 2 − 1 = 0 for the unknown function r(x4). this equation has a real solution r = r0 cosh x4 r0 ,where r0 is a some constant. the scalar curvature of the space is everywhere constant and has the form r = 12 r02 . proof. according to (30), the nonzero components of the covariant metric tensor gij have the form g11 = −r2, g22 = −r2sin2x1, g33 = −r2sin2x1sin 2x2, g44 = 1. (31) advances in systems science and applications (2012) vol.12 no.3 269 the components of the contravariant metric tensor gij have the form g11 = −r−2, g22 = − 1 r2sin2x1 , g33 = − 1 r2sin2x1sin 2x2 , g44 = 1. (32) substituting the metric components (31) and (32) into relation (1), we find all nonzero elements of the pseudo-riemannian connection; these are γ22 1 = − sinx1 cosx1,γ33 1 = − sinx1 cosx1sin 2x2,γ14 1 = 1 r dr dx4 , γ12 2 = cotx1,γ33 2 = − sinx2 cosx2,γ24 2 = 1 r dr dx4 , (33) γ13 3 = cotx1,γ23 3 = cotx2,γ34 3 = 1 r dr dx4 , γ11 4 = r dr dx4 ,γ22 4 = r dr dx4 sin2x1,γ33 4 = r dr dx4 sin2x1sin 2x2. using the components (33) of the connection and applying formula (2), we obtain all nonzero components of the ricci tensor rij : r11 = −2( dr dx4 ) 2 − r d2r dx42 − 2, r22 = sin2x1r11, (34) r33 = sin2x1sin 2x2r11, r44 = 3 r d2r dx42 . it follows from (10), (32), and (34) that the scalar curvature is r = 6 r d2r dx42 + 6 r2 ( dr dx4 ) 2 + 6 r2 . (35) by virtue of (31), (34), and (35), system (10) degenerates into a single equation of the form r d2r dx42 − ( dr dx4 ) 2 − 1 = 0. (36) we seek a solution of equation (36) in the form r = a1e β1x4 + a2e −β2x4 , where a1, a2, β1, and are some constants. then equation (36) transforms into the relation a1a2(β1 + β2) 2 = e(β2−β1)x4 , which implies β1 = β2 and a1a2 = 1 4β1 2 . since r > 0, it follows that a1 > 0 if β1 > 0, and we can set a1 = ea 2β1 and a2 = e−a 2β1 . if dr dx4 | x4=0 = 0, then a = 0. setting a1 = r0, we finally obtain r(x4) = r0 cosh x4 r0 . 270 n.n. popov:a geometric interpretation of gravity theory there exists yet another, complex, solution of equation (36), namely, r = ix4, but we are interested only in real positive solutions. substituting the solution r(x4) = r0 cosh x4 r0 into (35), we obtain the following expression r = 12 r02 for the scalar curvature, which proves the theorem. at first glance, the physical interpretation of the assertion of theorem 4 may seem rather contradictory. on the one hand, the hyperspherical space expands by an almost exponential law, i.e., the curvature radius increases as r = r0 cosh x4 r0 , while the scalar curvature r itself does not depend on the parameter x4 and remains everywhere constant. according to postulate 2, this means that the mass density of the matter uniformly filling the expanding space remains constant as well. in reality, these results involve no contradiction. at present, the existence of a new type of matter in our universe, which is known as dark energy, has been reliably established in cosmology; this matter fills uniformly whole space and is characterized by constant mass density not depending on the time parameter x4. the model suggested above can be regarded as an example confirming the existence of this type matter with such unusual physical properties. 7 conclusion the axiomatization of gravity theory proposed in this paper and the fundamental gravity equation obtained on the basis of these axioms make it possible to demonstrate the effectiveness of the suggested approach for a number of physical examples considered in the paper. it suffices to mention that the problem of constructing an everywhere continuous spherically symmetric stationary metric of a pseudo-riemannian space with discontinuous scalar curvature, which is solved in general form in section 4, still remains unsolved in the framework of grt, in which solving this problem involves fundamental difficulties. in the framework of the formalism suggested here, the concepts of a stationary black hole and dark energy are defined more rigorously from the mathematical point of view. importantly, dark energy arises in a natural way as one of the solutions of the system (10) of gravity equations, which, unlike in grt[9], does not require introduce any additional empirical constants into the main equation, such as the cosmological constant and its comparatively recent interpretation as the mass density of dark energy[10]. references [1] p. k. rashevskii. (1967), riemannian geometry and tensor analysis, (nauka, moscow) [in russian]. [2] b. a. dubrovin, s. p. novikov and a. t. fomenko. (1979), modern geometry, (nauka, moscow) [in russian]. advances in systems science and applications (2012) vol.12 no.3 271 [3] d. hilbert. (1915), “die grundlagen der physik (erste mitteilung), nachr. königl. gesellschaft d. wiss. göttingen”,math. phys. klasse, heft 3, pp.395407. [4] a. z. petrov. (1967), newest methods in general relativity theory, (nauka, moscow) [in russian]. [5] w. de sitter. (1917), “on einstein’s theory of gravitation and its astronomical consequences”, third paper, monthly notices roy. astron, soc. vol.78, no.3. [6] schwarzschild, k. (1916), “on the gravitational field of a point-mass, according to einstein’s theory sitzungsber”, preuss. akad. wiss., phys. math. kl. vol.189. [7] s. chandrasekhar. (1983), “the mathematical theory of black holes, research supported by nsf”, international series of monographs on physics, vol.69 (clarendon press-oxford university press, oxford-new york). [8] a. a. fridman. (1966), on space curvature, sselected works, (nauka, moscow) [in russian]. [9] a. einstein. (1917), kosmologische betrachtungen zur allgemeinen relativitätstheorie, ber. preuß. akad. wiss., berlin. [10] a.d. chernin. (2008), “dark energy and universal antigravitation”, uspekhi fiz. nauk, vol.178 no.3, pp.267-298 [physics-uspekhi, vol.51, no. 3, pp.253282]. corresponding author n.n. popov can be contacted at: nnpopov@mail.ru advances in systems science and applications (2012) vol.12 no.1 76-88 the performance study on two-stage vibration isolation system with active magnetic suspension control xiaojing liu and yefa hu mechanical and electronic engineering school, wuhan university of technology, wuhan 430070, china abstract two-stage vibration isolation system is normal equipment on naval vessel, which has good function of vibration isolation. active control on vibration isolation system is better than passive control. active magnetic suspension control has many virtues like alterable control methods and variable stiffness and damping. so that it is applicable to complex multi-interferes and multi-coupling floating raft. this paper set up mathematic models of two-stage vibration isolation with active magnetic suspension control system and solves the transfer functions. after simulating in matlab it analysis the effectiveness of kp ,τd,τi of pid controller to vibration isolation and the parameters which influence the vibration isolation performance including up and down layer stiffness, damping and mass ratio. the simulation results indicate that active magnetic suspension has good works to vibration isolation. at the same time, on the basis of these results optimize the construction. keywords two-stage vibration isolation system, active control of magnetic suspension 1 introduction two-stage vibration isolation systems is one kind of device of power equipment for damping vibration, which has better isolation vibration performance compared with one layer device. internal and external research focus on dynamic performance of two-stage vibration isolate system affected by design parameters[1], effects of nonlinear spring and damping on two stage vibration isolate system[2], effect of pedestal stiffness on two stage vibration isolate system[3] and study of flexible mass on two stage vibration isolate system[4], and so on. compared with passive control of isolation vibration system, active control can get better effectiveness by adjusting stiffness and damping. on this aspect, there are many researchs about methods[5], the fix position of active vibration isolator[6] and coupling vibration control[7]. active magnetic suspension control system has many virtues including stiffness and damping controllable and tunable. furthermore, it can apply classical or modern control theories according to different simulation force value and place to get the best effect, which make it become the best control system for multi-turbulent and multi-coupling floatin raft.this paper establish two-stage vibration isolation system with active magnetic suspenadvances in systems science and applications (2012), vol.12, no.1 77 sion control equipment and its mathematic models. then simulate in matlab and research how the three parapmeters of pid controller affect the peroformance of the system. furthermore, study the influence of parameters such as up and down layer stiffness, damping and mass ratio to vibration isolation. simulation results indicate that active magnetic suspension does good works to vibration isolation. in the end, on the basis of these results optimize the construction. 2 active magnetic suspension vibration isolation control system 2.1 theory of active magnetic suspension put active magnetic suspension vibration isolation system into the layer between middle mass m2and pedestal m3 .construction and theory is showed in fig.1 and fig.2.there are three masses —m1 is up layer mass,m2 is middle layer mass,m3is pedestal mass.k1and c1 is stiffness and daming of the first stage.k2 and c2 is stiffness and daming of the second stage.u is the system of active magnetic suspension.u is the system of active magnetic suspension.f = f0 sinϖt is outer simulate force act on up layer mass m1.x1 is displacement of m1.x2 is placement of m2.m3 is fixed to foundation firmly. so ignore its displacement. active fig.1model of active magnetic suspension in two-stage vibration isolation system magnetic suspension feed back system can provide damping force direct proportion to absolute speed of isolated object. when vibration happened and isolated vibration object elasticity mass increase or actual elasticity factor decrease, isolation vibration spring damping static deflection still keep constant. therefore active magnetic suspension feed back isolate vibration system is better than passive isolated vibration system, especially in low frequency area. middle mass m2 displacement is control variable quantity which is recorded by displacement sensor. after comparing with setting amount, pid controller calculate discrepant value and change electric current to produce magnetic force which help middle mass keep balance position to get the purpose about decrease the force from up layer to base m3. 78 xiaojing liu:the performance study on two-stage vibration isolation system... fig.2 theory of magnetic suspension system 2.2 transfer function establish system block plant (see fig.3).f(s) is outer simulation force,f1(s) is force from system to base. hk(s) is pid controller transfer function.h2(s) is power amplifier transfer function.h1(s) is the first stage transfer function.h3(s) is the second stage transfer function. fig.3 active magnetic suspension vibration isolation system block plant without magnetic force, system force transfer function is t (s) = f1(s) f(s) = h1(s)h3(s) when apply feed back control, system force transfer function is t1(s) = f1(s) f(s) = h1(s)h3(s) 1 +h2(s)hk(s) fe = k i2 x2 linearization force to fe = ki(i− i0)− kx(x− x0) advances in systems science and applications (2012), vol.12, no.1 79 hk(s) = kp + τi s + τds 1 + tfs h3(s) = k2 + c2s,h2(s) = ks substitute all parameters into system differential equation. ( m1 0 0 m2 ) .. x1 .. x2 + ( c1 −c1 −c1 c1 + c2 ) . x1 . x2  + ( k1 −k1 −k1 k1 + k2 )( x1(t) x2(t) ) = ( f fe ) use laplace transformation to equations and arrange to get h1(s) = x2(s) f (s) = s+ 440 (s+ 3.62)2(s+ 1.01)2 2.3 simulations in matlab simulink build block plant of system transfer function (see fig4.) system with and without active magnetic suspension control forcetime curves are list at fig.5. fig.4 block plant to simulate in simulink of matlab fig.5 showed that active magnetic suspension system indeed helps to improve the performance of two-stage isolate vibration. force transfer from up layer outer 80 xiaojing liu:the performance study on two-stage vibration isolation system... simulate to base become quit so small apparently and vibration curve become smooth. (a) nno feedback time-force transfer (b) with feedback time-force transfer fig.5 comparation with before and after using magnetic suspension feed back control 3 influence of pid controller parameters to isolate vibration 3.1 effective of proportion parameter kp without change other parameters, kp = 100, 500, 1000 result curve showed at fig.6 from curves we can see increase kp will decrease force to base tremedously. fig.6 different kp time-force curves 3.2 effective of differential constant τd without change other parameters, τd=100,500,1000. result curve showedat fig.7. form fig.7, τd has a little influnce to force max value and the speed of isolation vibriation. however, increase τd can make smooth of force vibraition and improve the performance of the system. advances in systems science and applications (2012), vol.12, no.1 81 fig.7 different τd time-force curves 3.3 effective of integral constant τi without change other parameters,τi=0.5,5,50. result curve showed at fig.8. fig.8 integral constant τd time-force curves curve tell that τi has no big influence to max force value but only effect to min force value slightly. 4 effectiveness of two-stage vibration isolation system ather parameters 4.1 natural frequency and force coefficient of transmission calculation without thinking of damping, system natural frequency will be calculated by listed formula(1) ωn 2 = k1 + k2 2m1 + k2 2m2 ± √ ( k1 + k2 2m1 + k2 2m2 )2 − k1k2 m1m2 substitute parameters into the formula: ξ2 = 0.06n/(m/s), ξ1 = 0.04n/(m/s) 82 xiaojing liu:the performance study on two-stage vibration isolation system... the natural frequencies are: ωn1 = 78hz,ωn2 = 2hz when force excited, the force coefficient of transmission is formula (2) tf = √ (α2 − 4ξ1ξ2αϖ2 1) 2 +ϖ2 1(2ξ1α 2 + 2ξ2α)2 a2 +b2 in the formula a = ϖ1 4 −ϖ1 2(α2 + 4ξ1ξ2α+ µ+ 1) + α2 b = ϖ1 3(2ξ2α+ 2ξ1µ+ 2ξ1)−ϖ1(2ξ1α 2 + 2ξ2α) µ = m1/m2, ϖ1 = ω/ω1, ϖ2 = ω2/ω1 ω1 2 = k1/m1, ω2 2 = k2/m2 ξ1 = c1/2 √ k1m1, ξ2 = c2/2 √ k2m2 when substitute the parameters into the formula, we get the curve of forcecoefficient transmission—simulation frequency (see fig.9). fig.9 curve of force-coefficient transmission-simulation frequency from the curve we can see the influence of simulation frequency to forcecoefficient transmission: • after 53hz, forcecoefficient of transmission is near zero that indicates the effectiveness of isolation vibration is good; • before 53hz, there are two peak values on resonance vibration that indicates this frequency region is the target we will research. to improve the effectiveness. advances in systems science and applications (2012), vol.12, no.1 83 4.2 influence of stiffness k2 to natural frequency according to formula (1), we can get the influence of stiffness to natural frequency (see fig.10). from the curve, we can see stiffness is almost linear direct proportion to natural frequency. so, the more stiffness is the more natural frequency is. on different working condition, we can change stiffness aim at natural frequency to avoid resonance vibration. fig.10 curve of stiffness to natural frequency fig.11 influence of stiffness to force-coefficient of 4.3 influence of stiffness k2 to force-coefficient transmission according to formula (2), choose three different value of stiffness k2500×102n/m, 500× 104n/m, 510×103n/m.we get three curves about k2 — force-coefficient of transmission (see fig.11). 84 xiaojing liu:the performance study on two-stage vibration isolation system... curves indicates that when stiffness is small the peak value move to left. if change up-layer stiffness, result is on the contrary that means when becomes larger the peak value move to left. so, increase up-layer or decrease down-layer stiffness all is good for isolation vibration. 4.4 influence of damping k1 ratio k1 to force-coefficient transmission without changing other parameters, damping ratio ξ1 = 0 ξ2 = 0 become ξ1 = 0.04 ξ2 = 0.06 we get the curve of simulation frequency—force e-coefficient of transmission (see fig.12).fig.12 indicates that damping ratio influence peak value directly but it doesn’t influence the position and tendency of peak value. the smaller damping ratio is the bigger peak value is. when ξ = 0.04, ξ2 = (0.01, 0.06, 0.2, 2) respectively, we get the curves (see fig.12 ξ1 = 0, ξ2 = 0; ξ1 = 0.04, ξ2 = 0.06 simulation frequency—force ecoefficient of transmission fig.13).when damping ratio is less than 1, it can’t influence peak value position and tendency, but influence peak value. the bigger damping ratio is the smaller peak value is, that is good for isolation vibration. while when damping ratio is more than 1, one hand it can decrease the peak value, on the other hand, peak value position move from left to right. advances in systems science and applications (2012), vol.12, no.1 85 fig.13 ξ1 = 0.04,ξ2 = (0.01, 0.06, 0.2, 2) simulation frequency—force e-coefficient of transmission fig.14 ξ2 = 0.06,ξ2 = (0.01, 0.04, 0.1, 2, 20) simulation frequency—force ecoefficient of transmission fig.15 change ξ2 and ξ1 simulation frequency—force e-coefficient of transmission 86 xiaojing liu:the performance study on two-stage vibration isolation system... when ξ2 = 0.06,ξ1 = (0.01, 0.04, 0.1, 2, 20) respectively, we get the curves (see fig.14).when damping ratio is less than 1, it can’t influence peak value position and tendency, but influence peak value. the bigger damping ratio is the smaller peak value is, that is good for isolation vibration. while when damping ration is more than 1, peak value position move from left to right and peak value become bigger when damping ratio is bigger. so, damping ratio than close to 1 is the best value to isolation vibration. 4.5 influence of mass ratio to force-coefficient transmission when m1 = 104kg, ξ1 = 0.04, ξ2 = 0.06, µ = 0.04042, 2, 5, 10 respectively, we get the curves showed in fig.16. isolation vibration system brings about two resonance points which position are determined by mass ratio. when µ ≥ 1, the first peak value position move to right but not very clearly and the second peak value position move to right quiet obvious. there is small influence to curve tendency in lower frequency region ,such as 40hz. but in low frequency region, the bigger mass ratio is the smaller peak value is. in middle frequency region, from 40hz to 180hz, curve change apparently, the second peak value decrease quickly. it is obviously that mass ratio influence middle frequency greatly. to make sure system work at frequency far away from resonance points we hope the distance between the first peak value and the second peak value is large, which means µ ≥ 1(mass m1 ≥ m1) is good for isolation vibration. fig.16different curve of simulation frequency—force e-coefficient of transmission 5 improve isolation vibration system according to above analysis, we take three proposals to improve the system performance. because mass has been chosen, then • without changing other parameters, decreasing under layer stiffness by readvances in systems science and applications (2012), vol.12, no.1 87 ducing spring from 6 to 4, then stiffness k2 = 430× 103n/m • without changing other parameters, raise up and under layer damping ξ1 = 0.4n/(m/s), ξ2 = 0.6n/(m/s) • decrease under layer stiffness and increase up layer stiffness at the same time, k2 = 430× 103n/m, ξ1 = 0.4n/(m/s), ξ2 = 0.6n/(m/s) results showed in fig.17, from which we can see that damping is the most important influence factor and only change stiffness is good for middle-high frequency isolation vibration. the best solution is decreasing under layer stiffness and increase up and under damping simultaneously. fig.17 change stiffness and dampin simulation frequency—force e-coefficient of transmission 6 conclusion this paper research on the performance of two-stage vibration isolation system and the effectiveness of system parameters including three parameters of pid controller,stiffness, damping and mass ratio. conlusion are • essence of active magnetic suspension control system act on two-stage vibration isolation system is change the second layer stiffness and damping to decrease vibration. • adjusting stiffness by active magnetic suspension to avoid system nature frequency is a good method to prevent resonance vibration. • decreasing the second stiffness by adjust active magnetic suspenion control has good effect. 88 xiaojing liu:the performance study on two-stage vibration isolation system... • adjusting damping by active magnetic suspension control to keep damping less than 1is good effect. • increasing proportion parameter kp is good for vibration isolation speed. • increasing differential constant τd ralatively is good for decreasing vibration. • integral constant τi has no apparent influence to prevent vibration. references [1] su ronghua, peng chenyu, ding wenwen. (2008), “research on two-stage isolation sysmtem’s dynamic performance affected by design parameters”, journal of basic science engineering, vol.12, pp.863-869. [2] meng quan, linjin tai, zhou yuxin. (2008), “the effects of nonlinear spring and a kind of semi-active control in hydraulic pressure damping on two stage vibration isolation system”, journal of academy of military transportation, vol.5, pp.89-91. [3] wang xiu-min, qiu yuan-wang, ding zhu-long. (2008), “effect of structure strength of pedestal on double isolation system”, i.c.e.aand powerplant, vol.2, pp.5-7. [4] zhu shi-jian, he lin. (2002), “study on the vibration-isolation effect of double-stage vibration isolation systems”, journal of naval university of engineering, vol.12, pp.6-8. [5] huang meizhi. (2008), “semi-active fuzzy control of diesel engine two-stage vibration isolation”, journal of heilongjiang institute of science and technology, vol.7, pp.292-294. [6] zhang chunliang, chen zicheng, mei deqing. (2003), “research on dynamics of two-layer active vibration isolation system”, china mechanical engineering, vol.7, pp.1232-1235. [7] yang tie-jun, chen yuqiang, huang jinge. (2001), “simulation research on active control of coupling vibration of a two-stage isolation system for diesel”, ship engineering, vol.3, pp.24-27. microsoft word 10 li guangyu wei fengying wang ke--some results for self-stabilization of stochastic differentialequation.d 280-285 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc. some results for self-stabilization of stochastic differential equation * li guangyu1, wei fengying 2 and wang ke 3 2,1 1dept of math, wenzhou univercity oujiang college, wenzhou 325027, zhejiang,pr china; 2college of mathematics and computer science, fuzhou university, fuzhou 350002, fujian, pr china; 3dept of math, harbin institute of technology, weihai 264209, shandong, pr china abstract a class of stochastic differential equation is studied about the self-stabilization in this paper. by constructing suitable hypothesis, sufficient criteria for the stochastic differential equation stabilized itself are established. keywords self-stabilization; stochastic neural networks; it’s formula; borel-cantelli lemma mr(2000) subject classification 92b20; 93e15 1. introduction the stability of solutions is very important in the theory of differential equations. many authors have done very excellent work [1-3] on the fields of stability. however, the stability for stochastic differential equations are somewhat difficult than that of deterministic differential equations, related literature can be found in recent work. for stochastic differential equations, [4] discussed a criterion that determines the region of mean square stability for second-order weak numerical schemes, [5] had given some sufficient conditions corcerning stability of solutions of stochastic differential evolution equations with general decay rate and [6] considered the one-step approximations of solutions, respectively. here we just mentioned the work for stability about mao xuerong, for example, he established stochastic versions of the well-known lasalle stability theorem in [7] and investigated exponential stability of paths for a class of hilbert space-valued non-linear stochastic evolutions in [8]. our work is motivated by mao [1], he had showed that the trivial solution of equation (1.1) in [1], namely, ( ) ( ( ), ) ( ( ), ) ( )dx t f x t t dt ug x t t db t= + (1) was almost surely exponentially stable for all sufficiently large u ,let 0u > be the noise intensity parameter and )(tb be an m -dimensional brownian motion. mao [1] had discussed if the intensity parameter u was replaced by 0 ( ) ( ) pt r s x s ds∫ , then the equation (1.1) becomes 0 ( ) ( ( ), ) ( ( ) ( ) ) ( ( ), ) ( ) pt dx t f x t t dt r s x s ds g x t t db t= + ∫ (2) where 0p > and ( )r s was a continuous $r^{n\times d}$-valued function defined on r+ satisfying ( ) tr t meγ≤ for all 0t ≥ ,which be called a convergence rate function. the standing hypothesis (h1) and (h2) are imposed in [1] as follows: *supported by nnsf of pr china(no.10701020) ; supported by nnsf of p.r.china (no.10726062). advances in systems science and applications (2011), vol.11, no.3-4 281 (h1.1) there exists a symmetric positive-definite d d× -matrixq and three positive constants , ,k α β with 2β α> , such that 2 2 2 ( , ) , ( ( , ) ( , )) , ( , ) , t t t t t x qf x t k x trace g x t qg x t x qx x qg x t x qx α β ≤ ≤ ≥ for all 0t ≥ and dx r∈ . now, we will prove that equation (4) stabilizes itself, that is a alternative theorem given by expressions (7)and (8). as far as the author's knowledge that there is no corresponding results. to make the statement more clear, mao [1] had stated the condition on the convergence rate function ( )r t as another hypothesis: (h1.2) there exists a pair of constants 0m > and 0γ ≥ such that ( ) tr t meγ≤ for all 0t ≥ 。 mao [1] had proved that if (h1.1) and (h1.2) hold, then for every 0 dx r∈ , the solution of equation (1.2), either 0 min 2( ) ( ) (2 ) ( ) p kr t x t dt qβ α λ ∞ ≤ −∫ (3) or ( )0 1limsup log ( ; ) 0 t x t x t→∞ < (4) holds for almost allω∈ω . if u is replaced by 0 sup ( ) ( ) s t r s x s ≤ ≤ , then equation (1.1) becomes 0 ( ) ( ( ), ) (sup ( ) ( ) ) ( ( ), ) ( ) s t dx t f x t t dt r s x s g x t t db t ≤ ≤ = + (5) for general 0 0t t≥ = and denote 0(0) dx x r= ∈ . mao [1] had also shown that if (h1) and (h2) hold, then for every 个 0 dx r∈ , the solution of equation (5) has the property 0 sup ( ) ( ) t r t x t ≤ <∞ < ∞ as. (6) furthermore, (i) if t →∞ as min ( ( ) ( ))r t x tλ →∞ , then lim ( ) 0 t x t →∞ = as. (7) (ii) if minliminf log[ ( ( ) ( ))] / 0t t r t r t tλ λ →∞ ≥ > ,then ( )1limsup log ( ) 2t x t t λ →∞ ≤ − as. (8) it is very interesting for us to analyze whether the system 0 ( ) ( ( ), ) (sup ( ) ( ) ) ( ( ), ) ( ) s t dx t f x t t dt r s x s g x t t db t ≤ ≤ = + (5) has the altermnative theorem like the results given by expressions (3) and (4). li: some results for self-stabilization of stochastic differential equation 282 to the author's knowledge, there is no corresponding results. now we will put our efforts on system (5), by use of exponential martingale formula, lyapunov function and some special inequalities in our paper, the trivial solution of equation (5) is self-stabilization will been shown in the next section. let us begin with our paper now. 2. main results lemma 2.1 let hypothesis (h1) hold. then the solution of equation (5) has the property that 0{ ( ; ) 0p x t x ≠ for all 0} 1t ≥ = provided 0 0x ≠ . we consider the following problem of stochastic self-stabilization in this section. suppose we are given a stochastic differential equation 0 ( ) ( ( ), ) (sup ( ) ( ) ) ( ( ), ) ( ) s t dx t f x t t dt r s x s g x t t db t ≤ ≤ = + (9) on 0 0t t≥ = with initial value 0(0) dx x r= ∈ (it is just for convenience to set 0 0t = and the theory clearly works for general 0 0t ≥ . we have the following result. theorem 2.1 let (h1) and (h2) hold. then for every 0 dx r∈ , the solution of equation (9), either 0 min 2sup ( ) ( ) (2 ) ( )s kr s x s qβ α λ≤ ≤∞ ≤ − (10) or ( )0 1limsup log ( ; ) 0 t x t x t→∞ < (11) proof. since hypothesis (1) guarantees ( ;0) 0x t ≡ , one only need to show the conclusions for all 0 0x ≠ . fix 0 0x ≠ arbitrarily, and write 0( ; ) ( )x t x x t= , by lemma 2.1 0t ≥ for all 0( ; ) 0x t x ≠ almost surely. suppose (10) is false, then there exists some 0 0x ≠ for which _ ( ) 0p ω > , where 0 min 2: sup ( ) ( ) . (2 ) ( )s kr s x s q ω β α λ − ≤ ≤∞ ⎧ ⎫⎪ ⎪ω = ∈ω >⎨ ⎬−⎪ ⎪⎩ ⎭ clearly, one only needs to show that (11) holds for almost allω∈ω , for each 1, 2, ,i = … define 1 1 2 0 min 2: sup ( ) ( ) (1 ) . (2 ) ( )i s kr s x s i i q ω β α λ − − − ≤ ≤∞ ⎧ ⎫⎪ ⎪ω = ∈ω > +⎨ ⎬−⎪ ⎪⎩ ⎭ now, 1 i i − −∞ = ω∈ ω∪ , and hence one only needs to show that for each 1i ≥ ,(10)holds for almost iω − ∈ω . fix any 1i ≥ let ( ( ), ) log( ( ) ( ))tv x t t x t qx t= , one then derives that advances in systems science and applications (2011), vol.11, no.3-4 283 2 2 ( ) 2 ( ) ( ) 2 ( ) 2 ( )( ( ), ) 0, ( ( ), ) , ( ( ), ) . ( ) ( ) ( ( ) ( )) t t t t t t x t q qx t qx t qx t x t qv x t t vx x t t vxx x t t x t qx t x t qx t − = = = i by it ô ’s formula, we obtain, we obtain 0 0 0 0 ( ( ), ) ( ( ), ) ( ( ), ) 1 [ ( ( ), )(sup ( ) ( ) ) ( ( ), )(sup ( ) ( ) ) ( ( ), )] 2 ( ( ), ) ( ( ), ) (sup ( ) ( ) ) ( ( ), ) ( ( ), ) ( ) 1 (sup ( ) ( ) ) 2 t x t xx s t s t x x s t s t dv x t t v x t t dt v x t t dx trace g x t t r s x s v x t t r s x s g x t t dt v x t t f x t t dt r s x s v x t t g x t t db t r s x s ≤ ≤ ≤ ≤ ≤ ≤ ≤ ≤ = + + = + + 2 0 2 2 0 0 ( ( ( ), ) ( ( ), ) ( ( ), )) 2 ( ) ( ( ), ) ( ) ( ( ), )2(sup ( ) ( ) ) ( ) ( ) ( ) ( ) ( ) 1(sup ( ) ( ) ) ( ( ( ), ) ( ) ( ) ( ( ), )) ( ( ) ( )) 2(sup t xx t t t t s t t t t s t trace g x t t v x t t g x t t dt x t qf x t t x t qg x t tdt r s x s db t x t qx t x t qx t r s x s trace g x t t qx t qx t g x t t dt x t qx t ≤ ≤ ≤ ≤ ≤ = + + − 2 t 2 1( ) ( ) ) ( ( ( ), ) ( ) ( ) ( ( ), )) , ( ( ) ( )) t t s t r s x s trace g x t t qx t x t qg x t t dt x t qx t≤ this yields that 0 0 00 0 2 00 2 0 log( ( ) ( )) 2 ( ) ( ( ), ) ( ) ( ( ), )log( ) 2 (sup ( ) ( ) ) ( ) ( ) ( ) ( ) ( ) ( ( ( ), ) ( ( ), ))(sup ( ) ( ) ) ( ) ( ) ( ) 2 (sup ( ) ( ) ) t t tt t t t t s t t t t s t t s t x t qx t x s qf x s s x s qg x s sx qx ds r s x s db s x s qx s x s qx s trace g x s s qg x s sr s x s ds x s qx s x s r s x s ≤ ≤ ≤ ≤ ≤ ≤ = + + + − ∫ ∫ ∫ 2 2 0 ( ( ), ) . ( ( ) ( )) t t qg x s s ds x s qx s∫ by hypothesis (h1.1) and the fact that 2 2 min max( ) ( ) ( ) ( ) ( ) ( )tq x t x t qx t q x tλ λ≤ ≤ for q is a symmetric d d× matrix, one can show that for any 0t ≥ 2 0 0 0min 0 2 2 2 00 log( ( ) ( )) 2log( ) ( ) (sup ( ) ( ) ) ( ) ( ) ( ( ), ) 2 (sup ( ) ( ) ) ( ( ) ( )) t t t v s tt t v s x t qx t ktx qx m t r v x v ds q x s qg x s s r v x v ds x s qx s α λ ≤ ≤ ≤ ≤ ≤ + + + − ∫ ∫ (12) where 00 ( ) ( ( ), )( ) 2 (sup ( ) ( ) ) ( ) ( ) ( ) t t t v s x s qg x s sm t r v x v db s x s qx s≤ ≤ = ∫ is a continuous martingale vanishing at 0t = , let 1, 2, ,k = … then by the exponential martingale inequality li: some results for self-stabilization of stochastic differential equation 284 1 1 2 0 2 4 (1 ) log 1( : sup ( ) ( ), ( ) ) , 4 (1 ) 2t k i kp m t m t m t i k β α βω β β α − − ≤ ≤ ⎡ ⎤− + − > ≤⎢ ⎥+ −⎣ ⎦ where 2 2 2 00 ( ) ( ( ), ) ( ), ( ) 4 (sup ( ) ( ) ) . ( ( ) ( )) tt t v s x s qg x s s m t m t r v x v ds x s qx s≤ ≤ = ∫ wherehence the well-known borel-cantelli lemma yields that for almost all ω∈ω there exists a random integer 1( )k ω uch that for almost all 1k k≥ , 1 1 0 2 4 (1 ) logsup ( ) ( ), ( ) 4 (1 ) 2t k i km t m t m t i β α β β β α − − ≤ ≤ ⎡ ⎤− + − ≤⎢ ⎥+ −⎣ ⎦ that is,for 0 ,t k≤ ≤ 1 1 2 1 2 1 2 00 4 (1 ) log 2( ) ( ), ( ) 2 4 (1 ) ( ) ( ( ), )4 (1 ) log 2 (sup ( ) ( ) ) . 2 (1 ) ( ( ) ( )) tt t v s i km t m t m t i x s qg x s si k r v x v ds i x s qx s β β α β α β β β α β α β − − − − ≤ ≤ + − ≤ + − + + − = + − + ∫ (13) substituting (13)into (12) and then applying (h1) one obtains that for each ^ ,ω∈ω−ω with ^ ω a p-null set, there exists a random integer 2 ( ),k ω such that, for every 个 iω − ∈ω there exists a random number 3( )k ω such that 1 1 2 0 min 2sup ( ) ( ) (1 ) (2 ) ( )s t kr s x s i i qβ α λ − ≤ ≤ ≥ + − for almost all 3.t k≥ it then follows from that for almost all ^ iω − ∈ω −ω, if 2 31 , ( 1)k t k k k k− ≤ ≤ ≥ ∨ + 3 1 2 0 0 1 0min 1 1 0 0 3 min min 1 3 0 0 min log( ( ) ( )) 2 4 (1 ) log 2log( ) (sup ( ) ( ) ) ( ) 2 (1 ) 2 4 (1 ) log 2 (1 )log( ) ( 1 ) ( ) 2 ( ) 2 ( 1) 4 (1 ) llog( ) ( ) t t t v sk t t x t qx t kt i kx qx r v x v ds q i i kt i k k ix qx k k q i q k k ix qx q β β α λ β α β λ β α λ β λ − − ≤ ≤ − − − + − ≤ + + − − + + + ≤ + + − − − − + + ≤ + + ∫ 3 min og 2 ( 1 ). 2 ( ) k k k k i qβ α λ − − − − this implies that 1 3 0 0 3 min min 2 ( 1)1 1 4 (1 ) log 2log( ( ) ( )) log( ) ( 1 ) . 1 ( ) 2 ( ) t t k k i k kx t qx t x qx k k t k q i q β λ β α λ −⎡ ⎤+ + ≤ + + − − −⎢ ⎥− −⎣ ⎦ it then follows that min 1 2limsup log( ( ) ( )) ( ) t t kx t qx t t i qλ→∞ ≤ − for almost all ^ .iω − ∈ω −ω (14) advances in systems science and applications (2011), vol.11, no.3-4 285 thus, for almost all ^ ,iω − ∈ω −ω there exists a random number 4 ( )k ω and 0δ > arbitrarily such that min 1 2log( ( ) ( )) 2 ( ) t kx t qx t t i q δ λ ≤ − + for almost all 4 ,t k≥ from the expression 2 min ( ) ( ) ( ) ( )tq x t x t qx tλ ≤ one derives that min min exp{( ) } ( )( ) ( ) k t i qx t q δ λ λ − + ≤ for almost all 4.t k≥ consequently, 0 min 1 2limsup log( ( ; ) ) ( )t kx t x t i q δ λ→∞ ≤ − + for almost all ^ .iω − ∈ω −ω since 0δ > is arbitrary, we must have that 0 min 1 2limsup log( ( ; ) ) 0 ( )t kx t x t i qλ→∞ < − < for almost all ^ .iω − ∈ω −ω the proof is now complete. references [1] x. mao, stochastic differential equations and there applications, horwood publication, chichester, 1997:100-141 [2] j. k. hale, theory of functional differential equations, springer-verlag, new york inc, 1977:50-70 [3] geng liu, study on capacity expansion problems on directed networks, advances in systems science and applications 2007, 7(4): 679-684 [4] xianguo wu, lieyun ding, xintian cai, deformation predication of deep excavation pit support structure based on fuzzy neural network, advances in systems science and applications 2006, 6(2), 119-126 [5] xiaoya hu, desen zhu, bingwen wang, delay analysis of switched ethernet for networked control systems, advances in systems science and applications 2006, 6(2) : 201-206 [6] chen wanyi, globally exponential asymptotic stability of hopfield neural network with time-varying delays, act scientiarwn naturaliwn universitatis nankaiensis, 2005, 5 (3):70-75 [7] jiang minghui, shen yi, liao xiaoxin, stability of stochastic neural networks with multi-delay, mathemics applicata, 2006, 3 (19): 61-65 [8] x. mao, some contribution to stochastic asymptotic stability and boundedness via multiple lyapunov functions, j. math. anal. appl. 2001 260: 325-340. the article e.v. korolyuk, e.v. mezentseva “program-target approach based on the institutional environment of entrepreneurship in the krasnodar territory”, advances in systems science and application (2016), vol.16, no.4, p. 13-21 was withdrawn by the editorial board due to violation of assa publication ethics (please, see publication ethics statement). the reason is that the article’s content lacks originality due to substantial intersections with the following paper: e.v. korolyuk, e.v. mezentseva "establishing programme-target approachbased institutional entrepreneurship environment in krasnodar krai” in mediterranean journal of social sciences (2015), vol. 6, no. 6, pp. 583-589 editorial board of advances in systems science and application, 22.08.17 http://ijassa.ipu.ru/ojs/ijassa/ethics advances in systems science and application (2016) vol.16 no.2 54-69 adaptive background modeling for dynamics background n.a. zainuddin, y.m. mustafah, a. a. shafie, a.w.azman, m.a. rashidan, and n.n.a. aziz department of mechatronics, kulliyyah of engineering, international islamic university malaysia abstract an increasing number of cctv have been deployed in public and crime-prone areas as demand for automatic monitoring system is increasing to counterbalance the limitation of human monitoring. to have a good monitoring system in such places, a good background model is needed in order to reduce amount of the video processing needed for tracking, classification, counting and etc. this paper proposes an adaptive background modeling that is able to model a scene under review at real-time. the proposed modeling system is also expected to be able to handle dynamic backgrounds and common problems in detection methods. a novel patch-based background reconstruction based on highest frequency of occurrences assumption and past pixel observation is proposed. contrast adjusting method is used to reduce the problem of incorrectly classified foreground which is shadow problem. the proposed algorithm is focused to be tested and analytically compared with the dynamic background at the indoor and outdoor environment. the main challenges of background subtraction such as illumination changes, geometrical changes, stationary moving object problem and high speed object problem are taken care of and extensively discussed in this paper. the experimental results show that the algorithm is able to reconstruct a background model and produce accurate and precise foreground that can be used for other processing stages. keywordsadaptive background modeling; background subtraction; surveillance; foreground segmentation; dynamic background. 1 introduction there are two basic elements in video analysis that need to be identified, which are background and foreground region. the background which is also known as the reference frame usually consists of non-interested objects, either static or dynamic. foreground image on the other hand, is the objects of interest that need to be identified, detected or further analyzed. the shape of a foreground region unnecessarily be a rectangle. it can be arbitrary in shape and the background is in the complementary shape of the foreground region. the most common technique used by many researchers in handling this problem is by using background subtraction[1,2]. there is a lot of study in background subtraction, but the problem is still yet to be properly solved. the problem of ghost, stationary advances in systems science and application (2016) vol.16 no.2 55 moving object and repetitive movement of complex background scenes have been attempted to be solved, however, the result is still unsatisfactory. in the early days, many researches have focused on improving the general background subtraction process such as temporal differencing, double temporal differencing and many more. the recent trend shows that researchers tend to work on the specific stage in the background subtraction algorithm, which is background modeling and background reconstruction. the reconstruction of the background model is essential in order to correctly identify the foreground. a good background model will benefit the analysis of the foreground and produce accurate and effective background subtraction. the simplest method is setting the image without any moving object as the background[3]. however, this assumption is too idealistic, and not possible to solve many problems especially in outdoor environments, such as on a crowded street with vehicles and pedestrians or an outdoor environment like in a street will experience various levels of illumination at different times of the day[4]. plus, outdoor environment has a tendency of experiencing adverse weather condition like fog, rain or strong wind that could modify the reference image[5]. in these cases, the background must be adaptively refreshed and updated. therefore, in this paper, an adaptive background modeling technique is presented. there are five sections in this paper. in section two, previous works related to the research are described. meanwhile, in section three, the proposed algorithm is presented and discussed extensively. in section four, the experimental results at various scenes are presented with qualitative and quantitative results from various situations are illustrated. finally, the conclusion is drawn in section five. 2 related works recently, a lot of works are concentrating on adaptive background reconstruction[69]. several notable methods were introduced including temporal smoothing, pixel intensity classification, running gaussian average, gaussian mixture model, hidden markov model and kernel density estimation. gaussian mixture model (gmm) is a common method used by researchers in the foreground detection. the main idea of the gaussian mixture model is by having n number of gaussian distribution function, in order to reconstruct the background pixel. the implementation of gmm has been proposed as the adaptive background reconstruction technique, especially in surveillance system since late 1990s by stauffer et al.[10]. as it is still in the early stage, a lot of unwanted small blobs are scattered all over the frame and the main blobs aren’t completely detected as foreground. since then, many of the researchers have continued to improve the proposed algorithm[11-13]. mukherjee et al. have incorporated the horpreset colour model in order to make the algorithm to be able to detect 56 n.a. zainuddin etal: adaptive background modeling for dynamics background shadows[11]. mukherjee et al. work focuses more on the shadow handling, the experimental results show that the algorithm is tested for a non-complex indoor scene with one foreground handling. literature[12,13] managed to reconstruct the background with shadow consideration better with a more complex scene than mukherjee et al.’s work. they use the train station dataset in order to have multiple object handling. they have implemented the use of window-based decision rules for the shadow and the background model is reconstructed by gmm. to sum up, many implementations of gmm are focusing on the indoor scene for the surveillance system. elgammal et al. claimed that gmm is ideal for indoor scenes only[14]. thus, they introduced the use of kernel density estimation (kde) for background reconstruction. in kde method, the background and foreground pixel are model of probability density function (pdf) and the pdf is estimated by a kernel function or also known as window function. however, the major drawback of this method is the expensive computational cost. later gao et al. improved the method by introducing the marr wavelet equation in estimating the pdf[5]. marr wavelet is a second derivative of the gaussian smooth function. they combine the gaussian and kernel method together. however, from the experimental results, the background model is not reconstructed correctly as some of the background pixels are counted as a target; thus in some cases, the target is not detected correctly. lee et al. managed to reduce the need of storage[15]. they initialized the first frame and update it at every frame by setting the learning rate. for dynamic background cases, they used threshold method. the first frame was set as the background model and later, updated by learning method. this method is only ideal for the video sequence at has no target or foreground the beginning of the video sequence. for the crowd and complex scene, it is quite impossible to get such video sequence. the other method that commonly used by researchers in reconstructing the background is temporal smoothing[16]. the main idea of temporal smoothing is combining the stored image with the new image on pixel-based. pixels on the smoothed image will be replaced by a part from previous value combined with a new value from new image at the same position. ridder et al.[17] have improved the method by modeling each pixel with kalman filter. they managed to make the system more robust to cope with the illumination changes, however, the algorithm update the background slowly. later, in recent year, hung et al. improved the method by combining with the median filtering[6]. by doing so, they managed to reduce the computing frequency of median operations. however, they focus on the computational performances of the algorithm with no real life situation is considered in their experiment. motivated by [6], asif et al. use the temporal smoothing and median filtering for modeling the background[7]. as advances in systems science and application (2016) vol.16 no.2 57 their research focus more on the human gait, the background and foreground are non-complex. based on their experimental result, the temporal smoothing seems to be ideal for the non-complex scene. thus, the implementation of temporal smoothing algorithm is ideal for surveillance system that involves non-complex scene (indoor surveillance). in early of 2000s, hou et al. have introduced the pixel intensity classification (pic) method as the background reconstruction technique[18]. in this approach, at every frame, the difference of pixel intensity is calculated. then, the classification is made based on the calculated difference. the background model is assumed to be the highest frequency in the intensity value. the simulation results of hou et al. literature[18] used the outdoor dataset with parking lot environment. the works on the same method is continued by xiao et al. in 2006 and 2008[2,8]. xiao et al. managed to handle more adverse weather by adapting the raining situation, however, their experimental results shown that the adaption of the algorithm to the rain is still at the early stage. next, cao et al. have employed the improved version of pic in reconstructing the background model for light flow traffic movement video sequence[19]. they managed to compare their experimental results with gmm and time averaging algorithm. the results however are only ideal for slow moving to medium moving foreground detection. some of the foreground in high speed foreground is not detected and small blobs appeared as a result from the ghost problem also known as the blending of high speed object movement in the video sequence. after analyzing these methods, several assumptions were made. in non crowded scenes, the background pixel would be the maximum frequency in the image sequence. for crowded condition, most of the pixels in the image frame are expected to be modified. from the assumptions, we developed a background reconstruction algorithm based on pixel intensity classification and mode filtering. however, the difference in the inter-frame pixel intensity value is not calculated in the first step, the mode filtering is done first in order to model the background. 3 methodology in order to produce a good background model, several techniques are deployed through several stages with pre-determined assumptions. fig.1 shows the overall process of proposed background modeling algorithm. an elaborated discussions on the proposed algorithm is presented in this section. 3.1 pre-processing the video will be converted into uniform size of image sequences. next, the image sequence will be gray-scaled and undergone median filtering process in order to get rid of noises, especially from the camera pixel noise and impulse noise. 58 n.a. zainuddin etal: adaptive background modeling for dynamics background fig. 1 process of proposed background modeling algorithm 3.2 background model in order to model the background, an assumption of background-foreground pixels needs to be made. thus, based on the assumption that the background pixel is the most frequent appeared in the entire video sequences, the proposed algorithm is as follows: step 1 : calculate the mean of each patch let f1, f2, f3, · · · , fn represent the frames from the same video sequence. first, each of the frames must be segmented into m × n patches, where m and n can be the same number. assume n patches are obtained, where we let patches be marked as p1, p2, p3, · · · , pn . experiments are conducted to find the optimum size of the patches. the result of the experiments shows that the size of the advances in systems science and application (2016) vol.16 no.2 59 patches is not as sensitive as the sampling frame value. it does affect the result, but in small percentage error. then, at each frame,f1, f2, f3, · · · , fn the intensity value of pixel (x,y) in the area of specified patch, pi are marked as i1, i2, i3, . . . , in, where i = 1, 2, 3, . . . , n and total number of pixels in one patch can be written as: q = m× n (1) then, mean mp,f at each frame for the specified patch is calculated as in the formula below: mp,f = ∑q j=1 ij(x, y) q (2) step 2 : calculate the preliminary mode in order to calculate the mode, the calculated mean is grouped into matrices called mean matrixp,f according to the specified sampling time. assume that the sampling time is denoted as t . thus, the preliminary mode at that time interval for patch 1 is: nmodepre = max(m1,1,m1,2, · · · ,m1,t−1) (3) this formula is repetitively used for all other patches. step 3 : check the preliminary mode there are two possibilities of the mean matrices. first, there is repetitive mean value and the second one is there is no repetitive mean value. for the former case, there is no issue as it will follow equation 3 correctly. however, for the latter case, the result of the mode is not correctly calculated. the mode algorithm will set the minimum mean value as mode if there is all frequencies of the matrix elements are 1. to avoid this, the checking algorithm is set. assume that the result is the final mode and marked as,nmodefinal and a threshold value, ε is introduced. meanwhile, subp,f matrices are introduced to store the result of subtraction. if the number of zero in the subp,f matrix equal to or greater than a threshold number, ε, the final mode, nmodefinal is set to the calculated preliminary mode, nmodepre indicating there is repetitive mean value in the mean matrixp,f . likewise, if the number of zero in the subp,f matrix less than a threshold number, ε, the final mode, nmodefinal is set to the previous frame values, as there is repetitive mean value in the mean matrixp,f . subp,f = |mean matrixp,f − nmodeprep,f | (4) nmodefinalp,f = { nmodeprep,f n(sub) = 0 ≥ ε nmodeprep,f−1 n(sub) = 0 < ε (5) 60 n.a. zainuddin etal: adaptive background modeling for dynamics background step 4 : construction of background model and adaptive update to build the background model, the calculated final mode values are utilized. it is important to note that, in the earlier stage, we have built storage to store the patches in rgb image form matrices. thus, the final stage to build the background model is by pointing the right rgb image of patch to the right patch according to the final mode value. 3.3 background subtraction and binarization in binarization process, the images will be transformed into black and white images. the binarization algorithm used in this research is motivated by global-used binarization method, otsus method, in which threshold algorithm is deployed. this method is computationally inexpensive thus ideal for real time applications. in this method, we need to set for an initial threshold, ti. we estimate average of the minimum and maximum pixel value of the image, ti, where mathematically it is calculated as: ti = max(i) +min(i) 2 (6) then, the whole pixels are segmented based on the ti value. for the intensity values that greater than or equal to the ti value, it is grouped into g1. otherwise, the intensity values will be grouped to g2 and the average intensity values of each group are calculated and noted as ave1 and ave2 respectively, as shown in equation 7,8,9 and 10 respectively. group = { g1 ij ≥ ti g2 ij < ti (7) ave1 = ∑ ig1 n(ig1) (8) ave2 = ∑ ig2 n(ig2) (9) based on the ave1 and ave2 value, new threshold value is computed: tfinal = ave1 + ave2 2 (10) 3.4 shadow removal shadow is the incorrectly classified foreground pixels that need to be eliminated. foreground mask without shadow removal will adversely affect further processing by providing the false information such as inaccurate blob size, centroid and etc. thus, the efficiency of the method will be affected. there are many proposed algorithms for shadow removal, however, as we already have a good background advances in systems science and application (2016) vol.16 no.2 61 model based on the proposed algorithm, we just need a simple shadow removal. thus, the selection of algorithms of the shadow removing is focused on the computational aspects. one of the simplest methods for shadow removal is by adjusting the contrast. to simplify that we just adjust the luminance value of the pixel by multiplying the current luminance value, ycurrent with a predefined contrast factor, c. thus, the new luminance value is: ynew(x, y) = ycurrent(x, y)× c (11) 3.5 morphological in order to filter the small unwanted pixel in the binary images of the foreground mask, the operation of mathematical morphological is used. there four are basic morphological operations, which are erosion, dilation, opening and closing. in this research, we used the closing morphological operator to close the gap of highly-deformed of interested object shape. closing morphological operator mathematical definition is basically the morphological dilation followed by erosion operation. thus, the equations are: h • i = (h ⊕ i)⊖h (12) where i is the binary image of the foreground and h is the structuring element. holes in the foreground those are smaller than h will be filled. thus, it will reduce the deformation on interested shapes. 4 experimental results in this section, evaluation of the methodology is presented in order to compare the proposed algorithm with several existing methods. both qualitative and quantitative perspectives are discussed. in order to prove the robustness of the proposed algorithm, several video sequences with various scenes are tested in which including the complex dynamic background at the indoor and outdoor environment. 4.1 comparison methods several works on background modeling and object detection techniques are chosen to be compared in terms of performance with our proposed method. the methods are stochastic approximation[20], bayesian learning-based[21], adaptive gaussian mixture model[22], gaussian mixture model[23], and pfinder[24]. all the methods are tested using matlab software and performed on windows 7 on intel r coretm i3-2330m cpu 2.2ghz processor. 4.2 performance measurements to quantitatively evaluate the performance, well-known parameters in gold standard test measurements[25,26] are used. in this method, there are four important 62 n.a. zainuddin etal: adaptive background modeling for dynamics background parameters that need to be defined which are true positive (tp), true negative (tn), false positive (fp) and false negative (fn). in this research, the true positive is defined as the number of pixels that correctly identified as foreground, conversely false positive is number of background pixels that incorrectly identified as foreground. meanwhile, true negative is defined as the number of pixels that correctly identified as background. on the contrary, false negative refers to number of foreground pixels that incorrectly identified as background[27]. from these four parameters, the performance measures are defined as follows: recall(tpr) = tp tp + fn (13) precision(ppv ) = tp tp + fp (14) accuracy(acc) = tp + tn tp + tn + fp + fn (15) f1score = 2tp 2tp + fp + fn (16) specificity(tnr) = tp tp + fn (17) negativepredictivevalue(npv ) = tn tn + fn (18) fallout(fpr) = fp tn + fp (19) falsenegativerate(fnr) = 1− tpr (20) falsediscoveryrate = 1− ppv (21) 4.3 results fig.2 shows the results of the studied methods with our proposed methods in movingcurtain, wavebeach, campusroad and fountain video sequences1. each row shows the foreground mask generated by the corresponding methods for each video sequence respectively. row 2, row 4, row 6 and row 8 of fig.2 show the overlapping results of the foreground generated after background subtraction and ground truth. the overlapping results are used to quantitatively measure the performances of each method. the true positive (tp), true negative (tn), false positive (fp) and false negative (fn) are defined based on the overlapped pixel colours. true positive denoted by white pixels and black pixels are true negative. meanwhile, magenta pixels are for false positive and green pixels indicate false negative. on the other hand, the first column shows the selected frames from the respective video sequences and the second column shows the ground truth advances in systems science and application (2016) vol.16 no.2 63 information of that frame. the other columns are the foreground mask results from the studied methods and our proposed results. fig. 2 experimental results on complex scenes of dynamic background frame number for movingcurtain, wavebeach, campusroad and fountain are 2774,1499, 2348 and 1196 respectively from fig.2, we can deduce that in the high variability scenes like movingcurtain and video sequences, all algorithms handle the continuous moving curtain quite well, except for pfinder[24] as several of the moving curtain pixels are falsely detected as foreground and shadow suppression also failed to be handled by pfinder[24] and uniform-gaussian[20] that causing the additional white pixel formation at the bottom of human foreground for both foreground masks. in the wavebeach video sequences, the variability of the background scene increases as the wave at the beach are continuously moving at a higher frequency than the moving curtain, plus as it is in an outdoor environment, the illumination condition is tested with the addition of stationary moving object condition. this is because the foreground (human) in this sequence stays for quite a long time before moving, thus, the moving object (human) tends to be falsely classified as background. in this video sequence, we can see that most of the algorithms could 64 n.a. zainuddin etal: adaptive background modeling for dynamics background not handle the continuously moving wave as several pixels are falsely classified as foreground except for our proposed algorithm and uniform-gaussian[20] that able to give a clean foreground mask. for the stationary moving object cases, most of the algorithms lost the information of bottom part of the human (foreground) however, the upper part of the foreground are perfectly detected. in the campusroad video sequences, the problem of similar pixel colours and intense noise are evaluated. out of five methods, two methods are able to separate the noise from the negative effect of the waving trees which are our proposed algorithm and the bayesian[21] method. as the dark-coloured car is quite similar to the colour pixels of background, the shape of the car in the foreground mask is deformed except for our proposed algorithm and uniform-gaussian[20]. meanwhile, in the fountain video sequences where the dynamic background (running water from the fountain) is more than 70 percent of the frame size and the pedestrian keep moving in the entire video sequences requires the algorithm to be able to update the background model in a short time. the problem of ’ghost’ detection or slow update can be detected in bayesian approach foreground mask as the foreground mask tend to be bigger in size than the actual size. meanwhile, the repetitive movements from the fountain fail to adapt by adaptive gaussian mixture model (agmm), gmm and pfinder. to quantitatively measure the performance of the studied methods and proposed method,the performance measurements discussed in section 4.2 are used and tabulated in table 1, 2, 3 and 4 for each of the test sequences. meanwhile, table 5 summarizes the results by averaging the result from table 1, 2, 3, and 4. in each table, the best performance is highlighted in bold. to sum up, in the campusroad test sequences, we can see that our proposed algorithm outperforms all other methods by having high value in all measurement parameters. in movingcurtain test sequences, bayesian learning-based is the best in term of accuracy, however, the difference is only by 0.0028 compared with our proposed method. in addition, in this test sequence, our proposed method has the highest precision value and harmonic mean of precision-recall, f1 scores which make our proposed algorithm surpassed the bayesian method holistically. meanwhile, in wavebeach test sequences, uniform-gaussian[20], performs in most measurement parameters compared with the rest of the studied methods, however, our proposed algorithm still performs the best in term of precision and only have a small difference in term of accuracy. in fountain test sequences, our proposed algorithm scores five best values out of nine test parameters and the highest precision value that is very huge difference compared to other methods. based on table 5, we can summarize that our proposed method is the best in term of accuracy, precision and recall. the consistent high precision value of our proposed algorithm is outstanding and make the algorithm as a reliable detection advances in systems science and application (2016) vol.16 no.2 65 t a b le 1 q u an titative resu lt for m ov in gc u rta in test seq u en ces a lg o rith m a ccu ra cy p recisio n r eca ll f n r f d r n p v t n r f 1 s co re f a llo u t u n ifo rm + g a u ssia n [2 5 ] 0 .9 8 4 0 .8 1 3 1 0 .8 3 0 4 0 .1 6 9 6 0 .1 8 6 9 0 .9 9 2 1 0 .8 3 0 4 0 .8 2 1 7 0 .0 0 8 9 b a y esia n [2 6 ] 0 .9 9 0 1 0 .9 1 3 0 .8 0 1 6 0 .1 9 8 4 0 .0 8 7 0 .9 9 2 6 0 .8 0 1 6 0 .8 5 3 7 0 .0 0 2 9 a g m m [2 7 ] 0 .9 8 7 0 .9 3 3 2 0 .8 0 1 9 0 .1 9 8 1 0 .0 6 6 8 0 .9 8 9 5 0 .8 0 1 9 0 .8 6 2 5 0 .0 0 3 1 g m m [2 8 ] 0 .9 9 0 .9 4 2 8 0 .8 2 3 7 0 .1 7 6 3 0 .0 5 7 2 0 .9 9 1 9 0 .8 2 3 7 0 .8 7 9 2 0 .0 0 2 3 p fi n d er[2 9 ] 0 .9 8 3 3 0 .8 0 0 7 0 .9 6 0 5 0 .0 3 9 5 0 .1 9 9 3 0 .9 9 7 4 0 .9 6 0 5 0 .8 7 3 3 0 .0 1 5 3 p ro p o sed 0 .9 8 7 3 0 .9 6 2 4 0 .8 4 6 5 0 .1 5 3 5 0 .0 3 7 6 0 .9 8 8 9 0 .8 4 6 5 0 .9 0 0 7 0 .0 0 2 4 t a b le 2 q u an titative r esu lt for w aveb ea ch test seq u en ces a lg o rith m a ccu ra cy p recisio n r eca ll f n r f d r n p v t n r f 1 s co re f a llo u t u n ifo rm + g a u ssia n [2 5 ] 0 .9 9 0 2 0 .9 2 4 8 0 .8 8 5 2 0 .1 1 4 8 0 .0 7 5 2 0 .9 9 3 7 0 .8 8 5 4 0 .9 0 4 5 0 .0 0 4 b a y esia n [2 6 ] 0 .9 7 2 0 .8 1 6 9 0 .7 3 7 8 0 .2 6 2 2 0 .1 8 3 1 0 .9 8 1 7 0 .7 3 7 8 0 .7 7 5 3 0 .0 1 1 6 a g m m [2 7 ] 0 .9 8 6 0 .9 4 9 5 0 .7 9 5 2 0 .2 0 4 8 0 .0 5 0 5 0 .9 8 7 8 0 .7 9 5 2 0 .8 6 5 5 0 .0 0 2 5 g m m [2 8 ] 0 .9 8 6 5 0 .9 5 1 9 0 .7 9 9 8 0 .2 0 0 2 0 .0 4 8 3 0 .9 8 8 3 0 .7 9 9 8 0 .8 6 9 2 0 .0 0 2 4 p fi n d er[2 9 ] 0 .9 8 3 4 0 .9 4 9 8 0 .7 4 5 9 0 .2 5 4 1 0 .0 5 0 2 0 .9 8 4 9 0 .7 4 5 9 0 .8 3 5 6 0 .0 0 2 4 p ro p o sed 0 .9 8 8 8 0 .9 6 3 6 0 .8 0 6 7 0 .1 9 3 3 0 .0 3 6 4 0 .9 8 9 9 0 .8 0 6 7 0 .8 7 8 2 1 .6 0 e -0 3 t a b le 3 q u an titative r esu lt for c am p u sr o a d test seq u en ces a lg o rith m a ccu ra cy p recisio n r eca ll f n r f d r n p v t n r f 1 s co re f a llo u t u n ifo rm + g a u ssia n [2 5 ] 0 .9 8 5 4 0 .8 2 0 4 0 .9 1 1 6 0 .0 8 8 4 0 .1 7 9 6 0 .9 9 5 3 0 .9 1 1 6 0 .8 6 3 6 0 .0 1 0 6 b a y esia n [2 6 ] 0 .9 7 9 9 0 .6 8 6 8 0 .6 9 3 7 0 .3 0 6 3 0 .3 1 3 2 0 .9 8 9 8 0 .6 9 3 7 0 .6 9 0 2 0 .0 1 0 6 a g m m [2 7 ] 0 .9 4 0 8 0 .3 1 7 2 0 .2 8 5 0 .7 1 5 0 .6 8 2 8 0 .9 6 6 8 0 .2 8 5 0 .3 0 0 3 0 .0 2 8 6 g m m [2 8 ] 0 .9 7 0 2 0 .6 9 2 7 0 .5 4 3 5 0 .4 5 6 5 0 .3 0 7 3 0 .9 7 9 8 0 .5 4 3 5 0 .6 0 9 1 0 .0 1 0 8 p fi n d er[2 9 ] 0 .9 3 7 3 0 .3 6 8 8 0 .5 0 .5 0 .6 3 1 2 0 .9 7 5 3 0 .5 0 .4 2 4 5 0 .0 4 1 5 p ro p o sed 0 .9 9 4 1 0 .9 2 8 0 .9 7 2 6 0 .0 2 7 4 0 .0 7 2 0 .9 9 8 3 0 .9 7 2 6 0 .9 4 9 8 0 .0 0 4 6 66 n.a. zainuddin etal: adaptive background modeling for dynamics background t a b le 4 q u an titative resu lt for f ou n ta in test seq u en ces a lg o rith m a ccu ra cy p recisio n r eca ll f n r f d r n p v t n r f 1 s co re f a llo u t u n ifo rm + g a u ssia n [2 5 ] 0 .9 8 5 8 0 .7 0 0 3 0 .5 6 5 4 0 .4 3 4 6 0 .2 9 9 7 0 .9 9 0 7 0 .5 6 5 4 0 .6 2 5 7 0 .0 0 5 2 b a y esia n [2 6 ] 0 .9 8 1 4 0 .6 3 4 1 0 .8 8 3 2 0 .1 1 6 8 0 .3 6 5 9 0 .9 9 6 4 0 .8 8 3 2 0 .7 3 8 2 0 .0 1 5 6 a g m m [2 7 ] 0 .9 8 6 3 0 .6 6 8 6 0 .6 2 4 3 0 .3 7 5 7 0 .3 3 1 4 0 .9 9 2 3 0 .6 2 4 3 0 .6 4 5 7 0 .0 0 6 3 g m m [2 8 ] 0 .9 8 9 2 0 .7 1 2 7 0 .7 1 6 7 0 .2 8 3 3 0 .2 8 7 3 0 .9 9 4 5 0 .7 1 6 7 0 .7 1 4 7 0 .0 0 5 6 p fi n d er[2 9 ] 0 .9 8 6 2 0 .7 9 1 5 0 .5 2 0 9 0 .4 7 9 1 0 .2 0 8 5 0 .9 8 9 1 0 .5 2 0 9 0 .6 2 8 3 0 .0 0 3 2 p ro p o sed 0 .9 8 9 9 0 .9 0 4 1 0 .7 9 5 4 0 .2 0 4 6 0 .0 9 5 9 0 .9 9 2 6 0 .7 9 5 4 0 .8 4 6 3 tex tb f3 .1 0 e -0 3 t a b le 5 a v erage of gold stan d ard m easu rem en ts p a ra m eters fo r a ll test seq u en ces a lg o rith m a ccu ra cy p recisio n r eca ll f n r f d r n p v t n r f 1 s co re f a llo u t u n ifo rm + g a u ssia n [2 5 ] 0 .9 8 6 4 0 .8 1 4 7 0 .7 9 8 2 0 .2 0 1 9 0 .1 8 5 4 0 .9 9 3 0 .7 9 8 2 0 .8 0 3 9 0 .0 0 7 2 b a y esia n [2 6 ] 0 .9 8 0 9 0 .7 6 2 7 0 .7 7 9 1 0 .2 2 0 9 0 .2 3 7 3 0 .9 9 0 1 0 .7 7 9 1 0 .7 6 4 4 0 .0 1 0 2 a g m m [2 7 ] 0 .9 7 5 0 .7 1 7 1 0 .6 2 6 6 0 .3 7 3 4 0 .2 8 2 9 0 .9 8 4 1 0 .6 2 6 6 0 .6 6 8 5 0 .0 1 0 1 g m m [2 8 ] 0 .9 8 4 0 .8 2 5 0 .7 2 0 9 0 .2 7 9 1 0 .1 7 5 0 .9 8 8 6 0 .7 2 0 9 0 .7 6 8 1 0 .0 0 5 3 p fi n d er[2 9 ] 0 .9 7 2 6 0 .7 2 7 7 0 .6 8 1 8 0 .3 1 8 2 0 .2 7 2 3 0 .9 8 6 7 0 .6 8 1 8 0 .6 9 0 4 0 .0 1 5 6 p ro p o sed 0 .9 9 0 .9 3 9 5 0 .8 5 5 3 0 .1 4 4 7 0 .0 6 0 5 0 .9 9 2 4 0 .8 5 5 3 0 .8 9 3 8 0 .0 0 2 9 advances in systems science and application (2016) vol.16 no.2 67 method. 5 conclusion this research presented a novel real time based background modeling technique for the scene with varied dynamic background. the proposed algorithm are tested in four complex scenes and compared with five recent algorithm studied. experimental results show that our proposed algorithm is able to give a consistent, accurate and precise detection results. the common problems in detection method are also successfully handled. acknowledgment this research work is supported by fundamental research grant scheme (frgs13030-0271) under the ministry of higher education of malaysia (mohe). references [1] m. piccardi.(2004), “background subtraction techniques: a review”, in ieee international conference on systems, man and cybernetics, the hague, netherlands. [2] m. xiao, c. han and x. kang. (2006), “a background reconstruction for dynamics scenes”, in ieee international conference on information fusion, florence, italy. [3] z. hou and c. han. (2004), “a background reconstruction algorithm based on pixel intensity classification in remote video surveillance system”, in 7th international conference on information fusion, stockholm, sweden. [4] c. kamath and s. c. cheung. (2004), “robust techniques for background subtraction in urban traffic video”, proceedings of spie, vol. 5308, no. 1, pp.881-892. [5] t. gao, j. zhang, w. gao and z. liu. (2009), “a robust technique for background subtraction in traffic video”, in 15th international conference on neural information processing, auckland, new zealand. [6] m. h. hung, j. s. pan and h. c. hsieh. (2014), “a fast algorithm of temporal median filter for background subtraction”, journal of information hiding and multimedia signal processing, vol. 5, no. 1, pp.33-41. [7] s. asif, a. javed and m. irfan. (2014), “human identification on the basis of gaits using time efficient feature extraction and temporal median background subtraction”, international journal image, graphics and signal processing, vol. 3, no. 2, pp.35-42. 68 n.a. zainuddin etal: adaptive background modeling for dynamics background [8] m. xiao, c. han and x. kang. (2006), “a background reconstruction for dynamic scenes”, in ieeee 9th international conference on information fusion. [9] e. l. rubio and r. m. baena. (2011), “stochastic approximation for background modelling”, computer vision and image understanding, vol. 115, no. 6, pp.735-749. [10] c. stauffer and w. e. l. grimson. (1999), “adaptive background mixture models for real-time tracking”, in ieee computer society conference, colorado, 1999. [11] s. mukerjee and k. das.(2013), “an adaptive gmm approach to background subtraction for application in real time surveillance”, international journal of research in engineering and technology, vol. 2, no. 1, pp.25-29. [12] m. nimse, s. varma and s. patil. (2014), “shadow removal using background subtraction and reconstruction”, international journal of emerging technology and advanced engineering, vol. 4, no. 4, pp.324-327. [13] h. asaidi, a. aarab, m. bellouki. (2014), “shadow elimination and vehicles classification approaches in traffic video surveillance contex”, journal of visual languages and computing, vol. 12. no.4, pp.333-345. [14] a. elgammal, r. duraiswami, d. harwood and l. davis. (2002), “background and foreground modeling using nonparametric kernel density estimation for visual surveillance”, proceeding of the ieee, vol. 90, no. 7, pp.1151-1162. [15] j. lee and m. park. (2012), “an adaptive background subtraction method based on kernel density estimation”, journal on the science and technology of sensors and biosensors, vol. 12, no. 9, pp.12279-12300. [16] j. b. cruz, a. t. ali and e. l. dagless. (1993), “a temporal smoothing technique for real-time motion detection”, in 5th computer analysis international conference, budapest, hungary. [17] c. ridder, o. munkelt and h. kirchner. (1995), “adaptive background estimation and foreground detection using kalman filter”, in proceedings of international conference on recent advances in mechatronics, istanbul, turkey. advances in systems science and application (2016) vol.16 no.2 69 [18] z. hou and c. han. (2004), “a background reconstruction algorithm based on pixel intensity classification in remote video surveillance system”, proceedings of the seventh international conference on information fusion, vol. 2. [19] l. cao and y. jiang. (2013), “an effective background reconstruction method for video objects detection”, in international conference on networking and distributed computing. [20] n. chaki , s. h. shaikh and k. saeed. (2014), exploring image binarization techniques, springer. [21] t. fawcett. (2006), “an introduction to roc analysis”, pattern recognition letters, vol. 27, pp.861-874. [22] d. powers. (2007), evaluation: from precision, recall and f-factor to roc, informedness, markedness & correlation, flinders university of south australia, adelaide. [23] k. kumar and s. agarwal. (2013), “an efficient hierarchical approach for background subtraction and shadow removal using adaptive gmm and color discrimination”, international journal of computer applications, vol. 75, no.7. [24] l. li, w. huang, i. y.-h. gu and q. tian. (2004), “statistical modeling of complex backgrounds for foreground object detection”,ieee transactions on image processing, vol. 13, no. 11, pp.1459-1472. [25] z. zivkovic and f. v. d. heijden. (2006), “efficient adaptive density estimation per image pixel for the task of background subtraction”,pattern recognition letters, vol. 27, no. 7, [26] c. stauffer and w. grimson.(2000), “learning patterns of activity using real-time tracking”, ieee transactions on pattern analysis and machine intelligence, vol. 22, no. 8, pp.747-757. [27] c. wren, a. azarbayejani, t. darrell and a. pentl. (1997), “pfinder: real-time tracking of the human body”, ieee transactions on pattern analysis and machine intelligence, vol. 19, no. 7, pp.780-785. corresponding author n.a. zainuddin can be contacted at: fiqahzainuddin@gmail.com adv syst sci appl 2017; 17(2); 14-28 published online http://ijassa.ipu.ru/ojs/ijassa/article/view/12 copyright ©2017 assa. adv. in systems science and appl. (2017) economic security under disturbances of foreign capital kurt schimmel 1) , sifeng liu 2,3) , jeananne nicholls 1) , nicholas a. nechval 4) , jeffrey yi-lin forrest 1) 1) school of business, slippery rock university, slippery rock, pa16057, u.s.a. e-mail: kurt.schimmel@sru.edu; jeananne.nicholls@sru.edu; jeffrey.forrest@sru.edu 2) institute for grey system studies, nanjing university of aeronautics and astronautics, nanjing 211106, pr china 3) centre for computational intelligence, de montfort university, leicester le1 9bh, uk; e-mail: sfliu@nuaa.edu.cn 4) department of mathematics, baltic international academy, lomonosov street 4, riga lv1019, latvia e-mail: nechval@junik.lv abstract: considering spillover effects of foreign capital, this paper establishes a method for the receiving economy to constantly monitor and forecast the movement of foreign capital within itself so that methods of regulation could be established to guarantee the health and stable development (or the security) of the domestic economy. methods of control theory are employed to model the movement of foreign capital, to estimate the initial state of foreign capital’s movement. on the method established and the systemic yoyo model, this paper introduces ways to prevent foreign capital from adversely impacting the receiving economy for establishing theoretical guidance for insuring the healthy development of the economy. this paper provides several practically useful strategies that could counter disturbances caused by foreign capital within the receiving economy and protect the economic security of the system in order to avoid the disastrous aftermath of the currency war that occurs along with large scale withdraw of foreign capital that was initially invested within the system in friendly terms. keywords: movement of money, economic stability, method of control, estimation, systemic yoyo model, strategies of regulation 1. introduction sequel to [7], this paper establishes a theoretical model on how to monitor the movement of foreign capital within a nation’s economy in order to keep the economic security of the economy intact. in particular, this paper is organized as follows: section 2 establishes a theoretical model for monitoring the dynamics of foreign capital within a nation’s economy. section 3 studies how to estimate the state of motion of foreign capital within the domestic economy. section 4 fills the need of estimating the initial state of foreign capital’s movement. section 5 explores some potential measures useful for countering the fluctuations of foreign capital within a national economy. section 6 concludes the presentation of this work. 15 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) 2. a model for monitoring dynamic foreign capital within an economic system the receiving economic system foreign investment is really a double-edged sword. it can not only provide sufficient capital, advanced technological equipment, and excellent management experience for the development of the economic system, but also seriously constrain the development of domestic enterprises, weaken the innovation capacity of the economic system, and create serious security risks for the economic system [7]. hence, practically speaking, there is an urgent need to appropriately formulate and implement strategic measures regarding foreign investments in order to accelerate the economic development by sufficiently taking advantage of foreign capital while lowering the security risk that accompanies the influx of foreign capital to a certain level or within the controllable range of the economic system. this section attempts to provide theoretically sound suggestions regarding the design of foreign capital policies and strategies for the receiving economic system by modeling the scale, regions, and areas of foreign investment, and by constructing a control-theory model for foreign investment. generally speaking, the investment within an economic system by foreign capital may be maintained at a certain stable speed without being affected by regional economic development cycles, financial shocks, and major changes in the environment of the economic system. for example, since 1992, china has been continuously attracting a large scale of foreign investment [1]. and only in some very particular situations, such as major financial shocks, drastic changes in the political environment of the economic system, etc., the invested foreign capital would show speculative, withdrawal and other behaviors. from what has been discussed in [7], it follows that foreign capital affects the economic security of the receiving economic system in three different ways: one is the magnitude of the invested foreign capital, two is the geographic distribution of the foreign investments, and three the difference between the actual distribution of invested foreign capital and the need distribution for foreign capital investment of the economic system. in terms of the magnitude of invested foreign capital, a rising magnitude increases the holding of foreign currencies of the economic system while it also increases the internal base money supply. that might affect adversely the effectiveness of the adopted monetary policies of the economic system. additionally, the rising amount of foreign investment also increases the uncertainty of how the economic system would be influenced by the external world. because of its drive for profit, foreign capital might leave the economic system in large scales within a short period of time when the economic system experiences shocks from the external environment. that exit could cause disastrous consequences for the healthy development of the economic system, and such effects had been sufficiently manifested in the southeast asian financial crises of the 1990s. in terms of the geographic distribution of foreign investments, when the imbalance in the geographic distribution of foreign investment goes up, the internal development imbalance within the economic system will expand, which might cause chaos in the regional development of the economic system so that the orderly domestic allocation of resources will be disrupted, such wasteful phenomena of resources as duplicated constructions, production overcapacity, etc., will likely to appear. in terms of the difference between the actual distribution of invested foreign capital and the need distribution for foreign capital investment of the economic system, it creates an imbalanced local distribution of regional industries within the economic system. that causes the internal industrial structure of the economic system to lose its balance so that when affected by external uncertainties, the development of the economic system will be badly economic security under disturbances of foreign capital 16 copyright ©2017 assa. adv. in systems science and appl. (2017) hindered by production imbalances and insufficient productions of some industries. for example, in some of the emerging economies foreign investment is relatively concentrated in a few industries. when such an economic system is affected by financial shocks, a quick withdraw of foreign capital generally leaves behind damaging effects for the development and stability of the economic system [3]. therefore, in the following, we use r(k) to stand for the magnitude of foreign investment at time tk within the economic system, (k) – the geographic distribution of foreign investment at time tk, and (k) for the difference between the actual industrial distribution of foreign investment and the actual distribution of areas that need foreign investment at time tk. here, we utilize the difference (k) to measure the effect of foreign capital on the economic security of the economic system at time tk . next, let us establish a control and monitor equation of these three variables in order to materialize the estimation and regulation of the state of foreign capital at any chosen time moment. additionally, in order to produce meaningful anticipation on the state of investment of foreign capital, let us use ( )r k , ( )k , and ( )k to respectively denote the rate of change of the investment magnitude r(k) of foreign capital, the rate of change of the geographic distribution of foreign investment (k), and the rate of change in the difference (k) between the actual industrial distribution of foreign investment and the actual distribution of areas that need foreign investment at time tk. if we let x(k) be the state vector of foreign capital within the economic system, where then we can use the following state equation to describe the movement of foreign capital within the economic system: ( ) ( 1)x k x k  (2.1) where  is a matrix defined below with t representing the time interval between consecutive observations, 1 0 1 1 0 1 1 0 1 t t t                     . however, due to the existence of various disturbances and errors in real life, the movement, when seen as a system, of foreign capital within the economy is constantly affected somehow. for example, no matter how we establish the dynamic equation, it will not be an accurate 17 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) expression of the realistic system; because of the interferences of the environment that is external to the economic system, exchange rate, politics, etc., the state of foreign capital movement is affected by countless many random factors. let us aggregate all these various known and unknown factors into one concept: process (or dynamic) noise. in order to more realistically reflect the true dynamics of the foreign capital movement, let us add a noise term ( )w k that describes all kinds of interferences into the previous equation. so, the state equation in eq. (2.1) can be rewritten as follows: ( ) ( 1) ( 1)x k x k tw k    (2.2) where 1 2 3( ) [ ( ), ( ), ( )]tw k w k w k w k is the random interference of the foreign capital movement system within the economic system with w1, w2, and w3 respectively being the random accelerations in the directions of r, , and , and matrix t is given below where t is the time interval between two consecutive observations 2 2 2 / 2 / 2 / 2 t t t t t t t                     . of particular interest is that the interference {w(k)} as introduced above is a stochastic sequence, which comprehensively (and approximately) reflects the interference as experienced by the foreign capital when it moves within the economic system. although there might be a large number of factors that interfere with the movement of foreign capital within the economic system, the effect of each factor might be quite small, which is likely unclear to the regulator of the system and cannot be readily measured. even so, the total effect of all the interfering factors on the system’s dynamics can be described by using a stochastic sequence (or a stochastic process if the movement system is modeled by a continuous model). however, it is still not enough for us to only employ eq. (2.2) to describe the movement of foreign capital within the economic system. for the general scenario, the economic system can acquire and observe some traces of the movement of the foreign capital by using various methods of statistics. however, what is acquired and observed tends to be only about some components of the state variable or the values of some linear transformations of the state variable. for example, when using econometric data to observe or measure the state of foreign capital, only the magnitude, locations of investment, and industries of investment of foreign capital can be observed and measured. however, the acceleration and other parameters of movement of foreign capital are difficult to measure directly. let us denote the data values that are actually observed as y(k), where economic security under disturbances of foreign capital 18 copyright ©2017 assa. adv. in systems science and appl. (2017) hence, the observation (or measurement) equation that describes the movement system of foreign capital within the economic system can be established as follows: ( ) ( )y k hx k (2.3) where 1 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 1 0 h           it can be expected that when econometric data are employed as observations of the movement of foreign capital, there is always error. such error could originate either from the inaccuracy of the observation equation or the method used during the statistical process and the random error that exists in the available data [6]. let us refer this sort of error as observation (or measurement) noise. to realistically reflect the relationship between the observed variables regarding the movement of foreign capital within the economic system, let us add a random variable v(k) into eq. (2.3) to describe such random interference. hence, the observation (or measurement) equation becomes: ( ) ( ) ( )y k hx k v k  (2.4) where v(k)=[v1(k), v2(k), v3(k)] t , and v1(k), v2(k), and v3(k) are respectively the observed values of r, , and  at time moment tk, and { v(k)} is a stochastic sequence. now, if we combine eq. (2.2), which is the state equation that describes the movement system of foreign capital within the economic system, and the observation equation (4), which describes how the movement system of foreign capital is measured, we then obtain the following state equation of the foreign capital that expresses how the capital moves within the economic system: ( ) ( 1) ( 1) ( ) ( ) ( ) x k x k tw k y k hx k v k         (2.5) because the movement system of foreign capital within the economic system, as described by equ. (5), contains random interferences, we will refer the system described by eq. (2.5) as a stochastic control system. for each of such systems, all obtained information (data) is “polluted” by noise. in order to obtain relatively more accurate information of the state, we will have to make our best estimation of the true state of the system from the available information (that is generally a sequence of observations) with interfering noise, which might even be incomplete. additionally, for stochastic control systems, the process of movement itself and the observation of the system’s output are all contaminated with noise [6]. so, to study the movement of the system quantitatively we must first describe the interfering noise quantitatively. from the knowledge of stochastic processes, it follows that the distribution of the given stochastic sequence determines all the statistical characteristics of the sequence [2]. however, generally it is difficult to fully know the distributional structure of a stochastic sequence. therefore, when solving practical problems, it will be mostly enough to simply know some of the main statistical characteristics of the given stochastic sequence. assume that {v(k)} is a stochastic sequence. if for any positive integer n and t, and n arbitrary time moments k1 < k2 <…< kn, v(k1),…, v(kn) and v(k1+t),…, v(kn+t) have the same joint distribution function, then the sequence {v(k)} is known as stationary. if the mean of the stationary stochastic sequence {v(k)} is 0, and its auto-covariance satisfies 19 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) 2( , ) ijr i j   , 0, 1, ij if i j if i j      (2.6) then {v(k)} is known as a white noise sequence. here, a white noise sequence can be fathomed as a pure stochastic sequence, satisfying that the values of any two consecutive terms are independent. if in eq. (2.5), which describes the movement of foreign capital within the economic system, we know the state equation and the statistical characteristics of the noise, then we can move ahead to estimate the state of the movement system of foreign capital. in order to establish a formula to estimate x(k) in the state equation in eq. (2.5), let us impose some conditions on the noises {w(k)} and{v(k)}. first, let us assume that the observation noise {v(k)} and the process dynamic noise {w(k)} are completely independent of each other; otherwise, they can no longer be known as noise. that is, we have [ ] 0t k je w v  , for any 0, 0k j  . (2.7) secondly, we assume that both {w(k)} and {v(k)} are zero mean white noise stochastic sequences. that is, these sequences satisfy the following: for {w(k)}, we have [ ( ) ( )] ( )t kje w k w j q k  , where 1 2 3 ( ) ( ) ( ) ( ) q k q k q k q k           , 1, 0, kj k j k j      which is known as the covariance matrix of the process noise, and 1( )q k , 2 ( )q k , and 3( )q k are respectively the stochastic acceleration variances of the objective along the directions r, , and  at time moment tk, which can be selected based on how freely foreign capital can move within the economic system and how accurate the estimate of the state needs to be. and for{v(k)}, we have [ ( ) ( )] ( )t kje v k v j r k  , where 1 2 3 ( ) ( ) ( ) ( ) r k r k r k r k           which is known as the covariance matrix of the measurement; and r1(k), r2(k), and r3(k) are respectively the observation noise variances of the objective along the directions of r, , and  at time moment tk. their values are related to the accuracy of the observation statistics and the state of foreign capital. as a matter of fact, if the noises are not of zero means or if {w(k)} and {v(k)} are not independent stochastic sequences, etc., then necessary mathematical treatments will be needed. of course, that will surely involve a lot of computational complexities. economic security under disturbances of foreign capital 20 copyright ©2017 assa. adv. in systems science and appl. (2017) other than the afore-mentioned basic assumptions about the random inference variables, we also need to assume that the initial state 0x is also a random vector satisfying 0 0[ ]e x x , 0 0 0 0 0[( )( ) ]te x x x x p   , and that 0x and kw , kv are independent. 3. estimate the state of motion of foreign capital at the initial time moment t=0, the ideal estimate of 0x is of course 0 0 0x̂ ex x  . if we assume that we have already obtained 0x , and let 0 0x̂ x , then the covariance matrix of the estimation error of the state is 0 0 0 0 0[( )( ) ]tp e x x x x   . by employing the thinking logic of mathematical induction, we can derive the needed recursive formula. assume that we have obtained the observations up until time moment k: y1,…,yk, and the optimal (unbiased) estimate ˆ kx of the state kx at time moment k. before the next new measurement data value yk+1 becomes available, to estimate the state xk+1 at time moment k+1 we have to start with ˆkx by making use of the evolutionary rule described in eq. (2.5). because wk is a zero-mean white noise, we cannot know what value it takes. so, the most reasonable choice is let wk=0, the mean value. on this basis, we obtain the predicted estimation as follows: 1| 1, ˆ ˆ k k k k kx x  . (3.1) because 0kew  , ˆ k kex ex , and 1| 1, 1, 1 ˆ ˆ k k k k k k k k kex ex t ew ex      , we conclude that 1| ˆ k kx  is an unbiased estimate of 1kx  . so, an optimal estimation is provided by eq. (3.1). and by using the measurement expression in eq. (2.4) and by letting 1 0kv   (for the same reason as above), we obtain the following predicted output value that corresponds to 1| ˆ k kx  : 1| 1 1| 1 1, ˆ ˆ ˆ k k k k k k k k ky h x h x      . (3.2) at the time moment t=k+1, we immediately obtain the actual output value yk+1 of the economic system. by comparing this actual value with the predicted value 1| ˆ k ky  , we can compute the error of the predicted output: (3.3) this expression includes not only the comprehensive information of the system’s noise kw , measurement noise vk, filtering error, and other random variables, but also some new information about the state xk+1 as the measurement value yk+1 at time moment k+1 becomes available. hence, let us refer to 1|k ky  as a vector of new information (innovation). it is now natural for us to think about employing the new information at moment k+1 to correct the predicted state 1| ˆ k kx  . that is, we can now derive a better estimate (filtering estimate) for the state at moment k+1. so, the estimation for the state at moment k+1 should consist of the following two parts: one is to derive the predicted estimate 1| ˆ k kx  of the state based on the observations up to time moment k, and the other is to correct the prediction 1| ˆ k kx  by employing the newly collected observation at moment k+1. when these two parts are put together, we establish our estimation for the state xk+1 at the time moment k+1. symbolically, we have 1 1| 1 1 1 1| ˆ ˆ ˆ[ ]k k k k k k k kx x k y h x        , (3.4) 21 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) where kk+1 is referred to as a gain matrix or a correction factor. it is a real matrix. the key for our optimal filtering is to select correctly the gain matrix kk+1 so that the estimation 1 ˆ kx  of the state obtained from eq. (3.4) has the minimum covariance matrix of error. that is, we have 1 1 1 1 1 ˆ ˆ[( )( ) ] mint k k k k kp e x x x x        . (3.5) next, we need to establish a formula from which we can produce the optimal gain matrix kk+1 that satisfies eq. (3.5) so that a complete recursive computational scheme can be developed for the purpose of estimating the state of movement of foreign capital within an economic system. let 1 1 1 ˆ k k kx x x    . then the covariance matrix pk+1 of the filtering error can be written as 1 1 1[ ]t k k kp e x x   . from 1 1 1 1 1| 1 1 1 1| 1| 1 1 1 1 1 1| 1| 1 1 1| 1 1 1 1| 1 1 ˆ ˆ ˆ[ ( )] ˆ[ ] [ ] [ ] k k k k k k k k k k k k k k k k k k k k k k k k k k k k k k k k k x x x x x k y h x x k h x v h x x k h x v i k h x k v                                           (3.6) where 1| 1 1| ˆ k k k k kx x x    , it follows that 1 1 1 1 1 1| 1 1 1 1 1| 1 1 1 1 1| 1| 1 1 1 1 1 1 1 1 1| 1 1 1 1 1| [ ] [(( ) )(( ) ) ] [ ] [ ][ ] [ ] [ ] [ ] [ t k k k t k k k k k k k k k k k k t t t t k k k k k k k k k k k k t t k k k k k k k k k p e x x e i k h x k v i k h x k v i k h e x x i k h k e v v k i k h e x v k k e v x                                             1 1][ ]t t k k ki k h  because 1| 1, ˆ ˆ k k k k kx x  only depends on 0x̂ and y1,y2,…,yk, from the knowledge of statistics, it follows that 1 1|[ ] 0t k k ke v x   and 1| 1[ ] 0t k k ke x v   . so, the formula for computing the covariance matrix pk+1 of the filtering error can be simplified as follows: 1 1 1 1| 1 1 1 1 1[ ] [ ]t t k k k k k k k k k kp i k h p i k h k r k            , (3.7) where 1| 1| 1|[ ]t k k k k k kp e x x   is referred to as the covariance matrix of the prediction error, and 1 1 1[ ]t k k kr e v v   . if we expand eq. (3.7) and regroup the terms, where, without causing confusion, we omit all the subscripts of 1kh  , 1kk  , and 1kr  , we have 1 1| 1| 1| 1| 1| 1| 1| 1|( ) t t t t t k k k k k k k k k t t t t k k k k k k k k p p khp h k khp p h k krk p k hp h r k khp p h k                    (3.8) our purpose here is to derive the gain matrix k such that pk+1 reaches its minimum. therefore, we rewrite eq. (3.8) as a complete square, such as the form of ( )( )tks a ks a  . because ( )( )t t t t t tks a ks a kss k aa ksa as k      , (3.9) by comparing eq. (3.9) and (3.8), we can see that because 1| t k khp h is a non-negative definite matrix and r is positive definite, 1| t k khp h r  is a positive definite matrix. so, from matrix theory [4], it follows that each positive definite matrix can be expressed as the square of a certain symmetric positive definite matrix. so, let us assume economic security under disturbances of foreign capital 22 copyright ©2017 assa. adv. in systems science and appl. (2017) 1| t t k khp h r ss   (3.10) where s is a symmetric positive definite matrix. next let sa t =hpk+1|k, where matrix a actually exists because the square matrix s is of full rank. by substituting eqs. (3.9) and (3.10) into eq. (3.8) we obtain the following: 1 1| 1| ( )( ) t t t t t k k k t t k k p p kss k ksa as k p ks a ks a aa             (3.11) from eq. (3.11), we see that only the second term on the right hand side has something to do with the gain matrix k, while this term is a non-negative definite matrix. hence, to make pk+1 reach its minimum, we only need to select such a matrix k so that the second term in eq. (3.11) is equal to zero. so, we have 1 1 1 1| 1| 1 1| 1| ( ) ( ) t t t t k k k k t t k k k k k as p h s s p h ss p h hp h r               (3.12) if we select k according to eq. (3.12), then the covariance matrix pk+1 of the filtering error is 1 1 1| 1| 1| 1| 1| 1|( ) t k k k k k k k k k k k k k p p aa p as hp p khp i kh p                 (3.13) next, let us derive a computational formula for the covariance matrix 1|k kp  of the prediction error. from the formula for pk+1|k, it follows that 1| 1| 1| 1| 1, 1| 1, 1, 1, 1, 1, 1, 1, 1, 1, [ ] {[ ][ ] } [ ] [ ] t t k k k k k k k k k k k k k k k k k k t t t t k k k k k k k k k k k k t t k k k k k k k k k k p e x x e x t w x t w e x x t e w w t p t q t                              (3.14) by combining what is obtained above, for the system in eq. (2.5), we have obtained all the recursive formulas necessary for estimating the state of movement of foreign capital within the economic system. however, to employ this theory to practically estimate the magnitude of foreign capital’s investment, regional distribution, and differences between industries within the economic system at any chosen time moment, we still need to estimate the initial state. 4. estimate the initial state of foreign capital’s movement from our discussions above, it follows that once we have the knowledge of the initial state of the movement of foreign capital within the economic system of our concern or a good estimation of that initial state, we then can estimate the state of foreign capital’s movement system step by step. to this end, the general method is to determine the initial state x(0) by using the prior (known) knowledge and then calculate the initial covariance matrix p(0) by employing the degree of accuracy of x(0). in this section, we assume that we have the initial observations at time moments t1 and t2. we then roughly estimate the state of the foreign capital’s movement system at moment t2 and its variance. these values will be used as the initial state of foreign capital’s movement within the economic system. in particular, assume that the initial two observations are respectively 23 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) 1 2 3(1) [ (1), (1), (1)]ty y y y and 1 2 3(2) [ (2), (2), (2)]ty y y y . then we calculate the following estimation of the initial state: 1 2 3 4 5 6 ˆ ˆ ˆ ˆ ˆ ˆ ˆ(0) [ (0), (0), (0), (0), (0), (0)]tx x x x x x x , where 2 1 ˆ (0) (2)i ix y  , 2 ˆ (0) [ (2) (1)] /i i ix y y t  , 1,2,3i  , (4.1) and 1 2 2 2 3 (0) ˆ ˆ(0) [( (0) (0))( (0) (0)) ] (0) (0) t p p e x x x x p p              , where 2 (2) (2) / (0) (2) / [ (1) (2)] / i i i i i i r r t p r t r r t        , 1,2,3i  , (4.2) in the following, we use a sub-matrix p1(0) as an example to illustrate the detailed steps of computation. because 1 1 1 1 ˆ (0) (2) (2) (2)x y x v   , we have 2 1 1 1 1 1 1 1 1 1 1 ˆ (0) [ (2) (1)] / [ (2) (2) (1) (1)] / {[ (2) (1)] [ (2) (1)]}/ x y y t x v x v t x x v v t           because 2 1 1(2) [ (2) (1)] /x x x t  , we have 1 1 1 ˆ(2) (0) (2)x x v   , 2 2 1 1 ˆ(2) (0) [ (2) (1)] /x x v v t    1 1 1 1 1 2 2 2 2 1 1 1 1 1 1 1 1 2 1 1 1 ˆ ˆ(2) (0) (2) (0) (0) ˆ ˆ(2) (0) (2) (0) (2) (2) [ (2) (1)] / [ (2) (1)] / (2) (2) / (2) / [ (1) (2)] / t t x x x x p e x x x x v v e v v t v v t r r t r t r r t                                                  similarly, we can obtain the other sub-matrices of eq. (4.2). as of this point in our presentation, we have established a complete theory on how to estimate the state of movement of foreign capital within the economic system. and on the basis of the estimations, we can design appropriate strategies and measures to counter the adverse effects of foreign investments while insuring the economic security of the economic system. economic security under disturbances of foreign capital 24 copyright ©2017 assa. adv. in systems science and appl. (2017) 5. how disturbances of foreign capital affect economic security because of the entry of foreign capital, the money supply within the receiving economic system is increased, making the system have additional resources to produce so that the economic strength of the system consequently grows rapidly. however, with the accumulation of foreign capital to a large magnitude, when the foreign capital suddenly withdraws, a severe shock to the economic activities of the economic system will be created. for example, let us look at a fictitious economic system whose financial system is open to the outside world. assume that at a certain time moment, the nation’s total money supply is $50,000, including domestically invested foreign capital of $7,500. that is, the foreign investment occupies 15% of the entire national money supply. with the additional money supply, as created by the foreign capital, the circulation speed of money within the economic system accelerates, which indicates that this nation spends more on various economic activities and on the promotion of its economic level. if at a later time moment the capital invested by foreign merchants is suddenly withdrawn from the economic system, then it will inevitably, to a great extent, cause the liquidity of money within the economic system to drop suddenly, affecting adversely over 15% of all the economic activities within the economy. hence, the economic system has to strengthen the monitoring of foreign capital in order to prevent the disastrous aftermath created by the currency war into which the initial friendly investment was evolved by sudden large scale withdraw. the essential reason why an economic system introduces foreign capital is to stimulate its economic development. that is to say, there is a pre-defined purpose within the economic system for how the entering foreign capital will develop. this pre-defined purpose is referred to as the ideal trajectory of the movement of foreign capital. on the other hand, the fundamental reason for foreign capital to enter the economic system is to make profit, which includes taking advantage of the promised preferential conditions, low-priced resources available from within the economic system, developing the market resources of the new economy, etc. in the process of materializing the desired profit, the actual state of movement of foreign capital most likely deviates from the ideal trajectory that is pre-defined by the economic system. when such deviation expands, it might gradually erode the economic security of the economic system that is additional to the significant damage that might occur to the economic system when a collective withdraw of foreign capital happens at certain time moment. based on what has been discussed and analyzed earlier, the economic system could acquire “truthful estimations” of the state of domestic movement of foreign capital and possible trends of future development. so, for the policy makers of the economic system, they must design and introduce a series of policies to guide the “truthful estimations” to approach the ideal trajectory in order to materialize the desired control of the adverse effect of foreign capital on the economic security of the economic system. such dynamic adjustment on the part of policy makers could also potentially prevent a currency war created by large scale sudden withdraws of foreign capital although the capital initially entered the economic system in friendly terms. when applying a strategy to regulate the state of movement of foreign capital, it is inevitable that some of the limited resources of the economic system will have to be occupied and exhausted. so, a control strategy is considered ideal if it could not only reach its purpose of control within a limited time interval but also consume as little amount of resource as possible. to this end, as a form of expression for energy consumption, quadratic functions will be used due to their excellent properties. these functions can describe not only how resources are wasted by implementing excessive control, but also the distance from the target as caused by insufficient 25 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) control. therefore, in the following, we will choose a quadratic function as our evaluation standard to measure control strategies. let us define the objective function as follows: 1 0 ˆ ˆ[( ( 1) ( 1)) ( )( ( 1) ( 1)) ( ) ( ) ( )] n t t n k j x k x k q k x k x k u k r k u k           where n stands for the final time moment when the actual state ( )x n could approach the predefined ideal state ( )x n “perfectly”, ( )q k the effect on the economic system as caused by the difference between the “truthful estimation” and the “ideal trajectory” at moment k, r(k) the price that has to be paid for adopting control strategy u(k) at moment k. because this objective function contains the effect on the economic system of the difference between the “truthful estimation” and the “ideal trajectory” and the price paid to implement the control strategy u, the policy makers should design an optimal control strategy that makes the previous objective function reach its minimum. however, the design of any policy strategy is constrained by the resources available within the economic system, that is, u(k)ω, where ω stands for the space of constraints of all strategies. so, the purpose of the entire control problem is to find a control sequence u*(k)ω such that ( *(0), *(1), , *( 1)) min nj u u u n j    because the effect of policies can lead to changes in the state of foreign capital movement system, the “truthful estimation” ˆ( 1)x k  is a function of policy u(k). according to the necessary conditions for a function to reach its extreme values, we have 0k k j u    . that is, ˆ( 1) ˆ( )( ( 1) ( 1)) ( ) ( ) 0 ( ) t x k q k x k x k r k u k u k         so, we have 1 ˆ( 1) ˆ( ) ( ) ( )( ( 1) ( 1)) ( ) x k u k r k q k x k x k u k          theoretically speaking, the feasible region of the constraint variables is bounded, the objective function is nonnegative and convex. therefore, generally there is an optimal control sequence u*(k)ω such that jn reaches its minimum. however, because it is difficult to quantitatively or analytically describe the feasible region of the constraint variables, it is also difficult to provide a formal expression for an optimal strategy. even so, we will try to develop a general description of the necessary control strategy by using the state equation and possible evolution tendencies of foreign capital movement within the economic system combined with the systemic yoyo model. therefore, in the following we will provide several strategies based on logical analysis and reasoning strategies that could counter disturbances caused by foreign capital within the economic system and protect the economic security of the system in order to avoid the disastrous aftermath of the currency war that occurs along with large scale withdraw of foreign capital that was initially invested within the system in friendly terms. first, moderately monitor and supervise the activities of foreign investments, and strengthen the control of foreign capital if it attempts to enter the virtual economy, such as the financial market, stock market, etc. although the assets of foreign capital that enters the virtual economy economic security under disturbances of foreign capital 26 copyright ©2017 assa. adv. in systems science and appl. (2017) could provide investment capital for the economic system, such capital can often perturb the operation of the domestic financial system through unconventional ways, disrupt the smooth operation of the system, and withdraw from the system as soon as a high rate of return is acquired within a short period of time by making use of its scale and other advantages. such short-term rapid movement of capital can cause turbulence to the otherwise normal operation of the domestic financial market of the economic system, which in practice materializes damage to the domestic financial system. the southeast asian financial storm of the year of 1998 was an example of how foreign capital acquired short-term high rates of returns by making use of the imperfections of the regional financial systems. this financial storm ransacked the economic achievements of the past several decades of the relevant regions within a short time period, while leaving the development of these economies on the verge of collapse [3]. therefore, the receiving economic system must strengthen its supervision of those foreign capital investments that intend to enter the domestic financial market or the virtual economy, crackdown any illegal entry behaviors of foreign capital, and improve the control mechanism over the foreign capital that flows into any of the non-real economies. at the same time, receiving economic systems must constantly perfect emergency plans and formulate control measures for temporary, abnormal capital movements by considering the potentially most severe scenarios based on the logic of balancing the abnormal inflow and outflow of foreign investment capital. second, implement a flexible exchange rate mechanism in order to maintain investment grade foreign capitals while expelling those that are speculative in nature. among all the capital of foreign investments, the portion that enters into the real economy is the one needed by the domestic economic system for its economic development. this capital intends to acquire relatively long-term profits by valuing the development potential of the economic system and by actively investing in the economic system. on the other hand, this capital could be relied on for the construction and development of the real economy and for its relatively long-term stay within the economic system. however, the entry of the speculative portion of foreign capital into the domestic economic system is mainly for the purpose of obtaining speculative profits by making use of the imperfections that exist in the market place and the financial system of the receiving economy. hence, the economic system should employ its flexible exchange rate mechanism to discourage the influx of such foreign capital and to convert the influx of such foreign capital into investments in the real economy. for example, assume that merchant c is an international speculator and that he is upbeat about the stock market of the particular economic system and enters the stock market with an amount of us$1 billion. if the exchange rate between the u.s. dollar and the domestic currency is 1:5, then his entry into the stock market represents 5 billion units of the local currency. after obtaining a short-term speculative profit in the stock market, c wants to close out his positions. if at this time moment the monetary authority of the economic system has somehow changed the exchange rate to 1:6 by making use of their flexible exchange rate policy, then c’s original speculation amount of 5 billion units of the local currency can only be converted to us$5/6 billion = us$0.833 billion, where us$0.167 billion has evaporated. that is, the flexible exchange rate mechanism has discouraged c’s speculation in the stock market, while the short-term varying exchange rate has little adverse effect on the foreign capital that is invested in the real economy. hence, when a relatively large amount of foreign capital enters into the economic system, the receiving economy could employ a flexible exchange rate mechanism to discourage speculative capital while retaining the capital that is invested in the real economy. 27 k. schimmel1, s. liu, j. nicholls, n.a. nechval, j. yi-lin forrest copyright ©2017 assa. adv. in systems science and appl. (2017) third, implement a flexible foreign investment policy to guide foreign investment capital into the target industries and geographic regions by making use of the economic system’s market demand, cost advantage, and other resources. the profit-driven characteristics of foreign capital determine where and which industries foreign capital will enter; from how foreign capital flows into the economic system to obtain its desired profit, the investment types of foreign capital can generally be classified into the following five categories: cost seeking, policy seeking, resource seeking, market seeking, and technology seeking. with increasing economic strength of the economic system, the accumulation of investment capital, and the rising of labor cost, the foreign invested enterprises of policy-seeking and cost-seeking types, and some of the market-seeking types with relatively weak competitiveness will have to leave the economic system or remake themselves into different types, while those foreign capital assets of market-seeking and technology seeking types might very well occupy the high end positions of their individual markets and supply chains. hence, the receiving economic system of foreign capital should sufficiently utilize its consumer market and related fundamental advantages, such as opening its domestic market, already established supporting industries, existing resource conditions, etc., to guide the investment capital of market-seeking, technology seeking, and other types into the high technology areas that are needed by the development of the economic system. these investments will not only stay within the economic system for the long term, but also strengthen their collaboration with the economic system and improve the position of the economic system within the machining and manufacturing chain. additionally, by making use of the differences between investment policies internal to the economic system, such as the various preferential policies in terms of taxation, land use, regulations, etc., as introduced by different regions with varying degrees of capital scarceness, the receiving economic system could direct the foreign invested enterprises of policy-seeking and cost-seeking types, and some of the investment assets of market-seeking type with less competitiveness to migrate into regions with less foreign investment capital. these different investment policies could also be used to improve the uneven investment distribution within the economic system while decreasing the disparity between the development gradient forces of different regions. last, promote the capability of innovation and creativity of the enterprises within the economic system and the public of the entire society. according to the systemic yoyo model, the competition between foreign capital and domestic business entities makes some domestic firms been absorbed by foreign capital, while some others been pushed and pulled (or bullied), where some of these bullied entities grow with the intensified competition and eventually become able competitors while some others are eliminated out of the existence. those economic organizations that are either absorbed by foreign capital or eliminated are mainly those that are less competitive, and the competitiveness of an organization is mainly originated from the organization’s creativity. when an economic organization possesses a strong capability of innovation, the yoyo field of this organization has a strong force of attraction and a steady ability to stand and take shocks from the environment. that strong capability of innovation can help promote the magnitude of the organization’s yoyo field, while carrying the field to a higher level. as soon as the magnitude of the yoyo field of the organization rises to a new level, it would generally have acquired additional energy to better counteract the attraction and pressure of the yoyo field of foreign capital, while it might conversely absorb the yoyo field of foreign capital and make this external field a part of itself. in other words, improving the endogenous creativity is the fundamental way for the economic system to increase the magnitude of its yoyo field. to this end, the economic system can improve its internal innovation environment by fostering a economic security under disturbances of foreign capital 28 copyright ©2017 assa. adv. in systems science and appl. (2017) good atmosphere for innovation, increasing investment in innovation, encouraging innovative behaviors, and providing innovators with conditions to materialize their creativities. on top of all these suggested remedies the economic system would be able to reach a new level of creativity of the entire society. 6. concluding remarks because foreign capital can adversely affect the security of the receiving economy [3,7], this paper addresses the following challenging question: while taking advantage of foreign investments, how can a nation protect itself against the potential disastrous aftermath in case some of the foreign investments were actually part of an active currency war against the nation? to this end, this paper establishes a theoretical model for monitoring the movement and dynamics of foreign capital within a nation’s economy. then potential measures that are useful for countering the adverse effects of foreign investments on the national economy are developed. references [1] cheng, l.k. & kwan, y.k. (2000). what are the determinants of the location of foreign direct investment? the chinese experience. j. int. economics, 51, 379-400. [2] cinlar, e. (2013) introduction to stochastic processes (reprint edition), new york, ny: dover publications. [3] forrest, j. (2014). a systems perspective on financial systems, new york, ny: crc press, an imprint of taylor and francis. [4] horn, r.a. & johnson, c.b. (2012). matrix analysis (2nd edition), cambridge, uk: cambridge university press. [5] lin, y. (2008). systemic yoyo model: some impacts of the second dimension, new york, ny: auerbach publications, an imprint of taylor and francis. [6] liu, s.f. & lin, y. (2006). grey information: theory and practical applications, london, uk: springer. [7] schimmel, k., liu, s.f., nicholls, j. & forrest, j. yl. (2017). effects of foreign capital on economic security, advances in systems science and applications. advances in systems science and applications (2013) vol.13 no.2 100-115 set-theoretic design and analysis of l systems xiaojun duan1, bing ju1 and yi lin2 1department of mathematics and systems science, national university of defense technology, changsha 410073, pr china 2department of mathematics, slippery rock university, slippery rock, pa 16057 usa abstract from the perspective of set theory, this paper establishes a rigorous mathematical definition of l systems, and proves the common characteristics of simple l systems by providing a theoretical framework for a systematic research of l systems. additionally, explanations for the characteristics of systems’ emergence based on l systems are provided, and some design methods of l systems are developed and interesting cases of design are constructed. this work is the first of its kind that investigates the l system using a rigorous mathematical definition based on set theory. keywords set theory, l system, characteristics of systems 1 introduction although the figures produced out of fractals are generally complicated, the description of fractals can be quite straightforward. among the commonly utilized methods are the l-system and the ifs (iterated function system). in terms of the fractals they respectively describe, the l-system is simpler than the if system, where the former contains simple iterations of character strings, while the latter is much more complicated in this regard. aristid lindenmayer, a hungary biologist, introduced the lindenmayer system, or l system in short, as the mathematics theory for describing the growth of a plant[1]. it is a kind of subsequent string replacement system; its theory focuses on the topology of plants, and attempts to describe the adjacency relations between cells or between larger plant modules. while lindenmayer and others proposed the initial solution[1], prusinkiewicz used turtle graphics to implement a lot of fractal shapes and herb model based on turtle shapes[2]. his works made turtle the most commonly used form and explanatory schemes of l systems. in order to avoid the models constructed using l systems inflexible, eichhorst and others proposed the concept of stochastic l-systems to enhance the flexibility of the earlier l-models[3]. herman generalized the concept of l-systems to a context sensitive model in order to establish associations between different modules of the plant model[4]. then, lindenmayer introduced parameters to make l-systems even more forthright and efficient. the most typical applications of l-systems are done by scholars at calgary university, canada. they were involved in the parametricalization of l-systems, the establishment of differential l-system, open advances in systems science and applications (2013) vol.13 no.2 101 l-system, and other relevant theoretical researches[2]. additionally, they developed the plant simulation software l-studio[2,5]. because there has not been any rigorous mathematical definition established for l systems, all published studies on the subject have stayed only at the level of innovative designs without much theoretical support. on the other hand, as shown in (lin, 1999)[6], set theory is a good tool for the theoretical framework of systemics. in this paper, starting from the basics of set theory, we will develop a rigorous mathematical definition for simple l systems. on the basis of this definition, we establish some of the general properties simple l systems satisfy so that our constructed theoretical framework is expected to provide the needed fundamental ground for further investigation of l systems. 2 the definition of simple l systems the fact is that l-systems represent a formal language; it can be divided into three classes: 0l system, 1l system, and 2l system. a 0l system stands for an l system that is context free. that is, the behavior of each element is solely determined by the rewriting rule; the current state of the system has something to do only with the state of the immediate previous time moment and has nothing to do with any surrounding element. among l systems are 0l systems the simplest; that is why each 0l system is also referred to as a simple l system. each 1l system stands for a context sensitive l system that considers only one single-sided grammatical relationship. the current state of the system has something to do with not only the state of the immediate previous time moment but also the state of the elements on one side, either left associated or right associated. each 2l system is also a context sensitive l system that, different of 1l systems, it considers grammatical relationships from both sides. that is, the current state of the system is related to not only the state of the immediate previous time moment, but also the states of the surrounding elements. it represents a method that is most sensitive to the context. these three classes of l systems are further divided into deterministic and random l systems depending on whether or not the rewrite rules are deterministic. let s = {s1, s2, . . . , sn} be a finite set of characters of a language, say, english, and s∗ the set of all strings of characters from s. because s ⊂ s∗, s∗ is a non-empty set. ∀α, β ∈ s∗, define that α = β ⇔ both α and β are identical, meanings that they have the same length and order of the same characters. now, define the addition operation ⊕ on s∗ as follows: α ⊕ β = αβ = the string of characters of those in α followed by those in β, ∀α, β ∈ s∗. the scalar multiplication on s∗ is defined as follows: ∀α ∈ s∗, k ∈ z+ = the set of all whole 102 xiaojun duan:set-theoretic design and analysis of l systems numbers, kα = αα . . . α︸ ︷︷ ︸ k times . then, the following properties can be shown: (1)both addition and scalar multiplication defined on s∗ are closed; (2)(α⊕ β)⊕ γ = α⊕ (β ⊕ γ),∀α, β, γ ∈ s∗; (3)k(lα) = (kl)α,(α⊕ β)⊕ γ = α⊕ (β ⊕ γ),∀α, β, γ ∈ s∗;and (4)(k + l)α = kα⊕ lα,∀α ∈ s∗, k, l ∈ z+. before we develop a rigorous definition of simple l systems, let us first look at a specific mapping s∗ → s∗, known as l mapping. 2.1 the l mapping definition 2.1(production rules). for any si ∈ s, 1 ≤ i ≤ n, if an ordered pair (si, α) ∈ s × s∗ can be defined, then this pair defines a relation pi from si to α, known as a production rule from s to s∗. definition 2.2 (same class production rules). for a given si ∈ s, if there are r ≥ 1 production rules p (1) i , p (2) i , . . . , p (r) i defined for si such that ∀j ̸= k(1 ≤ j, k ≤ r), p (j) i (si) ̸= p (k) i (si), then p (1) i , p (2) i , . . . , p (r) i are referred to as r same class production rules of the character si, and pi = {p(1)i , p (2) i , . . . , p (r) i } the set of same class production rules of the character si. definition 2.3 (l mapping). suppose that m(≥ n) production rules p = n∪ i=1 pi from s to s∗ are given, where each pi is non-empty and stands for the same class production rules of the character si ∈ s, 1 ≤ i ≤ n. for any group of production rules p1, p2, . . . , pn ∈ p , satisfying pi ∈ pi, i = 1, 2, . . . , n, the set φ = {p1, p2, . . . , pn} defines a mapping s∗ → s∗, still denoted φ, by φ : sk1sk2 . . . skr → pk1(sk1)pk2(sk2) . . . pkr(skr) (1) ∀sk1sk2 . . . skr ∈ s∗, ki ∈ {1, 2, . . . , n}, 1 ≤ i ≤ r, and r stands for the length of the character string. then, this mapping φ : s∗ → s∗ is referred to as an l mapping. note: from the definition of l mappings, it follows that the set of m(≥ n) production rules from s to s∗: p = {p(1)1 , . . . , p (r1) 1 , p (1) 2 , . . . , p (r2) 2 , . . . , p(1)n , . . . , p(rn)n } (2) can define (r1r2 . . . rn) many l mappings from s to s∗, where n∑ i=1 ri = m. the set of all l mappings determined by the set p is denoted by φ = {φ1, φ2, . . . , φr1r2...rn}. proposition 2.1 (properties of l mappings). for any α, β ∈ s∗ and m,n ∈ z+, each l mapping φ satisfies the following properties: (i) φ(α) ∈ s∗ is uniquely defined; (ii) if φ|s is surjective s → s, then φ|s must be bijective; advances in systems science and applications (2013) vol.13 no.2 103 (iii) if α = sk1sk2 . . . skp ∈ s∗, then φ(α) = φ(sk1)φ(sk2) . . . φ(skp); (iv) φ(α⊕ β) = φ(α)⊕ φ(β); (v) φ(mα) = mφ(α); (vi) ∀φ, ϕ ∈ φ, φϕ(mα⊕ nβ) = mφϕ(α)⊕ nφϕ(β); proof. because both (i) and (ii) are evident, it suffices to show (iii) (vi). (iii) for any α = sk1sk2 . . . skp ∈ s∗, the definition of the l mappings implies that φ(sk1sk2 . . . skp) = pk1(sk1)pk2(sk2) . . . pkp(skp) (3) and, ∀ski ∈ s∗(1 ≤ i ≤ p), we have φ(ski) = p(ski) (4) hence, φ(α) = φ(sk1sk2 . . . skp) = pk1(sk1)pk2(sk2) . . . pkp(skp) = φ(sk1)φ(sk2) . . . φ(skp). (iv) for any α, β ∈ s∗, from the addition operation on s∗ and property (iii), it follows that φ(α⊕ β) = φ(αβ) = φ(α)φ(β) = φ(α)⊕ φ(β) (5) (v) for any α ∈ s∗ and m ∈ z+, from the definition of scalar multiplication on s∗ and property (ii), it following that φ(mα) = φ(αα...α︸ ︷︷ ︸ m times ) = φ(α)φ(α)...φ(α)︸ ︷︷ ︸ m times = mφ(α) (6) (vi) for any φ, ϕ ∈ φ, by employing properties (iv) and (v), we obtain: φϕ(mα⊕ nβ) = φ(mϕ(α)⊕ nϕ(β)) = mφϕ(α)⊕ nφϕ(β) (7) qed. 2.2 simple l-systems with the concept of l-mappings in place, let us now look at how to define simple l-systems. according to the classification of simple l systems, deterministic and random simple l systems, we now establish the relevant definitions by using the concept of l mappings. definition 2.4(deterministic simple l systems). let s = {s1, s2, ..., sn} be a finite set of characters and s∗ the set of all strings of characters from s. assume that φ : s∗ → s∗ is a given l mapping. for any given initial string ω ∈ s∗, the system that is made up of the nth iterations φn(ω), n ≥ 1, is referred to a deterministic simple l (d0l) system (of order n), denoted by the ordered triplet ⟨s, ω, φ⟩. definition 2.5(random simple l systems). let s and s∗ be the same as in 104 xiaojun duan:set-theoretic design and analysis of l systems definition 2.4, φ = {φ1, φ2, ..., φk} a set of k l mappings from s∗ to s∗, and ξ the l mapping randomly drawn from φ such that the probability p (ξ = φi) for ξ to be φi is πi, where k∑ i=1 πi = 1. for a given initial string of characters ω ∈ s∗ , the system that is made up of the nth generalized iterations ξnξn−1...ξ1(ω), the composite of the mappings ξ1, ..., ξn−1, ξn, n ≥ 1, is referred to as a random simple l (r0l) system (of order n), where ξ1, ξ2, ..., ξn are independent and identical distributions. this system is written as the following ordered quadruple ⟨s, ω,φ, π⟩. 3 properties of simple l systems from the definitions, it follows that each simple l system is produced by iterations or generalized iterations of l mappings. hence, the properties of simple l systems are determined by those of l mappings. theorem 3.1(property of linearity). each simple l system is a linear system. proof. assume that the deterministic nth order simple l system is given as l1 = ⟨s, ω, φ⟩, and the random nth order simple l system l2 = ⟨s, ω,φ, π⟩, where n ≥ 1. to show both l1 and l2 are linear systems, it suffices to show that both l1 and l2 systems respectively satisfy the superposition principle. (i)we use mathematical induction to prove that φn satisfies the superposition principle. when n = 1, ∀α, β ∈ s∗, k1, k2 ∈ z+, properties (iv) and (v) of l-mappings imply that φ(k1α⊕ k2β) = k1φ(α)⊕ k1φ(β) (7) so, φ satisfies the superposition principle. assume that when n = k, k ≥ 1,φk satisfies the superposition principle. that is, ∀α, β ∈ s∗, k1, k2 ∈ z+, the following holds true: φk(k1α⊕ k2β) = k1φ k(α)⊕ k2φ k(β) (8) then when n = k + 1 , we have φk+1(k1α⊕ k2β) = φ(φk(k1α⊕ k2β)) = φ(k1φ k(α)⊕ k2φ k(β)) = k1φ k+1(α)⊕ k2φ k+1(β) (9) hence, ∀n ≥ 1, φn satisfies the superposition principle. (ii)we show that ξnξn−1...ξ1 satisfies the superposition principle. from the definition of random simple l systems, it follows that ξi , 1 ≤ i ≤ n, stands for the l mapping φj , 1 ≤ j ≤ n , where n is the total number of elements in the set φ , randomly selected from φ at step i. because each φi ∈ φ , 1 ≤ i ≤ n , satisfies the superposition principle, from advances in systems science and applications (2013) vol.13 no.2 105 property (vi) of l mappings, it follows that ξnξn−1...ξ1 satisfies the superposition principle. qed. from theorem 3.1 it follows that in terms of a simple l system, when its order n is relatively large, to shorten the computational time, one can decompose relatively long strings of characters into sums of several shorter strings, which can be handled using parallel treatments to obtain the state of the next moment. theorem 3.2 (property of fixed points). the restriction of the l mapping φ in a d0l system on s is bijective from s to s, if and only if for any α ∈ s∗, there is a natural number k such that φk(α) = α. before proving this result, let us first look at the following lemma. lemma 3.1 if the restriction of the l mapping φ on s is a bijection from s to s , then ∀si ∈ s , ∃ki ∈ z+ such that φki(si) = si , 1 ≤ i ≤ n. proof. by contradiction, assume that ∃s ∈ s, ∀k ∈ z+ , φk(s) ̸= s. then let si1 = φ(s), si2 = φ2(s), ..., sin = φn(s) so we have s ̸= si1 , s ̸= si2 , ..., s ̸= sin (10) step 1: from the first (n − 1) inequalities in equ. (10) and that fact that φ|s : s → s is bijective, it follows that φ(s) ̸= φ(si1), φ(s) ̸= φ(si2), ..., φ(s) ̸= φ(sin−1). that is, si1 ̸= si2 , si1 ̸= si3 , ..., si1 ̸= sin (11) step 2: from the first (n − 2) inequalities in equ. (11) and that fact that φ|s : s → s is bijective, it follows that φ(si1) ̸= φ(si2), φ(si1) ̸= φ(si3), ..., φ(si1) ̸= φ(sin−1). that is, si2 ̸= si3 , si2 ̸= si4 , ..., si2 ̸= sin (12) by continuing this procedure, we obtain at step (n− 1) that sin−1 ̸= sin (13) therefore, s ̸= si1 ̸= si2 ̸= ... ̸= sin . it means that there are n + 1 different elements in the set s. however, s has only n elements. a contradiction. so, the assumption that ∃s ∈ s, ∀k ∈ z+, φk(s) ̸= s does not hold true. in other words, ∀si ∈ s , ∃ki ∈ z+ such that φki(si) = si , 1 ≤ i ≤ n. qed. in the following, let us prove theorem 3.2. (⇒)for α = sm1sm2 ...smp ∈ s∗, from lemma 3.1, it follows that ∀smi ∈ s, ∃ki ∈ z+, φki(smi) = smi ,1 ≤ i ≤ p . let k =< k1, k2, ..., kp >. so, φk(α) = φk(sm1sm2 ...smp) = φk(sm1)φ k(sm2)...φ k(smp) = φk1(sm1)φ k2(sm2)...φ kp(smp) = α 106 xiaojun duan:set-theoretic design and analysis of l systems so, there is a natural number k such that φk(α) = α . (⇐)if ∀α ∈ s∗, there is natural number k such that φk(α) = α, let us take α = s1, s2, . . . , sn. then ∃k1, k2, ..., kn such that φk1(s1) = s1, φ k2(s2) = s2, ..., φ kn(sn) = sn (14) according to property (ii) of l mappings, to show the restriction of φ on s is a bijection from s to s, it suffice to prove that φ|s : s → s is surjective. again, we prove this end by contradiction. assume that ∃s ∈ s , ∀si ∈ s, 1 ≤ i ≤ n ,φ(si) ̸= s . then we have that ∀k ≥ 1 , φk(s) ̸= s , which contradicts with equ. (14). so, φ|s : s → s is surjective. hence, φ|s : s → s is bijective. qed. from theorem 3.2, it follows that in terms of d0l systems, if the restriction φ|s : s → s of l-mapping φ is a bijection, then there are at most k =< k1, k2, ..., kp > different states. therefore, to produce complicated fractal figures using d0l systems, one should avoid using any such l mapping whose restriction on s is a bijection from s to s. theorem 3.3(property of fixed points). the restriction of the l mapping φ in a r0l system on s is a bijection from s to s, if and only if for any α ∈ s∗ , there is a natural number k such that φk(α) = α. proof. although as a r0l system, the probability of different mappings at each step is different, the maximum number n of total states of each step is fixed. since the length p of the mapping is finite, the number of overall states is k = np at most. so there exists a integer k = np to satisfy φk(α) = α. other parts of the proof are similar to those of theorem 3.2 and omitted. qed. remark:in terms of numbers, the finite states of rol are more than those of dol. however, from theorem 3.3, to follows that in order to produce complicated fractal figures using r0l systems, one should still avoid using any such l mapping that its restriction on s is a bijection from s to s. 4 an explanation of holistic emergence of systems using dol although the mechanism for a simple l system appears is quite straightforward, it clearly shows the attribute of holistic emergence of systems. so, simple l systems can be employed as an effective tool to illustrate the emergence of systems’ wholeness. in the following, we will design a simple binary system to explain the emergence of systems’ wholeness. in the cartesian coordinate system r2, assume that the initial location of particle a is (0,0) and its initial angle is 0. let f stand for moving forward one unit step 1, and + turning an angle of π/3 counterclockwise. now, let us construct the following d0l system l2 = ⟨s, ω, φ⟩, where s = {f,+},ω = f , and φ = { f → f ++f ++f + → + advances in systems science and applications (2013) vol.13 no.2 107 the first iteration φ(ω) produces: f++f++f; the second iteration φ2(ω) leads to: f++f++f ++ f++f++f ++ f++f++f; . . .; when the nth order, n ≥ 1 , l system l2 is applied on the particle a, one can obtain the output state as shown in fig.1. as shown in fig.1, such an nth order l system, a binary sysfig.1 the output state of the l2 system tem, is very simple. its function can be comprehended as follows: drive particle a from the starting point at (0,0) to make 3n−1 counterclockwise turns along the triangular trajectory. now, let us employ this simple binary system to illustrate the emergence of systems’ wholeness: (1) the whole is greater than the sum of its parts; and (2) the function of the whole is more than the sum of the functions of the parts. for our purpose, by the word “sum”, it means the collected pile of parts without any interactions between the parts. as for the word “function”, it is understood as follows: as long as a system is identified, its functionality is a physical existence; however, any function of the system has to be manifested through the system’s act on a specific object that is external to the system. so, to investigate the overall function of a system and the sum of the functions of its parts, one needs to consider the respective effects of the system as a whole and each of its parts on an external object. through analyzing their effects on this object, one can compare the overall effect of the system and the aggregated effect of the parts. before we illustrate the emergence of the system’s wholeness, let us first establish the following assumptions: (a) the whole can always be divided into several distinguishable parts in terms of components, attributes, or functionalities; (b) similar parts of the system (or parts with similar attributes) satisfy the 108 xiaojun duan:set-theoretic design and analysis of l systems additive property, while different parts (or parts with different attributes) do not comply with this property. here, the additivity stands for the algebraic additivity and is different of the meaning of the “sum” in “the sum of parts”; and (c) when studying the sum of parts’ functionalities, if an identified part does not have any function, then the functionality of this part is seen as “0”. illustration 1: the whole is greater than the sum of its parts let us symbolically denote the constructed nth order simple l system l2 = ⟨s, ω, φ⟩, n ≥ 1 , as φn(f ), and sum of parts as ∑ i si , where si , i = 1, 2, ..., stands for a division of the whole φn(f ), including all characters. in the following, we will discuss from two different angles according to the following different divisions of the whole. (1) the whole is greater than the sum of all the elements (the whole is divided using components). if we treat the system l2 as one containing only two components (operations) f and +, then in the nth order l2 system, there are 3n components f and 2(3n -1) components +. from the basic assumptions, it follows that f and + respectively satisfy the additive property. their algebraic sums are respectively written as s1 = 3n(f ) and s2 = 2(3n − 1)(+) . then, the sum of the elements of the l2 system can be expressed as 2∑ i=1 si = s1s2 or 2∑ i=1 si = s2s1. applying the output of the elements’ sum 2∑ i=1 si on the particle a produces a line segment of length 3n , while applying the output of the whole φn(f ) on the particle a creates a regular triangle with each side’s length 1. the whole constitutes a figure of the 2-dimensional space, possessing a special structure, while the sum of the elements represents a line segment of the 1-dimensional space without any qualitative mutation. that is to say, the whole has a spatial structural effect that is not shared by the sum of the parts, therefore, φn(f ) > 2∑ i=1 si . (2) the whole is greater than the sum of its parts, where the whole is divided using attributes. let us treat the l2 system as being composed of 3n unit vectors of the plane: i⃗1, i⃗2, ..., i⃗3n . that is, we see the vectors i⃗1, i⃗2, ..., i⃗3n as having the same attributes. then this l2 system is made up only of these 3n components of the same attributes. now, we desire to show φn(f ) > 3n∑ k=1 i⃗k . evidently, 3n∑ k=1 i⃗k = 0. so, when the output of the sum 3n∑ k=1 i⃗k of the parts is advances in systems science and applications (2013) vol.13 no.2 109 applied on the particle a, it is like that no operation is ever applied on a so that the particle a is fastened at the origin without moving. because the output of the whole φn(f ) is a string of characters, which is equivalent to a sequence of operations with a before and after order, when φn(f ) is applied on the particle a, this particle will repeatedly travel counterclockwise along the triangle and eventually return to the origin. that is to say, the whole possesses a structural effect of time that is not shared by the sum of the parts, therefore, we have φn(f ) > 3n∑ k=1 i⃗k . in fact, the essential difference between the whole and the sum of parts is that the whole has some structural effect in terms of space or time, while the sum of parts does not have. for instance, when n bricks are used to build a house, the whole stands for a building along with the spatial structure of rooms, etc. however, when the components of this building are divided, because there is only one kind of component, the bricks, the sum of the parts satisfies the algebraic additivity and is equal to the n bricks, which do not have the spatial structure of the house. illustration 2: the functionality of the whole is greater than the sum of parts’ functionalities. let us first introduce a new kind of set [s], where each element is allowed to appear more than once. to avoid creating any conflict with the conventional set theory, we will only apply the operation of drawing elements out of the set [s] with the following convention: if a non-empty set [s] contains element s at least twice, then drawing one s from [s] means that we take any of the elements s’s. for example, [s] = {a, a, b, a, b} . then, drawing an a from [s] means that we take any one of the elements a’s. for the nth order simple l2 system, n ≥ 1 , the individual characters (operations) f and + stand for the smallest units of functionalities. divide this system into (3n+1−2) parts, which include 3n functional units f and 2( 3n-1) functional units +. the set of these (3n+1-2 ) functional units is written as set [s] . we first consider the effect of the whole on the particle a. the function of this l2 system is to order the elements of [s] according to some specific rules. then, the effect of the system on a is force the particle a to counterclockwisely travel along the triangle 3n−1 times. now, let us look at the effect on a of each part. when the set [s] is given, the sum of the parts’ functions can be understood as follows: there are a total of m operations, each of which takes mi arbitrary elements from [s] without replacement to act on a, until all the elements in [s] are exhausted. evidently, as long as the order of the elements taking out of the set [s] is different from that of φn(f ) , the total effect of these elements that are individually taken out of [s] will not reach that of the l2 system. therefore, the whole possesses a functionality the sum of the parts does not share. that is, the function of the whole is greater 110 xiaojun duan:set-theoretic design and analysis of l systems than the sum of the parts’ functions. in fact, the essential difference between the function of the whole and the sum of parts’ functions is that the whole has an organizational effect, while the sum of the parts does not have. 5 the design of simple l systems 5.1 the basic graph generation principle of simple l systems in essence each l system is a system that rewrites strings of characters. its working principle is quite simple. if each character is seen as an operation and different characters are seen as distinct operations, then strings of characters can be employed to generate various fractal figures. that is, as long as strings of characters can be generated, one is able to produce figures. the character strings of l systems that are used to generate figures can be made up of any recognizable symbols. for example, in the design of programs, the symbols f, -, and + can be used respectively so that “f” means move one unit length forward from the current location and draw a segment, “-” stands for turning clockwise from the current direction a pre-determined angle, and “+” turning counterclockwise from the current direction another pre-determined angle. when generating character strings, start from an initial string and replace the characters of this string by substrings of characters according to the predetermined rules. that completes the first iteration. then, treat the resultant character string from the first iteration as the mother string and replace each character in this string by strings determined by the rules. by continuing this procedure, one can finish the required iterations of an l system, where the length of the resultant string is controlled by the number of iterations. 5.2 a fractal structure design based on d0l in terms of a tree, it stands for a fractal structure. in particular, each trunk carries a large amount of branches, and each branch has an end point, representing a figure with one staring point and many ending points. this fact implies that when one draws a branch to its end, he has to return his drawing pen to draw other structures. let us take the following conventions: “f” stands for moving forward a unit length 1, and “+” turning an angle of π/8 clockwisely, “-” turning an angle of π/8 counterclockwisely, and the characters within “[ ]” represent a branch; when the characters within a pair of [ ] are implemented, return to the position right before “[” and maintain the original direction, and then carry out the characters after “]”. assume that the starting point is at (0,0) on the complex plane and the initial direction at π/2 . now, we design a d0l system g1 = ⟨s1, ω1, φ1⟩ as follows: s1 = {f,+,−, [, ]}; ω1 = f ;and advances in systems science and applications (2013) vol.13 no.2 111 fig.2 the fractal structures: a tree dancing in breeze (of different iteration steps) fig.3 the fractal structures: a standing tree (n = 5, α = π/8, with design form f→ff+[+f[−ff−]−f]− [−f[−ff+]f]) 112 xiaojun duan:set-theoretic design and analysis of l systems φ1 =  f → ff + [+f− f− f]− [− f + f + f] + → + − → − [ → [ ] → ] then, we can produce the fractal structures as shown in fig.2. we can obtain other illustrations by just changing the design forms as fig.3-5. fig.4 the fractal structures: a tree towards sun (n = 5, α = π/8, with design form f→ff+[+f[−f−f]−f]− [−f[f−ff+f]f]) fig.5 the fractal structures: a floating grass ball (n = 4, α = π/8, with design form f →ff+[[+f−ff−f]− [−f+ff+f]] ) 5.3 a fractal structure design based on r0l systems in nature, the forms plants take are not invariant. even for the same kinds of plants, their shapes can vary from one plant to another. such varieties are caused advances in systems science and applications (2013) vol.13 no.2 113 by the effects of the environment. fig.6 the different fractal structures for a r0l system with each iteration unchanged (n = 4, α = π/16 ) in terms of the simulation effects of plants, the figures created by using d0l systems seem to be quite stiff. under the prerequisite of maintaining the main characteristics of plants, in order to generate varieties on the details, we can utilize the plants? structures produced out of r0l systems. the advantage of these figures is that these simulated plants are more real-life like and much closer to the true forms of natural plants. to this end, let us design a r0l system g2 = ⟨s1, ω1,φ, π⟩ as follows: s1 = {f,+,−, [, ]}; ω1 = f ; φ = {φ1, φ2, φ3}; where φ1 =  f → f[+f]f[− f]f + → + − → − [ → [ ] → ] ,φ2 =  f → f[+f]f[− f[+f]] + → + − → − [ → [ ] → ] , 114 xiaojun duan:set-theoretic design and analysis of l systems φ3 =  f → ff[− f + f + f] + [+f− f− f] + → + − → − [ → [ ] → ] and π = p (ξ = φi) = 1/3, i = 1, 2, 3 . then the fractal structures can be produced. for instance, fig.6 shows the different fractal structures of a r0l system with each iteration unchanged. fig.7 shows the different fractal structures of a r0l system where each iteration is different. fig.7 the different fractal structures of a r0l system with each iteration different (n = 5, α = π/16 ) 6 summary although the design principle underlying the l systems is quite straightforward, these systems can be employed to produce many complicated fractal patterns. after many years of research, l systems have evolved from the original rewrite systems of characters to such capable systems that can describe complex 3-dimensional systems. they have evolved from the simplest d0l systems to random l systems, and then to open l systems. the l systems have provided simulations of fractal structures that have become over time much closer to real-life like, and have been employed as an important tool for creating virtual plants. from the point of view of set theory, this paper first establishes a rigorous mathematical definition of simple l-system, and then proves the common characteristics of simple l-systems. by doing so, this paper developed the badly needed fundamental ground for further investigation of l systems. it is expected to be applicable to form the theoretical framework of other l systems. acknowledgements advances in systems science and applications (2013) vol.13 no.2 115 this work is supported by the natural science foundation of china (60974124), the program for new century excellent talents in university and the projectsponsored by srf for rocs, sem in china. references [1] lindenmayer, a. (1968), “mathematical models for cellular interaction in development”, journal of theoretical biology, vol.18, pp.280-289. [2] prusinkiewicz, p, lindenmayer, a. (1996), the algorithmic beauty of plants, london: springer. [3] eichhorst, p, savitch, w. j. (1980), “growth functions of stochastic lindenmayer systems”, information and control, vol.45, pp.217-228. [4] herman, g.t, rozenberg, g. (1975), developmental systems and languages, amsterdam: north-holland publishing company. [5] prusinkiewicz, p, hammel, m, mjolsness, e. (1993), “animation of plant developmen”, acm siggraph, pp.351-360. [6] lin, y. (1999), “general systems theory: a mathematical approach”, new york: kluwer academic and plenum publishers. corresponding author xiaojun duan can be contracted at: xjduan@nudt.edu.cn microsoft word 6-刘秋梅.doc 34-40 advances in systems science and applications (2010), vol.10, no.1 issn 1078-6236 international institute for general systems studies, inc. the research on the early warning system of soybean import dependence degree based on decision tree∗ qiumei liu1 and jun meng*1,2 1science college, northeast agricultural university, harbin heilongjiang, 150030, china 2national study center of soybean engineering, northeast agricultural university, harbin heilongjiang, 150030, china email: merd@mail.neau.edu.cn(*corresponding author) abstract the early warning, which is usually built on the base of forecast, is the premise of remove-warning. the paper presents a new warning way-decision tree warning which can warn all kinds of conditions directly. first, the paper sets up a soybean early warning index system by analyzing soybean market of china in nearly the recent ten years. then, some regulations are extracted according to the warning sign index’s effects on the warning alert index in the way of decision tree. at last, the early warning system of the soybean import dependence degree is set up. keywords decision tree soybean import dependence degree early warning system 1. introduction looking back the soybean export and import conditions, china is always the soybean net export country before 1995, but from then on, china has became a soybean net import country until to now. at present, soybean market of china is in turmoil, and seriously dependents on the foreign soybean. foreign soybean has formed a great shock to the soybean production of china. soybean import dependence degree is identified to the ratio of soybean net import to soybean demand amount. according to the constructed warning sign index system, the soybean import dependence degree is warned, and the degree is divided into non-warning area, gentle warning area, medium warning area and serious warning area. decision tree, by which extracted regulations are accurate convenient and simple, is the basic content of data mining. so the soybean import dependence degree is researched in the way of decision tree, so that the serious condition in the soybean import can be found ahead, and scientific bases can be provided for macro regulations and control departments to make decision. 2. the establishing of warning index system the warning of soybean market in china mainly warns the factors which effect the market. the paper combines the soybean demand, soybean price and so on, warns the abnormal condition, and then provides the corresponding precautions. warning is the basis process to ensure the safe of the soybean market, and only the serious conditions and degrees are timely warned, can we ensure the safety of soybean market in china. 2.1 the choice of warning indexes whether the choice of warning indexes is reasonable affects directly the effect and accuracy of warning, so in the process of determining the warning index system of the soybean market, we must consider all kinds of factors which should reflect every side of the soybean market and have representative characters, especially consider the effect of international factors on soybean * this research has been supported by education ministry emphases science and technology project in heilongjiang (project number: 1153lz10) and irtneau. advances in systems science and applications (2010), vol.10, no.1 35 market of china, so we can present the development condition of soybean market more accurately. the paper considers overall the soybean production, demand, export, price, country policy and so on. the corresponding data can be seen in table 1. table 1 the corresponding data of soybean market in china from 1995 to 2006 year import dependent degree(%) soybean production (ten thousand tons) soybean import (ten thousand tons) soybean export (ten thousand tons) price ratio agriculture basis constructing expenditure (a hundred million) soybean production in usa (ten thousand tons) soybean sowing area (k hectare) 1995 0.02 1350.00 29.8 37.6 0.75 110.0 68444 8127 1996 0.08 1322.00 111.4 19.3 0.75 141.5 59174 7471 1997 0.17 1473.00 288.6 18.8 0.71 159.7 64780 8346 1998 0.18 1515.20 320.1 17.2 0.63 460.7 73176 8500 1999 0.24 1425.00 432.0 20.7 0.59 357.0 74598 7962 2000 0.41 1541.00 1041.9 21.5 0.64 414.4 72224 9307 2001 0.48 1540.70 1394.0 26.2 0.62 480.8 75055 9481 2002 0.41 1651.00 1131.5 30.5 0.69 423.8 78672 8720 2003 0.58 1539.32 2074.1 29.5 0.70 527.3 75010 9313 2004 0.54 1740.15 2023.0 34.9 0.65 542.3 66778 9589 2005 0.63 1634.78 2659.1 41.3 0.58 512.6 85013 9590 2006 0.64 1596.70 2827.0 39.5 0.57 504.3 83368 9279 the data come from web of statistics of national bureau of china directly or have been calculated. warning alert index is the ratio of soybean net import to soybean demand amount, named as soybean import dependence degree, which reflects the dependence degree of soybean market in china on foreign soybean market, in which, the soybean demand amount is calculated as follows: soybean demand amount = soybean production + soybean import – soybean export warning sign indexes are soybean production, the ratio of soybean price in usa to that in china which is named as price ratio, soybean demand amount and agriculture basic construction expenditure, in which, price ratio = soybean price in usa × exchange rate of money ÷soybean price in china. 2.2 decide the foregoing years of warning sign indexes the dividing of foregoing, synchronous and lag indexes is very important in the process of soybean warning, which is the key part of the whole warning system of soybean market. usually, the approach of time difference relevant analytics is used. the foregoing warning sign indexes, which are the most important sign indexes of the warning system, have a role of predicting the future of soybean market in china. the paper decides the foregoing years by many ways and comprehensive analysis, the foregoing years can be seen in the table 2. table 2 the foregoing years of warning sign indexes warning sign indexes foregoing years warning sign indexes foregoing years soybean production in china 3 soybean demand amount 1 price ratio 2 agriculture basic construction expenditure 2 soybean production in usa 3 soybean sowing area 3 liu: the research on the early warning system of soybean import dependence degree based on decision tree 36 2.3 dividing alert limits and degrees of warning alert index and warning sign indexes dividing the alert limits and degrees is the key part in the design of warning system. alert degrees reflect the level and intensify of warning contained in the actual values of warning alert index. in order to decide the warning degree of each index, first we ensure the alert limits of each index, and then ensure the prosperity station of each index each year according to the values of alert limits. consulting ge huiling’s paper ‘the forecast and early warning research about soybean market in china’, the station of soybean market import dependant degree and alert limits are described as table 3. table 3 status description of soybean market import dependent degree condition feature described alert limit non-warni ng area soybean market is stable and boom, soybean mainly comes from china, there are little import, doesn’t depend on foreign soybean. [0 0.2] gentle warning area soybean market is stable and boom, soybean mainly comes from china, there are a little of import, doesn’t depend on foreign soybean. [0.2 0.35] medium warning area soybean market is relatively stable, soybean mainly comes from china, there are a little of import, depends on foreign soybean in some degree. [0.35 0.5] serious warning area soybean market is turbulent, soybean coming from china is a very small part, there were mass import soybean, largely depends on foreign soybean. [0.5 1] making use of feedback theory, the warning sign indexes’ limits are decided according to the warning alert index’s limits. to represent in the boundary of two warning areas of warning alert index )(ix , and the one of warning sign index )(iy . so the boundary is between the non-warning and the gentle warning area, or between the gentle warning and medium warning area, or between the medium warning and serious warning area. to represent in the maximal of x and y of xmax and ymax , the minimal of x and y of xmin and ymin , then according to the radial principle, the warning area of x and y are decided by the following formula: yy xx y x iy ix minmax minmax )(max )(max − − = − − by calculating, the warning areas of warning sign indexes can be seen in table 4. table 4 the warning areas of warning sign indexes warning sign index serious warning area medium warning area gentle warning area non-warning area soybean production [1643 +∞) [1537 1643] [1441 1537] (-∞ 1425] soybean demand amount [3677 +∞) [2944 3677] [2211 2944] (-∞ 2211] price ratio [0 0.62] [0.62 0.66] [0.66 0.70] [0.70 +1] agriculture basic construction expenditure [441 +∞) [337 441] [233 337] (-∞ 233] usa soybean production [8021 +∞) [7364 8021] [6706 7364] (-∞ 6706] soybean sowing area [9098 +∞) [8587 9098] [8076 8587] (-∞ 8076] thought out the data and results, we set up the training table of soybean market import dependence degree in table 5. advances in systems science and applications (2010), vol.10, no.1 37 table 5 training table of soybean market import dependence degree year import dependent degree soybean production soybean demand amount price ratio agriculture basis constructing expenditure usa soybean production soybean sowing area 1995 non-warning non-warning non-warning non-warning non-warning gentle warning gentle warning 1996 non-warning non-warning non-warning non-warning non-warning non-warning non-warning 1997 non-warning gentle warning non-warning gentle warning non-warning non-warning gentle warning 1998 non-warning gentle warning non-warning medium warning serious warning gentle warning gentle warning 1999 gentle warning gentle warning non-warning serious warning medium warning medium warning non-warning 2000 medium warning gentle warning gentle warning medium warning medium warning gentle warning serious warning 2001 medium warning gentle warning gentle warning serious warning serious warning medium warning serious warning 2002 medium warning serious warning gentle warning gentle warning medium warning medium warning medium warning 2003 serious warning medium warning medium warning gentle warning serious warning medium warning serious warning 2004 serious warning serious warning serious warning medium warning serious warning non-warning serious warning 2005 serious warning medium warning serious warning serious warning serious warning serious warning serious warning 2006 serious warning medium warning serious warning serious warning serious warning serious warning serious warning 3. decision tree warning 3.1 decision tree arithmetic decision tree that is a tree structure is similar to a flow chart. each node designates an attribute, the branch designates the value of the attribute and the leaf designates a category. the highest level node is root. the way of setting up a tree is recursion from up to down. 3.1.1 the creation of decision tree (1) if all the examples in the trained sample set belong to the same type according to the category attribute, then the set can be called leaf node. the content of the leaf node is the sign of the type. (2) otherwise, according to some tactic (information gain ratio of each attribute), a non-category attribute is chosen whose information gain ratio is biggest. by the value of the attribute, the sample set is divided into several subsets. (3) then, deals with each subset by the way of recursion till all samples in the same subset have the same value of category attribute. 3.1.2 choice of determinant attribute about the recursion way of decision tree, usually the attribute which have the most information gain is the test attribute of present node, so the demand information is smallest for sorting the remained test sample set. that is, dividing the sample set by the attribute will make the mixture degrees of the subsets be slowest. however, this calculating way based on information theory is partial to the attribute which has more values, but the attribute is not always the best. so, information gain is represented by gain ratio.definition of entropy : ( ) ( )∑ = −= c i ii ppsentropy 1 2log s is the training set, c is the number of categories of category attribute and pi is the ratio of the number of samples of category i to that of s. definition of information gain ( ) ( ) ( ) ( ) ( ) ∑ ∈ −= avaluev vv sentropysssentropyasgain /, sv is the subset whose value of attribute a is v and entropy(s) is the entropy of the training set. the information gain of attribute is the reductive expected value after dividing. liu: the research on the early warning system of soybean import dependence degree based on decision tree 38 definition entropy of each non-category attribute. ( ) ( ) ( )∑ = = c i ii ssssassplitinfo 1 2 /log/, is is the subset which has the ith value of attribute a in s. for the more average of the attribute divided by attribute value, the bigger of splitinfo, information gain is represented by information gain ratio to avoid choosing that attribute. definition of information gain ratio ( ) ( ) ( )assplitinfo asgainasgainratio , ,, = gain ratio is information gain ratio,gain is information gain,splitinfo is the entropy of attribute. 3.2 actual application of decision tree putting the training data in table 5 into the decision tree arithmetic, the information gain ratio of non-category attribute is obtained in table 6. table 6 information gain ratio of each attribute soybean production in china soybean demand amount price ratio 0.5068 0.4856 0.6652 agriculture basis constructing expenditure soybean production in usa sowing area 0.5406 0.4192 0.6475 it can be seen from table 6, the information gain ratio of price ratio is biggest, so it is the root node, then deal with other attribute by the way of recursion. we obtain the following tree: figure 1 the decision tree obtained rules: if price ratio is non-warning, soybean import dependence degree is non-warning if price ratio is gentle warning and soybean demand is non-warning, soybean import dependence degree is gentle warning if price ratio is gentle warning and soybean demand is gentle or medium warning, soybean import dependence degree is serious warning if price ratio is medium warning and soybean production is gentle warning, soybean import price ratio soybean demand amount gentle warning serious warning agriculture basis constructing expenditure non-warning serious warning medium warning soybean production medium warning non warning medium warning gentle warning gentle warning non warning medium/ serious warning medium warning serious warning serious warning medium warning medium warning advances in systems science and applications (2010), vol.10, no.1 39 dependence degree is medium warning if price ratio is medium warning and soybean production is medium warning, soybean import dependence degree is serious warning if price ratio is serious warning and agriculture basis constructing expenditure is medium warning, soybean import dependence degree is medium warning if price ratio is serious warning and agriculture basis constructing expenditure is serious warning, soybean import dependence degree is serious warning. making use of this tree, the import dependence degree of the soybean market in china in 2007 is tested, and that of 2008 is warned ahead, corresponding warn degrees of foregoing warn sign indexes are in table 7. table 7 the training table of the import dependence degree of the soybean market in china in 2007 and 2008 year soybean production soybean demand price ratio agriculture basis constructing expenditure 2007 serious warn serious warn serious warn serious warn 2008 medium warn serious warn serious warn serious warn putting the training data in 2007 into decision tree, we obtain the result: the import dependence degree of soybean market in china in 2007 is serious warning. it has been known that the soybean production in china in 2007 is 1440 ten thousand tons, the soybean imports 3082 ten thousand tons and the soybean exports 39.2 ten thousand tons. by calculating, import dependence degree of soybean is 0.69 and in serious warning area. this accords with the result obtained by the way of decision tree. furthermore, the paper warns the import dependence degree of soybean market ahead in 2008, and the result is that it is in the serious warning. the tree comes true the warning by combining the soybean price, soybean production, soybean demand and country policy. usa supplies the most of soybean in the world, and china is one of the biggest import countries. the price ratio has a direct influence on china soybean market. the lower of the ratio, china is more partial to import foreign transgenic soybean in order to satisfy the demand, so resulting in warning affairs. at the same time, along with the rapid development of economy in china, people’s consumptions in albumen and oil are increasing gradually. because of the limit of the sowing area in china, the soybean production is limited. in recent years, the gap between inland soybean production and market demand amount is being enlarged gradually, following which the amount of soybean import is increasing, so resulting in warning affairs. besides, agricultural basis construction expenditure has an obvious effect on soybean market of china, in which storage and construction of traffic have a direct effect on soybean purchase price. if the costs of storage and transportation are reduced, the amount of the soybean import will be less. however, the ability of transportation in some areas of china is very low, which results in the rise of the cost of transportation and indirectly stimulates the rise of the inland soybean price, so the chinese soybean businessmen import the foreign soybean. therefore, the investment in agriculture basis construction expenditure is the direct factor influencing soybean import degree. the investment in agriculture basis construction expenditure of china hasn’t gone up for 4 years, which accords with the fact of serious warning of soybean import degree in recent years. in order to improve the present condition of china, first, the change of foreign and native soybean prices should be paid more attentions, and the inspecting degree of import soybean should be intensified at any moment. then, not only preventing dump of foreign country to soybean market in china, but also taking measures to retort the inferior position of domestic soybean price. at the same time, the policy is necessary and native soybean planting industry should also be stable. at last, the disadvantageous condition of depending foreign soybean should be changed overly by enhancing the completing ability. liu: the research on the early warning system of soybean import dependence degree based on decision tree 40 4. conclusion the basic functions of decision tree are prediction and classifying. both of the two functions are the basis of warning. first, the foregoing property of the warning sign indexes ensures the success of warning, and future import condition is estimated by the values of warning sign indexes, then the import condition is classified into 4 parts. the paper considers overall the factors affecting on the soybean import condition, so the non-category attributes are a little more, but there were only parts of the attributes which provides more information are used. therefore, the advantage of decision tree can be seen, that is the way can auto-filter factors and needn’t combining other ways. because of the limits of the foregoing years of warning sign indexes, the model can only warn the soybean import dependent degree of future in a year. besides, the number of data in this paper is small and taking use of every province’s data in china may be a good thing. the more of the data, the warning result is more accurate. references [1] jun meng. the analysis of soybean potential yield in jiansanjiang farm by logistic model. advances in systems sciences and applications, 2007, 7(1): 143-147. [2] jun meng. the research of agricultural producing industry structure’s evaluating and demonstration analysis. advances in systems sciences and applications, 2007, 7(2): 173-177. [3] fanliang kong, zhaowu ni. a martingale method of perpetual american options pricing with ddefault-risk. advances in systems sciences and applications, 2007, 7(1): 7-11. [4] minfen shen, bin li, qianhua zhan, p.j.beadle. estimation of visual erp signals with nonextensive entropy. advances in systems sciences and applications, 2007, 7(1): 26-31. [5] ming zhu. data mining[m]. he fei: china university of science and technology publishing house, 2002: 67-74. (in chinese) [6] jingmin chen. technology of data warehouse and data mining[m]. publishing house of electronics industry, 2007, 6: 203-206 (in chinese) [7] danyang cao, jinhong li, jinqiang wei, yanfang zhang. cet-4 grade performance analysis based on decision tree. journal of north china university of technology beijing china, 2007, 1: 38-41. (in chinese) [8] xue bai, fu duan. the study and application of decision tree on teaching access. the development and application of computer, 2007, 20(2): 24-26. (in chinese) advances in systems science and application(2015) vol.15 no.1 1-20 the paradoxes of the world’s progress(ii) sailau baizakov1 and jeffrey yi-lin forrest2 1jsc “iei” economic development and trade of kazakhstan, jsc “economic research institute”, str. temirkazyk 65, 0100000, astana, kazakhstan 2department of mathematics, slippery rock university, slippery rock, pa 16057, usa abstract an unbiased systematic view on the modern innovation allows you to see three key groups of innovations that are still very few people differentiate. the first is the technical and technological innovations that are the basis of development and change in technological ways of the world. secondly, it is monetary and financial innovation, the progress of which determines the change in monetary ways of life of the world. and third, it is the socio-political innovation, progress of which is in the basis of the change of socio-political models of the world. a clear distinction between these three “floors” of innovations is crucial for understanding the global crisis and the ways to quit it. since the essence of it lies in the tangle of clearly long overdue and contradictions between the rates of introduction of the world’s technical and technological, financial and socio-political innovations. as the global crisis clearly shows, technical and technological way today, is not decisive for the country’s prosperity and peace. exactly the countries with the highest level of technology development have become the main source and a key cause of the global crisis. keywords global crisis, sona analyzer, technical and technological innovation, monetary and financial innovation, socio-political innovation 3 the range of application of the analyzer sona 3.1 the main directions of analytics of macroeconomic policy area of practical applications for the sona analyzer in macroeconomics is very promising and productive. the basic formula that can be successfully used in the management of market economies are: 1.to estimate the rate of stp and modeling of balanced economic growth c(t) = gdp (t)/(qp (t) +gdp (t)) (1) 2. calculation of purchasing power of money and the price of world currencies pp(t) = (c(t) ∗ i2)/i1 (2) 3. to determine the index of market prices of goods and services 1/pp(t) = i1(c(t) ∗ i2) (3) 2 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) 4. to determine the index of balanced economic growth with ppp i3 = c(t) ∗ i2 (c1) i3 = pp(t) ∗ i1 (c2) c(t) ∗ i2 = pp(t) ∗ i1 (c) (4) 5. to determine the nature of the gdp deflator as the ratio of the nominal indices (i1) and real (i2) of growth p(t) = c(t)/pp(t) (5) 6. calculation of net growth rate of stp and comparative analysis of the competitiveness of the countries dc(t) = c(t)− 100% (6) since purchasing power of the national currency is picked for each country, based on the principles of calculation of sdr, we have: pp(sdr, t) = i=n∑ i=1 ngdp (i)/[pp(1) ∗ngdp (1) + pp(2) ∗ngdp (2) +...+ pp(n) ∗ngdp (n)] (7) where n total number of imf member countries willing to be engaged in foreign trade. 3.2 debates about economy management regulators: questions and answers during the debates of s. baizakov conducted with famous russian professor, doctor of economic sciences, corresponding member of the ras k.valtuh, the following questions and answers that can actually reveal the effectiveness of anticrisis measures in the countries of the world and the capability of market economy management had appeared. question 1. how is stp coefficient defined? answer. stp coefficient is defined by formula: c(t) = µ(t) 1 + µ(t) (8) after conversion, it turns out that c(t) = gdp/x, where x is output, which is the sum of costs of intermediate qp consumption and gdp for production qp +gdp . at any anti-crisis event involving the purchase and sale, a coefficient c(t) is always known value, it can be determined with an accuracy of gdp (t) and x(t). question 2. what determines the rate of scientific and technological progress (stp ) if it does not reflect the performance? advances in systems science and application(2015) vol.15 no.1 3 answer : indeed, the indicator called the coefficient of scientific and technological progress (stp ), and defined by the formula: c(t) = µ(t) 1 + µ(t) (9) where the indicator µ(t) = ngdp (t) qp (t) (10) expresses the organic composition used in the production of gdp of material resources, which are not directly related to labor productivity. comment by valtuh k.k. concept of organic structure is occupied. it does not applied here. this itself is substantively different (see “capital” t i) answer : in this formula, the money in one case serve as money capital (ngdp). in another case acts as a commodity-capital (x). in this formula, commoditycapital is represented by its denominator. according to the theory of emerson, it is the equivalent of the capital expended plus a normal profit [emerson, twelve principles of productivity. c.62]. comment by valtuh k.k. macroeconomic indicator of labor productivity exists and is systematically used (see, for example, the annual reports of the president of the united states). in general, the opposition of macroeconomics and the real economy is unacceptable. as for economics (not to be confused with the widespread literature, claiming without evidence for non-marxist economic theory), it is precisely the subject of the real economy (and for this literature largely fictional). answer: i am here for the organic composition of social labor meant the ratio of money capital (ngdp) to the commodity-capital (x). my statement follows logically from marx’s remarks on the work shtibelinga in “capital”, volume iii, part i, p. 25. question 3. what does the coefficient of µ = gdp/qp expresses, which is included in the formula for determining the rate of scientific and technological progress? answer : output(x) = qp + gdp is the sum of the intermediate material resources (qp ) and gdp . and coefficient µ = gdp/qp expresses the performance of material resources qp . comment by valtuh k.k. output is not the amount of material resources and gdp , but the amount of current material costs and gdp . answer : agree. question 4. scientific and technological progress (stp ) is always associated with labor productivity. why do not you have that connection! answer : if the level of macroeconomics goes to the level of the real economy, contribution directly to the dynamics of stp performance will appear, and then 4 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) its share of the output will be determined. advises by valtuh k.k.: representative calculations of this coefficient should be hold. it turns out that it varies in a narrow range. but the main quickly changing, usually growing the effect of technological progress does not change the value c(t) and the growth of productivity of human labor. this effect is not expressed in the formula. there are also a number of other ntp effects. answer: performance of living labor is reflected at the level of the coefficient of stp c(t). thus, the formula coefficient ntp c(t) = gdp (t)/x(t) can be replaced with an equivalent ratio of the productivity of living labor in gdp to performance of output x. but the growth of productivity of human labor will have a positive effect with domination of productivity growth for gdp over growth in productivity of the production. we are not interested in productivity growth, and its impact on the price of goods and services. as you know, the prices of goods and money change after the drop of productivity of labor and capital. it should be noted here that my opponent propose a coefficient of production to replace with the sources of real, not nominal growth. but, we are referring to the fact that the highest levels of productivity can be extinguished corresponding increase in average wages. the important point for a market economy it is not a productivity growth, the so-called living labor, but the importance of economic dynamics of labor productivity, which is defined as the ratio of labor productivity of living to the average annual salary. as has been justified above, the variation coefficient of stp in this case expresses the integrated expression of the effects of the anti-crisis measures and subjective actions of managers of production from a position of economic growth and stability in the country. table 1 shows the ntp coefficient calculated by formula c(t) = gdp (t)/x(t). as shown in the table 1, the range of the ratio of ntp does not vary within a narrow range, as k. valtuh thought, and its actual range of variation increased from 31% in cyprus and 54% in estonia to 148% in brazil and 162% in the czech republic. these range changes occurred in only 10 years, so we cannot speak about the narrowness of the range of variation of this indicator. do the obtained results allow us to provide answers to the question how to determine the ratio of scientific and technological progress (stp)? indeed, for a long time economists could not estimate a contribution of scientific technological progress, which remained exogenous, into the model of growth. mathematical models of economic growth, which appeared in the 80s, included the externality of economic growth.?only paul romer argued persuasively that the growth of capital by 10%increases the production of nearly 1%, rather than 0.25, according to calculations based on the solow’s model.?the costs of research and experimental development (niekr) provide the growing impact of scientific advances in systems science and application(2015) vol.15 no.1 5 table 1 coefficient of scientific and technological progress (stp), 2000 = 100% country/year 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 1 austria 98 97 97 96 94 91 92 88 90 88 2 belarus 104 114 117 115 117 116 115 115 108 105 3 belgium 99 85 71 65 74 87 87 89 95 94 4 bulgaria 100 100 99 94 93 90 87 89 93 98 5 brazil 98 96 95 94 94 103 111 116 113 148 6 united kingdom 100 101 101 101 98 101 100 103 77 96 7 hungry 101 106 107 108 106 103 105 99 109 106 8 germany 100 102 102 101 99 96 97 92 98 95 9 greece 102 102 103 104 105 106 105 120 105 107 10 denmark 98 99 100 100 98 97 94 100 95 96 11 india 101 100 99 97 96 96 96 88 89 114 12 ireland 76 69 69 68 68 66 65 70 64 64 13 spain 100 99 99 98 97 96 96 96 103 104 14 italy 100 100 100 100 99 98 98 95 103 101 15 kazakhstan 117 107 93 98 93 106 107 166 123 130 16 cyprus 95 87 70 60 55 51 26 31 31 31 17 china 101 101 100 100 100 99 102 100 121 131 18 latvia 100 101 98 94 91 93 96 100 77 65 19 lithuania 98 98 99 97 91 92 96 95 75 74 20 luxemburg 99 104 106 98 93 88 83 83 89 78 21 malta 107 107 104 103 103 96 109 108 106 107 22 netherlands 101 103 104 103 103 102 101 100 100 99 23 poland 100 100 99 96 97 97 99 97 101 110 24 portugal 100 102 102 101 101 100 97 96 91 114 25 russia 97 98 98 99 100 100 99 99 98 98 26 romania 98 98 99 97 100 100 101 94 75 74 27 slovakia 99 100 100 104 109 115 140 161 110 73 28 slovenia 101 101 102 101 104 102 100 100 105 103 29 united states 102 104 103 103 101 101 100 99 107 104 30 finland 100 102 102 101 101 100 87 85 98 97 31 france 100 98 99 98 97 96 97 96 99 98 32 czech republic 99 100 100 98 99 96 98 109 103 162 33 sweden 99 101 102 101 98 96 98 98 102 102 34 estonia 117 123 105 92 70 70 66 67 56 58 35 japan 82 89 101 102 92 88 91 100 102 109 36 turkey 100 86 100 100 100 100 101 101 101 101 37 iran, islamic republic 107 85 97 94 90 85 86 104 102 107 38 afghanistan 99 112 70 70 71 70 71 59 54 59 39 pakistan 101 102 103 104 105 106 107 108 109 111 40 uzbekistan 98 99 100 102 103 104 105 107 108 109 41 turkmenistan 100 100 100 101 101 101 101 101 101 102 42 tadzhikistan 86 86 87 92 94 97 101 102 93 93 43 azerbaijan 106 106 102 100 107 114 117 124 117 140 44 kyrgyzstan 102 102 102 101 101 101 100 85 92 94 45 south africa 98 105 108 105 105 107 111 109 116 127 6 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) and technological progress [1]. reversibility principle has allowed us for the first time in world economics make a significant step forward in measuring the contribution of scientific and technical progress in the development of the real economy through the exact formula c(t) = gdp (t)/x(t). question 5. the purchasing power of money, how is it determined? answer : the purchasing power of money is relative measure. ntp coefficient is directly related to the index of purchasing power of money. the index formula of purchasing power of money, as indicated above, is defined by the ratio of the coefficient of stp to the gdp deflator index of official statistics. table 2 shows the changes in indices of purchasing power of money obtained by calculation according to global statistics. it defined the purchasing power of money is the national currencies of 45 countries for 2000 2010, which have over 80% of the gdp of the world economy. estimated formula: pp(t) = a(t)/p(t), where c(t)-factor of stp, p(t) the gdp deflator. as it can be seen from table 2, the range of changes in the purchasing power of money is wide as the range at the rate of stp. it is changed over the past decade in the range of 0.12 to 1.20 in cyprus in japan. the nature of change in the index of purchasing power of money in this case expresses the effectiveness of monetary and financial management system in the countries.?only one-quarter of these countries has saved the purchasing power of their currencies at a level higher than 50% of their value in the base year 2000.this group includes japan (1.20), the u.s. (0.81), china (0.66) and brazil (0.53).in the same group there were majority of oecd countries.on the contrary, outsiders with weak national currencies are the half of the european union countries, including the ten countries that have a power of their currencies less than one-third of their base power: bulgaria with a coefficient of 0.25, hungary 0.30 ireland 0.22, cyprus?0.12, latvia 0.20, lithuania 0.22, luxembourg 0.30, romania 0.18, slovakia 0.17 estonia 0.27. these figures increase the likelihood of permanent crisis in europe and create difficulties for?these countries to implement the pact for economic growth and stability, adopted in mid-2012. question 6. can you comment on the formula rgdp (t) = ngdp (t)/p(t)? answer. formula ngdp = p ∗rgdp refers to the work of official statistics. statistics are watching over the production and determine the physical volume index (pvi) of goods and displays the gdp deflator (p) by the formula: p = ngdp/rgdp (11) the ratio of the growth rate of ngdp to rgdp growth rate represents the rate of inflation or gdp deflator. emphasize: the gdp deflator is a product of pv i. the more growth of pv i index, the lower the gdp deflator. therefore, the more stable, seems to be evolving economy. and vice versa. the lower the advances in systems science and application(2015) vol.15 no.1 7 table 2 the dynamics of the purchasing power of national currencies of the world for the period 2000-2010 (2000 = 1) country/year 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 1 austria 1,00 0,92 0,74 0,65 0,60 0,58 0,52 0,46 0,49 0,39 2 belarus 0,92 0,89 0,81 0,68 0,58 0,51 0,45 0,38 0,43 0,39 3 belgium 1,01 0,96 0,79 0,68 0,63 0,59 0,53 0,49 0,55 0,44 4 bulgaria 0,99 0,82 0,62 0,49 0,42 0,38 0,31 0,27 0,29 0,25 5 brazil 1,19 1,31 1,20 0,95 0,71 0,61 0,55 0,50 0,50 0,53 6 united kingdom 1,03 0,95 0,83 0,70 0,66 0,64 0,63 0,60 0,53 0,51 7 hungry 0,98 0,82 0,64 0,53 0,47 0,46 0,38 0,34 0,41 0,30 8 germany 1,03 0,98 0,81 0,72 0,70 0,67 0,61 0,54 0,59 0,47 9 greece 1,00 0,89 0,69 0,59 0,55 0,54 0,48 0,48 0,47 0,45 10 denmark 1,01 0,93 0,76 0,67 0,62 0,60 0,53 0,50 0,50 0,44 11 india 0,97 0,92 0,78 0,67 0,56 0,50 0,40 0,34 0,34 0,35 12 ireland 0,73 0,57 0,43 0,36 0,33 0,29 0,25 0,25 0,26 0,22 13 spain 0,99 0,87 0,68 0,57 0,52 0,47 0,40 0,37 0,41 0,35 14 italy 1,01 0,93 0,75 0,65 0,63 0,59 0,52 0,48 0,52 0,42 15 kazakhstan 1,09 0,99 0,75 0,62 0,49 0,43 0,37 0,46 0,40 0,35 16 cyprus 0,95 0,77 0,51 0,36 0,31 0,26 0,11 0,12 0,12 0,12 17 china 0,91 0,90 0,87 0,81 0,75 0,70 0,66 0,57 0,65 0,66 18 latvia 0,95 0,88 0,71 0,57 0,51 0,39 0,31 0,26 0,22 0,20 19 lithuania 0,92 0,80 0,61 0,51 0,45 0,38 0,31 0,26 0,22 0,22 20 luxemburg 0,99 0,95 0,78 0,60 0,54 0,44 0,35 0,32 0,36 0,30 21 malta 1,11 1,04 0,86 0,75 0,72 0,64 0,56 0,50 0,51 0,52 22 netherlands 1,02 0,94 0,76 0,66 0,63 0,59 0,51 0,45 0,48 0,39 23 poland 0,92 0,89 0,79 0,69 0,57 0,49 0,40 0,31 0,41 0,42 24 portugal 1,01 0,94 0,77 0,66 0,63 0,61 0,50 0,45 0,44 0,44 25 russia 0,89 0,82 0,67 0,53 0,47 0,37 0,30 0,30 0,32 0,28 26 romania 1,24 0,98 0,76 0,55 0,42 0,34 0,24 0,19 0,18 0,18 27 slovakia 1,00 0,86 0,62 0,53 0,47 0,42 0,39 0,35 0,25 0,17 28 slovenia 1,07 0,94 0,74 0,62 0,57 0,51 0,42 0,36 0,39 0,35 29 united states 1,00 1,00 0,97 0,94 0,89 0,86 0,83 0,81 0,85 0,81 30 finland 1,00 0,92 0,74 0,65 0,62 0,60 0,45 0,40 0,43 0,35 31 france 1,02 0,91 0,74 0,64 0,62 0,58 0,51 0,47 0,50 0,40 32 czech republic 0,98 0,81 0,65 0,52 0,46 0,39 0,33 0,29 0,30 0,31 33 sweden 1,11 1,02 0,83 0,70 0,67 0,62 0,54 0,50 0,60 0,47 34 estonia 0,95 0,81 0,63 0,50 0,46 0,39 0,32 0,31 0,27 0,27 35 japan 0,95 1,07 1,14 1,10 0,98 0,97 1,04 1,02 1,10 1,20 36 turkey 1,44 1,11 1,04 0,88 0,78 0,76 0,65 0,58 0,72 0,67 37 iran, islamic republic 0,98 0,82 0,86 0,73 0,62 0,54 0,45 0,48 0,49 0,40 38 afghanistan 1,13 0,75 0,49 0,47 0,45 0,42 0,38 0,26 0,25 0,23 39 pakistan 1,06 1,09 1,01 0,93 0,91 0,84 0,79 0,70 0,79 0,74 40 uzbekistan 1,23 1,53 1,54 1,41 1,29 1,17 0,99 0,88 0,82 0,75 41 turkmenistan 0,87 0,81 0,72 0,68 0,64 0,57 0,52 0,86 0,89 0,91 42 tadzhikistan 0,89 0,86 0,75 0,66 0,64 0,59 0,50 0,39 0,38 0,36 43 azerbaijan 1,08 1,09 0,99 0,90 0,80 0,72 0,63 0,50 0,54 0,57 44 kyrgyzstan 0,97 0,91 0,82 0,76 0,68 0,60 0,48 0,38 0,39 0,39 45 south africa 1,18 1,32 0,92 0,76 0,73 0,75 0,75 0,83 0,80 0,64 8 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) growth rates of pv i, the greater the rate of inflation, and inflation the scourge of the market economy. comment by valtuh k.k.: inflation index the concept is much broader than that of the gdp deflator. answer : we do not think so. we put the question differently: when earlier economists following a. marshall asked “what determines the price of the goods or services?” and what the rate of inflation, our initiative group asks “what determines the purchasing power of money and what is the level of balanced growth?” comment by valtuh k.k.: this question was discussed long before marshall. see also, in particular, my results in statistical test value theory. question 7. what is the relationship q of a particular product with a q in the fisher’s equation of exchange? do you want to aggregate the value of a specific product j and associate it with the indicators of macroeconomics? answer: in fact, the relationship q of a specific product j and q in the fisher’s equation does not exist. comment by valtuh k.k.: there is no direct connection. but it means in the “fisher’s equation” that the product of this equation is the sum of the prices of specific products. continued answer : as for the aggregation of indicators specific product j and link it to macroeconomic indicators, such problems in our economic management technology does not exist. it can be explained by the fact that we are willing to work with four indicators that are available in the statistical reports and businesses, and sectors of the economy, and in sna issue, the nominal gdp (gdp ) and real gdp (gdp ) deflator and v v p −ngdp/rgdp . aggregation of indicators is the work of statisticians, and the gdp deflator is defined by them alsongdp/rgdp . scientist has to work with analysis, and not the aggregation of indicators. conclusion by valtuh k.k.: i replied on your thoughts. further discussion between us is not necessary:in fact, the cause of our differences lies in the fact that we work in different ways, and not just have different understandings of a particular issue. conclusion by s. baizakov : for the conclusive answers to a system of all questions, i have reinterpreted presented the findings of the study. if to be precisely i introduced the concept of “balanced growth rate”(i3), which allowed me to remove many unclear answers to the questions in that discussion: i3 = pp (t) ∗ngdp (t) = c (t) ∗rgdp (t) (12) table 3 shows the comparative data rates of balanced growth in the european union and neighboring countries and regional organizations in developed countries. advances in systems science and application(2015) vol.15 no.1 9 table 3 the calculation of the rate of balanced economic growth in 2010, 2000 = 100 a coefficient of scientificreal growth balance nominal purchasing countries, technological rate growth rate growth rate power parity regions progress (stp) (i2 index) (i3 index) (i1 index) (ppp) 4=(2)*(3) 1 2 3 =(5)*(6) 5 6 u27 98,50 100,8 99,34 239,08 0,42 u15 97,13 100,1 97,21 233,39 0,42 u9 96,15 100,1 96,28 226,66 0,42 eec 99,83 165,0 164,74 566,25 0,29 bric 126,22 183,1 231,16 437,77 0,35 35 countries 105,46 121,8 128,43 200,97 0,64 eco 108,02 179,3 193,63 331,54 0,58 united states 103,70 114,5 118,74 146,22 0,81 japan 109,35 129,1 141,20 117,40 1,20 united kingdom 95,85 104,5 100,12 197,78 0,51 germany 94,51 109,3 103,28 219,53 0,47 france 97,55 101,2 98,75 245,13 0,40 as can be seen from the calculation formula of balanced growth rate and results of calculations based on this formula (table 3), the changes in coefficient in real gdp leads to changes of coefficient in nominal gdp . the source of any changes in rates of economic growth is scientific technological progress, in particular changing of its coefficient c(t). the overall conclusion of our discussions with valtuh k.k. valtuh k.k. believes that “in fact, the cause of our differences lies in the fact that we work in different ways, and not just have different understandings of this or that particular issue”. if so, then the k. valtuh’s desire to stop further discussion of the current issues in the world economy, i consider premature. my challenge is based on a k. valtuh’s study on value theory [2]. the results of this study are shown in section 4. 4 the mathematical substantiation of necessity of determining the rate of balanced growth 4.1 econometric models are meaningful if they relate to the economic laws of the market the modern theory of value is based, in addition to the works of marx and marshall on the major achievements of the twentieth century in the development of economic science. in this direction of economic theory a lot of work was done by novosibirsk school of economists led by corr. ras k. valtuh. k. valtuh said 10 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) the current prejudice against labor theory of value extremely negative affects the state of economic science, including formulated of practical recommendations on behalf [2,3]. more specifically, it is recommended to give a general theory of the cost a place in economic science, which it in its content deserves. theory of value responds well known concept of the theory in general meets the criteria of external justification and internal perfection. first of all, it is justified by the classical law of value in its findings related to pricing, which in measuring the outcome of the economy has not yet taken its rightful place. as we know, in modern economic science the difference between cash and trade flows is generally measured by the gdp deflator. according to these scientific and theoretical considerations, the official statistics determines it by the ratio (index) of the amount of the final internal current prices of products manufactured of a certain the country to amount of the same output as measured in the prices of some base year. subsequently, the gdp deflator is identically defined as function (the product) of the three variables identified in the labor theory of value: (1) reciprocal relationship of payroll to gdp , measured in current prices, (2) average payment of unit of labor and (3) direct labor intensity of gdp , measured in base year prices. accordingly, the dynamics of the gdp deflator is identically determined by the dynamics of these three variables (for example, the growth rate of the deflator for a certain year the growth rate of these quantities for the same year). however, as a result of this expansion, we have only the identity [2]. so to express deflator in motion, as a model of the gdp deflator valtuh has taken advantage of econometric analysis methods. but, in our opinion, he is trapped in a simple modeling technique and missed the point of research of the economic content of the gdp deflator, before engaging its modeling. valtuh just wrote “in modern economics the difference between cash and trade flows is generally measured by the gdp deflator” [ibid.]. further, he was in the wrong chain of reasoning, taking at face value the econometric model of the gdp deflator. one of the examples of failed economic modeling of inflation and gdp deflator in kazakhstan we have given above. these econometric models do not take into account the contributions aspirations, successes and risks of the entrepreneurs to the economy. “the first most important of the innate properties of matter karl marx wrote a movement, not only as a mechanical and mathematical movement, but even more as a desire, a life spirit, stress, or, to use the expression of jacob boehme, flour of [gual] matter” [4]. as you know, the driving force, the spirit of the economy is every desire, which is called the scientific and technological improvement of production, in short, scientific and technological progress (stp) in the economy. the level of scientific and technological progress in each moment advances in systems science and application(2015) vol.15 no.1 11 of time locked by the one or more indicators. and the result of this desire can also be described as the contribution of the ntp activities undertaken. in our system of models the result of stp is its ratio c(t), and the level of scientific and technological progress gdp (t)/x(t) : c(t) = gdp (t)/x(t). and the purchasing power of money, which sets the cost of money, therefore, the value of the currency is defined by the formula: pp(t) = c(t)/p(t). here, the gdp deflator is given by official statistics on all types of goods and services, industries and economic activities. its clearance from the influence of scientific and technological progress is precisely defined by the formula (t) = (t)/(t). these indicators of stp contribution and the purchasing power of money successfully complement the gdp deflator (inflation index) for analytical work and predictive calculations of economic development. the advantage of these indicators is to analyze the three indices of economic growth which reveal the essence of the true value and balanced economic growth rate (fig.2). usa 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 ngdp 100 103 107 112 119 127 134 141 144 143 146 rgdp 100 101 103 106 110 112 115 118 118 115 118 sgdp 100,0 108,3 101,4 119,2 128,3 137,4 146,6 157,1 168,1 176,0 193,6 japan ngdp 100 88 84 91 99 98 94 94 105 108 117 rgdp 100 100 99 100 105 102 101 105 105 115 126 sgdp 100,0 83,4 89,7 104,1 108,7 96,0 91,1 97,5 106,7 119,3 141,2 kazakhstan 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 ngdp 100 121 135 169 236 312 443 573 729 630 809 rgdp 100 114 125 136 149 164 181 197 204 206 221 sgdp 100,0 132,5 133,3 126,9 146,2 152,6 192,2 210,2 338,1 252,7 286,9 bric ngdp 100 103 110 127 153 185 229 291 342 360 438 rgdp 100 101 106 113 119 126 137 150 162 169 183 sgdp 100,0 102,3 109,9 118,1 128,6 138,5 150,1 161,8 169,9 156,9 164,7 12 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) united kingdom 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 ngdp 100 121 135 169 236 312 443 573 729 630 809 rgdp 100 114 125 136 149 164 181 197 204 206 221 sgdp 100,0 132,5 133,3 126,9 146,2 152,6 192,2 210,2 338,1 252,7 286,9 germany ngdp 100 103 110 127 153 185 229 291 342 360 438 rgdp 100 101 106 113 119 126 137 150 162 169 183 sgdp 100,0 102,3 109,9 118,1 128,6 138,5 150,1 161,8 169,9 156,9 164,7 fig.2 example of illustrations of the possible outcomes of the analysis of three levels of balanced economic growth on the fig.2 to illustrate the possible outcomes of the analysis there are the three levels of balanced economic growth. in the first row is the state of the economies of the u.s. and japan in 2000 and 2010. in these countries the rate of balanced growth is higher than the growth rate of nominal (second place) and real growth. on the second row there are the states of the bric and kazakhstan economies for the same years. their balanced growth rate was lower than the nominal, but higher than the actual growth rate. the third row shows the states of the economies of united kingdom and germany. as the chart above shows the development of the economies of these countries in the same years, both, the rate of balanced growth was lower not only advances in systems science and application(2015) vol.15 no.1 13 than the nominal, but the real rate of development. these analytical calculations confirm the findings of the president of kazakhstan nursultan nazarbayev made on the vi astana economic forum that definitely the developed countries caused the recession in some countries of the world[5]. in the practice of predictive calculations of economic development of countries stp contribution is still accepted, mainly exogenously. in the simplest case, non-decreasing function of time is introduced to the production function as a multiplier. these calculations in soviet times allowed us to determine the stp type (capital-intensive or labor-intensive) to take into account the structure of ntp, which is defined by technology used in the production, grouped according to their degree of progressivity [6,7]. but since the end of the twentieth century intensive work was started to integrate the stp function into models of economic analysis. thus, in the collective monograph sops at the academy of sciences of the russian federation, published under the imprint of the academic series “the problems of the soviet economy” in 1988 there were stress that the key “direction of account of stp progress is to use so-called stp functions (for example, the first equation of the model kaldor-mirrlees, the function proposed by baizakov s. in the institute of economic research). here, dependence of the relative indicators of the stp development is accepted as a function. according to kaldor-mirrlees, stp function is defined by dependence of growth of labor productivity on growth of capitallabor. in a model of s. baizakov it is determined by the dependence of growth in capital productivity per unit of labor embodied in the organic composition of production (capital-degree lead over costs to pay for it). in contrast to the first type of production functions, where ntp is autonomous, in these models, the latter a driving force of macroeconomic dynamics. similar analysis of unit of production functions as a method of accounting for ntp in the prediction was given in the works of s. baizakov, v. dadayan, n. kulbovskoy, etc.[8,9]”. indeed, as our russian colleagues point out, the implementation of stp functions in macroeconomic models would allow defining more precisely the most important synthetic indicators of the economy development (gross domestic product growth rate of labor productivity, capital productivity, etc.). if using a macroeconomic model the contribution of stp can be estimated, it would open the possibility for the treatment of price indices of goods and services from the influence of the gdp deflator. however, the analysis carried out by the institute for economic research in recent years has shown that the problem is not so simple. even if the amount of money corresponds to the number of produced commodity supply and circulation process is going well, the choice of indicators of measuring the final result pro14 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) duction in terms of money is not a simple technical problem. in fact measuring the final product by the value of nominal gdp , real gdp and the gdp deflator is only a necessary condition, but the system of indicators is not sufficient for the detection of defective mechanisms bottlenecks in economic management. indeed, the gdp deflator is one of the most important indicators of economic management. because it is determined by the ratio of nominal gdp this year to real gdp , the gdp deflator itself needs meaningful economic expertise. our conclusion is that before engaging in mechanical treatment the gdp deflator by econometric methods, its economic content should be thoroughly understood. econometrics will give meaningful results when economic models, based on it, meet the aspirations of entrepreneurs. it’s not a secret that the most widely used econometric models and production functions do not fully reflect the contribution of scientific and technological progress (stp ) in the economy. 4.2 limitations of monetary theory of economic growth there is given a new interpretation of monetarist model, based on which the limitations of the gdp deflator (p(t)) and its inadequacy are mathematically proved as a key instrument of financial management: qtxpt = vtxmt (13) here, qt is the product of the base year, which equals the nominal gdp , pt is inflation index, defined by the velocity of circulation of money of the base period vt = v0. for example, if according to the national economy one of the countries had rates qt = $14 366 700 000, mt= $12 460 100 000, and vt =v0 = q0 /m0 = 9 817 000/7 173 800=1,36845, then pt = vt x mt /qt =1,36845 x 12 460 100 / 14 366 700 = 1,187 or 118,7%. as can be seen from these calculations, they do not take into account change in counter of technological excellence of production c = µ/(1 + µ), where µ = gdp/qp , qp is a product of intermediate consumption in the system of national accounts or the cost of current material expenses in the amount of salesc = qp +gdp . increasing dynamics of indicator µ(t) by m. porter is an increase in productivity of current material costs with the cost of qp , used in the production of gdp , and their sum is constant at any given point in time -x = qp+gdp = const. from the standpoint of environmental protection and green economy indicator µ(t) is the most important criterion for competitiveness, as well as the gdp/qp at x = const means a maximum of gdp production at current low cost of energy and use of raw materials and other natural resources. in turn, the monetarists formula qtxpt= vtxmt at vt = v0 is represented as ytxp2,t =v0xmt, where yt = qtxpt nominal gdp , and p2,t = v0/vt money inflation index. now let’s consider the ratio of the data of values of the current advances in systems science and application(2015) vol.15 no.1 15 year t to cost in the base period of time: qt × pt q0 × p0 = vt ×mt v0 ×m0 (14) or yt y0 = vt ×mt v0 ×m0 (15) in another way, it appears as a product of growth of indicators q and m, as well as the growth rate of the control indicators p and v: qt q0 × pt p0 = vt v0 × mt m0 (16) or yt y0 = vt v0 × mt m0 (17) so, the product of the growth rate of real and gdp deflator is equal to the product of the growth rate of money turnover (= index value of one unit of money in relation to nominal gdp ) and growth rate of the money supply. or, in other words, we can say: the growth rate of nominal gdp is equal to growth rate of the product of the circulation of money multiplied by the money growth rate. hence the price index is equal to one unit of money: vt v0 = qt/q0 mt/m0 × pt p0 (18) or vt v0 = yt/y0 mt/m0 (19) that is, the growth index of the price of one unit of money equals to the ratio of real gdp growth and money supply multiplied by the gdp deflator. or, growth index of value of one unit of money equals to the ratio of growth rates of nominal gdp and the money supply. if the prices of goods and services for final use did not change, the rate of growth of the price of one unit of money would be determined only by the ratio of real gdp growth in the money supply. since the price index for goods and services in real-life is greater than one, then multiplying this ratio to the gdp deflator leads to an increase in index of the price of money. the real index of the “price” of money expressed by physical benchmark is the ratio of the index of price of money received by the gdp deflator and in other words is the ratio of the rate of real gdp to growth rate of the money supply: vt v0 / pt p0 = qt/q0 mt/m0 (20) 16 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) equation (4) shows how the physical contents of one unit of money changes when the real volumes of end-products and the money supply change. the reciprocal of the growth rate of money turnover represents an index of monetary inflation (relative to nominal gdp ): v0 vt = mt/m0 qt/q0 × p0 pt (21) or v0 vt = mt/m0 yt/y0 (22) if the prices of goods and services have not changed, then the monetary inflation would be determined by the ratio of the rate of growth of the money supply to growth rate of real gdp . and because the prices of goods and services change, the monetary inflation is determined by the ratio of the rate of growth of the money supply and nominal gdp . let us call the ratio of the rate of growth of the money supply to growth rate of real gdp in the base year price as index of real monetary inflation. it will be equal to the product of monetary inflation index for inflation in prices of goods and services for final use: mt/m0 qt/q0 = v0 vt × pt p0 (23) hence it is not difficult to get a formula to determine the nature of the gdp deflator, familiar to us from the official statistics: pt p0 = mt/m0 qt/q0 . vt v0 = yt/y0 qt/q0 = i1/i2 (24) wherei1 = yt/y0 is the rate of economic growth in the current year prices, andi2 = qt/q0 is the rate of economic growth in base year prices or the same as index of physical volume. as can be seen from (a1), the gdp deflator expresses the ratio of the rate of economic growth in the prices of the current year to the rate of growth in the prices of the base year. it does not take into account possible changes in the cost of the currency. as a result, the index of the gdp deflator is the ratio of a fixed state of the economy in two different time points, which does not react to changes in a rapidly changing reality, above all, to changes in the purchasing power of the national currency. 4.3 limitations of the duality theory of economic growth clarifying the roles of the gdp deflator is possible through the use of the principle of duality of kantorovich-koopmans, which is written as: qtxpt = ctxxt (25) advances in systems science and application(2015) vol.15 no.1 17 here, as noted above, indicator c represents the efficiency of the trade or expresses the contribution of scientific and technological progress in the real economy. at ct = c0 it is represented asytxp1,t = c0xxt, where yt = qtxpt is nominal gdp , and p1,t = c0/ct is the index of counter of technological excellence of production. now consider the formulas analogous to 1-6 with a contribution of scientific and technological progress in the real economy. first, we consider the ratio of the data values of the current year t to cost in the base period of time: qt × pt q0 × p0 = ct ×xt c0 ×x0 (26) or yt y0 = ct ×xt c0 ×x0 (27) in another way, it appears as a product of growth rates of indicators q and x, as well as the growth rate of the control indicators p and c: qt q0 × pt p0 = ct c0 × xt x0 (28) or yt y0 = ct c0 × xt x0 (29) that is the product of the growth rate of real gdp to the gdp deflator is equal to the product of the rate of counter of technological excellence of production of output of goods and services (including products for intermediate consumption). the term “technological improvements” for the first time in the economic cycle is entered, apparently, by alfred marshall. or, in other words, we can say: the rate of growth of nominal gdp is equal to the product of rate of counter of technological excellence of production to the rate of production of goods and services (including products for intermediate consumption). hence, index of counter of technological excellence of production (stp coefficient) equals: ct c0 = qt/q0 xt/x0 × pt p0 (30) or ct c0 = yt/y0 xt/x0 (31) that is, index of counter of technological excellence of production (stp coefficient) equals to the ratio of the rate of growth of real gdp and gross output multiplied by the gdp deflator. 18 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) or, index of counter of technological excellence of production (stp coefficient) equals to the ratio of nominal gdp growth rate and gross output. if the prices of goods and services for final use did not change, then stp coefficient would be determined only by the ratio of real gdp growth to gross output. since the price index for goods and services in real-life is greater than one, then multiplying this ratio to the gdp deflator leads to an increase in the rate of counter of technological excellence of production. real growth rate of stp coefficient expressed by physical benchmark is the ratio of the resulting index of growth rate of stp coefficient for the gdp deflator and in other way is the ratio of real gdp rate to gross output: ct c0 / pt p0 = qt/q0 xt/x0 (32) equation (10) shows how the physical contents of one unit of output (one unit of money needed to trade x) changes when the real volume of the final product and gross output changes. the reciprocal of the stp coefficient represents the index of excess of the gross output relative to nominal gdp : c0 ct = xt/x0 qt/q0 × p0 pt (33) or c0 ct = xt/x0 yt/y0 (34) if the prices of goods and services have not changed, then the index of exceeding the gross output of the nominal gdp would be determined by the ratio of the rate of growth of gross output to growth rate of real gdp. and because the prices of goods and services change, the index of exceeding the gross output relative to the nominal gdp determined by the ratio of growth rate of gross output and nominal gdp. let us call the ratio of the rate of growth of gross output to growth rate of gdp as real index of exceeding the gross output relating to the nominal gdp. it will be equal to the product of the index of exceeding the gross output and the nominal gdp deflator: xt/x0 qt/q0 = c0 ct × pt p0 (35) hence it is not difficult to obtain a new specification of the formula of the gdp deflator, which reflects the changes in the real economy, and the changes in the monetary and financial system: pt p0 = ct/c0 ppt/pp0 (36) advances in systems science and application(2015) vol.15 no.1 19 wherept/pp0 = qt/q0 xt/x0 means the rate of change of purchasing power of money, andpt and ct have the same values. as can be seen from (a2), the gdp deflator saves its money status and its quantitative value determined by the model of monetarism. but its economic nature has now become meaningful and richer. according to its new formula, the nature of the gdp deflator revealed the unity of the two mutually independent indicators of economic management. one of them represents the ratio of scientific and technological risk of entrepreneurs’ work in the real sector, and the other expresses the purchasing power of money in the monetary and financial system. one of them expresses the performance of the real sector; the other expresses the quality of the financial sector of the economy. thus, the “crossing” of keynesian and monetarists’ theory based on the duality theory of kantorovich-koopmans has been successful. 4.4 balanced economic growth as the key to the management of the market economy due to the formula (a2), the rate of balanced economic growth (growth index i3(t)) is defined as the product of the purchasing power of money and the growth rate of nominal gdp (growth index i1(t)): i3(t) = pp(t) ∗ i1(t) (37) on the other hand, the same rate of balanced economic growth (growth index i3(t)) is determined by the product of the ratio of stp coefficient and the real growth rate of gdp (growth index i2(t)): i3(t) = c(t) ∗ i2(t) (38) the equality of the given rate of nominal gdp growth pp(t) ∗ i1(t) with the given rate of real gdp growth c(t) ∗ i2(t) means that any point on the path of balanced economic growth t = t0 is an equilibrium point in prices of goods and services to the purchasing power of the national currency: pp(t) ∗ i1(t) =c(t) ∗ i2(t) (39) thus, the identification of “explosive feature” of the gdp deflator has allowed the initiative group of kazakhstan to present it by two conjugated indicators of scientific and technological progress of the real sector of the economy and the purchasing power of the currency. indicator of scientific and technological progress is a barometer of management of the real sector of the economy. managing the dynamics of change is a function of private sector entrepreneurs, and government agencies that monitor the development of the natural monopolies. indicator of the purchasing power of money is the barometer of management 20 sailau baizakov and jeffrey yi-lin forrest:the paradoxes of the world’s progress(ii) of the financial sector of the economy. managing the dynamics of change is a function of the representatives of monetary and financial system, and of the national bank, which oversees the activities of banks. the initiative group of economists of kazakhstan believes that the speech of the president of kazakhstan nursultan nazarbayev at the vi astana economic forum has given a new impulse to the study of sustainable development issues and served as a key to unlocking the essence of the gdp deflator. thus it opened the way to the justification of the rate of balanced growth as the criterion of economic governance of the world economy. references [1] m. ratner and d. ratner. (2004), “nanotechnology: a simple explanation of the next brilliant idea, translated from english”, williams, moscow, pp.240. [2] v. l. makarov. (1985), “on the performance of scientific and technical progress”, economics and math. methods, vol xxi, 2nd ed. [3] (2005), “theory of value: statistical verification. informational generalization. actual conclusions” ,herald of the ras, no.9. [4] k. k. valtuh. (1965), “public utility of product and labor costs of its production”, m:thought. [5] n. nazarbayev. (2013), speech on th vi astana economic forum. [6] k. marx and f. engels, op. t. 2. pp.142. [7] v.s. dadayan. (1973), “modeling of national economic processes”, economics, moscow, pp.433-438. [8] v. l. makarov and a. p. torzhevsky. (1986), “effect of changes in the technological level of production on macro indicators of economic development”, cemi as ussr, moscow, pp.4-7. [9] b.m. shtulberg, e.g. chistyakov and v.v kotilko etc. (1988), “problems and methods of study of the territorial plans”, science, moscow, pp.192. corresponding author sailau baizakov can be contacted at: baizakov37@mail.ru advances in systems science and applications (2012) vol.12 no.2 141-152 rotor, bearing and dynamic equations in energy storage flywheels for vehicles jinguang zhang1 and yefa hu1,2 1school of mechanical and electronic engineering,wuhan university of technology, wuhan, 430070,china 2hubei digital manufacturing key laboratory wuhan,university of technology, wuhan, 430070,china abstract energy storage systems for vehicles present significant challenges for rotor and bearing design. this paper discusses rotor and bearing design technology in energy storage flywheels for vehicles, with particular emphasis on orientation of flywheel rotors, rotor geometry and magnetic bearings. material, rotational speed and geometry are mainly factors of flywheel rotor design. in order to achieve an attractive specific energy, the rotor speed should be as high as possible. the bearings must be capable of extremely high speed, have very low friction, and be stiff to adequately constrain the rotor, have long life, and have high load capacity. these requirements frequently lead to choosing magnetic bearings for flywheel energy storage system applications. a flywheel energy storage system prototype with active magnetic bearings was designed. the flywheel was suspended by the permanent magnetic bearings and stabilized by the active magnetic bearing. finally, we deduce differential equations of the magnetic suspended flywheel for the design of control system. keywords flywheel, energy storage, active magnetic bearing, vehicle, dynamic equations 1 introduction traditionally, the energy storage requirement for vehicles has been satisfied by chemical batteries. however, batteries have a number of disadvantages such as limited cycle life, maintenance, conditioning requirements, and modest power densities which have been improved upon by newer technologies such as energy storage flywheels. flywheels in particular offer very high reliability and cycle life without degradation, reduced ambient temperature concerns, and construction free of environmentally harmful materials. the energy storage flywheel system mainly consists of flywheel rotor, motor/generator, magnetic bearings, housing and power transformation electronic system. one of the major advantages of flywheels is the ability to handle high power levels. this is a desirable quality in e.g. a vehicle, where a large peak power is necessary during acceleration and, if electrical breaks are used, a large amount of power is generated for a short while when breaking, which implies a more efficient use of energy, resulting in lower fuel consumption. individual 142 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles flywheels are capable of storing up to 500mj and peak power ranges from kilowatts to gigawatts,with the higher powers aimed at pulsed power applications. the flywheel energy storage packed in vehicle can operate at a nearly constant, optimum speed, reducing fuel consumption, air and noise pollution, and engine maintenance requirements, and extending engine life. short bursts of power, for climbing hills and acceleration, are taken from flywheel energy storage, which is replenished directly by the engine or by regenerative braking when the vehicle is slowed down. unlike friction brakes, which turn kinetic energy into waste heat, regenerative braking changes it to speed up the flywheel for subsequent acceleration[1-4]. the university of texas, center for electromechanics, developed and tested a high speed composite flywheel for an advanced technology transit bus. this flywheel operated at 40,000 r/min and could deliver 844 wh at a power rating of 150 kw. road testing of the bus revealed acceleration time to 75km/h was reduced by a factor of two with a simultaneous reduction in engine power of 25%,the design of a high performance flywheel energy storage system for use in a vehicle poses many challenges. in order to achieve an attractive specific energy (kwh/kg), it is necessary to construct the rotor from materials with high specific strength (ultimate stress/density), leading to selection of composite materials employing graphite fibers over metals. this allows a higher specific energy, and increases rotor tip speed. the high tip speed, in turn, leads to an enormous increase in parasitic windage loss. to reduce windage losses to acceptable levels requires spinning the rotor in a very tight vacuum[5-6]. this paper describes issues associated with rotor and bearing design technology in energy storage flywheels for vehicles, with particular emphasis on orientation of flywheel rotors, rotor geometries , magnetic bearings and differential equations of the magnetic suspended flywheel. 2 orientation of flywheel rotors the interaction between vehicle and flywheel dynamics produces many sources of bearing loads not present in a stationary application. the primary contributors to bearing loads are shown to be vehicle shock, vibration, maneuvering, and gyrodynamics[6]. among these loads, gyroscopic loads occur when the rotor is precessed as the vehicle angular velocity changes, for instance cornering, driving over a hill or through a dip. there is no vehicle axis in which angular velocity is avoided so the bearings must withstand the gyroscopic loads. a spinning flywheel has a relatively large angular momentum so changing its spin axis requires significant torque, which must be produced by the bearings. the required torque is: m = jθ̂ × ω (1) advances in systems science and applications (2012) vol.12 no.2 143 where j is the polar moment of inertia of flywheel rotor, is the spin speed of flywheel rotor, and θ̂ is the turning rate of the flywheel spin axis. on the basis fig.1 coordinate system of the coordinate system defined in fig.1, (1) can be expressed as follow: mx = j(θ̇yωz − θ̇zωy) my = j(θ̇zωx − θ̇xωz) mz = j(θ̇xωy − θ̇yωx) (2) in vehicle operating conditions,yaw rate ωz is greater than roll rate ωx and pitch rate ωy .if the flywheel is oriented vertically, flywheel rotor and the yawing (turning) axis of the vehicle in the same direction, and θ̇x = θ̇y = 0 ,so mx = −jθ̇zωy my = jθ̇zωx mz = 0 (3) this means that the flywheel spin axis should be vertical. in the case of a vehicle, the gyroscopic torque is too small to influence the motion of the vehicle. a way to reduce the impact is to employ two similar flywheels, each contra-rotating at the same speed. however, the torque is large enough to be a major contributor to loads on the radial bearings. in practice some mechanism or device must support the flywheel, isolating the flywheel from the motions of the bus and decrease the loads on the bearings, and allowing it to pitch and roll as freely as possible relative to the vehicle. the usual method to rigidly mounting the flywheel housing to the vehicle is to mount it in a two-axis gimbal. without the use of a gimbal, bearing loads induced by gyroscopic reactions to pitch and roll movements of the bus can significantly reduce bearing life and make the design impractical for long term use. 144 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles 3 rotor design in rotor design, there are mainly three fully-coupled design factors that have significant effect in the overall performance of flywheels, material strength, rotational speed and rotor geometry. the kinetic energy stored in a flywheel is proportional to the mass and to the square of its angular velocity. it is given as ek = 1 2 iω2 (4) where i is the mass moment of inertia and ω is the angular velocity. the moment of inertia for any object is a function of its shape and mass. it is obtained by the mass and geometry of the flywheel and given as, i = ∫ x2dmx (5) where x is the distance from rotational axis to the differential mass dmx . the way to increase the energy density and minimise the volume of the system is to use a flywheel with both a high rotational speed and a large inertia. the speed limit is set by the tensile strength of the flywheel material. the stored energy density with respect to mass is given by: em = kσ/ρ (6) where em is kinetic energy per unit mass, k is the shape-factor which relates the relative energy stored in the solid disk to that of a constant stress disk of infinite radius, σ is maximum stress in the flywheel and ρ is mass density. in case of planar stress, if the height of the disk is small compared with the diameter, and a homogenous isotropic material with poisson ratio of 0.3, i.e. steel, is used, the k factors are given in table 1[7] . in a three-dimensional flywheel there will be three-dimensional interaction of material stresses. for the flywheel design is based on a hollow cylinder and the outside radius is assumed to be large compared to the flywheel thickness, the two stresses of primary concern are the radial stress and the tangential stress. table 2 presents characteristics for common rotor materials. there are two basic classes of flywheels based on the material in the rotor. the first class uses a rotor made up of an advanced composite material such as carbon-fiber or graphite. these materials have very high strength to weight ratios, which give flywheels the potential of having high specific energy. the second class of flywheel uses steel as the main structural material in the rotor. the highest tensile flywheels are not made of steel, but of fiber-reinforced composites. as well as rotating faster and storing more energy than steel flywheels, advances in systems science and applications (2012) vol.12 no.2 145 these composite flywheels are much safer if the maximum safe speed is exceeded, since they tend to delaminate and disintegrate gradually from the outer circumference rather than explode catastrophically. the material at the outside diameter of the rotor is most effective in storing energy with energy storage of that material being proportional to square of the radius. the peripheral speed of the rotor should be as high as possible for maximum energy storage but this is limited by the stress levels the designer is willing to accept. for a chosen tip speed and rotor outer diameter, the rotational speed of the rotor is fixed and this also fixes the shape. there are three flywheel geometries were developed to meet the energy storage and power requirement needs for vehicles, shown in fig.2. table 1 shape-factor k for different planar stress geometries table 2 2 data for different rotor materials 146 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles fig.2 flywheel geometry disc flywheel is based on more traditional flywheel designs. the mass of the steel hub provides the main source of inertial energy storage in the system. the profile of the steel hub is based on modified equations for a constant stress disc which include a rim section for an increased radius of gyration. as long as the rotor speed is within acceptable limits and there are no rotor dynamics issues, a long tubular flywheel rotor is the preferred option for a number of reasons as opposed to a disc shaped flywheel. the flywheel designs were denoted by the name “powerbeams” for practical commercial application[8-12]. an arbor incorporates an inside out topography for the motor-generator. the inside-out topography makes better use of the available space and increases the specific energy and power densities of the design. the arbor flywheel shows significant advantages with respect to system size and energy storage capability[5]. 4 bearing design , prototype and dynamic equations 4.1 bearing design the spinning flywheel rotor must be supported on bearings. initially, both mechanical bearings and magnetic bearings were considered. the important parameters in assessing the use of bearings are weight, loss, cost, lifecycle life, and low losses. they also can isolate rotor and stiffness. if the rotor speed is within acceptable limit, mechanical bearings are ideal in that they can operate with low losses and have high life for the average load yet can accept high loads on an intermittent basis several times the average load. due to the high friction and short life, mechanical bearings cannot be adapted to modern high-speed flywheels. mechanical bearings have benefited greatly from material advances such as ceramics and very hard steels. the main life issues are not material fatigue life, but rather lubricant life. lubricant life depends primarily on temperature. instead magnetic bearing system is utilized, including permanent bearings, active magnetic bearings (electromagnetic bearings) and high temperature superconducting (hts) bearings. magnetic bearings do not have any contact with the shaft, has no moving parts, experience little wear and require no lubrication. active magnetic bearings present major advantages in terms of lifetime and advances in systems science and applications (2012) vol.12 no.2 147 rotational speed, and also favorably integrate into high-speed flywheel systems. unlike active magnetic bearings, the hts magnetic bearing can situate the flywheel automatically without need of electricity or positioning control system. however, hts magnets require cryogenic cooling by liquid nitrogen. it is not suitable for vehicles. due to the higher magnetic flux density reached by nd-fe-b magnets and their low cost, applications with permanent magnetic bearings have become attractive, in spite of being very unstable. therefore, permanent magnetic bearings can be used as an auxiliary bearing to reduce the load weight of the rotor and the flywheel and to increase the stiffness of the whole bearing system. 4.2 prototype of flywheel a flywheel energy storage system prototype with an arbor flywheel and a hybrid bearing set is shown in fig.3. fig.3 flywheel geometry according to the above discussion, high speed is desirable since the energy stored is proportional to the square of the speed but only linearly proportional to the mass. magnetic bearings can accommodate very high spin speeds and have theoretically unlimited imbalance induced vibrations. magnetic bearings offer very low friction enabling low internal losses during long-term storage. so, our design of a flywheel system using the magnetic bearing consists of a vertical arbor flywheel rotor, permanent magnetic bearings, and an active magnetic bearing. the flywheel axial stability is actively controlled by the active magnetic bearing 148 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles while the motions in other directions are restricted by other two pairs of active magnetic bearings. a motor/generator is located in the center region of the arbor flywheel rotor. 4.3 dynamic equations of flywheel active magnetic bearing usually use differential excitation. it uses a pair of symmetrical power amplifier circuit to drive electromagnet in differential model and get a pair of magnetic force in opposite direction. magnetic force of rotor is the difference value between upper magnets and lower magnets. assuming disturbance and control current are very small, in accordance with the taylor series expansion the magnetic force in static working point can be expressed as: fi = kyyi + kiii (7) here, ky = µ0a0n 2i0 2 y30 , ki = µ0a0n 2i0 y20 (8) where fi = total magnetic force, its direction and the y positive direction are consistent ky = displacement rigidity coefficient of magnetic bearing ki = current rigidity coefficient yi = current rigidity coefficient ii = displacement according to the balance position of rotor, its positive direction is upward i0 = offset current y0 = air-gap in balance position µ0 = air magnetic permeability a0 = area of electromagnet pole n = number of coil winding turns equation (7) is the linear model of resultant force in small deviation range.with the increase of distance of balance point, the precision of (7) is decrease. in some limit state, such as rotor contact with stator, strong current (iron-core saturation) or weak current in winding, (7) is incongruity. we consider the base as stationary firstly. inertial reference frames and flywheel reference frames were set up. fig.4 depicts the coordinate system and forced diagram. fi = magnetic force of radial direction fz = magnetic force of radial direction the centroid of flywheel is (xc, yc, zc) .the rotation angles of rotor in yz and xz plane are θx and θy . advances in systems science and applications (2012) vol.12 no.2 149 differential equations of the magnetic suspended flywheel are: mẍc = f1 + f3 mÿc = f2 + f4 mz̈c = fz −mg jθ̈x = −jzωθ̇y + h1f4 − h2f2 jθ̈y = jzωθ̇x + h2f1 − h1f3 t0 = t (8) t0 is motor torque. the above differential equations of the 5-dof magnetic suspended flywheel can decouple into a single dof differential equation and a 4-dof differential equation. the two equations can be separated into axial and radial direction and solved respectively. in this way, control system in axial magnetic bearing can be considered as a single dof magnetic suspension control system. control system in radial magnetic bearing can be considered as a multidof magnetic suspension control system. because θx and θy are small, the displacements may be defined as follows: y1 = xc + h2θy y2 = yc − h2θx y3 = xc − h1θy y4 = yc + h1θx (9) we have q(t) = [xc, yc, θx, θy] ′ i(t) = [i1, i2, i3, i4] ′ carrying (7) and (9) into (8) , we obtain:{ q̈ = p1q̇ + p2q + p3i y = p4q (10) in which p1 =  0 0 0 0 0 0 0 0 0 0 0 −jz j 0 0 jz j 0  p2 =  2ky m 0 0 kyh2−kyh1 m 0 2ky m kyh1−kyh2 m 0 0 kyh1−kyh2 j kyh1 2+kyh2 2 j 0 kyh2−kyh1 j 0 0 kyh1 2+kyh2 2 j  150 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles p3 =  ki m 0 ki m 0 0 ki m 0 ki m 0 −kih2 j 0 kih1 j kih2 j 0 −kih1 j 0  p4 =  1 0 0 h2 0 1 −h2 0 1 0 0 −h1 0 1 h1 0  we can set z = [ q q̇ ] equation (10) then becames ż = [ q̇ q̈ ] = [ 0 i p2 p1 ] z + [ 0 p3 ] i y = [ p4 0 ] z + [0] i (11) if flywheels is set up on vehiclesthe dynamic equations are given by the expression m(ẍc + ẍcar) = f1 + f3 m(ÿc + ÿcar) = f2 + f4 m(z̈c + z̈car) = fz −mg j(θ̈x + θ̈xcar) = −jzωθ̇y + h1f4 − h2f2 j(θ̈y + θ̈ycar) = jzωθ̇x + h2f1 − h1f3 t0 = t (12) where ẍcar, ÿcarandz̈car are accelerations of different directions of vehicle. θ̈xcarandθ̈ycar are angular acceleration of x and y directions of vehicle. (m+mcar)ẍcar = fxcar mcar ≫ m ẍcar = fxcar mcar and equation (11) then results in mẍc = f1 + f3 − m mcar fxcar mÿc = f2 + f4 − m mcar fycar mz̈c = fz −mg − m mcar fzcar jθ̈x = −jzωθ̇y + h1f4 − h2f2 − j jcar mxcar jθ̈y = jzωθ̇x + h2f1 − h1f3 − j jcar mycar t0 = t (13) advances in systems science and applications (2012) vol.12 no.2 151 here, mcar = mass of vehicle. fxcar, fycar, fzcar = applied forces of of different directions of vehicle, fzcar is connected with road surface profile spectral excitations. the differential equations of the magnetic suspended flywheel can be used for the design of control system. 5 conclusion in this paper, we presented a prototype of flywheel energy storage system for a vehicle. the prototype is viable and can achieve perfect robustness. more improvements in material, magnetic bearings and power electronics make flywheels a competitive choice for vehicle energy storage applications. the use of composite materials, optimized rotor shape and magnetic bearings enable high rotational velocity with power density greater than that of chemical batteries. finally, we deduce differential equations of the magnetic suspended flywheel for the design of control system. acknowledgements this work is funded by the national natural science foundation of china under grant no. 50675163. references [1] paul p acamnley, barrie c mecrow, james s burdess, et al. (1996), “design principles for a flywheel energy store for road vehicles”, in ieee transactions on industry applications, vol.32, no.6, pp.1402-1408. [2] hebner. r, beno. j, and walls. a. (2002), “flywheel batteries come around again”, spectrum, ieee, vol.39, no.4, pp.46-51. [3] c.c.chan. (2002), “the state of the art of electric and hybrid vehicles”, proceedings of the ieee, vol.90, no.2, pp.247-275. [4] j. vanmierlo, p. van den bossche, g.maggetto. (2004), “models of energy sources for ev and hev: fuel cells, batteries, ultracapacitors, flywheels and enginegenerators”, journal of power sources, vol.128, no.1, pp.76-89. [5] c.s hearn, m.m.flynn, m.c.lewis, et al. (2007), “low cost flywheel energy storage for a fuel cell powered transit bus”, vehicle power and propulsion conference, vppc 2007. ieee 9-12, pp.829-836. [6] b.t. murphy, d. a. bresie, and j. h. beno. (1997), “bearing loads in a vehicular flywheel battery”, sae special publications, 1243, feb, 1997, electric 152 jinguang zhang:rotor,bearing and dynamic equations in energy storage flywheels for vehicles and hybrid vehicle design studies, proceedings of the 1997 international congress and exposition, detroit, michigan, usa. [7] björn bolund, hans bernhoff, mats leijon. (2007), “flywheel energy and power storage systems”, renewable and sustainable energy reviews, vol.11, no.2, pp.235-258. [8] pullen k.r, ellis c.w.h. (2006),“ kinetic energy storage for vehicles, hybrid vehicle conference”, iet the institution of engineering and technology, pp.91-108. [9] yefa hu, zude zhou, zhengfeng jiang. (2006), basic theory and application of active magnetic bearing, china machine press. [10] jiqiang wang, feng-xiang wang, ming zong. (2007), “critical speed calculation of magnetic bearing-rotor system for a high speed machine”, proceedings of the csee, vol.27, no.27, pp.94-98. [11] amati n, brusa e, gianolio g. (2000), “effects on the dynamic behaviour of rotors on abm of the electromechanical coupling in induction motors. proceedings of the seventh international symposium on magnetic bearings”, eth zurich, switzerland, zurich: international center of magnetic bearing,eth-zurich, pp.595 -600. [12] kent davey. (2004), “new electromagnetic lift control method for magnetic levitation systems and magnetic bearings”, ieee transactions on magnetics, vol.40, no.3, pp.1617-1624. microsoft word 19 hong xu--group link gear locus modeling based on motion relativity.doc 348-354 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc group link gear locus modeling based on motion relativity hong xu mechanical engineering department, chengdu electromechanical college, chengdu, 610031, china abstract in order to solve the comprehensive analysis problem of the mechanism link locus and motion, this paper bring forward the parameter calculation of locus based on motion relativity and differential geometry theory. the problem of point-group calculation is thus solved, and the comprehensive analysis of motion and position is realized. moreover, the method is simplified and the calculation accuracy is also improved. in the paper, a moving frame of axes is set up by using the lever axes to describe the complex structure of linkage. motion relativity theory is adopted to study the overall motion rule of the structure. by using the calculation result of the linkage axes, the calculation model is set up to analyze its motion, which simplifies the point-group motion calculation. the point-groups form the same component have the same angle velocity, angle acceleration velocity, base point, and calculation model, which makes the calculation procedure modularized and the program shortened. owning to complete digital analysis, the characteristic parameter calculations of locus are very accurate, the data-base is easy to establish, the transfer and analysis of data become convenient. this also makes it easy to establish the shape chart of the linkage locus and to analyze their resemblance. it’s a useful method to analyze the locus and motion of the link gear comprehensively. keywords link gear motion motion analysing group locus multi-dimensional space 1.introduction the basic task of computer-aided analysis of link gear is to calculate the motion parameter. but it is difficulty because of the diversification of complex trajectory and motion parameter on every point of the link gear. especially the comprehensive analysis on motion and trajectory has been always important and difficult. so it’s very necessary to find a easy theoretical method to calculate and improve efficiency. francois isnard worked over the motion analysis of link gear, and has found out two arithmetic methods. john r m utilized graphics to analyze the characteristics of trajectory and motion with high speed. zahn c t, roskies r z put forward a method to fit linkage curve by fourier series. luokang proposed a easy method to process moving axes. by using two dimensional array to describe and store the coordinates of linkage curve, it has many merits in data processing and calculation, but too complicated. these methods ignored the physical characteristic linkage motion. even the study on the internal rule of motion didn’t utilize the integrated characteristic of the space configuration and the motion to simplify calculation. at present, there is no fundamental method to solve the unified problem on motion and geometry, also no method to establish and analyze linkage motion parameter accurately and swiftly. 2.our approach according to motion relativity and differential geometry theory, this paper proposes a motion trajectory analysis method based on moving axes of link gear. a frame of axes is fixed on link lever which produces a moving frame of axes by using lever axes as coordinate. it can be very convenient to express the complex structure of linkage, and also can take full advantage of the linkage axes calculation to simplify the motion calculation on point groups, thus realize the comprehensive study on linkage motion and geometry trajectory. it’s an efficient method to the project is funded by the sichuan educational commission in the area of nature science (contract no. 2004a163). advances in systems science and applications (2011), vol.11, no.3-4 349 analyze the motion characteristic of link gear... 3.new method for design of link mechanism 3.1.1 the basis problem for trajectory design generally, motion design on plane link gear can be arranged as three basis problems. first, component position, that is, the link gear can lead a certain component pass some preset positions accurately or approximately according to prescribed sequence. second, preset motion rule, that is the driving part and the follower meet some required relationship of positions, including definite quick-return characteristic. or, when the motion of driving part is defined, the follower can move accurately or approximately according to prescribed rule. third, prescribed motion trajectory, some point on the linkage component in plane motion can move accurately or approximately according to prescribed trajectory. the problem can be summarized as: present digital design needs to analyze the motion of some point or a group points on the component, which requires record the data of its trajectory characteristic on different point. the specific task of trajectory motion analysis is to analyze and calculate several adjacent points, define the most accordant articulated points which meet the preset trajectory motion requirements. on the other hand, the motion parameters on nonadjacent points are calculated, thus to establish the database on different trajectory motion characteristic, and prepare for the design. the description of the integrated motion rule of link gear it’s simple and flexible to utilize lever group method to analyze link gear, moreover, easy programming. paper [1] has worked out a general subprogram for second grade group of five basic type, and also several three and four grade lever group in common use. the kinematics analysis system of plane link gear based on fundamental lever group and single lever automatic identification is a general system for plane linkage motional performance analysis, which can analyze the motion of single and multifreedom system in a wide range. as to its mathematic model and solving process, it’s too complicated. normally the results are angle velocity and angle acceleration velocity. this paper works on the mathematic modeling of motional link gear. to simplify the process, the model is a four lever linkage. the links on the articulated points are the object to study the overall motion rule of linkage. fig. 1 shows a plane four lever link gear in xoy right-angle coordinate system. supposing the length of the framework is r1 and the angle θ1 (θ1=0); the length of the crank is r2 and the angleθ2; the length of the lever is r3 and the angleθ3; the length of the pendulum lever is r4 and the angle θ4. by using vector loop method to establish vector equation of linkage: →→→→ +=+ 4132 rrrr (1) supposing the angle between every lever and x axes is positive anticlockwise, and negative clockwise. the projections of every vector on coordinate system: b c d fig. 1 plane four lever mechanical linkage 350 xu: group link gear locus modeling based on motion relativity 0coscoscoscos 44113322 =−−+ θθθθ rrrr (2) 0sinsinsinsin 44113322 =−−+ θθθθ rrrr (3) 01 =θ tωθ =2 33 ωθ =′ 44 ωθ =′ 33 εθ =′′ 44 εθ =′′ 22441 2244 3 coscos sinsinarctan θθ θθ θ rrr rr −+ − = (4) cb cbaa − −+± = 222 4 arctan2θ (5) 2sinθ−=a 221 cosθ−= rrb 42142 2 3 2 4 2 2 2 )cos(2 1 rrrr rrrr c θ− −++ = 2 433 422 3 )sin( )sin( θ θθ θθ θ −− −− = r r 2 344 322 4 )sin( )sin( θ θθ θθθ ′ −− −− =′ r r )sin( )cos()cos( 343 2 2134 2 23422 3 θθ θθθθθθθ −− ′−−′+−− =′′ r rrcr )sin( )cos()cos( 341 34 2 44 2 3332 2 22 4 θθ θθθθθθθθ − −′−′+−′ =′′ r rrcr for displacement solution of linkage, the key problems are to solve angle, angle velocity and angle acceleration velocity, on the base of which to solve those of one point. when crank angle changes between0~360°,the displacement of that point changes accordingly. if the length of every lever and crank angleθ2 are preset, θ3 and θ4 can be produced by the equation (2) (3). because θ2 is the function of t, so are θ3 and θ4,supposing: )(13 tf=θ )(24 tf=θ differentiate θ3 and θ4,we can get the angle velocity and angle acceleration velocity ω3,ε3,ω4,ε4 from lever and pendulum lever. because the differentiation process is very complicated, so if the result cannot be in full use, it will be a big waste. we use these parameter in point group calculation. 4.motion analysis of moving axes on linkage 4.1.1 trajectory and displacement solution of any point m on link gear when link gear with diversified shape moves in a plane, the description will be very complicated in fixed coordinates. coordinates show in fig. 2 are established on lever, so relative position of any point on lever can be expressed by one group of local coordinate, which won’t x r fig. 2 vector rm point m on plane four lever linkage b c d m r m r r y r x m 1 m 2 m 3 m 4 m 6 m 7 m 5 advances in systems science and applications (2011), vol.11, no.3-4 351 change according to the moving of linkage. no matter how complex of linkage, the size of linkage can be expressed fully and normatively. the absolute position of any point can be expressed by relative coordinate and the origin of local coordinate. according to fig. 2, position vector equation of any point m is: rmm rrr += 2 (6) supposing the angle is positive anticlockwise, the coordinate of point m is: 322 coscos θθ rmm xrx += (7) 322 sinsin θθ rmm yry += (8) among which: xrm,yrm are the coordinates of point m in local axes, xm,ym are the absolute coordinates of point m. when preset a crank angle θ2, the displacements in both x and y direction can be produced on point m with angleθ2 by putting angle θ3,θ4 from equation(4), (5)into them. if preset:θ2=0~360°, step 5°, and make 72 step loop computing, then we can get the trajectory point of point m in this range. these coordinates are more precise than those data produced by graphic chart. differentiate equation(7), (8), the velocity and acceleration velocity on point m can be produced. 4.1.2 trajectory solution of point group on mechanical linkage by trajectory graphic chart, the size of lever and trajectory points can be defined approximately. component size can be determined by using the range of trajectory size, and point position can be chosen by trajectory shape. even lever size has been determined, but calculation and analysis are still necessary for determination whether point trajectory and motion parameter meet requirements. generally speaking, tiny difference on point position can produce huge one on trajectory and motion characteristic. trajectories form different points on lever are shown in fig. 3. the shape of trajectory on different point of lever differs. we can suppose that in xrbyr coordinate system, there be at least one point can meet trajectory requirements. the ideal point may be near the one picked by graphic chart, so we need to choose one group of points nearby for analysis and calculation. those points are all on one component, and share the same angle velocity and angle acceleration velocity ω3,ε3. as before, we take origin b as basing-point, the moving point for calculation. the motion of the moving point can be separated to translational moving of basing-point and rotary moving of moving point. as to moving points group, which have identical angle velocity, angle acceleration velocity and basing-point, the calculation load are no more than those of one point, but the results multiply. l n fig. 3 trajectory from different point on lever xr yr x y 352 xu: group link gear locus modeling based on motion relativity 5.data processing of lever’s moting 5.1.1 the establishing of trajectory data group of point group on link gear the calculation of adjacent point group solved the trajectory character problem of comparison and selection. calculation on different point can solve the moving character problem of analysis and storage for many trajectories. nonadjacent points on lever also have the same angle velocity, angle acceleration velocity and basing-point, the calculation load of point group are no more than those of one point, but the analysis can be very accurate for many kinds of trajectory motion rule. as to the plane link gear shown in fig.1, we can set the number of the lever as 1, other moving levers are 2, 3, … randomly. similarly, all the moving point for calculation can be set as 1,2,3,…,n2 randomly. form a three-dimensional data group grid1(n,q,l), which has n+l line, q+1 column and 2 page. suppose the coordinates of point m on link gear to be (xnm,ynm ) , grid(n,m,0)= xnm and grid(n,m,1)= ynm, according to the simple rule above, we can put all the point coordinates into data group grid. by accessing data group grid, we can get length between p and q on any component l: length=[(xlp-xlq)2+(ylp-ylq)2]1/2 among which, the local coordinates of point p on lever l are xlp=grid (l,p,0),ylq =grid (l,p,1), those of point q are xlq=grid (l,q,0),ylq=grid (l,q,1). 5.1.2 trajectory motion analysis based on local coordinate system by accessing data group grid, we can also calculate motion parameter of point group. the absolute coordinate of any point l equals to its local coordinate plus the absolute coordinate of point b. the calculation shows in (7), (8). as shown in fig.3, the motion of point m is an overlap of two motions (translational moving of basing-point b and rotary moving of point m around b). the velocities of b and m are vb and vm, the rotary velocity of point m relative to b is vmb: ( ) ( )jyixjyixrv bbbbb 22222 ωωωω +=+=×= ( ) jyixjyixrv mbmbmbmbrmmb 3333 ωωωω +=+=×= jyyixxvvv mbbmbbbmbm )()( 3131 ωωωω +++=+= the acceleration velocity of point m: 33 2 2 2 2 εωω τ ×+×+×= ++= rmrm mbmb n bm rrr aaaa the equations above are combination of coordinates, angle velocity and angle acceleration velocity. after the size of linkage and the angle velocity of crank have been given, we can calculate the position, velocity and acceleration velocity of moving point. equally, we can define two-dimensional data group, record position, velocity and acceleration velocity curve on trajectory. 5.1.3 trajectory programming of point group trajectory programming is computer-aided computing. because of the unification and standardization of equation, the programs for one motion parameter are the same model. given crank’s position (angle) and lever’s length, you can calculate trajectory (position) , velocity and acceleration velocity curve by accessing program model. this method can analyze point motion on component precisely. on the other hand, it can process many points together, and store the calculation results as a database, so as to analyze the motion character on different point with digital method. with fourier series proposed by zahn c t, roskies r z to fit lever’s curve, we can fit discrete point into mathematics equation. or, we can pick out the shape character of graphs, and with computer-aided function to analyze trajectory character of plane linkage curve and its distribution rule. advances in systems science and applications (2011), vol.11, no.3-4 353 the calculation process is the same for the same type of plane linkage, just to input different structure parameter and original condition, and you can get the trajectory and displacement. for other types of plane linkage such as crank block, we can deduce its overall motion equation with local axes, that is, the motion rule of moving axes, then the calculation equation of position, velocity and acceleration velocity on the component, then program according to process. the structure is different, but the method is the same. 6.analysis example for linkage motion for four-lever linkage n, imaginary line shown in fig.2 (rectangle with round angle), when move in a plane, we use local axes xrbyr, then the coordinate of any point m can be expressed as (xrmi,yrmi). θ3, ωn ,εn are angle movement of local axes xrbyr. all the points’ local coordinates are in data group grid. basing-point is b, solve position and motion parameters on point m1,m2,……,mn at any moment t: rmi,vmix,vmiy,amix,amiy。 solution: position vector equation at any moment 322 coscos 11 θθ rmnm xrx += 322 sinsin 11 θθ rmnm yry += …… 322 coscos θθ ii rmnm xrx += 322 sinsin θθ ii rmnm yry += velocity vector equation at any moment t: jyyixxvvv bmbbmbbbmm )()( 3131 1111 ωωωω +++=+= …… jyyixxvvv bmbbmbbbmm iiii )()( 3131 ωωωω +++=+= we can see from above, the equation for any point mi is the same as for point m1, so is the program. similarly, the calculation of acceleration velocity also has common calculation equation and program. the basic process to establish linkage curve database is: 1)set linkage basic size r1,r2,r3,r4, if only proportion factor matters for curve character calculation, then suppose r2=1; 2)set position parameter rrmi of point m; 3)use plane linkage kinematics based on basic lever group and single automatic identification to analyze system, and calculate the overall motion rule of component; 4)calculate motion trajectory parameter of point m; 5)solve character parameter t for trajectory graph; 6)store basic parameter p and character parameter t according to certain data group grid1(n,q,l) format. repeat the step above, we can get character parameter database of linkage curve for different linkage and point. 7.conclusion the trajectory parameter calculation method is based on motion relativity and differential geometry theory, which calculates points group motion parameter. it takes full advantage of the simple math relationship between local coordinates and the basing-point and overall motion rule, which makes the equation’s deduction more simple, flexible and systematic, the programming and accessing of subprogram more handy, and accessing of data and analysis of graph more convenient. local coordinate makes it more simple to construct storage model for linkage size information. it makes the best of the middle results for calculation, which reduced repetition in point group trajectory parameter calculation. it’s a very effective resolution method. acknowledgements the project is funded by the sichuan educational commission in the area of nature science (contract no. 2004a163). 354 xu: group link gear locus modeling based on motion relativity references [1] zahn c t,roskies r z. fourier descriptors for plane closed curves. ieee transaction on pattern analysis and machine intelligence. c-21(3) (1992):269~281. [2] francois isnard , gordon dodds. dynamic positioning ofclosed chain robot mechanisms in virtual reality environments. international conference on intelligent robots and systems , victoria , b. c. , canada ,1998. [3] luokang,xiongnan the structure identification of link gear based on connected matrix aggregtion operation, sichuan industry technology college scientific journal, 19 (2) (2000) 35∶ ~37. [4] wang shuang. functional analysis and optimization theory [m]. beijing: beijing aerospace university press, 2004. [5] isnard f , dodds g ,claude vallée. efficient multi arm closed chain dynamics computation for visualisation . international conference on intelligent robots and sys2 tems , grenoble , france ,1997. microsoft word 12 wang huajun, li yamin, zhang yang,huang jing,tang xuan--forming load and metal flow of rotary forging proc 294-300 advances in systems science and applications (2011), vol.11, no.3-4 forming load and metal flow of rotary forging process for spiral bevel gear wang huajun1, li yamin1, zhang yang1,huang jing1,2 and tang xuan1 1.school of material science and engineering, wuhan university of technology, wuhan, 430070, china 2the first aviation academy of chinese air force,xinyang,464000, china abstract rotary forging can overcome the shortcoming of cutting machining and reduces the forming load of the precision forging process. in this paper, the rotary forging process of spiral bevel gear was studied using finite element method, deformation and l-t (load–time) curve at every stage was analyzed. thereby tooth filling phase during rotary forging of spiral bevel gear was obtained. at the same time, the formation of the inside, mid-point and outside at the addendum and dedendum of the convex and concave surface on the blank was also analyzed. as a result, the reason that the outside is difficult to fill was explained. via theoretical calculation and the lead specimen test, the rotary forging process of spiral bevel gear was verified which demonstrated the validity of the rotary forging process of spiral bevel gear. keywords spiral bevel gear, rotary forging, punch, finite element, experiment 1.introduction spiral bevel gear is a core component of the differential,which has the following characteristics: high transmission efficiency, high carrying capacity, driving smoothly and space-saving. nowadays, its manufacturing mainly relies on special gear compound cutting machine[1] due to the complexity shape of spiral bevel gear. comparing with machining, plastic forming can reduce material waste, improve gear strength and service life. however, precision forging of large-diameter spiral bevel gear is very difficult. rotary forging, which is successive partial plastic forming, can decrease the forming load of spiral bevel gear[2] [3]. finite element simulation can be used to verify the validity of the initial design and confirm the new ideas. in recent years, some scholars make use of finite element (fe) methods to study on the plastic forming process of gears. soo-young kim studied on cold forging and heat treatment process for manufacturing of precision helical gear[3]. in order to reduce the forming load, sung-yuen jung studied on the two-step extrusion of helical gear [5]. the rotary forging process of straight bevel gear was studied by cheng peiyuan [6]. j.j.sheu studied on the rotary forging process of spur gear and found the ways to improve die life [7][8]. han xinghui studied the deformation characteristics of cold rotary forging of the ring workpiece[9]. finite element simulation has become an important tool to develop and improve gear plastic forming. because the tooth curve and metal flow of spiral bevel gear are complex,the die often fails from fracture. die failure and difficulty in filling become the key technology of rotary forging process of spiral bevel gear[10][11]. in order to investigate some problems about metal flow, die stress and load change during rotary forging of spiral bevel gear, fe model of rotary forging process is established and rotary forging course of spiral bevel gear is simulated and analyzed in this paper. 2.fe model of rotary forging process for spiral bevel gear taking an automobile differential device spiral bevel gear as the example, the rotary forging process of spiral bevel gear was studied. forging shape of gear components is shown in fig. 1 and the blank shape shown in fig. 2 .the blank material was aisi-1045. advances in systems science and applications (2011), vol.11, no.3-4 295 fig.1.the rotary forging model of spiral bevel gear fig.2.the blank of spiral bevel gear fig. 3 shows the finite element model of the rotary forging process for spiral bevel gear. the blank was fixed on the centre of die and the deformation on inner wall was constraint, so the gear was formed by the punch. the constant friction model was adopted for the contact between punch and blank as well as die and blank, and the friction coefficient was 0.12. fig.3.the fe model of spiral bevel gear for rotary forging 3.analysis of rotary forging load 3.1load-time curve the load-time curve during rotary forging process of spiral bevel gear was shown in fig.4. the figurowed that spiral bevel gear forming process is divided into three phases: free upsetting stage, tooth filling stage, tooth corner filling stage. 1) free upsetting stage: from the beginning up to 1.5sec; because of the existence of the gap between the blank and die, the deformation at this stage is equivalent to free upsetting process. the forming load increases quickly and the maximum load was 1×106n. fig.4.the load –time curve of rotary forging 2) tooth filling stage: from 1.5sec to 4.5sec. some metal close to the die tooth moves into the tooth impression. other metal flows back axially or horizontally due to the obstruction of tooth impression. deformation at this stage is equivalent to forward extrusion with small 296 wang: forming load and metal flow of rotary forging process for spiral bevel gear cone angle; the tooth profile forms gradually. the load increase slowly and the maximum load is 3×106n. 3) tooth corner filling stage: from 4.5sec to 5.6sec; metal overcome friction resistance and move into the tooth corner. the load surge and the maximum load is 7.5×106n. this stage is a decisive stage during forming, for it determines the tonnage of forging equipment, prefabricated blank and precision die shape and so on. 3.2the deformation process of rotary forging fig.5 shows the deformation of blank during the rotary forging process, which is obtained from the simulation. as shown, the blank deformation begins from the outside section near the middle, followed by a steady teeth body filling with a uniform flash at the outer and inner edge as shown in fig.5c. finally, while metal flows into the tooth corner, the flash becomes thinner and larger because the gap between the punch and the die reduces. (a) 0sec (b) 1.5sec (c) 4.5sec (d) 5.6sec fig.5.the deformation process of rotary forging 3.3the distribution of effective strain the distribution of effective strain on the blank at the final stage is shown in fig.11. it can be seen that greatest effective concentrates on the inner and outer lateral flashes. the width of the outer flash is 9.5mm and the inner flash is 8.5mm by measure. effective strain can reflect the forming load approximately. thus, large lateral flash consumes large forming road at the final stage of rotary forging process, which causes the load increase to 7.5×106n sharply. so reduce the flash area at the final forming stage is necessary. fig.11. effective strain of work piece 4.the metal flow 4.1 the overall flow of metal in the forming process, the metal flows along the circumferential direction shown in fig.6. the circumferential size becomes large and the axial size becomes small. at the area which rolling part contacted with punch , both sides metal was squeezed like a wedge then flow circumferentially and gathered to the opposite of the wedge, as shown in fig.6b. one side of the wedge, metal flows towards the die; the other side metal flows towards the punch, as shown in fig.6c. metal being rolled flows complexly and the tooth shape was changed, which advances in systems science and applications (2011), vol.11, no.3-4 297 lead to a deviation between the direction of metal flow and feeding, as shown in fig.6d. (a) cross-section (b) top view (c) side view (d) wedge section fig.6. total velocity 4.2 metal deformation at the addendum (a) the selected points on the formed part (b) the tracking points on the initial blank (c) the tangential strain (d) the effective strain fig.7. deformation of the selected points at the addendum the convex and the concave shapes of a spiral bevel gear are dissimilar. in order to study the deformation of the convex and concave surface at the addendum, some points are selected: p1, p2, p3, the inside, mid-point, outside points at the convex surface, respectively; corresponding to them , p4, p5, p6, the inside, mid-point and outside at the concave surface, as shown in fig.7a. after tracking their position and plastic deformation, the selected points at the addendum of the convex and concave can be found on the cone surface of the blank, as shown in fig.7b, and the points at the outside of the concave and convex produce the largest tangential displacement relatively than others, followed by the mid points, and the inner points are the smallest shown in fig.7c. effective strain of the selected points is shown in fig.7d which consists with tangential strain trend. the outside portion is difficult to fill except increasing the load. 298 wang: forming load and metal flow of rotary forging process for spiral bevel gear 4.3 metal deformation at the dedendum the same method was used to analyze the deformation at the dedendum, some points were selected: p1, p2, p3, the inside, mid-point, outside points at the convex surface, respectively, corresponding to them, p4, p5, p6, the inside, mid-point, outside points at the concave surface, as shown in fig. 8a. fig.8b shows the points on the initial blank, and it can be found that they are all under the cone surface of the blank. as fig.8c shows, the points at the outside of the concave and convex surfaces produce the largest effective strain, relatively than others, followed by the mid points, and the inside points are the smallest. (a)the selected points on the formed part (b)the tracking points on the initial blank (c)the effective strain fig.8.the deformation of the selected points at the dedendum 5.theoretical calculation and forming test 5.1 theoretical calculation the width of the outer flash is 9.5mm by measure and the inner flash is 8.5mm. so the outer radius of the blank is 95mm, and the inner radius is 44mm. radius values are inputted into the formula of contact surface area ratio of circular pieces during rotary forging process [2], and obtained (formula 1). 1527.0 ) 95 4431.001.1() 3952 5.114.0 3952 5.14.0( )31.001.1)(14.04.0( 1 1 = ×−× °×× ×+ °×× ×= −+= tgtg r rqqλ (1) inputting λ into the calculating load formulation [2], it was obtained: 2 2 6 0.1527 3.14 (95 44 ) 1.9 903 5.83 10 sp f k n λ σ= = × × − × × = × (2) whereλ is the contact surface ratio, f is the contact surface, q is the relative feeding, k is the coefficient of upsetting. load-time curve of rotary forging for spiral bevel gear (fig.4) shows that the stable forming force was about 6×106 n, the theoretical value is 5.83×106n according to formula (2). the forming force based on the fe simulation is close to the theoretical calculation. this fact shows that the fe simulation about the rotary forging process for spiral bevel gear is successful. 5.2 forming test in order to verify the validity of simulation, a test was done using lead as the test material. the blank shape is shown in fig.2, and its volume designed by equal volume method. the forming process during rotary forging was shown in fig.9. it can be found that the tooth filling process and the flash deformation obtained by simulation show good agreement with that advances in systems science and applications (2011), vol.11, no.3-4 299 obtained by the test. tooth is full-profile, which demonstrates that the rotary forging process of a spiral bevel gear is feasible. fig.9.the experiment of tooth filling process 6.conclusion spiral bevel gears have different convex and concave shape, long and spiral teeth, so the rotary forging process of spiral bevel gear is a complicated engineering. in this paper, the rotary forging process of spiral bevel gear was studied using deform-3d and test methods. through the study on the rotary forging process of spiral bevel gear, the deformation and load at every stage and the tooth filling characteristics are obtained. the main results of specific studies are as follows: (1) during the rotary forging process of spiral bevel gear, the course of metal-filled cavity is divided into three stages and at the final stage the flash is large and the load is excessive. (2) the largest strain at the outside points is the reason why it is difficult to fill totally at the outer portion. (3) via theoretical calculation and the lead specimen test, the feasibility and validity of the rotary forging process for spiral bevel gear was verified. acknowledgements the research is supported by the scientific foundation for hubei in china (2007aba214) and the fundamental research funds for the central universities (2010-zy-cl-046). references [1] zeng tao. spiral bevel gear design and processing harbin: harbin institute of technology press,1989 [2] pei xinghua, zhang meng , hu yamin. rotary forging.beijing: machinery industry press, 1991. [3] soo-young kim, satoshi kubota and masahito yamanaka. ‘application of cae in cold forging and heat treatment processes for manufacturing of precision helical gear part.’[j] journal of materials processing technology. 2008,201(1-3) : 25~31. [4] wang g.c, zhao g. q. ‘simulation and analysis of rotary forging a ring workpiece using finite element method’[j]. finite elem. anal. des.2002,38 (12):1151~1164 [5] sung-yuen jung, myung-chang kang and chul kim. ‘a study on the extrusion by a two-step process for manufacturing helical gear’[j]. the international journal of advanced manufacturing technology. 2008,41: 684~693. [6] cheng peiyuan, hu rui and hua lin. ‘finite element analysis of rotary forging in straight tooth bevel gear’[j]. hot working technology. 2006,35(21) : 65~68. [7] j.j. sheu and c.h.yu. ‘the die failure prediction and prevention of the orbital forging process’[j]. journal of materials processing technology. 2008,201(1-3) : 9~13. [8] j.j.sheu and c.h.yu. the cold orbital ‘forging die and process design of a hollow ring gear part’[c]. proc. of the the 35th international matador conference. taipei (2007) 111~114. [9] han xinghui, hua lin.‘3d fe modeling of cold rotary forging of a ring workpiece’ journal of materials processing technology.2009,209(12-13):5353~5362. 300 wang: forming load and metal flow of rotary forging process for spiral bevel gear [10] m.h.sadeghi,t. a. dean. ‘precision forging straight and helical spur gear’ [j]. journal of materials processing technology,1995,45(1-4) : 25-30.. [11] li yamin,wang huajun,wang xueyang,zhu chundong. ‘die fracture of bevel gear in the rotary forging process’[j].china metal forming equipment & manufacturing technology.2009,89(4):89-92. advances in systems science and application (2015) vol.15 no.3 242-253 image retrieval based on spatial organization of patches yan yu1,2 1hubei province key laboratory of systems science in metallurgical process, wuhan university of science and technology, china 2school of science, wuhan university of science and technology, china abstract this paper presents a novel approach for image retrieval by utilizing the spatial organization of image patches. firstly, each image is presented as a set of patches and each patch is descripted as a bag-of-visual-words (bow) feature. secondly, the matrix of matching costs between all pairs of patches that come from the two images respectively is computed by utilizing both the appearance feature and spatial information. finally, the distance between two images is obtained by minimize the total cost of matching using hungarian algorithm. the resulting distance is applied to image retrieval. the experimental results indicate that the proposed method outperforms the traditional bow method in image retrieval scenario. keywords image retrieval; bag-of-visual-words; image patch; image distance measure; hungarian algorithm. 1 introduction the most popular approaches for content-based image retrieval nowadays are based on the model of bag-of-visual-words (bow). they represent image as an orderless collection of local features and demonstrate impressive performance because of their compactness and invariance. the traditional approach based on bow contains the following steps. firstly, local features are extracted from a training set of images by interesting point sampling or dense sampling. secondly, all the feature vectors are quantized to construct a dictionary of visual words by k-means clustering. each cluster tends to contain image patches with similar appearance and its centroid represents a visual word. thirdly, each image is descripted as a histogram indicating the probability distribution of visual words by hard assignment or soft assignment. lastly, the distances between histograms are used for image retrieval. the major limitation of such approaches is that their image representation ignores encoding the spatial information or geometrical relationships among visual words, which has been discovered very important for understanding image content [1]. for solving the problem, some existing approaches leave the spatial verification as a post-processing step [2-3], however, they are only suitable for partial-duplicate image retrieval, as for semantic retrieval they are limited. spatial pyramid matching [4] is successful for representing the spatial information advances in systems science and application (2015) vol.15 no.3 243 of visual words, which subsquently turn into a general framework combined with various feature encoding methods [5-6] for image classification. for analogous approaches like spm, we uniformly call them patch-based methods. in methods like these, image is divided into several non-overlapping rectangular patches, and the feature histogram of each patches is computed and combined as the image feature. consequently, the distance between images is the weighted sum of distances between corresponding patches locating in the same spatial positions. but we know, the foreground objects of similar images maybe lie in different spatial locations, poses and orientations, which make the distance measure is not accurate in many situations. in this work, we present a more flexible patch-based distance measure method integrating with bow, aiming at increasing the precision of the traditional bow for semantic retrieval. each patch of the first image is assigned to exactly one patch of the second one considering both the bow feature and the spatial location of patches in the image, and the distance measure problem of images is transformed into square assignment problem. our idea is similar to [7] which works in pixel level and only makes use of image color feature. however, our method works in patch level and utilize image local features which have stronger descriptive power. the rest of this paper is organized as follows. section 2 discusses the related work. section 3 introduces the image feature representation as a set of patches. section 4 details the proposed image distance measure. section 5 gives the experimental results on image retrieval application. section 6 concludes the paper. 2 related work there are several recent works in the content-based image retrieval (cbir) area which present interesting advances for encoding spatial information of visual words. in the early days, researchers present color correlogram [8] to encode the spatial arrangement of colors. now the study object is turned into local feature rather than pixel color. the most popular approach to encode the spatial information of visual words is spatial pyramid matching, which involves repeatedly subdividing the image and computing histograms of local features at increasingly fine resolutions. it is succesful for image classification, however, it is not robust for some image transformations such as roation. many of the existing approaches leave the spatial verification as a post-processing step. they find the matching points between images, then compute the spatial representation which is used to identify and remove the false matches, compute the similarity between images by using the number of matched features and a penalty term from spatial verification. ransac exploits full geometric verification [9-10], which can greatly improve retrieval performance, but it is com244 yan yu: image retrieval based on spatial organization of patches putationally expensive. another recent approach explicitly encodes the spatial relationship among each pair of matching points by utilizing binary spatial maps [2]. this method is efficient and effective, but its representation is very specific for partial-duplicate search and it is not suitable for the semantic-search application. another recent approah learns the postions of visual words occurrences in an image dataset, called lp-norm pooling [11]. the method first transforms all the images into the same resolution, and use a dense sampling which generates m points in each image. each visual word w has a md vector, in which the m-th value representis the m-th point activate the visual word w. however, the method is designed for image classification and it represent the absolute postions of visual words, which is not robust to translation or rotation of the object in the image. a recent image content description is proposed for describing the spatial layout of visual words, called ∆−tsr [12]. triangle is used to describe the relationship among arbitrary three interesting points, which is represented as a 7d signature composed of labels of correspoind viusal words, angles of vertices, relative orientations and scales of points. after the strategy of triangle selection is applied, the similarity between images is defined as the ratio of the sum of similairity between the most similar triangles and the cardinality of the set of triangles. this method explicitly encodes the spatial relationship among visual words,but it only works with hard assignment. more recently, graph-based approaches for encoding spatial information are proposed. the nested multi-layered local graphs [13] is proposed by building upon sets of surf points with delaunay triangulation. a bow framework is applied on these graphs, giving birth to a bag-of-graphwords representation, visual dictionary is built for each layer of graphs and l1 distance function is used for image retrieval. this method explicitly encodes the spatial relationship of image points, however it has higher computational cost and only work with hard assignment. in this work, the content representation and similarity measure method of images can work with both hard assignment and soft assignment, furthermoe, it is suitable for semantic-retrieval. 3 proposed method 3.1 image as a set of patches for a given image, the set of patches can be represented as {fi}i=1,...,n, fi ∈ rm (1) where fi is the feature vector in the i-th image patch, n is the total number of image patches and m is the length of the feature vector. the pixel sizes of images in our dataset are various. in order to realize the assignment relation between advances in systems science and application (2015) vol.15 no.3 245 fig.1 original images and their sets of patches. patches of two images, all the images are divided into patches of the same amount, rather than limited to the same size of each patch. for example, figure1 shows that two images are divided into 8 rows and 8 columns respectively. although they have different aspect ratio, they all have a set of 64 patches in total. in our method, we encode the appearance feature for each patch rather than the global image. the feature vector of each patch is extracted via bow model according to the following steps. firstly, the images were all preprocessed into gray scale and local features are extracted from images in the training set by dense sampling. in this work, sift descriptor [14] is adopted as local feature descriptor, because it has averagely the best performance among local feature descriptors [15]. dense sampling is realized by a sliding window in the size of 16∗16 pixels, which sliding from top to down and left to right in an image by a step of 8 pixels. these regions are overlapping for obtaining enough abundant local features. secondly, all the feature vectors are quantized to construct a visual dictionary by standard k-means clustering algorithm [16]. lastly, each patch of each image is descripted as a histogram indicating the probability distribution of visual words by hard assignment. many more effective dictionary construction method such as hierarchical clustering, and more advanced coding schemes, such as soft assignment [17], sparse coding [5], fisher coding [18] et.al., all can be utilized in our method. however, they are not key for evaluating our method, so we adopt the easiest clustering algorithm and coding method. consequently, the signature for representing the content of an image is composed of n patch features. the i-th patch feature is represented as fi = {ai, pi}, where ai codes the appearance feature in the patch, whose dimension is the same as the size of visual dictionary. pi = (xi, yi) is composed of the horizontal and vertical spatial positions of the patch in the 2d plane space. note that xi and yi both are normalized to [0,1] for each patch. 3.2 distance between images the method to measure the distance between images is critical for retrieval precision. in this work, a flexible patch-based distance measure method is proposed. given the i-th patch of the first image u and the j-th patch of the second image 246 yan yu: image retrieval based on spatial organization of patches v, the distance between patches is given by d(fu i , f v j ) = wa · d(au i , a v j ) δa + wp · d(p u i , p v j ) δp (2) where d(au i , a v j ) is the distance between appearance features which can be computed by chi-square distance as (3) d(au i , a v j ) = d∑ k=1 (au i (k)−av j (k)) 2 au i (k) +av j (k) (3) where d is the length of the appearance feature vector,au i (k) represent the k-th component value of the vector. and d(p u i , p v j ) is the distance between the spatial positions of the two patches, then we have d(p u i , p v j ) = √ (xui − xvj ) 2 + (yui − yvj ) (4) parameters wa and wp are weight coefficients, which balance appearance and spatial contributions respectively and the sum of them always is 1. the proportion of them has a great influence on the performance of the method, which will be discussed in the next section. parameters δa and δp are chosen according to the appearance and spatial dynamics. typical values used in this paper are δa = 2 and δp = √ 2. the distance d(fu i , f v j ) between two patches can be seen as matching cost, which is denoted as cij . given the set of costs between all pairs of patches in the first image and in the second image, the minimum of the total matching cost can be seen the distance between two images, subject to the constraint that the matching be one-to one. then the distance between image u and image v is given by d(u, v) = min π n∑ i=1 d(fu i , f v π(i)) (6) where π is a permutation of {1, ..., n} to guarantee the constraint. the input to the assignment problem is a square cost matrix with entries cij . the result is a permutation π such that total matching cost is minimized. this is an instance of the square assignment problem, which can be solved using the hungarian algorithm [19]. given a cost matrix to find an optimal assignment using the hungarian algorithm is presented as follow: 1) subtract the smallest entry in each row from all the entries of its row. 2) subtract the smallest entry in each column from all the entries of its column. 3) draw lines through appropriate rows and columns so that all the zero entries of the cost matrix are covered and the minimum number of such lines is used. advances in systems science and application (2015) vol.15 no.3 247 4) test for optimality: (i) if the minimum number of covering lines is n, an optimal assignment of zeroes is possible and we are finished. (ii) if the minimum number of covering lines is less than n, an optimal assignment of zeroes is not yet possible.in that case, proceed to 5). 5) determine the smallest entry not covered by any line. subtract this entry from each uncovered row, and then add it to each covered column. return to 3). compared with the traditional patch based distance measure methods, our method can deal with the position change of the foreground objects in the image automatically. we should note that the distance that we proposed is not a rigid mathematical distance. however, it dose not matter because the distance can be used to represent the divergence between images and it can be utilized in ranking retrieval results. 4 experiments in this section, we evaluate the performance of our proposed method in image retrieval scenario. 4.1 experimental setups experiments are performed on a personal laptop in windows 8.1, with quad intel core i7-4500u running at 1.8ghz. most of our programs are implemented in matlab except dense sift and hungarian algorithm for which we use open source code in c. our dataset is composed of fifteen scene categories, including bedroom, kitchen, livingroom, store, coast, forest, et.al.. each category contains 200 to 400 images, and average image size is 300×250 pixels. the dataset is popular in image classification and retrieval scenario. fig.2 show examples of the dataset. we use 100 images per class randomly selected for generating visual words dictionary and the rest for testing. we use dictionary size 200 and hard assignment for testing our method. in practical application, the soft assignment can be used in our framework for improving retrieval performance. the experimental results are presented in terms of mean average precision (map) and precision for the top n retrieved images (p@n). although map is a very popular measure to assess the effectiveness of cbir methods, it dose not reflect the ranking quality in the first positions. the practical web users usually pay more attentions to the good set of top 10 or 20 retrieved images, so we are more interested in good p@n values than map values. precision measures the retrieval accuracy. it is defined as the ratio between the number of relevant images retrieved and the total number of images retrieved. relevant images are in the same category with the query example. 248 yan yu: image retrieval based on spatial organization of patches (a)livingroom (b)kitchen (c)bedroom (d)mitcoast (e)industrial (f)mitmountain fig.2 sample images from the dataset, highlighting 6 categories (livingroom, kitchen, bedroom, mitcoast, industrial and mitmountain). table 1 comparison between two kinds of partition patterns based on 10 queries. partition pattern p@5(%) p@10(%) average retrieval time(ms) 4× 4 62.00 63.00 23.64 8× 8 78.00 69.00 185.77 table 2 comparison between two kinds of weighting ratios based on 100 queries. weighting ratio p@5(%) p@10(%) p@15(%) p@20(%) 1:1 67.00 59.10 55.47 52.50 2:1 67.40 59.90 55.80 53.65 4:1 68.60 60.00 57.00 55.15 8:1 66.40 59.80 56.33 54.60 4.2 influence about size of patches in this section, the impact of the patch size is illustrated. we use n to present the total number of patches in each image. we divide each images into several rows and several columns, including two partition patterns which are 4 × 4 and 8 × 8 in our experiments. to compare the two pattern patterns, 10 images are randomly selected from the dataset and each of them is iteratively selected as a advances in systems science and application (2015) vol.15 no.3 249 query example, so the total number of queries in each pattern is 10. it is obvious that the size of each patch is decreasing when the total number of the patches is increasing. the details can be seen in table 1. it can be seen that increasing the total number of the patches has a improvent on the retrieval results, however, the computation become more time-consuming. in consideration of the efficiency, in the rest of the paper patches of n=16 are always used. 4.3 appearance-spatial weighting the weighting parameters wa and wp balance appearance feature and spatial contributions respectively. in this section, the impact of ratio of the weighting parameters is illustrated. to compare the influence of weighting parameters, 100 images are randomly selected from the dataset and each of them is iteratively selected as a query example, so the total number of queries in each experiment is 100. if we wish each patch to be matched with the patch which is close to it in spatial position, wp must be set at a high value. otherwise, if we wish each patch in the first image can be matches with any patch in the second image, wp must be set at a low value. we test the weighting parameters including four kinds of ratio patterns which are 1:1, 2:1, 4:1 and 8:1. the details can be seen in table2. it can be seen that increasing the proportion of wa has a slight improvement on results. but if the ratio keep going up and reach 8:1, the precision begin to decline. in the rest of the paper the ratio 4:1 are always used. 4.4 comparison with the traditional method the traditional bow method is also implemented for comparison. to compare our proposed method with it fairly, we iteratively select each image in the dataset as query example, so the total number of queries is 2985 in each query experiment. in table 3, it can be observed that the precision decreases while the number of retrieved images increases. in our proposed method, the p@n is always better than the traditional bow method. therefore, our proposed method is more suitable for image retrieval. table 3 comparison with the traditional bow based on 2985 queries. methods p@5(%) p@10(%) p@15(%) p@20(%) map(%) traditional bow 57.45 50.67 47.92 46.08 29.28 our method 66.11 58.87 55.30 53.15 32.54 in the end, we give some retrieval examples of the query image random selected from the dataset via the two kinds of methods in fig. 3, fig. 4 and fig.5. the query images come from “calsuburb”, “mittallbuilding”and “mitmountain” 250 yan yu: image retrieval based on spatial organization of patches category respectively. in each fig., the query image lies on the top left corner. retrieved top 12 images are arranged from left to right and top to bottom according to the increasing distances between the retrieved images and the query example. it can be seen that the vast majority of the retrieved images are in the same category with the query image in our proposed method. (a)traditional bow (b)our method fig.3 retrieval example for a query image taken from the “calsuburb” category. (a)traditional bow (b)our method fig.4 retrieval example for a query image taken from the “mittallbuilding” category. 5 conclusions in this paper we present a flexible patch-based distance measure method for images, which utilizes the abundant local feature of images and the spatial information of them. the proposed method is tested in image retrieval application. experiments show better retrieval results than the traditional bow. we use hard assignment for testing our method. in the future, we will test the retrieval performance of integrating the soft assignment with our method. in addition, we only utilize the absolute position of visual words in our method. the relative relationship of visual words is more robust to image transformation, which will advances in systems science and application (2015) vol.15 no.3 251 be our next research work. (a)traditional bow (b)our method fig.5 retrieval example for a query image taken from the “mitmountain” category. acknowledgments this work was supported by hubei province key laboratory of systems science in metallurgical process (wuhan university of science and technology):y201406, and the natural science foundation of hubei province: 22013cfa131. references [1] j. philbin, o. chum, m. isard, et al. (2007), “object retrieval with large vocabularies and fast spatial matching”, ieee conference on computer vision and pattern recognition. [2] w. zhou, y. lu, h. li, et al.(2010), “spatial coding for large scale partialduplicate web image search”, 18th acm international conference on multimedia, pp.511-520. [3] h. jégou, m. douze and c. schmid.(2010), “improving bag-of-features for large scale image search” international journal of computer vision,vol.87,pp.316-336. [4] s. lazebnik, c. schmid and j. ponce. (2006), “beyond bags of features: spatial pyramid matching for recognizing natural scene categories”, ieee conference on computer vision and pattern recognition, pp.2169-2178. [5] j. yang, k. yu, y. gong et al. (2009), “linear spatial pyramid matching using sparse coding for image classification”, ieee conference on computer vision and pattern recognition, pp.1794-1801. 252 yan yu: image retrieval based on spatial organization of patches [6] c. c. yan, l. li, z. wang, et al. (2014), “fusing multi-cues description for partial-duplicate image retrieval” journal of visual communication and image representation, vol.25, no.6, pp. 1726-1731. [7] t. hurtut, y. gousseau and f. schmitt. (2008), “adaptive image retrieval based on the spatial organization of colors”, computer vision and image understanding, vol.112, pp.101-113. [8] j. huang, s. r. kumar, m. mitra, et al. (1997), “image indexing using color correlograms”, ieee conference on computer vision and pattern recognition, pp.762-768. [9] h. e. j. egou, m. douze and c. schmid. (2008), “hamming embedding and weak geometric consistency for large scale image search”, 10th european conference on computer vision, pp.304-317. [10] o. chum, m. perdoch and j. matas. (2009), “geometric min-hashing: finding a (thick) needle in a haystack”, ieee conference on computer vision and pattern recognition, pp.17-24. [11] j. feng, b. ni, q. tian, et al. (2011), “geometric lp-norm feature pooling for image classification”,ieee conference on computer vision and pattern recognition, pp.2609-2704. [12] n. v. hoàng, v. gouet-brunet, m. rukoz, et al. (2010), “ embedding spatial information into image content description for scene retrieval ”, pattern recognition, vol.43, pp.3013-3024. [13] s. karaman, j. benois-pineau, r. megret, et al. (2012), “multi-layer local graph words for object recognition ”, 18th international conference on multimedia modeling, pp.29-39. [14] d. g. lowe. (1999), “object recognition from local scale-invariant features”, 7th ieee international conference on computer vision, pp.1150-1157. [15] k. mikolajczyk and c. schmid. (2005), “a performance evaluation of local descriptors”, ieee transactions on pattern analysis and machine intelligence, vol.27, pp.1615-1630. [16] c. m. bishop. (2006), pattern recognition and machine learning, springer, pp.424-430. [17] j. c. van gemert, j. geusebroek, c. j. veenman, et al. (2008), “kernel codebooks for scene categorization ”, 10th european conference on computer vision, marseille, pp.696-709. advances in systems science and application (2015) vol.15 no.3 253 [18] f. perronnin, j. s a nchez and t. mensink. (2010), “improving the fisher kernel for large-scale image classification”, 11th european conference on computer vision, pp.143-156. [19] c. h. papadimitriou and k. stieglitz. (1982), “combinatorial optimization: algorithms and complexity”, prentice hall. corresponding author yan yu can be contacted at: yuyan wust@163.com advances in systems science and application (2016) vol.16 no.1 19-46 research on extension strategy generation of sustainable utilization of regional water resources qiaoxing li1,2 and naiang wang2 1school of management, guizhou university, guiyang, guizhou province, 550025, pr china; 2college of earth and environmental sciences, lanzhou university, lanzhou, gansu province, 730000, pr china. abstract sustainable utilization of water resources has been a global issue of common concern. under the prerequisites to meet the development need of the region’s population, resources, environment and economy, we should solve simultaneously both water crisis and ecological environment construction, which may be more important for the arid and semi-arid regions. although a lot of results in other aspects of sustainable utilization of water resources have been achieved, it has not been discussed how to generate the strategies to deal with the contradictory problem between water crisis and ecological environment construction. in this paper, extension analysis and extension transformation were utilized to generate a variety of innovative schemes for sustainable utilization of water resources. by using an appropriate evaluation methodology, we chosen one or more excellent schemes and changed them into the specific programs that can be implemented in practice. the strategy generation process can be achieved by computer software, which makes the decision-making be intelligent. the extension generation method provides a formalized and quantitative way for effectively solving water crisis and ecological environment construction, and clarifies the relevant departments ideas. it is also fit for basin water resources. keywords sustainable utilization of water resources; water crisis; ecological environment; utilization strategy; extensible methods 1 introduction as society and economy developed, the demand of water increased. however, the available water provided by the natural world is limit. at present, the water resources problem in the world especially the arid and semi-arid regions is becoming more and more serious, and it mainly shows that water scarcity is increasing and water conflicts are of comprehensive intensification. initially, the water management only focused on some aspects of water utilization and protection such as supply, irrigation, hydropower and others. the traditional management ways of water supply and decentralized sectors ignored the integrity of natural environment and the diversity of water use. therefore, we could not coordinate the various water relations and did not achieve the sustainable utilization of water 20 qiaoxing li and naiang wang babu:research on extension strategy generation ... resources. in order to avoid this deficiency, some scholars have proposed integrated management of water resources which unified water, land and related resources into a jointly development[1-6]. other scholars believed that the fundamental way to ease the water crisis and the eco-environment destruction is to implement sustainable management of water resources[7,8]. they considered the development and utilization of water resources together with the complex system of social economy water resources ecological environment, and inquired the concrete pathway of water sustainable development, which means that we should support the coordinated development among population, resources, environment and economy, and meet the water needs of intrageneration and intergeneration of human being under the promise of maintaining the social continuity and the ecosystem integrity. sustainability is the most rational utilization pattern to integrate the development, utilization, protection, control and management of water resources, and the essence of sustainable utilization is to coordinate water utilization with environmental protection, economic growth and social development. many results of sustainable utilization of water resources have been achieved, such as efficiency evaluation of water use[9-11], water allocation problem[1214], ecological environment impacts[5, 6, 15], analysis for present situation and countermeasures[16], and rules establishment and construction which included effective system of water property, open water market, reasonable emerged mechanism of water price and complete management system of water resources[9-10], and so on. although scholars have made abundant achievements in many aspects of the sustainable utilization of water resources, we are still hard to coordinate the development among population quality, natural resources, living environment and national economy, and can not accurately grasp the quantitative relationship between the water crisis and the construction of ecosystem. to study the strategy generation methods of sustainable utilization of water resources is helpful for solving the relationship between the water crisis and the eco-environment construction in practice, and improving the decision-making capacity of water resources management. it has more practical significance to support the sustainable development of nature, economy and society. the following sections of this paper are below. we introduced some basic knowledge of extenics in section 2, and proposed extension strategy generation step of sustainable utilization of regional water resources in section 3. the extension planning process of sustainable utilization is introduced in section 4, and we give a modified case in section 5. at the end of this paper, a conclusion is given. advances in systems science and application (2016) vol.16 no.1 21 2 preliminaries of extenics extenics is an original interdisciplinary which was put forward by a chinese scholar in 1983. it discusses the rules and methods of extension and innovation of objective things by using the formalized models, which are utilized to solve the contradictory problems. contradictory means that people’s goals can not be reached in current conditions[17]. at present, the contents of extenics include basic-element theory, extension set theory and extension logic. for the sake of simplicity and convenience, we only introduce the knowledge of basic-element theory that will be utilized in this paper. basic-element theory includes extensible analysis theory (eat), conjugate analysis theory (cat) and extension transformation theory (ett). eat includes divergence analysis theory, correlative analysis theory, implication analysis theory and opening-up analysis theory. cat includes nonmaterial and material conjugate analysis, soft and hard conjugate analysis, latent and apparent conjugate analysis as well as negative and positive conjugate analysis. ett includes basic extension transformations, conductive transformation and conjugate transformation, calculation of extension transformations and nature of extension transformations. all of knowledge above will be introduced below and can be seen in [17]. firstly, we introduce the basic concept of basic-element, which includes matterelement, affair-element and relation-element. basic-element is a logical cell of extenics. definition 2.1: an ordered triple m = (om, cm, vm), which is as the fundamental element for matter description composed of matter om, the characteristic cm of omand the value vm of om about cm, is called as one-dimensional matterelement. furthermore, the following array composed of matter om, n-names of characteristics of cm1, cm2,· · · , cmn and the corresponding value vmi of om about cmi (i = 1, 2, · · · , n) m =  om cm1 vm1 cm2 vm2 ... ... cmn vmn  = (om, cm, vm) or m = [ om cm1 · · · cmn vm1 · · · vmn ] , is n-dimensional matter-element, where cm = [cm1, cm2, · · · , cmn] t and vm = [vm1, vm2, · · · , vmn] t . definition 2.2: the ordered triple a = (oa, ca, va,) as the fundamental element for affair description composed of the action oa, the characteristic cm of oa and the value va of oa about ca, is called one-dimensional affair-element. furthermore, the following array composed of action oa, n characteristics of ca1, 22 qiaoxing li and naiang wang babu:research on extension strategy generation ... ca2,· · · , can and the obtained value vai of oa about cmi (i=1,2,· · · ,n) a =  oa ca1 va1 ca2 va2 ... ... can van  = (oa, ca, va) or a = [ oa ca1 · · · can va1 · · · van ] , is n-dimensional affair-element, where ca = [ca1, ca2, · · · , can]t and va = [va1, va2, · · · , van]t . definition 2.3: the n-dimensional array composed of relative or relation symbol (refer to as relation name) or, n characteristics cri of or and the corresponding value vri of or about cri (i=1,2,· · · ,n) which is as the following: r =  or cr1 vr1 cr2 vr2 ... ... crn vrn  = (or, cr, vr) or r = [ or cr1 · · · crn vr1 · · · vrn ] , where cr = [cr1, cr2, · · · , crn]t and vr = [vr1, vr2, · · · , vrn]t , and r describes the relation between vr1 and vr2, is n-dimensional relation-element. because the relation-element r often contains the same characteristics such as antecedent cr1, consequent cr2, degree cr3, maintaining mode cr4, contact channel cr5, contact method cr6, location cr7, and so on, we simply denote it as r (or, vr1, vr2, · · · ). definition 2.4: matter-element, affair-element and relation-element are collectively referred to as basic-element which is expressed as b =  o c1 v1 c2 v2 ... ... cn vn  = (o,c, v ) or b = [ o c1 · · · cn v1 · · · vn ] , where o (object) indicates a certain object (matter, action or relation), and ci is a characteristic of object o, and vi is the corresponding value of o about ci (i=1, 2, · · · , n), and c = [c1, c2, · · · , cn]t and v = [v1, v2, · · · , vn]t . secondly, we introduce the principles of extensible analysis which include divergent, correlative, implication and opening-up ones below: principle 2.1 (principle of divergent analysis): from one basic-element, multiple basicelements with the same object (or characteristic) can be extended, and the set of basic-elements with the same object (or characteristic) must be non-empty. advances in systems science and application (2016) vol.16 no.1 23 inference 2.1: from one basic-element, multiple basic-elements with the same object and value (or with the same characteristic and value, or with the same object and characteristic) can be extended. definition 2.5: given two sets of basic-elements {b1} and {b2}, for any b1 ∈ {b1}, if there is at least one b2 ∈ {b2} to let b1 corresponds to b2 , then {b1} and {b2} are correlative and we denote them as {b1} ∼ {b2}. in particular, as to the sets of basic-elements {b1} and {b2} with c0 as their evaluated characteristic, for any b1 ∈ {b1}, if there is at least one b2 ∈ {b2} to let c0(b2) = f(c0(b1)), then {b1} and {b2} are correlative about the evaluated characteristics c0 and it is denoted as {b1} ∼ c0 {b2}. principle 2.2 (principle of correlative analysis): for a given matterelement m = (om, cm, vm), there is at least one matter-element with the same characteristic mc = (o′ m, cm, cm(o′ m)) or matter-element with the same matter mo = (om, c′m, c′m(om)) or matter-element with different matters m ′ = (o′ m, c′m, c′m(o′ m)), to let m ∼ mc, or m ∼ mo or m ∼ m ′. definition 2.6: suppose b1 and b2 are two basic-elements, and b1 is realized inevitably with the realization of b2, then we call that the basic-element b1 implies the basic-element b2, and it is denoted as b1 ⇒ b2. especially, if b1 ⇒ b2 under the condition l, then b1 ⇒ (l)b2. definition 2.7: let b, b1 and b2 be basic-elements, and (1) if both b1 and b2 are realized inevitably with the realization of b, then the basic-elements b1 and b2 imply the basic-element b, and it is denoted as b1 ∧b2 ⇒ b. (2) if either b1 or b2 is realized inevitably with the realization of b, then the basic-element b1 or b2 implies the basic-element b, and it is denoted as b1 ∨b2 ⇒ b. (3) if b is realized inevitably with the realization of both b1 and b2, then the basic-element b implies the basic-elements b1 and b2, and it is denoted as b ⇒ b1 ∧b2. (4) if b is realized inevitably with the realization of either b1 or b2, then the basic-element b implies the basic-element b1 or b2, and it is denoted as b ⇒ b1 ∨b2. principle 2.3 (principle of implication analysis): if b1⇒b2 and b2⇒b3, then b1⇒b3. definition 2.8: the possibilities of composing, decomposing and expanding/contracting that matter, affair and relation own are called as composability, decomposability and expandability/contractability, correspondingly, and they are identified as openness of basic-element. according to composability of basic-element, one matter can combine with other matter to generate new matter. by decomposability, one matter can be 24 qiaoxing li and naiang wang babu:research on extension strategy generation ... decomposed into several new matters with certain characteristics that may be different from ones of original matter. similarly, one matter can be expanded or contracted to provide possibility for solving contradictory problems. principle 2.4 (principle of opening-up analysis): for any basic-element, the opening-up analysis includes the analysis of composability, decomposability and expandability/contractability: (1) composability analysis: given a basic-element b1 = (o1, c1, v1), there is at least one basicelement b2 = (o2, c2, v2) to allow b1 and b2 to be composed into b, and b2 is called as the composable basic-element of b1, where b = b1 ⊕ b2 and 1) when o1 = o2 and c1 ̸= c2, b = (o1, c1 ⊕ c2, v1 ⊕ v2) = ( o1 c1 v1 c2 v2 ) ; 2) when o1 ̸= o2 and c1 = c2, b = (o1 ⊕ o2, c1, v1 ⊕ c1(o2)) ; 3) when o1 ̸= o2 and c1 ̸= c2, b = ( o1 ⊕ o2 c1 v1 ⊕ c1(o2) c2 c2(o1) ⊕ v2 ) ; (2) decomposability analysis: any basic-element b=(o,c,v) can be decomposed into several basic-elements b1, b2, · · · , bm with the same characteristics under certain condition l, where bi = (oi, c, c(oi))(i = 1, 2, · · · ,m), and we denote b//(l) {b1, b2, · · · , bm}. (3) expandability/contractability analysis: for a real positive number α , any basic-element b=(o,c,v) can be expanded (α > l) or contracted (l > α > 0) as αb = (o,c, αv) under certain condition. thirdly, we introduce the principle of conjugate analysis below: extenics help us to completely understand the matter from its physical, systematic, dynamic and antithetic properties. in terms of physical property of matter, all matters are composed of a physical part and a non-physical part, then the physical part is called material part and the nonphysical part is nonmaterial part by extenics, respectively. when considering a matters structure from its systematic property, the matters components as a whole are called the hard part of matter and the relations between the matter and its components as well as between the matter and other matters are the soft part of the matter. furthermore, from the dynamic property, any matter is changing continuously. stagnation is ever-relative while motion is permanent. the parts that have not been appeared are called as the latent parts of the matter and the appeared parts are the apparent parts of the matter. the latent part may become apparent under certain conditions. from the antithetic property, all matters have two parts that are antithetic. in extenics, the part producing the positive value in the measure of matter about certain characteristic is the positive part of matter and the other taking negative value is the negative part of matter. so matter’s conjugate advances in systems science and application (2016) vol.16 no.1 25 analysis includes nonmaterial-material, soft-hard, negative-positive and latentapparent ones. we should analyze a matter according to the following property. principle 2.5 (principle of conjugate analysis): all matters have four pairs of conjugate parts, i.e., the nonmaterial-material, the soft-hard, the negativepositive and the latent-apparent ones. each conjugate part of any matter has numerous characteristics and it can be denoted by an n-dimensional basic-element. at the same time, in each pair of conjugate part of any matter, one certain conjugate part has at least one characteristic that is relevant to certain characteristic of its corresponding conjugate. fourthly, we introduce the extension transformation below: the tool for solving the contradictory problem is extension transformation. by using certain transformations, an unfeasible problem can be transformed a feasible one. we only introduce the general concept and the basic types of transformation. definition 2.9: supposing that the object γ0 is a matter-element, affairelement, relationelement, criteria or any element in the universe of discourse, and the transformation from γ0 to the object γ or multiple objects γ1,γ2, · · · ,γn in the same class is called as extension transformation of γ and it is denoted as tγ0 = γ or tγ0 = {γ1,γ2, · · · ,γn}. definition 2.10: the object γ have five types of basic transformations below: (1) substitution transformation, i.e., tγ = γ0; (2) increasing/decreasing transformation, i.e., the increasing transformation tγ = γ ⊕ γ0 and the decreasing transformation tγ = γ− γ0; (3) expansion/contraction transformation: tγ = αγ, where it is expansion when α > 1 and contraction when 1 > α > 0; (4) decomposition transformation: tγ = {γ1,γ2, · · · ,γn}, where γ0 = γ1 ⊕ γ2⊕ · · · ⊕ γn; (5) duplication transformation: tγ = {γ,γ1}. principle 2.6 (principle of extension transformation): for any object γ , there must exist a certain transformation t to let tγ = γ0, where γ ̸= γ0. on the other hand, if there is a certain transformation t to let tγ = γ0, there should be another transformation t1 to let t1γ = γ0. furthermore, it can be utilization several transformations to successfully solve a contradictory problem. 3 extension strategy generation step of sustainable utilization of regional water resources during the process of sustainable utilization of water resources, we need to simultaneously solve two goals of water crisis and ecological environment construction in the planning area, and should meet the needs of population promotion, resource exploitation, environmental protection and economic development in the 26 qiaoxing li and naiang wang babu:research on extension strategy generation ... planning period. in order to effectively solve this problem, we propose the extension strategy generation of sustainable utilization of water resources according to the literature [17] below: 1) define the contradictory problem. combining with the regional goals and development needs, we draw up the basic-elements of goals (water crisis solution and ecological environment construction) and the basic-elements of conditions (population promotion, resource development, environmental protection and economic development) of sustainable utilization of water resources in the planning period. under the basic-elements of conditions, the two goals can not be achieved at the same time, and they constitute a contradictory problem. 2) assume that the targets are not changed and do extension analysis for the four basicelements of conditions. suppose that the premise of targets without change is that the tasks of targets are feasible, then we should obtain the appropriate basic-elements of goals by using correct investigation methods. at that moment, we get the extension analysis matter-elements through expanding the condition basic-elements by using correlative analysis, divergent analysis, conjugate analysis and opening-up analysis. 3) generate a variety of utilization strategies of water resources by extension transformation to the extension analysis matter-elements. the purpose of doing extension transformation is to resolve the contradictory problem and make two objectives be achieved simultaneously. extension transformation to the extension analysis matter-elements is mainly for the values of characteristics of matter-elements, and the transformations include substitution, increasing, decreasing, expansion, contraction and decomposition, as well as the integration of these basic transformations. 4) evaluate and select the utilization strategies of water resources. we may get a group of extension transformation matter-elements after doing a complete extension transformation, and then form an integral strategy. however, it is not true that every strategy can effectively solve the problem between water crisis and ecological environment construction, or that the solution effect is the same as others. therefore, we should evaluate and select the strategies by using the following steps: (1) at first, we should determine some indexes a1, a2, · · · , al to measure the strategies which make up a measurement condition set. the measurement condition set is utilized to judge whether the strategies are feasible or not. the measurement criterions of sustainable utilization of water resources in general are technical feasibility and economic viability, and the technical feasibility is usually viewed as the prerequisite which must be satisfied. (2) secondly, evaluate and select the strategies. first step, we abandon the strategies which do not meet the prerequisites, i.e., if the strategies whose technoladvances in systems science and application (2016) vol.16 no.1 27 ogy is not feasible, then they are removed, and we utilize the rest of measurement criterions to justify other strategies. next step, by using the correct methods of investigation and analysis to get the values of remaining strategies for other criterions, we select an appropriate function such as the comprehensive evaluation method to compute the superior degrees of these strategies, and choose the strategies whose degree is the maximum as the optimal ones. 5) draw up the specific programs of action according to the optimal strategies. the purpose to form the action programs is to make the decision-makers and executors understand the optimal strategies, so that they can correctly implement them and accomplish the goals. 4 extension planning process of sustainable utilization of regional water resources 1) define the contradiction problem of water utilization the process of water utilization involves the natural, economic and social systems, and induces many contradictory problems between the ecological environment and the economic development goal as well as the social development mode. in the arid and semi-arid regions, we should solve the contradiction between water crisis and ecological environment construction under the premise of regional development of population, resources, environment and economy. in extenics, the formalized model and the extension analysis methods are utilized to study how to resolve conflict issues, which are also suitable for selecting feasible and effective utilization strategies of water resources and provide a formalized approach for relevant departments. the extension analysis process of water utilization should be made efforts to solve the central issues of water crisis and ecological environment construction on the basis of development of population, resources, environment and economics in this regional which includes administrative region and watershed. in the arid and semi-arid regions, water crisis and ecological environment construction can not be achieved at the same time, so this is a contradictory problem. sustainable utilization of water resources means the synthesis between ecological environment construction and water utilization. therefore, the goal of contradictory problem that should be resolved within the region is ecological environment construction and water crisis, and the problems conditions are development needs of population, resources, environment and economics. then the regional government may formulate the general objective basic-elements g1 and g2 below: g1 = ( construct a1 a2 a3 a t b1 b2 b3 a t ) and g2 = ( w a c2 c3 t a w2 w3 t ) , where for the basic-element g1, construct is a verb, and the characteristics are a1=dominating object, a2 = acting object, a3 = receiving object, a = location 28 qiaoxing li and naiang wang babu:research on extension strategy generation ... and t = time, and the values of these characteristics are b1 = ecological environment, b2 = {administration, beneficiaries, contractors, collaborators, etc.}, b3 = {vegetation, land, species}, a = a certain region which may be an administrative region or watershed, and t = a planning period, such as three-year plan, etc.. apparently, g1 is an affair-element. for another basic-element g2, w=water resources, and the characteristics are c2 = ownership quantity and c3 = demand quantity, and the unit of the values w2 and w3 is hundred million cubic meters. here, the equation w3 > w2 holds, which indicates that the water demand within the region a is larger than the water owner, and the water crisis exists in this region. the quantity w2 is an average statistical value of water resources for recent years, and the quantity w3 is a forecasting value of water demand of the region a in the planning period. also apparently, g2 is a matter-element. in addition, the premise to resolve the contradictory problem between ecological environment construction and water crisis is to meet the development needs of regional population, resources, environment and economics, thus we establish the following condition basic-elements: l1 = ( upgrade a1 a2 a3 a t r11 r12 r13 a t ) , l2 = ( exploit a1 a2 a3 a t r21 r22 r23 a t ) , l3 = ( improve a1 a2 a3 a t r31 r32 r33 a t ) and l4 = ( develop a1 a2 a3 a t r41 r42 r43 a t ) , where the values of characteristic a1 are r11=population, which includes rural and urban ones, r21= natural resources, r31=living environment and r41=economics, and the values of characteristic a2 are corresponding administrations, and the values of characteristic a3 are r13=population qualities, which include population quantity, educational background, etc., r23= {land, minerals, wild beast}, r33={green belt, waste treatment, transportation belt}, and r43=economics. waste treatment mainly refers to exhaust, waste water and garbage, and transportation belt includes the road clean, green, construction and maintenance, etc.. because we can not simultaneously achieve the goals g1 and g2, it is a contradictory problem, and thus we establish the extension model p = (g1 ∧g2) ↑ (l1 ∧ l2 ∧ l3 ∧ l4). the purpose to do extension analysis is to transfer the contradictory problem into the co-existence issue by doing the correlative analysis, the divergent analysis, the conjugate analysis and the opening-up analysis on the goal basic-elements and the condition basic-elements. in general, there are five approaches to resolve contradictory problem with multi-objectives and multi-conditions: firstly, all targets are not changed, but make some (or all) conditions be transformed; secondly, all conditions are unchanged but some (or all) targets are transformed; thirdly, part of the targets are changed and some (or all) conditions are altered; fourthly, part of conditions advances in systems science and application (2016) vol.16 no.1 29 are changed and some (or all) targets are transferred; fifthly, all of goals and conditions are changed. in this paper, we only discuss the most simple issue, i.e., all targets are unchanged and some (or all) conditions are transferred. by transferring the conditions, we solve contradictory problem between water crisis and ecological environment construction. for the other issues, we will analyze them in other papers. firstly, we define the specific target basic-elements and simplify the condition basic-elements on the basis of the original problem p. according to the reality and overall targets of the region a, we decompose the goal affair-element g1 on the basis of the investigation and analysis results, then obtain the specific target matter-elements of sustainable utilization of water resources below: g′ 1 =  g1 a11 b11 a12 b12 a13 b13 a14 b14 a15 b15  , g′′ 1 =  g2 a21 b21 a22 b22 a23 b23 a24 b24 a25 b25  and g′′′ 1 =  g3 a31 b31 a32 b32 a33 b33 a34 b34 a35 b35  , where for the matter-element g′ 1, the object g1 is vegetation, and the characteristics of g1 are a11= administrative, a12= original acreage, a13= planning acreage, a14= original amount of water demand and a15 = incremental of water demand; for the matter-element g′′ 1, the object g2 is land, and the characteristics of g2 are a21= administrative, a22= remediation acreage, a23= incremental of vegetation, a24= incremental of building and a25 = incremental of water demand, as well as the equation b22= b23+ b24; for the matter-element g′′′ 1 , the object g3 is species, and the characteristics of g3 are a31= administrative, a32= original amount, a33= incremental of species, a34= original amount of water demand and a35 = incremental of water demand. then, the target affair-element g1 is transformed into three specific target matter-elements. furthermore, the land g2 mainly refers to the unutilized land in the region a. meanwhile, in the case of non-confusion, we can simplify the condition basicelements by omitting the characteristics location and time. therefore, we transform the condition basicelements l1, l2, l3 and l4 into the simple forms below: l′1 = r1 a1 r11 a2 r12 a3 r13  , l′2 = r2 a1 r21 a2 r22 a3 r23  , l′3 = r1 a1 r31 a2 r32 a3 r33  and l′4 = r1 a1 r41 a2 r42 a3 r43  , 30 qiaoxing li and naiang wang babu:research on extension strategy generation ... where r1=upgrade, r2=exploit, r3=improve and r4=develop. through the transformations above, the original problem p = (g1 ∧ g2) ↑ (l1 ∧ l2 ∧ l3 ∧ l4) is converted to p ′ = (g′ 1 ∧ g′′ 1 ∧ g′′′ 1 ∧ g2) ↑ (l′1 ∧ l′2 ∧ l′3 ∧ l′4). because the problem p is the specific form of the original one p, it can be resolved when p is achieved. thus we only do extension analysis for the condition basic-elements of the problem p. 2) extension analysis process of condition basic-elements the correlative analysis on condition basic-elements: the correlation of basicelement discusses the correlations among the same or different objects (matters, actions or relationships) as well as among their characteristics or values. it can make people more clearly understand the interactions between each of condition basic-elements and the sustainable utilization of water resources by using the formalized way. on the basis of the correlation between water resources and the development of population, resources, environment and economics, we get the correlative matterelements w1, w2, w3 and w4 of condition basic-elements below: w1 = w k1 w11 k2 w12 k3 w13  , w2 = w k1 w21 k2 w22 k3 w23  , w3 = w k1 w31 k2 w32 k3 w33  and w4 = w k1 w41 k2 w42 k3 w43  , where the matter w=water, and the characteristics k1, k2 and k3 represent user, use volume and use mode, respectively, and the values of k1 are w11 = population, w21 = natural resources, w31 = living environment and w41 = economy, and the values w12, w22, w32 and w42 of k2 are the average statistical value of regional water utilization in recent years which are consumed by population, natural resources, environment and economics, respectively, and the values of characteristic k3 are w13={drink, wash, rinse, etc.}, w23={produce, drink, etc.}, w33={irrigate, rinse, etc.} and w43= {irrigate, produce, wash, drink, rinse, etc.}. the new matter-elements w1, w2, w3 and w4, which come from the results after doing the correlative analysis on the condition affair-elements, are called the correlative matter-elements of condition basic-elements. divergent analysis on the correlative matter-elements: the divergence of matterelement includes: different characteristic with same matter, different value with same characteristic, different matter with same value, same characteristic with same matter, same value with same characteristic, same value with same matter, similar value with same characteristic, etc. by using the divergent analysis, we can understand the multiple aspects of correlative matter-elements which involve in the process of water utilization. doing the divergent analysis on the values of advances in systems science and application (2016) vol.16 no.1 31 characteristic k1 of correlative matter-elements according to same characteristic with same matter, we grasp all water users in region a and get following matterelements ( 7→ means divergence): w1 7→ w 1 1 = ( w k1 k2 k3 w1 11 w1 12 w1 13 ) , w 2 1 = ( w k1 k2 k3 w2 11 w2 12 w2 13 ) ; w2 7→ w 1 2 = ( w k1 k2 k3 w1 21 w1 22 w1 23 ) , w 2 2 = ( w k1 k2 k3 w2 21 w2 22 w2 23 ) , w 3 2 = ( w k1 k2 k3 w3 21 w3 22 w3 23 ) ; w3 7→ w 1 3 = ( w k1 k2 k3 w1 31 w1 32 w1 33 ) , w 2 3 = ( w k1 k2 k3 w2 31 w2 32 w2 33 ) , w 3 3 = ( w k1 k2 k3 w3 31 w3 32 w3 33 ) ; w4 7→ w 1 4 = ( w k1 k2 k3 w1 41 w1 42 w1 43 ) , w 2 4 = ( w k1 k2 k3 w2 41 w2 42 w2 43 ) , · · · , wm 4 = ( w k1 k2 k3 wm 41 wm 42 wm 43 ) , where w1 11 = rural residents, w2 11 = urban residents ; w1 21 = land, w2 21 = minerals, w3 21 = biology; w1 31 = green belt, w2 31 = waste treatment, w3 31 = transportation belt; and the matters w1 41 , w2 41 , · · · , wm 41 are m industrial departments in region a such as agriculture, tourism, catering, etc.; and the values w1 12 , w2 12 , w1 22 , w2 22 , w3 22 , w1 32 , w2 32 , w3 32 , w1 42 , · · · , wm 42 of characteristic k2 are the average statistical values of water utilization which consumed by corresponding users of region a in recent years, and the values of characteristic k3 are w 1 13 ⊆ w13 = {drink,wash, rinse} , w 1 23,w 2 23,w 3 23 ⊆ w23 = {produce, drink}, w 1 33,w 2 33,w 3 33 ⊆ w23 = {irrigate, rinse}, as well as w 1 43, · · · ,wm 43 ⊆ w43 = {irrigate, produce, wash, drink, rinse}. the basic elements which come from the correlative matter-elements by using the divergent analysis are called divergent matter-elements. conjugate analysis on divergent matter-elements: the properties of system, material, dynamic and opposition that things own are collectively known as conjugation. we can understand things in a more comprehensive perspective and reveal the development and variation nature of things in a more profound way by using the conjugate analysis. extension theory describes the structure of things 32 qiaoxing li and naiang wang babu:research on extension strategy generation ... from four pairs of conjugate and opposite concepts which are material and nonmaterial, hard and soft, apparent and latent, as well as positive and negative. in generally, we utilized to consider the problems from the material, hard, apparent and positive aspects of things. however, the nonmaterial, soft, latent and negative angles of things can make us obtain unexpected results. because the divergent matter-elements come from the water users which are population, resources, environment and economics through divergence analysis, we will do conjugate analysis by combining the user k1 with water resources w. from the nonmaterial part of things, we firstly determine the requirements of user k1 and get a new characteristic named water quality k4 whose values for all divergent matter-elements are vi4 ∈ {i ∼ v grade}, where the standard of water quality refers to the book named china’s environmental quality standard of surface water (gb3838-2002), and the subscript i denotes the serial number of the divergent matter-elements from w 1 1 to wm 4 . secondly, we justify the economic status and consumption of user k1 in the region a, and then get a new characteristic of water named importance k5, and the values vi5 of k5 satisfy vi5 ∈ {necessary, priority, current situation, decrease, no}, where no means that the certain user does not consume water and we may remove the corresponding matterelement. form the soft aspect of things, we view the relationship among the orders of water users, and get two characteristics of water resources named as former user k6 and later user k7 whose values vi6 and vi7 are the elements from the set composed by the users in matter-elements above. from latent aspect of things, we analyze the potential utilization values of the consumption and discharge of water and get a characteristic denoted as recyclability k8 whose value vi8 ∈ { yes, no}. furthermore, according to recyclability, we should determine the characteristic named recovered amount k9, and denote its value as vi9. from the negative aspect of things, we mainly view that whether waste water has negative impact for environment or not, and get a characteristic named emission k10 with value vi10 ∈ { direct emission, treatable emission, non-emission}, where nonemission means that the water has been completely consumed by the corresponding user. then we investigate the technology to treat the wastewater and get another characteristic named treatable technology k11 with the value vi11 ∈ { high, normal, low, no }, where no means that the value vi10 is direct emission or non emission. then we obtain the conjugate matter-elements by using conjugate analysis on divergent matter-elements. the opening-up analysis on conjugate matter-elements: the possibility of matter-elements combination or decomposition is called the opening-up property of matter-element which includes addition, multiplication and decomposition of matters, characteristics and values. the opening-up property of matter-element provides another approach to solve the contradictory problems. at first, we disadvances in systems science and application (2016) vol.16 no.1 33 cuss the decomposition of water resources, and obtain the characteristic named as source k12 of desirable water in region a whose value vi12 ∈ {surface water, groundwater, external water, recycled water, all water}, where recycled water is reutilized water while wastewater is retrieved, and all water contains surface water, groundwater, external water and recycled water. secondly, from the additive or multiplicative property of characteristics, we examine user k1, water quality k4 and former user k6, and get the characteristic named freshness k13 with the value vi13 ∈ {fresh water, circled water}, where fresh water includes surface water, groundwater and external water. finally, according to the additive property of values, we study on utilization volume k2, recyclability k8, source k12 and freshness k13 and obtain the characteristic increment k14 with the value vi14. from the values vi12 of utilization volume k2 and vi14 of increment k14, we get another characteristic named recycle volume k15 with the value vi15. at last, by using the values vi2, vi14 and vi15, we get a characteristic named total volume k16 with the value vi16. the new matterelements from the conjugate matterelements by using opening-up analysis are called opening-up matter-elements. condition matter-elements are now transferred into a group of new matterelements with the object named water and 16 characteristics through extension analysis which includes correlative analysis, divergent analysis, conjugate analysis and opening-up analysis, and they are called as the extension-analysis matterelements of condition basic-elements. from the analysis showed above, the final results, i.e., the opening-up matter-elements, are the extension-analysis matterelements. 3) the generation and optimal selection of sustainable utilization strategy of regional water resources and the program implementation of optimal strategy by using extension analysis, the condition-basic-elements are transferred into extensionanalysis matter-elements of sustainable utilization of water resources. for the sake of convenience, we re-order the characteristics of extension-analysis matter-elements as user k1, use mode k2, water quality k3, importance k4, former user k5, later user k6, recyclability k7, emission k8, treatable technology k9, source k10, freshness k11, use volume k12, recovered amount k13, increment k14, recycle volume k15 and total volume k16, where the characteristics with nonquantities are in the front and the others with quantities are at the back, and the quantitative characteristics satisfies the following relationship: (1) the value vi12 of use volume k12 ≤ the absolute value vi13 of recovered amount k13, and let vi13 ≤ 0 which represents an increase of water resources; (2) use volume k12, increment k14 ∈ {surface water, groundwater, external water} = {initial water}, and their values vi12 + vi14 = the volume of initial water, where vi12 ≥ 0 and vi14 ≥ 0, and they means the consumption; (3) the value vi15 of recycle volume k15 ≥ 0, which means that the amount 34 qiaoxing li and naiang wang babu:research on extension strategy generation ... of recycle water was consumed by user i, and vi15 ≤ |vi13|; (4) the value vi16 of total volume k16 satisfies vi16 = vi12 + vi14 + vi15, which represent the total amount of water resources consumed by user i. in general, extension transformation of extension-analysis matter-elements, which includes substitution, decomposition, addition, decrease, etc., is only in connection with those quantitative characteristics of water resources, then we obtain a set of new condition matter-elements with new values and we call them as extension-transformation matter-element below: t (w 1 1 ) = ( w k1 · · · k15 k16 w1 1,1 · · · w1 1,15 w1 1,16 ) , t (w 2 1 ) = ( w k1 · · · k15 k16 w2 1,1 · · · w2 1,15 w2 1,16 ) ; t (w 1 2 ) = ( w k1 · · · k15 k16 w1 2,1 · · · w1 2,15 w1 2,16 ) , t (w 2 2 ) = ( w k1 · · · k15 k16 w2 2,1 · · · w2 2,15 w2 2,16 ) , t (w 3 2 ) = ( w k1 · · · k15 k16 w3 2,1 · · · w3 2,15 w3 2,16 ) ; t (w 1 3 ) = ( w k1 · · · k15 k16 w1 3,1 · · · w1 3,15 w1 3,16 ) , t (w 2 3 ) = ( w k1 · · · k15 k16 w2 3,1 · · · w2 3,15 w2 3,16 ) , t (w 3 3 ) = ( w k1 · · · k15 k16 w3 3,1 · · · w3 3,15 w3 3,16 ) ; t (w 1 4 ) = ( w k1 · · · k15 k16 w1 4,1 · · · w1 4,15 w1 4,16 ) , t (w 2 4 ) = ( w k1 · · · k15 k16 w2 4,1 · · · w2 4,15 w2 4,16 ) , · · · , t (wm 4 ) = ( w k1 · · · k15 k16 wm 4,1 · · · wm 4,15 wm 4,16 ) . by carrying out a transformation t, we get a set of extension-transformation matter-elements, which corresponds to a set of sustainable utilization strategy of water resources. different transformation t corresponds to a different set of extension-transformation matter-elements and then the strategy is different. repeatedly doing transformation, we obtain a variety of sustainable utilization strategies. to doing extension transformation to extension-analysis matterelements is known as the sustainable utilization strategy generation of water resources. different strategies have different implementation effects in the process of sustainable utilization of water resources. the purpose of extension transformation is to find the best solutions as possible as can, so we need to select the optimal one in a variety of strategies. firstly, we should determine the evaluation characteristics, such as technological possibility and economic feasibility. secondly, we select an appropriate evaluation method to calculate the optimal degree of these matter-elements. finally, choosing the maximum in the optimal degrees and then viewing the corresponding set of matter-elements as the optimal strategy, we formulate the specific action plans according to the optimal strategy. the advances in systems science and application (2016) vol.16 no.1 35 action plan can achieve the sustainable utilization of water resources during the planning period t in the region a under the required conditions to meet the development of population, resources, environment and economics, and effectively resolve the contradictory problem between water crisis and ecological environment constructions, and then the overall goals can be achieved. the formalized generation process of sustainable utilization strategies of water resources can be shown in fig. 1. fig. 1 the flow chart of extension strategy 5 modified case study in order to illustrate the generation and optimal selection of sustainable utilization strategies of water resources and the program implementation process, we will utilize a specific case to describe it in this section. 1) the overview of socio-economics and water resources in the studied region 36 qiaoxing li and naiang wang babu:research on extension strategy generation ... one region a located in the middle of heihe river basin and the central place of hexi corridor is the political, economic and cultural center of zhangye city, gansu province, pr china. the total acreage is about 4000 square kilometers, where the propositions of mountainous, desert area and plain are 14.4%, 34.5% and 51.1%, respectively. the coverage rate of vegetation is only about 3%. the total population in the region is approximately 520,000, where the agricultural population is about 350,000 whose proposition is 65%. in recent years, gdp is about 6.8 hundred million rmb with about 15,000 rmb per person. the water which is utilized to maintain the basic life of population and the production in this region comes from the rivers flowing through the region and the groundwater. the region is a typical agriculture oasis and large irrigated agriculture area. the total volume of available water in the region is about 1.3 hundred million m3, and the average volumes per person and per acre are about 1330m3 and 570m3, respectively. so it is a relatively serious water shortage region. from the structure of water utilization, the ratio of agriculture, industry, people living and ecology is 83.7:2.7:6.4:7.2 in 2005, where the proportion of agricultural water is still relatively large. the region a is an important commodity grain base and one of the five bases named west vegetable to east in northwest of china, and thus the agricultural production plays an important role in the area. region a has superior condition with 30 kinds of mineral resources, where the reserve of coal has been proven for 1.05 hundred million tons, and the content of both tungsten and molybdenum ranks at the first order in north of china. according to the survey, region a has about 10 species of wild animals, where the total number is only 2,000 surplus and some of them are on the verge of extinction. the averages of annual rainfall and annual evaporation are 120mm and 2000-2350mm, respectively, so it is a typical temperate continental arid climate. there are many problems of water resources, such as serious water shortage, obvious contradiction between supply and demand, irrational utilization structure, low efficiency and effectiveness, thus the water resources is the most important constraints of sustainable development of society and economy in the region. 2) define contradictory problem and set goal basic-elements and condition matter-elements the local government of regional a viewed the eco-economic development as the main task. so the government transforms the economic development mode and highlights three goals which are ecological construction, modern agriculture and passage economy. then sustainable utilization of local water resources may be achieved. under the requirements to enhance the population quality, exploit the natural mineral resources, improve the living environment and promote economic development, local government makes its efforts to solve the two central goals of water crisis and ecological environment construction. then the govadvances in systems science and application (2016) vol.16 no.1 37 ernment develops a “five-year plan” according to local condition and obtains following total goal basic-elements g1 and g2: g1 = ( construct a1 a2 a3 a t b1 b2 b3 a t ) and g2 = ( w a c2 c3 t a 13 20 t ) , where the unit of c2 and c3 is hundred million m3, t=five years, and the meaning of others is the same with corresponding formers. because the value of c3 is larger than one of c2, the water crisis exists. then we develop the specific goal basic-elements and simplify the condition basic-elements according to g1 and g2 respectively. on the basis of the actual situation and total goals of region a, we decompose g1 and get following specific matterelements: g′ 1 =  g1 a11 b11 a12 120 a13 360 a14 0.9 a15 2.3  , g′′ 1 =  g2 a21 b21 a22 300 a23 240 a24 60 a25 3.1  and g′′′ 1 =  g3 a31 b31 a32 2000 a33 2300 a34 0 a35 0  , where the characteristics a11, · · · , a15, a21, · · · , a25, a31, · · · , a35 are the same as the formers, and the value of a13 =the value of a12+ the value of a23. because the species g3 usually live in vegetation g2 such as forest and grassland, its water requirement has been included in the vegetation water demand, so its characteristics a34 and a35 are 0. at the same time, a part of treated land g2 will be planted vegetation as woodland and grassland, and another part will be as construction land such as residential housing and factory building. so the value of a25 contains the value of a25. therefore, the total water increment of three specific targets is 3.1 hundred million m3. if it is together with the original demand 0.90 hundred million m3, the total water demand of target basic-elements is 4.0 hundred million m3. simplify the condition basic-elements l1, l2, l3 and l4 as follows: l′1 = r1 a1 r11 a2 r12 a3 r13  , l′2 = r2 a1 r21 a2 r22 a3 r23  , l′3 = r3 a1 r31 a2 r32 a3 r33  and l′4 = r4 a1 r41 a2 r42 a3 r43  . then we get the problem p′ = (g′ 1 ∧ g′′ 1 ∧ g′′′ 1 ∧ g2) ↑ (l′1 ∧ l′2 ∧ l′3 ∧ l′4)) . if the problem has been solved, then the total goal basic-elements g1 and g2 can be achieved. 38 qiaoxing li and naiang wang babu:research on extension strategy generation ... 3) extension analysis and extension transformation according to the process of extension analysis introduced above, we obtain the extensionanalysis matter-elements of condition basic-elements. firstly, we determine that the values of user k1 are w1 11 = rural residents, w2 11 = urban residents; w1 21 = land, w2 21 = minerals, w3 21 = biology; w1 31= green belt, w2 31 = waste treatment, w3 31 = transport belt. we also assume that there are four economic sectors in region a which are w1 41 = agriculture, w2 41 =livestock, w3 41 =building industry and w4 41 = catering industry, respectively; the values of other characteristics, such as use mode k2, water quality k3, importance k4, former user k5, later user k6, recyclability k7, emission k8, treatable technology k9, source k10 and freshness k11, are referred to the process of extension analysis in section 3. these characteristics only have important role for analyzing water utilization and need not do extension transformation. therefore, we list them separately as the innovation characteristics of extension analysis in table 1. the other characteristics of extension-analysis matter-elements involve the quantitative of water resources, such as use volume k12, recovered amount k13, increment k14, recycle volume k15 and total volume k16. doing extension transformation such as substitution, decomposition, increase, decrease, etc., to the quantitative of the five characteristics of extension-analysis matterelements, we get extension-transformation matter-elements with new value. suppose that some types of extension transformation, which are named as t1, t2, t3 and t4, respectively, have been done to the values of five quantitative characteristics of extension-analysis matter-elements, we get four sets of extension-transformation matter-elements, i.e., four types of innovation strategies of sustainable utilization of water resources, which can be seen in table 2 and table 3. for the transformation t1 in table 2, the value of use volume k12 is the original statistics of water resources in region a, and the value of increment k14 is water increases to meet the development needs of population, resource, environment and economics according to the current consumption structure of water, and the increases is from the fresh water. the value of recovered amount k13 is the recycled quantities from use volume k12 and increment k14 which have been consumed by corresponding users according to the current consumption structure. recycle volume k15 is the re-utilized amount of recovered amount k13. from table 2, we know that the water which is only consumed by minerals, livestock and catering industry has been retrieved, and part of the retrieved water has been utilized by waste treatment, agriculture and livestock. in table 2 and table 3, the transformations t2, t3 and t4 come from t1 in turns, and then produce three additional utilization strategies of water resources. transformations t2 and t3 only adopt the water-saving measures, namely that the corresponding users utilize the retrieved water at first for use volume k12 and increment k14 as posadvances in systems science and application (2016) vol.16 no.1 39 table 1 analysis table of extension innovation basic elememt user k1 use mode k2 water quality k3 importance k4 former user k5 later user k6 w1 1 rural residents drink, wash, rinse grade i-iii necessary, priority rural residents green belt, waste treatment, transportation belt, agriculture, animal w2 1 urban residents drink, wash, rinse grade i-iii necessary, priority urban residents green belt, waste treatment, transportation belt, agriculture, animal w1 2 land none grade i-v necessary agriculture none w2 2 minerals produce grade i-iv priority none green belt, waste treatment, transportation belt w3 2 biology drink grade i-iii priority none none w1 3 green belt irrigate grade i-v priority residents none w2 3 waste treatment rinse grade i-v necessary residents, transportation belt, minerals none w3 3 transportation belt irrigate, rinse grade i-v priority residents waste treatment w1 4 agriculture irrigate grade i-v necessary, decrease residents, animal land w2 4 animal drink, rinse grade i-iii necessary, decrease residents, catering agriculture w3 4 building produce grade i-v current situation, decrease decrease none waste treatment w4 4 catering drink, wash, rinse grade i-iii necessary, decrease none animal tabel 1 analysis table of extension innovation (continued) basic elememt recyclability k7 emission k8 treatable technology k9 source k10 freshness k11 w1 1 yes direct emission no surface water, groundwater, external water fresh water, recycled water w2 1 yes direct emission no surface water, groundwater, external water fresh water, recycled water w1 2 no non-emission no all water fresh water, recycled water w2 2 yes treatable emission high surface water, groundwater, external water fresh water w3 2 no non-emission no surface water fresh water w1 3 no non-emission no all water fresh water, recycled water w2 3 yes treatable emission high all water fresh water, recycled water w3 3 yes direct emission no all water fresh water, recycled water w1 4 no direct emission no all water fresh water, recycled water w2 4 yes direct emission no surface water, groundwater, external water fresh water w3 4 yes treatable emission normal surface water, groundwater, external water fresh water w4 4 yes direct emission no surface water, groundwater, external water fresh water 40 qiaoxing li and naiang wang babu:research on extension strategy generation ... table 2 the first innovation t1 and the second one t2 of extension-exchange matter-element (unit: hundred million m3 ) t1 k12 k13 k14 k15 k16 t2 k12 k13 k14 k15 k16 t1(w1 1) 0.53 1.2 1.73 t2(w1 1) 0.53 -1.4 0.8 0.4 1.73 t1(w2 1) 0.27 0.6 0.87 t2(w2 1) 0.27 -0.65 0.23 0.37 0.87 t1(w1 2) 0 0.06 0.06 t2(w1 2) 0 0 0.06 0.06 t1(w2 2) 0.25 -0.3 0.7 0.95 t2(w2 2) 0.25 -0.85 0.7 0.95 t1(w3 2) 0.01 0.02 0.03 t2(w3 2) 0.01 0.02 0.03 t1(w1 3) 0.03 0.05 0.08 t2(w1 3) 0.03 0 0.05 0.08 t1(w2 3) 0 0.6 0.32 0.92 t2(w2 3) 0 0 0.92 0.92 t1(w3 3) 0.01 0.03 0.04 t2(w3 3) 0.01 -0.01 0 0.03 0.04 t1(w1 4) 10.85 -1.53 0.1 9.42 t2(w1 4) 10.85 -2.73 -1.53 0.1 9.42 t1(w2 4) 0.58 -0.15 -0.08 0.2 0.7 t2(w2 4) 0.58 -0.6 -0.08 0.2 0.7 t1(w3 4) 0.9 -0.1 0.8 t2(w3 4) 0.9 -0.55 -0.1 0.8 t1(w4 4) 0.5 -0.2 -0.1 0.4 t2(w4 4) 0.5 -0.27 -0.1 0.4 total 13.93 -0.65 1.45 0.62 16 total 13.93 -7.06 -0.06 2.13 16 table 3 the first innovation t3 and t4 of extension-exchange matter-element (unit: hundred million m3 ) t3 k12 k13 k14 k15 k16 t2 k12 k13 k14 k15 k16 t3(w1 1) 0.53 -1.4 0.8 0.4 1.73 t4(w1 1) 0.53 -1.4 0.8 0.4 1.73 t3(w2 1) 0.27 -0.65 0.23 0.37 0.87 t4(w2 1) 0.27 -0.65 0.23 0.37 0.87 t3(w1 2) 0 0 0.06 0.06 t4(w1 2) 0 0 0.06 0.06 t3(w2 2) 0.25 -0.85 0.7 0.95 t4(w2 2) 0.25 -0.85 0.7 0.95 t3(w3 2) 0.01 0.02 0.03 t4(w3 2) 0.01 0.02 0.03 t3(w1 3) 0 0 0.08 0.08 t4(w1 3) 0 0 0.08 0.08 t3(w2 3) 0 0 0.92 0.92 t4(w2 3) 0 0 0.92 0.92 t3(w3 3) 0 -0.01 0 0.04 0.04 t4(w3 3) 0 -0.01 0 0.04 0.04 t3(w1 4) 9.95 -2.73 -1.53 1 9.42 t4(w1 4) 9.83 -2.73 -1.53 1 9.3 t3(w2 4) 0.18 -0.6 -0.08 0.6 0.7 t4(w2 4) 0.18 -0.6 -0.08 0.6 0.7 t3(w3 4) 0.9 -0.55 -0.1 0.8 t4(w3 4) 0.9 -0.55 -0.1 0.8 t3(w4 4) 0.5 -0.27 -0.1 0.4 t4(w4 4) 0.5 -0.27 -0.1 0.4 total 12.59 -7.06 -0.06 3.47 16 total 12.47 -7.06 -0.06 3.47 15.88 sible as they can. transformation t4 has adopted the water-saving technologies to agriculture on the basis of transformation t3. transformation t2 is substitution on the basis of transformation t1. firstly, if the value of later user k6 is not ”none”, then the corresponding user should take some measures to retrieve part of the utilized water, and then we get that the sum of recovered amount k13 is 7.06 hundred million m3. secondly, if the value of former user k5 is ”none”, then the value of corresponding increment k14 should be prior from recovered amount k13 that can be utilized, and we obtain the sum of recycle volume k15 is 2.13 hundred million m3. transformation t3 is a synthesis of substitution and decomposition on the basis of transformation t2, i.e. if the value of former user k5 is not ”none”, then part of consumption water of the corresponding use volume k12 is from recovered amount k13. therefore, the water resources of recycle volume k15 has been conadvances in systems science and application (2016) vol.16 no.1 41 sumed 3.47 hundred million m3 and the remaining part is 3.59 hundred million m3. apparently, the total consumption water of transformation t2 is the same as the one of transformation t3 which is 16 hundred million m3. adding the water consumed by ecological environment construction of target basic-elements, which is 4 hundred million m3, the total water consumed in region a is 20 hundred million m3. transformation t4 is a decrease on the basis of transformation t3, where the highest consumer of water resources agriculture adopts a certain water-saving technology to reduce the consumption of initial water. then, agricultural expends initial water 9.83 hundred million m3. transformation t4 consumed the sum of initial water 12.41 hundred million m3 and recycle water 3.47 hundred million m3, i.e., the actual amount of consumed water is 15.88 hundred million m3. 4) assessment and optimal selection of innovation strategies in order to evaluate these four innovation strategies, we establish a set of to measurement criterions: a1: technological feasibility; a2: economical feasibility and a3: to meet basic needs of regional water resources, where a1 and a2 are prerequisites, and a3 contains two meanings: first, the initial water utilized by target basic-elements is no less than 0.4 hundred million m3; second, the initial water consumed by condition basic-elements should meet the basic needs. furthermore, the first meaning is also prerequisite, which means that the initial water consumption of condition basic-elements is not more than 1.26 hundred million m3. in addition, because the total water demands of the target basic-elements is 4 hundred million m3, its recycle volume is not more than 3.6 hundred million m3. apparently, the four innovative strategies satisfy the conditions both a1 and a2, but t1 and t2 do not meet the first meaning of a3, so they can be deleted firstly. following, we select the optimal one between t3 and t4. in general, the pros and cons of extension strategies can be evaluated by using priority degree evaluation method, comprehensive evaluation method, or others. however, in this case, speaking on the water consumption of the two innovative strategies t3 and t4, only agriculture is different from each other. therefore, the optimal strategy selection is only to compare the costs of between 0.12 hundred million m3 water resources that t3 is surplus to t4 and the amount of saving water by doing the technological reform of t4. since the study area is located in the arid and semi-arid regions, the sustainable utilization of water resources occupies the prime location in natural, economic and social systems. therefore, the technological reform to save water in this region is the main goal of future work. then we choose t4 as the optimal strategy. 5) draw up the action programs according to table 1 and the strategy t4 of table 3, we formulate the specific 42 qiaoxing li and naiang wang babu:research on extension strategy generation ... five-year action programs of water resources in the study area below: (1) residents are limited the total amount of 1.83 hundred million m3 of water utilization, where the quotas of rural residents and urban residents are 1.33 and 0.5 hundred million m3, respectively. under the situation of water ratio systems, we should retrieve the wastewater. the amount of recycled water from residents is 2.05 hundred million m3, of which 0.77 hundred million m3 is utilized as non-drinking, such as washing (courtyard, toilet), etc. on the other hand, 1.28 hundred million m3 of water resources will be allocated to green belt, waste treatment, transportation belt, agriculture and livestock. (2) land should strictly limit their consumption of initial water, and only utilizes about 0.06 hundred million m3 of recycle water from agriculture. the retrieved water will be utilized to conserve water, prevent land from desertification and semi-desertification. (3) rationally develop mineral resources. within the 30 plus kinds of mineral resources in the studied area, we preferentially choose coal, tungsten and molybdenum, which have great development potentiality and mature technology, to be exploited. all water to develop mineral resources is of 0.95 hundred million m3, and the enterprises should adopt retrieved measures to collect wastewater about 0.85 hundred million m3, which will be utilized for green belt, waste treatment and transport belt. (4) protect the wildlife habitat and properly feed wild animals. the program will be carried out simultaneously with the ecological and environmental protection. at the same time, an additional water of 0.03 hundred million m3 is supplied to supplement wildlife habitat. (5) green urban and rural regions. in the towns and villages surrounding residential areas, plenty of trees and flowers should be planted. all required water of about 0.08 hundred million m3 should be from recycled water provided by the residents, and initial water is strictly prohibited for irrigation. (6) purify the living environment and let it be a healthy region, and vigorously clear up waste water, waste gas and waste disposal. the water of 0.92 hundred million m3 utilized to improve the health of towns and villages is provided by residents, transportation belt and minerals, and initial water is prohibited to purify the living environment. (7) protect and reinforce roads. shelterbelts will be planted and garbage should be collected on both sides along the roads. the water of 0.04 hundred million m3 to plant trees and clean roads is entirely provided by recovered water from residents. at the same time, the water which was secondly recovered by using water-saving measures will be utilized to clear up waste water, waste gas and waste disposal. (8) supported by scientific and technological progress and oriented by market, advances in systems science and application (2016) vol.16 no.1 43 the structure of agricultural production will be adjusted greatly. firstly, on the basis of ensuring the basic supply of grains, vegetables, fruits as well as other agricultural products, we should vigorously adjust the agriculture structure and develop water-saving agriculture; secondly, we may adopt water-saving technologies and measures on agricultural production in order to implement watersaving and recycling irrigations. the total of agricultural water would be less than 9.30 hundred million m3, where recycle volume is 1 hundred million m3 and initial water is not more than 8.30 hundred million m3. initial water has been saved 2.55 hundred million m3 than before, and recovered amount is 2.73 hundred million m3, where a small part was directly discharged to land and the majority is utilized for planting new vegetations. (9) appropriately reduce the development scale of livestock in order to save water resources and protect vegetation. arrange livestock for initial water of 0.10 hundred million m3 and recycle volume of 0.60 hundred million m3, where recycle water comes from recovered amount of residents and catering industry, and recovered water of 0.60 hundred million m3 from livestock by using watersaving measures is provided for agriculture. (10) control the development scale of building industry and pay attention to retrieve its utilized water. since building industry must utilize the initial water, we should appropriately control its scale of development. the water consumption amount of building industry was reduced from the original 0.90 hundred million m3 to 0.80 hundred million m3, and about 0.55 hundred million m3 of retrieved water by adopting saving-water measures is utilized for waste treatment. (11) maintain and reduce water consumption amount of catering industry. catering water should consume initial water, so we may pay attention to water conservation and against waste. the quota of initial water was reduced to 0.4 hundred million m3 and 0.27 hundred million m3 of water is retrieved for livestock by improving water-saving measurements. in order to effectively perform the schemes above, the local government needs to formulate the following supportable measures: (1) promote the establishment of water-saving society and encourage broad participation of resident. since the programs need to use a lot of recycle volume, residents and other water users must actively cooperate to retrieve the utilized water. full of recovered amount is a necessary condition for sustainable utilization of water resources. (2) develop relevant laws and regulations, and constraint the user behavior. under the rigid constraints to limit amount of available water resources, the integrated management of water resources should be comprehensively improved. regulatory reforms need to be clear the ways and patterns of water utilization, and determine the quota, and clear water rights, and implement the total control 44 qiaoxing li and naiang wang babu:research on extension strategy generation ... and quota management. (3) strengthen water conservation measures and transform water-saving technology, and vigorously popularize the water-saving infrastructure projects. by laying recovery line, we can provide the necessary conditions for residents to retrieve and re-utilize water resources. at the same time, by adopting water-saving technologies, such as the drip and the timing spray irrigations, we can improve water efficiency and effectiveness. (4) optimize the agricultural structure and enhance the development level of agricultural economics. agriculture is the largest water user and its effectiveness is also the lowest. to adjust the farming structure and enhance the water-saving technology is a sufficient condition of sustainable utilization of water resources. 6 conclutions the paper provided a formal approach to generate strategies for the sustainable utilization of regional water resources in a planning period. under the condition to maintain the development of regional population, resources, environment and economy, it can help the relevant departments to clarify their ideas and effectively resolve the contradictory problem between water crisis and ecological environment construction. in real applications, the relevant departments can adjust the contents and values of characteristics of the condition basic-elements. the extension analysis result of water utilization is an open thinking way. it can make the decision-making of relevant departments be more standardized and scientific. the approach is also compiled into software that enables the decisionmaking process be intelligent. the extension generation strategy of sustainable utilization of regional water resources alters the only method of qualitative analysis in the past to the perfect combination qualitative analysis with quantitative calculation. this method is also suitable for the strategy generation of sustainable utilization of river basin. references [1] icwe. (1992), “international conference on water and the environment: development issues for 21 century”, the dublin statement and report of the conference, dublin. [2] wei-hua zeng, zhi-feng yang and gen-suo jia. (2006), “integrated management of water resources in river basins in china”, aquatic ecosystem health & management, vol.9, no.3 , pp. 327-332. [3] hong yang and alexander zehnder. (2007), “ ‘virtual water’: an unfolding concept in integrated water resources management”, water resources research, vol.43, no.w12301. advances in systems science and application (2016) vol.16 no.1 45 [4] jordi gallego-ayala and dinis juzo. (2011), “strategic implementation of integrated water resources management in mozambique: an awot analysis”, physics and chemistry of the earth, vol.36, pp. 1103-1111. [5] brian r. cook and christopher j. spray. (2012), “ecosystem services and integrated water resource management: different paths to the same end”, journal of environmental management, vol.109, pp. 93-102. [6] shuang liu, neville d. crossman, martin nolan and hiyoba ghirmay. (2013), “bringing ecosystem services into integrated water resources management”, journal of environmental management, vol.129, pp. 92-102. [7] a. a. r.ioris, c. hunter and s. walker. (2008), “the development and application of water management sustainability indicators in brazil and scotland [j]”, journal of environmental management, vol.88, no.4, pp. 1190-1201. [8] changhao liu, kai zhang and jiaming zhang. (2010), “sustainable utilization of regional water resources: experiences from the hai hua ecological industry pilot zone (hheipz) project in china”, journal of cleaner production, vol.18, pp. 447-453. [9] tian aimin, jiang feng, dong ning, tian aijie and jiang anxi. (2010), “research on the sustainable utilization of water resources of jinan city”, chinese journal of population resources and environment, vol.8, no.4, pp. 55-60. [10] n. mahjouri and m. ardestani. (2010), “a game theoretic approach for inter basin water resources allocation considering the water quality issues”, environment monitor assess, vol.167, pp. 527-544. [11] m. sadegh and r. kerachian. (2011), “water resources allocation using solution concepts of fuzzy cooperative games: fuzzy least core and fuzzy weak least core”, water resources management, vol.25, no.10, pp. 2543-2573. [12] francisco assis souza filho, upmanu lall and rubem la laina porto. (2008), “role of price and enforcement in water allocation: insights from game theory”, water resources research, vol.44, no.w12420. [13] y.p. li, g.h. huang, y.f. huang and h.d. zhou. (2009), “a multistage fuzzy-stochastic programming model for supporting sustainable waterresources allocation and management”, environmental modelling & software, vol.24, pp. 786-797. 46 qiaoxing li and naiang wang babu:research on extension strategy generation ... [14] han mei, duhuan, yangxiaoyan and liuyuan. (2010), “research advances on water resources optimal distribution”, procedia environmental sciences, vol.2, pp. 1912-1918. [15] bossel h. (2000), “the human actor in ecological-economic models: policy assessment and simulation of actor orientation for sustainable development”, ecological economics, vol.34, pp. 337-355. [16] zhai jin-liang, feng ren-guo and xia jun. (2011), “constraining factors to sustainable utilization of water resources and their countermeasures in china”, chinese geographical science, vol.13, no.4, pp.310-316. [17] yang chunyan and cai wen. (2013), extenics: theory, method and application, science press, beijing. [18] pan hulin. (2009), “the evaluation of iwrm performance and the analysis of its influential factors in the arid area: a case study on the iwrm of ganzhouqu district”, northwest normal university, lanzhou. corresponding author qiaoxing li can be contacted at: gxqxli@163.com and liqx@lzu.edu.cn characterization of lp c -solutions for the dilation equations on r2 chun-tai liu1 and guo-tai deng2 1department of mathematics and physics, wuhan polytechnic university, wuhan 430023, p. r. china 2college of mathematics and statistics, huazhong normal university, wuhan 430079, p.r. china email: lct984@163.com, hilltower@163.com abstract in this paper, the author discussed the existence of compactly supported lp-solutions for the dilation equations on the plane. furthermore, two examples are given to illustrate the general theory. keywords dilation equation compactly supportedlp-solutions iteration function system 1. introduction a α-scale dilation equation is a functional equation of the form f(x) = n∑ n=0 cnf(αx− βn) where f : r → r(or c), α > 1, β0 < β1 < · · · < βn are real constants, and cn are real (complex) constants. the equation is called a lattice k-scale dilation equation if f(x) = n∑ n=0 cnf(kx− n) for an integer k ≥ 2. a special case of the functional equation (k = 3, n = 4, and cn = 1, 2/3, 1/3, 1) was first studied by de rham[1] as an example of a continuous nowhere differentiable function. recently this equation has attracted a lot of attention, especially for the lattice case with k = 2. in wavelet theory, the study of multiresolution and the search of various orthogonal, compactly supported wavelets has lead to the investigation of the existence, uniqueness, and smoothness of such continuous integrable solutions[2]. the equation also plays an important role in the “subdivision schemes” and “interpolation schemes” of constructing continuous spline curves, surfaces and fractal objects [3, 4] . there are two major approaches to the equation: the fourier method(the frequency domain approaches) and the iteration method(the time-domain approaches). using fourier transformation, daubechies and lagarias[3] proved that the equation has a nonzero integrable solution. by using the fourier transform of f and the paley-wiener theorem, it was proved in [3] that f has compact support in [0, βn/(α− 1)]. the fourier method, however, does not give sharp criteria for the existence of l1 − solutions in terms of the coefficients{cn}. some partial results are given in [5, 6]. the iteration method is restricted to the lattice case. it applies particularly well in the case of compactly supported solutions. the basic idea is to identify a given function f supported by issn 1078-6236 international institute for general systems studies, inc. advances in systems science and applications (2011), vol. 11, no. 1-2 61-70 [0, n ] with the vector-valued function f(x) = [f(x), · · · , f(x+ (n − 1))]t , x ∈ [0, 1], and to use the right side of the dilation equation to construct two n × n matrices t0 and t1. a constant vector v is used as the initial condition, followed by iteration with the matrices t0 and t1. the limit, if the sequence converges, will be the solution of the dilation equation. such an approach was used by daubechies and lagarias[4], and independently by michelli and prautzsch[7]. it was also used by collela and heil[8] and [9] and ka-sing lau and jianrong wang[10]. similarly, on the plane the dilation equation is defined as the form f(x) = m∑ m=0 n∑ n=0 cmnf(ax−qmn) (1) where f : r2 → r or c, a is an expand matrix, qmn are vectors, cmn are real (or complex) constants. the equation is called a lattice dilation equation if f(x) = m∑ m=0 n∑ n=0 cmnf(ax− ( m n ) ) (2) where a is an integer expand matrix. in this paper we will study the existence of the compactly supported lp-solutionof the equation(2) on the plane with cmn ∈ r and a = ( 2 0 0 2 ) . as usually the basic assumption on the coefficients is m∑ m=0 n∑ n=0 cmn = | deta| = 4. let d1 = ( 0 0 ) , d2 = ( 0 1 ) , d3 = ( 1 0 ) , d4 = ( 1 1 ) , ϕk(x) = a−1(x + dk), k = 1, 2, 3, 4, there exists an attractor t = [0, 1]× [0, 1] satisfing t = 4⋃ k=1 ϕk(t ) , 4⋃ k=1 tk. at the same time there exist vectors {eis = ( i s ) , 0 ≤ i ≤m − 1, 0 ≤ s ≤ n − 1} such that suppf ⊂ m−1⋃ i=0 n−1⋃ s=0 (t + eis). let pi0 = (ci,2s−t)0≤s, t≤n−1, pi1 = (ci,2s−t+1)0≤s, t≤n−1, i = 0, 1, · · · ,m. for example, p00 = (c0,2s−t)0≤s, t≤n−1 =  c0,0 c0,2 c0,1 c0,0 c0,4 c0,3 c0,2 c0,1 c0,0 · · · · · · · · · · · · · · · · · · 0 0 0 0 0 · · · c0,n c0,n−1  . 62 liu:characterization of lp c -solutions for the dilation equations on r2 and let m1 = (p2i−j,0)0≤i,j≤m−1 =  p0,0 p2,0 p1,0 p0,0 p4,0 p3,0 p2,0 p1,0 p0,0 · · · · · · · · · · · · · · · · · · 0 0 0 0 0 · · · pm,0 pm−1,0  , m2 = (p2i−j,1)0≤i,j≤m−1 =  p0,1 p2,1 p1,1 p0,1 p4,1 p3,1 p2,1 p1,1 p0,1 · · · · · · · · · · · · · · · · · · 0 0 0 0 0 · · · pm,1 pm−1,1  , m3 = (p2i−j+1,0)0≤i,j≤m−1 =  p1,0 p0,0 p3,0 p2,0 p1,0 p0,0 p5,0 p4,0 p3,0 p2,0 p1,0 p0,0 · · · · · · · · · · · · · · · · · · · · · 0 0 0 0 0 0 · · · 0 pm,0  , m4 = (p2i−j+1,1)0≤i,j≤m−1 =  p1,1 p0,1 p3,1 p2,1 p1,1 p0,1 p5,1 p4,1 p3,1 p2,1 p1,1 p0,1 · · · · · · · · · · · · · · · · · · · · · 0 0 0 0 0 0 · · · 0 pm,1  we define a vector function: f (x) = (f(x+e00), f(x+e01), f(x+e02), · · · , f(x+e0,n−1), f(x+ e10), · · · , f(x + e1,n−1), · · · , f(x + em−1,n−1)) t for x ∈ t = [0, 1] × [0, 1], then equation (2) will satisfy f (x) =  m1f (ϕ−11 (x)) x ∈ t1 = [0, 1/2)× [0, 1/2); m2f (ϕ−12 (x)) x ∈ t2 = [0, 1/2)× [1/2, 1); m3f (ϕ−13 (x)) x ∈ t3 = [1/2, 1)× [0, 1/2); m4f (ϕ−14 (x)) x ∈ t4 = [1/2, 1)× [1/2, 1); 0 x ∈ others. (3) let v is 4-eigenvector of (m1 + m2 + m3 + m4),we have (m1 + m3 − 2i)v = −(m2 + m4 − 2i)v.and let ṽ = (m1 + m3 − 2i)v, h(ṽ) be the subspace in rm×n spanned by {mσṽ : σ ∈ σ∗}. then the basic theorem is as follows. theorem 1.1. for 1 ≤ p ≤ ∞, the following are equivalent: (1) equation (2) has a nonzero compactly supportedlp-solution; (2) there exists a 4-eigenvector v of (m1 +m2 +m3 +m4) satisfying lim l→∞ 1 4l ∑ |σ|=l ‖mσṽ‖p = 0 advances in systems science and applications (2011), vol. 11, no. 1-2 63 (3) there exists a 4-eigenvector v of (m1+m2+m3+m4) such that there exists an integer l ≥ 1 such that 1 4l ∑ |σ|=l ‖mσu‖p < 1 for all u ∈ h(ṽ), ‖u‖ ≤ 1 2. preliminaries lemma 2.1. if equation(2) exists compactly supportedlp-solutionf , then suppf ⊂ [0,m ]× [0, n ]. proof let suppf ⊂ d, take x ∈ d with f(x) 6= 0 thenax− ( m n ) ∈ d, i.e. x ∈ a−1(d+ ( m n ) ). let e = {0, 1, 2 · · ·m} × {0, 1, 2 · · ·n}, then d ⊂ a−1(d + e) = a−1d +a−1e ⊂ a−1(a−1d +a−1e) +a−1e = a−2d +a−2e +a−1e ⊂ · · · ⊂ a−td +a−te +a−(t−1)e + · · ·+a−1e let t→∞, then d ⊂ { ∞∑ t=1 a−ty : y ∈ e} ⊂ [0,m ]× [0, n ] for the closed set e. proposition 2.2. let f be supported by [0,m ]× [0, n ], and let f be defined as above, then f is an lpc−solution of (2) if and only if f ∈ lp and f = mf , i.e. f satisfies equation(3). proposition 2.3. if m∑ m=0 n∑ n=0 cmn = 4, then 4 is an eigenvalue of (m1 +m2 +m3 +m4) with left eigenvalue [1, 1, · · · , 1]. proof obviously, the sum of each column is equal to 4 in the matrix (m1 +m2 +m3 +m4). � it follows that the right 4-eigenvector of (m1 +m2 +m3 +m4) exists also; it will play a central role in the existence of the solution of equation(2). let f4 be the average of f over 4, i.e., f4 = 1 l(4) ∫ 4 f . proposition 2.4. let f be an compactly supported lp-solutionof equation(2), v = [ft+e00 , ft+e01 · · · ft+em−1,n−1 ]t be the vector defined by the average of f on the m ×n subintervals as indicated. then v is 4-eigenvector of (m1 +m2 +m3 +m4). proof according to proposition2.2, f = mf , i.e., f (x) =  m1f (ϕ−11 (x)) x ∈ t1 = [0, 1/2)× [0, 1/2) m2f (ϕ−12 (x)) x ∈ t2 = [0, 1/2)× [1/2, 1) m3f (ϕ−13 (x)) x ∈ t3 = [1/2, 1)× [0, 1/2) m4f (ϕ−14 (x)) x ∈ t4 = [1/2, 1)× [1/2, 1) (4) 64 liu:characterization of lp c -solutions for the dilation equations on r2 when we integrate the expression over t1, t2, t3 and t4 separately, we have [f[0, 1 2 ]×[0, 1 2 ], · · · , f[m−1,n− 1 2 ]×[m−1,n− 1 2 ]] t = m1v [f[0, 1 2 ]×[ 1 2 ,1], · · · , f[m−1,n− 1 2 ]×[m− 1 2 ,n ]] t = m2v [f[ 1 2 ,1]×[0, 1 2 ], · · · , f[m− 1 2 ,n ]×[m−1,n− 1 2 ]] t = m3v [f[ 1 2 ,1]×[ 1 2 ,1], · · · , f[m− 1 2 ,n ]×[m− 1 2 ,n ]] t = m4v on the other hand, note that on each interval [i, i+ 1]× [i, i+ 1] the average satisfies f[i,i+ 1 2 ]×[i,i+ 1 2 ] +f[i,i+ 1 2 ]×[i+ 1 2 ,i+1] +f[i+ 1 2 ,i+1]×[i,i+ 1 2 ] +f[i+ 1 2 ,i+1]×[i+ 1 2 ,i+1] = 4f[i,i+1]×[i,i+1], hence we conclude that (m1 +m2 +m3 +m4)v = 4v. let σ = {1, 2, 3, 4}, σn = {(i1, i2, · · · , in) : ij ∈ σ}, σ0 = ∅, σ∗ = ∞⋃ n=0 σn, σ∞ = {(i1, i2, · · · ) : ij ∈ σ}. for each σ = (i1, i2, · · · ) ∈ σ∞, define σ|n = (i1, i2, · · · , in). let σ = (i1, i2, · · · , in) ∈ σ∗, τ = (j1, j2, · · · , jm) ∈ σ∗, define (σ, τ) := (i1, i2, · · · , in, j1, j2, · · · , jm).tσ := 4⋃ i=1 t(σ, i), and mσ := mi1mi2 · · ·min . so for any σ, τ ∈ σ∗, we have t(σ, τ) ⊂ tσ. lemma 2.5. let f0(x) = v for x ∈ t , and fk+1 = mfk for k ≥ 0. then fk = mσv for each x ∈ tσ. moreover, if f is an lpc −solution of equation(2) and v is the average vector of f defined in proposition 2.4, then fk = mσv = [ftσ+e00 , ftσ+e01 · · · ftσ+em−1,n−1 ]t , where (tσ + j) = {x+ j : x ∈ tσ}. also, fk → f in lp(t,rm×n ). proof we will use induction to show that fk(x) = mσv for x ∈ tσ with |σ| = k. suppose that fk(x) = mσv for x ∈ tσ. let x ∈ t(1,σ) = ϕ1(tσ); then ϕ−11 (x) ∈ tσ and fk+1(x) = mfk(x) = m1fk(ϕ −1 1 (x)) = m1mσv = m(1,σ)v. similarly, if x ∈ t(i,σ), then fk+1(x) = m(i,σ)v, i = 2, 3, 4. moreover, f = mf and f (x) = mσf (ϕ−1σ (x)) for x ∈ tσ. integrating this over the interval tσ, we obtain [ftσ+e00 , ftσ+e01 , · · · , ftσ+em−1,n−1 ]t = mσv. lemma 2.6. let v be a 4-eigenvector of (m1 + m2 + m3 + m4), and let fk be defined as above;then for each k, ∫ t fk(x)dx = v. (5) advances in systems science and applications (2011), vol. 11, no. 1-2 65 proof this follows from the following induction argument:∫ t fk+1 dx = ∫ t1 m1fk(ϕ −1 1 (x)) dx+ ∫ t2 m2fk(ϕ −1 2 (x)) dx + ∫ t3 m3fk(ϕ −1 3 (x)) dx+ ∫ t4 m4fk(ϕ −1 4 (x)) dx = 1 4 ( m1 ∫ t fk(x) dx+m2 ∫ t fk(x) dx+m3 ∫ t fk(x) dx+m4 ∫ t fk(x) dx ) = 1 4 ( m1 +m2 +m3 +m4 )∫ t fk(x) dx = 1 4 (m1 +m2 +m3 +m4)v = v. theorem 2.7. for 1 ≤ p ≤ ∞, the following are equivalent: (1) equation (2) has a nonzero compactly supportedlp-solution; (2) there exists a 4-eigenvector v of (m1 +m2 +m3 +m4) satisfying lim l→∞ 1 4l ∑ |σ|=l ‖mσṽ‖p = 0 (3) there exists a 4-eigenvector v of (m1 +m2 +m3 +m4) such that there exists an integer l ≥ 1 such that 1 4l ∑ |σ|=l ‖mσu‖p < 1 for all u ∈ h(ṽ), ‖u‖ ≤ 1 (6) proof let f0 = v and fn+1 = mfn. by lemma (2.5), for x ∈ tσ and |σ| = n, fn(x) = mσv. let gn = fn+1 − fn; then fn+1 = f0 +g0 + · · ·+gn, where gn(x) = { m(σ,1)v +m(σ,3)v − 2mσv = mσṽ if x ∈ t(σ,1) ∪ t(σ,3), m(σ,2)v +m(σ,4)v − 2mσv = −mσṽ if x ∈ t(σ,2) ∪ t(σ,4), and ‖gn‖p = 1 4n ∑ |σ|=n ‖mσṽ‖p. since (1) implies that ‖gn‖ converges to zero, (2) follows immediately. to prove that (2) implies (3), we note that h(ṽ) is finite dimensional and has a finite basis of mτ ṽ’s. let u = mτ ṽ with |τ | = k; then 1 4n ∑ |σ|=n ‖mσu‖p = 1 4n ∑ |σ|=n ‖mσmτ ṽ‖p ≤ 4k 1 4n+k ∑ |σ|=n+k ‖mσṽ‖p → 0 as n → ∞, and the convergence is uniform for all ‖u‖ ≤ 1. hence (6) follows by taking l = n for n sufficiently large. 66 liu:characterization of lp c -solutions for the dilation equations on r2 now assume (3) holds. sinceh(ṽ) is finite dimensional, there is a constant 0 < c < 1 such that for any u ∈ h(ṽ), 1 4l ∑ |σ|=l ‖mσu‖p ≤ c‖u‖p. for any |τ | = n, let u = mτ ṽ ∈ h(ṽ); then 1 4l ∑ |σ|=l ‖mσmτ ṽ‖p ≤ c‖mτ ṽ‖p. summing over all |τ | = n, we have 1 4l+n ∑ |σ|=l+n ‖mσṽ‖p = 1 4l+n ∑ |σ|=l ∑ |τ |=n ‖mσmτ ṽ‖p < c 4n ∑ |τ |=n ‖mτ ṽ‖p. it follows from the expression of ‖gn‖ given above that ‖gn+l‖p ≤ c‖gn‖p. for each fixed n, {‖gn+kl‖}∞k=1 is dominated by a geometric series, hence fn+1 = f0 +g0 + · · ·+gn converges in lp. the limit f is nonzero by lemma (2.6), and so by proposition (5), (1) follows. remark 2.8. we can also consider the equation g(x) = m∑ m=0 n∑ n=0 dmng(bx− ( m n ) ) with b = ( 1 1 1 −1 ) and m∑ m=0 n∑ n=0 dmn = |detb| = 2, since iterating this equation again, we obtain the equation (2). corollary 2.9. under the same hypotheses of theorem (2.7), assume that the solution f exists; then v /∈ h(ṽ), and the dimension of h(ṽ) is at most mn − 1. proof by theorem (2.7)(2), 1 4n ∑ |σ|=n ‖mσu‖p → 0 for any u ∈ h(ṽ). it follows that if v ∈ h(ṽ), then ‖v‖p = 1 4np ‖(m1 +m2 +m3 +m4) nv‖p ≤ 1 4n ∑ |σ|=n ‖mσv‖p → 0 as n→∞. this contradicts v 6= 0. advances in systems science and applications (2011), vol. 11, no. 1-2 67 3. some examples example 1: we consider a dilation equation : f(x) = 1∑ m=0 2∑ n=0 cmnf(ax− ( m n ) ) (7) where a = ( 2 0 0 2 ) , 1∑ m=0 2∑ n=0 cmn = 4. theorem 3.1. for 1 ≤ p <∞, equation(7) has a (nonzero)lpc−solution if and only if either c01 + c11 = 2 and 1 4 (|c00 + c10|p + |2− (c00 + c10)|p) < 1 or c00 + c10 = 2 and c02 + c12 = 2. proof note that m1 = ( c00 0 c02 c01 ) ,m2 = ( c01 c00 0 c02 ) ,m3 = ( c10 0 c12 c11 ) ,m4 = ( c11 c10 0 c12 ) , and m1 +m2 +m3 +m4 = ( c00 + c10 + c01 + c11 c00 + c10 c02 + c12 c01 + c11 + c02 + c12 ) . if (c00 + c10, c02 + c12) = (0, 0), then m1 + m2 + m3 + m4 = 4i . any nonzero vector v = [x, y]t will be a 4-eigenvector. it is a direct calculation that v ∈ h(ṽ) and, by corollary(2.9), no nonzero lpc−solution exists. we assume that (c00 + c10, c02 + c12) 6= (0, 0), then 4-eigenvector of m1 +m2 +m3 +m4 is v = [c00 + c10, c02 + c12] t , so that ṽ = (m1 +m3 − 2i)v = ( (c00 + c10)(c00 + c10 − 2) (c02 + c12)(2− (c02 + c12)) ) . (8) for a lpc−solution exists, h(ṽ) can only be {0} or one-dimensional(corollary(2.9)). in the first case, ṽ = 0, condition(6) is automatically satisfied. the only possible cases are (c00 + c10, c02 + c12) = {(2, 2), (0, 2), (2, 0)}. in the second case, ṽ 6= 0. since h(ṽ) is invariant under m1 +m3 and m2 +m4, (m1 + m3)ṽ = cṽ for some c. expression (8) yields the following cases: (a) c00 + c10 = 0 or c02 + c12 = 0. in this case v ∈ h(ṽ) and corollary(2.9) implies that (7) has no lpc−solution. (b) c00 + c10 = 2 or c02 + c12 = 2. in this case a direct calculation shows that (m1 + m3)ṽ, (m2 + m4)ṽ are independent. since h(ṽ) is 2-dimensional and by corollary(2.9) no lpc−solution exists. 68 liu:characterization of lp c -solutions for the dilation equations on r2 (c) c00 +c10 6= 0, 2 and c02 +c12 6= 0, 2. let a = c00 +c10, b = c02 +c12. by equating (8) and (m1 +m3)ṽ = ( a2(a− 2) b[(2− b)(4− a− b) + (a− 2)a] ) (9) with (m1 +m3)ṽ = cṽ, we have c = a, so that by (8) and (9), b[(2− b)(4− a− b) + (a− 2)a] = ab(2− b); that is (a+ b− 4)(a+ b− 2) = 0. hence, either (i) or (ii) below holds. (i) a+ b = 4. in this case v = [a, 4− a]t and ṽ = (a− 2)v. once again v ∈ h(ṽ) and no lpc−solution exists. (ii) a+b = 2. in this case a direct calculation shows that (m1+m3)ṽ = aṽ, (m2+m4)ṽ = bṽ. by theorem(6), equation(7) has an lpc−solution if and only if there exists an integer l ≥ 1 such that 1 4l (|a|p + |b|p)l‖ṽ‖p = 1 4l ∑ |σ|=l ‖mσṽ‖p < ‖ṽ‖p. this is equivalent to 1 4 (|a|p + |2− a|p) < 1, i.e., 1 4 (|c00 + c10|p + |2− (c00 + c10)|p) < 1. the theorem follows by summarizing all the cases. it follows directly from the theorem that if a+ b = 2 and if (i) a ∈ (−1, 3), then an l1 c − solution exists; (ii) a ∈ (0, 2), then an l2 c − solution exists; (iii) a = 1, then an lpc − solution exists for all 1 6 p <∞. example 2: considering a dilation equation as follows: g(x) = 1∑ m=0 1∑ n=0 dmng(bx− ( m n ) ) (10) where b = ( 1 1 1 −1 ) and 1∑ m=0 1∑ n=0 dmn = 2.. iterating the equation once, we obtain the suppg ⊂ [−1, 2]× [0, 3]. let f(x) = g(x− ( 1 0 ) ), a = b2 = ( 2 0 0 2 ) , and let (cmn)0≤m,n≤3 =  0 d00d10 d01d10 0 d200 d00d01 + d210 d00d11 + d10d11 d01d11 d00d10 d00d11 + d00d01 d201 + d10d11 d211 0 d01d10 d01d11 0  advances in systems science and applications (2011), vol. 11, no. 1-2 69 then (10) is rewritten as the form f(x) = 3∑ m=0 3∑ n=0 cmnf(ax− ( m n ) ) (11) with suppf ⊂ [0, 3]× [0, 3], so the discussion of the equation (11) is similar with the example 1. references [1] rham, g.de.. sur un example de fonction continue sans dérivée. enseign. math., 3(1957) 71-72. [2] heil,c. mathods of solving dilation equations, proc.1991 nato adv.sci.ins.on prob. and stoch. methods in anal.with appl.,j. byrnes,ed.,kluwer academic publishers,dordrecht, 1993. [3] daubechies,i. & lagarias j. two-scale difference equation i.existence and global regulatity of solutions. siam j. math.anal., 22(1991) 1388-1410. [4] daubechies,i. & lagarias j. two-scale difference equation ii. local regulatity, infinite products of matrices,and fractals, siam j. math.anal., 23(1992) 1031-1079. [5] lawton, w. necessary and sufficient conditions for constructing orthonormal wavelet bases, j.math.phys., 32(1991) 57-64. [6] mallalt, s. multiresolution approximation and wavelet orthonormal bases for l2(r), trans. amer. math. soc., 315(1989) 69-87. [7] michelli, c.a. & prautzsch,h. uniform refinement of curves, linear algebra appl., 114/115(1989) 841-870. [8] colella, d. & heil,c. the characterization of continuous,four-coefficient scaling functions and wavelets, ieee.trans.inform.theory, 30(1992) 876-881. [9] colella, d. & heil,c. characterizations of scaling functions, i.continuous solutions,j.math.anal.appl., 15(1994) 496-518. [10] lau k.l. & wang,j.r. characterization of lp−solutions for the two-scale dilation equations, siam j.math.anal., 26(1995),no.4 1018-1046. 70 liu:characterization of lp c -solutions for the dilation equations on r2 advances in systems science and applications (2014) vol.14 no.2 144-157 disease universe: visualisation of population-wide disease-wide associations max moldovan1,2, ruslan enikeev3, shabbir syed-abdul4, phung anh nguyen4, yo-cheng chang4 and yu-chuan li4 1australian institute of health innovation, university of new south wales, level 1 agsm building, sydney nsw 2052, australia 2school of population health, south australian health & medical research institute (sahmri), university of south australia, north terrace, adelaide sa 5000, ausralia 3the apac sale group, 3 petain rd., #05-07, singapore 208108 4graduate institute of medical informatics, college of medical science and technology, taipei medical university, taipei, taiwan abstract over a lifespan, a human organism is affected by multiple disorders of different origin and severity. we apply a force-directed spring embedding graph layout approach to electronic health records in order to visualise population-wide associations among human disorders as presented in an individual biological organism. the introduced visualisation is implemented on the basis of the google maps platform and can be found at http://disease-map.net. we argue that the suggested method of visualisation can both validate already known specifics of associations among disorders and identify novel, never noticed association patterns. keywords systems biology; phonemics; electronic health records; visualisation; graph layout 1 introduction it is known that many human disorders are positively associated, accompanying each other due to various, often unknown, genetic, bio-pathological or common risk factors [?]. there is also evidence that some disorders tend to be associated negatively, playing a preventative role against each other, or due to other hypothesised but not properly understood reasons [2-3]. we use population-wide electronic health records data to visualise how human disorders are positioned against each other in a population with respect to an individual biological organism. by doing so, we attempt to execute a systems biology approach in order to reveal the presence of common functional mechanisms influencing pathogenesis behind groups of human disorders through biological, epidemiological or environmental factors. it is important to note that, due to the specifics of electronic health records [?], together with biological and environmental mechanisms, the method may reflect certain aspects of a healthcare system. for example, closely related but distinct diagnoses are often recorded against the same medical condition. this would induce a positive association between disorders due to healthcare administration rather than biological reasons. advances in systems science and applications (2014) vol.14 no.2 145 so far the attempts to characterise interactions between human disorders, as observed in a population, have been mainly implemented through network approaches [5-6]. while being no doubt informative, network approaches predominantly focus on positive association patterns. one of the alternative approaches to the disease-wide association analysis has been reported by rzhetsky et al. [?]. among other things, the authors objectively characterised probabilities of a person to be affected by a certain disorder, say a, given that the same person has been actually affected by an alternative disorder, say b. this approach allowed us to identify not only positively associated disorders, but also disorders associated negatively – the disorders “competing for the same nucleotide site in the human genome”, as hypothesised by the authors. the shortcoming of the study by rzhetsky et al. [?] is that the authors utilised patient records obtained from a single hospital, also covering a limited pre-selected number of diseases. in the following study, we use electronic health records covering the entire population and much wider range of human disorders. a central objective of the method and its implementation presented below is to visualise association patterns, both positive and negative, among human disorders as observed in an entire population. this would bring to the surface not only already known empirical facts, but also information not previously available. given that the method implementation can reflect empirical information already known, e.g., a strong positive association between hypertension and diabetes mellitus, as well as unexpected association patterns never noticed before, the presented visualisation can serve as a starting point for formulating novel testable hypotheses in the areas of healthcare and medicine. this can further lead to a better understanding of the complex unobserved dynamics of human disorders intersecting in a single biological body. such understanding can practically improve the delivery of healthcare and medical treatments. 2 material and methods 2.1 force-directed spring embedding graph layout algorithm imagine a single pair of nodes, a and b, positioned on a plane and connected by a spring of a certain natural length, δab. when the distance between a and b is exactly dab = δab, the spring is in a state of equilibrium, creating neither attraction nor repulsion forces between the nodes (fig.1(a)). moving a and b further apart from each other would create an attraction force (fig.1(b)), while moving a and b closer to each other would create a repulsion force between the nodes (fig.1(c)). given the values of initial required distances δij between multiple pairs of nodes, it is rarely possible to locate more than three nodes on a plane such that all required distances between them are satisfied exactly. in fact, it is not even always 146 max moldovan: disease universe: visualisation of population-wide disease-wide ... fig.1 a single pair of nodes in three possible states. (a) equilibrium: the nodes neither attract nor repulse; (b) attraction force: the nodes are shifted far away from equilibrium and attempt to attract; (c) repulsion force: the nodes are closer than if they were in equilibrium and attempt to repulse. possible to locate three nodes, keeping the pairwise distances intact, see fig.2. when distances between the nodes are not satisfied, springs connecting them are not in equilibrium, creating a certain force – either attraction or repulsion. fig.2 a hypothetical system of three nodes. the initial distances δij , ij ∈ {ab,ac,bc} between the nodes are given by the theoretical lines a′b′, b′c ′ and a′c ′. the joint length of a′b′ and a′c ′ is less than the length of b′c ′, i.e., δab+δac < δbc . as a result, for the nodes to connect, one or more of the initial distances between the pairs have to be distorted. the three possible states of springs are equilibrium (ab), attraction (ac) and repulsion (bc). aggregated forces created by out-of-equilibrium springs can be expressed by a specific function, drawn from the principle of physics (hooke’s law), leading advances in systems science and applications (2014) vol.14 no.2 147 to system’s potential energy u . while a force can be positive (attraction) or negative (repulsion), an energy level is always non-negative irrespective of the sign of the force. the potential energy of a system of m nodes connected by springs of varying stiffness can be expressed as follows: u = 1 2 ∑ ij∈k ( (dij − δ̂ij) 2 · κij ) , k = ( m 2 ) , i ̸= j (2.1) where dij = √ (xi −xj)2 + (yi − yj)2 is an euclidian distance between nodes i and j with coordinates (xi, yi) and (xj , yj), respectively, δ̂ij is a natural length of a spring between nodes i and j, κij ≥ 0 is an arbitrary parameter that defines the stiffness of a spring between i and j, and k is a number of all possible springs connecting m nodes, with (· · ) being a binomial coefficient. by varying pairwise euclidian distances dij , the force-directed spring embedding graph layout algorithm [7-9] performs a search for the configuration of node locations such that system’s potential energy u is minimised. by finding the minimum energy u , we attempt to obtain a shape of a system of nodes in which competing forces largely compensate each other. minimising function (??) is a complicated task due to the presence of multiple local minima, and it can rarely be guaranteed that a true global minimum is reached, see appendix for details. however, we observed that most nodes have nearly constant “designated” locations with respect to other nodes across alternative local minima achieved when minimising (??). 2.2 defining natural distances between human disorders observing a (sub)-population of size n , suppose that over period t there were ca individuals with at least one occurrence of disorder a, and cb individuals with at least one occurrence of disorder b. further, cab individuals presented with both disorders a and b, each disorder observed at least once over the same period. then the information can be summarised as shown by table 1. table 1 occurrence counts of disorders a and b in population of size n . disorder a disorder b a present a absent total b present cab · cb b absent · · − total ca − n table 1 is an example of a 2 × 2 table with fixed margins. assuming that individuals are affected independently of each other (which can be violated, e.g., 148 max moldovan: disease universe: visualisation of population-wide disease-wide ... for infectious diseases), x follows a non-central hypergeometric distribution x ∼ hyper(n,ca, cb) given by [10-11]: pr(x = cab) = ( cb cab )( n−cb ca−cab )( n ca ) eθabcab (2.2) where max(0, ca+cb−n) ≤ cab ≤ min(ca, cb), θ ∈ (−∞,+∞) is a log-odds ratio, and e = 2.718 . . . is the base of a natural logarithm. for algorithm implementation, conditional maximum likelihood estimates of θ were approximated by unconditional log-odds ratios: θ̂ab = ln ( cab(n − ca + cab − cb) (ca − cab)(−cab + cb) ) (2.3) where ln(·) is a natural logarithm. switching the risk factor from being b for a to being a for b does not effect log-odds estimates. natural (equilibrium) lengths of springs between nodes i and j were obtained through the following reversed expit transform [?, p.121]: δ̂ij = exp(−θ̂ij) 1 + exp(−θ̂ij) (2.4) where δ̂ij ∈ [0, 1] by construction. note that the sign on log-odds estimate θ̂ was changed to the opposite (i.e., reversed), making stronger positive associations correspond to the smaller values of δ̂ij . we do so in order to make δ̂ij resemble euclidian distances between the nodes. due to estimation, there is uncertainty about δ̂ values obtained from the data. such an uncertainty is usually handled by reporting confidence intervals corresponding to δ̂. in the present version of method implementation, we intentionally avoided using confidence intervals, p-values or other statistical tools normally involved in hypothesis testing. we did so in order to reflect the empirical information contained in the data without any subjective interpretation that could otherwise be introduced through, for example, the choice of a significance level. 2.3 potential alternative implementations it should be noted that the force-directed spring embedding graph layout method for visualising relationships among multiple objects, human disorders in our case, is not the only approach available. together with network algorithms already mentioned above, multidimensional scaling [?] and biplot [?] methods are two more approaches for two-dimensional visualisation of relationships between multiple objects. however, it is important to emphasise that both methods, at least advances in systems science and applications (2014) vol.14 no.2 149 in their traditional form, focus on similarities between objects, which would correspond to positive associations between disorders as per our current settings. while an application of alternative methods to our empirical data set would be interesting and potentially informative, multidimensional scaling and biplot methods are unlikely to address negative associations between disorders in a proper way. at the same time, euclidian distances between disorders, as specified by (??) in our study, could be easily reversed, i.e., positive and negative associations can be made corresponding to longer and shorter euclidian distances between the nodes in figure 1, respectively. such an alternative vision of the problem could reveal an entirely new set of empirical information, not reflected by the current method implementation. we retain this research direction for later investigation. 3 practical implementation 3.1 empirical information the presented visualisation has been motivated by the internet map implementation [?]. we use electronic health records obtained from the taiwanese national health insurance research database covering the entire population of taiwan over the period of three years (2000-2002). the same three-year observation window of the maximum available length has been used to record the counts corresponding to table 1. disorder records are based on icd9-cm (international classification of diseases, ninth revision, clinical modification), five-digit version. the dataset has been stratified into male and female groups. each of these two groups has been further stratified into ten age sub-groups, i.e., 0-9, 10-19, . . . , 90+. each subject within a sub-group was noted by his of her first insurance claim starting from 01 january 2000, assigned to a certain age-gender group and followed for the rest of the period ending on 31 december 2002. codes corresponding to e and v categories of icd9-cm (external causes of injury and supplemental classification) were excluded from consideration. 3.2 prevalence threshold and spring stiffness we compute δ̂ij given by (??) for each observed pair of disorders i and j. an empirical examination of δ̂ij revealed that log-odds estimates θ̂ij that underlie δ̂ij , exhibit anomalous behaviour for smaller counts ci and cj , i.e., they tend to be much larger than it would be expected under a random process. such an anomaly would bias the attention of an optimisation algorithm applied to (??) towards diseases of a smaller prevalence. we attributed this anomaly to exceptionally high positive associations between certain low prevalence pairs of disorders as observed in the context of the entire population and reflected by odds ratio estimates. in particular, the expected value of x in the hypergeometric distribution function 150 max moldovan: disease universe: visualisation of population-wide disease-wide ... fig.3 broad disease categories as per icd9-cm classification and the corresponding colour codes as displayed at http://disease-map.net. (??) when θ = 0, i.e., there is no association between disorders, is given by [?]: recij = cicj n (3.5) where rec stands for random expected co-occurrence. we interpret recij as a value reflecting “visibility” of co-occurrences between i and j, with higher visibility (i.e., greater values of recij) leading to more reliable empirical outcomes cij in table 1. keeping this interpretation in mind, we have executed the following ad hoc solution for dealing with the identified anomaly. firstly, we imposed the threshold c = √ 2n on disease occurrence counts. this guarantees that recij > 2 for all possible ci and cj . the meaning behind this restriction is to ensure that only theoretically “visible” θ̂ij estimates are used for visualisation. the cost is that we dismissed smaller prevalence disorders that never exceeded recij = 2 in any of the age-gender groups. secondly to the imposed lower limit on the observed occurrence counts, we set the stiffness parameter of a spring between pairs i and j to κij = ln(recij). this modification makes sure that less theoretically “visible” co-occurrences cij are given less importance when minimising the energy function (??). 3.3 visualisation the google maps platform (https://developers.google.com/maps/) was used to visualise the outcomes. the sizes of the nodes are set to be proportional to the advances in systems science and applications (2014) vol.14 no.2 151 observed disease prevalence in the corresponding age-gender stratified sub-groups. the colour codes of the nodes correspond to the broad disease categories as per icd9-cm classification, see fig.3. all maps are displayed in the same coordinate system with the same scale so that they can be compared against each other. 4 some examples of using the maps 4.1 an accidental proximity? when exploring the maps, some regions and mutual locations can attract attention due to certain, often subjective, reasons. for example, observing the map for females age 40-49 (f40+), in the central region towards south-east we find that peptic ulcer (icd9-cm 533) is located side by side with neurotic disorders (icd9-cm 300). have these disorders fallen close together by chance? a search through the medical literature has quickly identified that an abnormal association between peptic ulcer and neurotic disorders was noticed years ago [15-16]. looking at a wider category of digestive system disorders (icd9-cm 520 to 579), this mutual location pattern remains largely the same over multiple maps (e.g., see m30+, m40+ and f30+), leading to several testable hypotheses. one example of such a hypothesis would be: “there is no direct psychosomatic link between digestive system disorders and neurotic disorders”. viewed from a different angle, the literature coverage on pairs of closely located disorders appears not to be accidental. syed-abdul et al. [?] found that the proximity of disorder pairs is positively correlated with the degree of literature coverage, the latter being represented by a number of hits for a pair of disorders returned from an appropriate query to the pubmed search engine. 4.2 unlikely neighbours: a cancer-schizophrenia association puzzle the topic of observed evidence of associations between schizophrenia and various cancers has been widely debated [?]. evidence tends to point towards the presence of a negative association between schizophrenia and several cancers, even though there is no absolute consensus [20-21]. a shared genetic architecture has been proposed as a reason for observed associations [20-21]. alternatively, there is evidence that the chances of schizophrenia patients being timely diagnosed with certain types of cancer, on average, are lower than for general population nonschizophrenic patients [?]. exploring the maps, it can be found that schizophrenic disorders (icd9-cm 295) consistently fall on the southern border of the maps, sometimes being the most distant points from the imaginary centre of a “galaxy”, e.g., see m50+. interestingly, various types of cancers also regularly fall on the same southern border even though, consistently with the literature, association estimates for schizophrenia-cancer pairs regularly cross to the negative side, i.e., δ̂ij given by (??) exceeds the value of 0.5. 152 max moldovan: disease universe: visualisation of population-wide disease-wide ... one potential explanation for such an anomaly would be that schizophrenia and some cancers have closely related underlying causes revealed through similar relationships with other disorders. in this way, schizophrenia and cancers are placed to their southern border locations by the forces generated within the system. often being negatively associated, schizophrenia and cancers are like two sides of the same coin, “competing for the same nucleotide site in the human genome” as per rzhetsky et al. [?] vision, but potentially disassociated due to deeper, not properly recognised and understood reasons which are still to be identified and investigated. the visualisation we introduced is a tool for originating and directing such investigations. 5 conclusion electronic health records have become an integral part of national healthcare systems worldwide, and it is essential to comprehensively utilise the information contained in the growing number of databases. the method we introduced is one of the effective and informative tools for doing so. while the current realisation of the method has its obvious limitations, the presented maps are the first implementation of this kind and intended to set a reference benchmark for further developments in the same direction. a formal empirical validation of the introduced visualisation is beyond the scope of this paper, but based on the broad examination of the resulted maps, we argue that the presented implementation can both assist with validation of already known phenomena as well as with identification of novel, previously never noticed, association patterns related to functional aspects of medicine and healthcare. we suggest that the maps be used for generating testable hypotheses and invite the reader to explore the vast amount of information contained in them at http://disease-map.net. acknowledgements: we are grateful to chris lloyd and nicholas nechval who provided their constructive feedback that helped to improve the manuscript. we thank hanna kalkova and andrey stepanov for preparing figures 1 and 2. author contributions: m.m., r.e., s.s-a. and y-c.l. invented and developed the concept (ipn-13-000062, newsouth innovations). s.s-a., m.m., y-c.l. and y-c.c. performed selective biological validation. r.e. wrote software, implemented the method and developed the website. m.m. and r.e. wrote the manuscript. y-c.l. obtained the data. pa.n. organised and validated the data. funding: no formal targeted funding was allocated towards the project. competing financial interests: the authors declared no competing financial interests. advances in systems science and applications (2014) vol.14 no.2 153 references [1] ferrannini, e. and cushman, w.c. (2012), “diabetes and hypertension: the bad companions”, lancet, no.380, pp.601-610. [2] rzhetsky, a, wajngurt, d., park, n., and zheng, t. (2007), “probing genetic overlap among human phenotypes”, proceedings of the national academy of sciences, no.104, pp.11694-11699. [3] chou, f.h.-c., tsai, k.-y., su, c.-y., and lee, c.-c. (2011), “the incidence and relative risk factors for developing cancer among patients with schizophrenia: a nine-year follow-up study”, schizophrenia research, no.129, pp.97-103. [4] hripcsak, g. and albers, d.j. (2013), “next-generation phenotyping of electronic health records”, journal of the american medical informatics association, no.20, pp.117-121. [5] hidalgo, c.a., blumm. n., barabási, a.-l., and christakis, n.a. (2009), “a dynamic network approach for the study of human phenotypes”, plos computational biology, vol.4, no.5, e1000353. [6] barabási, a.-l., gulbahce, n., and loscalzo, j. (2011), “network medicine: a network-based approach to human disease”, nature reviews genetics, no.12, pp.56-68. [7] kamada, t. and kawai, s. (1989), “an algorithm for drawing general undirected graphs”, information processing letters, no.31, pp.7-15. [8] tunkelang, d. (1999), “a numerical optimisation approach to general graph drawing”, phd thesis, carnegie mellon university: http://reportsarchive.adm.cs.cmu.edu/anon/1998/cmu-cs-98-189.pdf [9] hu, y. (2005), “efficient, high-quality force-directed graph drawing”, the mathematica journal, no.10, pp.37-71. [10] lloyd, c.j. (1999), statistical analysis of categorical data. wiley (new york). [11] agresti, a. (2002), categorical data analysis. 2d edition, john wiley & sons. [12] cox, t.f. and cox, m.a.a. (2001), multidimensional scaling. chapman and hall. 154 max moldovan: disease universe: visualisation of population-wide disease-wide ... [13] gabriel, k.r. (1971), “the biplot graphic display of matrices with application to principal component analysis.” biometrika, no.58, pp.453-467. [14] enikeev, r. (2012), the internet map. singapore, http://internet-map.net. [15] montgomery, h., schindler, r., underdahl, l.o., butt, h.r., and walters, w. (1944), “peptic ulcer, gastritis and psychoneurosis”, journal of the american medical association, no.125, pp.890-894. [16] stafford-clark, d. (1952), “peptic ulcer and the neurotic”, british medical journal, no.16, pp.390-391. [17] syed-abdul, s., enikeev, r., moldovan, m., jian, w.-s., nguyen, a., iqbal, u., chang, y.-c., hsu, m.-h., lin, s.c., and li, y.-c. (2013), “capturing and visualizing human diseasomic associations: a population-based observational study”, manuscript. [18] hodgson, r., wildgust, h.j., and bushe, c.j. (2010), “cancer and schizophrenia: is there a paradox?”, journal of psychopharmacology, no.24, pp.51-60. [19] hippisley-cox, j., vinogradova, y., coupland, c., and parker, c. (2007), “risk of malignancy in patients with schizophrenia or bipolar disorder”, archives of general psychiatry, no.64, pp.1368-1376. [20] catts v.s., catts s.v., o‘toole b.i., and frost, a.d.j. (2008), “cancer incidence in patients with schizophrenia and their first-degree relatives – a meta-analysis”, acta psychiatrica scandinavica, no.117, pp.323-336. [21] gal, g., goral, a., murad, h., gross, r., pugachova, i., barchana, m., kohn, r., and levav, i. (2012), “cancer in parents of persons with schizophrenia: is there a genetic protection?”, schizophrenia research, no.139, pp.189-193. [22] crump, c., winkleby, m.a., sundquist, k., and sundquist, j. (2013), “comorbidities and mortality in persons with schizophrenia: a swedish national cohort study”, american journal of psychiatry, no.170, pp.324-333. corresponding authors maxmoldovan (max.moldovan@gmail.com) and yu-chuan li (jaak88@gmail.com) appendix: energy minimisation method finding a global minimum of (??) is a complicated task due to the presence of multiple local minima of this function. different approaches of global minimisation can be applied, but it can be rarely known when and if the global minimum advances in systems science and applications (2014) vol.14 no.2 155 is reached, unless a minimum energy level is known in advance. our current implementation of energy minimisation is to use multiple local searches with the conjugate gradient algorithm from random starting positions in order to obtain a master map – the map that includes diseases across the entire spectrum of age groups and both genders. in each of multiple attempts, the nodes are dropped on the map with random positions (x,y ), and the conjugate gradient algorithm runs searching for the closest local minimum of u in (??). if the new local minimum is less than the best (smallest) minimum recorded over previous attempts, it becomes the new best minimum. the procedure is repeated until the best minimum stops changing even after a reasonably large (4000, in our implementation) number of random allocation attempts, see algorithm ??. the computational complexity of the algorithm is o(n2). the obtained master map served as a collection of starting points for the agegender stratified maps, see algorithm ??. minimising (??) from a single set of starting points leads to a local minimum that can be further improved by applying the minimisation approach used for obtaining the master map. however, we still used minimisation from the single set of starting points in order to make the maps comparable across age groups and genders. table a1 reports the achieved minimum energy levels using “partial” minimisation as per algorithm ?? compared to the “full” minimisation implemented through algorithm ??. 156 max moldovan: disease universe: visualisation of population-wide disease-wide ... table a1 minimum achieved energy levels from partial and (attempted) full minimisation approaches. group subjects followed (n) disorder numbers partial full per cent improve f 0-9 1,677,365 565 7,807.96 7,700.69 1.39 f 10-19 1,595,057 743 9,470.09 9,166.34 3.31 f 20-29 1,780,095 1041 22,897.04 22,268.39 2.82 f 30-39 1,765,866 1136 25,914.10 25,387.60 2.07 f 40-49 1,631,968 1243 31,126.22 30,913.37 0.69 f 50-59 930,496 1251 33,451.35 33,334.80 0.35 f 60-69 711,096 1271 36,129.76 36,056.87 0.20 f 70-79 427,821 1177 29,935.15 29,857.36 0.26 f 80-89 141,225 783 10,802.72 10,773.33 0.27 f 90-99 8,532 176 318.26 315.74 0.80 m 0-9 1,827,447 630 10,068.44 9,910.12 1.60 m 10-19 1,678,415 721 9,451.03 9,346.16 1.12 m 20-29 1,767,163 859 12,532.53 12,345.32 1.52 m 30-39 1,737,715 948 14,263.07 14,099.18 1.16 m 40-49 1,577,320 1090 19,485.22 19,454.02 0.16 m 50-59 898,150 1065 20,296.22 20,247.20 0.24 m 60-69 692,061 1163 26,737.58 26,563.97 0.65 m 70-79 532,308 1225 30,740.90 30,622.10 0.39 m 80-89 133,480 781 10,636.67 10,599.66 0.35 m 90-99 4,769 151 240.25 238.70 0.65 master 21,518,574 2298 − 130,381.91 − advances in systems science and applications (2014) vol.14 no.2 157 algorithm 1 energy minimisation for the master map. require: γ ← 0.01 /* tolerance for the change in objective function (??) require: s← 1 /* initial step size require: τ ← 0.9 /* step decrease rate require: smin ← 0.000001 /* minimum step tolerance require: δ̂ij for k = ( m 2 ) pairs, i ̸= j /* pairwise natural lengths given by (??) require: ucurrent ← +inf /* current energy level to be reduced require: cc← 0 /* random positions attempts counter require: ccmax ← 4000 /* maximum number of attempts with no energy reduction while (cc < ccmax) do (x0, y0)← random() /* drop nodes at random positions u0 ← fe(x0, y0) /* value of objective function (??) g← {−∇ (fe(x0, y0))} /* define antigradients for the first step (∆x,∆y )← fg(g) /* step direction (x,y )← (x0, y0) + (∆x,∆y ) · s /* current coordinates of nodes u ← fe(x,y ) /* current value of objective function (??) ∆u ← (u0 − u) /* change in energy while (∆u > γ) & (s > smin) do gc ← {∇c (fe(x0, y0;x,y ))} /* evaluate conjugate gradients (∆x,∆y )← fcg(gc) /* step direction (xtemp, ytemp)← (x,y ) + (∆x,∆y ) · s /* trial coordinates of nodes u ← fe(xtemp, ytemp) /* current value of the objective function if u < u0 then ∆u ← (u0 − u) /* update change in energy u0 ← u /* update preceding energy value (x0, y0)← (x,y ) /* update preceding coordinates (x,y )← (xtemp, ytemp) /* assign the values of current coordinates else s← s · τ /* reduce step size end if end while if u0 < ucurrent then ucurrent ← u0 /* update minimum energy value (xcurrent, ycurrent)← (x,y ) /* update coordinates cc← 0 /* set attempts count to zero else cc← cc+ 1 /* next attempt end if end while (xmaster, ymaster)← (xcurrent, ycurrent) return (xmaster, ymaster) /* nodes’ coordinates under minimum energy achieved 158 max moldovan: disease universe: visualisation of population-wide disease-wide ... algorithm 2 energy minimisation for age and gender stratified maps. require: γ ← 1e-5 /* tolerance for the change in objective function (??) require: s← 1 /* initial step size require: τ ← 0.9 /* step decrease rate require: smin ← 0.000001 /* minimum step tolerance require: δ̂ij for k = ( m 2 ) pairs, i ̸= j /* pairwise natural lengths given by (??) (x0, y0)← (xmaster, ymaster) /* use coordinates from the master map as starting points u0 ← fe(x0, y0) /* the value of objective function (??) g← {−∇ (fe(x0, y0))} /* define antigradients for the first step (∆x,∆y )← fg(g) /* step direction (x,y )← (x0, y0) + (∆x,∆y ) · s /* current coordinates of nodes u ← fe(x,y ) /* current value of objective function (??) ∆u ← (u0 − u) /* change in energy while (∆u > γ) & (s > smin) do gc ← {∇c (fe(x0, y0;x,y ))} /* evaluate conjugate gradients (∆x,∆y )← fcg(gc) /* step direction (xtemp, ytemp)← (x,y ) + (∆x,∆y ) · s /* trial coordinates of nodes u ← fe(xtemp, ytemp) /* current value of the objective function if u < u0 then ∆u ← (u0 − u) /* update change in energy u0 ← u /* update preceding energy value (x0, y0)← (x,y ) /* update preceding coordinates (x,y )← (xtemp, ytemp) /* assign values of current coordinates else s← s · τ /* reduce step size end if end while (xstratif , ystratif )← (x,y ) return (xstratif , ystratif ) /* nodes’ coordinates under minimum energy achieved advances in systems science and application (2016) vol.16 no.2 39-53 principal causes of financial frises in latvia for last 20 years valery i. roldugin baltic international academy, latvia abstract in the article the author analyses activity of the latvian banks with the purpose of search of ways from a crisis situation. in the middle of 2000th rapid development of crediting resulted in negative tendencies. it was assisted by not only superfluous liberalization of crediting but also decline of quality of supervision. successful development of economy of latvia in a great deal depends on stability of commercial banks that accumulated money resources. besides the commission of market of finances and capital, successful development of the banking system in a great deal depends on activity of central bank and government. therefore an exit from a financial crisis depends on latvian commercial banks will be able to tide over financial difficulties and pick up thread crediting of enterprises. secondly, on the basis of the analysis of events development in the banking area the author discloses the true reasons which brought the latvian economy to the deepest financial crisis among the countries of the european union. the period from 1995 till 2014 is being investigated. the author uses a wide range of research methods, such as: grouping method, method of comparison of financial ratios, etc. keywords central bank; commercial banks; financial market; bank lending; real estate market. 1 introduction after latvia had regained its independence the central bank began to reform the credit system. we can note both positive and negative moments of its activity. undoubtedly, one of the reasons of banking crises in 1995 and 1998 were the shortcomings of activity of the bank of latvia[1]. the fast price liberalization, cancellation of subsidies, manipulation with funds of the enterprises at exchange of old money to new one also contributed to the creation of certain problems. at the same time the essential changes in relation to streamlining of requirements for credit institutions occurred in the field of regulation of all banking system in latvia. this has allowed the strengthening of the position of the whole credit system of the country, increasing confidence in the banks and improving of their financial performance. prior to 1995 the number of failures did not exceed two or three cases per year. in a result of crisis, the number of licenses revoked by the bank of latvia reached 15: tautas banka, latgales komercbanka, latintrdes banka, depozitu banka, 40 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years centra banka, alejas komercbanka, polarzvaigzne, liepajas komercbanka, kredo banka, olimpija, oltibanka, talsu komercbanka, lettika, rigas starptautiska banka and banka baltija. in 1996, more seven banks have lost their licence: tukuma banka, daugavas banka, bauskas banka, ako banka, dinastija, jelgava, atmoda. despite the fact that the number of closed banks has been high, the total bank capital had resumed quickly. a considerable reduction of banks and the fact that the top ten banks were in this list, including the country’s largest bank banka baltija, significantly affected the main indicators of the industry, what gave a reason to speak about the banking crisis in latvia in 1995. 2 bank crisis in 1995 one of the main reasons of the banking crisis of 1995 was the fact that banks did not manage the risks in due manner. banks have been working in markets which had not been enough good explored. the recession, slow rate of reforms and the imperfection of legislation kept the already high degree of risk in bank activity. many of them enjoyed the demand for bank loans. they provided loans and attracted deposits at high rates, which had fallen harshly subsequently. a tough position of the bank of latvia in relation to the supervision over commercial banks had a considerable importance for stabilizing the banking system. only a year after the crisis, the bank of latvia issued a dozen of rules, guidelines and regulations. the most important of them included the toughening of regulatory requirements for commercial banks, providing the bank of latvia with information about shareholders, restrictions on attracting by banks of deposits from individuals, which had limitations in financial activities, mandatory annual audits by international audit companies, coordination with the bank of latvia of any changes in the capital, in management, etc. though all these measures have complicated the activity of commercial banks, streamlining of their activity was justifiable because of positive efficiency in development of banking system. some banks have had incompetent managers who failed to forecast the development of financial and currency markets. notwithstanding to the valid legislation the bank of latvia was approving people without economic education, or even without experience of work in the banking system, for the positions of chairmen of the board and their deputies. there was a shortage of qualified specialists in the field of lending, foreign exchange and bank marketing. as a result, the mistakes in lending and evaluating business plans were quite often. the banking regulatory authorities did not carry out a due control. advances in systems science and application (2016) vol.16 no.2 41 3 bank crisis in 1998 according to the law “on credit institutions”, which came into force in october, 1995, the bank of latvia got new supervising powers in relation to both licensing of banks and essential changes in structure of shareholders[2]. bank of latvia considerably strengthened this control, especially with acceptance of new rules of licensing in the beginning of 1996. the bank of latvia regularly checked banks records and events connected with the change of shareholders, and if the rules weren’t observed in full completeness, it applied the financial sanctions. the bank of latvia established a wide set of standards and regulations for credit institutions. developing them, it proceeded from principles and the standards valid in banking systems of the developed countries, considering eu regulations in the bank sphere and principles of basel committee of bank supervision. some of them were even stricter, for example, the requirement to the capital sufficiency (10%) to cover the defined risks. basel committee set a sufficiency of the capital at a level of 8%. it is necessary to notice that this standard had no big importance at that time since the capital sufficiency of bank industry at that time was more than 17%. at that time the bank of latvia has played an important role in the regulation of bank liquidity in order to limit the fluctuations in interest rates. as a matter of fact the value of interest rate is defined by the cash flows and the value of supply and demand in the financial market. the bank of latvia set a refinancing rate, which served as the initial point for the whole banking system. the short-term loans have been provided to the commercial banks under this rate. one of the most popular instruments of the bank of latvia was the credit refinancing which using began in 1993 by issuing the short-term loans to ensure the liquidity. in the late 1990-s there was a tendency in latvia of lowering of the refinancing rate. the interest rate of refinancing established in 1993 at level of 120% per annum, decreased to 4% by 1997. such policy made it possible to create more favorable conditions for banks lending, what made it possible to carry out a policy of credit expansion. the commercial banks lowered the interest rates, making their loans more available for economy development. the increased demand for loans thus caused the increasing of investments. initially, the loans were given to each bank within specific maximum limits which depended on the fulfillment of regulatory requirements stipulated by the bank of latvia. such order was necessary because the loans were given without any collateral. the amounts of the refinancing were relatively limited, as most of them have been used for foreign currency converting. though the loans were small, nevertheless they allowed banks keeping the liquidity in cases when those couldn’t get them at the interbank market which was poorly developed at that time. 42 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years the development of securities market has reduced the necessity to cover the domestic debt by the loans of the bank of latvia. as a result, there appeared a possibility to limit the additional emission and to reduce inflation. the internal state bonds were the main source of the budget deficit financing. in 1996, the demand was much higher than supply. the amount of reserve requirement established by the bank of latvia on july 15, 1993 was 8% of the balances involved in all types and terms of money. this meant that each bank having a license was obliged to reserve the average daily funds appropriate to the amount mentioned in the requirements in the period from 16th day of previous month until the 15th day of the reporting month. therefore, the amount of funds in different days could differ considerably. at the same time, 50% of the required reserves could be the average cash balances, while the rest part should be reserved on the correspondent account with the bank of latvia. as of december 1998, the volume of bank reserves amounted to 129.9 million lats, including 34.3 million in cash balances. cash balances in foreign currency were not taken into account in meeting the reserve requirements. assessment of standards of execution of reserve requirements under average value gave the banks chance to handle their funds more freely and to provide liquidity in settlements. in order to consider the reserve requirements as monetary instrument to control the money supply, in april 1, 1997 the bank of latvia gave the banks the right not to include the obligations related to the transactions with foreign central banks, local and foreign credit institutions into the calculation of reserve requirements. as a result, banks got an opportunity to increase the money supply in the frameworks of this conservative monetary policy instrument. after a year since the introduction of these changes, the volume of liabilities to foreign central banks and credit institutions increased by more than three times. this gave a chance to attract additional loan resources for a longer period of time. at the end of april 1997 they amounted to 57.1 million and increased to 183.5 million to the end of april 1998. in 1997, the lowering of credit rates promoted the lowering of discount rates for liabilities. in the beginning of 1997, they were 8-10.5 %, and decreased up to 3.5-5.5 % by the end of the year. the total amount of the state securities in 1998 fell to 18.4 %, and amounted to 127.0 million in the end of the year. respectively, the share of all term liabilities available in circulation decreased in comparison with the previous year, i.e. 12 month liabilities-36.5 %, 6 month-8.0 %, 3 month 1.7 %, 1 month-0.2 %, and the share of 2 year bonds was 53.6 %[1]. however, the sustainability of financial market at that time was achieved mainly due to foreign currency interventions. it was possible to reduce loans interest advances in systems science and application (2016) vol.16 no.2 43 rates owing to the operations of buying and selling of foreign currency, what favorably affected the formation of the business environment. a tough position in establishment of the rate of lat allowed limiting the inflation, increasing the fiscal discipline, creating the conditions for sustainability of business and the confidence of foreign investors. it is worth regretting about poor coordination between the government and the bank of latvia. should the national currency be supported timely by reasonable tax policy stimulating, it could be possible to cause a domestic investment demand, which would enable a more efficient structural change in the economy. in process of its development the monetary market and securities market would have a stimulating effect on production development, creating the conditions for linkage in a trade turnover of considerable monetary funds. since july 1, 1997 the “terms and conditions of trade (currency swap transactions) on the purchase with subsequent sale and sale with subsequent lats repurchase” came into force. the bank of latvia has developed them with the aim to introduce a new financial instrument. the terms of such transactions does not exceed three months. since may 12, 1998 the bank of latvia began organizing auctions of currency reverse transactions quarterly. the volume of currency “swap” transactions in december 1998 amounted to 23.1 million lats, including 7.5 million lats with 7 day period (average interest rate 6.8%), 6.9 million lats with 28 day period (average interest rate 7.5%), 8.7 million lats with 91 day maturity (average interest rate 8.6%). regarding the involved deposits the demand deposits dominated in the banking structure. in april 1998, their share in total deposits was 73.4%[1]. lets note, that the currency market traditionally has been one of the most developed and liquid sector of financial market in latvia. its average turnover per day amounted to 279.9 million lats as of april 1998. in contrast, the average daily turnover of interbank market was almost four times less (76.3 million lats) during this period. prior to 1994 it was impossible to carry out the operations on open market, as government securities werent available. after their issuance the volume of state bonds grew relatively slowly, and their liquidity was not enough high, because the secondary market of liabilities was inactive and poorly developed. such operations gained their significance after 1997. the volume of state securities in the secondary market amounted to 704.5 million lats in 1998. the share of transactions with residents has increased from 46.4% (in 1997) to 78.0% (in 1998), and non-residents transactions decreased from 18.0% to 2.3%. by carrying out the operations in the open market the regulation of the reserves has been made providing the fixed level of interest rate, defined by a tactical task in spite of the eventual changes of interest rates specified by impact of market 44 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years processes. the bank of latvia carried out the open-market purchases of securities as a stabilizing measure, thereby extending bank reserves and increasing the money supply to the extent that was sufficient to keep interest rates on the same level. this increase in money demand put pressure on the interest rate. as the interest rate exceeded the target level, the bank of latvia started buying the securities from commercial banks. the process of using the excessive loan resources increased the money supply in circulation. open-market purchases continued until the moment when sufficient amount of new money has been formed ensuring the compliance with the demand and supply at a fixed interest rate. thus, the latvian market has been achieving a certain balance. when the bank of latvia was selling the securities to the banks, sterilizing the money supply, the interest rates were increasing. the volume of transactions on purchase and sale of securities amounted to 109.5 million lats in 1998 (in previous year 89.7 million lats). the most volume of transactions was in august (purchase 30.2 million lats) and september (sale 11.5 million lats). the bank of latvia was buying mainly the government bonds, increasing its securities portfolio and replenishing the amount of active money by such operations. the securities portfolio of the bank of latvia in the middle of 1998 was 40.6 million lats (at par). since october, 1995 the bank of latvia began actively using “repo” auctions. the interest rates on “repo” auctions fixed by the bank of latvia were differentiated on terms and could be both above and below the level of a rate of refinancing. the bank of latvia carried out “repo” auctions daily, offering the credits for term of 7, 28 and 91 days. the interest rate at auctions exceeded a refinancing rate a bit. in 1997, the volume of the credits of the “repo” auctions with validity period of 7 days amounted to 55,7 million lats , for 28 days-33,1 million lats, for 91 days-5,7 million lats. in 1997, the bank of latvia credited the commercial banks in amount of 112.6 million lats that was more on 3.5 % than in 1996. their structure looked as follows: repo auctions-83.9 %, auto pawn-14.8 %, a pawn on demand-1.3 %[1]. the volumes of lending have been increasing every year. for example, in 1998, the loans to commercial banks reached 458.0 million lats, which was 4.1 times more than in 1997. from them: 58.8% was issued in repo auctions, 33.1% of pawn demand, 4.2% auto pawn and 3.9% loans in emergency situations. in 1998, repo loans for 194.4 million lats were issued with 7 day term, 40.6 million lats28 day term and 34.5 million lats with 91 day maturity. therefore, the commercial banks could solve the problem of short-term liquidity, and the bank of latvia knowing the maturity of loans, could make short-term forecasts of amendments in the liquidity of all banking sector. advances in systems science and application (2016) vol.16 no.2 45 while and today, one of the problems slowing down the development of a national economy of latvia[3, 4], was a deficit of balance of payments. foreign trade still remained a weakest point of national economy. in 1997, there was a growth of both in export and import. and the gap between them increased by high rates, creating a considerable deficit in the countrys trade balance. the volume of foreign trade in 1997 amounted to 2554.1 million lats that more on 23.2 % than in previous year. the deficit reached 610.7 million lats at export amounting to 971.7 million lats in1997. the payment of the international services already couldn’t compensate the missing foreign currency to pay the imported goods. the missing currency could be received generally by the means of loans. therefore, the trade surplus of the balance of payments at that time depended on capital flow more and more. it was clear, that since the producers of the goods don’t find the new markets and new buyers, the country will import on credit, live in the account of converting the foreign loans. the only one solution of successful development of economy was to search ways of increase the inflow of foreign investments into manufacture and infrastructure. the current account deficit of balance of payments could be financed either by foreign funds coming to the country or by using of foreign exchange reserves. the balance of payments summarizes the flows of goods, services, capital and related income interest payments, dividends of latvian residents and foreign countries. considering the limitation of the state budgetary funds and caution of commercial banks at crediting of the long-term projects, only the foreign investments could be the main source of long-term capital. however, the inefficient economic management, when the ministers of economic departments have been changed several times a year, volatile fiscal policy, the absence of free land market, the disorder in legislation and other factors prevented the attraction of foreign investments. since 1996, the government tried to limit the growth of external debt. the main reason lied in inefficient use of earlier received credits. in relation to state debt and the budget deficit, latvia focused on the implementation of the maastricht requirements. the external debt included loans from international financial organizations and the private sector. a number of loans was provided to latvia for a considerable period and under very favorable terms, i.e. the loan of the international bank for reconstruction and development-for 17 years (under 7.6% per annum), exim bank of japan for 16.5 years (4.7%), the european bank for reconstruction and development for 15 years, the credit of the american “commodity credit corporation” 30 years (terms of initial payments were extended to 5 years). the balance of state debt at the end of 1998 amounted to 231.5 million lats (6.1% of gdp) , and the government has issued guarantees in amount of 42.9 46 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years million lats in order the enterprises could get the foreign loans. in 1998, the government has spent 37.0 million lats to administrate the external debt. the amount spent for servicing the external government debt equaled to 3,5% of annual export amount. the domestic debt included the government borrowings in the local market. it grew every year. for example, it increased by 5.4% and amounted to 8.3 million lats during 1997[1]. contrary to the previous crisis, the bank crisis of 1998 was caused not only by internal shortcomings of activity of latvian banks, but also by many external factors. the financial crisis in russia burst in august, 1998 was the main reason. respectively, the russian crisis negatively affected the development of the banking sector of latvia in the second half of 1998, as the majority of banks formed the considerable part of their assets by investments to cis countries. the main aspect was that considerable monetary resources left latvia in the form of investments into foreign securities because the latvian banks had actively operated on securities markets of other countries. its known that the commercial banks way out from a crisis situation in 1995 was reached for the account of increase of the operations with investments and converting of currencies in russia and cis countries. the bank of latvia didnt see it. in november, 1996, the investment into foreign securities reached 121.5 million lats and amounted to 171 million lats in august, 1998. such migration of capital was stipulated by high interest rates on the russian government securities and securities of the entities what created a possibility to receive profit by the latvian banks. the interest rates nearly three times lagged behind their counterparts in the cis at that time in latvia. unfortunately, the bank of latvia didn’t take the timely actions to protect a national banking system from possible financial crashes in the monetary markets of cis countries. the financial crisis of 1998 brought serious changes to the development of bank sphere of latvia. it negatively affected the activity of all banks. the total losses of credit institutions according to financial reports of auditors for 1998 amounted to 56.6 million lats (only 8 banks got profit in amount of 3.5 million lats). two banks got bankrupt (latvijas kapitalbanka and komercbanka viktorija), and rigas komercbanka was in a state of insolvency with losses amounting to 9 million lats in the beginning of 1999. parekss banka had the greatest profit-1.21 million lats (the total declared profit before audit 5.6 million lats). the losses of latvijas unibanka were 15.1 million lats[1]. in spite of the fact that the bank of latvia revoked licenses of two banks during 1998: latvijas zemes banka (in connection with the merger with hansabanklatvija) and doma banka (due to insufficiency of the equity), we can state that the whole banking system has successfully overcome the negative consequences of crisis. the positive changes in key indicators of the banking industry were the advances in systems science and application (2016) vol.16 no.2 47 evidence as well as the fact that 19 banks from 24 finished the year of 1999 with a profit. the process of enhancement of the legislation continued. in 1998, two new laws came into force: “on the prevention of laundering of proceeds derived from criminal activity ” and “on guarantees of deposits of individuals”. the adoption of the law ”on the financial and capital market commission” was an important event in the process of improving the legislation base[5]. the introduction of unified supervision on the market of finance and capital in latvia became a significant event. as a result, agency on supervision of credit institutions of the bank of latvia, state inspection for supervision of the insurance market and the commission of the securities market were united into the financial and capital market commission. the number of banks in latvia table 1 bank and branches in latvia from 2005 to 2014 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 total number of the banks 22 21 21 21 21 21 20 20 19 17 branches of foreign banks 1 3 4 6 8 10 9 9 9 10 number of branches 215 224 235 247 231 223 406 342 313 average number of employees 10,520 11,611 13,334 14,381 12,628 11,616 10,256 9,845 9,362 customers current accounts, thousands 2,912 3,316 4,341 4,545 4,462 4,525 4,539 4,568 4,581 3,078 customer accounts available via internet, thousands 2,538 2,859 3,015 3,163 3,230 3,444 3,643 2,618 payment cards, thousands 1,711 2,107 2,389 2,518 2,478 2,424 2,325 2,381 2,369 2,305 number of atms 878 957 1,143 1,274 1,320 1,359 1,207 1,270 1,154 1,068 pos terminals accepting payment cards 18,495 17,571 20,367 23,350 24,381 24,366 25,430 26,259 29,066 31,121 has considerably reduced in a period from 1992 to 2005. thus, if there were 61 commercial banks in 1993, in 2005 the number of commercial banks reduced to 22 banks (in 2014 -17 banks). the licenses of commercial banks have been issued by the bank of latvia which carried out the supervision over banks activities. since 2001, the bank of latvia handed over these responsibilities to the financial and capital market commission. 4 financial crisis of 2008 the volume of loans increased up to 1,086.7 million lats at the end of 2000, 3.8 times more than in 1996. the year of 2001 has been the most successful for 48 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years the banking sector of latvia. the best indicator of the annual profit in amount of 46.7 million lats has been achieved this year. the indicators of all banking industry have been growing rapidly, i.e. the assets grew by 28.2%, deposits by 24.6% and shareholders’ equity 35.4%, and the volume of loans increased by 50.5%[1]. by this time, the main role in development to a bank system was played by three largest latvian banks. parekss banka was the leader of banking system in the beginning of 2000. it was the largest bank in latvia by size of total assets, equity and the attracted funds of the clients. from the beginning of 2000 to 2003 the assets of parekss banka increased by 74 % and reached 954 million lats. speaking about the dynamics of development among the largest banks, the leader was hansabanka which assets have been doubled for the similar period. latvijas unibanka, the second-top assets bank of latvia showed less aggressive dynamics at that period. its assets increased only by 57 %. at the same time, latvijas unibanka had the greatest amount of the loans among the largest banks of latvia 543 million lats. it was the largest lender among the largest latvian banks, the share of its credit portfolio in assets composed 74.5 %. this indicator clearly reflected the credit specialization of latvijas unibanka. it should be noted that credit-oriented latvijas unibanka had 135 ratio as regards of the issued loans and clients funds. it means that the bank allocated not only all client funds for loans, but also 35 % of other involved funds and own resources. accounting the supervision by the scandinavian bank group (seb), most likely the loans were financed from parent bank as well. hansabanka had the second rank by amount of the issued loans amounted to 497 million lats. hansabanka had essentially increased the credit portfolio for the last two years and bypassed of parekss banka under this indicator. hansabanka share of the issued loans in assets had been also high-70.5 %. parekss banka had relatively small credit portfolio among three largest banks of latvia (381 million lats), and its dynamics lagged behind the growth of credit portfolios of other banks in that triplet. in whole, the structure of assets of parekss was oriented on the investments into securities and interbank operations. after all, the credit portfolio even in structure of assets of parekss banka composed approximately 40 %. so, the triplet of system banks approached to today’s crisis with such indicators. regarding the prospects of development of the largest commercial banks of latvia it was possible to notice the following aspect: the affiliated structures of foreign banks (latvijas unibanka and hansabanka) still continued the dynamic development, using the resources of parent structures and favorable position on the market. these two banks were targeted to the increase of amounts of loans. parekss banka has been managed by major shareholders and constantly was advances in systems science and application (2016) vol.16 no.2 49 in process of search of the strategic investor. the parex financial group enlarged what fostered to attract the additional capital. with time, it was extended by parex bankas (lithuania); insurance companies parekss apdrosinasanas kompanija (latvia) and baltic polis (lithuania); leasing companies parekss lizings (latvia) and parex lizingas (lithuania). the activities of broker company parekss brokeru sistema and the pension fund parekss atklatais pensiju fonds effectively supported the activity of parex business group. the volumes of mortgage loans have formed a significant part of all loans issued in latvia. within 2007, the total amount of mortgage loans increased by 40.6% and their share in total loans was 52.9% that was 1.1% more than at the end of the previous year. in comparison with the end of 2006, the total loans to businesses for current assets increase, has grown by 24.4%, and its share in total loan portfolio of commercial banks amounted to 20.3%. the aggregate loan portfolio of commercial banks at the end of 2007 was divided between residents and non-residents of latvia as follows: residents-87.9 %, non-residents-12.1 %. similarly, the loan portfolio of borrowers-residents was divided as follows: private non-financial companies-49.8 %, households-41.2 %, financial institutions-6.7 %[1]. on the end of 2006, the latvian banks issued 916.5 thousand loans. in 2007, the number of the issued loans increased on 211.4 thousand (30 %). the number of the loans issued to private residents for acquisition, reconstruction or repair of housing has increased by 28.5 thousand (27.6 %) and reached 131.7 thousand loans; the consumer loans increased on 43.5 thousand (38.3 %) and reached 157 thousand loans. the number of the loans on individuals payments cards and accounts reached 583.9 thousand and increased on 131.9 thousand (29.2 %) in comparison with previous year[1]. the real estate in latvia was one of the basic instruments of savings. throughout more than ten last years it has been bringing high profits not only to businesses but also to individuals. a rare company did not invest their money into real estate transactions. these years were marked by prices explosion for all types of apartments throughout latvia. for example, there was 100% growth of prices in riga and its vicinity only in 2006. we can assume that the main reasons for such prices increases were the insufficient amount of new housing, the increase in solvent demand of people and growth of speculative transactions.the real estate in latvia was one of the basic instruments of savings. throughout more than ten last years it has been bringing high profits not only to businesses but also to individuals. a rare company did not invest their money into real estate transactions. these years were marked by prices explosion for all types of apartments throughout latvia. for example, there was 100% growth of prices in riga and its vicinity only in 2006. we can assume that the main reasons for 50 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years such prices increases were the insufficient amount of new housing, the increase in solvent demand of people and growth of speculative transactions. in a result of housing crediting the number of real estate transactions in the mortgage market increased considerably. if in 2000 there were registered 22,689 transactions of real estate sale and purchase, by the beginning of 2007 their number was doubled, and 65 % of them were financed by the bank loans[6]. the basic reasons of crediting growth were: first, availability of bank loan upon housing purchase, secondly, a rise in prices for real estate. the commercial banks themselves played a considerable role in real estate prices rise while demanding a security at crediting. therefore, even the inexpensive apartments were accepted as a security by the prices which was seemed unrealistic to the appraising firms even a year before. as well, the decrease in interest rates for long-term loans played its role on the latvian real estate market. at the end of 2004 the price explosion was observed on the market, the price grew even several times on specific segments. the greatest growth was observed on a segment of apartment houses, i.e. in riga they increased more than 7 times in average in comparison with 2000. such growth had forced bank of latvia to undertake the urgent measures on restriction of amounts of crediting, including increase in interest rates for longterm loans. as a result, if in 2004 the average interest rate in latvia decreased by 2.5 percent points compared to 2001, for the next 4 years it increased again by 6.9 percent points, and made 14.6 % per annum in 2007[1]. it is necessary to notice that the consumer loans have been issued not only by banks, but also by other companies having no relation to bank activity and not supervised by the financial and capital market commission. according to the research of dnb nord banka, the latvian real estate market was saturated by 1.8 billion lats of free money along with the bank mortgage loans since the beginning of 2007, i.e. about 150 million lats per month (5). the calculations were made on the basis of data of the real estate companies stating that a half of transactions on housing acquisitions for the last year were financed by such loans. in april, 2007, the government developed the program to fight against inflation. the most efficient of this 30-points program for chilling the market were the following points: the developer must brought to bank his/her own funds in amount of 10 % of a project cost prior to the beginning of construction as well as must submit the reference about the amount of the received income. as a result, the performance indicators of latvian real estate market changed extremely negative. the volumes of mortgage lending began to fall quickly even before world financial crisis. the prices for new housing gradually decreased as well as on secondary housing market. thus, according to the real-estate compaadvances in systems science and application (2016) vol.16 no.2 51 nies, the prices for serial housing in riga have decreased on 8-10 % in a year. the real estate agents werent able to find buyers for three thousand apartments in newly built houses. the number of real estate transactions fell considerably, the income of realtors and developers respectively fell. increasingly, they have started offering discounts on real estate aiming to attract new customers. the construction boom went on recession. the bankruptcies of construction campaigns began. the developers and other builders could hardly to obtain mortgage loans in banks. there passed the first auctions on sale of real estate which was mortgaged earlier in banks. the developers and real estate companies were not in a hurry to reduce the prices hoping that the situation in the market would change and demand for real estate will go up again. the potential buyers, in their turn, didn’t hurry to conclude the real estate transaction, assuming the prices for property to become even lower. such situation has coincided with world financial crisis which much more aggravated a situation in the latvian market. the hopes of developers, realtors and buyers of real estate couldn’t come true, and the fast development of a crisis situation equalized all sides regarding the responsibility for the situation occurred in the market. the predicted trend has been implemented into the actual figures. despite the bank of latvia in 2005 has been repeatedly raising the reserve rate and brought it to 8%, the banks continued to increase the lending capacities. according to the information of the financial and capital market commission, the loan portfolio of latvian commercial banks amounted to 4,380.6 million lats in december 31, 2004, and at the end of 2006-10,872.9 million lats[2]. the most significant increase in the volume of loans amounting to 16,588.9 million lats was in 2008 (see table 2). as the table shows, the growth of credit portfolio of the commercial banks was especially high when the financial crisis had begun. its share in total assets was 71.4% in the end of 2008. in comparison with 2008 the assets of banks have decreased by 2,775.2 million lats or 11.9 % in 2013. thus, the share of credit portfolio in total assets of commercial banks in comparison with 2008 decreased by 13.7 percent points and made 57,7 %. equities in capital of enterprises increased (29.6 %) for the same period. the due from other financial institutions increased on 738.2 million lats (29.6 %). analyzing the changes in total liabilities of commercial banks for the similar period, we can conclude that the balance of deposit accounts had a reverse tendency compared to the amounts of credit debt. thus, in 2013 the deposits amounted to 11,071,5 million lats and increased on 1207 million lats or by 12.5 %, in comparison with 2008, and their share in total liabilities of commercial banks reached 52 valery i. roldugin: principal causes of financial frises in latvia for fast 20 years 51.4 %[1]. table 2 structure of asset of latvian commercial banks from 2008 to 2013(the end of the year, in million lats) 2008 2013 share of items in total assets changes in interest points changes in 2013 compared to 2008 2008 2013 mln. % cash and balances with the bank of latvia 1,325.20 2,008.60 5.7 9.8 1.2 683.4 51.6 due from other financial institutions 2,496.90 3,235.10 10.7 15.8 6.3 738.2 29.6 loans 16,588.90 11,820.30 71.4 57.7 -13.7 -4,768.60 -28.7 bonds and other securities with fixed income 2,015.70 2,104.30 8.7 10.3 1.3 -88.6 4.4 shares and other securities with non-fixed income 51.9 210.6 0.2 1 0.8 158.7 305.8 equities in capital of enterprises 130.7 324.5 0.6 1.6 1 193.8 148.3 fixed assets and intangible assets 234.8 115.8 1 0.6 -0.4 -119 -50.7 other assets 275.9 490.7 1.2 2.4 1.2 214.8 77.9 prepaid expenses and accrued income 123.5 158.2 0.5 0.8 0.3 34.7 28.1 total assets 23,243.30 20,468.10 100 100 -2,775.20 -11.9 5 conclusion an effective banking system is one of the most important conditions of economic development of latvia. it is assumed that latvian economy will be able to implement positively the experience of monetary and credit regulation accumulated in the worlds practice. for the past years the modern two-level banking system was established and developed in the country. gradually, the competitive credit and financial infrastructure was formed, basic elements of which were the commercial banks. some of them have already received a high international rating. the association of the latvian banks turned into the national bank association. at the end of 2014, 27 commercial banks, including ten branches of foreign credit institutions were registered in the republic of latvia (see table 1). jsc ge money bank wound up its business in 2013. moreover, based on the applications submitted by the particular commercial banks, as of 1 january 2014 commercial banks licences were cancelled for jsc unicredit bank which left the baltic market due to changes in the group strategy and sjsc latvijas hipotku un zemes banka which was transformed into sjsc latvijas attistibas finansu institucija altum[1]. the latvian banks slowed the lending down and tried to reduce the credit risks, what had a negatively effect on national economy. the banks tried to increase the liquidity by transferring the bad credits to the structures specially advances in systems science and application (2016) vol.16 no.2 53 established for this purpose. at the same time, there is the unpredictability of actions of the latvian administrative bodies. a main goal of the bank of latvia and the financial and capital market commission is the ensuring of the general sustainability in the monetary and credit markets. thus, exercising all rights defined by the law, the bank of latvia should provide stability of the prices, and the financial and capital market commission should treat everyone who forms the instability in finances and capital market in whole or in its separate sectors by their activity or non-activity. references [1] the bank of latvia: annual report 1994-2014. available at: https://www.bank.lv/en/statistics/monetary-statistics/mfi-balance-sheetand-monetary-statistics [accessed: oct. 12, 2015]. [2] banks statistics. (2015), available at: http://www.lka.org.lv/ru/statistika/. [3] roldugin valery. (2007), international business explanatory dictionary. riga, jumava. [4] roldugin valery. (2006), international economic dictionary of states, riga, jumava. [5] law on financial and capital market commission: lr law. latvijas vstnesis. (2000), nr.230/232, available at: http://www.likumi.lv/doc.php?id=8172 [accessed: sep. 19, 2015]. [6] information about latvian market of real estate. available at: http://www.realty.lv [accessed: apr.11, 2015] corresponding author roldugin valery can be contacted at: vroldugin@hotmail.com microsoft word 7-王永红.doc advances in systems science and applications (2010), vol.10, no.1 issn 1078-6236 international institute for general systems studies, inc. 41-47 multidimensional structural regression model for causal inference under strongly ignorable treatment assignment∗ yonghong wang school of mathematics and computer, harbin university, harbin 150086, china email:wangyonghong418@163.com abstract in the study of epidemiology aetiology, we usually cannot measure exposed effect relative to an individual, but under some assumptions, we approximately replace the exposed effect by estimator of population average causal effect. a multidimensional structural regression model for causal inference is established to estimate the population average treatment effect under strongly ignorable treatment assignment. under the normal distribution, the maximum likelihood estimator for population average treatment effect is proved to be consistent, unbiased and asymptotically normal. keywords strongly ignorable treatment assignment causal inference population average treatment effect multidimensional structural regression model maximum likelihood estimator 1. introduction the objective of epidemiology study is to search aetiology, and to measure its causal effect in the light of quantity, thereby prevent from the occurrence of disease[1]. the case -control study, which attempts to find a contrasted or exposed group which can be compared with the treated or exposed one, is the common method that we employ in the research of epidemiology aetiology. except the difference of being exposed and unexposed, the ideal treated group and the contrasted group are expected to share the rest of the other features, which serves as the guiding principle in the epidemiology research[2-3]. the dummy truth model laid a theoretical foundation for us to demonstrate this principle[4]. let u be the study population, and denote a generic individual in u by u u∈ . a variable e is defined on each u u∈ so that ( ) 1e u = if u is exposed to the casual agent of interest and ( ) 0e u = if u is not so exposed. )(1 uy and )(0 uy indicate the diseased status of u as in the exposed and non-exposed cases respectively. thus we can define the exposed causal effect of the individual u as )(1 uy — )(0 uy .as we know, in any real epdidemilogic study, one does not observe )(1 uy and )(0 uy simultaneously. therefore the casual effect of the exposure to the individual is not attainable. however, in some circumstances, we can get the population average causal effect, which can be defined as 1 0 e(y (u)-y (u)) , while )(⋅e denotes expection or population average (over u ). under the circumstances of random tests, for the fact that the independence of the exposed e and other variants, the result should be e (y1,y0),which indicates that the exposed e is independent of the random vector(y1,y0). thus the population average causal effect is shown as )( 01 yye − = )()( 01 yeye − = )0()1( 01 =−= eyeeye .in other words, the causal effect can be estimated by the difference of the exposed population average diseased condition the non-exposed population average diseased condition[5]. the epidemiology, however, is an ∗this work has been supported by the science development foundation of harbin university(hxkq200701). wang: multidimensional structural regression model for causal inference under strongly ignorable treatment assignment 42 observatory science in nature, so the above formula is not to be established which necessitates the assumption of the state of being ignorable[5-6] . definition 1.1 given the covariant x , the exposed e is called being strongly ignorable. if (1)(y1,y0) xe , i.e. the observed covariant x is given, the response variant is independent of the exposed e. (1) (2) 1)pr(0 <=< xee , i.e. for each and every x, they may receive various treatment. when the e is strongly ignorable, by getting the matching-differencing-averaging between the exposed and non-exposed groups, we can get the unbiased estimate of population average causal effect[7]. but the above method is complicated in calculation and the additional sparse-data problems, excess covariant, though few, may arise. in addition, the method of ranking the individuals with the estimated propensity score is an advisable way to solve the problem, but the calculation is still very complicated[8]. [9] under strongly ignorable treatment assignment, a structural regression model for causal inference is established to estimates the population average treatment effect. they are the series of results achieved. with the response variant being one dimensional. the paper will discuss the inference problem with multi-dimensional response variants. 2. multidimensional structural regression model definition 2.1 in condition of strongly ignorable treatment assignment, denote multidimensional structural regression model as follows: ⎪⎩ ⎪ ⎨ ⎧ σ ++= tjtjttjxxtj tjtjtttj exvnenx exy oft independen is ),,0(~),,(~ μ βα (2) here t=0,1 means two treatment; tjy is the random vector of order 1×p ,indicated the response variant of the jth observation for the tth treatment; tjx is the covariant vector of order 1×k ; tje is the random error vector of order 1×p ; tα is the parameter vector of order 1×p ; tβ is the parameter matrix of order kp× ; xμ is the parameter vector of order 1×k ;the covariant matrix tv和xς of tjx and tje are both positive definite. under model (2) ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ + =⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ xtt x tj tj y x e μβα μ , ta ' ' tj x x t tj t x t x t t x cov y v β β β β ⎛ ⎞⎛ ⎞ σ σ = ⎜ ⎟⎜ ⎟ σ σ +⎝ ⎠ ⎝ ⎠ we can easily find txt va ⋅σ= ;because xς and tv are both positive definite, we can know 0≠σ x , 0≠tv ;so 0≠ta ;therefore 1− ta is valid, and we denote it as bt .again based on the principle of block matrix inversion, we can get bt= ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −+σ −− −−− 11 1'1'1 ttt tttttx vv vv β βββ so ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ +⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ t xtt x tj tj an y x ,~ μβα μ where: t=0,1 ; j=1,……,nt . we also get observation logarithm likelihood function of the sample ( )tjtj yx , with the advances in systems science and applications (2010), vol.10, no.1 43 capacity nt (t=0,1; j=1,……,nt ) ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −− − ⎟⎟ ⎠ ⎞ ⎜⎜ ⎝ ⎛ −− − +=− ∑∑∑ = == xtttj xtj t t n j xtttj xtj t tt y x b y x anl t μβα μ μβα μ ' 1 0 1 1 0 lnln2 furthermore, 1 1 ' 1 ' 1 0 0 1 2 ln (ln ln ) ( 2 tn t x t tj x tj tj x x t t j l n v x x x μ− − = = = − = σ + + σ − σ∑ ∑∑ tjtttjtjttttjxxx yvxxvx 1''1''1' 2 −−− −+σ+ βββμμ (3) )22 1'1'1'1'' ttttttjtjttjttttj vvyyvyvx ααααβ −−−− +−++ 3. conclusion 3.1 two lemmas in order to get the likelihood estimates of all parameters in model (2) and the population average causal effect, we give the following two lemmas, of which lemma3.1 is shown at [10]. first, suppose x is the matrix of order nm× , )(xf is the real-valued function of matrix x , x xf ∂ ∂ )( indicate real-valued function )(xf settling partial derivation for each element of matrix x . thus the following lemmas are valid: lemma 3.1(1) '')( ba x axbtr = ∂ ∂ (2) '' ' )( xbaaxb x axbxtr += ∂ ∂ (3) '1 )( ln −= ∂ ∂ x x x of which a and b are two matrixes that can make matrix calculation with x , ( )tr x means the trace of matrix x . lemma 3.2 if model (2) is valid, when the covariant matrix and the random error matrix are independent of each other, then ∑ = = tn j tj t t x n x 1 1 and ∑ = −−= t tt n j ttjttj t xx xxxx n m 1 '))((1 ' , ∑ = −−= t tt n j ttjttj t xe xxee n m 1 '))((1 ' are independent of each other, where ∑ = = tn j tj t t e n e 1 1 . proof: first to prove tx and ' tt xx m are independent. wang: multidimensional structural regression model for causal inference under strongly ignorable treatment assignment 44 let ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ ⎛ = ktntntn kttt kttt t ttt xxx xxx xxx x 21 22221 11211 set '1 tt x n x = 1 also '' 1(1 ' t t t t xx x n x n m tt −= 11 ' )( t t n x 1 − 11 ' tx ) = '1 t t x n (i t n nt 1 − 11 ' ) tx which 1 is the column vector of 1×tn . we can do the vector straight operations to matrix tx , then ⊗σ= xxveccov ))(( i tn , also, as for tx and each at the element ijm in ' tt xx m follow tx = tn 1 (ik⊗ 1 ' ) )( txvec ijm = ')( txvec [ ⊗+ )( 2 1 ' ijij ee (i t n nt 1 − 11 ' )] )( txvec of which⊗ indicates the kronecker product of matrix, ije indicates matrix of order kk × with (i,j)th is 1, and the rest is 0. now we can draw the following conclusion by making use of the characteristic of multidimensional normal random variant and the kronecker product of the matrix: tn 1 (ik⊗ 1 ' )( ⊗σ x i tn )[ ⊗+ )( 2 1 ' ijij ee (i t n nt 1 − 11 ' )] = t ijijx n ee 1[)]( 2 1[ ' ⊗+σ 1 ' (i t n nt 1 − 11 ' )]=0 thus we can prove ijm and tx are independent, and it’s valid to any i,j=1,……,k. so we have proved ∑ = = tn j tj t t x n x 1 1 and ∑ = −−= t tt n j ttjttj t xx xxxx n m 1 '))((1 ' are independent of each other. suppose ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ ⎛ = ptntntn pttt pttt t ttt eee eee eee e 21 22221 11211 ),( ttt exz = ,in the same way , and by mutual independence of tx and te , we can prove tx and ' tt xe m are also independent of each other. lemma3.2 is proved! advances in systems science and applications (2010), vol.10, no.1 45 3.2 estimate of parameter theorem 3.1 suppose model(2)is valid, we can get the maximum likelihood estimates of all the parameter: txxxytt xmmy tttt 1)(ˆ '' −−=α ; 1)(ˆ '' −= tttt xxxyt mmβ ; ∑∑ = =+ = 1 0 110 1ˆ t n j tjx t x nn μ ; ' 1 0 110 )ˆ)(ˆ(1ˆ xtjx t n j tjx xx nn t μμ −− + =σ ∑∑ = = ; ∑ = −−−−= tn j tjtttjtjtttj t t xyxy n v 1 ')ˆˆ)(ˆˆ(1ˆ βαβα of which ∑ = = tn j tj t t y n y 1 1 , ∑ = −−= t tt n j ttjttj t xy xxyy n m 1 '))((1 ' ; the form of tx and ' tt xx m is shown in lemma3.2. proof: first, we get the score function of tα : ∑ = −−− +− ∂ ∂ = ∂ −∂ tn j ttttttjttttj tt vvyvxl 1 1'1'1'' )22()ln2( ααααβ αα =∑ = −−− ∂ ∂ + ∂ ∂ − ∂ ∂tn j ttt t tttj t ttttj t vtrvytrvxtr 1 1'1'1'' )]()(2)(2[ αα α α α αβ α = )22( 11 1 tttjtt n j vxv t αβ −− = −∑ =0 we can get tttt xy βα −= . and the score function of tβ is as follows: 0)222()ln2( 1 '1'1'1 =+−= ∂ −∂ ∑ = −−− tn j tjtttjtjttjtjtt t xvxyvxxvl αβ β simplified as 0)( 1 ''' =+−∑ = tn j tjttjtjtjtjt xxyxx αβ fit tβ into above formula of tα , we set txxxytt xmmy tttt 1)(ˆ '' −−=α . in the same way, we can get the maximum likelihood estimates of the other three parameters. theorem3.1 is proved! in fact, as for the multidimensional structural regression model with the response variants as shown in model (2), we can see from the above demonstration that the maximum likelihood estimates of parameter tα and tβ do not depend on the selection of the covariant matrix and random error matrix. moreover, we can also see that if we want to estimate the parameter that measure relationship between the covariant and the jth(j=1,……,nt)response variable, we should only note the jth response variable. i.e. the model(2) can be divided into p one dimensional models, which is evident in jinhua’s paper(2000). 3.3 population average causal effect definition3.1 provides the covariant x, if the assumption of strongly ignorable condition (1) is fulfilled, then treatment e=1 to e=0 average causal effect ate is ))()(()()( 0101 xxyeyeeyeyeate x =−=−= wang: multidimensional structural regression model for causal inference under strongly ignorable treatment assignment 46 xx x xxe xxeyexxeyee μββααβαβα )()()( )),0(),1(( 01010011 01 −+−=−−+= ==−=== theorem3.2 suppose model (2) is valid, and treatment assignment variant e is strongly ignorable, then (1) the maximum likelihood estimator of population average treatment effect ate is ate = )( ˆˆ 01 10 0110 01 xx nn nn yy − + + −− ββ (2) ate is consistent unbiased estimator of ate. (3) with the case of the large sample, eta ˆ is close to normal distribution n(ate, γ). of which ' 0 1 0 1 01 0 1 0 1 ( ) ( )xv v n n n n β β β β− σ − γ + + + proof: (1) with model (2), if treatment assignment variant e is strongly ignorable, then the population average treatment effect ate= xμββαα )()( 0101 −+− based on theorem 3.1 and the invariance of maximum likelihood estimator, we can know ate xμββαα ˆ)ˆˆ()ˆˆ( 0101 −+−= )( ˆˆ 01 10 0110 01 xx nn nn yy − + + −−= ββ (2) consistency: we know based on the model (2) ttttt exy ++= βα , t=0,1 again based on theorem3.1, we get 1)(ˆ '' −= tttt xxxyt mmβ = 1))(( '' −+ tttt xxxet mmβ based on the law of large numbers and the nature of the random variant order converging by probability, we can know x p xx tt m σ⎯→⎯' , 0' ⎯→⎯p xe tt m then t p t ββ ⎯→⎯ˆ in the same way, x p tx μ⎯→⎯ , xtt p ty μβα +⎯→⎯ so far, again based on the nature of the random variant order converging by probability, we get pate ate⎯⎯→ , i.e. ate is consistent estimate of the population average treatment effect ate. unbiased:through the demonstration of the consistency, we know 1 1 ' )]())((1[ˆ ' − = ∑ −−+= tt t xx n j ttjttj t t mxxee n ββ whereupon })]())((1[{)ˆ( 1 1 ' ' − = ∑ −−+= tt t xx n j ttjttj t t mxxee n ee ββ = })]())((1[{ 1 1 ' ' ttxx n j ttjttj t xt xxmxxee n ee tt t =−−+ − = ∑β = tttxt xxe ββ ==+ )0( therefore xtttye μβα +=)( , so based on lemma3.2, we set advances in systems science and applications (2010), vol.10, no.1 47 0 1 1 0 1 0 1 0 0 1 ˆ ˆ ( ) ( ) ( ) ( _ )n ne ate e y y e e x x n n β β+ = − − + = atex =−+− μββαα )()( 0101 (3) asymptotically normality: based on above demonstration, we know ( )e ate ate= also because of 0 1 1 0 1 0 1 0 0 1 ( ) ( )n nate y y x x n n β β+ ≈ − − − + = ] )( [] )( [ 00 10 010 011 10 011 1 ex nn n ex nn n + + − +−+ + − + ββ α ββ α moreover x t t n xvar σ= 1)( , t t t v n evar 1)( = therefore 1 1 ' 01012 10 1 11 10 011 1 1)()( )( ) )( ( v nnn n ex nn n var x +−σ− + =+ + − + ββββ ββ α 0 0 ' 01012 10 0 00 10 010 0 1)()( )( ) )( ( v nnn n ex nn n var x +−σ− + =+ + − + ββββ ββ α so ' 1 0 1 0 1 0 1 0 0 1 1 1 1( ) ( ) ( )xvar ate v v n n n n β β β β= + + − σ − γ + so based on the central limited theorem, we know ~ ( , )ate an ate γ . theorem 3.2 is proved! references [1] rothman,k.j. modern epidemiology[m]. boston:litte, brown and company, 1986: 56. [2] guo,j.h. causal inference and confounding phenomenon. journal of northeast normal university, 2002, 34: 25-41. [3] habtemariam,t. and tameru,b. etc. epidemiologic modeling of hiv/aids: use of computational models to study the population dynamics of the disease to assess effective intervention strategies for decision-making. advances in systems science and applications, 2008, 8(1): 35-39. [4] tao,q.s. and li,l.m. aetiology effect model based on dummy truth theory. chinese journal of epidemiology, 2002, 23(1): 60-62. [5] rosenbaum,p. and rubin,d.b. the central role of the propensity score in observational studies for causal effects. biometrika, 1983, 70: 41-55. [6] holland,p.w. confounding in epidemiologic studies. biometrics, 1989, 45: 1310-1316. [7] kong,f.l. and wang,y.f. certain martingale methods in parameter estimation. advances in systems science and applications, 2004, 4(1): 1-6. [8] yu,j.w. and cheng,d.h. etc. factor analysis methods in the statistical literature of chinese medicine diabetes research applications. advances in systems science and applications, 2007, 7(2): 161-166. [9] jin,h. and fang,j.q. structural regression model for causal inference under strongly ignorable treatment assignment. journal of south china normal university, 2000, 4: 7-12. [10] wang,s.g. etc. theory of linear model[m]. beijing: science press, 2004: 44-47. microsoft word 8 lianhui jiang, guang yang--modeling and analysis of mechanical properties.doc 264-269 advances in systems science and applications (2011), vol.11, no.3-4 issn 1078-6236 international institute for general systems studies, inc modeling and analysis of mechanical properties for fdm parts* lianhui jiang and guang yang school of mechanical and electronic engineering, wuhan university of technology, wuhan 430070, china abstract fused deposition manufacturing has the potential to fabricate parts with locally controlled properties. modeling and analysis of the mechanical properties of the fdm parts are the foundation to produce this new class of products. a mechanics model is established for manufacturing fdm parts with required mechanical properties. a series of equations is proposed to determine the elastic constants that are used in the model. the results are evaluated by experiments. keywords solid freeform fabrication fused deposition manufacturing locally controlled properties 1.introduction solid freeform fabrication (sff), also referred to as rapid prototyping (rp) or layered manufacturing (lm), is a class of technologies that build objects in an additive fashion directly from a computer model. the ability of sff technologies to fabricate parts with locally controlled properties creates opportunities for manufacturing a whole new class of parts with functionally gradient structures [1,2,3]. the functionally graded structures could be useful in controlling the mechanical properties of parts and tools on a local scale. fused deposition manufacturing (fdm) has the potential to produce parts with locally controlled properties by changing deposition density and deposition orientation. in the fdm process, the deposit head that extrudes a semi-molten filament through a heated nozzle in a prescribed pattern follows the trajectory defined by the chosen deposition strategy of the layer. after the semi-molten material is deposited onto the worktable, it begins to cool and bond to the neighboring layer. the bonding between the individual roads of the same layer and of neighboring layers is driven by the thermal energy of semi-molten material. fdm processes use several types of materials, such as acrylonitrile-butadiene-styrene (abs), investment casting wax, ceramics and metals [5,6], to build conceptual or functional parts. most of the existing research efforts in fdm have been primarily directed to the development of new materials [2]. research efforts were also made to characterize fdm part quality. dutta and his associates investigated the relationship between deposition strategies and the resulting stiffness of fdm parts. matas characterized the meso-structural features of unidirectional extruded material as a function of fdm process variables. longmei li researched the relationship between parts’ properties and deposition parameters of fdm. [2,7] mechanical properties of fdm parts are governed by their meso-structures, which are determined by manufacturing parameters and bonding intension between filaments. manufacturing parameters include width of filaments, layer thickness, deposition orientation and gap sizes between filaments. by modulating manufacturing parameters, fdm processes can produce parts with desired properties. among the manufacturing parameters, deposition directions in layers and gap sizes between filaments are the most important parameters to control the mechanical properties. thus, it is essential to establish mechanical models of fdm parts in * this research has been supported by the national natural science foundation of china (60904073) advances in systems science and applications (2011), vol.11, no.3-4 265 relation to these two manufacturing parameters. the models can then be used to design parts with desired mechanical properties. the remainder of this article is organized as follows. section 2 is a study of the mechanical properties of fdm parts. experimental analysis is given in section 3. section 4 summarizes the conclusions of the paper. 2. modeling and analysis of mechanical properties of fdm parts 2.1 mechanical models. the fdm parts are composites of abs filaments, bonding between filaments, and voids. as composite materials, they can be viewed and analyzed at different levels and on different scales, depending on the particular characteristics and behavior under consideration. fdm parts are composed of layers, so lamination theory can be used to analyze the mechanical behaviors of fdm parts [8,9]. a unidirectional filament fdm part, fig. 1 definition of principal material axes and loading axes for one layer which can be regarded as a unidirectional fiber-reinforced composite, contains three orthogonal planes of material property symmetry, and is classified as an orthotropic material[8]. the coordinate system for analyzing mechanical behaviors of one layer of fdm parts was established as follows. two right-handed coordinate systems are defined, namely, 1-2-z and x-y-z, as shown in figure 1. both 1-2 and x-y axes are in the plane of the lamina, and the z-axis is normal to this plane. in the 1-2-z systems, axis 1, along the filament length direction, represents the longitudinal direction of the lamina. axis 2 is normal to the fiber length and represents the transverse direction of the lamina. in the x-y-z system, axes x and y represent the loading directions. the angle between positive axis x and axis 1 is called the fiber orientation angle, represented by θ . a laminate is constructed by stacking a number of laminas along the z direction. it should be pointed out that axes 1 and 2 are constant from one lamina to the next in a unidirectional laminate. for a general orthotropic ( 0≠θ or °90 ) lamina, the strain-stress relations can be expressed in matrix notation as (1): [ ] . 662616 262212 161211 ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ = ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ = ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ xy yy xx xy yy xx xy yy xx s sss sss sss τ σ σ τ σ σ γ ε ε (1) where [ ]s represents the compliance matrix for the lamina. the elements in [ ]s matrix can be expressed as (2): x yz 2 1 266 jiang: modeling and analysis of mechanical properties for fdm parts .cossin)22(cossin)22( .cossin)2(cossin)22( .coscossin)2(sin1 .cossin)()cos(sin .sincossin)2(cos1 3 661222 3 66121126 3 661222 3 66121116 4 22 22 6611 4 1122 22 662211 44 1212 4 22 22 6612 4 1111 θθθθ θθθθ θθθθ θθθθ ν θθθθ ssssssms ssssssms ssss e s ssss e s ssss e s y x yy yy xy xx −−−−−=−= −−−−−=−= +++== −+++== +++== (2) .111 ).cos(sincossin)422(21 12 66 22 22 22 21 11 12 2112 11 11 44 66 22 6612221166 g s e s ee ss e s sssss g s xy ==−=−=== ++−−+== νν θθθθ where 11e is longitudinal young’s modulus, 22e is transverse young’s modulus, 12ν is major poisson’s ratio and 12g is shear modulus. from (1) it appears that there are six elastic constants that govern the stress-strain behavior of a lamina. however, a close examination of these equations would indicate that these six elements are functions of the four independent elements. therefore, the four elastic constants can be used to determine the fully populated matrix of [ ]s . a laminate is constructed by stacking a number of laminas in the thickness direction. therefore, in macro mechanical approaches, the stiffness matrix of a laminate is composed of properties of every lamina according to the lamination theory [8], while the in-plane elastic behavior of a unidirectional lamina may be fully described in terms of four elastic lamina properties. for the design purpose, it is desirable to have reliable predictions of lamina properties as a function of constituent properties and geometrical characteristics. one of the objectives of composite analysis is to obtain such relationships. 2.2 modeling of elastic constants for fdm parts. when loaded in the longitudinal direction, a unidirectional fdm part is an aggregate of abs filaments. longitudinal young’s modulus 11e can be predicted by the rule of mixtures. therefore, longitudinal young’s modulus should be as follows: .)1( 111 ee ρ−= (3) where e is young’s modulus of abs filament and 1ρ is the area void density in the plane normal to filaments. the area void density can be calculated using the following equation: .1 2 11 1 dhv av −=ρ (4) where 1v is the velocity of the material supply system, 1a is the section area of the supply filament, 2v is the velocity of the manufacturing head, d is the distance between two filaments of the same layer and h is the thickness between two layers. advances in systems science and applications (2011), vol.11, no.3-4 267 in the case of transverse normal loading, the bonding among the filaments is the load carrier. because of imperfect bonding and the irregular interface between filaments, an accurate equation to calculate transverse modulus is impossible. an approximate equation is used to calculate transverse modulus: .22 ee ζ= (5) where 1~0=ζ is the coefficient that is decided by distance between filaments, bonding status, etc. to calculate ζ , the geometry of the cross section should be studied. fig. 2 shows the cross section image of road geometry [4]. fig. 3 shows the theoretical model of two adjacent roads. from the cross section images of road geometry, the shape of the hole between filaments can be assumed as lozenge. the horizontal cater-corner a of the lozenge can be calculated as follow: ).()32( dwra −−−= (6) where w is the deposition width as shown in fig. 3. the area of the lozenge is equal to the area of the cross section of the hole. so the perpendicular cater-corner of the lozenge b can be calculated as follow: ./2 1 adhb ρ= (7) the coefficient ζ can be calculated using the following equation: [ ] . 2)(/2 2/ 0 dxbxbhaahad d a ∫ +−+− =ζ (8) similarly, the prediction of the in-plane shear modulus takes the following form: .12 gg ζ= (9) the in-plane poisson’s ratio, 12ν , is the same as that of abs filaments, ν , as follows: fig. 2 image of cross section of road geometry fig. 3 theoretical model between two adjacent roads f w r table 1 fdm manufacturing parameters settings parameter value nozzle radius (r) 0.3[㎜] layer interval (h) 0.15[㎜] road width (w) 0.58[㎜] head temperature 220[℃] envelope temperature 60[℃] 268 jiang: modeling and analysis of mechanical properties for fdm parts .12 νν = (10) 2.3experimental analysis to evaluate the coefficient value in the equation, the experimental analysis was carried out. all the specimens were made using a fdm machine with the manufacturing parameters as shown in table 1. the designs of unidirectional specimens are shown in figure 4. the specimens with fiber orientation of °° 45,0 , and °90 , were tested to determine the elastic constants for different deposition orientation. the tests followed the standard procedures of gb 1040-79. in the 00 unidirectional specimen tests, the strain gages were mounted in the longitudinal and the transverse directions to determine the longitudinal young’s modulus 11e and poison’s ratio 12ν . transverse modulus 22e can be derived from the results of the °90 unidirectional specimen tests. with the resultant young’s modulus 45 xxe in the x direction, and with the other constants of ,, 1112 eν and 1222 ,ge can be determined as follows [8]: . 2114 1 11 12 2211 45 12 eeee g xx ν +−− = (11) the theoretic and experimental results are listed in table 2. table 2 theoretical results with experimental results 3.conclusion sff processes must become broadly used for manufacturing but not just prototyping. for fabrication of parts with locally controlled properties, a model of in-plane mechanical properties of fdm parts is established. a set of equations for calculating the elastic constants was found. the constants were used to determine the constitutive models of fdm parts. according to the model, different deposition densities, orientations and their combinations can be selected in producing constant, units experimental results theoretical estimation difference, % e11 [ mpa] 2238.7 2267.6 1.29% e22 [mpa] 2189.7 2156 1.54% 12ν 0.4211 n/a n/a g12 [mpa] 800 785.22 1.85% 1 20 0 11e 1 20 0 22e 1 20 0 45 xxe fig. 4 geometry of unidirectional i advances in systems science and applications (2011), vol.11, no.3-4 269 the parts with required stiffness properties. this research is useful for exploring fdm technology fabricating parts with locally controlled properties. 4.acknowledgement the authors would like to thank the national nature science foundation of china (grant no, 60904073) for its supporting for this research and to thank the people involved in the work. without their efforts and input, the completion of the research would not have been possible. references [1] cho.w, sachs,e.m., patrikalakis.n.m., cima.m.j., jackson.t.r, etc. “methods for distributed design and fabrication of parts with local composition control.” proceedings of the 2001 nsf design and manufacturing grantees conference. jan. 2001, tampa, florida. [2] longmei li, qiao sun, celine bellehumeur, peihua gu. “modeling and analysis for fabrication of fdm prototypes with locally controlled properties.” journal of manufacturing processes; 2002;vol4,no.2;pp129-141 [3] t.r. jackson, h. liua, n.m. patrikalakisa,u, e.m. sachsb, m.j. cimac “modeling and designing functionally graded material components for fabrication with local composition control” materials and design.2002, vol20, no.2-3,pp63-75 [4] wenbiao han, mohsen a. jafari, stephen c. danforth, ahmad safari, 2002, “tool path-based deposition planning in fused deposition processes”, transactions of the asme. journal of manufacturing science and engineering (v.124). pp462-472 [5] agarwala, m.k., van weeren, r., bandyopadhyay, a., whalen, p.j., safari, a., and danforth, s.c., 1996, ”fused deposition of ceramics and metals: an overview,” proceedings of solid freeform fabricaiton sypposium, austin, tx, pp.385-390. [6] jafari, m.a., han, w., mohammadi, f., safari, a., danforth, s.c., and langrana, n.a., 2000, “a novel system for fused deposition of advanced multiple ceramics.” rapid prototyping journal, 6, no.3, pp.161-174. [7] li, longmei (2002), “analysis and fabrication of fdm prototypes with locally controlled properties.” phd dissertation, calgary, alberta, canada: dept. of mecanical and mfg. engg, univ. of calgary. [8] mallick, p.k., 1988, fiber-reinforced composites: materials, manufacturing, and design, marcel dekker, inc, us. [9] p.gu, l.li, 2002, “fabrication of biomedical prototypes with locally controlled properties using fdm”, annals of the cirp, vol51(1), pp 181-184. advances in systems science and applications (2014) vol.14 no.1 50-65 investigation of the different type models for engine test data modelling m.h. wu school of technology, university of derby, marketon st., derby, de22 3aw, uk abstract the engine test bed has been used widely in automotive industry to test the protetypey developed engine in order to understand the engine operation process and find the optimum parameters for the engine operation. as a result, thousand and thousands data will be obtained from engine test bed.for this purpose. it is a time consuming task to get the data and it is hard to analyse those data directly from the huge number of the data. the best way to reduce the tested data from the test bed is to create a mathematics model from the limited engine tested data. the paper will list three different models used for the engine test data modelling, polynomial model, combined exponational model and neural network model. the principle and equations of each model has been introduced in the paper. the experiment results of each model have been listed as well. the results from last two developed models show that the models provide the achievement which can meet the customers requirement. keywords engine test data, engine performance, engine data modelling, neural network 1 introduction the engine test bed has been used widely in automotive industry to test the protetypey developed engine in order to understand the engine operation process and find the optimum parameters for the engine operation. as a result, thousand and thousands data will be obtained from engine test bed.for this purpose. it is a time consuming task to get the data and it is hard to analyse those data directly from the huge number of the data. the best way to reduce the tested data from the test bed is to create a mathematics model from the limited engine tested data. the model will be used to analyse the engine operation and to obtain the optimum operation data. figure 1 shows the structure of the engine test bed (etb). the newly developed or existed engines should be tested on it. fig.1 shows the format of the tested data from etb which are used to create the new model. the controled input variables and the responding out put are shows in fig.2. it includes: ignition time, ex valve timing, inlet valve timing , bmep and speed as the input variable and the bsfc as the output. there are three models developed for the engine data modelling. it is polyadvances in systems science and applications (2014) vol.14 no.1 51 fig.1 structure of engine test bed nomial model, combined exponational model and neural network model. fig.2 the form of the tested data from the etb 2 employing the polynomial model for engine test data modelling 2.1 the polynomial approximation model (1)shows the aplication of the polynomial approximation model in engine test data modelling. it called “quadratical model”(qm) which is denoted by voigt, lechner and hochschwarzer[1-3]. the qm is: y = xtqx+ atx+ b (1) here: y is an objective vector, x is an input variable vector, q is co-efficient matrix, a& b are constant vectors, 52 m.h. wu:investigation of the different type models for engine test data modelling the co-efficient of matrix q and the vectors a and b are calculated by leastsquares-method to minimize the deviation of measured values from the model values. 2.2 analysis of the application of the polynomial model this model has the advantage of simplicity, but is imperfect for real application. fig.3 shows the example which indicates the model’s imperfect. (a) draw by the original engine tested data (b) draw by the data calculated by qm fig.3 3d surface draws by the different data fig.3a shows the relationship between the bsfc and two input variables by the experiment data directly. fig.3 b shows the same relationship using the model qm. it is obviously that fig.3(a) is totally different to fig.3(b). in fig.3a it shows that the output will be varied as a irrgular wave when two inputs are varied. the fig.3b only provide a smooth raised convex area. in one word, the qm model can not trace the features in fig.3a. it is clearly that qm is not a suitable model to be used for engine parameters modelling and optimasation. the data from qm provides the large errors and losts nearly most of data features. another disadvantage of using quadratical equations is that the model only can be used to the case which has one convex (or concave) data set. 3 employing the combined exponational model for engine test data modelling 3.1 the combined exponational model (2) shows the aplication of the combined exponational model in engine test data modelling[4]. f(x1, x2, . . . xn) = a+ n∑ i=1 bixi + p∑ j=1 (hje(x1)e(x2) . . . e(xn)) (2) here: e(x) = exp(−8(x−m)2/w2); h : the heigth of the peak, it could be “+” or “-”; advances in systems science and applications (2014) vol.14 no.1 53 m : the position of the peak along the axis; w : the width of the peak; n = 5, (the model dealings the variables; f(x1, x2, ...xn) : break specific fuel consumption (bsfc); x1 : ignition time variable; x2 : inlet valve opening degree variable; x3 : exhaust valve opening degree variable; x4 : speed; x5 : toque (bmep); p : the number of peaks in nd space; it assumes p = 10; a : a constant; bi : a slope coefficient respect to input xi; fig.3a indicts that the output bsfc could be formed by constant element, linear element and multiconvex or concave peaks. the convex or concave peaks on the 3d surface can be expressed by the function of exponatioal. as a result, the (2) can be explained as follow: bsfc = constant + linear function + exponational function; 3.2 set up the coefficients of the combinated exponational model if the equation 2 has been set by 20 peaks (j=20) and 5 input variables (i=5), the total number of the coefficients of the equation, bi, hj , mi,j and wi,j , should be as following table 1. there is no suitable mathematics methods to set up the exactly values to 225 table 1 the number of the parameters of the combined exponational model with 5 inputs, 20 peaks coefficient bi hj mij wij total the numbers 5 20 100 100 225 coefficients through the existed tested data. as a result, genetics algorithm are used to define the appropriated values for each coefficient by the existed tested data. a vb based program is developed for this purpose. the software will be used to obtain the appropriate values of the coefficients in (2) by using the follow operations in ga: – random selection of the chromosome within certain range, – reproduction – crossover – evaluation – mutation in order to running the ga, the follow assumptions and definitions are re54 m.h. wu:investigation of the different type models for engine test data modelling quired: gene: the genes are used to form the chromosome in ga application. as a result, the number of the gene should be equal to the total numbers of the coefficients in the (2). for above example in table 1, the numbers of gene should be 225. population: the population is the numbers of the chromosome. it can be set from 1 to 100 in the software. the rule of setting up is that the more coefficients is , the larger population should be. in the case of table 1, the population is settled equal to 100. number of keep best members of the population: this is very important selection. if it is too small, too many poor chromosomes will be remained in the operator. on the other hand , it is easily tripped into the local best point. for the case in table 1, a range 10-20 is suggested. crossover method: crossover method can be selected from alternative methods, single point cross over, two points crossover, etc. type of mutation: random mutation hill climb and directional hill climb methods are available in the software. the first one is normally used for the fitness function is highly discontinuous. the 2nd one is for the continuous fitness function. our fitness function for ga is continuous so that the directional hill climb method is adopted. mutation probability of population: this percentage is employed to avoid trapped in local best. for engine test data modelling, the percentage is 10%. mutation probability of genes: if the percentage is too high, many “wild” mutations which have very poor fitness will be involved. if it is too low, some necessary genes will not be involved in crossover operator to improve the solution. it is normally 8-10%. fitness function in evaluation operation: fitness function is very important for chromosome parent selection in ga operation equation 3 is the fitness function used in the ga evaluation. min f(x1, x2, ...., xn) = ∑a+ n∑ i=1 bixi + p∑ j=1 hje(x1) · e(x2) · · · e(xn)− y0 2 (3) defined the ranges of variation of each variable: the ranges of variation of each variable in the engine test bed is settled as table 2, which will constraint the variables of x1, x2,...,xn; in (2) during the ga operation. 3.3 output results from the combined exponationa model as mentioned before, the parameters of the (2) are obtained from the existed test data from etb by ga system. as soon as the parametes is defined, the equation advances in systems science and applications (2014) vol.14 no.1 55 table 2 experimental data ranges variables range x1 input: speed 1237 to 4253 x2 input: bmep 0.9 to 6.8 x3 input: vanosin 80 to 132 x4 input: vanosex 70 to 135 x5 input: igtiming 4.5 to 51 y output: bfsc average: 346 2 can be used to calculate the engine output bsfc. fig. 4 shows the 3d surface drawings by the ordinary test data and the data calculated from (4) it is obviously that the major convex or concave peaks in (c) are repeated in (d). it means that the new model can generat the major features of the black box system (engine perfomance) and it is much more better than using the quadratical model in engine test data modelling. (a) draw by the original tested data (b) draw by the calculated data from new model fig.4 comparing the original tested data and the calculated data using the new model 4 employing the neural network model for engine test data modelling[56] 4.1 develop the radial basis function (rbf) networks for engine calibration modelling task[5] in 2003, dr. lin has published a paper. in the paper the neural networks used on the engine test data modelling was mentioned. on the application, radial basis function (rbf) networks is used for engine calibration modelling task. rbf networks are universal approximates, that is, given a network with enough hidden layer neurones, they can approximate any continuous function with any arbitrary accuracy [7-8]. an rbf is one whose output is symmetric around an associated centre µj . it is 56 m.h. wu:investigation of the different type models for engine test data modelling generally described as: ϕj(r) = ϕj(r)(∥x− µj∥); x ∈ rn; r ≥ 0 (4) where ϕj(r) is a continue function on(0,∞)and ∥ · ∥ denotes the euclidean norm. linear combination of rbfs represents a wide classes of functions: yj(x) = m∑ j=1 ωijϕj(∥x− µj∥) + ωi0 (5) the bias wi0 compensates for the difference between the average value over the fig.5 rbf network structure data set of the basis function activation and the corresponding average value of the targets. the rbf network is depicted as shown in fig.5. the bias wi0 can be incorporated into the summation by introducing an extra basis function ϕ0 and setting its activation to unity. thus the (2) can be written as yj(x) = m∑ j=1 ωijϕj(∥x− µj∥) (6) follows are the typical functions of the rbf: • the gaussian function: ϕ(r) = exp(−r2/2) (7) • the thin plate spline function: ϕ(r) = r2 × log r (8) advances in systems science and applications (2014) vol.14 no.1 57 • the multiquadric function: ϕ(r) = (r2 + 1)1/2 (9) • the inverse multiquadric function: ϕ(r) = 1 (r2 + 1)1/2 (10) • the pseudo cubic spline function ϕ(r) = r3 (11) • the logarithmic function: ϕ(r) = log(r2 + 1) (12) where r is a non-negative number and is the scaled distance from the input vector x to the rbf centre , which is defined as : rj = √√√√ n∑ l=1 (xl − µlj)2 σ2 j (13) in (13), n is the dimension of input vector and σj is scale factor or width of rbf. in engine calibration modelling task, all of six rbfs described above are applied. 4.2 develop neural network modelling tool for engine calibration modelling task[6] the fig.6 shows the structure of the neural network modeling tool used on the engine test bed. when the tested data is ready, they will be sending into the nn models tool box. in the nn tool box, there are three different types and totally ten nn structures. in it, the total tested data will be divided into three group. the first group data will take 50 % out of the total tested data. the second and third groups will take 25% out of the total data respectively. the first group data will be used to train the neural network and second group data will used to validate the nn model. final group data will be used to test the nn model. after nn tool box, the best fitted nn structural model is selected and sends to optimization section to defined the optimization data for engine management unit. in the optimization section, there are two options for optimization purposes according to the different requirement. soga is the model for single objective optimization and moga is one for multi objectives optimization. 58 m.h. wu:investigation of the different type models for engine test data modelling fig.6 the structure of the optimisation system of the engine performance as shown in figure 6 the nn model tool is employed to select a best suitable nn model out of ten nn models for the subsequence operation of optimization. fig.7 shows the common structure of the neural network in nn model tool and it is under the following assumptions: • it is a three layers nn model. • the numbers of the hidden layer note is defined by 2n + 1: here n is the number of inputs. the nn model tool includes three different type of neural networks, multi-layer perceptron (mlp); radial basis functions (rbf) and bar function (bar). a. mlp. the output of the mlp is denoted by yk(x) = m∑ j=0 f( n∑ i=0 xiωji)ωkj = m∑ j=1 f( n∑ i=1 xiωji + ωj0)ωkj + ωk0, k = 1, . . . , l where ωji and ωkj are the input-hidden weight and hidden-output weight, respectively. f is the activation function which has two types: (a) logistic function f(α) = 1 1 + e−α (15) advances in systems science and applications (2014) vol.14 no.1 59 fig.7 neural network structure (b) hyperbolic tangent function or tanh f(α) = eα − e−α eα + e−α (16) b. rbf. the output of the rbf is calculated using the following form in (5) and (6): yk(x) = m∑ j=1 ϕ(∥x− µj∥)ωkj + ωi0 = m∑ j=0 ϕ(∥x− µj∥)ωkj , k = 1, . . . , l where ωkj is the hidden-output weight and µj is the centre of j-th hidden unit. ϕ is the kernel function. to defining a distance from the input vector x to the rbf centre µj scaled by the scale factor or width σj , equation 13 is used. within the rbf networks, six kernel functions are widely used and defined by equations from (7) to (12). c. bar. the bar network has the same structure as the rbf network, but the kernel function is different. two kernel function for the bar network: (a) gaussian bar function ϕ(x) = n∑ i=1 exp [ − (xi − µji) 2 2σ2 ji ] (17) (b) sigmoidal bar function ϕ(x) = n∑ i=1 1/ { 1 + exp [ − (xi − µji) 2 2σ2 ji ]} (18) 60 m.h. wu:investigation of the different type models for engine test data modelling table 3 lists the summery of the neural network structure used in the nn model tool. table 3 nomenclature of neural network structures 4.3 experimental results 4.3.1 experiment results from the rbf nn model a set of engine test data with three inputs variables and one output variable was employed to evaluate the performance of the proposed approach. the input variables are “break mean efficient pressure”(bemp), “inlet valve opening degree”(vanosin), and “exhaust valve opening degree” (vanosex). the output is “break specific fuel consumption”(bsfc). the whole set of data are from 293 test points. the 196 data points out of 293 are used for training the ga-rbf network and another 97 data points are used for test the ga-rbf network. in the experiment, it run under the following conditions and definitions: • the number of rbfs is fixed at 20 • for comparison purposes, the six different rbf networks are executed 10 times • the normalised mean squares error (nmse) is used as a performance measure for the different rbfs: nmse = ( 1 k ∑k i=1(yi − ti) 2 )1/2 ( 1 k ∑k i=1(ti − t̄)2 )1/2 (19) where yi and ti are respectively the model output and the target value, and t̄ is the mean value of the target values on the training data set or test data set. this expression has the value 0 for a perfect match between model and target, and the value 1 if the model just outputs the target mean t̄. fig.8 shows the evolutionary process of different rbf. fig.8(a) is the average evolutionary process of 10 runs over the training data for six rbfs and fig.8(b) is the best process of 10 runs. the results review that advances in systems science and applications (2014) vol.14 no.1 61 the gaussian function provides the better results compare with other rbfs in terms of convergence speed and modelling accuracy. it is also shown that the local basis function has a faster convergence speed than non-local basis function. the table 4 summaries the final average nmse of 10 runs for different rbfs. the result indicates that the gaussian function has the smallest nmse while the pesudo cubic spline has the largest nmse. table 4 final average nmse over 10 runs for training data average nmse gaussian 0.1217 thin plate spline 0.1397 multiquadric 0.1304 inverse multiquadric 0.1368 pseudo cubic spline 0.1508 logarithmic 0.1274 table 5 compares the final best nmse of 10 runs for both training data and test data. it indicates clearly that the gaussian function performs better than other rbfs in both training data and test data. the figures from fig.5 to fig.10 plot the modelling performance of both table 5 final best nmse over training data and test data training data test data gaussian 0.1134 0.1119 thin plate spline 0.1231 0.1297 multiquadratic 0.1268 0.1245 inverse multiquadratic 0.1304 0.1281 pseudo cubic spline 0.1286 0.1419 logarithmic 0.1235 0.1240 training data and test data using the best final parameters of rbf networks by rcga run. these figures have shown that the rbf networks training by rcga are successfully applied in modelling of engine test data. fig.9 shows the results.of rbf network model for real engine test data analysis and modelling. the model provides a close fit to the ordinary tested data successfully. 4.3.2 experiment results from the nn-tool in this section, an experiment example for engine data modelling was carried out. the input data and output data are listed in fig.10 and fig.11. three nns, mlp, rbf and bar, with different activation functions, abbrevi62 m.h. wu:investigation of the different type models for engine test data modelling fig.8 evolutionary process of rbf networks over 10 runs fig.9 final results fig.10 input data set: speed (rpm), gn(%), ig(degree), vvt(degree) and egr(%) advances in systems science and applications (2014) vol.14 no.1 63 fig.11 output data set: bsfc (g/kwh) ated in table 3, were used throughout the example. the data was taken from engine test bed of the lander rover group, plc. in order to assess the goodness of modelling, the normalized mean squared error (nmse) over each output variable is used: nmse = ( 1 n ∑n i=1(yi − ti) 2 )1/2 ( 1 n ∑n i=1(ti − t̄)2 )1/2 (20) where n is the number of total data points. yi and ti are respectively the model output and the target value, and t̄ is the average value of the targets on the data set, defined as: t̄ = 1 n n∑ i=1 ti (21) nmse has the value 0 for a perfect match between model and target, and table 6 mean, standard (std. dev), minimum and maximum mse over 10 runs for each model 64 m.h. wu:investigation of the different type models for engine test data modelling fig.12 modelling performance of bsfc by sigmoidal bar model the value 1 if the model just outputs the target meant̄ .table 6 shows the result of the example. it is clearly that the model structure which has the minimum mse was the bar model with sigmoidal bar function. the modelling result by bar-sbar is plotted in fig.12. 5 conclusion in this paper, the models used by the engine test data modelling system is introduced. the details of rbf network and the neural network tool have been introduced and the experiment results show that the last two models are running successfully to catch the relationship between the input data and output data within the engine tested data. references [1] voigt k.u. (1993), model based on line optimization for modern engine management system, sae brazil. [2] lechner h. (1996), automated system for optimized calibration of engine management systems, rover group internal publication circa. [3] hochschwarzer h. (1992), fully automatic optimization of engine calibration, imece proceeding, pp.89-100. [4] w c lin & m h wu. (2003), optimization using genetic algorithm in engine calibration, icmr2003, glasgow, uk. [5] c lin, m h wu, s y duan. (2003), engine test data modelling by evolutionary radial basis function networks, imeche publication, journal of automobile eng, part d, vol.217, pp.489-497. advances in systems science and applications (2014) vol.14 no.1 65 [6] m h wu, wanchang lin, s y duan. (2008), develop a nn modelling tool for engine performance tested data modelling, icmr2008, brunel university, u.k. [7] park, j. and sandberg, i. w. (1991), universal approximation using radial basis function networks, neural computation, 3:pp.246-257. [8] poggio t. and girosi f. (1990), networks for approximation and learning, proceedings of ieee, vol.78, pp.1481-1497 corresponding author author can be contacted at: m.h.wu@derby.ac.uk advances in systems science and application (2015) vol.15 no.2 193-201 the physics of jet stream meandering walter e janach meggenhornstrasse 20, ch-6045 meggen, switzerland abstract large-amplitude jet stream meanders involve diabatic processes because conservation of angular momentum leads to vortex line stretching and shrinking. this requires latent heat release during poleward flow in contrast to cooling by sublimation and radiation during equatorward flow. moisture and its distribution, together with aerosols, play a key role. the increase of anthopogenic aerosols in the upper troposphere could be a cause for the more frequent large-amplitude jet stream meanders. keywords angular momentum; potential vorticity; latent heat; radiative cooling; sublimation cooling; rossby waves 1 introduction as the jet stream flows from west to east, rossby waves can form. large meridional excursions of the jet stream in form of meanders occur between alternating warm high-pressure ridges extending polewards and cold low-pressure troughs extending equatorwards. the balanced geostrophic flow of the jet stream is at right angle to the horizontal gradient of the pressure field, so that the resulting horizontal pressure force is balanced by the horizontal component of the coriolis force. at the same time the angular momentum of air parcels in the flow remains constant in the absolute system, which can be expressed by the conservation of potential vorticity. angular momentum remains constant also during horizontal convergence with vortex line expansion, visualized by a ballerina pulling her stretched arms close to her body in order to accelerate the spinning rate, or during horizontal divergence with vortex line shrinking. potential vorticity pv = −g(ζ + f)δθ/δp is the sum of relative vorticity ζ and planetary vorticity f (coriolis parameter) times the difference δθ of potential temperature between the bottom and the top isobaric surfaces of an air parcel divided by the corresponding pressure difference δp, with standard gravity g [1].the pressure difference δp between the bottom and top of an air parcel increases during horizontal convergence with vortex line stretching, while divergence lets the pressure difference δp decrease with vortex line shrinking. 2 conservation of angular momentum during meridional flow in synoptic-scale geostrophic flow of meandering jet stream, relative vorticity ζ is generally much smaller than the coriolis parameter f . as f changes with latitude, ζ of an air parcel cannot change sufficiently to conserve its angular momentum, 194 walter e janach: the physics of jet stream meandering so that a change of pressure difference δp between the bottom and top isobaric surfaces of the air parcel is needed in order to compensate the change of f . during poleward flow of an air parcel in the jet stream, the increase of coriolis parameter f (northern hemisphere) requires horizontal convergence to increase the pressure difference δp through a decrease of horizontal surface area of the air parcel (mass conservation). however the increase of the coriolis parameter f with latitude does not directly force the required increase of δp. what occurs is the forcing of a small ageostrophic deviation ζa of relative vorticity ζ in form of negative velocity shear (anticyclonic) parallel to the flow, superimposed on the velocity profile of the jet stream, as shown in fig.1. this causes a disturbance of the geostrophic balance in the jet stream, which is kept small by the existing pressure field. therefore this ageostrophic disturbance by ζa remains limited to a slightly faster flow on the left side of the jet stream and slower flow on the right side (looking in flow direction). fig.1 velocity profile of poleward jet stream and deviation caused by ageostrophic shear vorticity ζa. the result is a small ageostrophic imbalance between the horizontal component of the coriolis force and the horizontal pressure gradient across the jet stream. this leads to an up-gradient flow deviation on the left side of the jet stream and an opposite down-gradient deviation on the right side. the overall result is horizontal convergence from both sides so that mass continuity of the air parcel with mass δm leads to an increase of pressure difference δp between its bottom and top surfaces. and the increase of δp assures conservation of potential vorticity and angular momentum. this increase of δp gradually changes the pressure field. during poleward flow, the pressure difference δp of an air parcel must increase continuously in order to balance the increasing horizontal component of the coriolis force. by producing a small amount of superimposed velocity shear, which causes lateral convergence of the jet stream, the small ageostrophic deviation ζa of relative vorticity ζ acts as a forcing agent between the continuous increase of advances in systems science and application (2015) vol.15 no.2 195 f and the pressure difference δp. during equatorward flow, velocity shear with opposite sign causes lateral divergence, accompanied by a decrease of δp. in southern hemisphere, f is negative and ζa has opposite sign. the forcing between the change of coriolis parameter f and the pressure difference δp is stabilized by a negative feedback, which adapts the ageostrophic deviation ζa of shear vorticity ζ and the corresponding deviation of velocity in such a way that δp changes in proportion to f . fig.1 shows also the profile of this ageostrophic velocity deviation, whereby ζa is the derivative. the velocity deviation increases outwards from the jet stream center as long as the ageostrophic shear vorticity ζa is negative (anticyclonic). this lets the ageostrophic imbalance become larger with increasing distance from the center and therefore convergence with vortex line stretching also. as a result, starting from a certain distance from the center, the pressure difference δp becomes larger that required to compensate the increase of f . now conservation of potential vorticity forces a positive (cyclonic) correction of ζa in the outer region of the jet stream, accompanied by a reduction of the velocity deviation until it disapperars at the outer edge of the jet stream (fig.1). 3 heating in poleward flow and cooling in equatorward flow during meridional flow of the jet stream there is a continuous increase or decrease of pressure difference δp between the bottom and top isobaric surfaces of air parcels through vertical expansion or shrinking. because this affects an entire layer inside the jet stream, vertical expansion of the layer lifts the atmosphere above it and increases its geopotential height. this requires energy input to the upward expanding layer. in contrast, vertical shrinking of an entire layer reduces the geopotential above it and requires energy reduction. during poleward flow of the jet stream, the increase of δp from upward expansion requires latent heat release from condensation or deposition of water vapor in order to increase internal energy of the air parcel. in contrast during equatorward flow, cooling by radiation to space and by sublimation of ice crystals is needed to decrease internal energy. the upward expansion during poleward flow, which results from convergence and heating, increases the geopotential height of the air columns inside the jet stream as shown in fig.2. this results is westward propagation of the eastward pressure gradient and with it of the geostrophic flow and the jet stream, as a rossby wave. during equatorward flow, the downward shrinking caused by divergence and cooling lets the pressure gradient, which is westward now, propagate westwards also (fig.3). the jet stream meanders between warm high pressure ridges extending polewards and cold low pressure troughs extending equatorwards form a large-amplitude 196 walter e janach: the physics of jet stream meandering fig.2 vertical section of an isobar across poleward jet stream. fig.3 vertical section of an isobar across equatorward jet stream. rossby wave propagating westwards relative to the air mass. thereby air parcels enter into the jet stream from the west, are entrained by the meridional flow and slowly cross it until they exit to the east. this occurs in both poleward and equatorward jet stream. in poleward flow, the air parcels enter from the cold side with lower geopotential, are entrained by the flow and slowly traverse the jet stream towards the warm side with higher geopotential. in the equatorward jet stream, air parcels enter from the high geopotential side in the west and exit to the low geopotential side in the east. a stationary air parcel in the low geopotential region to the west of a poleward jet stream, displacing itself slowly sideways as a westward propagating rossby wave, is initially subjected to an increase of pressure difference δp (fig.2). conservation of potential vorticity at constant latitude and therefore constant f leads to an increase of relative vorticity in form of velocity shear, which accelerates the air parcel polewards. now the coriolis parameter f increases also, forcing a continuous increase of δp by means of a small negative ageostrophic shear vorticity ζa (fig. 1). this continues until the rossby wave and the jet stream have almost propagated past the air parcel and the pressure difference stops increasing. now δp becomes constant again as the air parcel comes to rest and reaches the high geopotential region in the east of the jet stream. heating during poleward flow by the release of latent heat becomes weaker with increasing height because the decrease of temperature strongly reduces the saturation water vapor concentration, a consequence of the nonlinearity of vapor pressure as a function of temperature. this lets the dry and moist adiabats almost coincide. towards the tropopause, long wave radiation to space by ice crystals in cirrus clouds increases strongly and its cooling offsets the heating. this prevents further upward expansion of the top cloud layer. in equatorward jet stream, ice crystals in cirrus clouds are instrumental for both radiative cooling to space and sublimation cooling. during downward shrinking, radiative cooling partly offsets compression warming with the effect that there is less sublimation and ice crystals can survive longer. as sublimation eventually lets ice crystals disappear, starting from the top cloud layer, lower advances in systems science and application (2015) vol.15 no.2 197 clouds get directly exposed to radiative cooling. heating and cooling in meridional flow of the jet stream are part of a dynamic system in which moisture and its distribution play a key role. heating and cooling however change the potential temperature, which violates conservation of potential vorticity. angular momentum can be conserved by replacing the difference of potential temperature between bottom and top isobaric surfaces of the air parcel by its mass m, so that now it is (ζ+ f)δm/δp of the air parcel that must be conserved. 4 changes of jet stream direction conservation of angular momentum of an air parcel requires that the ratio between absolute vorticity (ζ+f) and pressure difference δp must remain constant, δp being between the bottom and top isobaric surfaces of the air parcel with constant mass m. changes of coriolis parameter f or alternatively of pressure difference δp caused by convergence or divergence, lead to changes of relative vorticity ζ and subsequently to a change of the flow field. a change of ζ can be caused by either a change of f or a change of δp, depending upon the situation. diabatic processes play a key role in the evolution of the pressure and flow fields, including jet stream meandering. as an example, a poleward jet stream needs continuous release of latent heat to compensate the increasing f with latitude. an interruption of the moisture supply stops the convergence that increases δp. the result is that now ζ is forced to change, with opposite sign to the change of f . in northern hemisphere with positive f , the increase of f with latitude therefore produces a negative ζ (anticyclonic). this forces the poleward jet stream to veer to the right and eventually results in the cutting apart of the high pressure ridge to the right of the jet stream. thereby a blocking high can form to the north of the partition. in the southern hemisphere with negative f , all signs are opposite, including the sign of anticyclonic ζ, but the physics is the same. in a second example an oversupply of moisture and of latent heat release makes δp increase more than required to compensate the increase of f with latitude. the positive excess of δp causes a positive change of ζ and the formation of a cyclone to the left of the poleward jet stream. in an equatorward jet stream, moisture at higher level is in form of ice crystals that sublimate through compression heating during downward shrinking, which is needed to compensate the decrease of f . there is also radiative cooling to space, which is added to sublimation cooling. a slowdown or interruption of downward shrinking caused by a lack of ice crystals lets the decreasing f force an increase of ζ, which makes the equatorward jet stream veer to the left and east. at this point a cyclone can form at the cold front of the cold trough. now the warm conveyor belt of the cyclone can pick up moisture from the ocean and 198 walter e janach: the physics of jet stream meandering subsequently feed it into the jet stream, allowing it to flow polewards again in the east of the cold trough. 5 the role of aerosols aerosols form the nuclei of droplets and ice crystals by allowing condensation or deposition of water vapor. they play a key role in the formation of cirrus clouds in the upper troposphere [2,3] and in the jet stream. the size distribution and number concentration of aerosols and subsequently of ice crystals influence both the radiative properties and the evolution of cirrus clouds during upward expansion and downward shrinking in the jet stream. when aerosols are more numerous, smaller ice crystals with higher number concentration are formed. smaller ice crystals reduce the sedimentation rate, which influences their concentration at different levels. in addition smaller ice crystals increase the optical thickness for a given mass concentration with the effect that the top layer of the clouds which emits long wave radiation to space, becomes thinner. the higher optical thickness increases also the albedo of solar radiation by cirrus clouds, reducing the absorbed fraction and therefore solar heating of the affected top layer. the higher number density of the smaller ice crystals enhances the deposition of water vapor on ice nuclei because the average distance which a vapor molecule has to cover by diffusion decreases. the variety and subtlety of such processes can be observed in the formation and dissolving of contrails, the condensation trails of high flying airplanes. the dynamics of contrails varies strongly, depending on the properties of the air in which they form and to a smaller degree of the jet engines. 6 potential implications for climate jet stream meandering drives the advective meridional overturn of the troposphere between mid and high latitudes and influences the formation of extratropical cyclones. during meridional flow of the jet stream, aerosols play a key role in the formation of droplets and ice crystals, which in turn controls latent heat release, radiative cooling and sublimation cooling. aerosols have an influence upon the size distribution and number concentration of ice crystals. in the higher troposphere small size aerosols dominate because they have a longer lifetime. anthopogenic aerosols are generally smaller than natural aerosols. a technological reason for example are particle filters for the exhaust from coal fired power plants and diesel engines, which are less effective for small size particles. as a consequence the increase of anthropogenic aerosols at jet stream level is relatively higher than in the lower troposphere, where shortlived larger aerosols are more aboundant. the analysis of residual particles in cirrus crystals after sublimation of the ice by cziczo et al [3] shows that mineral advances in systems science and application (2015) vol.15 no.2 199 dust and metallic particles are the dominant ice nuclei. at present however the anthropogenic fraction of these ice forming aerosols is still uncertain. an increase of anthropogenic ice forming aerosols in the upper troposphere has several potential effects upon both heating and cooling during meridional flow of the jet stream: 1. the increase of the number concentration of small size aerosols leads to more numerous small ice crystals, which reduces their sedimentation speed. this increases their lifetime and in turn the mass concentration of the suspended ice crystals. 2. while this has no direct effect on latent heat release, the smaller size of ice crystals and their higher mass concentration increases the optical thickness, which concentrates the cooling by long wave radiation into a thinner top layer of cirrus clouds. this allows the decreasing amount of latent heat release with increasing height to remain effective up to a higher level and enhances the poleward flow of the jet stream. 3. once an air parcel exits from the poleward jet stream towards the high geopotential side on the east, upward expansion and latent heat release stop, allowing radiative cooling to become dominant. and the increased albedo of solar radiation resulting from the higher optical thickness of the more numerous small ice crystals reduces solar heating, which lets cooling by long wave radiation become more effective. 4. radiative cooling leads to subsidence, followed by compression warming, which lets ice crystals sublimate and generate additional cooling. continuous cooling by sublimation and radiation drives the process of downward shrinking during equatorward flow. and the increase of mass concentration of the smaller ice crystals, resulting from slower sedimentation, increases sublimation cooling during downward shrinking, which enhances the equatorward flow of the jet stream. because of these effects, the increase of anthropogenic aerosols has the potential to enhance both the upward expansion during poleward flow and the downward shrinking during equatorward flow of the jet stream. the overall result is an increase of jet stream meandering. evidence of the arctic jet stream becoming wavier, starting from the mid-1990s, has been presented by francis et al [4]. the regional effects of enhanced jet stream meandering are more frequent cold air outbreaks and the blocking of both low pressure troughs and high pressure ridges. francis [4] presents a linkage between arctic amplification and a wavier jet stream. the increase of anthropogenic aerosols through their role in jet stream meandering is a possible cause. starting from about the year 2000, a slowdown of global warming has appeared. more frequent cold air outbreaks, caused by the increase of jet stream 200 walter e janach: the physics of jet stream meandering meandering, could contribute to this in the following way. during the equatorward flow with downward shrinking of the cold air, high clouds sublimate and allow warmer low clouds and the warm ocean to radiate directly to space. this increases the heat loss by outgoing long wave radiation in a similar way as with the so called “iris effect” [5]. it is interesting to compare the slowdown of global warming during the past 15 years with a similar situation from the 40s to the 60s of the last century [6]. in both situations anthopogenic aerosols from coal burning, heavy industry and mining increased: in the first case during the second world war and the subsequent rebuilding of industry in europe and japan, while in the second case during the forced economic growth in china and india. finally it is possible that the enhanced cooling of a thinner top layer of cirrus clouds, resulting from the higher optical thickness of more numerous small ice crystals, creates a more effective barrier between the troposphere and the lower stratosphere. this could reduce the transfer of water vapor to the stratosphere, with the effect of lowering its concentration there. solomon et al. [7] found evidence of a 10 % decrease of water vapor in the lower stratosphere since the year 2000 and their simulations show a 25 % reduction of global warming. 7 summary and conclusions latent heat release is the main driver for convective overturn in the troposphere. the meridional advective overturn between mid and high latitudes depends also upon the presence of moisture and its distribution. we have shown that conservation of angular momentum and potential vorticity needs the release of latent heat during poleward flow and cooling by radiation and sublimation during equatorward flow of meandering jet stream. thereby aerosols form the nuclei for droplets and ice crystals and influence their size distribution. this in turn controls the rates of precipitation and sedimentation and influences the total mass of suspended ice crystals and their optical thickness, which in turn influences cooling by radiation and sublimation. we explain how the increase of anthropogenic aerosols can influence the meridional flow of the jet stream and suggest that this could increase both its meandering and the blocking of high pressure ridges and low pressure troughs between the meanders. characteristic for the physics of jet stream meandering are the multiple and complex interactions between rotational fluid mechanics, thermodynamics of moist air, aerosols and radiation. an example is the increase of stratification stability resulting from adiabatic convergence or divergence, which explains why latent heat release and cooling by radiation and sublimation are needed to reduce stability. often the distinction between cause and effect is difficult, a challenge for advances in systems science and application (2015) vol.15 no.2 201 numerical models. and a typical case is the evolution of the pressure field with synoptic-scale meridional flow, involving large changes of the coriolis parameter, which lead to the formation of rossby waves. acknoledgements the author would like to thank thomas stocker from university of bern and heini wernli from swiss federal institute of technology zurich for the critical and stimulating discussions. references [1] mcintyre m e. (2015), “potential vorticity”, encyclopedia of atmospheric science, second edition, pp.375-383. [2] lohmann u and feichter j. (2005), ”global indirect aerosol effects: a review atmos”, , atmos chem phys, vol.5, pp.715-737. [3] cziczo d j et al. (2013), “clarifying the dominant sources and mechanisms of cirrus cloud formation”, science, vol.340, pp.1320-1324. [4] francis a f and vavrus s j. (2015), “evidence for a wavier jet stream in response to rapid arctic warming”, environ. res. lett. vol.10. [5] mauritsen th and stevens b. (2015), “missing iris effect as a possible cause of muted hydrological change and high climate sensitivity in models”, tnature geoscience, vol.8, pp.346-351. [6] wilcox l highwood e j and dunstone n j. (2013), “the influence of anthropogenic aerosol on multidecadal variations of historical global climate”, environ. res. lett., vol.8.24-33. [7] solomon s. (2010),“ contributions of stratospheric water vapor to decadal changes in the rate of global warming”, science, vol.327, pp.1219-1223. corresponding author walter e janach can be contacted at: wjanach@gmx.ch the article y.m. tileubergenov, k.kh. shadiyev, b.m. koshpenbetov, y.a. buribayev, z.a. khamzina “fundamental causes of information systems vulnerability and their protection”, advances in systems science and application (2016), vol.16, no.4, p.53-61 was withdrawn by the editorial board due to violation of assa publication ethics (please, see publication ethics statement). the reason is that the article’s content lacks originality due to substantial intersections with the following paper: a.v. revnivykh, a.m. fedotov "main reasons of information systems vulnerability" in global journal of pure and applied mathematics (2016), vol. 12, no. 3, pp. 2133–2142 editorial board of advances in systems science and application, 18.08.17 http://ijassa.ipu.ru/ojs/ijassa/ethics advances in systems science and applications (2014) vol.14 no.1 76-83 exact conditional efficiency robust p-values from an arbitrary ranking of a sample space: an application to genome-wide association studied max moldovan1 and mette langaas2 1australian institute of health innovation, university of new south wales, level 1 agsm building, sydney nsw 2052, australia; 2department of mathematical sciences, norwegian university of science and technology, n-7491 trondheim, norway abstract we introduce a general method for computation of exact conditional efficiency robust enumeration p-values for detection of genotype-phenotype associations at a single bi-allelic genetic locus. our method can be based on any arbitrary ranking test statistics, such as efficiency robust test statistics or asymptotic p-values. the resulting p-values are exact conditional enumeration p-values and satisfy the basic statistical validity property pr(p ≤ α|h0) ≤ α for all parameters under the null hypothesis and all significance levels α. practically, the method allows performing statistically valid significance testing in genomic analyses with unknown modes of inheritance at individual bi-allelic genetic loci the situation typical in genome-wide association studies. we provide an open-source r code implementing the method. keywords: mode of genetic inheritance; efficiency robust statistics; exact conditional inference; enumeration; genome-wide association study. 1 introduction genome-wide association studies (gwas) consider hundreds of thousands of single nu-cleotide polymorphisms (snps) covering the entire human genome. each snp is nor-mally represented by a bi-allelic locus and assessed for association with a speci?c genetic trait, usually in the context of a case-control study. complex diseases such as asthma, diabetes and multiple sclerosis, among many others, are generally targeted by gwas in order to identify common genetic variations as potential disease risk factors. the number of genetic markers in a particular gwas can vary from several hundred thousands to several millions, depending on the platform used for genotyping and the type of genomes to be studied. for example, more snps are required for gwas that utilize african pop-ulations than for gwas involving european populations since the former are older, and thus being exposed to random gene recombinations for more genorations. see hirschhorn and daly[1] and manolio[2] for the interesting well illustrated introductions to gwas. at a given genetic locus, there is a pair of markers, called alleles, inherited from each of two parents. given a trait is passed through this locus, there are advances in systems science and applications (2014) vol.14 no.1 77 several modes of inheritance can be in effect. the dominant mode of inheritance requires the presence of a single “disease” allele from one of parents for a trait to be inherited. the recessive mode of inheritance requires “disease” alleles from both parents to be passed to the offspring for a trait to express. in case of the additive mode of inheritance, a trait is expressed only partly if a single disease allele is inherited, but express in full if both “disease” alleles are in place. there are a few more modes of inheritance can be specified depending on the degree to which a trait is expressed in an offspring, see visscher et al.[3] for the discussion of heritability concepts. when the mode of inheritance at a genetic locus is known, higher power of a test for genotype-phenotype association can be achieved through using the cochran-armitage trend test (catt) under the explicit assumption of the specific genetic model, see lettre et al.[4] and gonzlez et al.[5]. in practice, however, the mode of inheritance is usually unknown. under this typical scenario, the socalled efficiency robust tests (see podgor et al.[6]) can be used, the tests which remain sensitive to detection of genotype-phenotype associations even though the genetic model is either unknown or misspecified. there are several efficiency robust testing strategies. for example, the max test, first suggested by freidlin et al.[7], has been recommended by the several authors, see zheng and gastwirth [8] and gonzlez et al.[5]. this testing approach is implemented as aa sequential application of several statistical tests optimal for alternative genetic models with retaining the most significant result. the traditional version of the max test, normally referred to as max3, is based on the three catts with scores motivated by dominant, recessive and additive genetic models. alternatively, persons chi-square test (χ2) can be included within the same max testing strategy, leading to max4, see li et al.[9]. zheng et al.[10] demonstrated that χ2 test can be considered as a type of a trend test and also noted that this test is sensitive to detection of overdominant (underdominant) modes of inheritance. min2 is one more variant of the max test implemented as a combination of the additive catt and χ2 , see joo et al.[11]. a slightly different efficiency robust testing strategy is known as mert and is a weighted version of catt optimal for recessive and dominant models, see gastwirth[12] and freidlin et al.[7]. there are several more efficiency robust testing approaches can be specified and some authors even suggest applying a combination of different versions of efficiency robust tests within a single testing procedure, see joo et al.[11]. in finite sample settings, many of the currently known and used efficiency robust tests are not guaranteed to lead to statistically valid inference. this is because the under-lying computational procedures are based either on random sampling or on asymptotic distributions of efficiency robust statistics (gonzlez 78 max moldovan: exact conditional efficiency robust p-values from ... et al.[5], joo et al.[13], so anda sham [14]). for the methods that use random permutations (see sladek et al.[15]), the statistical inference will be valid, but a very high number of random permutations is needed to achieve the required precision for traditionally low gwas type significance levels (often in the order of 10−8 ). recently, loley et al.[16] attempted to unify the efficiency robust testing approaches by proposing a framework also leading to inference of unknown statistical validity. in the current paper, we introduce a computational procedure that takes as an input the ordering of a sample space imposed by any of test statistics or p-values, including the ones introduced above. the procedure outputs exact conditional enumeration p-values that satisfy the basic validity property pr(p ≤ α|h0) ≤ α, for all parameters under the null hypothesis and all significance levels α. 2 notation and the method let the information on a single snp be represented by the 23 contingency table given by table 1, where xi and yi are the counts of observed genotypes for n1 cases and n2 controls, respectively, with n = n1+n2 . we denote this empirically observed table by s ∗ (x1, x2|m1,m2, n1, n2), because all the other entries of the table can be calculated from these numbers. note that for given n1 , n2 , m1 and m2 , there is a finite number of possible contingency tables, called a reference set (verbeek[17]) and denoted here by (m1,m2, n1, n2). next let t be an arbitrary ranking statistic with the value t corresponding to the empirically observed table s ∗ (x1, x2|m1,m2, n1, n2). given a general hypothesis of ‘h0 : no association between genotypes and the case-control status of the subjects’ tested against ‘ha : there is association between genotypes and the case-control status of the subjects’, and larger values of t being more hostile to the null h0 , the set of tables ranked lower or equal than the observed table s ∗ (x1, x2|m1,m2, n1, n2) is given by the critical set: r(x1, x2|m1,m2, n1, n2) := {s(i, j|m1,m2, n1, n2) : t ≥ t} (1) by definition, a p-value is the probability of obtaining the outcome as extreme or table 1 genotype counts at a bi-allelic locus. aa ab bb total case x1 x2 x3 n1 control y1 y2 y3 n2 total m1 m2 m3 n worse than the empirically observed outcome s ∗ (·) under the null, which is just the probability of the critical set r(·). under the null and based on the assumed advances in systems science and applications (2014) vol.14 no.1 79 underlying hypergeometric sampling scheme (see lehmann[18] for descriptions of alternative sampling schemes), the probability of obtaining each individual table s(i, j|m1,m2, n1, n2) within the reference set can be computed as follows: f(i, j|m1,m2, n1, n2) = ( m1 i )( m2 j )( n−m1 −m2 n− i− j ) ( n n1 ) (2) see lloyd[19] for the generalization of the central multivariate hypergeometric probability function given by (2). the p-value ps∗;t corresponding to the empirically observed table s ∗ (x1, x2|m1,m2, n1, n2) is the probability of the critical set r(·) given by (1): ps∗, t (x1, x2|m1,m2, n1, n2) = ∑ s∈r f(i, j|m1,m2, n1, n2) (3) note that ps∗, t is a fisher-type conditional p-value by construction, inheriting positive(e.g. statistical validity and empirical relevance) as well as negative (e.g. conservatism and computational challenges) aspects of fisher’s p-values. 3 numerical illustration denote statistics obtained from catts optimal for dominant, recessive and additive models, respectively, by td, tr and ta (sasieni[20]): td = n(nx1 − n1m1) 2 n1m1(n− n1)(n−m1) tr = n(nx3 − n1m3) 2 n1(n− n1)(nm3 −m2 3) ta = n(n(x2 + 2x3)− n1(m2 + 2m3)) 2 n1(n− n1)(n(m2 + 4m3)− (m2 + 2m3)2) all three statistics asymptotically follow the chi-square distribution with one degree of freedom. the max3 test statistic is given by tmax3 = max(td, tr, ta) with the observed value tmax3 = max(td, tr, ta). for the empirically observed table s∗(0, 2|3, 4, 4, 5),the p-value ps∗, tmax3 can be computed as shown in table 2. specifically, there are 11 tables in the reference set s(3, 4, 4, 5) and only three tables in the critical set r(0, 2|3, 4, 4, 5) since only the tables with the values of tmax3 statistics equally or more extreme then observed are included in the critical set, i.e. tmax3 ≥ tmax3. the resulted exact conditional efficiency robust p-value ps∗, t (0, 2|3, 4, 4, 5) = 0 : 0952 and is the sum of f(·|m1,m2, n1, n2) given 80 max moldovan: exact conditional efficiency robust p-values from ... by (2) of the three tables in r(0, 2|3, 4, 4, 5). table 2 the illustrative example is based on (m1,m2, n1, n2) = (3, 4, 4, 5) with an observed value (x1, x2) = (0, 2). the critical region r is given by the lower part of the table under the horizontal line. table 2 genotype counts at a bi-allelic locus. x1 x2 td tr ta tmax3 f(x1, x2|m1,m2) 1 2 0.2250 0.0321 0.1636 0.2250 0.2857 2 1 0.9000 0.0321 0.2557 0.9000 0.1905 1 3 0.2250 2.0571 0.2557 2.0571 0.0952 2 2 0.9000 2.0571 2.0045 2.0571 0.1429 1 1 0.2250 3.2143 1.7284 3.2143 0.0952 2 0 0.9000 3.2143 0.1636 3.2143 0.0238 0 3 3.6000 0.0321 1.7284 3.6000 0.0635 0 4 3.6000 2.0571 0.1636 3.6000 0.0079 0 2 3.6000 3.2143 4.9500 4.9500 0.0476 3 0 5.6250 0.0321 2.0045 5.6250 0.0159 3 1 5.6250 2.0571 5.4102 5.6250 0.0317 ps∗, t = 0.0952 4 conclusion the method we suggested above is by no means new. the initial idea can be traced back to fisher [22] and ps;t given by (3) is based on the combinatorial results known for many decades, see freeman and halton [23]. our contribution to the original fisher’s methodology is the idea of ordering the sample space, given by the reference set s, based on any arbitrary chosen ranking statistics, the efficiency robust test statistics in our case. we have borrowed this approach from the unconditional exact testing literature, see barnard [24] and lloyd and moldovan [25] for the origination of the unconditional inference philosophy and one of the initial attempts to combine the conditional and unconditional types of exact inference, respectively. to conclude, it should be pointed out that only the basic form of the adjustment procedure has been given above. in practice, more special cases can arise, such as the presence of covariates (e.g. additional snps, environmental factors or baseline factors) or involvement of additional shifted parameters (e.g. in power studies). while this is clearly the limitation of the presented procedure, the basic general exact conditional method introduced above gives a solid basis for further investigations to these and possibly several more theoretical and applied research directions. we provide an open-source r code to encourage and facilitate such advances in systems science and applications (2014) vol.14 no.1 81 investigations. the r code is available upon request from the authors. references [1] hirschhorn, j.n., and daly, m.j. (2005), “genome-wide association studies for common diseases and complex traits”, nature reviews genetics, 6, 95108. [2] manolio, t.a. (2010), “genomewide association studies and assessment of the risk of disease”, new england journal of medicine, 363, 166-176. [3] visscher, p.m., hill, w.g., and wray, n.r. (2008), “heritability in the genomics era-concepts and misconceptions.”, nature reviews genetics, 9, 255266. [4] lettre, g., lange, c., and hirschhorn, j.n. (2007), “genetic model testing and statistical power in population-based association studies of quantitative traits” genetic epidemiology, 31, 358-362.. [5] gonzalez, j.r., carrasco, j.l., dudbridge, f., armengol, l., estivill, x., and moreno, v. (2008), “maximizing association statistics over genetic models” genetic epidemiology, 32, 246-254.. [6] podgor, m.j, gastwirth, j.l., and mehta c.r (1996), “efficiency robust tests of independence in contingency tables with ordered classi cations.” statistics in medicine, 15, 2095-2105. [7] freidlin, b., zheng, g., li, z. and gastwirth, j.l. (2002), “trend tests for casecontrol studies of genetic markers: power, sample size and robustness”, human heredity, 53, 146-152. [8] zheng, g., and gastwirth, j.l. (2006), “on estimation of the variance in cochranarmitage trend tests for genetic association using case-control studies.” statistics in medicine, 25, 3150-3159. [9] li, q., zheng, g., and yu, k. (2009), “robust tests for single-marker analysis in case-control genetic association studies.” annals of human genetics, 73, 245-252. [10] zheng, g., joo, j., and yang, y. (2009), “pearson’s test, trend test, and max are all trend tests with different types of scores.” annals of human genetics, 73, 133-140. [11] joo, j., kwak, m., ahn, k., and zheng, g. (2009), “a robust genome-wide scan statistic of the welcome trust case control consortium.” biometrics, 65, 1115-1122. 82 max moldovan: exact conditional efficiency robust p-values from ... [12] gastwirth, j.l. (1985), “the use of maximin efficiency robust tests in combining contingency tables and survival analysis”, journal of the american statistical association, 80, 380-384. [13] joo, j., kwak, m., and zheng, g. (2010), “improving power for testing genetic association in case-control studies by reducing the alternative space”, biometrics, 66, 266-276. [14] so, h.c., and sham, p.c. (2011), “robust association tests under different genetic models, allowing for binary or quantitative traits and covariates”, behavior genetics, 41, 768-775. [15] sladek, r., rocheleau, g., rung, j., dina, c., shen, l., serre, d., boutin, p., vincent, d., belisle, a., hadjadj, s., balkau, b., heude, b., charpentier, g., hudson, t.j., montpetit, a., pshezhetsky, a.v., prentki, m., posner, b.i., balding, d.j., meyre, d., polychronakos, c., and froguel, p. (2007), “a genome-wide association study identifies novel risk loci for type 2 diabetes”, nature, 445, 881-885. [16] loley, c.,, i.r., hothorn, l., and ziegler, a. (2013), “a unifying framework for robust association testing, estimation, and genetic model selection using the generalized linear model”, european journal of human genetics, to appear. [17] verbeek, a. (1985), “a survey of algorithms for exact distributions of test statistics in r × c contingency tables with fixed margins”, computational statistics and data analysis, 3, pp.159-185. [18] lehmann, e.l. (1986), testing statistical hypotheses, 2nd ed., new york: wiley. [19] lloyd, c.j. (1999), statistical analysis of categorical data, new york: wiley [20] sasieni, p.d. (1997), “from genotype to genes: doubling the sample size”, biometrics, 53, pp.1253-1261. [21] devlin, b., and roeder, k. (1999), “genomic control for association studies”, biometrics, 55, pp.997-1004. [22] fisher, r.a. (1935), “the logic of inductive inference (with discussion)”, journal of the royal statistical society, 98, pp.39-54. [23] freeman, g.h., and halton j.h. (1951), “note on an exact treatment of contingency, goodness of fit and other problems of significance”, biometrika, 38, pp.141-149. advances in systems science and applications (2014) vol.14 no.1 83 [24] barnard, g.a. (1947), “significance tests for 2 × 2 tables”, biometrika, 34, pp.123-138. [25] lloyd, c.j., and moldovan, m. (2007), “unconditional efficient one-sided confidence limits for the odds ratio based on conditional likelihood”, statistics in medicine, 26, pp.5136-5146. corresponding author authors can be contacted at: max.moldovan@gmail.com. advances in systems science and application(2015) vol.15 no.2 164-173 expert system for international crisis management -crisman rafael calduch cervera and nina wörmer nixdorf department of international law and international relations, madrid, spain abstract the prototype crisman is part of an ambitious project which claims to introduce computer simulation as a part of investigation labour fulfilled by spanish internationalists through the practicality of knowledge based systems and the establishment of multidisciplinary investigation groups. keywords expert system, international relations, rule-based, artificial intelligence, prediction, conflict. 1 computer simulation applied to international relations in the field of social sciences the discipline of international relations (ir) was included into degree studies in united kingdom and united states in 1919. nevertheless, its theoretical development and the use of quantitative techniques since the 1950s, allowed to incorporate mathematical models (game theory and probability theory) creating indicators which provided a more accurate description and explication of many phenomena of the international society: arms race, armed conflicts, “guerrillas” and terrorism, underdevelopment and conflicts of structural nature, mass media and public opinion, international bipolarity and multi-polarity or regional political integration, among others. progresses made in computational technology, artificial intelligence (ai) and communicational and informational media have accomplished to deal with international happenings through simulation[1]. in fact we can find important contributions of agent based simulations (abs) in social science[2]. however, the main obstacles for applying computer simulation in the field of ir are two: a) the vagueness of multiple theoretical concepts, used in this science, making quantitative data collection and its evaluation (using indicators or probability calculation) more difficult, and b) the existing ignorance among internationalists concerning the possibilities of new, both informational and computational, technologies. a fact that leads, quite often, to underestimate and undervalue the work carried out in the area of simulation. these two realities are more present among internationalists and investigators of the spanish speaking world due to the crucial influence of disciplines like legal or historical sciences during their training, being neither of those sciences really permeable concerning quantification and new technologies. crisman is the first step taken to introduce ai as an effective technique in the investigation labour fulfilled by spanish internationalists. we point out three consecutive phases of this project: 1st phase: creation of a prototype of a rule-based expert system (rbes) covering the analyse and foresight of international crises, with a direct application in different activities developed by investigators, intelligence analysts, political decision makers, high business executives, all of them guided in their international scope activity. advances in systems science and application(2015) vol.15 no.2 165 prepare a doctoral thesis focusing on the utility of expert systems (ess) teaching ir, including the development of a computer programme focusing on diagnostic, classification and training of international conflicts. 2nd.investigate about the applicability of fuzzy logic to ess created for international crisis management[3]. 3rd.develop a new ess concerning evaluation and international political risk management as well as training in international aid programmes through simulation. 2 prototype of expert system for international crisis management -crisman 2.1 problem framing the problem which is tried to be solved by using a rbex consists in defining the evolution and existing relations between two countries or group of countries a and b, taking as a starting point an initial situation (conflict or crisis) between those countries, in order to anticipate the more likely decisions which will be taken by political leaders and to provide explanations of those decisions. 2.2 technical specifications this es is constructed upon a theoretical model or simplified description of reality, in order to facilitate the learning of foreign policy analysis techniques. to achieve this goal, computer data and processing storage will be used with the aim of creating simulations of real-world cases through the use of rules or set of logical formulations which associate, in a conditional way, causal variables or attributes with consecutive variables or attributes. the shell or computer programme used for this expert system is clips, developed by nasa (http://www.ghg.net/clips) in its 6.23 version (2005). we also find an adapted version for java called jess ((http://www.herzberg.ca.sandia.gov/jess), as well as an application suitable for fuzzy logic called fuzzyclips, 6.10c version, developed by the national research council of canada. the knowledge base and the production rules have been carried out by professor rafael calduch cervera[4]. a forward chaining inference process allocated with certainty or trust coefficients will be applied to each of them. the case data, used as examples, and the results, obtained by executing the programme, will be stored in a data base using xml language. 2.3 theoretical model the theoretical model underpinning the knowledge base and the rules of the expert system are articulated on the existing relation between six basic attributes: 1) situations; 2) aims; 3) available means; 4) previous case experience; 5)future expectations; 166 rafael calduch cervera and nina wörmer nixdorf: expert system for international ... 6) relational strategy and 7)conduct. the user should define the variables choosing between the different options given by the es concerning the first four attributes, while the programme will deduce the categories 5, 6 and 7; which, in turn, become the initial situation for the new realisation cycle of the es. the biggest methodological problem that had to be solved to make the theoretical model applicable was to reduce the semantic options of each attribute category to its minimum, avoiding the exponential growth of number of rules. this has not been easy, mainly because of the huge amount of diversity of conditions, objectives, means and type of existing experiences concerning relations between two countries. graphic display of the theoretical model fig.1 graphic display of the theoretical model attributes 1.situations: advances in systems science and application(2015) vol.15 no.2 167 1.1).normalised 1.2).conflict 1.3).deepening conflict 1.4).crisis 1.5).armed conflict 2.aims: 2.1).compatible 2.2).perception of compatibility 2.3).incompatibility 2.4).perception of incompatibility 3.available means 3.1).equivalence 3.2).superiority 3.3).inferiority 4.previous case experience 4.1).certainty of trust 4.1.1).trust 4.1.2).mistrust 4.2).uncertainty 5.future expectations 5.1).continuation of the situation 5.2).escalating of disputes 5.3).deescalate of disputes 5.4).unpredictable 6.relational strategy 6.1).cooperative 6.2).negotiating 6.3).competing 6.4).deterrent 6.5).aggressive 7.conducts: 7.1).keep cooperation 7.2).intensify diplomacy 7.3).diplomacy with no military pressure measures 7.4).diplomacy with deterrent military measures 7.5).military deployment with limited use of force 7.6).widespread use of force 168 rafael calduch cervera and nina wörmer nixdorf: expert system for international ... 3 description of attributes 3.1 situations it is referred to the mutual relations between two countries in a limited period of time. we can choose between the following options: 1.1.normalised: the situation in which the relations between two countries (a and b) are mainly cooperative and respecting international legal standards. 1.2.conflict: the situation in which the relation between two countries (a and b) is of conflict, although there exists no explicit threat of violence or the use of force. 1.3.deepening conflict: the situation of conflict between two countries (a and b) in which one or both countries turn to military measures of pressure or with deterrent character, although there exist no direct nor express threat of violence or use of force. 1.4.crisis: the conflict situation in which, one or both countries (a and b), threaten with the use of force or make a limited use of it to condition the behaviour of the other country. 1.5.armed conflict: the situation in which the general use of force is the main tool concerning the relation of both countries (a and b). 3.2 aims we understand this topic as those interests or purposes which are being tried to achieve by each of the countries according to their capacities and available means in a particular situation. according to the relation between purposes and targets of the countries we can find the following possibilities: 2.1.compatible: it takes place when the purposes and targets of both countries (a and b) can be achieved simultaneously and, as well, this possibility is sensed as part of reality. 2.2.perception of compatibility: it takes place when the subjective perception of achieving interests or targets from one of the countries (a) can be achieved simultaneously with the satisfaction of others interests or targets of the other country (b), even though elements which prevent this compatibility exist. 2.3.incompatibility: it takes place when the accomplishment of interests or targets of one or both countries (a and b) cannot be achieved simultaneously and, also, this possibility is clearly sensed as part of reality. 2.4.perception of incompatibility: it takes place when the subjective perception makes believe that the achievement of interests or targets of one of the countries (a) cannot be achieved simultaneously with the satisfaction of the interests or targets of the other country (b), even though there do not exist any elements for this incompatibility in reality. 3.3 available means this topic includes all kind of capacities a country is willing to use in the relations with any other country to ensure the consecution of its targets. e can find the following possibilities: 3.1.equivalence: it exists when the means used by one country (a) in its relation with another country (b), whatever its nature, hold or at least are valued with the same level of effectiveness in the same manner, to ensure the achievement of their respective targets. advances in systems science and application(2015) vol.15 no.2 169 3.2.superiority: it exists when the means used by one country (a) in its relation with another country (b), are in fact, or at least are valued with a higher grade of effectiveness than those of the other country to ensure the attainment of their particular targets. 3.3.inferiority: it exists when the means used by one country (a) in relation with the other country (b) are in fact, or at least are valued with a lower grade of effectiveness than those of the other country to ensure the attainment of their particular targets. 3.4 previous case experience it is formed by the evolution of the political relations between two countries (a and b), during the period of time in which a generation exercises leadership in both countries (25 to 30 years) in the way it is perceived by their leaders. we have to work with the following options: 4.1.certainty of trust: when the leaders of one country (a), based on former identical or analogous experiences, are convinced that the leaders of the other country (b) will adopt the necessary decisions and actions to fulfil their compromises to execute their threats assuming the consequences associated to it. 4.1.1.trust: when the leaders of a country (a), based on former identical or analogous experiences, are convinced that the leaders of the other country (b) will fulfil the compromises reached with them, although they may imply losses of interests or additional winnings in case of infringement, and in case of escalation, they would formulate a clear and explicit threat. 4.1.1.mistrust: when the leaders of a country (a), based on former identical or analogous experiences, are convinced that the leaders of the other country (b) will not fulfil the compromises reached with them, because of putting the unilateral satisfaction of their targets or the achievement of additional winnings through infringement, and in case of escalation, they would not formulate a clear and explicit threat. 4.2.uncertainty: when the leaders of a country (a) have a lack of former identical or analogous experiences or the ones existing are contradictory and, in consequence, they have no deep-seated conviction concerning the degree of compliance or infringement of the other country (b) will make in relation with the achieved compromises or, in case of escalation, they do not know if a previous, clear and explicit threat will be formulated. 3.5 future expectations we consider future expectations as the evolution expected by political leaders of every country concerning the present situation with the other country. we have to choose between four possibilities: 5.1.-continuation of the situation: no significant changes between the two countries (a and b) are expected. 5.2.escalating of disputes: a new, more conflictive, situation between countries (a and b) is expected due to changes in their relation. 5.3.de-escalating of disputes: a new, less conflictive, situation between countries (a and b) is expected due to changes in their relation. 170 rafael calduch cervera and nina wörmer nixdorf: expert system for international ... 5.4.unpredictable: the political leaders of countries (a and b) lack the knowledge or enough former experiences not being able to make future expectations concerning the evolution of their relations. 3.6 relational strategy these kinds of strategies are formed by the planning and organization of the relations of one country with another through the exclusive, or clearly dominating, remit of specific types of behaviour in the relation. we can find the following categories: 6.1.cooperative: the strategy which resorts to collaborative relations between two countries (a and b) to achieve the established targets together. 6.2.-negotiating: the strategy which uses communication or political and diplomatic negotiation to achieve some kind of agreement or understanding between two countries (a and b) to satisfy their targets. 6.3.competing: it is the kind of strategy that combines communication or political and diplomatic negotiation with pressure builds or conflictive, but no military, behaviours, trying that country (a), the one making them, achieves unilateral advantages in relation with country (b) or achieves the exclusive satisfactions of its targets. 6.4.deterrent: this kind of strategy combines communication or political and diplomatic negotiation, deployment and/or increase of the military capacity of country (a), with the explicit intention of using this military capacity to defend its vital interests and targets or to avoid an armed aggression fulfilled by the other country (b). 6.5.aggressive: this strategy implies the unilateral and extensive use of force by one country (a) to achieve specific targets or exclusive advantages at the expense of the other country (b). 3.7 conducts conducts are the dominating actions which characterize the relations between two countries in a specific situation. there are the following possibilities: 7.1.keep cooperation: it means that the behaviour of one country (a) concerning the other country (b) keep the same so that the mutual interests or joint targets can be satisfied. 7.2.intensify diplomacy: it means that there is an increase of communication or political and diplomatic negotiation between one country (a) regarding the other country (b) in order to facilitate or achieve any type of understanding or agreement to achieve their interests or targets. 7.3.diplomacy with deterrent military measures: this conduct creates the use of a combination of communication or political and diplomatic negotiation with the use of behaviours, excluding the use of force, that are intended to provoke direct damage to the other country (b) or to hinder the accomplishment of its targets. 7.4.diplomacy with deterrent military measures: this conduct means the use of different communicational or political and diplomatic conducts with the use of demonstrations measures, display or increase of its military capacity by one country (a) but only with a defensive character concerning the other country (b). 7.5.military deployment with limited use of force: this kind of conduct means the use of a combination of measures from one country (a) that entails a direct threat for the other country (b) through the display of its military capacity or the use of these advances in systems science and application(2015) vol.15 no.2 171 military capacities in a limited period of time and space. 7.6.widespread use of force: this conduct involves developing of all kind of strategic, tactical and logistical conducts needed to warlike use of military capacities of one country (a) against another country (b) in order to achieve its defeat. 4 priority of conditioning of attributes to draw up rules the order of priority concerning the application of causal attributes to determine consistent attributes is the following: 1.initial situation 2.aims 3.available means/previous case experience the main conditioning resides in the initial situation of the relations between two countries because of two main reasons: firstly, because this situation will also be the final situation, product of former relations between those countries and this, of course, affects the previous case experience and also because the initial situation conditions the category of aims which are tried to be fulfilled by each country. the second most important conditioning is set-up by the aims every country tries to achieve through its relation with another country. concerning this topic we have to bear in mind the different strategies and conducts that can be developed in their relation. the third level of conditioning considers the effectiveness concerning the means every country uses compared to the ones used by another country, because this relation of effectiveness will condition possible strategies to be followed by every country as well as the effectiveness of conducts adopted by each of them concerning the other. nevertheless, this third conditioning can also match with previous case experience because this experience is decisive to determine future expectations made by the leaders of every country and this will, as well, condition the election of the judged to be more effective or probable strategies and conducts to achieve the targets set. the priority of conditionals concerning every single attribute will determine, in case of conflict, the effects that should prevail in the rule formulation, considering the following criteria: a).when the conditionals of different levels are complementary they cause a strengthening or intensification of the resultant effects. b).when there is a conflict or opposition between conditions of different levels, the effects of the hierarchical higher level will be chosen. c).if the conflict is between two conditionals of the same level their result concerning the effects will be neutralized, using those effects derived from higher hierarchical level. examples of rules used in crisman rules 1)normalized initial situation r 1).if initial situation = normalised and aims a) and b) = compatible 172 rafael calduch cervera and nina wörmer nixdorf: expert system for international ... and available means a) and b) = equivalence and previous case experience a) and b) = certainty of trust then future expectations a) and b) = continuation of the situation then relational strategy a) and b) = cooperative then conducts a) and b) = keep cooperation so final situation = normalised r 2).if initial situation = normalised and aims a) and b) = incompatibility and available means a) and b) = equivalence and previous case experience a) and b) = certainty of trust then future expectations a) and b) = escalating of disputes then relational strategy a) and b) = negotiating then conducts a) and b) = intensify diplomacy so final situation = conflict r 3). if initial situation = normalised and aims a) = perception of compatibility and aims b)= perception of incompatibility and available means a) and b) = equivalence and previous case experience a) and b) = certainty of trust then future expectations a) = continuation of the situation then future expectations b) = escalating of disputes then relational strategy a) = cooperative then relational strategy b) = negotiating then conducts a) = keep cooperation then conducts b) = intensify diplomacy so final situation = normalised 5 theoretical operating model of the es a).the initial situation between the countries a and b at the moment (t-0) is the starting point. b).the aims of each country (a and b) to identify the possibility/probability of simultaneous and joint or unilateral success of those aims by each of the countries are analyzed. c).the relation of existing effectiveness of each country (a and b) to achieve their targets is assigned. d).historical precedents that have existed in the relations between both countries (a and b) in identical or analogous situations are evaluated to define their influence concerning the perception of the leaders. taking the options chosen by the user for former attributes into account, the es will deduce: 1).expectations that the political leaders of each country (a and b) set on the future development of relations between the two countries. 2).the most likely strategies political leaders will develop in each country (a and b) in their relation with the other country. advances in systems science and application(2015) vol.15 no.2 173 3).the behaviour that will prevail in each country’s relations with the other based on the established relations strategy. 4).-the type of situation that will result at the end of reciprocal behaviour made by both countries. this final situation concludes the cycle and it is established as the initial situation of the next cycle in the operation of the expert system. 6 user profile crisman can be used in the teaching of graduate students in the universities and research centres, as well as by international analysts who may carry out an evaluation of political risk assessment. * carlos e. calduch cervera developed the prototype of the computer program to perform the expert system crisman. references [1] taber, ch. s. (1992), “poli: an expert system model of u.s. foreign policy belief systems”, american political science review, vol.86, no.4, pp.888-904. goldspink, ch. (2002), “methodological implications of complex systems approaches to sociality: simulation as a foundation for knowledge”, journal of artificial societies and social simulation, vol.5, no.1. see http://jasss.soc.surrey.ac.uk/5/1/3.html. [2] davisson, p. (2002), “agent based social simulation: a computer science view”, journal of artificial societies and social simulation, vol.5, no.1, see http://jasss.soc.surrey.ac.uk/5/1/7.html. [3] cioffi-revilla, c. a. (1981), “fuzzy sets and models of international relations”, american journal of political science, vol. 25, no.1, pp.129-159. [4] calduch, rafael. (1991),international relations , ediciones de ciencias sociales, madrid. [5] calduch, rafael. (1993),dynamics of international society, ediciones ceura, madrid. corresponding author nina wörmer nixdorf∗ can be contact at: nwormer@ccinf.ucm.es advances in systems science and application (2016) vol.16 no.1 95-101 vlsi implementation of lightweight cryptography algorithm hinpreet kaur and sakthivel r school of electronics engineering, vit university, vellore, tamil nadu, india. abstract lightweight algorithms for cryptography are popular for resource stringent devices. now days, radio frequency identification techniques are gaining popularity because of their small size and low cost applications. in this paper, hummingbird algorithm is used to provide security in such devices like rfid tags, smart cards, etc. it is a hybrid algorithm which provides security against most of the common attacks encountered such as linear attacks and differential attack. the hybrid combination of block and stream cipher makes this algorithm more secure using lesser number of clock cycles. the algorithm is implemented on both fpga and asic platform. the main aim is to reduce the number devices utilized and using lesser area on the chip. the algorithm is implemented in xilinx vertex-5 family. this design consumes the total standard cell area of 0.255 mm2. the design is placed and route using cadence soc encounter at tsmc90nm technology. keywords hummingbird algorithm; rfid tag; lightweight cryptography; asic; fpga. 1 introduction the motivation behind this work is the increase in demand of resource stringent devices such as smart cards and devices using the rfid (radio frequency identification) technology. rfid tags are now being used in libraries to keep the records of the book, in hospitals for medical equipment detection, etc. the small size and low cost of the tags makes it popular. rfid devices consist of mainly three elements, a tag, a reader and a database system. there exist a communication between the tag and the reader. this communication is affected by outside environment if a third person or attacker tries to leak the information. so it becomes essential to incorporate an algorithm in order to make the communication secure. this work proposes an area efficient architecture on both fpga and asic platform. 2 background hummingbird is proved to be very secure algorithm because it is a hybrid algorithm of stream cipher and block cipher. despite of presence of many cryptographic algorithms such as des (data encryption standards) and aes (advanced encryption standard)[? ], we need algorithm which is lightweight on both hardware 96 hinpreet kaur and sakthivel r: vlsi implementation of lightweight cryptography algorithm and software side, which is not fulfilled by the mentioned algorithms. since rfid devices are small in size, the cryptographic unit thus used should be hardware friendly and provides the same level of security as in other non-resource constrained devices. several work has been carried out till related to this algorithm. many other lightweight algorithms such as present, klein, hight, led, iceberg are implemented in fpga platform but hummingbird is proved to provide the highest degree of security and is resistant to many attacks such as birthday attacks, algebraic attacks, structural attack and cube attack. the work related to hummingbird are tabulated in table1.the first hummingbird algorithm was implemented in 4-bit microcontroller with low power consumption[? ]. many implementations are being carried out on fpga platform are shown in the table1. table 1 different implementations of hummingbird algorithm paper year of publication implementation platform key feature ref[? ] 2009 microcontroller 4-bit microcontroller , low power consumption ref[? ] 2010 fpga larger throughput ; smaller area ref[? ] 2011 fpga co-processor approach ref[? ] 2011 fpga throughput oriented and area oriented ref[? ] 2011 fpga gen2 protocols for mac ref[? ] 2014 fpga low power and high speed 3 algorithm steps the hummingbird algorithm consists of a 256-bit secret key shared between the tag and the reader. there are four internal registers of 16-bit and a 16-bit galois field lfsr. the registers are first initialized in the initialization process with some random initial vectors which are then being updated by the lfsr in the encryption process. the input to encryption is the plaintext and the output is the 16-bit cipher text. fig.1 shows the working of the algorithm. the 16-bit plain text is given as the input along with the 256-bit key, which is segmented into four 64-bit subkeys. there are three blocks which perform initialization, encryption and decryption. in the initialization process, the registers are updated for encryption. after each set of plaintext and cipher text, the internal status registers are updated by the 16-bit lfsr. the initialization block and the encryption block consist of the block substitution box or s-box and the linear permutation. the predefined s-box and the inverse s-box is shown in table 2. the message is encrypted by the tag using the encryption process and the resultant cipher text is sent to the reader. the reader decrypts the message using advances in systems science and application (2016) vol.16 no.1 97 the same key (as the key is being shared between tag and the reader). fig. 1 algorithm flow of hummingbird cryptography table 2 s-box and the inverse s-box used in the algorithm x 0 1 2 3 4 5 6 7 8 9 a b c d e f s(x) 2 e f 5 c 1 9 a b 4 6 8 0 7 3 d s(x) c 5 0 e 9 3 a d b 6 7 8 4 f 1 2 4 proposed architecture the proposed architecture is shown in fig.2. this architecture works on the encryption only and decryption only processor. the initialization block consists of four status registers and a ready pin. a 5-bit counter is used to count the number of clock cycles. when the data rdy pin goes high, the counter starts counting and the status registers are initialized with some random values. the registers are then updated by undergoing encryption round of block cipher in next 16 clock cycles. for the algorithm refer[? ]. the encryption block uses the initialized status registers to encrypt the plain 98 hinpreet kaur and sakthivel r: vlsi implementation of lightweight cryptography algorithm text into cipher text. encryption undergoes the modulo 216 addition of register rs1 and plain text and undergoes through the block encryption using the secret key. this is repeated four times and the resulting cipher text is given to the decryption module making the enc complete signal high. the decryption is just reverse of encryption. the input is cipher text and output is plain text. modulo 216 subtraction is used in decryption. thus when both dec complete and enc complete signals are high the output is given in one clock cycle. the registers are updated for the next process. table 3 device utilization summary sheet logic utilization used available utilization ref.[? ] this work ref.[? ] this work number of slice register 4242 74 12480 33% 1% number of slice luts 3504 2309 12480 28% 18% number of bonded iobs 278 117 172 161% 68% number of fully used lutff pair 1907 70 ref.[? ]5839 this work-2313 32% 3% number of bufg/ bufctrls 8 1 32 25% 3% fig. 2 proposed architecture of hummingbird encryption and decryption 5 simulation results the encryption block along with the initialization module of the hummingbird algorithm are designed and simulated using the s-box using the lut based apadvances in systems science and application (2016) vol.16 no.1 99 proach. lut based s-box uses less area and power. the simulation results are shown in fig.3 the 16-bit plain text is encrypted to 16-bit cipher text. when the reset is high, there will be no initialization process. after the reset signal changes to high, the initialization starts and the 16-bit plain text is converted to its cipher text using a 256-bit secret key. the first cipher text is obtained after 8-clock cycles and then at the subsequent clock cycles we will get the other cipher texts. the simulation carried out using modelsim 6.5b. in the simulation results shown in fig.3, a plaintext is encrypted using a 64-bit key and the respective cipher text and the decrypted data is obtained after 20-clock cycles. the result of the console window is given below: # hummingbird input data==0011000000111001 # hummingbird key==1110011011100111001000001101101110111000001101101011000111011001 # hummingbird encrypt data==0101100100001110 # hummingbird decrypt data==0011000000111001 fig. 3 simulation result of encryption process the entire algorithm is designed and implemented in xilinx 14.7 ise suite with vetex-5 xc5vlx20t in package ff-323 and speed grade -2. the results obtained were compared with the mentioned results in [? ]. the device utilization sheet is shown in table.3. the total number of slices occupied in our design is 917, which is very less as mentioned in [? ]. the design works at a frequency of 14mhz and the power consumption is 322.80 mw at 2.5v. the device utilization summary sheet mentioned in table.3 shows that our design utilizes less space and is hardware friendly. the lightweight algorithm such as hummingbird can be used in resource stringent devices and are now a days used in password identification, library management system, hospitals embedded in the rfid tag. the asic implementation of the hummingbird encryption and decryption core is done using cadence soc encounter using tsmc 90nm technology. the final chip layout is shown in fig.4. the design is successfully placed and route using cadence soc encounter tool using tsmc 90nm technology with no setup and hold violations. the cadence rc compiler results in the area, power, timing and the gate counts used in the design. the die area, power, area in gate counts and maximum frequency of 100 hinpreet kaur and sakthivel r: vlsi implementation of lightweight cryptography algorithm operation are tabulated in table 4. the design has worst case skew of 19psec and the best case skew of 16.1psec with no timing violations. after the place and route the design was verified for the geometry and there were no violations. this proposed design is thus area efficient at both fpga and asic platform. fig. 4 final chip layout table 4 specifications area .2556mm2 power 2.234mw gate counts 11363 max. frequency 2.129 ghz 6 conclusion and future scope the hummingbird encryption and decryption module is implemented in both fpga and asic platform. the design is coded using verilog hdl and implemented in xillinx 14.7 ise suite in vertex-5 family. the results show reduced device utilization in slices. the asic implementation is done using cadence soc encounter using tsmc90nm technology library. the design is successfully placed and route with no timing violations and the area power of the final chip is noted. thus this design is suitable for resource stringent devices and hardware friendly as compared to aes and other cryptographic algorithms. the algorithm used here s-boxes, which can be designed using boolean expression instead of lut based approach. so the area and power can be reduced by using bdd reduction unit. the variable reordering reduces the boolean functions according to the variables being selected. learning cad tools to perform bdd reduction module is further scope of study. references [1] p.chodowiec and k.gaj. (2003), “very compact fpga implementation of aes algorithm”, in cryptographic hardware and embedded systems cches 2003, c. walter, c. koy and c. paar(ed.), springer berlin heidelberg, vol. 2779, pp. 319-333. [2] f. xinxin, g. guang, k. lauffenburger and t. hicks. (2010), “fpga implementations of the hummingbird cryptographic algorithm”, in hardwareadvances in systems science and application (2016) vol.16 no.1 101 oriented security and trust (host), 2010 ieee international symposium on, pp. 48-51. [3] t. san and n. at.(2011). “compact hardware architecture for hummingbird cryptographic algorithm”, in international conference on field programmable logic and applications (fpl), pp. 376-381. [4] m. biao, r. c. c. cheung and h. yan. (2011). “fpga-based high throughput and area-efficient architectures of the hummingbird cryptography”, in iecon 2011 37th annual conference on ieee industrial electronics society, pp. 3998-4002. [5] x. mengqin, s. xiang, w. junyu and j. crop.(2011), “design of a uhf rfid tag baseband with the hummingbird cryptographic engine”, in 2011 ieee 9th international conference on asic (asicon) , pp. 800-803. [6] x. mengqin, s. xiang, w. junyu and j. crop.(2011), “design of a uhf rfid tag baseband with the hummingbird cryptographic engine”, in 2011 ieee 9th international conference on asic (asicon), pp. 800-803. [7] nikita arora and yogita gigras. (2014), “fpga implementation of low power and high speed hummingbird cryptographic algorithm”, international journal of computer applications , vol. 92, no. 16, pp. 0975-8887. [8] xinxin fan. (2010), efficient cryptographic algorithms and protocols for mobile ad hoc networks, canada: ontario. corresponding author hinpreet kaur can be contacted at: hinpreetkaur@gmail.com advances in systems science and applications (2013) vol.13 no.2 116-143 modified yoyo model and its applications in chinese history zelong wang1,yi lin2 and jubo zhu1 1school of science, national university of defense technology,changsha, 410073, china 2department of mathematics, slippery rock university, slippery rock, pa 16057, usa abstract this paper focuses on the modified yoyo model, a hot issue of the system research called the second dimension of science, and it supplies a new theory and methodology to address and solve the practical problems, which have been extremely difficult in modern science. moreover, the new explanations of the chinese history about the periods of spring and autumn and the warring states are introduced by the modified yoyo model. the systemic yoyo model is a new tool to describe the general systems, and it is effective to explain the existing phenomenon. the modified yoyo model, based on the systemic yoyo model, is constructed by the forms of the definitions and laws, which are demonstrated and explained by lots of the existing system theories and practical examples. the application of this model in chinese history is a new look at the historical events, compared with the traditional study of the chinese history. the modified yoyo model is proposed based on the systemic yoyo model and it emphasizes on explaining some theoretic problems of the system model, such as the general system structure model, system’s characteristics, systems’ interaction and so on. firstly, the yoyo model can not sufficiently depict the attributes of the general system, such as the behavior, the function, the size of the system and so on. all of these are the indispensability of a system and the system is half-baked without them. secondly, the basic characters of the yoyo model are not given enough to describe the system roundly. these matters include the complexity, diversity, stability of the system and how the system comes into being or perdition. thirdly, there have been many examples that refer to the interaction between yoyos, but the interaction process is not detailedly discussed and the final result can be ascertained by some laws. the modified yoyo model is applied in the chinese history about the spring and autumn and warring states periods. the military system of each country in the ancient time is modeled by the modified yoyo model and their interactions are analyzed by the characters of the modified yoyo, which lead to the same results with the history. the modified yoyo model is firstly applied in the chinese history, and the battle effectiveness of the military system is modeled as a yoyo. the wars between the countries are described as the interaction of the modified yoyos, and the rea117 advances in systems science and applications (2013) vol.13 no.2 sonable results of the wars are also explained. in fact, since the general systems are modeled by the modified yoyo model, this model can be applied in many practical systems, such as three-body movement system, economic system and so on. in theory, the modified yoyo model embodies and develops the systemic yoyo model, and it introduces some new concepts and characters of the yoyo as well as their interaction, which can be effectively used for the general system analysis. when referring to the application, the modified yoyo model can explain or predict lots of phenomenon, which may be the puzzles in modern science. keywords modified yoyo model, yoyo’s character, yoyos’ interaction, spring and autumn and warring states, military system 1 introduction as the developments of the system science and system thinking, new theories and methodologies have brought new understandings and discoveries to some of the major unsettled problems in modern science. moreover, the view of wholeness and interconnectedness in system thinking has greatly changed the tendency of modern science, such as synthesizing all areas of knowledge into a few major blocks, the appearance of the second dimension science and the cross-disciplinary studies, etc. system is not only an entity that can be touched and felt actually but also a thinking that is used to observe the world around us. then we would have a new idea to find the potential connection and disciplines that have been ignored for the divide of modern science. in fact, lots of so-called puzzles are caused by this parochialism. therefore, it is reasonable that these puzzles may be solved easily by system thinking, which causes the development of system science inversely. the yoyo model[1], cared and developed by yi lin, is a useful systemic tool to describe the real world and nearly all kinds of phenomenon, and it stands for the freshest achievement of system science. however, the renascence of yoyo model causes its imperfectness in theory and application. many entities and phenomenon have been modeled by the yoyo model soundly, but most of the basic properties of the system are not explained precisely, leading to an uncertain result. so this paper focuses on modifying the yoyo model and applying it in chinese history of the periods of spring and autumn and warring states. 1.1 a history review of system and the origin of yoyo model systems methodology is an important concept in system science, and it is understood in different ways by scholars at different periods and different science fields. however, the understandings are roughly the same and they are unified together gradually. quastler[2] said: “systems methodology is essentially the eszelong wang: modified yoyo model and its applications in chinese history 118 tablishment of a structural foundation for four kinds of theories of organization”. the opinion of zadeh[3] is that the main task of systems science is the study of general properties of systems without considering their physical specifics. even though the concept of systems has been a hot spot of discussion in almost all areas of modern science and technology, which was first introduced formally by von bertalanffy[4] in the second decade of the 20th century in biology, as all new concepts in science, the ideas and thinking logic of systems have a long history. although the concept of system is not emphasized in history, its idea was proposed thousands of years ago and “the whole is greater than the parts” is good evidence. during most time of the science history, the first dimension method has played a leading role in scientific research. as time going by, system attracted more and more attentions for the conflicts from the isolation of subjects and lots of new characters of system, such as the wholeness, interconnection, and occlusion, have been proposed in different fields. in last century, there have been many advances in technology: energies produced by various devices such as steam engines, motors, computers and automatic controllers, self-controlled equipment from domestic temperature controllers to self-directed missiles, and the information highway that has resulted in increased communication of new scientific results. all of these made the system a individual subject, and some scholars began to research the system from many points of view. the concept of the second dimension gradually came to the view, and system science came to a new level. at the turn of 21st century, with his profound insights, independent creativity, and courage, shoucheng ouyang proposed the blown-up theory of nonlinear evolution problems[5]. uneven structures are eddy sources, leading to eddy motions instead of waves, the mystery of nonlinearity, which has been bothering humankind for long time, is resolved at once both physically and mathematically. on the basis of the blown-up theory, the concepts of black holes, big bangs, and converging and diverging eddy motions are coined together in the model shown in fig.1[6]. each system or object considered in a study is a multidimensional entity that spins about its invisible axis. if such a spinning entity is fathomed in threedimensional space, a structure like that shown in fig.1 would be achieved. for the sake of convenience of communication, such a structure is called a yoyo due to its general shape. more specifically, what this model says is that each physical entity in the universe, be it a tangible or intangible object, a living being, an organization, a culture, a civilization, etc., can be seen as a kind of realization of a certain multidimensional spinning yoyo with an invisible spin field around it. it stays in a constant spinning motion. if it does stop spinning, it will no longer exist as an identifiable system. 119 advances in systems science and applications (2013) vol.13 no.2 fig.1 yoyo model (the side of a black hole sucks in all things, such as materials, information, energy, etc. after funneling through the short narrow neck, all things are spit out in the form of a big bang. some of the materials, spit out from the end of the big bang, never return to the other side, and some will.) 1.2 wide applications of yoyo model because spin is the fundamental evolutionary feature and characteristic of materials, the yoyo model can be used in figurative analysis method. it also gives us a chance to generalize all three laws of motion so that external forces are no longer required for these laws to work[7]. in terms of applications of the yoyo model in social science and humanity areas, it is shown that in a market of free competition, a concept as fundamental as demand and supply is about mutual reactions and mutual restrictions of different forces under equal quantitative effects[7]. hence, each economic entity can be naturally modeled and simulated as an economic yoyo or a flow of such yoyos. when looking at how people think, one can show the existence of the systemic yoyo structure in human thoughts[7]. so, the human way of thinking is proven to have the same structure as that of the material world. in a word, what can be thought as a system, what can be considered as a yoyo model. why does the yoyo model have so wide applications in our world? the most important is that it captures the essential character, i.e. nonlinearity and unevenness, of the general systems, and scientifically abstracts the main conflict between problems. the yoyo model describes the uneven resources, so it also describes the nonlinear motion since the motion comes from the uneven resources. the second reason lies in the materialism, which implies that the world is made up of materials. no one can talk anything without the materials, including the yoyo model. the yoyo model firstly pays attention on the materials, the basis of general system, and then has the ability to uncover the disciplines of much phenomenon about the materials. moreover, the yoyo model, proposed and developed by lots greats, has a unique structure that is fit to all kinds of real systems. all of these zelong wang: modified yoyo model and its applications in chinese history 120 make the yoyo model have a leading role in system research. 1.3 the imperfection of yoyo model as what’s known, the yoyo model has attracted more and more attentions since its appearance. this depends on its wide applications in our lives; however, due to the renascence of the yoyo model, it has some imperfections that need to be developed. there have been many examples that the yoyo model can explain partly, which show that the yoyo model is correct to describe our world; on the other hand, it can not give us a clear sense about the systems, their basic characters and their interactions. so we need to consummate the yoyo model to give a perfect explanation about the general system. fig.1 has gives us a direct vision about the yoyo model, which stands for any objects that can be considered as a real system. however, when a system is modeled as a yoyo, we could not know the basic information about the yoyo model, such as its exact shape, height, its spin direction, etc. in any case, we only know that the yoyo model exists with a system, and all the things, such as the materials, information and energy, come into the yoyo from the black hole and come out from the big bang. how shall we measure the stability of the yoyo? how does the yoyo spin? what is the yoyo’s size? nobody knows. if the yoyo model and its basic attributes are strictly defined, all of the puzzles above would be solved easily. moreover, the characters of the yoyo, such as the complexity and the diversity, should be explained by the yoyo model. firstly, the behavior of the yoyo is not paid enough attentions, because there are lots of cases that we can not capture the yoyo directly, but the behavior of the system is easy to learn, by which we can conclude the system indirectly, i.e. the yoyo and its behavior have some equivalences. secondly, the characters of the yoyo can tell us how a new yoyo comes into being, how an old yoyo is destroyed and how the yoyo keeps its steady structure. therefore, the characters of the yoyo should be studied carefully before its applications. another problem is the interaction between yoyos. as explained in examples, the interaction between yoyos is complex and the yoyos always release a ‘force’ to each others, so the phenomenon of the birth, the development and the perdition of yoyos occur during the interaction. for example, the competition between two yoyos, standing for two companies, will cause the bankruptcy of the weaker or their contemporary existence, but how shall we know which would happen? what we have finished only is that both of them can be explained by yoyo model and their interaction leads to two results. therefore, the interaction between yoyos should be analyzed more precisely and interaction disciplines should be uncovered. then based on the interaction laws, we can easily predict what will happen after the interaction between yoyos. 121 advances in systems science and applications (2013) vol.13 no.2 therefore, the yoyo model should be modified and some indispensable parts should be added to the yoyo model. since lots of systems have been successfully explained by the original yoyo model, the modified yoyo model would be more efficient in applications. in the following, the modified yoyo model is proposed and it would be proved by some simple examples. 1.4 organization of this paper this paper focuses on the modified yoyo model and its application in chinese history. in section 2, the modified yoyo model is introduced, including the background of the modified yoyo model, the characters of modified yoyo and their interaction. this section mainly consummates the yoyo model and makes the yoyo model have enough information to stand for the system exactly. the disciplines of the interaction between yoyos are helpful to predict the possible interaction result. section 3 gives us a new example about modified yoyo model, which has not been illuminated by the yoyo model. chinese history about spring and autumn and warring states is a wonderful period, and the transition from slave society to feudalism society causes the complexity and the diversity of the politics, militancy and ideology. these can be all explained by the systemic yoyo theory, but only the military system is modeled by the modified yoyo model. the last section is the conclusions about the whole paper and two open problems about the yoyo model are introduced for the continuing research. 2 modified yoyos and their interaction in order to give a more refined explanation about the system, the modified yoyo model and their interaction laws are proposed in this section. 2.1 background of modified yoyo model the key concept of system science is “system” and it has been a hot issue since its birth. although it is still not clear enough, the basic properties, such as the wholeness, interaction and so on, has been accepted by lots of scholars. now if we define a system, every one will understand what we mean. the yoyo model, a new tool as well as other system models, is proposed to describe the general system, and it has explained lots of properties of the system, such as the blown-ups and shrinking. it uncovers the essential characters of the general systems, i.e. the nonlinearity and unevenness cause the eddy motion. it is a great discovery that many unsettled problems in modern science have been solved by the yoyo model. however, the yoyo model is not perfect, because another important factor, i.e. the behavior of the system, is not included in the model. most of the time, we not only focus on the system itself, but also put emphasis on the interaction between systems. if a system can not have any influence on the environments, its existence will not make any sense. the system and its behavior should be zelong wang: modified yoyo model and its applications in chinese history 122 trussed together and modeled by yoyo model. except for the system itself, the behavior of the system is another aspect that can help us to understand the system’s properties. for example, even if we never hear of or visit some company, we can know its size easily by its raw material or production scope. the raw material and production are considered as the input and output of the company, which can also be seen as the behavior of the company. when we research the company, the input and output are also important factors except for the company itself. sometimes, we even can not capture the system; instead, we can only describe its behavior. taking the black hole as an example again, nothing can escape from the black hole, including the light, let alone the scientists. so nobody can look into it, but we can conclude its properties by its behavior. therefore, the behavior of the system supply us another way to describe the system. on one hand, the basic characters of the yoyo depend on its behavior and can also be depicted by the behavior. for example, the stability, complexity and diversity of the yoyo can be reflected by its behavior. on the other hand, the interaction between yoyos is also based on their behavior. when a system receives an action from others or it has an action on others, the acceptation and the impartment are both the behavior of the system. for example, there is gravitation among three-body problem, and the competition exists between two similar companies. when a system has no behavior, it would not have any connection with the environments and then it would not make any sense for the development of the human, which would not be the research context of human. as to the close system, it actually would not exist in nature, i.e. the systems in nature are open and the close ones are ideal. therefore, the yoyo model should include not only the system but also their behavior, from which the basic characters and the interaction laws of the yoyos. in the following, the modified yoyo model and its new properties are proposed. 2.2 modified yoyo model the basic concept of yoyo model has been introduced in detail, and every system can be considered as a yoyo model. in this subsection, the modified yoyo model is introduced, including its structure, intensity, polarities, size and its representation. all of them are reconsidered from a new point of view. 2.2.1 spin field the yoyo model has been shown in fig.1, where the eddy stands for the system, reflecting its nonlinearity and unevenness. what are the things that come into or out the yoyo from the black hole or the big bang? as to the company, they may be the input and output of the company; they can also be gravitation field for the celestial bodies; or they can be air current for the fan. they are actually the behavior ability of the system and can not be separated from the system, just as 123 advances in systems science and applications (2013) vol.13 no.2 fig.2. the behavior of the system is also important for the interaction between systems, and they either receive the behavior from others or give off behavior to others. if a system does not have any behavior, maybe it would not exist any more. therefore, the system and its behavior are trussed together and the yoyo model should also reflect both of them instead of the system alone. in the yoyo model, they are spin field through the yoyo, as along as the yoyo spins all the time. without the spin field, the yoyo would not spin and the yoyo would be destroyed. what’s more, the spin field is the medium that the yoyos connect fig.2 spin field of yoyo model with each other. it is just like that the systems have connection with others by their behavior, instead of the system themselves. for example, the company changes its raw material and production as along with the outside market, and the output and input are medium for connection. without these things, the company nearly has no behavior and no connection with the market. therefore, the spin field is a necessary for the yoyo model; the yoyo will be destroyed without the spin field. the spin field is not the part of the yoyo, i.e. the system, but it denotes all the behavior ability of the yoyo, including the action that the yoyo receives or gives off. since the function of the system is carried out by its behavior, every system should have the behavior ability to realize some function; otherwise, the existence of the system would not make any sense. therefore, as to the yoyo, there must be spin field through it, permeating the whole space. 2.2.2 spin line in order to make the spin field exact and visible, the spin line (sl) is assumed to describe the spin field of the yoyo model and it does not exist actually. as to a company, each spin line stands for a flow of input or output. generally speaking, a truss of spin lines of the yoyo reflects a kind of behavior ability by which the system has connection with other systems. all the spin lines of the yoyo form the whole behavior ability of the system, so the spin field is made up of all the spin lines of the yoyo, i.e. as to discretionarily given point in the spin field, there is a zelong wang: modified yoyo model and its applications in chinese history 124 spin line through it, shown in fig.3. since the spin lines denote the behavior fig.3 spin lines of yoyo model ability of the system, the behavior intensity, i.e. the spin field intensity (sfi), should also be defined by the spin lines. the sfi at some point in the spin field is defined as: sfi(p) = lim a→0 1 a ∫ s(p) sl(s)ds (1) where p is the observed point; sl(s) is the spin line that pass through the point s; the s(p), covering the observed point p, is the observed area and a denotes the acreage of the observed area, shown in fig.4. returning to the system archetype, sfi reflects the behavior intensity of the system, and where the sfi is larger, where the density of the spin lines is bigger, meaning that the system has a greater behavior. fig.4 spin field intensity there are two important properties of the spin lines, i.e. the partial occlusion and the non-intersection. generally speaking, the partial occlusion means that some of the spin lines that depart from some points would return to these points after period of time and their tracks are just like some close loops, but others not. as to a system, the close spin lines denote that the action that this system takes to the environments has influence on the system itself actually. however, not 125 advances in systems science and applications (2013) vol.13 no.2 all the action from the system has influence on itself, so there also are non-close spin lines except for the close ones. for example, if a company is considered as a system, i.e. a yoyo model, the spin lines would be the flow of the input and output of the company. some stuff of the input maybe just part of its production, which indicates some of the output returns to the company as part input of the company; and other output not. therefore, the spin lines of the yoyo are classified as the close and the non-close, i.e. the spin lines are partial occlusion. the second property of the spin lines is the non-intersection, i.e. each two spin lines in spin field of the yoyo do not have the same point. it means that the behavior of the system (a yoyo) would not have influence on its other behavior or each two categories of the behavior of a system are independent of each other. actually it may be impossible that two categories of the behavior of a system do not have any connection with each other, but the puny connection is always ignored and they are supposed to be independent for a convenient analysis. this method of assumption has been used in nearly all the modern science, and it is also so in system science. taking the three-body problem as an example, the popular method used in astrophysics is that the two heaviest bodies are firstly researched without the third body and then the third one is added to the two-body system. what to do next is just some modification of the system. the analysis results by the independent assumption are very close to the factual instance, although the modification is not discussed in this paper. 2.2.3 polarity of the yoyo we have known that all the things come into the yoyo by the black bole and out from the yoyo by the big bang, so the black hole and the big bang are two basic properties or polarities of the yoyo. however, the black hole and the big bang do not have universality, and they are replaced by the polarity of the yoyo. in fact, each yoyo, i.e. each system, has two polarities, such as the input and output of the system. generally speaking, each system has the ability of giving off influence on and receiving influence from the environments, and these different directions of behavior abilities of a system are considered as the system’s polarities. moreover, the two polarities of the yoyo can not be separated, such as the positive and negative polarities of a cell, the south and north poles of a magnet, etc. as to general system, if it only has influence on the environments and can not receive action, the system can not be controlled; if it only receives action from the environments and can not produce any action, the system would be useless and not exist any more. since two polarities of yoyo exist, they should be defined and their meanings should also be explained clearly. they are defined as the following. every yoyo has two polarities, the positive and the negative polarities; the positive polarity is located in the side where the spin lines of the yoyo come out from the yoyo, zelong wang: modified yoyo model and its applications in chinese history 126 and the negative one is located in the other side of the yoyo, shown in fig.6. the polarities can reflect the direction of the behavior of the yoyo, i.e. the yoyo takes action on others or receives action from others. for example, when a car is running by the control of the motorman, the motorman is located in the negative polarity and the running of the car is located in the other polarity. both of the two polarities are important for the car, the car would make no sense without anyone. fig.5 two polarities of the yoyo 2.2.4 size of the yoyo when we focus on a yoyo (a system), its size is an important content. actually it has great influence on the behavior intensity of the yoyo and even the behavior category, so it is discussed here. generally speaking, the size is a relative concept and it is hard to define this concept among the systems belonging to different layers. in fact, the size of the yoyo depends on the research intention and content, and the size of even the same system would be different if it is researched by different intention and content. if two yoyos are researched for different contents, it would make no sense to compare the size of a yoyo with the other. for example, when we experiment on the mass of the earth and the velocity of the light respectively, the mass and the velocity can be seen as a measure of their sizes. however, they do not have any comparableness for the different dimensions. as to the yoyos with the same research content, how shall we define the sizes of them by the yoyo model? when studying a yoyo, mostly we do not directly analyze the system itself; we can put an emphasis on its behavior instead. the definition of the size can also come from this thinking. if we study the mass of the earth, we can also measure it from the eart’s gravitation to another known system, i.e. the outside behavior of the earth. and sometimes it is the only method to measure the yoyo’s size, because the yoyo has not been mastered by the human in detail, such as the black hole, the universe, some particles and lots of social phenomenon. in sub-subsection 2.2.1 and 2.2.2, we have defined the spin field and the spin 127 advances in systems science and applications (2013) vol.13 no.2 lines of the yoyo, and both of them depict the behavior of the yoyo. so it is natural to define the size of the yoyo by its spin field or spin lines. as to a given research content , the behavior of a yoyo is only some of all the spin lines except for studying any aspects of the yoyo, shown in fig.5. in other words, the size of a yoyo is its behavior intensity at the research aspect. fig.6 size of the yoyo for a given research content therefore, the size of a yoyo is defined as: size(γ) = ∫ s(γ) sl(s)ds (2) where γ is the research content and s(γ) is the total spin lines that stands for the whole behavior at the aspect γ. if the spin lines denote the input or output flow of a company, then the total input or the output, defined as the size of the yoyo, can reflect the manufacture size of the company sufficiently. 2.2.5 denotation symbol of the yoyo in the book “systemic yoyos” written by professor lin, the denotation symbol of the yoyo model has been given in detail, just like figure 7. they are very visual, but there is some discommodiousness by these symbols. first of all, they can not contain enough information of the yoyo, such as the spin field, the size and the polarity of the yoyo, so they can not denote yoyos sufficiently. secondly, these symbols are inconvenient to paint on the papers, because there is a long line round some fixed point and it is hard to paint roundly. moreover, they take up a larger area of the paper. thirdly, when two yoyos have interaction with each other, their shapes change to elliptic type. if these symbols denoted the yoyos, does that mean the yoyos themselves have been changed? as we known, the interaction between yoyos need not change the yoyos, but we can not see the interaction clearly if the shape does not change. then a contradiction occurs. fourthly, as to a given yoyo model, we can see different spin directions from different sides, so how shall we make a decision about the spin direction when painting them? zelong wang: modified yoyo model and its applications in chinese history 128 fig.7 denotation symbol of the yoyo given by professor lin therefore, a simple and convenient symbol is given in fig.8. this symbol is a line with the same direction to the spin lines of the yoyo; its length is decided by the size of the yoyo; the line points from the negative polarity to the positive polarity and the jumping-off point of the line is located in the neck of the yoyo. moreover, when a yoyo line is given, the spin yoyo can be only ascertained by the representation laws. the representation law can be described as the following: the direction of the spin lines of the yoyo is accordant to that of the symbol, and spin direction of the yoyo has dextro rotation relation with that of the symbol, i.e. when the yoyo line is grasped by the right hand and the thumb points to the direction of the spin lines, the spin direction of the yoyo is along to the direction of other four fingers; the size of the yoyo is also ascertained by the length of the line. this representation symbol is not only simple but also contains nearly all the information of the yoyo. what’s more, they have interaction with each other by the spin field, i.e. their behavior, instead of the yoyos. fig.8 yoyo model and its denotation symbol 2.3 properties of modified yoyo model since the yoyo model has been modified and some dimensions have been added to the yoyo model, the new properties of the yoyo model are discovered naturally. these include the characters of the yoyo and the interaction between yoyos, which are described lively by the new dimensions. 129 advances in systems science and applications (2013) vol.13 no.2 2.3.1 basic characters of the modified yoyo as to the simplex yoyo, the most interesting things are how the yoyo comes to being and perdition, how the yoyo keeps steady and why the yoyo can endure the little change of the outside environments. all the answers come from the modified yoyo model and they can be explained in the following by the form of properties. property1 : the cause and effect relation leads to the continuity of the spin line through the yoyo, so the fluxes through two polarities are equivalent. the cause and effect relation widely exists in the world. for example, we are living for food; the output is produced for hard work; and the circle movement of celestial bodies keeps for their gravitation. so the cause and effect relation is general in the word, including the modified yoyo model. every effect of a system is begotten by some cause that has been known completely or incompletely. maybe there is more than one reason that leads to a selfsame result; to analyze these complex reasons conveniently, all of them are summed up to a spin line that links to the spin line of the result at the neck of the yoyo. therefore, the spin lines that have been painted in the figures continuously pass through the yoyo, which indicates the cause and effect relation of the system. since every cause spin line leads to an effect spin line and every effect spin line is begotten by a cause spin line, it is concluded that the quantities of spin lines through each polarity, i.e. the two fluxes of spin lines through two polarities correspondingly, are equivalent. when the “cause” spin lines change, the “effect” spin lines would change correspondingly. for example, if a factory is modeled as a yoyo, the change of the producers’ goods must leads to a change of the yield. the equivalence of the spin lines through two polarities can bring the convenience for analyzing the characters and interaction of the yoyos, so it is a basic and important property. for example, according to the research content and intuition, the size of the yoyo can be defined by each flux of spin lines through two polarities. another example is that the direction of the spin lines can be ascertained by each polarity of the yoyo for their equivalence. property2 : the behavior size of a yoyo has equipollence with its tsme size. the size of a yoyo is defined by its behavior ability, i.e. its spin field, and this definition is one of many possible definitions about the yoyo’s size. it can be used to describe the size of a yoyo if the yoyo’s structure and component are not learned clearly by us or we pay our attentions to the yoyo’s behavior ability instead of itself. when a yoyo has some invariability with regard to its structure or component, maybe it also can be used to describe the yoyo’s size. because for each given yoyo there must be a positive number such that: at ×bs × cm ×de = a (3) zelong wang: modified yoyo model and its applications in chinese history 130 where a,b,c,d and d are some constants determined by the structure and attributes of the system of our concern,t is the time as measured in the system,s is the space occupied by the system, and m and e are the total mass and energy contained in the system. if the yoyo keeps steady, the positive constant a, called tsme size in this paper, measures the essential size of the yoyo. moreover, the behavior size has some equivalence with the tsme size, i.e.: a ⇔ ∑ γ∈ω size(γ) (4) where γ is a sort of behavior ability, and ω is the set of all behavior abilities. property3 : the diversity and complexity of a yoyo are reflected by that of its behavior abilities, i.e. the spin field. now we have known that the world is wonderful for all kinds of species, so why the world is composed of so many objects, including the organics and the inorganics? the great natural historian darwin gives the answer by his darwinism, stating that all species of organisms arise and develop through the natural selection of small, inherited variations that increase the individuals ability to compete, survive, and reproduce. in other words, the national selection of behavior abilities causes the biology evolution as well as the inorganics. another visible example is that lots of artificial objects, such as the car, the computer, the airplane and so on, arise as the development of the human civilization, which increases the diversity and complexity of the world. their appearance lies in that their various behavior abilities can service for the human. we can not affirmatively conclude that the diversity and complexity of the system are brought to by their various behavior abilities, but it is affirmative that the diversity and complexity of the system can be reflected by that of its various behavior abilities. as to the yoyo model, its structure, component and attributes are not discussed in detail, so it is hard to talk about its diversity and complexity by the yoyo itself; however, its spin field, i.e. the various behavior abilities, can reflect its diversity and complexity. so the different behavior abilities and intensities stand for the difference between different yoyos. property4 : the stability of the yoyo is equivalent to that of its spin field. firstly, let us discuss the stability of the system. when a system is in the state of stability, its component, structure and function all do not change as time going by. of course its behavior is also steady. if the system’s component or structure changes, its behavior would change correspondingly, which even can not be discovered easily by us. moreover, if the behavior of a system keeps steady, the system must also keep steady. in other words, the behavior of a system can reflect the stability of the system. for example, when a cup is broken up, it can not be used for drinking, i.e. the 131 advances in systems science and applications (2013) vol.13 no.2 behavior of the cup has changed greatly. sometimes the destruction of a system can not bring a timely change of its behavior, but this change would come sooner or later. taking the assumptive vanishing of the sun as an example, all the planets in solar system maybe can not experience this change at once, i.e. they can still receive the gravitation and light for about 8 minutes[8], but their orbits must change for the disappearance of the sun finally. the change of the behavior of a system must lie in the change of the system; otherwise, we even can not control the simplest system to service for the human. in modified yoyo model, the system and its behavior are modeled as the yoyo and its spin field, so the stability of the yoyo is equivalent to that of its spin field. property5 : the change of the spin field leads to the corresponding size change of the yoyo. the change of the spin field means that the behavior of the yoyo has changed, so the size of the yoyo changes for its definition. when an additional “force” is operated on a system and its function is changed, the behavior of the yoyo must be changed correspondingly. for example, the input and output of a company, i.e. the spin field of a yoyo, are reduced for the furious market competition, and then the company has to reduce its size, such as reducing the employees, decreasing the salary and saving the cost. otherwise, the company would take the opposite action. property6 : the yoyo has inertia to keep the original state. the concept of inertia, explicitly proposed by great physical scientist newton, is used to describe the kinetic character of all the objects. the inertia counterworks the change of the movement of an object, but it can not stop this change. does a system (yoyo) have similar inertia to counterwork the change of its state or stability? yes, it not only can counterwork this change but also can stop this change within some range, i.e. when the yoyo is influenced by some outside factors, the yoyo, including its structure and its component, would not change if the influence is limited in some extent. the influence that begins to make the yoyo change is defined as the stability boundary. as to a relationship between two persons, denoted by a and b, when one of them, supposed as a, satirizes the other person for some reason, b would accept the advance and their relationship holds out in all probability if b is given enough prestige; however, if a speaks ill of the person b now and then instead of proposing his advance by courtesy, his action would go beyond the stability boundary of their relationship and the relationship would break up. therefore, the relationship system has the inertia to keep the original state. the inertia is important for the stability of the yoyo; otherwise, the yoyo would be changed as long as it is affected by the environment, no matter how little the effect is. the influence boundary describes the stability degree, and the yoyo is zelong wang: modified yoyo model and its applications in chinese history 132 steadier if its influence boundary is larger. the inertia helps the yoyo to have a steady spin field, i.e. the yoyo has a steady behavior, and then the yoyo has a steady function to servicing for the human. property7 : the yoyo comes into being or perdition both for the poignant change of the spin field of the environment, i.e. all the systems outside the observed system. the development of the yoyo is important for the system research and it also reflects the phenomenon of everything’s movement. the development of a yoyo means the processing of the yoyo from an old state to a new state, and these change including the variety of the components, structure or attributes of the yoyo. in other words, the development of a yoyo can also be considered as the perdition of the “old” yoyo and the being of a “new” yoyo. therefore, if the perdition and the being of a yoyo are researched clearly, the development of the yoyo is also known by us. how does the yoyo come into the being or perdition? property 7 gives us the answer. when the spin field that the observed yoyo is located in changes greatly, the yoyo would receive great “force” from the spin field and it would be destroyed if the “force” exceeds the stability boundary of the yoyo. when two or more yoyos have interactions with each others, their composed spin field would greatly change, and the new yoyo would come into being among their poignant interaction. for example, in the chinese history of the spring and autumn and warring states, the country jin was overthrown under the condition of unstable political, economic and cultural positions, which had exceeded the stability boundary of the country jin, and three new countries came into being from the “old” country jin. there are two basic processing by which the old yoyo is destroyed and the new yoyo comes into being. one of them is the composition, i.e. two or more yoyos form a new yoyo, such as the composition of two companies; the other is the decomposition of an old yoyo into small new yoyos, such as the dividing of a country. in a word, the yoyo’s being or perdition lies in the poignant change of the spin field. 2.3.2 interaction between modified yoyos the modified yoyo model and its some basic properties have been introduced, and then the interaction between modified yoyos is studied, because all the objects in the world have connections with each other generally and they lead to all the change instead of the isolated system. moreover, the birth of new yoyo and the perdition of the old yoyo both are created by the interaction between yoyos. the clear knowledge of the interaction between yoyos is helpful for us to grasp the properties of the yoyo. property8 : the interaction between the yoyos depends on the spin field. 133 advances in systems science and applications (2013) vol.13 no.2 as is argued above, the spin field of the modified yoyo is used to describe the behavior of the yoyo and sfi scales the behavior intensity. when yoyo a is in the spin field of yoyo b, a would receive the action from b. for example, the competition between two companies means that their productions scrabble for the market, i.e. their spin fields fight with each other instead of the companies. another example is about the sun and the earth, both of which are considered as two systems. their spin fields, i.e. their gravitation fields, are the medium that they attract each other. property9 : the locations and the polarity directions of spin field where a yoyo is located determine the influence that the yoyo receives. the interaction between the yoyos depends on their spin field, and this thesis tells how a yoyo receives the influence by the spin field of another yoyo. as to the same two yoyos, if they have little connection with each other, their behavior on each other would have small intensity. as we known, where the behavior intensity is smaller, where the density of spin lines is smaller and the distance between two yoyos is larger, vice versa. because the polarity of the spin lines stands for the direction of the behavior ability, the direction of a yoyo is determined by its behavior. if two yoyos have the same behavior abilities, they have the same directions, i.e. the polarities, vice versa. if some of their behavior abilities are the same, some are opposite and the other are different, their polarities would have a angle that ranges from zero to π. for example, the workers work in the factory and the farmers work on the farm, so their behavior abilities, i.e. their work skills, nearly are different from each other. they are assumed to be vertical to each other, shown in fig.9. in fact, there also are farkers in our country in recent years, who work on the farm when the harvest is coming and in the factory without harvest, so its yoyo is located between the worker yoyo and the farmer yoyo. since the size of farkers, i.e. the product of their quantity and their working load, is less than that of the workers and that of the farmers, the length of the farker yoyo is less correspondingly. fig.9 the polarities of different yoyos property10 : superposition law if there are only two yoyos, they would be in the spin field of each other, which zelong wang: modified yoyo model and its applications in chinese history 134 is simple to analyze. when a yoyo is in the spin field of two or more yoyos at the same time or two or more yoyos are considered as a whole, how shall we seek the spin field of the whole from the yoyos? superposition theorem would tell us the answer. superposition law tells that the spin field of the whole yoyos is equivalent to the superposition of the spin field of each yoyo. the law means that the spin field of yoyos is independent with that of each others. please note that this law does not mean that the spin field of the yoyos does not have interaction with each other; instead the whole spin field is composed of each independent spin field of the yoyos. for example, the earth is in the gravitation field of the sun and the moon, and the gravitation of them independently effects on the earth; the whole gravitation felt by the earth is composed of the gravitation of the sun and the moon together, and the composition method of the spin lines is introduced by the spin line composition law in the following. property11 : spin line composition law the whole spin field is achieved as long as its spin lines are gotten. suppose that there are two yoyos in the field space, such as yoyo a and yoyo b in fig.10. at a randomly given point p in the field space, the spin lines of yoyo a and yoyo b are a and b with regard to and respectively. by vector composition laws, the whole spin line c at point p is composed by vectors a and b, shown in fig.10. fig.10 composition of spin field the spin line composition law reflects the nonlinearity of the interaction of the yoyos, i.e. the systems. the whole behavior of two yoyos is not the simple linear sum of each, and the direction and the intensity of the whole behavior has parallelogram relation with the original spin field. now another problem that seems hard to solve is the phenomenon “1 + 1 > 2”, i.e. the whole behavior 135 advances in systems science and applications (2013) vol.13 no.2 seems stronger than the sum of each individuals. in fact, there has been another new yoyo produced by the poignant interaction of the yoyos. property12 : yoyo composition law in fact, the yoyo composition law can be concluded by the superposition theorem and spin field composition law. however, the yoyo composition law uncovers the discipline of the composition of yoyos from a high level and it let us see the processing of new yoyo’s birth clearly. the law can be used to explain the combination of the factories, the combined armies, etc. the yoyo composition law tells that the composition of yoyos satisfies the vector composition law, shown in fig.11. yoyo a and b from the new yoyo c, and their spin field are piled in the field space; the spin field of the new yoyo c is gotten by the spin field composition law, and then we get the new yoyo c. the yoyo c also has parallelogram relation with the original yoyos. because the composition of yoyos does not simply sum the yoyos linearly, this composition reflects the nonlinear relation between yoyos. therefore, this law can be fit to the complexity of the systems. fig.11 composition of yoyos 3 modified yoyos of chinese history spring and autumn and warring states periods, from 770 bc to 221 bc, is in the interim from ancient civilization to middle ages civilization and is also a transformation with sharp social movements, complex political dispute, endless warfare and wonderful academic culture. spring and autumn began with the east-transfer of the king ping, and then the western zhou dynasty became weaker and weaker. because the vassals contended for the slavery and the strong country oppressed the weak, many wars began to act on the historical arena. the partition of the country jin by three vassals indicated the end of spring and autumn and the beginning of warring states. during this period, except for the appearance of reformation, the wars between countries were more and more for zelong wang: modified yoyo model and its applications in chinese history 136 more ground and wealth. in 238 bc, the first emperor of qin began to prepare for overthrowing other six countries; finally he established the first centralized multinational country in chinese history with the perdition of qi in 221 bc. the chinese history came into a new era. during the period of spring and autumn and warring states periods, there was a homology between different countries, because they were all located at the era from slave society to feudalism society. one of them was the wars that happened between numerous countries. during the 242 years of spring and autumn, there had been more than 170 vassal countries recorded in history; 36 vassals were killed ferociously; 52 vassal countries were overthrown by more than 480 wars, and the vassal meetings were more than 450. the wars became more and more and the size of wars were much larger than that in spring and autumn as the history came into warring states. in the ancient slave society with the underdeveloped science and technology, the number of the soldiers is an important factor for the war, so the total belligerent soldiers of both countries maybe reached tens of thousands even hundreds of thousands. therefore, the history of spring and autumn and warring states periods also can be called a martial history. in this section, both sides of the war are considered and modeled by the modified yoyo model as well as their interaction correspondingly. since these matters occurred more than two thousands years ago, there were more word depictions instead of integrated and detailed record data. however, the chinese history is modeled by the modified yoyo model to the best of our abilities. 3.1 military yoyos every new overlord was established with bleedings and sacrifice by the wars. during the period of spring and autumn and warring states, there were so many wars that lots of famous generals and specific example of a battle, such as the battle of zhongyuan, the battle of handan and the battle of changping, had been studied by scholars. how shall we predict the result of a war by the war factors? the yoyo model supplies a useful tool to study the war and the interactions between the both sides of the war can help us to know the result. generally speaking, there are many factors that commonly decide the result of the war and these factors can be changed at different periods. economy of a country is a key factor that has influence on its munitions, such as the food supplies and armament, that service for the army; moreover, in the ancient time, the country population is also decided by the economy. the science and technology, deciding the power of the weapon, are universal underdeveloped, so the weapons of different countries during spring and autumn and warring states are almost the same. although the daily training had been formed at that time, the abilities of a single soldier of different countries are also nearly the same. the morale of the soldiers is important for the war, and it is decided by the prompting actions 137 advances in systems science and applications (2013) vol.13 no.2 of the generals. heroical historical view pays great attentions on the heroes or the generals of a war, so they are considered as a factor of the war. according to the records in history[9], the situations of the war were always turned by the wisdom of the generals. now the military organization of a country in spring and autumn and warring states is modeled by the yoyo model. according to the analysis above, the number of soldiers and the wisdom of generals are the most important factors that decide the result of the war; and other factors are ignored for their similarities between different countries. assume that the population of the country is n and the number of its soldiers is n1, where n1 ≤ n . therefore, the number of the people that are arranged to work for the economy is n2, where n1 +n2 = n (5) generally speaking, n2 affects the battle effectiveness indirectly by affecting the economy of the country, because the food, the clothes and the weapons of the soldiers come from the economy contributed to by its people. so the change rate of the battle effectiveness p can be expressed as: dp dn1 = an1n2 (6) where α is a constant, which means that the battle effectiveness will increase with the growth of the numbers of the soldiers and the workers. please note that equ.(6) is only fit for the ancient wars instead of modern wars, because the former is much simpler than the latter. combining equ.(5) and equ.(6), we have: dp dn1 = an1(n −n1) (7) as we known, the army is used to protect its country and overcome other countries. so an important rule of the war is that the battle effectiveness aspired after by the sovran of the country is to overcome other countries and protect the security of its country instead of maximization of the power. therefore, the number of the soldiers is also affected by its battle effectiveness, i.e. when the battle effectiveness is strong, the number of the soldiers would decrease, vice versa. so we have: n1 = γ(p0 − βp ) (8) where p0 > 0, which avoids the appearance of negative battle effectiveness; and γ > 0, β > 0, which means that the country would decrease its soldiers if its battle effectiveness has became strong enough. taking equ.(7) into equ.(8), we have: dp dn1 = ap 2 +bp + c (9) zelong wang: modified yoyo model and its applications in chinese history 138 where a = −αγ2β2 < 0, b = 2αγ2p0β − αnγβ and c = αγnp0 − αγ2p 2 0 . the wisdom of generals is always a key factor. if one side of the war does not have enough stratagems, it sometimes would be defeated by the other side with a relative weak power. the battle effectiveness should be affected by a factor k ∈ [0, 1] about the wisdom of the generals. when the generals have great providence, stratagems and brightness, such as lianpo, limu and sunwu, the factor k=1; if the general is an armchair strategist, such as zhaokuo, the factor k=0; otherwise, the factor k ∈ (0, 1]). therefore, equ.(9) can be rewritten as: dp dn1 = k(ap 2 +bp + c) (10) how does the battle effectiveness keep steady between all the countries? when the battle effectiveness is balanceable, the possibility of wars is small, vice versa. so the stability is the focus that most of the scholars pay attentions to. the discriminant of equ.(9) is: ∆ = b2 − 4ac (11) so there are three cases: ∆ = 0, ∆ > 0 and ∆ < 0. the evolution models are also different at different cases. in the following, let us take the first case as an example. if ∆ = 0, there is only one balance, i.e. p1 = −2a/b, that all the countries do not invade each others. so we have: p (n1) = k ( p1 − 1 p2 +an1 ) (12) where p2 is the initial constant determined by the country condition. if p2 < 0, so n0 = −p2/a < 0. the battle effectiveness decreases with the increasing of the soldier number. when p2 > 0, n0 = −p2/a > 0. if c, the larger the soldier number, the stronger battle effectiveness. if n1 = n0, the discontinuity occurs to the battle effectiveness. if n1 > n0, the battle effectiveness would decrease as the soldier number’s rising. this processing indicates that the battle effectiveness could not blindly increase as the soldier number. by the restriction of economy of the country, the army would keep a balance or fluctuate at some size. now let us analyze the representation of the yoyo model of the military organization. the spin field, depicted by the spin lines, naturally denotes the behavior ability, i.e. the battle effectiveness and the resistance. the spin field is important for the yoyo to have connection with others, because the behavior abilities of a yoyo is given off or received by its spin field, just like invading other countries. except for the soldiers and the generals, there should be other factors that affect the military organization, such as the weapons, logistics and geography. when modeling the yoyo, they are all ignored for the puny differences between them. 139 advances in systems science and applications (2013) vol.13 no.2 this disposal is consistent with the known word records, because most of the wars recorded in history only come down to the soldiers, generals and the stratagems[10]. what’s more, the stronger the behavior ability is, the larger its spin field is, which leads to a larger size of the yoyo. when interacting with other yoyos, the yoyo with larger size would be the leading role. for example, the victory of the country zhao in the war of fei[9], with the leader of general limu, showed that the country zhao had stronger behavior ability than the country qin for the wonderful stratagem. the polarities of the yoyo reflect the directions of the behavior ability, i.e. the military organization gives off its action or receives the force from others. if two countries had the same battle goal and defeated their common enemy, their yoyos would have the same polarities. according to the definition of the size of the yoyo, the size of a yoyo means the battle effectiveness; in fact, it also can be measured by its munitions, i.e. its food supplies and armament, although this measurement is inconvenient in this paper. therefore, the yoyo of a military organization can be expressed in fig.12. fig.12 the yoyo of military organization 3.2 explanations of the yoyo’s characters the explanations of the yoyo’s characters are considered by the qualitative analysis, such as the figurative analysis method, of the historical events in spring and autumn and warring states. that the soldiers and the generals determine the military effectiveness means that the spin lines at different polarities have a corresponding relation. taking the country wu and chu as examples, when the king helv of wu paid great attentions to the development of the army and selected the militarist sunwu as general, the military effectiveness of wu became larger and larger and defeated the chu finally[10]. so we can conclude that the investment and the harvest, i.e. the received behavior and the external behavior, are equivalent. moreover, the stability of the military effectiveness is equivalent to that of the zelong wang: modified yoyo model and its applications in chinese history 140 yoyo, i.e. only if the military organization has a steady military effectiveness, the organization would be also steady, and vice versa. when the military organization was not defeated by others, its military effectiveness was not changed, i.e. it was steady. it was sure that the organization was still steady by the protection of the steady military effectiveness. if the military effectiveness of a organization was decreased by other organizations, the military effectiveness would be unsteady and so was the military organization. the country qi was nearly conquered thoroughly in the war of yan and qi, because its military effectiveness was extremely unsteady, only 7000 persons left after that war[10]. the change of the effectiveness would lead to the change of the size of the yoyo. when we select the military organization as the yoyo model, its behavior ability will be the research context and intuition. therefore, the military effectiveness is used to describe the size of the yoyo exactly. it seems that the number of the soldiers and the generals is fit to measure the yoyo’s size, but it can not capture the essentials of military organization. as to a given military organization, it has inertia to keep steady and it would be steady if the outside force acted on the yoyo does not exceed the stability boundary. it is well explained by the phenomenon that not all the wars lead to perdition of the military organization. for example, during the war of yan and qi, the country qi did not be conquered finally although there were only 7000 persons left; this means that the violent military shock of yan still did not exceed the boundary of qi’s perdition. on the other hand, this boundary of a yoyo can reflect the inertia of the stability of the yoyo, returning the example above, the qi had great attributes to keep its country, i.e. the structure and component of the yoyo. the furious change of the spin field, exceeding the inertia boundary, would lead to the birth of a new yoyo and the perdition of an old one. for example, the war of wu and yue had been lasting for more than 30 years[10]; it showed that when the change of the spin field could not exceed the inertia boundary, the yoyo, i.e. the country, could not be changed. finally the perdition of the country wu shows that the force acted on it by the country yue had exceeded its inertia boundary and the yoyo came to perdition. at the same time, the country yue became a new yoyo with the addition of wu’s land and wealths. 3.3 the interaction of the yoyos what about the interactions between different countries? the combination of countries reached the climax at the end period of warring states, such as the combination of six countries for resisting qin[9]. let us analyze this phenomenon by the yoyo model. the combination of six countries means that the six yoyos with different sizes are united together; by the composition laws of the yoyos, the new yoyo of the combination should be stronger than that of the qin, shown in 141 advances in systems science and applications (2013) vol.13 no.2 fig.13, why would the qin be the winner finally? fig.13 the perfect combination of six countries and qin the problem lies in the polarities of the six countries, and the fig.13 shows that the polarities of the six countries are the same. however, in history the six countries had never been actually united together, and they acted by the motivation of their own benefits instead of defeating their enemy qin. therefore, the polarities of the six yoyos were not consentaneous, shown in fig.14, and their combination was still weaker than the qin. moreover, the country qin tried itself to destroy the combination of the six countries by alienating these countries one by one, so there even were some countries that attacked their allies surreptitiously, such as the country wei. when the number surviving countries was fewer and fewer, they could not defeat the strongest country qin even if they shared a bitter hatred of the enemy qin for their puniness finally. in 221 bc, the country qin defeated the six countries and established a centralized multinational country in chinese history, i.e. a new yoyo appeared by their complicated interactions. fig.14 the factual combination of six countries and qin 4 conclusions this paper focuses on modifying the yoyo model and proposes the modified yoyo model. firstly, the modified yoyo model enriches the yoyo model. the concept of the spin field is endowed with new meanings, i.e. the behavior abilities of the yoyo, and the spin lines are used to depict the spin field exactly, including the intensity and polarity of the spin field and the size of the yoyo. the symbol denotation of the modified yoyo is convenient to describe the yoyos and their zelong wang: modified yoyo model and its applications in chinese history 142 interaction. the modified yoyo model has more basic concepts about the system. secondly, the characters of the modified yoyo are researched in detail. the stability, diversity and complexity of the system can be reflected by the modified yoyo model; the equipollence of the stability between the yoyo and its spin field, the spin line between two different polarities and the yoyo’s size between the behavior and the tsme is also discussed; finally, the perdition and the birth of yoyo are explained by the inertia of the stability effectively. at last, the interaction between yoyos is studied and the composition laws of the spin lines and the yoyos are proposed. in order to prove the practical usefulness of the modified yoyo model, the chinese history of spring and autumn and warring states, mainly about the military affairs, is explained by this model. according to the analysis results, the military organization of each country can be considered as a yoyo, and the basic characters of the yoyos and their interaction can be reflected by the word records in history. therefore, the modified yoyo model has the rationality, to some degree, as the complementarities of the yoyo model. the modified yoyo model has referred to many basic properties of the system, but these are only restricted to the word description, i.e. qualitative analysis, or the figurative analysis method instead of the quantitative analysis, which should be an open problem for the development of the yoyo model. another open problem is that since almost any objects can be seen as systems, how these systems are modeled by the yoyo model exactly and effectively. references [1] yi lin. (2008), systemic yoyos, impacts of the second dimension, published by taylor and francis. [2] quastler h. (1965), “general principles of systems analysis”, in theoretical and mathematical biology, ed. t. h. waterman and h. j. morowitz. new york: blaisdell publishing. [3] zadeh l. (1962), “from circuit theory to systems theory”, proc. ire, vol.50, pp.856-65. [4] von bertalanffy l. (1924), einfuhrung in spengler’s werk, literaturblatt kolnische zeitung, may. [5] ouyang s. c, k. zhang, l. p. hao, and l. r. zhou. (2005), “structural transformation of irregular time-series information and refined analysis on evolutions”, eng. sci. china, vol.7, pp.36-41. [6] yi lin. (2008), systemic yoyos, impacts of the second dimension, published by taylor and francis, pp.9-10. 143 advances in systems science and applications (2013) vol.13 no.2 [7] yi lin. (2008), systemic yoyos, impacts of the second dimension, published by taylor and francis, pp.11-14. [8] http://www.thenakedscientists.com/forum/index.php?topic=22668. [9] shouqian wang. (1992), zhanguoce full translation, guizhou press (china). [10] qingxiang meng. (1986), zhanguoce translation, heilongjiang people’s press (china). corresponding author zelong wang can be contacted at: zelong wang@163.com. advances in systems science and applications (2014) vol.14 no.2 101-128 the theory of parametric control of macroeconomic systems and its applications(ii) a. ashimov, zh. adilov, r. alshanov, yu. borovskiy and b. sultanov kazakh national technical university named after k. satpayev, 22 satpaev street, almaty, 050013, kazakhstan abstract this work consists of three parts and presents the recent results of development of the theory of parametric control of macroeconomic systems and some its applications for solving a number of concrete problems. keywords mathematical model, structural stability, parametrical identification, parametric control part 2. mathematical foundations of the parametric control theory 2.1 sufficient conditions for the existence of solutions for the problems on synthesis and choice of optimal parametric control laws 2.1.1 conditions for the existence of solution for the variational calculus problem on synthesis of optimal parametric control law of continuous dynamical system consider continuous controllable system* ẋ(t) = f(x(t), µ(t), a(t)), t ∈ [0, t ] (1) x(0) = x0 (2) where t is time; x = x(t) = (x1(t), ..., xm(t)) is a vector function of system’s state; µ = µ(t) = (µ1(t), ..., µq(t)) is a vector function of control; a = a(t) = (a1(t), ..., as(t)) is a known vector function; s0 = (x10, ..., x m 0 ) is an initial condition of the system, a known vector; f-a known vector function of its arguments. the problem of synthesis of optimal economic tools values consists in finding the extremum of the following criterion k = ∫ t 0 f (t, x(t))dt (3) where f is a known function, at the phase constraints x(t) ∈ x(t), t ∈ [0, t ] (4) where x(t) is a given set and at explicit control constraints u(t) ∈ u(t), t ∈ [0, t ] (5) * all of the formulated in this part results for non-autonomous dynamical systems remain true for autonomous dynamical systems as well, when an exogenous vector function a(·) is taken as constant. 102 a. ashimov: the theory of parametric control of macroeconomic systems and ... where u(t) is a given set. define the variational calculus problem on synthesis of optimal parametric control laws for the continuous dynamical system. problem 2.1. at given function a(·) to find the control u(·), satisfying the condition (5), so that dynamical system (1), (2)solution meets the condition (4) and gives the maximum (the minimum) to functional (3). to prove the solubility of the problem 2.1 we first need the single-valued solubility of cauchy problem (9), (10). to obtain this result we use known result from the theory of ordinary differential equations. let be given: number t > 0, metrizable compact u and continuous function φ : [0, t ] × rm × u → rm such that for any ρ ≥ 0 exists such σ ≥ 0, that the following inequality is true |φ(t, y, µ)− φ(t, y′, µ)| ≤ σ|y − y′|∀t ∈ [0, t ], y, y′ ∈ ρb, µ ∈ u (6) where is a unit ball in rm and there is such a constant η ≥ 0 that the following inequality is true |yφ(t, y, µ)| ≤ η(1 + |y|2)∀t ∈ [0, t ], y ∈ rm, µ ∈ u (7) consider cauchy problem ẏ(t) = φ(t, y(t), µ(t))∀t ∈ [0, t ] (8) y(0) = y0 (9) where y0 ∈ rm. the following statement is true ([1], lemma 4.1): lemma 2.1. given above mentioned constraints (6), (7) for any measurable mapping µ : [0, t ] → u the problem (8), (9) has the only solution y : [0, t ] → rm, satisfying the estimate |y(t)| ≤ ( |y0|2 + 2ηt )1/2 eηt∀t ∈ [0, t ] (10) the following statement is true. theorem 2.2. let function a(·) be continuous on segment [0, t ], u is a compact in rq, function f is continuous, for any ρ ≥ 0 there is such σ ≥ 0, that the following inequality is true |f(x, µ, a(t))− f(x′, µ, a(t))| ≤ σ|x− x′|∀t ∈ [0, t ], x, x′ ∈ ρb, µ ∈ u (11) and exists such a constant η ≥ 0 that the following inequality is true |xf(x, µ, a(t))| ≤ η(1 + |x|2)∀t ∈ [0, t ], x ∈ rm, µ ∈ u (12) advances in systems science and applications (2014) vol.14 no.2 103 then for any measurable mapping µ : [0, t ] → u , cauchy problem (2.1), (2.2) has the only solution x : [0, t ] → rm satisfying the estimate |(t)| ≤ ( |x0|2 + 2ηt )1/2 eηt∀t ∈ [0, t ] (13) proof. let us introduce notation: φ(t, x, µ) = f (x, µ, a(t)) and f functions continuity results in continuity of function φ. (6), (7) inequalities truth follows from relations (11), (12). thus the theorem statements follow from lemma 2.1. the theorem is proved. for proof of solubility of the problem 2.1 we use known result from the optimal control theory for systems, described by differential equations. let us specify a closed subset e ⊂ [0, t ] × rm × u . consider the system, described by cauchy problem (8), (9) when lemma 2.1 conditions are satisfied, herewith by control we mean a measurable mapping µ : [0, t ] → u . acceptable pair for the system in question is such a pair “state-control”, which satisfies the relations (8), (9) and inclusion (t, x(t), u(t)) ∈ e (14) let us specify carathéodory function φ with non-negative values on the set [0, t]× (rm × u). let us specify a functional i = t∫ 0 φ (t, x(t), u(t)) dt. problem 2.2. to find the minimum for the functional i on the acceptable pairs set of the system (8), (9). for any point (t, x) ∈ [0, t]×rm define a section et,x = {u ∈ u : (t, x, u) ∈ e}. let us specify a set γt,x = {φ(t, x, u) : (t, x, u) ∈ et,x} and a function g : [0, t ]× rm×rm → r using an equality g(t, x, y) = min{φ(t, x, u)| (t, x, u) ∈ e,φ(t, x, u) = y}. it is known the following statement about solubility of optimization problem ([1], statement 4.2), where r̄ implies the set of real numbers, supplemented with the values −∞ and +∞: lemma 2.3. let, when lemma 2.1 condition is satisfied, the set γt,x be convex for all t ∈ [0, t ].x ∈ rm, and the function g(t, x, ·) : rm → r̄ be convex for all (t, x) ∈ [0, t ]×rm. then the problem 2.2has a solution. let be a closure of the sum ∪ t∈[0,t ] x(t). the following statement is true. lemma 2.4. let be a compact, function f be continuouson [0, t ]×x. then function φ, specified by equality φ(t, x, u) = f0 − f (t, x) where f0 is the maximum of function f on the set [0, t ]×x satisfies lemma 2.3 conditions. proof. existence of the maximum for f0 follows from the weierstrass theorem. non-negativity of the determined above function φ is obvious. it is carathéodory function because of the continuity of f. at last, the convexity of function g(t, x, ·) 104 a. ashimov: the theory of parametric control of macroeconomic systems and ... for all (t, x) ∈ [0, t ]×rm is implemented, as function f does not depend on control u. the lemma is proved. let u be a closure of the sum ∪ t∈[0,t ]. for reduction the constraints (4), (5) to the form (14) we specify the set e ⊂ [0, t ] × rm × u the way to hold the relation: et0 = {(t, x, u) ∈ e| t = t0} = x(t0)× u(t0), t0 ∈ [0, t ]. lemma 2.5. let the mappings t → x(t), t → u(t) be continuous at each point t ∈ [0, t ] in the following sense: if the inclusions xk ∈ x(tk), uk ∈ u(tk) are true, where tk ∈ [0, t ], k = 1, 2, ... and convergence of the sequences is met tk → t, xk → x, uk → u, then the inclusions x ∈ x(t), u ∈ u(t) are true.then the set is closed. proof. consider such an sequence {(tk, uk, xk)} of elements of the set , that there is a convergence (tk, uk, xk) → (t, u, x). from the inclusion (tk, uk, k) ∈ e, because of specifying the set , follow the inclusions tk ∈ [0, t ], xk ∈ x(tk), uk ∈ u(tk), from closure of segment [0, t ] follows the inclusion t ∈ [0, t ]. using lemma conditions, ascertain that x ∈ x(t), u ∈ u(t) therefore, the inclusion (t, u, x) ∈ e is true, that is the set is closed. the lemma is proved. let us make sure that lemma conditions are not excessively restrictive. consider, for instance, a typical situation for scalar case, when the set x(t) = {x| a(t) ≤ x ≤ b(t)} , t ∈ [0, t ] is given, where functions and b are continuous. let the inclusion xk ∈ x(tk) be met, where tk ∈ [0, t ], k = 1, 2, ... and there is the convergence tk → t, xk → x, ipso facto, the inequalities a(tk) ≤ xk ≤ b(tk), k = 1, 2, ... are true. by proceeding here to the limit allowing for continuity of functions and b, we get the result a(t) ≤ x ≤ b(t) what results in that x ∈ x(t) thereby, the mapping t ∈ x(t) is continuous on the segment [0, t ]. analogously ascertains the truth of lemma 2.5 conditions and in more general cases, when boundaries of sets of acceptable control and condition values are continuous functions of time. for reduction the equation (1) to the form (8) it is enough to specify φ(t, x, u) = f(x, u, a(t)) then the set γt,x, appearing in lemma 2.2 conditions, defines in the following way γt,x = {f (x,w, a(t))|w ∈ u(t)} (15) denote by v the set of acceptable pairs “state-control” of the system (1), (2) in question at given known function , that is such vector function pairs (x, u) which satisfy the relations (1), (2), (4), (5). from lemma 2.3 directly follows the next statement. theorem 2.6. let function (·) be continuous on the segment [0, t], u be a compact in rq, function f be continuous in x × u × a and for any ρ ≥ 0 there is such σ ≥ 0, that the following inequality is true∣∣f(x, u, a(t))− f(x′, u, a(t)) ∣∣ ≤ σ ∣∣x− x′ ∣∣ , x, x′ ∈ ρb, u ∈ u (16) advances in systems science and applications (2014) vol.14 no.2 105 and there is such a constant η ≥ 0, that the following inequality is true |xf(x, u, a(t))| ≤ η ( 1 + |x|2 ) ∀t ∈ [0, t ], x ∈ rm, u ∈ u (17) let x be a compact, function f be continuous on [0, t ]×x. let, moreover, the mappings t → x(t), t → u(t) be continuous for t ∈ [0, t ] in the following sense: if the inclusion xk ∈ x(tk), uk ∈ u(tk) are true, where tk ∈ [0, t ], k = 1, 2, ... and there is convergence of sequences tk → t, xk → x, uk → u, then the inclusions x ∈ x(t), u ∈ u(t) are true. then, in the case of non-emptiness of the set va and convexity of the set γt,x for all t ∈ [0, t ], x ∈ x(t) the problem 2.1 has a solution in the class of measurable functions. 2.1.2 conditions for the existence of solution for the variational calculus problem on choice (among given finite algorithms set) of optimal parametric control law of continuous dynamical system the continuous controllable system in question (1), (2). control u is chosen here among the set of given control laws: uj(t) = gj (v, x(t)) , t ∈ (0, t ), j = 1, ..., r (18) here gj is known vector function of its arguments; υ = (υ1, ..., υl) is vector of control parameters. on control parameters lay the constraints like v ∈ v (19) where v is some subset of the space rl. moreover, it is assumed that control parameters should be such that corresponding control law (18) satisfies the condition (5), that is the following inclusion is met gj (v, x(t)) ∈ u(t), t ∈ (0, t ) (20) where u(t) is given set. there are phase constraints on the system: x(t) ∈ x(t), t ∈ (0, t ) (21) where x(t) is given set. consider optimality criteria kj = kj(a, v) = ∫ t 0 f [t, xj(t)]dt (22) where xj = xj(t) = ( x1j (t), ..., x m j (t) ) is solution of cauchy problem (1), (2) at given function (·) and control u = uj(t) = ( u1j (t), ..., u q j(t) ) , that is for chosen 106 a. ashimov: the theory of parametric control of macroeconomic systems and ... j-th control law (26). consider the next subsidiary extremal problem: problem 2.3*. at given function a(·) for each of r control laws to find such a vector of control parameters v, that corresponding it solution x = xj of the problem (1), (2) with control law u = uj , determining by formula (18), satisfies the conditions (19)-(21) and gives the maximum for the functional (22). define the following variational calculus problem on choice (among given finite algorithms set) of optimal parametric control law for non-autonomous continuous system. problem 2.3. given known function a(·) among all optimal control laws in sense of the problem 2.3* to choose the one, which corresponds to the maximal optimality criterion value (22). substituting into the equation (1) control value from formula (19), we get ẋ(t) = f ( x(t), gj (v, x(t)) , a(t) ) , t ∈ (0, t ) (23) where for short we omit the index j in denoting condition function of the system, corresponding to given control law. denote by w j a a set of acceptable pairs “state-control parameter” of the system in question, that is such pairs (x, v), which satisfy both the equalities (23), (2), and the inclusions (19)-(21). thus, the problem 2.3* comes to the maximization of functional k = ∫ t 0 f [t, x(t)]dt on the set w j a . ascertain first a solubility of cauchy problem (23), (2). the following theorem directly results from lemma 2.1. theorem 2.7. let the function (·) is continuous on the segment [0, t ], sets u and v are compact, functions f and gj are continuous, for any ρ ≥ 0 there exist such σ ≥ 0, χ ≥ 0, that the following inequalities are true |f(x, u, a(t))−f(x′, u′, a(t))| ≤ σ(|x−x′|+|u−u′|)∀t ∈ [0, t ], x, x′ ∈ ρb, u, u′ ∈ u (24)∣∣gj (v, x)−gj ( v, x′ )∣∣ ≤ χ ∣∣x− x′ ∣∣∀x, x′ ∈ ρb, v ∈ v (25) and exists such a constant η ≥ 0 that the following inequalitiy is true |xf(x,gj(v, x), a(t))| ≤ η(1 + |y|2)∀t ∈ [0, t ], x ∈ rm, v ∈ v (26) then for any v ∈ v the problem (2), (2) has the only solution x : [0, t ] → rm,satisfying the estimate |(t)| ≤ ( |0|2 + 2ηt )1/2 eηt∀t ∈ [0, t ] (27) ascertain solubility of the problem 2.3*. the following theorem is true. theorem 2.8. assume that when the conditions of the theorem 2.7 are satisfied for given function (·) and given value of j ∈ {1, ..., r} the setw j a is not empty, advances in systems science and applications (2014) vol.14 no.2 107 the sets v, u, x are compact, the sets x(t), u(t) are compact for all t ∈ [0, t ], and function f is continuous. then the problem 2.3* has a solution. proof. from the condition (21) and continuity of function f follows that this function is bounded. then, because of non-emptiness of the set w j a , upper boundary sup k of function k exists on set w j a of acceptable pairs of the system in question. then there exists such an elements sequence {(xk, vk)} from the set w j a , that convergence takes place (here kk = ∫ t 0 f [t, xk(t)]dt) kk → supk (28) the inclusion (xk, vk) ∈w j a assumes the following relations to be true ẋk(t) = f ( xk(t), gj (vk, xk(t)) , a(t) ) , t ∈ (0, t ) (29) xk(0) = x0 (30) xk(t) ∈ x(t), t ∈ (0, t ) (31) vk ∈ v (32) gj (vk, xk(t)) ∈ u(t), t ∈ (0, t ) (33) from the theorem 2.7 results the estimate |xk(t)| ≤ (|x0|2 + 2ηt ) 1 2 eηt∀t ∈ [0, t ] (34) thus, using bolzanocweierstrass theorem allowing for boundedness of set v, after subsequence separation we get the convergences k(t) → (t)∀t ∈ [0, t ] (35) vk → v (36) limiting value v satisfies the inclusion (19) because of closure of set v and condition (32). using the conditions (35), (36) and continuity of function gj , we get the convergence f ( xk(t), gj (vk, xk(t)) , a(t) ) → f ( x(t), gj (v, x(t)) , a(t) ) , t ∈ (0, t ) (37) function xk satisfies the integral relation xk(t) = x0 + ∫ t 0 f ( xk(τ), gj (vk, xk(τ)) , a(τ) ) dτ, t ∈ (0, t ) (39) by proceeding here to the limit allowing for the conditions (35), (38), we get the equality x(t) = x0 + ∫ t 0 f ( x(τ), gj (v, x(τ)) , a(τ) ) dτ, t ∈ (0, t ) (40) 108 a. ashimov: the theory of parametric control of macroeconomic systems and ... by differentiating both parts of the equality (40) with respect to t, we ascertain that the equation (1) is true. consequently, after proceeding to the limit in equality (30) taking into account the condition (35), we ascertain that the initial condition (2) is true. thus, the limiting pair (x, v) is acceptable. continuity of function f from the condition (35) results in convergence f [t, xk(t)] → f [t, x(t)] , t ∈ [0, t ] (41) therefore ∫ t 0 f [t, xk(t)]dt→ ∫ t 0 f [t, x(t)]dt (42) from (42) and (28) follows the equality t∫ 0 f [t, x(t)]dt = supk. thus, we found the pair (x, u) ∈ w j a for which the upper boundary of functional k(on the set of all acceptable pairs of the system in question) is obtained. the theorem is proved. obviously, the problem 2.3 comes to the search of maximum of function φa(j) = max (x,v)∈w j a kj on finite set {1, ..., r} at given function (·). the maximal value of considered functional at given function (·) and j-th control law is placed on the right side of previous equality. the existence of this value is proved in the theorem 2.8. as the maximum of function on finite set is obtained always, we get the following statement. theorem 2.9. when theorem 2.8 conditions are satisfied, the problem 2.3 has a solution. 2.1.3 conditions for the existence of solution for the variational calculus problem on synthesis of optimal parametric control law of discrete dynamical system consider discrete non-autonomous controllable system x(t+ 1) = f(x(t), u(t), a(t)), t = 0, 1, ..., n− 1 (43) x(0) = x0 (44) where t is time. here x = x(t) = ( x1(t), ..., xm(t) ) is vector function (of discrete argument) of systems state; u = u(t) = ( u1(t), ..., uq(t) ) is control, vector function of discrete argument; a = a(t) = ( a1(t), ..., as(t) ) is known vector function of discrete argument; x0 = ( x10, ..., x m 0 ) is initial condition of the system, known vector; f is known vector function of its arguments. the problem of choosing optimal economic tools values consists in finding the extremum of the following criterion k = n∑ t=1 f [t, x(t)] → max (min) (45) advances in systems science and applications (2014) vol.14 no.2 109 where f is known function, at phase constraints on system (43)-(44) solution like x(t) ∈ x(t), t = 1, ..., n (46) where x(t) is given set, and following constraints on control: u(t) ∈ u(t), t = 0, 1, ..., n− 1 (47) where u(t) is given set. define the variational calculus problem on synthesis of optimal parametric control laws for discrete dynamical system. problem 2.4. given known function a(·), to find control u(·), satisfying the condition (47), so that dynamical system solution (43), (44), corresponding it, satisfies the condition (46) and gives the maximum (minimum) for functional (45). denote by va the set of acceptable pairs “state-control” of the system in question at given known function a(·), that is such pairs of vector functions (x, u), which satisfy the relations (43), (44), (46), (47). introduce notations: x = n∪ t=1 x(t), u = n−1∪ t=0 u(t). the following statement is true. theorem 2.10. let for given function a(·) the set va be not empty; sets x(t) and u(tc1) are closed and bounded for all t = 1, ..., n; mapping f is continuous by the first two arguments on set x × u , and function f is continuous by the second argument on set x. then the problem 2.4 has a solution. proof of this theorem is based on properties of functions which continuous on compact. 2.1.4 conditions for the existence of solution for the variational calculus problem on choice (among given finite algorithms set) of optimal parametric control law of discrete dynamical system consider discrete non-autonomous controllable system (43), (44). the following phase constraints are imposed on this system: x(t) ∈ x(t), t = 1, ..., n (48) where x(t) is given set. in equation (43) control is chosen among the set of given control laws: uj(t) = gj (v, x(t)) , t = 1, ..., n, j = 1, ..., r (49) where gj is known vector function of its arguments, v = (v1, ..., vl) is vector of control parameters. the following constraints are imposed on control parameters v ∈ v (50) 110 a. ashimov: the theory of parametric control of macroeconomic systems and ... where v is some subset of space rl. moreover, we will assume that control parameters should be such that corresponding control law (49) satisfies the condition gj (v, x(t)) ∈ u(t), t = 0, ..., n− 1 (51) where u(t) is given set. consider the following optimality criteria: kj = kj(a, v) = n∑ t=1 f [t, xj(t)] (52) where xj = xj(t) = ( x1j (t), ..., x m j (t) ) is problem (43), (44) solution at given function (·) and control u = uj(t) = ( u1j (t), ..., u q j(t) ) , that is for chosen j-th control law. define the next subsidiary extremal problem: problem 2.5*. given known function (·) for each of r control laws to find such a vector of control parameters v, that corresponding it solution x = xj of problem (43), (44) with control law u = uj , determining by formula (2.49), satisfies the conditions (48), (50), (51) and gives the maximum for functional (52). the following problem is called as variational calculus problem on choice among given finite set of algorithms of optimal parametric control laws for discrete non-autonomous system. problem 2.5. given known function (·) among all optimal control laws in sense of the problem 2.5* to find the one, which corresponds to the maximal optimality criterion value (52). ascertain first the solubility of the problem 2.5* for fixed control law. substituting into equation (43) the control value from formula (49), we get x(t+ 1) = f ( x(t), gj (v, x(t)) , a(t) ) , t = 0, 1, ..., n− 1 (53) where for short we omit the index j in denoting condition function of the system, corresponding to j-th control law. denote by w j a the set of acceptable pairs “state-control parameter” of the system in question, that is such pairs (x, v) which satisfy both the equalities (44), (53), and the inclusions (48), (49), (50). thus, the problem 2.5* comes to the maximization of functional kj = on the set w j a . let the sets and u be determined by relations x = n∪ t=1 x(t), u = n−1∪ t=0 u(t). the following theorem is true. theorem 2.11. let given known function (·) and j-th control law the set w j a is not empty, the sets v, x(t) and u(t − 1) are closed and bounded for all t = 1, ..., n, function f is continuous by the first two arguments on the set x × u advances in systems science and applications (2014) vol.14 no.2 111 function gj is continuous on the set v ×x and function f is continuous by the second argument on the set . then the problem 2.5* has a solution. proof of this theorem is based on property of continuous on compact functions to reach their maximal and minimal values. as on finite set the maximum of function is obtained always, we get the following statement. theorem 2.12. when the theorem 2.11 condition is satisfied, the problem 2.5 has a solution. 2.1.5 conditions for the existence of solution for the variational calculus problem on synthesis of optimal parametric control law of discrete stochastic dynamical system consider discrete stochastic controllable system like x(t+ 1) = f(x(t), u(t), a) + ξ(t), t = 0, 1, ..., n− 1 (54) x(0) = x0 (55) here x = x(t) ∈ rm is function of the (54), (55) systems state, random vector function of the discrete argument (vector random process); x0 is initial condition of the system, deterministic vector; u = u(t) ∈ rq is vector of controllable parameters, vector function of the discrete argument; a ∈ rs is vector of uncontrollable parameters, a ∈ a,a ⊂ rs is given set. ξ = ξ(t) = (ξ1(t), ..., ξm(t)) is known vector random process, expressing noises (as such a noise can be, for instance, gaussian noise); f is known vector function of its arguments. specify optimality criterion, which is subject to maximization for the present problem is of the form k = e { n∑ t=1 ft(x(t)) } (56) here ft are known functions, e is mathematical expectation. there are phase constraints on the system: e[x(t)] ∈ x(t), t = 1, ..., n (57) where x(t) is given set. determined earlier constraints on control hold in the problems considered further: u(t) ∈ u(t), t = 0, 1, ..., n− 1 (58) where u(t) ∈ rq is given set. define the following variational problem, called as the variational calculus problem on synthesis of optimal parametric control law for discrete stochastic dynamical system. problem 2.6. given known vector of uncontrollable parameters a ∈ a to find the parametric control law u, satisfying the condition (58), so that dynamical 112 a. ashimov: the theory of parametric control of macroeconomic systems and ... system (54), (55) solution corresponding it satisfies the condition (57) and gives the maximum for functional (56). determine the set of acceptable controls uad for system under study in terms of the set of such control laws u(t), satisfying the constraint (58), for which the mathematical expectation e[x(t)] of corresponding solution of stochastic system satisfies the inclusion (57). the following theorem is true. theorem 2.13. let in the problem 2.6 given a ∈ a for any t = 1, ..., n random variables ξ(t) are absolutely continuous and has zero mathematical expectations, the sets x(t), u(t) are closed and bounded for all t, function f satisfies the lipschitz condition, and functions ft are continuous by lipschitz. functions f (for u ∈ u and a ∈ a) and ft by absolute value does not exceed some linear relatively |x| functions. then, if the set of acceptable controls uad is not empty, the problem 2.6 will be soluble. proof. according to the weierstrass theorem, continuous function on nonempty bounded set reaches its maximum. thus, it is enough to show that the multivariable function k = k(u), determined by (56), is continuous, and the set uad is closed and bounded. its non-emptiness is one of the conditions of the theorem. we will show that mathematical expectations of the magnitudes, entering into phase constraint (57) are exist. indeed, according to equation (54), we have e[x(t+ 1)] = e[f(x(t), u(t), a)] + e[ξ(t)] the second summand of the right-hand-side of this equality has sense under the theorem conditions, and the first is calculated by formula e[f(x(t), u(t), λ)] = ∫ rn f(ω, u(t), a)px(t)(ω)dω if the last integral absolutely converges (here by px(t) is denoted probability density function of random variable x(t)). the latter fact is really takes place under constraints on increase of function f and existence of mathematical expectation of the magnitude x(t) for any t = 1, ..., n (this fact is tested using the mathematical induction method). existence of mathematical expectation on the right-hand-side of equality (56) results from constraints on increase of function ft and existence of mathematical expectation of the variable x(t). let convergence of vectors uk → u, uk ∈ uad takes place. from equation (54) follows the equality |xk(t+ 1)− x(t+ 1)| = |f(xk(t), uk(t), a)− f(x(t), u(t), a)| advances in systems science and applications (2014) vol.14 no.2 113 where uk and x are solutions of the problem (2.54), (2.55) at controls uk and u, correspondingly. then the following relation is true |xk(t+ 1)− x(t+ 1)| ≤ lf [|xk(t)− x(t)|+ |uk(t)− u(t)|] where lf is the lipschitz constant of function f . by repeating analogous reasoning and taking into account that under the condition (55), xk(0) = x(0) we will have |xk(t+ 1)− x(t+ 1)| ≤ (lf ) 2|xk(t− 1)− x(t− 1)|+ (lf ) 2|uk(t− 1)− u(t− 1)|+ lf |uk(t)− u(t)| ≤ t∑ s=0 (lf ) s+1|uk(t− s)− u(t− s)| ≤ εk where εk ≤ 0 under k → ∞. by denoting the maximal of the lipschitz constants of functions ft by lf for t = 1, ..., n we get the estimate |ft[xk(t)]− ft[x(t)]| ≤ lf εk after calculating mathematical expectations of both of parts of this inequality, we get inequality e{|ft[xk(t)]− ft[x(t)]|} ≤ lf εk this implies that e{|ft[xk(t)]−ft[x(t)]|} → 0 and convergence of the sequence in question e{|ft[xk(t)] − ft[x(t)]|}. hence, function k = k(u) is continuous because of (56). boundedness of the set uad results from boundedness of the set u(t). closure of the set uad results from continuity of the mapping uad → x, determined by defining of the set uad and compactness of the set x(theorem about closure of complete preimageof compact at continuous mapping). now, existence of solution for the problem under study follows from the weierstrass theorem. the theorem is proved. 2.1.6 conditions for the existence of solution for the variational calculus problem on choice (among given finite algorithms set) of optimal parametric control law of discrete stochastic dynamical system controllable discrete dynamical system with given additive noise, described by equations (54), (55) with phase constraints (57) is considered again in the following parametric control problem for discrete dynamical system. in this problem control is chosen among the set of given control laws: uj(t) = gj(v, x(t)), t = 0, ..., n− 1, j = 1, ..., r (59) 114 a. ashimov: the theory of parametric control of macroeconomic systems and ... where gj is known vector function of its arguments, v = (v1, ..., vl) is parameters vector of control law gj . on adjustable coefficients v impose the constraints v ∈ v (60) where v is compact in space rl. moreover, it is assumed that parameters of control law should be such that corresponding control law (59) satisfies the condition (58), that is the following inclusion would be met e [ gj(v, x v j (t)) ] ∈ u(t) (61) here xvj is solution of problem (54), (55) at chosen v coefficient values, uncontrollable parameter a and j-th parametric control law. consider optimality criteria kj = kj(v, a) = e { n∑ t=1 ft(x v j (t)) } (62) define the following variational problem, called as the variational calculus problem on choice of optimal parametric control law for discrete stochastic dynamical system. problem 2.7. given known vector of uncontrollable parameter a ∈ a to find each of r control laws to find such a vector of adjustable coefficients v, so that corresponding it solution x = xj of problem (54), (55) with control law u = uj , determined by formula (59), satisfies the conditions (60), (61) and gives the maximum for functional (62) with consequent choice of the best of found optimal control laws, i.e. the one, to which corresponds the maximal value of optimality criterion. now we get sufficient conditions for the existence of solution of the problem 2.7. denote by xvj solution of the system (54), (55) for chosen j-th parametric control law (59), its adjustable coefficient v and parameter α: xvj (t+ 1) = f ( xvj (t), gj ( v, xvj (t) ) , a ) + ξ(t), t = 0, 1, ..., n− 1 (63) xvj (0) = x(0) (64) for considered problem denote the set of acceptable values of adjustable coefficients as the set v j ad, consisting of such values v ∈ v satisfying the condition (60), for which corresponding solution of the problem (63), (64) satisfies the inclusions e [ gj(v, x v j (t)) ] ∈ u(t), t = 0, 1, ..., n− 1 (65) advances in systems science and applications (2014) vol.14 no.2 115 e [ xvj (t) ] ∈ x(t), t = 1, ..., n− 1 (66) we will call the problem 2.7 as nontrivial, if the set v j ad, corresponding it, is not empty and contains some open set for each j = 1, .., r. theorem2.14. let in the problem 2.7 a ∈ a, the sets u(t), x(t), v are compact, functions f , gj , ft satisfy the lipschitz condition. these functions satisfy constraints on increase as well: functions |f(x,gj(v, x), a)|, |ft(x)| do not exceed linear relatively |x| functions uniformly by v ∈ v . random variable ξ(t) is absolutely continuous and has zero mathematical expectation. then in the case of non-emptiness of the sets v j ad the problem 2.7 has a solution. proof. it is enough to ascertain that all of the functions (62) kj are continuous, and all of the sets v j ad are closed and bounded, where j = 1, ..., r. the existence of all used below mathematical expectations is proved in the same way as when proving the theorem 2.13. allowing for mathematical expectation additivity, we find the values of kj = kj(v) = n∑ t=1 e { ft[x v j (t)] } whence it follows inequality for v, w ∈ v j ad |kj(v)−kj(w)| = | n∑ t=1 e { ft[x v j (t)] } − n∑ t=1 e { ft[x w j (t)] } | ≤ n∑ t=1 |e { ft[x v j (t)] } − e { ft[x w j (t)] } | from relations (63), (64) follows that |xvj (t+ 1)− xwj (t+ 1)| = ∣∣f (xvj (t), gj ( v, xvj (t) ) , a ) − f ( xwj (t), gj ( w, xwj (t) ) , a )∣∣ ≤ lf [ |xvj (t)− xwj (t)|+ |gj ( v, xvj (t) ) −gj ( w, xwj (t) ) | ] where lf is the lipschitz constant of function f . after denoting by la the maximal one of the lipschitz constants of function gj , we will get inequality |xvj (t+1)−xwj (t+1)| ≤ lf (1+la)|xvj (t)−xwj (t)|+lflg|v−w|, t = 0, 1, ..., n−1 taking into account equality |xvj (0)− xwj (0)| = 0, we get the estimate |xvj (t+ 1)− xwj (t+ 1)| ≤ lflg t∑ l=0 [lf (1 + lg)] l |v − w| ≤ β|v − w|∀v, w ∈ v j ad 116 a. ashimov: the theory of parametric control of macroeconomic systems and ... where β = lflg t∑ l=0 [lf (1 + lg)] l after denoting by lf the maximal one of the lipschitz constants of function ft, we will have∣∣f [xvj (t)]− f [xwj (t)] ∣∣ ≤ lf ∣∣xvj (t)− xwj (t) ∣∣ ≤ lfβ|v − w|∀v, w ∈ v j ad thus, in the case of sufficient smallness of difference of adjustable coefficients v and w the values of xvj (t) and x w j (t) (and the same f [xvj (t)] and f [x w j (t)]) will be arbitrary close to each other. determine converged sequence w = vk → v. then, after finding mathematical expectations of lhs and rhs of the last inequality, we will get the following inequality: e [∣∣f [xvj (t)]− f [xwj (t)] ∣∣] ≤ lfβ|v − w| hence results the following convergence e { ft [ xvkj (t) ]} → e { ft [ xwk j (t) ]} from which follows continuity of function kv j . as v j ad ⊂ v , then all of the sets v j ad are bounded. closure of the sets v j ad results from closure of the sets u(t), x(t + 1), of proved above continuity of mappings v → e[xvj (t + 1)], v → e[gj(v, x v j (t))] and determining the set v j ad as complete preimageof pointed sets at continuous mappings. the theorem statement follows from the weierstrass theorem about reaching continuous on compact function of its upper boundary. 2.2 sufficient conditions for continuous dependence of optimal criteria values of the parametric control problems on uncontrollable functions in this section, within the framework of developing the 4th component of the parametric control theory, there are derived sufficient conditions of continuous dependence on uncontrollable functions a(·) of optimal criteria values for all of considered above parametric control problems for non-autonomous deterministic dynamical systems and sufficient conditions of continuous dependence on uncontrollable parameters of optimal criteria values for considered above parametric control problems for autonomous stochastic dynamical systems. all of the results defined for non-autonomous systems remain true for autonomous dynamical systems as well, where exogenous vector function a(·) is taken as constant. the following theorem determines sufficient conditions of continuous dependence of optimal criterion values for the problem 2.4. theorem 2.15. assume that when conditions of the theorem 2.10 are met, advances in systems science and applications (2014) vol.14 no.2 117 in neighborhood (in euclidean topology) of function a(·) function f is continuous by the third argument and satisfies the lipschitz condition by the first argument on uniformly by the second and third arguments.then optimal criterion value for the problem 2.4 continuously depends on uncontrollable function at the point a(·). proof. let the following convergence occurs ak → abrsn (67) according to the theorem 2.10 the problem 2.4 given the values of uncontrollable functions and ak has a solutions, which we denote, correspondingly, by u and uk. denote by x[b, w] solution of condition equation (43), (44), corresponding to uncontrollable function b and control w. ipso facto the function x[b, w] satisfies relations x [b, w] (t+ 1) = f (x [b, w] (t), w(t), b(t)) , t = 0, 1, ..., n− 1 (68) x [b, w] (0) = x0 (69) then the following inequalities are true 0 ≥ k (x [a, uk])−k (x [a, u]) ,k (x [ak, uk])−k (x [ak, u]) ≥ 0 where by k(y ) is denoted the value of criterion k at given value of system state . consequently, we get the following relations: 0 ≤ k (x [a, u])−k (x [a, uk]) ≤ k {(x [a, u])−k (x [ak, u])}+ {(x [ak, u])− (x [k, uk])}+ {k(x [ak, uk])−k (x [a, uk])} ≤ 2 sup w∈u |k(x [a,w])−k (x [ak, w])| (70) from the conditions (68) we derive inequalities |x [ak, w] (t+ 1)− x [a,w] (t+ 1)| = |f (x [ak, w] (t), w(t), ak(t))− f (x [a,w] (t), w(t), (t))| ≤ |f (x [ak, w] (t), w(t), ak(t))− f (x [ak, w] (t), w(t), (t))| + |f (x [ak, w] (t), w(t), a(t))− f (x [a,w] (t), w(t), a(t))| ≤ sup ∈x,φ∈u |f (y, φ, ak(t))− f (y, φ, a(t))|+ l |x [ak, w] (t)− x [a,w] (t)| , t = 0, 1, ..., n− 1, where l is the lipschitz constant of function f by the first argument, which does not depend on w. consequently, we get the following relation |ψk(t+ 1)| ≤ ηk + l |ψk(t)| , t = 0, 1, ..., n− 1 118 a. ashimov: the theory of parametric control of macroeconomic systems and ... where ψk(t) = x [ak, w] (t)− x [a,w] (t), t = 0, 1, ..., n− 1, ηk = max t=0,1,...,n−1 sup ∈x,φ∈u |f (y, φ, ak(t))− f (y, φ, a(t))| then we determine relations |ψk(t+ 1)| ≤ ηk + l|ψk(t)| ≤ ηk + l(ηk + l|ψk(t− 1)|) = (1 + l)ηk + l2|ψk(t− 1)| ≤ (1 + l)ηk + l2(ηk + l|ψk(t− 2)|) = (1 + l+ l2)ηk + l3|ψk(t− 2)| ≤ ... ≤ t∑ r=0 lrηk + lt+1|ψk(0)|, t = 0, 1, ..., n− 1 from initial condition (69) follows that |ψk(0)| = 0. then the following estimate is true |ψk(t+ 1)| ≤ ηk t∑ r=0 lr, t = 0, 1, ..., n− 1 from which follows that |x[ak, w](t+ 1)− x[a,w](t+ 1)| ≤ ηk t∑ r=0 lr, t = 0, 1, ..., n− 1 note that right-hand-side of this inequality does not depend on w. from the condition (67) follows the convergence ηk → 0 under continuity of function f by the third argument, as well as closure and boundedness of the sets x and u . then we derive from the last inequality that x [ak, w] (t) → x [a,w] (t) uniformly by w, t = 1, ..., n− 1. hence under continuity of function f follows that f [t, x [ak, w] (t)] → f [t, x [a,w] (t)] uniformly by w, t = 1, ..., n− 1.a therefore, k(x[ak, w]) → k(x[a,w]) uniformly by w, t = 1, ..., n− 1.a whence the convergence sup w∈u |k(x [a,w])−k (x [ak, w])| → 0 (71) follows. taking the limits in inequality (70) and taking into account the conditions (58), we will have k (x [a, uk]) → k (x [a, u]) (72) advances in systems science and applications (2014) vol.14 no.2 119 the following estimate is true. |k(x [ak, uk])− (kx [a, u])| ≤ |k(x [ak, uk])− (kx [a, uk])| + |(kx [a, uk])− (kx [a, u])| taking the limits and taking into account the conditions (71), (72), we will determine the convergence k (x [ak, uk]) → k (x [a, u]) hence under defining controls u and uk results that the maximal value of criterion k, corresponding to uncontrollable function ak, converges to its maximum, corresponding to the extreme value of uncontrollable function . the theorem is proved. determine now sufficient conditions of continuous dependence on uncontrollable functions of optimal criterion values for the variational calculus problem 2.5 on choice (among given finite algorithms set) of optimal parametric control laws based on discrete non-autonomous dynamical system. determine first corresponding result for the problem 2.5*. theorem 2.16. assume that when conditions of the theorem 2.11 are met in neighborhood of point function f is continuous by the third argument and satisfies the lipschitz condition by the first two arguments on x×u uniformly by the third argument, and function gj satisfies the lipschitz condition by the second argument on uniformly by the first argument.then optimal criterion value for the problem 2.5* continuously depends on uncontrollable function at point . proof. let the convergence occurs ak → abrsn (73) according to the theorem 2.11 the problem 2.5* at values of uncontrollable functions and ak has a solution, which we denote, correspondingly, by v and vk. denote by x[b, w] solution of state equations (53), (44), corresponding to uncontrollable function b(·) and controllable parameter w. ipso facto function x[b, w] meets relations x [b, w] (t+ 1) = f ( x [b, w] (t), gj (v, x [b, w] (t)) , b(t) ) , t = 0, 1, ..., n− 1 (74) x [b, w] (0) = x0 (75) analogously with inequality (70) the following relation is determined 0 ≤ k (x [a, v])−k (x [a, vk]) ≤ 2 sup w∈v |k(x [a,w])−k (x [ak, w])| (76) 120 a. ashimov: the theory of parametric control of macroeconomic systems and ... after denoting a = x [a,w] , ak = x [ak, w] from equalities (74) we derive inequalities |xk(t+ 1)− x(t+ 1)| = ∣∣f (xk(t), gj (w, xk(t)) , ak(t) ) − f ( x(t), gj (w, x(t)) , a(t) )∣∣ ≤ ∣∣f (xk(t), gj (w, xk(t)) , ak(t) ) − f ( xk(t), gj (w, xk(t)) , a(t) )∣∣ + ∣∣f (xk(t), gj (w, xk(t)) , a(t) ) − f ( x(t), gj (w, x(t)) , a(t) )∣∣ ≤ sup y∈x,φ∈u |f (y, φ, ak(t))− f (y, φ, a(t))| + l [ |xk(t)− x(t)|+ ∣∣gj (w, xk(t))−gj (w, x(t)) ∣∣] ≤ sup y∈x,φ∈u |f (y, φ, ak(t))− f (y, φ, a(t))| + l(1 +m) |xk(t)− x(t)| , t = 0, 1, ..., n− 1, where l is the lipschitz constant of function f by the first two arguments, and is the lipschitz constant of function gj by the second argument. consequently, we get inequality |ψk(t+ 1)| ≤ ηk +n |ψk(t)| , t = 0, 1, ..., n− 1 where n = l(1 +m) ψk(t) = xk(t)− x(t), t = 0, 1, ..., n− 1, ηk = max t=0,1,...,n−1 sup y∈x,φ∈u |f (y, φ, ak(t))− f (y, φ, a(t))| using technique, described above in proof of the theorem 2.15, we determine the estimate |ψk(t+ 1)| ≤ ηk n∑ r=0 n r, t = 0, 1, ..., n− 1 and the convergence ηk → 0. thereby, |ψk(t)| → 0, t = 1, ..., n therefore, x [ak, w] (t) → x [a,w] (t) uniformly by w, t = 1, ..., n− 1. hence under continuity of function f follows that f [t, x [ak, w] (t)] → f [t, x [a,w] (t)] uniformly by w, t = 1, ..., n− 1. therefore, k (x [ak, w]) → k (x [a,w]) uniformly by w. advances in systems science and applications (2014) vol.14 no.2 121 then under finite dimensionality of the space of controls and boundedness of the set u we get sup w∈v |k(x [a,w])−k (x [ak, w])| → 0 (77) taking the limits in inequality (76) and taking into account the condition (77), we will have k (x [a, vk]) → k (x [a, v]) (78) the following estimate is true |k(x [ak, vk])−k (x [a, v])| ≤ |k(x [ak, vk])−k (x [a, vk])| + |k(x [a, vk])−k (x [a, v])| taking the limits here and taking into account the conditions (77), (78), we determine the convergence k (x [ak, vk]) → k (x [a, v]) hence under definition of control parameters v and vk follows that the maximal value of criterion k, corresponding to uncontrollable function ak, converges to its maximum, corresponding to extreme value of uncontrollable function. the theorem is proved. as the maximum of two (therefore, any finite number as well) of continuous functions is continuous, then from the theorem 2.16 follows similar result for the problem 2.5. theorem 2.17. assume that when conditions of the theorem 2.11 are met in neighborhood of point , function f is continuous by the third argument and satisfies the lipschitz condition by the first two arguments on x × u uniformly by the third argument, and function gj satisfies the lipschitz condition by the second argument on uniformly by the first argument.then optimal criterion value for the problem 2.5 continuously depends on uncontrollable function at point a. the purpose of the next study is to determine sufficient conditions of continuous dependence on uncontrollable functions of optimal criterion values for the problem 2.1 based on continuous non-autonomous dynamical system. theorem 2.18. assume that when conditions of the theorem 2.6 are met in neighborhoods of point , function f is continuous by the second argument and satisfies the lipschitz condition by the first and third arguments on x × a uniformly by the second argument.then optimal criterion value for the problem 2.1 continuously depends on uncontrollable function at point . proof. let the convergence occurs ak → a (c[0, t ])s (79) 122 a. ashimov: the theory of parametric control of macroeconomic systems and ... according the theorem 2.6 the problem 2.1at values of uncontrollable functions and ak has solutions, which we denote, correspondingly, by u and uk. denote by x[b, w] solution of state equations (1), (2), corresponding to uncontrollable function b and control w. ipso facto function x[b, w] meets relations ẋ [b, w] (t) = f (x[b, w] , w(t), b(t)) , t ∈ (0, t ) (80) x [b, w] (0) = x0. (81) then, by repeating reasoning from proof of the theorem 2.13, analogously with relation (70) we determine inequality 0 ≤ k (x [a, u])−k (x [a, uk]) ≤ 2 sup w∈u |k(x [a,w])−k (x [ak, w])| (82) after denoting x = x [a,w] , xk = x [ak, w] from the conditions (80) we derive inequalities ẋk(t)− ẋ(t) = f (xk(t), w(t), ak(t))− f (x(t), w(t), a(t)) , t ∈ (0, t ) by integrating by t and taking into account equalities (81), we get |xk(t)− x(t)| ≤ ∫ t 0 |f (xk(τ), w(τ), ak(τ))− f (x(τ), w(τ), (τ))| dτ ≤ l ∫ t 0 |xk(τ)− x(τ)| dτ+l ∫ t 0 |ak(τ)− (τ)| dτ ≤ l ∫ t 0 |xk(τ)− x(τ)| dτ+lt∥ak−∥θ, t ∈ (0, t ) where l is the lipschitz constant of function f , not depending on w. using the gronwall lemma, we will have |xk(t)− x(t)| ≤ ∥ak − a∥θ, t ∈ (0, t ) where positive constant depends only on l and on . using the condition (79), we get that xk(t) → x(t), t ∈ (0, t ) and therefore, x [ak, w] (t) → x [a,w] (t) uniformly by w, t ∈ (0, t ) hence under continuity of function f follows the convergence k (x [ak, w]) → k (x [a,w]) uniformly by w advances in systems science and applications (2014) vol.14 no.2 123 consequently, we get sup w∈u |k(x [a,w])−k (x [ak, w])| → 0 (83) taking the limits in inequality (82) and taking into account the conditions (83), we will have k (x [a, uk]) → k (x [a, u]) (84) the following estimate is true |k(x [ak, uk])−k (x [a, u])| ≤ |k(x [ak, uk])−k (x [a, uk])|+ |k(x [a, uk])−k (x [a, u])| taking the limits here and taking into account the conditions (83), (84), we will get the convergence k (x [ak, uk]) → k (x [a, u]) hence under definition of controls u and uk it follows that the maximal value of criterion k, corresponding to uncontrollable function ak, converges to its maximum, corresponding to extreme value of uncontrollable function. the theorem is proved. determine sufficient conditions of continuous dependence on uncontrollable functions of optimal criterion values for the variational calculus problem on choice (among given finite algorithms set) of optimal parametric control laws based on continuous non-autonomous dynamical system. theorem 2.19. assume that when conditions of the theorem 2.8 are met in neighborhood of point , function f satisfies the lipschitz condition on the set x×u ×a, and function satisfies the lipschitz condition by the second argument on uniformly by the first argument.then optimal criterion value for the problem 2.3* continuously depends on uncontrollable function at point . proof. let the convergence occurs ak → ab(c[0, t ])s (85) according the theorem 2.8 the problem 2.3* at values of uncontrollable functions and ak has a solutions, which we denote, correspondingly, by v and vk. denote by x[b, w] solution of state equations (23), (2), corresponding to uncontrollable function b and control w. ipso facto function x[b, w] meets relations ẋ [b, w] (t) = f ( x [b, w] (t), gj (w, x(t)) , b(t) ) , t ∈ (0, t ) (86) x [b, w] (0) = x0 (87) 124 a. ashimov: the theory of parametric control of macroeconomic systems and ... then, by repeating reasoning from proof of the theorem 2.15, analogously with relation (70) we determine inequality 0 ≤ k (x [a, u])−k (x [a, uk]) ≤ 2 sup w∈v |k(x [a,w])−k (x [ak, w])| (88) after denoting x = x [a,w] , xk = x [ak, w] from the conditions (86) we derive inequalities ẋk(t)−ẋ(t) = f ( xk(t), gj (w, xk(t)) , ak(t) ) −f ( x(t), gj (w, x(t)) , a(t) ) , t ∈ (0, t ) by integrating by t and taking into account equalities (87), we get |xk(t)− x(t)| ≤ ∫ t 0 ∣∣∣f (xk(τ), gj (w, xk(τ)) , ak(τ) ) − f ( x(τ), gj (w, x(τ)) , a(τ) )∣∣∣ dτ ≤ l ∫ t 0 [ |xk(τ)− x(τ)|+ ∣∣∣gj (w, xk(τ))−gj (w, x(τ)) ∣∣∣+ |ak(τ)− a(τ)| ] dτ ≤ l(1+) ∫ t 0 |xk(τ)− x(τ)| dτ+lt∥ak − a∥θ, t ∈ (0, t ), where l is the lipschitz constant of function f , and is the lipschitz constant of function by the second argument. using the gronwall lemma, we will have |xk(t)− x(t)| ≤ ∥ak − a∥θ, t ∈ (0, t ) where positive constant depends only on l and on . using the condition (85), we get that x [ak, w] (t) → x [a,w] (t) uniformly by w, t ∈ (0, t ) hence under continuity of function f follows that f [t, x [ak, w] (t)] → f [t, x [a,w] (t)] uniformly by w therefore, k (x [ak, w]) → k (x [a,w]) uniformly by w consequently, we get sup w∈v |k(x [a,w])−k (x [ak, w])| → 0 (89) taking the limits in inequality (88) and taking into account the condition (89), we will have k (x [a, uk]) → k (x [a, u]) (90) advances in systems science and applications (2014) vol.14 no.2 125 the following estimate is true |k(x [ak, uk])−k (x [a, u])| ≤ |k(x [ak, uk])−k (x [a, uk])|+ |k(x [a, uk])−k (x [a, u])| taking the limits here and taking into account the conditions (89), (90) we will get the convergence k (x [ak, uk]) → k (x [a, u]) hence under definition of controls u and uk follows that the maximal value of criterion k, corresponding to uncontrollable function ak, converges to its maximum, corresponding to extreme value . the theorem is proved. using the theorem 2.19 analogously with the theorem 2.17 (in this case system continuity is not a matter of principle), we come to the following statement. theorem 2.20. when for all j = 1, ..., r conditions of the theorem 2.19 are met, optimal criterion value for the problem 2.5 continuously depends on uncontrollable function at point . present statements of the theorems about continuous dependence on uncontrollable parameter of optimal criterion values for the parametric control problems of stochastic dynamical systems. proves of these theorems are similar to those presented above. theorem 2.21. let given any a ∈ a conditions of the theorem 2.13 are met. then optimal criterion value for the problem 2.6 are continuous function of the parameter . now, study the conditions of continuous dependence of optimal criterion value for the variational calculus problems on choice of parametric control laws on uncontrollable parameters. theorem 2.22. let given any a ∈ a conditions of the theorem 2.14 are met. then for chosen number value of the law j kj optimal criterion values for the problem 2.7 are continuous function of the parameter . consequence 2.23. when conditions of the theorem 2.22 are met for all j = 1, ..., r, optimal criterion value k = maxj=1,...,rkj for the problem 2.7 are continuous functions of the parameter a ∈ a. obtained results will be used in the next section in proving the existence of corresponding bifurcation points of extremals of the variational problems. 2.3 sufficient conditions for the existence of extremals’ bifurcation points of the problems on choice of optimal parametric control laws let us introduce a notion of extremals’ bifurcation point of the variational calculus problems on choice (among given finite algorithms set) of optimal parametric control laws.existence of such a bifurcation point for some uncontrollable function a(·) means that in its neighborhood for the problem in question occurs transfer from one optimal parametric control law to another. 126 a. ashimov: the theory of parametric control of macroeconomic systems and ... consider abstract variational calculus problem on choice (among given finite algorithms set) of optimal parametric control laws, generalizing the problems 2.3, 2.5, and 2.7. given: the set of uncontrollable functions (or parameters), set of the acceptable controls sets (adjustable coefficients values) v j a , a ∈ a, j = 1, ..., r and set of functionals (optimality criteria) kj = kj(a, v) where v ∈ v j a , a ∈ a, j = 1, ..., r. give definition for extremals’ bifurcation point for set of the maximization problems for given functionals on corresponding sets of acceptable controls, i.e. abstract variational calculus problem on choice (among given finite algorithms set) of optimal parametric control laws. definition. the value a ∈ a call asextremals’ bifurcation point for the maximization problem for mappings u→ kk(a, v) on the sets v k a , k = 1, ..., r if there are exist two different numbers i, j ∈ 1, ..., r such that the following relation is true max v∈v l a ki(a, v) = max v∈v j a ki(a, v) = max k=1,...,r max v∈v k a ki(a, v) and in any neighborhood of point there is such a point b ∈ a for which the value max k=1,...,r max v∈v j b kj(b, v) reachs for the only value of k. theorem 2.24. assume that when conditions of the theorem 2.17(or 2.20, or 2.23) are met, is a connected set, and there are two different values a0, a1 ∈ a, such that the maximal values by j = 1, ..., r of function maxima v → kj(a, v) on the sets are reached for the different values j0, j1 that is the following inequalities are true: max j=1,...,r,j ̸=j0 max v∈v j a0 kj(a0, v) < max v∈v j0 a0 kj0(a0, v) max j=1,...,r,j ̸=j1 max v∈v j a1 kj(a1, v) < max v∈v j1 a1 kj1(a1, v) then there exists a bifurcation point of extremals of the problem 2.3 (or 2.5, or 2.7) on choice (among given finite algorithms set) of optimal parametric control law. proof. under connectivity of the set points a0, a1 can be connected by continuous line a = a(s), s ∈ [0, 1], lying in the set , and the following equalities are true a(0) = a0, a(1) = a1 denote kj(s) = max v∈v j a(s) kj1(a(s), v), s ∈ [0, 1] advances in systems science and applications (2014) vol.14 no.2 127 from the theorems 2.17 (or 2.20, or2.23) follows continuity of functions s → kj(s) in the segment [0, 1], and therefore, continuity in this segment of function as well, where determine the set ∆(j) = {s ∈ [0, 1]|kj(s) = k∗(s)} , j = 1, ..., r it is closed, being complete preimageof closed set, consisting of the only point (zero) for continuous function y = y(s) where y(s) = kj(s) − k ∗ (s) thereby, we present the segment [0, 1] in terms of the following sum [0, 1] = ∪ j=1,...,r ∆(j) consisting, according to the theorem conditions, as minimum, of two non-empty closed sets. from theorem conditions it follows the following relations as well: 0 ∈ ∆(j0), 1 /∈ ∆(j0) then the set of boundary points of the set ∆(j0) which are situated in the interval (0, 1), is not empty. consequently, there exists lower boundary s0 of such boundary points. this value is a boundary point of some another set ∆(j2) and is a part of it as well. thereby, at a = a(s0) the maximum by j = 1, ..., r the values is reached, as minimum, for two different numbers j0 and j2. at the same time, at 0 ≤ s ≤ s0 this maximum is reached for the only value j0. thus, a(s0) actually corresponds to desired bifurcation point. the theorem is proved. the following statement is direct consequence of the theorem 2.24. consequence 2.25. assume that when conditions of the theorem 2.17 (or 2.20, or 2.23) are met, is connected set, and at value control using the law provides solution of the problem 2.3 (or 2.5, or 2.7), and at , ( control using this law does not provide solution of the problem in question, that is the following inequalities are true max j=1,...,r,j ̸=j0 max v∈v j a0 kj(a0, v) < max v∈v j0 a0 kj0(a0, v) max j=1,...,r,j ̸=j1 max v∈v j a1 kj(a1, v) > max v∈v j0 a1 kj0(a1, v) then there is at least one bifurcation point of extremals of mentioned problem 2.3 (or 2.5, or 2.7). at the end we present description of the numerical algorithm for finding bifurcational value of function (or parameter) a of one of the problems 2.3 or 2.5 or 128 a. ashimov: the theory of parametric control of macroeconomic systems and ... 2.7 on choice (among given finite algorithms set) of parametric control laws and when conditions of the theorem 2.24 are met. connect the points and by smooth curve. divide this curve to equal parts with sufficiently small step. for obtained values (points) are determined the numbers of parametric control lawsbringing solution of the problem 2.3 or 2.5 or 2.7 given values. then we find the first value i, at which corresponding law number differs from. in this case bifurcational value lies on arc of the curve. for found part of the curve, the algorithm for determining bifurcation point with given accuracy consists in using the method of bisections. consequently we find the point, on the one hand from which on this arc within the range of deviation from the value the optimal law is, and on the other hand-within the range of deviation from the value this law is not optimal. from the consequence 2.25 follows that extremal’s bifurcation point for the problem under solution lies on mentioned arc, and as its estimate can be taken any point of the are. references [1] ekeland i., temam r. (1976), “convex analysis and variational problems”, amsterdam: nord-holland publishing company. corresponding author a. ashimov can be contacted at: ashimov37@mail.ru adv syst sci appl 2018; 1; 85-91 published online at http://ijassa.ipu.ru. copyright ©0000 assa. adv. in systems science and appl. (0000) management of the development of the region: the attractiveness of the region and human potential vadim kleparskiy 1 , victor sheinis 2 1) v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: kleparvg@ipu.ru 2) v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: scheynis@yandex.ru abstract: using a comparative analysis of the data of the "passive" experiment, it was shown that an increase in the attractiveness of the region can be achieved by long-term, purposeful increase in human potential and by attracting corresponding additional investments keywords: predictive control, “passive” experiment, attractiveness, human potential, investments. 1. introduction the top-level control of the region – a large-scale socio-economic system (ses) – is the complicated process of realizing adequate administrative measures aimed at maintaining commensurate flows of material, labor, information and other resources in a regime that ensures a sustainable increase in the level of well-being. such a goal can be achieved by a steady increase in the total productive potential (tpp) the main driving force for the evolution of the system. from the positions of the top level of the regional control hierarchy, such a multidimensional and multilevel process is reduced to the activity of forming the desired channel of "attracting" solutions (the attraction channel in terms of nonlinear dynamics). at the same time, the top level of regional governance should be taken into account that the total productive potential is determined not only by material assets (physical capital, investments), but also by intangible (nonmaterial) assets. and the main component of intangible assets is human potential. it should also be noted that for developed economies of the world, the share of human potential in the total volume of tpp is 50-60% (see, for example [1]). increasing the attractiveness of the region in physical (material) capital (for investment), as well as the desire to increase the human potential of the system is therefore the main objective of long-term predictive control of the region. in the competitive environment for resources, this task is quite difficult to fulfill. success is largely determined by the long-term control activity to create favorable conditions for ensuring the socialeconomic merits of the governed region using the cultural and historical heritage. in the proposed letter, using the modified poincare transversal surface method, using the "passive" experiment, the possibilities of improving the population welfare of the region by attracting qualified human resources and additional investments are considered. 2. investigation concepts given the relatively slow process of establishing a dynamic equilibrium in the flow of a dissipative administrative environment, the rate of change in the chosen "order parameter" c 86 v. kleparskiy, v. sheinis copyright ©2018 assa. adv. in systems science and appl. (2018) (the level of social-economic well-being) determined largely by the specific value of the regional gross product (of gdpreg per person) сan be determined with sufficient accuracy in the gradient approximation by the expression )( )( )( ),( t c pp t c rcw dt dc phnonmat         (1) here, w(c,r) is the total productive potential (tpp) of the system, r(t) is the vector of control actions, pph the physical potential of the ses (based assets, investments), pnonmat – is nonmaterial assets (human potential and its components). the second term – ξ(t) on the right side of the equation expresses the possibility of the influence of random perturbations on the trajectory of the ses development. dynamic equation (1) allows to study the process of long-term planned approximation of the region to the realities of the outside world from the results of observations of the course of socio-economic development using the "passive" experiment using the basic provisions of the qualitative theory of nonlinear dynamics and physics of society analyzing (1) it can be seen that with the increase in the share of nonmaterial assets in the tpp system the dynamics of socio-economic development of the region will increasingly be determined by a skilful increase in the human potential. and, simultaneously, increasing the effectiveness of its use at all levels of the hierarchy of management of the system. the strategy of long-term predictive control therefore requires an adequate assessment of the main components of the human potential of the system and the identification of possible ways to increase them. a qualitative assessment of the value of the human potential components of ses can be performed using the human development index (hdi) – until 2013 as an index of human development. the determination of the hdi value is performed taking into account three groups of basic indicators.. these groups of indicators are estimated: life expectancy – longevity (determined mainly by the level of development of medicine); the level of literacy of the population (determined by the availability of educational institutions); the standard of living, estimated through the value of gdpreg per capita at purchasing power parity.. it should be expected that it is the first two components of the hdi (the state of medicine and, accordingly, the level of literacy) that will determine the level of the third component of the hdi gdpreg/capita. creation of conditions ensuring the increase of human.potential in the predicted control can be ensured by the steady increase of its social and cultural components increasing in this way the attractiveness of the region for a growing population (and, accordingly, the saturation of the region by initiative, efficient skilled personnel) therefore seems to be the main task of the top level of the governance hierarchy in planning long-term development of the region. as the object of the "passive" experiment, the federal lands of germany were chosen. earlier it was noted that for germany there are two groups of federal lands: "new" (former gdrs) and "old", more developed in industrial relation (see, for example, [2]). the newest statistical data confirm the sharp heterogeneity of the level of regional well-being achieved by various regions (federal lands and individual cities in germany), estimated by the level of gdpreg per person employed. (see fig.1). 3. results and discussion analyzing the spstial distributions (gdpreg/person empl. data) presented in fig. 1, it can be seen that the largest values of gdpreg/person employed. were achieved in the "old" federal lands of germany: bavaria, baden-württemberg, north rhine-westphalia, hessen, hamburg). this level of the state of the objects of research makes it possible to compare the effectiveness of long-term planned control of the development of the region more fully based on the representative results of the "passive" experiment management of the development of the region 87 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.1. spatial distribution of regional gdpreg/person employed. the brightest areas are the gdpreg/person empl.< 52.7 10 3 euro. the darkest areas are gdpreg/person empl. > 65.2 10 3 euro. fig. 1 obtained on the base of [3]. 3. results and discussion analyzing the spstial distributions (gdpreg/person empl. data) presented in fig. 1, it can be seen that the largest values of gdpreg/person employed. were achieved in the "old" federal lands of germany: bavaria, baden-württemberg, north rhine-westphalia, hessen, hamburg). this level of the state of the objects of research makes it possible to compare the effectiveness of long-term planned control of the development of the region more fully based on the representative results of the "passive" experiment as it follows from the definition of the human development index (hdi), the influence of nonmaterial assets (human potential and its components) on the dynamics of the region's long-term development is largely realized through skilled and highly motivated work at all levels of regional control activity (including also and the institutional stage of governance). it is the achieved level of social and cultural development of the region that determines the possibility of increasing the share of scientific and productive activities in the process of ses-control operation. the process of this increase is determined by the requirements of the modern (global) economy of high technology production. at the same time, regions with previously developed industrial infrastructure have the advantage of reinforcing the transition to a new "postindustrial" period – to the knowledge economy – with the corresponding additional investments in physical capital and human potential. such a 88 v. kleparskiy, v. sheinis copyright ©2018 assa. adv. in systems science and appl. (2018) complex relationship between necessity and obligation (the need for a transition to a knowledge economy, the necessety for additional investment and the need to increase the effectiveness of human potential use) largely determines the planned increase in well-being – and the value of the specific regional gross product gdpreg/capita. at the same time, the historical past of the region is of great importance. for example, the agglomeration of large cities (duisburg, dusseldorf, essen, cologne, bonn) in north rhine-westphalia (based on the ruhr industrial region) is characterized by a close relationship between the gdpreg per worker (parameter, which characterizes the labor.productivity achieved) – fig.1 – and the share of highly skilled workers engaged in high technology production – fig. 2. fig.2. proportion (in %) of workers employed in the high-technology industry. the brightest areas the share of workers employed in science-intensive production is less than 1%.the darkest areas – this share is more than 10%.fig. 2 obtained on the base of [3]. comparison of the spatial distributions of gdpreg per worker and the share of highly skilled workers (figure 1 and figure 2) makes it possible to visually compare and estimate the effectiveness of predicted control implemented in the course of evolution using the planned increase in the components of human potential for improving the welfare of the population of the region. comparative analysis presented in fig. 1 and fig. 2 spatial distributions of labor productivity (gdpreg per worker employed in production) and the share of workers. employed in science-intensive production allows you to notice numerous local deviations from the mean (by federal land) values. these local positive outbursts of gdp are characterized not only by certain regions, but also by individual cities, attractive both for investment in physical capital and for increasing human potential. this ability of the region to social and economic development is largely determined by the long-term activity of the top level of the control hierarchy of the region (city) in creating and maintaining an attractive image of a controlled ses. the hastily developing metropolitan cities (munich, stuttgart, düsseldorf, hamburg, bremen) and urban agglomerations (for example, the urban management of the development of the region 89 copyright ©2018 assa. adv. in systems science and appl. (2018) agglomerations of north rhine-westphalia and bavaria) are a good example of the implementation of such positive predictive control. one of the tools for creating an attractive image of the region is the predictive control activity associated with employment in scientific and technical research and innovation (stria) this control activity is impossible without the corresponding costs for highly qualified scientific and engineering specialists (stria-expenses). in their origin, these stria-expenses are due to the costs of training and saturating the region's economy by graduates of the higher school of science and technology. the values of stria-expenses largely depend on the region's economics possibilities – on region’s values of gdpreg/capita. there are peculiar feedbacks: the dependence of gdpreg/capita on the amount of investment (the specific value of investment) and the amount of expenditure on stria, and, at the same time, the dependence of striareg/person on gdpreg/capita. this dependence for two groups of german federal lands was detected using a modified poincare method of a transversal cutting surface (see, for example, [2]).. the striareg/person dependences on the gdpreg/capita for the "old" and the "new" federal lands of germany are shown in fig. 3. . it can be seen that the highly developed bavaria could have specific costs for stria/person  900 euro – within  2.4% of gdpreg/ person. such opportunities bavaria could have due to the implementation of a multi-year development program to create the conditions for a more effective exchange between scientific achievements and the economy (see, for example, [4])..at the same time, most of the funds for development programs were invested in education and training, research and development, in the development of information and communication technologies. as a result of the chosen strategy, bayern was able to reach the front lines in the aviation and space industries, biotechnology and medical technology, and other high-tech industries. this provided bavaria with an average gdpreg/capita of about 40,000 euros and, accordingly, a sufficiently high attractiveness. 0 0,2 0,4 0,6 0,8 1 1,2 1,4 1,6 20 25 30 35 40 45 gdp reg/person, 10^3 euro s t r ia r e g /p e r. , 1 0 ^ 3 e u ro 5 4 7 3 2 5 4 6 1 3 1 2 fig. 3. influence of the value of the regional gross domestic product gdpreg/person on the specific value of expenditure for striareg/person. the upper line is the "old" lands of germany. . (1bavaria, 2badenwürttemberg, 3hessen, 4lower saxony, 5north rhine-westphalia, 6rhineland-pfalz, 7-schleswigholstein). lower (left) line new "lands" of germanylower (left) line new "lands" of germany. (1thuringia, 2saxony-anhalt, 3saxony, 4mecklenburg, 5brandenburg). for a relatively backward thuringia (as an example), it was possible to provide striareg/person only in the amount of 220 euros – within 1% of the achied level of the 90 v. kleparskiy, v. sheinis copyright ©2018 assa. adv. in systems science and appl. (2018) gdpreg/person. insufficiently large gdpreg/person. values in germany's “new” lands cause relatively small financial allocations for scientific and technological development. and innovations (striareg/person expenditures). under such circumstances, the region's achievement of attractiveness for additional investments in physical capital and human potential becomes difficult. accordingly, the subsequent achievement of a high standard of living for the “new” german lands (for lands as a whole) is also difficult. as a result, the achievement of a high standard of living in the “new” federal lands of germany was mainly realized only in the regional metropolises and individual centers of the previously achieved scientific and technological development. illustrative examples of such metropolitan cities are potsdam (brandenburg land), leipzig and dresden (saxony). in these cities, the leading scientific centers, knowledge-intensive industries are concentrated. these large cultural and historical centers have an anomalous attracting effect. these conglomerate cities have.the advantage of competing for skilled labor, which gives an additional impetus to their development. as a result, these cities were included in the top ten cities of germany with the largest population growth (potsdam – 6.3%, leipzig – 7%, dresden – 5.8%) for the federal land of thuringia (the average gdpreg per capita for the country = 23,000 euros), cities – islands of local development – are (as can be seen in figure 2) jena, erfurt, weimar. the city of jena – has a university and is the center of electronic-optical engineering – and has a regional gdpreg/person employed = 58,2 10 3 euro. the city of erfurt – the capital of the federal land and the university sity – has gdpreg/person employed = 54,3 10 3 euro. at the same time in erfurt there is a well-established system of social insurance of employees, which contributes to the inflow of labor.the city of weimar known for its university and its adjoining research institutions (precision instrumentation making)– has gdpreg/person employed = 54,3 10 3 euro. the listed cities have gdpreg per capita roughly twice the average for the federal state of thuringia [5]. but in the nearby city of ingolstadt (bavaria) – a flourishing center of the automobile industry – the value of the gdpreg per capita is 123 10 3 euro. such cities in bavaria as erlangen and regensburg have the value of gdpreg per capita = 83 10 3 euro. naturally, the "old" federal lands of germany continuously "suck" the qualified specialists from the "new" lands. especially noticeable is this effect in small provincial towns and rural areas [3, 6].. 3. results and discussion the study of the level of social and economic development of the "old" and "new" german lands (using a comparative analysis of the data of the "passive" experiment – figures 1, 2 and figure 3) shows that a balanced growth of regional welfare requires from the regional governance structures a long-term program for switching the economy to knowledge-based industries, aimed at creating a new value. the implementation of such a program requires, in turn, the expansion and modernization of the relevant educational institutions and the creation of an advanced training system at all levels of the regional management hierarchy. in the conditions of the existence of the federal land, as an open ses, the fulfillment of all these requirements for the “new” lands of germany is quite a difficult task.to prevent the outflow of human and financial resources into the "old" federal lands, it is therefore necessary to develop and implement appropriate preferences and benefits. references [1] meljanzev, v. o a. (2005) problemy i faktory formirovanija sovremennogo (intensivnogo) ekonomnitsheskogo rosta v stranah zapada, vostoka i v rossii [problems and factors of formation of modern (intensive) economic growth in the management of the development of the region 91 copyright ©2018 assa. adv. in systems science and appl. (2018) countries of the west, the east and in russia]. moscow, ussr. history and synergetics. methodology of the study m.: urss, 2005.[in russian] [2] kleparskiy, v. o g. & scheinis v. o e. (2016) upravleniye dolgosrotchnym razvitijem regiona: investicii, tchelovetcheskiy kapital, institucionalnye struktury. management of long-term development of the region: investments, human potential, institutional structures. proc. of the 6th int. conf. mlsd’2016. moscow, 2: 67-76. [in russian]. [3] albrech, j., fink, ph. & tiemann, h. (2015) ungleiches deutschland: sozioökonomischer disparitätenbericht 2015. [the disparate germany: the researchreport for social and economy disparate 2015]. 1-56. friedrich-ebert-stiftung [in german] www.fes-2017plus.de [4] grinberg, r. o s., knogler, m. & zedilin, l.i. (2008) ekonomitcheskaya i sozialnaya politika v bavarii: uroki dlya rossiyskih regionov [economic and industrial policy in bavaria: lessons for russian regions]. moscow, russia. econ-inform, [in russian] [5] http://www.statistik.thueringen.de/datenbank/tabauswahl [6] schwaldt, n. (2015, august 14) pulsirende metropolen, veroedete doerfer. [the trob capitals, the devastated country-sides]. die welt, [in german] http://www.fes-2017plus.de/ http://www.statistik.thueringen.de/datenbank/tabauswahl advances in systems science and applications (2012) vol.12 no.2 153-161 effect of laser shock processing residual stress on pellet traveling grate surface crack x.d.ren, y.k.zhang, d.w.jiang, a.x.feng and t.zhang jiangsu key laboratory of laser manufacture science and technology ministry, jiangsu university,zhenjiang,jiangsu 212013 abstract the effect of the laser shock processing on pellet traveling grate surface fatigue crack growth performance were investigated from the theory of the fracture mechanics. a stress formula was set up to indicate the function of the residual stress on fatigue life. the effect of the compressive stresses was deemed responsible for increasing the resistance to fatigue crack growth of the pellet traveling grate. it was observed that the dynamic stress intensity factor of the crack and the stress ratio were declined due to laser shock processing, which would make the spread doorsill of the fatigue crack increasing. the results indicate a significant reduction in fatigue crack growth rates using laser shock processning specimens. it is shown that the near-surface microstructures, which in pellet traveling grate consist of a layer of work hardened nanoscale grains, play a critical role in the enhancement of fatigue life by mechanical surface treatment. keywords laser shock, residual stress, pellet traveling grate, crack expand 1 introduction in recent years, the pellet has obtained favor and been taken seriously as the high quality raw material, and has the tendency of replacing the agglomerate gradually. the chain fire grate machine is a core equipment in the production of pelletizing mineral aggregate, and as a non-sign design large-scale equipment, which core part movement chain system operates in an alternation temperature environment from normal temperature to 1050◦c for a long time. the main bottleneck question is that the alternation high temperature environment forms the complex heat expansion and the alternation thermal load, which causes the heat-resisting service life to be shorter and the pelletizing chain fire grate machine requires extremely high performances of heat shock fatigue, strength at high temperatures and wear-resisting ability of spare part material[1]. therefore, enhancing the fatigue resistance ability of the crucial element and guaranteeing the reliability of the movement are the difficult problems of the entire equipment technology. laser shock processing is to use the high power density (gw/cm2 magnitude), short pulse (ns magnitude) to impact the energy conversion body in the metal surface. after laser energy absorption, the temperature is elevated rapidly, forming the shock-wave with a high peak-to-peak value, which means that the 154 x.d.ren:effect of laser shock processing residual stress on pellet traveling... energy of light is transformed into the shock-wave mechanical energy. the shockwave pressure which reaches as highly as counts gpa causes that the microscopic plastic deformation happens in the material surface layer, and forms the residual compressive stress layer, lsp improves the mechanical property of the metallic material effectively, which would enhance the fatigue life and the anti-stress corrosion performance of the material in a large scale specially. compared with the conventional method, this kind of high rate of strain strengthening technology has a unique superiority [2-3]. regarding to the stress destruction in the surface of essential spare part axis in chain fire grate machine, this article uses the technology of laser shock processing to enhance the fatigue life and the wear resistance of the axis, and the results indicate that using advanced surface treatment technology of the shock-wave mechanics effect induced by the strong laser can effectively improve the surface layer stress condition of the axis material, and specially can enhance the fatigue life performance of the material obviously. 2 residual stress analysis in the laser shock processing, the coating raises the laser energy coupling efficiency, and the coating is able to enhance the absorption of the laser energy, which causes the absorption coating gasification to form the plasma to explode, producing the high-pressured shock-wave separately disseminates to the metal target and the restraint level. the high-pressured shock-wave plays the strengthening and the distortion roles to the metal target, and the superficial residual compressive stress is enhanced greatly. researching the fatigue short fracture growth problem in laser shock processing residual compressive stress field has also solved the problems of the crack rate of expansion caused by residual stress and life prediction[4-6]. the model of metal surface’s influence layer depth and superficial residual compressive stress after laser shock processing proposed by ballard[7] is based on in this article, and the shock wave parameters of peak pressure is optimized in condition of ideal surface stress. a residual stress model which is used to calculate the longitudinal and plane shock-wave system’s ideal elastic-plasticity response for metallic material is proposed, and some basic suppositions are made as fellow: (1) distortion induced by the laser shock processing is single axle and plane; (2) distribution of pressure pulse induced by laser is even in space; (3) the material obeys von the mises yield criterion, and the elastic deformation and the strain strengthening effect of the material are neglected. the residual stress’s production has two stages, as shown in fig.1. (a) during the time of laser pulse, a pure uniaxial stress is generated along the shock-wave propagation direction because of rapid spraying by plasma, while a tensile stress is generated in the plane parallel to the material surface ; (b) after laser pulse advances in systems science and applications (2012) vol.12 no.2 155 vanishing, the plastic strain happens to the laser impact zone’s volume, and two axle compressive stress field is produced in the plane parallel to the impact surface due to that the plastic strain is limited by the metallic material around. fig.1 schematic of the residual stress induced by the lsp according to the hooke’s law, the metal surface plasticity strain εp can be written as[8], εp = −2hel 3λ+ 2µ ( p hel − 1 ) (1) wherehel is the hugoniot limit of elasticity, p is the laser shock-wave pressure, and λ and µ are the lame’s parameters. to determine the residual stress field in the materials impact area, regarding any impact condition assigned, the plastic influence depth was calculated as, lp = celcplτ cel − cpl (p −hel 2hel ) (2) cel and cpl represent the elasticity speed and the plastic speed separately, and τ is the duration of laser pulse. and ρ represents the target material’s density, while cel and cpl can be defined as, cel = √ λ+ 2µ ρ ,cpl = √ λ+ 2µ/3 ρ (3) therefore, according to the influence level depth, the compressive residual stress produced in the elastic-plasticity semi-infinite body can be shown as, σsurf = σ0 − [ µεp 1 + ν 1− ν + σ0 ][ 1− 4 √ 2 π (1 + ν) lp r0 √ 2 ] (4) where ρ0 represents the laser facular’s radius, and σ0 represents the initial superficial residual stress which may be zero in the ordinary circumstances. the formula (4) indicates that the metal surface residual stress’s peak-to-peak value increases with the laser peak pressure in the laser impact process, while the peak pressure is related to the incident laser power density. 156 x.d.ren:effect of laser shock processing residual stress on pellet traveling... 3 laser shock experiment the laser treatment was carried out at the jiangsu university strong laser laboratorys high power and nd:glass laser implement, with its laser pulse wave is 1.06µm,and syntory laser stick is made by ϕ6mm × 80mm yttrium aluminum garnet. one stair nd:glass laser magnify beforehand and four steps main amplifier by two ways, and big caliber quarter wave piece and high strength big caliber polarize film was adopt, q switch, the output laser film wave at half maximum is about 20ns, the pulse repletion rate is 0.5hz and the pulse energy was 28j to 36j corresponding to laser energy densities at the metals surface of 56j/cm2 and 72 j/cm2. the experiment material is 12crmov, whose thickness is 5mm. before the experiment, the test sample surface was polished with the ethyl alcohol. the axis material was made into the standard test specimen. fig.2 residual stress of the staff in different conditions the depths of the impact zone indicate that after laser shock processing, its superficial residual stress is transformed from tensile stress to compressive stress, as shown in fig.2. the surface residual stress is qualitatively changed from the 135.7mpa tensile stress to 230.6mpa compressive stresses, which scope is reached as highly as 3 times. and the hardness is obviously enhanced, which has an important value to enhance fatigue life of the work piece. 4 results and discussions 4.1 material performance influence after laser shock processing, the residual stress of the 12crmov is qualitatively changed from the 135.7mpa tensile stress to 230.6mpa compressive stress, correspondingly surface hardness is enhanced from hv252 to hv530. fig.3 shows the result of metallography observed, and the big arrow shows the crack that the naked eye sees, which assumes the bending shape and branches out many small cracks; the small arrow shows the small crack which branches out from two sides advances in systems science and applications (2012) vol.12 no.2 157 of the big crack. by magnifying part of the crack that naked eye sees in fig.3, it can be seen that the small cracks along with both sides of the crack are along crystal cracks. obviously these along crystal small cracks indicate that the crack that can be seen by naked eye is also to expand along the crystal. it is only because the latter is a little wider, and its characteristic of along crystal crack is not clearer than the small cracks.fig.4 shows the microstructure of the axis test specimen by chemical reagent corrosion. it is the axis high-temperature steel normal tissue. and fig.4 demonstrates that crack is to expand along the crystal. fig.3 metallographic grinding in the surface morphology of the cracks fig.4 microstructure of the small axis 4.2 influence to crack growth after laser shock processing, residual compressive stress remains in the surface, and the residual stress will weaken in the process of test specimen surface withstanding cyclic loading, and the residual stress on the test sample section is a 158 x.d.ren:effect of laser shock processing residual stress on pellet traveling... distribution rather than a definite value. in order to analyze that simply, the effect of the residual stress was estimated with mean stress’s viewpoint, and the influence of mean stress to fatigue limit is often described with the goodman relations [9], which is shown in fig.5. in fig.5, σm is the mean stress, and σb is the corresponding tensile strength, while σ0 p is fatigue limit whenσm = 0 . as average pressure exists, the fatigue limit may be represented as. σm p = σ0 p − (σ0 p/σb)/σm = σ0 p −mσm (5) fig.5 goodman sketch map of the relationship between stress and metal fatigue limit in the formula(4.1),m = σ0 p/σb denotes the slope of the line between σ0 p and σb in fig.5, which is called the mean stress sensitive coefficient. when the residual stress σr exists and also it is thought to be equivalent with mean stress, the formula (5) can be rewritten as. σr+m p = σ0 p −m(σm + σr) (6) compare formula (5) with formula (6), it is known that the change of materials fatigue limit induced by the residual stress is. ∆σr p = σr+m p − σm p = −mσr (7) it could be seen that the materials fatigue limit would decrease if the residual tensile stress remains in test specimens surface, while the materials fatigue limit would increase when the residual compressive stress remains in test specimens surface, and m may also be called as the residual stress function coefficient. therefore, the residual stresss influence may be estimated quantificationally if advances in systems science and applications (2012) vol.12 no.2 159 the materials residual stress function coefficien m and the test samples residual stress value are known. the problem of crack growth in residual compressive stress field, besides the influence of residual stress to fatigue limit, and also influence of the residual stress to the crack growth threshold should be considered. according to the revision relations proposed by ei haddad [10] and so on, the fatigue limit of the short crack in the metal surface can be calculate as ∆σ = ∆kth y[π(r+ a0)]0.5 (8) where ∆σ is the fatigue limited stress, and ∆kth is the cracks threshold stress intensity factor, and y is the scoop channel’s shape factor, and a0 is the materials critical crack length. ∆kth = ∆σy[π(r+ a0)] 0.5 (9) after laser shock processing, the existence of the residual compressive stress in the metallic material surface greatly enhances the fatigue limit in the formula (9), and its result will cause the threshold value to obtain the enhancement, and in this condition the crack is not easily generated; the cracks that have extended will stop extending in the high residual compressive stress region and become the non-extension cracks, and a more universal influence is that the fatigue crack growth rate is reduced obviously in the residual compressive stress field. this is because that the residual compressive stress reduces the mean stress in alternating load, and reduces the alternating tensile stress that the test sample surface actual withstands. meanwhile, laser shock processing enhances the plastic deformation resistance of the sample surface hardened layer, which enhances the fatigue strength of the material. when ratio of materials fatigue strength (σr) and value of test sample surface withstanding stress (σω)σr/σω > 1, he probability of the formation of fatigue crack is greatly reduced [11]. the influence of the residual stress to the fatigue crack growth rate can be described with forman formula, da dn = c(∆k)m/[(1−r)k −∆k)] (10) in the formula (10), ∆k is the cracks threshold stress intensity factor, and r is the stress ratio, while c and m are constants. when residual compressive stress exists after laser shock processing, it can be known that the stress ratio r is smaller than the stress ratio rw when residual compressive stress doesnt exist. obviously the residual compressive stress not only could reduce the expansion speed of the fatigue cracking, but also enhance the fatigue cracking expansion resisting force of the test sample. 160 x.d.ren:effect of laser shock processing residual stress on pellet traveling... 5 conclusions the compressive residual stress is formed in the impact metal surface during laser shock processing,. the residual compressive stress is equivalent to negative average residual stress, and it can enhance the anti-fatigue strength of the workpiece. the compressive residual stress increases the locking force of the crack and reduces the expansion speed of the fatigue cracking obviously, and also can close the metal plate crack effectively.after a large number of experiments, the empirical datum is summarized and theories are consummated unceasingly, and the high rate of strain strengthening theory and the laser shock processing database are established. acknowledgements support provided by the state key program of national natural science of china (grant no. 50735001) and project supported by the natural science foundation of the jiangsu higher education institutions of china (grant no. 07kjb460012 ) and the ph.d. programs foundation of jiangsu university(cx07b-02x). references [1] zhang xiliang, zhang jian, feng ai-xin, zhu shen-cai, wu ti-chang, wan xue-gong. (2007), “reliability design on key parts of chain grate”, sintering and pelletizing, vol.32, no.4, pp.5-9. [2] ren xudong, zhang yongkang, zhou jianzhong, feng aixin, kong dejun. (2006), “study of the effect of coatings on mechanical properties of tc4 titanium alloy during laser shock processing”, materials science forum, no.532-533, pp.73-76. [3] ren xu-dong, zhang yong-kang, zhou jian-zhong, feng ai-xin. (2006), “effect of laser shock processing on residual stress and fatigue behavior of 6061t651 aluminum alloy”, transactions of nonferrous metals society of china, vol.16, pp.1305-1308. [4] li fukai. (1998), “fatigue behavior of the metal materials in remanent stress field”, xi’an university of science & technology journal, vol.18, pp.154157. [5] cui zhenqi, xu kewei, hu naisai. (1994), “model and experiments on fatigue short crack growth in residual stress field”, acta metallrugica sinica, vol.30, no.2, pp.a65-a69. [6] li hangyue, hu naisai, he jiawen, zhhou huijiu. (1998), “an analytical model of compressive residual stress effect on closure”, acta metallrugica sinica, vol.34, no.8, pp.847-851. advances in systems science and applications (2012) vol.12 no.2 161 [7] dubrujeaud b, vardavoulias m, jeandia m. (1974), “dry sliding wear behavior of a p/m ferrous alloy superficially densified by laser shock proeessing”, surface and coatings technology, vol.67, pp.125-132. [8] g.banas et al. (1990), “laser shock induced mechanical and microstructural modification of welded marging steel”, j. appl. phys, vol.67, no.5, pp.23892401. [9] zhang ding quan. (2002), “the effects of residual stresses on the fatigue strength of metal”, physical testing and chemical analysis parta physical testing, vol.38 no.6, pp.231-235. [10] ei haddad m h, topper k j, pook l p. (1974), metal fatigue, london: oxford univ. press, pp.130-195. [11] wang shuqin, wang jingyi. (1997), “effect of laser hardening treatmeng on fatigue stength of 34crnimo steel”, ordnance material science and engineering, vol.20 no.4, pp.29-34. corresponding author corresponding author: renxd@ujs.edu.cn advances in systems science and application (2015) vol.15 no.4 338-350 simulation modelling for computer aided design of secondary aerodynamic wing surfaces aleksandr a. gorbunov, aleksej d. pripadchev, irina s. bykova and valerij v. elagin federal state educational government-financed institution of higher professional education, orenburg state university, orenburg region, russia abstract the simulation modelling technique of secondary aerodynamic wing surface of the long-haul aircraft using a high-precision mathematical simulation in the aerodynamics and hydrodynamics computer programme has been formulated in the present article. the method proposed is based on the modern methods in the field of computer-aided design of aircraft using catia three-dimensional simulation and simulation modelling in salome environment. the use of highprecision mathematical simulation in computer-aided design of secondary aerodynamic wing surfaces allows us to determine the aerodynamic characteristics of the secondary aerodynamic surface of the model developed and to verify the previous results of the project definition. keywords secondary aerodynamic surfaces; high-precision computer mathematical simulation; simulation model; a long-haul aircraft; synthesis; design automation; project procedures; three-dimensional modelling 1 introduction building the new long-haul aircraft (a/c) with improved performance characteristics and enhancing the currently used aircraft is effected in various ways. one of them is improving its aerodynamics depending largely on the a/c look. in our view, the most rational way to improve the aerodynamic characteristics of the a/c is to install the secondary aerodynamic surfaces (sas) at the wingtip. the use of sas reduces the induced drag of the aircraft, enhances the effective wing aspect-ratio and ascensional power at the wingtip, improves a/c longitudinal/transverse stability, reduces specific fuel consumption, cuts takeoff run and landing roll of the aircraft. currently, there are many wing sas designs installed at the long-haul a/c differing in geometrical and aerodynamic characteristics [1-3]. however, despite years of experience based on the empirical approach, there is no uniform method and engineering tools of computer-aided design of sas both individually and as a part of the wing in the available scientific matter. in this regard, development of the scientific foundations and the method of computer-aided designing of sas wing as a component of the simulation modelling technique is a priority task. sas designing involves numerous aerodynamic, energetic, structural and geadvances in systems science and application (2015) vol.15 no.4 339 ometrical, technological and operating characteristics. thus, modern computer technologies are to be used for the analysis, synthesis and appropriate design solutions. in this connection, the design and construction of sas is reasonable to do using a high-precision mathematical simulation in the computer programs of computational aeroand fluid dynamics (cfd). 2 methods the purpose of this research is to develop methods of simulation modelling of the designed sas of the wing of the long-haul aircraft using the latest methods of high-precision mathematical simulation and engineering analysis. the object of the research is simulation modelling of aircraft sas of the wing within the proposed method of simulation. the subject of the research is the methodology of high-precision mathematical simulation of aircraft sas wing. the methodological provision of the study is in using the latest methods of mathematical simulation in the computer programs of computational aeroand fluid dynamics. due to the fact that designing and construction of the sas wing is a complex and modifying with time process producing a large number of iterations, there is a necessity to describe rationally the design process with simulation and physical models. we shall consider the process of simulation modelling of the secondary aerodynamic wing surface of the long-haul aircraft. simulation modelling means a method of a high-precision mathematical simulation in the computational aeroand hydrodynamics programme [4-6]. currently, to solve cfd problems, several software products can be used, for example: cosmosfloworks; nafems efd. lab; efd.v5 forcatiav5; salome; ansys. we propose salome product to be used. it is intended for solving problems of computational aeroand hydrodynamics. this is an open integrated software platform for numerical computations and simulation modelling. the salome user interface of the computational aeroand hydrodynamics provides the following options: it sets the master data and displays the results immediately in the graphic design window; it has a user friendly interface; it takes minimal time to prepare the data and view the results of the experiment. cad used is a finite-element pre-post processor, being the computational environment kernel surrounded by the united multitude of cae solvers. with 340 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... salome, one can also develop proprietory software solutions, which has been done by us, program for investigation and calculation of aircraft aerodynamic characteristics being the resulting product [7]. this programme based on cae solver open foam (liquids and gases flow analysis) helped us to carry out simulation of sas of the long-haul aircraft. the revised open foam package uses parasolid geometric simulation kernel which is a platform for the latest cad such as: ansys, icem-cfd, femap, solid works having full technical support. 2.1 discussion development of the wing sas 3d model was performed in catia high-level cad system. use of catia cad is conditioned by the fact that the system possesses a hybrid design function, combining both surface and solid elements in a single model. another important thing is the fact that the system has a function of free parameterization and constructs models with the help of the earlier created drawings, fig.1. fig.1 general view of sas, patent no.2481242 ni the number of working surfaces; χ1 leading edge sweep angle of the upper surface; χ2 leading edge sweep angle of the lower surface; bk tip chord advances in systems science and application (2015) vol.15 no.4 341 of the upper surface; b0 root chord of the upper surface; bk tip chord of the lower surface; b0 root chord of the lower surface; γ angle between the working surfaces; α installation angle end plate relative to the end rib axis; βleading edge sweep angle of the end plate. based on the previously developed sas drawing, its 3d model was created, fig.2. the 3d model developed was converted into step format to support to carry out simulation modelling in cad salome programme. step format makes it possible to describe the aerodynamic surfaces obtained using mathematical model of cad programme. step data format selection was done according to the requirements of the software used for simulation modelling. fig.2 3d sas model of the wing in catia system the figure shows that all the coupling parts of sas, namely the interface between the working surfaces have smooth fairings with the given radii. obtaining accurate aerodynamic shapes has become possible owing to the high-level catia cad system. in the dialogue box, a sas 3d model is mounted on the main wing of the long-haul aircraft on the left side in the flight direction [8]. the problem of computational aeroand hydrodynamics being solved leads to the analysis of the impact of air on the body. to make calculations it is necessary to set the initial data in the form of the initial and boundary conditions. the 342 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... body is restrained by the surfaces which limit the range of the fluid. such conditions are called a wall. due to the fact that for sas the case of external flow is used, which is characterized by such parameters as speed, pressure, temperature, etc., this type is called ambient conditions (environmental conditions). in accordance with the specified conditions, the external type of task will be solved. as a result, the design area is limited on one side by the field of the set boundaries in the form of the walls forming a rectangular hollow parallelepiped and solid sas model, placed in a hollow parallelepiped. the size of the design area is set either automatically or manually by the user. in this case, the design grid had been done by the user, which allowed us to obtain a more accurate mathematical problem solution by means of a smaller calculation grid in the design area, fig.3. the design area is divided by the grid with a different step: a smaller one immediately in the design area for more accurate solutions, further with a larger step as the high accuracy of the calculation is not required for the remaining area. the design area construction method used results in the efficient use of the computation resources leading to the significant reduction of the computation time [9, 10]. fig.3 shows the view of a 3d sas model on the right located in the design area. fig.3 design area thus, the design area shown in fig.3 is a spatial cube, which contains a 3d sas model of the wing of the long-haul aircraft, in what connection, boundary conditions wall in the form of four faces on axes oy and ox are set around the model, and the initial conditions graphically displayed in the form of two faces on axis oz are set at the front and at the back of the 3d model. the boundary advances in systems science and application (2015) vol.15 no.4 343 and initial conditions are set in accordance with the conditions, which sas in natural scale will be found under. the initial condition in the form of the face, located in front of the sas in the program is referred to as (inlet), and a face located behind the sas is called (outlet) i.e. an output parameter. setting inlet and outlet conditions, we thereby define the mode of the fluid flow in the design area. to eliminate the influence of the boundary conditions wall on the model sas to neutralize vortex formation around wall faces, the movement of the flow rate in the direction of the fluid has been set. the parameters of the fluid are set by the following initial data given in table 1. in addition to the parameters listed in table 1 and initial conditions, we set the table 1 initial data for the software experiment implementation item no name unit of measurement value 1 name air 2 density kg/v3 1.205 3 dynamic viscosity pa·s 0.000019137 4 specific heat capacity l/kg· 1.006 5 kinematic viscosity m2/s 0.0000158813 6 laminar number 0.9 7 turbulent prandtl number 0.85 8 thermal conductivity w/m· 0.024 9 reference (absolute) pressure pa 101.325 10 thermal expansion coefficient k1 0.00333 11 reference temperature k 300 turbulence model which suits in the best way to the current design case and select the model of turbulence (k-e), where the calculation of two additional equations for kinetic energy turbulence transport and turbulence dissipation transport. other turbulence models implemented in the program can be used, but in this design case it is advisable to use the very (k-e) model [11]. all further calculations will be performed for the case when the body is at rest while the medium is moving. the selected design case is applicable for finding ascensional power, drag force, induced drag and fluid force acting on the body located therein. the problem of computational aeroand hydrodynamics for sas wing being solved is made to study the movement of air around the body. solving the problem using cad gives us the parameters describing the fluid flow, namely: speed, pressure, temperature, etc. simulation modelling of sas of the a/c wing in cfd in salome programme has been performed within the time equal to 200 calculation steps. the number of the calculated steps is set by the user and selected in the view of the following consideration: to obtain the sufficient experimental data under the steady flow process. 344 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... eventually, repeatedly conducted purges in cfd programme allowed us to establish that the steady flow process starts with the 65th design step when the interference processes become steady, and the number of iterations of 200 design steps selected from the obtained information volume of 250 gb. a larger volume of information is not required, since the available data is sufficient for making design decisions. fig.4 fluid distribution rate change over sas fig.5 pressure distribution in sas section the results of simulation modelling presented further feature the same moment of time corresponding to design step no.180 in various design sections. fig.4 shows sas of the wing with the air distribution rate change in the form of relief advances in systems science and application (2015) vol.15 no.4 345 growths that change their colour with the distribution rate over sas. flow distribution value is found using the graduated scale changing its colour from red to blue and graph curves shown below. the field of low rates is marked with the red colour and the field of high rates is marked with the blue colour. fig.6 pressure distribution in the wingtip section. fig.7 pressure distribution over sas of the wing fig.5 shows sas section using which we can find the change in pressure over the sas. this section clearly shows pressure rise at the leading edge of sas and also a marked increase at the trailing edge, and further decrease due to the distance from the model owing to air circulation. fig.6 presents a section of the wingtip, which shows pressure distribution over 346 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... the surface of the wing airfoil located close to sas. fig.7 is a 3d sas model showing the pressure distribution fields and the set cutting plane at the wingtip shown by the section in fig.6. the pressure distribution over sas makes it possible to trace the character of pressure change over time. fig.8 shows the enlarged section of the wingtip close to sas displaying air particles vector movement and rate distribution fields. fig.8 pressure distribution over sas of the wing a graphic curve in the form of the bar graph in fig.9 shows the change in the reynolds number (re) and kinematic speed over sas conforming to fig.4. fig.9 graphic curve of re number change and kinematic speed over sas advances in systems science and application (2015) vol.15 no.4 347 fig.10 a curve of a temperature change, kinematic speed and pressure change over sas (a) (b) (c) fig.11 secondary aerodynamic surface, patent no. 2481242 348 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... fig.10 shows a curve describing air parameters, namely, a temperature change, kinematic speed and pressure change over sas conforming to fig.4. simulation modelling for sas wing made it possible to find the aerodynamic characteristics of the designed sas and to confirm flow processes around sas. let us consider the flow for sas wing, patent no.2481242. while the air is flowing the wing, the air overflows from the lower plane of the wing to the upper one, at the same time, at the endplate 1 supplied with the additional wind swept surface 2 of low aspect with a sharp front edge 3, mounted on the outer side of the endplate 1, there formed a vertical chamfering field transforming into a stable vortex flow with a conical vortex formation on the leading edge 3 of the secondary aerodynamic surface 2 mounted on the endplate 1, fig 11. at lower vertical aerodynamic surface 8, the vertical chamfering field due to the low aspect the lower surface, does not result in the early formation of the vortex at the leading edge 9 but it is transformed at the surface end into an end conical vortex. 3 results the sas proposed improves aerodynamic efficiency of the long-haul a/c and provides a maximum effect from air overflowing throughout the entire field of the effective values [12, 13]. in this way, it makes it possible to enlarge the effective wingspan reducing the induced drag generated by the vortex flying from the sweptwing tip and, consequently, increasing the ascensional power at the wingtip; to increase the effective aspect of the wing, without changing its span; to improve fuel efficiency of the a/c and the flying range. 4 conclusions 1. high-precision computer mathematical simulation and simulation modelling techniques in the analysis of the flow process around the designed sas model provides a means to find out the aerodynamic characteristics for given initial and final conditions affecting the induced drag of the aircraft and ascensional power generated by the system of lifting surface areas of the a/c. 2. using a simulation modelling technique can be used for complex multivariate iterative calculations of the new sas designs, high quality design solutions being secured due to the accumulation of output data followed by statistical analysis with the time from the previous to the next point in time set with the predefined step. 3. the behaviour of the flow processes around the sas model designed identified as a result of the high-precision mathematical simulation modeling enables us to confirm or deny (in the case of unsound designing) aerodynamic characteristics obtained at the preceding stages and to develop physical sas models. advances in systems science and application (2015) vol.15 no.4 349 acknowledgement the project has been performed under contract no.14.z56.15.5527-mk of february 16, 2015, grant of the president of the russian federation for the state support of young russian scientists to carry out research on the subject of “computer-aided design of secondary aerodynamic wing surfaces as an element of the aircraft life cycle”. the outcomes of the scientific research have been accepted for implementation in jsc. kapo tupolev (kazan). references [1] v.a. komarov et al (2013), conceptual design of the aircraft, house of samara state aerospace university, samara. pp.120. [2] llc novaya tekhnika enterprise(2013), ontology of design, novaya tekhnika publishing,samara. no.1. [3] daniel p. raymer.,(1992), aircraft design: a conceptual approach, american institute of aeronautics and astronautics, washington. pp.391. [4] shannon, r. (1978),“systems simulation modelling”, art and science, pp.268. [5] pavlovsky, yu.n. (2000), simulation models and systems, fasis: exhibition centre of russian academy of sciences, pp.134. [6] clayton bargsten winglets. (2011), striving for wingtip efficiency, national aeronautics and space administration, washington dc: uncl, pp.68. [7] a.v. gordiyenko, a.a. gorbunov,a.d. pripadche. (2013), “program for research and calculation of the aircraft aerodynamic characteristics: a certificate of the official registration of the software”, certificate no. 2013616240 the russian federation.applicant and patent holder orenburg state university. no. 2013616240. [8] a.a. gorbunov,a.d. pripadchev (rf).(2013), “aircraft wingtip”, patent ru no. 2481242, ecc 643/10.no. 2011148436. [9] jeppe johansen and niels n. sorensen, (2006), “aerodynamic investigation of winglets on wind turbine blades using cfd”, roskilde, denmark: riso r-1543. [10] hepperle, m.(2005),“aerodynamic optimisation of a flying wing transport aircraft”, new results in numerical and experimental fluid mechanics v 350 simulation modelling for computer aided design of secondary aerodynamic wing surfaces... (notes on numerical fluid mechanics and multidisciplinary design), vol.92, pp.69-76 [11] g. j. kennedy.(2012), “aerostructural analysis and design optimization of composite aircraft”, phd thesis, university of toronto. [12] schiktanz, daniel.(2011), “conceptual design of a medium range box wing aircraft”. department fahrzeug technik und flugzeugbau, master thesis, hamburg, haw hamburg. [13] joseph katz, allen plotkin.(1991), “from wing theory to panel methods”, low speed aerodynamics, singapore, pp.632. corresponding author tatiana can be contacted at:yal05@mail.ru adv sist sci appl 2017; 17(2); 52-62 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/253 copyright ©2017 assa. adv. in systems science and appl. (2017) a system approach to the regional sustainable management gennady a. ougolnitsky southern federal university, rostov-on-don, russia e-mail: ougoln@mail.ru abstract. a system approach to the regional sustainable management based on game theoretic models and information technologies is considered. a notion of the regional active system is specified. the author's approach uses hierarchical dynamical game theoretic models formalizing and integrating many different aspects of the control in active systems. a concept of the regional sustainable development is added by the requirement of system compatibility. sustainable development means that a human activity should provide sufficiently high values of indicators of the social-economic development, and in the same time sustain an environmental equilibrium. formally, all indicators of the state of a regional socio-ecological-economic system must belong to a given range. in turn, system compatibility means that interests of the active agents should be considered by the center in a maximal degree for the achievement of the center's purposes. for a quantitative measurement of the system compatibility based on the proposed model an index of system compatibility is used. administrative and economic control mechanisms providing the system compatibility are described. on the regional level the problem of coordination of interests has the following specific features. first, a regional administration which should be a leader of sd, in a higher degree is interested in the approval from the federal center that estimates the regional activity by other criteria. second, the principal control levers, both legislative and economic, are also situated in the federal level. third, the municipal districts that should provide a practical implementation of the requirements of homeostasis, have not enough resources for it. a hierarchical structure of the model system in this domain is considered. a practical implementation of the methodology of sustainable management on the regional level requires a regional information-analytical system, the structure of which is proposed. key words: computer simulation, game theoretic modeling, homeostasis, regional development, sustainable management, system approach, system compatibility. 1. introduction system approach and its applications were extremely popular in 1960-1980s [25]. since that time up to the moment this approach has been developed not so actively though the system methodology is not completely exhausted [6]. it should be noticed that a system, complex approach was always used by russian scientists starting from michael lomonosov. vassily dokuchaev, kliment timiryazev, dmitry mendeleev, konstantin tsiolkovsky, alexander chizhevsky, alexander bogdanov and especially vladimir vernadsky belong to that glorious cohort. a great role in the system research based on mathematical modeling was played by works of andrey kolmogorov, leonid kantorovich, andrey lyapunov, nikita moisseev, alexander samarsky, sergey kurdyumov and many others. an essential contribution to the system methodology was made by the notion of active system proposed by vladimir burkov and his school in the institute of control sciences [5]. this notion formalizes relations in active systems of different nature including human beings (organizational, economic, social systems). an active system consists of active elements (agents) having their own objectives and interests and possibilities to achieve them. as a rule, one of the agents is separated as the center reflecting the objectives and interests of the whole system. the interests of the agents and the center do not coincide but are not also antagonistic that allows for some trade-off solutions acceptable for everybody. 53 g. a. ougolnitsky copyright ©2017 assa. adv. in systems science and appl. (2017) an informational interaction of the agents with the center and each other may include a strategic manipulation (a deliberate misrepresentation of the reported information in private purposes). the main problem of the theory of active systems is a concordance of the interests of the center (the whole system) and the active agents as elements of the system. for this purpose special administrative, economic, and informational control mechanisms are used [17]. close ideas are developed in the information theory of hierarchical systems [8], theory of incentives and mechanism design [1,12]. a widely spread since 1980s concept of sustainable development has also a system character. thus, a key approach of "three pillars" in the analysis and implementation of the sustainable development requires to consider social, economic, and environmental factors in their totality as a whole system. it should be also noticed an important idea that a global sustainable development is formed by the sum of local efforts in this direction [21]. an important class of active systems is generated by territorial formations. here the problem of coordination of private and public interests is also of the main importance and may be formulated in terms of state regions, regions municipal districts, municipal districts economic agents depending on the level of consideration. in russia this problem was especially essential in 1990s in the connection with the development of federalism, and it is still actual [15,16]. the author's approach to the analysis of regional active systems is based on the general complex theory of sustainable management in active systems [20]. the approach uses hierarchical dynamical game theoretic models formalizing and integrating all the named aspects. the complexity of these models and big data volumes make necessary an application of modern information technologies of the research support. the regional management is an example of the possible applications of the developed universal theory. the rest of the paper is organized as follows. in the section 2 a notion of a regional active system is considered. a specified concept of the regional sustainable development with consideration of the condition of system compatibility is proposed in the section 3. in the section 4 mechanisms of system compatibility and monitoring on the regional level are described. in the section 5 a hierarchical structure of the model system of the sustainable management is considered. section 6 concludes. 2. a regional active system a general structure of the regional active system is presented in fig.1. fig.1 a regional active system a system approach to the regional sustainable management 54 copyright ©2017 assa. adv. in systems science and appl. (2017) on the higher control level a regional administration (ra, center) is situated. the middle control level is presented by municipal districts (md). a more complicated variant is also possible when urban municipal formations and their districts are considered separately. the lower control level is formed by local active agents (la): enterprises, organizations, firms, individual entrepreneurs. the control object is the regional socio-ecological-economic system considered as a hierarchically controlled dynamic system [20]. in the majority of cases la are economic agents having certain financial, human, and other resources and tending to maximize their income (profit). however, non-commercial organizations and local control agencies can also be considered as la having non-economic objectives. economic la can have social and ecological objectives too. mds solve the problems of social-economic development of the respective territories with consideration of the environmental requirements using certain budget resources [15]. the objectives and possibilities of ra are structurally similar to the ones of md but differ from them by the volume and sources of financing [16]. ra can establish its own laws, and have additional means of influence to md and la. a control object is characterized by three groups of indicators describing the social sphere, economic potential, and environmental state respectively. for determination of the specific values of the indicators the data of official statistics are used which may be complemented by expert estimates. notice that only las exert a direct influence to the regional socio-ecologicaleconomic system. the higher control levels just create some administrative, legislative, and economic framing conditions of the influence. the laws of change of the state indicators due to natural factors and la's impact are described by dynamic balance relations. the left hand side contains a new value of the indicator, the right hand side contains an old one plus increase minus decrease generated by the considered processes. it is more appropriate to set the problems of sustainable management in an infinite period of time. however, a finite, long enough period may be more acceptable in practice. some discount rate should reflect a relation between incomes and expenses in different moments of time. most typical political technologies of regional elites include:  a concentration of financial resources;  a power pressure to the opponents;  an establishment of control and strategic partnership with the most powerful regional economic agents;  using of information flows from the federal center for the protection of the elites' interests;  negotiation process with the political leaders;  a direct or indirect control on the information flows and media, an active construction of a positive image of the regional administration by media;  a formation of the regional mythology;  a harder selection of the new candidates to the regional elites [13]. technically, the administrative controls of ra and md include legislative restrictions of the activity of subordinated control agents (norms, quotas, quality standards). the economic controls represent tax rates, parameters of privileges, subsidies, grants. the laws of state dynamics are determined by natural processes and social-economic activity. for their description specific models of the mathematical ecology and economics, and dynamic forecasting models of the regional level are used [18]. the respective three-level game theoretic model of regional control has the form     t t ttttttt xuqpsrgej 1 00 max),,,,,( (1) ;; ssrr tt  (2) 55 g. a. ougolnitsky copyright ©2017 assa. adv. in systems science and appl. (2017)     t t tt i t i t i t ii t i xuqprgej 1 max),,,,( (3) );();( t ii t i t ii t i sqqrpp  (4)     t t tt i t ij t iij t ij xuprgej 1 max),,,( (5) );,( t ij t iij t ij qsuu  ;,...,1;,...1 nikj i  (6) ;),,( 0 01 xxuxfxx tttt  (7) );,...,();,...,( 11 t n ttt n tt sssrrr  );,...,();,...,();,...,();,...,( 1111 t ik t i t i t ik t i t i t n ttt n tt ii qqqpppqqqppp  ),...,,...,,...,(),...,( 11111 1 t nk t n t k tt n tt n uuuuuuu  . here t is a period of consideration (in years);  is a discount rate; iji jjj ,,0 and iji ggg ,,0 summary and current payoff functions of ra, mds, and las respectively; tt sr , are economic and administrative controls of ra in the year t; sr, are the respective sets of feasible controls; t i t i qp , are economic and administrative controls of the i-th md in the year t; ii qp , are the respective sets of feasible controls; t iju is a control of the ij-th la in the year t; iju is the respective set of feasible controls; tx is a state vector of the regional socio-ecological-economic system in the year t; 0x a vector of initial values of the state indicators on a base year; f a set of models that describe the state dynamics; n a number of mds; ik a number of las in the ith md. the model (1)-(7) represents a hierarchical differential game, a discrete form of which is oriented to simulation analysis based on the scenario method [14]. solutions of the game (1)-(7) are understood in the sense of stackelberg [3]. as a rule, the ra chooses only ts while tr is fixed (administrative control, or compulsion) or vice versa (economic control, or impulsion). similarly, given ts or tr the mdi choose t iq while t ip are fixed or vice versa. at last, given t iq and t ip as well as ts and tr the laij choose their control parameters t iju . in fact, simulation modeling is almost unique possibility to solve the game (1)-(7) in the general setup. some simplified versions may be investigated by means of standard techniques like hamilton-jacobibellman equations or pontryagin' s maximum principle together with numerical methods. 3. sustainable development a serious attention to the problem of sustainable development (sd) was attracted by the report “our common future” (1987) prepared by the un world commission on environment and development, known as “brundtland commission”. conclusions of the brundtland report formed a base for the decisions of the un conference on environment and development (rio de janeiro, 1992). the brundtland report defines sustainable development as the development which satisfies the needs of present generation and does not undermine the possibility of future generations to satisfy their needs [21, p.43]. the declaration on environment and development accepted at the rio-92 conference includes 27 principles, such as principle 3 “a right to the development should be realized to provide a just satisfaction of the needs of the present and future generations in the domains of development and environment” and principle 4 “to provide the sustainable development an environmental protection should be an integral part of the development process and cannot be considered without it”. a system approach to the regional sustainable management 56 copyright ©2017 assa. adv. in systems science and appl. (2017) in spite of the active discussion of the concept of sustainable development and a number of accepted official documents, a unity in the definition and interpretation of the notion is still absent. already in the early work [23] more than 60 definitions of sustainable development given by different authors are cited. the general transformations of the environmental scene are shown in table 1 [26, p.84]. table 1 evolution of the views on sd in 1970 2000s environmental policies 1970-1980s 2010 iconic policy instrument command and control collaborative and market-based instruments key group relied upon government government and stakeholders dominant mode of action work with industry, mostly through technology work with industry and consumers, through technology, economy, finance knowledge about the environment superficial, or limited to specialists extensive, diffused in many realms of society discourse on green issues view of green as important (before 1973) and then as rather marginal acknowledgment of the seriousness of some problems: a trend of “going green” and resulting business opportunities social issues superficial or neglected concerns concerns for environmental justice and the influence of environmental degradation on poverty based on the initial notion of sd of the society it is possible to say about sd of any active system [20]. according to the existing approach, the main role in sd is played by the notion of homeostasis. it means that a human activity should provide sufficiently high values of indicators of the social-economic development, and in the same time sustain an environmental equilibrium. formally, all indicators of the state of a regional socio-ecological-economic system must belong to a given range that can be written as the condition *,...,1 xxtt t  . (8) a stronger formulation is also possible when the condition of homeostasis is treated as an asymptotic approximation of the values of indicators to their ideal values [20]. it seems that the main problem of the practical implementation of the concept of sd consists in the absence of interested and plenipotentiary agents of the implementation. the criteria used by acting politicians not always coincide with the conditions of homeostasis of the territories governed by them. that's why the concepts of transition to sd accepted by many countries and regions remain only as declarations in a great part. it should be noticed in the same time that from an objective point of view all politicians who wish to enlist the support of the population and secure the power for a long time, should be interested in the conditions of homeostasis. it concerns also the owners of companies thinking about conservation of their family business for many generations. on the level of a firm the considered problem is tightly connected with strategic delegation (see [24]), i.e. sharing of the powers between owners, topmanagers, and managers of lower levels. as for territories, the problem is in the distribution of resources and powers between federal, regional, and local agents of power and control (the problem of federalism). it is especially important that the values of sd be shared by active agents of the lower level whose activity has a direct impact to the state of a controlled system. that's why the concept of homeostasis, playing a key role in the traditional approach to sd, should necessarily be complemented by a requirement of system compatibility [2,7]. it means that interests of the 57 g. a. ougolnitsky copyright ©2017 assa. adv. in systems science and appl. (2017) active agents should be considered by the center in a maximal degree for the achievement of the center's purposes. for a quantitative measurement of the system compatibility based on the model (1) (7) it is expedient to use an index of system compatibility in the form * 0 max 0 jjsci  , (9) where );,,,,,(maxmaxmax 0 ),( )( )( max 0 xuqpsrjj qpuu sqq rpp ss rr       (10) ),,,,,(minminmax 0 ),(),(, * 0 xuqpsrjj qpneusrneqp ss rr     . (11) the substrahend in the formula (9) gives a guaranteed payoff of the center in the worst nash equilibrium in the game of agents regulated by her, and the minuend shows her payoff in the case of complete cooperation of the agents with the center (globally maximal value of the payoff in the team solution). the value 0sci indicates the complete (ideal) system compatibility. this approach generalizes an idea of the price of anarchy proposed for network games [22], and its development in other works [4]. an indicator of the ideal compatibility is proposed in [8] as well. given (9) a requirement of the system compatibility is expressed by the condition *scisci  , (12) where *sci is an expert threshold value of the system compatibility index. the condition (12) may be used additionally after the solution of the game (1)-(7) with phase constraints (8) to check whether the system compatibility holds. the relations (1)-(12) determine a holistic formal description of the problem of sustainable management in any active system [19]. on the regional level the problem of coordination of interests has the following specific features. first, a regional administration which should be a leader of sd, in a higher degree is interested in the approval from the federal center that estimates the regional activity by other criteria. second, the principal control levers, both legislative and economic, are also situated in the federal level. third, the municipal districts that should provide a practical implementation of the requirements of homeostasis, have not enough resources for it. thus, the problem of coordination of interests on the regional level has a special importance and deserves a detailed investigation by means of systems analysis, mathematical modeling, and information technologies. 4. control mechanisms in the case of a low system compatibility ( 0sci ) the center should use special control mechanisms to provide it. to develop a classification of the control mechanisms three attributes characterizing the center's strategy can be used: 1) absence/presence of a feedback of the center's strategy on the state of a controlled dynamic system (cds). this attribute has two basic values: open-loop strategies (ol) which depend only on the instant of time t , and closed-loop strategies (cl) which depend on the game position ))(,( txt [3]; 2) absence/presence of a feedback of the center's strategy on the agents' strategies. in the first case we deal with stackelberg games (st), and games of the second type we propose to call germeier games (ger), or incentive stackelberg games [9-11]; 3) methods of hierarchical control. here we differentiate compulsion, when the leader influences the followers' sets of feasible strategies, and impulsion, when she influences the followers' payoff functionals [20]. so, in the case of compulsion the leader chooses ts in the a system approach to the regional sustainable management 58 copyright ©2017 assa. adv. in systems science and appl. (2017) higher control level and tq in the middle control level while impulsion means the respective choice of tr and tp . a classification of the control mechanisms on example of the russian traffic rules (rtr) is given in table 2. here an illustrative active system is considered. in this system the center is the federal state that establishes rtr and the penalties for their violation exposed in the administrative codex (ac) and the criminal codex (cc); an active agent is a driver and his vehicle; a controlled object is a road and its environment. the respective clauses of the legislative documents are presented in brackets. administrative control mechanisms explicitly forbid some driver's actions, and economic ones charge penalties for their violation. from the mathematical point of view, a type of the center's control mechanism determines an information structure of the difference game (1) (7). for example, а/st/cl is a stackelberg game in closed-loop strategies in which the center restricts the agents' sets of feasible strategies. the solutions of respective differential games are defined in [20]. table 2 classification of the control mechanisms on example of the russian traffic rules st/ol st/cl ger/ol ger/cl administrative mechanisms it is forbidden to injure or pollute a road cover (1.5 rtr) in populated areas it is allowed to drive a vehicle with the speed not more than 60 km/h (10.2 rtr) it is forbidden to drive a vehicle with a working brake system disrepair (2.3.1 rtr) it is forbidden to drive a vehicle with not burning (absent) head-lights and back marker lights in the dark period (2.3.1 rtr) economic mechanisms absence of an insurance policy 800 roubles (12.37 p.2 ac) exceeding of the allowed speed: more than 20, but not more than 40 km/h 500 r. (12.9 p.2 ac); more than 40, but not more than 60 km/h 1000-1500 r. (12.9 p.3 ac) driving a vehicle by a driver in the state of intoxication 30000 r. (12.8 p.3 ac), repeatedly 200000 300000 r. (cc 264) non-compliance of the requirement to stop before the stop line: first time 800 r. (12.12. p.2 ac); repeatedly 5000 r. (12.12 p.3 ac) a very important role in the practical implementation of the control mechanisms is played by the monitoring system that provides to the control agents information about the state of the object and therefore ensures a feedback. a general form of the monitoring in a simplified hierarchically controlled dynamic system with one agent is shown in fig. 2 [20]. here slx is a center's information about the state of cds; fslx is a center's information about the impact of the agent on cds; flx is a center's information about the agent's state; sfx is the agent's information about the state of cds. in conformity with a regional active system the presented model can be used both on the level ra md, and on the level md la (see fig. 1). in the former case cds is treated as a socio-ecological-economic system of the respective md, and in the latter one as a part of that system controlled by the given la. a union of those systems forms a regional monitoring system, the visual presentation of which is huge. the information is collected by means of official statistics which can be complemented by special sample audits and expert estimates. 59 g. a. ougolnitsky copyright ©2017 assa. adv. in systems science and appl. (2017) fig.2 monitoring in a hierarchically controlled dynamic system (hcds) 5. hierarchical structure of the model system the regional level has an intermediate position in the structure of state control (fig. 3). here fc corresponds to ra in fig. 1, rai corresponds to mdi, and mdij corresponds to laij. the dynamics of a regional socio-ecological-economic system is described by a system of nonlinear equations (7), and the interests of agents and their possibilities by the relations (1)-(6). the respective system connections are presented in table 3. fig.3 hierarchical structure of the model system table 3 connections of the regional control level with federal and local levels inputs outputs fc federal laws and instructions informal directions of the federal power federal budget financing tax revenues to the federal budget information about the state of the socioecological-economic system of a region (indicators of the official statistics) md tax revenues to the regional budget information about the state of the socioecological-economic system of a region by separate md regional laws and instructions informal directions of the regional administration regional financing of the local budgets a system approach to the regional sustainable management 60 copyright ©2017 assa. adv. in systems science and appl. (2017) a practical implementation of the methodology of sustainable management on the regional level requires a regional information-analytical system, the structure of which is shown in fig.4. fig. 4 structure of the regional information-analytical sustainable management support system here the optimization subsystem includes mathematical models like the game (1)-(7) while the simulation subsystem allows for their solution. the expert subsystem additionally uses an expert knowledge. 6. conclusion the system approach plays an important methodological role in the solution of difficult complex interdisciplinary problems of the regional sustainable management. a specific character of the regional level of control consists in its intermediate, dual nature. from one side, a regional administration plays a role of the center relating to the municipal formations within the region, from the other side it is itself an active agent in the interrelation of the regions with their federal center. in spite of it, ra can and must be an agent of the sustainable management. it follows from the hierarchical nature of sustainable development, the achievement of which on the global level is provided by the totality of efforts of all agents on the lower control levels. this fact determines a necessity of the delegation of an essential volume of administrative and economic powers on the ra level. it should be noticed that the need of solution of the problems of sustainable management is not always realized by the agent, and the mission of science and public opinion formed by the science is of great importance here. a key role in the solution of the sustainable management problems is played by the coordination of interests of the center as the subject of sustainable development and the active agents who implement it directly. the requirements of homeostasis will not be satisfied by themselves unless specific agents be interested in it. in the great majority of cases the interests of active agents contain only maximization of their economic payoff in short periods of time. that's why it is necessary to design and implement administrative and economic control mechanisms that incorporate the conditions of homeostasis in the agents' payoff functions and feasible sets of strategies. later it is planned to analyze horizontal connections in a regional active system that leads to the cooperative games as conflict control models and investigation of the time consistency of their solutions. it is supposed particularly to apply the described methodology for decision support of the sustainable development of the rostov region (russian federation). 61 g. a. ougolnitsky copyright ©2017 assa. adv. in systems science and appl. (2017) acknowledgements the paper is supported by the russian science foundation, project #17-19-01038 references [1] nisan, n., roughgarden, t., tardos, e. & vazirani v. (eds.) (2007). algorithmic game theory. new york, ny: cambridge university press. [2] antonenko, a.v., gorbaneva, o.i. & ougolnitsky, g.a. (2016). concordance of private and public interests: dynamic graph representation, identification and simulation modeling, adv. in systems science and applications, 16 (4), 43-52. [3] basar, t. & olsder, g. j. (1999) dynamic noncooperative game theory. philadelphia, pa: siam. [4] basar, t. & zhu q. (2011). prices of anarchy, information, and cooperation in differential games, dynamic games and applications, 1(1), 50-73. [5] burkov, v. n. & opoitsev, v. i. (1974). metagame approach to the control in hierarchical systems, automation and remote control, 35(1), 93-103. [6] cabrera, d. & cabrera, l. (2015) systems thinking made simple: new hope for solving wicked problems. ithaca, ny: odyssean press. [7] gorbaneva, o.i. & ougolnitsky, g.a. (2013). purpose and non-purpose resource use models in two-level control systems.adv. in systems science and applications, 13(4), 379-391. [8] gorelik, v., gorelov, m. & kononenko a. (1991). analiz konflictnyh situacii v sistemah upravlenia [analysis of conflict situations in control systems]. moscow, russia:radio i sviaz [in russian]. [9] gorelov, m.a. & kononenko, a.f. (2014). dynamical conflict models. i. language of modeling. automation and remote control, 75(11), 1996-2013. [10] gorelov, m.a. & kononenko, a.f. (2014). dynamical conflict models. ii. equilibria. automation and remote control, 75(12), 2135-2151. [11] gorelov, m.a. & kononenko, a.f. (2015). dynamical conflict models. iii. hierarchical games. automation and remote control, 76(2), 264-277. [12] laffont, j.-j. & martimort, d. (2002). the theory of incentives. the principal-agent model. princeton , ny: princeton university press. [13] lapina, n. & chirikova, a. (2000). strategii regionalnykh elit: economika, modeli vlasti, politicheskii vibor [strategies of regional elites: economics, power models, political choice]. moscow, russia: inion ras, [in russian]. [14] law, w. & kelton, a. (2000). simulation modeling and analysis. new york, ny: mcgraw-hill. [15] leksin, v. & shvetzov, a. (2000). municipalnaia rossia. socialno-ekonomicheskaia situacia, pravo, statistika [municipal russia. social-economic situation, law, statistics] moscow, russia: urss, [in russian]. [16] leksin, v. & shvetzov, a. (2016). gosudarstvo i regioni: teoria i praktika gosudarstvennogo regulirovania territorialnogo razvitia [state and regions: theory and practice of government regulation of territorial development]. moscow, russia: librocom, [in russian]. [17] novikov, d. (ed.) (2013). mechanism design and management: mathematical methods for smart organizations. new york, ny: nova science publishers. [18] gurman, v. & ryumina, e. (ed.) (2001). modelirovanie socio-ekologo-ekonomicheskoi sistemy regiona [modeling of a socio-ecological-economic regional system], moscow, russia: nauka, [in russian]. [19] ougolnitsky, g. a. (2015). sustainable management as a key to sustainable development in reyes d. (ed.) sustainable development: processes, challenges and prospects (pp.87128). new york, ny: nova science publishers. a system approach to the regional sustainable management 62 copyright ©2017 assa. adv. in systems science and appl. (2017) [20] ougolnitsky, g. (2016). ustoichevoe razvitie v aktivnyh sistemah [sustainable management in active systems]. rostov-on-don, russia: izd. yuzhn. fed. univ., [in russian]. [21] world commission on environment and development (1987). report of the world commission on environment and development: our common future. [online], available http://www.un-documents.net/wced-ocf.htm. [22] papadimitriou, c. h. (2001). algorithms, games, and the internet. proc.33th acm symp. theory of computing, heraklion, greece, 749-753. [23] pezzey, j. (1989). economic analysis of sustainable growth and sustainable development. the world bank. [24] stamatopoulos, g. (2016). cournot and stackelberg equlibrium under strategic delegation: an equivalence result. theory and decision, 81, 553-570. [25] emery, f. e. (ed.) (1972) systems thinking. vol. 1-2. lincoln, uk: penguin. [26] zaccai, e. (2012). over two decades in pursuit of sustainable development: influence, transformations, limits. environmental development, 1, 79-90. advances in systems science and applications (2017) vol.17 no.1 9 regional security: analysis of the emergency management effectiveness based on the scenario approach vladimir l. schultz a , vladimir v. kulba 1b, oleg a. zaikin c, alexei b. shelkov b, igor v. chernov b a) institute of socio-political research of the russian academy of sciences, 119333, fotievoi street, 6, building 1, moscow, russia. b) v.a.trapeznikov institute of control sciences of the russian academy of sciences, 117997, profsoyuznaya street, 65, moscow, russia. c) warsaw school of computer science, 00-169, lewartovskiego street, 19, warsaw, poland. abstract. the paper analyses the effectiveness of the scenario analysis and modeling methods for solving the planning and operational control problems of in the processes of prevention and elimination of the causes and effects of man-made disasters, industrial catastrophes and emergencies. the scenario approach belongs to a class of object-oriented methods used to represent information about the state of a plant and an environment, which is necessary to generate control actions for the rescue measures and activities, as well as disaster management. the main feature discussed in this paper is a class of industrial safety management tasks based on the scenario approach which have to be developed as well as the general normative framework that allows to fundamentally change the approach to building models of an emergency. the research shows that in this case the procedure of forming the base model using a comprehensive analysis of existing technical regulations, norms and standards were quite effective. the results of modeling and scenario analysis of the processes, elimination of technological accidents consequences on transport infrastructure of metropolises are presented in the paper. key words: management, scenario analysis, man-made disaster, emergency, simulation, signed graphs. 1. introduction in the last decades, the world has witnessed a stable trend of substantial growth of material losses as a result of industrial catastrophes, emergencies and natural disasters. one of the main reasons for this, apart from global climate changes accelerating in the coming century and the saliently manifested synergetic nature of many industrial catastrophes, is the lack of willingness of warning management systems to the prompt, adequate and effective response to such events. the main peculiarities of functioning of management processes in conditions of industrial disasters consist in the fact that an emergency (disaster) arises and develops suddenly and unexpectedly. since its appearance in the management system poses the problems which essentially aren’t inherent for the stationary mode of operation. thus, that is especially important that the measures to counter the emergency development and elimination of its consequences should be taken without delay and been as efficient as possible. at the same time the fundamentally new problems arise in management systems, which have been complicated by the powerful stream of information that is required to examine and analyze promptly. analysis of management processes in emergency situations (es) has allowed to identify a number of features compared to the daily activities mode. the most typical of them are given in the table 1 [1]. a management of prevention and dealing with consequences of technological disasters should cover the whole range of issues relating to emergencies, the most important of which are the stages of emergency forecasting, proactive planning and operational management of the liquidation of the causes and consequences of the emergency under high uncertainty. the research is focused on the effectiveness of different methods for scenario analysis and modeling in the planning and operational management measures to prevent and eliminate causes and consequences of man-made disasters and emergencies. 1 corresponding author. email: kulba@ipu.ru 10 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach table 1. comparative characteristics of management systems daily activity mode management in emergency situations permanent mode of operation different operating modes strict structure and clear allocation of functions for a long period the lack of a strict structure and a clear allocation of functions for a long period, flexibility, aggressiveness the narrow functional focus broad and partially unpredictable scope monostructure polystructure regulated information flows the dependence of the information flows on the current situation accurate information uncertain information redundant information insufficient information the low rate of change a high rate of change unpredictability of the situations unpredictability of the situations; orientation on past experience usually does not make sense the principle of unity of powers and responsibilities combination of the principles of unity of command and the distribution of powers and responsibilities functional potential organizational capacity the predominance of mostly socio-economic objectives and performance criteria the objectives – efficiency and effectiveness in liquidation the causes of emergencies and their consequences; criteria minimizing the time to achieve the objectives, a minimum loss (victims) in liquidation of emergencies 2. planning of prevention and liquidation the causes and consequences of manmade disasters and emergencies unlike traditional planning and control systems [2], which are intended to solve strategic tasks for a long time period, the control system in emergency should operate in real time. strategic tasks should be solved in a limited time interval promptly and continuously. in practice this also means periodic adjustment of the list of key strategic tasks and continuous monitoring of new emergency events. moreover, a control system in emergency situations must be ideologically focused on the need, the possibility and the ability to operate in extreme conditions and to adapt quickly to changing conditions. the most important task of the control system under these conditions is not only operational, but also long-term planning for prevention and elimination of consequences of emergency situations on the transport infrastructure of urban agglomerations. the principal features of the planning and management processes of eliminating the consequences of manmade disasters and emergencies are: (1) partial predictability of serious problems and possibilities of their solution; (2) partial predictability of places of appearance and development of emergencies; (3) low predictability of the scale and the time of a disaster; (4) unpredictability of adverse events and situations arising from the emergence and development of emergencies, i.e. availability of strategic surprises. these conditions dictate the need for pre-emptive use of planning and management methods, based on anticipation of problems, situations and events, making flexible emergency solutions, oriented to the external system environment (natural environment, living conditions of population and staff of enterprises and organizations, socio-political situation etc.). such methods can be effectively implemented in the unified modern concept of strategic and tactical planning and operational management. full planning and control cycle of risk of industrial catastrophes and emergencies, as well as the liquidation of their consequences include:  forecast risk and severity of the consequences of emergency situation by generating scenarios; advances in systems science and applications (2017) vol.17 no.1 11  forming objectives and risk management criteria;  strategic (long-term) planning of preventive measures;  tactical (current) planning of alternative response to the emerging threat of disaster;  strategic and operational management in emergency situations. the risk of occurrence and development of es of natural and industrial type is forecasted on the basis of constructed scenarios of these situations. the concept of the scenario approach in control theory is to some degree new, though now has been used widely, especially in the analysis of strategic decisions in organizational management [3]. during the study of management processes the scenario of prevention of es and liquidation the consequences of industrial catastrophes can be considered as a tool for formal analysis of alternative variants of es development on a control object for the given objectives and criteria under conditions of uncertainty, when it is impossible to directly form a specific and detailed plan of measures on liquidation of es consequences under the existing time constraints. at its core, the scenario approach belongs to a class of object oriented methods of presenting information on control object’s state and environment, which is necessary to generate control actions for rescue measures and activities, as well as disaster management. in other words, the main task to be solved in terms of the scenario approach is to form the necessary data for preparation and making effective strategic and operational decisions, as well as a comprehensive analysis of the consequences of these decisions’ implementation in a variety of conditions. thus, a scenario of a problem situation development is needed as an intermediary between the stage of goal setting and the stage of formation and implementation of specific management decisions aimed at achieving the objectives for the prevention and liquidation of es consequences. in its content, a scenario of a plant’s behaviour is a model of changes, which is related to appearance and development of an es and is determined in discrete time with a given step. the scenarios as a tool for management of complex systems belong to a class of so-called partial mathematical models, i.e. such models which include only the essential factors that can be formalized with an acceptable degree of accuracy. the main application of such models is a class of tasks that can be reduced to finding both optimistic and pessimistic assessments of key quantitative and qualitative parameters of the objects and the processes when a particular set of management decisions is implemented. in general, the scenario construction task may be stated as follows: a complex, dynamic, open, controlled, not fully observable system is given. describe the possible ways of change in several alternative directions thus to provide the most complete picture of the possible future states and paths of the studied system development. the process of scenario construction of industrial catastrophes and emergencies includes the following main stages [3]. (1) the possible moment of emergency and / or its consequences on the interval [0, t] is divided into discrete times ti with steps i. the period t describes the time during which the damage is liquidated and the emergency stops its further development, depending on the scale of the damage and the disaster. time instances correspond to control points of the es development when control actions aimed to improve the situation are implemented. the step in a scenario is determined on the basis of the condition of effective using of forces and means, as well as the preparation of the necessary resources for countermeasures to the development of emergency. (2) formation of the initial conditions (t = 0) and the conditions under which emergency occurs and develops and the evaluation of losses (damage). on this stage the development of es is simulated and the possible conditions of its course are described on the basis of initial and observed data. concretize targets of countering, assess their effectiveness, determine limitations on the course and consequences of emergency situations as well as resources support. define pessimistic preliminary assessment of 12 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach losses and damage of the disaster. develop measures to counter the emergency in different variants of its development. (3) organization and putting into operation forces and means necessary to eliminate causes and consequences of emergencies, rescue and other urgent works in industrial catastrophes, depending on scale, character and amount of damage. it describes the process of mobilization and deployment of forces and resources, providing materials and food, emergency medical care, including the preparation for the reception of additional operational units, vehicles of all types, hospitals deployment points for the rehabilitation of victims, etc. thus, at this stage briefly all the activities and resources are described as well as the forces and means necessary for an effective response to emergency situations. (4) formation of the current situation (at the time ti) and the conditions of emergency situations, as well as the adjustment of the expected damage assessment, i.e. at this stage the current range of activities is described (similar to p.2). the situation and the most important indicators obtained at the time ti-1 are taken as initial conditions. further the measures considering new conditions prevailing at a given moment are described. this process has to be described until the complete liquidation of the es. (5) verify the completeness and reality of the obtained scenario and its correction for maximum adequacy to the real development of the es. scenario is checked by experts and is included into education and training. (6) formation for a given scenario of the actions course and operational plans for countering an es. based on the developed scenario a course of actions is being defined (a graph of goals and tasks with logic ‘and’), preventive and operational action plans to reduce risk and damage in the form of many types of documents and descriptions (at the time ti) which are distributed across the entire organizational structure for the execution. thus, based on a scenario the general objective of the operation is formed, providing creation of the desired situation in the future. the formulated general objective is represented in the hierarchy of goals and called course of actions. course of actions is a graph of goals and tasks, formed as a result of the decision-making on a complete graph of alternatives. there are two ways to form general objectives with the appropriate hierarchy of goals and tasks, and therefore two ways to form a course of action: from all the alternative scenarios the most probable one is selected and the general goal and the course of actions are produced in accordance with it; for each alternative scenario its own general goal is built with its hierarchy of goals and tasks (the variant planning). however, in all cases for each scenario several alternative general goals and courses of actions are considered. a scenario formed this way allows to reflect the development of emergencies, develop a strategy of the organization and implementation of preventive and operational measures against emergencies, create strategic and tactical plans of actions to carry out a qualitative analysis of the impact, as well as to forecast data on the estimated losses and damage. 3. usage of scenario analysis in management of prevention and liquidation of emergencies the basic feature of use the scenario approach in the planning and management of the prevention and liquidation of consequences of man-made emergencies is the need to expand its capabilities to conduct advanced comprehensive analysis of the current situation. one of the most important features of man-made emergencies, as a research domain, is the sufficiently powerful normative base. its development and improvement are continuously expanding over a number of areas, such as industry, transport, fire, radiation, chemicals, energy, environmental, social, public safety, security of life activity, certification of potentially dangerous objects, medicine disaster etc. it allows to improve efficiency of the scenario analysis methodology in the management of the liquidation of consequences of emergency situations and to change traditional methodology and technology of research of the simulation models. advances in systems science and applications (2017) vol.17 no.1 13 at present the functional approach is widely used to study the problems of increasing efficiency of socio-economic systems. to study, for example, one or a few ‘close’ control functions is often not effective enough for solving the control tasks for the liquidation of consequences of man-made es. the main reason for this is the need for a broader, comprehensive analysis of the situation and taking into account the synergistic nature of adverse processes and phenomena, which significantly expands the boundaries of the study subject area and thus reduces the efficiency of the traditional approach. using the functional approach in these conditions creates objective difficulties for combining expert knowledge in various subject areas into a unified picture. this doesn't allow to carry out a comprehensive analysis of states and tendencies of the situation development. this is caused by the difference of professional languages (terminology) of experts, contradictory procedures and discrepancy of expert assessments of the results, as well the lack of common tools and mechanisms for pooling expertise. this leads to great time required to collect basic data and expert assessments under sufficiently rigid restrictions on the time of decision making. in addition, the traditional approach focuses on expert assessment of the current situation on a plant. this results in that a simulation model or at least its methodology basis is used. in this case, substantial expansion of domain area for a comprehensive analysis of the situation leads to difficulties with the assessment of the model’s adequacy, as well as with the validity of its borders and the level of detail. the presence of a general and developed normative base allows to fundamentally change the approach to formation of es patterns. it is more effective to form a basic model analyzing existing procedures and regulations, and further modify by the use of detailed information about the object’s specifics and the operational information about the development of the situation. this approach provides several quite apparent advantages: (7) the use of technical regulations and norms as an information base can significantly improve the adequacy of the developed multi-factor model, since it is based on reliable data about the object of research. the quality of such model is guaranteed by multiple checks in the development, coordination and approval of technical regulations and norms. (8) efficiency of diagnosis of sources of vulnerability in the control object is increased for different kinds of threats, as well as accuracy of risk assessment. the risk is characterized by a relatively low probability of es occurrence and by catastrophic consequences when it is realized. (9) the complexity of model development is substantially reduced, because it is done on the agreed and (at least partially) formalized documents that contain extensive, reliable and, most importantly, relevant information for the research scenario. (10) there are almost no serious problems of coordination and integration of expertise in the development of the base model, because, in fact, the most complex procedures of this type have been conducted in the development of regulations and their results can be directly used in developing of the model. (11) time spent on diagnosis and detailed analysis in the development of multi-factor model is reduced substantially, because the procedures of expert assessments, needed in the description of domain, carried out not on the full range of problems, and serve only to clarify the necessary details or analysis of emerging unexpected circumstances in the development of the situation. (12) using of normative base with a high level of detail gives possibility to conduct a study of a simulation model based on quantitative assessments and absolute scales. a simulation model allows to conduct experiments in real time and provides (1) increase of the validity of the generated scenarios of situation development, (2) precision of forecasts, formed on their basis, (3) as well as reliability of the evaluation of the effectiveness of management decisions. 14 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach (13) the most complex procedure of modification of a model is considerably simplified. it refers to an expansion and covering of adjacent domains and also allows to obtain complex strategic decisions, taking into account synergies in the development of es. (14) using a normative base allows to simplify solutions of many technological problems of scenario study for complex systems, associated with the laborious procedures of operative change of the models of a managed object in a limited time, taking into account the dynamics of change of key factors and parameters, as well as their relationships. (15) at last, perhaps the most important advantage of the widespread use of a normative base in the course of the scenario study is the possibility of an integrated approach to a solution of management tasks of prevention and liquidation of consequences of es, allowing at the same time to consider the related but different in nature phenomena and processes. this allows effectively integrate in a single model the factors and threats of terrorism, as well as fire, radiation, chemical, energy and environmental security. 4. scenario analysis of safety underground station in an emergency situation let us consider the task of scenario analysis of safety for potentially dangerous objects of urban transport infrastructure. the solution of this task is necessary in the process of planning and decision making on the prevention and liquidation of the causes and consequences of possible es. here we must briefly recall the main principles and approaches to the design and study of models of the situation development, which provide the implementation of management decisions [3-5]. the process of modeling and synthesis of alternative scenarios of the situation development is carried out using the apparatus of functional signed graphs. the mathematical model of these graphs is an extension of the classical models of oriented digraphs. in addition, the model of digraph g (x, e), where the x – a finite set of vertices, and e – a set of arcs of the graph includes additional components. so, each vertex of the graph has the parameters  xnivv i  , and each arc of the graph has a sign, weight or functional transformation f (v, e). as known, underground facilities of subways are the objects of high potential danger, that is related by a number of factors. the most important of them are:  a large number of passengers served by the subway;  high concentration of people inside trains, station premises and on inter-exchange transitions during peak hours;  the possibility of panic and need of evacuation of large numbers of passengers;  considerable depth of the tunnels and station facilities;  a limited number of inclined tunnels and vertical shafts, coming to the surface;  a large extent and a limited capacity of escape routes;  difficult maintenance of smooth and safe evacuation of passengers from a train stopped inside a tunnel;  presence of power grids under high voltage;  high-speed of smoke pollution of tunnels and space stations, the complexity of reconnaissance and control of fires;  the need for laying hose lines over long distances, taking into account the complexity of the layout and the availability of wagons;  the impact of ventilation on gas exchange of fire;  rapid increase in dangerous factors of fire values to a critical level;  threat of deformation and loss of structural elements carrying capacity station etc. in addition, metro stations and their environment are complex socio-technical systems, characterized by a significant amount of power, escalator, control-crossing, ventilation, climate, advances in systems science and applications (2017) vol.17 no.1 15 lighting and other kinds of process equipment; availability of premises of different categories for explosion and fire hazard, and fire resistance of load-bearing structures; a significant number of businesses for various purposes (kiosks, stalls, trade and promotional stands) to be placed in ground and underground hallways, in the transitions and on the platforms of underground stations, etc. consider the complex scenario models to predict the prevention planning and management, and liquidation of the causes and consequences of technological disasters and es in the underground facilities. as an example, we will use a hypothetical es initiated by a conventional explosion in a typical subway station of deep foundation, which moscow has more than 70. this type of station is most typical for the central administrative district (cad) of moscow, which will be considered in the following model example. when constructing model the risk zone is limited to only outside of the station. to ensure generality of the model we also exclude from consideration peak hours of the metro, which represent the greatest danger from the point of view of the consequences of es, as well as the busiest station and interchange nodes with the total daily load, considerably more than 50 million passengers (is known that the busiest station daily serves 100-150 thousand people). we assume that the number of people at the station at the time of a es corresponds to its average occupancy and is about 800 people. the designed simulation model consists of three parts: (16) the security model of the station and the people; (17) model of the work of the moscow city warning and es liquidation system; (18) model of the power and means involved for liquidation of es. the model can be further expanded to include information about the route to the nearest metro station firehouse (fh). in particular, one can consider a situation related to difficulties in extension of the given fh due to the congestion of highways or other reasons. also, data for the location of hospitals, in which the suffered people can come can additionally be used, as well as major socio-relevant and potentially dangerous objects, located near the considered metro station (if any). 4.1. the security model of station and people the model includes the following security issues: (19) the safety of people; (20) the safety of stay conditions; (21) mechanical safety; (22) fire safety. if necessary, the number of security issues may be increased. each of these aspects includes several related factors. 4.2. model of the work stages of the moscow city warning and liquidation of es system (mces). the main stages of the simulated operation of the district-level mces in the event of a es in the subway stations are as follows. (23) the first stage – the adoption of emergency measures:  alerting the population;  adoption of emergency measures to protect the public, victim assistance and localization of the accident;  arrival of the operational group (og) mces;  organization of a reconnaissance in the disaster zone;  notification about es and organization of an operational headquarter and the commission on prevention and liquidation of es and fire safety;  clarification of tasks for mces forces;  alerting forces of mces on the district level. (24) the second stage – operational planning of emergency. rescue and other urgent works:  collection of information for the decision-making mces; 16 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach  organization of work planning;  organization of interaction between the involved forces. (25) the third stage – liquidation of the es:  deployment of es staff, organization of communication with the management bodies operating in the area of the emergency and the parent body;  collection of executors and interacting of the management bodies deployed in the disaster area, clarifying the situation, the composition of forces, the action plan, the report for the elimination of es;  develop proposals of mces for operational decisions;  ensuring timely delivery of tasks to subordinates and interacting controls bodies;  organization of information exchange about the situation, measures taken to interacting, neighboring authorities. (26) the fourth stage – the implementation of measures on social protection and rehabilitation of areas affected by emergencies factors:  complete deactivation, decontamination and disinfection facilities;  repair of electrical, water, and gas utilities, communication lines, etc. in addition to these activities in the first stage rescue and other urgent works are carried out (asdnr), which include:  search of victims, the evacuation of the focus of infection;  conduct health suffered partial processing;  delivery of the victims at the place of loading on ambulances;  pre-hospital medical care, evacuation – transport sorting;  delivery of the victims to special medical institutions;  providing of skilled care. the structure of the model takes into account that the activities of the individual steps are directly related to each other. 4.3. model of forces and means involved for liquidation of emergency situations. the list of factors is based on the information about the presence of forces and means of liquidation of es, normative data about grouping mces forces in 10 minutes after the disaster, and includes:  technical services of subway;  squad of the subway protection;  operational groups of mces;  search and rescue team;  police units of internal affairs (uvd) cad;  units of road – patrol uvd cad;  public joint stock company for energy and electrification of moscow (mosenergo);  joint stock company ‘mosvodokanal’ (a company providing services in the sphere of water);  supply and sanitation in the city of moscow;  emergency gas service of moscow (mosgaz);  forces of emergency medical aid center;  forces of the specialized enterprises of fuel-energy sector of moscow operation communication;  collectors (moskollektor);  volunteer emergency rescue teams. additional factors in the model are the following peaks:  activation of forces and means;  arrival time of fire and rescue services; advances in systems science and applications (2017) vol.17 no.1 17  the number of people at the station. the main arcs in the model are circular, which means a continuous activity of forces and means. if necessary, the model can be developed, replacing the factors of new structures (for example, drive factor and hrc can be replaced with a model fire fighting, etc.). 4.4. joint model of the emergency liquidation. topology of unified situational model, which is the unification of the previously discussed models, is shown in figure 1. consider baseline scenarios obtained by simulation of an emergency at the facility underground. scenario 1: ‘the fire caused by an explosion and subsequent quenching’. at the first stage fire-fighting forces is done by underground workers (valid up to 20 steps). at first the human safety of the plant is growing. after 20 steps, gradually the all important characteristics become negative trend. then, starting with the 100 steps of action of fire brigades and other forces alter the situation in a favorable direction (figure 2). scenario 2: ‘the spread of panic in liquidation of emergencies’. we consider the situation in which the occurrence of emergency situations and fire causes panic among the passengers in the station. panic prevents the fire-rescue units and makes it difficult to search for survivors. this situation corresponds to the following modification of the original model:  increased weight feedback relationship ‘panic’ --> ‘access firefighters’ 10 times;  addition of a new arc ‘access firefighters’ --> (+1) --> ‘fire propagation’;  increased weight feedback relationship ‘panic’ --> ‘search of victims’ twice. dynamics of the major factors is presented in figure 3. fig.1. joint model of the emergency liquidation 18 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach fig.2. the fire caused by the explosion and subsequent quenching (scenario 1) fig.3. spread of panic in liquidation of emergencies (scenario 2) advances in systems science and applications (2017) vol.17 no.1 19 initially, there is a decrease in the safety of people and station structures. outbreaks of fire cannot be located and the fire spread process continues. significant interference to the undertaken measures to extinguish the fire are the presence of people in the station, and not to reduce panic. a further negative development of the situation can lead to the appearance of the victims of the man-made disaster. over time, due to held fire-rescue actions the fire safety factor first stabilized. it reduces the risk of destruction of bearing structures of plant facilities while reducing the spread of fire intensity. accelerate the arrival of additional fire-rescue services accelerates the growth of the security of stations (in terms of the developed model), however, that is especially important, given the operational situation is after the onset of the fire victims. scenario 3: ‘efficiency of the actions the fire-rescue service on liquidation of emergency’. under this scenario, the variant of the situation, in which the strong and well-coordinated actions of rescuers help reduce the panic, that is reflected in the following modification of the model:  added a negative relationship ‘the effectiveness of the fire and rescue services’ --> ‘panic’;  added a nonlinear relationship ‘panic’ --> ‘the efficiency of the fire and rescue services’ (link is positive if there is the panic growth and negative otherwise). the simulation results are presented in figure 4. as seen from represented in figure 4 graphics dependencies, shortly after the arrival of the additional fire and rescue services the panic reduced. this increases the efficiency of fire suppression and leads to subsequent growth of human safety. scenario 4: ‘modeling of losses during working escalators’. the aim of this scenario is an analysis of the casualties among the passengers of the station as a result of a disaster. the basis for the building a model are the characteristics of the underground station. carry out temporary normalization of modeling steps as follows: 100 steps = 10 minutes, that is 1 step = 6 sec. accordingly, the fire and rescue services come into place on the emergency regulations in 10 minutes, i.e. after 100 steps. fig.4. efficiency actions the fire rescue services on liquidation of emergencies (scenario 3) 20 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach the activity of metro services unit for the protection of underground rescue people is 20 steps = 2 min. in the period from 20 to 100 steps (4.8 min.) the assistance in rescuing on their own is limited. as before, it is assumed that metro stations are around 800 people. this corresponds to an average workload of the station. assume that the escalators on all saving time is not damaged, and the only way to save. the speed of the escalator is the motion on the rise – 3.0 thousand people / hour. thus, the rate of evacuation is 5 people / sec. it should also be noted that the speed of evacuation decreases with a decrease in the safety of people. moreover, the evacuation of injured people is slower than healthy people. the evacuation of everyone injured takes twice more time. every 6 sec. (step 1) add one injured. under these initial conditions the following modification of the model is conducted: (27) new factors are added:  ‘the number of people at the metro station’,  ‘evacuation speed’,  ‘the rate of destruction of people’,  ‘the number of evacuees’,  ‘the number of injured’,  ‘the number of rescued without injuries’. (28) new relationships are added:  ‘evacuation speed’ --> (-1) --> ‘number of people at the station’,  ‘evacuation speed’ --> (+1) --> ‘number of evacuees’,  ‘evacuation speed’ --> (+1) --> ‘number of rescued without injuries’,  ‘destruction speed’ --> (+1) --> ‘number of affected’,  ‘destruction speed’ --> (-1) --> ‘number of rescued without injuries’,  ‘the rate of evacuation’ --> (5, if the safety of the people is increasing 2 or otherwise) --> ‘evacuation speed’ (circular arc). (29) the conditions of factors activity:  ‘evacuation speed’ is active, if the number of people at the station is greater than zero,  ‘destruction speed’ is active, if the number of people at the station is greater than zero. (30) initial impulses:  --> ‘evacuation rate’ +1;  --> ‘destruction speed’ +1. the simulation results are presented in figure 5 (qualitative results) and figure 6 (quantitative results). the simulation results show that as a result of works on es liquidation 800 people are evacuated, 271 is injured (get damaged), and 529 people escaped without injury. scenario 5: ‘modeling of losses during non-working escalator’. the modification of the previous model: the speed of evacuation is 3 person / sec. the rate of evacuation of injured persons 2 / sec. evacuation speed cannot be negative. qualitative and quantitative simulation results are shown in figure 7 and figure 8, respectively. as a result of work on es liquidation on the results of the simulation of the 800 evacuated people were injured 447 people, 353 people escaped without injury. the results were almost twice as bad. furthermore, the evacuation time increased to 12 minutes. advances in systems science and applications (2017) vol.17 no.1 21 fig.5. modeling of losses during working escalators (scenario 4, the qualitative results) fig.6. modeling of losses during working escalators (scenario 4, the quantitative results) 22 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach fig.7. simulation of losses during nonworking escalators (scenario 5, the qualitative results) fig.8. modeling of losses during nonworking escalators (scenario 5, the quantitative results) advances in systems science and applications (2017) vol.17 no.1 23 5. conclusions the main feature discussed in this paper, a class of industrial safety management tasks based on the scenario approach is to have developed and the general normative framework that allows to fundamentally change the approach to building models of emergency. studies have shown that in this case the procedure of forming the base model using a comprehensive analysis of existing rules and norms were quite effective. in the second stage – modifications of the model taking into account the detailed information about the specifics of the object of study and the incoming operational information about the nature of the emergency. using the normative base in the process of forming the simulation model allowed in the generation of alternative scenarios of situation development to base not only on qualitative but also on quantitative indices. this improves the accuracy and reliability of the scenario analysis results, and the quality and efficiency of decisions made in the process of liquidation the es consequences. the developed methodology of advanced scenario analysis allows to solve a wide range of process management tasks of prevention and liquidation of consequences of technogenic catastrophes and emergencies. the proposed methods provide the complex multidimensional study of alternatives of an emergency at the given target criterion under conditions of uncertainty. the main advantage of this approach is the ability to predict the behavior of the simulated object through the formation of alternative development scenarios. this approach allows to make conclusions about the most likely and appropriate directions of dynamic processes development, their stability, and other significant characteristics on the basis of information about the structural features of the study domain. the practical application of scenario approach enables a comprehensive analysis of the current situation at a given time horizon, to form the short-term and long-term forecasts of its development and to counter the emerging threats, to assess the effectiveness and consistency of distributed in time and space strategic and tactical management decisions for the prevention and elimination of consequences of emergencies and man-made disasters and catastrophes. modeling and scenario analysis of a wide range of safety of potentially dangerous for transport and infrastructure, residential and industrial buildings, etc. it comprises:  diagnostic analysis and evaluation of the situation;  the development of the object model, the choice of performance criteria and an assessment of their relative importance;  generation of possible development scenarios;  assess the scenarios developed (primarily management processes of prevention and liquidation of consequences of es), and selecting the best of them for a given criterion of efficiency  continuous analysis of monitoring information on the situation and making the appropriate changes in the structure of the models on the basis of the data obtained;  assessment and selection of management actions on liquidation of emergency situations and improve safety;  dynamic analysis of the possible consequences of control actions;  collection of the results of the implementation of scenarios and assessment of data. the opportunity to develop the methodology of scenario analysis and modeling is certainly not limited to the examples given in this paper – they are much wider. currently, one of the most promising directions of development of the proposed methodology is unification of planning and scenario modeling. this area, which requires a separate study, involves the selection, analysis and certification of standard subclasses of potentially dangerous objects on the basis approved by the ministry of the russian federation for civil defense, emergencies and liquidation of consequences of natural disasters (emercom of russia) classification, which includes five basic classes, technological accidents which may are the source of federal (cross-border), regional, territorial and local emergencies respectively. unification (the reduction of all the variety of plans, control and monitoring actions to quite a limited set) should enable the development and use of standard tools and resources planning and 24 v.l. schultz et.al.: regional security: analysis of the emergency management effectiveness based on the scenario approach management processes of prevention and liquidation of consequences of es, as well as a large community of semantic and information content of the standard (basic) plans. on the one hand, it provides the ability to use quite limited set of information elements, and on the other hand – a significant invariance for the different levels of government and departmental affiliation, which should simplify the implementation of agreed plans of the system in various modes. development of theoretical and applied research in this field will provide an opportunity to solve a wide range of practical problems of planning and management processes of prevention and liquidation of consequences of emergency situations at the facility, municipal, regional and federal levels. references [1] arkhipova n., kulba v. emergency management. moscow, russian state humanitarian university,1998 (in russian). [2] arkhipova n., kulba v., kosyachenko s., chanhieva f., shelkov a. organizational management. moscow, russian state humanitarian university, 2007. (in russian). [3] schultz v., kulba v., kononov d., kosyachenko s., shelkov a., chernov i. models and methods of analysis and synthesis of scenarios of socio economic systems. moscow, nauka, 2012. (in russian). [4] schultz, v., kulba v., shelkov a., chernov i. methods of the planning and management of technogenic safety based on the scenario approach. national security / nota bene, 2013, vol.2, no. 25, p.198-216 (in russian). [5] kulba, v., zaikin o., shelkov a., chernov i. scenario analysis in the management of regional security and social stability. in “new frontiers in information and production systems”, eds. p. rozewski, d. novikov, n. bakhtadze and o. zaikin. intelligent systems reference library, switzerland: springer international publishing, 2016, vol. 98., p.249–268. adv syst sci appl 2018; 4:39–51 http://ijassa.ipu.ru/index.php/ijassa/article/view/572 estimation of stress-strength reliability for quasi lindley distribution m.m. mohie el-din1, a. sadek 1∗, shaimaa h. elmeghawry2 1 dep. of mathematics, faculty of science, al-azhar university, egypt 2 faculty of engineering, benha university, egypt received april 14, 2018; revised october 17, 2018; published december 31, 2018 abstract: this paper discussed the problem of estimating of the stress-strength reliability r = pr(y < x). it is assumed that the strength of a system x, and the environmental stress applied on it y, follow the quasi lindley distribution(qld). stress-strength reliability is studied using the maximum likelihood, and bayes estimations. asymptotic confidence interval for reliability is obtained. bayesian estimations were proposed using two different methods: importance sampling technique, and mcmc technique via metropolis-hastings algorithm, under symmetric loss function (squared error) and asymmetric loss functions (linex, general entropy). the behaviors of the maximum likelihood and bayes estimators of stress-strength reliability have been studied through the monte carlo simulation study. finally analysis of a real data set has also been presented. keywords: quasi lindley distribution; stress-strength reliability; maximum likelihood estimation; asymptotic confidence interval; bayesian estimation; importance sampling technique; mcmc technique via metropolis-hastings algorithm. 1. introduction the stress-strength models have been widely used for reliability design of systems. in these models the reliability is defined as the probability that the strength is larger than the stress r = pr(y < x). as the strength x is larger than the stress y, the system work efficiently, otherwise the the system fails. estimation of stress-strength reliability was studied by several authors, for example, stress-strength model and its generalizations has been discussed in [10]. the estimation of r when x and y are normally distributed was introduced by church and harris [7]. krishnamoorthy et al. [11] introduced an inference on reliability in two-parameter exponential stress-strength model. al-mutairi et al. [1, 2] presented the stress-strength reliability for lindley and weighted lindley distributions respectively. stress-strength reliability estimation for generalized lindley distribution has been introduced by singh et al. [18]. recentely khan and jan [9] studied the estimation of stress-strength reliability model using finite mixture of two parameter lindley distributions. this paper is focused upon studying the problem of the estimation of the stress-strength reliability for the quasi lindley distribution (qld) introduced by shanker et al. [17] of which the lindley distribution is a particular case. we estimated the parameter of the stressstrength reliability r using the maximum likelihood, and bayesian estimation methods. we ∗corresponding author: a sadek@azhar.edu.eg 40 m. mohie el-din, a. sadek, shaimaa elmeghawry construct the asymptotic confidence interval of r based on the asymptotic distribution of the mle of r. in bayesian estimation we introduced two sampling methods (importance sampling and mcmc). the qld has the following probability density function (pdf): f(x) = θ α + 1 (α + θx) e−θx , (1.1) and the following cumulative distribution function (cdf): f (x) = 1− [(1 + α + θx) α + 1 e−θx ] , (1.2) where; x > 0, θ > 0, α > −1. this paper is organized as follows. in section 2, stress-strength reliability issue is studied to obtain the reliability function of the parameters of qld distribution. maximum likelihood estimation for stress-strength reliability is discussed in section 3. asympototic confidence interval of reliability is proposed in section 4. in section 5, a general procedure for deriving the bayesian estimator of reliability is introduced, wherein we applied the importance sampling and mcmc techniques to compute the approximation of this estimator. section 6 presented simulation study to investigate and compare the performance of each method of estimation. section 7 presented analysis of a real data set for illustrative purposes. finally, conclusions appear in section 7. 2. stress-strength reliability assume x ∼ qld(θ1, α1) and y ∼ qld(θ2, α2) are independent random variables with pdf f(x) and g(y), respectively. then the stress strength reliability can be obtained as: r = pr(y < x). = ∫ ∞ 0 ∫ x 0 f(x)g(y) dydx. = ∫ ∞ 0 f(x)g(x) dx. = ∫ ∞ 0 θ1(α1 + θ1x) α1 + 1 e−θ1x [ 1− [ (1 + α2 + θ2x) α2 + 1 e−θ2x ]] dx. = ∫ ∞ 0 θ1(α1 + θ1x) α1 + 1 e−θ1x − ∫ ∞ 0 θ1(α1 + θ1x)(1 + α2 + θ2x) e−(θ1+θ2)x (α1 + 1)(α2 + 1) dx. = 1− θ1 ( 2θ1θ2 + (θ1 + θ2) (α2θ1 + α1θ2 + θ1) + α1 (α2 + 1) (θ1 + θ2) 2) (α1 + 1) (α2 + 1) (θ1 + θ2) 3 . (2.3) from eq.(2.3), we noticed that r is a function of parameters θ = (θ1, α1, θ2, α2). therefore, for maximum likelihood estimate (mle) of r we need to obtain the mles of these parameters. copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 41 3. maximum likelihood estimation for reliability suppose that x1, x2, · · · , xn is random sample from qld(θ1, α1), and y1, y2, · · · , ym is random sample from qld(θ2, α2), then the log likelihood function can be written as: logl(x, y; θ) = n log (θ1) +m log (θ2)− n log (α1 + 1)−m log (α2 + 1) −θ1 n∑ i=1 xi − θ2 m∑ j=1 yj + n∑ i=1 log (α1 + θ1xi) + m∑ j=1 log (α2 + θ2yj) . (3.4) the mle of θ = (θ1, α1, θ2, α2) can be obtained as a solution of the following equations: ∂l ∂θ1 = n θ1 − n∑ i=1 xi + n∑ i=1 xi α1 + θ1xi = 0, ∂l ∂θ2 = m θ2 − m∑ j=1 yj + m∑ j=1 yj α2 + θ2yj = 0, ∂l ∂α1 = − n α1 + 1 + n∑ i=1 1 α1 + θ1xi = 0, and ∂l ∂α2 = − m α2 + 1 + m∑ j=1 1 α2 + θ2yj = 0. solving these equations numerically using an iterative process as newton raphson to get θ̂1, α̂1, θ̂2, α̂2, then the mle of r can be obtained as following: r̂ = 1− θ̂1 ( 2θ̂1θ̂2 + ( θ̂1 + θ̂2 )( α̂2θ̂1 + α̂1θ̂2 + θ̂1 ) + α̂1 (α̂2 + 1) ( θ̂1 + θ̂2 ) 2 ) (α̂1 + 1) (α̂2 + 1) ( θ̂1 + θ̂2 )3 . (3.5) 4. asymptotic confidence interval of r the asymptotic variance-covariance matrix of all parameters can be approximated by the inverse of observed information matrix, and then derive the asymptotic distribution of r̂. based on the asymptotic distribution of r̂, we obtain the asymptotic confidence interval of r. the fisher information matrix of θ = (θ1, α1, θ2, α2) is given as: i(θ) = −  e( ∂ 2l ∂θ1 2 ) e( ∂2l ∂θ1∂α1 ) e( ∂2l ∂θ1∂θ2 ) e( ∂2l ∂θ1∂α2 ) e( ∂2l ∂α1∂θ1 ) e( ∂ 2l ∂α1 2 ) e( ∂2l ∂α1∂θ2 ) e( ∂2l ∂α1∂α2 ) e( ∂2l ∂θ2∂θ1 ) e( ∂2l ∂θ2∂α1 ) e( ∂ 2l ∂θ2 2 ) e( ∂2l ∂θ2∂α2 ) e( ∂2l ∂α2∂θ1 ) e( ∂2l ∂α2∂α1 ) e( ∂2l ∂α2∂θ2 ) e( ∂ 2l ∂α2 2 )  . = i11 i12 i13 i14 i21 i22 i23 i24 i31 i32 i33 i34 i41 i42 i43 i44  . copyright c© 2018 assa. adv syst sci appl (2018) 42 m. mohie el-din, a. sadek, shaimaa elmeghawry where: i13 = i31 = 0; i14 = i41 = 0, i23 = i32 = 0; i24 = i42 = 0, i11 = − n θ1 2 − n∑ i=1 xi 2 (α1 + θ1xi)2 , i22 = n∑ i=1 1 (α1 + θ1xi)2 − n (α1 + 1)2 , i33 = −m θ2 2 − m∑ j=1 yj 2 (α2 + θ2yj)2 , i44 = m∑ j=1 1 (α2 + θ2yj)2 − m (α2 + 1)2 , i12 = i21 = n∑ i=1 xi (α1 + θ1xi)2 , and i34 = i43 = − m∑ j=1 yj (α2 + θ2yj)2 . using the central limit theorem, we obtain the following theorem : theorem 1: as n→∞, m→∞; then ( √ n(θ̂1 − θ1), √ n(α̂1 − α1), √ m(θ̂2 − θ2), √ m(α̂2 − α2)) d→ n(0, i−1(θ)). where d→ means converge in distribution, and i−1(θ) is the inverse of the matrix i(θ). in order to establish the asymptotic normality of r, we first define: d(θ) = ( ∂r ∂θ1 , ∂r ∂α1 , ∂r ∂θ2 , ∂r ∂α2 )t = (d1, d2, d3, d4) t , where t is transpose operation, and d1 = −θ1(2α1(α2 + 1)(θ1 + θ2) + α1θ2 + (α2 + 1)(θ1 + θ2) + α2θ1 + θ1 + 2θ2) (α1 + 1)(α2 + 1)(θ1 + θ2)3 + 3θ1 (α1(α2 + 1)(θ1 + θ2) 2 + (θ1 + θ2)(α1θ2 + α2θ1 + θ1) + 2θ1θ2) (α1 + 1)(α2 + 1)(θ1 + θ2)4 −α1(α2 + 1)(θ1 + θ2) 2 + (θ1 + θ2)(α1θ2 + α2θ1 + θ1) + 2θ1θ2 (α1 + 1)(α2 + 1)(θ1 + θ2)3 , copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 43 d2 = θ1 (α1(α2 + 1)(θ1 + θ2) 2 + (θ1 + θ2)(α1θ2 + α2θ1 + θ1) + 2θ1θ2) (α1 + 1)2(α2 + 1)(θ1 + θ2)3 −θ1 ((α2 + 1)(θ1 + θ2) 2 + θ2(θ1 + θ2)) (α1 + 1)(α2 + 1)(θ1 + θ2)3 , d3 = 3θ1 (α1(α2 + 1)(θ1 + θ2) 2 + (θ1 + θ2)(α1θ2 + α2θ1 + θ1) + 2θ1θ2) (α1 + 1)(α2 + 1)(θ1 + θ2)4 −θ1(2α1(α2 + 1)(θ1 + θ2) + α1(θ1 + θ2) + α1θ2 + α2θ1 + 3θ1) (α1 + 1)(α2 + 1)(θ1 + θ2)3 , and d4 = θ1 (α1(α2 + 1)(θ1 + θ2) 2 + (θ1 + θ2)(α1θ2 + α2θ1 + θ1) + 2θ1θ2) (α1 + 1)(α2 + 1)2(θ1 + θ2)3 −θ1 (α1(θ1 + θ2) 2 + θ1(θ1 + θ2)) (α1 + 1)(α2 + 1)(θ1 + θ2)3 . hence; using theorem 1, the asymptotic distribution of r̂, the mle of r is defined as √ n+m(r̂−r) d→ n(0, b), where b = v ar(r̂) = dt (θ)i−1(θ)d(θ). (4.6) therefore, using eq.(4.6), an asymptotic 100(1− α)% confidence interval for r can be obtained as: r̂± zα 2 √ v ar(r̂), where zα 2 is the upper α 2 precentile of the standard normal distribution. 5. bayesian estimation in this section, we provide the bayes estimate of r where θ1, θ2, α1, α2 are unknown parameters and all of these parameters having independent gamma prior distributions as following: π(θ1) ∼ gamma(a1, b1), π(θ2) ∼ gamma(a2, b2), π(α1) ∼ gamma(a3, b3), and π(α2) ∼ gamma(a4, b4). copyright c© 2018 assa. adv syst sci appl (2018) 44 m. mohie el-din, a. sadek, shaimaa elmeghawry the joint posterior pdf is defined as g(θ1, θ2, α1, α2/data) = l(x, y/θ1, θ2, α1, α2)π(θ1)π(θ2)π(α1)π(α2)∫∞ 0 ∫∞ 0 ∫∞ 0 ∫∞ 0 l(x, y/θ1, θ2, α1, α2)π(θ1)π(θ2)π(α1)π(α2)dθ1dθ2dα1dα2 . then g(θ1, θ2, α1, α2/data) ∝ θ1 n (α1 + 1)n θ2 m (α2 + 1)m n∏ i=1 (α1 + θ1xi) m∏ j=1 (α2 + θ2yj) e −θ1 ∑n i=1 xi × e−θ2 ∑m j=1 yj θ1 a1−1e−b1θ1θ2 a2−1e−b2θ2α1 a3−1e−b3α1α2 a4−1e−b4α2 . (5.7) 5.1. bayes estimators under symmetric and asymmetric loss function: the bayes estimate of reliability formula depending on the choice of the loss function. two different loss functions are used, symmetric and asymmetric loss function. if the amount of loss assigned by a loss function to a positive error is equal to the negative error of the same magnitude, then the loss function is called a symmetric loss function. in most of the studies on estimation and prediction problems, authors prefer to use the squared error loss function which is symmetric in nature. however, the use of the squared error loss function is not appropriate particularly in these cases, where the losses are not symmetric. thus in order to make the statistical inferences more practical and applicable, we often needs to choose an asymmetric loss function. a number of asymmetric loss functions proposed for use, first is the linex loss function which suggested by varian [20], and studied by several others including basu and ebrahimi [3], zellner [21]. second is the general entropy loss function which introduced by calabria and pulcini [5]. these asymmetric loss functions are also studied by braess and dette [4], pandey and rao [14], parsian and kirmani [15], sanku dey [16] and soliman [19], who have used these loss function in different estimation and prediction problem. then the equation of the bayes estimate of the reliability is depend on the loss function, here is the general equation for each type: -the bayes estimate under the squared error loss function, which is the posterior mean of r, is given by: r̂se = ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 r g(θ1, θ2, α1, α2/data)dθ1dθ2dα1dα2. -the bayes estimate under the linex loss function is given by: r̂lx = −1 c ln [ ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 e−cr g(θ1, θ2, α1, α2/data)dθ1dθ2dα1dα2 ] , where c is constant, c > 0, see [21]. -the bayes estimate under the general entropy loss function is given by: r̂ge = [ ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 ∫ ∞ 0 r−q g(θ1, θ2, α1, α2/data)dθ1dθ2dα1dα2 ]−1/q , where q is constant, q > 0, see [5]. it is impossible to compute these integrals analytically. two approaches can be used to approximate these integrals, namely, importance sampling technique and mcmc technique. copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 45 5.2. importance sampling technique importance sampling technique has suggested by chen and shao [6]. in statistics, importance sampling is the name for the general technique of determining the properties of a distribution by drawing samples from another distribution. the focus of importance sampling here is to determine as easily and accurately as possible the properties of the posterior from a representative sample from the second distribution. using importance sampling technique, eq.(5.7) can be written as g(θ1, θ2, α1, α2/data) ∝ g1(θ1/data)g2(θ2/data)g3(α1/data)g4(α2/data)h(θ1, θ2, α1, α2/data), where: g1(θ1/data) ∝ gamma(n+ a1, b1 + n∑ i=1 xi), g2(θ2/data) ∝ gamma(m+ a2, b2 + m∑ j=1 yj), g3(α1/data) ∝ gamma(a3, b3), g4(α2/data) ∝ gamma(a4, b4), and h(θ1, θ2, α1, α2/data) = ∏n i=1 (α1 + θ1xi) ∏m j=1 (α2 + θ2yj) (α1 + 1)n (α2 + 1)m . as shown, all the above functions from g1(θ1/data) to g4(α2/data) follow gamma distributions with different parameters, so it is quite simple to generate qld parameters from them. assuming that a1, · · · , a4 and b1, · · · , b4 are known, and assuming initial values for θ1, θ2, α1, α2. we can use the following importance sampling algorithm: • step1: generate θ11 from g1(θ1/data). • step2: generate θ21 from g2(θ2/data). • step3: generate α11 from g3(α1/data). • step4: generate α21 from g4(α2/data). • step5: repeat steps from 1 to 4, n times to obtain the vector (θ11, θ21, α11, α21), · · · , (θ1n , θ2n , α1n , α2n). then -an approximate bayes estimate of r under squared error loss function can be obtained as r̃impse = ∑n i=1ri h(θ1i, θ2i, α1i, α2i/data)∑n i=1 h(θ1i, θ2i, α1i, α2i/data) , an approximate bayes estimate of r under linex loss function can be obtained as r̃implx = −1 c log [ ∑n i=1 e −cri h(θ1i, θ2i, α1i, α2i/data)∑n i=1 h(θ1i, θ2i, α1i, α2i/data) ], -an approximate bayes estimate of r under general entropy loss function can be obtained as: r̃impge = [ ∑n i=1ri −q h(θ1i, θ2i, α1i, α2i/data)∑n i=1 h(θ1i, θ2i, α1i, α2i/data) ]−1/q, where ri = r(θ1i, θ2i, α1i, α2i). as defined in eq.(2.3), for i = 1, · · · , n . copyright c© 2018 assa. adv syst sci appl (2018) 46 m. mohie el-din, a. sadek, shaimaa elmeghawry 5.3. mcmc technique the most general mcmc algorithm is the metropolis-hastings (mh) algorithm, which was originally introduced by metropolis et al. [13], and hastings [8]. the metropolis-hastings (mh) algorithm simulates samples from a probability distribution by making use of the full joint density function and (independent) proposal distributions for each of the variables of interest. the joint posterior density function of θ1, θ2, α1, and α2 is given in eq.(5.7). it is easily seen that the posterior density functions of θ1, θ2, α1, and α2 are, respectively: π1(θ1/data) ∝ gamma ( n+ a1, b1 + n∑ i=1 xi ) , (5.8) π2(θ2/data) ∝ gamma ( m+ a2, b2 + m∑ j=1 yj ) , (5.9) π3(α1/θ1, data) ∝ α1 a3−1e−b3α1 ∏n i=1 (α1 + θ1xi) (α1 + 1)n , (5.10) and π4(α2/θ2, data) ∝ α2 a4−1e−b4α2 ∏m j=1 (α2 + θ2yj) (α2 + 1)m . (5.11) therefore, easily samples of θ1 and θ2 can be generated by using gamma distribution as shown in eqs.(5.8), and (5.9) respectively. however, the posterior distribution of α1 , α2 cannot be generated from a well known distributions. the metropolis-hastings algorithm, can be used to solve this problem, as shown in the following algorithm. • step1: start with initial value of α1,α2 such that α1 (0) = α̂1, and α2 (0) = α̂2. • step2: set i = 1. • step3: generate θ1(i) from π1(θ1/data). • step4: generate θ2(i) from π2(θ2/data). • step5: generate α1 (i) from π3(α1/θ1, data) using the metropolis-hastings algorithm with the proposal distribution q1 as following: – generate α1 (∗) from the proposal distribution q1 = n(α1 (i−1), v ar(α1 (i−1))). – calculate the acceptance probability r1(α1 (i−1), α1 (∗)) = min[1, π3(α1 (∗)/θ1 (i),data) π3(α1 (i−1)θ1 (i),data) ]. – generate u from uniform(0, 1). – if u ≤ r1(α1 (i−1), α1 (∗)), accept the proposal distribution and set α1 (i) = α1 (∗) , otherwise set α1 (i) = α1 (i−1). • step6: generate α2 (i) from π4(α2/θ2, data) using the metropolis-hastings algorithm with the proposal distribution q2 as following: copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 47 – generate α2 (∗) from the proposal distribution q2 = n(α2 (i−1), v ar(α2 (i−1))). – calculate the acceptance probability r2(α2 (i−1), α2 (∗)) = min[1, π4(α2 (∗)/θ2 (i),data) π4(α2 (i−1)θ2 (i),data) ]. – generate u from uniform(0, 1). – if u ≤ r2(α2 (i−1), α2 (∗)), accept the proposal distribution and set α2 (i) = α2 (∗) , otherwise set α2 (i) = α2 (i−1). • step7: compute r(i) at (θ1 (i), θ2 (i), α1 (i), α2 (i)) using eq.(2.3). • step8: set i = i+ 1. • step9: repeat steps from (3− 8) n times. then; -an approximate bayes estimate of r under squared error loss function is given as: r̃mhse = 1 n −m n∑ i=m+1 r(i). -an approximate bayes estimate of r under linex loss function is given as: r̃mhlx = −1 c log [ 1 n −m n∑ i=m+1 e−cr (i) ] . -an approximate bayes estimate of r under general entropy loss function is given as: r̃mhge = [ 1 n −m n∑ i=m+1 ( r(i) )−q]−1/q . where m is the burn-in units, n is the mcmc samples. 6. simulation study in this section, we mainly present some simulation experiments to see the performance of the mentioned methods for different sample sizes, (n,m) = (10, 10), (20, 20), (30, 30), (50, 50), (70, 70), (100, 100). we simulated 1000 complete samples from quasi lindely distribution with the parameter values; θ1 = 0.2, θ2 = 1.5, α1 = 2, α2 = 0.8 with true reliability value is 0.87399. we also compute the 95% confidence intervals of r based on the observed fisher information matrix. we compared the performances of the mle and the bayes estimates in terms of mean squared errors (mse’s). also two different techniques of bayesian estimation (importance, mcmc) are compared for different loss error functions. bayesian estimation for different loss error functions was proposed with many values of c, q such that; c1 = −3(lx1), c2 = 5(lx2), q1 = −3 (ge1), q2 = 5 (ge2). bayesian estimation studied under the informative gamma priors. for choosing suitable hyper-parameters, the experimenters can incorporate their prior guess in terms of location and precision for the parameter of interest. the gamma distribution for the priors has mean = a/b, and variance = a/b2. we assume a small value of prior variance (0.01), and take the mean equal to the true value of the parameter of interest. for each parameter prior we solve the two equations of the mean and the variance, we obtain the following values of copyright c© 2018 assa. adv syst sci appl (2018) 48 m. mohie el-din, a. sadek, shaimaa elmeghawry hyper-parameters : a1 = 4, a2 = 225, a3 = 400, a4 = 64, and b1 = 20, b2 = 150, b3 = 200, b4 = 80. we also computed the bayes estimates based on 11000 samples and discard the first 1000 values as burn-in. the maximum likelihood estimator and asymptotic confidence intervals of r for different (n,m) are obtained in table 6.1. bayes estimates of r using different techniques under different loss error functions are obtained in table 6.2. therefore, from this study of the simulation results we observed that: • the performance of the bayes estimators is better than maximum likelihood for all different sample sizes. • mean squared error(mse’s) for all estimation methods decreased as sample size increased. • as sample size increased, the asymptotic confidence intervals for r are improving, and their lengths are decreasing. that means the estimated reliability becomes in the most accurate interval. • when the sample size increased, both bayesian and maximum likelihood results become close to each other. • for bayes estimators, importance sampling technique gives less mse’s values, so it is better than mcmc technique for the same priors values, and same number of generated samples. • general entropy, and linex loss error functions gave less mse’s at specified values of c, q. as shown lx2, ge2 acheived the best results for mcmc, but for importance sampling technique lx1, ge1 are the best methods . table 6.1. average estimate (mean squared error) for mle, and average confidence length of the simulated 95%confidence intervals of r. all mse values are multiplied by 10−3 estimator mle c.i.l c.i.u c.i. length (10,10) 0.866095 0.734952 0.997237 0.262 (3.2129) (20,20) 0.878225 0.787222 0.969227 0.182 (1.1262) (30,30) 0.874831 0.797586 0.952077 0.155 (0.7609) (50,50) 0.875794 0.816741 0.93706 0.120 (0.459333) (70,70) 0.875149 0.823703 0.926595 0.103 (0.3349) (100,100) 0.874276 0.834959 0.919893 0.084 (0.1554) 7. real data analysis in this section we present the analysis of real data, introduced by singh et al. [18]. the data represent the waiting times (in minutes) before customer service of two banks a and b, respectively. the use of lindley distribution for the waiting times (bank a) data has been originally discussed by lindley [12]. since then, many authors have suggested the data under different set-up for lindley distribution. we are interested in estimating the stress-strength copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 49 table 6.2. average estimates(mean squared error ) of r for different bayes estimators under different error loss functions. all mse values are multiplied by 10−3 est. importance sampling mcmc technique (m,n) se lx1 lx2 ge1 ge2 se lx1 lx2 ge1 ge2 (10,10) 0.86713 0.86862 0.86457 0.86829 0.86346 0.88787 0.88908 0.88577 0.8888 0.88494 (0.7516) (0.7058) (0.8421) (0.7134) (0.8934) (0.7611) (0.7748) (0.7459) (0.77) (0.7477) (20,20) 0.87063 0.87147 0.86919 0.87128 0.86861 0.89199 0.89268 0.89081 0.89251 0.89038 (0.4769) (0.4611) (0.5075) (0.4637) (0.5239) (0.6747) (0.692) (0.6479) (0.6870) (0.6403) (30,30) 0.87073 0.87135 0.86967 0.87121 0.86925 0.89157 0.89208 0.89071 0.89196 0.8904 (0.3591) (0.3491) (0.3784) (0.3508) (0.3883) (0.5684) (0.5822) (0.5464) (0.5785) (0.5396) (50,50) 0.87168 0.87206 0.87104 0.87198 0.87079 0.88997 0.89032 0.88938 0.89023 0.88916 (0.2536) (0.2486) (0.2626) (0.2496) (0.2670) (0.4366) (0.4458) (0.4216) (0.4433) (0.4167) (70,70) 0.873434 0.873697 0.872995 0.873636 0.872827 0.88874 0.889014 0.88828 0.888946 0.888114 (0.2189) (0.217) (0.2226) (0.2173) (0.2244) (0.3755) (0.3823) (0.3645) (0.3805) (0.3609) (100,100) 0.875249 0.875418 0.874965 0.875378 0.874857 0.885949 0.886165 0.885586 0.886112 0.885455 (0.1384) (0.1387) (0.138) (0.1386) (0.138) (0.2432) (0.2477) (0.2359) (0.2465) (0.2334) parameter r = p (y < x) where x and y denotes the customer service time in bank a and b (data set 1, 2) respectively. the data sets are presented below: data set 1: x (n=100) 0.8, 0.8, 1.3, 1.5, 1.8, 1.9, 1.9, 2.1, 2.6, 2.7, 2.9, 3.1, 3.2,3.3, 3.5, 3.6,4.0, 4.1, 4.2, 4.2, 4.3, 4.3, 4.4, 4.4, 4.6, 4.7, 4.7, 4.8, 4.9, 4.9, 5.0, 5.3, 5.5, 5.7, 5.7, 6.1, 6.2, 6.2, 6.2, 6.3, 6.7, 6.9, 7.1, 7.1, 7.1, 7.1, 7.4, 7.6, 7.7, 8.0, 8.2, 8.6, 8.6, 8.6, 8.8, 8.8, 8.9, 8.9, 9.5, 9.6, 9.7, 9.8, 10.7, 10.9, 11.0, 11.0, 11.1, 11.2, 11.2, 11.5, 11.9, 12.4, 12.5, 12.9, 13.0, 13.1, 13.3, 13.6, 13.7, 13.9, 14.1, 15.4, 15.4, 17.3, 17.3, 18.1, 18.2, 18.4, 18.9, 19.0, 19.9, 20.6, 21.3, 21.4, 21.9, 23.0, 27.0, 31.6, 33.1, 38.5. data set 2: y (m=60) 0.1, 0.2, 0.3, 0.7, 0.9, 1.1, 1.2, 1.8, 1.9, 2.0, 2.2, 2.3, 2.3, 2.3, 2.5, 2.6, 2.7, 2.7, 2.9, 3.1, 3.1, 3.2, 3.4, 3.4, 3.5, 3.9, 4.0, 4.2, 4.5, 4.7, 5.3, 5.6, 5.6, 6.2, 6.3, 6.6, 6.8, 7.3, 7.5, 7.7, 7.7, 8.0, 8.0, 8.5, 8.5, 8.7, 9.5, 10.7, 10.9, 11.0, 12.1, 12.3, 12.8, 12.9, 13.2, 13.7, 14.5, 16.0, 16.5, 28.0. first,we checked the suitability of quasi lindley distribution for the considered real data sets. we, therefore, have provided the kolmogorov-smirnov (k-s),anderson-darling(a-d) and cramér-von mises statistics to test the goodness-of-fit of above data sets to the quasi lindley distribution. the fitting summary has been presented in table 7.1, which indicates that the qld fits well to data set 1 and data set 2. table 7.3. p-value (statistic) of different goodness-of-fit tests for data set 1, 2. k-s a-d cramér-von data set 1. 0.0654 (0.1290) 0.0217 (3.2033) 0.0501 (0.4610) data set 2. 0.9287 (0.0677) 0.965 (0.2597) 0.9310 (0.0404) based on the mles θ̂1, α̂1, θ̂2, α̂2 the point estimate of r is 0.59 and the 95% confidence interval of r is (0.25, 0.93). for real data sets, the maximum likelihood and bayes estimates of the stress-strength parameters and reliability are summarized in table 7.2. copyright c© 2018 assa. adv syst sci appl (2018) 50 m. mohie el-din, a. sadek, shaimaa elmeghawry table 7.4. the mles and bayes estimates of stress-strength parameters and reliability r from real data sets est. θ̂1 θ̂1 α̂1 α̂2 r̂ mle 0.1 0.27 84.09 0.41 0.59 bayesimp 0.13 0.59 1.44 0.66 0.81 bayesmh 0.1 0.54 22.13 0.69 0.77 8. conclusion in this paper, maximum likelihood and bayesian estimation methods for stress-strength reliability r were discussed, when x and y both follow a quasi lindley distribution with different parameters. we obtained the 95% confidence intervals of r based on the observed fisher information matrix. we proposed the bayesian estimation based on independent gamma priors under different error loss functions(squared, linex, and general entropy). we suggested the importance sampling, and mcmc techniques to generate samples from the posterior distributions and then compute the bayes estimates. simulation study has been introduced to investigate the performance and compare among all mentioned methods. simulation results suggest that the performance of the bayes estimator is better than maximum likelihood for all different sample sizes. also, maximum likelihood method provides very satisfactory results as the sample size increased. references 1. al-mutairi, d.k., ghitany, m.e. & kundu, d. (2013). inferences on stress-strength reliability from lindley distributions. communications in statistics-theory and methods, 42, 1443-1463. 2. al-mutairi, d.k., ghitany & m.e.,kundu, d. (2015). inferences on stress-strength reliability from weigthed lindley distributions. communications in statistics-theory and methods, 44, 4096-4113. 3. basu, a.p & ebrahimi, n.(1991). bayesian approach to life testing and reliability estimation using asymmetric loss function. jour. stat. plann. infer, 29, 21-31. 4. braess, d. & dette, h.(2004). the asymptotic minimax risk for the estimation of constrained binomial and multinomial probabilities. sankhyaind.j.statist.66(4), 707732. 5. calabria, r. & pulcini, g. (1994). an engineering approach to bayes estimation for the weibull distribution. microelectronic reliability, 34 (5),789-802. 6. chen, m.h. & shao, q.m. (1999). monte carlo estimation of bayesian credible and hpd intervals. journal of computational and graphical statistics, 8(1), 69-92. 7. church, j.d. & harris, b. (1970). the estimation of reliability from stress-strength relationships. technometrics, 12, 49-54. 8. hastings, wk. (1970). monte carlo sampling methods using markov chains and their applications. biometrika, 57, 97-109. 9. khan, adil h. & jan, t. r. (2015). estimation of stress-strength reliability model using finite mixture of two parameter lindley distributions. journal of statistics applications and probability, 4(1), 147-159. copyright c© 2018 assa. adv syst sci appl (2018) estimation of stress-strength reliability 51 10. kotz, s. , lumelskii,y. & penskey,m. (2003). the stress–strength model and its generalizations and applications. world scientific, singapore. 11. krishnamoorthy, k., mukherjee, s. & guo, h. (2007). inference on reliability in twoparameter exponential stress-strength model. metrika, 65, 261-273. 12. lindley, d.v.(1958) fiducial distributions and bayes theorem. journal of the royal statistical society, 20, 102-107. 13. metropolis, n., rosenbluth, aw., rosenbluth, mn., teller, ah. & teller, e., (1953). equations of state calculations by fast computing machine. journal of chemical physics, 21, 1087-1092. 14. pandey, h. & rao, a. k. (2009). bayesian estimation of the shape parameter of a generalized pareto distribution under asymmetric loss functions. hacettepe journal of mathematics and statistics, 38(1), 69-83. 15. parsian, a. & kirmani, s. n. u. a. (2002). estimation under linex loss function. hand book of applied econometrics and statistical inference, statistics textbook and monograph.165. new york: marcel dekker inc, 53-76. 16. dey, s. (2010). bayesian estimation of the shape parameter of the generalized exponential distribution under different loss functions. pak. j. statist. oper. res, 62, 163-174. 17. shanker, r., sharma, s. & shanker, r. (2013). a quasi lindley distribution. african journal of mathematics and computer science research, 6(4), 64-71. 18. singh, s. k., singh, u., & sharma, v. k. (2014). estimation on system reliability in generalized lindley stress-strength model.journal of statistics applications and probability, 3(1), 61-75. 19. soliman, a.a. (2000). comparison of linex and quadratic bayes estimators for the rayleigh distribution. commun.statis.theory meth, 29(1), 95-107. 20. varian, h.r.(1975). a bayesian approach to reliability real estate assessment. amsterdam, north holland, 195-208. 21. zellner, a. (1986). bayesian estimation and prediction using asymmetric loss functions. journal of american statistical associations, 81(394), 446-451. copyright c© 2018 assa. adv syst sci appl (2018) advances in systems science and applications (2012) vol.12 no.1 67-75 the research on effective video scene character extraction algorithm in natural scene images jinghua hu1 and huafeng kong2 1wuhan university of technology, wuhan 430070 2key laboratory of information network security, the ministry of public security, the third research institute of ministry of public security, shanghai 201204 abstract pcb technology is particularly important in electronic industry. however, the increasing of technology complexity makes it difficult to inspect the quality of pcb. aoi system greatly improves the pcb inspection efficiency, and it has become one of the most instructive subjects that how to identify the text information in the chip image rapidly and accurately. since the chip images have complex natural background, and are greatly affected by light, shadow, noise, font and size, texture, color, position and arrangement., it is often difficult to inspect, extract and identity the text. this paper presents an effective algorithm of scene character segmentation and recognition in natural scene images. the algorithm segments lines of character, then segments every line of character into individual words for further processing, such as feature extraction and character recognition according to the known features of character. after the segmentation of character, we use an advanced bp algorithm to recognize the character. it improves bp mainly through restructuring gradient in the sigmoid. the experiment shows a significantly improvement in pc: the recognition accuracy rate of this system is above 96% and the response time is 5ms/100 words. keywords scenes character segmentation, feature extraction, bp algorithm, and image processing 1 introduction as the optical character recognition (ocr) technology comes in vogue, many scholars begin to research on the character extraction from the document images. till the 1990s, during the rapid development of computer technology and multimedia technology, the content-based multimedia retrieval has become a research hotspot. then, the character extraction in the natural scene images has again gradually become one of the research hotspots. usually, the natural scene image character has substantial changes in font, size, color, alignment and arrangement; also with the complex character background, low resolution and high noise. furthermore, many systems also require the algorithm has higher processing speed in the application. all of these causing difficulties to effectively extract the characters from the natural scene images, especially for the natural 68 jinghua hu:the research on effective video scene character extraction algorithm... scene images based on the videos. many domestic and foreign scholars have made beneficial explorations and attempts in this field. y.zhong[1] firstly put forward the resolution that position the characters in the complex images. however, this resolution mainly aims at the character positioning in the image scanning for the colorful disc covers, and fails to be directly applied into the natural scene images. a.k.jai[2] and others proposed a kind of character positioning method, which applies to the newspaper, web pages and general images and video frames, but not ideal to identify the small character fonts. m.a.smith[3] and others developed a kind of method detecting the characters on the images. however, because of the limitation in size, this method can only detect the characters within the specific font scope. it can not use the feature that the same characters will appear in multi frame to further enhance the character detection performance. also, the word segmentation prepared for the ocr is not done. sato[4] and others developed a character segmentation and recognition system aiming at the static low resolution headlines. the system will take advantage of the method mentioned in literature[3] to recognize the character. then magnify as 4 times as the recognized characters, which will use the smallest time-based image pixed value to conformify the headline. such system has a good effect on news program. r.lienhart[5-6] and others have developed systems in character detecting, identifying, recognizing successively. the color-based algorithm system of inter infiltration segmentation adopted in literature[5], only deals with image segmentation and recognition in single frame, without consideration in successive images in multi frame. however, the system in literature [6] takes further consideration on texture feature in title character, and tracking and conforming in successive images in multi frame, which makes it a better recognition in ocr. character recognition has a close connection with character segmentation which is one of critical factors in character recognition[7]. since the threshold process in natural scene image tends to bring about noise and low quality problems, and probably results in character conglutination. such traditional character segmentation will not be satisfied with people. in order to identify characters in the natural scenes fast and accurately, localizing the chunks should be taken as first priority, and later extracting smaller chunks. therefore, the small chunk could be used respectively to sharpen the noise tremendously, without the complex impact from other parts in images.as a result, this paper presents a algorithm of scene character segmentation and recognition in natural scene images, based on feature feedback, which will segment chinese, english and adherent character. after the segmentation of character, an improved bp algorithm is afterwards used to recognize the character. advances in systems science and applications (2012), vol.12, no.1 69 2 feature feedback based algorithm of character segmentation & recognition between chinese, english and adherent character several traditional methods in character segmentation & recognition[8] 1⃝ imagebased analysis. search for rational segment point between characters, chiefly adopt static projection analysis; 2⃝ recognitionbased. select various kinds of current segmentations via identifying ability before actual segmentation. 3⃝ synthetical based on image analysis and recognition. reduce and filter vertical segmentation hypothesis by image analysis. 4⃝ unity recognition. use the whole word as recognizing object, according to the features of the whole word, to avoid segmentation harm on character. however, in the case of word conglutination, those methods can not segment adherent characters accurately. this paper deals with this kind of situation by an algorithm of feature feedbackbased character segmentation and recognition between chinese, english and adherent characters. first segment each line of character, and then search and extract the adherent character in each line, finally transfer small segmenting module to discover the most reliable segment point, so as to achieve the segmentation purpose.key points as follows: 1.identify the category of adherent characters after initial identification, the recognition and length of each image as follows: 1⃝ suppose the length of each character image is w, and the average length of each character image is w. when w is much wider than wv, and pick up a threshold via experiment, we could identify whether it is adherent character or not. 2⃝ after identifying adherent character image, according to english character image length is wider than the chinese character image length, and the separation distance is small between the adjoining characters, we could identify whether the adherent character image is chinese or english; according to average english adherent character length is shorter than the chinese one, we could identify those character images with wrong structure as adherent character image. 2.identify the category of adherent characters for the adherent chinese character, we can segment the chinese character by judging the potential character image length. we use identifying module to judge the border to segment the character. in this process, we confirm the most reliable segmenting point by right-to-left and left-to-right searching methods.algorithm as follows: step1, record the value of left border bl, right border br and line height hl. step2, cobfirm x0 value (x0 = hl)from the left border, choose one threshold value t0, the segmenting point should be picked up from ( x0-t0 , x0+t0), modulate it according to steplong. 70 jinghua hu:the research on effective video scene character extraction algorithm... step3, transfer identifying module, and judge whether the result is in the range or not. if not , do step2; if is , do step 4. step 4, save segmenting result br1, take br1 as the left border , and then do step2 and the following steps successively, till the right border br. to compare the validity of the final segmenting result, do all steps in turns from the right border, and save the best segmenting result. 3.segment adherent english number for the adherent english number, first use border searching algorithm to segment these adherent character with gaps in between, which can not be segmented by projection method. then suppose the potential height of each adherent character string is x, according to x, transfer the identifying module to segment the adherent character string so as to locate the segmenting point. border searching algorithm as follows: step1, take note of left rect urlli ,tiand right rect urlri ,bi, and pick up a threshold value t,find a p point whose gray-scale value is 1, and take it as initial point of the contour line. step2, start with the fund point and continue with the searching job. if the gray-scale value is 1, please carry on to the left; if the gray-scale value is 0, please carry on from the right turn. once meet the point whose gray-scale value is 1, these points would be contour point. step3, repeat step 2 till the contour line point coincides with the initial point; step4, confirm the right border via contour line searching. if br1 coincides with the right border, there is no adherent character; if not,for identified character we can pick up a suitable threshold via experiment, and change it to the left or eight. then use it as the segmenting position and save the result.if character width to height ratio is large, the character is adherent; if the ratio is relatively small, we could transfer the identifying module to check it is or not, after it being segmented. if it is adherent we can calculate its potential height value, according to whose height we adjust a threshold value and use it as its possible width. then we take advantage of identifying module to segment the adherent character and confirm the segmenting position via identified result. save the correct final segmenting result. the english character baseline feature as follows: for the adherent english number, if the line height is close to 5/6 of chinese fig.1 english character baseline feature line height, there are up-protruding letters (e.g. t) as well as down-protruding letters (e.g. p). we confirm the baseline approximately as much as 1/4 as adheradvances in systems science and applications (2012), vol.12, no.1 71 ent character height, and height x is approximately as much as 1/2 of adherent character height; if the line height is close to 2/3 of chinese line height, there are up-protruding letter or down-protruding letter. we confirm the baseline approximately as much as 1/3 as adherent character height, and height x is approximately as much as 2/3 of adherent character height; if the line height is close to 1/3 of chinese line height, there are neither upprotruding letter nor down-protruding letter. we confirm the baseline approximately as much as 1/2 as adherent character height, and height x is approximately as much as adherent character height. and we can adjust the potential segmenting position to meet height x via threshold-adjusting. and the same method could be used in english number adhesion and chinese adhesion. feature feedback based algorithm of character segmentation & recognition between chinese, english and adherent character identify result as fig.2 and fig.3. fig.2 english character segmenting result fig.3 chinese character segmenting result 3 character size normalization after being segmented,each character is in different size; therefore the segmented character region is different.in order to extract the character feature,we need to normalize the segmented character region. suppose height and width to represent respectively normalized character vertical and horizontal length. and h and w respectively represent height and width of the character rectangle. therefore, the ratio of vertical to horizontal hr and wr would be: hr = height h ,wr = width w 72 jinghua hu:the research on effective video scene character extraction algorithm... then we make coordinate mapping, and we suppose coordinate (i new,j new) is the mapped coordinate (i,j) after it has been normalized. similarly we map the following coordinates: i new = top+ (i− top)/hr, j new = left+ (j − left)/wr note: top is the vertical coordinate of the point a in the mornalized character region, and left is the horizontal coordinate of the point a in the mornalized character region. the character size normalized result is shown in fig.4. assign the gray-scale value of the character coordinate (i,j) before being normalized to the normalized character coordinate (i new,j new) to accomplish the character normalization. (a) before normalization (b) after normalization fig.4 contrasts two normalization results 4 improvement of bp algorithm base on the traditional bp algorithm, this paper presents an improved algorithm which improves bp mainly through restructuring gradient in the sigmoid. if the value is supposed too much, the input 0 or 1 of each layer would be scattered, and the study result is worse. otherwise, if too little, the system linearity is enhanced whereas nonlinearity is weakened. therefore, the best value is between the above both. the most important point is to amend the gradient of each allnodes in the sigmoid so as to find a best value. in the improved bp algorithm, sigmoid is turned into binary driving function on the connection of gradient and the potential value of the junction point: φ(α, υ) = 1 1 + e−∞ we can easily deduce that after all the amendment, adjusting method of the weight and threshold is the same with bp algorithm. suppose the actual output vector o = {o0o1, . . . , ol}, teacher vector t = {t0, t1, . . . , tl}, mean square error e: e = 1 2 ∑ k (tk −ok)2 advances in systems science and applications (2012), vol.12, no.1 73 gradient α is changed according to e ′ negative gradient change, where mean square error e reduces the least. hence the improved gradient algorithm would be: ∆α∞ = ∂e ∂α the node point of the output layer is: ok = 1 e−αk(µk−θk) ,then ∆αk = −η ∂e ∂αk = −η ∂e ok · ∂ok ∂αk = −η(tk −ok) ∂ ∂αk 1 1 + e−αk(µk−θk) = η(tk −ok)(µk − θk)ok(1−ok) η is studying step length,okis the actual output of output layer neuron k,tkis teacher signal,µkis the input signal linear combination of the output layer node, then µk = ∑ k hjwkj ,wkj is weight;θkis threshold;hj is the input signal of output layer nodepoint, that is the output of intermediate layer neuron. for the node point of ∆αj = −η ∂e ∂αj = −η ∂e ∂hj · ∂hj ∂αj hidden layer ∂hj ∂αj = ∂ αj · 1 1 + e−αj(µj−θj) = (µj − θj)hj(1−hj) ∂e ∂hj = ∑ k ∂e ∂µk · ∂µk hj = ∑ k ∂e ∂µk ·wkj ∂e ∂µk = ∂e ∂ok · ok µj = −(tk −ok)αkok(1−ok) substitute in turns would be: ∆αj = η(µj − θj)hj(1−hj) ∑ k (tk −ok)αkok(1−ok)wkj using driving function via the improved gradient, we can adjust the gradient of sigmoid to its best in the process. to study the photo sample, we memorize the character feature in the sample and find the weight and threshold, and then we establish and improve bp neural network. then we can input a photo, using bp neural network, to find a closest character to the input chinese character in the sample. the steps as follows: (1)network topologymost characters to be identified are chinese characters. because the chinese character structure is quadrel, we divide them into 26×26 unit. each unit corresponds to one input; hence the output nodes in the neural 74 jinghua hu:the research on effective video scene character extraction algorithm... network would be 26×26=676. the numbers of the output nodes in the neural network is the same with the characters to be identified. (2)output node confirmnesscalculation of node output is equivalent to calculation of the function sigmoid output. because of the functions nature, we would be aware of its value range [0,1]. the closer the value is to the limits, the less sensitive the function value changes to its independent variable. therefore, we can enhance the nodes precision by its output-to-input sensitivity. we suppose t is 0.9, and f is 0.1; for input, if with stroke, then input 0.9, if not, we input 0.1. and we deal with output the same way. (3)edit of training samplethe more samples the collection owns, the more corresponding sample each chinese character has. also the better identifying ability the neural network is after being trained, however, the more time the training takes. here we only take the sized 18 thick song ti character as sample. there are english characters, numbers, 3,775 chinese characters of one-level, and basic punctuation in the sample collection. (4)experiments and analysis the identifying result of bp-algorithm-based character identifying algorithm is shown in fig.5. this system is designed to deal with images including character, whose recognition rate is above 96%, and its response time is 5ms / 100 words. compared to the traditional recognition method it has been improved, whose recognition result is better. however, the worse the images quality will result in the worse the recognition result. in the image treatment phase, the wiping-out of useless information is partial, which has an effect on character recognition and the whole recognition result. fig.5 results of character identifying algorithm 5 conclusion according to the above discussion, the identifying result of bp-algorithm-based character identifying algorithm can improve the efficiency of ic identifying in aoi system. advances in systems science and applications (2012), vol.12, no.1 75 references [1] yu zhong, k.karu, a.k.jain. (1995), “locating text in complex color images”, in proceedings of the third international conference on document analysis and recognition, vol.1, no.14, pp.146-149. [2] a.k.jain, b.yu. (1998), “automatic text location in images and video frames”, pattern recognition, vol.13, no.12, pp.2005-2076. [3] m.a.smith, t.kanade. (1995), “video skimming for quick browsing based on audio and image characterization”, 3technology report cmu-cs-95-186, pp.13-15. [4] t.sato, t.kanade. (1995), “video ocr: indexing digital news libraries by recognition of superimposed caption”, multimedia systems, vo.l7, no.5, pp.385–395. [5] h.j. zhang, c.y. low, s.w. smoliar, and j.h. wu. (1995), “video parsing, retrieval and browsing: an integrated and content-based solution”, proc. acm multimedia , pp.15-24. [6] r.lienhart, w.effelsberg. (2000), “automatic text segmentation and text recognition for video indexing”, multimedia systems, vo.l8, no.2, pp.69-81. [7] huiping li, omid kia, david doermann. (2004), “text enhancement in digital videos”, proc. spie99-document recognition and retrieval, pp 3651:2-9 [8] casey r g, lecolinet e. (1996), “a survey of methods and strategies in character segmentation”, ieee transactions on pattern analysis and machine intelligence, vol.18, no.7, pp.690-706. advances in systems science and applications (2014) vol.14 no.4 302-324 the paradoxes of the world’s progress (i) sailau baizakov1 and jeffrey yi-lin forrest2 1jsc “iei” economic development and trade of kazakhstan, jsc “economic research institute”, str. temirkazyk 65, 0100000, astana, kazakhstan 2department of mathematics, slippery rock university, slippery rock, pa 16057, usa abstract an unbiased systematic view on the modern innovation allows you to see three key groups of innovations that are still very few people differentiate. the first is the technical and technological innovations that are the basis of development and change in technological ways of the world. secondly, it is monetary and financial innovation, the progress of which determines the change in monetary ways of life of the world. and third, it is the socio-political innovation, progress of which is in the basis of the change of socio-political models of the world. a clear distinction between these three “floors” of innovations is crucial for understanding the global crisis and the ways to quit it. since the essence of it lies in the tangle of clearly long overdue and contradictions between the rates of introduction of the world’s technical and technological, financial and socio-political innovations. as the global crisis clearly shows, technical and technological way today, is not decisive for the country’s prosperity and peace. exactly the countries with the highest level of technology development have become the main source and a key cause of the global crisis. keywords global crisis, sona analyzer, technical and technological innovation, monetary and financial innovation, socio-political innovation 1 analyzer sona the system weapon of independent analysts of kazakhstan 1.1 why the global crisis can not be regarded as finished ? the tone of the work on the vi astana economic forum was given by the president of kazakhstan. his conclusion is that the global crisis can not be regarded as finished. nazarbayev believes that the crisis now acquires explosive features, with local manifestations, such as in cyprus. what is the reason ? according to the president of our country, this is the artificial “inflation” of financial bubbles and desire for “easy money”. this is the lack of proper responsibility of national financial institutions, the ineffectiveness of anti-crisis measures because of the weakness of the mechanisms of global financial governance. anti-crisis measures taken at national level by the dictating of the imf and other un institutions themselves become the cause of recession. thus, the pact of growth and stability, adopted recently in europe, did not give real results. a year advances in systems science and applications (2014) vol.14 no.4 303 has not passed since the pact ; there is a new pact on the promotion of growth and competitiveness preparing. headache for leaders of european countries is unemployment, a large proportion of non-employed people among youth. to accept radical solutions to the global economy, sufficient will and responsibility have not yet been manifested. there are neither effective global anti-crisis mechanisms, nor reliable global reserve currency. the rules of international consensus were not worked out, meeting the mutual interests of global financial institutions and nation-states actors of the financial sector and the real economy of the world. any pact of economic growth and stability will be reduced to practice of the “patching holes up”, if the boundary of the regional union, as in the european union, expanding blindly, mechanically, without regard to the interests of each country separately. the conclusion of the president of the republic of kazakhstan is original : the world needs a new economic model of governance, which should be based on welldeveloped tools of their integration. thus, the financial sector can not and should not develop in isolation from the real production at all levels of management. and, it is important to study and solve the problem of monetary system off-shore in some countries. according to n. nazarbayev, the current global crisis is multidimensional. its speed and dynamics are determined not only by the economic but also the political, humanitarian and moral-value causes. in the end, the president of kazakhstan proposed to begin work on the pact of global regulation. it is said about the multidimensional innovative tool for construction a new global financial system and the creation of global regulator, which determines a uniform playing rules. the basis of its design is capacity of five simple and clear principles : i evolution and rejection of revolutionary change in policy ; i justice, equality, consensus ; i global tolerance and trust ; i global transparency ; i constructive multipolarity. 1.2 three types of innovation in the economy of the world : analysis and problem there is an urgent objective need for a number of analytical researches, including but not beyond one-sided judgments about the benefits of the existing models of management of the national economies of individual countries, such as anglosaxon. today’s rapidly changing world of market forces in the world economy makes innovation management to lift to new heights and to achieve consistency with the development of its real and financial sectors. the steady growth of the national economies of individual countries and the building of its respective management model, as pointed out by n. nazarbayev in his article “the fifth way”, are closely related to the harmonization of the rates 304 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress of development of the three types of innovations : type 1 innovations in the production of real goods ; type 2 innovations in the field of renovation of monetary system ; type 3 innovations in the field of social and political life of people, business, organization, region, etc. of course, the economic model of the future, that will enable implementing the principles of mutual coordination of rates of development of these types of innovations, defines a roadmap for sustainable post-crisis economic growth of not only individual countries but for all countries in the world[1-3]. market economy itself is based on the principles of competitiveness and innovation development in the production of real goods and services. due to the modern information and computer technology there is a scope for the smooth implementation of the first type of innovation in the real sector of the economy. there are no obstacles to the development of innovation type 1, except legal obstacles and restrictions on private sector development and further liberalization of its work. as for the second type of innovation, now in most countries of the world unlimited possibilities for the development of innovative technologies in the monetary and financial system are created. here, same as in production, there are no obstacles in the way of innovation type 2. the main obstacle, which is now holding back the potential of both of these types of innovative development, is the third type of innovation. innovations in the field of social and political life of people, business, organization, region, etc. in most countries are based on a mechanical copying the anglo-saxon model and implementation of the principles of the washington consensus. but these models and principles of its construction have already reached their limits and were inadequate for preventing the crisis manifestations and overcome the consequences of the crisis. this explains the incompleteness of the global crisis, the permanent duration of local manifestations and long term duration of its consequences. thus, the stagnation in the development of economic science and the lack of technical technological innovation in governance does not enable control on the gaps between turnover of money and goods flows between the rates of development of the real and financial sectors of the economy, not to mention the innovative rate of development of tools of economic management. the rate of the third type of innovation and the development level is the weakest link, which is an obstacle in the whole chain types of innovation development of economy and ensuring its financial stability. at the moment, the only obstacle in the way of innovation development of the world economy is exactly the type of innovation type 3, innovations in the field of socio-political life of people, business, organization, region, etc. an innovative tool type 3 analysis and management is needed not only at the advances in systems science and applications (2014) vol.14 no.4 305 level of macroeconomics : it is needed in each area of the economy from home economics, including the economics of the business sector and ending with the national economy. otherwise making the systematic and right decisions excluding the formation of imbalances in the implementation phase of projects and programs are not possible. to date, the initiative group of kazakh analysts developed a system tool with innovative components that meet the principles of “fifth way” realistic “estimate, measure, exchange, transfer the true cost of goods and services”[1]. this versatile tool of economic management is analyzer sona, which is fully belongs to the “new financial instruments of a new quality : the real measuring instruments of cost of goods and services”[ibid]. 1.3 three indexes of economic development and financial stability sona analyzer is able to determine the rate of balanced growth (growth index i3, which determines the level of innovation of the type 3), which will connect the index of the nominal growth rate of financial stability i1, which determines the level of innovation of type 1 with an index of the rate of real growth in the sphere of production i2, which determines the level of innovation type 2. thus, the establishment of a balanced growth rate measures the gaps between the development of the real sector and the financial system. and measurable economy is manageable economy. what does the index for balanced growth mean ? the answer to this question, just like the question of what the index of the real growth means is directly related to the concept of index of nominal growth. so first you should understand the index of real growth. if the index of the nominal growth is determined by the ratio of nominal gdp in the prices of the current year t to nominal gdp of base year 0 by the formula i1(t) = ngdp (t)/ngdp (0), then the index of real growth rate is determined by the index of the physical volume of a good or service by formula i2(t) = rgdp (t)/rgdp (0). the first index is the gdp growth rate in prices of current year. but the economic content of the index is determined not only by market prices of the current year, but the nominal value of the national currency of the country concerned. this is important, the fundamental advantage of a market economy compared with a directive command economy, which is focused on the principles of hard pricing. it should be noted that the index of growth of nominal gdp is expressed in nominal charge of the national currency of current year. thus, the nominal price of the national currency of kazakhstan in 2008, according to official statistics was 120.30 tenge per u.s. dollar, and three years later, in 2010 it was 147.35 tenge per u.s. dollar. its rate of growth for only three years was 122.5%. the second index, which is called the index of the real growth, is the growth 306 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress rate of gdp in 2010, but prices of the base year 2008. the official statistics, using special methods for calculating, transfers nominal gdp of current year to gdp of the same year, restated to prices of the base year. in our example, the nominal gdp in 2010 prices is transferred to gdp-2010 at 2008 prices. so far, to ascertain the nature of the gdp deflator, we have two gdp growth rates of 2010 relative to the same, letąŕs say, a base 2008. one is index of gdp growth in the nominal value of the national currency of the current year, the other is in the nominal value of the national currency of the base year. if we define the growth rate of nominal gdp relative to 2008, we have growth index i1(t) = ngdp (t)/ngdp (0), and if we define the growth rate of real gdp relative to 2008, we have index i2(t) = rgdp (t)/rgdp (0). basis for calculating both indexes of growth is 2008 and their performances in the base year equal to each other : rgdp(0) = ngdp(0). both of these figures represents a monetary phenomenon, and except tenge, kazakhstan national currency, does not contain any ounces of gold, or a watt of energy, not an ounce of food or a basket of rational budget or one unit of another substance in nature, and not even a single unit foreign currencies, including the u.s. dollar. one more step to determining the rate of balanced growth brings us the concept of the nature of the gdp deflator. denoting index of gdp deflator to official statistics p(t), we can express its meaning in classical method : p(t) = i1(t)/i2(t), which measures the gap between the nominal and real indexes of growth. analysis of the formula of the gdp deflator shows that at equal rates of growth of nominal gdp and real gdp (i1(t) = i2(t)), the gdp deflator is constant and equal to one. the rate of growth of the real sector and the financial system are in an equal distance from the bisector of the coordinate system (x, y). this equilibrium line of growth rates in both sectors of the economy represents an ideal balanced growth dynamics the index i3(t). of course, the nature of the change of each index of growth is influenced by its source of innovation development. thus, the source of market forces development of the growth rate of nominal gdp of financial stability i1(t) is the potential of innovative development of the monetary system. and source of market forces development of the growth rate of real gdp i2(t) is the potential of the innovative development of the production in the real sector of the economy. on the equilibrium trajectory of the development rates of the real and financial sectors of the economy equality of not two, but three indexes of growth is attained. this third index, as already indicated as i3(t), is the index of balanced growth. the function of this third type of innovation is to ensure the comparability of development the real and financial sectors of the economy not only in the indexes of prices of goods and services, but also on the purchasing power of the national advances in systems science and applications (2014) vol.14 no.4 307 currency, to provide scientific management by the first two types of innovation. however, in reality a balanced of growth rate (i3(t)) does not coincide with any index of real or nominal growth, and may be freely deflected to either side of the line of the ideal dynamic equilibrium. let his deviation from the growth rate of nominal gdp is α : α ∗ i1(t) = i3(t), and the growth rate of real gdp is β : β ∗ i2(t) = i3(t). the growth rate of the economy i3(t) on the equilibrium line αi1(t) = βi2(t) is said to be the dynamics of the rate of balanced growth. how accurately does a triple measurement tool for economic growth meet the criteria of democratic governance and the principles of liberalization of the market economy ? the section of the article“the fifth way” n. nazarbayev “the paradox of the global progress” will help to answer this question : “an unbiased systematic view on the modern innovation allows you to see three key groups of innovations that are still very few people differentiate. the first is the technical and technological innovations that are the basis of development and change in technological ways of the world. secondly, it is monetary and financial innovation, the progress of which determines the change in monetary ways of life of the world. and third, it is the socio-political innovation, progress of which is in the basis of the change of socio-political models of the world. a clear distinction between these three “floors" of innovations is crucial for understanding the global crisis and the ways to quit it. since the essence of it lies in the tangle of unsolvable contradictions between the rates and levels of development of technological, financial and monetary, social and political structure of each country and world as a whole. or, in other words, it is the tangle of clearly long overdue and contradictions between the rates of introduction of the worldąŕs technical and technological, financial and socio-political innovations. as the global crisis clearly shows, technical and technological way today, is not decisive for the country’s prosperity and peace. exactly the countries with the highest level of technology development have become the main source and a key cause of the global crisis ”[1]. 1.4 the transformation of the gdp deflator into indicators of balancing real growth rates of production areas with the rate of monetary and financial system remains to find out what forces withdraw the initial equity indices of growth in both sectors of the economy in 2010 from balance, when in the base 2008 they were equal : rgdp(2008) = ngdp(2008). no force will make us to consider nominal gdp balanced to real gdp in 2010, when at the point p (2010) = 113.4% economy has become an equilibrium with only two indices of growth i1(t) and i2(t) : 113.4 ∗ i2(t = 2010) = i1(t = 2010). at this point, equilibrium in the face 308 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress of the national currency of kazakhstan has not yet reached because in 2008 the price was 120.30 tenge per u.s. dollar, and three years later, in 2010 it was 147.35 tenge per u.s. dollar. what kind of balance can be discussed if the price of the national currency in only three years has deviated from its base by 22.5%, and the magnitude of the gap was even 8 percentage points more than the value of the gdp deflator (inflation rate) for the same years that’s why n. nazarbayev in his “paradox of the global progress” highlights the need to overcome “tangle of clearly and long overdue contradictions between the rate of introduction of the world’s technical and technological, financial and socio-political innovation” [1]. according to claims of training manual of sachs and laren on the economy, the gdp deflator used to calculate the inflation rate in 2010 is not the same as the index of growth of actual prices of goods and services of that year : “note j. sachs and f. laren support our conclusion that we calculate the price index indirectly. at first, we take the nominal gdp (ngdp) in current prices, then find a real rgdp in constant prices, ie q = rgdp. therefore price deflator calculated in this way is sometimes called implicit price deflator of gdp ”[13]. only the determination of the nature of the gdp deflator (inflation index) will help to understand the “tangle of unsolvable contradictions between the rates and levels of development of technological, financial and monetary, social and political structure of each country and world as a whole” which was mentioned by president [1]. to assess those market forces that unbalance the economy, we assume that the gdp deflator in 2010 represented the integrated expression of two opposite market forces. the first of these forces α = pp(t)represents purchasing power of national currency, and the second one β = c(t)represents market force of technical and technological structure of modern production sphere, expressed by the ratio of ntp in the real economy. as a result of this assumption, the rate of balanced growth is represented by two components of the gdp deflator, one of which serves as the weighting factor of equilibrium with the rate of nominal gdp growth (growth index i1(t)), and the other is with the rate of growth of real gdp (growth index i2(t) ). now easy to prove that the rate of growth of the real sector with a weighting factor ntp c(t) is balanced with the rate of growth of the financial sector of the economy on the line i3(t) : c(t) ∗ i2(t) = pp(t) ∗ i1(t). thus, the objective necessity of studying the nature of the gdp deflator does not derive only from the assessment of purchasing power of the national currency, but also from the assessing the contribution of scientific and technological developments in real growth rate of production of material goods. advances in systems science and applications (2014) vol.14 no.4 309 1.5 the function of scientific and technological progress however, the theory of economic growth does not allow us to calculate the size of the contribution of scientific and technical progress, the more the contribution of innovations to economic growth. thus one of the most important conclusions drawn from the theory of “romer and lucas according to russian sources for economic research is the fact that the economy, which manages large resources of human capital and the development of science in the long run is more likely to increase than economy, not having these benefits” [5]. it emphasizes only opportunities and points to the overall usefulness of technical and technological improvements. an attempt to solve the problem of estimating the contribution of stp, independently as the effect of investments in fixed assets at the time was made by n. kaldor and j. mirrlees [6]. james mirrlees is an active member of astana economic forum in recent years, the nobel prize in 1996. thus, the formulation of this problem in the model of kaldor mirrlees is stated as follows : let the balanced growth path is described by the interrelated functions of exponential following type [7-8] : gt(x) = g0(x)e z(x)t where g0(x) required initial level of indicator x to the balanced growth path ; gt(x) level of indicator x on the balanced growth path at time t ; z(x) required parameter value that determines the rate of increase in the index x on the same trajectory. kaldor mirrlees model forms a system of 11 equations with 11 unknowns. the peculiarity of this system is the presence within it of the equation that determines the function of scientific and technical progress. the function of scientific and technological progress is a performance increase depending on the growth of capital-labor ratio. in the simplest case, the function of scientific and technical progress is linear. in the latter case, the rate of balanced growth λ = z (x) is determined from the linear regression equation : φ t φt = σ1 + σ2 i′t it ;λ = it it = w′t wt = σ1 1− σ2 (1) σ1, σ2 > 0, σ2 < 1 where σ1, σ the coefficients of the regression equation. the economic content of the regression equation (1) reduces to the fact that with increasing intensity of labor productivity increases, but to a lesser extent, because it is assumed σ2 < 1. in fact, this limitation restricts the border of the 310 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress use of stp when the σ2 < 1 condition is not satisfied. models of this type do not take into account the explosive effects of individual scientific and technological activities. and because we believe that the contribution of scientific and technological progress allows us to understand more deeply the following two indicators. the first of them is the technical and economic level of production, which expresses the qualitative characteristics of scientific and technical activities at any given point in time. the second is effectiveness of the scientific and technical solutions that characterizes his return [13]. due to this mating pair of economic indicators a realistic model of the contribution of scientific and technical progress in the economy can be built. and the models as of the kaldor-mirrlees and other macroeconomic models are mostly theoretical and unlikely in the near future will find practical expression in economic management. however, the nbrk in its analytical work practices using the following econometric models, including : inflation equation, which is determined by the model : dlog(cpi)=0.3*diog(cpi(-1))+0.2*d(ulc)+0.02*gap+ +0.2*dlog(p_imp)-0.3*(log(cpi)-1.3-0.4*ulc-0.2*log(p_imp)), • cpi consumer price index,%, december1999=100 • ulc labor costs per unit of output • p_imp import price index, 2000q1=100 • gap deviation of gdp, equation of the gdp deflator by the other model : dlog(pgdp)=-1.3*(log(pgdp)+6.5-0.9*ulc-0.3* log(p_imp(-1)))+0.4*dlog(pgdp(-1))+0.3*dlog(pgdp(-4)) • pgdp gdp deflator,%, 2000 q=100 • ulc labor costs per unit of output • p_imp import price index, 2000q1=100 since the nbrk for determining an indicator of inflation applies one econometric regression equation, and to determine the gdp deflator applies other economic regression equation, it is easy to estimate the difference between them and it can be represented as a contribution to the scientific and technological activities. but the difference between two related indicators obtained in this way will be far from its real value. shown here a brief overview of the features of scientific and technological progress and econometric models of the gdp deflator (inflation index) allows us to understand the nature of the gdp deflator. so, in terms of content gdp deflator is different from the inflation index. inflation, as we know, is an immediate harm to the sustainable development of the market economy. however, the methodology for determining the inflation index and the gdp deflator is still single. in teaching aids and textbooks of such famous authors as sachs, dornbusch, advances in systems science and applications (2014) vol.14 no.4 311 mcconnell, menkyu these terms are used interchangeably and even are written together “gdp deflator (inflation rate).” however, the gdp deflator and the inflation rate needs to be systematically studied as indicators having different roles in determining the rate of balanced growth. thusin the construction of the analyzer sona, the contribution of stp is determined by a formula, the purchasing power of money to a different formula, and the gdp deflator by the third formula. these formulas of three mutually independent indicators allow to estimate the contribution of each of the real and financial sectors on the dynamics of qualitative and quantitative parameters of the development of the national economy. 1.6 analyzer sona is a versatile tool to support economic management analyzer sona is versatile tool that allows assessing, firstly, the contribution of scientific and technological developments in the real economy, and secondly, the purchasing power of the national currency, not only of kazakhstan but also the u.s. dollar, as well as the national currencies of other countries. the strength of analyzer sona is that the foundations of its construction are the laws of political economy. in the words of british economist john stuart mill, the laws of political economy act like the law of gravity, which is “without any compunction breaks neck to the best and dear man”, if he does not care about the consequences of the law of nature. unfortunately, in most countries of the world, the economy is dominated by legal laws that do not always conform to the objective economic laws. in some countries, political reform is not linked to the level of economic development, iron principle of scientific management of the market economy is disturbed, “economy first, then politics”. the main advantage of the analyzer sona is the use of an entire system of economic laws to ensure the sustainable development of a market economy. the analyzer shows that in a market economy is not market forces rule, but the system of economic laws by which equilibrium is established between the rate of nominal, real and balanced growth in the economy. the principle that has allowed them to balance is the principle of reversibility between the purchasing power of money and the market prices of goods and services. due to the principle reversibility indicators of the real and financial sectors are equally involved in the management of the market economy. both indicators conform to the methodological provisions of the statistics operating in the republic of kazakhstan. it is encouraging that in these methodological positions on statistics the difference between the deflators and price indices of goods and services is clearly indicated : “ comparing with the price index for goods and services, gdp deflator measures the change of wages, earnings (including mixed income) and consumption of fixed capital as a result of changes in prices and nominal net of 312 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress taxes” [9]. the formula for calculating the difference between the gdp deflator and the price indices of goods and services in kazakhstan is defined as clearly defined as ohm’s law in physics. the formula of this law is mathematically proved in section 4 of this paper, and it is economically justified that it is meaningful. the difference between them (its designation, c(t)) was equal to the production of gdp gdp (t), divided by gross output x (t), defined by amount of current material costs for the production of gdp qp (t) and of the gdp at current prices gdp (t) : c(t) = gdp (t)/x(t) = gdp (t)/(qp (t) +gdp (t)) thus, the impact of stp defined by its coefficient c (t), represents the effectiveness of the scientific and technical solutions. and the techno-economic level of production that defined the impact of stp, is gdp (t) / x (t). the peculiarity of the construction rate of stp is its expression ratio of two key indicators of national accounts : gdp and gross output. these indicators form the foundation of leontief table “input output”. it is considered that the total cost of production of specific types of products are gross output gross output (x). if you subtract from it a work in progress, it is determined by commodity products, expressed in money. if the cost of raw materials is deducted from commercial products, the newly created value or gdp will be determined. at one time, the index of gross output, called “gross social product” and the indicator “gross domestic product (gdp)” served as an apple of discord between the supporters of the private and public modes of production. the fact remains that the economists of the former soviet union, for the most part, considered gross output as the main indicator of economic development of the country, it is output x (the total social product x), giving parity to its naturalmaterial structure, the european union economists believe that gdp is the main indicator of the economy, giving the parity to its monetary cost structure. in the first case the goods were deficit in nature ; in the second case the money were deficit. the above formula for the contribution of the technical and technological progress shows that the dynamics of the growth rate of the indicator c(t) is expressed in a well-defined mathematical function defined by the ratio of gdp to be released. thus, it is strictly proved that the parity status should be fixed for the output and gdp, without mutual exclusion : either gdp or gross output. moreover, the gdp deflator is just one of the equivalent indicators of control and itself can serve as the function of the indicator ratio c (t) to the index of the purchasing power of money (pp(t)) : p(t) = c(t)/pp(t) advances in systems science and applications (2014) vol.14 no.4 313 in general, the originality of the principle of reversibility of commodity prices and the purchasing power of money is that the difference between the gdp deflator and price indices of goods and services determines the level of scientific and technological competitiveness (stc) of the real sector. and this difference between the gdp deflator and the price indices of goods and services can be called “contribution” of scientific and technological improvements aimed to increase the competitiveness of the real economy. the information necessary to calculate the difference between them is in the official statistics of the world, which is in line with international standards. however, the “contribution” of scientific and technological competitiveness (growth factors ntc) can be either positive or negative. negative growth of the stc rate in the real sector of the economy is formed when the current material costs of production qp (t) per unit of product will be greater than in the base year. conversely, the positive contribution of the stc in the real economy means lower production costs per unit of output of goods on the market conditions of development. in the case of the positive contribution of stp rate of economic growth is incremented by the value : ∆c(t) = c(t)− 100 > 0 and the index of the gdp deflator is reduced by the same amount. in the second case, on the contrary, when the contribution of stp is negative (∆c(t) < 0), the index of the gdp deflator is incremented, and it increases on this multiplicity −(∆c(t)). at the same time the rate of economic growth decreases by the same amount. thus, the objective the need to define index of market prices of goods and services is brewing other than the level of the gdp deflator. to save the gdp deflator as an independent regulator of the economy and deal with anti-crisis measures leads to a waste of money and time, as its value can be larger or smaller than the index of prices of goods and services by the amount of the contribution of stp. as evidenced by international experts, the losses of the world, gone with the wind of the global crisis of 2008-2009 exceed 10-12 trillion u.s. dollars. it is recognized that n. nazarbayev was right, pointing back in 2009 at a bottleneck in the chain of development of the three types of innovation as the paradoxes of world progress. despite the huge losses of the world economy, a tangle of “ unsolvable contradictions between the rates and levels of development of technological, financial and monetary, social and political structure of each country and world as a whole” remains unsolved to this day. the initiative group of economists kazakhstan believes that the gdp deflator is “explosive substance” containing in its composition “contribution” of scientific and technological improvements. the more real the positive effects of scientific 314 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress and technological progress is, the more powerful the explosive force of the gdp deflator. purification of the real price index of goods and services from this “explosive mixture” of the gdp deflator is impossible to carry out without the knowledge of the forces of economic laws and their practical application skills. its purification is performed only by those who know the power of these laws. otherwise, as shown by the economic crises of recent years it is impossible to pass the permanent shocks in different countries of the world. 2 the function of the gdp deflator in defining the principles of balanced growth 2.1 reciprocity principle is the core of the theory of balanced growth of the real economy and financial stability it is recognized that inflation is the scourge of the market economy. to control inflation, this section studied the nature of the gdp deflator, as a factor determining the trajectory of change in inflation. schematic diagram of the presentation of the gdp deflator by two mutually independent indicators and the definition of balanced economic growth is shown in fig.1. the designations : p(t) = i1/i2 − gdp deflator, where i1 = ngdp and ris. 1. ill�straci� vzaimosv�zi treh indeksov rosta. i2 = rgdp c(t)=ngdp(t)/x(t) risk management index (rmi), where x(t)= nqp + ngdp, where nqp(t) current material production costs ngdp(t). rmi is a synonym of stp coefficient or rate of scientific and technological competitiveness. pp(t)=c(t)/p(t)index of the purchasing power of money. as can be seen from figure 1, the gdp deflator (inflation index) is expanded by a factor of stp c (t) and the indicator of purchasing power of the national currency (pp (t)). by the coefficient of stp further integrated expression of risk management in manufacturing is understood, which is synonymous with the level of scientific and technological excellence of production. more precisely, the coefficient of stp is hiding all the integrated effect of all risk management decisions of the current period of analysis. main thing in theory of reversibility of indices of commodity prices and the advances in systems science and applications (2014) vol.14 no.4 315 purchasing power of money is that it opens the door to reveal the entire system of economic laws that govern the development of the market economy. thus, based on the above formula economic laws are defined, which are very useful for the stability of the market economy. so if the appropriate formulas of these laws are used in the analysis of the market economy not the market elements will dominate, but the system of economic laws by which the principle of market equilibrium is realization. and the indicators of of real and financial sectors in equal measure will participate in the analysis and management of the market economy. basic economic laws that meet the above logic, analysis, and outlined in the terms of the indices of growth of major management indicators, the following : i law of determining the overall impact of the adopted incentives for innovation and investment in the economy of scientific and technological improvement c(t) : c(t) = gdp (t)/(qp (t) +gdp (t)) i law of determining the purchasing power of money pp(t) : pp(t) = (c(t) ∗ i2(t))/i1(t) i law of determining prices of goods and services 1/pp(t) : 1/pp(t) = i1/(c(t) ∗ i2(t)) i leading law of determining the real growth of the economy i3(t) : i3(t) = pp(t) ∗ i1(t). i control law of definition of the real growth of the economy i3(t) : i3(t) = c(t) ∗ i2(t) i law of general price deflation p(t) : p(t) = c(t)/pp(t) = i1(t)/i2(t) i law of definition of net benefits from the stimulation of scientific and technological improvements -∆c(t) : ∆c(t) = c(t)− 100 it should be recalled that, in our notation, i1 (t) is the index of gdp growth in the prices of the current year, i2 (t) the index of growth of gdp at comparable prices. and qp (t) indicated the current material costs necessary for the production of gdp (gdp (t)). 316 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress if earlier political economy as the ideological basis of private or socialized economy disconnected the world into two camps, but now it is special science freed from ideology and will be used for the benefit of all, starting with the economy of the individual entrepreneur and completing with the individual economies of the world. all participants of production will have to study and use these laws in economic activity, just as natural scientists in their work using ohm’s law, or the law of gravity without regard to ideological awnings hanging over them. learning about and using them in their work is forced not the directive from above or below, and their own economic interests. they are equally beneficial and at the enterprise level and at the level of the economy. 2.2 investment in fixed assets as carriers of innovation the relationships between the key indicators of economic will change under the influence of capital investments, innovations in technology, science and other technical or organizational and technological measures. the main carrier of innovation is an investment that has a direct impact on the change in the structure of the market forces of the real economy, hence, on the purchasing power of the currency. of course, these investments will be effective if they are focused on the major discoveries in the natural sciences, which are expected in the areas of space technology and nano technology. space technology is only in the field of telecommunications has opened doors to new scientific and technological revolution, which is already being implemented in almost all human activities, ranging from educational processes, ending municipal services to the population. it is widely implemented in the management process and monetary institutions. it is expected that within the framework of e-government in the near future high technology breakthroughs of intelligent systems of state management of the economy and finance will be created. based on the current intensity of research and development work carried out at universities in the u.s. and western europe in the field of nanotechnology, breakthrough technology in the near future will be implemented in the real sector rapidly. as m. ratner and d. ratner write [11], nano technology will become the foundation of many advanced technologies. feature of this technology is that it can be developed in university laboratories and individual scientific research centers. and, in these institutions intellectual work comes to the forefront. even today, without noticing it, we use “smart materials” obtained through the use of nanotechnology. under the “smart materials” the latest developments in the field of materials consumer or industrial purposes in the nano world of electrons and neutrons are meant. the basic technology of their production is realized using the theory of quantum mechanics. from the point of view of science, here is the integration and sharing of existing knowledge of various branches of basic advances in systems science and applications (2014) vol.14 no.4 317 sciences chemical, biological, physical, mathematical and others. in developing such a “smart material” skilled professionals and financial capital are involved, the integration of which will give a new impetus to a high-performance production. sona analyzer is highly effective and innovative product that allows transferring responsibility for the managerial decisions, including, forecasting and planning at the appropriate levels of microeconomics. macroeconomic policy stops focusing on the development of planning and forecast indicators and, thus, increases the analytic function of public service staff of macroeconomic management. once management functions of microeconomics remains at a real sector, the main condition for profit maximization at the enterprise level is the dynamics of total factor productivity and attraction of additional labor. what is the reason for changes in the dynamics of total factor productivity and capital ? first of all, they are due to the intensity of capital investment in the real economy and its effectiveness. investments in fixed assets contribute to changes in the labor armament with basic production assets and return of capital. in the analyzer sona, armament of labor with fixed assets is measured by a special indicator that expresses the level of industrialization of the economy and the price of capital is measured by other special indicator that expresses the change in the level of innovation and contribution to innovations, carrier of which is the newly introduced in the production of new capacity and replacement of fixed assets. but the dynamics of changes in total factor productivity and fixed capital is associated by productivity of economic labor, which is represented by the ratio of the total factor productivity to the average annual wage per worker. these special indicators in the analysis of management decisions are of great practical meaning, since not all the “innovative activities and innovations”, in practice, increase the level of productivity of economic labor and often work in the opposite direction due to miscalculations in the stages of their design and implementation. thus, the two key indicators of interrelated capacities total factor productivity and productivity of economic labor determine the levels of innovation development of any part of the real, and after aggregation, also the national and global economy. 2.3 resource productivity of intermediate consumption as a measure of growth of production of the final product both labor productivities are arguments of the function of the third productivity the resource productivity of intermediate consumption for the production of the final product. this includes all products and services, including, first and foremost, energy, environmental, and other natural resources that are annually involved in the turnover of the current material costs of the real economy and are constantly in circulation. they are annually deducted from turnover. however, these resources of inter318 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress mediate destination reserve the mark elusive at first sight in the measurement of progress, and the money in circulation, and the mass of commodities in circulation. and the growth of the total number of employed people and other factors of increasing the working time can only affect the quantitative growth of economic indicators. productivity of intermediate product for the production of the final product (gdp) is a function, which reflects the dynamics of total factor productivity of labor and capital, hence, productivity of economic labor. as it is known, natural resources, primarily fuel and energy resources and mining and smelting clusters that are involved in the production, form the bulk of intermediate product values. therefore, this key indicator, defined by the relation between the final and intermediate products is the relationship between the natural environment and human society. and therefore the relevant performance indicator of intermediate resources of consumption is a crucial tool of economic management in a single technology of analyzer sona. what caused the changes in the productivity of intermediate product for the production of the final product (gdp) ? capital expenditures on scientific, technical, organizational, technological, and any other events, including, events aimed at resource conservation, improvement and updating of the commodity nomenclature lead to changes in productivity of the intermediate product for the production of the final product (gdp). these changes take place mainly under the influence of scientific and technological developments, including technical, technological and organizational measures that are used by certain people’s knowledge and monetary resources. but sona analyzer does not work out and does not offer a specific administrative decision and, therefore, does not control the quality of science and technology and other measures implemented in the real economy. it only establishes a de facto existing levels of three key performance indicators in the annual real time, and assesses changes in the level of armament of capital for labor (indicator of industrialization) and in the level of prices of fixed capital (the indicator of innovation) due to management decision-making and the definition of economic laws. that is, it does not replace the persons taking management solutions. subordinated system of indicators of sona analyzer is used only for audit and examination of the quality of the management decisions they made. a key indicator of the analyzer along with the number of employed people in the economy allows us to determine the quantitative parameters of the growth of its main indicators. it includes : • productivity of intermediate product for the production of the final product ; • total factor productivity of labor and capital ; • productivity of economic labor ; • the number of employed people in the economy. advances in systems science and applications (2014) vol.14 no.4 319 all other indicators needed to analyze the economic development of the country are derived from these four basic concepts and categories. official annual statistics of the national economy has background information on the definition of these key and additional indicators of analysis and management. thus, global practice of measurement of purchasing power of money is limited to assessment the physical volume index (pvi) of goods and services and the establishment of the gdp deflator in the real sector of the economy, as the ratio of the index of growth in prices of the current year to an index of growth in the prices of the base year. and the theory of balanced growth is still out of sight of management practices. until now, application tools of analysis of key sectors of the national economy are not synchronized, there are no direct and inverse relationships between indicators of the real and financial sectors, as part of a single economic system. as a result of the lack of mechanisms to control deviations of interrelated indicators of development of market forces, the gap “between those who are doing business and those who make money” is increasing. the proposed technology of work based on the analytical formulas of the system of economic laws is highly effective. and the appropriate navigation system for the analysis of planned, project and other management decisions gives concrete results in the form of a system of the control indicators based on data from official statistics. in particular, the indicator of growth of the price of a particular product or service in the real sector is clearly defined, which is the reciprocal of the purchasing power of money in relative terms. 2.4 the index of prices of goods and services as the inverse of purchasing power of money the scientific basis of the principle of reversibility as well as the principle of determining the sona analyzer is based on the use of the global economic thought and practice. it is based on a procedure that meets the basic rule of international consensus :the equality of the indices of growth of the sum of the sellerąŕs commodity prices and the reciprocal of the growth indexes of purchasing power the buyer’s sums of money. this means that all price indicators of economic analysis and management of the real economy are under the authority of the power unit of the national currency, determined, in this particular case, in accordance with the principles of international consensus : ⨿ᵀ ∗kᵀ/kπ = 1 (2) where ⨿ᵀ ∗kᵀ as the numerator of this equation expresses the value of commodity mass of the seller, and kπ is the money supply from the buyer at face value of the national currency. but this is only the rule ; “one” in the right-hand side of (1) is a relative quantity, necessary to compare purchasing power of the national currency of the 320 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress current time with its level in the base year. since the gross profit ⨿ᵀ ∗kᵀ is actually output in monetary terms -x= ⨿ᵀ ∗kᵀ, and money supply m = kπ, then velocity of money circulation at face value of the national currency will be vx : vx = x/m, or by the conditions of international consensus : x = vx ∗m. hence : x/(vx ∗m) = 1 (3) it is clear that the output x contains the repeated recording of individual components of products that are sold and resold, then nominal gdp is defined as : ngdp = x qp, where qp is intermediate consumption, and ngdp is gdp for the production, the product without double counting. since the equation of exchange of the monetarists : ngdp/vn ∗m = 1 (4) then replacing ngdp by x qp we have the equality that defines the difference between the three speeds : vx − vz = vn (5) which are clearly determined by the gross profit x= ⨿ᵀ ∗kᵀ. this rule is fully consistent with international and kazakhstani practice of analysis of nominal and real gdp and gdp deflator. introduction to the analyzer sona of three key performance indicators allows focusing the purchasing power of money at one point. and this point of consensus core of world currency can not be determined without consideration of gross profit x= ⨿ᵀ ∗kᵀ, on the basis of which three velocities of money are determined by the equation of the vn = vx − vz. the initiative group of kazakhstani analysts, after the analogical settlement of the economies of states of the customs union and the economy of the customs union as a whole, has come to the conclusion that there was an opportunity to focus their national currencies to a single point. in this case, the currency of the customs union will have its consensus core, other than the core of the national currencies of countries of the customs union. the point of this consensus itself, adopted as a basis for currency of cu, as well as world currency itself is in constant motion. thus, the price of the u.s. dollar in 2000 was equal to 0.75, in 2008 0.62, and in 2010 regained the relatively high level 0.81. similarly, as the u.s. dollar, all other national currencies will have their equilibrium prices different from world currency prices. price of world currency is needed for countries of the world to support the effectiveness of the exchange rate of its currency and the favorable choice for the country’s economic policy. in this case, the relative value of the power unit of the nominal value of world currency at the point defined by the application tool of the analyzer such as sona, will become advances in systems science and applications (2014) vol.14 no.4 321 a concentrated expression of the purchasing power of all national currencies of the states concerned un member states, including the united states. once for each country purchasing power of its currency is determined, then the basis for calculating the sdr can be taken the following formula, which does not contradict the general methodological principles of the imf’s definition of “price” of world currency : pp (sdr, t) = ∑i=n i=1 ngdp (i)/[pp(1) ∗ngdp (1) + pp(2) ∗ngdp (2) + ... +pp(n) ∗ngdp (n)] where n expresses the number of countries included to the zone of foreign trade. the cost of one million of kazakhstan tenge, as well as other national currencies in their official rate, which were included in the determination of the value of world currencies, is an equivalent product, and can be freely exchanged for a certain amount of other goods and services at their market prices. but its purchasing power, as well as other national currencies, undergoes a change with changes in the velocity of circulation not just one product, say, gold or oil. it varies with the velocity of circulation of all goods and services, including the prices of intermediate goods and services. 2.5 function of the velocity of circulation of goods and services in determining the rate of balanced growth all indicators of the analyzer sona are a function of main performance indicator of intermediate product for the production of the final product, which determines the economic content of the relative velocities of the key indicators of macro-and microeconomics. they include : • dynamics of intermediate product performance (qp) for gdp at face value (gdp(t))−gdp (t) qp (t) = µ(t), the reciprocal of material consumption of gdp ; • rate of stp, the value of which is equal to gdp(t)/x(t) in macroeconomics in this case is expressed by formula — c(t) = µ(t) 1+µ(t) . this indicator defines the contribution of innovation factors in the development of the economy and serves as a general indicator of technical, technological and other organizational measures implemented in the real economy as a whole ; • rate of turnover (x) on the face value of gdp, in this case, is defined by the formula — vx(t) = x(t) gdp (t) = 1+µ(t) µ(t) = 1 c(t) ; • turnover rate of the money supply m on face value of gdp is expressed by the formula — vn(t) = x(t) m(t) ∗ ( µ(t) 1+µ(t) ) = gdp (t) m(t) . as can be seen from this equation turnover rate indicator of the money supply does not take into account the contribution of all innovations in the real economy. • turnover rate of the money supply vn from the perspective of the circulation of goods and services is determined by a different formula : vn(t) = vx(t)− vz(t) 322 sailau baizakov, jeffrey yi-lin :the paradoxes of the world’s progress where vzthe turnover rate of the current material costs used in the production of each unit sold goods and services is determined by the formula — vz(t) = qp (t) m(t) . • the purchasing power of money is determined directly proportional to the ratio of ntp and inversely proportional to the gdp deflator : pp(t) = c(t)/p(t) • balanced economic growth, defined by multiplication of the purchasing power of money and nominal gdp : rngdp (t) = pp (t) ∗ngdp (t) • balanced economic growth, defined by multiplication of the coefficient ntc and real gdp : nrgdp (t) = c (t) ∗rgdp (t) • the principle of invertibility : nrgdp (t) = rngdp (t) or pp (t) ∗ngdp (t) = c (t) ∗ngdp (t) a key indicator in determining the efficiency of the financial sector of the economy is the value of money, expressed as purchasing power of the currency. the value of money, as the prices of goods and services, changes over time. but the value of money does not contain a single atom of natural ingredients : no gold, no coal, no oil, no labor, no capital, no bread, no meat and other goods. and therefore it makes no sense to look for another form of money, other than paper form. there are no prospects, for example, of the return to the gold standard, the more no prospects of energy (watt, joule) and other material substitutes for paper money. only deviation of index of prices of goods and services relative to the gdp deflator can accurately estimate the purchasing power of money, and set the real price and exchange rates of the national currencies of the world. its formula is defined from equation p(t) =c(t)/pp(t), where p(t) is gdp deflator, c(t) is the coefficient of stp and pp(t) is an indicator of purchasing power of money. analyzer sona, as the technology of management of unity development of the real, financial and monetary sectors of the economy is used to measure the purchasing power of money and the gap between the real and nominal gdp. as a result, the third dimension of economic growth is defined as the product of nominal gdp and purchasing power of money (pp(t)*ngdp) and a double of this measurement, defined by the product of real gdp of and the stc ratio advances in systems science and applications (2014) vol.14 no.4 323 (c(t)*rgdp). exactly this dual pair is the economic content of the principle of mutual convertibility, with the help of which the indices of growth of prices of goods and services (1/pp(t)) are defined as the reciprocal of the index of the purchasing power of money pp(t). the analytical model for determining the threshold levels of indicators of economic management based on official statistics of gdp deflators (p (t)) and on the human dimension of the purchasing power of money is composed of the following recursive system of equations : ngdp (t) = ngdpn (t) · ln (t) ·n(t). tw (t) = γ(t) · ln (t) ·n(t). tr(t) = ngdp (t)− tw (t). x(t) = ( 1 u(t) + 1) · q(t) · γ(t) · ln (t) ·n(t). c(t) = gdp (t) x(t) . p(t) = ngdp (t) ngdp (0) / rgdp (t) rgdp (0) . pp(t) = c(t) p(t) . pp(t) ∗ngdp (t) = c(t) ∗rgdp (t). (6) additionally accepted designations : n number of the country’s population,ln the proportion of people employed in the economy of the total population,tw -fund salaries,tr income on equity, ngdpn gdp per capita. equality between gdp by income, gdp by production and gdp by end-use remains in this analytical model. pp(t) ∗ngdp (t) = c(t) ∗rgdp (t). references [1] n. nazarbayev. (2009), “fifth way” , izvestia, sep. pp.22. [2] n. nazarbayev. (2009), “ keys of crisis”, russian newspaper, feb. pp.2. [3] n. nazarbayev. (2009), speech at a business forum in india, respublika.kz,no.3, jan. pp.27. [4] n. nazarbayev. (2013), speech on th vi astana economic forum, may. pp.23. [5] e-book http://www.monographies.ru/ 324 sailau baizakov, jeffrey yi-lin:the paradoxes of the world’s progress [6] n. kaldor, g. a. mirrlees. (1962), “a new model of economic growth”, review of economic studies, june, pp.174-192. [7] kaldor-mirrlees model. (1976), “in the book of a. pesenti”, essays on the political economy of capitalism, 2, translated from italian, moscow: progress publishers, pp.837-870. [8] s. baizakov. (1985), “scientific and technological progress a key lever of production efficiency (analysis of the theory of economic growth and efficiency)”, almaty, pp.89. [9] “scientific and technological progress a key lever of competitiveness and economic growth: 2nd ed”, ext, almaty: nc sti, 2007, pp.76. [10] “methodological guidelines on statistics. agency of the republic of kazakhstan on statistics”, astana, 2009, pp.198. [11] m. ratner, d. ratner. (2004), “nanotechnology: a simple explanation of the next brilliant idea, translated from english”, moscow: publishing house “williams”, pp.240 [12] e-book http://www.monographies.ru/ [13] j. d. sachs, f. b. laren. (1996), “macroeconomics” ,business, pp.52-54. [14] v. l. makarov. (1985), “on the performance of scientific and technical progress”, economics and math. methods, vol.xxi, no.2. [15] “theory of value: statistical verification. informational generalization. actual conclusions”, “herald of the ras” magazine, no.9, 2005. [16] k. k. valtuh. (1965), “public utility of product and labor costs of its production”, m: thought. [17] k. marx and f. engels. op. t. 2. pp.142. [18] editor v.s. dadayan. (1973), “modeling of national economic processes”, moscow: economics, pp.433-438. [19] v. l. makarov, a. p. torzhevsky. (1986), “effect of changes in the technological level of production on macro indicators of economic development”, moscow:cemi as ussr, pp.4-7. [20] b.m. shtulberg, e.g. chistyakov, v.v kotilko etc. (1988), “problems and methods of study of the territorial plans”, moscow:science, pp.192. corresponding author sailau baizakov can be contacted at: baizakov37@mail.ru. adv syst sci appl 2017; 4; 22-33 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/255 hybrid papr reduction scheme for universal filter multicarrier modulation in next generation wireless systems pooja rani, silki baghla, dr. himanshu monga* j.c.d.m. college of engineering, sirsa, haryana, india e-mail: poojamehta0193@gmail.com, silky.er@gmail.com, himanshumonga@gmail.com abstract: universal filter multi carrier (ufmc) is one of the promising multi carrier modulation techniques for next generation wireless communication systems. ufmc seems to be most attractive because it provides better sub carrier separation like fbmc (filer bank multi carrier) and less complexity like ofdm (orthogonal frequency division multiplexing). but this technique suffers from limitation of higher peak to average power ratio (papr). in this paper a hybrid papr reduction technique scufmc have been proposed using slm (selective mapping) and clipping. the performance of proposed technique is evaluated for various design parameters including filter length, fft size and bits per sub carrier. the simulation results show that hybrid technique provides better papr reduction as compared with conventional slm and clipping techniques. keywords: ufmc, ofdm, fbmc, papr, slm. 1. introduction orthogonal frequency division multiplexing (ofdm) is the most popular multi-carrier modulation technique which is being used in 4th generation wireless communication [1]. but in the last few years, the number of users and the demand for higher data rates has increased exponentially so, next generation wireless communication systems must be able to deal with large numbers of users and provide a much higher data transmission rate using less complex systems. in order to serve all these requirements, various new multi carrier modulation techniques like filter bank multi carrier (fbmc), universal filter multi carrier (ufmc) and generalized frequency division multiplexing (gfdm) have been introduced [2,3]. in fbmc, each subcarrier is individually filtered and provides robustness against inter-carrier interference (ici) effects [4]. however, fbmc systems utilize filters, whose length is multiple times of samples per multi-carrier symbol resulting in increased complexity of the system. universal filtered multi-carrier (ufmc) is a novel multi-carrier modulation technique, which combines the features of fbmc and ofdm. ufmc filters groups of subcarriers instead of per subcarrier like fbmc or a complete signal in a single shot like ofdm. this allows reducing the filter length considerably as compared to fbmc. so, it is less complex like ofdm and provides better subcarrier separation like fbmc [5]. the main drawback of all these multicarrier modulation techniques is high peak to average power ratio [7,8]. various papr reduction techniques have been proposed in literature and are implemented on ofdm and fbmc. the conventional techniques for papr reduction include selective mapping (slm), companding, * corresponding author: himanshumonga@gmail.com http://ijassa.ipu.ru/ojs/ijassa/article/view/255 mailto:poojamehta0193@gmail.com mailto:silky.er@gmail.com mailto:himanshumonga@gmail.com mailto:himanshumonga@gmail.com 23 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) tone reservation, clipping and filtering, partial transmit sequence (pts) and active constellation extension (ace) [9]. these techniques have been implemented with orthogonal frequency division multiplexing, but little efforts have been done for papr reduction in universal filtered multi-carrier [10]. clipping is the simplest and extensively used method for papr reduction. selective mapping is a suitable match for the subcarrier nature of ufmc. in this paper a novel hybrid papr reduction technique sc-ufmc for ufmc systems has been proposed. this technique is implemented by using slm and clipping papr reduction techniques. it is observed that this hybrid technique provides better results when compared with the individual performances of slm and clipping techniques. 2. ufmc waveform generation for next generation wireless communication systems, a new waveform is required which should achieve the asynchronous reception and transmission, non-orthogonal waveforms for better spectral efficiency and low latency. ufmc has been introduced as a new waveform design representing a generalization of this principle targeting to collect the advantages while avoiding the disadvantages of other modulation techniques [11,12]. ufmc is the method which combines the advantages of orthogonality of ofdm and concept of filter bank in fbmc. instead of filtering each carrier like in fbmc, blocks of carriers called sub-bands are filtered. each subband contains a number of carriers and the filter length will depend upon the width of the subband [13]. fig. 2.1 shows the process of transmission and reception in a ufmc system. here, the complex symbols generated from the modulator (qpsk or qam) are applied to serial to parallel converter resulting in a block of streams and fed as input to their respective ifft. the length of n points ifft output is converted back to serial per block and that output will be filtered with a pulse shaping filter of length l. the transmitted signal can be given as: 𝑥𝑘 = ∑ 𝐹𝑝,𝑘𝑉𝑝,𝑘 𝐵 𝑝=1 𝑋𝑝,𝑘 (2.1) as shown in fig. 2.1, the overall bandwidth and the total number of subcarriers are divided into number of sub bands. now, the input data stream of kth user, 𝑥𝑘 is divided into multiple sub streams denoted by 𝑋𝑝.𝑘 for pϵ{1,…, b}. here b is the total number of sub streams or sub bands. a sub band in ufmc may also correspond to physical resource block (prb) in lte. then, the signal of each sub band is applied to individual n point ifft represented by matrix v. the output of ifft is converted to the serial form and applied to the respective filter represented by matrix f. f is a toeplitz matrix, composed of the filter impulse response, performing the linear convolution in equation (2.1) [14]. for pth sub band, where p varies from 1 to b, 𝑋𝑝,𝑘 , 𝑉𝑝,𝑘 and 𝐹𝑝,𝑘 represents data blocks, ifft matrix and the filter respectively. for filter length l and ifft matrix of size n, the symbol xk of duration n+l-1 is generated at the output of transmitter section. hybrid papr reduction scheme for universal filter multi-carrier modulation 24 copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 2.1. ufmc transceiver fig. 2.2 shows the spectrum for ufmc waveform with sub bands b=10 and each sub band is having 20 sub carriers. the complete process of data transmission and reception is shown in fig. 2.1. at the receiver, after passing through rf-link section the signal is applied to time domain pre-processing window to suppress interference. after windowing, the signal will be converted into 2n parallel streams; here n is the number of subcarriers. the demodulated signal is sent to the de-mapper, which is a demodulator to retrieve the data bits from the received symbols. the generated waveform is shown in fig. 2.2. the block-wise filtering provides flexibility to the system and may be used to avoid the main drawbacks of fbmc. ufmc supports short bursts data transmission, as well as operation in fragmented bands. the filter provides protection against inter-symbol interference (isi), as well as robustness for supporting multiple access users which are not perfectly time-aligned. due to the possibility to reduce guard bands, and to avoid need of cp, ufmc is spectrally more efficient than cp-ofdm [15]. serial to parallel converte r idft filter filter p/s p/s p/s idft idft filter base band to rf channel + noise rf to baseband time domain preprocessin g (e.g. windowing) + s/p fft 2n point symbol demapping complex symbols 25 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 2.2. ufmc waveform the receiver processing can still be similar to cp-ofdm, single-tap per-subcarrier frequency domain equalizers can be used which equalize the joint impact of the radio channel and the respective sub band-filter. this leads to similar complexity order as cp-ofdm. so, it is thereby clear that ufmc provides advantages of both ofdm and fbmc system. 3. peak to average power ratio ufmc has numerous advantages over other modulation techniques, but it also suffers from high peak to average power ratio (papr).the papr is the relation between the maximum power of a sample in a given transmitted symbol divided by the average power of that symbol. papr occurs when in a multicarrier system the different sub-carriers are out of phase with each other. there are a large number of independently modulated subcarriers in a multicarrier system which are different with respect to each other at different phase values. when all the subcarriers achieve the maximum value simultaneously, this will cause the output envelope to suddenly increase which causes a 'peak' in the output, and when they are added up coherently for transmission purpose give a large peak value which is very large as compared to average value of the sample. the ratio of the peak to average power value is termed as peak-to-average power ratio [16]. the mathematical valuation of papr is defined in equation (3.1). papr = ( 𝑚𝑎𝑥{|𝑥[𝑛]|2} 𝐸{|𝑥[𝑛]|2} ) (3.1) where, |x[𝑛]| is the amplitude of x[n] and e denotes the expectation of the signal. this higher papr causes saturation in the power amplifier which produces inter modulation products among sub bands and also increases out of band radiation (oob). -0.2 -0.1 0 0.1 0.2 0.3 0.4 0.5 -120 -100 -80 -60 -40 -20 0 ufmc normalized frequency p s d (d b w /h z ) hybrid papr reduction scheme for universal filter multi-carrier modulation 26 copyright ©2017 assa. adv. in systems science and appl. (2017) 3.1.conventional papr reduction techniques there are various papr reduction techniques which can be used to improve performance of multicarrier modulation systems. some of them are: 3.1.1. selective mapping (slm) technique in slm, from the original data block several candidate data blocks are generated and all the data blocks are having the same information. after this a phase rotation is applied to each block and passed from its respective idft and a block with minimum papr is selected for transmission [17]. 3.1.2. companding companding is an easy and less complex method of papr reduction, the basic idea is to expand the small signal in the transmitter section and compression is carried out at the receiver side. in this technique, we enlarge the small signals while compressing the large signals to increase the immunity of small signals from noise. this compression is carried out at the transmitter end, after the output is taken from ifft block. there are two types of companders: μ−law and a-law companders [18]. 3.1.3. partial transmit sequence in pts original data block is partitioned into n disjoint sub blocks. the subcarriers in each sub block are rotated by the same phase factor such the papr of the combination can be minimized. pts scheme reduces papr with some additional complexity and it also affects spectral efficiency of the system because side information is also required to be transmitted. it does not produce any distortion in system [19]. 3.1.4. clipping and filtering this is one of the simplest techniques for papr reduction. the principle is to define a clipping level for data transmission above which the input signal is clipped off and peaks of signal are reduced [20]. let, there is a signal y[𝑛] which is to be transmitted and 𝑦𝑐[𝑛] is its clipped version which can be denoted as: 𝑦𝑐[𝑛] = { −a 𝑦[𝑛] ≤ −a 𝑦[𝑛] |𝑦[𝑛]| < a a 𝑦[𝑛] ≥ a (3.2) where a is the clipping level. after clipping out of band radiations are produced this can be reduced by using filtering after clipping. 3.1.5. tone reservation in this scheme, some subcarriers are reserved within the transmitted bandwidth and appropriate value is assigned to these reserved tones [21]. these reserved subcarriers don’t carry any data information, are only used for reducing papr. 3.1.6. active constellation extension (ace) in ace, at each block, some of the outer signal constellation points are extended towards outside of the constellation such that the papr of the resulting block is reduced. it is transparent to the receiver. there is no loss of data rate and no side information is required. however, all these papr reduction techniques were introduced for ofdm systems and due to different frame structure of ufmc signal, these techniques cannot be effectively utilized in ufmc systems. 27 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) 4. proposed hybrid scheme various papr reduction techniques introduced for ofdm system are discussed in the previous section and we tried to apply all these techniques on ufmc systems. it has been observed that slm and clipping are suitable papr reduction techniques for ufmc systems and can reduce papr very effectively with a small increase in system complexity. on the other hand, schemes like pts, ace and tr makes the system very complex and it is very difficult to apply these techniques on the frame structure of ufmc systems. in this work, we have proposed a novel hybrid technique for papr reduction in ufmc by using slm and clipping. unlike existing slm and clipping schemes for ofdm systems, the proposed method exploits the nature of sub-block data transmission of ufmc. in ufmc data is generated in the form of sub blocks which are groups of subcarriers. the advantages of slm and clipping have been exploited in this hybrid technique. in clipping some distortion is produced but in slm there is no any distortion while on the other hand slm produces data rate loss but in clipping there is no any data rate loss. so, in this work we have combined these two schemes so that we can exploit the advantages of both schemes. both techniques have the advantage that the power of the system doesn’t increase. hence, we can decrease papr with the same power which is used by the system when no papr reduction technique is applied. 4.1. hybrid (sc-ufmc) papr reduction technique the basic block diagram of hybrid technique is shown in fig. 4.1. the waveform generated by the ufmc modulator is given to the serial to parallel converter where several candidate blocks are generated from original data block. this is done to find a block with minimum papr for transmission. after generating candidate blocks a phase rotation is applied to each block and applied to ifft of their respective. because of the varying assignment of data to the transmit signal, it is called selective mapping. the core is to choose a particular signal which is having desired properties out of n signals representing the same information. finally, we select a block with minimum papr, the signal generated by this selector is applied to clipper for removing the higher peaks. for this a perticular thershold value is defined above which all the signal is clipped of so that peak to average power ratio can be reduced. the signal generated by ufmc is given in equation 2.1 and after applying a phase rotation the signal become as given in equation (4.1). 𝑥 = ∑ 𝐹𝑃,𝑘𝑉𝑝,𝑘 𝐵 𝑝=1 𝑋𝑝,𝑘𝑝(𝑛) (4.1) fig. 4.1. block diagram of hybrid (sc-ufmc) technique clipping data from ufmc modulator serial to parallel conversion select sequence with minimum papr ifft p (n) ifft p (1) ifft p (2) hybrid papr reduction scheme for universal filter multi-carrier modulation 28 copyright ©2017 assa. adv. in systems science and appl. (2017) where, 𝑝(𝑛) denotes the phase rotation of the signal, after this signal after fft is applied to clipper. let the threshold for the signal is a. then the final signal will be: 𝑥𝑐 = { 𝑥, |𝑥| < 𝐴 𝐴, |𝑥| > 𝐴 (4.2) hence, by using this technique we selected a block with minimum papr and then a clipper circuit is used to clip of the peaks of that block so that we can reduce peak to average power ratio of signal to a great extent. 5. simulation setup and results to compare the proposed hybrid (sc-ufmc) technique with conventional slm (s-ufmc) and clipping (c-ufmc) techniques, matlab platform has been used. all the simulation work, including comparison and performance evaluation of proposed technique with variation in design parameters was done by using matlab platform. 5.1. simulation setup table 5.1 provides the simulation set up to evaluate the performance of proposed papr reduction technique. the proposed technique is also compared with three other techniques as original (ufmc), with selective mapping (s-ufmc), with clipping (c-ufmc) and hybrid (sc-ufmc) to analyze the effectiveness in papr reduction. the performance of the proposed method is evaluated and compared on the basis of variation in fft size, bits per sub carrier, filter length and modulation order. table 5.1. simulation setup parameter values fft size 2048, 1024, 512, 256 sub band size 20 number of sub bands 10 modulation order qam (4, 16, 64) bits per sub carrier 2, 4, 6 filter length 43, 63, 83 in fig. 5.1 peak to average power ratio of ufmc, slm scheme (s-ufmc), clipping scheme (c-ufmc) and hybrid scheme (sc-ufmc) with fft size 1024, filter length 43, bits per subcarrier 2 and modulation order 4 are shown. it can be concluded that the proposed hybrid schemes have improved papr reduction performance as compared with the two conventional schemes for papr reduction. 29 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 5.1. papr reduction of different schemes 5.2. performance evaluation of sc-ufmc with variation in design parameters this section presents performance evaluation of sc-ufmc with variation in design parameters (fft size, bits per sub carrier and filter length). 5.2.1. variation in fft size fig. 5.2 shows the papr of ufmc signal and performance of three papr reduction techniques at fft size 2048. it can be observed that sc-ufmc is providing better papr reduction than other two techniques. fig. 5.2. performance of various papr reduction techniques with different fft sizes -60 -40 -20 0 20 40 60 80 0 0.2 0.4 0.6 0.8 1 papr in db ufmc s-ufmc c-ufmc sc-ufmc 0 5 10 15 20 25 ufmc s-ufmc c-ufmc sc-ufmc p a p r in d b papr reduction techniques papr with varition in fft size 2048 1024 512 256 hybrid papr reduction scheme for universal filter multi-carrier modulation 30 copyright ©2017 assa. adv. in systems science and appl. (2017) as shown in fig. 5.2, c-ufmc shows better performance than s-ufmc, so effectiveness of sc-ufmc is analyzed with different values of design parameters with respect to c-ufmc. table 5.2 is showing that sc-ufmc performs more effectively with larger fft size. table 5.2. effectiveness with variation in fft size 5.2.2 variation in bits per sub carrier performance of various papr reduction techniques with 2 bits per sub carrier is shown in fig. 5.3 from here it can be observed that proposed scheme is showing better results than conventional schemes. all the schemes are showing different values when we change any of the design parameters as shown in fig. 5.3, here we are changing bits per sub carriers. fig. 5.3. performance of various papr reduction techniques with different bits per sub carrier as shown in table 5.3 with increase in bits per sub carrier, sc-ufmc is being more effective and at 6 bits per sub carrier it is having maximum papr reduction. table 5.3. effectiveness with variation in bits per sub carrier 0 5 10 15 20 25 ufmc s-ufmc c-ufmc sc-ufmc p a p r in d b papr reduction techniques papr with variation in bits per sub carrier 2 4 6 fft size % effectiveness of sc-ufmc 256 4.23% 512 27% 1024 28.15% 2048 30.14% bits per sub carrier %effectiveness of sc-ufmc 2 30.14% 4 37.62% 6 58.60% 31 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) 5.2.3 variation in filter length with changing filter length, papr of original ufmc signal as well as performance of papr reduction techniques also get affected as shown in fig. 5.4. fig. 5.4. performance of various papr reduction techniques with different filter lengths it can be observed from table 5.4 that at higher filter length proposed scheme performed less effectively and a better papr reduction can be obtained with lower filter length. table 5.4. effectiveness with variation in filter length filter length %effectiveness of sc-ufmc 43 36.79% 63 14.57% 83 8.40% 6. conclusion in this paper, a novel hybrid scheme sc-ufmc for papr reduction of ufmc signals is proposed. simulation results and analysis shows that the hybrid scheme is an efficient papr reduction method for ufmc systems, and it can provide better papr reduction performance than the conventional slm and clipping schemes. further, the effect of various design parameters on papr reduction in sc-ufmc has been analyzed and it is concluded that the proposed technique provides effective papr reduction than conventional schemes. references [1] 5gnow, (2013, november 21) d3.1: 5g waveform candidate selection. tech. rep. [online]. available: http://5gnow.eu/wp-content/uploads/2015/04/5gnow_d3.1_v1.1_1.pdf [2] rohde & schwarz (2016). 5g waveform candidates. application note. [online]. available: https://www.rohde-schwarz.com/us/applications/5g-waveform-candidatesapplication-note_56280-267585.html 0 5 10 15 20 25 ufmc s-ufmc c-ufmc sc-ufmc p a p r in d b papr reduction techniques papr with variation in filter length 43 63 83 hybrid papr reduction scheme for universal filter multi-carrier modulation 32 copyright ©2017 assa. adv. in systems science and appl. (2017) [3] schaich, f. & wild, t. (2014). waveform contenders for 5g: ofdm vs. fbmc vs. ufmc. 2014 6th international symposium on communications, control and signal processing (isccsp), ilissia athens, greece, 457–460. [4] farhang-boroujeny, b. (2011). ofdm versus filter bank multicarrier. ieee signal process. mag., 28, 92–112. [5] vakilian, v., wild, t., schaich, f., ten brink, s. & frigon, j.-f. (2013). universalfiltered multi-carrier technique for wireless systems beyond lte. 2013 ieee globecom workshops (gc wkshps), atlanta, ga, 223–228. [6] wunder, g., jung, p., kasparick, m., wild, t., schaich, f., et. al. (2014). 5gnow: nonorthogonal, asynchronous waveforms for future mobile applications. ieee communications magazine, 52(2), 97–105. [7] gangwar, a. & bhardwaj, m. (2012). an overview: peak to average power ratio in ofdm system & its effect. international journal of communication and computer technologies, 1(2), 22-25 [8] mehta, p., baghla, s., monga, h. (2016) papr reduction methods for multicarrier modulation schemes used in next generation wireless networks – a review. international journal of broadband cellular communication, 2(2), 35-44. [9] sharma, s., kumar gaur, p. (2015). survey on papr reduction techniques in ofdm system. international journal of advanced research in computer and communication engineering, 4(6), 271-274. [10] praneeth kumar, p., krishna kishore, k. (2016). ber and papr analysis of ufmc for 5g communication. indian journal of science and technology, 9(s1), 1-6. [11] chen, y., schaich, f. & wild, t. (2014). multiple access and waveforms for 5g: idma and universal filtered multi-carrier 2014 ieee 79th vehicular technology conference (vtc2014 spring), seoul, korea, 1–5. [12] wild, t., schaich, f., chen, y. (2014). 5g air interface design based on universal filtered (uf-) ofdm. proceedings of the 19th international conference on digital signal processing, 699-704. [13] an, c., kim, b., & ryu, h.-g. (2016). waveform comparison and nonlinearity sensitivities of fbmc, ufmc and w-ofdm systems. eighth international conference on networks & communications, sydney, australia, 83– 90. [14] dahlman, e., parkvall, s., &. skold, j (2011). 4g: lte/lte-advanced for mobile broadband: lte/lte-advanced for mobile broadband. oxford, uk: academic press. [15] rani, p., baghla, s. & monga, h. (2017). performance evaluation of multi-carrier modulation techniques for next generation wireless systems. international journal of advanced research in computer science, 8(5), 508-511. [16] tellado, j., & cioffi, j. m. (1999). papr reduction in multicarrier transmission system. ph.d. thesis, stanford university. [17] bauml, r. w., fischer, r.f.h. & huber, j.b. (1996). reducing the peak to average power ratio of multi carrier modulation by selected mapping. ieee electronics letters, 32(22), 2056-2057 [18] wang, x., tjhung, t.t., ng, c.s. (1999). reduction of peak to average power ratio of ofdm system using a companding technique. ieee transactions on broadcasting, 45(3), 303-307. 33 p. rani, s. baghla, h. monga copyright ©2017 assa. adv. in systems science and appl. (2017) [19] muller, s.h. & huber, j.b. (1997) ofdm with reduced peak-to-average power ratio by optimum combination of partial transmit sequences. ieee electronics letters, 33(5), 368-369. [20] o'neill, r. & lopes, l.b. (1994). performance of amplitude limited multitone signals. proceedings of ieee 44th vehicular technology conference, stockholm, sweden, 1675-1679. [21] krongold, b.s. &. jones, d.l. (2002). a new tone reservation method for complex baseband par reduction in ofdm system. 2002 ieee international conference on acoustics, speech and signal processing (icassp), orlando, fl, usa, 2321-2324, doi: 10.1109/icassp.2002.5745110 microsoft word 18 xiushan wang, youzhou jiao, yongchang yu--thermal error modeling analysis and compensation based on the gre advances in systems science and applications (2011), vol.11, no.3-4 339-347 issn 1078-6236 international institute for general systems studies, inc thermal error modeling analysis and compensation based on the grey system theory on two turntable 5-axis nc machine tools xiushan wang, youzhou jiao and yongchang yu college of mechanical & electrical engineering, henan agricultural university, zhengzhou china abstract in this paper a novel concept of thermal error mode analysis is proposed in order to develop a better understanding of the thermal deformation on two turntable 5-axis nc machine tools. the correlation analysis theory is used as finding the correlation between temperature data and thermal errors. four key temperature points of two turntable 5-axis nc machine tools were obtained based on the temperature field analysis. a novel thermal error model on the basis of grey system theory was proposed. the experiments proved that there was averagely a 40% increase in machine tool precision. keywords thermal error, modeling, grey theory, compensation 1. introduction because inaccuracy of machine tools is a major source of workpiece errors, control of machine error sources is critically important. among the sources of machine error, thermally induced errors and geometric errors are known to be key contributors [1]. these errors could be reduced with structural improvement of the machine tool itself through design and manufacturing technology. however, there are, in many cases, physical limitations to accuracy improvement that cannot be overcome solely by production and design techniques. therefore, error compensation is a cost effective way to solve the problem [2]. accurate modeling of errors is a key part of error compensation. the thermal errors of a machine tool origin from the non-linear and time-varying thermal deformations caused by the non-uniform temperature variations in the machine structure. the temperature variations are related to the heat source location, heat source intensity, thermal resistance coefficient and the machine system configuration. therefore, the thermal error model is usually achieved by non-linear empirical modeling approaches which correlate machine thermal errors to temperature measurements of a machine. in this paper a novel concept of thermal error modeling analysis is proposed in order to develop a better understanding of the thermal deformation on two turntable 5-axis nc machine tools. the thermal error prediction model, which based on grey system theory, was built on the x, y, and z directions. the prediction precision of model is very high. the robustness of model is very strong. experiments proved that the model is more optimal model among available thermal error modeling of machine tools. 2. temperature field identification and thermal error measurement of machine tools 2.1 sensor location distribution on two turntable 5-axis nc machine tool two turntable 5-axis nc machine tools is a precise machine tool, with high geometric and position accuracy. so, the thermal error is the important factor affecting the machine tool accuracy. to detect the temperature field of two turntable 5-axis nc machine tools, a total of 24 temperature sensors were installed on the machine tools as shown in fig. 1. the locations of the temperature sensors can be divided into 8 groups [1], as shown in table 1. as shown in fig. 1, the spindle house of the machine tools is installed on the z direction slide, so the temperature field from the area is very complicated. in order to analyze the complex area, the more temperature 340 wang: thermal error modeling analysis and compensation based on the grey system theory on two turntable 5-axis nc machine tools sensors are applied. table 1 temperature sensor location & function group number sensor serial installment place & function 1 1,2,3,4,5,6,7,19,23 measuring the temperatures of the ball screw bearings and nuts of z axis 2 8,12,15,18 measuring the temperatures of the machine tool bed 3 21,22,24 measuring the temperatures of the spindle housing 4 10,11 measuring the temperatures of the ball screw bearings and nuts of x axis 5 13,14 measuring the temperatures of the ball screw bearings and nuts of y axis 6 9,17 measuring the temperatures of the ball screw bearings and nuts of a axis 7 16 measuring the temperatures of the ball screw bearings and nuts of c axis 8 20 measuring the ambient temperature fig.1. sensor locations of the thermal sensing system 2.2 temperature field under a cutting cycle in order to investigate the thermal behavior of the machine tool, simulation working conditions has been tested and the temperature field was recorded by the compensation control and temperature sensing systems. firstly, simultaneously moving the x, y, and z axes at low speed (800mm/min), with a spindle speed of 1500rpm for 2.5 hours. secondly, stopping and cooling down the machine for 1.5 hours. thirdly, moving the x, y, and z axes at high speed (1200mm/min) for another 2.5 hours, with a spindle speed of 3000rpm at the same time. in addition, during the testing moving a and c axes with 50rpm. advances in systems science and applications (2011), vol.11, no.3-4 341 fig. 3(a)–(d) show the temperature variations of the 8 groups of temperature sensors over time series respectively. fig.2. measurement of thermal errors for machine tool 2.3 thermal error measurement of machine tool the experimental set-up for measuring thermal errors is shown in fig. 2. three displacement sensors mounted on the turntable were used for measuring the thermal drift errors of the spindle in the x, y, and z directions. the inclination error was neglected because the workpiece is short. measurement results of thermal errors can be seen in fig. 4. 2.4 measurement result analysis of temperature and thermal error the conclusions are drawn from figs. 3figs.4: (1) temperature change is slow with machining process. the highest measuring temperature is about 45 0c . (2) the temperatures of the spindle are about 13 0c during the foregoing 2.5 hours and 21 0c the latter 2.5 hours. the temperatures of the nut and screw bearing of the x, y, a, and c-axis rise about 10 0c during the foregoing 2.5 hours and 18 0c the latter 2.5 hours. (3) the thermal error range is about 28 mμ , which is larger than the expected value. (4) the part error decreases with the machine temperature rise but having some time-lag. (5) when stopping cutting, the machine tool temperature does not drop immediately. thermal error does not decrease with stopping machine tool. 2 4 3 1 5 25 30 35 40 45 0 0 1 2 5 6 73 4 run2.5 hours stop 1.5 hours stop 6 t/sec 3106.3 ×× t/ ( ) run2.5 hours (a) 1-6 342 wang: thermal error modeling analysis and compensation based on the grey system theory on two turntable 5-axis nc machine tools 3106.3 ×× (b) 7-12 25 30 35 40 45 0 0 1 2 5 6 73 4 t/sec 3106.3 ×× t/ (º ) cutting 2.5 hours stop 1.5 hours cutting 2.5 hours stop (c) 13-18 3106.3 ×× t/ ( ) (d) 19-24 号 fig.3. temperature field and temperature collection systems 3. choice of the optimal temperature measuring point the amount of the temperature variable is a key factor to the accuracy of the thermal error model. too fewer amount induces that the accuracy of the thermal error model is very low, which can’t insure compensation precision. however, too much amount leads that the calculating time and additional cost are high. therefore, the temperature variable amount should be determined before the thermal error modeling. in this study, the temperature variables are determined based on correlation analysis between the thermal errors and temperature variables [3,4]. the optimal sensitive temperature measurement point should meet two factors (1) one is that the location should lie in the thermal sources (2) the other is that the correlation between the advances in systems science and applications (2011), vol.11, no.3-4 343 thermal errors and temperature variables is highest. correlation analysis is one of the most effective data processing methods on mathematic statistics, which describe correlative degree among multi-variables. the correlation among variables is often expressed by related coefficient ,x yρ : ( , ) ( ) ( )xy cov x y d x d x ρ = ⋅ (1) 33.6 10× × fig.4. output of displacement sensors from three directions table 2 the related coefficients among the temperature and error sensor coefficient error t21 t2 t24 t15 x 0.693909 0.985567 0.704328 0.969573 y 0.622667 0.710333 0.983336 0.958945 z 0.987637 0.692565 0.732559 0.985667 in this article, the related coefficients between temperature and thermal error variables are acquired through data processing program and matlab software. through calculating and analyzing the related coefficients between 24 temperature data series and x, y, and z direction thermal error series, the correlation coefficients 21, 2, 24, 15,, , ,x x x xρ ρ ρ ρ , as shown in table 2, are more sensitive than other points. what’s more, the thermal deformations of installment locations of four temperature measuring points seriously affect the machining precision of machine tool in the x, y, and z directions [1,3]. so the data series variable from 21, 2, 24, 15 temperature sensors are used for the thermal error modeling. 4. error modeling based on grey system theory 4.1 data prediction and process of grey system theory if (0){ ( )}( 1, 2, , )x i i n= is an original data series [5], new data series can be gotten by once accumulation calculation: (1) (0)( ) ( ) 1 k x k x j j = ∑ = 1, 2, ,k n= (2) then the differential equation relative to (1,1)gm model is: (1) 2 (1)dx ax u dt u v ∂ ω + = ∂ ∂ (3) 344 wang: thermal error modeling analysis and compensation based on the grey system theory on two turntable 5-axis nc machine tools here, a is development grey parameter; u is the endogenous control grey parameter. it is assumed that a is the estimating parameter, a a u ⎡ ⎤ = ⎢ ⎥ ⎣ ⎦ is acquired by the least square method. ( ) 1t t na b b b y − = (4) here ( ) ( ) (1) (1) 1 1 (1) (1) 1 (1) (2) 1 2 1 (2) (3) 1 2 1 ( 1) ( ) 1 2 x x x x b x n x n ⎡ ⎤⎡ ⎤− +⎢ ⎥⎣ ⎦ ⎢ ⎥ ⎢ ⎥⎡ ⎤− +⎣ ⎦⎢ ⎥= ⎢ ⎥ ⎢ ⎥ ⎢ ⎥ ⎡ ⎤− − +⎢ ⎥⎣ ⎦⎣ ⎦ (5) ( ) ( ) ( )( )0 0 0(2), (3), , ( ) t ny x x x n= (6) solving (3), grey model can be gotten: ( )0( 1) (1) aku ux k x e a a −⎡ ⎤+ = − +⎢ ⎥⎣ ⎦ 0,1, 2, ,k n= (7) 4.2 improvement of grey system theory during the thermal error modeling, the prediction model is acquired both by the collected all data and only a part of data. generally, the model from different data series is different. so the parameter a is not same. if (0) (0) (0) (0)( (1), (2), , ( ))x x x x n= , (0) ( )x n can be seen the origin of the time axis, and t i∈ , then t n< is considered as past, t n= as present, t n< as future. so (1) the grey model from (0) (0) (0) (0)( (1), (2), , ( ))x x x x n= is called as all data model (1,1)gm . (2) if 0 1k∀ > , the grey model from (0) (0) (0) (0) 0 0( ( ), ( 1), , ( ))x x k x k x n= + is called as part data model (1,1)gm . (3) if (0) ( 1)x n + is the new information, which is added to (0)x , the grey model from (0) (0) (0) (0) (0)( (1), (2), , ( ), ( 1))x x x x n x n= + is called as new model (1,1)gm . (4) adding new information (0) ( 1)x n + and eliminating the old information (0) (1)x , the grey model from (0) (0) (0) (0)( (2), , ( ), ( 1))x x x n x n= + is called as the metabolism (1,1)gm . 4.3 application of the grey system modeling on the basis of grey system modeling analysis and selected temperature variables above sections, the metabolism (1,1)gm of grey model is individually used as building thermal error model in the x, y, and z directions. detailed modeling process is neglected for the paper length limitation. this paper only offers built model based on experiment data, which include measurement values, forecast values, and residuals. experiment data graph, model predication graph, and residuals graph of three directions have been drawn as shown in fig.5. fig. 5 shows that experiment data graph and model predication graph are very coincident, the residuals of any data point is less 5 mμ . it can be seen that the prediction models is precise, which can be written in compensation control device. advances in systems science and applications (2011), vol.11, no.3-4 345 fig.5. the output results of real measurement and model 5. real time compensation of the thermal error in order to evaluate the performance of the metabolism grey model (1,1)gm , other experiments are carried out according to the standard iso.232. the experimental parameter setting is similar with the experimental setting described in section 2. four temperature sensors (no. 2, 15, 21 and 24) and the displacement sensors are only used in the experiments. two group comparison experiments have been done, one group is 60 parts without compensation, and the other group is also 60 parts with compensation. the material of experiments parts is 45# fe. the fig.6 is the trial result. the averagely thermal error of two turntable 5-axis nc machine tools is 20 mμ before compensation, and is 7 mμ after compensation. there is averagely a 60% 346 wang: thermal error modeling analysis and compensation based on the grey system theory on two turntable 5-axis nc machine tools increase on precision. compensation experiments proved the metabolism grey model (1,1)gm is with higher precision, and the modeling method is accurate. part num. e/ μ m 10 20 30 40 50 600 -5 0 5 10 15 20 25 before compensation after compensation (b) y direction fig. 6. x, y, z direction experiments before and after compensation 6. conclusions in this paper, the metabolism grey model (1,1)gm was proposed for on-line prediction of the dynamic and highly non-linear thermal errors on two turntable 5-axis nc machine tools. the proposed model not only enhances the prediction accuracy of the thermal error but also reduces the computation cost. relative analysis method is used as assuring the optimal temperature measurement point location of machine tool. the experimental results demonstrated that the machining accuracy was improved significantly after implementation of the compensation of the thermal errors. advances in systems science and applications (2011), vol.11, no.3-4 347 acknowledgments the authors are thankful to the education department of henan province of china for supporting this research under grant science research project 2010a460012. references [1] yang, j., yuan, j., ni, j.. thermal error mode analysis and robust modeling for error compensation on a cnc turning center , int. j. mach. tools manuf. vol.39, 1999, 1367–1381. [2] wu h., zhang, h., wang, x.s.. thermal error optimization modeling and real-time compensation on a cnc turning center, j. mater. process. technol. vol.207, 2008, 172–179. [3] wang xiushan, yang jianguo. synthesis error modeling and thermal error compensation of five-axis machining center. mater, sci. forum. vol.532-533, 2006, 49-52. [4] zhang zhiyong. master matlab 6.5 version . bhu publishing company, beijing, 2004. [5] li yongxiang, researchment and application of high-effective precise testing and reeor modeling for nc machine tools. shanghai: shanghai jiaotong university, 2007. advances in systems science and applications (2013) vol.13 no.1 68-79 optimization of statistical decision for personnel management problems in tourism konstantin n. nechval1, nicholas a. nechval2, gundars berzins2 and maris purgailis2 1applied mathematics department, transport and telecommunication institute,lomonosov street 1, lv-1019 riga, latvia 2statistics department, evf research institute, university of latvia,raina blvd 19, lv-1050 riga, latvia abstract a large number of problems in production planning and scheduling, location, transportation, finance, and engineering design require that decisions be made in the presence of uncertainty. in the present paper, for improvement or optimization of statistical decisions under parametric uncertainty, a new technique of invariant embedding of sample statistics in a performance index is proposed. this technique represents a simple and computationally attractive statistical method based on the constructive use of the invariance principle in mathematical statistics. unlike the bayesian approach, an invariant embedding technique is independent of the choice of priors. it allows one to eliminate unknown parameters from the problem and to find the best invariant decision rule, which has smaller risk than any of the well-known decision rules. in order to illustrate the application of the proposed technique for constructing optimal statistical decisions under parametric uncertainty, we discuss the following personnel management problem in tourism. a certain company provides interpreter-guides for tourists. some of the interpreter-guides are permanent ones working on a monthly basis at a daily guaranteed salary. the problem is to determine how many permanent interpreter-guides should the company employ so that their overall costs will be minimal? we restrict attention to families of underlying distributions invariant under location and/or scale changes. a numerical example is given. keywords personnel management problem, invariant embedding technique, optimization 1 introduction most of the operations research and management science literature assumes that the true distributions are specified explicitly. however, in many practical situations, the true distributions are not known, and the only information available may be a time-series (or random sample) of the past data. analysis of decisionmaking problems with unknown distribution is not new. several important papers have appeared in the literature. when the true distribution is unknown, one may either use a parametric approach (where it is assumed that the true distribution belongs to a parametric family of distributions) or a non-parametric approach advances in systems science and applications (2013) vol.13 no.1 69 (where no assumption regarding the parametric form of the unknown distribution is made). under the parametric approach, one may choose to estimate the unknown parameters or choose a prior distribution for the unknown parameters and apply the bayesian approach to incorporating the past data available. parameter estimation is first considered in [1] and further development is reported in [2]. scarf [3] considers a bayesian framework for the unknown demand distribution. specifically, assuming that the demand distribution belongs to the family of exponential distributions, the demand process is characterized by the prior distribution on the unknown parameter. further extension of this approach is presented in [4]. within the non-parametric approach, either the empirical distribution [2] or the bootstrapping method (e.g. see [5]) can be applied with the available past data to obtain a statistical decision rule. a third alternative to dealing with the unknown distribution is when the random variable is partially characterized by its moments. when the unknown demand distribution is characterized by the first two moments, scarf [6] derives a robust min-max inventory control policy. further development and review of this model is given in [7]. in the present paper we consider the case, where it is known that the true distribution function belongs to a parametric family of distributions. it will be noted that, in this case, most stochastic models to solve the problems of control and optimization of system and processes are developed in the extensive literature under the assumptions that the parameter values of the underlying distributions are known with certainty. in actual practice, such is simply not the case. when these models are applied to solve real-world problems, the parameters are estimated and then treated as if they were the true values. the risk associated with using estimates rather than the true parameters is called estimation risk and is often ignored. when data are limited and (or) unreliable, estimation risk may be significant, and failure to incorporate it into the model design may lead to serious errors. its explicit consideration is important since decision rules that are optimal in the absence of uncertainty need not even be approximately optimal in the presence of such uncertainty. the problem of determining an optimal decision rule in the absence of complete information about the underlying distribution, i.e., when we specify only the functional form of the distribution and leave some or all of its parameters unspecified, is seen to be a standard problem of statistical estimation. unfortunately, the classical theory of statistical estimation has little to offer in general type of situation of loss function. the bulk of the classical theory has been developed about the assumption of a quadratic, or at least symmetric and analytically simple loss structure. in some cases this assumption is made explicit, although in most it is implicit in the search for estimating procedures that have the “nice” statistical properties of unbiasedness and minimum variance. such procedures are usually satisfactory if the estimators so generated are to be used 70 konstantin n. nechval: optimization of statistical decision for personnel management ... solely for the purpose of reporting information to another party for an unknown purpose, when the loss structure is not easily discernible, or when the number of observations is large enough to support normal approximations and asymptotic results. unfortunately, we seldom are fortunate enough to be in asymptotic situations. small sample sizes are generally the rule when estimation of system states and the small sample properties of estimators do not appear to have been thoroughly investigated. therefore, the above procedures of the statistical estimation have long been recognized as deficient, however, when the purpose of estimation is the making of a specific decision (or sequence of decisions) on the basis of a limited amount of information in a situation where the losses are clearly asymmetric-as they are here. in this paper, we propose a new technique to solve optimization problems of statistical decisions under parametric uncertainty. the technique is based on the constructive use of the invariance principle for improvement (or optimization) of statistical decisions. it allows one to yield an operational, optimal information-processing rule and may be employed for finding the effective statistical decisions for many problems of the operations research and management science. the illustrative application of the invariant embedding technique to personnel management problems in tourism is given below. 2 invariant embedding technique this paper is concerned with the implications of group theoretic structure for invariant performance indexes. we present an invariant embedding technique based on the constructive use of the invariance principle for decision-making. this technique allows one to solve many problems of the theory of statistical inferences in a simple way. the aim of the present paper is to show how the invariance principle may be employed in the particular case of improvement or optimization of statistical decisions. the technique used here is a special case of more general considerations applicable whenever the statistical problem is invariant under a group of transformations, which acts transitively on the parameter space [8-13]. 2.1 preliminaries our underlying structure consists of a class of probability models (x, a, p), a one-one mapping ψ taking p onto an index set θ, a measurable space of actions (u,b), and a real-valued function r defined on θ× u. we assume that a group g of one-one a-measurable transformations acts on x and that it leaves the class of models (x,a,p) invariant. we further assume that homomorphic images g and g̃ of g act on θ and u, respectively. (g may be induced on θ through ψ; g̃ may be induced on u through r). we shall say that r is invariant if for every (θ,u) ∈ θ×u r(gθ, g̃u) = r(θ,u), g ∈ g. (1) advances in systems science and applications (2013) vol.13 no.1 71 given the structure described above there are aesthetic and sometimes admissibility grounds for restricting attention to decision rules φ: x → u which are (g,g̃) equivariant in the sense that φ(gx) = g̃φ(x),x ∈ x, g ∈ g (2) if g is trivial and (1), (2) hold, we say φ is g-invariant, or simply invariant. 2.2 invariant functions we begin by noting that r is invariant in the sense of (1) if and only if r is a g•invariant function, where g• is defined on θ×u as follows: to each g ∈ g, with homomorphic images g, g̃ in g, g̃ respectively, let g•(θ,u) = (ḡθ, g̃u), (θ,u) ∈ (θ× u). it is assumed that g̃ is a homomorphic image of g. definition 1 (transitivity). a transformation group g acting on a set θ is called (uniquely) transitive if for every θ,ϑ ∈ θ there exists a (unique) g = g such that gθ = ϑ. when g is transitive on θ we may index g by θ: fix an arbitrary point θ ∈ θ and define gθ1 to be the unique g = g satisfying gθ = θ1. the identity of g clearly corresponds to θ. an immediate consequence is lemma 1. lemma 1 (transformation). let g be transitive on θ. fix θ ∈ θ and define gθ1 as above. then gqθ1 = q gθ1 for θ ∈ θ, q ∈ g. proof.the identity gqθ1θ=qθ1=q gθ1θ shows that gqθ1 and q gθ1 both take θ into qθ1, and the lemma follows by unique transitivity. theorem 1 (maximal invariant). let g be transitive on θ. fix a reference point θ0 ∈ θ and index g by θ. a maximal invariant m with respect to g• acting on θ× u is defined by m(θ,u) = g̃−1 θ u, (θ,u) ∈ θ× u (3) proof. for each (θ,u) ∈ (θ× u) and g ∈ g m(gθ, g̃u) = (g̃−1 gθ )g̃u = (g̃g̃θ) −1g̃u = g̃−1 θ g̃−1g̃u = g̃−1 θ u = m(θ,u) (4) by lemma 1 and the structure preserving properties of homomorphisms. thus m is g•− invariant. to see that m is maximal, let m(θ1,u1) = m(θ2,u2). then g̃−1 θ1 u1 = g̃−1 θ2 u2 or u1 = g̃u2, where g̃ = g̃θ1 g̃ −1 θ2 . since θ1 = gθ1θ0 = ḡθ1 ḡ −1 –θ2 θ2 = ḡθ2, (θ1,u1) = g•(θ2,u2) for some g• ∈ g•, and the proof is complete. corollary 1.1 (invariant embedding). an invariant function, r(θ,u), can be transformed as follows: r(θ,u) = r(g θ̂ −1θ, g θ̂ −1u) = r̈(v,η) (5) where v = v(θ, θ̂) is a function (it is called a pivotal quantity) such that the distribution of v does not depend on θ; η = η(u, θ̂) is an ancillary factor; θ̂ is 72 konstantin n. nechval: optimization of statistical decision for personnel management ... the maximum likelihood estimator of θ (or the sufficient statistic for θ). corollary 1.2 (best invariant decision rule). if r(θ,u) is an invariant loss function, the best invariant decision rule is given by φ∗(x) = u∗ = η−1(η∗, θ̂) (6) where η∗ = arg inf η eη {r̈(v, η)} . (7) corollary 1.3 (risk). a risk function (performance index) r(θ,φ(x)) = eθ{r(θ,φ(x))} = eη{r̈(v0,η0)} (8) is constant on orbits when an invariant decision rule φ(x) is used, where v0 = v0(θ,x) is a function whose distribution does not depend on θ; η0 = η0(u,x) is an ancillary factor. for instance, consider the problem of estimating the locationscale parameter of a distribution belonging to a family generated by a continuous cdf f : p = {pθ : f ((x− µ)/σ), x ∈ r,θ ∈ θ}, θ = {(µ, σ) : µ, σ ∈ r, σ > 0} = u . the group g of location and scale changes leaves the class of models invariant. since g induced on θ by pθ → θ is uniquely transitive, we may apply theorem 1 and obtain invariant loss functions of the form r(θ,φ(x)) = r[(φ1(x)− µ)/σ, φ2(x)/σ] (9) where θ = (µ, σ) and φ(x) = (φ1(x), φ2(x)) (10) let θ̂ = (µ̂, σ̂) and u = (u1, u2), then r(θ,u) = r̈(v,η) = r̈(ν1 + η1ν2, η2ν2) (11) where ν = (ν1, ν2), ν1 = (µ̂− µ)/σ, ν2 = σ̂/σ (12) η = (η1, η2), η1 = (u1 − µ̂)/σ̂, η2 = u2/σ̂ (13) 3 application to personnel management problem in tourism personnel management forms a significant proportion of overall costs in hotels, tourism companies and fast food restaurants. a reduction in this by even 1% represents considerable cost savings. demand for services is not generally known with certainty before hand and management often relies on a combination of intuition, software systems and local knowledge (particularly of marketing campaigns, events and attractions). staff scheduling is a key element of management planning in such circumstances. there have been a number of general survey advances in systems science and applications (2013) vol.13 no.1 73 papers in the area of personnel management; these include [14] and [15]. the latter survey concentrates on general labour scheduling models. a survey of crew scheduling is given in [16]. surveys of the literature in airline crew scheduling appear in [17-18]. a good survey of tools, models and methods for bus crew scheduling is [19]. a survey of the nurse scheduling literature is provided in [2021]. as can be seen from this review, a large amount of work has already been done in the area of personnel scheduling. nevertheless there is still significant room for improvements in this area. we see improvements occurring not only in the area of tools, models and methods for personnel management, but also in the wider applicability of these tools, models and methods. in this paper, we consider the following personnel management problem in tourism. a certain company provides interpreter-guides for tourists. the number of permanent interpreterguides employed by the company is such that u of them are permanently working on a monthly basis at a daily guaranteed salary c1 (in terms of money); when the demand for their services exceeds u, supplementary interpreter-guides or extras are taken on at a daily salary c2(> c1). sometimes the shortage of extras will necessitate canceling a tour, and when this happens, the loss is reckoned at c3(> c2). how many permanent interpreter-guides should the company employ so that overall costs will be minimal? following kaufmann and faure [22], we review the personnel management model and provide a broader interpretation to the structure of its solution. in development of the personnel management model, we will assume that the daily demand for tours x is a continuous nonnegative random variable with the probability density function fθ(x) and cumulative distribution function fθ(x). the notation, we use for the personnel management model, is given below. x random variable representing the daily demand for tours fθ(y) probability density function of a demand x fθ(y) cumulative distribution function of a demand x θ parameter (in general, vector) y random variable representing the daily supply of extras p(y) probability of a supply y, where y=0, 1, ...,∞ c1 daily guaranteed salary for the permanent interpreter-guide c2 daily salary for the supplementary (or extra) interpreter-guide c3 shortage cost per unit of x u variable representing the number of the permanent interpreter-guides u∗ optimal quantity of the number of the permanent interpreter-guides c(u) expected overall costs as a function of u 74 konstantin n. nechval: optimization of statistical decision for personnel management ... thus, the function of overall costs is given by c(u,x, y ) =  c1u, 0 ≤ x ≤ u c1u+ c2(x − u), u ≤ x ≤ u+ y c1u+ c2y + c3(x − u− y ), u+ y < x < ∞ (14) we write the expected overall costs as c(u) = e {eθ {c(u,x, y )}} = ∞∑ y=0 p(y) ∫ ∞ 0 c(u, x, y)fθ(x)dx = ∞∑ y=0 p(y)c(u, y), (15) where c(u, y) = ∫ ∞ 0 c(u, x, y)fθ(x)dx = c1u+ c2 ∫ u+y u (x− u)fθ(x)dx + ∫ ∞ u+y [c2y + c3(x− u− y)]fθ(x)dx. (16) the function c(u) can be shown to be convex in u, thus having a unique minimum. taking the first derivative of c(u) with respect to u and equating it to zero, we get ∞∑ y=0 p(y) ( c1 − c2 ∫ u+y u fθ(x)dx− c3 ∫ ∞ u+y fθ(x)dx ) = 0. (17) the value of u that minimizes (17) is the one that satisfies c2f̄θ(u ∗) + (c3 − c2) ∞∑ y=0 p(y)f̄θ(u ∗ + y) = c3 − c1, (18) where fθ(x) = 1− fθ(x) (19) if p(y = 0) = 1,then fθ(u ∗) = c3 − c1 c3 . (20) in this case, we should choose the u∗ such that the cumulative distribution function of u∗ equals the ratio of the difference of the underage and overage costs to the underage cost. a relatively high underage cost results in a higher number of the permanent interpreter-guides, whereas a relatively high overage cost leads to a lower number of the permanent interpreter-guides, as one would expect. if the advances in systems science and applications (2013) vol.13 no.1 75 daily demand for tours x follows the exponential distribution with the probability density function fσ(x) = σ−1exp(−x/σ), σ > 0 (21) and the cumulative distribution function fσ(x) = 1− exp(−x/σ) (22) where σ is the scale parameter (σ > 0), then c(u) = ∞∑ y=0 p(y)c(u, y) (23) where c(u, y) = σ[c1 u σ + c2exp(− u σ ) + (c3 − c2)exp(− u+ y σ )] (24) and the value of u that minimizes (23) is the one that satisfies c2exp(− u∗ σ ) + (c3 − c2) ∞∑ y=0 p(y)exp(−u∗ + y σ ) = c1 (25) if p(y = 0) = 1, then u∗ = σln( c3 c1 ) (26) and c(u∗) = c1[1 + ln( c3 c1 )]σ (27) parametric uncertainty. consider the case when the parameter σ is unknown. let x1 ≤ ... ≤ xn be the past observations (of the daily demand for tours) from the exponential distribution (21). then s = n∑ i=1 xi (28) is a sufficient statistic for σ; s is distributed with gσ(s) = [γ(n)σn]−1sn−1exp(−s/σ)(s > 0) (29) to find the best invariant decision rule ubi , we use the invariant embedding technique [8-14] to transform (24) to the form, which depends on the pivotal quantity υ = s/σ, the ancillary factor η = u/s and y/s, c(u, y) = σ[c1 u s s σ + c2exp(− u s s σ ) + (c3 − c2)exp(− u+ y s s σ )] = σ[c1ην + c2exp(−ην) + (c3 − c2)exp(−η + y s )ν] = c(η, y, ν|s) (30) 76 konstantin n. nechval: optimization of statistical decision for personnel management ... we find the expected overall costs for the statistical decision u = ηs as c(η|s) = ∞∑ y=0 p(y)c(η, y|s) (31) where c(η, y|s) = ∫ ∞ 0 c(η, y, ν|s)g(ν)dν = σ[c1ηn+ c2 1 (1 + η)n + (c3 − c2)(1 + η + y s )−n] (32) g(ν) = [γ(n)]−1νn−1exp(−ν)(ν > 0) (33) the value of η that minimizes (31) is the one that satisfies c2 1 (1 + η∗)n+1 + (c3 − c2) ∞∑ y=0 p(y)(1 + η∗ y s )−(n+1) = c1 (34) thus, ubi = η∗s (35) if p(y = 0) = 1, then η∗ = ( c3 c1 )1/(n+1) − 1 (36) and c(η∗|s) = σ[c1η ∗n+ c3 1 (1 + η∗)n ] = c1[( c3 c1 )1/(n+1) − n]σ (37) comparison of decision rules. for comparison, consider the maximum likelihood decision rule that can be obtained from (26) as uml = σ̂ln( c3 c1 ) = ηmls (38) where σ̂ = s/n is the maximum likelihood estimator of σ, ηml = ln( c3 c1 )1/n (39) since ubi and uml belong to the same class c = {u : u = ηs} (40) it follows from the above that uml is inadmissible in relation to ubi . if, say, c1 = 50, c3 = 3500 (in terms of money), and n=1, we have that rel.eff.c(η|s){uml, ubi , σ} = c(η∗|s)/c(ηml|s) = c1η ∗n+ c3 1 (1+η∗)n c1ηmln+ c3 1 (1+ηml)n = 0.9 (41) advances in systems science and applications (2013) vol.13 no.1 77 thus, in this case, the use of ubi leads to a reduction in the expected overall costs of about 10% as compared with uml. the absolute expected overall costs will be proportional to σ and may be considerable. predictive inference. it will be noted that the predictive probability density function of the daily demand for tours, x, which is compatible with (15), is given by f(x|s) = n+ 1 s (1 + x s )−(n+2)(x > 0) (42) using (42), the predictive overall costs are determined as c(p)(u|s) = ∞∑ y=0 p(y)c(p)(u, y|s) (43) where c(p)(u, y|s) = c1u+ c2 ∫ u+y u (x− u)f(x|s)dx+ ∫ ∞ u+y [c2y + c3(x− u− y)]f(x|s)dx = s n [c1 u s n+ c2(1 + u s )−n + (c3 − c2)(1 + u s + y s )−n] (44) which can be reduced to c(p)(η, y) = s n [c1ηn+ c2 1 (1 + η)n + (c3 − c2)(1 + η + y s )−n] (45) thus, it follows from (32) and (45) that ubi can be found immediately from (43) as ubi = argmin u c(p)(u|s). (46) 4 conclusions and directions for future research in this paper, we propose a new technique to improve or optimize statistical decisions under parametric uncertainty. the method used is that of the invariant embedding of sample statistics in a performance index in order to form pivotal quantities, which make it possible to eliminate unknown parameters (i.e., parametric uncertainty) from the problem. it is especially efficient when we deal with asymmetric performance indexes and small data samples. more work is needed, however, to obtain improved or optimal decision rules for the problems of unconstrained and constrained optimization under parameter uncertainty when: (i) the observations are from general continuous exponential families of distributions, (ii) the observations are from discrete exponential families of distributions, (iii) some of the observations are from continuous exponential families of distributions and some from discrete exponential families of distributions, (iv) the observations 78 konstantin n. nechval: optimization of statistical decision for personnel management ... are from multiparametric or multidimensional distributions, (v) the observations are from truncated distributions, (vi) the observations are censored, (vii) the censored observations are from truncated distributions. references [1] conrad s.a. (1976), “data and the estimation of demand”, oper. res. quart, vol.27, pp.123-127. [2] liyanage l.h, shanthikumarj.g. (2005), “a practical inventory policy using operational statistics”, operations research letters, vol.33, pp.341-348. [3] scarf h. (1959), “bayes solutions of statistical inventory problem”, ann. math. statist, vol.30, pp.490-508. [4] chu l.y, shanthikumar j.g, shen z.j.m. (2008), “solving operational statistics via a bayesian analysis”, operations research letters, vol.36, pp.110-116. [5] bookbinder j.h, lordahl a.e. (1989), “estimation of inventory reorder level using the bootstrap statistical procedure”, iie trans, vol.21, pp.302-312. [6] scarf h. (1958), a min-max solution of an inventory problem. studies in the mathematical theory of inventory and production (chapter 12), stanford: stanford university press. [7] gallego g, moon i. (1993), “the distribution free newsvendor problem: review and extensions”, j. oper. res. soc, vol.44, pp.825-834. [8] nechval n. a, vasermanis e. k. (2004), improved decisions in statistics, riga: izglitibas soli. [9] nechval n.a, berzins g, purgailis m, nechval k.n. (2008), “improved estimation of state of stochastic systems via invariant embedding technique”, wseas transactions on mathematics, vol.7, pp.141-159. [10] nechval n.a, berzins g, purgailis m, nechval k.n, zolova n. (2009), “improved adaptive control of stochastic systems”, advances in systems science and applications, vol.9, pp.11-20. [11] nechval n.a, nechval k.n, danovich v, liepins t. (2011), “optimization of new-sample and within-sample prediction intervals for order statistics”, in: proceedings of the 2011 world congress in computer science, computer engineering, and applied computing, worldcomp’11, 18-21 july, 2011, las vegas, nevada, usa, pp.91-97. advances in systems science and applications (2013) vol.13 no.1 79 [12] nechval n.a, nechval k.n, purgailis m, rozevskis, u. (2012). “optimal prediction intervals for order statistics coming from location-scale families”, engineering letters, vol.20, pp.353-362. [13] nechval n.a, purgailis m. (2012), “stochastic control and improvement of statistical decisions in revenue optimization systems”, in: stochastic modeling and control, ivan ganchev ivanov (ed.). croatia: sciyo, pp.185-210. [14] bechtold s, brusco m, showalter m. (1991), “a comparative evaluation of labor tour scheduling methods”, decision sciences, vol.22, pp.683-699. [15] tien j, kamiyama a. (1982), “on manpower scheduling algorithms”, siam review, vol.24, pp.275-287. [16] bodin l, golden b, assad a, ball m. (1983), “routing and scheduling of vehicles and crews the state of the art”, computers and operations research, vol.10, pp.63-211. [17] arabeyre j, fearnley j, steiger f, teather w. (1969), “the airline crew scheduling problem: a survey”, transportation science, vol.3, pp.140-163. [18] gamache m, soumis f. (1998), “a method for optimally solving the rostering problem”, in: or in airline industry, g. yu (ed.). boston: kluwer academic publishers, pp.124-157. [19] wren a. (1981), “a general review of the use of computers in scheduling buses and their crews”, in: computer scheduling of public transport, urban passenger vehicle and crew scheduling, a. wren (ed.). north-holland, amsterdam, pp.3-16. [20] bradley d., martin j. (1991), “continuous personnel scheduling algorithms: a literature review”, journal of the society for health systems, vol.2, pp.2-8. [21] sitompul, d., radhawa, s. (1990), “nurse scheduling: a state-of-the-art review”, journal of the society for health systems, vol.2, pp.62-72. [22] kaufmann a, faure r. (1968), introduction to operations research, new york: academic press. corresponding author author can be contacted at: nechval@junik.lv. adv syst sci appl 2018; 1; 20-40 published online at http://ijassa.ipu.ru. the contract theory: a 3-dimensional reflection on commodities & capital conversions sailau baizakov 1 , azamat ryskulovich. oinarov 2 , dana akylbekovna eshimova 3 , jeffrey yi-lin forrest 4 1) scientific supervisor, jsc "economic research institute", 65 temir-kazyk street, astana 010000, kazakhstan; e-mail: 37baizakov@gmail.com; 2) chairman of the board of the jsc "kazakhstan center for public-private partnership", 65 temir-kazyk street, astana 010000, kazakhstan; e-mail: azamat.oinarov@gmail.com; 3) deputy chairman of the board of the jsc "kazakhstan center for public-private partnership",65 temir-kazyk street, astana 010000, kazakhstan; e-mail: eda.07@mail.ru; 4) school of business, slippery rock university, slippery rock, pa 16057, usa; e-mail: jeffrey.forrest@sru.edu abstract: contract theory is a relatively young field of the economic science. the uniqueness of this branch of investigation is established on the basis that it is crucial to study the origin of microeconomic and macroeconomic indicators. in particular, because of this theory of contracts, it becomes possible to properly evaluate the performance of the real balanced growth. the study of the theoretical foundations of contracting paves the way for further development of particular tools for analyzing economic policies. keyword: trade flow, cash flow, reproduction, tio, circulation, input-output. 1. introduction the analytical techniques developed for evaluating the domestic product fail to ensure the assessment accuracy of the cost of consumed goods and services. therefore, measuring the real final product is only based on the inflation indicator that is entailed from the nominal gdp indicator. at the same time, (piketty, 2015, p.592) discovers that the indicators of inflation and real economy growth, as analyzed by using the currently available analytical techniques, may not always be accurate. the members of kazakhstan’s economists’ interest group are currently developing a program to enable the process of decomposing intra-industry input-output tables into smaller regional components. such regionalization of the intra-industry balance sheet is deemed to clearly define the deviations between the indicators of the real and financial sectors in the development of the market economy. specifically, the decomposition of the country’s intraindustries balance sheet, as reflected in the country’s input-output tables, into smaller regional components could better mirror economic activities of the subordinate administrative and territorial subdivisions within the common national management system, and would mailto:37baizakov@gmail.com mailto:azamat.oinarov@gmail.com mailto:eda.07@mail.ru mailto:jeffrey.forrest@sru.edu 21 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) enable the setup of a new analyses model for qualitatively overseeing regulatory policies for sustainably developing the national economy. the concept that lies in the foundation of the program was initiated by the president of the republic of kazakhstan nursultan a. nazarbayev (2009). the core of the concept focuses on the right “choice of the innovation types, on which the ‘currency-and-financial system’ of the country and of the rest of the world would be based. such choice is viewed as the nucleus of the socio-political, technical, and technological milieu, in which every citizen would want to live. that is being desired by not only every citizen but also every family, every country, and the entire world. the development of the analytical tools and techniques in line with nazarbayev’s concept suggests that wreathing of benefits that item from technical and technological innovations in the real economy should be commensurate to socio-political innovations, especially in managerial decision-making. the latter has been set on a sound footing based on currency and financial innovations. measuring the true costs of goods and services, as reflected in (nazarbayev, 2009), is a distinctively new concept when compared to the existing approaches in the development analysis of a market economy. the unique content of the concept is linked to such a methodology that the key is the assessment of the performance quality of the institutions that implement the following three major innovational technologies within the cycle of reproducing capital in its monetary form, and capital in its commodity-form:  the real sector technical and technological innovations,  currency-and-financial innovations in managerial decision making, and  socio-political innovations in managerial decision making the specificity in measuring the true cost of goods and services under nazarbayev’s concept is given through the fact that production growth is closely aligned with conservation of resources. likewise, it is in line with the output increase of intermediate consumer goods. the latters are in essence directly involved in the production of output. the socio-economic effect of nazarbayev’s concept is in its negligence of the idea of generating profit at any cost. at the same time, according to (nazarbayev, 2009), profit may be obtained through efficient utilization of material, technical, and financial resources. such efficiency needs to be addressed in the production of each and every unit of the final product. 2. the purpose of this work the analysis of the macroeconomic dynamics using the three-component reproduction model, developed by russian academician alexander g. granberg (1985), has shown that economic growth indicators, specifically identified for sectors of an economy, and their growth rates have to be defined by decision makers. the key value of the granberg model has been attached to the reduction of a weighted average of material intensity in the gross domestic product. however, such contraction, according to the granberg model, “may be a cause of a more compounded impact on the dynamics of the three components” (granberg, 1985, p. 109). it is because the changes in the expenditure coefficient of one component necessitate changes in the other component. the solution as for how to satisfy the granberg-identified need for a reduction of material intensity in the gross product may be found in the comments by f. engels and were reflected in the supplement by f. engels to capital. on this matter, f. engels wrote the following: “the development of the productive power of labour reacts also on the original capital already engaged in the process of production. a part of the functioning constant capital consists of instruments of labour, such as machinery, which are not consumed, and therefore not reproduced, or replaced by new ones of the same kind, until after long periods of time. but the contract theory: a 2-dimensional reflection on commodities & capital 22 copyright ©2018 assa. adv. in systems science and appl. (2018) every year a part of these instruments of labour perishes or reaches the limit of its productive function. if the productiveness of labour has, during the using up of these instruments of labour, increased (and it develops continually with the uninterrupted advance of science and technology), more efficient and (considering their increased efficiency), cheaper machines, tools, apparatus, replace the old. the old capital is reproduced in a more productive form, apart from the constant detail improvements in the instruments of labour already in use. the other part of the constant capital, raw material and auxiliary substances, is constantly reproduced in less than a year. every introduction of improved methods, therefore, works almost simultaneously on the new capital and on that already in action. like the increased exploitation of natural wealth by the mere increase in the tension of labour-power, science and technology give capital a power of expansion independent of the given magnitude of the capital actually functioning. they react at the same time on that part of the original capital which has entered upon its stage of renewal. this, in passing into its new shape, incorporates gratis the social advance made while its old shape was being used up. of course, this development of productive power is accompanied by a partial depreciation of functioning capital. labour transmits to its product the value of the means of production consumed by it. on the other hand, the value and mass of the means of production set in motion by a given quantity of labour increase as the labour becomes more productive. though the same quantity of labour adds always to its products only the same sum of new value, still the old capital value, transmitted by the labour to the products, increases with the growing productivity of labour.” 1 by carefully reading through engels’ comments on k. marx’ capital, the accurate measurements of capital in its monetary form ( ), and capital in its commodity form ( ), one may find that these equations may well serve the basis for solving the granberg urge, targeting at the reduction in material intensity of the gross product. in other words, the three dimensional measurements of the indicators of the nominal gdp (ngdp = ) help to define the cost of the final product, i.e., by means of the indicators of the gross aggregate product ( ). that represents the sum of costs, including materials, in the form of the annual income ( ). without using the three dimensional method to measure the final product indicators, the solution of the granberg puzzle is deemed impossible or difficult. the research subject matter of this work, as first revealed by granberg, relates to the subjective need to track down the material cost and the efficiency of production resources. as such it echoes the following definition of this problem, as formulated by m. porter (2002, p. 496, 220): “any motion in a developed economy requires developing a sound local competitiveness. competition should be in line with the shift of the major focus from low wages to low costs. that would require improvements in the efficiency of production and of services”. 3. granberg’s contribution of unveiling the leontief paradox a. granberg not only discovered the function of the costs of production resources, but also led his followers to a search of the true cost of goods and services. moreover, he succeeded in explaining the leontief paradox. the approach developed by w. leontief consisted of numerical measurements of not only production, but also distribution of the common good. he constructed the reproduction schemes of the gross product. according to granberg, leontief discovered a new area in the economic science by blending the economic functions theory with mathematical modeling, systemic techniques, and processing of the economic information (suslov v.i., 2016). granberg named leontief as the most pragmatic economist-theoretician. leontief’s research methodology was built on practical observations and analyses of structural shifts in 1 marx, k., engels, f. collection. ed.2., v. 25, p. i, p. 286. 23 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) economy and trade. unexpected conclusions were quite frequent outcomes of his research. the leontief paradox is now well studied academically. granberg acquiesced to leontief’s thought that the united states of america must interact with the rest of the world in such a way that its exports are capital intensive and imports labor intensive. such stance is stemmed from the stature of the u.s.a. as an excessively capitalized country with a relatively limited highly paid labor market. in this regard, leontief found that the u.s. did export labor and import capital. such a statement though had found no reflection in the well-known theories on foreign trade (suslov v.i., 2016, p. 163-164). for instance, the heckscher-olin theory explained it differently. the scientific essence of leontief’s approach lies in applying systemic thinking and logic to find economic solutions. then, what is the leontief paradox? the assessment of changes in the structure of the costs of materials, labor, and capital resources per a unit of production of the final product is the core of the leontief paradox. absence of juxtaposition between the methods of market-based planning and those of planned economic management, as shown in leontief’s approach, is important. leontief viewed an economy as a ship where private initiative is the wind, while planning is the steering wheel, which shows the direction. he pointed out the need for setting an equalizing balance between market and state (suslov v.i., 2016, p. 181). in our view, the leontief concept fully complies with the three dimensional measurement of capital in its monetary form, as a mechanism of pushing the market-based economy forward, and capital in its commodity form, as a mechanism making the management of state-run economy with limited resources effective. 4. the foundation of the contract theory the contract theory is based on the principle of gaining mutual benefits from a trade deal. in this theory, market relations are represented as a cycle of reproduction schemes of goods and services that mainly consist of the conversion of the capital in its monetary form, to the capital in its commodity form. the contractual relations between economic actors form the foundation for mutual conversions within the common cycle of capital and commodity. the basis of the contract theory is similar to the cyclical cost of capital in its monetary form, and the cost of capital in its commodity form. in formulaic terms, measuring the indicators of the cost of capital in its money form ( ), and capital in its commodity form ( ), has been confirmed by the fact that the cost of capital in its money form ( ) has originally been determined in monetary terms. however, the cost of production ( ) has been determined in monetary terms, and also in terms of time, which has been utilized for labor. let the average price of a unit of a national currency, as measured by purchasing power, in relative terms, be indicated as purchase power of the national currency (ppnc) following a balanced equation, which, in legal terms, has been notary-verified and signed as a contract. such a contract-based equation has close linkages between aggregate labor costs ( ) in the form of a reward of labor, and final outcomes of labor ( ) in the form of a chain of surplus value. as with the account of ‘ppnc’ and ‘wages’, such legally binding act of purchase and sale may be presented by the following equation of the contract theory: (1) since the contract has the power of a law, equ. (1), as a theoretic reflection of the noted specific law, has a legally binding force. this described pattern works in such a manner that is analogous to the power of either the ohm’s law in physics or the natural law of gravitation. however, equ. (1), upon having been converted to economic terms, is itself an economic law. the contract theory: a 2-dimensional reflection on commodities & capital 24 copyright ©2018 assa. adv. in systems science and appl. (2018) when this particular economic law, as represented in equ. (1), is extrapolated onto various sectors of the economy and various types of economic activities, the economic law, well known to experts in the area of analysis of intra-industries input-output tables, may easily be obtained as follows: (2) where is the direct labor intensity of the product, and the overall labor intensity of the product, as complied with the cost of final product ( ) and the aggregate sum of the costs of materials and financial resources that have been utilized in producing product ( ). let the coefficient ‘ ’ be the proportionality of in monetary terms. by measuring the variables in labor terms, we obtain the following equation: . both of these proportions are economically essential functions of time. in fact, they are equal to each another. here, the first proportion represents the cost of final product ( ), the sum of the costs of materials, labor, and capital resources, vs. the production ( ). it changes over time. each increase in the relevant coefficient ‘ ’ over time reflects an increase in the volume of the final product ( ) that items from the utilization of resources for production ( ). any reduction, of this coefficient over time indicates a decrease in the volume of the final product. the above-noted changes over time occur under the pressure of new, innovative technologies, etc. in a series of turnovers of capital and commodities, the growth rate of the technological potential of any economy may be determined. therefore, the indicator, referred to in this paper as ‘ ’, shall be named as the ‘coefficient of science-and-technology potential’. the second proportion also tends to change over time. however, its reverse measurement ( ), which represents the proportion of full (direct and indirect) labor costs vs. direct costs ( ) may be interpreted differently. in this regard, the leontief paradox is based on this particular economic law. any theory is enlivened only when it is positively tested in real time practices. for that matter, the contract theory is not an exception. and, the reciprocal technological cycle of the mutual conversion of physical measurements of commodity masses into monetary masses has clearly been observed in practice. for example, we may admit the hypotheses relating to the annual volumes of goods and services as ones that have been realized in-kind, denoted as ‘ ’, of the currency unit. those that constitute gross proceeds are marked as ‘ ’. then, the price of commodities ( ) is given by , and the proceeds from sales are . in that case, the velocity of money in the national currency unit will be equal to ‘ ’. the marxian reproduction schemes formula may be rewritten as follows: . (3) from the this, the following equation is derived: , (4) or . (5) the afore-described hypothesis, as accepted by this paper, fully complies with the clark concept where any realized commodity-based product is represented as a sum of the elementary utilities of the material wealth, which has been utilized in producing the materialized commodity-driven product (clark, 2000, p.220). 25 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) in spite of the opportunity readily provided in evaluating the effectiveness of regulatory policies by the three-dimensional measurements of the cost of capital in its monetary form and capital in its commodity form, in compliance with the marxian reproduction scheme, additional analyses are currently undertaken by means of a more simplified technique following a. smith’s one dimensional income method (smith, 2007, p. 960). under the smithian income method, the nominal gdp is determined by simply deducting the costs of materials ( ) from the proceeds ( ): – . such calculations do not account for the contract theory, of which the main processes have been represented in equ. (1). in this regard, equ. (4) loses its original macroeconomic content and acquires quite a new setting. in the end, it converts to a macroeconomic equation incapable of solving the granberg puzzle: , (6) where ‘ ’ stands for the velocity of money. equ. (6) contains no physical dimension and is therefore represented by its monetary form, only. a. smith who defined labor as a sole method of measuring the annual income turns out to contradict himself. his followers are using a specific indicator, namely, the gdp deflator, to determine the physical volume of the final product. in reality, however, the gdp deflator, which has been derived from physical indices of goods and services, does serve as both a universal economic indicator, which is necessary to install the required dynamics of market prices on goods and services, and an indicator of the dynamics of the purchasing power of money in the financial system. since the latter indicator supports the balance in the nominal gdp as well as in real gdp ( ), equ. (6) is then helpful in determining . as with ceteris paribus, we have the following equation: (7) where the velocity of money ‘ ’ is itself a function in the velocity of money in the proceeds ‘ ’, and also dependent on the pace of money turnover in the intermediate product ‘ ’: and so, we have: . (8) therefore, equ. (8) relates the velocity of money ‘ ’ to the pace of the turnover of the nominal gdp and the pace of the turnover of intermediate commodities. the indicators of proceeds are therefore directly linked to the indicators of consumption of the intermediate products qp and the nominal gdp. based on equ. (8), emerges an opportunity for developing an operational technique, capable of transforming real economies into the engine of sustainable economic growth, meaning that the gdp price deflator (inflation indicator) and purchasing power of money may be transformed into the key indicators of managing innovational investments. the gdp price indicator (inflation indicator) is an integral factor consisting of such components that possess not only destructive, but also, constructive forces in developing the market economy. investigating its structure may help analyze the differences between the pace of implementation of technical and technological, ‘currency-related’ and financial, and socio-political innovations. the contract theory: a 2-dimensional reflection on commodities & capital 26 copyright ©2018 assa. adv. in systems science and appl. (2018) 5. contracts-based methods of analysis for economic policies the sraffa’s model: production of commodities by means of commodities. the main idea behind this model (sraffa, 1960) is found in the intra-industries nature of reproduction. the final product of one type of production serves the raw resource for the other type. in reality, the reproduction cycle forms a cycle of reproduction, where the final product is fully defined, and wages may contain the component of a surplus value, because the profit margin is defined by the interest rate. and that feature of sraffa’s model is distinct from the marginal utility theory, where the correlation between supply and demand may serve as the law on distribution. it also differs from the labor theory, where wages are defined by the means that provides the living to the worker and his/her family. similarly, the profit margin is defined by production technologies. the ronald coase conjecture (1960). the theory (kapelushnikov, 2017) stated that firms are created when transaction costs are lower inside the firm than outside it in the open market. it means that if the property rights of all parties are carefully defined and transaction costs nullified, then the final outcomes do not depend on the changes in the distribution of property rights. however, in the process of accounting the transaction costs, the desirable outcomes may not always be attainable. in this regard, high costs of obtaining the required information, conducting negotiations, and settling disputes may exceed potential benefits of the deal. besides, when accounting losses, there may appear essential differences in propensities of contractors on recording the losses. a reference to the potential impact of the income (that later became the core of the neo-institutional theory) has been introduced to capture all the afore-described differences. nevertheless, the coase theory fails to reflect the details of why some firms grow owing to the integration of sequential production stages, while others focus more on just one or a few production stages. the energy sector serves as an example of such integration where coal mines are paired with hydropower stations that work on coal. the williamson’s theory of transaction costs 2 . conceptually, this theory, established in 2009, is closely associated to the costs of contract settlements and regaining the right to property, or other services within the acts of interaction between two or more participants of a contract. according this theory, the hierarchical structure prevails in the market until it ensures an inexpensive and expeditious method of resolving conflicts. if the three agents cannot resolve their disputes relating to work load distribution and income, the manager steps in to resolve those disputes. not only from the economic stance has williamson’s concept been thoroughly studied, but also from the legal standpoint where a contract is regarded as an organizational component in harmonizing work processes. the theory ensured the unity of economies and legal acts, backed up by contracts and, thus, led to a more profound understanding of the aims and objectives of a collaborative work. the williamson’s firm has been depicted through the prism of institutional notions, not industrial. a firm and a market are being judged by their capacities to conduct various transactions, enabling resource minimization. the theory also pointed out that production facilities standing afar from one another are more likely to group together under the same owner. oliver hart and bengt holmstrom (kornelyuk, 2016): market economy is the economy of contracts. the theory is developed by hart and holmstrom to define the parameters of contracts. the theory has been valued for its advocacy of mutually beneficial decisions 2 the theories of transaction costs. chapters on economics. history of the economic science. retrieved from: (http://ecouniver.com/economik-rasdel/istekuz/215-teorii-transakcionnyx-izderzhek.html) 27 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) relating to contracts’ parties. it is of practical value for any of the macroeconomic models geared to balance economic growth. in macroeconomics, the theory may serve as the basis for analyses of varied contracting models, such as rewards of top managers, franchises, surplus insurance payments, and privatization of public enterprises. it has been effective in decision-making on investments and incentivizing economic activities through a choice of optimal efforts. the sagadiev theory of surpluses’ exchange. among kazakh economists, the specific subject matter of a surplus exchange has been thoroughly studied by nurlan sagadiev. sagadiev (2004) concluded that neither special effects nor production costs, nor income, determine the final outcomes of market relations. the sagadiev approach resonated with the theories developed at different times by williamson and hart. sagadiev developed linkages in the formation of the surplus value and connected those to the sales profit. to show the appropriateness of his theory’s stance, sagadiev made a reference to the research by d. rosenberg, who provided his valuable comments relating to k. marx’ capital as follows: “it may seem that k. marx did not provide all the components of the whole chain that are necessary to transit from simple commodities turnover to the production of surplus value. it may also seem that he did not explain the details pertaining to the trade. those have been named by prof. n. sagadiev as a perfected cycle of commodities, which sets the grounds for capital. according to him, profit is a category, which is unknown for a casual turnover of commodities. it has first emerged in trade. in this regard, the then k. marx had proceeded to investigating the details of the surplus value, which was based on the mature capitalist method of production. the profit faced by sellers at different times, relating to the existence of trade, had various sources. those, sometimes, were mere robberies” (rosenberg, 1984). sagadiev (2004) concluded that “consumption, which delivers wealth, and labor, which creates wealth, in fact, are the characteristic features of wealth. accordingly, there exists the difference between the cost of the commodity and its consumer cost. and, that may be named as the surplus value in order to be abstracted from the postulates formulated by the marxian surplus theory. in other words, it is that component of the surplus value of wealth, which surpasses the cost”. additionally, sagadiev considered that condillac not marx was the first to discover the true source of creation of the surplus. it was condillac who questioned the surplus, the phenomenon, which came to be the subject of exchange. according to sagadiev (2004, p. 90), the latter should be regarded as an alternative to the paradigm of the cost, which has governed economic thinking since the times of aristotle. the results of sagadiev’s work and the logic of his methodological approach captured researchers’ attention, especially, when he presented his theory of contracts. in particular, he stated the following: “by purchasing a commodity, a buyer enters into the relationship with a producer of the commodity. the producer wants to be rewarded for his labor. therefore, his cost is based on the volume of those benefits that are necessary for maintaining his labor continually. after the sales of his commodity, the producer becomes a buyer and enters into the relationship of exchange, possibly, with the same seller of a consumer commodity. as a buyer, the producer is interested in the consumer cost of the commodity. by means of the consumer cost of the commodity, he should reward the costs of his own labor. no doubt, the demand of a producer of commodities is not for one specific commodity but for a set of commodities. the labor costs are rewarded by some set of consumer costs. at the same time, if we consider the aggregate of all commodities as one commodity, and also, all sellers of commodities as one seller, then we can consider the act of the exchange as one deal where the aggregate consumer cost of the aggregate commodity is greater than the labor cost of the commodity. the cost, which the producer receives, in exchange of his produced the contract theory: a 2-dimensional reflection on commodities & capital 28 copyright ©2018 assa. adv. in systems science and appl. (2018) commodity, in the form of money, becomes the means of the evaluation of the consumer cost of the commodities that he needs for his life. the metamorphosis of commodities for the producer concludes by the fact that the producer receives one good instead of the other: т-d-т. the producer materialized it in the capacity of the cost, i.e. as that volume of consumption, which is necessary to reward his expenses, in exchange of the consumer cost of another commodity, which satisfies his needs. eventually, for the producer, the three different commodities have similar costs of exchange because those have been exchanged for the similar quantities of money. however, in one instance of exchange, the subject of exchange has reflected the true cost of the commodity and, in another instance, the subject of exchange has turned out to be the consumer cost of the commodity. the final metamorphosis of commodities for the seller acquires the reflection of the following sequence: ‘d-т-d’. here, all the evolutions went upside down. one and the same commodity had different exchange value. in buying the commodity, the buyer should have paid its cost. however, the same buyer was selling the commodity at its consumer cost. the difference in money had been the consumer surplus cost. at first, such method of creating the surplus value may seem to depend on the will or whim of certain individuals. in the real world, things do happen in accordance with this particular method. however, things are not solely confined to such a method. a buyer and a seller are presented in the personified actors of economic activities. the role played by a seller has been in converting the production costs to consumer costs. in this regard, the producers have been in need of sellers. the sellers have been in need of producers. to perform his duties, a seller has to have money and commodities. to ensure such possessions, he has to accumulate excess commodity and excess money in the form of consumer surplus costs. as mass volumes of commodities and money are being exchanged through sellers, the aggregate seller in the eyes of an autonomous producer is associated with the world of monies and commodities. the wider the exchange of commodities and monies is, the wider the division of labor, and the more powerful the role of the seller in the life of the society. by carefully looking at the evolution of the capitalist production relationship, specifically, the example of england, one may observe that trade surplus has been the first ever to create the initial form of the accumulation of capital (sagadiev, 2004, p. 90). by substituting ‘economics of the buyer’ in sagadiev’s theory with ‘economics of the currency-financial sector’, we may see the full picture of the cycle of the reproduction processes of capital in its money form ( ), and capital in its commodity form, ( ). 6. a geopolitical model of developing countries: evaluating real sector growth as mentioned earlier in this paper, thomas piketty noted that the concepts of ‘inflation and growth’ have not always been accurately defined: the division of the nominal growth into ‘real’ and ‘inflation-derived’ components are arbitrary, and thus may be disputed. the reasons underneath the inaccurate definition of growth and inflation, in our view, are not related to the low quality of the measurement tools of the balanced economic growth. instead, they are linked to the necessity of balancing them against real time conditions of the globalization process of the world economy. what piketty reflected was about the models developed on the basis of one dimensional measurement of macroeconomic indicators of the sectors of developed countries, where money of the current year has been measured by money in the previous year. according to 29 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) piketty, the traditional tools of economic analyses do not ensure full compliance of macroeconomic interests in the development of the financial sector with the real sector interests. that entails certain risks. the development the economies of the uk and france are examples that show that such incompliances occur in spite that these economies move along the trajectories of balanced economic growth. moreover, such risks tend to increase. so, any application of the traditional analysis techniques is subject to effects of risks. the need for qualitatively new analytical tools in measuring and evaluating indicators of balanced economic growth is there. the construction of growth models for developing countries is linked to the marker capacities of all types in managing their growth rates, or in reducing the material intensity of the gross product. under the condition of the globalizing world economy, the real economy mainly depends on real proceeds as well as on the efficient utilization of intermediate consumer goods and natural resources. as reflected earlier in this paper, macroeconomic models for defining the nominal gdp are incapable of solving the current challenges related to accounting the material intensity of the gross product in view that they are not defined as part of the whole after deductions of current costs of materials and costs of wear and tear of fixed assets. potential solutions to the outstanding challenges regarding the regulation of the costs stemming from material expenditures and the costs of utilizing natural resources lay beyond the domain of macroeconomic regulators. problems in this area cover the linkages of the macroeconomic indicators and are directly linked to the solution of environmental problems and the issue of green economy. that means, if we want to develop green economy and maintain a healthy environment, we should not only be directly involved in growth of current incomes and profits, but also responsibly engage in efficient utilization of materials and natural resources. stating the problem: the reasons behind the derivation of unbalanced indicators of growth in the above-noted sectors of economy are linked to the use of the outdated tools of assessment and evaluation of the indicators of the balanced economic growth that do not allow full compliance of interests of the real sector economy with interests in the financial sector. according to the monetarists, the keynesian model of balanced growth was immature because in it capital in its money form played a secondary role. a much improved model, built later by keynes, is presented as follows: . in the initial formula, ‘рр’ stood for purchasing power of the national currency unit. in the improved model, the equation has been upgraded as follows: where ‘b’ stood for the gdp deflator. both of these models suffer from one-sidedness in measuring the indicators of the real and financial sectors of an economy, which define the volume of the real final product. the methodology of constructing these models is based on the assumption that prices for goods and services remain unchanged as well as the velocity of money. that followed the one-sided principle of measuring the quality of balanced indicators of economic growth. and that principle paved the way to the introduction of arbitrary definitions. there also is a third model named after mandell and fleming. however, its close similarity to the keynesian models can easily be spotted. the only difference between them is in that the latter are derived from the unit of a world’s reserve currency instead of the units of individual national currencies. the grave one-sidedness of measurements in the afore-mentioned models has been formed by the three concurrent economic theories. the first is the marginal inutility theory, the contract theory: a 2-dimensional reflection on commodities & capital 30 copyright ©2018 assa. adv. in systems science and appl. (2018) which defines the price for goods and services by supply and demand. the second is the labor cost theory, which defines prices for goods and services by the costs required for producing the commodity following the marxian expenditure method. the first theory presents income based on annual product without accounting for materials and capital costs required for production ( ). the second presents annual product by balancing the costs of production ( ). in short, both approaches suffer from limitations of one-sidedness and one-dimensionality in measuring the indicators of balanced economic growth. one of them focuses on income in the annual production of capital in its monetary form, which is defined by using the macroeconomic approach based on . the other focuses on the capital in its commodity form, which is defined by the macroeconomic approach based on . the piketty model of balanced economic growth is developed based on economic laws of development. to resolve the problem of measurements inefficiencies of inflation and growth indicators, piketty attempted to use the cobbs-douglas production function by taking the aggregate elasticity factor within the limits between 1.3 and 1.6. however, these limits do not conform to the two economic laws given in his work (piketty, 2015, p. 225). the first of piketty’s laws is related to the dynamics of proportionality of the national capital and national income (piketty, 2015, p. 67): , where denotes the capital profitability, β the accumulated capital expressed in years of the national income α. the second law, named as the law of cumulative growth and of cumulative profitability, defines the size of the cumulative capital as the proportion of the form of the accumulation ‘ ’ to the pace of the economic growth ‘ ’: (piketty, 2015, p. 171). according to these laws and laws on population growth, any insignificant increase in the capital profitability (which prevails in the growth of the economy over the lengthy period of time) leads to significant increase in the growth of capital and, thus makes a destabilizing impact on the structure and dynamics of the social inequality (piketty, 2015, p. 90). however, these piketty’s laws do not account for linkages between macroeconomic and microeconomic indicators of growth, while the linkages are, in fact, products of reproduction of capital in its income form, and capital in its commodity form within cycles of the development of common good. the geopolitical model for analyzing the management efficiency of public-private partnership projects: our discussion above shows that the evolution of piketty’s economic laws, as the principle of defining the model structure of balanced economic growth based on input-output tables, enables the use of the three-dimensional method to measure the indicators of the market equilibrium. in the following macroeconomic analyses, with due reference to equ. (2), which represents part of the three-component equ. (1) of the macroeconomic contracts theory, we will justify the macroeconomic origin of the contents of equs. (3) (8). the macroeconomic contents of equ. (2) have been revealed by using the algorithm of scaling (akimov, 2014), where the algorithm of scaling vectors and matrices represents the measurements of the model of intra-industries balance reflected in input-output tables. let represents the labor intensity of the gross output, and the labor intensity of the gross product. then the concurrent measurements of aggregate labor intensity may be written respectively as and , which indicate respectively the labor intensity of the final and gross product to the i th extent of the sector ( ) of the economy. so , and can be written as follows (akimov, 2014, p. 176): and (9) 31 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) then we obtain the set of equations for the full cost of physical labor that has been utilized for the production of the unit of the final product for each of industries (akimov, 2014, p. 176): (10) thus, the full labor cost of producing the final product is defined by scaling the product of direct labor intensity of the gross output of all industries by relevant columns of the matrix that reflect the full labor cost in the model of intra-industries balance, which is represented in monetary form. the true cost of money is obtained by solving the concurrent problems of the threedimensional measurement of the balanced economic growth, on the one hand, and the product’s labor intensity on the other, where the unit of money is where ‘ ’ stands for the coefficient of the science-and-technology potential of the country, of which the increment: ( , ) is defined by either the margin of changes in the price for goods and services or the measurement of rent stemming from the utilization of natural resources. in that case, the real growth rate of the science-and-technology potential of any country may be measured by the difference between the growth rate of the production cost of the final product and the cost of resources that have been utilized in the production. it is because both had been sourced from the common pool of time, spent for labor ( ): (11) since equ. (1) defines the difference between the three growth rates of different indicators of the national economy, given the equality of the time, spent for labor that had been utilized for production ‘l’, it may be represented as an indicator of acceleration of the economic growth . so, the whole of the science-and-technology potential may be defined by . 7. a model for the science-and-technology potential by using gdp deflator models of the science-and-technology potential may be built on the basis of converting the monetarists’ equation of exchange. let the gdp deflator be expressed by based on the monetarists’ formulation. by multiplying both sides with the purchasing power of the national currency ‘рр’, we obtain the following: which is a qualitatively new expression of the balanced economic growth and defines the real volume of the final product – : which represents the nominal gdp that defines the cost of the final product and its product with the true cost of money represents the real final product. if , then the purchasing power of money may be defined under the new economic law as follows: the contract theory: a 2-dimensional reflection on commodities & capital 32 copyright ©2018 assa. adv. in systems science and appl. (2018) in this case, the coefficient ‘с’ may be defined as the proportion of the direct labor intensity to the full labor intensity ( ), or as the proportion of the cost of the utilized final product to the aggregate cost of resources utilized in the production ( ). however, the assessment of the coefficient is subject to that measuring the real gdp is done by following the method of three approaches. the first is the income method of a. smith, according to which, capital has only one dimension, its money form. under the one-dimensional approach to measuring the indicators of the balanced economic growth, the coefficient of the science-and-technology potential ‘ ’ is defined as follows: under the smithian theory, the purchasing power of money is defined by: , and the gdp deflator by: . the second approach is represented by the marxian expenditure approach, according to which capital has three dimensions. in this case, capital in its money form has the money dimension and capital in its commodity form has the labor dimension. so under the threedimensional measurement of indicators of balanced economic growth, the coefficient of the science-and-technology potential ‘ ’ is defined by : where represents the price index for goods and services, and the indicator index ‘ ’, and the material cost of producing the final product . since the real final product ‘ ’ is equal to , all the previously discussed cases present an opportunity to re-assess the true cost of the quasi-real gdp by using the following formula: here, the main equation of assessment of the real final product ‘ ’ at the macroeconomic level is fully defined owing to the coefficient ‘ ’, as a result of analyzing the inputs and outputs at the microeconomic level: the result of analyzing the assessment and evaluation of the impacting efficiency of regulatory policies pertaining to the development of the national economy is presented in the equation, which defines the mutual convertibility of the cost of capital in its commodity form, back to its monetary form as follows: (12) 33 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) 8. economically assessing the efficiency of innovative technologies in the monetary dimension the difference between the marginal growth rates of capital in its money form in macroeconomics, and capital in its commodity form, is defined by the following formula: (13) for simplicity, let (standing for acceleration), representing the difference between the three growth rates of the cost of the final product as an indicator of capital in its monetary form and the output of goods and services as an indicator of capital in its commodity form . it may be named as the marginal coefficient of the science-and-technology change. this indicator of the acceleration ‘a’ emerged as ‘deus ex machina’, given equality of labor and capital that have been utilized for producing the final product ( ) in macroeconomics, and output of goods and services ( ) in microeconomics. so, in formulaic terms the above stated may be presented as follows: (14) the final result evolves from inside of the socio-economic system, which is defined by juxtapositioning the outcomes of the intra-industries balance models of the development of a country, as reflected in labor and monetary dimensions. such an effect of the science-andtechnology potential of a country, as portrayed j. clark, relating to the macroeconomic dimension has been named as the enterprisers’ profit by (baizakov and oinarov, 2015), and the surplus profit in the reproduction schemes, as lenin formulated, with accounting for the scientific-technical progress (baizakov and oinarov, 2015, p. 172). 9. convert science and technology effect in its monetary form to its labor form equ. (14) may be rewritten by representing it as the difference between the growth rates of capital profitability in its form of the cost of the final product , and capital in its form of the resources utilized for production . such this difference represents the acceleration of the input of science-and technology potential in the development of a national economy ‘ ’, defined by equating the costs of labor time in hours, days, years ‘ ’. accordingly, the overall potential of science-andtechnology innovations in the aggregate expression ‘ ’ is defined by the product of the total labor time ‘ ’ and the indicator of the acceleration of economic growth such effect is defined by the difference between the three growth rates of the varying indicators of the development of the national economy. if the input of science-andtechnology potential of the country was previously defined in obscure terms, a scaling effect or a solow residual now enables an accurate assessment by employing the cost of the actually utilized labor time in production, labor productivity, capital profitability, and the coefficient of the science-and-technology potential. additionally, each entrepreneur may readily perform formulae-derived calculations corresponding to types of his/her economic activities. the contract theory: a 2-dimensional reflection on commodities & capital 34 copyright ©2018 assa. adv. in systems science and appl. (2018) 10. convert the science and technology effect in its labor form to its energy components the measurement ‘ ’, which represents the product of the number of people employed in an economy ‘ ’ and the indicator of the acceleration ‘ ’, may also be considered as the capacities in the energy units. according to the fao data, an average person needs a minimum 1,800 kcal (7,500 kilojoules) per day. each country has its own conversion coefficients on the efficiency of labor time in energy unit, which takes into account of climatic conditions of living in the country, the type of economic activities the citizens are engaged in, ages, genders, etc. example, in the uk, an average aged female would consume 2,200 kcal per day, a man 2,500 kcal. it adds up to 2,350 kcal per day, on average. these indicators in the usa would be: female 2,200 kcal, man 2,700 kcal, which adds up to 2,400-2,500 kcal per day on average. the energy formula used to define the science-and-technology potential becomes clear and comparable with formulae on measuring based on the thermodynamics theory and relativity theories. those measurements with the above-noted impacts in money, labor time, and energy units, turn out to be authentic with each other. they clearly define the level of the development of the production forces and that of capital. assessment of the energy impact, as well as that of money and labor, is defined by comparing the growth rates of key indicators of the macroeconomic dynamics, capital in its money form, and capital in its commodity form. the formula for the energy efficiency of innovations in the production and in the innovations technology is measured by , where ‘ ’ stands for the energy impact measured in kcal, kilojoules, кwт, and ‘ ’ the number of people employed in the economy, also measured in (kcal, kilojoules, кwт), and the indicator ‘ ’ the acceleration, defined as the difference between labor productivity and capital measured by the final product, and labor productivity as well as capital profitability measured by costs of the gross product. the unit of the indicator of acceleration is the relative measurement defined as the difference between the three productivities measured by their growth rates. the novice of the present research is in its ability to enable the measurement to be interpreted as the net input of the science-and-technology potential of the economy, in developing the formula where labor time utilized for the production of product ( ) and of goods and services ( ) enables the measurement to be considered as a net contribution of the science-and-technology potential in the economy, which in turn is defined as the difference between the marginal labor productivity in the final product and the marginal labor productivity in the aggregate costs of production. the afore-described formula may serve as a solution to the granberg puzzle. such solution has been obtained by science-driven measurment of inputs and outcomes. in measuring the macroeconomic dynamics, if the indicators of one of the threecomponent dimensions, either capital in its money form or capital in its commodity form, are missing, the resolving the granberg puzzle will become impossible. an application of the full, three-component matrix of the intra-industries input-output tables is the necessary prerequisite for solving the granberg puzzle. no doubt, not always, the science-and-technology potential of a country yields positive impact. it is because not every investment can ensure sustainable enterprisers’ profits evenly across all industries. in this regard, the negative surplus profit is often being registered at the national economy level. moreover, not always, the invested capital works at its maximum capacity, and often the expected impact proves to be nil. in order to avoid such unexpected 35 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) outcomes, the assessment of the input of the science-and-technology potential into an economy at the country level is needed. the cycles theory versus the contract theory: the literature suggests the following three approaches that are applicable in explaining natural cycles. a) a cycle is a phenomenon that is external to the economic system; and its evolution is influenced by noneconomic factors. such factors may be solar radiation, political shake-ups, revolutions, demographic booms, inventions and innovations, influencing the economic and environmental milieu for production processes; b) a cycle is an internal phenomenon, driven by causes endogenous to the economic system, initiated by the self-reproduction of economic cycles, such as demand and benefits, consumption and investments; and c) a cycle is a multispectral phenomenon, driven by a synthesis of both internal and external factors that occasionally but sustainably impact the economic system. according to american researchers, 1,380 types of economic cycles have been identified. the duration of these cycles vary from 20 to 700 years (klinov, et al., 1989, p. 7). from the practical point of view, there exist different classifications of cycles. the following two criteria are given in the core of the above-referred classification: (1) the duration of cycles; and (2) the driving forces of cycles. according to the first criterion, we suggest the following classifications: 1) seasonal cycles with the duration of one week to several months; 2) business cycles with the duration of one to several years; 3) the kitchin cycles with the duration of 3 to 5 years; 4) the jugular cycles with the duration of 7 to 11 years; 5) the kuznets cycles with the duration of up to 20 years, and 6) the kondratiev cycles with the duration of 40 to 60 years. the driving forces of business cycles are primarily linked to investment activities. the mechanism underneath the market dynamics is vested on principles of acceleration and multiplier effects. the principle of acceleration suggests that the scale of investments depends on the increment and changes in the demand versus the final product. business cycles have well been studied in the western economic science, for details please consult with the works by p. samuelson, j. hicks, and others. short-term cycles with the duration of 3 to 5 years evolve out of the dynamics proportionate to the size of the reserves, which consist of materials and commodities stocked by firms. they have been named in honor of d. kitchin, an english researcher who first stidied how business cycles work. he particularly noted: “first, they emerge as outcomes of investing in materials, raw resources, and stock capital in an intention to make the best use of the market demand. gradually, the demand wanes off, and capital, invested in stocks, becomes excessive. stock investments sharply diminish in size, and the balance between stocks and demand slowly restores. thus, the cycle helps restore the market equilibrium of supply and demand" (menshikov, 1989). mid-term cycles, also known as industrial cycles, having the duration of 7 to 11 years, are related to renewal of fixed assets, i.e. inventory, equipment, facilities, cars. the life cycle of those assets depend on the extent of their wear and tear, while the latter defines the factual duration of the mid-term cycle. the important role in sustaining the cycles of this type is played by the scientific-technical progress. along similar lines, other types of mid-term cycles evolve as outcomes of mass revolutionary innovational technologies. those cycles are transmitted from one industry to another across the entire chain of industries in a domino effect, thereby completely changing the very foundation of the production (klinov, et al., 1989, p. 26). the contract theory: a 2-dimensional reflection on commodities & capital 36 copyright ©2018 assa. adv. in systems science and appl. (2018) s. kuznets developed the construction cycles theory for the cycles lasting 15 20 years. these cycles are linked to periodical mass renewals in housing and dwellings. y. rostow, h. biskhar, a. kleinknekht, among others consider the kuznets cycles as specific characteristics of the american economy only because they reflect huge migrants’ inflows and their construction-related activities. the super cycles theory or k-waves, established by kondratiev (1989), is one of the most promising directions in exploring long-term tendencies in the development of world economy. this area, by some reasons, has remained a less investigated domain compared to short-term and mid-term cycles. even so, the practical interest has gravitated towards the study of long-term changes in an economy since relatively recent past. of interest is also the account of managing the processes of production and sales of produce. such tendencies, often repetitive in nature, are clearly subjective. however, the lengthier is the tendency the slower is the process of accumulating statistical data and information required for uncovering and investigation of such tendencies. the economic policies that are currently adopted in developing countries are mainly designed on the bases on the monetarist’s concepts that are strictly aligned with the requirements of international monetary institutions and organizations. they, however, might lead to continual stagnation of developing economies and ineffective utilization of economic potential of these countries. based on the afore-discussed observations, equipped with the common aim of painlessly overcoming economic hardships imposed by a series of economic and financial crises in order to create prerequisites for the subsequent more sustainable economic growth of developing countries, a timely and effective change of the models of economic development for developing countries is much needed. in this regard, this paper specifically recommends to replace the monetarists-developed model for regulating economies by geopolitical model, which is primarily oriented at developing the function of the science-and-technology potential (stp) of a developing economy. if such model is accepted, the main indicators of the stp may objectively signal about the accumulation of negative consequences of any given economic policy in a developing country, undertaken at any given time interval, and identify the need to make remedial corrections in order to prevent an economic decline. the growth rate of labor productivity in terms of the cost of the final product and the growth rate of labor productivity increments in relation to the aggregate cost of productizing the final product may serve as the key indicators of the stp. the need for proceeding with active research in this direction is acute with the reasons listed below: the more is explored in the area of cyclical economic dynamics, the more accurate shall market economic forecasts be, and the more effective shall the impact of state be on the given economy. with appropriate knowledge in this research direction the governments of developing countries will be able to timely implement required economic policies in order to remove inefficiencies and work out action-based implementation plans, inclusive of investment, financial, credit, tax policies, as relevant. such sets of measures will help reduce negative and disastrous consequences of global economic and financial crises, and smooth out at least some of the cyclical nature of economic development. improving the pace of economic development towards its balanced and thus sustainable growth will ensure favorable conditions of the long-term prospective development and provide increasing rates of economic growth, which thereby improves the quality of life styles and wealth of people from around the world. as the main set of variations in assessing the short-, mid-, and long-term perspectives of economic growth, the growth model with kazakhstan’s regional specificities is suggested in this paper to be used as a pattern. the main criterion, which serves as the basis for replacing 37 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) one model by another is the economy’s potential of ascending economic growth. it may be defined, according to equ. (14), as follows: . indeed, the development of market-driven forces of certain sectors of a developing economy does not underscore essential and insignificant deviations in the growth rate of another economy of either developed and developing countries. however, the contract theory, as a reflection of economic laws, enables one to timely react to such deviations not only at the macroeconomic level, but also at the level of individual firms, be they small or medium sized. the point is that the dynamics of proportions between the annual cost of the final product ( ) to that of the commodities of intermediate consumption (qp) that have been utilized in the process of producing the output represents the productivity of commodities of intermediate consumption in terms of the final product. if this indicator is denoted by ‘ ’, we then obtain the following function of time: . the system of national accounts of world countries reflects the costs of final product ( ) as well as of intermediate consumption ( ). in other words, both these indicators are known for calculations. that means that one can easily derive the dynamics of change in the productivity of commodities of intermediate consumption by using the final product. analyses of economies of both developed and developing countries reveal the fact that this particular function complies with the definition given by michael porter in terms of the productivity of the materials resources 3 . it defines the cyclical nature of the economic development of national economies of the world as well as the duration of the cycle, and also, the driving forces of the cycle, which are defined by the input of the stp in the overall economic growth rate. thus, in an ideal case, possible values of the productivity of intermediate consumption goods in terms of the final product, μ, are positioned within the range from 0.1 to 9.0. in a concrete case, one may define equations as functions of time, and identify the form of possible deviations. 11. growth rate of stp as a function of the productivity of intermediate consumption goods what is surprising is the fact that the growth rate of the stp, which is the coefficient ‘ ’, is a function of the productivity of the intermediate consumption goods ( : . which can be rewritten as follows: . if we let , , and , then the function of the stp takes the standard form of . in this particular case, the equation, involving the stp may be expressed as follows: . to chart of the function of the stp, which is equal to , we may involve the range of indicators of the independent variable ‘ ’ and then derive from the relevant values of the dependent variable ‘ ’. 3 it is michael porter who introduced the term ‘productivity of resources’, of which the reverse meaning defines the efficiency of production resources. special attention needs to be paid to the shift of ‘the focus from low wages to the overall low costs’ as the most important factor of the ‘development of strong local competition’ along the path of ‘moving towards the developed economy’. porter, m. e. (2002). on competition. in russian. m: williams publishing house. st. petersburg, moscow, kiev. (p.496), (p.220). the contract theory: a 2-dimensional reflection on commodities & capital 38 copyright ©2018 assa. adv. in systems science and appl. (2018) by using , , and , we can construct the following table by deriving the values of productivity of intermediate consumption goods from final product ‘ ’ within the range from 0.1 to 9.0: table 1. the stp function 0.45 1.45 -0.69 0.31 3.60 4.60 -0.22 0.78 0.50 1.50 -0.67 0.33 4.10 5.10 -0.20 0.80 0.55 1.55 -0.65 0.35 4.60 5.60 -0.18 0.82 0.60 1.60 -0.63 0.38 5.10 6.10 -0.16 0.84 0.65 1.65 -0.61 0.39 5.60 6.60 -0.15 0.85 0.10 1.10 -0.91 0.09 6.10 7.10 -0.14 0.86 0.60 1.60 -0.63 0.38 6.60 7.60 -0.13 0.87 1.10 2.10 -0.48 0.52 7.10 8.10 -0.12 0.88 1.60 2.60 -0.38 0.62 7.60 8.60 -0.12 0.88 2.10 3.10 -0.32 0.68 8.10 9.10 -0.11 0.89 2.60 3.60 -0.28 0.72 8.60 9.60 -0.10 0.90 3.10 4.10 -0.24 0.76 9.10 10.10 -0.10 0.90 as a result, the stp function has been derived from intermediate consumption goods , which, by the nature of changes, is fully dependent on the productivity of intermediate consumption goods. figure 1 showed the three instances of development of the stp. in the first instance, the stp function has been reflected at the background of the growth rate of productivity of intermediate consumption goods (function 1a), which acquires values from 0.10 to 9.0 and shows the ascending trend. this curve is shown in diamond symbols in figure 1. fig. 1. the function of the stp and intermediate consumption goods in the second instance, when the function of the stp , given that the growth rate of productivity of intermediate consumption goods ‘ ’ (curve 1б) acquires values 0,0000 0,1000 0,2000 0,3000 0,4000 0,5000 0,6000 0,7000 0,8000 0,9000 1,0000 0,00 2,00 4,00 6,00 8,00 10,00 с μ с=μ/(1+μ) 39 s. baizakov, a.r. oinarov, d. eshimova, j. y.-l. forrest copyright ©2018 assa. adv. in systems science and appl. (2018) from 2.0 to 9.0, we assume, has the descending trend, which at point 2.0, acquires the value of 0.89 and, at point = 9.0 equals 0.07. this trend has been shown in square symbols in function 1б. three more possible instances, as reflected in figure 1, have been shown by direct lines that are parallel to the axis of the productivity of intermediate consumption goods . the first of them characterizes the real time situation in the developing market given that the stp coefficient remains stable (curve 1с). thus, figure 1 reflects the individual case that for any . the second trend line parallel to the axis of the productivity of interdemediate consumption goods transcends the points of intersection of both curves (for details, see function 1d). above this trend line, a set of the indicators of balanced economic growth of different countries has been reflected. as per the trajectory of changes in the productivity of interdemiate consumption goods ( ), as derived from the cost of the final product, it may serve as a forecast indicator for the analysis of causes of distortions in any economy, since the stp growth rate is defined depending on the dynamics of the changes in the productivity of intermediate consumption goods. 12. some final remarks generally speaking, the results of recent economic discussions are mostly derived out of either keynesian type models or monetary-policy type models, developed respectively by the followers of these respective schools. although the thoughts of these schools are still used in the practical management of the world’s economies, many of their results are not consistent with the modern realities of the globalizing economy. as a matter of fact, both of these types of models are special cases of our generalized model of market equilibrium, developed on the basic ideological positions of the “fifth way”, proposed by nursultan nazarbayev, the president of kazakhstan. even so, this does not means that the theories of these schools should be sent to the archive of the theories on economic growth and equilibrium. in particular, keynesian type models are more in line with the economic interests of development of real economic sectors so that they can continue to provide service as the theoretical basis for relevant decision makings. and the models of monetary-policy type can continue to serve the economic interests of developing the financial sector. since our generalized model proposed in this paper is obtained by integrating three different market equilibrium indicators used in the keynesian type and monetary-policy type models, it is expected that this new model can help harmonize relevant economic interests and therefore acts as an instrument to ensure the implementation of regulatory policies. references [1] akimov, n. (2014). ot capitalisma k capitalismu [from capitalism to capitalis. volume i]. almaty, kazakhstan: raritet. [in russian] [2] baizakov, s., oinarov, a.r. (2015). systema modeley sbalansirovannogo razvitiya rynochnoi economiki [the system of models of the balanced development of market economy: theory of construction]. astana, kazakhstan: center of public-private partners. [in russian] [3] clark, j. (1908). the distribution of wealth: a theory of wages, interest and profits. new york, usa: macmillan. [4] granberg, a.g. (1985). dinamicheskiye modeli narodnogo khozyaystva [dynamic models of a national economy]. мoscow, ussr: ekonomika. [in russian]. the contract theory: a 2-dimensional reflection on commodities & capital 40 copyright ©2018 assa. adv. in systems science and appl. (2018) [5] coase, r. h. (1960) the problem of social cost. journal of law and economics, 3, 1–44. [6] klinov, v.g. manukovskii, a.b., hartukov, e.p., tsygichko, l.p. (1989). voprosy teorii ekonomicheskoy konyunktury: ucheb. posobiye [the issues of the theory of economic environment]. moscow, ussr: msiir. [in russian] [7] kondratiev, n. (1989). problemy ekonomicheskoy dinamiki [the problems of economic dynamics]. moscow, ussr: ekonomika. [in russian] [8] kornelyuk, r. (2016, october 11). teoriya kontraktov: v chem sut otkryty nobelevskikh laureatov-2016 [the contracts theory: what if the core of the discoveries of the nobel laureates]. [online]. available: http://forbes.net.ua/nation/1422320-teoriya-kontraktov-v-chem-sut-otkrytijnobelevskih-laureatov-2016. [in russian] [9] menshikov, s. k., klimenko, l. a. (1989). dlinnye volny v ekonomike. kogda obshchestvo menyaet kozhu [long waves in economy. when a society renews its skin]. moscow, ussr: mezhdunarodnye otnosheniya. [in russian] [10] nazarbayev, n. (2009, september 29). pyaty put. [the fifth path]. izvestia. [online]. available: http://izvestia.ru/news/353298. [in russian] [11] piketty, t. (2014). capital in the twenty-first century. cambridge, ma, usa: harvard university press. [12] porter, m. (2008). on competition. usa: harvard business review. [13] rosenberg, d. (1984). kommentarii k «kapitalu» k. marksa [the comments on the capital of k. marx]. moscow, ussr: ekonomika. [in russian] [14] sagadiev, n. (2004). ponyatie stoimosti v contexte probleme universaliy v nauke. perspektivy nereduktivnoi filosofii. [the notion of the cost in the context of the problem of the universality of science. perspectives of the non-reductive philosophy]. almaty, kazakhstan. [in russian]. [15] smith, a. (2007). issledovaniye o prirode i prichinakh bogatstva narodov. [an inquiry into the nature and causes of the wealth of nations]. moscow, russia: eksmo. [in russian] [16] sraffa, p. (1960). production of commodities by means of commodities: prelude to a critique of economic theory. cambridge, uk: cambridge university press. [17] granberg, a. g. (2016). vasily leontyev v mirovoy i otechestvennoy ekonomicheskoy nauke [wassily leontief in the world and national economic science]. in suslov, v.i. & suspitsyn, s. a. (eds.) k 80-letiyu so dnya rozhdeniya aleksandra grigoryevicha granberga: ucheny, uchitel, chelovek, (pp. 153-184). novosibirsk, russia: ieie so ras. [in russian] http://forbes.net.ua/nation/1422320-teoriya-kontraktov-v-chem-sut-otkrytij-nobelevskih-laureatov-2016 http://forbes.net.ua/nation/1422320-teoriya-kontraktov-v-chem-sut-otkrytij-nobelevskih-laureatov-2016 http://izvestia.ru/news/353298 adv syst sci appl 2018; 1; 132-141 published online at http://ijassa.ipu.ru. copyright ©2018 assa. adv. in systems science and appl. (2018) neuroprotection and timely troubleshooting of electric drive equipment viktor m. buyankin bauman moscow state technical university, moscow, russian federation e-mail: viktor12buyankin@yandex.ru abstract:. the article discusses the neurodiagnostics drives with the prediction of the future fault is a relevant task to avoid failure of electric motors and actuators. such systems neurodiagnostics increase the reliability and safety of operation, anticipating contingencies. an important aspect in the management of complex dynamic objects designed to operate in harsh operating conditions, such as aggressive, explosive and dusty environments, extreme temperatures, increased vibration, is the use of real sensors. the use of state observers in such facilities will improve the operational reliability of the electric drive, avoid shock currents, reduce weight and size characteristics, etc. in recent years, there are high requirements for modern control systems of electric drives, namely: accurate speed control, maintaining high torque at low control speeds, limiting the starting and shock currents, high dynamic characteristics, high signal processing accuracy, improved coordinate accuracy. keywords: neural network, extended kalman filter, fuzzy logic, genetic algorithm, induction motor, asynchronous electric drive. 1.introduction today there are various methods of identification of parameters and variables of the state of the asynchronous electric drive which include the extended kalman filter; the observer created on the basis of a neural network and using fuzzy logic; genetic algorithms [12-16]. most well-known observers do not provide for maintaining the necessary accuracy of identification of parameters and variables of the state in the entire range of speed control at different modes of operation of the electric drive and non-sinusoidal forms of stator currents [9-11]. the last remark is typical when driving an asynchronous motor from a thyristor voltage regulator and an autonomous voltage inverter [2-5]. the purpose of the study presented in this paper is a comparative analysis of the most common in practice, the observer of the state of the asynchronous electric drive on the simplicity of their implementation (in terms of mathematical description) and reliability. on the basis of this analysis, practical recommendations on their application in this type of electric drive in terms of their efficiency are developed [1]. today, the use of real sensors in complex dynamic systems, and in particular, drives operating under severe operating conditions, is undesirable for a number of reasons, including a significant increase in installation, increase in weight and size characteristics, decrease in operational reliability, etc. [17-20] methods of identification of parameters and state variables attracted both domestic and foreign scientists [6-8]. to determine the parameters and state variables used by the observers based on extended kalman filter, fuzzy logic, neural networks and genetic algorithms. http://ijassa.ipu.ru/ neuroprotection and timely troubleshooting of electric drive equipment 133 copyright ©2018 assa. adv. in systems science and appl. (2018) 2. literature review the ann used to identify the parameters and state variables of complex dynamic objects is composed of three main layers: input, hidden and output. more than one hidden layers are possible. to date, the most commonly used activation functions of the neuron include threshold, linear, sigmoidal, tangential, radial-basis activation functions. in practice, a linear function is used as the activation function of neurons in the output layer. the first layer of the neural network is a relay. the activation function of the neurons of the hidden layer is mainly non-linear. according to the work [15], the most suitable activation function for neurons of the hidden layer is the tangential activation function. to identify the parameters and state variables of complex dynamic objects it is necessary to use a dynamic neural network, which includes a dynamic neural network with delay at the input [16], a jordan network [17], an elman network [18], the combination of a dynamic neural network [19, 20]. the peculiarity of these neural networks is the presence of signal delays at the input, output and both input and output, which allows to provide the best learning and filter out strong pulse interference. before training a neural network, the developer must decide on the array of data needed for its training, and the choice of training algorithm. an excessively large array of data from each of the input signals can lead to a retrain effect, as a result of which, within the training sample, the neural network gives a minimum evaluation error, and when working with a test sample (different from the training sample), the evaluation error is very large. to date, there are a large number of training algorithms, the main of which are the gradient descent algorithm, gradient descent algorithm with perturbation, moller's learning algorithm and levenberg–markvardt learning algorithm. the presented first three methods of training require small computing power of the computer, but are not able to find a global minimum learning errors. the levenberg-markvardt learning algorithm requires significant computing abilities of the computer, but is able to get out of the local minimum and find a global one. 3. materials and methods rapidly expanding the range of functional requirements for automated systems the electromechanical energy conversion in industry, transport, special equipment, stringent requirements for the dynamic characteristics of such systems and their energy efficiency would require the construction of control algorithms, which for some features can be called "intelligent". it is generally recognized that the synthesis of such algorithms is currently the central problem of the modern theory of automatic control of electric drives. the conflict between obviously insufficient methodical basis of construction of "intellectual" algorithms of management of ep and growing opportunities of hardware becomes more and more acute as the modern condition of means of power and information electronics already allows to realize elements of "intelligence" (intelligent control) even in rather inexpensive serial sau. according to the author, the sign of "intelligence" of the controlled electric energy converter for electric drive systems is not the use of "exotic" methods (fuzzy logic, neural networks, genetic algorithms, etc.), as it is often presented, but functional completeness in solving the main problems of modern control theory in the appendix to such complex objects as general industrial electric drive. "intelligence" in this sense – is to provide a comfortable interface between a person and a microprocessor control system, a minimum of manually adjustable parameters of the ep, but in any case not the ability to independently set and solve new non-standard problems. particularly important tasks of identification and adaptive control become in the construction of common industrial frequency-controlled electric drives, one of the main requirements for which is the rejection of the use of external in relation to the controlled source of electrical energy (pm) sensors, including sensors directly controlled coordinates of mechanical motion. 134 v.m. buyankin copyright ©2018 assa. adv. in systems science and appl. (2018) 4. results and discussions electric drives are used in many branches of production such as machine-tool building, mechanical engineering, in the mining and oil-extracting industry. the motors work with a variety of loads in various modes: long lasting, intermittent, short-term modes. during operation, the drive equipment wears out over time, which leads to deterioration of static and dynamic characteristics, and sometimes to emergency situations. therefore, the prediction and timely detection of faults of electric equipment is an urgent task. however, predicting the malfunctions of the electric drive equipment is quite difficult and time-consuming task. a large number of diagnostic systems have been developed for fault detection, but it is not possible to determine the whole range of faults of the electric drive equipment by one hundred percent. to improve the quality of diagnostics, we propose systems with neural networks that have proven themselves quite well in pattern recognition and approximation of complex nonlinear dependencies [1]. fault diagnosis by many criteria coincides with pattern recognition and therefore, using neural networks, it is possible to achieve higher results of fault detection of electric drive equipment compared to other diagnostic systems. there is a wide variety of neural networks that differ from each other in their advantages and disadvantages. when designing a system neurodiagnostics stop your choice on neural networks like: newff and anfis [2]. that is, we will develop a combined neuroprotection system with an expert neural network approach. for initial identification of faulty components of the electric drive is advisable to use neuroprogenitor on the basis of neural networks like: newff and anfis, and for a more detailed examination of the health elements of the drive it is advisable to use the system neuroprogenitor with expert neural network approach fig. 1 fig. 1. system neuroprogenitor neuroprotection and timely troubleshooting of electric drive equipment 135 copyright ©2018 assa. adv. in systems science and appl. (2018) the purpose of building a neural network expert system is the initial prediction of faulty electric drive units, which consists of a neuroregulator, a power converter, an electric motor, a load mechanism fig. .2 fig. 2. block diagram of electric drive the operation of the actuator must comply with the rated static and dynamic characteristics. at rated load, the actuator must have a nominal armature voltage, rated current, nominal speed, which is displayed on the family of mechanical characteristics fig. 3 deviation from the nominal parameters leads to malfunctions, and sometimes to emergency situations of the electric drive equipment. we will consider as faulty the electric drive of deviation of parameters which exceed the maximum deviations. 136 v.m. buyankin copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 3. mechanical characteristics of the electric drive will develop a system of neurodiagnostics in the beat of a forward current in the excitation winding iob(n+1) (and voltage in the armature of the motor (s) using a neural network newff [3]. suppose there is a database containing the current values of the iov table 1, fig. 4. table 1. data structure iов(n-9) iов(n-8) iов(n-7) iов(n-6) iов(n-5) iов(n-4) iов(n-3) iов(n-2) iов(n-1) iов(n) iов(n+1) 203ма 192ма 203ма 230ма 200ма 198ма 205ма 199ма 190ма 206ма ? fig. 4. value of current iов neuroprotection and timely troubleshooting of electric drive equipment 137 copyright ©2018 assa. adv. in systems science and appl. (2018) six lines of training samples for a neural network consists of four inputs i1, i2, i3, i4 data, and the desired output is the data of the subsequent clock i5 (table 2). the neural network training was carried out in the matlab environment. table 2. matrix for training the neural network 1i 2i 3i 4i 5i 203 192 203 230 200 192 203 230 200 198 203 230 200 198 205 230 200 198 205 199 200 198 205 199 190 198 205 199 190 200 the newff predictive neural network has four inputs and one output (fig. 5). fig. 5. newff predictive neural network system neurodiagnostics in tact ahead of the current in the excitation winding i5, iов(n+1) described by the following system of neuromante [4]. ,tan , ,tan ,tan ,tan ,tan , , , , ' 05 ' 1 ' 44 ' 33 ' 22 ' 11 ' 0 44 33 22 11 44444334224114 334433332213113 22442332222112 11141131121111 sigyi bwrwrwrwry siger siger siger siger bwiwiwiwie bwiwiwiwie bwiwiwiwie bwiwiwiwie = ++++= = = = = ++++= ++++= ++++= ++++= (1) 138 v.m. buyankin copyright ©2018 assa. adv. in systems science and appl. (2018) where i1, i2, i3, i4 the input signals of the neural network; i5 the output signal of the neural network; x0 the input signal of the neural network; y0 the output signal of the neural network; y1, y2, y3 the output signals of the neural network were detained in 1,2,3 tact; e1…e10 output signals of the first layer of neurons; w1…w215 weight of the first layer of neurons; b1…b10 displacements of the first layer of neurons; b1…b10 signals at the output of blocks of activation of the first layer of neurons; y0` signal at the output of the second layer of neurons; w1`…w10` weight of the second layer of neurons; b1` the offset of the second layer neurons, the signals output from the neural network were detained in 1,2,3 tact; displacements of the first layer of neurons; r1…r10 signals at the output of blocks of activation of the first layer of neurons; y0` signal at the output of the second layer of neurons; w1`…w10` weight of the second layer of neurons; b1` displacement of the second layer of neurons. after modeling and neural network learning algorithm fig. 6 we obtain the required weights and biases [5]. fig. 6. neural network modeling and learning algorithm weights of the first layer of neurons: w11= 0.0218, w12= 0.0203, w13= 0.0186, w14= 0.0213, w21=-0.0019, w22=0.0213, w23=0.0074, w24=0.0202, w31=-0.0097, w32=-0.0171, w33= -0.0092, w34=-0.0329, w41=0.0148, w42= -0.0029, w43=0.0252, w44=0.0145. weights of the second layer of neurons w’1=40.5136, w’2= 40.4765, w’3= -39.8220 w’4= 40.4293. displacements for the first layer of neurons b1= -5.9110, b2=-0.1183, b3=1.4517, b4=-0.9472 neuroprotection and timely troubleshooting of electric drive equipment 139 copyright ©2018 assa. adv. in systems science and appl. (2018) displacements for the second layer of neurons b’1=38.7585 for fig. 7 the dependence of learning errors depending on the number of epochs is given. fig. 7. the dependence of the training error depending on the number of epochs when testing and predicting neural network, we get the future value i5=200.0000 after predicting the future value iов(n+1) compare it with iовн. if the difference iовн и iов(n+1) greater than the maximum value allowed, the variable x1 assigns 1 otherwise 0. a value of 1 indicates an emergency condition in the motor excitation winding. 5. conclusion based on the data provided about the state observers, it can be concluded that it makes no sense to allocate a specific identifier. each of them has its advantages and disadvantages. everything will depend on the area in which the observer will be used and what performance criteria he or she should support. robustness have the majority of observers state. state observers, implemented on the basis of a mathematical model, require knowledge of the internal parameters of the object of identification. for creation of observers of a state at the same computing abilities of the computer the observers constructed on the basis of the genetic algorithm and the extended kalman filter are the most laborconsuming. the work of such observers is possible for most identifiers, except for the genetic algorithm. the scope of application of state observers is quite wide – from social sciences to technical sciences (up to the construction of new computer architectures). 140 v.m. buyankin copyright ©2018 assa. adv. in systems science and appl. (2018) references [1] abbas j.j., chizeck h.j. (1993). neural network control of functional neuromuscular stimulation systems. ann biomed eng., 21(4), 459-460. doi:10.1007/bf02368636. [2] asakawa s, kyoya i. (2013). hopfield neural network model for explaining double dissociation in semantic memory impairment. bmc neurosci, 14(1), p233. doi:10.1186/1471-2202-14-s1-p233. [3] baddeley b, graham p, husbands p, philippides a. (2012). a neural network based holistic model of ant route navigation. bmc neurosci. 13(1), o1. doi:10.1186/14712202-13-s1-o1. [4] chik d, borisyuk r. (2009). spiking neural network models for memorizing sequences with forward and backward recall. bmc neurosci. 10(1), p211. doi:10.1186/1471-220210-s1-p211. [5] jaramillo-avila u, rostro-gonzález h. (2015). spiking neural network configuration designed for switching between basic forms of movement in a biped robot. bmc neurosci, 16(1), p104. doi:10.1186/1471-2202-16-s1-p104. [6] kuo r.j., tseng y.s., chen z-y. (2016). integration of fuzzy neural network and artificial immune system-based back-propagation neural network for sales forecasting using qualitative and quantitative data. j intell manuf., 27(6), 1191-1207. doi:10.1007/s10845-014-0944-1. [7] kutschireiter a, surace sc, sprekeler h, pfister j-p. (2015). approximate nonlinear filtering with a recurrent neural network. bmc neurosci., 16(1), p196. doi:10.1186/1471-2202-16-s1-p196. [8] li c, xu j, xue l. (2001). knowledge-based artificial neural network models for finline. int j infrared millimeter waves, 22(2), 351-359. doi:10.1023/a:1010760707665. [9] li y, pu y, xu d, qian w, wang l. (2017). image aesthetic quality evaluation using convolution neural network embedded learning. optoelectron lett., 13(6), 471-475. doi:10.1007/s11801-017-7203-6. [10] lin w, liao x, deng j, liu y. (2016). land cover classification of radarsat-2 sar data using convolutional neural network. wuhan univ j nat sci., 21(2), 151-158. doi:10.1007/s11859-016-1152-y. [11] manoj k, charul b. (2017). hybrid tracking model and gslm based neural network for crowd behavior recognition. j cent south univ., 24(9), 2071-2081. doi:10.1007/s11771-017-3616-4. [12] miner d, triesch j. (2015). self-organization of complex cortex-like wiring in a spiking neural network model. bmc neurosci., 16(1), p265. doi:10.1186/1471-2202-16s1-p265. [13] pomerleau d.a. (1995). a reply to towell’s book review of neural network perception for mobile robot guidance. mach learn., 18(1), 121-122. doi:10.1007/bf00993825. [14] turner j.p., nowotny t. (2015). estimating numerical error in neural network simulations on graphics processing units. bmc neurosci., 16(1), p182. doi:10.1186/1471-2202-16-s1-p182. neuroprotection and timely troubleshooting of electric drive equipment 141 copyright ©2018 assa. adv. in systems science and appl. (2018) [15] yilmaz i, gullu m. (2012). georeferencing of historical maps using back propagation artificial neural network. exp tech., 36(5), 15-19. doi:10.1111/j.17471567.2010.00694.x. [16] yong l, xiu-fen z. (2003). from designing a single neural network to designing neural network ensembles. wuhan univ j nat sci., 8(1), 155-164. doi:10.1007/bf02899473. [17] yuan c-w, leibold c. (2011). capacity measurement of a recurrent inhibitory neural network. bmc neurosci., 12(1), p196. doi:10.1186/1471-2202-12-s1-p196. [18] zhao y, qin b, liu t. (2017). encoding syntactic representations with a neural network for sentiment collocation extraction. sci china inf sci., 60(11), 110101. doi:10.1007/s11432-016-9229-y. [19] zhilin v v, filist sa, rakhim ka, shatalova o v. (2008). a method for creating fuzzy neural-network models using the matlab package for biomedical applications. biomed eng (ny), 42(2), 64-66. doi:10.1007/s10527-008-9019-y. [20] zhong x, wang b-z, wang h. (2001). artificial neural network model for the gap discontinuity in shielded coplanar waveguide. int j infrared millimeter waves, 22(8), 1267-1276. doi:10.1023/a:1015079619009. adv syst sci appl 2017; 4; 34-45 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/257 copyright ©2017 assa. adv. in systems science and appl. (2017) development of the agent-based demography and migration model of eurasia and its supercomputer implementation valery. l. makarov, albert. r. bakhtizin, elena. d. sushko, gennady. b. sushko central economics and mathematics institute cemi ras, moscow, russia e-mail: albert.bakhtizin@gmail.com abstract: in this work we describe the development of a scalable agent-based modelling framework for simulation of eurasia population described in terms of demography, migration and transport flows. the simulated system will consist of agents representing individuals and sets of links to other agents, which represent the social interactions of individual. the individual agents in the model will participate in several independent processes, for which different sets of social links is important such as family and neighbors. as a base for our simulation system we have used a combination of a base native layer implemented using c++ language which uses mpi library, and microsoft .net platform as an environment for model code written in high-level c# programming language. to perform a load balancing of agents between processes the metis/parmetis algorithms were used. these algorithms allow to split the graph of agents and links into parts of similar size with the least possible number of links between them. a number of numerical experiments were carried out for test model to estimate the influence of the parameters of the model on its performance and parallel scalability. for each combination of parameters a number of simulations were performed to average the results. keywords: agent-based modelling, demography, numerical modelling, parallel computing. 2 1. introduction in this work we describe the development of a scalable agent-based modelling framework for simulation of population of eurasia described in terms of demography, migration and transport flows. the goal of the simulation is to describe the influence of large transport infrastructure projects on the development of population of the region. the simulated system consists of agents who represent individuals and a set of links to other agents, which represent the social interactions of individual. the individual agents in the model participate in several independent processes, for which different sets of social links is important such as family and neighbors. the agents of the system participate in two processes: 1) the process of reproduction of the population, and 2) the process of migration. in the first process they use messages exchange to search for the partner to form a family. in the second process the message exchange mechanism is used to obtain information about available jobs in different regions to determine the direction of migration. there are several tools for high performance computing for abm. microsoft axum is a domain-specific concurrent programming language, based on the actor model that was under active development by microsoft between 2009 and 2011. it is an object-oriented language based on the .net common language runtime using a c-like syntax which, being a domain-specific language, is intended for development of portions of a software application that is well-suited to concurrency. but it contains enough generalpurpose constructs that one need not switch to a general-purpose programming language (like c#) for the sequential parts of the concurrent components. http://ijassa.ipu.ru/ojs/ijassa/article/view/257 35 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) the main idiom of programming in axum is an agent (or an actor), which is an isolated entity that is being executed in parallel with other agents. in axum parlance, this is referred to as the agents executing in separate isolation domains; objects instantiated within a domain cannot be directly accessed from another. agents are loosely coupled (i.e., the number of dependencies between agents is minimal) and do not share resources like memory (unlike the shared memory model of c# and similar languages); instead a message passing model is used. to coordinate agents or having an agent request the resources of another, an explicit message must be sent to the agent. axum provides channels to facilitate this. the axum project reached the state of a prototype with working microsoft visual studio integration. microsoft had made a ctp of axum available to the public, but it was removed later. although microsoft decided not to turn axum into a project, some of the ideas behind axum are used in tpl dataflow in .net 4.5 (more information at [5]). repast for high performance computing (repast hpc) 2.2.0, released on 30 september 2016, is a next generation agent-based modeling and simulation (abms) toolkit for high performance distributed computing platforms. repast hpc is based on the principles and concepts development in the repast simphony toolkit. repast hpc is written in c++ using mpi for parallel operations. it also makes extensive use of the boost [2] library. repast hpc is written in cross-platform c++. it can be used on workstations, clusters, and supercomputers running apple mac os x, linux, or unix. portable models can be written in either standard or logo-style c++. repast hpc is intended for users with: • basic c++ expertise. • access to high performance computers. • a simulation amenable to a parallel computation. simulations that consist of many local interactions are typically good candidates. models can be written in c++ or with a “logo-style” c++ [1]. cybergis toolkit is a suite of loosely coupled open-source geospatial software components that provide computationally scalable spatial analysis and modeling capabilities enabled by advanced cyberinfrastructure. cybergis toolkit represents a deep approach to cybergis software integration research and development and is one of the three key pillars of the cybergis software environment, along with cybergis gateway and gisolve middleware [3]. the integration approach to building cybergis toolkit is focused on developing and leveraging innovative computational strategies needed to solve computingand dataintensive geospatial problems by exploiting high-end cyberinfrastructure resources such as supercomputing resources provided by the extreme science and engineering discovery environment and high-throughput computing resources on the open science grid. a rigorous process of software engineering and computational intensity analysis is applied to integrate an identified software component into the toolkit, including software building, testing, packaging, scalability and performance analysis, and deployment. this process includes three major steps: 1. local build and test by software researchers and developers using continuous integration software or specified services; 2. continuous integration testing, portability testing, small-scale scalability testing on the national middleware initiative build and test facility; and 3. xsede-based evaluation and testing of software performance, scalability, and portability. by leveraging the high-performance computing expertise in the integration team of the nsf cybergis project, large-scale problem-solving tests are conducted on various supercomputing environments on xsede to identify potential computational bottlenecks and achieve maximum problem-solving capabilities of each software installation. the implementation of the scalable modelling framework 36 copyright ©2017 assa. adv. in systems science and appl. (2017) pandora is a novel open-source framework created by the social simulation research group of the barcelona supercomputing centre. this tool is designed to implement agentbased models and to execute them in high-performance computing environments. it has been explicitly programmed to allow the execution of large-scale agent-based simulations, and it is capable of dealing with thousands of agents developing complex actions. pandora has full geographical information system support, to cope with simulations in which spatial coordinates are relevant, both in terms of agent interactions and environment. the results of each simulation are stored in hierarchical data format (hdf), a popular format that can be loaded by most gis. this feature is particularly useful, as we will also use gis to analyze simulation results. pandora is complemented by cassandra, a program developed to analyze the results generated by a simulation created with the library. cassandra allows the user to visualize the complete execution of simulations using a combination of 2d and 3d graphics, as well as statistical figures (more information at [6]). swages [10], a distributed agent-based life simulation and experimentation environment that uses automatic dynamic parallelization and distribution of simulations in heterogeneous computing environments to minimize simulation times. swages allows for multi-language agent definitions, uses a general plug-in architecture for external physical and graphical engines to augment the integrated simworld simulation environment, and includes extensive data collection and analysis mechanisms, including filters and scripts for external statistics and visualization tools. moreover, it provides a very flexible experiment scheduler with a simple, web-based interface and automatic fault detection and error recovery mechanisms for running large-scale simulation experiments [10]. a hierarchical parallel simulation framework for spatially-explicit agent-based models (hpabm [11]) is developed to enable computationally intensive agent-based models for the investigation of large-scale geospatial problems. hpabm allows for the utilization of highperformance and parallel computing resources to address computational challenges in agentbased models. within hpabm, an agent-based model is decomposed into a set of sub-models that function as computational units for parallel computing. each sub-model is comprised of a subset of agents and their spatially-explicit environments. sub-models are aggregated into a group of super-models that represent computing tasks. hpabm based on the design of superand sub-models leads to the loose coupling of agent-based models and underlying parallel computing architectures. the utility of hpabm in enabling the development of parallel agent-based models was examined in a case study. results of computational experiments indicate that hpabm is scalable for developing large-scale agent-based models and, thus, demonstrates efficient support for enhancing the capability of agent-based modeling for large-scale geospatial simulation [11]. the growing interest in abm among the leading players in the it industry (microsoft, wolfram, esri, etc.) definitely shows the relevance of this instrument and its big future, while exponential growth of overall data volumes related to human functioning and the need for analytical systems to obtain new-generation data needed to forecast social processes, call for the use of supercomputer technologies. in march 2011, an abm was launched at the lomonosov supercomputer to simulate the development of russia’s socio-economic system for the next 50 years [8]. the implemented abm was based on the interaction of 100 mln agents who conditionally represented russia’s socio-economic milieu. the behavior of each agent was specified by a set of algorithms that described the agent’s actions and interaction with other agents in the real world. the adevs library for multiagent simulation, which the authors had already tested during the multisequencing of russia’s demographic model in 2011, showed itself quite well. in addition, the latest adevs versions support java to a certain extent, which is also a plus. however, the adevs developers have not yet implemented multisequencing on 37 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) supercomputers except for the openmp technology for multiprocessors; therefore, our previous work required many updates for mpi support. when multisequencing the previous, quite simple, model, it was fully rewritten in c++, which was superfluous: the preand postprocessing of data, as well as the creation of the initial state of the multiagent environment, are not time-critical operations. usually, it is quite enough for supercomputers to multisequence only the algorithm’s computing core, i.e., the phase of converting the population’s state in this case. the stages and methods of the efficient reflection of the computing core of a multiagent system on the architecture of the state-of-the-art supercomputer using the supercomputer technology for agent-oriented simulation (stars) developed by the authors are analyzed in later article [9]. analysis of the latest software technologies has shown that embeddable tools are being actively developed lately to execute java programs that use the so-called ahead-of-time (aot) compilation. in this case, the result of the aot compiler is a typical self-contained executable module that contains a machine code for the target platform. it is interesting to note that this approach is used in the new versions of the android operating system, which, in our opinion, is not accidental: the efficiency of code execution is the main factor both for embeddable systems and for supercomputers. experiments with a similar product – the avian aot compiler – allowed us to conclude that, first, it helps obtain a selfcontained executable module as an mpi application for supercomputers; in addition, a random additional code, including initialization and binding to the mpi communication library, is easily implemented in c++; second, the operating speed of the obtained program module is close to the speed of the adevs operation. this made it possible to shift a large part of the work to the aot compiler and to implement only the most necessary in c++, fixing the function of supporting accelerated stages with complex interagent communication to adevs. 3 2. the model description in this work we describe the development of an agent-based modelling framework for simulation of large scale societies which consists of large number of agents. the described technology is to be applied for the implementation of the large-scale agent based model of countries of eurasia describing economy, migration and the results of implementation of large infrastructural projects. the main types of agents in these simulations are individuals, enterprises, regions and governments. the main processes described by the model are demographical evolution of the society, education, jobs and career of individuals and the migration of the workforce according to changing economic conditions. individual agents in the system are described in terms of their age, sex, education level, income and the set of social links to other agents. this set of parameters defines the formation of families, awareness about working conditions in different areas and the possibility of the labor migration. the regions in the model are characterized by the transport connectivity graph, the level of economic development and the labor market conditions. due to that structure of the model agents form the following graph: each agent is linked with a dozen of other individuals in the same or other regions. an individual agent is also linked to his place of work and a region. agents describing enterprises are linked with each other by trade contracts. the described structure of the model leads to formation of a largescale graph of agents of different types which can be partitioned into connected blocks linked to regions. the model uses a change in the transport connectivity graph as an input which leads to the change in economic activity in neighbor regions and the migration of the population. starting from the implementation of the infrastructural project the living conditions are the implementation of the scalable modelling framework 38 copyright ©2017 assa. adv. in systems science and appl. (2017) changing: new jobs are created in neighbor areas which leads to a change in incomes and migration of workforce. 2.1 the technology on the computational level the model should be scalable for system up to 109 agents. in order to perform efficient simulations of such systems the model should support running on modern supercomputers. to simplify the development of the model we use the high-level microsoft .net platform which has become available for running on supercomputers. the most usual architecture of modern supercomputers is a cluster of multicore computing nodes connected by high-performance low-latency network. to run efficiently on such cluster the program should be split into multiple separate processes exchanging messages through the network. the most common way of writing such programs is to use c++/fortran language and mpi library which are available on all supercomputers. as a base for our simulation system we have used a combination of a base native layer implemented using c++ language which uses mpi library, and microsoft .net platform as an environment for model code written in high-level c# programming language. as most of modern supercomputers run on linux operating system, we have decided to use microsoft .net core and mono [4] implementations of .net platform available for this os. the choice of these technologies was determined by the following criteria: 1) the system has to be scalable across multiple computational cluster nodes (i.e. use resource of multiple nodes for speedup) therefore the multithreading calculation model was not suitable as it is limited to single cluster node. 2) the model should be easy to develop and maintain and therefore the high-level programming language c# was used. 3) the system should be efficient and therefore the native mpi library was used instead of tcp/ip sockets or .net libraries like windows communication foundation as these technologies are not optimal for supercomputers and hpc applications. mpi libraries installed on each cluster computer are usually tuned for particular proprietary network system which is used on the cluster such as infiniband. on the level of individual agents the simulation of agent’s internal state evolution, the formation of constant and temporary links between agents, message exchange and formation and destruction of agents in the system are supported. to carry out these simulations efficiently the system has to implement a dynamic load balancing mechanism for agents taking into account their links to neighbors. the proper description and taking into account of links of agents in the process of decomposition is crucial for reduction of network exchange traffic which is necessary for scalability of the model up to hundreds and thousands of cpu cores. to perform a load balancing of agents between processes the metis/parmetis [7] algorithms were used, which are commonly used for decomposition of big graphs (up to 109), computational grids and matrices. these algorithms allow to split the graph of agents and links into parts of similar size with least possible number of links between them. the algorithm can be applied recursively in order to calculate hierarchical splitting of the system in efficient way. the use of the algorithm allows both initial decomposition of the system and refinement of the decomposition in the process of calculation which is necessary to maintain load balancing as new agents are added to the system and some of old agents are being removed. the dynamic decomposition and redistribution of agents should allow us to use efficiently up to 1000 cpu cores. 2.2 implementation of the model the implementation of the agent-based simulation platform requires the proper definition of classes for agents, messages, model, time-steps and utility classes for file input and calculation of characteristics. 39 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) an efficient mechanism of message exchange between agents was implemented by means of message queue and native mpi collective operations. the message queue accumulates a buffer of messages to different processes and then uses mpi alltoall exchange operation to deliver contents of messages. this operation delivers the buffers of arbitrary size from each mpi process to all other processes in most efficient way by splitting the buffer into chunks of optimal size for network transfer and hiding the latency of network operations by performing simultaneous several send and receive operations. to use native operations with managed c# objects operations of binary serialization and deserialization of objects were implemented and c# wrappers for native functions were written. fig. 1. the agents of the model are implemented using c# programming language and run in .net virtual machine. the interaction between vm instances is implemented through the native library and mpi library. the implementation of native wrappers for c++ library was carried out using interopservices library and “dllimport” annotations in managed c# code. in fig. 1 the scheme of interaction of model agents through native library and mpi operation is shown. 2.3 the test model description in order to test the message delivery mechanism the test model was implemented and a set of simulations were carried out to estimate the scalability of the designed model. the model is characterized by the following parameters: the total number of agents n. the ratio of agents participating in message exchange s. message exchange intensity i. total simulation time t. the number of mpi processes u. as it is usual for all mpi programs the program starts as a set of separate processes running the same program. at the initial stage of the model the initial set of agents is created with the number of agents n. each agent has the only numerical parameter the number of messages it should send at each simulation time (a). the set of agents is distributed over all mpi processes, i.e. each mpi process creates only n/m agents according to its process identifier. the process with index 0 creates agents 0...n/m, the process with index 1 creates agents n/m+1...2n/m and so on. on each simulation step for each agent of the system the random number is generated which determines if the agent will participate in message sending process (the probability of the event is s/n). if the agent is sending messages on this simulation step, that the random number of messages a is generated with uniform distribution of probability between 1 and i. after that the implementation of the scalable modelling framework 40 copyright ©2017 assa. adv. in systems science and appl. (2017) the agent sends a messages to agents with random numbers and then receives replies. each simulation step consists of 5 stages: 1. the loop over all agents on the current process, execution of the performstep method of each agent which results in generation of random messages according to corresponding sending probabilities. all messages are put into outgoing message queue. 2. after the generation of all initial messages the method sendreceivemessages is called which initiates the exchange of the parts of message queue between all mpi processes using collective alltoall operation. each process is sending data buffers to all other processes and receives corresponding buffers from all other processes. at this stage the messages in the queue are serialized into binary arrays, these arrays are transmitted into the native library which uses mpi library to perform the exchange, after that new buffers are transmitted to c# part of the program and messages are deserialized. 3. after the exchange of buffers and deserialization of all messages the delivery of messages to corresponding agents on each mpi process is performed which results to the generation of new set of reply messages. 4. the exchange of the parts of message queue between all mpi processes using collective alltoall operation. each process is sending data buffers to all other processes and receives corresponding buffers from all other processes. 5. the delivery of messages to corresponding agents on each mpi process. the use of the delayed delivery of messages through the message queue allows us to optimize the message exchange which is now bound not to the latency of the network but to its bandwidth. after processing all agents in the population, the output characteristics are calculated and put into output file. the following output parameters of the model were written: 1. step number; 2. the average number of message recipients for agents participating in message exchange; 3. the total calculation wall time. 2.4 messages exchange procedure the message queue implemented in the program is a managed c# object which receives objects of abstract class message from agents, each mpi process contains one instance of the message queue object. the simulation model consists of agents of different types and messages can also have different types derived from the message class. for each message type the operations of reading and writing to binary stream are implemented which use binaryreader and binarywriter classes of c# standard library. these operations encode and decode the type of message object, the numbers of the sender and receiver objects and all additional message data fields into binary format. each message queue contains a set of binarywriter objects which are used as buffers for outgoing messages. for each outgoing message put into queue the number of destination process is determined, that the message is written to the corresponding binary stream. the message queue implements a sendreceivemessages method which performs the delivery of all messages to the corresponding agents. in order to deliver messages message queue objects of all mpi processes exchange the corresponding message buffers, for that all binarywriter objects are written to memorystream object which generates an outgoing array of bytes. in order to deliver the outgoing array of bytes first the mpi_alltoall function is used to exchange the sizes of all message buffers between all processes. after that the mpi_alltoallv function is used to deliver parts of the outgoing byte array to the corresponding processes. 41 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) after the exchange of the message buffers each message queue has separate buffers with messages from all other processes, i.e. process 0 has k1 bytes from process 1, k2 bytes from process 2… for all of these buffers the binaryreader objects are created and the deserialization process is started. the program reads messages from all incoming buffers and generates an array of message objects. for each of these messages the receiver agent index is determined, and for the corresponding agent the notify method is called with the corresponding message object passed as argument. the process of the delivery of the message is illustrated in fig. 2. fig. 2. the message sending procedure. the message from agent 0 is transmitted first to message queue of process 0, then the message is serialized to outgoing buffer for sending to process 1, then the buffers are exchanged between processes and the outgoing buffer number 1 becomes the incoming buffer number 0 on mpi process number 1, then the message is deserialized and delivered to agent 1 by the message queue. 3. numerical experiments a number of numerical experiments was carried out to estimate the influence of the parameters of the model on its performance and parallel scalability. for each combination of parameters a number of simulations was performed to average the results. 3.1 test cluster configuration to test the performance and scalability of the model we have used the cluster consisting of 4 dual-processor nodes using amd opteron 6172 (12 cores) processors. the total number of cores in each node was 24 and the same number of mpi processes on each node was used. the high-performance network (qdr infiniband) was used to connect nodes of the cluster. 3.2 the results of the numerical experiments to study the parallel efficiency of the model the following test configuration was used: n = 10000000 (the number of agents) s = 200000 (the number of agents sending messages on the simulation step) i = 10 (the maximal number of messages for one agent on each step) t = 3000 (the number of simulation steps) the simulations were carried out for the number of mpi processes m = 1,24,48,96 and results of these simulations were compared in terms of parallel speedup (the ratio of total computation time for parallel and serial cases) and parallel efficiency (the speedup divided by number of cpu cores used). the implementation of the scalable modelling framework 42 copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 3. the dependence of the parallel speedup of the test model on the number of cpu cores. in fig. 3 the dependence of the parallel speedup on the number of cpu cores is shown. the increase of the number of cpu cores leads to nearly linear speedup of the calculations. the use of 96 cores results to speedup of calculations by factor of 60 which means 65% efficiency of the cluster use. the dependence of the parallel efficiency on the number of cpu cores is shown in fig. 4. fig. 4. the dependence of the parallel efficiency of the simulation on the number of mpi processes. to estimate the influence of the parameters of the model on scalability the simulations were performed with parameter values n = 10000000, s = 200000, t = 3000 using m = 96 mpi processes with different values of message exchange intensity i = 30, 50, 100, 200, 1000. the increase of intensity of message exchange leads to linear increase of the network traffic on each simulation step, and also increase of computations (random number generation) on each simulation step. in fig. 5 the dependence of the speedup on message 43 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) exchange intensity shows that these factors are balanced and increase of intensity doesn’t lead to degradation of parallel efficiency. fig. 5. the study of the influence of the message exchange intensity (i) on the scalability of the simulation. to study the influence of parameter s the parameter i was fixed (i = 10) and the calculations were carried out for the values s = 300, 500, 1000000, 10000000. the increase of the parameter s also leads to the linear increase of both network traffic and calculations of the random numbers, which shouldn’t affect much the scalability. in the case of lower values of the parameter the size of the data is rather small and the speedup is determined more by the latency of the exchange network. this effect leads to the decrease of efficient for low values of s and much better efficiency for larger values of the parameter. 3.3 the influence of the interprocess communications in the second variant of the test model the agents are divided into groups of size g, the exchange of the messages is done only between agents inside the group. for that each agent has additional parameter the number of group ng, which is defined at the beginning of the simulation to establish the uniform distribution of agents between groups. the use of such groups corresponds to the case of ideal decomposition of the graph of agents where most of the interprocess communications are removed. the implementation of the scalable modelling framework 44 copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 6. the dependence of the parallel speedup on the number of agents participating in message exchange (s). the number of mpi processes m = 96. in order to study this effect a set of simulations was performed and the dependence of the parallel speedup on the number of participating agents s was plotted. the results of these simulations are plotted in fig. 6. the plot shows that the influence of the presence of blocks is much higher for high values of s when the number of agents participating in message exchange is high. 4. conclusion in this work a new framework for parallel calculations of agent-based models was presented and tested. the framework links the use of the high-level c# programming language and high-performance platform for messages exchange written using native c++ library and native mpi library of supercomputer. the provided results of the test simulations show good scalability of the program across multiple computational nodes. acknowledgements this work was supported by the russian science foundation (grant # 14-18-01968). references [1] collier n. (2013, august). repast hpc manual. [online] available: http://repast.sourceforge.net [2] boost c++ libraries (2017) [online]. available: http://boost.org [3] cybergis (2017) [online]. available: http://cybergis.cigi.uiuc.edu [4] mono (2017) [online]. available: http://www.mono-project.com [5] axum_(programming_language) (2017), [online]. available: https://en.wikipedia.org/wiki/axum_(programming_language) [6] pandora: an hpc agent-based modelling framework (2017), [online]. available: https://www.bsc.es/research-and-development/software-and-apps/softwarelist/pandora-hpc-agent-based-modelling-framework [7] karypis, g. & kumar, v. (1995). metis-unstructured graph partitioning and sparse matrix ordering system, version 2.0. technical report. http://boost.org/ http://cybergis.cigi.uiuc.edu/ http://www.mono-project.com/ https://en.wikipedia.org/wiki/axum_(programming_language) 45 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©assa. adv. in systems science and appl. (2017) [8] makarov v.l., bakhtizin a.r., vasenin v.a., roganov, v.a. & trifonov i.a. (2011). capacities of a supercomputer system for work with agent-based models, programnaya injeneriya [software engineering], 3, 2-14 [in russian] [9] makarov, v.l., bakhtizin, a.r., sushko, e.d. et al. (2016) supercomputer technologies in social sciences: agent-oriented demographic models, her. russ. acad. sci, 86 (3), 248-257. doi:10.1134/s1019331616030047 [10] scheutz, m., connaughton, r., dingler, a., & schermerhorn, p. (2006). swages an extendable distributed experimentation system for large-scale agent-based alife simulations. in proc. of artificial life x, 412-419. [11] tang w. & wang s. (2009). hpabm: a hierarchical parallel simulation framework for spatially‐explicit agent‐based models, transactions in gis 13(3)315-333. advances in systems science and applications (2017) vol.17 no.1 25 a novel defense solution towards currency wars qionghao chen a) , yirong ying a) , jeffrey yi-lin forrest b) a) college of economics, shanghai university, shanghai, 200444, china; e-mail: qionghao_chen@sina.com, yrying@staff.shu.edu.cn; b) school of business, slippery rock university, slippery rock, pa16057, u.s.a.; e-mail: jeffrey.forrest@sru.edu. abstract. this paper investigates the following problem: how could a nation possibly design a measure to counter large-scale sudden flight of foreign investments in order to avoid the undesirable disastrous consequences? continuing [2], and based on how a currency war could be potentially raged against a nation, all results herein are established by making use of the results of feedback systems. based on theoretical reasoning and systemic analysis, this paper develops a self-defense mechanism that could conceivably protect the nation under siege. when a nation tries to accelerate its economic development, large amounts of foreign investments would generally be welcomed. and at the same time, a lot of such foreign investments would strategically rush into the nation in order to ride along with the forthcoming economic boom. however, recent financial events from around the world indicate that how to avoid the disastrous aftermath when a large-scale flight of foreign capital appears suddenly is still not well understood and well planned out. this fact vividly shows theoretical and practical value of this work. key words: purchasing power, feedback system, monetary policy 1. introduction 1.1. currency war since world war ii, the form of war has changed. the main battlefield of modern warfare has quietly shifted from direct military operations to the economic maneuvers. fundamentally, all modern forms of warfare are currency related. although there is no physical battleground, both the scales and benefits the conflicts eventually generate out of the competition for the financial highlands are no less any those of any war in history. each financial crisis can be seen as the signal of a currency war. our present world on the average experiences about 10 massive financial crises per years. the results are that the relevant countries lose their leadership, if there was any, and have to stay in the consequent economic shadows for years to come, or might be worse, they can no longer recover and return to their previous glory. for example, the british sterling crisis, japan's decade of recession after the plaza accord, southeastern asia's financial crisis, and so on. these related countries did not experience any military conflict, and did not have the time to use any advanced weaponry that had been prepared for use. compared to a military warfare, currency wars had made these countries pay a much greater economic cost. the reality is that most nations today have no need, and are unable to resolve conflicts by employing conventional wars, because by using the idea of currency wars one can achieve his desired goals. klein [11] showed in his empirical analysis that capital account liberalization has different effects on economic growth in different countries. henry [8], based on event study, discussed the effects of stock market liberalization on the emerging market countries’ economic reform and equity market prices. li and zhang [12], based on a dynamic model of aghion, analyzed the impact of capital account liberalization on economic and financial stability. they also used cross-sectional data model involving 57 countries, further studied the economic consequences of opening direct investment for different country samples. zhang [31] analyzed the conduction of a financial crisis as the breakthrough point based on the demonstration effect of the conduction, and tried to investigate financial crises from one side. li [13] established the currency substitution vec model and made a dynamic analysis of extent of china's currency substitution and the relationship between its influence factors. after reviewing previous theoretical analyses with recent cases of speculative attacks in the arena of international finance, we surely see the following predicament. when a nation tries to mailto:qionghao_chen@sina.com mailto:yrying@staff.shu.edu.cn mailto:jeffrey.forrest@sru.edu 26 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars accelerate the economic development, due to its capital account, liberalization and loose monetary policies, large amounts of foreign investments would be welcomed; and at the same time, a lot of such foreign investments would strategically rush into the nation in order to ride along with the forthcoming economic boom. therefore, a serious question at this junction is: how could a nation possibly design a measure to counter such foreign investments to leave suddenly in order to avoid the undesirable disastrous consequences? in this work, based on the studies and discussions of how china’s capital account liberalization affects international capital flows and the economy and finance, we study the evolution mechanism of a currency war by using systems theory. 1.2. literature review 1.2.1. importance of capital account liberalization capital account liberalization is an important proposition of international finance, and has been a focus of attention of economists. as a result of the liberalization in industrialized countries, the issue of capital account liberalization is almost entirely concentrated on developing countries. in the economic globalization today, with the influence of liberalization of external and induced current accounts, implementation of capital account liberalization in developing countries in fact has become an inevitable choice. but since the 1990s, the world has been prone to financial crises, such as the mexican financial crisis in 1994, the asian financial crisis in 1997, and, a few years after that, the outbreak of the brazil's financial crisis, venezuela’s financial crisis, and argentina's financial crisis. the varying degrees of the crises reflect negative impacts on economies and finances as caused by international capital flows consequent to the liberalization of capital accounts [1]. 1.2.2. inevitability of currency wars with the development of civilization, economies, and societies, after world war ii, several of the world's major military powers have found that the use of force, including nuclear weapons, has become increasingly handicapped in resolving disputes. so, the economy has become a battleground, currency wars have already been referred to as a frequent agenda through news outlets. removing potential threats in the global monetary system has been the goal of the globalized economy of the modern world. previous reports on defense strategy towards currency wars can be categorized into active defense and passive defense. gagnon [5] indicated that 22 countries had boosted their economy and created employment opportunities by intervening in foreign-exchange markets. gagnon and bergsten [6] estimated that 91 economies increased their external deficits as some other countries manipulated their currencies so that they had to devalue their currencies as a hedge. the capitalist world is bound to the outbreak of economic crises, which represent the best time to repudiate debts. for example, in 1847, the then wealthy british even got out of its debts in china and the united states by using bankruptcy, creating a precedent example of how debts could be repudiated [18]. this technique can be surely employed today more readily than any time in history, creating more pains and uncertainties in the world. mathematically, an economic system can be written approximately in the following form of an n-dimensional constant coefficient linear system [29]: dx ax bu dt y cx       , where its state space is assumed to be x, and all relevant terms in these equations were defined in [2,3]. assume that xc is the subspace of x that consists of all controllable states and xno the subspace that consists of all non-observable states. then the state space x can be decomposed into the following four subspaces: 1 no cx x x  , advances in systems science and applications (2017) vol.17 no.1 27 2x such that 1 2cx x x  , 3x such that 1 3nox x x  , and 4x such that 1 2 3 4x x x x x    . if we assume the dimensionality of ix is in , then we have 1 2 3 4n n n n n    . by choosing a basis 11{ ' ' }ne e, , in 1x , 1 1 21{ ' ' }n n ne e , , in 2x , 1 2 1 2 31{ ' ' }n n n n ne e   , , in 3x , and 1 2 3 1{ ' ' }n n n ne e   , , in 4x , we can introduce the following coordinate system: 1 1 1 2 1 2 1 2 3 1 2 31 1 1 1{ ' ' , ' ' , ' ' , ' ' }n n n n n n n n n n n n nco e e e e e e e e         , , , , , , , , . for any x x , let t be the transformation from the original coordinate system  to the new system co so that 'x tx , where the i-th column is the coordinates of 'ie in the original system  . so, t is obviously non-singular. now, in co we denote 1 2 3 4' [ , , , ]t t t t tx x x x x , where 11 1[ ' ' ]nx x x , , , 1 1 22 1[ ' ' ]n n nx x x  , , , 1 2 1 2 33 1[ ' ' ]n n n n nx x x    , , , and 1 2 34 1[ ' ' ]n n n nx e e   , , . with this coordinate transformation, the kalman canonical decomposition theorem of the control theory [10] implies the following: the previous constant coefficient linear system approximation of an economic system can be written in the coordinate system co into the following canonical form   11 12 13 14 1 22 24 2 33 34 44 2 4 0 0' ' 0 0 0 0 0 0 0 0 0 ' a a a a b a a bdx x u a adt a y c c x                              and the original system can be decomposed into the following four subsystems: 1 :s 1 11 1 1 1 1 10 0 dx a x b u f dt y x         , where 1 12 2 13 3 14 4f a x a x a x   2 :s 2 22 2 2 2 2 2 2 dx a x b u f dt y c x        , where 2 24 4f a x 3 :s 3 33 3 3 3 30 0 dx a x f dt y x        , where 3 34 4f a x 4 :s 4 44 4 4 4 4 dx a x dt y c x      , where the subsystem s1 does not have any observable output, the subsystem s2 can be completely observed and controlled, the subsystem s3 does not have any output and cannot be affected by any control variable, and the subsystem s4 is not influenced by any control variable. what this theoretical result implies is that in the most general circumstance, each economic system is vulnerable to external influences through first the subspace s1, and then the subspaces s3 and s4. and, domestic monetary and fiscal policies can only effectively affect the subsystem 28 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars s2 and the consequent results on s2 can also be readily observed. the rest of the presentation of this paper turns to the topic of how to defend oneself against currency attacks without much symbolic connection with what is presented in this subsection. 1.2.3. currency superiority of the united states in order to maintain the central position of power in the world economy, the u.s. is ready to fight against the threat of any currency system. as of this writing the u.s. has deployed three strategic fronts centered on the dollar. (1) expansion of global influence in history, the first to the dollar's dominance as the center of the strategic front formed shortly after world war ii. on june 5, 1947, george marshall, the 50th secretary of state of the u.s.a., who had a long-term strategic mind, delivered a historic speech, declaring that the u.s. was ready to help with the european recovery that was to be funded by the u.s. with the most potentially far-reaching marshall plan. this program aimed to make europe american. all allies receiving the u.s. economic aids had to be centered on the dollar; and the recovery from the aftermath of world war ii made europe a mixture consisting largely of american interests, values and informal contact networks. and european countries also undertook political and economic commitments and obligations. that warranted that the united states would have a good hand in the following contest with the economy of the soviet union, and greatly curbed the forces of soviet communist in europe to advance further from the east to the west. however, the dollar, whose value tumbled in europe, was used to purchase additional u.s. goods to enter europe. the consequence was a substantial devaluation of the dollar, leading to a confidence crisis in the dollar. european countries then began to hedge the value of the dollars in their hands with gold. in october 1960, the first dollar crisis erupted. on august 15, 1971, nixon announced a "new economic policy". the u.s. provisionally stopped the convertibility of the dollar into gold, and refused to sell gold to foreign central banks, which not only made the gold-linked dollar in name only but also caused a complete collapse of the bretton woods system. however, as the sole oil pricing currency, the dollar gained a qualitative leap and drew for the first time the boundaries, beyond which the u.s. would employ public actions and deploy military forces. this strategic front of protecting the unique relationship between oil and the dollar, as the time goes on, has become an integral part of american global power. the establishment of cultural, economic and political networks around oil countries has become a global force for american expansion of global influence [26]. (2) expansion of its economic system and political will the second strategic front of consolidating the center position for the dollar was soon employed after the formation of the first one. because liberal economic reforms, as proposed by chicago school of economics, took place successfully in latin america, the u.s. was successful in promoting its capitalist form of economy in terms of geopolitics and in radiating the particular economic form to other socialist countries. the u.s. subversively tried to make chile, bolivia, mexico, and some other countries to abandon the influence of the soviet economic model. with the collapse of soviet union and economies in eastern europe, the united states achieved its important strategic goal to changing the political map of europe that was formed since world war ii, and become a dominant force in the global monetary system and acquired its political advantage over others [25]. it should be noted that since the start of the new century, with the initial regionalization of the rmb, sino and u.s. currencies have been in competition for global influence; and because of the marginalization of the dollar by china, the u.s. has determined to exclude china from the economic power center [26]. in 2011, the most important strategic goal of the united states for the next 10 years consisted of: the united states must expand to the west, dramatically increase diplomatic, economic, strategic and other investments in asia pacific to play a leading role in the entire 21 st century, because the pacific region will serve as the center of nation’s prosperity and global leadership [26]. (3) globally comprehensive containment capability advances in systems science and applications (2017) vol.17 no.1 29 this third strategic front was formed much later in time. the previous two strategic fronts made the united states maintain a multifaceted, cross-cutting, integrated containment capability: even if there were such a country that could surpass the united states in terms of military might, that country would still be hampered by its economic capacity, technological innovation, and social development so that it would no longer be a potentially threaten. more importantly, the advantage of technological innovations of the u.s. led to strategic advantages in many areas, such as the deployment of conventional forces around the world, competition in energy production, climate warming, space network, computer technology, and so on. military force becomes less critical in terms of defending national interests, while its importance only lies in helping the united states to establish a decisive advantage in the world. weak dollar monetary policy was used to hinder economic influence of developing countries. however, unexpectedly, such monetary policies that might shift the u.s. debt crisis also made the economic strengths of these developing countries both constrained and strengthened at the same time. in other words, america’s massive qe monetary policy was introduced at the expense of its national political influence, while its long-term strategic goals are immeasurable. the inconsistency between the monetary dominance of the dollar and the health of the world economy will continue to evolve. in order to reduce the risk of foreign exchanges, various countries will put aside the dollar, and settle their international trades with its own currency. a new currency pattern will be formed [25]. 1.2.4. increasing influence of chinese currency to reduce the pain created by large-scale quantitative easing in the u.s.a., and to maintain its currency sovereignty, china has adopted a strategy to marginalize the dollar while expanding the influence of its rmb [26]. (1) the center front of marginalizing the dollar strategy in the center front, to support and strengthen the rmb’s global influence, china in recent years relied mainly on signing rmb’s bilateral currency swap agreements, expanding the regionalization of trade settlement in rmb, and selectively deepening multilateral cooperation (e.g. free trade agreements). in order to ensure the safety of all its assets that are priced in the dollar, china exercises strategic relations with other developing countries in order to promote its rmb’s global ambition while restrict the effect of the dollar. due to the increasing global influence of rmb, china indeed poses a potential and visible threat to the u.s. through extensive global-scale deployments of its rmb while providing a new currency alternative to the world. china’s activities have resulted in such a directional trend in the international arena that many governments, including some in europe, tend to "marginalize the dollar" rather than oppose to the expanding global influence of rmb. in addition, because china's overall economic level has been below that of the united states, the trend has positively contributed to the implementation and expansion of china’s strategy of marginalizing the dollar [26]. (2) the international front of marginalizing the dollar strategy in the international front, china relies more heavily on political and diplomatic methods. some countries have recognized the fact that heavy dependency on the united states is not appropriate in terms of international trades and currencies [25,26]. this is the political foundation for china to implement the rmb trade settlements and bilateral currency swap agreements in the world. and, it represents a time to establish a broad, united front of currency union in politics, which would include mainly developing currencies and emerging countries in the currency markets, secondly all of the capitalist countries that are seriously, and adversely affected by the dollar hegemony. the union would establish a multilateral currency swap fund for its member countries in order for them to resolve all crises as caused by the instability of the current international currencies and the world financial system due to the instability of the dollar. (3) the outside front of marginalizing the dollar strategy in the outside front, the u.s. is clearly aware of the growing, worldwide movement of marginalizing the dollar. this will certainly encourage the u.s. to exploit any existing weaknesses or economic friction in order to slow down china's economic growth and its 30 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars expanding global currency influence. as a currency policy tool, the exchange rate of the dollar has shown its dual effect. on one hand, depreciation of the dollar leads to a sharp rise in the trading price of the dollar-denominated commodities, such as oil and gold, forces countries from around the world to increase their demand for the dollar, triggers global inflation, and exploit people including the american people themselves. on the other hand, in order to avoid the dollar exchange rate to rise from the increased demand for the dollar, the u.s. government increased the supply of the dollar through massive qes, which not only reduce the fiscal deficit and valuation of foreign debts, but also enables the government to devalue the dollar to achieve the relative appreciation of other currencies so as to realize comparative advantage of the weak currency in exports. to counter such uncertainties, china has maintained the stability of its exchange rate with the dollar and appropriate trade scale [25,26]. 2. framework description 2.1 preliminaries lemma. for any positive definite symmetric matrix w and any constant 0y  , the following inequality holds true 0 ( ) ( ) ( ( ) ( )) ( ( ) ( ))t t t y y x t wx t d x t x t y w x t x t y            (1) proof: because w is a positive definite symmetric matrix, there is 0d  such that tw d d . let 1(g , ,g ) r , 1n ng g   , be a constant vector g of length one. then in the light of cauchy inequality, we have: 0 0 0 0 ( ) ( ) ( ( ) )( ( ) )t t t t y y y y dx t dx t d g gd g dx t d dx t gd                    (2) therefore, we have 0 0 01 ( ) ( ) ( ( ) )( ( ) )t t t y y y x t wx t d g dx t d dx t gd y                        ytxtxdggdytxtx y tttt  ()( 1 1 ( ( ) ( )) ( ( ) ( ))t tx t x t y w x t x t y y      .q.e.d. 2.2 the main result consider the following situation with a polynomial lag 1 1 1 1 ( ) ( ) ( ) ( ) 3(a) ( ) ( ) ( ) ( ) 3( ) ( ) ( ), ,0 n i i i n i i i x ax t a x t h bw t b u t z cx t c x t h dw t d u t b x t t t h                          (3) let 1 1 ( ) ( ) i i m mt t t t t i i i t h t h i i v x px x q xdt h t x hr xdt            (4) then we have the following theorem. theorem 1. if the decision matrix m = m1 + m2 + m3 of system (3) satisfies m1+m2+m3 < 0, then system (3) is stable. proof. a stability condition of system (3) is that the eigenvalues of m are negative numbers. this condition of equivalence means that the decision matrix m must be negative definite, that is, m = m1 + m2 + m3 < 0. in fact, we can do the following symbolic calculation: advances in systems science and applications (2017) vol.17 no.1 31 1 1 ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) m m t t t t i i i i i i v t x t px t x t px t x t q x t x t k q x t k          2 1 1 ( )( ) ( ) ( )( ) ( ) i m m t t t i i i i t k i i x t k r x t x k q x d         (5) 1 1 2 ( ) ( ) 2 ( ) ( ) 2 ( ) ( ) ( ) m m t t t t t t t i i i i i x t a px t x t k a px t w b px t x t q x t         2 1 1 1 ( ) ( ) ( )( ) ( ) ( )( ) ( ) i m m m t t t t i i i i i i i t k i i i x t k q xt k x t k r x t x k q x d              . notice that (a) 2 1 ( )( ) ( ) m t i i i x t k r x t     2 1 2 1 1 1 ( ) ( ( ) ( ) ( ) ) ( ) ( ) m t t t t t t t t m i i m i m m a x a x t h x t x t k x t k w k r a a a w a x t h b w                                     1 1 1 ( ) ( ( ) ( ) ( ) ) ( ) t t t t m m x x t h x t x t k x t k w m x t h w                   where 2 2 2 2 1 1 1 1 1 2 2 2 2 1 1 1 1 1 1 1 1 1 1 2 2 2 2 1 1 1 1 1 ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) m m m m t t t t i i i i i i m i i i i i i m m m m t t t t i i i i i i m i i i i i i m m m m t t t t i i i i i i m i i i i i i a k r a a k r a a k r a a k r w a k r a a k r a a k r a a k r w m b k r a b k r a b k r a b k r w                                                  (6) (b) ( )( ) ( ) ( ( ) ( )) ( )( ( ) ( )) i t t t i i i i i i t k x k r x d x t x t k k r x t x t k         1 1 2 1 ( ) ( ) ( )( ) ( ) ( ( ) ( ) ( ) ) ( ) i m t t t t t t i i m t k i m x t x t h x k r x d x t x t k x t k w m x t h w                          where 32 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars 1 1 1 1 2 0 0 0 0 0 0 0 0 m m m m m m r r r r r r m r r r r                 (7) plugging eqs. (6) and (7) into eq. (5) gives 1 1 ( ) ( ) ( ) ( ( ) ( ) ( ) ) ( ) t t t t m m x t x t h v t x t x t k x t k w m x t h w                   because m = m1 + m2 + m3, 1 2 1 1 1 2 23 0 0 0 0 0 0 0 0 0 0 0 m t i m i t t t m m t a p pa q pa pa pa pb a p q a p qm a p q b p                            (8) now, the stability condition of the general system with polynomial lag is: m1 + m2 + m3<0. (9) q.e.d. 4. case study as in [2,3], we continue to use w to represent the vector  1 2 3 t w w w of categorized monetary policies 1 2, , nw w w , which are accordingly grouped into three categories as follows: 1w = the set of all those monetary policies that deal with the population meeting the minimum need to maintain the basic living standard; 2w = the set of all those monetary policies that deal with the population’s need for acquiring desired living conditions; and 3w = the set of all those monetary policies that deal with the population’s need for enjoying luxurious living conditions. similar to the concept of overall balance of international payments, we introduce an economic index vector  1 2 3 t z z z z such that iz measures the state of the economic sector i, i = 1, 2, 3. when the purchasing power increases, people will purchase more assets and products from foreign countries, the overall balance of international payments will decrease (foreign exchange expenditure increase); when the purchasing power decreases, people will sell more assets and products from the domestic country, and the overall balance of international payments will increase (foreign exchange revenue increase). we established the systemic model (eq. (3)) with polynomial lag. in this model, we use symbol z to represent the state of the national economy, w1, w2, and w3 the positive and negative effects of the monetary policies on the performance of the economy directly, or on the currency demand and supply to have an impact on the economy indirectly. here u(t) is a random vector with a nonzero mean. due to the fact that economic development can be seen as a continuous process, the current change in the money stock is determined by the current monetary policies, advances in systems science and applications (2017) vol.17 no.1 33 money stock, and the previous money stock. and the current performance of the economy is also determined by the current monetary policies, money stock, and the previous money stock. here x is the 3  1 matrix [d1 – s1 d2 – s2 d3 – s3] t of the categorized difference of demand and supply of money, referred to as the three parts of the state of the economic system. specifically, our systemic model divides the economy into three sectors e1, e2, and e3. compared to sector e2, which consists of such goods, services that are used by citizens to acquire desired living conditions, sector e1 consists of living necessities. sector e3 consists of such goods, services that are used by the citizens for their enjoyment of luxurious living. our systemic model of the national economy indicates that our separation of the economy into these three sectors can help properly manage the market reaction to the monetary policies. when the monetary policies have positive effect on the performance of the economy, people in every economic sector will purchase more assets and products with the increase of the purchasing power of their income, (foreign exchange expenditure increases); when monetary policies have negative effect on the performance of the economy, people in every economic sector will sell more assets and products with the decrease of the purchasing power of their income (foreign exchange revenue increases). we obtained the stability criterion m1 + m2 + m3 < 0 (eq. (9)) for the general time-delay system based on the systemic model structure with the first-order lag. let 1 1 1 1 2 2 2 2 3 3 3 3 0 0 0 0 0 0 0 0 0 0 , 0 0 , 0 0 , 0 0 0 0 0 0 0 0 0 0 i i i i a b c d a a b b c c d d a b c d                                            because airi = ki, biri = qi, and w are positive definite matrix, from eqs. (6), (7), and (8) we have 3 3 3 3 3 2 2 2 2 2 1 2 3 1 1 1 1 1 3 3 3 3 3 2 2 2 2 2 1 1 1 1 2 1 3 1 1 1 1 1 1 3 3 2 2 1 2 2 1 2 1 1 ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( t t t t t i i i i i i i i i i i i i i i t t t t t i i i i i i i i i i i i i i i t t i i i i i i a k r a a k r a a k r a a k r a a k r w a k r a a k r a a k r a a k r a a k r w m a k r a a k r a a                          3 3 3 2 2 2 2 2 3 2 1 1 1 3 3 3 3 3 2 2 2 2 2 3 3 1 3 2 3 3 3 1 1 1 1 1 3 3 3 3 2 2 2 2 2 1 2 3 1 1 1 1 ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( t t t i i i i i i i i i t t t t t i i i i i i i i i i i i i i i t t t t i i i i i i i i i i i i i i k r a a k r a a k r w a k r a a k r a a k r a a k r a a k r w b k r a b k r a b k r a b k r a b k r                         3 1 ) t i w                                 1 1 2 3 1 1 2 3 2 2 1 2 3 3 3 3 3 0 0 0 0 0 0 0 0 0 r r r r r r r r m r r r r r r r r                  3 1 2 3 1 1 1 3 2 2 3 3 0 0 0 0 0 0 0 0 0 0 0 0 0 t i i t t t t a p pa q pa pa pa pb a p q m a p q a p q b p                        34 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars so we can obtain an expression for m1 + m2+ m3 explicitly. if the sum is negative, then the systemic model above is stable. to facilitate the detailed calculation, we only select a one-dimensional case for explanation so that the three sectors in [2] become one sector. by substituting the demand and supply of money, x is defined as an exchange rate. we will still use the same symbol w to represent the vector [w1 w2 w3] t of categorized monetary policies. eq. (3a) indicates that the current exchange rate is not only determined by the current monetary policies, but also by the previous monetary policies. in this paper, our model established by empirical studies shows that when the financial crisis occurred (from 2008 to 2010), the government made the exchange rate of the rmb against the u.s. dollar remaining at around 6.8 through the implementation of a series of effective policies and instruments. based on the systemic model structure for the second-order lag, the fitting degree of the model that contains parameters for policy implications increases 16.8% from that of the model without any parameters for policy implications, for details see figure 1. this result means the effectiveness of the policy parameters. fig. 1: delayed effects of monetary policies (fitting after the second order difference in 68 weeks from 2008 to 2010) when we define z as the overall balance of international payments, eq. (3b) indicates that the overall balance of international payments is determined by the current exchange rate and the previous exchange rates. the result also shows the fitting degree of model that contains parameters for policy implications is better than that of the model without any parameter for policy implications. the policy parameters are useful and necessary in the fitting process. we also know that z is determined by the current exchange rate and the previous exchange rates directly, and determined by the current monetary policies and the previous monetary policies indirectly. so the nature of the changing z is determined by quantitative continuous-deferred monetary policies. 5. implications of the established theory in the internationalization process of a currency, government policy is an extremely important factor. to this end, let us consider the policy implications process of some major currencies from around the world. gbp: britain was the first country in the history to build modern financial institutions that grew the fastest and developed with the most perfection. british national order passed a bill to establish a bank of england in 1694 so that britain became the first country in the world to have a central bank. from 1816 to 1819, british government issued a series of regulations and policies about mint and exchange, and implemented a true gold standard, which also made britain the world's first to implement such a standard. from the middle ages to the 19th century when britain became the "sun" empire, the british had dominated the world's finance. after world advances in systems science and applications (2017) vol.17 no.1 35 war i, in the "dollar bloc" and "franc bloc" supplant, pound was no longer used as an international currency. after world war ii, with the establishment of the bretton woods system, pound was degraded to a national currency [30]. usd: after a century of dormancy, with the establishment of the bretton woods system, the dollar became the world currency. however, as of 1971, a deficit, not seen since 1893, in the overall balance of payments of the united states emerged, and the gold reserve of the united states amounted to less than 1/5 of the foreign short-term liabilities. to prevent countries to exchange their holdings of the dollar into gold, president nixon announced the new economic policy on august 15th, 1971, and his administration issued policies and laws to save the crumbling bretton woods system; and the group of ten reached the smithsonian agreement in december, 1971. however, these efforts failed to curb the selling wave of the dollar and the buying of gold and other currencies of the world. the bretton woods system collapsed in 1973. since then, german mark, french franc, britain pound, and other currencies began to enter the international currency system. even since, the dollar began to embark on its long downward spiral [30]. jpy: after the meiji restoration, japan established the bank of japan, the central bank of the country, in october 1882. diverted to the gold standard in 1897, japan became the asian financial pacesetter. after world war i japan began its dominance of the far east and the pacific region. after world war ii japan was taken over by america's "allied command". since then, japan implemented a series of democratization reform measures. as a result of the war on the korean peninsula as well as the dodge plan in june 1950, japanese economy quickly recovered to the pre-war levels. in 1952, japan recovered its sovereignty and joined the international monetary fund and the world bank. however, the nationalization and free convertibility of the yen is not synchronized. since the early 1970s, the yen has become an international currency, but japan did not issue a decree to allow foreigners to issue bond in japan until the mid-1970s. in the meanwhile, japanese investment in foreign securities began to liberalize. during 2013-2014, the proportion of japanese yen traded in new york foreign exchange market was hovering around 23% [14]. mark: after world war i, germany became the country in the world that suffered the most from inflation. since the end of world war ii, germany has always stressed the independence of its central bank; and the german territories occupied by the west followed the united states and established a two-stage system to avoid the government from manipulating the center bank. on july 26, 1957, the federal german parliament enacted the bundesbank act, making bundesbank a unified central bank. after the 1973 collapse of the bretton woods system, the german decree introduced a floating exchange rate system. since then, german mark became the second-largest international currency in the world only after the dollar of the united states [30]. eur: the euro was controversial. based on the euro zone's gdp, import and export volume when comparable to that of the u.s., many european scholars held their widespread optimism that the euro will challenge the dollar with weight tilting to the euro. but scholars of the u.s. held bearish view on the euro (samuelson, 2000; frankel, 2015; soros, 2010). on january 4, 1999, the euro came into the world. since then, the history seems to suggest that although the european union has introduced a series of policies and regulations, the effectiveness of the policies has impacted the international community slightly [14]. china is now the second largest economy and the largest exporter in the world. with its growing economy and deepening financial reform, its rmb has the ambition to become an international currency. however, the internationalization will be a long process due to the following reasons. the chinese government deems free trade agreements (ftas) as a new platform to further open up the country to the outside world and speed up domestic reforms, an effective approach to integrate the country into the global economy and to strengthen economic cooperation with other economies, as well as particularly an important supplement to the multilateral trading system. currently, china has 20 ftas under construction or implementation. the rmb aimed at 36 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars eventually becoming a regional financial settlement instrument and commodity with the ongoing development of the multilateral economic cooperation. it is seen as a natural and regional monetary integration which will eliminate the cost of currency trading and the risk of exchange rate fluctuations, and gradually expand the trading network, promote the economic cycle. secondly, the internationalization of rmb will provide favorable conditions for the development and implementation of china’s economic policies. this factor has a strong appeal to the chinese government. the internationalization of rmb will enhance the autonomy of monetary policy in china and help china to get rid of problems that exist in developed countries, such as the "mead dilemma” that to maintain external balance developing countries will have to sacrifice internal balance. this end will help china reduce the risk of economic instability and safeguard the national economic security, which is particularly important in the course of economic globalization. finally, the internationalization of rmb will help china to guard against financial crises [28]. within the imperfect dollar-dominated international monetary system, the internationalization of rmb will improve china’s capacity of finance. when financial assets are priced (at least partially) in rmb in the world markets, china would be able to decide prices and control risk to some extent. and in the process of internationalization of rmb, china must face two types of risks: (1) the liquidity risk. because of the integration of the global finance, when the assets china holds are priced in a foreign currency, the government will have no means of control to provide liquidity as lender of the last resort. this risk is difficult to manage once the market becomes volatile. the assets and liabilities held by either the government or the citizens in the current account will lead a basic process of money creation. we believe that a contraction of the money supply indicates a greater liquidity problem. even for those countries that have enough foreign exchange reserves, like china, trying to provide adequate liquidity within a brief moment of time will cause a huge impact on assets markets and create a vicious cycle. (2) the risk of price fluctuations under currency mismatch. under the floating exchange rate regime, one has to face the risk of price fluctuations. however, if assets are denominated in the local currency, the risk of price fluctuations will to certain extent be reduced. the fact that the assets held by china are denominated in foreign currencies means that china has surrendered part of its power to regulate its financial market to some foreign governments. so, whether it is for the purpose for china to maintain the stability of its economy and the security of its financial assets or for the purpose for it to balance the stability of its international trades, it is important for china to accelerate the internationalization of its currency, the rmb. the recent rise of emerging-market economies provides a special opportunity for the internationalization of rmb, because diversified international currency represents the interests of the majority of countries. because of these particular circumstances, china enjoys many favorable conditions to promote the internationalization of its rmb. first, china's comprehensive national strength has increased significantly. its fast growing economy in the past decades has laid a firm foundation for people from around the world to use rmb. due to the introduction of prudent monetary policies, china has enjoyed favorable social reputation. for example, china has maintained the value of its rmb by curbing its domestic inflation. rmb has played a notable role in combating the recent global financial crisis. the relatively stable exchange rate has also laid a solid, reliable foundation for promoting the internationalization of rmb. secondly, it is the relatively liberal regime environment in china and improvement of the convertibility of rmb. in china’s neighboring countries and regions, rmb has been employed in similar fashions as the dollar and other hard currencies that are relatively easy to acquire. in order to promote rmb in cross-border trade settlements, some border provinces have developed innovative institutions. asean free trade area for overseas circulation of rmb has provided a broad space for promoting rmb. under the premise of stable exchange rate with rmb, some countries and regions, which trade and invest frequently or with large amounts with china, have advances in systems science and applications (2017) vol.17 no.1 37 been willing to accept rmb as the denominated settlement currency. for example, in trades with vietnam, thailand, myanmar, and other countries, rmb has in fact become one of the monetary instruments for settlements [26]. thirdly, hong kong, a recognized center of international finance, represents a unique advantage for china. in the process of internationalizing rmb, hong kong can play and has played an important role. the city has a complete and mature legal system, perfect market infrastructure, and the hkma (hong kong monetary authority) operates independently from the government. so, hong kong can help with the implementation of rmb’s regionalization, provide useful experience for the chinese central government. before the chinese financial market could become mature enough to handle the excessive volatility of the international financial market, the offshore rmb market in hong kong will play the role of "firewall" for the mainland. and, the financial supervisory system in hong kong can help the mainland government and currency authority to grasp market trends and to detect potential threats. although many favorable conditions exist, the regionalization of rmb is still plagued with various internal and external difficulties. first, the degree of regional economic integration affects the scale of rmb’s regionalization. in areas where trades and personnel exchanges happen frequently, especially in border areas, rmb has seen greater liquidity, and even enjoyed more popularity than the relevant national currencies. however, if a neighboring country did not reach any institutional arrangement with china, then the existing economic integration measures along the border between the two countries would not in general be extended to the whole region, where the regionalization of rmb could only be confined to the border areas. in such areas, financial services are generally in short supply and there would be no unified basis for the exchange rate of the rmb. in some areas of southeastern asia that have liberalized trades and traded heavily with china, including cafta (china and asean free trade area), many countries suffer from their underdeveloped banking systems so that these countries have few financial institutions that handle the business of rmb exchange. in addition, many of these countries do not have their official exchange rates for direct conversion of their local currencies into rmb other than the black market. also, the current security issues and conflicts in northeastern and southeastern asia might very well place a damp on the internationalization of rmb and affect the confidence for rmb to be widely used in trade settlements and investments in the region. because of the constraints of these listed factors, the internationalization of rmb will be a gradual and long drawn-out process. that surely posts challenges for china’s authorities to formulate and to implement wise currency policies. with the increase in capital liquidity and overseas business of chinese commercial banks, the authorities better have an ability to predict, to manage, and to control risks. 6. conclusion each currency war represents a battle without involving gunpowder. during the fight, countries try to compete for acquiring biggest economic benefits mainly by adjusting the supply and value of their domestic money. today, more than 100 countries have been actively or passively involved in currency wars. although some countries can benefit from currency competitions in the short term, it is doubtful whether this benefit is sustainable. and one thing is for sure: most countries, especially those involuntarily involved in the wars, will suffer great economic and social unrest. by taking the impact of capital account liberalization of china on international capital flow and economic and financial system as an entry point, this paper establishes a dynamic systemic model with lag variables, and the stability condition of the dynamical system. then, a onedimensional case is developed to explain the significance of this work. through the model we can know the following. first, stability condition of the dynamical system shows monetary policy and its subsequent effects can play an important role in the economic stability of the country under the free flow of capital, and this condition has certain warning effect. second, how 38 q. chen, y. ying, j.y-i. forrest: a novel defense solution towards currency wars monetary policy regulates the economy in the system. third, the impact of monetary policy and its subsequent effects will play a decisive role in regulating economic equilibrium. based on what has been accomplished in this paper, we have following important suggestions on defense solutions towards currency wars. first, it is comparatively limited in theory to study simply the dynamic systems model between two countries. such study should be expanded to a much bigger dynamic system involving many mutually reciprocal feedback countries so that more convincing results with real policy effects can be established. second, the following is truly a quite complex problem: how can one improve the accuracy of assessing and quantifying the impact of different monetary policies on the economy? this problem and related issues need to be further investigated. acknowledgements. this research was supported by national natural science foundation of china (71301064; 71171128), and research fund of program foundation of education ministry of china (10yja790233). references [1] forrest, j. a systems perspective on financial systems, crc press, balkema, the netherlands, 2014. [2] forrest, j. hopkins, z. and liu, s. f. “currency wars and a possible self-defense (i): how currency wars take place”, advances in systems science and application, vol. 13, no. 3, pp. 198-217, 2013. [3] forrest, j. hopkins, z. and liu, s. f. “currency wars and a possible self-defense (ii): a plan of self-protection”, advances in systems science and application, vol. 13, no. 4, pp. 298-315, 2013. [4] frankel, j., “the euro crisis: where to from here?”, journal of policy modeling, vol. 37, no. 3, pp. 428-444, 2015. [5] gagnon, j. e., “currency wars”, the milken institute review, vol. 15, no. 1, pp. 47-55, 2013. [6] gagnon, j. e., and bergsten, c. f., “currency manipulation, the us economy, and the global economic order", policy briefs, pp.1-25, 2012. [7] he, h. g., “rmb internationalization: mode selection and path arrangements”, finance & economics, vol. 2, pp. 10 – 15, 2007. [8] henry, p. b., “do stock market liberalization cause investment booms?”, journal of financial economics, vol. 58, pp. 301 – 334, 2000. [9] jiang, b. k. and zhang, q. l., “currency internalization: academic review of its terms and impact”, new finance, vol. 8, pp. 6 – 9, 2005. [10] karnopp, d. c. margolis, d. l. and rosenberg, r. c., system dynamics: modeling, simulation, and control of mechatronic systems (5th edition), wiley: new york, 2012. [11] klein, m. w. “capital account openness and the varieties of growth experience”, nb er working paper, pp. 9500, 2003. [12] li, w. and zhang, c., “the impact on fluctuations in real exchange rate and domestic output fluctuations by opening fdi”, management world, vol. 20 no. 6, pp. 11-20, 2008. [13] li, q., “vec model of currency substitution in china: 1994 – 2005”, modem economic science, vol. 29, no. 1, pp. 10-14, 2007. [14] lu, q. j. and zhu, l. n., “an analysis on the effect of monetary policy instruments on money base and money multiplier: the data from 2003 to 2011 in china”, journal of shanghai university of finance and economics, vol. 13, no.1, pp. 50-56, 2011. advances in systems science and applications (2017) vol.17 no.1 39 [15] li, x. and ding, y. b., research on rmb regionalization. tsinghua university press, beijing, 2010. [16] li, x. and kamikawa, t., rmb and yen cooperate with asian currency, tsinghua university press, beijing, 2010. [17] liu, l.z. and xu, q.y., exploration of rmb internationalization, people’s publishing house, beijing, 2006. [18] marx, k. and engels, f., complete works of marx and engels, people’s publishing house, beijing, 2009. [19] qiu, y.l., “europe’s economic future and destiny”, international economic review, no. 2, pp. 49-56, 1997. [20] samuelson, p.a., “japan’s future financial structure”, japan and the world economy, vol. 12, no. 2, pp. 185-187, 2000. [21] soros, g., “the crisis and the euro”, the new york review of books, 2010, available at http://www.nybooks.com/articles/2010/08/19/crisis-euro/ (accessed on march 21, 2016). [22] su, j., “on development of off shore rmb market”, china opening journal, vol. 3, pp. 62-64, 2014. [23] sun, j. wei, x.h. and tang, a.p., “strategic path choice of internationalization of rmb based on the developmental process of three main currencies”, asia-pacific economic review, vol. 2, pp. 69-71, 2005. [24] wang, s.c., “on rmb's internationalization”, contemporary internal relations, vol. 8, pp. 29-33, 2008. [25] wang, y.l., “2012 global outlook: chinese and the united state to compete for currency influence”, available at http://www.docin.com/p-985514784.html (accessed on march 21, 2016), 2012. [26] wang, y.l. “on the strategy of global currency influence of both chinese and united state”, academic journal of zhongzhou, no. 4, pp. 44-48, 2012. [27] xu, m.q., “the internationalization and regionalization of rmb: drawing the experiences and lessons from the internationalization of japanese yen”, world economy study, vol. 12, pp. 39-44, 2005. [28] xu, n.n., “china-asean cooperation will have a major breakthrough in 2015”, the official archive located at http://fta.mofcom.gov.cn/article/shidianyj/201412/ 19716_1.html (accessed on dec 21, 2014), 2014. [29] yang, b.h., valencia, j., li, q.x., and forrest, j. yl., “systemic representation of economic entities”, a presentation at the 2016 annual conference of pennsylvania economic association, 2016. [30] yu, l.n. and xie, h.z., “spillover effect of monetary policy: causes, influence and strategies”, journal of graduate school of chinese academy of social sciences, no. 1, pp. 51-57, 2011. [31] zhang, s.m., “an analysis on demonstrative effect of financial crisis and devaluation effect of competitiveness – the revelation of china’s transitional economy with the opening condition”, world economy study, vol. 3, pp. 36-40, 2003. http://fta.mofcom.gov.cn/article/shidianyj/ adv syst sci appl 2018; 1; 92-101 published online at http://ijassa.ipu.ru. copyright ©2018 assa. adv. in systems science and appl. (2018) a two-stage algorithm for generating a set of paretooptimal trajectories of an object elena l. kulida 1 1) institute of control sciences, russian academy of sciences, moscow, russia e-mail: lenak@ipu.ru abstract: a two-stage algorithm is proposed for generating trajectories and parameters of the object's motion based on searching for a set of pareto-optimal paths from the initial vertex to the terminal vertex. the features of graph construction in the first and second stages of the algorithm are described. at the first stage, the set of vertices of the graph uniformly cover the area of the object's motion, and the edges connect only the nearest neighboring vertices. at the second stage, the set of vertices of the graph includes only those vertices through which the paths constructed in the first stage pass. for this set of vertices, a complete graph is constructed: the edges connect each vertex to all the others. an algorithm for constructing a set of pareto-optimal characteristics of graph paths is described. the two-stage approach allows to significantly reduce the calculation time in comparison with the application of the described algorithm for a complete graph in one stage. two variants of the algorithm are considered. the first algorithm requires a longer calculation time, but allows to obtain more trajectories that are diverse. in particular, it is possible to obtain trajectories of different classes, for example, avoiding obstacles from different sides, etc. for the case when time for calculations is not enough, an abridged algorithm is proposed. the results of computational experiments are presented. keywords: approximation of trajectories by graph paths, a set of pareto-optimal characteristics, search for a path in a graph, algorithm for generation of trajectories. 1. introduction algorithms for generating trajectories that satisfy specified conditions are needed to effectively control the movement of an object moving in a conflict environment. it is necessary that flight management systems are able to comply with 4d trajectories, which needs to be done in computationally efficient ways due to the limited computational resources available [3, 5]. in [6], analytical and discrete approaches are considered to optimize the trajectory of an aircraft in a dangerous environment. the discrete approach allows solving the problem in case of the detection risk of an aircraft by several radars. however, the use of discrete optimization methods requires very large computational costs to generate optimal trajectories, especially if optimization by several criteria is required. in this paper, we propose a decomposition of the optimization process into two stages, which allows to reduce the time required for solving the problem. in [4], an algorithm for controlling the trajectory and parameters of the motion of an object in a conflict environment is given. the risk of detecting an object and the possibility of collision with the earth's surface are taken into account. optimization is carried out in accordance with specified quality criterion of optimizing the trajectory and traffic parameters with a restriction on the time of motion. for this purpose, the specified quality criterion and the criterion of the minimum time of completion of a route are considered. a graph is formed, and its paths approximate all possible trajectories of the object's motion in the given region. the graph must include vertices that cover the region of motion tightly enough to approximate the trajectories with sufficient accuracy. it is desirable that each vertex is a two-stage algorithm for generating a set of pareto-optimal trajectories of an object 93 copyright ©2018 assa. adv. in systems science and appl. (2018) connected by edges to all other vertices to move to an arbitrary direction if the motion along such an edge does not violate restrictions associated with the terrain [1]. two-component characteristics are set on the edge of the graph. the characteristic consist the values of two optimization criteria. sets of pareto-optimal characteristics of paths from a given initial vertex are constructed for the vertices of a graph. however the calculation of the set of pareto-optimal characteristics of the vertices of such a graph requires time-consuming unsatisfactory large for the practical use of the algorithm. the paper proposes an algorithm to solve this problem more effectively in two stages. 2. formulation of the problem the area of motion of the object m is set on the map. when the object moves in this area conflicts with the relief may occur. a known height matrix of the relief for the area is assumed. we consider the set z of all possible trajectories of the object motion from the initial point a m with coordinates and height to the endpoint with coordinates and height . where are the coordinates of the object, is the height of the object, is the speed of the object's motion at time t, and t is the trajectory z passing time. the trajectory should not have conflicts with the relief, at any point of the trajectory the height of the object above the relief must be greater than the specified value . the functional of the cost (the risk) of passing the point with the speed is the cost of the trajectory z is the task is to generate trajectories of the object from point a to point b of the paretooptimal by two criteria: the minimum cost and the minimum time of passing of the route. the problem cannot be solved analytically due to the complexity of the cost functional and the need to take into account the terrain. discrete optimization methods are used to solve the problem. a graph is constructed to approximate the trajectories from the set z, and the problem of finding the pareto-optimal trajectories of the graph by two specified optimality criteria is solved. 3. discrete optimization method for solving the problem 3.1. the construction of the graph in the first stage of optimization we proceed to a discrete solution to obtain an approximate solution from a continuous problem. the solution is sought in the form of a piecewise linear trajectory with piecewise constant speed and height of motion (speed and height is constant on segments of the trajectory and can be changed when moving from one segment to another). we consider a finite set of heights (depths) of the motion , l is the number of heights, and a finite set of velocities of the object , m – the number of considered speeds of movement of the object. 94 e.l. kulida copyright ©2018 assa. adv. in systems science and appl. (2018) the trajectories of the object's motion are approximated by the paths of the graph , n is the set of vertices of the graph, and e is the set of edges of the graph. the path is sought in the form of a sequence of edges where is the speed of motion along the edge . the edge either lies in the horizontal surface at height , then , or is intended to go to another horizontal surface, then . the given area of the object's motion is covered by a uniform grid of points at a distance d from each other to construct the set of vertices of a graph. a separate grid layer is created for each of the l different heights of the object. the vertices are projected onto all layers. if there is an obstacle to the relief: , the corresponding vertex is excluded from the set of vertices of the graph n. in addition the points a and b are included in the set n. it is assumed that the constant velocity is ascribed to the edge of the graph. this speed determines the cost and time of movement along the edge. m edges connect each vertex from the set n with adjacent (nearest) vertices in the horizontal layer. one of the m speeds corresponds to one of these m edges. in addition, each vertex is connected by edges to vertices in adjacent horizontal layers; these vertices are assigned vertical velocities (ascending or descending) of the object. the construction of the graph in the first stage of optimization is illustrated in fig. 3.1. in this figure in the horizontal layers the vertices form a square grid, each vertex is connected by edges with four neighboring vertices in the horizontal layer and with vertices under and above it in adjacent horizontal layers. the number of vertices of the obtained graph can be approximately estimated by the number , where s is the area of the region of motion, d is the distance between neighboring vertices in the layer, and l is the number of layers. the number of edges . these are estimates from above, because some vertices and edges are excluded from consideration due to the obstacles of the relief fig. 3.1. graph construction in the first stage of optimization 3.2. the algorithm for constructing the set of pareto-optimal characteristics each edge is assigned a characteristic in the resulting graph. definition 3.1: the characteristic of the edge at the speed of motion is the two-component vector , where is the cost of passing the edge with the speed , is the time of the passing with the speed , . a two-stage algorithm for generating a set of pareto-optimal trajectories of an object 95 copyright ©2018 assa. adv. in systems science and appl. (2018) definition 3.2: the characteristic of the trajectory z is the two-component vector , where is the cost of the trajectory z, is the trajectory z passing time, defined by the formulas: the problem is to construct in the graph the set of paths from vertex a to vertex b, pareto-optimal by two criteria: and . an algorithm is used to construct the sets of pareto-optimal characteristics of paths from the initial vertex to other vertices of the graph to solve this problem. then the paths corresponding to the pareto-optimal characteristics obtained are constructed for the final vertex. since there are two criteria, each vertex is associated with a set of incomparable twocomponent characteristics. consider two arbitrary characteristics and . definition 3.3: the characteristic and are equal: , if . definition 3.4: the characteristic is smaller than the characteristic : , if . definition 3.5: the characteristics and are incomparable if . definition 3.6: we call the set q of path characteristics from the vertex a to the vertex b pareto-optimal if all the characteristics of this set are pairwise incomparable and for the characteristic of an arbitrary path z from the vertex a to the vertex b . the algorithm for constructing the set of pareto-optimal characteristics notation: is an intermediate set of pareto-optimal characteristics of paths from vertex a to vertex n; is the final set of pareto-optimal characteristics of paths from vertex a to vertex n; q is the set of vertices n for which the set . step 1: the initial vertex a is considered. the characteristic (0,0) is stored in the set , the vertex a is stored in the set q. cycle: a cyclic process is performed until the set q is not empty. 96 e.l. kulida copyright ©2018 assa. adv. in systems science and appl. (2018) an iteration of the cycle consists of two stages. the first stage is to select the vertex for processing: among all vertices , we look for a vertex whose set contains the minimal characteristic . the characteristic is excluded from the set and is included in the set . if after this the set becomes empty, then the vertex n is excluded from the set q. the second stage is processing the selected vertex: all vertices connected by edges with vertex n are considered. for each vertex m, the characteristic , where is the minimal characteristic in the set , is the characteristic of the edge that connects the vertices n and m. if in the sets and there is no characteristic smaller than , then is added to the set . all the characteristics , are removed from – this ensures the pareto optimality of the set . if before the insertion of in this set was empty, then the vertex m is stored in q. after the completion of the algorithm, the set is the set pareto-optimal characteristics of paths from vertex a to vertex b. 3.3. the algorithm for generation of pareto-optimal trajectories the goal of the first stage of optimization is to construct a set n of vertices of the graph for the second stage of optimization. the first stage of optimization step 1: the construction of a graph in the form of a uniform grid. calculation of the characteristics of edges of a graph. step 2: construction of sets of pareto-optimal path characteristics from the vertex a to vertices of the graph. step 3: construction of the set of pareto-optimal paths from vertex a to vertex b based on the constructed set . we denote by n * all the vertices through which the paths corresponding to the set pass. examples of the pareto-optimal trajectories generated in the first stage are shown in figures 3.2 – 3.4. dangerous areas for the movement of the object are represented in the figures with filled rectangles. the color intensity characterizes the degree of danger of the zone. optimization according to two criteria allows obtaining different trajectories, including bypassing dangerous zones from different directions or passing through dangerous zones, as this degrades the quality of the trajectory by one criterion, but improves by the other – it allows reducing the trajectory transit time. a two-stage algorithm for generating a set of pareto-optimal trajectories of an object 97 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 3.2. trajectory bypasses dangerous areas fig. 3.3. this trajectory is more dangerous but the travel time is less 98 e.l. kulida copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 3.4. the trajectory is passing through dangerous zones but the travel time is less the second stage of optimization step 4: construction of a graph for the second stage of optimization. n * is the set of vertices through which the pareto-optimal paths constructed in the first stage pass. each of the vertices of the set n * is projected onto all l layers of the grid to construct the set n of vertices of the graph for the second stage of optimization: a complete graph is constructed in each horizontal layer in the second stage of optimization; each of the vertices of the layer is joined by edges with all the other vertices of this layer to construct real trajectories of the object's motion. the vertices are connected by edges between the horizontal layers one under the other. a characteristic is calculated for each edge in the resulting graph. in the first step, the set of vertices of the graph includes vertices evenly covering the area of motion, and the edges of the graph connect only to the nearest neighboring vertices. at the second stage, the set of vertices of the graph contains only those vertices through which the paths built at the first stage pass. each of the vertices is connected by edges to all other vertices of the graph. a set of pareto-optimal paths is constructed in the complete graph. this two-step approach allows to significantly reduce the computation time in comparison with the use of the algorithm for constructing a set of pareto-optimal paths in the initial graph. step 5: construction of sets of pareto-optimal characteristics of paths from vertex a for the vertices of the graph. step 6: a two-stage algorithm for generating a set of pareto-optimal trajectories of an object 99 copyright ©2018 assa. adv. in systems science and appl. (2018) an analysis of the constructed set of pareto-optimal characteristics; the choice of characteristics satisfying the requirements and constraints [2]; the construction of appropriate paths. 4. simulation results a series of experiments was performed to evaluate the efficiency and accuracy of the proposed algorithm using the example of generating object trajectories moving in a conflict environment, where it is detected by means of detection of different nature. below are the results characterizing the use of the proposed optimization algorithm in one of such experiments. as can be seen from the table 4.1, a decrease in the distance d between vertices in the construction of the grid at the first stage of optimization predictably allows to build more diverse trajectories and substantially improve the accuracy of the solution, but the calculation time increases dramatically. the evaluation of the trajectory is the probability of not detecting an object during the trajectory, so the evaluation lies in the range [0,1]. the best trajectory has greater evaluation. table 4.1. results for different length of edge of graph length of edge of graph, km evaluation of the trajectory travel time, hours calculation time hour:min:sec number of trajectories 32 0,4078 28,29 00:01:50 254 16 0,5612 21,05 00:26:23 1140 8 0,6148 18,4 06:20:21 2639 when analyzing the set of generated trajectories, it turns out that the parameters of incomparable trajectory characteristics can differ by negligibly small values. in order to reduce the number of generated trajectories and to shorten the calculation time, the following method was used: during the calculation, the values of the parameters of the segments characteristics were rounded up to four decimal places. the results of the computational experiment are shown in the table 4.2. table 4.2. results for different length of edge of graph with/without rounding length of edge of graph, km evaluation of the trajectory travel time, hours calculation time hour:min:sec number of trajectories without rounding 32 0,4078 28,29 00:01:50 254 with rounding 32 0,4078 28,29 00:01:12 245 without rounding 16 0,5612 21,05 00:26:23 1140 with rounding 16 0,5613 21,05 00:19:44 1064 without rounding 8 0,6148 18,4 06:20:21 2639 with rounding 8 0,6149 18,4 04:04:00 2420 as follows from the table, the characteristics of the generated trajectories did not change to two decimal places, and the calculation time decreased. in order to reduce the calculation time significantly, one can use a shorter algorithm. the algorithm for generating pareto-optimal trajectories can be reduced as follows: step 2: instead of constructing sets of pareto-optimal path characteristics, calculate the length of the optimal path by the main criterion for each level of the graph, for example using the dijkstra algorithm. 100 e.l. kulida copyright ©2018 assa. adv. in systems science and appl. (2018) step 3: for each layer of the graph construct a path from vertex a to vertex b, which is optimal by the main criterion, i.e. with a minimal cost. we denote by n* all the vertices through which constructed paths pass. the other steps of the algorithm remain unchanged. the results of the comparison of the algorithm 1 (source) and the algorithm 2 (reduced) are presented in the table 4.3. table 4.3. the results of the comparison of the algorithm 1 and the algorithm 2 length of edge of graph, km evaluation of the trajectory travel time, hours сalculation time hour:min:sec тumber of trajectories algorithm 1 32 0,4078 28,29 00:01:12 245 algorithm 2 32 0,332 27,79 00:00:14 65 algorithm 1 16 0,5613 21,05 00:19:44 1064 algorithm 2 16 0,5291 20,67 00:01:33 262 algorithm 1 8 0,6149 18,4 04:04:00 2420 algorithm 2 8 0,5968 19,16 00:12:16 553 5.conclusion the article presents two algorithms for constructing pareto-optimal trajectories of object motion. optimization by the two criteria makes it possible to obtain trajectories that are optimal by the main criterion and satisfy a given time limit. the results of numerical experiments are presented. the following result is obtained by comparing the two proposed algorithms. in algorithm 2, the number of generated trajectories is reduced significantly, while the accuracy of the best solution decreases insignificantly, and the calculation time decreases very significantly. as can be seen from the table 3, in the case where the edge of the graph is 8 km, the calculation time has decreased by more than 20 times! references [1] bazhenov, s.g., egorov, n.a., kulida, e.l., & lebedev, v.g. (2016). control of aircraft trajectory and speed to avoid terrain and traffic conflicts during approach maneuvering, automation and remote control 77(10), 1827-1837. [2] bazhenov, s.g., korolyov, v.s., kulida, e.l., & lebedev, v.g. (2014). simulation of on-board model of airliner to evaluate capability of trajectories and flight safety. 29th congress of the international council of the aeronautical sciences, icas 2014. [3] diaz y., lee s., egerstedt m., & young s. (2013) optimal trajectory generation for next generation flight management systems. 32nd digital avionics systems conference. [4] dobrovidov, a.v., kulida, e.l., & rudko, i.m. (2015) path optimization for a moving object in an anisotropic environment using the probabilistic criterion in the passive sonar mode. automation and remote control, 76(7), 1271-1281. [5] hehn, m., & andrea r.d. (2011) quadrocopter trajectory generation and control. 18th ifac world congress, vol. 18, 1485–1491. a two-stage algorithm for generating a set of pareto-optimal trajectories of an object 101 copyright ©2018 assa. adv. in systems science and appl. (2018) [6] zabarankin, m., uryasev, s. and pardalos p. (2002). optimal risk path algorithms. cooperative control and optimization (r. murphey and p. pardalos ed.), kluwer academic publishers, dordrecht, vol. 66, 271–303. microsoft word 13 ge zhenghao, zhang kaikai, jiang meng , yang fulian--an exact reverse design approach for disk cam mechanis advances in systems science and applications (2011), vol.11, no.3-4 301-308 issn 1078-6236 international institute for general systems studies, inc. an exact reverse design approach for disk cam mechanisms ge zhenghao, zhang kaikai, jiang meng and yang fulian shaanxi university of science & technology, xi’an, shaanxi, 710021, china abstract in the view of the difficult problem of receiving the original design of the cam and its follower motion specification, an exact new reverse design method for disk cams is provided. the essential difference between the proposed method and other existing approaches is its ability to make the cam profile smooth while still exactly satisfying boundary condition of follower displacement,velocity and acceleration. it takes computer movement simulation technology as the main instrument. firstly,in order to reverse the follower motion specifications accurately, an new method is proposed in the paper, first of all, smooth disposal is processed to the cam profile which is formed by the equalized measurement data,and then the cam profile is processed into a series of discrete data. secondly,this paper does research on motion specification of disk cam follower through establishing mathematical modeling and analysis treatment,finally it works out the expression of follower motion specifications. in addition,the way of distinguish the real motion specification of cam mechanism is put forward. thirdly, the simulation modeling is established and the follower motion specifications which is the accurate gauge of the reversed disk cam is also achieved. an example shows that the proposed method can be a powerful tool of cam profile smoothing, which verifies the method can realize exact reverse design for disk cam mechanisms rapidly. the approach is not only suitable for own program but also for cad/cae software, and it can be used for the spatial cam reverse design as well. keywords reverse design motion simulation motion specification cam mechanism 1. introduction a cam is a mechanical device for transmitting mechanical work to another component (the follower) according to the transmission,direction and control function. a planar cam mechanism is the common type to transform rotary motion of the cam into translating or oscillating motion of the follower. since the motions of a follower depend on the cam profile and the follower type, the exact profile of the cams must be given to obtain the prescribed output of the follower after reasonable choice of a follower type and parameters for the cam follower systems. therefore, how to achieve process of the introduced products without drawing, and how to test for processed products, which not only play an important rule on import machinery accessories, but also can speed up the digestion and absorption of foreign advanced technology. so as to create and develop our own new products has very important practical significance. 2. relevant research many scholars make a great deal of reverse investigation on disk cam mechanisms. for example, the reverse design equation of follower motion specification is founded basing on the data which are measured by cmm and the meshing relation of cam and roller. using cubic spline function, the interpolation functions of actual cam contour is obtained, and then academic contour line equation is worked out, finally the motion specification of follower and the implementation are reversed. nowadays, there remain lots of difficulties on reverse design of disk cam mechanisms, such as realize measurement and evaluation reversed disk cam fleetly and precisely. as the fig. 1 shows,a new approach was introduced. first achieved the cam profile by the measured data and processed smooth disposal, second calculated the follower motion specification via motion 302 ge: an exact reverse design approach for disk cam mechanisms simulation,finally the new method of detection and evaluation for the cam profile fitting result using the motion specification of cam mechanism was put forward. this method of reserve design is not only suitable for own program conveniently, but also for existing cad/cae software. fig. 1 the flow of cam reverse design 3. curve fitting and discrete of disk cam contour the data which were measured by cmm can not be used directly in reverse design of disk cam motion specification. it should process smooth disposal. however, it is complicate to dispose the profile formed by the measured data smooth directly. thereby, in this paper, it took the method that firstly fitting profile via the measured data and then transforming the smooth curve into discrete points. nowadays, nurbs curve is widely used for various curves fitting .because it is convenient to adjust the curve slightly due to the localized performance of basis function of . its k-th power curve equation is: ( ) ( ) ( ) ( ) ( ) ( ) ( ) [ ]1,0 0 , , , 0 , 0 , 0 , ∈ ⎪ ⎪ ⎪ ⎪ ⎩ ⎪⎪ ⎪ ⎪ ⎨ ⎧ = == ∑ ∑ ∑ ∑ = = = = u wub wub ur ur wub wub u n j jkj iki ki n i ikin i iki n i iiki , v v p (1) iv –control vertex; iw weighted factor; ( )ub ki, k-th power basic function of b-spline. first the cam contour curve was processed into discrete points after fitted, and then the discrete point-group of cam contour curve can be obtained. ( )ii yx ′′′′, (i= 1, 2, 3……n). advances in systems science and applications (2011), vol.11, no.3-4 303 4. reverse design of motion specification as figure 2 shows, o xyz− was basic coordinate system, 1 1 1 1o x y z− was variation coordinate system which was consolidated with cam. 2 2 2 2o x y z− was variation coordinate system firmed with follower. roller and cam are meshed in points . basing on conjugate meshing principle, it can be concluded that the normal of cam which get across must get across roller-center point ifo , . fig.2 cam mechanism and its coordinate system 4.1 calculate the coordinate of roller center point when cam rotates counterclockwise, the angle between coordinate system and coordinate system can be described as: n i i ⋅ = πθ 2 (2) suppose the theoretic coordinate of cam contour curve is ( ), the roller center point can be concluded as follow: ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ ′′′+′′′− ′′′+′′′ = ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ ′′′ ′′′ ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ −= ⎥ ⎥ ⎥ ⎦ ⎤ ⎢ ⎢ ⎢ ⎣ ⎡ 1 cossinsin sincos 1100 0cossin 0sincos 1 , , iii iiii i i ii ii if if yx yx y x y x θθ θθ θθ θθ (3) the slope of cam counter curve normal via meshing point can be calculated by formula (4): 11 11 −+ +− ′′−′′ ′′−′′ =′′ ii ii i yy xx n (4) then the expressions of theoretic cam contour point can be worked out. ( ) ( )⎪ ⎪ ⎩ ⎪ ⎪ ⎨ ⎧ +′′ ⋅′′ ±′′=′′′ +′′ ±′′=′′′ 1 1 2 2 i fi ii i f ii n rn yy n r xx (5) the coordinate of roller center can be deducted putting formula (5) into formula (3). 4.2 calculate the swing angle of rocker 304 ge: an exact reverse design approach for disk cam mechanisms in coordinate system o xyz− the original point of 2 2 2 2o x y z− coordinate system is known, the original angle between o xyz− and 2 2 2 2o x y z− coordinate system is ϕ , hypotheses the angle between rocker and x axis is iϕ′ . then the swing angle of rocker can be described as follow while the cam rotating. ii ϕϕπϕ ′−−= 2 (6) according to the triangle geometric relations: if if i xc y , ,tan − =′ϕ (7) then put formula (7) into formula (6),obtained formula (8): ⎟ ⎟ ⎠ ⎞ ⎜ ⎜ ⎝ ⎛ − −−= − if if i xc y , ,1tan 2 ϕπϕ (8) the swing angle of rocker can be constructed by formula (3), (5) and (8). 4.3 calculate velocity and acceleration of follower using is 、 iv and ia represented discrete actuality displacement, velocity, and accelerate of follower separately, the expressions were constructed as follow: ⎪ ⎪ ⎪ ⎩ ⎪ ⎪ ⎪ ⎨ ⎧ − − = − − = = −+ −+ −+ −+ 11 11 11 11 ii ii i ii ii i ii vva v s θθ θθ ϕϕ ϕ (9) 4.4 calculate dimensionless motion specification of follower the actual motion specification curve of follower can be obtained by processing the follower actual motion specifications into dimensionless. because the follower actuality displacement, velocity, and accelerate are discrete, so they can be dimensionless processed directly. the time , displacement , velocity , and accelerate can be expressed as follow: ⎪ ⎪ ⎪ ⎪ ⎩ ⎪ ⎪ ⎪ ⎪ ⎨ ⎧ = = = = h ta a h tv v h s s t tt hi i hi i i i h 2 (10) in the expressions: iii avst ,,, ——dimensionless discrete time, displacement, velocity, and acceleration of follower; ht ——intervals of cam lifting or returning time; h ——position corresponding to of follower. advances in systems science and applications (2011), vol.11, no.3-4 305 5. example of cam mechanism reverse design a piece of cam which belongs to conjugating cam is shown in fig.3. for this disk cam mechanism, the distance between a and b is 228mm, the length of pendulum rod is bc=145 mm, and radius of roller circle is 65 mm. fig. 3 disk-cam mechanism with swinging follower the discrete data is measured by cmm, the diameter of probe-radius is 1mm. cam contour curve fitting shown as in figure 4 was obtained by equalized data. the curvature quality of contour curve directly fitted via compensation data is poor in fig. (a). however, the curvature of contour curve shown in fig. (b), which was fitted by points and processed smoothing, is more smooth and high-quality. (a) curvature of cam contour curve which is fitted directly(b) curvature of cam contour curve which is treated smoothly fig. 4 curvature of cam contour curve according to fitting cam contour curve, the paper built a motion model and simulate motion,shown in fig. 5. the follower actual motion specification curves were obtained. (fig.6) then compared the motion specification of the follower with the desired curve of design requirements(the expected tract of the actual follower is known ),if the difference is in the scope of permissible error of the track, the reverse cam quality is eligibility, otherwise it need to re-reverse design cam. fig.5 mechanism simulation modeling 306 ge: an exact reverse design approach for disk cam mechanisms (a) displacement curve (b) velocity curve (c) acceleration curve fig. 6 actual motion specification curve advances in systems science and applications (2011), vol.11, no.3-4 307 according to the actual motion specification curve, the cam rotates 90°during lifting and returning time separately. processed actual motion specification curve into dimensionless motion specification, the dimensionless motion specification of lifting and returning curve were constructed in fig.(7) and fig.(8). (a) lifting dimensionless displacement curve (b) lifting dimensionless velocity curve (c) lifting dimensionless acceleration curve fig. 7 lifting curve of dimensionless motion specification (a) returning dimensionless displacement curve (b) returning dimensionless velocity curve 308 ge: an exact reverse design approach for disk cam mechanisms (c) returning dimensionless acceleration curve fig. 8 returning curve of dimensionless motion specification basing on lifting motion specification curve, it can be concluded that the lifting motion specification is polynomial motion of 4-5-6-7. while it can be concluded that the returning motion specification is polynomial motion of 4-6-8-10, basing on returning motion specification curve. 6. conclusion an exact approach to profile generation of disk cams and evaluation of the reverse designed cam is presented for cam-follower systems. the way of identification the real motion specification of cam mechanism was put forward also. the approach is not only suit for own program but also for cad/cae software, and it can be adopt to the reverse design of spatial cam mechanisms as well. references [1] liu changqi, muye yang and cao xijing,. the design of cam mechanism.beijing:china machine press, (2005). [2] der-min tsay and meng-hung huanghui-chun ho. producing follower motions through their digitized cam coutours. journal of mechanical transmission, vol. 2, issue 2(2003) 98~105. [3] yang dongming. the studies of reverse design,error analysis and test for cam mechanism. machine design and research, (1999) 39-62. [4] wang yiming and zhao jibin. the reverse design for nozzle body of paper. machine design, 22(6) (2005) 51-53. [5] wang shuang, yen guo fu, luo zhong xian and wang ying – guo. development of planar cam mechanism cad system. coal mine machinery,27(6) (2006)999-1000. [6] yu xiaoyan and lan zhaohui. a method of determining the curvature radius of a planar cam profile by cubic spline differentiation. mechanical transmission,32(1) (2008)50-51. 188-200 advances in systems science and applications (2011), vol. 11, no. 1-2 issn 1078-6236 international institute for general systems studies, inc. analysis for gene networks of colon cancer based on logical relationships * 1 college of information science and engineering, shandong university of science and technology,qingdao, shandong 266510, china 2 college of applied science, beijing university of technology, beijing 100124, china 3 department of mathematics, southern illinois university carbondale, makanida 62958, usa email: wangshd2008@yahoo.com.cn abstract genome forms gene networks in terms of complicated interaction to realize its functions. the fundamental principle that “the structure of system determines its function” inspires researchers to contrast the differences between the structure of gene networks for experimental and control groups to get useful information of molecular mechanism of the formation of cancer. in this research, the higher-order logical networks are constructed through reverse network modeling based on the expression data of phylogenetic profiles of cancer-related genes in the normal and diseased stages. through contrasting of the network structure between experimental and control groups, such obvious differences in the structural components of networks are found as the number of non-isolated nodes and the logical relationships, the distribution of logic motifs of 2-order. furthermore, the dynamical behaviors of logical networks of the commander genes of experimental and control groups are simulated and analyzed respectively and significant differences in dynamical stability are discovered. when the above method is applied to the expression data of different stages of other cancers, revelatory differences of the variation of reaction modes for the formation and development of cancer can be given, which is useful for the study of the mechanism of the development of cancer. keywords systems biology gene network logical network dynamics 1. introduction the invasion and metastasis of colon cancer is an important reason that would influence the prognosis of the patients and lead to death. in recent years, through studying for the etiology and pathogenesis of colon cancer, it is generally accepted by the biomedical scientists that the formation and development of human‟s colon cancer is a complicated process which involves the alteration of many genes and stages. the process of canceration of the cells contains the stages of initiation, development and diffusion, and every stage involves the activation of oncogenes and the inactivation of the cancer suppressor genes. hence, finding the functional genome related with disease characteristics from many virulence genes, the reaction modes among genes and the dynamic of oncogenes etc. is of great significance to the diagnosis and cure of the cancer and drug design. this is also an important project in the research of *supported by the national natural science foundation of china (grant nos. 60874036, 60503002) and sdust research fund xiuzhu xu 1, shudong wang 1, kaikai li 1, dashun xu 3, dazhi meng1,2 mailto:wangshd2008@yahoo.com.cn app:ds:fundamental%20principle javascript:showjdsw('jd_t','j_') app:ds:contrast http://dict.cnki.net/dict_source.aspx?searchword=profiles app:ds:contrast app:ds:structural javascript:void(0) javascript:void(0) javascript:showjdsw('jd_t','j_') javascript:showjdsw('jd_t','j_') javascript:showjdsw('jd_t','j_') javascript:showjdsw('jd_t','j_') app:ds:process advances in systems science and applications (2011), vol. 11, no. 1-2 189 bioinformatics [1-2] . in recent years, the theoretical methods for studying the pathogenesis of cancers are mostly based on the expression data of phylogenetic profiles [3−11] , and the computational methods have just been developed to detect functional linkages between proteins or genes, such as pearson correlation coefficient, euclidean and hamming distances, mutual information, the hypergeometric distribution and shortest-path analysis [12−14] . however, traditional pairwise relationships cannot adequately illustrate the complexities that arise in cellular networks because of branching and alternative pathways. for instance, shikimate 5-dehydrogenase (cog0169) is present if and only if 3-dehydroquinate dehydratase (cog0757) or 3-dehydroquinate dehydratase (cog0710) is present [15] . in order to further reveal the ubiquitous logical relationships mentioned above in biology, a computational approach-logic analysis of phylogenetic profiles (abbreviated as lapp) was proposed for identifying detailed relationships among proteins on the basis of genomic data, which was applied into 4873 distinct orthologous protein families of 67 fully sequenced organisms, and identified 750, 000 triplets from analysis of the original, unshuffled biological protein profiles [15] . lapp can reveal the causal or restricted relationships in complex network, help biologists to understand the unknown logical relationships among genes or proteins, and further discover the new biological mechanisms through experiments. as an application of triplet logic analysis to 85 diffuse infiltrating gliomas quantified using oligonucleotide arrays, the meaningful sets of genes that matched clinical outcomes were explored [16] . a bayesian modeling framework for combining phylogenetic profile data via a likelihood with rosetta stone data via a prior probability was described [17] . a three-way gene interaction model was proposed that captured the dynamic of co-expression relationships between two genes [18] . a logical network with 16 active genes of shoot in different external stimuli was constructed, and the dynamics of the network was analyzed [19] . in this work, the logical networks of the normal and diseased stages are constructed through lapp method based on the expression data of phylogenetic profiles of cancer-related genes. through contrasting of the network structure between experimental and control groups, such obvious differences in the structural components of networks are found as the number of non-isolated nodes and the logical relationships, the distribution of logic motifs of 2-order. the differences inspire biomedical researchers: the reason that organism is normal in the normal stage might be that the biochemical reaction mode, the extent of gene relationships and most of the genes play normal roles. furthermore, the dynamical behaviors of logical networks of the commander genes of experimental and control groups are simulated and analyzed respectively and significant differences in dynamical stability are discovered that the number of attractors for control group is far less than the one of the experimental groups. this paper is organized as follows: in the first part, the background of our work and the present researching situation of logical network of genes. in the second part, the method of lapp is clarified briefly and the data sources and the processing methods of the data are given in the third part. the logical networks corresponding to experimental and control groups are constructed using lapp method respectively and the structural components of networks are contrasted and analyzed. the dynamical attractors of logical networks of the commander genes in the experimental and control groups are contrasted respectively in the fifth part. the above results are analyzed and discussed, and the problems to be solved further are put forward in the sixth part. 2. the method of lapp in order to understand and predict the functions of biological systems, we must first identify the structure of the system. for example, in order to show the regulatory relationships among the genes, we must identify all the components, the functions of every component, their interrelations and all the parameters related in the system. now, there are two main methods to http://dict.cnki.net/dict_source.aspx?searchword=profiles http://dict.cnki.net/dict_source.aspx?searchword=profiles app:ds:contrast app:ds:structural javascript:void(0) javascript:void(0) javascript:void(0) app:ds:structural app:ds:contrast app:ds:contrast build the model of biological system structure: forward and reverse network modelings. reverse logical network modeling is a method of reverse network modeling and can determine the logical relationships and types among elements of systems. in this work, we identify the logical structure of cancer-related genome using lapp method based on the data of gene expression profile. we entitle the sample data of gene obtained in many different experiments as its expression profile. for the convenience of writing and computing, we also denote the sample data by the genes. for example, the logic of genes a and b indicates the logical relationships of expression data of genes a and b. lapp method is determining the logical relationships of 1-, 2-, and 3-order by calculating the uncertainty coefficient (abbreviated as u ) of logical functions of every order among the genes based on the expression profiles of genes. the 1-order logical relationship between genes a and b is determined as follows:            )( , | 11 1 bh afbhafhbh afbu   , (1) where h refers to the entropy of the individual or joint distributions. )( xyu denotes the uncertainty coefficient of the influence of x on y . the size of u value denotes the statistical possibility of the uncertainty logic of 1-order of x to y . 1f is one of the proper functions of 1-order logic of a to b. the proper functions of 1-order logic are divided into two types (table 1): ab  , namely, the presence of a leads to the presence of b , which is called the synchronization (equivalent with ab  ); and ab  , namely, the presence of a leads to the absence of b , which is called the asynchronization (equivalent with ab  ). table 1 the list of the proper functions and logic types of 1-, 2-order. obviously, the proper functions of 1-order and logic types are the same. order the proper function and its serial number the logic type and its serial number 1-order 1.b a 1.b a 2.b a 2.b a 2-order 1. bac  1. bac  2. )( bac  2. )( bac  3. bac  3. bac  4. )( bac  4. )( bac  5. bac  6. bac  5. bac  bac  7. bac  8. bac  6. bac  bac  9. )( bac  7. )( bac  10. bac  8. bac  similarly, the expressions of 2-, 3-order logics are as follows:             ch bafchbafhch bafcu ,,, ,| 22 2   , (2)             dh cbafchcbafhdh cbafdu ,,,,, ,,| 33 3   , (3) where 32 , ff are one of the proper functions of 2-, 3-order logics respectively. the logics of 190 xu:analysis for gene networks of colon cancer based on logical relationships 190 xu:analysis for gene networks of colon cancer based on logical relationships 190 xu:analysis for gene networks of colon cancer based on logical relationships 190 xu:analysis for gene networks of colon cancer based on logical relationships 190 xu:analysis for gene networks of colon cancer based on logical relationships http://www.iciba.com/expression/ 191 2-order among a, b and c contain 10 proper functions and 8 logic types (table 1). the logics of 3-order among a, b, c and d contain 218 proper functions and 68 logic types. of course, the higher order logics can be calculated, but the complexities are higher. in this research, we only consider 1-, 2-, and 3-order boolean logics among genes. the detailed can be found in references [15]-[17] . 3. data source and processing 3.1 data source national center for biotechnology information (ncbi) and the european bioinformatics institute (embl-ebi) provide the database generally used in field of bioinformations. for description convenience, we refer to the sample data of stage a, b and c of colon cancer as experimental groups, normal stage as control group. the sample data is selected from gpl570 in ncbi. the sample numbers of stage a, b and c are 39, 103, and 92 respectively. they are all from gse2109 in gpl570. the 53 samples of normal stage consist of 10 samples of gds2609, 11 samples of gse10715 and 32 samples of gse8671. the data mentioned above include p_value and p-m-a, where p, a and m indicate presence, absence and margin respectively. we represent p with 1, a and m with 0. every sample database includes p_values and 0-1 expression profiles of 20827 genes of human (corresponding to 54676 probes). the four databases (normal, stage a, b and c) mentioned above form the original database used in this work. 3.2 selection of cancer-related genes in the original database mentioned above, every database includes the sample data of more than 20,000 genes of human genome. to construct and analyze the logical networks of more than 20,000 genes, the computational complexity is so high that it is beyond our computational ability. in order to simplify the computation, take the genes corresponding to 801 probes into account according to the known cancer-related genes [21] . if several probes correspond to the same gene, then we choose the data expressing the most in probes as the expression profile data of the gene. thus, 286 cancer-related genes in the original database are obtained. because we are concerned about the structure of the gene network, among these 286 cancer-related genes, the data almost 0 or 1 contribute very little to the difference of the network structure. hence, we delete these genes in whose expression profile the numbers of 1 are less than 15% or more than 90%. and then, the numbers of the remaining cancer-related genes of these four databases is normal: 91, stage a: 79, b: 70, c: 60 respectively. after the above processing, there are some genes with the same 0-1 expression profiles, which contribute equally the construction of network. hence, we preserve one of them, and delete other genes with the same expression profiles. thus, the number of genes in normal database is reduced to 79. after simplifying the samples treated as above, the obtained databases are the working databases. 4. numerical experiments, methods and results 4.1 determination of every order logics based on the above working databases, all the 1-order u values between genes are calculated respectively by the formula (1). obviously, there are 1-order u values of two directions between any two genes, denoted by  abu and  bau respectively. if the absolute value of the ratio of the difference between the two directions and the average value is less than advances in systems science and applications (2011), vol. 11, no. 1-2 http://www.iciba.com/contribute/ http://www.iciba.com/absolute%20value/ http://www.iciba.com/less/ certain threshold d , namely,         2 b a a b b a a b u u d u u    , then by the certainty of logical relationship, there is not 1-order logic between the two genes. only when         2 b a a b b a a b u u d u u    and the u value is greater than 1-order threshold, there is 1-order logic between genes a and b. if    baab uu  , then we consider that a regulates b, otherwise, b regulates a. by the formula (1) in the algorithm of u value,    a|ba|b uu , which means that the logic type between a and b might be ba  or ba . then we determine the logic type between a and b according to the support of ba  and ba , namely, the logic type with greater support is the one between a and b. here, the support refers the ratio of the number of the two states being 1 in the expression profile data of the genes and the total number of samples. for instance, suppose the expression profiles of the genes m and n are  0,0,0,1,1 and  1,0,0,1,1 respectively, then the supports of logic type m=n and m= n are 5 2 and 5 1 respectively. take the threshold of 1-order 3.01 u , then 143.0)m|n()m|n( uuu  , the support m=n m= n 2 1 5 5 s s    , thus we think the logic type of genes m and n is m=n . the proper functions of 2-, 3-order among genes in every database are determined by the formula (2), (3) and the support. that is to say, for any three genes a, b and c, calculating the u values corresponding to the 10 proper functions, then, the function 2f corresponding to the maximum u value is the proper function of a, b to c. similarly, for any four genes a, b, c and d, calculating the u values corresponding to the 218 proper functions, then, the function 3f corresponding to the maximum u value is the proper function of a, b , c and d. if the u values corresponding to several proper functions are the maximum value, then take the proper function with greater support as the one among genes (similarly to determining the logic type (or proper function) of 1-order). by the above numerical calculation, the complete directed logical networks corresponding to four databases of experimental and control groups, namely, each network contains all the proper functions of 1-, 2-, 3-order of its own database. 4.2 determination of threshold of every order in order to highlight the characteristics of the network and obtain the valuable biological information, we need to choose a threshold to carry out coarse graining to the complete directed logical networks mentioned above. the threshold is the expression of coarse-grained level and the structural difference between logical network structures should be contrasted in the same coarse-grained level. in order to make the logical network structures for experimental and control groups comparable, we normalize the u values of 1-, 2-, 3-order of the four networks, namely: logic u values of every order of each database are replaced with u max u , where max u is the maximum u value of each order in the own database. for simplicity, the normalized values are also denoted by the u values. let the thresholds of 1-, 2-, 3-order be 321 ,, uuu respectively. the 2-order logic of a , b to c is considered only when the 1-order logics of a to c and b to c do not exist, and claim that 192 xu:analysis for gene networks of colon cancer based on logical relationships http://www.iciba.com/consider/ 193 2u is greater than 1u :                12 11 11 22 )( )( )),(( uu ubfcu uafcu ubafcu , (4) the 3-order logic of a , b , c to d is considered only when the 1-order logics of a to d , b to d , c to d , and the 2-order logics of any two of a , b ,c to d do not exist, and 3u is not less than 2u :                  23 111111 222222 33 )(&)(&)( )),((&)),((&)),(( )),,(( uu ucfduubfduuafdu ucbfduucafduubafdu ucbafdu , (5) through the above numerical calculation, the four logical networks for experimental and control groups are obtained (fig. 1). the detailed situations can be seen in table 2. (a) normal (b) stage a (c) stage b (d) stage c fig. 1 taking 0.42, 0.75 and 0.85 as the threshold of 1-, 2-, 3-order respectively, the logical networks of 111 genes of the four databases are located, where (a), (b), (c) and (d) represent the logical networks corresponding to normal, stage a, b and c respectively. the nodes in the networks are divided into two types: the node labeling i denotes gene i, the node labeling “j(. . .)” (j = 2, 3) denotes the intermediate node of j-order logical relationships. for example, “2(5_1)” denotes logic type 5_1 of 2-order. the out-degree and in-degree of the intermediate nodes are fixed, i.e. when j =2, the in-degree and out-degree of the node is 2 and 1, respectively; when j = 3, the in-degree and out-degree of the node is 3 and 1, respectively. for instance, the node “2(3)” in fig. 1(b) denotes that gene cxcl and mmp2 regulate gene extl3 through logic type 3 of 2-order. table 2 the detailed list of the logical networks for control and three experimental groups stage no. of 1-order logics no. of 2-order logics no. of of 3-order logics no. of all the logics no. of genes no. of the non-isolated nodes normal 90 1614 39 1743 79 74 stage a 22 33 18 73 79 60 stage b 37 51 2 90 70 55 stage c 4 32 9 45 60 35 4.3 distribution of logic motifs of 2-order the subgraph with higher frequency in complex network is called the network motif [22] . the identification to the motifs of complex network helps to identify the typical local connections advances in systems science and applications (2011), vol. 11, no. 1-2 of network. previous studies have showed that: the motif in complex biological network has a direct biological significance. for example: the motif in the protein interaction networks of yeast highly evolves to protect the components, the transcription regulatory networks of different species have an evolutionary trend to the same motif [23] . in the logical networks, each of 8 logic types of 2-order is corresponding to a motif, called the logic motif. each type of logic motifs is corresponding to a typical mode of biochemical reactions. in order to explore the differences of reaction modes in the internal mechanism of normal and diseased organisms, we contrast and analyze the logic motifs of 2-order of the logical networks for experimental and control groups. fig. 2 and table 3 are the distribution histogram and the detailed list about logic motifs of 2-order of the four logical networks for experimental and control groups respectively. from fig. 2 and table 3, we can discover that the logic motifs of 2-order in control network cover six logic types 1, 2, 3, 4, 5, 6. however, in the networks for experimental groups, there only contain three logic types 1, 3, 6 (table 3) except that the network corresponding to stage a contains logic types 1,2,3,5,6,7 of 2-order. moreover, three logic motifs of types 1, 3 and 6 appear more in the network for control group and the corresponding proper functions are bac  , c a b  and bac  or c a b  , respectively; one logic motif of type 1 appears more in the network for stage a and the corresponding proper function is bac  ; two logic motifs of types 1 and 3 appear more in the networks for stage b and c and the corresponding proper functions are bac  , respectively. 1 2 3 4 5 6 7 8 0 100 200 300 400 500 600 logical type t h e n u m b e r o f l o g ic t y p e stage a stage b stage c normal 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 x 1 0 -3 logical type l o g ic a l ty p e f re q u e n c y stage a stage b stage c normal fig. 2 the distribution histogram of logic types of 2-order in the logical networks for control and three experimental groups table 3 the list of logic types of 2-order in the logical networks for control and three experimental groups logic type 1 2 3 4 5 6 7 8 normal 275 177 311 78 199 574 0 0 stage a 13 1 5 0 2 7 5 0 stage b 21 0 28 0 0 1 1 0 stage c 12 0 17 0 0 3 0 0 4.4 dynamical analysis of logical network although the study and analysis of topological structure of the above logical networks can obtain some biological principles and information, the structure of actual biological system is a changing network over time and the dynamical behaviors cannot be reflected by topological structure in most situations. therefore, we research the evolution of gene regulatory logical network from the dynamical point of view. because the number of the non-isolated nodes are too great in the logical networks corresponding to experimental and control groups to analyse the dynamical transfer laws, “the structural key genes” must be selected appropriately in every stages. the logical network is directed and every node in the network has out-degrees and in-degrees. if node i has an out-arc to j , then node i has a logical regulatory relationship to j . on the other hand, if node i has an in-arc, then some node has a logical regulatory relationship to i . the degree-difference of node is defined as the difference of out-degree and in-degree. the nodes 194 xu:analysis for gene networks of colon cancer based on logical relationships app:ds:on app:ds:the app:ds:other app:ds:hand 195 whose degrees and degree-differences are relatively great are called “the commander genes”, and the changes of states of these genes can influence most of genes of the network. in this work, the method of discovering the commander genes is as follows: (1) calculate the maximum degree and the maximum degree-difference of network; (2) normalize the the maximum degree and the maximum degree-difference of the non-isolated nodes respectively, namely: the degree and degree-difference of each non-isolated node divide the maximum degree and the maximum degree-difference of the network respectively. (3) set the parameters 1 0.05c  and 0.15c 2 . take these genes with degree 1c times greater than the maximum degree and degree-difference c2 times greater than the maximum degree-difference as the commander genes. when taking the parameters 1c and c2 , the number of commander genes must be appropriate, because some key genes might be lost if the commander genes are too few, and we cannot calculate if the commander genes are too many. the detailed situations of the commander genes of control and three experimental groups can be seen in table 4 and 5. table 4 taking the parameters 1 0.05c  and 0.15c 2 , the detailed list of the commander genes of control and three experimental groups stage the maximum degree the maximum degree-difference no. of the commander genes no. of the non-isolated nodes normal 855 109 12 74 stage a 30 15 12 60 stage b 42 6 12 55 stage c 32 5 13 35 table 5 the list of the commander genes of every stages normal stage a stage b stage c faslg araf araf ar axl bok bak1 bak1 etv3 esr2 esr1 elk1 gli3 ptk2b ptk2b erbb4 il1a gli2 fgf2 etv1 plag1 mafg hnf4a fgf2 rbl1 mdm2 ihh ihh tal1 pdgfb abcb1 mapk7 vav1 plag1 shh sele zap70 stk11 vav1 stk11 tcl1a tp73 arhgef5 tp53bp1 bag2 hpse hpse zap70 chek2 using the method of constructing the logical networks in (1) (2) of section 4, we can obtain the networks of the commander genes of control and three experimental groups (fig. 3) advances in systems science and applications (2011), vol. 11, no. 1-2 app:ds:relatively app:ds:influence app:ds:maximal app:ds:maximal app:ds:normalization app:ds:maximal app:ds:maximal app:ds:divide app:ds:maximal app:ds:maximal app:ds:maximal app:ds:maximal app:ds:maybe app:ds:maximal app:ds:maximal (a) normal (b) stage a (c) stage b (d) stage c fig. 3 the logical networks of the commander genes of control and three experimental groups. now we begin to analyse the dynamics of the above logical networks of the commander genes. assume the network have n genes. let the state of gene i at moment t be }1,0{)( txi , and the state-vector of gene i at moment 1t be ))1(),...,1(),1(()1( 21  txtxtxts n . without loss of generality, the triple set comprised of the u values, the corresponding proper functions of 1-order and the genes affecting gene i at moment t is 1 1 {( , , ) {1,2,..., }, }k i kz k u f k n k i   , which denotes that gene k regulates gene i by the proper function 1 kf and the uncertainty coefficient is iku  ; the triple set consisting of the u values, the corresponding proper functions of 2-order and the genes affecting gene i is 1 2 1 2 2 2 1 2 , , 1 2 1 2 1 2{(( , ), , ) , {1,2,..., }, , , }k k i k kz k k u f k k n k i k i k k     , which denotes that gene 21,kk regulate gene i by the proper function 1 2 2 ,k kf and the uncertainty coefficient is ikku 21 , ; the triple set consisting of the u values, the corresponding proper functions of 3-order and the genes affecting gene i is 1 2 3 1 2 3 3 3 1 2 3 , , , , 1 2 3 1 2 3 1 2 1 3 2 3{(( , , ), , ) , , {1,2,..., }, , , , , , }k k k i k k kz k k k u f k k k n k i k i k i k k k k k k        , which denotes that gene 321 ,, kkk regulate gene i by the proper function 1 2 3 3 , ,k k kf and the uncertainty coefficient is ikkku 321 ,, . in this research, we consider that the state of gene i at moment 1t is related to the states of itself and its adjacent nodes. at moment t , let genes 1 2, , , jk k k regulate gene i by proper function 1 2, , , j j k k kf  (  3,2,1j ). if the truth-value of proper function 1 2, , , j j k k kf  with independent variable being the state of genes jkkk ,,, 21  at moment t is 1, then these j genes jkkk ,,, 21  activate the expression of gene i . let 1 2, , , sign ( ) 1 jk k k i t    . otherwise, these j genes jkkk ,,, 21  repress the expression of gene i . let 1 2, , , sign ( ) 1 jk k k i t     . at moment t , for gene i , let )(sign)(sign)(sign)()( 3213 321 212 21 1 ,, ,, , , tutututxtp ikkkz ikkk ikkz ikk ikz ikii         . (6) then, the transition rule is denoted by 196 xu:analysis for gene networks of colon cancer based on logical relationships app:ds:consist app:ds:of app:ds:consist app:ds:of app:ds:adjacent%20points app:ds:truth-value app:ds:independent%20variable 197       ,5.0)(,0 ,5.0)(,1 )1( tp tp tx i i i ni ,...,2,1 . (7) according to the above transition rules, the state transition configuration network for control and three experimental groups are obtained. the detailed situations are listed in table 6, fig. 4 and fig. 5. in the networks corresponding to control and three experimental groups, the numbers of nodes (genes) have certain differences. in order to eliminate the differences caused by the number of attractors, the stability of system is measured by the frequency of attractors instead of the number of attractors. table 6 the situations of attractors of control and three experimental groups stage no. of genes no. (frequency) of 1-periodic attractor no. (frequency) of other attractor no. (frequency) of all the attractors normal 12 262(6.40%) 22(0.54%) 284(6.94%) stage a 12 594(14.50%) 6(0.15%) 600(14.65%) stage b 12 407(9.94%) 0(0%) 407(9.94%) stage c 13 1812(22.12%) 116(1.42%) 1928(23.54%) fig. 4 2-periodic domain of attraction with the most nodes (46 nodes) of control group, where the node is labeled by the decimal number corresponding to the binary state vertor. (a) normal (b) stage a (c) stage b (d) stage c fig. 5 in the state transition configuration networks for control and experimental groups, the 2-periodic domain of attraction with the most nodes, where the transition networks for normal, stage a, b and c contain 111, 42, 40, 16 nodes respectively. 5. conclusions and discussions in this work, we construct the logical networks using lapp method based on the expression data of phylogenetic profiles of cancer-related genes in normal, stage a, b and c of colon cancer . through contrasting and analyzing of the structures of four networks corresponding to advances in systems science and applications (2011), vol. 11, no. 1-2 app:ds:eliminate app:ds:attractor app:ds:stability app:ds:measure app:ds:decimal%20number app:ds:binary http://dict.cnki.net/dict_source.aspx?searchword=profiles app:ds:colon%20cancer app:ds:colon%20cancer app:ds:colon%20cancer app:ds:contrast experimental and control groups, we find the significant differences in the structural components (fig. 1 and table 2). (1) the significant differences in the number of non-isolated nodes: in the network for control group, the number of non-isolated nodes is 74; in stage a, b and c of colon cancer, the numbers of non-isolated nodes are reduced obviously, moreover, the fewer the numbers of non-isolated nodes are, the more cancer deteriorates. (2) the significant differences in the the number of logical relationship: the number of logics of the network for control group is far more than the one of the three experimental groups. in the complex networks, the number of non-isolated nodes reflects the intensity of correlation relationship of nodes and the number of the relationships of nodes indicates the richness of the relationship of nodes of network. in the complex networks, most of genes in the normal stage play cancer suppression roles with other genes cooperating perfectly. for some (internal and external) reasons, the regulatory relationships of several genes with some other genes are lost completely or the regulatory pathways malfunction seriously. the normal tissue appears abnormal; the organism suffers from the colon cancer. with the intensifying of the failure of regulation and broadening of the failure propagation, the illness increase gradually. (3) the significant differences in the distribution of logic motifs of 2-order (see fig. 2 and table 3): in the network for control group, there are three logic motifs: type 1, 3 and 6; in the network corresponding to stage a, there are one logic motif: type 1; in the networks for stage b and c, there are two logic motifs: type 1 and 3. we find from the above results that the logic type 1 appears in the networks for control and experimental groups, yet the logic type 6 appears only in the network for control group. in this research, we consider that logic types (motifs) of 2-order are corresponding to typical modes of biochemical reactions in life actions. the above phenomena might show that logic type 1 ( bac  ) is the basic mode maintaining human existence and physiological metabolism reactions, yet logic type 6 (c a b  or c a b   ) is the reaction mode suppressing the normal organism not suffering from the colon cancer. the materials [24] have showed that logic type 1 does appear frequently in the complex biochemical reactions of the human body. for example: the coupling protein x and y plays a regulatory role in the expression of protein z, where x, y and z are the specific proteins. in conclusion, the logical networks for control and experimental groups are the models of structural mechanism of the normal organism suffering from the colon cancer gradually. the organism is normal in the normal stage because the biochemical reaction mode (type 6), the extent of genes relationship (the number of logical relationship) and most of the genes (the number of non-isolated nodes) play normal roles. with the change of the reaction mode of the organism and the decreasing of the number of acting genes and their relationships, many intrinsic biological functions are weaken or lost gradually, thus the organism suffers from the colon cancer. however, the significant differences between the networks for control and experimental groups mentioned above are the causes or the consequences of the organism suffering from the colon cancer, which will also need to be verified by the clinical experiments and annlysis of biomedical researchers. table 5 shows the commander genes of control and three experimental groups. we can find that: three genes plag1, vav1 and zap70 appearing in control group are also active in the experimental groups, and the three genes might regulate the basic survival activities of human. the material shows: gene vav1 as a kind of signal transduction molecule can happen quickly tyrosine phosphorylation with ifn-α, il-3, gm-csf, growth factor or antigen receptors after the activation and plays an important role in the process of cell differentiation, proliferation [21] . gene zap70 plays a very important role of transfer signal in the process of activation of t-cell mediated by cd3 and/or cd3 chain [21] . other genes faslg, axl, etv3, gli3, il1a, rbl1, tal1, tcl1a, bag2 of control group do not appear in the experimental groups and they might be the suppressor genes of colon cancer in human genome. furthermore, genes araf,bak1,ptk2b,fgf2,ihh,stk11,hpse appear in at least two stages of the 198 xu:analysis for gene networks of colon cancer based on logical relationships app:ds:significant app:ds:difference app:ds:significant app:ds:difference app:ds:significant app:ds:difference app:ds:reflect app:ds:indicate app:ds:tissue app:ds:organism app:ds:significant app:ds:difference app:ds:biological%20phenomena app:ds:suppressor app:ds:organism app:ds:material app:ds:specific app:ds:mechanism app:ds:organism app:ds:gradually app:ds:organism app:ds:extent app:ds:normal app:ds:change app:ds:organism app:ds:decrease app:ds:acting app:ds:biological app:ds:weaken app:ds:gradually app:ds:organism app:ds:significant app:ds:difference app:ds:mention app:ds:organism app:ds:clinical%20experiment app:ds:active app:ds:material app:ds:growth%20factor 199 experimental groups and not in control group, thus the seven genes might be the oncogenes of colon cancer. the facticity and reliability predicted in the above will also need to be verified by biological experiments. in the process of the dynamic analysis of logical network of the commander genes, we find from the numerical experiments that the number of attractors of control and three experimental groups has obvious difference (table 6). the former is far less than the latter. according to the principle and theory of dynamics: the less the number of attractors is, the more stable the system is. from the view of biology, the construction should be more stable in the normal organism than the stages of colon cancer. the views from two points are just coincident. thus, we conclude that the instability of the commander genes of organisms might be an important reason or result in the process of the colon cancer disease. we utilize lapp to search 1-, 2and 3-order logical relationships and logic types among the genes or proteins. using lapp to determine the logical relationships of genes and proteins, there might appear „false-positive‟ and „false-negative‟. that is to say there exist some errors during determining logical relationships. we do not analyze the errors caused by lapp in this paper. as to how to overcome the shortcomings of lapp method and seek the better algorithm of u value, we are making further study. references [1] lander e s, weinberg r a. genomics: journey to the center of biology. science, 2000, 287(5459): 1777-1782. [2] zhenfang liao, dean lei, wei shen. the research progress of genes related to the metastasis of colon cancer. modern oncology, 2006, 14(9): 1174-1176. [3] zhang h, yu c y, si nger b, et al.. recursive partioning for tumor classification with gene expression microarray data. pnas, 2001, 98: 6730-6735. [4] xia li, shaoqi rao, tianwen zhang, zheng guo, qingpu zhang, k. l. moser, e. j. topol. integrated decision method of mining genes related to complex disease appling the dna chip data. science in china, ser. c-life science, 2004, 34(2):195-202. [5] tibsh irani r, ha s ti e t, narasim han b, et al.. diagnosis of multiple cancer types by shrunken centroids of gene expression. pnas, 2002, 99(10): 6567-6572. [6] khan, j.,wei, j., ringner, m., saal, l., ladanyi, m.,westermann, f., berthold, f., schwab, m., antonescu, c., peterson, c., et al.. classification and diagnostic prediction of cancers using gene expression profiling and artificial neural networks. nat. med., 2001, 7(6): 673-679. [7] alon u, barka in, notterman d a, et al.. broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays. pnas, 1999, 96: 6745-6750. [8] huang ds, zheng ch. independent component analysis-based penalized discriminant method for tumor classification using gene expression data. bioinformatics, 2006, 22 (15): 1855-1862. [9] ambroise, c, mclachlan, g j.. selection bias in gene extraction on the basis of microarray gene-expression data. pnas, 2002, 99: 6562-6566. [10] hastie t, tibshirani r, eisen m b, et al.. „gene shaving‟ as a method for identifying distinct sets of genes with similar expression patterns. genome biol., 2000, 1: research0003.1-21. [11] li xia, rao shao qi, wang yadong, et al.. gene mining: a novel and powerful ensemble decision approach to hunting for disease genes using microarray expression profiling. nucleic acids research, 2004, 32 (9): 2685-2694. [12] marcotte em, pellegrini m, ng hl, rice dw, yeates to, eisenberg d. detecting protein advances in systems science and applications (2011), vol. 11, no. 1-2 app:ds:oncogene app:ds:the%20latter app:ds:organism app:ds:point%20of%20view app:ds:coincident app:ds:organism app:ds:algorithm function and protein–protein interactions from genome sequences. science, 1999, 285: 751-753. [13] strong m, graeber tg, beeby m, pellegrini m, thompson mj, yeates to, eisenberg d. visualization and interpretation of protein networks in mycobacterium tuberculosis based on hierarchical clustering of genome-wide functional linkage maps. nucleic acids res., 2003, 31: 7099-7109. [14] zhou x, kao mc, wong wh. transitive functional annotation by shortest-path analysis of gene expression data. pnas, 2002, 102(38): 12783-12788. [15] bowers p m, cokus s j, eisenberg d, et a1.. use of logic relationships to decipher protein network organization. science, 2004, 306(5706): 2246-2249. [16] bowers p m, o'connor b d, cokus s j, et al.. utilizing logical relationships in genomic data to decipher cellular processes. the febs journal, 2005, 272 (20): 5110-5118. [17] xin zhang, seungchan kim, tie wang and chitta baral. joint learning of logic relationships for studying protein function using phylogenetic profiles and the rosetta stone method. ieee transactions on signal processing, 2006, 54 (6): 2427-2435. [18] zhang j x, ji y, zhang l. extracting three-way gene interactions from microarray data. bioinformatics, 2007, 23(21): 2903-2909. [19] shudong wang,yan chen, qingyun wang, eryan li, yansen su, dazhi meng. analysis for gene networks based on logical relationships. journal of systems science and complexity, 2010, 23(5): 999-1011. [20] qingyun wang, hui yi, laifu liu, dazhi meng. progress in gene logic networks. progress in biochemistry and biophysics, 2008, 35(11): 1239-1246. [21] lǚsheng si, xu li. oncogene, cancer suppressor gene, cancer-related gene. shanxi technology and science press, xi‟an, 2002.12. [22] milo r, shen orr s s, itzkovitz s, et al.. network motifs: simple building blocks of complex network. science, 2002, 298: 824-827. [23] barábasi a l, oltvai z n. network biology: understanding the cell‟s functional organization. nature reviews-genetics, 2004, 5: 101-114. [24] fuchu he, pengyuan yang, yunping zhu. systems biology in practice: concepts, implementation and application. fudan unversity press, shanghai, 2007.12. 200 xu:analysis for gene networks of colon cancer based on logical relationships http://www.cqvip.com/qk/94860x/200811/28883778.html adv syst sci appl 2019; 03; 140-162 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/676 scada based remote monitoring and control of pressure & flow in fluid transport system using imcpid controller e.b.priyanka1*, c.maheswari2, m.ponnibala3, s.thangavel4 1) senior research fellow, department of mechatronics engineering, kongu engineering college, perundurai, india. e-mail: priyankabhaskaran1993@gmail.com 2) faculty, department of mechatronics engineering, kongu engineering college, perundurai, india. e-mail:maheswari@kongu.ac.in 3) faculty, department of mechatronics engineering, kongu engineering college, perundurai, india. e-mail:ponnibala@kongu.ac.in 4) faculty, department of mechatronics engineering, kongu engineering college, perundurai, india. e-mail:thangavel.mts@kongu.ac.in received may 5, 2018; revised september 9, 2019; published october 1, 2019 abstract: ubiquitous sensing enabled for monitoring and control in the industrial sector are mostly wsn (wireless sensor network) based systems or scada systems. wsn based systems are homogenous or compatible systems and they support coordinated communication and transparency among regions and processes. hence in order to enhance the proprietary of remote monitoring and control technology, this paper emphasizes the development of architecture with local intelligent using imc based pid controller can be integrated with dcs/scada and this research work spectacle only the development and application of local intelligence architecture for remote monitoring and control of concerned field parameters. this proposed local intelligence tactic governs the input and output from a hardware platform positioned on a process plant, organize added processing power for meticulous analysis in the control center by the use of software integration, data acquisition and logging on a created hybrid database, and spectacle significant processed post-data statistics to operators via a standard dcs/scada hmi. the established local intelligence is experimentally validated in the lab scale experimental fluid transport system to monitor and control pressure and flow rate parameters remotely by incorporating with centum cs 3000.the simulation and experimental results of local intelligence using an imc-pid controller on a lab scale fluid transport system are conveyed with its numerical data. keywords: local intelligence, imc based pid controller, remote monitoring & control. 1. introduction now-a-days, supervisory control and data acquisition (scada) systems [1, 8] are not climbable due to derisory memory and processing time, offers high cost of hardware drivers and management, not malleable when a process plant hardware needs to be updated, inflexible when there is a need for protocol change and software updating, and provides data and result with long delay. hence, human interfaced scada [4, 5] and the engineering internet are reconnoitered towards the improvement of modules used to identify the manifestation of impairment. in this upgradation era, all renovated industries has been assured using the emerging methodology remote monitoring and automation [19] which permit a pervasive bonding between the process plant hardware (sensors, actuators, control panel, etc.), smart hybrid objects and networked implanted devices to accumulate applicable data that will be * corresponding author: priyankabhaskaran1993@gmail.com scada based remote monitoring and control of pressure & flow rate 141 copyright ©2019 assa. adv. in systems science and appl. (2019) communicated wireless manner to the remote information analytics center to be examined and modulated in order to construe the process and afford the appropriate decision in case of calamitous circumstances [7, 11]. this paper emphasizes a reliable monitoring and control with local intelligence using imcpid controller modular architecture design along with scada/dcs. this architecture with local intelligence using imc-pid controller facilitates minimum human intervention will provide better workplace safety, maintenance of assets and will perform predictive maintenance of various industrial assets by analyzing various parameters (sensed data) and detecting failure modes either before they are going to take place or when the equipment will likely to fail or need service. this paper enlightens the brief information on the development of local intelligence using an imc-pid controller which is present in the proposed iot architecture. to overcome the limitations of conventional remote monitoring and control systems [2, 3, 6], the proposed design of local intelligent offers a new utility monitoring and control system can amenably put up variations made to progress the functioning efficacy using an imc-pid controller. the imc pid tuning rubrics have the benefit of utilizing only a solitary tuning parameter to accomplish a vibrant trade-off between the closed-loop enactment and robustness. over the unified incorporation of self-governing practicalities in a integrated checking and control system, an complete process system can be supervised and functioned in real time from the dcs/scada hmi through this local intelligence incorporated with imc-pid controller which creates conceivable based automatic activation of tuning window block with trends, substantively enlightening engineering proficiency and working safety by exploiting centum cs 3000. 2. lab scale experimental setup of fluid transport system the lab scale experimental fluid transport system consists of the pumping unit, the transmitting pipes having a diameter of 1 inch about 15m long for the well-organized and extended transmission and manifold analog sensors at different distinctive locations with its corresponding control valves installed to compensate its drops/loss during the fluid transportation in the pipelines. the block diagram with local intelligence for the lab scale investigational setup envisaging the physical categorizations of apparatus positioned and also exhibits all the piping with directional tracks and equipment’s facts with their controls is shown in figure 1. the process plant consists of two sections consisting of an electric pump, differential pressure transmitter holding a range of 03-15psi (0.1-3kg/cm2), and pressure control valve at one section along with orifice flow meter having the limit of 0-1800lph, pump and flow control valve at another section. when the process plant is on track to run, initially reservoir tank fills up to 20% of its capacity, and then the electric pump is actuated to suck the fluid from the tank in order to transmit the fluid to each section. when the fluid starts to pass through the pipelines its corresponding pressure and flow transmitter send its present pressure and flow rate data to the i/o hub module station. during transportation to regulate the pressure and flow rate of the transmitting fluid, the opening and closing of the control valves are operated by the i/p converter in order to maintain its preferred effective range limits until the destination of the long run, till it drains from the process tank as shown in figure 2. the local intelligence will take the role to regulate and control the process plant parameters before it reaches the state of over limit/threshold. since this paper encompass on the development and application of local intelligence using imc-pid controller, the designed local intelligence can remotely monitor and control the pressure and flow rate as a separate loop through dcs/scada hmi which can be operated as stand-alone station without depending to a central server with mutual backup configuration by regulating the operation of the corresponding control valves to reach its desired operating points of pressure and flow rate. 142 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) fig. 1.block diagram with local intelligence for the process plant. fig. 2.the lab scale experimental setup of the fluid transport system 3.functional deployment of local intelligence scada frame work the scada incorporating local intelligence using an imc-pid controller is developed for the experimental setup of the fluid transport system is shown in figure 3 to deal with remote surveillance and control of flow and pressure and it also encompasses of individual interface displays for flow rate and its respective pressure parameters of the fluid as shown in figure 3 input/ output hub module input ports pressure-ios flow-ios pressure transmitter pneumatic actuated control valve i/p converter flow transmitter pneumatic actuated control valve i/p converter pump 2 pump 1 output ports field control station (fcs) high level engineering interface with controller engineering station oil station 2 lab scale setup scada based remote monitoring and control of pressure & flow rate 143 copyright ©2019 assa. adv. in systems science and appl. (2019) and 4. the developed scada provides more functional utility options like system message banner, graphic view with graphics and control attributes, trend view, browser bar and tuning window. the system message indication window articulates the alarm existence eminence visually. the alarm existence eminence is presented by colors and blinking of action buttons, and the message presentation. the system message indication window is constantly shown at the header of the scada hmi, so will never be veiled overdue by other windows. the browser bar is utilized to call up process and manipulative windows to entact the scheduled tasks. it also spectacles a list of task oriented manipulating windows and process hierarchical configurations in a tree-like manner, enabling the whole process plant to be effortlessly functioned. the graphics attributes view tool parades plants conditions graphically and can be instinctively activated and monitored. the control attribute graphic window shows the function block eminences by elaborating instrument faceplates. fig. 3.scada view of the lab scale experimental fluid transport system 4. system modeling for pressure and flow control loop 4.1.open loop response analysis for modeling the fluid transport system, a transient response curve is chronicled by modifying the control valve opening in order to acquire the equivalent liquid pressure and flow rate changes on the pipeline in the open loop structure. this open loop experimentation reveals pressure of the liquid is at a maximum rate and the flow is at a minimum value when the 144 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) opening of the control valve is around 10% of the overall opening and vice versa when the control valve opening reaches its full stretch of 100%. figure 4.individual scada view of pressure and flow control loop the open loop test run is carried out by linearly adjusting the percentage of control valve opening. the experimental analysis call ups with the initial flow rate of 179 lph and pressure of 2.2 kg/cm2 and unremitting readings were documented until the flow rate and pressure influences a balanced/fixed state. the result divulges that for 100% control valve opening the ultimate flow rate and pressure accomplished are 1792 lph and 0.22 kg/cm2. open loop readings were noted for the percentage of control valve opening versus flow rate and pressure through centum cs 3000 as displayed in figure 5 and obtained data are presented in table 1 by which the first order model considerations (process gain kp and process time constant p) are calculated. the evaluated model parameters for pressure and flow rate are given in table 2. . table 1.open loop response analyses of pressure and flow for a different level of control valve openings to enumerate first order strictures such as process gain kp, process time constant ʈp and time delay 𝜃, with 𝐾𝑝 = [initial value − final value of parameter]/ [maximum value of parameter] [initial value of opening − final value of opening]/ [maximum value of control valve opening] percentage of control valve opening (in %) flow rate (in lph) pressure (in kg/cm2) 10 179 2.20 20 230 1.92 30 373 1.83 40 468 1.71 50 585 1.48 60 757 1.21 70 919 0.91 80 1137 0.55 90 1429 0.36 100 1792 0.22 scada based remote monitoring and control of pressure & flow rate 145 copyright ©2019 assa. adv. in systems science and appl. (2019) ʈ𝑃 = 1.5 ∗ [t2 − t1] (4.1) fig. 5. pressure and flow rates of the liquid for the different level of control valve opening t1 at a1 = [initial value of parameter((initial value of parameter– final value of parameter)*0.632)] t2 at a2 = [initial value of parameter ((initial value of parameter– final value of parameter)*0.283)] θ = (𝑇2 − ʈ𝑃) (4.2) 146 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) table 2.recognized model parameters at different % of control valve opening percentage of control valve opening (%) flow rate pressure kp ʈp (sec) 𝛉 (sec) kp ʈp (sec) 𝛉 (sec) 20 0.329 11.941 7.91 1.681 4.95 13.58 40 0.485 8.317 12.83 1.364 6.43 22.84 60 0.914 9.084 15.72 1.649 4.026 35.59 80 0.046 13.51 9.89 1.024 9.88 21.91 100 0.871 17.46 5.16 0.995 14.32 29.92 4.2.model identification for pressure and flow control loop from table 2, it shows that the behavior of a process is nonlinear and stable. hence, first order plus time delay (foptd) transfer function ( g(s) = kp ps+1 𝑒−s ) model is used to exemplify the pressure and flow rate maintenance of fluid transport system, where kp = process gain, p = time constant and  = process delay. then, the worst case model with the leading process gain and lowest time constant is designated to epitomize the process plant model [12, 13]. the identified foptd model for the flow control loop in the process plant is represented as, g(s) = 0.914 8.317s+1 𝑒−5.16s (4.3) similarly the foptd model for the pressure control loop is exemplified as, g(s) = 1.681 4.026s+1 𝑒−13.58s (4.4) the system model identification is arrived by formulated the real-time experimental data obtained from the fluid transport system in open loop performance analysis on process plant [16]. 5. robust controller design the pid (proportional-integral-derivative) controller serves as one of the popular and extensively applied controllers in the industrial sector due to its uncomplicatedness, robustness and affords wide applicability to near-optimal enactment. even though innovative control techniques can deliver substantial improvements, an elegant-characterized pid controller has ascertained to be suitable for an enormous quantity of engineering control loops. a pid type controller is used to optimize the performance of the control valve in edict to regulate and uphold pressure and flow-rate of the fluids being transported in the oil pipeline system. the proportional term ensures the trade of fast-acting modification which makes necessary alteration of output as rapidly as the error ascends. the integral part proceeds after a finite time but has the proficiency to create the steady state error as zero and the derivative term can advance proficiently the control loop performance. 5.1 ziegler-nicholas pid (zn-pid) controller the ziegler-nichols pid (zn-pid) controller is the utmost universally employed heuristic technique of fine-tuning a pid controller in all the engineering oriented feedback control application. the ability to predict the future errors in the process is possible in the pid scada based remote monitoring and control of pressure & flow rate 147 copyright ©2019 assa. adv. in systems science and appl. (2019) controller, meanwhile, it can eliminate oscillations and can decrease the rise time in the performance [9]. since the pressure and flow control process is the first order based system with time delay characteristics, hence the implementation of zn-pid controller is the benchmark of conventional techniques used for the comparative purpose. the parameters of zn-pid controller is exposed in table 3. table 3.zn-pid controller tuning parameters with values tuning rules tuning parameters for pressure tuning parameters for flow rate k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) kc = ap θkp a∈ [1.2,2] i = 2θ and d = 0.5θ 1.84 3.77 1.42 0.490 2.629 1.28 5.14 0.699 0.25 0.9 where kc is controller gain, i, d indicate integral time and derivative time. 5.2 cohen-coon pid (cc-pid) controller cohen-coon tuning system is the more difficult and complex form of open loop ziegler nicholas method and its tuning parameter relations were established empirically to afford closed loop reaction for a feedback system which sustains ¼ decay ratio. the reformed cohencoon tuning procedures [9] is an outstanding scheme for attaining fast response on practically all control loops with self-adapting practices and to use on a non-interactive controller algorithm to provide a fast response on the process plant. the tuning parameters relations are given and their corresponding controller parameter of cc-pid is presented in table 4. 𝐾𝑐 = 1 𝐾𝑝 . ʈ𝐩 θ ( 4 3 + θ 4ʈ𝐩 ) (5.1) 𝑖 = 𝑑 32+6θ/ʈ𝐩 13+8θ/ʈ𝐩 (5.2) 𝑑 = 𝑑 4 11+2θ/ʈ𝐩 (5.3) where kc is controller gain,i and d are integral time and derivative time respectively. table 4.cc-pid controller parameters controller tuning parameters for pressure tuning parameters for flow rate k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) cohencoon (ccpid) controller 1.84 3.77 1.42 0.490 2.629 1.28 5.14 0.699 0.25 0.9 148 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) 5.3. internal mode pid (imc-pid) controller the efficiency of the internal model control (imc) strategy norm [9,11] has prepared it eye-catching in the process industrial sectors, where numerous endeavors have been done to accomplish the imc standard to develop pi/pid controllers in cooperation with stable and unstable process system. the eminent internal model control-pid (imc-pid) tuning rules have the improvement of solitary using of single tuning constraint to succeed a clear transaction between closed-loop presentation and robustness to exemplary inexactitudes and also serve as a newfangled online controller fine-tuning technique in closed-loop approach. the direct synthesis upon disturbance (ds-d) technique proposed by edghar and seborg [14] can be considered into the identical class of the imc-pid procedures, in that they achieve the pi/pid controller strictures by calculating the epitome feedback controller which contributes a predefined anticipated closed-loop retort. even though the ultimate feedback controller upholding both the imc and ds is habitually more problematical than the pid controller for time delayed practices, the controller arrangement can be condensed to that of either a pi or a pid based controller cascaded along with lower order filter by execution of applicable guesstimates of the dead time in process plant model. it is important to accentuate that the pi/pid controller premeditated according to the imc procedure affords first-rate setpoint tracking and consents the supervision of stable, unstable, and integrating system processes. this closed-loop fine-tuning method overwhelms the inadequacy of the renowned ziegler nichols unremitting cycling routine and contributes a reliably enhanced performance and forcefulness for a wide-ranging scheme of control process. 5.3.1 imc-pi/pid controller design for fluid transport process plant figure 6(a) and (b) show the block diagrams of the imc control [15-17] and corresponding traditional feedback control configurations, respectively, where 𝐺𝑝 points the process, 𝐺�̃� denotes the process model, q and 𝑓𝑟indicates the imc controller, the set-point filter, and 𝐺𝑐 corresponds to the comparable feedback controller. for the trifling case (i.e.,𝐺𝑝 =𝐺�̃�), the setpoint and disturbance reactions in the imc scheme can be streamlined as: 𝑌 = 𝐺𝑝𝑞𝑓𝑟 + (1 − 𝐺𝑝𝑞)𝐺𝑝𝑑 (5.4) conferring to the imc characterization (morari and zafiriou [10]), the process model 𝐺�̃� is sub-mani folded as factor of two parts: 𝐺�̃� = 𝑃𝑚𝑃𝐴 (5.5) where pm is the percentage of the model reversed by the controller, pa is the ration of the model not overturned by the controller and pa(0)= 1. the non-reversible part typically embraces the dead time and right side occupied half plane zeros and is elected to be all-pass. to acquire a enhanced reaction for unstable or stable processes with poles situated near zero, the imc controller q must gratify subsequent conditions: if the process gp possess unstable poles or poles positioned near zero at z1, z2… zm, then (i) q must ensure zeros at z1, z2, …, zm (ii) 1− gpq would also obligate zeros at z1, z2, …, zm scada based remote monitoring and control of pressure & flow rate 149 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 6.block diagram of the imc and corresponding conventional feedback control systems. for a fluid transport system process the methodology is encouraged by the enactment enhancement of the disturbance rejection predominantly. the overhead design principles are obligatory for the internal steadiness of the unstable method. the supplementary profit of utilized criteria is that they support in the performance progress of the control loop systems. meanwhile the imc controller q is intended as 𝑞 = 𝑃𝑚 −1𝑓, the paramount condition is fulfilled inevitably. the second condition can be contented by conniving the imc filter (f) as 𝑓 = ∑ 𝛼𝑖𝑠𝑖𝑚 𝑖=1 (𝜏𝑐𝑠+1)𝑟 (5.6) where τc is an modifiable parameter of the fluid pipeline system which reins the adjustment between the enactment and robustness; r is nominated to be great ample to sort the imc controller (semi-)accurate; αi are evaluated by equation (12) to abandon the poles 1 − 𝐺𝑃𝑞|𝑠=z1,z2,…,z𝑚 = |1 − 𝑃𝐴(∑ 𝛼𝑖𝑠𝑖𝑚 𝑖=1 ) (𝜏𝑐𝑠+1)𝑟 | 𝑠=z1,z2,…,z𝑚 = 0 (5.7) situated near zero around 𝐺𝑝. then, the imc controller emanates to be 𝑞 = 𝑝𝑚 −1(∑ 𝛼𝑖𝑠𝑖𝑚 𝑖=1 +1) (𝜏𝑐𝑠+1)𝑟 (5.8) as a result, the resultant set-point and disturbance retorts are attained as: y r = (gpq)fr = pa(∑ αisi+1m i=1 ) (τcs+1)r fr (5.9) y d = (1 − gpq)gp = (1 − pa(∑ αisim i=1 ) (τcs+1)r )gp (5.10) the numerator countenances ∑ 𝑎𝑖 𝑚 𝑖=1 𝑠𝑖 + 1 in equation (5.9) reason to be occasionally an unwarranted overshoot in the servo response that can be eradicated by presenting the set-point 150 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) filter fr to recompense for the overshoot in the servo response. as of the beyond design practice, a stable, closed-loop reaction can be reached by exhausting the imc controller. the ultimate feedback controller that is correspondent to the imc controller [18] can be articulated for the fluid transport system in expressions of the internal model 𝐺�̃� and the imc controller q: gc = q 1−gpq (5.11) relieving equations. (10) and (13) into (16) contributes the ideal feedback controller: gc = pm −1∑ αisim i=1 (τcs+1)r 1− pa(∑ αisim i=1 ) (τcs+1)r (5.12) the resultant controller in equation (5.12) is substantially realizable for the fluid process plant, but it does not possess the customary pi/pid usage. the anticipated form of the controller can be attained by actuating the guesstimate of the dead time strategy in the process plant. meanwhile, the final derivation exploits both easiness and estimate error due to dead time mode deliberated carefully during the pi/pid controller strategy. first order plus dead time (fopdt) form is a descriptive model normally followed in the control process industries. on the source of the above detailed strategy principle, the fopdt process plant has been reflected as g(s) = kp ps+1 𝑒−s (5.13) where kp gives process gain, p and  denotes process stimulated time constant = process delay, the imc filter nominated is f = αs+1 (𝜏𝑐𝑠+1)2 (5.14) after exploiting the above examined principle the ultimate feedback controller is specified as 𝐺𝑐 = (𝜏𝑠+1)(𝛼𝑠+1) 𝐾[(𝜏𝑐𝑠+1)2−𝑒−𝜃𝑠(𝛼𝑠+1) (5.15) subsequently the ultimate feedback controller in equation (5.15) does not contain pi controller term; the residual assignment is to plan the pid controller that estimates the ideal feedback controller best meticulously. the ultimate feedback controller, gc, corresponding to the imc controller [20], can be attained later by the calculation of the dead time allocated for the poles by taylor series expansion, 𝑒−s =1−θs and outcomes in 𝐺𝑐 = (𝜏𝑠+1)(𝛼𝑠+1) 𝐾[(𝜏𝑐𝑠+1)2−(1−𝜃𝑠)(𝛼𝑠+1) (5.16) subsequently reorganizing of equation (5.16) gives 𝐺𝑐 = (𝛼𝑠+1) 𝐾(2𝜏𝑐−𝛼+𝜃)𝑠 (𝜏𝑠+1) [ (𝜏𝑐 2+𝜏𝑐𝜃) (2𝜏𝑐−𝛼+𝜃) 𝑠+1] (5.17) from equation (5.17), the consequential pid controller can be achieved after vulgarization as kc= 𝛼 𝐾𝑝(2i −α+) ; i=α ; d=2 (5.18) where kc points controller gain, ti and d are integral time and derivative time of the model designed process [21-24]. the significance of α (non-minimum phase element) is designated so that it terminates out the pole located at s=−1/τ and the rate of α is attained as 𝛼 = 𝜏{1 − (1 − 𝜏𝑐 𝜏 )2 𝑒− 𝜃 𝜏 ; c=2θ (5.19) by implementing this tuning procedure, it is conceivable to develop the enriched disturbance rejection action by adjusting the single fine-tuning of parameter in the controller. the important feature of this controller is that it compacts with the nonlinear and stable oriented plant system in a incorporated way. the parameters of the imc-pid controller are revealed in table 5. scada based remote monitoring and control of pressure & flow rate 151 copyright ©2019 assa. adv. in systems science and appl. (2019) table 5.imc-pid controller parameters controller tuning parameters for pressure tuning parameters for flow rate k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) k c i (sec) d (sec) k i = (k c /i) k d = (k c *d) imc-pid controller 6.43 2.64 0.67 2.428 4.323 1.89 1.26 0.158 1.5 0.3 6. results and discussion the simulation and real-time investigational work have been conceded on the flow and pressure control loop outfit to examine the performance of the zn tuned pid, cc-pid and imc tuned pid controllers. to contrivance the closed loop control purpose, a pid block is established in the centum cs 3000 in which feedback regulating signal is programmed to pass through the control valve to form closed loop feedback structure in fluid transport process plant. 6.1.simulation results in the simulation work, the recitations of different controllers are equated by setting the peak of maximum uncertainty (ms) value for comparison. it is noteworthy to ensure iae, tv and ms (integral absolute error, total variation of the input, uncertainty margin value) to be lesser, but for a best characterized tuned controller causes the existence of trade-off, which indicates a declination in iae infers an escalation in tv and ms (and vice versa) [25, 26]. in the mean of the imc and direct synthesis implied tuning mode τc is an changeable parameter and hence it can be altered it for the anticipated robustness (ms) level. in the simulation work of the current paper, recitals of various controllers are associated by setting the similar ms value for a fair-minded comparison (i.e. ms=1.26). the controller structure of the pid block for the fluid transportation system to monitor pressure and flow rate is developed using simulink in matlab platform is presented in figure 7 and 8. fig. 7. matlab platform based created simulink model for pressure control loop 152 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) fig. 8.matlab platform based created simulink model for flow control loop 6.1.1. performance analyses for pressure and flow control loop of process plant the performance of zn, cc and imc-pid controller are investigated at required setpoint ranges in order to normalize both pressure and flow rate of the fluid being transmitted. the closed-loop simulated transient responses for pressure and flow rates at the operational choice of 500 lph and 3kg/cm2 for the time interval of t = 0–50 s are unveiled on figures 9a and 9b. from the figures 9a and 9b, it is clear that the imc-pid controller is enforced to trail the fixed operating point of pressure and flow rate at short duration of time of about 17 seconds and 15.5 seconds and maintain the steady state with overshoot as compared to zn-pid and cc-pid control techniques which settle at about 34 seconds and 26.5 seconds for flow rate and 38.59 seconds and 47.987 seconds for pressure respectively. it confirms that imc-pid controller offers 16.72% of improved quality indices on comparing with zn tuned pid and cc-pid controllers. the presence of overshoot in the imc-pid controller performance ensures high rise time with quickly settling on its desired set point in order to improve the experimented operating condition of the fluid transport system. scada based remote monitoring and control of pressure & flow rate 153 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 9a .performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for flow control loop at a set point of 500 lph. fig 9b. performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for pressure control loop at a set point of 3kg/cm2. 6.1.2.setpoint tracking for pressure and flow control loop closed loop simulated transient responses attained at different operating points such as 0.7 and 1.6 kg/cm2 for pressure and 900 and 1400 lph for the flow rate to confirm the robustness of the imc-pid controller are shown in figure 10 and 11. the table 6 and table 7 reveal that the performance of imc-pid on comparing with zn tuned pid and cc tuned pid controllers affords superior performance with the similar settings for altered operating points. among the controller tuning rules, imc-pid sustains the trepidations in the model parameters and pay for the extreme reliable and robust response when the operating point diverges. figure 10 and 11 also ensures that imc-pid controller endows least possible error indices of 20.94% as compared with zn tuned pid and cc-pid controllers. this pragmatic imc-pid controller confirms the boosted performance meanwhile it has the talent to trajectory on faults online. 154 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) fig. 10. performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for pressure control loop at different setpoint tracking. table 6.pressure-performance measures of the controllers at different operating points operating points in kg/cm2 error and quality indices measures zn-pid cc-pid imc-pid 0.7 ise 0.00031 0.000435 0.001375 iae 0.0434 0.0405 0.037 itae 0.1077 0.0749 0.04332 tr (sec) 1.8 nil nil ts (sec) 39.9 47.9 15.5 %mp 36.214 nil nil 1.6 ise 0.00031 0.000441 0.001375 iae 0.04342 0.040529 0.037077 itae 0.10772 0.040529 0.037077 tr (sec) 1.76 nil nil ts (sec) 39.102 46.32 16.2 %mp 36.347 nil nil 3 ise 0.00034 0.000491 0.000253 iae 0.03590 0.024387 0.01444 itae 0.01447 0.013514 0.012359 tr (sec) 1.91 nil nil ts (sec) 38.59 47.987 15.37 %mp 37.102 nil nil scada based remote monitoring and control of pressure & flow rate 155 copyright ©2019 assa. adv. in systems science and appl. (2019) figure 11.performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for flow control loop at different setpoint tracking. table 7.flow rate-performance measures of the controllers at different operating points operating points in lph error and quality indices measures imc-pid cc-pid zn-pid 500 ise 0.000417 0.00044 0.000435 iae 0.02471 0.027021 0.028951 itae 0.03691 0.039539 0.047324 tr (sec) 2.1 6.3 8.99 ts (sec) 17 26.5 34 %mp 15.36 3.075 22.98 900 ise 0.01667 0.0176 0.001741 iae 0.04943 0.054043 0.057901 itae 0.02967 0.030249 0.030684 tr (sec) 2.3 6.41 9.14 ts (sec) 17 27.01 35.102 %mp 14.89 3.152 21.39 1400 ise 0.01887 0.000356 0.018872 iae 0.02409 0.02658 0.025112 itae 0.03895 0.042130 0.043897 tr (sec) 2 5.98 8.93 ts (sec) 16.5 26.12 34.89 %mp 14.57 2.936 21.134 6.1.3. disturbance rejection test for pressure and flow control loop the disturbance rejection performance is investigated at the operating point of pressure at 2.1kg/cm2 and flow rate at 1300 lph. a step disturbance is introduced into the process by way of increasing the pressure to 3kg/cm2 and flow rate to 1600 lph after the time period of 100 and 150 seconds as shown in figure 12a and 12b. as of the analysis result endorses the merit by pointing only imc-pid controller damp the disturbance in a smaller time period of 13 seconds with less undershoot of 11.33% as compared to the cc-pid and zn-pid pid controllers on pressure control loop. the application of imc-pid controller on flow control 156 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) loop divulges the sluggish disturbance rejection of zn tuned pid and cc-pid controllers where imc-pid tolerates the disruption with a short span of 21 seconds with enhanced quality indices. the error and quality indices of the output signal are applied to estimate the disturbance rejection presentation of the controllers for pressure and flow rates are given in table 8 and 9. fig. 12a.performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for pressure control loop after a disturbance at a set point of 2.1kg/cm2. fig. 12b.performance analyses of zn tuned pid, cc tuned pid, and imc tuned pid controllers for flow control loop after a disturbance at a set point of 1300 lph. scada based remote monitoring and control of pressure & flow rate 157 copyright ©2019 assa. adv. in systems science and appl. (2019) table 8.performance measures after a disturbance at the operating point of 2.1 kg/cm2 performance measures zn-pid cc-pid imc-pid ise 0.000453 0.000418 0.000389 iae 0.014475 0.013514 0.012359 itae 0.007482 0.008391 0.006453 tr (sec) 1.792 nil nil ts (sec) 39.02 48.13 16.19 %mp 40.127 nil nil table 9. performance measures after a disturbance at the operating point of 1300 lph error and quality indices measures imc tuned pid cc tuned pid zn tuned pid ise 0.039257 0.043302 0.042816 iae 0.026801 0.028617 0.027843 itae 0.0141 0.018 0.0268 tr (sec) 5.34 8.921 13.973 ts (sec) 24.137 31.856 52.965 %mp 19.357 3.954 33.045 hence by simulation results, imc-pid controller contributes optimal smallest settling time with least error integral value and highest robustness on comparing with zn-pid and cc-pid controllers and hence affords 34.18% enhanced performance in monitoring and regulating the pressure and flow rate variations through the control valve opening and closing to attain the preferred operating choice at the destination in the fluid transport system process plant. 6.2.real-time experimentation of local intelligence on fluid transport system the proposed local intelligence using imc-pid controller design performance is experimentally substantiated in real-time on lab-scale experimental set up of the fluid transport system. the control drawing for the process plant architecture is developed by creating two separate control loop blocks of pid controller for pressure and flow rate as shown in figure 13. the resultant tuning values of pid controller using imc technique which is confirmed through simulation is put on to the established local intelligence tune window using centum cs 3000 r4 as shown in figure 14. through this control drawing builder option, hardware configuration gets synchronized with the flexibility of software customized application to the process plant in order to provide monitoring and control capabilities through online at a remote location from the process plant. the process plant is starting to run by enabling the auto mode initiated by the local intelligence, when the fixed operating point for flow rate and its pressure is given in the corresponding monitoring particular field parameters faceplate present in the scada front end panel of the fluid transport system. after fixing the required operating range, the pump will be on track to run which is enabled by an operator remotely. the real-time successive data of both pressure and flow rate parameters of the fluid being transported is displayed continuously in local intelligence trend view window in pic100.pv/ fic100.pv tag tab which is present below the trend graph and these data can be exported to excel by exhausting local utility data box option. 158 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) fig. 13.control drawing for fluid transport system process plant fig. 14.developed local intelligence tune window of pid controller for the process plant. scada based remote monitoring and control of pressure & flow rate 159 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 15.real-time monitoring and control of pressure through local intelligence using an imc-pid controller. fig. 16.real-time monitoring and control of flow rate through local intelligence using an imc-pid controller. figure 15 and 16 reveals the performance of local intelligence using an imc-pid controller on remote surveillance and control of flow rate and its pressure on the implemented process plant. the setpoint of pressure and flow rate is given as 0.99kg/cm2 and 1200 lph respectively. based on the operating set point, the developed imc-pid controller running on the back end of the local intelligence scada adjusts the feedback signal going from the remote master control panel to the i/p converter incorporated with corresponding pressure control valve and flow control valve to regulate its opening and closing installed on the process plant control loops. the real-time experimentation discloses that when the percentage of control valve opening gets increased, the field parameters as like pressure and flow rate of the fluids passing through the pipelines get decreases and increases consistently. the developed local intelligence using 160 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) imc-pid controllers authorizes the enriched performance meanwhile it has the capability to trajectory of errors occurrences through online by making real-time process plant information easily accessible through the online applications by its connectivity and interoperability. 7. conclusion and future scope the development and investigation of local intelligence using an imc-pid controller on the real-time fluid transport process plant are carried out in this paper. the real-time experimental results show that this local intelligence using imc-pid controller acts with provisioning smart decision making on the real-time data of the field parameters in order to optimize the process plant with minimum losses by means of its affordability and the reliability to conduct measurements remotely and in real-time of monitoring and control of field parameters. the experimental validation emphasizes has been engineered to condensed downtime, upgraded system accessibility, enriched control consistency, and unremitting system access. i/o remotehub module with control server, operator and engineer stations handling data organization, and gateway utilities are disseminated on an ethernet linkage to confirm system veracity and welltimed data broadcast communication. the evaluation of this local intelligence ensures that condition based maintenance with the proactive event-driven computing paradigm for fully exploiting its capabilities by enabling proactive maintenance decisions ahead of time on the process plant, so that the operator can react pre-emptive prior to sort out occurrences of any failure on the process plant. regarding the future work, an internet of things (iot) based reliable monitoring and control with local intelligence using imc-pid controller modular architecture design along with scada/dcs which is outlined in this paper will be implied and investigated on this lab-scale fluid transport system process plant. it is evident that the proposed iot architecture will offer a promising and a novel infrastructure to remote monitoring and control in any industrial sectors. conflict of interest all the contributors in this research work have no clashes of attention to announce and broadcasting this article. acknowledgment this research work is carried out under the senior research fellowship received from csir (council for scientific and industrial research) with grant no.678/08(0001)2k18 emr-i. references 1. galloway, b. & hancke, g.p (2013) introduction to industrial control networks, ieee communications surveys & tutorials, 15(2), 860. 2. davis, p., sullivan, e., marlow, d. & marney, d. (2013) a selection framework for infrastructure condition monitoring technologies in water and wastewater networks. expert syst. appl. 40 (6) 1947– 1958. 3. el-darymli, k., f. khan & m.h. mohammed (2007), reliability modeling of wireless sensor networks for oil and gas pipelines monitoring, sensors and transducers journal 108 (7), 6–26. 4. kim, m. (2016). a quality model for evaluating iot applications. international journal of computer and electrical engineering, 8(1), 66. scada based remote monitoring and control of pressure & flow rate 161 copyright ©2019 assa. adv. in systems science and appl. (2019) 5. liu, x., liu h., wan, z., chen, t., & tian, k. (2015) application and study of the internet of things used in rural water conservancy project, journal of computational methods in sciences and engineering, 15(3), 477–488. 6. liu, z & kleiner, y. (2014) computational intelligence for urban infrastructure condition assessment: water transmission and distribution systems, ieee sensors journal, 14(12), 4122–4133,. 7. marlow, d., beale, d. & mashford j (2012), risk-based prioritization and its application to the inspection of valves in the water sector. reliable engineering system safety 100 (1) 67–74. 8. meribout m., (2011) a wireless sensor network based infrastructure for real-time and online pipeline inspection, ieee sensor journal, 11 (11), 2966–2972. 9. shamsuzzoha, m. (2015). a unified approach for proportional-integral-derivative controller design for time delay processes. korean journal of chemical engineering, 32(4), 583-596. 10. morari, m. & zafiriou,. e., robust process control. nj: prentice hall englewood cliffs, 1989. 11. mohammed, n. & jawhar, n. (2008) a fault tolerant wired/wireless sensor network architecture for monitoring pipeline infrastructures, in: 2nd international conference on sensor technologies and applications, 13 178–184. 12. priyanka, e.b. & maheswari, c. (2016) parameter monitoring and control during petrol transportation using plc based pid controller, journal of applied research and technology, 14 (5) 125-131. 13. priyanka, e.b, maheswari, c. & thangavel, s. (2018), remote monitoring and control of an oil pipeline transportation system using a fuzzy-pid controller, flow measurement and instrumentation, 62 144-151. 14. priyanka, e.b., maheswari, c. & thangavel, s., (2018) iot based field parameters monitoring and control in press shop assembly. internet of things,(3),1-11. 15. priyanka, e.b., maheswari, c. & thangavel, s., (2019). remote monitoring and control of lqr-pi controller parameters for an oil pipeline transport system. proceedings of the institution of mechanical engineers, part i: journal of systems and control engineering, p.0959651818803183. 16. priyanka e.b., maheswari c & thangavel s. proactive decision making based iot framework for an oil pipeline transportation system. in international conference on computer networks, big data and iot 2018 dec 19 (pp. 108-119). springer, cham. 17. priyanka, e.b., krishnamurthy, k. & maheswari, c., 2016, november. remote monitoring and control of pressure and flow in oil pipelines transport system using plc based controller. in 2016 online international conference on green engineering and technologies (ic-get) (pp. 1-6). ieee. 18. subramaniam, t. & bhaskaran, p. (2019). local intelligence for remote surveillance and control of flow in fluid transportation system. advances in modelling and analysis c, vol. 74, no. 1, pp. 15-21. https://doi.org/10.18280/ama_c.740102. https://doi.org/10.18280/ama_c.740102 162 e.b.priyanka, c.maheswari, m.ponnibala, s.thangavel copyright ©2019 assa adv. in systems science and appl. (2019) 19. priyanka, e.b., maheswari, c. & thangavel, s., 2017, hmi-plc automation for pressure and flow control in oil pipelines, isbn: 978-3-659-97459-5, lap lambert academic publishing, germany. 20. priyanka e.b., maheswari, c. & thangavel, s. a mini reviewinvestigation and study of risks in oil pipeline construction substations. trends in civil engineering and its architecture 3(4)-2019. tceia.ms.id.000170. doi: 10.32474/tceia.2019.03.000170. 21. seborg, d.e., edgar, t.f. & mellichamp, d.a., process dynamics and control, 2nd edition, john wiley & sons, new york, usa, 2004. 22. shamsuzzoha, m. & lee, m. (2008) analytical design of enhanced pid·filter controller for integrating and fist order unstable processes with time delay chemical engineering science, 63(2008) 2717-273. 23. shamsuzzoha, m., skliar, m & lee, m. (2012) design of imc filter for pid control strategy of open-loop unstable processes with time delay, asia-pacific journal of chemical engineering, 7, 93-97. 24. tan, w., marquez, h. & chen, t. (2003) imc design for unstable processes with time delays, journal of process control, 13, 203-213. 25. yang, s.h., chen x., l. yang, b. chao, & cao, j. (2015) a case study of the internet of things: a wireless household water consumption monitoring system, internet of things (wf-iot), 2015 ieee 2nd world forum on. ieee, 681–686. 26. yang, x. p., wang, q. g., hang, c. c., & lin, c. (2002). imc-based control system design for unstable processes. industrial & engineering chemistry research, 41(17), 4288-4294. word template for assa manuscript adv syst sci appl 2019; 02; 1-7 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/758 technology of quality and reliability of complex technical systems characteristics stage-by-stage improvement on examples of rocket-space technology objects v.k. zawadzkii1* , v.p. ivanov1 , e.b. kablova1 , l.g. klenovaya1 1) v.a. trapeznikov institute of control sciences russian academy of sciences moscow, russia e-mail: vladguc@ipu.rssi.ru received june 20, 2019; revised july 1, 2019; published july 10, 2019 abstract: the principles of the evolutionary process of creating complex technical systems are formulated. the possibility of using the technology of stage-by-stage manufacturing of each specific sample of the system is considered. in the process of stage-by-stage manufacturing, the system improves its properties and characteristics, moving from relatively simple versions to more complex and effective versions. based on the results of the control, the system adapts to the real operating conditions. keywords: reliability, improvement of quality characteristics, rocket and space technology, complex technical system 1. introduction when improving the quality and reliability characteristics of complex technical systems (cts), which include modern objects of rocket and technical equipment (rte), problems have arisen that affect the conceptual basis of technology of their creation [1]. these systems are characterized by a multi-level hierarchical structure, a variety of "behaviors", a high dimension of the state space and a large number of elements, the ability to interact with the environment (adapt to and effect on the environment). for such objects, the incompleteness of physico-chemical processes knowledge underlying in the basis of their functioning, manufacturing technology and application conditions is characteristic. for the method of manufacture based on the decomposition of the product into the simplest elements, and production on the sequence of the simplest technological operations, these factors serve as the causes of design errors, production defects and violations in the operational procedures. in connection with the mentioned factors, the creation of complex systems is not a onestage act of manufacturing the finally chosen variants according to previously known technologies, but occurs through trial and error on simpler product prototypes [1, 2]. by experimental development of each new version, the missing information is filled and the correctness of design, technological and other innovations is confirmed. improvement of the technical characteristics of each new version of the system involves the use of new more effective principles of the system elements operation and a deeper understanding of the physical processes that underlie its functioning. as a result of a number * corresponding author: vladguc@ipu.rssi.ru mailto:vladguc@ipu.rssi.ru 2 v.k. zawadzkii , v.p. ivanov , e.b. kablova , l.g. klenovaya copyright ©2019 assa. adv. in systems science and appl. (2019) of iterations, the system approaches a certain limiting level of quality characteristics achievable within the product of this type. an example is the on-board control systems of launcher rockets, which in their development went from the forms of construction based on disintegration to the simplest subsystems and the use of the element base in the 1950s. to modern integrated control systems [3, 4], which include high-performance on-board computational facilities. let us stop on the fundamental principles of the stage-by-stage technology of creating cts: at all stages cts should retain the basic properties inherent in products of this type; transition to a new, more sophisticated product modification is associated with innovations; new cts maintains continuity with old, prior product samples; new product is created on the basis of a workable, reliable prototype; each product with innovations passes the stage of development test. the principles of stage-by-stage (additive) manufacturing are the basis of the additive manufacturing technology. this technology is associated with the use of three-dimensional printing and is considered by several authors as one of the important directions of the unfolding new production revolution. the most obvious possibility is the technology of stage-by-stage manufacturing in the production of dynamic systems and software (control algorithms) for these systems. elements of formalization of this technology within the life cycle of a dynamic system were considered in [5]. the improvement of the characteristics of the system was linked to the possibility of implementing a wider range of movements [6]. the existing technology of creating a system's software (control algorithm) resembles the assembly of a technical product from individual parts. with respect to the algorithm, the assembling elements are computational operations and the data of aprior and current information used in them. all the disadvantages of this method of manufacturing indicated above are also present in this case. a new, additive technology involves a stage-by-stage improvement of the properties and capabilities of the algorithm due to the sequential complication of the synthesis problem. the experiment in this case is simulation modeling. 2. formulation of a steady-state control problem let us consider the problem of complex multiply connected objects controlling, which include liquid-propellant launcher rockets and booster blocks. the control object is described by a differential equation of the form: )),(),(( ttwtxfx  , k extx ⊂∈)( ,  ew , r eutu ⊂∈)( , r , (2.1) ], 0 [∈ ttt , 0 ∈ 0 ) 0 ( xxtx  . there x(t) is the combined vector of the state coordinates of the object and the equations of the perturbation factors model; )( 0 tx extended vector of unknown initial conditions; f known vector function, differentiable with respect to the set of its arguments; )(tw control, which is a known vector function of time t; t the moment of achievement of the required final state (terminal time moment). for the final state of the object, boundary conditions are designated:  з)),((  ttx , l,1 , kl ≤ , lr  . (2.2) the value of the terminal time moment can be specified in advance. in general case, its magnitude is unknown in advance, and a certain degree of freedom in the assignment of the quality and reliability of complex technical systems 3 copyright ©2019 assa. adv. in systems science and appl. (2019) moment t is allowed. under perturbed conditions, the terminal time moment can be considered as a control parameter, chosen together with the parameters u from condition (2.2). for equation (2.1), we shall assume that the existence conditions and the uniqueness of the solution conditions are fulfilled. taking into account the possibility of control errors in solving problem (2.1)(2.2), we introduce the notion of boundary conditions discrepancies and reformulate the goal of control (2.2):  kllttxzz ≤ ,,1 ,з-)),((   , z , (2.3) where δ the set of admissible values of the vector of discrepancies, characterizing the accuracy of the solution of the problem (2.1)(2.2). in addition to the terminal conditions (2.3), we also consider indicators of integral type:  dwx t t ngni ),,,( 0 , nn ,...,2,1 , (2.4) where gn known function of fixed sign, the physical content of which can compose the cost of energy resource, time, loss due to control. to control the object, x(t) process measurements are made, defined by equation (2.1). the measuring devices are usually represented by static links. in the general case, the equation of the measuring part of the system for the control object (2.1) is written in the form: ))()(()( t,htx t =ty  , (2.5) where y(t) – dimension measurement vector p, h(t) measurement error random vector. equations (2.1) describe the processes of various physical nature occurring on a rocket carrier: the mechanics of rigid body motion, fluid hydrodynamics, heat exchange and thermodynamic processes, etc. the boundary conditions (2.1) and the discrepancies (2.3) are determined by the problems of launching the spacecraft and the conditions for the safe shutdown of the engine. criterion (2.4) is formed, based on the constant need to improve the energy characteristics of launching missile systems. the control problem formulated above should be solved taking into account the probabilistic nature of the perturbing factors, the relationship between the coordinates of the equations of various physical processes in the object, the diversity of the launching problems and the modes of operation of the control system. due to the complexity of the overall task, the object management is decomposed into specific tasks: navigation, guidance and stabilization, fuel consumption, tank pressurization, engine control. the mathematical support of each subsystem is formed from separate functional blocks and computational operations. at the same time, the understanding of the unity of the management of the object is lost, and the interconnections are not properly taken into account. when upgrading and improving the characteristics of the control system, developers constantly faced with the difficulties of making changes to the existing, proven technology of forming the software of the system. 3. decomposition of the synthesis problem by the degree of use of aprior and current information and definition of the main classes of terminal control it is advisable to consider the possibility of building a new technology based on the procedure for the stage-by-stage solution of partial problems of synthesis and continuity of these solutions. 4 v.k. zawadzkii , v.p. ivanov , e.b. kablova , l.g. klenovaya copyright ©2019 assa. adv. in systems science and appl. (2019) in this connection, let us consider possible ways of decomposition of the above terminal control problem into a sequence of simpler synthesis problems. as a result of solution of these problems, various classes of terminal control can be formed. by the nature and efficiency of aprior information and current information in the formation of physically realizable control, there are four main classes of terminal control: 1. program control (or open-loop control), where the formation of the control completely ignores the current information about the measured coordinates and realized controls. the program control is formed as a function of the time wпр(t) for the undisturbed motion of the object under given initial conditions and the terminal time t . under these conditions, the nominal (undisturbed) trajectory of the object xн(t) is realized. for the formation of software control, aprior information about the equations of the object and the specified boundary conditions for the nominal operating mode of the system is used. as a rule, optimal control is chosen as program control, which ensures the maximization of criterion (2.4), for example, on the basis of the maximum principle [7]. in accordance with this principle, a hamiltonian is constructed, which is the scalar multiplication of the vector of the right-hand sides of the equations of the object by the coordinate vector of the conjugate system. as a result of maximization, the control is defined as the vector function w(u,t) of the parameters u and the current time. parameters u are determined as a result of the solution of the main and conjugate system under given initial conditions. the initial conditions for the conjugate system must be linked to the terminal conditions of the terminal problem. as an example of a control program, it is possible to give an optimal law of variation of the pitch angle θ at given boundary conditions with respect to the position and velocity of the rocket's motion in a homogeneous plane-parallel gravitational field: tgθ = u1+ u2t. the law in the form w(u,t) turns out to be convenient for control with feedback. 2. control with feedback on the full vector of coordinates. when forming the control, the data of the current information arriving via the feedback loop is used. in the synthesis of control, a deterministic aprior description of the system is used. in this case, a considerably idealized formulation of the synthesis problem for an undisturbed object is considered. the deviations of the coordinates from the nominal trajectory and the discrepancy of the boundary conditions are caused by variable initial conditions, previously unknown before the time of measurement. the discrepancies are predicted for boundary conditions under a known vector of the current values of the object coordinates and given, for example, program control. the synthesis problem is to choose control with feedback in the form w(u,t). in the previous section, this vector-function was determined when solving the optimal control problem. the dynamics of the control process is characterized by a change in the vector of discrepancies of the boundary conditions. on that basis, it is advisable to choose a control from the lyapunov stability condition of the object's motion in the part of the vanishing of the predicted value of the residual vector at tt. stability of motion can be considered from the linear approximation of the object for deviations from the nominal trajectory: )()()()( twtbtxtax  , ], 0 [∈ ttt , (3.6) )()(),()( twtbtttxz  , з-)(),()( xtxtttxz  , where a(t), b(t) – matrix of order kk, k, (t,t)=x(t)x-1(t), x(t) – fundamental system of solutions. for the quadratic lyapunov function, the asymptotically stable motion and the condition zx(t)=0 are provided by the control of the following form: quality and reliability of complex technical systems 5 copyright ©2019 assa. adv. in systems science and appl. (2019) )(),( т )( т -)( tutttbtw  , )()( -1 )( txztctu  ,  dd t t tc )(∫)( , ),( т )( т )(),()( tttbtbtttd  . consider the case when the dimension of the control vector w(t) coincides with the dimension of the coordinate vector of the state x(t) (=k), and the matrix b(t) is such that the components of the control vector w(t) directly affect the derivatives of all coordinates of the state vector. as a rule, restrictions are imposed on the deviations of coordinates. in connection with this, for such deviations, it is possible to determine the values that are minimally needed for the solution of the terminal problem. the minimum required deviations are provided for a pulsed control action w() on a small time interval  in the t surroundings, for the remaining time interval w() =0: w()=w(t), tt+ , w() =0, >t+. here >0 is a small value in comparison with t. impulse control action on the object (3.6) from the condition 0) 1 (ˆ t x z for all t1: tt1>t+, is defined as follows:  /)(),( -1 -)( txztt b tw , )(),(),( tbtttt b  . such control is used in fuel consumption control systems for liquid-propellant launcher rockets. 3. deterministic control with feedback on the incomplete vector of object coordinates measurements. when measuring )()( thxty  , h – is a constant matrix of order kn , nk  , to recover x(t) let’s use the values of y()n a limited interval of prehistory ],-[ ttt  . as a criterion for the quality of recovery x(t) we take a definite-positive scalar function of the form:    t tt dyyqyyts ))(ˆ-)()(( т ))(ˆ-)((5,0)( , where q() diagonal matrix of positive coefficients. let )()-( 0 tttx x . at this )()-,()( 0 txttttx  . for )(ˆ 0 tx equation is specified: )()()()()( 000 ˆˆ tuttwttbttta xx  , ],[∈ 0 tt t , k etu ∈)(0 , (3.7) )( 0 ˆ∂ )(∂ )( 1 -)( 0 tx ts tktu  ,    dyyqhtt t tttx ts ))(ˆ-)()(( т )-,( т ∫ -)( 0 ˆ∂ )(∂ , where k1(t) diagonal matrix of algorithm parameters of order kk. algorithm (3.7) provides an asymptotically stable process of convergence )(ˆ 0 tx to )(0 tx . 4. stochastic control with feedback by a disturbed object. probability distribution laws and interference are assumed to be known. to solve the synthesis problem, the terminal quality control criterion is represented as a conditional mathematical expectation of the loss function, which depends on the discrepancies of the boundary conditions. for a wide class of problems, the synthesis of stochastic control is divided into independent estimation and control problems by estimating the extended coordinate vector of the object, including the parameters of disturbance models. to solve the estimation problem it is convenient to use the discrete analogue of the control object and the measurement channel in the form: i w i b i x i a i x  1, 6 v.k. zawadzkii , v.p. ivanov , e.b. kablova , l.g. klenovaya copyright ©2019 assa. adv. in systems science and appl. (2019) i=0,1,2,…,i+1, i hx i y  . as a result of solving a discrete equation, a linear function of the following form can be obtained: ) 1-,, ( jri w ri x jj x   , ) 1,..., ( 1-,j w ri w jri w   , irij ,...,1 , kr  . the estimation algorithm is presented in the following form: iri y i l iri x ri a iri x ,1-1-/ˆ 1-/ˆ      , (3.8) ),..., 1( ,1i y ri y iri y       , ) 1-,, 1-/ˆ(jri w iri x j h j y j y   , irij ,...,1 . the resulting estimation algorithm is a discrete analogue of the algorithm for reconstructing the full vector of state coordinates. in the case of an excessive number of measurements (r>k) with an appropriate choice of weighting matrix li in the algorithm (3.8) measurements error filtering is performed. conclusion in conclusion, we note the following. based on the analysis of a number of known, more complicated problems of synthesis, continuity of solutions has been revealed from program control to stochastic control with feedback. in the process of stage-by-stage improvement, approaching the real conditions of functioning, the possibilities of control are expanding, it acquires new useful properties. the carried out analysis can be used at creation of additive technology of control synthesis in complex technical systems. references 1. ivanov v.p. & portnov-sokolov iu. p. (1998) the issues of creating complex technical objects from the point of their safety (on the example of rocket and space technology objects), reliability and quality control, 1, 24–35. 2. martino d. (1977). technological forecasting. – m.: progress. 3. ivanov v.p., zavadskii v.k., guskov a.d., dishel v.d., vasyagina i.v., et.al. (2008). terminal control of launcher rocket and fuel consumption in the mode of its full development. international scientific and technical conference "systems and complexes of automatic control of aircrafts" dedicated to the 100th anniversary of the birth of academician of the russian academy of sciences nikolai alekseevich pilyugin. p. ii. proceedings of the plenary meeting (reports and announcements) mirea, april 23, 2008m.: “research and publishing center "engineer", 56-65. 4. zavadskii v.k., ivanov v.p., kablova e.b. & klenovaya l.g. (2013). integration of on-board control systems to improve the energy and reliability characteristics of quality and reliability of complex technical systems 7 copyright ©2019 assa. adv. in systems science and appl. (2019) launch vehicles, materials of iii all-russian scientific and technical conference "actual problems of rocket and space equipment": iii kozlov readings, september 16 – 20, 2013, samara) / samara scientific centre russian academy of sciences. – samara, 122–126. 5. ivanov v.p., kablova e.b. & klenovaya l.g. (2015). elements of formalization of the evolutionary process of improving the characteristics of complex dynamic systems, control problems, 4, 41-50. 6. krasnoshchekov p.s., fedorov v.v. & flerov iu.a (1997). elements of the mathematical theory of making design decisions, design automation, 1, 15–23. 7. pontriagin l.s., boltianskii v.g., garmkrelidze r.v. & mishchenko e.r. (1969) mathematical theory of optimal processes. – m.: nauka. advances in systems science and applications (2013) vol.13 no.4 317-333 latvia’s participation in the exchange rate mechanism of the european monetary system valery i. roldugin baltic international academy, latvia abstract the first objective is to investigate the on-going monetary policy of the bank of latvia, to analyse the basic principles of its operations and influence on national economic growth. the second objective of this article is to follow the main trends in the development in the latvian’s monetary policy and the european monetary system accession process, with a focus on the local currency stability problems. it discusses the process and strategies for choice of the strategy as well as the main issues that have arisen in the accession process. both objectives fully corresponding to the article’s research object, i.e. to a monetary policy of the bank of latvia. regarding the developed countries the general monetary policy objectives deals not only with maintenance of stability of the exchange rate and general price level, but also with stimulation of economic development, growth of employment and incomes of the citizens. the period from 2004 till 2011 is being investigated. the author uses a wide range of research methods, such as : grouping method, method of comparison of financial ratios, etc. keywords central bank, economic growth, monetary aggregate, monetary instruments, monetary policy. 1 introduction the monetary policy instruments of the bank of latvia are already in line with those used in the euro area. like the european central bank, the bank of latvia also uses the reserve requirements, market operations, as well as standing facilities of lending and deposit of funds. assets of bank of latvia including gold and exchange currency reserves, serve as maintenance of money issue in latvia. external reserves of bank of latvia which include gold reserves and foreign currency, and also currency from basket sdr in the end of 2012 has reached 7366,1 mill. lats (in 2003-830,5 mill. lats). the choice of the strategy depends on the size of country, level of openness, economic growth, relations between the key objectives and the intermediate target, features of financial and capital markets and other economic factors. the key objective of the central bank’s monetary policy is to facilitate the favourable macroeconomic environment for growth of the national economy in the long term. the course of global economic development assumes that monetary policy, employment and financial stability can foster the economic growth most of all by ensuring the low inflation rate. economic growth has stopped and even decreased in latvia, as well as baltic states and euro area in general thus restricting pos318 valery i. roldugin: latvia’s participation in the exchange rate mechanism... sibilities of profitable transactions carried out by latvian economic basic unitsenterprises in both domestic and export markets. rapid economic growth in recent years in latvia was mostly based on private consumption increase and large credit resource injections mainly in activities related to real estate market development. 2008 marked a turning point in latvia’s economic development, which began to decelerate after several years of buoyant growth. insufficient improvement in performance and efficiency of national economy, public administration and public services structure reduced the overall economic competitiveness, which was particularly influenced by recession in the export market. first signs of a slowdown of the economic growth became apparent in the second half of the year, when the implementation of the anti-inflation plan produced by the government, moderation of funding from parent bank and tightening of the banks’ lending policies resulted in a rather abrupt deceleration of the domestic lending growth. the reasons of failures of monetary policy include the methods of monetary regulation used by the bank of latvia. there is a danger of traps in the course of application of the monetary policy instruments. it is necessary that the bank of latvia should carefully chose the necessary financial tools of monetary policy. the author offers the ways of the solving of these problems on the basis of main principles using the achievements of world economic thought. in the conditions of a financial crisis the bank of latvia is able to achieve its targets while using the various financial instruments. 2 preparation for the participation in the european monetary union accession to the eu implies preparation for the participation in the economic and monetary union and introduction of the euro, for the new member states of the eu may not choose to stay outside the euro area. hence, after joining the eu, latvia will have to demonstrate its capability to meet the criteria for joining the economic and monetary union. the cabinet of ministers and responsible public administration institutions will take steps necessary to comply with maastricht convergence criteria to ensure euro changeover as soon as possible. in the analysis of any definition there is the first important question, the answer to which is connected to all systems definitions and classifications. this question will be worded as follows : “does a system really exist ?” not many researchers ask this question, but those who are wondering, as a rule, respond to it positively. it seems to us that the question canąŕt be answered at all. why ? let’s think constructively. economic convergence is a serious pre-requisite for successful operation of a monetary union ; therefore, prior to joining the euro area the eu member states have to ensure compliance with the maastricht convergence criteria : achieve price advances in systems science and applications (2013) vol.13 no.4 319 stability, sustainability of the government’s financial position, exchange rate stability against the euro and long-term interest rate convergence. maastricht convergence criteria : price stability is measured as the average consumer price increase of the last 12 months in three eu member states with the lowest inflation plus 1.5 percentage points ; sustainability of the government’s financial position means that the government deficit and the government debt do not exceed 3% and 60% of gdp, respectively ; long-term interest rates convergence criterion is interpreted as the average yield rate of 10-year government bonds within a 12 months period in three countries with the lowest inflation plus 2 percentage points ; exchange rate stability criterion provides for participation in the exchange rate mechanism ii (erm ii) without severe tensions, which, in turn, primarily means without currency devaluation against the euro. if required, other adequate indicators may be used to assess stability. in order to achieve the key objective as well as join the economic and monetary union successfully, the bank of latvia has chosen its exchange rate strategy for the implementation of monetary policy. now latvia almost complies with the majority of the maastricht convergence criteria necessary for joining the economic and monetary union (except for a parameter of interest rates on government long-term securities). table 1 the compliance with the maastricht criteria [1-4] criterion latvia’s criterion in latvia’s criterion latvia’s in performance in performance in performance 2005 in 2005 2005 in 2008 2011 in 2011 budget deficit -3.0 -1.0 -3.0 -4.0 -3.0 -4.1(% of gdp) government 60.0 14.7 60.0 19.5 60.0 42.6debt (% of gdp) average annu2.5 6.7 4.1 15.3 2.4 4.0al inflation rate (%) interest rates 5.4 3.96 6.24 6.43 7.80 6.00on government long-term securities(%) fixed fixed fixed fixed fixed fixed exchange exchange exchange exchange exchange exchange rate against rate against rate against rate against rate against rate against exchange the euro and the sdr the euro and the euro and the euro and the euro and rate regime participation since 1994 participation participation participation participation in erm ii in erm ii in erm ii in erm ii in erm ii for at least for at least for since for at least for since two years two years may 2005 two years may 2005 statistical data available for the last years indicate that the initial plan of latvia to adopt the euro in 2008 cannot be implemented due to the high inflation 320 valery i. roldugin: latvia’s participation in the exchange rate mechanism... rate. the schedule has not been revised as yet but according to the information released by the ministry of finance, in 2007 the government would discuss a new target for the changeover to the euro, tentatively in 2012-2014. the introduction of the euro in latvia will be an issue of the eu multilateral relations affecting common interests of all eu countries. therefore, the projected time frame for the introduction of the euro is merely tentative and will gain an official status only after the completion of all negotiations and other formal procedures. the high rate of lat stimulated a wide import of consumer goods displacing the latvian producers from domestic market that brought to the decrease of internal accumulation. in 1936, the gold content of lat was 0.29032254 grams of gold, and its parity to dollar equaled usd/lvl 3.03. the rate of lat at the riga exchange quoted on the same level. however, after adoption of law «on currency reform» on september 28, 1936 which provided the devaluation of lat on 40%, its rate at the riga exchange decreased to usd/lvl 5.04 [5]. the latvian exporters returned what they have lost due to the growth of lat rate before the currency reform, i.e. the same 40%. considering the indicators of national economy of latvia after lat devaluation in 1936, we see that it promoted the improving of economic situation. in 1937, in comparison with 1935 the amount of bank deposits increased on 67%, employment grew on 18%, a number of the unemployed people was reduced by 50%. the u.s. scientists carried out a study which stated that the fixed rate would bring latvian economy to such a decline, which’s incomparable with the u.s. recession during the great depression years of 1929-33 [6]. therefore, it is not casual that the rating agency moody’s and british analytical center capital economics point out the risk of devaluation of the lat. the news agency bloomberg referring to the experts of the u.s. bank brown brothers harriman & co. reported that the latvian currency may be devalued against the euro by 50% [7]. many western economists spoke about the necessity of devaluation of national latvian currency. kenneth rogoff, the professor of the harvard university and the former chief economist of the international monetary fund, considers that latvia should devaluate lat to avoid the toughening of economic crisis. it should be noted that the imf has no univalent position concerning devaluation of lat. according to prime minister valdis dombrovskis at discussions about financial aid to latvia, the imf wouldn’t object the devaluation of national currency but without insisting on it [8]. ilmars rimshevich, the president of the bank of latvia expressed opinion that lat devaluation would pull down the economy of latvia for one night. it is necessary to agree with him as it is necessary to treat devaluation of lat extremely carefully. the situation is similar to the tense spring. the devaluation must be based on exact calculations and shouldn’t be carried out without any damage to advances in systems science and applications (2013) vol.13 no.4 321 the population. there required the thorough economic calculations for search of an optimal rate which would correspond to the parity of purchasing power of lat to foreign currency more precisely than today. the price reform will demand many efforts because the standard of price expressed in a new rate of lat will change that should affect the size of salaries, pensions and other social payments, the state budget and its external balance of payments. the volume of export of the latvian industrial and agricultural output will increase. accordingly, the taxes and also the sums required for repayment of received earlier foreign loans will increase. all these points should be considered while starting such painful operation as devaluation of national currency. 3 problems of devaluation of lat the bank of latvia should be regarded as a state administrative body which carries out the main responsibility for development of events of the monetary sphere. its priority direction was always the external stabilization of the national currency which strengthening shall promote the internal sustainability as well. over the past 20 years, the latvian currency had a fixed exchange rate. since june 1993 lat had a floating rate, and in february 1994, the bank of latvia pegged the latąős exchange rate to the sdr basket (1 sdr=0.7997 lvl). on january 1, 2005 it was pegged to the euro (1 eur=0.702804 lvl). the fluctuations are admitted within +/-1% of the pegged rate. on may 1, 2004, when latvia has joined eu, the bank of latvia became a member of the european central banks system and began the preparations on euro adoption that was planned to take place in early 2008 according to the plans at those times. no big changes in the activities of the bank of latvia have happened. bank continued to perform the former economic functions, including monetary policy to ensure the operation of the payment system, the circulation of cash, preparation of financial statistics and balance of state payments and perform the macroeconomic analysis and research including the identification of trends in economic development in the scales of the european union. the administrative functions including licensing and control over the banks have been handed over to the financial and capital market commission in 2001. since may 1, 2004, the bank of latvia became the holder of the shares of the capital of the european central bank (ecb). the total share in the registered capital of the european central bank was established at 0.2978%, or 16,571.6 thousand euro. the share of the bank of latvia in the ecb’s capital was calculated on the basis of volume of the country’s economy and population. the bank of latvia brought 750 thousand lats and 1763 thousand lats into the fixed capitals of the european central bank and bank respectively for international payments. the bank of latvia uses the available foreign currency for the corre322 valery i. roldugin: latvia’s participation in the exchange rate mechanism... sponding installments. the bank of latvia aspired to direct the currency rate policy on achievement of low inflation and convergence of the prices for the goods and services and level of inflation similar the countries, which national currencies are included into sdr basket. latvia enjoys one of the most liberal currency regimes in the world. both the citizens and foreigners can open accounts in lats and in other currencies without restrictions and freely sell and buy lats by other currencies. the commercial banks credited in foreign currency not only the entities, but also the individuals. as a result, the real estate transactions have been carried out not in foreign currency rather than in national one. the same picture could be observed in the automobile market. the bank of latvia have been buying and selling the basket of sdr currencies on requests of commercial banks without restrictions. the primary long-term goal of approximately 65% of the world’s central banks is the prices stabilization ; the second most common goal is the sustainability of national currency. the small countries with open economies (like latvia) usually stick to the second long-term goal. it is impossible to achieve both goals at once because the central bank can control just one direction : either inflation or exchange rate. initially, the activity of the bank of latvia in the conditions of high inflation was directed on fight against inflation by the means of strict monetary policy. in 1992-1994, the monetary base ğő0 was the main indicator of bank of latvia because it was easier to supervise the monetary base, rather than other monetary aggregates in the conditions of high inflation (58.7% in 1992, 34.9% in 1993). the situation analysis after inflation decrease showed that the fiscal policy and monopolists’ price regulation incl. power industry, transport etc., became the determining factors. the internal assets of bank of latvia became the target parameter ; and the minimum limit of net external assets of the bank of latvia was set to maintain the stability of lat what was connected with an unofficial pegging of lat to sdr. the baltic states started their programs of sustainability of national currency in the middle of 1992. the economists had faced a problem when choosing the exchange rate mode : to accept the floating exchange rate or to operate the fixed rate. though the programs of sustainability of the baltic states differed from each other, however, all of them came nearly to the same result, i.e. estonia stuck the national currency to the fixed rate of dm, lithuania stuck to us dollar and latvia stuck to sdr. the main criteria were : to prevent the national economy from external shocks in a best possible way and to ensure the financial and economic sustainability of the restructuring of the economy. advances in systems science and applications (2013) vol.13 no.4 323 in 1992, estonia exchanged the soviet rubles at the market rate of 10 :1 and equated krone at a rate of dm/eek 8. lithuania replaced the intermediate currency in the form of coupon by lit in june, 1993, and pegged it to us dollar at a rate of usd/ltl 4 since april, 1994. it is interesting that in 1936 the rates of estonian krone, lithuanian lit and latvian lat against us dollar were following : usd 1/eem 6.95/ltl 4.38/lvl 5.04. in 1999, the rates of three baltic currencies to us dollar were respectively : usd 1/eek 11.2/ltl 4/lvl 0.59. there is a question how the current rate of lat against us dollar could be established almost ten times higher in comparison with 1936. on september 28, 1936 one british pound had been cost 25.11 lvl [5]. presently, the latvian lat is the main competitor for the pound. the latvian lat was heavier than pound, i.e. its rate was gbp/lvl 0.7910 as of march 1, 2010 [6]. so, the purchasing power of lat was approximately set on the same level as in 1936. let’s view the chronology of events related to the establishment of the new national currency rate in 1990-ies. it is known that latvia lost independence on july 17, 1940. the latvian lat was in circulation as a legal payment instrument till november 25, 1940 when the soviet ruble had been introduced in parallel at a rate of 1 lat=1 ruble. since march 25, 1941 all lat cash and deposits over 1000 lats (rubles) have been cancelled without any warning. such monetary confiscation has been carried out in three baltic republics. two currencies were functioned in latvia from 1941 to 1945. there were the german occupational money reichsmark (rm), and also the soviet rubles at a rate 10 rubles=1 reichshsmark. after the second world war till 1992 only soviet rubles were in circulation. only after latvia had regained its independence and the law “on the bank of latvia” had been adopted on may 12, 1992, there was an opportunity to carry out the monetary reforms targeted to establishing of a national currency circulation. to do this, the bank of latvia has specially established a committee for monetary reform (latvijas republikas augstākās padomes naudas reformas komiteja) that decided to introduce a temporary currency-the latvian ruble on may 2, 1992. the first stage of reform began on may 7, 1992, when the bank of latvia issued a temporary money the latvian ruble, which were in circulation along with the soviet rubles at a rate 1 latvian ruble = 1 ruble. since july 20, 1992 the parallel circulation was stopped, and the latvian ruble became the sole legal payment instrument. finally, the first five-banknotes came into circulation on march 5, 1993, which were used in parallel with the latvian rubles at a rate of 1 lat=200 latvian rubles. the latvian ruble was in parallel circulation with the lat till october 18, 1993. within 1992 there was a sharp shortage of cash in circulation in the state what 324 valery i. roldugin: latvia’s participation in the exchange rate mechanism... was threatening to cause a serious social crisis. so, if in february, 1992 the amount of cash exported from latvia exceeded the quantity of soviet rubles imported to latvia on 122 million banknotes (5.9%), in april of the same year this amount equaled to 686 million banknotes (29.2%) [10]. moreover, at the beginning of 1993 the latvian parliament discussed a case connected with unauthorized export of cash in amount of 665 million soviet rubles in railway car to russia. the commodity exchange carried out its operations in the building of the ministry of agriculture. commodity brokers have been changing the meat to paper, shoes to cereal, etc. the employees have been paid often by the enterprises products. in november 1993, there were 66 registered banks. the annual interest rate for deposits was as follows : banka baltija-90%, latvijas depozitu banka-45%, latvijas industriala banka-60%. now it is difficult to apprehend, but the liter of milk cost 14 centimes at the end of 1992, and the us dollar buying rate fixed by the bank of latvia was 0.8350 lvl and 0.5919 lvl already at the end of 1993. for january 1, 1995 the official rate of us dollar was fixed by the bank of latvia on a level 0.5480 lvl. as we can see, the bank of latvia took as a basis the rate of soviet ruble while establishing a rate of lat. therefore, on a moment national currency introduction, its rate initially was artificially overestimated. the high rate of lat allowed supporting one of the highest rates of economic growth among the countries of eastern and central europe (see tab. 2). the nominal gdp has been growing on more than 20% a year in pre-crisis years. for example, its growth was 23.3% in 2006, and 32.3% in 2007. 4 trends of latvia’s monetary policy the major quantity indicator of monetary circulation is the money supply representing a total purchasing volume and legal tenders, which serve for economic circulation of financial recourses of private persons, enterprises and state. the analysis of structure and dynamics of money supply has a great significance for development of reference points of the monetary policy of central bank. the monetary aggregates are applied to money supply definition. the cash assets are the basis of all monetary aggregates. the share of cash in money offer can vary depending on the circumstances. so, the more economy of the state is developed, the more banking system is stable and wide and the fewer shares of cash. the bank of latvia applies the following main monetary aggregates to the characteristic of monetary system and banking : (m0) monetary base monetary base is calculated on the basis of the bank of latvia’s methodology and comprises the lats banknotes and coins issued by the bank of latvia and demand deposits of resident mfis and financial institutions (overnight deposits) with the bank of latvia. the central bank directly creates advances in systems science and applications (2013) vol.13 no.4 325 this part of the monetary offer. by increasing the assets the central bank creates the money of high efficiency, by reducing it destroys them. the opportunities of central bank towards the creation of money of high efficiency are extremely great, as its liabilities (passives) are the money itself. m2x (wide money) broad money aggregates are calculated on the basis of the bank of latviaąŕs methodology and comprise the lats banknotes and coins issued by the bank of latvia (less vault cash of mfis) and overnight deposits and time deposits in lats and foreign currencies (including deposits redeemable at notice and repurchase agreements) held with mfis by resident non-financial corporations, financial institutions, households and non-profit institutions serving the households. m2x incorporates the deposits placed by local governments as a net position on the demand side. this parameter characterizes the total money in national economy. as foreign currency is of great significance for latvian economics, this parameter includes both the deposits of the enterprises and private persons in foreign currency [11]. let’s review the monetary parameters presented in table 2[1-4]. table 2 the basic indicators of monetary policy in latvia from 2004 till 2011(mill. lats) 2004 2005 2006 2007 2008 2009 2010 2011 gdp (current 7434.5 9059.1 11171.6 14779.8 16243.2 13070.6 12784.1 14275.2prices) growth gdp 116.3 121.0 123.3 132.3 109.9 80.5 97.8 111.7(%) m0 957.2 1350.7 2248.8 2471.2 2111.5 1645.8 1755.2 2168.9 m2x 2816.5 3905.8 5479.9 6171.3 5931.4 5796.2 6390.0 6486.1 growth m2x 127.0 138.7 140.3 112.6 96.1 97.7 110.2 101.5(%) m2x/m0 294.2 289.2 243.6 249.7 280.9 352.2 364.1 299.1(%) gdp/m2x 2.7 2.3 2.0 2.3 2.7 2.3 2.0 2.2 the velocity of money testifies the communication between monetary circulation and processes of economic development. aggravation of the macroeconomic risks and lower savings induced acceleration of the velocity of money, growing from 2.3 in 2007 to 2.7 in 2008. resident financial institution, non-financial corporation and household deposits with mfis decreased by 205.9 million lats or 3.9% in 2008 in comparison with an increase of 16.9% in 2007. in 2004, the velocity of money made 2,7 times per year. the economic situation stimulated the decrease of rate of the money turnover from 2.7 in 2008 till 2.2 in 2011. in relation to this parameter 326 valery i. roldugin: latvia’s participation in the exchange rate mechanism... latvia approaches to the developed countries where the speed of money turnover does not exceed 1.5 times. the decrease of rate of turnover of monetary volume (in 1.5 times) for the last 8 years cannot lower the negative influence of prompt monetary growth which increases more than by 40 % in 2006. the behaviour of the monetary aggregates in 2008 mirrored the sharp downturn of the economic development with both domestic and external demand shrinking, as well as the impact of the global financial crisis on the latvian banking system and money market. in 2009 m2x decreased by 2.3% (a growth of 10.2% in 2010) and amounted to 5796.2 million lats at the end of 2009 (see table 2). with the economic development coming to a halt in the second half of the 2008, banks cutting down on their lending business remarkably and confidence with regard to the financial sector deteriorating. the negative rate of the monetary expansion was primarily a result of the decelerating growth of mfi loans to the private sector, with the total loans outstanding shrinking in the last three year of the 2000s. the monetary situation in latvia is characterized also by other parameters : monetary multiplicator (m2x/m0) and velocity of money (nominal gdp/m2x). the great importance has monetary multiplicator a parameter describing the opportunities of economy as a whole and banking system in particular to increase a money stock in a turnover. its size pays off as the attitude of m2x to monetary base(m0). the monetary multiplicator is necessary to control over the monetary volume dynamics and the rate of inflation in latvia. in 2011, the monetary multiplicator has decreaseed essentially, and by the end of the year it consisted 299.1% (in 2010-364.1%). monetary base m0 decreased by 14.6% in 2008 and totalled 2111.5 million lats at the end of the year, whereas the cash component of the monetary base grew to 48.2% in comparison with 42.5% at the end of 2007. deposits from credit institutions and other financial institutions held by the bank of latvia declined by 328.3 million lats or 23.1% in 2008 as opposed to a 21.0% increase in 2007. for the second consecutive year, the demand for cash decreased, and currency in circulation shrank by 31.4 million lats or 3.0% (by 2.3% in 2007). the development trends of monetary aggregates were influenced by drying up capital inflows and foreign exchange interventions of the central bank reflected in the changes of net external assets. negative net foreign assets of mfis grew by 31.9% during the year, amounting to 5914.6 million lats, whereas the respective indicator of the bank of latvia decreased by 16.0% and totalled 2332.3 million lats at the end of 2008. thus the growth of the negative net foreign asset position decelerated considerably in comparison with 2007, when it expanded 1.7 times. nevertheless, latvia experience significant economic growth beginning in 2004 with eu accession. this followed the u.s. strategy to prevent an economic recession through asset inflation following the collapse of its stock markets in 2000. advances in systems science and applications (2013) vol.13 no.4 327 the usa and other developed countries flooded a world economy cheap credit resources. thus, developed countries created credit found its way into latvia through swedish banks. combine with eu structural funds, the latvian economy headily increased of gdp, until the inevitable global economy crisis led to latvia’s disastrous fall. but, latvia’s economy was not purely the victim of induced events. the bank of latvia and other state regulators are also responsible. the basic problem of latvia for the last 20 years was a balance of payments deficit of current transactions. balance of payments deficit arises as a result of excess of import over export. within last 8 years the volume of foreign trade of latvia has considerably increased in 3 times. the volume of foreign trade has been increasing year from a year, except 2009 when the consequences of the financial crisis have occurred (table 3) [1-4]. table 3 parameters of foreign trade of latvia from 2004 till 2011(mill. lats) 2004 2005 2006 2007 2008 2009 2010 2011 import 2989.2 4866.9 6378.5 7780.2 7484.4 4709.8 5911.9 7719.4 export 1650.6 2888.2 3293.2 4040.3 4406.0 3602.0 4695.0 5998.5 balance -1338.6 -1978.7 -3085.3 -3739.9 -3078.4 -1107.8 -1216.9 -1720.9 in 2008, the foreign trade dynamics was affected by the global financial crisis and weakening domestic and external demand. exports of goods expanded by 9.1% (a 15.1% rise in the first nine months of the year and 7.2% drop in the fourth quarter). imports of goods continued to shrink gradually, recording a 3.8% downslide on an annual basis. the excess of imports of goods over exports of goods decreased to 69.9% (92.6% in 2007) and the foreign trade deficit narrowed by 17.7%. the role of the eu in latvia’s foreign trade is strengthening : 76.1% of total export from latvia went to the eu countries in 2011 (73.1% in 2008). let’s consider the main financial instruments and principles of monetary policy in latvia. the choice of monetary instruments is wide enough. the basic tools of monetary regulation are : the standing facilities of lending and deposit of funds (refinancing) ; the reserve requirements ; the market operations. the choice and combination of instruments of monetary regulation depends, first of all, on issues, which are settled by central bank at the stage of economic development. the regulation of discount rate relates to market instruments of monetary regulation. the mechanism of regulation is simple enough ; therefore it is widely used in developed and developing countries. the official discount rate is a reference point for other market rates. the above level of official discount rate is the above cost of central bank refinancing. it means that the policy of change of discount rate represents a variant of regulation of qualitative parameter of money market the cost of bank credits. 328 valery i. roldugin: latvia’s participation in the exchange rate mechanism... the refinancing of commercial banks is carried out by holding of credit auctions, granting of lombard loans, etc. the bank of latvia has started using of refinancing as monetary policy instrument in 1993 only by granting the short-term credits to commercial banks for liquidity maintenance. originally the credits were granted to each bank within the limits depending on the bank’s performance in accordance with regulatory requirements established by the bank of latvia. such an order was necessary because the credits were granted without collateral. since november 1993, when the demand for credits exceeded the supply of credit resources, the bank of latvia has started carrying out of the credit auctions. as earlier, the credits were granted without collateral, therefore the quantity of the participants of auction was defined depending on the size of the equity capital of the bank, liquidity, and the bank performance in accordance with regulations of the bank of latvia. the bank of latvia has started granting the lombard loans in september 1995. it is the form of refinancing when the central bank grants the credit under pledge. the bank of latvia grants two types of lombard loans : the automatic lombard loans and lombard loans on demand. the commercial banks can exceed the balance of the correspondent account within the limit of the lombard loan during a payday for maintenance of efficiency of the interbank payment system. the bank of latvia grants the credits to commercial bank for one day in amount of the debit balance of corresponding account, in case the commercial bank could not involve the money resources till the end of a payday to liquidate the lack of resources on a corresponding account. the bank of latvia grants the lombard loans to commercial banks automatically in the end of a payday without special requirement. the basis for the lombard loan issuance is the special agreement between the bank of latvia and commercial bank. usually the lombard loan interest rate is higher than credits repo interest rate. it is a kind of “penalty" for the usage of resources of the central bank. the interest rate of lombard loans can vary depending on the terms of drawdown. as the lombard loans interest rate is higher than the interbank market interest rates and the interest rate of refinancing of the bank of latvia, the demand for these credits is usually small. the commercial banks use the lombard loans only in case of emergency. the bank of latvia supports the money volume in the set parameters and adjusts the level of liquidity of commercial banks by changing the minimal reserve requirements. it is assumed that credit institutions have to hold a certain share (currently 3.5%) of the attracted non-bank deposits with the bank of latvia. in event the reserve requirements are increased, these credit institutions will have to hold more funds with the central bank. it means that the amount of funds attracted by credit institutions, which is at their disposal and could be freely placed in the economics, thus increasing the level of credit and broad money, will advances in systems science and applications (2013) vol.13 no.4 329 decrease. the reserve requirements as a monetary policy instrument ensure the higher stability in the monetary base demand and facilitate the effectiveness of market operations, preventing the excessive interbank interest rate daily fluctuations. 5 three dogmatic “rule” there is a dogmatic “rule” of the monetary policy : the actual rate of refinancing should be positive (though it was negative in many developed countries in specific years, and in 2007 the refinancing rate was below the rates of inflation in the majority of these countries). in latvia, the monetary market rate is below the inflation rates. it is obvious there is an imported inflation in the conditions of prices rise which is underestimated by the central bank. accordingly, it overestimates the measurements of the base inflation depending only on the monetary factors. therefore, the refinancing rate should be minimal, i.e. at a level of correctly estimated base inflation. in the first half of 2007, the bank of latvia continued to pursue the tough monetary policy and in two occasions raised the refinancing rate by a total of 100 basis points (to 6.0%), thereby dampening the excessive domestic demand. later, when the signs of economic overheating abated, the bank of latvia left its interest rates unchanged, but in the first half of 2008 due to the slowdown in the growth of lending and the associated deterioration in the banks’ role in fuelling the domestic demand, it lowered the minimum reserve requirement for bank liabilities with agreed maturity of over 2 years by 2.0 percentage points (to 6.0%). the reasons of failures of monetary policy include the methods of monetary regulation used by the bank of latvia. they are reduced by application of several “rules” which are considered to be suitable for any condition in any country. the first “rule” : in order to decrease the rates of inflation it is necessary to limit the monetary offer or to apply the quantitative credit restrictions or to overestimate the refinancing rate. one more dogmatic “rule”-in order to decrease the inflation it is necessary to strengthen the rate of national currency ; on other hand, it reduces the price competitiveness of national commodity producers. it could be raised by depreciation of the credits, but it is forbidden by the first “rule”. the application of such “rule” of monetary policy by no means is rather offensive : the economics can get to a condition called by the “trap”. in such conditions the measures of state regulation do not bring positive results. these traps are known : “the debt”, “the liquidities” and “the negative effect rendered by strengthening of actual exchange rate of national currency on economic development”. the debt trap occurs at excessive debts of the state and private sector at small duration of debt. by involving the short-term and intermediate term loans under 330 valery i. roldugin: latvia’s participation in the exchange rate mechanism... low interests the borrowers count them as refinance by new loans. in case of steep increase of interest rates in the financial markets its long service sharply rises. if the borrower is not in position to extinguish it, the avalanche growth of debts begins, and it is impossible to extinguish it even after decrease of the interest rates. now, there is a growth of interest rates which were on rather low level for the long period. the latvian enterprises borrowed the financial resources abroad and can shortly face the refinancing problem and service of debts. the liquidity trap occurs at too low nominal interest rates when the central bank decreases refinancing rate, but it do not lead to expansion of the credits and stimulation of economic growth. the third trap is caused by the increase in balance of the international payments. it leads to the strengthening of national currency and shifting of employment to sphere of services. the bank of latvia can use the maintenance of the exchange rate by buying up the foreign currency, and it conducts to the growth of monetary offer and inflation strengthening. at high inflation the strengthening or stability of the nominal exchange rate leads to strengthening of the actual exchange rate. in other words, the efforts of the central bank do not achieve the object on maintenance of price competitiveness. thus, there is a necessity of sterilisation of superfluous liquidity. the escalating of the gold and exchange currency reserves by latvia basically is crediting of the usa and the european union countries. the inflation is mostly supported by excessive inflows of a variety of financing : the credit resources from foreign parent banks, the foreign direct investment, the eu funding, the workers’ remittances from abroad. to avoid the imported inflation the high validity of monetary policy is required. for example, the most effective remedy of the negative effect rendered by strengthening of actual exchange rate is crediting of the enterprises by the government and the central bank by replacing the foreign loans. certainly, such replacement of credits is better than the accumulation of superfluous currency provisions. however, from the macroeconomic point of view the replacement of foreign loans of the latvian enterprises by local credits is similar to repayment of external debt. now, the latvian commercial banks involve the foreign loans, and the central bank is compelled to buy up the foreign currency, generating the superfluous monetary offer which needs to be sterilised immediately. with the view of restriction of money growth the bank of latvia does not refinance sufficiently the credit organisations by establishing the refinancing rate at high level. it leads to the overestimate of credit resources cost for banks and enterprises. therefore, the latvian enterprises increase the foreign loans. the vicious circle turns out-the bank of latvia is compelled to get the additional volumes of the foreign currency arriving in the form of foreign credits to private sector advances in systems science and applications (2013) vol.13 no.4 331 of economy, increasing thereby the monetary base. the government is compelled to “freeze” huge budgetary funds on accounts of the bank of latvia in order not to admit an excessive monetary issue and inflation strengthening. the failure includes the preservation of high average rates of inflation which exceed 15.4% in 2008. the ways of their decrease in conformity with anti-inflationary program developed by the government are not enough clear. at low technological efficiency of the majority of latvian economic branches and growing world prices for energy sources it is impossible to stop inflation, including the attempts to limit the growth of monetary weight. however, the high inflation at stable or raising exchange rate of lat leads to the strengthening of the actual rate of exchange, and consequently-to corresponding decrease of competitiveness of the latvian enterprises. though the indicator of parity of purchasing capacity is considerably underestimated regarding the currencies of the developed countries, the actual exchange rate strengthening brakes the economic growth and promotes the advancing growth of import. at high inflation the investors of banks receive negative real percent on deposits, and the enterprises pay the overestimated income and added cost taxes. 6 conclusions for the past years the modern two-level banking system was established and developed in the country. gradually, there formed the competitive credit and financial infrastructure which basic elements were the commercial banks. some of them have already got a high international rating. the association of the latvian banks turned into national bank association. at the same time, there is an unpredictability of actions of the latvian administrative bodies. a main goal of the bank of latvia and the financial and capital market commission is the ensuring of the general sustainability in the monetary and credit markets. thus, exercising all rights defined by the law, the bank of latvia should provide a stability of the prices, and the financial and capital market commission should treat everyone who forms the instability in finances and capital market in whole or in its separate sectors by their activity or non-activity. the results of research testify the mistakes of the central bank made at the moment of establishment of national currency rate. the country had a possibility to provide unreasonably high rates of gdp growth due to the overestimated currency especially in the middle of the 2000th. in order to solve these problems it is necessary to boost the access of latvia into euro zone since the adaption of euro will help the bank of latvia not to carry out a number of aforementioned actions. an effective banking system is one of the most important conditions of economic development of latvia. it is assumed that latvian economy will be able to implement positively the experience of monetary 332 valery i. roldugin: latvia’s participation in the exchange rate mechanism... and credit regulation accumulated in world practice. in spite of the fact that devaluation of lat is matured, euro adoption is more of big modern importance. therefore, referring to a world experience, we should not abandon the possibility of joining the euro zone de facto, although the process of co-ordination with the european central bank may not be less complex than the joining of euro zone. in latvia, a market undergoes a process of euroization for a long time . euro has pressed lat long ago in the conclusion of credit and trading agreements. in order to find the correct options it is possible to make use of experience of the european and other countries which use euro as national currency, but didn’t enter the euro zone : montenegro, kosovo, andorra, monaco, san marino, mayotta, etc. references [1] the bank of latvia: annual report 2003-2010, available at: http://www.bank.lv/statistika/datu-telpa/apraksti/apraksti (accessed mar.28, 2012). [2] data of the central statistical bureau of latvia, available at: http://www.csb. gov.lv/en (accessed dec.12, 2011). [3] data of eurostat, available at: http://epp.eurostat.ec.europa.eu/portal/page/portal/eurostat/home/ (accessed may.11, 2012). [4] ārzemju valūtu tirgus latvijas republikā(2011), available at: http://www.bank.lv/lat/main/all/pubrun/lbgadaparsk/lb1993gadparsk/valstekon/arvvalutas/ (accessed nov.8, 2011). [5] ducmane k. (2011). latvijas nacionālā valūta-vēsture uns̆odiena, available at: http://www.bank.lv/nauda/latvijas-nacionala-valuta (accessed jan.16, 2012). [6] asv pētnieki: lata fiksētās piesaistes dēl. latvija nonāks dzil.as recesijas slazdā (2011). bns. 2010. gada 4. februāris, available at: http://www.diena.lv/lat/business/hotnews/asv-petnieki-lata-fiksetaspiesaistes-del-latvija-nonaks-dzilas-recesijas-slazda (accessed dec.23, 2011). [7] mvf ne trebuet ot latvii deval~vacii nacional~no@i vai�my, ria\novosti". kategorii: latvi� mvf. [8] glavny@i �kspert mvf: lamvi� dop�na deval~virovat~ lat, qtoby izbe�at~ u�estoqeni� �konomiqeskogo krizisa. [9] exchange rates of the bank of latvia, available at: http://www.bank.lv/lat /main/all/finfo/notvalkur/ (accessed nov.2, 2011). [10] the latvian ruble versus the russian ruble, available at: http://www. bank.lv/eng/main/all/pubrun/lbgadaparsk/lb1992gadparsk/v-alstekon1992 /rublelvvsrus/ advances in systems science and applications (2013) vol.13 no.4 333 [11] bel.kovskis k.(2008), “short-term forecasts of latvia’s real gross domestic product growth using monthly indicators”, riga., pp.45-54. [12] roldugin v.(2002), “nav izslēgta latvijas tirgus eirozācija”, latvijas ekonomists, no.4(88), pp.91-97. corresponding author author can be contacted at: vroldugin@hotmail.com. advances in systems science and applications (2014) vol.14 no.1 66-75 towards a new philosophical foundation for physics – speculations on the nature of space, gravity,inertia and mass philip j. tattersall1 and benjamin p. sidebottom2 1lenborough st, beauty point, tasmania, australia 7270 2apartment 5 block a, albion mill, pollard street, manchester. m4 7aj. united kingdom abstract recognizing that physics is now at an important turning point, the authors put forward ideas relating to the nature of space and its role in the emergence of gravity, inertia, mass and, ultimately, the ‘reality’ that derives from (scientific) observation and measurement. the essay cites relatively recent experiments and observations relating to phenomena such as the casimir effect,unruh effect, zitterbewegung and the results from the ‘moving mirror’ experiment. it argues that, when combined with older problems such as quantum entanglement,these phenomena provide new evidence that might inform a better understanding of the role of space in its interaction with particle/field entities. furthermore, it suggests that space may have a significant role in the creation of these entities. the authors suggest that fresh creative insight will be needed for physics to address the scale of the challenge implicit in this new, and exciting, territory. however, this is unlikely to emerge without revising the philosophical framework that underpins physics. this would need to reconcile quantum ontologies with non-quantum ontologies that may be scale dependent. in order to meet the many emerging challenges a more open, participatory and permissive physics is envisioned. keywords entropy, gravity, inclusionality, inertia, mass, quantum vacuum, reality, space. 1 introduction recent theoretical work and experimental findings each refute the idea that space is a benign nothingness, or void. our essay on the significance of space moots a philosophy that we see as a first step in enhancing current theoretical frames leading to a more comprehensive and compatible system of exploration and understanding. while all sciences represent their observations in domains that are, at least, implicitly 3-dimensional, the philosophical framework underpinning the more intensely theoretical sciences, such as physics, make it necessary to theorize in higher dimensions1 in order to ‘explain’, or to account for, the results of the new 1for example time as a 4th dimension in general relativity theory. advances in systems science and applications (2014) vol.14 no.1 67 experiments and observations. as experiments become more sophisticated and sensitive the resultant observations lead to ever more intense debate and speculation as to the ‘real’ nature of reality. while some see this as evidence of a perplexing universe, we are more inclined to see it as an epistemological overhead that stems from continuing to observe and to measure in only 3-dimensions when physics is increasingly multi-dimensional. this would explain why current approaches seem to offer only partial, or illusory, glimpses of what we are exploring. it accounts for the growing and, possibly, unparalleled sense of mystery and ‘weirdness’ that many scientists experience. in the past, science has usefully drawn upon a vital source of ideas and valuable insights from the ‘philosophical spring’ in order to make advancements. today, a similar process is no less important. this paper therefore advances a number of ideas as a way to initiate a conversation about how physics might refocus its philosophical base, in order to invite and encourage creative and informed debate, and research, on the above topics. we believe that the present interest in the nature of space and, in particular, quantum vacuum activity is an early indicator as to the future direction of physics. a successful continuation of this new trajectory in physics will require a (not unprecedented) process of reinvention, or possible revolution, that includes developing a new philosophical foundation. to that end we acknowledge the significant role natural inclusionality [1] has played in influencing our thinking during the development of this paper. according to natural inclusionality the classical description of space requires a reinterpretation: space is regarded merely as the distance over which mass, force and energy are stretched (or stretch themselves), such that they have variable density or frequency, and has no other influence beyond their limits. in this default condition, matter is inert and space passive. the very possibility of motion is therefore made ultimately dependent on some inscrutable external forceful agency or ‘unmoved mover’ to get it going. but if such agency can only be contained or applied locally, where is it? there is clearly something, or rather somewhere, missing from this classical description, which leads energy in the guise of mass and force paradoxically to be mentally confined within and excluded from the boundaries of discrete, completely quantifiable units-i.e. as atomic particles in material bodies, photons in electromagnetic radiation and phonons in heat. that missing somewhere, according to natural inclusionality, is everywhere, without limit-the intangible receptive presence of space. with the dynamic inclusion of this non-local omnipresence within, throughout and beyond local form, movement and change become understood in terms of processes of flow as a continuous energetic reconfiguration of space, not as the travel of independent particles or waves through space. by the same token, massy bodies and electromagnetic radiation are un68 philip j. tattersall: towards a new philosophical foundation for physics ... derstood as distinctive energetic configurations of space, neither solely ‘particles’ nor ‘waves’, but ‘flow-forms’ [1]. 2 inertia, mass and gravity the relatively recent discoveries of a physical manifestation in ‘empty space’ have shown space to be a sea of virtual2 activity3. the casimir and davies-unruh effects [2], the results of the ‘moving mirror’ experiments [3], fermion chirality and zitterbewegung [4-5] all suggest that space plays an active role in the manifestation of certain physical phenomena. we venture to suggest that the virtual sea of activity resident in space may well be the domain of so called ‘hidden variables’ that mediate and, perhaps, cause the emergence of ‘phenomena’ including non-locality, mass conference, inertia and, perhaps, plays a pivotal role in the emergence of ‘gravitational’ influence. rueda and haisch [5-6] have suggested that inertial and gravitational mass each arise from interactions of the electric charges and quarks of matter with the quantum vacuum. they suggest that matter distorts or polarizes the quantum vacuum, leading to an attraction of virtual particles with opposite charges and repulsion of those with like charges (cited by chown [5]). this idea resonates with the idea of matter interacting with ‘space’, as has been proposed elsewhere [1,7-8]. whilst offering a different interpretation of the nature of space itself, we agree that the quantum vacuum somehow interacts with mass and this is what causes the emergent properties of gravity and inertia; hence our proposed mechanism is somewhat similar to that proposed by rueda and haisch [6]. as matter is mostly ‘empty space’, and as space is considered to be a sea of quantum vacuum activity, it would seem that matter itself is permeated with quantum vacuum activity. it follows that the quantum vacuum interactions within matter may play a significant role in the emergence of mass and inertia. when energetic bodies are in uniform motion, the quantum vacuum permeates as laminar flow throughout them, whereas under acceleration we suggest there is a disturbance4 of the quantum vacuum activity, which is manifest as a davies-unruh effect [9-10]. as has been suggested [2,6], the result is an increase in inertia, which we term virtual mass. by way of analogy this could, perhaps,be understood as some kind of induction phenomena whereby quantum vacuum energy becomes stored within the body during acceleration5. radiation emission 2virtual in the sense that such activity is not easily measurable, is barely observable and does not have extension in 3 dimensional ‘space’. 3also termed quantum vacuum noise. 4the quantum vacuum when moving through mass under acceleration induces radiation effects. 5under acceleration there is a kind of ‘induction effect’ (with photon emission) in which the quantum vacuum interacts with fundamental matter fields [6] causing the affected body to advances in systems science and applications (2014) vol.14 no.1 69 during acceleration suggests virtual particle conversion[11], which in our view is not unlike the hawking black hole radiation effect6[2]. it is only at the large accelerations in the vicinity of black holes that we see effects that are barely detectable in our everyday experiences (e.g. at low accelerations, photon numbers are small and their wavelengths are very large [11]). davies sums up the present paradoxes in the following way: a further set of unsolved problems concerns the deeper significance of the relationship between acceleration and quantum vacuum noise. does the existence of “acceleration radiation” suggest a link between the quantum vacuum and inertia? haisch et al.41 claim that the very existence of inertia can be traced to the activity of vacuum noise on an accelerating particle. although this claim has not received widespread support, it is tempting to believe that the distinction between inertial and accelerated motion provided by acceleration radiation is telling us something fundamentally new about the principles of dynamics[2]. davies is suggesting that there is something new, maybe at a deeper level, that has been missed or perhaps misunderstood. in any case, researchers still find themselves having to make sense in the 3 dimensional domain of reality. verlinde [12], in making a conclusion on the origin of gravity, touches on the dynamic role of space in the creation of inertia and gravity. he posits that differences in entropy is the primary cause of gravity: other authors have proposed that gravity has an entropic or thermodynamic origin, see for instance [14]. but we have added an important element that is new. instead of only focussing on the equations that govern the gravitational field, we uncovered what is the origin of force and inertia in a context in which space is emerging. we identified a cause, a mechanism, for gravity. it is driven by differences in entropy, in whatever way defined, and a consequence of the statistical averaged random dynamics at the microscopic level. the reason why gravity has to keep track of energies as well as entropy differences is now clear. it has to, because this is what causes motion! the presented arguments have admittedly been rather heuristic. while we agree that entropic considerations are important and, indeed, relevant we view entropic effects as emergent and, therefore, only indicators of ‘causes’ resident at a deeper level, that is, beyond 3-dimensions. in what follows we attempt to explore these ideas further. become progressively ‘saturated’ with a form of quantum vacuum activity, thus progressively retarding the movement of quantum vacuum through it. this causes a progressive resistance to increasing acceleration. 6see also new scientist. “hawking radiation glimpsed in artificial black hole”. accessed december 21, 2012. http://www.newscientist.com/article/dn19508-hawking-radiation-glimpsedin-artificial-black-hole.html?full=true&print=true for parallels with the unruh effect. 70 philip j. tattersall: towards a new philosophical foundation for physics ... 3 the role of entropy in the quantum vacuum it has been suggested that gravity, rather than a ‘physical field’, emerges from quantum field theory (sakharov cited in [13]). entropy has been suggested to be an element of a similar ‘gravitational’ induction effect, emergent from the energy flux of unobservable degrees of freedom (jacobson cited in [13-14]). whilst entropy is not directly measureable [15] in the 3-dimensional domain it is nonetheless produced there. it is not a physical element of the thermodynamic equilibrium itself (as conceptualized in 3-dimensions); rather it is a virtual effect or flux produced during a reaction, which remains hidden. we propose that entropy, as understood by thermodynamic theory, and when emitted by bond resonances at equilibrium, is a disturbance or flux in quantum vacuum activity caused by the local presence of energetic flux and mass. entropy viewed in this way would be analogous to the casimir effect and is a part of the underlying mechanism for the induction of mass, gravity and inertia, as described later in this paper, and as explored in previous work [7]. we propose that the quantum vacuum of space, rather than being the source of gravitational influence via a polarisation mechanism in itself, also includes ‘hidden variables’, ‘dark’ forms of energetic flux. these ‘hidden variables’ would exist in states that, whilst not being accessible to direct detection or measurement in 3 dimensions when ‘entropic’, none-the-less play a role in the energetic interactions of mass when ‘gravitational’ or ‘inertial’. we suspect that what are currently theorized as dark energy, or dark matter,are varieties of these virtual manifestations permeating the quantum vacuum of space. it is possible that they will provide a route for the investigation of further dimensions and provide evidence of the so-called ‘hidden’ variables that would explain quantum phenomena, such as non-locality. in the case of non-locality, its apparent manifestation in 3 dimensions might suggest that, at a deeper level, ‘distance’ as such, may not exist; furthermore, that the notion of definitive ‘locality’ (as distinct from dynamic locality) is purely a manifestation of 3-dimensional ‘reality’. 3.1 some ideas on the phenomenon of gravity we contend that, in the vicinity of a body, there exists a disturbance in the quantum vacuum proportional to its mass7. in our view, in its interaction with matter, the quantum vacuum induces the emergence of gravity (see sakharov cited in [13]). this disturbance results from a casimir-like effect such that, in the vicinity of mass, there is an imbalance of quantum vacuum activity8. this in turn causes bodies in close proximity to move together, but not necessarily via 7mass is an extension in 3-d of interactions from within the quantum vacuum. mass ‘soaks’ up quantum vacuum activity. 8this is similar to the active gravitational mass idea of haisch and rueda (see [6]). advances in systems science and applications (2014) vol.14 no.1 71 the same mechanism as that proposed by le sage or brush cited in edwards [16]. when two bodies approach each other in ‘free space’, the strength of apparent ‘attraction’ is directly proportional to the magnitude of disturbance of the quantum vacuum activity in the intervening space, caused by the presence of the bodies. within this disturbance, the vacuum bodies induce a ‘suction-like’ reaction that causes them to move together. it is almost as though two bodies that appear to undergo gravitational attraction are pushed together by the higher vacuum activity in the surrounding space. although mass is mostly space there are still residual field, or energy centres observable at any particular scale (we loosely term these ‘horizons’ or ‘points of diminution’ or ‘diminished mass’). this explains why gravitational affect is related to size and mass. it also follows that the interaction of mass and the quantum vacuum can provide an explanation for the equivalence of inertial and gravitational mass. thus, two bodies of unequal mass will fall at the same acceleration in a ‘gravitational field’ because the casimir effect is proportionally influenced by an opposing davies-unruh effect. a comparatively large mass will experience a higher casimir effect, compared with a less massive body, but its ‘fall’ will also be proportionally retarded by an increase in its virtual mass (inertia) due to the davies-unruh effect. on the other hand while the less massive of the two bodies will experience a lower casimir ‘push’, it will also experience a proportionally lower inertial increase or retardation. 4 suggestions for physics – some tentative conclusions the observations and theorising that led to the idea of a quantum vacuum have now taken physics to an important new horizon and we are beginning to glimpse a ‘reality’ far stranger than the one we have become used to. this poses new challenges, because theorizing is tending to take the place of experiments and observations due to the constraints of working in 3 dimensions. on the other hand, certain empirical observations, such as those from the moving mirror experiment, casimir, unruh, the double slit experiment and from non-locality experiments in general are providing tantalizing glimpses of what appears to be a deeper world. however, while these sophisticated experiments are being used to explain the many new and exciting ideas, such as the many world theory, string theory, and so on, their conclusions are limited by being interpreted in only 3 dimensions. in order to explain their strangeness, physicists are forced into making their theorizing processes more elaborate, thus creating a new virtual reality as a surrogate. in the meantime the search continues for the ground-breaking observations and or experiments that might tell us how it all works and what it all means. the excitement over the work on the higgs boson is a recent example. the experiments and observations employed by physics are cosmic in scale, now 72 philip j. tattersall: towards a new philosophical foundation for physics ... that we can take excursions, via our telescopes or interstellar vehicles, into black holes, neutron stars and far-off galaxies. despite some wonderful observations and measurements made from these ‘natural’, massively powerful laboratories, we still try to explicate the data using only 3 dimensions. thus, the problem remains. while we are getting tantalizing glimpses of a mysterious, ‘hidden world’, how can we ‘explain’ them without having to devise increasingly complex experiments and ever more complex theories? in this essay, our missions not to offer a clear solution but, rather, to invite a philosophical conversation about the nature of space. this, we argue,is needed in order to invite the much-needed creative input and ideas that will be needed to define the new order. such a platform is already being suggested by many researchers and distinguished commentators including davies [2], rueda and haisch [6], johnson & walker[17], wuthrich [13] and david tong [18]. for instance, tong’s recent article in scientific american has already initiated an important debate that is at the heart of physics. moreover, its proposition, namely, that reality is a non-quantized continuum is germane to this essay. in our view, this is an exciting and much-needed first step, as we continue to move into a new era of physics. the paper by rafelski et al [19] demonstrates the depth of interest in the problem of the quantum vacuum and in our view straddles an important physicsphilosophy coupling. in their discussion of the ‘three riddles’ at the nexus of quantum theory, particle physics and cosmology the authors bring to the fore not only plausible arguments, but as important suggest new directions for inquiry into the nature of space. the authors argue: contemporary physics faces three great riddles that lie at the intersection of quantum theory, particle physics and cosmology, they are 1. the expansion of the universe is accelerating the extra factor of two appears in the size. 2. zero-point fluctuations do not gravitate a matter of 120 orders of magnitude. 3. the “true” quantum state does not gravitate. the latter two are explicitly problems related to the interpretation and physical role and relation of the quantum vacuum with and in general relativity. their resolution may require a major advance in our formulation and understanding of a common unified approach to quantum physics and gravity. to achieve this goal we must develop an experimental basis. so not only is more research needed, but moreover a new perspective from which to formulate ideas and problems is also needed. we suggest that new and perhaps alternative ontologies are required that encourage new, innovative and creative insights. advances in systems science and applications (2014) vol.14 no.1 73 the maturing philosophy of natural inclusion (ni) offers one such perspective. alan rayner [8] the founder, describes ni in these words: a term introduced by alan rayner and ted lumley, in conversation with others, intended to distinguish a form of reasoning that includes intangible presence and so is more comprehensive, comprehensible and realistic than abstract rationality. eventually it became necessary for alan rayner to distinguish his understanding of inclusionality as ‘natural inclusionality’, which takes account of local influence and identity, from ted lumley‘s understanding, which considers only nonlocal influence and regards locality as illusory. space in the context of ni is described as [8]: according to the logic of natural inclusionality, ... space cannot be pluralized into discrete particularities; it can only be distinguished into distinct, dynamically and permeably bounded regions. this is because a presence that has no resistance can neither be cut nor resisted by a tangible frame. it is inescapably present throughout and beyond the boundaries of tangible figures. a tangible frame is an inclusion of and is included in space but the frame is not the space. the tangible frame can move (or be moved) and be cut, but not the space. when the frame moves the space stays where it is: in relative terms by remaining still space permeates freely through the frame, the frame does not cut through the space. moreover, if the frame is to move without being forced to do so by a force situated somewhere outside of it, it must have the capacity for movement within itself, i.e. the frame is itself a manifestation of energy, not inert structure-it is a variably fluid ‘framing’, not a permanent, absolutely rigid ‘framework’. this tangible ‘framing’, or ‘dynamic interfacing’, has to be present for form to be distinguishable in a feature-full cosmos, but it can neither ‘occupy’ nor ‘exclude’ the space that it includes and is included in. ni accommodates continuousness and therefore perhaps at certain scales posits that reality is in fact non-quantum in nature, as suggested by david tong [18]. the introduction of such philosophical innovations would, in our view, invite new and perhaps very productive conversations leading to ideas and maybe new insights that might create conditions for thinking about old and current problems in new ways. after all that is what has happened many times before as ‘breakthroughs’ have arisen in the most unexpected of ways. in this essay metaphor and analogy have been employed as they help to invite new ideas and imaginings. our approach is therefore very much in keeping with the inquiry trajectory proposed by bohm [20] who suggested that future scientists would be less dependent on mathematics and modelling as they begin to draw upon new approaches that in the end would lead to a merging of art and science. acknowledgements 74 philip j. tattersall: towards a new philosophical foundation for physics ... the authors acknowledge the support of emeritus professor john wood who offered several very helpful suggestions and recommendations during the final stages of manuscript preparation. also our thanks to dr. alan rayner who offered useful suggestions during the early drafting of this paper. references [1] rayner a.d.m. (2011), “space cannot be cut: why self-identity naturally includes neighbourhood”, integrated psychological behaviour, vol.45, pp.161184. [2] davies p.c.w. (2001), “quantum vacuum noise in physics and cosmology”, chaos, vol.11, no.3, pp.539-547. [3] brumfiel g. (2011), “moving mirrors make light from nothing”, accessed november 15, 2012. http://www.nature.com/news/2011/110603/full /news.2011.346.html [4] huang, k. (1952), “on the zitterbewegung of the dirac electron”, american journal of physics, vol.20, no.8, pp.479-484. [5] chown,m. (nd), “mass medium why are loaded fridges difficult to budge? because empty space impedes them”, accessed december 15, 2012. http://www.calphysics.org/articles/chown2007.html [6] rueda a. and haisch b. (2005), “gravity and the quantum vacuum inertia hypothesis”, annals of physics (leipzig), vol.14, no.8, pp.479-498. [7] rayner a. d. m., b., peleshok d. and tattersall p. (2012), “place-time: the flow geometry of space”, accessed december 1, 2012, http://www.bestthinking.com/articles/science/math/place-time-theflow-geometry-of-space. [8] rayner a. d. m. (2012), “a natural inclusional glossary of terms”, accessed january 15, 2013. http://www.bestthinking.com/articles/society and humanities/languages/english language/a-natural-inclusional-glossary -of-terms. [9] davies p.c.w. (1975), “scalar particle production in schwarzschild and rindler metrics”, journal of physics a: mathematical and general, vol.8, no.4, pp.609-616. [10] unruh w.g. (1976), “notes on black-hole evaporation”, physical review d, vol.14, no.4, pp.870c892. advances in systems science and applications (2014) vol.14 no.1 75 [11] mcculloch m. e. (2012), “testing quantised inertia on galactic scales”, accessed january 2, 2013. http://arxiv.org/abs/1207.7007v1 [physics.genph] [12] verlinde e. p. (2010), “origin of gravity and the laws of newton”, accessed november 30, 2012. http://arxiv.org/abs/1001.0785 [hep-th] [13] wuthrich, c. (2005), “to quantize or not to quantize: fact and folklore in quantum gravity”, philosophy of science, vol.72, pp.777-788. [14] jacobson t. (1995), “thermodynamics of spacetime: the einstein equation of state”, physical review letters,vol.75, no.7, pp.1260-1263. [15] angrist s. w., and helper, l.(1973), laws of energy and entropy: order and chaos, ringwood, victoria: pelican books. [16] edwards m. r. (2007), “photon-graviton recycling as a cause of gravitation”, apeiron, vol.14, no.3, pp.214-233. [17] johnson g.w., and walker, m.(2005), “sir michael atiyah‘s einstein lecture: the nature of space”, notices of the american mathematical society, vol.53, no.6, pp.674-678. [18] tong d. (2012), “the unquantum quantum”, scientific american, vol.307, no.6, pp.32-35. [19] rafelski j., labun l.,hadad y., and chen p.(2009), “quantum vacuum structure and cosmology”, accessed december 23, 2012. http://arxiv.org/pdf/0909.2989v1.pdf [20] bohm‘s alternative to quantum mechanics, david albert scientific american may 94. last words of a quantum heretic, john horgan, new scientist 29 feb 93 accessed august 12, 2013. http://www.dhushara.com/ book/quantcos/bohm/bohm.htm. corresponding author philip j. tattersall can be contacted at: soiltechresearch@bigpond.com. adv syst sci appl 2020; 03; 24-35 published online at https://ijassa.ipu.ru. original russian text © m.a. gorelov, 2018, published in upravlenie bol’shimi sistemami / large-scale systems control, 2018, no. 72, pp. 6-26. the “value at risk” principle in hierarchical game mikhail gorelov1* 1) dorodnicyn computing centre, federal research center “computer science and control”, russian academy of sciences, moscow, russia e-mail: griefer@ccas.ru abstract: a hierarchical game of two persons with random factors is considered. it is assumed that the top-level player has the right of the first move. it is believed that the lower level player at the time of decision making knows exactly the realization of the random factor and the choice of partner. and the top-level player at the time of decision making knows only a probabilistic measure on the set of values of an uncertain factor. the principle of optimality is new: it is believed that a top-level player is ready to neglect some of the “unpleasant” events, the total probability of which is given, but otherwise he is careful. under these assumptions, the maximum guaranteed result of the top-level player is calculated. the structure of strategies providing such a result is clarified. two cases were investigated: a game with and without feedback. to solve the problem, an original definition of the maximum guaranteed result is proposed. it is equivalent to the classical definition, but is simpler. using this technique, solving of the problem reduces to identical transformations of the formulas for predicate calculus. as a result of the solution, the optimal strategy and the set of “unpleasant” cases which are excluded from consideration search task is reduced to calculating multiple maximins on finite-dimensional spaces. in this case, the operation of calculating the expected value with respect to given probabilistic measure is considered to be “elementary”. models of this type can have different interpretations. one can use them for methodological justification of the principle of maximum guaranteed result. one can use them when solving risk management tasks. one can consider them as models for managing the “customer base” of the service company. the proposed method allows to study such models at a qualitative level, and in some cases to obtain quantitative results. keywords: informational theory of hierarchical systems, hierarchical games, decision making under risk, maximal guaranteed payoff, risk management 1. introduction a study of hierarchical games with uncertain factors was started in the early seventies of the last century [8,9,15]. around the same time, similar models were investigated in the theory of active systems [4,5,18] and in the theory of contracts [2,3,16]. in this case, two methods of eliminating uncertainty were mainly studied. in one of them, it is assumed that the players know only sets of possible values of uncertain factors and they are careful i.e. reckons upon the worst option for themselves. another considered that the probability measure is defined on the set of uncertain factors, and the players are risk-neutral, that is they are ready to be guided by the expectations of their payoffs. a third, in a sense, intermediate way of eliminating uncertainty is also possible. in the financial engineering it received title principle “value at risk” [6,17]. the development of these ideas can be found in the monograph [1] and other works of the same author. at a meaningful level, its essence can be explained as follows. i do not think that making decisions in everyday life, people are very often appreciate some probability distributions, * corresponding author: griefer@ccas.ru the “value at risk” principle in hierarchical game 25 copyright ©2020 assa. adv. in systems science and appl. (2020) and even more so calculate mathematic expectations. and on a professional level, i have never had to deal with the customer who formulates the problem in probabilistic terms. but the tendency to the principle of guaranteed results, customers several times clearly formulated. but the principle of maximum guaranteed result has one not a very attractive property. if it is conducted quite consequentially, the possibility cannot be ruled out that the player “does not fall down a brick on his head” or something else unpleasant happens. and constantly focusing on such cases, it is hardly possible to make really effective decisions. in practice and in theory [7], from this situation, the next way out is used. a part of the completely “fatal” values of the uncertain factor is excluded from consideration, and with respect to the remaining part, the principle of maximal and minimal guaranteed result is used. on what basis are some possibilities excluded from consideration? probably, it happens because the operating party regards them as “unlikely”. of course, there is the temptation to postpone the work for exclusion of unfavorable factors from the level of the model building to the level of its research. for this purpose, it is necessary the operating party accept some probability distribution and said that with a given probability  it wants to get a payoff not less than the value of . of course, it is preferably to the value of  to be larger. thus, we come to the statement of the problem considered below. it is unlikely that all this makes sense in the analysis of the problems at the household level. but in the analysis of business decisions risk management problems has recently become very relevant, if not fashionable. moreover, the risk assessment method discussed above is one of the two most popular (as far as i know, game-theoretic models in which risk is estimated using dispersion also have not yet been investigated). but this method is usually used to investigate problems of centralized decision making under risk. game-theoretic models of this kind, apparently, have not yet been studied. further presentation is constructed as follows. the next section gives a formal statement of the problem. the rest of the article is devoted to its “solution”. by tradition the solution of the problem of calculating of maximal guaranteed result in hierarchical game with a “complicated” structure is regarded as it’s reduction to some “elementary” operations such as computation of maxima and minima on the “simple” sets. we will follow this tradition. in a problem under consideration “non-elementary” operation of selection of a set of uncertainties to be excluded from consideration arises already in games not endowed with additional structure. section 3 is devoted to their study. in the fourth paragraph, the more traditional problem of finding the maximal guaranteed result in a “feedback games” is solved. 2. games with random factors so, we begin to describe the simulated conflict. game with random factors will be hereinafter called a six-tuple  =  u , v , a , g , h ,  . here, u , v and a are the sets, g – function which maps the cartesian product u  v to the set of real numbers , : ,h u v a  → and  is probability measure on the set a. these constructions are interpreted as follows. it is assumed that two participants take part in the game, which we will call the first and second players. a set u is interpreted as the set of controls of the first player, a set v – the set of controls of his partner. we assume that the interests of the first and second players are described by the desire to maximize the functions g and h, respectively. the value of the indefinite factor   a is chosen by some third party – nature. this choice is made randomly in accordance with the distribution . we make the traditional technical assumptions, which significantly simplify the further narration. the sets u, v, and a are assumed to be endowed with topologies and compact ones. the functions g and h are assumed to be continuous. measure  will be considered to be borel. 26 m.a. gorelov copyright ©2020 assa. adv. in systems science and appl. (2020) unfortunately, not all results can be obtained in such a general assumptions. additional conditions on game under consideration, if required, will be indicated when corresponding results will be formulated. game  describes the possibilities and interests of the players. for the model completion one must also describe the dynamics of decision-making and the attitude of the players to the existing uncertainty. let’s start with the interpretation. we believe that all the game  options are exactly known to the first player. we assume that events take place as follows. at first, the first player chooses his control u  u. then, the specific value of the indefinite factor   a is realized (in accordance with the given probabilistic measure ). the values of u and  become known to the second player. thus, for the second player, no uncertainty remains. consequently, for him all the controls v  v will be divided into “reasonable”, in case of the choice of which he will receive a payoff greater than or equal to a certain number of , and “unreasonable” , the choice of which promises him payoff smaller then  (of course, number  depends on u and ). this principle of behavior is known to the first player. but he is careful and therefore he reckons upon the worst result that can happen when “reasonable” choice of partner will be made. but since the value  is not known to the first player, for him this result is a random variable. the attitude of the first player to this uncertainty is as follows. he agrees to exclude from consideration a certain number of “force majeure” events, the total probability of which does not exceed a given value 1 – . but other than, he focuses on the worst case for himself and wants to get the maximal guaranteed result. formally, the above is described as follows. definition 2.1: let a real number   [0,1] be given . number  is -guaranteed result of the first player in the game , if there exist a measurable set b  a, measure (b) of which is greater than or equal to , and such strategy u  u, that for every   b there exist a number , for which the following conditions are hold true: 1 . there exists w  v for which h(u,w,) ≥ ; 2 . for any v  v, either g(u,v) ≥  or h(u,v,) < . supremum of -guaranteed results of the first player is called its maximal -guaranteed result. remark 2.1: it would be possible to refuse the assumption of measurability of set b in this definition, replacing measure (b) with the corresponding outer measure. it will be seen from what follows that, under the assumptions that the measure  is borel, and the function h is measurable, such a modification of the definition do not essentially change anything. remark 2.2: the definition 2.1 is a modification of definition of maximal guaranteed result, proposed by the author in [10]. in more conventional terms the maximal -guaranteed result can be defined by the formula ( , ) supsupinf min ( , ) b v br ub u u g u v    wherein the outer supremum is taken over all measurable subsets b of the set a, for which (b) ≥ , and  ( , ) : ( , , ) max ( , , ) . w v br u v v h u v h u w    =  = the “value at risk” principle in hierarchical game 27 copyright ©2020 assa. adv. in systems science and appl. (2020) the proof of the equivalence of these definitions only insignificant differs from the reasoning in [10]. a significant part of this work will, in fact, be done in the next section. for more complex problems definition 2.1 is more convenient, therefore we will use it. remark 2.3: all probabilities in this paper can be regarded as subjective, namely, as an evaluation of the operating party (first player) the feasibility of certain events. in the future, no results such as the law of large numbers are used. therefore, there is no need to take care of any kind of statistical stability. indeed only readiness of operating party to eliminate uncertainty by the method described in the definition 2.1 is important. 3. game without feedback calculation of maximal -guaranteed result, for example, by the formula from the remark 2.2 involves computation of least upper bound on the class of subsets of the set a. standard methods for calculating of such supremum not exists even for the relatively simple case where the set a is a segment. in this section, we simplify the solution of the problem by replacing the operation of computing such an upper bound with the operation of calculating the expected value. traditionally such an operation is considered to be “elementary”. introduce the following notation ( , ) max ( , , ). v v m u h u v   = let  be -guaranteed result of the first player in the game . choose a set b  a and strategy u  u, the existence of which is provided for the definition 2.1. fix an arbitrary   b. for strategy w  v, the existence of which postulate the item 1 of definition the inequality h(u,w,) ≥  is satisfied, the more this inequality must be satisfied for the strategy w0 determining by the equality h(u,w0,) = m(u,). consequently the number , appearing in definition 2.1, must satisfy the condition   m(u,). therefore, if the item 2 of definition is satisfied for someone value , it moreover holds for  = m(u,). but with such a value of  item 1 also is obviously satisfied: it suffices to choose, for example, w = w0. thus, the number of  is -guaranteed results of the first player in the game , if there is a measurable set b  a , measure (b) of which is greater than or equal to , and such a strategy u  u , that for every   b and any v  v either g(u,v ) ≥  or h(u,v,) < m(u,). consider the set  ( , ) : ( , , ) max ( , , ) . w v br u v v h u v h u w    =  = if v  br(u,), then h(u,v,) = m(u,); therefore , by the second item of definition 2.1, the inequality g(u,v) ≥  must hold. conversely, if the last inequality holds for all v  br(u,), then the number  is -guaranteed result. indeed, for w  br(u,) and  = m(u,) the first item of definition is satisfied. for the same values of  and v  br(u,) the item 2 is met because of in this case h(u,v,) < m(u,), and for v  br(u,) it satisfies as by the assumption g(u,v) ≥ . thus, the number of  is -guaranteed result of the first player in the game , if there exist a measurable set b  a, measure (b) of which is greater than or equal to , and such a strategy u  u, that for every   b and any v  br(u,) the inequality g ( u , v ) ≥  is true. the same condition can be formulated equivalently: the number  is -guaranteed result of the first player in the game , if there exist a measurable set b  a, measure (b) 28 m.a. gorelov copyright ©2020 assa. adv. in systems science and appl. (2020) of which is greater than or equal to , and such a strategy u  u, that for any   b, the inequality ( , ) min ( , ) v br u g u v     (3.1) holds. consider the set   ( , ) ( ) : min ( , ) v br u c u a g u v     =   . since condition (3.1) holds for all   b, the set b must be contained in the set c(u), and hence , the condition (c(u)) ≥ (b) ≥  must holds. conversely, if the condition (c(u)) ≥  is valid for some strategy u, then condition (3.1) will be satisfied for all   b = c(u) , and hence  is -guaranteed result. let define the function (x) by the condition 1, if 0, ( ) 0, if 0. x x x   =   the measure of the set c(u) is equal to ( ) ( , ) min ( , ) , v br u g u v      − where the symbol  designate operator of calculating mathematical expectations with respect to measure . this immediately yields the following result. theorem 3.1: in order for the number  to be -the guaranteed result of the first player in the game , it is necessary and sufficient that either ( ) ( , ) max min ( , ) , v br uu u g u v       −  (3.2) or ( ) ( , ) sup min ( , ) , v br uu u g u v       −  if the upper bound in the last formula is not reached. remark 3.1: in the papers [11] and [12] in the same manner was formulated conditions characterizing maximum guaranteed result for games with undefined interval uncertainty and games with risk-neutral first player. in earlier papers [14] and [15], although with some additional assumptions explicit formulas were obtained for the maximum guaranteed results in these problems. in the problem considered in this paper, such an alternative is not visible even for simpler analogue problem of optimization corresponding the game , wherein a set v consists of one point. remark 3.2: formula (3.2) could be a starting point in an attempt to give a more traditional definition of maximal -guaranteed result. but such a definition in this case would require clarification. and apparently, this explanation inevitably would be similar to the definition 2.1. in addition, quite a “classical” form of this definition can’t be given, because the number  is part of the argument of function . thus, in this case it is not very convenient to the “value at risk” principle in hierarchical game 29 copyright ©2020 assa. adv. in systems science and appl. (2020) follow the traditions. in the next section, the advantage of the new definition will become even more obvious. in order not to be distracted by the technical details in the proof of theorem 3.1 was left a gap. to fill it, it is necessary to prove the following statement. lemma 3.1: for every u  u, the set c(u) is measurable. proof. fix u  u. it is sufficient to prove that the function ( , ) ( ) min ( , ) v br u g u v     = is measurable. to do this, it sufficient to prove that the function –() = 0 – () is measurable. therefore, it suffices to prove that for any  the set    0 ( , ) ( ) : ( ) : min ( , ) v br u c u a a g u v         =  −  − =   is measurable (for convenience, this part of the proof uses only the facts explicitly stated in [13]). since the measure  is assumed to be borel, it is sufficient to prove that the set c0(u) is open. suppose the contrary. then there exist   c0(u) and a sequence , ... such that lim k k   → = and  k  c0(u) for all k = 1,2, ... the set br(u, k) is defined by the condition of equality type, so it is closed due to the continuity of function h. since a set v is assumed to be compact the set br(u,k) also will be compact. hence, at some point vk  br(u,k) the minimum of the function g(u,v) on the set br(u,k) is achieved. since by assumption k  c0(u), the inequality g(u,vk)   holds. since the set v is compact it is possible without loss of generality to assume that the sequence v1, v2, ... converges to an element v  v. fix an arbitrary w  v. since vk  br(u,k), the inequality h(u,vk,k) ≥ h(u,w,k) holds. since the function h is continuous, going to the limit in this inequality, we obtain h(u,v,) ≥ h(u,w,). since w is arbitrary, it follows that v  br(u,). and going to the limit in the inequality g(u,vk)  , we get g(u,v)  , and even more so ( , ) min ( , ) v br u g u v     which contradicts the condition   c0(u). the obtained contradiction proves the lemma. □ thus, theorem 3.1 is completely proved. remark 3.3: the assumptions formulated in the previous section about the topological and metric structure of the game  can be changed, and, perhaps, simply weakened. the question of the weakest assumptions under which lemma 3.1 and further results of this type remain true, is of some interest, but beyond the scope of this article. the most interesting model examples certainly satisfy the conditions formulated above, so we omit further discussion. remark 3.4: the proof of lemma 3.1 is quite standard, but rather long. for this reason, further, the proof of similar results is omitted. a question of search optimal strategy of the first player in this game is quite meaningful. given the results obtained, the answer to it is not difficult to find. first of all, we note that the supremum in the definition of maximal -guaranteed result can’t be achieved. therefore, there may not exist a strategy allowing one to obtain such a result with probability . this effect is understandable because similar fact takes place already in the game with no uncertainty (i.e., with single-point set a). if the number  is a -guaranteed result, then any solution of the inequality 30 m.a. gorelov copyright ©2020 assa. adv. in systems science and appl. (2020) ( ) ( , ) min ( , ) , v br u g u v       −  if the upper bound in formula (3.2) is reached, or the inequality ( ) ( , ) min ( , ) , v br u g u v       −  otherwise, is the desired strategy. in both cases, the inequalities have solutions, since by assumption  is -guaranteed result. one of sets b “suitable” for this strategy can be defined by the condition b = c(u). of course, such selection of the set b is not the only possible. however, in the general case, the same applies to the choice of the optimal strategy u. 4. game with feedback consider another game * = u*, v*, a, g*, h*, , in a certain way related to the game . denote by (x,y) the class of all functions from the set x into the set y. let u* = (v  a,u), v* = v  a, and the functions g* and h* are determined by the conditions g*(u*,v*) = g(u*(v,),v), h*(u*,v*,) = h(u*(v,),v,), where v* = (v,). the set a and the measure  on it are the same as in the game . these constructions can be interpreted as follows. players choose their “physical” controls from the sets u and v. but by the time of the selection of its control u  u the first player receives reliable information about control v  v selected by his partner. in addition, the second player can transmit to the first player information on the realized value of the uncertain factor. however, this information does not have to be reliable, that is, the second player has the right to select some message   a, which he will transmit to the partner. his physical control u*(v,) the first player chooses on the basis of all the information received, and both players’ payoffs depend only on the physical choices made by them, and do not depend on information they exchanged. game  has the same structure as the game , so the question can be put of finding the maximal -guaranteed result in this game. we will deal with this task. when analyzing this problem, using the english language is already quite inconvenient. therefore, we turn to the language of predicate calculus. for the game  the definition of -guaranteed result  will look as follows:     * * * * ( , ) : ( ) & & : ( ( , ), , ) & & ( ( , ), ) ( ( , ), , ) . b u v a u b b w v a h u w w v v a g u v v h u v v                                   (4.1) this formula is not “elementary”, because it contains one existential quantifier refers to a class of all measurable subsets b of the set a, and another – to the class of all functions u from the set v  a to the set u. it can be simplified, but one has to make the following assumption. hypothesis 4.1: the function h is such that there exists such control up  u, that for any v  v and any   a the equality ( , , ) min ( , , )p u u h u v h u v   = holds. in fact, it assumes the existence of a universal (independent of ) strategy of punishing the second player by first. the “value at risk” principle in hierarchical game 31 copyright ©2020 assa. adv. in systems science and appl. (2020) now we can start converting the formula (4.1). as in the previous section let’s start with the specification of the value . put  ( ) ( , ) : ( , ) ,h u v u v g u v =    ( , ) ( ) ( , ) max ( , , ). u v h l h u v      = in meaningful terms, h() is the set of “acceptable” outcomes for the first player. the number l(,) characterizes the maximum payoff that the second player can get, provided that the first one somehow ensures that he gets an “acceptable” result (of course, this maximum payoff depends on the indefinite factor ). according to the condition (4.1), there exist a set b and a function * such that     * * * : ( ) & : ( ( , ), , ) & & ( ( , ), ) ( ( , ), , ) , b b w v a h w w v v a g v v h v v                                  (4.2) fix such a set and a function. then there exists a function u*, for which the condition is satisfied:     * * * ( ) & & : ( ( , ), , ) ( , ) & & ( ( , ), ) ( ( , ), , ) ( , ) . b b w v a h u w w l v v a g u v v h u v v l                               (4.3) let’s prove it. let the condition (4.2) be satisfied. then the set h() is not empty. indeed, let’s fix any   b. then by virtue of condition (4.2) we have h(*(w,),w,) ≥ . then, due to the same condition g(*(w,),w) ≥  and therefore (*(w,),w)  h(). for each   a let’s fix a pair (u,v)  h() such that h(u,v,) = l(,). put * * , if and , ( , ) ( , ) in other cases. u v v u v v         = = =   then for any   a a condition *: ( ( , ), , ) ( , )w v a h u w w l         is satisfied (one can take w = v and  = ). in addition, from the inequality h(*(w,),w,) ≥  it follows that   l(,), so condition (4.2) implies * *( ( , ), ) ( ( , ), , ) ( , ).g v v h v v l          (4.4) if *(v,)  u*(v,), then by construction inequality g(u*(v,),v) ≥  is true, hence the condition * *( ( , ), ) ( ( , ), , ) ( , )g u v v h u v v l        (4.5) is satisfied. otherwise, conditions (4.4) and (4.5) are equivalent. thus, it is proved that condition (4.2) follows condition (4.3). the reverse implication is obvious. therefore, conditions (4.2) and (4.3) are equivalent. exactly the same “modification” of the strategy of the first player proves that the condition     * * * * ( , ) ( ) & & : ( ( , ), , ) ( , ) & & ( ( , ), ) ( ( , ), , ) ( , ) b u v a u b b w v a h u w w l v v a g u v v h u v v l                                   is equivalent to a simpler condition 32 m.a. gorelov copyright ©2020 assa. adv. in systems science and appl. (2020)   * * * ( , ) ( ) & & ( ( , ), ) ( ( , ), , ) ( , ) . b u v a u b b v v a g u v v h u v v l                         the relevant reasoning is practically the same as the above, so we omit them. change the order of the generality quantifiers in the last formula:   * * * ( , ) ( ) & & ( ( , ), ) ( ( , ), , ) ( , ) . b u v a u v v a b g u v v b h u v v l                         now, one can use the structure of the set of strategies of the first player to change the order of quantifiers of existence and generality:  : ( ) & ( , ) ( , , ) ( , ) .b v v a u u b g u v b h u v l                    the variable  in square brackets has “disappeared”, so this formula can be further simplified:  : ( ) & ( , ) ( , , ) ( , ) .b v v u u b g u v b h u v l                 denote  ( ) : max ( , ) . u u e v v g u v   =   then the previous condition can be rewritten in an equivalent form: ( ) : ( ) & ( , , ) ( , ),b v e u u b b h u v l               or ( ) ( ) & : ( , , ) ( , ).b v e b u u b h u v l               now let us use hypothesis 4.1 to change the order of the generality and existence quantifiers: ( ) ( ) & : ( , , ) ( , ).b v e b b u u h u v l               once again changing the order of the generality quantifiers, we get ( ) & ( ) : ( , , ) ( , ).b b b v e u u h u v l               let 1, if 0, ( ) 0, if 0. x x x   =   replacing the quantifiers of generality and existence with the operators of maximum, minimum and expectation, how was it done in the previous section, we get the following result. theorem 4.1: let hypothesis 4.1 be fulfilled. then in order for the number  to be a -guaranteed result, it is necessary that ( )( ) ( ) inf max ( , ) ( , , ) , v e u u l h u v          −  and it is sufficient that ( )( ) ( ) inf max ( , ) ( , , ) . v e u u l h u v          −  (4.6) the “value at risk” principle in hierarchical game 33 copyright ©2020 assa. adv. in systems science and appl. (2020) remark 4.1: there are games  for which the least upper bound of the numbers  satisfying the necessary condition in the theorem differs from the least upper bound of the numbers  satisfying the sufficient condition. for such games, the results obtained do not give a definitive answer to the question, what is the maximum -guaranteed result? but it does not make sense to refine the obtained necessary and sufficient conditions in this case, since it is clear that for such games the problem of calculating the maximal -guaranteed result is not stable with respect to small changes in the parameters of the game . therefore, additional study this problem is required, which is beyond the scope of this article. however, such games are in a sense “exceptional”. remark 4.2: the acceptance of hypothesis 4.1 from a formal point of view seems rather restrictive, since it is essentially assumed that all functions from some parametric family have saddle points. but in many meaningful models, its use seems justified. for example, if the first player chooses a price, he can choose it as the minimum possible, if he allocates a resource to a partner, he can allocate it “at a minimum” under all conditions. a. f. kononenko generally believed that in economic models, hypothesis 4.1 is always fulfilled. i do not quite share this view, for the principle of proportion of the severity of the punishment to the severity of the offence must be considered. but in this case it is impossible to abandon this hypothesis. in this sense, the problem considered in this paper is more complicated than the problem with a risk-neutral player, where, as shown in [12], a similar hypothesis can be abandoned by introducing a “gauge” additive to the payoff function of the second player. analysis of the proof shows that hypothesis 4.1 can be replaced by the following assumption. hypothesis 4.2: there exist a control u  u such that the inequality max ( , , ) ( , ) v v h u v l     holds for all   a. since the inequality in hypothesis 4.2 is strict, it cannot be argued that it is weaker than hypothesis 4.1. however, it is quite possible to expect that there are quite a lot of meaningful models in which hypothesis 4.2 holds and hypothesis 4.1 does not. hypothesis 4.1 is accepted as the main one, since it is easier to interpret. if the sufficient condition (4.6) of theorem 4.1 is satisfied for some number , then the results obtained above allow us to construct one of the strategies that allow us to obtain the result  with probability . as the set b, which appears in the definition of the maximum guaranteed result, we can take the set ( )  ( ) : inf max ( , ) ( , , ) 0 . v e u u b a l h u v        =  −  for any  from the set b thus chosen let’s choose an arbitrary pair (u,v) from the set h() satisfying the condition h(u,v) = l(,). put * , if и , ( , ) in all other cases.p u b v v u v u      = =   it is directly verified that the so-defined strategy u* is the desired one. the interpretation of these constructions is standard. the second player is asked to select control v, if the value of the undefined factor   b has been realized, and to report the true information about this factor. in this case, the first player promises to use the “encouraging” control u. otherwise, the second player faces punishment. controls u and v are chosen so 34 m.a. gorelov copyright ©2020 assa. adv. in systems science and appl. (2020) that the message of reliable information is really beneficial to the second player. cases   b are excluded from consideration by the first player. therefore, in particular, for such values , the set h() can be empty. in these cases, the strategy of punishment is used for greater reliability. 5. conclusion in addition to the “methodological” interpretation described in the introduction, the studied model has another, perhaps more interesting one. the value 1 –  in this model can be naturally considered as a measure of risk. thus, the model explicitly describes both the “yield”, estimated by the value of the payoff g(u,v), and the risk. this seems important enough. it is quite natural to assume that the value  is the control of the operating party (the first player), along with u. here we can assume that the order of decision-making is as follows. first, the first player fixes the value  and the strategy u (or u*, respectively), then the value of the uncertain factor  is realized, then the second player chooses his control. in this case, it does not matter whether the second player receives information about the selected value , since his payoff does not depend on him. however, the model is not fully formed, because in this case it is natural to assume the presence of two criteria: the risk measure 1 –  and the corresponding -guaranteed result. it follows directly from the definition that, as  increases, the corresponding -guaranteed result does not increase. the choice of balance between the two criteria is left to the operating party. but if the researcher of the operation has an effective way of calculating the -guaranteed result, it will be a serious help in solving this problem. the described method of risk accounting, of course, is not the only one possible. but already on the basis of the studied model it is possible to construct other meaningful problem statements. for example, one can assume that the operating party selects a number of values  1, 2,…, n. for each strategy u  u it is possible to find the payoff of the first player  i(u) which with a probability of  i is provided with the choice of strategy u under the rational actions of a partner. and then a multi-criteria problem is solved with the criteria  1(u), (u)2,…, (u)n. to these criteria, one can add the expected value of a guaranteed payoff of the first player. thus, we get a fairly wide range of statements, each of which can be “tried on” for the simulated situation. apparently, the key step in the study of the corresponding problems is made in this article. references 1. agasandyan, g.a. (2011). primeneniye kontinual'nogo kriteriya var na finansovykh rynkakh [application of the var continuum criterion in financial markets]. moscow, russia: ccas, [in russian]. 2. bolton, p. & dewatripont, m. (2005). contract theory. mass.: mit press. 3. bremzen, a.s. & guriev, s.m. (2005). konspekty lektsiy po teorii kontraktov [lecture notes on contract theory]. moscow, russia: resh, [in russian]. 4. burkov, v.n. (1977). osnovy matematicheskoy teorii aktivnykh system [fundamentals of the mathematical theory of active systems]. moscow, ussr: nauka, [in russian]. 5. burkov, v.n. & novikov, d.a. (1999). teoriya aktivnykh sistem: sostoyaniye i perspektivy [the theory of active systems: state and prospects]. moscow, russia: sinteg, [in russian]. the “value at risk” principle in hierarchical game 35 copyright ©2020 assa. adv. in systems science and appl. (2020) 6. dempster, m.a.h. (ed.). (2002). risk management. value at risk and beyond. cambridge: cambridge university press. 7. germeĭer, yu.b. (1971). vvedeniye v teoriyu issledovaniya operatsiy [introduction to operations research theory]. moscow, ussr: nauka, [in russian]. 8. germeĭer, yu.b. (1986). nonantagonistic games. dordrecht, germany: d. reidel publishing co. 9. gorelik, v.a., gorelov, m.a. & kononenko, a.f. (1991). analiz konfliktnykh situatsiy v sistemakh upravleniya [analysis of conflict situations in control systems]. moscow, ussr: radio i svyaz', [in russian]. 10. gorelov, m.a. (2011). maximal guaranteed result for limited volume of transmitted information. automation and remote control, 72(3), 580–599, doi: 10.1134/s000511791103009x. 11. gorelov, m.a. (2016). iyerarkhicheskiye igry s neopredelennymi faktorami [hierarchical games under uncertainty]. upravleniye bol'shimi sistemami, 59, 6–22, [in russian]. 12. gorelov, m.a. (2016). iyerarkhicheskiye igry so sluchaynymi faktorami [hierarchical games under risk]. upravleniye bol'shimi sistemami, 63, 87–105, [in russian]. 13. kolmogorov, a.n. & fomin, s.v. (1961). elements of the theory of functions and functional analysis, volume 2. albany, ny: greylock press. 14. kononenko, a.f. (1973). the role of information on the opponent's target function in a two-person game with a fixed sequence of moves. ussr computational mathematics and mathematical physics, 13(2), 49–56, https://doi.org/10.1016/0041-5553(73)90130-4. 15. kononenko, a.f., khalezov, a.d. & chumakov, v.v. (1991). prinyatiye resheniy v usloviyakh neopredelennosti [decision making under uncertainty]. moscow, ussr: vc ran, [in russian]. 16. laffont, j.-j. & martimort d. (2002). the theory of incentives: the principal–agent model. princeton: princeton university press. 17. marshall, j.f. & bansal, v.k. (1992). financial engineering. ny: allyn & bacon. 18. voronin, a.a., gubko, m.v., mishin, s.p. & novikov, d.a. (2008). matematicheskiye modeli organizatsiy [mathematical models of organizations]. moscow, russia: lenand, [in russian]. https://www.sciencedirect.com/science/journal/00415553 https://www.sciencedirect.com/science/journal/00415553 https://doi.org/10.1016/0041-5553(73)90130-4 1. introduction 2. games with random factors 3. game without feedback 4. game with feedback 5. conclusion 14. kononenko, a.f. (1973). the role of information on the opponent's target function in a two-person game with a fixed sequence of moves. ussr computational mathematics and mathematical physics, 13(2), 49–56, https://doi.org/10.1016/0041-5553(73)9013... adv syst sci appl 2018; 03; 144-153 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/634 on classification and assessment of regional economic and social effects of the implementation of projects for mining of solid minerals deposits andrey m. valuev1, maria a. lozinskaya2 1) mechanical engineering research institute, russian academy of sciences, moscow, russia e-mail: valuev.online@gmail.com 2) the national university of science and technology misis, moscow, russia e-mail: mary-loz@mail.ru abstract: the main regional economic, ecological and social effects resulting from the implementation of design solutions in the field of mining of solid mineral deposits are systematized in general. two cases are distinguished, for which the new enterprise production is mainly consumed within the region and outside it, resp. typical impact on transport branch and social sphere is studied for both cases. the paper mainly concentrates on the case of russian federation, since differences in legal system (in particular, the fiscal system), culture of everyday life, territorial distribution and mobility of population, as it is shown in scientific literature, lead to different form of interrelation between these factors in involved countries and regions. their influence on incomes and expenses of local budgets depending on the economic conjuncture and the state of the manpower market is considered. direct and indirect effects are taken in to account, the latter resulting from additional amount of services due to increase of total wages in the local mining industry. a formula for budget increment because of the new production including the indirect effects is derived. particular attention is paid to the problem of selecting the level of completeness of the stock extraction in the project. economic and ecological consequences of the latter are considered in detail. the problem rests to reveal significance of impact of expected mineral industry activities on different sectors of the local economy, for which probably the statistical method by toda and yamamoto may be applied. keywords: solid minerals, mining projects, budget, tax revenue, economic effects, social effects. 1. introduction the problem of regional economic and social effects caused by development and expansion of mining production as well as by its reduction is important for all mining countries. many factors have universal significance and are to some extent characteristic for the development processes of other branches of heavy industry, but their combination has significant national and even regional specifics. in this regard, the vast majority of studies are focused on the problems of a specific country and a particular region, and generalizing works that could formulate patterns for the world as a whole are absent. e.g., in the paper [7], despite its title, only the case of canadian province british columbia is studied and its fiscal system is addressed; nevertheless, some way of quantitative assessment of dependency between local mining industry activities and regional economic development for more general use is proposed. the lack of universality is caused by significant differences in legal system, culture of everyday life, territorial distribution and mobility of population of involved countries and regions. despite that, conclusions of researches may be useful for analysis of situations lying far beyond the scope of the papers. 145 a.m.valuev, m.a.lozinskaya copyright ©2018 assa. adv. in systems science and appl. (2018) the impact of rise or decline of local mining industry to most economical sectors of regional economy of different countries (canada, poland, turkey, sweden) is investigated in [1, 3, 4, 7]. most effects are seen via changes of employment in main sectors of regional economy. economic and social effects of migration caused by mining industry development in peripheral areas “characterized by remoteness, a scattered population, and limited access to infrastructure” [4] are studied in [4, 9, 11]. it must be emphasized that such conditions are typical for russian mining industry as well to most other large northern countries. the paper [10] as well as many other ones concentrates on development of knowledgeintensive service activities directly serving australian mining sector. most principal features of impact of mining industry activities on public health that are typical not only to australia are studied in [8]. taking into account results of the above mentioned (and other) researches, we restrict our attention to the case of russian mining industry. russian mining industry, on the one hand, is a very important part of the economy, and on the other, is extremely unevenly distributed. enterprises that do not belong themselves to the leading enterprises of the national mining industry, nevertheless, can make a very significant contribution to the regional economy and social sphere. when selecting and approving mining projects for solid mineral deposits (smd), the interests of the state-proprietor of subsoil (subsoil owner) and subsoil user organization are taken into account [14, 18]. the significance of a particular mining enterprise being designed at the federal and regional levels is not the same, thus it is not possible to speak of a single state interest in connection with the mining project. in some cases, the national interest dominates, because of which local interests may not be sufficiently taken into account when selecting a project, for local enterprises it is vice versa. when considering the project variant, one should try to take into account all projected revenues to budgets, as well as budget expenditures (for social support measures, environmental protection). the indicators that determine the commercial efficiency of a mining project are basically the same as for other industries (adjusted for the more pronounced random nature of their estimates), but the national and regional economic effects are much more diverse and much less accountable. thus, in [14, p. 235], in addition to tax revenues, payments for the use of subsoil and customs duties received directly from the subsoil user, “additional tax revenue (or losses) from third parties, due to the impact of the project on their financial situation”. in this article, an attempt is made to reveal more details of the situation in question and to take into consideration, in addition to revenues, also budget expenditures, as well as indirect economic effects. 2. general characteristics of regional economic and social effects of progects of industrial enterprises creation emergence of the new production in itself creates several principal effects for the region in which it is located. if its products are consumed mainly within the region (which is typical, for example, for the production of steam coal that provides local power stations and population needs as well as for construction materials quarries), then its consumers pay to the new manufacturer instead of the former suppliers (presumably from other regions). in addition to increasing the profitability of production for consumers, which justifies shift to new suppliers, the effect of the latter is that the payment for the products they supply remains in the region. as a result, this leads to an increase in the profits of regional producers and the total wages of workers. for some transport enterprises, on the contrary, owners loose profits and employees their salary; however these losses are obviously much less than the increase for consumers and new manufacturers. regional economic and social effects of the implementation of projects for mining 146 copyright ©2018 assa. adv. in systems science and appl. (2018) if the produced products are mainly intended for consumers outside the region, then the income of the created enterprise is added to the income of the regional economy. in this case, additional income can also be received by transport companies in the region, if they are engaged in the delivery of these products. at the same time, a number of industries, including mining enterprises, can have a negative impact on the environment, mainly through pollution. penalties that companies pay for identified cases of pollution cannot always compensate this negative effect. the consequences of pollution, first of all, are the increase in the morbidity, which requires additional investments in the health care system. at the same time, only the maintenance of the health of the employees of polluting enterprises can be taken on the responsibility of the company and carried out at the expense of its revenues. 3. the main aspects of the implementation of mining projects taking into account the economic and social environment before considering the issue of comparing alternatives for a specific mining company’s project taking into account indirect effects, we note the following fundamental aspects of the impact of such enterprises activities on the region. 3.1. market for sales of the new enterprise production and its relations with project indicators production at the new enterprise can either compensate for the outgoing capacities of other mining enterprises, or create the opportunity to replace imported raw materials (but thereby reduce the burden on transport infrastructure and employment in transport), or serve as a means for creating new industries in the region (supply of raw materials or energy) or in other regions, or, at last, for exports increase. thus, for projects in question, it is necessary to take into account their effect on the functioning or development of industry and transport, and in the case of exports, on payment of the export duty. industrial production (or power generation) on a regional scale can depend heavily on the extraction of a particular raw material at a particular location due to high transport expenses, or, on the contrary, it may be economically acceptable to replace this production with imported raw materials or substitute raw materials or energy transfer from other regions. in the latter case, the possible volume of sales in the local market (and beyond it, indeed) is closely related to the offer price. project alternatives differ not only in terms of the estimated level of prime cost, and consequently, in the level of prices of sales of products, but also in terms of the achieved quality values of products. it would be more accurate in many cases to suppose the relationship between cost and the main (or cumulative) quality indicator. the latter refers to cases where the use of raw materials is permissible in a fairly wide range of quality, but the efficiency of the production usage depends on it. this is typical, in particular, for power stations generating energy by coal combustion. thus, project alternatives differ in economic indicators not only for the mining enterprise itself, but also for consumers of its products, including local consumers. with regard to the objectives of maintaining existing industries or energy companies, one should take into account the sensitivity of their profits and other elements of the tax base to the costs of raw materials and to assess the additional effect as the difference between budget payments for various project options, provided that the raw material in the corresponding project is purchased at the projected enterprise, and the missing volume — on the external market beyond the region. in the case when the supply of raw materials from other sources is unprofitable, one 147 a.m.valuev, m.a.lozinskaya copyright ©2018 assa. adv. in systems science and appl. (2018) must take into account the sensitivity of the profit of the consumer of raw materials to the volume of production. similarly, the contribution of a mining enterprise to the economy of the region (and even of the country) needs to be assessed when its creation is a prerequisite for the creation of new industries. 3.2. significance of transport and energy transfer infrastructure for most types of mineral products, significant physical volumes are typical, and as a result, its transportation creates a considerable, sometimes the main load of transport system of the region. handling this load may impose restrictions on the economic activities of the region. the existing road network and transport fleet provides a certain amount of transport work, and its increase requires investments, transport infrastructure development demanding sufficient time for its expansion. if the enterprise being created is sufficiently profitable, it can afford, in case of a shortage of transport resources, to pay a higher price for transport services. but in this way transport needs for other, less profitable activities will not be fully met. in this case, the costs of the regional transport system development for satisfaction of requirements for the transportation of the products of the company being established or expanded should be attributed to its account. on the contrary, in the case of a significant reserve of the transport infrastructure capacity, investments in its development are not required, but transport fleet increase should be provided at the expense of the created enterprise except for the situation of prevalence of transport services by interregional organizations. the most important processes of mining, namely excavation, conveyor transportation and processes of primary processing (crushing, sorting) are energy-intensive and carry a high load on the energy system of the region. the latter is especially important for energy-deficient regions and can have a negative effect on other economic activities. 3.3. ecological aspects the environmental effects of mining of smd are very diverse and may not always be expressed in monetary terms. among other things, the territory of a mining allotment is withdrawn from other types of activities, including agriculture and forestry. obtaining a land tax does not necessarily compensate for the loss from the shortfall in payments to the budget from other commercial activities that have not taken place on this territory. on the other hand, a significant part of mining enterprises is located where there are no conditions for efficient agriculture. unfavorable environmental effects can have varying degrees of significance. for example, in the pavlodar region of kazakhstan, where the largest ekibastuz coal deposit is located, there is also a fairly developed grain farming, but not in the vicinity of the field itself. therefore, the pollution of soils by coal and rock dust, which dissipates from the faces and dumps, does not have a significant negative economic effect. quite the reverse situation is in the zone of the kursk magnetic anomaly, abundant with fertile black soil. unfavorable ecological environment can affect not only other types of economic activity, but also the health status of the population. consequently, nature protection measures have different costs depending on the state of the local economy, especially agriculture and forestry, and the resettlement of residents. according to the authors of [20], the existing penalties do not fully cover the costs of compensating for the negative impact of industrial activities on the state of the environment and suggest a formula for correcting the grp (gross regional product) taking into account the following:  costs for wastewater treatment in accordance with the cost of cleaning; regional economic and social effects of the implementation of projects for mining 148 copyright ©2018 assa. adv. in systems science and appl. (2018)  the cost of greenhouse gas emissions (co2 equivalent) in accordance with prices in the carbon market;  costs for removal and disposal of production and consumption wastes in accordance with the fee for negative impacts. but the above mentioned factors do not cover the entire pollution cost. perhaps the most significant impact of them is on the health of the population. it should be noted that in some cases, employees of an enterprise exposed to harmful working conditions and the population not directly related to production, but experiencing its negative impact, are in an unequal position. the mining company is to a certain extent interested in maintaining the health and physical fitness of its employees and, at the expense of its revenues, can pay for recreational activities and other health measures, and, if necessary, also treatment. at the same time, it will hardly be engaged in charity with respect to the entire population of the territory. a typical example of this is the main subdivision of pjsc mmc norilsk nickel its polar division with a center in norilsk. by common belief, the environmental situation in norilsk is very unfavorable. meanwhile, in respect of the employees, there is a whole system of measures. employees are provided with various social guarantees, benefits and compensations aimed at health improvement, treatment and recreation on russian and foreign resorts [12]. as for the rest of the population, additional health care costs, along with the mandatory medical insurance fund, must be paid by the local budget. therefore, regional interests should be taken into account in the form of a requirement to implement a sufficiently high level of environment protection measures in the project. however, this level may be the result of a compromise between the requirements of maintaining a favorable ecological environment and the economic interests of the regional budget and the population in obtaining additional revenues as a result of the projected enterprise. 4. social importance of a mining enterprise an enterprise or a branch of industry that is a consumer of a particular type of raw material may have high social importance being the main (city-forming) enterprise in some territory. in this case, the raw material supply of such production becomes an important social and economic task. the region should seek the adoption of a project for a mining enterprise that maximizes the volume of break-even production in a supported enterprise and offers a compromise option for the selling price of this raw material. regional labor market, including occupations related to mining operations, may have a different ratio between supply and demand, as established it the level of wages. different versions of the project, the differences in volume and profitability differ, and what wages they can offer to employees. but it must be estimated from the actual level in the region — the idea put forward by f. w. taylor who was one of the founders of the management theory [16]. 5. the effect of completeness of deposit extraction as it was noted in [2], one of the key parameters distinguishing project variants proposed for selection and approval is the completeness of extraction of balance reserves, in other words, the percentage of their losses during extraction. in the same works, a fact is noted that seems paradoxical for authors, that in typical cases, the most profitable project variant for both the subsoil user and owner is the variant with the highest loss level. the authors of [2] link this fact to insufficient consideration of budget revenues not only directly from the project (i.e., paid by the subsoil user), but also “in connection with the project” (paid by consumers of mineral raw 149 a.m.valuev, m.a.lozinskaya copyright ©2018 assa. adv. in systems science and appl. (2018) materials). we accept this position, but we do not absolutize the importance of the completeness of the extraction of mineral reserves in the light of the side effects discussed in this paper. the main factors associated with the completeness of extraction of reserves: 1. the volume of production and sales. 2. the cost of extraction and primary processing. 3. the effect on cost efficiency among consumers of raw materials. 4. the need for workers and their level of payment. the increase in the completeness of the extraction of reserves either requires an increase in the degree of selectivity of their development, which leads to a complication of technology, poor performance of mining machinery (or the use of less productive and thus less economical machines), or leads to an increase in the dilution of mined minerals. an example of the first case is the project of using layer-by-layer cutters instead of singlebucket excavators when working out complex-structural coal deposits in eastern siberia. clogging of coal with rock and loss of coal in this case, indeed, decrease, but the cost of production also increases. it should be noted that the increase in the completeness of the extraction of reserves at the expense of their dilution has an ambiguous effect on the total profit across the entire production chain, and thus on budget payments. when burning high-ash coal, the output of electricity per unit of calorific value of fuel falls, so to ensure the power generation process it is necessary to add fuel oil to the furnace, which increases the cost of production, which generally depends on the technologies of fuel combustion [6, 15, 19]. the cost of ash removal influences the growth of the cost price. diluted ore not only increases the cost of its enrichment, but also cause losses in the process of enrichment, so the concentrate amount may be even smaller than when enriching a smaller amount of unsoiled ore. therefore, the loss of a certain portion of the reserves, accompanied by an increase in the amount of the final product, is beneficial both for producers and for the budget. in the case of a certain reduction in the volume of the final product, but a significant reduction in the cost of production along the entire production chain, the total effect depends on the correlation between economic indicators. 6. direct and indirect relationship between project indicators and budget revenues and expenditures alternatives associated, in particular, with the different completeness of the extraction of balance reserves (and, consequently, with the different quantity and quality of the produced raw materials) can be viewed from several viewpoints, and the role of the factors listed below depends on the type of mineral, the place of extraction, the state of the budget and the labor market at both the federal and the local scale. the mining enterprise, depending on the project variant, employs a different number of workers in extraction, primary processing (sorting, enrichment) and transportation to its consumers (including transportation within the region). in addition to the main payments to the budget from the subsoil user (the tax on the extraction of minerals and the tax on profit), their employees also pay income tax. the total wage non-linearly depends on the number of hired employees. with an excess of labor supply, an increase in the number of employees, together with a decrease in the enterprise’s income per employee, entails a decrease in wages, and in the case of a deficit, on the contrary, the employment of new workers becomes possible only with an increase in wages, even if the income per employee decreases. the change in the unemployment rate changes the costs of benefits, retraining, housing subsidies and other measures of social support. regional economic and social effects of the implementation of projects for mining 150 copyright ©2018 assa. adv. in systems science and appl. (2018) the social consequences of the adoption of a particular design decision take place, first of all, at the regional level. it was noted above that the variants differ in the different number (and composition) of the labor force involved in the development of the deposit and the accompanying and related business processes. the increase in the aggregate wage of these workers makes it possible to develop the services sector to “absorb” a part of this increase, which in turn increases payments to the budget from the services sector. such an indirect effect can be identified for any production, not just mining. the state statistics provide publicly available data on the structure of the nation-wide gross product. obviously, the same data can be obtained for the regions. among the highlighted classification headings, along with mining, processing industries, construction, as well as various categories of services. in turn, the main ways of using the gross domestic product (gdp) are assessed quantitatively. data from the source “on production and use of (gdp) for 2016” [13] that are of the greatest interest to us, are shown in table 6.1. table 6.1. data on the structure of the gnp of the russian federation in 2016 components of gdp in basic prices share (year), % share for quarters, % i ii iii iv agriculture, hunting and forestry 4.5 2.3 3.4 7.2 4.5 fishery, fish farming 0.3 0.4 0.2 0.3 0.2 mining 9.4 8.6 10.1 9.7 9.1 manufacturing industries 13.7 12.5 14.1 13.6 14.4 generation and distribution of electricity, gas and water 3.1 3.9 2.7 2.5 3.4 construction 6.2 4.5 5.5 6.4 7.8 wholesale and retail trade; repair of motor vehicles, motorcycles, household goods and personal items 15.9 16.9 15.9 15.3 16.0 hotels and restaurants 0.8 0.8 0.9 0.9 0.8 transport and communication 7.8 8.4 8.1 7.8 7.2 health and social services 3.8 3.9 3.8 3.7 3.7 other communal, social and personal services 1.7 1.8 1.8 1.6 1.7 other important data from [13] are: as to the structure of gdp use: final consumption expenditure as a whole it is 76.6% and 55.8% for consumption of households; as to the structure of gdp by source of income: compensation of employees (including labor remuneration and mixed incomes not observed by direct statistical methods) is 46.7%. as it can be seen, the share of final consumption varies significantly even within a year. analysis of the correlation between such indicators allows to predict the change in the volumes of various types of services, accompanying changes in the volume of industrial activity and total wages in this area. according to the budget code of rf [15], taxes received by the regional budget directly from the activities of a mining enterprise are: 151 a.m.valuev, m.a.lozinskaya copyright ©2018 assa. adv. in systems science and appl. (2018) 1) the tax on the extraction of minerals (for general-distributed solid minerals and natural diamonds — 100% of the amount of tax, for other solid minerals — 60%); 2) the tax on profit of organizations — 18% of its value; 3) the tax on income of individuals (company employees) at a standard of 85%; 4) the tax on property of organizations — 100%; 5) transport tax. an increase in the total wages of workers in both the created and consuming and servicing organizations will be spent in large part within the region. to do this, the total volume of services (in particular, trade, which, obviously, in the amount of grp, as well as in the amount of gnp is a very significant share). in turn, some of these revenues turn into wages of organizations and individual private entrepreneurs working in the service sector. let the increment of the total wages in the non-consumer sector be amounted to w1. on the basis of statistical processing of the data in the form of linear regression, we assume that the elasticity of local consumption with respect to wage k12w1. in turn, we will assume that for workers in the services sector (the consumer sector) the share of their wages spent on services is k22 and that a unit of income in the consumer sector yields the amount of profit kp2 and the amount of wages kw2. to define the unknown increase in the volume of services in monetary terms, caused by w1, we employ the balance of revenues in the services sector be expressed by the ratio kw2(k12 w1+k22w2)= w2 (6.1) thus, we obtain for the increment of the total wage the formula w2=kw2k12w1/(1-kw2k22) (6.2) in addition, the increment in profit in the services sector will be p2=kp2(w1+w2). (6.3) this means that the cumulative increase in the regional budget revenues from the profit tax of the listed organizations of the non-consumer sector and all consumer sector organizations, as well as the personal income tax of workers of these organizations and individual private entrepreneurs will be 0.18(p1+p2)+0.130.15(w1+w2). (6.4) this is the main budget effect of the new production, which will be different for different project variants. here, the effect of increased consumption by the owners of the organizations in question is not taken into account, which is more difficult to assess, in particular, because some of them do not live permanently in the region. but in any case, it can only increase the revenues of the regional budget. other effects, such as a change in the cost of social support due to a possible reduction in the level of unemployment or to eliminate the effects of environmental pollution, are less predictable and require special study of the current factors. this is the main budget effect of the new production, which will be different for different project variants. here, the effect of increased consumption by the owners of the organizations in question is not taken into account, which is more difficult to assess, in particular, because some of them do not live permanently in the region. but in any case, it can only increase the revenues of the regional budget. 7. the time factor in the assessment of project options regional economic and social effects of the implementation of projects for mining 152 copyright ©2018 assa. adv. in systems science and appl. (2018) all that was said above did not take into account the time factor. as a general rule, in order to determine the cumulative effect for a subsoil user, the company’s estimated profit is summed by years, taking into account discounting; the same is done with respect to budget revenue. the net discounted income expresses (approximately, due to fluctuations in the economic conjuncture and technological changes) the real economic interest of the entrepreneur and the investor. for a state or region that, in general, spends the budget not for the economic benefit and is not a creditor even in the case of a budget surplus (although potentially it can be a creditor and even an investor), the goal is rather the guaranteed receipt of budget revenue as permanently (or, better, in everincreasing volumes). on the other hand, not being a creditor, the region is more often a debtor who is interested, if possible, to pay off loans more quickly. this justifies the use of the traditional formula of net discounted income for the project's budget contribution, but with a lower discount rate than for the subsoil user, because the state or a region usually takes a loan at a lower interest rate than the business. 8. conclusion the proposed approach yields the preliminary estimation of main regional economic effects, direct or indirect, that may result from a certain project for the development of a smd. it does not, however, reveal significance of impact of expected mineral industry activities on different sectors of the local economy. with the use of more sophisticated statistical methods, such as the method by toda and yamamoto [17], it looks likely that the most significant effects may be found in advance on the base of the study of analogs. in that case it would be possible to explore these effects in more detail for more perfect decision making. acknowledgements the work was partially supported by the ministry of education and science of the russian federation within the framework of the basic part of the state task no. 2014/113 for nust “misis” (project no. 952). we thank the members of the organizational committee of mlsd’2017 for their kind invitation to contribute the paper presented at the conference to the journal as well as mikhail goubko, associate editor-in-chief, and andrey shevlyakov, the secretary of advances in systems science and applications, for technical assistance. references [1] aksoy, e. (2013). relationships between employment and growth from industrial perspective by considering employment incentives: the case of turkey, international journal of economics and financial issues, 3(1), 74-86. [2] ashikhmin, a. a. & sytnik, ju. v. (2013). mnogokriterial'naja jekonomicheskaja ocenka i vybor variantov realizacii proektov na razrabotku mestorozhdenij tverdyh poleznyh iskopaemyh. [multicriteria economic evaluation and choice of options for the implementation of projects for the development of deposits of solid minerals]. racional'noe osvoenie nedr, 2, 15-19, [in russian]. [3] baranyai, n. & lux, g. (2014). upper silesia: the revival of a traditional industrial region in poland. regional statistics, 4(2), 126-144. 153 a.m.valuev, m.a.lozinskaya copyright ©2018 assa. adv. in systems science and appl. (2018) [4] brown, t. c., bankston, w. b., forsyth, c. j. & berthelot, e. r. (2011). qualifying the boom-bust paradigm: an examination of the off-shore oil and gas industry. sociology mind, 1(3), 96–104. [5] budget code of russian federation. [6] geisbrecht, r. & dipietro, p. (2009). evaluating options for us coal fired power plants in the face of uncertainties and greenhouse gas caps: the economics of refurbishing, retrofitting, and repowering. energy procedia, 1(1), 4347–4354. [7] gunton, t. (2007). natural resources and regional development: an assessment of dependency and comparative advantage paradigms. economic geography, 79(1), 67-94. [8] kinnear, s., kabir, z., mann, j., & bricknell, l. (2013). the need to measure and manage the cumulative impacts of resource development on public health: an australian perspective. in current topics in public health, al. j. rodriguez-morales, ed., rijeka (croatia): intech, 125–148. [9] knobblock, e. a. (2013). organizational changes and employment shifts in the mining industry: toward a new understanding of resource-based economies in peripheral areas. journal of rural and community development, 8(1). [10] martinez-fernandez, c. (2010). knowledge-intensive service activities in the success of the australian mining industry. the service industries journal, 30(1), 55-70. [11] martinez-fernandez, c., wu, c. t., schatz, l. k., taira, n. & vargas-hernández, j. g. (2013). the shrinking mining city: urban dynamics and contested territory. international journal of urban and regional research, 36(2), 245–260. [12] official site of ojsc mmc norilsk nickel, http://www.nornik.ru. [13] on production and use of gross domestic product (gdp) for 2016, http://www.gks.ru/bgd/free/b04_03/isswww.exe/stg/d02/64.htm. [14] shestakov, v. a. (1995). proektirovanie gornykh predpriyatii [design of mining enterprises]. textbook for universities. moscow: izdatel'stvo moskovskogo gosudarstvennogo gornogo universiteta, 508, [in russian]. [15] stenin, v. a. (2013). sravnenie jekonomichnosti teplovyh jelektrostancij [comparison of the economics of thermal power plants]. nauchnoe obozrenie, 3, 182, [in russian]. [16] taylor, f. w. (1911). the principles of scientific management, new york; london: harper, 144. [17] toda, h. y. & yamamoto, t. (1991). statistical inference in vector autoregressive with possibly integrated processes. journal of econometrics, 66(1), 225-250. [18] trubetskoi, k.n., krasnyanskii, g.l. & khronin, v.v. (2001). proektirovanie kar'erov [design of quarries], textbook for universities, vol. 1. moscow: izdatel'stvo akademii gornykh nauk, 519, [in russian]. [19] wang, w., jizhe, l., deliang, z., zhongwei, l. & can, c. (2012). variable-speed technology used in power plants for better plant economics and grid stability. energy, 45(1), 588594. [20] zabelina, a. & klevakina, e.a. (2011). otsenka ekologicheskikh zatrat v proizvedennom valovom regional'nom produkte [estimation of environmental costs in the produced gross regional product]. region: ekonomika i sotsiologiya, 2, 223–232, [in russian]. adv syst sci appl 2017; 17(2); 1-13 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/491 copyright ©2017 assa. adv. in systems science and appl. (2017) new analytics of international relations: system forecast of cold war’s outcomes victor a. svetlov, nikolay m. sidorov 1 1) emperor alexander i st. petersburg state transport university, 190031, 9 moskovsky ave., st. petersburg, russian federation e-mail: victsvetlov@yandex.ru abstract. systemic forecasts of international relations evolution for quite a long time were quite a rare phenomenon. the main reason for this is the lack of independent of authors' ideological and political predilections reliable analysts, and this fact determines relevance of the current study. the main goal of the article is to develop new analytics that allows prediction of long-term trends in the evolution of international relations world system. therefore, the algebra of relations and the corresponding section of predicate logic are used. the authors proved sixteen basic theorems on the properties of the world system. as the initial opposition, a pair of relations "dependence-independence" was chosen. the empirical conditions of the current state of affairs make it possible from the outset to exclude from the analysis the state of independence of states as a long-term factor of international politics. it was established that the world system of international relations can be strictly in one of three states – conflict, synergistic or antagonistic. the authors also carried out the forecast of the world states system dynamics after the end of the cold war. in regards with impossibility of achieving by the international relations world system in the next thirty years any of two possible attractors – states of synergism and antagonism, predicts – its stable oscillation between these points of stability until at least the middle of the nineteenth century is forecasted. in practice, this means, depending on the direction of the trend, the emergence of a variety of waves of instability, primarily in the field of international security keywords: relational algebra, conflict, synergism, antagonism, one-pole state, double-pole state. 1. introduction problem statement. in the late 20th century the warsaw treaty organization dissolved and then the soviet union collapsed. the cold war ended, and the world system’s bipolar division into two antagonistic military and political blocs, established after the world war ii, disappeared. the end of the age of the bipolar division of the world system makes the search for common regularities of international relations’ structural changes relevant more than ever. without theoretical solution of this problem, including its mathematical simulation, it’s impossible to comprehend, for example, what trend currently dominates and what configuration of relations between states is most probable in the nearest future. international relations play an important role in the implementation of the foreign policy of states, since they contribute to solving many economic, environmental issues, as well as issues related to the settlement and prevention of conflicts. regional issues which may be related to a military conflict in a particular territory and to affect the interests of many states are also resolved at the international level. international relations are also viewed as human activity where individuals from more than a single state interact individually or in groups. international relations can be also presented as interaction between two or more states, and foreign policy as an external action of a nation that proves the relevance of the current topic [1]. there is very little agreement among international relations specialists about estimation of the most probable structure of international relations after the cold war’s end. huntington s. 2 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) distinguished the following possible configurations based on the analysis of most authoritative research paradigms [2]. • there will be one relatively harmonious world. – fukuyama’s forecast [3]. • despite inevitable economic interaction, the cultural opposition of west and east will never disappear. – northrop’s forecast [4]. • the world system will split into many autonomous nation states, each of them will strive for survival, entering alliances with other states or augmenting its potential independently. – waltz’s forecast [5]. • the world system will sink into an utter chaos, national states will dissolve, tribal and ethnic conflicts will escalate, terrorism will become a common phenomenon, international criminal groups will appear. – brzezinski’s forecast [4]. • a polycivilizational world divided into two global groups will appear – west civilization and a small array of non-western civilizations not related to each other, in total about 7-8. national states will remain leading players on the world stage, but conflicts between them will be caused by their belonging to two specified groups. – huntington’s forecast [2]. each of the forecasts distinguished by huntington, including his own one, predicts a certain result of rebuilding the system of states’ international relations started after the cold war’s end. there is no doubt any one of them has its more or less ample grounds. however, it’s crucial that none of them is based on the system nature of changing international relations. this circumstance explains analysts’ dispersed opinions on the forming historical trend and the world system’s future configuration. it is commonly known that the choice of forecast point has a substantial effect on its result. at this point we associate ourselves with american historian and theoretician lipschutz r. who clearly demonstrated how the end of the cold war led to the paralysis of an idea of national state and collapse of the previous system of international safety [6]. the study of issues related to international relations was a subject of high interest for scientists since ancient times: e.h. carr, g. morgenthau, l. waltz, g. clark, j.s. nye, r. cohen, h. bull, a. rapoport, j. burton, e. haas, a. walfers, k. wright, o. holsty. a lot of modern scientists also deeply study the issues of international relations: r. powell, s.a. lantsov, m.a. muntean, b. buzan, j.b. mannheim, r.k. rich, m.a. khrustalev, f. moreau-defarge, m. kaplan and others. researchers of the concepts of dependent development (a. cordova, o. zunkel, f.kh. cardoso, etc.) believe that the main reason for most countries of the world is the "covert" use of the poorly economically developed countries by more advanced ones [7]. the study of international relations by k. wright is based on the fact that they represent a body of knowledge with the help of which it is possible to assess and control the relationship between states. a. kaminsky considers this from the position that international relations are decisive factors, levers, mechanisms of mutual relations and finds regularities and randomness in these relationships. concept headings. taking into account a total absence of analogues of mathematical solution to the stated problem in literature, new analytics of system forecast of the world system’s dynamics of the international relations after the cold war’s end is justified below. to this end a special discipline of predicate logic, often named logic of (binary) relations, is used. necessary theorems are stated and proved. our main hypothesis is that the world system cannot have more than two points of stability represented by its synergetic and antagonistic states. the system dynamics is surprisingly simple: it either reaches one of the points and remains there for a long time or oscillates between them [8, 9]. strange as it may seem, the suggested system forecast of cold war’s outcomes doesn’t depend on the content of relations between states recorded empirically. it takes into account only structural and dynamic properties of the relations themselves [10]. that's its advantage and at the same time its drawback. the forecast’s – as any mathematical model’s – strength is a high new analytics of international relations: system forecast of cold war’s outcomes 3 copyright ©2017 assa. adv. in systems science and appl. (2017) degree of credibility. the weakness is that its general conclusions require particular historical, economic and political details while explaining and forecasting specific events. 2. materials and methods let ws = (a, b, c, …) designates the world system of states denoted with symbols a, b, c, … . first of allб we’re interested in possible relations between states. to this end the framework of relational algebra is used. assuming this comment, term ws can be interpreted as a world system of international relations [11]. the analysis’ long-run objective is statement of general laws the change of relations between states follow and determination of dominating trend after the cold war’s end. let x, y, z, … – individual variables, running the elements ws; (х), (eх) – generality and existential quantifiers, respectively. let signs , , , ,  denote complementary operations (to complete relation), entailment, equivalence, multiplication and relation addition, respectively. let’s denote probability measure determined on the set of all subsets of the ws system with рr. let’s set positive (р), negative (n), conflict (с), relevant (r) and irrelevant (ir) relations (impact, dependence) on cartesian product ws  ws according to the following definitions. definition 1: state a has a positive impact on state b if and only if probability of existence of b providing existence of a is more than 0.5: рr(b/a)  0.5. definition 2: state a has a negative impact on state b if and only if probability of existence of b providing existence of a is less than 0.5: рr(b/a)  0.5. definition 3: state a has no impact on state b (а is related to b irrelevantly) if and only if probability of existence of b providing existence of a is 0.5: irаb = рr(b/a) = 0.5. definition 4: state a has an impact on state b (а is related to b relevantly) if and only if state а has a positive or negative impact on state b: rаb = раb  nаb. definition 5: state a is in conflict with state b if and only if a relates to b both positively and negatively: саb = раb  nаb. definition 6: states a and b are in synergetic dependence if and only if they’re both related to each other only positively. definition 7: states a and b are in antagonistic dependence if and only if they’re both related to each other only negatively. the categories of conflict, synergism and antagonism have a special role to play in the building of qualitative forecasts of system transformations. conflict denotes the inner cause of system change, synergism and antagonism – stable, although opposite outcomes of conflict solution, its points of stability (attractors). the outbreak of a system conflict indicates the beginning of a system change, synergism and antagonism – achieving steady state of stability by the system. a conflict with the course of time either transfers the system to a higher or lower level of conflict or transforms it into a conflict-free – synergetic or antagonistic state. the essence of synergism as a steady conflict-free system state reveals the following significant dynamic properties: (1) states, which are elements of a synergetic system, all together either progress or regress. (2) all synergetic systems with the course of time and continuous generation of energy from outside only increase their synergism and strive to remain, therefore, conflict-free. (3) a synergetic system ceases to exist only when strengthening or weakening of all or several states becomes incompatible with ensuring general synergism of its components. antagonism as a steady conflict-free system state is characterized by the following distinctive dynamic properties: 4 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) (1) the progress of some states within antagonistic systems always takes place at the expense of other states’ regress. this means that the one pole of antagonistic system prospers with the course of time, while the other one certainly comes down. (2) all antagonistic systems with the course of time and continuous generation of energy from outside only increase its antagonism and strive to remain, therefore, conflict-free. (3) an antagonistic system ceases to exist only when the degrading pole cannot stand its antagonism with the stronger pole anymore. let’s name positive, negative relations and their additions to complete relations initial. multiplicative unification of positive and negative relations with their additions can give rise to both relevant, i.e. positive or negative relations, and conflict and irrelevant relations as derivatives. all possible kinds of derivatives relations between states a and b based on their initial relations’ multiplication is summarized in table 1. table 1. all possible kinds of derivatives relations nab  nab pab сab pab pab nab irab let pl denotes a pole (alliance, bloc, coalition) of the world system of states composed by its elements based on the combinations of certain kind of relations with ws. definition 8: state a forms with state b a one-pole system if and only if a and b relate to each other only positively: plab = pab  nаb. the one-pole system’s extensive definition is as follows: plxy = (x)(y)[(x  y)  (pxy  nxy )]. it follows from def. 6 and def. 8 that one-pole system can be synergetic only. definition 9: state a forms with state b an antagonistic double-pole system if and only if a and b relate to each other only negatively: pla plb = nab  pаb. the double-pole system’s extensive definition is as follows: plx  ply = (x)(y)(z){[( x  y) & (x  z)]  [(nxy  pxy)  ((nxz  pxz pyz  nxz)  (nyz  pyz  pxz  nxz))]}. it follows from def. 7 and def. 9 that double-pole system can be antagonistic only, while each pole consists of synergetically interacting elements. relations introduced using def. 15 can be summed up and multiplied, forming more complex relations. for the purpose of the present paper it’s sufficient to determine the matrix of multiplication of the relations of various modalities (see table 2). table 2. relations of various modalities  p n c r ir p p n c r ir n n p c r ir c c c c c ir r r r c r ir ir ir ir ir ir ir for example, the following system of relations, despite even number of negative relations, is nonetheless conflict: new analytics of international relations: system forecast of cold war’s outcomes 5 copyright ©2017 assa. adv. in systems science and appl. (2017) pnnpppс = [pn]npppc = [nn]prpc = [pp]rpc = [pr]pc = [rp]c = rc = c it follows from table 2 that the relation of irrelevance ir is the most stable: being multiplied by any relation, it always remains. it can be considered a peculiar null (dominant) relation in the logic of studied relations between states. gaining complete independence as distinct from other types of relations offers a solution to all international problems: a state can have no difficulties in relations with neighbors, if only it exists independently of them. the second most stable is с conflict relation. it dominates all types of relations, except for irrelevance relation ir. conflict stability before all other dependence relations confirms a worldly wisdom: it’s easy to come into conflict, but difficult to come out of it [12]. the relation of relevant (positive or negative) relationship r is the next in the hierarchy of stability. this combined relation dominates only positive and only negative relations, giving up in stability only to irrelevance and conflict relations. the relationships of positive relation р and relevant relation r are reflexive, symmetrical and transitive, i.e. they’re a relation of equivalence. it is two types of relations (and only them) that create a basis for emergence of a stable system of international relations. strictly speaking, each of them is only necessary, but they can become adequate grounds for international stability if conditions are right. the relation of negative relation n is neither reflexive nor transitive, but it is symmetrical. the relations of irrelevant relation ir and conflict с are not reflexive, but symmetrical and transitive. these types of relations do not form equivalence classes and in no event can ensure uprising – let alone ensuring – of stable world order. it follows from the above that the problem of searching for stable architecture of the world system ws is mathematically comes to breakdown of all states into many non-crossing and jointly exhaustive classes (poles) offering equivalence. only two relations (from the considered above) have the equivalence property – the relations of relevance r and positive p relation. it means that each of them may become not only necessary but adequate grounds for dividing the world system ws into non-crossing and jointly exhaustive classes if conditions are right. this result is notable on its own, since it indicates relevance (positive or negative dependence) and positive dependence of states as two indispensable and, notably, only conditions of the international relations’ system stability [13]. the relations of positive (p) and negative (n) relevance represent peculiar atoms, various combinations of which give rise to all the rest types of relations (compare tables 1 and 3). table 3. the relations of positive (p) and negative (n) relevance u = all relations r = relevant ir = irrelevant (pn) conflict conflict-free c = (pn) c = ( pn ) ( p n ) let’s state the main theorems of the logic of international relations necessary for system forecast justification. however, it should be noted that in a nonformal sense the logic of relations between states is principally the same as the logic of interpersonal relations that are governed by four well-known rules of “the golden rule of morals” [1]: 1. the friend of my friend is my friend. 2. the friend of my enemy is my enemy. 3. the enemy of my friend is my enemy. 6 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) 4. the enemy of my enemy is my friend. тheorem 1: in the world system ws states are dependent on each other if and only if it’s false that they’re independent on each other: (x)(y) (rxy  irxy). proof: 1. (x)(y) rxy (assumption of direct proof) 2. (x)(y)   rxy (1) 3. (х)(y)   [(pr (y/x)  0.5)  (pr (y/x)  0.5)] (2, def. 4) 4. (х)(y)  [pr (y/x) = 0.5] (3, def. 3) 5. (х)(y) irxy (4, def. 3) 6. the proof of converse implication is analogous. qed according to theorem 1, dependent states form an equivalent class, the common feature (relevance of relations) of which is not typical for none independent state (see table 3). from this perspective, being independent means having neither positive nor negative, nor conflict relations with any other state of the world system ws. тheorem 2: in the world system ws dependent states are positively related to each other if and only if it’s false that they’re negatively related to each other: (х)(y) (рху  nxy). proof: 1. (х)(y) pxy (assumption of direct proof) 2. (х)(y)   pxy (1) 3. (х)(y)   [pr (y/x)  0.5] (2, def. 1) 4. (х)(y)  irxу   [pr (y/x)  0.5] (3, assumption of dependency) 5. (х)(y)  [pr (y/x)  0.5] (4, t 1) 6. (х)(y) nxy (4, def. 2) 7. the proof of converse implication is analogous. qed according to theorem 2, positively related states among dependent states form their own equivalent class, the common feature of which (positive relevance) is not typical for none negatively dependent state. this theorem states that positive dependence is not the only type of dependence. there are also such types as conflict and antagonism, including not only positive but also negative relations of states. тheorem 3: in the world system ws pairs of dependent states are conflict-free if and only if they’re related in each pair either only positively or only negatively: (х)(y) {(rxy   сxy)  [(рxy   nxy)  (рxy  nxy)]}. proof: 1. (х)(y) (rxy   сxy) (assumption of direct proof) 2. (х)(y) [(рxy  nxy)   (рxy  nxy)] (1, def. 4 and 5) 3. (х)(y) [(рxy  nxy)  (рxy  nxy)] (2) 4. (х)(y) [(рxy  nxy)  (рxy  nxy)] (3) 5. the proof of converse implication is analogous. qed according to theorem 3, two interdependent states form an elementary dynamic cycle, which is conflict-free in two cases: either two ways of cycle are positive (synergism occurrence) or they’re both negative (antagonism occurrence). it is obvious that pairwise conflict-free nature of states does not generally guarantees the conflict-free nature of the whole world system ws. тheorem 4: in the world system ws, which is in conflict, each state is in a negative selfreference: new analytics of international relations: system forecast of cold war’s outcomes 7 copyright ©2017 assa. adv. in systems science and appl. (2017) (х)(y) сxy  nxч. proof: 1. (х)(y) сxy (assumption of direct proof) 2. (х)(y) (рxy  nxy) (1, def. 5) 3. (х)(y) (рxy  nyх) (2, symmetry of relation nxу) 4. (х)(y) [(рxy  nyх)  nxх]( theorem of logic of relations) 5. (х)nxх (3, 4) qed theorem 4 indicates a required feature of the world system’s conflict state: each its element is a relation of negative converse relation with itself. it means that whatever measures a nation in such a state may take, it will only aggravate its situation, thereby, escalating the system conflict. тheorem 5: in the world system ws, which is conflict-free, each state is in a positive selfreference: (х) рхх. proof: (cpr – calculus of probability) 1. (ех)  рхх (assumption of indirect proof) 2. (ех) (nхх  irxx) (1) 3. (ех) [(pr (x/x)  0.5)  (pr (x)  0)] (1, 2, cp, def. 2) 4. (ех) [pr (x/x)  0.5] (3) 5. (ех) [pr (x)  0] (4) 6. (х) [is pr (x)  0, pr (x/x) = 1] (theorem cpr) 7. (х) pr (x/x) = 1 (5, 6) 8. contradiction (4, 7) 9. (х) рхх (1, 8). qed theorem 5 indicates a required feature of the world system’s conflict-free state: each its element should be in a relation of positive converse relation with itself. тheorem 5 expresses a peculiar principle of state (self) preservation. in order to prosper each state must be able to support itself, advocate its interests, maintain consistent relations with its friends and enemies. тheorem 6. the world system ws is conflict-free if and only if it has none negative relation: (х)(y) ( nxy   cxy). proof: 1. (х)(y) nxу (condition) 2. (х)(y) nуx (1, symmetry of relation nxу) 3. (х)(y) (nxу  nуx) (1, 2) 4. (х)(y) [(nxy  nуx)  pxх] (theorem of logic of relations) 5. (х)(y) pxх (3, 4) 6. (х) nxх (5, t 2) 7. (ех)(еy) cxy (assumption of indirect proof) 8. (ех)(еy) (nxy  рху) (2, def. 5) 9. (ех)(еy) [(nxy  рух)  nxх] (theorem of logic of relations) 10. (ех) nxx (8, 9) 11. contradiction (6, 10) 12. (х)(y)  cxy (7, 11) 13. the proof of converse implication is analogous. qed the absence of negative relations in the world system ws is an indispensable and sufficient conditions for its conflict-free environment. this statement is fair for the admitted above assumptions on positive and negative relations as primary ones for systems of any type. however, a deeper analysis shows that a conflict is possible even when a system has only positive relations, which vary in its impact (see [1]). 8 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) тheorem 7: the world system ws is conflict-free if and only if it has only positive relations: (х)(y) (рху   cxy). proof: follows from the combination of тheorem 2 and theorem 6. qed a system that has only positive relations is synergetic. it is known that synergism can have both positive and negative effects. in the former case the system with its elements progresses, in the latter – degrades. thus the system’s conflict-free state doesn’t guarantee its positive development trend. тheorem 8: the world system ws is conflict if and only if none of the states is independent on each other: (х)(y) (сху   irxy). proof: 1. (х)(y) сху (assumption of direct proof) 2. (х)(y) (nxy  рху) (1, def. 5) 3. (х)(y) [(nxy  рху)  (nxy  рху)] (theorem of logic of relations) 4. (х)(y) (nxy  рху) (2, 3) 5. (х)(y) rxy (4, def. 4) 6. (х)(y)  irxy (5, т1). qed according to theorem 8, a conflict is possible not only between dependent states. in other words, dependence is an indispensable, although not sufficient condition of conflict. dependence between states is potential of not only conflict-free but also conflict development. if several states become members of the same economic, political or war system, they not only benefit from mutual cooperation but also increase chances for conflicts between themselves. it explains why every union, stable as it may be conceived, dissolves sooner or later. тheorem 9: the world system ws is conflict free if it is consists of states dependent on each other: (х)(y) (irxy  сху). proof: theorem т9 is a contraposition to theorem т8 and, therefore, is equivalent to it. qed according to theorem 9, independence and conflict environment are incompatible system features. theorems 8 and 9 highlight state’s independence as the only guaranteed solution to any international conflict. however, this formula for establishing a stable world order can scarcely be put into practice worldwide. achieving complete independence by all states in the foreseeable future seems a utopian project because of the evident scarcity of resources necessary for progressive development. it means that strengthening of world economy’s globalization tendencies will certainly enhance the likelihood of international conflicts among its members. тheorem 10: if each state of the world system ws is in a positive self-reference, the system is conflict-free: (х)(y) (рхх   сху). proof: 1. (х)рхх (condition) 2. (ех)(еy) сху (assumption of indirect proof) 3. (ех)(еy) (nxy  рху) (3, def. 5) 4. (ех)(еy) (nxy  рух) (symmetry of relation рху) 5. (ех)(еy) [(nxy  рух)  nxх] (theorem of logic of relations) 6. (ех) nxх (4, 5) 7. (ех) рxх (6, т 2) 8. contradiction (1, 7) new analytics of international relations: system forecast of cold war’s outcomes 9 copyright ©2017 assa. adv. in systems science and appl. (2017) 9. (х)(y)  сху (2, 8). qed according to theorem 10, positive self-reference of each state is a sufficient feature of a conflict-free state of the world system ws. if we translate the “positive self-reference” term into the language of international relations, it means a positive effect of state’s self-regulation resulting from successful combination of system-wide and national interests. тheorem 11: the world system ws is conflict-free if and only if all states are dependent and each of them is in a positive self-reference: (х)(y) [(рхх  rху)   сху)]. proof: 1. (х)(y) (рхх  rху) (assumption) 2. (х)(y) [(рхх  rху)   сху] (т10) 3.  сху (1, 2) 4. the proof of converse implication is analogous. qed according to theorem 11, in case of dependence positive self-reference is not only a sufficient but at the same time necessary condition of conflict-free existence. collapses of coalitions always start when some of its members loose positive self-reference for one reason or another. for example, brexit took place when about a half of british population stopped to feel self-identification regarding membership in the eec. тheorem 12: if the world system of interdependent states ws has exactly one pole, it is conflict-free: (х)(y) (plxy   сху). proof: 1. (х)(y) plхy (assumption of direct proof) 2. (х)(y) (рху  nxy) (1, def. 6) 3. (х)(y) [(рху  nxy)   сху] (conclusion of theorem т6 and т7) 4. (х)(y) сху (2, 3). qed theorem 12 states that the absence of conflict in the community of dependent states is a necessary condition of one-pole and, therefore, synergetic systems. if the world economy achieved a complete and universally beneficial state of globalization for all ws states, and their political regimes and institutes corresponded to and supported it, the mentioned one-pole and synergetic word order would appear. in such a world order conflict would be impossible. тheorem 13: if the world system of interdependent states ws has exactly two poles, it is conflict-free: (х)(y) [(plx  ply)  сху]. proof: 1. (х)(y) (plх  ply) (assumption of direct proof) 2. (х)(y) (nху  pху) (1, def. 7) 3. (х)(y) (nxy  nxy) (2, т2) 4. (х)(y) (nxy  nxy)  pху (logic of relations theorem) 5. (х)(y) pxy (3, 4) 6. (х)(y) [pxy  (nxx  pxx  nyy  pyy)] (logic of relations theorem) 7. (х)(y) (nxx  pхx  nyy  pyy) (5, 6) 8. (х)(y) [(nxx  pxx  nyy  pyy)  сху] (conclusion of theorem т6 and т7) 9. (х)(y) сху (7, 8). qed theorem 13 is of important methodological value. it proves that antagonism, which everyone’s prone to identify with conflict, is actually one of its (in addition to synergism) 10 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) oppositions. like synergism, antagonism features high resistance and ceases to exist until one of antagonists dies out or stops fighting because of eclipse of powers and resources depletion. the cold war between the soviet union and usa and its allies that lasted a little more than 40 years and ended only with the collapse of the soviet union, can be a case in point, demonstrating antagonism stability. тheorem 14: if the world system of interdependent states ws is polarized into one or two poles, it is conflict-free: (х)(y) [(plxy  (plx ply))  сху] proof: follows from the combination of theorems 12 and 13. qed the meaning of theorem 14 is that the absence of conflict in the system of international relations of interdependent states can be a consequence of only two causes – its one-pole (synergetic) or double-pole (antagonistic) organization. тheorem 15: the world system of interdependent states ws is conflict-free if only it is polarized into one or two poles: (х)(y) [сху  (plxy  (plх ply))]. proof: 1. (х)(y) сху (condition) 2. (ех)(еy)  [plхy  (plх ply)] (assumption of indirect proof) 3. (ех)(еy) [ plхy   (plх ply)] (2) 4. (ех)(еy) [ (pхy  nху)   (nху   pхy)] (3) 5. (ех)(еy) [( pхy  nху )  (nху  pхy)] (4) 6. (ех)(еy) [(nхy  nху )  (рху  pхy)] (5, т 2) 7. (ех)(еy) (nхy  pхy) (6) 8. (ех)(еy) схy (7, def. 5) 9. contradiction (1, 8) 10. (х)(y) [plхy  (plх ply)] (2, 9). qed theorem 15 is converse to theorem 14. it establishes that if the system of international relations is conflict-free, it means it is either one-pole (synergetic) or double-pole (antagonistic). тheorem 16: the world system of interdependent states ws is conflict-free if and only if it’s polarized into one or two poles: (х)(y) [(plхy  (plх ply))  сху] proof: follows from the combination of theorems 14 and 15. qed the important meaning of theorem 16, which can be called central one for our forecast justification, is that it determines necessary and sufficient conditions of conflict-free development of the system of international relations – synergism within one pole or antagonism between two synergetic poles. all other combinations of relations between states are knowingly conflict, i.e. induce the system to move to one of the stated points of stability. 3. results and discussion forecast and its justification. let’s divide all system forecasts into qualitative and quantitative. qualitative forecasts do not depend on historical, political, economic and cultural factors and are more significant to this effect in the degree of community of their conclusions. the credibility degree of a qualitative system forecast depends on three conditions: (1) the choice of forecast base point; (2) identification of dynamics or the main trend of the system’s new analytics of international relations: system forecast of cold war’s outcomes 11 copyright ©2017 assa. adv. in systems science and appl. (2017) change; (3) calculation of most probable outcomes of the system dynamics in a chosen period of time. we choose the cold war’s end as base point of suggested forecast. the reasons are the following. from the system perspective, the cold war represented antagonism of two superpowers (together with their allies) – the soviet union and usa. the world system ws was divided into two poles and according to theorem 13 was conflict-free and, therefore, a steady point of stability. each superpower strongly tried to maintain power balance in the military, first of all, nuclear field, which gave rise to the stability effect. however, the logic of antagonism development appeared stern and one of the antagonists – ussr eventually ceased to exist. antagonism and the resulting division of the states’ world system ws into two poles disappeared. however, the following 15-year domination of usa didn’t make the world system ws more stable and predictable for one reason only: usa didn’t become a powerful source of synergism for the whole world system neither in economic nor in political and military terms. on the contrary, the usa strategy focused on its isolated domination led and has been leading till now to strengthening and expansion of regional wars, emergence of new centers and waves of terrorism, and destruction of traditional states. according to theorem 16, it means that the world system ws reached not the “history end” but a new system conflict and consequently uprise of whirling motion to a new point of stability. chaotic and therefore poorly predictable movement of the world system ws from the burnt-out antagonism of two superpowers to a new point of stability – the main trend of the modern history after the cold war’s end. it follows from theorem 16 that the world economy can have only two points of stability– association of all (or the major part) states into one synergetic system or their dissociation no longer on the ideological basis into two antagonistic poles. before discussing the chances of each of these outcomes, it’s reasonable to assess the resiliency of an idea popular in political and diplomatic quarters – the idea of “multipolar world”, “polycentric architecture of international relations” that appeared as counterbalance to the idea of usa unconditional domination. let’ consider the simplest example of multipolarity. let’s assume the world system ws consists of three states – russia, china, and usa. pure algebra shows that in this case four conflict-free and four conflict states are possible. conflict-free states, in their turn, include one synergetic (all three states are friends) and three antagonistic (any two countries confront the third one) states. therefore, the world system multipolarity is possible only as purely quantitative expansion of synergism or antagonism of its members with binding preservation of its pole number (one or two). otherwise, when it has no poles or there are two poles, the world system becomes conflict, the sole destiny of which is to drift to new points of stability. thus, our theory states that there are no any other conflict-free states of the world system ws except oneor double-pole ones. the appearance of configurations with three and more poles in a conflict-free world system or the lack of any poles as long-term and stable state is impossible on the purely formal grounds. if we take into account that negative and positive relations can vary in their impact (lovefriendship-sympathy; hatred-enmity-antipathy), there are the following new opportunities [1]. firstly, any pole of synergetic or antagonistic system without conflict occurrence can be divided into several subsets (subpoles), the members of which are positively related to each other more strongly than all the other members. relations of incompatibility of the members of different poles can also vary within antagonistic systems, but in this case from strongly negative to weakly negative. secondly, any synergetic or antagonistic system can coexist conflict-free with any other number of poles symmetrically independent on them. one-pole system of international relations is a synergetic system (follows from def. 6 and def. 8). if synergism as a universal tendency has no any strong opposition, then what is now called “global community” would quickly emerge. however, the major obstacle to achieving this outcome by the world system is not so much significant differences in economic, political, ethnic, 12 v.a. svetlov, n.m. sidorov copyright ©2017 assa. adv. in systems science and appl. (2017) religious and cultural conditions of different nations’ existence, sometimes coming to their total incompatibility, as the state’s need to defend its national interests, which always fatally counteracts the unconditional “victory” of panhuman interests [13]. in view of the above stated, achieving the state of universal synergism by the humanity appears an extremely unlikely outcome of the transformation of the world system of relations after the cold war’s end in the foreseeable future. a double-pole system of international relations is an antagonistic system (follows from def. 7 and def. 9). so far, today there are no any prerequisites of military, economic or cultural nature, which indicate the imminence of arising two opposing world centers of forces. the economic competition of china and usa will never result in worldwide economic antagonism because of their economic interdependence, i.e. it will always be local and temporal. the similar conclusion can be made in relation to the military competition of russia and usa. any tendency of the world system division into two new poles will always be opposed by the tendencies of global labour and capital market formation, scarcity of resources, and consequently inevitable mutual dependence of all countries. thus, antagonism as a point of stability and inevitable division of the world system into two poles now and in the foreseen future can be considered not less problematic solution to its conflict of the world system of relations than universal synergism [14]. in view of the foregoing, our forecast is based on the following system regularity. if neither of the two possible points of stability is reachable in a definite period of time for a changing system, then the system most likely starts steadily oscillating (fluctuating) between them within its limits. forecast: after the cold war’s end the world system of states ws reached a stage of sustained oscillation between two conflict-free and equally unreachable states – unipolarity (synergism) and bipolarity (antagonism). taking into account present tendencies counteracting to the achievement of each of the stated points of stability, this stage will last at least till the middle of the 21 st century. empirical characteristics of the oscillation of the world system of states in a given time period will be determined by particular correlation of two counter forces, forming the basic conflict of any system, – a tendency to universal association of states into one global community and a tendency to protection of their national interests [10]. 4 conclusion ● there are two stable prejudices among analysts and practitioners of international relations. the both of them are related to misestimation of the conflict essence and functions in the evolution of the world system of states. ● according to the first prejudice, conflict is a solely deconstructive state, and all disputes, as a rule, unfold around searching for efficient methods of its prevention or resolution. ● according to the second prejudice, antagonism and such its important kind as war, are identified with conflict or are considered its type. ● the both prejudices can be overcome with a new theoretical approach, called new analytics. conflict is considered the only condition of transformation of the world system of international relations between the points of stability (attractors); synergism and antagonism – two (and only two) steady conflict-free states. references [1] svetlov, v. a. (2013). vvedenie v edinuyu teoriyu analiza i razresheniya konfliktov [introduction to a unified theory of conflict analysis and solutions]. moscow, russia: librokom. [in russian] [2] huntington, s. (1996). the clash of civilizations and the remaking of world order , new york, ny: simon & schuster. new analytics of international relations: system forecast of cold war’s outcomes 13 copyright ©2017 assa. adv. in systems science and appl. (2017) [3] solomon, t. & steele, b. j. (2017). micro-moves in international relations theory, european j. of int. relations, 23(2), 267-291. [4] northrop, f., the meeting of east and west: an inquiry concerning world understanding. new york, ny: macmillan, 1947. [5] abraham, k. j. & abramson, y. (2017). a pragmatist vocation for international relations: the (global) public and its problems. european j. int. relations, 23(1), 26-48. [6] lipschutz, r. d. (2000). after authority. war, peace, and global politics in 21 st century. albany, ny: state university of new york press. [7] lebow, r. n. (2014). constructing cause in international relations. cambridge, u.k.: cambridge university press. [8] matti, j. (2017). new national organization of europe: nationalism and minority rights after the end of the cold war. int. relations, 31(1), 21-41. [9] wolff, j. & spanger, h.-j. (2017). the interaction of interests and norms in international democracy promotion. j. int. relations and development, 20(1), 80-107. [10] svetlov, v.a. (2012) konflict i evolucia. ot geneticheskih konfliktov k konfliktu pokolenii [conflict and evolution. from genetic conflicts to generation gap]. moscow, russia: librokom. [in russian] [11] glencross, a. (2015) from ‘doing history’ to thinking historically: historical consciousness across history and international relations. int. relations, 29(4), 413-433. [12] maslikov, v.a. (2015) aktualnie factory neustoichivisti obshestvennogo razvitia [actual factors of the instability of social development]. materialy afanasievskih chtenii, 1(13), 210-214. [in russian] [13] zadihailo, d. (2013) economic power in the context of the legal regulation of economic relations. j. the national academy of legal sciences of ukraine, 4(75), 163-171. [14] buyarov, d., kireev, a. & druzyaka, a. (2017) russian-chinese relations: from the decision of interstate problems to the strategic partnership. j. advanced research in law and economics, 6(4), 746-752. https://doi.org/10.14505//jarle.v6.4(14).03. advances in systems science and applications (2013) vol.13 no.1 37-52 static models of corruption in hierarchical systems andrey v. antonenko and guennady a. ougolnitsky southern federal university, russia abstract the principles of modeling and control of corruption in the hierarchical systems are formulated. a system of the theoretical static models of administrative and economic corruption as well as models of corruption in real estate development is built. the dependence of corrupted behavior on model parameters is investigated, and the analytical conditions in which corruption is not profitable for the agent or can be controlled by the principal for ensuring sustainable development of the system are received. keywords corruption, hierarchical systems, game theory, optimization, sustainable development 1 introduction the pioneering work on the mathematical modeling of corruption is susan roseackerman’s paper [1] with the consequent development of the proposed approach in her monograph [2]. the monograph specifies the ideas by gary becker [3] concerning the modeling of arbitrary crimes and punishments. later the static models of corruption were developed in two principal directions. the first one includes mathematical models that investigate corruption inside an organization (internal corruption) and between organizations (external corruption). within the framework of this class of models the problems of interaction of competence and corruption, non-optimal resource distribution, the principal’s impact to the model equilibrium and some other ones are studied. the paper [4] may be taken as an example in which the problems of federal resources distribution by bureaucrats and the forms of impact of the competence between both the bureaucrats and the agents on the bribery scope are investigated. models of corruption in tax authorities also belong to this class, for example, [5-6]. the second research direction is connected with modeling of corruption in social and political life (particularly the corruption patterns presented in the electoral process). models of the principal agent type that analyze such problems as anti-corruption incentives development, building of decision making rules that facilitate the indication of bribery facts, providing the hierarchical structure and control mechanisms for the principal that compel or impel the agent to abstain from corruption. one of the principal papers in this research domain is [7] that describes the social costs from corruption in the bureaucratic chain of an arbitrary length with consideration of the competence between agents throughout their hierarchy. the paper by m.bac [8] continues the investigations started by f.kofman and j.lawarree [9] and targeted to the analysis of corruption in the 38 andrey v. antonenko: static models of corruption in hierarchical systems hierarchy principal controller agent. as an extension of the basic model m.bac has built a derivative model in which a hierarchy of homogeneous agents is considered instead of one agent. the papers [10-12] also refer to this class. this research direction includes also the works dealing with studies of the political-economic corruption, especially in the voting systems [13-15]. problems of endemic corruption are discussed in the papers [16-17]. in [18] a game theoretic model predicts that resource rents tend to an increase in corruption if the quality of the democratic institutions is relatively poor, but not otherwise. the paper [19] investigates the role of guilt aversion in public administration. the authors’ approach to the modeling of corruption is represented in [20-21]. its essence consists in the following principles. 2 principles of modeling twelve principles are used in building of the models of corruption. 1). the basic modeling pattern is a hierarchical structure “principal supervisor agent object” in different modifications and its investigation by means of the optimization theory and stackelberg games theory. in the dynamic models the state of the object is described explicitly while in the static models only the impact of the agent to the object is considered. the supervisor may be corrupted, while the principal is not corrupted and controls corruption. so the elements of the above structure are bribe-controller, bribe-taker, bribe-giver and object of impact respectively. 2). certain requirements of the sustainable development of the controlled system (object) are supposed to be known. in the dynamic models they are formulated in terms of the object’s state while in the static ones in terms of the agent’s impact to the object. if the requirements of sustainable development are satisfied then the principal’s control problem is assumed to be solved even if corruption exists. 3). both pairs “principal supervisor” and “supervisor agent” are in the relations “leader follower”. the leading player on any level (principal or supervisor respectively) uses methods of compulsion (administrative or legislative impacts) or impulsion (economic impacts) for achievement of his/her objectives. the mathematical formalization of compulsion means an impact of the leader to the set of admissible strategies of the follower (as a rule, without a feed back), and impulsion means the impact to the follower’s payoff function (as a rule, with a feed back) [20]. 4). the cases of administrative corruption when administrative requirements or constraints are weakened for a bribe and the economic corruption when the economic ones are weakened by corruption are differentiated. the model of administrative corruption describes compulsion of the agent by the supervisor with advances in systems science and applications (2013) vol.13 no.1 39 a feed back on bribe, and the model of economic corruption describes impulsion in that pair with an additional feed back on bribe. 5). corruption threatens to the object’s sustainable development because it is profitable for the bribe-taker to weaken the requirements of sustainable development in exchange for a bribe. from the other side, corruption is a specific form of a feedback in the hierarchical systems subject to which the control variables become the functions of bribe. 6). corruption exists in the forms of capture and extortion. in the case of capture a basic set of administrative or economic services is guaranteed while additional indulgences are provided for a bribe. in the case of extortion a bribe is required already for the basic set of services, otherwise the requirements are enhanced. 7). the bribe-taker’s behavior is characterized by tractability (a willingness to weaken the administrative or economic constraints in exchange for a bribe) and greed (a price of the weakening). a set of quantitative indicators of tractability and greed is developed. 8). the bribe may represent a part of the payoff received by the agent subject to the administrative or economic indulgences or an absolute sum. in both cases it is convenient to suppose the bribe variable to be a part of the corrupted payoff or the whole agent’s income respectively. 9). for studying corruption in hierarchical control systems with consideration of the requirements of sustainable development both descriptive and normative approaches are applicable. in the case of descriptive approach the functions of administrative and economic corruption are given, and the main problem is to identify their parameters on statistical data. in the case of normative approach the corruption (bribery) function is found as the solution of an optimization or game theoretic problem. 10). the investigation of corruption in the system “principal-supervisor-agentobject” is possible from three points of view. if the bribery function is known then from the point of view of the agent the corruption can be described by an optimization model. from the supervisor’s point of view a hierarchical parametrical game of the class γ2 arises which solution in the form of bribery function with a feedback on the value of bribe is known in a general form [22, see appendix]. from the point of view of the principal the problem of corruption control consists in seeking of such values of control variables that the found optimal strategy of the supervisor satisfies the requirements of sustainable development. 11). the identification problem in this domain is not at all trivial and requires special investigations and expertises. each data set determines a specific social, economic and political system exposed to corruption. 12). it makes sense to build “genetic” series of sequentially complicated mod40 andrey v. antonenko: static models of corruption in hierarchical systems els that more and more precisely describe the real phenomena of corruption in hierarchical control systems. the principal logical pattern of this sequential complication has a form “optimization models hierarchical two-person games hierarchical three-person games”. with consideration of the possible modifications of the models of each type the “series” become the “genetic networks”. it is the last principle that determines the rest of the paper. 3 system of theoretical models of corruption the principles formulated above are used to build a system of static models of corruption in hierarchical control systems. 3.1 static models of economic corruption the basic optimization model of economic corruption has the form g(b) = b+ r(b) → min (1) 0 ≤ b ≤ 1 (2) where b is a part of the bribe, r(b) ∈ [0, 1] is a given function of the economic corruption (for example, a real diminishing of the tax rate, i.e. absence of sanctions in case of non-payment for the bribe). thus, the function g(b) means the total costs for tax payments and bribe that are to be minimized by the agent. in case of the linear parameterization r(b) = r0 −ab the model (1) takes the form g(b) = r0 + (1−a)b → min, 0 ≤ b ≤ 1 (3) here r0 is an official tax rate (0 ≤ r0 ≤ 1), a is a model parameter. considering that the function of economic corruption r(b) = r0 −ab monotonically decreases when 0 ≤ b ≤ 1 then a > 0. from the other side, the total costs g(b) are non-negative, therefore a ≤ 1 + r0. thus 0 < a ≤ 1 + r0. the parameter determines the qualitative characteristics of the bribe-taker’s behavior. if a = 0 then corruption is completely absent. as the value of a increases, the bribetaker’s tractability also increases and his greed decreases. the threshold value is a = r0: in this case r(1) = 0, i.e. the maximal greed ensures the maximal tractability. if a < r0 then the greed is over-limited and the tractability does not reach the maximal value (i.e. a positive tax is paid for any bribe). when a > r0 the agent can avoid the tax payments completely in exchange for a moderate bribe (maximal tractability and small greed). return to the solution of the optimization problem (1). having that dg(b) db = 1−a the function monotonically increases when 0 < a < 1 and its minimal value is reached in the left end of the admissible range: gmin = g(0) = r0. respectively when 1 < a < 1+r0 the function g monotonically decreases and its minimal value advances in systems science and applications (2013) vol.13 no.1 41 is reached in the right end: gmin = g(1) = 1 + r0 − a < r0. in the degenerate case a = 1 we get g(b) ≡ r0(the bribe is useless and corruption is absent). so, the parameter a again plays the key role and determines two qualitatively different behavior strategies of the agent. if 0 < a < 1 then the total costs of the agent g(b) increase, it is rational to abandon from bribe and to pay taxes equal to the legislative rate r0. if 1 < a < 1 + r0 then the costs g(b) diminish and an economic incentive arises to give the bribe and to pay in total 1 + r0 −a < r0. now consider the function of economic corruption in a more general form g(b) = b+ r(b) = b+ r0 −abk(k > 0). it is true that dg db = 1− kabk−1 = 0 ⇒ b∗ = (ka) 1 1−k ; d2g db2 = k(1− k)abk−2; d2g(b∗) db2 = (1− k)(ka) 1 k−1 . therefore, d2g(b∗) db2 { > 0, 0 < k < 1 ⇒ b∗ is the point of min imun; < 0, k > 1 ⇒ b∗ is the point of max imun. if b∗ is a point of maximum then min b g(b) = { g(0) = r0, 0 < a < 1, g(1) = r0 + 1−a, 1 < a < 1 + r0. finally bmin  (ka) 1 1−k , 0 < k < 1; 0, k > 1 ∧ 0 < a < 1; 1, k > 1 ∧ 1 < a < 1 + r0; (the case k = 1 is studied separately). so, if 0 < k < 1 then it is always profitable for the agent to give a bribe; if k > 1 then the reason to pay the bribe depends on the parameter . now consider a hierarchical game supervisor agent in the form gs (r, b) = b+ pr → max, 0 ≤ r ≤ r0; (4) ga(r, b) = b+ r → min, 0 ≤ b ≤ 1. (5) here the parameter p designates a part of the collected tax payments transferred by the principal (considered in the model implicitly) to the supervisor as a reward. in this model the function r = r(b) is not given and is found as an optimal strategy 42 andrey v. antonenko: static models of corruption in hierarchical systems of the leader in the game (4)-(5). using germeyer’s theorem (see appendix), we find the ε-optimal strategy in the form r̃ε(b) = { 0, b = r0 − ε ∧ p < 1− ε r0 , r0, otherwise. so, it is almost always profitable for the agent to give the bribe b = r0 − ε and to receive an arbitrary small but positive tax economy ε. the condition of effectiveness of the economic control of corruption is given by the inequality p ≥ 1 − ε r0 . however, it is hardly possible because in this case almost all tax payments should be assigned to the supervisors reward. thus let’s consider a principal’s problem of the administrative (compulsive) control of corruption as a hierarchical three-person game in the form gp = c(q) +k(r0 − r) → min, 0 ≤ q ≤ r0 (6) gs = b+ pr → max, q ≤ r ≤ r0 (7) ga = b+ r → min, 0 ≤ b ≤ 1 (8) here q is a variable of the principal’s administrative control that constraints from below the ability of supervisor’s corrupted behavior; c(q) is an increasing convex principal’s control cost function, c(0) = 0, c(r0) = ∞ is a parameter of the penalty charged on the supervisor if the condition of sustainable development r = r0 is violated. now the solution of the hierarchical game (7)-(8) takes the form r̃ε(b) = { q, b = r0 − ε ∧ p < 1− ε r0 , r0, otherwise. the first-order condition for the problem (6) gives q̂ = (c ′)−1(k). the values of the objective function in this point and in the ends of the admissible segment are equal to gp(q̂) = c ′(k) + k(r0 − r), gp(0) = kr0, gp(r0) = c(r0). having that k ≫ 1 or even k → ∞, i.e. the condition of sustainable development is unalterable for the principal, we get that his objective function reaches its maximum when q = r0 therefore the supervisor’s optimal strategy is identically equal to r0, the condition of sustainable development is satisfied and corruption is absent. thus, in this model the administrative control of corruption is more effective than the economic one. 3.2 static models of administrative corruption the basic optimization model of administrative corruption has the initial form ga(u, b) = (1− b)f(u) → max (9) advances in systems science and applications (2013) vol.13 no.1 43 0 ≤ u ≤ s(b), 0 ≤ b ≤ 1 (10) where b is a part of the bribe, u is the agent’s action, f(u) is the agent (bribegiver)/s production function, s(b) is a quota that constraints the agent’s action from above and may be extended for the bribe. having that the production function increases its maximum is always reached in the right end of the admissible segment. therefore the model (9)-(10) can be represented as an optimization problem with one variable g(b) = (1− b)f(s(b)) → max, 0 ≤ b ≤ 1 (11) in case of the linear parameterization of the function of administrative corruption s(b) = s0 + ab, where s0 is the official value of quota (notice that the function monotonically increases when 0 ≤ b ≤ 1 because it describes the quota extension in exchange for the bribe) and the linear production function f(x) = x the model (11) takes the form g(b) = (1− b)(s0 +ab) → max, 0 ≤ b ≤ 1 (12) as in the case of economic corruption, the parameter determines the qualitative characteristics of the bribe-taker’s behavior. if = 0 then corruption is completely absent. as the value of increases, the bribe-taker’s tractability also increases and his greed decreases. the threshold value is a = 1 − s0: in this case s(1) = 1, i.e. the maximal greed ensures the maximal tractability. if a < 1 − s0 then the greed is over-limited and the tractability does not reach the maximal value (i.e. any bribe nevertheless requires to obey a quota strictly less than 1). when a > 1−s0, the agent can ignore the quota completely in exchange for a moderate bribe (maximal tractability and small greed). return to the solution of the problem (12). we have g(0) = s0, g(1) = 0, dg(b) db = a− s0 − 2ab, d2g(b) db2 = −2a < 0, therefore b∗ = a−s0 2a is the point of maximum, g(b∗) = (a+s0)2 4a ≥ g(0). notice that b∗ { > 0, a > s0, < 0, a < s0, , and gmax = { g(b∗), a > s0, g(0), a < s0. so, the parameter again plays the key role and determines two qualitatively different behavior strategies of the agent. if a < s0 then there is no reason to give a bribe because the agent’s payoff is maximal when b = 0 and is equal to s0. however, if a > s0 then the optimal part of bribe is equal to b∗ = a−s0 2a that leads to the agent’s payoff (a+s0)2 4a ≥ s0. 44 andrey v. antonenko: static models of corruption in hierarchical systems in more general case g(b) = (1−b)(s0+ab)k, k ≤ 1, the maximal agent’s payoff is equal to gmax = { g(0), a ≤ s0 k , g(b∗), a > s0 k , where b∗ = ka− s0 (1 + k)a . so, it is disadvantageously to give a bribe when s0 ≥ ka. adding of a supervisor leads to the two-person hierarchical game in the form gs(s, u, b) = bf(u) → max, s0 ≤ s ≤ 1; ga(s, u, b) = (1− b)f(u) → max, 0 ≤ u ≤ s, 0 ≤ b ≤ 1. taking the agent’s production function in the form f(u) = aum and considering that its maximum is reached in the right end of the admissible segment, we get the following game: gs(s, b) = absm → max, s0 ≤ s ≤ 1; (13) ga(s, b) = a(1− b)sm → max, 0 ≤ b ≤ 1; (14) using germeyers theorem (see appendix) and considering that interests of the players coincide in s, we get the solution of the game (13)-(14) in the form s̃ε(b) = { 1, b = 1− ε− s0 m, s0, otherwise. so, it is profitable for the agent to give the bribe and to avoid quota completely. to prevent corruption it is necessary to add to the game (13)-(14) a principal with the control problem gp = c(q) +k(q − s0) → min, s0 ≤ q ≤ 1; (15) that is similar to the problem (6) in the model of economic corruption; the supervisor’s problem (13) takes the form gs(s, b) = absm → max, s0 ≤ s ≤ q; (16) i.e. the supervisor’s ability of corrupted behavior is restricted from above by the variable of non-corrupted principal’s administrative control q. then the solution of the game (16), (14) is s̃ε(b) = { q, b = 1− ε− ( s0 q )m , s0, otherwise. advances in systems science and applications (2013) vol.13 no.1 45 and the solution of the problem (15) has the form q∗ = { s0, c(s0) < c(1) +k(1− s0) 1, otherwise. considering that as in the model of economic corruption k ≫ 1 or even k → ∞, we get q∗ = s0, i.e. the principal ensures the condition of sustainable development s = s0, and corruption is absent. 4 models of corruption in real estate development in a general form an optimization model of the economic corruption in real estate development may be represented as g(b) = (1− b)[r(b) + ξ(1− r(b))] → max, 0 ≤ b ≤ 1, (17) where g(b) is the agent (developer)’s payoff function having the sense of his income from a real estate development project considering a bribe cost; b is an economic bribe as a part of the agent’s income; r(b) is a part of the social real estate redeemed with guarantee by the state at a fixed price (the price may be augmented for a bribe); ξ is a part of the social real estate that can be sold by the agent himself. the function of economic corruption r(b) is supposed to be known according to the descriptive approach. according to the economic sense it increases monotonically in the segment [0,1] and in the considered case of capture (extortion is analyzed similarly) r(0) = r0, where r0 is the legislative value of r (it is supposed that r = r0 is the condition of sustainable development of the real estate project). it is natural to use the function of economic corruption in the form r(b) = r0 +abk (a > 0), 0 ≤ r(b) ≤ 1, (18) restrict ourselves by the linear parameterization (k = 1). considering (18) we get r(b) = min{r0 +ab, 1}, a > 0. the parameter a determines the qualitative characteristics of the bribe-taker’s behavior. if a = 0 then corruption is completely absent. as the value of increases, the bribe-taker’s tractability also increases and his greed decreases. the threshold value is a = 1− r0: in this case r(1) = 1, i.e. the maximal greed ensures the maximal tractability. if a < 1− r0 then the greed is over-limited and the tractability does not reach the maximal value (i.e. the agent should sell a part of the social real estate himself with any bribe). when a > 1− r0 the agent can sell the total amount of the social real estate to the state at a fixed price in exchange for a moderate bribe (maximal tractability and small greed). the optimization problem (17) takes the form g(b) = (1− b)[r0 +ab+ ξ(1− (r0 +ab))] → max, 0 ≤ b ≤ 1, 46 andrey v. antonenko: static models of corruption in hierarchical systems we have g(0) = r0+ξ(1−r0), g(1) = 0, dg(b) db = −2a(1−ξ)b+a−aξ−r0−ξ+ξr0. from the first-order condition dg(b) db = 0,we find b∗ = (a−r0)(1−ξ)−ξ 2a(1−ξ) that is the point of maximum because dg2(b∗) db2 = −2a(1− ξ) < 0. by the structure gmax = { g(0) = r0 + ξ(1− r0), a < r0(1−ξ)+ξ 1−ξ . g(b∗) > g(0), otherwise. (19) thus, the agent’s optimal strategy is determined by the parameters a, r0, ξ subject to (19): in the first case it is the zero bribe, in the second one the bribe b∗ > 0 giving the payoff g(b∗) > g(0) . now consider the models of administrative corruption in real estate development. in a general case the optimization problem has the form g(b, u) = (1− b)[ru+ ξ(1− r)u+ γη(γ)(1− u)] → max, s(b) ≤ u ≤ 1, 0 ≤ b ≤ 1, (20) where g(b, u) is the agent(developer)’s payoff function that means his income from selling all types of the real estate considering corruption; b is a part of the administrative bribe; u is a part of the social type in the total amount of real estate; r – a part of the social real estate redeemed by the state at a fixed price with a guarantee; ξ – an estimated part of the social real estate that can be sold by the agent himself at the same price; γ > 1 – an increasing factor of the price of more expensive types of the real estate; η(γ) – an estimated part of the agent’s own sale of more expensive types of the real estate; s(b) – an obligatory minimal quota of the part of social real estate in the total amount (can be diminished for a bribe). notice that in the optimization problem (20) the payoff function depends on two variables. investigate the problem: ∂g ∂b = −[(r + ξ(1− r)− γη)u+ γη]; ∂g ∂b = 0 ⇒ µ∗ = γη γη − (r + ξ(1− r)) > 0, r + ξ(1− r) < γη; ∂2g ∂b2 = 0, ∂2g ∂b∂u = γη − (r + ξ(1− r)); ∂g ∂u = (1− b)(r + ξ(1− r)− γη); ∂g ∂u = 0 ⇒ b∗ = 1; ∂2g ∂u2 = 0; ∂2g ∂b∂u = γη − (r + ξ(1− r)). advances in systems science and applications (2013) vol.13 no.1 47 the hesse matrix has the form h = ∣∣∣∣∣∣∣∣ 0 γη − (r + ξ(1− r)) γη − (r + ξ(1− r)) 0 ∣∣∣∣∣∣∣∣ , its determinant is equal to |h| = −[γη−(r+ξ(1−r))]2 < 0. therefore (b∗, u∗) is a saddle point and maximum of the function g(b, u) is reached in the boundary of the set of admissible values. we have g(1, u) = 0;g(0, 1) = r + ξ(1− r)− γη; g(0, s0) = r + ξ(1− r) + (1− s0)γη > g(0, 1);g(0, s0) > 0. thus in any case maximum of the function g(b, u) is reached when u = s(b). then one-variable optimization problem arises in the form g(b) = (1− b)[(r + ξ(1− r)− γη)s(b) + γη] → max, 0 ≤ b ≤ 1. denote c = γη > 0, d = γη− (r+ ξ(1− r)). notice that if d < 0 then it is more profitable to sell the social real estate than more expensive one, and to give a bribe is not rational. therefore the agents payoff function can be written in the form g(b) = (1− b)(c −ds(b)), d > 0. as earlier, represent a function of administrative corruption in the form s(b) = s0 −abk (a > 0), 0 ≤ s(b) ≤ 1, where s0 is the legislative value of quota, and restrict ourselves by the case of linear (k = 1) parameterization of the function. s(b) = s0 −ab (a > 0), g(b) = (1− b)(c −ds0 +adb) = c −ds0 + (d(a+ s0)− c)b−adb2; ∂g ∂b = ad(1−2b) = 0 ⇒ b∗ = 1/2; ∂2g ∂b2 = −2ad < 0 ⇒ b∗is the point of maximum; g(0) = c −ds0; g(1) = 0; g(b∗) = 2c +d(a− 2s0) 4 . thus, gmax = { g(b∗), d(a+ 2s0) > 2c, g(0), otherwise i.e. the advantageousness of giving a bribe depends on the relation between the model parameters. notice that in the case b = b∗ the agent’s payoff is positive if s0 < 2c+ad 2d , i.e. an obligatory quota of the social real estate is not very big. 48 andrey v. antonenko: static models of corruption in hierarchical systems now consider a game theoretic model of the administrative corruption in real estate development in the form gp(q, s) = p1(s− s0)− q s0 − q → max, 0 ≤ q ≤ s0. gs(s, b) = p2(s− s0) + bc[(r + ξ(c)(1− r))u+ γη(γ)(1− u)] → max, q ≤ s ≤ s0. ga(b, u) = c(1− b)[(r + ξ(c)(1− r))u+ γη(γ)(1− u)] → max, s ≤ u ≤ 1, 0 ≤ b ≤ 1. (21) here b is a part of the administrative bribe; u is a part of the social type in the total amount of real estate; c – a price of the social real estate; r – a part of the social real estate redeemed by the state at the fixed price c with a guarantee; ξ(c) – an estimated part of the social real estate that can be sold by the agent himself at the same price; γ > 1 – an increasing factor of the price of more expensive types of the real estate; η(γ) – an estimated part of the agent’s own sale of more expensive types of the real estate; s(b) – an obligatory minimal quota of the part of social real estate in the total amount (can be diminished for a bribe); q – a parameter of the principal’s administrative control; p1, p2 > 0 – penalty factors charged on the principal and the supervisor respectively if the quota (condition of sustainable development s ≥ s0) is violated. it is supposed that r = const, ∼ s = s(b). the problem (21) is solved using a heuristic two-stage algorithm. on the first stage the hierarchical game between the supervisor and the agent is considered. notice that the function ga achieves its maximum in u independently from b, therefore subject to the linearity of the function ga in u the solution of the agent’s problem has the form u∗ = { 1, r + ξ(c)(1− r) ≥ γη(γ), s, otherwise consider these two cases separately. 1). r+ ξ(c)(1− r) ≥ γη(γ), i.e. it is more profitable to sell the social real estate. in this case a game does not arise because when u = 1 then the function ga does not depend on s. an evident solution of the optimization problem is b = 0, i.e. corruption is absent. 2). r + ξ(c)(1 − r) < γη(γ), i.e. it is more profitable to sell the expensive real estate. then a standard hierarchical two-person game arises in the form gs(s, b) = p2(s− s0) + cb[r + ξ(c)(1− r)− γη(γ)] → max, q ≤ s ≤ s0; ga(s, b) = c(1− b)[(r + ξ(c)(1− r)− γη(γ))s+ γη(γ)] → max, 0 ≤ b ≤ 1. advances in systems science and applications (2013) vol.13 no.1 49 application of germeyer’s theorem (see appendix) gives sp (b) ≡ s0; s d(b) = { s0, p2 > bc(r + ξ(c)(1− r)− γη(γ)), q, otherwise; la = c[(r + ξ(c)(1− r)− γη(γ))s0 + γη(γ)]; ea = {b = 0};da = {(s, b) : ga(s, b) > la} = ∅, s.t.; (1− b)[(r+ ξ(c)(1− r)− γη(γ))s+ γη(γ)] ≤ (r+ ξ(c)(1− r)− γη(γ))s0 + γη(γ); when 0 ≤ b ≤ 1, q ≤ s ≤ s0. it is assumed in this case that k1 = −∞ < k2, thus ∼ s ∗ (b) = { sd(b), b = 0, s0, otherwise, = { q, p2 < bc(r + ξ(c)(1− r)− γη(γ)) ∧ b = 0, s0, otherwise. however the conditions p2 < bc(r + ξ(c)(1 − r) − γη(γ))and b = 0 are disjoint, therefore ∼ s ∗ (b) ≡ s0. so, the solution of the agent’s optimization problem has the form b∗ = 0 (in fact, corruption is absent), and u∗ = s0, therefore it is not required to involve the principal to ensure the condition of sustainable development. respectively, in the second stage it is sufficient for the principal to solve the control costs minimization problem q s0−q → min, 0 ≤ q ≤ s0,the trivial solution of which has the form q∗ = 0. at last, consider a model of the economic corruption in the three-level system of real estate development gp = h(c) +m(r0 − r) → min, 0 ≤ c ≤ 1; gs = f(1)(b+ cr) → max, 0 ≤ r ≤ r0; ga = f(1)(1− b− r) → max, 0 ≤ b ≤ 1; where h(c) – an increasing convex principal’s cost function; m – a penalty factor (the penalty is charged if the condition of sustainable development r = r0 is violated). the supervisor’s optimal guaranteeing strategy has the form ∼ r ∗ (b) = { 0, b = r0 − ε ∧ c < 1− ε r0 , r0, otherwise, if b ̸= r0 − ε then r ≡ r0 and the evident solution of the principal’s optimization problem is c = 0. if b = r0 − ε then the principal can ensure the condition of 50 andrey v. antonenko: static models of corruption in hierarchical systems sustainable development only by choosing c = 1. therefore the solution of his optimization problem has the form c∗ = { 1, h(1) < mr0, 0, otherwise, i.e. the principal is obliged to compare the penalty for the violation of sustainable development condition and the costs required for its control. 5 conclusion the principles of modeling and control of corruption in the hierarchical systems are formulated. particularly, a problem of corruption control is set as the problem of realization of certain requirements to the state of controlled system (conditions of sustainable development). if the conditions are satisfied then the corruption control problem is supposes to be solved even if corruption in the system exists. this setting differs from the proposed by g.becker especially economic approach based on commensurability of the damage caused by corruption and its control costs [3]. another important methodical principle is building of “genetic” series of sequentially complicated models that more and more precisely describe the real phenomena of corruption in hierarchical control systems. this principle is realized in the paper by building of series of the static theoretical models of administrative and economic corruption as well as models of corruption in real estate development. the dependence of corrupted behavior on model parameters is investigated, and the analytical conditions in which corruption is not profitable for the agent or can be controlled by the principal for ensuring sustainable development of the system are received. appendix germeyer’s theorem [22]. suppose that functionsm1(x1, x2),m2(x1, x2) are continuous on the compacts x1, x2. introduce the following functions: a punishment strategy xp1 (x2) by the rule m2(x p 1 , x2) = min x1∈x1 m2(x1, x2), and the dominance strategy xd1 (x2) that satisfies the condition m1(x d 1 , x2) = max x1∈x1 m(x1, x2). besides, the following quantities and sets are introduced: l2 = max x2∈x2 m2(x p 1 , x2);e2 = {x2 ∈ x2 : m2(x p 1 , x2) = l2}; d2 = {(x1, x2) ∈ x1 ×x2 : m2(x1, x2) > l2}; k1 = sup (x1,x2)∈d2 m1(x1, x2) ≤ m1(x ε 1, x ε 2) + ε (d2 = ∅ ⇒ k1 = −∞); advances in systems science and applications (2013) vol.13 no.1 51 k2 = min x2∈e2 max x1∈x1 m1(x1, x2). then the guaranteed payoff of the player 1 (leader) in the game γ2 is equal to ω1 = max(k1,k2), and the respective ε-optimal strategy has the form x̃ε1(x2) =  xε1, x2 = xε2, k1 > k2, xd1 (x2), x2 ∈ d2, k1 ≤ k2, xp1 (x2), otherwise. acknowledgements the work is supported by rfbr (project 12-01-00017). references [1] rose-ackerman s. (1975), “the economics of corruption”, journal of public economics, no.4, pp.187-203. [2] rose-ackerman s. (1978), corruption: a study in political economy, n.y.: academic press. [3] becker g. (1968), “crime and punishment: an economic approach”, journal of political economy, vol.76, no.2, pp.169-217. [4] shleifer a, vishny r.w. (1993), “corruption”, the quarterly journal of economics, vol.108, no.3, pp.599-617. [5] besley t, mclaren j. (1993), “taxes and bribery: the role of wage incentives”, the economic journal, vol.103, pp.119-141. [6] chandler p, wilde r. (1992), “corruption in tax administration”, journal of political economy, vol.49, no.2, pp.333-349. [7] hillman l, katz e. (1987), “hierarchical structure and the social costs of bribes and transfers”, journal of political economy, vol.34, pp.129-142. [8] bac m. (1996), “corruption and supervision costs in hierarchies”, journal of comparative economics, vol.22, pp.99-118. [9] kofman f, lawarree j. (1993), “collusion in hierarchical agency”, econometrica, vol.61, no.3, pp.629-656. [10] lambert-mogiliansky a. (1996), “essays on corruption”, department of economics, stockholm university, pp.101-138. [11] olsen t.e, torsvik g. (1998), “collusion and renegotiations in hierarchies: a case of beneficial corruption”, international economic review, vol.39, no.2, pp.143-157. 52 andrey v. antonenko: static models of corruption in hierarchical systems [12] hindriks j, keen m., muthoo a. (1999), “corruption, extortion and evasion”, journal of public economics, vol.74, no.1, pp.395-430. [13] myerson r.b. (1993), “effectiveness of electoral systems for reducing government corruption: a game-theoretic analysis”, journal of economic literature, vol.5, no.1, pp.118-132. [14] dudley l, montmarquette c. (1987), “bureaucratic corruption as a constraint on voter choice”, public choice, vol.55, pp.127-160. [15] rasmusen e, ramseyer j. (1992), “trivial bribes and the corruption ban: a coordination game among rational legislators”, public choice, vol.78, pp.305-327. [16] levin m, satarov g. (2000), “corruption and institutions in russia”, european journal of political economy, vol.16, pp.113-132. [17] kahana n, qijun l. (2010), “endemic corruption”, european journal of political economy, vol.26, pp.82-88. [18] bhattacharia s, hodler r. (2010), “natural resources, bureaucracy and corruption”, european economic review, vol.54, pp.608-621. [19] balafoutas l. (2011), “public beliefs and corruption in a repeated psychological game”, journal of economic behavior and organization, vol.78, pp.5159. [20] ougolnitsky g. (2011), sustainable management, n.y.: nova science publishers, pp.287. [21] antonenko a.v, ougolnitsky g.a, usov a.b. (2012), “static models of corruption in hierarchical control systems”, contributions to game theory and management. vol.v. collected papers presented on the fifth international conference game theory and management / ed. leon a. petrosyan, nikolay a. zenkevich. spb.: graduate school of management spbu, pp.20-32. [22] gorelick v.a, kononenko a.f. (1982), game theoretic models of decision making in ecological-economic systems, moscow, pp.145(in russian). corresponding author guennady a. ougolnitsky can be contacted at: ougoln@gmail.com. advances in systems science and applications (2013) vol.13 no.4 369-378 in-solid acoustic sources extraction using a multi-sensor architecture d.t. pham, m. yang and m. al-kutubi manufacturing engineering centre cardiff university, cardiff cf24 3aa, uk abstract this paper presents a novel signal feature extraction method based on location pattern matching (lpm) for improving the reliability of pattern matching of received acoustic signals by using multi-sensor architecture. the proposed solution is to solve the reliability problem of signal matching by extracting specific features from the signal that is only associated with the source location not the source information with multi-sensor layout. experiments showed that this method can increase the tolerance of lpm on various types of impacts, thus the matching accuracy. keywords signal feature extraction, pattern matching, multi-sensor architecture 1 introduction fig.1 lpm system layout. according to the well described acoustic time-reversal theory [1], it is possible to reconstruct an acoustic signal at its initial excitation location in a scatter370 d. t. pham:in-solid acoustic sources extraction using a multi-sensor architecture ing medium by recording the received signal and sending back the time reversed version of the signal through the medium. this implies that the received signal carries its source location signature as a result of scattering in the transmission medium and reflections from its complex boundaries. in the lpm approach, the uniqueness of features embedded in a received acoustic signal, as a representation of its source location, is employed to determine if it matches one of the predefined locations on a tangible object. this is achieved by creating an array of template of received signals generated from taps on preknown locations [2-5]. the source location can then be determined by a cross correlation analysis from the best similarity match with a signal in the templates. fig.1 shows a block diagram of an lpm system comprising a board, one sensor located at the bottom right of the board, signal conditioning hardware, data acquisition device and a host pc. a tapping on the hardboard is detected by the sensor, conditioned by the signal conditioning hardware, digitised by the data acquisition card and processed by the pc. the matching result can be achieved by a cross correlation analysis which is widely used to measure similarity between two signals or images. 2 pattern-based localisation the theory which location identification from the received signal pattern is based on is that an acoustic wave propagating from its source to a destination carries a specific signature in its pattern associated with the source location as a result of scattering. when a driving force is applied into a medium, a mechanical wave is generated transporting energy away from the source of disturbance. in a bounded medium the waves propagate until they reach the boundaries and are reflected or absorbed. the reflections cause reverberation on the received signal acquired by a transducer at certain location in the medium. to comprehend the effect of boundary on the magnitude and phase of a received signal, a plane wave is assumed propagating in the x direction, the acoustic wave equation is given by: ∂2p ∂x2 = 1 ν2p ∂2p ∂t2 (1) where p is the acoustic pressure as a function of time t and distance x.the solution of the differential equation (1) is the propagating wave equation given by [6]: p = aej(ωt−kx) (2) where a is the amplitude constant, νp phase speed and k is the wave number. with this wave equation, it can be shown how a transmission medium with one reflection affects the received signal. assume a simple model of two signal paths advances in systems science and applications (2013) vol.13 no.4 371 from the transmitter to the receiver. one direct path with unity gain and a delay td , and the other is reflected with attenuation α and a delay td+∆t resulting from the path length difference. the overall transfer function of such a transmission medium h(ω) can be expressed by: h(ω) = e−jωtd + αe−jω(td+∆t) = |h(ω)|︷ ︸︸ ︷√ 1 + α2 + 2α cosω∆t) θh(ω)︷ ︸︸ ︷ e −j(ωtd+tan−1 α sinω∆t 1+α cosω∆t ) therefore, multi-path reflections cause the distortion in the magnitude |h(ω)| and the phase θh(ω) characteristics of the transmission medium. since the time delay is the product of the path length difference by the wave-number, it is clear from (1) that the magnitude and phase of the received signal at a certain location will vary as the source location changes. in random medium the received signal is a combination of the direct wave and multiple delayed scattered waves went through reflections and refractions plus the effect of non isotropic material, and therefore the received signal in reality has much more complicated relation to its source location than that given in equation (1). with this phenomenon, signals received from different source locations will have a distinctive feature that can be used to localize an unknown source signal if there is some knowledge about this feature. practically, location features are obtained from the received signal in the training stage. 3 focusing in time reversal theory when a source is applied into a medium at certain location and a received signal is recorded by an array of transducers as shown in fig.1. the received signals are reversed in time and then re-emitted into the medium. the re-emitted acoustic energy wave propagates back through the same medium and goes through all the multiple scattering, reflections and refraction that they underwent in the forward direction and refocuses on the source location. if only an aperture of limited area, called time-reversal mirror, is adopted in a time reversal operation, a small part of the field radiated by the acoustic source is captured and time reversed, thus limiting focusing quality [7]. however, in a bounded medium, multiple reflections along the medium boundaries significantly increase the apparent aperture of the time reversal mirror which means that the transducers can equivalently replaced by reflecting boundaries that redirect part of the incident wave towards the aperture. thus spatial information is converted into one dimension representation of the time domain and the reversal quality depends crucially on the duration of the time-reversal window, i.e. the length of the recording to be reversed. the heterogeneity of the medium 372 d. t. pham:in-solid acoustic sources extraction using a multi-sensor architecture or the boundaries which produces multi-paths contributes to have an aperture that is much larger than its physical size. it has been shown experimentally that in a cavity with specific geometrical property focusing with time reversal can be obtained using one transducer [7]. the time-reversal approach is clearly connected to the inverse source problem. (a) impulse transmission and reception. (b) time reversed transmission and impulse localization fig.2 sketch of time reversal focusing in random medium they both deal with propagation of a time-reversed field, but the propagation is real in the time reversal experiment and simulated in the inverse problem. moreover, the most important distinction is that the time reversal approach doesnt need knowledge of the propagating medium while the inverse problem method does. as any linear and time-invariant process, wave propagation through a multiple scattering medium may be described as a linear system with a certain impulse response. if the source sends a dirac pulse δ(t) function,the jth transducer of the time-reversal-mirror will receive a signal hj(t),which is the propagation impulse response from the source to transducer j. moreover, due to spatial reciprocity,hj(t) is also the impulse response describing the propagation of a pulse from the hj(t) transducer to the source. thus, if the transducer is able to record and time reverse the whole impulse response, the signal generated at the source can be represented by the convolution hj(t) ∗ hj(−t).this convolution product, in terms of signal analysis, is a typical matched filter which is a linear filter whose output is optimal in some sense. for whatever the impulse response hj(t), the temporal result is the convolution between this response and its time advances in systems science and applications (2013) vol.13 no.4 373 fig.3 time-reversed wave field observed at different times around the central point on a square of 15 15mm2 ([fp01]). reverse version hj(t) ∗ hj(−t).which is maximal at time t = 0 . this maximum is always positive and equals ∫ h2j t dt , i.e. the energy conveyed by hj(t) . for an n-element array, the signal recreated on the source can be written as, s(t) = j=n∑ j=1 hj(t) ∗ hj(−t) (3) even if hj(t) are completely random and apparently uncorrelated signals, each term in this sum reaches its maximum at time t = 0. so all contributions add constructively around t = 0 , whereas at earlier or later times uncorrelated contributions tend to interfere destructively one with another. thus the recreation of a sharp peak after time reversing on n-element array can be viewed as an interference process between n outputs of n matched filters. 4 realization of source localization from time reversal focusing as described in the previous section, it is possible with time-reversal theory to reconstruct an acoustic signal in its original location in a scattering medium by recording the received signals and sending back the time reversed version of these signals through the medium. this implies that the received signal carries its source location signature as a result of scattering in the transmission medium and reflections from its complex boundaries. with the same assumption of dirac delta source excitation, the response term hj(t) of the temporal correlation as given by (1) can be interpreted in lpm as a template obtained in the learning stage. here hj(−t) is the applied test signal with negative sign turns the convolution into a 374 d. t. pham:in-solid acoustic sources extraction using a multi-sensor architecture cross correlation operation. therefore, cross-correlation is a focusing process in time reversal but a similarity measure in template matching localization. although source reconstruction in time reversal is a transmission process or active, it is comparable to passive source localization with template matching by the operation of cross correlation as illustrated in the previous section. however, time reversal can help to provide explanation of physical limitations for source localization problems. ideally, the source reconstruction is achieved with array of sensors surrounding the source origin with element spacing of at least half a wavelength. practically a limited aperture area is used on the cost of focusing resolution. the smaller the array, the larger the focal spot. as a result of wave diffraction, the waves will refocus to a spot not smaller than the shortest wavelength [5]. accordingly, the achievable localization resolution using template matching can be increased with more sensors but still limited to the smallest wavelength, and since wavelength is inversely related to frequency, higher accuracy can be obtained with interactions that generate higher frequency signals such as using nail clicks and metallic object than those who generate lower frequencies as finger tap and damping material. 5 location feature extraction for enhanced reliability ideally to comply with the time reversal theory, an impulsive source is needed for both learning and recognition stages. practically, it is found that cross correlation is sensitive to the template type which means that similar type of impacts should be used in both stages since the variation in the signal will appear as variation in the location signature. one option to make the system more reliable is to use multiple templates for different type of interactions, but this is impractical as the calibration work will be intensified. signal filtering improves the resolution but found to be not effective to improve reliability because it filters the frequency components not the location feature. therefore an attractive novel solution is proposed here to solve the reliability problem by extracting specific features from the signal that is associated with the source location not the source information. let an unknown source signal given by s(t) emitted from location i on the surface of a tangible object as shown in fig.4, and two sensors are receiving the signals g1(t) and g2(t). the propagation path from the given source to sensor-1 and sensor-2 can be expressed by a specific transfer function denoted by h1(t) and h2(t) respectively. the transfer function is characterized by the complex propagation path and independent on the source signal information. accordingly, the transfer function for a specific source to receiver path represents the actual source location signature. treating the transmission medium as a black box of a single input/multiple output time invariant system, the output signal received advances in systems science and applications (2013) vol.13 no.4 375 by the ith sensor can be expressed analytically the convolution integral given by: gi(t) = ∫ ∞ −∞ hi(t)s(t− τ)dτ (4) for instance considering the output from one sensor only, it is possible to measure the transfer function for a given location by applying an impulse δ(t) t that location and measure the output. in that case the received signal is the transfer function which in turn can be used as location signature in the template. if the test signal is also an impulse, then the matching process is just the comparison of location signatures and therefore high accuracy is anticipated. otherwise, if the signal used for test, or for the template, are not an impulse, the resulting received signal will include source signal information plus location information, which accordingly will result in estimation error. this is why lpm works better with impulsive type of impacts. the task now is to develop a technique to extract the only source information from any type of interaction by employing two sensors. let the input in the system shown in fig.5 is a stationary random signal. the fig.4 received signals from two different paths. fig.5 black box lpm model 376 d. t. pham:in-solid acoustic sources extraction using a multi-sensor architecture input/output relationship in the frequency domain is given by fourier transform of equation (5) is given by: gi(f) = s(f)hi(f) (5) the hypothesis of extracting the location signature involves utilizing a measurable quantity that doesnt require any knowledge about the input excitation or medium transfer function. thus from the output/output relationship that is given by the cross spectral density function between the two outputs as given by: pg1g2 = g1(f)g2(f) (6) then a hypothetical transfer function can be defined as: hg1g2(f) = pg1g2(f) pg1g1(f) (7) where pg1g2 is the autocorrelation function of g1(t). it can be seen from the above three equations hg1g2(f) is a function of the two transfer functions h1(t) and h2(t) which still represents an independent location signature. a related subject in literature is the binaural localization in human which is simulated by the head related transfer function cue and defined by the ratio of the two output spectrums [7]. by rewriting the complex equation (7) in the form of magnitude and phase as: hg1g2(f) = ag1g2(f)∠φg1g2 (8) either the magnitude or phase patterns can be used as a location signature information. to use both of amplitude and phase information, the pattern given by: ĥg1g2(t) = ∫ ∞ −∞ ψ(f)pg1g2(f)/pg1g1(f)e −j2πft df (9) can be used, where φ(f) is a weighting function introduced to compensate the bandwidth variation and depends on the interaction method. with φ(f) = ag1g2(f), only phase information are extracted. obviously, utilizing only the phase or the magnitude pattern is computationally faster than using (9) since it is not necessary to convert them into time domain. an experimental result was carried out by registering a template from impacts at defined locations generated by pen tip hits on a glass sheet. the test database consists of different interaction types as pen tip hits, nail clicks and finger tapping. then with the evaluation procedure it is found that highest percentage of correct estimations were obtained using (9), than using the phase information only and lastly when magnitude information were used only. advances in systems science and applications (2013) vol.13 no.4 377 mdf board with 4 piezo disk sensors 3-d interactive object conditioning circuit for piezo sensors piezoelectric microphone on mdf board fig.6 experimental setups and sensors for lpm localization 6 experimentation the experimental setup consists of the interactive object, sensors, signal amplifier, data acquisition card and a pc to process the signals. different object materials and shapes have been tested including metal, glass, plastic, fibre boards and 3-d objects. the suitable sensors were the piezoelectric discs, electret microphones and accelerometers. for data acquisition, four channel pci card is used for evaluation and two channel sound card used for demonstrations such as the portable usb sound card and the wireless audio transmitter. some of them are pictured in fig.6. the lpm system found to be working well on variety of materials. the piezo ceramic sounders and electret microphones are the cheapest but can only pick up low frequencies when firmly attached to the surface. the piezoelectric microphone is very sensitive with wide bandwidth response and the most expensive among others. the piezoelectric shock sensor from murata is the best sensor with sufficient frequency response and a very reasonable price. 378 d. t. pham:in-solid acoustic sources extraction using a multi-sensor architecture 7 conclusion in this paper, a new method for in-solid acoustic source localisation based on learning and template match is presented. with this location pattern matching technique, it is possible to convert virtually any solid objects into an interactive interface by simply attaching sensors on the objects surface to transfer the acoustic signals resulting from natural interactions through signal conditioning hardware to data acquisition device for sampling before delivering the digital data to a pc where the localization algorithm is running. the reliability problem caused by the variance of impact patterns to the matching process has been improved by the concept of extracting the location signature pattern from received signals using multi-sensor layout. acknowledgements this work was financed by the european fp6 ist project “tangible acoustic interfaces for computer-human interaction (tai-chi)”. the support of the european commission is gratefully acknowledged. references [1] h.jeong, y.jang. (2000), “wavelet analysis of plate wave propagation in composite laminates”, composite structure, vol.49, pp.443-450. [2] r.ing, n.quieffin, s.catheline, m.fink. (2004), “tangible interactive interface usingacoustic time reversal process”, applied physics letters. [3] g.lai, p aarabi. (2003), “active object localization using speaker arrays”, proceedings of the ieee sixth international conference on information fusion, vol.1, pp.70-73. [4] e dijk. (2004), “indoor ultrasonic position estimation using a single base station, phd thesis”, eindhoven university of technology, netherlands. [5] bousseljot r. and kreiseler d. (2000), “waveform recognition with 10,000 ecgs”, ieee computers in cardiology proceedings, pp.24-27, cambridge, ma, pp.331-334. [6] o’hagan r and zelinsky a. (1997), “finger track-a robust and real-time gesture interface”, advanced topics in artificial intelligence, tenth australian joint conference on artificial intelligence proceedings, pp.475-484. [7] m.fink. (1999), “time reversed acoustics”, scientific america, pp.91-95. corresponding author author can be contacted at: my@cf.ac.uk. advances in systems science and applications (2012) vol.12 no.2 194-206 prevalence of clustering for current smoking, current drinking and components of metabolic syndrome among adults of lanxi, heilongjiang, china yuling duan1, jingbo zhao2,yujun zhao3, shiying fu3, fuman wang2 and liting yang2 1harbin center for disease control and prevention, harbin, heilongjiang provence, 150056, china 2department of epidemiology, school of public health, harbin medical university, harbin 150081, china 3department of cardiology, first clinical medical college of harbin medical university, harbin 150086, china abstract a total of 2967 participants were surveyed in a poor county of heilongjiang by a stratified randomized cluster sampling. information on cvd risk factors was collected with standardized questionnaires and laboratory measurements during 2006 to 2007.overall, 93.65%, 73.71%, 46.29%, 21.55% and 6.99% of rural residents had ≥ 1,≥ 2,≥ 3,≥ 4 and ≥ 5 modifiable cvd risk factors (serum triglycerides(tg), high density lipoprotein cholesterol(hdl-c), hypertension, diabetes, current smoking, current drinking and overweight. both hypertension prevalence and diabetes prevalence increased with age, but hdl-c prevalence decreased with age among both men and women (each p for trend, < 0.001). tg, current drinking and current smoking prevalence decreased with age (each p for trend, < 0.05, except for overweight and ms, each p for trend, > 0.05) among men. while, tg, current drinking, current smoking and ms prevalence increased with age (each p for trend, < 0.001) among women. in a multivariate model including age and sex, the odds ratio (95% confidence interval [ci]) of having ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 cvd risk factors versus none of the studied risk factors were 3.15(95%ci, 1.92-5.16), 3.73(95%ci, 2.27-6.13), 4.31(95%ci, 2.59-7.17), 5.24(95% ci, 3.04-9.03), and 9.29(95% ci, 4.78-18.06), respectively, for chinese adults 65 to 74 years old versus 35 to 44 years old; 1.77(95%ci, 1.282.45), 1.98(95%ci, 1.42-2.74), 2.39(95%ci, 1.70-3.35), 2.40(95% ci, 1.66-3.74), and 2.79(95% ci, 1.77-4.40), respectively, for men compared with women. keywords smoking, drinking, metabolic syndrome, clustering, prevalence 1 introduction the term “metabolic syndrome”(ms) refers to the clustering of visceral adiposity, hyperglycemia, high blood pressure(bp) and dyslipidemia[1-2] ,associated with increased risk of cardiovascular disease(cvd). drinking and smoking are considadvances in systems science and applications (2012) vol.12 no.2 195 ered as risk factors of cvd. cvd accounts for half of noncommunicable diseases deaths worldwide, 16.7 million in 2002[3]. nowadays, cvd are not exclusive in developed countries. the majority of those deaths (≈ 9.1 million) occurred to economically developing countries, and much of the burden of cvd in developing countries occurs to china[4-5]. since the reform and opening-up in 1978, infectious diseases has declined, but changes in lifestyle and diet have led to an increase in life expectancy, a greatly increased frequency of cvd and other chronic diseases[6-7] in china. heilongjiang province located in the northeast of china. lanxi county is about 60 kilometers in the north of harbin city, heilongjiang. although at the beginning of opening-up and reform in china, lanxis economy was among poor counties of nation, and its economic is still lagging behind other developed regions, the income and lifestyle of residents have greatly changed during recent decade. although prevalence and risk factors of cvd in the developed country or developed population were paid much attention to, but little we known about prevalence and risk factors of cvd in the poor population in the rural. in order to provide important information and health education programs to policy makers in economic poverty county of heilongjiang province, it is important to quantify the proportion of the population at high risk for cvd. such data provide an understanding of the size of the population in need of targeted interventions to lower the population burden of illness due to cvd. the goal of this study was to quantify the proportion of adult residents in poor economic county who had 1 or more of the following major modifiable cvd risk factors: components of metabolic syndrome(which was defined by the international diabetes federation(idf), including overweight, hdl-c, serum triglycerides(tc), hypertension, diabetes),current smoking and current drinking, to determine the percentage of adults with a clustering of 2 or more, 3 or more, 4 or more and 5 or more of these risk factors. this was completed with the use of data collected from rural residents in lanxi county of heilongjiang province during 2006 to 2007. 2 methods 2.1 study population multi-stage stratified sampling method was used to select a representative sample of the rural residents of heilongjiang. in the first stage, lanxi county was selected as poor county, according to economic condition and residents income of lanxi during 1978 to 2006. secondly, each town was chosen from lanxi county if its economy was among middle, and then pingshan was selected as its economy was in the middle level of lanxi county. lastly, 11 villages were chosen by randomized method from 35 villages of pingshan. total residents, who aged 35 year and lived in the locate villages more than 5 years were considered as subjects, were 3480 persons. 3012 individuals completed the survey and examination, 196 yuling duan:prevalence of clustering for current smoking, current drinking and... and others were excluded because of absence or denying to answer questions, as well as exclusion criterion: type diabetes, fever, acute infection disease and other factors which influenced measurement of blood pressure, blood glucose or blood lipids. the response rate is 86.55%. at last, 2967 participants aged 35 to 74 (1324 males and 1643 females ) were used for analysis. 2.2 data collection data collection was conducted in examination centers at local health stations in the participants residential area. in a few instances when participants were unable to go to the examination center, the interview and examination were conducted in their homes. during clinic or home visits, trained research staff administered a standard questionnaire. information on demographic characteristics include age, gender, education, ethnicity, occupation, household income, cigarette smoking, drinking, a self-reported history of stroke, myocardial infarction and congestive heart failure, and the previous diagnosis and treatment of hypertension, high cholesterol, and diabetes. smoking prevalence was calculated using data obtained from the self report questionnaire. current smoker is defined as a person who has smoked cigarettes continuously or accumulatively for over 6 months during his or her life and smoked at least once during the last month before the investigation[8]. current drinker was defined as a person who has drunk beer or wine or distilled spirits at least once a week, apart from drinking in festivals and holidays[9]. 2.3 blood pressure measurement for each participant, after 15 min of sitting, 3 blood pressure measurements were obtained by a trained nurse using a standard sphygmomanometer, which procedures recommended by the american heart association[10]. before their blood pressure measurement for at least 30 minutes, all subjects were advised to avoid alcohol, cigarette smoking, coffee, tea, and excessive exercise. the mean value of three blood pressure measurements was used for this study. hypertension was defined as an average systolic blood pressure (sbp) ≥ 140mm hg and/or an average diastolic blood pressure (dbp)≥ 90mm hg and/or self-reported current treatment for hypertension with antihypertensive medication[11-12]. 2.4 measurements of weight, height and waist circumstance body weight, height and waist circumstance were measured by trained and certified observers according to a standard protocol. subjects wore light clothing and no shoes. weight was measured to the nearest 0.1kg by using calibrated balance scales placed on a solid horizontal surface. height was measured to the nearest 0.1cm with a frankfort plane positioned at a 90◦ angle against a wall-mounted metal tape by using calibrated stadiometer. advances in systems science and applications (2012) vol.12 no.2 197 bmi was calculated as weight divided by height squared (kg/m2). waist circumference was taken at the midway point between the inferior margin of the last rib and the crest of the ilium in a horizontal plane and measured to the nearest 0.1 cm. overweight was defined as having a waist circumference≥ 90 cm (man),≥ 80cm (woman)[13]. 2.5 laboratory measurements participants were asked to fast overnight before their study visit, and blood samples, to measure serum lipids and plasma glucose, were drawn by venipuncture. blood specimens were processed at the field center and then sent to a central clinical laboratory of first affiliated hospital of harbin medical university in harbin by bus, where specimens were stored at -70◦c until laboratory assays were performed. hdl-c and serum tc were analyzed enzymatically with commercially available reagents[14]. lipid measurements were standardized according to the criteria of the centers for disease control and prevention-national heart, lung, and blood institute lipid standardization program[15]. abnormal tc was defined as tc ≥ 1.7 mmol/l. abnormal hdl-c was defined as hdl-c < 0.91 mmol/l or self-reported current treatment with cholesterol-lowering medication. for the glucose measurement, whole blood was collected in evacuated tubes containing naf. plasma glucose was measured with a modified hexokinase enzymatic method. diabetes was defined as having a fasting plasma glucose level≥ 7.0 mmol/l and/or self-reported current treatment with antidiabetes medication (insulin or oral hypoglycemic agents). 2.6 the idf consensus worldwide definition of the metabolic syndrome according to the new idf definition[13], for a person defined as having the metabolic syndrome they must have central obesity (defined as waist circumference ≥ 90cm for chinese men or ≥ 80cm for chinese women) plus any two of the following four factors: raised tg level: ≥ 1.7 mmol/l, or specific treatment for this lipid abnormality, reduced hdl cholesterol: < 1.03 mmol/l in males and < 1.29 mmol/l in females, or specific treatment for this lipid abnormality, raised blood pressure: systolic bp ≥ 130 or diastolic bp ≥ 85 mm hg, or treatment of previously diagnosed hypertension, raised fasting plasma glucose (fpg) ≥ 5.6 mmol/l, or previously diagnosed type 2 diabetes, if above 5.6 mmol/l , ogtt is strongly recommended but is not necessary to define presence of the syndrome. in addition, ethics committees of harbin medical university in china approved the study. written, informed consent was obtained from each participant before data collection. participants with untreated conditions, identified during the study, were referred to their usual primary healthcare provider. 198 yuling duan:prevalence of clustering for current smoking, current drinking and... 2.7 statistical methods database was set up by epidata 3.0, and all analyses were conducted with sas 9.1.3. analyses were conducted in these survey participants without a history of myocardial infarction, stroke, or congestive heart failure. crude rate were standardized according to the age distribution for chinese adults in the year 2000. participants missing key variables measurements such as smoking, drinking, height, weight, waist circumstances, hdl-c, tg, glucose, and blood pressure were also excluded from the analyses, leaving a final study population of 2967. the prevalence of each cvd risk factor (hdl-c, tc, hypertension, diabetes, current smoking, current drinking and overweight, ms) was determined for men and women separately, by age group (35 to 44, 45 to 54, 55 to 64, and 65 to 74 years), all from rural residence. the prevalence of 0, 1, 2, 3, 4, and 5 cvd risk factors was determined for the overall study population as well as for men and women separately. then, the percentage of the population with ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 cvd risk factors was determined by age group and sex separately. the significance of the differences in the prevalence of ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 cvd risk factors across subgroups was compared with the wald-χ2 test. the adjusted odds ratios(or) and 95% confidence intervals (ci) of having ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 major cvd risk factors versus no cvd risk factor were determined from multivariable logistic-regression models that included age group and sex. 3 results 3.1 prevalence of cvd risk factors in rural residents in heilongjiang in the 11 villages of lanxi county, population aged 35 to 74 years, the age standardized prevalence of hdl-c, tg, hypertension, diabetes, current drinking, current smoking, overweight(wc≥ 90(man),≥ 80(woman)) and ms was 31.67%, 13.82 %, 39.26%, 6.49%, 24.65%, 41.81%, 32.86%, 19.84% respectively (table 1). prevalence ± se ∗hdl−c < 0.91mmol/l, serum tg ≥ 1.7 mmol/l, and/or current cholesterollowering medication use. †sbp ≥ 140mm hg and/or dbp ≥ 90mm hg and/or current antihypertensive medication use. ‡fasting plasma glucose ≥ 7.0mmol/l and/or current antidiabetes medication use. §overweight was defined as waist circumference (waist circumference ≥ 90(man), ≥ 80(woman)) se indicates standard error. total sr indicates total standard rate, standardized on the basis of the year 2000 age distribution of the chinese population. the age standardized prevalence of overweight, tg and ms was higher in women than in men (each p < 0.001). while hypertension, diabetes, current advances in systems science and applications (2012) vol.12 no.2 199 table 1 age-standardized prevalence of cvd risk factors among rural study participants in heilongjiang province population no tg∗ hdl∗ hypertension† diabetes‡ current current overweight§ ms groups drinking smoking (90(man)) (80(woman)) total sr 2967 31.67±0.85 13.82±0.63 39.26±0.90 6.49±0.45 24.65±0.79 41.81±0.91 32.86±0.86 19.84±0.73 total 2967 32.69±0.86 13.38±0.63 41.32±0.90 6.81±0.46 25.45±0.80 42.67±0.91 33.50±0.87 20.63±0.74 man overall 1324 31.80±1.28 14.34±0.96 43.96±1.36 6.90±0.70 53.22±1.37 53.59±1.37 21.06±1.12 14.17±0.96 sr 1324 31.34±1.27 13.52±0.94 48.04±1.37 7.48±0.72 53.32±1.37 53.10±1.37 20.77±1.11 14.12±0.96 overall age, y 35321 31.78±2.60 17.45±2.11 25.86±2.44 4.67±1.18 52.02±2.79 53.89±2.78 21.81±2.30 14.02±1.94 45401 38.90±2.43 14.21±1.74 47.38±2.49 6.23±1.21 59.35±2.45 56.61±2.47 23.19±2.11 15.96±1.83 55331 29.31±2.50 10.88±1.71 56.19±2.73 9.37±1.60 56.50±2.72 54.08±2.74 21.15±2.24 14.20±1.92 65-74 271 22.14±2.52 11.07±1.91 65.31±2.89 10.33±1.85 42.07±3.00 45.76±3.03 15.50±2.20 11.44±1.93 woman overall 1643 33.36±1.16 13.24±0.84 35.39±1.18 6.30±0.60 2.91±0.41 38.03±1.20 43.01±1.22 25.35±1.07 sr 1643 33.78±1.17 13.27±0.84 35.91±1.18 6.27±0.60 2.98±0.42 34.27±1.17 43.76±1.22 25.87±1.08 overall age, y 35524 20.80±1.77 17.18±1.65 19.47±1.73 2.86±0.73 1.15±0.47 24.05±1.87 34.16±2.07 15.27±1.57 45533 36.77±2.09 12.76±1.45 35.65±2.07 6.00±1.03 3.38±0.78 39.77±2.12 48.78±2.17 27.95±1.94 55376 40.96±2.54 12.23±1.69 46.81±2.57 7.71±1.38 3.46±0.94 36.44±2.48 48.67±2.58 33.51±2.41 65-74 210 45.71±3.44 6.67±1.72 58.10±3.40 12.86±2.31 5.71±1.60 41.90±3.40 46.19±3.44 35.24±3.30 drinking and current smoking prevalence were higher in men than in women (each p < 0.001). there were no significant difference on hdl-c prevalence between woman and man(p > 0.05). hypertension and diabetes prevalence increased with age among both men and women (each p for trend, < 0.001). among men, tg, hdl-c, current drinking and current smoking prevalence decreased with age (each p for trend, < 0.05). among women, hdl-c prevalence decreased with age (p < 0.001). while, tg, current drinking, current smoking and ms prevalence increases with age(each p for trend, < 0.001). 3.2 prevalence of ≥ 1,≥ 2, and ≥ 3 cvd risk factors in rural residents in heilongjiang overall, 4.32% of rural men and 7.98% of rural women, respectively, did not have any of the risk factors investigated(tg, hdl-c, hypertension, diabetes, current drinking, current smoking and, overweight; (figure 1 and table 2). cvd risk factors: current smoking, current drinking and component of metabolic syndrome (tg, hdl-c, hypertension, diabetes, overweight) in contrast, 17.07%, 25.95%, 27.47%, 15.86%, and 9.33% of men and 22.11%, 28.75%, 22.53%, 13.52%, and 5.12% of women had 1, 2, 3, 4 and ≥ 5 of these risk factors, respectively. overall, 6.35%, 19.86%, 27.50%, 24.73%, 14.56%, and 6.99% of rural adults had 0, 1, 2, 3, 4 and 5 of these risk factors, respectively. in total, 93.65% of the population had 1 or more cvd risk factors (table 2). the prevalence of ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 major modifiable cvd risk factors was mainly higher at older ages (table 3; each p < 0.001). also, the age standardized prevalence of ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 risk factors were higher among men compared with women respectively (all p < 0.05)(table 3 ). prevalence±se; se indicates standard error; total sr indicates total standard 200 yuling duan:prevalence of clustering for current smoking, current drinking and... fig.1 age-standardized prevalence of cvd risk factors among men and women of poor county in heilongjiang in china table 2 standardized prevalence of 0, 1 ,2 ,3 ,4 and ≥ 5 risk factors among total and gender rural participants in heilongjiang province 0risk factors 1risk factors 2risk factors 3risk factors 4risk factors ≥ 5risk factors n(%se) n(%se) n(%se) n(%se) n(%se) n(%se) age group 35102(12.07±1.12) 227(26.86±1.52) 225(26.63±1.52) 178(21.07±1.40) 85(10.06±1.03) 28(3.31±0.62) 4534(3.64±0.61) 172(18.44±1.27) 266(28.51±1.48) 234(25.08±1.42) 151(16.18±1.21) 76(8.15±0.90) 5532(4.55±0.78) 107(15.22±1.35) 192(27.31±1.68) 190(27.03±1.67) 130(18.49±1.46) 52(7.40±0.98) 65-74 20(4.18±0.91) 82(17.12±1.72) 131(27.35±2.04) 130(27.14±2.03) 65(13.57±1.56) 51(10.65±1.41) woman 137(7.98±0.67) 363(22.11±1.02) 472(28.75±1.11) 370(22.53±1.03) 222(13.52±0.84) 84(5.12±0.54) man 57(4.32±0.56) 225(17.07±1.03) 342(25.95±1.20) 362(27.47±1.23) 209(15.86±1.00) 123(9.33±0.80) total 188(6.35±0.45) 588(19.86±0.73) 814(27.50±0.82) 732(24.73±0.79) 431(14.56±0.65) 207(6.99±0.47) total sr 6.92±0.47 20.80±0.75 27.63±0.82 24.24±0.79 13.83±0.63 6.59±0.46 rate, standardized on the basis of the year 2000 age distribution of the chinese population. prevalence±se; n indicates number of participants with risk factors; % indicate prevalence; se indicates standard error the adjusted odds ratio of having ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 major modifiable cvd risk factors versus none decreased progressively with increasing age (table 4). the adjusted odds ratios (95% ci) of ≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 cvd risk factors for persons 65 to 74 years of age compared with their counterparts 35 to 44 years of age were 2.43 (95% ci, 1.46 to 4.04), 2.96 (95% ci, 1.77 to 4.96), 3.78 advances in systems science and applications (2012) vol.12 no.2 201 table 3 prevalence of 0,≥ 1,≥ 2,≥ 3,≥ 4, and ≥ 5 risk factors among rural participants in heilongjiang province 0risk factors ≥ 1risk factors ≥ 2risk factors ≥ 3risk factors ≥ 4risk factors ≥ 5risk factors n(%se) n(%se) n(%se) n(%se) n(%se) n(%se) age-group 35102(12.07±1.12) 743(87.93±1.12) 516(61.07±1.68) 291(34.44±1.63) 113(13.37±1.17) 28(3.31±0.62) 4534(3.64±0.61 899(94.36±0.61) 727(77.92±1.360) 461(49.41±1.64) 227(24.38±1.41) 76(8.15±0.90) 5532(4.55±0.78) 671(95.45±0.78) 564(80.23±1.50) 372(52.92±1.88) 182(25.89±1.65) 52(7.40±0.98) 65-74 20(4.18±0.91 ) 459(95.82±0.91) 377(78.71±1.87) 246(51.36±2.28) 116(24.22±1.96) 51(10.65±1.41) woman 131(7.98±0.67) 1511(92.02±0.67) 1148(69.91±1.13) 676(41.17±1.21) 306(18.64±0.96) 84(5.12±0.54) man 57(4.32±0.56) 1261(95.68±0.56) 1036(78.60±1.13) 694(52.66±1.37) 332(25.19±3.27) 123(9.33±0.80) total 188(6.35±0.45) 2772(93.65±0.45) 2184(73.78±0.81) 1370(46.28±0.92) 638(21.55±0.75) 207(6.99±0.47) (95% ci, 2.24 to 6.40), 4.62(95% ci, 2.62 to 8.15) and 7.39(95% ci,3.68 to 14.84), respectively. in addition, after multivariable adjustment, men were more likely to have ≥ 1,≥ 2, and ≥ 3 cvd risk factors compared with women, respectively (all p < 0.001)(table 4). table 4 standardized prevalence of 0, 1 ,2 ,3 ,4 and ≥ 5 risk factors among total and gender rural participants in heilongjiang province ≥ 1risk factors ≥ 2risk factors ≥ 3risk factors ≥ 4risk factors ≥ 5risk factors age group 351.00(ref) 1.00(ref) 1.00(ref) 1.00(ref) 1.00(ref) 453.56(2.38-5.31) 4.19(2.80-6.29) 4.86(3.20-7.39) 5.85(3.71-79.23) 7.28(4.01-13.22) 552.76(1.83-4.17) 3.41(2.25-5.17) 4.07(2.65-6.25) 5.07(3.17-8.10) 5.10(2.73-9.52) 65-74 2.43(1.46-4.04) 2.96(1.77-4.96) 3.78(2.24-6.40) 4.62(2.62-8.15) 7.39(3.68-14.84) male-gender 1.77(1.28-2.45) 1.98(1.42-2.74) 2.39(1.70-3.35) 2.40(1.66-3.47) 2.79(1.77-4.40) ∗all variables (age group, gender) were included in the same models. for example, the odds ratios for age group were adjusted for gender, the odds ratios for gender were adjusted for age group. †risk factors include components of metabolic syndromeoverweight, wc≥90 cm(man) or ≥80 cm(woman), hdl-c ≤0.91 mmol/l, tg≥1.7 mmol/l, or on cholesterol-lowering medication, hypertension(sbp≥140 and/or dbp≥90 or on high blood pressure medication), diabetes (fasting plasma glucose ≥7.0 mmol/l or on antidiabetes medication), current smoking and current drinking. 4 discussion the metabolic syndrome is a common risk factor for cardiovascular diseases. individuals with this syndrome have an increased risk of developing cardiovascular disease(cvd)[16]. sandra costa fuchs et al[17] revealed that hypertension, diabetes mellitus, obesity, low fruit and vegetable intake, and lack of vigorous or moderate physical activity were clustered into a combination of risk factors, which were independently associated with self-reported cardiovascular disease. a populationcbased, cross-sectional national health examination survey(1998) of 8,816 subject aged 15-79 [18] showed that clustering of 3 or more cvd risk 202 yuling duan:prevalence of clustering for current smoking, current drinking and... factors was 22.7% in man and 21.7% in women. using < 21kg/m2 as a referent, subjects with bmi of 23kg/m2 and 27kg/m2 had an odds ratio of 3.5 and 10.2 in men, and 3.1 and 6.7 in women, respectively, for clustering of cvd risk factors. using <65 cm as a referent, subjects with wc of ≥90 cm in man and ≥80 cm in women had an odds ratio of 13.4 and 13.6, respectively, for clustering of cvd risk factors. a study[19] explored how to influence effect of risk factors on coronary artery disease(cad) when the age gap between men and women narrows, indicated that hypertension, diabetes, dyslipidemia and family history were independent risk factors for women with stable cad. clustering of traditional risk factors may explain the precocity of cad in women who are near in age to men. a study from souss, tunisia[20] for clustering of cardiovascular risk factors among obese urban school children showed that obese children were found to have higher blood pressure, higher tg levels and lower hdl-c than children of normal weight. this results indicated that clustering of cardiovascular risk factors among obese is no longer limited to industrialized countries. another study from japan[21] for multiple risk factors clustering and risk of hypertension, which 5275 japanese male office workers aged 23-59 years were involved, showed that after controlling for potential risk factors of hypertension, the odds ratio of hypertension compared with the absence of risk factors was 1.91, 2.65, 3.88, 6.54, and 8.18 for the presence of 1, 2, 3, 4, and 5 risk factors, respectively ( each p for trend < 0.001). the results indicated that the accumulation of risk factors is highly associated with the increased risk of hypertension in japanese men. multicentre collaborative study showed that high levels of cvd risk factors are common in many economically developing countries[22]. in a study of 7 economically developing countries, as many as 78%, 46%, 50%, and 20% of adults in 1 or more countries were current cigarette smokers, had a high cholesterol level, were overweight, and had hypertension, respectively[22]. these risk factors have emerged as important characteristics in predicting cvd morbidity and mortality in economically developing countries, including china[23]. result from international collaborative study of cardiovascular disease in asia[24] showed that 80.5%, 45.9%, and 17.2% of chinese adults had ≥ 1,≥ 2, and ≥ 3 modifiable cvd risk factors (dyslipidemia, hypertension, diabetes, cigarette smoking, and overweight), respectively. by comparison, 93.1%, 73.0%, and 35.9% of us adults had ≥ 1,≥ 2, and ≥ 3 of these risk factors, respectively. in a multivariate model including age, sex, and area of residence, the odds ratio (95% confidence interval [ci]) of having ≥ 1,≥ 2, and ≥ 3 cvd risk factors versus none of the studied risk factors were 2.61 (2.09-3.27), 3.55 (2.77-4.54) and 4.97 (3.67-6.74), respectively, for chinese adults 65 to 74 years old versus 35 to 44 years old; 3.65 (3.21-4.15), 4.67 (4.06-5.38), and 5.60 (4.70-6.67), respectively, for men compared with women; 1.18 (1.07-1.30), 1.34 (1.21-1.50), and 1.84 (1.60advances in systems science and applications (2012) vol.12 no.2 203 2.12), respectively, for urban compared with rural residents; and 1.98 (1.76-2.22), 2.75 (2.42-3.13), and 4.36 (3.68-5.18), for residents of northern compared with southern china, respectively. the present study indicates that 93.65% rural adults aged 35 to 74 years have at least 1 of the following cvd risk factors: current smoking, current drinking and component of syndrome (tg, hdl, hypertension, diabetes, and overweight). in addition, clustering of 2 or more, 3 or more, 4 or more or 5 or more of these risk factors was noted in 73.78%, 46.28%, 21.55% and 6.99 of rural adults, respectively. the significantly higher prevalence of ≥ 1,≥ 2,≥ 3,≥ 4 and ≥ 5 risk factors in men compared with women may be due to the fact that 43.96% of rural men versus 35.39% of women were hypertension, 6.90% of rural men versus 6.30% of women were diabetes, 53.22% of rural men versus 2.91% of women were current drinker and 53.59% of rural men versus 38.03% of women were current smokers. although the prevalence of current smoking and current drinking of men was higher than those of women, the prevalence of current smoking and current drinking decrease with age in man. in contrast, the prevalence of current smoking and current drinking increase with age in woman. the prevalence of ms in woman was higher than that in man. in this study, we selected components of metabolic syndrome which is defined by idf, current smoking and current drinking to estimate prevalence of cardiovascular disease. indexes we selected were different from international collaborative study in asia[24]. there were some reasons. firstly, components of ms are considered as risk factors of cardiovascular diseases, and prevalence of ms increase with age. secondly, prevalence of cvd, which involved components of metabolic syndrome, is limited. lastly, drinking should be considered as risk factors of cardiovascular diseases in our clustering analysis. this study contains the facts that its results are based on findings in 11villages, representative sample of the residents in the poor county, which allows for calculation of rural representative estimates. otherwise, standard protocols and instruments were used, 86.55% response rate was achieved, and the trained investigators were very careful to collect data. this is a cross-sectional study, which assesses the prevalence of several cvd risk factors. a limitation of the study was its reliance on estimates derived from a cross-sectional study. cross-sectional study has demonstrated the importance of these risk factors to the development of cvd, but cross-sectional study do not allow for quantification of the importance of risk factor clustering in the incidence of cvd. further design and statistics methods are needed[25-27]. in all, 93.65%, 73.78%, 46.28%, 21.55%, 6.99% have ≥ 1,≥ 2,≥ 3,≥ 4,≥ 5 of the cvd risk factors investigated in the current study, respectively. it was noted that hypertension and diabetes prevalence increased with age among both men 204 yuling duan:prevalence of clustering for current smoking, current drinking and... and women (each p for trend, <0.001). at the same time, among women, current drinking, current smoking and ms prevalence increases with age(each p for trend, <0.001). effective interventions for farmers such as improved diet, suitable drinking, smoking cessation and increased physical activity can safely and effectively lower the risk of cvd[28-29]. local government should take measure to prevent, detect and treat metabolic syndrome, and improve lifestyle in order to decrease the burden of cvd in the rural residents in the poor county. references [1] guidelines subcommittee. (2001), “executive summary of the third report of the national cholesterol education program (ncep) expert panel on detection, evaluation, and treatment of high blood cholesterol in adults (adult treatment panel iii)”, jama, vol.285, pp.2486-2497. [2] meigs jb, d’agostino rb sr, wilson pw, cupples la, nathan dm, singer de.(1997), “risk variable clustering in the insulin resistance syndrome”, the framingham offspring study. diabetes, vol.46, pp.159-1600. [3] lopez ad. (1993), “assessing the burden of mortality from cardiovascular diseases”, world health stat q, vol.46, pp.91-96. [4] he j, gu d, chen j, wu x, kelly tn, huang jf, chen jc, chen cs, bazzano la, reynolds k, whelton pk, klag mj. (2009), “premature deaths attributable to blood pressure in china: a prospective cohort study”, lancet, vol.374, no.9703, pp.1765-72. [5] murray cjl, lopez ad. (1997), “mortality by cause for eight regions of the world: global burden of disease study”, lancet, vol.349, pp.1269-1276. [6] popkin bm, horton s, kim s, mahal a, shuigao j. (2001), “trends in diet, nutritional status, and diet-related noncommunicable diseases in china and india: the economic costs of the nutrition transition”, nutr. rev., vol.59, pp.379-390. [7] yusuf s, reddy s, ounpuu s, anand s. (2001), “global burden of cardiovascular diseases. part i. general considerations, the epidemiologic transition, risk factors, and impact of urbanization”, circulatio, vol.104, pp.2746-53. [8] ma guang-sheng, kong ling-zhi,lun de-chun, et al. (2005), “the descriptive analysis of the smoking pattern of people in china”, chin. j. prev contr. chron. non-commun. dis., vol.13, no.5, pp.195. [9] ma guan-sheng, zhu dan-hong, hu xiao-qi, et al. (2005), “the drinking practice of people in china”, acta nutrimenta sinica, vol.17, no.5, pp.363. advances in systems science and applications (2012) vol.12 no.2 205 [10] perloff d, grim c, flack j, frohlich ed, hill m, mcdonald m, morgenstern bz. (1993), “human blood pressure determination by sphygmomanometry”, circulation, vol.88, pp.2460-2470. [11] guidelines subcommittee. (1997), “the sixth report of the joint national committee on prevention, detection, evaluation, and treatment of high blood pressure”, arch intern med, vol.157, pp.2413-2446. [12] guidelines subcommittee. (1999), “world health organization-international society of hypertension guidelines for the management of hypertension. guidelines subcommittee”, blood press suppl, vol.1, pp.9-43. [13] alberti kg, zimmet p, shaw j. (2005), “idf epidemiology task force consensus group. the metabolic syndrome a new worldwide definition”, lancet, vol.366, pp.1059-1062. [14] allain cc, poon ls, chan cs, richmond w, fu pc. (1974), “enzymatic determination of total serum cholesterol”, enzymatic determination of total serum cholesterol, vol.20, pp.470-475. [15] myers gl, cooper gr, winn cl, smith sj. (1989), “the centers for disease control-national heart, lung and blood institute lipid standardization program: an approach to accurate and precise lipid measurements”, clin lab med, vol.9, pp.105-135. [16] dorairaj prabhakaran. (2004), “the metabolic syndrome:an emerging risk state for cardiovascular disease”, the metabolic syndrome: an emerging risk state for cardiovascular disease, vol.9, no.1, pp.55-68. [17] fuchs sc, moreira lb, camey sa, moreira mb, fuchs fd. (2008), “clustering of risk factors for cardiovascular disease among women in southern brazil: a population-based study”, cad saude publica, vol.24, no.2, pp.285293. [18] park hs, yun ys, park jy, kim ys, choi jm. (2003), “obesity, abdominal obesity, and clustering of cardiovascular risk factors in south korea”, asia pac j clin nutr, vol.12, no.4, pp.411-418. [19] mansur ap, gomes ep, avakian sd, favarato d, césar la, aldrighi jm, ramires ja. (2001), “clustering of traditional risk factors and precocity of coronary disease in women”, int j cardiol, vol.81, no.2-3,pp.205-209. [20] ghannem h, harrabi i, ben abdelaziz a, gaha r, mrizak n. (2003), “clustering of cardiovascular risk factors among obese urban schoolchildren in sousse”, tunisia, east mediterr health j, vol.9, no.1-2, pp.70-77. 206 yuling duan:prevalence of clustering for current smoking, current drinking and... [21] nakanishi n, li w, fukuda h, takatorige t, suzuki k, tatara k. (2003), “mnltiple risk factor clustering and risk of hypertension in japanese male office workers”, ind health, vol.41, no.1, pp.327-331. [22] inclen multicentre collaborative group. (1992), “risk factors for cardiovascular disease in the developing world: a multicentre collaborative study in the international clinical epidemiology network (inclen)”, j clin epidemiol, vol.45, pp.841-847. [23] liu j, hong y, d’agostino rb sr, wu z, wang w, sun j, wilson pw, kannel wb, zhao d. (2004), “predictive value for the chinese population of the framingham chd risk assessment tool compared with the chinese multi-provincial cohort study”, jama, vol.291, pp.2591-2599. [24] dongfeng gu, md, msc; anjali gupta, mph; paul muntner, phd, mhs et al. (2005), “prevalence of cardiovascular disease risk factor clustering among the adult population of china”, circulation, vol.112, pp.658-665. [25] weihu cheng and zhenhai yang. (2008), “multivariate logistic distribution”, advances in systems science and applications, vol.8, no.3, pp.415-420. [26] liangping hu and tianming zhang. (2008), “statistics thought and its value in biomedical research”, advances in systems science and applications, vol.8, no.3, pp.430-436. [27] fanliang kong, hua xin. (2007), “advance research of generalized poisson risk model”, advances in systems science and applications, vol.7, no.1, pp. 17-21. [28] miller yd, dunstan dw. (2004), “the effectiveness of physical activity interventions for the treatment of overweight and obesity and type 2 diabetes”, j. sci med sport, vol.7, pp.52-59. [29] thomson cc, rigotti na. (2003), “hospitaland clinic-based smoking cessation interventions for smokers with cardiovascular disease”, prog cardiovasc dis, vol.45, pp.459-479. corresponding author corresponding author:zhaojb168@sina.com,zyjddm@sohu.com advances in systems science and applications (2013) vol.13 no.1 21-36 systems defined on dynamic sets and analysis of their characteristics xiaojun duan1 and yi lin2 1department of mathematics and systems science, national university of defense technology,changsha 410073, pr china 2department of mathematics,slippery rock university slippery rock, pa 16057 usa abstract (lin, 1999) describes the static structures of general systems and investigates systems dynamics using the concept of time systems without clearing pointing out where systems attributes would come into play. to resolve this problem, in this paper, we introduce the concept of dynamic sets so that the evolutions and changing structures of systems can be adequately described by using object sets, relation sets, attribute sets, and environments of systems. on the basis of this background, we revisit some of the fundamental properties of systems, including systems emergence, stability, etc. keywords dynamic set, system description frame, attribute set, system property 1 introduction systems research represents such a science that it investigates the structure, evolution, and control of systems. its goal is to study the interactions of the parts or relevant matters or events that constitute a whole and the consequent emerging holistic properties and behaviors of the whole. this whole that is made up of the parts, relevant matters or events is known as a system. in order to investigate different behaviors of systems, as of this writing, there has been a good amount of published literature, constituting a well developed theory, see (lin, 1999)[1] and listed references there[2-4]. zadeh (1962)[5] believes that the main goal of systems science is to study the organization and structure of systems, that is, how parts of a system are organized together to form the whole. from the basis of set theory, (lin, 1999)[1] develops a multi-relation systems theory of rigor by establishing and analyzing the elementary concepts and fundamental characteristics of systems. that is, lin establishes the basic concepts and properties of systems science on the solid foundation and rigor of mathematics. the general systems theory maintains (lin, 1999)[1] that each system stands for such a whole that it consists of a set of objects and a set of interactions between the objects. symbolically, a system can be written as s = (m,r) (1.1) where the set m = {mi : i ∈ i} contains all the elements that make up of the system s, known as the object set, i is an index set, r the set of all relations 22 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics between the objects in m that describe the system s. for any relation r ∈ r, there is an ordinal number n(r), which is a function of r, such that r is a subset of the cartesian product of n(r) copies of m : r ⊆ mn(r). both the object and relation sets m and r together completely describe the given system s, its component parts, and the interactions between the parts. on this basis, leveled systems, time systems, etc., are effectively described and studied in general systems theory. because the objects of systems are highly abstract, the only difference between sets of objects will be their index sets. that is why such general theory can be powerfully employed to investigate systems with parts of identical properties. that is, the general system, as given in equ.(1.1), grasps the commonality of objects by ignoring their individual specifics. however, in terms of a realistic system, the objects’ specifics might play important roles in the operations of the system. they influence not only the composite of the system, but also the structure and functionality of the system, producing the relevant dynamic behaviors of evolution of the system. for example, a system that is made up of pure oxygen gas or of pure hydrogen gas is fundamentally different of that consisting of both oxygen and hydrogen gases. in particular, the interactions of the objects of the former system are mainly the repulsive and attractive forces between the gas molecules, while in the latter system, other than the similar repulsive and attractive forces, there are also following chemical reactions when the environment provides the needed condition: 2h2 +o2 → 2h2o (1.2) the resultant system after the chemical reactions is different of the system that existed before the reactions in terms of the object set and also the relation set. what is more important is that the property of further chemical reactions in the system resulted from the first round of chemical reactions is essentially different from that of the system that existed before the first round of chemical reactions. if the system that existed before the chemical reactions is written as s0 = (m0, r0) and the system that existed after as s1 = (m1, r1), then it is obvious that m0 ̸= m1 and r0 ̸= r1. systems s0 and s1 are very different due to that the objects in s0 have undergone substantial change with 2h2o molecules added. additionally, the interactions between the objects of s1 are totally different of those in s0. other than repulsive and attractive reactions between the gas molecules, the interactions of the objects of s1 also include those between liquid molecules and gas molecules, and those between liquid molecules. from this example, it follows that for certain circumstance, it is just not enough to employ only object and relation sets to describe the systems and their behaviors, especially if changes and evolutionary behaviors of the systems are concerned with. it is because in the current general systems theory, both h2 and advances in systems science and applications (2013) vol.13 no.1 23 o2 are treated as abstract objects in s0, while ignoring their differences in other aspects. when such differences do not affect the evolution of the systems much, the description of systems in equ.(1.1) will be most likely sufficient. however, the undeniable fact is that there are many such systems that the interactions between the objects of specific attributes great affect the structures, functionalities, and evolutions of the systems, just as the differences described by the two gases in equ.(1.2). therefore, we need to develop such a systems theory to deal with this situation. we have to further deliberate the description of the object and relation sets of systems so that the consequent evolutions of systems can be adequately investigated. 2 attributes of systems because the usage of rigorous mathematical language in the discussion of properties of general systems is very advantageous, we will continue to employ this approach. in this paper, we will not consider the case when a set is empty; and all sets are assumed to be well defined. that is, we do not consider such paradoxical situations as the set containing all sets as its elements. to this end, those readers who are interested in set theory and the rigorous treatment of general systems theory are advised to consult with (lin, 1999)[1]. as discussed earlier, our main focus here is the difference between various parts of a system, that is, the differences between the system’s objects and between the relations of the objects. we will employ the concept of attributes to describe such differences. by attribute, it means a particular property of the objects and their interactions and the overall behavior and evolutionary characteristic of the system related to the property. if speaking in the abstract language of mathematics, the attributes of a system are a series of propositions regarding the system’s objects, relations between the objects, and the system itself. due to the differences widely existing between systems’ objects, between systems’ relations, and between systems themselves, these propositions can vary greatly. for example, a proposition might hold true for some objects, and become untrue for others. in this case, we say that these objects have the particular attribute, while the others do not. as another example, a proposition can be written as a mapping from the object set to the set of all real numbers so that different objects are mapped onto different real numbers. in this case, we say that the system’s objects contain differences. additionally, the discussion above also indicates that each system evolves with time so that its objects, relations, and attributes should all change with time. so, when we consider the general description of systems, we must include the time factor. by doing so, the structure of a system at each fixed time moment is 24 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics embodied in the object set, relation set, and attribute set, while the evolution of the system is shown in the changes of these sets with time. summarizing what is analyzed above, we can define the system of our concern as follows: definition 2.1. assume that t is a connected subset of the interval [0,+∞), on which a system s exists. then, for t ∈ t , the system s is defined as the following order triplet: st = (mt, rt, qt) (2.1) where mt stands for set of all objects of the system s at the time moment t, rt the set of the relations between the objects in mt, and qt the set of all attributes of the system s. for the sake of convenience of communication, t is referred to as the life span and st the momentary system of the system s. what needs to be emphasized is that we study not only the evolution of systems with time, but also the evolution of the system along with continuous changes of some conditions, such as temperature, density, etc. in such cases, t will be understood as one of those external conditions. the object set of the system of our concern is made up of the system’s fundamental units. so, for each t ∈ t , the objects in mt are fixed. let us write m̂ = {mt,a : a ∈ it, t ∈ t} = ∪ t∈t mt (2.2) where mt = {mt,a : a ∈ it} is the object set of the momentary system st, and it the index set of mt as a function of time t. if for any t ∈ t , mt = m , for some set m , then we say that the object set of the system s is fixed. one of the most elementary relations between objects is binary, relating each pair of objects. such a relation can be written by using the 2-dimensional cartesian product of the system’s object set: r0 t,2 = {(mt,a,mt,b) ∈ mt ×mt : φ 0 t,2(mt,a,mt,b)} ⊆ m2 t (2.3) for each t ∈ t , where φ0 t,2(, ) is a proposition that defines r0 t,2. corresponding to different properties, the system st might contain different binary relations as subsets of the 2-dimensional cartesian product m2 t of the object set. the set of all the binary relations in st is written as follows: rt,2 = {rk t,2 : k ∈ k2} (2.4) where k2 is the index set of all binary relations of st. similarly, the system st might contain trinary relations: r0 t,3 = { (mt,a,mt,b,mt,c) ∈ m3 t : φ0 t,3 (mt,a,mt,b,mt,c) } ⊆ m3 t (2.5) advances in systems science and applications (2013) vol.13 no.1 25 for each t ∈ t , where φ0 t,3(, , ) is a proposition that defines r0 t,3 . the set of all trinary relations of st is written as follows: rt,3 = {rk t,3 : k ∈ k3} (2.6) where k3 is the index set of all trinary relations of st. higher order relations of st can be introduced similarly. for convenience, let us define unitary relations of st as follows: r0 t,1 = {(mt,a) ∈ mt : φ 0 t,1(mt,a)} ⊆ mt (2.7) for each t ∈ t , where φ0 t,1() is a proposition that defines r0 t,1. the set of all unitary relations of st is written as follows: rt,1 = {rk t,1 : k ∈ k1} (2.8) wherek1 is the index set of all unitary relations of st. now, the set rt of relations of st can be written as follows: rt = ∪ α∈ord ri,α (2.9) where ord stands for the set of all ordinal numbers. when ord is taken to be n = the set of all natural numbers, equ. (2.9) stands for the set of all finite relations the momentary system st contains. from the discussion above, it can be seen that all kinds of algebras and spaces studied in mathematics are special cases of equ. (2.9) with ord replaced by n . without loss of generality, let us assume that mt ∈ rt,1. with this convention, it can be seen that in the following when we talk about the attributes of a system st, we only need to mean the attributes of the relations of st, because the object set is also considered a relation, a unitary relation and the attributes of objects are now also those of relations. the attribute set qt of the momentary system st are a series of propositions about the relations. symbolically, we can write: qt = {qt(r1, r2, ..., rα, ...) : qt(...) is a proposition of rα ∈ rt, α = 1, 2, 3...} (2.10) without any doubt, with this notation in place, these propositions in qt embody all aspects of the momentary system st. for example, the concept of mass of object in physics is a mapping: µt : rt → r+ (2.11) which assigns each element in an unitary relation the mass of the element, and each element x⃗ = (x1, x2, ..., xα, ...) in an n-nary relation the sum of the masses 26 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics of the objects contained in the element, if the sum exists. that is, for any r ∈ rt, and any x⃗ = (x1, x2, ..., xα, ...) ∈ r, µt(x⃗) = µt(x1, x2, ..., xα, ...) = ∑ α µt(xα) (2.12) assuming that the sum on the right hand side converges. let us look at the network model of systems as an example, where only unitary and binary relations are considered. in this model, all objects of the system of concern are treated as nodes; each binary relation is modeled as the set of edges of the network. by doing so, each binary relation of the system corresponds to a network or a graph; and different binary relations correspond to different sets of edges. such correspondence can be seen as an attribute of the edges. accordingly, different unitary relations of the system correspond to different attributes of nodes in the network, such as size of the nodes, flow intensities of the nodes, etc. if different types of edges are treated as identical, then the network can be expressed as g = (v,e,q), where v stands for the set of all nodes, e the set of edges, and q some attributes of either the nodes or the edges or both, such as weights of the edges. the ordered pair (v,e) completely describes the topological structure of the network, which is sufficient for some applications. however, if the system we investigate is quite specific, for example, it is a network of railroads, a network of human relationships, etc., we may very well need to model multiple relations. in this case, the attribute set q can be employed to describe the scales of the stations in the railroad network, the traffic conditions or transportation capabilities between stations, etc. if the system is a network of human relationships, then q can be utilized to represent the social status of each individual person, the intensity of interaction between two chosen persons, etc. as a matter of fact, the concept of systems, as defined in equ.(2.1), is a generalization of that as defined in equ.(1.1) (lin, 1999)[1]. in other words, we can rewrite equ.(1.1) in the format of equ.(2.1) as follows. let all object sets be static. so, for any t ∈ t , we have mt = m and rt = r. and what is interesting is how an attribute is introduced. to this end, we can introduce a proposition q0 on the cartesian product m̂ = ∑ α∈ordm α of the object set m so that for any r ∈ m̂ , q0(r) = { 1, if r ∈ r 0, otherwise (2.13) then, we take q = {q0}. that is, the attribute set is a singleton. now, each system written in the format of equ. (1.1) is rewritten in the format of equ. (2.1). here, the attribute q0 describes if an arbitrarily chosen relation in the cartesian product m̂ belongs to the system’s relation set r or not. in essence, it restates advances in systems science and applications (2013) vol.13 no.1 27 the membership relation to the relation set r from the angle of attributes. in particular, because the relation set r of the system s is a subset of m̂ ; now the membership in the relation set r is determined by a proposition q0, while such a description is an attribute of the system s. that is, the general systems theory developed on set theory (lin, 1999)[1] has already implicitly introduced the concept of attributes. what we do here is to make this fact explicit. and because the systems we are interested in can have multiple attributes, our contribution to the general systems theory is to make the concept of attributes more general as a set q of attributes, including more than just the particular attribute q0. 3 subsystems just like each set has its own subsets, every system has subsystems, which can be constructed from the object set, relation set, and the attribute set of the system. in short, a system s is a subsystem of the system s, provided that the object set, relation set, and attribute set of s are corresponding subsets of those of s so that the restrictions of the attributes of s on s agree with the attributes of s. symbolically, we have definition 3.1. let st = (mt, rt, qt) and st = (mt, rt, qt) are two systems, for any t ∈ t . if the following hold true: mt ⊆ mt, rt ⊆ rt|st , and qt ⊆ qt|st ,∀t ∈ t (3.1) then st is known as a subsystem of st, denoted st < st, ∀t ∈ t, or s < s (3.2) where rt|st and qt|st represent respectively the restrictions of the relations and attributes in rt and qt on the system st and are defined as follows: rt|st = {r|mt : r ∈ rt} and qt|st = {q|mt∪rt : q ∈ qt}. in this definition, the notation of less than of mathematics is employed for the relationship of subsystems, because the relation of subsystems can be seen as a partial ordering on the collection of all systems. similarly, when equ. (3.1) does not hold true, we say that s is not a subsystem of s, denoted st ≮ st, ∀t ∈ t , or s ≮ s. let the set of all subsystems of s be s , then we have s ≡ {s : s < s} = {st : st < st, ∀t ∈ t} (3.3) for any given system s = (m,r,q), where m = {mt : t ∈ t}, r = {rt : t ∈ t}, and q = {qt : t ∈ t}, as defined by equ. (2.1), take a subset set of its object set 28 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics a = {at ⊆ mt : t ∈ t}. then by restricting the relation set r and the attribute set q on a, we obtain the following subsystem of s induced by a: s|a = {st = (at, rt|at , qt|(at,rt|at ) ) : t ∈ t} (3.4) where the restrictions rt|at and qt|at are assumed respectively to be rt|at = {r|at : each element in r|athas the same length, r ∈ rt} (3.4a) and qt|(at,rt|at ) = {q|(at,rt|at ) : qt|(at,rt|at ) is a well defined proposition on st, q ∈ qt} (3.4b) proposition 3.1. the induced subsystem s|a is the maximum subsystem induced by a. proof. let s = {(at, rt,a, qt,a) : t ∈ t} < s be an arbitrary subsystem with the entire a as its object set. according equ. (3.1) we have rt,a ⊆ rt|s, and qt,a = qt|s = qt|(at,rt,a), ∀t ∈ t so, from equ. (3.4a), it follows that s ≤ s|a only when rt,a = rt|s, the equal sign holds true. qed. proposition 3.2. assume that a and b are subsets of m satisfying that b ⊆ a ⊆ m , then s|b ≤ s|a. proof. according equ. (3.4), we have s|b = {st = (bt, rt|bt , qt|(bt,rt|bt ) ) : t ∈ t}.b ⊆ a implies that bt ⊆ at, ∀t ∈ t . so, it follows that rt|bt ⊆ rt|at and consequently rt|bt ⊆ (rt|at)|bt and qt|(bt,rt|bt ) = (qt|(at,rt|at ) )|(bt,rt|bt ) . therefore, s|b ≤ s|a. qed. proposition 3.3. let s be a system. then the collection of all subsystems of s forms a partially ordered set by the subsystem relation “<”. proof. this result is a straightforward consequence of proposition 3.2. qed. assume that s1, s2, and s are systems with the same time span such that s1 < s and s2 < s. let si, i = 1, 2, denote the set of all subsystems of si. then each element in the set s1 − s2 = {s : s < s1, s ≮ s2} is a subsystem of s1 but not a subsystem of s2. similarly, each element in the set s2 − s1 = {s : s < s2, s ≮ s1} is a subsystem of s2 but not a subsystem of s1. let s1∆s2 = {s : s < s1, s ≮ s2} ∪ {s : s < s2, s ≮ s1} be the union of the previous two sets of subsystems of either s1 or s2; and s1 ∩ s2 = {s : s < s1, s < s2} the set of all subsystems of both s1 and s2. proposition 3.4. assume that s1 < s and s2 < s. then s1 ∪ s2 = {s : s < s1} ∪ {s : s < s2} is a subset of s|m1∪m2 . proof. s|m1∪m2 is a maximal element in the partially ordered set (s|(m1∪m2),⊆) , satisfying ∀s ∈ s|m1∪m2 , s < s|m1∪m2 . on the contrary,∀s < s|m1∪m2 , we advances in systems science and applications (2013) vol.13 no.1 29 have s ∈ s|m1∪m2 . so, ∀s ∈ s1 ∪ s2, we have s ∈ s|m1∪m2 . qed. what this result indicates is that the union s1 ∪ s2 of the sets of subsystems of two subsystems s1 and s2 is a subset of the s|m1∪m2 of the subsystems of the induced system on the union m1 ∪m2. proposition 3.5. given two arbitrary systems s1 and s2, there is always a system s12 such that s1 < s12 and s2 < s12. proof. without loss of generality, assume that s1 = {(m1 t , r 1 t , q 1 t ) : t ∈ t 1} and s2 = {(m2 t , r 2 t , q 2 t ) : t ∈ t 2}. to construct the system s12 = {(m12 t , r12 t , q12 t ) : t ∈ t 12}, t 12 = t 1 ∪ t 2, we first assume that the object sets m1 t and m2 t are disjoint, that is, m1 t ∩ m2 t = ∅, for any t ∈ t 1 ∩ t 2. then, the desired system s12 is defined as follows: for t ∈ t 1 ∩ t 2,  m12 t = m1 t ∪m2 t r12 t = r1 t ∪r2 t q12 t = q1 t ∪q2 t (3.5) for t ∈ t 1−t 2, define m12 t = m1 t , r 12 t = r1 t , and q12 t = q1 t , and for t ∈ t 2−t 1, define m12 t = m2 t , r 12 t = r2 t , and q12 t = q2 t , and is denoted as s12 = s1 ⊕ s2. now, if m1 t ∩ m2 t ̸= ∅, for some t ∈ t 1 ∩ t 2, we simply take two systems ∗s1 = {(∗m1 t , ∗r1 t , ∗q1 t ) : t ∈ t 1} and ∗s2 = {(∗m2 t , ∗r2 t , ∗q2 t ) : t ∈ t 2} with m1 t ∩ m2 t = ∅, for any t ∈ t 1 ∩ t 2, such that ∗si is similar to si, i = 1, 2, where similar systems are defined in the same fashion as in [1]. then, we define s12 =∗ s1⊕∗s2. up to a similarity, the system s12 is uniquely defined. therefore, it can be seen as well constructed such that s1 < s12 and s2 < s12, where the time spans t 1 and t 2 are seen as the same as t 12 such that when t ∈ t 1 − t 2 (respectively, t ∈ t 2 −t 1), we treat s2 t (respectively,s1 t ) as a system with empty object set. qed. from this proposition, it follows that when the interactions of some given systems are considered, these systems can always be seen as subsystems of a larger system. on the other hand, this proposition also shows that there is always some kind of interaction between two given systems, which is embodied in the fact that they are subsystems of a certain system. 4 interactions between systems when there is an interaction between two objects of a system s = (m,r,q) (of time span t ), where m = {mt : t ∈ t}, r = {rt : t ∈ t}, and q = {qt : t ∈ t}, it can be described by using a binary relation of the system. we say that objects m1 and m2 ∈ m , which means either mi = (mt)t∈t such that mt ∈ mt, for each t ∈ t , or mi = mt ∈ mt, for a particular t ∈ t, i = 1, 2, interacts with respect to an attribute q ∈ q, it means that the proposition q holds true for a binary relation 30 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics r ∈ r that contains either (m1,m2) or (m2,m1) or both. the idea of interactions between systems is a natural generalization of that between two objects. based on the discussion of the previous section, we will discuss interactions of systems in the framework of subsystems. assume that si < s, i = 1, 2. now, let us look at how these subsystems could interact with each other. definition 4.1. the system s1 is said to have a weak effect on the system s2 with respect to an attribute q ∈ q, provided that for any r2 ∈ r2 there is r1 ∈ r1 such that the proposition q holds true for the ordered pair (r1, r2). when no confusion is caused, we simply say that the system s1 affects the system s2 weakly without mentioning q. if the system s2 also exerts a weak affect of s1, then we say that these systems interact with each other weakly. definition 4.2. the system s1 is said to have a strong effect on the system s2 with respect to an attribute q ∈ q, provided that for any r2 ∈ r2 and any r1 ∈ r1, the proposition q holds true for the ordered pair (r1, r2). when no confusion is caused, we simply say that the system s1 affects the system s2 strongly without mentioning q. if the system s2 also exerts a strong affect of s1 , then we say that these systems interact with each other strongly. from these definitions, it follows that strong interaction requires interactions between every ordered pair of objects, which is a more rigorous requirement than that of weak interactions. also, to maintain the intuition behind the concepts of interactions, in definitions 4.1 and 3.2, we only look at two relations ri ∈ ri, i = 1, 2. in order to capture the general spirit, these individual relations should be replaced by subsets {ri ∈ ri : φi(ri)} , where φi() stands for the proposition that defines the set, for i = 1, 2. by doing so, what are discussed in definitions 4.1 and 4.2 become special cases. definition 4.3. given a subsystem s = (m, r, q) of a system s = (m,r,q), the totality of all objects in m −m, each of which interacts weakly with at least one object in m, is known as the environment of the subsystem s in s, denoted es . 5 systems properties based on dynamic set theory 5.1 basic properties for two given systems s1 = {s1 t = (m1 t , r 1 t , q 1 t ) : t ∈ t 1} and s2 = {s2 t = (m2 t , r 2 t , q 2 t ) : t ∈ t 2} , let us consider definition 5.1. these systems s1 and s2 are equal, provided that m1 t = m2 t , r 1 t = r2 t , q 1 t = q2 t , and t 1 = t 2, (5.1) for each t ∈ t 1 = t 2. the systems s1 and s2 are said to be identical on the time period t ⊆ t 1∩t 2, it means that m1 t = m2 t , r 1 t = r2 t , q 1 t = q2 t , for each t ∈ t. (5.2) advances in systems science and applications (2013) vol.13 no.1 31 definition 5.2. the system s1 is said to be homomorphically embeddable into the system s2, provided that there is a non-decreasing mapping f : t 1 → t 2 such that for any t1 ∈ t 1, if t2 = f(t1) ∈ t 2, then s1 t1 = s2 t2 , or equivalently, m1 t1 = m2 t2 , r 1 t1 = r2 t2 , and q1 t1 = q2 t2 (5.3) the mapping f is referred to as an embedding mapping from s1 into s2. if the system s1 can be homomorphically embeddable into s2 and s2 into s1, then the systems s1 and s2 are said to be homophorhically equivalent. evidently, equal systems are homomorphically equivalent with the identity mapping on the time set as the canonical embedding mapping. proposition 5.1. if the embedding mapping f : t 1 → t 2 from the system s1 into s2 is bijective, then the systems are homomorphically equivalent. proof. it suffers to show that the inverse mapping f−1 : t 2 → t 1 is an embedding mapping from s2 into s1. because f is bijective, it is strictly increasing from t 1 into t 2; and its inverse f−1 is also a strictly increasing mapping from t 2 into t 1 satisfying for any t2 ∈ t 2, if t1 = f−1(t2) ∈ t 1, then m2 t2 = m1 t1 , r 2 t2 = r1 t1 , and q2 t2 = q1 t1 . therefore, f−1 : t 2 → t 1 is an embedding mapping from s2 into s1. qed. definition 5.3. a system s = {st = (mt, rt, qt) : t ∈ t} is said to be cyclic or periodic, provided that there is time tc > 0 such that st+tc = st, for any t ∈ t . evidently, in this case, for any natural number n ∈ n , ntc is also a period of the system s. the minimum period is named the period of s, denoted tc. 5.2 systemic emergence by systemic emergence, it means the properties of the whole system that parts of the system do not have. such a property might suddenly appear at a particular time moment. in terms of the system, the holistic emergence is mainly created and excited by the system’s specific organization of its parts and how these parts interact, supplement, and constrain on each other. it is a kind of effect of relevance, the organizational effect, and the structural effect. to be specific, let p represent such a property. it is defined on the entire system s. because of the system’s dynamic characteristics, the interactions between the system’s objects, relations, attributes, and the environment change constantly. so, the value of p also varies accordingly. define p (st) = { 1, if s has this property 0, otherwise , ∀t ∈ t (5.4) it satisfies the following properties: p (st) = 1 → p (st) = 1, ∀s < s, ∀t ∈ t (5.5) 32 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics that is, as long as a subsystem s has this property, the overall system s also has the property. that is another way to say that the whole is greater than the sum of parts. of course, there are also such properties that some subsystems have, while the overall system does not have. for instance, let p be a property the system s does not have. define p−1(1) = {s < s : s ̸= s} be the collection of all proper subsystems of s, where the superscript (−1) stands for the inverse operation. then, for any s ∈ p−1(1), s < s and p (s) = 1. however, p (s) = 0 does not satisfy equ. (5.5). when a property p is said to be system s’s holistic emergence, provided that for any subsystem s < s, p (st) = 0, ∀t ∈ t , and p (st0) = 1 for at least one t0 ∈ t . in the traditional static description of systems, there is another definition of systemic emergence. in particular, consider a monotonically increasing sequence of subsystems ∅ ̸= s1 < s2 < ... < sk < ... < s, there is a k0 such that p (sk) = 0, ∀k < k0, p (sk0) = 1 (5.6) in this case, the system s is said to have the emergent property p , while its parts sk, k < k0, do not share this property until they reach the whole sk0 of certain scale. from a detailed analysis, it follows readily that what is just presented is a special case of the systemic emergence of dynamic systems, where we can surely treat the collection of all the superscript k as a subset of the time index t so that the sequence {sk : k = 1, 2, 3, ...} of subsystems a subsystem of a dynamic system by letting st = st, for t = 1, 2, 3, ... from equ. (5.6), it follows that p (st) = 0, for t = 1, 2, 3, ... < k0, and p (sk0) = 1 that is, the property p emerges at the time moment k0. 5.3 stability of systems by stability, it means the maintenance or continuity of a certain measure of the dynamic system on a certain time scale. that is, under small disturbances, the measure does not undergo noticeable changes. let us look at the trajectory system of a single point, where the focus is how the point moves under the influence of an external force. if a disturbance is given to a portion of the trajectory that is not at a threshold point, then the disturbed trajectory will not differ from the original trajectory much. however, if the same disturbance is given to the trajectory at an extremely unstable extreme point, then a minor change in the disturbed value could cause major deviations in the following portion of the trajectory. thus, only the trajectory of motion at stable critical extrema is stable, while the trajectory systems with instable critical points, such as saddle point, are instable. advances in systems science and applications (2013) vol.13 no.1 33 in general, assume that an attribute q ∈ q of the system s satisfies that q is a real-valued function defined for each momentary system st, for any t ∈ t . let t0 ∈ t . if for any ε > 0, there is a δt0,ε > 0 such that |q(st0)− q(st)| < ε, ∀t ∈ (t0 − δt0,ε, t0 + δt0,ε) (5.7) then the attribute q of the system s is regionally stable over time at t0 ∈ t . if for any ε > 0, there is δ = δ(ε) = δε > 0 such that |q(st1)− q(st2)| < ε, ∀t1, t2 ∈ t such that |t1 − t2| < δε (5.8) then the attribute q is said to be holistically stable over time or uniformly stable over time. when no confusion can be caused, the previous concepts of stability of the attribute q are respectively referred to as that the system s is regionally stable at t0 ∈ t , or uniformly stable over t . in addition, the structural stability of systems can also be defined. in particular, let d ∈ q be such that d : st → r+ is a positively real-valued function defined for each subsystem of the momentary system st, t ∈ t , satisfying (1)d(∅) = 0,where ∅ stands for the the subsystems with the empty set as their object set; (2)∀s1, s2 < s, if s1 < s2, then d(s1) ≤ d(s2); (3)∀s1, s2 < s, if s1△s2 = ∅, then d(s1 ∪ s2) = d(s1) + d(s2). it can be readily seen that such an attribute d can be employed to measure the difference between subsystems of s; and for any chosen s1 < s, and for any s2 ∈ s, the greater the attribute value d(s1△s2) is, the more different the systems s1 and s2 are. by making use of such a d ∈ q, which satisfies the previous properties, for any chosen s ∈ s and any real number δ > 0, we can define the δ -neighborhood nbrdd(s, δ) of s as follows: nbrdd(s, δ) = {s′ ∈ s : d(s′△s) < δ} (5.9) the attribute q ∈ q of the system s is said to be structurally stable in the neighborhood of a subsystem s < s with respect to attribute d ∈ q, provided that for any ε > 0, there is δ = δε > 0 such that |q(s)− q(s′)| < ε, ∀s′ ∈ nbrdd(s, δε) (5.10) similarly, the concepts of regionally structural stability and uniformly structural stability of a system s at all of its subsystems can be defined. if a system s is both uniformly stable over time and uniformly structural stable, then the system is referred to as a uniformly stable system. 34 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics 5.4 evolution of systems when systems’ evolution is concerned with, the focus is how the system develops over time. speaking rigorously, each system can be described by using the dynamic format in equ.(2.1). as for those systems, which do not seem to change with time, when seen from the angle of dynamics, they are either evolving with time extremely slowly or considered within a very short period of time. that is how they project a mistakenly incorrect sense of being static. the evolution of systems takes two main forms: one is the transition of the system from one structure or form to another structure or form, and the other is that a system appears from its earlier state of non-existence and grows from a immature state to a more mature state. both of these two forms of evolution can be described and illustrated by using the language of the dynamic systems in equ.(2.1). assume that the system of our concern is indeed described in the form of equ.(2.1). then for the first form of evolution, the relation set rt and the attribute set qt change with time, while for the second form of evolution, the elements in the object set mt develop over time together of course with changes of the relation and attributes sets rt and qt. they represent two specific cases of the evolution of dynamic systems. 5.5 boundary of systems by boundary, it stands for the separation between a system and its environment. the boundary is a part of the system and interacts with the environment more closely than the interior of the system. by reviewing the definition of environments in definition 4.3, we can define the boundary of a system as follows. assume that s < s is a subsystem of the system s over the time span t . let us denote the subsystem s and the system s respectively as st = (ms t , r s t , q s t ) and s = (m,r,q), for t ∈ t . for the sake of convenience of communication, we write s = (ms, rs, qs) for a fixed time moment. then, the (external) environment es of s within the system s is defined to be: es = {(m −ms, r|me , q|me s ,re ) : r ∈ r(r|ms /∈ rs), q ∈ q(q|ms /∈ qs)} = (me , re , qe) where me = m − m, re = r|me , and qe = q|me ,re , for r ∈ r(r|me /∈ rs) and q ∈ q(q|me /∈ qs). assume that there is an attribute µ ∈ qt of the system s such that it assigns each relation to a positive real number µ(r) : rt → r+. intuitively, this attribute µ is an index that measures the intensity of each relation of the objects of s. now, let us fix a threshold value µ0 for the relational intensity, then by combining with the concept of weak interactions between systems (definition 4.1), we can obtain the set of all relations in s of intensity at least µ0 that relate the subsystem s advances in systems science and applications (2013) vol.13 no.1 35 and its environment es as follows: rb = {r ∈ r : r|ms ∈ rs, supp(r) ∩ms ̸= ∅ ̸= supp(r) ∩me , µ(r) ≥ µ0} (5.11) where supp(r) stands for the support of the relation r, which is the set of all objects that appear in the relation r. the set of all objects of s that interact with the environment es of intensity of at least µ0 is given by the following: mb = ∪ r∈rb supp(r)−me (5.12) then the system ∂s = (mb, rb|mb , qb|mb,rb|mb ) (5.13) satisfies that ∂s < s, and is referred to as the boundary (system) of s within the system s. the intensity of its interaction with the environment is no less than µ0, while any other subsystem of s interacts with the environment es with strictly less intensity than µ0. if in the evolution of the system, ∂s has good stability, then we say that the boundary of the system is clear. otherwise, we say that the boundary of the system is fuzzy. symbolically, let q ∈ q be an attribute and t0 ∈ t chosen. if for any ε > 0, there is δ = δt0,ε > 0 such that |q(∂st0)− q(∂st)| < ε, ∀t ∈ (t0 − δt0,ε, t0 + δt0,ε) (5.14) then we say that the boundary of the system s at time t0 ∈ t is definite. otherwise, the boundary is said to be fuzzy at t0 ∈ t . it is not hard to see that the definiteness of a system’s boundary is defined by using the stability of the boundary system ∂s. therefore, the stability of the boundary system at one time moment corresponds to the definiteness of the system’s boundary. similarly, we can study the concepts of regional definiteness of boundaries over time and structural definiteness of boundaries. 6 some final words by using the traditional set theory, the static structures and relations of systems are described. then, the dynamic changes of systems are studied by using the concept of time systems. in this paper, on the basis of the concept of systems, we describe the dynamic structure and change of a general system from a more delicate angle by determining the objects, relations, and attributes. by considering the affects of environment, we establish the basic concepts of systems using the idea of dynamic sets that are related to the variable of time. then on top of these developed concepts, we revisit the description and analysis of some of the most fundamental properties of systems, while making the necessary comparisons 36 xiaojun duan: systems defined on dynamic sets and analysis of their characteristics with the studies of systems developed on the classical set theory. as for the analysis of systems properties under the new setting, additional rigorous deduction of various results of systems is badly needed in order to enrich the relevant systems theory. acknowledgements this work is supported by the national science foundation of china (60974124), the program for new century excellent in university, and the project-sponsored by srf for rocs, sem in china. references [1] lin. y. (1999), general systems theory: a mathematical approach, new york: kluwer academic and plenum publishers. [2] mickens. r. e. (1990), mathematics and science, singapore: world scientific. [3] quastler. h. (1965), general principles of systems analysis, in: t. h. waterman and h. [4] j. morrowits (eds.), theoretical and mathematical biology, new york: blaisdell publishing. [5] zadeh. l. (1962), “from circuit theory to systems theory”, proc. ire, vol.50, pp.856-865. corresponding author xiaojun duan can be contacted at: xjduan@nudt.edu.cn. adv syst sci appl 2017; 4:46–60 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/510 threshold analysis of a stochastic epidemic model with delay and temporary immunity rui xue, fengying wei ∗ college of mathematics and computer science, fuzhou university, fuzhou 350116, p.r. china abstract: a stochastic susceptible-infected-recovered model is formulated and investigated when the temporary immunity is fixed for the population in this paper. the existence and uniqueness of the global positive solution has been checked with probability one for any initial value. and the sufficient conditions for the extinction and the persistence of the stochastic epidemic model with temporary immunity are derived by constructing lyapunov functions and the generalized ito’s formula, where the threshold of the persistence does not depend on the temporary immunity, while the densities of the infected and recovered are obviously dependent on the temporary immunity when given a perturbation. illustrative examples and simulations show that the perturbations make the properties of the stochastic epidemic model different from the deterministic one. keywords: stochastic epidemic model; delay; persistence; extinction; threshold 1. introduction the susceptible-infected-removed epidemic model is one of the most important models in epidemiological patterns and disease control. epidemic models with delay always make the description of the epidemiological patterns more realistic, more interesting and more complicated. some authors investigated the effects of disease latency or immunity, in which the models are described by systems of ordinary differential equations with delay, for instance [1–10], and systems without delay [11–15]. among these recent results we would like to mention the work by wen and yang [10], in which they considered a type of delayed susceptible-infected-recovered epidemic model with temporary immunity: ṡ(t) = λ− µ1s(t)− βs(t)i(t) + γi(t− τ)e−µ3τ , i̇(t) = βs(t)i(t)− (µ2 + γ)i(t), ṙ(t) = γi(t)− γi(t− τ)e−µ3τ − µ3r(t), (1.1) where s(t) is the number of the susceptible individuals to the disease, i(t) represents the number of the individuals who are infected and r(t) denotes the recovered individuals who have been removed from the possibility of infection, λ denotes a constant input of new members into the population, µ1, µ2, µ3 represent the death rates of the susceptible, the infected, the recovered, and they assume that µ1 ≤ min{µ2, µ3} from the biological point of view, β denotes the transmission coefficient between compartments s and i , γ denotes the recovery rate of the infected individuals, i.e., the rate at which the individuals move from the infected compartment to the recovered compartment, τ > 0 is the length of immunity ∗corresponding author: weifengying@fzu.edu.cn http://ijassa.ipu.ru/ojs/ijassa/article/view/510 threshold analysis of a stochastic epidemic model with delay and temporary 47 period. wen and yang [10] assume that λ, µi, β, γ are all positive constants in model (1.1), and they obtain the expression of the basic reproduction number in terms of parameters, that is, r0 = βλ µ1(µ2+γ) , which measures how fast the diseases spread in the deterministic model provided that the epidemics take place. some common phenomena in the real world, such as the weather fluctuations, temperature changes and perturbations caused by human beings, can not been ignored when it comes to the dynamics of the infectious diseases. from both biological and mathematical perspectives, the stochastic models have more reasonable patterns and have been investigated by different approaches presented in the recent literatures [16–20]. motivated by model (1.1) and the approaches mentioned in [16–20], we try to investigate several properties of model (1.2) when transmission coefficient β in (1.1) is perturbed by the white noise β + σξ(t): ds(t) = ( λ− µ1s(t)− βs(t)i(t) + γi(t− τ)e−µ3τ ) dt− σs(t)i(t)db(t), di(t) = ( βs(t)i(t)− (µ2 + γ)i(t) ) dt+ σs(t)i(t)db(t), dr(t) = ( γi(t)− γi(t− τ)e−µ3τ − µ3r(t) ) dt, (1.2) where ξ(t) = db(t) dt , andb(t) is the independent brownian motion, and σ is the intensity of the white noise when model (1.2) works on a complete probability space (ω,f , {ft}t≥0,p) with a filtration {ft}t≥0 satisfying the usual conditions (i.e., it is right continuous and increasing function while f0 contains all p-null sets). what we concern about in model (1.2) in this paper are listed below, of course, for any positive initial value herewith: • existence and uniqueness of a global nonnegative solution with probability one; • the sufficient conditions that guarantees the extinction of the diseases; • figuring out the threshold presented in terms of model parameters; • demonstrating illustrative examples and their realizations to support the validity of the main results. 2. existence and uniqueness of the positive solution for model (1.2), the sum of three equations gives d(s(t) + i(t) +r(t)) dt ≤ λ− µ1(s(t) + i(t) +r(t)), (2.3) then one can obtain that s(t) + i(t) +r(t) ≤  λ µ1 , s(0) + i(0) +r(0) ≤ λ µ1 , s(0) + i(0) +r(0), s(0) + i(0) +r(0) > λ µ1 . (2.4) we denote n = max { λ µ1 , s(0) + i(0) +r(0) } , (2.5) therefore, s(t) ≤ n, i(t) ≤ n,r(t) ≤ n . we notice that the first two equations in model (1.2) do not depend on the third equation, we omit them without loss of generality. thus, we copyright c© 2017 assa. adv syst sci appl (2017) 48 r. xue, f.y. wei only discuss the simplified model of having two equations:{ ds(t) = ( λ− µ1s(t)− βs(t)i(t) + γi(t− τ)e−µ3τ ) dt− σs(t)i(t)db(t), di(t) = ( βs(t)i(t)− (µ2 + γ)i(t) ) dt+ σs(t)i(t)db(t). (2.6) in order to investigate other properties, we firstly study the fundamental property of the solution of model (2.6). next we will present the existence and uniqueness of the solution of model (1.2). theorem 2.1: for any initial value s(0) > 0 and i(ζ) ≥ 0 for all ζ ∈ [−τ, 0) with i(0) > 0, model (2.6) admits a unique positive solution (s(t), i(t)) on t > 0 and the solution will remain in r2 + with probability one, that is to say, (s(t), i(t)) ∈ r2 + for all t > 0 almost surely. proof according to the approach mentioned in [13–15, 21]. the proof will go as follows. since the coefficients of model (2.6) satisfy the local lipschitz conditions, then for any initial value s(0) > 0 and i(ζ) ≥ 0 for all ζ ∈ [−τ, 0) with i(0) > 0, there is a unique local solution (s(t), i(t)) on the half-closed interval [−τ, τe), where τe represents the explosion time of the solution. to verify the solution is global, what we prepare to check that explosion time τe =∞ holds almost surely. to this end, let k0 ≥ 1 be sufficiently large such that each component of the initial value (s(0), i(ζ)) always lies within the interval [ 1 k0 , k0]. for each integer k ≥ k0, let us define the following stopping time τk = inf { t ∈ [0, τe) ∣∣∣∣s(t) /∈ ( 1 k , k ) or i(t) /∈ ( 1 k , k )} . (2.7) throughout this paper, we set inf ∅ =∞ (as usual ∅ represents the empty set). obviously, τk is an increasing function as k →∞. let τ∞ be the limit of τk as k →∞. it is obvious that τ∞ ≤ τe is valid almost surely. in order to complete the proof, we only need to show τ∞ =∞ a.s.. we start the proof by contradiction, if this assertion is false, then there exists a pair of constants t > 0 and ε ∈ (0, 1) such that p{τ∞ ≤ t} > ε. (2.8) thereby, there is an integer k1 ≥ k0 such that p{τk ≤ t} ≥ ε for all k ≥ k1. (2.9) let us define a c2-function v : r2 + → r+ as follows: v (s(t), i(t)) = s(t)− a− a ln s(t) a + i(t)− 1− ln i(t) + γe−µ3τ ∫ t t−τ i(r)dr, (2.10) where a is a positive constant to be determined later. the nonnegativity of µ− 1− log µ is obvious for all µ > 0. let k ≥ k0 and t > 0 be arbitrary positive constants. the application of itô’s formula gives dv (s(t), i(t)) = ( 1− a s(t) ) ds(t) + a 2s2(t) (ds(t))2 + ( 1− 1 i(t) ) di(t) + 1 2i2(t) (di(t))2 + γi(t)e−µ3τdt− γi(t− τ)e−µ3τdt := lv (s(t), i(t))dt+ σ(ai(t)− s(t))db(t), (2.11) copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 49 where lv (s(t), i(t)) is defined from r2 + to r and computed as lv (s(t), i(t)) = ( 1− a s(t) )( λ− µ1s(t)− βs(t)i(t) + γi(t− τ)e−µ3τ ) + aσ2 2 i2(t) + ( 1− 1 i(t) )( βs(t)i(t)− (µ2 + γ)i(t) ) + σ2 2 s2(t) + γi(t)e−µ3τ − γi(t− τ)e−µ3τ = λ− µ1s(t)− (µ2 + γ)i(t) + γi(t− τ)e−µ3τ − aλ s(t) + µ1a+ aβi(t)− aγe−µ3τ i(t− τ) s(t) − βs(t) + µ2 + γ +γi(t)e−µ3τ − γi(t− τ)e−µ3τ + σ2 2 (s2(t) + ai2(t)) ≤ λ + µ1a+ µ2 + γ + (aβ + γe−µ3τ − (µ2 + γ))i(t) + σ2 2 (1 + a)n2. (2.12) choosing a positive constant a = µ2 + γ(1− e−µ3τ ) β (2.13) such that aβ + γe−µ3τ = µ2 + γ, then we have lv (s(t), i(t)) ≤ λ + µ1a+ µ2 + γ + σ2 2 (1 + a)n2 := k, (2.14) where k > 0 is a constant. the remainder of the proof follows the similar approach given in [22]. remark 2.1: theorem 2.1 implies that the stochastic model (1.2) admits a unique solution (s(t), i(t), r(t)) for any initial value (s(0), i(0), r(0)) ∈ r3 +, i(ζ) ≥ 0 for all ζ ∈ [−τ, 0). each component of the solution (s(t), i(t), r(t)) is positive for all t ≥ 0 almost surely. according to the statement of theorem 2.1, if the sum of each component with the initial value is almost surely bounded, say s(0) + i(0) +r(0) ≤ λ µ1 , so the region γ∗ = { (s(t), i(t), r(t)) ∈ r3 + : s(t) + i(t) +r(t) ≤ λ µ1 , t ≥ 0 } (2.15) is the positively invariant set of model (1.2). 3. the sufficient conditions of the extinction of the diseases let r̃0 = βλ µ1(µ2 + γ) − σ2λ2 2µ2 1(µ2 + γ) , (3.16) which can be presented in terms of the basic reproduction number of the deterministic model (1.1): r̃0 = r0 − σ2λ2 2µ2 1(µ2 + γ) . (3.17) copyright c© 2017 assa. adv syst sci appl (2017) 50 r. xue, f.y. wei for the stochastic model (1.2), we will investigate the sufficient conditions of the extinction of the disease. in other words, the infected individuals will eventually recover from the attack of the disease, and the capacity of the total population will keep kind of constant in the long run. theorem 3.1: let (s(t), i(t), r(t)) be the solution of model (1.2) with initial value (s(0), i(0), r(0)) ∈ γ∗. if the intensity of the white noise satisfies σ2 > β2 2(µ2 + γ) (3.18) or r̃0 < 1, σ2 ≤ βµ1 λ , (3.19) then the density of the infected declines exponentially lim sup t→∞ ln i(t) t ≤ −(µ2 + γ) + β2 2σ2 < 0 a.s. (3.20) or lim sup t→∞ ln i(t) t ≤ (µ2 + γ)(r̃0 − 1) < 0 a.s. (3.21) respectively; and lim sup t→∞ s(t) = λ µ1 a.s.. (3.22) proof adding up those three equations of model (1.2) gives d ( s(t) + i(t) + γe−µ3τ ∫ t t−τ i(r)dr ) = [λ− µ1s(t)− (µ2 + γ(1− e−µ3τ ))i(t)]dt. (3.23) we integrate the above from 0 to t and divided by t, which implies λ− µ1〈s(t)〉 − (µ2 + γ(1− e−µ3τ ))〈i(t)〉 = 1 t ( s(t) + i(t) + γe−µ3τ ∫ t t−τ i(r)dr − s(0)− i(0)− γe−µ3τ ∫ 0 −τ i(r)dr ) , (3.24) and then derives that 〈s(t)〉 = λ µ1 − µ2 + γ(1− e−µ3τ ) µ1 〈i(t)〉+ ϕ(t), (3.25) where ϕ(t) = 1 µ1t ( s(t) + i(t) + γe−µ3τ ∫ t t−τ i(r)dr − s(0)− i(0)− γe−µ3τ ∫ 0 −τ i(r)dr ) . (3.26) the strong law of large numbers and the fact s(0) + i(0) ≤ λ µ1 show that lim t→∞ ϕ(t) = 0 a.s.. (3.27) copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 51 for the second equation of model (1.2), the generalized itô’s formula leads to d ln i(t) = ( βs(t)− (µ2 + γ)− σ2s2(t) 2 ) dt+ σs(t)db(t). (3.28) integrating (3.28) from 0 to t and divided by t gives ln i(t) t = β〈s(t)〉 − (µ2 + γ)− σ2 2 〈s2(t)〉+ σ t ∫ t 0 s(r)db(r) + ln i(0) t . (3.29) we denote m(t) := σ ∫ t 0 s(r)db(r), (3.30) where m(t) is a local continuous martingale with m(0) = 0. moreover, lim sup t→∞ 〈m(t),m(t)〉t t ≤ σ2λ2 µ2 1 <∞ a.s.. (3.31) the strong law of large numbers for martingales presented in [23] thus implies that lim sup t→∞ m(t) t = 0 a.s.. (3.32) if σ2 > β2 2(µ2+γ) , rewriting (3.29) gives ln i(t) t = −σ 2 2t ∫ t 0 ( s(r)− β σ2 )2 dr − (µ2 + γ) + β2 2σ2 + σ t ∫ t 0 s(r)db(r) + ln i(0) t ≤ −(µ2 + γ) + β2 2σ2 + σ t ∫ t 0 s(r)db(r) + ln i(0) t . (3.33) taking superior limit on both sides of (3.33), then we get that lim sup t→∞ ln i(t) t ≤ −(µ2 + γ) + β2 2σ2 < 0 a.s.. (3.34) copyright c© 2017 assa. adv syst sci appl (2017) 52 r. xue, f.y. wei if r̃0 < 1 and σ2 ≤ βµ1 λ hold, the schwarz inequality and substituting (3.25) into (3.29) yield that ln i(t) t ≤ β〈s(t)〉 − (µ2 + γ)− σ2 2 〈s(t)〉2 + σ t ∫ t 0 s(r)db(r) + ln i(0) t = βλ µ1 − (µ2 + γ)− β(µ2 + γ(1− e−µ3τ )) µ1 〈i(t)〉+ βϕ(t) + σ t ∫ t 0 s(r)db(r)− σ2 2 [ λ µ1 − µ2 + γ(1− e−µ3τ ) µ1 〈i(t)〉+ ϕ(t) ]2 = βλ µ1 − (µ2 + γ)− β(µ2 + γ(1− e−µ3τ )) µ1 〈i(t)〉+ βϕ(t) + σ t ∫ t 0 s(r)db(r)− σ2 2 [ λ2 µ2 1 + ( µ2 + γ(1− e−µ3τ ) µ1 〈i(t)〉 )2 +ϕ2(t)− 2λ(µ2 + γ(1− e−µ3τ )) µ2 1 〈i(t)〉 + 2λ− 2(µ2 + γ(1− e−µ3τ ))〈i(t)〉 µ1 ϕ(t) ]2 , (3.35) which gives that ln i(t) t ≤ βλ µ1 − σ2λ2 2µ2 1 − µ2 + γ(1− e−µ3τ ) µ1 ( β − σ2λ µ1 ) 〈i(t)〉 −(µ2 + γ) + ψ(t), (3.36) where ψ(t) = βϕ(t) + σ t ∫ t 0 s(r)db(r) −σ 2 2 [ ϕ2(t) + 2λ− 2(µ2 + γ(1− e−µ3τ ))〈i(t)〉 µ1 ϕ(t) ] . (3.37) it is easy to check that lim t→∞ ψ(t) = 0 a.s.. (3.38) we take superior limit on both sides of (3.35) and derive that lim sup t→∞ ln i(t) t ≤ (µ2 + γ)(r̃0 − 1) < 0 a.s.. (3.39) further, the density of the infected individuals declines to zero, and the density of the susceptible individuals keeps constant in a long time run, mathematically speaking, these two expressions lim t→∞ i(t) = 0, lim t→∞ s(t) = λ µ1 (3.40) hold almost surely when condition (3.18) or condition (3.19) is satisfied. remark 3.1: wen and yang [10] had shown that the disease-free equilibrium e0( λ µ1 , 0, 0) of the deterministic model (1.1) was globally asymptotically stable if the basic reproduction number r0 < 1 (i.e., the disease disappeared under the condition r0 < 1). while for the stochastic copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 53 model (1.2), we obtain that the disease disappears when conditions r̃0 < 1 and σ2 ≤ βµ1 λ are being satisfied in this chapter. this indicates that the condition of extinction to the stochastic model (1.2) is weaker than that of the corresponding deterministic model (1.1). 4. the threshold for the persistence of the diseases theorem 4.1: let (s(t), i(t), r(t)) be any solution of model (1.2), (s(0), i(0), r(0)) ∈ γ∗ be any initial value with i(ζ) ≥ 0 for all ζ ∈ [−τ, 0). we assume that r̃0 > 1, and also assume that the intensity of the white noise satisfies σ2 ≤ βµ1 λ , (4.41) then the densities of the infected and the recovered individuals have the following properties: ĩ∗ ≤ lim inf t→∞ 〈i(t)〉 ≤ lim sup t→∞ 〈i(t)〉 ≤ ĩ∗ a.s. (4.42) and r̃∗ ≤ lim inf t→∞ 〈r(t)〉 ≤ lim sup t→∞ 〈r(t)〉 ≤ r̃∗ a.s., (4.43) where ĩ∗ = µ1(µ2 + γ) β(µ2 + γ(1− e−µ3τ )) (r̃0 − 1), r̃∗ = γ(1− e−µ3τ ) µ3 ĩ∗, (4.44) and ĩ∗ = µ2 1(µ2 + γ) (βµ1 − σ2λ)(µ2 + γ(1− e−µ3τ )) (r̃0 − 1), r̃∗ = γ(1− e−µ3τ ) µ3 ĩ∗. (4.45) proof from (3.28), the second equation of model (1.2) implies that d ln i(t) ≥ ( βs(t)− µ2 − γ − σ2λ2 2µ2 ) dt+ σs(t)db(t), (4.46) integrating (4.46) on both sides and substituting (3.25) into the integration, which derive that ln i(t)− ln i(0) t ≥ β〈s(t)〉 − µ2 − γ − σ2λ2 2µ2 + σ t ∫ t 0 s(r)db(r) = βλ µ1 − ( µ2 + γ + σ2λ2 2µ2 ) − β(µ2 + γ(1− e−µ3τ )) µ1 〈i(t)〉 +βϕ(t) + σ t ∫ t 0 s(r)db(r). (4.47) inequality (4.47) can be rewritten as 〈i(t)〉 ≥ µ1 β(µ2 + γ(1− e−µ3τ )) ( βλ µ1 − ( µ2 + γ + σ2λ2 2µ2 ) − ln i(t)− ln i(0) t + βϕ(t) + σ t ∫ t 0 s(r)db(r) ) . (4.48) copyright c© 2017 assa. adv syst sci appl (2017) 54 r. xue, f.y. wei according to (2.15), we have −∞ < ln i(t) < ln λ µ1 , lim t→∞ σ t ∫ t 0 s(r)db(r) = 0, (4.49) and lim t→∞ ϕ(t) = 0. (4.50) taking the inferior limit on both sides of (4.48), we have lim inf t→∞ 〈i(t)〉 ≥ µ1(µ2 + γ) β(µ2 + γ(1− e−µ3τ )) (r̃0 − 1) := ĩ∗. (4.51) on the other hand, integrating the first equation of (4.46) from 0 to t on both sides yields ln i(t)− ln i(0) t = β〈s(t)〉 − µ2 − γ − σ2 2 〈s2(t)〉+ σ t ∫ t 0 s(r)db(r). (4.52) we rewrite (3.35), then we get that 〈i(t)〉 ≤ µ2 1 (µ2 + γ(1− e−µ3τ ))(βµ1 − σ2λ) × ( βλ µ1 − ( µ2 + γ + σ2λ2 2µ2 1 ) − ln i(t)− ln(0) t + ψ(t) ) , (4.53) taking the superior limit on both sides of above equation, then one can obtain that lim sup t→∞ 〈i(t)〉 ≤ µ2 1 (µ2 + γ(1− e−µ3τ ))(βµ1 − σ2λ) (r̃0 − 1) := ĩ∗. (4.54) the last equation of model (1.2) yields 〈r(t)〉 = γ µ3 〈i(t)〉 − γe−µ3τ µ3 〈i(t− τ)〉 − r(t)−r(0) µ3t , (4.55) then lim inf t→∞ 〈r(t)〉 = γ(1− e−µ3τ ) µ3 lim inf t→∞ 〈i(t)〉 ≥ γ(1− e−µ3τ ) µ3 ĩ∗ = r̃∗, (4.56) lim sup t→∞ 〈r(t)〉 = γ(1− e−µ3τ ) µ3 lim sup t→∞ 〈i(t)〉 ≤ γ(1− e−µ3τ ) µ3 ĩ∗ = r̃∗. (4.57) remark 4.1: both theorem 2 and theorem 3 have the common condition σ2 ≤ βµ1 λ therewith. while the opposite properties of these two theorems depend on the value of r̃0, that is, r̃0 < 1 indicates the extinction of the disease and r̃0 > 1 means the persistence of the disease. the expression r̃0 plays the role of threshold of model (1.2). copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 55 5. illustrative examples and their realizations example 1 let the parameters of model (1.2) be λ = 0.2, µ1 = µ2 = µ3 = 0.2, β = 0.8, γ = 0.5, σ = 0.6, and the initial value be (s(0), i(0), r(0)) = (0.1, 0.6, 0.3). we can verify that r̃0 = r0 − σ2λ2 2µ2 1(µ2 + γ) = 0.8857 < 1, σ2 = 0.36 ≤ βµ1 λ = 0.8. (5.58) condition (3.19) of theorem 1.2 is satisfied, the solution of model (1.2) has the properties as follows: lim sup t→∞ ln i(t) t ≤ (µ2 + γ)(r̃0 − 1) = −0.08 < 0 a.s. (5.59) and lim sup t→∞ s(t) = λ µ1 = 1 a.s. (5.60) figure 5.1 indicates that the infected vanishes exponentially, and the susceptible individuals reach their maximum when given a long time run. obviously, the basic reproduction number of the deterministic model (1.1) can be computed as r0 = 1.1429 > 1 in this case. and the endemic equilibrium e∗ of model (1.1) is globally asymptotically stable according to theorem 5.2 in [10]. we conclude that in the case of medium perturbation, say σ = 0.6, the density of the susceptible approaches one almost surely, and is much higher than that of the deterministic model. compared with the deterministic model, the density of the infected declines fast to zero with the exponential rate −0.08 at early time scale 1.5× 104 days. and the density of the recovered is somehow affected by the infected and ends up at zero at 2.5× 104 days instead of persistence for the deterministic model. example 2 we keep the initial value and other parameters same as shown in example 1 except for σ = 0.9. here σ2 ≥ β2 2(µ2 + γ) = 0.4571, (5.61) and condition (3.18) of theorem 1.2 is being satisfied, then the solution of model (1.2) admits the following property: lim sup t→∞ ln i(t) t ≤ −(µ2 + γ) + β2 2σ2 = −0.3349 < 0 a.s. (5.62) the corresponding simulations would be shown in figure 5.2 to support the main results of theorem 1.2 we got in the previous section. under large perturbation, say σ = 0.9 in this case, we find that the density of the susceptible behaves the similar dynamics. while the curve of the infected shows more sharper than that in example 1, and still decays exponentially with a larger rate−0.3349 at early time scale 1× 104 days. and the density of the recovered is close to zero at 2× 104 days compared with the deterministic model. example 3 let the intensity of the white noise be σ = 0.1, and the initial value and other parameters be the same as shown in example 1. here r̃0 = r0 − σ2λ2 2µ2 1(µ2 + γ) = 1.1357 > 1, σ2 = 0.01 ≤ βµ1 λ = 0.8. (5.63) copyright c© 2017 assa. adv syst sci appl (2017) 56 r. xue, f.y. wei 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the stochastic system s(t) i(t) r(t) 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the deterministic system s(t) i(t) r(t) fig. 5.1. realizations for model (1.2) with r̃0 < 1 and model (1.1) with r0 > 1 respectively. by theorem 4.1, the following property of model (1.2) holds: 0.0817 ≤ lim inf t→∞ 〈i(t)〉 ≤ lim sup t→∞ 〈i(t)〉 ≤ 0.0827 a.s. (5.64) if we keep the initial value and other parameters same as shown in example 1 except for µ2 = 0.3, γ = 0.4. we easily check that examples 1 and 2 still keep the same conclusions. while the persistence level of model (1.2) is lower than that in (5.64) by theorem 4.1. that is, 0.0638 ≤ lim inf t→∞ 〈i(t)〉 ≤ lim sup t→∞ 〈i(t)〉 ≤ 0.0646 a.s. (5.65) figure 5.3 reveals that the prevalence of the disease takes place under small perturbation of the white noise. the properties of the solutions for the stochastic model (1.2) and the deterministic model (1.1) demonstrate the similar behaviors when set a small perturbation. we would like to conclude that the density levels for the susceptible, the infected and the recovered are all alike when given a small perturbation environment. especially, the higher the recovery rate for the infected individuals, the more the infected, such as, the density copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 57 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the stochastic system s(t) i(t) r(t) 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the deterministic system s(t) i(t) r(t) fig. 5.2. realizations for model (1.2) with σ = 0.9 and model (1.1) with r0 > 1 respectively. of the infected varies from [0.0817, 0.0827] when γ = 0.5 (the green line in figure 5.4) to [0.0638, 0.0646] when γ = 0.4 (the blue line in figure 5.4). 6. conclusions in this paper, we work on the susceptible-infected-recovered model, where the individuals stayed in the recovered compartment finally lost temporary immunity returned to the susceptible compartment. the research results of this paper demonstrate that the existence and uniqueness of the global positive solution of model (1.2) has nothing to do with the temporary immunity due to the construction of lyapunov function (2.10) therewith. while no matter how big the temporary immunity period is, the sufficient condition of the extinction of the diseases merely depends on the parameters of model (1.2), say condition (3.18) or (3.19) is λ, µ1, µ2, β, γ, σdependent, and τ -dependent instead in theorem 2. copyright c© 2017 assa. adv syst sci appl (2017) 58 r. xue, f.y. wei 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the stochastic system s(t) i(t) r(t) 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the deterministic system s(t) i(t) r(t) fig. 5.3. realizations for model (1.2) with σ = 0.1 and model (1.1) with r0 > 1 respectively. we therefore conclude that, in theorem 3, the expressions of ĩ∗ and ĩ∗ are the inverse functions of factor 1− e−µ3τ , where r̃0 is a τ -independent expression; and r̃∗ and r̃∗ are respectively the saturated functions of 1− e−µ3τ . the illustrative examples have shown that the larger the perturbation, the sharper the infected individuals decline exponentially, the earlier the recovered individuals meet extinction. acknowledgements this work is supported by the national natural science foundation of china (grants no. 11201075 and 61773122), natural science foundation of fujian province of china (grant no. 2016j01015). copyright c© 2017 assa. adv syst sci appl (2017) threshold analysis of a stochastic epidemic model with delay and temporary 59 0 0.5 1 1.5 2 2.5 3 3.5 4 t,days ×104 0 0.2 0.4 0.6 0.8 1 1.2 the stochastic system i(t) i(t) fig. 5.4. realizations for model (1.2) with γ = 0.5 (the green) and γ = 0.4 (the blue) respectively. references 1. kyrychko y. & blyuss k. (2005) global properties of a delayed sir model with temporary immunity and nonlinear incidence rate. nonlinear anal: real world appl., 6, 495–507. 2. xu r. & ma z. (2010) global stability of a delayed seirs epidemic model with saturation incidence rate. nonlinear anal: real world appl., 61, 229–239. 3. muroya y., enatsu y. & nakata y. (2011) global stability of a delayed sirs epidemic model with a non-monotonic incidence rate. j. math. anal. appl., 377, 1–14. 4. lahrouz a. (2015) dynamics of a delayed epidemic model with varying immunity period and nonlinear transmission. int. j. biomath., 8, 1550027. 5. ma w., song m. & takeuchi y. (2004) global stability of an sir epidemic model with time delay. appl. math. lett., 17, 1141–1145. 6. guo h. & li m. (2006) global dynamics of a staged progression model for infectious diseases. math. biosci. eng., 3 (3), 513–525. 7. beretta e. & kuang y. (2001) modeling and analysis of a marine bacteriophage infection with latency period. nonlinear anal: real world appl., 2, 35–74. 8. beretta e. & takeuchi y. (1995) global stability of an sir epidemic model with time delays. j. math. biol., 33, 250–260. 9. beretta e. & takeuchi y. (1997) convergence results in sir epidemic models with varying population sizes. nonlinear anal: theory methods appl., 28, 1909–1921. 10. wen l. & yang x. global stability of a delayed sirs model with temporary immunity. chaos solitons fractals, 38, 221–226. 11. melnichenko o.a. & romanyukha a.a. (2008) a model of tuberculosis epidemiology: estimation of parameters and analysis of factors influencing the dynamics of an epidemic process. russian j. numer. anal. math. model., 23 (1), 63–75. 12. melnichenko o.a. & romanyukha a.a. (2009) a model of tuberculosis epidemiology: data analysis and estimation of parameters, math. models comput. simul., 1 (4), 428–444. 13. liu j.m., wei f.y. dynamics of stochastic seis epidemic model with varying population size. physica a. 2016; 464: pp. 241-250. 14. wei f.y., liu j.m. long-time behavior of a stochastic epidemic model with varying population size. physica a. 2017; 470: pp. 146-153. copyright c© 2017 assa. adv syst sci appl (2017) 60 r. xue, f.y. wei 15. chen l.h., wei f.y. persistence and distribution of a stochastic susceptible-infectedrecovered epidemic model with varying population size. physica a. 2017; 483: pp. 386397. 16. dalal n., greenhalgh d., mao x. a stochastic model of aids and condom use. j. math. anal. appl. 2007; 325: pp. 36-53. 17. lahrouz a., omari l., kiouach d., belmaati a. complete global stability for an sirs epidemic model with generalized non-linear incidence and vaccination. appl. math. comput. 2012; 218 (11): pp. 6519-6525. 18. zhang x., jiang d., alsaedi a., hayat t. stationary distribution of stochastic sis epidemic model with vaccination under regime switching. appl. math. lett. 2016; 59: pp. 87-93. 19. yu j., jiang d., shi n. global stability of two-group sir model with random perturbation. j. math. anal. appl. 2009; 360: pp. 235-244. 20. xue r., wei f.y. persistence and extinction of a stochastic sis epidemic model with double epidemic hypothesis. ann. appl. math. 2017; 33(1): pp. 77-89. 21. dalal n., greenhalgh d., mao x. a stochastic model of aids and condom use. j. math. anal. appl. 2007; 325: pp. 36-53. 22. mao x., marion g., renshaw e. environmental brownian noise suppresses explosions in population dynamics. stoch. process appl. 2002; 97; pp. 95-110. 23. mao x. stochastic differential equations and applications (2nd ed.), horwood, chichester, uk, 2007. copyright c© 2017 assa. adv syst sci appl (2017) introduction existence and uniqueness of the positive solution the sufficient conditions of the extinction of the diseases the threshold for the persistence of the diseases illustrative examples and their realizations conclusions advances in systems science and applications (2014) vol.14 no.3 230-243 robust design and its challenge for manufacturing system han-xiong li and xinjiang lu department of manufacturing eng & eng management, city university of hong kong, hong kong, pr china abstract robust performance is one of the most important concerns in design of any system in manufacturing industry. this performance can be achieved by the robust design. the paper will provide a brief overview of the robust design. it includes discussions about how to account for design uncertainty, and how to measure and evaluate robustness for both static and dynamic systems. by reviewing the strengths and weaknesses of different design methods, the challenges in this area will be discussed. keywords robust design, sensitivity analysis, stability, uncertainty 1 introduction in order to design and manufacture high quality productions at a minimal cost, increasingly accurate systems are required in practical industry. one serious problem in these systems is the inconsistent performance due to uncontrollable variations existing in the real-world, including manufacturing operations, variations in material properties and the operating environment. if these variations are not considered, they will degrade the performance and may result in a failure in practice [1]. thus, robust performance is one of the most important concerns in design of any system. the concept of the robust design was introduced by taguchi. the fundamental principle in the robust design is to improve the quality of a product by minimizing the effects of variations without eliminating the causes [2]. by making a design more robust to variations, it is possible to improve number of the eligible parts or use less experiment [3]. in past decades, much effort has been dedicated to the robust design. in general, robust design may be classified two groups, static model based design and dynamic model based design. • static model based robust design is relatively simple since it is irrelevant with time and only considers the effect of the static variations on performance ; • dynamic model based robust design is complex because it must consider the stability, flexibility and robustness of the system in the whole operating process. the aim of this paper is to review and compare the applicable methods, so that their underlined philosophy can be clearly shown and easily understood. though the review, the challenges in this area can be discussed. advances in systems science and applications (2014) vol.14 no.3 231 2 uncertainty in a robust design problem, the system includes three kinds of variables : • design variables (or control variable) ĺc since their nominal values can be selected between the range of upper and lower bounds, they are controllable ; • uncertainties that can not be adjusted by the designer, so they are uncontrollable ; • performances that are the objective of design, they depend on the system model, design variables and uncertainties. these uncertainties, which should be first identified to obtain the robust performance, mainly include : • noise it is caused by changes of operating conditions, such as, environmental temperature, pressure, humidity and material change, etc. thus, it is of stochastic nature. in a robust design problem, the system includes three kinds of variables : • model parameters it is often caused by manufacture error. parameters of a product can only be realized to a certain degree of accuracy due to machinery limitation. it could be either stochastic or non-stochastic. • model uncertainty it is often caused by approximations in the modelling process. basically, these uncertainties can have following three different natures : • the deterministic type defines the domains in which the uncertainties can vary ; • the probabilistic type defines probability measure describing the likelihood by which a certain event occurs ; • the possibilistic type defines fuzzy measures describing the possibility or membership grade by which a certain event can be plausible or believable. these three different types of uncertainties are usually modelled by crisp sets, probability distributions and fuzzy sets, respectively. 3 robust design the robust design framework is shown in figure 1. the key issue in this framework is the robust design strategy, which optimizes the design variables to make the system less robust to uncertainties. all existing robust design methods depend on the characteristics of a system given in table 1 and table 2. if a system model is unknown and design parameters are non-probabilistic type, then no robust design method is applicable to it. the indirect method for this case is data-based modelling methods that transfer it into the case of the known model. 232 han-xiong li : robust design and its challenge for manufacturing system 3.1 static model based robust design this robust design includes three different approaches : deterministic robust design, probabilistic robust design, and fuzzy analysis. table 1 existing robust design methods for static system system static model model known unknown uncertainty non-probabilistic probabilistic non-probabilistic probabilistic type type type type •monte carlo simulation •taguchi robust •deterministic •first and second method design robust methods order moment n/a •monte methods •fuzzy analysis methods carlo •probabilistic simulation sensitivity analysis • performance index table 2 existing robust design methods for static system system static model model known unknown uncertainty non-probabilistic probabilistic non-probabilistic probabilistic type type type type •monte carlo simulation •taguchi robust •deterministic •first and second method design robust methods order moment n/a •monte methods •fuzzy analysis methods carlo •probabilistic simulation sensitivity analysis • performance index 3.1.1 deterministic robust design the design parameters are deterministic. (1) euclidean norm method and conditional number method the methods obtain the system robustness by analytically measuring the sensitivity of a design using the gradient information of parameters. the sensitivity analysis is based on the taylor-series expansion of the performance. the euclidean norm method [1, 3] is to minimize the largest singular value σmaxof the sensitivity matrix. the condition number method [1, 4] is to minimize the condition number σmax σmin of the sensitivity matrix. caro, et al. [1] has compared these two methods when designing the damper, and shown that the euclidean norm is more suitable advances in systems science and applications (2014) vol.14 no.3 233 as the robust index than the condition number. (2) the sensitivity region measures method gunawan & azarm [5] proposed the sensitivity region measures for the single objective robust design optimization. this method first projects the desirable objective function space into the design parameter space, where the sensitivity region is constructed. then the most sensitivity direction in the sensitivity region can be found. this most sensitivity direction is actually a measurement of a robust performance. li, azarm & boyars [6] proposed another sensitivity region measures for robust optimization. this method first projects the design parameter space into the desirable objective variations domain, where the objective sensitivity region is constructed. then the maximum performance variation in the sensitivity region is found. this variation is used as a measurement of a robust performance. moreover, the comparison with the gunawanąŕs method is carried out. 3.1.2 fuzzy analysis the possibility methods are proposed to apply in areas where it is not possible to obtain accurate statistical data dueto restriction of resource or conditions. its foundation is possibility theory. the extension principle calculates the possibility distribution of the fuzzy response from the possibility distribution of the fuzzy input variables. recently, this method was applied to the robust optimal design to deal with the epistemic uncertainty [7]. also, this method was applied for the modeling of tolerances and clearances in the mechanism analysis [8]. the comparison of probability and possibility for design against catastrophic failure under uncertainty was presented by [9]. the review of this application can be found in he & qu [10] and beyer & sendhoff [11]. 3.1.3 probabilistic robust design this probabilistic robust design will be less conservative than the deterministic robust design since it makes use of the probabilistic information of parameters. (1) monte carlo simulation monte carlo methods are a class of computational algorithms that rely on repeated random sampling from probabilistic density function of each parameter to compute their results. then, an experiment is executed under the generated samples and the process data is collected. finally, the statistical method is used to estimate its probability density function from the process data. (2) first and second order moment methods the most widely used non-statistical uncertainty analysis method is moment methods. a taylor series expansion is employed to estimate response variance 234 han-xiong li : robust design and its challenge for manufacturing system based on variance of model parameters x as σfirstordery = ( ∂y ∂x ) |µx2σ(x) (3.1) σsecondordery = ( ∂y ∂x ) |µx2σ(x) (3.2) (3) probabilistic sensitivity analysis this method evaluates the effect of design variables on performances using the sensitivity information. probabilistic sensitivity analysis methods have been developed to provide insight into the probabilistic behavior of a model, which can be used to identify those non-significant variables and reduce the dimension of random design space. a review about probabilistic sensitivity analysis was presented by [12]. the sensitivity analysis includes : variance-based methods [13], probabilistic sensitivity coefficients [14], kullback-leibler entropy based probabilistic sensitivity analysis (psa) method [15]. (4) performance index various performance indexes are constructed to measure the performance and capability of a system. • suhąŕs information content the information axiom is used to evaluate quality of designs so that an appropriate design can be chosen from available design alternatives. according to the information axiom proposed by suh [16], a design candidate that has minimum information content should be selected. thus, the information content is regarded as a robust index. • design capability indices the six standard deviations(±3σ)is commonly a measure of process capability, which compares the variation of a process to the customer specifications through the equation cdl = usl− lsl 6σ (3.3) where usl is the upper specification limit, lsl is the lower specification limit, is the standard deviation of the process. we hope that cp is greater than one so that the process variation is less than the specification limits and the performance can satisfy the requirement. chen, et al. [18] extended the concept of measuring process capability to measure the approximate degree between the mean of the process and the target value. they proposed design capability indices (dci) as metrics for system performance and robustness. the indices cdu, cdl and cdk means that smaller is better, larger advances in systems science and applications (2014) vol.14 no.3 235 is better and target is better respectively, and are defined as cp = usl− lsl 6σ ,cdu = url− µ 3σ ,cdk = min(cdl, cdu) (3.4) where url and lrl are upper and lower requirement limits. the index is expected to be greater than unity so that the design will meet the requirement satisfactorily. forcing the index larger to unity is achieved by reducing performance deviation and/or locating the mean of performance deviation farther from requirement limits. (5) taguchi method taguchi robust design, also known as parameter design, is an approach to identify design variable values that satisfy a set of performance requirements despite variation in noise factor [2]. since 1960, taguchi methods have obtained a great success in improving the quality of products and design robustness. this method is based on experimental data and includes three parts : experimental design, quality loss function and signal-to-noise ratio. taguchi robust design approach for the variable design starts from the experimental design, where the orthogonal array is used for design. control factor resides in an inner array and noise factors conditions in an outer array. the experimental results in all combinations of control factors and noise factors are recorded. then, taguchi proposed a signal-to-noise ratio for measuring sensitivity analysis of response to variation of noise factors. based on the signal-to-noise ratio, the robust design is obtained. although taguchiąŕs method has obtained great success, there are certain assumptions and limitations associated with his methods. use of the taguchi method will not yield an accurate solution for design problems that embody highly nonlinear behavior [19]. the taguchi method has been criticized by the statistical community [20]. many of taguchiąŕs statistical methods, e.g., orthogonal arrays, linear graphs and accumulation analysis, are not statistically efficient [20]. shoemaker et al., [21] presented a combined single array for both control and noise factors instead of orthogonal array. 3.2 dynamic model based robust design these uncertainties are, in general, also dynamic in nature and correspond to variations in either external variables or internal process parameters [22]. 3.2.1 stability based design the recent integration of the steady-state design and the dynamic stability is to explicitly consider dynamic elements in the process design by use of the eigenvalue theory. blanco and bandoni [23] proposed the multi-period program to solve the lyapunovąŕs stability matrix equality that can guarantee the system stability 236 han-xiong li : robust design and its challenge for manufacturing system under uncertainty and disturbance. however, the accuracy of this method will depend on the degree of discretization. since lyapuovąŕs criteria for asymptotic stability can not be easily implemented within a design optimization framework, mohideen, perkins & pistikopoulos [24] and kokossis & floudas [25] proposed an alternative robust stability criteria based on the matrix measures, which can avoid the tedious calculation of all eigenvalues. matrix measures can provide a single upper bound for all eigenvalues. however, this bound should not be used because it is typically not tight, therefore, may result in an overestimation of the stability boundary. monnigmann & marquardt [26] and grosch, monnigmann & marquardt [27] proposed the stability design in the steady-state process optimization. this design employed manifolds method to figure out the bound of the parameter variations that guaranteed all eigenvlues of the process smaller than zero when the parameter variations were limited in this bound. 3.2.2 flexibility analysis the flexibility, which defines the ability to maintain feasible operation over a range of uncertain conditions, is a vital important characteristic for the operation of these plants. as the state of dimitriadis & pistikopoulos [22], the flexibility analysis problem generally consists of two tasks which are complementary to each other. • the first task is to determine if a given design can feasibly operate over the range of uncertainty considered. this problem is known as the flexibility test problem. • the second task is to calculate a measure to quantify the ability of the design to operate in the presence of uncertainty. this is known as the flexibility index problem and is usually tackled by establishing the maximum parameter range over which the design can operate feasible. the flexibility measure is often used to select the suitable design by comparing different design alternatives with respect to their flexible operation [28]. swaney and grossmann [29] defined the flexibility index for measuring the flexibility of steady-state processes where the uncertain parameters are described by bounds of a specified range of operation. this approach was also extended to the analysis of dynamic systems under time-varying uncertainties [22]. the stochastic flexibility index is a metric for quantifying the ability of a process to maintain feasible operation in the face of stochastic uncertainties [30]. 3.2.3 robustness index various robustness indexes can be designed to measure the degree to which a system can meet its design objectives despite external disturbance and uncertainties in its design parameters. since adequate robust index is a necessary part of the optimal process design, it is desirable to consider process resiliency assessment when determining the process structure and establishing the operating range advances in systems science and applications (2014) vol.14 no.3 237 the resiliency indexmeasures the effect of the disturbance on the control input u. skogestad & morari [31] and lewin [32] considered the ratio of the control input u to disturbance d as the resiliency index based on the linear process model. cao, rossiter and owens [33] applied it to select the control inputs. solovyev & lewin [34] extended this resiliency index into the nonlinear system. the condition number is defined as the ratio between the maximum and minimum singular values. the condition number provides a direct measure of the directionality of the system. a large condition number indicates that the gain of the plant changes significantly with the input direction and that the system is sensitive to input uncertainty [35]. the disturbance condition number is a measure of the input magnitude which is needed to reject a disturbance in the given direction, relative to rejecting a disturbance with the same magnitude, but in the direction requiring the least control effort. a small disturbance condition number are most effective for disturbance rejection [36]. the relative gain matrix (rga)was originally proposed by bristol [37]. its objective is to provide a measure of interactions for multivariable square systems [31, 35]. if the plant has large rga elements within the frequency range where effective control is desired, then it is not possible to achieve good reference tracking with feedforward control because of strong sensitivity to diagonal input uncertainty. manousiouthakis et al. [38] generalized the concept of the rga to block relative gain which is capable of handling partially decentralized control systems. chen and yu [39] extended this method to non-square multivariable systems for selection of square subsystem from non-square system. a dynamic relative gain was proposed by avoy, et al. [40]. 3.2.4 operability index the operability measure can quantify the inherent ability of the process to move from one steady state to another and to reject any of the expected disturbances. the operability index is defined by vinson and georgakis [41] to effectively capture the inherent operability of continuous processes. vison and georgakis [41] applied this index to analyze the steady state of the linear system. the technique has also been proven to be effective for nonlinear processes [42]. it was also extended to dynamic systems by uztĺźrk and georgakis [43]. a brief survey paper about the operability index was presented by georgakis, et al. [44]. 4 challenge any method has its strength and weakness. the fundamental difficulties in robust design are related to model/parameter uncertainties and nonlinearity of the system. most of the existing methods can deal with linear system with the known model, and are not capable to handle model uncertainty or nonlinearity 238 han-xiong li : robust design and its challenge for manufacturing system that in turn generates extra uncertainties to the system. the weakness of the existing design methods poses the challenges to unsolved problems as shown in table 3. 4.1 robust design for static system 1) there is still no method that can consider the model uncertainties in the deterministic robust design. the accurate model is needed for euclidean norm method and conditional number method to obtain the gradient information. the sensitivity region measure methods are based on the projection between parameter space and performance space, which requires the system model. 2) the model uncertainties still can not be handled in the probabilistic robust design, where all the existing methods can be classified two categories : modelbased robust design and data-based robust design. the data-based robust design includes monte carlo method and taguchi method. • since monte carlo method is a simulation method, all system knowledge, which has to be known beforehand, need to be translated into computer code. thus, the accurate system model is critical. moreover, it costs huge computational time that will limit its application. • taguchi method obtains the system robustness based on the experiment data. however, it is only suitable for the stochastic environment and not suitable to minimize the effect of parameter variations. the model-based robust design includes firstand secondorder moment methods. since these methods are based on the taylor series expansion, which requires the accurate system model. thus, the model uncertainty will lead to the significant performance degradation. so far, there is still no solution. 3) the membership function in fuzzy analysis is defined according to human experience, which could be too subjective and causes uncertainties. it would be a challenge if the experimental data can be used to reduce subjective uncertainties of the membership. 4.2 robust design for dynamic system in difference to the static system, another problem for the dynamic system is that the design optimization becomes extremely difficult when the system has both continuous and discrete design variables. 1) the current stability design can work for the linear system, where its eigenvalues are critical to the system stability. the weakness of the existing methods have not considered two influences : one is model uncertainties on the eigenvalues and their variations, and the other is parameter perturbation on the eigenvaluesąŕ variations. 2) the stochastic flexibility index is only applicable to the linear dynamic system. so far, there is still no study on its application to the nonlinear dynamic system. advances in systems science and applications (2014) vol.14 no.3 239 4.3 potential solutions for the static system, a novel model-based robust design is proposed to design the system using the nominal model. the system-model mismatch, model uncertainties, can be properly considered in the design. for the dynamic system, a novel stability based robust design is proposed to guarantee the robust performance as well as the system stability, so that the method can also be applied to the weak nonlinear system. 5 conclusion this paper presents a brief overview about advances in robust design. different approaches in robust design are reviewed and compared, upon which challenges have been proposed to the unsolved problems. two novel approaches are proposed to the unsolved problems. table 3 summary of evaluation of existing methods design classification existing method challenges for unsolved problems deterministic system deterministic robust design these methods need an accurate model and can not handle the model uncertainty. probabilistic data-based methods all these methods can not handle model uncertainties properly, furthermore, static • monte carlo • monte carlo simulation costs huge computational time, and requires the accurate model ; robust system • taguchi method • taguchi method is not effective to parameter variations ; model-based methods • model-based methods require the accurate model ; design fuzzy system • fuzzy analysis membership function strongly depends on human experience that causes subjective uncertainties. dynamic robust design • stability based design all existing methods can not work for the system with model uncertainties, and • feasibility design hybrid system withboth continuous and discrete design variables, furthermore, • other design • stability design has not been applied to the nonlinear system. • feasibility design has not been applied to nonlinear dynamic system. 240 han-xiong li: robust design and its challenge for manufacturing system références [1] caro s., bennis f., and wenger p. (2005). “tolerance synthesis of mechanisms : a robust design approach”, journal of mechanical design, vol.127, no.1, pp.86-94. [2] taguchi, g., (1993). “aguchi on robust technology development : bringing quality engineering upstream”, asme press, new york. [3] zhu j.m., and ting k.l. (2001), “performance distribution analysis and robust design”, journal of mechanical design, vol.123, no.1, pp.11-17. [4] ting k.l., and long y.f., (1996), “performance quality and tolerance sensitivity of mechanisms”, journal of mechanical design, vol.118, no.1, pp. 144-150. [5] hong, w.m. (2003), “large deviations for the super-brownian motion with super-brownian immigration”, j. theoret. probab. vol.126, pp.395-402. [6] li, m., azarm, s., and boyars, a. (2006), “a new deterministic approach using sensitivity region measures for multi-objective robust and feasibility robust design optimization”, journal of mechanical design, vol.128, pp.874883. 1218-1224. [7] huang, h.z., and zhang x.d. (2009), “design optimization with discrete and contrinuous variables of alteatory and epistemic uncertainties”, journal of mechanical design, vol.131, pp.1-7. [8] wu, w.d., and rao, s.s., (2004), “interval approach for the modeling of tolerances and clearances in mechanism analysis”, journal of mechanical design, vol.126, pp.581-592. [9] nikolaidis, e., chen, s., cudney, h., haftka, r.t., and rosca, r. (2004), “comparison of probability and possibility for design against catastrophic failure under uncertainty”, journal of mechanical design, vol.126, pp.386395. [10] he, l.p., and qu, f.z., (2008), “possibility and evidence theory-based design optimization : an overview”, kybernetes, vol.37, pp.1322-1330. [11] beyer h.g., and sendhoff b.,(2007), “robust optimization ĺc a comprehensive survey”, computer methods in applied mechanics and engineering, vol.196, pp.3190-3218. [12] liu h.b., chen w., & sudjianto, a., (2004), “probabilistic sensitivity analysis methods for design under uncertainty”, aiaa-2004-4589, 10th aiaa/issmo multidisciplinary analysis and optimization conference, albany, new york. [13] soboląŕ, i.m., (1993), “sensitivity analysis for nonlinear mathematical models”, mathematical model & computational experiment, vol.1, pp.407-414. advances in systems science and applications (2014) vol.14 no.3 241 [14] melchers, r.e., (1999), structural reliability analysis and prediction, john wiley & sons, chichester, new york. [15] krzykacz-hausmann, b., (2001), “epistemic sensitivity analysis based on the concept of entropy”, in proceedings of samo 2001, ciemat, pp.31-35. [16] suh, n.p., (2005), complexity : theory and applications, new york : oxford university press. [17] chen, w., and yuan, c., (1999), “a probabilistic-based design model for achieving flexibility in design”, journal of mechanical design, vol.31, pp.615639. [18] chen w., allen j.k., tsui k.l, and mistree f., (1996), “a procedure for robust design : minimizing variations caused by noise factors and control factors”, journal of mechanical design, vol.118, no.4, pp.478-493. [19] tsui, k.-l., (1992), “an overview of taguchi methods and newly developed statistical methods for robust design”, iie transactions, vol.24, no.5, pp.44-57. [20] shoemaker, a. c., tsui, k. l., and wu, j., (1991), “economical experimentation methods for robust design”, technometrics, vol.33, no.4, pp.415-427. [21] suh, n.p., (2005), complexity : theory and applications, new york : oxford university press. [22] dimitriadis, v.d., and pistikopoulos, e.n. (1995), “flexibility analysis of dynamic system”, ind. eng. chem. res, vol.27, no.8, pp.1291-1301. [23] mohideen, m.j., perkins, j.d., pistikopoulos, e.n., (1997), “robust stability considerations in optimal design of dynamic systems under uncertainty”, journal of process control, vol.7, no.5, pp.371-385. [24] kokossis, a., floudas, c., (1994), “stability in optimal design : synthesis of complex reactor networks”, aiche journal, vol.40, no.5, pp.849-861. [25] monnigmann, m., and marquardt, w., (2003), “steady-state process optimization with guaranteed robust stability and feasibility”, aiche journal, vol.49, no.12, pp.3110-3126. [26] grosch, r., monnigmann, m., and marquardt, w., (2008), “integrated design and control for robust performance : application to an msmpr crystallizer”, journal of process control, vol.18, no.2, pp.173-188. [27] grossmann, i.e. and straub, d.a., (1991), “recent developments in the evaluation and optimization of flexible chemical processes”, in proceedings computer oriented process engineering, barcelona. [28] swaney, r.e., and grossmann, i.e., (1985), “an index for operational flexibility in chemical process design. part i : formulation and theory”, aiche journal, vol.31, pp.621-630. 242 han-xiong li : robust design and its challenge for manufacturing system [29] straub, d.a., and grossmann, l.e., (1993), “design optimization of stochastic flexibility”, comp. chem. eng, vol.17, pp.339-354. [30] skogestad, s., and morari, m., (1987), “the effect of disturbance directions on closed loop performance”, ind. eng. chem. res, vol.26, pp.2029-2035. [31] lewin, d.r., (1996), “a simple tool for disturbance resiliency diagnosis and feedforward control design”, computers & chemical engineering, vol.20, pp.13-25. [32] cao, y., rossiter, d., and owens, d., (1997), “input selection for disturbance rejection under manipulated variable constraints”, computers& chemical engineering, vol.21, pp.s403-s408. [33] lewin, d.r., (1996), “a simple tool for disturbance resiliency diagnosis and feedforward control design”, computers & chemical engineering, vol.20, pp.13-25. [34] solovyev, b.m., and lewin, d.r., (2003), “a steady-state process resiliency index for nonlinear processes : 1. analysis”, ind. eng. chem. res, vol.42, pp.4506-4511. [35] skogestad, s., and havre, k., (1996), “the use of rga and condition number as robustness measures”, computers chem. engng, vol.20, pp.s1005-s1010. [36] wal, m.v.d., and jager, b.d., (2001), a review of methods for input/output selection, automatica, vol.37, pp.487-510. [37] bristol, e.h., (1996), “on a new measure of interaction of multivariable process control”, ieee trans. auto. control, vol.11, pp.133-134. [38] manousiouthakis, v., savage, r., and arkun, y., (1986), synthesis of decentralized process control structures using the block relative gain, aiche, vol.32, pp.991-1003. [39] chen, j.w., and yu, c.c., (1990), “the relative gain for non-square multivariable systems”, chemical engineering science, vol.45, pp.1309-1990. [40] avoy, t.m., arkun, y., chen, r., robinson, d., schnelle, p.d., (2003), a new approch to defining a dynamic relative gain, control engineering practical, vol.11, pp.907-914. [41] vinson, d.r., and georgakis, c., (1998), “a new measure of process output controllability”, proceedings of the fifth ifac symposium on dynamics and control of process systems, in georgakis, pp.700-709. [42] subramanian s., uztĺźrk d., and georgakis, c., (2001), “an optimizationbased approach for the operability analysis of continuously stirred tank reactors”, industrial and engineering chemistry research, vol.40, pp.4238-4252. advances in systems science and applications (2014) vol.14 no.3 243 [43] uztĺźrk, d., and georgakis, c., (2002), “inherent dynamic operability of processes : general definitions and analysis of siso cases”, industrial and engineering chemistry research, vol.41, pp.421-432. [44] georgakis, c., uztĺźrk, d., subramanian, s., and vinson, d.r., (2003), “on the operability of continuous processes”, control engineering practice, vol.11, pp.859-869. microsoft word 8-门可佩.doc 48-54 advances in systems science and applications (2010), vol.10, no.1 issn 1078-6236 international institute for general systems studies, inc. research on comprehensive evaluation of harmonious society in east china* kepei men, liangyu jiang and jing liu nanjing university of information science and technology, nanjing 210044, china email: menkp@yahoo.com.cn abstract the harmonious society is the general goal of social development. we select 5 subsystems and 35 evaluation index to establish an system for evaluating harmonious society, with the method of principle component-cluster analysis, according to 2007 statistical yearbook of china and some other statistical data. meanwhile, we use topsis method as an assistant certification to analyze the harmonious society development of east china. the result shows that the six provinces in east china and shanghai city could be divided into 3 parts. the first part includes shanghai. the second part includes zhejiang, jiangsu and shandong. and the third part includes fujian, jiangxi and anhui. the result is identical to practice developing of east china. keywords east china harmonious society comprehensive evaluation principle component-cluster analysis (pcca) 1. introduction realizing social harmony and building a nice society is the social ideal that human being assiduously seeks. it is also the social ideal of all the marxist party including the communist party of china. to construct a socialist harmonious society is a great task that put forward in whole under the new situation of initiating the socialist road with the specific practice in china. mr hu jintao proposed to realize harmony zone on the 5th anniversary forum of the foundation of shanghai cooperation organization. east china is the most active economic growth zone. the six provinces and shanghai city of it account for less than 1/4 of the whole country, the population is 28.74% of the whole nation and the land area only accounts for 8.13%. but it has made great contribution for china, because of its wonderful economy. according to the statistical bulletins of nation and east china in 2007, we knew that the resident income of the six provinces and shanghai city was kept growing, and the living standard was improved further. the per capital dominating income of urban residents had exceeded 11 thousand yuan. among them, shanghai, jiangsu, zhejiang, fujian and shandong were over 14 thousand yuan, which were higher than the national average 13.786 thousand yuan. the gdp of this area in 2006 and 2007 were reached 8826.513 billion yuan and 10406.21 billion yuan respectively, which accounted for 41.86% and 42.20% of china’s gdp. therefore, study on the problem about the construction of the harmonious society in east china is strongly representative and has a profound significance. in this paper, we focus on the construction idea of socialist harmonious society, to put forward a harmonious society index system with great operative, for the purpose of providing reference for the construction of harmonious society. 2. the introduction to research methods 2.1 the method of principal component and clustering analysis[1-4] suppose there are n samples, every sample has p indices, we get a original data matrix ∗ this work is supported by the key projects of national statistics research program (2008lz022). advances in systems science and applications (2010), vol.10, no.1 49 ( )ij n px x ×= , ( 1, 2, , ; 1, 2, )i n j p= = and we make linear combination (comprehensive index) with p vectors 1 2, , , px x x , that is 1 1 2 2 , 1, 2, ,i i i pi pf a x a x a x i p= + + = then we limit the combination coefficients 1 2( , , , )i i i pia a a a= with 2 2 2 1 2 1, 1, 2, ,i i pia a a i p+ + + = = in which, ia is a unit vector and 1i ia a = . the comprehensive index is determined by the following rules: (1) if and jf ( , , 1, 2, ,i j i j p≠ = ) are uncorrelated, that is ( , ) 0i jcov f f = . (2) 1f is the biggest variance in all the linear combinations of 1 2, , , px x x , namely, ' 1 1 ( ) max p i i c c i var f var c x = = ⎛ ⎞ = ⎜ ⎟ ⎝ ⎠ ∑ , in which ' 1 2,( , , )pc c c c= . as to the rule (2), we will get the similar conclusions. 2f , uncorrelated with 1f , is the biggest variance in all the linear combinations of 1 2, , , px x x . pf , uncorrelated with 1 2 1, , pf f f − , is the biggest variance in all the linear combinations of 1 2, , , px x x , and so on. the comprehensive vectors 1 2, , , pf f f , satisfied above requirements, are the principal components. the information extracted from the information content of original indices decrease successively. we use variance to measure information extracted by every principal component, and the contribution of principal component variance is equal to corresponding eigenvalue iλ to correlation matrix of the original data. 1 2( , , , )i i i pia a a a= , the combination coefficient of every principal component, is the eigenvector it to corresponding eigenvalue iλ . the contribution rate of variance is 1 / p i i j j α λ λ = = ∑ . the more iα , the more information corresponding principal component explains. when the variance contribution rate of one principal component is very small, we think the information it provides is little, then we may delete it. in general case, if the cumulative variance contribution rate of first q principal components reaches 85%, we just consider the first q principal components, then explain properties of random vector x with them. other principal components are the random errors caused by incorrect observation. research on comprehensive evaluation has made process[5-7]. papers [1-3] point the method popular in recent year is to structure a comprehensive evaluation function of principal components for ranking, which is based on the contribution rate of variance iα , but it is wrong. when the contribution rate of variance of the first principal component 1f is quite high (over 85%), we may think this principal component can almost reflect the information provided by the original variables. then we may rank and evaluate according to the scores of the first principal component. to the ranking problems of multiple index system, when the contribution rate of variance of the first principal component is not over 85%, that is, the original data information expressed by the first principal component is not enough, it has one-sidedness, if we still rank and evaluate only by the scores of the first principal component. at this time, we combine principal men: research on comprehensive evaluation of harmonious society in east china 50 component analysis with cluster analysis, which is named “principal component-cluster analysis (pcca)”. table 1 the index system subsystem index unit directivity m at er ia l c iv ili za tio n 1 2 3 4 5 6 7 8 9 10 yuan/person yuan yuan % % square meter/person % part/10 thousand agriculture=1 % positive positive positive negative negative positive positive positive medium positive po lit ic al ci vi liz at io n 11 12 13 14 person/10 thousand female=100 % part/10 thousand positive medium negative negative sp iri tu al ci vi liz at io n 15 16 17 18 19 % person/100 thousand % % kind/10 thousand positive positive positive positive positive so ci al ci vi liz at io n 20 21 22 23 24 25 26 27 person/10 thousand person/100 thousand old % % part/10 thousand % % positive negative positive positive positive positive negative positive h ar m on io us so ci et y ec ol og ic al ci vi liz at io n 28 29 20 31 32 33 34 35 ton/10 thousand yuan % % % % % 10 thousand stere % negative positive positive positive positive positive positive positive note: 1 means per capital gdp; 2 means per capital disposable income of urban residents; 3 means per capital net income of rural residents; 4 means the engle's coefficient of urban residents; 5 means the engle's coefficient of rural residents; 6 means housing areas of rural residents; 7 means the proportion that the tertiary industry accounts for gdp; 8 means numbers of patents application; 9 means consumption level of urban and rural residents; 10 means the proportion that r&d accounts for gdp; 11 means lawyer’s numbers that per ten thousand people have; 12 means sex ratio of senior middle school graduates; 13 means registered unemployment rate of urban residents; 14 means numbers of criminal case; 15 means the proportion that education and culture entertainment account for total consumption expenditure; 16 means average students number of colleges and universities; 17 means excellent rate of product quality; 18 means the proportion that education operating expenses account for fiscal expenditure; 19 means book publishing kinds; 20 means doctor’s numbers that per ten thousand people have; 21 means death toll from traffic accidents; 22 means average life expectancy; 23 means mobile subscription; 24 means gross enrollment ratio of higher education; 25 means numbers of urban community service facilities; 26 means gross divorce rate; 27 means the proportion that numbers of participate in advances in systems science and applications (2010), vol.10, no.1 51 medical insurance account for total amount; 28 means energy consumption in every unit; 29 means forest coverage; 30 means industrial wastewater treatment level; 31 means per capita park greenery area; 32 means water subscription; 33 means gas subscription; 34 means daily treatment ability of domestic sewage; 35 means harmless treatment rate of domestic waste. [11] “principal component-cluster analysis” is such a method. first we do the principal component to get some principal components, with which we do cluster analysis to samples, then we rank and classify the samples by the scores of the first principal component. specific ways are the following: (1) select first r principal components by cumulative contribution rate and calculate the scores of the principal components 1 1 2 2 , ( 1,2, , )l l l pl pf a x a x a x l r= + + = (2) do the systematic cluster analysis to the selected new matrix ( 1 2, , , rf f f ); (3) calculate the scores’ average value of the first principal component to determine ranking of every class; (4) determine ranking of every sample in the class to get comprehensive evaluation result, according to every sample’s score. 2.2 the method of topsis main steps of topsis: (1) select p evaluation index to n evaluation units for comprehensive evaluation, then get a evaluation matrix ( )ij n px x ×= , in which ijx is observation data in unit i . (2) make original data being dimensionless, then get a normalized evaluation matrix ( )ij n pz z ×= , in which 2 1 / ( 1,2, , ; 1,2, , ) n ij ij kj k z x x i n j p = = = =∑ . (3) determine the positive ideal solution z + and negative ideal solution z − of matrix z , then get 1 2( , , , )qz z z z+ + + += , 1 2( , , , )qz z z z− − − −= , in which 1 2max{ , , , } ,j j j njz z z z+ = 1 2min{ , , , }, ( 1, 2, , )j j j njz z z z j q− = = . (4) calculate the distances between every unit and positive ideal solution, and the distances between every unit and negative ideal solution: 2 1 ( ) q i ij j j d z z+ + = = −∑ , 2 1 ( ) q i ij j j d z z− − = = −∑ ( 1, 2, , )i n= (5) calculate relative approach degrees between every evaluation unit and the optimal solution: 100%, ( 1, 2, , )i i i i dc i n d d − + −= × = + (6) rank according to the relative approach degree. the more it closes to 100, the more ic . it indicates that unit i is more close to the optimal level. otherwise, unit i is less close to the optimal level. 3. establishment of index system and empirical analysis to build a harmonious society is a great social system engineering. from the horizontal point of view, it is closely associated with society, politics, economy and culture. from the vertical point of view, it involves in macro-view, medium-view and micro-view. the concept of harmonious society is extensive, comprehensive and systematic. it includes every aspect of men: research on comprehensive evaluation of harmonious society in east china 52 economy, politics, society, culture, environment and people’s life. therefore, the index system of harmonious society should embody the comprehensive and systematic connections among index, and reflect progress condition of construction in many aspects. in this paper, we will divide it into 5 subsystems to make it have good operability, that is, material civilization system, political civilization system, spiritual civilization system, social civilization system and ecological civilization system, (details in table 1) [8-11] . 3.1 calculation with the method of principal component and cluster analysis first we do pretreatment to the original data. then we make principal component analysis to the harmonious index of six provinces and shanghai city with sas program. finally, we get the following eigenvalues of the correlation matrix. eigenvalues of the correlation matrix eigenvalue difference proportion cumulative 1 18.6926827 12.4745482 0.5341 0.5341 2 6.2181345 2.0407184 0.1777 0.7117 3 4.1774161 1.4626788 0.1194 0.8311 4 2.7147372 0.8307516 0.0776 0.9087 5 1.8839856 0.5709418 0.0538 0.9625 6 1.3130439 1.3130439 0.0375 1.0000 ………… in which, prin1, prin2, prin3, prin4 denote the first four principal components. the second list is eigenvalues of sample correlation matrix, the fourth list is contribution proportions of variance, and the fifth list is cumulative contribution proportions. as cumulative contribution proportion of the first four eigenvalues reaches 90.87%, which is higher than 85%, we only need to select the first four principal components to summarize the all data, then get the scores of the principle components (see table 2). table 2 the scores of the principle components analysis cities scores the first principal component the second principal component the third principal component the fourth principal component shanghai 8.52652 -2.49745 -0.61749 0.58853 jiangsu 0.68228 1.08935 1.92012 -1.20837 zhejiang 1.99413 4.24675 -0.908082 -0.56337 anhui -3.85481 -2.3237 -2.88072 -1.98804 fujian -1.67628 1.85291 -0.83024 1.03237 jiangxi -3.62903 -0.86869 -0.02778 2.92133 shandong -2.04282 -1.49017 3.34413 -0.78246 the contribution proportion of first principal component’s variance is 53.41%, which is the largest of all linear combinations, and the information of the first principal component is the biggest. we get samples’ ranking the first time by calculating the scores of the first principal component. if we rank only according to the scores of the first principal component, as contribution proportion is not over 85%, the information will not be big enough and the result will have one-sidedness. thus, we do the cluster analysis to the first four principal components’ score matrix with sas program, then we get the cluster graph (see figure 1). from figure 1, we know that the six provinces and shanghai city can be divided into three classes, that is, {shanghai}; {jiangsu, shandong, zhejiang}; {fujian, jiangxi, anhui}. then we rank according to the scores of first principal component. that is, {shanghai}; {jiangsu, shandong, zhejiang}; {fujian, jiangxi, anhui}. finally, we get the ranking by the scores of the first principal component in every class. that is, {shanghai, zhejiang, jiangsu, shandong, fujian, jiangxi, anhui}. from table 3, we get the rankings by principal component-cluster analysis and the scores of the first principal component. we find that they are almost the same, the only different is ranking between shandong and fujian. however, 20 index in shandong are better than fujian, shandong should be in front of fujian. therefore, it is more reliable to get the ranking by principal advances in systems science and applications (2010), vol.10, no.1 53 component-cluster analysis. figure 1 cluster graph table 3 the compositors of the principal component cluster and the first principal component shanghai jiangsu zhejiang anhui fujian jiangxi shandong ranking by principal component-cluster analysis 1 3 2 7 5 6 4 ranking by the scores of the first principal component 1 3 2 7 4 6 5 3.2 calculation with the method of topsis we calculate relative approach degree ic of every subsystem by matlab program. the more it closes to 100, the more ic . it indicates that unit i is more close to the optimal level. otherwise, unit i is less close to the optimal level. details in table 4. the construction of material civilization, political civilization, spiritual civilization, social civilization and ecological civilization are closely linked with each other in the process of building a socialist harmonious society and a well-off society. they are very important, so we should treat them equally. we get the final scores of the six provinces and shanghai city after doing the equal weight. then we rank by the final scores, which is listed in table 4. the result is the same as the ranking by principal component-cluster analysis, which verifies rationality of the method of principal component-cluster analysis. table 4 the development level of six provinces in east china and shanghai city material civilization political civilization spiritual civilization social civilization ecological civilization score ranking shanghai 86.6131 46.87 80.45 71.89 32.28 63.62062 1 jiangsu 38.4696 32.37 19.50 32.84 57.44 36.12392 3 zhejiang 53.6959 23.66 20.32 47.55 52.83 39.61118 2 anhui 12.3374 31.65 9.48 17.73 24.29 19.09748 7 fujian 21.9262 28.24 23.78 22.12 49.08 29.02924 5 jiangxi 8.7141 35.27 27.04 27.24 41.41 27.93482 6 shandong 24.8297 60.62 14.78 26.26 44.15 34.12794 4 4. conclusions and discussions 2006 was the beginning of eleventh five-year plan. under the leadership of provincial and municipal governments in east china, we had acquired remarkable achievement in the field of social economic development, by insisting in carrying out macro-control policy and guiding with scientific development view. according to the methods of principal component-cluster analysis men: research on comprehensive evaluation of harmonious society in east china 54 and topsis, we divide six provinces and shanghai city into 3 classes on the level of social harmonious development. (1) shanghai city is the first class. the score of shanghai is far higher than any other province. it shows that shanghai city is superior to other places in the construction of harmonious society. since reform and opening up to the outside world, shanghai city has got great achievement as the first metropolitan. the index of per capital gdp, per capital disposable income of urban residents and per capital net income of rural residents are in the first place. (2) zhejiang, jiangsu and shandong are in the second class, which are high in harmony. the scores of zhejiang province and jiangsu province are high in material civilization system, ecological civilization system and social civilization system, but low in spiritual civilization system. these places have strong economic strength, good security systems and harmonious living condition. the scores of shandong province are low in material civilization system, spiritual civilization system and social civilization system, but high in political civilization system. shandong province is one of the best provinces in public security. (3) fujian, jiangxi and anhui are in the third class. it has significant gap to others. the relatively backward social economy restricts social development, leading to the development of society unbalanced. from above conclusions, we know that shanghai, zhejiang, jiangsu and shandong are in the leading positions of the whole nation. in the future, we should not only focus on the quantity of economic operation, but also focus on the quality. to other relative undeveloped regions in east china, we should speed up the development to shorten the difference. six provinces and shanghai city in east china will obtain new development in the new century through cooperation and exchange and will create a more bright future. references [1] xueming wang. applied multivariate analysis (2rd edition). shanghai: shanghai university of finance and economics press, 2004: 1-358. (in chinese) [2] xueming wang. query to the applied method of principal component. statistics and decision, 2007, 8: 31-32. (in chinese) [3] jingya xu, wang yuanzheng. improvement to the applied method of principal component. mathematics in practice and theory, 2006, 6: 68-75. (in chinese) [4] daoyuan zhu, chengou wu, weiliang qin. applied multivariate analysis and sas software. nanjing: southeast university press, 1998: 1-408. (in chinese) [5] chuncheng wu, shengbao yao, chaoyuan yue. the model of the evaluation of the comprehensive progress for large project. advances in system science and applications, 2006, 6(3): 410-415. [6] pu gong, qiang si, jianling meng. the evaluation method of multi-stage compound real option on human capital valude. advances in systems science and applications, 2006, 6(1): 101-106. [7] wei wang, zhuangzhi liu. comprehensive evaluation of project risk based on fuzzy network analytical methods. advances in systems science and applications, 2008, 8(2): 270-275. [8] state statistic bureau studying team. research on the index system of harmonious society. statistical research, 2006, 5: 23-29. (in chinese) [9] mei song, qi xin. the construction of harmonious society index system. beijing social science, 2006, 1: 62-66. (in chinese) [10] jianguo ouyang. research on comprehensive evaluation index system of harmonious society. zhejiang social science, 2006, 2: 16-22. (in chinese) [11] prc national statistics bureau. china statistical yearbook-2007. beijing: china statistics press, 2007: 1-1028. (in chinese) adv syst sci appl 2017; 4:61–77 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/518 price of anarchy for maximizing the minimum machine load yulia. v. chirkova institute of applied mathematical research, karelian research centre ras, petrozavodsk, russia abstract: the maximizing the minimum machine delay game (or cover game) with uniformly related machines is considered. players choose machines with different speeds to run their jobs trying to minimize job’s delay, i.e. the chosen machine’s completion time. the social payoff is the minimal delay over all machines. for the general case of n machines we found the lower bound for price of anarchy (poa), and for the case of 3 machines we found its exact value. we proved that the poa does not change or increases when an additional third machine is included into the system with two machines. also we propose a method of computation the poa value and illustrate it for 3 machines. keywords: nash equilibrium, maximizing the minimum load, cover game, price of anarchy 1. introduction load balancing represents a major problem in networks and distributed computing systems, since load optimization guarantees efficient resource utilization. modern systems such as telecommunication networks, cloud computing systems, grid, etc. consist of independent components, in many cases without their centralized control. particularly, users located at nodes and data transmission protocols do not interact with each other for maintaining a certain load level. furthermore, in practice they demonstrate egoistic behavior with respect to free resources. and global optimization methods often become inapplicable due to the infeasibility of realizing optimal resource utilization plans in such systems (server request schedules, capacity norms of data transmission channels and so on). the game-theoretic approach allows treating load balancing as a game, where players have egoistic behavior and can reach some equilibrium state such that none of them benefits from unilateral deviation from a chosen strategy. system efficiency is assessed by comparing the above equilibria with the global optimum. the present paper focuses on the maximizing the minimum machine delay game (or cover game) [1–3] also known as the scheduling problem [4] in the form of a game equivalent to the kp-model (see [5, 6]) with parallel different-capacity channels where system optimization is the maximizing the minimum machine delay [1–3] instead of the minimization the maximum machine delay (makespan). it is necessary to distribute several jobs of various volumes among machines of nonidentical speeds. the volume of a job is its completion time on a free unitspeed machine. machine load is the total volume of jobs executed by a given machine. the ratio of machine load and speed defines its delay, i.e., the job completion time at this machine. each player chooses a machine for its job striving to minimize job’s delay. players have egoistic behavior and reach a nash equilibrium, viz., a job distribution such that none of them benefits from unilateral change of a chosen machine. in the sequel, we study pure strategies nash equilibria only; as is well-known [7, 8], such an equilibrium always exists ∗corresponding author: julia@krc.karelia.ru http://ijassa.ipu.ru/ojs/ijassa/article/view/518 62 y.v. chirkova in the described class of games. the system payoff (also called the social payoff) is the minimum delay over all machines for an obtained job distribution. the price of anarchy [1] (poa) is defined as the maximum ratio of the optimal social payoff and the social payoff in the worst-case nash equilibrium. the problem where the system tries to maximize the minimum delay over all machines appears from the concept of a fair resource sharing and efficient routing of traffic. the paper [1] first in the equilibrium efficiency studying for such model gives motivations coming from issues of quality of service, fair resource allocation, and fair queuing. the base idea is that each system component must be loaded as much as possible and not to idle. consider an example where each player pays the system a value which equals his delay for his job processing. fair system should not have privileged players who pay rather less than others due to successful machine choose. also such system should not have machines providing small or zero payoff. according to the earlier publications, the poa in the maximizing the minimum machine delay games with pure strategies can be estimated by • for n ≥ 2 machines with speeds 1 ≤ · · · ≤ s [1] the price of anarchy is not limited if s ≥ 2; • the price of anarchy is closed to and no more than 1.7 for any number of homogeneous machines [1, 9]; • the price of anarchy equals{ 2+s (1+s)(2−s) for 1 ≤ s ≤ √ 2, 2 s(2−s) for √ 2 < s < 2 for two machines with speeds 1 ≤ s [2]; • the price of anarchy equals 2 + s 2(2− s) for 1 ≤ s < 2 for three machines with speeds 1 = 1 ≤ s [2]; • the price of anarchy equals 1+s s for 1 ≤ s ≤ s0, 2+s (1+s)(2−s) for s0 < s ≤ √ 2, 2 s(2−s) for √ 2 < s < 2 in the hierarchical model of two machines with speeds 1 ≤ s and two types of jobs where first machine can process both types jobs and second machine can process only second type jobs [3]. here s0 is the largest root of the equation 1+s s = 2+s (1+s)(2−s) . in what follows, we derive a lower estimate for the poa in the case of n ≥ 3 machines. also we present the exact value of the poa for 3 machines with speeds 1 ≤ r ≤ s < 2:{ 2+s (1+r)(2−s) for rs ≤ 2, 2 r(2−s) for rs > 2. moreover we show that the poa increases or does not change under new machine inclusion into the system of two machines. in the case of n machines, a computing algorithm of the exact poa value is proposed based on solving a series of linear programming problems. the algorithm is described for the case of 3 machines and is implemented numerically in the form of a program which draws the curves of the poa as a function of the fastest machine and compares them with the curves of the corresponding estimates. copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 63 2. the model consider a system s = s(n, v) composed of n machines operating with speeds v1 = 1 ≤ · · · ≤ vn = s. note that such choose of machine speeds does not contradict with generality since one always can normalize speeds dividing them by the speed of the slowest machine. the system is used by a set of players u = u(n,w): each of n players chooses an appropriate machine for its job execution. for player j, the volume of job equals wj , j = 1, . . . , n. denote by w = n∑ j=1 wj the total volume of all jobs. free machine i with speed vi executes a job of volume w during the time w/vi. study the following pure strategies game γ =< s(n, v), u(n,w), λ >. each player can choose any machine. the strategy of player j is machine lj selected by this player for its job execution. then the strategy profile in the game γ represents the vector l = (l1, . . . , ln). the load of machine i, i.e., the total volume of all jobs assigned to the machine is defined by δi(l) = ∑ j=1,...,n:lj=i wj . the delay of machine i takes the form λi(l) = ∑ j=1,...,n:lj=i wj/vi = δi(l) vi . actually, this quantity is the same for all players selecting a given machine. we suppose that the goal of the system is minimizing of the least busy machine idling, that is maximizing of its working time or delay on it. the social payoff is described by the minimum delay over all machines: sc(l) = min i=1,...,n λi(l). designate by opt = opt (s, u) = max l is a profile in γ(s,u,λ) sc(l) the optimal payoff (the social payoff in the optimal case) where maximization runs over all admissible strategy profiles in the game γ(s, u, λ). a strategy profile l such that none player benefits from unilateral deviation (change of the machine chosen in l for its job execution) is a pure strategies nash equilibrium. to provide a formal definition, let l(j → i) = (l1, . . . , lj−1, i, lj+1, . . . , ln) signify the profile obtained from a profile l if player j replaces machine lj chosen by it in the profile l for another machine i, whereas the rest players use the same strategies as before (remain invariable). definition 2.1: a strategy profile l is said to be a pure strategies nash equilibrium iff each player chooses a machine with the minimum delay, i.e., for each player j = 1, . . . , n we have the inequality λlj(l) ≤ λi(l(j → i)) for all machines i = 1, . . . , n . definition 2.2: the price of anarchy in the system s is the maximum ratio of the social payoff in the optimal case and the social payoff in the worst-case nash equilibrium: poa(s) = max u opt (s, u) min l is a nash equilibrium in γ(s,u,λ) sc(l) . copyright c© 2017 assa. adv syst sci appl (2017) 64 y.v. chirkova 3. the general case of n machines in this section we give the following assumptions and results which will be employed in further analysis. consider a system composed of n ≥ 2 machines operating with speeds v1 = 1 ≤ · · · ≤ vn = s. if the number of jobs n is less than the number of machines n then obviously the social payoff is zero in any profile. in this case we assume by definition that the ratio of an optimal payoff to an equilibrium payoff is 1. further we suppose that n ≥ n . if s ≥ 2 then the price of anarchy is infinite [1]. therefore, we assume 1 ¡ s ¡ 2 in the following. if the number of jobs n is more than or equals the number of machines n then obviously all machines are loaded in an optimal profile. moreover in this case all machines are loaded in any equilibrium. the optimal social payoff is not larger than the social payoff in the case when the whole volume of jobs is distributed among machines proportionally to their speeds so that all machines have an identical delay: opt ≤ w n∑ i=1 vi . (3.1) further we determine estimates for equilibrium delays and volumes for some jobs processed on machines. we also restate the proof for results taken from cited papers for the sake of completeness. lemma 3.1: [2] if the number of jobs n is not less than the number of machinesn then in any equilibrium the load of any machine is more than zero. proof consider an arbitrary equilibrium profile l. suppose that some machine i has zero load. then there is a machine k receiving not more than two jobs. since v1 = 1 ≤ · · · ≤ vn = s < 2 then vi > vk 2 . let wk be the minimal job volume on k. if it moves to an idle machine i then its load becomes equal wk vi < 2wk vk ≤ λk(l), that is less comparing with its load in the profile l. denote the number of jobs on some machine k in a profile l by nk. lemma 3.2: [2] suppose that l is a nash equilibrium profile and sc(l) = λi(l). if nk > vk vi then λk(l) ≤ nkvi nkvi−vk λi(l) for any machine k. proof let w be the job with the smallest volume on some machine k. then w ≤ vk nk λk(l). since l is an equilibrium then λk(l) ≤ λi(l) + w vi ≤ λi(l) + vk nkvi and thus λk(l) ≤ nkvi nkvi−vk λi(l). lemma 3.3: suppose that l is an equilibrium profile and sc(l) = λi(l) and consider an arbitrary machine k. if nk ≥ 2 and 1 ≤ vk vi < 2 then the volume wj of any job j on the machine k is at most vivk 2vi−vk λi(l). moreover the total volume of remaining jobs on k is also no more than vivk 2vi−vk λi(l). proof let the machine k receive two or more jobs and w be the minimal job volume on k. then copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 65 the total volume of remaining jobs on k equals vkλk(l)− w. since l is an equilibrium then λk(l) = vkλk(l) vk ≤ λi(l) + w vi and thus vkλk(l)− w ≤ vkλi(l) + ( vk vi − 1 ) w ≤ vkλi(l) + ( vk vi − 1 ) wj:lj=k ≤ vkλi(l) + ( vk vi − 1 ) (vkλk(l)− w). then w ≤ wj:lj=k ≤ vkλk(l)− w ≤ vivk 2vi−vk λi(l). the next theorem determines the lower estimate for the price of anarchy in the system of n ≥ 3 machines. the estimate is determined by speeds of 3 machines in the system: the first one and the second one which are the slowest, and the last one which is the fastest. theorem 3.1: for the system composed of n ≥ 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 ≤ · · · ≤ vn = s < 2 the price of anarchy is at least est(r, s) = min{ 2 + s (1 + r)(2− s) , 2 r(2− s) }. (3.2) proof as far as we need to prove the lower estimate it suffices to present examples of systems providing ratios of the optimal payoff and the worst-case equilibrium payoff given in the theorem condition. suppose that in the system each machine i has a speed vi for each i = 1, . . . , n . 1. let first rs ≤ 2. then est(r, s) = 2+s (1+r)(2−s) . consider the set of jobs: w1 = w2 = (1 + r)s, wi3 = vi(2 + s), where i = 3 . . . , n , w4 = 2r − s, w5 = 2− rs. in the equilibrium lmachinen receives jobsw1 andw2, each machine i− 1 receives each jobwi3, i = 3 . . . , n , jobs w4 and w5 are assigned to machine 1. we need to show that l is really an equilibrium and find the system payoff. the loads of machines n and 1 equal 2s(1 + r) and (1 + r)(2− s) respectively. the load of each machine i = 2, . . . , n − 1 equals vi+1(2 + s). since λn(l) = 2(1 + r) > (1 + r)(2− s) = λ1(l) and λi(l) = vi+1(2+s) vi ≥ (1 + r)(2− s) = λ1(l), i = 2, . . . , n − 1, due to 2 + s > 1 + r, vi+1 ≥ vi and 2− s ≤ 1, then machine 1 has the smallest delay which equals its load. denote the delay of machine i as λji (l) = λi(l) + wj vi in the case where some job j deviates from the profile l and moves to machine i from another machine. no one of jobs w1 or w2 moves to machine i, i = 2, . . . , n − 1, since λn(l) = 2(1 + r) ≤ (2 + s) + (1 + r) ≤ vi+1(2+s)+s(1+r) vi = λ1 i (l) = λ2 i (l). also no one of them moves to machine 1, since λn(l) = 2(1 + r) = (1 + r)(2− s) + s(1 + r) = λ1 1(l) = λ2 1(l). each of jobs wi3, i = 3, . . . , n , has not reason to move to machine n due to λi−1(l) = vi(2+s) vi−1 ≤ 2(1 + r) + vi(2+s) s = λi3n(l), that is equivalent to the inequality (s− vi−1)vi(2 + s) ≤ 2svi−1(1 + r) which holds true since s− vi−1 < 1, 2 + s < 4 2 s vi vi−1(1 + r) ≥ 4. also no one of jobs wi3, i = 3, . . . , n , moves to machine j > i− 1 since λi−1(l) = vi(2+s) vi−1 < 2vi(2+s) vj ≤ (vi+vj)(2+s) vj = λi3j (l). moreover, job wi3 does not move to slower machine 1 or j < i− 1, and no one job from machine 1 moves to another machine because the delay on machine 1 is minimal. therefore, the given profile is an equilibrium with the social payoff (1 + r)(2− s). consider the profile where each job wi3, i = 3, . . . , n belongs to machine i, jobs w1 and w4 are assigned to machine 2 and machine 1 receives jobsw2 andw5. the social payoff equals 2 + s in this profile, so, opt ≥ 2 + s. 2. let now rs > 2. then est(r, s) = 2 r(2−s) . consider the set of jobs: w1 = w2 = rs, wi3 = 2vi, i = 3, . . . , n , w4 = r(2− s). in the equilibrium l jobs w1 and w2 belong to copyright c© 2017 assa. adv syst sci appl (2017) 66 y.v. chirkova machine n , each job wi3, i = 3, . . . , n is assigned to machine i− 1, machine 1 receives job w4. we show that it is an equilibrium indeed and find the system payoff. since λn(l) = 2r > r(2− s) = λ1(l) and λi(l) = 2vi+1 vi ≥ r(2− s) = λ1(l), i = 2, . . . , n − 1, due to 2 vi ≥ 1, vi+1 ≥ r and 2− s < 1, then machine 1 has the smallest delay which equals r(2− s). job w1 or w2 does not move to machine i, i = 2, . . . , n − 1, since λn(l) = 2r = r + r ≤ 2vi+1+rs vi = λ1 i (l) = λ2 i (l), and also to machine 1 due to λn(l) = 2r = r(2− s) + rs = λ1 1(l) = λ2 1(l). no one of jobs wi3, i = 3, . . . , n , moves to machine n since λi−1(l) = 2vi vi−1 ≤ 2r + 2vi s = λi3n(l) due to 2vi(s−vi−1) s ≤ 2rvi. also no one of jobs wi3, i = 3, . . . , n , moves to machine j > i− 1, since λi−1(l) = 2vi vi−1 ≤ 4vi vj ≤ 2(vi+vj) vj = λi3j (l). moreover, job wi3 does not move to slower machine 1 or j < i− 1. no one job on machine 1 moves to other machines with not smaller delay. hence, the given profile l is an equilibrium with the social payoff r(2− s). consider the profile where each job wi3, i = 3, . . . , n , is assigned to machine i, jobs w1 and w4 belong to machine 2, and job w2 to machine 1. the social payoff equals 2 for this profile, thus, opt ≥ 2. in both considered cases the ratio of the optimal payoff and the equilibrium payoff equal est(r, s), hence, the price of anarchy is not less than this estimate. from the obtained estimate (3.2) we see that when the speed of the fastest machine increases and comes closer to the value of 2, the lower estimate of the price of anarchy grows infinitely. hence, we obtain the following corollary from the theorem 3.1. corollary 3.1: for the system composed of n ≥ 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 ≤ · · · ≤ vn = s < 2 the price of anarchy tends to infinity as s→ 2− 0. according to the following result, for poa evaluation it suffices to consider only games, where the optimal social payoff equals 1. theorem 3.2: for the system s, the price of anarchy constitutes poa(s) = max u1:opt (s,u1)=1 1 min l is a nash equilibrium in γ(s,u1,λ) sc(l) . proof we show that one can normalize job volumes in any game γ(s, u, λ) such that the optimal social payoff becomes equal 1 and the ratio of the optimal social payoff and the worst-case equilibrium payoff does not change its value. assume that l is the worst-case equilibrium in the game γ(s, u, λ) with an arbitrary set of players u(n,w). for each player j, the volume of its job equals wj , and the vector lopt gives the optimal strategy profile in this game. let sc and opt be the social payoff in the profile l and the optimal social payoff, respectively. the ratio of the optimal and worst-case equilibrium social payoff is defined by opt sc . so long as l represents an equilibrium, then for any player j we obtain that ∑ k=1,...,n:lk=lj wk vlj ≤ ∑ k=1,...,n:lk=i wk+wj vi for any machine i. now, explore the game with the same set of machines and players, where each player j has the job of volume wj opt . the social payoff in the profiles l and lopt constitutes sc opt and 1, respectively. by virtue of the linear homogeneity of machine delays in their loads, the profiles l and lopt form the worst-case equilibrium and optimal profiles, respectively, in the new game. particularly, the profile l is an equilibrium in the new game, since for any player copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 67 j the inequality ∑ k=1,...,n:lk=lj wk vljopt ≤ ∑ k=1,...,n:lk=i wk+wj viopt holds true for any machine i. imagine that l is any non-worst-case equilibrium in the new game. then the game admits an equilibrium l′ with social payoff sc′ opt such that the social payoff in the profile l′ is less than that in the profile l, i.e., sc′ opt < sc opt . however, in the initial game the profile l′ corresponds to the social payoff sc ′ < sc, and the equilibrium l′ is worse than its counterpart l. similarly, lopt gives the optimal profile in the new game. then the ratio of the optimal and the worstcase equilibrium social payoff in the new game also equals opt sc . consequently, any game γ(s, u, λ) corresponds to a game γ(s, u1, λ) with normalized job volumes such that opt (s, u1) = 1. moreover, the ratio of the optimal and the worstcase equilibrium social payoff is same in both games. hence, for poa evaluation it suffices to consider only games with unit optimal social payoff. 4. the case of 3 machines as a matter of fact, the exact poa value in the two-machine model was found in the paper [2]. consider the case of 3 machines in the system s. without loss of generality, throughout this section we believe that the machines have speeds v1 = 1 ≤ v2 = r ≤ v3 = s, i.e., machine 1 is the slowest one, machine 2 has medium speed and machine 3 is the fastest one. lemma 4.1: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s the inequality opt ≤ w−wk 1+r holds true for any job k with the volume wk. proof suppose that there is such job of the volume wk assigned to machine i in the optimal profile l, that opt > w−wk 1+r . then all optimal delays on machines exceed w−wk 1+r . moreover, it is clear that λi(l) ≥ wk vi . hence, w = viλi(l) + vjλj(l) + vlλl(l) > wk + (vj + vl) w−wk 1+r ≥ wk + (1 + r)w−wk 1+r = w . lemma 4.2: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s, if two jobs of volumes wk1 and wk2 are assigned to the same machine in the optimal profile, then opt ≤ w−wk1−wk2 1+r . proof assume for the sake of contradiction that opt > w−wk1−wk2 1+r and jobs of volumes wk1 and wk2 are assigned to machine i in the optimal profile. then all optimal delays on machines exceed w−wk1−wk2 1+r and λi(l) ≥ wk1+wk2 vi . then w = viλi(l) + vjλj(l) + vlλl(l) > wk1 + wk2 + (vj + vl) w−wk1−wk2 1+r ≥ wk1 + wk2 + (1 + r) w−wk1−wk2 1+r = w . theorem 4.1: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s < 2 the price of anarchy does not exceed est(r, s) = min{ 2+s (1+r)(2−s) , 2 r(2−s)}. proof in the proof we consider cases with a certain number of jobs assigned to each of two the most loaded machines having the largest delay. for each case we show that the price of anarchy estimate presented in the theorem condition is true. suppose that l is an equilibrium profile copyright c© 2017 assa. adv syst sci appl (2017) 68 y.v. chirkova and sc(l) = λi(l), that is machine i has the smallest delay. explore different cases of an equilibrium l. 1. each of machines j and l receives one job. in the optimal profile these two jobs occupy at most two machines. thus, there is some machine k in the optimal profile, taking partially or wholly the equilibrium load of machine i and nothing else. that is opt ≤ viλi(l) vk ≤ sλi(l). by lemma 8.1 s ≤ est(r, s). 2. machine j receives nj ≥ 2 jobs, machine l receives nl = 1 job. by lemma 3.2 λj(l) ≤ 2vi 2vi−vjλi(l). by lemma 4.1 opt ≤ viλi(l)+ 2vivj 2vi−vj λi(l) 1+r = λi(l) 2v2i+vivj (1+r)(2vi−vj) . a) assume first that vi ≥ vj . then 2v2 i + vivj ≤ 3v2 i , since this expression increases by vj . also 2vi − vj ≥ vi, so long as this expression decreases by vj . then opt ≤ λi(l) 3vi 1+r ≤ λi(l) 3s 1+r ≤ λi(l)est(r, s) by lemma 8.2. b) suppose now that vi < vj . then by lemma 8.3 2v2i+vivj (1+r)(2vi−vj) < 2+s (1+r)(2−s) . explore now two cases. in the first case assume that vi = r and vj = s. here by lemma 8.4 2r2+rs (1+r)(2r−s) < 2 r(2−s) . in the second case vi = 1. by lemma 3.3 wk ≤ vk 2−vk λi(l) ≤ s 2−sλi(l) and vjλj(l)− wk ≤ s 2−sλi(l) for any job of volume wk assigned to machine j. if all jobs assigned to machine j in the profile l remain there in the optimal profile, two cases are possible. if a single job assigned to machine l in the equilibrium keeps its position in the optimal profile, then the load of machine i can only decrease with system’s transition from the equilibrium to the optimal profile. if this single job leaves machine l, then in the optimal profile machine l receives the load at most λi(l) coming from machine i. in both cases opt ≤ λi(l). if jobs move from machine j only to machine l with system’s transition from the equilibrium to the optimal profile, similar two cases are possible. if a single job assigned to machine l in the equilibrium remains there in the optimal profile, then the load of machine i can only decrease with transition to the optimal profile. then opt ≤ λi(l). if this single job leaves machine l, then in the optimal profile machine l can receive the load at most λi(l) + s 2−sλi(l) coming from machines i and j. then opt ≤ λi(l) 1+ s 2−s vl ≤ λi(l) 1+ s 2−s r = λi(l) 2 r(2−s) . if some jobs move from machine j to machine i with system’s transition from the equilibrium to the optimal profile, we obtain the same two cases. if a single job assigned to machine l in the equilibrium remains there in the optimal profile, then machine j receives the load at most λi(l) + s 2−sλi(l) consisting from the load remaining on j and possibly coming from i. then opt ≤ λi(l) 1+ s 2−s vj ≤ λi(l) 1+ s 2−s r = λi(l) 2 r(2−s) . if this single job leaves machine l, then in the optimal profile machine l can receive the load at most λi(l) + s 2−sλi(l) coming from machines i and j. 3. each of machines j and l receives exactly two jobs: nj = nl = 2. the total number of jobs assigned to machines j and l is four, the number of machines is three, so in the optimal profile at least two of these jobs (wk1 and wk2) become assigned to the same machine. then by lemma 4.2 opt ≤ w−wk1−wk2 1+r = viλi(l)+wk3+wk4 1+r , where wk3 wk4 remaining two jobs assigned to machines j and l. consider machine k ∈ {j, l}. if vi ≤ vk, then by lemma 3.3 the volume of any of jobs assigned to machine k does not exceed λi(l) vivk 2vi−vk . let now vi > vk. l is an equilibrium, therefore λk(l) ≤ λi(l) + w vi , where w is the smallest volume job assigned to machine k. thus, w ≥ viλk(l)− viλi(l). by lemma 3.2 copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 69 λk(l) ≤ λi(l) 2vi 2vi−vk ≤ 2λi(l), 2vi − vk ≥ vi, w ≤ λi(l). another job of the largest volume assigned to machine k has the volume vkλk(l)− w ≤ vkλk(l) = viλi(l)− (vi − vk)λk(l) ≤ viλi(l)− (vi − vk)λi(l) = vkλi(l) ≤ rλi(l) ≤ λi(l) rs 2r−s ≤ λi(l) s 2−s . a) let vi = s, then opt ≤ λi(l) s+2r 1+r ≤ λl(l) 3s 1+r ≤ λi(l)est(r, s) by lemma 8.2. b) let vi = r, then opt ≤ λi(l) r+2 rs 2r−s 1+r = λi(l) 2r2+rs (1+r)(2r−s) ≤ λi(l)est(r, s) by lemma 8.3. c) let vi = 1, then opt ≤ λi(l) r+2 s 2−s 1+r = λi(l) 2+s (1+r)(2−s) . from the other side so long as the number of machines equals three there are surely two machines α and β receiving at most one job from considered four jobs and, perhaps, some part of the load of machine i. then opt does not exceed the minimal delay over these machines: opt ≤ min α6=β {λi(l) 2 vα(2−s) , λi(l) 2 vβ(2−s)} ≤ min{λi(l) 2 1(2−s) , λi(l) 2 r(2−s)} = λi(l) 2 r(2−s) . 4. machines j and l receive the following job allocation: nj ≥ 2, nl ≥ 3. by lemma 3.2 λj(l) ≤ λi(l) 2vi 2vi−vj and λl(l) ≤ λi(l) 3vi 3vi−vl . thus in accordance with the estimate (3.1), opt ≤ λi(l) vi+ 2vivj 2vi−vj + 3vivl 3vi−vl 1+r+s ≤ λi(l)est(r, s) by lemma 8.5 and lemma 8.6. the next theorem is a special case of theorem 3.1 for the system of three machines. theorem 4.2: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s < 2 is at least est(r, s) = min{ 2+s (1+r)(2−s) , 2 r(2−s)}. then we obtain from theorems 4.1 and 4.2 an exact value of the price of anarchy for the tree-machine system. theorem 4.3: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s < 2 the price of anarchy exactly equals { 2+s (1+r)(2−s) for rs ≤ 2, 2 r(2−s) for rs > 2. the exact poa value allows establishing a possibility for poa increase under new machine inclusion into the system, i.e., in a situation resembling the braess paradox [10–13] when system’s power increasing leads to its performance characteristic degradation. the next statement illustrates that the price of anarchy increases or does not change under new machine inclusion into the system of two machines. theorem 4.4: for the system s composed of two machines having speeds 1 ≤ s the price of anarchy does not decrease with adding a new machine of speed 1 ≤ q < 2. proof 1. suppose that new machine has a speed of q ≤ s. if qs ≤ s2 < 2 then the price of anarchy does not decrease since 2+s (1+s)(2−s) ≤ 2+s (1+q)(2−s) . if s2 > 2 and qs ≤ 2 then that does not decrease due to 2 s(2−s) ≤ 2+s (1+s)(2−s) ≤ 2+s (1+q)(2−s) . consequently, we have the same if s2 > 2 and qs > 2 since 2 s(2−s) ≤ 2 q(2−s) . 2. suppose now that new machine is more powerful than existing in the system, s < q < 2. if qs ≤ 2, then s2 ≤ 2, and the price of anarchy does not decrease since 2+s (1+s)(2−s) ≤ 2+q (1+s)(2−q) . if qs > 2 s2 ≤ 2, then that does not decrease due to 2+s (1+s)(2−s) ≤ 2 s(2−s) ≤ 2 s(2−q) . also if qs > 2 and s2 > 2, we obtain the same because of 2 s(2−s) ≤ 2 s(2−q) . copyright c© 2017 assa. adv syst sci appl (2017) 70 y.v. chirkova 5. evaluating the price of anarchy in the previous section, we have derived an analytic expression for the price of anarchy in the three-machine model, where the fastest machine possesses a rather high speed. in what follows, we suggest a computing method for the price of anarchy on the example of the system of 3 machines which is similar to a corresponding method for the load balance game [14]. this method can be generalized to systems composed of more machines. but such generalization increases the number of linear programming problems to-be-solved and the number of associated variables and imposed constraints. particularly the n -machine model requires n ! linear programming problems each of which includes (2n − 1)n−1 subproblems with n2 variables. consider the following system of linear equations in the components of the vectors a = (a1, a2, a3), b = (b1, b2, b3), c = (c1, c2, c3). a1+a2+a3 vi ≤ b1+b2+b3+ min k=1,2,3:ak>0 ak vj a1+a2+a3 vi ≤ c1+c2+c3+ min k=1,2,3:ak>0 ak vl b1+b2+b3 vj ≤ c1+c2+c3+ min k=1,2,3:bk>0 bk vl a1+a2+a3 vi ≥ b1+b2+b3 vj ≥ c1+c2+c3 vl max k=1,2,3 ak > 0 max k=1,2,3 bk > 0 ak, bk, ck ≥ 0, k = 1, 2, 3. (5.3) this system describes a set of hyperplanes passing through the point (0, 0, 0, 0, 0, 0, 0, 0, 0) in the 9-dimensional space, and the solution set represents a domain in the space bounded by the hyperplanes. the above system is feasible, as far as, e.g., the triplet a1 = a2 = a3 = αsi, b1 = b2 = b3 = αsj and c1 = c2 = c3 = αsl makes its solution for all α > 0. furthermore, the solution set is unbounded, since α can be arbitrarily large. study the system s composed of 3 machines having speeds 1 ≤ r ≤ s and n players. let l indicate a nash equilibrium in the system s such that machine i is slowest in this profile having the greatest delay, machine j has a medium delay and machine l is fastest. suppose that in the equilibrium l machine i receives the total volume of jobs defined by∑ k=1,...,n:lk=i wk = a1 + a2 + a3 and the corresponding volumes for machines j and l equal∑ k=1,...,n:lk=j wk = b1 + b2 + b3 and ∑ k=1,...,n:lk=l wk = c1 + c2 + c3, respectively. the volume of jobs on each machine is somehow divided into three parts so that each component of the three-dimensional vectors a, b and c is either zero or positive and includes at least one job. lemma 5.1: let l be a nash equilibrium in the game involving three machines i, j and l and n players such that λi(l) ≥ λj(l) ≥ λl(l),∑ k=1,...,n:lk=i wk = a1 + a2 + a3,∑ k=1,...,n:lk=j wk = b1 + b2 + b3,∑ k=1,...,n:lk=l wk = c1 + c2 + c3. here for all k = 1, 2, 3 component ak equals zero or the volume of at least one job on machine i, component bk equals zero or the volume of at least one job on machine j, and component copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 71 ck equals zero or the volume of at least one job on machine l. then the set of the vectors a, b and c is the solution of the system (5.3). proof suppose that l represents a nash equilibrium and λi(l) ≥ λj(l) ≥ λl(l). the lemma 3.1 claims that λk(l) > 0, k = i, j, l. in this case, the following inequalities take place: ∑ k=1,...,n:lk=i wk vi ≤ ∑ k=1,...,n:lk=j wk+ min k=1,...,n:lk=i,wk>0 wk vj∑ k=1,...,n:lk=i wk vi ≤ ∑ k=1,...,n:lk=l wk+ min k=1,...,n:lk=i,wk>0 wk vl∑ k=1,...,n:lk=j wk vj ≤ ∑ k=1,...,n:lk=l wk+ min k=1,...,n:lk=j,wk>0 wk vl∑ k=1,...,n:lk=i wk vi ≥ ∑ k=1,...,n:lk=j wk vj ≥ ∑ k=1,...,n:lk=l wk vl . since each nonzero quantity ak (k = 1, 2, 3) equals the volume of at least one job on machine i, then we naturally have that min k:ak>0 ak ≥ min k:lk=i,wk>0 wk that provides satisfaction of the first and the second inequality of the system (5.3). similarly, min k:bk>0 bk ≥ min k:lk=j,wk>0 wk. this means satisfaction of the system (5.3). lemma 5.2: any solution of the system (5.3) defines a nash equilibrium l in the game involving the system s composed of 3 machines i, j and l and players whose jobs correspond to the nonzero components of the vectors a, b and c and the delays are sorted in the order λi(l) ≥ λj(l) ≥ λl(l). proof assume that the set of the vectors a, b and c gives the solution of the system (5.3). consider the game with 3 machines i, j and l. let each nonzero component of the vectors a, b and c specify the job volume of a regular player. consider a profile l such that the jobs of volumes ak > 0, bk > 0 and ck are assigned to machines i, j and l, respectively. so long as all inequalities (5.3) hold true, the profile l gives the desired nash equilibrium. the following result is immediate. theorem 5.1: any nash equilibrium l in the game involving the system s composed of 3 machines i, j and l and n players corresponds to a nash equilibrium l′ in the game involving the same system s and at most 9 players, where each machine receives no more than 3 jobs and the delays on all machines in l and l′ do coincide. proof consider a nash equilibrium l in the game with the system s of 3 machines and n players. number the machines so that λi(l) ≥ λj(l) ≥ λl(l). according to lemma 5.1, for any nash equilibrium in the game involving the system s and any number of players there exist a corresponding solution a, b, c of the system (5.3). by virtue of lemma 5.2, this solution determines a nash equilibrium l′ in the game with the system s such that the nonzero components of the vectors a, b and c specify the job volumes on machines i, j and l, respectively. by definition, the element sum of the vector a represents the load of machine i in a profile l. hence, delays on machine i coincide in both equilibria l and l′. similarly, for machines j and l the delays in the equilibrium l coincide with the corresponding delays in the equilibrium l′. copyright c© 2017 assa. adv syst sci appl (2017) 72 y.v. chirkova this theorem claims that it is sufficient to consider only games, where in an equilibrium each machine receives at most three jobs and the equilibrium solves the system (5.3). and the domain of the social payoff coincides with the value domain of games with an arbitrary number of players. imagine that the components of the vectors a, b and c are chosen as follows. in the optimal profile yielding the maximum social payoff, machines i, j and l receive the total volumes of jobs a1 + b1 + c1, a2 + b2 + c2 and a3 + b3 + c3, respectively, and the lowest delay can be on each of them. furthermore, by theorem 3.2, the volumes of jobs are assumed to be normalized so that in the optimal profile the maximum delay among all machines equals 1. in our case, this means that a1 + b1 + c1 ≥ vi, a2 + b2 + c2 ≥ vj, a3 + b3 + c3 ≥ vl, and at least one of these inequalities holds as an equality. lemma 5.3: solution of the linear programming problem lpp (vi, vj, vl) :  c1 + c2 + c3 → min (r1) a1+a2+a3 vi ≤ b1+b2+b3+ min k:ak>0 ak vj (r2) a1+a2+a3 vi ≤ c1+c2+c3+ min k:ak>0 ak vl (r3) b1+b2+b3 vj ≤ c1+c2+c3+ min k:bk>0 bk vl (r4) a1+a2+a3 vi ≥ b1+b2+b3 vj ≥ c1+c2+c3 vl (r5) max k=1,2,3 ak > 0 (r6) max k=1,2,3 bk > 0 (r7) ak, bk, ck ≥ 0, k = 1, 2, 3 (r8) a1 + b1 + c1 ≥ vi (r9) a2 + b2 + c2 ≥ vj (r10) a3 + b3 + c3 ≥ vl (5.4) with respect to the components of the vectors a, b and c provides the minimal social payoff in a nash equilibrium among all games, where in an equilibrium at most 3 jobs are assigned to each machine, i, j and l indicate the numbers of the machines in the descending order of their delays and the optimal social payoff makes up 1. proof due to lemma 5.2, any solution of inequalities (r1)− (r7) in the problem lpp (vi, vj, vl) defines an equilibrium in the game with 3 machines, where each machine receives at most 3 jobs and i, j, and l are the numbers of machines in the descending order of their delays the goal function in this game is bounded above only by the hyperplanes corresponding to inequalities (r8)− (r10). actually, inequalities (r1)− (r7) admit arbitrarily small nonnegative values of the goal function, including zero. therefore, the minimum is reached on one of the boundaries answering to the last three inequalities. this guarantees that one of them is satisfied as an equality, ergo the optimal payoff in the game corresponding to the solution of the problem lpp (vi, vj, vl) equals 1. consequently, exact poa evaluation for the system s composed of 3 machines calls for solving a series of linear programming methods lpp (vi, vj, vl) for all permutations (1, r, s). and the minimum solution among them yields the value of poa(s). in other words, it is possible to establish the following fact. copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 73 theorem 5.2: for the system s composed of 3 machines having speeds v1 = 1 ≤ v2 = r ≤ v3 = s < 2, the price of anarchy constitutes poa(s), which is the inverse value of 1 poa(s) = min (vi,vj ,vl) are permutations of (1,r,s){ c1+c2+c3 vl |a, b, c is a solution of lpp (vi, vj, vl) } , where lpp (vi, vj, vl) is the linear programming problem (5.4). proof according to lemma 5.3, the solution of the problem (5.4) gives the minimum social payoff in a nash equilibrium, where i, j and l are the numbers of the machines in the descending order of their delays, among all games such that in an equilibrium each machine receives at most 3 jobs and the optimal payoff equals 1. the minimum solution among the problems for all admissible permutations (1, r, s) as the values of (vi, vj, vl) provides the minimum social payoff in a nash equilibrium among all games, where in an equilibrium at most 3 jobs are assigned to each machine and the optimal payoff equal 1. by theorem 5.1, for any equilibrium in the game involving the system s of 3 machines and an arbitrary number of players, it is possible to construct a corresponding equilibrium in the game with the same machines and a set of at most 9 players, where each machine receives no more than 3 jobs and the social payoff coincides for both equilibria. thus, for poa evaluation it suffices to consider only games, where in an equilibrium each machine has at most 3 jobs. using theorem 3.2, we finally obtain that for poa evaluation it suffices to consider only games, where the social payoff in the optimal profile equal 1. 6. computing experiments to estimate the price of anarchy in the three-machine model, we have developed a program implementation of poa evaluation method presented in the previous section. this program allows to compare visually the theoretic poa value and its exact value constructed by solving a series of linear programming problems. moreover the program provides the possibility to see the poa dynamics for the machine number n > 3 where no any theoretic poa estimates are obtained. the parameters of the system s act as the options in the program; by assumption, the speed of machine 1 equals 1, whereas an exact value and a certain range are specified for the speeds of machines 2 and 3, respectively. in this case, users can study the poa dynamics under variations in the speed of one machine. the figures 6.1 and 6.2 present examples of poa estimates for different values of speeds of machine 2 and 3. at the fig. 6.1 the speed of machine 2 is r = 1.1, and the speed of machine 3 is s ∈ [r, 2). at the fig. 6.2 the speed of the fastest machine 3 is s = 1.7 and the speed of machine 2 is r ∈ [1, s]. here we can see that theoretical and computed values of poa coincide. the next example is more interesting. consider the system of four machines with speeds v1 = 1 ≤ v2 = q ≤ v3 = r ≤ v4 = s < 2. figures 6.3 and 6.4 present poa comparing with the lower poa estimate (3.2), which in fact is poa for the system composed of 3 machines with speeds 1 ≤ r ≤ s < 2. fig. 6.3 presents poa for the following cases. at the area a the value of q changes in the range [1, r], r = 1.3, s = 1.5. at the area b q = 1.3, the value of r changes in the range [q, s], s = 1.5. at the area c q = 1.3, r = 1.5, s value changes in the range [r, 2). in these cases poa value for four-machine system coincide with its lower estimate (3.2). fig. 6.4 presents the poa dynamics for those systems where machine speeds differ rather little, that is normalized speeds are closed to 1. in this case one can see that the poa value copyright c© 2017 assa. adv syst sci appl (2017) 74 y.v. chirkova 1 2 3 4 5 6 7 8 9 10 1.1 1.2 1.3 1.4 1.5 1.6 1.7 1.8 1.9 2 s fig. 6.1. poa for the system s, where r = 1.1, s ∈ [r, 2). 4 4.5 5 5.5 6 6.5 1 1.1 1.2 1.3 1.4 1.5 1.6 1.7 r fig. 6.2. poa for the system s, where s = 1.7, r ∈ [1, s]. 1 2 3 4 5 6 7 8 9 10 1 1.2 1.4 1.6 1.8 2 q r s a b c fig. 6.3. poa for the four-machine system s. presented by the thin curve exceeds its estimate (3.2) presented by the bold curve. both curves coincide under machine speeds increasing. at the area a the value of q changes in the range [1, r], r = 1.05, s = 1.1. at the area b q = 1.05, the value of r changes in the range [q, s], s = 1.1. at the area c q = 1.05, r = 1.1, the value of s changes in the range [r, 1.3). copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 75 1 1.5 2 2.5 3 1 1.05 1.1 1.15 1.2 1.25 1.3 q r s a b c fig. 6.4. poa for the four-machine system s with small speeds 7. conclusion this paper has explored the service system composed of n machines and n players and derived the lower estimate for the price of anarchy in the maximizing the minimum machine delay game (or cover game). the three-machine model has been analyzed in detail. here we have determined the exact value of the price of anarchy and showed that the poa increases or does not change under new machine inclusion into the system of two machines. also we have proposed a computing algorithm of the exact poa value. the algorithm can be generalized to systems with more machines, but this increases the number of linear programming problems to-be-solved and the number of associated variables and imposed constraints. and finally, we have implemented the algorithm as a program and conducted numerical experiments for comparing the obtained estimates of the poa with its exact value. the results of these experiments have demonstrated the correctness of the derived estimates. for the case of fourmachine system computing experiments demonstrate partial poa coinciding for three and four-machine systems, the analytic confirmation of this fact needs further investigations. acknowledgements this work was supported by the mathematical sciences branch of the russian academy of sciences and by the russian foundation for basic research (project no. 16-51-55006), and by the russian humanitarian scientific foundation (project no. 15-02-00352 a). 8. appendix this appendix contains supporting lemmas which are used in proofs in section 4. lemma 8.1: for any real r, s such that 1 ≤ r ≤ s < 2 we have s ≤ min{ 2+s (1+r)(2−s) , 2 r(2−s)}. proof 2 r(2−s) ≥ 2 s(2−s) ≥ s, since s3 − 2s2 + 2 = s(s− 1)2 + (2− s) > 0. 2+s (1+r)(2−s) ≥ 2+s (1+s)(2−s) ≥ s, since s3 − s2 − s+ 2 > s3 − 2s2 + 2 = s(s− 1)2 + (2− s) > 0. copyright c© 2017 assa. adv syst sci appl (2017) 76 y.v. chirkova lemma 8.2: for any real r, s such that 1 ≤ r ≤ s < 2 we have 3s 1+r ≤ min{ 2+s (1+r)(2−s) , 2 r(2−s)}. proof first, 3s ≤ 2+s 2−s , due to 3s2 − 5s+ 2 = (s− 1)(3s− 2) > 0. second, 3s 1+r ≤ 2 r(2−s) , since 6rs− 3rs2 − 2− 2r = r(6s− 3s2 − 2)− 2 = r(1− 3(s− 1)2)− 2 ≤ r − 2 < 0. lemma 8.3: for vi < vj , vi, vj ∈ {1, r, s}, where real r, s are such that 1 ≤ r ≤ s < 2, we have 2v2i+vivj 2vi−vj ≤ 2+s 2−s . proof if vi < vj , then 2v2i+vivj 2vi−vj decreases by vi and increases by vj , since 4v2 i − 4vivj − v2 j < 0 and vi(2vi − vj) + 2v2 i + vivj > 0. lemma 8.4: for any real r, s such that 1 ≤ r ≤ s < 2 we have 2r2+rs (1+r)(2r−s) < 2 r(2−s) . proof the inequality in the condition is equivalent to f(r, s) = −r2s2 − 2s(r3 − r2 − r − 1) + 4(r3 − r2 − r) < 0, chech if it holds true. show that f ′r(r, s) = −2rs2 + 2(2− s)(3r2 − 2r − 1) < 0, then f(r, s) ≤ f(1, s) = −s2 + 4s− 4 = −(2− s)2 < 0. for each fixed s the function f ′r(r, s) is a parabola with branches directed upwards. thus, its largest value is achieved on one of two ends of the interval r ∈ [1, s]. at the left end f ′r(1, s) = −2s2 < 0. at the right end f ′r(s, s) = −8s3 + 16s2 − 6s− 4 = −8s(s− 1)2 − 2(2− s) < 0. lemma 8.5: for vi 6= vj 6= vl, vi, vj, vl ∈ {1, r, s}, where real r, s are such that 1 ≤ r ≤ s < 2, we have f(vi, vj, vl) = vi + 2vivj 2vi−vj + 3vivl 3vi−vl ≤ 1 + 2s 2−s + 3s 3−s . proof the function f(vi, vj, vl) obviously increases by vj and vl, hence f(vi, vj, vl) ≤ vi + 2svi 2vi−s + 3svi 3vi−s = g(vi). show that g(vi) decreases by vi. the derivative g′vi(vi) = 1− 2s2 (2vi−s)2 − 3s2 (3vi−s)2 increases by vi and, therefore, does not exceed g′vi(s) = 1− 2− 3 4 < 0. then g(vi) ≤ g(1) = 1 + 2s 2−s + 3s 3−s . lemma 8.6: for any real r, s such that 1 ≤ r ≤ s < 2 we have 1+ 2s 2−s+ 3s 3−s 1+r+s ≤ min{ 2+s (1+r)(2−s) , 2 r(2−s)}. proof we show first that 1 + 2s 2−s + 3s 3−s ≤ (1+r+s)(2+s) (1+r)(2−s) . the right part of the inequality decreases by r, thus it suffices to show that 1 + 2s 2−s + 3s 3−s ≤ (1+2s)(2+s) (1+s)(2−s) . this is equivalent to s ≤ s2, that holds true under s ≥ 1. show now that 1 + 2s 2−s + 3s 3−s ≤ 2(1+r+s) r(2−s) . the right part of the inequality decreases by r, so it suffices to show that 1 + 2s 2−s + 3s 3−s ≤ 2(1+2s) s(2−s) . this inequality is equivalent to −4s3 + 11s2 − 4s− 6 = −s(2s− 3)2 − (2− s)(3− s) < 0. copyright c© 2017 assa. adv syst sci appl (2017) price of anarchy for maximizing the minimum machine load 77 references 1. epstein l., kleiman e., & van stee r. (2009) maximizing the minimum load: the cost of selfishness. in proceedings of the 5th international workshop on internet and network economics, lncs 5929, 232-243. 2. tan z., wan l., zhang q. & ren w. (2012) inefficiency of equilibria for the machine covering game on uniform machines. acta informatica, 49 (6), 361–379. 3. wu y., cheng t. c. e. & ji m. (2015) inefficiency of the nash equilibrium for selfish machine covering on two hierarchical uniform machines. information processing letters, 115 (11), 838–844. 4. andelman n., feldman m. & mansour y. (2007) strong price of anarchy. in proc. of the 18th annual acm-siam symposium on discrete algorithms (soda), 189–198. 5. koutsoupias e., papadimitriou c. h. (1999) worst-case equilibria. in proc. of stacs 1999, 1563, 404–413. 6. lücking t., mavronicolas m., monien b., rode m., spirakis p. & (2003) vrto i. which is the worst-case nash equilibrium? in proc. of the 26th international symposium on mathematical foundations of computer sciencem, lncs 2747, 551–561. 7. even-dar e., kesselman a. & mansour y. (2003) convergence time to nash equilibria. in proc. of the 30th international colloquium on automata, languages and programming (icalp2003), 502–513. 8. fotakis d., kontogiannis s. c., koutsoupias e., mavronicolas m. & spirakis p. g. (2002) the structure and complexity of nash equilibria for a selfish routing game. in proc. of the 29th international colloquium on automata, languages and programming (icalp2002), 123–134. 9. chen x., epstein l., kleiman e. & van stee r. (2013) maximizing the minimum load: the cost of selfishness. theoretical computer science, 482, 9–19. 10. murchland j. d. (1970) braess’s paradox of traffic flow. transportation research, 4, 391–394. 11. roughgarden t. & tardos é. (2002) how bad is selfish routing? journal of the acm, 49 (2), 236–259. 12. korilis y. a., lazar a. a. & orda a. (1999) avoiding the braess paradox in noncooperative networks. journal of apllied probability, 36 (1), 211–222. 13. mazalov, v. v. (2014) mathematical game theory and applications. new york, ny: wiley. 14. chirkova yu. v. (2015) price of anarchy in machine load balancing game. automation and remote control, 76 (10), 1849–1864. copyright c© 2017 assa. adv syst sci appl (2017) introduction the model the general case of n machines the case of 3 machines evaluating the price of anarchy computing experiments conclusion appendix advances in systems science and applications (2012) vol.12 no.3 220-237 creativity & biodiversity: towards a synergy-of-synergies john wood emeritus professor of design, goldsmiths, university of london, london se14 6nw abstract the paper argues that, in order to curb the current acceleration of species losses, a massive systemic change is needed. however, we are unlikely to achieve this with current methods. paradigm change includes the need to rethink the prevailing ‘realities’ and governances upon which the current paradigm depends. as designers are expressly trained to ‘think creatively’ this paper argues that their methods might usefully augment those used by scientists and politicians. however, as these design methodologies evolved as part of the old paradigm they, too, need to be re-designed. the paper explores some research into ‘metadesign’ a non-hierarchical, self-organizing framework for practice that helps designers to re-design their own mission and working vernacular, when required. in work that began in 2002, instead of seeing design as the process of creating products and services, we sought to cultivate ‘synergy’ at many levels. however, synergies tend to operate outside familiar boundaries of language and custom, which means that metadesigners must develop a metadiscourse, with which to explore each new predicament. this process encourages the radical ‘relanguaging’ of purposes, roles and self-identities. although, in the hierarchical sense, this makes these tasks ‘unmanageable’, politicians can facilitate local change by encouraging citizens to cultivate a ‘diversity-of-diversities’. this would act as an attractor to hitherto unforeseen orders, entities and configurations. it would offer an interoperable framework, within which communities could self-orchestrate a more complex, and emergent, ‘synergy-of-synergies’. keywords biomimicry, biodiversity, co-sustainment, creativity, ecomimicry, languaging, metadesign, paradigm, sympoiesis 1 the 6th great extinction ‘economies of scale’ is a term that emerged from many imperialistic enterprises over the last five or ten thousand years. while it probably started with the invention of agriculture, humans later applied it to the building of heavy structures and the management of large bureaucracies and manufacturing plants. but the discoveries that led to these achievements engendered a presumptuous mode of formalism that migrated to many other disciplines, including engineering, education, business and economics[1]. while it has facilitated rapid population growth and made many lifestyles more comfortable, it has also led to environmental pollution, habitat destruction and biodiversity contraction. hence, while politicians advances in systems science and applications (2012) vol.12 no.3 221 compete with one another to squeeze yet more ‘efficiency’ and ‘productivity’ out of a damaging and dysfunctional economic system, the rate of species extinctions is rising. some distinguished scientists, such as e.o. wilson, warn that, over the next few decades, this could rise to 10,000 times the ‘background’ rates. although an event of this magnitude would probably be fatal for human beings, it would not be without precedent. richard leakey’s term ‘6th extinction’(leakey & lewin, 1996) gives a 600 million year context to what is happening now. if these experts are right, our offspring could witness a halving of the number of living species by 2100, thus putting the survival of homo sapiens in grave doubt. this should not be surprising. from this perspective, systems of biodiversity seem less manageable, useful or attractive than monocultures. this approach has made a global force majeure more likely, probably in the form of violent climatic changes and agricultural failures. known remedies are likely to be too weak or inapplicable because they reflect, and uphold, the paradigm that caused the malady. in short, whereas nature is a system that operates on an ‘ecologies-of-scale’ basis, we have sought to manage it using an ‘economies-of-scale’ approach. 2 dysfunctional governance in order to address these problems, a more systemic and ecomimetic model is required. as systemic changes would need to be self-inclusively regulated, at virtually every level, this calls for a more joined-up, self-reflexive and innovative approach. in other words, we need to introduce a second-order cybernetic system[2] in place of the, largely first-order, system that is in place. rather than following existing approaches, our ‘metadesign’ methodology is intended to help us understand how to adopt an ‘ecology-of-scale’ approach[3-4]. most human societies are accustomed to a top-down system of government in which the management of major crises are seen, primarily, as the duty of politicians, scientists and civil servants. in the 1990’s, concerned by the apparent inability of ngos to influence radical changes in behaviour, donella meadows (1997) found that the most familiar measures were the least effective. this is probably because those most commonly used were bureaucratic (i.e. targets, subsidies, taxes, legislation) and, therefore, the least direct. again, the formalism implicit in these processes can be traced back to aristotelean and pythagorean logic[1]. while many regard these as axioms of truth, their dependability is largely epistemological. by contrast, living systems are less predictable because they are a great deal more complex and distributed. even the setting, and attaining, of targets can be a dispiriting process, especially when they are politically ambitious or controversial. if they are set lower than what is required, they may not easily be reached within the agreed timescale. moreover, when targets are missed (this is not uncommon), the morale of the parties involved may be reduced, thus creating a 222 john wood:creativity & biodiversity: towards a synergy-of-synergies positive feedback loop that can trigger, or amplify, a shared sense of futility in the mission. in seeking to think beyond this kind of sub-optimal system meadows looked for what she called the ‘levers of change’[5]. she noted that it is important to identify, and to reframe, the agreed purpose of a given paradigm in order to change it. this lesson has yet to be adequately acknowledged and applied by governments, who persist in applying a targets-based approach. however, in 2010 the nagoya world biodiversity summit sought to move beyond the customary agenda of deepening scientific knowledge and raising public awareness. it also agreed to assign large regions of land and the sea to act as regions of wilderness, in order to give endangered species a chance to recover. unfortunately, it soon became clear that the plan would fail[6] because the natural rate of species replenishment in the wild is too low to reverse the current rate of losses. 3 the role of science it seems unlikely that applying current, hierarchical, top-down remedies for biodiversity losses will be effective within the timescale we have. instead, we might start by differentiating between, then merging, the functions of scientists, politicians and designers, as their respective professions reflect complementary approaches. from the designer’s perspective, for example, governments tend to deal with major problems by listening to ‘scientific’ advice, then trading expediencies in a way that maintains the political status quo. whereas designers are trained to deliver desirable, pragmatic solutions within a short timescale, scientists are trained to regard evidence-based knowledge and reasoning as vital prerequisites to prudent practice. dr. brundtland, former director-general of the world health organization famously likened the problem of biodiversity to a major fire: “the library of life is burning and we do not even know the titles of the books”[7]. perhaps it is because i am more of a designer than a scientist i find this to be a peculiar analogy. if i am watching the library burn, my instinct is to put out the flames, not to lament my lack of knowledge. however, a scientist would probably argue that finding sensible remedies cannot work unless we can differentiate between species that look identical, yet are at very different levels of risk. more than three centuries after his birth, carl linnaeus (1707-1778) is still highly praised for simplifying the elaborate nomenclature that previous scholars used to differentiate between species. he is also celebrated for his clear taxonomy of classification. nonetheless, while his efforts may have made it easier to document all life on earth, the task has proved much too big to accomplish. recent studies estimate that 86% of all land species and 91% of all sea species are unclassified or undiscovered[8]. advances in systems science and applications (2012) vol.12 no.3 223 4 can scientists work more closely with designers? much of the problem lies not with science, but with the way that scientific knowledge tends to be managed. in the fishing industry, for example, just as politicians may be tempted to exaggerate figures when setting catch quotas, so fisheries are equally tempted to exceed these quotas. this led to the scandalous process of ‘discards’, trawled fish that exceed the official quotas[9]. with all due respect to scientists, we cannot wait for ‘robust data’, evidence-based reasoning, or so-called ‘rigorous’ analysis of the problem in order to address it. in other words, we cannot expect to name and classify all of the relevant species. therefore, while we know that a ban on bottom trawlers would probably save some species of marine life, it is difficult to know which ones. on the other hand, it is enough to know that some ‘fish fingers’ (or ‘fish sticks’) may contain a proportion of unnamed and/or unknown species. this is closer to design practices, in which designers often have to make formative decisions before there is clear, evidence-based knowledge. moreover, it would be politically difficult to do so, unless alternative technologies and business models can be offered. this is where design thinking is useful, especially if it can be brought into synergy with the gathering of scientific truths, and the making of political decisions. these hybrid systems of governance would need to embrace what i have called ‘auspicious reasoning’[1], which is contingent and outcome-focused, rather than knowledge-based, or truth-focused. a well-known example of ‘inauspicious reasoning’ is the ‘tragedy of the commons’[10], which reminds us that humans often choose actions with selfish and, or, short-term gratification, rather than settling for small, immediate reductions in well-being that will ensure greater long-term safety, or abundance, for all. 5 the importance of synergy since darwin, the life sciences have offered a worldview in which living systems (appear to) defy entropy and create emergent forms of abundance. unfortunately, the dominant paradigms of governance appear to reflect a less optimistic view that emerged from physics, inspired by the exploitative logic of mining. both traditions depict the world as a materials and energy system that is ‘running out’, or ‘running down’. this unhelpful mindset can also be found in school curricula, and in the non-ecological terms used by most environmentalists (e.g. in the notions of ‘resources’, ‘entropy’ and ‘sustainability’). there are very few ‘resources’ that are beneficial when consumed by themselves. indeed, peter corning claims that synergy is the cause behind the evolution of complexity in living systems[11]. buckminster fuller was probably the first person to show that, by designing for synergy, rather than for better products or services, designers would achieve more with less. he defines it as “...the behaviour of whole systems unpredicted by the separately observed behaviours of their parts taken separately”[12]. in a simple 224 john wood:creativity & biodiversity: towards a synergy-of-synergies example, combining nickel, iron and manganese to makes stainless steel, which is up to 35% stronger than any of its constituents. ‘synergy’ also became a key term within our metadesign research. this is partly because it applies equally to systems, whether they are seen as animate, or inanimate. but, despite its invaluable benefits, the task of seeking, harvesting and harnessing synergies may raise complex epistemological, technological, or managerial difficulties. this is probably because it appears in such a diverse and, often, elusive forms. for example, spontaneous humour, or a spiritually uplifting experience, are no less synergistic than the combinatorial benefits of stainless steel. in practice, many are hard to disentangle from the processes with which one thinks, experiences, shares and acts. fig.1 laufrad campa vento bicycle wheel fig.2 millstone 6 creating synergies-of-synergies the bicycle is said to be the most energy-efficient mode of transport on the planet. we can learn a great deal from how its vast array of subtle alignments (i.e. a ‘synergy-of-synergies’) is configured. as it is easier to analyse some of its simpler synergies, we will use these as a basis from which to reflect on more complex features. fig.1 shows a modern design whose metallurgical synergies combine with synergies of form. the use of stainless steel (see previous paragraph) for making wheel spokes ensures that they are considerably lighter than a more solid design. this is because stainless steel is far stronger when under tension, than in compression. these, and a totality of other synergies, make it capable of supporting a load that is up to seven hundred times its own weight. as this example may sound like the traditional ‘economies-of-scale’ logic, how might we advances in systems science and applications (2012) vol.12 no.3 225 retrieve ‘ecologies-of-scale’? here, we should not forget that, like any ecosystem (or ‘paradigm’) the bicycle is part of a much larger system that includes, in this case, a suitable cyclist, a transport culture and a network of flat roads. 7 creating synergies-of-synergies let us look at a more qualitative mode of synergy. for example, when brought together in the appropriate way, the poisons, chlorine and sodium, become nutritionally useful and palatable in the form of salt. unfortunately, simple, ‘freestanding’ synergies are less common than complex, messy ones, because most synergies have more than one property (e.g. most stainless steels last far longer than ordinary steels because they do not rust). hence, we probably fail to notice new synergies[13] even though they may be abundant and ubiquitous. in many cases they operate outside what can easily be described. this underlines the subjective nature of synergy, and its dependence on our creative ability to imagine conditions beyond the current affordances of ones language. this is similar to the challenge of supporting hitherto unnoticed, unnamed or untaxonomised species. however, it also suggests that we can adapt to the universe in a myriad of unforeseen ways. how might we set about creating spiritual states, or even miracles[14]? the first step is almost an act of faith, because observation is an emotional, as well as a rational faculty. this can begin with the logical assertion that ‘unthinkable’ things may be attainable, once we begin to notice them. after all, until quite recently, ships made of iron, flying machines, and quantum logic were all ‘unthinkable’ or ‘impossible’. 8 learning from synergistic complexity for example, by exploring synergies, a good cook can combine ordinary, local, inexpensive cooking ingredients and turn them into an extraordinary experience for her, or his, guests. this example includes very many types of synergy from the way heat melts cheese, or butter (i.e. simple physics) to the more complex nutritional benefits of combining, say, broccoli and sprouts[15]. the way that people co-create a unique, indefinable atmosphere together is even more complex and intangible. one thing we can learn from food synergies is the way that complementary flavours work. some molecular gastronomists predict that any two foods with a common (molecular) ingredient will taste good when combined. also, some apparently incompatible substances, such as oil and water can be made to mix when an emulsifier is added. this works if the emulsifier’s molecular chain has a water-compatible atom at one end and an oil-compatible atom at the other. these lessons from food technology can be applied, via metaphors, to other systems. for example, ‘immiscible’ ingredients, such as oil and water, can be combined when an emulsifier is added. the same principle could be applied 226 john wood:creativity & biodiversity: towards a synergy-of-synergies to teams containing antagonistic members. hence, differences might be resolved (‘emulsified’) via a third party who seems benign to each of the antagonists. ecosystems can be used as the model for a new order of innovation. for one reason, their complexity and interdependency that makes them resilient within a community context. we know, for example, that ‘keystone species’ are critical to the survival of other species, so we might set about designing ‘keystone synergies’ that would engender, and sustain, subordinate synergies within a given environment or context(see fig.2). another way to explain this process is by thinking of it as a singular act that behaves as a ‘manifold innovation’ that brings many benefits to many stakeholders. a good historical example is the millstone, which attracted many subsequent social, cultural, nutritional and technological synergies. 9 continuously re-languaging nature while we might assume, hypothetically speaking, that ecological diversity can be enhanced, it may be hard to see how this can be achieved. this is partly because the ‘language’ we use to enframe nature is always less than adequate to the situation. systems theory helped us to recognise the important interdependencies between meanings and actions in ecosystems. however, this relationship embodies a paradox. although language shapes our ‘realities’ and guides our actions[16-17], nature is ineffable and emergent, and will therefore defy clear and enduring definition[18]. hence, while new ideas may only become popularly understood when suitable terminology is found, this may have a limited era of usefulness. sometimes, seemingly simple keywords prove too difficult, or unpopular, to be accepted. for example, the rather narrow term ‘ecological footprint’[19] soon gave way to the (even narrower) idea of ‘carbon footprint’[19], thus giving the false impression that the complex phenomena that cause climate change can be curbed by simple economic transactions that ‘offset’ carbon emissions, or by geoengineering solutions that aim to ‘remove’ it from the environment. arguably, in living systems, the task of creating, modulating and switching meanings cannot be managed successfully in a top-down, external, or hierarchical way, especially when the hierarchy grows and becomes many-layered. this follows from ross ashby’s law of requisite variety, which warns of the dangers of arbitrarily reducing the number of variables within subsystems[20]. maturana and varela offer a useful ecomimetic account of living systems, by explaining that they balance the ‘meaning’ of their internal and their external identities[21]. 10 re-languaging sustainability the above description suggests that the survival of a given species is, at least, explicable within systemic terms that transcend ethics. no organism deserves advances in systems science and applications (2012) vol.12 no.3 227 longevity because of an a priori moral right or because it enjoys a privileged status. this is because there a seldom a simple, intuitive logic of cause-and-effect, or a fixed hierarchy of relations. here, it is the ambiguity of the popular term ‘sustainable’[22] that has rendered it unhelpful. the verb ‘to sustain’ denotes a continuation over time, but it may also carry the non-temporal meaning of something ‘holding together’. using the verb ‘to sustain’ transitively and temporally (e.g. ‘b’ sustains the continued existence of ‘a’) would clarify the direction of causation. however, merely saying that something is ‘sustainable’ does not make clear who ‘sustains’ what. in any case, reciprocal (2-way) relationships are prevalent in ecosystems, even though this may seem counterintuitive. the death of a predator species, for example, may seem to ‘cause’ the death of its prey. i therefore use the term ‘co-sustainment’ instead of ‘sustainability’, as it acknowledges the mutual dependency of all living systems[23]. evolution adjusts relationships all the time, which means that everything is always changing. in order for a species to endure i.e. to remain a viable part of the whole it must always be ready to adapt to its changing habitat. this means adjusting our actions and identities, rather than trying to re-design nature in accordance with our expectations. inviting designers to work at this level would also mean re-designing the design paradigm that is part of the problem. 11 changing the change in practical terms, we have found it helpful to change other key terms that were being applied in what we saw as a dubious, or unhelpful way. for example, we stopped using certain terms (e.g. ‘creativity’, ‘sustainability’ and ‘biomimicry’) and designed more useful alternatives (i.e. ‘sympoiesis’, ‘co-sustainment’ and ‘ecomimicry’). the idea of ‘languaging’ change is not only a theory about the seamlessness of interplay between action and thought. it is, also, a practical way for design teams to co-create their ‘survival’ as whole, living systems. however, this active, radical, consensual reframing of meaning means that, of necessity, the participant’s perceived reality will also change (we are aware that this also applies to us, as ‘metadesigners’). we have, therefore, found it useful to follow maturana and varela’s practice of using the word ‘language’ as a verb. as this change would be both shared and self-reflexive, our new, shared self-identity cannot be separated from our behavioural culture. this pluralisation of identity, led us to coin the term ‘sympoiesis’[24], based on the idea of ‘autopoiesis’[25], in which living systems ‘create themselves’. this is also reminiscent of freud’s early notion of ‘polymorphous perversity’[26], which describes how an infant steers its habits and identities by making situated choices based on its personal experience of gratification or displeasure. these habits may later be guided by the values and responses of the society, which are also mediated by the framework of language 228 john wood:creativity & biodiversity: towards a synergy-of-synergies within which the infant will, in theory, be free to co-create, as a member of that society. it may also be that there insufficient creative optimism, or opportunism, within the metadesign team. this also has useful implications for the future of democracy. indeed, the work of donella meadows (1997) implies that the collective imagination is more powerful than the counting of individual votes. 12 how can we do it better? unless we decide to see our fate as a largely self-inflicted (and well-deserved) inevitability, we need some decisive action that will have the requisite effect on species decline. i have been arguing for many years that design professionals might be a helpful complement to the range of experts currently consulted by the most influential government agencies and corporations. this would mean commissioning designers to be servants of humanity, rather than working for the profit of a few. unfortunately, most are trained to be specialist mercenaries who apply their ‘creative’ visual intelligence to service the insatiable demands of a consumer-led economy. while designers have shown themselves to be indispensible to the commercial system, they are less influential for their thinking. western thought has tended to value logic and truth so many of the major professions are educated within a learning culture that tends to see knowledge as a set of truthclaims, backed-up by evidence. designers are different because they are trained to immerse themselves in the pleasurable immediacy of forms and images. but they also think in terms of anticipatory affordances (i.e. the optimistically hypothetical, the creatively provisional and the contingently possible). these not only include skills cultivated in schools and universities, but they also reflect a set of capabilities that evolved over the last million years, or so[27]. at the end of the 1960s, the scientist, herbert simon and the designer victor papanek[28] both made the claim that everyone is a designer. in simon’s version, we ‘design’ whenever we create a particular course of action that is likely to change an existing situation into a preferred one[29]. more recently, the sociologist, zigmunt bauman pushed the idea further by arguing that the similarity between design and management has existed since the late 1880s, when governments invited designers and managers to bring science and technology into everyday life[30]. recently, tim brown[31] has popularised what herbert simon[29] called ‘design thinking’, and what nigel cross called ‘designerly ways of knowing’[32]. before this, many non-designers had not fully appreciated the distinctiveness of these approaches. 13 design for survival according to john thackara (2005) the decisions we take at the design stage are the ‘cause’ of 80% of the environmental impact of the products. yet, historically, although designers are encouraged to explore alternative futures in a creative advances in systems science and applications (2012) vol.12 no.3 229 way, their status, as a ‘minor’, or relatively junior, profession has given them few opportunities to work at a more strategic level. this may help to explain why we are stuck in a paradigm of bad habits that perpetuates confused thinking. to be able to tackle highly complex problems, such as biodiversity losses, designers would need to rethink their traditional role as catalysts of the consumer society. rather than working as specialists, they would need to cultivate a more selfreflexive and comprehensive approach that enables them to intervene at many simultaneous points within a whole system. they might need to learn how to create synergies for all, rather than new products and services. the pivotal role of designers within the fashion industry highlights this, in quite an ironic way. in europe, the habits of the fashion industry cause around 10% of all the waste produced, yet the role assigned to designers makes them seem more like the problem than the solution. this is a systemic problem that cannot be blamed on one particular group or method. as mathilda tham, the trends forecaster, put it, ‘fashion thrives on innovation but resists change’[33]. in other words, although fashion designers may be capable of re-thinking long-term business futures, they are trained to ignore these possibilities, and to focus onto next season’s stylefutures. but, if designers were trained to envision long-term business futures, they would be able to deliver waste-free rewards for corporations. the ‘cradleto-cradle’ movement is an excellent example of this approach. by creating an adaptive, circular economy, rather than one that looks for unlimited innovation and growth, it will be possible to design complex synergies of combination, rather than offering short-life products that cost additional money for their dispose. fig.3 some of the co-sustaining paradigms that maintain the status quo. 230 john wood:creativity & biodiversity: towards a synergy-of-synergies 14 re-designing paradigms what is stopping the implementation of systems that will produce longer-term outcomes? to a great extent, the obstacles to change are caused by the poor economic thinking behind technological development. if designers are not paid to look beyond their short-term cycle of design, production and re-design they cannot take full responsibility for what happens in the longer-term. one way to reconcile design thinking with ecological thinking is to invent a systemic discourse within which complex systems are seen as ‘paradigms’[34], because they resemble ecosystems. they also contain both animate, and inanimate agencies, similar to james lovelock’s gaia theory[35]. and these paradigms sustain themselves by attracting new elements that serve to support them. while designers tend to be taught how to (re)design existing products, we also need more radical solutions at the paradigmatic level. for example, although architects are trained to come up with a variety of iterations to high-rise offices, few are able to re-think the underlying paradigm of concrete, steel and glass, even though these materials have an inordinately large carbon footprint. the high-rise office is ubiquitous to london, new york or beijing. like all paradigms, it resists change because it is sustained by an interconnected array of subsidiary systems (also paradigms, or quasi-species), such as insurance protocols, commercial habits, technological habits, economic assumptions each of which appears to depend on it for their own survival[36]. 15 design for biodiversity in order to encourage designers to achieve better levels of ‘co-sustainment’, we would need to develop a more ‘ecomimetic’ approach[37]. ecomimicry is significantly different from what is known as ‘biomimicry’[38]. although ‘biomimicry’ was defined within an admirably broad philosophy, the way that designers tend to apply it seems disappointingly narrow. perhaps this reflects the prevailing business mindset, in which projects are seldom seen as part of a circular[39], or long-term economic vision. even today, it is hard to find examples of biomimetic innovation that do much more than behave as discrete technological ‘fixes’, gadgets or products. instead of seeing the designer’s role as the fixer of links in a chain of transactions, ‘ecomimetic designers’ would be expected to design the “conditions” that support interdependencies within whole systems. this is not a simple, or trivial shift. for example, the massive complexity and ineffability of these interdependencies mean that we can no longer expect to work predictively, as we might have done with, say, product design. many other aspects of design, as we know it, will also need to change, therefore we will refer to the new approach as metadesign. one of the radical changes we need to make is to re-purpose what we understand as ‘creativity’ and acquire a wiser understanding advances in systems science and applications (2012) vol.12 no.3 231 of relations, rather than separate ‘things’. 16 re-inventing creativity integrating hitherto disconnected methodologies is useful, but it is not enough. we also need to modify the assumptions that motivate designers, especially where ‘creativity’ is concerned. this means analysing and, where necessary, challenging the received assumptions behind this term. in the last decade or two, we have seen particular models of creativity as tools for regenerating urban communities [40], for stimulating economic growth, or for transforming businesses to make the economy more ‘efficient’[41]. here, ‘creatives’, such as designers, have come to be seen as part of a dependable toolbox that can bring success to prestigious real estate deals, national re-branding exercises, or other tourist attractions, such as the hosting of an olympic games events. how might we think of ‘creativity’ if the main aim is to generate wellbeing and prosperity by achieving optimum biodiversity? i believe we must develop creativity’s capacity to help us to ‘adapt’ better to our changing habitat. this may include applying it, less to the defeat of existing, or rival, plans but to look for new ‘synergies’ that will deliver additional benefits that emerge from combining existing things. 17 the problem of ‘creative genius’ one of the problems with the modern idea of ‘creativity’ is that it became part of the culture of competitive advantage, rather than a way to adapt to new situations. often, one finds ecological terminology applied, shamelessly, to an economic context whose autonomy is seldom framed within the bigger picture. and it is increasingly associated with destructive terms, such as ‘disruptive innovation’, or ‘disruptive technologies’. these tendencies remind one of the aggressive, pitiless language of francis bacon, whose crude scientific methodologies are perpetuated in some global corporations, whose products continue to threaten the survival of certain key species. the same, swaggering, solipsistic stance can also be noted from books on creativity and innovation, such as ‘ignore everybody’ or ‘relentless innovation’. redefining creativity is not a simple step. the destructive and egoistic models that we inherited from the romantics remain part of the popular culture. after the enlightenment it came to valorize individual power, originality, and to be blunt arrogance. david hume (1711-1776) and others, such as arthur schopenhauer (1788-1860), strongly admired the notion of ‘genius’ as someone who is so extraordinarily self-styled and unfathomable that he would struggle to adapt to the ‘normal’ world around him. this also echoes nietzsche’s idea of ‘der übermensch’, which developed in the 1880’s. this theory asserted that homo sapiens has the potential to create a new order, provided there is a sufficient will to power, and a readiness to reject societal ideals and moral codes. 232 john wood:creativity & biodiversity: towards a synergy-of-synergies in the 20th century, inspired by psychoanalysis and by concepts such as ‘positive self-regard’, ‘presentation of self’, and ‘self-actualisation’ creativity came to be associated with the right to personal self-expression. 18 unreasonable men in a sense, then, the modern sense of the word ‘creativity’ may mean a cynical, or arrogant refusal to adapt to anything. perhaps this is what the architect, frank gehry, meant when he said, in 2005, “i don’t do context”. today, the invitation to ‘be unreasonable’, or to ‘think different’ (presumably, a fusion of ‘be different’ and ‘think differently’) has become familiar to today’s consumers. indeed, the apple corporation have used it many times to advertise their computers. however, sustaining and enhancing biodiversity will require us to reflect more deeply on the shareable and distributed nature of creativity and how can help us adapt to our habitat. our stridently humanistic idea of creative may be traced back several thousand years to horace, who inspired kant’s famous phrase ‘dare to know’ (1784). and, while it had encouraged a long development of careful reason and inquiry, this was not what we would understand as ‘creativity’, as we understand it today. by the 17th century the idea of ‘daring to know’ had inspired john locke’s radical insight (1689) that “the mind can furnish the understanding with ideas”. while the notion of ‘creative thinking’ may now be commonplace for agnostics in the 21st century, it had remained virtually unthinkable to philosophers before hume and locke. the idea that an individual can choose to think what s/he wants to, later acquired a powerful framework of thinking, with coleridge’s term ‘self-consciousness’, as celebrated by a series of narcissistic ‘genius’ figures who dominated the post-romantic era in art, literature and music. 19 can we find creativity in nature? in 1877, ten years before the first sherlock holmes publications made it popular, charles peirce announced his idea of ‘abductive reasoning’. this turned deductive argumentation on its head, by enabling thinkers to begin the process with what might, hitherto, have seemed like a conclusion. while we may see this type of logic as a characteristically human, or even modern, mode of thought, gregory bateson has suggested that abductive reasoning is part the natural order: “all thought would be totally impossible in a universe in which abduction was not expectable......”[42]. he believed that evolution is responsible for the parallels between the way we think, and the way nature works. arthur koestler’s ‘bisociation’ method is interesting in this respect, as it is designed to elicit a new idea when two things are combined[43]. this resembles sexual reproduction in that different ‘parent’ ideas are brought together to create a new hybrid outcome. when the two ideas (or creatures) creatures, ‘a’ and ‘b’, are combined we may advances in systems science and applications (2012) vol.12 no.3 233 find that we also have a third idea ‘c’. abductive reasoning is a way to ‘reverse engineer’ evolutionary logic. instead of taking the 2 parents and seeing what type of child we will get, we start with the child (‘c’) and look for an ‘a’ and a ‘b’ that might have been its parents. one lesson we can learn from evolution is that creativity is only important when it works to produce new opportunities. and one reason why the romantic era created so much interest and excitement is because it was less concerned with the strict languages of scientific ‘truth’ and more to do with enriching the semantic discourse that may, subsequently, become applicable to science. however, unlike scientific ‘truths’, artistic propositions are usually hard to fathom, without reflecting upon their aesthetic status. 20 the role of aesthetics it will be important to harness aesthetic criteria when changing the paradigm, because aesthetics acts as an interface between thoughts and actions. aesthetic discourse can change consensus because we share and modify our perceptions when we discuss personal tastes. arguably, beauty tells us what is good, or, rather, what used to be good. aesthetics is not only a philosophical discourse, it is also a multi-dimensional field of awareness, so it cannot be quantified using a small number of dimensions. claims to beauty exist as combinations of elements, or as patterns of sensory awareness, that are deemed to work together. the fact that we have fads in food tastes, or fashions in clothing, is evidence that aesthetic judgements are seldom fixed by a rigid biological ‘need’, or aesthetic code. however, it seems likely that certain predilections and habits are strongly ‘wired’ into the human body, especially when they remind it of tastes, flavours, or colours that have proved beneficial to us over many hundreds, or thousands, of years. this is useful to designers (and to artists), because creating new tastes and aesthetic forms will change appetites & habits. a good example of this is the pleasure that we take when they see varieties of the rosaceae (rose) plant family. humans have enjoyed products of the rose family for many thousands of years. they have give us many varieties of nutritious fruits, such as apples, apricots, plums, cherries, peaches, pears, raspberries, and strawberries. during this time we have made them bigger and sweeter, thus creating a strong level of co-sustainment between people and the plants themselves. in my opinion it is not a meaningless coincidence that the largest and most successful brand in the world is an ‘apple’. 21 re-languaging new realities in previous papers i have argued that designers have, traditionally, overlooked the importance of language in ‘design thinking’. this is not meant to imply that designers should try to think in a more ‘critical’ way, but that they could be more 234 john wood:creativity & biodiversity: towards a synergy-of-synergies creative with the terms that we use to describe things, experiences, values and ideas. how many colours are there in a rainbow? even though science tells us that it contains countless wavelengths of light, most people answer with a small number (usually 7) they learned at school. but how does this answer affect our perceptions? our reality is lived out in metaphors, adjectives, images and categories. these, in turn, shape our beliefs, actions and assumptions. if one language has more words for flavours, and for colours, than another, it seems likely that the speaker of this language will have a bigger horizon of experience. the pioneer eco-semiotician, jakob von uexküll (1864-1944), used the term ‘umwelt’ to describe the phenomenological ‘reality’ of different creatures. this is an important idea, because no two creatures experience the world in quite the same way. even though symbiotic relationships may develop within the affordances of an existing language system (i.e. what maturana & varela call ‘structural coupling’) there may be a very tenuous overlap between the two sides of the ‘conversation’ for example, one may be able to see but not to hear, the other may have the opposite capabilities. this suggests that, by augmenting the naming system we have, we can expand our umwelt, perhaps in order to re-map vested interests, empathetically, in conjunction with other species. hypothetically speaking, redefining the agreed purpose of a paradigm should make it easy to change[5]. however, social inertial is co-sustained by the language and customs of the old paradigm, and these tend to mask opportunities, making them seem difficult, unthinkable or impossible. crudely speaking, the answer is to create new words that would afford new concepts and, perhaps, serve to facilitate new ecological paradigms that enable the homo sapiens to survive a little longer. references [1] backwell j. and wood j. (2009), “catalysing network consciousness in leaderless groups: a metadesign tool”, in consciousness reframed 12, art, identity and the technology of the transformation, editors roy ascott & luis miguel girão, university of aveiro, portugal, pp. 36-41. [2] ranulph glanville. (2004), “the purpose of second-order cybernetics”, kybernetes, vol.33, no.9, pp.1379-1386. [3] wood j. (2007), “design for micro-utopias: thinking beyond the possible”, ashgate: uk. [4] wood j. (2007), “synergy city, planning for a high density, super-symbiotic society”, (editors p. jones, h. jones), landscape and urban planning, vol.83, no.1, pp.77-83. advances in systems science and applications (2012) vol.12 no.3 235 [5] meadows d. (1999), “leverage points: places to intervene in a system”, the sustainability institute, http://www.sustainer.org/pubs/leverage points.pd f, accessed 14 august, 2009. [6] gross m. and williams n. (2010), “missed targets”, in current biology, vol.20, no.12, pp.496-497. [7] valiverronen e. and hellsten i. (2002), “from ‘burning library’ to ‘green medicine’ the role of metaphors in communicating biodiversity”, science communication, vol.24, no.2, pp.229-245. [8] mora c., tittensor d.p., adl, s., simpson, a.g.b. and worm. b. (2011), “how many species are there on earth and in the ocean?”, plos biol vol.9, no.8: e1001127. doi: 10.1371/journal.pbio.1001127. [9] brown p. (2011), guardian articles, http://www.guardian.co.uk/environme nt/2011/sep/30/uk-cod-collapse-overfishing? accessed 30 september 2011. [10] hardin g. (1977), “the tragedy of the commons. in garrett hardin and john baden (eds.)”, managing the commons, w.h. freeman and co, san francisco. [11] corning p. (1983), the synergism hypothesis, institute for the study of complex systems, palo alto. [12] fuller, richard buckminster. (1969), operating manual for spaceship earth, southern illinois university press: carbondale, il. [13] diaconis p. and mosteller f. (1989), “methods of studying coincidences”, journal of american statistical association, vol.84, pp.853-861 (cited in weisstein, eric w, law of truly large numbers from mathworld-a wolfram web resource). [14] harrop, s. (2011), “living in harmony with nature? outcomes of the 2010 nagoya conference of the convention on biological diversity”, in journal of environmental law, vol.23, no.1, pp.117-128. [15] magee e. (2007), food synergy, unleash hundreds of powerful healing food combinations to fight disease and live well, rodale, new york. [16] whorf b.l. (1956), language, thought and reality: selected writings of benjamin lee whorf, mit press: cambridge, ma. [17] lakoff g. and johnson m. (1980), metaphors we live by, university of chicago: chicago and london. 236 john wood:creativity & biodiversity: towards a synergy-of-synergies [18] wood j. (2011), “languaging change from within: can we metadesign biodiversity?”, in journal of science and innovation, vol.1, no.3, pp.27-32. [19] wackernagel m. and rees w.e. (1996), our ecological footprint, reducing human impact on the earth, philadelphia: new society publishers. [20] ashby w.r. (1956), introduction to cybernetics, chapman & hall, london. [21] maturana h. and varela f. (1992), the tree of knowledge: biological roots of understanding, shambhala: boston. [22] brundtland e. (1987), “our common future”, report of the world commission on environment and development, available at: http://www.undocuments.net/wced-ocf.htm (accessed 18th march 2012). [23] wood, j. (2002), “(un)managing the butterfly: co-sustainment and the grammar of self”, in international review of sociology: revue internationale de sociologie, vol.12, no.2, 2002, p.1. [24] wood j. and nieuwenhuijze o. (2005), “synergy and sympoiesis in the writing of joint papers; anticipation with/in imagination”, paper presented at seventh international conference on computing anticipatory systems, hec management school, university of liege, liège, belgium, 8-13 august. [25] maturana h., varela f. (1980), “autopoiesis and cognitionthe realisation of the living”, in boston studies in philosophy of science, reidel: boston. [26] freud s. (1991), “on sexuality: three essays on the theory of sexuality and other works”, edited by angela richards, 1991, penguin,harmondsworth. [27] dilnot, clive. (2008), “the critical in design (part one)”, journal of writing in creative practice, 1.2 pp. 177-189. [28] papanek v. and fuller, r.b. (1972), design for the real world, thames and hudson. [29] simon h. (1969), the sciences of the artificial, 3rd edition, mit press: cambridge, ma. [30] bauman z. (2006), “design, ethics and humanism, cumulus conference”, warsaw. nantes, france 5-17 june 2006. [31] brown t. (2009), change by design: how design thinking creates new alternatives for business and society: how design thinking can transform organizations and inspire innovation, harper collins, ny. advances in systems science and applications (2012) vol.12 no.3 237 [32] cross n. (2010), designerly ways of knowing, springer-verlag: london. [33] tham m. and h. jones. (2008), “metadesign tools: designing the seeds for shared processes of change”, 10-12 july, 2008 at the ‘changing the change’ conference, turin, italy. [34] wood j. (2012), “in the cultivation of research excellence, is rigour a no-brainer?”, journal of writing in creative practice, 5:1, pp.11-26, doi:10.1386/jwcp.5.1.11 1. [35] lovelock j. (1979), gaia: a new look at life on earth, oxford university press: oxford. [36] kauffman s. (1995), at home in the universe, the search for the laws of self-organization, oxford university press: new york. [37] fairclough k, ed. jones, h. ecozen . (2005), in agents of change: a decade of ma design futures, goldsmiths college: london, pp.42. [38] benyus j. (1997), innovation inspired by nature: biomimicry, william morrow and co.: new york. [39] mcdonough w. and braungart m. (2002), cradle to cradle: remaking the way we make things, north point press: new york. [40] landry, c. (2000), the creative city: a toolkit for urban innovators, earthscan ltd, uk & usa, isbn-10:1853836133, isbn-13:978-1853836138. [41] cox g. and dayan z. (2005), cox review of creativity in business: building on the uk’s strengths, tso: norwich, chapter 1. [42] bateson g. (1980), mind and nature: a necessary unity, new york: bantam books. [43] koestler a. (1967), the ghost in the machine, penguin: london (reprint 1990). corresponding author john wood can be contacted at: maxripple@gmail.com advances in systems science and application (2015) vol.15 no.4 316-325 computer investigation of a game theoretic model of social partnership in the system of continuing education vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko southern federal university, russia. abstract a game-theoretic model of social partnership in the system of continuing education is proposed and investigated in the simulation mode. results of the model identification and investigation based on simulation modeling are considered. a comparative analysis of egoistic and cooperative approaches to the social partnership is conducted. keywords social partnership; continuing education; simulation modeling; identification; difference games. 1 introduction a system of higher education should be flexible enough to prepare qualified and competitive specialists capable to improve their knowledge and skills in the changing environment. one of the important mechanisms providing the solution of this problem is social partnership that allows for the cooperation of employers, universities, and students. social partnership relations are very important in the system of higher education [1,2]. an overview of the papers concerned with mathematical modeling of the problems of social partnership is proposed in some previous studies of the authors [3]. social partnership in continuing education is a specific system of joint activities of the education system agents characterized by trust, common objectives and values, and providing highly qualified, competitive, and mobile specialists for the labor market. the main research hypothesis is that social partnership permits to increase a level of professional competence of the specialists. it seems natural to use the formalism of differential games [4] for description of social partnership relationships. due to the high complexity of the differential game model the techniques of simulation modeling [5] are applied for its investigation in the difference form. the paper develops authors approach exposed in [3]. in the section 2 a general description of the model is given and the new moments in comparison with the previous studies [3] are shown. in the section 3 the model identification is described. in the section 4 planning and implementation of the computer simulation experiments are discussed. section 5 deals with processing and analysis of the modeling results. section 6 concludes. advances in systems science and application (2015) vol.15 no.4 317 2 a general description of the model in the proposed model social partnership relations among employers, students, and university are considered. the payoff functions are as follows: jp = t∑ t=0 gp (up (t), ub(t), uc(t), x(t) → max, up (t) ∈ up ; jb = t∑ t=0 gb(up (t), ub(t), uc(t), x(t) → max, ub(t) ∈ ub; jc = t∑ t=0 gc(up (t), ub(t), uc(t), x(t) → max, uc(t) ∈ uc ; where n = {p, b, c} is a set of players, namely: p – employer; b – university; c – student. the systems dynamics are given by the equation x(t+ 1) = x(t) + f(x(t), up (t), ub(t), uc(t)), x(0) = x0; here up (t), ub(t), uc(t) are strategies of the players describing their efforts directed to the development of social partnership relations; up , ub, uc – domains of feasible strategies; jp , jb, jc – payoff functionals of the players; gp , gb, gc – instantaneous payoff functions; p= {p1, ... ,pr} – a finite set of employers; b ={b1, ... ,bm} – a finite set of universities; c = {c1, ... ,cs} – a finite set of students; t = 4 (period = 5 years). the strategies determine a share of the annual budget assigned by a player to the needs of continuing professional education (cpe): upi(t) – a share of the annual budget assigned to cpe programs by an employer, upi = [0,rp ]; ucj (t) – a share of the annual budget assigned to cpe programs by a student, ucj = [0,rc ]; ub(t) – a share of the annual budget assigned to cpe programs by the university, ub = [0,rb]; rp , rc , rb – the annual budgets of the players. the main contribution of this paper in comparison with [3] is that the payoff functionals are taken in the form jp = t∑ t=0 e−ξt[gp (rp − up (t)) +gp (x(t))] → max, up (t) ∈ up ; jb = t∑ t=0 e−ξt[gb(rb − ub(t)) +gb(x(t))] → max, ub(t) ∈ ub; jc = t∑ t=0 e−ξt[gc(rc − uc(t)) +gc(x(t))] → max, uc(t) ∈ uc ; 318 vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko:computer ... this idea was firstly proposed in [6] and means that each player shares her resources between common (ui) and private (ri ui) interests, i=p, b, c. the first summand describes the private interests, and the second one describes the common interest. in this model it is supposed that the common interest consists in the development of the cpe system and each player gains her own gain gi of it. the private interests of the players are different and describe their payoffs from investments beyond the cpe system. discounting is also used, and the random variable ξ reflects a state of the social environment. in the current investigation ten university departments, five employers and ten students are considered. the strategies of generalized players in any moment of time are calculated as an arithmetic mean of the strategies of those agents, namely: strategy of p (employer): up (t) = 1 r r∑ i=1 upi(t);up (t) ∈ up , r = 5; strategy of c (student): uc(t) = 1 s s∑ j=1 ucj (t);uc(t) ∈ uc , s = 10; strategy of b (university): ub(t) = 1 z z∑ k=1 ubk (t);ub(t) ∈ ub, z = 10; the state variable of the model considered as a time function x(t) is also specified in comparison with [3] and now characterizes a level of professional training of the student; f – a function of the system dynamics depending on the players’ strategies. it is assumed that the function of system dynamics increases in respect of all arguments (the efforts of players positively influence to the results of social partnership). to give the system dynamics the modified verhulst-pearl model is used, i.e. f is taken in the form f(x(t), up (t), ub(t), uc(t)) = h(up (t), ub(t), uc(t))x(t)(1− x(t) k ); where k – maximal feasible value of the state variable in the given conditions; h – function of impact of the players’ strategies, namely: h(up (t), ub(t), uc(t)) = 3∑ i=1 aiui(t); ai ≥ 0; 3∑ i=1 ai = 1; i = p,b,c; advances in systems science and application (2015) vol.15 no.4 319 ai – relative weights of the strategies. for the estimation of the relative weights the following considerations are used. the sum of the weights is equal to 1, and all weights are positive. the most important influence is made by students as key agents of the cpe system. the two other weights are approximately equal to each other and don’t differ from the former weight (students) too significantly. so, the weights are chosen as follows (table 1). table 1 relative weights of the strategies (ai) relative weight i employer university student ai 0.3 0.3 0.4 3 identification of the model suppose that the state variable x(t) characterizes a level of professional training of the students. for the estimation of the initial training level a known approach proposed by donald kirkpatrick [7] is used. d. kirkpatrick has developed a model of evaluation of the training effectiveness which considers four levels: 1) reaction: to what degree participants react favorably to the training. for the estimation the results of sociological polls conducted among pce students of the southern federal university are used. the sample consisted of 2,110 respondents, including 827 female (39%) and 1283 male (61%). 2) learning: to what degree participants acquire the intended knowledge, skills, attitudes, confidence and commitment based on their participation in a training event. here the data about the students’ progress together with data characterizing the material base of education are considered. 3) behavior: to what degree participants apply what they learned during training in their professional practice. the results of sociological polls conducted among employers in the rostov region are used. the sample consisted of 2,580 respondents, including 785 female (30, 4%) and 1795 male (69, 6%). 4) results: to what degree targeted outcomes occur as a result of the training event and subsequent reinforcement. the data of polls among employers on the topic “a level of professional knowledge and skills of the graduates” are used. as a result, an initial value of the professional training level is evaluated as x0 in the scale [0, 10]. also, the maximal feasible value of the state variable in the given conditions is taken equal to k = 10. the identification of payoff functions was conducted as follows. let gi be the i’s player payoff from the investments which are not concerned with cpe. the most important type are bank deposits; suppose that vklad i is a 320 vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko:computer ... dividend function: v kladi = αi · (ri − ui(t)) · st; i = p,b,c; αi –a share of the investments; st – an average annual percent rate. use also a function ini, i=p,b,c , that determines some fixed payoffs or income of a university professor from other sources: gb(rb − ub(t)) = v kladb +rpt+ inb; rpt – a gain from additional training services: rpt = time · (stavka− β · (rb − ub(t))); time – a total amount of lectures; stavka – an average price of a lecture; β – a share of expenditures required for a lecture; gp (rp − up (t)) = v kladp + ent+ inp ; ent – a function of income business (for an employer): ent = ∑ j productj · (pricej − γj · (rp − up (t))); product j – an amount of the good j ; pricej – a price of the good j ; γj – expense for purchasing and storage of the good j ; gc(rc − uc(t)) = v kladc + earn+ inc ; earn – an income function for student: earn = ∑ k (servk · valuek)− ε · (rc − uc(t)); servk – a volume of services; valuek – an average price of a service; ε – a share of the expenditures. farther,gi – a payoff of each player from increase of the students level of professional training: gi(x(t)) =(x(t) − x(t − 1)) · ( ∑lengthmati k=1 matik · pricematik + ∑lengthmati j=1 nemi j · pricenemi j), i = p,b,c; advances in systems science and application (2015) vol.15 no.4 321 matk i – a set of evaluations of the material payoffs of the player i from cpe; nemj i – a set of evaluations of the non-material payoffs of the player i from cpe; the evaluations are made in the segment [0, 10]. pricematk i – a price of one point of the material payoff of the player i ; pricenemj i – a price of one point of the non-material payoff of the player i ; lengthmati – a number of evaluations of the material payoff of the player i ; lengthnemi – a number of evaluations of the non-material payoff of the player i ; material and non-material payoffs for different players are represented by different factors and determine the characteristics and types of the payoffs. a preliminary expert estimation of the points is made. a maximal possible value of gi for each player is found. a logically substantiated relation between material and non-material payoffs is deduced. 4 planning and implementation of the computer simulation experiments the model investigation was conducted by computer simulation on the base of scenario method [5]. the scenarios are formed according to plausible behavior patterns of the players. for simplicity the arithmetic mean strategies are described. it is supposed that for all scenarios for any moment of time the values of strategies are equal: up (t) = uc(t) = ub(t) = uo(t), t = 0, ... ,4. six scenarios are considered: 1) maximal (max) one corresponds to the maximal possible financing when the whole budget of a player is assigned for training: uo(t) = ri; 2) medium (med) one assigns a half of the budget for training: uo(t) = 0.5ri; 3) minimal (min) allows for training only a small part of the budget: uo(t) = 0,2ri; 4) absence of financing (abs) is clear: uo(t) = 0; 5) decreasing of financing (dec) means that initially an eminent part of the players budget is assigned for training but then the share decreases to a small value, namely: uo(t) = (0.8 0.15t)ri; 6) increasing of financing (inc) describes the opposite strategies: uo(t) = (0.2 + 0.15t)ri. thus, in the former four scenarios the strategies are constant in time, while in the fifth and sixth scenarios the shares of financing are time-dependent. 5 processing and analysis of the modeling results the processing of modeling data includes a comparative analysis of the graphs of state variable and payoff functionals for different scenarios. for example, in fig.1 the graphs of all payoff functions for the scenario of maximal financing are shown. in contrary, in fig.2, the graphs for students payoffs for all scenarios are 322 vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko:computer ... given. fig. 1 a comparison of the graphs of payoff functions of employer (jp(t)), student (jc(t)), and university (jb(t)) for the scenario of maximal financing fig. 2 a comparison of the graphs of the students payoff function for different scenarios advances in systems science and application (2015) vol.15 no.4 323 the graphs of payoff functions show a moderate growth considering discounting. instead of abstract utilities in [3], in this paper financial estimations of the payoffs were made. the comparative analysis has shown that the biggest payoff goes to employer, a less to student and the smallest to university. this result is new and was absent in [3]. preliminary analysis of compensation of the investments has shown that in the first period only 71 % of the expenses of student, 63% of the expenses of employer and 18% of the expenses of university are returned. considering the future investments, only the expenses of student and employer for two scenarios: minimal one and increasing of financing – are compensated at all. probably, a more precise identification of the model parameters is needed. the received values of payoffs are very sensitive to the amounts of annual budgets of the players. as test data were used for calculations, the results are still qualitative. nevertheless, the investigation is very important for debugging of the new model and preparation of the special sociological research. more precise values are expected to be received after this research. the results about the impact of strategies on the payoff functions are similar to the previous ones [3]: the best results are provided by the maximal financing scenario, and the worst ones – for the case of no financing. a comparison of the level of professional training for different scenario is made in fig.3 (for periods). fig. 3 comparison of the level of professional training for different scenarios this graph is qualitatively similar to its counterpart in [3]. accordingly to the new questionnaires, the players will be able to estimate the changes in the level of professional training themselves. 324 vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko:computer ... 6 conclusion a new version of the model of social partnership in continuing education system is considered. the payoff functions were specified as sums of private and common payoffs, where common payoff depends on the level of professional training of students. the transitive model has determined a set of questions to be answered by a special sociological research among students, professors, and employers. the next model version will be implemented after the research. the current version of the model shows that even a small investment to the development of continuing education system gives a certain gain to its participants. a detailed analysis should be made for the problem of compensation: not all proposed scenario are substantiated from the point of view of investments. for the university the compensation is problematic. in the model the players receive the maximal payoffs if the financing is maximal. it is evident that to increase the level of professional training is possible only if investments are made. in general, the model emphasizes the necessity of development of the continuing education system by all its participants. acknowledgements the work is supported by the russian foundation for humanities, project # 14-03-00236. references [1] keith n. (2011), engaging in social partnerships: a professional guide for successful collaboration in higher education, routledge, pp.288. [2] siegel d. (2010), organizing for social partnerships: higher education in cross-sector collaboration, routledge, pp.224. [3] dyachenko v.k., ougolnitsky g.a. and tarasenko l.v. (2014), “computer simulation of social partnership in the system of continuing professional education”, system science and application, vol.14, no.4, pp.27-37. [4] dockner e., jorgensen s., long n.v. and sorger g. (2000), differential games in economics and management science, cambridge university press,pp.382. [5] law a.m. and kelton w.d. (2000), simulation modeling and analysis, mcgraw-hill, pp.784. [6] germeyer yu.b. and vatel i.a. (1974), “games with hierarchical vector of interests”, j journal of computer systems international. advances in systems science and application (2015) vol.15 no.4 325 [7] kirkpatrick d.l. and kirkpatrick j.d. (2007), implementing the four levels: a practical guide for effective evolution of training programs, berrett koeler publishers, pp.153. corresponding author guennady a. ougolnitsky can be contacted at: ougoln@gmail.com advances in systems science and applications (2014) vol.14 no.4 396-403 stability of solution maps for a η-parameter weak vector variational inequality changchang bu, yuqiang feng and hui li school of science wuhan university of science and technology wuhan 430065 china abstract in this paper, we introduce a new class of η-parameter weak vector variational inequality (for short, η-pwvvi) in banach space, which extends the existing parameter weak vector variational inequality. we use the concepts of η(y, x) function, invex set, η-hemicontinuous and η-strongly c pseudomonotone mapping to study (η-pwvvi) and we obtain new η-generalized linearization lemma. the stability of solution maps for (η-pwvvi) is obtained by this lemma. finally, we present an example to illustrate our results. keywords: η-parameter weak vector variational inequality, η-generalized linearization lemma, invex set, η-hemicontinuous, η-strongly c pseudomonotone. 1 introduction as very powerful and important tools in the study of nonlinear sciences, variational inequalities and vector optimization have attracted so much attention. over the last decades, variational inequality and vector optimization techniques have been applied extensively in such diverse fields as biology, chemistry, economics, engineering, game theory, management science and physics. in 1980, f. giannessi [1] introduced a well-known inequality in finite dimensional space, which was called vector variational inequality (for short,vvi). thereafter, g. y. chen and g. m. chen [2] discussed this problem in infinite dimensional space. g.y. chen [3] presented a new class of vector variational inequality problem in real banach space. the new vector variational inequality problem not only extended the classical vector variational inequality problem, but also related to the existence of non-dominated solution for vector optimization problem. inspired by g. y. chen’s result, a. h. siddiqi et al. [4] and k. l. lin et al. [5] introduced and researched general vector variational inequality problem. l. n. wang et al. [6] proved the stability of solution maps for a weak vector variational inequality in 2013. it is worth pointing out that the convexity plays a significant role while studying the continuity of the solution for vector variational inequality. the concepts of invex set and generalized convexity have been given by s. r. mohan, s. k. neogy [7] and x. m. yang [8] respectively. motivated by the work reported in [1]-[8], the aim of this paper is to introduce a new class of η-parameter weak vector variational inequality (for short, η-pwvvi) in banach space, which extends the existing parameter weak vector variational advances in systems science and applications (2014) vol.14 no.4 397 inequality. our results unify, generalize and complement various known comparable results from the current literature. the rest of the paper is organized as follows. in sect.2, we recall some basic definitions and notations which will be used in the sequel. in sect.3, we use the concepts of η(y;x) function, invex set, η-hemicontinuous and η-strongly c pseudomonotone mappings to study (η-pwvvi) and we obtain new η-generalized linearization lemma. as a consequence, the stability of solution maps for (ηpwvvi) is obtained in theorem 3.3. finally, we present an example to illustrate our results in sect.4. 2 preliminaries let x, y and w (parameter space) are banach s-paces, c ⊆ y is a non-empty closed convex cone with int c ̸= ∅. l(x; y) denotes the space which consist of all the continuous linear operators, define the value of linear operator t ∈ l(x, y) at x ∈ x by ⟨t, x⟩. throughout this paper, assume that η(y, x) : x × x → x satisfies all the conditions as follows: (c1) η(x, x+ λη(y, x)) = −λη(y, x); (c2) η(x, y) + η(y, x) = θ; (c3) η(y, ·) is continuous. here, θ denotes the zero element of x. consider the following η-weak vector variational inequality problem (for short, η-wvvi) of finding x ∈ k such that ⟨t (x), η(y, x)⟩ /∈ −int c, ∀y ∈ k where k ⊆ x is non-empty, t : x → l(x,y ) is a vector value function. when the operator t perturbed by the parameter µ with µ ∈ λ ⊆ w and λ is non-empty, for fixed µ, we deal with the following η-parameter weak vector variational inequality problem (η-pwvvi) of finding x ∈ k such that ⟨t (x, µ), η(y, x)⟩ /∈ −int c, ∀y ∈ k where k ⊆ x is non-empty, t : x×λ → l(x,y ) is a vector value bifunction. for any µ ∈ λ, sη(µ) denotes the solution set of (η-pwvvi), that is, sη(µ) = {x ∈ k | ⟨t (x, µ), η(y, x)⟩ /∈ −int c, ∀y ∈ k} in this paper, we assume that for any µ ∈ λ, sη(µ) is non-empty. now, we give some basic definitions and some properties needed in the following 398 changchang bu: stability of solution maps for a -parameter weak vector ... sections. definition 2.1. (see [7]) a set k ⊆ x is said to be invex with respect to a given η(y, x) : x ×x → x if ∀x, y ∈ k,λ ∈ [0, 1] ⇒ x+ λη(y, x) ∈ k. definition 2.2. let k ⊆ x and k is invex with respect to η(y, x), the operator t : k → l(x,y ) is said to be η-hemicontinuous if and only if for any x, y ∈ k,λ ∈ [0, 1], the mapping λ → ⟨t (x + λη(y, x)), η(y, x)⟩ is continuous at 0+. definition 2.3. let k ⊆ x and k is invex with respect to η(y, x), the operator t : k → l(x,y ) is said to be η-weakly c pseudomonotone on k if for any x, y ∈ k, ⟨t (x), η(y, x)⟩ /∈ −intc implies ⟨t (y), η(y, x)⟩ /∈ −intc. definition 2.4. let k ⊆ x and k is invex with respect to η(y, x), the operator t : k → l(x,y ) is said to be η-strongly c pseudomonotone on k if there exists λ > 0 such that for any x, y ∈ k, ⟨t (x), η(y, x)⟩ /∈ −intc implies ⟨t (y), η(y, x)⟨+λ ∥ η(y, x) ∥2 by ∈ c, where by denotes the unit, closed ball in y . remark 2.1 if we take η(y, x) = y − x in definition 2.1, then invex set run into convex set. similarly, in definition 2.2, η-hemicontinuous reduce to νhemicontinuous (see [6]) with η(y, x) = y − x. remark 2.2 it is evident from definition 2.3 and definition 2.4 that if an operator is η-strongly c pseudomonotone, then it is η-weakly c pseudomonotone. 3 main results the following lemma 3.1 (η-generalized linearization lemma) extends the generalized linearization lemma (see [3]) through the concepts of η(y, x) function and invex set. lemma 3.1. (η-generalized linearization lemma) let k ⊆ x and k is invex with respect to η(y, x). moreover, assume the operator t : k → l(x,y ) is η-weakly c pseudomonotone and η-hemicontinuous, then the following two problems (i) and (ii) are equivalent: (i) there exists x ∈ k, such that for any y ∈ k, ⟨t (x), η(y, x)⟩ /∈ −intc; (ii) there exists x ∈ k, such that for any y ∈ k, ⟨t (y), η(y, x)⟩ /∈ −intc; proof. in view of definition 2.3, it is obvious that (i) implies (ii). now, assume that (ii) holds, then there exists x ∈ k, such that for any y0 ∈ k, ⟨t (y0), η(y0, x)⟩ /∈ −intc (1) note that k is invex with respect to η(y, x), thus for any x, y ∈ k,λ ∈ [0, 1], x+ λη(y, x) ∈ k (2) advances in systems science and applications (2014) vol.14 no.4 399 by (2) we can take y0 = x+ λη(y, x) ∈ k and combine the result with (1), we have ⟨t (x+ λη(y, x)), η(x+ λη(y, x), x)⟩ /∈ −intc (3) in view of the conditions (c1) and (c2) of η(y, x), we obtain that η(x+ λη(y, x), x) = −η(x, x+ λη(y, x)) = λη(y, x) thus, we claim that (3) is equivalent to ⟨t (x+ λη(y, x)), λη(y, x)⟩ /∈ −intc dividing by λ, we get ⟨t (x+ λη(y, x)), η(y, x)⟩ /∈ −intc let λ → 0+, take into account the η-hemicontinuity of t , we have ⟨t (x), η(y, x)⟩ /∈ −intc therefore, (i) holds. the proof is complete. lemma 3.2. let k ⊆ x and k is invex with respect to η(y, x). if for fixed µ ∈ λ, t (·, µ) is η-strongly c pseudomonotone, then the solution set of (η-pwvvi) is single valued, i.e., for fixed µ ∈ λ, sη(µ) is a singleton. proof. suppose, to the contrary, that there exist x1, x2 ∈ sη(µ), but x1 ̸= x2. by definition of sη(µ), ⟨t (x1, µ), η(y, x1)⟩ ∈ y \ − intc, ∀y ∈ k (4) ⟨t (x2, µ), η(y, x2)⟩ ∈ y \ − intc, ∀y ∈ k (5) in particular, take y = x2 in (4) and y = x1 in (5), respectively, we obtain ⟨t (x1, µ), η(x2, x1)⟩ ∈ y \ − intc (6) ⟨t (x2, µ), η(x1, x2)⟩ ∈ y \ − intc (7) since (6) holds and t is η-strongly c pseudomonotone, we claim that there exists λ > 0, such that ⟨t (x2, µ), η(x2, x1)⟩+ λ ∥ η(x2, x1) ∥2 by ∈ c 400 changchang bu: stability of solution maps for a -parameter weak vector ... consider that x1 ̸= x2, we have ⟨t (x2, µ), η(x2, x1)⟩ ∈ intc (8) in view of the condition (c2) of η(y, x), (8) is equivalent to ⟨t (x2, µ),−η(x1, x2)⟩ ∈ intc that is to say ⟨t (x2, µ), η(x1, x2)⟩ ∈ −intc this is a contradiction to (7). therefore, for fixed µ ∈ λ, sη(µ) is a singleton. theorem 3.3. let k is a non-empty, compact subset of x. assume k is invex with respect to η(y, x). if the following conditions hold: (i) for fixed µ ∈ λ, t (·, µ) is η-hemicontinuous on k; (ii) for fixed µ ∈ λ, t (·, µ) is η-strongly c pseudomonotone on k; (iii) for fixed x ∈ k,t (x, ·) is continuous on λ. then, sη(·) is continuous on λ. proof. first of all, we invoke lemma 3.2 to conclude that sη(·) is single valued. thus, we assume, without loss of generality, that sη(µ) = x(µ), ∀µ ∈ λ. take µ0 ∈ λ, in order to obtain that sη(·) is continuous at µ0, we only need to show that x(µ) → x(µ0) as µ → µ0. for any sequence {µn} ⊆ λ that satisfies µn → µ0, we can find a solution set sequence x(µn) ∈ k. further, note that k is compact, there exists a convergent subsequence {x(µnk )} such that x(µnk ) → ν. next we prove that ν = x(µ0). note that x(µn) is the solution of (η-pwvvi), we obtain ⟨t (x(µn), µn), η(y, x(µn))⟩ ∈ y \ − intc, ∀y ∈ k (9) in view of lemma 3.1, (9) is equivalent to ⟨t (y, µn), η(y, x(µn))⟩ ∈ y \ − intc, ∀y ∈ k (10) we claim, bear in mind that for fixed x ∈ k,t (x, ·) is continuous on λ and η(y, ·) is continuous, that ∥ ⟨t (y, µn), η(y, x(µn))⟩ − ⟨t (y, µ0), η(y, ν)⟩ ∥ ≤ ∥ ⟨t (y, µn), η(y, x(µn))⟩ − ⟨t (y, µ0), η(y, x(µn))⟩ ∥ + ∥ ⟨t (y, µ0), η(y, x(µn))⟩ − ⟨t (y, µ0), η(y, ν)⟩ ∥ ≤ ∥ t (y, µn)− t (y, µ0) ∥ · ∥ η(y, x(µn)) ∥ + ∥ t (y, µ0) ∥ · ∥ η(y, x(µn))− η(y, ν) ∥ advances in systems science and applications (2014) vol.14 no.4 401 thus, let n → ∞, we have ⟨t (y, µn), η(y, x(µn))⟩ → ⟨t (y, µ0), η(y, ν)⟩ (11) note that y \ − intc is closed and (11), we get ⟨t (y, µ0), η(y, ν)⟩ ∈ y \ − intc, ∀y ∈ k combine this with condition (ii) and use lemma 3.1 again, we find ⟨t (ν, µ0), η(y, ν)⟩ ∈ y \ − intc, ∀y ∈ k hence, v ∈ sη(µ0). recall that, by lemma 3.2, sη(·) is single valued. thus, ν = x(µ0). the proof is complete. remark 3.1 note that in theorem 3.3, t is required to be a η-hemicontinuous operator which extend ν-hemicontinuous in [6] and hence it weaken the continuous condition in [13]. 4 example in this section, we present an example to show that there exists η(y, x) function that satisfies the condition (c1)-(c3). let k = r, take the function η(y, x) = { y − x, x ≤ 0, y ≤ 0 and x ≥ 0, y ≥ 0, x− y, x ≤ 0, y ≥ 0 and x ≥ 0, y ≤ 0. it is evident that k is invex with respect to η(y, x). the function η(y, x) above satisfies condition (c2) and (c3) is obvious. next we verify that it also satisfies condition (c1). (i) for x ≤ 0, y ≤ 0 and any λ ∈ [0, 1], x+ λ(y − x) = (1− λ)x+ λy ≤ 0, η(x, x+ λη(y, x)) = η(x, x+ λ(y − x)) = η(x, (1− λ)x+ λy) = −λ(y − x) = −λη(y, x). (ii) for x ≥ 0, y ≥ 0 and any λ ∈ [0, 1], x+ λ(y − x) = (1− λ)x+ λy ≥ 0, η(x, x+ λη(y, x)) = η(x, x+ λ(y − x)) = η(x, (1− λ)x+ λy) = −λ(y − x) = −λη(y, x). 402 changchang bu: stability of solution maps for a -parameter weak vector ... (iii) for x ≤ 0, y ≥ 0 and any λ ∈ [0, 1], x+ λ(x− y) = (1 + λ)x− λy ≤ 0, η(x, x+ λη(y, x)) = η(x, x+ λ(x− y)) = −λ(x− y) = −λη(y, x). (iv) for x ≥ 0, y ≤ 0 and any λ ∈ [0, 1], x+ λ(x− y) = (1 + λ)x− λy ≥ 0, η(x, x+ λη(y, x)) = η(x, x+ λ(x− y)) = −λ(x− y) = −λη(y, x). the analysis above shows that the η(y, x) function satisfies the condition (c1)(c3) but η(y, x) ̸= y − x. that is to say η-parameter weak vector variational inequality(η-pwvvi) extends the existing parameter weak vector variational inequality and the results we obtained is reasonable. acknowledgements this research is supported by the doctoral fund of education ministry of china (20134219120003), the natural science foundation of hubei province (2013cfa131) and the nature science foundation of china (f030203). references [1] f. giannessi. (1980), “theorems of alternative, quadratic programs and complementarity problems”, variational inequalities and complementarity problems, edited by r. w. cottle, f. giannessi and j. l. lion, new york: john wiley and sons, pp.151-186. [2] y. chen, g. m. chen. (1987), “variational inequalities and vector optimization”, lecture notes in economics and mathematical systems, berlin:springer-verlag, 285, pp.408-416. [3] g. y. chen. (1992), “existence of solutions for a vector variational inequality: an extension of the hartman-stampacchia theorem”. j. optimiz. theory appl., 74, pp.445-456. [4] a. h. siddiqi, q. h. ansari, a. khaliq. (1995), “on vector variational inequalities”, j. optimiz. theory appl., 84, pp.171-180. [5] k. l. lin, d. p. yang, j. c. yao. (1997), “generalized vector variational inequalities”, j. optimiz. theory appl., 92, pp.117-125. [6] l. n. wang, z. m. fang, j. z. li. (2013), “stability of solution maps for aweak vector variational inequality”, journal of southwest university (natural science edition), vol.35, no.1, pp.95-98. (in chinese). advances in systems science and applications (2014) vol.14 no.4 403 [7] s. r. mohan, s. k. neogy. (1995), “on invex sets and preinvex functions”, j. math. anal. appl., 189, pp.901-908. [8] x. m. yang, x. q. yang, k. l. teo. (2003), “generalized invexity and generelized invariant monotonicity”, j. optimiz. theory appl., 117, pp.607625. [9] s. j. yu, j. c. yao. (1996), “on vector variational inequalities”, j. optimiz. theory appl., vol.89, no.3, 749-769 . [10] s. chang, b. s. lee, y. q. chen. (1995), “variational inequalities for monotone operators in nonreflexive banach space”, appl. math. lett., vol.8, no.6, pp.29-34. [11] n. j. huang, y. p. fang. (2005), “on vector variational inequalities in reflexive banach space”, j. global optim., vol.32, no.4, pp.495-505. [12] b. s. lee, g. m. lee. (1999), “variational inequalities for (η; θ)pseudomonotone operators in nonreflexive banach space”, appl. math. lett., vol.12, no.5, pp.13-17. [13] a. barbagallo, m. g. cojocaru. (2009), “continuity of solutions for parametric variational inequalities in banach space”, j. math. anal. appl., vol.351, no.2, pp.707-720. corresponding author yuqiang feng can be contacted at: yqfeng6@126.com. adv syst sci appl 2018; 02; 1-10 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/620 copyright ©2018 assa. adv. in systems science and appl. (2018) organization of traffic flows simulation aimed at establishment of integral characteristics of their dynamics anatoliy a. solovyev1,2, andrey m. valuev1 1) mechanical engineering research institute, russian academy of sciences, moscow, russia e-mail: aa.solovjev@yandex.ru 2) moscow institute of physics and technology (state university), dolgoprudny, moscow region, russia e-mail: valuev.online@gmail.com abstract: the paper presents an approach aimed at the study of various aspects of traffic passage through a multi-lane road that is based on computational experiments. microscopic simulation based on models of the type “leader-following models” is considered to be the most adequate and the most accurate means for performing them. the leader-following model of the entire traffic flows in a city road network acquires the more holistic form via the formalism of hybrid dynamic systems or, in other words, event-switched systems. the general formalization of this approach is presented and the corresponding representation of multi-lane traffic according to this approach is presented. the latter includes the detailed formal description of the traffic flow carrier with a certain traffic organization as well as the description of conditions for drivers’ choice of acceleration/braking and lane change according to incentive motives for it and positions and speeds of neighboring vehicles and the vehicle itself. organization of computational experiments allowing to establish the dependence of the average speed of traffic on the density of incoming flows, the distribution of various types of vehicles and their drivers, road organization and other factors in a multi-lane road is considered. it is demonstrated in what way they allow to evaluate quantitatively the dependence of the road throughput on the above factors. results of calculations are presented and analyzed. keywords: hybrid systems; traffic flows; leader-following model; city road network; multi-lane traffic; computational experiments. 1. introduction the city road network (crn) as a whole and its individual segments limit satisfaction of transport needs of the population, which is expressed in slow motion, time waste in traffic jams, rejection of some desirable trips or choosing an uncomfortable time for them. effective use of the throughput of the existing crn and its fragments, as well as substantiation of measures to increase it, requires knowledge of the characteristics of individual roads and crossroads with the existing traffic organization and alternative options. meanwhile, these characteristics are approximately known only for the most elementary cases. so, for singlelane traffic, estimates of road throughput have been established, ranging between 2000 and 2790 vehicles per hour [6, 8]. the variance of estimates is related to the difference in driving culture and the use of a different fleet of cars and trucks, and these differences can manifest even within one megacity. the use of road throughput is also affected by the distribution of trips between transport correspondences, almost always preventing the full load at the maximum level of all roads and their lanes. 2 a.a. solovyev, a.m. valuev copyright ©2018 assa. adv. in systems science and appl. (2018) the busiest highways carrying the bulk of the traffic flow (tf) are characterized by multi-lane movement. multi-lane traffic also results in complex organization of motion at crossroads, for which a number of types are distinguished [3]. integral characteristics of the flow through homogeneous segments of roads and crossroads depend on many circumstances and numerical characteristics of the tfs, including the proportions in which the flows are divided between directions in the nodes of the road network, as well as location of segments of allowed passage between the lanes and parking along the roads. so it quite impossible to establish integral characteristics of tf dynamics directly from treatment of observed data because of multidimensionality of affecting factors. however, systematical observations must be the source of needed data concerning elementary components of tfs (such as motion of a pair of successive cars) on a representative set of drivers and vehicles. physical-style theories, like widely known three-phase traffic theory by b.s. kerner [4] yield the qualitative picture of the traffic phenomena but hardly may take into account all the affecting details. another way to establish the necessary integral characteristics of tf is realized by mathematical modeling (analytical or, more often, by simulation). in this area both macroscopic models (e.g., the modern versions of hydrodynamic lwr models [2]) and microscopic ones exist. the latter, in turn, may be relatively rude cellular automata based on discretization of space and time [5] or much more realistic models in terms of dynamic systems that unite the movements of individual vehicles described with odes subject to additional servo-links. the conceptual basis for modeling tf as a dynamic system is the “leader-following model”, the latter being a formal description of the vehicle control choice by a driver depending on the movement of the previous one [6]. the paper [3] put forward the way of integration of these particular models together with conditions resulting from traffic light regulation in the in the holistic model. the next step of evolution of this approach is the expression of such models via the formalism of hybrid systems [7]. this way of tf modelling may be applied to multi-lane traffic as well, but in that case the “leader-following model” must be coupled with a “lane-changing model” [6]. the paper presents the construction of such a model and the way to organize computational experiments with it and treat their results to obtain integral characteristics of tf dynamics. 2. general approach to traffic modeling via the formalism of hybrid dynamical systems the analytical study of the tf dynamics within the framework of the “leader-following model” is limited to simple cases, mainly related to the so-called homogeneous congested traffic along one lane with a number of limiting assumptions, the main one being the identity of vehicles and their drivers (for example, [6]). the same models can be used via simulation. advantages of such models of private tf modes consist in the explicit formulation of all interrelations. undoubtedly, modeling by computer simulation has repeatedly been used to investigate more complex regimes, but only in the form of agent modeling, which treats relationships between the movements of neighboring vehicles as some forms of interaction between them. however, an integrated description of such an interaction may be absent; in any case, it is clearly not formulated in publications and therefore gives no possibility for critical analysis. with respect to verification of any agent modeling system, if it was fulfilled, one cannot be sure whether it is the result of a really correct reproduction of the nature of the transport processes or a successfully performed parametric identification (in other words, fitting) of the working model for specific conditions. in accordance with common sense traffic members are considered as “vehicle–driver units” (vdu) [6], i.e., specific vehicles that are controlled by specific drivers. the proposed approach does not postulate a specific model for a separate vdu and its choice of motion direction (in the case of a shift between lanes, a stop for parking, or a crossroad pass) in organization of traffic flows simulation 3 copyright ©2018 assa. adv. in systems science and appl. (2018) interaction with the movements of neighboring vehicles. such elementary relationships, obtained empirically, can be integrated into the resulting tf model for the road network segment, provided they are expressed in the required form. the model also includes a formal representation of the traffic organization on the segment of the road network being studied. in this sense, not a specific model of the tf is proposed, but a metamodel, which, however, is uniquely concretized to a specific tf model under specific conditions when using the required model elements characterizing the conditions being considered definitely or presumably. components of the model are: 1. formal description of the tf carrier furnished with the means of traffic regulation (traffic lights, barriers) and the conditions for the permissibility of trajectories along tf carriers; 2. formal description of the dependencies that determine the drivers’ choice of control when driving on a standard trajectory; 3. formal description of the conditions for drivers’ choice of the direction of movement; 4. description of the input streams (inflows) and their links with outflows that means distribution of vdus entering the road network in consideration in a certain on-ramp between out-ramps. 1. the tf carrier is a segment of the road network in a detailed description. permissible directions of motion on crossroads must be indicated; for a multi-lane road they must be determined for each lane. each crossroad is described by its own network, or a “route web”, uniting permissible trajectories that connect the all permitted pairs of entering lanes ends and beginnings of exit lanes. points of their intersection must be indicated as well. for regulated crossroads, it is indicated for each trajectory at which phase of the traffic light cycle movement along it is permitted. if parking is allowed on the road segment at the road edge or in special parking pockets, the totality of such parking spaces is also treated as a lane. in general, the carrier is represented by an oriented labelled graph g with some peculiarities. they consist in the fact that each arc (oriented edge) of g, representing a certain lane, is assigned to a certain road segment and it is determined which lanes (if any) of the same segment are adjacent to it on the left and on the right. an arc may allow (along a certain zone of its border) the shift of vdus moving along it to the adjacent lane. zones of permissible transitions must be determined; they are treated as labels of arcs. arcs of a crossroad “route web” connecting the same road segments are also grouped into a segment of the crossroad network. the transition between arcs that do not belong to the same road segment is possible only at the points of their joining, i.e. at the end of the entering arc, coinciding with the beginning of the outgoing arc. these points are the vertices (nodes) of the graph g. intersection points of admissible trajectories at crossroads are considered as vertices (nodes) of the graph g only if through such a point the vdus travelling along more than one admissible trajectory may pass within certain time intervals. the latter is characteristic only for unregulated crossroads. in view of the inability to simulate the tfs in their natural limits (then we would have to consider the road network from vladivostok to lisbon), a fragment is cut from the network into which the vdus enter from outside through the in-ramps and exit outside through the out-ramps. 2. each edge of g at each instant of time contains a chain of vehicles, treated as a moving, and sometimes a motionless queue. the concept of a hybrid system means mixed, discretecontinuous dynamics. discrete dynamics in models of this type include the change of the set of vdus in some queue, as well as the change of the traffic light phases and the mentioned below modes of motion of individual vdus. the dependencies defining the control laws (in the sense that this concept has in the control theory) for vdus are based on safety conditions. the most important and mandatory condition of safety is aimed to exclude the collision with the previous car and can be 4 a.a. solovyev, a.m. valuev copyright ©2018 assa. adv. in systems science and appl. (2018) expressed as a formula of safe distance depending on the speed of this and, possibly, the previous car. it is supplemented by the conditions of speed reduction on curved sections of the road, as well as by the conditions of braking on the red phase of the traffic light (for the vehicle at the head of the queue). in the transition between the lanes, parking on the roadside, the safety condition is modified. another condition for control choice is determined by the restriction on the speed and the desired (for public transport means — the normative) speed (they are distinguished if the latter is less than the maximum allowed speed). it is assumed that the forms of the required dependencies are the same for all vdus, but have parameters whose values are different for different types of vdus. various researchers developing “leader-following models” treat in the same way the general nature of these dependencies: the safe distance measured from the front bumper of the follower to the leader’s rear bumper depends on the speeds of both. the parameters of the dependence are the characteristics of the driver (first, its reaction time) and the vehicle of the follower. it is more convenient to consider the safety conditions with respect to the distance between the front bumpers of both (their positions relative to the current lane start are denoted by si-1(t), si(t) and serve as phase variables in the dynamic system of the tf). then we need to add the length of the vehicleleader to the “natural” distance. it is assumed, therefore, that the basic safety conditions si-1(t)-si(t)≥ssafe(vi-1(t),vi(t),pi,pi-1) (2.1) are satisfied permanently for all pairs “leader-follower”. together with (2.1) for curved sections of the road, restrictions on the speed are considered depending on the local curvature crv(s): �̇�𝑖 ≤ 𝑉safe crv(𝐶𝑅𝑉(𝑠𝑖(𝑡))). (2.2) in addition, there is a restriction similar to (2.2), on the maximum speed dictated by the traffic organization; they both can be combined in the form (2.2). the prohibition of entry to a crossroad when the traffic light is red means zero speed of the first vehicle in the lane at the moment of reaching the crossroad border and further until the green light turns on. obviously, this condition is one of those that make hybrid the dynamic system of a tf: the leader of the queue changes the mode of motion from braking (until reaching the border of the crossroad) to a stop and then to acceleration from the moment when the green light turns on. such qualitatively different situations for choosing control are various but few). they are: acceleration at free movement when a speed is less than the desired speed; maintaining the latter; in the case when the minimum safe distance to the consequent vdu is reached — maintaining this distance; braking before a traffic light, or when stopping for parking, or when approaching a turn; transition between the lanes. for each of these cases the definite law of control acts, in other words, a definite motion mode. for each vdu, the current mode is treated as a variable of the qualitative state; other such variables characterize the position on the graph g (arc number) and the sequence number in the chain on the arc. the variables that characterize the current and forthcoming directions of motion may also be treated as variables of the qualitative state as well. together with the regulator states, they form the vector of the qualitative state d of the simulated tf as a whole. the movement (or its expectation) of all vdu in constant modes for some time eventually leads to a change in the mode of one or several vdus or a change of the vdus set for some chains. the transition from one mode to another is caused by the achievement of a certain boundary (hypersurface) in the combined phase space of the vdu (and perhaps of the adjacent vdus) extended by the variable of the current time. for example, the condition for reaching a certain speed, approaching the minimum safe distance, starting the braking before the red traffic light signal from the stopping condition at the crossroad boundary under normal acceleration of braking (similarly when trying to occupy a free parking place) are all obviously conditions of that type. the same type of conditions determine the change in the of the vdus set for chains: reaching the end of the lane by the head of the chain; termination of the shift of a certain vdu between lanes at the moment of when the lateral organization of traffic flows simulation 5 copyright ©2018 assa. adv. in systems science and appl. (2018) coordinate reaches its final value; entering the network by a next vdu from the on-ramp at a certain time instant. all these changes are traditionally known for hybrid systems as switches. it should be emphasized that the conditions of each switch, in addition to the switching hypersurface, are also characterized by the set of values of some components of the vector d. so, braking before the crossroad at red light begins only with the shift to the corresponding phase of the traffic light; formally it means the change of the proper component of d. to start a shift between lanes or a parking process, the conditions are not limited to one component of d. all previously listed switches — changing the driving mode by a separate car, entering or leaving one tf in question, switching traffic lights at the end of the road — are considered as events of a stage change. within the stage the phases of all traffic light cycles, the number of vehicles on each arc, the mode of each vehicle remain unchanged. as moments of switches are not known in advance, the same holds for time limits [t(l-1), t(l)) for each (l-th) stage. the discrete state variables in the l-th stage are combined into the vector d(l), the phase variables characterizing the positions and speeds of the vehicles at the time moment t during the l–th stage — into the vector x(t,l), its components for the i-th vehicle form vector xi(t,l). dimensions of both x(t,l) and d(l) and even of certain xi(t,l) vary from stage to stage, since the latter includes the coordinate and the speed of movement in the transverse direction only during the transition between lanes. constant values characterizing individual vehicles and driving them, including the purpose of movement, are also treated as components of d(l). to describe the dynamics of an individual vehicle, the generic representation is used ))),(),((),,(),((/),( ltxldultxldfdtltdx iii  (2.3) in which one or more components of d(l) characterize the motion mode, and the formula expresses the control law for this mode. it seems most natural to use acceleration as a control variable. the conditions for determining the type j(l) and the moment t(l) of the switch that completes the l-th have the general form ,0))(),,(),(()( ltltxldg lj jj(d(l)), ))(),,(),(( ltltxldg j <0, jj(d(l))\{j(l)}, (2.4) and follow from restrictions of type (2.1), (2.2). the result of switching is the change of the set of qualitative state variables ))(()),(()1( )()( ldiilddld ljdlji  (2.5) together with, maybe, some phase variables )).(()),(),),((),(()1),(( )()( ldiiltlltxldxlltx ljxlji  (2.6) formula (2.6), in particular, refers to the longitudinal coordinates that change when passing from one arc to another ( 0)(( ltsi ), since they are now measured from the other point. the similar change takes place for the transverse coordinate at the time of completion of the transition to a new lane. however, a mode change may require certain values of several components of d(l). in this case, it occurs at the moment when the last one receives the required value. this moment is also determined by the condition (2.4) for a component of d(l), but only for the case when all other ones already have the required values (and do not change them at the last switching). however, it doesn’t not matter which component of d(l) is the last to get the required value. this does not require the modification or addition of the relations (2.4)–(2.6), but only their detailed specification. namely: a change in the regime of a specific vehicle can be made for different combinations of j(l) and d(l), but in all cases, the value of the corresponding component di(l+1) expressing the motion mode of this vehicle, in all cases, will be the same. 3. the current direction of movement of a specific vdu is characterized by the goal of stopping at the current section or continuing to move on reaching the end of the current segment for a certain new segment. hence, according to the current scheme of traffic 6 a.a. solovyev, a.m. valuev copyright ©2018 assa. adv. in systems science and appl. (2018) organization, this goal results in the task of transition to the lane from which it can proceed to this new segment. the overall goal of a particular vdu is to reach a specific target point (off-ramp) or to park in a certain area. with the exception of public transport, in most cases it is possible to choose a route. in practice, it is possible either to pre-select a route or to select it step by step when receiving information about the status of tfs, which, however, does not guarantee the choice of the most effective route. if we take as a basis any method of forecasting the duration of the route according to the current and historical data, then its recommendation will be reduced to the proposal to select the same next section for all vdus having the same target node and currently passing a certain node. each such recommendation is expressed by a variable that has the properties of a component of d and, apparently, can somehow be calculated within the same formalism, although it is unlikely that the method of calculating it will be relatively simple (for some forecasting methods, the computations that realize them are guessed, but confirm the assumption about their complexity and high amount of calculation). we retain such an option as possible in principle, but at the present stage of research we confine ourselves to a more realistic way of modeling the behavior of drivers, namely by assigning each of them a certain route (with detailing to the sequence of passing the sections of the road network). 3. modeling of traffic on a multi-lane road in accordance with the objectives of the study, a one-way multi-lane road from the crn is selected for study, it is bounded at both ends by crossroads (junctions) and not contains them within itself. the traffic organization can allow vehicles to leave the tf and enter the tf only from parking places along a part of the length of the road. parking spaces are collectively considered as the rightmost lane of a road, traffic on which is not allowed. regarding the transitions between the lanes, the segments of each lane on which passages through its left or right border are permitted (including for parking purposes) are defined. in addition, the destination of the end segment of each lane is defined as to carry out further movement in a certain direction or directions (right, left, right). limitations on traffic for each lane include: maximum speed and category of vehicles (passenger car, bus, freight with a certain weight limit) for which movement along it is allowed. the definite number of categories is introduced. for each individual participant (vdu) of the simulated tf, the following parameters are specified: 1) category; 2) the purpose of the movement, namely the direction of movement when leaving the lane or the intention to park along the road; 3) the length; 4) a set of indicators that determine the dynamics and driving, including perception of the minimum safe distance to the leader. it is assumed that to achieve its goal, each driver uses several driving modes. in the process of moving along one lane these modes are: 1) acceleration up to the maximum (desired) speed, 2) maintaining the desired speed; deceleration aimed at 3) stopping at a given place or (for a curved road) at 4) non-exceeding the safe speed; 5) maintaining the minimum safe distance to the leader. in the process of transition between the lanes, the variants of the same modes are implemented, but the conditions for the maintenance of the minimum safe distance are modified, since it is calculated to the nearest leader on the old and new lane. each mode is assumed to be expressed by the corresponding model law of control (acceleration/deceleration). the change of driving mode by a separate vehicle, entering or leaving the tf in question by a vehicle, switching traffic lights at the end of the road are considered as events of phase change; within the phase of the phase of the traffic light cycle, the number of vehicles, the mode of each remains unchanged. for the 1st mode, the “normal” (for a specific vdu) acceleration ( ii altxldu norm)),(),((  ) and for the third one “normal” deceleration organization of traffic flows simulation 7 copyright ©2018 assa. adv. in systems science and appl. (2018) ( ii bltxldu norm)),(),((  ) are accepted, for the 2nd one the uniform motion ( 0)),(),(( ltxldui ) takes place. for the fourth mode, the acceleration is determined from the motion condition at a safe speed ( dttsdvltxldu iii /))(()),(),(( safe ). in all these cases, there is no dependence of the control of the vdu on the phase coordinates of other vdus, which expresses the specifics of a tf as the interconnected motion of many of its participants. in the fifth mode, this specifics is present, but can be expressed in different ways. the movement of a vehicle is considered, first of all, along the axes of the road lanes. even in the passage to another lane, the trajectory is directed at a small angle to the axes of adjacent lanes and the longitudinal velocity is practically equal to the actual velocity; the change in velocity during the transition time can be neglected. therefore, with respect to the quantitative dynamics, we may reduce representation of the vdu dynamics to the equations of longitudinal motion with respect to the current position of the vehicle si measured along the axis of the lane, namely dsi(t, l)/dt=vi(t, l); dvi(t, l)/dt=ai(t, l), ai(t, l)= )),(),(( ltxldui . (3.1) for any representation and for any of the above modes, switches between them occur when the condition represented with one equation is satisfied, which parameter may be components of the qualitative state in the mode before switching. so, to complete the set of speed and transition to a uniform motion, this is just a condition vi(t)=vmax i, (3.2) for inclusion in the cluster and transition to the fifth mode — si-1(t)-si(t)=ssafe(vi-1(t),vi(t),pi), (3.3) similarly for switching between the fifth mode options. to start the transition, the transition conditions are somewhat more complicated. let us first formulate the conditions for the beginning of the transition at the content level. the condition for the possibility of a transition is: the admissibility of such a transition (the intersection of the dividing line in the direction of the new lane); presence of sufficient advance of the potential "follower" on the adjacent strip s(k); presence of motivation for transition under conditions of speed advantage, consisting in that the safe distance to the leader on the neighboring strip exceeds the safe distance to the leader on the current lane by at least; the presence of motivation for the transition to the conditions of the timely occupation of the lane, at which the vehicle under consideration is required to be located at the end of the present road segment. the moment of the onset of each of the conditions, as well as the moment when the previously fulfilled condition is violated, are determined, as above, from relations of the form (2.4), the parameters of which relate to the considered vdu, preceding on the strip and conditional leaders and followers for it on adjacent lanes. the moment of the shift beginning begins a stage that is characterized with additional coordinates and equations for lateral motion and additional relations between (security conditions are associated with a tuning vehicle with its "leader" and "follower" on both strips). with respect to the vdus the lateral coordinate, the condition for the termination of the transition is recorded, the result of which is the change in the order of the vdu on both lanes and the restoration of the old system of constraints with new parameters. 4. organization of computational experiments in accordance with the research objectives, in a separate computational experiment, the tf dynamics are calculated with definite and on average constant values: 1) the intensity of the input flow (and exit from parking places along the road, if it is allowed); 2) the proportions between the target directions in the input stream; 3) the composition of the input stream (in the simplest case the shares of the introduced categories of vehicles). in accordance with the selected values of the experimental constants listed above, the characteristics of the 8 a.a. solovyev, a.m. valuev copyright ©2018 assa. adv. in systems science and appl. (2018) individual vehicles of the input sequence and the time intervals between their appearances are randomly generated. a separate experiment begins with an empty road gradually filled with vehicles. flow characteristics are calculated in steady state when the road is completely filled. average characteristics and their dispersion for a sufficiently long period are determined. in the case of simulation of the application of traffic light regulation, the dynamics of the tf for a period after reaching the steady state, a multiple of the duration of the traffic light cycle is considered. the result of the calculation is the sequence of values of the components of the vectors x(t,l) and d(l). at the same time, these quantities can be stored in the database, but only integral indicators of flows along the lanes. initial data for the computational experiments are recorded in the database tables (see tables 4.1 and 4.2) and extracted from them for tf simulation. they characterize both the content conditions of the experiments and calculation parameters (see the maximal step tmax for integrating motion equations in table 4.1). table 4.1. general data for computational experiments expt. tmax, s tmax, s averaging interval, s qlane, vdus per hour t in tst011 0,5 300 10 1500 0,5 tst011a 0,5 600 10 1500 0,5 tst021 0,5 300 10 2000 0,5 tst021a 0,5 420 10 2000 0,5 tst021b 0,5 420 10 2000 0,5 table 4.2. data оn vdu types in computational experiments expt. vdu type anorm i; bnorm i length; width vopt, m/s treac, s sin type quote in onramp flow tst011 1 4;;7 3; 1,8 25 0,3 0,1 75 tst011 2 2; 4 10; 3 17,5 0,2 0,15 25 tst021 1 4; 7 3; 1,8 25 0,3 0,1 75 tst021 2 2; 4 10; 3 17,5 0,2 0,15 25 the character of the tf dynamics is ambiguous and depends on the characteristics of the atp (see tables 4.3, 4.4). at a moderate tf intensity (in the examples it is 1500 vdu/h per one lane), flow characteristics are stable, although deviations from the mean values take place permanently preserved. with the equivalence of the lanes, neither one gets an advantage. table 4.3. statistics on atp, a series of experiments “tst011” expt. tbeg tend mean vdus count mean vdu speed vdus in shift, % tst011 20 300 11,75 21,36 8,81 tst011a 20 300 12,18 20,78 9,94 tst011a 300 600 12,55 20,85 10,97 tst011b 20 300 11,92 21,08 9,67 tst011b 300 600 11,72 21,18 8,00 organization of traffic flows simulation 9 copyright ©2018 assa. adv. in systems science and appl. (2018) table 4.4. statistics on atp, a series of experiments “tst011” expt. tbeg tend mean vdus count mean vdu speed 1st lane 2nd lane 1st lane 2nd lane tst011 20 300 6,02 5,73 21,50 21,18 tst011a 20 300 5,95 6,23 20,71 20,81 tst011a 300 600 6,46 6,10 20,64 21,01 tst011b 20 300 6,26 5,66 21,47 20,51 tst011b 300 600 5,74 5,88 20,93 21,44 on the contrary, the flow approaching its maximum intensity is unstable (see tables 4.5, 4.6), queues at the entrance and supersaturation of the road section can be formed. in this case, the speed can drop sharply. these phenomena are well known, but modeling by simulation allows determining the conditions of their occurrence (depending not only on the intensity of the flows, but also on their composition with respect to vdu types). table 4.5. statistics on atp, a series of experiments “tst021” expt. tbeg tend mean vdus count mean vdu speed vdus in shift, % tst021 20 300 19,18 17,36 12,72 tst021b 20 300 16,81 19,53 11,13 tst021b 300 400 22,35 9,62 24,95 table 4.6. statistics on atp, a series of experiments “tst021” expt. tbeg tend mean vdus count mean vdu speed 1st lane 2nd lane 1st lane 2nd lane tst021 20 300 9,79 9,38 17,25 17,52 tst021b 20 300 8,07 8,74 19,31 19,68 tst021b 300 400 9,75 12,61 11,59 8,30 repetition of a series of experiments with different values makes it possible to obtain a series of data for the characteristics of the unknown dependencies that are interpolated to the range of interest. 5. conclusion the paper presents the general approach to establishment of integral characteristics of traffic flow dynamics, notably for multi-lane roads, and illustrates it in the case of two-lane traffic with a relatively simple road organization. with the development of this approach more sophisticated integral characteristics of multi-lane tfs must be obtained, including the influence of the structure and parameters of the traffic light cycle at the end of the road in question and the distribution of vdus between desired directions of subsequent passage of the road network. acknowledgements we thank the members of the organizational committee of mlsd’2017 for their kind invitation to contribute the paper presented at the conference to the journal and andrey 10 a.a. solovyev, a.m. valuev copyright ©2018 assa. adv. in systems science and appl. (2018) shevlyakov, the secretary of advances in systems science and applications, for technical assistance. references [1] bando, m., hasebe, k., nakayama, a., shibata, a. & sugiyama, y. (february 1995). dynamical model of traffic congestion and numerical simulation. physical review e, 51(2), 1035–1042. [2] garavello, m., & piccoli, b.a. (2013). multibuffer model for lwr road networks. complex networks and dynamic systems. (2), 143-161. [3] glukharev, k.k., ulyukov, n.m., valuev, a.m., & kalinin, i.n. (2013). on traffic flow on the arterial network model. traffic and granular flow’11. berlinheidelberg, germany: springer-verlag, 399–412. [4] kerner, b.s. (2009). introduction to modern traffic flow theory and control. the long road to three-phase traffic theory. springer science & business media. [5] ma, k., yan, b., & luo, x. (june 2014). a cellular automata simulation for traffic flow on multi-lane freeways under various control rules. 11th ieee world congress on intelligent control and automation (wcica). shenyang, china, 455–460. [6] treiber, m., & kesting, a. (2013). traffic flow dynamics: data, models and simulation. berlin-heidelberg, germany: springer-verlag. [7] valuev, a.m. (2014, june) modelirovanie transportnykh protsessov v formalizme gibridnykh sistem [modeling of transport processes in the formalism of hybrid systems]. xii vserossiiskoe soveshchanie po problemam upravleniya vspu-2014, proceedings (electronic resource), moscow, russia, 5033–5043. [in russian]. [8] velmurugan, s., errampalli, m., ravinder, k., sitaramanjaneyulu, k., & gangopadhyay, s. (2010, october). critical evaluation of roadway capacity of multilane high speed corridors under heterogeneous traffic conditions through traditional and microscopic simulation models. journal of indian roads congress, 71(3), 235– 264. adv syst sci appl 2020; 02:71–81 published online at https://ijassa.ipu.ru. comparison of two dynamic models of economic growth alexander p. chernyaev1* 1moscow institute of physics and technology (state university), dolgoprudnyi, russia abstract: two dynamic models of economic growth with the same balance equation are considered. first, we establish the solution of the harrod-domar model with time-dependent coefficient of the capital intensity of income growth. (previously, the only constant coefficients were considered.) second, we show that in the solow model with the cobb-douglas production function, the capital intensity of income growth depends on time. comparing these models, we demonstrate the effectiveness of the setting optimal control problems (maximization of the integral discounted utility function) in the extended the harrod-domar model. keywords: economic growth, production function, harrod-domar model, solow model 1. introduction models of economic growth became very popular due to their universality. they were applied to various objects in economic structures of many kinds. in mathematical economics, there are two widely acknowledged models of economic dynamics: the harrod-domar model and the solow model, which are presented in scientific and educational literature. see [1] – [12]. in both mentioned models, the total income is the sum of the total investment and the total consumption. following [13], we establish the exact solution of the cauchy problem for the differential equation in the harrod-domar model of macroeconomic dynamics with the time-dependent coefficient of the capital intensity of income growth (ciig). previously, the only constant coefficients of ciig were considered; see [12]. to confirm economic validity of the assumption that the coefficient of ciig depends on time, we investigate the exact solution of the solow model with the cobb-douglas production function. see, e.g., [3, 4] and [7] – [10]. calculation of the income growth capital intensity factor of this solution shows that this coefficient is a time-dependent function. in the present paper, we also show that the harrod-domar model is quite convenient for using the apparatus of the optimal control theory and the calculus of variations. using these methods, one can find the maximum of the integral discounted utility function of consumption. there are also formulated and investigated several optimal control problems with various constrains that follows from natural economic conditions. problems of this type are to find the maximum of a functional that expresses the integral discounted utility function in the presence of a differential relation. consumption and phase constraints are investigated, extremal problems in the pontryagin and dubovitsky-milyutin forms are considered. ∗corresponding author: chernyaev49@yandex.ru 72 a.p. chernyaev 2. the extended harrod-domar model in the harrod-domar model, the differential equation of the macroeconomic dynamics with exogenous dynamics of the consumption of arbitrary character [12, 13] has the form y (t) = c(t) +by ′(t). (2.1) in this model, time t is continuous. the income y (t) is equal to the sum of the consumptionc(t) and the investment i(t). usually, y (t) refers to the gross domestic product, which is identified with the national income. the economy is supposed to be closed, therefore, the net exports are zero and the government expenses are not considered in the model. the main factor of the growth – the speed of income growth – is proportional to the investment. see [6, 12, 13]. that is, i(t) = by ′(t), where b is the coefficient of the capital intensity of income growth (ciig) and 1/b is the limit product of capital at the macroeconomic level. previously, the coefficient of ciig was supposed to be a positive constant [12, 14]: b = const > 0. (2.2) in the case (2.2), the solution of differential equation (2.1) is given by the formula y (t) = y0e t−t0 b − 1 b ∫ t t0 c(τ)e t−τ b dτ. (2.3) we consider the cauchy problem: differential equation (2.1) with the initial condition y (t0) = y0 > 0. (2.4) the main assumption is that b = b(t). (2.5) the solution of this cauchy problem is given by the formula (see [13]): y (t) = y0e ∫ t t0 ds b(s) − e ∫ t t0 ds b(s) ∫ t t0 c(τ) b(τ) e − ∫ t t0 ds b(s)dτ. (2.6) obviously, in the case (2.2) formula (2.6) becomes (2.3). the economic validity of the assumption (2.5) follows from the comparative analysis of the harrod-domar model and the solow model. from the economic viewpoint, it reflects the rate of the technical progress and the rapidly changing of the economic conditions. 3. the solow model there are several different ways of the presentation of the solow macroeconomic model. see, e.g., the original works [3, 4] and the papers [6] – [11]. in the solow model, the average per capita capital k satisfies a first-order nonlinear differential equation, which follows from the balance equation for funds: dk dt = −λk + ρf(k), (3.7) with the initial condition k(0) = k0 > 0. (3.8) equation (3.7) with condition (3.8) yield cauchy problem we shall deal with. here k = k/l, where k is the capital or funds, l means human resources (or labour) resources, time t is measured in years, ρ is the norm of accumulation (the share of gross copyright © 2020 assa. adv syst sci appl (2020) dynamic models of economic growth 73 investment in the gross domestic product). following [7], we shall assume that ρ = const, 0 < ρ < 1. the function f(k) = f (k,l) l = f (k, 1). (3.9) here f (k,l) is the neoclassical production function; see [7, 8, 15]. the coefficient λ = µ+ ν, where µ is the share of the annual decrease of main production funds. similarly to [7], we assume that µ = const. the second term ν means the annual growth rate of the labour force, i.e. 1 l dl dt = ν. (3.10) here µ and ν satisfy the restrictions 0 < µ < 1, −1 < ν < 1. the balance equation for funds reads dk dt = ρy − µk. (3.11) then for the average per capita capital we have the following chain of the equalities dk dt = d dt ( k l ) = k ′tl−kl′t l2 = k ′t l −kl′t l2 = ρy − µk l − νk l . this yields equation (3.7). the transition mode [7, 8] was investigated under assumption that the production function f (k,l) is the cobb-douglas function [7, 8, 14]. indeed, in this case f(k) is the power function with a constant factor; see [7, 8]. then equation (3.7) is explicitly integrated, and its solution is represented via elementary functions. if ν is constant, equation (3.7) has the stationary solution k = k∗, where k∗ is the positive root of the equation ρf(k)− λk = 0. (3.12) here we assume that k0 and k∗ belong to the interval of average per capita capital under consideration. for k 6= k∗, integrating (3.7) with the initial condition (3.8), we obtain the integral equation ∫ k k0 ds ρf(s)− λs = t (3.13) which determines the function k = k(t). remark that formula (3.13) is correct if the function (ρf(s)− λs)−1 is integrable on the corresponding interval. sufficient conditions for this can be formulated in several ways. one can impose some conditions on f (as it is done in [11]) or impose conditions on the neoclassical production function f (k,l) and use (3.9). to satisfy the basic condition of the transient mode in the solow model [7, 8] k∞ = lim t→+∞ k(t) = k∗, (3.14) equation (3.12) needs to have one positive root k∗ on the interval under consideration. in addition, the improper integral ∫ k∗ k0 ds ρf(s)− λs (3.15) needs to diverge. this follows from the limit transition in (3.13) as t→ +∞ and (3.14). therefore, for the existence of a transition regime in the solow model one need to assume that function f(k) generated by the neoclassical production function f (k,l) satisfies to copyright © 2020 assa. adv syst sci appl (2020) 74 a.p. chernyaev conditions mentioned above: the existence of a unique positive root k∗ 6= k0 of equation (3.12) in the interval under consideration and the divergence of the integral (3.15). it is worth observing that if the solution to the cauchy problem (3.7), (3.8) is known, then all endogenous variables can be found from the equality y = f (k,l) = c + i. (3.16) here, the final product y is used for the non-productive consumption c and the investment i . now let us discuss the solow model with the cobb-douglas production function [7,8,14]: f (k,l) = akαl1−α, a = const > 0, 0 < α < 1. (3.17) substituting (3.17) in formula (3.9), we obtain f(k) = f (k,l) l = f (k, 1) = akα. (3.18) taking into account (3.18), one can bring equation (3.7) to the bernoulli form: dk dt = −λk + ρakα. (3.19) solving equation (3.19) with the initial condition (3.8) and the additional assumption λ = const, (3.20) we get k(t) = e−λt [ ρa λ eλ(1−α)t − ρa λ + k0 1−α ] 1 1−α . (3.21) let us find the capital intensity of the income growth for (3.21), that is, for the average per capita capital in the solow economic growth model with the cobb-douglas production function (3.17) and the additional condition (3.20). using the basic premise of the harrod-domar model, the definition of the solow model accumulation norm and the balance formula (3.16), we have i(t) = by ′(t) = ρy, (3.22) that is, y ′(t) = ρ b(t) y (t). (3.23) integrating (3.23), we get y (t) = e ρ ∫ t t0 ds b(s) , where the exponent is the indicator of the income growth. theorem 3.1: under the above assumptions, the capital intensity of the income growth is given by the following formula: b(t) = i(t) y ′(t) = ρf (k,l) d dt [f (k,l)] = ρk (ν − αν − αµ)k + αρakα . (3.24) copyright © 2020 assa. adv syst sci appl (2020) dynamic models of economic growth 75 proof from (3.22) and (3.16), we have b(t) = i(t) y ′(t) = ρy (t) y ′(t) = ρf (k,l) d dt [f (k,l)] . using formula (3.17) and multiplying the numerator and the denominator of the right hand side by a−1lαk1−α, one can write the latter equality in the form b(t) = ρkl αldk dt + (1− α)k dl dt . applying the formulas (3.10) and (3.11) to the right hand side of this equality, we obtain b(t) = ρkl αl(ρy − µk) + ν(1− α)kl . finally, taking into account (3.16), (3.17) and multiplying the numerator and denominator of the right hand side by l−2, we have b(t) = ρ(k/l) αρa(k/l)α + (ν − αµ− αν)(k/l) . to complete the proof, recall the definition of the average per capita capital: k = k/l. remark 3.1: formula (3.24) justifies the assumption that the capital intensity of the income growth depends on time. remark 3.2: equation (3.19) can be considered without assumption (3.20), i.e., λ = λ(t) is an arbitrary integrable function. this makes sense, because the annual growth rate of the employment (growth rate of the labour force (3.10)) is not constant. remark 3.3: in the case λ = λ(t), the solution of bernoulli equation (3.19) with the initial condition (3.8) is given by the formula k(t) = e− ∫ t 0 λ(τ)dτ [ ρa(1− α) ∫ t 0 e(1−α) ∫ s 0 λ(τ)dτds+ k0 1−α ] 1 1−α . 4. optimal consumption in the extended harrod-domar model in the extended harrod-domar model, formula (2.6) determines the income through the consumption. a natural question: how to find the consumption? from the mathematical viewpoint, this question can be formulated as an optimal control problem. we consider this problem is the most general form, as the problem of maximizing the integral discounted utility of the consumption [17, 18]:∫ t1 t0 u(c(t)) exp(−δt)dt⇒ max, (4.25) where δ > 0 is the discount factor. here we use the analogy with problems of optimal consumption management in the household economy, see [19] – [23]. following [19] – [21], copyright © 2020 assa. adv syst sci appl (2020) 76 a.p. chernyaev we suppose that the utility of consumption is represented by the function u(c), which reflects constant aversion to risk by arrow-pratt: a = −u ′′(c)c u′(c) ≥ 0. (4.26) the economic meaning of (4.26) is clear from the notation g(c) = u′(c). (4.27) then g(c) is the limit utility of consumption. further, taking into account (4.26) and (4.27), we have the equalities: g′(c) = u′′(c), ec(g) = g′(c) g(c)/c = u′′(c)c u′(c) = −a. here, ec(g) is the elasticity of g with respect to c. further, we assume that g(c) is a monotonically decreasing function, therefore, a ≥ 0. conversely, the assumption that g(c) is increasing, is associated with the risk. it is natural to call the condition that g(c) does not increase the disgust to risk. lemma 4.1: the derivative of the utility function u describing the constant aversion to risk by arrow-pratt (4.26) is given by the formula u′(c) = g(c) = γ ca = γc−a, γ = const > 0. (4.28) proof considering (4.26) as a differential equation with the unknown function u, we get a c = −u ′′(c) u′(c) . integrating this equation, we obtain a ∫ dc c = α ln |c| = − ∫ u′′(c) u′(c) dc = − ln |u′(c)|+ const = ln ∣∣∣∣ c0 u′(c) ∣∣∣∣ . since logarithm is a monotonic function, this yields |c|a = ∣∣∣∣ c0 u′(c) ∣∣∣∣ . since the consumption is positive, one can put |c0| = γ > 0. taking into account (4.27), the last equality implies (4.28). corollary 4.1: the utility function u satisfying the condition of lemma 4.1 has the form u(c) =  γc1−a 1− a + χ, a 6= 1; γ lnc + χ, a = 1; γ = const > 0, χ = const. (4.29) proof integrating equation (4.28), we obtain (4.29). copyright © 2020 assa. adv syst sci appl (2020) dynamic models of economic growth 77 theorem 4.1: the consumption function c(t) = (u′) −1 [ 1 b(t) c1 exp { δt− ∫ t t0 dτ b(τ) }] (4.30) gives the maximum in the variation problem (4.25) with fixed boundaries. proof substituting the expression of the consumption from (2.1) into (4.25), we obtain the functional j(y ) = ∫ t1 t0 u(y −b(t)y ′) exp(−δt)dt, (4.31) whose increment j(y + h)− j(y ) = ∆j(y, h) has the form ∆j(y, h) = ∫ t1 t0 [u(y + h−b(t)(y ′ + h′))− u(y −b(t)y ′)] exp(−δt)dt. (4.32) from the equality y −b(t)y ′ = c(t), it follows that u(y + h−b(t)(y ′ + h′))− u(y −b(t)y ′) = u(c(t) + h−b(t)h′)− u(c(t)), ∆j(y, h) = ∫ t1 t0 [u(c(t) + h−b(t)h′)− u(c(t))]e−δtdt. (4.33) substituting the taylor expansion u(c(t) + h−b(t)h′) = u(c(t)) + u′(c(t))[h−b(t)h′] + 1 2 u′′(c(t))[h−b(t)h′]2 +r(t), where r(t) = o[h−b(t)h′]2 as h−b(t)h′ → 0, into (4.33), we obtain ∆j(y, h) = ∫ t1 t0 u′(c(t))[h−b(t)h′]e−δtdt+ 1 2 ∫ t1 t0 u′′(c(t))[h−b(t)h′]2e−δtdt+ ∫ t1 t0 r(t)e−δtdt. (4.34) then, integrating by parts, we have − ∫ t1 t0 u′(c(t))b(t)h′e−δtdt = − ∫ t1 t0 u′(c(t))b(t)e−δtdh = − u′(c(t))b(t)e−δth ∣∣∣∣t1 t0 + ∫ t1 t0 h d dt [u′(c(t))b(t)e−δt]dt. (4.35) since we consider the variation problem (4.25) with fixed boundaries, in addition to the initial condition (2.4), i.e., the boundary condition on the left edge, we have the boundary condition copyright © 2020 assa. adv syst sci appl (2020) 78 a.p. chernyaev on the right edge: y (t1) = y1 > 0. (4.36) from (2.4) and (4.36), it follows that for the function h = h(t) the equalities h(t0) = h(t1) = 0. (4.37) hold true. from (4.37), it follows that the first term in the right hand part of (4.35) is zero. therefore, the equality (4.35) reads − ∫ t1 t0 u′(c(t))b(t)h′e−δtdt = ∫ t1 t0 h d dt [u′(c(t))b(t)e−δt]dt. (4.38) from (4.38) and (4.34) we conclude that ∆j(y, h) = ∫ t1 t0 h ( u′(c(t))e−δt + d dt [u′(c(t))b(t)e−δt] ) dt+ 1 2 ∫ t1 t0 u′′(c(t))[h−b(t)h′]2e−δtdt+ ∫ t1 t0 r(t)e−δtdt. (4.39) the sign of the left hand side of (4.39) coincides with the sign of the first term in the right hand side of (4.39). indeed, replacing h with βh, where β = const, we get ∆j(y, βh) = j(y + βh)− j(y ) = = β ∫ t1 t0 h ( u′(c(t))e−δt + d dt [u′(c(t))b(t)e−δt] ) dt+ 1 2 β2 ∫ t1 t0 u′′(c(t))[h−b(t)h′]2e−δtdt+ ∫ t1 t0 r̃(t)e−δtdt, (4.40) where r̃(t) = o(β2[h−b(t)h′]2) as β(h−b(t)h′)→ 0. passing to the limit β → 0, one can see that the first term in the right hand side of (4.40) tends to zero with the first order of smallness, while the second term has the second order of smallness and the third term has the order greater than two. this proves the statement. now let us prove that the first term in the right hand side of (4.39) is zero. suppose the contrary. then, replacing β with−β, we change the sign of the first term in the right hand side of (4.40), while the sign of the left hand side of (4.40) remains the same. this contradiction shows that ∫ t1 t0 h ( u′(c(t))e−δt + d dt [u′(c(t))b(t)e−δt] ) dt = 0. by the fundamental lemma of the calculus of variations, this yields the euler equation u′(c(t))e−δt + d dt [u′(c(t))b(t)e−δt] = 0. (4.41) taking into account (4.41), one can simplify (4.39) as follows: ∆j(y, h) = 1 2 ∫ t1 t0 u′′(c(t))[h−b(t)h′]2e−δtdt+ ∫ t1 t0 r(t)e−δtdt. (4.42) let us check that the solution of the euler equation (4.41) with the boundary conditions (2.4) and (4.36), and consequently, (4.37), give the maximum of the functional (4.25), or copyright © 2020 assa. adv syst sci appl (2020) dynamic models of economic growth 79 equivalently, (4.31). first, we remark that the sign of the left hand side of (4.42) coincides with the sign of the first term in the right hand side of (4.42). indeed, replacing in (4.42) h with βh, where β = const, we have ∆j(y, βh) = 1 2 β2 ∫ t1 t0 u′′(c(t))[h−b(t)h′]2e−δtdt+ ∫ t1 t0 r̃(t)e−δtdt. (4.43) here r̃(t) = o(β2[h−b(t)h′]2) as β(h−b(t)h′)→ 0. passing to the limit β → 0, one can see that the first term tends to zero with the second order of smallness, while the second term the order greater than two. therefore, the sign of the left side of (4.43) coincides with the sign of the first term in the right hand side of (4.43). it remains to use the inequality u′′(c(t)) ≤ 0, which follows from (4.26). taking into account (4.27), we can write equation (4.41) in the form g(c(t))e−δt + d dt [g(c(t))b(t)e−δt] = 0. (4.44) the change of variables w = g(c(t))b(t)e−δt, (4.45) i.e., g(c(t))e−δt = w/b(t), transforms equation (4.44) into w b(t) + dw dt = 0. integrating the latter equation, we obtain w = c1 exp { − ∫ t t0 dτ b(τ) } , c1 = const. substituting the obtained equality in (4.45), after obvious transformations we obtain g(c(t)) = 1 b(t) c1 exp { δt− ∫ t t0 dτ b(τ) } , c1 = const > 0. (4.46) finally, recall that g(c) monotonically decreases. therefore, it is invertible, and c(t) = g−1 [ 1 b(t) c1 exp { δt− ∫ t t0 dτ b(τ) }] , c1 = const > 0. (4.47) taking into account (4.27), from (4.47) it follows (4.30). the proof is complete. remark 4.1: the consumption function (4.30) can be also written in the form c(t) = [ γb(t) c1 ] 1 a exp { 1 a [∫ t t0 dτ b(τ) − δt ]} . (4.48) proof from (4.28) and (4.46) we have the equality [c(t)]a = γb(t) c1 exp {∫ t t0 dτ b(τ) − δt } , (4.49) which can be resolves by c and it gives (4.48). copyright © 2020 assa. adv syst sci appl (2020) 80 a.p. chernyaev remark 4.2: the consumption function (4.30) or, equivalently, (4.48) gives the maximum in the variation problem (4.25) with fixed boundary conditions (2.4) and (4.36). formula (4.49) expresses the consumption under the condition (4.26), which means that the utility function satisfies the constant risk aversion according to arrow-pratt. it is very convenient to formulate the above problems using the terminology from the control theory. the problem of maximization of the functional (4.25) under constraints (2.1), (2.4), (4.36), and the additional restriction 0 < c ≤ c(t) ≤ c < +∞ (4.50) is called the pontryagin problem. in the inequality (4.50), the lower bound c means the total subsistence minimum and the upper bound c is the total subsistence maximum. in turn, the pontryagin problem can be also formulated with an additional phase constraint, for example, y (t) ≥ const ≥ 0. such a problem is often called the dubovitskymilyutin problem. see [19] – [21]. 5. conclusion we presented the comparative analysis of two models of economic dynamics: the model of the economic growth by harrod-domar and the model by solow. there were several earlier models of the economic growth, but they are not widely acknowledged; see, e.g., [24]. at present, more advanced models are gaining popularity, however, they are based on the models discussed in the present paper; see [25] – [27]. the comparative analysis confirms the economic viability of the assumption that the ciig depends on time. comparing these models and using the analogy with household economies [19] – [23], we demonstrate the efficiency of the control theory approach in the extended harrod-domar model. the obtained results shaw that despite significant differences between these models, their comparative analysis is substantial. it is worth observing that using the approach [28] – [31] based on the theory of covering mappings, the both considered models can be generalized to the market of many goods with various production functions. references 1. harrod, r.f. (1939) an essay in dynamic theory, economic jornal, 49, 14–33. 2. domar, e. (1946) capital expansion, rate of growth and employment, econometrica, 14 (2), 137–147. 3. solow, r.m. (1956) contribution to the theory of economic growth, the quarterly journal of economics, 70 (1), 65–94. 4. solow, r.m. (1957) technical change and the aggregate production function, the review of economics and statistics, 39 (3), 312–320. 5. hamburg, d. (1981) early growth theory of the domar and harrod. moscow: progress. 6. zamkov, o.o., tolstopyatenko, a.v., & cheremnykh, yu.n. (1998) mathematical methods in economy. moscow: lomonosov moscow state university, dis publishing house. 7. malychin, v. (2001) mathematics in economics: tutorial. moscow: infra-m. 8. kolemayev, v.a. (2002) mathematical economics: textbook for higher education institutions. moscow: unity-dana. 9. samarov, k.l. (2009) economic and mathematical models. moscow: resolventa. 10. samarov, k.l. & samarova, s.s. (2014) robert solow’ s model of economic growth in the course of differential equations, information and technological journal, 2, 81–84. copyright © 2020 assa. adv syst sci appl (2020) dynamic models of economic growth 81 11. meerson, a.y. & chernyaev, a.p. (2010) integral method of research of transition regime in solow model, economics of nature management, 3, 105–109. 12. meerson, a.y. & chernyaev, a.p. (2011) exact solution of the macroeconomic model of harrod-domar with exogenous dynamics of the volume of consumption of arbitrary character, russian economic university bulletin, 1, 142–147. 13. meerson, a.y. & chernyaev, a.p. (2013) the exact solution of koshi problem for the differential equation of the harrod-domar macroeconomic model with a variable coefficient of capital intensity of income growth, the journal of mgup, 3, 252–255. 14. gracheva, m.v., fadeeva, l.n., & cheremnych, y.n. (2005) simulation of economic processes: textbook for university students studying in the fields of economics and management. moscow: unity-dana. 15. drogobytsky, i.n. (2006) economic and mathematical modeling. moscow: examination. 16. chernyaev, a.p. (2019) comparison of two dynamicmodels of economic growth. proc. first int. conf.: math. physics, dyn. syst., infinite-dimensional anal., dolgoprudny, russia: mipt, 161–161. 17. meerson, a.y. & chernyaev, a.p. (2014) variation problem of optimization of consumption of the macroeconomic model of harrod-domar with variable coefficient of capital intensity of income growth, mgtu mami bull., 4 (3), 77–80. 18. meerson, a.y. & chernyaev, a.p. (2014) variation problem of optimization of consumption of the model of economic dynamics of harrod-domar with variable coefficient of capital intensity of income growth, proc. free economic soc. russia, 186, 502–506. 19. dikusar, v.v., meerson, a.y., & chernyaev, a.p. (2004) problems of optimal distribution of resources on the example of households. moscow: dorodnitsyn computing centre ras. 20. dikusar, v.v., meerson, a.y., & chernyaev, a.p. (2005) consumption patterns and questions of optimal control, theoretical and applied problems of nonlinear analysis, moscow: dorodnitsyn computing centre ras, 46–61. 21. dikusar, v.v., meerson, a.y., & chernyaev, a.p. (2005) problems of optimal consumption management in households, dynamics of heterogeneous systems, 9, moscow: isa ras, 212–229. 22. guriev, s.m., & pospelov, i.g. (1994) model of general equilibrium of economy of transition period, mathematical modeling., 6 (2), 3–21. 23. guriev, s.m. (1994) model of formation of savings and demand for money, mathematical modeling., 6 (7), 15–40. 24. ramsey, f.p. (1928) a mathematical theory of saving, economic journal, 38 (152), 543– 559. 25. zamulin, o.a. & sonin, k.i. (2019) economic growth of 2018 and lessons for russia, questions of economic, 1, 11–36. 26. romer, p. (1987) growth based on increasing returns due to specialization, american economic review, 77, 56–62. 27. nordhaus, w. (2017) integrated assessment models of climate change, nber reporter, 3, 16–20. 28. arutyunov, a.v., zhukovskij, s.e., & pavlova, n.g. (2013) equilibrium price as a coincidence point of two mappings, comput. math. math. phys., 53 (2), 158–169. 29. arutyunov, a.v., pavlova, n.g., & shananin, a.a. (2018) new conditions for the existence of equilibrium prices, yugosl. j. oper. res., 28 (1), 59–77. 30. pavlova n.g. (2019) study of the continuous-time open dynamic leontief model as a linear dynamical control system, diff. equations, 55 (1), 113–119. 31. pavlova n.g. (2018) necessary conditions for closedness of the technology set in dynamical leontief model, proc. 11th int. conf.: management large-scale system develop., moscow, russia: ieee. copyright © 2020 assa. adv syst sci appl (2020) introduction the extended harrod-domar model the solow model optimal consumption in the extended harrod-domar model conclusion microsoft word 1155 article text, copyedited.doc adv syst sci appl 2021; 04; 57-64 published online at https://ijassa.ipu.ru. the multidimensional network models method of developing discrete microfluidics andrey v. balabanov 1*, asim m. kasimov 2 v. a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia 1) e-mail: fca07@mail.ru 2) e-mail: kasimov@ipu.ru abstract: the performances of the microfluidic units may vary within a wide range depending on the design, even if the operational geometry remains the same. in order to obtain the most effective results of the designing, it is reasonable to use the formal procedures of systematizing possible variants of the design, analyzing these variants, and selecting the best design variant based on the preset criteria. in this context, the problem of creating the effective formal methods of designing the microfluidics is topical. the main purpose of this paper is to propose a decision of this problem. to solve the problem mentioned above, the authors devised a method of analyzing and synthesizing the designs based on the multidimensional network models (mnmmethod). the mnm-method provides the possibility of generating and in-computable-formanalyzing qualitatively different variants of the design on the principle of the unity of the geometrical, structural-hierarchical, and functional characteristics, as well as the possibility of selecting the best design variant on the basis of the preset criteria by means of analyzing both the structure and features of the multidimensional networks. keywords: multidimensional network model, design analysis and synthesis, microfluidics, reserve control system 1. introduction the experimentally validated high performances of microfluidic control elements, in combination with the small sizes, in large part, set the trend for the implementation of microfluidic devices in the form of integrated circuit units – monolithic (indivisible) designs, without using discrete components. the integrated form of the control microfluidics provides possibilities to avoid (or to significantly minimize) mounted interconnections, reduce the energy loss in transmitting the data, and provide the high performances like those of the discrete microfluidic elements. therefore, the main directions of developing the control microfluidics, at present, are improvement of the weight-size parameters as well as transition to the integrated technology. to effectively develop these trends it is necessary to bring the gap between, at the one hand, the conceptual, schematic and algorithmic decisions, and, at the other hand, the design implementation. as a solution, this paper represents the formal method of designing of the microfluidics, which is based on the multidimensional network models – the mnm-method. the mnm-method of designing the microfluidics is a new formal method of the design analysis and synthesis, with the effectiveness proved by practice. the method uses a new kind of the generalized models – the multidimensional network models – that are based on the ‘structure class’ entity. in order to design the microfluidics, the formal definitions of the geometrical, structural-hierarchical, and functional structure classes have been developed. * corresponding author: fca07@mail.ru 58 a.v. balabanov, a.m. kasimov copyright ©2021 assa. adv. in systems science and appl. (2021) in this paper, the key stages of applying the mnm-method are represented by the example of designing the microfluidic generator of the 100-µm feature size. this is very important from the practical viewpoint, because the microfluidics is increasingly used to create high-technology products. for instance, the microfluidics provides potential to create promising non-electric reserve control systems (rcs). the prospectivity of rcs is mainly determined by the resistance of the microfluidics to multiple destabilizing factors resulting in failures of electronics (for instance, radioactive and corpuscular emissions, electromagnetic emission, high temperatures, etc.). the microfluidics has its own approaches to generating, saving, transforming, and transmitting the data. therefore, the microfluidics allows new original cybernetic systems to be developed by means of the non-electronic element base. at the present time, the fields taking advantage of the microfluidics also are microanalytics, micromechanics, biotechnology, bioengineering and other complex scientific areas. this paper represents one of the possible applications of the mnm-method – designing the microfluidics. in general, the field of application is wider than the above one. however, not to disturb the monographic composition of the paper, no alternative was considered within the research undertaken. therefore, the authors do not declare the universality of the mnm-method, despite its large potential to be used as a development tool. 2. background there are a number of the researches devoted different aspects of creating the microfluidics. these are the schematic development, the investigation of microfluidics features, the modelling, and the manufacturing. in this context, what should be noted firstly are [2-11]. however, having performed the literature overview, the authors revealed no formal methods to aid the microfluidics designing. that the microfluidics has no formal tools to be designed is a circumstance that has been inhibiting and would have inhibited the development in this field – if this had been ignored. this paper is an attempt to make the microfluidics designing both more comprehensive and more scientific-proved, at least, in relation to the devices similar those represented bellow in order to give an instance of the mnm-designing. 3. mnm-method analyzing and synthesizing the microfluidic design, it is necessary to take into consideration the set of the possible variants of this design in order to select the best design structure based on the requirements for its characteristics. the mnm-method supposes that the main origin of generating the implementation set is the decomposition of the characteristics into the functional, geometrical, and structurehierarchical classes, with these classes systematized in the form of mnms. identifying the classes within the microfluidics (mf) is performed by means of the following formal procedures of the decomposition: (3.1) the procedure of the geometrical decomposition ( ) is represented by (3.1) and allows the set of the geometrical implementations of the design to be generated. the result of is the set of the structure classes, such that for any pare of the instances ( ) of any of these classes, the value areas of the geometrical functionals ( ) of the instances can differ by up to . the completion of the procedure results in the set of structure classes defining the geometrical structure of the microfluidic unit being designed: { } ( ) ( ) ( ) g ib g ia g i g ib g ia g m i i gg fefessssmfo d£-î"® = ,,|: 1 go go i gs ib g ia g ss , ib g ia g ff , gd go the multidimensional network models method of developing discrete… 59 copyright ©2021 assa. adv. in systems science and appl. (2021) (3.2) the formula (3.2) reflects the procedure of the structure-hierarchical decomposition. the result of is the set of the structure classes, such that for any pare of the instances ( ) of any of these classes, the definition domains of the geometrical functionals ( ) of the instances can differ by up to . thus, setting different requirements for , it is possible to obtain qualitatively different sets of the structure classes. for instance, the structure-hierarchical decomposition can be performed by the criteria of the standardization, the product kind (assembly, detail, kit, and complex), and so on. it is the requirement setting that is the main origin of formation of a variant set of the hierarchical structure of the microfluidics: (3.3) the procedure of the functional decomposition is described by (3.3). the result of is the set of structure classes, such that for any pare of the instances ( ) of any of these classes, the definition domains and the value areas of the geometrical functionals ( ) of the instances can differ by up to and , correspondingly. the functional decomposition provides formation of a set of the classes defining the functional structure of the microfluidics. mnms are built on the basis of the sets of the structure classes resulting from the decomposition procedures. to design a microfluidic unit with required parameters it is necessary to systematize the sets of the instances of the structure classes by means of mnms. these models have the i, j, and k dimensions. the subranges of i, j, and k depend on the cardinalities of the classes sets obtained as the result of decomposition. for example, let the two functional classes would be generated. then, the number of i-subranges is equal to two (i.e. one has and ). thus, the mnm dimensions correspond to the classes sets. given below is the case of mnm-designing involving the i, j, and k simple (single-member) subranges. in general, the number of the dimension subranges may be varied, and it is chosen depending on the complexity of both analyzing and synthesizing the design being created. 4. practice in this section, the discrete-microfluidics designing based on mnms is considered by the example of the microfluidic three-stage generator. at the first stage of the designing, it is necessary to carry out the analysis and the classification of the product requirements in terms of the preset features, with the number of these features depending on the task being solved. in relation to the generator, three groups of the features are specified: geometrical, structuralhierarchical, and functional. corresponding to the three groups of the features, the set of the structure classes (in this case, these are the geometrical, structural-hierarchical, and functional classes) is to be formed. for each of the structure classes, the fixed-cardinality set of the instances representing possible physical implementations of the class within the design to be created shall be generated and rated by the implementation costs. given in fig. 4.1 are the schematic diagram of the generator as well as the common operational geometry of its elements. the operational element of the generator under design is a microfluidic trigger. it is the trigger function that represents the main functional class. as the generator is a three-stage one, it consists of three triggers. these may be both identical and different in implementation. each original implementation of the trigger is an instance of the corresponding structure class. { } ( ) ( ) ( ) h ib h ia h i h ib h ia h n i i hh fdfdssssmfo d£-î"® = ,,|: 1 ho i hs ib h ia h ss , ib h ia h ff , hd ho { } ( ) ( ) ( ) ( ) ( )( )feib f ia ffd ib f ia f i f ib f ia f k i i ff fefefdfdssssmfo d£-ùd£-î"® = ,,|: 1 fo i fs ib f ia f ss , ib f ia f ff , fdd fed 1i 2i 60 a.v. balabanov, a.m. kasimov copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4.1. schematic diagram (a) and operational geometry (b) of three-stage microfluidic generator fig. 4.2. example of possible implementations of microfluidic trigger to form the sets of the geometrical and structural-hierarchical classes, one considers the three available implementations of the microfluidic trigger. these are represented in fig. 4.2. they are different in the operational geometries, but have the same functions (memory cell) and the same design kinds (in the form of a sheet detail). it means that t1, t2, and t3 belong to same functional and structural-hierarchical classes, but to different geometrical classes. it should be noted that this example of dividing into the classes relates to the elements, but not to the generator as a whole. however, it is the classification by the feature groups specified that is the basis to form the classes sets. at the second stage of designing, the sets of the classes instances shall be systematized in the form of a multidimensional network model (mnm). building the mnm is to arrange the instances by placing the corresponding nodes along the i, j, k axes in order of increasing the implementation costs. this process is equivalent to generating a set of the possible variants of the design. each of these design variants is represented by a unique path from the input to the output of the mnm. fig. 4.3. mnm of microfluidic generator (a) and mnm-level example (b) the multidimensional network models method of developing discrete… 61 copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4.3a demonstrates the three-dimensional mnm of the microfluidic generator, built on the basis of the previous procedures of designing. the number of the dimensions (three) of this mnm is determined by that, as an example, one class of each of the feature groups has been specified: 1) the functional class (i) is defined as the trigger function, 2) the geometrical class (j) is defined as the operational profile core, and 3) the structuralhierarchical class (k) is defined as the design kind. in general case, the number of the dimensions is determined by the sets of the structure classes, and can be varied. thus, the nodes of the mnm are the instances of the structure classes, i.e. the elements of the generator design. each of the ellipsoids includes the set of physical implementations of one of the functional-class instances. the loops of the mnm correspond to the operations of building the instances of the structure classes. in fig. 4.3a, the loops reflect the assembly operations involving the operational elements (triggers). other edges (excepting the input and output ones) of the mnm represent the transition operations between the pairs of the class instances (for example, making interconnections of the triggers). the end nodes of the transition edges shall be matched by the input and output characteristics (flow-rate, frequency, pressure, etc.). the input and output edges of the mnm reflect the complex assembly operations of mounting the terminal elements (base, covers, external interconnections) of the design. the mnm having been built, the edges shall be weighted. a required design variant within the mnm is selected by means of the procedures of analyzing the mnm features, evaluating the weights of the edges, and calculating the mnm shortest path in terms of the preset criteria. given that the generator design must provide the maximally possible frequency, the main criterion to be used in selecting would be reasonably assumed as being defined by the hydraulic resistance of the connection channels. consequently, the selecting is to evaluate the different combinations of the cross-section and the route, which are available in the mnm. as an example, fig. 4.3a explicitly shows the set of the edges, which renders the shortest path calculated by the criterion of the minimal coefficient of the hydraulic resistance of the connection channels. calculating of the resistance as depended on the design parameters is described in [1]. the determination the hydraulic resistance coefficient aids mnm-modelling the generator design in the part which is concerned with forming and estimating the physical implementations involving different sets of the mnm-levels. this part of the mnm-modelling is very important due to its large structure forming influence on the design. the nodes corresponding to the physical implementations are placed in the mnm-levels and are, in fact, reference decisions to design. as an example, given in fig. 4.3b is one of the levels of the mnm. the instances of the functional structure class (trigger) represented in fig. 4.3b are assigned as t1, t2, and t3. the figure shows that these instances are three different operating geometries (the i-dimension). each of these geometries is represented by four implementations (the j-dimension) differing in relative positions within the reference system of the generator. provided the structure-class instances are those represented in fig. 4.3b, possible variants of the generator design could be built with different sets of the mnm-levels (in this case, from one to three – the k-dimension) and, accordingly, would be different in interconnections (lengths, cross-sections, trajectories). the mnm given in fig. 4.3a encapsulates a set of the design variants of the generator, with each of the variants modelled by an original path from the input to the output of the network (a set of the nodes and the edges). the path to be selected should meet the criteria of the design quality most advantageously. the influence this path has on the performances is reflected by the weights of the mnm edges. for example, if two paths would represent two different variants of the allocation of the operational elements, the interconnections should be routed differently. the differences would influence such performances of the generator as the frequency, the branching factor, the hydrodynamic resistance, and others. therefore, the total weights of the two paths should differ in value. if some of the interconnections cannot be implemented because of design and technological factors, or the implementation results in 62 a.v. balabanov, a.m. kasimov copyright ©2021 assa. adv. in systems science and appl. (2021) inoperable state of the device, the weight of the corresponding edge is set equal to infinity (given that the best design variant is to provide the minimal total path weight). thus, the selection of the shortest path is defined as depended on the criteria of estimating the quality of the device being designed. the generator is given to be designed in order to evaluate the maximal frequency response of the operational elements. therefore, it is this response that is the main criterion of estimating the design. in general, the criteria of estimating the design quality by using the mnm-method may be both elementary and complex (taking into account technical, economic, ergonomic and other characteristics). the mnm shown in fig. 4.3a is a generalized design model of the generator. this model represents the generator in the form of the set of the structure class instances (nodes) and the operations (edges) on these instances. this set is ordered by the specified (meaningful) characteristics of the generator. the i and j dimensions define, correspondingly, functional and geometrical characteristics of the elements (triggers) at the mnm-levels of the design. the k dimension defines the number of these levels. as a result of analyzing the mnm, the following is revealed. allocation of the elements in the only planar level provides a possibility of manufacturing the generator as a detail, i.e. in the form of a monolithic design, without using assembly operations. however, the planar design does not allow the length of the interconnections between the generator stages to be reduced to minimum, which decreases the generator frequency. therefore, allocation of all the operational elements on the only plane does not provide a possibility of experimentally investigating the extreme frequency response. in addition, the use of the initial operational elements (see fig. 4.2) in the form of discrete units causes the need for assembly operations. the element designs given in fig. 4.2 provide the minimal length of the feed-back channels under the condition of the three-level generator design, with the structure class instances of this design outlined with the thick circle lines in fig. 4.3b. the figure shows that the orientations of the operational elements t1, t2 coincide. the element t3 is rotated by 180° in relation to the two others. therefore, the j-coordinates of t1 and t2 instances coincide, and the j-coordinate of t3 is different. it should be noted that fig. 4.3b demonstrates the differences in the i and j coordinates, but it does not determine allocation of the instances to the mnm-levels. as it is mentioned above, the design must be three-level. hence, the outlined instances are to be allocated to the three different levels, i.e. they are to have different k-coordinates. fig. 4.4. schematic and assembly diagrams of generator fig. 4.4 represents the schematic and assembly diagrams of the generator. the output signal of the first stage is transmitted to the input of the second stage directly, without using the multidimensional network models method of developing discrete… 63 copyright ©2021 assa. adv. in systems science and appl. (2021) mounted connections. the same is truth in relation to the output of the second stage and the input of the third stage. then, the signal is looped from the output of the third stage to the input of the first stage. the thicknesses of the matching boards are negligible. in this case, the lengths of the feed-back channels between the first and third stages are approximately equal to the thickness of the operational element (200 µm), which allows the design frequency to be considered as close to the maximum. thus, the design variant represented in fig. 4.4 is the best one from the view point of investigating the frequency response. the shortest path rendered in fig. 4.3a in the form of the set of the mnm edges and their end nodes corresponds to this design. fig. 4.5. experimental model of generator in accordance to the design variant selected by means of the mnm, the experimental model of the three-stage microfluidic generator has been created (see fig. 4.5). the mnm edge 1 in fig. 4.3a corresponds to the technological operation of mounting the top cover featuring the in/outlets in order to supply the power and transmit the data signals from the monitoring points of the first stage to a pneumoelectric converter. the further operations are concerned with mounting the three operational elements (edges 2, 4, 6) and their interconnections in the form of matching boards (edges 3, 5). the edge 7 corresponds to the technological operation of mounting the bottom cover, as well as the fixings and the output connectors. 5. conclusion the undertaken research has given rise to the development of the mnm-method regarding the discrete microfluidics. mnms are proved by the authors to be applicable to generate and compactly represent sets of the possible variants of microfluidics designs, as well as select the best design variant through calculating the shortest path within the network. using mnms, the designs may be synthesized, comprehensively analyzed, variously modified and repeatedly transformed, estimated, and ultimately properly structured. the number of the possible variants and the best design variant may be varied with requirements for the quality of the microfluidic device being created. these requirements may cover multiple aspects concerned with the creating, the exploitation, the repair, and the service. focused in this paper is the frequency response of the generator. more concretely, it is the maximal performance. this characteristic is provided by the corresponding design decision made by means of the mnm involving the structure and technology aspects. thus, the results reported in the paper prove both the reasonability and the feasibility of using the mnm-method to create devices of the discrete microfluidics. 64 a.v. balabanov, a.m. kasimov copyright ©2021 assa. adv. in systems science and appl. (2021) references 1. balabanov, a. v., kasimov, a. m. & dolgov, i. v. (2018). method of calculating design parameters of communication and throttle channels of microfluidic system, sensors & systems, 5, 39–44. 2. channon, r. b., menger, r. f., wang, w., carrão, d. b., vallabhuneni, s., et. al. (2021). design and application of a self-pumping microfluidic staggered herringbone mixer. microfluidics and nanofluidics, 25(4), 31, https://doi.org/10.1007/s10404-02102426-x. 3. hamilton, e. s., ganjalizadeh, v., wright, j. g., pitt, w. g., schmidt, h., et. al. (2019). 3d hydrodynamic focusing in microscale channels formed with two photoresist layers. microfluidics and nanofluidics, 23(11), 122, https://doi.org/10.1007/s10404-0192293-z. 4. huang, c., kuo, y., chen, y., huang, p., & lee, g. (2021). a miniaturized, dna-fet biosensor-based microfluidic system for quantification of two breast cancer biomarkers. microfluidics and nanofluidics, 25(4), 33, https://doi.org/10.1007/s10404021-02437-8. 5. lai, r. l., & huang, n. (2019). dimensional analysis and parametric studies of the microwell for particle trapping. microfluidics and nanofluidics, 23(11), 121, https://doi.org/10.1007/s10404-019-2289-8. 6. liu, b., li, m., tian, b., yang, x., & yang, j. (2018). a positive pressure-driven pdms pump for fluid handling in microfluidic chips. microfluidics and nanofluidics, 22(9), 94, https://doi.org/10.1007/s10404-018-2112-y. 7. liu, x., dong, z., zhao, q., & li, g. (2020). optimization of micromilled channels for microfluidic applications using gas-blowing-assisted pdms coating. microfluidics and nanofluidics, 24(1), 11, https://doi.org/10.1007/s10404-019-2315-x. 8. şahin, m. a., çetin, b., & özer, m. b. (2020). investigation of effect of design and operating parameters on acoustophoretic particle separation via 3d device-level simulations. microfluidics and nanofluidics, 24(1), 8, https://doi.org/10.1007/s10404019-2311-1. 9. shum, c., rosengarten, g., & zhu, y. (2017). enhancing wicking microflows in metallic foams. microfluidics and nanofluidics, 21(12), 177, https://doi.org/10.1007/s10404-017-2018-0. 10. wu, j., cui, y., xuan, s., & gong, x. (2018). 3d-printed microfluidic manipulation device integrated with magnetic array. microfluidics and nanofluidics, 22(9), 103, https://doi.org/10.1007/s10404-018-2123-8. 11. yin, z., & zou, h. (2017). multilayer patterning technique for microand nanofluidic chip fabrication. microfluidics and nanofluidics, 21(12), 174,https://doi.org/10.1007/s10404-017-2013-5. advances in systems science and applications (2013) vol.13 no.2 167-181 e-market segmentation for internet-mediated fashion brands: a conceptual framework with mean-variance consideration tsan-ming choi, pui-sze chow and jin-hui zheng institute of textiles and clothing, the hong kong polytechnic university, hung hom, kowloon, hong kong. abstract e-market segmentation takes a crucial role in marketing for fashion brands. its applications for collaborative functions such as estimating the advertisement budget, retaining customers, carrying out direct tailored marketing, and implementing dynamic pricing are essentially important and critical for their e-business. unlike the bricks-and-mortar traditional store retailers, fashion brands which operate online can keep track of the details of the online customers easily and precisely. how to make use of these customers details in segmenting the e-market for each particular function becomes an important issue. as a result, we propose and discuss in this paper a conceptual model for carrying out e-market segmentation and we focus on the areas of dynamic pricing and advertisement budget estimation. through extensive discussions with mean-variance consideration, we believe that the model can be incorporated into other existing market segmentation analyses. managerial implications are discussed. keywords fashion branding, e-market segmentation, internet marketing, meanvariance 1 introduction for fashion brands which operate solely in the traditional bricks-and-mortar retailing mode, keeping track of the buying behaviours and preferences of each specific customer is a difficult, if not impossible, task. even if the fashion brands can collect some customers’ data by observation, the results are relatively sketch [1]. unlike the bricks-and-mortar stores, fashion brands which have its online channel (we call them efbs (e-fashion-brands)) can precisely keep track of the online customers’ buying behaviours easily. details such as the surfing habit, buying preference, price sensitivity, loyalty, etc can all be recorded and estimated. obviously, information of these details can help the efbs in making good business decisions. one of the most intuitive uses of these observed customers buying data is for marketing purpose (see [2] for the discussion of the issue of marketing on the internet). market segmentation∗ , as defined in chaffey [4], is the “identification of different groups within a target market in order to develop different offerings for ∗see [3] for a recent review and discussion on a framework for market segmentation and its applications. tsan ming choi : e-market segmentation for internet-mediated fashion brands... 168 the groups”. it involves identifying segments in the market, grouping customers into segments, and targeting, positioning and developing a differential advantage over competitors [5]. it also represents a rational and precise adjustment of the products and services for customers [6]. a good market segmentation scheme implies a good analysis of the target market and an identification of the true requirements of each specific group of customers [4, 7-9]. in modern business world, market segmentation has been realized as an essential part in nearly all marketing projects [4,10-11]. according to dibb and stern [8], the literature of market segmentation mainly goes to two distinct streams. the first stream treats market segmentation as a technique in determining and studying different segments in the market. the second stream views market segmentation as an approach for an efficient allocation of resource to different segments in the market. this paper basically follows the viewpoint of the second stream. in both the academic literature [10-12] and the industrial practice [13], in general, we need two types of information for market segmentation. the first type is called the “classification variables” which include four types of variables: demographic variables (e.g. age, gender, etc), geographic variables (e.g. city, country, etc), psychographic variables (e.g. risk attitude, lifestyle, etc) and behavioral variables (e.g. brand loyalty, usage level, etc). the second type is the descriptor variables which describe each particular segment and are used to distinguish one segment from another. many classification variables would function as the descriptor variables, too. market segmentation can be a complicated process in business. for example, suppose that after adopting the conventional segmentation steps with the classification according to the geographical and demographic variables, we have obtained a number of different segments. an important question to ask is: “should we continue to break down the segments into smaller segments?” no matter the answer is “yes” or “no”, we need to have a good reason for it. in other words, a good market segmentation scheme requires a good stopping condition. it is true that there are a few rules of thumbs for decision makers to decide the number of segments and when to stop the segmentation process. some factors [13] that decision makers will bear in mind include: (1). the size of the segments must be large enough. (2). the segments must be reachable by the company’s marketing strategies (e.g. promotion, pricing, etc). (3). the segments must be relevant to the company’s products and different segments should be clearly and sufficiently different. these guidelines are important and practical but they do not provide a precise decision model for decision makers for making a wise and optimal decision for a specific collaborative function with market segmentation. 169 advances in systems science and applications (2013) vol.13 no.2 with the advance of science and technology, market segmentation can now be done by a computerized decision support system. methods which rely on data mining and artificial intelligence [14-15], advanced database system [16], operations research and management science optimization techniques [6,10,17-18] , and others [11-12,14,19-20] have been widely tested and applied. however, none of these methods dominate the literature and each of them has its pros and cons. as many recent research projects revealed and discussed, market segmentation is very important for internet-enabled e-commerce [4,16,21-22]. this also gives rise to the term “e-market segmentation” which represents the segmentation of the electronic market. in fact, good e-market segmentation is generally believed to be not only important for maintaining the company’s competitive edge, but it is also essential and crucial, for efbs in the knowledge-based economy. when we look deeper into the potential collaborative applications of the e-market segmentation scheme for efbs, we can identify several major functionality areas for it. some of them include the following: 1. estimating the advertisement budget for acquiring different groups of customers: different customers have different values for the efbs. some of them only surfed around and did not buy a single item at all; some made one purchase and then did not come back; some made repeated purchases and showed a strong sense of trust on the efb (see moe and fader [23] and betts [24] for the identifications of the four types of online shopping visits). since it can be very expensive to acquire new customers on the internet [25], the advertising strategies should be more focused. the expenses spent on advertisement for different groups (i.e. segments) of customers should be made different, too. this can be achieved by a good e-market segmentation scheme. on the other hand, the information from the estimation of advertisement budget can be useful for e-market segmentation. 2. direct marketing by providing tailor-fit services and products to the interested customers: different customers have different preferences and requirements. if an efb can provide the right product to the right customer at the right price at the right time in the right place, then a transaction results. even though it is impossible for the efb to exactly predict all of these aspects, good market segmentation can help to narrow down the error in terms of the direct marketing offers of product type, selling price, etc. it can also help in turning browsers into buyers [23-25] and improve customer relationship management [26]. 3. dynamic pricing: different consumers have different sensitivities towards prices and other attributes for the products [1,27]. some consumers are highly price-sensitive while some care more about service and reliability. surveys also show that many customers purchasing online do not shop around and they just buy from the efb which they first visit. according to a survey as reported in baker et al. [27], the percentage of consumers who buy online from the first tsan ming choi : e-market segmentation for internet-mediated fashion brands... 170 website they visit ranges from 76% to 89% for various products including cds, books and electronics. in light of this, efbs should try to attract the customers at the very beginning by bringing the products in front of the consumers when they are in need (it is mentioned in the last paragraph). at the same time, how to set the selling prices dynamically for different groups of customers so that the profit for the efb is maximized becomes vitally important. failing to do so can lead to the collapse of the business (see the example of sun country in dutta et al.[28]). in order to exercise a profit-making dynamic pricing scheme for different groups of customers, we need to segment the market properly. 4. retaining customers: for all kinds of retail businesses, there are always some loyal customers. for every loyal customer, there is a sense of trust between the retailer and herself. however, the importance of customer’s loyalty varies among different retailers. for instance, a retail store whose customers are mainly the tourists from overseas countries care relatively less about the loyalty of the customers since it is unlikely for the overseas tourists to come back again in the near future while a cosmetics retailer always wants to maintain a loyal customer base (since the loyal customers are more likely to try some other products and have repeated purchase’). since most of the customers buying online have to contribute some private data to the efb (e.g. the credit card number) but they cannot touch the product and cannot visit the store to talk with the sales staff face-to-face, they simply won’t make a purchase without trusting the efb. as a consequence, compared with the bricks-and-mortar stores, the trust and loyalty of customers is especially important for efbs. moreover, having the loyal customers can help the efbs in at least two ways. first, a loyal customer is more likely to repeat her purchase. second, a loyal customer tends to refer new customers to the efb she is loyal to. it is why reichheld and schefter [1] have proposed that “price does not rule the web; trust does”. in order to gain the trust and keep the loyalty of the customers, we have to focus on their needs. without proper market segmentation, building and keeping the loyalty of the customers becomes more difficult. from the above description, we can see that a good e-market segmentation scheme is undoubtedly crucial for the success of efbs and it can be applied to different collaborative functionality areas. however, what is a good e-market segmentation scheme for each particular collaborative function? obviously, a good e-market segmentation scheme depends on its specific targeted functionality area. for example, a good e-market segmentation scheme for dynamic pricing may not be good for the estimation of advertisement budget. as a result, we propose in this paper a decision model for carrying out e-market segmentation for different specific functionality areas. we focus our attention on two important functionality areas: advertisement budget estimation, and dynamic pricing. 171 advances in systems science and applications (2013) vol.13 no.2 these are important functions for efbs and they can be collaborated with the conventional e-market segmentation scheme. through the incorporation of the performance measure (for the collaborative function) into the e-market segmentation decision model, we can provide more precise e-market segmentation results for the marketing managers to make an optimal decision (with each particular function). the idea behind this proposed decision framework is inspired by the classical markowitz’ mean-variance theory in financial portfolio management [29] with which we quantify and control the uncertainty associated with a decision by the variance of that measure (this point will be discussed in the next section). with the illustrative examples, we demonstrate the applicability of the proposed model. the organization of the rest of this paper is as follows. we first present the mean-variance decision framework which combines the conventional e-market segmentation and the collaborative function in section 2. the detailed e-market segmentation schemes for estimation of advertisement budget and dynamic pricing are proposed in sections 3 and 4. the e-market segmentation schemes for other functionality areas are discussed in section 5. we conclude with the discussion of managerial insights in section 6. 2 mean-variance decision models before we present each particular e-market segmentation model for the specific function area, we propose in this section the basic general decision model. in performing e-market segmentation, as we mentioned earlier, it is usual that the marketing managers would make use of the demographic variables, geographic variables, psychographic variables and behavioral variables. despite the intuitive physical meanings behind these variables, some of the classifications with these variables may not be very helpful for all specific collaborative functions. for example, when the objective of a particular market segmentation project is to decide the advertisement budget (and hence the advertisement strategy), the profit that can be generated by each customer becomes a key measure. as a result, we should incorporate a measure of the profit generated by the customers in the market segmentation scheme. however, different customers can carry different profit-values to the company, a precision control rule is hence essential for building a good market segmentation decision. in the advertisement budget estimation example we mentioned above, a precision control rule can be imposed on the degree of uncertainty of the profit. thus, a mean-variance consideration with which the average profit is used as the performance measure variable and the variance of profit is applied as a precision control variable. in this paper, we call the market segmentation scheme which includes the average objective performance measure and the variance of this performance measure for a particular tsan ming choi : e-market segmentation for internet-mediated fashion brands... 172 market segmentation project the “mean-variance performance measure market segmentation scheme (mvpm)”. with mvpm, we can carry out market segmentation following a mean-variance consideration with which the segmentation is performed with a measure by the “mean” and its precision is controlled by the “variance”. since the variance is an absolute measure, we make use of the relative measure of the coefficient of variation, defined as “the standard deviation divided by the mean” as the precision measure. the general idea behind mvpm is that: the e-market segmentation scheme follows the conventional type of segmentation policy and yields a number of emarket segments. this segmentation process is treated as an initial and basic e-market segmentation. after that, depending on different collaborative functions, the manager of each function would impose an additional measure on each of the segments. systematic and precise evaluation is carried out and further segmentation or re-segmentation may be required depending on the evaluation results. by doing so, tailor-fit e-market segmentation results are provided to each collaborative function and optimal decision can hence be made. in the following sections, we outline the use of the concept of mvpm for several important collaborative functionality areas with e-market segmentation. 3 segmentation for estimation of advertisement budget the expenses companies spent on online advertisement is expected to increase in the coming years. as reported in bhatnagar and papatla [21], forrester research has estimated that the spending on online advertisement in the united states will reach us$22 billion by 2004, which is more than 8% of the total spending for advertising in the united states. in fact, acquiring customers on the internet can be very expensive. as estimated and shown in hoffman and novak [25], many efbs have spent more than us$100 to acquire a new customer and some have even spent us$500! however, the “values” of most new customers, as measured by the expected lifetime spending on the efb, are less than these advertising expenses. in fact, some consumers shopping around the internet only surf and buy nothing; some of them may only make one purchase and never come back; some may have repeated purchases and are loyal customers. as we mentioned above, since it is expensive to acquire new customers on the internet, the advertising strategies should be more focused. recalling from the well-agreed pareto rule (or called the 80-20 rule), the majority of profit is actually generated by a relatively small amount of customers. it is thus a wise decision to focus the company’s resource on promoting to the customers which can bring higher profits. furthermore, the expenses spent on advertisement for different groups (i.e. segments) of customers should be made different. this can be achieved by a good e-market segmentation scheme. in the 173 advances in systems science and applications (2013) vol.13 no.2 literature, there are a number of methods proposed to help. they include the use of the customer’s searching behavior [21] and the use of personalized advertisement [9,14]. in this section, we apply the mvpm to build a decision model for the estimation of the advertisement budget for different market segments. under our proposed mean-variance model, mvpm, the efb should incorporate the performance measure and the precision control measure for “the profit that can be generated by the customer” into the decision model for market segmentation. thus, when efbs try to evaluate the “value” of the customers purchasing online, they should investigate the e-market segments with respect to the profit generated by the customers in each of these segments with a precision control. to be specific, we have the following market segmentation decision model: “the efb first segments the customers according to the conventional segmentation scheme by, for example, the geographic and demographic variables. then, the efb checks (and/or further segments) each group with respect to the average generated profit and the variance of profit from the members of that group. the objective is to ensure that the coefficient of variation of profit, defined as the standard deviation of profit divided by the average profit, for each segment is under the efb’s precision control and the segments size is large enough. after that, the final outcome from each market segment will have an average value for the customers in that segment and this average value can be a good representative measure since its variation is under the efb’s control.” to illustrate the above statement, let us consider a simple example. suppose that an efb has classified his customers according to the conventional segmentation scheme with respect to the customers’ variables of location, gender, age and education, and under a constraint on the size of each segment. after this segmentation scheme, he has obtained different market segments. when the efb looks deeply into each obtained market segment, he can identify the profit that has been generated by each customer inside each market segment. he can then obtain the average profit and the variance of profit generated by all the members in each market segment. for a particular segment, if the coefficient of variation of profit is larger than a certain threshold (decided by the efb), the level of uncertainty of the profit generated by the members inside this market segment is too large. further market segmentation should be carried out by adding another attribute (e.g. the purchear frequency of the customers). if the coefficient of variation of profit is less than a certain threshold (decided by the efb), the market segmentation that has been done is good enough and the efb can stick with it. however, if the coefficient of variation of profit is too large but the market segments size is also relatively small, further market segmentation should not be carried out. in this case, the efb needs to reconsider carrying out the market tsan ming choi : e-market segmentation for internet-mediated fashion brands... 174 segmentation in another way with consideration of other variables at the very beginning. we summarize this approach in the following proposition. proposition 1. the e-market segmentation scheme for estimating the advertisement budget can be stopped when the coefficient of variation of profit generated by the members of that segment is less than a threshold α. if the coefficient of variation of profit is larger than α, then further segmentation or re-segmentation is needed. to give a better picture of the proposed decision model in proposition 1, let us have an illustrative numerical example below. example 1. an efb has performed a market segmentation scheme by using the conventional variables of city, gender and age and identified 18 e-market segments. suppose that this efb looks into two distinct segments: segment 1 and segment 2, where segment 1 refers to the group of customers who live in city 1, male, and aged between 22-25, and segment 2 refers to the group of customers who live in city 2, female, and aged between 18-21. the desirable minimum size of each segment is 500. for segment 1, there are 1600 customers and for segment 2, there are 850 customers. the profit generated by each one of the customers can be found from the efb’s database. when the efb calculates the average profit per head (ap), variance of profit (vp), standard deviation of profit (sdp), and coefficient of variation of profit (cvp) generated by the members of each segment, he has the following results, table 1.1 example 1 segment 1 segment 2 average profit (in $) 50 180 variance of profit (in $2) 352 802 standard deviation of profit (in $) 35 80 coefficient of variation of profit 0.70 0.44 obviously, even though the vp of the members in segment 1 is smaller than the members in segment 2, the cvp is much larger. in fact, the profit uncertainty for segment 1 is too large to be ignored. thus, if the precision threshold of the efb (α) is 0.5, then the e-market segmentation for segment 2 is good enough because its cvp is less than 0.5. however, the e-market segmentation for segment 1 is not good enough because segment 1’s cvp is larger than α. since the size of this segment is 1600 and the minimum segment size is 500, the efb can consider carrying out further segmentation on segment 1 by using another classification variable. after the e-market segmentation scheme with the precision control over the profit generated by the members of each segment, the efb can make use of the estimated average profit for each segment as an indicator to decide the 175 advances in systems science and applications (2013) vol.13 no.2 amount of budget for advertisement on each segment. in this example, the value for each customer in segment 2 is $180 and a reasonable budget (say $100 per head) for advertising towards customers in this market segment can hence be estimated. as summarized in proposition 1 and illustrated by example 1, we can see that a decision model which provided a tailor-made stopping condition for the e-market segmentation process has been proposed. as we mentioned earlier, advertisement for e-commerce can be very expensive. nowadays, on-site banner, pup-up screen, emailing, affiliate program, tv and radio commercials, etc are all popular means of advertisement. however, obviously, different means of advertisement carry different costs. as a result, it is a wise decision to decide the advertisement budget for each specific group of customers before considering the specific means of advertisement. by having the market segmentation scheme as described in proposition 1 where the members of each segment are grouped together with the consideration of the average profit generated under precision control, we can identify precisely the “value” of each member of that particular market segment. as a consequence, the efb can allocate the optimal advertisement resource to focus on the most profitable customers, and decide the most appropriate advertisement scheme for them. 4 segmentation for dynamic pricing the online consumers have different sensitivities towards price and other nonprice factors. some consumers are highly price-sensitive and they like to use the shop bots for finding the efbs which offer the lowest prices while some care more about service, reliability and trust. as a result, how to set the right selling prices dynamically for different groups of customers so that the efb’s profit is maximized becomes very important and, in fact, crucial (the failure stories due to the lack of good pricing policy can be found in [28]). one of the effective ways for dynamic pricing is to carry out dynamic price testing. the idea of the dynamic price testing [27,30] is that: from changing the listed price showing on an efb’s website and keeping track of the customers’ purchasing rates at that price, the efb can know the expected profitability of each listed price. for example, when the efb sets the product’s listed price as $10, it is observed that 2 out of 10 visitors will buy the product; when the efb changes and tests the price at $9.5, it is found that 3 out of 10 visitors will buy the product. the purchasing rates are 20% and 30% for $10 and $9.5, respectively. the efb can then decide whether $9.5 is better than $10 or not based on the corresponding observed purchasing rates and profit margins. however, in order to exercise an effective dynamic price testing scheme (and hence a good dynamic pricing), the efb needs to segment the e-market properly. from the sales record of the customers for that product tsan ming choi : e-market segmentation for internet-mediated fashion brands... 176 (or a closely related product if the data for that product is not sufficient or available) in the past, the efb can keep track of the exact purchasing price of each customer. as a result, the efb can first segment the e-market for that product according to the conventional approach by the customers’ age, gender, etc. after that, the efb can check the average purchasing price and the variation of the purchasing price per each segment following the concept from mvpm. similar to the proposed method for the segmentation for advertisement budget estimation, we have the following model: “from the database of the customers, the efb first segments the customers according to the conventional segmentation scheme by, for example, the geographic and demographic variables. then, the efb further segments each group with respect to the average purchasing price and the variance of purchasing price from the customers in that group. the objective is to ensure that the coefficient of variation of the purchasing price, defined as the standard deviation of purchasing price divided by the average purchasing price, for each segment is under the efb’s precision control and the segments size is large enough.” with the above method, the efb can effectively identify the group of customers with a specific average purchasing price. after that, a tailor-made dynamic price testing scheme can be arranged for that particular market segment. it is thus more effective than performing the dynamic price test blindly. we summarize this proposed method in proposition 2 below. proposition 2. the e-market segmentation scheme for dynamic price testing (and hence dynamic pricing) can be stopped when the coefficient of variation of the purchasing price of the members in that segment is less than a threshold θ. if the coefficient of variation of profit is larger than θ, then further segmentation or re-segmentation is needed. example 2 below gives an illustrative numerical example for the proposed emarket segmentation model presented in proposition 2. example 2. an efb has performed conventional market segmentation (by using the variables such as age, city, gender, etc) for the customers of a specific product and has identified 20 e-market segments. suppose that this efb looks into two distinct segments: segment a and segment b. the desirable minimum size of each segment is 500. for segment a, there are 1300 customers and for segment b, there are 1500 customers. when the efb calculates the average purchasing price, variance of purchasing price, standard deviation of purchasing price, and coefficient of variation of purchasing price generated by the members of each segment, he has an khown in table 1, from table 1, the coefficient of variation of the purchasing price for customers in segment b is much larger than the customers in segment a. suppose that the efb has set the precision threshold θ(for the coefficient of variation of the pur177 advances in systems science and applications (2013) vol.13 no.2 table 1.2 example 2 segment a segment b average purchasing price (in $) 100 90 variance of purchasing price (in $2) 252 502 standard deviation of purchasing price (in $) 25 50 coefficient of variation of purchasing price 0.25 0.56 chasing price) to be 0.5. by proposition 2, the e-market segmentation result for segment a is good enough because the coefficient of variation of the purchasing price is less than θ. thus, for consumers in segment a, the efb can carry out the dynamic price test with a reference price of $100 and the dynamic testing prices can be set within a reasonable range. on the other hand, the e-market segmentation for segment b is not good enough because segment b’s coefficient of variation of the purchasing price is larger than θ. since the size of this segment is 1500 and the minimum segment size is 500, the efb can consider further segmenting segment b by using another classification variable. 5 segmentation for retaining customers, direct marketing and others in sections 3 and 4, we have discussed the e-market segmentation schemes, following the concept of mvpm, for two important collaborative functions for efbs. we will discuss more functionality areas where mvpm can be applied for emarket segmentation in this section. as we mentioned earlier, trust and customers’ loyalty are two important issues for efbs doing business on the internet. without proper e-market segmentation, building and keeping the loyalty of the online customers becomes more difficult. as a result, when the efb performs the e-market segmentation, it is important for him to bear in mind that he has to be able to identify precisely the need of the customers inside that market segment and be focused [1]. this objective follows exactly the mvpm where the need of the customers inside each e-market segment refers to the average measure of that need and the precision for the understanding of this need in the corresponding e-market segment is reflected by the coefficient of variation of that need measure. this is what the mvpm captures. as an example, suppose the efb would like to retain the customers by providing a membership system. in order to attract the customers, the efb would like to provide bonus points for the customers who join the membership and make some purchases. obviously, in order to attract more customers to be members, the amount of the bonus points and the amount of required purchases should be offered differently to customers in different segments. the concept of mvpm can hence be applied for providing a measure for this purpose. tsan ming choi : e-market segmentation for internet-mediated fashion brands... 178 similarly, for an efb who wants to provide the right product and service to the right customer for the right price at the right time in the right place with direct marketing strategy [31-32] , the mvpm also works where it provides a mechanism for e-market segmentation which helps to narrow down the error in terms of the product type, selling price, etc. for instance, the efb can make use of the data on the products bought by the customers in the past to help in segmenting the customers and measuring the customers likelihood of buying the product that will be direct marketed. 6 conclusion and managerial implications we have proposed in this paper a conceptual decision model, called the “meanvariance performance measure market segmentation scheme (mvpm)”, for emarket segmentation for different specific collaborative functions. as we all know, the traditional market segmentation relies on the classification variables like the demographic, geographic, psychographic and behavioral variables. however, in spite of the intuition behind the classification by using these variables, researchers have questioned the reliability of the available market segmentation techniques (e.g. [8]). furthermore, the classifications with these conventional variables may not be very helpful for a specific collaborative function. as a result, we propose to incorporate the collaborative function’s performance measure and its precision into the e-market segmentation scheme for effective decision making by the managers of the corresponding functions. a mean-variance consideration with which the average or expected performance measure (e.g. the average profit) is used as the performance measure variable and the degree of variation of the performance measure (e.g. the coefficient of variation of profit) is applied as a precision control variable. this market segmentation scheme is thus called the mean-variance performance measure market segmentation scheme (mvpm). with mvpm, efbs can carry out e-market segmentation for each particular collaborative functionality area following a mean-variance consideration. we have discussed the use of mvpm for efbs with the estimation of advertisement budget, dynamic pricing, retaining customers and direct marketing. since brand managers are involved with the resource allocation and decision making for the efbs, effective and precise market segmentation schemes can provide them with good assistance in making a sound and optimal decision. under mvpm, the e-market segmentation decision is under a control on the precision of the specific performance measure. as a consequence, better resource allocation decisions for different collaborative operations can be made by the fashion brand managers optimally and precisely. we can thus view mvpm as an uncertainty control model which targets at reducing and constraining the degree of uncertain179 advances in systems science and applications (2013) vol.13 no.2 ty associated with the performance measure. we believe that mvpm is especially important for e-commerce (and mobile commerce) because the e-market is highly volatile with a large amount of uncertainty sources and customers details can easily be recorded and analyzed. notice that, even though we focus on the use of mvpm for internet enabled e-market segmentation schemes, the conceptual framework of mvpm can actually be applied to the general market segmentation systems. the model of mvpm can also be implemented into an intelligent computerized decision support system which automatically helps managers in making wise and scientifically sound decisions. acknowledgements the authors are indebted to the anonymous referees for their critical comments (on the earlier version of the paper) which improve this paper substantially. this paper is supported in part by the hong kong polytechnic university with research funding account of g-yj23. references [1] reichheld f and schefter p. (2000), e-loyalty: your secret weapon on the web, harvard business review, vol.78, no.4, pp.105-113. [2] kiang m.y, raghu t.s and shang k.h.m. (2000), “marketing on the internet who can benefit from an online marketing approach?”, decision support systems, vol.27, pp.383-393. [3] liu y, kiang m and brusco m. (2012), “a unified framework for market segmentation and its applications”, expert systems with applications, vol.39, pp.10292-10302. [4] chaffey d. (2002), e-business and e-commerce management, 1st edition, prentice hall. [5] dibb s, simin l, pride w and ferrell o. (2000), marketing concepts and strategies, 4st edition, houghton mifflin, boston, ma. [6] smith w.r. (1956), “product differentiation and market segmentation as alternative product strategies”, journal of marketing, vol.20, no.7, pp.3-8. [7] gary r.k. (2001), “new dimension of internet buyer behavior: strategic marketing implications”, the proceedings of the first international conference on electronic business, hong kong. [8] dibb s and stern p. (1995), “questioning the reliability of market segmentation techniques”, omega, vol.23, no.6, pp.625-636. tsan ming choi : e-market segmentation for internet-mediated fashion brands... 180 [9] raghu t.s, kannan p.k, rao h.r and whinston a.b. (2001), “dynamic profiling of consumers for customized offerings over the internet: a model and analysis”, decision support systems, vol.32, pp.117-134. [10] haley r.i. (1968), benefit segmentation: “a decision-oriented research tool”, journal of marketing, vol.32, pp.30-35. [11] johnson r.m. (1971), “market segmentation: a strategic management tool”, journal of marketing research, vol.8, pp.13-18. [12] darden w.r and perreault w.d. (1977), “classification for market segmentation: an improved linear model for solving problems of arbitrary origin”, management science, vol.24, no.3, pp.259-271. [13] dss research, understanding market segmentation, downloadable from http://www.dssresearch.com/marketsegment/library/segment/understanding.asp [14] kim j.w, lee b.h, shaw m.j, chang h.l and nelson m, (2001), “application of decision-tree induction techniques to personalized advertisements on internet storefronts”, international journal of electronic commerce, vol.5, no.3, pp.45-62. [15] shaw m.j, subramaniam c, tan g.w and welge m.e. (2001), “knowledge management and data mining for marketing”, decision support systems, vol.31, pp.127-137. [16] montgomery l. (2001), “applying quantitative marketing techniques”, interfaces, vol.31, no.2, pp.90-108. [17] brusco m.j, cradit j.d and stahl s. (2002), “a simulated annealing heuristic for a bicriterion partitioning problem in market segmentation”, journal of marketing research, vol.39, pp.99-109. [18] novak t.p, hoffman d.l and yung y.f. (2000), “measuring the customer experience in online environments: a structural modeling approach”, marketing science, vol.19, no.1, pp.22-42. [19] moorthy k.s. (1984), “market segmentation, self-selection, and product line design”, marketing science, vol.3, no.4, pp.288-307. [20] winter f.w. (1989), “market segmentation using modeling and lotus 1-23”, interfaces, vol.19, no.6, pp.83-94. 181 advances in systems science and applications (2013) vol.13 no.2 [21] bhatnagar a, and papatla p. (2001), “identifying locations for targeted advertising on the internet”, international journal of electronic commerce, vol.5, no.3, pp.23-44. [22] turban e, lee j, king d and chung h.m. (2012), electronic commerce a managerial perspective, 7th edition, pearson. [23] moe w. w and fader p.s. (2001), “dynamic conversion behaviour at ecommerce sites”, management science, vol.50, pp.326-335. [24] betts m. (2001), “turning browsers into buyers”, mit sloan management review, winter, pp.8-9. [25] hoffman l and novak t.p. (2000), “how to acquire customers on the web”, harvard business review, vol.78, no.3, pp.179-188. [26] chatranon a, chen j.c.h, chong p.p and chen y.s. (2001), “customer relationship management (crm) and e-commerce”, the proceedings of the first international conference on electronic business, hong kong. [27] baker w, marn m, and zawada c. (2001. february), “price smarter on the net”, harvard business review, vol.79, no.2, pp.122-127. [28] dutta s, bergen m, levy d, ritson m and zbarachi m. (2002), “pricing as a strategic capability”, mit sloan management review, spring, pp.61-66. [29] markowitz h.m. (1959), portfolio selection: efficient diversification of investment, new york: john wiley & sons. [30] choi t.m, chow p.s and xiao t. (2012), “electronic price-testing scheme for fashion retailing with information updating”, international journal of production economics, vol.140, pp.396-406. [31] kenny d and marshall j.f. (2000), contextual marketing: the real business of the internet, harvard business review, vol.78, no.6 pp.119-125. [32] rust r.t and lemon k.n. (2001), “e-service and the consumer”, international journal of electronic commerce, vol.5, no.3, pp.85-101. corresponding author tsan-ming choi can be contacted at jason.choi@polyu.edu.hk. advances in systems science and application(2015) vol.15 no.2 186-192 systemic complexity of empires and the globalization hermínio duarteramos new university of lisbon, lisbon, portugal abstract generally, very wide systems are complex, and they afford particular methodology implementations to obtain reasonable operation outputs. here we are looking for a strategy to rational interpretations on empire data collected from different civilizations in the universal history, in order to understand complex globalization trends as it occurs in our days. keywords systemic complexity; empire system; globalisation; worldwide integrated system 1 introduction the concept of complexity is linked to real systems through the integration of multiple interactive components.from such unitary structure emerges a systemic output called telonomy, which is significant for the system intentionality and for the human interpretation of the observed reality [1]. the system operation response may be difficult to be interpreted, and so it will be impossible to warranty a sure comprehensive operation prediction for the future.this reasoning conducted us to define complexity on the basis of difficulties we experiment to point out the essential features of the system functionalities. these systemic essentials are the following: acrony or structural components, axony or interactions among components, aquadry or real and virtual boundaries of the functional set, and adaptacy or the evolution system to the optimum working point in order to get the best telonomy, according to the system intentionality. whenever we don’t know the right conditions of one or more systemic essentials, we can say that the system is complex (even if we recognise it is simple), and otherwise it will be simplex (even we say it is complicated).in current human lives, we can handle trivial and sophisticated cases of simplexity, using the science and art or the societal common knowledge, but instances of complexity do offer some unknown process data. 2 worldwide societal systems a system in any society includes natural and technological components, which are inserted in all normal social activities. each worldwide system integrates human actions and natural objects or signals and artificial manufactured products. human and technology and universe phenomenology are three classes of general sets operating partial subsystems within societal systems. we know that very well in our current life, for instance it happens in remote communication systems advances in systems science and application(2015) vol.15 no.2 187 using the trivial internet. we can see it in everyday activity, and also regarding the historical evolution in all geographic regions until the contemporary globalization perspective of the world. in fact, today we have the possibility to interact in very short time at long distances as if we were remotely present. nevertheless, such remote action is very ancient. primitive people did migrate from africa to asia and europe and to america, they moved away with some adventure purposes and building several social structures to better profit from natural goods, organizing different human communities to live according to their own collective wishes. some of them tried to exert forced influences imposing empire charges impelled by ambitious and audacious adventurers, using newer power materials (as weapons or animal powers) to submit neighbour populations to their rules. the empires did grow up, and fell down after terrible and declining experiences. however nations hope to be free, living according to their cultures and developing freely own traditions. after the 15th century, maritime navigations over the atlantic ocean did carry europeans to the indic ocean and later to the pacific, and the entire world became known for all peoples. the time did pass along, and the 20th century developed sufficient scientific knowledge and new technologies to amplify some relations between countries over the world to a very high extent. the basic idea states that all organized societies will change global interactions by wide networks of relationships. is that the real trend for the future? what are the necessary global conditions to live happy in peace? how to defeat any possible threat by an eventual empire outburst? these questions inspire us to analyse the complexity of global interlinked systems over the world in comparison to local forced interconnected countries composing empires. between empires and the globalization we have colonization systems, with features from both basic social system types. for that we endeavour to understand evident systemic essentials associated to empires and the globalization. 3 what is an empire? typically any empire is an extensive territory governed by a single supreme authority, subjugating several peoples and different cultures to a dominant power. analysing the ancient history, since four millenniums behind, mostly at east and far east territories, we note the birth of several empires, as qin and han in china, cyrus in persian, alexander the great, roman empire, maurya and gupta in india, genghis khan, ming, and inca and aztec in america, the ottoman empire and others [2-4]. from such information we can extract some common features to identify the general empire concept under a systemic standpoint, neglecting political or so188 hermínio duarteramos: systemic complexity of empires and the globalization cial references. an empire sets down in forced blocs of populations with different cultures, using dominant strategies and a serial of tactic adaptations in order to harmonise suitable rules over the space and the time. imposing authority to subjugated people the empires use powerful technologies (transports, weapons, new inventions) and human capabilities (brave warriors, smart advisers, loyal governors). the empire evolution always strengthens higher hegemony levels, forcing social value uniformities, including a vehicular language, and destroying culture diversities. nevertheless, the empire implementation requires a convenient intellectual background warranting somewhat social stability. 4 what is the globalization? much more than world relations, the globalization do fall upon the free integration of several organizations inside many countries. each organization implements proper structures and pursues its own objectives. it is a singular societal organization within a general independent sovereignty endeavouring social aims without external impositions, but it grows by mutual acceptable requirements. today we observe a tremendous increase on a commercial globalization, practicing business everywhere through international corporations (are they only a powerful finance globalization?) operating as national enterprises in many independent countries and working in special industrial sectors or as services providers [5]. so, we detect a few trends on the globalization reality, extending actions over all human activities (science, art, literature) [6] and social environments (university, finance, energy) [7].in such global processes, the globalization will induce value uniformities when the time elapses, but without programmed forcibleness, preserving culture diversities by local behaviour adaptations to general imported influences. 5 systemic essentials of an empire history data gathered from asian empires, which has been created and died several thousands of centuries ago, can give us important information about their main features as systems. we observe similar properties in european empires and also in primitive american civilizations. the following description summarises empire systemic essentials as we did notice it. acrony: the composition structure of an empire depends on the occupied planetary space and on living people inside their boundaries, showing an authority expansion to contiguous spaces by subjugation of resident populations to external dictated rules. an empire has always certain heterogeneity among subsystems, and each one advances in systems science and application(2015) vol.15 no.2 189 exhibits a particular culture following their old traditions. that is to say the empire acrony can be definite, and it is not very much homogeneous. axony: distinct people do live integrated and in peace only if the empire system reveals an interactivity to aim at getting a minimum harmony, pursuing same accepted ways of live. but the empire domination compels behaviours against people desires, otherwise the dominator will be immediately defeated. this means the empire axony is not always well known being incomplete. aquadry: each partial country belonging to an empire tries to maintain its community frontier according tradition although it will be not clear. however the empire external boundaries are perfectly defended, separating the empire territory to neighbour spaces, and may be extended if the empire decides to struggle for greater dimensions. for that reason we can assert the empire aquadry may be certain in some history periods. adaptacy: the law code of an empire must consolidate a dominant culture, forcing it to be followed everywhere by the integrated people. guide lines to control all subsystems claim to attenuate culture differences, imposing a common paradigm to general behaviours even against people reactions. therefore empire adaptacy is mostly indeterminate. telonomy: some outputs from the empire process will go to the external world, but the majority results emerging from the organic activity are self oriented to maintain the auto-authority. the empire system must assure the sovereignty and independence, spending a lot of resources to continue alive, and the external actions try to create an image of internal paradise, pretending to influence neighbours to adhere to the expanded project, even under several threats. but the empire telonomy is almost isolated from other wide systems in the world, and outputs are not very much accurate due to some secret and not transparent actions. 6 systemic essentials of the globalization basic globalization features are very different from empire ones, because the global aim is quite distinct from the empire intentionality. as a corollary we note the consensual existence of multiple interconnected systems and so they can design an intricate societal network over the world. acrony: the organization structure of a global system is homogeneous in each functional subsystem. partial components are distributed over separated regions in the 190 hermínio duarteramos: systemic complexity of empires and the globalization world, instead of adjacent territories. following we can tell without any doubt that the global acrony is well definite. axony: all parts of a global system interact in different spaces, and they are perfectly controlled converging to the same global intentionality. this network paradigm means a fundamental feature of global systems. as a consequence the global axony is considered complete in general cases. aquadry: a global system has necessarily a virtual frontier, because concrete boundaries of all partial subsystems are distant one from others, and the global system boundaries are not physical, being distributed over many countries. this means that a global boundary may be not rigid, depending very much from exogenous variables, following political or culture environments in each country of implementation. we affirm that the global aquadry may be not exactly known and thereafter will be uncertain. adaptacy: operating subsystems optimise global work functionalities in an easy way, owing to the structure homogeneity and to the interaction accuracy. in fact, the network operation by agent interconnections will make easier to optimize working points. consequently global adaptacy appears to be perfectly determinate. telonomy: multiple outputs from a global system may be active in several subsystems to various geographical regions, giving all capabilities to interact in the internal network at remote locations. the result is a reasonable efficient control over the expected system intentionality. finally, we may say the global telonomy interacts accurately with many environments. 7 comparing systemic essentials although brief, the foregoing analysis reveals how far the systemic complexity of empires and globalization is to be considered. to better see the problem we summarize the most important features of the systemic essentials in the comparison table 1. the table teaches us that the empire system is a complex system (and also very complicated) denoting a very difficult scientific approach, giving the incompleteness of the axony and the indetermination of the adaptacy, although the acrony definiteness and the aquadry certainty can ease necessary specifications for a right system description. this means that in general empire systems have a 2nd or 3rd degree of complexity. on counterpoint, a concrete globalization represents a complex system (even advances in systems science and application(2015) vol.15 no.2 191 it could be very complicated) with a 1st or 2nd degree of complexity, seeing that we can get higher quality for scientific approaches, giving the definiteness of the acrony, the completeness of the axony and the determination of the adaptacy, although somewhat aquadry uncertainty may difficult system specifications. table 1 systemic essential comparison systemic empire system global system essential acrony definite contiguous components definite separated components axony incomplete forced interaction complete free programmed interaction aquadry certain rigid external boundary uncertain soft internal boundaries adapatacy indeterminate the operation depends on many unknown factors determinate the operation depends on many known factors telonomy accurate multiple internal on many unknown few external outputs inaccurate multiple external outputs and few internal outputs 8 conclusions a worldwide process requires induced perspectives from very high observation positions, and paradoxically we notice much better a global system far away from their details. in fact, we see better the earth from the sky, because the distance allows us to have a clearer vision on the part under global observation. really, we get a finer understanding about ourselves if we go outside from us and look at our profound inner states. any global system must be regarded from outside interpreting external outputs of its telonomy, but we must know inner components and interactions to understand the best operation modes inside their organization. an example can be the global finance system and its consequences for the natural world, requiring a suitable global control to eliminate perversion and inequity, and avoiding dangerous austerities [6]. all this is very complex to face it. but all this must be rationally workable, processing it by science and ethics in order to get a superior knowledge on reality to better survive and being happy. references [1] hermínio duarteramos. (2008), “hard and soft systems intentionality”, 7 th congress of the ues, res-systemica, afscet, lisbon, vol.7. [2] edward gibbon. (1998), the history of the decline and fall of the roman empire,wordsworth classics of world literature , london. [3] laurent testot (ed.). (2013), sigma: a new open economy model for policy analysis,les grands dossiers des sciences humaines, paris. 192 hermínio duarteramos: systemic complexity of empires and the globalization [4] dietmar pieper (ed.). (2014),another world: the inca, maya, aztecs, the mysterious kingdoms, der spiegel geschichte, nr.2, hamburg. [5] jacques adda. (1966), economic globalization, la découverte, paris. [6] zygmunt bauman. (1998), globalization: the human consequences, blackwell publishers, cambridge/oxford. [7] world energy council. (2014), 2014 world energy issues monitor, world energy council, london. [8] jean-pierre warnier. (1999),cultural globalization, éditions la découverte et syros, paris. corresponding author hermínio duarteramos can be contacted at: hduarteramos@gmail.com microsoft word 11-wu jianhua.doc advances in systems science and applications (2010), vol.10, no.1 67-72 issn 1078-6236 international institute for general systems studies, inc. characteristics of item replacement in weibull distribution jianhua wu, dexin tao and hongxiang li wuhan university of technology, wuhan, china abstract the aim of this paper is to study preventive replacement in order to increase system’s mtbf by replacing item following the weibull distribution. here, we discuss the periodic preventive replacement and random preventive replacement as preventive replacement. according to item preventive replacement following weibull distribution, based on the mtbf evaluation of item to study the characteristics of item replacement. keywords weibull distribution preventive replacement mtbf 1. introduction outline and background of the preventive replacement theory notation used in this paper. f(t): the failure distribution function of replacement item. dt/)t(df)t(f);t(f1:)t(f =− g(t): the distribution function of the preventive replacement dt/)t(dg)t(g);t(g1:)t(g =− : first moment or mtbf under the preventive replacement : n moment μ: the replacement factor of preventive replacement distribution function g(t)=1-exp(-μt) t : time interval of preventive replacement г(·): gamma function; г(·,·): in-complete gamma function; δ(·): derta function β: shape parameter of weibull distribution η: scale parameter of weibull distribution cv: coefficient of variation τ: reliability improvement rate now let’s gather up the result of fundamental and general theory about preventive replacement (1) first moment or mtbf under preventive replacement mtbf or the first moment is given as: 0 0 ( ) ( ) 1 f(t)g(t)dt f t g t dt t ∞ ∞< >= − ∫ ∫ (1) (2) second moment and variance under preventive replacement the second moment is given as: 68 wu: characteristics of item replacement in weibull distribution 2 0 0 2 0 0 0 2 0 2{1 f( ) ( ) }{ tg( ) ( ) } (1 ( ) ( ) ) 2{ g( ) ( ) }{ tf( ) ( ) } (1 f( ) ( ) ) t g t dt t f t dt t f t g t dt t f t dt t g t dt t g t dt ∞ ∞ ∞ ∞ ∞ ∞ − < >= + − − ∫ ∫ ∫ ∫ ∫ ∫ (2) thus the varianceδ2 is easily got fromδ 2 =- 2 2. periodic preventive replacement and random preventive replacement the periodic preventive replacement is the most general way in the preventive replacement, so that we study on this case, and then take up the random preventive replacement as extreme example 2.1 the average p and the 2nd moment p of the periodic preventive replacement if theb periodic preventive replacement is done at time t, nonpreventive replacement )t(g is as figure 1, then 1 0 ( ) 0 t t g t theother ≤ ≤⎧ = ⎨ ⎩ (3) g(t)=δ(t-t) (4) figure 1 consequently, the molecule and the denominator of (eq.1) is described as following ∫∫ ∫∫ ∞∞ ∞ =−δ= = 00 t 00 )t(fdt)tt()t(fdt)t(g)t(f dt)t(fdt)t(f)t(g and )t(f)t(f1dt)t(g)t(f1 0 =−=− ∫ ∞ therefore the 1st moment p 0 ( ) ( ) t p f t dt t f t < > = ∫ (5) on the other hand, the 2nd moment p can be also obtained in the same way. 2 0 0 2 2 ( ) 2 ( ) ( ) ( ) ( ) t t p tf t dt t f t f t dt t f t f t < > = +∫ ∫ (6) so the variance of preventive replacement can be obtained through formula (eq.5) and (eq.6). advances in systems science and applications (2010), vol.10, no.1 69 2.2 the 1st moment and the 2nd moment of the random replacement in the case of the random preventive replacement, as maintenance function we put g(t)=1-exp(-μt) into formula (eq.1) and (eq.2), then 1st moment r and 2nd moment r are obtained by 0 0 ( ) exp( ) 1 ( )exp( ) r f t t dt t f t t dt μ μ μ ∞ ∞ − < > = − − ∫ ∫ (7) 2 0 0 0 0 2 0 2 ( )exp( ) 1 ( ) exp( ) 2( ( )exp( ) )( ( ) exp( ) ) (1 ( ) exp( ) ) r tf t t dt t f t t dt f t t dt tf t t dt f t t dt μ μ μ μ μ μ μ ∞ ∞ ∞ ∞ ∞ − < > = + − − − − − − ∫ ∫ ∫ ∫ ∫ (8) 3. applying to weibull distribution when applying these theoretical, we consider preventive replacement characteristic from average value increase of the failure interval. we suppose the weibull distribution as ( ) exp{ ( / ) }f t t βη= − (9) 3.1 the periodic preventive replacement we insert it into (eq.5), the average value is as follow: 0 exp{ ( / ) } 1 exp{ ( / ) } t p t dt t t β β η η − < > = − − ∫ (10) moreover, it can be rearrange as [(1/ ), ( / ) ] [1, ( / ) ]p tt t β β η β η β η γ < > = γ (11) on the other hand, the 2nd moment can be got by (eq.6) 2 2 2 2 [(2 / ), ( / ) ] [1, ( / ) ] 2 exp{ ( / ) } [1/ , ( / ) ] [1, ( / ) ] p tt t t t t t β β β β β η β η β η η η β η β η γ < > = + γ − γ γ (12) 3.2 the random preventive replacement we apply the weibull distribution into (eq.7) and (eq.8) yield the following: average value i.e. 0 0 exp( ) 1 exp( ) r t dt t t dt λ μ λ ∞ ∞ − < > = − − ∫ ∫ (13) 70 wu: characteristics of item replacement in weibull distribution 2 0 0 0 0 2 0 2 exp( ) 1 exp( ) 2 ( exp( ) )( exp( ) ) (1 exp( ) ) r t t dt t t dt t dt t t dt t dt λ μ λ μ λ λ μ λ ∞ ∞ ∞ ∞ ∞ ⋅ − < > = + − − − ⋅ − − − ∫ ∫ ∫ ∫ ∫ (14) simplify, here (t/η)β +μt =λt is used 4. consideration 4.1 condition for calculation first, for convenience and simple expression, we suppose average of the weibull distribution e(t)=ηг(1/β+1)=1 and normalize the real time. if the periodic normalized time is t=0.1, the real time is 0.1ηг(1/β+1). more once we define the replacement rate μof the preventive replacement distribution function g(t)=1-exp(μt) in the random replacement because reverse of the μ is average replacement intervals. suppose t is 0.1, μ is 10, it show random maintenance had 0.1 intervals in average, the real time is 0.1ηг(1/β+1).moreover in order to compare the preventive replacement characteristics, we use the evolution reliability improvement rate in mtbf and coefficient of variation in dispersion. the reliability improvement rate is defined as following (1/ 1) 1 (1/ 1) t tη βτ η β < > γ + = =< > − γ + (15) 1 t t t tt ff f − >< = −>< =τ and the coefficient of variation cv is defined as 2 2 2 2 2 var 1iance t t tcv average t t < > − < > < > = = = − < > < > (16) 4.2 reliability improvement rate in the case of the periodic replacement, according to formula (eq.15) andηг(1/β+1)=1 improvement rate of periodic replacement is [(1/ ), ( / ) ] 1 1 exp( ( / ) )p t t β β η β ητ η γ < > = − − − (17) about random replacement use the same way. improvement rate of random replacement is 0 0 exp( ) 1 1 exp( ) r t dt t dt λ τ μ λ ∞ ∞ − < > = − − − ∫ ∫ (18) 4.3 coefficient of variation the cv value of the periodic replacement is advances in systems science and applications (2010), vol.10, no.1 71 2 2 { } 1 [2 / , ( ( )) ](1 exp( ( / )) ) [1/ , ( ( )) ] ( ) exp( ( ( )) ) [1/ , ( ( )) ] pcv a b t ta t t tb t β β β β β β β η β β = + − γ γ ⋅ − − = γ γ ⋅ γ ⋅ − γ ⋅ = γ γ ⋅ (19) here, г(·) is г(1/β+1),η=1/г(1/β+1) on the other hand, the cv value of the random replacement is 0 2 0 2 exp( ) 1 ( exp( ) ) r t t dt cv t dt λ λ ∞ ∞ ⋅ − < > = − − ∫ ∫ (20) 4.4 numerical results table 1 shows an example of calculated results in the periodic preventive replacement β= 2.5 β= 3.0 β = 3.5 β= 4.0 t τp cvp τp cvp τp cvp τp cvp 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 479 84 30 14 7 4 2 1 1 0 0.664 1.305 1.688 1.986 2.250 2.510 2.791 3.126 3.557 4.144 1571 195 57 23 11 6 3 2 1 0 0.973 1.605 2.000 2.305 2.575 2.844 3.147 3.526 4.052 4.837 5085 448 107 38 17 8 4 2 1 0 1.234 1.882 2.287 2.595 2.865 3.136 3.451 3.866 4.481 5.486 16344 1020 200 62 25 11 5 2 1 0 1.469 2.141 2.554 2.862 3.128 3.397 3.716 4.158 4.857 6.099 comparing to above table, we can get two picture as following: figure 2 improvement rate of periodic replacement from the figure 2 of above, we can see that: the interval t of exchange is smaller, reliability improvement rate τp is bigger. turn over, the interval t of exchange is bigger, reliability improvement rate τp is smaller. this kind of trend has nothing to do with the size of β. the interval of exchange around one, the reliability improvement rate will become very small, the effect of preventive replacement will disappear. in a word, it is very important to exchange 72 wu: characteristics of item replacement in weibull distribution with the interval which is smaller to the average of item. the shape parameter β is bigger, the reliability improvement rate is bigger, the interval between breakdown and breakdown will become bigger. figure 3 cv value of periodic replacement from the figure 3 of above, we can see that: the interval t of exchange is bigger, the coefficient of variation cvp is bigger. turn over, the interval t of exchange is smaller, the coefficient of variation cvp is smaller. this kind of trend has nothing to do with the size of β. the shape parameter β is bigger, the coefficient of variation is bigger. if we want to get the optimum value, we don’t only consider the reliability improvement rate but also consider the coefficient of variation cvp synthesize: it will obtain the good effect when short the interval time of exchange. this is in accordance with our experience. 5. discussion and conclusion in this report, we obtain the results of theory about periodic preventive replacement and random preventive replacement and then introduced clearly the formula of average and variance in the both preventive replacement, more applied these result to the weibull distribution, and showed the example of replacement characteristics about preventive replacement. references [1] barlow. r, proschan. f, hunter. l. “mathematical theory of reliability”, john wiley&sons inc., pp.61-63, 1965. [2] weiss. g. h. on the theory of replacement of machinery with a random failure time. naval research logistics quartely, 1956, vol.3, no.4, pp.279-293. [3] jinming su, shenyong ruan. matlab6.1 practical guide, publishing house of electronics industry, 2002, pp.212-213. [4] jianjun zhang, xiaoping wu. the c-k class of estimators of the coefficients. advances in systems science and applications, 2004, vol.4, no.3, pp.393-399. [5] jun fei, mianyun chen, chuanfeng zheng. audit inherent risk assessment based on fuzzy ahp. advances in systems science and applications, 2004, vol.4, no.4, pp.557-561. [6] jun meng. the analysis of soybean potential yield in jiansanjiang farm by logistic model. advances in systems science and applications, 2007, vol.7, no.1, pp.143-147. [7] songfa xie , jiaxiong peng, nanzhong he. automatic recognition of elevation values in scanned topographical maps. advances in systems science and applications, 2006, vol.6, no.3, pp.361-368. advances in systems science and applications (2014) vol.14 no.4 388-395 optimal location model and algorithm of the emergency exits hong-jie wang and yan gao school of management, university of shanghai for science and technology, 334 jun gong road shanghai 200093, china abstract in evacuation, the emergency exits are important. the exit’s location is a key which decides the success of evacuation. so how to build the emergency exits is worth to researching. in this paper, the effect of the emergency exit and main door are considered. the emergency exit’s ordinate is defined as the independent variable. then a model of exits is proposed and the golden section algorithm is used to find the most suitable location. at last, it is verified that the method is effective and feasible. keywords evacuation emergency exits golden section algorithm simulation. 1 introduction emergency evacuation is a mass movement of people from areas affected by disasters, such as earthquake, hurricane, and fire, to safer in a timely manner to mitigate disastrous consequences. it involves the choice of routes. if people can quickly find the exit and through the door, then hope of escape is big in the danger. so the location of emergency exit is important. and the right location of the exit can increase the rate of evacuation. in recently, the researchers have found many factors which affect the evacuation [1-5] and provide many evacuation models [6-9]. although some researches show that the building’s structure can affect the evacuation [10-11], the most of the research presupposes the building’s layout. they didn’t consider the effect of the exit’s position. it means the emergency exit, main door and access are designed and don’t change in the researches. obviously, the effect of layout is ignored. if we don’t consider the main door in evacuation process, obviously the middle sides is the most suitable exit’s location in some symmetrical architect. but if main door can be used well in the evacuation, the new evacuation route is added. so the exits and main door is in some suitable locations, maybe it can shorten the evacuation time. it’s useful to evacuation. so it’s important to finding the suitable locations. in this paper, we try to find the suitable location of exit’s location. so we only consider the effect of exit’s location. and in the research, we also consider the main door in the evacuation. in order to finding the suitable location of exit, we use the golden section algorithm. at last, we discuss and conclude the numerical results by numerical examples. advances in systems science and applications (2014) vol.14 no.4 389 2 description the authors consider a population of n individuals and assumed that each individual nk(k = 1, 2, · · · , n) moving by the speed vkin the emergency condition and randomly select the evacuation route. generally, the evacuation time is that how many times is used between the beginning of escape and all people arriving at the safe place. because we hope that we can save more people in shortest time, so the evacuation problem can be described by the object function as follows: min f = max{ti(ni) | i = 1, 2, · · · , n } (1) where ti(ni) is the evacuation time of the ith individual. in the articles, we consider the situation as follow. all individuals were randomly distributed in space according to the population density, and their initial velocities equal to zero. everyone has two choose, main door or emergency exits, to escape. we assume that each individual is a particle without quality and each individuals speed is same. we set up a coordinate system as fig.1 and use the array (xi, yi) to show the location of the ith particle. (a) the rectilinearly shaped floor (b) the vertical shaped floor fig.1 the initial distribution of the people on the floor we assume the number of main door is l,and the one of exits is w.then the kth particle’s evacuation time by the each door can be as:(k = 1, 2, · · · , n) ti(nk) = ski vk i = 1, 2 · · · , l (2) tj(nk) = skj vk j = l + 1, l + 2, · · · , l + w (3) which the ski is the shortest distance between theith main door and the kth particle,the skj is the shortest distance between the jth exit and the kth particle.so the kth particle′s evacuation time tk is satisfied the equation: tk(nk) ≤ max{ti(nk) |i = 1, 2 · · · l, l + 1, l + 2, · · · , l + w} (4) 390 hong-jie wang: optimal location model and algorithm of the emergency exits summary, the evacuation problem can be as: min f = max{tk(nk) | k = 1, 2, · · · , n } ≤min max 1≤k≤n {ti(nk) |i = 1, 2, · · · l, l + 1, l + 2, · · · , l + w} (5) obviously, when the each people choose the shortest path, the equation is established. the people’s location is fixed. perhaps, the each evacuation time maybe change, if the emergency exit’s location is changed. because the distances between the exits and the kth particle are changed.so the location of emergency exit effects the evacuation. as follow, we consider the exit’s location. we assume(x, y) is the coordinate of emergency exit and f(x) is the max evacuation time. in articles, we only consider the situation that there are two emergency exits which are distributed on the right and left sides of building respectively. in the situation, the coordinate y is 0 and the x is the independent variable. so we only consider the coordinate of emergency exit x.absolutely, the f(x) is decided by x. because we hope to find the most suitable location of emergency exit, which the evacuation time is the shortest, so we can setup the function as follows: min f(x) = max{tk(nk, x) | k = 1, 2, · · · , n } s.t. a ≤ x ≤ b (6) where a is the index of wall-width,b is the superscript of wall-width, and x is the coordinate of emergency exit’s center point. 3 algorithm according to the shortest path theory, when everyone selects the shortest path, the each individual’s evacuation time and total time are the shortest. we select the shortest path as the each individual’s escaping path. let each individuals distance function as: (k = 1, 2, · · · , n) sk(x) = min{ski(nk, x) |i = 1, 2, · · · , l, l + 1, l + 2, · · · , l + w} (7) absolutely, the function sk(x) is the linear and the f(x) = max{tk(nk, x) |k = 1, 2, · · · , n} is the linear, too. we can solve the problem by the optimization algorithm. we hope that we can exclude the unsuitable position in the process of searching. so we select the golden section algorithm to solving the problem. the golden section algorithm is a kind of optimization algorithms for solving one-dimensional problem as follow: min f(x) s.t. a ≤ x ≤ b (8) advances in systems science and applications (2014) vol.14 no.4 391 where f(x) is a uni-modal descent function. the method is a iteration algorithm and find the best solution by shorting the independent variables interval in each iteration, which reduced probability is the 0.618. the algorithms step is as follows: step1:define the a,b,ε and (xi, yi), i = 1, 2, · · · , n ; step2:define: x2 = a+ 0.618 ∗ (b− a) (9) sk(x2) = min{ski(nk, x2) |i = 1, 2, · · · , l, l + 1, l+2, · · · , l + w} (10) f2 = max{tk(x2) = sk(x2) vk | k = 1, 2, · · · , n } (11) go to step3; step3:define x1 = a+ 0.382 ∗ (b− a) (12) sk(x1) = min{ski(nk, x1) |i = 1, 2, · · · , l, l + 1, l + 2, · · · , l + w} (13) f1 = max{tk(x1) = sk(x1) vk | k = 1, 2, · · · , n } (14) go to step4; step4: if | b− a | ≤ ε,define x∗ = a+ b 2 (15) and stop; else go to step5; step5: if f1 < f2, define b = x2, x2 = x1, f2 = f1 then go to step3; if f1 = f2, define b = x2, x2 = x1,f2 = f1, then go to step3; if f1 > f2, define a = x1, x1 = x2, f1 = f2,then go to step6; step6: define x2 = a+ 0.618 ∗ (b− a) (16) sk(x2) = min{ski(nk, x2) |i = 1, 2, · · · , l, l + 1, l + 2, · · · , l + w} (17) f2 = max{tk(x2) = sk(x2) vk | k = 1, 2, · · · , n } (18) then go to step4. 392 hong-jie wang: optimal location model and algorithm of the emergency exits 4 simulation results 4.1 architectural attributes to design the suitable position of exit, the authors have designed the building layout to be two shaped. one is the rectilinearly shaped building, which each floor is a 10×30 units orthogonal area. the another is the vertical shaped building, which major semi-axis of floor is 14units, and the minor semi-axis is the 7units. in the two layouts, we consider the one kind of scenario that main entrance is placed on the bottom middle-south side of the floor. because the exits are usually placed on the right and left sides in most buildings, we consider the exits on the right and left sides, and they are opposite. the space occupancy levels vary in square meter per person ranging within the limits of space and occupancy density standards. in the two scenarios, the authors used occupancy densities of 0.15, 0.25, 0.5, 0.7, and 1.2 person/units 4.2 simulation and numerical tests to verify the feasible and effective of the way, it’s run each scenario for six hundred times, take the average of the evacuation time, and analyze the all results. in the all tables, where x, is the center point’s coordinate of emergency exit,p is occupancy densities, and p is the percentage of people selecting the emergency exit. in order to clearly verify our conclusion, we compare the number results to the evacuation time of building with emergency exits in center sides. simulation numerical results reveal the following: table 1.1 the results of rectilinearly shaped with one main entrance situation of emergency exit p number of number of number of number of number of number of( person/m2 ) 0 < x < 1 1 ≤ x < 2 2 ≤ x < 3 3 ≤ x < 4 4 ≤ x < 4.9 4.9 ≤ x < 5.1 0.15 1 7 4 176 44 3 0.25 0 1 5 98 55 14 0.5 0 0 0 128 39 8 0.7 0 0 0 100 49 12 1.2 0 0 0 87 34 12 table 1.2 the results of rectilinearly shaped with one main entrance. situation of emergency exit p number of number of number of number of number of( person/m2 ) 5.1 ≤ x < 6 6 ≤ x < 7 7 ≤ x < 8 8 ≤ x < 9 9 ≤ x < 10 0.15 40 171 50 81 53 0.25 0.45 191 63 70 32 0.5 82 208 50 50 35 0.7 103 201 53 44 38 1.2 128 200 66 40 33 advances in systems science and applications (2014) vol.14 no.4 393 table 2 the results of vertical shaped with one main entrance. occupancy building with emergency exits densities in center sides results of simulations the percentage of the percentage of p average-time people selecting the average-time people selecting the( person/m2 ) (unit) emergency exit p (unit) emergency exit p 0.15 6.51 0.51966 6.43 0.51049 0.25 6.66 0.51896 6.56 0.51319 0.5 6.78 0.52090 6.67 0.51368 0.7 6.82 0.51860 6.71 0.51167 1.2 6.89 0.51890 6.77 0.51450 table 3.1 the results of vertical shaped with one main entrance. situation of emergency exit p number of number of number of number of number of number of number of( person/m2 ) 0 < x ≤ 1 1 < x ≤ 2 2 < x ≤ 3 3 < x ≤ 4 4 < x ≤ 5 5 < x ≤ 6 6 < x ≤ 6.9 0.15 13 9 20 8 10 15 2 0.25 11 5 10 13 4 7 1 0.5 2 0 7 6 2 3 0 0.7 0 2 8 3 9 4 1 1.2 0 0 1 2 0 2 0 table 3.2 the results of vertical shaped with one main entrance. situation of emergency exit p number of number of number of number of number of number of number of( person/m2 ) 6.9 < x ≤ 7.1 7.1 ≤ x < 9 9 ≤ x < 10 10 ≤ x < 11 11 ≤ x < 12 12 ≤ x < 13 13 ≤ x < 14 0.15 7 19 22 115 99 209 52 0.25 6 11 11 103 115 256 47 0.5 1 2 8 101 65 341 62 0.7 0 2 3 87 52 378 60 1.2 0 2 4 79 22 411 77 table 4 the compare of vertical shaped with one main entrance. occupancy building with emergency exits densities in center sides results of simulations the percentage of the percentage of p average-time people selecting the average-time people selecting the( person/m2 ) (unit) emergency exit p (unit) emergency exit p 0.15 7.78 0.64214 5.62 0.68798 0.25 8.22 0.64292 5.94 0.67642 0.5 8.62 0.65034 6.18 0.68156 0.7 8.81 0.64968 6.26 0.68131 1.2 9.01 0.64573 6.37 0.68303 394 hong-jie wang: optimal location model and algorithm of the emergency exits 5 discussions and conclusion from the table1.1, table 1.2, table 3.1, and table 3.2, it can be known that the x arent distributed randomly, but most in some areas. the occupancy density is more, the number of x’s appearance is larger in this areas. in the rectilinearly shaped building, there are about 50% x in the area from 5.1 to 7. in the vertical shaped building, the x is the most in the areas from 12 to 13. and both of numbers are rising with the occupancy density’s raising. in the centre area, the percentage of x’s appearance is less than 2.5% in each simulation, both of two shaped. obviously, the more the occupancy density is, the smaller the individual’s distributive space is. so the regularity is clearer as the density raising. these show that the effect of emergency exit is existent and the center point is not the suitable place for buildings of this size and shape. because that if the effect isn’t existent, the distribution of x is random and no regular. from the table 2 and table 4, the evacuation times of simulations are shorter than the one of situations which emergency exit is in the center of right and left sides, although the p are close in each comparison. it shows the location of emergency exits is effect the evacuation, and the center point is not necessarily the most suitable position. the suitable location of emergency exit can shorten the evacuation time. so the authors conclude the conclusions as follow: 1. the location of emergency exit effects the evacuation in some way. 2. the suitable place of emergency exit is related to the shape of building, occupancy density, and the layout and people distribution. the center point of left and right sides is not necessarily the most suitable position in different situation. 3.golden section algorithm about searching the emergency exit is effective. because the distribution of people is random and irregular in each simulation, so the results are not clustering some point. but when a building is designed and built, the building’s shape, the buildings layout and the distribution of people can be estimated. so for buildings, we can use the golden section algorithm to search the emergency exit in base of some constraint, such as the shape of building, the width of exit, the pass-rate per unit of exit and the estimable distribution of person. references [1] y. lim, and s. rhee. (2010), “an efficient dissimilar path searching method for evacuation routing”, ksce journal of civil engineering, vol.14b, pp.61-67. [2] r. stamatina, s. constantinos. (2010), “escape dynamics in office buildings: using molecular dynamics to quantify the impact of certain aspects of human behavior advances in systems science and applications (2014) vol.14 no.4 395 during emergency evacuation”, environmental modeling assessment, no.15, pp.411418. [3] g. y. jeon, j.y. kim, w.h. honget al. (2011), “evacuation performance of individuals in different visibility conditions”, building an environment, vol.46,1094-1103. [4] h. frantzich. (2001), “occupant behavior and response time cresults from evacuation experiments”, in: proceeding of 2nd international symposium on human behavior in fire,pp.159-165. [5] r. a. kady. (2012), “the development of a movement-density relationship for people going on four in evacuation”, safety science , no.50, pp.253-258. [6] z.x. fang, q.q. li, q.p. li, et al. (2011), “a proposed pedestrian waiting-time model for improving space-time use efficiency in stadium evacuation scenariosa”, building an environment, no.46, pp.1774-1784. [7] p. a . thompson, e. w. marchant. (1995), “a computer model for the evacuation of large building populations”, fire safety journal, no.24, pp.131-148. [8] t.s. shen. esm. (2010), “a building evacuation simulation model”, regional studies, vol.45, no.6, pp.733-754. [9] buch, c.m. (2005), “why do banks go abroad? evidence from german data”, building and environment, no.40, pp.671-680. [10] w j. lei, a g. li, r. gao et al. (2012), “influences of exit and stair conditions on human evacuation in a dormitory”, physica a no.391, pp.6279-6286. [11] z.m. fang, w.g. song, j. zhang, et al. (2010), “experiment and modeling of exitselecting behaviors during a building evacuation”, physica a, no.389, pp. 815-824. [12] m. kobes, i. helslootb. de vries, et al. (2010), “way finding during fire evacuation: an analysis of unannounced fire drills in a hotel at night”, building an environment, no.45, pp.537-548. [13] y. yuan, d.w. wang, z.z. jiang, et al. (2008), “mult-i objective path selectionmodel for emergency evacuation taking into account the path complexity”, operations researchand management science,, no.17, pp.73-79. [14] jf. yang, y. gao, lh. ling. (2010), “emergency evacuation model and algorithm in the building with severalexits”, systems engineering-theory & practice, no.31, pp.147-153. corresponding author author can be contacted at: gaoyan@usst.edu.cn microsoft word 1151 article text, copyedited.docx adv syst sci appl 2021; 04; 45-56 published online at https://ijassa.ipu.ru. choosing directions for investments in the development of companies under uncertainty valery akinfiev* v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: akinf@ipu.ru abstract: the problem of choosing investment decisions of companies in the oligopoly market in conditions of uncertainty in demand for products is considered. investment decisions include the choice of the size and ratio of funds for the implementation of projects of two types: projects to expand production and, accordingly, to increase the supply of products on the market; as well as projects aimed at reducing production costs, which only affect the profitability of production and the free cash flow of companies. we propose an approach based on the joint use of the dcf model and algorithms that describe the investment behavior of companies in the market. the model takes into account the relationship between the choice of investment decisions by companies and the dynamics of the market price. the solution of the problem is reduced to the analysis of a matrix game in which the payoff matrix is formed as a result of numerical simulation. an illustrative example of using the proposed approach is given. keywords: investment decisions, oligopoly, demand uncertainty, company behavior model 1. introduction we consider the problem of analyzing and choosing investment strategies for companies under conditions of uncertainty in market demand. we are looking at oligopolistic markets, where two or more companies compete with each other for market share and therefore profit share. companies can choose different investment strategies knowing that the future dynamics of market demand is uncertain. different investment strategies, depending on the implementation of the demand dynamics scenario, can lead to both profits and losses for the company. in addition, when setting the problem, an important factor of mutual influence of the investment strategies of companies and the dynamics of prices for products will be taken into account [24]. the investment strategy determines the amount of funds directed by the company to invest in expanding production and (or) reducing production costs. the cost of products (below the industry average or above the industry average) is of great importance in the competitive struggle. consequently, an investment strategy aimed at reducing costs is often preferable to a production expansion strategy. the modern theory of investments in conditions of uncertainty has been developing over the past twenty to twenty-five years and is associated with the development of analysis methods using the ideology of evaluating the value of real options of various types in continuous time. the fundamentals of the methodology of this approach are outlined in a number of publications [6, 10, 11], which had a great influence on the development of this area of research. currently, it is characterized by a wide arsenal of methods and many problems statements [5-11, 13]. * corresponding author: akinf@ipu.ru 46 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) for most variants of these problems, qualitative or quantitative results were obtained under the following fairly strong assumptions: 1. it is assumed that investments occur instantly and immediately after that there is an infinite cash flow, which is the main factor of uncertainty. in a number of works, the uncertainty of the cash flow is determined by the stochastic demand and supply of products on the market. 2. it is also assumed that cash flow or demand behaves as a random process of a certain kind. the overwhelming majority of works use geometric brownian motion (gbm) to model this uncertainty, which was originally used in the black-scholes-merton option pricing model in financial markets, and consists of the following: a random process st is a geometric brownian motion (gbm) if it satisfies the following stochastic differential equation [12]: (1) in this model, the parameter µ is a non-random trend, σ is the degree of volatility, and wt is the standard wiener process (brownian motion). note that this approach assumes that the parameters µ and σ in the stochastic equation are constant. this approach, according to its supporters, describes well the volatility of asset prices in financial markets and can be applied with great caution to the valuation of real assets, in particular, to the assessment of investment projects and investment strategies. there is some pessimism regarding the practical value of this approach [12]. in addition, these models provide only a first approximation to the complex processes taking place in modern financial and especially commodity markets. it should be emphasized that in equation (1) the parameter µ (non-random trend) is considered constant and does not depend on time. this assumption significantly reduces the realism of modeling the dynamics of demand in the market, the rate of change of which can also vary significantly over time if we consider time intervals of several years, which is a characteristic feature of investment projects in real sectors of the economy. and, of course, the function µ (t) is also unknown in advance. for large investment projects, the forecasting horizon is usually 10-15 years. demand shocks can occur during this period, when demand can change dramatically. for example, “demand shocks” in the markets occurred in 2008-2009. (world economic crisis) and in 2014-2015. (falling oil prices). for example, the demand for metal on the russian market fell by 20-25% in 2009 and began to partially recover only by 2010-2011. in 2015, domestic demand for metal also decreased by 10-15% due to a significant decrease in production and sales in the automotive industry (-30%) and a decrease in the construction industry. another important note. for the purposes of evaluating the investment decisions of companies, the parameter µ (t) is more important than the second stochastic term. the influence of the second term in formula (1) decreases if we take into account that time-averaged integral characteristics are used to assess the efficiency of investment decisions, such as, for example, the npv indicator. regarding the first assumption, it is well known that the investment process of a company is continuous. the investment effect (cash flow) occurs with a lag, which depends on the duration of the investment phase of the project. in addition, cash flow can change significantly over different periods of time, including under the influence of random demand, the results of the choice of investment strategies of companies and product prices. as shown, there is a significant gap in the degree of realistic modeling of investment processes between theoretical works based on the methodology of real options and traditional methods of modeling investment projects based on the dcf methodology. dcf models more realistically allow simulating cash flows at different stages of the implementation of investment projects, which explains their wide application in practice. however, these methods do not work well in the face of uncertainty in the model input data, including price and demand volatility associated with company competition in the market. this t t t tds s dt s dwµ s= + choosing directions for investments in the development of companies... 47 copyright ©2021 assa. adv. in systems science and appl. (2021) is one of the main arguments against them on the part of the proponents of the real options approach [6]. the article attempts to mitigate the disadvantages of both approaches. for this, it is proposed to use jointly methods of scenario modeling of market uncertainty and aggregated dcf models, which are supplemented by the behavioral models of companies making investment decisions based on changing market information. 2. the model consider a market with n companies. each company can make investment decisions in accordance with the chosen strategy and based on information from the market. the criterion for the success of the company's investment strategy is the npv indicator (total discounted free cash flow of the company for the forecast period). as noted earlier, the choice of investment decisions is significantly influenced by the forecast of demand and prices for products. we will consider a situation of high demand volatility, including when market demand can change the trend several times during the forecast period. this situation is the most difficult to assess and analyze the effectiveness of investment decisions. the difficulty lies in the fact that companies cannot predict changes in demand for the entire forecast period and are limited only to assessing the trend that is observed during the period in which the company's investment budget is formed. the question arises what should be the investment strategy of companies in the face of uncertainty in demand and uncertainty in the behavior of competitors? the situation becomes even more complicated if we take into account the influence of companies' investment activity on the dynamics of market parameters. excessive investment activity of companies, as a rule, leads to the emergence of "extra" production capacity and, during periods of declining demand, to a significant decrease in the price of products [2, 3]. we consider a situation where campaigns at the beginning of each period form investment budgets based on some rational rules, using the results of analyzing trends in the dynamics of demand and product prices. in this paper, the term "choice of an investment strategy" includes: the choice of the company's financial resources allocated for investment, and the choice of the investment direction, which determines the ratio of funds allocated for the implementation of projects of two types: § projects to expand production and, accordingly, increase the supply of products on the market; § projects aimed at reducing production costs, which only affect the profitability of production and the free cash flow of companies. consider a time period (forecast period) equal to t periods, . free cash flow of company i in period t, equal to the company's net profit for this period minus investments, and is calculated by the formula: (2) where: the market price of products in the period t. in each period of time, the price is formed based on the ratio of demand for products and the total supply of manufacturers . is determined in each period t as , where is the production capacity of company i. then , where the parameter is the price elasticity on the value of the excess of demand over supply. here is the market price at the beginning of the forecast period (initial conditions). if there is a shortage of supply in the tt ,1= ( )tncfi ),1( ni = )()1()())()(()( tiptbtctptncf iiii --××-= )(tp ( )td ( )ts ( )ts ( ) å = = n i i tsts 1 )( )(tsi ( ) ( ) ( ) ( ) ( ) 0 (1 ( ) d t s t p t p d t g = × + × g ( )0p ( ) ( ) 0³tstd 48 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) market, and the price rises, otherwise, there is an excess of supply and, accordingly, the price falls. it should be noted that may also depend on price dynamics. an increase can lead to a decrease , which is taken into account by introducing negative feedback into the model. the degree of price influence on demand is specified through the parameter of demand elasticity in relation to changes in the market price. further, markets with inelastic demand are considered, which, in particular, include the metallurgical industry. let, further, be the total sales in the market in period t, which is calculated in the model using the following formula: . suppose that the capacity utilization of all companies is the same, then the sales revenue of company i is calculated as follows: is the cost of production of company i in period t, p is the income tax rate. note that all the quantities included in the calculation formula depend on random demand and the choice of investment strategies of companies that determine the dynamics of their production capacity and production costs . next, we will consider a model that allows us to evaluate the effectiveness of investment strategies of companies, taking into account the market uncertainty of demand. the general structure of the developed model for the duopoly market is shown in fig. 1. demand dynamics . exogenous variable of the model, the graph, which sets various external macroeconomic scenarios in relation to the model. fig. 1. model structure ( )td ( )tp ( )td ( )tb ( ) ( ) ( ){ }tstdtb ;min= )( )( )()( ts ts tbtb i i ×= )(tci ( )tncfi ( )td )(tsi )(tci ( )td ( )td choosing directions for investments in the development of companies... 49 copyright ©2021 assa. adv. in systems science and appl. (2021) 2.1 company behavior model we assume that companies make investment decisions in the face of market volatility and uncertainty in the dynamics of demand for companies' products. companies can observe in each period t only the change in their financial indicators (net profit, product price, sales volume) and (or) predict their change for the next several periods. the model allows varying the depth of a “reliable” forecast of market dynamics available to market participants. this makes it possible to take into account the factor of "foresight" of companies in the analysis. obviously, in case of low market volatility, the depth of the “reliable” forecast can be increased. at the beginning of each period t, the company makes an investment decision based on this available information in accordance with some predetermined algorithm, which will be described below. let further, if or (which signals the company about an upward trend in the market), then part of the company's net cash flow accumulated over period t in a share equal to this can be directed to investments in its development . the value determines the investment activity of the company i. the higher the value (share), the higher its investment activity. thus . as noted earlier, a company can direct total investments in projects of two types: projects aimed at increasing production capacity (projects of the first type) and projects aimed at reducing costs (projects of the second type) in a certain ratio and , ( ). the quantities , and are the parameters of the model that companies can choose based on their forecasts of market dynamics. thus, the company chooses investments based on the analysis of actual data and a possible forecast of market dynamics. cannot exceed , which is determined, in turn, through a variable parameter of investment activity. , where: and investments in projects of the first and second types, respectively, which are determined in accordance with some algorithms described below. 2.2 first type investment company i invests in projects of the first type according to the following algorithm: if during the period t there is an upward trend in the market (supply-demand> 0), then the company invests in projects of the first type as follows: , where is the maximum allowable level of investments in projects of the first type for the period. if there is a downward trend in the market during period t, then this signals the company about the appearance of excess production capacity and, in accordance with this, the company does not invest in projects of the first type, i.e., . suppose there is a lag between the investment period and the period of increasing production capacity . let also the value characterize the increase in production capacity per unit of investment. then , production capacity is calculated using a recurring formula . at t = 1, the initial production capacity of the company i is set. 0)1()( >-tptp 0)1()( >-tbtb ii ia ( )ti i * ia ( ) å = ×= 1 1 * )( t t iii tncfti a 1 ia 2 ia 12 1 ii aa -= ia 1 ia 2 ia ( )ti i ( )ti i ( )ti i * )()()( 21 tititi iii += )(1 ti i )(2 ti i }),(min{)( 1*11 прiii ititi ×= a прi 1 0)(1 =ti i 1t ( )tvi 1e )()( 1 1 1 t-×= tietv ii ( ) ( ) ( )tvtsts iii +-= 1 ( )0is 50 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) 2.3 second type investment. let us consider the influence of the company's investment decisions on the dynamics of changes in the cost of production. the cost of production in period t is calculated using the following formula: (3) where: the specific efficiency of investments in projects to reduce production costs, the lag between the investment period and the period of the corresponding change in the cost. for t = 1, set is the production cost at the beginning of the forecast period. the value characterizes the reduction in the cost of production per unit of investment in projects of the second type. the calculations are based on the “decline in investment efficiency” model. in accordance with this model, it decreases as the cost of production changes and approaches a certain threshold value . assessment of the maximum possible reduction in production costs, upon reaching which the investment efficiency becomes equal to zero . (4) the company invests in projects of the second type according to the following algorithm: , where is the maximum allowable level of investments in projects of the second type for the period. in accordance with the described algorithm, if starting from period t becomes equal to zero, then the amount of investment will also become equal to zero. in accordance with this algorithm, the volume of investments of company i in period t depends on the accumulated cash flow for this period, the selected parameters and , as well as on the ratio of the current investment efficiency indicator to the investment efficiency at the beginning of the period. forecast period. 3. the problem of choosing an investment strategy the presented model makes it possible to calculate the free cash flow of companies based on the choice of parameters of investment strategies and scenarios of market demand dynamics . let be the free cash flow of company i , which depends on the choice of investment strategies by all companies and the implementation of the scenario of market demand dynamics. then the indicator of the effectiveness of the chosen investment strategy is calculated as the difference between the free cash flow of the company implementing the investment strategy with parameters and the free cash flow of the company in the absence of investments, i.e., . (5) where: d is the discount rate. companies make decisions to invest in expanding production or reducing production costs in accordance with the rules described in the previous sections. for each scenario of market )()()1()( 2 2 2 t-×--= titetctc iii )(2 te 2t )0(ic )(2 te )(2 te )(tci прc прc 0)(2 =te ) )()0( 1()0()( 22 пр ii c tcc ete -×= }, )0( )()1()(min{)( 2 2 21*2 прiii i e tetiti ×-×= a прi 2 )(2 te )(2 ti i ia 1 ia ),1,,( 1 niii =aa y )(tncfi ),1( ni = 1, ii aa 0,0 1 == ii aa t tt t iiiii d tncftncfnpv )1( 1))0,0,(),,(( 1 1 + ×-=å = = aa choosing directions for investments in the development of companies... 51 copyright ©2021 assa. adv. in systems science and appl. (2021) demand dynamics , companies choose investment strategies that maximize (5), taking into account the possible choice of investment strategies of competitors. in the above formulation of the problem, each company solves its own maximization problem (5). the sought-for task variables are parameters and . the difficulty in solving this problem is that the dependence cannot be specified in the form of an explicit analytical expression on and . the value of the optimality criterion (5) is calculated each time for given and for all agents using the simulation model described in the previous section. this problem belongs to a class of optimization problems that is quite difficult to solve, in which the values of the criterion and constraints are set using a simulation model. standard mathematical programming methods are not applicable here. a grid search algorithm for finding a solution to the problem is proposed, which consists in discretizing the set of solutions. for each point, a simulation experiment is carried out, and on the basis of the data obtained, matrices of criteria values are constructed for each agent. the general solution of the problem for n agents must satisfy the nash equilibrium conditions (the saddle point of function (5) in the space of the sought variables). the final stage of the search for a solution is reduced to the analysis and search for a solution to the game, which is described below. the problem can be reduced to the study of a model of a one-step continuous non-zero sum game, in which the payoff functions of players (companies) are specified by a simulation model. therefore, numerical modeling methods will be used as the main method for studying this problem. note that the set of possible investment strategies for each company coincides with the set of points in the unit square on the plane. without loss of generality, one can consider a finite set of strategies using, for example, grid search methods. let, further, for greater clarity and simplicity of presentation, companies use several basic investment strategies. consider, for example, the following set of strategies. strategy 1 ( ). invest in production expansion (intensive development option). this strategy allows you to increase production volumes and, in a growing market, leads to an increase in sales and, accordingly, cash flow. in a falling market, this strategy leads to a decrease in the utilization of production equipment and an increase in production costs and, accordingly, a decrease in cash flow. strategy 2. ( ). invest in lower production costs (extensive development option). this strategy allows you to reduce production costs without increasing production. in a growing market, this leads to maintaining sales and, accordingly, increasing cash flow by increasing profitability. in a falling market, this strategy allows you to maintain the amount of cash flow due to the lower "break-even point". often, companies use more cautious strategies in which the estimated volume of investment is evenly distributed between investments of the first and second types ( ). at the same time, only investment activity changes, which depends on the predicted dynamics of the market situation. as a rule, with the expected improvement in market conditions, companies increase their investment activity ( ). strategy 3. ( ). high investment activity strategy 4 ( ). moderate investment activity strategy 5 ( ). low investment activity. y ),1,,( 1 niii =aa ia 1 ia 1( , , )i i incf t a a ia 1 ia ia 1 ia 1,1 1 == ii aa 0,1 1 == ii aa 5,01 =ia ia 5,0,1 1 == ii aa 5,0,5,0 1 == ii aa 5,0,0 1 == ii aa 52 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) next, we'll look at the duopoly market. let the set of investment strategies of company 1 ( ) and company 2 ( ). based on the results of a series of calculations for each scenario, it is possible to construct companies' payoff matrices in this formulation, the solution of the problem is reduced to the analysis of a bimatrix game with payoff matrices . this problem has been well researched. as is known, the condition for the existence of at least one nash equilibrium point in pure strategies (k0, j0) is the fulfillment of the following inequalities: (6) (7) if such a point exists, then this is considered a solution to this problem. the possibility of obtaining a solution (nash equilibrium point) of a bimatrix game in pure strategies, as a rule, is not guaranteed and depends on the properties of the matrix . the method for solving this problem includes carrying out a series of numerical calculations on a simulation model, constructing a payment matrix, analyzing it, and finding a solution. as an example of using the proposed approach, the following are the results of calculations and analysis of solutions to the problem for the case of three investment strategies and two scenarios of market demand dynamics. it will be shown that in many cases there are nash equilibrium points in pure strategies; the analysis of game solutions allows us to draw a number of conclusions that are interesting for practice. the result of this stage of the analysis of the investment strategies of companies is the answer to the following question: if the scenario is implemented, then what should be the companies' strategies that are optimal in the sense of criterion (5). the solution to the problem assumes that both companies consider the most likely implementation of the same scenario of market demand dynamics. 4. modeling and analyzing results this section provides an illustrative example of the application of the proposed approach. consider the duopoly market in the steel industry. a simplified situation is considered: there are two companies on the market that produce one type of metal products. the following model parameters were used in the calculations: it is assumed that in the period t = 0 supply and demand in the market are balanced, i.e., b (0) = d (0) = s (0), and equal to 10 million tons per year. p (0) is the market price of the product, equal to usd 500 per tonne; c (0) the cost of producing one ton of products is usd 400. the parameter of price elasticity g is equal to 0.5. this means that if the value of the imbalance deviates in the period t from the equilibrium value (in the period t = 0) by b%, respectively, the price of the product will change in relation to its equilibrium value by 0.5b% (with the same sign). e1(t)=0,03, i.e., with an investment of us $ 100 million, the company's production capacity is increased by 3.0 thousand tons over the period. is equal to 2 years, which corresponds to the average duration of the implementation of investment projects in metallurgy aimed at increasing production capacity. the maximum allowable volume of investments in projects of the first type for the period was adopted at the level of usd 100 million per year. e2 (0) = 0.015, i.e., with an investment of usd 100 million, the production cost of 1 ton of products is reduced by usd 1.5. the value of in the calculations is taken to be 1 year. these parameters of the model were obtained on the basis of an analysis of data on the implementation of investment projects in a large metallurgical company (nlmk group) and, naturally, reflect the assessment of the averaged values of these parameters. this information kk ,1= jj ,1= 2 , 1 , , jkjk npvnpv 2 , 1 , , jkjk npvnpv kknpvnpv kjjk ,1,11 000 =³ jjnpvnpv jkjk ,1,22 000 =³ 2 , 1 , , jkjk npvnpv y 1t прi 1 2t choosing directions for investments in the development of companies... 53 copyright ©2021 assa. adv. in systems science and appl. (2021) is presented in detail in [1]. for other sectors of the economy, these parameters of the model should be revised. the forecast period from 2020 to 2035 is considered. let both companies have the same initial capacity (5 million tons of products per year). scenario 1 (steady growth in demand). in accordance with this scenario, market demand for products throughout the forecast period grows at a steady 4% per year (fig. 2). the payoff matrix is presented in table 1. table 1. payoff matrix (scenario 1) j=1 j=2 j=3 k=1 7120\7120 8097\6414 9221\4262 k=2 6414\8097 7481\7485 9056\4801 k=3 4262\9221 4801\9056 6965\6965 the first element of the matrix corresponds to the winning of the first player, and the second element corresponds to the winning of the second player. in the table, the maximum elements of the columns of the matrix of the first player and the maximum elements of the rows of the matrix of the second player are highlighted in bold. analysis of the resulting matrix shows the presence of a single nash equilibrium point, which corresponds to the choice of investment strategy 1 by both companies (k = 1 and j = 1). at this stage, the elements of the matrix of the first and second players are highlighted in bold. figure 2 shows the dynamics of the expected market demand for products for scenario 1 and the dynamics of the growth of production capacity and, accordingly, the supply from companies 1 and 2. fig. 2. demand production capacity (scenario 1) it is important to note that a solution to a bimatrix game can have multiple nash points. this situation is illustrated by an example (table 1). at the first nash point, the winnings of both players are the same and equal to usd 7120 million. however, if companies agree to choose a different point that corresponds to investment strategy 2 (k = 2 and j = 2), then the players' gains will amount to usd 7481 million, which exceeds their winnings at the first nash point. the way out in such situations is the cooperation of the players. in real economic situations, competitors can interact with each other, entering into negotiations and concluding agreements, 54 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) often informal. if the payoff matrix of the game has a set of points that are pareto optimal (i.e., others do not dominate), then such a set is called negotiated. players can agree to jointly select a point from the negotiating set in a variety of ways, including various arbitration schemes. the analysis of the game matrix (table 1) illustrates the usefulness of such cooperation for both companies. if cooperation between players is impossible, it is more profitable for them to adhere to equilibrium strategies. scenario 2 (2020-2024 is the period of growth in demand, then, 2025-2026 is a sharp drop in demand, and, further, 2027-2035 is a slow steady recovery of demand to the level of 2024) (fig. 3). the payoff matrix for scenario 2 is shown in table 2. in the table, the maximum elements of the columns of the matrix of the first player and the maximum elements of the rows of the matrix of the second player are highlighted in bold. analysis of the resulting matrix shows the presence of a single nash point, which corresponds to the choice of investment strategy 1 by both companies (k = 1 and j = 1). the elements of the matrix of the first and second players are highlighted in bold at this point. table 2. payoff matrix (scenario 2) j=1 j=2 j=3 k=1 4750\4750 4285\3507 5764\2980 k=2 3507\4285 4087\4087 5941\3715 k=3 2980\5764 3715\5941 5098\5098 note that the nash point for both scenarios is the same. this suggests that in conditions of competition and lack of cooperation, regardless of the scenario, the most profitable strategy for both companies are the strategy of maximum investment activity. the negotiation set consists of a single point that corresponds to the choice of investment strategy 3 by both companies (k = 3 and j = 3), which corresponds to the strategy of minimum investment activity. this set of player strategies provides them with the maximum payoff in case of their cooperation. figure 3 shows the dynamics of the expected market demand for products for scenario 2 and the dynamics of the growth of production capacities of company 1 and 2. fig. 3. demand production capacity (scenario 2) choosing directions for investments in the development of companies... 55 copyright ©2021 assa. adv. in systems science and appl. (2021) 5. conclusion the problem of analysis and selection of investment strategies of a company in the duopoly market, aimed at increasing their competitive advantages (increasing production capacities and reducing production costs), has been investigated. the problem is reduced to the analysis of a bimatrix game, in which the payoff matrix is formed as a result of numerical simulation. the method for solving this problem includes carrying out numerical calculations on a simulation model, constructing a payoff matrix and analyzing it. it is shown that in many cases there is a solution of a given game (nash equilibrium point) in pure strategies. the analysis of decisions taking into account possible coalitions of players and various types of agreements between them made it possible to draw a number of qualitative conclusions that are interesting for practice. it should be noted that the proposed approach to the study of the problem of choosing the investment strategies of companies opens up ample opportunities to study various modifications of this problem. for example, by varying the parameters of the model, one can consider various types of asymmetry in the market: § one of the companies is the industry leader in terms of production volumes and market share (d1 (0)> d2 (0)); § one of the companies is the technology leader in the industry. has lower production costs (s1 (0)> s2 (0)); § one of the companies has the best management, which translates into more efficient use of funds allocated for investments. e1 (0)> e2 (0). the methodology presented in the article was used to study the investment strategies of companies in the rolled metal market, as well as in the oil market [3, 4]. references 1. akinfiev, v. (2010) management of the development of integrated industrial companies: theory and practice (on the example of ferrous metallurgy). moscow: lenand, 224 pages. [in russian] 2. akinfiev, v. (2017) a modeling and choice of strategic investment decisions in oligopoly markets / proceedings of the 10th international conference "management of large-scale system development" (mlsd). moscow: ieee, [online]. available: https://ieeexplore.ieee.org/document/8109587/ 3. akinfiev, v. (2017) a model of competition between oil companies with conventional and unconventional oil production. large-scale systems control. issue 67. moscow: ics ras. 52-80, [in russian] 4. akinfiev, v. (2019) modeling and estimating the impact of the opec agreement on oil production in russia. advances in systems science and applications, 19(3), 131-139. doi.org/10.25728/assa.2019.19.3.718 5. bolton, p., yang, n. & wang, n. (2014) investment, liquidity, and financing under uncertainty // working papers/ columbia university. [online]. available: https://www0.gsb.columbia.edu/faculty/pbolton/papers/ssrn-id2364067 6. dixit, a. k. & pindyck, r. s. (1994) investment under uncertainty // princeton university press, princeton, р. 488. 7. grenadier, s. r. & wang, n. (2005) investment timing, agency, and information // journal of financial economics 75, 493–533. 8. grenadier, s.r. (2002) option exercise games: an application to the equilibrium investment strategies of firms // review of financial studies, vol. 15, no. 3, 691–721. 56 v. akinfiev copyright ©2021 assa. adv. in systems science and appl. (2021) 9. mason, r. & weeds, h. (2010) investment, uncertainty and pre-emption // international journal of industrial organization, 28, 278–287. 10. schwartz, e. s. & trigeorgis l. (2001) real options and investment under uncertainty: classical readings and recent contributions // the mit press, cambridge, usa, р. 261 11. smit, h. t. & trigeorgis, l., (2004) strategic investment: real options and games // princeton university press, princeton, nj, usa. 12. subbotin, a. (2009) modeling of volatility: from conditional heteroscedasticity to cascades on multiple horizons // applied econometrics, no. 3 (15), 94 138 [in russian]. 13. xiumei, l., shiqin, x. & xiaoling, t. (2014) investment timing and capacity choice under uncertainty // hindawi publishing corporation. abstract and applied analysis. volume 2014, article id 801862. adv syst sci appl 2018; 02; 53-58 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/582 use of neural network models in market risk management mariya radosteva1, vladimir soloviev1*, vera ivanyuk1,2, anatoliy tsvirkun2 1) financial university under the government of the russian federation, moscow, russia e-mail: vsoloviev@fa.ru 2) institute of control sciences ras, moscow, russia e-mail: tsvirkun@ipu.ru abstract: this topic is of high relevance due to the fact that many market risk assessment mathematical models currently available contain many limitations for their effective use. however, these limitations are often not feasible, which leads to a decrease in forecast accuracy. to avoid this, more accurate models are necessary. neural network-based models can show a more accurate result due to their basic property – nonlinearity. the goal of this paper is to build a model that can enable us to assess a market risk for a company. keywords: risk, forecasting, neural networks 1. introduction fig. 1. the primary goal of this paper is to determine a lower bound of the yield to be forecast by the neural network model with a certain level of significance. current actual yields will be fed to the neural network output, and some factors will be fed to the neural network input. starbucks stock prices for the period from 2011 to 2016 were taken as the factors fed to the neural network input. * corresponding author: vsoloviev@fa.ru 54 m. radosteva, v.soloviev, v.ivanyuk, a.tsvirkun copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 2 however, there is no point in using this data as is, so the values of stock yield were fed to the neural network input. it was decided to take weekly yields with a lag of one day as the yield basic indices. in other words, the yield values were calculated as follows: 𝑟𝑡 = 𝑃𝑡 − 𝑃𝑡−5 𝑃𝑡−5 , where the t is a time interval assumed as a day (the t assumes the values of 6,7,8,... from the beginning of the counting period taking into account that the initial value is indexed by the one); the 𝑃𝑡 is the closing price on the t day without taking into account dividends (this choice is due to the fact that we cannot forecast dividends upon stocks and we do not need this because we assess the market risk of a company); the 𝑟𝑡 is the earned weekly yield which has been used for modeling. fig. 3 to improve the model quality volatility was also used as input and output data, so the next step was to calculate daily yields and weekly volatilities (standard deviations) on their basis with a lag of one day similarly with the yields. the volatilities finding procedure can be described as follows: 𝑟𝑘 = 𝑃𝑘 − 𝑃𝑘−1 𝑃𝑘−1 𝜎 = √∑(𝑟𝑘 − 𝑟�̅�)2 4 𝑘=1 another input parameter was trade volume change values. use of neural network models in market risk management 55 copyright ©2018 assa. adv. in systems science and appl. (2018) 𝑣𝑜𝑙𝑢𝑚𝑒𝑡 = 𝑉𝑜𝑙𝑢𝑚𝑒𝑡 − 𝑉𝑜𝑙𝑢𝑚𝑒𝑡−5 𝑉𝑜𝑙𝑢𝑚𝑒𝑡−5 , where 𝑉𝑜𝑙𝑢𝑚𝑒𝑡 is an absolute trade volume value on the t day; 𝑣𝑜𝑙𝑢𝑚𝑒𝑡 is a trade volume relative change on the t day. 2. distribution of the neural network errors when forecasting the yield the model is based on the architecture of the artificial neural network of a multilayer perceptron or a direct propagation network with error back propagation learning method. the model is aimed at forecasting yield lower bound for the time being, which can be crossed by the actual yields graph not more than α in 100% of cases. to solve this problem, it was decided that the neural network should learn based on historical data. it was also assumed that the current yield and volatility depended on lagged yield and lagged volatility values. an empirical distribution function was constructed and a quantile of 0.95 level was taken. the error distribution density is shown in fig. 4. fig. 4 the resulting value was calculated as a correction level to the initially specified var curve. that is the final formula for the var based on the neural network is as follows: 𝑉𝑎𝑅𝑛𝑒𝑢𝑟 = 𝑟𝑛𝑒𝑢𝑟 − 𝜎𝑛𝑒𝑢𝑟 − 𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑖𝑜𝑛 this curve was constructed in two ways. the first way of constructing is based on the fact that the neural network learns once and then a forecast is made using it. that is essentially a model with static weights in that the weights inside the network are calculated once and no longer change. this approach to the var curve constructing gives an advantage in terms of program running time, since the learning occurs once. for the sake of convenience this model will be hereinafter referred to as a neural network model with static weights. the second way is based on the assumption that the market structure is changeable (which is actually the case) and to make a forecast more accurate the neural network learns each time when calculating the var. that is, to forecast the yield for the time being the neural network learns using n size sample, which includes the previous n values of this index. this is a model with dynamic weights. at each calculation of the var value the neural network learns again. of course, such method of the var constructing is inferior to the first one, but it should be more accurate. for the sake of convenience this model will be hereinafter referred to as a neural network model with dynamic weights. 56 m. radosteva, v.soloviev, v.ivanyuk, a.tsvirkun copyright ©2018 assa. adv. in systems science and appl. (2018) kupic test consists in checking the following statistical hypothesis: 𝐻0: 𝛼0 = 𝛼 against alternative 𝐻𝑎𝑙𝑡: 𝛼0 ≠ 𝛼, where the 𝛼 = 𝐾 𝑁 , 𝐾 is the number of the var line breaks, the 𝑁 is the quantity of forecast data, and the 1 − 𝛼0 is a specified level of significance. checking is performed using statistics 𝑆𝐶𝑢𝑝 = −2 ln((1 − 𝛼) 𝑁−𝐾𝛼𝐾) + −2 ln((1 − 𝛼0) 𝑁−𝐾𝛼0 𝐾), which has the 𝜒2(1) distribution if the null hypothesis is true. another quality index that helps to choice between the models that passed the kupic test is the loss function value, which measures the average value of the var level excess by actual losses. the smaller the loss function value, the more adequately risk is assessed by the considered model. the loss functions most commonly used are the lopez and blanco-ihle’s ones: 𝐿𝐿𝑜 = 1 𝐾 ∑(𝑦𝑡 − 𝑉𝑎𝑅𝑡) 2𝐼(𝑦𝑡 < 𝑉𝑎𝑅𝑡) 𝑇 𝑡=1 , 𝐿𝐵𝐼 = 1 𝐾 ( 𝑦𝑡 − 𝑉𝑎𝑅𝑡 𝑉𝑎𝑅𝑡 ) 𝐼(𝑦𝑡 < 𝑉𝑎𝑅𝑡) the lopez’s loss function is different in that it gives a greater weight to significant deviations. this is justified from a substantive point of view, since single large excesses are usually more dangerous than a few small ones. when constructing the var curve, the programming language r with the neuralnet library was used, which allowed to construct direct propagation neural networks with the error back propagation methods in different versions. in particular, the rprop learning method was used. during the research, experiments to change the number of layers in the multilayer perceptron, as well as the number of neurons in each layer in order to improve forecast accuracy were carried out. the learning sample-based forecast error formula built into the neural network was used as a guide. such error formula was calculated as follows: 𝐸 = 1 2 (𝑦𝑓𝑎𝑐𝑡 − �̃�) 2 and was output by the program itself. since the neural network showed different results at each learning session, it was decided to take the arithmetic mean of this error in order to generalize this parameter for the neural network with such layers and quantity. 𝐸ср = 1 𝑛 ∑𝐸𝑖 𝑛 𝑖=1 , where the n was taken equal to the order of 30. the result of the research is the error! reference source not found.. table 1. number of neurons in layers yield volatility with a relative change in trade volume without a relative change in trade volume without a relative change in trade volume (6, 3) 0.14399 0.163759 0.00813 use of neural network models in market risk management 57 copyright ©2018 assa. adv. in systems science and appl. (2018) (8, 3) 0.135139 0.157891 0.00813 (11, 6) 0.12661 0.14632 0.00813 (12, 3) 0.12789 0.139695 0.00813 (12, 6) 0.13096 0.132457 0.00813 (22, 6) 0.1176 0.14441 0.00801 (22, 11) 0.10761 0.15441 0.00771 (24, 3) 0.107 0.14177 0.00713 (24, 6) 0.11826 0.17130 0.00712 (22, 11, 3) 0.12405 0.1789 0.01519 (22, 11, 6) 0.10574 0.17111 0.02907 (24, 12, 6) 0.11581 0.1753 0.01813 from the implementation point of view the back propagation method converges for a long time, therefore its various modifications are often used. in particular, during the neural network learning the rprop method has been used which is as follows. the algorithm is based on the partial derivative sign. new weights are updated according to: 𝑤𝑖𝑗 (𝑡+1) = 𝑤𝑖𝑗 (𝑡) + ∆𝑤𝑖𝑗 (𝑡) { ∆𝑤𝑖𝑗 (𝑡) = −𝑠𝑖𝑔𝑛 ( 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) )∆𝑖𝑗 (𝑡), if 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡−1) 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) ≥ 0 ∆𝑤𝑖𝑗 (𝑡) = −∆𝑤𝑖𝑗 (𝑡−1) и ( 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) ) = 0, if 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡−1) 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) < 0 . where the 𝑠𝑖𝑔𝑛() is a function that assumes the value of +1 or -1 depending on which sign has a value in parentheses; ∆𝑖𝑗 (𝑡)= { min(𝜂+∆𝑖𝑗 (𝑡−1) , δmax ) , if 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡−1) 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) > 0 max(𝜂−∆𝑖𝑗 (𝑡−1), δmin ) , if 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡−1) 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) < 0 ∆𝑖𝑗 (𝑡−1), if 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡−1) 𝜕𝐸 𝜕𝑤𝑖𝑗 (𝑡) = 0 where the zero values ∆𝑖𝑗 (0) and 0 < 𝜂− < 1 < 𝜂+ are chosen arbitrarily. such an algorithm converges faster than a usual back propagation method. the result of this study is the var curve constructed on the basis of the neural network model. one of the ways to implement the neural network model for the var level forecasting is shown in figure 5. 58 m. radosteva, v.soloviev, v.ivanyuk, a.tsvirkun copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 5 the var curve constructed on the basis of the neural network model with static weights and a significance level of 0.95 (the black line is the yield curve, the red line is the var level). in this case, the actual percentage of break is at the level of 4.65%. the kupic test has shown a p-value= 0.6809912418, which is a good result. the lopez and blanco-ihle’s tests for such a situation assumed the following values: 𝐿𝐿𝑜 = 0.0001344051456, 𝐿𝐵𝐼 = 0.1513434227, which generally evidences a good quality of the model, since the break depth is relatively small. in particular, the lopez’s index indicates that existing breaks are light, which is a good sign. the result of the paper is the market risk assessment neural network model, which has shown good results according to the tests performed. references [1] ben, k. & van der smagt, p. (1996). an introduction to neural networks. university of amsterdam. [2] haykin, s. (1999) neural networks: a comprehensive foundation, prentice hall [3] ivanyuk, v. & tsvirkun, a. (2013). intelligent system for financial time series prediction and identification of periods of speculative growth on the financial market. ifac proceedings volumes, 46(9), 1128-1133. [4] ivanyuk, v. & pashchenko, f. (2015) methods and models for the forecasting and management of time series. // itise 2015, granada, 283-292 [5] simonov, b., ivanyuk, v., & simonova, i. (2016). existence of best approximation elements in the spaces l. journal of mathematical sciences, 217(5), 1-21. [6] koroteev, m., terelyanskii,, p. & ivanyuk v. (2016) approximation of series of expert preferences by dynamical fuzzy numbers. journal of mathematical sciences, 216(5), 692-695. [7] rumelhart, d. e., hinton, g. e., & williams, r. j. (1985). learning internal representations by error propagation (no. ics-8506). california univ san diego la jolla inst for cognitive science. adv syst sci appl 2017; 3:22–33 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/497 conformal kernel expected similarity for anomaly detection in time-series data aleksandr safin1,2, evgeny burnaev2,3* 1 national research university higher school of economics, moscow, russia 2 skolkovo institute of science and technology, skolkovo, moscow region, russia 3 institute for information transmission problems, moscow, russia abstract: the problem of anomaly detection arises in many practical applications. currently it is highly important to be able to detect outliers in data streams, as recent years have seen a rapid growth in the amount of such data. only a few techniques are applicable to real-time data and even fewer could provide an interpretable anomaly score. probabilistic interpretation of the anomaly score could allow an analyst to choose the anomaly threshold based on the desired false alarm rate, which is highly important in a number of real-life applications. we propose a modification of the expose algorithm for anomaly detection in time series data, which produces a probabilistic score of abnormality. the proposed algorithm is developed within the framework of conformal anomaly detection and utilizes the expected similarity as a measure of non-conformity. keywords: anomaly detection, conformal prediction, time series, kernel methods, expected similarity 1. introduction there are many cases in which it is highly important to determine whether a new observation comes from the same distribution or not. this problem is referred to as outlier or anomaly detection. e.g. when a fitted model is applied to new data, it should be checked whether a test data set belongs to the same population as the training data set. to address the issue of novelty detection, anomaly detection techniques can be used. anomaly detection has proven to be helpful for certain medical purposes, fraud detection and machine diagnostics, to name but a few. for instance, in [1] failure prediction for aircrafts is considered. the definition of an anomaly varies between algorithms and applications. in general, an anomaly “is an element whose properties differ from the majority of the other elements under consideration which is called as normal data” [14]. in [15] anomaly detection is described as follows: “anomaly detection refers to the problem of finding patterns in data that do not conform to expected behavior”. summing up, the problem of anomaly detection can be formulated as follows: the task is to determine for every object in a test set whether it is a normal or abnormal instance in comparison with observations from a training set. anomaly detection approaches can be divided in the following three groups [15]: • unsupervised approaches use only the assumption that most observations are normal. such assumption favours incremental and autonomous learning in data streams. ∗corresponding author: e.burnaev@skoltech.ru http://ijassa.ipu.ru/ojs/ijassa/article/view/497 conformal kernel expected similarity for anomaly detection in time-series data 23 • supervised approaches require availability of a labelled training set containing instances of both normal and abnormal objects. • semi-supervised approaches require a small amount of labeled data with a large amount of unlabeled data. in many practical cases a number of outliers is significantly smaller than a number of target observations, and thus usual classification methods may yield unsatisfactory results as classes in a dataset are very imbalanced. the significant dominance of target instances over outliers is a natural property of real-life data: e.g. in case of air traffic safety problems accidents happen very rarely. another reason for that is the impossibility or very high costs of reproducing faulty conditions when we consider a machine diagnostic task. in the light of known and outlined difficulties, classical methods are not applicable to solving these problems, thus a variety of outlier detection methods have been developed. such problems justify the need for specialized approaches to anomaly model selection [2], learning with privileged information [6], construction of ensembles of non-parametric anomaly detectors in data streams [3,7,8], usage of specific time-series models [9–13], and explicit rebalancing of normal and abnormal classes [4, 5], among others. unsupervised anomaly detection does not require the training dataset to be labelled, thus it is applicable to various problems, as in general it is not feasible to collect labels. therefore a number of applications adopts unsupervised approaches, e.g. based on density estimation or clustering [16]. according to the surveys [15] and [17] unsupervised anomaly detection techniques could be generally categorized as probabilistic (distributionand density-based), prediction-based, distance-based, classification-based, clustering-based and information-theoretic approaches. distribution-based methods estimate parameters of a target data distribution and determine whether a test object comes from the same distribution that generated samples from the training set. the main drawback of these methods is the necessity to select some parametric class of data distributions. one of the tricks is to model the target distribution as a mixture of gaussians, however the number of gaussians still have to be determined. to mitigate this problem, other non-parametric techniques could be utilized, for instance histogram-based or kernel density estimator (kde). however, the number of bins should be firstly specified, and the performance is highly sensitive to this hyper-parameter. for multivariate problems a basic approach is to estimate a histogram per each input feature. however some features could be correlated, in that case the information about such dependency will be lost. prediction-based techniques predict future observations based on previous items and then compare predicted and real data to identify anomalies. another type of approaches is based on the distance to the k-th nearest neighbour (knn). one of such techniques is the knn-based outlier detector [18]. all objects are sorted w.r.t. the average distance to k nearest neighbours and top n of them with the highest average distance are claimed to be anomalies. lof method, proposed in [19], exploits density based approach: it uses the distance to the k-th nearest neighbour as an inverse estimate of a local density value. however data could contain clusters with different densities leading to significantly increased false anomaly detections. main drawback of distances-based anomaly detection methods is poor interpretability of their output. to address this issue conformal anomaly detector (cad) was proposed by laxhammar [16]. having a probabilistic interpretation of the degree of anomalousness allows choosing a threshold with a false alarm rate guarantee. zhao and saligrama stated in [20] that “while [modern anomaly detection] approaches provide impressive computationally efficient solutions on real data, it is generally difficult to precisely relate tuning parameter choices to desired false alarm probability”. at the same time according to burnaev and nazarov, conformal prediction could be used for constructing non-parametric confidence intervals with a specified confidence probability [21, 22]. copyright © 2017 assa. adv syst sci appl (2017) 24 aleksandr safin, evgeny burnaev some techniques adopt approaches used for classification tasks. tax and duin proposed support vector data description [23] and later it was refined in [24]. the task of data description is formulated as follows: given the unlabelled training data, construct a closed boundary (or a set of them) that contains predominantly the target data, and outliers are outside this boundary. in a simple case, boundary is supposed to be spherical, but in general it is possible to determine an arbitrary-shaped flexible boundary by using kernel functions. moreover, svdd is robust against the training data containing outliers and also is capable of improving the accuracy by incorporating additional information about negative examples, in case when the training dataset is labelled. according to the results of their study, svdd is shown to yield mostly comparable or even better results for sparse and complex multidimensional datasets. an extension of svm to the case of unlabelled data is outlined in [25]. this approach which is referred to as one-class svm has been adapted by ma and perkins [26] for time series. the problem of anomaly detection has many dimensions. in particular, data could be not fixed but represented as a stream. a variety of techniques could be used for anomaly detection. however, only a few can be used for data streams. classical algorithms for anomaly detection are not applicable due to their computational complexity and memory consumption, since the number of elements in the set is growing and therefore it is not possible to store all previously observed data. it is worth emphasizing that frequently the concept of anomaly could change in the course of time. this phenomenon is called as concept drift. in general, a streaming version of an anomaly detection algorithm should have an ability of adaptation to a concept drift; therefore, such algorithms are usually based on one of the following strategies of accumulating information about current changes in data: namely windowing and exponential smoothing, which is also referred to as decay. the above-mentioned issues are considered in the paper [14]. the proposed expose algorithm is developed within the framework of reproducing kernel hilbert space (rkhs) and exploits the concept of kernel mean embedding. in a nutshell, the estimator used in the algorithm could be described as a dot product of a kernel mean map and a feature map of the observed data point. however, the produced anomaly score has a lack of intuitive interpretability and therefore the false alarm rate could not be guaranteed when anomaly threshold is chosen. in order to eliminate this drawback we utilize lazy drifting conformal detector procedure proposed in [32] to construct the algorithm with probabilistic interpretation of the anomaly score while based on the idea of kernel mean embedding proposed in [14]. the structure of the paper is the following. sections 2 and 3 shed light on kernels and conformal anomaly detection respectively. the proposed algorithm is described in section 4. in section 5, the results of empirical evaluation (using numenta anomaly benchmark) of the proposed algorithm are outlined. finally, section 6 describes achieved results and the work still to be done. 2. overview of kernel-based methods for anomaly detection in machine learning kernels are broadly used for handling data of diverse nature. therefore it is not surprising that a number of anomaly detection methods are based on the kernel framework. in this section we give some necessary definitions and provide an overview of such anomaly detection methods which uses kernels and therefore are applicable to data of various types. 2.1. introduction to kernels reproducing kernel hilbert space (rkhs) is a hilbert space (h, 〈·, ·〉) of functions f: x → r if the evaluational functional δ̄x : f → f(x ) is continuous. copyright © 2017 assa. adv syst sci appl (2017) conformal kernel expected similarity for anomaly detection in time-series data 25 reproducing kernel of h is a function k: x × x → r which satisfies the reproducing property: 〈f,k(x , ·)〉 = f(x ), 〈k(x , ·), k(y , ·)〉 = k(x ,y). the map φ: x → h with the property that k(x ,y) = 〈φ(x ), φ(y)〉 is referred to as a feature map. definition 2.1 (expected similarity estimation): the expected similarity [14] of z ∈ x given the probability distribution p(x) is defined as: η(z) = ex [φ(z)] = ∫ x k(z, x)dp(x). definition 2.2 (kernel embedding): kernel embedding of the distribution p has the form µ[p] = ∫ x k(x, ·)dp(x). expectation of any f ∈ h ex [f ] = 〈f, µ[p]〉h. thus, η(z) = 〈φ(z), µ[p]〉h. given the empirical distribution pn(x) by observing n realizations {x1, . . . xn} independently sampled from p, one could approximate µ[p] as follows: µ[p] ≈ µ[pn] = 1 n n∑ i=1 φ(xi). this approach is referred to as empirical kernel embedding [29] and given that ‖φ(x)‖ ≤ c,c > 0 the following guarantee has been proved by schneider [30] for all ε > 0: p (‖µ[p]− µ[pn]‖ ≥ ε) ≤ 2e− nε2 8c2 . considering above mentioned, having observed {x1, . . . xn}, the expected similarity estimation for z ∈ x is η(z) = 〈φ(z), µ[p]〉h ≈ 〈φ(z), µ[pn]〉h = 1 n n∑ i=1 k(z, xi). 2.2. expose expected similarity estimation (expose) that was proposed in [14] is the method for anomaly detection which could handle data streams. for every new observation z the algorithm computes an anomaly score η(z) based on computed empirical kernel mean map wt of previously observed items η(z) = 〈φ(z), wt〉 ‖wt‖2 . copyright © 2017 assa. adv syst sci appl (2017) 26 aleksandr safin, evgeny burnaev kernel mean map could be evaluated using one of the following strategies. the first strategy is to use a sliding window of length l: wt = 1 l t∑ i=t−l+1 φ(xi). more flexible approach is to apply exponential smoothing to all previous observations: wt = γφ(xt) + (1− γ)wt−1, t > 1. the parameter γ reflects the influence of a new data item. however, as already was indicated, the anomaly score provided by this algorithm could not be interpreted in a probabilistic manner, therefore we propose an approach to transform the anomaly score produced by expose. to that end, the idea of conformal anomaly detection is adopted to build an anomaly detector for online data. 3. conformal anomaly detection laxhammar [16] proposed a conformal anomaly detection (cad) which is a distributionfree procedure for probability-like confidence measure estimation based on non-conformity measure (ncm) provided by some detector. the ncm a(x, y) reflects how different the investigated object y is from other observations x. ncm could be for instance the average distance to k neighbours, the distance to the k-th neighbour, residual in a regression model, to name just a few. let us consider a time series xt, then compute scores ats = a(x−s:t , xs), s = 1, . . . , t, where a(x, y) is an ncm used by the algorithm. then the empirical p-value is defined as: p(xt, x:(t−1), a) = 1 t |{s = 1, . . . , t : ats ≥ att}|. the lower it is, the lower the probability of falsely rejecting the null hypothesis (xt is anomaly) is, thus the more likely xt is an anomaly instance. shafer and vovk proved [31] the fact that cad could provide the following guarantee when xt is i.i.d: px∼d(p(xt, x −t, a) < ε) ≤ ε,x = (xs) t s=1. it is clear that cad could be computationally heavy as it requires computations of a(x−s:t , xs) for s from 1 to t. to mitigate this problem, an inductive conformal anomaly detection (icad) was proposed by laxhammar and falkman in [27]. this approach relies on scores computed on training set x̄ for every instance of the calibration set. for further simplicity, let us consider relabelled sequence xt that starts from−n+ 1. then, icad has the following setup for every t ≥ 1: x−n+1, . . . ,x0︸ ︷︷ ︸ x̄ training , calibration︷ ︸︸ ︷ x1,x2, . . . ,xt−1,xt. in that setup, the conformal p-value of a test object xt is computed on the basis of modified scores: {ats = a(x̄,xs), s = 1, . . . , t}, x̄ = (x−n+1, . . . ,x0). however, by relaxing deterministic guarantee to probably approximately correct guarantee, it is achievable to adapt icad to use only fixed size calibration set. offline icad copyright © 2017 assa. adv syst sci appl (2017) conformal kernel expected similarity for anomaly detection in time-series data 27 was developed to use a calibration set only with fixed size m, sliding along the time series, as illustrated: x−n+1, . . . ,x0︸ ︷︷ ︸ x̄ training , . . . , calibration︷ ︸︸ ︷ xt−m,xt−m+1, . . . ,xt−1,xt. it should be highlighted that the conformal p-value in this case uses a subsample of the icad non-conformity scores: p(xt, x:(t−1), a) = 1 m+ 1 |{s = 0, . . . ,m : att−s ≥ att}|. vovk proved [28] the following guarantee for the offline icad: px∼d(p(x, x, ā) < ε) ≤ ε+ √ log 1 δ 2m . 4. proposed approach for anomaly detection in time series data in this section we outline the proposed algorithm for anomaly detection in time series. it is worth emphasising that the developed approach does not require any assumption about the data distribution and it outputs a probabilistic measure of anomality based on nonconformity scores. such measure of anomality is calculated using an adaptation of icad to the case of potentially non-stationary and quasi-periodic time series. cad and online icad are computationally complex, therefore offline icad seems much suitable for the task. nevertheless, as offline icad uses a fixed training set, it should be noticed that one could face problems in case of non-stationary time series. in the light of the discussed details and difficulties, lazy drifting conformal detector (ldcd) has been proposed in our paper [32]. for simplicity we consider a univariate time-series x = (xt)t≥1 ∈ r, although our approach is valid for multivariate data as well since we are going to use kernel-based non-conformity measure. to begin with, the time series x should be embedded into l-dimensional space. to that end, we further consider the sequence of xt = (xt−l+1, . . . , xt) ∈ rl constructed by moving window of the width l on the time series x: . . . , xt−l−1, xt−1 xt−l, xt−l+1, . . . , xt−1, xt xt , xt+1, . . . . it should be noticed that such approach obviously produce t− l+ 1 embeddings from the sequence of the length t. in other words, to produce the first such embedding, we need to observe l instances initially. as ncm we are using the expected similarity: a(tt,xt) = 1 n n∑ i=1 k(xt,xt−m−i), where tt = {xs : s = t−m− n, . . . , t−m− 1}, k(x, y) is a kernel function. data: . . . , tt training︷ ︸︸ ︷ xt−m−n, . . . ,xt−m−1, xt−m, . . . ,xt−1, test xt, . . . scores: . . . , at−m−n, . . . , at−m−1, at−m, . . . , at−1︸ ︷︷ ︸ at calibration , at test , . . . copyright © 2017 assa. adv syst sci appl (2017) 28 aleksandr safin, evgeny burnaev the proposed approach could be described as follows: 1. construct time series embedding in a sliding window, 2. compute an+s = a(tn+s,xn+s), s = 1, . . . ,m+ 1, 3. evaluate the empirical p-value of its non-conformity score: p(xt, tt, a) = 1 m+ 1 |{i = 0, . . . ,m : at−i ≥ at}|. the algorithm is depicted in the figure 4.1. fig. 4.1. flowchart of the algorithm it is worth emphasising that the proposed approach could be implemented with time complexity equal to o(n) +o(logm) in the case of using red-black tree for the calibration set. 5. results on numenta anomaly benchmark the results of the expose ldcd comparison with several other algorithms is presented in this section along with the testing methodology. the numenta anomaly benchmark (nab) is utilized to test the proposed algorithm. 5.1. datasets the nab corpus consists of 58 both real-world and artificial time series datasets. real-world data are obtained from such sources as aws server metrics, twitter volume, advertisement clicking metrics, traffic data, to name just a few. we also conduct experiments using numenta anomaly benchmark on yahoo! s5 dataset [34] which has been created to gauge the anomaly detectors performance on different types of anomalies. this corpus is divided into 4 groups: first one contains real production metrics from different yahoo! properties, and the rest are synthetic time series. copyright © 2017 assa. adv syst sci appl (2017) conformal kernel expected similarity for anomaly detection in time-series data 29 5.2. scoring algorithm commonly used metrics for performance evaluation such as accuracy, precision and recall do not suit well for anomaly detection, since they do not consider time. the nab proposes such an approach for scoring which rewards only early true detection, meanwhile penalizes late detections and punishes false alarms very hard. to capture early detections, nab considers the area which is centred around the anomaly point which is referred to as anomaly window. the window length is defined as 10% of the length of the time series. all detections within this window are true positives, but only the earliest one contributes in the total score, the others will be ignored. the detections outside the anomaly window are false positives, missed anomalies are false negatives. true negatives are not considered in the scoring mechanism. bellow we described scoring scheme used in nab [33]. an example of time series is provided in figure 5.2. the first 15% of the time series is considered as probationary period and during this period an algorithm learns patterns from the data and is not required to do any detections. then the algorithm is evaluated on the remaining part of time series. the weights for accuracy calculation is evaluated using the smooth sigmoid function as depicted in figure 5.3. fig. 5.2. the purple shaded area is the probationary period. anomalies are depicted as red points and red shaded regions represent anomaly windows [33]. as the costs of true positive (tp), false positive (fp) and false negative (fn) vary among distinct applications, in nab this is captured by an application profile which reflects the contribution of weights for tp, fp an fn detections. the “standard” application profile reflects scenarios in which misdetections have identical costs. the “reward low fp” and “reward low fn” profiles penalize harder for fp and fn respectively. profile atp afn afp atn standard 1.0 -1.0 -0.11 1.0 reward low fp 1.0 -1.0 -0.22 1.0 reward low fn 1.0 -2.0 -0.11 1.0 table 5.1. the detection rewards on nab application profiles the reward for the detection depends on the relative position t of the alarm (about possible anomaly) to the left side of the anomaly window: σa(t) = (atp − afp ) ( 1 1 + e5t ) − 1. the raw performance score on the dataset x with respect to application profile a is the sum of the scores over all detections plus the impact of missed anomalies (false negatives) copyright © 2017 assa. adv syst sci appl (2017) 30 aleksandr safin, evgeny burnaev captured by the number of anomaly windows with no detections fdet: sadet(x) = ∑ y∈ydet σa(y) + afnfdet. the overall performance of the algorithm is the sum of raw performance scores over the all datasets d: sadet = ∑ x∈d s a det(x). the final normalized performance score is determined by: sanab = 100 sadet − sanull saperfect − sanull . fig. 5.3. nab weighted scores: detections outside the anomaly window are false positives and punished; only earliest detection inside window is true positive and it will be counted, other will be ignored [33]. 5.3. results since proposed expose ldcd is conservative and demonstrates high level of false alarms, we have applied the following simple pruning strategy to reduce the false alarm rate: we output 1− p as anomaly score for the observation xt and if p is greater than 99.65%, then output of the detector is fixed at 0.5 for the next n 5 observations (n is the length of probationary period). the proposed approach has been validated on both the numenta anomaly benchmark corpus and the yahoo! s5 dataset. tables 5.2 and 5.3 reflect the results of the algorithms comparison. 5.4. automated kernel bandwidth tuning kernel-based methods are sensitive to the choice of bandwidth, therefore we modify the algorithm to choose the bandwidth of the kernel based on the best value of the bandwidth for kernel density estimator obtained by 3-fold cross-validation. the proposed modification demonstrates significantly better results on nab dataset and is able to increase the score on both low fn and low fp profiles, meanwhile it results in slight score decrease on standard profile. 6. conclusion in this paper we propose an algorithm for anomaly detection in time series data, utilizing the concept of expected similarity and applying framework of conformal anomaly detection. copyright © 2017 assa. adv syst sci appl (2017) conformal kernel expected similarity for anomaly detection in time-series data 31 table 5.2. results on numenta anomaly benchmark detector profile standard reward low fp reward low fn numenta htm 70.1 63.1 74.3 expose ldcd +tuning 45.53 25.77 54.78 windowed gaussian 39.6 20.9 47.4 expose ldcd 37.93 20.14 45.11 etsy skyline 35.7 27.1 44.5 bayesian changepoint 17.7 3.2 32.2 expose 16.4 3.2 26.9 table 5.3. results on yahoo! s5 dataset detector profile standard reward low fp reward low fn expose ldcd 51.88 38.76 58.95 expose ldcd +tuning 49.79 43.73 61.45 numenta htm 41.0 37.5 44.4 bayesian changepoint 35.7 17.6 43.6 expose 32.09 7.00 45.45 windowed gaussian 31.1 25.8 40.7 etsy skyline 23.6 18.0 28.9 table 5.4. average running time performance on nab dataset detector performance items per second ms per item windowed gaussian 1984.862 0.504 expose ldcd 1500.224 0.667 bayesian changepoint 428.639 2.333 expose 398.496 2.51 numenta htm 98.012 10.202 etsy skyline 4.582 218.229 table 5.5. average running time performance on yahoo! s5 dataset detector performance items per second ms per item expose ldcd 2548.293 0.392 windowed gaussian 2348.041 0.426 bayesian changepoint 1217.888 0.821 expose 383.711 2.606 numenta htm 103.777 9.636 etsy skyline 4.656 214.773 this approach has been rigorously validated on nab corpus and yahoo! s5 dataset using numenta anomaly benchmark. on both datasets the proposed approach excel the expose, which produces expected similarity as anomaly score and expected similarity is used as nonconformity measure in the ldcd procedure. moreover, the developed algorithm shows great running time performance, which is important for online detectors and achieves high results on standard profile on yahoo dataset. also, the implementation of the algorithm could be enhanced, as it has not been thoroughly optimised and it could be one of the directions for future research. we also propose a tuning procedure for the kernel bandwidth parameter, however there is still a significant room for improvements. acknowledgements the work was supported by the ministry of education and science of russian federation, grant no. 14.606.21.0004, grant code: rfmefi60617x0004. copyright © 2017 assa. adv syst sci appl (2017) 32 aleksandr safin, evgeny burnaev references 1. alestra s., bordry c., brand c., burnaev e., erofeev p., papanov a. & silveirafreixo c. (2014) application of rare event anticipation techniques to aircraft health management advanced materials research, 1016, 413–417. 2. burnaev e., erofeev p. & smolyakov d. (2015) model selection for anomaly detection. proc. spie9875, eighth international conference on machine vision (icmv 2015), 987525, http://dx.doi.org/10.1117/12.2228794 3. artemov a. & burnaev e. (2015) ensembles of detectors for online detection of transient changes. proc. spie9875, eighth international conference on machine vision (icmv 2015), 98751z, http://dx.doi.org/10.1117/12.2228369 4. burnaev e., erofeev p. & papanov a. (2015) influence of resampling on accuracy of imbalanced classification. proc. spie9875, eighth international conference on machine vision (icmv 2015), 987521, http://dx.doi.org/10.1117/12.2228523 5. burnaev e., erofeev p. & papanov a. (2017) meta-learning for construction of resampling recommendation systems. arxiv e-prints, 1706.02289, [online]. available:https://arxiv.org/abs/1706.02289 6. burnaev e & smolyakov d. (2016) one-class svm with privileged information and its application to malware detection. 2016 ieee 16th international conference on data mining workshops (icdmw), pp. 273–280. 7. burnaev e., ishimtsev v., bernstein a. & nazarov a. (2017) conformal k-nn anomaly detector for univariate data streams. proceedings of machine learning research, 60, 213–227. 8. volkhonsky d., burnaev e., nouretdinov i., gammerman a. & vovk v. (2017) inductive conformal martingales for change-point detection. proceedings of machine learning research, 60, 132–153. 9. artemov a. & burnaev e. (2016) optimal sequential estimation of a signal, observed in a fractional gaussian noise. theory of probability and its applications, 60(1), 126–134. 10. artemov a. & burnaev e. 2016) detecting performance degradation of softwareintensive systems in the presence of trends and long-range dependence. 2016 ieee 16th international conference on data mining workshops (icdmw), pp. 29–36. 11. artemov a., burnaev e. & lokot a. (2015) nonparametric decomposition of quasiperiodic time series for change-point detection. proc. spie 9875, eighth international conference on machine vision, 987520. 12. burnaev e. (2009) disorder problem for poisson process in generalized bayesian setting. theory probab. appl., 53(3), 500–518. 13. burnaev e., feinberg e. & shiryaev a. (2009) on asymptotic optimality of the second order in the minimax quickest detection problem of drift change for brownian motion. theory probab. appl., 53(3), 519–536. 14. schneider m., ertel w. & ramos fabio t. (2016) expected similarity estimation for large-scale batch and streaming anomaly detection. machine learning, 105(3), 305– 333, https://doi.org/10.1007/s10994-016-5567-7 15. chandola v., banerjee a. & kumar v. (2009) anomaly detection: a survey. acm comput. surv., 41(3), 15:1–15:58. 16. laxhammar r. (2014) conformal anomaly detection. detecting abnormal trajectories in surveillance applications. ph. d. thesis. university of skövde, skövde. retrieved from http://urn.kb.se/resolve?urn=urn:nbn:se:his:diva-8762 17. pimentel m. a. f., clifton d. a., clifton l. & tarassenko l. (2014) review: a review of novelty detection. signal process, 99, 215–249. 18. ramaswamy s., rastogi r. & shim k. (2000) efficient algorithms for mining outliers from large data sets. sigmod rec., 29(2), 427–438. 19. breunig m.m., kriegel h.-p., ng r. t. & sander j. (2000) lof: identifying densitybased local outliers. sigmod rec., 29(2), 93–104. copyright © 2017 assa. adv syst sci appl (2017) http://dx.doi.org/10.1117/12.2228794 http://dx.doi.org/10.1117/12.2228369 http://dx.doi.org/10.1117/12.2228523 https://arxiv.org/abs/1706.02289 https://doi.org/10.1007/s10994-016-5567-7 http://urn.kb.se/resolve?urn=urn:nbn:se:his:diva-8762 conformal kernel expected similarity for anomaly detection in time-series data 33 20. zhao m. & saligrama v. (2009) anomaly detection with score functions based on nearest neighbor graphs. nips’09 proceedings of the 22nd international conference on neural information processing systems, vancouver, canada, 2250-2258. 21. burnaev e. & nazarov i. (2016) conformalized kernel ridge regression. 2016 15th ieee international conference on machine learning and applications (icmla), 45– 52, https://doi.org/10.1109/icmla.2016.0017 22. burnaev e. & vovk v. (2014) efficiency of conformalized ridge regression. proceedings of the twenty seventh annual conference on learning theory. jmlr: workshop and conference proceedings, 35, 605–622. 23. tax d.m.j. & duin r.p.w. (2004) support vector data description. machine learning, 54 (1), 45–66. 24. chang w.-c., lee c.-p. & lin c.-jen. (2013) a revisit to support vector data description (svdd). documents in the citeseerx database [online]. available: http://ai2-s2-pdfs.s3.amazonaws.com/a244/422ba339713d0c9eaa153b378e9f9fc08263.pdf 25. schölkopf b., platt j.c., shawe-taylor j. c. et al. (2001) estimating the support of a high-dimensional distribution. neural comput., 13 (7), 1443–1471. 26. ma j. & perkins s. (2003) time-series novelty detection using one-class support vector machines. proceedings of the international joint conference on neural networks, 2003, 3, 1741–1745. 27. laxhammar r. & falkman g. (2015) inductive conformal anomaly detection for sequential detection of anomalous sub-trajectories. annals of mathematics and artificial intelligence, 74 (1), 67–94. 28. vovk v. (2012) conditional validity of inductive conformal predictors. proceedings of the asian conference on machine learning, in pmlr, 25, 475-490. 29. smola a., gretton a., song l. & schölkopf b. (2007) a hilbert space embedding for distributions. algorithmic learning theory: 18th international conference, alt 2007, 13–31, https://doi.org/10.1007/978-3-540-75225-7 5 30. schneider m. (2016) probability inequalities for kernel embeddings in sampling without replacement. proceedings of machine learning research, 66–74. 31. shafer g. & vovk v. (2008) a tutorial on conformal prediction. j. mach. learn. res., 9, 371–421 32. ishimtsev v., nazarov i., bernstein a. & burnaev e. (2017) conformal k-nn anomaly detector for univariate data streams. arxiv e-prints, [online]. available: https://arxiv.org/abs/1706.03412 33. lavin a. & ahmad s. (2015) evaluating real-time anomaly detection algorithms the numenta anomaly benchmark. 14th international conference on machine learning and applications (ieee icmla), 38-44, https://arxiv.org/abs/1510.03336 34. yahoo! webscope (2017, december 26) s5 a labeled anomaly detection dataset, version 1.0. [online]. available https://webscope.sandbox.yahoo.com/catalog.php?datatype=s&did=70. copyright © 2017 assa. adv syst sci appl (2017) https://doi.org/10.1109/icmla.2016.0017 http://ai2-s2-pdfs.s3.amazonaws.com/a244/422ba339713d0c9eaa153b378e9f9fc08263.pdf https://doi.org/10.1007/978-3-540-75225-7_5 https://arxiv.org/abs/1706.03412 https://arxiv.org/abs/1510.03336 https://webscope.sandbox.yahoo.com/catalog.php?datatype=s&did=70 introduction overview of kernel-based methods for anomaly detection introduction to kernels expose conformal anomaly detection proposed approach for anomaly detection in time series data results on numenta anomaly benchmark datasets scoring algorithm results automated kernel bandwidth tuning conclusion adv syst sci appl 2018; 02; 63-83 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/576 copyright ©2018 assa adv. in systems science and appl. (2018) wgw: a hybrid approach based on whale and grey wolf optimization algorithms for requirements prioritization amjad hudaib1, raja masadeh2, abdullah alzaqebah2 1computer information systems department, the university of jordan, amman, jordan e-mail: ahudaib@ju.edu.jo 2computer science department, the world islamic sciences and education university, amman, jordan e-mail: raja.masadeh@wise.edu.jo , abdullah.zaqebah@wise.edu.jo abstract: requirement engineering is the base phase of any software project, since this phase is concerned about requirements identification, processing and manipulation. the main source of these requirements is the project stakeholders with considering the project constraints and limitation. number of requirement is varying for each project, so the requirements prioritization term comes for prioritizing the order of execution for software requirements according to the stakeholder's opinions and decisions. various proposed optimization algorithms are employed to solve optimization problems; recently whale optimization (wo) algorithm is proposed in 2016 by mirjalili which mimics the main characteristic of humpback whales which is the foraging method that is called bubble-net technique. on the other hand grey wolf optimization (gwo) algorithm was proposed in 2014 in order to solve optimization problems by imitating the grey wolves hunting behavior. in this paper, a hybrid approach based on whale and grey wolf optimization algorithms (wgw) is proposed by combining the advantages of each algorithm in order to prioritize the software requirements. moreover, the data set that used in this paper is ralic which a real software project’s requirements is in order to evaluate the proposed method. thus, the proposed method shows 91% accuracy of requirements prioritization comparing with ralic data sat. keywords: requirement prioritizations (rp), whale optimization algorithm (woa), grey wolf optimization (gwo), replacement access, library and id card project (ralic), rp-woa. i. introduction requirement engineering (re) in one of the most significant branch in the domain of software engineering. in addition, it is considered as the most important phase in software development life cycle [1]. this phase contains identification and elicitation of requirements, analysis and requirements validation and documentation. in almost software projects, there are restrictions on development process like budget and time to market production, this leads to deliver the software projects in consecutive releases, hence large projects have more than one of stakeholders that makes a difficulty in decision making about what release should be developed firstly. this difficulty contributes the software engineers to prioritize the requirements in efficient way to make the right decision about project's delivery and development. [1, 2, 3] thus, requirement prioritization (rp) is the most important section of re that comes under the requirement analysis stage. rp is considered as one of the most significant activities in the process to construct software project and deliver good system as the customer need. in case the project has strict execution plan, insufficient resources and the expectations of the consumers are so high, then it must publish the most important characteristics as early as possible. thus, this reason leads the requirements to be prioritized. 64 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) requirement prioritization process increases the stakeholder's involvement by including them in deciding what requirements should be included in the software with its importance and effects on development process, and this involvement aids the stakeholders to really understand the restrictions and limitations on project's resources, and negotiate the conflicts between viewpoints that really effect in software development, and these conflicts come from different objectives and roles of stakeholders [43]. thus, in projects that have large number of requirements, the prioritization become a base of software success or failure according to the project's constraints and limitations, these participation and involvement aims to prioritize the requirement according to their importance in efficient manner and useful order of execution [1, 2, 4]. according to that, the clustering technique could be used in order to classify the requirements based on their importance. data clustering is one of the most important functions in data mining, which has attracted many authors, researchers and experts recently. clustering is an unsupervised ranking of types in groups [5]. the major purpose of data clustering depends on the notion and the idea of grouping the objects into groups that based on the similarity and dissimilarity between these objects [6, 7]. in general, clustering is categorized into two main classifications; the hard clustering and the soft clustering. in hard clustering, the data objects only belong to one cluster, also in the soft clustering; all the individual data objects belong to the individual classes to some range [6, 8]. the major goal of data clustering is which it optimally sorting all n data into k clusters in the opinion that the whole unseen types in the data are visible [9, 10]. thus, each cluster contains the similarity data objects, as well as, the clusters are various from others. one of the most significant methods that are employed in the data exploration process, neural computing, image segmentation and other engineering is cluster analysis [11]. for the time being, many clustering techniques are proposed by researchers which categorized into model-based clustering algorithm, hierarchical clustering algorithm, partition-based clustering algorithm, grid-based clustering algorithm and density-based clustering algorithm [9, 13]. the data that are divided into k clusters utilizing the euclidean distance as a measure in the partition-based clustering algorithms, as well the tree of groups that are created in the hierarchical clustering algorithms. recently, many researchers have proposed much of meta-heuristics and heuristic algorithms in order to solve the issues which happen as a result of complicated datasets. however, the most techniques that proposed in order to resolve the issue of the optimization relies on the meta-heuristic algorithms [11, 40-42]. the meta-heuristic algorithms main goal is to define the optimal solutions for fulfilling data cluster and reduces the issue of the local minima [14, 15]. the latest meta-heuristic optimizations algorithms are grey wolf optimization (gwo) algorithm and whale optimization (wo) algorithm. grey wolf optimization (gwo) algorithm is proposed by mirjalili et al in 2014 [12]. this mimics the hunting behavior of grey wolves in nature. grey wolves are one of the well-known predators in the nature. usually, they live in pack within group size is 5 to 12 on average. these wolves have robust rules in social dominant hierarchy. according to [12] grey wolves include alpha, beta, omega and delta wolves. alpha wolf represents the leader of the pack and it is responsible for making decision about hunting and other activities. while the beta wolf helps the higher level to make decision. omega wolves are responsible to submit any information to the highest levels. all other wolves in the pack are called delta. whale optimization (wo) algorithm is recently suggested by mirjalili and lewis in 2016 [13]. where its main objective is to define the global optimal solution for any given optimization problem. the major distinction between this algorithm and other algorithms is the principle that develops the candidate solutions in each iteration of optimization. in addition to that, bubblenet feeding process represents the hunting process in humpback whales in order to find and attack the prey. a hybrid approach based on whale and grey wolf optimization algorithms 65 copyright ©2018 assa adv. in systems science and appl. (2018) the reset of this paper is organized as follows: section ii contains backgrounds of requirement prioritization, gwo algorithm and wo algorithm in details. while section iii describes the proposed wgw and how it works. section iv outlines the suggested algorithm “rp-wgw”. the experimental results are presented in section v. finally, section vi draws the conclusion and future work. ii.background a.requirements prioritization the meaning of requirements prioritization is seen from several angles. summerville defined the requirements prioritization as one of the most significant task for decision makers [16]. while, firesmith defines it as the major process in software engineering as it gives perfect implementation order of the requirements in order to planning software versions and supplying desirable functionality as early as possible or the process to define the priority of the requirements to stakeholders based on the requirements importance [17]. therefore, we can conclude that the requirements prioritization denotes the prioritization by importance or by implementation. implementing the most significant operations that leads to get incremental feedback from the customer, set schedule, solve mistakes and resolve any misunderstand between the customer and the corporation in premature phases that lead up to customer contentment. moreover, it is valuable by eliminating needless requirements which may be inefficiently costly and choosing the most suitable requirements for each version; which leads to assist in future planning, reduce the risk of cancellation, evaluate the benefits, prioritize the investments and determine the financial effect with regards to the implementation of each requirement [18]. requirement prioritization is the most significant and critical portion of requirements analysis due to the restrictions in project resources. in other words, it is so hard to implement the whole requirements simultaneously due to the restrictions in resources whence of schedule, staff and budget. moreover, to improve some projects may require many months or often several years, wherefore it is important to determine the requirement that should be implemented at the beginning. furthermore, budget plays an important factor, especially when transaction with requirement prioritization process because budget is considered as small activity with regards to requirement engineering compared to other activities in software engineering. lastly, as mentioned before concludes that requirements have various levels with regards to their importance and it is complicate to determine which one is the most important. as mentioned above, the project stakeholders are the base of the prioritization process with respect to business and regulations factors because they have various viewpoints and each one must determine the highest priority for requirements in order to impose stakeholders to clearly gather all the relative importance requirements that guides to raise the communication between stakeholders, supplies a reasonable base for requirement negotiation and enables engineering to schedule the development activities in reasonably. b.grey wolf optimization grey wolf optimization algorithm (gwo) is one of the latest bio-inspired optimization techniques that proposed by mirjalili et al in 2014 [12]. which imitate the hunting behavior of grey wolves in nature. whereas the major purpose of gwo algorithm is determining the optimal for a given problem by using a population of search agents. these wolves usually live in groups; their members between five and twelve. the major difference between gwo algorithm and the other optimization algorithms is the social dominant hierarchy which develops the candidate solution in each iteration of optimization. in facts, the gwo simulates the foraging behavior of the wolves in finding and attacking the victims [12, 19]. the social hierarchy of the wolves pack is shown in fig.1 [12]. 66 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) fig.1. hierarchy of grey wolf. alpha displays the leader which is the best candidate solution. in addition to that, alpha are dominant wolves which followed by the other wolves. while beta represents the second candidate solution which helping alpha in decision making and represent bridges between the leader and the rest of the pack that plays the lowest levels. delta displays the third candidate solution that responsible to submit information to the two higher levels (alpha and beta). while omega's displays the rest of solutions. moreover, its responsible to submit information to the three higher levels. in fact, hunting mechanism contains three steps: tracking, encircling and attacking the prey. thus, gwo represents the hunting technique of the grey wolves mathematical that is used to resolve complicated optimization problem. thus, the optimal solution of the given problem is considered as victims. the three top levels movement simulates the victim encirclement by grey wolves, which is the following formula that is suggested in this regard [12]. �⃗⃗� = |�⃗⃗� . �⃗⃗� 𝒑(𝒕) − �⃗⃗� (𝒕)| (1) where t displays the current iteration, xp indicates to the prey position vector, x represents the grey wolf position and c is a coefficient vector. thus, the result of vector d is used to move the particular element toward or away from the area that the best solution is located which represents the prey by using the following equation [12]: �⃗⃗� (𝒕 + 𝟏) = |�⃗⃗� 𝒑(𝒕) − �⃗⃗� . �⃗⃗� | , with �⃗⃗� = 𝟐�⃗⃗� . �⃗� 𝟏 − �⃗⃗� (2) where r1 is selected randomly in [0, 1] and a is minimized from 2 to 0 through predetermined number of iterations. in case |a| < 1, this matches to the exploitation behavior and simulates the behavior of attacking the prey. otherwise, if |a|> 1, this matches exploration behavior and imitates the wolf spacing from the victim. the suggested values for a are in [-2, 2]. thus, three higher levels α, β and δ will be computed using the following mathematical expressions [12]. a hybrid approach based on whale and grey wolf optimization algorithms 67 copyright ©2018 assa adv. in systems science and appl. (2018) �⃗⃗� 𝜶 = |�⃗⃗� 𝟏. �⃗⃗� 𝜶 − �⃗⃗� | with �⃗⃗� 𝟏 = �⃗⃗� 𝜶 − �⃗⃗� 𝜶. (�⃗⃗� 𝜶) (3) �⃗⃗� 𝜷 = |�⃗⃗� 𝟐. �⃗⃗� 𝜷 − �⃗⃗� | with �⃗⃗� 𝟐 = �⃗⃗� 𝜷 − �⃗⃗� 𝜷. (�⃗⃗� 𝜷) (4) �⃗⃗� 𝜹 = |�⃗⃗� 𝟑. �⃗⃗� 𝜹 − �⃗⃗� | with �⃗⃗� 𝟑 = �⃗⃗� 𝜹 − �⃗⃗� 𝜹. (�⃗⃗� 𝜹) (5) in order to mathematically imitative the hunting process of grey wolf, assume that α, β and δ have enough knowledge about the possible position of the victim. moreover, the first three best solutions that gained are saved and force the other agents to update their locations according to the best agents α, β and δ. this behavior is represented mathematical by the following expression [12], and the pseudocode of the gwo is shown in algorithm 1 [12]. �⃗⃗� (𝒕 + 𝟏) = �⃗⃗� 𝟏+ �⃗⃗� 𝟐+�⃗⃗� 𝟑 𝟑 (6) algorithm 1: pseudocodes of gwo [12] c.whale optimization whale optimization algorithm (woa) is a recently suggested randomly optimization algorithm by mirjalili and lewis in 2016 [13]. this algorithm purposes to determine the global optimum for the problem by using a population of search agents (whales). the search process is the first step begins with generating a collection of candidate solutions that are selected randomly for the given problem. then, it ameliorates this collection during many numbers of iterations until the satisfaction of an end condition. in fact, whales imitative private hunting technique that was called bubble-net feeding method as shown in fig. 2 [13]. 68 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) fig. 2. bubble-net hunting behavior [13] it is obvious from fig. 2 that a humpback whale generates ambush with shifting in a spiral route around the victims, then generates bubbles all the way ahead. thus, this search process is the major inspiration of the whale optimization algorithm. in addition to that, the encircling method is another simulated technique in woa. whereas the humpback whales surrounding around the victims in order to begin hunting them using the foraging mechanism which is called the bubblenet technique. this behavior is represented mathematical by the following equations [13]: 𝑋 (𝑡 + 1) = { 𝑋 ∗ (𝑡) − 𝐴 . �⃗⃗� 𝑖𝑓 𝑝 < 0.5 𝐷’. 𝑒𝑏𝑙 . 𝐶𝑜𝑠 (2 𝜋 𝑙) + 𝑋 ∗ (𝑡) 𝑖𝑓 𝑝 ≥ 0.5 (7) where p is a random number in [0, 1], b is a constant for determining the shape of the logarithmic spiral, and l indicates to a random number in [−1, 1], t presents the current iteration, and d’ = |x∗ (t) – x (t)| which mentions the distance between the ith whale and the victim. in other words, the first phase is presented in this equation is the foraging mechanism that mimics the encircling technique, whereas the second phase simulates the bubble-net mechanism. the variable p exchanges between these two phases with similar probability. the potential cases those using these equations are shown in fig.3 [13]. fig.3.mathematical models for prey encircling and bubble-net hunting [13] a hybrid approach based on whale and grey wolf optimization algorithms 69 copyright ©2018 assa adv. in systems science and appl. (2018) the main two phases of the optimization algorithm by using population based algorithms are exploration and exploitation phases. whereas both of them ensured in woa by adaptively setting a and c in the major equation. in case, a problem is given, the woa begins optimizing this problem by generating collection of random solutions. in each iteration, the search agents update their location depends on the randomly chosen search agent or the best search agents that will gain so far. to assure the exploration phase, the other agents update their positions based on the best solution that represents the pivot point when |a|>1. in other case, when |a|<1 the best solution plays another role with the pivot point. the pseudocode of the woa is shown in algorithm 2 [13]. algorithm 2: pseudocodes of woa [13] iii.related work as mentioned before, the requirements prioritization phase is an operation that presenting the priority of a requirement over other requirements. moreover, it is the most motivating field in the requirement engineering domain [21]. many techniques are proposed in order to give the precedence to a software requirement over other requirements at the same project. fig.4 shows the classification of requirements prioritization techniques according to [22]. 70 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) fig.4. classification of requirement prioritization techniques [22]. as shown in fig.4 the requirement prioritization techniques are classified as ordinal scale, ordinal scale and ratio scale in terms of powerful prioritization. the most powerful prioritization is ratio scale that is giving how much more significant a requirement than the others requirements [42, 44]. while the least powerful prioritization is the ordinal scale which is responsible to give ranked list of requirements and giving which requirement is more important than others requirements but without giving how much more significant. the analytic hierarchy process (ahp) is the most popular and traditional technique that is joined ratio scale class. thus, it is mentioned by large number of studies [23-28]. furthermore, it is considered as a systematic decision making technique which has been performed for precedence of requirements for a particular project [29, 30]. it compares all potential pairs of requirements to order them and define which has higher priority. ahp is not suitable for a project that has large number of requirements [31, 32]. thus, many researchers have attempted to minimize the number of comparisons such as [33, 34] and different techniques are presented in order to minimize the number of comparisons by approximately 75% such as [35]. the most precise level software requirements are placed at the base of the hierarchy while the most complicated requirements are placed on the top of the hierarchy [23]. hierarchy ahp is another type of requirements prioritization technique that prefers the requirements that are presented at the same level. cumulative voting is called the 100-dollar test which is a straightforward technique; where 100 imaginary units are given to the stakeholders in order to divide among the given software requirements [25]. the finding of the prioritization is given on a ratio scale. numerical assignment is considered as one of the most popular mechanism that is proposed in [24, 27, 47]. this technique depends on clustering requirements into various priority classes. however, the number of classes are vary and there are three classes are very popular [25, 28]. thus, in this technique it is significant that each class displays something which the stakeholder can communicate to such as optional, critical, etc. while moscow (museum of a hybrid approach based on whale and grey wolf optimization algorithms 71 copyright ©2018 assa adv. in systems science and appl. (2018) soviet calculators on the web) is considered as a type of numerical assignment mechanism that depicted in [36, 37]. moreover, this technique has four priority classes; must have, could have, should have and wont have. must have means the requirements that joined this class must be implemented at the first then goes to version, while could have means the requirement that joined this class exist then it will be great for the software product. should have means the requirements that present in this class are implemented then will be great for the software product and finally wont have means that requirements join this cluster cannot be implemented in the present iteration as other requirements that are of low precedence. bubble sort mechanism is used to rank the element such as requirements. in this technique, two requirements are taken to compare with each other. in case the requirement is not in series then exchange between them and then compare it with another requirement and continue until get ranked list of requirements descending (from higher to lower) as used in these studies [23, 37]. binary search tree technique is another type that is used for ranking which suggested by [37]. furthermore, this technique was presented at the first time by [23] for requirement prioritization. each node in this technique indicates to a software requirement where all requirements that located in left side of the tree are of lower priority comparing with other nodes while the other requirements that are located at the right side of the tree are of higher priority. however, at the first one requirement is selected to be the root node then will compare with unsorted requirement. in case this requirement is lower priority than the root, it searches the left side of the tree. otherwise, it searches the right side of the tree. the operation is repeated until get sorted tree. in top ten requirements technique, the stakeholders select top ten requirements in terms of their importance for them without giving inner rank among these requirements. this leads the technique to be suitable for many stakeholders of equal importance [38]. grey wolf optimization (gwo) is applied by [40] which is one of the most recently metaheuristic algorithms; in order to prioritize the software project’s requirements. in addition; it is evaluated and compared with analytical hierarchy process (ahp) mechanism; where the proposed work performed better than the traditional technique (ahp) by approximately (30%). while, [41] is applied whale optimization algorithm (woa) which is recently used in optimization problems since it imitates the humpback whale hunting behavior by employing bubble net hunting technique. it is likewise evaluated with ahp; where the results shown the proposed work outperform the ahp mechanism by approximately (40%). iv.the proposed algorithm (wgw) as shown in fig.5 the approach that suggested in this study is mainly designed by combining the woa and the gwo algorithm in order to design a hybrid approach that called "wgw". the combination process is done by combining the advantages of each algorithm that are utilized in this study. the main advantage of gwo that is from each cluster the top three highest values will be considered as α, β and δ and these values represents the first three best solutions from that cluster, with keep into account the group that represents the cluster is limited 5-12 nodes on average. on the other hand the woa has no limitation in group members and finally the woa did not have a top three highest solution instead of one. from theses points the wgw is proposed to make a combination between these two algorithms by employing the woa algorithm to skip the limitation on the group team members and make it unlimited and take the top three highest solutions as gwo works. [12, 13] as any meta-heuristic optimization algorithms, the proposed algorithm is composed of two stages. the first is the exploration stage while the second is the exploitation stage. 72 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) in this study, the proposed wgw algorithm is used to search in the overall search space in order to find all possible xrands, these xrands play an important roles in exploration phase. on the other hand, any search agent sees prey, it is considered to be xrand, then this foraging process is done by generating bubbles along the path that forms spiral shape and it called the bubble-net hunting technique. thus, when any search agent sees the bubbles that mean it belongs to this spiral. in case, there are many spirals at the same time from different xrand, each one is considered as cluster in the search space and each one of them has first three best solutions xα, xβ and xδ. whereas this advantage is driven from gwo algorithm, whereas, the previous scenario represents the exploitation phase.[12, 13] in each iteration, after determining the best solutions xα, xβ and xδ, the other search agents update their positions in the cluster. taken in account the value of absolute a, if it is greater than one means the search agent goes beyond the cluster (spiral) and it must search to another xrand to belong to new cluster. otherwise, the absolute value of a is less than one means the search agent stills in the cluster and closes to the prey. fig.5. flowchart of the proposed wgw algorithm. a hybrid approach based on whale and grey wolf optimization algorithms 73 copyright ©2018 assa adv. in systems science and appl. (2018) v.algorithm “rp-wgw” in this work, an improvement approach is suggested by combining the recently bio-inspired optimization techniques that proposed by mirjalili et al which are gwo and wo algorithms which imitate the hunting techniques in nature. then, the wgw algorithm is used for requirements prioritization. to get the prioritization of the requirements, the wgw algorithm is applied. fig.6 presents the suggested pseudo code for “rp-wgw” algorithm. “rp-wgw” algorithm begin 1. 2. 3. 4. 5. 6. 7. 8. 9. 10. initialize the agents’ population xi (i = 1, 2, 3…..… n) initialize c, r and a select xrand randomly calculate the distance between each whale (i) and all xrand by eq. (1) if the agent (i) is not assigned assign agent (i) to its closet xrand calculate the fitness for each agent (i) invoke cluster function invoke requirement prioritization function return the best solution end fig. 6. pseudo – code for “rp-wgw” algorithm a.initialization stage at the beginning, initialize the agents’ population as shown in fig.6. number of agents is selected randomly to generate the bubble that performs as cone (spiral) shape. then, measure the distance between all agents and all xrands to assign to the closet one and join that group to be a member of it. b.fitness function based on eq. (8), xrand generates bubble-net when it sight the victim, all agents that see the bubble-net will associate to the cluster (spiral). thus, the other agents who joined the cluster must update their locations towards the location of xrand as illustrated in fig. 7. in other hand, these agents will update their locations depending on the value of absolute a. in case, |a| is less than one, this means the agent is still entering the cluster and updates its location by the following eq. (8). otherwise, |a| is greater than or equal one, which means the agent is not belong to the cluster, so it search for another xrand to associate by using eq. (9). this behavior is represented mathematical by eq. (8, 9) [13] and eq. (10). 𝑋 (𝑡 + 1) = �⃗⃗� . 𝑒𝑏𝑙 . 𝐶𝑜𝑠 (2𝜋𝑙) + 𝑋 (𝑡) (8) 𝑋 (𝑡 + 1) = 𝑋𝑟𝑎𝑛𝑑⃗⃗ ⃗⃗ ⃗⃗ ⃗⃗ ⃗⃗ ⃗⃗ ⃗ − 𝐴 ⃗⃗ ⃗. 𝐷 ⃗⃗ ⃗ (9) 𝑉 = ℎ ∗ 𝑟 ∗ 𝜋 ( 1 3 ) (10) where r is the radius of the bubble-net which is constant value (r = 15) [20], h is the height of bubble-net, which is selected randomly between 6 and 12 [12, 20], and π is constant value that equal 3.14. 74 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) fitness function begin 1. 2. 3. 4. 5. 6. 7. 8. 9. 10. 11. xrand creates bubble-net by eq. (8) for each search agent (i) update a, a, c and l calculate the distance between each agent (i) and xrand by eq. (1) if (a<1) update the position of the current agent (i) by the eq. (8) else select new xrand randomly update the position of the current agent (i) by the eq. (9) end if end for end fig. 7. fitness function c.clustering as shown in fig.8, each cluster k has xrand that selected randomly. the fitness function will be calculated for each agent in order to check if this agent sights the bubble that generated by xrand. in other words, this agent is still joining this spiral. in case the agent didn’t see the bubble, it searches for another xrand to belong. thus, each cluster k will get xα, xβ and xδ that represent the first three highest priority of all requirements in this cluster. cluster function begin 1. 2. 3. 4. 5. 6. 7. 8. 9. for each cluster k select xrand randomly select xα, xβ and xδ while (t< = max_iteration) invoke the fitness function invoke the rp function end while return xα, xβ and xδ end for end fig.8. cluster function d.requirements prioritization function as discussed in clustering step, each agent in the search space represents a requirement with its importance; this importance represents its importance in the development process of the target project. this importance calculated according to the stakeholders ranking for each requirement with respect the importance of that stakeholder and their effect on the project. since each stakeholder belongs to a role in the target environment that the system built for, the role's rank effect directly to the stakeholder importance so when calculating the stakeholder importance the role's rank plays an important role in it. according to the project environment, first the importance of the role based on the stakeholders ranking on it will be calculated as equation (11) [39]. 𝐼𝑛𝑓𝑙𝑢𝑛𝑐𝑒 𝑟𝑜𝑙𝑒(𝑖) = 𝑅𝑅𝑀𝑎𝑥 + 1 − 𝑅𝑎𝑛𝑘(𝑅𝑜𝑙𝑒(𝑖)) ∑ 𝑅𝑅𝑀𝑎𝑥 + 1 − 𝑅𝑎𝑛𝑘(𝑅𝑜𝑙𝑒(𝐽))𝑛 𝐽=1 (11) a hybrid approach based on whale and grey wolf optimization algorithms 75 copyright ©2018 assa adv. in systems science and appl. (2018) where rrmax is the maximum rank of roles list, rank (role (i)) is the rank of the i's role and n is the total number of roles, rrmax+1 is used to invert the value of rank since the lowest rank is highly effect. after that the influence of each stakeholder in each role will be calculated to find the effect of that stakeholder on the project, equation (12) [39] shows the stakeholder's importance calculation. 𝐼𝑛𝑓𝑙𝑢𝑛𝑐𝑒 (𝑖) = 𝑅𝑆𝑀𝑎𝑥 + 1 − 𝑅𝑎𝑛𝑘(𝑖) ∑ 𝑅𝑆𝑀𝑎𝑥 + 1 − 𝑅𝑎𝑛𝑘(𝐾)𝑛 𝐾=1 (12) where i represents a specific stakeholder, rsmax is the maximum rank of stakeholder in that role, rank(i) is the rank of the i'th stakeholder in the role and n is the total number of stakeholders in the same role, rsmax+1 is used to invert the value since the lowest rank is the highest effect. then the influence of that stakeholder on the project at all will be calculated by multiplying the role influence and the stakeholder influence as equation (13) [39]. 𝑃𝑟𝑜𝑗𝑒𝑐𝑡 𝐼𝑛𝑓𝑙𝑢𝑛𝑐𝑒 (𝑖) = 𝐼𝑛𝑓𝑙𝑢𝑛𝑐𝑒 𝑟𝑜𝑙𝑒(𝑖) × 𝑖𝑛𝑓𝑙𝑢𝑛𝑐𝑒 (𝑖) (13) since each stakeholder ranked the requirement list, the importance of that requirement is calculated by summation of all stakeholders' influence on the project multiplied by its rating on that requirement, equation (14) shows the requirement importance calculation )[39]. 𝐼𝑚𝑝𝑜𝑟𝑡𝑎𝑛𝑐𝑒 𝑅 = ∑𝑃𝑟𝑜𝑗𝑒𝑐𝑡 𝑖𝑛𝑓𝑙𝑢𝑛𝑐𝑒(𝑖) × 𝑟(𝑖) 𝑛 𝑖=1 (14) where r (i) represents the i'th stakeholder's rank on that requirement and n is the total number of stakeholders that rating the requirement r (i). as shown in fig. 9, the first three top values will be chosen from each cluster. thus, this process will be executed for each cluster iteratively, as well as a result from this process the set of best solutions from each cluster will be gained in a temporary list, so this list contains the α’s, β’s and δ’s from the clusters. since this list contains the best solution, the top three solutions will selected as first result of prioritization process. the selected results will be moved to the final ranked list as best three solutions, since they are already taken and prioritized; these requirements will be removed from original clusters. this process will be repeated until all clusters have no requirements inside, as mention above the selected best solutions will be removed from the original clusters, so the number of requirements inside the cluster will be decreased while iteration keep running. finally the proposed approach will return the final ranked set of requirements according to the importance of each requirement. 76 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) requirements prioritization function begin 1. 2. 3. 4. 5. 6. 7. 8. 9. while each cluster is not empty (1,2,3,…k) for each cluster k sort the requirements based on eq. (14) select xα, xβ and xδ and move it to temp_list end for select xα, xβ and xδ from the sorted temp_list and move them to final ranked list remove xα, xβ and xδ from original clusters end while return the final results set end fig.9. requirements prioritization function vi.experimental results in order to evaluate the proposed method for requirements prioritization in term of accuracy, the ralic project was selected from a soo ling lim phd thesis as a case study in this paper; ralic stands for the replacement access, library and id card project. it was a software project to enhance the existing access control system at university college london (ucl). located in the central part of london, ucl is concerned with its security [39]. many ucl buildings require authorized access, such as the libraries, academic departments, and computer clusters. in 2005, ucl had various methods of identification and access control, such as swipe cards, contactless cards, photo id cards, library barcode, digital security code, and metal door keys. ucl staff and students had to use different mechanisms to access different buildings, which meant they had to carry various cards with them [39]. furthermore, some of the security systems were already obsolete; others would cease to be operable in a few years’ time. ralic’s aim was to replace the obsolete access control systems, consolidate the various existing access control mechanisms, and at the minimal, combine the photo id card, access card, and library card [39]. ralic was a combination of development and customization of an off-the-shelf system [39]. the project started in 2005 and its duration was 2.5 years. the system has been deployed in ucl. the project scope is summarized in table 1 [39]. a hybrid approach based on whale and grey wolf optimization algorithms 77 copyright ©2018 assa adv. in systems science and appl. (2018) table 1. ralic project scope [39] ralic is a large-scale software project. it had more than 60 stakeholder groups and approximately 30,000 users. some of the stakeholder groups included students, academic staff, short course and academic visitors, administrators from academic departments, security staff, developers, managers, and front line staff from supporting divisions such as the estates and facilities division that manages ucl’s physical estate, human resource division that manages staff information, registry that manages student information, information services division, and library services. ralic has approximately 30,000 students, staff, and visitors, who use the system to enter ucl buildings, borrow library resources, use the fitness center, and gain it access. ralic had a complex stakeholder base, with different stakeholders having conflicting requirements. for example, members of the ucl development & corporate communications office preferred the id card to have ucl branding, but the security guards preferred otherwise for security reasons in case the cards were lost. some administrators were worried that the new system, which promised to reduce manual labor, would threaten their job. the project involved many divisions in ucl, some of which had low stake in the project but were critical to its success. in particular, the project team found it difficult to engage with the divisions that manage the interfacing systems that supply data to ralic, such as the student registry that provide student data, and human resources that provide staff data, because they had little stake in the project. some representatives from these departments were often absent in project meetings. the author presents a ground truth in order to evaluate her proposed method; this ground truth contains a real list that prioritized in the real project after deploying it, the data set contains 10 project objectives, 49 requirements and 80 specific requirements after cleaning and processing the data which described in [39]. in this research; the work of [39] was implemented specifically on requirements with keep into account the rank of project objectives and discard her work on specific requirements since the proposed method concerns about requirements prioritization. the tested data set contains 57 requirements with each importance of them, the result of [39] work is shown in table 2 which shows a prioritized list of requirements. 78 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) table 2. prioritized list of requirements [39] id requirement description priority a.3 all in 1 card 1 a.1 easier to use 2 a.2 use the same access control for library entrance 3 c.3 enable visual checking 4 c.4 access control to include movement tracking/logs 5 c.5 increase access control to buildings 6 c.2 control access to ucl buildings 7 c.1 ensure appropriate access for each individual 8 d.5 granting access rights 9 d.1 faster issue of cards 10 d.2 reduce queuing time 11 d.3 id card status: ability to check if a user has collected an id card 12 d.6 able to create access reports 13 d.4 easy to replace lost cards 14 d.7 procedures for dealing with fraud 15 b.3 card should be secure 16 b.1 card to include user details 17 b.4 card to include ucl branding 18 b.5 easy identification/card is clear looking 19 b.6 card should be sturdy/robust 20 b.2 card to include barcode 21 b.7 card should be attractive 22 g.1 centralized management of access and identification information 23 g.2 export data to other systems 24 g.3 import data from other systems 25 g.4 data access: able to view, update, delete remotely and securely 26 g.5 clear policies on use of access data 27 e.1 save money on cards 28 e.2 save processing time 29 e.3 reduce paper trials 30 f.5 compatible with current network infrastructure 31 f.6 impact to other systems 32 f.4 compatible with upi 33 f.2 compatible with library systems 34 f.1 compatible with bloomsbury system (gladstone mrm) 35 f.3 compatible with hr system 36 h.3 used for computer logon 37 h.2 include payment mechanism 38 h.4 upgradable (software revisions) 39 h.1 include digital certificate 40 h.5 increase security 41 j.2 conform to standards and legislations 42 j.6 available 43 j.5 reliable 44 a hybrid approach based on whale and grey wolf optimization algorithms 79 copyright ©2018 assa adv. in systems science and appl. (2018) j.3 technology 45 j.1 fail safe 46 j.7 network infrastructure 47 j.4 lifecycle 48 j.9 the chosen manufacturer must have a proven tracker record within institutions with access control, id pass production, odbc, smart card technologies. 49 j.10 photo id pass software must be an embedded feature of the access control software and only require software/license upgrades. 50 j.11 the system manufacturer must be a microsoft™ certified partner. 51 j.12 the system must utilize microsoft™ windows 2000 and/or xp operating system. 52 j.13 the database platform must support microsoft™ sql server and/or oracle server. 53 j.8 be capable of having direct printing to both sides of the card, which will include the library barcode 54 i.3 project management activities 55 i.2 technical documents 56 i.1 supplier support 57 in order to evaluate the proposed technique, the theoretical approach was implemented and tested on the same requirements data set with same parameters, the proposed technique prioritize the set of requirements with 52 matching prioritization values and 5 mismatching values. the error rate is calculated using equation (15) and the accuracy is calculated using equation (16) 𝑬𝒓𝒓𝒐𝒓 𝑹𝒂𝒕𝒆 = 𝒎𝒊𝒔𝒎𝒂𝒕𝒄𝒉𝒊𝒏𝒈/𝒔𝒆𝒕 𝒔𝒊𝒛𝒆 ∗ 𝟏𝟎𝟎 (15) so the error rate of proposed method is approximately 9% so the approximate accuracy is 91% while comparing the proposed method's result with the used requirement list. table 3 shows the ordering comparison between the proposed method and [39] work, and fig.10 shows the error rate and accuracy between the two methods. table 3. comparison of requirements ordering between wgw-rp and work of [39] wgw ordering (req id) original ordering (req id) matching order wgw ordering (req id) original ordering (req id) matching order a.3 a.3 yes e.3 e.3 yes a.1 a.1 yes f.5 f.5 yes a.2 a.2 yes f.6 f.6 yes c.3 c.3 yes f.4 f.4 yes c.4 c.4 yes f.2 f.2 yes c.5 c.5 yes f.1 f.1 yes c.2 c.2 yes f.3 f.3 yes c.1 c.1 yes h.3 h.3 yes 𝑨𝒄𝒄𝒖𝒓𝒂𝒄𝒚 = 𝟏𝟎𝟎% − 𝑬𝒓𝒓𝒐𝒓 𝑹𝒂𝒕𝒆 (16) 80 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) d.5 d.5 yes h.2 h.2 yes d.1 d.1 yes h.4 h.4 yes d.2 d.2 yes h.1 h.1 yes d.3 d.3 yes h.5 h.5 yes d.6 d.6 yes j.2 j.2 yes d.4 d.4 yes j.6 j.6 yes d.7 d.7 yes j.5 j.5 yes b.3 b.3 yes j.3 j.3 yes b.1 b.1 yes j.1 j.1 yes b.4 b.4 yes j.7 j.7 yes b.5 b.5 yes j.4 j.4 yes b.6 b.6 yes j.9 j.9 yes b.2 b.2 yes j.11 j.10 no b.7 b.7 yes j.8 j.11 no g.1 g.1 yes j.10 j.12 no g.2 g.2 yes j.12 j.13 no g.3 g.3 yes j.13 j.8 no g.4 g.4 yes i.3 i.3 yes g.5 g.5 yes i.2 i.2 yes e.1 e.1 yes i.1 i.1 yes e.2 e.2 yes fig.10. error rate and accuracy for wgw-rp vii. conclusion requirement engineering is the most important phase in software development since it dealing with stakeholders and other activities. since the number of requirements is varying for each project the requirements prioritization is an important process in order to deliver a good ordering of project's phases with satisfying stakeholders and end users. in this paper a hybrid approach was suggested by combining the advantages of gwo and woa algorithms as metaheuristic approach in order to prioritize the software project requirements. 0% 20% 40% 60% 80% 100% error rate accuracy error rate accuracy a hybrid approach based on whale and grey wolf optimization algorithms 81 copyright ©2018 assa adv. in systems science and appl. (2018) ralic project's requirements are used to evaluate the proposed method (wgw). the wgw shows 91% accuracy and 9% error rate of prioritizing these requirements compared to work of author [39]. references [1] hasan, m. s., mahmood, a. a., alam, m. j., hasan, s. n., & rahman, f. (2010). an evaluation of software requirement prioritization techniques. international journal of computer science and information security (ijcsis), 8(9). [2] duan, c., laurent, p., cleland-huang, j., & kwiatkowski, c. (2009). towards automated requirements prioritization and triage. requirements engineering, 14(2), 73-89. [3] paetsch, f., eberlein, a., & maurer, f. (2003, june). requirements engineering and agile software development. in proceedings of the twelfth ieee international workshops on enabling technologies: infrastructure for collaborative enterprises (wetice’03) (pp. 308313). ieee. [4] wiegers, k. (1999). first things first: prioritizing requirements. software development, 7(9), 48-53. [5] jain, a. k., murty, m. n., & flynn, p. j. (1999). data clustering: a review. acm computing surveys (csur), 31(3), 264-323. [6] emami, h., & derakhshan, f. (2015). integrating fuzzy k-means, particle swarm optimization, and imperialist competitive algorithm for data clustering. arabian journal for science and engineering, 40(12), 3545-3554. [7] yang, f., sun, t., & zhang, c. (2009). an efficient hybrid data clustering method based on k-harmonic means and particle swarm optimization. expert systems with applications, 36(6), 9847-9852. [8] nayak, j., naik, b., & behera, h. s. (2015). fuzzy c-means (fcm) clustering algorithm: a decade review from 2000 to 2014. in computational intelligence in data miningvolume 2 (pp. 133-149). springer, new delhi. [9] kumar, y., & sahoo, g. (2015). hybridization of magnetic charge system search and particle swarm optimization for efficient data clustering using neighborhood search strategy. soft computing, 19(12), 3621-3645. [10] alpaydin, e. (2004). introduction to machine learning. cambridge: mit press. [11] zhang, q. h., li, b. l., liu, y. j., gao, l., liu, l. j., & shi, x. l. (2016). data clustering using multivariant optimization algorithm. international journal of machine learning and cybernetics, 7(5), 773-782. [12] mirjalili, s., mirjalili, s. m., & lewis, a. (2014). grey wolf optimizer. advances in engineering software, 69, 46-61. [13] mirjalili, s., & lewis, a. (2016). the whale optimization algorithm. advances in engineering software, 95, 51-67. [14] chander, s., vijaya, p., & dhyani, p. (2018). multi kernel and dynamic fractional lion optimization algorithm for data clustering. alexandria engineering journal, 57(1), 267-276. [15] çomak, e. (2016). a modified particle swarm optimization algorithm using renyi entropy-based clustering. neural computing and applications, 27(5), 1381-1390. [16] greer, d., & bustard, d. w. (1997). serum-software engineering risk: understanding and management. the international journal of project & business risk, 1(4), 373-388. 82 a. hudaib, r. masadeh, a. alzaqebah copyright ©2018 assa adv. in systems science and appl. (2018) [17] ma, q. (2009). the effectiveness of requirements prioritization techniques for a medium to large number of requirements: a systematic literature review (doctoral dissertation, auckland university of technology). [18] ibrahim, o., & nosseir, a. (2016). a combined ahp and source of power schemes for prioritising requirements applied on a human resources. in matec web of conferences (vol. 76, p. 04016). edp sciences [19] goldbogen, j. a., friedlaender, a. s., calambokidis, j., mckenna, m. f., simon, m., et al. (2013). integrative approaches to the study of baleen whale diving behavior, feeding performance, and foraging ecology. bioscience, 63(2), 90-100.. [20] humpback whale. (n.d.). in wikipedia. retrieved august 20, 2018, from https://en.wikipedia.org/wiki/humpback_whale [21] daneva, m., damian, d., marchetto, a., & pastor, o. (2014). empirical research methodologies and studies in requirements engineering: how far did we come?. journal of systems and software, 95, 1-9. [22] achimugu, p., selamat, a., ibrahim, r., & mahrin, m. n. r. (2014). a systematic literature review of software requirements prioritization research. information and software technology, 56(6), 568-585. [23] karlsson, j., wohlin, c., & regnell, b. (1998). an evaluation of methods for prioritizing software requirements. information and software technology, 39(14-15), 939-947. [24] brender, s. key words for use in rfc’s to indicate requirements levels. rfc 2119. [25] d. leffingwell & d. widrig. (1999) managing software requirements: a unified approach, upper saddle river:addisonwesley. [26] hatton,s. (2007). early prioritisation of goals. in international conference on conceptual modeling (pp. 235-244). springer, berlin, heidelberg. [27] ieee std 830-1998 (1998) ieee recommended practice for software requirements specifications. ieee computer society, los alamitos [28] sommerville, i., & sawyer, p. (1997). requirements engineering: a good practice guide. john wiley & sons, inc.. [29] regnell b, höst m, natt och dag j, beremark p, hjelm t (2001) an industrial case study on distributed prioritization in market-driven requirements engineering for packaged software. requirements engineering 6(1):51-62 [30] saaty tl (1980) the analytic hierarchy process. mcgraw-hill, new york [31] lehtola, l., & kauppinen, m. (2004). empirical evaluation of two requirements prioritization methods in product development projects. in european conference on software process improvement (pp. 161-170). springer, berlin, heidelberg. [32] karlsson, j., wohlin, c., & regnell, b. (1998). an evaluation of methods for prioritizing software requirements. information and software technology, 39(14-15), 939-947. [33] shen, y., hoerl, a. e., & mcconnell, w. (1992). an incomplete design in the analytic hierarchy process. mathematical and computer modelling: an international journal, 16(5), 121-129. [34] harker, p. t. (1987). incomplete pairwise comparisons in the analytic hierarchy process. mathematical modelling, 9(11), 837-848. [35] karlsson, j., olsson, s., & ryan, k. (1997). improved practical support for large-scale requirements prioritising. requirements engineering, 2(1), 51-60. https://en.wikipedia.org/wiki/humpback_whale a hybrid approach based on whale and grey wolf optimization algorithms 83 copyright ©2018 assa adv. in systems science and appl. (2018) [36] dsdm public version 4.2, from www.dsdm.org, tech. rep., retrieved, 6 june, 2009, last visited, april, 5, 2018. [37] aho, a. v., hopcroft, j. e., & ullman, j. (1983). data structures and algorithms. addison-wesley longman publishing co., inc.. [38] lauesen, s. (2002) software requirements – styles and techniques. pearson education, essex [39] lim, s. l. (2011). social networks and collaborative filtering for large-scale requirements elicitation (doctoral dissertation, university of new south wales). [40] alzaqebah, a., masadeh, r., & hudaib, a. (2018). whale optimization algorithm for requirements prioritization. in proceedings of the 9th international conference on information and communication systems (icics), (pp. 84-89). ieee. [41] masadeh, r., alzaqebah, a. & hudaib, a. (2018). grey wolf algorithm for requirements prioritization. modern applied science, 12(2), 54. [42] hudaib, a., masadeh, r., qasem, m. h., & alzaqebah, a. (2018). requirements prioritization techniques comparison. modern applied science, 12(2), 62. [43] masadeh, r., alzaqebah, a., & sharieh, a. (2018). whale optimization algorithm for solving the maximum flow problem. journal of theoretical & applied information technology, 96(8). [44] masadeh, r., sharieh, a., & sliet, a. (2017). grey wolf optimization applied to the maximum flow problem. international journal of advanced and applied sciences, 4(7), 95100. [45] yassien, e., masadeh, r., alzaqebah, a., & shaheen, a. (2017). grey wolf optimization applied to the 0/1 knapsack problem. international journal of computer applications, 169(5). [46] tarhini, a., ammar, h., & tarhini, t. (2015). analysis of the critical success factors for enterprise resource planning implementation from stakeholders’ perspective: a systematic review. international business research, 8(4), 25. [47] vestola, m. (2010). a comparison of nine basic techniques for requirements prioritization. helsinki university of technology. advances in systems science and applications (2012) vol.12 no.4 388-398 optimal designs in random intercept model with heteroscedastic errors jing cheng1 and rongxian yue2 1department of mathematics, chaohu college, anhui 238000, china 2department of mathematics, shanghai normal university, shanghai 200234, china scientific computing key laboratory of shanghai universities, and division of scientific computation of e-institute of shanghai universities abstract this paper considers optimal designs based on the d-, g-, a-, iand ds-optimality criteria for a random intercept model with heteroscedastic errors. it is shown that the search of optimal approximate designs can be confined at extreme settings of the design region if heteroscedastic structure satisfies specified conditions. closed expressions for the optimal proportions are given. keywords optimal design, random intercept model, heteroscedastic errors, identical design 1 introduction random coefficient models have been widely used for the researching in the area of biosciences, psychology and population pharmacokinetics, where repeated measurements are available from different individuals. these models have been introduced by longford[1], for recent researching we refer to pena and yohai[2] and yu[3]. in recent years, the problem of optimal designs for random coefficient models has attracted growing interest. schmelter[4-5] showed that optimal designs in the linear mixed models could be restricted to the class of group-wise identical designs, and optimal designs in the class of single-group designs were also optimal designs in the larger class of more group designs when the design criteria satisfied some assumptions. schwabe and schmelter[6], schmelter et al[7] and luoma et al[8] . investigated optimal designs in random intercept model, random slope model and random coefficient cubic regression model, respectively. entholzner et al[9] obtained optimal and efficient designs in mixed models. debusho and haines[10] provided v-optimal and d-optimal designs with longitudinal data in linear regression models with a random intercept. there are many other results of optimal designs are obtained, such as wang et al[11], yu[12] and wen et al[13]. in this article, we investigate the problem of optimal designs based on some common optimality criteria for a random intercept model with heteroscedastic errors. in section 2, we introduce the model with necessary notations. section 3 provides a lemma which makes it sure that we can confine the search of optimal designs at extreme settings of the design region if the optimality criteria satisfy an assumption. simple expressions of these optimal advances in systems science and applications (2012) vol.12 no.4 389 designs are given in this section. section 4 introduces some examples. proof of lemma 1 is given in appendix. 2 the random intercept model with heteroscedastic errors we investigate a linear regression model on the unit interval with a random intercept and heteroscedastic errors. it is assumed that there are individuals with observations each, and the jth observation of ith individual is described by yij = µi + xijβ + e(xij), i = 1, ..., n; j = 1, ...,mi. (1) where, xij ∈ [0, 1] is the experimental setting; µi denotes the ith individual effect with unknown mean µ and known variance σ2 µ; β is the unknown slope parameter; observational errors e(xij) are assumed to be heteroscedastic with zero mean and variance σ2/λ(xij), here σ2 is known and λ(xij) is a positive real-valued continuous function defined on [0,1]. we assume that cov (µi, µi′) = 0, i ̸= i′ cov (µi, e(xi′j)) = 0, ∀i, i′; cov (e(xij), e(xi′j′)) = 0, (i, j) ̸= (i′, j′). for the ith individual, denote yi =  yi1 ... yimi  , xi =  xi1 ... ximi  , e(xi) =  e(xi1) ... e(ximi)  , fi = (1mi , xi) here 1mi is a vector of length mi with all entries equal one. then the model (1) can be expressed by yi = fi ( µ β ) + 1mi(µi − µ) + e(xi) , fiθ + 1mi(µi − µ) + e(xi), i = 1, . . . , n. by the assumptions we have (µi − µ) ∼ (0, σ2 µ) and vi , cov (yi) = σ2diag{1/λ(xi1), . . . 1/λ(ximi)}+σ2 µ1mi1 t mi , σ2(di+d1mi1 t mi ). here di = diag{1/λ(xi1), . . . 1/λ(ximi)} and d = σ2 µ/σ 2. for all n individuals, the vector of all observations can be expressed by y =  y1 ... yn  =  f1 ... fn  θ +  1m1 0 . . . 0 1mn   µ1 − µ ... µn − µ +  e(x1) ... e(xn)  (2) 390 jing cheng:optimal designs in random intercept model with heteroscedastic errors the design matrix for random intercepts is block diagonal, e.g., 1m1 0 . . . 0 1mn  . consequently, the covariance matrix of y is cov (y ) = diag{v1, . . . , vn}. the best linear unbiased estimate of θ is given by θ̂ = ( n∑ i=1 f t i v −1 i fi )−1 n∑ i=1 f t i v −1 i yi. (3) and we can get cov (θ̂) = ( n∑ i=1 f t i v −1 i fi )−1 . 3 optimal designs in this section, we investigate the optimal designs based on d-, g-, a-, iand dsoptimality criteria for the models described in previous section. the d-optimal design minimizes the generalized variance of parameter estimates, the g-optimal design minimizes the maximum variance of the predicted value of the response over the design region, the a-optimal design minimizes the total variance of the parameter estimates, the i-optimal design minimizes the integrated mean squared error and the interest of ds-optimal design is in estimating the slope. in some practical situations like human or animal pharmaceutics studies or medical diagnostics there are often restrictions, e.g., technical implementations, which force the experiment to be performed with identical regimes for all individuals. this means that for each individual the number mi of repeated measurements equals m and experimental settings xij = xj are identical across all the individuals. so we only consider identical designs in the following, i.e., mi = m, xi = x1 and hence, fi = f1, vi = v for all i. then the best linear unbiased estimate of θ can be written as θ̂ = ( nf t 1 v −1 1 f1 )−1 f t 1 v −1 1 n∑ i=1 yi. here v −1 1 = 1 σ2 ( d1 + d1m1tm )−1 = 1 σ2 ( d−1 1 − dd−1 1 1m1tmd−1 1 1 + d1tmd−1 1 1m ) = 1 σ2 ( diag{λ(xj)} − dd−1 1 1m1tmd−1 1 1 + d ∑n j=1 λ(xj) ) . advances in systems science and applications (2012) vol.12 no.4 391 without loss of generality, we assume σ2 = 1 in the followings. furthermore, we will consider approximate designs. for any approximate design ξ of the following form ξ = ( x1, . . . , xp ω1, . . . , ωp ) , 2 < p < m, p∑ j=1 ωj = 1. (4) denote νk = ∫ 1 0 xkλ(x)dξ(x) = p∑ j=1 ωjx k jλ(xj), k = 0, 1, 2. then the information matrix corresponding to the design ξ of the form (4) can be expressed by m(ξ) = mn 1 + γν0 ( ν0 ν1 ν1 ν2 + γ(ν0ν2 − ν21) ) . (5) here we note γ = md. for regression model without any random effects, optimal designs are obtained at extreme settings of the design region and schwabe et al[6] discussed optimal designs of random intercept models. we can’t use the conclusions in schwabe et al[6] directly in the random intercept model with heteroscedastic errors, but we have the following lemma. lemma 1 in the model (2), assume that mi = m, (i = 1, . . . , n) and λ(x) satisfies the following condition 1 λ(x) ≥ 1− x λ0 + x λ(1) , x ∈ [0, 1]. (6) where λ0 = λ(0) and λ1 = λ(1). then for any approximate design ξ of the form (4), there exists an approximate design of the form ξ∗ = ( 0, 1 1− ω, ω ) , 0 < ω < 1. such that m(ξ∗) ≥ m(ξ). the proof of lemma 1 can be found in the appendix. the criteria, φ(·), considered in this paper are functions of the information matrices which are required to satisfy the following assumptions: a1 φ(·), is a real-valued function defined on the whole set m of 2×2 symmetric non-negative definite matrices, φ : m → (−∞,∞]; a2 φ(·) is monotone (the loewner order (e.g., pukelsheim[14], p.101)) on m in the sense that m1,m2 ∈ m,m∞ ≥ m∈ ⇒ ⊕(m∞) ≤ ⊕(m∈). 392 jing cheng:optimal designs in random intercept model with heteroscedastic errors these assumptions are satisfied for most of the common criteria including the d-, g-, a-, iand ds-optimality. so, by majorization we can confine the search of optimal designs at extreme settings x = 0 and x = 1 if λ(x) satisfies the condition (6) in lemma 1. therefore, in what follows we only consider approximate designs ξ of the form ξ = ( 0 1 1− ω ω ) . (7) for approximate designs of the form (7), we have ν0 = λ1ω + λ0(1− ω), ν1 = ν2 = λ1ω. first, we consider the d-optimality φ(m(ξ)) = ∣∣m−1(ξ) ∣∣. note that |m(ξ)| , |m(ω)| = (mn)2λ0λ1ω(1− ω) 1 + γ[ωλ1 + (1− ω)λ0] . it is easy to verify that |m(ω)| is maximized at ω = √ 1 + γλ0/( √ 1 + γλ0 + √ 1 + γλ1) and hence ∣∣m−1(ω) ∣∣ is minimized. therefore we have theorem 1 for the model (2) with mi = m (i = 1, . . . , n) and λ(x) satisfying (6) the d-optimal design is ξ∗d = ( 0, 1 1− ωd, ωd ) , ωd = √ 1 + γλ0√ 1 + γλ0 + √ 1 + γλ1 for the ds-optimality, note that the covariance matrix of θ̂ can be calculated by m−1(ω) = 1 mn(ν0 − ν2) ( 1 + γ(ν0 − ν1) −1 −1 ν0 ν1 ) = 1 mn ( 1 λ0(1−ω) + γ − 1 λ0(1−ω) − 1 λ0(1−ω) 1 λ0(1−ω) + 1 λ1(ω) ) the variance of the estimate for β is given by cov (β̂) = [m−1(ω)]22 = 1 mn [ 1 λ0(1− ω) + 1 λ1ω ]. the variance of β̂ is minimized at ω = √ λ0/( √ λ0 + √ λ1). therefore we have theorem 2 for the model (2) with mi = m (i = 1, . . . , n), and λ(x) satisfying (6) the ds-optimal design is ξ∗ds = ( 0, 1 1− ωds , ωds ) , ωds = √ λ0√ λ0 + √ λ1 . advances in systems science and applications (2012) vol.12 no.4 393 consider the g-optimality, then φ(m(ω)) = max x∈[0,1] d(ω, x). here d(ω, x) is the variance of the predicted value of the response, which is given by d(x, ω) = ( 1 x ) m−1(ω) ( 1 x ) = 1 mn {[ 1 λ0(1− ω) + 1 λ1ω ] x2 − 2x λ0(1− ω) + 1 λ0(1− ω) + γ } . as d(ω, x) is a polynomial of degree 2 with positive leading term, its maximum is attained either x = 0 or x = 1 or both, i.e.,max d(ω, x) = max x∈[0,1] {d(0, ω), d(1, ω)} . note that d(0, ω) = 1 mn { 1 λ0(1− ω) + γ } is strictly increasing in ω, d(1, ω) = 1 mn { 1 λ1(ω) + γ } is strictly decreasing in ω. thus min ω∈[0,1] max x∈[0,1] d(x, ω) is attained when d(0, ω) = d(1, ω), i.e., 1 λ0(1− ω) = 1 λ1(ω) . so we have theorem 3 for the model (2) with mi = m (i = 1, . . . , n), and λ(x) satisfying (6) the g-optimal design is ξg ∗ = ( 0, 1 1− ωg, ωg ) , ωg = λ0 λ0 + λ1 . for the i-optimality, φ(m(ω)) = ∫ 1 0 d(x, ω)dx = 1 mn [ γ + 1 3λ0(1− ω) + 1 3λ1ω ] . which is minimized at ω = √ λ0/( √ λ0 + √ λ1) = ωds . so we have theorem 4 for the model (2) with mi = m (i = 1, . . . , n), and λ(x) satisfying (6) the i-optimal design is ξ∗i = ( 0, 1 1− ωi , ωi ) , ωi = √ λ0√ λ0 + √ λ1 . 394 jing cheng:optimal designs in random intercept model with heteroscedastic errors for the a-optimality φ(m(ω)) = tr ( m−1(ω) ) = 1 mn [ 1 λ1ω + 2 λ0(1− ω) + γ ] . it is easy to verify that tr ( m−1(ω) ) is minimized at ω = √ λ0√ λ0+ √ 2λ1 .therefore we have theorem 5 for the model (2) with mi = m (i = 1, . . . , n), and λ(x) satisfying (6) the a-optimal design is ξa ∗ = ( 0, 1 1− ωa, ωa ) , ωi = √ λ0√ λ0 + √ 2λ1 . from above discussion, we observe that the g-, ds-, iand a-optimal designs only depend on the variances at extreme settings; the d-optimal design depends repeated times m and variance proportion d and error variances at the extreme settings. note that the particular shape of λ(x) is immaterial for the results, but only its values at 0 and 1, as long as condition (6) is satisfied. specially, when heteroscedastic structure satisfies λ0 = λ1 = max x∈[0,1] λ(x) , the optimal designs discussed above are independent of the variance ratio d. these optimal designs are the same as the corresponding optimal designs in the linear regression model without any random effects, i.e., ωd = ωg = ωi = ωds = 1 2 , ωa = √ 2− 1. 4 examples in this section, we consider three random intercept models with the following heteroscedastic errors λ(x) = x2 + 1, λ(x) = 1 1 + x , λ(x) = x4 + 1 x2 + 1 . it is easy to verify that these three λ(x) satisfy the condition (6). these heteroscedastic structures are also considered in chang[15] for d-optimal designs in weighted polynomial regression models. we will give the optimal designs for the three models in terms for the results given in section 3. we also compare the dand g-optimal designs for the three models with the equireplicated design ω0 = 0.5 which is simultaneously dand g-optimal for the fixed effects only model (d = 0) in terms of the dand g-efficiency which are defined as following effd(ω0) = ( |m(ω0)| |m(ω)| ) 1 2 , effd(ω0) = max x∈[0,1] d(x, ωg) max x∈[0,1] d(x, ω0) (8) advances in systems science and applications (2012) vol.12 no.4 395 example 1 for the model (2) with mi = m(i = 1, . . . , n), and λ(x) = x2 + 1 , from the theorems in section 3, we obtain the d-, g-, a-, iand ds-optimal proportions as follows: ωd = √ 1 + γ√ 1 + γ + √ 1 + 2γ , ωg = ωa = 1 3 , ωi = ωds = √ 2− 1. the dand g-efficiencies defined by (8) of the equireplicated design ω0 = 0.5 are as follows: effd(ω0) = √ 1 + 2γ + √ 1 + γ√ 4 + 6γ , effd(ω0) = γ + 1.5 γ + 2 . it is clear that effd(ω0) decreases strictly in γ and ultimately tends to ( √ 2 + 1)/ √ 6 , and effd(ω0) increases strictly in γ and ultimately tends to one. fig.1 shows the plots of these two efficiencies. fig.1 the efficiencies of effd(ω0) and effg(ω0) with different λ example 2 for the model (2) with mi = m(i = 1, . . . , n),and λ(x) = 1 1+x , the d-, g-, a-, iand ds-optimal proportions as follows: ωd = √ 2 + 2γ√ +γ + √ 2 + 2γ , ωg = 2 3 , ωa = 1 2 , ωi = ωds = 2− √ 2. the dand g-efficiencies defined by (8) of the equireplicated design ω0 = 0.5 are as follows: effd(ω0) = √ 2 + 2γ + √ 2 + γ√ 8 + 6γ , effd(ω0) = γ + 3 γ + 4 . it is clear that effd(ω0) decreases strictly in γ and ultimately tends to ( √ 2 + 1)/ √ 6 = 0.9856, and effd(ω0) increases strictly in γ and ultimately tends to one. fig.2 shows the plots of these two efficiencies. 396 jing cheng:optimal designs in random intercept model with heteroscedastic errors fig.2 the efficiencies of effd(ω0) and effg(ω0) with different λ example 3 for the model (2) with mi = m (i = 1, . . . , n), and λ(x) = x4+1 x2+1 , the d-, g-, a-, iand ds-optimal proportions as follows: ωd = ωg = ωi = ωds = 1 2 , ωa = √ 2− 1. that is, the d-, g-, iand ds-optimal designs are all the equireplicated designs. appendix proof of lemma 1 from liski et al[16], we get m−1(ξ) = m−1 0 (ξ) + ( d n 0 0 0 ) here m0(ξ) is the corresponding generalized information matrix when there are no individual intercepts, i.e., m0(ξ) = mn ( ν0 ν1 ν1 ν2 ) . let the proportion ω in ξ∗ be of the form ω = ν1/λ1. it follows that m0(ξ ∗) = mn ( ν∗0 ν∗1 ν∗1 ν∗2 ) , and m0(ξ ∗)−m0(ξ) = mn ( ν∗0 − ν0 0 0 ν∗2 − ν2 ) . here ν∗0 = λ1ω + λ0(1− ω) and ν∗2 = ν∗1 = λ1ω . since 1 λ(x) ≥ 1− x λ0 + x λ1 ≥ x λ1 , advances in systems science and applications (2012) vol.12 no.4 397 so λ1 ≥ xλ(x). it implies 0 < ω < 1. [ m0(ξ ∗)−m0(ξ) ] 11 = mn [ λ1ω + λ0(1− ω)− p∑ j=1 ωjλ(xj) ] = mn p∑ j=1 ωj [ λ(xj)xj(λ1 − λ0)− λ1λ(xj) + λ1λ0 ] condition (6) implies λ(x)x(λ1 − λ0)− λ1λ(x) + λ1λ0 ≥ 0. so we obtain [ m0(ξ ∗)−m0(ξ) ] 11 ≥ 0 by ν∗2 = ν∗1 = ν1 ≥ ν2 , we have[ m0(ξ ∗)−m0(ξ) ] 22 ≥ 0 so we get m0(ξ ∗) ≥ m0(ξ) and hence m(ξ∗) ≥ m0(ξ). acknowledgements this work was partially supported by a nsfc grant (11071168), special funds for doctoral authorities of education ministry (20103127110002), e-institutes of shanghai municipal education commission (e03004), shanghai leading academic discipline project (s30405), the innovation program of shanghai municipal education commission (11zz116), and the scientific research foundation of chaohu college. references [1] n.t. longford (1993), random coefficient regression models, clarendon press, oxford. [2] d. pena, v. yohai (2006), “a dirichlet random coefficient regression model for quality indicators”, journal of statistical planning and inference, vol.136, pp.942-961. [3] f. yu (2007), “quadratic design criterion for nonlinear models”, advances in systems science and applications, vol.7, no.2, pp.155-160. [4] t. schmelter (2007), “the optimality of single-group designs for certain mixed models”, metrika, vol.65, pp.183-193. 398 jing cheng:optimal designs in random intercept model with heteroscedastic errors [5] t. schmelter (2007), “consideration on group-wise identical designs for linear mixed models”, journal of statistical planning and inference, vol.137, pp.4003-4010. [6] r. schwabe, t. schmelter. (2008), “on optimal designs in random intercept models”, tatar mountains mathematical publications, vol.39, pp.189-195. [7] t. schmelter, et al. (2007), “some curiosities in optimal designs for random slopes”, in: contributions to statistics. moda 8 advances in modeloriented design and analysis,physica-verlag heidelberg, pp.189-195. [8] a. luoma, et al. (2007), “optimal designs in random coefficient cubic regression models”, journal of statistical planning and inference, vol.137, pp.3611-3617. [9] m. entholzner, et al. (2005), “a note on designs for estimating population parameter”, listy biometryczne-biometrical letters, vol.42, pp.25-41. [10] l.k. debusho, l.m. haines (2007), “vand d-optimal population designs for the simple linear regression model with a random intercept term”, journal of statistical planning and inference, vol.138, pp.1116-1130. [11] g. z. wang, j. zhao, j. b. chen (2006), “the uniform design modeling and analysis”, advances in systems science and applications, vol.6, no.4, pp.699-703. [12] s.h. yu (2007), “the linear minimax estimator of stochastic regression coefficients and parameters under quadratic loss function”, statistics and probability letters, vol.77, pp.54-62. [13] g.y. wen, et al. (2008), “orthogonal experimental design in matrix form and its application to the introduction selection of thymus genus”, advances in systems science and applications, vol.8, no.3, pp.437-446. [14] f. pukelsheim (1993), optimal design of experiment, johnwiley, new york. [15] f.c. chang (2005), “d-optimal designs for weighted polynomial regression– a functional approach”, annals of the institute of statistical mathematics, vol.57, no.4, pp.833-844. [16] e.p. liski, n.k. mandal, k.r. shah, b.k. sinha (2002), topics in optimal designs, springer, new york. corresponding author nicholas nechval can be contacted at:yue2@shnu.edu.cn advances in systems science and applications (2013) vol.13 no.4 334-354 agent based modeling of integration of organizational cultures in mergers and acquisitions albert r. bakhtizin1 and svetlana v. denisova2 1central economics and mathematics institute of russian academy of sciences, moscow, russia 2moscow aviation institute (national research university), moscow, russia abstract in the article a new approach to the analysis of compatibility of organizational cultures during merges and acquisition using the agent based simulation model (abm) is considering. the choice of this method is substantiated. using abm makes the considerable economy of time possible. usually companies spend a lot of time on estimation of possible results of merge. also abm allows promptly change the initial parameters of the model. results of the research experiments are presented in the article. they showed that the agent based simulation model can be successfully used as a tool of prompt analyzing and prediction of the result of organizational culturesąŕ integration. in addition to this it can take into account features of each culture separately. keywords agent based model, modeling organizational culture, merger and acquisition, culture integration, corporate mutual relations 1 introduction the number of mergers and acquisitions is growing up every year. russian companies also are using potential of such transactions for development of their successful businesses. mergers and acquisitions open opportunities to enter new markets, gain access to promising modern technologies, and transition to the new, more promising economic sectors. regular changes of economic environment, together with increase of information flows with fast access to data and acquisition of new knowledge, create preconditions for development of internal sources of economic growth, which allow the company to operate progressively in rapidly changing environment. organizational culture of company provides resources designed to support flexible, adaptive and efficient business systems. organizational culture determines how and in what way created and controlled business processes are. organizational culture creates foundation for joint activities of company team and formation of organizational culture, corresponding to the current external environment, and it also provides effective organizational development of the company. on the other hand, study of organizational culture itself makes it possible to obtain objective assessment of many processes occurring in the company, which becomes especially important in the process of implementing mergers and acquisiadvances in systems science and applications (2013) vol.13 no.4 335 tions, when merging of companies with already established organizational cultures greatly increase resistance of their personnel to conducted organizational changes. influence of organizational culture upon results of merger or acquisition deal is often underestimated. organizational culture, as a cohesive core of company, can produce significant impact on many aspects of company activities, promoting or delaying its development. it depends primarily on characteristics of the culture itself and its compliance with current situation and goals, in the context of which it manifests. creating of organizational culture of such type, which will be more consistent with structure and goals of the company formed as the result of merger or acquisition, is one of challenges that must be met at an early stage of planning of the deal. large-scale changes in mergers and acquisitions initiate resistance of the personnel of merging companies to these changes, which may result in substantially lower to expected effectiveness of merger or acquisition deal. solution to this problem can become establishment of informational support of the coming changes, when employees are regularly informed about all events inside the company, which creates a favorable working atmosphere and greatly impedes spread of rumors and thus mitigates negative reaction of team. as far as possible, company employees should be involved in reorganization process, which significantly increase the role of communications. but if resistance of personnel to changes in the period of integration is a special case of reaction of employees to any organizational changes, and the ways to solution of this problem are indicated, than solution to wide-range problems still is not fully determined. in this way, the potential conflict of organizational cultures should be analyzed at the early stage of merger or company takeover deal. but in practice, to conduct such analysis is very problematic case. the main difficulty is that it is impossible to predict the future culture of a new company after merger deal and to evaluate its effectiveness. options may vary depending on many factors, like : organizational cultures type prior to merging, their strength and level, similarity and ability to changes. according to consultants any change of organizational culture requires at least three years. in other words, it takes a long time for assessment of the results of integration of organizational cultures. while if in the process of cultural integration were made mistakes, than negative results show up only with time, and to make adjustments and get the desired results would take more time and additional capital inputs. the purpose of all the mergers and acquisitions transactions, without any exception, is integration carried-out successfully. integration is understood as element joining-up resulted in forming a single whole. as a rule, the process of merging and acquiring companies, integration, results in creation of a new organizational structure capable to dispose available resources in a way more efficient 336 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... and rational, to optimize material, financial, labor and informational flows of the integrated companies. statistical data show that 70% cases of prospectively advantageous transactions fail because of the low-quality preparation and carrying out integration. so, selection of an organizational form for companies integration in compliance with the purpose and objectives stated is an important step that can facilitate to unify resources and flows efficiently. integration is carried out in various directions. though, whereas corporate strategy integration, product series management systems, product distribution and delivery management systems are predictable and subject to numerical computation on the base of data available, organizational culture integration is often unpredictable. integration of united enterprises organizational cultures is a crucial component in forming the corporate interaction efficiently. resistance exhibited by the personnel to the process in the period of companies mergers and acquisitions is conditioned, first of all, by the difference among various (in the majority of cases) corporate values and organizational cultures. the mergers of ąőequalsąŕ finishes frequently with the matter that a stronger group imposes their organizational culture on a unilateral basis. there is no doubt, mergers and acquisitions propose a lot of advantages for business development. nevertheless, some erroneous illusion might be created that such transactions are relatively easy, inexpensive and represent the only way to increase business considerably. however, most investigations on mergers and acquisitions efficiency evidence that 60 to 80% companies even armed with potentially advantageous strategy do not accomplish the objectives. it is often concerned with mistakes committed in the course of enterprise integration as well as incorrect organization of the transaction itself. mistakes and inadvertences can appear in every phase of mergers and acquisitions. thus, in addition to the incorrectly chosen object for merger and acquisition, low-quality preparation for transactions, and erroneous financial estimation, there could be chosen erroneous ways to implement integration. particularly, serious mistakes can be made in the course of changing and forming the enterprise organizational culture resulted from mergers or acquisitions. in this phase, the mistakes committed can be conditioned by the lack of detailed integration plan and the lack of appropriate approach. the main cause for most failed mergers is inconsistence and incompatibility of organizational cultures. the cause for most failed mergers and acquisitions consists in the joining companies inability to overcome organizational culture contradictions. cultural problems are to be solved even for most successful mergers. in other words, cultural and organizational problem solving acquires the utmost importance for each integration, both successful and failed one. although this area has not been investigated advances in systems science and applications (2013) vol.13 no.4 337 properly yet, experts working in this area stated that enterprises which completed integration successfully had paid a lot of attention to the following aspects : management team forming : how to target the management ranks to the tasks issued by the director general and the board ; organizational structure : how to create a structure that would mostly correspond to the new enterprise strategy ; highly-efficient culture : how to work-out and develop a culture that would facilitate for efficiency increase and would help the new enterprise to realize their long-term objectives ; expert employee administration : how to reveal the most various employees in both enterprises and what actions to undertake for involving them into the process of new enterprise creation. in order to avoid mistakes, it is necessary to elaborate a plan for mergers and acquisition procedure, and, as it has already been noted and never before stated, the question on organizational culture and emerging companies personnel is to be formulated in the phase of choosing the object for mergers and acquisitions. in case the merged enterprises remain existing independently from each other, as a rule, there are no big problems with the personnel. however, if the enterprises start functioning as a single whole, the question on organizational cultures integration becomes exceptionally acute. the new organizational culture could not be acquired by simply joining two old cultures. in theory, if promoting people, introducing new values, orientations, introducing new behavioral models, remunerating them, forming models to emulate, these behavioral models will be repeated and will be fixed in the personnel minds. this way, a new culture is being created. though, in practice, organizational culture integration is rather complicated, as far as it is a multiphase complex process that needs comprehensive planning and accurate implementation. detailed calculations along with competent and thoughtful actions of the managing personnel will guarantee the successfully implemented integration. nevertheless, the potential of organizational culture confrontation is rarely analyzed in this phase, preceding the mergers or acquisition transaction. as a result, the culture confrontation creates a serious problem for merger procedure. from this we can conclude that during the period of mergers and acquisitions the most important is issue of compatibility of organizational cultures of the merging organizations. and at planning stage of the deal, for future development of effective merging strategy, it is very important to be aware of possible results of integration of cultures. at the moment there is no tool of diagnostics of organizational cultures compatibility, which can during short period of time to play off a large number of possible scenarios for merging organizational cultures and get expected results of 338 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... integration, in order to subsequently select among all possible variants the most optimal one. in this regard, the main objective of study is building of tool that allows during short period of time to study large number of possible scenarios of merging of organizational cultures targeted on getting of expected integrated culture, and then choose from all options the most suitable one. in this way, one of the objectives of study is task of taking into account individual characteristics of micro level individual agents in order to obtain more realistic assessment of impact of organizational changes on production indicators. as the above mentioned tool, we decided to develop agent-oriented computer simulation model, related to the class of models based on individual behavior of agents. due to its advantages, namely : 1) ability to simulate close to reality system, 2) emergency, 3) flexibility, 4) possibility of specification of model parameters without knowledge of global dependencies within the frameworks of simulating of relevant subject area, we assume that the agent based model will allow us to achieve objectives of study. 2 literature review it should be noted that in russia an agent-oriented models have been developed relatively recently and their development is mainly concentrated in the central economics and mathematics institute (cemi ras) (under leadership of academician v.l. makarov), as well as the company’s “xj technologies” (st. petersburg). for more details about theoretical aspects of this model class see articles of v.l. makarov and a.r. bakhtizin [1]. in the world practice there is some experience in developing agent-oriented models of organizational culture and corporate relations. examples of the most advanced works are shown below. a multi-agent simulation platform for modeling perfectly rational and boundedrational agents in organizations [2]. this paper presents an agent-based simulation framework for the analysis of the equilibria that emerge in a complex structure such as an organization ; we can think of some of these equilibria as corporate culture. authors concentrate on modeling the effort exerted by heterogeneous agents in an organization, and how the interaction between them may lead to a common level of effort (corporate culture). the simple model authors propose is a system in which agents interact in a dynamic, adaptive and evolving way. the model shows how different compositions of the population may lead the system to different common behaviours ; the implications of result findings are both descriptive and normative, and shed advances in systems science and applications (2013) vol.13 no.4 339 light on some core problems of the economics of organization design. how groups can foster consensus : the case of local cultures [3]. this work is based on an idea that a local culture denotes a set of rules on business behaviour among firms in a cluster. similar to social norms or conventions, it is an emergent feature of interaction in an economic network. to model its emergence, authors consider a distributed agent population, representing cluster firms. the model introduces a feedback mechanism of agent behaviour and in-group structure. studying its consequences by means of agent-based computer simulations, authors find that for narrow-minded agents the feedback mechanism helps find consensus more often, whereas for open-minded agents this does not necessarily hold. overall, the dynamics of agent interaction in clusters as modelled here, are conducive to consensus among all or a majority of agents. computer mediated communication and organizational culture : an agentbased simulation model [4]. this paper examines the mutual relationship between the organizational use of computer mediated communication and organizational culture. computer mediated communication supplements communication among members of an organization to maintain the culture, especially when those persons cannot communicate by other means. on the other hand, a strong organizational culture allows a more effective use of computer mediated communication by providing members with some of the necessary common ground to better understand the information exchanged. these relationships are investigated using an agent-based model. this agent-based model incorporates many partial theories into a coherent and fully defined model, which helps formalize and integrate those theories. in this paper, authors present some of the results of the agent-based model that show that organizational culture can influence the effectiveness of computer mediated communication and that computer mediated communication can help maintain and stabilize a culture. social construction of organizational culture : an agent-based model [5]. this model, called “orgnorms”, assumes that culture is important to organizations, and companies in particular, on two levels. first, the homogeneity of an organization’s culture affects communication and efficiency. second, the ‘cultural fitness’ of an organization to local society affects its competitive advantage. orgnorms is an agent-based model designed to simulate the development of and changes in organizational culture in a culturally changing society, tracking the organization’s internal homogeneity and its external fitness to its societal environment. a central assumption is that homogeneity demands that new members adopt the organizational norms, while fitness demands that the organization adopts the views of the new members. that connection breeds similarity, that knowledge is local, and that agents take after those who are similar and the local 340 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... majority, are other key assumptions. these examples of agent-oriented simulation models related to organizational culture show that aom can be successfully applied in studies and forecasting of various processes of transformation of organizational culture. 3 agent-based modeling of organizational mergers with regard to experience of our colleagues from russia and other countries, in 2010 we developed an agent-based model of organizational mergers, as is described below. 3.1 characteristics of agent . table 1 the characteristic parameters of agent n/n parameter description parameter value 1 age from 18 to 60 years 1.1 first age group 20-25 years 1.2 second age group 25-40 years 1.3 third age group 40-50 years 1.4 fourth age group 50-60 years 2 marital status 0 agent has no family, 1 agent has a family 3 professionalism of agent 3.1 education 0 secondary, 1 higher 3.2 experience from 1 to 40 years 3.3 work experience in particular organization before merging from 1 to 10 years 4 loyalty from 1 to 10 points, where minimum value of parameter means intolerance to the values of company, and 10 points means his strong commitment 5 ability to adaptation from 1 to 10 points 6 satisfaction with working conditions after integration 1 to 10 points 7 labor market demand of agent in times of integration from 1 to 10 points 8 ability to work from 1 to 10 points (1). age. for these parameters, a distribution among four age groups : 20-25 years ; 25-40 years ; 40-50 years ; 50-60 years was specified. (2). marital status. it includes : has no family, has a family. advances in systems science and applications (2013) vol.13 no.4 341 (3). professionalism of agent, which consists of concepts such as : education(c1), experience(c2), work experience in particular organization before merging(c3). these variables during initializing of models have random values with standard deviations enclosed in brackets. in order to create the model, 8 main parameters were chosen : age, marital status, professionalism of agent, loyalty, ability to adaptation, satisfaction with working conditions after integration, labor market demand of agent in times of integration, ability to work. input parameter data are given in table 1. professionalism is determined by value within intervals from 0 to 100 points, as function of three components (c1, c2, c3) in the following as p = 33.3 · c1 + 33.3 · c2 40 + 33.3 · c3 10 (1) i.e. in case of maximum values of all components the level of professionalism of agent is also maximum close to 100 points. (4). loyalty. in this case, a loyal employee should share the core beliefs and values of the company (from 1 to 10 points, where minimum value of parameter means intolerance to the values of company, and 10 points means his strong commitment). (5). ability to adaptation. it also can be divided from 1 to 10 points. under adaptation we understand mutual adjustment of employee and company, which is based on gradual involvement of worker into labor activities in new professional, psychophysiological, psychosocial, organizational, administrative, and economic conditions. (6). satisfaction with working conditions after integration. it also can be divided from 1 to 10 points. (7). labor market demand of agent in times of integration. it also can be divided from 1 to 10 points. (8). ability to work. it also can be divided from 1 to 10 points. 3.2 characteristics of environment for functioning of agents and organizational cultures we come to the description of environment of agent model. the agent functioning medium comprises four organizational culture types clan culture, adhocratic culture, hierarchical culture and market culture. for determination of environment we used a simplified version of methodology for assessing of organizational culture, and below is brief description of four types of organizational cultures, used in probability function, which graph is shown in fig.5. type 1. clan culture. it means very friendly working place where people have much in common. companies are like big families. leaders or chiefs of companies 342 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... are perceived as educators, and perhaps even as parents. the company is cemented with loyalty and traditions. there is high responsibility of the company. it is focused on a long-term benefits for improving of individual, pays attention to a high degree of team unity and morale. success is measured in terms of good relations to consumer and care about people. the company encourages teamwork, participation of people in business and mutual consent. type 2. adhocratic culture. it means a dynamic, entrepreneurial and creative working place. people are willing to offer their own support and stand the risks. the leaders are considered innovators and persons who are ready to stand the risks. cohesive spirit of this company is devotion to experiments and innovation. necessity of work on the business forefront is emphasizes as need. the general long term policy of such company is its growth and acquiring of new resources. success means production and performance of unique and new products and / or services. it is important to be a leader on the products and services markets. such organization encourages individual initiative and freedom. type 3. hierarchical culture (bureaucratic). it is very formalized and structured place of work. procedures dominate activities of employees. leaders are proud of the fact that they are rationally minded coordinators and organizers. it is critical to maintain smooth running of the company. company is united with formal rules and official policies. long-term concern of organization is to provide stability and smooth running performance of cost-effective operations. success is measured in terms of supply, smooth schedules and lower costs. management is concerned about employment status and long-term predictability of employees. type 4. market culture. this kind of company is focused on results, the main concern of which is performance of task. people are ambitious and compete with each other. chiefs are hard leaders and tough competitors. they are unshaken and demanding. this company is united together with emphasis on the desire to win. reputation and success are things of common concern. focus of strategy is targeted to a specific action, achievement of tasks and measurable goals. success is measured in terms of markets penetration and increase of market share. important is competitive pricing and leadership on the markets. the working style of this company is hard line targeted on competition. 3.3 agent behavior behavior of agents is specified by diagram of state (state chart), transitions inside of which depend on the values of probability functions listed below. since within the frameworks of model occurs absorption of one company by another, the agent has two options : adapt to the new conditions or leave (for simplicity it is assumed that after reorganization the absorbing company does not change its type, and, on the other hand, the absorbed company takes leading style of absorbing organization). this chart shows process in the following way : agents advances in systems science and applications (2013) vol.13 no.4 343 of absorbed organization (fig.1 “organization 2”) through conversion pass to the absorbing organization (“organization 1”) or have to leave (this transition is shown along the arrow directed towards the ring with dot in center). in the process of work this model simulates the process of absorption, and some time after reorganization, when agent may resign (i.e. it is another transition along the arrow directed towards the ring). fig.1 state chart of model agent for simplicity, we do not consider optimization of personnel, i.e. possible redundancy of employees by company. next, we go to the more detailed description of agent state chart. first of all in the state chart of transition may work out transition 1, depending on agent’s loyalty towards values of company (in this case loyalty parameter of all agents is relevant only to absorbing company). if company’s values are alien to the agent, than he can adapt to them, depending on the values of corresponding parameter (i.e. may work out transition 2). the behavior of agent may be adjusted depending on other parameters. for example, if qualification of agent is in high demand on labor market, than probability of his resignation is high (transition 3). otherwise “thing, which can change up his mind,” may be his age (transition 4), as well as having a family (transition 5). after the process of absorption the agent may stay unsatisfied with new con344 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... fig.2 probability (y-axis) of transition of agent into absorbing organization based on his level of loyalty (x-axis) fig.3 probability (y-axis) of transition of agent into absorbing organization based on his adaptation ability (x-axis) fig.4 probability (y-axis) of transition of agent into absorbing organization based on the level of labor market demand in times of integration (x-axis) advances in systems science and applications (2013) vol.13 no.4 345 fig.5(a) probability (y-axis) of transition of agent (working in company with organizational culture of first type) into absorbing company, depending on type of organizational culture (x-axis) fig.5(b) probability (y-axis) of transition of agent (working in company with organizational culture of second type) into absorbing organization, depending on type of organizational culture (x-axis) fig.5(c) probability (y-axis) of transition of agent (working in company with organizational culture of third type) into absorbing company, depending on type of organizational culture (x-axis) 346 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... fig.5(d) probability (y-axis) of transition of agent (working in company with organizational culture of fourth type) into absorbing company, depending on type of organizational culture (x-axis) fig.6 probability (y-axis) of transition of agent into absorbing company based on age (x-axis) fig.7 probability (y-axis) of transition of agent into absorbing company based on satisfaction with working conditions after integration (x-axis) ditions, and in this case he may initiate transition 6 (this applies to employees of advances in systems science and applications (2013) vol.13 no.4 347 absorbed organization, only). transitions in state chart can work out depending on values of probability functions, defined by experts. fig.2 shows a graph of probability function, where the argument (x-axis) is level of loyalty and value of function (y-axis) probability of transition 1. fig.3 and 6 show graphs with probabilities of working of transitions 2 and 4, depending on capacity for adaptation and age of agent, respectively. for determination of probability of transition 3 we must calculate average probability on basis of functions, whose graphs are shown in fig.4 and 5(a)-5(d). 3.4 adequacy of model for testing of model adequacy we have conducted three experiments on merging of different types of companies, whose parameters are provided in tables 2-4. table 2 parameters of absorbing and absorbed companies for experiment no.1 parameter absorbing company absorbed company type of organizational culture bureaucratic market number of employees 20 000 20 000 age of employees 20-25 years 25% 20-25 years 60% 25-40 years 40% 25-40 years 20% 40-50 years 20% 40-50 years 10% 50-60 years 5% 50-60 years 10% table 3 parameters of absorbing and absorbed companies for experiment no.2 parameter absorbing company absorbed company type of organizational culture bureaucratic market number of employees 300 30 age of employees 20-25 years 25% 20-25 years 15% 25-40 years 50% 25-40 years 50% 40-50 years 20% 40-50 years 30% 50-60 years 5% 50-60 years 5% for experiments, were collected data on transactions of mergers and acquisitions during the period from 2004 to 2009. companies involved in this experiment represent large retail sector and it-sphere. in the first experiment, two organizations having manpower strength of about 20 000 persons each, were represented ; each company had associated companies. the experiment took into account the total number of employees for each enterprise including associate companies and representative offices. 348 albert r. bakhtizin : agent based modeling of integration of organizational cultures in ... table 4 parameters of absorbing and absorbed companies for experiment no.3 parameter absorbing company absorbed company type of organizational culture bureaucratic market number of employees 50 30 age of employees 20-25 years 20% 20-25 years 18% 25-40 years 60% 25-40 years 70% 40-50 years 16% 40-50 years 10% 50-60 years 4% 50-60 years 2% in the second experiment, the absorbing company before reorganization had 300 people, the merged company had 30 people. in the third experiment, small enterprises took part ; the absorbing company personnel counted 50 persons for when transaction started and the merged company personnel counted 30 persons. bellow in listing 1 adduces a code generated by anylogic program which is responsible for agent initialization and model initial state initialization. all the agent populations (people.size()) are distributed by groups (age groups, as well as agent groups pertaining to the first or the second organization) depending on the values preset (initial number of employees working for the first (agentfirm1) and the second organization and, also, four age frames. (age1, age2, age3, age 4)). for (int i = 0 ; i m1,m2,m3,m4,m5 t14 +z ,−z m1,m2,m3,m4,m5 t11 +z ,−z opt 7 fig.2 the hierarchical representation of process information based on an input part and stock used. advances in systems science and applications (2014) vol.14 no.1 41 3.1.3 production monitor agent (pma) due to the dynamic and stochastic characteristic in nature, a job shop production always faces different kinds of uncertainties (e.g. machine breakdown and rush order), leading to a predictive schedule not applicable. thus, in order to minimize the impact of the disruption on the production, we need to quickly react to these disturbances and revise the schedule in a cost-effective manner. the task of pma is to monitor any disruption occurred in the production. once a disruption is detected, a pre-processing procedure is triggered to pre-process the disruption so that the rescheduling procedure can further handle it. generally, with respect to the different disruptions, the pre-processing procedure will perform in different ways. for instance, when a rush order comes, it will inform the ppa to generate a feasible process plan for the input part based on the current available machines. when a machine breaks down, it will identify the affected jobs, which are to be processed on this breakdown machine, and then inform the ppa to generate new process plans for them. note that the stability of the schedule will be maintained and measured in terms of the sequence deviation when revising the disturbed schedule. 3.1.4 process planning agent (ppa) the ppa is used to obtain a high quality of process plan for a part, with its input process information from the pia. the objective of ppa is to select a set of opms for each vf and place them into an ordered sequence such that the sequence satisfies the precedence constraints and process plan achieves the predefined objective function. obviously, process planning problem is a combinatorial optimization problem. typically, the criteria for process plan evaluation include minimum number of setups, shortest processing time, minimum manufacturing cost, etc. from the economic point of view, the minimum manufacturing cost is taken as the objective function in this work, which can be considered from the following five cost aspects. (1) machine cost (mc) mc = n∑ i=1 mcii (1) where n is the total number of opts and mcii is the cost index for the machine used to perform opti, which is a constant for a particular machine. (2) tool cost (tc) tc = n∑ i=1 tcii (2) where tcii is the tool cost index for the tool used to perform opti, which is a constant for a particular tool. 42 y.f. wang: an agent-based distributed process planning and scheduling system (3) machine change cost (mcc): a machine change cost is required when two adjacent opts are performed on different machines. mcc = mcci n−1∑ i=1 (1− ω(mi+1,mi)) (3) ω(x, y) = { 1 if x = y 0 otherwise (4) where mcci is the machine change cost index and mi is the identity of the machine used to perform opti. (4) set-up change cost (scc): a setup change cost is required when two adjacent opts performed on the same machine have different tads. scc = scci n−1∑ i=1 ω(mi+1,mi)(1− ω(tadi+1, tadi)) (5) where scci is the setup change cost index. (5) tool change cost (tcc): a tool change cost is required when two adjacent opts performed on the same machine use different tools tcc = tcci n−1∑ i=1 ω(mi+1,mi)(1− ω(ti+1, ti)) (6) where tcci is the tool change cost index. the above five cost items can be taken either individually or collectively as a cost compound based on the actual requirement and data availability of the job shop. in this paper, all these five items are considered in the objective function, i.e., the total manufacturing cost (tmc): tcm = mc + tc +mcc + scc + tcc (7) two steps are followed to find an optimal process plan. the first step is to generate the process plan solution space of the input part, formed by all feasible opms that can be used for the fabrication of the part subject to the precedence relationship. the second step is to find the best solution based on a given optimization objective, i.e., minimum manufacture cost, as formulated in eq.(7). due to its vast solution space, it is difficult to find the optimal solution for this combination problem in a reasonable amount of time. two optimization methods based on genetic algorithm (ga) and simulated annealing algorithm (saa), respectively, have been developed to resolve this intractable problem and promising results were reported [12-13]. recently, we have developed an alternative algorithm advances in systems science and applications (2014) vol.14 no.1 43 based on particle swarm optimization [14]. particle swarm optimization (pso), being one of evolutionary computation approaches, was firstly proposed by kennedy and eberhart [15]. it is a class of population-based optimization algorithm that imitates the social swarm behaviours. members in the population interact with one another by learning from their own experience and gradually individuals move into better regions of the problem space. the attractive features of pso include inexpensive computation, individual improvement, and the ability of effective exploration and exploitation search. due to the characteristic of discrete process planning solution space and the continuous nature of the original pso, a novel solution representation scheme is introduced for the application of pso in solving the process planning problem. moreover, a local search algorithm is incorporated and interweaved with pso evolution to improve the best solution in each generation. the numerical experiments and analysis have demonstrated that the proposed algorithm (pso-ls) is capable of gaining a good quality of solution in an efficient way. one example optimal process plan in terms of xml format is shown below: 1 m1,t1,+x 2 m1,t1,+x 3.1.5 scheduling agent (sa) the sa is used to generate a high quality of schedule based on the involved process plans from the different ppas, each specified with the information of weight (1-10), due date, and batch size. a set of dispatching rules [16], including earliest due date (edd), shortest processing time (spt), are developed for the generation of schedule. since the scheduling problem is a typical non-polynomial (np) hard problem, the quality of solution obtained by dispatching rules may not be good. we therefore developed an approximation algorithm by incorporating 44 y.f. wang: an agent-based distributed process planning and scheduling system a tabu search in the pso for minimizing the tardiness in a flexible job shop scheduling. by this integration, pso provides a diversity of initial solutions for the tabu search while tabu search performs an exploitation search so as to affect the particles swarm search behavior. this hybrid procedure, taking the advantage of the difference between these two algorithms, proves to be an effective algorithm. the proposed algorithm (pso-ts) has been tested on various experiments and the results have demonstrated its robustness, efficiency and efficacy in all sets of experiments. 3.1.6 facilitator agent (fa) the objective of the fa is to improve the performance measure of the schedule through the coordination with the other agents. the fa evaluates the schedule obtained from sa, followed by invoking a heuristic rule, which is able to automatically issue a modification suggestion by identifying a particular operation of one job from the scheduled jobs and the resource to be modified. although the rule can choose several jobs for process plan modification, the strategy employed here is for fine-tuning, thus only one job is chosen in each round of iteration. the generated suggestion will be fed back to the corresponding ppa that the job belongs to and impose the constraints on its solution space. based on the update solution space, the affected ppa will then re-generate a process plan using an optimization algorithm. this generated optimal process plan, together with unchanged process plans, forms a new schedule in the sa. the coordination process continues until a satisfactory result is achieved. it is noted that the heuristic rules are in association with the optimization scenario and thus performs in different ways. currently, the system is capable of balancing the machine utilization and minimizing the number of tardy jobs and their tardiness, as well as accommodating the disruption of machine breakdown. the carried out simulation results have proven its effectiveness to achieve a satisfactory plans/schedule solution in a reasonable amount of time [3-5]. 3.2 agent structure in a multi-agent system, each agent is essentially an autonomous cognitive entity, which is able to communicate, reason, and react to the events from the external environment. to start, the agent receives a message and stores it in the receiving message queue. by decoding a message from the message list, an agent resolves the problem using its domain knowledge. once the problem is resolved, the agent will encode the solutions in a well-defined message and return it to the sender agent. the internal structure of an agent typically includes the following components: (1) network interface: it couples the agent to the network. (2) communication interface: it enables the agent to exchange and understand the messages. advances in systems science and applications (2014) vol.14 no.1 45 (3) interaction interface: it allows the agent to have conversation with other agents. (4) domain knowledge: it provides enterprise expertise knowledge required to perform the specific task. (5) problem solving module: owning the reasoning and decision making capability, it enables the agent to resolve the problem according to the received message and domain knowledge. 3.3 system architecture fig.3 presents the system architecture for the agent-based distributed process planning and scheduling system, which is based on the multi-tier application model, consisting of presentation tier, business tier, and eis (enterprise information system) tier. moreover, in order to ensure the effective coordination and communication among the involved agents, each agent sits on a platform based on jade (java agent development framework), which is a software framework to develop distributed agent-based applications (http://jade.tilab.com). in the presentation tier, a user can operate one or more agents to accomplish the required tasks in the local machine through the graphic user interfaces (gui). depending on the role of the user, one may have different privileges to access different kinds of agents. as such, these agents configured at geographically dispersed locations forms a peer-to-peer network, where the loosely coupled agents can communicate and cooperate with each other to fulfill the task. being located at the jade platform, each agent has a lifecycle. some agents are shut down when its behavior is completed while some are always active to repetitively execute the specific task. to manage these agents, a directory facilitator agent (dfa) is used to search and modify the description of registered agents. when an agent starts up (shuts down), it will be registered (deregistered) to the dfa. the business-tier located on the server side provides a set of functionality for business logic processing. it comprises two parts. the first part is to create an instance of runtime container for all active agents based on jade platform. the agent in the container is taken as an independent and autonomous process with a unique identity. in order to fulfil the task, it requires communicating with other agents. such a communication is realized through asynchronous message passing with an agent communication language (acl) in terms of a well-defined semantics. to implement it, some kernel packages are utilized, including jade core, jade acl, jade content, jade domain, jade mtp (message transport protocol), and jade protocol. however, the modeled agents cannot perform the specific task, since they are not endowed with capabilities except those of communication and interaction. we therefore develop different kinds of algorithms placed in the business tier, which are eligible for each agent to invoke in order to accomplish the assigned task. note that the feature extraction algorithm is written in c++ 46 y.f. wang: an agent-based distributed process planning and scheduling system fig.3 system architecture language. since the framework is developed based on the java virtual machines (jvm) and remote method invocation (rmi), agents cannot directly invoke this algorithm. we therefore utilize the java native interface (jni) technology, which allows a java application running in the jvm to operate with other applications or libraries written in other languages. the eis-tier plays a critical role in providing information infrastructure to the business process of an enterprise, which may include data integration, the existing application system integration, and legacy system integration. in this work, we only concern about the data integration, focusing on integrating the existing data with the developed system. in this way, the business process and data can be easily shared. the connection to the data stored in a relational database is realized by the jdbc technology. based on this architecture, the agents in the client tier communicate and exchange the data with the eis-tier through business tier while the business-tier functions manipulate the data from eis-tier and client-tier. by logically separating the presentation tier, the business tier, and the eis tier, the presented architecture brings the following benefits: (1) it only needs to define the business logic once within the business tier, which can be shared by any agent in the presentation tier. (2) it is possible to change the contents of any tier without having to make corresponding changes in the others. (3) it enables parallel development of the different tiers for a complex application. advances in systems science and applications (2014) vol.14 no.1 47 4 system implementation a prototype system has been developed to realize the integration of distributed process planning and scheduling based on the proposed architecture, using java language. mysql database is used to manage the production entities so as to ensure timely information sharing and maintain the data consistency among different functionalities. moreover, database connection pool is designed to improve the performance of the data retrieval and manipulation. in this way, users in geographically dispersed departments are able to cooperate with each other in a distributed, transactional, and portable environment. fig.4 illustrates the guis of the ppa, sa, and fa. fig.4 system interfaces of the process planning agent, scheduling agent, and facilitator agent. 5 conclusion in this paper, an interoperable intelligent multi-agent system is implemented to realize the integration of process planning and scheduling in a distributed manner. since both process planning and scheduling are known to be np-hard in nature, the combined solution space therefore makes it more difficult to find a 48 y.f. wang: an agent-based distributed process planning and scheduling system good solution. we therefore developed an intuitive approach based on the agent technology. through the iterative coordination among the modelled functional agents, the user delivery requirement will be finally satisfied while the costs of process plans are maintained as low as possible. to develop the prototype system, multi-tier system architecture and the jade platform, taking the advantage of its flexibility, scalability, reusability, and interoperability, were employed. although the prototype system based on the agent technology has gained certain success to find a satisfactory solution, more efforts still need to realize its applicability to the actual manufacturing environment. firstly, the functionalities in a manufacturing system usually interact with each other to give a better product design and achieve a higher production efficiency. likewise, the design process and shop floor control, which are respectively performed before the process planning and after the scheduling, have a significant impact on the quality of the obtained process plans and schedule. thus, to achieve a more robust solution, it would be better to incorporate the design agent and the shop floor agent. secondly, due to various unexpected events in the actual manufacturing, future work is expected to extend its capability to handle more disruptions. moreover, to meet different users requirement, more performance measures in the scheduling agent should be included. references [1] [1] jennings, n. r., and wooldridge, m. (1998), “applications of intelligent agents”, in: jennings, n.r., and wooldridge, m.j., agent technology: foundations, applications, and markets, springer, pp.3-28, . [2] wooldridge, m., and jennings, n.r. (1995), “intelligent agents: theory and practice”, the knowledge engineering review, vol.10, no.2, pp.115-152, [3] zhang, y.f., saravanan, a.n. and fuh, j.y.h. (2003), “integration of process planning and scheduling by exploring the flexibility of process planning”, int. j. prod. res., vol.41, no.3, pp.611-628, . [4] wang y.f., zhang y.f., fuh j.y.h., zhou z. d., xue l.g. and lou p. (2008), “a web-based integrated process planning and scheduling system”, ieee int. conf. on automation sci. and eng., aug. pp.23-26, . [5] wang y.f., zhang y.f., fuh j.y.h., zhou z.d., lou p., and xue l.g. (2008), “an integrated approach to reactive scheduling subject to the machine breakdown”, ieee int. conf. on automation and logistics, sep. pp.13, . [6] mcguire, j.g., kuokka, d.r., weber, j.c., tenenbaum, j.m., gruber, t.r., and olsen, g. r. (1993), “shade: technology for knowledge-based collabadvances in systems science and applications (2014) vol.14 no.1 49 orative engineering”, concurrent engineering: research and application, vol.1, no.3, pp.137-146, . [7] peng, y., finin, t., labrou, y., chu, b., long, j., tolone,w.j., and boughannam, a. (1998), “a multi-agent system for enterprise integration”, proc. of paam 98, london, uk, pp.155-169. [8] shen, w., maturana, f., norrie, d.h. (2000), “metamorph ii: an agentbased architecture for distributed intelligent design and manufacturing”, journal of intelligent manufacturing, vol.11, no.4, pp.237-251. [9] jia, h.z., ong, s.k., fuh, j.y.h., zhang, y.f., and nee, a.y.c. (2004), “an adaptive and upgradable agent-based system for coordinated product development and manufacture”, robotic and computer-integrated manufacturing, vol.20, no.2, pp.79-90. [10] wong, t.n., leung, c.w., mak, k.l. and fung, r.y.k. (2006), “integrated process planning and scheduling/rescheduling-an agent-based approach”, int. j. prod. res., vol.44, no.18-19, pp.3627-3655. [11] ahmadi, h. (2008), “automated volumetric feature extraction from the machining perspective”, master of engineering thesis, national university of singapore. [12] zhang, f., zhang, y.f., and nee, a.y.c. (1997), “using genetic algorithms in process planning for job shop machining”, ieee transactions on evolutionary computation, vol.1, no.4, pp.278-289. [13] ma, g.h., zhang, y.f., and nee, a.y.c. (2000), “a simulated annealingbased optimization algorithm for process planning”, international journal of production research, vol.38, no.1), pp.2671-2678. [14] wang y.f., zhang y.f., and fuh j.y.h. (2009), “using hybrid particle swarm optimization for process planning problem”, ieee int. conf. on computational science and optimization, april pp.24-26. [15] kennedy, j. and eberhart, r. (1995), “particle swarm optimization”, proceedings of ieee int. conf. on neural networks, piscataway, nj, ieee press, pp.1942-1948. [16] baker, k.r. (1974), “introduction to sequencing and scheduling”, new york: wiley publications. corresponding author author can be contacted at: mpefuhyh@nus.edu.sg advances in systems science and applications (2013) vol.13 no.2 182-197 finding unbiased simultaneous prediction limits for order statistics of future samples with applications nicholas a. nechval1, konstantin n. nechval2, gundars berzins1, juris krasts1, maris purgailis1, uldis rozevskis1 and vladimir f. strelchonok3 1 statistics department, evf research institute, university of latvia, raina blvd 19, lv-1050 riga, latvia 2 applied mathematics department, transport and telecommunication institute, lomonosov street 1, lv-1019 riga, latvia 3 informatics department, baltic international academy, lomonosov street 4, lv-1019 riga, latvia abstract this paper provides procedures for finding unbiased simultaneous prediction limits on the observations or functions of observations of all of k future samples using the results of a previous sample from the same underlying distribution belonging to invariant family. the results have direct application in reliability theory, where the time until the first failure in a group of several items in service provides a measure of assurance regarding the operation of the items. the simultaneous prediction limits are required as specifications on future life for components, as warranty limits for the future performance of a specified number of systems with standby units, and in various other applications. prediction limit is an important statistical tool in the area of quality control. the lower simultaneous prediction limits are often used as warranty criteria by manufacturers. the initial sample and k future samples are available, and the manufacturer wants to have a high assurance that all of the k future orders will be accepted. it is assumed throughout that k + 1 samples are obtained by taking random samples from the same population. in other words, the manufacturing process remains constant. the results in this paper are generalizations of the usual prediction limits on observations or functions of observations of only one future sample. in the paper, attention is restricted to invariant families of distributions. the technique used here emphasizes pivotal quantities relevant for obtaining ancillary statistics and is applicable whenever the statistical problem is invariant under a group of transformations that acts transitively on the parameter space. applications of the proposed procedures are given for the two-parameter exponential and weibull distributions. the exact prediction limits are found and illustrated with a numerical example. keywords future samples, order statistics, simultaneous prediction limits 1 introduction statistical intervals used by engineers and others include confidence intervals on a population parameter, such as the mean, and tolerance intervals. confidence intervals give information about parameter of the population or a function of 183 advances in systems science and applications (2013) vol.13 no.2 population parameters such as a percentile; tolerance intervals give information about a region which contains a specified proportion of a population. often one desires to construct from the results of a previous sample an interval which will have a high probability of containing the values of all of k future observations. for example, such an interval would be required in establishing limits on the values of some performance variable for a small shipment of equipment when the satisfactory performance of all units is to be guaranteed, or in setting acceptance limits on a specific lot of material, when acceptance requires the values of all items in a future sample to fall within the limits. an interval which contains the values of a specified number of future observations with a specified probability is known as a prediction interval. such an interval need be distinguished both from a confidence interval on an unknown distribution parameter, and from a tolerance interval to contain the values of a specified proportion of the population. research works on prediction intervals related to a single future statistic are abundant (see hahn and meeker [1], patel [2], and references therein). in many situations of interest, it is desirable to construct lower simultaneous prediction limits that are exceeded with probability γ by observations or functions of observations of all of k future samples, each consisting of m units. the prediction limits depend upon a previously available complete or type ii censored sample from the same distribution. for instance, two situations where such limits are required are: 1. a customer has placed an order for a product which has an underlying timeto-failure distribution. the terms of his purchase call for k monthly shipments. from each shipment the customer will select a random sample of m units and accept the shipment only if the smallest time to failure for this sample exceeds a specified lower limit. the manufacturer wishes to use the results of a previous sample of n units to calculate this limit so that the probability is γ that all k shipments will be accepted. it is assumed that the n past units and the km future units are random samples from the same population. 2. a system consists of n identical components whose times to failure follow an underlying distribution. initially one component is operating and the remaining n-1 components are in a standby mode; a new component goes into operation as soon as the preceding component has failed. the system is said to fail when all n components have failed. thus, the system time to failure is the total of the failure times for the n components. a simultaneous lower prediction limit to be exceeded with probability γ by the system time to failure of all of k future systems is desired. this limit is to be calculated from the times to failure of n previously tested components. similar problems also arise in various product maintenance and servicing problems. nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 184 prediction limits can be of several forms. hahn [3] dealt with simultaneous prediction limits on the standard deviations of all of the k future samples from a normal population. hahn [4] considered the problem of obtaining simultaneous prediction limits on the means of all of k future samples from an exponential distribution. in addition, hahn and nelson [5] discussed such limits and their applications. mann, schafer, and singpurwalla [6] gave an interval that contains, with probability γ, all m observations of a single future sample from the same population. fertig and mann [7] constructed prediction intervals to contain at least m − k + 1 out of m future observations from a normal distribution with probability 1− β. they considered life-test data, and the performance variate of interest is the failure time of an item. their lower prediction limit constitutes a “warranty period”. in this paper we give an expression for obtaining unbiased simultaneous prediction limits on order statistics of all of k future samples. in order to obtain the unbiased simultaneous prediction limits, attention is restricted to invariant families of distributions. in particular, the case is considered where a previously available complete or type ii censored sample is from a continuous distribution with cumulative distribution function (cdf) f ((x−µ)/σ) and probability density function (pdf) 1/σf((x − µ)/σ), where f (.) is known but both the location (µ) and scale (σ) parameters are unknown. for such family of distributions the decision problem remains invariant under a group of transformations (a subgroup of the full affine group) which takes µ (the location parameter) and σ (the scale) into cµ+ b and cσ, respectively, where b lies in the range of µ, c > 0. this group acts transitively on the parameter space and, consequently, the risk of any equivariant estimator is a constant. among the class of such estimators there is therefore a “best” one. the effect of imposing the principle of invariance, in this case, is to reduce the class of all possible estimators to one. in the present paper we investigate this question for the problem of constructing the unbiased simultaneous prediction limits on order statistics in future samples. the technique used here emphasizes pivotal quantities relevant for obtaining ancillary statistics. it is a special case of the method of invariant embedding of sample statistics into a performance index [8-11] applicable whenever the statistical problem is invariant under a group of transformations which acts transitively on the parameter space (i.e., in problems where there is a unique best invariant procedure). the exact unbiased simultaneous prediction limits on order statistics of all of k future samples are obtained via the technique of invariant embedding and illustrated with numerical example. 185 advances in systems science and applications (2013) vol.13 no.2 2 mathematical preliminaries the main theorem, which shows how to construct lower (upper) simultaneous prediction limit for the order statistics in all of k future samples when prediction limit for a single future sample is available, is given below. theorem 1 (lower (upper) simultaneous prediction limit under complete information). let (y1j , ..., ymj) be the jth random sample of mj “future” observations from the cdf fθ(.), where θ is the parameter (in general, vector), j ∈ 1, ..., k, and let y(rj ,mj) denote the rjth order statistic in the jth sample of size mj . assume that all of k samples from the same cdf are independent. then a lower simultaneous (1− α) prediction limit h on the rjth order statistics y(rj ,mj), j = 1, , k, of all of k future samples may be obtained from pθ{y(r1,m1) ≥ h, ..., y(rj ,mj) ≥ h, ..., y(rk,mk) ≥ h} = r1−1∑ i1=0 ... rj−1∑ ij=0 ... rk−1∑ ik=0 ( m1 i1 ) ... ( mj ij ) ... ( mk ik ) × pθ{y(iς+1,mς) ≥ h} − pθ{y(iς,mς) ≥ h}( mς iς ) = 1− α (1) where iς = k∑ j=1 ij ,mς = k∑ j=1 mj (2) (observe that an upper simultaneous α prediction limit h may be obtained from a lower simultaneous prediction limit by replacing 1− α by α.) proof. we have: pθ{y(r1,m1) ≥ h, ..., y(rj ,mj) ≥ h, ..., y(rk,mk) ≥ h} = k∏ j=1 pθ{y(rj ,mj) ≥ h} = k∏ j=1 rj−1∑ ij=0 ( mj ij ) [fθ(h)] ij [1− fθ(h)] mj−ij = r1−1∑ i1=0 ... rj−1∑ ij=0 ... rk−1∑ ik=0 ( m1 i1 ) ... ( mj ij ) ... ( mk ik ) [fθ(h)] iς [1− fθ(h)] mς−iς (3) since nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 186 [fθ(h)] iς [1− fθ(h)] mς−iς = ( mς iς )−1 [ iς∑ i=0 ( mς i ) [fθ(h)] i[1− fθ(h)] m∑−i− iς−1∑ i=0 ( mς i ) [fθ(h)] i[1− fθ(h)] mς−i ] = pθ{y(iς+1,mς) ≥ h} − pθ{y(iς,mς) ≥ h}( mς iς ) (4) the joint probability can be written as pθ{y(r1,m1) ≥ h, ..., y(rj ,mj) ≥ h, ..., y(rk,mk) ≥ h} = r1−1∑ i1=0 ... rj−1∑ ij=0 ... rk−1∑ ik=0 ( m1 i1 ) ... ( mj ij ) ... ( mk ik ) × pθ{y(iς+1,mς) ≥ h} − pθ{y(iς,mς) ≥ h}( mς iς ) (5) this ends the proof. corollary 1.1 if rj = 1, ∀j = 1(1)k, then pθ{y(1,m1) ≥ h, ..., y(1,mj) ≥ h, ..., y(1,mk) ≥ h} = pθ{y(1,mς) ≥ h} = 1− α (6) theorem 2 (lower (upper) unbiased simultaneous prediction limit under parametric uncertainty). let (x1 ≤ ... ≤ xr) be the r smallest observations in a random sample of size n from the cdf fθ(.), where the θ is the parameter (in general, vector), and let (y1j , ..., ymj ) be the jth random sample of mj “future” observations from the same cdf, j ∈ {1, ..., k}. assume that (k + 1) samples are independent and the parameter θ is unknown. let h = h(x1, ..., xr) be any statistic based on the preliminary sample and let y(rj ,mj) denote the rjth order statistic in the jth sample of size mj . then an unbiased lower simultaneous (1− α) prediction limit h on the rjth order statistics y(rj ,mj), j = 1, , k, of all of k future samples may be obtained from 187 advances in systems science and applications (2013) vol.13 no.2 eθ { pθ{y(r1,m1) ≥ h, ..., y(rj ,mj) ≥ h, ..., y(rk,mk) ≥ h} } = r1−1∑ i1=0 ... rj−1∑ ij=0 ... rk−1∑ ik=0 ( m1 i1 ) ... ( mj ij ) ... ( mk ik ) · eθ { pθ{y(iς+1,mς) ≥ h} } − eθ { pθ{y(iς,mς) ≥ h} }( mς iς ) (7) proof. for the proof we refer to theorem 1. corollary 2.1. if rj = 1, ∀j = 1(1)k, then eθ { pθ{y(1,m1) ≥ h, ..., y(1,mj) ≥ h, ..., y(1,mk) ≥ h} } = eθ { pθ{y(1,mς) ≥ h} } = 1− α (8) remark. in this paper, in order to find the unbiased lower simultaneous (1−α) prediction limith on the rjth order statistics y(rj ,mj), j = 1, ..., k, of all of k future samples, the technique of invariant embedding [8-11] is used. 2.1 weibull distribution in this paper, the two-parameter weibull distribution with the pdf fθ(x) = δ β ( x β )δ−1exp [ −( x β )δ ] , x > 0, β > 0, δ > 0 (9) indexed by scale and shape parameters β and δ is used as the underlying distribution of a random variable x in a sample of the lifetime data, where θ = (β, δ). we consider both parameters β, δ to be unknown. let (x1, ..., xn) be a random sample from the two-parameter weibull distribution (9), and let β̂, δ̂ be maximum likelihood estimates of β, δ computed on the basis of (x1, ..., xn). in terms of the weibull variates, we have that v1 = ( β̂ β )δ, v2 = δ δ̂ , v3 = ( β̂ β )δ̂ (10) are pivotal quantities. further more, let zi = (xi/β̂) δ̂, i = 1, ..., n (11) it is readily verified that any n-2 of the zi’s, say zi, ..., zn−2 form a set of n-2 functionally independent ancillary statistics. the appropriate conditional approach, first suggested by fisher [12], is to consider the distributions of v1, v2, v3 nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 188 conditional on the observed value of z(n) = (zi, ..., zn). (for purposes of symmetry of notation we include all of (zi, ..., zn) in expressions stated here; it can be shown that zn, zn−1, can be determined as functions of zi, ..., zn−2 only.) theorem 3. (joint pdf of the pivotal quantities v1, v2 from the two-parameter weibull distribution) let (x1 ≤ ... ≤ xr) be the first r ordered observations from a sample of size n from the two-parameter weibull distribution (9). then the joint pdf of the pivotal quantities v1 = ( β̂ β )δ, v2 = δ δ̂ (12) conditional on fixed z(r) = (zi, ..., zr) (13) where zi = ( xi β̂ )δ̂, i = 1, ..., r (14) are ancillary statistics, any r-2 of which form a functionally independent set,β̂ and δ̂ are the maximum likelihood estimates for β and δ based on the first r ordered observations (x1 ≤ ... ≤ xr) from a sample of size n from the two-parameter weibull distribution (9), which can be found from solution of β̂ = ([ r∑ i=1 xδ̂i + (n− r)xδ̂r ] /r )1/δ̂ (15) and δ̂ = ( r∑ i=1 xδ̂ i lnxi + (n− r)xδ̂ rlnxr )( r∑ i=1 xδ̂ i + (n− r)xδ̂ r )−1 − 1 r r∑ i=1 lnxi −1 (16) is given by f(v1, v2|z(r)) = ϑ•(z(r))vr−2 2 r∏ i=1 zv2i vr−1 1 exp ( −v1 [ r∑ i=1 zv2i + (n− r)zv2r ]) = f(v2|z(r))f(v1|v2, z(r)), v1 ∈ (0,∞), v2 ∈ (0,∞) (17) where ϑ•(z(r)) = [∫ ∞ 0 γ(r)vr−2 2 r∏ i=1 zv2i ( r∑ i=1 zv2i + (n− r)zv2r )−r dv2 ]−1 (18) 189 advances in systems science and applications (2013) vol.13 no.2 is the normalizing constant, f(v2|z(r)) = ϑ(z(r))vr−2 2 r∏ i=1 zv2i ( r∑ i=1 zv2i + (n− r)zv2r )−r , v2 ∈ (0,∞) (19) ϑ(z(r)) = [∫ ∞ 0 vr−2 2 r∏ i=1 zv2i ( r∑ i=1 zv2i + (n− r)zv2r )−r dv2 ]−1 (20) f(v1|v2, z(r)) = [ ∑r i=1 z v2 i + (n− r)zv2r ]r γ(r) vr−1 1 exp ( −v1 [ r∑ i=1 zv2i + (n− r)zv2r ]) = 1 γ(r) ( v1 [ r∑ i=1 zv2i + (n− r)zv2r ])r−1 exp ( −v1 [ r∑ i=1 zv2i + (n− r)zv2r ]) ×[ r∑ i=1 zv2i + (n− r)zv2r ] , v1 ∈ (0,∞) (21) proof. the joint density of x1 ≤ ... ≤ xr is given by fθ(x1, ..., xr) = n! (n− r)! r∏ i=1 δ β ( xi β )δ−1exp(−( xi β )δ)exp(−(n− r)( xr β )δ) (22) using the invariant embedding technique [8-11], we transform (22) to fθ(x1, ..., xr)dβ̂dδ̂ = n! (n− r)! r∏ i=1 x−1 i δr r∏ i=1 ( xi β )δexp ( − r∑ i=1 ( xi β )δ − (n− r)( xr β )δ ) dβ̂dδ̂ = n! (n− r)! β̂δ̂r r∏ i=1 x−1 i ( δ δ̂ )r−2 r∏ i=1 ( xi β̂ )δ̂( δ δ̂ )( β̂ β )δ(r−1)× exp ( −( β̂ β )δ [ r∑ i=1 ( xi β̂ )δ̂( δ δ̂ ) + (n− r)( xr β̂ )δ̂( δ δ̂ ) ])( δ β ( β̂ β )δ−1dβ̂ ) (− δ δ̂2 dδ̂) = n! (n− r)! β̂δ̂r r∏ i=1 x−1 i vr−2 2 r∏ i=1 zv2i vr−1 1 exp ( −v1 [ r∑ i=1 zv2i + (n− r)zv2 r ]) dv1dv2 (23) normalizing (23), we obtain (17). this ends the proof. theorem 4. (lower (upper) unbiased prediction limit h for the lth order statistic yl in a new (future) sample of m observations from the two-parameter weibull distribution on the basis of the preliminary data sample) let x1 ≤ ... ≤ xr be the first r ordered observations from the preliminary sample of size n from the nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 190 two-parameter weibull distribution (9). then a lower unbiased (1−α) prediction limit h on the lth order statistic yl from a set of m future ordered observations y1 ≤ ... ≤ ym also from the distribution (9) is given by h = arg[eθ{pθ{y1 ≥ h}|z(r)} = 1− α] = z 1/δ̂ h β̂ (24) where eθ{pθ{yl ≥ h}|z(r)} = ∞∫ 0 vr−2 2 r∏ i=1 zv2 i l−1∑ k=0 ( m k ) k∑ j=0 ( k j ) (−1) j ( (m− k + j)zv2h + r∑ i=1 zv2 i + (n− r)zv2r )−r dv2 ∞∫ 0 vr−2 2 r∏ i=1 zv2i ( r∑ i=1 zv2i + (n− r)zv2 r )−r dv2 (25) zh = ( h β̂ )δ̂ (26) zi = (xi/β̂) δ̂, i = 1, ..., r; β̂ and δ̂ are the maximum likelihood estimates for β and β based on the first r ordered observations (x1 ≤ ... ≤ xr) from a sample of size n from the two-parameter weibull distribution (9). (observe that an upper unbiased α prediction limit h on the lth order statistic yl from a set of m future ordered observations y1 ≤ ... ≤ ym may be obtained from a lower unbiased (1− α) prediction limit by replacing 1− α by α.) proof. if there is a random sample of m ordered observations y1 ≤ ... ≤ ym from the two-parameter weibull distribution (9) with the pdf fθ(y) and cdf fθ(y), then for the lth order statistic yl we have pθ{yl ≥ h} = l−1∑ k=0 ( m k ) [fθ(h)]k[1− fθ(h)]m−k = l−1∑ k=0 ( m k )[ 1− exp ( − ( h β )δ )]k[ exp ( − ( h β )δ )]m−k (27) writing (27) as 191 advances in systems science and applications (2013) vol.13 no.2 pθ{yl ≥ h} = l−1∑ k=0 ( m k )[ 1− exp ( − ( h β )δ )]k exp ( −(m− k) ( h β )δ ) = l−1∑ k=0 ( m k )1− exp − ( h ⌢ β )⌢ δ ( δ ⌢ δ )(⌢ β β )δ   k exp −(m− k) ( h ⌢ β )⌢ δ ( δ ⌢ δ )(⌢ β β )δ  = l−1∑ k=0 ( m k ) [1− exp(−zv2 h v1)] k exp(−(m− k)zv2 h v1) = l−1∑ k=0 ( m k ) k∑ j=0 ( k j ) (−1)j exp[−v1(m− k + j)zv2h ] = p{zl > zh |v1, v2} (28) where zl = ( yl ⌢ β )⌢ δ (29) we have from (17) and (28) that eθ{pθ{yl ≥ h}|z(r)} = e{p{zl ≥ zh |v1, v2}|z(r)} = ∞∫ 0 ∞∫ 0 p{zl ≥ zh |v1, v2}f(v1, v2|z(r))dv1dv2 (30) now v1 can be integrated out of (30) in a straightforward way to give (25). this completes the proof. corollary 4.1. if l = 1, then h = arg  ∞∫ 0 vr−2 2 r∏ i=1 zv2 i m (h ⌢ β )⌢ δ v2 + r∑ i=1 zv2 i + (n− r)zv2r −r dv2 ∞∫ 0 vr−2 2 r∏ i=1 zv2 i ( r∑ i=1 zv2 i + (n− r)zv2 r )−r dv2 = 1− α  (31) theorem 5 ((lower (upper) unbiased prediction limit h for the lth order statistic yl in a new (future) sample of m observations from the left-truncated weibull distribution on the basis of the preliminary data sample) let x1 ≤ ... ≤ xr be the first r ordered observations from the preliminary sample of size n from the left-truncated weibull distribution with the pdf fθ(x) = δ σx δ−1 exp[−(xδ − µ)/σ], (xδ ≥ µ, σ, δ > 0) (32) nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 192 where θ = (µ, σ, δ), δ is termed the shape parameter,δ is the scale parameter, and µ is the truncation parameter. it is assumed that the parameter δ is known. then a lower unbiased (1−α) prediction limit h on the lth order statistic yl from a set of m future ordered observations y1 ≤ ... ≤ ym also from the distribution (32) is given by h = ( xδ 1 + whs )1/δ (33) where wh =  arg ( 1− m!(m+n−l)! (m−l)!(m+n)! (1− nwh) −(r−1) = 1− α ) , if α < m!(n+m−l)! (m−l)!(n+m)! , arg nl ( m l ) l−1∑ i=0  l − 1 i (−1)i[1+wh(m−l+i+1)]−(r−1) (n+m−l+i+1)(m−l+i+1) = 1− α  , if α ≥ m!(n+m−l)! (m−l)!(n+m)! , (34) s = r∑ i=1 (xδ i −xδ 1) + (n− r)(xδ r −xδ 1) (35) (observe that an upper unbiased α prediction limit h on the lth order statistic yl may be obtained from a lower unbiased (1 − α) prediction limit by replacing 1− α by α.) proof. it can be justified by using the factorization theorem that (xδ 1 , s) is a sufficient statistic for (µ, δ). we wish, on the basis of the sufficient statistic (xδ 1 , s) for (µ, δ), to construct the predictive density function of the lth order statistic yl from a set of m future ordered observations y1 ≤ ... ≤ ym. by using the technique of invariant embedding [8-11] of (xδ 1 , s), if x1 ≤ yl, or (y δ l , s), if x1 ≥ yl, into a pivotal quantity (y δ l − µ)/σ or (xδ 1 − µ)/σ, respectively, we obtain an ancillary statistic wl = ( y δ l −xδ 1 )/ s (36) it can be shown that the pdf of wl is given by f(wl) =  n(r − 1)l ( m l ) l−1∑ i=0 ( l − 1 i ) (−1) i [1 + wl(m− l + i+ 1)]−r n+m− l + i+ 1 , if wl ≥ 0, n(r − 1) m!(n+m−l)! (m−l)!(n+m)! (1− nwl) −r , if wl < 0. (37) 193 advances in systems science and applications (2013) vol.13 no.2 it follows from (37) that p (wl > wh) =  nl ( m l ) l−1∑ i=0 ( l − 1 i ) (−1) i [1 + wh(m− l + i+ 1)] −(r−1) (n+m− l + i+ 1)(m− l + i+ 1) , if wh ≥ 0, 1− m!(m+ n− l)! (m− l)!(m+ n)! (1− nwh) −(r−1) , if wh < 0. (38) where wh = ( hδ −xδ 1 )/ s (39) this ends the proof. corollary 5.1. if l = 1, then a lower (1−α) prediction limit h on the minimum y1 of a set of m future ordered observations y1 ≤ ... ≤ ym is given by h =  ( xδ 1 + s m [( n (1−α)(n+m) ) 1 r−1 − 1 ])1/δ , if α ≥ m n+m , ( xδ 1 − s n [( m α(n+m) ) 1 r−1 − 1 ])1/δ , if α < m n+m . (40) 2.2 two-parameter exponential distribution theorem 6 ((lower (upper) unbiased prediction limit h for the lth order statistic yl in a new (future) sample of m observations from the two-parameter exponential distribution on the basis of the preliminary data sample) let x1 ≤ ... ≤ xr be the first r ordered observations from the preliminary sample of size n from the two-parameter exponential distribution with the pdf fθ(x) = 1 σ exp[−(x− µ)/σ],(xδ ≥ µ, σ > 0) (41) where θ = (µ, σ), σ is the scale parameter, and µ is the shift parameter. it is assumed that these parameters are unknown. then a lower unbiased (1 − α) prediction limit h on the lth order statistic yl from a set of m future ordered observations y1 ≤ ... ≤ ym also from the distribution (41) is given by h = x1 + whs (42) where wh =  arg ( 1− m!(m+n−l)! (m−l)!(m+n)! (1− nwh) −(r−1) = 1− α ) , if α < m!(n+m−l)! (m−l)!(n+m)! , arg nl ( m l ) l−1∑ i=0  l − 1 i (−1)i[1+wh(m−l+i+1)]−(r−1) (n+m−l+i+1)(m−l+i+1) = 1− α  , if α ≥ m!(n+m−l)! (m−l)!(n+m)! (43) nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 194 s = r∑ i=1 (xi −x1) + (n− r)(xr −x1) (44) (observe that an upper unbiased α prediction limit h on the lth order statistic yl may be obtained from a lower unbiased (1 − α) prediction limit by replacing 1− α by α.) proof. for the proof we refer to theorem 5. corollary 6.1. if l = 1, then a lower (1−α) prediction limit h on the minimum y1 of a set of m future ordered observations y1 ≤ ... ≤ ym is given by h =  ( x1 + s m [( n (1−α)(n+m) ) 1 r−1 − 1 ]) , if α ≥ m n+m , ( x1 − s n [( m α(n+m) ) 1 r−1 − 1 ]) , if α < m n+m . (45) remark 2. let us assume that the parent distributions are the two-parameter exponential fθ(x) = 1− exp ( −x− θ2 θ1 ) , x ≥ θ2, θ1 > 0 (46) where θ = (θ1, θ2) and the pareto distribution fθ(x) = 1− (θ2/x) 1/θ1 , x ≥ θ2 > 0, θ1 > 0 (47) let x be a random variable with the pareto distribution (47), and define y = lnx. then y becomes a random variable with the exponential distribution (46), where θ2 is replaced by lnθ2. therefore it is enough to consider only the exponential distribution, because the results for the pareto distribution are easily obtained from those for the exponential distribution. 3 numerical example an industrial firm has the policy to replace a certain device, used at several locations in its plant, at the end of 24-month intervals. it doesn’t want too many of these items to fail before being replaced. shipments of a lot of devices are made to each of three firms. each firm selects a random sample of 5 items and accepts his shipment if no failures occur before a specified lifetime has accumulated. the manufacturer wishes to take a random sample and to calculate the lower prediction limit so that all shipments will be accepted with a probability of 0.95. the resulting lifetimes (rounded off to the nearest month) of an initial sample of size 15 from a population of such devices are given in table 1. goodness-of-fit 195 advances in systems science and applications (2013) vol.13 no.2 table 1 the resulting lifetimes statistical results x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 x13 x14 x15 8 9 10 12 14 17 20 25 29 30 35 40 47 54 62 lifetime (in number of month intervals) testing. it is assumed that xi ∼ fθ(x) = δ σx δ−1 exp[−(xδ − µ)/σ], (x ≥ µ, σ, δ > 0), i = 1(1)15 (48) where the parameters µ and δ are unknown; (δ = 0.87). thus, for this example, r = n = 15, k = 3,m = 5, 1−α = 0.95, xδ 1 = 6.1,and s = 170.8. it can be shown that the uj = 1−  j+1∑ i=2 (n− i+ 1)(xδ i −xδ i−1) j+2∑ i=2 (n− i+ 1)(xδ i −xδ i−1)  j , j = 1(1)n− 2 (49) are i.i.d. u(0, 1) rv’s (nechval et al. [13]). we assess the statistical significance of departures from the left-truncated weibull model by performing the kolmogorovsmirnov goodness-of-fit test. we use the k statistic (muller et al. [14]). the rejection region for the α level of significance is k ≥ kn;α. the percentage points for kn;α were given by muller et al. [14]. for this example, k = 0.220 < kn=13,α=0.05 = 0.361 (50) thus, there is not evidence to rule out the left-truncated weibull model. it follows from (8) and (40), for α = 0.05 < km n+ km = 0.5 (51) that h = ( xδ 1 − s n [( km α(n+km) ) 1 n−1 − 1 ]) 1 δ = ( 6.1− 170.8 15 [( 15 0.05(15+15) ) 1 14 − 1 ]) 1 0.87 = 5 (52) thus, the manufacturer has 95% assurance that no failures will occur in each shipment before h = 5 month intervals. nicholas a. nechval : finding unbiased simultaneous prediction limits for order ... 196 4 conclusion and future work in this paper we propose the technique of constructing unbiased simultaneous prediction limits on observations or functions of observations in all of k future samples under parametric uncertainty of the underlying distribution. these unbiased simultaneous prediction limits are based on a previously available complete or type ii censored sample from the same distribution. we present an equation for this type of unbiased simultaneous prediction limits which holds for any distribution and any statistic from the previous sample when a prediction limit for a single future sample is available. the exact prediction limits are found and illustrated with a numerical example. the methodology described here can be extended in several different directions to handle various problems that arise in practice. we have illustrated the proposed methodology for the two-parameter exponential and weibull distributions. application to other distributions could follow directly. acknowledgements this research was supported in part by grant no.09.1544 from the latvian council of science and the national institute of mathematics and informatics of latvia. references [1] hahn g.j and meeker w.q. (1991), statistical intervals – a guide for practitioners, john wiley, new york. [2] patel j.k. (1989), “prediction intervals – a review”, communications in statistics theory and methods, vol.18, pp.2393-2465. [3] hahn g.j. (1972), “simultaneous prediction intervals to contain the standard deviations or ranges of future samples from a normal distribution”, journal of the american statistical association, vol.67, pp.938-942. [4] hahn g.j. (1975), “a prediction interval on the means of future samples from an exponential distribution”, technometrics, vol.17, pp.341-345. [5] hahn g.j, nelson w. (1973), “a survey of prediction intervals and their applications”, journal of quality technology, vol.5, pp.178-188. [6] mann n.r, schafer r.e, singpurwalla j.d. (1974), methods for statistical analysis of reliability and life data, john wiley, new york [7] fertig k.w, mann n.r. (1977), “one-sided prediction intervals for at least p out of m future observations from a normal population”, technometrics, vol.19, pp.167-177. 197 advances in systems science and applications (2013) vol.13 no.2 [8] nechval n.a, berzins g, purgailis m, nechval k.n. (2008), “improved estimation of state of stochastic systems via invariant embedding technique”, wseas transactions on mathematics, vol.7, pp.141-159. [9] nechval n.a, purgailis m, berzins g, cikste k, krasts j, nechval k.n. (2010), “invariant embedding technique and its applications for improvement or optimization of statistical decisions”, in: al-begain. k., fiems. d., knottenbelt. w. (eds.), analytical and stochastic modeling techniques and applications, lncs, springer, heidelberg, vol.6148, pp.306-320. [10] nechval n.a, purgailis m, cikste k, berzins g, nechval k.n. (2010), “optimization of statistical decisions via an invariant embedding technique”, in: lecture notes in engineering and computer science: proceedings of the world congress on engineering 2010, wce 2010, london, 30 june 2 july 2010, pp.1776-1782. [11] nechval n.a, purgailis m, nechval k.n, strelchonok v.f. (2012), “optimal predictive inferences for future order statistics via a specific loss function”, iaeng international journal of applied mathematics, vol.42, pp.40-51. [12] fisher r.a. (1934), “two new properties of mathematical likelihood”, proceedings of the royal society a, vol.144, pp.285-307. [13] nechval n.a, nechval k.n. (1998), “characterization theorems for selecting the type of underlying distribution”, in: proceedings of the 7th vilnius conference on probability theory and 22nd european meeting of statisticians, tev, vilnius, pp.352-353. [14] muller p.h, neumann p, storm r. (1979), tables of mathematical statistics, veb fachbuchverlag, leipzig. corresponding author nicholas a. nechval can be contacted at: nechval@junik.lv мягкая вероятность ошибки при передаче информации adv syst sci appl 2018; 04; 136-150 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/584 systems methodology and model tools for territorial sustainable management t.yu. anopchenko1, o.i. gorbaneva1, e.i. lazareva1, a.d. murzin1, g.a. ougolnitsky*1 1)southern federal university, rostov-on-don, russia received may 7, 2018; revised november 14, 2018; published december 31, 2018 abstract: no doubt, this research is topical in the context of new challenges arising for different russian territories due to intensive globalization processes and jumping trends of world economy. these challenges are stressing the need for an innovation-oriented improvement of territorial management systems in accordance with the realities of cyclicity and the conditions of fundamental resources to stimulate positive social and economic dynamics. the goal of the present paper is to justify a methodology and real ways of territorial sustainable development management system design based on efficient model tools and information technologies. for achieving this goal, several problems have been solved such as the dynamic modeling of a territorial social-ecological-economic system and coordination of public (social) and private interests, the elaboration of reasonable management approaches to territorial development risks, as well as the design of an appropriate methodology to integrate the suggested models and approaches within an information-analytical sustainable management support system. the novelty of an original scientific outlook introduced by the authors consists in the elaboration of a well-grounded systems methodology and model tools for territorial sustainable management and also in the justification and verification of qualitative and quantitative approaches to the analysis of development risks and their integration into a unified territorial sustainable management system.the methodology and model tools presented below can be used for enhancing the scientific and practical components of a territorial sustainable management system under dynamic and conflicting external conditions. keywords: homeostasis, simulation, econometric methods, territorial sustainable development, system compatibility, welfare capital reproduction, risk management, sustainable development management. 1. introduction this paper extends the results of the earlier publication [22], in which a systems approach to regional sustainable management was described. the methodology suggested below integrates territorial social-ecological-economic system simulation, econometric assessment of its innovative sustainable development potential and coordination of social and private interests into a unified territorial sustainable management system. a conceptual platform of the original modeling approach is the evolutionary-cyclic paradigm of subject-object relations in the “welfare capital reproduction –– territorial innovative sustainable development” system. in accordance with this paradigm, the network reality makes it necessary to elaborate a collectively compatible management strategy relying on the coordination and balancing of personalized interests for different subjects of innovative sustainable development, mutual benefit, trust, and public-private partnership (fig. 1.1). as shown in [12], an alternative way to implement the coordination principle is to identify an ideal hierarchical chain of interests of economic subjects, associating it with an adequate incentive policy. * corresponding author: ougoln@mail.ru model tools for sustainable management 137 copyright ©2018 assa adv. in systems science and appl. (2018) the models that produce analysis tools for analyzing the object in strategic management systems – the trajectories of territorial sustainable development must have the following major features [14]:  nonlinearity of the model and orientation for a long time interval;  taking into account the impact of economic activity on natural processes and the state of the social sphere;  •inclusion of feedback flows between the ecological, social and economic subsystems of the territorial system;  not only material or monetary valued services of social and natural systems should be included, but also the rest, "intangible", such as favorable conditions for production and residence;  taking into account the concern for future generations, which can be expressed in restrictions on natural capital, social and economic capital to ensure an even distribution of social, natural and economic potential between successive generations;  the possibility of describing qualitative structural changes;  reflection of the limited availability of resources, including the assimilation potential of the environment. there are two methods to add the sustainable development conditions of dynamic trajectories in the model as follows: – establishing temporal constraints for welfare level (at any time, welfare level must exceed a given threshold or welfare dynamics must have a uniform nondecreasing character); – establishing physical constraints for resources (stocks and flows). an important class of dynamic models of territorial economic development is formed by the so-called computable general equilibrium (cge) models, see [6, 23]. these models have good microeconomic grounds and also provide a complete description for the sectoral structure of economy and mutual effects of different economic sectors. however, the cge models suffer from a major drawback: in fact, they are mostly economic-mathematical models with a superficial treatment of ecological and social aspects and all the phenomena connected with dynamics and uncertainty. besides, their identification is also a complex procedure. the modeling framework of regional social-ecological-economic development using optimal control methods was suggested by gurman et al. [19]. regional development problems are studied by analytical methods in combination with simulation modeling. in the monograph [8], the well-known solow growth model [17] was modified taking into account the spatial aspect and environmental pollution. a detailed survey of the models and decision support systems for sustainable management was given in [21]; various economic growth models were described in [5]. the concept of homeostasis plays a key role for sustainable management. it can be formalized using aubin’s viability theory [4]. in particular, the so-called capture-viability kernels are of assistance here. the capture-viability kernel of a closed set k with closed target x under a set-valued dynamic f is the set of all initial states from which there exists at least one solution, remaining in k and reaching x at a given finite time horizon [7]. in the earlier works, the authors of this paper introduced an original complex approach to model the processes of accumulation and productive use of tangible and intangible assets – resources for territorial sustainable development as well as to model the coordination processes of social and private interests for resource allocation in hierarchical control systems. an implementation of this approach allowed to establish the system compatibility conditions for different control problem setups; to justify a control strategy design methodology for balancing the social and economic interests of national, regional (local) and global economic agents in the reproduction and utilization of welfare resources for the sustainable development of a regional system; to propose system coordination mechanisms and study their properties; to initiate application of the coordination models of social and private interests to real regional 138 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) management problems, more specifically, to the design of administrative and economic coordination mechanisms (engines) for the interests of territorial subjects [3, 9, 10, 12]. fig. 1.1 the conceptual model of subject-object relations [12] model tools for sustainable management 139 copyright ©2018 assa adv. in systems science and appl. (2018) the problems of rationalizing the mutual influence of the subjects of environmental and economic relations within certain territorial boundaries have been repeatedly considered by the authors [1, 2, 16, 20]. in accordance with the concept of sustainable development, which emphasizes the vital importance of current needs satisfaction taking into account the interests of future generations, risk is a significant quantitative indicator of balanced decisions. for these purposes, a control algorithm for the ecological and economic risks of urbanized territories development was suggested in [20]. for managing the social and economic risks of territorial development, earlier the authors introduced the economic and mathematical models based on specification of the latent mutual influences of welfare capital accumulation and the pace of innovative sustainable development of a territory [15] and also on the use of an econometrically identified relationship between population health and the variations of environmental parameters as a primary functional characteristic of social and economic damage [2]. an implementation of the sustainable management methodology at the regional level requires a regional information-analytical support system [22] as a basic technological tool of management. this system has a hierarchical structure, and the subsystems at the lower levels of management can be used independently. the remainder of this paper is organized as follows. in section 2, the structure of the dynamic model of a regional social-ecological-economic system, its main variables and processes as well as operation of this model are presented. section 3 is dedicated to the coordination mechanisms (engines) of interests in the public-private partnership and their use in the model. next, in section 4, risk management methods at territorial level are characterized. section 5 is to synthesize the above-mentioned methods and models for territorial sustainable management system design. finally, in section 6, some concluding remarks are given. 2. dynamic model of regional social-ecological-economic system the model of a regional social-ecological-economic system has the form )())(()()( 1 tlrtktaty ii iiiii    ; (2.1) )()()( tytsti iii  ; (2.2) )()](1[)( tytstc iii  ; (2.3) )()1()1( trtr iii  ; (2.4)    n j jjiiii tittktk 0 )()()()1()1(  ; (2.5) )()1()1( tlmbtl iiii  ; (2.6) )]()()][(1[)( tlbtkbtvctp i a lii a ki a i a i a i  ; (2.7) )]()()][(1[)( tlbtkbtvctp i w lii w ki w i w i w i  ; (2.8) ;)0(;)0(;)0( 000 iiiiii rrllkk  (2.9) 140 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) ;0)(;0)(;0;1)(0;1)()()( 0   tvtvtstvtvt w i a iiji w i a i n j ij  (2.10) ,...2,1,0;,...,1,0,  tnji as a rule, index i denotes a municipal unit within a region (e.g., within a federal subject). it can be associated with an enterprise but, in this case, productive and economic activity should be described by the firm model. in the sequel, it will be referred to as the agent’s index. the model has discrete times t = 0, 1, 2, ... with a step of 1 year. the model also includes other variables and parameters as follows. )(tyi as the final output of agent i in year t (in financial terms); )(tk i as the agent’s basic production assets (capital) in year t; )(tli as the agent’s labor resources in year t; ( )ir t as the efficiency of the agent’s labor resources in year t; )(tai as the influence function of the agent’s innovative activity on the final output in year t; i as the parameter of the agent’s cobb-douglas production function; )(ti i as the agent’s production investments in year t; )(tci as the agent’s nonproduction consumption in year t; )(tsi as the share of the agent’s production investments in its final output in year t; i as the efficiency growth parameter of the agent’s labor resources; i as the depreciation factor of the agent’s basic assets; )(tij as the share of the investments of agent i in the activity of agent j (the cooperation coefficient of these agents); here index j=0 describes an external agent for the whole system; ib and im as the reproduction and retirement coefficients of the agent’s labor resources, respectively; ( )a ip t and ( )w ip t as the agent’s pollutant emissions into air and water in year t, respectively; ( )a iv t and ( )w iv t as the agent’s allocations to prevent air and water pollution in year t, respectively; a ic and w ic as the efficiency coefficients of these allocations; a kib and w kib as the specific rates of industrial pollution into air and water, respectively; a lib and w lib as the specific rates of human pollution (labor resources) into air and water, respectively; finally, 0 0, ,i ik l and 0 ir as given initial values of the model variables. therefore, the agent’s state vector is ))(),(),(),(),(),(),(),(()( tptptrtltktctitytx w i a iiiiiiii  ; the control vector is ))}({),(),(),(()( 0 n jij w i a iii ttvtvtstu   ; and the parameter vector is ),,,,,,,,,,( w li a li w ki a ki w i a iiiiiii bbbbccmbz  . the whole system can be written as ),...,()),(),...,(()()),(),...,(()( 111 nnn zzztutututxtxtx  . (2.11) the innovative activity function )(tai is considered separately [17]. with these notations, model (2.1)-(2.10) takes the form )),(),(()()1( ztutxftxtx iii  ; (2.12) model tools for sustainable management 141 copyright ©2018 assa adv. in systems science and appl. (2018) ii tu )( ; (2.13) ,...2,1,0,,...,1,)0( 0  tnixx ii (2.14) the structure of model (2.1)-(2.10) has three blocks, namely, economy (2.1)-(2.5), demography (2.6), and ecology (2.7)-(2.8). the transboundary interaction of municipal units is described using the control variables )(tij . also note that innovative activity can be modeled by through model parameters instead of the function )(tai . more specifically, the parameter i characterizes labor productivity; the parameter i resource-saving technologies; the parameters , , ,a w a i i kic c b and w kib nature protection technologies; the parameter ib demographic policy; the parameter im innovations in public health; the parameters a lib and w lib the ecological consciousness of population. to identify the parameters of model (2.1)-(2.10), it is necessary to define the numerical values of all elements of the vector z. in a rough approximation, consider just two values of each parameter },{ h i l ii zzz  , where the second value corresponds to higher technological level. for further refinement, the parameter values can be taken from a discrete set },...,{ 1 ik iii zzz  . the sustainable development conditions (homeostasis) of a regional social-ecologicaleconomic system in model (2.1)-(2.10) can be defined by ,...2,1;,...,1;)(;)(;)( maxmaxmin  tniptpptpyty w i w i a i a iii (2.15) the first condition in (2.15) states the requirements to the agent’s economic growth; the other, the maximum permissible emissions of pollutants into the environment. these conditions can be interpreted as the goal of control. the agent’s objective function in model (2.1)-(2.10) takes the natural form max)( 0     dttcej i t i  , (2.16) where )(/)()( tltctc iii  is the current specific consumption of agent i;  denotes the discount factor. since the model has discrete time and will be studied using simulation, the objective function (2.16) should be replaced by its discrete analog    t t i t i tcej 0 )( max . (2.17) simulation scenarios for model (2.1)-(2.10) include some trajectories of control variables of vector (2.13). here it seems reasonable to perform optimization by solving a certain control problem. a classification of such problems is given in table 2.1. table 2.1. control problem setups within the model decentralized centralized optimal control optimization performed by each agent independently (optimal control problem) global optimization performed by an external agent (optimal control problem) conflict control conflict resolution with interaction of agents (differential normal form game) conflict resolution with the optimal response of agents (hierarchical differential normal form game) 142 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) the decentralized control problem setups treat all agents equally. if the situation is considered from the viewpoint of a single agent, then the optimal control problem with the objective function (2.16) or (2.17) subject to constraints (2.1)-(2.10) arises accordingly. attempts to study the interaction of several or all agents lead to differential normal form games, with nash equilibrium as a standard solution concept. on the other hand, the centralized setups interpret one of the agents as a principal responsible for the system goals. in practice, the principal’s role can be played by federal government, regional administration or even a special coordination body for the activity of several agents or the whole set of agents established on voluntary basis (e.g., coordinating committee, project directorate, etc.). as before, the principal may perform global optimization or take into account the agents’ response to its actions. in the latter case, hierarchical differential games are immediate, with stackelberg equilibrium or the germeier principle of guaranteed result as possible solution concepts, see [22]. 3. coordination models for social and private interests in [9] and other publications, the authors constructed and analyzed coordination models for social and private interests (spice-models). the suggested approach consists in the following. 1. consider n agents, each distributing his/her/its resource between private activity and the production of some social good. 2. in turn, the produced social good is allocated among the agents in given or controlled shares. this defines the distinction between social good and pure public good. 3. the payoff function of each agent includes two terms, the first reflecting his/her/its income from private activity while the second his/her/its share in social good. 4. the concept of system compatibility is defined to characterize in quantitative terms the degree of concordance of social and private interests in a corresponding organizational and technical system. perfect system compatibility is achieved if the aggregate of the individually optimal strategies of all agents (dominant strategies or nash equilibria) maximizes the social welfare function of the system. 5. in real economic organizations, perfect system compatibility is a rare thing due to the individualism of agents. therefore, a special agent (principal) should be assigned to represent the social interests (social welfare maximization) and ensure system compatibility. 6. the principal controls the agents in two ways. first, he/she/it may restrict the share of resource allocated by the agents to their private activity (administrative mechanisms, compulsion). second, the principal may determine the shares of all agents in social good depending on their actions (economic mechanisms, impulsion). speaking mathematically, these control mechanisms are formalized as the germeier games г1 and г2 (the stackelberg and inverse stackelberg games, respectively). an implementation of this approach yielded the following. – system compatibility conditions were obtained for different control problem setups. the major result is that, without external control, perfect system compatibility can be achieved only by partitioning the set of all agents into pure individualists (distributing all their resources to private activity) and pure collectivists (distributing all their resources to social good production). in addition, the spice-models were compared in the cases of independent equal agents (nash equilibrium), hierarchically organized agents (stackelberg equilibrium), and full cooperation of agents (the pareto maximal value of the total payoff function). it was demonstrated that equality is preferable to hierarchy in the sense of social welfare maximization; – system coordination mechanisms were suggested and their properties were studied. a mechanism is system compatible if it ensures perfect system compatibility. here empirical and theoretical approaches are possible as follows. in accordance with the former, it is necessary to analyze the mechanisms widely used in practice (e.g., proportional allocation). for instance, it was shown that the proportional allocation mechanism is system compatible only under the model tools for sustainable management 143 copyright ©2018 assa adv. in systems science and appl. (2018) linear social income function. the theoretical approach suggests to construct the control mechanism as the ε-optimal solution of the germeier game г2 (the inverse stackelberg game). the economic mechanism based on the germeier game г2 is system compatible if social and private incomes represent power functions with an exponential less than 1; – an application of the spice-models to real regional management problems was also initiated; more specifically, the administrative and economic coordination mechanisms for the interests of regional subjects were explored. a control problem was studied in which two or more neighbor subjects distribute their funds between the development of their own and common (transboundary) territory or between private activity and a joint project. a special control authority (principal) was introduced for activity coordination. the economic mechanism was considered in two modifications (financial participation in the development of a transboundary territory via income share control, and resource allocation). a detailed analysis of these mechanisms was carried out and their organizational and economic interpretation for specific regional management problems was given (public-private partnership, euroregions), see [3,10]. the static spice-model has the form max),...,()(),...,( 11  niiiini uucsurpuug (3.1) .,...,1, ,0,0 ,0:,1 ,0,0,0 1 ni si si ssrru n j i i jiiii          (3.2) the notations are as follows: },...,1{ nn  as a finite set of active agents; ],0[ ii ru  as the set of admissible strategies of agent i ; ir as an amount of resource available to agent i ; ),...,( 1 ni uug as the payoff function of agent i ; ;...,: 1 ni uuurug  )( iii urp  as the private interest function of agent i ; ),...,( 1 nuuc as a social income function; is as the share of social income allocated to agent i ; ),...,( 1 ni uucs as the social component of the payoff function of agent i . the following assumptions are applied to this model: c is a monotonically increasing function in all iu such that ;0)0,...,0( c ip are monotonically increasing functions in )( ii ur  and monotonically decreasing in iu such that 0)0( ip (for ii ru  ); 0is  if 0iu . construct the social welfare function    n j njjj n j njn uucurpuuguug 1 1 1 11 ).,...,()(),...,(),...,( (3.3) denote by },...,{ )()1( ne k ne uune  the set of nash equilibria in game (3.1)–(3.2). also let )}(),...,(min{ )()1(min ne k nene ugugg  and ).()(max max max ugugg uu   then the price of anarchy in model (3.1)–(3.2) (an indicator of system compatibility) is max min g g pa ne  . (3.4) obviously, 1pa . if pa is close to 1, then the equilibria have high efficiency and there is little need for system coordination in model (3.1)–(3.2) (for 1pa , even no need at all). the need for system coordination grows as pa is decreased. as a matter of fact, in itself the condition of system compatibility ( 1pa ) holds merely in some cases. to ensure this condition, it seems reasonable to use control mechanisms [18]. 144 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) suppose that maximization of the social welfare function (3.3) is the goal of a certain subject (principal, leader, or mechanism designer) that may influence the sets of admissible controls and/or payoff functions of agents to achieve it. denote by )( iii quu  the first possibility and by ),( iiii usgg  the second. then the following control mechanisms are immediate, see table 3.1. the principal may influence the sets of admissible controls of agents (administrative mechanism) or their payoff functions (economic mechanism). both types of influence are based on the germeier games г1 and г2 (the stackelberg and inverse stackelberg games, respectively). therefore, there are four types of control mechanisms illustrated in table 3.1. table 3.1. control mechanisms principal influences: based on germeier games г1 (stackelberg games) based on germeier games г2 (inverse stackelberg games) the sets of admissible controls of agents compulsion, or administrative mechanism, without feedback constqi  compulsion, or administrative mechanism, with feedback )(uqq ii  the payoff functions of agents impulsion, or economic mechanism, without feedback constsi  impulsion, or economic mechanism, with feedback )(uss ii  the economic control mechanisms in the spice-model (3.1)-(3.2) are implemented as the result of choosing values is by the principal. for the administrative mechanisms, an additional assumption is that the principal may restrict the admissible controls of agents, i.e. niquq iii  ,~ . (3.5) in the continuous-time setup with finite horizon, the agent’s objective function (2.16) takes the form                  t n i ij iijii t i dttifttcej 0 1 max])()()([  ; (3.6) in the discrete-time setup with the same horizon,                   t t n i ij iijii t i tifttcej 0 1 max])()()([  . (3.7) in both cases, f denotes the social welfare function while ρ is the discount factor. note that, by choosing a control trajectory, the agent defines the logical chain iiiiii ccyk  in model (2.1)–(2.10), which corresponds to the choice of allocations to private activity in the static spice-model. therefore, the first term in the integrand (summand) describes the agent’s model tools for sustainable management 145 copyright ©2018 assa adv. in systems science and appl. (2018) income from private activity; the second, his/her/its consumption of social good depending on the variable )(ti . the administrative and economic mechanisms can be described for the dynamic spicemodel with a principal. using administrative mechanisms, the principal explicitly restricts the “egoism” of agents by imposing the conditions )()( max tt iiii   . (3.8) note that conditions (3.8) can be the result of a voluntary agreement of the agents. if they are established by the principal, then his/her/its objective function also includes administrative control cost. economic mechanisms may have a share-based motivation of the agents, i.e. ))(()( tt iiii   , (3.9) or a resource allocation among them performed by the principal. 4. approaches to risk factors management in territorial development sustainable development is impossible without proper consideration of risks, which have to be identified first. a detailed identification of territorial risks facilitates the objective formalization and modeling of territorial development, yielding adequate assessments for possible consequences of decisions using the integral risk indicator of a territory. territorial risk assessment methods are mostly reduced to score calculation in some rating scales for comparing the development levels of different regions. speaking formally, the existing approaches to territorial risk analysis can be divided into several groups, namely, qualitative analysis, quantitative analysis, combined analysis, and structural analysis. qualitative analysis considers the weights of different factors affecting risks. note that the list and significance of such factors are defined via expertise. major disadvantages of this approach consist in the subjectivism and qualification of experts. therefore, qualitative analysis makes sense only in the case of well-defined goals of study and available highly qualified experts, who have rich experience and sufficient familiarity with regional situation. quantitative analysis is often performed for most important risk indicators: consideration of all risk factors would dramatically increase the complexity of calculations and further examination. this approach allows constructing a multifactor function for explicit quantitative assessments. there are two standard methods to define regional risk functions as follows. in accordance with the first method, it is necessary to consider only those factors that yield an objective quantitative characterization for the current state of economical, ecological, and social spheres. then the territorial risk function takes the form r = f (x1, x2, …, xn) = r (xi), i = 1,...,n. (4.1) the second approach involves a set of numerical qualitative assessments, i.e. r = f (r1, r2, …, rm) = r (ri), i = 1,...,m. (4.2) econometric modeling is intended for statistical data processing and prediction based on quantitative data arrays of key indicators. however, in real conditions, some significant factors 146 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) and criteria (e.g., social and ecological situation in a region, or the synergetic effect of several environmental conditions) cannot be assessed in quantitative terms. in this context, a reasonable solution is to integrate quantitative and qualitative assessment methods, which makes the essence of combined approaches. in this case, the integral risk indicator includes heterogeneous parameters of different risks in numerical form obtained by objective calculations and also by subjective qualitative assessments, which are considered in some scales or criteria. the identified factors are used for assessing the risks of comparable level. the integral risk indicator r (4.1), (4.2) can be expressed as a linear relationship of the corresponding indicators calculated for the risk of smaller level, i.e. r = f (x1, x2, x3, x4, x5) = a0 + a1x1 + a2x2 + a3x3 + a4x4 + a5x5. (4.3) here the notations are as follows: х1 = y (b1,…,bn) as the function of environmental and ecological factors; х2 = y(c1,…,cm) as the function of social and demographic factors; х3 = y (t1,…, tk) as the function of anthropogenic and production factors; х4 = y (s1,…,sj) as the function of financial and economic factors; finally, х5 = y (h1,…,hg) as the function of organizational and managerial factors. these variables characterize the specific levels of corresponding risks; their weights а1, …, а5 can be defined using qualitative methods in order to consider a basic level of risk formed by the associated factors. the described approach to territorial risk assessment is rather flexible and convenient for practical implementation. moreover, it takes into account any changes in structural parameters through the coefficients of the corresponding variables, which gives an objective description for the dynamics of external and internal territorial environment. the sustainable development programs of different territories can be elaborated using the scenario approach. the number of scenarios under study (alternatives) may vary depending on regional capabilities and resources. as a rule, developmental prospects are considered in view of three alternatives. the first regional development scenario is the preservation of current social and economic dynamics at the same level. this scenario implies that territorial authorities have no significant interference into social and economic development. their major concern and funding are focused on maintaining the current volumes of regional gross domestic product (gdp) and the existing sectors of regional industry; economy is developing with orientation towards export of resources and raw materials. the second territorial development scenario is based on the implementation of continuous investment projects and programs, not only in the sphere of raw materials and semi-products but also in financial and industrial groups, resource processing and resource supply. for this scenario, of crucial importance are the industries and productions oriented towards import substitution, high technologies, finished products and services. the third territorial development scenario is a logical continuation of the second; it can be implemented in several fields simultaneously through initiation of long-term projects (clusters and growth drivers in industries, high technologies, science and education). in contrast to the productions oriented towards export of resources and raw materials, these projects can be implemented in form of medium and even small enterprises, hence with smaller capital investments and higher rates of return. in the long run, this scenario yields the multiplicative effect owing to the creation of new jobs, the development of adjacent, auxiliary and supporting industries, and natural clustering. the structural method of territorial risk assessment is based on the expertise of given quantitative parameters––the probability and amount of losses. for the identified risks, this approach involves probabilistic weighting of each development scenario to distribute the final result. the territories with similar level and initial conditions of development can be compared using the integral risk indicator that includes several structural components, namely, (a) the ratio model tools for sustainable management 147 copyright ©2018 assa adv. in systems science and appl. (2018) of regional gpd and average gdp for a corresponding group of regions; (b) the ratio of the deviation of regional gdp from average gdp for a corresponding group of regions; (c) the ratio of regional gpd increase rate and average gdp increase rate for a corresponding group of regions. also this approach is a snap analysis tool for territorial risks in a given group of regions. note that the identification of possible risks and the assessment of average territorial risk must agree with the goals of development. for strategic development programming, such procedures allow to find territorial fields and spheres in which risk management is applicable and also reasonable. consequently, the resulting assessments can be used to refine territorial development programs taking into account risk management methods and tools, territorial risk management system optimization and the prediction of possible consequences of different risk events. in this regard, there exists an obvious need for integrating the risk management subsystem directly into the territorial sustainable management system: otherwise, the whole complex of well-known approaches to risks identification, assessment and management would be fruitless, not gaining an expected effect. integration must cover all levels of the management system, relying on available administrative resources and the support of all structural elements. this is the only way towards the required efficiency of risks monitoring and optimization with territorial sustainable development. the territorial risk management subsystem is deployed in the following way. the functions and authorities are divided by management levels, industries, and spheres of development. necessary management information is exchanged by the structural units of the system in the online mode. all necessary resources of each management level and any sphere of activity are immediately mobilized. the goals of risk management must agree with the goals of territorial development for achieving strategic stability based on the prediction of possible threats and negative factors in the long run. modern science has produced numerous methods to affect risks. almost any territorial risk can be managed properly. one may control the probability of risks and also the amount of incurred losses by preventive and protective measures. evidence suggests that purely preventive methods do not completely eliminate the negative factors of risk in form of damage and loss conditions. therefore, it is crucial to plan complex measures towards the abolition and/or absorption of territorial risks at the level of territorial administration, particularly with the engagement of large territorial taxpayers. the external and global risks that are uncontrollable at the territorial level must be transferred to other management levels or secured using efficient financial mechanisms and tools (e.g., insurance). in this case, possible damage is compensated at necessary level (in terms of time and amount), which makes the territory independent of environmental disasters and anthropogenic accidents. so the strategic goals of territorial development can be achieved. 5. territorial sustainable development management system design based on mathematical models and information technologies the territorial sustainable management system covers organizational and political as well as technical aspects. the former concerns the existence of well-defined sustainable development regulations and procedures for territorial (legislative and executive) authorities. in particular, of crucial importance is territorial monitoring that implements feedback in the management system through the acquisition of necessary information about the dynamics of the territorial socialecological-economic system. the technical part of the system is represented by the regional information-analytical sustainable development management support system, see fig. 5.1. 148 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) fig. 5.1 structure of regional information-analytical sustainable development management support system [22] this information-analytical system has the client-server architecture. the client part includes the interfaces of models and data; the other blocks belong to the server part. operative databases contain information about separate municipalities and large enterprises, which is provided by the monitoring system. for regional management, it is necessary to integrate these (often heterogeneous) data. this problem is solved by importing all available data to the regional data warehouse. the latter also implements other functions such as data refinement, aggregation and security, processing of large data arrays, creation of multilevel metadata directories, execution of user requests, and generation of different reports. the analytical block consists of three subsystems––simulation, optimization and expert ones. the simulation subsystem implements model (2.1)–(2.10) using data from the regional data warehouse. the optimization subsystem solves the control problems described in table 3.1 using the dynamic optimality criteria (3.7) subject to constraints (2.1)–(2.10). the expert subsystem is optional: this subsystem allows to consider the control rules used in practice. the internal system interface is responsible for the interaction of models and data as well as for the implementation of different calculation schemes. a major role is played by the external user interface, which provides a user-friendly environment for common users of the regional information-analytical sustainable management support system. 6. conclusions the intrinsic character of a territorial system as a multilayer polystructural complex of heterogeneous subsystems (economic, social, and ecological) suggests the idea to consider such systems as the subjects of sustainable economic growth and to review the methodological approaches to territorial management. these approaches must be focused on the territorial reproduction of qualitative resources (first of all, human capital) based on proper coordination of interests of different subjects and public-private partnership [11]. overcoming the shortcomings of modern territorial management and emerging risks in its system is inextricably linked with the enhancement of the effectiveness of the mechanism for managing sustainable development of the territory on the basis of the current model tools and information technologies model tools for sustainable management 149 copyright ©2018 assa adv. in systems science and appl. (2018) no doubt, the methodology and model tools suggested in this paper will accelerate the positive processes of the innovation-oriented transformation of territorial sustainable development management systems. for the new strategic policy and its efficiency, a significant methodological aspect of dynamic systems analysis is the gradual transition to the management processes of functional-spatial territorial development, with highest priority assigned to the formation of innovative clusters and behavioral rules (economic, social, ecological, etc.) for the elements of social environment, see [13]. this work was supported by the russian foundation for basic research, project no. 18-010-00594. references [1] anopchenko, t.yu. & murzin, a.d. (2012). upravlenie riskami razvitiya urbanizirovannykh territorii [risk management in development of urbanized territories]. rostov-on-don: rostov state university of civil engineering [in russian]. [2] anopchenko, t.yu. & murzin, a.d. (2014). economic-mathematical modeling of social and environmental risks management of projects of urbanized territories development, asian social science, 10(15), 249–254. [3] anopchenko, t.yu., murzin, a.d. & ougolnitsky, g.a. (2017). modelirovanie soglasovaniya interesov v zadachakh upravleniya ustoichivym razvitiem territorii [modeling of interests coordination in regional sustainable management problems], ekonomika prirodopol'zovaniya, 6, 35-47. [4] aubin, j.-p. (1991). viability theory. birkhauser: springer-verlag. [5] barro, r.j. & sala-i-martin, x. (2004). economic growth. massachusetts: mit press. [6] bohringer, c. & loschel, a. (2006). computable general equilibrium models sustainability impact assessment: status quo and prospects, ecological economics, 60 (1), 49-64. [7] bonneuil, n. & boucekkine, r. (2017). viable nash equilibria in the problem of common pollution, pure and applied functional analysis, 2(3), 427-440. [8] druzhinin, a.g. & ougolnitsky, g.a. (2013). ustoichivoe razvitie territorial'nykh sotsial'no-ekonomicheskikh sistem: teoriya i praktika modelirovaniya [sustainable development of regional social and economic systems: theory and practice of modeling]. moscow, russia: vuzovskaya kniga [in russian]. [9] gorbaneva, o.i. & ougolnitsky, g.a. (2015). system compatibility: price of anarchy and control mechanisms in the models of concordance of private and public interests, advances in systems science and applications, 15(1), 45-59. [10] gorbaneva, o.i., murzin, a.d. & ougolnitsky, g.a. (2018). mekhanizmy soglasovaniya interesov pri upravlenii proektami razvitiya territorii [mechanisms of interests’ coordination in project management of regional development], upravlenie bol'shimi sistemami, 71, 61-97. [11] lazareva, e. & karaycheva, o. (2017). human oriented reframing of the territories of innovative sustainable development system management model, sgem 2017 proceedings, book 4, vol. 2, 672-670. [12] lazareva, e.i. (2013). features of national welfare innovative potential parametric indication information-analytical tools system in the globalization trends' context, ceur workshop proceedings 9, integration, harmonization and knowledge transfer, 339-351. [13] lazareva, e., anopchenko, t. & lozovitskaya, d. (2016). identification of the city welfare economics strategic management innovative model in the global challenges conditions, sgem 2016 proceedings, book 2, vol. 4, 3–11. https://elibrary.ru/item.asp?id=23979708 https://elibrary.ru/item.asp?id=23979708 https://elibrary.ru/item.asp?id=23979700 https://elibrary.ru/item.asp?id=23979700 150 t.yu. anopchenko, o.i. gorbaneva, e.i. lazareva, a.d. murzin, g.a. ougolnitsky copyright ©2018 assa adv. in systems science and appl. (2018) [14] lazareva, e.i. (2011). metody modelirovaniya innovatsionno-orientirovannykh ekonomicheskikh strategii ekologoustoychivogo razvitiya [modeling methods of innovationoriented economic strategies of ecological sustainable development]. rostov-on-don: southern federal university [in russian]. [15] lazareva, e.i. (2013). osobennosti modelirovaniya traektorii prirashcheniya kapitala natsional'nogo blagosostoyaniya v perspektive ustoichivogo innovatsionno-orientirovannogo razvitiya [modeling specifics for capital increment trajectories of national welfare capital subject to sustainable innovation-oriented development], partnerstvo tsivilizatsii, 4, 234-244. [16] lazareva, e.i. et al. (2012). modeli ekologicheskoi orientatsii gosudarstvenno-chastnykh strategii investirovaniya chelovecheskogo kapitala v innovatsionnoi ekonomike [models of ecological orientation of public-private strategies of human capital investment in innovative economy]. rostov-on-don: southern federal university [in russian]. [17] lotov, a.v. (1984). vvedenie v ekonomiko-matematicheskoe modelirovanie [introduction to economic-mathematical modeling]. moscow, russia: nauka [in russian]. [18] mechanism design and management: mathematical methods for smart organizations, ed. by prof. d. novikov. new york: nova science publishers, 2013. [19] modelirovanie sotsio-ekologo-ekonomicheskoi sistemy regiona [modeling of socialecological-economic system of region], ed. by gurman, v.i. & ryumina, e.v. moscow, russia: nauka (2001). [20] murzin, a.d. (2015). algorithmization of ecologo-economic risk-management in urban areas, asian social science, 11(9), 312-319. [21] ougolnitsky, g.a. (2015). sustainable management as a key to sustainable development, in sustainable development: processes, challenges and prospects, ed. by d. reyes. new york: nova science publishers, 87-128. [22] ougolnitsky, g.a. (2017). a system approach to the regional sustainable management, advances in systems science and applications, 17(2), 52-62. [23] partridge, m.d. & rickman, d.s. (2010). cge modeling for regional economic development analysis, regional studies, 44(10), 1311-1328.  2017 г adv syst sci appl 2018; 1; 59-84 published online at http://ijassa.ipu.ru. analytical study of the antitumor viral vaccine introduction regimens based on mathematical modeling nina a. babushkina1, ekaterina a. kuzina1 1) v.a. trapeznikov institute of control sciences of russian academy of sciences, 117997, profsoyuznaya street, 65, moscow, russia e-mail: babushkina_na@mail.ru, kate_k93@mail.ru abstract. the paper presents the algorithm for calculating the maximal effective antitumor viral vaccine introduction regimens, using a computing experiment method (in silico) based on the software implementation of two mathematical models in the matlab-simulink system. the first model of antitumor vaccine therapy describes a two-stage mechanism of the tumor cells’ death as a result of the immune response. the effectiveness of immune response is measured in the number of antibodies formed by the immune system against virus-infected tumor cells. the second model of antitumor therapy with discontinuous trajectories of tumor growth is designed to evaluate the rate of the tumor cells’ death after the introduction of the viral vaccine. the effectiveness of the therapy is measured in the number of dying tumor cells after the introduction of a viral vaccine. keywords: mathematical model, experimental oncology, tumor cells, kinetic curves of tumor growth, tumor growth delay, virus, vaccine therapy, immune response, number of antibodies, vaccine efficacy, time-frame for recurrent vaccine introductions, in silico. 1. introduction vaccine therapy is one of the methods of immune therapy of the oncological diseases. it implies vaccination against the malignant cells formed in the body. as with any vaccination, the introduction of viral vaccines stimulates the body's immune system and causes the formation of antibodies specific to the tumor type. experiments in tumor immunotherapy can be traced back to ehrlich [1], and since the middle of the twentieth century the studies aiming to find the viruses that are able to affect malignant tumors without causing any harm to the humans commenced in the soviet union and were further continued in russia [2-6]. one of approaches to developing effective antitumor vaccines implies using the viruses that are able to identify malignant cells and do not cause any damage to the normal tissues. the viruses selected for therapy should be clinically harmless and epidemiologically safe. therefore, oncotropic viruses must have a genetically fixed absence of pathogenic properties and be unable to cause an acute infectious disease in a human body. among the viruses that meet these criteria are the venezuelan horse encephalomyelitis virus (vee) [5-7] and rat parvovirus h-1 [15-23]. both are not harmful to humans because they do not affect healthy tissues since they cannot replicate in the non-dividing cells [16-18]. thus these viruses are able to identify the tumor cells based on their high proliferation rates, and specifically target them without attacking the normal tissues [18-19]. experimental studies of rat parvovirus h-1 indicate that there are two possible mechanisms of the tumor cells’ death as a result of vaccine introduction. firstly, the virus blocks the cycle of the tumor cells [16-18]. secondly, it produces specific protein formations on the tumor cell surface that stimulate the immune system to develop antibodies against these cells, resulting in their death [13-14, 19-21]. this corresponds to the results of the experimental studies of the impact of the vee virus on tumor growth [4-6]. however, since such experiments require mailto:babushkina_na@mail.ru mailto:kate_k93@mail.ru 60 n.a. babushkina, e.a. kuzina extensive funding, these studies did not determine the maximal and minimal effective dosages and the optimal moment of vaccine introduction. thus, the present study presents an algorithm that allows to determine the optimal dosage of the vee virus-derived vaccine and the moment of its introduction that will ensure the maximal rate of tumor cells’ death, building on the mathematical models of anti-tumor therapy and vaccine therapy developed in [7-9]. 2. the mathematical model of the vee virus vaccine therapy the mathematical model of vaccine therapy describes the mechanism of two-stage death of tumor cells caused by the virus itself and as a result of the immune system stimulation against the infected tumor cells. based on the experimental data on the growth of erlich adenocarcinoma in mice after a single introduction of the vee virus vaccine, one day after the tumor was transplanted to animals (figure 1), the parameters of the model were estimated and its adequacy was tested using experimental dosages of the vee virus vaccine [8-11]. the values of the model parameters are shown in appendix 1. fig. 1. experimental curves of the ehrlich adenocarcinoma growth without vaccine administration (control) and after single vaccine administration the mathematical model of the two-stage death of tumor cells after the viral vaccine introduction is described by the following system of differential equations (1-8) [7-10]: ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) v v v v v av v v v v n n an n n n dn t t n t t t t t k v t t k a t t n t t t t dt t t k a t t n t t       =   − + − − − − −  − − +   + − − −  −  (1) where ( )v t is the number of viruses, ( )n t is the number of tumor cells without vaccine introduction, described by the following differential equation: ( ) ( ) ( ) dn t t n t dt =  , at 0 0( )n t n= (2) where  (t) is the rate of tumor cells’ proliferation, analytical study of antitumor viral vaccine 61 ( )vn t is the number of infected tumor cells after the vaccine introduction, 1v cvt z= + is the moment of the start of immune response against viruses, 1n cnt z= + is the moment of the start of immune response against tumor cells, 1 is the moment of vaccine introduction, cvz and cnz are the time lag of immune response against the virus and infected tumor cells, ( )v va t t− is the number of antibodies against the virus, ( )n na t t− is the number of antibodies against infected tumor cells, ( )vt t − and ( )nt t − are the heaviside step function, avk , ank , and vk are the dimension factors. the first component of equation (1) describes the dynamics of the tumor cells’ growth before the introduction of the viral vaccine. the second component of equation (1) describes the death of infected tumor cells caused by the virus. the first stage of the immune system stimulation is production of antibodies against the virus (figure 2 – ( )vc t , ( )va t ). at this stage the death of tumor cells is caused by the penetration of the virus inside the tumor cells (figure 3). the third component of the equation describes the death of tumor cells caused by antibodies produced against the infected tumor cells (figure 2 – ( )nc t , ( )na t ). reproducing in tumor cells, the viruses induce the development of new protein formations on the surface of the tumor cell membranes. such infected tumor cells are perceived by the immune system as foreign to the body. as a result, immune system produces antibodies specific to the tumor cells, and they destroy the infected tumor cells (figure 3). the patterns of immune response to the foreign cells are based on marchuk's mathematical model of infectious diseases [24-27]. this model treats the immune response (production of t and b cells) as a combined action. therefore in the present study the immune response means the formation of antibodies and plasma cells specific to the vee virus and infected tumor cells. using marchuk’s model, the dynamics of the number of viruses is described by the following equation: ( ) ( ) ( ) ( ),v v v dv t v t a t v t dt  = − (3) where 0 1( )v v = is the initial dose of the viral vaccine, 1 is the moment of the first administration of the viral vaccine, v is the rate of viruses’ reproduction within the cell, v is the death rate of viruses when they interact with antibodies ( )va t . in the vaccine therapy model, the initial condition 0 1( )v v = denotes the dosage of the vaccine. the immune response of the body to the virus introduction results in production of antibodies and plasma cells. their numbers are calculated using the following four differential equations [24-25] for each of the two stages of the immune response – against the virus and against the infected tumor cells. the dynamics of the number of antibodies to the virus ( )va t : 1 ( ) ( ) ( ) ( ) ( ),v a v v av v v v v v da t c t t a t t v t t a t dt    = − − − − − − (4) 62 n.a. babushkina, e.a. kuzina where a is the rate of formation of antibodies from a single plasma cell, av is the death rate of antibodies due to interaction with viruses ( )va t , v is the rate of decrease in the number of antibodies due to natural destruction. the dynamics of formation of plasma cells cv(t): 1 1 ( ) ( ) ( ) ( )v vc v cv v vn dc t v t a t t c t c dt         = − − − − − , at 1 ,( )v vnc c = (5) where с is the rate of formation of plasma cells, cv is the dimension factor, cvz is the time lag of the immune response to the formation of a plasma cell clone. the second term of this equation defines the deviation of the actual number of plasma cells from the norm cvn. the equation for the dynamics of the number of antibodies ( )na t acting against infected tumor cells is as follows: ( ) ( ) ( ) ( ) ( ),n an n n an n n v n nn n n da t c t t a t t n t t a t t dt   = − − − − − − (6) where an is the rate of formation of antibodies from a single plasma cell, an is the death rate of antibodies ( )na t due to interaction with infected tumor cells ( )vn t , nn is rate of decrease in the number of antibodies due to natural destruction. the dynamics of formation of plasma cells ( )nc t : , ( ) ( ) ( ) ( )n cn v n n cn n n nn dc t n t a t t c t t c dt        = − − − − (7) at 1( ) ,an nnc c = where cn is rate of formation of plasma cells, cn is the dimension factor, zcn is the time lag of the immune response to the formation of a plasma cell clone against the infected tumor cells. dynamics of antibodies ( )va t against the virus and against infected tumor cells ( )na t are shown in figure 2. fig. 2. population dynamics of antibodies ( )va t against the virus and against infected tumor cells an(t) for experimental dosage 0 0,015v = on 1 = day 1 (the first and second stages of the immune response) analytical study of antitumor viral vaccine 63 in the equations that describe the dynamics of the formation of antibodies (6) and plasma cells (7), the number of infected tumor cells was calculated as ( ) ( ) ( ),v n nn t n t p t=  where ( )np t accounts for the proportion of rapidly proliferating cell fraction as the size of the tumor increases [14]: 2 2 21 ( ) 1 ( ) , 1 p p p p t p t arctg k t       = −   −    (8) where p and kp are constant parameters, t is the current time of the tumor cell population growth, * 1 ,p t  = where *t is the moment when the numbers of the rapidly and slowly proliferating cells are equal. the fraction of rapidly proliferating tumor cells is located near the blood vessels (oxygenated fraction), while the fraction of slowly proliferating cells is pushed to the tumor periphery (hypoxic fraction). the process of the change in the proportion of the rapidly and slowly proliferating cells with the increase in the tumor size is described in the skipper model [31]. as the number of the fast-proliferating tumor cells decreases, the virus ceases to affect them. this causes the decrease of the tumor sensitivity to the viral vaccines and of the number of infected tumor cells ( )vn t that can induce the immune response. as a result, the number of the antibodies developed against tumor cells is reduced, and the effectiveness of the viral vaccines declines. thus, the mathematical model of vaccine therapy describes the mechanism of two-stage death of tumor cells caused by the virus itself and as a result of the immune system stimulation against the infected tumor cells. the calculated curves that describe the dynamics of the two-stage death of infected tumor cells ( )vn t and the production of antibodies and plasma cells at the each of the two stages of the immune system stimulation are shown in figure 3 [8-10]. fig. 3. calculated growth curve of the total number of tumor cells n(t) (full line) that approximates experimental data (+) based on the mathematical model of antitumor therapy (dosage 0 0,015v = on 1 = day 1); the dotted line indicates the dynamics of the number of infected tumor cells ( )vn t . 64 n.a. babushkina, e.a. kuzina calculations show that the curve that describes the dynamics of the number of infected tumor cells is located below the experimental curve of the tumor growth. this indicates that when the viral vaccine is administered, it does not infect the entire tumor cell population ( ).n t some cells remain uninfected ( )rn t and continue to multiply, causing a repeated growth of the tumor, as observed in the experiment (figure 1). some cells survive ( 0 vrn – after the 1st stage of immune response, 0 nrn – after the 2nd stage), causing a repeated growth of the tumor as observed in the experiment (figure 1). 3. the mathematical model of antitumor therapy with discontinous trajectories for assessing the tumor growth delay after the two-stage death of tumor cells the mathematical model of antitumor therapy with discontinuous trajectories was developed to control and optimize the regimes of applying different antitumor therapy methods based on kinetic curves of tumor growth [8-10]. figure 3 shows the calculated growth curve of the total number of tumor cells n(t) (full line), which was obtained based on the mathematical model of antitumor therapy by approximating the experimental data. the duration of tumor growth delay was estimated based on the following assumptions adopted in the mathematical model of antitumor therapy [8-10]: 1. the cell death occurs instantly at each of the two stages of the immune system stimulation. the abrupt decrease in the tumor size occurs at the moment when the maximal number of antibodies to the virus and infected tumor cells is produced. 2. reduction in the tumor size occurs only due to the death of the virus-infected tumor cells. 3. the surviving tumor cells continue to grow at the same reproduction rate. the kinetic curves of their growth are described by the gompertz function, keeping the parameters of the tumor growth without treatment (control), but accounting for the time shift caused by tumor growth delay 0( )v v and 0( )n v after each stage of their death (figure 4). according to the basic antitumor therapy model [8-10], the kinetic curves of tumor growth without treatment are described by a simple differential equation ( ) exp( ) ( )n n n dn t t n t dt   = − , at 0(0)n n= , (9) where 0n is the initial number of tumor cells at the time of tumor transplantation to animals at t = 0, n(t) is the number of cells in the tumor at time t, n and n are the parameters of the gompertz function, which approximates the experimental kinetic curve of ehrlich adenocarcinoma growth without treatment (figure 1) [5-7, 10-14]. based on the assumptions listed above, tumor growth after the introduction of viral vaccines follows a differential equation describing the two-stage death of tumor cells: ,0 1 0 , 1 0 ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ( )) ( ) ( ) ( ) ( ) ( ) ( ( )), v v v v v v n n v n n n n dn t t n t t t s v n t t t t t n t v t t dt s n n t t t t t n t v            = − − − + − − − − − − + − − (10) where ( )n t is the total number of tumor cells at time t, ( )t is the growth rate of a given type of a tumor, ( )vt t − and ( )nt t − is the heaviside step function, 1v cvt z= + and 1n cnt z= + are the moments of the tumor cells' death at each of the two stages of the immune system stimulation, 1 is the moment of the viral vaccine introduction, analytical study of antitumor viral vaccine 65 cvz and cnz denote the time lag of the immune response against the virus and infected tumor cells, ( )vt t − and ( )nt t − are the dirac delta function describing instant death of tumor cells at each of two stages of the immune system stimulation, ,0( )v vs v t and ,0( )v ns v t are the relative decrease in the tumor size at the time the tumor cells’ death at each of the two stages, ,0( )v vv t and ,0( )n nv t are the tumor growth delay after cells’ death at each of the two stages of the immune system stimulation (figure 3). in the vaccine therapy model, it was assumed that the relative decrease in tumor volume occurs only due to the death of the virus-infected tumor cells nv(t). thus in this model the number of dying cells was calculated as the difference between the maximum and minimum number of infected cells: 1 2( ) ( ) ( )v v v vv vn t n t n t = − and 1 2( ) ( ) ( )n n v vv nn t n t n t = − , (11) where 1 vt and 1 nt are the starting moments for the immune response against the virus and against the infected tumor cells, 2 vt and 2 nt are the ending moments of the immune response. according to the model of antitumor therapy, the relative decrease in tumor size at the time of the tumor cells’ death at each of the two stages was determined from the following equations: 0 ( ) ( , ) ( ) v v v v v n t s v t n t  = и 0 ( ) ( , ) , ( ) v n n n n n t s v t n t  = (12) where ( )vn t and ( )nn t denote the total number of cells in the tumor at the time points vt and nt . the death of tumor cells causes tumor growth delay, and its duration serves as a quantitative measure of the cells’ death depending on the dosage of the viral vaccine. the tumor growth delay ,0( )v vv t and ,0( )n nv t equals the time interval between the cells’ death and their subsequent recovery to the original size at the moment of vaccine introduction (figure 3): 0( ) ( ( , ))v v v vn t n t v t= + и 0( ) ( ( , )).n n n nn t n t v t= + (13) then the tumor growth after each stage of the infected cells' death is described by gompertz equations with a time shift that accounts for the duration of the tumor growth delay as follows: 0 0( ) exp( (1 exp( ( ( ))))),v v n n v vn t t rn t v  − = − − − (14) 0 0( ) exp( (1 exp( ( ( )))))n n n n n nn t t rn t v  − = − − − (15) where 0 vrn and 0 nrn is the number of surviving cells after the 1st and the 2nd stages of immune response. figure 3 shows the calculated curve that describes the total number of tumor cells n(t) (upper curve) derived from the mathematical model of antitumor therapy with discontinuous trajectories of tumor growth by approximating the experimental data. this curve describes the tumor growth n(t) following three gompertz equations (one for a control group) with a time shift that accounts for the tumor growth delay after each stage of the cells’ death [8-10]. appendix 1 shows the values of the parameters for the vaccine therapy model and the antitumor therapy model, estimated based on the experimental data provided in [10-14]. 66 n.a. babushkina, e.a. kuzina 4. algorithm for calculating effective administration regimens for viral vaccines mathematical model of vaccine therapy and mathematical model of antitumor therapy with discontinuous trajectories of tumor growth were developed in matlab-simulink to conduct a computing experiment (in silico). the experiment aimed to identify the efficient vaccine therapy regimes for a wide range of vee virus dosages and for different moments of vaccine introduction. the efficiency of the vaccine therapy was assessed after each of the two stages of immune response based on two parameters: 1 – number of produced antibodies ( )va t and ( )na t , 2 – number of dying tumor cells ( )v vn t and ( )v nn t . the efficiency of vaccine therapy method was estimated by two criteria. the first was the number of antibodies formed by the immune system at each of the two stages of its stimulation. we needed to determine the dosage that results in the maximum number of antibodies against infected tumor cells [10-13]. when this condition is met, the tumor-bearing organism is able to restrain the growth of the surviving tumor cells. the second criterion was the tumor growth delay. it depends on the number of dying cells on each of the two stages of immune response. therefore, the algorithm for calculating optimal regimen of the viral vaccine is based on defining a maximal effective dosage when the immune system produces the biggest number of antibodies. 4.1. calculating the viral vaccine dosage using the criterion of the production of antibodies against infected tumor cells based on the model of vaccine therapy with a single introduction of the varying dosages of the viral vaccines at the varying moments (from the 1st to 35th day), the graphs that show the changes in the number of antibodies against the virus ( )va t (figure 4) and antibodies against infected tumor cells ( )na t (figure 5) were developed [14]. these graphs reflect two fundamentally different types of the production of antibodies against the virus and against the tumor cells, depending on the vaccine dosage. the number of antibodies against the virus increases in proportion to the dosage increase, i.e. there is a linear dependence between these variables (figure 4). however, the increase in the number of antibodies against infected tumor cells 0 1( , )na v  is non-linear, and depends on the tumor size at the moment of vaccine introduction (figure 5). this non-linear dependence allows determining the virus dose max 0v = 0,024 that results in the maximal production of antibodies against the infected tumor cells. analytical study of antitumor viral vaccine 67 fig. 4. maximum number of antibodies against the virus max 0( )va v depending on the dose 0v administered on day 1 (the first stage of the immune system stimulation) fig. 5. maximum number of antibodies against infected tumor cells max 0( )na v depending on the dose of 0v administered at different moments (the second stage of the immune system stimulation); 1 = 1, 5, 15, 25, 35 days the non-linear dependence of production of antibodies on the dosage increase 0 1( , )na v  may be due to the changes in the number of infected tumor cells on the 1st and 2nd stages of immune response. when the dosage is small, the number of the dying cells is low, but many cells get infected by the virus. with the further increase of the dosage exceeding max 0v = 0,024, the number of cells dying at the 1st stage of immune response increases. this leads to reduction of the number of infected cells, and the consequent decrease in the number of antibodies against these cells on the 2nd stage of immune response. thus, with the administration of significant doses of the virus, almost the entire population of tumor cells is killed already at the 1st stage of immune response. this leads to a rapid and effective destruction of the tumor, but doesn’t stimulate the 2nd stage of immune response; therefore the antibodies against the infected tumor cells are not produced. 68 n.a. babushkina, e.a. kuzina the effectiveness of production of the antibodies against the infected tumor cells depends on the tumor size. as indicated by the calculations [14], after 35 days the number of fastproliferating tumor cells starts to decrease, and this leads to the decline in the number of infected tumor cells and antibodies against them at the later stages of vaccine therapy. thus, the calculations using the mathematical model of vaccine therapy allowed to define the maximal effective dosage of the vee virus-derived vaccine and to identify the moment of its administration that results in the maximal number of infected tumor cells maxt [14]. based on the model of vaccine therapy the graphs that show the number of infected tumor cells that die at the 1st and 2nd stages of immune response were developed based on equation (11) (figures 6 and 7). fig. 6. calculated curves describing the dynamics of the tumor cells’ death 0( )v vn v depending on the dosage on 1 = day 1, 15, 25 (1st stage of immune response) fig. 7. calculated curves describing the dynamics of the infected cells’ death 0( )n vn v depending on the dosage on 1 = day 1, 15, 25 (2nd stage of immune response) analytical study of antitumor viral vaccine 69 there are fundamental differences in the dynamics of the tumor cells’ death on each stage of the immune response. on the 1st stage the number of dying cells increases with the increase in the dosage, and depends on the tumor size at the moment of vaccine introduction (figure 6). on the 2nd stage the number of dying cells declines due to the decreasing number of produced antibodies, as shown by the 0 1( , )na v  graphs (figures 5 and 7). 4.2. kinetic curves of tumor cells’ growth after 1st and 2nd stages of immune response a model of antitumor therapy with discontinuous trajectories (equations (9)-(16)) was used to forecast the dynamics of the tumor growth after the introduction of the viral vaccines in different regimes. calculated curves of tumor growth are shown in appendix 2. the model includes a threshold value 0 cut offn − that indicates the minimal number of the tumor cells that can trigger the renewed experimental tumor growth. after each of the two stages of tumor cells’ death the tumor can resume its growth if the number of surviving cells 0 vrn and 0 nrn exceeds the threshold 0 cut offn − . threshold value depends on the normal level of plasma cells vnc and nnc (equations (5) and (7)). threshold value was defined based on proportion 0 0 0,76 0,4 2 2 cut off n n − = = = , where 0n is the initial number of tumor cells (equation (2)). if the number of surviving cells is below the threshold, then the tumor will not be able to resume its growth. in the model this is taken as complete elimination of tumor cells. in this case the lifetime of the vaccine-treated animals will be equal to the average lifetime of the experimental animals without treatment. the resulting kinetic curves that describe tumor growth after vaccine introduction on 1 = 1 day (figures 8-15) allow to identify two dosage ranges that determine two possible regimens for administering antitumor vaccines. the first regimen (with dosages exceeding exp 0 03 0,045v v=  = ) makes it possible to ensure that at the 1st stage of immune response the number of tumor cells drops below the threshold 0 cut offn − (figures 8-15), which is equivalent to elimination of tumor cells. the second regime (with dosages between exp 0 0,015v = and exp 03 0,045v = ) does not produce the decline in the number of tumor cells below the threshold level 0 cut offn − . however, these dosages cause the death of the tumor cells at both stages of immune response, and this initiates the production of antibodies against infected cells. when these antibodies exist, the process of tumor growth can be controlled through recurrent introductions of the viral vaccine. this regimen will stabilize the initial tumor size. if the initial tumor size is small, this regimen makes it possible to avoid the surgical intervention, with vaccine therapy acting as a regular vaccination against autologous tumor cells. the calculations show that if the vaccine is introduced on 1 =25 day (appendix 2) with dosages between exp 0 0,015v = to 0 0,1v = , the decrease of the number of tumor cells below the threshold level is not possible. therefore if the tumor size is big, it needs to be surgically removed. however, if the surgical intervention is contraindicated, the tumor size can be stabilized through recurrent vaccinations. in any case, this decision rests with the physician. 5. time frame for recurrent introduction of the viral vaccines a model of antitumor therapy with discontinuous trajectories of tumor growth (equations (9)(16)) was used to identify the time frame for recurrent viral vaccine introductions. the time length between the recurrent vaccine introductions t is equal to the interval between the moment of initial introduction 1 and 2 – the moment when the tumor size n( 1 ) = n( 2 ). 70 n.a. babushkina, e.a. kuzina kinetic curves of tumor growth after the viral vaccine is introduced on 1 =1 day, with dosages between exp 0 0,015v = to exp 0 02,67 0,04v v=  = , indicate that the time frame for subsequent introduction lies in the range of t = 14,5 days to t = 18,9 days depending on the dosage. therefore, if the tumor size is small, it can be stabilized through recurrent introduction of these dosages within 14,5 – 18,9 days. when the tumor is already big, its size can be stabilized while it is dominated by the fraction of the fast-proliferating cells ( * 35t = days). if the vaccine is initially introduced on 1 =25 day, stabilizing regimen can be implemented through subsequent administration of dosages that produce two-stage immune response. introduction of dosages in the range of exp 0 0,015v = to exp 0 03 0,045v v=  = within the time frame of t =30,5 can stabilize the tumor size on its pretreatment level as of day 25. this allows to avoid surgical removal of the tumor. 6. prediction of the frequency of repeated viral vaccine administration for the clinical setting the model used in this study allows to calculate the time frame for recurrent viral vaccine introductions for experimental animals. in humans, these time intervals will be different. however, it is possible to calculate their approximate duration by applying allometric relationship (16) that describes the dependence of the metabolic processes in mammals on their body weight [32, 33]: bm a =  (16) where m is the mammal's body weight,  is the speed of the mammal's metabolic processes, a and b are constant parameters. research indicates that body weight and lifespan also follow an allometric relationship for both humans and laboratory mice [32-34]. if we find the logarithm of equation (16), we obtain a linear equation that connects the average body weight of various mammals and humans with their average life span tl (figure 8): ln ln ln lnm b tl a b tl c=  + =  + (17) regression analysis allowed to define the value of regression coefficient b = 2,4981, which characterizes the slope of the regression line. figure 8. dependence of the life span on body weight for ○ mice, □ rats, + guinea pigs, х rabbits, ◊ dogs, and * humans analytical study of antitumor viral vaccine 71 allometric relationships that define the tumor growth delay for mice and humans are as follows: 0( )human human bm a t v=  and 0( )mouse mouse bm a t v=  (18) where 0( )humant v is the duration of the tumor growth remission in a human after the viral vaccine administration in clinical oncology, 0( )mouset v is duration of tumor growth delay in a mouse in the experiment. the duration of tumor growth delay for a mouse and a human can be calculated using the following proportion 1 0 0 ( ) ( ) human human b mouse mouse t v m t v m   =      this proportion makes it possible to calculate the duration of tumor remission in humans based on the estimated tumor growth delay in mice (experimental group), average body weight of mice (experimental group), and human body weight: 1 0 0( ) ( ) human b human mouse mouse m t v t v m    =      (19) for the vaccine dosage max 0 0 0,024v v= = that produces the maximal number of antibodies against infected tumor cells, the periods between recurrent introductions for a mouse and a human will be as follows: for small tumors, when the vaccine is initially administered on 1 =1 day: 0( )mouset v = 16,7 days, 0( )humant v =1,3 years; for bigger tumors, when the vaccine is initially administered on 1 =25 day: 0( )mouset v = 30,5 days, 0( )humant v =2,1 years. the period for the repeated vaccine introduction can be treated as a time when a patient needs to make a return visit to the therapist for examination. examination will indicate whether the repeated administration of the vaccine is needed. 7. conclusion the software developed in matlab-simulink based on two mathematical models describing tumor growth after the introduction of antitumoral viral vaccines makes it possible to study the effectiveness of vaccine therapy for a wide range of dosages and times of introduction. the proposed algorithm for defining the optimal regimens of viral vaccine administration can be used for various types of experimental tumors and types of antitumor viral vaccines. the results of the computing experiment presented in this paper make it possible to assess the effectiveness of different regimes of vaccine therapy. they suggest two strategies for the viral vaccine treatment. the first strategy allows a complete elimination of the tumor cells. this strategy implies a single-shot administration of a high dosage that produces the death of tumor cells at the 1st stage of immune response. the second strategy does not lead to the complete elimination of the tumor, but makes it possible to stabilize its size through recurrent introduction of the viral vaccine. if the initial tumor size is small, this stabilization gives the possibility to avoid the surgical intervention, with vaccine therapy acting as a regular vaccination against autologous tumor cells. the present study also proposes the method of using experimental results to predict the duration of remission for the humans, based on allometric relationships. this method makes it 72 n.a. babushkina, e.a. kuzina possible to define the time when a patient needs to make a return visit to the therapist for repeated examination and decision on whether further antitumor vaccine administration is required. the development of high-tech methods for the treatment of oncological diseases is directly related to the cost of experimental research. utilization of mathematical modeling at the different stages of experimental studies and on the stage of transferring the results to the clinical setting can reduce the cost of experiments and the length of research. references [1] dermime, s., armstrong, a., hawkins, r.e. & stern, p.l. (2002) cancer vaccines and immunotherapy. british medical bulletin, 62(1), 149-162. [2] mutsenietse, a. ya. (1972) onkotropizm virusov i problema viroterapii zlokachestvennyh opuholej [oncotropism of viruses and the problem of viral therapy of malignant tumors]. riga, ussr: zinatne, [in russian]. [3] mutsenietse, a. ya., bumbieris, ya. v., bruvere, r. zh. (1982) immunologija opuholej [immunology of tumors]. riga, ussr: zinatne, [in russian]. [4] moiseenko, v. m. (1999) primenenie monoklonal'nyh antitel dlja lechenija zlokachestvennyh solidnyh opuholej [the use of monoclonal antibodies for the treatment of malignant solid tumors]. voprosy onkologii, 45(4), 458-462, [in russian]. [5] gromova a. yu. (1999) protivoopuholevye svojstva vakcinnogo shtamma virusa venesujel'skogo jencefalomielita i ego onkolizata [antitumor properties of the vaccine strain of the venezuelan encephalomyelitis virus and its oncolyte]. phd thesis (biology), st. petersburg, [in russian]. [6] urazova, l. n. (2003) jeffektivnost' i mehanizmy protivoopuholevogo dejstvija virusnyh vakcin pri jeksperimental'nom onkogeneze [efficiency and mechanisms of the antitumor action of viral vaccines in experimental oncogenesis]. phd thesis (biology), st. petersburg, [in russian]. [7] vidyaeva i. g. (2005) virusnye vakciny i ih onkolizaty v terapii jeksperimental'nyh opuholej [viral vaccines and their oncolytes in the therapy of experimental tumors]. phd thesis (medicine), tomsk, [in russian]. [8] babushkina, n. a. (2005) ispol'zovanie matematicheskogo modelirovanija dlja optimizacii rezhimov himioterapii na jeksperimental'nyh opuholjah [the use of mathematical modeling for the optimization of chemotherapy schedules on experimental tumors]. proceedings of the iv int. conf. sicpro-05 "identification of systems and control tasks", мoscow, [in russian]. [9] babushkina, n. a. (2011) upravlenie processom himioterapii s ispol'zovaniem ferromagnitnyh nanochastic [control of chemotherapy process with the use of ferromagnetic particles]. problemy upravlenija [control sciences], 3, 56-63, [in russian]. [10] babushkina, n. a. (2013) ocenka upravljajushhih dozovyh vozdejstvij protivoopuholevoj vakcinoterapii s pomoshh'ju matematicheskogo modelirovanija [evaluation of controlling dosage of vaccine therapy with the use of mathematical modeling]. problemy upravlenija [control sciences], 5, 60-65. [in russian]. [11] babushkina, n. a. & glumov, v. m. (2014) matematicheskoe modelirovanie mehanizmov protivoopuholevogo dejstvija virusnyh vakcin [mathematical modeling of mechanisms of the antitumor action of viral vaccines]. proceedings of the ix int. conf. “physics analytical study of antitumor viral vaccine 73 and radioelectronics in medicine and ecology” phreme-2014. vladimir, russia, 1, 153-158, [in russian] [12] babushkina, n. a. & kuzina, e. a. (2015) komp'juternye tehnologii na osnove matematicheskogo modelirovanija v sistemnoj jeksperimental'noj onkologii [computer technologies based on mathematical modeling in the system of experimental oncology]. proceedings of the viii int. conf. “management of large-scale system development” (mlsd2015). moscow, 272-284. [in russian] [13] babushkina, n. a., glumov, v. m. & kuzina, e.a. (2016) primenenie komp'juternyh tehnologij pri jeksperimental'nom izuchenii jeffektivnosti protivoopuholevyh virusnyh vakcin [the use of computer technology in the experimental study of the efficacy of antitumoral viral vaccines]. proceedings of the xii int. conf. “physics and radioelectronics in medicine and ecology” phreme-2016. vladimir, russia, 1, 116-121. [in russian] [14] babushkina, n. a., glumov, v. m. & kuzina, e. a. (2017) primenenie matematicheskogo modelirovanija dlja ocenki jeffektivnosti metoda protivoopuholevoj terapii [application of mathematical modeling for the evaluation of the effectiveness of the antitumor therapy method]. problemy upravlenija [control sciences], 3, 49-56. [in russian] [15] loktev, v.b., ivankina t.yu., netesov s.v. & chumakov p.m. (2012) onkoliticheskie parvovirusy. novye podhody k lecheniju rakovyh zabolevanij [oncolytic parvoviruses. new approaches to the treatment of cancer diseases]. vestnik rossijskoj akademii medicinskih nauk [bulletin of the russian academy of medical sciences], 2, 42–47. [in russian]. [16] lezhnin yu.n., kravchenko yu.e., frolova e.i., chumakov p.m., chumakov s.p. (2015) onkotoksicheskie belki v protivorakovoj terapii: mehanizmy dejstvija [oncotoxic proteins in anticancer therapy: mechanisms of action]. molekuljarnaja biologija [molecular biology], 49(2), 264-278. [in russian]. [17] hristov g., krämer, m., li, j., el-andaloussi, n., mora, r., daeffler, l., zentgraf h., rommelaere j., marchini, a. (2010) through its nonstructural protein ns1, parvovirus h-1 induces apoptosis via accumulation of reactive oxygen species. journal of virology, 84(12), 5909-5922. [18] rommelaere j., geletneky k., angelova a.l., daeffler l., dinsart c., kiprianova i., schlehofer j.r., raykov z. (2010) oncolytic parvoviruses as cancer therapeutics. cytokine & growth factor reviews, 21(2), 185-195. [19] cotmore s. f. & tattersall p. (2007) parvoviral host range and cell entry mechanisms. advances in virus research, 70, 183-232. [20] moehler m. h., zeidler m., wilsberg v., cornelis j. j., woelfel t., rommelaere j., gallep. r., heike m. (2005) parvovirus h-1-induced tumor cell death enhances human immune response in vitro via increased phagocytosis, maturation, and cross-presentation by dendritic cells. human gene therapy, 16(8), 996-1005. [21] grekova s. p., aprahamian m., daeffler l., leuchs b., angelova a., giese t., galabov a., helle ra., giese n. a., rommelaere j., raykov z. (2011) interferon γ improves the vaccination potential of oncolytic parvovirus h-1pv for the treatment of peritoneal carcinomatosis in pancreatic cancer. cancer biology & therapy, 12(10), 888–895. [22] raykov z., grekova s., galabov a.s., balboni g., koch u., aprahamian m., rommelaere j. (2007) combined oncolytic and vaccination activities of parvovirus h-1 in a metastatic tumor model. oncology reports, 17(6), 1493-1500. [23] angelova a. l.,aprahamian m., balboni g.,delecluse h. j., feederle r., kiprianova i., grekova s., galabov a., witzens-harig m., ho a. d., rommelaere, j., raykov z. (2009) 74 n.a. babushkina, e.a. kuzina oncolytic rat parvovirus h-1pv, a candidate for the treatment of human lymphoma: in vitro and in vivo studies. molecular therapy, 17(7), 1164–1172. [24] marchuk g.i. (1991) matematicheskie modeli v immunologii. vychislitel'nye metody i jeksperimenty [mathematical models in immunology. computational methods and experiments]. moscow: nauka, [in russian] [25] romanyukha a.a. (2011) matematicheskie modeli v immunologii i jepidemiologii infekcionnyh zabolevanij [mathematical models in immunology and epidemiology of infectious diseases]. moscow: binom. laboratorija znanij, [in russian] [26] bolodurina i.p., lugovskova yu.p. (2009) optimal'noe upravlenie immunologicheskimi reakcijami organizma cheloveka [optimum management of immunological responses of the human body]. problemy upravlenija [control sciences], 5, 44–52, [in russian] [27] rusakov s.v., chirkov m.v. (2012) matematicheskaja model' vlijanija immunoterapii na dinamiku immunnogo otveta [the mathematical model of the effect of immunotherapy on the immune response dynamics]. problemy upravlenija [control sciences], 6, 45–50, [in russian] [28] kogan, y., halevi–tobias, k., elishmereni, m., vuk-pavlović, s., & agur, z. (2012) reconsidering the paradigm of cancer immunotherapy by computationally aided real-time personalization. cancer research, 72(9), 2218 2227. [29] de pillis, l. g., radunskaya, a. e., & wiseman, c. l. (2005) a validated mathematical model of cell-mediated immune response to tumor growth. cancer research, 65(17), 7950-7958. [30] palladini, a., nicoletti, g., pappalardo, f., murgo, a., grosso, v., stivani, v., ianzano, m.l., antognoli, a., croci, s., landuzzi, l., de giovanni, c., nanni, p., motta, s., lollini, p.-l. (2010) in silico modeling and in vivo efficacy of cancer-preventive vaccinations. cancer research, 70(20), 7755-7763. [31] skipper, h.e. (1971) kinetics of mammary tumor cell growth and implications for therapy. cancer, 28(6), 1479-1499. [32] monichev a.y. (1984) dinamika krovetvorenija [dynamics of hematopoiesis]. moscow: medizina, [in russian] [33] monichev, a. y. (1987) a mathematical model of the spatial structure of bone marrow in hemopoietic dynamics. cybernetics and systems analysis, 23(2), 274-280. [34] kuzina e.a., babushkina n.a. (2016) programmnaja realizacija metoda prognozirovanija jeffektivnosti vakcinoterapii ot jeksperimenta v kliniku [software implementation of the method for predicting the efficacy of vaccine therapy from experiment to the clinic]. proceedings of the xii int. conf. “physics and radioelectronics in medicine and ecology” phreme-2016. vladimir, russia, 1, 125-130, [in russian] analytical study of antitumor viral vaccine 75 appendix 1. values of the models' parameters after a single introduction of the viral vaccine on the first day of the tumor growth equation parameter description ( )dn t dt – (9) n = 3.3613, n = 0.0332 n= 23 0n = 0.75 parameters of the gompertz function approximating the experimental growth curves of the tumor cell population without the vaccine administration (control sample) ( )vdn t dt – (1) 1 = 1 moment of the first introduction of the viral vaccine cvz = 4.5 time lag of the immune response against the virus cnz = 10.5 time lag of the immune response against the infected tumor cells 1v cvt z= + starting moment of the immune response against the virus 1n cnt z= + starting moment of the immune response against infected tumor cells ( )dv t dt – (3) v = 0.1 rate of virus replication v = 15 death rate of viruses due to their interaction with antibodies 0v = 0.015 dosage of the viral vaccine – the initial condition of equation (2) ( )vda t dt – (4) a = 100 rate of antibodies’ formation from a single plasma cell av = 70 rate of decrease in the number of antibodies due to interaction with viruses a = 5 rate of decrease in the number of antibodies due to natural destruction max va = 1.05 maximum calculated number of antibodies ( )vdc t dt – (5) c = 100 rate of antibodies’ formation from a single plasma cell cv = 4.5 dimension factor vnc = 0.001 initial number of plasma cells ( )nda t dt – (6) an = 30 rate of antibodies’ formation from a single plasma cell an = 6.2 rate of decrease in the number of antibodies due to interaction with tumor cells nn = 6.3 rate of decrease in the number of antibodies due to 76 n.a. babushkina, e.a. kuzina natural destruction max na = 4.6445 maximum calculated number of antibodies ( )ndc t dt – (7) cn = 76.677 rate of formation of plasma cells cn = 38 dimension factor nnc = 0.0001 initial number of plasma cells ( )vdn t dt – (1) vk = 0.25 avk = 0.8 ank = 0.8 constant coefficients of the dynamic equation for infected cells after a single introduction of the vaccine ( )p t – (8) p = 0.3 pk = 0.95 p = 1/t* parameters of function p(t), describing the dynamic equation for reduction of the fast-proliferating fraction of tumor cells t*= 35 days moment when the number of fractions of rapidly proliferating cells is equal to the number of fractions of slowly proliferating cells ( )dn t dt –(10 ) v = 4.3 days length of the tumor cells’ growth delay after their death at 1st stage of immune system stimulation cn = 9.8 days length of the tumor cells’ growth delay after their death at 2nd stage of immune system stimulation sum = 14.1 days total length of the tumor cells’ growth delay (1st and 2nd stage of immune system stimulation) 1 vt = 5.82 1 nt = 12.32 starting moments of the 1st and the 2nd stage of immune response 2 vt =7.13 2 nt =14.288 ending moments of the 1st and the 2nd stages of immune response 1( )v v vn t =1.3125 1( )n n vn t =1.1458 maximal numbers of infected tumor cells before the start of the 1st and the 2nd stages of immune response 2( )v v vn t = 0.84 2( )n n vn t =0.1031 minimal numbers of infected tumor cells before the end of the 1st and the 2nd stages of immune response 1( )v v vn t = 0.472, 1( )n n vn t = 1.043 number of dead infected cells at the 1st and the 2nd stages of immune response analytical study of antitumor viral vaccine 77 appendix 2. kinetic curves of tumor growth kinetic curves are calculated for varying dosages on two moments of introduction: 1 = day 1 (figures (a) – (b)) and 1 = day 25 (figures (c) – (d)). for each dosage, figure (a) shows the dynamics of tumor growth ( )n t after viral vaccine administration (dotted line denotes the dynamics of the number of infected tumor cells ( )vn t ), figure (b) shows the dynamics of the number of antibodies ( )va t against the virus (red line), and ( )na t against infected tumor cells (black line) at two stages of the immune response. dosage 0v = 0,015 а) c) b) d) 78 n.a. babushkina, e.a. kuzina dosage 0v = 0,024 а) c) b) d) analytical study of antitumor viral vaccine 79 dosage 0v = 0,03 а) c) b) d) 80 n.a. babushkina, e.a. kuzina dosage 0v = 0,04 а) c) b) d) analytical study of antitumor viral vaccine 81 dosage 0v = 0,045 а) c) b) d) 82 n.a. babushkina, e.a. kuzina dosage 0v = 0,06 а) c) b) d) analytical study of antitumor viral vaccine 83 dosage 0v = 0,075 а) c) b) d) 84 n.a. babushkina, e.a. kuzina dosage 0v = 0,1 а) c) b) d) advances in systems science and applications (2013) vol.13 no.1 80-99 an extension of a logistic model for microbial kinetics anne talkington1, floyd inman iii2, leonard d. holmes2 and guo wei2 1duke university, durham, nc 27708, usa 2university of north carolina at pembroke, pembroke, nc 28372, usa abstract in contrast to the traditional logistic model, a series of asymmetrical models have been proposed for modeling bacterial growth. these models are similar to the logistic model for the lag-phase and exponential phase of the population growth, but quite different in the stationary phase − the growth becomes remarkably slower after the inflection point. at this point, limiting factors within the population such as competition or environmental stress inhibit further exponential growth. moreover, models as variations of the traditional exponential model are also proposed. the models presented demonstrate more general patterns and representative properties, from which relevant algorithms can be developed for calculating the population specific growth rate occurring in the exponential phase. keywords maclaurin series, logistic growth, exponential growth, point of inflection, carrying capacity, specific growth rate 1 introduction the study of microbial kinetics is of particular significance to research in the fields of microbiology and biotechnology [1]. although existing models and methods have been developed, an issue arises over the precision of potentially subjective, traditional methods of culture analysis. calculating population growth rate is essential to microbial kinetic studies of substrate-culture interactions, and the introduction of a new model for population growth increases the precision of the method. pinpointing this microbial specific growth rate, represented as µ or r, with accuracy and precision is at the heart of consistency in understanding laboratory results. furthermore, extending calculations to model the data uses growth patterns to provide insight into characteristics of the microbe. an exploration of four models incorporates a review and extension of the traditional exponential and logistic growth function, and introduces the potential for exponential-like and logistic-like functions using a modified form of the exponential maclaurin series. several models have been fit for microbial growth curves, modifications of the idealized exponential (malthusian) and logistic (limited) growth to fit the reality of microbial patterns. among growth models, competition models, and nutrient uptake models are included theta logistic functions, trans-theta logistic functions, time-delay logistic functions, general logistic functions. michaelis-menten kinetics, gompertz, von bertalanffy and general von bertalanffy, verhlust, and advances in systems science and applications (2013) vol.13 no.1 81 lotka-volterra are specific cases of these models [2-11]. exponential growth, or unlimited and ever-increasing growth, is represented by the general differential form dp dt = rp , and its solution p (t) = p0e rt[8]. logistic growth is sigmoidal and limited. it is represented by the general differential form dp dt = rp (1 − p m ), and its solution p (t) = mp0ert m+p0(ert−1) [8]. previous extensions to the logistic model involve the general logistic model and the time-sensitive carrying capacity logistic model [8-11]. the general logistic model takes the differential form dp dt = rpα[1 − ( p m )ν ]γ , which is simplified to p (t) = m [1+[(γ−1)βrm(α−1)t+[( m p0 )β−1](1−γ)]1/(1−γ)]1/β or p (t) = a + m−a (1+qe−r(t−t1))1/ν , (t1 is defined as time of maximum growth, a is the value of the lower asymptote, q is an initial condition constant, ν is a constant determining the point of inflection, and α and γ represent constant inflection value parameters) [8]. this introduces new parameters so that the curve can idealize scenarios previously considered “less ideal” by existing models. the time-varying model introduces the upper bound itself as a function of time, and it therefore cannot be definitely determined as an analytical solution. this model further complicates the function with the introduction of a variable parameter. it is represented as dp dt = rp (1− p m(t))[11]. the introduction of a series incorporates the concept of a traditional model modified through extension. this new function is explored as a differential, and characterized by its non-integrable form. the exponential and logistic models represent the first two terms of a maclaurin series. previously, the remaining terms were not considered significant, and the use of a series was disregarded [8]. however, discovering and incorporating the terms as an expanded series removes inconsistencies between methods of calculation for parameters, such as rate of increase. the convergence of the series represents an ideal model, that can be characterized by the properties of several functions, and to which these functions converge. the first terms reflect its most basic exponential and logistic properties. as it continues to expand and the series converges, truncation produces a series of polynomial functions with exponential and logistic-like (sigmoidal) growth patterns. the final closed model itself appears to grow infinitely, like an exponential function, but at a slower rate, like a hyperbolic pattern. indeed, elements of the series resemble the taylor series expansions for the hyperbolic sine (sinh) and cosine (cosh) functions, and the geometric series hyperbolic expansion, in addition to the inherently exponential elements. the potentially hyperbolic properties connect this model to the monod and michaelis-menten models [2], while they provide a basis for parametric data analysis through the function. however, it is not a distinct hyperbola because the values of dp dt approach infinity as p approaches infinity, without any distinct limit or asymptote. this new series relies on the parameters of growth limit, rate of increase, and 82 anne talkington: an extension of a logistic model for microbial kinetics population size related through time. it is unique in that it does not introduce new factors or parameters with each additional term, but builds upon known parameters. manipulation of the series model over its interval of convergence opens applications in both symmetrical and asymmetrical (semi-logistic) growth patterns. analysis of the curve to the upper bound of its convergence (below and including the inflection point) produces a precise estimate for the initial takeoff of the microbe; the rotation of the existing curve and intersection of another curve are two approaches for determining the upper bound. evaluating the series as a whole, or as an individual polynomial function at any point of truncation, allows for the ability to fit both limited and unlimited growth patterns. the models, likewise, are consistent with numerical (nonparametric) methods for determining microbial specific growth rate [12-13]. numerical methods analyze the data without assumption, whereas the models provide a framework for the data. each introduces a degree of precision and understanding of the data. 2 procedure to determine the value of the specific growth rate numerically, an algorithm was developed to analyze the data through elimination. it removes all subjectivity and assumptions found in previous methods. the new algorithm focuses on eliminating non-significant data points. initially, all points past the point of inflection are eliminated. this is defined as the point at which the concavity of the data changes from up to down, or the second derivative changes from positive to negative. thus, the deceleration phase is disregarded in analyzing growth. the next step is to isolate the growth phase from the lag and acceleration phases. further elimination is accomplished through regression. the natural logarithm of the function is taken, and the slope of the linearized function is determined from the line of best fit. points from the left that lower the slope are eliminated, and another regression line is fitted to the remaining points. the process is repeated as the slope increases and iterations continue until either the change in slope (indicative of growth rate) is not statistically significant, or until the number of remaining points becomes too small. the slope of the regression line for the final remaining points is the specific growth rate. 3 cases of the series discussion there are four notable models that serve appropriate data set fit. fitting an appropriate model to the data set is critical for the most accurate approximations of data parameters and descriptions of the populations. the most basic is the exponential model [8, 14-15]. also known as the malthusian model, it depicts unlimited growth. biologically, there is no carrying capacity of the population, and it therefore multiplies infinitely. mathematically, there is advances in systems science and applications (2013) vol.13 no.1 83 no asymptote to describe limiting factors. these may include competition, lack of nutrients, or density-dependent inhibition in the microbial population [16]. the model is represented as a direct proportion between the rate of population change and population size (equation 1, fig.1). dp dt = rp (1) fig.1 graphical representation of the exponential model for microbial population growth of course, the simplest model given above does not reflect most biological growth systems. more appropriate models were thus explored by researcher, as stated below. microbial growth may be logistic [17]. this model is a special case of limited growth in which rate of population decrease mirrors rate of increase. the population grows exponentially but encounters one or more obstacles that prohibit it from increasing indefinitely. as the definite limit m approaches infinity, growth patterns approach the exponential model. logistic growth, often used as a standard in population studies, is represented by the differential equation 2 (see also fig.2). dp dt = rp (1− p m ) (2) the logistic growth model introduces the first two terms of a maclaurin series that can be used to model microbial growth. limiting the series to two terms introduces truncation error [14-15]. this error can be reduced with the introduction of additional terms. successive terms alternate in sign, and each is smaller in value than the term it follows [14-15]. extending the series another step produces a differential function to the fourth power of p . this function is similar to logistic growth in the lower portion of the curve, up to the point of 84 anne talkington: an extension of a logistic model for microbial kinetics fig.2 graphical representation of the logistic growth model inflection. however, the point of inflection is no longer the midpoint between the upper and lower limits of growth. beyond the point of inflection, growth reduces rapidly and tapers until it reaches its actual limit. the population may have been subjected to extreme pressure over a short period of time, so realized growth does not reach the theoretical, ideal estimate of m . this reduced upper bound is characteristic of higher even degrees of p in the extension of the formula due to factorial growth in the denominator. the end behavior of odd degrees of p mimics exponential-like growth. the fourth-degree formula is represented as equation 3, and its graphical interpretation as fig.3. dp dt = rp − 2rp 2 m(2!) + 8rp 3 m2(3!) − 48rp 4 m3(4!) (3) fig.3 graphical representation of the fourth-degree model the most idealized model is the use of the full infinite series, represented in advances in systems science and applications (2013) vol.13 no.1 85 its closed form for precision. it states that the rate of change of the population and a function of the population size are directly proportional. the function of population size is the maclaurin series. the function, or the series, incorporates the exponential model when m (the upper limit) approaches infinity. it incorporates the logistic model as the ideal symmetrical case in which the point of inflection is the midpoint between the upper and lower limits of growth. if the lower limit is 0, p represents population size at the point of inflection, and m represents the upper growth limit, then this is represented as 2p = m . all of the previous models converge, as a series of overestimates and underestimates, to the model presented by the complete series. it is an exponential-like function, but its predictions for population (p (t)) lie between the exponential and logistic predictions. the modified alternating maclaurin series is convergent for all points up to and including the point of inflection. it satisfies the conditions of the ratio test and leibniz’s theorem for this domain and range [14-15]. past the point of inflection, divergence becomes more extreme with additional terms. as a model, therefore, it accurately produces the lower portion of the curve. if the growth pattern is symmetrical, a 180o rotation about the point of inflection projects the upper portion. if the growth is semi-logistic, it appears as portions of intersecting logistic curves, and the point of intersection is the point of inflection. in this scenario, the bottom portion of the curve is accurately described by the model (equation 4, fig.4). fig.4 graphical depiction of the modified alternating maclaurin series model dp dt = ∞∑ n=0 (−1)n(2nn!)rpn+1 mn(n+ 1)! (4) its expanded form appears as equation 5. 86 anne talkington: an extension of a logistic model for microbial kinetics dp dt =rp − 2rp 2 m(2!) + 8rp 3 m2(3!) − 48rp 4 m3(4!) + 384rp 5 m4(5!) − 3840rp 6 m5(6!) + ...+ (−1)n(2nn!)rpn+1 mn(n+ 1)! (5) its simplified form appears as equation 6. dp dt = rp ∞∑ n=0 (−1)n (n+ 1) × ( 2p m )n (6) its closed form appears as equation 7. dp dt = rln(1 + p m 2 ) m 2 (7) the similarities of the function to the exponential function are evident in the simplified form of the model. the factor rp represents exponential growth. in addition, the behavior of the series is similar to that of exponential growth for sufficiently small values of p (p << 1). dp dt = rln(1 + p m 2 ) m 2 dp dt = r × m 2 × ln(1 + 2p m ) ln(1 + 2p m ) ≈ 2p m dp dt = r × m 2 × 2p m = rp this property does not apply when 2p m is greater than 1, because ln(1 + x) = x− 1 2x 2 + 1 3x 3 − 1 4x 4 + ... holds only for −1 < x ≤ 1. also like an exponential function, the graph of the population p over time is concave up, as justified by the always positive second derivative: d2p dt2 = r m 2 1 1 + p m 2 (1 + p m 2 ) ′ t = r m 2 1 1 + p m 2 1 m 2 dp dt = r2 m 2 1 1 + p m 2 ln(1 + p m 2 ) > 0 however, the rate of growth as p approaches infinity is much slower than that of an exponential function since by limit comparison: lim p→∞ rp r × m 2 × ln(1 + p m ) = lim p→∞ 2p m × ln(1 + p m ) = ∞ advances in systems science and applications (2013) vol.13 no.1 87 it is known that the numerator becomes infinitely large at a faster rate than the denominator. l′hôpital′s rule confirms this assertion by removing the infinite term from the function, to arrive at the definite conclusion: lim p→∞ rp r × m 2 × ln(1 + p m ) = lim p→∞ (rp )′ lim p→∞ (r × m 2 × ln(1 + p m ))′ = r lim p→∞ (r × m 2 × 1 m 1+ p m ) = ∞ this property is the result of the alternating series succession of positive and negative terms. the positive and negative terms indicate the properties of the series at any point of truncation. if truncated up to an even degree of p (say up to pn+1 where n is odd), then the solution p has an upper bound, and an inflection point (see theorem 2). the inflection point will always occur when p = m 2 as justified by the second derivative of p (t) as follows. dp dt = n∑ n=0 (−1)n2nrpn+1 mn(n+ 1) (8) d2p dt2 = n∑ n=0 (−1)n2nr(n+ 1)pn mn(n+ 1) ( dp dt ) = r( dp dt ) n∑ n=0 (−1)n( 2p m )n (9) in the case 2p = m , we have d2p dt2 = r( dp dt ) n∑ n=0 (−1)n hence, when n is odd, it holds that d2p dt2 |p=m 2 = 0 it follows from the chain rule of derivatives [14-15]: d dt ( dp dt ) = d dp ( dp dt )× dp dt that the derivative of dp dt with respect to t is equal to the derivative of dp dt with respect to p . the value for dp dt decreases at it nears its upper bound but never becomes 0. it is always positive (theorem 3). continued simplification of the modified alternating maclaurin series and its truncations, as shown below, reveals its properties and its relationship to the form of a transcritical bifurcation. 88 anne talkington: an extension of a logistic model for microbial kinetics 4 simplified series dp dt = rp ∞∑ n=0 (−1)n n+ 1 ( 2p m )n = rp [1− 1 2 ( 2p m ) + 1 3 ( 2p m )2 − 1 4 ( 2p m )3 + 1 5 ( 2p m )4 − 1 6 ( 2p m )5 + ...] (10) theorem 1 for equation (10), (i) the closed form of the series on the right-hand side is dp dt = rln(1 + p m 2 ) m 2 ; (ii) it holds that dp dt > 0 at any time t. proof (i) equation (10) is simplified to dp dt = r ∑∞ n=0 (−1)n2npn+1 mn(n+1) (by cancelling n!). then we have dp dt = r ∞∑ n=0 (−1)n2npn+1 mn+1(n+ 1) = rp ∞∑ n=0 (−1)n (2pm )n n+ 1 = rp ∞∑ n=0 (−2p m )n n+ 1 = rp [1 + x 2 + x2 3 + x3 4 + ...+ xn n+ 1 + ...](for x = −2p m ) = rp x [x+ x2 2 + x3 3 + x4 4 + ...+ xn+1 n+ 1 + ...] = rp x (−1)ln(1− x) = r m 2 ln(1 + 2p m ) by the taylor expansion of natural logarithm: ln(1 − x) = ∑∞ n=0 xn n for −1 ≤ x < 1. hence, we have r = dp dt ln(1+ p m 2 ) m 2 , and this holds whenever −1 ≤ −2p m < 1, i.e., p ≤ m 2 . thus, the closed form of the series is given by dp dt = rln(1 + p m 2 ) m 2 . (11) (ii) this is implied by the closed form given in (i). theorem 2: for the following tail-truncated equation dp dt = rp n∑ n=0 (−1)n n+ 1 ( 2p m )n (12) advances in systems science and applications (2013) vol.13 no.1 89 (i) if n is odd (so the highest power of p is even), then d2p dt2 is positive when p < m 2 , is zero when p = m 2 , and is negative when p > m 2 . this is an asymmetrical model. (ii) if n is even (so the highest power of p is odd), then d2p dt2 is always positive. since the second derivative is always positive, the first derivative is always increasing. therefore, if the first derivative is positive initially, the first derivative is never zero or negative. (the least upper bound of p is finite and determined as described in theorem 3.) proof it has been established above that in the case 2p = m d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt )[1− 1 + 1− 1 + ...] for an even number of terms, or odd n : case 2p = m . the truncation of the above equation results in an even pairing of terms that cancel out, and the derivative is equal to 0. case 2p < m . d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt )[1−( 2p m )+( 2p m )2−( 2p m )3+( 2p m )4−( 2p m )5+ ...] a grouping of the first 6 terms of the second derivative results in d2p dt2 = r( dp dt ) [ [1− ( 2p m )] + [( 2p m )2 − ( 2p m )3] + [( 2p m )4 − ( 2p m )5] ] > 0. case 2p > m . d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt ) [ 1−( 2p m )+( 2p m )2−( 2p m )3+( 2p m )4−( 2p m )5+... ] a grouping of the first 6 terms results in[ [1− ( 2p m )]+ [( 2p m )2− ( 2p m )3]+ [( 2p m )4− ( 2p m )5] ] < 0. (each bracket is negative.) for an odd number of terms, or even n : the truncation of d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt )[1− 1 + 1− 1 + 1− 1 + ...] 90 anne talkington: an extension of a logistic model for microbial kinetics results in a pairing with an additional positive term that is not cancelled. so this second derivative is positive. case 2p < m . d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt ) [ 1−( 2p m )+( 2p m )2−( 2p m )3+( 2p m )4−( 2p m )5+... ] a grouping of the first 5 terms of the second derivative results in[ [1− ( 2p m )] + [( 2p m )2 − ( 2p m )3] + ( 2p m )4] > 0. case 2p > m . d2p dt2 = r( dp dt ) ∞∑ n=0 (−1)n = r( dp dt ) [ 1−( 2p m )+( 2p m )2−( 2p m )3+( 2p m )4−( 2p m )5+... ] a grouping of the first 5 terms results in[ 1 + [−( 2p m ) + ( 2p m )2] + [−( 2p m )3 + ( 2p m )4] ] > 0. theorem 3: for the above tail-truncated equation (12), (i) dp dt is always positive (i.e., the right-hand side is always positive). (ii) when n is odd (i.e., including up to an even power of p in the partial sum), the equation rp ∑n n=0 (−1)n n+1 (2pm )n = 0 has a unique positive solution and the supremum value of p is determined by this positive zero of the equation. moreover, it holds that rp ∑n n=0 (−1)n n+1 (2pm )n = rm 2 ∫ 2p m 0 1−xn+1 1+x dx. these models are all asymmetrical. proof (i) if n is even, this is clearly true. let gn (x) = n∑ n=0 (−1)n n+ 1 xn. then dp dt = rp n∑ n=0 (−1)n n+ 1 ( 2p m )n = rpgn ( 2p m ). since xgn (x) = ∑n n=0 (−1)n n+1 xn+1, we have (xgn (x)) ′ = ∑n n=0(−x)n = 1−(−x)n+1 1+x , and hence it holds that[ xgn (x) ] | 2p m 0 = ∫ 2p m 0 ( xgn (x) )′ dx = ∫ 2p m 0 1− (−x)n+1 1 + x dx, advances in systems science and applications (2013) vol.13 no.1 91 implying 2p m gn (2p m ) − 0 =  ∫ 2p m 0 1−xn+1 1+x dx, if n + 1 even,∫ 2p m 0 1+xn+1 1+x dx, if n + 1 odd. therefore, dp dt = rpgn (2p m ) = rp×m 2p ×2p m gn (2p m ) = rm 2  ∫ 2p m 0 1−xn+1 1+x dx, if n + 1 even,∫ 2p m 0 1+xn+1 1+x dx, if n + 1 odd. it is obvious that dp dt = rm 2 ∫ 2p m 0 1+xn+1 1+x dx > 0 when n is even (n + 1 is odd). next we consider the case that n is odd (or n + 1 is even). it follows from the above proof that dp dt = rm 2 ∫ 2p m 0 1−xn+1 1+x dx. when p ≤ m 2 (2pm ≤ 1), dp dt is positive because 1 − xn+1 is always positive except at the single point x = 2p m , at which 1 − xn+1 is zero. we claim that dp dt never becomes 0. in fact, if dp dt becomes 0 at some infimum time t0, it must strictly be after the time at which the inflection point occurs. (at this point, dp dt changes from positive with upward concavity and begins to decrease. the value of the integral, likewise, begins to decrease after the shift in dp dt at this point, corresponding to the concavity, but the value still must be positive. it follows that dp dt cannot be 0 at the inflection point.) since d2p dt2 is still negative at t0, p (t) by definition would have attained its local maximum p (t0) (at a specific p (t0) > m 2 ). now, let t1 be a nearby point of t0, on the right of t0 such that m 2 < p (t1) < p (t0). (since dp dt |t=t0 = 0 and d2p dt2 < 0 after t = t0, dp dt is negative after t = t0). then the result would become 0 = dp dt |t=t0 = rm 2 ∫ 2p (t0) m 0 1−xn+1 1+x dx < rm 2 ∫ 2p (t1) m 0 1−xn+1 1+x dx = dp dt |t=t1 , a contradiction. because the value of p (t1) is less than the value of p (t0), the inequality is inconsistent with the fact that dp dt at t1 must be negative. hence, dp dt never reaches 0, and it follows from the continuity of dp dt that dp dt can never be negative. therefore, dp dt is always positive. proof of (ii) method for determining the supremum of p : let us illustrate the method using two (logistic), four, or six terms. if n = 1, n+1 = 2 (the traditional logistic model), then dp dt = ... = rp (1− p m ). hence, dp dt > 0 if and only if p < m . observe that if p could exceed m , then dp dt would become negative and consequently p would be immediately lowered making dp dt positive. therefore, dp dt is always positive. again, continuity indicates that dp dt must reach and cross through a point of 0 slope to become negative, which is impossible. 92 anne talkington: an extension of a logistic model for microbial kinetics if n = 3, n + 1 = 4 (four terms), then dp dt = rm 2 ∫ 2p m 0 1−x4 1+x dx = rp (1− 1 2 2p m + 1 3( 2p m )2− 1 4( 2p m )3). the equation 1− 1 2x+ 1 3x 2− 1 4x 3 = 0 has a real solution 1.621. for example, this gives the maximum of p as 81.05, assuming m = 100. if n = 5, n + 1 = 6 (six terms), then dp dt = rm 2 ∫ 2p m 0 1−x6 1+x dx = rp (1− 1 2 2p m + 1 3( 2p m )2− 1 4( 2p m )3+ 1 5( 2p m )4− 1 6( 2p m )5). the equation 1− 1 2x+ 1 3x 2− 1 4x 3+ 1 5x 4− 1 6x 5 = 0 has a real solution 1.462. this gives the maximum of p as 73.1, assuming m = 100. generally, for any odd n , dp dt is always positive, and the supremum value of p is determined by a positive real solution (that exists uniquely) of the equation dp dt = rm 2 ∫ 2p m 0 1−xn+1 1+x dx = rp ∑n n=0 (−1)n n+1 (2pm )n. the supremum values decrease as n is increased. rp ∑n n=0 (−1)n n+1 (2pm )n = rm 2 ∫ 2p m 0 1−xn+1 1+x dx cannot be zero when p ≤ m 2 ; the equation ∑n n=0 (−1)n n+1 xn = 0 does not have a solution less than 1 (hence divergence as the partial series extends to infinite form). ∫ α 0 1−xn+1 1+x dx > 0 if α ≤ 1 and ∫ β 0 1−xn+1 1+x dx < 0 if β is (sufficiently) large. there exists γ between α and β such that ∫ γ 0 1−xn+1 1+x dx = 0; ∫ α 0 1−xn+1 1+x dx > ∫ β 0 1−xn+1 1+x dx for α ≤ 1 < β. more precisely, an arbitrary α beyond the point of inflection presents a scenario that is consistent with the definition of a maximum point and downward concavity in this domain. hence, the equation n∑ n=0 (−1)n n+ 1 xn = 0 (13) has a unique positive solution x0. the formula for calculating the supremum of p is psup = x0p = x0 m 2 . in summary, each term is smaller than the term preceding it until p is equal to the least upper bound. at p = m 2 , dp dt has reached its maximum value and will decrease but will still be positive, so the negative terms cannot be larger than the positive terms before them. the slopes remain positive and decrease as dp dt approaches 0 at the least upper bound without ever reaching it. therefore, dp dt will always be positive despite negative concavity (reduction in value) past the point of inflection, and the limit of dp dt as p approaches a set of arbitrary values reveals an approximation of the least upper bound of the function. if the limit reveals that dp dt is less than 0 at a particular value of p , then the least upper bound is below that value; the value can never be reached. at the upper limit, a horizontal asymptote, dp dt is 0. therefore, dp dt approaches this value at the transition from a small positive value to a small negative value. when limit comparison reveals that one value results in a positive dp dt and the next value for p results in a negative dp dt , the bound must lie between these two points. in the advances in systems science and applications (2013) vol.13 no.1 93 logistic case, this limit is m . by contrast, a function with an odd degree of p resembles the properties of an exponential function. it is concave up as well as always increasing; the sum of the alternating terms is positive, plus the addition of a positive term. it has neither a point of inflection nor an upper bound. in terms of the second derivative defined above, the incomplete pairings ensure positive concavity (theorem 2). the series and its truncations can be analyzed as a transcritical bifurcation [18]. the concept of “harvesting”, initially applied to logistic growth, remains. for even degrees of p , the model exhibits logistic-like properties and therefore the potential for two positive zeros. in logistic growth (power of p is 2), zeros are found as the solutions to a simplified, basic quadratic form of the equation. in the general situation (power of p is even and larger than or equal to 4), by the integral representation given in theorem 3 (ii), dp dt = 0 has two positive zeros, one positive zero, or no positive (real) zeros. such zeros in the forms of the series can be determined by utilizing numerical methods. as the degree increases, the amount of symmetry as found in logistic models decreases. as the number of terms in the polynomial increases, and approaches the closed form of the series, the equation becomes increasingly complex to solve. in this parametric approach, fixed points are determined as the variable values that set the equation at equilibrium. a harvesting term introduces a parameter that affects the location of such points, and their position for approaching or separating from one another as the differential increases or decreases. the equation for transcritical bifurcation is as follows (equation 14): dp dt = rp ∞∑ n=0 (−1)n n+ 1 × ( 2p m )n −h (14) because the model is idealized, it can be used as a method to solve for rate of growth. such calculations are consistent with the numerical estimates. like theta logistic models, this model introduces precision through specific terms added to the traditional logistic model. it is unique in that it achieves this precision without the introduction of new parameters and accounts for lack of symmetry (semi-logistic scenarios) through the model itself rather than a term in the formula. this form of the series is appropriate to solve for growth rate of any indicator (equation 15): µ = r = dp dt∑∞ n=0 (−1)n(n!2n)pn+1 mn(n+1)! (15) 94 anne talkington: an extension of a logistic model for microbial kinetics closed form of the series (equation 16): r = dp dt m 2 ln(1 + 2p m ) . (16) consider that: ln(1 + 2p m ) = ln(2) when m = 2p . the modified alternating maclaurin series accurately models the data as it converges, over the domain of initial data collection to the point of inflection p . the model converges with the condition that p ≤ m 2 . the applicability of the maclaurin series as a method for determining growth rate is categorized as one of three cases, based on the relationship between p and m . the following procedure applies for the first case of the modified alternating maclaurin series: if data is symmetrical (ideally logistic) and p = m 2 . 1) the y-coordinate of the point of inflection (value substituted for p ) is the midpoint between the upper and lower limits of growth (asymptotes). it is calculated as lower limit + (1/2) (upper limit lower limit). 2) dp dt represents change in population (growth) over time. it should reach its maximum value at the point of inflection. dp dt should therefore reflect the change at the point of inflection as a tangent line. it is determined more precisely as the mean value of the rates of change surrounding the point of inflection. for even greater precision, arbitrary ranges chosen to include the point of inflection can be averaged and compared; the greatest value among these is taken as dp dt . 3) m is the upper limit of growth. it is identified as a horizontal asymptote. 4) values for p , dp dt and m are substituted into the series to solve for r (µ). example of the first case : fig.5 graphical traditional method: µ=1.0374 advances in systems science and applications (2013) vol.13 no.1 95 graphical traditional method: µ=1.0374 maclaurin series method: p = 3; dpdt = 2.2;m = 6 µ = 1.0580 error (disparity between the methods): 1.99 % the following procedure applies for the second case of the modified alternating maclaurin series: if the point of inflection is less than half of the carrying capacity, or p < m 2 . 1) the y-coordinate of the point of inflection (value substituted for p ) is less than the midpoint between the upper and lower limits of growth (asymptotes). it is calculated as the point in the center of the range of maximum change in population over time (largest values of dp dt ). when averages are taken for dp dt , p should be at or near the center of the range. 2) dp dt represents change in population (growth) over time. it should reach its maximum value at the point of inflection. dp dt should therefore reflect the change at the point of inflection as a tangent line. it is determined more precisely as the mean value of the rates of change surrounding the point of inflection. for even greater precision, arbitrary ranges chosen to include the point of inflection can be averaged and compared; the greatest value among these is taken as dp dt . 3) m is the upper limit of growth. it is identified as a horizontal asymptote. 4) values for p , dp dt , and m are substituted into the series to solve for r (µ). 5) because p is low with respect to m , convergence of the series occurs more rapidly than in the symmetrical first case. the result is more precise at a smaller number of terms. example of the second case : fig.6 graphical traditional method: µ=0.2817 graphical traditional method: µ=0.2817 96 anne talkington: an extension of a logistic model for microbial kinetics maclaurin series method: p = 0.15; dpdt = 0.04;m = 1.3 µ = 0.2964 error (disparity between the methods): 5.22 % the following procedure applies for the third case of the modified alternating maclaurin series: if the point of inflection is greater than half of the carrying capacity, or p > m 2 . 1) the y-coordinate of the point of inflection (value substituted for p ) is greater than the midpoint between the upper and lower limits of growth (asymptotes). it is calculated as the point in the center of the range of maximum change in population over time (largest values of dp dt ). when averages are taken for dp dt , p should be at or near the center of the range. 2) dp dt represents change in population (growth) over time. it should reach its maximum value at the point of inflection. dp dt should therefore reflect the change at the point of inflection as a tangent line. it is determined more precisely as the mean value of the rates of change surrounding the point of inflection. for even greater precision, arbitrary ranges chosen to include the point of inflection can be averaged and compared; the greatest value among these is taken as dp dt . furthermore, the use of means contributes to the robustness and applicability of the symmetrical case. 3) m is the upper limit of growth. it is identified as a horizontal asymptote. in this case, use of the horizontal asymptote results in divergence of the series. therefore, ratios are implemented to calculate r or µ. let the horizontal asymptote be denoted as m0. set m1, an alternate limit, as equal to 2p . 4) values for p , dp dt , and m1 are substituted into the series to solve for r (µ). the series converges as in the symmetrical first case. 5) the use of m1 results in an inflated value for r (µ). to account for the substitution of m1, the result of the series is divided by the value of m1 m0 . example of the third case : graphical traditional method: µ=0.1704 maclaurin series method: p = 0.38; dpdt = 0.05;m0 = 0.61;m1 = 0.76 µ = 0.1523 error (disparity between the methods): 10.62 % 5 limitations irregular data is the greatest limitation to any method of determining microbial kinetics. environmental factors such as ph inconsistency, culture contamination, advances in systems science and applications (2013) vol.13 no.1 97 fig.7 graphical traditional method: µ=0.1704 or other laboratory error introduce complications to the data [16]. furthermore, microbial cultures are inherently not ideal. the use of an ideal model provides a basis of comparison by which inconsistencies may be determined. a numerical method strictly relies upon the data. however, irregular data will produce an irregular result. neither method will produce an accurate or precise estimate. the expansion of the maclaurin series is currently under investigation to extend the model beyond the point of inflection. successful extension would eliminate divergence as a limitation to the microbial growth model and its applications. 6 conclusion microbial growth kinetics is the numerical analysis of a complex living system. a data-based approach is significant as it determines microbial growth rate through a procedure designed to fit the data exactly. the use of models is essential to understanding the population structure through the relationship of population, population growth rate, carrying capacity, and time. population, and the population growth rate, change as a function of time. the models relate this through direct variation and a more complex function. the explored models are built upon exponential and logistic growth, and related through the transcritical bifurcation. each model is unique in its ability to predict and determine the growth of a particular microbe. culturing technique is also a factor in that it imposes or prevents natural limiting factors. batch cultures, and secondary metabolite fed batch, for example, display a pattern that may fall into the logistic, the symmetrical series, or the fourth-degree model. primary metabolite and continuous cultures tend towards exponential-like growth patterns, as modeled by exponential growth or the closed form of the modified alternating maclaurin series. for calculations using these models, the point of manipulation of culture conditions 98 anne talkington: an extension of a logistic model for microbial kinetics represents the point of environmental change, indicated graphically as the point of inflection. in full expansion, the discussed models converge to the curve of the modified alternating maclaurin series, up to the point of inflection. the partial curve is precise for this domain and range, and presents a scenario of semi-logistic analysis viewing the growth curve as the intersection of separate curves. in all cases, the models can be solved for the microbial specific growth rate, and yield precise results consistent with traditional methods of analysis. acknowledgements the authors would like to thank farm bureau robeson for their support. references [1] brock td. (1971), “microbial growth rates in nature”, bacteriol rev, vol.35, no.1, pp.39-58. [2] monod j. (1949), “the growth of bacterial cultures”, ann rev microbiol, vol.3, pp.371-394. [3] zwietering mh, jongenburger i, rombouts fm, riet kv. (1990), “modeling of the bacterial growth curve”, app env microbiol, vol.56, no.6, pp.18751881. [4] contois de. (1959), “kinetics of bacterial growth: relationship between population density and specific growth rate of continuous cultures”, j gen microbiol, vol.21, no.1, pp.40-50. [5] grijspeerdt k, vanrolleghem p. (1999), “estimating the parameters of the baranyi model for bacterial growth”, food microbiol, vol.16, no.6, pp.593605. [6] gilpin michael, e ayala, francisco j. (1973), “global models of growth and competition”, proceedings of the national academy of sciences of the united states of america, vol.70, no.12, pp.3590-3593. [7] kozusko f, bourdeau m. (2011), “trans-theta logistics: a new family of population growth sigmoid functions”, acta biotheor, vol.59, no.3-4, pp.273289. doi 10.1007/s10441-011-9131-3. [8] tsoularis a. (2001), “analysis of logistic growth models”, res. lett. inf. math. sci., vol.2, pp.23-46. [9] birch c.p.d. (1999), “a new generalized logistic sigmoid growth equation compared with the richards growth equation”, annals of botany, vol.83, no.6, pp.713-723 advances in systems science and applications (2013) vol.13 no.1 99 [10] yin x, goudriaan j, lantinga e.a, spiertz h.j. (2003), “a flexible sigmoid function of determinate growth”, annals of botany, vol.91, no.3, pp.361371. [11] yukalov v.i, yukalova e.p, sornette d. (2009), “punctuated evolution due to delayed carrying capacity”, physica d, vol.238, no.17, pp.1752-1767. [12] panikov ns. (1995), microbial growth kinetics, chapman & hall, london. [13] perni s, andrew pw, shama g. (2005), “estimating the maximum specific growth rate from microbial growth curves: definition is everything”, food microbiol , vol.22, no.6, pp.491-495. doi: 10.1016/j.fm.2004.11.014 [14] stewart, j. (2008), calculus, early transcendentals(6th ed.), belmont, california: thomson learning. [15] demana f. d, finney r. l, kennedy d, & waits b. k. (2003), calculus: graphical, numerical, algebraic, upper saddle river, new jersey: prentice hall. [16] gause g. f. (1934), struggle for existence, hafner, new york. [17] gause g. f. (1932), “experimental studies on the struggle for existence. i. mixed populations of two species of yeast”, j. exp. biol., vol.9, no.4, pp.389-402. [18] guckenheimer j. and holmes p. (1997), nonlinear oscillations, dynamical systems, and bifurcations of vector fields, 3rd ed. new york: springerverlag. corresponding author guo wei can be contacted at: guo.wei@uncp.edu. adv syst sci appl 2018; 04; 74-91 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/617 modeling and simulation of frequency response masking fir filter bank using approximate multiplier for hearing aid application raghavachari ramya1, sridharan moorthi*2 1) vlsi systems research laboratory, department of electrical and electronics engineering, national institute of technology, tiruchirapalli, india e-mail:407114003@nitt.edu 2) vlsi systems research laboratory, department of electrical and electronics engineering, national institute of technology, tiruchirapalli, india e-mail: srimoorthi@nitt.edu received july 20,2018; revised december 15,2018; published december 31,2018 abstract: the tremendous increase in the use of portable electronic devices is due to the development in the fields of signal processing and electronic technology. these battery operated devices needs reduction in power consumption with increased performance and long battery life. since cmos technology scaling fast approaches its physical limit of minimum supply voltage and smaller feature size, the hardware designer has to opt for new multiplier architectures for achieving low power and high speed performance. this paper proposes an area and power efficient approximate multiplier architecture. the error metrics and circuit characteristics are estimated to verify its performance advantage over other approximate multipliers. using frequency response masking approach, a 6-band non-uniform digital fir filter bank is developed using approximate multiplier for hearing aid application. audiogram matching is done with audiograms of two different types of hearing losses and the matching error is computed. simulation results show that the audiogram matching error falls within +/4 db range. keywords: approximate multiplier, audiogram, filter bank, hearing aid, wordlength 1. introduction audio signal processing to devise low power and low cost hearing aids become one of the major application areas in digital signal processing. it should also be noted that a large population affected by hearing defects of different levels requires medical assistance in the form of hearing aids. the hearing aid selectively amplifies the audio sounds such that the processed sound matches with that of one’s audiogram. the signal processing block involved in the design process is the digital filters. digital filter banks are networks of digital filters. the filter bank separates an input signal into many sub-band signals. these sub-band signals can be independently processed according to the requirements and can also be combined together to generate the desired output signal. non-uniform filter banks are more desirable because they can better match the physiological properties or the perceptual properties of the human ear. the major arithmetic block in digital filters is the multiply and accumulate (mac) unit. they are the largest power consuming units in digital filters. since the speeding up and energy reduction in vlsi technology fast approaching its physical limits [1-2], new approaches in digital design of arithmetic circuits are mandated for high speed, low power vlsi systems. thus, to achieve improved performance in terms of speed and power consumption it is essential to adopt power efficient * corresponding author: srimoorthi@nitt.edu mailto:srimoorthi@nitt.edu modeling and simulation of frequency response masking fir filter bank 75 copyright ©2018 assa. adv. in systems science and appl. (2018) hardware design of multiplier and adder to enhance the overall system performance. the adder and multiplier unit decide the speed, size and power dissipation of the mac unit and in turn of the filter bank module. approximate computing represents a paradigm shift in low-power vlsi design for error resilient signal processing applications. for multiplication involving large numbers, truncation and allowance of a small magnitude of error in the generation of partial products will improve speed and reduce power consumption in the multiplier to a large extent. digital signal processing for speech processing, multimedia, graphics, and computer vision can accommodate approximation methods to generate the outputs of multiplication and accumulation process. this is possible because of the inherent error resilience of above mentioned application areas. this paper proposes an area and power efficient approximate integer multiplier architecture. both unsigned and signed multiplier architectures are developed and their performances were studied. a 6-band non-uniform finite impulse response (fir) filter bank is developed using the frequency response masking (frm) technique and used to demonstrate the efficacy of the approximate multiplier in applications like hearing aids. the computationally intensive mac operation of fir filtering is performed using the proposed approximate multiplier. the approximate multiplier architecture is modeled using verilog hdl. the circuit characterization is done by evaluation of the area, power, and delay performance of the circuit. the error characterization of the proposed design is performed using standard error characterization techniques and compared with other similar approximate designs. the paper is organized as follows. the related research in approximate multiplier design and filter bank design for hearing aid are reviewed in section 2. the description of the proposed multiplier architecture and its error characteristics are presented in section 3. section 4 discusses the hardware implementation of the proposed approximate multiplier and its circuit characteristics are discussed. section 5 deals with the implementation of frm based fir filter bank for error tolerant hearing aid application using the proposed multiplier unit and audiogram matching for two different hearing losses are analyzed. finally, the conclusions are presented in section 6. 2. related works in this section, we describe some prior research that focuses on approximate multiplier design and filter bank design for hearing aid applications. the difficult design problem in fir filter design is to design filters with sharp transition band with less hardware complexity. this can be overcome by using frm design method [34]. frequency response masking technique is the most efficient method for the realization of digital fir filters with sharp transition bandwidth. the advantage of frm technique is that the filter has a very sparse coefficient vector and hence the hardware complexity is low. also, it has guaranteed stability and linear phase response. ying wei and yong lian [5] proposed an 8-band non-uniform digital fir filter bank for hearing aid applications using frm approach. two half-band filters are used as prototype filters and the filter bank provides a minimum stop band attenuation of 80 db. the filter bank is tested with audiograms for different hearing losses and matching errors are within +/5 db. nisha haridas and elizabeth elias proposed a set of variable bandwidth filters for hearing aid design using farrow structure [6]. a fixed number of bands are generated from the variable bandwidth filters by spectral shifting of the required bandwidth response. each filter of the bank is tuned to match different categories of audiograms. the reported matching errors are within 3 db for a 6-band filter bank and better matching is achieved with higher order filter banks. deng [7] proposed a three-channel variable filter bank for digital hearing aid applications. a normalized analog chebyshev type-i low pass filter is used as a prototype filter. using analog frequency 76 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) transformations and a modified bilinear transformation approach, variable low pass filter, variable band pass filter, and variable high pass filter are obtained from the prototype filter. these filters have adjustable gains and band edge frequencies which can be independently tuned for matching various hearing loss patterns. the variable filter bank was designed using infinite impulse response (iir) digital filters, which introduces non–linear phase to the system. the audiogram matching experiments shows that the maximum matching error is below 3.44 db. a 16-band low power non-uniform fir filter bank designed using multiplier less approach for a digital hearing instrument was proposed by setiawan, lesmana, and gwee [8]. the power reduction is mainly due to the replacement of multipliers with multiplexers and adder arrays and restricting the wordlength of filter coefficient and data word. minimum attenuation of 60 db is provided in the stop band for high programmability of magnitude gain. ying wei and debao liu [9] proposed a reconfigurable fir filter bank with adjustable sidebands. the frm technique used reduces the computational complexity and is able to show better results for the matching error. interpolation and decimation techniques are used to make the filter bank reconfigurable. a quasi-ansi s1.11 1/3 octave filter bank with 18 bands is proposed in [10] for hearing aid design with significant reduction in group delay and number of multiplications per sample. interpolated fir technique is adopted for reducing the computational complexity. ying wei and yinfeng wang [11] proposed an adjustable filter bank for personalized hearing aids. fractional interpolation technique along with symmetric filters and complementary filters are utilized to reduce the complexity. the filter bank can meet different types of hearing loss with acceptable delay. moving away from the exact arithmetic towards the approximate arithmetic for filter bank design gives enough opportunity to develop power and area efficient circuits and systems for error tolerant audio signal processing applications like portable hearing aids. discussions of the recent developments in the field of approximate multipliers are elaborated. several error tolerant approximate adders and multipliers have been proposed in the literature [12-15]. most of these multiplier designs use truncated multiplication method. truncation reduces the complexity of the multiplier unit by computing the most-significant bits of the product only. in order to reduce the error introduced by the truncation process, a correction factor may be added [16]. as more columns are eliminated, the error introduced by truncation also increases. kyaw et al. [17] developed an error tolerant multiplier (etm) in which the operands are split into a multiplication part with higher order bits and a nonmultiplication with lower order bits. every bit position from left to right is checked for the lower order bits and if either or both the operands are ‘1’, all the bit positions are set to ‘1’ from that bit onwards. for the higher order bits, normal multiplication operation is performed. the major drawback of the etm is the relatively large magnitude of the relative error. parag kulkarni et al [18] proposed a 2 × 2 underdesigned approximate multiplier block and used this block to construct large inaccurate multipliers. the inaccurate multipliers have shown to achieve an appreciable power savings over an accurate multiplier. the performance of the underdesigned multiplier decreases in terms of area and power with increasing wordlength. reza zendegani et al [19] presented three hardware implementations of approximate multipliers based on rounding of the inputs in the form of 2n was proposed. one unsigned and two signed rounding-based approximate multiplier (roba) architectures are implemented. in [20], two variants of signed 16-bit approximate radix-8 booth multipliers are developed using approximate recoding adder logic with and without truncation of a number of less significant bits in the partial products. venkatachalam & ko [21] proposed two variants of approximate multipliers. the partial products of the multiplier are altered using generate and propagate signals and the accumulation of generate signals is done column wise. approximate 4:2 compressors and adders are used to accumulate the remaining partial products. in the first variant, approximation is applied to all the columns of the partial product, and in variant 2, approximation is not applied to most significant column of partial products. even though multiplier design2 has lower relative error compared to multiplier design1, it offers less area modeling and simulation of frequency response masking fir filter bank 77 copyright ©2018 assa. adv. in systems science and appl. (2018) and power savings as compared to multiplier design1. narayanamoorthy et al [22] proposed dynamic segment method (dsm), static segment method (ssm), enhanced static segment method (essm) based approximation multipliers for various dsp and classification applications. a scalable dynamic range unbiased multiplier (drum) for approximate applications such as image filtering, jpeg compression, perceptron classifier is proposed in [23]. both [22] and [23] follows an approach for approximating the input operands based on the output of a complex leading one detection (lod) circuit. even though drum offers better accuracy than [22], the power and area gets increased due to the complex steering logic employed. the steering logic circuit extracts a predefined number of consecutive bits starting from the leading ‘1’ bit to be retained for further processing. as the size of the input grows, the complexity of the steering logic also increases. we propose an approximate multiplier circuit that has the same accuracy as that of dsm and better circuit characteristic than both dsm and drum. in this paper we target power and area reduction by reducing the wordlength of the input operands. the wordlength reduction is performed by right shifting and the amount of right shift required to truncate the operands is derived from the most-significant n/2-bits of the inputs. the approach retains a block of most significant information carrying n/2-bits beginning with the leading one for multiplication. this ensures much accurate result compared to a truncation process. the reduced circuit complexity and error characteristics thus achieved in the design of the approximate multiplier makes it very attractive for asic implementation. 3. proposed approximate multiplier the data wordlength affects the design parameters like speed, area, and power [24] in vlsi circuits. increased speed and reduced power consumption in dsp circuits can be achieved using data wordlength reduction in various arithmetic operations. this reduces the switching activity of cmos circuits and in turn the power consumption. the multiplier unit decides these performance parameters for dsp circuits. the data wordlength reduction can be applied to one or both the inputs of the multiplier and this greatly reduces the area and power in the multiplier unit. the main idea in the design of proposed multiplier lies in the fact that simple right shifting operation is exploited to reduce the wordlength of the input operands for multiplication with good accuracy along with reduced area and power compared to similar designs. the approximate multiplier architecture can be represented by three sub-blocks as shown in fig 3.1. fig.3.1 general block diagram of approximate multiplier the first sub-block is the wordlength reduction logic which reduces the input operand wordlength. second sub-block is an arithmetic unit which is an exact multiplier block of n/2bit wordlength and the last sub-block is a correction logic block to compensate for the reduction of input operand wordlength. in the proposed approximate multiplier, the 78 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) wordlength of the n-bit input operands are reduced to n/2-bits and multiplication is done with a single n/2-bit multiplier instead of n-bit multiplier. the block diagram of unsigned approximate multiplier is shown in fig.3.2. the wordlength reduction logic is composed of following blocks: n/2-bit priority encoder, a binary to excess one converter, 2:1 mux, and n-bit right barrel shifter. the priority encoder receives only the most-significant n/2-bit as the input. the location of one bit position present in the mostsignificant n/2-bit of the input i.e. a[n-1:n/2] and x [n-1:n/2] are encoded by the n/2-bit priority encoder. in order to obtain the amount of shift needed to obtain the truncated input operands, a binary-to-excess one converter and a 2:1 multiplexer (mux) are used. the select lines needed for the 2:1 mux are obtained by performing the bit-wise or operation of the most significant n/2-bit of the input operand. in case of input operands with wordlength less than n/2-bits, the select value will be set as ‘0’ and the 2:1 mux select the input as ‘0’ and the right barrel shifter will not provide any shift and input operands are not truncated. for operand size greater than n/2-bits, then select value will be set as ‘1’ and the 2:1 mux will select the required shift amount from the binary-to-excess one converter. the input operands are truncated to n/2-bit wordlength using right barrel shifters and multiplication is performed using n/2-bit multiplier. the correction logic is made up of 2n-bit left-barrel shifter which expands the truncated product to a 2n-bit number by left-shifting. the barrel shifter left shifts the product by an amount equal to the number of bit positions which is the sum of right shifts applied to both the input operands. the design can be extended for signed numbers by inserting a two’s complement block at the input of each branch. at the output the product value may be negated if necessary . for further simplifying the circuit complexity, the maximum negative input of magnitude -2n is left out from the computation and the resulting simplified architecture of proposed signed approximate multiplier is shown in fig.3.3. the signed multiplier comprises of a two’s complement block at the input and output, n/2-bit sign-extension encoder, inverter based control logic, right and left barrel shifters, unsigned exact multiplier block of n/2-bit wordlength. the multiplier receives the inputs in signed two's complement format. the sign of the input operand is determined and if the sign bit is set, the two’s complement block determines the absolute value of the negative operand. a modified priority encoder is used as a sign-extension encoder to calculate the effective size of the positive operands. the effective size of the positive operands is computed by finding the number of sign-extension bits (zeros) immediately to the right of the most-significant sign bit. for a n-bit multiplier, n/2 to log2(n/2) priority encoder is required since the multiplier block is of n/2-bit wordlength. the count of the sign-extension bits is passed to the control logic block which is used to calculate the amount of shift needed to truncate the inputs. if the size of the input operand is below n/2bits, then shift operation need not be performed and operands are not truncated. in case of operand wordlength above n/2-bits, the right barrel shifter right shifts the input operands to the required number of places as decided by the control logic to obtain the n/2-bit input operand. now the effective size of the input operands is limited to wordlength of size n/2-bit instead of n-bits and the multiplication is performed as an n/2×n/2-bit multiplication. unsigned exact wallace tree multiplier is adopted for the fundamental multiplier block. in order to compensate for the n/2-bit truncated multiplication, the correction logic block is used. the obtained truncated product is left shifted appropriately to compensate for the truncation performed. depending on the sign of the input operands, the unsigned product is negated to obtain the final signed product as the output. modeling and simulation of frequency response masking fir filter bank 79 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 3.2 block diagram of unsigned approximate multiplier fig. 3.3 block diagram of signed approximate multiplier 3.1. illustration of approximate multiplication of two 16-bit signed numbers as an example, a multiplication of two signed 16-bit numbers with multiplicand a= -259 and multiplier x = 517 is shown in fig. 3.4. the msb of the inputs are used to determine the sign of the operands. since the sign bit of a is ‘1’ the operand is negative and hence the absolute value of a is calculated and used for computation. hence the value of a becomes 259 and x =517. the control logic will decide the amount of shift needed for both the inputs based on the wordlength of the input. the right barrel shifter provides 1-bit right shift for operand a and 2-bit right shift for operand x. after right shifting, the operands are truncated 80 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) to 8-bit wordlength i.e., a′ =129 and x′=129 and process of multiplication is illustrated in fig 3.4. truncated product p′ 0100 0001 0000 0001 ap 1111 1111 1111 1101 1111 0111 1111 1000 ep 1111 1111 1111 1101 1111 0100 1111 0001 fig. 3.4 illustration of multiplication of two 16-bit signed numbers a= -257 and x=513; ap: approximate signed product , ep: exact product the truncated product p′ is expanded to 32-bits by left-shifting three bits to obtain the unsigned approximate product. since the multiplicand is negative, two’s complement of the unsigned product is taken and approximate signed product , ap = 1111 1111 1111 1101 1111 0111 1111 1000 = -133128 is obtained in place of the exact product value, ep = 1111 1111 1111 1101 1111 0100 1111 0001 = -133903. 3. 2. error characteristics of the proposed system the error metrics such as error rate, normalized mean error distance (nmed), absolute value of mean relative error distance (mred) and percentage mean accuracy are used to evaluate the performance of approximate multipliers [15] [25]. the definitions of the various error metrics are given as follows: 1. error distance (ed): for adders and multipliers, ed is the absolute difference between the accurate output (m) and the approximated output (m’). ed = | m – m’ | 2. mean error distance (med): it is computed by taking the average value of all possible eds . med = 1 n ∑ edi n i=0 where n is the total number of samples, and edi is the error distance in the ith value. 3. normalized mean error distance (nmed): it is the normalization of the mean error distance by the maximum output of the accurate multiplier. nmed = med/ mmax, where mmax is the maximum accurate product 4. mean relative error distance (mred): mred is computed as the average value of all possible relative error distances and is defined as: mred = 1 n ∑ edi mi ⁄n i=0 where edi and mi are the error distances and the accurate output of the ith input a 1111 1110 1111 1101 |a| 0000 0001 0000 0011 x 0000 0010 0000 0101 truncated input a′ 1000 0001 truncated input x′ 1000 0001 modeling and simulation of frequency response masking fir filter bank 81 copyright ©2018 assa. adv. in systems science and appl. (2018) 5. mean accuracy: it is defined as 100-mred. 6. error rate (er): it is defined as the probability of producing incorrect outputs for different combination of inputs. error rate = number of incorrect outputs/ total number of outputs in order to evaluate the error performance of the proposed multipliers two sets of hundred thousand random numbers are generated with uniform probability and the multiplication is performed. the error metrics for different 16-bit approximate multipliers were evaluated and summarized in table 3.1. table 3.1 arithmetic accuracy comparison of proposed 16-bit multiplier with state-of art designs unsigned designs error rate (%) nmed (%) mred (%) mean accuracy (%) proposed unsigned approximate multiplier 99.95 0.13 0.53 99.47 underdesigned multiplier udm[18] 80.85 1.37 3.33 96.67 unsigned roba [19] 99.96 0.69 2.93 97.07 venkatachalam and ko multiplier design1 [21] 99.80 1.78 7.63 92.37 dsm8×8 [22] 99.95 0.13 0.53 99.47 drum8 [23] 99.98 0.09 0.36 99.64 signed designs error rate (%) nmed (%) mred (%) mean accuracy (%) proposed signed approximate multiplier 99.87 0.032 0.52 99.48 signed roba [19] 99.90 0.172 2.88 97.12 approximate signed roba [19] 99.94 0.173 2.89 97.11 the results tabulated in table 3.1 shows, the current design of the approximate multiplier provides highest accuracy in terms of various error metrics. the proposed unsigned multiplier achieves the same error performance as that of the dsm8. this is due to the same functionality of the wordlength reduction logic even though the hardware implementation differs in them. due to the unbiased nature of error distribution of the drum8 design, it shows a better error 82 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) performance than the proposed multiplier. comparing with the unsigned roba design, the error performance of dsm8 is shown to be better [19]. the proposed multiplier also shows better error performance comparing with roba designs. smaller the values of nmed and mred shows that the proposed approximate multiplier gives higher accuracy over underdesigned multiplier, roba, multiplier design1 architectures. the relative error is also calculated for inputs of various wordlength and is tabulated in table 3.2. table 3. 2 variation of relative error (re) of outputs in percentage for various wordlength for the proposed approximate multiplier multiplier type input wordlength re <0.5% re <1% re <2% re <5% re <10% re <20% u n si g n ed 8-bit 3.79 % 5.4% 9.6% 31.7% 79.8% 100% 16-bit 47% 97 % 100% 32-bit 100% s ig n ed 8-bit 9.3% 10.72% 16.1% 41.7% 85.6% 100% 16-bit 48.6% 97.23% 100% 32-bit 100 % table 3.2 gives the percentage of outputs with the relative error less than the specified percentage for various wordlength such as 8-bit, 16-bit, and 32-bit proposed approximate designs. it must also be noted that the relative error reduces as wordlength of the operand increases. for a 16-bit design, the relative error is less than 2% and that for a 32-bit multiplier, it is less than 0.5%. the difference between the exact and approximate multiplier almost vanishes for 32-bit signal processing applications and it is negligible for 16-bit applications. 4. hardware implementation of the proposed approximate multiplier the proposed 16-bit approximate multiplier architectures are modeled using verilog hdl and synthesized in cadence rtl compiler using generic pdk (gpdk) 90-nm cmos technology with typical library settings. the functionality of the proposed multipliers is verified using cadence ncsim and all the designs are synthesized in rtl compiler with proper timing constraints. the post-synthesis circuit performance characteristics such as power, area and critical-path delay are tabulated in table 4.1. the power-delay-product (pdp) is a measure of energy and is defined as the product of average power and the corresponding delay of the circuit. since leakage currents are also contributing to the total power, better parameters for characterizing circuit performance are energy or power-delay product (pdp), and area-delay product (adp) [15] [26]. hence the compound metrics such as pdp, and adp were also computed and tabulated in table 4.1 for performance comparison. table 4.1 gives the post-synthesis circuit characteristics of various 16-bit multipliers. the circuit characteristics of the proposed design are compared against exact multiplier architectures like wallace tree (exact unsigned), and baugh-wooley multiplier (exact signed) architectures. for the comparison, a few of the state-of-art approximate multiplier architectures are considered and implemented using gpdk 90-nm cmos technology with the same timing constraints. the synthesized gate level netlist is used to extract the layout of the proposed approximate multiplier using cadence soc encounter and the physical layout of the proposed 16-bit multiplier is shown in fig.4.1. modeling and simulation of frequency response masking fir filter bank 83 copyright ©2018 assa. adv. in systems science and appl. (2018) table 4.1. post synthesis performance characteristics of various 16-bit multipliers unsigned designs power (µw) delay (ns) pdp (pj) area (µm2) adp (µm2. ns) proposed unsigned multiplier 303.47 3.97 1.21 3533 14026 underdesigned multiplier [18] 760.49 4.05 3.07 6241 25276 unsigned roba [19] 235.97 4.84 1.14 4522 21886 venkatachalam & ko multipier 1[21] 425.75 4.38 1.86 4527 19828 dsm 8 ×8 [22] 402.71 4.44 1.78 3548 15753 drum8 segment [23] 424.83 4.64 1.96 3806 17659 wallace tree (unsigned exact) 871.72 4.08 3.56 7012 28609 signed designs power (µw) delay (ns) pdp (pj) area (µm2) adp (µm2. ns) proposed signed multiplier 414.02 5.34 2.21 4012 21424 signed roba [19] 541.17 5.24 2.83 5640 29553 approximate signed roba[19] 537.40 5.15 2.76 5210 26831 baugh-wooley (signed exact) 769.07 4.56 3.51 6679 30456 the tabulated result in table 4.1. shows that the proposed unsigned and signed multiplier architectures have smaller values of area and power consumption (except that unsigned roba has lower power consumption) than other existing state-of-art multipliers, due to hardware efficient wordlength reduction logic employed, thereby their pdps and adps are also small. the area requirement of the proposed approximate unsigned (signed) multiplier is 50% (40%) lower than that of exact wallace (baugh-wooley) multiplier. the proposed 16-bit unsigned (signed) multiplier consumes 65% (45%) less than the total power consumed by exact wallace (baugh-wooley) multiplier. the above results show the efficiency of the multiplier in terms of area and power compared to standard designs. the energy or power-delay product (pdp), and adp of the proposed unsigned (signed) approximate multiplier are about 66% (37%), 51% (30%), lower than that of exact wallace (baugh-wooley) multiplier. table 4.2 illustrates the ranking of approximate multipliers in terms of both circuit metrics and error metrics such as pdp, adp, nmed and mred. the results reveal that the proposed unsigned multiplier gives minimum area, power, and adp; while drum8 gives lowest nmed and mred values and unsigned roba design gives lowest pdp value in comparison to all multipliers. but the nmed and mred error metrics of unsigned roba is higher than the proposed, dsm8, and drum8 designs. the other two multipliers have comparatively less performance in terms of design and circuit metrics. 84 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) table 4.2 ranking of approximate multipliers in terms of design metrics and error metrics unsigned designs pdp adp nmed (%) mred (%) proposed unsigned multiplier 2 1 1 2 underdesigned multiplier 6 6 4 4 unsigned roba 1 5 3 3 venkatachalam & ko multiplier design 1 4 4 5 5 dsm 8×8 3 2 2 2 drum8 segment 5 3 1 1 signed designs pdp adp nmed (%) mred (%) proposed signed multiplier 1 1 1 1 signed roba 3 3 2 2 approximate signed roba 2 2 3 3 the proposed signed multiplier gives lowest pdp, adp, nmed, and mred values in comparison with signed-roba and approximate signed roba architectures. the lowest value of pdp and adp is due to the reduced complexity wordlength reduction logic employed in the multiplier design. fig 4.1 layout of proposed 16-bit approximate multiplier 5. applicationapproximate fir filter bank the design method followed in the design of filter banks is the frequency response masking approach introduced in [3]. the method finds wide acceptance in the design of fir filters with sharp transition band. the filter bank proposed covers the frequency ranges from 0 to 8 khz with 6 non-uniform bands. the sampling frequency selected as 16 khz. the prototype filter is h(z) and the interpolated filters h(z4) and h(z2) are used for the filter design. the modeling and simulation of frequency response masking fir filter bank 85 copyright ©2018 assa. adv. in systems science and appl. (2018) higher order filters appropriately cascaded to generate the desired frequency response when masked with h(z). the transfer functions of different sub-bands in the 6-band filter bank are listed in the table 5.1. table 5.1 transfer functions for each sub-bands the prototype filter h (z) is designed in matlab using the least square method. the dependency of stop band attenuation on transition bandwidth was verified and is shown in fig.5.1. for a 40 tap filter. the simulation study to identify the dependence of transition bandwidth on stop band attenuation shows that a minimum transition bandwidth of 0.2 radians should be maintained to achieve the 50 db stop band attenuation. fig. 5.1 the variation of minimum stop-band attenuation with increasing transition bandwidth the normalized transition bandwidth of the prototype h(z) is fixed as 0.2 and the order of the filter is selected as 40 to ensure a minimum stop band attenuation of 50 db which is sufficient for hearing aids to compensate for the hearing loss satisfying the dynamic range requirements of the hearing impaired person. the structure of the 6-band filter bank is shown in fig.5.2. the lower bands are complementary bands of the upper bands formed by replacing h(z) with its complement hc(z) as the masking filter. the multipliers used in the circuit are integer multipliers. hence the integer filter coefficients are generated by scaling the real filter coefficients by a scaling factor 2n and rounded off to the nearest integer. the filtering process is the convolution of impulse response of the filter with the input audio samples. sub-band transfer function b1 h(z4)h(z2)h(z) b2 h(z2)h(z)h(z4)h(z2)h(z) b3 h(z)h(z2)h(z) b4 hc(z)-h(z2)hc(z) b5 h(z2)hc(z)-h(z4)h(z2)hc(z) b6 h(z4)h(z2)hc(z) 86 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 5.2 structure of 6-band filter bank approximate multipliers are used to compute the mac operation of the convolution sum which in effect can be interpreted as the convolution of approximate input samples with an approximate impulse response. the approximate filter thus results in a frequency response with less attenuation in the stop band. the frequency response of the 6-band filter bank designed in matlab is shown in fig.5.3. fig. 5.3 the frequency response of the 6-band approximate filter bank 5.1. audiogram matching the 6-band approximate fir filter bank is simulated in matlab and the gains of the different bands are adjusted to match the audiogram of hearing impaired. fig. 5.4a shows the audiogram matching result of the experiment. the audiogram of a patient with noise induced hearing loss (nihl) is selected for the experiment. the results show that the proposed 6-band approximate filter is a suitable one for the implementation of the hearing aids. the matching error is plotted in fig. 5.4b and is within +/3db. modeling and simulation of frequency response masking fir filter bank 87 copyright ©2018 assa. adv. in systems science and appl. (2018) (a) (b) fig.5.4 matching example 1 (a) nihl audiogram matching for the proposed approximate filter bank (b) plot of matching error another matching experiment is performed using old age related presbycusis hearing loss. fig 5.5a shows the plot of audiogram matching and the corresponding matching error plot are shown in fig. 5.5b. results show that matching error can be adjusted within +/4db for the particular case considered. 88 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) (a) (b) fig. 5.5 matching example 2 (a) presbycusis audiogram matching for the proposed approximate filter bank (b) plot of matching error 5.2. area and power savings the area and power savings achieved for the approximate filter bank are evaluated when using proposed approximate multiplier for the digital hearing aid application. table 5.2 gives the area and power savings achieved in the 6-band digital filter bank designed using with and without the utilization of proposed approximate multiplier. the filter-bank is synthesized in rtl compiler and post-layout design parameters are taken from cadence soc encounter. modeling and simulation of frequency response masking fir filter bank 89 copyright ©2018 assa. adv. in systems science and appl. (2018) table 5.2 area and power savings achieved for the proposed approximate filter bank exact design proposed approximate design savings achieved area (µm2) power (mw) area (µm2) power (mw) area (%) power (%) 487233 7.170 374268 6.17 23 14 the power and area results show the advantage of the proposed approximate design in comparison with exact design. the approximate filter bank consumes 14% less power and 23 % less area than the filter bank designed using exact multiplier. 6. conclusion in this paper, we proposed an area and power efficient approximate multiplier based on simple shifting operation to approximate a number for multiplication. the proposed multiplier consumes less area, power, and in energy compared to exact and other recently published approximate multipliers. by adopting frequency response masking approach, a 6-band nonuniform digital fir filter bank is developed using approximate multiplier for hearing aid application. audiogram matching is done with audiograms of two different types of hearing losses and the matching error is computed. it is found that the design offer comparable matching error performance with that of hearing aids implemented with exact arithmetic filter banks reported elsewhere. the large attenuation in the stop band ensures high programmability of the magnitude response of the filter bank, which may be utilized to attain arbitrary reduction of the matching error. the customary low clock rates of digital audio devices makes allowance for time sharing of the resources which may further improve the hardware efficiency. references 1. itoh, k., yamaoka, m., & oshima, t. (2010). adaptive circuits for the 0.5-v nanoscale cmos era. ieice transactions on electronics, 93(3), 216-233, https://doi.org/10.1587/transele.e93.c.216 2. nowak, e. j. (2002). maintaining the benefits of cmos scaling when scaling bogs down. ibm journal of research and development, 46(2.3), 169-180, https://doi.org/ 10.1147/rd.462.0169 3. lim, y. (1986). frequency-response masking approach for the synthesis of sharp linear phase digital filters, ieee transactions on circuits and systems, 33(4), 357-364, https://doi.org/10.1109/tcs.1986.1085930 4. lian, y., zhang, l., & ko, c.c. (2001). an improved frequency response masking approach for designing sharp fir filters, signal processing, 81(12), 2573-2581, https://doi.org/10.1016/s0165-1684(01)00149-9 5. lian, y., & wei, y. (2005). a computationally efficient non uniform fir digital filter bank for hearing aid, ieee transactions on circuits and systems i: regular papers, 52 (12), 2754-2762, https://doi.org/10.1109/ tcsi.2005.857871 6. haridas, n., & elias, e. (2016). efficient variable bandwidth filters for digital hearing aid using farrow structure, journal of advanced research, 7(2), 255-262, https://doi.org/10.1016/j.jare.2015.06.002 https://doi.org/ https://doi.org/10.1109/tcs.1986.1085930 https://doi.org/10.1016/s0165-1684(01)00149-9 https://doi.org/10.1109/ https://doi.org/10.1016/j.jare.2015.06.002 90 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) 7. deng, t.b. (2010). three-channel variable filter-bank for digital hearing aids, iet signal processing, 4(2), 181-196, https://doi.org/10.1049/iet-spr.2008.0164 8. setiawan, r., lesmana, v.p., & gwee, b.h. (2005). design and implementation of a low power fir filter bank, journal of the institution of engineers, singapore, 45(5), 77-87. 9. wei, y., & liu, d. (2013). a reconfigurable digital filterbank for hearing-aid systems with a variety of sound wave decomposition plans. ieee transactions on biomedical engineering, 60(6), 1628-1635, https://doi.org/10.1109/tbme.2013.2240681 10. lin, c. h., chang, k. c., chuang, m. h., & liu, c. w. (2012). design and implementation of 18-band quasi-ansi s1.11 1/3-octave filter bank for digital hearing aids. in 2012 ieee international symposium on vlsi design, automation, and test (vlsi-dat), pp. 1-4, https://doi.org/10.1109/vlsi-dat.2012.6212620 11. wei, y., & wang, y. (2015). design of low complexity adjustable filter bank for personalized hearing aid solutions. ieee/acm transactions on audio, speech, and language processing, 23(5), 923-931, https://doi.org/10.1109/taslp.2015.2409774 12. jiang, h., han, j., & lombardi, f. (2015). a comparative review and evaluation of approximate adders. in proceedings of the 25th edition on great lakes symposium on vlsi, pp. 343-348, https://doi.org/ 10.1145/2742060.2743760 13. zhu, n., goh, w. l., zhang, w., yeo, k. s., & kong, z. h. (2010). design of lowpower high-speed truncation-error-tolerant adder and its application in digital signal processing. ieee transactions on very large scale integration (vlsi) systems, 18(8), 12251229, https://doi.org/10.1109/tvlsi.2009.2020591 14. jiang, h., liu, c., maheshwari, n., lombardi, f., & han, j. (2016) a comparative evaluation of approximate multipliers. in 2016 ieee/acm international symposium on nanoscale architectures (nanoarch), beijing, china, 191-196, https://doi.org/10.1145/2950067.2950068 15. jiang, h., liu, c., liu, l., lombardi, f., & han, j. (2017). a review, classification, and comparative evaluation of approximate arithmetic circuits. acm journal on emerging technologies in computing systems (jetc), 13(4), 60:1–60:34, https://doi.org/10.1145/3094124 16. schulte, m.j., & swartzlander, e.( 1993). truncated multiplication with correction constant [for dsp]. in proceedings of ieee workshop on vlsi signal processing vi, veldhoven, netherlands, 388–396, https://doi.org/10.1109/vlsisp.1993.404467 17. kyaw, k. y., goh, w. l., & yeo, k. s. (2010). low-power high-speed multiplier for error-tolerant application. in 2010 ieee international conference on electron devices and solid-state circuits (edssc), pp.1-4, https://doi.org/10.1109/edssc.2010.5713751 18. kulkarni, p., gupta, p., & ercegovac, m. (2011, january). trading accuracy for power with underdesigned multiplier architecture. in proceeding of 24th international conference on vlsi design, chennai, india, 346–351, https://doi.org/ 10.1109/vlsid.2011.51 19. zendegani, r., kamal, m., bahadori, m., afzali-kusha, a., & pedram, m. (2017). roba multiplier: a rounding-based approximate multiplier for high-speed yet energyefficient digital signal processing. ieee transactions on very large scale integration (vlsi)systems, 25(2),393-401, https://doi.org/10.1109/tvlsi.2016.2587696 20. jiang, h., han, j., qiao, f. & lombardi, f. (2016). approximate radix-8 booth multipliers for low-power and high-performance operation. ieee transactions on computers, 65(8), 2638-2644, https://doi.org/10.1109/tc.2015.2493547 https://doi.org/10.1049/iet-spr.2008.0164 https://doi.org/10.1109/tvlsi.2009.2020591 https://doi.org/10.1145/2950067.2950068 https://doi.org/10.1145/3094124 https://doi.org/10.1109/vlsisp.1993.404467 https://doi.org/ https://doi.org/10.1109/tc.2015.2493547 modeling and simulation of frequency response masking fir filter bank 91 copyright ©2018 assa. adv. in systems science and appl. (2018) 21. venkatachalam, s., & ko, s.b. (2017). design of power and area efficient approximate multipliers. ieee transactions on very large scale integration (vlsi) systems, 25(5), 1782-1786, https://doi.org/10.1109/tvlsi.2016.2643639 22. narayanamoorthy, s., moghaddam, h.a., liu, z., park, t., & kim, n.s. (2015). energy-efficient approximate multiplication for digital signal processing and classification applications. ieee transactions on very large scale integration (vlsi) systems, 23(6), 11801184, https://doi.org/10.1109/tvlsi.2014.2333366 23. hashemi, s., bahar, r., & reda, s. (2015). drum: a dynamic range unbiased multiplier for approximate applications. in proceedings of the ieee/acm international conference on computer aided design (iccad), austin, tx, usa, 418–425, https://doi.org/10.1109/iccad.2015.7372600 24. chandrakasan, a. p., potkonjak, m., mehra, r., rabaey, j., & brodersen, r. w. (1995). optimizing power using transformations. ieee transactions on computer-aided design of integrated circuits and systems, 14(1), 12-31, https://doi.org/10.1109/43.363126 25. liang, j., han, j., & lombardi, f. (2013). new metrics for the reliability of approximate and probabilistic adders. ieee transactions on computers, 62(9), 1760-1771, https://doi.org/10.1109/tc.2012.146 26. sengupta, d., & saleh, r. (2007). generalized power-delay metrics in deep submicron cmos designs. ieee transactions on computer-aided design of integrated circuits and systems,26(1), 183-189, https://doi.org/ 10.1109/tcad.2006.883926 https://doi.org/10.1109/tvlsi.2016.2643639 https://doi.org/10.1109/tvlsi.2014.2333366 advances in systems science and application (2015) vol.15 no.2 176-187 making a causal contextualization with the four causes of aristotle j.-f. vautier french society of systems science (afscet), paris, france abstract this article presents an empirical study which proposes a chained contextualization based on the four causes of aristotle: material, formal, moving and final ones. this contextualization is a causal one, dedicated to help an analyst looking for root causes of a problem or an unwanted event. it consists of four different chains of causes beginning by the problem or the event and ending by root causes. a “why” question is asked at each step of questioning. two examples are studied: “why is this part brittle?” and “why is this valve opened?”. an aristotelian causal contextualization aims at helping an analyst to deepen the chains of causes and to extend the range of causes which can be identified with a traditional method. then, the first main interest of an aristotelian causal contextualization is to help an analyst to see, during an interview, the field of responses of the interviewee and to detect responses which stay on the same “plateau” of explanation (reformulations of symptoms or causes, or description of different steps of an activity). moreover, since the aristotelian causal contextualization permits to see the kind of field of causality which is used by an interviewee, it is possible to guide him more easily in other fields of causality by using appropriate questions related to other fields of responses. it is the second main interest of this contextualization. keywords causality, aristotle, contextualization, root cause 1 introduction a communication related to an approach of contextualization, based on a musical metaphor (the concept of systemic score), was presented at the 7th congress of eus in 2008 [1]. the four causes of aristotle [2-5] were used and grouped in a disc-shaped which was named a note. in this article, the four causes of aristotle, still grouped as a note, are also used, but to present a chained contextualization aimed at helping an analyst looking for root causes of a problem or an unwanted event. this contextualization consists of different chains of causes beginning by a problem or an event, occurring in a working situation of an industrial company, and ending by root causes. the root causes are defined as causes that condition the occurrence of more direct causes of a problem or an unwanted event. they are more general than direct causes. it means that the work on a root cause may lead to a larger variety of effects than the work on a direct cause. the general purpose of this article is to discuss the interest of this causal conadvances in systems science and application (2015) vol.15 no.2 177 textualization. then, the first part will present how to take place an aristotelian causal contextualization. secondly, two examples of this contextualization will be proposed and finally a discussion will be presented. 2 how does a causal contextualization using the four causes of aristotle take place ? 2.1 an aristotelian note and the four causes in the concept of systemic score, a note is represented as a disc-shaped set of the four causes of aristotle [1]. fig.1 representation of the four causes as an “aristotelian note” the description of the different blocks is as below. the material causes represent the components of a system. the formal causes represent the principles of functioning of the system (which explain also its shape). the moving causes represent the actors dedicated to the process of designing and managing the system. according to aristotle, the notion of movement is more general than the notion that is considered nowadays (which is only related to kinematic). in this way, the aging of a person is for example a movement too. the final causes represent the purposes which are aimed by the system. then, if the system is for example a car: the material causes represent the components of the car: the body, the engine, the wheels, the tyres...; the formal causes represent the principles of functioning: the propulsion of the car resulting from the movement of the four wheels on the ground...; 178 j.-f. vautier: making a causal contextualization with the four causes of aristotle the moving causes represent the build of the car by the manufacturer and the way of driving and maintaining the car by its driver(s) (indeed that strongly explains also the state of the car...); the final causes represent for what uses the car was designed and also the current uses of the car by the driver(s): to transport people, goods... in a sportive way or not, just in town or not? 2.2 principles and pratical aspects of contextualization in the present contextualization, the process for identifying the causes is based on “why” questions. the four causes are used to answer, in four different ways, questions which begin by “why”. for example, with a question like: “why is this car red ?”, the four responses may be: because a red paint is on the body of the car (material cause); because a red color is emitted by the painting of the body of the car (formal cause); because somebody painted the car in red (moving cause); because the manufacturer wanted to enjoy the buyer (final cause). in this paper, the two initial “why” questions of the two examples (cf. §3), which will permit to start the contextualization, are “why is this part brittle ?” and “why is this valve opened ?”. then, it means that the system which is firstly considered, to start the contextualization, is a local one: a working situation in a company for example. afterwards, for moving and final causes, a more general system is progressively considered to point out more general causes. in practice, an aristotelian causal contextualization is supported by a process of iterative question-answer which starts with an aristotelian note. new notes are built from each of the four causes, hence the previous denomination. the contextualization starts with a question which begins by “why”. the first answer is represented as an aristotelian note, whose the four parts are filled in, i.e. a set of four responses which correspond to the four parts of the note. afterwards, a new question (which still begins by “why”) built from each previous response is asked. the four new answers are also represented as aristotelian notes but just one part of these notes is filled in. the part dedicated to one cause is related to the same part of the previous note which is considered. this general process is repeated several times. to find the responses, a specific “connector” is used for each kind of previous responses. thus, this contextualization works with keys of questioning to go advances in systems science and application (2015) vol.15 no.2 179 back step by step to root causes. a connector is a set of few words which begins by “because”, introduces a response and connects each kind of response to the “why” question. then, to contextualize, it is used iteratively the connector: • “because it consists of ...” for the “material causes” (an analytical approach); • “because there is/are ... (the principle of functioning of this phenomenon)” for the “formal causes”; • “because there was... (a project manager of it)” for the “moving causes”; • “because the project manager wanted to...” for the “final causes”. the connector “because the project manager wanted to...” was preferred rather than “in order to” for the final causes since the focus is here on the process inside a company and the decisions which were taken. then, the notion of project manager was introduced since this function is very common in a lot of companies (and besides it is often a part of the job of a manager). for the moving causes, the focus is also on the role of the project manager of the system since the purpose is to present the root causes on which it can be acted on. indeed, the designer of the current system is generally not still present in the plant. in this way, as aristotle suggested, there are a lot of relationships between moving cause and final cause (hence the curved dotted lines in the two examples below). 3 examples of contextualization the first question which is considered is: why is this part brittle? the second one is: why is this valve opened? these two questions are “why” questions, respectively related to a problem or an unwanted event. in these examples, the problem or the unwanted event are related to the occurrence of decision failures. it means that, for instance, an operator did not use the right procedure... considering the part, which is brittle, there was a dysfunction in the process of making it, since the part ought to be resistant to breaking. furthermore, there was no sabotage (otherwise the final chain of causes would be different). considering the activity of closing the valve in a room of a plant, there is an unwanted event since the valve ought to be closed. 3.1 why is this part brittle ? the responses of this question is in the middle of the figure 2 and the other answers are built from these previous ones (look at the arrows). for example, to find the first response of “why is this part brittle?” in the “material causes” axis, an interviewee may use the connector “because it consists 180 j.-f. vautier: making a causal contextualization with the four causes of aristotle of” to find the response “brittle matters”. afterwards, during an interview, an analyst has to ask “why” and to propose that the interviewee uses again the same connector for the same axis or another one for the other axis. fig.2 example of an aristotelian causal contextualization 3.2 why is this valve opened ? the responses of this question are in the middle of the figure 3 and the other answers are built from these previous ones (look at the arrows). the valve is a butterfly valve which is manually opened and closed. it consists of a core and a disc which rotates around an axis and there is generally a fluid in the pipe when the valve is opened. 3.3 some remarks about the contextualization 1) representation of an aristotelian causal contextualization in the two examples, only one chain of aristotelian notes was presented in each field of causality. other chains of notes could be carried out with the same answer-formulation. for example, it could be possible to go back to root material causes from each components and afterwards sub-components of the system. nevertheless, our objective is not to propose another method of determination advances in systems science and application (2015) vol.15 no.2 181 of root causes. it is to discuss the interests of this kind of contextualization to help the analyst when he carries out a fact gathering and uses methods, like, for example, the 5 whys [6-7] or the causal tree analysis [8]. fig.3 example of an aristotelian causal contextualization 2) expression of the moving and final causes the focus was on two kinds of examples which are related to two domains: quality and safety. in the first one, the notion of quality, as a result [9], may be evaluated in terms of scope (of the product), cost and time/duration (to obtain the product) [10]. it is why the final causes are indicated with one objective and one or two characteristics of it (in terms of scope, cost or time/duration) which may mainly contribute to explain the occurrence of the problem (of quality). concerning safety, it is not provided, for final causes, characterizations of an objective in terms of scope, cost or duration. indeed, “to be opened” is not, for the valve which is considered, a defect by itself (sometimes the application of the procedures leads to open it). on the contrary, “to be brittle” is, for the part of 182 j.-f. vautier: making a causal contextualization with the four causes of aristotle the first example, a defect by itself (it is never a quality for it). it is why it is necessary to add some characteristic in terms of scope, cost or time/duration. note that for other things like for example some kind of cakes, “to be brittle” is a quality... in our context of looking for the root causes, it was considered that “to be brittle” or “to be opened” are respectively a problem or an unwanted event. for the part, it could have been related to some problems concerning raw materials used to make it. but, in this case, the final causes would have been different. it means that, in the two examples, operators made mistakes. more precisely, these mistakes are rule-based or knowledge-based errors [11]: rule-based error: for example, an operator opened the valve instead of keeping it closed. he did not analyze correctly the situation and then did not use the right procedure; knowledge-based error: in this case, the operator did not know he had to close the valve to achieve this operation. a skill-based error (for example an operator wanted to close the valve but he kept the valve opened: it was a slip) may be here considered only for the part since in our example the operator wanted to open the valve. the explanation of the occurrence of the human failures comes from the examination of the moving causes. in this axis, the explanation of the problem, the unwanted event or a direct moving causes results from the lack of adaptation between four sets of factors: local work organization, team and competences, technical devices and work environment (like premises, noise) which can be identified in each root moving causes. these groups of factors appear in the classical methods using “why” questions or in human and organizational factors methods. the aristotelian causal contextualization proposes here a framework to locate these factors in the working situation, the unit or the plant. the factors are very similar to those which can be found in a mto approach (man, technology, organization) [12]. then, in a few words, if there is a mistake in the operation which is considered, the explanation (the lack of interaction between the previous factors) is found in the moving causes which are identified. 4 discussion: interests of this aristotelian causal contextualization to help an analyst even if more examples would be necessary to confirm these main findings, here are some first elements. two points are proposed: deepen the chains of causes, and extend the range of causes. advances in systems science and application (2015) vol.15 no.2 183 4.1 deepen the chains of causes after the occurrence of a problem, an unwanted event the first step of an analysis is generally to carry out a fact gathering. it is often the step before an iterative question-answer session using a more or less complicated “why” question to find root causes and then to identify chains of causes that explain the occurrence of the event, the problem· · · afterwards, the objective is to propose solutions to remove the root causes. nevertheless, sometimes analysts may not really go back to root causes, even if there is a presentation of different causal chains in the analysis. in fact, there may be a treadmill effect at one or several levels of explanation of the event or the problem. there is a “plateau” in which, for example, it may be found reformulations or descriptions of different steps of a human activity or a mistake. it could partly explain that analysts may just consider the symptoms or very direct causes. it means also that the number of question-answer is not a sufficient criterion to provide an in-depth analysis and to find the real root causes. this problem may be observed especially when the focus is on the material or formal fields of causality. two phenomena may be mainly described: concerning the material causality, the evocation of a succession of steps of the activity of an operator instead of detailing one step into sub-steps... (for example a detailing process would be: make a mistake / press the wrong button / push the button a with one finger / push of 0.5 cm on the button a) concerning the formal causality, the reformulation of the cause instead of going back to a more general principle of explanation (for a mistake for example, a more general principle indicates the type of mistake and afterwards the type of characterization of the mistake...: make a mistake / realize an operation in a different way that is assigned / consider the results of the activity to evoke the mistake) to go back really to root causes, in the material causality and formal causality fields (cf. the two examples of §3), it is necessary to identify, at each step of questioning (with a “why” question), respectively, the basic components and the general principle of functioning (which underlies the previous formal cause which is considered). 1) material causality: succession of elements as some steps of the activity of a person here is an example which illustrates this phenomenon of treadmill effect: “why did he press the wrong button? because he confused the two buttons why did he confuse the two buttons? because he had a wrong mental representation why did he have a wrong representation? because he did not analyze correctly the working situation 184 j.-f. vautier: making a causal contextualization with the four causes of aristotle why did he not analyze correctly the working situation? because he did not detect the pertinent information” in this succession of question-answer, the focus is on a mistake. first of all, the movement of a finger of the operator is evoked, next some dysfunctions of the mental process are explored and, in the end, the problem of perception is considered. in other words, the model of j. rasmussen [13] is followed in the reverse direction (from execution of action to the perception). thus, it is taken into account several elements which are just different steps of the mental information processing and execution of action! 2) formal causality: some kinds of reformulation • example 1: going from action (or human failure) to emotion and vice versa “why did he move these objects nervously? because he was angry” or (in the opposite way) “why was he happy? because he was watching the sun set over the sea” • example 2: going from an expression of a mistake to another expression of the same mistake “why did he make a careless mistake ? because he did not pay attention” “why did he not pay attention ? because he worked too quickly” “why did he work too quickly ? because he used a mode of thought “system 1” (fast thinking) [14]” then, in cases of reformulations or successions of different steps of an activity, the analysts do not go back to root causes. they stand still! this phenomenon seems to occur particularly with material and formal causes. thus, the first interest of the aristotelian causal contextualization is both to help the analyst (during an interview) to see the field of responses of an interviewee and to avoid some current treadmill effect. when the analyst-interviewer detects responses which stay on the same “plateau” of explanation, for example reformulations of symptoms, or direct causes, or the description of different steps of an activity, it means that the interviewee stands still at one level of causality without going back really to root causes. 4.2 extend the range of causes a little variety of causes is sometimes taken into account in the analyses [6-7]. it may mean that analysts cannot go beyond their current knowledge about the working situations. the interests of an aristotelian causal contextualization is to propose four ways of questioning a problem and force the analysts to multiply the points of view. it leads analysts not to focus on a single set of root causes. the example of §4.1 about material causality, with a 4 why’s iterative questioning, shows that it is easy to stay in the “material causes” field of a mistake without going to another causality field like moving causes. for example, considering the question “why did he press the wrong button?” the response could advances in systems science and application (2015) vol.15 no.2 185 also have been: “because the buttons of displays were too small” or “because he did not know that the procedure had changed”. thus, since the aristotelian causal contextualization permits to see the kind of field of causality which is used by an interviewee, it is possible to guide him more easily in other fields by using appropriate questions related to the wanted field of responses. for example, if an analyst wants to identify some causal factors related to moving causes, he can turn to questions like “what technical devices or competences are necessary to make this product?” or “what are the objectives of the project manager of the unit?”. in other words, the range of fields proposed by an aristotelian causal contextualization is very large since, as r. caratini (2012) [15] indicated, the material and formal causalities are immanent (they depend only on the object) and the moving and final causalities are external. it is an important interest of this kind of formalization to permit to enlarge the set of points of view! 5 conclusion an aristotelian causal contextualization is an implementation of the notion of contextualization [16]. it is a support to identify and categorize the causes of a problem or an event [17-18]. then, it may be a way to help an analyst looking for root causes. to end this article, here are some perspectives of works for the future: the possibility to distinguish different ways of causality and the existence of specific connectors for each field of causality are perhaps a way to find, in a chain, a “good distance” from one cause to another. indeed, with a lot of approaches, the different causes may be more or less “distant” one to another. it means that, with these latest approaches, another analyst may often identify an intermediate cause between two successive causes of a chain; the use of an aristotelian causal contextualization aims at completing a common investigation about the root causes. it is perhaps a way to make results more repeatable from one analyst to another (each of them using their own method completed by this aristotelian contextualization). do we see the same evolution than the airplane traffic? in this latest domain, the air corridors were created so that all flights can be repeatable... may these air corridors be compared to the 4 causality fields of an aristotelian causal contextualization? then, going back in the past to the basics of aristotle will be perhaps a way to take a better jump forward! references [1] j.-f. vautier. (2008), “a systemic approach to question complexity: the systemic scores”, 7th congress of the european union for systemics (eus186 j.-f. vautier: making a causal contextualization with the four causes of aristotle ues), lisbon, portugal. [2] aristote. (2002), physics, book ii, chapters 3 and 7, translation of pierre pellegrin, editions garnier flammarion, paris. [3] aristote. (2008), metaphysics (the original is in french: métaphysique), book ∆, chapter 2, translation of marie-paule duminil et annick jaulin, editions garnier flammarion, paris. [4] aristotle. (1930), physics, book ii, chapters 3 and 7, translation of r.p. hardie and r.k. gaye , oxford press. [5] p. pellegrin. (2007), dictionary aristotle (the original is in french: dictionnaire aristote), ellipses. [6] fs. anderson. (2009), “root cause analysis: addressing some limitations of the 5 whys”, see http://www.qualitydigest.com/inside/fda-compliancenews/root-cause-analysis-addressing-some-limitations-5-whys.html. [7] t. minoura. (2007), “talks about problems with 5-whys”, see http://www.taproot.com/archives/710. [8] t. meyer and g. reniers. (2013), engineering risk management, walter de gruyter. [9] r. atkinson. (1999), “project management: cost, time and quality, two best guesses and a phenomenon, its time to accept other success criteria”, (international journal of project management, vol.17, no.6, pp.337-342, elsevier science ltd and ipma. [10] k. schwalbe. (2009), an introduction to project management, course technology cengage learning. [11] j. reason. (1990), human error, cambridge university press. [12] o. andersson and c. rollenhagen. (2002), “the mto concept and organisational learning at forsmark npp“, sweden (iaea international conference on safety culture in nuclear installations, rio de janeiro, brazil. [13] j. rasmussen, a. m. pejtersen and l. p. goodstein. (1994), cognitive systems engineering, wiley. [14] d. kahneman. (2011), thinking, fast and slow, macmillan. [15] r. caratini. (2012), initiation to philosophy (the original is in french: initiation à la philosophie), editions archipoche. advances in systems science and application (2015) vol.15 no.2 187 [16] e. morin. (2006), for a reform of the thinking (the original is in french: pour une réforme de la pensée), entretiens nathan, editions nathan. [17] j.-f. vautier. (2014), “the birth of systems according to a new approach: the birth triangle”, 9th congress of the european union for systemics (eusues), valencia, spain. [18] j.-f. vautier. (2011), “the interests of the lists of categories to deal with diversity and systemics” (the original is in french:“de lintérêt des listes de catégories pour appréhender la diversité et aborder la systémique) the interests of the lists of categories to deal with diversity and systemics”, acta europeana systemica, vol.1, pp.1-8. corresponding author j.-f. vautier. can be contacted at: jean-francois.vautier@cegetel.net advances in systems science and applications (2014) vol.14 no.4 378-387 computer simulation of social partnership in the system of continuing professional education vladimir k. dyachenko, guennady a. ougolnitsky and larissa v. tarasenko southern federal university, russia abstract a game-theoretic model of social partnership in the system of continuing professional education is proposed. some results of the model identification and investigation based on simulation modeling are considered. a comparative analysis of egoistic and cooperative approaches to the social partner-ship is conducted. keywords: social partnership, continuing professional education, simulation modeling, identification, dynamic games. 1 introduction a system of professional education should be flexible enough to prepare qualified and competitive specialists capable to improve their knowledge and skills in the changing environment. one of the important mechanisms providing the solution of this problem is social partnership that allows for the cooperation of employers, universities, and students. traditionally, social partnership is considered in conformity with relations between workers, employers, and trade unions[1]. a wider view is presented in[2]. social partnership relations also exist in the system of higher education. in order to be successful in these partnerships, agents must look carefully at how they are influenced by factors such as race, class, gender, age, culture, histories and other differences[3]. unequal power relations prevailing in such partnerships affect the ability to share information, solve problems and form honest relationships. the most complex social challenges such as post-secondary access and success for under-represented students, diversification of the workforce, poverty, environmental degradation, and global health exceed the problem-solving capacity of single organizations or societal sectors. social partnership provides colleges and universities, corporations, government agencies, non-profits, and other organizations with a model for how to effectively address these and other pressing social issues through strong, effective collaboration[4]. there are only a few papers that contribute to the mathematical modeling of the problems of social partnership. d. talman and z. yang[5] propose a general model of partnership formation. let a number of agents want to conduct some activities. they may act alone or seek a partner for cooperation and need in the latter case to consider with whom to cooperate and how to share the profit in a cooperative or competitive environment. necessary and sufficient conditions under which an equilibrium exists are given. in[6] the cobb-douglas production function is used to measure synergy effects of a public social private partneradvances in systems science and applications (2014) vol.14 no.4 379 ship project. the microeconomic approach opens up a possibility for allocating the synergy effects to partners according to a cooperative principle of optimality. analysis of a social partnership among a complex network of stakeholder organizations is fulfilled in[7]. m. diaconu and a. pandelica examine the partnership relationship between economic academic and business environment and propose a series of measures regarding how this relationship can shape the modern university[8]. the authors of[9] propose a synergy model of a smart tri-partite partnership among polytechnics industry students in the industrial training program in malaysia. social partnership in continuing professional education is a specific system of joint activities of the education system agents characterized by trust, common objectives and values, and providing highly qualified, competitive, and mobile specialists for the labor market. the main research hypothesis is that social partnership permits to increase a level of professional competence of the students. it seems natural to use the formalism of differential games for description of social partnership relationships[10]. due to the high complexity of the differential game model the techniques of simulation modeling are applied for its investigation[11]. some previous results of the authors’ approach in social systems modeling are presented in[12, 13]. 2 a general description of the model in the proposed model social partnership relations among employers, students, and university are considered. the payoff functions are as follows: jp = t∑ t=0 gp (up (t), ub(t), uc(t), x(t)) → max, up (t) ∈ up jb = t∑ t=0 gb(up (t), ub(t), uc(t), x(t)) → max, ub(t) ∈ ub jc = t∑ t=0 gc(up (t), ub(t), uc(t), x(t)) → max, uc(t) ∈ uc (1) where n = p,b,c is a set of players, namely: employer; university; student. the systems dynamics are given by the equation x(t+ 1) = x(t) + f(x(t), up (t), ub(t), uc(t)), x(0) = x0 (2) here up(t), ub(t), uc(t) are strategies of the players describing their efforts directed to the development of social partnership relations; up, ub, uc domains 380 vladimir k. dyachenko:computer simulation of social partnership in the system of ... of feasible strategies; jp , jb, jc payoff functionals of the players; gp , gb, gc instantaneous payoff functions; p = {p1, . . . , pr} a finite set of employers; university (only one is considered); c = {c1, . . . , cs} a finite set of students; = 4 (period = 5 years). the strategies determine a share of the annual budget assigned by a player to the needs of continuing professional education (cpe): upi (t) a share of the annual budget assigned to cpe programs by an employer, upi = [0, 1]; ucj(t) a share of the annual budget assigned to cpe programs by a student, ucj = [0, 1]; ub(t) a share of the annual budget assigned to cpe programs by the university, ub = [0, 1]. in the current investigation five employers and ten students are considered. the strategies of generalized players in any moment of time are calculated as an arithmetic mean of the strategies of those agents, namely: strategy of p (employer): up (t) = 1 r r∑ i=1 upi(t);up (t) ∈ up , r = 5 (3) strategy of (student) uc(t) = 1 s s∑ j=1 ucj (t);uc(t) ∈ uc , s = 10 (4) strategy of (university): ub(t);ub(t) ∈ ub (5) the state variable of the model considered as a time function x(t) characterizes a quantitative factor which determines relations in the social partnership system; f − a function of the system dynamics depending on the players strategies. it is assumed that the function of system dynamics increases in respect of all arguments (the efforts of players positively influence to the results of social partnership). to give the system dynamics the modified verhulst-pearl model is used, i.e. f is taken in the form f(x(t), up (t), ub(t), uc(t)) = h(up (t), ub(t), uc(t))x(t)(1− x(t) k ) (6) where k maximal feasible value of the state variable in the given conditions; h function of impact of the players’ strategies, namely: h(up (t), ub(t), uc(t)) = 3∑ i=1 aiui(t);ai ≥ 0; 3∑ i=1 ai =1; i = p,b,c; (7) advances in systems science and applications (2014) vol.14 no.4 381 ai relative weights of the strategies. for the estimation of the relative weights the following considerations are used. the sum of the weights is equal to 1, and all weights are positive. the most important influence is made by students as key agents of the cpe system. the two other weights are approximately equal to each other and dont differ from the former weight (students) too significantly. so, the weights are chosen as follows (table 1): table 1 relative weights of the strategies (ai) relative weight i employer university student ai 0.3 0.3 0.4 accordingly to the research hypothesis it is rational to consider two variants of parameterization of the payoff functions. 1) egoistic approach. speaking about the real situation in the cpe system it is natural to suppose that gi decreases on ui and increases on other arguments (“ free-rider principle ”). this variant determines an egoistic approach when all players save their personal efforts. thus, a problem of coordination of private (efforts saving) and common (social partnership development) interests in the cpe system arises. in this case the payoff functions are parameterized as follows: gi(up (t), ub(t), uc(t), x(t)) = bjuj(t) + bkuk(t) + bxx(t) 1 + biui(t) ; i, j, k = p,b,c (8) bi relative weights of the strategies. 2) cooperative approach. this parameterization describes a desirable (ideal) state of the relationships of social partnership in the cpe system when its elements voluntarily and consciously contribute to the development of social partnership relationships. this variant defines a cooperative approach when payoff functions become increasing on all arguments: gi(up (t), ub(t), uc(t), x(t)) = bipup (t) + bibub(t) + bicuc(t) + bixx(t),i = p,b,c (9) bj i a relative value of the factor for the player i (i=p , b, c; j=p , b, c, x). 3 identification of the model suppose that the state variable x(t) characterizes a level of professional training of the students. for the estimation of the initial training level a known approach 382 vladimir k. dyachenko:computer simulation of social partnership in the system of ... proposed by donald kirkpatrick[14] is used. d.kirkpatrick has developed a model of evaluation of the training effectiveness which considers four levels: 1) reaction: to what degree participants react favorably to the training. for the estimation the results of sociological polls conducted among students of the southern federal university are used. 2) learning: to what degree participants acquire the intended knowledge, skills, attitudes, confidence and commitment based on their participation in a training event. here the data about the students progress together with data characterizing the material base of education are considered. 3) behavior: to what degree participants apply what they learned during training when they are back on the job. the results of polls conducted among employers in the rostov region are used. 4) results: to what degree targeted outcomes occur as a result of the training event and subsequent reinforcement. the data of polls among employers on the topic “a level of professional knowledge and skills of the graduates” are used. therefore, an initial value of the professional training level is evaluated by an evident formula x0 = 1 4 4∑ i=1 xi (10) to determine the values of xi the results of polls are interpreted as follows. the answers “satisfied”, “rather satisfied”, “high” and “very high” are treated as satisfaction and their values for a question are added. the answer “i dont know” means that the respondent may have either positive or negative opinion. supposing their ratio as 1:1, the value is taken equal to 1/2 in this case. the summing up gives an array of the values of xi as follows: x = [0.911; 0.823; 0.559; 0.617]. also, the maximal feasible value of the state variable in the given conditions is taken equal to k = 1. the relative weights are identified accordingly to the following reasoning. 1) egoistic approach. this approach supposes an economy of the personal efforts (“free-rider problem”). the principal role belongs to the student whose efforts and professional training are of crucial importance. nevertheless, other players are important too. the next is university as the training base and then employer who determines the requirements on a labor market. if this approach is chosen then the players are more interested in resource economy than in the improvement of professional training. the respective relative weights of influence factors bi are presented in table 2. 2) cooperative approach. student. the aim of partnership for the student is knowledge acquisition and improved professional training. respectively, the factor of professional level is the principal one for the student and it has the biggest value. the factor of personal advances in systems science and applications (2014) vol.14 no.4 383 table 2 relative weights of influence factors (bi) relative weight employer university student professional level bi 0.25 0.3 0.4 0.1 table 3 relative values of the factor j for the player i, ( bij , i = p,b,c; j = p,b,c, x) relative values bj i employer university student professional level employer 2/8 1/8 2/8 3/8 university 1/8 3/8 2/8 2/8 student 1/7 1/7 2/7 3/7 effort is evaluated a bit smaller. the two other factors (significance of university and employer) are equal and have minimal values. university. the university gives a maximal value to the significance of its efforts because they provide a base for the professional training. the significance of student and professional level are evaluated by the university with equal and a bit smaller values. at last, the minimal value goes to the employer who is as a rule not very active participant of the training process. employer. this player is interested above all in the professional level of the potential employees, and efforts of the student and his own receive equal values. the last place belongs to the university because its training programs rarely satisfy the employers needs completely. the relative values of the factors are presented in table 3. 4 planning and implementation of the computer simulation experiments the model investigation was conducted by computer simulation on the base of scenario method [11]. the scenarios are formed according to plausible behavior patterns of the players. for simplicity the arithmetic mean strategies are described. it is supposed that for all scenarios for any moment of time the values of strategies are equal: up (t) = uc(t) = ub(t) = uo(t), t = 0, ..., 4. six scenarios are considered: 1) maximal (max) one corresponds to the maximal possible financing when the whole budget of a player is assigned for training: uo(t) = 1; 2) medium (med) one assigns a half of the budget for training: uo(t) = 0.5; 3) minimal (min) allows for training only a small part of the budget: uo(t) = 0.2; 4) absence of financing (abs) is clear: uo(t) = 0; 5) decreasing of financing (dec) means that initially an eminent part of the 384 vladimir k. dyachenko:computer simulation of social partnership in the system of ... players budget is assigned for training but then the share decreases to a small value, namely: uo(t) = 0.8− 0.15t; 6) increasing of financing (inc) describes the opposite strategies: uo(t) = 0.2 + 0.15t. thus, in the former four scenarios the strategies are constant in time, while in the fifth and sixth scenarios the shares of financing are time-dependent. 5 processing and analysis of the modeling results the processing of modeling data includes a comparative analysis of the graphs of state variable and payoff functionals for different scenarios. the graphs of the employer’s payoff functional for different scenarios are shown in fig.1-fig.2. the functional j1p(t) corresponds to the egoistic approach, and j2p(t) to the cooperative one. as it could be expected, the best results are achieved for the maximal scenario fig.1 comparison of graphs of the employers payoff functional for the scenarios “dec” and “inc” , and the worst results for the minimal one. the values of payoff functionals and state variable decrease in the following order of scenarios: maximal, decreasing, medium, increasing, minimal, absence. in fig.3 it is shown the comparative dynamics of the state variable (professional level) respec-tively to the six scenarios. the dynamics are described by the verhulst-pearl function. let‘s notice that for the decreasing scenario the values of payoff functionals are initially big but then they decrease in time, and for the increasing scenario the picture is opposite. it means that if the players make an eminent contribution from the very beginning then they can create a good base which allows for a certain decreasing in the future, and the level of development of social partnership advances in systems science and applications (2014) vol.14 no.4 385 fig.2 comparison of graphs of the employers payoff functional for the scenarios “max” and “abs” fig.3 comparison of graphs of the state variable for different scenarios will be still satisfactory. the comparative analysis shows that all players win from the social partnership. however, for any specific scenario the payoff functional of a player can achieve both maximal and minimal value in respect to the payoffs of other players depending on the players strategy. so, the strategic choice is critical for a player. for two variants of payoff functionals (egoistic and cooperative approach) a comparison of the payoffs is fulfilled. a difference between the approaches is clearly observable in time. the difference grows when financing increases, and the comparison demonstrates advantages of a higher level of the social integration. in the case of cooperation payoffs are greater than in the case of egoism: for the employer in 3.75 times (absence of financing) and 1.64 times (maximal scenario); for the student in 4.29 and in 2.04 times for the same scenarios respec386 vladimir k. dyachenko:computer simulation of social partnership in the system of ... tively; for the university in 2.5 times (absence of financing) and in 1.76 times (medium scenario). 6 conclusions in the paper a mathematical model of the social partnership in the cpe system is built using differential games techniques and investigated by computer simulation. the model identification is made on the base of sociological polls conducted in the southern federal university and rostov region (russia). one of the main results is a confirmation of the thesis concerning necessity to join efforts of the subjects of social partnership. in this case all considered scenarios lead to better results than in the egoistic approach. it is clear that more financing allows for more convincible results. acknowledgments the work is supported by russian foundation for humanities, project #14-0300236 references [1] frege c.m. (1999), “social partnership at work: workplace relations in post-unification germany”, routledge, pp.272. [2] eds. m. m. seitanidi, a. crane. (2013), “social partnership and responsible business: a research handbook”, routledge, pp.432. [3] keith n. (2011), “engaging in social partnerships: a professional guide for successful collaboration in higher education”, routledge, pp.288. [4] siegel d. (2010), “organizing for social partnerships: higher education in cross-sector collaboration”, routledge, pp.224. [5] talman d., yang z. (2011), “a model of partnership formation”, journal of mathematical economics, no.47, pp.206-212. [6] fandel g., giese a., mohn b. (2012), “measuring synergy effects of a public social private partner-ship (pspp) project ”, int. j. production economics, no.140, pp.815-824. [7] wilson e.j., bunn m.d., (2010), “savage g.t. anatomy of a social partnership: a stakeholder pers-pective ”, industrial marketing management, no.39, pp.76-90 advances in systems science and applications (2014) vol.14 no.4 387 [8] diaconu m., pandelica a. (2012), “the partnership relationship between economic academic and business environment, component of modern university marketing orientation ”, procedia social and behavioral sciences, no.62, pp.722-727. [9] zaharatul a.a.z., abd shukor h., ir. ghazari a.a. (2012), “smart tripartite partnership: polytechnic industry student ”, procedia social and behavioral sciences, no.31), pp.517-521. [10] dockner e., jorgensen s., long n.v., sorger g. (2000), “differential games in economics and man-agement science”, cambridge university press, pp.382. [11] law a.m., kelton w.d. (2000), “simulation modeling and analysis”, mcgraw-hill, pp.784. [12] ougolnitsky g. (2011), “sustainable management”, nova science publishers, pp.287. [13] antonenko a.v., ougolnitsky g.a. (2013), “static models of corruption in hierarchical systems”, advances in systems science and applications. vol.13, no.1, pp.32-46. [14] kirkpatrick d.l., kirkpatrick j.d. (2007), “implementing the four levels: a practical guide for ef-fective evolution of training programs”, berrett koeler publishers, pp.153. corresponding author authors can be contacted at: ougoln@gmail.com. adv syst sci appl 2017; 17(2); 29-42 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/483 copyright ©2017 assa. adv. in systems science and appl. (2017) attitude to entrepreneurship in russia: three-dimensional institutional approach irina a. petrovskaya, sergey m. zaverskiy and elena s. kiseleva lomonosov moscow state university business school, 1-44 leninskie gory, moscow 119234, russia. e-mail: petrovskaya@mgubs.ru abstract. this paper aims to analyze the impediments to the development of entrepreneurship in russia from the institutional perspective. to describe the institutional environment we use a concept of a three-dimensional institutional profile which classifies the institutions into three types: regulatory, cognitive and normative. these three dimensions imply three bases of legitimacy: entrepreneurship can be legitimized if it conforms to legal requirements (regulatory dimension), if it is seen as legitimate through a common frame of reference (cognitive dimension) and if it conforms to the existent moral base (normative dimension). we argue that one of the impediments to entrepreneurship development in russia is that it is not seen as legitimate enough by the society at large. we explore the foundations for this through the regulatory dimension (the dynamic of the legal legitimation of entrepreneurial activity from the soviet epoch to the present times), in the cognitive dimension (the stereotype of entrepreneur and its origins), and in the normative dimension (basic assumptions which relate to the fundamental moral dimensions of entrepreneurial activity: assumptions about money, wealth, and work). key words: russia; entrepreneurship; institutional environment 1. introduction at the end of xx century russia faced a new reality which brought the new concepts, new words and new meanings. “market economy”, “free trade”, “competition”, “entrepreneurship” and “entrepreneur” – these new terms were understood by many very vaguely, and the perception of their meaning was largely shaped in the preceding historical period. russia was not the only country that faced this situation, and the variations across the success levels of economic transition in the ex-communist european countries show that the transition to market economy is highly dependent on the historical and social context. in this paper, we try to look at the context where one of the key actors of the market system – the entrepreneur – exists in russia. the market economy in general is grounded in the set of fundamental beliefs about the benefits of free market and competition, but the perceived meanings of these words may be different and the market actors may behave according to their own understanding. thus different elements of market economy need to be considered not only from the economic perspective, but also through the lens of deep-lying beliefs about what is good and what is bad. in this respect, the entrepreneurial activity should be also placed in this context if we attempt to explore its drivers and impediments. in russia entrepreneurial activity is lower as compared to countries with the same level of economic development. the global entrepreneurship monitor shows that in russia in 2014 30 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) (latest available data) entrepreneurial activity measured as a share of population at ages 18-64 who are early-stage or established entrepreneurs was 8,6%, while for instance in argentina it was 25,5%, in brazil – 34,3%, in slovakia – 18,7%, in poland – 16,5% [1]. since entrepreneurial activity is one of the contributors to economic performance, the issue of what aids or impedes entrepreneurship has received much attention. we approach this issue from the institutional perspective [2] which implies that the level of entrepreneurial activity in a country is influenced by its institutional environment. north [2] defines institutions as formal and informal rules of the game in a society; they act as constraints and guidelines for human behavior, and their main role is to reduce uncertainty in human interaction. north’s definition of institutions is a very broad one, and consequently the research on entrepreneurship and institutions adopts multiple and varying views of what institutions are. one stream of research focuses on formal institutions and economic environment, such as secure property rights [3]. another stream of research focuses on culture [4-6] and is largely based on the models for cross-cultural comparison such as developed by hofstede [7] and inglehart [8]. however, the complexity of the issue calls for the broader framework for its analysis. therefore in this paper we will employ a concept of a threedimensional institutional profile [9], which follows scott’s concept of institutional pillars [10] and allows for analyzing country institutional environment in the context of entrepreneurship [11]. the concept of a three-dimensional institutional profile classifies the institutions into three types: regulatory (laws), cognitive (cognitive categories) and normative (values, beliefs and other culture-related notions), and therefore makes it possible to encompass a variety of institutions, including both formal and informal. the existent institutional research on entrepreneurship mostly focuses on how institutions influence entrepreneurial activity [12]. however, the focus of the present study is not on how institutions influence entrepreneurial activity per se, but the perception of entrepreneur and entrepreneurship in the society. following etzioni’s concept of legitimation of entrepreneurship [13], we argue that one of the impediments to entrepreneurship development in russia is that it is not seen as legitimate enough by the society at large. the notion that negative attitude to entrepreneurship affects its development in russia is voiced largely by business experts [14]. however, this issue didn’t receive sufficient coverage in the academic research. this paper aims to address this gap. 2. three-dimensional institutional profile as the legitimation framework etzioni [13] argues that the legitimacy of entrepreneurship in a society is a key factor that determines the level of entrepreneurship: “the extent to which entrepreneurship is legitimate, the demand for it is higher; the supply of entrepreneurship is higher; and more resources are allocated to the entrepreneurial function”. legitimacy is defined as a “generalized perception or assumption that the actions of an entity are desirable, proper, or appropriate within some socially constructed system of norms, values, beliefs, and definitions” [15]. accordingly, legitimation occurs through placing a phenomena within a framework through which it is viewed as right and proper [16]. the three dimensions of the institutional profile imply three bases of legitimacy [10,17]. entrepreneurship can be legitimized if it conforms to legal requirements (regulatory dimension), if it is seen as legitimate through a common frame of reference (cognitive dimension) and if it conforms to the existent moral base (normative dimension). in order to explore these sources of attitude to entrepreneurship in russia: three-dimensional institutional approach 31 copyright ©2017 assa. adv. in systems science and appl. (2017) legitimation, it is necessary to specify the contents of each dimension of the institutional profile. since the set of regulatory, cognitive and normative institutions within the institutional profile is issue-based [9], this implies that the contents of each dimension are defined by researchers according to the specifics of their study, and such is indeed the case [11,18-21]. adapting this framework to the aims of our research, we will specify the contents of each dimension as follows. the regulatory component includes the “existing laws and rules in a particular national environment that promote certain types of behaviors and restrict others” [9]. as the detailed revision of the russian legislation is not the aim of this paper, within this dimension we will focus on the two specific points: the dynamics of the legal legitimation of entrepreneurial activity from the soviet epoch to the present times, and the policy statements made by the authorities regarding the current attitude to business legitimation. the cognitive component in the context of entrepreneurship was defined as “the widely shared social knowledge and cognitive categories (for instance, schemata, stereotypes) used by the people in a given country that influence the way a particular phenomenon is categorized and interpreted” [9]. following this broad definition, we will focus specifically on the stereotypes. stereotypes are psychological representations of the characteristics of people that belong to particular groups which serve as aids to cognition by simplifying the person perception process [22]. consequently, in our study of this component of the institutional environment we will focus on the stereotype of entrepreneur defined as a set of widely-shared beliefs about his or her personal attributes and qualities. the normative dimension of the institutional environment reflects the values, beliefs, norms, and assumptions about human nature and human behavior held by the individuals in a given country [9]. the research on connection between values and various dimensions of entrepreneurial phenomena is extensive [4-6] and includes russian culture as well [23,24]. as mentioned above, this stream of research is mostly based on the models for cross-cultural comparison such as developed by hofstede [7], inglehart [8] and schwartz [25]. the application of these frameworks to the russian culture was researched by many authors [26-30], so we do not see the need to further explore this perspective. instead, we will build our discussion of the normative dimension around the concept of basic assumptions. trompenaars and hampdenturner [31] and schein [32] place the basic assumptions at the deepest layer of culture. schein [32] argues that when value leads to a certain behavior which solves the problem, it is transformed into “underlying assumption about how things really are”. these are assumptions about relationship to environment, nature of reality, time and space, nature of human nature, nature of human activity, and nature of human relationships – a classification very similar to the value orientations of kluckhohn and strodtbeck [33]. trompenaars and hampden-turner [31] describe basic assumption as “absolute presupposition about life”, “what is taken for granted, unquestioned reality”. there is no clear boundary in the existing literature between values and basic assumptions. however, since the domain of values rests mostly with cross-cultural research on entrepreneurship (etic approach), for the study of the meanings implicit to the russian culture we will employ the emic approach [34] and use the term “basic assumption” to avoid conceptual confusion. in view of our focus on entrepreneurship, we will discuss the assumptions which relate to the fundamental moral dimensions of entrepreneurial activity: assumptions about money and wealth, and about work as a means to acquire it. the rationale behind this set of assumptions lies in the notion of entrepreneurship and business. in essence, entrepreneurial behavior is aimed at profit generation [35] – the same can be said about business behavior in general. therefore the concept of money becomes central to the discussion of legitimation of entrepreneurship: if money-seeking 32 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) is encouraged in the specific culture, we may expect that the activity aimed at acquisition of money will be legitimized. the important point here is that money is not just the medium of exchange – it is a symbol imbued with meaning [36]; money and wealth signify deepseated and complex values [37]. the idea that fundamental assumptions about the morality of striving for profit and wealth accumulation, as well as work ethic are connected to the entrepreneurial activity can be traced back to weber [4]. therefore the assumptions about money, wealth and work will be the focal point of our discussion of the normative dimension of the legitimation of entrepreneurship. finally, we need to add a caveat that the three institutional dimensions of legitimation are in interplay. the interaction of normative and regulatory components was demonstrated by edelman [38], of cognitive and normative dimensions – by aldrich & fiol [39]. cognitive dimension may interact with the regulatory one – e.g. the way of treating entrepreneurs as criminals can lead to the creation of the cognitive scheme that implies that entrepreneurs are indeed criminals. although for the purposes of the present study we will draw boundaries between the different components of the institutional profile, it is nevertheless important to keep in mind that they do not exist separately from each other. 3. attitude to entrepreneurship in the society attitudes to entrepreneurship in a society influence the level of entrepreneurial activity, determine peoples’ intentions to become entrepreneurs and influence the business climate. in this section we will explore the specifics of the general attitude to entrepreneurship in russia. as a method of research we use comparative analysis of sociological data revealing respondents’ attitudes to entrepreneurship activity and entrepreneurs themselves, respondents’ opinions about main characteristics of entrepreneurs and their role in economic and social development. we draw on the data from the surveys conducted by the leading russian opinion research companies such as levada-center and fom. samples of these surveys are representative for the russian adult population as a whole (at age of 18 and older). we also use data of international surveys on entrepreneurship, such as global entrepreneurship monitor and eurobarometer. in the soviet era entrepreneurial activity was a criminal offence – private property and private business became legitimate only in the late 1980s, and due to the rapid transition from command to market economy the attitude of russians towards entrepreneurs and entrepreneurial activity is now contradictory. fom surveys conducted in december, 2001 (10 years after the dissolution of ussr) and in march 2013, show that the majority of respondents reported having positive attitude towards entrepreneurs (58% in 2001 and 68% in 2013), and the recent data indicates that 42% of population agrees that entrepreneurship is beneficial for the country [40]. eurobarometer surveys [41] show that a very similar share of respondents in russia, european union and the usa agree that entrepreneurs bring economic benefits through creating jobs (87% of respondents in eu27, 88% in the usa and 89% in russia). if the public surveys show that the attitude to entrepreneurship is quite positive – much less positive than in the usa, but still comparable to eu27 – why the concerns about the negative attitude to entrepreneurship are voiced by the business itself [14]? to answer this question we need to look deeper into the subject and consider not only on the general attitude to entrepreneurship and to its outcomes, but the entrepreneurial behavior. in the fom 2001 survey, when asked an open question about “what do entrepreneurs do? what is their job about?”, the significant part of respondents characterized entrepreneurship attitude to entrepreneurship in russia: three-dimensional institutional approach 33 copyright ©2017 assa. adv. in systems science and appl. (2017) largely as a selfish profit-seeking activity. among the respondents 26% think that entrepreneurs make money: they “grab money”, “coin money”, “enrich themselves”. another 23% of the respondents point out that entrepreneurs speculate, repurchase goods: “buy cheaper, sell at higher price”, “in soviet times it was called “speculation”, “rip-off the consumers”. at the same time positive evaluation of entrepreneurial activity is less common. only 6% of respondents point out that entrepreneurs “launch manufacturing”, 5% – that entrepreneurs “develop the economy”, 2% – “work hard” and 2% – “create jobs” [42]. the eurobarometer survey results [41] also indicate that when it comes to entrepreneurial behavior, the inclination to ascribe negative properties to entrepreneurs in russia becomes more apparent. the share of respondents agreeing that “entrepreneurs take advantage of other people’s work” in russia exceeds the eu27 and the usa levels and amounts to 76%, and the same applies to the statement “entrepreneurs think only about their own pockets”. levada-center has conducted several surveys in 1991-2013 when the respondents were asked to describe the qualities of the russian and western entrepreneurs [43], and these indeed support the concerns about the negative image of an entrepreneur: according to the survey results, the respondents see the primary quality of the russian entrepreneur as the “itch for gains”, or profitseeking (with negative connotations), followed by “inclination for deception and fraud”, “cornercutting” and “unwillingness to work honestly”. moreover, when compared to the qualities attributed to the western entrepreneur, the negative image of the russian entrepreneur becomes even more pronounced: while a western entrepreneur is characterized by the business acumen (51%), rationality (44%) and competence (39%), the russian entrepreneur looks for gains mostly of the illicit quality (53%) [43]. thus, even though the public surveys show the general attitude to entrepreneurship in the russian society being rather positive than negative, the image of entrepreneurs, the motives of their behavior and their personal qualities are characterized in a primarily negative way, which supports the claim that entrepreneurship is not seen as a legitimate activity. in the following section we will discuss the sources of legitimizing entrepreneurship through the perspective of a three-dimensional institutional profile. 4. three-dimensional approach to legitimization of entrepreneurship in russia as discussed above, the three dimensions, or pillars, of the institutional profile imply three bases of legitimacy. entrepreneurship is legitimized through laws and regulations (regulatory dimension), through a common frame of reference (cognitive dimension) and through the existent moral base (normative dimension). we will start with the regulatory dimension and focus on the development of the legal legitimation of entrepreneurial activity from the soviet epoch to the present times, and on the policy statements made by the authorities regarding the current attitude to business legitimation. 4.1. the regulatory dimension in the soviet times entrepreneurship was officially qualified as a criminal offence. the criminal code of russian soviet federal socialist republic of 1960 imposed the punishment for “private entrepreneurial activity” of up to 5 years of imprisonment (article 153). profiteering (article 154) was punished with imprisonment for up to 7 years. currency transactions implied 34 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) punishment ranging from imprisonment for up to 15 years to death penalty (by the special law) – actually, that was an equivalent of punishment for aggravated homicide (article 102). the legal emergence of entrepreneurs became possible only after 1986, when selfemployment and creation of joint ventures with foreign capital were allowed. after that it became possible to establish cooperative organizations (“cooperatives”) – prototypes of private companies able to rent their premises from state enterprises. at the same time, the prosecution of “underground entrepreneurs” was strengthened. however, those entrepreneurs who became rich overnight were distrusted by the society due to the dishonest nature of these earnings. partly, this was also facilitated by government policies [44]. the sources of immediate enrichment were often connected with certain regulations, providing exceptional opportunities to a small number of admitted individuals. the examples are provision of export licenses, import subsidies, preferential industrial loans and restrictions on domestic prices in some sectors. however, despite the formal legitimation of entrepreneurship in the legal sense, an ambiguous attitude of public authorities to business and, in particular, to entrepreneurs, was still publicly shown. even in 1990s, at the time of large-scale liberalization and promotion of entrepreneurship (at least formal), the state acted as a “grabbing hand” [45]. an entrepreneur in russia in 1990s de-facto was treated by government officials as an object for contrived arraignments and extortions, and not as an object for care and support. as a result, many entrepreneurs could not make their business in a law-abiding way, and had to “go into the shadows” [46]. since 2000 a number of initiatives were implemented to improve the image of an entrepreneur in russia and enhance the government support of business (at least in terms of lowering the administrative barriers). however, despite these initiatives, the laws that regulate business activity currently appear to inhibit it to a certain extent, primarily because of their inconsistency: “even when laws and regulations do not obstruct firms’ entry and exit, application and enforcement of rules often remain inconsistent” [47]. thus the analysis of the regulatory dimension of the institutional profile shows that even though the entrepreneurship is officially legalized, the issue with the regulatory legitimation is an environment which makes it difficult for entrepreneurs to follow the existent laws, thus often forcing them to become illegitimate by default. 4.2. the cognitive dimension in this study the cognitive dimension of the institutional profile deals with the stereotype of an entrepreneur – a set of widely-shared beliefs about his or her personal attributes and qualities. the beliefs about the qualities of an entrepreneur and the negative stereotype of an entrepreneur have already been discussed above. in the following section we will look at how this stereotype developed and how it was transferred to the population. specifically, two aspects of stereotype development should be considered: the soviet ideological attack at the entrepreneurs, and the stereotypes which developed in the transition period. 4.2.1. entrepreneurship and the soviet ideology stereotypes are formed largely through education. in the soviet school education the ideological tasks dominated over the educational ones [48]. a special role in the spread of ideology in the soviet time was played by primary and secondary schools. elements of ideology were present in many academic disciplines, but the main role in this process was given to history and literature. the history textbooks in the soviet union were the mirror of public policy [49]. the idea that history textbooks should play the central role in the ideological construction was formulated in attitude to entrepreneurship in russia: three-dimensional institutional approach 35 copyright ©2017 assa. adv. in systems science and appl. (2017) the ussr in 1930. in march 1934, the current state of history teaching in schools has been the subject of two sessions of the politburo, and stalin was personally involved in the editing of general history textbook [50]. one of the reviews on a history textbook, written in 1948, was rather typical and stated that “the idea of irreconcilable class struggle of the proletariat against all its enemies should permeate all textbook exposition. as a result of modern history studies, soviet schoolchildren should feel the hatred of capitalism and its’ political leaders, should feel contempt and disgust for the social-democratic lackeys of capitalism” [51]. later, in the post-stalin times, the situation has changed only slightly – in the resolution of the central committee of the cpsu and the ussr council of ministers of 1959 it was stated that the course of history in the secondary school should form the belief in the inevitable collapse of capitalism and the victory of communism. in line with this, one of the central places of the soviet history textbooks was devoted to a crisis of capitalism, and futility of this way of development was strongly substantiated. a capitalist was portrayed almost exclusively as an immoral oppressor of the working people. in line with this, it was stated that only centralized, authoritarian state is necessary for ussr. noteworthy, the key findings on these topics in new history textbooks of the post-soviet era were largely preserved, although the grounds have changed [49]. teaching of literature also occupied a significant place in soviet schools. this was largely due to the fact that russian literature was a synthetic phenomenon, exercising at one time the functions of philosophy, humanities, the social and political platform [52]. accordingly, the image of an entrepreneur as shown in the literature, especially in the classical literature of the xix century, has produced a number of stereotypes that are deeply embedded in the consciousness of many generations. moreover, the educational curricula focused only those works from the variety of classical literature that supported existence of such stereotypes. as a result, the idea of entrepreneurship as a free creativity was undermined. description of business was focused not on the business itself, but only on money acquisition, while business activities were associated with a number of negative characteristics: moral decay, deceit, exploitation of other people, even crime [53]. the origins of entrepreneurial success, vitality and energy of businessmen were seen in the illegibility in achieving the goals – people moved by a passionate desire to create value often found themselves unable to resist the temptation to “accelerate” it through fraud, corruption, forgery [54]. referring to the surveys quoted above [43] this set of characteristics remarkably coincides with the personal qualities attributed to entrepreneurs in the post-soviet times. 4.2.2. the image of entrepreneur in transition period even though in soviet times entrepreneurship was a criminal offence, some individuals were still involved in these activities, and some of them were convicted and sentenced to prison. thus in the post-soviet era it turned out that business experience and spirit of entrepreneurship can be often attributed to ex-prisoners. while being in prison, they had time to develop connections in the criminal world, and thus, later on, the newborn entrepreneurs were surrounded by former criminals. in a situation when property protection and law enforcement were extremely poor, the private businesses had to resolve problems with the help of the so-called "roof" (“krysha”) that provided protection of their lives and interests [55]. as a result, the public perception of business, even after it was legitimized by the authorities, was still associated with illegal and criminally bound activities. the other category of new-born businessmen which emerged in the transition period was former communist party officials, managers of public enterprises and government officials. their 36 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) credibility among ordinary people was rather low at all times, but after legalization of business it turned out that what had previously been considered as public property suddenly became someone's undeserved, but legal private property, which created a strong feeling of unfairness in the general public. this was also stipulated by numerous cases of administration officials legally obtaining a privileged position during the privatization of enterprises. the large-scale fund withdrawals from these enterprises (for the benefit of privatizers) were held directly in front of workers who were paid with significant delay [56]. the privatization process largely affected the public attitude to entrepreneurship. after privatization the majority of russian citizens did not get any income-generating property, in contrast to propaganda of privatization. moreover, they felt themselves robbed. this became especially evident after the “shares for loans” auctions used for privatization of major state enterprises that invited questions among the public – why did somebody become the owner of resources and enterprises that were built by thousands of people and worked for decades almost for free [57]. and even if the new owners paid something for the assets, there still was the question about origins of the money. the media reports about dollar billionaires further led to negative and even hateful attitude to them. “itching for gains” as an attribute of an entrepreneur may be largely traced back to the privatization as well as to the soviet era. thus the complicated history and heritage of the transition period still influences the perceptions of entrepreneurship, and combined with the issues in the regulatory dimension contributes to the persistence of the negative general stereotype of the entrepreneur in the society. 4.3. the normative dimension as mentioned above, the normative dimension of the institutional environment deals with basic assumptions relate to the fundamental moral dimensions of entrepreneurial activity: assumptions about money and wealth, and about work as a means to acquire it. in essence, business behavior is aimed at revenue generation through seeking opportunities and rational allocation and utilization of resources. this contrasts sharply with the superiority of spiritual over material inherent in the russian culture. this opposition is well illustrated in chekhov’s “the cherry orchard”, where lopakhin, whose character doesn’t display any negative traits, nevertheless invites antipathy by his triumph after the acquisition of the cherry garden. his behavior, which makes perfect sense from the business perspective, doesn’t invite any compassion as opposed to the behavior and attitudes of the impractical and passive landowners living in the past. similar opposition can be found in dostoevsky’s protagonist’s speculation on “which is the worst of the two russian ineptitude or the german method of growing rich through honest toil”, followed by stating his preference for the former: “i would rather live a wandering life in tents, … than bow the knee to a german idol” [58]. why does “ineptitude” emerge as a more worthy option than “honest toil”? we argue that the roots of this can be found in the specific attitude to money, wealth and work shaped partly by religion and work ethic. 4.3.1. assumptions about money and wealth the concept of money and wealth is one of the focal points in the discussion of legitimation of entrepreneurship: if money-seeking is encouraged in the specific culture, we may expect that the activity aimed at acquisition of money will be legitimized. in russia, the specific attitude to money and wealth has developed under the influence of the religion and the general course of economic development. these will be considered further on. attitude to entrepreneurship in russia: three-dimensional institutional approach 37 copyright ©2017 assa. adv. in systems science and appl. (2017) the orthodox religion promotes the ideals of simplification and humility. the “mundane way of escape and pilgrimage”, “high estimate of begging and poverty” was preached already in ancient religious poetry [52,59]. accordingly, the entrepreneurs saw their activities not only and not so much directed towards accumulation of wealth, but rather as a kind of a mission entrusted by god or fate. a distinctive feature of orthodoxy is that the owner is not the master of his estates, but is the manager of the god’s belongings that he was given for temporary use during his life [60]. it was typically said that the wealth was granted by god for use, and god will require a report on it, which contributed to the development of philanthropy that was regarded as the fulfillment of a duty [61,60]: by 1900, in moscow there were more donations produced than in paris, berlin and vienna combined [60]. thus money and wealth were not considered as fundamental values, and were treated as perishable and incidental, subject to a high risk of loss. this may have continued into the modern period and manifest itself in the desire to immediately make a handsome fortune [62]. insecurity, lack of guarantees of the irreversibility of reforms, instability in the economy lead to the spread of “a one-time” business psychology, that has nothing to do with the care for reputation or business ethics. 4.3.2. assumptions about work assumptions about work – its meaning and its purpose – are as central to the discussion of entrepreneurship as are the assumptions about money. the soviet period undermined the meaning of work and the incentives for productive work. higher salaries were not necessarily aligned with productivity growth, the rates were normalized, and the employees knew that they get paid primarily for the time spent for work, and not for the result. moreover, motivation was restricted by the limited availability of career opportunities. as noted in the popular saying, “they pretend that they pay us and we pretend that we work” [44]. however, even in the presoviet period work had specific connotations that shaped the assumptions about work and its meaning. the etymology of the word “work” (“trud”) in russian implies that the work is seen as a necessary burden or necessary evil rather than the source of joy or fulfillment: work is ”everything that requires effort, diligence and care, any tension of bodily or mental powers, all that makes tired” [63]. religion and the work ethic it implies also have the fundamental influence on the assumptions about the meaning of work and its purpose. in pre-revolutionary period the work ethic that was dominant among peasants and industrial workers was minimalist or traditional [64], focused on meeting the modest, almost minimal needs for food, clothing, shelter [65]. minimalist consumption standards allow individuals not to worry about the accumulation of wealth [66]. this type of work ethic existed in europe as well, but with the start of industrialization it began to transform – the process which weber associated with the emergence of protestantism [67]. while the protestant ethic, according to weber, suggests that the criteria chosen for salvation is the welfare, and the path to welfare is multiplication of wealth through work [67], the orthodox church suggests that economic activity has nothing to do with the salvation of souls. in orthodoxy, not any work is useful, but only work that contributes to the improvement of the soul. orthodox religion doesn’t denounce work per se – on the contrary, work is seen as a natural mode of life [68]. however, the ultimate goal of any activity, including economic activity, is spiritual improvement, and material well-being is not connected with the prospects of salvation. moreover, the extent of this material well-being is set at the level which provides for meeting the basic needs, and anything exceeding this level is not seen as moral. the situation has partly changed in the soviet era. marxist ideology stated that work “transformed apes into men”, and the love for work was a measure of moral maturity. working 38 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) was a primary activity, and those who were not engaged in public production without good reason were considered as “parasites” and were subject to criminal prosecution [66]. due to this emphasis on working soviet model is sometimes compared with protestantism [69] – however, the important difference was that in terms of the soviet ideology the purpose of work was not to obtain individual benefits, but contribute to meeting the goals of the society, as stipulated in 1977 constitution of the ussr. the assumptions about work and its meaning which constitute the normative dimension of the institutional profile allow an observation that the soviet experiment didn’t develop in a vacuum, but to some extent built on the existent normative environment. it can be seen that some aspects of these assumptions are deep-lying and even based on etymology. recent research suggests that work values have been changing since the start of the transition period, and the changes happen very quickly [70] – specifically, money comes to be seen as a primary reason for work, which is typical for the emerging economies [71]. however, research also suggests that the general attitude to entrepreneurship is defined primarily by the morality of aspiration for wealth [72], which supports the concept of moral legitimation as essential to legitimation of entrepreneurship [13]. 5. conclusion the issue of legitimacy of entrepreneurship in russia dates back many centuries and cannot be attributed only to the soviet period. however, the soviet times delivered probably the strongest blow to the legitimation of entrepreneurship. entrepreneurship was a criminal offense, and the system of education transferred a negative image of entrepreneur to the population. thus currently one of the problems that have to be dealt with is the generational gap in assessing entrepreneurial activity and its role in the economy. while the positive attitude towards entrepreneurs is typical for the young and educated population groups, the older generation sees it mainly through the negative lens [42]. however, the younger generation is placed in the complicated institutional environment which often provides conflicting messages, especially its regulative component. what is needed now is the change regarding all the three dimensions of the institutional profile. clear messages about the legitimacy of entrepreneurship should be sent by the state and the government, and the rules of the game should be set and observed (regulatory dimension). the positive stereotypes of entrepreneurs should be developed and transmitted (cognitive dimension). the normative dimension, which deals with the basic assumptions and with the deepest layers of culture, is the most difficult to change, however it should also be addressed. the globalizing world and the contact with different cultures that it brings may gradually influence the change of these basic assumptions. in any case, what is required for the further development of entrepreneurship is the consistent policy and the consistent effort that recognizes the interconnection between the different dimensions of the institutional environment and their specifics. entrepreneurship is one of the key drivers of the economic growth, and the current level of entrepreneurial activity in russia undermines its potential for diversifying the economy, thus making it susceptible to the fluctuations in oil and gas prices that can be subject to the market manipulations. development of the institutional environment with an aim to support and promote entrepreneurship is thus becoming a key to opening the unlimited dimensions of economic growth [73]. attitude to entrepreneurship in russia: three-dimensional institutional approach 39 copyright ©2017 assa. adv. in systems science and appl. (2017) references [1] global entrepreneurship monitor. (2017). country profiles, 1999-2016. [online]. available: http://www.gemconsortium.org/country-profiles [2] north, d. (1990). institutions, institutional change and economic performance. cambridge, uk: cambridge university press. [3] acemoglu, d. & verdier, t. (1998). property rights, corruption and the allocation of talent: a general equilibrium approach. the economic j., 108(450), 1381–1403 [4] thomas, a. s. & mueller, s. l. (2000). a case for comparative entrepreneurship: assessing the relevance of culture, j. int. business studies, 31(2), 287-301. [5] uhlaner, l. & thurik, r. (2007). postmaterialism influencing total entrepreneurial activity across nations, j. evolutionary economics, 17(2), 161-185. [6] könig, c., steinmetz, h., frese, m., rauch, a. & wang z. m. (2007). scenario-based scales measuring cultural orientations of business owners. j. evolutionary economics, 17(2), 211-239. [7] hofstede, g. (2001). culture's consequences: comparing values, behaviors, institutions and organizations across nations. thousand oaks, ca: sage publications. [8] inglehart, r. (1997). modernization and post-modernization: cultural, economic, and political change in 43 societies. princeton, nj: princeton university press. [9] kostova, t & roth, k. (2002). adoption of an organizational practice by subsidiaries of multinational corporations: institutional and relational effects. academy of management j., 45(1), 215-233. [10] scott, w. r. (2003). institutional carriers: reviewing modes of transporting ideas over time and space and considering their consequences. industrial and corporate change, 12(4), 879-894. [11] busenitz, l., gomez, c. & spencer, j. w. (2000). country institutional profiles: unlocking entrepreneurial phenomena. academy of management j., 43(5), 994-1003. [12] baumol, w. j. (1990). entrepreneurship: productive, unproductive, and destructive. j. political economy, 98(5), 893-921. [13] etzioni, a. (1987). entrepreneurship, adaptation and legitimation: a macro-behavioral perspective. j. of economic behavior and organization, 8(2), 175-189. [14] verkhovskaia, o. r. & dorokhina m. v. (ed.) (2011). global entrepreneurship monitor. russia 2011. [online], available: http://www.gemconsortium.org/docs/2407/gem-russia2011-report [15] suchman, m. c. (1995). managing legitimacy: strategic and institutional approaches. academy of management review, 20(3), 571-610. [16] tyler, t. r. (2006). psychological perspectives on legitimacy and legitimation. annual review of psychology, 57, 375-400. [17] veciana, j. m. & urbano, d. (2008). the institutional approach to entrepreneurship research. introduction. int. entrepreneurship and management j., 4(4), 365-379. [18] manolova, t. s., eunni, r. v. & gyoshev, b. s. (2003). institutional environments for entrepreneurship: evidence from emerging economies in eastern europe. entrepreneurship theory and practice, 32(1), 203–218. [19] eden, l. & miller, s. r. (2004). distance matters: liability of foreignness, institutional distance and ownership strategy, adv. int. management, 16, 187-221. [20] munir, k. a. (2002). being different: how normative and cognitive aspects of institutional environments influence technology transfer, human relations, 55(12), 1403-1428. 40 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) [21] kostova, t. & zaheer, s. (1999). organizational legitimacy under conditions of complexity: the case of the multinational enterprise. academy of management review, 24(1), 64-81. [22] mcgarty, c., yzerbyt, v. y., spears, r. (eds.) (2005). stereotypes as explanations: the formation of meaningful beliefs about social groups. cambridge, uk: cambridge university press. [23] stewart jr., w. h., carland, j. c., carland, j. w., watson, w. e. & sweo, r. (2003). entrepreneurial dispositions and goal orientations: a comparative exploration of united states and russian entrepreneurs, j. small business management, 41(1), 27-46. [24] ardichvili, a. & gasparishvili, a. (2003). russian and georgian entrepreneurs and nonentrepreneurs: a study of value differences, organization studies, 24(1), 29-46. [25] schwartz, s. h. (1992). universals in the content and structure of values: theoretical advances and empirical tests in 20 countries, adv. experimental social psychology, 25(1), 1-65. [26] naumov, a. & puffer, s. (2000). measuring russian culture using hofstede’s dimensions. applied psychology, 49(4), 709-718. [27] naumov, a. & petrovskaia, i. (2011). izmeneniia v rossiiskoi delovoi culture v period 1996-2006 gg. [business culture changes in russia (1996-2006)]. vestnik moskovskogo universiteta. seriia 24: menedzhment, 2, 65-97 [in russian]. [28] latov, iu. & latova, n. (2007). otkrytiia i paradoxy etnometricheskogo analiza rossiiskoi khoziastvennoi kultury po metodike g. hofstede [discoveries and paradoxes of the ethnometric analysis of russian economic culture based on g. hofstede’s method]. mir rossii [world of russia], 16(4), 43-72 [in russian]. [29] grachev, m. v. & bobina, m. a. (2006). russian organizational leadership: lessons from the globe study, international journal of leadership studies, 1(2), 67-79. [30] lebedeva, n. & tatarko, a. (2012) values of russians: the dynamics and relations towards economic attitudes. higher school of economics research paper no. wp brp, 2012, 3, [online]. available: http://www.hse.ru/data/2012/02/18/1263036400/03soc2012.pdf [31] trompenaars, f., hampden-turner, ch. (1993). riding the waves of culture: understanding cultural diversity in business. london, uk: the economist books ltd.. [32] schein, e. h. (1984). coming to a new awareness of organizational culture., sloan management review, 25(2), 3-16. [33] kluckhohn, f. r. & strodtbeck, f. l. (1961). variations in value orientations. westport, ct: greenwood press. [34] morris, m. w., leung, k., ames, d. & lickel, b. (1999). views from inside and outside: integrating emic and etic insights about culture and justice judgment. academy of management review, 24(4), 781-796. [35] hull, d. l., bosley, j. j. & udell, g. g. (1980). renewing the hunt for the heffalump: identifying potential entrepreneurs by personality characteristics. j. small business, 18(1), 11-18. [36] furnham, a. (1984). many sides of the coin: the psychology of money usage. personality and individual differences, 5(5), 501-509. [37] kornhauser, m. e. (1994). morality of money: american attitudes toward wealth and the income tax. indiana law j., 70(1), 119-169. [38] edelman, l. b. (1992). legal ambiguity and symbolic structures: organizational mediation of civil rights law. american j. sociology, 97(6), 1531-1576. attitude to entrepreneurship in russia: three-dimensional institutional approach 41 copyright ©2017 assa. adv. in systems science and appl. (2017) [39] aldrich, h. e. & fiol, c. m. (1994). fools rush in? the institutional context of industry creation. academy of management review, 19(4), 645-670. [40] fom. (2016 july, 4). otnoshenie k predprinimatel’stvu i predprinimateliam [attitude to entrepreneurship and entrepreneurs], fom, [online], available: http://fom.ru/ekonomika/12735 [in russian]. [41] eurobarometer (2012, august). entrepreneurship in the eu and beyond. european commission. flash eurobarometer 354, june august 2012, [online], available: http://ec.europa.eu/public_opinion/flash/fl_354_en.pdf [42] fom. (2001, december 6). predprinimatel’stvo i predprinimateli [entrepreneurship and entrepreneurs]. fom, [online], available: http://bd.fom.ru/report/cat/ec_bus/small_business/dd014631 [in russian]. [43] levada-center. (2013, june 10). rossiiane o predprinimateliakh. press-vypusk [russians about entrepreneurs. press release]. levada-center, [online]. available: http://www.levada.ru/old/10-06-2013/rossiyane-o-predprinimatelyakh [in russian]. [44] zubov, a. (ed.) (2011). istoriia rossii. xx vek: 1939-2007 [history of russia. xx century: 1939-2007]. moscow, russia: astrel, ast, [in russian]. [45] shleifer, a. & vishny, r. (2002). the grabbing hand: government pathologies and their cures. cambridge, ma: harvard university press. [46] latov, iu. (2003). antikapitalisticheskaia mental’nost’ rossiian – bar’er na puti k legalizatsii [anti-capitalist mentality of the russians – barrier on the way to legalization], neprikosnovenniy zapas, 3(29), 64-70 [in russian]. [47] world bank group (2013, october 23) doing business 2013: smarter regulations for small and medium-size enterprises. [online], available: http://www.doingbusiness.org/reports/global-reports/doing-business-2013 [48] dobrenko, e. (1997). formovka sovetskogo chitatelia. sotsialnye i esteticheskie predposylki rezepcii sovetskoi literatury [formation of a soviet reader. social and aesthetical antecedents of receiving soviet literature]. saint petersburg, russia: academicheskii proekt [in russian]. [49] zhukovskaia, i. (2002). chemu my uchim, prepodavaia istoriyu [what we teach at history lessons]. prepodavanie istorii ishchestvoznaniia v shkole, 9, 35-47 [in russian]. [50] chapkovskii, f. (2011). uchebnik istorii i ideologicheskii defitsit [history textbook and ideological deficit]. pro et contra, 2011, january-april, 117-123 [in russian]. [51] tikhonov, v. (2012). ideologicheskie kampanii “pozdnego stalinizma” i shkol’nyi uchebnik po novoi istorii [ideological campaigns of “late stalinism” and school textbook on modern history]. istoriya: elektronnyi nauchno-obrazovetelnyi zhurnal, 5(13),[online], available: http://cliohvit.ru/view_post.php?id=106 [in russian] [52] kondakov, i. (1998). russkaia kultura: kratkii ocherk istorii i teorii [russian culture: brief outline of history and theory]. moscow, russia: knizhnyi dom “universitet”, [in russian]. [53] milekhina, t. (2005). obraz predprinimatelia v russkoi slovesnosti [the image of entrepreneur in russian literature], vestnik ogu, 11, 208-114 [in russian]. [54] zarubina, n. (2003). rossiiskii predprinimatel’ v hudozhestvennoi literature xix – nachala xx veka [russian entrepreneur in belles-lettres of xix and early xx century], obschestvennye nauki i sovremennost’, 1, 101-115, [in russian]. [55] chazan, g. (2011, november 2). russian oligarch seeks to distance himself from rival, the wall street journal, [online], available: http://www.wsj.com/articles/sb10001424052970203707504577012132373208086 http://bd.fom.ru/report/cat/ec_bus/small_business/dd014631 http://www.doingbusiness.org/reports/global-reports/doing-business-2013 http://cliohvit.ru/view_post.php?id=106 42 i. a. petrovskaya, s. m. zaverskiy & e. s. kiseleva copyright ©2017 assa. adv. in systems science and appl. (2017) [56] kapeliushnikov, r. (2001). prichiny zaderzhek zarabotnoi platy: mikroeconomicheskii podhod [the reasons for late payment of wages: microeconomics approach]. problemy prognozirovaniya, 3, 117-132, [in russian]. [57] fom. (2005, january 1) privatizatsiia: kak eto bylo i k chemu eto privelo. opros naseleniia (otchet) [privatization: how it proceeded and what were the results. population survey (report)], fom, [online], available: http://bd.fom.ru/report/cat/pow_pec/dd050322 [in russian]. [58] dostoevski, f. (2009, march 1) the gambler. the project gutenberg ebook. [online], available: http://www.gutenberg.org/files/2197/2197-h/2197-h.htm [59] berdiaev, n. (1997) russkaia ideia; sud’ba rossii [russian idea; russian destiny]. moscow, russia: izdatel’stvo v. shevchuk . [60] sharapov, s. & ulybysheva, m. (2011) bednost’ i bogatstvo. pravoslavnaia etika predprinimatel’stva [poverty and wealth. orthodox ethics of entrepreneurship]. moscow, russia: kovcheg [in russian]. [61] prokhorov, a. (2002). russkaia model’ upravleniia [russian model of management]. moscow, russia: zhurnal “expert”, [in russian]. [62] rakovskaia, o. (1993). sotsial’nye orientiry molodezhi: tendentsii, problemy, perspektivy [social attitudes of youth: trends, problems, perspectives] moscow, russia: nauka, [in russian]. [63] dal, v. (2017) tolkovyi slovar’ zhivogo velikorusskogo iazyka [explanatory dictionary of the live great russian language], [online], available at: http://slovardalja.net [in russian]. [64] tiazhelnikova, v. (2001). otnoshenie k trudu v sovetskii i postsovetskii period [attitude to work during soviet and post-soviet period], sotsial’no ekonomicheskaya transformatsiya v rossii, 131, 99-123 [in russian]. [65] mironov, b. (2001). otnoshenie k trudu v dorevoliutsionnoi rossii [attitude to work in pre-revolutionary russia], soziologicheskie issledovaniia, 10, 99-108 [in russian]. [66] okolskaia, l. (2006). rossiiskaia formula truda: istoricheskii ekskurs [russian formula of labor: historical perspective], chelovek, 4, 16-30 [in russian]. [67] weber, m. (2005). the protestant ethic and the spirit of capitalism. london and new york, ny: routledge . [68] zabaev, i. (2007). motivaziia hoziaistvennoi deiatel’nosti v etike russkogo pravoslaviia [motivation for business activities in ethics of russian orthodoxy], monitoring obschestvennogo mneniya, 1(81), 149-160 [in russian]. [69] magun, v. (1998). rossiyskie trudovye tsennosti: ideologiia i massovoe soznanie [russian work values: ideology and public consciousness], mir rossii, 4, 113-144 [in russian]. [70] magun, v. (2006). dinamika trudovykh tsennostey rossiskikh rabotnikov, 1991-2004 [changes in the work values of russian employees, 1991-2004], rossiiski zhurnal menedzhmenta, 4(4), 45-74 [in russian]. [71] lynn, r. (1991). the secret of the miracle economy: different national attitudes to competitiveness and money. london, uk: social affairs unit. [72] petrovskaya, i., zaverskiy, s. & kiseleva, e. (2016). putting assumptions into words: money and work beliefs and legitimacy of entrepreneurship in russia, european j. int. management, 10(2), 157-180. [73] wen, y. (2016). the making of an economic superpower: unlocking china’s secret of rapid industrialization. singapore: world scientific. http://bd.fom.ru/report/cat/pow_pec/dd050322 http://slovardalja.net/ adv syst sci appl. 2019; 03;11–22 published online at https://ijassa.ipu.ru/index.php/ijassa/article/view/751 hybrid particle swarm optimization and pegasos algorithm for spam email detection lamiaa m. el bakrawy∗ faculty of science, al-azhar university, cairo, egypt e-mail: dr lamiaa el bakrawy@azhar.edu.eg received may 25, 2019; revised september 9, 2019; published october 1, 2019 abstract: email is one of the most popular communication tools for most internet users nowadays. it has become fast and an effective method to share and exchange information all over the world. despite the great advantages of emails, its usage is facing problem which is spam emails. spam emails are the huge presence of bulk and unsolicited emails which are expensive for the companies, consume a huge amount of mail servers, network bandwidth and waste of time. isolating and detecting these emails is known as spam detection. many spam detection methods have been proposed but there is still need to detect the email spam effectively with high accuracy. in this paper, hybrid particle swarm optimization and pegasos algorithm, which is called (psopegasos) is proposed for spam email detection. particle swarm optimization is employed as a search strategy to determine the optimal parameters for pegasos algorithm in order to achieve higher performance. the proposed algorithm has been applied on spambase dataset downloaded from uci machine learning repository. experimental results demonstrate that the proposed algorithm outperforms the performance of all the earlier proposed algorithms, considering the accuracy, recall, precision and f-measure on the same dataset. keywords: particle swarm optimization, pegasos algorithm, email spam, spam detection, accuracy. 1. introduction nowadays, email is one of the most popular communication tools for most internet users because of its free availability and efficiency [1, 2]. email is a method of receiving and sending information over electronic networks such as the internet. however, the major problem is the presence of bulk and unsolicited email which is known as spam. spammer is the person who sends mass quantity of spam emails and collects email addresses from chatrooms, viruses, customer lists and websites. spam email consumes a huge amount of mail servers, network bandwidth and wastes users’ time to remove all spam emails which causes lower productivity. thus, how to isolate and detect spam email in efficient way with high accuracy becomes an important study. spam email detection can be considered as classification problem which is used to detect the spam emails one by one to classify email as spam or non-spam [3]. in recent years, most of the spam detection algorithms based on machine learning techniques is used, but still the reported accuracy requires more work to accomplish better accuracy. sabri et al. in [4] presented continuous learning approach based on artificial neural network (cla ann) for spam email detection. they made core modifications in the input layer of artificial neural network to substitute the useless layers with new favorable layers and to be varied with time. ∗corresponding author: dr lamiaa el bakrawy@azhar.edu.eg 12 l.m. el bakrawy the results showed that applying cla ann using 300 input layers succeeded in achieving 3.668 % false negative and 0.534 % false positive. zhang et al. in [5] presented naive bayes model for spam email detection by applying cost-sensitive multi-objective genetic programming for feature extraction and achieved an accuracy of 79.3%. renuka et al. in [6] proposed spam classification algorithm using hybrid ant colony optimization and naive bayes classifier and applied it on spambase dataset. the accuracy obtained was 84% which indicated that the hybrid algorithm outperformed hybrid genetic algorithm and naive bayes tested on the same dataset. özgür et al. in [7] used artificial neural network and bayesian filter for spam email detection. they considered two artificial neural network structures, multi layer perceptron and single layer and the inputs are specified based on probabilistic and binary models. experimental results for 750 e-mails (410 spams and 340 non-spam), achieved 90% accuracy. temitayo et al. in [8] used genetic algorithm to optimize the support vector machines (svm) classification parameters. the hybrid algorithm achieved 90% accuracy for the testing set. liu et al. in [9] proposed a new learning method (pso-lm ) for process propagation neural networks (pnns) based on particle swarm optimization (pso) and gaussian mixture functions. experiments results showed that applying (pso-lm) on spambase dataset achieved 90.5% accuracy for the testing set which is better than back propagation neural networks (bpnns) and basis function expansion based learning method (bfe-lm). moreover, idris and selamat in [10] presented a hybrid model of negative selection algorithm (nsa) and particle swarm optimization (pso). they worked on spambase dataset and achieved 91.22% accuracy for the testing set. awad and foqaha in [1] proposed a hybrid algorithm of rbf neural network and particle swarm optimization (hc-rbfpso) for spam email classification. they used particle swarm optimization algorithm to optimize the parameters of radial basis function neural networks (rbfnn) based on the evolutionary heuristic search of pso. they divided spambase dataset into 70% training set and 30% testing set. experiments are measured by using a different number of hidden layer starting from 10 to 50. the accuracy obtained was 91.4% for the testing set which was concluded that the hybrid approach had good performance compared to other algorithms tested on the same dataset. olatunji in [11] proposed support vector machines-based model for spam detection. he used a systematic parameter search in order to achieve better spam detection accuracy. the accuracy obtained was 94.06% for the testing set. experimental results show that the proposed scheme outperformed other published algorithms tested on spambase dataset used in this work. considering the performance accuracy achieved till now, there is still need to try to achieve better results on the same dataset. the main aim of this paper is to propose an alternative algorithm that can accomplish a performance higher than previous algorithms. in this paper, hybrid particle swarm optimization and pegasos algorithm (pso-pegasos) is proposed to achieve better accuracy of spam email detection. pegasos algorithm is applied to solve the optimization problem cast by support vector machines (svm) while particle swarm optimization is used as a search strategy to select the optimal parameters (the weights) for pegasos algorithm, which means in each iteration of pso, the weights (w-parameters) are changed based on the fitness function (mean squared error). after running pso algorithm a number of iterations, it will obtain the best optimal w-parameter for pegasos algorithm. in this paper, comparison of performance measures of pegasos and hybrid algorithm (pso-pegasos) for training and testing sets is presented for spam email detection. the rest of this paper is structured as follows: the fundamentals of particle swarm optimization and the principles of the original pegasos algorithm are introduced in section copyright c© 2019 assa. adv. in systems science and appl.(2019) hybrid particle swarm optimization and pegasos algorithm for spam email detection13 2. section 3 describes the details of the proposed algorithm. experimental results and discussions are demonstrated in section 4. finally, section 5 concludes the paper. 2. preliminaries 2.1. particle swarm optimization particle swarm optimization (pso) was invented by kennedy and eberhart in 1995 [12]. pso is a widely used population-based stochastic optimization technique since it has strong global search capability, high convergence speed, high robustness and is conceptually very simple [13, 14, 15]. pso is still attracted the attention of a lot of researchers over nearly a quarter century. particle swarm optimization simulates the social behavior among species such as fish schools, bird flocks. the set of particles represent a population of the possible solutions. in canonical pso algorithm, particles are initialized with a population to get a random solution. then, the particles fly iteratively around in d-dimension search space to search the optimal solution, where the proper fitness function can be calculated according to the problem. each particle is indicated by a row vector ~xi, where i is the index of the particle, and a velocity indicated by ~vi. the best position of the particle (pbest) is indicated by vector ~x#i , and its j-th dimensional value is x#ij , while the best position among the swarm (gbest) is indicated by a vector ~x∗, and its j-th dimensional value is x∗j . in each iteration t, the velocity updating formula of particle is calculated by eq. (2.1) and the position updating formula of particle is determined by the sum of the previous position and the new velocity by eq. (2.2). vij(t+ 1) = { wvij(t) + c1r1(x # ij(t)− xij(t)) +c2r2(x ∗ j(t)− xij(t)) (2.1) xij(t+ 1) = xij(t) + vij(t+ 1). (2.2) where c1 and c2 are nonnegative constants called as learning factors, r1 and r2 are random numbers uniformly distributed in u(0,1) for the j-th dimension of the i-th particle. w is the inertia weight, which can increase the algorithm search capability and control the process of algorithms searching. eq. (2.1) makes each particle tends to move across the design space, considering its own experience, which is the memory of its best fitness function value achieved by the particle in the past, and the experience of its most successful particle in the swarm. in pso algorithm, the particles tend to search the solutions in the problem space with a range [−s, s] to prevent the particle from flying away out of the search space. if the range [−s, s] is not symmetrical, it will be changed to the corresponding symmetrical range and the maximum velocity during one iteration must be limited on the interval [−vmax, vmax] given in eq.(2.3) vij = sign(vij)min(|vij| , vmax). (2.3) where the value of vmax is p× s, with p ∈ [0.1, 1] but vmax is usually selected to be s, i.e. p = 1. the termination criterion for iterations will be determined according to whether the maximum number of iterations or minimum fitness function error is reached. 2.2. pegasos: primal estimated sub-gradient solver for svm pegasos was described and analyzed by shalev-shwartz et al. in [16] for solving the optimization problem cast by support vector machine (svm). it performed a stochastic subgradient descent based on the primal objective by chosen step size carefully to improve copyright c© 2019 assa. adv. in systems science and appl.(2019) 14 l.m. el bakrawy convergence [17, 18, 19]. pegasos has attracted research interest because it has better convergence bounds and robustly convex optimization objective. it uses theory of strongly convex optimization problems and hinge loss instead of the original linear constraints which makes the objective of svm unconstrained. given a binary classification problem with training set s = (xi, yi), (i = 1, . . . , n), where xi is a d-dimensional feature vector and yi = ±1 is the class label. the goal of linear support vector machines is to find a classifier in the following form h(x) = sign(wtx), (2.4) where w is the weight vector which can be learnt from training set to solve the following optimization problem after number of iterations t . min w = λ 2 ‖w‖2 + 1 n ∑ (x,y)∈s l(w, (x, y)), (2.5) where l(w, (x, y)) = max(0, 1− y(w, x)), and λ ≥ 0 is the regularization parameter. in each iteration t , pegasos algorithm aims to update w by choosing a random training set at ⊆ s with size k, where k is the number of training examples used for calculating sub-gradient through the following approximate objective function f(w,at) = λ 2 ‖w‖2 + 1 k ∑ (x,y)∈at l(w, (x, y)). (2.6) the sub-gradient of the approximate objective function f(w,at) at wt is calculated by ∇t = λwt − 1 |at| ∑ (x,y)∈a+ t yx, (2.7) where a+ t is the set of examples when w suffers a non-zero loss. finally, the sub-gradient is used to update the weight by using a step size of ηt = 1 |λt| as wt+1 = wt − ηt∇t, (2.8) where ηt is the learning rate. the last vector wt+1 is the output of pegasos algorithm after number of iterations t . 3. the proposed algorithm in this section, we describe the proposed pso-pegasos algorithm to determine the optimal values of pegasos parameters as shown in figure 3.1. the detailed description is as follows: 3.1. data preprocessing in this paper, the popular and often used corpus benchmark spambase dataset is utilized to classify email as spam or non-spam. the dataset is available in numeric form and the features are frequencies of various characters and words in emails. the main tasks in preprocessing are transformation, reduction, cleaning, integration and normalization. normalization is an important pahse to fast the algorithm, convergence and decrease the influence of imbalance in data. in spambase dataset, normalization is done before running pso-pegasos algorithm. each feature of spambase dataset is normalized in the range [0, 1] through the following function copyright c© 2019 assa. adv. in systems science and appl.(2019) hybrid particle swarm optimization and pegasos algorithm for spam email detection15 figure 3.1: flowchart of the proposed algorithm. a = a−min max−min (3.9) where a is the scaled value, a is the original value, max and min are the maximum and minimum bounds of the feature value. 3.2. pso-pegasos algorithm in this research, hybrid particle swarm optimization and pegasos algorithm (pso-pegasos) is proposed for spam email detection. particle swarm optimization has been utilized to optimize the parameters of pegasos algorithm. in pso each solution is called a particle. fitness function ( mean squared error) is used to evaluate the particles for the optimal solution. particle swarm optimization is used as a search strategy to determine the optimal parameters (weights ) for pegasos algorithm, which means in each iteration of pso, the weights (w-parameters) are updated depending the fitness function. no assumptions are needed about the w-parameter in pegasos algorithm since pso algorithm can help us to identify automatically the best optimal w-parameter (bw) that utilized to obtain the highest classification accuracy for pegasos algorithm. the major steps of the hybrid particle swarm optimization and pegasos algorithm (pso-pegasos) are shown as follows: copyright c© 2019 assa. adv. in systems science and appl.(2019) 16 l.m. el bakrawy 1. initialize the population for w-parameter individuals (particles) in a random manner from spambase dataset. suppose that, each particle swarm position is xi = {ai,j, j = 1, 2, ..., k}, where ai,j is j th w-parameter for the i th individual, k is the number of features (attributes) of spambase dataset and the value of the w-parameter for each individual is vector of k random numbers in range from -10 to 10. 2. initialize velocity of particle swarm optimization randomly in range from -100 to 100 3. calculate the fitness function for each particle which is acquired by pegasos algorithm to classify non-spam and spam emails correctly by fitness =mse = 1 n n∑ i=1 (xi − yi)2. (3.10) where mse is the mean squared error , x is a vector of n predictions, and y is the vector of true values. 4. if the fitness function is better than the best fitness function of the particle (pbest) then the current position will be (pbest) 5. select the best position among all particles (gbest) in current iteration 6. update the velocity of each particle depending on eq. (2.1). 7. update the position of each particle (w-parameter) depending on eq. (2.2). 8. search the the pbest of particle as (w-parameter) of pegasos algorithm in same iteration. 9. repeat steps 3 to 8 until obtaining the best optimal w-parameter (bw) which leads to get the highest accuracy for spam email detection with more exploration in the search space. 4. experimental results and discussions in this paper, the experiments were performed on a system with a 2.40 ghz intel(r) core(tm)i7 processor and 16 gb memory using written codes in matlab 15. 4.1. dataset description in this work, the dataset utilized is spambase dataset which is used to evaluate the proposed algorithm. hopkins et al. [20] presented spambase dataset in their colleagues. it has been collected from uci machine learning repository site. in the spambase dataset, the total email instances is 4601. 1813 from these email instances are characterized as spam (39.4%) and the remaining are non-spam. spambase dataset consists of 57 features and 1 classification attribute, which is the label of class indicating the status of each email instance whether it is spam (1) or non-spam (0). most of the features (1-54) show particular characters or words were repeatedly occurring in an email or not. the features from 55 to 57 present the measurement for length of consecutive capital letters. the definitions of the features can be shown as follows • features from 1 to 48 are real continuous features which are equal to the percentage of words in the e-mail that match word. • features from 49 to 54 are real continuous features which are equal to the percentage of characters in the e-mail that match char. • feature 55 is real continuous feature which is equal to the average length of continuous sequences of capital letters. • feature 56 is an integer continuous feature which is equal to the length of longest continuous sequence of capital letters. • feature 57 is an integer continuous feature which is equal to the total number of capital letters in the e-mail. copyright c© 2019 assa. adv. in systems science and appl.(2019) hybrid particle swarm optimization and pegasos algorithm for spam email detection17 4.2. evaluation measures in this research, the evaluation of the proposed algorithm is carried out based on popular and commonly performance measures such as accuracy, recall, precision, f-measure [21, 22]. the information about these measures is done depending on the confusion matrix presented in table 4.1. table 4.1: confusion matrix actual class spam non-spam predicted class spam tp fp non-spam fn tn brief overview of each performance measure is shown below. • accuracy is defined as the fraction of all emails (non-spam and spam emails) that are classified correctly by the algorithm. it can be represented by the following equation: accuracy = tp+tn fp+fn+tp+tn (4.11) where tp and tn are the number of spam emails and non-spam emails correctly classified, respectively. fp and fn are the number of spam emails and non-spam emails incorrectly classified, respectively. • recall stands for the proportion of spam emails being recognized and can be represented as follows: recall = tp fn+tp (4.12) • precision stands for the fraction of spam emails that are correctly classified as spam. precision = tp fp+tp (4.13) • f-measure (f-score), denotes the harmonic average of precision and recall and can be written as follows: f −measure = 2∗precision∗recall precision+recall (4.14) 4.3. results and discussion the experimental method applied here followed carefully the computational intelligence technique. spambase dataset was first divided into two phases, training set and testing set in the ratio 7:3, respectively. the data was chosen randomly for training and testing sets in order to exclude any particular behavior of the dataset. then, the training set ( 70% of data) was first entered to the algorithm for training and validation and the rest of dataset ( 30% of data) was used to test the algorithm to ensure the performance accuracy of the proposed algorithm. to evaluate the proposed algorithm, the parameters settings for original pegasos algorithm are regularization parameter (λ) with different values from 0.0001 to 0.1 and number of iterations t =1000. the original pegasos algorithm is applied on spambase dataset with different values of λ and the four performance measures are recorded for training and testing sets as shown in table 4.2. it can be observed in table 4.2 that using a small value of λ (0.0001) in a large dataset increases the accuracy, recall, precision and f-measure, respectively for spam detection by pegasos algorithm. according to this result, we fixed the copyright c© 2019 assa. adv. in systems science and appl.(2019) 18 l.m. el bakrawy table 4.2: performance measures for pegasos algorithm with different values of λ for spambase dataset (λ) accuracy recall precision f-measure 0.0001 training set 93.62 % 0.9360 0.9370 0.9360 testing set 92.71 % 0.9270 0.9280 0.9270 0.001 training set 93.04 % 0.9300 0.9310 0.9310 testing set 92.51 % 0.9250 0.9250 0.9250 0.01 training set 92.60 % 0.9260 0.926 0.926 testing set 91.67 % 0.9170 0.9180 0.9160 0.1 training set 90.43 % 0.9040 0.9060 0.903 testing set 89.19 % 0.8920 0.8970 0.8900 value of λ as 0.0001 in proposed algorithm (pso-pegasos). in pso-pegasos, the parameters of particle swarm optimization were set as learning factors c1 = c2 = 1.4, vmax = 4 and inertia weight (w) was linearly decreased from 0.9 to 0.4. the population size was fixed to 20 particles to reduce the computational cost and fast the convergence process of the algorithm. tuning the parameters for particle swarm optimization is important in designing the algorithm. figure 4.2 shows the effect of the number of iterations on the accuracy of the proposed algorithm pso-pegasos using different number of iterations from 5 to 30. as shown in fig. 4.2, we can observe that when the number of iterations was increased, the accuracy was increased until it accomplished an extent (number of iterations =20) at which increasing the number of iterations did not affect the accuracy of the proposed algorithm. figure 4.2: effect of the number of iterations on the accuracy of pso-pegasos algorithm for training and testing sets. according to parameter analysis and paper results, we put number of iterations in pso = 20 to run the proposed algorithm, therefore, the computational cost is small. figures 4.3 and 4.4 show the performance measures accuracy, recall, precision and f-measure of pegasos and pso-pegasos algorithms for training and testing sets, respectively. experimental results in figures 4.3 and 4.4 show that the accuracy of the proposed algorithm (pso-pegasos) for training and testing sets are higher than the accuracy of pegasos algorithm by about 3.39 and 3.48 respectively. it also shows that the proposed algorithm outperforms pegasos copyright c© 2019 assa. adv. in systems science and appl.(2019) hybrid particle swarm optimization and pegasos algorithm for spam email detection19 algorithm in terms of recall, precision and f-measure for training and testing sets due to the existence of particle swarm optimization, which has strong global search capability and high convergence speed to optimal solution. figure 4.3: comparison of performance measures of pegasos and pso-pegasos for training set. figure 4.4: comparison of performance measures of pegasos and pso-pegasos for testing set. finally, in order to indicate that the improvement obtained by the proposed algorithm (pso-pegasos) clearer, its accuracy compared with earlier used algorithms implemented on the same dataset is presented below. experimental results in table 4.3 show that the results of the proposed algorithm outperforms the results of other published classifier called svm-based spam detector [11]. the proposed pso-pegasos presented improvement of 2.13% over svm-based spam detector model, which is the best among the other earlier published classifiers for spam email copyright c© 2019 assa. adv. in systems science and appl.(2019) 20 l.m. el bakrawy table 4.3: comparison of accuracy of the proposed algorithm and other published classifiers on spambase dataset classifiers classification accuracy ga-naive bayes [6] 77% aco-naive baye [6] 84 % pso-lm [9] 90.5 % nsa [10] 68.86 % pso [10] 81.32 % nsa-pso [10] 91.22 % hc-rbfpso [1] 91.4 % svm-based spam detector [11] 94.06% pso-pegasos (proposed) 96.19 % detection. it also presented an accuracy improvement of 4.79 % over a hybrid approach (hcrbfpso), that combines radial basis function neural network (rbfnn) and particle swarm optimization (pso) algorithm [1]. the proposed algorithm also presented improvement of 4.97% over a hybridized negative selection algorithm and particle swarm optimization (nsapso) [10], yet the proposed algorithm in this paper outperformed all the three algorithms nsa-pso, nsa and pso including the hybrid schemes. it also presented an accuracy improvement of 5.69 % over learning method for process neural networks based on particle swarm optimization (pso-lm) [9] and an accuracy improvement of 12.19 % over hybrid ant colony optimization and naive bayes (aconaive bayes), while presenting an accuracy improvement of 19.19% over hybrid genetic algorithm and naive bayes (ganaive bayes) [6]. 5. conclusion primal estimated sub-gradient solver for svm (pegasos) algorithm was utilized to solve the optimization problem cast by support vector machines (svm). it is characterized by better convergence bounds and robustly convex optimization objective. in this paper, hybrid particle swarm optimization (pso) and pegasos algorithm, called (pso-pegasos) is proposed for spam email detection. pso is used to identify automatically the best optimal w-parameter for original pegasos algorithm. the proposed algorithm has been trained and tested using popular and often used spambase dataset, which consists of collection of spam and non-spam emails with 57 features and 1 classification attribute. excremental results indicated that the proposed pso-pegasos algorithm outperformed original pegasos algorithm and other recently published algorithms tested on the same popular dataset used in this paper. the need for more accurate spam email detection method cannot be overemphasized, the proposed pso-pegasos algorithm provides improvement of 2.13% over svm-based spam detector model, which is the best among the previous reported schemes for spam email detection. the results show that pso-pegasos improves the convergence accuracy and it is an effective algorithm, which is a powerful alternative for spam email detection. we can conclude that the aim of this paper has been achieved through training and testing proposed pso-pegasos algorithm on spambase dataset. this algorithm has enabled build an improved spam email detection system based on hybridization of particle swarm optimization and pegasos algorithm. references copyright c© 2019 assa. adv. in systems science and appl.(2019) hybrid particle swarm optimization and pegasos algorithm for spam email detection21 1. awad m, foqaha m (2016) email spam classification using hybrid approach of rbf neural network and particle swarm optimization, international journal of network security and its applications (ijnsa) vol.8, no.4, pp. 17-28. 2. saad o, hassanien a, darwish a, faraj r (2013) a survey of machine learning techniques for spam filtering, ijcsns international journal of computer science and network security, vol.13 no.1, pp. 103-110. 3. zhiwei m, singh m, zaaba z (2017) email spam detection: a method of metaclassifiers stacking, proceedings of the 6th international conference on computing and informatics, icoci, pp. 750-757.757. 4. sabri a, mohammads a, al-shargabi b, hamdeh m (2010) developing new continuous learning approach for spam detection using artificial neural network (cla ann), european journal of scientific research, 42(3), pp. 525-535. 5. zhang y, li h, niranjan m, rockett p (2008) applying costsensitive multiobjective genetic programming to feature extraction for spam e-mail filtering. springer, berlin, pp. 325-336. doi:10.1007/978-3-540-78671-9 28 6. renuka d, visalakshi p, sankar t, improving e-mail spam classification using ant colony optimization algorithm, international journal of computer applications (0975 8887)international conference on innovations in computing techniques (icict 2015) 7. özgür l, güngör t, gürgen f (2004) spam mail detection using artificial neural network and bayesian filter. pp. 505-510. doi:10.1007/978-3-540-28651-6 74. 8. temitayo f, stephen o, abimbola a (2012) hybrid ga-svm for efficient feature selection in e-mail classification, computer engineering and intelligent systems, vol 3, no.3, pp. 17-29 9. liu k, tan y, he x (2010) particle swarm optimization based learning method for process neural networks, in advances in neural networks-isnn 2010 (pp. 280-287). springer berlin heidelberg. 10. idris i, selamat a (2014) improved email spam detection model with negative selection algorithm and particle swarm optimization, applied soft computing, 22, pp. 11-27. 11. olatunji s (2017) improved email spam detection model based on support vector machines, neural computing and applications, 20131(3), pp. 691-699. 12. kennedy j, eberhart r (1995)”particle swarm optimization”. in proceedings international conference on neural networks (icnn 95) perth, australia, pp. 19421948. 13. modares h, alfi a, sistani m (2010) parameter estimation of bilinear systems based on an adaptive particle swarm optimization, engineering applications of artificial intelligence, 23(7), pp. 1105-1111. 14. liu z, li h, zhu p (2019) diversity enhanced particle swarm optimization algorithm and its application in vehicle lightweight design, journal of mechanical science and technology, 33 (2), pp. 695-709. 15. cheng s, lu h, lei x, hi y (2018) a quarter century of particle swarm optimization, complex and intelligent systems, january 2018, accepted, 22 march 2018 16. shalev-shwartz s, singer y, srebro n, (2007) pegasos: primal estimated sub-gradient solver for svm, in proceedings of the 24th international conference on machine learning, pp. 807 -814. 17. shalev-shwartz s, singer y, srebro n, cotter a (2011) extended version: pegasos: primal estimated sub-gradient solver for svm, mathematical programming, series b, 127(1), pp. 3-30, springer and mathematical optimization society. 18. lu s, jin z (2017) improved stochastic gradient descent algorithm for svm, international journal of recent engineering science (ijres), issn 2349-7157, vol 4, pp. 39-42. 19. v. jumutc, x. huang, j. a. k. suykens,(2013) fixed-size pegasos for hinge and pinball loss svm, in proceedings of the 2013 international joint conference on neural networks (ijcnn), pp. 1122-1128, 2013. copyright c© 2019 assa. adv. in systems science and appl.(2019) 22 l.m. el bakrawy 20. hopkins m, reeber e, forman g, suermondt j (1999) spambase dataset. hewlett-packard labs, 1501 page mill rd., palo alto, ca 94304. https://archive.ics.uci.edu/ml/datasets/spambase. 21. el bakrawy l (2017) grey wolf optimization and naive bayes classifier incorporation for heart disease diagnosis, australian journal of basic and applied sciences, 11(7) m, pp. 64-70. 22. liu p, moh t (2016) content based spam e-mail filtering, in proceedings of the international conference on collaboration technologies and systems, 978-1-5090-23004/16 31.00, ieee, pp. 2018-2024. copyright c© 2019 assa. adv. in systems science and appl.(2019) introduction preliminaries particle swarm optimization pegasos: primal estimated sub-gradient solver for svm the proposed algorithm data preprocessing pso-pegasos algorithm experimental results and discussions dataset description evaluation measures results and discussion conclusion adv syst sci appl 2021; 02:8–19 published online at https://ijassa.ipu.ru. uncertain controllability and observability of an optimal control model tolulope latunde1*, adam a. ishaq2, adedayo f. adedotun3, joel o. ajinuhi4, olumuyiwa j. peter5 1department of mathematics, federal university oye-ekiti, oye-ekiti, nigeria 2department of physical sciences, al-hikmah university, ilorin, nigeria 3department of mathematical sciences, olabisi onabanjo university, ago-iwoye, nigeria 4department of mathematics, federal university of technology, minna, nigeria 5 department of mathematics, university of ilorin, ilorin, nigeria abstract: the interest of this paper is to examine the controllability and observability of a control system in the configuration state-space of an uncertain optimal control system. the control system is designed based on the realization of capital asset values where a special case of asset management is modelled and optimized. thus some necessary and sufficient conditions of the controllability and observability of the deterministic systems and the corresponding uncertain systems for the case of the uncertain optimal control system with application in capital asset management are considered. keywords: controllability, observability, uncertain systems, optimal control, capital asset management 1. introduction controllability and observability are important properties in control systems. they represent the ability to move a system around its entire configuration space using certain manipulations. the controllability and observability of a system are mathematical duals that play important roles in control problems such as optimal control. these have played important roles in control theories such as in [1–5]. recently, researchers such as [6–8] and a host of others have been considering controllability and observability problems in dynamic systems. however, most works done in this area have concentrated on the deterministic and stochastic controllability and observability problems. in this work, uncertain controllability and observability of dynamic systems are carried out by formulating a capital asset management control problem for the uncertain dynamic system such that the uncertain dynamic system is limited to systems involving uncertain processes. the choice of uncertainty theory over the conventional probability theory exists when the sample size is small to estimate a probability distribution and degree beliefs are ascertained from experts to work in place of frequency since human beings always over-weigh unlikely events. here, the general controllability and observability for the uncertain system in uncertainty theory are presented based on klamka and mahmudov works. ∗corresponding author: tolulope.latunde@fuoye.edu.ng uncertain controllability and observability of an optimal control model 9 2. preliminaries uncertainty theory is a branch of mathematics for modelling belief degrees. the theory is based on some concepts which may be referred to [9]. for easy interpretation, some of the concepts are given. let γ be a nonempty set and l be a σalgebra over γ such that (γ, l) is a measurable space. each element λ ∈ l is called an event. definition 2.1: a set function m defined on the σ-algebra over l is called an uncertain measure if it satisfies the following axioms: axiom 1. (normality axiom): m{λ} = 1 for the universal set γ. axiom 2. (duality axiom): m{λ} + m{λc} = 1 for any event λ. axiom 3. (subadditivity axiom): for every countable sequence of events, λ1,λ2, . . . , we have m { ∞⋃ i=1 λi } ≤ ∞∑ i=1 m{λi}. axiom 4. (product axiom): let (γk, lk,mk) be uncertainty spaces for k = 1, 2, . . . the product uncertain measure m is an uncertain measure satisfying m { ∞∏ k=1 λk } = min 1≤k≤∞ mk{λk}, where λk are arbitrarily chosen events from lk for k = 1, 2, · · · , respectively; see [9]. definition 2.2: let (γ, l,m) be an uncertainty space and let t be a totally ordered set. an uncertain process is a function xt(γ) from t × (γ, l,m) to the set of real numbers such that {xt ∈ b} is an event for any borel set b of real numbers at each time t; see [10]. definition 2.3: an uncertain process cσ is said to be a liu process if (i) c0 = 0 and almost all sample paths are lipschitz continuous, (ii) cσ has stationary and independent increments, (iii) every increment cs+σ − cs is a normal uncertain variable with expected value 0 and variance σ2. the uncertainty distribution of cσ is φσ(x) = [ 1 + exp ( −πx√ 3σ )]−1 , x ∈ r, (2.1) and the inverse distribution is φ−1 σ (y) = σ √ 3 π ln y 1− y , y ∈ r; (2.2) see [11]. definition 2.4: let ξ be an uncertain variable. then the expected value of ξ is defined by e[ξ] = ∫ +∞ 0 m{ξ ≥ x}dx− ∫ 0 −∞ m{ξ ≤ x}dx provided that at least one of the two integrals is finite; see [9]. copyright © 2021 assa. adv syst sci appl (2021) 10 t. latunde, a.a. ishaq, a.f. adedotun, j.o. ajinuhi, o.j. peter definition 2.5: an uncertain process xt is said to have independent increments if xt1 −xt0 , xt2 −xt1 , · · · , xtk −xtk−1 are independent uncertain variables, where t0 < t1 < · · · < tk. that is, an independent increment process means that its increments are independent uncertain variables whenever the time intervals do not overlap. it is noted that the increments are also independent of the initial state; see [10]. definition 2.6: suppose ct is a canonical liu process, and f and g are two functions. then dxt = f(t,xt)dt+ g(t,xt)dct is called an uncertain differential equation. a solution is a liu process xt that satisfies (2.1) and (2.2) identically in t; see [10]. definition 2.7: let xt be an uncertain process. then for each γ ∈ γ, the function xt(γ)is called a sample path of xt; see [10]. definition 2.8: an uncertain process xt is said to be sample-continuous if almost all sample paths are continuous functions with respect to time t; see [9]. definition 2.9: uncertainty distribution of solution. let α be a number with 0 < α < 1. an uncertain differential equation dx(t) = f(t,x(t))dt+ g(t,x(t))dc(t) is said to have an α-path x(t)α if it solves the corresponding ordinary differential equation dx(t)α = f(t,x(t)α)dt+ |g(t,x(t))|φ−1(α)dt, where φ−1(α) is the inverse uncertainty distribution of a standard normal uncertain variable, that is, φ−1(α) = √ 3 π ln α 1− α , α ∈ r; see [11]. 3. the asset management model asset management problem is mainly based on decision making and the understanding of probable asset degradation and trading-off capital investments, maintenance costs, risks and other uncertainties to optimize decisions made by investors. however, it is assumed that an individual invests his/her wealth in a capital asset, a(t), of a large business for a time, t, ranging from t0 to tn. supposing he/she starts with a known initial net worthx0(t). at the time t, what fraction of his/her net worth, ψ, must he/she choose to utilize on the capital asset and what fraction of his net worth, τ , must he/she choose to be incurred on the liability of the business such that the expected present value of the utility of asset, j(x), is maximized. table 3.1 represents the definition of the formulated model’s parameters. a dynamic optimization model of the expected present value of assets over a given life cycle based on uncertainty theory is herein presented following the study of portfolio copyright © 2021 assa. adv syst sci appl (2021) uncertain controllability and observability of an optimal control model 11 table 3.1. definition of parameters to the model. parameter description x(t) net worth at time t (state variable) τ(t) liability ratio (control) at time t, τ ∈ r σr(t) diffusion volatility of liability (with variance σ2 r per unit time) ψ(t) capital asset ratio at time t (control) ψ ∈ r σb(t) diffusion volatility of asset (with variance σ2 b per unit time) κ(t) capital gain on asset due to inflation at time t σp(t) diffusion volatility on asset price (with variance σ2 p per unit time) β(t) mean rate of return on the asset at time t ω(t) mean interest rate of liability at time t c(t) liu canonical process at time t µ(t) consumption level at time t j(t) tax ratio at time t g(t) depreciation ratio at time t h(t) asset supplies ratio at time t η subjective discount rate, e.g., a η+1 = present value λ degree of relative risk, where (1− λ) is the risk aversion u utility function selection by [12]. it is assumed that the goal of the asset management is to choose the optimal utilization and asset allocation policies for maximizing a value function that discounts exponentially future uncertain values of hyperbolic absolute risk aversion (hara) utility function over a given time horizon with the net worth of tangible assets as the state variable. the risky asset is assumed to earn an uncertain return and an uncertain gain with the mean rate of return and capital gain. furthermore, we express the change in liability as the sum of liability service with an assumption of uncertainty, consumptions, investment and net foreign supply, less taxation, depreciation and revenue over a period of time; see [13]. thus, we have j(x) = max ψ ec  tn∫ t0 1 λ e−ηt(ψx(t))λdt  subject to dx(t) = [(κ+ β)ψ − (ω(ψ − 1) + µ+ h− j − g)]x(t)dt +[ψσp + ψσb − σr(ψ − 1)]x(t)dc(t) . this model has been solved, characterised, analysed and applied to some real-life situations of sustainable finance; see [14–17]. 3.1. optimality of the solution it is important to derive the optimal solutions of the capital asset management problem as it would help in selecting the best available values. the following are utilised in deriving the optimality of the proposed model. definition 3.1: (principle of optimality) [19]: for any (t, x) ∈ [0, t )× r and ∆t > 0 with t+ ∆t < t , we have j(t, x) = sup d e [∫ t+∆t t f(xs, d, s)ds+ j(t+ ∆t, x+ ∆xt) ] , copyright © 2021 assa. adv syst sci appl (2021) 12 t. latunde, a.a. ishaq, a.f. adedotun, j.o. ajinuhi, o.j. peter where x+ ∆xt = xt+∆t. theorem 3.1: (equation of optimality, [18]) let j(t, x) be twice differentiable on [0, t )× r, then we have −jt(t, x) = sup d [f(x,d, t)) + jx(t, x)v (x,d)] , where jt(t, x) and jx(t, x) are the partial derivatives of the function j(t, x) in t and x respectively. proof see ( [19], pp. 15-16) 3.2. optimal control of the model the above equation of optimality is applied to the uncertain optimal control problem to evaluate the optimal controls analytically. applying equation (3.1), we obtain −jt = max ψ { 1 λ e−ηt(ψx)λ − ψ(κ+ β)xjx + (µ+ j + g + h− (ψ − 1)ω)xjx } = max ψ h where h stands for terms in the braces (condition the optimal ψ satisfies), ∂h ∂ψ = 0, ∂h ∂ψ = e−ηt(ψx)λ−1x − (κ+ β − ω)xjx = 0, ψ = 1 x [ (ω − κ− β)jxe ηt ] 1 λ−1 . hence, by solving the above equations, we obtained the optimal ratio of the net worth in capital assets as ψ∗ = (µ+ j + g + ω − h)λ− η (1− λ)(κ+ β − ω) . however, the optimal liability ratio, τ ∗ can also be obtained as a control to the system. since τ = ψ − 1 τ ∗ = [ (µ+ j + g + ω − h)λ− η (1− λ)(κ+ β − ω) ] − 1 or τ ∗ = (µ+ j + g − h)λ− (1− λ)(κ+ β) + ω − η (1− λ)(κ+ β − ω) . 3.3. solution to the model here, the analytical and numerical solutions are derived. for the analytic solution, the required problem under consideration is j(ψ) = min ψ ec  tn∫ t0 1 λ e−ηt(ψx(t))λdt  copyright © 2021 assa. adv syst sci appl (2021) uncertain controllability and observability of an optimal control model 13 subject to dx(t) = [(κ+ β)ψ − (ω(ψ − 1) + µ+ h− j − g)]x(t)dt +[ψσp + ψσb − σr(ψ − 1)]x(t)dc(t) with α-path equation dx(t)α = [(κ+ β)ψ − (ω(ψ − 1) + µ+ h− j − g)]x(t)αdt +|[ψσp + ψσb − σr(ψ − 1)]x(t)α|φ−1(α)dt. the analytical solution to the constraint is x(t) = x0 exp ([(κ+ β)ψ − (ω(ψ − 1) + µ+ h− j − g)]t +[ψσp + ψσb − σr(ψ − 1)]c(t)) and its inverse uncertainty distribution is ψ(t)−1(α) = x0 exp ( [(κ+ β)ψ − (ω(ψ − 1) + µ+ h− j − g)]t + [ψσp + ψσb − σr(ψ − 1)]t √ 3 π ln α 1− α ) . hence, ψ(t)−1(α) = e(x(t)α). 3.4. multifactor model the multifactor model can be expressed in the following form: j(x) = max ψ ec  tf∫ t0 1 λ e−ηt(uλ)tx1−λdt  (3.3) subject to dx = fxdt+ upxdt+ uqxdc(t), (3.4) where x =  x1t x2t ... xnt  , u =  ψ1 0 · · · 0 0 ψ2 · · · 0 ... . . . ... . . . 0 0 · · · ψn  , f =  µ1 + h1 − j1 − g1 − ω1 0 · · · 0 0 µ2 + h2 − j2 − g2 − ω2 · · · 0 ... . . . . . . ... 0 0 · · · µn + hn − jn − gn − ωm  , p =  κ1 + β1 − ω1 0 · · · 0 0 κ2 + β2 − ω2 · · · 0 ... . . . . . . ... 0 0 · · · κn + βn − ωn  , copyright © 2021 assa. adv syst sci appl (2021) 14 t. latunde, a.a. ishaq, a.f. adedotun, j.o. ajinuhi, o.j. peter q =  σ1p + σ1b − σ1r + 1 0 · · · 0 0 σ2p + σ2b − σ2r + 1 · · · 0 ... . . . . . . ... 0 0 · · · σnp + σnb − σnr + 1  . see, for example, [18]. equations (3.3) and (3.4), which are the model of risky capital assets, is an uncertain optimal control system whereby x(t) is the state, u is the control, f is n× n dimensional constant matrix while p and q are n×m dimensional constant matrices, c(t) is the liu process and j(x) is the objective functional. 4. controllability and observability of the uncertain system let (γ, l,m) be a complete uncertainty space with uncertain measurem on γ and a filtration {l(t)|t ∈ [0, t ]} generated by n-dimensional uncertain process {c(t) : 0 ≤ t ≤ t} defined on the uncertainty space (γ, l,m). let l2(γ, l(t),rn) represent the hilbert space of all l(t)-measurable square integrable uncertain variables with all values in rn. also, let ll2 ([0, t ],rn) represent the hilbert space of all square integrable and l(t)-measurable process with the values in rn. let x(t) = x(t+ s) for s ∈ [0, t ] represent the segment of the trajectory, that is, x(t) ∈ ll2 ([0, t ], l2(γ, l(t),rn)). letr be a linear operator on the hilbert space l2(γ, l(t),rn) with domaind1(r).r ≥ 0 if 〈rc, c〉 ≥ 0 for all c ∈ d1(r), r > 0 if 〈rc, c〉 > 0 for all nonzero c ∈ d1(r) and r is coercive (r− γi ≥ 0) if there exists a γ > 0 such that 〈rc, c〉 ≥ γ ‖c‖2for all c ∈ d1(r). now, using the proposed model with the multidimensional constraint of uncertain differential equation dx(t) = fx(t)dt+ upx(t)dt+ uqx(t)dc(t) (4.5) for t ∈ [0, t ] with the function initial condition x0 ∈ ll2 ([0, t ], l2(γ, l(t),rn)), where the state x(t) ∈ l2(γ, l(t),rn) and control ψ(t) ∈ rm = u . u and f are n× n dimensional constant matrix while p and q are n×m dimensional constant matrices. thus, suppose the admissible controls u = ll2 ([0, t ],rm), then for any given initial condition x0 ∈ ll2 ([0, t ], l2(γ, l(t),rn)) and any admissible control ψ ∈ u for t ∈ [0, t ], there exists a unique solution x(t;x0, ψ) ∈ l2(γ, l(t),rn) of the constraint uncertain differential state equation (4.5); see [20]. 4.1. controllability definition 4.1: the uncertain dynamic system (4.5) is said to be relatively exactly controllable on [0, t ] if r(t)(u) = l2(γ, l(t),rn). this implies that if all the points in l2(γ, l(t),rn) can exactly be reached at time t from any arbitrary initial condition x0 ∈ ll2 ([0, t ], l2(γ, l(t),rn)). definition 4.2: the uncertain dynamic system (4.5) is said to be relatively approximately controllable on [0, t ] if r(t)(u) = l2(γ, l(t),rn). copyright © 2021 assa. adv syst sci appl (2021) uncertain controllability and observability of an optimal control model 15 this implies that if all the points in l2(γ, l(t),rn) can approximately be reached at time t from any arbitrary initial condition x0 ∈ ll2 ([0, t ], l2(γ, l(t),rn)). the relationship between the controllability concepts for the uncertain dynamic system (4.5) and the controllability of the related deterministic dynamic system below dx(t) = fx(t)dt+ upx(t)dt for t ∈ [0, t ], (4.6) where the admissible control ψ ∈ l2([0, t ],rm). firstly, the deterministic system (4.6) is defined according to [1]. let qj(t) = fqj−1(t) for j = 1, 2, 3, . . . and t > 0, with the initial condition q(t) = q0(0) = j, t = 0, q(t) = q0(t) = 0, t 6= 0. for instance, the sequence of the n× n dimensional matrices qj(t) deduced from the determining equation gives: q0(0) = p, q1(0) = fp, q2(0) = f 2p. these can be written in a general notation as qj(t;t ) = {q0(t), q1(t), q2(t), . . . , qj−1(t) fort ∈ [0, t ]}. the following lemma is given with respect to [1], relating to the controllability of the deterministic system (4.6) in the time interval [0, t ]: lemma 4.1: the following conditions are equivalent: 1. the deterministic system (4.6) is relatively controllable on [0, t ], 2. the relative controllability matrix b(t) is non-singular, 3. rank qj(t;t ) = k. hence, the following lemma is formed with respect to [4, 20, 21] which will be useful in the proof of the uncertain controllability and observability. lemma 4.2: for every c ∈ l2(γ, l(t),rn), ∃ a process q ∈ ll2 ,rn×n such that the controllability operator is expressed as: g(t)c = b(t)ec+ ∫ t 0 b(t)(s)q(s)dc(s). lemma 4.3: the uncertain system (4.5) is relatively controllable on [0, t ] if and only if one of the following conditions holds true: 1. e 〈g(t)c, c〉 ≥ γe ‖c‖2 for some γ > 0 and all c ∈ l2(γ, l(t),rn), 2. r(λ1, g(t)) converges as λ1 → 0+ in the uniform operator topology, copyright © 2021 assa. adv syst sci appl (2021) 16 t. latunde, a.a. ishaq, a.f. adedotun, j.o. ajinuhi, o.j. peter 3. λ1r(λ1, g(t)) converges to the zero operator as λ1 → 0+ in the uniform operator topology, 4. ker(l(t))∗ = {0} and im(l(t))∗. lemma 4.4: the uncertain system (4.5) is approximately controllable on [0, t ] if and only if one of the following conditions holds true: 1. g(t) > 0, 2. λ1r(λ,g(t)) converges as λ1 → 0+ in the strong operator topology, 3. λ1r(λ,g(t)) converges to the zero operator as λ1 → 0+ in the weak operator topology, 4. ker(l(t))∗ = {0}. theorem 4.1: the following conditions are equivalent: 1. the deterministic system (4.6) is relatively controllable on [0, t ], 2. the uncertain system (4.5) is relatively exactly controllable on [0, t ], 3. the uncertain system (4.5) is relatively approximately controllable on [0, t ]. proof condition (i) implies condition (ii). suppose the deterministic system (4.6) is relatively controllable on [0, t ], then the relative controllability matrix b(t)(s) is invertible and strictly positive definite for all s ∈ [0, t ]. hence, for some γ > 0, 〈b(t)(s)x,x〉 ≥ γ ‖x‖2 for all s ∈ [0, t ] and x ∈ rn. in order to prove the relative exact controllability of the uncertain system (4.5) on [0, t ], the relationship between the controllability operator g(t) and the controllability matrix b(t) given in lemma 4.2 as e 〈g(t)c, c〉 to be expressed in terms of 〈g(t)ec,ec〉. firstly, e 〈g(t)c, c〉 = e 〈 b(t)ec+ ∫ t 0 b(t)(s)q(s)dc(s),ec+ ∫ t 0 q(s)dc(s) 〉 = 〈b(t), ec, ec〉+ e ∫ t 0 〈b(t)(s)q(s), q(s)〉 ds ≥ γ ( ‖ec‖2 + e ∫ t 0 ‖q(s)‖2 ds ) = γe ‖c‖2 . thus, in the view of the controllability operator, g(t) ≥ γi, which implies that the relative controllability operator g(t) is strictly positive definite and, the inverse operatorg(t)−1 is bounded, [1]. hence, the uncertain relative exact controllability of uncertain dynamic system (4.5) on [0, t ] is proved from the relative controllability of deterministic system (4.6) on [0, t ] . condition (ii) implies condition (iii). since the state space for the uncertain dynamic system (4.5) is finite-dimensional, which implies that the exact and approximate controllability coincide; see [1]. hence, it is easy to conclude that the uncertain system (4.5) is relatively approximately controllable on [0, t ]. condition (iii) implies condition (i). suppose the uncertain dynamic system (4.5) is uncertainly relatively approximately copyright © 2021 assa. adv syst sci appl (2021) uncertain controllability and observability of an optimal control model 17 controllable on [0, t ] and its controllability operator is positive definite, that is, g(t) > 0, then applying the resolvent operator λ1r(λ1, g(t)), where λ > 0, [20], gives e ‖λ1r(λ1, g(t))c‖2 → 0. this implies e ‖λ1r(λ1, g(t))c‖2 = ‖λ1r(λ1, b(t))ec‖2 + e ∫ t 0 ‖λ1r(λ1, b(t)(s))q(s)‖2 ds→ 0. therefore, e ∫ t 0 ‖λ1r(λ1, b(t)(s))q(s)‖2 ds→ 0 for all q(s) ∈ ll2 [0, t ],rn×n, and consequently there exists a subsequence λn such that for every c ∈ l2(γ, l(t),rn) we have ‖λnr(λk, b(t)(s))c‖ → 0 almost everywhere on [0, t ]. thus, the property holds for all 0 ≤ s < t because of the continuity of r(λ,b(t)(s)). this implies that the deterministic system (4.6) is relatively approximately controllable on [0, t ]. however, considering the state space for the deterministic system (4.6) is finitedimensional, that is, exact and approximate controllabilities coincides, [1]. therefore, it is concluded that the deterministic system (4.6) is relatively controllable on [0, t ]. 4.2. observability here, the dual concepts of observability for the uncertain dynamic system (4.5) are treated. suppose f and j generate c0-semigroup s1 and s2 on the hilbert space l2(γ, l(t),rn), that is, f and j generate a continuous representation of semigroup s. then the dual of the uncertain system (4.5) is: dx(t) = [f ∗ + j∗ +w ∗ 1 +w ∗ 2 ]x(t)dt+ uqx(t)dc(t). thus, the following concepts are described: • the observability map of the uncertain system (4.5) on [0, t ] is expressed as the linear operator o(t) : l2(γ, l(t),rn)→ ll2 ([0, t ]) which is defined by o(t)c = (w1 +w2)(s1 + s2)(t − s)e{c|l}. • the observability gramian of the uncertain system (4.5) on [0, t ] is defined by: θ = (o(t))∗o(t). definition 4.3: the uncertain system (4.5) is said to be relatively observable on [0, t ] if the operator o(t) is injective and its inverse bounded on the range o(t). the means that the initial state can be uniquely and continuously constructed from outputs in ll2 (γ, l(t),rn). definition 4.4: the uncertain system (4.5) is said to be approximately observable on [0, t ] if o(t) = {0}, that is, the initial state uniquely depends on the knowledge of the result in ll2 (γ, l(t),rn). copyright © 2021 assa. adv syst sci appl (2021) 18 t. latunde, a.a. ishaq, a.f. adedotun, j.o. ajinuhi, o.j. peter theorem 4.2: for the uncertain dynamic system (4.6), the following duality results hold true: 1. the uncertain system (4.5) is relatively observable on [0, t ] if and only if the dual system (4.2) is relatively controllable on [0, t ]. 2. the uncertain system (4.5) is approximately controllable on [0, t ] if and only if the dual system is approximately controllable on [0, t ]. proof since f and j generate a co-semigroup s1(t) and s2(t) respectively on ll2 (γ, l(t),rn), then f ∗ and j∗ generate the co-semigroup s∗1(t) and s∗2(t) respectively. also, since o(t) ∈ ll2 ([0, t ], l2(γ, l(t),rn)), (o(t))∗q = ∫ t 0 (s∗1 + s∗2)(t − s)(w ∗ 1 +w ∗ 2 )q(s)ds. thus, the range of o(t) implies that of the controllability operator of the dual system (4.2). let the controllability of the dual system (4.2) be represented by n(t), then (o(t))∗ = n(t), (n(t))∗ = o(t). 1. suppose the deterministic system (4.5) is relatively observable, there exists an inverse (o(t))−1 on the range of o(t). then,∥∥(o(t))−1q ∥∥ ≤ l ‖q‖ for all q ∈ ll2 ([0, t ],rn×n) and l > 0. hence, ‖c‖ = ∥∥(o(t))−1o(t)c ∥∥ ≤ l ∥∥ot c ∥∥ = l ‖(n(t))∗c‖ . therefore, the relative controllability of the uncertain system (4.5) follows from lemma 4.3 as thus. suppose that the uncertain system is relatively controllable, then (n(t))∗ is injective and has a closed range. this implies that from (n0)∗ = o(t), o(t) is injective and has a closed range. thus, by the closed graph theorem, the inverse of o(t) is bounded on the range of o(t). 2. however, by definition, the deterministic system (4.6) is approximately observable if and only if kero(t) = ker(n(t))∗ = {0}. therefore by lemma 4.4, ker(n(t))∗ = {0} if and only if the uncertain system (4.5) is approximately controllable. hence the proof of the equivalence. 5. conclusion the necessary and sufficient conditions for the uncertain controllability and observability of a finite-dimensional uncertain dynamic control system have been established and proved. thus, the effectiveness of each input and output in the general operation of the control system can be determined. the application of this work is to measure the controllability and observability degree of the system by using certain admissible input factors to determine how the system can move around its configuration space. subsequently, we can observe the analytic and numerical solutions to the multifactor system. copyright © 2021 assa. adv syst sci appl (2021) uncertain controllability and observability of an optimal control model 19 references 1. klamka, j. (1991) controllability of dynamical systems. dordrecht: kluwer academic, 1991. 2. klamka, j. (1993) controllability of dynamical systems a survey. archives of control sciences, 2(3/4), 281–307. 3. klamka, j. (2007) stochastic controllability of linear systems with state delays. international journal of applied mathematics and computer science, 17(1), 5–13. 4. mahmudov, n.i.& denker a. (2000) on controllability of linear stochastic systems. international journal of control, 73(2), 144–151. 5. mahmudov, n.i. (2003) controllability and observability of linear stochastic systems in hilbert spaces. progress in probability, 53, 151–167. 6. guo, t.l. (2012) controllability and observability of impulsive fractional linear timeinvariant system. computer and mathematics with applications, 64 3171–3182. 7. wang, l & chen, y. (2017) physical controllability of complex networks. scientific reports, 7, 1–14. 8. leitold, d., vathy-fogarassy, a. & abonyi, j. (2017) controllability and observability in complex networks the effect of connection types. scientific reports, 7(151), 1–9. 9. liu, b. (2016) uncertain theory (5th ed.), uncertainty theory laboratory, china. 10. liu, b. (2008) fuzzy process, hybrid process and uncertain process. journal of uncertain systems 2(1), 3–16. 11. yao, k. & chen, x.a. (2013) numerical method for solving uncertain differential equations. journal of intelligent and fuzzy systems, 25(3), 825–832. 12. merton, r.c. (1971) optimal consumption and portfolio rules in a continuous time model. journal of economic theory, 42 (3), 373–413. 13. latunde, t. & bamigbola o.m. (2016) uncertain optimal control model for management of net risky capital asset. iosr journal of mathematics, 12(3), 22–30. 14. latunde, t. & bamigbola, o.m. (2016) on existence and uniqueness of solution to a special case of asset management. journal of the nigeria association of mathematical physics (namp), 35 269–274. 15. latunde, t. & bamigbola, o.m. (2016) on the stability analysis of uncertain optimal control model of management of net risky capital asset. international journal of applied sciences and mathematical theory, 2(1), 29–36. 16. latunde, t. & bamigbola o.m. (2018) parameter estimation and sensitivity analysis of an optimal control model for capital asset management. advances in fuzzy systems, 2018 1–11. 17. latunde, t. (2019) optimal values in an uncertain optimal control model with application to capital asset management. advances in systems science and applications (assa), 19(3), 52–64. 18. latunde, t. (2020) multifactor modelling in asset management. international journal of mathematics in operational research, 17(3), 333–352. 19. zhu, y. (2010) uncertain optimal control with application to a portfolio selection model. cybernetics and systems: an international journal, 41, 535–547. 20. mahmudov, n.i. (2001) controllability of linear stochastic systems in hilbert spaces. journal of mathematical analysis and application, 251, 64–82. 21. mahmudov, n.i. & zorlu, s. (2003) controllability of nonlinear stochastic systems. international journal of control, 76(2), 95–104. copyright © 2021 assa. adv syst sci appl (2021) introduction preliminaries the asset management model optimality of the solution optimal control of the model solution to the model multifactor model controllability and observability of the uncertain system controllability observability conclusion adv syst sci appl 2018; 04; 92-120 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/536 design of robust stabilizing pi/pid controller for time delay interval process plants using particle swarm optimization d.srinivasa rao1, m.siva kumar*2, m. ramalinga raju3 1,2) department of electrical and electronics engineering, gudlavalleru engineering college, gudlavalleru, ap, india. e-mail: dsrinivasarao1993@gmail.com, profsivakumar@gmail.com. 3) department of electrical and electronics engineering, university college of engineering, jntu university kakinada, ap, india. e-mail: rajumanyala@yahoo.com received december 16,2017; revised november 11,2018; published december 31,2018 abstract: generally, most of the real plants operate in a wide range of unknown operating conditions, but bounded parameter uncertainties in the control system. these uncertainties in the control system cause degradation of system performance and destabilization. in general, it is not easy to design a controller for interval time –delay process plant, because of interval dead time therefore, robust control of these uncertainties is a vital to operate the plant under stabilized condition. with a view to conquering the uncertainty, in this paper a new stability conditions are developed for determining the stability of interval process plants based on rouche’s theorem and then a robust pi/pid controller is designed for the interval process plant with and without time delay based on these newly developed stability conditions for stability of interval polynomial by using particle swarm optimization algorithm. a set of inequalities for a closed loop characteristic polynomial of an interval process plant in terms of controller parameters are derived from these newly developed stability conditions. these inequalities are solved to obtain controller parameters with the help of pso algorithm. the pi/pid controller designed in this proposed method stabilizes the given interval plant with and without time delay at all operating conditions. the proposed method has the advantage of having less computational complexity and easy to implement on a digital computer. the viability of the proposed methodology is illustrated through numerical examples of its successful implementation. the efficacy of the proposed methodology is also evaluated against the available approaches presented in the literature and the results were successfully implemented. keywords: kharitonov’s theorem, parametric uncertainty, robust controller, interval polynomial, particle swarm optimization. 1. introduction generally, many of the real plants operate in a wide range of unknown operating conditions bounded under parametric uncertainties called interval plants, in control systems. the large uncertainty present in the control system causes degradation of system performance and destabilization. therefore, robust control of these uncertainties is vital to operate the plant under stabilized condition. this necessitates a robust controller design which could stabilize the plant for all the operating conditions. hence designing a robust controller for the parametric uncertain plants having unknown, but bounded parameter uncertainties has become the problem of research nowadays. with a view to minimizing the stated * corresponding author: profsivakumar@gmail.com mailto:dsrinivasarao1993@gmail.com mailto:rajumanyala@yahoo.com design of robust stabilizing pi/pid using particle swarm optimization 93 copyright ©2018 assa. adv. in systems science and appl. (2018) uncertainties, many solutions are proposed in the literature for the simulation, design and tuning of controllers [1-3]. recently, affordable results have been reported on computation of all stabilizing p, pi and pid controllers which are mentioned here. the problem in [4] of stabilizing a linear time-invariant plant using a fixed order compensator was considered by using the hermite-biehler theorem. a feasible robust pid controllers have been developed in [5] using the minimax search, a co-evolutionary algorithm based on particle swarm optimization. the problem of designing robust and optimal pid controllers for a given linear time-invariant plant was proposed in[6-7]. design of a robust pid controller for a first-order lag with pure delay (folpd) model in [8] using pso enabled automated quantitative feedback theory (qft) and compared with manual graphical techniques. a design method was proposed for calculating the optimum values pid controller for interval plants using the pso algorithm. most of the practical systems operate based on approximate polynomial models; the parameters of these models would lie within an interval but not have specific values and are unknown. therefore, the stability analysis of polynomials subjected to parameter uncertainty has received considerable attention after the celebrated theorem of kharitonov [10], which assesses robust stability under the condition that four specially constructed extreme polynomials, called kharitionov polynomials are hurwitz. robust stability of interval polynomial is also discussed here by many researchers. among these discussions, some important methods have been presented here from the literature. a robust controller has been designed [11] for interval plants based on kharitonov’theorem and the results nie of [12] for fixed polynomials. a systematic optimization approach was proposed [13] to design a robust controller using the well-known kharitonov and hermite-biehler stability theorems for single-input/single-output process systems in the presence of unknown but bounded parameter uncertainties. the problem of robust stabilization [15] of a linear timeinvariant system was considered subject to variations of a real parameter vector used to design a robust controller. the design of a robust course controller was proposed in [16] for a cargo ship interacting with an uncertain environment using pso enabled automated quantitative feedback theory. a fractional-order proportional-integral controller was proposed and designed in [17] for a class of nonlinear integer-order systems to guarantee the desired control performance and the robustness of the designed controllers to the loop gain variations. a robust controller was designed in [18-19] for interval plants based on the result of kharitonov’theorem. with a view to reducing the test of hurwitz stability of the entire family, several investigations have been presented in the literature. among these, a few imperative investigations are discussed here, including; an algorithm has been presented in the design of a robust pi and pid controller [20]. this method is based on approximating the fuzzy coefficients by the nearest interval system and then a robust controller is designed using the necessary and sufficient conditions for stability of the interval systems. the inverse bilinear transformation (ibt) is proposed in [21] to design a robust controller using the necessary and sufficient conditions for a discrete-time interval plants. they designed a robust controller using the necessary and sufficient conditions for a chemical process plant with delay subjected to unknown, but bounded parameter uncertainties referred to an interval process plant with interval time delay. srinivasa rao et.al [22] presented a new algorithm for the design of the robust pi controller for a process control interval plant using routhe’s theorem and kharitonov’s theorem. a robust pi controller design approach was discussed in [23] by finding the controllers using pole placement method for active suspension system with parametric uncertainty. a pure gain compensator c(s) = k stabilizes the entire interval plant family such that a distinguished set of eight of the extreme plants are stable [24]. the first order controller is made by the experimental setup in [25] which was developed by ghosh. they prove that to robustly stabilize the extreme plants which are obtained by taking 94 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) all possible combinations of extreme values of the plant numerator has degree m and the plant denominator is monotonic with degree n, the number of extreme plants can be high as next =2m+n+1 in [26]. an explicit equation of control parameters defining the stability boundary in parametric space was derived based on the plant model in time domain and by using the extraordinary feature results from the kronecker sum operation [27]. the stabilizing values of the parameters of a pi controller were computed based on plotting the stability boundary locus method in [28]. a complete survey of these extreme points is given in [29]. the necessary and sufficient conditions in [18] and [30] for interval polynomials are proposed using the results of [12] for fixed polynomials. in process industries due to the presence of transportation lag, recycle loops, and dead time corresponding to composition analysis, time delays frequently occur. the mathematical model of uncertain processes has described by the interval time-delay model in the presence of time delay. unfortunately, as compared with the successful development of controller design for the rational interval model, much less effort has been devoted to the processes described by an interval time-delay model; this is because the process delay is a source of instability and can render many established techniques inadequate. a new approach [32] was proposed to determine the entire set of stabilizing pi/pid parameters for time delay process with bounded uncertainties using the combination of the generalized kharitonov theorem and the hermit biehler theorem. the design of a smith predictor for the operation of processes under the variation of process gain, time constant and dead time based on the concept of an inferential control framework was presented in [32]. by considering the interval time-delay process, an alternative smith predictor design in [33] was proposed for the purpose of ensuring robust performance. the use of the structured singular value to design robust controllers was presented in [34] for interval time-delay processes. the designers of the robust stabilizing controller and construction of pre-filter with interval time delay have been considered in [30] to guarantee both robust stability and performance. particle swarm optimization (pso), first introduced by kennedy and eberhart [35] is one of the modern heuristic algorithms. it was developed through simulation of a simplified social system, and has been found to be robust in solving continuous nonlinear optimization problems [9] and [36-37]. the pso technique can generate a high-quality solution within shorter calculation time and stable convergence characteristic than other stochastic methods. in this note, a pi/pid controller is designed for an interval process plant with and without time delay based on the newly developed necessary and sufficient stability conditions. these conditions are used to derive a set of inequalities in terms of controller parameters. the inequality constraints from the characteristic polynomial are solved consequently to obtain the controller parameters with the help of pso algorithm. the efficacy of the proposed method is demonstrated by implementing with typical numerical examples available in the literature. in comparison with the method available in the literature [5], [8], [18], [30] and [31] the proposed method in this paper is simple and involves less computational complexity. the paper is organized as follows: section 2 describes development of stability conditions for robust stability of interval polynomial. section 3 gives the design of robust stabilizing pi/pid controller with and without time delay process plant. section 4 proposes pso algorithm to find the controller parameters. in section 5, the proposed method is applied to design a robust pi/pid controller for an interval process plant. 2. development of stability conditions for interval polynomial according to anderson et.al [38] the necessary and sufficient condition for robust stability of interval polynomials of order 3n  is positive lower bounds on the coefficients of an interval polynomial. design of robust stabilizing pi/pid using particle swarm optimization 95 copyright ©2018 assa. adv. in systems science and appl. (2018) therefore, consider an interval polynomial of order n=1 ].b,a[pwhere,sp)s(p iii i1 0i i   ].b,a[s]b,a[psp)s(p 001101  therefore, as per anderson [38], the robust stability condition is 0aand0a 01  i.e. 0a i for i=0,1. similarly for order n=2 ].b,a[s]b,a[s]b,a[pspspsp)s(p 0011 2 2201 2 2 i2 0i i   therefore, the robust stability condition is 0a,0a 12  and 0a0  i.e. 0a i  for i =0,1,2. lemma 2.1 consider a real hurwitz polynomial q(s) of the form 01 i i 1n 1n n n qsq....sq...sqsqq(s)    (2.1) n,.......,2,1,0i  where iq is real and positive, 0q0  . if any complex number z such that ,)z(f)z(f,0re  moreover, ,)z(f)z(f conzconz  where c is a closed contour, then, according to routhe’s theorem [39] the following two polynomials can be formulated. x2s0 ])s(q)s(q[ 2 1 q   (2.2) x2s1 ])s(q)s(q[ s2 1 q   (2.3) theorem 2.1: for stability of )s(q the two polynomials 0q and 1q formed by the alternate coefficients of a hurwitz polynomial in accordance with equations (2.2) and (2.3) must have negative real zeros. the proof of this is given in [39]. 2.1 necessary conditions for stability of an interval polynomials consider an interval polynomial of order n > 3 of the form ,psp...sp...spsp)s(p 01 i i 1n 1n n n    where ]b,a[p iii  for .n,.....,3,2,1,0i  the necessary conditions for an interval polynomial to be stable is given as 0ab ii  for n,......,3,2,1,0i  (2.4) 2.2 sufficient conditions for stability of interval polynomial for n > 3 2.2.1. for fourth-order interval polynomial (n= 4) consider the fourth-order interval polynomial as 01 2 2 3 3 4 4 pspspspsp)s(p  (2.5) where ]b,a[pand]b,a[p],b,a[p],b,a[p],b,a[p 444333222111000  using lemma 2.1, the p(s) can be represented into two polynomials p0 and p1 as given below. ]b,a[x]b,a[x]b,a[)]s(p)s(p[ 2 1 p 0022 2 44x2s0   (2.6) 96 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) ]b,a[x]b,a[)]s(p)s(p[ s2 1 p 1133x2s1   (2.7) according to the theorem 2.1, for robust stability of interval polynomial p(s) the above polynomials p0 and p1 must have negative real zeros i.e. ]b,a][b,a[4]b,a[ 4400 2 22  (2.8) .0 ]b,a[ ]b,a[ 33 11   (2.9) apply interval arithmetic to the above equations (2.8) and (2.9), the stability conditions for the interval polynomial are 40 2 2 bb4a  (2.10) 40 2 2 aa4b  (2.11) 0 b a 3 1   (2.12) 0 a b 3 1   (2.13) from the above four equations, the sufficient conditions for the robust stability of fourth order interval polynomial p(s) are 40 2 2 bb4a  (2.14) 0 b a 3 1   (2.15) 2.2.2. for fifth-order interval polynomial (n = 5) consider the fifth-order interval polynomial as 01 2 2 3 3 4 4 5 5 pspspspspsp)s(p  (2.16) where ]b,a[p and]b,a[p],b,a[p],b,a[p],b,a[p],b,a[p 555 444333222111000   using lemma 2.1, the p(s) can be represented into two polynomials p0 and p1 as given below. ]b,a[x]b,a[x]b,a[)]s(p)s(p[ 2 1 p 0022 2 44x2s0   (2.17) ]b,a[x]b,a[x]b,a[)]s(p)s(p[ s2 1 p 1133 2 55x2s1   (2.18) according to the theorem 2.1, for robust stability of interval polynomial p(s) the above polynomials p0 and p1 must have negative real zeros i.e. design of robust stabilizing pi/pid using particle swarm optimization 97 copyright ©2018 assa. adv. in systems science and appl. (2018) ]b,a][b,a[4]b,a[ 4400 2 22  (2.19) ]b,a][b,a[4]b,a[ 5511 2 33  (2.20) apply interval arithmetic to the above equations (2.19) and (2.20), the stability conditions for the interval polynomial are 40 2 2 bb4a  (2.21) 40 2 2 aa4b  (2.22) 51 2 3 bb4a  (2.23) 51 2 2 aa4b  (2.24) from the above four equations, the sufficient conditions for the robust stability of fifth order interval polynomial p(s) are 40 2 2 bb4a  (2.25) 51 2 3 bb4a  (2.26) in a similar manner, the robust stability conditions for interval polynomial of degree 4n  can be determined. the robust stability conditions for higher-order interval polynomials are represented in a tabular form in table 2.1. . table2.1. robust stability conditions for various higher order interval polynomials order of the polynomial robust stability conditions necessary conditions sufficient conditions 3n  0ai  where i = 0,1,2,3. 20 2 1 3 bba  4n  0ai  where i = 0, 1,2,3,4. 40 2 2 4 bba  and .0 3 1   b a 5n  0ai  where i = 0, 1,..,4, 5. 40 2 2 4 bba  and 51 2 3 4 bba  . 6n  0ai  where i = 0, 1,...5, 6. 40 2 2 3 bba  and .4 51 2 3 bba  7n  0ai  where i = 0, 1,.... 6, 7. 40 2 2 3 bba  and 51 2 3 3 bba  . 8n  0ai  where i = 0, 1,2,..,7,8. 51 2 3 3 bba  , 80 2 4 4 aab  and 0 6 2   b a . 9n  0ai  where i = 0, 1,2,..,8,9.. 80 2 4 4 aab  91 2 5 4 aab  , 0 6 2   b a and 0 7 3   b a 80 2 4 4 aab  , 91 2 5 4 aab  , 98 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) using these developed stability conditions, the stability of interval polynomials can be determined easily without formulating the four kharitonov’s polynomials, unlike kharitonov’s theorem. a robust pi/pid controller (which can stabilize the given plant under large uncertainty) can be designed easily using these stability conditions. the design procedure is given in following section. 3. design of robust stabilizing pi/pid controller 3.1 interval plant without time delay consider a plant with parametric uncertainty without time delay represented by its transfer function as n n 1n 1n10 m m 1m 1m10 sdsd...sdd scsc...scc )d,s(d )c,s(n )d,c,s(g        (3.1) where ]c,c[cc ii  for m,....,2,1,0i  ]d,d[dd ii  for n,....,,,i 210 and the bounds  iii d,c,c and  id are specified a priori and mn  . let the stabilizing pi/pid controller transfer function of the form given below )s(d )s(n s k k)s(c c ci ppi  for pi controller (3.2) )s(d )s(n sk s k k)s(c c c d i ppid  for pid controller (3.3) where pk = proportional gain, ik =integral gain and dk = derivative gain 3.2. interval plant with time delay consider a plant with parametric uncertainty with time delay represented by its transfer function as sdte )d,s(d )c,s(n )d,c,s(g   (3.4) where the numerator and denominator polynomials are of the form m m 1m 1m10 scsc...scc)c,s(n    n n 1n 1n10 sdsd...sdd)d,s(d    with the parameters being specified by their lower and upper bounds as follows: ]c,c[cc ii  for .m,....,2,1,0i  ]d,d[dd ii  for .n,....,2,1,0i    ddd ttt . 0td  it is not easy work to design the stabilizing controller for interval time-delay process plants, because of interval dead time. in order to extend the technique, the firstorder rational function of the approximation of the interval delay plant is obtained using the procedure given in corollary 3.1 to solve the design of the robust stabilizing controller problem of an interval time-delay plant. 10n 0ai  where i = 0, 1,2,..,10. 102 2 6 4 aab  and 0 7 3   b a design of robust stabilizing pi/pid using particle swarm optimization 99 copyright ©2018 assa. adv. in systems science and appl. (2018) 3.2.1. an interval approximate time delay model the firstorder rational function, )t,s( d of the approximation of the interval delay part is given by .)t,s(e d sdt   for   ddd ttt (3.5) the simplex form for )t,s( d is given by the following corollary which is expressed in[13]. corollary 3.1 : the interval function sdte  , where   ddd ttt can be approximated by a first-order interval rational function ),,s(  which is given by s 2 t t, 2 t 1 s 2 t t, 2 t 1 )t,s( d d d d d d d                      (3.6) with the following properties: 1 j 2 t 1 jt 2 t 1 )t,j( max d d d maxd                       (3.7) 1 jt 2 t 1 j 2 t 1 )t,j( min d d d mind                       (3.8) .e)t,j( jdt maxd    (3.9) .e)t,j( jdt mind    (3.10) where <  and  is the limited frequency that the inequalities in equations (3.7) to (3.10) hold. the arbitrarily extended approximation of a desired degree is the first-order approximation, since it covers the properties of phase frequency of original function for a wider range of frequencies, whereas the approximate original function can be in the other forms of approximations for control in the loss of some system information within a very limited frequency range. the rational interval function, then the approximate system model is given by ).t,s( )d,s(d )c,s(n )d,c,s(g d (3.11) after simplification the (3.11) will become 100 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) )d,s(d )c,s(n )d,c,s(g  (3.12) where m m 1m 1m10 scsc...scc)c,s(n    n n 1n 1n10 sdsd...sdd)d,s(d    ]c,c[cc ii  for .m,....,2,1,0i  ]d,d[dd ii  for .n,....,2,1,0i    ddd ttt . and the bounds ddiiii tandt,d,d,c,c   are specified a priori and mn  . now the system with robust stabilizing controller for parametric uncertainty is as shown in fig.3.1. fig.3.1. block diagram of interval plant with a robust controller let the stabilizing pi/pid controller transfer function of the form given below )s(d )s(n ) s 1 1(k)s(c c c i cpi   for pi controller (3.13) )s(d )s(n ) s 1 1(k)s(c c c d i cpid    for pid controller (3.14) where kc = proportional gain, i =integral gain and d = derivative gain. then the closed loop transfer function with a pi / pid controller can be defined as )d,s(d)s(d)c,s(n)s(n )c,s(n)s(n )d,c,s(g)s(c1 )d,c,s(g)s(c )s(t cc c pi pi     for pi controller (3.15) )d,s(d)s(d)c,s(n)s(n )c,s(n)s(n )d,c,s(g)s(c1 )d,c,s(g)s(c )s(t cc c pid pid     for pid controller (3.16) the characteristic equation of this system with a pi / pid controller is given as )d,s(d)s(d)c,s(n)s(n)d,c,s(g)s(c1 ccpi  for pi controller (3.17) )d,s(d)s(d)c,s(n)s(n)d,c,s(g)s(c1 ccpid  for pid controller (3.18) where )c,s(n and )d,s(d are the numerator and denominator polynomials of the plant considered respectively, and )s(nc and )s(dc are the numerator and denominator polynomials of pi/pid controller transfer function respectively. this pi/pid controller robustly stabilizes the interval plants family, if for all cc and dd  , then the characteristic polynomial of a closed loop transfer function given in equations (3.17) and (3.18) has all zeros have negative real values. now apply the necessary and sufficient design of robust stabilizing pi/pid using particle swarm optimization 101 copyright ©2018 assa. adv. in systems science and appl. (2018) conditions of robust stability conditions given in table.1 to the closed-loop polynomial )d,s(d)s(d)c,s(n)s(n cc  which leads to a set of inequalities in terms of controller parameters. then these inequalities can be solved by using pso with matlab optimization [40] programming so as to minimize the objective function to obtain controller parameters. then after obtaining the controller parameters, form four sets of kharitonov’s polynomials to check the stability and the closed-loop step response to verify the results. the pso algorithm for the proposed method is given in section 4. fig.3.2. flowchart for the proposed algorithm. 4. application of particle swarm optimization algorithm kennedy and eberhart [35] first introduced the pso method. it is one of the optimization algorithms and a kind of evolutionary computation algorithm. the method has been found to be robust in solving problems featuring nonlinearity and no differentiability, multiple optima, and high dimensionality through adaptation, which is derived from the social-psychological theory. pso is inspired by social and cooperative behaviour displayed by various species to fill their needs in the search space. the algorithm is guided by personal experience (p best), the overall experience (g best) and the present movement of the particles to decide their positions in the next space. further the experiences are accelerated by two factors c1 and c2, and two random numbers generated between [0, 1]. whereas the present movement is multiplied by an inertia factor w varying between [wmin,wmax]. 102 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) application of pso algorithm for determining the controller parameters is as follows: step 1: initialization: pso parameters are chosen as population size (p) = 100 number of iterations (n) =1000 cognitive coefficients c1 =2 and c2 =2 inertia weight n/)ww(xiterww minmaxmax  where wmax =0.9, wmin =0.4 step 2: initial search space limits of control variables in pi/pid controller are selected as for pi controller: 10k0 p  and 10k0 i  .(without time delay) 1k0 c  and 150 i  (with time delay) for pid controller: ,10k0 p  10k0 i  and 10k0 d  (without time delay) ,1k0 c  150 i  and 100 d  (with time delay) step 3: initial search space populations of xi are generated from specified intervals using the given below equation ),xx).((randxx min,imax,imin,i 0 i  (4.1) where i =1, 2......, n, and rand () represents a uniformly distributed random number within the range of [0,1] step 4: initialize the iteration index n=1 during the initialization, parameters of a pi / pid controller are randomly generated within allowable limits. step 5: evaluate the fitness function to get effective performance, in this paper, the fitness function j is defined as minimize j for pi controller: the fitness function j is chosen as integral square error (ise)   0 2 dt)t(ej (without time delay) (4.2) where output1)t(e  2 0 i 0 ii 2 0 c 0 cc k kk j       (with time delay) (4.3) for pid controller: the fitness function j is chosen as integral square error (ise)   0 2 dt)t(ej (without time delay) (4.4) where output1)t(e  2 0 d 0 dd 2 0 i 0 ii 2 0 c 0 cc k kk j           (with time delay) (4.5) j is determined when the controller parameters are subjected to ;kkk maxppminp  ;kkk aximiinim  and ;kkk maxddmind  ;maxmin cpc kkk  ;maxmin iii   and ;maxddmind   step 6: update velocity. for each particle, the velocity can be updated by design of robust stabilizing pi/pid using particle swarm optimization 103 copyright ©2018 assa. adv. in systems science and appl. (2018) )xgbest()(randc)xpbest()(randcvwv n i n i2 n i n i1 n i 1n i  (4.6) step 7: update position. each particle changes its position by adding the updated velocity to the previous position and it is represented as. 1n i 0 i vx)j,i(x  (4.7) step 8: repeat steps 5 to 7 until maximum generations are completed pso algorithm is run for each particle to evaluate fitness function several times and better results are saved and applied to the proposed pi/pid controller. 5. design of robust stabilizing controller in this section, a design procedure for a robust pi/pid controller of a plant with parametric uncertainty is illustrated. example 1 consider a wing aircraft [18] whose transfer function with parametric uncertainty is given by ]1.0,1.0[s]9.33,1.30[s]8.80,4.50[s]6.4,8.2[s ]166,90[s]74,54[ )d,c,s(g 234    (5.1) as the necessary conditions 0ab ii  (for i=0, 1, 2, 3, 4) are not satisfied for the above characteristic polynomial, hence the given interval plant is unstable. thereby, it is required to design a robust controller, which stabilizes the given plant. design of pi controller: the transfer function of the pi controller is given by )s(d )s(n s k k)s(c c ci ppi  then the closed loop transfer function with a pi controller becomes ]k166,k90[s]1.0k74k166,1.0k54k90[ s]9.33k74,1.30k54[s]8.80,4.50[s]6.4,8.2[s]1,1[ ]k166,k90[s]k74k166,k54k90[s]k74,k54[ )s(t iiipip 2 pp 345 iiipip 2 pp     (5.2) from the above equation, the characteristic equation of the closed loop interval system with pi controller can be taken as 0]k166,k90[s]1.0k74k166,1.0k54k90[ s]9.33k74,1.30k54[s]8.80,4.50[s]6.4,8.2[s]1,1[ iiipip 2 pp 345   (5.3) the step response pi controller ( pk = 1.3172 and ik = 1.9378) using ziegler-nichols settings [41] are shown in figure 5.1. from this figure 5.1, it has been observed that the designed pi controller from ziegler-nichols tuning method cannot stabilize the given interval plant at all operating conditions. hence it is necessary to redesign the pi controller to stabilize given interval plant. by applying the necessary and sufficient conditions from table 1 to the above 5th order polynomial (5.3), the following set of inequality constraints are obtained. in order to make this set of constraints into the feasible closed set, a small positive number ‘ε’ is introduced into the constraints. hence the optimization problem can be stated as to find pk and ik such that the objective function dt)t(ej 0 2   is minimized, subjected to the following constraints. inequality constraints for proposed method: necessary conditions: 104 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) 0k90 i   01.05490  ip kk 01.3054  pk sufficient conditions: 04.305401.9068.32502916 2  ipp kkk 04.029666416.2540  ip kk fig.5.1. closed loop step response with pi controller for all extreme plants using the ziegler-nichols the linear programming problem consists of two decision variables and five constraints. the controller parameters pk and ik are restricted to small values by choosing the objective function j properly. the purpose of using a small positive number ε is to formulate a feasible set closed. in this work, the pso algorithm proposed in section 4 is used to minimize objective function. it attempts to explain the problems of minimization, subjected to linear as well as nonlinear constraints. by applying the proposed algorithm, then the values of controller parameters are obtained as pk = 0.5766 and ik = 0.01. the closed loop step response of the system with a pi controller for both proposed methods ( pk = 0.5766 and ik = 0.001) and the method given in [18] ( pk = 0.5 and ik = 0.1) are shown in figures 5.2 and 5.3 for ε = 1 respectively. the nominal pi controller parameters 655.11and6.0k 0 i 0 c   and nominal pid controller parameters 993.668.0k 0 i 0 c   and 74825.10 d are designed by the ziegler-nichols settings. the time domain specifications of figures 5.2 and 5.3 are shown in table 5.1 and which describes the efficacy of the proposed method for the design of the robust pi controller. the step response comparison of four extreme plants with a pi controller obtained by the proposed method and the method given in [18] is shown in figure design of robust stabilizing pi/pid using particle swarm optimization 105 copyright ©2018 assa. adv. in systems science and appl. (2018) 5.4. fig. 5.2. closed loop step response to the pi controller for all extreme plants using the proposed method. in order determine the robustness of proposed method, the given interval plant g(s,c,d) from equation (5.1) can be written as s32s64s7.3s 128s64 )s(g 234    (5.4) here g(s) is obtained by averaging the upper and lower bound of s coefficients of g(s,c,d). the closed transfer function of wing aircraft with pi controller is given by iip 2 p 345 iip 2 p1 k128s)k64k128(s)32k64(s6.65s7.3s k128s)k64k128(sk64 )s(t    (5.5) the closed loop step response of the system with pi controller ( pk = 0.5766 and ik = 0.001) from proposed methods is shown in figure 5.5. the pi controller designed from proposed method can stabilize the given plant if any changes occur in the uncertain parameters within the bounds . 106 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.3. closed loop step response to the pi controller for all extreme plants using the method given in [18]. fig.5.4. closed loop step response to the pi controller for all extreme plants for proposed method and the method given in[18]. design of robust stabilizing pi/pid using particle swarm optimization 107 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.5. closed loop step response to the pi controller for changes in the uncertain parameter to determine the robustness for proposed method table 5.1. time domain specifications for proposed method and existing method in [18] it has been observed from figures 5.2 and 5.3 that the designed pi controller, which uses the proposed stability conditions, robustly stabilizes the plant very quickly when compared to the method given in [18]. from table 5.1, the designed pi controller stabilizes the plant with lesser time domain parameters than the existing method. it has been observed from figure 5.5, that the designed pi controller from the proposed method is used to determine the robustness in any changes in the uncertain parameters of the given plant. in our proposed method, the controller parameters are obtained based on the minimization of the objective function (ise). this integral square error obtained from our proposed method is less compared to other methods available in the literature. this shows the efficacy of the proposed method in terms of time domain specifications and the ise. the proposed method involves five sets of equations for nlp to solve. whereas the method in [18] requires eight set of equations. thus, the proposed method requires less computational complexity than the method given in [18]. it is also observed that the computation time required for solving the name of the kharitonov polynomial proposed method existing method in [18] % peak overshoot mp peak time tp(sec) rise time tr(sec) settling time ts(sec) ise 10-4 % peak overshoo t mp peak time tp(sec) rise time tr(sec) settling time ts(sec) ise 10-3 first 17.492 3.514 1.501 7.0534 4.22 32.4147 4.008 1.508 10.808 1.101 second 20.826 1.707 0.878 5.0156 0.093 24.3994 2.488 0.880 8.2771 0.158 third 29.702 3.683 1.349 14.229 1.53 57.6729 3.820 1.378 26.828 16.00 fourth 31.419 1.988 0.718 7.7175 0.631 42.659 2.166 0.741 8.326 0.022 108 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.6. closed loop step response with pid controller for all extreme plants using the ziegler-nichols nlp problem with minimum number of equations using proposed stability conditions and pso algorithm is much less than the method given in [18]. thus, the developed pi controller using necessary and sufficient conditions of interval polynomial robustly stabilizes the wing aircraft. these stability conditions can be implemented easily for determining the stability of higher order interval plants. design of pid controller the transfer function of the pid controller is given by )s(d )s(n sk s k k)s(c c c d i ppid  then the closed loop transfer function with pid controller becomes ]k166,k90[s]1.0k74k166,1.0k54k90[ s]9.33k74k166,1.30k54k90[ s]8.80k74,4.50k54[s]6.4,8.2[s]1,1[ ]k166,k90[s]k74k166,k54k90[ s]k74k166,k54k90[s]k74,k54[ )s(t iiipip 2 pdpd 3 dd 45 iiipip 2 pdpd 3 dd       (5.6) from the above equation, the characteristic equation of the closed loop interval system with pid controller can be taken as 0]k166,k90[s]1.0k74k166,1.0k54k90[ s]9.33k74k166,1.30k54k90[ s]8.80k74,4.50k54[s]6.4,8.2[s]1,1[ iiipip 2 pdpd 3 dd 45    (5.7) the step response pid controller ( pk = 1.3172, ik = 1.9378 and kd=0.1791) using design of robust stabilizing pi/pid using particle swarm optimization 109 copyright ©2018 assa. adv. in systems science and appl. (2018) ziegler-nichols settings [41] are shown in figure 5.6. from this figure 5.6, it has been observed that the designed pi controller from ziegler-nichols tuning method cannot stabilize the given interval plant at all operating conditions. hence it is necessary to redesign the pid controller to stabilize given interval plant. by applying the necessary and sufficient conditions for the above 5th order polynomial (5.7), six sets of inequality constraints are obtained. the controller parameters dip kandk,k are obtained by solving these inequality constraints using pso such that the objective function dt)t(ej 0 2   is minimized. then the controller parameters pk = 0.5933, ik = 0.001 and dk =0.252 are obtained. the closed loop step response of system with pid controller using the proposed method for all extreme plants is shown in figure 5.7 for ε = 1 respectively. the time domain specifications of figure 5.7 are tabulated in table 5.2 which describes the efficacy of the proposed method. from the equation (5.4) the closed transfer function of wing aircraft with pid controller is given by ]k128s)k64k128(s)32k64k128(s)6.65k64(s7.3s,1[ k128s)k64k128(s)k64k128(sk64 )s(t iip 2 pd 3 d 45 iip 2 pd 3 d1    (5.8) the closed loop step response of the system with pid controller ( pk = 0.5933, ik = 0.001 and dk =0.252) from proposed methods is shown in figure 5.8.the pid controller designed from proposed method can stabilize can stabilize the given plant if any changes occur in the uncertain parameters within the bounds. table 5.2. time domain specifications for proposed method name of the kharitonov polynomial % peak overshoot mp peak time tp(sec) rise time tr(sec) settling time ts(sec) ise *10-6 first 4.934 4.179 1.4975 6.3771 10.381 second 15.79 2.096 0.9614 5.0017 0.1072 third 12.41 4.015 1.8559 6.9367 56.76 fourth 15.26 2.189 0.9886 5.4216 0.4.348 it has been observed from the simulation results of figure 5.7 that the designed pid controller robustly stabilizes the plant very quickly. from table 5.2, it is evident that the designed pid controller stabilizes the plant with lesser time domain parameters. the proposed method involves six sets of equations for nlp to solve. thus, the proposed method requires less computational complexity than the methods in the literature it has been observed from figure 5.8, that the designed pid controller from the proposed method is used to determine the robustness in any changes in the uncertain parameter of the given plant this shows the efficacy of the proposed method in terms of time domain specifications. it is also observed in the computation time required for solving the nlp problem using stability conditions and pso algorithm is much less. thus, the developed pid controller using necessary and sufficient conditions of interval polynomial robustly stabilizes the wing aircraft. 110 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.7. closed loop step response with pid controller for all extreme plants using the proposed method. fig.5.8. closed loop step response to the pid controller for changes in the uncertain parameter to determine the robustness of the proposed method example 2 consider the dynamics of a vast range of chemical processes which is described by the interval first-order plus dead time (fopdt) model [30] is given by sdte 1s k )s(g     (5.9) where ,]9,3[k ]6,1[t,]18,10[ d  in this, the nominal value variations in the process gain k, time constant and dead time td are taken as ±50%, ±28.57%, and ±71.43%. the nominal design execution is continued based on the nominal values of .5.3,14,60  dtk  the nominal pi controller parameters 655.11and6.0k 0 i 0 c   and nominal pid controller parameters design of robust stabilizing pi/pid using particle swarm optimization 111 copyright ©2018 assa. adv. in systems science and appl. (2018) 993.668.0k 0 i 0 c   and 74825.10 d are designed by the ziegler-nichols settings [41]. the step response pi/pid controller using ziegler-nichols is shown in figure 9 and 10. from this figure 9 and 10,it has been observed that the designed pi/pid controller from zieglernichols cannot stabilize the given interval plant at all operating conditions. hence it is necessary to redesign the pi/pid controller to stabilize given interval time delay plant. from the corollary 3.1, se ]61[ can be written as . s]5.65.0[1 s]5.65.0[1   then the following delay-free model for original interval fopdt is given as 1s]5.24,5.10[s]117,5[ ]9,3[s]5.58,5.1[ )d,c,s(g 2    (5.10) fig. 5.9. : closed loop step response to the pi controller for all extreme plants using ziegler-nichlos method with time delay. fig. 5.10: closed loop step response to the pid controller for all extreme plants using ziegler-nichlos method with time delay. design of pi controller: 112 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) consider the pi controller of the form )s(d )s(n ) s 1 1(k)s(c c c i cpi   then the closed loop transfer function with a pi controller becomes ]k9,k3[s]k5.1k9,k5.58k3[ s]k5.15.24,k5.585.10[s]117,5[ ]k9,k3[ s]k5.1k9,k5.58k3[s]k5.58k5.1[ )s(t cccicicici 2 iciici 3 ii cc ciccic 2 ic,ic         (5.11) from the above equation, the characteristic equation of the closed loop interval system with pi controller can be taken as 0]k9,k3[s]k5.1k9,k5.58k3[ s]k5.15.24,k5.585.10[s]117,5[ cccicicici 2 iciici 3 ii     (5.12) by applying the necessary and sufficient conditions for the above 3rd order polynomial (5.12), the following set of inequality constraints is obtained. in order to make this set of constraints into the feasible closed set, a small positive number ‘ε’ is introduced into the constraints. hence the optimization problem can be stated as to find ck and i such that the objective function 2 0 i 0 ii 2 0 c 0 cc k kk j       is minimized, subjected to the following constraints. necessary conditions: 0k3 c   0k5.58k3 icic   0k5.585.10 ici   05 i   sufficient condition: 0)k5.124(k27)k5.58k3( icic 2 cici   the linear programming problem consists of two decision variables and five constraints. the controller parameters kc and i are restricted to small values by choosing the objective function j purposely. in this work, the pso algorithm proposed in section 4 is used to minimize the objective function .j by applying the proposed algorithm, then the values of controller parameters are obtained. as shown in table 5.3, the values of the controller parameters ck and i increase as ‘ε’ are increased which shows the sensitivity of the controller parameters with respect to the nlp parameter ‘ε’. the closed loop step response of the system with the pi controller by the proposed method ( kc = 4.0432 and i = 0.00926) and the method given in [30] ( kc = 0.0684 and i = 12.5102) are shown in figures 5.11 and 5 .12 for ε = 0.05 respectively. the time domain specifications of figures 5.11 and 5.12 are shown in table 5.4 which describes the efficacy of the proposed method. in order determine the robustness of proposed method, the given interval plant g(s,c,d) from equation (5.10) can be written as 1s5.17s61 6s30 )s(g 2    (5.13) here g(s) is obtained by averaging the upper and lower bound of s coefficients of g(s,c,d). the closed transfer function of wing aircraft with pi controller is given by design of robust stabilizing pi/pid using particle swarm optimization 113 copyright ©2018 assa. adv. in systems science and appl. (2018) ccici 2 ici 3 i ccic 2 ic k6s)k30k6(s)k305.17(s61 k6s)k30k6(sk30 )s(t      (5.14) the closed loop step response of the system with pi controller ( kc = 4.0432 and i = 0.00926) from proposed method is shown in figure 5.13. the pi controller designed from proposed method can stabilize the given plant if any changes occur in the uncertain parameters within the bounds . table 5.3. variation of ck and i for different values of ε controller set  proposed method existing method ck i ck i 1 0.05 0.00926 4.0432 0.0684 12.5102 2 1 0.01389 8.0934 0.0683 12.5110 3 1.5 0.01803 10.4032 0.0683 12.5116 4 2 0.02059 11.9023 0.0683 12.5121 table 5.4. time domain specifications for proposed method it has been observed from the figure 5.11 and 5.12 that the designed pi controller, which uses the proposed stability conditions robustly stabilizes the plant when compared to the method given in [30]. from figure 5.12, it is observed that the closed loop interval time delay system with a controller is unstable by the method in [30]. whereas closed loop interval time delay system with proposed method is stable. it has been observed from figure 5.13, that the designed pi controller from the proposed method is used to determine the robustness in any changes in the uncertain parameters of the given plant. this shows that proposed method robustly stabilizes the interval time delay process plant. the proposed method has five inequality constraints which are less compared with the method (which has seven inequality constraints) in [30]. hence the proposed method is simple and requires less computation as it uses a lesser number of inequality constraints than the method given in [30]. it is also observed that the computation time required for solving the nlp problem using proposed stability conditions and pso algorithm is 3.872603 seconds, which is much less than the method given in [30]. thus, the developed pi controller using necessary and sufficient conditions of interval polynomial robustly stabilizes the interval process time delay system. these stability conditions can be implemented easily for determining the stability of higher order interval plants. name of the kharitonov polynomial rise time tr (sec) settling time ts (sec) first 222.2482 422.5395 second 294.5659 526.1731 third 43.3195 139.2935 fourth 92.4901 166.2970 114 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.11.closed loop step response with a pi controller for all extreme plants using proposed method with time delay. fig.5.12. closed loop step response to the pi controller for all extreme plants using the existing method [30]. design of robust stabilizing pi/pid using particle swarm optimization 115 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.13. closed loop step response to the pi controller for changes in the uncertain parameter to determine the robustness for proposed method design of pid controller: consider the pid controller of the form )s(d )s(n ) s 1 1(k)s(c c c d i cpid    then the closed loop transfer function with pid controller becomes ]9,3[]5.19,5.583[]5.195.24 ,5.5835.10[]5.1117,5.585[ ]9,3[]5.19,5.583[ ]5.19,5.583[]5.58,5.1[ )( 2 3 23 ccciciciciicdici icdicidicidici ccciccic icdicicdicdicdic kkskkkkskk kkskk kkskkkk skkkkskk st          (5.15) from the above equation, the characteristic equation of the closed loop interval system with pid controller can be taken as 0]k9,k3[s]k5.1k9,k5.58k3[ s]k5.1k95.24,k5.58k35.10[s]k5.1117,k5.585[ cccicicici 2 icdiciicdici 3 dicidici     (5.16) by applying the necessary and sufficient conditions of equations for the above 3rd order polynomial (5.16), set of inequality constraints are obtained. hence the optimization problem can be stated as to find ,kc i and d such that the objective function 2 0 d 0 dd 2 0 i 0 ii 2 0 c 0 cc k kk j           is minimized. by using the proposed method given in section 4 to the above nlp problem, the controller parameters ,kc i and d are obtained. as shown in table 5.5, the values of the controller parameters ,kc i and d increase as ‘ε’ are increased which shows the sensitivity of the controller parameters with respect to the nlp parameter ‘ε’. they are given by 1520.3,00914.0k ic   and d = 1.2040. the closed loop step response of system with pid controller for proposed method is shown in figure 5.14 for ε = 1. the time domain 116 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) specifications of figure 5.14 are shown in table 5.6 which describes the efficacy of the proposed method in terms of the time domain specifications. table 5.5. variation of ,kc i and d for different values of ε table 5.6. time domain specifications for proposed method from equation (5.10) the closed loop transfer function of wing aircraft with pid controller is given by ccici 2 icdici 3 dici ccic 2 icdic 3 dic k6s)k30k6[s)k30k95.17(s)k3061( k6s)k30k6(s)k30k6(sk30 )s(t      (5.17) the closed loop step response of the system with pid controller ( 1520.3,00914.0k ic   and d = 1.2040.) from proposed methods is shown in figure 5.15. the pid controller designed from proposed method can stabilize the given plant if any changes occur in the uncertain parameters within the bounds. controller set  ck i d 1 0.05 0.00914 3.1520 1.2040 2 1 0.00925 6.8461 1.7481 2 1.5 0.00930 7.1864 1.8998 4 2 0.00977 7.4979 2.1075 name of the kharitonov polynomial rise time tr (sec) settling time ts(sec) first 152.3465 283.6998 second 230.6303 381.7915 third 29.0473 131.9949 fourth 72.5310 119.3718 design of robust stabilizing pi/pid using particle swarm optimization 117 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5.14. closed loop step response with pid controller for all extreme plants using the proposed method with time delay. fig.5.15. closed loop step response to the pid controller for changes in the uncertain parameter to determine the robustness with time delay for the proposed method it has been observed from the simulation results of figure 5.14 that the designed pid controller robustly stabilizes the plant very quickly. from table 5.6, it is evident that the designed pid controller stabilizes the plant with lesser time domain parameters. this shows that proposed method robustly stabilizes the interval time delay process plant. it has been observed from figure 5.15, that the designed pi controller from the proposed method is used to determine the robustness in any changes in the uncertain parameters of the given plant. the proposed method has five inequality constraints which are less compared with the method (which has seven inequality constraints) in [30]. hence the proposed method is simple and requires less computation as it uses a lesser number of inequality constraints than 118 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) the method given in [30]. this shows the efficacy of the proposed method in terms of time domain specifications. it is also observed that the computation time (4.63 seconds) required for solving the nlp problem using the proposed stability conditions and pso algorithm is much less. thus, the developed pid controller using necessary and sufficient conditions of interval polynomial robustly stabilizes the interval process time delay system. 6. conclusions a robust stabilizing pi/pid controller is designed for interval process plant with and without time delay based on the newly developed necessary and sufficient conditions using pso algorithm. a set of inequality constraints in terms of controller parameters are derived from interval polynomial based on the new necessary and sufficient conditions. consequently, these inequalities are solved using pso to obtain controller parameters. the proposed pi/pid controller procedure is also applied and demonstrated through typical numerical examples. it is observed that the designed pi/pid controller robustly stabilizes the plant with and without time delay with lesser time domain parameters than the existing methods. the pi/pid controller designed from proposed method can stabilize the given plant if any changes occur in the uncertain parameters within the bounds the simulation results are evidence for its robustness in stabilizing the interval process time delay plant using a pi/pid controller. also, the proposed method is simple and involves less computational complexity in comparison with the methods available in the literature. references 1. astrom, k. j., hagglund, t., hang, c. c. & ho, w.k. (1993). automatic tuning and adaptation for pid controllersa survey, control engineering practice, 1(4), 699-714. 2. astrom, k. j. & hagglund, t. (1995). pid controllers: theory, design, and tuning, isa: the instrumentation, systems, and automation society, research triangle park, nc. 3. ho, w. k., hang, c. c. & cao, l. s. (1995). tuning of pid controllers based on gain and phase margin specifications, automatica, 31(3), 497-502. 4. ho, m. t., datta, a. & bhattacharyya, s. p. (1996). a new approach to feedback stabilization, proc. of the 35th ieee conf. on decision and control, 4(7), 4643-4648. 5. yongling zheng, longhua, ma, liyan zhang & jixin qian. (2003). robust pid controller design using the particle swarm optimizer, intelligent control,ieee international symposium, paper-12, 974-979. 6. ho, m. t., datta, a. & bhattacharyya, s. p. (1997). a linear programming characterization of all stabilizing pid controllers, proc. of american control conf, 6(23,). 3922-3928. 7. ho, m. t., datta, a.& bhattacharyya, s. p. (1997). a new approach to feedback design part ii: pi and pid controllers, dept. of electrical eng., texas a&m univ., college station, tx, tech. report tamuece 97001-b. 8. satpati, biplab, chiranjib koley & subhashis datta.(2014). robust pid controller design using particle swarm optimization-enabled automated quantitative feedback theory approach for a first-order lag system with minimal dead time, systems science and control engineering, 2 (1), 502-511. 9. kalyani, t. & siva kumar, m. (2013). computational of optimal stabilizing pid controller for interval plants using pso, international journal engineering research and technology,2(10), 2993-2998. design of robust stabilizing pi/pid using particle swarm optimization 119 copyright ©2018 assa. adv. in systems science and appl. (2018) 10. barmish, b. r. (1989). a generalization of kharitonov’s four polynomial concept for robust stability problems with linearly dependent coefficient perturbations, ieee trans. automatic control, 34(2),157165. 11. deore, p.j. & patre, b.m. (2005). design of robust compensator for jet engine: an interval analysis approach, proc. of the ieee conf. on control, toronto, ont., canada,376381. 12. nie, y.y. (1976). a new class of criterion for the stability of the polynomial, act.mechnica sinica, 15(1),110-116. 13. chen, c.t. & wang, m.d. (1997). robust controller design for interval process system', computer and chemical engineering, .21(7), 739-750. 14. chen, ct. fan, ck. & peng, st. (2004). robust control system design for interval time delay process systems. chin inst chem engrs, 35(2), 213–223. 15. bhattachrya,s.p.(1987). robust stabilization against structured perturbations', lect. notes in control and information sciences, springer verlag, berlin, 99, 78-99. 16. satpati, b., & sadhu, s. (2008). course changing control for a cargo mariner class ship using quantitative feedback theory. ie(i), marine divisional journal, 88, 1–8. 17. patil, m. d., nataraj, p. s. v., & vyawahare, v. a. (2012). automated design of the fractional pi qft controller using interval constraint satisfaction technique (icst). nonlinear dynamics, 69(3), 1405–1422. 18. patre, b.m. & deore, p.j. (2003). robust stabilization of interval plants’, european control conf. ecc-03, university of cambridge (uk), pp.1660-1663. 19. patre, b.m. & deore, p.j. (2007). robust stability and performance for interval process plants, isa transactions, 46(3),343-349. 20. siva kumar, m., ramalinga raju,m., srinivasa rao, d. & bala bhargavi, t. (2015). design of robust pi and pid controller for dc motor fuzzy parametric approach, global journal of mathematical sciences (gjms), special issue for recent advances in mathematical sciences and applications-13, 2(2),77–83. 21. siva kumar, m. & bala bhargavi, t. (2013). design of robust pi/pid controller discrete-time parametric uncertain system', international journal of scientific research,.2(10),5-9. 22. srinivasa rao,d., siva kumar,m. & ramalinga raju,m.(2018). new algorithm for the design of robust pi controller for plants with parametric uncertainty, trans. of the institute of measurement and control, 40(5), 1481-1489. 23. rahul mittal & manisha bhandari. (2015). design of robust pi controller for active suspension system with uncertain parameters, ieee conf. on signal processing, computing and control at waknaghat isbn: 978-1-4799-8435-0, 333-337. 24. ghosh b. k. (1985). some new results in the simultaneous stabilization of a family of single input, single output systems, systems and control letters, 6(1), 39-45. 25. hollot, c.v. & yang, f. (1990). robust stabilization of interval plants using lead or lag compensators, systems and control letters, 14(1), 9-12. 26. barmish, b. r., hollot, c. v., kraus, f. j. & tempo, r. (1992). extreme point results for robust stabilization of interval plants with first order compensators, ieee trans. on automatic control, 37(6),707714. 120 d.srinivasa rao, m. siva kumar and m. ramalinga raju copyright ©2018 assa. adv. in systems science and appl. (2018) 27. jian’an fang, da zheng, & zhengyun ren. (2009). computation of stabilizing pi and pid controllers by using kronecker summation method, energy conversion and management, 50(7),1821–1827. 28. nusret tan, ibrahim kaya, celaleddin yeroglu & derek p. atherton. (2006). computation of stabilizing pi and pid controllers using the stability boundary locus, energy conversion and management, 47(18–19),3045–3058. 29. barmish, b. r. & kang, h. i. (1993). a survey of extreme point results for robustness of control systems', automatica, 29(1),13-35. 30. patre, b.m. & deore, p.j. (2011). robust stability and performance of interval process plant with interval time delay', transactions of the institute of measurements and control, 34 (5), 627-634. 31. farkh, rihem, kaouther laabidi & mekki ksouri. robust pi/pid controller for interval first order systems with time delay, international journal of modelling, identification and control13.1-2 (2011): 67-77. 32. brosilow cb. (1979. the structure and design of smith predictors from viewpoint of inferential control, in: joint automatic control conference, 16, 288. 33. laughlin dl, rivera de & morari m. (1987). smith predictor design for robust performance, international journal of control, 46(2), 477-504. 34. wang, zq. & skogestad, s. (1993). robust control of time delay systems using the smith predictor, international journal on control 57: 1405. 35. kennedy, j. & eberhart, r.c.(1995) ’particle swarm optimization’, proc. ieee international conf. on neural networks iv, piscataway, nj, 1942-1948 36. gaing, z-l. (2004) ‘a particle swarm optimization for optimum design of pid controller in avr system', ieee trans. on energy conversion,19 (2), 384-391. 37. angeline p.j. (1998). using selection to improve particle swarm optimization, proc.ieee int. conf. evol. comput., anchorage, a k, 84–89. 38. anderson, bdo., jury, ei. & mansour, m. (1987). on robust hurwitz polynomials', ieee trans. on automatic control, 32 (10), 909-913. 39. guillemin, e.a. (1962). the mathematics of circuit analysis, 7th ed. calcutta, india: oxford and ibh publishing, 327-330. 40. optimization toolbox: math works inc, user guide 2010. 41. ziegler, j.b. and nichols, n.b. (1942). optimum settings for automatic controllers, trans. asme, 64,759-768. advances in systems science and applications (2013) vol. 13 no. 1 1-20 mechanisms of organizational behavior control: a survey v.n. burkov, m.v. goubko, n.a. korgin and d.a. novikov institute of control sciences, russian academy of sciences, moscow, russia abstract we survey basics of a version of the mechanism design theory tailored to solve management problems, and introduce the core concept of a mechanism of organizational behavior control. we discuss methodological grounds of the theory, give classification of mechanisms, and characterize the mechanism implementation process. finally, we sketch basic mechanisms which help solving important management problems on all stages of management cycle and provide an example of the mechanism of incentive-compatible planning. keywords mechanism design, management theory, organizational behavior 1 introduction in xxi century economics is transforming into the economics of knowledge. contemporary trends in management science consider workforce not as just a “yet another production factor” (in line with land and capital), but promote it at least to the status of intangible asset, which possesses rational economic behavior with the abilities of self-regulation and self-development. the reason is that employees’ decisions, based on their skills and highly influenced by their motivation, now play the crucial role in a value added chain. business process management (bpm) [1] and business process reengineering (bpr) [2] techniques did not cover motivational and decision-making aspects of business processes and, thus, they alone were insufficient to drive corporate reengineering efficiently. we claim that it is lack of attention paid to employees’ behavioral response to changes was the main reason of bpr projects failure1. the major differences between a complicated technical system and a human (or a collective) as a control object are condensed in the concept of activeness. the first aspect of activity consists in human ability of independent goalsetting. a technical system has no interests beyond interests of its designer. but employees in an organization do pursue their own objectives, which can be inconsistent with objectives of the organization. the second aspect lies in human ability to choose actions independently; in particular, an employee can deliberately manipulate information (when he or she finds it profitable) and/or not fulfill the assigned plans and orders (again, if this promises some benefit). 1the citation attributed to m. hammer “i was reflecting my engineering background and was insufficient appreciative of the human dimension. i’ve learned that’s critical” is in line with the above proposition. 2 v.n. burkov: mechanisms of organizational behavior control: a survey the third important aspect of activity is the ability of reflexion regarding his or her personal activity and activity of the other subjects (including the ability of forecasting their behavior). in this paper we survey the approach, which allows accounting systematically for the phenomenon of activity of employees in organizations. the approach is based on the methodology of systems and control sciences, and widely employs results of mechanism design – a branch of game theory, which deals with conflict situations involving the principal and a set of active agents (usually, in the presence of asymmetric information). mechanism design theory delivers a solution to management problems in the form of a control mechanism, i.e., a formalized routine of decision-making. for a control mechanism to keep efficiency in the presence of active agents, it must be robust to information manipulation, plan non-fulfillment and other aspects of activeness [3]. the main challenge of mechanism design in management is search of the so-called correct mechanisms, which assure both interests coordination for all organization employees including stakeholders, management, and workers, as well as motivate them to report true information, which forms the basis of any efficient decision. to adopt mechanism design to practical problems of business administration we developed a tailored and, in some aspects, more straightforward theory. the main goal of this paper is to provide the reader with basic methodological and technical grounds of control mechanisms in organizations. in section 1 we briefly explain a formal scheme of accounting for behavioral response by the principal (a manager in an organization). the art of management is positioned as ensuring the desired behavior of subordinates. in section 2 models of organizational behavior are considered in the context of the production and management activity. in section 3 we introduce different types of mechanisms in an organization classified by the element of the organizational system model being affected by the mechanism. in section 4 we describe an applied technology of mechanism analysis, synthesis, and implementation in an organization. in section 5 the toolkit of mechanisms developed during last decades is briefly sketched, while in section 6 an example of a mechanism to control employees’ behavior is described in more detail. 2 principal-agent model consider a formal scheme of making a management decision by the principal (a manager in an organization). the problem of the principal is to analyze the available information and to choose the decision relevant to the current situation; so, the management decision appears to be the action of a principal. the principal compares possible options (feasible decisions) and their consequences and advances in systems science and applications (2013) vol.13 no.1 3 then chooses the best one. if the decision touches interests of the other persons (employees or some external parties), to predict properly the consequences of his or her decisions the principal has to forecast the response of such counterparties, with subordinates being the most important ones, a certain decision. a subordinate (an individual or a team) possesses exactly the same abilities and comprehension as a principal. indeed, they have personal preferences and interests; they make decisions and perform actions. the only difference is that the situation for a subordinate includes a decision made by the principal. the action of the subordinate, being it predicted or real, is accounted by the principal when making decisions. so, we obtain the simplified scheme of interaction between the principal and the subordinate (see fig. 1). fig. 1 the scheme of interaction between the principal and the subordinate in terms of control theory [4], we have a control subject (a manager) and a controlled system, i.e., a control object (a subordinate). state of the control object depends on external disturbances, actions of the control subject and, probably, on actions of the object, see fig. 1 (if the control object appears active, especially in organizational systems). a problem of the control subject consists in performing control actions (or, management decision, see fig. 1) to ensure a required state of the control object. this is done on the basis of information on external disturbances (or external situation). generally speaking, the art of management lies in ensuring the desired behavior of subordinates. consequently, the principal wants to select a mechanism such that subordinates choose desired actions. the problem is decomposed into, first, the analysis problem (given a behavioral model of a subordinate, find actions he or she would select under a certain control mechanism), and, second, the design problem (find a mechanism assuring the desired subordinates’ actions according to a given behavioral model). solving the design problem requires the ability to solve the analysis problem. hence, the principal should forecast behavior of subordinates as their response 4 v.n. burkov: mechanisms of organizational behavior control: a survey to his or her management decision. interaction demonstrated in fig. 1 represents a “closed loop”. let us replace the principal with a control mechanism, viz., a management decision procedure used by the principal. the inputs of this procedure are given by external environment and an action of a subordinate, while the output is a specific management decision. thus, we have the structure presented by fig. 2; here the control mechanism determines management decisions made by the principal. fig. 2 the scheme of interaction between the principal and the subordinate 3 models of employees’ behavior control theory considers a control object as passive one, possessing no individual preferences and information. the constraints imposed on his or her activity are unique and defined by the planned state, while the action (actual state) coincides with the plan. control of a passive object consists in assigning “plans” or in specifying the requirements to actual state, i.e., to the results of activity. in fact, there is no need in control subject’s pondering whether a passive control object performs the plan or not! the model of a passive object seems natural for technical systems. but, more surprisingly, many management theories imply that subordinates surely strive to implement the orders (in other words, it is assumed that the actual state equals to the planned one). planned and actual states may not coincide due to the following reasons: • presence of uncertainty (uncontrolled external factors that influence the result of the employee’s activity, making it differ from his or her action); • employee’s activity. both aspects should be considered when solving problems of management in organizations. the following fact has been emphasized earlier. the ability of employees to set goals, to choose actions independently, is reflected by the concept of activity of the control objects (agents) in organizational systems. it implies that the agents have certain opportunities and scope for independent goal-setting and making advances in systems science and applications (2013) vol.13 no.1 5 decisions regarding their actions. this is done, first, within the framework of the full-fledged structure of activity. next, we list the effects of activity: data manipulation, choosing a state varying from the plan, a negligent behavior, etc. these effects should be considered when designing and applying control mechanisms. efficient control makes it necessary to model agents’ behavior, i.e., to predict their reaction to specific control actions. thus is done using a concept of utility maximization, known in economics as the model of “economic man behavior” (we will use the term “agent” as a synonym). an agent acts (reports information, chooses actions, and so on) to maximize his or her utility. the stated concept has turned out fruitful and gained a dominating role in mathematical economics, decision theory and other scientific directions focused on the models of human behavior. as a rule, two types of economic agents (economics subjects) are identified, notably, • an economic subject of the market (the examples are an organization, a holding company, a firm, a corporation); • employee(s) of an economic subject. for both types of agents, the utility is typically represented by “economic profit”, that is, the difference between “revenue” and “costs”. yet, the content and methods used to represent the income and costs depend on the agent’s type. suppose that a department or a separate organization is studied; its revenue and costs are calculated in a management accounting system. the utility could be derived from the operational income according to the results of business activity (under the adopted policy of management accounting). if an agent is an individual employee, the management accounting system stores income, but not efforts exerted by the agent to perform the operational activity and corresponding costs. moreover, the information system does not weight costs against the income. thus, it is necessary to compare efforts and costs with rewards received by the agent. this sort of relationship is set by virtue of different supervision techniques, timing, as well as by the search for analogs. a specific situation being analyzed, one should choose an approach enabling the simplest and the most accurate definition of the agent’s effort function and utility function. the principal is also an “economic man”. he or she also has a utility function, but has different possibilities to implement his or her activity (in particular, management authorities). the principal also has internal constraints, utility, awareness and actions. the principal’s action consists in choosing a control mechanism, i.e., the relationship between control actions and actions of the agent. the major difference between the principal and agent lies in that the former possesses management authority; notably, he or she has the right “to make the 6 v.n. burkov: mechanisms of organizational behavior control: a survey first move” by establishing activity actions (a control mechanism) for the agent. the principal’s utility generally depends on the actions of the agent, i.e., control efficiency is determined by the utility of the principal (representing the interests of the whole organizational system) gained as the result of agent’s activity. therefore, the mechanism in the described environment specifies a “legislative, regulating and mandatory base”, i.e., the rules and procedures involved to perform actions of the principal and the agent during implementation of processes and projects. moreover, control rises to the level of meta-control, to formation of the “rules of the game”, conditions of the agent’s functioning and motivation. in a certain sense, the principal appears an architect of social and economic rules and a meta-player. according to the general model [4], an economic agent is described by four basic parameters aggregating the procedural components of his or her activity: 1) constraints and norms of activity; 2) an utility function; 3) awareness; 4) action of the agent (it is chosen on the basis of his or her awareness under existing constraints and utility). the action indicates of the agent’s state and, to a considerable degree, determines the result of his or her activity. under the existing constraints, economic agent has two types of possible actions. they are: • reporting information on uncertain parameters to the principal and/or to other agents; • choosing an action (production output, time consumed, etc). in fact, what is the incentive-compatible control of an economic agent? the matter concerns, e.g., the fact that the principal never exactly knows the agent’s subjective estimate function of his or her actions. in other words, the principal does not know how much the agent is willing to do for $1 reward. low level of motivation results in no productivity growth. starting from a certain moment (exactly when the reward exceeds the agent’s subjective estimate of his or her efforts), productivity increases. next, a moment comes when the principal no more benefits from raising the reward (costs to motivate additional efforts are greater than gains of productivity growth). roughly speaking, three intervals of rewards appear; the principal’s task lies in balancing the interests to make the situation convenient for him or her and for the agent (this is done over the admissible interval with respect to the costs and productivity). evaluating such balance of interests requires joint actions of economically rational persons (the principal and the agent). their common aim is making a decision that would be more efficient than in the case of independent actions, based on their individual interests. thus, a mechanism of control in an organization is a sort of a “visible advances in systems science and applications (2013) vol.13 no.1 7 hand”, which drives self-interested actors toward socially desirable outcomes. consider the case of a single principal and several agents. now the agents immediately choose their actions given the mechanism set by the principal. as soon as interests of agents differ, they appear to participate is a certain conflict. the model of such conflict situations is given by the game theory [5], which uses different concepts of equilibrium (e.g. nash equilibrium) to predict the outcome of conflict. so, the set of equilibrium outcomes of a game is considered as a behavioral response of a controlled system to a management decision in the case when a system consists of several agents. the assumption that the agents’ action coincides with the output of his or her activity makes a sort of simplification. actually, the result may vary from the action due to uncontrolled factors, namely, actions of the rest agents, the state of external environment, etc. in other words, the result of the agent’s activity generally depends on his or her action, actions of the other agents and impact exerted by the external environment. then, to predict agent’s behavior in the case of a single agent the principal employs the hypothesis of expected utility maximization, which implies that an agent chooses his or her action to maximize the expected utility function obtained by averaging agent’s utility over realization of random states of external environment. in the case of several agents the concept of bayesian nash equilibrium [5] or informational equilibrium [6] is employed to account for unknown external parameters. the case of asymmetric information in principal-agent models is also studied in detail by the contract theory [7]. so, game-theoretical and optimization models of human rational behavior allow analyzing the response of a control object to certain control action and form the basis of the synthesis phase of control mechanism design. 4 methods of control in organization above we defined control as purposeful impact on control object. yet, organization represents an intricate control object. hence, one should clarify the internal structure of an organization as a control object; in other words, it is necessary to find out what entities inside an organization can be influenced by the control actions (so as to change state, more specifically, the behavior of the organization). those components of organizational system being modified during the process (and as the result) of control are known as objects of control. expressing any complicated system as a complex of interacting elements makes a certain model. therefore, below we introduce the model of organizational system [4]. organizational system (os) is described via specifying: • staff (employees, their groups and collectives, its members); • structure (a set of informational, control, technological and other relations among the os members); 8 v.n. burkov: mechanisms of organizational behavior control: a survey • constraints and norms of activity imposed on os members; they are institutional, planned, technological and other of constraints and norms of individual and joint activity; • goals and preferences of os members; • information, i.e., data regarding essential parameters being available to os members at the moment of decision-making (choosing the strategies); • sequence of operations (of data acquisition and choice of actions by os members). the staff determines “who” is included into the system; the structure describes “who interacts with whom, who is subordinate”, etc. finally, feasible sets define “who can do what”, goal functions represent “who wants what”, and information states “who knows what”. control in os, being interpreted as an impact on the controlled system to ensure its required behavior, may affect each of the listed parameters. hence, taking the focus of control (the parameter of os, which is modified during the process of control and as a result of control) as the basis of classification of control in os, we obtain the following methods (types) of control [4]; staff control deals with the following issues: who is included into the organization or department, who should be dismissed or recruited? generally, staff control either includes the problems of personnel training and development. for example, the well-known model of signaling [8], one of the seminal models in contract theory, when considered from the principal’s point of view, can be treated as a mechanism of staff control, which allows a principal to implement a rational hiring policy in the presence of hidden information. structure control is as a rule performed in parallel to staff control. it provides answers to several questions, viz, what functions should be performed by whom, what participants are subordinate to whom, who should control and be controlled, what information should be transferred and acquired, etc. mechanisms to control structure are not well studied formally at this moment, although a number of competing models exists of a management hierarchy in a firm (see the surveys in [9-10]). institutional control appears to be the most stringent; it consists in that the principal restricts the sets of feasible actions and results of activity of the subordinates (in a purposeful way). such restriction may be implemented via explicit or implicit influence (legal acts, directives, orders, etc) or mental and ethical norms, corporate culture and so on. numerous variations of resource allotment mechanisms (with auctions being a very special case) commonly met in public institutions give us a good example of institutional control mechanisms [4, 11]. incentive control is in a certain sense “softer” than institutional one, and consists in purposeful modification of the preferences of control object (the subordiadvances in systems science and applications (2013) vol.13 no.1 9 nates). the described modification could be performed through introduction of a certain scheme of penalties and/or rewards for choosing specific actions and/or attaining definite results of activity. job contracts for individual employees and teams are a good example of incentive schemes, which lie in the very core of any organizational activity [4, 7]. against institutional and motivational counterparts, informational control appears the “softest” (indirect) type. it lies in formation of control object’s awareness such that the decisions made on its basis are the most beneficial to control subject. this sort of mechanisms is used in organizations from ancient times, but in recent years they started a new life with development of social networks in internet [12]. 5 technology of control mechanism design: from models to policies the version of mechanism design we suggest for organizational management technically is close to the mechanism design, which forms the grounds of modern social choice theory and agency theory. the only difference is a systematically promoted pragmatic normative approach – in contrast to a standard descriptive approach of economic theory any conflict is a priori considered in the view of a single part, i.e., a client. not-withstanding formal resemblance of mathematical models, motivation of the research often varies, mostly limited to optimal mechanism design (i.e., construction of the best mechanisms in the view of a principal), whereas efficiency (in the sense of welfare economics) is overshadowed. under a given control mechanism, one may separate out three major stages of designing and implementing management decisions, notably, 1) data acquisition for decision-making; 2) decision-making process; 3) implementation of the decisions made. let us characterize each stage in greater detail. 1) data acquisition for making decisions. information required for making decisions may be acquired by the principal from the agents. the principal should be on the watch for untrue information (the agents may either “color up the truth” or “get dramatized”). no doubt, the principal would desire to have a certain control mechanism when economic agents (as rational “economic men” with individual preferences striving to maximize their utility function) benefit from being fair. the stated mechanisms are known as fair play mechanisms (strategyproof mechanisms). however, when the agents benefit from truth-telling? the only situation is when information reported by no way harms them. this is the underlying property of fair play mechanisms. 2) making decisions. at the second stage, the principal should make an effi10 v.n. burkov: mechanisms of organizational behavior control: a survey cient decision. if the principal is unable to succeed, he should improve himself (otherwise, he will be definitely replaced by another principal). thus, a key component here consists in management capabilities of the principal. 3) implementing the decisions made. finally, the third stage is intended for implementation of the decisions made by the principal. the principal is then exposed to the danger of non-fulfillment (or incomplete fulfillment) of the decisions. mechanisms where agents benefit (again, according to their individual preferences) from implementing the decisions of the principal are said to be incentive-compatible. an incentive-compatible strategy-proof mechanism is a correct mechanism. surely, an ideal situation is when a correct mechanism appears optimal (i.e., the most efficient). the mechanisms studied are adapted to managerial practice. for instance, despite their formal reducibility to a social choice problem (in a certain statement), resource allocation problems and auctions are traditionally considered separately, as the ones arising at different stages of organizational control cycle. now turn from the theory to the practice. who should solve the above-stated problems of mechanism analysis and synthesis? the first alternative is for the principal to solve both problems involving the templates and personal experience. the second alternative is to design a control mechanism with the aid of some sort of external consulting. almost any control problem in an organization can be stated formally in the following way. find feasible control actions ensuring maximum efficiency (such control is called optimal). to succeed, one should solve an optimization problem, notably, choose optimal control (optimal control actions. this is the most general setting. consider a general technology of solving the problems of this sort, which covers all the stages (from organizational system modeling to implementation of the model in concrete regulations); see fig. 3 where inverse links between the stages are omitted for better clarity. • the first stage (model construction) consists in description of an organizational system (first and foremost, control object) and in its modeling, i.e., specification of staff, structure and functions of the modeled system. • the second stage (model analysis) lies in studying the behavior of control object under different control actions. • the analysis stage being completed, one may apply control theory methods to solve the formulated control problem. first, solving the direct control problem, i.e., synthesis problem for optimal control actions (find a feasible control ensuring the maximal efficiency). second, solving the inverse control problem (find a set of feasible controls rendering the os to the desired state). it should be emphasized that generally this stage causes major theoretical difficulties and seems the most advances in systems science and applications (2013) vol.13 no.1 11 time-consuming for the researcher. • having obtained the set of solutions to the control problem, one should move to the fourth stage, notably, study their stability. stability analysis implies solving (at the very least) two problems. the first problem is to study the dependence of optimal solutions on parameters of the model; in other words, this is the analysis problem for solution stability. the second problem turns out specific for mathematical modeling. it consists in theoretical study of model adequacy with respect to the real system; such study means efficiency evaluation for those solutions derived as optimal within the model when they are applied to real os (due to modeling errors the solutions may vary from the model). • thus, the above-mentioned four stages constitute general theoretical study of os model. using the results of theoretical study to control real os requires adjustment of the model (i.e., identifying the modeled system and carrying out a series of simulation experiments; these are the fifth and sixth stages, respectively). in many cases the stage of simulation appears necessary due to several reasons. first, far from always one is able to obtain analytical solution to optimal control synthesis problem and to study its dependence on parameters of the model. note that simulation may represent a certain tool to derive and assess the solution. second, simulation allows for verifying the validity of hypotheses adopted to construct and analyze the model. in other words, simulation gives additional information on the adequacy of the model without conducting natural experiment. finally (and this is the third reason), employing management games and simulation models for training aims lets the managing staff master and test the suggested control mechanisms. • the closing is, in fact, the seventh stage, or the stage of implementation; it includes training of the managing staff, implementation of control mechanisms (designed and analyzed at the previous stages) in real os, with subsequent efficiency assessment of their application, correction of the model, and so on. 6 toolkit of control mechanisms management activity is traditionally divided into stages, which form the cycle of management. one of most popular divisions belongs to classic of regular management henry fayol, who detached five stages: planning, organizing, commanding, coordinating, and controlling. in the course of management theory development the role of proper employees’ motivation was also emphasized. every stage gives rise to a number of complex management problems. most of them involve principal-agents interactions of some sort, and, thus, require adequate support in the form of organizational behavior control mechanisms. the following list of mechanisms summarizes briefly the authors’ long experience of both academic studies and consulting projects. the mechanisms are 12 v.n. burkov: mechanisms of organizational behavior control: a survey separated into four stages of management cycle (fayol’s stages of commanding and coordinating are replaced to motivating, which is more important for employees’ behavior management) – see fig. 3 – and address the major management problems, which arise on the appropriate stage. this brief description is intended to give the reader general expression about the set of problems covered by mechanisms of employees’ behavior control and some flavor of mechanism design approach to address these problems (see the detailed description of mechanisms in [4]). the list presented below is not complete, but, of course, basic mechanisms do not cover all the variety of management problems. at the same time, these basic mechanisms can serve as bricks in designing numerous complex mechanisms of organizational management, which are based on combinations of a comparably small set of ideas. fig. 3 the complex of control mechanisms within a fayol’s management cycle 6.1 mechanisms of planning the resource allocation mechanism allows distributing scared resource (typically, budget funds, but also production plans, water, etc) among agents to maximize total efficiency of resource usage in the presence of lack of information about agents’ abilities to use resource efficiently (depending on situation this may be advances in systems science and applications (2013) vol.13 no.1 13 real demand for resource, local production capacities, etc). the properly designed mechanism ensures credibility of agents’ reports. we suggest using variations of a sequential resource allocation scheme [4, 11] to guarantee maximum efficiency while keeping truth-telling property. the mechanism of active expertise supports managers in many common situations when a decision is made on the basis of opinions of some experts. it is always a big problem to find a competent expert in the field, but very often no external experts can be involved at all, while independency and equity of internal experts is open to questions. the mechanism of active expertize minimizes the consequences of possible opinions distortion by motivating experts’ truth-telling. a class of median schemes [4, 13-14] is shown to solve the problem of truth-telling, in contrast to typically employed variations of an averaging scheme. transfer prices (also known as internal or corporate prices) help to distribute profit between production units in a corporation, but also provide a tool for high performance compensation and incentive compatible planning. units report their production plans, while the principal balances external demand with local production capabilities adjusting internal prices. the mechanism of transfer pricing establishes relation between the total manufacturing plan and the internal price per unit of manufactured good. when the number of agents is large enough the properly designed mechanism of transfer prices ensures, first, reporting real production plans (which units are interested to fulfill) and, second, efficient plan allocation between production units. an idea of a contest is used in a rank-order tournament mechanism to select most efficient project portfolios. the simplest tournament reduces to the following procedure. first, for every candidate project its total discounted cost is calculated and the effect of the project is estimated. second, projects are ordered by effect to cost ration in the descending order, and included in a portfolio until the budget runs out. this tournament procedure is also called the cost-benefit analysis. tournaments differ in the procedure of winners’ selection. for example, in a more complex mechanism a knapsack problem [15] is solved for the set of project cost-effect pairs. 6.2 mechanisms of organization the mechanism of joint financing helps to share federal funds with that of local authorities to implement complex regional programs. mechanisms of this sort are used in public-private partnership to finance socially oriented projects. the same idea works well in corporations, when, for instance, some local projects of performance improvement are jointly financed both by corporate center (by the principal) and at the local level (by the agents). project costs are reported by the agents, and the whole amount of a central fund is allocated proportionally to the reports, while the remainder of the reported cost must be covered by the 14 v.n. burkov: mechanisms of organizational behavior control: a survey agent. cost-saving mechanisms are intended to motivate an agent to improve efficiency of his or her activity as much as possible. the mechanism stimulates an agent to keep high quality of the output with minimum costs. cost-saving mechanisms are based on the following general concept. assume that agent’s profit depends on variables of two types, viz., parameters selected by an agent (e.g., costs) and parameters specified by the principal (e.g., a flexible norm of profitability, pricing coefficients, taxation coefficients, and so on). the mechanism design problem is to choose the values of principal-driven parameters for agent’s profit to grow with the costs decrease; this should be either accompanied by price reduction. in other words, net profit of an agent should rise when costs fall; in contrast, the product price should decrease. 6.3 incentive mechanisms an individual incentive scheme serves, firstly, for motivating agents to choose actions being beneficial to a principal and, secondly, for increasing the employees intensity of labor and motivating them to achieve better production results. the idea is that the employee’s wage includes a tariff (a fixed part) and a bonus (a variable part). the latter depends either on (measurable) intensity of his or her labor, or on the results of his or her activity (the intensity being immeasurable). adjusting the mechanism allows for attaining the required performance of the employees with minimum payments. when an agent is paid merely for fulfillment (or overfulfillment) of the plan assigned by a principal, the agent is not interested in having a high (“tight”) plan. the reason is performing it would require additional efforts (costs) of the agent. for instance, an agent may inform the principal of his or her preferences, eo ipso reporting his or her estimate of the plan (referred to as a “counter plan”). within the framework of incentive scheme for counter plans, an agent is given rewards for reporting counter plans that better meet the principals interests (yet, are “tighter” for the agent). a collective incentive scheme is designed for situations when a principal turns out unable to separately observe the action of every agent (the principal merely knows a certain aggregated rate, e.g., the result of collective activity). imagine that the principal can evaluate the minimum costs to-be-incurred by the agents to achieve the required result of collective activity. in this case, the efficient incentive scheme takes the following form. the minimum costs of each agent are compensated (provided that the result of collective activity agrees with the requirements of the principal). moreover, sometimes the principal has no costs related to observing individual actions of every agent; thus, the principal’s workload to acquire and process information is substantially reduced. the unified incentives mechanism is employed in situations when a principal advances in systems science and applications (2013) vol.13 no.1 15 has to motivate large groups of agents, to involve “democratic” management methods and to decrease the amount of processed data. under unified incentives, the relationship between the reward and labor intensity of the agents (alternatively, the results attained by them) is identical for all agents. in several cases, the described unification leads to no loss of efficiency, while the wages fund is spent in an optimal way. yet, unified control may be inefficient, when non-consideration of individual features of the agents results in inefficient spending of financial resources. a special case of unified incentive scheme is represented by competition incentives mechanism. the mechanism of informational control in incentive problems is involved when a principal possesses complete information (in contrast to agents). since the latter choose their actions based on individual awareness, the principal may impact on their actions by manipulating information; in other words, the principal may modify awareness of the agents, e.g., by means of informing every agent of the intensity of labor and/or plans of the rest agents. the team incentive mechanism serves to motivate groups (e.g., a production area, a department, a shop floor, a team, etc) with collective organization of work (taking into consideration individual contribution of every employee). incentive procedures for team members are based on bonus fund distribution according to their labor participation factor. 6.4 controlling mechanisms in integrated rating mechanisms one passes from a detailed description of a complex object (involving numerous indicators and parameters) to an aggregated description based on few generalized characteristics of the object. these mechanisms allow for regular monitoring and timely estimating (taking into account priorities of a principal) the results of the object’s activity and changes within the object. note that such changes take place during operation of the object or depending on the impact exerted by an external environment. the mechanism of consent is intended for coordinated distribution of financial resources among several possible directions of investment. for this, expert commissions are created; each commission generates a coordinated decision concerning the ratio of financial funds to-be-spent on a given direction to that of a fixed (basic) direction. we emphasize that the number of commissions should be by unity less than the number of directions. using information acquired from all expert commissions, a distribution of financial resources is defined among all directions; the idea is making the ratio of financial funds (to-be-spent on a specific direction) to that of the basic direction equal to the estimate of the corresponding expert commission. the mechanism ensures truth-telling of the commissions. the primary concept of a two-channel mechanism lies in simultaneous involving two channels of decision-making. motivating actions are based on comparative 16 v.n. burkov: mechanisms of organizational behavior control: a survey analysis of the efficiency levels of the decisions suggested by different channels. notably, rewards are evaluated for every channel. the first channel makes a decision. certain alternatives exist for the second channel; it could be either an active channel (decisions are worked out by human beings) or an advising (a computer-type, a “normative”) channel. in the latter case, suggestions are used to establish the norms of control efficiency (to compare the actual efficiency of the decisions made by the first channel). mechanism of predictive self-control is intended for well-timed informing a principal of possible deviations (from a plan) in the agents’ activity. the earlier the principal gets aware of possible deviations from the plan (e.g., in due dates, financial investments, etc), the more efficient and well-timed would be his or her decision (e.g., additional measures to eliminate deviations and reduce losses, or plan correction); note the agents report deviations. the matter is that the penalties of the agents (in the case of plan correction) depend on the moment the agents report of the correction (they are smaller if the report is early); moreover, these penalties are less than in the case of plan non-fulfillment. 7 an example of the mechanism: incentive-compatible planning due to limited space we leave the detailed description of the whole toolkit of organizational control mechanisms for future papers, and just give references to the papers with detailed description of the models and techniques used to solve various problems of control mechanism analysis and synthesis. below we give just an example of a single incentive mechanism used to boost planning accuracy due to accounting for employees’ incentives. as it has been mentioned, an incentive-compatible mechanism is remarkable for that the agent benefits from performing the plan. the following questions arise then: 1) does the principal need exact fulfillment of the plan by the agent? 2) if yes, how could it be ensured? the first question appears not as easy as it could seem. for instance, consider the utility function of the agent defined by the difference between his or her income 0.5y (the agent gains 50% income) and the penalty for plan nonfulfillment: 1 4(x−y)2; here y denotes production output of the agent expressed in money (the sales volume), while x means the plan. this function is shown in fig. 4. simple computations indicate that the agent benefits from choosing the action (x+1), i.e., he or she would strive for overfulfilling the plan by the unity. hence, if the principal requires y = 5 units of products (exactly this quantity, since the residual amount exceeds demand), he or she should assign the plan x = 4. we underline that the plan differs from the action which is desired by the advances in systems science and applications (2013) vol.13 no.1 17 principal! such situation is common in control of active agents. the principal has to predict their behavior and assign the plan based on the forecast. the described mechanism is not incentive-compatible (the agent does not perform the plan). nevertheless, sensus communis suggests that it would be good to have incentive-compatible mechanisms. fortunately, there exists a wide range of the incentive/penalty functions such that optimal plan is proven to be incentive-compatible. an example is provided by fig. 5. fig. 4 the utility function of the agent in an incentive-compatible planning mechanism fig. 5 an example of the penalty function such functions have been shown to possess optimal (in the view of the principal) incentive-compatible plan. thus, the ideology of optimal planning is being replaced with the ideology of optimal incentive-compatible planning. the plan should be optimal exactly over the set of incentive-compatible plans (beneficial to the agents). fig. 6 demonstrates an example of the set of incentive-compatible plans, where the agent is equally penalized for the plan non-fulfillment or overfulfillment. 18 v.n. burkov: mechanisms of organizational behavior control: a survey fig. 6 an example of the set of incentive-compatible plans one easily observes that any plan belonging to the set of incentive-compatible plans is beneficial to the agent. if the principal is interested in maximal production output, the optimal plan corresponds to the point a in fig. 6. for further increasing the output, the principal has to either enlarge the penalty or pay a reward for the plan fulfillment. in the latter case optimal incentive-compatible plan is defined by the point b in fig. 6. increasing the penalties or rewards corresponds to growing centralization of control, since the set of incentive-compatible plans is enlarged and the principal possesses better opportunities for assigning the plans. therefore, growing centralization of control may improve the efficiency of functioning in an organizational system. however, an infinite growth of centralization is clearly impossible; first, it requires infinite resources of control and, second, it may contradict the existing social norms, as well as the democratic and autonomous norms of control. 8 conclusion thus, the theory of behavior control mechanisms reviewed in this paper is positioned as a branch of control science (more specifically, cybernetics) which deals with control in the so-called active systems. the elements of active systems are people possessing individual interests, being able to choose actions independently and to manipulate information. in fact, the subject of the theory is systematic accounting for activity phenomenon in control problems based on systems approach and using the methods and results of operations research and game theory. the primary task is designing efficient mechanisms of organizational control. game-theoretic modeling appears one of the basic research methods used by the theory. as compared to modern economic theory, application of game theory in the theory of control mechanisms is different. pragmatic normative approach (a game conflict is a priori considered in the view of a single part, i.e., a client) is typical, in contrast to standard descriptive approach in economics (when a conflict is viewed indirectly). advances in systems science and applications (2013) vol.13 no.1 19 therefore, the basic idea consists in combining maximum usability in formulation of organizational control problems with wide adoption of formal models (including game-theoretic ones). still, to solve a complex applied problem of improving the efficiency of project management, we supplement the mechanisms of employees’ behavior control with optimal project planning or supply network optimization techniques. the subject of the theory coincides with that of management theory; nevertheless, the corresponding methods have dramatic differences. management theory seems much flexible in the description of psychological aspects, while mechanism design brings all psychological factors to the concept of rational behavior based on utility theory. many implications of management theory and organizational theory per se supplement formal analysis (performed by mechanism design theory) with empirical components that could not be embedded in modern formal models, yet are relevant during implementation of theoretical results. for instance, incentive mechanisms with monetary rewards are successfully supplemented with motivation theories, such as f. herzberg’s motivator-hygiene theory [16] or w.g. ouchi’s theory z [17] and others. integrated rating mechanisms provide a certain tool to design control systems for company efficiency based on different concepts, such as management by objectives by p. drucker [18] or balanced scorecards by norton d.p. and kaplan r.s. [19]. thus, we are sure that blending mechanism design with management theory may serve as a bridge of the formal theoretical results towards managerial practice. references [1] thiault, d. (2012), managing performance through business processes, createspace independent publishing platform. [2] hammer, m. and champy, j.a. (1993), reengineering the corporation: a manifesto for business revolution, new york, harper business books. [3] novikov, d.a. and rusjaeva, e. (2012), “foundations of control methodology”, advances in systems science and application, vol.12, no.3, pp. 33-52. [4] novikov, d.a. (2013), theory of control in organizations, new york, nova science publishers inc. [5] myerson, r.b. (1991), game theory: analysis of conflict, london, harvard university press. 20 v.n. burkov: mechanisms of organizational behavior control: a survey [6] novikov, d.a. and chkhartishvili, a.g. (2003), “information equilibrium: punctual structures of information distribution”, automation and remote control, vol.64, no.10, pp. 1609-1619. [7] bolton, p. and dewatripont, m. (2005), contract theory, mit press. [8] spence, m. (1973), “job market signaling”, the quarterly journal of economics, vol.87, no.3, pp. 355-374. [9] gubko, m.v. (2008), “mathematical models of formation of rational organizational hierarchies”, automation and remote control, vol.69, no.9, pp. 1552-1575. [10] forrest, j. and novikov, d. (2012), “modern trends in control theory: networks, hierarchies and interdisciplinarity”, advances in systems science and application, vol.12, no.3, pp. 1-13. [11] shoham, y. and leyton-brown, k. (2008), multiagent systems: algorithmic, game-theoretic, and logical foundations, cambridge university press. [12] barabanov, i. n., korgin, n.a., novikov, d.a., and chkhartishvili, a.g. (2010) “dynamic models of informational control in social networks”, automation and remote control, vol.71, no.11, pp. 2417-2426. [13] arrow, k., sen, a., suzumura, k. (2002), handbook of social choice and welfare, north holland. volume. 1. [14] arrow, k., sen, a., suzumura, k. (2011), handbook of social choice and welfare, north holland. volume. 2. [15] kellerer, h., pferschy, u., and pisinger, d. (2004), knapsack problems, springer. [16] herzberg, f. and mausner, b. (1993), the motivation to work, new brunswick, transaction publishers. [17] ouchi, w.g. (1981), theory z. how american business can meet the japanese challenge, massachusetts, addison-wesley. [18] drucker, p. (2006), the effective executive: the definitive guide to getting the right things done, collins business. [19] kaplan, r. s. and norton, d. p. (1996), balanced scorecard: translating strategy into action, harvard, harvard business school press. corresponding author d.a. novikov can be contacted at: novikov@ipu.ru. microsoft word 825 network redundant dcs configuration management adv syst sci appl 2020; 01; 83-90 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/825 network redundant dcs configuration management n.a. zakharov1, v.i. klepikov1, d.s. podkhvatilin1 1) “dozor” subdivision jsc “kemz” moscow, russia e-mail: nazakharov@npp-dozor.ru abstract: an approach to the configuration management of a distributed control system with redundant components is proposed. the approach is based on a formalized vector-matrix representation of sets of system components, possible system configurations and functional relationships between them. keywords: distributed control system, configuration, configuration management, configuration supervisor, redundancy, fault detection 1. introduction the evolution of distributed control systems (dcs) at the present stage leads to the creation of models of complex automation of processes and production with automatic reconfiguration of their structure without the participation of personnel in order to continue normal operation in the event of failure in their components. dcs with digital communication channels can be represented as a set of jointly and purposefully functioning distributed dynamic objects and, in accordance with the modern theory of systems, is classified as a complex dynamic system. currently, a promising concept for the construction of dcs based is on maintenance-free modular electronics [1, 2]. such a dcs consists of redundant equipment and redundant communication network. this concept implies that the dcs possesses comprehensive system of collection and generalization information about equipment functioning. it uses highperformance algorithms for detection and isolation observable and unobservable failures. the key feature of the dcs under consideration is its capability of deep system reconfiguration based on component redundancy of functional systems and communication network. 2. redundancy the most promising way to parry failures in dcs is to use the redundancy of system resources. in general, there are the following types of redundancy:  system redundancy of individual components and subsystems;  structural (architectural) – adaptive reconfiguration of the system structure;  information addition to the main signal information on which it is possible to check its reliability (the checksum in the communication channel, estimates on model in dynamic control systems);  communication use of duplicated and different types of communication channels;  algorithmic using different algorithms to complete the same tasks;  software – use of different software tools;  technological use of a variety of software, information and other technologies; 84 n.a. zakharov, v.i. klepikov, d.s. podkhvatilin copyright ©2020 assa adv. in systems science and appl. (2020)  time re-execution of operations with subsequent processing of results;  semantic use of redundant semantic structures, digital images of controlled parameters;  organizational use of additional systems or subsystems that duplicate functions or information flows of the main system. 3. schedule-triggered protocol a promising solution to build a communication network is the use of the scheduletriggered protocol (stp) [3, 4]. the network distributed architecture of the control system based on the stp provides a high level of hardware, software, time, communication and information redundancy. this redundancy can be used not only to improve the reliability and fault tolerance of the control system, but also to improve the accuracy and quality of regulation and control. the large bandwidth of the stp channel allows to organize a single information space for all nodes of the distributed system. it allows several nodes to simultaneously perform calculations of control algorithms and transmit the results to the schemes of majority voting or data aggregation. 4. configuration supervisor method to control dcs redundancy a modified method for configuration supervisor (cs) is offered. in the initial method [5], cs refers to software and hardware modules used to monitor the operability of its configuration, participate in inter-service arbitration to activate its configuration either in case of winning the arbitration, or for parallel work together with other configurations on a common actuator. in the proposed modified cs method, the determination of the integrity of the configuration based on the information about the operability of its components is supplemented with estimates of the integrity of the configurations, formed on the basis of the analysis of the configurations output data. the obtained estimates of the integrity of the configurations are used as feedback to obtain estimates of the operability of the components forming the configuration. let’s consider dcs based on stp. the dcs has an excessive number of hardware, software, communication and software components (resources). resources in the general are:  sensors and input devices;  actuators and signal output devices;  computing nodes (controllers);  communication lines between the nodes and interconnecting systems;  built-in mathematical models. in this system, n variants of the sets of the available m components can be organized to perform the specified system functions. each such variant of the set will be called a configuration. different configurations can implement either different or the same system functions, configurations can have completely disjoint component sets, or they can share resources. the designed approach to management of redundancy of the distributed system is providing:  formalized representation in vector-matrix form of all components, used configurations and sets of components involved in each configuration;  formalized method for working configurations vector y direct calculation based on the original vector x component of operability;  calculation of the vector ŷ configurations integrities estimates based on configurations operation results;  formalized the procedure of reverse computations of the vector of estimated components operability based on vector serviceable configurations estimates ŷ; network redundant dcs configuration management 85 copyright ©2020 assa. adv. in systems science and appl. (2020)  replacement of existing faulty configurations by new or replacement in existing configuration of the failed components by the same serviceable. 5. gas turbine unit distributed control system let’s consider operation of the proposed method on the example of configuration management of a gas turbine unit drive distributed control system (gtu dcs). the block diagram of the gtu dcs is shown in fig. 1. the system contains k=46 components ci, i=1...46, which have the following functionality:  c1...c5 – u1...u5 – gtu dcs controllers;  c6, c12 – l1, l2 – redundant stp-bus;  c7...c11 – l11...l15 – controllers u1...u5 stp-bus l1 ports;  c13...c17 – l21...l25 – controllers u1...u5 stp-bus l2 ports;  c18 – l3 –controller u5 with the upper level control system communication channel;  c19, c32 – n11, n12 – main and backup speed sensors of the low-pressure compressor (lpc);  c20, c33 – n21, n22 – main and backup speed sensors of the high-pressure compressor (hpc);  c21, c34 – n31, n32 – main and backup speed sensors of the free power turbine (fpt);  c22, c35 – t11, t12 – main and backup air temperature sensors;  c23, c36 – p11, p12 – main and backup air pressure gauges;  c24, c37 – t41, t42 – main and backup temperature sensors of the gas before fpt;  c25, c38 – p21, p22 – main and backup combustion chamber pressure sensors;  c26, c39 – q11, q12 – fuel control valve position main and backup sensors;  c27, c40 – q21, q22 – lpc guide vanes position main and backup sensors;  c28, c41 – q31, q32 – hpc guide vanes main and backup position sensors;  c29, c42 – z11, z12 – main and backup control signals to the fuel control valve;  c30, c43 – z21, z22 – main and backup control signals to the lpc guide vanes;  c31, c44 – z31, z32 – main and backup control signals to the hpc guide vanes;  c45 – t13, p13, ns3 – information from the upper level control system: air temperature and pressure, set point to the free power turbine speed controller;  c46 – n14, n24, n34, t44, p24 – built-in engine model that produces real-time estimates of speed, gas temperature and pressure in the combustion chamber. the gtu dcs must provide:  depending on the current free turbine speed setpoint (ns) and the actual values of inlet temperature (t1) and air pressure (p1), by regulating the fuel consumption (q1) by the control signal (z1), ensure that the required value of the free turbine speed (n3) is maintained, while preventing the output of the combustion chamber pressure (p2) and gas temperature (t4) parameters for permissible values, which are also functions of temperature (t1) and air pressure (p1).  depending on the actual values of the engine rotor speeds (n1, n2) and the pressure in the combustion chamber (p2) adjust by means of control signals (z2, z3), the guide vanes position (q2, q3), which provides the required margin of gas-dynamic stability of the engine. the control system operation can be described by three functions: f1 – fuel consumption control, f2 – lpc guide vanes position control and f3 – hpc guide vanes position control: 86 n.a. zakharov, v.i. klepikov, d.s. podkhvatilin copyright ©2020 assa adv. in systems science and appl. (2020) z1 = f1(ns, n1, n2, n3, t1, p1, p2, t4, q1), z2 = f2(n1, p2, q2), (1) z3 = f3(n2, p2, q3). analysis of the gtu dcs block diagram (fig. 1) indicates that the values of the arguments of the functions f1, f2 and f3 can be obtained from various sources (redundant sensors, model, information links with the upper level system). calculation of functions values can be performed on one or several processors of various controllers, in-system information exchange can be carried out via one of two or simultaneously via both buses of the stp duplicated channel. issue of control signals both to the fuel control valve and to the guide vanes drives can be made via any of two or simultaneously on both control channels. during gtu dcs operation failures of individual components may occur, there may be failures in the processes of measurement, calculation, data transmission. due to the hardware, computing, information and time redundancy of the dcs, the gtu can continue operation using remaining serviceable components. fig. 1. dcs block diagram 6. configurations let's call configuration a set of components of the system c = {c1, c2, ..., cm}, providing execution of a certain function. the configuration can provide both the execution of the object control function, i.e. end with the calculation of the output value of the system, and perform intermediate calculations that determine the values of the parameters necessary for other configurations, for example, to calculate the most reliable values of the input parameters of the dcs based on the readings of several sensors. the same function in the control system can be implemented in different configurations depending on which components are currently in good condition. in complex objects distributed control systems can consist of hundreds of components, which, depending on their current state can be combined into dozens of different configurations to perform certain functional tasks. the formation of such configurations must be performed either in advance at the design stage of the system, or generated automatically in real time, depending on the current situation. without the use of formal design methods, both approaches are very time-consuming, require a lot of "manual" work at the stages of design and testing of the system. for this reason there is a need to develop analytical methods for representing the sets of configurations of available resources and managing these configurations in real time. network redundant dcs configuration management 87 copyright ©2020 assa. adv. in systems science and appl. (2020) the value of the fi function formed in some j-th configuration of cij will be denoted as zij. in the given system, for the functions f1, f2 and f3 can be formed, for example, three different variants of the calculation of the output values: z1 1 = f1(ns3, n11, n21, n31, t11, p11, p21, t41, q11), z1 2 = f1(ns3, n12, n22, n32, t12, p12, p22, t42, q12), z1 3 = f1(ns3, n14, n21, n32, t13, p13, p24, t44, q11), z2 4 = f2(n11, p21, q21), z2 5 = f2(n12, p22, q22), (2) z2 6 = f2(n14, p24, q21), z3 7 = f3(n21, p21, q31), z3 8 = f3(n24, p22, q32), z3 9 = f3(n22, p24, q31). in terms of the functional designation of the component, these variants can be implemented in the following configurations: c1 1 = {ns3, n11, n21, n31, t11, p11, p21, t41, q11, u1, u2, l1, l11, l12, l3, z11}, c1 2 = {ns3, n12, n22, n32, t12, p12, p22, t42, q12, u3, u4, l2, l23, l24, l3, z12}, c1 3 = {ns3, n14, n22, n32, t13, p13, p24, t44, q11, u1, u2, u5, l1, l15, l12, l2, l25, l22, l3, z11}, c2 4 = {n11, p21, q21, u1, u2, l1, l11, l12, z21}, c2 5 = {n12, p22, q22, u3, u4, l13, l14, l2, l23, l24, z22}, (3) c2 6 = {n14, p24, q21, u2, u5, l1, l15, l12, l25, l22, z21}, c3 7 = {n21, p21, q31, u1, u2, l1, l11, l12, z31}, c3 8 = {n24, p22, q32, u3, u4, u5, l2, l23, l24, z32}, c3 9 = {n22, p24, q31, u2, u3, u5, l1, l12, l13, l15, l2, l25, l22, z31}. the same configurations in terms of component numbers can be written as: c1 1 = {c45, c19, c20, c21, c22, c23, c25, c24, c26, c1, c2, c6, c7, c8, c18, c31}, c1 2 = {c45, c32, c33, c34, c35, c36, c38, c37, c39, c3, c4, c12, c15, c16, c18, c42}, c1 3 = {c45, c46, c33, c34, c26, c1, c2, c5, c7, c11, c8, c12, c17, c14, c18, c31}, c2 4 = {c19, c25, c27, c1, c2, c6, c7, c8, c30}, c2 5 = {c32, c38, c40, c3, c4, c9, c10, c12, c15, c16, c43}, (4) c2 6 = {c46, c27, c2, c5, c6, c11, c8, c17, c14, c30}, c3 7 = {c20, c25, c28, c1, c2, c6, c7, c8, c31}, c3 8 = {c46, c38, c41, c3, c4, c5, c12, c15, c16, c44}, c3 9 = {c33, c46, c28, c2, c3, c5, c6, c8, c9, c11, c12, c17, c14, c31}. components in formulas (3, 4) can be written in arbitrary order because these formulas are intermediate. they are used to clarify formal representation of configurations. expressions (4) for sets of ck components present in ci j configurations can be written as a matrix of k dimension (n×m), where n is the number of all configurations considered in the system, m is the number of all components of the system involved in the formation of configurations. the kij element of the matrix k is 1 if the i-th configuration uses the cj dcs component (see fig. 1). 7. state and availability vectors we define a vector x that characterizes the current state of all components of the dcs. if all components of the dcs are operable, the components of the vector x are represented as: 88 n.a. zakharov, v.i. klepikov, d.s. podkhvatilin copyright ©2020 assa adv. in systems science and appl. (2020) xs = 1, s = 1…m. (5) note that all components involved in the configuration are equally involved in the implementation of the system function, regardless of their physical nature and indicators of their own availability. for the example under consideration, m = 46. if there are faulty components in the system, the corresponding components of the vector x are zero. the value of this vector is formed in real time according to the results of the functioning of dcs software and hardware self-diagnostics. next, we define the availability vector of configurations y, which characterizes the readiness for operation of all considered configurations. yi = 1 if the i-th configuration is healthy, otherwise yi = 0, i = 1...n. configurations availability is a function of the components involved in configurations and their operability. y = f1(k, x), (6) where f1 is defined by the expression:   m s ssj xkf 1 ,1   , (7) the symbol  denotes the logical "and" function, the symbol  denotes the implication function. the meaning of the expression (7) is that each j-th element of the vector y will be equal to 1, i.e. the j-th configuration will work if only serviceable components are used in this configuration. in the stp-based dcs, several or all configurations can be executed in parallel or sequentially, which allows you to parry the failures of its components. the number of executable configurations is limited by the computing power of the dcs. if at execution of several configurations different values of the same output parameter are received, it is required to solve two problems: 1) reject incorrect results and mark the configurations which have formed them as faulty; 2) on the basis of the results obtained from the recognized working configurations to form the final agreed output value. 8. dcs components diagnostics an effective way to diagnose the dcs components operability is the inclusion in its composition of the built-in real time model of the object under control [6]. this allows you to get a number of additional features to improve the quality and reliability of control, improve the performance of the system: noise and fault filtering; the recovery of unmeasured parameters for the diagnosis and management; detection of abnormal states of the object and the control system; diagnostics of the state and parametric degradation of the object. some methods of rejection of incorrect values are considered in [7, 8]. next, consider the analysis of configurations based on tolerance control. let's write down the vector of output values obtained from the results of all n configurations operation. z = [z1, …zj, …zn] t (8) (upper indexes are omitted for clarity) and the initial configuration availability vector y = [y1, …yj, …yn] t. (9) network redundant dcs configuration management 89 copyright ©2020 assa. adv. in systems science and appl. (2020) using the vectors zmin and zmax, we set the minimum and maximum output values for each configuration, respectively. let’s define a tolerance control function as f2(y, z, zmin, zmax) = 1, if (yj=1) & (zmin j ≤ zj ≤ zmax j); (10) else 0. the evaluation of the configurations availability vector ŷ is defined by the function f2: ŷ = f2(y, z, zmin, zmax), (11) the zero value of any component of the configuration availability vector indicates that there is one or more failed components of the dcs in the corresponding configuration. next, we define the function f3 as   n j isj ykf 1 ,3 ˆ   , (12) the symbol  denotes a logical or function; the symbol  denotes an implication function. for the vector that characterizes the serviceability evaluation of the dcs components, we can write:  ykfx ˆ,ˆ 3 , (13) the meaning of the expression (13) is that for each s-th element of the state vector, the operability of all j = 1...n configurations, in which it is involved, is checked. if at least one configuration in which this component is involved, i.e. the condition cj,s ≤ ŷj is satisfied for at least one j = 1...n, then the s-th component is identified as serviceable. if all configurations in which this component is involved are found to be faulty, the component is identified as failed. 9. configuration management algorithm configuration management algorithm is the following. the algorithm is executed cyclically. the first step of the algorithm is the formation of the initial vector x of the dcs components serviceability. the vector ν of dimension m formed by the built-in dcs components self-diagnostics means and the vector x̂ of components serviceability estimates calculated on the previous cycle are used. at the initial cycle, all components of the vector x̂ are assumed to be equal 1. in the second step, based on the original serviceability vector x and the configuration matrix k the vector of configurations availability y is formed. next, according to the configurations availability verification method (in this example – tolerance control) configurations availability evaluations vector ŷ is formed. in the fourth step the dcs components serviceability evaluations vector x̂ is calculated. in case of faulty configurations detection, a new configurations matrix k is formed by replacing the used faulty configurations with new ones or replacing the faulty components with serviceable ones in the existing dcs configurations. 90 n.a. zakharov, v.i. klepikov, d.s. podkhvatilin copyright ©2020 assa adv. in systems science and appl. (2020) 10. conclusion in conclusion it should be noted that both of configurations set and their component composition can be optimized according to various criteria, such as the dcs computing and communication resources load, results interpretation unambiguity. in particular, it is obvious that if each of the system components will be involved in only one specific configuration and absent in all the others, its fault is clearly detected when recognizing this configuration failed. if two or more components are present in only one configuration (or a group of configurations) and are not present in all the others, it is not possible to distinguish their failures due to the failure of this configuration. optimization of the number and component composition of dcs configurations should be performed taking into account the depth and reliability of the built-in software and hardware self-diagnostics of dcs components. references 1. jin, h. lee s., han s., jo, h. kim d. (2012). wip abstract: challenges and strategies for exploiting integrated modular avionics on unmanned aerial vehicles. ieee/acm third international conference on cyber-physical systems, beijing, china, 211-211. https://doi.org/10.1109/iccps.2012.34. 2. wang, l. sun y., guo p., zhang y. (2015). an enhanced reconfiguration method for the second generation integrated modular avionics, 11th international conference on computational intelligence and security (cis), shenzhen, china, 433-436. https://doi.org/10.1109/cis.2015.110. 3. kopetz, h. (2011) real-time systems. design principles for distributed embedded applications. heidelberg, germany: springer. https://doi.org/10.1007/978-1-44198237-7. 4. zakharov, n.a., klepikov, v.i., podkhvatilin, d.s. (2013) sinkhronno-vremennoj protokol dlja raspredelennykh sistem upravlenia [schedule triggered protocol for distributed control systems], avtomatizatsiya v promyshlennosti, 2, 37-39. [in russian]. 5. ageev, a.m., bronnikov, a.m., bukov, v.n., gamayunov, i.f. (2017). supervisory control method for redundant technical systems, journal of computer and systems sciences international, 56 (3), 410-419. https://doi.org/10.1134/s1064230717030029. 6. klepikov, v.i., kalin, s.v., zakharov, n.a., podkhvatilin, d.s. (2008). algoritmicheskoe obespechenie otkazoustojchivosti raspredelennykh system upravlenia [algorithmic provision of distributed control systems fault tolerance], radioelektronni i komp'uterni sistemi, 34 (7), 43–48. [in russian]. 7. klepikov, v.i., podkhvatilin, d.s., dudorov, y.n. sharapov, g.v., zakharov, n.a. (2011). information-measuring diagnostics complex for technical maintenance. autom remote control, 72 (5), 1089-1094. https://doi.org/10.1134/s0005117911050171. 8. klepikov, v.i. (2014) otkazoustoychivost' raspredelennykh sistem upravleniya [fault tolerance of distributed control systems]. moscow, russia: zolotoe sechenie [in russian]. advances in systems science and application(2016) vol.16 no.3 52-75 load-aware congestion adaptive multipath multicasting h. santhi and n. jaisankar school of computing science and engineering (scope), vit university, vellore, india abstract in recent years, all communication system becomes wireless due to vast development and advancements in wireless technology. such network offers a platform to deploy a wide range of multimedia and other services with different data types. in mobile ad hoc networks one of the most prominent factors which affect the overall performance of the network is congestion. congestion is one of the most misinterpreted concepts in the context of the wireless network. in general, excessive data traffic in the network referred to as congestion. however, in the wireless network, the congestion occurs due to unavailability of resources such as insufficient bandwidth, low battery power, and so on which leads to high packet losses, bandwidth reduction, and wastage of energy and time in recovering congestion. in addition to this, another factor which degrades the performance is improper load balancing due to certain routing metrics limitation. from the literature, it is observed that many of the existing solutions handle the above problems separately in a not-adaptive manner. however, addressing these issues together in an adaptive manner provides an effective solution to balance the load in the network as well as congestion. the proposed work addresses congestion and load balancing problems parallel with the aim to improve and enhance the overall network performance and lifetime. the proposed load aware congestion adaptive multipath multicast (lacamm) routing approach adapt to current changes in the load and congestion level to find a suitable path even in the case of congestion scenario and node resource constraints. the proposed scheme measures the node resources such as residual bandwidth and the residual battery to predict the node stability. redirects the data transmission under congested scenario from the congested node through the non-congested alternate path thereby improves the overall network performance. this work performs well under burst traffic with the harsh environment. keywords multicasting, congestion, multipath, adaptive routing 1 introduction a mobile ad hoc wireless network consists of a set of wireless mobile nodes shaped dynamically with none central administration or existing network infrastructure[1]. every node in this network will act as a host and additionally as a router and contains a capability to move freely and at random in a direction at any speed. owing to limitation of mobile nodes transmission varies the packets advances in systems science and application(2016) vol.16 no.3 53 are forwarded to the destination during a multi-hop fashion with the assistance of intermediate nodes. one among the key problems in multihop mobile ad hoc network is congestion. the crucial factors that influence congestion throughout multi-hop relay are shared restricted wireless information measure, low device power, dynamically ever-changing configuration, and so on[2,3]. congestion will cause packet loss, degradation of information measure, waste of resources on congestion recovery. it’s troublesome to beat congestion problem however it’s potential that the congestion is often avoided by adapting bound appropriate mechanism and rules for the flow. normally routing algorithms in manets are generally classified into proactive and reactive routing. in proactive routing, the routes are established as like wired network approach and are updated either sporadically or on a progressive update fashion. this approach isn’t appropriate once the network is just too massive and, therefore, the nodes are extremely mobile. in reactive routing, the routes are created as and once required and is a lot of economical than the proactive routing approach[4-7]. however, the matter is that the reactive routing protocol creates high management overhead just in case of frequent path breaks. this drawback is self-addressed by use of multipath routing approach. the routing protocols in manets are often classified in our own way as congestion-aware routing and congestion adaptive routing. several existing solutions belong to congestion aware approach solely only a few are congestion adaptive. in congestion aware approach the congestion is taken into thought solely throughout route discovery and maintains a similar standing till the trail breaks. however in congestion adaptive routing the routes are adaptive to the present congestion standing of the network. congestion non-adaptiveness can cause the following[8-9]: long delay: congestion-aware routing takes long-standing to observe congestion. upon congestion, it’s quite essential to use a brand new route. but, the matter with congestion aware on demand routing protocol is that it takes longstanding to search out a much better non-engorged route this ends up in high delay. high overhead: upon congestion invoking re-route discovery involves with flooding of control packets to search out a brand new route to forward the information. the flooding of control packets creates high control overhead that degrades the performance of the network. many packet losses: as mentioned congestion could happen to any node at any time as a result of lack of resources that ends up in several packet losses. a typical congestion control try and scale back the traffic load, either by decreasing the rate at the sender or dropping packets at the intermediate nodes or doing each. these problems become additional visible throughout transmission of enor54 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting mous multimedia system applications during a large mobile accidental network and supply a negative impact on the network performance still as within the quality of service. 1.1 congestion congestion may be a drawback that happens on shared networks, once multiple users access to constant resources (bandwidth, buffers, and queues). once numbers of packets are present in a network is larger than the capacity of the network then this case is termed as congestion[10]. congestion in a network could occur once the load on the network i.e. the amount of packets sent to the network is bigger than the capacity of network[6,11]. 1.1.1 congestion control congestion control mechanism is performed once the network faces congestion. congestion control mechanism sometimes enhances network overall performance based on the load condition of the network. the congestion control mechanism is completed through controlling the sending rate of information streams of every source and conjointly results in high utilization of the offered bandwidth. the most objective of congestion control is to attenuate the delay and buffer overflow caused by network congestion and, therefore, alter the network to perform higher. as congestion is directly associated with the problem of dropping the packet, it’s needed that some technique is applied on the network so the drop of the packet can decrease. however to regulate on the quantity of dropping rate is tougher in manets as compared to the wired network due to the following characteristics [7,12-13]. a. dynamic topology as in manet, there’s no central point or base station, to regulate the entire network. each device will move freely in manet, therefore, the topology of the network isn’t mounted. thus, it can’t be expected whether or not a node that participates throughout some transmission can collaborate in the whole transmission or not[14]. a node will move any time instance thus a path detected by the source node to transfer its information will be a break at any time. if no path is found by the intermediate node to forward the information it’ll begin to drop the packet once a while. b. multi-hop routing each node in manet will receive and forward the data towards the destination nodes. however node forwarding capacity is restricted to its transmission range; it suggests that it will deliver the data packets to solely that node that come beneath its transmission range. if any 2 nodes that not come back beneath the transmission range of each other than the forwarding node depends on intermediadvances in systems science and application(2016) vol.16 no.3 55 ate nodes to relay data in a multi-hop fashion[10, 12]. a route has been detected by a routing protocol then sender begins to transfer the data to a node that comes beneath its transmission range this node referred to as an intermediate node, every intermediate node further transmitted data to its neighbor node and this process is repeated till information reach to the destination. arrival rate of packets at this specific node are often larger than its forwarding capacity so this node begins to drop the packet. c. heterogeneous environment in manet, any device will participate if it’s able to forward the data. these participating devices are totally different of various kind having a different storage capability and different resource. the transmission rate of every device could stay completely different. in manet addition of recent device is extremely simple if it comes beneath the transmission range of different node it becomes the part of that network. thus, it should be possible that a brand new device comes back and begin to transmit its own data on the route that is already detected by a different node. all devices taking part in communication are of various kinds and will become unavailable at any time that makes the period of communication not so long[5,15]. in such kind of condition, packets are dropped by the precursor node. sometimes a particular node becomes the intermediate node between several nodes. a scenario will arise at this node that several of its neighbor nodes forward the data to that the same time, thus there’ll be an excessive quantity of packets inward at these intermediate node. if the arrival rate of data on the nodes is bigger from its transmission rate node can begin to drop the packets. d. density of node the number of neighbor nodes of every node might also the reason in manet, as a result of if a node cannot deliver the data on to the receiver node then use another intermediate node to forward the data packet. in manet for every node, a lot of neighbor nodes mean a lot of link connections between the nodes and their neighbors. itll become the rationale of a lot of arrival rate of packets at a specific node, therefore, a lot of neighbor node of any intermediate node become the rationale of coming significant load as compared to the node reception capacity. in such kind of condition, a node can begin to drop the packet[13,16]. e. presence of malicious node reliability of the node is going to decrease as a result of the presence of malicious packet dropper node. in manet participating devices have restricted resource sometimes routing protocol select the path within which packet dropper node work as an intermediate node. a packet dropped node is self-seeking nodes that really not forward the data packets to next node however in place of this it 56 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting simply drops the packet to save lots of the resources[3,17]. the presence of packet dropper node could be a severe downside in manet and that they don’t seem to be the sole reason for the massive delay, however additionally become the reason of heavy traffic load on the network because the sender might become involved in causing packets again and again if no acknowledgment is received from the receiver. f. absence of physical protection in manet it’s impossible to guard a node type numerous kinds of threats because the node position isn’t fixed, a node will move in any direction within the network[18]. the nodes will be attacked from any direction wherever fixed physical protection like firewall and gateways can’t be applied. it means that for securing itself a node should be equipped to fulfill an offender directly or indirectly. however because of the absence of physical protection like in hard wired network, there’s a lot of likelihood for a node to become unreliable, and begin to drop the packet. 1.1.2 congestion prevention it is the mechanism to handle the network from congestion that involves play before network faces congestion. for this purpose nodes got to monitor their status and that they negotiate with the neighbor node within the network so no a lot of traffic than the required amount, the node will handle, are allowed to return to the network so no congestion can occur. congestion affects the performance of the network. therefore, some necessary congestion control technique is needed to stop the network from the congestion. prevention from congestion in manets is far difficult as compared to wired networks because of its specific characteristics. the subsequent are a number of the most qos provisioning and maintenance issues in manets[19]. a. stable route to prevent the network from the congestion it’s better to decide on a reliable path. for this purpose route are going to be analyzed so a perfect error free totally coverage path with high transmission delivery ratio is select. it needs data of the nodes which can be remain offered all the time, however because of the dynamic environment of manet choice of such node isn’t possible. b. reservation of bandwidth bandwidth reservation is a technique to stop the network from congestion, during which nodes reserve bandwidth for future communication through negotiation between the neighbors nodes which come back among 2 to 3 hops. it needs communication, and exchanges of a message between them because the channel is shared between the nodes. in manet environment, a node will moves from the reservation space of the node at any time even communication goes on. thus, advances in systems science and application(2016) vol.16 no.3 57 reservation of bandwidth means that additional overhead for communication and releasing messages. therefore, bandwidth reservation isn’t attainable in manet. c. service level agreement (sla) in manet, every participating node works as a host and as a router. any node isn’t responsible for performing some specific task. since all the nodes within the network work to produce services, there’s no clear definition of a service level agreement (sla). whereas in, an infrastructure network the services to the users within the network are provisioned by one or additional service providers. thus, estimation of the node behavior isn’t possible that is needed for prevention from the congestion. d. channel reliability since the wireless bandwidth and capacity in manets are suffering from interference, noise and multi-path attenuation, the channel isn’t reliable. moreover, the offered bandwidth at a node can’t be estimated precisely as a result of it involves large variations based on the quality of the node and other wireless device transmission within the neighborhood etc. e. routing difficulty routing is troublesome in manet because link breakage occurs overtimes. once any link of a path breaks, it got to find the other offered link or replaced with a new found path. this rerouting operation costs the scarce radio resource and battery power whereas rerouting additionally increase delay that additionally have an effect on quality of service of applications and degrade the network performance. thus, the routing operation has got to take care of such variety of challenge that is tough to handle. 2 previous work there are many congestion algorithms are proposed for mobile ad hoc networks some of them are explained below. in [19] developed a method for detecting congestion well in advance in order to prevent the network from the congestion. their work is based upon the calculation of approximate queue length in advance. for this purpose, they calculate the average queue length at the node level. network characteristics like congestion and route failure need to be monitored and resolved with a reliable mechanism. to solve the congestion problem, a novel dynamic congestion estimation technique has proposed that could analyze the traffic fluctuation. by the assessment of average queue length, a node is able to find that there is some probability of congestion so it sends a warning message to its neighbors. upon receiving the warning message they try to search some alternative congestion free path to the destination and resumes communication through an alternate path. so this dy58 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting namic congestion estimation procedure tries to provide a reliable communication within the manets by controlling upon the congestion in ad hoc networks. in [20] a new technique to detect the packet dropper nodes in the network by using a reliability factor. in manet each node has limited resources like limited battery power, a packet dropper node is that node in the network which may not cooperate properly in network operations as they not forward the coming data packets to the next node but instead of this they drop the data packet to save their resources. such nodes are called selfish or misbehaving nodes and these nodes are also the reason of congestion. the dropping of data packet not only affects the network connectivity but also can widely waste the network resources. to handle this situation a scheme based on mac-layer acknowledgments is used to detect the packet dropper nodes. to eliminate such nodes from the network its reliability is evaluated during the packet transformation. in this work the field of reliability factor is increased on the basis of acknowledgment received from the receiver, and all senders making the decision to send a packet to a node having higher reliability factor. the reliability factor identifies the packet dropper nodes based on the acknowledgment. hence, on the basis of node reliability factor, a packet dropper node can be detected and also can be isolated from the network. in [21], proposed congestion-aware routing (carm) to adapt to the congestion. the high throughput non-congested routes to any node in the network are selected based on the weighted channel delay (wcd) value. second, the proposed work adapts mismatched data-rate routes using effective link data-rate categories (eldc). in general, the protocol tackles congestion by switching between the above-said approaches to compact congestion in the network and efficiently increases the overall network performance. a method for reliability analysis for manet is presented by sreedhar and damodaram[22]. they proposed that the node performance is also influenced by the number of neighbor nodes of that node. in their work effect of node mobility and reliability in a real manet platform is proposed and analyzed. they proved that the wireless network has limited capacity, and the throughput of the wireless network granted to each user can be decreased to zero if the number of users increased. as the transmission capacity of the wireless network affect the throughput and it will affect the terminal reliability of manet. congestion means the arrival of an excessive amount of packets at a network which leads to many packet drops. a node can communicate with many nodes which are its neighbor nodes. as they come under the communication range of that node then there will always the chance that at the same time many neighbor nodes send their data packets to the same node, so there will be an excessive amount of packets arriving at these nodes become the reason of packets drop. hence, congestion is related to the density of the node in some area, and it will influence advances in systems science and application(2016) vol.16 no.3 59 the terminal reliability by reducing the intermediate node reliability. this work focuses on upon identifying the relationship between the number of link connections and the node reliability to reduce the congestion problem. type of service aware routing protocol (tsa) proposed in [23] is an improvement to aodv. this approach uses only a hop count as a metric for route selection. tsa is a cross-layer congestion-avoidance routing protocol in which the routes are used for extended periods of delay sensitive traffic. avoiding busy nodes alleviates congestion, leads to fewer packets drop and in a short end-to-end delay. in addition, tsa distributes the load on a large area, so by increasing the spatial reuse. a simulation study reveals that tsa significantly improves the throughput and reduce packet delay during high congestion state. to handle the network dynamics an optimized reliable ad-hoc on-demand distance vector (oraodv) scheme proposed in [24]. the proposed protocol (oraodv) is meant for best route discovery and reliability of packet delivery. a new idea of blocking expanding ring search (blocking-ers) is employed in it to avoid network wide broadcasting. the blocking-ers doesn’t begin its route search procedure from the source node whenever a broadcast is needed. the broadcast is initialized by any acceptable intermediate nodes on behalf of the source node that acts as a relay or an agent node. 3 proposed lacamm approach this work is adaptive to the current load and the tries to prevent congestion well in advance by warning its upstream and downstream forwarding group nodes. each node on the primary path generates a warn message when it is prone to be congested. upon receiving warn message the upstream node uses an alternate non-congested path along the primary path for avoiding the potential congestion area. traffic is distributed eventually over the available routes, thus, efficiently decrease the chance of congestion. lacamr is on-demand multipath multicast routing protocol which comprises the following components: 1.resource monitoring 2.congestion monitoring 3.construction of resource full node list 4.congestion free route primary route discovery 5.congestion adaptively and traffic redistribution and 6.route failure recovery these components are explained in detail in the forth-coming subsections. fig.1 depicts the proposed load aware congestion adaptive multipath multicasting approach. 60 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting fig. 1 proposed ladamr approach 3.1 resource monitoring 3.1.1 link stability in mobile ad hoc networks, the mobility induced by nodes as well as the propagation effects cause a packet to suffer fading effect. the link stability can be measured using signal-to-noise ratio (snr) and can be determined with the help of hardware component. if the estimated snr value is less than the threshold limits the packets will contain excessive errors due to high noise. this result in retransmission their by increases the overall delay thus degrades the performance significantly. the link stability can be estimated as follow using an equation (1). ber = 0.5 ∗ func( √ rsp ∗ cb np ∗br (1) where, ber = bit error rate, rsp = received signal power, cb = channel bandwidth, np = noise power and br = bit rate and func = error function the signal to noise ratio (snr) for multiple packet transmission can be estimated using the equation (2) as: snr = 10log rsp np + ∑n i=1rspi (2) where, ∑n i=1rspiis the signal strength of packets at the receiver.n is the number of packets received instantaneously. when a node is sending the data packet it appends its signal strength i.e., a transmitted power then the receiving nodes advances in systems science and application(2016) vol.16 no.3 61 estimates the received signal strength using the free-space propagation model using the wavelength of the medium,the distance between the communicating nodes and unity gain of sending and receiving antennas as shown in equation (3). rss = tss ( λ 4πd )2 gsgr (3) 3.1.2 available bandwidth the bandwidth availability is one of the very important factors which determine the connectivity of the network. in general, packet forwarding between the source and the destination follows a multi-hop communication. hence, it is very much essential to ensure that whether the intermediate forwarding nodes have sufficient bandwidth to forward the data or not. the nodes in the wireless networks rely on the shared wireless links and the links are severely affected by fading, inference, and path loss[23]. it is estimated by measuring the idle periods of the wireless channel. each node in the network listens to the channel and obtains the status to estimate the channel idle period using the channel observed time interval (coti). then the channel idle time (cii) can be estimated by increasing the count from the previous busy time to the start of the next busy time. let us consider the total channel idle time consists of several channel idle slots, say n. total channel idle time (citi) is the sum of all n idle times. thus, available bandwidth at a node is estimated using equation (4): ab = ∑n i=1cii coti ∗bwtotal (4) 3.1.3 estimation of residual battery battery lifetime (bli) of a node i is estimated using the residual energy (rei) and new and old drain rate (dr) of a node i and is estimated as shown in equations (5) and (6): bli = rei ∝ ∗drold(j, k) + (1− α)drnew(j, k) (5) drnew(j, k) = ej,k (1− perror)n (6) where, drold and drnew represents old and current calculated drain values and represents a constant value between 0 and 1. the number of neighbors of node says n is xk where n knows its entire neighbors total current load, tcl(i) then the probability of data forwarding can be computed as shown in equation (7): p (n) = 1− ( tcl(n)i macqi(i) ) (7) 62 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting 3.2 congestion monitoring detecting congestion in a reactive manner will produce longer delay, high packet loss, and high control overhead. unlike wired and high-speed networks congestion in manet becomes more viable during transmission of large-scale multimedia data. hence, eliminating congestion in such dynamic networks produce excessive overhead, delay, and a waste of resources. many solutions have adapted active queue management strategies to eliminate the congestion problems. this work aims to propose an approach which is adaptive to the incoming traffic and anticipates congestion by redistributing the traffic along the available congestion free path. each node in the network estimates the incoming traffic and updates the same in the neighbor table periodically. this helps to find out the neighbor node current load status. the current data traffic(cdt) is estimated to find out the congestion level of the node. each node ni samples the queue length in the mac layer periodically.suppose qj(k) is the kth sample value, and x is the overall sampling time period, then the current data traffic of node ni can be estimated using the equation (8). cdt (i) = ∑x k=1qj(k) x (8) the total length of the queue of node ni is the maximum capacity of the queue in the mac layer is macqi(i); then the total current load is defined as follows using equation (9). tcl(i) = cdt (i) macqi(i) (9) to monitor congestion well in advance, the average queue size is estimated by setting the static maximum and minimum threshold value for the queue length as qminth = 0.25 ∗ size of buffer and qmaxth = 0.75 ∗ size of buffer the current queue size can be estimated using equation (10) as: currentavgqsize(i) = (1− wq) ∗avgqold+ tcl(i) ∗ wq (10) where wq is the queue weight, is a constant (wq = 0.002) from red queue results in floyd, (1997). the current congestion status is computed as shown in equation (11). currentcs(i) = tcl(i)− currentavgqsize(i) (11) if the currentcs(i) is less than qminth ,then the incoming data traffic is below the buffer size and hence, a node can handle the traffic. if currentcs(i) ≤ advances in systems science and application(2016) vol.16 no.3 63 qminth and ≥ qmaxth , then the buffer overflow likely to take place perform packet drop probability to avid the packet loss. finally, if currentcs(i) ≥ qmaxth , then the node is congested, invoke redistribute the route through the available non-congested alternate path. table 1 list of symbolizations used in this work symbol description rsp received signal power cb channel bandwidth np bit rate br error function func interference ranges of the nodes snr signal to noise ratio rss received signal strength gs sending antennas gain gr receiver antennas gain propagation wavelength of the medium tss transmitted signal strength d distance between any two communicating nodes coti channel utilized time interval cii channel idle time tcii total channel idles time ab available bandwidth bli battery lifetime of node i dri drain rate of node i rei residual energy of node i xk set of neighbor nodes of k tcl total current load p(n) the probability of data forwarding cdt the current data traffic x overall sampling time macqi(i) current load on mac layer of node i qminth the minimum queues threshold limit qmaxth the maximum queue threshold limit wq the queue weight currentcs the current congestion status of node currentavgqsize the current average queue size rfn the set of resource full node 64 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting 3.3 construction of resource full node (rfn) list fig.2 shows sample network scenario. each node periodically updates its one-hop neighbor resource information. the periodic interval is set to 1sec. it helps each mobile node in the network to know about its one-hop neighbor resource information for constructing one-hop and two-hop congestion free resource-full node (rfn) list. after that, this set of congestion free (rfn) nodes will be used as a subset of a forwarding node to forward the datagram from the corresponding source to a destination. fig. 2 sample network scenario each mobile node updates its one-hop and two-hop neighbor list in its routing table and the same is used during the route discovery process to build congestion free primary path. the routing table contains the following fields information for each route entry: multicast rt { src addr is the source mobile node address, dst addr is the destination node address, grp addr is the multicast group address, hop cnt is the number of intermediate hops i.e. hop count, rfn node addr is the resource-full node address, rfn set is the list of resource-full node set, and advances in systems science and application(2016) vol.16 no.3 65 con status is the neighbors congestion status } table 2 resource full node list node id one-hop resource full node id two-hop resource full node id s 1, 2 4, 7 1 s, 4 5, 7 2 s, 7 4, r3 4 1, 5 7, r3 5 4, r1 7, 9, r3 7 2, r3 5,9 9 r2, r3 r1 r1 5 9 r2 9 5 r3 7,9 4,5 3.4 congestion free primary route discovery ldamm is an on-demand protocol initiates route discovery when a source mobile node has data to send. it first checks from its resource full node list whether the multicast receiver is in two-hop resource full node list or not. if the multicast receiver is in the two-hop resource full list, then it forwards the jrreq using the existing path in its routing table. if not, then the source node initiates a route discovery process by just forwarding the jrreq packet through its one-hop and two-hop resource full node set rather than flooding the jrreq packet into the network. this procedure helps in minimizing the control overhead to a certain extent. upon receiving this packet, the receiver node checks its two-hop resource full node list. if the multicast receiver found, then it forwards the jrreq packet directly to it. the multicast receiver then responds to the first received jrreq packet and sends back jrrep packet. if not found, it updates the received information and forwards jrreq packet to its one-hop resource-full node. this process repeats until the multicast receiver node found. the first jrrep path considered as a primary path between the source and the multicast receiver. finally, the source finds a resource full non-congested primary path to the destination. a primary route found in this case, are s→1→4→5→r1,s→2→7→r3,and s→2→7→r3→r2. these routes are used by the source to transmit a datagram towards the multicast receivers. thus, the proposed work finds a resource full congested free primary path from source to destination with the help of resource full node list and controls the overhead by avoiding unnecessary flooding of packets. table 3 describes the overall procedure involved in congestion free primary route discovery process after resource full node list selection. 66 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting table 3 procedure for congestion free primary route discovery process input: g= (v, e) output: the congestion free multicast tree begin 1) the source mobile node s checks its one-hop rfn list (1, 2) and a two-hop rfn list (4, 7) to find whether a multicast receiver is available or not. 2) if the multicast receiver nodes r1, r2, and r3 is not in one-hop and two-hop rfn list, then the source mobile node forwards jrreq packet to its one-hop resource full nodes 1 and 2. 3) now upon receiving jrreq node 1 and node 2 will check its one and two hop rfn list to find whether r1, r2, and r3 is available or not. if not 1 and 2 forwards a jrreq packet to its one-hop rfn nodes i.e.,(4, 7) 4) this process is repeated until jrreq reaches a multicast receiver. in this case, node 2 and node 4 finds the multicast receivers is its two-hop resource full node list and forwards the jrreq through intermediate nodes 5 and 7. 5) the multicast receiver node now sends a jrrep along the reverse path of the jrreq to reach the source. 6) the source mobile node now fixes the first jrrep as a congestion free primary path and starts the transmission along this path. end 3.5 congestion adaptive alternate route discovery each node in the primary path periodically estimates and finds its congestion status. if it is likely to be congested then warns its upstream and downstream node by sending congestion warning packet (cwp). upon receiving this packet, the upstream node checks it updated resource full node (rfn) list to find whether a multicast receiver is in it or not. if exists, exchange the new rfn list with its neighbors and resume the transmission along the newly available congestion free path. this new alternate route gets updated in its routing table. if not forwards the cwp to its previous node. if no rfn list found on the congestion free primary path, then the cwp, send to the source node. the source now assigns another alternate path if it is available otherwise initiates a new route discovery process to find a congestion free primary path. this alternate path finding process does not incur any significant overhead, because of the availability of one-hop and two-hop resource full node list in each node. for example, if the resource-full node 9 detects the congestion, sends a cwp to its neighboring nodes on the primary path in this case r2 and r3 and updates the rfn list in the routing table. in response, the upstream node r3 checks its routing table new rfn list along the primary path. if exits, traffic will be resumed through the available rfn nodes otherwise it forwards the cwp to advances in systems science and application(2016) vol.16 no.3 67 its previous node. in this case, there is no such alternate congestion free alternate path exists for node r2 it forwards towards its source node. the source node then initiates a new route discovery process. table 4 presents the overall procedure involved in finding the congestion free alternate routes. table 4 procedure for congestion free alternate route input: g=(v, e),multicast sessions. output: the congestion free of alternate route v ∈ rms begin 1) initialize the current queue buffer size, average queue size new and old as 0. 2) set minimum queue threshold limit to 0.25 * current queue buffer size and maximum queue threshold limit to 0.75 * current queue buffer size and queue utilization. also, set queue weight as 0.002 3) check if current avg que size is half of queue size. 4) for each arriving packet in queue increment the instantaneous queue size. 5) if it is a non-empty queue size, then apply the formula and if queue average new is less than queue minimum and queue average new together which is less then congestion warning limit, then set queue status as safe. 6) else if queue average new is greater than queue minimum and queue average new put together which is less than queue maximum then set queue status as likely to be congested 7) if the instantaneous queue size is greater than queue maximum and alternate path be false together then update queue maximum. 8) else queue status is congested 9) update queue average old as queue average new and queue weight. end 4 results and discussion a comparison of ladamm performance with that of maodv is done for various network scenarios using the network simulator (ns2.34) . the observations are presented below 4.1 simulation configuration and performance metrics the network consists of 100 nodes in a 1700 * 1700 m terrain size. the radio range is set to 250 m with bandwidth 2 mbps. to detect the link breaks using feedback mechanism ieee802.11 dcf is used. the channel propagation model used is two-way ground propagation model[24]. a queue size at each node is set to hold only 50 data packets and a routing buffer size is set to 64 data packets. the queue and buffer value is initially set to zero until the route discovery process. 68 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting the routing protocols used for performance analysis is maodv. the data flow used constant bit rate (cbr), which varies from 1 packet/sec to 50 packets/sec. the mobility speed of the node is varied from 5m/s to 30 m/s and each scenario is simulated for 600 s. table 5 describes a list of simulation configuration parameters used for performance analysis for various network scenarios. table 5 list of simulation parameters used simulation parameters simulation parameters node placement scheme random propagation model two-way ground propagation environment size 1500mx1500 m number of nodes 100 transmitter range 250m bandwidth 1mbps simulation time 600s traffic type constant bit rate (cbr) packet size 512bytes number of packets transmitted by sources 100 mobility model random way point model packet rate 5-50packets/s 4.1.1 end-to-end delay the average end-to-end delay is a measure of time consumed to deliver a packet from the source to the destination due to buffering of packets, transmission, retransmission and propagation delays. 4.1.2 packet delivery ratio percentage of data packets received at the receivers out of the number of data packets generated by the cbr traffic sources. pdr(%) = total number of packets received total number of packets transmitted ∗ 100 4.1.3 routing control overhead is the ratio of a total number of control packets received to the total number of control packets generated during the simulation time. 4.2 overall performance evaluation the simulated results discuss the different network scenarios. various performance metrics such as end-to-end delay, packet delivery ratio, and control overadvances in systems science and application(2016) vol.16 no.3 69 head are evaluated to facilitate the performance of the proposed lacamm protocol. 4.2.1 impact of lacamm with maodv the end-to-end delay, packet delivery ratio and control overhead results with respect to varying cbr packet rates are shown below from fig.3 to 5. these figures clearly show that the proposed lacamm yields better results when compared to maodv. fig.3 represents the end-to-end delay for lacamm and aodv with respect to varying the cbr packets rates from 5 packets/s to 55 packets/s. this figure shows that the end-to-end delay for proposed lacamm is much smaller that of maodv for all values of packet rates. the delay variation is lacamm was less than that of madov enables the proposed work more suitable for real-time multimedia applications. fig.4 shows the obtained packet delivery ratio with respect to varying the cbr packet rate is much higher than that of maodv. this is because the proposed lacamm has an ability to adapt to the load and congestion. but when the cbr packet rate increase maodv fails to handle congestion and hence leads to poor packet delivery ratio. fig.5 shows the routing control overhead for maodv and lacamm with respect to varying cbr packet rates. the figure reveals that proposed lacamm has less routing control overhead that of maodv. this is because lacamm uses suppressed flooding concept during route discovery process with the help of rfn list and redistributes the traffic in case of congestion or route failure using an alternate congestion free path along the primary path. the re-route discovery takes place only if no alternate congestion free paths exists along the primary path and thus make the lacamm superior to the modav. fig. 3 attained end-to-end delay with respect to various cbr packet rates 70 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting fig. 4 attained packet delivery ratio with respect to various cbr packet rates fig. 5 attained routing control overhead with respect to various cbr packet rates 4.2.2 impact of lacamm with edaodv fig.6 to fig.7 shows the results obtained for the end-to-end delay, packet delivery ratio and routing control overhead of the proposed lacamm with edaodv. fig.6 shows the attained end-to-end delay for lacamm and edaodv with respect to varying the cbr packet rates. both protocols attain somewhat same end-to-end delay when the data rate is between 5 packets/s to 15 packets/s but at higher data rates form 25 packets/s to 55 packets/s the proposed lacamm attain lower delay that of edaodv. this is because the availability of alternate congestion-free routes at each node along the primary path. fig.7 show the attain advances in systems science and application(2016) vol.16 no.3 71 packet delivery ratio for lacamm and edaodv with regard to the packet rates. the result reveals that for lower data rates 5 packets/s to 15 packets/s both protocols attain somewhat same packet delivery ratio but for higher data rates from 25 packets/s to 55 packets/s lacamm improves the packet delivery ratio. fig.8 shows the attained routing control overhead with respect to varying packet rates. from the figure, it is clearly understood that the proposed lacamm has lower routing control overhead that of edaovd for all packet rates. thus, these results reveal that the proposed lacamm achieves better result when compared with other two protocols madov and edaodv improving network performance. fig. 6 attained end-to-end delay with respect to various cbr packet rates fig. 7 attained packet delivery ratio with respect to various cbr packet rates 72 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting fig. 8 attained routing control overhead with respect to various cbr packet rates 5 conclusion congestion control techniques have been specifically made for multimedia applications in manets. a suitable mechanism needs to be implemented such that network characteristics like congestion and route failure need to be found out and an apt solution needs to be supplied. the proposed approach could analyze the fluctuation in traffic and categorize the congestion status accurately to solve the congestion problem which is robust and dynamic in nature to estimate congestion. the lacamm controls the congestion by using an alternative path after estimating the congestion status at the node levels along a path. the dcd-ammrp algorithm shows considerable performance over the maodv and edaodv. the ns-2-based simulation confirms that the lacamm outperforms in terms of delay, packet delivery ratio and routing overhead than that of maodv and edaodv. references [1] avokh, a. and g. mirjalily. (2013), “load-balanced multicast tree routing in multi-channel multi-radio wireless mesh networks using a new cost function”, wireless personal communications, vol. 69, no. 1, pp. 75-106. [2] baolinr, s. and l.l yuan. (2006), “distributed qos multicast routing protocol in ad hoc networks”, journal of systems engineering and electronics, vol. 17, no. 3, pp. 692c698. [3] bonmariage, n and g. leduc. (2006), “a survey of optimal network conadvances in systems science and application(2016) vol.16 no.3 73 gestion control for unicast and multicast transmission”, computer networks, vol. 50, no. 3, pp. 448-468. [4] bawa, o. s. and s. banerjee. (2013), “congestion based route discovery aomdv protocol”, international journal of computer trends and technology, vol. 4, no. 1, pp. 54-58. [5] chen, l. and w. b. heinzelman. (2007), “a survey of routing protocols that support qos in mobile ad hoc networks”, ieee network, vol. 21, no. 6, pp. 30-38. [6] floyd, s. and v. jacobson. (1993), “random early detection gateways for congestion avoidance”, ieee/acm transactions on networking, vol. 1, no. 4, pp. 397-413. [7] gawas, m. a., l. j. gudino and k.r. anupama. (2015), “cross-layer best effort qos-aware routing protocol for ad hoc network”, international conference on advances in computing, communications and informatics (icacci), kochi, india, pp. 999-1005. [8] yi y. and s. shakkottai. (2007), “hop-by-hop congestion control over a wireless multi-hop network”, ieee/acm transactions on networking, vol. 15, no. 1, pp. 133-144. [9] gulati, m. k. and k. kumar. (2013), “a review of qos routing protocols in manets”, international conference on computer communication and informatics (iccci), coimbatore, tamil nadu, india, pp. 1-6. [10] jetcheva, j. g. and d.b. johnson. (2001), “adaptive demand-driven multicast routing in multi-hop wireless ad-hoc networks”, in proceedings of the 2nd acm international symposium on mobile ad hoc networking and computing,long beach, ca, the usa, pp. 33-44. [11] yu, y. and g. giannakis. (2008), “cross-layer congestion and contention control for wireless ad hoc networks”, ieee transactions on wireless communications, vol. 7, no. 1, pp. 37-42. [12] kumar, n., n. chilamkurti and j. h. lee. (2012), “a novel minimum delay maximum flow multicast algorithm to construct a multicast tree in wireless mesh networks”, computers and mathematics with applications, vol. 63, no. 2, pp. 481-491. [13] li, j., m. yuksel and s. kalyanaraman. (2006), “explicit rate multicast congestion control”, computer networks,vol. 50, no. 15, pp. 2614-2640. 74 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting [14] karunakaran, s. and p, thangaraj. (2010), “a cluster based congestion control protocol for mobile ad-hoc networks”, international journal of information technology and knowledge management, vol. 2, no. 2, pp. 471-474. [15] narayan, d. g., r. nivedita, s. kiran and m. uma. (2012), “congestion adaptive multipath routing protocol for multi-radio wireless mesh networks”, international conference on radar, communication and computing (icrcc), tiruvannamalai, india, pp. 72-76. [16] xie, f., g. feng and c.k. siew. (2006)., “the impact of loss recovery on congestion control for reliable multicast”, ieee/acm transactions on networking (ton),vol. 14, no. 6, pp. 1323-1335. [17] lochert, c., b. scheuermann and m. mauve. (2007), “a survey on congestion control for mobile ad hoc networks”, wireless communications, and mobile computing,vol. 7, no. 5, pp. 655-676. [18] lucas, v., j.j pansiot, d. grad, and b.hilt. (2013), “robust and fair multicast congestion control (m2c)”, computer networks, vol.57, no. 3, pp. 699-724. [19] rishiwal, v., s. verma and s. k. bajpai. (2009), “qos based power-aware routing in manets”, international journal of computer theory and engineering, vol. 1, no. 1, pp. 47-57. [20] senthilkumaran, t. and v. sankaranarayanan. (2011), “early congestion detection and optimal control routing in manet”, european journal of scientific research, vol. 63, no. 1, pp. 15-31. [21] sheeja, s., and r.v. pujeri. (2013), “effective congestion avoidance scheme for mobile ad hoc networks”, international journal of computer network and information security, vol. 5, no. 1, pp. 33-40. [22] sreedhar, g. s and a. damodaram. (2012), “malmr: medium access level multicast routing for congestion avoidance in multicast mobile ad hoc routing protocol”, global journal of computer science and technology, vol. 12, no. 13, pp. 23-30. [23] tang, k. and m. gerla. (2003), “congestion control multicast in wireless ad hoc networks”, computer communications, vol. 26, no. 3, pp. 278-288. [24] tran, d. and h. raghavendra. (2006), “congestion adaptive routing in mobile ad hoc networks”, ieee transactions on parallel and distributed systems, vol. 17, no. 11, pp. 1294-1305. advances in systems science and application(2016) vol.16 no.3 75 corresponding author n. jaisankar can be contacted at: hsanthi@vit.ac.in adv syst sci appl 2019; 2:80–89 http://ijassa.ipu.ru/index.php/ijassa/article/view/707 aodv-based key management in vanet chaima bensaid1∗, sofiane boukli-hacene 1 1eedis laboratory , computer science department, djillali liabes university at sidi bel abbes , sidi bel abbes , algeria e-mail: chaimaa184@hotmail.fr , boukli@gmail.com received february 23, 2019; revised june 21, 2019; published july 10, 2019 abstract: vehicular ad hoc network (vanet) is a self-organized multi hop system comprised by multiple vehicles. this kind of network offers different kind of communication such as vehicle to vehicle (v2v) and vehicle to infrastructure (v2i) communication. the v2v communication a folk of mobile ad hoc network(manet). it is characterized by a high mobility, dynamic topology, lack of centralized control infrastructure and a non secured shared media. similar to manet, v2v is vulnerable to a many types of attacks such as blackhole, sybil and denial of service attack. security in such cases is primordial. many studies have addressed this issue and many solutions have been proposed. one of the most promising studies used distribution certification authority (dca) to secure communication within vanets networks. in this paper, we propose a new method of certification based on the study of the behavior of each node based on aodv routing protocol. the operational performance is evaluated with rigorous analysis and extensive simulation study. our approach reduces certificate with 6.01 to 69.42% less certificates. keywords: vanet,aodv, clustering , trust based mechanism. 1. introduction vanet is a network in which a mobile node is a smart vehicle equipped with a sensors. it offers three types of communication: vehicles to vehicles (v2v) and vehicles to road side unit infrastructure (v2i) and hybrid communication. many researches state that v2v vanet communication are vulnerable to a lot of attacks because the dynamic topology and the use of wireless links. to secure these communications a security mechanism is essential. a public key infrastructure assists that the users obtain the necessary public keys, these public keys are used to perform cryptographic operations. the ca (certificate authority) issues the digital certificates and provides the means operations to verify the validity of the certificates. however, this solution is costly and not suitable for the different requirements of the vanets. lightweight solutions have been also proposed such as dca and mobile certificate authority (moca). cluster based ca have been proposed for cluster based routing protocols (cluster based routing protocol). these solution suffer from cluster head (ch) overload because these nodes issue new certificates when a mobile cluster number moves between clusters. in our paper, we propose a new approach to adopt the pki in vanet based on the well known routing protocol aodv with a trusted system when an extension is used to predict node migration to adjacent clusters. the trusted system proposed is lightwight and it is based only on neighboring node behavior and the prediction process predict the destination cluster ∗corresponding author: chaimaa184@hotmail.fr aodv-based key management in vanet 81 and anticipate certificate issuing which minimize the number of issued certificate and allow certification renewal with in all clusters. our improvement gives a satisfactory result, where the number of issued certificates has declined by about 69%. 2. aodv routing protocol the aodv protocol is a reactive routing protocol [1, 2]. it operates using three types of messages: request messages (rreq) and route reply messages (rrep) and route error messages (rerr). when a node wants to send a data packet to a destination node, it looks in its routing table if a fresh route exists to the destination node. if there exist already a valid path it is use to route packets, else, it launches a route discovery process by broadcasting a rreq. each intermediate nodes check if it is the destination node or it has a fresh route to the destination node. if so, it sends a rrep to the source node. upon receiving the rrep, the source node uses the discovered route to exchange packets with the destination [3]. these route are not secured, for that many studies have focused on security in vanet. in the next section, we’re going to discuss security mechanism in vanet. fig. 2.1. rreq and rrep packets in aodv 3. key distribution in this section, we present key distribution mechanism present in literature. zhou and hass [4] proposed a partially distributed ca where the key management service (k, n) consists of n servers. each node has a public/private key, the private key is divided into n shares, one share for each server,and the public key is known to all nodes in the network. each ca generates a portion of the certificate using its private key and sends it to the combiner. with k correct partial signatures the combiner can construct the complete certificate for member node. it is always possible for a combining node to be compromised by an adversary or be unavailable . authors have not paid too much attention to the certificate revocation; they proposed a simple approach which is a certification revocation list (crl). yi and kravet [5] proposed a distributed certificate authority based on threshold cryptography where the signature scheme does not require a combiner c. the distribution certification authority is called mobile certificate authority (moca). all nodes are equipped with mp (moca certification protocol), the combination of different parts of the signatures is done by each node. this approach is unsuited for vanet because all certificates must be known by the dca servers certificates must be known by the dca servers before providing any access to certification . a fully distributed certification service based on clustering have proposed by kong and al [6] . the establishment of certificates in the network is provided by the nodes themselves, copyright c© 2019 assa. adv syst sci appl (2019) 82 c. bensaid, s. boukli-hacene which will be stored by a particular node called cmn (certificate management node) for each cluster. all nodes must request the cmn of the same cluster to collect the certificate strings for each authentication. this system converges to the centralized models where the certificates are stored by a set of special nodes, which puts into question the availability of the certification service especially when cmn became unreachable. in [7], a based cluster certification authority architecture is proposed. each cluster-head (ch) has a ca information table, which contains a list of ca nodes. when a member node requires a certificate, it sends a request to its cluster-head ch to get information about the ca servers. the ch collects information, and forwards to the member node. in this time the node sends the certification request to them. this approach generates lower certificate maintenance overhead by resolving certificate transaction problems. chaining is done only with trusted nodes that are cluster head and gateways. mukherjee et al in [8] proposed an extention of aodv protocol by adding a new trust mechanism and considering the successful cooperation frequency and average encounter rate as factors to compute for direct trust. modified dempster-shafer theory is used to build the recommended trust based on multiple pieces of trust evidence. the trust model is composed of two phases : route discovery and trusted route selection which selects the most reliable next hop for routing discard nodes with high mobility and high drop packets. ahmed et al. [9] proposed a modified route discovery algorithm to efficiently and securely route data to its destination. this flooding algorithm is used to define the link failure probability of misbehaving nodes and normal nodes and an enhanced multi-swarm optimization is used to optimize the discovered route. 4. proposed approach we propose a cluster based trust model (cbtm) to secure data exchanges in vanets networks. to degrade network efficiency in a vanet network, a malicious node sends a great number of rrep to intercept the data packets of its neighbors or to overload the network. for that it uses a very high sequence number in the rrep packet to attract the data packets of its neighbors to go through him to edit or delete them later. in our model, each node in the network is equipped by a cache memory, in which it saves the number of data packets, route request packets (rreq) and the number of route reply packets (rreps) received from this node and the sequence number. each node will update its cache memory with the following functions : • recv rreq () cache [0] [@src] ++; • recv data () cache [1] [@src] ++; • recv reply () cache [2] [@src] ++; copyright c© 2019 assa. adv syst sci appl (2019) aodv-based key management in vanet 83 n d: the number of data packets received from a node; n request: the number of rreq packets received from a node; n reply: the number of rrep packets received from a node; if n d +n request == 0 and n seq dest >>>>> n seq src then trusted value=0; else trusted value=1; end if n d +n request ! = 0 and n rply < n d +n request then trusted value=1; end if n d +n request == 0 and n reply ! = 0 then trusted value=0; end algorithm 1: computing trust value if the received packet is a rrep, it looks in its cache memory to check the stored values. it will use the following algorithm to decide whether the node is a trusted node or not. 4.1. clustering to divide the networks into clusters each node uses its neighbors table. in our implementation, the size of each cluster is empirically fixed to 5, and within each cluster there is a single ch which is the node that has the highest trust value and the higher sequence number in the cluster. in addition, only a trust node can be a cluster member. 4.2. certification hahn and al [10] proposed a model for manet, where the cluster-head acts as a ca. the certificate chain allows the exchange of session keys and the encryption / decryption of data. however, any node can be selected as a ca. due to mobility which is a important feature of the v2v communication, a member node request a new certificate each time it passes from one cluster to another which will overload the number the certificate generating. our approach is a trust model based on the study of the total behavior of the member node and if the member node passes of cluster to another, the certificate is broadcast by the gateways. in this paper, we develop this proposal with detailed simulations study by the well known network simulator ns2.35 [11]. when a node enters in the preemptive region, three signal values are collected and we used the lagrange interpolation to predict node mobility. the formula of the interpolation is: y = n∑ i=0 [ ∏n j=0,j 6=i (x− xj)∏n j=0,j 6=i (xi − xj) × yi ] (4.1) we store three signals values of received data packets and their corresponding receiving time. when two consecutive measurements give the same signal, we store the time of the second occurrence. the expected signal strength p of the packets received from the ch node is computed as follows [12, 13]: p = ( (t− t1)× (t− t2) (t0 − t1)× (t0 − t2) × p0 ) + ( (t− t0)× (t− t2) (t1 − t0)× (t1 − t2) × p1 ) + ( (t− t0)× (t− t1) (t2 − t0)× (t2 − t1) × p2 ) (4.2) where p0 , p1 , p2 are the measured power strengths at the times t0 , t1 , and t2 respectively. the time t is the sum of time required to send the certificate to cluster adjacent copyright c© 2019 assa. adv syst sci appl (2019) 84 c. bensaid, s. boukli-hacene (inonde period) and the difference between t2 and the average value of the measurement; this value has been determined empirically [14]. t = 2× t2 − ( t0 + t1 + t2 3 ) + inonde period (4.3) when p is less than the minimum acceptable power (81dbm) a warning message is sent to the ch. the ch sends the certificate to adjacent cluster [12]. when the join the adjacent cluster, the cluster-head compares the nodes address with the received one and saves its certificate. fig. 4.2. certficate generation 5. performance evaluation to evaluate the performance we used ns2 simulator . in our approach we used two mobility scenarios, the first is of the city of malaga and the second is generated by the simulator vanetmobisim . 5.1. malaga city scenarios fig.5.3 present a geographic card of urban vanet scenarios from the downtown of malaga, spain [15, 16] . it is composed of three areas u1 , u2 and u3. detailed parameters of the simulation area is presented in the table 5.1. copyright c© 2019 assa. adv syst sci appl (2019) aodv-based key management in vanet 85 table 5.1. vanet scenarios details scenario area size number of vehicules number of connections u1 120000m2 60 10 15 u2 240000m2 60 20 u3 360000m2 60 30 40 fig. 5.3. malaga urban areas in this scenario 60 vehicles are simulated. each of them use aodv routing protocol. a cbr application over udp is used to simulate data transmission between pairs of communicating nodes where packet of 1 kb are sent using a rate of 100 kpbs during 180s. table 5.2 summarize all used parameters. table 5.2. simulation paramaters parameters value propagation model nakagami phy layer ieee 802.11p mac layer ieee 802.11p routing layer aodv transport layer udp cbr packet size 1024 bytes cbr packet rate 100kbps simulation time 180 s 5.2. vanetmobisim scenarios the vehicular ad hoc networks mobility simulator (vanetmobisim) [17] is a set of extensions to user mobility modeling framework canumobisim, used by the canu (communication in adhoc networks for ubiquitous computing) research group, university of stuttgart. it includes a visualization module, mobility models, as well as various formats parsers for geographic data sources. this framework is easily extendable and it is based on the concept of pluggable modules. the set of extensions provided by vanetmobisim consists mainly on a vehicular spatial model using gdf-compliant data structures, and a set of vehicular-oriented mobility models. copyright c© 2019 assa. adv syst sci appl (2019) 86 c. bensaid, s. boukli-hacene table 5.3 details all parameters used in this scenario where protocol is under attack. table 5.3. simulation paramaters parameters value propagation model nakagami phy layer ieee 802.11p mac layer ieee 802.11p area size 1000*1000 m vehicule speed from 8.33 to 13.89 m/s cbr packet size 1024 bytes cbr packet rate 100kbps simulation time 900 s number of malicious nodes 3 the well-known metrics to evaluate routing protocol is used in our study : • packet delivery ratio (pdr): represents the percentage of packets delivered to their destinations . • the average latency of data packets (delay): this is the average time required to deliver data packets from the source to the destination successfully. • additive costs (overhead): this criterion illustrates the amount of additives cost required for each received data packet. • dropped packet: number of dropped packets due to either link failure or by the malicious node. • certificate overhead: the number of certificates sent in network. simulation scenarios are : • aodv vanetmobisim denotes aodv under attack. • cbtm vanetmobisim represent our approach under attack used vatenmobisim scenario. • aodv malaga denotes aodv under attack. • cbtm malaga represent our approach under attack used malaga scenario. fig. 5.4. certificate overhead fig.5.4 shows the certificate overhead. we observe that the overhead increase high because many nodes request a new certificate from ca and from adjacent ca when a node moves. however, we depict that, our approach outperforms the original with 6.01% to 44.88% less certificates by malaga mobility model and 40.40% to 69.42% less certificates by the model generated by vanetmobisim. copyright c© 2019 assa. adv syst sci appl (2019) aodv-based key management in vanet 87 fig. 5.5. packet delivery fraction the fig. 5.5 shows the decreasing evolution of pdf in protocol aodv under attack against to our approach with two mobility models. when the number of vehicles is high, the pdf of our approach register a small degradation. this is due to frequent network topology changes with a high number of connections. we also observe that if the number of vehicles increases the pdf increases for our solution. in the aodv protocol under attack the pdf reach 19.98% . fig. 5.6. dropped packet the fig. 5.6 presents the number of dropped packets depending on the number of vehicles. in our proposal, with 10 sources of connection, the number of lost packets is 69 and almost no significant against to the aodv under attack 9971 packets in malaga mobility model. whereas the number of connections increases, it means that there is a lot of data packet sent. in this case, the malicious node intercepts a large quantity of packet so the rate of dropped packet, but in our proposal is a little minimal. fig. 5.7. normalized routing load copyright c© 2019 assa. adv syst sci appl (2019) 88 c. bensaid, s. boukli-hacene the simulation results in fig. 5.7 show that our proposal is a small routing head relative to aodv under attack. this is due to the fact that when the packet is not received in the right destination, the source node attempts to repair the paths with rrer packets. fig. 5.8. communication delay the fig. 5.8 shows the evolution of the average end to end delay, depending on number of vehicles in our approach and the aodv under attacks with two mobility models. it is found that the time required by our proposal is higher than the aodv under attack. this can be justified by using an extra process in our solution to create the certificates and update him. 6. conclusion in this paper, we presented the different approaches proposed based on pki in vanet. we also presented a new proposal based on trusted model. in our proposal, we interested to v2v communication system where the cluster-head node acts as a virtual ca the main idea is to compute a confidence level for each node by traking its behavior and avoiding the issue of new certificate request in case a node moves. our approach reduces certificate with 6.01% to 44.88% less certificates by malaga mobility model and 40.40% to 69.42% less certificates by the model generated by vanetmobisim. references 1. gupta, p., goel, p., varshney, p., & tyagi, n. (2019). reliability factor based aodv protocol: prevention of black hole attack in manet. in smart innovations in communication and computational sciences (pp. 271-279). springer, singapore. 2. kandali, k., & bennis, h. (2018, july). performance assessment of aodv, dsr and dsdv in an urban vanet scenario. in international conference on advanced intelligent systems for sustainable development (pp. 98-109). springer, cham. 3. perkins, c. e., belding-royer, e. m., & das, s. r. (2002). mobile ad hoc networking working group-internet draft. 4. zhou, l., & haas, z. j. (1999). securing ad hoc networks. ieee network, 13(6), 24-30. 5. yi, s., & kravets, r. (2004). moca: mobile certificate authority for wireless ad hoc networks. 6. kong, j., zerfos, p., luo, h., lu, s., & zhang, l. (2001, november). providing robust and ubiquitous security support for mobile ad hoc networks. in icnp (vol. 1, pp. 251-260). copyright c© 2019 assa. adv syst sci appl (2019) aodv-based key management in vanet 89 7. dong, y., sui, a. f., yiu, s. m., li, v. o., & hui, l. c. (2007). providing distributed certificate authority service in cluster-based mobile ad hoc networks. computer communications, 30(11-12), 2442-2452. 8. mukherjee, s., chattopadhyay, m., chattopadhyay, s., & kar, p. (2018). eaer-aodv: enhanced trust model based on average encounter rate for secure routing in manet. in advanced computing and systems for security (pp. 135-151). springer, singapore. 9. ahmed, m. n., abdullah, a. h., chizari, h., & kaiwartya, o. (2017). f3tm: flooding factor based trust management framework for secure data transmission in manets. journal of king saud university-computer and information sciences, 29(3), 269-280. 10. g. hahn, t. kwon, s. kim, & j. song, ”cluster-based certificate chain for mobile ad hoc networks,” in international conference on computational science and applications (iccsa) , pp. 769-778, 2006. 11. issariyakul, t., & hossain, e. (2011). introduction to network simulator ns2. berlin: springer 12. boukli-hacene, s. , lehireche, a., & meddahi, a. (2006). predictive preemptive ad hoc on-demand distance vector routing. malaysian journal of computer science, 19(2), 189-195. 13. laroui, m., sellami, a., nour, b., moungla, h., afifi, h., & boukli hacene, s. (2018, december). driving path stability in vanets. in 2018 ieee global communications conference (globecom) (pp. 1-6). ieee. 14. boukli-hacene, s., ouali, a., & bassou, a. (2014). predictive preemptive certificate transfer in cluster-based certificate chain. international journal of communication networks and information security, 6(1), 44. 15. toutouh, j., & alba, e. (2011, july). an efficient routing protocol for green communications in vehicular ad-hoc networks. in proceedings of the 13th annual conference companion on genetic and evolutionary computation (pp. 719-726). acm. 16. malaga city downtown scenario . http://neo.lcc.uma.es/staff/jamal/vanet/?q=node/11 17. vanetmobisim manuel, (2006) institut eurcom/politecnico di torino copyright c© 2019 assa. adv syst sci appl (2019) adv syst sci appl 2019; 04; 45-57 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/748 evaluating comparative total factor productivity in the russian industry svetlana v. orekhova1*, evgeniy v. kislitsyn2 1) business economics dept., ural state university of economics, yekaterinburg, russia e-mail: bentarask@list.ru 2) information technology and statistics dept., ural state university of economics, yekaterinburg, russia e-mail: kev@usue.ru received may 14, 2019; revised november 28, 2019; published december 31, 2019 abstract: searching for domestic reserves of economic growth has lately become one of the central problems in russia. the paper examines the role of small industrial enterprises in stimulating economic growth. in contrast to the stereotype that small business serves as a driving force for the economic development and encourages innovation, the authors hypothesizes that in russia this statement is false. the conceptual and methodological framework of the study rests upon neoclassical models of economic growth. the authors investigate the existing approaches to the analysis of factors influencing economic growth of the state and choose the tools of total factor productivity analysis. total factor productivity is calculated using the translog production function, which allows determining the effect of the technological level on value added of the object under study. the choice of the type’s function is due to the low elasticity between the factors of production, as well as the imperfect competition in the industrial markets under review. the information base of the research includes the data of small, medium-sized and large enterprises of 10 industrial macro-sectors of the russian economy for 2013–2017. we use spark-interfax database to assess production functions. the results of the analysis prove that small enterprises operating in the russian industry demonstrate much lower values of average and weighted average total factor productivity than medium-sized and large enterprises. the general trend for such businesses is a decline in total factor productivity. only single leading companies produce a gain in value added in small entrepreneurship. thus, the economic situation in russia rejects the hypothesis about a higher entrepreneurial potential of small businesses, business models and technological innovations emerging on its basis. our further studies are assessing institutional traps and general context in the development of small business. keywords: total factor productivity, small enterprises, russian industry, economic development. 1. introduction harrod [20] and domar [12], who described a one-factor model for determining growth rates, provided fundamentals of the theory of economic growth. since the paradigmatic works by solow [30; 31] and swan [33], who proposed endogenous growth model, contemporary economic thought has been developing towards substantiating sources and refining growth models. the basic exogenous growth factors include scientific and technological progress and labor [27]. arrow [4] и uzawa [35] examine the role of knowledge and human capital in economic growth. the mechanisms of innovation growth are clarified in the schumpeterian growth model [1; 18] for russian economic development strategy, the key issue is the effective economic structure (and also institutional environment). this aspect “solves the question of the number of types of activities, markets, the size of the general welfare and its distribution (income * corresponding author: bentarask@list.ru mailto:kev@usue.ru 46 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) inequality)” [32, p.26]. equally important is the question of the contribution of different types of enterprises to economic growth. the history and practices of small business have been among the most controversial issues in economic development in different countries. science postulate that small business is an important object of economic growth [15; 17; 38; 39]. job creation, production of innovations and technologies, and finally, profits are the outcome of small businesses activity due to their entrepreneurial potential, flexibility and maneuverability of managerial decisions, optimization and red tape reduction of business processes. small businesses make a huge contribution to national prosperity. on the other hand, empirical studies demonstrate that the size of enterprises can influence their growth rate differently (the positive correlation is highlighted in [6; 13; 37] and the negative one is discussed in [3; 10; 15]). the role of small business in an economy has frequently been undermined and even misinterpreted. in the past, small businesses were believed to impede economic growth by attracting scarce resources from their larger counterparts [5]. these inconsistencies between theory and practice encourage more detailed studies to be conducted on the role of small enterprises in promoting economic development of the russian industry, and later – on the factors determining this role. the purpose of the research is to analyze total factor productivity (tfp) of small enterprises of russia’s industrial sector in comparison with medium-sized and large companies. such a research agenda is set for the first time and implies a consistent solution of a number of tasks. the first task is to perform a critical analysis of the existing theoretical approaches and russian empirical studies on measuring economic growth rates taking into account various factors. the second task is to substantiate the choice of a method for calculating tfp and its empirical testing. it is reasonable to carry out a comparative analysis of data on small and other enterprises in the context of individual industries, which allows taking into account the organizational and production specifics of enterprises. the authors state that industrial segments, serving as the basis for the rf government’s economic growth policy, are of the greatest academic interest. at the same time, regional, technologic and economic disproportions of industrial markets are very serios [2]. the third task of the research is to interpret the obtained results of calculating the total factor productivity from the standpoint of small industrial enterprises’ development in the russian economy. 2. theoretical approaches and empirical studies on factor productivity neoclassical economics theory showed that the level of technological development was the primary indicator of the degree to which society masters the forces of nature. the combination of productive forces and the type of production relationships constitutes a unique mode of production. neo-institutional economics largely associates the rate of technological growth with institutional environment for business development. [11; 28] total factor productivity is the most integrated indicator of technological progress and growing economic efficiency. it is calculated as the ratio of output to the volume of production factors used [23]. the logic of tfp calculation is as follows: an enterprise uses a certain set of factors of production (labour and capital) to produce the final product. then the growth of the factors of production used or a change in the combination of their use (technology) causes an increase in output of the final product. the dynamics of factor productivity may indicate the degree to which the growth of an enterprise, industry or economic sector is sustainable. measuring tfp means discovering the correlation between output (production quantity) and factors of production, such as resource evaluating comparative total factor productivity in the russian industry 47 copyright ©2019 assa. adv. in systems science and appl. (2019) costs and the level of technology. such a quantitative dependence was dubbed “production function”. there are numerous types of production functions and the most famous of them are the models of cobb–douglas, leontiev, ces, etc. the choice of these models is determined by a combination of factors that embrace data access, market specificity, a dynamic or static aspect of research. methods for assessing tfp are widely used throughout the world, including as a way of comparing the economic state of countries [14; 16; 21; 25] and the reasons for their differentiation [19; 26; 29]. in russia, the evaluation of the effectiveness of economic subjects are conducted on a regular basis using certain types of production functions. this fact is due to a number of reasons. first, the production function is an adequate tool for measuring the influence of various factors on economic growth. second, this tool allows analyzing intra-industry, crossindustry and inter-regional differentiation. third, in the context of uncertainty and volatility of incomes and expenditures of enterprises, such an assessment allows performing a real-time monitoring of the reasons behind the rise or decline in economic development and determining its trend. for example, bessonov [7] show that industries with relatively safe output dynamics and lacking sufficient incentives to increase productivity demonstrated the worst tfp dynamics in russia in 1989–2002. the tornqvist index was also calculated by orekhova [24] to analyze the growth sustainability in metallurgy. tornqvist index formula allows characterizing the change in the efficiency of resource use by factors. the tornqvist index deals with the change in the volume of two resources – labor (data on the average number of employees) and capital (data on the nominal value of fixed assets). the results of the analysis illustrate the overall technological and technical underdevelopment of the industry. timmer and voskoboynikov [34] prove that the growing role of capital expenditures in the growth of value added in russia resulted in the use of an extensive model of economic growth, which is typical of rental economies. bessonova [9] demonstrates that there is a significant gap in the total factor productivity of enterprises within the industry. such a gap may arise due to both organizational-efficiency-based reasons and resource or institutional constraints of the russian market. all evaluations of the russian industry’s effectiveness illustrate spasmodic and chaotic shifts, but also display a declining total factor productivity. at the same time, such assessments are more integrated and not explaining the behavior of individual firms, their contribution to the total efficiency of the industry or the entire economy. on the basis of the recent research [10; 39], we assume that small entrepreneurship, a priori not so rich in resources and power in comparison with large business, will be less efficient. 3. tfp research method when it comes to the method for calculating total factor productivity, our study rests upon a series of works by bessonova [8; 9]. total factor productivity is value added of the final product minus changes in labour and capital costs. empirically tfp growth can be calculated as an unexplained residue of the final product’s growth. this residue encompasses the effects from technological or organizational innovations that determine technological progress in the industry and influence the shift of the production function [9, p. 99]. however, such a calculation method is possible to be applied only on the basis of correct data on the share of costs of production factors by industry. the previous approach to the evaluation of tfp within the framework of the solow growth model [30] makes an assumption about competitive factors of production. but nowadays we knows that industrial markets are characterized by a hybrid form of organization, which obliges a researcher to first evaluate their production functions, and only after that calculate indicators of production factors and growth of tfp. 48 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) this fact rationalizes the choice of the translog production function for calculating tfp. the function allows one not to make prerequisites about absolute elasticity of substitution between factors of production and perfect competition in the markets of these factors [22]: ln 𝑌𝑖𝑡 = 𝛼0 + 𝛼𝐿 ln 𝐿𝑖𝑡 + 𝛼𝐾 ln 𝐾𝑖𝑡 + 𝛼𝑡𝑡 + 𝛼𝐿𝐿(ln 𝐿)2 + 𝛼𝐾𝐾(ln 𝐾)2 + 𝛼𝑡𝑡𝑡2 + 𝛼𝐿𝐾 ln 𝐿𝑖𝑡 ln 𝐾𝑖𝑡 + 𝛼𝐿𝑡 ln 𝐿𝑖𝑡𝑡 +𝛼𝐾𝑡 ln 𝐾𝑖𝑡𝑡 + 𝜀𝑖𝑡 (3.1) where yit denotes value added at the enterprise i for the period t; kit is fixed assets of the enterprise i for the period t; lit is wage at the enterprise i for the period t; t denotes a time factor ranging from 1 to n (where n is the number of observed periods). based on the proposed production function, the growth of tfp is calculated by formula: 𝐴𝑖,𝑡 = ln ( 𝑌𝑖,𝑡 𝑌𝑖,𝑡−1 ) − �̅�𝐿 ln ( 𝐿𝑖,𝑡 𝐿𝑖,𝑡−1 ) − �̅�𝐾 ln ( 𝐾𝑖,𝑡 𝐾𝑖,𝑡−1 ) (3.2) where a denotes a total factor productivity; is average elasticity of value added of labour; is average elasticity of value added of capital. average elasticity is calculated as average value of the elasticities of the added value of labour and capital for the periods (t – 1) and t, which in turn are measured as a partial derivative of the corresponding factor: �̅�𝐿,𝑡 = 𝜕 ln 𝑌𝑖𝑡 𝜕 ln 𝐿𝑖𝑡 = �̂�𝐿 + 2�̂�𝐿𝐿 ln 𝐿𝑖𝑡 + �̂�𝐿𝐾 ln 𝐾𝑖𝑡 + �̂�𝐿𝑡𝑡, (3.3) �̅�𝐾,𝑡 = 𝜕 ln 𝑌𝑖𝑡 𝜕 ln 𝐾𝑖𝑡 = �̂�𝐾 + 2�̂�𝐾𝐾 ln 𝐾𝑖𝑡 + �̂�𝐿𝐾 ln 𝐿𝑖𝑡 + �̂�𝐾𝑡𝑡, (3.4) to estimate value added and factors of labour and capital, it is required to use the following indicators: fixed assets, the revenue volume, total costs and wage. value added was calculated using formula: yit = volit-(tcit-wageit) (3.5) where volit is the revenue volume of the enterprise i for the period t; tcit is total costs of the enterprise i for the period t; wageit is labour cost borne by the enterprise i for the period t. the amount of capital was estimated as the average annual value of fixed assets, and labour – as the enterprise’s costs incurred in labour remuneration. 4. evaluating tfp of industrial enterprises in russia: small vs. large enterprises the study aims at primarily identifying the contribution of small enterprises to the economic growth of the russian industry. in accordance with the federal law no. 209-fz “on the development of small and medium-sized entrepreneurship in the russian federation” of july 24, 2007, a small enterprise is that which employs no more than 100 people and its revenue is below 800 million rubles. the empirical testing of the proposed method for calculating tfp was divided into several stages (fig. 4.1). evaluating comparative total factor productivity in the russian industry 49 copyright ©2019 assa. adv. in systems science and appl. (2019) stage 1. description of the research object stage 2. constructing production functions stage 3. evaluating total factor productivity stage 4. analysis and interpretation of results choosing the industries for investigation descriptive statistics data upload calculating value added forming panel data estimating production function s coefficients checking coefficients for reliability, unbiasedness and consistency calculating elasticities of value added of labour and capital for each period calculating total factor productivity for each period calculating average tfp of industry for each period calculating tfp of small enterprises for each period calculating average tfp of small enterprises in the industry for each period fig. 1. algorithm for comparative evaluation of tfp of small and large businesses hence, the authors use spark-interfax database to assess production functions; the period under study was from 2013 to 2017. the research object was industrial sectors of the russian federation. the entire set of industries is combined into 10 industrial sectors (tab. 4.1). this grouping is based on similar industry characteristics within groups. the group “small enterprises” embraces all the companies meeting the aforementioned criteria (micro-enterprises included). since tfp is an average of the indicators’ values for two years, tfp was calculated for 4 periods. table 4.1. number of enterprises in the study sample industry 2014 2015 2016 2017 enterprises large and medium -sized smal l and micr o large and medium -sized smal l and micr o large and medium -sized smal l and micr o large and medium -sized small and micro 50 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) mining 501 432 537 518 502 608 510 666 food industry 886 995 933 1223 1012 1489 1064 1817 light industry 77 317 83 373 87 428 94 497 woodworking , pulp and paper 240 464 242 574 260 664 264 817 petrochemical industry 525 570 569 742 591 913 697 1033 mineral industry 266 618 288 736 308 840 306 998 metallurgy 365 475 395 633 384 798 391 993 mechanical engineering 693 912 725 1099 749 1304 790 1622 industrial service 79 321 102 435 100 537 99 662 energy industry 631 2167 360 1927 850 1696 1296 1502 total 4263 7271 4234 8260 4843 9277 5511 1060 7 at stage 2, production functions were constructed for each enterprise belonging to the 10 industrial sectors. the data were retrieved from spark-interfax database and included the following details: company size, fixed assets, revenue, total costs and wages for 2013–2017. calculation of value added for each enterprise was carried out according to formula (3.5). to construct the production function, the authors used panel data, which contain statistical information about the same number of enterprises for several consecutive periods of time (2013–2017). the use of panel data allowed us to enhance the size of the sample under review, which provided greater efficiency in estimating the regression model’s parameters. to construct the production function, we applied the method of least squares. the obtained results were checked for reliability, unbiasedness and consistency through calculating the expected value, the durbin–watson statistic, constructing the correlation matrix and holding the white test. we obtained a total of 10 production functions for various groups of industries. it is worth noting that for different groups of industries the indicators’ coefficients vary significantly. nevertheless, we can notice that in all production functions, coefficients with cross-section indicators have a negative value, and in more than half of them – the coefficient with a logarithm of the labor factor. all the rest coefficients have a positive value. for example, for mechanical engineering enterprises, the production function is as follows: ln(𝑌) = 9.65 − 0,10 ln(𝐿) + 0,17 ln(𝐾) + 0.04(ln(𝐿))2 + 0.01(ln(𝐾))2 + 0.01𝑡2 − 0.03 ln(𝐿) ln(𝐾) − 0.01 ln(𝐿) 𝑡 + 0.01 ln(𝐾) 𝑡 (4.1) once the production functions are constructed, we proceed to stage 3: calculating total factor productivity. based on the coefficients of production functions, we calculate the elasticity of value added of labour and capital for each period through identifying partial derivatives (formulas (3.2) – (3.3)). the result is presented in table 4.2. evaluating comparative total factor productivity in the russian industry 51 copyright ©2019 assa. adv. in systems science and appl. (2019) table 4.2. elasticity of value added of labour and capital for industrial sectors for 2013–2017 year / industry m in in g f o o d in d u st ry l ig h t in d u st ry w o o d w o rk i n g , p u lp an d p ap er p et ro ch em i ca l in d u st ry m in er al in d u st ry m et al lu rg y m ec h an ic al en g in ee ri n g in d u st ri al se rv ic e e n er g y in d u st ry 2013 0,605 0,916 0,848 0,846 0,912 0,889 0,854 0,884 0,844 1,04 0,222 0,134 0,132 0,143 0,1 0,112 0,099 0,036 0,03 0,008 2014 0,51 0,902 0,84 0,83 0,899 0,863 0,829 0,869 0,819 1,015 0,31 0,135 0,128 0,151 0,099 0,111 0,1 0,041 0,041 0,019 2015 0,519 0,906 0,84 0,838 0,892 0,854 0,809 0,855 0,808 1,003 0,297 0,131 0,12 0,147 0,097 0,104 0,101 0,045 0,047 0,025 2016 0,522 0,904 0,852 0,839 0,89 0,836 0,798 0,845 0,794 0,998 0,294 0,13 0,108 0,148 0,095 0,099 0,101 0,048 0,053 0,028 2017 0,523 0,886 0,794 0,798 0,852 0,787 0,742 0,792 0,702 0,993 0,29 0,118 0,13 0,153 0,092 0,078 0,103 0,063 0,093 0,027 over the entire period, the greatest elasticity of value added of capital was typical of the industrial sectors, such as energy, food and petrochemical industries. for the mining industry, on the contrary, this indicator varies between 51–60%. most industries exhibit the same level of the elasticity of labor value added – 9–15%. however, in 2013, this indicator for the mechanical engineering and industrial service was quite low; but by 2017, there was a positive trend, i.e. an increase of 6% and 9% respectively. the lowest level of labor elasticity was observed in the energy sector. taking into account the elasticity values, we used formula (3.1) to calculate a rise in tfp for the industrial sector as a whole, as well as for large/medium-sized enterprises and small enterprises taken separately. tfp growth rate was analyzed using the indicator of simple average tfp growth rate. the method developed by bessonova [9] also implies the calculation of another indicator, i.e. weighted average tfp growth rate by the volume of value added. however, within the scope of the present research, the calculation of this indicator does not make sense. one of the study’s tasks is to compare tfp growth rates for large and small industrial enterprises. at the same time, the volume of value added of large and medium-sized enterprises will be greater, which means a strong preponderance of such businesses and, as a result, a distortion of the tfp indicator. for this reason, we deliberately exclude the calculation of this indicator from our research. dynamics of average tfp growth rates for 2013–2017 illustrates that their values for large and medium-sized enterprises significantly exceeds similar values for small enterprises (fig. 4.2). 52 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) note: based on the authors’ calculations. fig. 4.2. average tfp growth rates by size of enterprises in comparison with 2013, large and medium-sized enterprises’ tfp growth rates in 2014 was 4.4%, whereas small companies demonstrated a 1.4% drop in growth. in the post-crisis period (after 2014), small industrial enterprises experienced an increase of up to 1.2%, but it was still significantly lower than that of other enterprises. let us look at the dynamics of changes in tfp for the 10 industries. figure 4.3 shows the values of tfp average growth in 2014. note: based on the authors’ calculations. fig. 4.3. tfp average growth in 2014 the period under consideration is characterized by both the decline in the global economic situation and the worsening national institutional environment in russia. this results in the fact that more than 50% of the industries experience a negative tfp growth rate, regardless of the size of enterprises. however, for the mining, food, woodworking, pulp and paper industries and metallurgy, tfp growth of large and medium-sized industrial enterprises significantly exceeds that of small businesses. it is also noteworthy that tfp decline rate of large enterprises 0,044 -0,014 0,069 -0,023 -0,030 -0,020 -0,010 0,000 0,010 0,020 0,030 0,040 0,050 0,060 0,070 2017 2016 2015 2014 -8,0% -6,0% -4,0% -2,0% 0,0% 2,0% 4,0% 6,0% 8,0% large and medium-sized enterprises small enterprises mining food industry light industry woodworking, pulp and paper petrochemical industry mineral industry metallurgy mechanical engineering industrial service energy industry evaluating comparative total factor productivity in the russian industry 53 copyright ©2019 assa. adv. in systems science and appl. (2019) operating in industrial service and petrochemical industry are much lower than that of small ones. in general, in 2014, small enterprises of two industries only had a positive dynamic of tfp growth. note: based on the authors’ calculations. fig. 4.4. tfp average growth in 2015 figure 4.4 demonstrates that, compared with the previous period, tfp growth rate in 2015 changed dramatically. for all the industries, with the exception of energy production, tfp growth rate of large and medium-sized enterprises significantly exceeded that of small enterprises. the biggest gap – from 6 to 14% – was characteristic of mining, food, light and petrochemical industries. in 2014–2015, the strongest tfp growth was achieved by small enterprises engaged in woodworking, pulp and paper industry – 3.6%, petrochemical industry – 3.5% and light industry – 3.2%. the same indicators for large and medium-sized businesses equaled 6.1%, 9.1% and 8.6% respectively. note: based on the authors’ calculations. fig. 4.5. tfp average growth in 2016 -8,0% -6,0% -4,0% -2,0% 0,0% 2,0% 4,0% 6,0% 8,0% large and medium-sized enterprises small enterprises mining food industry light industry woodworking, pulp and paper petrochemical industry mineral industry metallurgy mechanical engineering industrial service energy industry -8,0% -6,0% -4,0% -2,0% 0,0% 2,0% 4,0% 6,0% 8,0% large and medium-sized enterprises small enterprises mining food industry light industry woodworking, pulp and paper petrochemical industry mineral industry metallurgy mechanical engineering industrial service energy industry 54 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) in 2016, the absolute growth of average factor productivity of large and medium-sized enterprises was typical of only 6 out of 10 industries under study (fig. 4.5): mining, woodworking, pulp and paper, mineral industry, mechanical engineering, industrial service and energy industry. moreover, in all sectors, excluding industrial service, the growth exceeded that of small and micro-enterprises. other industries experienced a decline in tfp, but only in light industry this decline was greater for small enterprises. the greatest tfp growth among small enterprises was observed in mining (2.6%), energy industry (2.4%) and industrial service (1.1%). note: based on the authors’ calculations. fig. 4.6. tfp average growth in 2017 in 2017, only 50% industries encountered an increase in tfp of large and medium-sized enterprises, which dominated over the same indicator for small businesses. only small enterprises operating in metallurgy demonstrated a positive dynamic of tfp growth (0.9%), in contrast to total factor productivity of large businesses (–2.2%). the research results indicate that small industrial enterprises are characterized by low tfp growth rates. the general trend for such businesses is a decline in total factor productivity. in 2017, the decrease in tfp for small enterprises in the petrochemical industry was 11.3%; food industry – 4.7%; woodworking, pulp and paper industry – 2.2. however, in 2016–2017, small enterprises of some industrial sectors exhibited a slight gain in tfp (mining – 2.7%; engineering – 1.7%; industrial services – 1.3%; metallurgy – 0.9%). most industrial sectors experienced an increase in total factor productivity of large and medium-sized enterprises. a sharp fall in tfp growth rates in 2017 was recorded only for large enterprises engaged in food and woodworking industries. having compared the data, we found that small enterprises’ tfp growth exceeded that of the rest of the enterprises only in 2016–2017 in metallurgy (small businesses’ tfp growth was 0.4 and 0.9% respectively). in all other cases, even if tfp growth rate of small businesses in absolute value is greater than that of large and medium-sized enterprises, the value of this indicator is negative, that is, only if there is no tfp growth per se. 5. conclusion the period under consideration is not long enough to project further tfp dynamics. however, it is possible to draw a number of important conclusions. -8,0% -6,0% -4,0% -2,0% 0,0% 2,0% 4,0% 6,0% 8,0% large and medium-sized enterprises small enterprises mining food industry light industry woodworking, pulp and paper petrochemical industry mineral industry metallurgy mechanical engineering industrial service energy industry evaluating comparative total factor productivity in the russian industry 55 copyright ©2019 assa. adv. in systems science and appl. (2019) firstly, the hypothesis that large enterprises develop faster than small ones was confirmed. there are several possible explanations for this phenomenon and each of them has to be checked and verified in further studies. on the one hand, there are objective reasons behind poor efficiency of small businesses. one of them is a vague possibility to get increasing returns due to economies of scale. since we only examine the industrial sector, this factor can be of considerable importance. in addition, industry in russia is vertically or horizontally integrated structures, where auxiliary production is outsourced. thus, small industrial enterprises in russia are poorly performing divisions of large businesses. on the other hand, there are subjective factors associated with small industrial business in russia lagging behind in terms of technology and innovation. it means that the technological factor of tfp growth is more attributed to the development of technical equipment, rather than technological innovation. large businesses possess more advanced and productive types of fixed assets, which supports the conclusions about the rental, extensive growth of the russian industry. secondly, the industrial sectors under review are characterized by a significant companies differentiation, which leads to a slowdown in the economic growth of the entire industrial production, and consequently, the processes of re-innovation of the russian economy. the results obtained are confirmed by other studies on total factor productivity of the russian industry. however, the previous research made conclusions “about a large group of inefficient enterprises that continue functioning in the market ...” [8, p. 31], but failed to reveal what kind of enterprises they were and what features they had. we suppose that the current paper lays the foundation for the study of such enterprises’ behavior. according to our study, small enterprises are among those companies forming the inefficient segment of the russian economy. yudin and cherkasov [39] state that “in russia, small business [...] does not fulfill the functions of diversifying production and introducing effective innovation processes. small enterprises develop primarily in the sphere of rapid capital turnover and are not involved in research and development”. the reason behind this is that flexibility and entrepreneurship skills are insignificant resources unable to provide competitive advantages in the russian institutional space. assessing institutional traps in the development of small business is the primary avenue for the authors’ further studies. references 1. aghion, p. & howitt, p. (1992) a model of growth through creative destruction, econometrica, 60(2), 323–351. 2. akberdina v.v. (2018) digitalization of industrial markets: regional characteristics, upravlenets – the manager, 6, 78–87. 3. almus m. & nerlinger e. a. (1999) growth of new technology-based firms: which factors matter? small business economics, 13(2), 141–154. 4. arrow k. (1962) the economic implications of learning by doing, review of economic studies, 29(3), 155–173. 5. audretsch d. b., carree m. a., van stel a. j. & thurik a. r. (2000) impeded industrial restructuring: the growth penality, research paper, and center for advanced small business economics. rotterdam: erasmus university. 6. beccetti l. & trovato g. (2002) the determinants of firm growth for small and medium sized firms. the role of the availability of external finance, small business economics, 19(4), 291–306. 7. bessonov v.a. (2004) o dinamike sovokupnoj faktornoj proizvoditel'nosti v rossijskoj perekhodnoj ehkonomike [on the dynamics of total factor productivity in the russian transition economy]. moscow: institut ehkonomiki perekhodnogo perioda. [in russian]. 56 s.v. orekhova, e.v. kislitsyn copyright ©2019 assa adv. in systems science and appl. (2019) 8. bessonova e.v. (2007) ocenka ehffektivnosti proizvodstva rossijskih promyshlennyh predpriyatij [assessment of production efficiency of russian industrial enterprises], prikladnaya ehkonometrika, 2(6), 13-35, [in russian]. 9. bessonova e. v. (2018) analiz dinamiki sovokupnoj proizvoditel'nosti faktorov na rossijskih predpriyatiyah (2009-2015 gg.) [analysis of the dynamics of total factor productivity at russian enterprises], voprosy ehkonomiki, 7, 96-118, [in russian]. 10. davidsson p., kirchoff b., hatemi a. & gustavsson h. (2002) empirical analysis of business growth factors using swedish data, journal of small business management, 40(4), 332–349. 11. djankov s., glaeser e., la porta r. & lopez-de-silanes f., shleifer a. (2003) the new comparative economics, journal of comparative economics, 31, 595–619. 12. domar e. (1946) capital expansion, rate of growth and employment, econometrica, 14(2), 137–147. 13. dunne p. & hughes a. (1994) age, size, growth and survival: u. k. companies in the 1980s, journal of industrial economics, 42(2), 115–140. 14. easterly w. & levine r. (2001) it`s not factor accumulation stylized facts and growth models, the world bank economic review, 15(2), 177-219. 15. fishman a., don-yehiya h. & schreiber a. (2018) too big to succeed or too big to fail? small business economics, 4, 811-822. 16. fuentes r. & morales m. (2011) the measurement of total factor productivity: a latent variable approach, macroeconomic dynamics, 15(02), 145-159. 17. gebremeskel h. gebremariam g. h., gebremedhin t. g. & jackson r.w. (2004) the role of small business in economic growth and poverty alleviation in west virginia: an empirical analysis, research paper, 2004-10. 18. grossman g. m. & helpman e. (1991) innovation and growth in the global economy. cambridge, ma: mit press. 19. hall r. & jones c. (1999) why do some countries produce so much more output per worker than others, the quarterly journal of economics, 114 (1), 83-116. 20. harrod r.f. (1939) an essay in dynamic theory, economic journal, 49 (march), 14–33. 21. hulten c. (2001) total factor productivity: a short biography. in charles r. hulten, edwin r. dean and michael j. harper (eds.), new developments in productivity analysis (pp. 1-54). university of chicago press. 22. klacek j., vošvrda m. & schlosser š. (2007) kle translog production function and total factor productivity, statistika, 4, 261–274. 23. oecd productivity manual (2001): a guide to the measurement of industrylevel and aggregate productivity growth. washington, d. c.: oecd. 24. orekhova s. (2017) economic growth quality of metallurgical industry in russia, journal of applied economic science, 5, 1377–1388. 25. prescott e. (1997) needed: a theory of total factor productivity. research stuff report, federal reserve bank of minneapolis. minneapolis. 26. rodric d., subramanian a. & trebbi f. (2004) institutions rule: the primacy of institutions over geography and integrations in economic development, journal of economic growth, 9, 131-165. 27. romer p.m. (1986) increasing returns and long-run growth, the journal of political economy, 94 (october), 1002–1037. 28. rothstein b. (2012) good governance. in david levi-faur (ed.). the oxford handbook of governance, 143–154. oxford, oxford university press. 29. smith a. (2008) an inquiry into the nature and causes of the wealth of nations. oxford: oxford university press. 30. solow r. m. (1956) a contribution to the theory of economic growth, quarterly journal of economic, 70, 65–94. https://www.nber.org/books/hult01-1 https://www.nber.org/books/hult01-1 evaluating comparative total factor productivity in the russian industry 57 copyright ©2019 assa. adv. in systems science and appl. (2019) 31. solow r. m. (1957) technical change and the aggregate production function, the review of economics and statistics, 39(3), 312–320. 32. sukharev o. s. (2018) strukturnyy analiz tekhnologicheskikh izmeneniy i strategiya ekonomicheskogo rosta [structural analysis of technological changes and the strategy of economic growth], izvestiya uralskogo gosudarstvennogo ekonomicheskogo universiteta – journal of the ural state university of economics, 3, 26−41, [in russian]. 33. swan t. w. (1956) economic growth and capital accumulation, economic record, 32 (november), 334–361. 34. timmer m. & voskoboynikov i. (2016) in mining fuelling long-run growth in russia? industry productivity growth trends in 1995-2012. in: d. jorgenson, k. fucao, m. timmer (eds.) the world economy: growth or stagnation? (pp. 281-318). cambridge: cambridge university press. 35. uzawa h. (1965) optimum technical change in an aggregative model of economic growth, international economic review, 6(1), 18–31. 36. wagner j. (1992) firm size, firm growth, and persistence of chance: testing gibrat’s law with establishment data from lower saxony, 1972–1989, small business economics, 4(2), 125–131. 37. wagner, m. & zidorn, w. (2017) effects of extent and diversity of alliancing on innovation: the moderating role of firm newness, small bus econ, 4, 919–936. 38. yang j.s. (2017) the governance environment and innovative smes, small business economics, 48, 525-541. 39. yudin n.s. & cherkasov v.a. (2016) analiz specifiki razvitiya malogo predprinimatel'stva v rossii v ramkah razrabotki metodologii ocenki social'noehkonomicheskoj ehffektivnosti malyh predpriyatij [analysis of the specifics of small business development in russia in the framework of the methodology for assessing the socioeconomic efficiency of small enterprises], social'no-ehkonomicheskie yavleniya i process, 3, 117-123, [in russian]. advances in systems science and application (2016) vol.16 no.3 76-93 optical mems sensor for measurement of low stress using ptolemy ii i. mala serene, rajasekhara babu m, zachariah.c.alex school of computing science and engineering, vit university, school of electronics engineering, vit university, vellore-632 014, tamil nadu, india abstract modeling and simulation plays vital role in the micro electro mechanical systems (mems) field. optical mems comprise of three domains namely optical, electrical and mechanical. the existing mems software for modeling is very expensive. this cost of modeling software increases the design and development of optical mems sensors. this paper proposes the design and development of a novel optical read out mechanism. this mechanism is used to measure the maximum stress applied on the cantilever and its corresponding deflection of the cantilever. the experiments have been carried out using ptolemy ii software for design and simulation of mems optical sensors. laser actor, a photo detector and force actor have been created using ptolemy ii. comsol software has been used to model cantilever. a comparative study has been done for cantilever with three modes of eigen frequencies using comsol. the experimental result shows that the parylene optical mems force sensor can sense less range of stress 0.0003 n/m to 0.272 n/m when compared to the polyimide optical mems sensor. keywords microcantilever, optical mems, ptolemy, sensor, laserdiode, comsol 1 introduction optical mems can be defined as micro devices with three functionalities like electrical, mechanical and optical at the same time and can be fabricated using batch processing techniques developed from microelectronic fabrication [1]. for integrated micro-systems composed of electrical, optical and mechanical components, the need to model large numbers of linear and non-linear components with sufficient accuracy to analyze cross-talk, noise and tolerance in an interactive environment leads to the requirement of an efficient yet accurate mixed-technology simulation technique[2]. stevan p. levitan et al reported a computer aided design tool for free-space optoelectronic systems and achieved system-level modeling[3]. the advantages of optical mems sensors over electrical sensors are high adaptability in harsh environments high temperature, chemical corrosion, strong electromagnetic interference and high-energy radiation exposure[4]. currently, no single cad tool completely models the complexity of these mixed tools to model, simulate, and analyze each stage of the design[4] . hence we have chosen advances in systems science and application (2016) vol.16 no.3 77 ptolemy as our framework for developing optical mems based sensors. ptolemy ii is a system level design environment that supports heterogeneous modeling and design of concurrent systems. for simulating optical mems devices it is essential to integrate tools with different models of computation to simulate the whole system [5]. the ptolemy ii software provides an infrastructure that allows designers explore and integrate the different models of computation [6]. it is system levels tool it. it does not provide the functionality for implementation-level simulation. but external tools based on different model of computation can be integrated into each domain and ptolemy ii can serve as semantic glue. in this present work, the simple component of mems, a microcantilever is used to sense the stress. it can be operated in two modes: static and dynamic mode. in static mode, the bending of microcantilever depends upon the force or stress on the cantilever. in dynamic mode, the resonant frequency of microcantilever changes when the mass added to it. the different read out mechanisms of the microcantilever are optical readout, piezoelectric and piezoresistive [7]. many researchers reported that the microcantilever is made of materials like silicon, silicon nitride and polysilicon [8-9]. but the fabrication cost of the silicon based cantilevers is expensive. so silicon can be replaced by a polymer which offers a shining future for the development of chemical and biological sensors. the merits of the polymer microcantilever over silicon microcantilever are low cost, more flexibility, transparency to visible uv, easily mouldable capability, improved bio-compatibility[10]. in this paper, polyimide and parylene are identified as suitable polymers for microcantilevers given their low youngs modulus, high planarity, chemical resistance and biocompatibility [11]. 2 expermental and simulation laser source emits the light of wavelength (λ=850nanaometers).this laser beam is then passing through the two optical fibers separated apart axially. the cantilever structure is fixed at one end and free at other end. a slit is connected at the free end of the cantilever moves between the two optical fibers when force is applied. the deflection of the beam will be in y direction and by virtue of this deflection the output power detected at one of the fiber ends is varied continuously from maximum to minimum though the slit arrangement as shown below. this output power variation can be calibrated according to change in minute force variation over the cantilever which in turn will constitute an accurate optical mems sensor. the light coming out of the second optical fiber is detected by the photo detector. 78 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... fig. 1 flowchart of individual particle update 3 details of software development in the present investigation, we have developed software codes for various actors that make an optical mems sensor in a software platform called ptolemy. the various actors are laser, a photo detector and a force actor. the individual figures of the various actors like laser actor, force actor and photodiode actor are given in the fig 2 (a)-(c) below: fig. 2 (a) laser actor (b) force actor and (c) photodiode actor 3.1 laser actor the abbreviation of laser is light amplification by stimulated emission of radiation. laser operates on the principle called stimulated emission. it was postulated by albert einstein before 1920. this is a semiconductor laser diode (gaas) which emits light when we apply a forward biased across the p-n junction. the laser diode actor is modelled using the mathematical equations which include internal power of the laser, external power of the laser and reverse leakage current of the diode. the external power of the laser diode is given by po = pint n(n+ 1)2 (1) where n, pint, po is refractive index of the gaas, internal power and external power of the laser. advances in systems science and application (2016) vol.16 no.3 79 3.2 force actor the force actor made of a cantilever beam and two optical fibers. the deflection of the cantilever is modelled using stoneys equation, spring constant and the three modes of the resonant frequency. the optical fiber actor is created using the power output detected at the second fiber and the loss of light due to the force applied on the cantilever. 3.2.1 cantilever beam micro cantilever is a widely used component in micro electro mechanical system devices [12]. cantilever is a type of beam fixed at one end and suspended freely at the other end and the beam is originally straight. the equation (2) is the stoneys formula [13], which relates cantilever end deflection δ to applied stress σ: δ = 3σ(1− ν) e ( l t )2 (2) where δ,σ,l,t,e,ν are deflection, stress, length of the cantilever beam, youngs modulus, poissons ratio. the spring constant (k) of the cantilever beam is given by k = ewt3 4l3 (3) where e, w, t and l are the youngs modulus, width , thickness and length of the cantilever beam. the frequency at which a cantilever tends to oscillate in the absence of any force is the eigen frequency .the eigen frequency of a cantilever beam [14] can be find out from the optimized cantilever geometry for the l and t and density , for the two sensors is given by f = αn t l2 √ e ρ (4) αn = 1 4π √ ε λ2 n (5) where λn=1.8751, 4.6941, 7.8547 ...... 3.2.2 optical fibre we have designed two fibres with core diameter 2a= 175m coupled longitudinally such that the free end of the cantilever will move the slit vertically down between the fiber ends as force is applied on it. as a result the light coupled from fiber1 to fiber2 decreases gradually as the amount of force increases. there are two formulas used for calculating loss and power detected at the second at the second fiber is given below. pout = pin [ 1− ( w 2a )] (6) 80 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... loss = 20log10a (7) where a stands for ratio of pout/pin. loss is a function f(w), where w is the cantilever deflection, which is numerically equal to w/2a, where 2a is the fiber diameter. 3.3 detector actor a photodiode is a semiconductor device, with a p-n junction and an intrinsic layer between p and n layers. the photo detector used is a reverse biased photodiode (pd) which converts the input optical power into the photo current (ip). the following formulas are applied to create a detector actor: the photocurrent is given by i = rpout (8) the responsivity measure the electrical output per optical input of the photodiode is given by r = ηqλ hc (9) where η, q, h, c, λ are internal quantum efficiency, charge of electron, plancks constant, velocity of light in vacuum, wavelength of light. using the above actors, the optical mems sensor model are created in the ptolemy framework as shown in the fig 3 and fig 4. in the present work, two force actors were created using the same geometrical parameters but the cantilever beam is made of different polymer materials like polyimide and parylene. the maximum stress sensed by the cantilever is measured for two different materials of the cantilever beam. the material properties of the cantilever beam include the youngs modulus (e), poisson ratio (ν) and density (ρ) is given in the table 1 : table 1 material properties of the cantilever beam material properties polyimide parylene youngs modulus (gpa) 3.2 2.8 poison ratio 0.42 0.4 density (kg/m3) 1300 1289 advances in systems science and application (2016) vol.16 no.3 81 fig. 3 model of the optical mems sensor(polyimide material) using ptolemy ii fig. 4 model of the optical mems sensor (parylene material) using ptolemy ii 3.4 model the mems cantilever beam using comsol comsol multiphysics version 5.0, a commercial fem tool for mems was used to develop a finite element model [15] of the polymer cantilevers. in the present work, the cantilever beam modelled using cost effective open source ptolemy software and its eigen frequency of the first three modes are compared with the rectangular beam of two different materials polyimide and parylene using the using comsol software. the free tetrahedral meshing is applied. 82 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... fig. 5 5(a) model of the rectangular cantilever beam using comsol, fig 5(b) mesh model of the cantilever beam 4 results and discussion 4.1 optical mems sensor using ptolemy in the optical mems sensor model, the laser diode is modelled the fig.6 (a) and fig 7(a) represents the output power of the laser and output current of laser diode. the output power increases linearly with the applied current, when the applied current is larger than the threshold current. when the force applied on the cantilever, the cantilever bends and light passing from the optical fiber 1 to optical 2 is blocked based on the amount of force applied. the range of the force applied and the deflection of the cantilever is recorded for the two optical sensors are tabulated in table 2. in fig.6 (b)-(d) and fig. 7(b)-(d), the sample of the force applied in the cantilever and corresponding deflection of the cantilever is recorded, then the deflected laser power is converted into current by the photodiode and plotted in the graph. 4.2 comsol cantilever beam result the results of the first three modes of the cantilever beam of two materials modeled using comsol software are shown in the fig 8(a)-(f). the analytical values of the eigen frequencies are compared with eigen frequencies of the two cantilevers modelled using comsol are tabulated in the table 3 and the same is represented using bar chart is shown in fig 9(a)-(b). 5 optimization of the geometrical parameters the different lengths (200 µm, 300 µm, 400 µm, 450 µm, 500 µm) of the two different materials of the cantilever are kept constant and the thickness of the cantilever is varied from 0.5 µm to 3.0 µm. for each length and the maximum stress/force is recorded for each simulation is shown in table 4 and table 6 and the results are plotted is shown in figure 10. (a)-(f). for different thickness (t=0.5 µm, 1.0 µm,1.5 µm,2.0 µm,2.5 µm and 3.0 µm), the length is varied from 200 µm to 500 µm for each thickness and the maximum stress/force is recorded for each advances in systems science and application (2016) vol.16 no.3 83 fig. 6 simulation result of the polyimide optical mems force sensor (a) output power of laser diode (b), (c) and (d) deflected laser power vs output current at the photodetector fig. 7 simulation result of the parylene optical mems sensor1 (a) output power of laser diode (b), (c) and (d) deflected laser power vs output current at the photo detector 84 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... table 2 stress/force applied vs. cantilever deflection of the two optical mems sensor using ptolemy polyimide optical mems sensor parylene optical mems sensor stress/force applied cantilever stress/force applied cantilever (n/m) deflection (n/m) deflection 0.0005 2.72e-07 0.0003 1.93e-07 0.05 2.72e-05 0.05 3.21e-05 0.2 1.09e-04 0.1 6.43e-05 0.321 1.75e-04 0.272 1.75e-04 0.322 1.75e-04 0.273 1.76e-04 fig. 8 eigen frequency of the result of the rectangular cantilever beam1 (a)-(c) and (d)-(f)cantilever beam2 using comsol table 3 comparison of eigen frequency values usingptolemy andcomsol software modes of polyimide cantilever parylene cantilever eigen frequency ptolemy(khz) comsol(khz) ptolemy(khz) comsol(khz) 1 0.5071 0.5074 0.4764 0.4678 2 3.178 3.2392 2.986 3.037 3 8.899 9.1789 8.3597 8.602 advances in systems science and application (2016) vol.16 no.3 85 fig. 9 (a-b): comparison of the modes vs. eigen frequency of the polyimide and parylene cantilever beam using ptolemy and comsol table 4 different length of the cantilever beam vs. maximum stress/force applied at constant thickness of the polyimide optical mems force sensor polyimide optical mems sensor cantilever max. stress/force applied (n/m) thickness (µm) l=200 µm l=300 µm l=400 µm l=450 µm l=500 µm 0.5 2.01 0.88 0.502 0.397 0.321 1 8.045 3.57 2.011 1.588 1.287 1.5 18.1 8.04 4.52 3.575 2.89 2 32.1 14.29 8.04 6.35 5.14 2.5 50 22.3 12.56 9.93 8.04 3 72.3 32.17 18.08 14.3 11.58 table 5 different thickness of the cantilever beam vs. maximum stress/force applied at constant length of the polyimide optical mems sensor polyimide optical mems sensor cantilever max. stress/force applied (n/m) length (µm) t=0.5µm t=1.0µm t=1.5µm t=2.0µm t=2.5µm t=3.0µm 200µm 2.01 8.045 18.1 32.1 50 72.3 300µm 0.88 3.57 8.04 14.29 22.3 32.17 400µm 0.502 2.011 4.52 8.04 12.56 18.08 450µm 0.397 1.588 3.575 6.35 9.93 8.04 500µm 0.321 1.287 2.89 5.14 8.04 11.58 86 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... fig. 10 (a)-(f) different length of the cantilever beam vs maximum stress/force applied of the polyimide optical mems sensor at constant thickness advances in systems science and application (2016) vol.16 no.3 87 fig. 11 (a)-(f) different length of the cantilever beam vs maximum stress/force applied of the polyimide optical mems sensor at constant thickness 88 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... fig. 12 (a)-(f) different thickness of the cantilever beam vs. maximum stress/force applied of the polyimide optical mems sensor at constant length advances in systems science and application (2016) vol.16 no.3 89 fig. 13 (a)-(f) different length of the cantilever beam vs. maximum stress/force applied at constant thickness of the parylene optical mems sensor 90 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... fig. 14 (a)-(e) different thickness of the cantilever beam vs. maximum stress/force applied at constant length of the parylene optical mems force sensor advances in systems science and application (2016) vol.16 no.3 91 table 6 different length of the cantilever beam vs. maximum stress / force applied at constant thickness of the parylene optical mems sensor parylene optical mems sensor cantilever max. stress/force applied (n/m) thickness (µm) l=200µm l=300µm l=400µm l=450µm l=500µm 0.5 1.7 0.755 0.425 0.336 0.272 1 6.78 3.02 1.7 1.342 1.088 1.5 15.3 6.8 3.824 3.02 2.448 2 27.2 12.07 6.8 5.37 4.35 2.5 42.5 18.84 10.63 8.39 6.8 3 61.1 27.16 15.3 12.09 9.79 table 7 different thickness of the cantilever beam vs. maximum stress/force applied at constant length of the parylene optical mems sensor parylene optical mems sensor cantilever max. stress/force applied (n/m) length (µm) t=0.5µm t=1.0µm t=1.5µm t=2.0µm t=2.5µm t=3.0µm 200 µm 1.7 6.78 15.3 27.2 42.5 61.1 300 µm 0.755 3.02 6.8 12.07 18.84 27.16 400 µm 0.425 1.7 3.824 6.8 10.63 15.3 450 µm 0.336 1.342 3.02 5.37 8.39 12.09 500 µm 0.272 1.088 2.448 4.35 6.8 9.79 simulation is shown in table 5 and table 7 and the results are plotted is shown in figure 11. (a)-(f). from the recorded values, low stress/force is achieved at the length 500 µm and the thickness is 0.5m for both the sensors. 6 conclusion different actors like laser actor, force actor and photodetector have been developed and added in ptolemy framework. the physical functioning of each component of the optical mems force sensor device has been simulated using these actors. the results have been presented. the two optical mems force sensor are simulated in ptolemy ii. the parylene optical mems sensor can sense the low stress/force in the range of 0.0003 n/m to 0.272 n/m with the present fiber optic setup. the eigen frequency of the two cantilever beam is modelled using open source ptolemy ii and the three modes of eigen frequencies are compared with 92 i. mala serene, rajasekhara babu m and zachariah.c.alex:optical mems sensor for ... the results of the comsol. the low stress/force is measured for the optimized length (500 µm) and the thickness (0.5 µm) is found by varying the thickness and length of the cantilever using ptolemy. 7 acknowledgement we are thankful for support by the npmass mems design centre, sense, vit university for doing this project. references [1] e. ollier(2002), “optical mems devices based on moving waveguides”, ieee j.sel. topics quantum elect, vol. 8, no. 1, pp. 155c62. [2] selvarajan, a and pattnaik, prasant kumar and badrinarayana, t and srinivas, t (2006), “a comparative study of moem pressure sensors using mzi, dc, and racetrack resonator io structures”, in proceedings spie international conference on smart structures and materials 2006:smart electronics, mems, biomems, and nanotechnology, pages vol.61726, 61721a, san diego, california, usa. [3] levitan, s.p., kurzweg, t.p., marchand, p.j., rempel, m.a., chiarulli, d.m., martinez, j.a, bridgen, j.m.,fan, c. and mccormick, f.b.(1998), “ chatoyant: a computer-aided design tool for free-space optoelectronic systems”, applied optics,vol. 37, no. 26, pp. 6078-6092. [4] t.p. kurzweg, j.a. martinez, s.p. levitan, p.j. marchand and d.m. chiarulli(2001), “dynamic simulation of optical mem switches”, optics in computing (oc’01),lake tahoe, nv, pp. 35-37. [5] j. davis, etc.(1999), “heterogeneous concurrent modeling and design in java”, ucb/erl memorandum m99/ 40, dept. of eecs, university of california, berkeley, ca 94720. [6] j. liu , b. wu, x. liu and e. a. lee(2001), “interoperation of heterogeneous cad tools in ptolemy ii”, department of electrical engineering and computer science, university of california, berkeley, ca 94720 u.s.a.journal of modeling and simulation of microsystems,vol. 2, no. 1, pp. 1-10. [7] p.sangeetha and a.vimala juliet, “mems cantilever based immunosensors for biomolecular recognition ”, international journal of computer technology and electronics engineering (ijctee), vol. 2, no. 1. advances in systems science and application (2016) vol.16 no.3 93 [8] m. maute, s. raible, f.e. prins, d.p. kern, h. ulmer, u.weimar and w. gopel(1999), “detection of volatile organic compounds vocs/ with polymer-coated cantilevers”, sensors and actuators , pp.505c511. [9] denis spitzer, thomas cottineau, nelly piazzon, sbastien josset, fabien schnell, sergey nikolayevich pronkin, elena romanovna savinova and valerie keller(2012), “research on the sustainable utilization of water resources of jinan city”, bio-inspired nanostructured sensor for the detection of ultralow concentrations of explosives,angew. chem. int. ed. pp.5334c5338. [10] b g sheeparamatti, m s hebbal, r b sheeparamatti, v b math and j s kadadevaramath(2006), ‘simulation of biosensor using fem”, journal of physics: conference series, vol.34, pp.241c246. [11] richardson r, j.a. miller and w.m. reichert(1993), biomaterials., vol.(14), pp.627c635. [12] subhashini. s and vimala juliet. a(2013), “resonance based micromechanical cantilever for gas sensing”, international journal of network security and its applications (ijnsa), vol.44, no.w12420. [13] stoney g(1909), the tension of metallic films deposited by electrolysis, proc r soc lond ser a 1909; 82: 172-5.. stoney, g. gerald. “the tension of metallic films deposited by electrolysis”, proceedings of the royal society of london,pp.40-43. [14] suryansh arora, sumati, arti arora and p.j.george (2012), “ design of mems based microcantilever using comsol multiphysics”, international journal of applied engineering research,vol.7 no.11. pp.1-3. [15] g. louarn and s. cuenot(2008), “ finite element modelling of microcantilevers used as chemical sensors”, recent advances in modelling and simulation,pp.207-222. corresponding author rajasekhara babu m can be contacted at: rajababu.m1@gmail.com adv syst sci appl 2019; 04; 14-24 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/754 distributed data gathering system to analyze natural gas composition ivan brokarev1*, sergei vaskovskii2 1) national university of oil and gas «gubkin university», moscow, russia e-mail: brokarev.i@gubkin.ru 2) v. a. trapeznikov institute of control sciences of russian academy of sciences, moscow, russia e-mail: v63v@yandex.ru received june 07, 2019; revised september 30, 2019; published december 31, 2019 abstract: problems of development of gas analysis system based on computing fabrics have been studied. the structure of the data gathering system that will be used for the natural gas composition analysis is proposed. the most significant features and the main advantages of the proposed system are described. the most appropriate statistical models that can be used in solving the task of natural gas composition analysis are presented. the consecutive stages for statistical model development and correlation analysis results are shown. algorithms have been developed for effective solution for a number of gas analysis tasks for distributed systems. the mathematical basis of the learning algorithm and the architecture of the proposed neural network model are described. the different training completion cases for the developed neural network model are shown. recommendations are made for effective application of neural network tools as the main models for solving the task of the natural gas composition analysis. opportunities for further development of the proposed system are considered. keywords: distributed systems, data gathering system, composition analysis, neural network analysis. 1. introduction currently a variety of different gas analysis systems are under development [2]. distributed small size computerized systems are widely used to develop automatic process control systems including gas analysis systems. real time systems are the most effective systems to analyze gas. the modcon systems wireless gas analysis system is one of them; the system monitors the technological process online [7]. however, in spite of the past gains in this field, it is time-consuming to develop distributed real-time control and gas analysis systems. existing systems have a number of disadvantages: it is necessary to use expensive specialized equipment, the system is oriented mainly to monitor the technological process, the system detects limited amount of gas components. so, it is worthwhile to develop distributed gas analysis systems to overcome these shortcomings. it is one of the most complicated task to develop a distributed hardware and software for distributed microcomputerized systems. the main task solved by a distributed data gathering system is to provide access to a large amount of data in real time. this paper describes difficulties the engineers face while developing a distributed data gathering system to analyze natural gas composition; we take into account the specific features of such systems and how to apply neural network tools to investigate the features. * corresponding author: ivan brokarev (brokarev.i@gubkin.ru) mailto:v63v@yandex.ru distributed data gathering system to analyze natural gas composition 15 copyright ©2019 assa. adv. in systems science and appl. (2019) 2. development of the distributed data gathering system for natural gas composition let us see how to develop a distributed data gathering system to analyze natural gas composition. we want the system to be automated and based on the proposed data gathering system. the automated control system can be both information-logical system to monitor the technological process of gas transport and information-computing system to calculate a number of qualitative natural gas parameters, for example, the gas calorific value. we suggest to use a multiplex data gathering system for the task. in particular, multiplex systems, that utilize individual signal processing techniques at each measuring channel, are currently the most widespread systems. the most important features of distributed systems are the following: • transparency of the system – ability to hide physical distribution of resources, errors concerning access to resources and operation itself, duplication of resources, difficulties concerning concurrent operation of several users with a single resource; • scalability of the system – dependence of system characteristics on the number of its users and involved resources; • system security – data integrity protection, proper fall-over protection and error recovery capability; • openness of the system – availability of full description of interfaces and services of system operation; • system real-time operation – ability to respond to unpredictable flow of external events within predictable time. it should be noted that such features are inherent for the majority of distributed systems that can be applied to analyze gas. the main differences of the proposed system from the existing systems are commercially available and relatively inexpensive sensors as hardware components. in addition, our system is able to compute in real-time a number of natural gas energy characteristics, including its calorific value. finally, our system can detect main natural gas components (for example, eight components should be detected for russian natural gas [1]). all these can be detected and computed in real-time. when developing the data gathering system the process is considered to be carried out in several stages. the main stage is dedicated to independent development and debugging of interconnected components of applied software and hardware. the software of the designed data gathering system consists of the following main components: • programming of the low level of data gathering (sensor level); • interaction of the low and high levels (sensor level and data analysis level); • data analysis after monitoring, data documenting and data archiving. the structure of the proposed distributed system is shown in fig. 2.1. two measurement chambers with connected sensors are located at the first level of our system. two chambers are necessary to verify the collected measurement data. an important subsystem of our data gathering system is the sensor system. we suggest to use commercially available and relatively inexpensive sensors to measure physical parameters of gas mixtures – speed of sound [4], thermal conductivity [11], and molar fraction of carbon dioxide [3]. our data collecting system is located at the second level of the proposed system. the main functions of that subsystem are to scan for the information and transit the information to the higher level. the system processes the gathered data at the third level. to process the data is to eliminate errors from the measuring results, to average the collected measurement data, to compute the accessory parameters and visualize the collected data. the algorithms to analyze and evaluate the collected data are implemented at the fourth level of the system. these algorithms will be described later. at that level we compute how accurate the designed neural network is. this accuracy output shows if our model is adequate. 16 i. brokarev, s. vaskovskii copyright ©2019 assa adv. in systems science and appl. (2019) fig. 2.1. structure of the proposed distributed system 3. neural network tools for natural gas composition analysis currently, artificial intelligence solves a variety of applied problems. one of the important areas of artificial intelligence are statistical models. among a large variety of these models, we distinguish artificial neural networks (ann) mathematical models built on the principles of how biological neural networks are organized and function. ann are trained to detect complex dependencies between input and output. this ability is one of the main advantages of neural networks over traditional algorithms. one of the promising tasks, where anns are applied, is to analyze component composition of natural gas. if gases from different fields and sources are mixed during transportation and storage, the natural gas composition changes from that on the operation sites. this makes it harder to account the natural gas. since gas composition determines gas quality indicators for the equipment to operate properly and optimally, we need to measure or calculate certain properties of natural gas to transport the natural gas via a pipeline. this makes it very important to determine component composition of gas in real time; to solve this problem, correlation methods are currently being developed to analyze gas quality [2, 5]. statistical models, in particular ann, are often used to determine the desired properties and composition of natural gas by measuring the physical parameters of the gas. the fourth level of the shown distributed system carries out algorithms for determination of the target composition by measuring input physical parameters of natural gas. a statistical method is proposed as the main method for determination of the target composition. simplicity of mathematical calculations is the main advantage of the statistical method over distributed data gathering system to analyze natural gas composition 17 copyright ©2019 assa. adv. in systems science and appl. (2019) traditional algorithms. the main disadvantage the need for a large amount of initial data is eliminated by conducting the measurements via two channels. the following models can be used as statistical models for solving the task under discussion: • multi-parameter linear regression can serve as a reference model. its results can be used to compare accuracy of other regression models. for this model, it is necessary to calculate each component of the gas mixture separately. this is inefficient in comparison to other models; • model based on the support vector machine (svm). this is a set of similar supervised machine learning algorithms used to classify and regress analysis tasks. the main idea of the method is to translate source vectors into a higher dimension space and to search separating hyperplanes with a maximum gap in this space. we don't choose this model because of its significant shortcomings, in particular, complex interpretation of model parameters; also the model is applicable only to solve problems with two classes; • gaussian process regression (gpr); • ridge regression; • neural network model. on the basis of the previous research [5], it can be concluded that neural network is most effective as the main statistical model in the discussed task because of its learning capability and scalability in comparison with other models. to develop a model to analyze gas mixture composition, we implement these consecutive stages: • select data to train the model; • choose model architecture; • select a model training method; • assess accuracy of the model. the first stage consists in selecting input and output data for the model. we select a physical parameter as input, if there is a relationship between this gas parameter and the gas composition, as well as if the parameter can be measured with commercially available and relatively inexpensive measurement devices [3,4,11]. to determine if the parameters and the composition of the natural gas are related, a correlation analysis is carried out to verify inter dependencies within the sets of the input and output parameters and if the parameters are connected with each other. all data is cross-validated before it is used to train the model. cross-validation is a method to evaluate an analytical model and its behavior on independent data. this procedure is as follows: available data is divided into k parts, then the model is trained on the k 1 parts of the data, and the rest of the data is used to test the model. the procedure is repeated k times; each of the k pieces of the data is used to test the model. as a result, we assess effectiveness of the model if the available data is used uniformly. for this task, the sample of n = 700000 elements was divided into k = 10 parts. moreover, the training sample was normalized before training to improve results of prediction of the designed neural network model. the ranges of gas mixture components for the training samples are shown in table 3.1. table 3.1. ranges of components molar fractions for training sample component molar fraction, % methane 70 – 100 ethane 0 – 10 propane 0 – 5 carbon dioxide 0 – 10 nitrogen 0 – 10 18 i. brokarev, s. vaskovskii copyright ©2019 assa adv. in systems science and appl. (2019) the aim is to eliminate possible multicollinearity of the parameters linear interrelation of two or several variables. multicollinearity can lead to undesirable consequences, since the parameter estimates become unreliable. this means that the standard error increases, and it becomes impossible to isolate how the factors influence the effective indicators. we use correlation analysis (see table 3.2) with carbon dioxide molar fraction, sound velocity, and thermal conductivity coefficient as the input parameters. the output parameters are molar fractions of the gas mixture components. table 3.2. the correlation analysis results s p ee d o f so u n d , m /s t h er m al c o n d u ct iv it y , m w /( m * k ) m et h an e m o la r fr ac ti o n , % n it ro g en m o la r fr ac ti o n , % c ar b o n d io x id e m o la r fr ac ti o n , % e th an e m o la r fr ac ti o n , % p ro p an e m o la r fr ac ti o n , % speed of sound, m/s 1 0,99 0,93 -0,24 -0,61 -0,33 -0,67 thermal conductivity, mw/(m*k) 0,99 1 0,91 -0,15 -0,53 -0,45 -0,71 methane molar fraction, % 0,93 0,91 1 -0,52 -0,52 -0,48 -0,49 nitrogen molar fraction, % -0,24 -0,15 -0,52 1 0 0 0 carbon dioxide molar fraction, % -0,61 -0,53 -0,52 0 1 0 0 ethane molar fraction, % -0,33 -0,45 -0,48 0 0 1 0 propane molar fraction, % -0,68 -0,75 -0,49 0 0 0 1 the results of correlation analysis are shown on fig. 3.2 for parameters that have not been chosen as input parameters because of low correlation (speed of sound is shown as reference parameter for comparison). the results of correlation analysis are shown on fig. 3.3 for parameters that have been chosen as input parameters because of high correlation (methane concentration is shown as reference parameter for comparison). distributed data gathering system to analyze natural gas composition 19 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 3.2. the correlation analysis results for unchosen as input parameters fig. 3.3. the correlation analysis results for chosen as input parameters this neural network model is proposed as the main statistical model in the task [9]. to solve the problem, a three-layer network (multilayer perceptron) was chosen, which includes two neurons in the input layer, where two is the number of components of the input vector, eleven neurons in the second layer and four neurons in the output layer, where four is the number of components of the output vector. the input layer is a vector x = [x0 x1 x2]. the neurons of each next layer are connected to the neurons of the previous layer, and each input signal has a certain weight. the weight is the same for all input neurons since all input variables are equally important. each neuron has an activation function; the function argument is the input signal of the neuron. for the neurons of the hidden layer, the activation function is sigmoidal in the form of a hyperbolic tangent. for the neurons of the output layer, the activation function is linear. for training the network the levenberg-marquardt algorithm was chosen. this algorithm optimizes parameters of nonlinear regression models. note that in this algorithm, the optimization criterion is the root-mean-square error of the model on a training sample. the basic idea of the algorithm is as follows: to minimize the error locally, the initial values of the parameters are approximated. let a regression sample be a set of pairs of an independent variable x and a dependent variable y, and the regression model is a continuously differentiable function f. it is 20 i. brokarev, s. vaskovskii copyright ©2019 assa adv. in systems science and appl. (2019) necessary to find the value of the parameter vector w, where the error function fε reaches its local minimum: 1 ( ( , ) n i i i f = y f x w   (3.1) on the first iteration of the algorithm, the initial vector of parameters w0 is specified. on each following iteration, the vector is replaced by the vector w0 + δw. to estimate the increment δw, the following approximation of f is used: 0 0( , ) ( , ) *f w w x f w x j w     (3.2) where j is the jacobian of f. the increment δw at the minimum of fε is zero. so, to find the subsequent value of the increment δw, it is necessary to set the vector of partial derivatives of fε over w to zero. then we differentiate this expression over w and set the partial derivative to zero. after all transformations, we get the expression for δw: 1( * ) * *( ( ))t tw j j j y f w   (3.3) for this algorithm, the condition number of the matrix is important; the number shows how close the matrix is to the partial rank matrix (for square matrices it shows how close the matrix is to degeneracy). since the condition number of the matrix jt*j is equal to the squared condition number of the matrix j, the matrix jt*j may turn out to be degenerate. for this reason, in this algorithm we introduced a regularization parameter λ; the parameter is greater than or equal to zero. this parameter is selected on each iteration of the algorithm. given the regularization parameter, expression for δw takes the following form: 1( * * ) * *( ( ))t tw j j e j y f w     (3.4) where e is a unity matrix. there is a modification of this method, where the regularization parameter is multiplied by the matrix d, a diagonal matrix with the elements that coincide with the diagonal elements of the matrix jt*j. this approach is used to reduce the effect of the regularization parameter on the δw value. it is necessary, however, to note that this algorithm converges slower if the step is constant, which is a disadvantage. the problem is solved by introducing a coefficient k; the coefficient determines the step length and makes the method converge faster. given the two amendments described, the expression for δw takes the following form: 1*( * * ) * *( ( ))t tw k j j d j y f w     (3.5) the value of the vector w at the last iteration of the algorithm is the target. it is reached either if the calculated increment δw is less than the specified value, or if the error function fε is less than the specified value for the vector w. the architecture of the neural network model is shown in fig. 3.4. the number of neurons in the input layer is n = 3 for the case when the vector of input parameters contains carbon dioxide molar fraction, sound velocity, and thermal conductivity. the number of neurons in the hidden layer of our model is 11; to get this number, many different models were analyzed. the number of neurons in the output layer is m = 4 for the case of a five-component gas mixture. distributed data gathering system to analyze natural gas composition 21 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 3.4. the architecture of the proposed neural network model the next step is to train the network on the previously obtained data using the levenbergmarquardt algorithm. before that it is necessary to set the stopping criteria; for example, the criteria can be the maximum deviation value, when training is considered complete. in our case, to estimate the value, the sum of the squared deviations of the network outputs from the true values is calculated. also, when training the model, the maximum number of training cycles may be selected as the stopping criteria. however, as can be seen from fig. 3.5, this choice makes the network longer to train. fig. 3.5. cases of training completion for the developed neural network model after successful training of the network, the network is tested on the data, that were not in the training set. then, to assess the model accuracy, the following accuracy parameters are calculated: 22 i. brokarev, s. vaskovskii copyright ©2019 assa adv. in systems science and appl. (2019) • maximum absolute error; rgmax[| |]out ta ety y y   (3.6) where yout are the values of the neural network output, ytarget are target parameter values. • mean absolute error; rg[| |]out ta ety avg y y   (3.7) • maximum relative error (in %); rg rg | | max[ ] out ta et y ta et y y y    (3.8) • average relative error (in %); rg rg | | [ ] out ta et y ta et y y avg y    (3.9) • standard deviation; 2 rg 1 ( ) n out ta et i y y mse n     (3.10) • determination coefficient is a parameter that shows how the investigated model corresponds to the data. if the determination coefficient equals 1, the data is exactly described by the model. 2 rg 2 1 21 1 ( ) 1 ( ) n out ta et i n outn i out i y y r y y n           (3.11) according to the listed parameters, we conclude that our model can be used to solve the problem. if the model has low accuracy, it is trained again, or a different model architecture is used. 4. application of the neural network model according to the results obtained in this study, the neural network model can be used to calculate the component composition of natural gas, where the composition is measured by physical parameters. to check how accurate the method is, it was tested on a data sample. this sample included data that was not used for training the neural network model. the test sample includes 2218 gas mixtures, the ranges of its components are shown in table 4.3. to calculate the physical chemical parameters of the simulated sample, we used nist refprop software [8]. to calculate the speed of sound, the gerg-2008 gas state equation was used at standard temperature and pressure conditions. to calculate the thermal conductivity of the gas mixture, models for individual components and extended models for the corresponding states were used; the models are implemented in nist refprop. table 4.3. ranges of components molar fractions for test sample component molar fraction, distributed data gathering system to analyze natural gas composition 23 copyright ©2019 assa. adv. in systems science and appl. (2019) % methane 80,5 – 99,5 ethane 0 – 4,5 propane 0 – 2 carbon dioxide 0 – 9,5 nitrogen 0 – 9,5 table 4.4 and fig. 4.6 show how accurately the model predicts the natural gas composition of the test sample. carbon dioxide deviation was set to zero, because the content of this component in the gas sample is known since the non-dispersive absorption of infrared radiation (ndir) makes it easy to measure carbon dioxide molar fraction. calculation of the parameters, as well as constructing, training, and testing the neural network model was implemented by means of the matlab 2018a package with the deep learning toolbox [6]. table 4.4. accuracy of the test sample composition prediction using neural network model component maximum absolute error, % mean absolute error, % standard deviation determination coefficient methane 0,77 0,33 0,35 0,9999 ethane 0,38 0,16 0,16 0,9999 propane 0,27 0,12 0,12 0,9999 nitrogen 0,23 0,09 0,11 0,9999 carbon dioxide 0 0 0 1 fig. 4.6. neural network prediction deviation (in absolute scale) for the test sample 5. conclusion the proposed conception of distributed data gathering system for natural gas composition analysis was used to determine gas composition by values of the gas measurable parameters. however, expansion of such systems is limited by huge complexity of distributed software debugging [10]. the main tasks of further research are to develop a neural network model with a more complicated architecture, to optimize the number and set of input measurable parameters to make this model more accurate. 24 i. brokarev, s. vaskovskii copyright ©2019 assa adv. in systems science and appl. (2019) acknowledgements the authors are grateful to brisk ii ta and erasmus + 2017-1-se01-ka107-034292 staff mobility for providing the opportunity to conduct this study. references 1. altfeld k., schley p. (2012) development of natural gas qualities in europe, heat processing, 3, 77-83. 2. dorr h., koturbash t., kutcherov v. (2019) review of impacts of gas qualities with regard to quality determination and energy metering of natural gas, measurement science and technology, 30(2), 1-20. 3. dynament infrared gas sensors website [online]. available: https://www.dynament.com. 4. koturbash t., bicz a., bicz w. (2016) new instrument for measuring velocity of sound and quantitative characterization of binary gas mixtures composition, measurement automation monitoring, 3, 254-258. 5. koturbash t.t., brokarev i.a. (2018) metod opredeleniya svoystv i sostava prirodnogo gaza po izmereniyam ego fizicheskikh parametrov [the method of determination of the properties and the composition of natural gas by measuring of its physical parameters]. sensors & systems, 6(1), 43-50, [in russian]. 6. matlab 2018b software [online]. available: https://www.mathworks.com. 7. modcon systems ltd website [online]. available: https://www.modconsystems.com. 8. refprop software [online]. available: https://www.nist.gov/srd/refprop. 9. stepin u.p., trahtengerts e.a. (2007) kompyuternaya podderzhka upravleniya neftegazovymi tekhnologicheskimi protsessami i proizvodstvami.kniga 1. metody i algoritmy formirovaniya upravlencheskikh resheniy [computer-aided management of oil and gas technological processes and industries. book i. methods and algorithms of managerial decisions]. 221-223, [in russian]. 10. vaskovskii s.v. (2009) sredstva otladki raspredelennykh mikroprotsessornykh kompleksov realnogo vremeni [ways of debugging distributed microcomputerized real time systems]. sensors & systems, 1, 44-45, [in russian]. 11. xensor integration website [online]. available: http://www.xensor.nl advances in systems science and applications (2012) vol.12 no.2 186-193 a novel grey-fuzzy predictive control algorithm based on physical mode for roller kiln liang tang and mingzhong yang school of mechanical and electronic engineering, wuhan university of technology, wuhan 430070, china abstract this paper proposes a novel grey-fuzzy predictive control (gfpc) strategy using an on-line dynamic switching mechanism for controlling the sintering process of ceramic roller kiln. the grey predictor is applied to extract key information such as temperature parameters and gas pressure, and send the prediction information to the controller of actuator. the mathematical model which based on physical and heat conduction theory is derived and the real-time fuzzy controller is designed to make the output temperature in the sintering area follow the temperature-curve accurately. to achieve excellent transient performance and steady-state response, an on-line switching mechanism is adopted to regulate the forecasting step size of the grey predictor appropriately according to the error between reference value and real temperature of the kiln. additionally, this kind of method can control the temperature of kiln without accurate mathematical models for the kiln effectively. to verify the gfpc strategy, the control module based on matlab is presented. when the sintering operating condition is stable, it is a kind of effective control strategy for temperature of roller kiln. keywords roller kiln, grey-fuzzy predictive control, firing process, physical mode, heat conduction 1 introduction chinese construction industrial is booming, the ceramic tile manufacturing industry is one important industry related to this field. ceramic tile manufacturing line is a very complicated process which contained pressing machine, vertical dryer, enameling, firing kiln, sorting line, and packaging. roller kiln have important applications in ceramic industry and shows superiority over other intermittent kilns in process flexibility, production efficiency and energy consumption. roller kiln is nearly the last station and its performance is key process which control final tile quality. therefore, the finished ceramic tile has to satisfy the engineering parameters such as shape, size, durability and deformation as much as possible. if temperature values in roller kiln do not follow the reference curve accurately, then many serious defects, such as cracks and fractures, black heart, insufficient firing and over firing, tonality variations of the glaze will occur in the finished ceramic tiles [1]. however, if there are any flaws at the firing process all efforts of the preceding stations will become the waste. in view of these factors, in order to avoid the defects and obtain better quality, good controller is required to improve advances in systems science and applications (2012), vol.12, no.2 187 the control accuracy of temperature process. an appropriate control system designed to make temperature process control in roller kiln as accurate as possible. so far, researches on roller kiln have concentrated most on the kiln construction, linging material, loading and unloading machinery, and other installations. only a few studies considered the how to make the output temperature of roller kiln track the set-point values closely is a key factor affecting the firing process and the products quality. huang and fang [2] design the distributed intelligent control system for roller kiln, which using auto-tuning method based on relay feedback theory in order to decide the pid’s control parameter, and using the fuzzy control theory on this foundation. chen et al. [3] present logic control algorithm based on panboolean algebra, the control scheme relies on the correlation between e and ec of temperature. however, to be able to control this temperature process, in this study such as pid controller and fuzzy logic controller, for the best response characteristic of the system in time domain, you need choose the precision parameters. therefore, fuzzy logic controller improves the intelligent behavior, robustness, response speed, and accuracy of control system. due to the complexity of physical behaviors of temperature process in roller kiln, one of the ways to improve the system is by collecting the input-output data of plant and these details used to train nn or neuro-fuzzy model. thus, nguyen et al. present a comparative study using anns and co-active neuro-fuzzy inference system (canfis) in modeling a real, complicated multi input multi output (mimo) nonlinear temperature process of roller kiln used in ceramic tile manufacturing line[1]. in consideration of previous studies, to improve the systems performance, we proposed a novel control algorithm based on grey-fuzzy theory. the reason for using grey prediction in a fuzzy control system is that the grey predicted output from an unknown plant could always provide us some useful information for better control of the system before the system behavior runs into bad situations. the rest of this paper is organized as follows. in section 2 and 3, the structure and the mathematical model for roller kiln based on grey-fuzzy predictive control is presented. in section 4, matlab program with the proposed control scheme are obtained. finally, conclusion remarks and future work are presented in section 4. 2 grey prediction model it has been more than twenty years since the grey system theory proposed by a chinese professor julong deng in the 1980s. the grey predictive method has been successfully used to model the dynamic systems in different fields such as agriculture, ecology, economy, statistics, meteorology, industry, environment, and so on [4]. grey system theory has classified its research objects into three kinds of ‘black’, ‘white’ and ‘grey’, according to some cognitive hierarchy. 188 liang tang:a novel grey-fuzzy predictive control algorithm based... the general form of a grey differential model is gm (i, j), where i is the order of the ordinary differential equation and j the number of grey variables. the grey model gm (1, 1) with a single parameter has been most widely applied, therefore the model can be derived by the following basic steps in the following [5-6]. step 1 : rc and ago (accumulated generating operation) considering a temperature sequence as the original data, we have η s(0) = {t (0)(1), t (0)(2), . . . , t (0)(m)} (1) where m=2,. . . ,η,η ≥ 4 wherein (0) represents original data of grey prediction. the data sequence proceeded by ago once is s(1) = ago ◦ s(0) = {t (1)(1), t (1)(2), . . . , t (1)(m)},m = 1, 2, . . . , η, η ≥ 4 (2) which t (1)(1) = t (0)(1) t (1)(m) = t (1)(m− 1) + t (0)(m),m = 2, 3, . . . , η, η ≥ 4 (3) step 2: build grey model gm (1,1) the gm (1,1) model can be constructed by establishing a first order differential equation as follows: dt (1)(x) dx + at (1)(x) = b (4) where a and b are estimation parameters.z(1)(m) is affected by t (1) z(1)(m) = αmt (0)(m) + t (1)(m− 1) (5) m = 2, 3, . . . , η, η ≥ 4 however,the selection of αm value will influence the accuracy of predicted value. the solution for eq.(3) with discretization is t̂ (1)(m+ 1) = (t (0)(1)− b a )e−am + b a (6) m ≥ 0 wherein̂represents grey forecast and t̂ (1)(m+1) is the grey predicted value of the s(1) . advances in systems science and applications (2012), vol.12, no.2 189 here, we set b =  −z(1)(2) 1 −z(1)(3) 1 ... ... −z(1)(m) 1  , t =  t (0)(2) t (0)(3) ... t (0)(m)  (7) m = [ a b ] ,m = 2, 3, . . . , η, η ≥ 4 so, eq.(4) can be substituted as tn = bm (8) then the optimal parameter m can be obtained by using the minimum leastsquare estimation algorithm [ a b ] = (btb)−1btt (9) fig.1 measured points and predicted points for the temperature in roller kiln 3 modeling of grey-fuzzy predictive control 3.1 the derivation of physical model for roller kiln as shown in fig.1, in the ceramic roller kiln, dozens of nozzles distributed upper locations and lower locations alone roller kiln which controlled the temperature 190 liang tang:a novel grey-fuzzy predictive control algorithm based... at considered point within a limited range. the kiln shows non-linearity, due to the different behavior in the heating and in the cooling phases. however, it is very difficult, even impossible to build an accurate and perfect mathematical model for temperature process in ceramic roller kiln. so, with reference to fig. 1, our research assumes that the kiln has been wrapped up by the insulated materials and there is no any heat source besides nozzle inside such target. the governing equation for the one dimensional transient heat conduction is given by ∂t ∂t = ∂ ∂x (k ∂t ∂x ) (10) namely, ∂t ∂t = k ∂2t ∂x2 + ∂k ∂x ∂t ∂x (11) t means the temperature, t indicates the time, x indicates the positional parameter, and k indicates the thermal conductivity. wherein, eq.(11) can be discretizated into n-1 units as ti j+1 − ti j ∆t = ki j ti j+1 − 2ti j + ti−1 j (∆x)2 + ki+1 j − ki−1 j 2∆x ti+1 j − 2ti−1 j 2∆x (12) from manuscript [7], we can deduce at = dc (13) and, t = a−1dc = ec,e = a−1d (14) c = (ete)−1ettgrey (15) wherein tgrey indicates the temperature acquired by the grey prediction. according to eq.(13)-(15),it is dispensable to measure the temperature at each dispersed point. 3.2 grey-fuzzy predictive controller of roller kiln thus, in our control system, the control procedures are summarized as follows: (1) measure the temperature signal from thermocouples (sensors). (2) calculate output by using the control algorithm and switch the responses instantaneously. the temperature control system of the ceramic roller kiln consists of thermocouples (sensors), valves, actuators, electrically-driven conveyors and other advances in systems science and applications (2012), vol.12, no.2 191 components to form realtime control. schematic diagram of the proposed gfpc (greycfuzzy predictive control) with an on-line dynamic switching mechanism is shown in fig. 2. r(t) is the reference value, t(t) indicates the measurement signal, t ∗(t) = tgrey is the grey prediction value, u(t) is the output of the fuzzy controller, and e∗(t) is the offset: e∗(t) = r(t)− t ∗(t). fig.2 the structure of grey-fuzzy predictive controller in this paper the whole sintering will be dispersed into 8 units. according to eq. (6), thus t̂ (1)(5) = (t (0)(1)− b a )e−4a + b a (16) since the forecasting step size decides the predictive value and finally affects the control performance, an on-line switching mechanism is adopted to regulate the appropriate forecasting step size of the grey predictor [7]. this paper, we have also presented the comparison on selecting the measuring points for temperatures in order to reduce the synergistic effect on the grey prediction with errors of measurement. therefore, the grey predictor has the ability of predicting the ”previous” behavior of the system based on different states of feedback responses by either increase or decrease of control signal applied to actuator. we have the following switching mechanism: p =  p1 < 0 if e(t) > e1 p2 > 0 if eξ < e(t) < e1 p3 < 0 if e(t) < eξ (17) in kiln temperature control system, e and ec (e means error between reference value and real temperature of the kiln, and ec means the differential coefficient of e) are typical variables. any control rules of this temperature control system 192 liang tang:a novel grey-fuzzy predictive control algorithm based... should be based on these variables [3]. then, in our research, the input variables of the fuzzy controller are e∗(t) = r(t)− t ∗(t) ∆e∗(t) = [e∗(t)− e∗(t− t )]/t (18) where e∗(t) and ∆e∗(t) are the output error and the error change, respectively. 4 results and discussion due to the characteristics of long time-delay, nonlinearity, coupling and incompletion of parameter information in sintering process, the grey predictive controller is established by using the matlab. the algorithm is shown in fig.3. fig.3 algorithm of grey predictive controller 5 conclusions in this paper, a control algorithm for temperature in the roller kiln has been proposed. from the study we concluded that this temperature process is too complicated to build a certain mathematical model for identification purpose. this work focused on analyzes the heat conductive feature. basically, from the viewpoint of developing an effective control system, several key technologies were analyzed and the structure of controller was developed successfully in this paper. besides, we created matlab project to implement the predictive controller. however the limitation so far that all the control system are obtained by theory and simulation and in future we are planning to test the system for a pilot ceramic plant. advances in systems science and applications (2012), vol.12, no.2 193 references [1] nguyen quoc dinh, nitin v. afzulpurkar. (2007), “neuro-fuzzy mimo nonlinear control for ceramic roller kiln”, industrialsystems engineering, asian institute of technology, (ait), p.o. box 4, klong luang, pathumthani 12120. [2] huang yi-xin, fang yi-bin. (2002), “reaearch and application of distributed intelligent control system for temperature in rolling kiln”, proceedings of the csee, vol.22, no.5, pp.148-151. [3] chen jing, xu guocheng, xiao chun, yuan youxin, xiang kui, lang jianxun. (2008), “logic control algorithm based on panboolean algebra and its application for temperature control of ceramic roller kiln”, 2008 ieee pacificasia workshop on computational intelligence and industrial application. [4] g.r. yu, c.w. chuang, r.c. hwang. (2001), “fuzzy control of brushless dc motors by grey prediction”, ninth ifsa world congress and 20th nafips international conference, vol.5, pp.2819-2824. [5] lisheng wei, minrui fei, huosheng hu. (2008), “modeling and stability analysis of grey-fuzzy predictive control”, neurocomputing 72, pp.197-202. [6] f. cupertino, v. giordano, d. naso, l. delfine. (2006), “fuzzy control of a mobile robot”, ieee rob. autom. mag, no.13, pp.74-81. [7] jaw-yeong chiang, cha’o-kuang chen. (2008), “application of grey prediction to inverse nonlinear heat conduction problem”, international journal of heat and mass transfer, vol.51, pp.576-585. advances in systems science and applications (2014) vol.14 no.4 325-345 dynamic stackelberg games with requirements to the controlled system as a model of sustainable environmental management sergey a. kornienko and guennady a. ougolnitsky southern federal university, russia abstract a definition of sustainable environmental management based on the hierarchical game theoretic formalization is proposed. the definition includes requirements both to the state of the considered environmental system and to the control impact on it. a classification of leaders in the model context is given, a model example is described, and a qualitative assessment of the proposed game theoretic models is fulfilled. keywords dynamic stackelberg games, sustainable environmental management 1 introduction the term “sustainable development” was introduced by the international commission for environment and development in 1987. it determines a development that “meets the needs of the present without compromising the ability of future generations to meet their own needs” [1]. an apparently simple intergenerational rule is that development is sustainable “if it does not decrease the capacity to provide non-declining/capita utility for infinity” [2]. the weak sustainability rule requires that total net capital investment, or the rate of change of total net capital wealth, not allowed to be persistently negative. this definition entails the assumption that natural capital is similar to produced capital and can be substituted for it. proponents of the strong sustainability concept argue that natural capital is to a greater or lesser extent non-substitutable [3]. the concept of sustainable development (sustainability) is very vague and fuzzy. pezzey listed 60 published definitions of sustainable development[4]. a detailed analysis of the modern state of the art is made by zaccai [5]. glavic and lukman have proposed a hierarchical classification of concepts and terms concerned with sustainability[6]. some researchers argue for a special science about sustainable development (sustainability science)[7-9]. necessity of sustainability economics as a complement to ecological economics is discussed by baumgartner and quaas [10]. there are many principles of sustainability assessment and measurement [11]. an example of the assessment is given in literature[12]. nooteboom puts the problem of environmental assessment procedures in the context of complexity theory[13]. some authors tried to give the concept of sustainability a formal nature, either symbolic [14], discursive [15], or reflexive [16]. the axiomatic foundation of sustainable development based on the concept of sustainable preferences has 326 a. ashimov: the theory of parametric control of macroeconomic systems and ... been launched by chichilnisky [17]. asheim and mitra have introduced sustainable discounted utilitarianism, allowing to resolve intergenerational conflicts while satisfying the two main chichilnisky axioms[18]. d’albis and ambec addressed the question of fair intergenerational sharing of scarce natural resources[19]. a very interesting approach to the mathematical formalization of ecosystem sustainability based on optimal control theory is described in literature[20-28]. in these papers, fisher information as a sustainability measure for dynamic systems is proposed, and the sustainability hypotheses with particular focus on the natural ecosystems are formulated. it was also marked that strong efforts on different levels are necessary to provide sustainability. the sustainable development of an environmental system could be defined as: 1) an integral and balanced development of all aspects of the system functioning; 2) concordance of interests of all agents associated with the system; 3) a trade-off between short-term and long-term criteria of the system efficiency. it is very important to notice that to ensure sustainability it is necessary to provide some requirements not only to the state of the environmental system, but also to the control impact on it. as far any environmental system is impacted by many associated agents, the environmental management is a conflict process that should be described by game theoretic models. in many practical situations a traditional model of controlled dynamical system is not sufficient. for example, let the control subject be an industrial enterprise situated on the bank of a river which is the control object. the enterprise objective is to maximize its profit without consideration of the river water pollution by industrial sewage. so, in many cases the actions of control subject determined by his/her private interests and objectives are able to bring the control object in a state which is not acceptable from the point of view of sustainability. this suggests that an additional (higher) control level should be introduced to provide sustainability requirements to the state of control object. this new control subject is able to have his/her own interests too. to achieve his/her objective the new control subject can exert an impact to the initial control subject. for example, an environmental protection agency can control water quality by establishing sewage limits and charging penalties if an enterprise exceeds them. in this case a new subject of the hierarchical control arises which is a complexly structured system with its internal links and relations. dynamic stackelberg games [29] is a relevant model in this case. from the other side, sustainable development requires strong collaborative efforts of states, corporations, social organizations and individuals. from this point of view, a cooperative game theoretic formalization is required and the main problem is time consistency of the trade-off solutions. literature[30-35] presented classes of transferable-payoff cooperative games with solutions which advances in systems science and applications (2014) vol.14 no.4 327 satisfy group optimality and individual rationality. in literature[36-43] are presented solutions satisfying group optimality and individual rationality at the initial time in cooperative games with nontransferable payoffs. in literature[38-39] threats are used to ensure that no players will deviate from the initial cooperative strategies throughout the game horizon. the problem of time consistency in differential games has been intensively explored in the past decades [44]. haurie [42] raised the problem of instability when the nash bargaining solution is extended to differential games. petrosyan [45] formalized the notion of time consistency in differential games. kidland and prescott [46] introduced the notion of time consistency related to economic problems. petrosyan [47-48]and petrosyan and zenkevich [49-50] presented a detailed analysis of time consistency in cooperative differential games, in which a method of regularization was used to construct time consistent solutions. yeung and petrosyan [51] designed time consistent solutions in differential games and characterized the conditions that the allocation-distribution procedure must satisfy. petrosyan and zenkevich [52] proposed the conditions of sustainable cooperation. environmental game theoretic applications are considered in literature [3235,39,44,53-63] and others. in the author’s publications [64-67] an approach to the mathematical formalization of sustainable development and sustainable management is developed. our contribution in this paper is to describe a model of sustainable environmental management that includes requirements both to the state of the considered environmental system and to the hierarchical control impact on it. in the section 2 of the paper a mathematical formalization of sustainable environmental management is proposed. in the section 3 dynamic compulsion and impulsion stackelberg games with requirements to the controlled system are introduced. in the section 4 a classification of leaders in the games is given. section 5 is dedicated to an illustrative model example. in section 6 the principles of assessment of decision making models of sustainability [68] are used for the proposed model. section 7 concludes. 2 sustainable environmental management: mathematical formalization thus, to define sustainable environmental management it is necessary to formulate certain requirements both to the state of the environmental system and to the control impact on it. to some extent of conventionality the first group of requirements is called “homeostasis”, and the second group “compromise” [65]. requirements of homeostasis. highly organized systems of the real world should resist to the external impacts or accommodate to them providing a conservation of the conditions of their existence and goal-oriented development. as a french physiologist claude bernard has said: “the constancy of the internal 328 a. ashimov: the theory of parametric control of macroeconomic systems and ... milieu is a condition of the free life of organism” (1878). the term “homeostasis” was introduced in 1932 by walter cannon who treated it as a relative dynamic constancy of the whole organism [69]. there is no escape from the conclusion about a similarity between the concepts of homeostasis as a dynamic constancy of the organism and sustainable development as a combination of the economic development (dynamics) and the environmental stability (constancy) on the biosphere level. we think that the concept of homeostasis could be considered not only on the level of a separate organism but also on the levels of populations, ecosystems, environmental-economic systems, and arbitrary dynamic systems including human beings. let’s name the homeostasis of a system the domain of values of the essential parameters of the system in which its normal existence and development are possible. functioning of any dynamic system is characterized by a set of parameters the values of which change in the time. the condition of homeostasis means that all parameters of the system functioning during a considered period of time (long enough or even infinite) take their values from a given range, and in the specific case take given point values. for example, the point requirements to the physiological parameters of a human organism are well known: temperature 36.6 degrees centigrade, blood pressure 120 on 80 millimeters of mercury column, and so on. in the same time certain deviations from the standard values are allowable: they form the admissible ranges of functioning of a healthy organism. the point and interval requirements of homeostasis of any dynamic system including human beings can be given similarly. in the mathematical modeling a set of the essential parameters of functioning of a dynamic system is called its vector of state (phase vector) and is denoted x(t) = (x1(t), ..., xn(t)). its components (state variables) xi(t)(i = 1, ..., n) are the values of parameters which characterize the system state in the moment of time t from the point of view and with the degree of detail which are determined by the objectives and resources of the research. for example, a population can be characterized by one parameter (its number or biomass) or, more precisely, by dozens of parameters considering its sex, age, genetic, and other structuring. then a requirement (condition) of homeostasis can be formulated in two forms: weak and strong. in the weak form the condition can be written as (lagrange stability) ∀t ∈ [0, t ] : x(t) ∈ x∗ (1) where x∗ is the domain of homeostasis, t the length of the considered time period. for example, ∀t ∈ [0, t ] : x(t) > 0 (population does not extinct), or ∀t ∈ [0, t ] : x(t) > xcr (number of population is not less than a dangerous threshold), or ∀t ∈ [0, t ] : x(t) ≤ x̄ (concentration of a pollutant does not exceed the maximal allowable one). advances in systems science and applications (2014) vol.14 no.4 329 often it is possible to define the domain of homeostasis as x∗ = [x∗−ε, x∗+ε], where x∗i is an “ideal” value of the i-th state variable, ±εi is an allowable deviation from the value (i = 1, ..., n). in this case the condition (1) takes the form ∀t ∈ [0, t ] : x(t) ∈ [x∗ − ε, x∗ + ε] (2) and becomes in fact the condition of (neutral) lyapunov stability of a stationary point x∗ of the controlled environmental system. the strong form of requirement of homeostasis means satisfaction of (2) together with the additional condition lim t→∞ x(t) = x∗ (3) i.e. asymptotic lyapunov stability of the stationary point x∗ of the controlled environmental system [70]. let’s notice that it is possible to require not only stability of a stationary point, but also of another trajectory of a controlled dynamic system (periodical oscillations, linear or exponential growth, and so on). the choice of the specific form of the requirement of homeostasis (lagrange stability, neutral or asymptotic lyapunov stability of a stationary point or another trajectory) is determined by a specific environmental sustainable management problem under consideration. requirements of compromise. it is extremely important to notice again that the requirements of homeostasis are necessary but not sufficient to provide sustainable environmental management. many agents are associated with any environmental system. from one side, their objectives and interests are determined to an extent by the state and dynamics of the system. from the other side, the agents impact the system and exert some effect on its functioning. thus the sustainable development is possible only if a condition of the consideration and coordination of interests of the associated agents is satisfied. as the objectives and interests of agents do not coincide in the majority of cases then their interaction is conflict. but in the same time the objectives and interests are not antagonistic, therefore a compromise is possible. in the light of the requirements of sustainable development the compromise should consider the condition of homeostasis. it is this compromise between all agents associated with the system that forms another condition of its sustainable development. in the mathematical formalization of a conflict interaction by game theoretic models compromises are described by optimality principles for different classes of games. thus from the mathematical point of view the condition of compromise means an existence of solution of the game describing a conflict interaction of the agents associated with the system. the condition of homeostasis reflects the requirement to the state of the system meanwhile the condition of compromise formalizes the demands to the impact on 330 a. ashimov: the theory of parametric control of macroeconomic systems and ... it. if a compromise is not achieved then the system will be permanently threatened by not appropriate impacts violating the homeostasis. a simple question should be answered: who will ensure the condition of homeostasis and why does (s)he need it? note that any solution of a game (optimality principle) has some properties that provide a stability of the compromise. so, nash equilibrium is stable towards individual deviations, i.e. no player has incentives to violate an initial agreement. in stackelberg equilibrium the follower chooses an optimal answer to the leaders strategy which in turn is chosen such as to provide the leader the maximal payoff on the set of the optimal answers. in the case of a cooperative solution the key property is pareto optimality due to which a players payoff can be increased only at the cost of other players. another key problem of the sustainable management is a possible nonconformity of the short-term and long-term criteria of optimality. the condition of homeostasis is a long-term one because it should be satisfied along the whole period of existence and functioning of the environmental system. in the same time the agents associated with the system are often guided by short-term criteria with much smaller character times. as a result the compromise is under the threat of violation by those participants of the initial agreement for which it could be more advantageous in the current moment of time to take another strategy corresponding to their short-term interests. in this connection a realization of the requirements of sustainable management is possible only for those compromises which keep their optimality for all associated agents along the whole period of existence of the environmental system. in the game theoretic formalization this principle was called a time consistency [51-56, 70]. the property of time consistency of the solution of a game means that a truncation of the solution is still optimal in all subgames arising along the optimal trajectory of the system development. this property provides a practical realization of the compromise solution from which it is not advantageous for any agent to deviate all along the time of system functioning. thus, the weak form of the requirement of compromise is existence of a solution of a game theoretic model of conflict interaction of the agents associated with the controlled environmental system. in the case of a hierarchically controlled environmental system that is under consideration in this paper the stackelberg solution is used. the strong form of the requirement of compromise means additionally the time consistency of the solution. it could be stated that separately the characterized requirements of homeostasis and compromise are necessary, and in their totality also sufficient conditions of the sustainable development of any environmental system. the condition of homeostasis expresses basic requirements to all aspects of the system functioning, advances in systems science and applications (2014) vol.14 no.4 331 the condition of compromise provides the adequacy of impacts by all agents associated with the system with acceptable consideration of their interests, including a coordination of short-term and long-term optimality criteria of the agents and the consequent non-advantageousness for them to deviate from the initially agreed compromise solution all along the time. if the homeostasis is provided and a dynamic consistent compromise between all the associated agents is achieved then a sustainable environmental management of the system takes place. the sustainable environmental management ensures both conditions of sustainable development and means of their realization. the conditions of sustainable environmental management are characterized in table 1. table 1 conditions of sustainable environmental management. homeostasis compromise weak form lagrange (1) or neutral lyapunov stability (2) of a phase trajectory (for example, stationary point) of the controlled environmental system existence of a solution of the game theoretic model which has some strategic stability (for example, nash of stackelberg solution) strong form asymptotic lyapunov stability (2)-(3) of a phase trajectory (for example, stationary point) of the controlled environmental system time consistency of the solution 3 dynamic stackelberg games with requirements to the controlled system. in many practically important cases the stakeholders of an environmental system are organized hierarchically. then three methods of hierarchical control [65] can be used (table 2). their game theoretic formalization is proposed in [64]. consider the example from introduction in more details. the higher level control subject (leader) is an environmental protection agency, the lower level control subject (follower) is an industrial enterprise, and the control object is river ecosystem. it is natural to treat the desirable strategy for leader as such a strategy in which the industrial sewage doesnt exceed the maximum allowable concentration of pollutants in the river. in the case of compulsion the objective is reached by establishing some sewage limits and license recall from the enterprise if the limits are violated. in the case of impulsion if the enterprise exceeds the maximum allowable concentration of pollutants in the river then a penalty is charged. at last, in the case of conviction the enterprise administration is environmentally 332 a. ashimov: the theory of parametric control of macroeconomic systems and ... table 2 characteristic of the methods of hierarchical control. compulsion impulsion conviction general description of the method leader provides choosing the desirable strategy of follower by force the strategy desirable for leader is more profitable for follower than undesirable ones follower chooses the strategy desirable for leader voluntarily and consciously nature of impact administrative or legislative economic socialpsychological type of relationships subject-to-object subject-to-object with partial consideration of follower’s interests subject-to-subject mathematical formalization leader’s impact on the follower’s set of admissible strategies leader’s impact on the follower’s payoff function transition of leader and follower to cooperation and maximization of the summarized coalitional payoff function conscious and it provides the required sewage refinement voluntarily. to find the strategies of sustainable environmental management it is necessary to build game theoretic models with requirements to the state of controlled system. we will consider dynamic stackelberg games which formalize hierarchical relations between the stakeholders of an environmental system based on compulsion or impulsion. a possible transition to conviction (cooperation) is also considered. dynamic compulsion stackelberg game with requirements to the controlled system. the game can be written as follows: jl = ∫ t 0 e−αt[gli(u(t), x(t))− glc(q(t))−mρ(x(t), x∗)]dt → max (4) q(t) ∈ q (5) jf = ∫ t 0 e−αtgf (u(t), x(t))dt → max (6) u(t) ∈ u(q(t)) (7) ẋ = f(x(t), u(t)), x(0) = x0 (8) advances in systems science and applications (2014) vol.14 no.4 333 where jl, jf – payoff functionals of the leader and the follower respectively; t – period of consideration; α – discount factor; x(t) – state vector of the hierarchically controlled environmental system; u(t) – vector of controls of the follower (impact on the controlled environmental system); q(t) – vector of compulsive controls of the leader; gli – function of the leader’s personal interest (for example, income); glc – function of control cost of the leader (glc(0) = 0); m – penalty constant; ρ(x(t), x∗) = { 0, x(t) ∈ x∗, 1, otherwise; – indicator function; q,u – sets of admissible controls of the leader and the follower respectively; gf – instant payoff function of the follower; f – known function of controlled dynamics of the environmental system; x0 – given initial condition. compulsion in the game (4) – (8) consists in that the leader by controls q(t) exerts an impact to the set of the follower’s admissible controls u(q(t)) (without control dependence). requirements to the state of the controlled environmental system (without loss of generality in the form of lagrange stability) are described by means of an indicator function ρ and a penalty constant m in the leader’s payoff functional. therefore, if the requirements to the controlled system (treated as its homeostasis conditions) are violated then the leader is charged a penalty m . a stackelberg equilibrium [29] is accepted as solution of the dynamic game (4) – (8). its specific feature is presence of the indicator function ρ in the leader’s payoff functional subject to which the functional can become discontinuous when m → ∞. therefore a key role belongs here to the solution of a parametrical inverse optimal control problem, i.e. the problem of building of the set of “homeostatic” controls of the follower u∗ = {u(t) ∈ u(q(t)) : x(t) ∈ x∗} (9) if ∃q(t) : u∗ ̸= ∅ then solution of the game (4) – (8) is reduced to the solution of an ordinary dynamic stackelberg game without requirements to the controlled system, otherwise the game (4) – (8) has no solution. time consistency of a solution of the game (4) – (8) depends on its information structure. it is known [29] that in the class of open-loop strategies a stackelberg solution is not time consistent, i.e. in this case only the weak form of the requirement of compromise can be discussed. to provide time consistency (the strong form of compromise) closed-loop strategies should be considered. dynamic impulsion stackelberg game with requirements to the controlled system. the game can be written as follows: jl = ∫ t 0 e−αt[gli(p(t), u(t), x(t))− glc(p(t))−mρ(x(t), x∗)]dt → max (10) p(t) ∈ pu (11) 334 a. ashimov: the theory of parametric control of macroeconomic systems and ... jf = ∫ t 0 e−αtgf (p(t), u(t), x(t))dt → max (12) u(t) ∈ u (13) subject to (8), where in comparison with (4) – (7) p(t) – vector of impulsive controls of the leader. impulsion in the game (8), (10) – (13) consists in that the leader by controls p(t) exerts an impact to the follower’s payoff function. a control dependence (feedback) also takes place, namely p̃(t) = p(u(t)). therefore the set of the follower’s admissible strategies is a set of maps pu = {p : u → p}. considerations about the stackelberg solution are the same as in the case of compulsion. transition to conviction. both in the case of compulsion and in the case of impulsion a transition to conviction is possible. this transition means cooperation between the leader and the follower for joint solution of the problem of sustainable environmental management. from the mathematical point of view, conviction means a coalition of the leader and the follower and their joint maximization of the summarized payoff functional, or solution of the optimal control problem jl+f = ∫ t 0 e−αt[gli(u(t), x(t))− glc(q(t))−mρ(x(t), x∗) + gf (u(t), x(t))]dt → max (14) q(t) ∈ q; u(t) ∈ u(q(t)) subject to (8). suppose that u∗ ̸= ∅ (otherwise the sustainable environmental management problem has no solution). then it is evident that the follower chooses u(t) ∈ u∗. thus q(t) ≡ 0 (compulsion is not required), respectively glc(0) = ρ(x(t), x∗) = 0, and the problem (14) takes the form jl+f = ∫ t 0 e−αt[gli(u(t), x(t)) + gf (u(t), x(t))]dt → max, u(t) ∈ u (15) subject to (8), i.e. only the follower maximizes the coalitional payoff functional. after solution of the optimal control problem (15) the received payoff is shared between the leader and the follower accordingly to a cooperative optimality principle [36-37]. in the case of impulsion a transition to conviction is similar. 4 classification of leaders in the proposed conceptual framework the leader is treated as a regulator who is responsible for the requirements to the controlled environmental system, or a subject of the sustainable environmental management. the following classification attributes can be used accordingly to the form of the payoff functional (4) or (10). advances in systems science and applications (2014) vol.14 no.4 335 importance of the requirements to the controlled environmental system. from the point of view of the leader, the requirements can be: mandatory (m → ∞): in this case the leader can’t solve her control problem without ensuring the requirements because she is charged an arbitrarily big penalty otherwise; desirable (m ≈ gli , glc): in this case the leader may compare what is more profitable: to ensure the requirements, to reduce cost or to increase her income; neglectable (m ≈ 0): in this case the leader is in fact not interested to ensure the requirements. the first variant can be considered as an “ideal” setting of the problem of sustainable environmental management, the second one as a more practical model of the real life. the third variant is a degenerated one because the requirements of homeostasis are neglectable (in fact, absent). notice that the specific problem of sustainable environmental management arises only if the requirements are mandatory. presence of the personal interest. the leader can be disinterested (gli = 0) or interested (gli > 0). in the former case the only leaders objective is to ensure the requirements of homeostasis (considering her control cost) while in the latter case the leader has also her personal interest (for example, to get a share from the collected taxes or fines). control efficiency. a process of control generates cost for the leader. so, the efficiency of control can be high (glc ≈ 0) or low (glc >> 0). in the former case the costs are very small, and the leader can concentrate on the requirements of homeostasis (probably considering her personal interest). in the latter case the control costs become an essential factor. 5 a model example for the sake of parsimony, let’s consider as an illustrative example the most simplistic malthus model of a controlled environmental system ẋ = (a− u(t))x(t), x(0) = x0 (16) where x(t) – number (biomass) of a population; u(t) ∈ [0, 1] – share of its exploitation (harvesting). lets require that ∀t ∈ [0, t ] x0 − ε ≤ x(t) ≤ x0 + ε (17) be the weak condition of homeostasis (given that the initial population value is optimal for the habitat). then the strong form of homeostasis is (17) plus lim t→∞ x(t) = x0 (18) for each u(t) the solution of cauchy problem (16) takes the form x(t) = x0e (a−u(t))t (19) 336 a. ashimov: the theory of parametric control of macroeconomic systems and ... it is evident that both conditions (17) and (18) are satisfied for (19) when u(t) ≡ a (20) (the solution of the inverse optimal control problem). now lets investigate whether the leader can ensure (20). compulsion stackelberg game has the form jl = ∫ t 0 [kpu(t)x(t)− cq2(t)−mρ(x(t), x∗)]dt → max (21) 0 ≤ q(t) ≤ 1 (22) jf = ∫ t 0 (1− p)u(t)x(t)dt → max (23) 0 ≤ u(t) ≤ 1− q(t) (24) subject to (16). here discounting on the finite time interval [0,t] is omitted for simplicity, p – constant tax rate, k – share of the collected taxes in the leader’s income, c – factor of the leader’s control efficiency, m → ∞ – penalty constant. the leader can ensure the condition (20) by choosing q(t) ≡ 1− a (25) in this case is the stackelberg solution for the game (16), (21) c (24), and the players’ payoffs are jcomp l = (kpax0 − c)t, jcomp f = (1− p)ax0t (26) notice that jcomp l { > 0, kpax0 > c (the leader′s control is efficient), < 0, otherwise. . in the case of transition to conviction the coalitional control problem is jl+f = ∫ t 0 [(1− (1− k)p)u(t)x(t)− cq2(t)−mρ(x(t), x∗)]dt → max 0 ≤ q(t) ≤ 1, 0 ≤ u(t) ≤ 1− q(t) the team solution has the form ; in this case the condition (20) is satisfied and jconv l+f = (1− (1− k)p)ax0t > ((1− (1− k)p)ax0 − c)t = jcomp l + jcomp f impulsion stackelberg game has the form jl = ∫ t 0 [kp(t)u(t)x(t)− cp2(t)−mρ(x(t), x∗)]dt → max (27) advances in systems science and applications (2014) vol.14 no.4 337 0 ≤ p(t) ≤ 1 (28) jf = ∫ t 0 (1− p(t))u(t)x(t)dt → max (29) 0 ≤ u(t) ≤ 1 (30) subject to (16). here the leader can ensure the condition (20) by choosing the impulsive strategy p̃(t) = { 1− δ, u(t) = a, 1, otherwise, (31) and the follower’s optimal reaction is also u(t) ≡ a. thus, (1 − δ, a) is the stackelberg solution in the game (16), (27)–(30), and the players’ payoffs are j imp l = (kax0 − (1− δ)c)t, j imp f = δax0t (32) as in the case of compulsion, jcomp l { > 0, kax0 > (1− δ)c (the leader′s control is efficient), < 0, otherwise. let’s notice also that jcomp l < j imp l , but in the same time jcomp f >> j imp f . transition to conviction entails the following optimal control problem: jl+f = ∫ t 0 [(1− (1− k)p(t))u(t)x(t)− cp2(t)−mρ(x(t), x∗)]dt → max 0 ≤ p(t) ≤ 1, 0 ≤ u(t) ≤ 1. the team solution is also ; in this case the condition (20) is satisfied and jconv l+f = ax0t > ((k + δ)ax0 − (1− δ)c)t = jcomp l + jcomp f i.e. cooperation is more advantageous for both players. 6 qualitative assessment of the model in the paper [68] five methodological criteria for sustainability models are proposed: an interdisciplinary approach, uncertainty management, a long-range or intergenerational point of view, “glocality” and participation. let’s use these criteria for the qualitative assessment of our model of sustainable environmental management. interdisciplinary approach. this approach is followed in two aspects. first, state variables of the modeled controlled environmental system characterize its ecological, economic, social and other components. second, in the game theoretic models an equation of dynamics can describe environment, while payoff functions 338 a. ashimov: the theory of parametric control of macroeconomic systems and ... and sets of admissible strategies of the players determine socio-economic interests and possibility of their implementation. uncertainty. this requirement is naturally ensured by building and investigation of stochastic control (game theoretic) models. note that random variables can enter both in the dynamics equation and in the payoff functions. long-term perspective. this criterion is considered by the very setting of a dynamic game theoretic model. the game can be defined both on finite (long enough for the character scale of functioning) and on infinite period of time. an additional comparison of the temporal effects is provided by discounting of the payoffs. global-local perspective. this requirement means a consideration of the hierarchical nature of modeled processes that is ensured by using stackelberg games. in more complicated model settings the systems of equations describing the modeled system dynamics could include small parameters that reflect different scale of the modeled processes. participation. as in the case of “glocality”, this criterion is considered by the very game theoretic setting that means a compromise concordance of different interests of the stakeholders. thus, the qualitative analysis demonstrates that the proposed sustainable environmental management models satisfy the methodological criteria introduced by boulanger and brechet [68]. 7 conclusion in spite of the long-term discussion, the concept of sustainability remains vague and fuzzy. in this paper a mathematical formalization of the concept based on requirements of neutral or asymptotic lyapunov stability of an ideal trajectory in the state space of an environmental system is used. this approach can be considered as a formalization of the requirement of strong sustainability. however, this is not sufficient. while speaking about sustainable environmental management, it is necessary to add some requirements concerning the control process. it is provided by the description of interests of all stakeholders in a game theoretic model. as far in many practical important situations relations between the stakeholders are asymmetric, dynamic stackelberg games are used as the model. we consider compulsion stackelberg games in which the leader exerts an administrative impact to the set of the followers admissible controls without control dependence (feedback), and impulsion stackelberg games in which the leader exerts an economic impact to the followers payoff function and control dependence (feedback) takes place. in both cases a transition to conviction (cooperation) is possible that means a coalition of the leader and the follower and their team maximization of the joint payoff function. conviction is the most advances in systems science and applications (2014) vol.14 no.4 339 perspective and advantageous for both players if they are able to overcome different obstacles on the way to cooperation. all proposed game theoretic models include requirements to the controlled environmental system that are treated as conditions of its homeostasis. the proposed models are successfully tested by criteria described in [68]. it seems worthwhile to develop the proposed methodology for different classes of stackelberg games and environmental systems. acknowledgments the work is supported by southern federal university, project #213.01-07.2014/07 references [1] united nations. (1987), “report of the world commission on environment and development”, general assembly resolution, 42/187, 11 december. [2] neumayer e. (2003), “weak versus strong sustainability: exploring the limits of two opposing paradigms”, edward elgar, northampton, ma. [3] dietz s., and neumayer e. (2007), “weak and strong sustainability in the seea: concepts and measurement”, ecological economics, vol.61, pp.617626. [4] pezzey j. (1989), “economic analysis of sustainable growth and sustainable development: environmental department working” paper no.15, world bank, washington. [5] zaccai e. (2012), “over two decades in pursuit of sustainable development: influence, transformations, limits”, environmental development, vol.1, pp.79-90. [6] glavic p., and lukman r. (2007), “review of sustainability terms and their definitions”, journal of cleaner production, vol.15, pp.1875-1885. [7] kates r. et al. (2001), “sustainability science”, science, vol.292, pp.641642. [8] clark w.c., and dickson n.m. (2003), “sustainability science: the emerging research program”, proceedings of the national academy of science, vol.100, pp.8059-8061. [9] clark w.c. (2007), “sustainability science: a room of its own”, proceedings of the national academy of science, vol.114, pp.1737-1738. 340 a. ashimov: the theory of parametric control of macroeconomic systems and ... [10] baumgartner s., and quaas m. (2010), “what is sustainability economics?”, ecological economics, vol.69, no.3, pp.445-450. [11] pinter l., hardi p., martinuzzi a., and hall j. (2012), “bellagio stamp: principles for sustainability assessment and measurement”, ecological indicators, vol.17, pp.20-28. [12] shmelev s.e., and rodrigez-labajos b. (2009), “dynamic multidimensional assessment of sustainability at the macro level: the case of austria”, ecological economics, vol.68, pp.2560-2573. [13] nooteboom s. (2007), “impact assessment procedures for sustainable development: a complexity theory perspective”, env. impact assessment review, vol.27, pp.645-665. [14] baker s. (2007), “sustainable development as symbolic commitment: declaratory politics and the seductive appeal of ecological modernization in the european union”, environmental politics, vol.16, no.2, pp.297-317. [15] rumpala y. (2010), “recherche de voies de passage au ‘developpement durable’ et reflexivite institutionelle”, revue francaise de socio-economie, vol.2, no.6, pp.47-63. [16] kemp p., and martens p. (2007), “sustainable development: how to manage something that is subjective and never can be achieved?”, sustainability: science, practice and policy, vol.3, no.2, pp.5-14. [17] chichilnisky g. (1996), “an axiomatic approach to sustainable development”, social choice and welfare, vol.13, pp.231-257. [18] asheim g., and mitra t. (2010), “sustainability and discounted utilitarianism in models of economic growth”, mathematical social sciences, vol.59, no.2, pp.148-169. [19] d’albis h. and ambec s. (2010), “fair intergenerational sharing of a natural resource”, mathematical social sciences, vol.59, no.2, pp.170-183. [20] cabezas h., and fath b.d. (2002), “towards a theory of sustainable systems”, fluid phase equilibria, pp.194-197, 3-14. [21] cabezas h., pawlowski c.w., mayer a.l., and hoagland n.t. (2003), “sustainability: ecological, social, economic, technological, and systems perspectives”, clean technology and environmental policy, vol.5, pp.167-180. advances in systems science and applications (2014) vol.14 no.4 341 [22] cabezas h., pawlowski c.w., mayer a.l., and hoagland n.t. (2005), “simulated experiments with complex sustainable systems: ecology and technology”, resources, conservation and recycling, vol.44, pp.279-291. [23] cabezas h., pawlowski c.w., mayer a.l., and hoagland n.t. (2005), “sustainability systems theory: ecological and other aspects”, j. of cleaner production, vol.13, pp.455-467. [24] cabezas h., whitmore h.w., pawlowski c.w., and mayer a.l. (2007), “on the sustainability of the integrated model system with industrial, ecological, and macroeconomic components”, resources, conservation and recycling, vol.50, pp.122-129. [25] shastri y., and diwekar u. (2006), “sustainable ecosystem management using optimal control theory: part 1 (deterministic systems)”, j. of theoretical biology, vol.241, pp.506-521. [26] shastri y., diwekar u. (2006), “sustainable ecosystem management using optimal control theory: part 2 (stochastic systems)”, journal of theoretical biology, vol.241, pp.522-532. [27] shastri y., diwekar u., and cabezas h. (2008), “optimal control theory for sustainable environmental management”, environmental science and technology, vol.42, pp.5322-5328. [28] shastri y., diwekar u., cabezas h., and williamson j. (2008), “is sustainability achievable? exploring the limits of sustainability with model systems”, environmental science and technology, vol.42, pp.6710-6716. [29] basar t., and olsder g.j. (1999), “dynamic noncooperative game theory”, siam, philadelphia. [30] haurie, a., and zaccour, g. (1986), “a differential game model of power exchange between interconnected utilizers”, proceedings of the 25th ieee conference on decision and control, athens, greece. [31] haurie, a., and zaccour, g. (1991), “a game programming approach to efficient management of interconnected power networks”, springer-verlag, berlin. [32] kaitala, v., and pohjola, m. (1988), “optimal recovery of a shared resource stock: a differential game with efficient memory equilibria”, natural resource modeling, vol.3, pp.91-118. 342 a. ashimov: the theory of parametric control of macroeconomic systems and ... [33] kaitala, v., and pohjola, m. (1995), “sustainable international agreements on greenhouse warming: a game theory study”, annals of the international society of dynamic games, vol.2, pp.67-87. [34] kaitala, v., maler, k.g., and tulkens, h. (1995), “the acid rain game as a resource allocation process with an application to the international cooperation among finland, russia, and estonia”, scandinavian journal of economics, vol.97, pp.325-343. [35] jorgensen, s., and zaccour, g. (2001), “time consistent side payments in a dynamic game of downstream pollution”, j. of economic dynamics and control, vol.25, pp.1973-1987. [36] leitmann, g. (1974), “cooperative and non-cooperative many players differential games”, springer-verlag, n.y. [37] leitmann, g. (1975), “cooperative and non-cooperative differential games”, d.reide, amsterdam. [38] tolwinsky, b., haurie, a., and leitmann, g. (1970), “cooperative equilibria in differential games”, journal of math. analysis and applications, vol.119, pp.182-202. [39] hamalainen, r.p., haurie, a., and kaitala, v. (1986), “equilibria and threats in a fishery management game”, optimal control applications and methods, vol.6, pp.315-333. [40] haurie, a., and pohjola, m. (1987), “efficient equilibria in a differential game of capitalism”, journal of economic dynamics and control, vol.11, pp.65-78. [41] gao, l., jakubowski, a., klompstra, m.b., and olsder, g.j. (1989), “timedependent cooperation in games”, springer-verlag, berlin. [42] haurie, a. (1991), “piecewise deterministic and piecewise diffusion differential games with model uncertainties”, springer-verlag, berlin. [43] haurie, a., krawczyk, j.b., and roche, m. (1994), “monitoring cooperative equilibria in a stochastic differential game”, j. of optimization theory and applications, vol.81, pp.73-95. [44] yeung, d.w.k. (2009), cooperative game-theoretic mechanism design for optimal resource use. in: contributions to game theory and management. petrosyan, l.a., zenkevich, n.a. (eds.). vol. ii. collected papers presented on the second international conference game theory and management. spb.: graduate school of management spbu, pp.483-513. advances in systems science and applications (2014) vol.14 no.4 343 [45] petrosyan, l.a. (1977), “the stability of solutions in differential games with many players”, vestnik leningradskogo universiteta, ser.1, iss.4, vol.19, 4652 (in russian). [46] kidland f.e., prescott e.c. (1977), “rules rather than decisions: the inconsistency of optimal plans”, journal of political economy, 85, pp.473-490. [47] petrosyan, l.a. (1991), “the time-consistency of the optimality principles in non-zero sum differential games”, in: lecture notes in control and information sciences (eds. hamalainen r.p. and ehtamo h.k). vol.157, pp.299-311. springer-verlag, berlin. [48] petrosyan, l.a. (1993), “the time consistency in differential games with a discount factor”, game theory and applications, vol.1. [49] petrosyan, l.a., and zenkevich, n.a. (1996), “game theory”, world scientific publ.co., singapore. [50] petrosyan, l.a., and zenkevich, n.a. (2007), time consistency of cooperative solutions. in: contributions to game theory and management. petrosyan, l.a., zenkevich, n.a. (eds.). collected papers presented on the international conference game theory and management. spb.: graduate school of management sbbu, pp.413-440. [51] yeung, d.w.k., and petrosyan, l.a. (2001), “proportional time-consistent solution in differential games”, saint petersburg state university. [52] petrosyan, l.a., and zenkevich, n.a. (2009), conditions for sustainable cooperation. in: contributions to game theory and management. petrosyan, l.a., zenkevich, n.a. (eds.). vol. ii. collected papers presented on the second international conference game theory and management. spb.: graduate school of management spbu, pp.344-354. [53] andre f.j., sokri a., and zaccour g. (2011), “public disclosure programs vs. traditional approaches for environmental regulation: green goodwill and the policies of the firm”, european journal of operational research, vol.212, pp.199-212. [54] bailey m., rashid sumaila u., and lindroos m. (2010), “application of game theory to fisheries over three decades”, fisheries research, vol.102, pp.1-8. [55] breton m., zaccour g., and zahaf m. (2006), “a game-theoretic formulation of joint implementation of environmental projects”, european journal of operational research, vol.168, pp.221-239. 344 a. ashimov: the theory of parametric control of macroeconomic systems and ... [56] breton m., sokri a., and zaccour g. (2008), “incentive equilibrium in an overlapping-generations environmental game”, european journal of oper. research, vol.185, pp.687-699. [57] dockner, e.j., and long, n.v. (1993), “international pollution control: cooperative versus non-cooperative strategies”, journal of environmental economics and management, vol.24, pp.13-29. [58] fanokoa p.s., telahigue i., and zaccour g. (2011), “buying cooperation in an asymmetric environmental differential game”, j. of economic dynamics and control, vol.35, pp.935-946. [59] fredj k., martin-herran g., and zaccour g. (2004), “slowing deforestation page through subsidies: a differential game”, automatica, vol.40, pp.301309. [60] legras s. and zaccour g. (2011), “temporal flexibility of permit trading when pollutants are correlated”, automatica, vol.47, pp.909-919. [61] madani k. (2010), “game theory and water resources”, j. of hydrology, vol.381, pp.225-238. [62] mazalov, v., and rettieva, a. (2004), “a fishery game model with migration: reserved territory approach”, game theory and applications, vol.10. [63] mazalov, v., and rettieva, a. (2007), cooperative incentive equilibrium. in: contributions to game theory and management. petrosjan, l.a., zenkevich, n.a. (eds.). collected papers presented on the international conference game theory and management. spb.: graduate school of management spbu, pp.316-325. [64] ougolnitsky, g.a. (2002), game theoretic modeling of the hierarchical control of sustainable development. in: petrosjan, l.a., mazalov, v.v. (eds). game theory and applications. vol.8. nova science publishers, n.y., pp.8291. [65] ougolnitsky, g.a. (2011), “sustainable management”, nova science publishers, n.y. [66] ougolnitsky, g.a. (2012), “game theoretic formalization of the concept of sustainable development in the hierarchical control systems”, annals of operations research, doi 10.1007/s10479-012-1090-9 [67] ougolnitsky, g.a., and usov, a.b. (2009), problems of the sustainable development of ecological-economic systems. in: cracknell, a.p., krapivin, advances in systems science and applications (2014) vol.14 no.4 345 v.p., varotsos, c.a. (eds.). global climatology and ecodynamics: anthropogenic changes to planet earth. springer-praxis, pp.427-444. [68] 8. boulanger p.-m., and brechet t. (2005), “models for policy-making in sustainable development: the state of the art and perspectives for research”, ecological economics, vol.55, pp.337-350 [69] cannon w. (1932), the wisdom of the body. c london. [70] coppel w.a. (1965), “stability and asymptotic behavior in differential equations”, heath mathematical monographs, boston. corresponding author author can be contacted at: ougoln@gmail.com adv syst sci appl 2020; 04; 45-59 published online at https://ijassa.ipu.ru. computational models of trajectory investigation of marine geophysical fields and its implementation for solving problems of map-aided navigation l.v. kiselev 1*, v.b. kostousov 2, a.v.medvedev1, a.e. tarkhanov2, k.v. dunaevskaya2 1) imtp feb ras, vladivostok, russia e-mail: imtp@marine.febras.ru 2) imm ub ras, ekaterinburg, russia e-mail: dir-info@imm.uran.ru abstract: comprehensive investigations of geophysical fields (gpf) of the ocean are among the top priority problems in underwater robotics. problems reside in developing automated systems capable of real-time operation for gathering, accumulating, and processing diverse geophysical information. the data’s total volume is applied for long-term or real-time monitoring of marine areas, environment or objects surveillance, mapping certain regions, or anomalies in geographical coordinates. one of the main elements of such systems is the computational models sensitive to properties of geophysical fields, features of search routes of underwater vehicles, and particular aspects of missions performed during trajectory measurements and features of navigational support tasks. the work considers computational models of comprehensive interpretation of the results of trajectory measurements of geophysical fields using an autonomous underwater vehicle (auv), and estimation of accuracy of map-aided navigation on the reconstructed map of geophysical fields. algorithms and software consider distinctive features of representation of bathymetry, magnetometry, and gravimetry data visualized in 2d and 3d. algorithm of mapaided navigation by the reconstructed fields is discussed. the final assessments of the considered models’ accuracy consider averaged errors of measurements, mapping, and inertial navigation. obtained assessments are based on theoretical investigations, results of model experiments, and experimental trials of auv’s systems under actual operating conditions. keywords: autonomous underwater robot (autonomous unmanned underwater vehicle), motion control, navigation, marine geophysical fields, trajectory measurements, mapping, computational models, bathymetry, magnetometry, gravimetry. 1. introduction nowadays, geophysical measurements using aerospace, ground, and marine means are widely used for precise navigation of moving objects and investigation of the earth’s geological structure. particular importance is attached to comprehensive research of anomalous fields in different regions of the earth’s surface, including aquatic areas promising in terms of its geological structure, exploration, long-term and real-time monitoring, and performing of various underwater missions. herewith, tasks measuring the parameters of fields with precise navigational referencing, 3d visualization of surveying results, bathymetry, magnetometry, and gravimetry of different marine areas are relevant for practical applications. prospects of using auvs for geophysical measurements at sea repeatedly served as subjects of research and development over the number of years. originally this topic was * corresponding author: imtp@marine.febras.ru 46 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) raised in the monograph [1] and concerned the estimation of advantages, which can be introduced using auvs. later, similar problems were discussed in the works of russian [217] and non-russian specialists [18-22]. bathymetry measurements always accompany echo-ranging surveys of the sea bottom terrain, usually performed using modern side-scan sonars, interferometric and multi-beam sonars, and synthetic aperture sonars. studying marine geomagnetic fields and its anomalies and solving magnetized object search and detection tasks are carried out using highly sensitive field-level meters and gradiometers. some issues of solving search tasks utilizing auvs equipped with magnetic anomalies meters were considered in works [1-6, 10]. the use of auvs seems promising for trajectory measurements and mapping of local gravity fields or gravity anomalies (ga) with the prospect of using the results for navigation purposes. an important tool, along with the creation of equipment and conducting field tests, is the development and research of information and computational models [23-24] of specific processes for measuring geophysical fields, their recording and mapping, including models for the use of auv and the use of mission results. gravimetric computational models and their capabilities in solving navigational adjustments tasks were studied in research [12-17]. problems of mapping geophysical fields using auv are essential by itself, and besides that are directly related to the solving tasks of navigation by distinctive gpf elements. these tasks may include the follows: • search for the position of a fragment of the field map, recorded during the auv movement, on the field reference map, which is related to the task of matching the fragment to the reference field; • the best approximation of the field map in terms of accuracy of navigation for economical storage of the reference field onboard the auv; • estimating the informativeness of the field and developing the most meaningful route, i.e., the best route in terms of the accuracy of navigation by the gfp while moving along the route [25, 26]. let’s take a closer look at possible solutions to this kind of task. 2. search tasks during mapping and reconstructing the map of anomalous geophysical field based on trajectory measurements from auv let be the function describing the cross-section of the field in a coordinate plane , which can be denoted by the isoline map the field meter, while moving alongside the trajectory, records a representation of the field with some random error (2.1) suppose that coordinates and projections of the speed vector of the auv (field meter) are given by an internal navigation system (ins) with errors then, coordinates related to the measurement of differs from true ones (2.2) let’s denote the field’s measure of the variability alongside the trajectory by the value of field gradient or an increment , where is the moving velocity, is the time interval. ( , ) f x y ( , ),x y ( , ) .f x y const= ( ) f t ( ) : tx ( ) ( ( ), ( )) ( ). f t f x t y t tx= + ( , )x y ( , ) x yv v ( , ), ( , ).а а x yx y v vd d d d ( ( ), ( ))x t y t ( ) f t ( ) ( ) .( ) , ( )) a x a yx x t x v t y t y v tt y t= + d +d = +d +d! ! fñ f f v td = ñ d v td computational models of trajectory investigation of marine geophysical fields 47 copyright ©2020 assa. adv. in systems science and appl. (2020) the following interrelated tasks are of practical interest: task 1 consists of recovery of the field map in the form of an intensity matrix, defined on the regular orthogonal grid, or as a set of isolines. the recovery of the field map must be performed by covering the given area by the net of trajectories and measurements (2.1) of the field with referencing to points (2.2.) obtained by inaccurate navigational data. as usual, the control of an underwater robot consists of an assignment and correction of a navigational program in characteristic points of the trajectory. the points can be derived from conditions of passing the field local extremums or condition of maximal field gradient: (2.3) task 2 consists of the search of field anomalies, which appear in the function f as areas of local extremums. for that, it is necessary to arrange a search motion trajectory for an underwater robot for detection and contouring the anomaly by the fixed level and variability of the field meter signal under the masking influence of external noises. the search of coordinates with specific (extremal in particular) values the field parameters is vital for the determination of purposed motion cue. the search procedure can be performed based on orthogonal or gradient descent with an orientation towards a chosen target, such as an anomaly source. accumulation of the data about the extremal values of the field allows outlining the borders of anomaly or boundaries of area (background) with no anomalies. task 3 is connected with tracking the given isoline delimiting the boundaries of the anomaly by the given (extremal) field value. the motion is then organized by the orientation of the speed vector by the field behavior proportional to the value . generally, the search mission scenario includes the following processes: • approaching the area with the given field level using orthogonal tacks and course correction according to field gradient sign; • covering the fount area by the net of squarewave like rectilinear trajectories (tacks) with continuous measurements of the field parameters and mapping them to navigational data for further reconstruction of the field map; • performing search movements by orthogonal descent for determination of the probable position of the extremum; • tracking the given isoline delimiting the boundaries of the anomaly by pointing the speed vector according to the “curvature” of the field alteration or moving by calculated coordinates, which corresponds to the given field level. the primary scenario process consists is a restoration of the field map be the net of discrete measurements obtained by covering the target area by rectilinear trajectories (tacks). measured in nodes of the irregular net values are then interpolated with step depending on the value of the field gradient value in nodes. based on obtained data, isolines with given values of the field can be created. motion control in the field of anomaly is a unique but yet practically important case. while developing the particular algorithms of auv motion control during field anomaly investigation, it is necessary to consider features of the spatial structure of the field, including the presence of the natural background field, time and spatial distortions, and significant nonlinearities in their formal description. examples of modeling the search motions and restoration of the map of investigating areas of the field are presented in [7-9, 11]. ( , )f x y ( , )k kx y { } .( , ) ( , )loc loc loc max min maxk k k k f f f ff x y f x yº º ñ º ñ! ! ( )tf ( )tfd ( , ) ,f x y const= /x yf fñ ñ 48 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) 3. features of the visual representation of geophysical information and results of its comprehensive processing as was mentioned before, the idea of using auv for solving tasks of marine geodesy and navigation appeared at the very beginning of their development and then continuously evolved in theoretical and applied studies. let’s consider some of the results of research and computational experiments in the field of marine gravimetry. first of all, it is worth mentioning that auv is a well-stabilized platform abled of compensating external forces’ actions, such as vertical accelerations of the carrier [2-4, 7-9]. examples of gravimeter signal records obtained by auv “tiflonus” when working on the depth of 70 m are shown in the work [14]. the procedure of recording and restoration of the map of the gravity anomalies (ga) field based on trajectory measurements while covering the given area by a planned grid of trajectories with predetermined measurement accuracy are illustrated in fig. 3.1. here, particular estimations use a geographical map of residual gravity anomalies in the the peter the great gulf, obtained in the laboratory of gravimetry of the v.i. il'ichev pacific oceanological institute (poi) feb ras using shipboard precision gravimetry chekan-am made by csri elektropribor [10]. the map contains a selected fragment, which has a ceased volcano with a volatile ga gradient in the lower part of it (fig. 3.1,a). a 50x40 km part of the field in this area was investigated using squarewave-like trajectories (fig. 3.1,b). the computational experiment assumed that the processing time of the small tack is 20 min, which equals to the distance of 1.2 km on the speed of 1 m/s. correspondingly, the processing time of the big tack is 14 hours, and the distance is 50 km. trajectories were created considering auv dynamics in modes of depth and heading control. obtained digital data then were used for the restoration of the field map using scripts of matrix transformation and functions of interpolation for 2d and 3d images (fig. 3.1,b,c). (c) fig. 3.1. the fragment of the ga geographical map with plotted squarewave-like trajectories (a); the 3d image of the recovered field fragment (b); a horizontal slice of the ga map recovered fragment (c) a subject of undoubted interest is extensive use of all the available geophysical information, including bathymetry, magnetometry, and gravimetry. it allows firstly to find a possible correlation in measurements of heterogeneous fields, and secondly to obtain more certain and precise data for navigational correction by the cumulative results of mapping. let’s have a closer look at the possibility of such an approach on the example of combining the data of mapping of sea depth field, geomagnetic field, ga field for the predetermined aquatic area. we will use an integrated mapping data for the studied area of the strait of tartary provided by the laboratories of gravimetry and magnetometry of the feb ras. these initial maps originate from the expedition using high-precision shipboard computational models of trajectory investigation of marine geophysical fields 49 copyright ©2020 assa. adv. in systems science and appl. (2020) measurement devices. there are no similar available maps obtained with the help of auv. therefore, real maps are used, on which it can work out all the procedures simulating the trajectory measurements results of the field from the board of auv. recovered maps are virtual maps obtained using computational procedures based on the achievable results of trajectory measurements from the auv, taking into account methodological errors (model errors) and measurement errors. these maps are actually analogs of those maps that can be obtained by using of auv in real work. further these maps are used as a reference ones to evaluate the accuracy of navigation by the gfp of certain moving objects, in the structure of which the map data is embedded. in the accepted computational model these maps are used to work out the evaluation algorithm described below. the initial 2d maps of these three fields are demonstrated in fig. 3.2,a,b,c, and the reconstructed maps on fig. 3.2,d,e,f. two-dimensional images of three fields shown in fig. 3.2 can be transformed into 3d images in the form of corresponding reliefs of the spatial structure (an example of it can be fount in fig. 3.1,b). additional capabilities of the visual representation in the form of threedimensional images can be obtained by a combination of the ga field, anomalous geomagnetic field (agmf), and the sea depth field (sdf). here, the sea depths axis is a vertical coordinate, and the other two form a projections of their 3d images on the upper (zero) level of the sea terrain field. such representation can be got in the form of images obtained from any angle and any step of the image rotation angle. note that according to the bathymetry matrix, the average depth of the selected area in the tatar strait is 2737 m. fig. 3.3 demonstrates examples of this visual representation with the covering the entire zone with squarewave-like trajectories when auv is in motion on the depth of 1000 m. during movement along a given trajectory, simultaneous and spatially combined measurements of three fields are made. a) b) c) d) e) f) fig. 3.2. fragments of the initial maps of the sea depths (а), ga field (b), agmf (c), map of the recovered fragmets of the sea depths field (d), ga field (e), agmf (f) 50 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) a) b) fig. 3.3. examples of the angle pictures combining the sea bottom’s terrain field with ga field (a) and with the field of magnetic anomalies (b). the redline shows the coverage trajectory. the images are obtained using matlab graphic toolkit (functions surf(), mesh(), plot3()), where one space unifies data about physical fields of two types and trajectories with different color schemes (colormap()) 4. estimation of the navigational informativeness of an anomalous geophysical field works [10-16, 27-29] consider issues of the informativeness of the anomalous ga field and estimating the accuracy of navigational correction by recovered field map. here the estimations of navigational informativeness of the fields (sdf, ga fields, and agmf) recorded in the tatar strait are considered. thereunder, the method of estimation of errors of map-aided navigation [32-35] proposed in [30, 31] is studied by the example of the joint using of three geophysical fields. the task of the method of map-aided navigation is correction of errors of required navigational parameters (for example, coordinates (latitude and longitude) of the carrier) by the matching the measured field fragment to the reference map [33, 34, 36, 37]. the matching of the fragment and the reference map can be performed using the search of the global extremum of some matching functional of measured fragment and reference fragments obtained from the reference map. the navigational informativeness of the geophysical field is measured by the level of navigation errors when using this method. as a rule, the main source of errors is misclosures of mapmaking and errors in field measurements while the object is moving. however, the sensibility of the navigation results to these errors is mainly determined by the gradient characteristics of the field itself. here, the task is considered in the probabilistic statement, as in works [33-35]. computational models of trajectory investigation of marine geophysical fields 51 copyright ©2020 assa. adv. in systems science and appl. (2020) 4.1. theoretical estimation of errors of coordinates correction by the maps of geophysical fields since the map-aided method of navigation is used for the correction of errors of inertial navigation system, the task the navigation by geophysical field is often called the correction task of navigation system. here, the task of coordinates error correction in a horizontal plane is considered. let’s denote the reference map of the geophysical field defined in a certain rectangular area of the plane by . formally, the field map is a smooth function of the two-dimensional vector with values on the axis of real numbers. practically, the function is represented by samples in nodes of a regular 2d grid with a given step size. if necessary, the values of the function between the nodes can be obtained using a proper interpolation method, for example, bilinear interpolation [38, pp. 123-128]. let’s consider that measurements of the field are performed in points of a rectilinear interval, which begins in the unknown point in a given fixed direction defined by a unit vector . the model of the field measurement onboard the moving object can be expressed as: (4.1) where is the value of the geophysical field measured in the discrete moment . then, let’s denote the set of values (m-vector) by a measured fragment of the field or just a fragment. the expression (4.1) consists of the following notions: • is an area of a priori location of the initial point of the fragment in the area of the given reference map ; • is the actual location of the fragment; • is a given directing unit vector of the measurement route; • is a vector of the fragment distortion. equality (4.1) can be rewritten in the vector form as: (4.2) where m-vector is a set of samples recorded from the field map in time points : (4.3) at its simplest, the vector is a centered gaussian vector with a zero mean and covariance matrix (where is an identity matrix, is a given dispersion). in reality, the distortion of the fragment along with the field meter random noises include the inaccuracies of mapping and errors of the relative positions of measurement points, which are caused by navigational errors of the carrier. in the correction task, it is required to obtain the best estimation of the unknown vector x using given initial data and evaluate the accuracy of this estimation, for example, using the covariance matrix . in the framework of the bayesian method, vectors x and are assumed to be random vectors with known distribution , . here, the main research subject is a correlation-extremal search method of the navigational error correction [24-26], which is added with the new approach for the correction error estimation based on the analysis of matching functional. w 2 1 2( , ).r x x= x ( )g x ( )g x x ( )g x ( )g × x p ( ) ( ) , , 1,... ,k k k kt g t x q k mj j x= + î =! x x + p kj kt 1( ,... )tmj j=φ g q ìw qîx φ w g x p 1( ,... )mx x=ξ ( ) ,= +φ s x ξ s(x) kt 1( ) ( ( ),... ( )) .tmg t g t= + +s x x p x p ξ 2 xs i i 2 xs ( )!x φ ( )p φ 1( ,... )mx x=ξ ( )xf x ( )fx ξ 52 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) let’s consider the following quadratic functional of matching measurements and field map as an example: (4.4) in this case, the estimation can be found as a result of the search of minimum: : (4.5) the calculation of the correction error estimation relies on the following inequality: , (4.6) where is a point of the global minimum, is a point of the second-biggest local minimum spaced apart from the first minimum on a significant distance, is a threshold value. in the case of a single extremum, it is possible as take the minimum value of the functional at the boundary of the uncertainty domain. inequality (4.6) shows how strongly the global minimum in (4.5) is pronounced against other local minimums, as the smaller the relation in (4.6), the more reliably the correction problem is solved in errors presence. in this context, this inequality describes the informativeness of the field in the . the threshold is derived from a statistical experiment using the neumannpearson criterion and corresponds to the minimal level of the errors of the second kind at a fixed level of the errors of the first kind. here the errors of the first kind are understood to be taking a wrong algorithm result as an accurate correction, and the errors of the second kind are understood to be a false rejection of the correction while the algorithm result is correct. in a statistical experiment, the correctness of algorithm output was defined by exceeding a predetermined threshold of the distance between the actual location of the fragment and the solution of the task (4.5). at a predetermined threshold , it is possible to calculate the threshold and then to find a diameter of the next set: (4.7) in this case, the diameter of the set is defined as the maximum distance between the points of the set. thus, the result of the correction algorithm is taken as an estimate (4.5), and the value (4.7) serves as an estimate error of the search algorithm, which can be calculated during the process work of the algorithm in real time without knowing the actual position x. in this case, the value of is an estimate of the maximum possible radial error of the correction. upon the availability of several spatially connected maps of different geophysical fields and a possibility of simultaneous measurement of these fields during motion, it is possible to co-use them to improve the accuracy of navigational error correction using a search algorithm (4.5). to do so, we will consider the matching functional of the following form: (4.8) where k is a number of used geophysical fields, is a map of the i-th field, is the kth sample of the measured fragment of the ith field, is a weight influence coefficient of the measurements of the i-th field on the matching functional. [ ]2 1 ( ) ( ) m k k k g t = f = + -åx x p j ( )!x j ( )f x ( ) argmin ( ) qî = f! x x xj min_1 min_ 2 ( ) ( ) trp f < f x x min_1x min_ 2x min_1x trp min_ 2( )f x q trp trp min_1( ) tr tr x p f f = max { : ( ) }.trd diam x x= f £f ( )!x j maxd / 2max maxdr = computational models of trajectory investigation of marine geophysical fields 53 copyright ©2020 assa. adv. in systems science and appl. (2020) further coefficients 𝛼�coefficients � are accepted to be equal to inverse values of dispersions of summary errors of the measurement and the reference map: (4.9) 4.2. results of computational experiments the most accurate and practically usable method of assessing the informativeness of geophysical field maps is a method of statistical imitational modeling of the process of navigational error correction, including repeated forming the sequence of measurements and their referencing to the reference field maps. recovered maps of three fields recorded in the tatar strait in coordinates e137º…e138º, n43,8º...n44,3º (fig. 3.2,d-f) were used as reference maps. maps were developed as matrices of samples assigned in nodes of the regular net with a step of 0.001° on both axes. herewith, experiments involved forming of model measurements of the field by the initial maps of corresponding fields. experimental conditions in correction zones for every three fields (their maps can be found in fig. 3.2), we perform the forming of the model measurements of fields with the step of 100 m alongside rectilinear routes of 40 km directed along an x-axis. the correction was made by the recovered maps, shown in fig. 3.2,d,e,f. confidence rectangle q (i.e., an area of the a priori location of the route beginning point, where a search of the functional minimum was performed (4.5)) had a size of 48x55 km. the size of the researched zones was approximately 85×55 km. the grid of the experiment containing the nodes, where the measurement paths began, covered the confidence rectangle with a step of 0.01 degrees (about 1 km). there were nэ = 2880 total experiments of correlation-extremal correction of navigation errors performed for each map. in each experiment, a set of measurements was modeled on the initial map, and the functional was formed (4.4) or (4.8) based on the reconstructed maps, the search for the functional minimum and the calculation of the radial -estimate were performed. the full generation model of the measurements of the sdf, ga field, and agmf included errors of autonomous navigation, random errors of the meters: echosounder, gravimeter, magnetometer correspondingly. the approach for error generation of the ga field measurements and inertial navigation system is provided in work [14]. fig. 4.1 comparatively demonstrates the results of statistical experiments of modeling the correction and obtaining the estimation of correction errors . the results in fig. 4.1 have the form of matrices with a size of 55×85 according to the grid in increments of one kilometer covering the research area. herewith, ga field experiments specified the rms deviation of the fluctuating part of the measurement error = 3 mgal, the computational threshold of the estimation was chosen to be =0.8. for the sdf, these parameters were taken equal to = 60 m, =0.6. and for the magnetic field, they were equal to = 7 nt, = 0.9. sdf: actual error sdf: error of dmax estimation (a) (d) maxd maxd xs maxd trp xs trp xs trp 54 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) ga field: actual error ga field: error of dmax estimation (b) (e) agmf: actual error agmf: error of dmax estimation (c) (f) fig. 4.1. a comparison between distributions of actual errors along the correction zone (a,b,c) and estimations using dmax (d,e,f). the predetermined green-red color scale (the scale is shown on the left in each figure) shows the values of the actual radial errors of coordinate correction and their -estimations related to the initial points of the measurement routes located in the nodes of the kilometer grid. one of the measurement routes is shown in the center of every figure. all distributions are shown against the background of the corresponding halftone images, which represent the average radial modulus of the gradient of the study area in the gray scale (on the right). all the matrices shown in the figures correspond to the 85×55 km zones of the experiments. in the presented results, the estimation error vector for convenience is brought to the scalar value – radial error , where , are the squares of the vector coordinates. fig. 4.2 uses the same color map and shows distributions of the actual error of the comprehensive correction along with all three fields according to the algorithm (4.7) and distribution of the corresponding -estimation with the threshold =0.8. (a) combining fields: actual error (b) conbining ields: dmax-estimations fig. 4.2. a comparison between the distribution of the actual errors along the correction zones (a) and dmaxestimation (b) when correcting by all three fields. the given colormap shows the values of coordinate correction errors and their dmax-estimates related to the initial points of the measurement routes located in the nodes of the kilometer grid. one of the measurement routes is shown in the center of every figure. the distributions of the errors are shown against the matrix of the integral informativeness, obtained by the weighted sum of the matrices of component fields. the statistical mean values and statistical rms of estimation (4.7) of the correction errors of the considered geophysical fields and their combination are shown in table 4.1 in comparison with the actual errors for different locations of the fragment (nэ = 2880) and for different rms deviations of measurement errors . maxd ( )= !δ x x φ 1 2 2 2( ) x xr d d= +d 1 2 xd 2 2 xd δ maxd trp maxd qîx ξ computational models of trajectory investigation of marine geophysical fields 55 copyright ©2020 assa. adv. in systems science and appl. (2020) table 4.1. results of a statistical experiment for the separate fields and their combination correction by the ga field: = 3 mgal, ptr = 0.8 correction by the sea depths field: = 60 m, ptr = 0.6 correction by the magnetic anomalies field: = 7nt, ptr = 0.9 comprehensive correction by three fields: ptr = 0.8 actual error: / , m 980 / 811 761 / 752 1223 / 1015 671 / 737 radial estimation: / , m 1300 / 837 1175 / 710 966 / 842 737 / 528 in the table, we accepted the following notations: average actual radial error of the search algorithm, where – solution of task (4.5) in the i-th experiment, – the actual position of the estimated parameter in the i-th experiment, the expression described above; – statistical standard deviation of the actual radial error from its mean value; the average value of the error estimate of the search algorithm according to the method (4.7), где – where is the radial estimation in the i-th experiment; – statistical standard deviation of the radial – estimation from its mean value. matching the results shown in table 4.1 and distributions of the errors shown in fig. 4.1, and fig. 4.2 allow concluding that -estimation is close to the actual estimation of the correction accuracy, which proves the applicability of the method for solving the tasks of this kind. regarding the navigational informativeness of considered geophysical fields, it is mentioning that the most informative (or the most accurate in terms of correction) field is an sdf, and the least informative is an agmf. this actually shows the property of corresponding measuring systems and methods of measurement processing. it can be summarized that the complexing the measurements in three fields significantly increases the average accuracy of map-aided navigation: for the sea depths field of at least 10%, for the ga field of at least 30% and for anomalous geomagnetic field of at least 45%. 5. conclusion 1. the preceding experience of works connected with the use of auv for real-time and long-term monitoring of the geophysical fields of the ocean proves its viability for solving tasks of this kind. but on the other side, it demonstrated a number of theoretical and practical issues regarding the development of the proper computational models. these xs xs xs sd srms maxd maxr max rmsr 1 1( ( ) ) ( ( ) ) эn s s s i i iэ r r x x n d j = = = -åx φ x! ! ( )s ix j! ix ( )r × srms ( ( ) )s ir x xj -! max max, 1 1 эn i iэn r r = = å max,ir maxd max rmsr maxd maxd 56 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) models are based on methods of trajectory investigation of the areas and objects with respect to the features of the motion control tasks, information properties of the fields itself, and precision properties of the means of navigational correction by the recovered field map. nowadays, gravimetry tasks are of particular importance thanks to the opportunity to perform it onboard the auv, which is the precision stabilized platform with a minimum of the vertical accelerations. that establishes close relationships with tasks of inertial navigation. the results of natural and computational experiments show that this is a way for providing high-precision gravimetric measurements with the cumulative error of the ga field mapping of units of mgal, which corresponds to the geographical referencing error of units of meters. 2. critical issues of the described research are: • development of the models of trajectory investigation of the fields, anomalies search, and reconstruction of the map using the results of trajectory measurements; • visual representation of the recovered fragments of the map in 2d and 3d formats and estimate the map reconstruction’s accuracy regarding the dynamic errors, measurement, and navigation errors; • development of computational models for estimating the information properties of reconstructed fragments of the field map for solving the task of navigational correction using a map-aided navigation approach. 3. accuracy evaluation of correlation-extreme correction by the maps of geophysical fields is considered using specially developed a new method, which does not impose excessive requirements for onboard computer. 4. one of the main advantages of auv equipped with the complex of meters of the geophysical fields is the joint processing of all the available information for the mapping and following use of it for navigational correction. using software processing and 2d and 3d visual representation of digital formats, it is possible to find a correlation in the fields’ scalar-vector properties, which can be an additional means for developing more precise computational models of map-aided navigation. references 1. ageev m.d., kasatkin b.a., kiselev l.v., molokov yu.g., nikiforov v.v., rylov n.i. (1981). avtomaticheskie podvodnye apparaty [automated underwater vehicles]. leningrad: sudostroenie, 223, [in russian]. 2. ageev m.d. (1994). auv – a precise platform for underwater gravity measurement. proc. of ieee oceans’94, brest, france. 3. ageev m.d, kiselev l.v., matvienko yu.v., etc. (2005) avtonomnye podvodnye roboty. sistemy i tekhnologii [autonomous underwater robots. systems and technologies]. moscow: nauka, 400, [in russian]. 4. ageev m.d. (2009) avtonomnyj podvodnyj apparat – ideal’naya precizionnaya platforma dlya podvodnyh gravimetricheskih izmerenij [autonomous underwater vehicle – an ideal precise platform for underwatger gravimetric measurements]. podvodnye issledovaniya i robototekhnika [underwater research and robotics], 1(7), 4-8, [in russian]. 5. kiselev l.v. (2009) upravlenie dvizheniem avtonomnogo podvodnogo robota pri traektornom obsledovanii fizicheskih polej okeana [motion control of autonomous underwater robot during trajectory measurements of the oceans physical fields]. avtomatika i telemekhanika [automatics and telemechanics], 4, 141-148, [in russian]. 6. kiselev l.v. (2011) kod glubiny [depth code]. vladivostok: dal'nauka, 331, [in russian]. computational models of trajectory investigation of marine geophysical fields 57 copyright ©2020 assa. adv. in systems science and appl. (2020) 7. kiselev l.v., medvedev a.v. (2011) modeli dinamiki i algoritmy upravleniya dvizheniem avtonomnogo podvodnogo robota pri traektornom obsledovanii anomal’ nyh fizicheskih polej [dynamic models and motion control algorithms of the autonomous underwater robot during trajectory measurements of the physical fields]. podvodnye issledovaniya i robototekhnika [underwater research and robotics], 1(11), 24-31, [in russian]. 8. kiselev l.v. (2017) prioritetnye zadachi podvodnoj robototekhniki v oblasti morskoj geodezii [priority tasks of underwater robotics in the marine geodesy]. proc. of the xi conf. «analiticheskaya mekhanika,utojchivost’ i upravlenie» [proc. of the xi conf. «analytical mechanics, adaptability and control»], kazan’, russia, 138-151, [in russian]. 9. inzarcev a.v., kiselev l.v., kostenko v.v., matvienko yu.v, pavin a.m.,shcherbatyuk a.f. (2018) podvodnye robototekhnicheskie kompleksy: sistemy, tekhnologii, primeneniya [underwater robotic complexes: systems, technologies, applications]. vladivostok: dal’press, 368, [in russian]. 10. peshekhonov v.g., stepanov o.a., avgustov l.i., etc. (2017) sovremennye metody i sredstva izmereniya parametrov gravitacionnogo polya zemli [modern methods and tools of measuring the parameters of the earth's gravitational field.]. csri elektropribor, itmo university. pod obshchej red. v.g. peshekhonova; nauch. redaktor o.a. stepanov. sankt-peterburg, [in russian]. 11. kiselev l.v., medvedev a.v., hmel’kov d.b. (2017). o nekotoryh zadachah podvodnoj robototekhniki v oblasti morskoj geofiziki [about some tasks of underwater robotics in marine geophysics]. proc. of the 7 conf. «tekhnicheskie problemy osvoeniya mirovogo okeana» [proc. of the 7 conf. «technical problems of the development of the world ocean»], 386-393, [in russian]. 12. kostousov v.b., tarhanov a.e. (2017). ocenka informativnosti geofizicheskogo polya s tochki zreniya korrelyacionno-ekstremal’ noj navigacii [estimates of the informativeness of geophisical field in perspectives of map-aided navigation]. proc. of the 7 conf. «tekhnicheskie problemy osvoeniya mirovogo okeana» [proc. of the 7 conf. «technical problems of the development of the world ocean»], 394-398, [in russian]. 13. kiselev l.v., medvedev a.v., kostousov v.b., tarkhanov a.e. (2017) autonomous underwater robot as an ideal platform for marine gravity surveys. proc. of the 24-th saint petersburg international conference on integrated navigation systems. csri «elektropribor», 605-608. 14. kiselev l.v., kostousov v.b., medvedev a.v., tarhanov a.e. (2019). o gravimetrii s borta avtonomnogo podvodnogo robota i ocenkah eyo informativnosti dlya navigacii po karte [about gravimetry onboard the autonomous underwater robot and estimation of its applicability to map-aided navigation]. podvodnye issledovaniya i robototekhnika [underwater research and robotics], 1(27), 21-30, [in russian]. 15. vitalii i. berdyshev_ lev v. kiselev, victor b. kostousov. (2018) mapping problems of geophysical fields in ocean and extremum problems of underwater objects navigation / ifac papers on line 51-32, 189–194. 16. kiselev l.v., kostousov v.b. (2019) on interrelation and similarity in solution of navigation and gravimetric tasks in underwater robotics. proc. of the 26-th saint petersburg intern.conf. integrated navigation systems. csri "elektropribor". 395397. 17. kiselev l.v., v.b.kostousov, k.v. dunaevskaya, a.e.tarhanov. (2020) ocenka oshibok korrelyacionno-ekstremal’ noj navigacii po karte anomalij sily tyazhesti na osnove traektornyh izmerenij s borta avtonomnogo podvodnogo robota [error estimation of map-aided navigation by the map of gravity anomalies based on 58 l.v. kiselev, v.b. kostousov, a.v. medvedev, a.e. tarkhanov, k.v. dunaevskaya copyright ©2020 assa. adv. in systems science and appl. (2020) trajectory measurements onboard autonomous underwater robot]. podvodnye issledovaniya i robototekhnika [underwater research and robotics], 1(31), 13-20, [in russian]. 18. zumberge m.a., hildebrand j.a., stevenson j.m. (1991) submarine measurements of the nevtonian gravitational constant // physical review letters. vol. 67. iss. 22. 19. zumberge m.a., sasagawa g., zimmerman r., ridgway j. (2010) autonomous underwater vehicles borne gravity meter. patent application pub.no.: us 2010/0153050 a1, jun.17. 20. zumberge m.a., sasagawa g., zimmerman r., ridgway j. patent no. us 2010/0153050, a1. (jun.17, 2010) autonomous underwater vehicles borne gravity meter. 21. kinsey j., tivey m., yoerger d. (2009) toward high-spatial resolution gravity surveying of the mid-ocean ridges with autonomous underwater vehicles. whoi deep ocean exploration institute and a whoi green innovation technology award. 22. melo j., matos, a., (2017) survey on advances on terrain based navigation for autonomous underwater vehicles, ocean engineering, vol. 139, 250–264. 23. calder m et al. 2018 computational modelling for decision-making: where, why, what, who and how. r. soc. open sci. 5: 172096. http://dx.doi.org/10.1098/rsos.172096 24. gost r 57412-2017 komp'yuternye modeli v processah razrabotki, proizvodstva i ekspluatacii izdelij. obshchie polozheniya (gost r 57412-2017 computer models of products in design, manufacturing and maintenance. general), [in russian]. 25. stepanov o.a., nosov a.s., toropov a.b. (2018) navigacionnaya informativnost' geofizicheskih polej i vybor traektorij v zadache utochneniya koordinat s ispol'zovaniem karty [navigational informativeness of geophysical fields and trajectory selection in the problem of coordinate refinement using a map]. proceedings of the tula state university. technical sciences. № 5, 74-92, [in russian]. 26. berdyshev v.i., kostousov, v.b. (2007) ekstremal'nye zadachi i modeli navigacii po geofizicheskim polyam [extreme problems and models of navigation in geophysical fields]. ekaterinburg: ub ras, 2007. p. 270, [in russian]. 27. zheleznyak l.k., koneshov v.n. (2007) izuchenie gravitacionnogo polya mirovogo okeana [investigations of the gravitational field of the world ocean]. vestnik ran, 77(5), [in russian]. 28. koneshov v.n., nepoklonov v.b., avgustov l.i. (2016) ocenka navigacionnoj informativnosti anomal'nogo gravitacionnogo polya zemli [estimation of navigational informativeness of the earth's anomalous gravitational field]. giroskopiya i navigaciya [gyroscopy and navigation]. t. 24. № 2 (93). 95-106, [in russian]. 29. kopytenko yu.a., petrova a.a., avgustov l.i. (2017) analiz informativnosti magnitnogo polya zemli dlya avtonomnoj korrelyacionno-ekstremal'noj navigacii [analysis of the informativeness of the earth's magnetic field for autonomous mapaided navigation]. fundamental and applied hydrophysics [fundamental and applied astrophysics]. t. 10. № 1. 61-67, [in russian]. 30. kostousov v.b., dunaevskaya k.v. (2018). metod korrekcii navigacionnyh oshibok po polyu vysot ob “ektov mestnosti [methods for navigation error correction using the field of terrain objects height field]. proc. of the xxxi conf. in memoriam of n.n.ostryakova, 218-227, [in russian]. 31. kostousov v.b., tarhanov a.e. (2019) novyj metod ocenki oshibok korrekcii koordinat po karte geofizicheskogo polya [new estimation method for coordinates error correction by the map o geophysical field]. proc. of the 8 conf. «tekhnicheskie computational models of trajectory investigation of marine geophysical fields 59 copyright ©2020 assa. adv. in systems science and appl. (2020) problemy osvoeniya mirovogo okeana» [proc. of the 8 conf. «technical problems of the development of the world ocean»]. 347-351, [in russian]. 32. beloglazov i.n., dzhandzhgava g.i., chigin g.p. (1985). osnovy navigacii po geofizicheskim polyam [basic of navigation by the physical fields]. moscow: «nauka», 328, [in russian]. 33. stepanov, o.a, toropov, a.b.: nonlinear filtering for map-aided navigation, part 1: an overview of algorithms. gyroscopy and navigation, 6(4):324–337 (2015) 34. stepanov, o.a, toropov, a.b.: nonlinear filtering for map-aided navigation, part 2: trends in the algorithm development. gyroscopy and navigation, 7(1):82–89 (2016) 35. stepanov o.a. (2017). osnovy teorii ocenivaniya s prilozheniyami k zadacham obrabotki navigacionnoj informacii. chast’ 1. vvedenie v teoriyu ocenivaniya [basics of estimation theory with applications to the tasks of navigational information processing. part 1. introduction to estimation theory]. saint-peterburg: csri «elektropribor», [in russian]. 36. stepanov o.a., nosov a.s. a map-aided navigation algorithm without preprocessing of field measurements gyroscopy and navigation, 2020, 11(2), 162– 175. 37. nosov, a.s., stepanov, o.a. (2018) the effect of measurement preprocessing on the accuracy of map-aided navigation. 25th saint petersburg international conference on integrated navigation systems, icins 2018 proceedings, 1–4. 38. press, william h.; teukolsky, saul a.; vetterling, william t.; flannery, brian p. (1992) numerical recipes in c: the art of scientific computing (2nd ed.). new york, ny, usa: cambridge university press. advances in systems science and applications (2014) vol.14 no.4 346-360 research and development of process knowledge management system in okp company w. l. chenli1 and s. q. (shane) xie2 1department of mechanical engineering,university of auckland, auckland,new zealand, f. f. zeng 2school of mechanical science and engineering, huazhong university of science and technology, wuhan, china abstract the coming of knowledge-based economic era has greatly changed manufacturing environment. knowledge has become the most important and valuable asset for manufacturing companies. for most of one-of-a-kind production (okp) companies, innovation based on knowledge is their opportunity and the sharpest competitive weapon to compete and thrive in current global market-driven environment. process knowledge is a special type of knowledge that controls how products are manufactured. the knowledge may help okp companies manufacture high value-added products with best quality and shorter times at competitive cost. managing process knowledge is very important, and its benefits are enormous. this paper proposes a framework and definition of process knowledge management system (pkms) after analysis of the characteristics of process knowledge and requirements in okp companies. a hybrid approach aiming at representation of three main types of process knowledge is discussed. one of the ways for representing dynamic process flow knowledge is introduced in detail, which is solved using an approach based on parameter flow charts (pfc). the proposed approach supports experts to create knowledge repository themselves. a prototype system is developed to validate the proposed framework and approach. keywords knowledge management, knowledge representation, one-of-a-kind production, knowledge-based engineering, process planning 1 introduction we are now in a knowledge-based economic era[1], in which manufacturing situation has changed greatly and can be characterized by the following buzzwords: customer-oriented, globalization, and time-driven competition. this is not only a challenge, but also an opportunity for manufacturing companies, especially for one-of-a-kind production (okp) companies to compete and thrive in this global market environment. these okp companies might be small and middle-sized, but their survival and success depend on producing newer, better, quicker and more innovative products and services, not simply rely on size and strength as before. innovation plays an important role in their strategies and is becoming advances in systems science and applications (2014) vol.14 no.4 347 one of the sharpest competitive weapons for them to manufacture high quality products at lower costs with shorter development times. the foundation and source of innovation is knowledge. knowledge and intellectual capital have become the most important and valuable asset for organizations. this has attracted much research interest in the areas of knowledge based engineering (kbe) and knowledge management (km) over the past years [2-7]. process knowledge in manufacturing industry is a special type of knowledge. most of them are the kernel technical secret that may help companies manufacture high value-added product with best quality and shorter times at competitive cost. the process knowledge may exist in the mind of experts or skilled employees. survey showed that to train an expert or a skilled employee of process planning might take at least 5 years or more. for most okp companies, the shortage of knowledge workers is one of their biggest obstacles to grow and the problem has to be solved. hence, exploiting and managing process knowledge is very important for okp companies and its benefits are enormous. process knowledge management system (pkms) can help okp companies develop, build, and deploy process knowledge assets effectively and profitably. considerable research effort has been placed in this area [8-13]. however, most of existing work are based on fixed process knowledge representation schema and fixed system architecture. they were applied to a particular kind of manufacturing environments and are very specific, and are difficult to adapt to continuous technology developments and changing manufacturing environment. to resolve these issues, a tiered framework of pkms is proposed in this paper. a hybrid and open approach for process knowledge representation is discussed in detail. the prototype version of system has been developed and demonstrated with case studies. 2 definition and framework of pkms 2.1 analysis of process knowledge in okp companies before discussing the framework of pkms, the characteristics of process knowledge in okp companies are analyzed firstly. for most okp companies, their market strategies are either make-to-stock, or assembly-to-order, or make-toorder. they are facing special critical problems including high customization, ‘once’ successful approach, loose or fatter production, and complicated product data and information flow, etc[14]. this determines their manufacturing environment is customer-oriented and has been undergoing changes often. hence, process knowledge in these okp companies also has special different characteristics that can be summed up as followings: (a) dynamic and changing. process knowledge should be updated along with the emergence of new techniques and new methods. 348 w. l. chenli: research and development of process knowledge management system ... (b) form diversity. there is a variety of process knowledge format, including description sentence, tables, pictures, drawings, formulas, and numerous other representations. (c) universality. process knowledge comprises of not only internal knowledge in a company, but also external knowledge of suppliers and customers. internal knowledge includes all the knowledge from market survey to product maintenance after sale. (d) uncertainty. process knowledge comes from experiences and skills of experts, and it is implicit and hard to represent by computer. 2.2 definition and framework of pkms the definition of pkms here refers to a system based on information technology for managing process knowledge in okp companies, supporting creation, capture, representation, storage, dissemination and share of information. pkms aims to enhance the innovation ability of okp companies by accumulating and sharing of personal, organizational, and industrial process knowledge. based on above analysis, the most important requirements and characteristics of pkms are: (a) supporting management of all kinds of process knowledge. not only the knowledge based on rules, but also the knowledge embedded in drawings, documents, formulas and other representations needs to be managed. (b) open and extensible. when manufacturing situation and technology change, the knowledge database can be updated and reconstructed. (c) people oriented. the source of knowledge is people. the knowledge database in pkms should be maintained and used by these skilled employees, not just by other it employees or system developers. companys own experts should be able to create, capture and storage their process knowledge in pkms. (d) combination of technological advancement and practicability. as process knowledge is tacit and implicit, its difficult to capture and transform this kind of knowledge clearly and automatically, new methods need to be developed. based on the requirements and characteristics mentioned above, a proposed framework of pkms has been constructed and is shown in fig.1. the framework includes five layers. each layer is devoted to a specific task and is designed to link the interface with the framework from different levels of abstractions and functionalities. (1) support layer: it provides the necessary computer net environment that supports the system running. support layer includes a computer network, a database management system (dbms), operation systems (os), desk pcs, and servers. (2) database layer: it is the core knowledge repository of pkms that stores all kinds of process knowledge of a okp company. the process knowledge is classified into four main types. process resource knowledge refers to information of advances in systems science and applications (2014) vol.14 no.4 349 fig.1 framework of pkms in okp companies machine tools, fixtures, cutters, materials, and quantity standard used in cutting, etc. product process knowledge is typical process knowledge used regularly in companies.feature process knowledge includes recommended process routes of all sorts of machining features. other process knowledge refers to other often used process planning information, such as basic concepts in process planning, formulas of working hour calculation, etc. (3) core layer: it is comprised of knowledge representation, knowledge discovery, and knowledge reasoning and explaining. knowledge representation is to provide environment and methods for representing various types of process knowledge. knowledge discovery is used to search for appropriate process knowledge in repository. knowledge reasoning and explaining deal with the search result and give helpful output process knowledge. (4) function layer: it is the functional implement layer for the purpose of process knowledge management. function modules in this layer include the creation, presentation, modification, searching, and browsing of all the process knowledge stored in the database. for the purpose of managing diverse forms of process knowledge, a process resource management module, a process flow knowledge definition module, and a formula management module are developed. (5) application layer: this layer provides industrial solutions for a given com350 w. l. chenli: research and development of process knowledge management system ... fig.2 design support strategy for pkms pany or industry, such as equipment, machine tools, mould, boat and ship, etc. the knowledge repository is deployed flexibly according to the requirements of a particular company or industry. 2.3 design support strategy for pkms according to the proposed system framework as shown in fig.1, the design support strategy is proposed to implement pkms as shown in fig.2. the following summarizes the key issues of the design support strategy. (1) requirements investigation. to know the current situation and requirements of process knowledge management in okp companies, collect process knowledge of personal, company and industry. the investigation results are used as input and foundation for pkms development and implementation (2) technical solutions planning. based on the result of investigation, the technical route, function models, and task schedule for pkms are planned and determined. (3) knowledge establishment. classify all kinds of collected process knowledge, create, modify, update knowledge database using function modules of system. (4) application framework deployment and releasing. configure pkmss running environment according the situation of a particular company and release it. (5) system deployment and implementation. establish implementation team, install and deploy the system, train companys employees to use the system. (6) system assessment and continuous optimization. after running for a period of time, collect feedback, assess, improve and maintain the system. to develop this proposed pkms, one of the key issues is the representation and management of a variety of process knowledge. as mentioned above, the diversity and dynamic characteristics of process knowledge makes it hard to be represented effectively through only one approach. a hybrid and open approach for process knowledge representation is required. advances in systems science and applications (2014) vol.14 no.4 351 3 hybrid approach for representation of process knowledge 3.1 classification of process knowledge process knowledge is classified into three main types (see fig.3) according to the form of process knowledge described as following: (1) knowledge of process flow is rule-based knowledge with process routes or working operations. it comprises knowledge of feature process/product process/typical process. feature is the definition of components basic geometry entity for manufacturing. popular form features include cylinder, hole, plane, etc, which have typical recommended process scheme for each of them. knowledge of product process refers to process route information of product family or similar products, which may change according to the input manufacturing data. knowledge of typical process is the mature process route information which has been validated by practice and is used frequently (2) knowledge of resource refers to static manufacturing resource information, which includes all kinds of process resource, such as machine tools, fixtures, cutters, machining data, materials, etc. (3) knowledge of calculation refers to information obtained through calculation. in process planning, the selection of working hour and material quota is a regular activity. for these three types of knowledge, process flow knowledge is most difficult to capture and represent, because they are rule-based, dynamic, determined by many factors. most of them come from experience of experts or skilled employees. in the proposed pkms, a visual expressing method based on parameter flow charts (pfc) is used to represent them. this method combines the parameter technology, flow chart technology, and visual technology. for resource knowledge, an approach of classification tree is used to represent them. a formula management tool is used to represent calculation knowledge. the focus of this paper is placed on the approach of pfc. 3.2 representation approach based on pfc 3.2.1 definition of pfc definition 1: parameter (p) p is an element that has effects on the description or reasoning of the knowledge, e.g., when machining a hole in a part, the diameter, machining precision, and surface roughness will influence the machining route, therefore the elements can be extracted as: p = f(pg, pa, pr, ccons) f()is the definition function. is the geometric information of p. is the attributes of p, such as name, id, data type. is the action range of p, such as global parameter, local parameter. defines the constraints. p has two types, one is 352 w. l. chenli: research and development of process knowledge management system ... fig.3 process knowledge composition and classification model individual, and the other is dependent. the value of the individual p is directly from designers input or is derived from the superior process; the value of the dependent p is from the calculation result among other individual p according to the constraints. definition 2: parameter table (pt) pt is a set of information that arrays the parameters of the knowledge using a table form. it can efficiently support parameter storage, classification, discovery, and comparison. pt = f(pta, p, ui,m) ptais the attributes of pt, p is the parameter set in pt; ui is the user interface set of the pt; m is the manipulation set of pt, e.g, add, delete and edit. for process planning, the evaluations of the parameters often have mixed number and string operations, including the illegibility operation. definition 3: chart unit (cu) cu is the basic node that makes up of the pfc, which presents an action of the knowledge expression simulating the human beings. in different application domains, the actions are different, and the tasks and the aims of the actions are different too, so cus are different. according to the requirements of ke in different domains, the types and the attributes of cus can be extended. generally, advances in systems science and applications (2014) vol.14 no.4 353 a cu consists of seven types. cu = f(bu,eu, v u,ru,ou, iu, spfc) bu refers to the beginning of the pfc(see fig.4,bu1 ). refers to the end of pfc (see fig.4, bu1 ). sets value for parameters (see fig.4,vu1,vu2 ). ru depicts the reasoning logics according to the parameters. ru has single branch (see fig.4, ru1 ) or multiple branches (see fig.4, ru2 ).the ru with single branch denotes the judgment of yes or no logic. the with multiple branches defines several decisions according with several conditions. oudefines the output (see fig.4,ou1, ou2, ou3 ). is the attention graphic unit (see fig.4,iu1 ). iu1 assists in judging and validating for ke, which can help the designers find the incorrectness and remind designers to revise it. spfc represents a sub-process and nesting is available in the (see fig.4,spfc1 ). by establishing the database of and calling the spfc in the main process, the whole knowledge is available. fig.4 is a simple example of representation form for process flow knowledge based on pfc. it’s a visual and direction knowledge chart, starts from chart unit bu1 and ends at chart unit eu1. vu1 and vu2 are used to set values for parameters. ru1 and ru2 are responsible for judging the input data and determining the next flow step. spfc1 is a sub-pfc that is defined in another place ou1,ou2,ou3 , and are chart units of output, such as work operations. through drawing such flow chart with predefined chart units, a rule-based process knowledge can then be represented. fig.4 an example of representation form based on pfc 354 w. l. chenli: research and development of process knowledge management system ... definition 4: parameter flow chart (pfc) pfc is a direction chart for knowledge description and reasoning that consists of parameter tables, chart units and the logical routes among the chart units (see fig.4). pfc presents the contents and the structure of the knowledge. pfc =< a,pt,cu,cr > a, pt, cu and cr refers to the set of attributes, parameter tables, chart units and the relations among the chart units of pfc respectively. cr is classified as joined cr, unjoined cr, paratactic cr and branchless cr. for representing complicated knowledge, it can be represented by several sub-pfc. 3.2.2 model of the software system based on pfc fig.5 model of the pfc software system based on the analysis of the process flow knowledge and the definitions of pfc above, the model of the software system based on pfc is shown in fig.5. the parameters are defined via interactive interfaces, including extraction and semantic description of the parameters. the process flow chart is defined via microsoft visio, a popular flow design software system based on the component object. visio is customized and integrated with the proposed software system. hence, the capabilities of extendibility, visualization and readability of system are achieved. when explaining and reasoning using the pfc, the system will firstly read the related information from the pfc file (vsd file), then call the main module to load corresponding file via automation interface, finally explain and reason the flow chart. the system provides two operations, reasoning in step and reasoning advances in systems science and applications (2014) vol.14 no.4 355 continuously, to meet the requirements of debugging and output of process planning knowledge. during the process of explaining and reasoning, if errors occur, the system will notify the user and ask for solutions. after the process flow chart has been explained and reasoned successfully, the results will be formalized with a table form via neutral file. 3.3 other representation approach resource knowledge comprises machine tools, fixtures, machining data, etc. they are static and have diverse forms. classification tree is used to represent them. the node of the classification tree indicates the type of resource knowledge, and its self-defined attributes show the description and information of resource knowledge. such hierarchical structure of tree is very similar to practical process resource management. the structure of node is defined as following: class cproceeresnodetree { public: int nnodekey; //node id int nparentnodekey; //parent node id int nsibling; //sequence number in sibling int nchildren; //whether has child node cstring snodename; //node name int nnodetablekey; //database tale id of node int nnodegraphicskey; //document id of node cstring snodeeditaccess; //permission of edit cstring snodeviewaccess; //permission of browse } for calculation knowledge, the key issues include creating user-defined formula, definition of variable value source, and definition of suitable condition for the formula. the data structure of calculation knowledge is abstracted as following: class cformulaknowledge { public: int nformulakey; // formula id cstring sformulatype; // type of formula cstring sformula; // formula expression cstring sdescription; // formula description int nconditiontablekey; // table id of condition int nvariabletablekey; // table id of variable value source } to represent the calculation knowledge, a formula management tool based on the above data structure is developed and is shown in following case study. 356 w. l. chenli: research and development of process knowledge management system ... 4 case study fig.6 manage machine tools in pkms fig.7 manage cutting tools in pkms a prototype system of pkms has been developed to validate and demonstrate the feasibility and the compatibility of the proposed pkms. microsoft visual c++ is employed to develop the framework and functional modules. microsoft sql server 2005 is utilized to construct the systems repositories. microsoft visio advances in systems science and applications (2014) vol.14 no.4 357 2003 is customized and used to draw process flow charts. a variety of standard industrial process knowledge has been established using functional modules of the prototype system. fig.6 and fig.7 are screenshots of process resource management module using classification tree method. the left side of fig. 6 shows the resource classification tree that comprises of process resource nodes. the data structure of node is defined as mentioned above, including its name, parent node, database table and related document, etc. the nodes of tree can be created, deleted, searched according companys situation given authorization. the right side of fig. 6 shows the information of current selected resource node, a list of horizontal lathes with their technical parameters, which may guide to select appropriate lathe when carry on process planning. fig.7 is an interface of managing standard cutting tools. the drawings of current selected end mill are also captured and displayed. fig.8 manage calculation knowledge in pkms fig.8 is the screenshots of formula management module, which takes charge of managing calculation knowledge. two types of formula used often in process planning are built in case study, one is for material calculation, and the other is for working hour calculation. a whole formula may have its expression, variables, applied conditions, and categorization, etc. its data structure is defined as mentioned above. fig.9 is the screenshots of representation approach based on pfc. the example 358 w. l. chenli: research and development of process knowledge management system ... fig.9 manage process flow knowledge in pkms flow chart is drawn in microsoft visio 2003 using predefined chart units. mentioned above. there are two, four and one in fig.9, which are pointed by arrows, is used to judge the next information flow direction according to the input conditions. stands for working operation of process. is a sub-pfc defined elsewhere. the directed flow chart represents knowledge of a parts process route. 5 conclusion in this paper, the characteristics of process knowledge in okp companies have been analyzed; a definition and framework of pkms in okp companies were introduced as well. to overcome the difficulty of representing diverse type of process knowledge, a hybrid approach aiming at three main types of process knowledge was proposed. a prototype system has been developed to validate the proposed ideas. from this research work, the following conclusions are drawn: (1) the proposed pkms is open and extensible, the process knowledge database can be established and updated through developed interfaces. (2) the introduced hybrid approach can represent all kinds of process knowledge effectively; especially approach based on pfc can represent dynamic and rulebased process flow knowledge successfully, which has been validated in our case study. advances in systems science and applications (2014) vol.14 no.4 359 (3) the proposed approach is people oriented. microsoft visio, which is the most popular diagramming tool and easy to learn and use, is utilized to define process flow knowledge. skilled employee or experts can build process knowledge themselves without having to have much knowledge on information technology and without software programming. future research work is still required to improve the proposed system. this includes the development of new tools for process knowledge discovery, demonstration of the prototype system in different okp companies, and integration of other new approaches of knowledge representation, etc. references [1] oecd.(1996), the knowledge-based economy: france: head of publications service, pp.1-21. [2] sapuan, s.m. and h.s. abdalla.(1998), “a prototype knowledge-based system for the material selection of polymeric-based composites for automotive components”, composites part a: applied science and manufacturing, vol.29, no.7, pp.731-742. [3] aziz, e.s. and c. chassapis.(2002), “a knowledge-based approach to spur gear fabrication in precision forging process”, in proceedings of the asme design engineering technical conference. [4] huin, s.f., l.h.s. luong, and k. abhary.(2003), “knowledge-based tool for planning of enterprise resources in asean smes”, robotics and computerintegrated manufacturing, vol.19, no.5, pp.409-414. [5] mendikoa, i., m. sorli, j.i. barbero, and a. carrillo.(2005), “knowledge based distributed product design and manufacturing”, in proceedings of the 9th international conference on computer supported cooperative work in design. [6] cheung, c.f., y.l. chan, s.k. kwok, w.b. lee, and w.m. wang.(2006), “a knowledge-based service automation system for service logistics”, journal of manufacturing technology management, vol.17, no.6, pp.750-771. [7] halevi, g. and k. wang.(2007), “knowledge based manufacturing system (kbms)”, journal of intelligent manufacturing, vol.18, no.4, pp.467-474. [8] huang, s.h., x. hao, and m. benjamin.(2001), “automated knowledge acquisition for design and manufacturing: the case of micromachined atomizer”, journal of intelligent manufacturing, vol.12, no.4, pp.377-391. 360 w. l. chenli: research and development of process knowledge management system ... [9] paiva, e.l., a.v. roth, and j.e. fensterseifer.(2002), “focusing information in manufacturing: a knowledge management perspective”, industrial management and data systems, vol.102, no.7, pp.381-389. [10] shehab, e. and h. abdalla.(2002), “an intelligent knowledge-based system for product cost modelling”, international journal of advanced manufacturing technology, vol.19, no.1, pp.49-65. [11] aziz, e.s. and c. chassapis.(2004), “development of process optimization for an intelligent knowledge-based system for spur gear precision forging die design”, in proceedings of the asme design engineering technical conference. [12] pan, x., j. wang, and l. liu.(2004), “research on framework of knowledge management integration and key technologies in digital manufacturing enterprise”, jisuanji jicheng zhizao xitong/computer integrated manufacturing systems, cims, vol.10(suppl.), pp.90-95. [13] mahl, a. and r. krikler.(2007), “approach for a rule based system for capturing and usage of knowledge in the manufacturing industry”, journal of intelligent manufacturing, vol.18, no.4, pp.519-526. [14] xie, s.q. and y.l. tu.(2006), ‘rapid one-of-a-kind product development”, international journal of advanced manufacturing technology. vol.27, no.56, pp.421-430. corresponding author author can be contacted at: s.xie@auckland.ac.nz advances in systems science and application(2015) vol.15 no.2 119-133 obsolescence of organizations: a modeling approach in system dynamics didier cumenal1,2 1 dynamic systems, group afscet the french association of systems sciences 2 system dynamics society abstract economic crisis as well as human crisis is well alive. bankruptcy filings, the collapse of financial/capital markets, strategy failures, the death of some organizations accuse the turbulent economic environment. but what about their potential activities, their internal capacities that change over time. it is unfortunate that some literature areas dealing with the decline of organizations focus on their disease at a specific time by giving importance to factors related to age, size and performance. this is wrong to underestimate the evolutionary aspect of tangible and above all intangible resources. the purpose of this communication is to highlight that before the decline and death of an organization, there is a preliminary stage, often misperceived: the obsolescence.obsolescence defines a condition that prevents the organization to perform correctly its vital functions and that without her being fully conscious. our model of system dynamics has the primarily objective of defining the state of obsolescence and secondly to highlight the importance of two strongly linked subsystems meaning: organizational capacity and managerial ability. these two dimensions, coupled together, are often misidentified or misunderstood by the top management, whereas they may be able to characterize and anticipate an emergency situation. keywords managerial capacity; organizational capacity; dynamics; obsolescence; model; system 1 why, what field and what subject for this research? 1.1 why this study our past experience has led us to lead a group of “sme”, to straighten firms in distress as a senior consultant and, many years, to conceptualize these achievements with a thesis phd at the sorbonne in the field of organizations dynamics (evolution and change of state). it was during this career and after in becoming a full faculty member and researcher that we often asked ourselves some questions about the future of organizations. the sole ambition of this study is to identify and teach the concept of obsolescence as a general phenomenon often unknowed for executives and managers of the company. yet it is a step often announcing the decline or disappearance of the organization. our research on the phenomenon of obsolescence is trying to reach an understanding which is not limited to specific cases, but that is applicable 120 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics across time and space and to different types of organization. 1.2 the research question is there a persistent step which would herald the decline and death of organizations? what would this step? is it possible to identify the precursor weak signals of the organization collapse? can we model this prior step announcing disaster? while there are many researches on organizations mortality, yet few qualitative studies have been conducted about symptoms often considered marginal, but harbingers of economic and social tragedy. because the model is based on concepts such as the organizational capacity or the managerial capacity (cognitive factors) it belongs to qualitative areas. we then discard ourselves from quantitative parameters such as the age of the organization, of the staff, the size, the strength of the company that abound in the studies. 1.3 the problem the above questions can be synthetized by the finding and the question that follows, giving rise to thought and discussion: some organizations becoming mature lock themselves increasingly into their adaptations and are pushed inexorably such ”lava”, unable to change course, heralding a likely senescence index. therefore, how to spot early enough the gradual slowing of vital functions of an organization? what are the warning signs? to address this problem, we will define semantically the concepts of complex organization, organizational capacity and obsolescence.then we will depict the literature on our research field. finally we will present our model, the methodology underlying it and results of simulations will be debated during the ”discussion” which will close this presentation. 2 definitions and foundations of the model on the obsolescence 2.1 organization this is an intangible concept. for the organization is a myth[1], “it exists only events linked by causal loops” that creates meaning (’sense-making’). it is the interactions between the components of the organization that create a coherent and stable set. it is also a system of collective problem solving. in this spirit, the organization is the product of the choices of leaders and of their social representation. this collaborative system attempts to combine, to arrange the components of the work to be done to achieve the best organizational capacity, concept we will define below. organizations move, change over time1. the uncontrolled change creates asym1this topic has previously been the subject of dynamic simulations. cumenal d[3,4]. advances in systems science and application(2015) vol.15 no.2 121 metry, disrupts the harmony and curbs the progression of the organization. this one changes shape both structurally and in terms of the functional organization due to its history, its environment, internal resources and its management. thus, we enter in a field named “dynamics of organizations”. we think we can characterize this one by different stages: a conservative state where the weight of history, customs of the past pushes the organization to inertia despite active economic context; an adaptive state due to the pressure of the environment causing the competitive intelligence, to imitation; a cooperative state characterized by the maintenance of cohesion, the desire to avoid conflicts, the development of inter-organizational networks; finally, a state called progressing state which is defined by the desire to innovate and operate the organizational inimitable assets. of course, these states may be overlapped or juxtaposed. to move from one state to another, there is always energy dissipation, a cost of change. 2.2 the organizational capacity the end of 20th century saw the development of a school of thought based on organizational resources and on organizational skills named resource based view. organizational capacity is defined as a collective ability. we explain that the performance of the company benefits from a human capital (knowledge, skills) but also from a good organizational architecture (optimized process, formal and informal network of relations and communications, etc.) and from technologies. thus, the results can be explained by the presence of intangible resources shared by the players in the organization, but also difficult to imitate by competitors. according to laurent and gilles renard e st-amant “the concept of organizational capacity defines the ability or the skill of the organization to deliver activities with efficiency and effectively in deploying, combining and coordinating its resources and its expertise across various processes creating value, depending on the objectives defined previously [2].” 2.3 the definition of obsolescence from “obsolere”, “to fall into disuse”, in latin, the obsolescence reflects a behavior, an processing that is no more adapted to needs, to current and desired methods and to changing desired by the management . it turns out, over time, by slowing the functions and vital activities of the company (the unsuitable routines and processes business and support, etc.). the behaviors, the actions of an organization are no longer adapted to environmental change because the processes, routines, professional reflexes that have developed during the learning have become rigid or unable to agree to new situations. symptoms of obsolescence are often veiled, because they grow moderately over 122 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics time. for example, a product loses slowly market share; a business process is becoming more expensive and less efficient than a new process developed by competitors; staff engagement in the business of the company gradually becomes weak; the indifference of the staff for customers is becoming more marked; the leaders do not move in order to avoid achanges; information flows less and less. in general the reactions are increasingly inadequate and unsuitable for external events and circumstances. 3 literature review existing approaches for understanding the decline and death of organizations are fertile, although empirical studies on the spot are leaner. however, works on the obsolescence are not common and do not define properly the concept. seventy five percent of the literature on the decline of the firm were written during the last decade of the 20th century. we can mention books and papers by david a. whetten and kim s. cameron, authoritative on the subject of decay and mortality of the organizations[5][6]. most recently, yitzhak samuel describes in his book, models for predicting bankruptcy[7]. it also addresses the age, size and niche factors, performance and even corruption. stewart thornhill and raphael amit in an article named ”bankruptcy, firm age, and the resource-based view conducted a literature review on organizational mortality[8]. through the number of publications, it is clear for many researchers, that the age of the organization is a critical factor in mortality. however, stewart thornhill and amit raphael, learning , establish that the firms failure is also caused by a lack of managerial knowledge and in general by the lack of tangible and intangible resources (such as skills and organizational capacity)[9]. these firms are often paralyzed by a change in their environment. barry m. staw, lance e. sandelands, jane e. dutton show how organizational rigidity evolves faced with a challenge[10]. faced with a threatening environment, the organization tends to respond based on past experience deemed relevant, but unfortunately inappropriate given the new circumstances. moreover, in an urgent and tense situation, policymakers can hide certain information by discarding solutions they consider abnormal with regard to the opinions and beliefs of the group they belong to. somehow, filtered information must be in resonance with the unique thought of the leaders! erik larsen and alessandro lami operate a systems dynamics model by demonstrating how the causal relationships and feedbacks between variables, simulate over time, evolution of organizations inertia[11]. these authors show the opposition between resistance to change and the ability of organizations to develop new capacities to generate performance. the model, rather schematic is, according to advances in systems science and application(2015) vol.15 no.2 123 these authors, based on ecological theories and evolution of organizations. inertia is mentioned and briefly explained by the following factors: size, age, cumulative experience. figure 1 shows their model. it will be seen that size is correlated with age, which is an exogenous (external) variable to the system. fig.1 the organnization evolution model 4 model 4.1 methodology the model we have developed is not intended to generate assumptions to be tested. it seeks to show the stages of evolution of an organization to explain the causes of obsolescence thereof. it is therefore not limited to analyze each step separately considered. our model generates situations, rather than describing them. it strives to provide an pedagogical insight into the obsolescence onset. the model aims to study the evolution of organizational behavior and to identify early signs of wear and the weakening of main functions and activities of the organization. our approach is keyed on systems dynamics that studies how things, objects evolve and change state over time. john d sterman , in his book“business dynamics ”, sets principles2 of this approach[12] the approach is structured around the concept of causal network that establishes a cognitive diagram highlighting structural feedback loops. these are the relationships between the variables that 2for a criticism of the system dynamics, we carefully read the book by david berlinski[12]. 124 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics explain the system behavior over time, and not the system parameters considered separately. system dynamics highlights the counter-intuitive effects in the medium and long term, related to decisions taken in the short term, taking into account the effects of random events that can occur within, or in the environment of, the system. the model breaks down the system in state variables (reservoir) and flow variables (transit rules from one state to another). delayed effects and adjustment time are integrated into the model structure. the model uses numerical algorithms to produce simulated results. we can present the results as time series plots and measure the sensitivity of a variable to evolution of other one. system dynamics can be modeled using a set of equations with that standard form: x(t+∇t) = x(t) + x(t) • ∇t (1) in which is the model integration time step. however, the choice of value greatly influences the results of the numerical simulation as shown with the three graphs below (figure 2 : the population development; time on x-axis and quantities on y-axis). clearly, there are analytical solutions, but as soon as the model becomes complex, it is difficult to solve the equations in a conventional way. fig.2 the population development we will simulate our organization over several years to observe the evolution of the properties that characterize obsolescence. but for that, and before its use, we must expose the assumptions that characterize the structure of the model. advances in systems science and application(2015) vol.15 no.2 125 4.2 the causal structure of the model our model is structured into two subsystems or two causal diagrams: 1) the organizational capacity (c.0), 2) the cognitive ability of management (ccm). the following properties are not included in our model, voluntarily: age of the organization, age and size of staff. these attributes are often considered by the authors as exogenous variables. we believe that there is not necessarily a correlation between size and age of the firm, for example. likewise, these variables influence very little the inertia of the organization as we think sometimes. rather, we focus on capabilities of the organization, i.e. the potential3 and cognitive factors of the “top management” that ensure organizational transformation. it is the hypothesis that we assume. 4.3 the causal diagram of organizational capacity we defined above the organizational capacity. we modeled it (see figure 3) using the following properties: fig.3 organizational capacity causal diagram (o.c.) -organizational learning4 i.e. collective acquisition of skills by recurrent actions, -the skills produced by both organizational learning and training croutines5 built and distributed by internal collaboration within the orga3tangible and intangible resources provide the firm a competitive benefit. 4for a study on organizational learning, reference may be made to argyris, c[14]. 5routines represent individual and collective, repetitive work habits which may be driven by operational capitalized modes. 126 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics nization the flow of information and communication that contribute to this synergy between actors. this causal diagram shows many feedback loops that explain the results that we will comment during the presentation of results and discussion. for example, the more the widely disseminated (and unfiltered) information in the organization, the more it contributes through exchanges, to enhance coordination between employees or vice versa (direct relationship). on the other hand, the more the turnover is increasing, the more organizational skills6 are diminishing, due to departures and overall impact the organizational capacity. 4.4 the causal diagram of the managerial capacity collective rationalization is a cognitive process of sorting and selection of information, but this process also products resonance, answers already learned, as consistent with prevailing views in the firm. this mental state generates a collective logic that justify the decisions. necessarily wanting to be in agreement with the “others”, and especially with the “top management”, can be a powerful engine of collective errors. this rationalization taken to the extreme, leads to a denial, an occultation of the environment, filtering of information, but also a refusal to take a look on the past and on current state of the organization. the strategic blindness with unrealistic goals, can lead the entire organization to a thought, a single depiction. we have, unfortunately, observed this fact many times when a department persists in its judgment error. therefore excess of hierarchical pressure (the authority) was included as a variable in our model. beyond a pressure threshold, it acts to discipline, to structure responses of the organization. there is a “mahatma” effect by which governance behaves as a spiritual leader! its possible that problem could be identifyed, but the illusion that what we’ve always done is the best thing, trumps reality, generating actions that may be the wrong answers. our model in figure 4, below, reflects this divergence of representation with a gap that symbolizes a “cognitive dissonance”7 . with the latter, the information, leading indicators of entrepreneurial tragedy are ignored because they are too discordant with the successes until recently. the intensity of this difference is as important to the overall performance8 of the organization. moreover, poor re6organizational skills are generally depicted as coordination of resources and capacity[15]. they are linked to an accumulated experience (organizational learning). 7pioneered the concept of cognitive dissonance is leon festinger.according to this theory, when events lead a person or organization to act at odds with his beliefs, these acts produce a state of discomfort andtensioncalleddissonance.to reduce this condition, the person or organization is allowed to exculpate himself by constructing an argument near a deception! and management can persist in his error justifying its decisions and actions despite the good sense of “field”[16]. 8we define the concept of performance as a business development index value created by taking into account the quality of the responses of the“top management” facing a problem. advances in systems science and application(2015) vol.15 no.2 127 sults reinforce hierarchical control by systematically operating (routines) modes, but also impacting organizational agents behavior of the organization, themselves sensitive to performance. fig.4 managerial capacity causal diagram (m.c.) 4.5 results and behavior of the model it is interesting to study properties of all variables only on a qualitatively basis. we focused on the behavior, i.e. on shape of the curves rather than looking numerical values hard to estimate when the concepts of organizational and managerial capacity is mentioned. the range of variation of each variable extends for most of them from 0 (minimum) to 1 (maximum). it corresponds to the assumptions, and common sense “on the ground”. for instance, increased internal information filtering reduces synergy, knowledge sharing or informal9 kills formal knowledge and consequently performance. the simulation performed using a system dynamics software (stella, from isee systems, v.10.0.6) produces a typical organization behavior over a period of 10 years. you can of course change the duration of the simulation. watching graphical patterns on charts : 5, 6 and 7, we note a performance in sharp decrease. the diagnosis of the curves shows a slowdown in performance from grade 8 (curve 4 in chart 1 below) and then decreased. it can be seen also from the sixth year a slowdown in the level of organizational skills (curve 3), that of organizational routines10 (curve 5) and finally a breathlessness of organizational learning (curve 1). these are state variables which gradually accumulate flows acting as drivers 9it is, for example, short circuits in the official channels of communication in the organization. 10as a reminder, the routine reflects the regularity of collective behavior induced by procedures. 128 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics of change over time.chart 2 below shows a bad fit between dominant and learned responses that are an expression of “top management” thought and timely responses (curve 2) produced by the community in light of the problems identified. the gap persists for many several years, peaking in the third year of the firm. then, the two curves are similar from grade 611 to deviate from each other the 7th grade. we believe this difference is a revealing leading indicator of future performance degradation. finally, it takes place in companies when some people say that top management persists in his error by maintaining an axis of inappropriate development. chart1 however, the entire management team adopts this guided behavior and looking for one or several arguments to justify their position according to the cognitive dissonance theory, which we mentioned earlier. the chart 3 bellow, analyzes several behaviors:the impact of information filtering (curve 1), the degree of formalization and rationalization of the weighted procedures organization (curve 2) and intra-organizational synergy (curve 3). the latter also depends on the information that flows and the level of technical and managerial skills move through the organization. these are also indicators revealing symptoms which anticipate possible resources shortage and organization’s capabilities. actually, we hear, for example, that information does not flow, or the procedures are too heavy, reducing more and more autonomy in the work of each. 11is this rapprochement leads to the reason? advances in systems science and application(2015) vol.15 no.2 129 chart 2 chart 3 5 discussion and conclusion our model consists of five stock variables12(levels) corresponding to five integrations or state equations described as follows: e(tk) = e(tj) + ∫ tk tj [input(t)− output(t)]dt (2) e(tk)= thestateatktime. e (tj) = thestateatpreviousjtime. (j = k − 1)∑ tj tk = algebraic sum of inputs and outputs between t time and t time + increment time step dt. inputs and outputs over time depends on “auxiliary 12skills-learning-behavior-routine sand performance variables that have been previously defined. 130 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics variables” that changes according to causality with other variables (stocks, flows or auxiliaries). auxiliary variables are expressed with mathematical functions often based on non-linear correlations. our systemic model is built on feedback processes with, often, a delay between cause and effect. our model can help us to understand the emergent properties of the obsolescence and to analyze the results of an a concept not easily accessible through statistical interpretations of reality. we support the following proposal: analysis of organizational capacity and of managerial ability, i.e. symbolic representations and the answers given by the leaders facing a problem, can anticipate a resources and vital functions slowdown of the organization. we did not cover theories that emphasize the organizations ecology (size, age, etc.) to help us focus on intangible properties. of these ,we may recall the important role of learning and organizational skills, filtering and retention of information, the coordination and collaboration process(organizational synergy), collective rationalization of responses from the leaders giving the illusion that what they do is the best thing (further reinforced by the weight routines or past habits). the simulations show that weakly filtered information combined with a normal hierarchical pressure (more or less) can achieve a good level of performance. in this case, the staff has the opportunity to provide possible answers to the context. obviously, these responses are determined by the skills level and by the received or perceived information quality. censorship or a bad collection of information weakens staff responses and strengthens, conversely, the traditional answers provided by the “top management”. these are related to the rationalization of produced responses and to the organizational routines that capitalize collective behaviors generated by the procedures and by the taken decisions. in the latter situation, the operational process integrates finally the dominant responses produced by the “top management”, in finding good reasons to accept (see the concept of “cognitive dissonance”). our simulation explain the importance of the gap between the collective responses learned by staff and dominant responses supplied by the “top management”. this gap is an important performance gear. 5.1 model limitations we may recall the important role of learning and organizational skills, filtering and retention of information, the coordination an collaboration process (organizational synergy), collective rationalization of responses from the leaders giving the illusion that what they do is the best thing (the rest reinforced by the routines weight or the habits of past). the dynamics systems’ mathematics is widely used in physics. however, social sciences resist using them. indeed, formalize mathematically organizational behavior is very optimistic! advances in systems science and application(2015) vol.15 no.2 131 the model quickly becomes a black box for an external observer who wants understanding the results produced (non-linearity is one of the reasons). another limit is the validation of the model. we don’t have neither historical nor quantitative data about concepts presented. as we have previously stated: our model is more generative than descriptive of existing situations. it suggests events and behaviors. it allows us to anticipate, and to have a better understanding. in conclusion, we have been always convinced that an organization gets its best performance when the representations of a problem, the supplied answers, aims of employees and those of top management are similar. fig.5 bayesian model outline when societies and civilizations are becoming complex and prosperous, they have difficulty adapting to their environment. their potential assets appear to be inadequate to move to their next evolution stage. 5.2 the potential model evolution we are currently developing a bayesian network based model (figure 5) to anticipate the phenomenon of obsolescence. we intend to test this tool on several professional situations. the result will be the topic of a future communication. references [1] weick, karl e. (1995), “making sense of the organization”. blackwell publishing, pp.179-224. 132 didier cumenal: obsolescenceof organizations:amodeling approachinsystem dynamics [2] st-amant gilles et renard laurent. (2003), “capacity, organizational capacity and dynamic capacity: a proposal for definitions”. research paper, uqam. [3] cuménal, d. (2005), “simulating organizational change: moving and shaking”. in: 23rd international conferenceof the system dynamics society, july, pp.17-21. [4] cuménal, d. (2010),“can we model and simulate the changes of state of the organization over time?” . management & sciences sociales,vol.5, pp.91122. [5] whetten david a. (1979), “the organizational life cycle, chapitre 10”. josseybass publishers, pp.343-373. [6] cameron kim s.,. kim myung u,. whetten david a.(1987), “organizational effects of decline and turbulence”. administrative science quarterly, vol.32, pp.222-240. [7] yitzhak samuel. (2010), organizational pathology,transactions publishers. [8] thornhill stewart, amit raphael. (2003), “bankruptcy, firm age, and the resource-based view”. organization science volume,vol.14. [9] thornhill stewart, amit raphael. (2003), “learning about failure: bankruptcy, firm age, and the resource-based view”. organization science,vol.14. [10] staw barry m., sandelands lance e., dutton jane e. (1981), “threat rigidity effects in organizational behavior: a multilevel analysis. ”. administrative science quarterly,vol.26, pp.501-524. [11] larsen erik., lami alessandro. (2002), “representing change : a system model of organizational inertia and capabilities as dynamic accumulation processes ”. elsevier simulation modelling pratice and theory,vol.10, pp.271-296. [12] sterman john d. (2000), business dynamic,mcgraw-hill. [13] berlinski david.(1978), “on system analysis: an essay concerning the limitations of some mathematical methods in the social”. political, and biological sciences, mit press. [14] argyris, c, and schön d. a. (1996), “apprentissage organisationnel”. théorie, méthode, pratique, de boeck université. advances in systems science and application(2015) vol.15 no.2 133 [15] grant r.m. (1991), “ the resource-based theory of competitive advantage : implications for strategy formulation ”. california management review,vol.33, pp.114-135. [16] festinger léon. (1957), a theory of cognitive dissonance,standford university press. [17] amit r., schoemaker p.j. (1993). “strategic assets and organizational rent”. strategic management journal, vol.14, pp.33-46. [18] bossuet jacques bénigne. (1722),organizational effects of decline and turbulence, hachette réédition. corresponding author dr. didier cumenal can be contacted at: cumenald@wanadoo.fr sv-lncs adv syst sci appl 2018; 02; 107-120 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/539 copyright ©2018 assa. adv. in systems science and appl. (2018) an approach for developing context-aware adaptive information systems mahmoud hussein1 1) faculty of computers and information, menoufia university, egypt e-mail: mahmoud.hussein@ci.menofia.edu.eg abstract. context-awareness and adaptability are highly desirable features for modern information systems that are operating in dynamic environments. unfortunately, such information systems are still difficult to build. issues like (a) lack of an approach to compose a system from a set of functional services and context providers and (b) lack of a mechanism to enable runtime adaptation of the system in response to changes in its operating context need to be solved for easy and effective development of such information systems. in this paper, we introduce a novel approach to developing context-aware adaptive information systems. in composing the system, our approach explicitly separates but relates the system model and the context model, so that they and their relationships and changes can be clearly captured and managed. we also support runtime changes to the system, the context model, and their adaptation logic in response to the context changes. we do so by maintaining runtime representations of their models, and then we use these representations to realize the required changes. we have developed a tool to enable modelling the system and generate its implementations from their models. to demonstrate the viability of our approach, we used it to develop a context-aware adaptive travel guide system. we also assessed the performance of our approach through measuring the overhead caused by performing the system adaptation at runtime. the results demonstrate that our approach is capable of performing runtime adaptation with a small overhead. keywords: information systems engineering, context-awareness, system adaptation, context and system modelling, model driven development. 1. introduction there is an increasing demand for information systems that are aware of their contexts and dynamically adapt themselves at runtime in response to changes in these contexts [19]. consider, for example, a context-aware travel guide system that a tourist can use to plan his trip to visit some attractions. this system needs to be composed of a set of functional services (e.g. route planner and attraction finder) that takes the context information (e.g. weather and traffic information) into account to operate effectively (e.g. give the tourist better suggestions for routes and attractions). existing approaches focus on composing the system’s functionality (e.g. [4, 6-8, 10, 14, 25-26, 29]), but they leave the task of capturing the context and relating it with the system to the system developers. as such, the system implementation stage become more complex and requires a developer with experience in building such type of systems. in practice this expertise is usually lacking, and it will be a tedious task for the developers. while the travel guide system is in operation, it needs to adapt itself to keep achieving the tourist needs. for example, when the tourist wants to include a route planning service that was not provided to him free initially, the system should adapt itself by adding such service (i.e. changing system’s functionality), the context providers needed by that service to operate effectively such as a traffic information provider (i.e. adapting the context model), and some adaptation logic to enable the switching between different route planning services based on 108 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) their availability and provided qualities (i.e. changing the adaptation logic itself). most of existing approaches supports system’s functionality changes. however, existing approaches intertwined the context with the system as discussed above. as such, the addition of a context provider and relating it to the system becomes a difficult and error prone task. in addition, they use goal-policy [14, 29], utility-policy [7-8, 26], or action-policy [4, 6, 10] to capture the system’s adaptation logic. most of these approaches are not able to change the adaptation logic at runtime except the approach proposed by andrade et al. [2] that has the ability to change the adaptation logic manually at runtime by the system developer. however, the communications between the system developer and the running system are usually infrequent and sometimes impossible [18]. as such, the process of changing the system adaptation logic needs to be done automatically without the developer involvement. to address the above drawbacks, in this paper, we propose a model-driven approach to developing information systems that are able to dynamically adapt themselves in response to changes in their contexts, which we call context-aware adaptive systems. our approach has the following novel features. firstly, in composing the system, we explicitly separate but relate the system model and the context model, so that their relationships, changes, and changes impact across the system and its contexts can be clearly captured and managed. secondly, our approach enables the runtime changes to the system, its contexts, and their adaptation logic. we do so by maintaining runtime representations of their models and having two sets of rules: adaptation rules and adaptation meta-rules and. in response to context changes, the adaptation rules are used to decide the required changes to system and its contexts while the adaptation meta-rules specify changes that need to be applied into the adaptation rules themselves. in addition, our runtime environment enables applying the required adaptation actions. finally, our tool can be used to model a context-aware adaptive system and generate its implementations from their models. the remainder of the paper is organized as follows. we start by introducing a motivating scenario in section two. our approach to modelling context-aware adaptive information systems is described in section three. in section four, we discuss our tool support for modelling and realizing context-aware adaptive systems. in this section, we also measured the runtime overhead of adding the context-awareness and adaptability features to an information system to assure our approach’s applicability. section five analyzes existing work with respect to our approach. finally, we conclude the paper in section six. 2. motivating scenario and requirements analysis the travel guide information system helps a tourist to find attractions, plan his trip by providing suitable routes, and locate a restaurant. below are a few scenarios the tourist experiences in using this application during his visit to melbourne one day. scene 1: in the morning, the tourist starts to plan his trip. based on his preferences (e.g. outdoor attractions) and the weather forecast for that day (e.g. sunny), the travel guide suggests to him a number of attractions such as melbourne aquarium, royal botanic gardens, melbourne zoo, etc. he selected some of these attractions to visit using his rented car, and then a set of routes are displayed to him. these routes are calculated based on his current location, his attractions list, his driving preferences (e.g. shortest route), and current traffic information (e.g. blocked and congested roads). he selected a suitable route and started to explore the attractions. scene 2: at lunch time, the application suggested to him a number of nearby restaurants that matches his food preferences while taking into account the locations of the remaining attractions he is planning to visit. when he selects a restaurant, the trip route is re-planned automatically to take into account the restaurant location. an approach for developing context-aware adaptive information systems 109 copyright ©2018 assa. adv. in systems science and appl. (2018) to develop the travel guide system that meets the tourist needs, a set of general requirements need to be considered: compose a context-aware system (req. 1): the travel guide application need to be consisted of a set of functional services (i.e. attractions finder, restaurant locator, and route planner) that interact with each other to meet the tourist needs, while considering some quality requirements (e.g. fast route planner). in addition, these services should take into account the context information with a certain quality (if required) to give the tourist better suggestions. for example, the route planner needs the traffic information updated to the last minute to provide accurate estimations for the possible routes travel times and to display the routes that are less congested first. as such, there is a need for an approach that can be used to compose a set of functional services and context providers to provide the required context-aware functionalities while considering their quality requirements. adapt the system while it is in operation (req. 2): the travel guide provider may want to provide the attractions finder service free, while the tourist should pay to use the other services. as such, while the travel guide system is in operation, the tourist may wants to include the route planner service that is not provided free to him initially. to include such service, several changes need to be applied into the running system. firstly, the system needs to be adapted by incorporating the route planner service (i.e. adding a functional service). secondly, to find a suitable route for the tourist, there is also a need to acquire the tourist driving preference and current traffic information and use them in calculating and suggesting the routes (i.e. changing the context model by including context providers and their relationships with the system functionality). finally, the traffic information may become unavailable for a period of time (due to communication failure with road side units, for example), and then the travel guide provider need to have two route planners. one of them considers the traffic information in calculating the routes while the other does not take it into account. to switch between these two route planners while the system is running, the system adaptation logic needs to be changed, so that a suitable route planner can be selected based on the traffic information availability. 3. the approach to support the development of context-aware adaptive systems, we propose a model-driven approach. model-driven development is the notion of constructing a model of the system that can be transformed to a real system [9]. in this section, we introduce an organisational approach and associated notation that can be used to model such systems and in the next section we discuss how to transform a system model created in this notation to a real system. 3.1. composing context-aware adaptive systems: an organizational approach to develop a context-aware adaptive information system such as the travel guide system, it need to be composed of a set of functional services and context providers that interact with (related to) each other to meet the user needs (req. 1). in addition, while the system is in operation it needs to adapt itself in response to context changes to preserve the achievement of the user goals (req. 2). a system as an organisation is “a set of dynamic relationships between its roles to maintain the system viability in a changing environment” [21]. the relationships can be seen as defining the system, where they are used to specify the system roles position descriptions. these descriptions specify what tasks the system roles should do, while there are players who actually perform the tasks by playing these roles. in addition, in response to environment changes, the system manager may change the system roles, their relationships, and their players’ bindings to maintain the system viability. for example, a business organization is a 110 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) collection of roles (e.g. public officer, secretary, etc.) that are related to each other through contracts (relationships) [12]. these contracts define the permissible interactions between the organization roles and their mutual obligations (i.e. what tasks a role can request from others). in addition, the organization roles are played by employees or outsourced to external organizations. furthermore, to maintain the business viability in response to business market changes, the business manager may add roles, hire employees, etc [24]. similarly, we can see the travel guide system as an organization. in this organization, there are a set of functional and contextual roles that interact with each other to provide the system required functionally while taking the context information into account. in addition, these functional and contextual roles are played by functional services and context providers respectively (i.e. req. 1). furthermore, there is an organizer role that is able to adapt this organization in response to context changes to keep achieving the user needs. for example, add the route planning service and the other elements related to it when the tourist needs such service (i.e. req 2). because of the above correspondences, we followed the organizational approach in modelling context-aware adaptive information systems. following the organizational approach, a meta-model for a context-aware adaptive system is shown in figure 1. the system composition is consisted of a set of roles that are related to each other through contracts and each role can be played by one or more players. we have three types of roles: functional, context, or organizer as shown in figure 1. the functional roles represent the systems functionality while context roles capture the context model. the organizer role bound to its player is used to manage the system by adapting it in response to context changes. in addition, to capture the system elements’ different relationships, we have two types of contracts: functional and contextual. the functional contract is formed between two functional roles (i.e. role a and role b) to capture their functional interactions and mutual obligations. on the other hand, the contextual contract is formed between a context source and a context consumer to capture the contextual requirements and their required quality. furthermore, for each type of role, we have a corresponding player (i.e. context provider, organizer player, and functional player). in the following, we describe these concepts in details with examples from the travel guide system. fig. 1. a meta-model for context-aware adaptive software systems 3.2 modelling system’s functionality and context model using the above meta-model concepts, we introduced a modelling language that can be used by the software engineer to model a context-aware adaptive system. the basic elements of this language are shown in figure 2 and listings 1 and 2. we introduced these specific graphical notations and textual representation to ease the system modelling task compared to using general purpose languages (e.g. uml) and to enable the system code generation from its model [17, 23]. a uml profile was an option for defining our domain-specific language [11], however the contracts in our model are more complex (see listings 1 and 2) and it cannot an approach for developing context-aware adaptive information systems 111 copyright ©2018 assa. adv. in systems science and appl. (2018) be captured easily using the uml profile concepts. in the following, we describe our modelling language elements and use them to model the travel guide system. the functional system model: the system’s functionality is modelled as a set of functional roles that interact with each other though functional contracts. in addition, each role can be played by one or more functional players. functional contracts: the functional contracts are used to capture the relationships between the system functional elements and they have the following items. first, each contract has an identifier and it is formed between two functional roles. for example, the contract “fc4” is formed between the user and route planner roles as shown in figure 2. second, it has a set of permissible interactions between the contracted roles. each interaction as shown in listing 1 has (a) an identifier (e.g. i2) and a name (e.g. request routes), (b) zero or more input parameters (e.g. destination and current location), (c) a direction to specify who is responsible for providing the operation included in that interaction (e.g. “atob” which mean the route planer role is responsible for providing route calculation operation), (d) zero or more context parameters (e.g. blocked and congested roads), and (e) a return type (e.g. routes). third, the contract defines a set of conversion clauses that specifies the acceptable sequences of interactions between the functional roles. we used interaction rule specification (irs) language to specify these temporal constraints as shown in listing 1 [15]. for example, “c1” specifies that “request routes” operation must be invoked before invoking “select a route” operation. forth, to define the system non-functional requirements (e.g. response time, availability, and reliability), we specify a set of obligations on the functional interactions and we followed the web service level agreement (wsla) language in defining these obligations [20]. for example, the obligation “o1” in listing 1 specifies that the request route operation “i1” should not take more than 5 seconds in calculating the routes. finally, the interactions between the functional roles can be permitted or blocked based on the context information. for example, the interaction “i2” cannot be performed when the traffic information is not available and then it needs to be blocked. to do so, each contract has a set of regulation rules to regulate its interactions. we used event-condition-action rules to decide permitting or blocking an interaction [22]. an example of such rules is the rule shown in listing 1. this rule is triggered when the interaction “i2” is requested (i.e. i2.activated). when the traffic information is not available, the interaction “i2” is blocked (r2) while it is permitted otherwise. fig. 2. the context-aware travel guide system model 112 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) interaction clauses: i1: {requestroutes (destination, currentlocation), atob, routes}; i2: {requestroutes1 (destination, currentlocation), contextparameter (blockedroads, congestedroads), atob, routes}; i3: {selectroutes (routeid), atob}; conversion clauses (temporal constraints): c1: {i1 precedes i3 globally} c2: {i2 precedes i3 globally} obligations: o1: {i1, responsetime, lessthan , 5 seconds} o2: {i2, responsetime, lessthan , 7 seconds} regulation rules: r1: {block the interaction i2: when i2.activated; if trafficinformationavailability == false; do i2.block( )}; functional roles and their players: the functional roles represent the system’s functionality where each role position description is derived from its relationships (contracts) with other functional roles (i.e. an abstract definition of the tasks that need to be performed by a player who will play that role). the functional role also has one or more functional players to provide its actual functionality at runtime. for example, there are two route planning algorithms that can play the route planner role as shown in figure 2. routeplanner1 considers the traffic information in calculating the routes while routeplanner2 do not take the traffic information into account. the context model: the context model is represented through a set of contextual contracts that are formed between context sources and context consumers to capture the system contextual requirements. in addition, there are a set of context providers to make the context information available to the system. contextual contracts: the contextual contract defines what context information required by a system element (i.e. the context consumer in figure 1 such as functional contract, organizer role, or functional role) and the quality of this required context (e.g. accuracy, freshness, etc.). for example, in listing 2, a contextual contract “cc6” is used to specify that the functional contract four needs to know (a) the congested roads with freshness up to the last minute and accuracy greater than 80%, and (b) the blocked roads with accuracy greater than 95% so that the route planner can calculate the routes effectively. in addition, it needs to know the traffic information availability to decide permitting or blocking the interaction “i2”. listing. 2. the contextual contract between the traffic information role and the contract fc4 contextual contract id cc6: trafficinformation_fc4 parties: context source:trafficinformation; context consumer: fc4; context attributes: a1: congestedroads; a2: blockedroads; a3: trafficinformationavailability; context attributes quality: q1: {a1, freshness, lessthan , 1 minute} q2: {a1, accuracy, greaterthan, 80%} q3: {a2, accuracy, greaterthan, 95%} context roles and context providers: the system contextual requirements are captured through a set of contextual contracts as discussed above. these contracts are then used to drive the context roles position description, where each context role bound with a context provider (each context role can have one or more context providers) is responsible for providing some context information. for example, a weather role bound to its provider (i.e. the weather service) listing. 1. the functional contract between the user and the route planner functional roles functional contract id fc4: user_routepalnner parties: role a: user; role b: routeplanner; an approach for developing context-aware adaptive information systems 113 copyright ©2018 assa. adv. in systems science and appl. (2018) is responsible for providing current temperature and rain level to the attraction finder service, so that a correct suggestion is given to the user based on current weather conditions. 3.3 engineering system’s adaptability through organizer role and its player to make the system able to adapt itself in response to context changes, there is a need for a mechanism to first decide when and what to adapt and then apply the decided adaptation actions. we do so through the system organizer role and its player. we modelled the organizer player as a set of event-condition-action rules that can be used to decide when and what to adapt [22]. the event and condition part of a rule specifies when to adapt, and the action part of the rule defines what to adapt. the events that activate the adaptation rules are usually context changes where the system needs to adapt itself in response to these changes. the rule condition is used to specify the context situation that needs a system reaction(s). the rule action is a set of adaptation actions to cope with the context changes. in general, the adaptation actions are to add, remove, or modify a system element. for example, to change the system functional and context roles, we have three adaptation actions: add role, remove role, and change role-player binding. in the same manner, we have actions to add, remove, and change a functional contract and a contextual contract. to apply the adaptation actions, the organizer role is engineered with a set of standard methods that are corresponds to the adaptation actions so that the organizer player can invoke these methods to adapt the system while it is in operation. an example of an adaptation rule is shown blow. this rule is activated when the tourist wants to include the route planning service in his application (i.e. event). in response to the changes in the tourist needs (i.e. he wants the route planning service, condition), the application adapts itself (i.e. actions) by adding a set of roles (e.g. route planner), adding a set of contracts (e.g. fc3 and fc4), and bind players with the added roles (e.g. bind route planner role with route planner one). rule “adaptationrule1”: { when valuechanges (routeplannerselected); if routeplannerselected == true; do addrole(“routeplanner”), addcontract (“fc3”), addcontract (“fc4”), addrole (“trafficinformation”), addcontract (“cc6”), addcontract (“cc8”), bind(“routeplanner”, “routeplanner1”), bind(“trafficinformation”, “roadsideunit”)}; when the running system is changed, its adaptation rules may need to be adapted too. for example, after performing the adaptation actions in the above rule, we can see that the route planner role has two players and the traffic information role has two context providers as shown in figure 2. as such, there is a need for some rules to decide the switching between the added roles-players. to do so, we specified another set of rules (i.e. adaptation meta-rules) that adapt the adaptation rules themselves. these rules have the same structure of the adaptation rules describe above, but they have different adaptation actions (i.e. add rule and remove rule). an example of such meta-rules is show blow. in this rule, four adaptation rules are added to enable a correct selection of a route planner algorithm and a traffic information provider when the route planning service is included into the running system. rule “adaptationmetarule1”: { when valuechanges (routeplannerselected); if routeplannerselected == true; do addrule(“selectrouteplanner1”), addrule(“selectrouteplanner2”)}; addrule(“selectroadsideunit”), addrule(“selecttrafficinfoprovider”)}; 4. implementation and approach evaluation in this section, we describe the tool that has been developed to support the modelling and realization of context-aware adaptive information systems, and how to use this tool to develop 114 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) the travel guide system described in section two. we also measure the performance overhead of adding the adaptability feature to a software system using our approach to assure its applicability. 4.1 tool support for modelling context-aware adaptive system we have developed a tool to enable the modelling of a context-aware adaptive system. it enables the software engineer to specify the system roles, players, contracts, and adaptation rules. screenshots from our tool during the system modelling are shown in figure 3. fig. 3. screenshots from our tool during the modelling of the travel guide application to simplify the process of specifying the adaptation rules, we provide a gui that helps the engineer in codifying these rules. this gui is used to specify the rule conditions and adaptation actions. the rule events are directly inferred from the rule conditions, where they usually are the changes in the context attributes that are used in the rule conditions. an example is shown in figure 4, where the engineer can specify the current situation i.e. the user wants to include the route planning service and a set of adaptation actions need to be performed in this situation such as add route planner role (figure 4-a), bind the route planner one player to route planner role (figure 4-b), and add functional contract “fc4” (figure 4-c). fig. 4. specifying the travel guide system adaptation rules using our tool 4.2 realizing context-aware adaptive systems to realize context-aware adaptive systems, we used road framework1 where it follows the organizational approach as our approach does. this framework is an extension to the apache axis22 to realize adaptive software systems [16]. to use this framework, we used our tool to transform the model described in the previous section to a model that is compatible with the road framework. in the following we describe the major transformations we did while the others are one-to-one mapping. 1 http://www.swinburne.edu.au/ict/research/cs3/road/ 2 http://axis.apache.org/axis2/java/core/ http://www.swinburne.edu.au/ict/research/cs3/road/ http://axis.apache.org/axis2/java/core/ an approach for developing context-aware adaptive information systems 115 copyright ©2018 assa. adv. in systems science and appl. (2018) first, the use of context information as extra parameter in a functional interaction is not considered in road model. but, during the system execution, the functional contracts are used to mediate the interactions between the system roles. when an interaction reaches a contract a set of rules are executed to decide permitting or blocking it (in road these rules are codified as drools rules3). as such, we added a rule to these rules that is activated when a contextualized interaction is received. this rule is responsible for updating the context information required by this interaction before sending it to the destination role. an example of such rule is shown blow. this rule updates the context information (i.e. congestedroads and blockedroads) of the request route interaction when it is received at the functional contract “fc4”. second, in road model, the context information is maintained as a set of facts. each fact contains one or more context attributes. these facts can be provided or consumed by the system roles or its functional contracts. in our model, the contextual contracts are used to drive context roles descriptions, and then each context role can be seen as a collection of context attributes. this makes a correspondence between a fact in road terms and context role derived from contextual contracts in our model. as such, we use the contextual contracts to drive context roles descriptions, and then we transform each role position description to a fact in road model. third, to enable the execution of the adaptation rules and meta-rules described above, we transform them to drools rules, so that the drools rule engine can be used for their execution to decide the required adaptation actions while the system is in operation. the result of transforming the “adaptatiorule1” described above is shown blow. in transforming the rules, we used the rule “when” part to specify the rule event (e.g. the user need of the route planning service is changed). in addition, the rule “then” part is used for capturing both the rule condition and action. to evaluate the rule condition part, we created a class called “conditionevaluator”. this class has a method called “evaluate” that takes a condition as an input, and then it replace the condition variables with current context values and returns true or false based on the condition evaluation. when the condition is evaluated to true (i.e. the user wants the route planning service for example), a set of adaptation actions are added to the adaptation script (e.g. actions.addaction(“addrole_routeplanner”). by the same manner, the adaptation meta-rule “adaptatiometarule1” can be generated. the adaptation rules are used to generate an adaptation script. this script is then used by the organizer player to adapt the running system by invoking the organizer role standard adaptation methods that are corresponding to the required actions. these standard methods require some details to execute. for example, to add a role there is a need for the role name, 3 http://www.jboss.org/drools http://www.jboss.org/drools 116 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) identifier and description. here comes the role the maintained runtime representation of the system models where they have these required details. as such, the organizer player parses these representations to get the required details and generate the executable actions. for example, adding the route planner role in the above script is transformed to organizer.addnewrole(“fr2”, “routeplanner”, “role represents route planning service”). the organizer variable is a reference to the running system organizer role. a similar mechanism is used for changing the adaptation rules, where we have a reference of the loaded adaptation rules (i.e. instance of knowledgebase 4 class) and the methods addknowledgepackages and removerule are used to add and remove adaptation rules respectively. the above transformation process is automated in our tool. when the software engineer completes the system modelling, he can press a button that generates the files required by the road framework to deploy an instance of the system. this instance contains the system roles, their contracts and the generated organizer player. to have a fully running system, we have developed a set of functional players and context providers. for example, we used google maps5 services to develop he route planners, attractions finder, and restaurant locator players. we also have developed a gui to enable the user interactions with the provided services as shown in figure 5. in figure 5-a, the application only includes the attraction finder service which is provided free initially. this service suggested to the tourist a set of attractions based on his preference, his current location, and the weather forecast. he can select some of them to be included in his attractions list. to plan a route to see these attractions, the tourist requests the route planning service to be included in his application. as such, the application is adapted to include such service by performing the adaptations described above. after the internal adaptations are performed, the application gui is changed also by including the route planning service. when this service becomes available, it acquires tourist location, his attractions list, and traffic information to suggest a suitable route (see figure 5-b). fig.5. the context-aware adaptive travel guide application in action 4.3 performance evaluation6 the overhead in our approach is the extra time needed to adapt the system while it is in operation. this can be calculated by the time required to monitor the context, decide the required adaptation actions, and act these actions. monitor the context. in our approach, there is a need to keep track of some context variables that cause system adaptation. when any of the variables that the adaptation rules are interest in changes, this change is notified to the organizer player to decide the needed 4 http://docs.jboss.org/jbpm/v5.1/javadocs/org/drools/knowledgebase.html 5 http://code.google.com/apis/maps/documentation/webservices/ 6 a pc with intel core 2 due 3 ghz cpu and 3 gb ram is used as the test-bed and drools-5.1is used as the rule engine. http://docs.jboss.org/jbpm/v5.1/javadocs/org/drools/knowledgebase.html http://code.google.com/apis/maps/documentation/webservices/ an approach for developing context-aware adaptive information systems 117 copyright ©2018 assa. adv. in systems science and appl. (2018) adaptations. the time required to notify the system organizer with a context variable change equals to 14.59 milliseconds in average. decide required adaptation actions. when the context is changed, the adaptation rules need to be evaluated to decide the required adaptation actions. the overhead in the decision making process is laid in rules loading time at the beginning and their execution time in response to context changes. to measure that overhead, we used sets of rules with sizes 10, 20, 30, 40, and 50. figure 6-a shows that the rules loading time varies from 1.81 to 2.05 seconds based on the adaptation rules size. this is not much overhead where it is usually performed once at the system start-up. in addition, the time to execute the rules is between 14.8 to 29.9 milliseconds (see figure 6-b) which cannot be considered as an overhead also. fig.6. the adaptation rules loading and execution times apply the adaptation actions. we have different adaptation actions that can be performed to adapt the system in response to context changes. table 1 summarises the average time needed to apply some of the required adaptation actions in milliseconds. in table 1, we only show the actions for adding some elements to the system, where they are of interest from the user point of view (i.e. they are added where the user wants them). the adaptation actions for removing parts of the system have a small overhead and are performed while it is running. as such, they do not affect the user interactions with the system. due to space limitation, we do not present them here. table 1. the time required to apply the required adaptation actions in milliseconds adaptation action required time adaptation action required time add functional/context role 198.953 change role-player binding 0.0348 add functional contract 0.204 add adaptation rule 0.0096 add contextual contract 1.565 a delay that the tourist can experience in using the travel guide system is happened when he wants to include the route planning service. to include this service, it takes around 422 milliseconds which cannot be considered as a delay to the tourist. 5. related work and discussions a number of approaches have been proposed to support the process of developing contextaware adaptive systems. some of them propose a way to model and realize such systems [4, 8, 14, 25-27, 29] while the others provide a framework or a middleware to help the software engineer in realizing the context acquisition and interpretation [5, 13, 28] or the system runtime management [1-3, 7, 10]. in this section, we analyse existing and our approaches in relation to the requirements we have identified in section two. compose a context-aware system: existing approaches to developing context-aware adaptive systems can be classified into two categories. on the one hand, a set of approaches consider the system functionality as a single service and the context information is used to 118 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) adapt this service operational parameters [27, 29]. on the other hand, a set of approaches has been proposed to compose a system from a set of components that are changeable at runtime to cope with context changes [4, 8, 14, 25-26]. however, they capture the relationships between the system and its context implicitly during the system development except few who only consider these relationships explicitly at design time [4, 27, 29]. as such, the system implementation complexity is increased and the system-context relationships changes become difficult and error prone. in addition, in composing system’s functionality, existing approaches connect the composed components directly (i.e. linking the components required and provided interfaces directly) [4, 8, 14, 25-26]. as such, in large scale systems, the composed elements complex interactions and their mutual obligations become difficult to capture and manage. in our approach, we explicitly represent the relationships between the system and its context and between the system functional elements themselves. as such, we can clearly compose and realize a context-aware adaptive information system. similar to our approach sheng et al. [27] propose approach who only consider these relationships explicitly only at design time and they also consider the system as a single service runtime adaptation of the context model: at runtime the context model may change by including a context attribute or remove an existing one as shown in our motivating scenario. however, most of existing approaches have only a design time context model and usually it disappears during the system implementation where it is intertwined with system’s functionality and/or management. few approaches keep the context model explicit at runtime [26, 28], so that they can switch between different context providers or include new providers while system is in operation. however, they do not have the ability to change the context model itself by adding, removing, or modify the system required context attributes. our approach has an adaptable runtime representation of the context model and its management enables its runtime changes. adaptation logic runtime changes: existing approaches use goal-policy [14, 29], utilitypolicy [8, 26], or actionpolicy [1-2, 13] to capture the system adaptation logic. most of these approaches do not support the runtime changes to the system adaptation logic except andrade et al. [2] who enable the system developer to change this logic manually. however, the adaptation logic of the travel guide system need to be changed automatically without the developer involvement as the system may include or exclude parts of its adaptation rules while it is running. to do so, our approach maintains a runtime representation of the adaptation rules and we have a set of adaptation meta-rules that decides the required changes to the adaptation rules in response to context changes. in this paper, we adopted the action-policy approach in capturing the adaptation logic because of its expressiveness and availability of tools that supports runtime changes of the rules. 6. conclusion in this paper, we have proposed a model-driven approach to developing context-aware adaptive information systems. we have considered the system model, the context model, and their relationships explicitly from modelling to realization and to runtime execution and adaptation. in addition, we have developed a prototype tool for modelling the system and generating its implementations from their models. furthermore, we have demonstrated our approach through the development of the context-aware adaptive travel guide system. we also measured the overhead of adding the runtime adaptability feature to a software system using our approach. compared to existing approaches, our approach has the following key contributions. firstly, we explicitly separate but relate the system model and the context model, so that their relationships and changes can be clearly captured and managed. secondly, the relationships an approach for developing context-aware adaptive information systems 119 copyright ©2018 assa. adv. in systems science and appl. (2018) between the system functional elements are represented explicitly, so that the functional elements interactions, mutual obligations, and changes can be clearly captured and managed. thirdly, our approach supports the runtime adaptation of the system, the context model, and their adaptation logic in response to context changes. finally, our tool supports the system modelling and the generation of its implementations from their models. as future work, our approach can be enhanced in several directions. first, software systems are usually deployed in environments which are not totally anticipated at the system design time. while we have a runtime representation of the system aspects (including rule-based management) to be able to cope with unanticipated changes, runtime system management strategies and decision-making techniques are required to fully realize this capability. second, we have applied our approach to the tourist travel guide case study, and the results were promising. we will perform more validations to assess the applicability and practicality of our approach. references [1] adamczyk, j., chojnacki, r., jarząb, m., & zieliński, k. (2008). rule engine based lightweight framework for adaptive and autonomic computing international conference on computational science (vol. 5101, pp. 355-364). [2] andrade, s. s., & de araujo macedo, r. j. (2009, 18-19 may 2009). a non-intrusive component-based approach for deploying unanticipated self-management behaviour. paper presented at the software engineering for adaptive and self-managing systems, 2009. seams '09. icse workshop on. [3] asadollahi, r., salehie, m., & tahvildari, l. (2009, 18-19 may 2009). starmx: a framework for developing self-managing java-based systems. paper presented at the software engineering for adaptive and self-managing systems, 2009. seams '09. icse workshop on. [4] ayed, d., delanote, d., & berbers, y. (2007). mdd approach for the development of context-aware applications. paper presented at the proceedings of the 6th international and interdisciplinary conference on modeling and using context, roskilde, denmark. [5] capra, l., emmerich, w., & mascolo, c. (2003). carisma: context-aware reflective middleware system for mobile applications. software engineering, ieee transactions on, 29(10), 929-945. [6] david, p.-c., & ledoux, t. (2006). an aspect-oriented approach for developing selfadaptive fractal components software composition (sc'06) (vol. lncs 4089, pp. 8297). [7] elkhodary, a., esfahani, n., & malek, s. (2010). fusion: a framework for engineering self-tuning self-adaptive software systems. paper presented at the proceedings of the eighteenth acm sigsoft international symposium on foundations of software engineering, santa fe, new mexico, usa. [8] floch, j., hallsteinsen, s., stav, e., eliassen, f., lund, k., & gjorven, e. (2006). using architecture models for runtime adaptability. ieee softw., 23(2), 62-70. doi: http://dx.doi.org/10.1109/ms.2006.61 [9] france, r., & rumpe, b. (2007). model-driven development of complex software: a research roadmap. paper presented at the 2007 future of software engineering. [10] garlan, d., cheng, s. w., huang, a. c., schmerl, b., & steenkiste, p. (2004). rainbow: architecture-based self-adaptation with reusable infrastructure. computer, 37(10), 4654. [11] giachetti, g., marín, b., & pastor, o. (2009). using uml as a domain-specific modeling language: a proposal for automatic generation of uml profiles advanced information systems engineering (vol. lncs 5565, pp. 110-124). http://dx.doi.org/10.1109/ms.2006.61 120 m. hussein copyright ©2018 assa. adv. in systems science and appl. (2018) [12] governatori, g. (2005). representing business contracts in ruleml. international journal of cooperative information systems, 14(2), 181-216. [13] gu, t., pung, h. k., & zhang, d. q. (2005). a service-oriented middleware for building context-aware services. j. netw. comput. appl., 28(1), 1-18. doi: http://dx.doi.org/10.1016/j.jnca.2004.06.002 [14] heaven, w., sykes, d., magee, j., & kramer, j. (2009). a case study in goal-driven architectural adaptation software engineering for self-adaptive systems (pp. 109-127): springer-verlag. [15] jin, y., & han, j. (2005, 15-17 dec. 2005). consistency and interoperability checking for component interaction rules. paper presented at the software engineering conference, 2005. apsec '05. 12th asia-pacific. [16] kapuruge, m., colman, a., & king, j. (2011, aug. 29 2011-sept. 2 2011). road4ws - extending apache axis2 for adaptive service compositions. paper presented at the enterprise distributed object computing conference (edoc), 2011 15th ieee international. [17] kelly, s., & tolvanen, j. p. (2008). domain-specific modeling: enabling full code generation: wiley-ieee computer society pr. [18] kramer, j., & magee, j. (2007). self-managed systems: an architectural challenge. future of software engineering, 2007. fose'07, 259-268. [19] liaskos, s., litoiu, m., jungblut, m., & mylopoulos, j. (2011). goal-based behavioral customization of information systems. in h. mouratidis & c. rolland (eds.), advanced information systems engineering (vol. 6741, pp. 77-92): springer berlin / heidelberg. [20] ludwig, h., keller, a., dan, a., king, r. p., & franck, r. (2003). web service level agreement (wsla) language specification. ibm corporation, 815-824. [21] maturana, h. r., & varela, f. j. (1987). the tree of knowledge the biological roots of human understanding. 1st ed edn. boston: new science library. distributed in the united state by random house. [22] mccarthy, d., & dayal, u. (1989). the architecture of an active database management system. acm sigmod record, 18(2), 215-224. [23] mernik, m., heering, j., & sloane, a. m. (2005). when and how to develop domainspecific languages. acm computing surveys (csur), 37(4), 316-344. [24] mintzberg, h. (1994). rounding out the manager's job. sloan management review, 36, 11-26. [25] morin, b., barais, o., nain, g., & jezequel, j.-m. (2009). taming dynamically adaptive systems using models and aspects. paper presented at the proceedings of the 31st international conference on software engineering. [26] rouvoy, r., barone, p., ding, y., eliassen, f., hallsteinsen, s., lorenzo, j., . . . scholz, u. (2009). music: middleware support for self-adaptation in ubiquitous and serviceoriented environments. in b. cheng, r. de lemos, h. giese, p. inverardi & j. magee (eds.), software engineering for self-adaptive systems (vol. 5525, pp. 164-182): springer berlin / heidelberg. [27] sheng, q. z., jian yu, segev, a., & liao, k. (2010). techniques on developing contextaware web services. international journal of web information systems, 6(3). [28] taconet, c., kazi-aoul, z., zaier, m., & conan, d. (2009). ca3m: a runtime model and a middleware for dynamic context management. paper presented at the proceedings of the confederated international conferences, coopis, doa, is, vilamoura, portugal. [29] zhang, j., & cheng, b. h. c. (2006). model-based development of dynamically adaptive software. paper presented at the proceedings of the 28th international conference on software engineering, shanghai, china. http://dx.doi.org/10.1016/j.jnca.2004.06.002 advances in systems science and application (2015) vol.15 no.4 351-365 high throughput area efficient architecture for light weight cryptography vanitha m and subha school of information and technology, vit university, vellore 632014, tamilnadu, india abstract hummingbird algorithm is one of the recently proposed light weight cryptographic algorithms targeted for resource constrained devices like rfid (radio frequency identification), smart cards and majority of wireless sensor nodes. the main advantage of this algorithm is that it provides adequate security with smaller block size. as per the previous works on this algorithm, area and performance are two main design tradeoff of this algorithm. performance is increased by loop unrolling and area is optimized by looping. so optimization in both area and performance is a big challenge. this work, proposes efficient hardware architecture for the hummingbird algorithm using partial loop unrolling which concerns both the area and performance of hardware implementation. the overall architecture was modeled using verilog hdl and synthesized using cadence rtl compiler with 45nm technology from tsmc. proposed design also implemented in low cost spartan 3 fpga board and the results are compared with the existing implementations. results show that there is an area reduction of around 6% and throughput almost get doubles. keywords hummingbird; lightweight cryptography; rfid. 1 introduction the importance of low cost devices like rfid (radio frequency identification), smart cards and various wireless sensor devices is increasing in our present day life. so the security of such devices is very important. since these devices are extremely resource constrained in terms of computing power, battery power supply and memory, the standardized cryptographic algorithms such as aes(advanced encryption standard) , des(data encryption standard), which are well focused on software implementation rather than hardware, cannot be used as it is in these devices. so a certain class of cryptographic algorithms known as light weight cryptography is evolved. the design metrics of light weight cryptography are security, cost, and performance. practically it is difficult to optimize all the three design goals. 1.1 related works many lightweight cryptographic algorithms are presented until now, a light weight algorithm named hight is proposed with 3048 gate equivalents (ge) which is much faster than aes[1]. scalable encryption algorithm (sea), with a block 352 vanitha m and subha s:high throughput area efficient architecture for light ... size and key size of 96 bits, and word size of 8 bits and 93 rounds operation can encrypt one data block within 1428 clock cycles and 3758 ge[2].slight modifications is done on classical block cipher lightweight des variant called desl (des lightweight)[3]. asic implementation of the same requires 1848 ge and it can encrypt one 64-bit data block in 144 clock cycles. the implementation of desl has approximately 20% smaller chip size than des. key whitening technique can be useful to improve the security of cipher in desxl algorithm. for encrypting the plaintext, it requires 2, 170 ges and 144 clock cycles. present algorithm is a lightweight substitution, permutation based block cipher, which operates on 32 rounds, key size of 80 or 128 bits and 64 bit block size. present serial version can be implemented with 1000ges[4]. the hummingbird algorithm is the one of the recently presented ultra-light weight cryptographic algorithm[5]. the size of the key and the internal state of hummingbird provides adequate security level for many embedded applications. the present researches are going on the development and different implementation of the same because of the simplicity in the architecture. until now the hummingbird has been implemented on different target platform, software as well as in hardware and shows good efficiency in both. xinxin fan implemented the algorithm on 16 bit as well as 4-bit microcontroller.[6] daniel engels implemented the architecture on 8-bit microcontroller.[7] xinxin fan implemented both area oriented and throughput oriented design of hummingbird on low cost fpga.[6,8] biao min proposed another efficient hardware implementation of the same fpga.[9] ismail san proposed yet another implementation on fpga using coprocessor approach.[10] all these implementation shows that this algorithm works well on different target platforms. as per the previous works, hummingbird can be implemented in different platforms but each of which is focused on either optimizing area or optimizing the speed. here we are presenting an efficient architecture for the hummingbird algorithm focusing to optimize the area as well as the speed. comparison is made with the existing implementations. the synthesized result shows the increase in throughput by 111% and reduction in area by 6.6%. the remaining portion of this paper is organized as follows. section 2 presents the standard hummingbird algorithm, with its initialization, encryption and decryption steps. next, section 3 discuss about the proposed architecture for increasing the throughput and reducing the area. section 4 presents the simulation results, synthesis outcome in fpga and asic platform and their comparison with the existing architecture section 5 concludes this work. advances in systems science and application (2015) vol.15 no.4 353 2 standard hummingbird algorithm hummingbird is the recently proposed ultra-light weight cryptographic algorithm, which is the combination of block cipher and a stream cipher. the design of hummingbird consists of 16-bit block size, 256-bit key size, and 80-bit internal state. the main advantage of this algorithm is, it is having smaller block size compared to other algorithms and provide sufficient security even though the block size is small. the overall structure of the hummingbird cryptographic algorithm includes four 16-bit block ciphers eki(i=1,2,3,4), four 16-bit internal state registers rsi (i= 1,2,3,4), and a 16-bit lfsr (linear feedback shift register). the 256-bit key with 4 internal registers and lfsr provides adequate security level. for each block eki., 256-bit secret key is divided into four 64-bit sub keys ki(i=1,2,3,4). fig. 1 initialization notations used rs1rs4 internal state registers to hold nonce value ek1ek4 state registers to hold the modulo addition values of 16 bit register rs with ek + modulo addition tv register to store the final output after 4 rounds of initialization 2.1 initialization the initialization process shown in fig.1 initialize the four internal states registers ek1 to ek4 and get the lfsr initial value before encryption starts. the four internal state registers are first loaded with four 16-bit random nonce values. taking rs1 + rs3 as input data, four block ciphers are consecutively executed four times and the states are updated accordingly as shown in fig.1. the final output after four iteration is shown in the register tv, which is used to get the initial value of lfsr and used to update the state rs3 in encryption process. 354 vanitha m and subha s:high throughput area efficient architecture for light ... the lfsr is used not only for rs3 updating but also to ensure that period of the internal states are at least 216. fig. 2 encryption notations used rs1rs4 internal state registers ek1ek4 16 bit register to hold the block cipher value after modulo addition + modulo addition pt plain text ct cipher text fig. 3 decryption notations used rs1rs4 internal state registers dk1dk4 16 bit register to hold the value after modulo subtraction modulo subtraction + modulo addition pti plain text cti cipher text lfsr linear feedback shift register advances in systems science and application (2015) vol.15 no.4 355 fig. 4 block cipher 2.2 encryption/decryption after the initialization process, encryption starts by taking the input as the plain text(pt) followed by the modulo 216 addition (shown in fig.2 as + ) of plain text with internal state register rs1 and the result is applied to the block cipher ek1. the whole process of encryption is shown in fig.2. 356 vanitha m and subha s:high throughput area efficient architecture for light ... this process is repeated for four times and produces the cipher text. when all the four block ciphers are completed, the rsi state register is updated accordingly. decryption process is just reverse operation of the encryption as shown in fig.3. 2.3 block cipher block cipher used in encryption process as shown in fig.4. it consists of four rounds of operation and a final round. one regular round operation consists of a key mixing step, a substitution step and a linear transformation step. the sub-key of 64-bit is separated into four 16-bit round keys which will be used in the next corresponding rounds respectively. in the key mixing process, plaintext block uses an exclusive-or with the round-key. the s-box produces the results step by step. the substitution round uses 4 serpent-type s-boxes with the input and output of 4-bits. table 1 shows the substitution value of 4bits used for substitution process. here each 4-bit input is substituted with another 4-bit to make confusion in output. for the final step it uses one substitute step. there is no linear transformation step in the final round. the linear transformation step is defined in equation (1) where an xor operation is performed between din and 6 times right shifted value of din and 10 times right shifted value of din. dout = din∧(din << 6)∧(din << 10) (1) table 1 substitution s box values 0 1 2 3 4 5 6 7 8 9 a b c d e f s1 8 6 5 f 1 c a 9 e b 2 4 7 0 d 3 s2 0 7 e 1 5 b 8 2 3 a d 6 f 6 2 9 s3 2 e f 5 c 1 9 a b 4 6 8 0 7 3 d s4 0 7 3 4 c 1 a f d e 6 b 2 8 9 5 3 proposed architecture until now, hummingbird is implemented across different target platform. each of that implementation is either focusing area optimization by looping technique or throughput optimizing by loop unrolling of block cipher. the proposed block cipher is similar to the previous models with a modification that cipher loop is partially unrolled. this architecture is a compromise between earlier looped and loop unrolled architecture so that it tries to optimize both area and throughput. it requires 6 xors, 12 s-boxes, a linear transform and 2 multiplexers. here a plain text can be encrypted in 8 clock cycles. so the proposed architecture is an optimization of area and speed. fig.5 shows the proposed block cipher architecture. the four rounds of the encryption are unrolled and the three operation in a round like substitution, permutation and linear transformation is executed advances in systems science and application (2015) vol.15 no.4 357 in sequence. since the four rounds are executed in parallel it requires 4 times duplication of hardware but at the same time the number of clock cycle required is only 1 and for the last round it require one more extra. this encrypted block of 16 bit data is passed for next 5 rounds in a looped pipelined manner with pipelining register r1 and r2. the final encrypted data is taken in a 64 bit register at every clock cycle so it require 5 more clock cycle so in total it require 7 clock cycle for getting the encrypted data. the encryption block (ee) is represented from ee1-ee4 which operates in parallel. instead of using the four sbox as such in the design, a hardware friendly sbox is selected from the sbox given in the table 1 and it is repeated four times to improve the area and performance. from the table 3, it is found that minimum hardware required for s3 implementation. so it is selected as the hardware friendly sbox’s and used in the further implementation. the overall encryption core architecture is shown in fig.6. system is made reset before the first encryption starts. upon reset, the control signals data sel, rss el, key sel,init encr and the counters reset to zero. data sel is used to select corresponding input data to the block cipher, rs sel is used to select required internal register during operation, key sel is used to select corresponding sub keys and init encr is used to differentiate between initialization and encryption. at first the internal registers are loaded with the 16-bit random nonce values and core starts encrypting the rs1+rs3 for four iterations. each iteration takes 7 clock cycles as well as internal state updating and same block cipher architecture is reused each time as shown in the fig.5 which ensures maximum architecture reuse and hence better area and performance. after four iteration, the initialization is completed and the control signal init encr set to 1. once the initialization is completed, i th plain text is taken as the input and after 7 clock cycles, we get i th cipher text. during the above procedure, the corresponding sub keys, internal register and input data are selected according to the control signals key sel, rs sel and data sel respectively. the control signals are generated based on two counters. control signals rs sel and data sel should be updated after each block cipher operation so it is controlled by a block counter. the init encr signal is updated after initialization process so it is controlled by a round counter. after each iteration, the internal states are updated accordingly. the lfsr is initialized after initialization process and updated after each encryption. the proposed architecture is not altering the algorithm; it is an optimized hardware implementation of the same. 358 vanitha m and subha s:high throughput area efficient architecture for light ... fig. 5 proposed block cipher advances in systems science and application (2015) vol.15 no.4 359 table 2 cipher comparison looped loop unrolled proposed (partial loop unrolled) area 5xor 8sbox 1linear transform 2mux 8xor 20sbox 4linear transform no mux 6xor 12sbox 2linear transform mux clock cycles (encryption) 16 4 7 fig. 6 overall architecture 360 vanitha m and subha s:high throughput area efficient architecture for light ... table 3 area requirement of different sbox implemented on spartan-3 xcs200 fpga platform sbox #lut #flipflop total occupied slices s1 187 16 101 s2 184 16 99 s3 184 16 98 s4 186 16 99 fig. 7 proposed architecture against scan attack 3.1 cryptanalysis of proposed hummingbird architecture cryptographic processors are subjected to side channel attack. if the processor uses the intermediate registers, the probability of easy hacking is more. the scan chain based flip-flop (ff) used in the design synthesis is more vulnerable to scan attack. the hackers can easily run the processor for 2 to 3 rounds by operating in normal mode and then switch over to test mode which can retrieve all the data stored in the intermediate register made of scan based ff. the hackers can hack the data by running the processor for few clock cycles by applying all possible inputs and outputs vectors in a reversal fashion. based on the hamming distance between the input pairs the plain text can be hacked. usage of scan based ff cannot be avoided because it facilitates the easy testability of a ics and this opens an easy way for the hackers to steal the data. a response compactor register is used instead of using normal ff to store the intermediate results. the response compactor is a multiple input signature advances in systems science and application (2015) vol.15 no.4 361 register (misr) which compresses the scan ff intermediate data in a compact register by xor operation as shown in the fig.7. this will reduce the area and simultaneously make the architecture resistance to scan attack. we could see this architecture provide two hamming distance (1 and 8) which are at extreme so that we get 40 ≈ 25.32 possible values for each key byte. therefore, on average, the final key hypotheses is (25.32)16 = 285.15, which cannot be brute-forced easily. basically this architecture is resistance to attacks like birthday attack, differential attack and linear attack. 3.2 validation of the proposed architecture the encryption and decryption block for 16 bit is modeled with verilog hdl and simulated using modelsim. the functional verification is shown in fig.8 and 9. the asic implementation is done by synthesizing the design with rtl compiler with 45nm technology file from tsmc. the backend of the design is done using soc encounter. the complete chip layout is shown in fig.10 with design specification. the cryptanalysis on the proposed architecture is performed, since we have used the iterative architecture for compressing the scan chain ff response in a response compactor it shows that it is very difficult for brute-forced attach on our architecture. 4 results and discussion the architecture for encryption core as well as decryption core is modeled using verilog hdl and simulated in modelsim. the architecture is synthesized and implemented using cadence rtl complier and encounter. for comparison with the previous implementations, the same architecture is also implemented using xilinx ise 13.2 by taking low cost fpga sparten-3 xc3s200 in package ft-256 with a speed grade of 5. fig. 8 simulation result of encryption fig. 9 simulation result of decryption 362 vanitha m and subha s:high throughput area efficient architecture for light ... table 4 cadence synthesis report in 45nm item area power(uw) frequency 4900(mhz) cell cell area encr cipher 535 1126 794.09 218.91 decr cipher 463 1025 735.92 244.61 encryption 1393 3996 845.86 164.17 decryption 1956 5494 1623.29 216.07 table 5 performance comparison of fpga implementations design algorithm key size block size fpga area (slices) memory blocks frequency (mhz) through put (mbps) efficiency (mbps/ #slice) x.fan et al.[6,8] hummingbird (throughput opt) 256 16 spartan3 xc3s2005 273 0 40.1 160.4 0.59 hummingbird (area opt) 253 0 66.1 66.1 0.26 ismailsan et al.[10] hummingbird 256 16 spartan3 xc3s2005 40 2 260.8 55.64 1.38 biaomin[9] hummingbird (throughput opt) 256 16 spartan3 xc3s2005 230 0 39.8 157.6 0.68 hummingbird (area opt) 183 0 72.9 72.9 0.39 poschmann et al.[4] present 80 64 spartan3 xc3s4005 176 0 258 516 2.93 128 64 202 0 254 508 2.51 kaps et al.[11] stea 128 64 spartan3 xc3s505 254 0 62.6 36 0.14 yalla et al.[12] hight 128 64 spartan3 xc3s505 91 0 163.7 65.48 0.72 mace et al.[13] sea 126 126 virtex-ii xc2v4000 424 0 145 156 0.368 this work hummingbird 256 16 spartan3 xc3s2005 255 0 69.96 140 0.54 table 4 shows the 45nm synthesis report in cadence. no comparison is made with this report because it is done for the asic implementation of the algorithm. comparison is done with fpga implementation of the same with existing implementations. table 5 shows the comparison of our implementation with other existing implementations as well as other few cryptographic algorithm implementations. x.fan proposed throughput optimized implementation of hummingbird with 273 slices with a throughput of 160.4mbps,[8] x.fan implemented area opadvances in systems science and application (2015) vol.15 no.4 363 timized hummingbird with 253 slices and 66.1mbps throughput.[6] comparing with [8] our implementation has better area performance (-6.6%) but throughput is slightly lower. comparing with [6] our implementation has a good throughput performance (+111%) but area is slightly higher. strictly speaking, our implementation is a compromise between area optimized design and throughput optimized design. fig. 10 simulation result of decryption this clearly shows that the proposed design is optimized in terms of both area and throughput. ismail san has very good area performance but it utilizes the embedded memory in the fpga and it utilizes instruction stored in the embedded memory.[10] so this implementation cannot be assumed as a complete hardware implementation. biao min throughput oriented design has better performance compared to our design.[9]but the proposed design have better throughput compared to area oriented design of the same [9]. the algorithm present [4] has better area and throughput performance than hummingbird implementation but it required comparatively larger fpga xc3s400 which intern increases the cost. xtea[11] and hight[12] have very less throughput compared to hummingbird implementation and for sea[13] requirement of area is very high. the loop unrolling can increase the throughput but simultaneously increase the area so an optimum of around 50% rounds is unrolled to have a compromising solution for area and performance. 5 conclusions this paper presented an efficient vlsi architecture for ultra-light weight cryptographic algorithm named hummingbird. previous implementation of this algorithm was focusing on optimizing either the throughput or area. but this paper presented a design which is optimized in terms of both in area as well as throughput. the proposed design is a compromise between both area optimized and 364 vanitha m and subha s:high throughput area efficient architecture for light ... throughput optimized schemes. the hummingbird algorithm has good efficiency compared to all other algorithms and also it has smallest block size compared to all. the experimental results shows that the hummingbird algorithm is very much useful for resource constrained embedded devices. the initialization of internal registers is done by loading some random nonce values which is highly influences the security of the algorithm. the generation of these random values on the hardware is done using an lfsr. references [1] d. hong, j. sung, s. hong, j. lim, s. lee, b. s koo, c. lee, d. chang, j. lee, k. jeong, h. kim, j. kim and s. chee. (2006), “hight: a new block cipher suitable for low-resource device”, proceedings of ches 2006, volume 4249 of lncs, pages 46-9, springer. [2] f. x standaert, g. piret, n. gershenfeld and j.-j. quisquater. (2006), “sea: a scalable encryption algorithm for small embedded application”, smart card research and applications, proceedings of cardis 2006, volume 3928 of lncs, pages 222-236, springer-verlag. [3] g.leander, c.paar, a. poschmann and k. schramm. (2007), “new lightweightdes variants”, fast software encryption. [4] a.poschmann, a.bogdanov and l.r knudsen. (2007), present: an ultralightweight block cipher, springer. [5] d. engels, x. fan, g. gong, h. hu and e. m. smith. (2010), “hummingbird:ultra-lightweight cryptography for resourceconstrained devices”, toappear in the proceedings of the 14th international conference on financial cryptography and data security fc 2010, berlin, germany:springer-verlag. [6] x.fan, g. gong, k.lauffenburger and t.hicks. (2010), design space exploration of hummingbird implementations on fpgas, techincal report. [7] xinxin fan, honggang hu, guang gong1, eric m. smith and daniel engels. (2009), “lightweight implementation of hummingbird cryptographic algorithm on 4-bit microcontrollers”, institute of electrical and electronics engineers. [8] xinxin fan, guang gong, ken lauffenburger and troy hicks. (2010), “fpga implementations of the hummingbird cryptographic algorithm”, 9781-4244-7812-5/10/, ieee. advances in systems science and application (2015) vol.15 no.4 365 [9] biao min, ray c.c. cheung and yan han. (2011), “fpga-based highthroughput and area-efficient architectures of the hummingbird cryptography”, 978-1-61284-972-0/11/, ieee. [10] ismail san and nuray at. (2011), “compact hardware architecture for hummingbird cryptographic algorithm”, 21st international conference on field programmable logic and applications, 978-0-7695-4529-5/11, ieee. [11] j.-p. kaps. (2008), “chai-tea, cryptographic hardware implementations of xtea”, indocrypt 2008, lncs, vol.5365, pp.363-375, springer. [12] p. yalla and j.p. kaps. (2009), “lightweight cryptography for fpgas”, international conference on re-configurablecomputing and fpgas reconfig’09. [13] f. mace, f.x. standaert and j.j. quisquater. (2007), “fpgaimplementation(s) of a scalable encryption algorithm”, ieeetrans, very large scale integ, (vlsi) syst.vol.16, no.2, pp.212-216. corresponding author vanitha m can be contacted at: mvanitha@vit.ac.in adv syst sci appl 2019; 03; 1-10 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/753 test bench algorithms for catamaran roll simulation ivan lipko* sevastopol state university, sevastopol, russia e-mail: ivanlipko@yandex.ru received may 28, 2019; revised august 21, 2019; published october 1, 2019 abstract: creation of catamaran seaworthiness system requires two stages: computer simulation of models and prototyping of software and hardware. to implement second stage we made test bench to perform catamaran roll simulation, computation and collecting experiment data. possibility of making computations like on real onboard vessel deck is required. thus test bench consists of a stewart platform with microcontrollers and computer connected to them. first performs catamaran deck roll and second performs calculations. 3-dof sensor mounted on movable plate of platform gives us roll and pitch measurements. using transfer function approximation of pierson-moskowitz spectrum and runge-kutta 4-order method with constant step size we developed algorithms for sea-wave simulation, roll generation, data filtering and common calculations. we made a set of command-line software based on these algorithms. utilizing all possibilities of test bench allows us to collect statistics and compare recorded data. keywords: stewart platform, catamaran roll, test bench, simulation, motion platform. 1. introduction this article is devoted to developing of algorithms for test bench. test benches are complex devices which consist of several parts and designed to test any electronic devices or software in natural environment. in comparison with numerical simulation researcher has conditions when his theoretical models are not limited by math representation of real world. the industrial scope application of the test bench is the testing of electronic devices for catamarans. simulation is a one of important steps in developing any electronic devices and systems. at first one performs a numerical simulation. the goal of this step is research of the system, synthesis of controls and algorithms. at second one performs a hardware testing on conditions close to real. the goal of this step is proof of theoretical results. then engineers carry out onboard experiments. and final is exploitation. the test bench we discuss aims at the simulation of the catamaran roll and to test and to debug onboard measurement systems like on real environmental conditions and collect data to verify theoretical results. it is known that in vessel researching and simulation procedures researches use wave generation pool [6]. in this pool a small prototype of real vessel is swinging in the water. special wave generator creates waves similar to a real sea. researcher watches on vessel position, roll angles, etc. this method is very good to research vessel seaworthiness but has several restrictions: difficult to repeat conditions of experiment, require special laboratory with expensive equipment. in our case, we need to reproduce vessel roll as many times as possible in a limited time. mathematical descriptions of the real sea disturbances and catamaran are considered sufficient. that is why we develop test bench instead of using the special pool. usually the necessary algorithms for the bench include methods to prepare, to conduct experiments and to collect the results. preparation of experiment includes a generation of theoretical path of events which must be registered. conduction and collection include a control of all actions executed by test bench, registration and saving of a data. in our case there *corresponding author: ivanlipko@yandex.ru mailto:ivanlipko@yandex.ru 2 i. lipko copyright ©2019 assa adv. in systems science and appl. (2019) are respectively: trajectory generation of catamaran roll, filtering the data, collecting data to files. the test bench consists of the stewart platform with bottom fixed part and top mobile part which corresponds to vessel deck and pc-based operator workspace and microcontrollers. there are known several examples of stewart platform [15] usage to simulate roll of vessels and sea-offshore construction under sea disturbances [13]. some of them use two merged platforms to simulate sea-waves (bottom platform) and vessel deck (top platform) [9]. the same situation is in the other fields: automobile [1, 8, 14], aircraft and helicopters [16], where simulators are used to train machine operators or to test devices. to move mobile part of the platform most of them use linear actuators such as hydraulic, electro-mechanic or pneumatic. these actuators have many advantages, but very expensive. in contrast to them we use rotated step motors which less expensive. we have several conceptual requirements for test bench and operator workspace: 1. simple intuitive inputs and outputs. there are many situations when one requires interpreting results according to reality: real vessel has horizontal velocity, vessel hull is influenced by wind and sea waves which have an angle to vessel direction, wind velocity, etc. therefore appropriate mathematical models of catamaran and external disturbances are required. 2. repeat experiment as many times as one needs. ideal actions sequence is a loop: set input, run experiment, collect, show and explore data, export data to other applications. e.g. we developed several predictive algorithms [2, 3, 12] which must be tested on the real applications. repeating of experiments allows us to verify the quality of the algorithms. 3. test bench must support mounting of any available microelectronic or other sensors. due to these requirements the key development points are: 1. compute trajectory of catamaran roll on a microcontroller or a pc. so we use linear models when it possible. 2. reproduce catamaran roll trajectory by stewart platform. 3. collect angle and velocity data from sensors mounted on mobile part of platform. using these measurements estimate full state vector and produce special-task computations (e.g. stability, overpitch probability, seaworthiness, etc.). the paper is organized as follow. test bench based on the stewart platform is briefly described in section 2, where block diagram of its construction is discussed. the mathematical models of simulated catamaran, sea-wave disturbances and state observation are presented in section 3. results of simulation and experimental study are given in section 4. concluding remarks and the future work intentions are presented in the last section. 2. construction this section will describe a bench construction and its function diagram. test bench (fig. 2.1) is designed for debugging of predictive algorithms of catamaran roll based on large deviation theory. physically it consists of stewart platform and pc connected to microcontroller which drives motors (fig. 2.2). at the same time there are two logical parts which allow to reproduce catamaran roll and to collect and to process data from the sensors. the first part is the stewart platform that reproduces the movement of the ship deck disturbed by sea wave. the bottom part of platform is fixed and top part is movable. this version of stewart platform uses rotated step motors and legs connecting bottom and top parts of the platform (details in [4]). ball joints in the sole of the legs are a weakness of this construction because it leads to small rattling in motion. noise is not visible by eyes but noticeable by sensor. to reduce this noise we use ballast weight on top surface. motivation of use it is cheap cost and availability, simple control. the pitching trajectories are generated on the computer and converted into data for the microcontroller which controls platform motors. test bench algorithms for catamaran roll simulation 3 copyright ©2019 assa adv. in systems science and appl. (2019) fig. 2.1.test bench the second part (blocks “sensor”, “filter and estimate”, “general computations”) is a prototype of device for reading the roll data and calculating probability estimates. roll data (roll angle and velocity, pitch angle and velocity) are taken by a sensor mounted on the movable surface of platform and transmitted to the computer where calculations are made and displayed on the screen: current state of the vessel, estimates of the probability of large deviations. the blocks described above are modules of developed programs and perform certain tasks. all programs (unless otherwise written) are written in c++ in the qt creator and executed under the console environment. console usage is convenient for creating a test bench user interface application. block “roll generation”. the pitching trajectory generator is program designed to create trajectories of the catamaran that the platform will reproduce. the program started using command line parameters or a special configuration file. the parameters include initial state of the catamaran, seed-number for the random number generator, service flags. as a result the program creates a csv-file containing in each line the time countdown and the current catamaran state: angle and velocity of roll, level of heave, angle and velocity of pitch. to generate trajectories we use 4-order runge-kutta integration method with a discrete step. the system of differential equations describing the catamaran under the influence of wind-wave disturbances discussed in section 3.2. block “motor calculation”. this block solves inverse kinematics problem and translates roll trajectory motion to the motors control signals of the stewart platform. csv file with trajectory from previous block is an input of the program (written in c#). output of this block is csv file with lines which contain the angles of platform motors. a special program implements the execution of platform movements [4]. block “filter and estimate”. this block performs data reading and restoring the state of catamaran. 3dof-sensor (mpu6050) and microcontroller nucleo-f401re are used for reading angles and their velocities. the values of the observable state vector are passed to the input of the filter and recovery module. filtration and estimation of the full state vector is performed in accordance with the section 3.3. the restored full state vector is saved to a file and passed into next block. 4 i. lipko copyright ©2019 assa adv. in systems science and appl. (2019) fig. 2.2. block diagram block “general computations”. this block implements the predictive algorithms which necessary for other research. input of the block is full state estimates the resulting vector is a task-special data displayed and saved to a file. 3. algorithms this section describes the mathematical model of the simulated catamaran and its control system aimed to reduce the amplitude of the roll and the angle of the roll; the mathematical description of the wind-wave disturbances acting on the catamaran, used in the generation of the trajectory of the catamaran and the estimation of state vector. 3.1. catamaran model the catamaran is a vessel with two identical symmetrical hulls connected by a deck [11]. to reduce pitching of the boat it have t-foil and thruster flaps, which create a forces (𝐹𝑇 and 𝐹𝐹 respectively) and moments (𝑀𝑇 and 𝑀𝐹, respectively). the main dimensions of the catamaran are presented in table 3.1, block diagram on fig 3.1. table 3.1 the catamaran dimensions name value length, m 90 width, m 25.96 draught, m 2.6 displacement, t 734.54 in general there is a relationship between all types of ship pitching in the motion model [7]. in this case we are interested in roll, pitch angles and heave of the catamaran. the space-state equations for the catamaran will be as follows: { �̇� = 𝐴𝑥 + 𝐵(𝑢 + 𝑤) 𝑦 = 𝐶𝑥 , (3) stewart platform roll generation motors sea state path computer sensor motor calculation motor angles filter and estimate general computations roll, pitch state vector results test bench algorithms for catamaran roll simulation 5 copyright ©2019 assa adv. in systems science and appl. (2019) fig. 3.1.catamaran controller and signals where 𝑥 = [η̇, ζ̇, η, ζ, θ̇, θ] is state vector including heave velocity, pitch velocity, heave, pitch, roll velocity, roll; 𝑢 = [ 𝐹𝑇 + 𝐹𝐹 𝑀𝑇 + 𝑀𝐹 ] is control vector including forces and moments of actuators; 𝑤 = [ 𝐹𝑤 𝑀𝑤 ] is wind-wave disturbances including forces and moments acting on vessel hull; state and input matrices are: a = [ −0.9073 −25.1097 −14.1503 −17.4945 0.001 0 0.0514 −0.503 0.2442 −12.4 0.001 0 1 0 0 0 0 0 0 1 0 0 0 0 0.001 0.001 −0.01 −0.01 −5 −15 0 0 0 0 1 0 ] , 𝐵 = [ 0.0082 0.0000083 −0.00016 0.000017 0 0 0 0 0.002 0.0000083 0 0 ] , and output matrix is 𝐶 = [ 1 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 1 ]. for simulation and path generation matrix 𝐶 is identity matrix, but in real situation measuring a heave is very difficult task. we have actuators that will be controlled by feedback controller. in this article we use lqregulator which will reduce pitching amplitudes. so matrices 𝐴, 𝐵, 𝐶 are suitable to make controller with feedback matrix 𝐾 = [ 59.3026 26.6427 18336.7 2006.9 −1807.6 −2950.1 −1.2133 2.5576 14.2237 −2.8566 2.1435 93.5241 ] and control 𝑢 = −𝐾𝑥. as a result we have a closed-loop system { �̇� = (𝐴 − 𝐵𝐾)𝑥 + 𝐵𝑤 𝑦 = 𝐶𝑥 . (4) equation (4) is integrated via 4-order runge-kutta method with constant step size 0.01. wind-wave disturbances 𝑤 is input of the system. sea-wind forces and moments controller catamaran actuators t-foil signal flaps signal x moments forces forces moments 6 i. lipko copyright ©2019 assa adv. in systems science and appl. (2019) 3.2. wave-wind disturbances to describe sea-wind disturbances we use pierson-moskowitz spectrum in common form 𝑆(ω) = 8.1 ∙ 10−39.82 ∙ ω−5 ∙ exp (−0.74 ∙ 0.92 ∙ ω−4). above numbers and parameters corresponds to sea state with 2.14 m significant wave height and 10 m/s wind velocity [5]. pitching of the catamaran arises as a result of the external forces and moments actions created by sea waves. the simulation of these forces and moments is obtained by using two transfer functions (see fig. 3.2). transfer functions are the linear approximations of the corresponding nonlinear relations (e.g. spectrum in fig. 3.3). first, white noise is applied to the input of the pierson-moskowitz spectrum linear approximation block and then to the block of linear approximation of forces and moments. pierson-moskowitz spectrum linearized by 2-order transfer function [5] has form 𝐻𝑃𝑀(𝑠) = 𝐾ω𝑠 𝑠2+2λω0𝑠+ω0 2, where𝐾𝜔 = 2λω0σ is amplifier constant, σ2 = max 0<ω<∞ 𝑆(ω) is wave disturbance intensity constant, σ is damping factor, ω0 is dominating wave frequency. damping factor 𝜎 is calculated using nonlinear least square method. fig. 3.2.sea-wind forces and moments generation fig. 3.3. pierson-moskowitz (pm) spectrum (dashed) and linear approximation (solid) block “linear approximation of forces and moments” approximates “wave to force” function by constant 𝐾𝑓. it is tuned so that the amplitude of roll catamaran satisfies sea state statistics [5]. resulting signal is an input 𝑤 to system (4). 3.3. state estimation in this section we describe in short about filtering the data from sensor. to do this we use kalman filter which performs data filtering and estimation of full state vector. remind that from sensor we receive only roll, pitch and their velocities and we also need to know heave and heave velocity. linear approximation of spectrum white noise linear approximation of forces and moments sea state, spectrum wave amplitude forces, moments p o w er , ra d /s frequency, rad/s test bench algorithms for catamaran roll simulation 7 copyright ©2019 assa adv. in systems science and appl. (2019) to programming filter we need to discretize object system (4) and made discrete version of kalman filter so called “delayed estimator” [10]. delayed estimator is very simple to programming, but it is enough to get sufficient quality: �̂�𝑛+1,𝑛 = (𝐴𝑒 − 𝐿𝑒𝐶𝑒)�̂�𝑛,𝑛+1 + [𝐵𝑒 − 𝐿𝑒𝐷𝑒 𝐿𝑒] ∙ [𝑂4,1 𝑦𝑟𝑎𝑤]𝑇, [ �̂�𝑛,𝑛−1 �̂�𝑛,𝑛−1 ] = [ 𝐶𝑒 𝐼6 ] �̂�𝑛,𝑛−1 + [ 𝐷𝑒 𝑂10,4 𝑂6,4 𝑂6,4 ] ∙ [𝑂4,1 𝑦𝑟𝑎𝑤]𝑇, where�̂�𝑛,𝑛−1 is catamaran state estimate, �̂�𝑛,𝑛−1 is catamaran output estimate, 𝑦𝑟𝑎𝑤 is raw sensor data vector. matrices 𝐴𝑒, 𝐵𝑒, 𝐿𝑒, 𝐷𝑒 are the following: 𝐴𝑒 = [ 0.8770 −0.0091 −0.0762 −0.0119 0.0053 0.0244 −1.0854 −0.0195 −0.7805 −0.1685 −0.0077 0.4085 0.0067 0.0975 0.5590 0.0360 −0.0055 0.0040 −0.1307 −0.7889 −5.9193 −0.0658 −0.0661 −0.0626 0.0004 −0.0007 0.0075 0.0013 0.9390 0.0930 0.0013 0.0231 0.0866 0.0200 −1.1861 0.8465] , 𝐵𝑒 = [ 0.0597 0.0857 0.0001 0.0003 −0.0674 0.5716 0.0026 0.0359 −0.0077 −0.0980 −0.0004 −0.0067 0.1209 0.7826 0.0026 0.0231 −0.0004 0.0007 0.0001 0.0003 −0.0013 −0.0229 −0.0002 0.0004] , 𝐶𝑒 = [ [ 𝐼2 02,4 02,4 𝐼2 ] 𝐼6 ], 𝐷𝑒 = 𝑂10,4, 𝐿𝑒 = [ 0.0597 0.0857 0.0001 0.0003 −0.0674 0.5716 0.0026 0.0359 −0.0077 −0.0980 −0.0004 −0.0067 0.1209 0.7826 0.0026 0.0231 −0.0004 0.0007 0.0001 0.0003 −0.0013 −0.0229 −0.0002 0.0004] . matrices above include 𝐼𝑛 that is identity matrix with size n x n; 𝑂𝑛,𝑚 is zero matrix with n rows and m columns; and computed via kalman filter synthesis procedure with parameters 𝑄 = 𝐸(𝑤𝑤𝑇) = 𝐼2 ∙ 107, 𝑅 = 𝐸(ζζ𝑇) = 𝐼4 ∙ 10−1, 𝑁 = 𝐸(𝑤ζ𝑇) = 𝑂2,4, where ζ is measurement noise. 4. results, experiments in this section we show several examples of test bench experiments. to make sure in quality of results we perform experiments in range of input parameters. here some of them. all test bench experiments are configured for better compare and interpretation via natural parameters: wind velocity, relative bearing, initial state vector of catamaran and seed-number. fig. 4.1 and fig. 4.2 show examples where test bench filter outputs are compared with numeric simulation results. raw data from sensors were scaled and filtered. the test bench reproduces catamaran roll with satisfactory quality for our applications. the initial state vector of catamaran is equal to zero in next examples. 8 i. lipko copyright ©2019 assa adv. in systems science and appl. (2019) fig. 4.1.errors in catamaran roll (top) and pitch (bottom) between simulated and experimental angles corresponding to a) wind 13.45 m/s, seed-number 11622; b) wind 13.23 m/s, seed-number 5; c) wind 12.96 m/s, seed-number 25 a) b) c) fig. 4.2.test bench and runge-kutta outputs a) wind 13.45 m/s, seed-number 11622; b) wind 13.23 m/s, seed-number 5; c) wind 12.96 m/s, seed-number 25 test bench algorithms for catamaran roll simulation 9 copyright ©2019 assa adv. in systems science and appl. (2019) experiment data in fig. 4.1 and 4.2 have some rattling and noise influenced by ball joints in the sole of the legs connecting the bottom and the top parts of platform. paths from these figures can be associated via indices “a”, “b”, “c”. 5. conclusion the algorithms for reproducing a catamaran roll on test bench are developed. for better interpretation of results inputs of algorithms have nature-like parameters: wind velocity, direction, catamaran initial state and other. mathematical models of algorithms include the catamaran and sea-wave disturbances. full state of catamaran is estimated by kalman filter. to improve estimating one should use another filter, e.g. extended kalman filter the test bench for reproducing a catamaran roll includes the stewart platform, sensors, microcontrollers and pc. the top part of the platform moves like a catamaran deck and results of experiments are collected by sensors. thus, it turns out to reproduce the same oscillations of the catamaran and to debug device prototypes that will be installed on the real vessel. software for test bench has been developed based on these algorithms. it helps to create the trajectories of the catamaran deck roll, to filter data obtained from the sensors, and to estimate the full state vector of the catamaran. these developments can also be used in any other similar area where it is necessary to accurately reproduce the specified motion of the object, for example, any other vessels, submarines or aircraft. acknowledgements studies were supported by the ministry of science and higher education of russian federation (project rfmefi57817x0259) mechanic and electro-mechanic components were assembled by osadchenko a.e. and lazarev v.b. we would also like to show our gratitude to the reviewer for his useful insights and for comments on an earlier version of the manuscript. references 1. andrievsky, b. et al. (2014). control of pneumatically actuated 6-dof stewart platform for driving simulator, 2014 19th international conference on methods and models in automation and robotics (mmar), miedzyzdroje, 663-668, doi: 10.1109/mmar.2014.6957433. 2. dubovik, s. a. & kabanov, a. a. (2017). quasipotentials in synthesis of control systems based on knowledge. proceedings of international conference on industrial engineering, applications and manufacturing (icieam), st. petersburg, 1-4. doi: 10.1109/icieam.2017.8076123. 3. dubovik, s. a. & kabanov, a. a. (2019). funkcional'no ustojchivye sistemy upravleniya: asimptoticheskie metody sinteza [functionally stable control systems: asymptotic methods of synthesis].moscow, infra-m, [in russian]. doi: 10.12737/monography_5b446a985cf9a5.11626044. 4. dubovik, s. a., kabanov, a. a. & lipko, i.u. (2017). sistema kontrolya krena sudna na baze analiza bolshih ukloneniy [control system of vessel roll, based on large deviations analysis]. перспективные системы и задачи управления, south federal university, rostovon-don, 196-209, [in russian]. 5. fossen, t. i. (2011). handbook of marine craft hydrodynamics and motion control. jhon wiley & sons, isbn 978-1-119-99149-6. 10 i. lipko copyright ©2019 assa adv. in systems science and appl. (2019) 6. hassani, v., alterskjær, s. a., fathi, d., selvik, o. & sæther, l. o. (2014). experimental results on motion regulation in high speed marine vessels, 22nd mediterranean conference on control and automation, palermo, 487-492, doi: 10.1109/med.2014.6961420. 7. kramar, v. (2018). the construction principle and system architecture of the automatic hold system of an unmanned vessel. 2018 international multi-conference on industrial engineering and modern technologies (fareastcon), vladivostok, 1-4, doi: 10.1109/fareastcon.2018.8602791. 8. krasnov, e. i., mikhaylov v. v., sergeev, s. l. & stuchenkov, a. b. (2015). development of a motion system for a training simulator based on stewart platform, international conference "stability and control processes" in memory of v.i. zubov (scp), st. petersburg, 99-101, doi: 10.1109/scp.2015.7342074. 9. lebrón, r., valente, v., sobczyk, m. & perondi, e. (2016). control of an electrohydraulic stewart platform manipulator as vessels motion simulator, 9th fpni ph.d. symposium on fluid power, v001t01a033. doi: 10.1115/fpni2016-1553. 10. lewis, f. l., xie, l. & popa, d. (2017). optimal and robust estimation: with an introduction to stochastic control theory, second edition. crc press. isbn: 9780849390081. 11. liang, l., yuan j. & zhang, s. (2016). application of model predictive control technique for wave piercing catamarans ride control system, 2016 ieee international conference on mechatronics and automation, harbin, 726-731, doi: 10.1109/icma.2016.7558652. 12. lipko, i. u. (2018). forecasting of catamaran roll threshold exceeding.2018 international russian automation conference (rusautocon), 9-16 sept. 2018, sochi, russia, ieee, isbn: 978-1-5386-4938-1, doi: 10.1109/rusautocon.2018.8501764. 13. rekdalsbakken, w. (2006). design and application of a motion platform for a high-speed craft simulator, ieee international conference on mechatronics, budapest, 38-43.doi: 10.1109/icmech.2006.252493. 14. shiong, c. y., jalil, m. k. a. & hussein m. (2009). motion visualization and control of a driving simulator motion platform, 2009 6th international symposium on mechatronics and its applications, sharjah, 1-5. doi: 10.1109/isma.2009.5164839 15. stewart, d. (1965-66).a platform with six degrees of freedom, proceedings of the institution of mechanical engineers, vol. 180, 371-378. 16. villacís, c. et al. (2017). real-time flight simulator construction with a network for training pilots using mechatronics and cyber-physical system approaches, 2017 ieee international conference on power, control, signals and instrumentation engineering (icpcsi), chennai, 238-247, doi: 10.1109/icpcsi.2017.8392169. adv syst sci appl 2017; 4:1–13 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/481 predictive feedback control method for stabilization of continuous time systems lina a. shalby1,2 1moscow institute of physics and technology, dolgoprudny, moscow region, russia 2faculty of women, ain shams university, cairo, egypt abstract: recent researchers have studied the chaotic phenomenon and introduced methods to control chaos. one of these methods is the predictive feedback control (pfc). pfc method has been used to stabilize only the discrete chaotic maps. in this paper, we generalize the pfc method by extending it to stabilize continuous time systems. the numerical simulations are carried out to solve the system using euler method for its simplicity, and then the pfc method is applied to the discretized system. the controlled trajectory converges to an unstable equilibrium point via small control action. the stability analysis is shown compared to that of the continuous system in details. also the choice of the controlling input of the pfc method and its properties are discussed. lorenz system and rössler system are the well known chosen chaotic examples. the pfc method is applied to each of them, showing stability at each of their equilibrium point. with a very small control, the behaviour of the system is completely changed from chaos to stable. keywords: continuous time systems, chaos control, euler method, predictive feedback control, unstable equilibrium point 1. introduction chaos theory is one of the recent fields of study that attracts the interest of scientists. two most important properties of chaos are the high sensitivity of trajectories to initial conditions and to parameter values. in the last decades, many researchers concerned with how to control chaotic phenomena depending on these two properties. they aimed to bring a trajectory to a small neighborhood of a desired location in the chaotic attractor, using a small perturbation. pioneering work ott-grebogi-yorke [12] proposed ogy method for stabilizing chaos. it could only be applied to discrete data, so continuous time systems should first be made discrete time by using the poincare map. then it was followed by many other publications adjusting and extending the ogy method as in [6] and [7], see also survey [1]. pyragas [14,15] proposed the delayed feedback control (dfc) to stabilize the unstable periodic orbits (upo). the control input was given as the multiplication of a gain with the difference between current system state and a state with the time delay as it is discussed in [4]. dfc method is firstly constructed for continuous time systems and extended to the discrete-time case in [18]. many applications, such as in [3, 10], were controlled by using dfc method. also, many developments for dfc method were performed, see [8, 11]. one of the recent surveys on dfc is presented in [16]. ushio [19] introduced the idea of predicting control in his paper and it was extended in [2] to stabilize continuous time systems using predictive-based control method. predictive ∗corresponding author: lina.khamis@gmail.com, lina.khamis@women.asu.edu.eg http://ijassa.ipu.ru/ojs/ijassa/article/view/481 2 lina a. shalby feedback control (pfc) [13] is an easily implemented method to stabilize unknown upos in discrete time dynamical systems. in pfc, the control input was given as the multiplication of a gain with the difference between two predicted future system states. pfc depends on applying small controls that completely change the nature of system’s behavior and devotes its attention to the most important problem of stabilizing or controlling chaos. detailed survey on existing approaches and methods for chaos control can be found in the handbook of chaos control [17]. in the next section of the paper, we extend pfc to stabilize chaotic continuous time dynamical systems solved numerically by euler method and the pfc method is analyzed. in section 3, the numerical simulations results are given for the most popular examples of chaos: lorenz system and rössler system. 2. predictive feedback control method we start with introducing the pfc method [13] for discrete-time system which depends on the system state and a control action u(x) as follows xn+1 = f (xn)− u(xn), (2.1) where f, u : rn → rn and u(x) = e(f p+2(x)− f p+1(x)), for a chosen control gain matrix e and any arbitrary forward predicted state p (p > 0 ), by denoting f p+1(x) = f (f p(x)), f 1(x) = f (x) — prediction of the map f for p+ 1 iterations forward. the idea is to choose e to stabilize unstable orbits or unstable equilibria points via small control action u(x). indeed state x and control u are of the same dimension n , and this strongly limits the application of the proposed technique to real-life problems. however there are various situations where such controls are applicable. for instance if one deals with classical analytical examples (e.g. rössler and lorenz systems) it is of interest to deal with their fixed points and orbits which are unstable. then pfc allows to stabilize such points and orbits. and the main peculiarity of the approach is the arbitrary small control action required. suppose n = 1, x∗ is an unstable equilibria point point of f (x), that is x∗ = f (x∗), µ = |f ′(x∗)| > 1. then for e = 1 µp+1(µ−1) , iterations of (1) converge to x∗. the most important fact is obtaining control u(x) small enough. indeed, if we deal with chaos control, iterations f p(x) remain bounded, while e is small for p large, because µ > 1. it is effective to choose p large enough (because larger is p, smaller is the control action). however, large predictive horizons imply computational error. thus the choice of p is some trade-of. in [13] the result is extended for n > 1 and for unstable orbits (i.e. for x∗ such that x∗ = f k(x∗) with k > 1). similar results are available for µ known approximately. the goal of the present paper is the application of the above results for continuous-time systems. let us consider a system of nonlinear autonomous differential equations ẋ = f(x), (2.2) where x ∈ rn , and f is assumed to be differentiable. an equilibrium point x∗ necessarily satisfies f(x∗) = 0. the following result on asymptotic stability is standard (see e.g. [5]). the equilibrium point x = x∗ of ẋ = f(x) is asymptotically stable if all eigenvalues of the jacobian matrix a = f ′(x∗) satisfy re(λi) < 0. from now on, we reduce a continuous time system to a discrete time map. this map is created from sampling the flow at discrete times tn = t0 + nh, n = 0, 1, 2, . . . , where the sampling interval h can be chosen on the basis of convenience. thus, a continuous time trajectory x(t) yields a discrete time trajectory xn = x(tn). by using euler method, the simplest numerical method to solve a differential equation, the system will be discretized easily to take the form xn+1 = xn + hf(xn) = f (xn), (2.3) copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 3 where f (x) = x+ hf(x). equation (2.3) is the simplest discrete form corresponding to the continuous time system. lemma 2.1 highlights the relation between the discretized function f and the continuous system function f . let µi denote the eigenvalues of the jacobian matrix m = f ′(x∗) = i + hf ′(x∗), f(x∗) = 0, while λi are eigenvalues of a = f ′(x∗). lemma 2.1: f (x) satisfies the following: (i) f(0) = 0, and f (x∗) = x∗ + hf(x∗) = x∗. (ii) let re(λi) = ui and im(λi) = νi, if ui < 0 and 0 < h < 2|ui| u2i+ν 2 i , then |µi| = |1 + hui| < 1, and if ui ≥ 0, then |µi| ≥ 1. lemma 2.1 shows that the discrete form for the continuous system has the same equilibrium point, and unstable equilibrium points of the continuous system imply instability of the discrete system. thus for the discretized system, the predictive feedback control method in [13] can be used to stabilize the system as illustrated for the vector case. thus we arrive to the following stabilization scheme for continuous system (2), where the matrix e could be considered to be the diagonal with entries εi. the choice of εi depends on the value of the eigenvalues µi of the jacobian matrix m , which could be real or complex. firstly, for | µi |< 1, we take εi = 0. secondly, if | µi |≥ 1 and µi ∈ r , then εi = 1 µp+1 i (µi−1) . thirdly, if | µi |≥ 1 and µi ∈ c, such that µi = ai + bij, j = √ −1, then the control input εi will be a 2× 2 matrix: εi = d−p−1(d − i2)−1 , where d = [ ai bi −bi ai ] . we restrict our analysis with fixed points, because orbits are fixed points of the iterated maps. the next theorem and its proof illustrate local stability of the system for these choice of the control gain. theorem 2.1: let f ∈ c1, x ∈ rn and assume x∗ is an unstable equilibrium point of continuous system (2.2) and its jacobian matrix m is not hurwitz. then x∗ is stable fixed point of the predictive feedback control equation (2.1) after discretizing system (2.2) numerically by using euler method (2.3) provided h is small enough. proof at x = x∗, matrix a has eigenvalues λi, with ui = re(λi) > 0 for some i. suppose that f (x) is h-step sized euler approximation of equation (2.2), given by equation (2.3), then the jacobian matrix m has the eigenvalues µi = 1 + hλi. by the chain rule, the jacobian matrix of f p(x) is associated with the eigenvalues µpi . let j is the jacobian matrix of the right hand side in equation (2.1) at x = x∗, which will have the form, j =m − e {mp+2 −mp+1} = m − emp+1 {m − i}, and j has eigenvalues ηi , i = 1, ..., n . then | ηi |=| µi − δ {µi − 1} |< 1∀i, for the following cases of µi : for all | µi |< 1, εi = 0 and ηi = µi. if | µi |≥ 1 and µi ∈ r, for εi = δh|λi| µp+1 i (µi−1) = δ µp+1 i , 1 < δ < 1 h|λi| , then | ηi |=| µi − δ(µi − 1) |=| 1 + hλi − δhλi |=| 1− (δ − 1)hλi |< 1. while if | µi |≥ 1 and µi ∈ c, for εi = δh | λi | (d−p−1(d − i2)−1), 1 < δ < 1 h|λi| where d = [ ai bi −bi ai ] has the same eigenvalues µi and dp has the eigenvalues µpi . then | ηi |=| µi − δh|λi| µp+1 i (µi−1) µp+1 i (µi − 1) |=| 1 + hλi − δh | λi ||< 1. hence the system is stable since the eigenvalues of j are less than unity. copyright c© 2017 assa. adv syst sci appl (2017) 4 lina a. shalby the next section illustrates the examples of chaotic systems and their stabilization achievement after adding the control term. the examples are clearly showing easy implementation and the efficiency of the pfc method. 3. examples in this section, we provide two examples of the pfc method applied to classical chaotic system. lorenz model is 3d system example, and rössler system is 4d example. 3.1. lorenz model lorenz [9] was the first who introduced the strange attractor notion and coined the term “butterfly effect”. lorenz system is described by differential equations ẋ = −σ(x− y) ẏ = −xz + rx− y ż = xy − kz (3.4) equation (3.4) has two nonlinear terms, and exhibits both periodic and chaotic motion depending upon the values of the control parameters σ, r and k. by using euler method the system will be discretized to the form xn+1 = xn − hσ(xn − yn) yn+1 = yn + h(−xnzn + rxn − yn) zn+1 = zn + h(xnyn − kzn) (3.5) where h is the step size for discretization of time t. the stability analysis of equation (2.1) can be performed by studying the discrete 3d function corresponding to it, f (x, y, z) = x− hσ(x− y)y + h(−xz + rx− y) z + h(xy − kz). (3.6) the system (3.4) has three fixed points s1 = (−b √ r − 1,−b √ r − 1, r − 1), s2 = (0, 0, 0) and s3 = (b √ r − 1, b √ r − 1, r − 1), which are the same as the equilibrium points of the continuous system (we denote 3d vector (x, y, z) as s). set σ = 10 and k = 2.67, and make r the adjustable control parameter. varying the values of r reveals a critical value at rc = 24.74. below rc the system decays to steady, non-oscillating, state. once r increases beyond rc, the continuous oscillatory behavior occurs and the system shows aperiodic behavior which lorenz called deterministic non-periodic flow which refer to chaos. for r = 28, the chaotic case, the equilibrium points are s1 = (−8.4906,−8.4906, 27), s2 = (0, 0, 0) and s3 = (8.4906, 8.4906, 27) and the eigenvalues corresponding to them are λs1 = λs3 = {−13.8569, 0.0934± 10.2001j} and λs2 = {−22.8277, 11.8277,−2.67}. let us use the step size h to be a small value equal to 0.001 which satisfies the condition due to the complex value of λs1 (i.e h < 2×0.0934 (0.0934)2+(10.2001)2 = 0.00179), and then the solution of the system is chaotic as shown in fig. 3.1. for the discrete form, the eigenvalues are µs1 = µs3 = {0.9861, 1.00009± 0.0102j} and µs2 = {0.99733, 1.0118, 0.9771}. the complex eigenvalues have the absolute value 1.00014. let us choose p = 5, the initial condition to be (5, 5, 5), and step size h = 0.001. starting with stabilizing the system at the origin, at which ε1 = ε1 = 0 for | µ(1,3) s2 |< 1, and otherwise we copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 5 fig. 3.1. phase space plot for lorenz system. choose δ = 2, ε2 = 2 (µ (2) s2 )p+1 . the matrix e is taken to be in the form e = [ 0 0 0 0 1.8638 0 0 0 0 ] and the first iterate of the decaying additive control term is −0.2040. fig. 3.2. the distance between s and 0 before and after stabilization. fig. 3.3. phase space plot of the controlled trajectory. copyright c© 2017 assa. adv syst sci appl (2017) 6 lina a. shalby fig. 3.2 illustrates the difference between the equilibrium (0, 0, 0) and the values of xtrajectory, y-trajectory and z-trajectory respectively, before stabilizing the system in blue color compared to those after applying the controlling term to the system that is coloured by red. fig. 3.3 is the phase space plot of the stabilized system. both figures show that the system is stable and all variables time series are converging to the origin starting from initial point (5, 5, 5). in case of the complex eigenvalues, µ (1) s1 = µ (1) s2 = 0.9861 and µ (2) s1 = 1.00009± 0.0102j, ε1 = 0, and d = [ 1.00009 0.01020 −0.01020 1.00009 ] ,ε2,3 = δh | λ2,3 | d−p−1(d − i2)−1 =[ 0.01030 −0.1978 0.1978 0.01030 ] , and e = [ 0 0 0 0 0.0103 −0.1978 0 0.1978 0.0103 ] . fig. 3.4. comparison between the controlled and chaotic trajectories for the equilibrium point (-8.4906, -8.4906, 27). fig. 3.5. phase space plot for the controlled trajectory for the equilibrium point (-8.4906, -8.4906, 27). in fig. 3.4 and fig. 3.5, the equilibrium point (−8.4906,−8.4906, 27) is stable. to stabilize the system at (8.4906, 8.4906, 27), the negative sign of the imaginary part of the complex eigenvalue is reversed, so that the matrix d will be [ 1.00009 −0.01020 0.01020 1.00009 ] . both fig. 3.6 and fig. 3.7 show the slow convergence of the trajectory after stabilization due to the point (8.4906, 8.4906, 27). it is important and interesting to test the effect of different initial conditions. for 50 random initial conditions, the convergence has been tested. at t = 10, the trajectories succeeded to reach the fixed points (0, 0, 0) (see fig. 3.8) copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 7 fig. 3.6. comparison between the controlled and chaotic trajectories with respect to each dimension with the equilibrium point (8.4906, 8.4906, 27). fig. 3.7. phase space plot for the controlled trajectory due to the equilibrium point (8.4906, 8.4906, 27). and (8.4906, 8.4906, 27) (as shown in fig. 3.9) for all initial points x0, y0 ∈ [−10, 10] and z0 ∈ [0, 30]. all the statements rely on local analysis. but in some cases, there was an effect similar to global stabilization, as it is seen in fig. 3.8 and fig. 3.9. this fact is a motivation for research of global behavior in future works. 3.2. rössler model rössler model of 4d phase space is one of the most famous hyperchaos models. for x ∈ r4, its defining equations are ẋ = −y − z ẏ = x+ αy + w ż = β + xz ẇ = −γz + τw (3.7) the chaotic behavior of the system (3.7) is observed for the parameter values α = 0.25, β = 3, γ = 0.5 and τ = 0.05. in this case, the equilibrium points are eq1 = (5.40833, 0.55470, −0.55470,−5.54700) and eq2 = (−5.40833,−0.55470, 0.55470, 5.54700), which are associated with the eigenvalues λeq1 = {5.504039, 0.103929, 0.0501791± 0.971052j} (the step size h must be less than 0.1061) and λeq2 = {0.101890, 0.0493731± 0.9986873j,−5.308963} (the step size h must be less than 0.0987) respectively. by solving the system with euler method, copyright c© 2017 assa. adv syst sci appl (2017) 8 lina a. shalby fig. 3.8. the stability is reachable at (x(10), y(10), z(10)) = (0, 0, 0) for 50 random initial conditions. fig. 3.9. the stability is reachable at (x(10), y(10), z(10)) = (8.4906, 8.4906, 27) for 50 random initial conditions. adjusting the step size h = 0.001(suitable for stability conditions stated in lemma 2.1), the continuous system translated to the discrete form with eigenvalues µeq1 = {1.005504, 1.000103929, 1.00005017± 0.00097105j} and µeq2 = {1.000101890, 1.000493731± 0.0009986873j, 0.9946910}. let us use the initial condition (−10,−6, 0, 10), then the system (3.7) is hyperchaos as shown in fig. 3.10 and fig. 3.11 for the phase space xyw and x− y plot, respectively. by setting p = 3, δ1,2 = 1.5 and δ3,4 = 5 for ε1 = δ1 (µ (1) eq1) 4 = 1.46743, ε2 = δ2 (µ (2) eq2) 4 = 1.49937, and ε3,4 = δ3,4 × h | λ(3)eq1 | ×d−4(d − i2)−1 = [ 0.01243 0.26011 −0.26011 0.01243 ] , then eeq1 =  0.01243 0.26011 0 0 −0.26011 0.01243 0 0 0 0 1.46743 0 0 0 0 1.49937 . while δ1 = 1.5 and δ2,3 = 40 for ε1 = δ1 (µ (1) eq2) 4 = 1.49939 , ε2,3 = δ2,3 × h | λ(2)eq2 | ×d−4(d − i2)−1 = [ 0.09053 1.99260 −1.99260 0.09053 ] , then eeq2 =  0.09053 1.99260 0 0 −1.99260 0.09053 0 0 0 0 0 0 0 0 0 1.49939 . the first iterate of the control term values are -0.00031, -0.00158, 0.00419, and 0.00074 to stabilize the system at eq1. and its first iterate values to stabilize the system at eq2 are -0.00237, -0.01207, 0, and 0.00074. the control term values are small negligible in the real life applications but they are enough to change the behavior of the system from ergodicity to be stable. fig. 3.12 and fig. 3.13 are showing the controlled phase space xyw due to eq1 and eq2, respectively. and fig. 3.14-3.17 are the trajectories convergence of x, y, z and w to the equilibrium points compared to the uncontrolled trajectories, respectively. copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 9 fig. 3.10. the strange attractors of rössler hyperchaos system xyw phase space plot. fig. 3.11. x− y plot for rössler hyperchaos system. for each example, the pfc method is successfully stabilized all the equilibrium points. it is also important to discuss the relation between the continuous system and its discretized form numerically, as they are difficult to be solved analytically. at this point, a question may arise. what is the continuous system that has stable attractor while its numerical solution is chaotic? copyright c© 2017 assa. adv syst sci appl (2017) 10 lina a. shalby fig. 3.12. the xyw phase space after control at eq1. fig. 3.13. the xyw phase space after control at eq2. fig. 3.14. the x trajectory behavior before control blue colored trajectory and after control red colored trajectory compared to eq1 and eq2 in green respectively. copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 11 fig. 3.15. the y trajectory behavior before control blue colored trajectory and after control red colored trajectory compared to eq1 and eq2 in green respectively. fig. 3.16. the z trajectory behavior before control blue colored trajectory and after control red colored trajectory compared to eq1 and eq2 in green respectively. fig. 3.17. the w trajectory behavior before control blue colored trajectory and after control red colored trajectory compared to eq1 and eq2 in green respectively. copyright c© 2017 assa. adv syst sci appl (2017) 12 lina a. shalby 4. conclusion in this paper, we have extended pfc method to stabilize chaotic continuous time systems by discretizing the system numerically using euler method, and then applying the pfc method. the pfc method is used by adding small control term to the iterative solution of the system, this term is the multiplication of a controlling matrix by the difference between two consecutive predicted iterates. we have also discussed the stability analysis of the continuous system compared to its numerical discrete form, and the choice of the control matrix due to the eigenvalues of the discretized system type either real or complex, and we have proved that the method is stable. the method has been applied to the most popular systems, lorenz 3d system and rössler hyperchaos 4d system. their trajectories have been controlled to each equilibrium point after a few iterations. the extended pfc method is easy implemented and highly efficient method for all chaotic systems to be stabilized, as we have generalized pfc method for all chaotic systems either discrete or continuous time system. acknowledgements it is my pleasure to have the supervision of prof. b. t. polyak. and i would like to express my gratitude to him, for the patient guidance, support and advice he has provided me in this work. also we would like to express our thankful to the editor and anonymous reviewers of the journal, for their great efforts and valuable revision. references 1. andrievskii, b. r. & fradkov, a. l. (2003) control of chaos: methods and applications. i. methods. automation and remote control, 64 (5), 673–713. 2. boukabou, a. & mansouri, n. (2008) predictive control of continuous chaotic systems. int. j. bifurcation chaos, 18 (2), 587–592. 3. celka, p. (1994) experimantal verification of pyragass chaos control method applied to chuas circuit. int. j. biforcation chaos appl. sci. eng., 4 (6), 29–36. 4. de sousa vieira, m. & lichtenberg, aj. (1996) controlling chaos using nonlinear feedback with delay. phys. rev. e, 54 (2), 1200–1207. 5. khalil, h. k.(2002) nonlinear systems. upper saddle river, nj: prentice hall. 6. kostelich, e. j., grebogi, c., ott, e. & yorke j. a.(1993) higher-dimensional targeting. phys. rev. e, 47 (1), 305–310. 7. lai, y.-c., ding, m. & grebogi, c. (1993) controlling hamiltonian chaos. phys. rev. e, 47 (1), 86–92. 8. leonov, g. a., zvyagintseva, k. a. & kuznetsova, o. a. (2016) pyragas stabilization of discrete systems via delayed feedback with periodic control gain. ifac-papersonline, 49 (14), 56–61. 9. lorenz, e.(1963) deterministic non-periodic flow. journal of the atmospheric sciences, 20 (2), 130–141. 10. mitsubori, k. & aihara, k. (2002) delayed–feedback control of chaotic roll motion of a flooded ship in waves. in proceedings of the royal society of london a: mathematical, physical and engineering sciences, 458 (2027), 2801-2813. 11. morgul, o. (2003) on the stability of delayed feedback controllers. phys. lett. a, 341(4), 278–285. 12. ott, e., grebogi, c. & yorke, j. a.(1990) controlling chaos. phys. rev. lett., 64(11), 1196–1199. 13. polyak, b. t. (2005) stabilizing chaos with predictive control. automation and remote control, 66(11), 1791–1804. copyright c© 2017 assa. adv syst sci appl (2017) pfcm for stabilization of continuous time systems 13 14. pyragas, k. (1992) continuous control of chaos by self-controlling feedback. phys. lett. a, 170 (6), 421-428. 15. pyragas, k. (1995) control of chaos via extended delay feedback. phys. lett. a, 206(5-6), 323-330. 16. pyragas, k. (2006) delayed feedback control of chaos. philosophical transactions of the royal society of london a: mathematical, physical and engineering sciences, 364(1846), 2309–2334. 17. schöll, e. & schuster, h. g.(2008) handbook of chaos control, john wiley & sons. 18. ushio, t. (1996) limitation of delayed feedback control in nonlinear discrete-time systems. ieee trans. circ. syst., 43(9), 815–816. 19. ushio, t., and yamamoto, s. (1999) prediction-based control of chaos. phys. lett. a, 264(1), 30–35. copyright c© 2017 assa. adv syst sci appl (2017) introduction predictive feedback control method examples lorenz model rössler model conclusion adv syst sci appl 2018; 04; 121-135 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/584 chaotic analysis and improved finite-time adaptive stabilization of a novel 4-d hyperchaotic system edwin a. umoh*1, ogechukwu n. iloanusi1 1) department of electronic engineering, university of nigeria, nsukka, nigeria e-mail: edwin.umoh.pg76768@unn.edu.ng, ogechukwu.iloanusi@unn.edu.ng received july 27, 2018; revised november 11, 2018; published december 31, 2018 abstract: chaotic systems are evidently very sensitive to slight perturbations in their algebraic structures and initial conditions, which can result in unpredictability of their future states. this characteristics has rendered them very useful in modelling and design of engineering and nonengineering systems. using the burke-shaw chaotic system as a reference template, a special case of a novel 4-d hyperchaotic system is proposed. the system consists of 10 terms and 9 bounded parameters. in this paper, after the realization of a mathematical model of the novel system, we designed an autonomous electronic circuit equivalent of the model and subsequently proposed an improved adaptive finite-time stabilizing controller which incorporates some augmented strength coefficients in the derived controller structures. these augmented coefficients greatly constrained transient overshoots and resulted in a faster convergence time for the controlled trajectories of the novel system. this novel system is suitable for application in the modelling and design of information security systems such as image encryption and multimedia security systems, due to its good bifurcation property. keywords: chaotic analysis, adaptive finite-time stabilization, hyperchaos, lyapunov stability. 1. introduction chaos is a unique phenomenon that often occur in a class of dynamic systems that are sensitive to perturbation in their initial conditions or mathematical models, consequently resulting in unpredictability of their future states [1]. chaos has been found to exists in natural and man-made systems, and its dynamics have been used in the modelling and studies of practical and hypothetical systems in communications engineering [2], medicines [3], robotics [4], thermodynamics [5], waste water treatment plant [6], electric power system [7], amongst others. however, for chaos to be of practical use, it must be controllable. consequently, during the past decades since the seminal works of ott, grebogi and yorke [8], extensive research on chaos control has resulted in the proposition of different methods in the literature and their applications by researchers in controlling and stabilizing chaotic dynamics such as active control [9], sliding mode control [10], fuzzy control [11] and adaptive hybrid control [12], amongst others. classical stability concepts such as the lyapunov stability, bibo stability and asymptotic stability concepts [13], [14] do not take into consideration the time intervals of stabilization of systems. rather, the main objective is the stabilization which can continue into infinite time. in recent years however, the concept of finite time stability and stabilization has gained wide attention due to its usefulness in time-critical system designs including secure communication systems, image and data encryption systems. a dynamic system is said to be finite-time stable (or synchronized), if given a bound on the initial conditions, its states does not exceed a certain threshold during a specified interval. finite-time stability is essential to ensure that the trajectories of a targeted system does not overshoot a specified bound in order to avoid undesirable consequences * corresponding author: edwin.umoh.pg76768@unn.edu.ng mailto:ogechukwu.iloanusi@unn.edu.ng 122 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) during applications. there is a distinction between finite time stability and lyapunov stability. a system which is finite-time stable may be lyapunov asymptotically stable, whereas a system that is lyapunov asymptotically stable may not be finite-time stable, if during transient it uncontrolled dynamics exceeds a prescribed threshold of time [15]. many works on finite time control have appeared in recent year [16]–[18]. in some of the reference works listed under section four, however, the settling times are comparably large, even when the controllers featured the use of sign function. in the proposed adaptive controller in this paper, we proposes a control structure that offers a faster rate of convergence of the controlled trajectories of a novel special case of a hyperchaotic system [19], which was derived from the burke – shaw atmospheric model [20]. certain characteristics distinguished the novel system from several others in the literature. these include its unique bifurcation properties where the plots of control parameters depict bifurcation diagrams with fewer periodic windows, thus making the system a good candidate for application in chaos-based cryptosystem modelling and design. the elegance of the algebraic structure makes it possible to further reduce the algebraic structure to consists of fewer parameters, yet would still exhibit hyperchaotic features with some degrees of novelty in its qualitative features. the novel system is also highly sensitive to slightest change in its initial conditions and system parameters. thus, several attractors may be evolved through key sensitivity. 1.1 concepts and preliminaries throughout this paper, the following existing definitions and lemmas are used in the analysis and synthesis of the improved finite-time controller design for stabilization. consider a class of nonlinear systems given by 0 0( , ), ( ) (0)x f x t x t x x   (1.1) where 0( , ), ( ) nx f x t x t  is the system state and : n nx f    is a nonlinear function. assume that the origin o is an equilibrium point, 0(0,0,0,0) of (1). if there exist a constant 0t  where t may be influenced by the initial condition such that lim ( ) 0, t t x t   lim ( ) 0, t t x t   if t t then the system (1) is finite-time stable. definition 1.1 [15], [21] the origin of (1.1) is a finite-time stable equilibrium, if the origin is lyapunov stable and there exist a function : nt  called the settling time function, such that for every 0 nx  , the solution 0( , )x t x of (1.1) is defined on 0 0 0[0, ( ), ( , ) [0, ( )]]nt x x t x t t x   and 0 0 ( ) lim ( , ) 0 t t x x t x   . lemma 1.1 [22], [23] if there exist a continuous positive definite function ( ) : n nv t  , such that ( )v t is radially bounded (i.e. ( ( ))v x t  as ( )x t  ), and satisfies the following differential inequality: 0 0( ) ( ( )) , 0, , ( ) 0,qv t v t t t t v t      (1.2) where 0  and 0 1q  are two positive numbers. it follows that for any 0t , ( )v t satisfies the inequality: 1 1 0 0 0( ) ( ) (1 )( ),q qv t v t q t t t t t       (1.3) chaotic analysis of a novel 4-d hyperchaotic system 123 copyright ©2018 assa. adv. in systems science and appl. (2018) and ( ) 0,v t t t   . we conclude that the origin of (1.1) is globally stable in finite time t and the settling time t is given by the relationship: 1 0 0 1 ( ) (1 ) qt t v t q    (1.4) assumption 1 [24], [25] assume 1 2, ,..., n na a a  and 0 1  are all real numbers. then the following inequality holds: 1 2 1 2( ... ) ... rr rr n na a a a a a       (1.5) lemma 1.2 [25] for the dynamic system (1.1), if there exists a continuous differentiable function :[0, ) ,v d   class k function (.)a and (.)b , a function :[0, )k   , such that ( ) 0k t  for almost all [0, ),t  a real number (0,1)  and an open neighbourhood m d of the origin, such that: ( ,0) 0, [0, ) ( ) ( , ) ( ), [0, ), , ( , ) ( )( ( , )) , [0, ), , v t t a x v t x b x t x m v t x k t v t x t x m              (1.6) holds, then the equilibrium point ( ) 0x t  of the system (1.1) is uniformly finite-time stable with settling time function 0 0( , )t t x satisfying 1 1 0 0 0 0 ( , ) ( , ) [ ],( ), 1 v t x t t x k k        where 0 ( ) ( ) t t k t k s ds  . if nm d  , then the equilibrium ( ) 0x t  of the system (1.1) is globally uniformly finite-time stable. 2. main results it is well known in practical system applications that system parameters are not always known in advance due to the uncertainties that inevitable arises during operations. as a result, practical controllers are designed with uncertainties in focus. in this section, the aim is to design a finite-time adaptive stabilizing control laws that will stabilizing the unstable dynamics of the novel system and update the unknown parameters in a uniform finite time. 2.1 mathematical model of the novel system consider a general structure of a novel 4d hyperchaotic system which is inspired by the burke-shaw atmospheric model template [19], 1 1 1 2 2 3 2 3 1 3 4 2 5 4 3 6 1 2 7 4 8 1 9 2 10 3 11 4 ( )x x x x x x x x x x x x x x x x x                              (2.1) where 1 11, , 0i i      are bounded parameter vectors and 1 2 4[ , ,..., ]x x x x are the state variables. system (2.1) may be considered as a generic structure of a family of the novel hyperchaotic systems which can produce special cases that still exhibit hyperchaos, yet possessing qualitative and quantitative properties which are conditioned by nullifying one or 124 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) more system parameters, in addition to appropriate selection of system parameter values. in this paper, we nullified two terms 2 11 0   , resulting in a new 10 terms and 9 parameter system given by 1 1 1 2 2 3 1 3 4 2 5 4 3 6 1 2 7 4 8 1 9 2 10 3 ( )x x x x x x x x x x x x x x x                          (2.2) when 1 37, 3.1,   4 53.5, 0.95,   6 71, 13,   8 9 100.01, 3, 0.01     , system (2.2) exhibit hyperchaotic behaviour depicted in fig. 1 (a) – (f) after 200s. fig. 2.1. 2-d phase portraits of the novel hyperchaotic system -15 -10 -5 0 5 10 15 -20 -15 -10 -5 0 5 10 x1 x 3 (a) -15 -10 -5 0 5 10 15 -20 -15 -10 -5 0 5 10 15 20 x1 x 4 (d) -25 -20 -15 -10 -5 0 5 10 15 20 25 -20 -15 -10 -5 0 5 10 x2 x 3 (b) -25 -20 -15 -10 -5 0 5 10 15 20 -20 -15 -10 -5 0 5 10 15 20 x2 x 4 (e) -15 -10 -5 0 5 10 15 -30 -20 -10 0 10 20 30 x1 x 2 (c) -15 -10 -5 0 5 10 -20 -15 -10 -5 0 5 10 15 20 x3 x 4 (f) chaotic analysis of a novel 4-d hyperchaotic system 125 copyright ©2018 assa. adv. in systems science and appl. (2018) 2.2 qualitative and quantitative properties of the hyperchaotic system 2.2.1 dissipativity we can examine whether the system is dissipative or otherwise by applying the liouville divergence theorem [26], [27]. let the vector notation of the 4-d hyperchaotic system (2.2) be denoted by 1 1 1 21 2 3 1 3 4 2 5 42 3 3 6 1 2 7 4 4 8 1 9 2 10 3 ( ) ( )( ) ( )( ) ( ) ( ) ( ) ( ) ( ) i i f x a x xf x f x a x x a x a xf xdx f x f x f x a x xdt f x f x a x a x a x                             (2.3) suppose  is a region in the phase space 4 with a smooth boundary and ( ) tt  where t is the flow of the system (2.3). if v is a hypervolume of the phase space at time 0t  , then by liouville’s divergence theorem [28], 1 2 3 4 ( ) . t dv t fdx dx dx dx dt    (2.4) where 4 1 ( ) . i i i df x f dx   is the divergence of the vector field f of the system (2.3). using the theorem, the rate of volume contraction is given by the lie derivative [33] 1 1,2,3...i i i dv i v dt     (2.5) where 1 1 2 2 3 3 4 4, , ,x x x x        are the state variables of the system (2.3). the divergence of the vector field f on 4 can be obtained by using the relationship 31 2 4 1 2 3 4 ( )( ) ( ) ( ) . df xdf x df x df x f dx dx dx dx      (2.6) by using (2.3), (2.4) and (2.6), the divergence is 1 4 1 4( )f           (2.7) for 1 47, 3.5   , 3.5f   . since the divergence is negative, it implies that the hyperchaotic system is dissipative. 2.2.2 equilibria and local stability hyperchaotic systems are essentially nonlinear models, and are therefore linearized in order to study their local stability for different parameters. the stability is determined by the sign of the real part of the eigenvalues of the jacobian matrix. the jacobian matrix is the matrix of the partial derivatives of the right-hand side with respect to state variables where all derivatives are evaluated at the equilibrium point ex x and is expressed by 126 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) 1 1 1 1 1 2 3 4 2 2 2 2 1 2 3 4 3 3 3 3 1 2 3 4 4 4 4 4 1 2 3 4 ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) ( ) f x f x f x f x x x x x f x f x f x f x x x x x j x f x f x f x f x x x x x f x f x f x f x x x x x                                                   (2.8) when 0x  , the equilibrium is at the origin, i.e. 1 2 3 4( 0, 0, 0, 0)e x x x x    . the equilibrium point of (2.3) at any point 4x r (other than the origin), that is, * * * * 1 2 3 4( 0, 0, 0, 0)e x x x x    is calculated by using the matrix (2.8). 1 1 * * 3 3 4 3 1 5 * * 6 2 6 1 8 9 10 0 0 ( ) 0 0 0 x x j x x x                              (2.9) analyzing (2.9), the equilibrium points are located at 10( , , , )e b b ab cb da   10( , , , )e b b ab cb da     (2.10) where 8 9 7 6 4 5 3 7 5 6 10 a b c d                 (2.11) to test for the type of stability associated with each equilibrium point, eq. (2.10) were computed as follows: 3.61 3.61 3.61 3.61 , 1078 1078 12670.28 12696.88 e e                             (2.12) using e and e respectively in the matrix (2.9) gives an indication of the nature of stability of these equilibrium points i.e. 7 7 0 0 3341.8 3.5 11.191 0.95 ( ) 3.61 3.61 0 0 0.01 3 0.01 0 j e                  (2.13) the eigenvalues of (2.13) are 1 2 3 40.000, 154.911, 150.886, 0.0251        , which implies that it is a saddle and unstable. next, we have 7 7 0 0 3341.8 3.5 11.191 0.95 ( ) 3.61 3.61 0 0 0.01 3 0.01 0 j e                (2.14) chaotic analysis of a novel 4-d hyperchaotic system 127 copyright ©2018 assa. adv. in systems science and appl. (2018) the eigenvalues of (2.14) are 1 2 3 40.000, 155.032, 151.000, 0.025        which implies it is a saddle and unstable. 2.2.3 lyapunov exponents and kaplan-yorke dimension the lyapunov exponent measures qualitatively the rate of exponential divergence or convergence of nearby trajectories of the system in state space, while the kaplan-yorke dimension (lyapunov dimension) gives a quantitative measure of this divergence. the lyapunov exponents was calculated based on the wolf algorithm [29], in conjunction with the procedure which implements the gram-schmidt ortho-normalization in the matlab environment. the numerically computed values are 1 2.216,le  2 1.280,le  3 0.000,le  and 4 6.997le   . the system has two positive lyapunov exponents, a null and a negative lyapunov exponent (+, +, 0, -), which confirmed its hyperchaoticity. moreover, 1 2 3 4le le le le   and 1 2 3 4 0le le le le    , hence the novel system is dissipative. the kaplan-yorke dimension [30] is given by 11 1 , 1,2,3 d ky j jd d d le j le     (2.15) where d is the topological dimension of the attractor, and must satisfy 1 0 d j j le   . for regular chaos with three dimensions, the topological dimension is 2, while for hyperchaos, it is 3. by using the numerically generated values of lyapunov exponents, the kaplan-yorke dimension is calculated as 1 2 3 4 3 3.4996ky le le le d le      2.2.4. bifurcation diagrams bifurcation diagram shows how the dynamics of a system changes with variation of control parameters. in this paper, we explored the parameter space to discover which parameters influence the dynamic characteristics of the system. two parameters 1 and 3 influences the route to chaos. fig. 2.2 depicts the bifurcation diagrams plotted by varying the two control parameters for the range 16.3 6.5  and 13.5 4.15  . fig. 2.2. bifurcation diagrams of control parameters 1 , 3 3. finite-time adaptive stabilization the objective of finite-time adaptive stabilization is to design an adaptive control law and parameter update laws such that the state and parameter update trajectories of the controlled system (2.16) converged in uniform finite time for any initial condition. let the controlled form of system (2.3) be given as follows: 128 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) 1 1 1 2 1 2 3 1 3 4 2 5 4 2 3 6 1 2 7 3 4 8 1 9 2 10 3 4 ˆ ( ) ˆ ˆ ˆ ˆ ˆ ˆ ˆ ˆ x x x u x x x x x u x x x u x x x x u                              (3.1) where 1 3 10 ˆ ˆ ˆ,   are the unknown parameters to be estimated, , ( 1,2..4) i ii x xu d a i   is the finite-time adaptive control functions associated with each coupled equation and comprises of ixd , the model equation-based derivative of the control functions and sgn( ) ,( 1,2..4) i ix i u i ia x x x i      , the augmented control functions and iu are the augmented controller strength coefficients. lemma 3.1 the unknown parameters of system (3.1) can be estimated if there exists parametric error functions ,( 1,3,...10)i i  , where : 1 1 1 3 3 3 4 4 4 5 5 5 6 6 6 7 7 7 8 8 8 9 9 9 10 10 10 ˆ ˆ ˆ, , , ˆ ˆ ˆ, , , ˆ ˆ ˆ, ,                                              (3.2) 1 3 10 ˆ ˆ ˆ,   are the unknown parameters to be estimated. taking the derivatives of (3.2) yields the following time-varying functions: 1 1 3 3 4 4 5 5 6 6 7 7 8 8 9 9 10 10 ˆ ˆ ˆ ˆ ˆ, , , , ˆ ˆ ˆ ˆ, , ,                                     (3.3) remark 3.1 several works have featured the use of sign function in finite-time controller design due to its good tracking properties [18], [31] etc. in these references, the designed controllers exhibit good noise rejection and good convergence time than controllers that do not feature the signum variables. however, their convergence time are still appreciably large. in this present work therefore, we introduced coefficient terms called augmented controller strength coefficients and augmented parameter update strength coefficients respectively, in conjunction with signum variables to produces controller and parameter update structures that drastically cut down on the uniform convergence time of the dynamics of the controlled system. 3.1 the proposed controller and parameter update structures the controlled hyperchaotic system (3.1) can be finite-time adaptive stabilized and the unknown parameters can be accurately estimated in uniform finite time, if the following adaptive stabilizing control laws and parameter update laws are applied: chaotic analysis of a novel 4-d hyperchaotic system 129 copyright ©2018 assa. adv. in systems science and appl. (2018) 1 2 3 4 1 3 4 1 1 1 2 1 1 1 2 3 1 3 4 2 5 4 2 2 2 3 6 1 2 7 3 3 3 4 8 1 9 2 10 3 4 4 4 1 1 1 2 1 3 1 2 3 3 2 4 2 4 5 ˆ ( ) sgn( ) ˆ ˆ ˆ sgn( ) ˆ ˆ sgn( ) ˆ ˆ ˆ sgn( ) ˆ ( ) ˆ ˆ ˆ u u u u u x x x x x u x x x x x x x u x x x x x x u x x x x x x x x x x x x x                                                          5 6 7 8 9 10 2 4 5 6 1 2 3 6 7 3 7 8 1 4 8 9 2 4 9 10 3 4 10 ˆ ˆ ˆ ˆ ˆ x x x x x x x x x x x x                                                              (3.4) where 1  are the augmented parameter update strength coefficients ( i  has the same definition with iu ) and 0 1  is a non-negative index. proof. firstly, by using (3.2) and (3.4) in (3.1), the controlled system (3.1) becomes 1 2 3 4 1 1 1 2 1 1 1 2 3 1 3 4 2 5 4 2 2 2 3 6 1 2 7 3 3 3 4 8 1 9 2 10 3 4 4 4 ( ) sgn( ) sgn( ) sgn( ) sgn( ) u u u u x x x x x x x x x x x x x x x x x x x x x x x x x x x x                                            (3.5) secondly, based on lemma 1.2, a lyapunov function candidate is given by 4 10 2 2 1 1 2 2 2 2 2 2 2 2 2 2 2 2 2 1 2 3 4 1 3 4 5 6 7 8 9 10 1 1 ( ) 2 2 1 1 ( ) ( ) 2 2 x p i i i i v t v v x x x x x                               (3.6) the partial derivative of (3.6) produces 1 1 2 2 3 3 4 4 1 1 3 3 4 4 5 5 6 6 7 7 8 8 9 9 10 10 ( ) ... ... v t x x x x x x x x                                 (3.7) by making use of (3.3), (3.4) and (3.5) in (3.7), the derivative becomes 1 2 3 4 1 1 1 2 1 1 1 2 3 1 3 4 2 5 4 2 2 2 3 6 1 2 7 3 3 3 4 8 1 9 2 10 3 4 4 4 1 1 3 3 4 4 5 5 ( ) ( ( ) sgn( ) ) ( ... ... sgn( ) ) ( sgn( ) ) ... ˆ ˆ... ( sgn( ) ) ... ˆ ˆ... u u u u v t x x x x x x x x x x x x x x x x x x x x x x x x x x x x                                                       6 6 7 7 8 8 9 9 10 10 ˆ ˆ ˆ ˆ ˆ              (3.8) 130 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) using the following convention 1 sgn( ) ; sgn( ) sgn( ) i i i i i i i i i i i i i x x x x x x x x x x x x x        (3.9) eq. (3.8) is then transformed to 1 2 3 4 1 1 1 1 1 2 3 4 1 2 3 4 2 1 1 1 2 3 1 2 3 4 2 5 2 4 6 1 2 3 7 3 8 1 4 9 2 4 10 3 4 1 1 3 3 4 4 5 5 6 6 7 7 8 8 9 9 ( ) ... ... ( ) ... ˆ ˆ ˆ ˆ ˆ ˆ ˆ ˆ... u u u uv t x x x x x x x x x x x x x x x x x x x x x x x x x x x x                                                                     10 10̂ (3.10) and by using (3.9) in (3.10), the derivative reduces to 1 2 3 4 1 3 4 5 6 7 8 9 10 1 1 1 1 1 2 3 4 1 2 3 4 1 3 4 5 6 7 8 9 10 ... ( ) 0 ... u u u ux x x x x x x x v t                                                                  (3.11) eq. (3.11) is negative definite in 17 , thus lemma 1.2 is satisfied, and the equilibrium point ( ) 0x t  is uniformly finite-time stable about the origin because ( ) 0v t  and also implies that (0) 0v  and lim 0 t t x   . suppose 1 2 3 4 cu u u u u        (where cu is known as the controller strength coefficient) and 1 3 4 10 ... p             (where pu is known as the parameter update coefficient), then by setting c pu u g    , ( g is known as the global strength coefficient), (3.11) reduces to: 4 4 9 11 1 1 1 ( ) 0i g i g i i i i v t x x                (3.12) using assumption 1, it can be deducted that 14 4 1 12 2 1 1 ( )g i i i i x x x             (3.13) substituting (3.3) in (3.12) yields 4 11 11 1 1 ( ) i g i i i v t x x            (3.14) by virtue of assumption 1, it is easy to see that 1 1 2 2 21 1 1 1 2 2 2 2 ( ) ( ) 2 ( ) 2 2 v t x x x v                     (3.15) setting 1 2 1 2 , 2 p        , reduces (3.15) to chaotic analysis of a novel 4-d hyperchaotic system 131 copyright ©2018 assa. adv. in systems science and appl. (2018) ( ) pv t v  (3.16) also, by substituting  and p into (1.4), we have   1 12 0 20 0 01 2 ( ) 1 2 ( ) 11 2 2 v t t t t v t                  (3.17) where 22 2 2 0( ) (0) (0) (0) (0)x p i iv t v v x     remark 3.2 eq (3.17) implies that the finite-time stabilization of the controlled system (3.1) depends on the initial condition  and rational number p , while the uniform convergence time is either increasing or decreasing as the parameter p is varied as can be observed in the works of [18], [32] . in our case however, due to the complexity introduced by the global strength coefficient g , it was observed that the bounded value of g , min max [ , ]g g g   has a strong constraining effect on trajectory overshoot and also has domino effects on the uniform convergence time. accordingly, the following cases were observed when min max    and min maxg g g    . a. as ming g  and min  , the convergence rate was relatively slower and the uniform convergence time is an increasing function b. as maxg g  and max  , the convergence rate was relatively faster and the uniform convergence time is a decreasing function. 4. numerical simulation results the controlled hyperchaotic system is simulated using matlab for the following parameters 1 37, 3.1   4 5 6 73.5, 0.95, 1, 13       , 8 9 100.01, 3, 0.01     , with initial conditions set as 1 2 3 4[ (0), (0), (0), (0)] [4,2, 4, 2]x x x x    . the following results were obtained for two different cases. case 1: minmin 0.8, 1000g   0 0.001 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.01 -4 -2 0 2 4 t(s) s ta b il iz e d t r a je c to r ie s x1 x2 x3 x4 fig.4.1. stabilized trajectories of the controlled system 0 0.001 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.01 -3000 -2000 -1000 0 1000 2000 3000 4000 t(s) c o n tr o ll e r d y n a m ic s u1 u2 u3 u4 fig.4.2. uniformly converged dynamics of the adaptive controller 132 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) case 2: maxmax 0.99, 8000g   remark 4.1 the rate of asymptotic convergence of the system dynamics in case 1 is slower than those in case 2 due to the effects of the global strength coefficients. this is in consonance with the remark 3.2 on the effects of the global strength coefficient. it can be observed that the larger g , the faster the rate of uniform convergence of the system’s trajectories. when compared to controllers that featured in related works listed in table i, it is obvious that the designed adaptive controller provides a faster convergence time than those in these works. 0 0.001 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.01 -5 0 5 10 15 t(s) e s ti m a te d p a r a m e te r s a7 a8 a9 a10 fig.4.3 (b). estimated parameters of the controlled system 0 0.001 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.01 -10 -5 0 5 10 t(s) e s ti m a te d p a r a m e te r s a1 a3 a4 a5 a6 fig.4.3 (a). estimated parameters of the controlled system 0 0.5 1 1.5 x 10 -3 -4 -2 0 2 4 t(s) s ta b il iz e d t r a je c to r ie s x1 x2 x3 x4 fig.4.4. stabilized trajectories of the controlled system 0 0.5 1 1.5 x 10 -3 -4 -2 0 2 4 x 10 4 t(s) c o n tr o ll e r d y n a m ic s u1 u2 u3 u4 fig.4.5. uniformly converged dynamics of the adaptive controller 0 0.5 1 1.5 x 10 -3 -10 -5 0 5 10 t(s) e s ti m a te d p a r a m e te r s a1 a3 a4 a5 a6 fig.4.6 (a). estimated parameters of the controlled system 0 0.5 1 1.5 x 10 -3 -5 0 5 10 15 t(s) e s ti m a te d p a r a m e te r s a7 a8 a9 a10 fig.4.6 (b). estimated parameters of the controlled system chaotic analysis of a novel 4-d hyperchaotic system 133 copyright ©2018 assa. adv. in systems science and appl. (2018) 4.1. comparison with related works the rate of uniform convergence of system dynamics using the proposed finite-time adaptive controller was compared with related works in the literature. the different rates reported in the papers are given in table i. table i. comparison with related works related works approximated convergence time of error dynamic systems proposed work 0.005s ref [33] 0.2s ref [34] 0.3s ref [35] 0.5s ref [36] 1.0s ref [37] >1.0s ref [38] >1.0s ref [31] 2.0s ref [39] >1.0s ref [40] >2.0s ref [41] 8.0s 5. conclusion in this paper, we proposed an improved finite-time adaptive controller with a fast rate of uniform convergence of system trajectories of hyperchaotic systems. we examined the performance of the controller on a novel special case of a hyperchaotic system which was originally evolved by adding feedback control to the classical burke-shaw chaotic system that models some atmospheric phenomena. numerically obtained results shows that the proposed control structure offered better rate of convergence of the controlled system dynamics when compared to related works in the literature. overall, these systems can contribute to improve the design characteristics of chaos-based cryptosystems. references [1] e. n. lorenz, “deterministic nonperiodic flow,” j. atmos. sci., vol. 20, no. 2, pp. 130–141, 1963. [2] l. m. pecora and t. l. carroll, “synchronization in chaotic systems,” phys. rev. lett., vol. 64, no. 8, pp. 821–824, 1990. [3] a. kumar and b. m. hegde, “chaos theory: impact on and applications in medicine,” nitte univ. j. heal. sci., vol. 2, no. 4, pp. 93–99, 2012. [4] r. paper, x. zang, s. iqbal, y. zhu, x. liu, and j. zhao, “applications of chaotic dynamics in robotics,” vol. 1, 2016. [5] d. vitali and p. grigolini, “chaos, thermodynamics and quantum mechanics: an application to celestial dynamics,” phys. lett. a, vol. 249, no. 4, pp. 248–258, 1998. [6] j. qiao, z. hu, and w. li, “soft measurement modeling based on chaos theory for biochemical oxygen demand (bod),” water, vol. 8, no. 12, p. 581, 2016. [7] h. r. abbasi, a. gholami, m. rostami, and a. abbasi, “investigation and control of unstable chaotic behavior using of chaos theory in electrical power systems,” iran. j. electr. electron. eng., vol. 7, no. 1, pp. 42–51, 2011. 134 e. a. umoh, o. n. iloanusi copyright ©2018 assa. adv. in systems science and appl. (2018) [8] e. ott, c. grebogi, and j. a. yorke, “controlling chaos,” phys. rev. lett., vol. 64, pp. 1196–1199, 1990. [9] e. a. umoh, “generalized synchronization of topologically-nonequivalent chaotic signals via active control,” int. j. signal process. syst., vol. 2, no. 2, pp. 139–143, 2014. [10] m. roopaei, b. r. sahraei, and t.-c. lin, “adaptive sliding mode control in a novel class of chaotic systems,” commun. nonlinear sci. numer. simul., vol. 15, no. 12, pp. 4158–4170, 2010. [11] d. chen, w. zhao, j. c. sprott, and x. ma, “application of takagi-sugeno fuzzy model to a class of chaotic synchronization and anti-synchronization,” nonlinear dyn., vol. 73, no. 3, pp. 1495–1505, 2013. [12] e. a. umoh, “adaptive hybrid synchronization of lorenz-84 system with uncertain parameters,” telkomnika indones. j. electr. eng., vol. 12, no. 7, 2014. [13] s. schauland, j. velten, and a. kummert, “insufficiencies of the practical bibo stability concept with regard to signal processing systems,” pp. 0–4, 2009. [14] t. binazadeh and m. h. shafiei, “a novel approach in the finite-time controller design,” syst. sci. control eng., vol. 2, no. 1, pp. 119–124, 2014. [15] s. p. bhat and d. s. bernstein, “finite-time stability of continuous autonomous systems,” siam j. control optim., vol. 38, no. 3, pp. 751–766, 2000. [16] f. amato, m. ariola, and p. dorato, “robust finite-time stabilization of linear systems depending on parametric uncertainties,” in 37th ieee conference on decision and control, 1998, pp. 1207–1208. [17] j. wang, x. chen, and j. fu, “adaptive finite-time control of chaos in permanent magnet synchronous motor with uncertain parameters,” nonlinear dyn., vol. 78, no. 2, pp. 1321–1328, 2014. [18] j. wang, t. gao, and g. zhang, “adaptive finite-time control for hyperchaotic lorenz – stenflo systems,” phys. scr., vol. 90, no. 2, p. 25204, 2015. [19] e. a. umoh and o. n. iloanusi, “algebraic structure, dynamics and electronic circuit realization of a novel reducible hyperchaotic system,” in 2017 ieee 3rd international conference on electro-technology for national development (nigercon 2017), owerri, nigeria, 2017, pp. 483–490. [20] r. shaw, “strange attractor, chaotic behaviour and information flow,” zeitschrift fur naturforsch a, vol. 36, pp. 80–112, 1981. [21] m. defoort, k. veluvolu, m. djemai, a. polyakov, and g. demesure, “leaderfollower fixed-time consensus for multi-agent systems with unknown non-linear inherent dynamics,” iet control theory appl., vol. 9, no. 14, pp. 2165–2170, 2015. [22] i. ahmad, a. b. saaban, a. . ibrahim, and m. shahzad, “robust finite-time antisynchronization of chaotic systems with different dimensions,” mathematics, vol. 3, pp. 1222–1240, 2015. [23] g. hardy, j. littlewood, and g. polya, inequalities. cambridge, uk: cambridge university press, 1952. [24] m. p. aghababa, s. khanmohammadi, and g. alizadeh, “finite-time synchronization of two different chaotic systems with unknown parameters via sliding mode technique,” appl. math. model., vol. 35, no. 6, pp. 3080–3091, 2011. chaotic analysis of a novel 4-d hyperchaotic system 135 copyright ©2018 assa. adv. in systems science and appl. (2018) [25] w. m. haddad, s. g. nersesov, and l. du, “finite-time stability for time-varying nonlinear dynamical systems,” proc. am. control conf., no. 3, pp. 4135–4139, 2008. [26] f. hoppensteadst, analysis and simulation of chaotic systems, 2nd ed. new york: springer, 2000. [27] s. vaidyanathan, “a novel 4-d hyperchaotic thermal convection system and its adaptive control,” in advances in chaos theory and intelligent control, studies in fuzziness and soft computing 337, a. t. azar and s. vaidyanathan, eds. switzerland: springer international publishing, 2016, pp. 75–100. [28] m. godina and p. matteucci, “reductive g structures and lie derivatives,” j. geom. phys., vol. 47, no. 1, pp. 66–86, 2003. [29] a. wolf, j. b. swift, h. l. swinney, and j. a. vastano, “determining lyapunov exponents from a time series,” phys. d nonlinear phenom., vol. 16, no. 3, pp. 285– 317, 1985. [30] p. grassberger and i. procaccia, “characterization of strange attractors,” phys. rev. lett., vol. 50, no. 5, pp. 346–349, 1983. [31] r. li, w. chen, and s. li, “finite-time stabilization for hyper-chaotic lorenz system families via adaptive control,” appl. math. model., vol. 37, no. 4, pp. 1966–1972, 2013. [32] t. ren, z. zhu, and h. yu, “design of finite-time synchronization controller and its application to security communication system,” appl. math. inf. sci., vol. 8, no. 1, pp. 387–391, 2014. [33] c.-z. chen, p. he, t. fan, and c. jing, “finite-time chaotic control of unified hyperchaotic systems with multiple parameters,” int. j. autom. control, vol. 8, no. 8, pp. 57–66, 2015. [34] w. xiong and j. huang, “finite-time control and synchronization for memristor-based chaotic system via impulsive adaptive strategy,” adv. differ. equations, vol. 2016, pp. 101–109, 2016. [35] h. lin, j. cai, and j. wang, “finite-time combination-combination synchronization for hyperchaotic systems,” j. chaos, no. id304643, pp. 107, 2013. [36] x.-t. tran and h.-j. kang, “synchronization and stabilization for hyperchaotic systems via a new modified finite-time control,” in icmce’16, 2016, pp. 113–117. [37] z. ma, y. sun, and h. shi, “finite-time stabilization of dynamical system with adaptive feedback control,” j. appl. math. phys., vol. 5, pp. 412–421, 2017. [38] a. abooee, “a robust finite-time hyperchaotic secure communication scheme based on terminal sliding mode control,” in 2016 24th iranian conference on electrical engineering (icee), 2016, pp. 854–858. [39] h. wang, z. han, and q. xie, “finite-time chaos control of unified chaotic systems with uncertain parameters,” nonlinear dyn., vol. 55, pp. 323–328, 2009. [40] f. gao and f. yuan, “adaptive finite-time stabilization for a class of uncertain high order nonholonomic systems,” isa trans., vol. 54, pp. 75–82, 2015. [41] l. liu, x. cao, z. fu, and s. song, “guaranteed cost finite-time control of fractional-order positive switched systems,” j. control sci. eng., vol. 2017, p. 10pp, 2017. adv syst sci appl 2019; 01; 44-60 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/647 protecting a low voltage direct current system using solid-state switching devices for dc grid applications ibrahim almutairy1*, johnson asumadu2, zhang xingzhe3 1) majmaah university, western michigan university, michigan, usa e-mail: ibrahim.almutairy@wmich.edu 2,3) electrical and computer engineering, western michigan university, michigan, usa e-mail: xingzhe.zhang@wmich.edu received september 20, 2018; revised february 24, 2019; published april 15, 2019 abstract: low voltage dc (lvdc) distribution system has less distributional losses than ac grid, integrating renewable sources, greatly increasing the mix of clean energy sources in the standard grid in the next decade and this has become a trend that is likely to continue. as an additional benefit, dc electrical power is oftentimes seen as beneficial to applications that use cleanly generated, renewable power. considering lvdc power system design, and focusing on fault protection and the systems dc circuit breaker, are system priorities. only solid-state circuit breakers (sscb) should be considered to obtain advanced system topography. a new type of solid state circuit breaker is being developed as a new power device for lvdc power networks to replace the emcb. the only suitable candidate for this task is the insulated gate bipolar transistor (igbt). development of dc grid protection allows for the development and improvement of new power electronic devices, focusing in on the application for medium to low voltage dc grids, a rapidly acting switching function as well as fault current limiting features. to avoid nuisance tripping, fault current limiting function can be satisfactorily accomplished by extending the elapsed time using the same control circuit. this paper presents a novel lvdc circuit breaker model, which uses a coupled inductors circuit breaker, a designed model of igbt showed and discussed and developed an lvdc system for grid applications. keywords: dc protection, insulated-gate bipolar transistors (igbts), lvdc power system, solid-state circuit breakers (sscb). 1. introduction lvdc electrical energy using renewable environmentally friendly power is frequently seen in many applications as helpfully modern. considering the design of dc power systems, it is important to focus on fault protection and maintain the power system's dccb. only solidstate circuit breakers should be considered for installation and applied to these types of systems because these are the only types breakers that respond quick enough to faults to prevent damages even though they tend to have a greater rate of power loss than other breaker types. sscbs are pure semiconductor devices that monitor all the behaviors of the igbt which are tremendously important to switching transitions. solid state lvdccbs make use of a completely manageable power semiconductor device from the primary branch to interrupt electrical current and doesn’t require a mechanical switch. they interrupt dc current very effectively the idea being introduced in this system is a new type of dc solid state circuit breaker with transformer coupling that uses coupled inductors to detect and isolate faults. these types of circuit breakers gain advantages over other types of circuit breakers because they * corresponding author: ibrahim.almutairy@wmich.edu mailto:xingzhe.zhang@wmich.edu protecting a low voltage direct current system 45 copyright ©2019 assa. adv. in systems science and appl. (2019) don’t have as many parts and has an adjustable current threshold that can be set to operate when the level of current rises past the set threshold causing the circuit breaker to trip [1-6]. normally, the circuit breaker is in the system between the source and the load. sscbs are used in high voltage systems because they have a swift switching speed and interrupt high voltage dc current very efficiently. during fault interruption, the rapidly dropping current will expose the sscb to serious strain due to overvoltage. if this is left uncorrected, this overvoltage can easily cause the sscb to become unresponsive and might possibly damage the sensitive devices inside the sscb [7], [8]. to address this issue, it is necessary to install a snubber circuit to protect the system by suppressing overvoltage. having an installed snubber circuit becomes essential to the safe turn off of the system. there are some requirements that should be able to be met by the snubber circuit. it should be highly reliable for medium capacity, low voltage applications. overvoltage should cause a minimal level of stress to the sscb at turn off. the presence of a snubber circuit that has been designed and specialized for precise application in sscbs is of extreme importance. generally, methods of snubber design anticipate that snubbers be must be suited for converter switches and thus, they cannot be applied directly to the sscb snubber. there are many reasons for this. reduction of snubber loss as well as rapid suppression features do not get lots of attention in sscb snubber design because of inconsecutiveness as well as infrequency of sscb switching. suppression of overvoltage as well as ability to withstand fault current are the features which receive the most attention in sscb snubber design due to the magnitude of fault current that is met and, as a consequence, sscbs become exposed to high overvoltage [7-10]. this paper is organized as follows: section two covers related work, section three present the background of the lvdc system and protection using sscb for lvdc grid applications. in section four, analytical analysis of proposed dc protection circuit is presented. in section five, examination of igbts for implementation in lvdc sscb is covered. section six presents modeling the lvdc system under study and section seven conclusion. 2. related work certain advancements in the field of power electronics have been ongoing and lvdc systems are looking more advantageous than they ever and have also become increasingly dependable since the application of grid distribution, advanced circuit breaker technology, power converters and other lvdc protective devices. new microgrids will be perceived in the dc power format and many power systems or, power conversion components are readily available but, as it applies to dc circuit breakers, many different designs are still in the experimental stages. circuit breakers have certain functions and duties which must be performed. a circuit breaker must be capable of performing three basic duties. it must be capable of opening the faulted circuit and breaking the current. it must be able to be closed to a fault. lastly it must be able to carry fault current for a short period of time as part of a hierarchal protection coordination scheme. in addition, it must be completely controllable. a circuit breaker must have a rapid switching speed as well as low conduction losses, no arcing, and must ensure definitive tripping of the whole breaker system. detecting short circuits as quickly as possible is the key to minimize power dissipation in system components during fault events so that overheating and damages to a switching device, such as an igbt, can be prevented. various works address the need for a new and fast circuit breaker model for lvdc system applications. one recent research proposed a bidirectional magnetically coupled t-source inverter (tsi) for lvdc systems [11]. the paper investigated the performance of a magnetically coupled t-source inverter for lvdc systems. considering that building an ideal transformer is not possible, the coupling coefficient of the transformer, in respect to the leakage inductance, 46 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) determines its performance, settling upon practical lvdc application of the tsi and its limitations were exposed in the experimental results. analytical circuit analysis and derivation of mathematical equations were performed under ideal conditions. another study highlights protection of lvdc systems using a bidirectional z-source circuit breaker [12]. modified topologies of a z-source lvdc with bidirectional power flow capability are introduced in the study. conventional z-source breaker topologies make it possible to protect unidirectional power flow. proposed topologies in the study are a bidirectional zsb and bi directional zsb with a coupled inductor and can clear faults in any direction that is needed. both topologies make use of a scr as a switch for commutation. coupled inductors help reduce the size of inductors, the cost and also to eliminate any need for an additional z-source breaker capacitor. other recent research focuses on the design of solid-state circuit breaker-based protection for dc shipboard power systems [13]. protection is an important aspect of dc shipboard distribution system design. it should be thought about from the very beginning design stage for reliable, low cost and efficient system operation quick action is an essential requirement for dc shipboard distribution protection. sscb based dc protection offers very rapid protection speed and reduces the requirements on the fault withstanding limitations and fire hazards but has a high cost and losses of power because of semiconductor switching devices. dc protection methods are evaluated based upon speed and other criteria. any optimized sscb-based dc protection design is a compromise of the right levels of reliability, fast speed, reasonably low cost and good overall system efficiency. further work addresses dc microgrid protection using the coupled inductor sscb [14]. large scale implementation of dc microgrids depends upon the development of reliable dc protection options. the proposed dc breaker design uses an scr to interrupt the fault current instantly. a bidirectional version of the breaker is also provided so multiple dc breakers can be installed in a notional dc microgrid. the breaker has a control switch in the design so the breaker can be opened manually without creating a large disturbance. traditionally, a central control was proposed to supervise operations of multiple breakers in a microgrid. two other control strategies are studied. one involves independent, local control of breakers and the other is to control breakers in pairs. the merits of all three strategies are studied. it can be concluded that within certain limitations, the paired method of control provides the desired results either a much simpler communication architecture. the next research introduced a new dc sscb concept using a small inductor like that used in electro-mechanical circuit breakers [15]. the new sscb scheme uses the voltage inductor to activate short circuit protection. it is in this way that the circuit starts to open when the fault condition to prevent overload current is present [16]. some sscb technologies for dc are reported and can be classified into three main types, mechanicals, solid state and hybrids. each of them have their advantages and disadvantages in respect to the others however, there are some common characteristics that are extremely important in this device, such as, time of response, electric arc suppression, high efficiency, and low cost among others. the development of new semiconductor devices with better features is entirely feasible and may improve sscb technologies in many aspects. the new sscb is based on two power mosfets, having one latch circuit, one inductor and one high-pass filter within a voltage divider. 3. the idea of low voltage dc electrical distribution low voltage dc grids are not very different from their ac counterparts. even though some components are different in nature, characteristics, or function, the basic concept is always the same. the objective is to transport electricity through cables and components at low voltage to the end-user. the main difference between dc and ac systems are the way in which they are configured. the microgrid itself, is a distributed power system on a small protecting a low voltage direct current system 47 copyright ©2019 assa. adv. in systems science and appl. (2019) scale, that is made of distributed energy sources and associated loads. microgrids eliminate the need for connection to the main utility grid [12-15]. there are numerous advantages of dc microgrids over ac microgrids such as the absence of a skin effect, lower on-state losses as well as the ability to provide 1.41 times more power than ac, no problems with reactive power control and no proximity effect. dc also has unique issues of its own like arc interruption, fewer protection devices and is more difficult to control than an ac system. lvdc microgrids are a new concept in distribution systems. it is well suited for installation in offices that have sensitive computer loads, isolated or rural power systems. the grid that we have today in the united states of america is mostly ac and most of loads we consume using various household electronic devices which consume the energy in dc, requiring power converters to connect to the grid. if these loads are fed through the dc bus, then they will require fewer power conversion stages. since there are fewer power conversion stages, losses resulting from conversion are also lower. most of the resistive loads can be connected to either the ac or the dc bus. fig. 2.1. conceptualization of the lvdc bus microgrid system. the above diagram shows the concept of a lvdc bus microgrid system. the active power consumed by dc loads is the same as ac loads but, there is no reactive power function in dc. a dc microgrid has a wind turbine, dc/dc converter for its battery and the solar pv. generally, wind turbine and micro turbine uses both dc/ac and ac/dc converters but, only an ac/dc converter is used in the methodology of the latest technology. ac power grid systems possess various advantages that are easily applicable but, certain issues like synchronization, control of reactive power and stability of the ac bus are persistent problems in the ac bus microgrid. as a lvdc bus microgrid system is small, it can neglect the loss of transmissions within the system and provides a solution to these problems. the dc distribution protection system is of considerable benefit to a lvdc bus microgrid as the fault is detected and located within the system autonomously by the protection system of the lvdc bus microgrid. besides the benefits, dc bus microgrids also face challenges like lack of standard practices and guidelines, as well as breaking of dc arcs and dc protective devices. a lvdc microgrid is dynamic because operation of supply and demand equipment so, suitable power-flow control, voltage regulation, protection, coordination and adequate filtering are necessary to guarantee an acceptable quality of supply. over the last few decades, power electronics have been one of the main contributing factors behind the paradigm shift that started in the 1960s and have become increasingly power-dense while decreasing in physical volume. such a trend was made possible, mainly, by increasing switching frequency. these days, power electronics that are very small, can be found in all types of electrical appliances that have an efficiency ranging from 50% up to 99%. the dialog regarding voltage levels seems unending because voltage levels will always be the greatest difference between the two types of systems. proposals about voltage level were specifically limited to voltages below 120vdc [16] [17]. some recent pilot projects put 48 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) voltages from 24vdc to 48vdc demonstrating that this may be something worthy of investigation in future work. an approach such as this one is consistent with the definition of iec 60038 for extra low voltage to assure the protection of the user, it should be possible to have a lvdc system that runs that level of voltage even if the system had plug-in electric vehicles as well as the household’s kitchen loads. 3.1. lvdc short circuit currents characteristics a lvdc microgrid that has a new, complex topology of mixed ac and dc currents become dangerous as the system becomes more complex and new types of faults are introduced and different responses from the system are to be expected when faults occur on a lvdc system. the energy stored in the smoothing capacitors and storage devices begin to feed a large discharge current in a very short amount of transient time. this causes dc fault currents to undergo a transient discharge current that has high frequency oscillations between the smoothing capacitors and system inductances as well as short circuit current in the steady state. these transient fault currents cause a very high level of stress to the operation of protection and performance of lvdc systems [18] [19]. characterizing lvdc short-circuits is therefore, of extreme importance. ratings of appropriate equipment and correct protection settings, as well as selectivity, need accurate short-circuit characterization. the effectiveness of using standards that are already in place such as iec61660 for characterizing lvdc faults will be investigated. when an external dc fault in a lvdc microgrid begins, the charged smoothing capacitor acts instantly as a considerable source of dc fault current and begins to feed the fault. when the discharge peak current is reached, the capacitor is totally discharged and the dc terminal voltage become very small. in some cases, it is almost zero. when the antiparallel diodes of the converter become forward biased, the vsc loses control and the igbt switches become blocked for the purpose of self-protection. this results in the anti-parallel diodes functioning as a bridge rectifier and continuing to feed the fault during the transient. after fault occurrence, transient fault current ends, steady-state dc fault current is fed by the grid and local generation. 3.2. dc fault types in general, the most common causes of short-circuiting are component failure, lighting surges, high magnitude ground fault or line to line fault, external environmental stressors such as fires or insulation failure caused by operating at excessive temperatures that result from overload stress. in lvdc microgrids, converters like vscs igbt-based converters are sturdier to ac faults and more sensitive to dc faults. two types of faults may occur on the dc side. the first is an internal dc fault inside of the main converter and the other is an external dc fault, either on the converter terminals or at isolated places on the downstream dc feeders. internal faults aren’t as frequently occurring when compared to external faults [20], [21]. the most extreme external dc fault happens at the converter terminals. remote faults on dc feeders also have a considerable impact on converter performance. dc faults are classified as follows: line-to-ground fault: this type of fault occurs when a positive or negative pole is shorted to the ground. in this case, the voltage of the faulted dc line drops down depending on the fault impedance. line-to-line fault: this type of fault happens if the positive and negative dc poles are shorted. this could be the result of the dc cable insulation breaking down or a direct short circuit between the lines. table 3.1 shows the differences between dc fault types. protecting a low voltage direct current system 49 copyright ©2019 assa. adv. in systems science and appl. (2019) table 3.1 differences between line-to-line and line-to-ground fault line-to-line fault line-to-ground fault the positive and negative lines are normally shortcircuited whenever line-to-line faults happen. line will be shortcircuited to the ground when line-to ground faults happen. low fault impedance is the typical characteristic of line-to-line faults. both low and high impedance are the typical characteristics of line-toground faults. although overvoltage protection is vital during line to ground faults, when faults happen, the capacitor swiftly discharges to the ground [22]. current flows through the ground and back to the non-faulted line then finally goes back to the source, causing voltage on the healthy pole to rise. overvoltage also is concerning during the loss of an inverter. the loss of an inverter causes voltage to rapidly spike because of extreme unnecessarily high power, rapidly charging the dc link capacitors. a rectifier loss, only becomes concerning when the loss is temporary. when the rectifier abruptly returns, it causes overvoltages like an inverter loss does. 3.3. problems associated with lvdc protection system typically, power distribution systems have not been considered important when they are viewed from the perspective of protection complexity because of the high cost as well as the radial characteristic and nature of distribution networks, that have fault current directions that are always known. the principles of operating most protection systems are based on non-unit overcurrent protection schemes in which operating time decreases as the fault current increases. protecting lvdc distribution networks does not involve installing hardware of any kind. additionally, interrupting dc fault current is more difficult than interrupting ac fault current. dc fault current and voltage waveforms do not have a natural zero crossing ability, so they must be forced to zero meaning that the risk of fires or arcing are expected more often with dc systems than ac because of problems that develop from forcing dc currents to zero. because of this, standard circuit breakers and fuses have difficulty in protecting dc systems. if emcb are going to be utilized in a dc system, a larger size and heavier weight as well as slow performance can be expected. there is a solution. interruption of dc fault current is possible if current-limiting fuses or current limiting circuit breakers, such as sscbs, are used and their ratings are adjustable for application in lvdc systems [23]-[25]. these types of devices don’t require zero-crossing to extinguish dc fault arcs. fault current will be interrupted much sooner than the time needed for fault current to reach its peak. dc fault arcs are more difficult than with ac because they must reduce the voltage across the arc by increasing its length which is accomplished by increasing the length between the two contactors and using an arc splitter to split the arc. in addition, dc short circuit profile includes two main forms. one is a high transient discharge short circuit current path and the second is a steady state short circuit current path. transient discharge fault current can be high when compared with steady state fault current as well as causes high stress levels to network components and negatively affects protection system performance [26], [27]. two big problems are expected to happen because of high discharging dc fault current. the first is an increase in the risk of physical damages to the sensitive equipment because of high magnitude and the second is the negative impact upon selectivity of non-unit protection. during transient current discharge, dc voltages near the fault become very small, almost zero, and sometimes the dc voltages can even become negative because of oscillation between cable inductances and filter capacitors. 50 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) 4. analytical analysis of proposed lvdc protection circuit observe the suggested lvdc circuit breaker design shown in fig. 4.1 during normal steadystate operation, current flows from the source to the load passing through the igbt and coupled inductors. an inductor is a passive two-terminal electrical component that holds electrical power in a magnetic field when electrical current is flowing through it. in this circuit, the inductor is considered as an inductive value in the distribution line. fig. 4.1. proposed lvdc circuit breaker mode. insulated-gate bipolar transistor (igbt) is a three-terminal power semiconductor device primarily used as an electronic switch. when compared with a mechanical switch, igbt needs several microseconds to switch off the circuit. consider the dc model of igbt [28], shown in fig. 4.2. fig. 4.2. lvdc model of igbt. in the model above, the voltage transfer function is expressed as: (4.1) (4.2) in the above equations, the parameters are: vcet : collector emitter voltage k : boltz mann constant m : modulation factor to : temperature of the junction of diode protecting a low voltage direct current system 51 copyright ©2019 assa. adv. in systems science and appl. (2019) vg : gate emitter voltage vd : dc bus voltage (collector emitter voltage) a coupled inductor transfers electrical energy between two circuits during current or voltage alteration. during the igbt shutdown, the voltage rapidly spikes because the circuit’s inductance (v=ldi/dt). the transformer transfers voltage to the snubber circuit which absorbs and dissipates energy and is used to suppress the voltage spikes that are caused by the circuit's inductance [29]. the capacitor absorbs the voltage spikes and the resistor releases any energy stored in the capacitor. the equivalent circuit of the proposed lvdc breaker is shown in fig. 4.3. fig. 4.3. proposed lvdc circuit breaker mode. vd is the voltage for igbt, it is acknowledged as a negligible value; so it can be ignored, in order for the transfer function of the equivalent circuit of the suggested lvdc circuit breaker is expressed as: (4.3) (4.4) kcl at node v1 gives: (4.5) (4.6) vo=v1-vg (4.7) (4.8) in this model, the voltage to current relation is expressed as: (4.9) the suggested lvdc circuit model (without snubber) shows the circuit with only the igbt, without a snubber circuit. from the simulation with no snubber, it shows that there is obviously a voltage spike happening at the load during igbt. parameter calculations of custom designed snubber circuit resistor and capacitor value. current across igbt: = (4.10) 52 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) (4.11) then choose the value for (4.12) = (4.13) by calculation, when c = 1mf is used. the power dissipates in the resistor (4.14) the lvdc circuit (with snubber) is the circuit model with both igbt and snubber circuit installed. from the simulation with a snubber, fig. 4.4. illustrates how the voltage spikes are reduced to zero, because the spike is being absorbed and dissipated by the snubber. the proposed lvdc circuit breaker has a rapid switching time of 30µs. as displayed in fig. 4.5. fig. 4.4. proposed lvdc cb model overvoltage with snubber. fig. 4.5. proposed lvdc cb model overvoltage with snubber. 5. examination of igbts for application in lvdc sscb connected systems the igbt is the most frequently used switching device for the sscb. meaning that the reliability and durability of igbts are ceaselessly being confirmed by several converters that protecting a low voltage direct current system 53 copyright ©2019 assa. adv. in systems science and appl. (2019) are installed in real-world power systems. most sscbs are based upon implementation of three semiconductor devices in some type of way. semiconductor devices are used in sscbs because they have extensive commercial accessibility, low power requirements for operation of gate drives as well as, high current ratings [30-33]. igbts have many benefits in contrast to its meagre number of drawbacks. the benefits are low switching losses, minor snubber circuitry requirements and high input impedance. because of these things, igbts are reliable and durable enough to be used for the circuit breaker role. there are many topologies for the dc circuit breaker using igbts. based on fig. 5.1, the igbt model under study comprises of an nchannel mosfet and two bjts, bjts one and two. the first one is a p+n-p bjt and the second one is an n-pn+ bjt. r1 is the resistance that is given by the drift region. r2 is the resistance given by the body region. it has been noted that the collector of bjt1 is the same as the base of the collector in bjt2 and vice versa. an equivalent circuit model of igbt is shown below. fig. 5.1. p igbt model equivalent circuit. the simulation result displays the igbt voltage waveform in fig. 5.2. the switching time for igbt model shown in fig. 5.3, is 0.593 ms. to make the switching time of the igbt better for protecting the dc system, capacitors are introduced to the model in parallel with the resistor. fig. 5.4, shows the improved igbt model equivalent circuit with the added capacitors. fig. 5.2. voltage result of igbt model. 54 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) fig. 5.3. switching time of igbt model. fig. 5.4. improved igbt model equivalent circuit. fig. 5.4, shows the improved equivalent circuit of igbt model. the capacitor increases the speed of the switch’s response. the simulation result from fig. 5.5, shows the voltage of the improved igbt during the on-state of the igbt. the switching time for the improved igbt model is 0.553ms as shown in fig. 5.6 because of the shorter response time than the original model, the improved igbt model has been chosen for implementation in sscb for protecting lvdc microgrids. fig. 5.5. switching time of igbt model. protecting a low voltage direct current system 55 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 5.6. switching time of the improved igbt model. 6. modeling the lvdc system under study a circuit schematic has been developed for a post-regulated, isolated dc-dc converter that has two dc loads. a dc-dc converter, specifically, a unidirectional post-regulated isolated dc-dc converter is assembled from a solid-state transformer that has a controlled halfbridge on its primary side and a diode-based half-bridge on the secondary side, followed by a buck converter as displayed in fig 6.1. another important system component to have is a voltage follower. a voltage follower delivers high impedance transformation from one circuit to another, with the purpose of preventing the source from being affected by currents or voltages that the load produces. the corresponding parameters and their respective values are displayed in table 6.1. both loads have an equal rating, protected by sscbs and the load-side has freewheeling diodes that provide current continuity because of load side inductances when the load sscb opens. with the improved igbt model, the capacitor increases the speed of the switch’s response with the improved equivalent circuit of igbt. table 6.1 parameters used in the proposed model and their values parameters specifications c2,c3 c4,c5 v z2, z3, z5 l1 c11 l11 r10 d1-d8 r13, r15, r14, r16 c9, c12, c13, c16 r8, r9 l9, l10 l3, l4, l12, l13 c1, c14 r3, r17 10uf 100uf 60v igbts(ixgh40n60) 0.1h 1u 0.1h 0.003ω diode(d1n4002) 0.5ω 0.01f 0.0015ω 30.25uh 1h 5nf 1kω 56 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) fig. 6.1. modified circuit schematic based on existing lvdc sscb without snubber circuit. fig. 6.2. shows high levels of positively charged overvoltage during switching on times, and negatively charged overvoltage during switching off times which negatively effects equipment, causing overheating which damages electronic devices. fig. 6.2. simulation of modified circuit based on existing lvdc sscb model overvoltage without snubber circuit. the equations are based on capacitor voltage and inductor currents in the distribution system as shown below and, the complete system parameters can be obtained thru consideration of the following equations. (6.1) (6.2) (6.3) (6.4) the laplace transform protecting a low voltage direct current system 57 copyright ©2019 assa. adv. in systems science and appl. (2019) (6.5) (6.6) where: the following schematic of proposed lvdc sscb include improved igbt model and custom design snubber circuit in each load. the parameters of snubber circuits were previously discussed. fig. 6.3. circuit schematic of proposed lvdc sscb with snubber circuit. fig. 6.4. protected sscb model with snubber circuit. 58 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) fig. 6.4. shows a simulation of a circuit schematic of proposed lvdc sscb with a snubber circuit. when compared with fig. 6.2, which is without a snubber circuit, an extremely small overvoltage during switching on times, and negative overvoltage is reduced to zero during switching off times. the custom designed snubber circuit suppresses the voltage spikes caused by the circuit's inductance when operation modes switch. fig. 6.5. displays the switching time of the igbt, the time duration is preformed faster than the existing igbt model, as discussed above. fig. 6.5. proposed sscb model switching time. 7. conclusion this paper offers an illustration of the workings of sscbs in lvdc applications. the motives for the selection of sscbs are discussed along with the design of the suggested lvdc sscbs. advancement in semiconductor technologies have experienced substantial improvement these last few decades and many devices have been designed. in the future, these technologies are only going to get even better. sscbs are swift to respond to faults experienced in the lvdc grid. this makes the dc power system of the future a more reliably sturdy and advanced system that is much better protected than the systems already in place today. further studies on the use of different semiconductor devices such as igcts and gtos enables greater understanding of the difference between the systems of the past and the modern systems of today as well as future systems as new technologies become available. this will enable protection of many different configurations of lvdc distribution systems and microgrid applications. references [1] d. salomonsson, l. soder & a. sannino (2009). protection of low-voltage dc microgrids, ieee transactions on power delivery, 24(3), 1045 – 1053. [2] meyer, c. & de doncker, r. w. (2006). solid-state circuit breaker based on active thyristor topologies, ieee transactions on power electronics, 21(2), 450 –458. [3] magnusson, j., saers, r., liljestrand, l. & engdahl, g. (2014). "separation of the energy absorption and overvoltage protection in solid-state breakers by the use of parallel varistors," in ieee transactions on power electronics, vol. 29, no. 6, pp. 2715-2722, june 2014. [4] corzine, k. a. (2017). a new-coupled-inductor circuit breaker for dc applications, in ieee transactions on power electronics, 32(2), 1411 – 1418. protecting a low voltage direct current system 59 copyright ©2019 assa. adv. in systems science and appl. (2019) [5] almutairy, i. & asumadu, j. (2018). a novel dc circuit breaker design using a magnetically coupled-inductor for dc applications, 2018 ieee international conference on consumer electronics (icce), las vegas, nv, usa, 1 – 5. [6] sano, k., & takasaki, m. (2014). a surgeless solid-state dc circuit breaker for voltage-source-converter-based hvdc systems, in ieee transactions on industry applications, 50(4), 2690 – 2699. [7] zha, x., ning, h., lai, x., huang, y. & liu, f. (2014). suppression strategy for short-circuit current in loop-type dc microgrid, 2014 ieee energy conversion congress and exposition (ecce), pittsburgh, pa, 758 – 764. [8] lai, x., liu, f., deng, k., gao, q. & zha, x. (2014). a short-circuit current calculation method for low-voltage dc microgrid, 2014 international power electronics and application conference and exposition, shanghai, 365–371. [9] hassanpoor, a., häfner, j. & jacobson, b. (2015). technical assessment of load commutation switch in hybrid hvdc breaker, ieee transactions on power electronics, 30(10), 5393–5400. [10] magnusson, j., saers, r., liljestrand, l. & engdahl, g. (2014), "separation of the energy absorption and overvoltage protection in solid-state breakers by the use of parallel varistors," in ieee transactions on power electronics, 29(6), 2715--2722. [11] baier, t. & piepenbreier, b. (2016). bidirectional magnetically coupled t-source inverter for extra low voltage application, 2016 ieee applied power electronics conference and exposition (apec), long beach, ca, 2897--2904. [12] savaliya, s., singh, s. & fernandes, b. g. (2016) protection of dc system using bidirectional z-source circuit breaker, iecon 2016 42nd annual conference of the ieee industrial electronics society, florence, 4217--4222. [13] qi, l. l., antoniazzi, a., raciti, l. & leoni, d. (2017). design of solid-state circuit breaker-based protection for dc shipboard power systems, ieee journal of emerging and selected topics in power electronics, 5(1), 260 – 268. [14] maqsood a. & corzine, k. (2016). dc microgrid protection: using the coupledinductor solid-state circuit breaker, ieee electrification magazine, 4(2), 58–64. [15] capilla, l. a., rodríguez, e. j. j., león, f. j., gordillo-tapia, c. & martínez, j. j. (2017). dc ultra-fast solid state circuit breaker using the voltage inductor to activate short circuit protection, ieee second international conference on dc microgrids (icdcm), nuremburg, 2017, 24–25. [16] gu, c., effah, f., castellazzi, a., waston. a. j. & wheeler, p. (2016) novel currentlimiting strategy for solid-state circuit breakers (sscb) without additional impedance, 2016 18th european conference on power electronics and applications (epe'16 ecce europe), karlsruhe, 1–10. [17] salonen, p., nuutinen, p., peltoniemi, p. & partanen, j. (2009). lvdc distribution system protection — solutions, implementation and measurements, 13th european conference on power electronics and applications, barcelona, 1–10. [18] cuzner, r. m., & venkataramanan, g. (2008) the status of dc micro-grid protection, 2008 ieee industry applications society annual meeting, edmonton, alta., 1–8. [19] park j. d., & candelaria, j. (2013) fault detection and isolation in low-voltage dc-bus microgrid system, ieee transactions on power delivery, (28)2, 779–787. 60 i. almutairy , j. asumadu, z. xingzhe copyright ©2019 assa adv. in systems science and appl. (2019) [20] saeedifard, m., graovac, m., dias, r. f. & iravani, r. (2010) dc power systems: challenges and opportunities, ieee pes general meeting, minneapolis, mn, 1–7. [21] li, w., mou, x., zhou, y., & marnay, c. (2012) on voltage standards for dc home microgrids energized by distributed sources, proceedings of the 7th international power electronics and motion control conference. [22] williamson, b. j. et al., (2011) project edison: smart-dc, 2011 2nd ieee pes international conference and exhibition on innovative smart grid technologies, manchester, 1–10. [23] salonen, p., nuutinen, p., peltoniemi, p., & partanen, j. (2009) protection scheme for an lvdc distribution system, iet conference publications. [24] salomonsson, d., soder, l. & sannino, a. (2009), protection of low-voltage dc microgrids, ieee transactions on power delivery, 24(3), 10451053. [25] yang, j., fletcher, j. e. & o'reilly, j. (2010) multiterminal dc wind farm collection grid internal fault analysis and protection design, ieee transactions on power delivery, 25(4), 2308–2318. [26] almutairy, i. & alluhaidan, m. (2017). protecting a low voltage dc microgrid during short-circuit using solid-state switching devices, in ieee green energy and smart systems conference (igessc), 1–6. [27] candelaria, j. & park, j. d. (2011). vsc-hvdc system protection: a review of current methods, 2011 ieee/pes power systems conference and exposition, phoenix, az, 1–7. [28] almutairy, i., & asumadu, j. & alluhaidan, m. (2017), investigation of igbt behavior in dc systems with advanced circuit breaker technology, 2017 ieee 8th annual ubiquitous computing, electronics and mobile communication conference (uemcon), new york city, ny, 504–508 [29] almutairy, i. & asumadu, j. (2018). a novel dc circuit breaker design using a magnetically coupled-inductor for dc applications, 2018 ieee international conference on consumer electronics (icce), las vegas, nv, usa, , 1–5. [30] magnusson, j., saers, r., liljestrand, l. & engdahl, g. (2014). separation of the energy absorption and overvoltage protection in solid-state breakers by the use of parallel varistors, ieee transactions on power electronics, 29(6), 2715–2722. [32] almutairy, i. & asumadu, j. (2018). examination of breaker-based protection systems for implementation in lvdc sscb applications, 2018 ieee 8th annual computing and communication workshop and conference (ccwc), las vegas, nv, usa, 509–514. [33] yang, j., fletcher, j. e. & o'reilly, j. (2010) multiterminal dc wind farm collection grid internal fault analysis and protection design, ieee transactions on power delivery, 25(4), 2308–2318. adv syst sci appl 2021; 01; 76-85 published online at https://ijassa.ipu.ru. on uniform convergence property of solutions for periodic differential inclusions with asymptotically stable sets mikhail morozov v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: granmiguel@mail.ru abstract: the paper considers a periodic differential inclusion with an asymptotically stable set. the uniform character of convergence of solutions to an asymptotically stable set is established. an exponential estimate is obtained for solutions of a periodic differential inclusion homogeneous in state vector. examples of control systems leading to consideration of periodic differential inclusions are given. these results can find applications in the stability analysis of control systems with periodic parameters, in particular, servomechanisms whose elements operate on ac, control systems with pulse amplitude modulation, and systems used to solve problems related to investigating vibrations of milling machines. keywords: time-invariant differential inclusion, periodic differential inclusion, homogeneous differential inclusion, asymptotically stable set, control system 1. introduction studying control systems has led to the use of differential inclusions theory. consider a control system ),,,( uxtfx = (1.1) where nrx , dtdxx /= is the velocity vector, t is time, and )(tuu = is the control on which the next constraint is imposed ,)( utu  (1.2) u is an arbitrary set, nru  . under fairly general assumptions control system (1.1) with constraint (1.2) is equivalent to the differential inclusion ),,( xtfx (1.3) where ),( xtf denotes a multivalued mapping, i.e., a function that assigns a set nrxtf  ),( to each time t and each point nrx in the state space. differential inclusion (1.3) can be adopted not only for a control system with given control constraints but also for other objects. for example, such objects include systems of differential inequalities, implicit differential equations, control systems with state-space constraints, systems with variable structure and with sliding modes, and differential equations with discontinuous right-hand sides. theory of differential inclusions is well developed and presented in systematic form in [1,4,17]. the use of differential inclusions in control systems theory is covered in the monograph [6]. the examples leading to differential inclusions are given below in section 4. in some cases, e.g., such as the problem of absolute stability, the study of linear nonstationary systems, the matrix of the right-hand side of which satisfies interval on uniform convergence property of solutions for periodic differential… 77 copyright ©2021 assa. adv. in systems science and appl. (2021) constraints, and the stability analysis of control systems that contain elements with incomplete information linear-selectionable differential inclusions can be used. a linearselectionable inclusion is an inclusion of the form ),,( xtfx  ,)()( ,)(:),( ttaxtayyxtf == (1.4) where nryx  , and )(t is a set in the space of nn matrices. an inclusion of the form (1.4) is called a linear-selectionable inclusion, because the multivalued mapping ),( xtf in (1.4) is the union of linear single-valued mappings (selectors) ,)( xta )()( tta  . in the case of a time-invariant linear-selectionable inclusion, the right-hand side f of this inclusion − the matrix a and the set  − are time-invariant. a linear-selectionable inclusion is said to be asymptotically stable if its trivial solution 0x is asymptotically stable. time-invariant linear-selectionable inclusions are studied in a number of publications. for time-invariant linear-selectionable inclusions for which the set  is a compact or convex polyhedron, necessary and sufficient conditions of the zero solution asymptotic stability were obtained in [12,13] on the base of lyapunov functions method. the papers [7,12,13] give various algebraic criteria for the asymptotic stability of time-invariant linear-selectionable inclusions. the publications on time-periodic (in short, periodic) differential inclusions (e.g., [2,8,9]) were mostly devoted to the existence of periodic solutions. few investigations were focused on the analysis of solutions of periodic differential inclusions and their properties. for example, the weak asymptotical and weak exponential stability of an equilibrium of a periodic differential inclusion were studied in [5,18]. in accordance with the definitions introduced therein, an equilibrium of a given periodic differential inclusion is weakly asymptotically (weakly exponentially) stable if there exists at least a single solution satisfying the standard definitions conditions of the asymptotical (exponential) stability for a differential inclusion equilibrium. the method consists in the design of a firstapproximation inclusion and further analysis of the properties of its solutions. in addition, the theorems of the weak asymptotical (weak exponential) stability of an equilibrium of an original inclusion using the corresponding properties of an equilibrium of its first approximation-inclusion were established. it was demonstrated that the proposed method can be used to study the weak asymptotical (weak exponential) stability of the inclusions equivalent to control systems. the problems of absolute and robust stability of control systems with periodic variable parameters were solved in [10,11,14,15]. in particular, control systems with periodic parameters under consideration were proved to be equivalent to a periodic differential inclusion in the sense of the coincidence of the sets of absolutely continuous solutions. as was demonstrated in [16], in some cases solutions of periodic differential inclusions with the asymptotically stable trivial solution have the same properties as solutions of autonomous differential inclusions. this paper continues the research of [16]. it considers periodic differential inclusions with an asymptotically stable set. the remainder of this paper is structured as follows. in section 1 we consider periodic differential inclusion of general form and give preliminary remarks. the definition of an asymptotically stable set is also given. in section 2 the uniform character of convergence of solutions to an asymptotically stable set is established. for solutions of periodic differential inclusion that is homogeneous in state vector we derive an exponential estimate. in section 4 examples of control systems leading to the periodic differential inclusions are given. in the final section we offer concluding remarks. 2. statement of the problem consider the dynamic systems described by periodic differential inclusion 78 m.v. morozov copyright ©2021 assa. adv. in systems science and appl. (2021) 0.t const,t , 0, t),,(),( ),,( =+ nrxxttfxtfxtfx (2.1) everywhere below we assume that in some domain }},:{ , ,0{ 00 rxxggxttg rr == the multivalued function ),( xtf satisfies the main conditions [4, p. 60]; i.e., for all gxt ),( the set nrxtf ),( is nonempty, bounded, closed, and convex, and the function ),( xtf is upper semicontinuous [4, p. 52] with respect to ).,( xt a solution of inclusion (2.1) is understood as an absolutely continuous vector function )(tx defined on an open or closed interval i and satisfying (2.1) almost everywhere on i . by virtue of periodicity of the multivalued function ),( xtf in t , when studying the properties of solutions ),,( 00 xttx of inclusion (2.1), we can assume without loss of generality that ],0[0 tt  . by the definition of solutions and the periodicity of the right-hand side of (2.1) in t the solutions of this inclusion have two properties as follows. if a function )(tx is a solution of inclusion (2.1) (for   t ), then 1) the function ),( kttx + where kttkt −−  and k denotes any integer, is also a solution of inclusion (2.1); moreover, the solutions )(tx and )( kttx + have the same trajectory; 2) for any tttt and , ],,0[ 10  such that ,10 ttt  the equality =))(,,( 11 txttx ),,( 00 xttx holds, where ),,()( 0011 xttxtx = . let nra , nrb be points (vectors) with coordinates ia and ib respectively, ni ,1= , and also let nrb  be a set. the distance  between two points or between a point and a set is interpreted as the nonnegative values ,)(),( 2/1 2 1       −=−=  = n i ii bababa ),(inf),( baba bb = . as is well-known, the function ),( bx is uniformly continuous and for any points nrx and nry )y,(),(y-),(x xbb   . a closed  -neighborhood m of a set m is a set of such points x that  m),(x . let gm 0  , .00  definition 2.1: a set m is asymptotically stable for inclusion (2.1) if for any 0 there exists a value 0)(  such that, for each 0x satisfying the inequality )(),( 0  mx , there exists a solution with the initial condition 00 )( xtx = and all solutions with the abovementioned property are extendable on the interval  tt0 and also satisfy the conditions  )),(( mtx for  tt0 and 0)),(( →mtx as →t . the problem is to study solutions of inclusion (2.1) with an asymptotically stable set .m 3. results theorem 3.1: if inclusion (2.1) has an asymptotically stable set ,m then there exists a value 00  such that all solutions ),,( 00 xttx of inclusion (2.1) satisfy the condition on uniform convergence property of solutions for periodic differential… 79 copyright ©2021 assa. adv. in systems science and appl. (2021) 0)),(( →mtx as →t (3.1) uniformly with respect to ),( 00 xt for any ],0[0 tt  and . 0 0  mx  proof. first, let us show that there exists a value 00  such that all solutions of inclusion (2.1) with initial condition =)( 0tx ,0x , 0 0  mx  satisfy condition (3.1) uniformly with respect to 0 x for any given ].,0[0 tt  suppose the contrary. then for any ,0 there exists ,0)(  ],0[0 tt  , a sequence of solutions )(txk of inclusion (2.1), a numerical sequence ktttk + 0 , and a sequence of vectors ,...,2,1 ,0 =kxk such that mxk 0 and . tas ,...,2,1 ,0)()),,,(( k00 →= kmxttx k kk  since the set m is asymptotically stable, it follows that for any 0 we can choose a value  small enough to ensure that all solutions with mtx )( 0 satisfy the relations  ))),(,,(( 00 mtxttx ),( 0  tt 0)),(( →mtx ).(t → (3.2) consequently, there exists a number 0 such that all solutions of inclusion (2.1) with mtx )( 0 satisfy the inequality )())),(,,(( 00  mtxttx ).( 0  tt (3.3) let us show that all ),,( 00 k k xttx satisfy the inequalities , 0 mxk  1,2,...k ,k1,m ,)),x,,(( k 000 ==+  mtmttxk (3.4) indeed, otherwise there would exist k ~ and ) ~ (km ( kkm ~ ) ~ (1  ) such that ,)),x,,) ~ ((( k ~ 000~  + mttkmtx k . ~ 0 mxk  then the solution =+++=+= )),,) ~ ((,) ~ (,) ~ (()),,) ~ ((,,()( ~ 000~0~ ~ 000~0 k kk k k xttkmtxtkmttkmtxxttkmtxttztz ),,) ~ (( ~ 00~ k k xttkmtx += )( 0  tt of inclusion (2.1) would satisfy the relations =)),(( 0 mtz ,)),x,,) ~ ((( k ~ 000~  + mttkmtx k ),()),,,(())),,,) ~ ((,,)tk ~ ((( ~ 00~~ ~ 000~0~  =+− mxttxmxttkmtxtmtz k kk k kk which contradict (3.3). the sequence of segments of solutions )(txk contains a subsequence uniformly convergent for 10 ttt  ; in turn, the subsequence contains a new subsequence uniformly convergent for 20 ttt  , and so on. since ),( xtf satisfies the basic assumptions, it follows that the limit function )(tx is a solution of inclusion (2.1) [4, p. 60], which, together with (3.4), implies that mtx )( 0 and  + ))),(,,(( 000 mtxtkttx , 1,2,...,k = and we have arrived at a contradiction with (3.2). therefore, the assertion started at the beginning of the proof of the theorem is valid. 80 m.v. morozov copyright ©2021 assa. adv. in systems science and appl. (2021) let us now prove the assertion of theorem 2.1. suppose that is fails. then for any ,0 there exist a number ,0)(  a sequence )(txk of solutions of inclusion (2.1), a sequence of numbers kt0 ( ],0[0 tt k  and kttt k k + 0 ), and a sequence of vectors ,...,2,1 ,0 =kxk such that mxk 0 , . t,...;2,1 ,0)()),,,(( k00 →= kmxttx kk kk  (3.5) without loss of generality, we can assume that ,0  where 0 is a number for which the above-proved assertion holds. since , 0 mxk  ],0[0 tt k  , ,...,2,1 =k it follows that the sequence ,...2,1 =k contains a subsequence ),( }{ → kk such that the limits 00lim xxk k =  → ( m0x ) and 00lim tt k k =  → ( ],0[0 tt  ) simultaneously exist. to simplify the notation, we assume that kk tt 00 =  , kk xx 00 =  ,...,2,1 k == k and ,lim 00 xx k k = → .lim 00 tt k k = → (3.6) let .0m),( 00 =−  x by (3.6), there exist a 1k such that 2/)x,( 00  kx (3.7) for all 1kk  since all solutions of inclusion (2.1) are equicontinuous [2, p. 61] in any closed bounded domain lying in ,g it follows that they are equicontinuous in the domain , { 0mx ]},0[ tt as well. therefore, for the point 0t , there exist a neighbourhood ]},,0[ ,2/:{)( 00 tttttts −=  ],[)( 0 bats = ( ta 0= if 00 =t and 0tb = if tt =0 ), such that .2/))t ~ x(,)) ~ (, ~ ,(( 000  txtbx (3.8) for all solutions )) ~ (, ~ ,( 00 txttx of inclusion (2.1) with ],[ ~ 0 bat  and .) ~ ( 0 mtx  then, taking into account relations (3.7), (3.8) and uniform continuity of the function ),()( mxx  = we obtain two relations 2/)x,(),(x-),(x),(),( 000 k 000  − kk xmmmxmx , 2/)x),,,((),(x-)),x,t(b,(x),()),,,(( k 000 k 0 k 0 k 0k000  − kk k kkk k xtbxmmmxmxtbx for all solutions )(txk with .1kk  adding the last two inequalities, we obtain 0000 ),()),,,((  =+ mxmxtbx kk k , .1kk  this, together with (3.5) and (3.6), means that there exist a 2k such that the relations ,0 tbt k  000 )),,,((  mxtbx kk k , )()),,,(( 00  mxttx kk kk (3.9) are simultaneously valid for all .2kk  taking into account relations (3.9), from the sequence of solutions )(txk of inclusion (2.1) satisfying inequalities (3.5) for 2kk  , we pass to the sequence of solutions )(tyk ),,( 00 kk k xttx= defined by the relations ),()( 0 bxyby k k k == ,),( 00  myk ),()),y,,(( k 0  mbty kk .k →t (3.10) on uniform convergence property of solutions for periodic differential… 81 copyright ©2021 assa. adv. in systems science and appl. (2021) the existence of a sequence ),(tyk ,2kk = ,12 +k ,...22 +k ,(  tb ),0 tb  satisfying conditions (3.10) contradicts the already known fact that relation (3.1) is valid uniformly with respect to 0x (for 0 0  mx  and ],0[0 tt  ) for all solutions of inclusion (2.1). the proof of theorem 3.1 is complete.  in what follows, we consider differential inclusions homogeneous with respect to .nrx if b is a set in nr and c is a number, then cb stands for a set of points of the form cx for all .bx definition 3.2: a multivalued function ),( xtf is homogeneous (of degree one) in x if ),( cxtf ),( xtcf for all 0c . definition 3.3: a differential inclusion )0 ),,(),(( ),(  cxtcfcxtfxtfx (3.11) is homogeneous in .x homogeneous differential inclusion (3.11) is preserved under the substitution 1cxx = with arbitrary .0c it means that if a function )(tx = is a solution of inclusion (3.11), then the function )(tcx = with arbitrary 0c is also a solution. let us consider the differential inclusion ), 0,( ),,( nrxtxtfx  0),t const,(t ),,(),( =+ xttfxtf ),0( ),(),(  cxtcfcxtf (3.12) periodic in t and homogeneous in .x since inclusion (3.12) is homogeneous in ,x we have the following assertion. corollary 3.1: if a bounded set m is asymptotically stable for inclusion (3.12), then all solutions ))(,,( 00 txttx of inclusion (3.12) with the initial conditions ],0[0 tt  and rgx 0 (where r is an arbitrary positive number) satisfy condition (3.1) uniformly with respect to ).,( 00 xt for solutions of periodic homogeneous differential inclusion (3.12) with an asymptotically stable set m an exponential estimate is valid. theorem 3.2: if a bounded set m is asymptotically stable for inclusion (3.12), then there exist numbers 0 ,0 10  cc such that any solution ),,( 00 xttx of inclusion (3.12) satisfies the estimate ).(t )exp()),x,,(( 010000 − ttcxcmttx (3.13) for any 0t and .0tt  proof. by virtue of corollary 3.1, for some ,0 there exists a 0 (independent of ),( 00 xt ) such that tk ~ = ( k ~ is some positive integer) and ,2/)),x,,(( 00  mttx + t0t , for all solutions ),,( 00 xttx with the initial condition .0 x by virtue of theorem 3 in [4, p. 62] the set of solutions of inclusion (3.12) is compact on the closed interval ]t,[ 00 +t in the metric of ]t,[ 00 +tc , whence it follows that  200 )),x,,(( cmttx  )0( 2 c for .t00 + tt if ),,( 00 xttx is a solution with an arbitrary 00 x and ,/ 0xc = then the function ),,(),,( 0000 xttcxytty = is also a solution and ,0 =y whence it follows that 82 m.v. morozov copyright ©2021 assa. adv. in systems science and appl. (2021) 2/)),y,,(( 00  mtty )t( 0 + t and  200 )),y,,(( cmtty  ).t( 00 + tt returning from ),,( 00 ytty to ),,( 00 xttx , we obtain 0200 )),x,,(( xcmttx  ),t( 00 + tt 2/)),x,,(( 000 xmttx  ).t( 0 + t (3.14) for any 0t and .0x since after replacement of t by ktt + ( k is an arbitrary integer), a solution of inclusion (3.12) remains a solution, it follows from (3.14) that a solution ),,( 00 xttx with an arbitrary 0x satisfies the relations 000 2)),x,,(( xmttx i− ),t( i  t (3.15) where ,...2,1 ,t 0i =+= iit  for any ,0tt  we choose an i such that .t 1-i itt  then itt + 0 and ./)( 0 tti − therefore, the inequality )/2lnexp()/2lnexp()/)(2lnexp()2lnexp(2 00  ttttii −=−−−=− is fulfilled. the last inequality, together with (3.15), implies (3.13) with )/2lnexp( 00 tc = and ./2ln1 =c the proof of theorem 3.2 is complete.  4. exapmles example 4.1: in the paper [13] the nonlinear control system is considered ,),( 1  = += m j jj j tbaxx   = == n i i j i j j xcxc 1 ,, ,,1,0),,0( mjtj = (4.1) where nn ii rxxx = = ,)( 1 is n -dimensional vector, characterizing the deviation of the system from the mode prescribed by control objective (the zero solution otx )( of system (4.1) corresponds to this mode), a is a constant square matrix of order ,n and jb and ),1( mjc j = are constant n -dimensional vectors. the brackets .,. denote the scalar product. it is also assumed that nonlinear functions ,,1),,( mjtjj = that define the characteristics of nonlinear elements, satisfy conditions for the existence of an absolutely continuous solution of system (4.1) for any initial conditions and inequalities ( )mjkkktk jjjjjjjjj ,1, ),( 21 2 2 2 1 =+−  for all j and .0t system (4.1) is equivalent, see the paper [13], to time-invariant linear-selectionable inclusion (1.4), where  is a compact set in the 2n -dimensional matrix space. here and in the following examples, equivalence is understood in the sense of solutions sets identity for the system under consideration and the corresponding inclusion under the same initial conditions. example 4.2: in the paper [10] the linear nonstationary control systems ,)()( 1 xtatx k m k k = =  0)( tk , 1)( 1 = = t m k k (4.2) on uniform convergence property of solutions for periodic differential… 83 copyright ©2021 assa. adv. in systems science and appl. (2021) are considered, where nт ii rxx = =1)( is n -dimensional state vector of the system, =)(tak (t))(a n 1ji, k ij = are given continuous periodic matrices of the period ,0t =+ )( ttak )(tak , mk ,1= , and )(tk ( mk ,1= ) are bounded measurable functions. system (4.2) is equivalent to the periodic differential inclusion ),,(),( ),,( xttfxtfxtfx +  == === m kk m k k tttatyyxtf 1k k 1 }.1)( ,0)( ),()(:{),(  (4.3) example 4.3: in the paper [14] the linear nonstationary control system nrxxtax = ,)( (4.4) is considered, where n jiij tata 1,))(()( == is an arbitrary matrix ( )(taij , nji ,1, = , are measurable functions, generally speaking, not periodic). system (4.4) almost everywhere satisfies the inequalities ))(()( ,))(((t) (t),(t) (t) 1,1j, n jiij n nij tatataaaaa == == (4.5) on any finite interval of the semi-axis ).[0, matrix inequalities (4.5) are understood elementwise, that is ,,1, (t),)((t) ijij njiataa ij = , where ,)(a tij , ,1, (t), (t), ijij njiaa = are arbitrary measurable functions. it is assumed that the given bounded "extreme" matrices (t) a and (t)a are periodic with a period 0t  , i.e. conditions are valid (t) )( atta + , ).()( tatta + thus, by virtue of (4.5), not one fixed system (4.4) is considered, but a set of linear nonstationary systems (4.4) with periodic interval constraints (4.5). a series of examples leading to systems (4.2) and (4.4) with constraints (4.5) were considered in [19]. among such systems, note tracking systems with ac elements, control systems with pulse amplitude modulation and also the systems arising in the vibration analysis of milling machines. if the condition )()( tatta =+ is additionally satisfied, the set of linear nonstationary systems (4.4) with periodic interval constraints (4.5) is equivalent to the linear-selectionable periodic inclusion ),,( xtfx  )()( ,)(:),( ttaxtayyxtf == ,  )()()()()( :)()( 21 tattattatat  +== , )()( ttt =+ , where arbitrary bounded and measurable functions 2,1),( =ktk satisfy the conditions ,1)( ,0)( 2 1 =  =k kk tt  ),()( ttt kk  =+ .0tt  5. conclusion the paper considers solutions of periodic differential inclusions with asymptotically stable sets. theorem 3.1 establishes the uniform character of convergence to an asymptotically stable set. homogeneous in state vector periodic differential inclusions are also considered. corollary 3.1 is obtained for such inclusions. theorem 3.2 gives the exponential estimate for solutions. theorem 3.1 is a generalization of lemma 2, proved for autonomous differential inclusions [3]. theorem 3.2 generalizes the estimate well known for solutions of an autonomous homogeneous (first degree) differential inclusion [4]. the examples of control systems leading to consideration of periodic differential inclusions are given. the results obtained can find applications in the stability analysis of control systems with periodic parameters, in particular, servomechanisms whose elements operate on ac, control 84 m.v. morozov copyright ©2021 assa. adv. in systems science and appl. (2021) systems with pulse amplitude modulation, and systems used to solve problems related to investigating vibrations of milling machines. further research into the inclusions considered in the present paper can be related to producing weak asymptotic and weak exponential stability conditions. in addition, it seems interesting to distinguish the classes of lyapunov functions establishing necessary and sufficient conditions for the asymptotic stability of periodic differential and difference inclusions. references 1. aubin, j.-p., & cellina, a. (1983). differential inclusions. berlin heidelberg: springerverlag. 2. filippakis, m., michael e., & papageorgiou, nikolaos s. (2006). periodic solutions for differential inclusions in ,nr archivum mathematicum., 42(2), 115−123. 3. filippov, a. f. (1979). ustoychivost’ dlia differentsial’nyh uravneniy c razryvnymi i mnogoznachnymi pravymi chastiami [stability for differential equations with discontinuous and multivalued right-hand sides], differets. uravn., 15(6), 1018-1027, [in russian]. 4. filippov, a. f. (1985). differentsial’nye uravneniya s razryvnoi pravoi chast’yu [differential equations with discontinuous right-hand side]. moscow, ussr: nauka, [in russian]. 5. gama, r. & smirnov, g.v. (2013). weak exponential stability for time-periodic differential inclusions via first approximation averaging, set-valued and variational analysis, 21(2), 191-200, https://doi.org/10.1007/s11228-012-0216-1 6. han, z., cai, x., huang, j. (2016). theory of control systems described by differential inclusions. berlin heidelberg: springer. 7. ivanov, g.g., alferov, g.v., & efimova, p.a. (2017). ustoychivost’ selektornolineynyh differentsial’nyh vklucheniy [stability of selector-linear differential inclusions]. vestnik permskogo universiteta, matematika, mekhanika, informatika. 2(37), 25-30, [in russian], https://doi.org/10.17072/1993-0550-2017-2-25-30 8. jack w., macki, paolo, nistri. & pietro, zecca (1988) the existence of periodic solutions to non-autonomous differential inclusions, proc. of the amer. math. soc., 104(3), 840-844, https://doi.org/10.2307/2046803 9. li, g.&xue, x. (2002). on the existence of periodic solutions for differential inclusions, j. math. anal. appl., 276(1), 168−183. 10. molchanov, a. p. & morozov, m. v. (1997). algoritmy analiza robastnoy ustoychivosti lineynih nestatsionarnyh system upravleniya c periodicheskimi ogranicheniyami [algorithms for robust stability analysis of linear time-varying control systems with periodic constraints]. automation and remote control, 58(5(2)), 795-804, [in russian]. 11. molchanov, a. p. & morozov, m. v. (1992). absolutnaya ustoychivost’ nelineynih nestatsionarnyh system upravleniya c periodicheskoy lineynoy chast’u [absolute stability of nonlinear nonstationary control systems with periodic linear sections]. automation and remote control, 53(2(1)), 189-198, [in russian]. 12. molchanov, a. p. & pyatnitskii, e.s. (1989). criteria of asymptotic stability of differential and difference inclusions encountered in control theory, systems & control letters., 13, 59−64. 13. molchanov, a. p. & pyatnitskii, e.s. (1986). funktsii lyapunova opredelyaushie neobhodimye i dostatochnie usloniya absolutnoy ustoychivosti nelineynyh nestatsionarnyh system upravlenia. [lyapunov functions defining the necessary and sufficient conditions for absolute stability of the nonlinear nonstationary control systems.]. avtom. & telemekh., 3, 63−73, [in russian]. on uniform convergence property of solutions for periodic differential… 85 copyright ©2021 assa. adv. in systems science and appl. (2021) 14. morozov, m. v. (2016). kriterii robastnoy ustoychivosti nestatsionarnih system interval’nimi ogranicheniyami [robust stability criteria for nonstationary systems with interval constraints]. proceedings of the institute for systems analysis, 66, 4, 4-9 [in russian]. 15. morozov, m.v. (2014). kriterii robastnoy absolutnoy ustoychivosti diskretnyh system upravleniya s periodicheskimi ogranicheniyami [criteria of robust absolute stability for discrete control systems with periodic constraints]. proceedings of institute for systems analysis, 64, 2, 13-18. [in russian]. 16. morozov, m. v. (2000). svoystva resheniy periodicheskih differentsial’nyh vklucheniy [properies of solutions of periodic differential inclusions]. differets. uravn., 36(5), 677682. [in russian], https://doi.org/10.1007/bf02754225 17. smirnov, g. v. (2002). introduction to the theory of differential inclusions. providence, rhode island, usa: amer. math. soc. graduate studies in mathematics, 41. 18. smirnov, g. v. (1995). weak asymptotic stability at first approximation for periodic differential inclusions, nonlinear differential equations and applications., 2(4), 445 461, https://doi.org/10.1007/bf01210619 19. shil'man, s.v. (1978). metod proizvodiatshikh funktsiy v teorii dinamitcheskikh sistem. [the method of generating functions in the theory of dynamic systems]. moscow, ussr: nauka, [in russian]. 1. introduction microsoft word 10-zhang yong.doc advances in systems science and applications (2010), vol.10, no.1 61-66 issn 1078-6236 international institute for general systems studies, inc. research on dynamic pricing model based on e-supply chain management* yong zhang1,2, yingjin lu1 and xianglan jiang1 1school of management, university of electronic science & technology of china, chengdu 610054, china 2 college of information science and technology, chendu university, chengdu 610054, china email: luyingjin@uestc.edu.cn, yingjin.lu@changhong.com abstract with the influence of the uncertain factor in business trade, the research on dynamic pricing considering the random fluctuation of price has become an important project in managerial economics. in this paper, we introduce the uncertain factor, which produced by the random errors, into the pricing model of the dominant manufacturers. by introducing the expectation of retail price, variance and transfer price, we change the game of incomplete information into the game of complete information. then we resolve the problem of optimal pricing about transfer price with extremum of function according to the optimal production and storage model. keywords random fluctuation dynamic pricing incomplete information game extremum of function 1. introduction based on the network and electronic commerce (ec), the internet can arrive any place of the world. it results in the substantial increase of demand or supply under non-prediction. at the same time, lower menu cost supports traders to change the price more frequently according to the condition of market. so it is less and less efficient to confirm the optimal price according to establish the static demand price model under the random fluctuation of demand and price. many scholars researched the dynamic pricing model. the article[4] discussed the dynamic pricing model of the dominant manufacturer after the time factor was added to the demand function. the article[5] suggested introducing uncertain factor into economic model and establishing the price model of rational expectation. the article[6] displayed the optimization of dynamic game and nonlinear pricing and the dominant firm price leadership model completely. based on the above articles, this paper discuss the pricing model of the dominant manufacturer further with uncertain factor, which produced by the random errors in ec. by introducing uncertain expectation and variance, it changes the game of incomplete information into the game of complete information. then it solves the problem of optimal pricing about transfer price with the extremum function method according to the optimal production and storage model. 2. analyzing and hypothesizing of the model assumed that a new kind of production is in monopoly position, and there is only one manufacturer and one seller. because of the new production, the manufacturer is in dominance and can predict the demand of market. the manufacturer confirms the transfer prices of its production according to the maximum profit. then, the seller confirms the order quantity according to the maximum profit. this process can be considered as the incomplete information game with the decentralized decision. this paper considers uncertain factor as an endogenous variable and introduce it into the model. then it assumes the demand in consumption market is * this work was supported by fund item for the doctoral program of high education (no. 20030614011), national scientific fund item for excellent youth (no. 79725002), chinese science fund item for post doctor (no. 79725002). zhang: research on dynamic pricing model based on e-supply chain management 62 related not only to the price, but also to the price fluctuation which is the uncertain factor of price, so it expresses price fluctuation with price variance. 3. supply chain model 3.1 optimal order strategy of sellers after manufacturer price the production 1p with maximum profit according to some information, sellers choose the optimal ordering quantity. supposed the lead time t k− is fixed, and the period of order t is constant. when sellers estimate the ordering quantity according to the profit maximization, they order t k− days in advance. then, the products will arrive after k days. after sellers introduce uncertain factor into economic model as inner variable, they can price according to the demand of the market in internet. the demand function of consumers in this period is: 1 1( ) ( ( ), , ) ( )u u s s s s sq t d p t p u a bp t p uλ= = − + + , (1) where ( )sp t is the sell price at point t, and ( )sq t is the demand amount at point t. both of them are random process. u sp is the variance of price (fixed), which reflects the fluctuation of price. if λ is negative, consumers hate the price undulation. if λ is positive, consumers like the price undulation. 1u is a random disturbance. and the parameters , 0a b > . (and the parameters a and b are greater than zero.) because of the predomination of manufacturer, transferring price 1p is assured. seller assures sell price ( )sp t , and market gives the amount of demand ( )sq t . assumed all the demands are satisfied under ( )sp t , the order quantity in this period is the accumulation of instantaneous demand ( )sq t , that is, 0 ( ) t sq q t dt= ∫ , (2) the storage at point x is 0 0 0 ( ) ( ) ( ) ( ) x t x t s s s s sx q q q t dt q t dt q t dt q t dt= − = − =∫ ∫ ∫ ∫ , (3) if the storage cost of unit production is 2c at unit time, the storage cost is 2 2 20 0 ( ) ( ( ) ) ( ) t t t s s s sx d c c q x dx c q t dt dx c q t dtdx= ⋅ = =∫ ∫ ∫ ∫∫ , 2 20 0 0 ( ( ) ) ( ) t t t s sc q t dx dt c t q t dt= = ⋅∫ ∫ ∫ , (4) where d is a triangle, surrounded by 0, , ,x x t t x t t= = = = . the purchase cost is 1 1 1 10 ( ) t b sc p q c p q t dt c= + = +∫ , (5) where 1c is fixed purchase cost, so total cost is s b stc c c= + , (6) the revenue of the seller is si 1 0 0 1( ) ( ) [ ( ) ] ( ) t t u s s s s s uap t q t dt q t p q t dt b b b b λ = = − + +∫ ∫ advances in systems science and applications (2010), vol.10, no.1 63 2 1 0 1[ ( ) ( ) ( ) ( )] t u s s s s s ua q t q t p q t q t dt b b b b λ = − + +∫ , (7) and the profit of the seller in order period is s s s b sr i tc i c c= − = − − 2 1 1 2 10 0 0 1[ ( ) ( ) ( ) ( )] ( ) ( ) t t tu s s s s s s s ua q t q t p q t q t dt p q t dt c t q t dt c b b b b λ = − + + − − × −∫ ∫ ∫ 2 1 1 2 10 1[ ( ) ( ) ( ) ( ) ( ) ( )] t u s s s s s s s ua q t q t p q t q t p q t c t q t dt c b b b b λ = − + + − − × −∫ . (8) assumed supply quantity equal demand quantity in every period, the seller predicts and chooses supply quantity ( )sq t and its sell price ( )sp t to get maximize the profit sr . if 2 1 1 2 1( ) ( ) ( ) ( ) ( ) ( )u s s s s s s s uaf q t q t p q t q t p q t c t q t b b b b λ = − + + − − × . we will get maximum profit sr with euler equation. then ( ) 0 [ ] '[ ]s s f d f q t dt q t ∂ ∂ − = ∂ ∂ (9) solving this differential equation, we can get the sell function at point t when the profit is maximum: 1 2 1 1( ) ( ) 2 u s sq t a bp p bc t uλ= − + − + (10) *q is the optimal order quantity for sellers after the wholesale price is determined, that is, 2 2 1 2 1 1 10 2 1 1 0.51( ) ( ) 4* 0.5 0 . u t u s s s u s a p bc t u q t dt bc t a bp p u t p bq a p bc t u p b λ λ λ ⎧ + − + = − + − + + ≤⎪⎪= ⎨ + − +⎪ >⎪⎩ ∫ (11) 3.2 rational expectations of the best product plan and transferring price that manufacturer determined assumed the demand function of the market and the order interval of the seller are open. then, the order quantity of seller can be predicted, and the transferring price 1p can be determined under the maximum profit. the manufacturer signs a contract for goods with the seller. according the contract, the seller purchase *q productions with 1p and the manufacturer deliver at point k . for the manufacturers, they must consider sales revenue, production cost and storage cost. sales revenue is multiplication of price and order quantity. production cost depends on production rate, which is the quantity in unit time. and the production rate is higher, the production cost is more. storage cost is determined by the finished productions and expiration time. ( )x t is the cumulative production of product plan until point t without consideration of assets’ time value. because the product rate at point t (marginal product) is '( )x t , marginal product cost is ( '( ))f x t , marginal storage cost is ( ( ))g x t and marginal revenue is ( '( ))x tϕ . then, the total cost ( ( ))c x t is zhang: research on dynamic pricing model based on e-supply chain management 64 0 ( ( )) [ ( '( )) ( ( ))] k c x t f x t g x t dt= +∫ . (12) in order to confirm the actual expression of this function, we assume the following items: (1) the cost of increasing one production in uniform time is in proportion to the product rate, where the coefficient is 12k . so, we get 1 ( '( )) 2 '( ) '( ) df x t k x t dx t = , or, 2 1( '( )) [ '( )]f x t k x t= . (13) (2) the storage cost is in proportion to storage amount in unit time, where the coefficient is 2k , that is, 2( ( )) ( )g x t k x t= (14) the total revenue ( ( ))i x t from point 0 to point k is: 10 0 ( ( )) ( '( )) '( ) k k i x t x t dt p x t dtϕ= =∫ ∫ (15) the profit r from point 0 to point k is: 0 0 ( ( )) ( ( )) ( ( )) ( ( )) [ ( '( )) ( ( ))] k k r r x t i x t c x t x t dt f x t g x t dtϕ= = − = − +∫ ∫ 2 1 1 20 [ '( ) ' ( ) ( )] k p x t k x t k x t dt= − −∫ (16) (0) 0, ( ) *x x k q= = (17) the problem may be attributed to find the maximum of functional ( ( ))r x t , which can be solved by variation method. if 2 1 1 2( , , ') ' 'mr t x x p x k x k x= − − , we will get the maximum profit r according to euler equation, we have to let: ( ) 0 ' mr d mr x dt x ∂ ∂ − = ∂ ∂ then we can get the differential equation of second order: 2 12 ''( ) 0k k x t− + = under the constraint condition (17), its solution is: * 2 22 1 2 1 1 4 ( ) 4 4 k k q k k x t t t k k k − = + . (18) it’s the product plan of maximum profit. obviously, with ( ) 0,0x t t k≥ ≤ ≤ , we can get 2 * 2 14 k k q k ≥ . (19) that is, when the order quantity of the sellers satisfies condition (19), the product plan confirmed by condition (18) is optimal. if 2 * 2 14 k k q k < , we can delay the start time which make the time difference k satisfying equation 2 * 2 14 k k q k = from t1 to k, so the product plan is better. after we substitute ( )x t in the equation (16), the maximum profit is solved according to 1p and ( )x t : advances in systems science and applications (2010), vol.10, no.1 65 3 2 * * 2 *2 2 1 1 1 ( ) 48 2m k k kk q k q r p q k k = − + − (20) take non-zero of *q in equation (11) into equation (20), we get 3 2 22 2 1 2 1 1 1 1 ( ) 48 2 4 u m s k k kk r p bc t a bp p u t k λ⎛ ⎞ ⎡ ⎤= + − − + − + +⎜ ⎟ ⎢ ⎥⎣ ⎦⎝ ⎠ 2 21 2 1 1 1 ( ) 4 u s k bc t a bp p u t k λ⎡ ⎤− − + − + +⎢ ⎥⎣ ⎦ (21) obviously, r is a quadratic and convex function of price 1p . if we get maximum r with appropriate 1p , that is 1 0dr dp = , then 2 2 2 2 1 1 1 2 2 1 1 1 1 2 4 ( 2 ) 4 ( ) 2 8 ( ) 8 ( ) u u s sb c k t a k bk t k p u bk k bc kt bk t p u p b k bk t λ λ− + + + + + − + + = + (22) 1 ep is the expectation of the best price 1p , 1 1( )ep e p= 2 2 2 1 2 1 2 2 1 1 1 4 ( 2 ) 2 2 2 ( ) 8 ( ) 2 ( ) u s a k bk t b c k t bk k bc kt k bk t e p b k bk t b k bk t λ + − + − + = + + + . assumed variance u sp of transferring price sp is a constant. then 2 2 2 1 2 1 2 2 1 1 1 1 4 ( 2 ) 2 2 2 8 ( ) 2 ( ) e u s a k bk t b c k t bk k bc kt k bk t p p b k bk t b k bk t λ + − + − + = + + + (23) so, 1 ep is made up of two terms. one is determined by the structure of model, the other involves retail price undulation. when 0λ > , 1 ep is in proportion to u sp . u sp is higher, 1 ep is higher. it is coincident with the theory, which is the expectation revenue of asset increased with the risk. in a word, under the incomplete information, the expectation of economical people is scaled to uncertainty rather than related to it only. 4. calculation examples 1( ) 970 3 ( )s sq t p t u= − + assumed 1000, 3, 2a b λ= = = − and 15u sp = in equation (1), the demand function of the consumer in order cycle time is: 1( ) 970 3 ( )s sq t p t u= − + if 7t = , we can confirm the lead time is two days according to some conditions, such as transportation, so 2 5k t= − = . the coefficient between rate of change and production rate 1 2 k is 1 when the product rate increase unit product cost, and the coefficient between storage cost and storage amount 2k is 0.05. and the purchase cost 1c is 100 yuan, holding cost per unit production in unit time 2c is 1 yuan. above data can be adjusted and modified according zhang: research on dynamic pricing model based on e-supply chain management 66 to the actual statistic and forecasting. because zero is the average of 1u , we substitute mathematical expectation ( )sq t for ( )sq t when we calculated. according to equation (23), we can get 1 314.78 0.32 305.31e u sp pλ≈ + = . and with 1 321.58ep ≤ from equation (11), the optimal sale *q is 360 when the manufacturer determine price. according to equation (21) and (8), the maximum profit of the manufacturers rm is 58028.6, and the maximum profit of the wholesalers sr is 1295.17. from the data simulation, we can know that the manufacturers have the active power and control the profit mostly because of the open and dominant price under the monopoly and competition. 5. conclusions comparing to definite situation,the hypothesis containing fluctuate price is more close to the fact. with the development of the electronic commerce, the forecast and decision should be more reliable and it is more important to study the random situations. we have made some effort on the related aspects in this paper and got the optimal strategy of pricing under the situation of the fluctuating price in random. (see formula (23)) the application of this strategy is important to resolve the questions in commodity exchange with uncertainty factors references [1] luo zhanghua, liu lu. research on dynamic pricing mechanism in transactions[j]. beijing university of aeronautics and astronautics journal (social science edition), 2005,18 (1). [2] kannan p k, praveen k kopalle. dynamic pricing on the internet: importance and implications for consumer behaviour[j]. international journal of electronic commerce, 2001, 5 (3): 63283. [3] gallego g, van ryzin g j. a multiple product dynamic pricing problem with applications to network yield management [j]. operations research, 1997, (45): 24241. [4] liu jian, lai ming-yong, zhang han-jiang. the optimization of supply chain ordering and pricing decisions based on the trend of demand [j]. system engineering, 2003, 21 (5). [5] cao caixia, zhang yan. rational expectations model with uncertainties [j]. journal of the beijing union university (natural science) 2004, 18 (1). [6] tang xiaowo, zeng yong, li shiming, et al.. management economic analysis -theory and application [m]. chengdu: university of electronic science and technology of china press, 2000. [7] jiang qiyuan, xie jinxing, ye jun. mathematical model [m]. beijing: higher education press, 2003. adv syst sci appl 2021; 03:40–62 published online at https://ijassa.ipu.ru. a controllability problem for causal functional inclusions with an infinite delay and impulse conditions maria afanasova1, valeri obukhovskii1, garik petrosyan1,2* 1faculty of physics and mathematics, voronezh state pedagogical university, voronezh, russia 2research center, voronezh state university of engineering technologies, voronezh, russia abstract: in this paper we study the controllability problem in a banach space for various classes of functional inclusions with causal operators with an infinite delay, and impulse effects. basing on the topological degree theory for condensing multimaps, we prove a global theorem on the existence of trajectories for systems governed by functional inclusions. as an application, we obtain generalizations of existence theorems for the controllability problem for a semilinear first order functional differential inclusions of this type and a semilinear functional differential inclusions of a fractional order 0 < q < 1. keywords: causal operator, functional inclusion, controllability problem, functional differential inclusion, fractional derivative, measure of non-compactness, fixed point, topological degree, condensing multioperator 1. introduction it is well known that the contemporary approach in the theory of control systems and in mathematical physics leads to models which are convenient to be described by using differential equations and inclusions. recently, the attention of many researchers (see [1], [2], [3] and references therein) has been attracted to generalizations of differential equations and inclusions, namely to the class of functional equations and inclusions with causal operators. the term of a causal or volterra operator in the sense of a.n. tikhonov (see [4]), was used in mathematical physics to solve problems of differential equations, integro-differential equations, functional-differential equations with a finite or infinite delay, integral equations of volterra type, functional equations of a neutral type, etc. (see, for example, [5]). the papers [6], [7], [8], [9] among others are devoted to the study of equations and inclusions with causal operators of various types, theorems on the existence of solutions, description of qualitative properties of solutions and various applications. at the same time, in recent decades the interest to the theory of fractional-order differential equations has significantly increased, thanks to applications in various branches of applied mathematics, physics, engineering, biology, economics, etc. (see, for example, monographs [10], [11] papers [12], [13], [14], [15], [16], etc.). the boundary value problems of various types for fractional differential equations and inclusions were considered in the works [17], [18], [19], [20], [21], [22], [23]. in this paper we develop and generalize the results of papers [2] and [3], and we study the controllability problem in banach spaces for various classes of functional inclusions with causal operators with an infinite delay and impulse effects. basing on the topological degree ∗corresponding author: garikpetrosyan@yandex.ru a controllability problem for causal functional inclusions with an infinite delay41 theory for condensing multimaps we prove a global theorem on the existence of trajectories for systems governed by functional inclusions. as an application, we obtain generalizations of existence theorems for solutions of the controllability problem for a first order semilinear functional differential inclusions and a semilinear functional differential inclusions of a fractional order 0 < q < 1. 2. preliminaries 2.1. multivalued maps and measures of noncompactness let x be a metric space and y a normed space. introduce the following notation: p (y ) denotes the collection of all non-empty subsets of y ; pb(y ) denotes the collection of all non-empty and bounded subsets of y ; c(y ) denotes the collection of all non-empty and closed subsets of y ; cv(y ) denotes the collection of all non-empty, closed and convex subsets of y ; k(y ) denotes the collection of all non-empty and compact subsets of y ; kv(y ) denotes the collection of all non-empty, compact and convex subsets of y. let us recall some notations (see, for example, [24], [25]). definition 2.1: a multivalued map (multimap) f : x → p (y ) is said to be upper semicontinuous (u.s.c.) at a point x ∈ x, if for every open set v ⊂ y such that f(x) ⊂ v, there exists a neighborhood u(x) of x such that f(u(x)) ⊂ v. definition 2.2: a multivalued map (multimap) f : x → p (y ) is called closed if its graph gf = {(x, y) : x ∈ x, y ∈ f(x)} is a closed subset of x × y. definition 2.3: a multivalued map (multimap) f : x → p (y ) is called quasicompact if its restriction to each compact subset a ⊂ x is compact. lemma 2.1: ( [24], theorem 1.1.12). if f : x → k (y ) a closed quasicompact multimap, then f is u.s.c. definition 2.4: for a given p ≥ 1, a multifunction g : [0, t ]→ k(y ) is called: • lp–integrable if it admits an lp–bochner integrable selection, i.e., there exists a function g ∈ lp ([0, t ];y ) such that g(t) ∈ g(t) for a.e. t ∈ [0, t ]; • lp–integrably bounded if there exists a function ξ ∈ lp([0, t ]) such that ‖g(t)‖ ≤ ξ(t) for a.e. t ∈ [0, t ]. let e be a banach space lemma 2.2: (see [24], theorem 4.2.1.) let a sequence of functions {ξn} ⊂ l1([0, t ];e) be l1–integrably bounded. suppose that χ({ξn} (t)) ≤ α(t) a.e. t ∈ [0, t ] for all n = 1, 2, ..., where α ∈ l1 +([0, t ]). then for every δ > 0 there exist a compact set kδ ⊂ e, a set mδ ⊂ [0, t ] of a lebesgue measure mδ < δ, and a set of functions gδ ⊂ l1([0, t ];e) with values in kδ, such that for every n ≥ 1 there exists a function bn ∈ gδ for copyright © 2021 assa. adv syst sci appl (2021) 42 m. afanasova, v. obukhovskii, g. petrosyan which ‖ξn(t)− bn(t)‖e ≤ 2α(t) + δ, t ∈ [0, t ] \mδ. moreover, the sequence {bn}may be chosen so that bn ≡ 0 onmδ and this sequence is weakly compact. definition 2.5: let (a,≥) be a partially ordered set. a function β : pb(e)→ a is called the measure of noncompactness (mnc) in e if for each ω ∈ pb(e) we have: β(co ω) = β(ω), where co ω denotes the closure of the convex hull of ω. a measure of noncompactness β is called: 1) monotone if for each ω0,ω1 ∈ pb(e), from ω0 ⊆ ω1 follows β(ω0) ≤ β(ω1). 2) nonsingular, if for each a ∈ e and each ω ∈ pb(e) we have β({a} ∪ ω) = β(ω). if a is a cone in a banach space, the mnc β is called: 3) regular, if β(ω) = 0 is equivalent to the relative compactness of ω ∈ pb(e); 4) real, if a is the set of all real numbers r with the natural ordering. as the example of a real mnc obeying all above properties, we can consider the hausdorff mnc χ(ω): χ(ω) = inf{ε > 0, for which ω has a finite ε-net in e }. as other examples, consider the measures of noncompactness defined in the space of continuous functions c([a, b];e) with values in the banach space e: (1) the modulus of fiber noncompactness: ϕ(ω) = sup t∈[a,b] χe(ω(t)), where χe is the hausdorff mnc in e and ω(t) = {y(t) : y ∈ ω}; (2) the fading modulus of fiber noncompactness: γ(ω) = sup t∈[a,b] e−ltχe(ω(t)), where l > 0 is a given number; (3) the modulus of equicontinuity: modc (ω) = lim δ→0 sup y∈ω max |t1−t2|≤δ ‖y (t1)− y (t2)‖ . these measures of noncompactness satisfy all the above properties, except for the regularity. definition 2.6: a multimap f : x ⊆ e → k(e) is called condensing with respect to a mnc β (or β– condensing) if for each bounded set ω ⊆ x which is not relatively compact, we have: β(f (ω)) 6≥ β(ω). let d ⊂ e a non-empty closed convex subset, v a non-empty bounded open subset of d, β a monotone nonsingular mnc in e and f : v → kv (d) be a u.s.c. β-condensing map copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay43 such that x /∈ f (x) for all x ∈ ∂v , where v and ∂v denote the closure and the boundary of the set v in the relative topology of d. in such a setting, the (relative) topological degree degd ( i−f , v ) of the corresponding vector field i−f , satisfying the standard properties is defined (see, for example, [24], [25]). in particular, the condition degd ( i−f , v ) 6= 0 implies that the fixed points set fixf = {x : x ∈ f(x)} is a nonempty subset of v. application of topological degree theory leads to the following fixed point principles, which will be used in the what follows. theorem 2.1: ( [24], corollary 3.3.1). let m be a convex closed bounded subset of e and f :m→ kv(m) a β–condensing multimap, where β is a monotone nonsingular mnc in e . then the fixed point set fixf is non-empty. theorem 2.2: ( [24], theorem 3.3.4). let v ⊂ d be a bounded open neighborhood of a point a ∈ v and f : v → kv(d) a u.s.c. β-condensing multimap, where β is a monotone nonsingular mnc in e , satisfying the boundary condition x− a /∈ λ(f(x)− a) for all x ∈ ∂v and 0 < λ ≤ 1. then fixf 6= ∅ is a non-empty compact set. 2.2. phase space we will use the axiomatic definition of the phase space b, introduced by j.k. hale and j. kato (see. [26], [27]).the space b we will be considered as a linear topological space of functions defined on (−∞, 0] with values in a banach space e endowed with the seminorm ‖ · ‖b. for all function x : (−∞, t ]→ e, where t > 0, and every t ∈ (−∞, t ], xt is a function from (−∞, 0] to e, defined as xt(θ) = x(t+ θ), θ ∈ (−∞, 0]. we will be assume that b satisfies the following axioms: (b1) if the function x : (−∞;t ]→ e is continuous on [0;t ] and x0 ∈ b, then for each t ∈ [0;t ] : (i) xt ∈ b; (ii) the function t 7→ xt is continuous; (iii) ‖xt‖b ≤ k(t) sup 0≤τ≤t ‖x(τ)‖+h(t)‖x0‖b, where the functions k,h : [0;∞)→ [0;∞) are independent of x, k is strictly positive and continuous and h is locally bounded. (b0) there exists l > 0 such that ‖ψ(0)‖e ≤ l‖ψ‖b, for all ψ ∈ b. notice that under these conditions the space c00 of all continuous functions from (−∞, 0] to e with compact support into phase space b( [27], proposition 1.2.1). in addition, we will assume that the following condition is satisfied: (bc1) if a uniformly bounded sequence {ψn}+∞ n=1 ⊂ c00 converges to a function ψ compactly (i.e. uniformly on each compact subset (−∞, 0]), then ψ ∈ b and lim n→+∞ ‖ψn − ψ‖b = 0. copyright © 2021 assa. adv syst sci appl (2021) 44 m. afanasova, v. obukhovskii, g. petrosyan the condition (bc1) implies that the banach space of bounded continuous functions bc = bc((−∞, 0];e) is continuously embedded into b. more precisely, the following assertion is true. theorem 2.3: ( [27], proposition 7.1.1). (i) bc ⊂ c00, where c00 denote the closure of c00 in b; (ii) if a uniformly bounded sequence {ψn} in bc converges to a function ψ compactly on (−∞, 0], then ψ ∈ b and lim n→+∞ ‖ψn − ψ‖b = 0; (iii) there exists l > 0 such that ‖ψ‖b ≤ l‖ψ‖bc for all ψ ∈ bc. finally, we will assume that the following condition is satisfied: (bc2) if ψ ∈ bc and ‖ψ‖bc 6= 0, then ‖ψ‖b 6= 0. this assumption implies that the space bc, endowed with ‖ · ‖b is a normed space. we will denote it as bc. we may consider the following examples of phase spaces satisfying all the above properties: (1) for γ > 0 let b = cγ be the space of continuous functions ϕ : (−∞; 0]→ e, having a limit lim θ→−∞ eγθϕ(θ) with ‖ϕ‖b = sup −∞<θ≤0 eγθ‖ϕ(θ)‖. (2) (spaces of ”fading memory”) let b = cρ be the space of functions ϕ : (−∞; 0]→ e such that (a) ϕ is continuous on [−r; 0], r > 0; (b) ϕ is lebesgue measurable on (−∞; r) and there exists a nonnegative lebesgue integrable function ρ : (−∞;−r)→ r+ such that ρϕ lebesgue integrable on (−∞; r); moreover, there exists a locally bounded function p : (−∞; 0]→ r+ such that, for all ξ ≤ 0, ρ(ξ + θ) ≤ p (ξ)ρ(θ) a.e. θ ∈ (−∞;−r). then, ‖ϕ‖b = sup −r≤θ≤0 ‖ϕ(θ)‖+ −r∫ −∞ ρ(θ)‖ϕ(θ)‖dθ. a simple example of such a space is given by ρ(θ) = eµθ, µ ∈ r. 2.3. causal multioperators with infinite delay let e be a separable banach space. by lp ([0, t ];e) , 1 ≤ p ≤ ∞, we denote the banach space of all bochner summable functions f : [0, t ]→ e with the usual norm. for each subset n ⊂ lp ([0, t ];e) and τ ∈ (0, t ) we define restriction n on [0, τ ] as n |[0,τ ]= {f |[0,τ ]: f ∈ n}. we split the segment [0, t ] by points 0 < t1 < ... < tm < t,m ≥ 1. for a function c : [0, t ]→ e we will denote c(t+k ) = lim ξ→0+ c(tk + ξ), c(t−k ) = lim ξ→0− c(tk + ξ), copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay45 for 1 ≤ k ≤ m. for a function g : (−∞, t ]→ e, we will assume that its restriction to [0, t ] belongs to the space pc([0, t ];e) of functions z : [0, t ]→ e, continuous on [0, t ] \ {t1, · · · , tm} and such that the left and right limits z(t+k ) and z(t−k ), 1 ≤ k ≤ m, exist and z(t−k ) = z(tk). it is easy to see that the space pc([0, t ];e), endowed with the norm ‖z‖pc = sup t∈[0,t ] ‖g(t)‖e, is a banach space. the classical space of continuous function c([0, t ], e) is its closed subspace. we denote by c((−∞;t ];e) the normed space of piecewise continuous functions x : (−∞;t ]→ e, endowed with the norm ‖x‖c = ‖x0‖bc + ‖x |[0;t ] ‖pc. definition 2.7: a multivalued mapq : c ((−∞, t ];e) ( lp ([0, t ];e) is said to be a causal multioperator, if for each τ ∈ (0, t ) and for every u(·), v(·) ∈ c ((−∞, t ];e) the condition u |(−∞,τ ]= v |(−∞,τ ] implies that q(u) |[0,τ ]= q(v) |[0,τ ] . let us give examples of causal multioperators. example 2.1: we assume that the multimap f : [0, t ]× bc × e → kv (e) satisfies the following conditions: (f1) for each (ψ, φ) ∈ bc × e the multifunction f (·, ψ, φ) : [0, t ]→ kv (e) admits a measurable selection; (f2) for a.e. t ∈ [0, t ] the multifunction f (t, ·, ·) : bc × e → kv (e) is u.s.c.; (f3) there exists a function α ∈ lp+[0, t ], 1 ≤ p ≤ ∞, such that ‖f (t, ψ, φ)‖e := sup{‖z‖e : z ∈ f (t, ψ, φ)} ≤ α(t)(1 + ‖ψ‖bc + ‖φ‖e) for a.e. t ∈ [0, t ] and (ψ, φ) ∈ bc × e. from above conditions (f1)− (f3) and (b1) it follows that the multimap pf : c((−∞;t ];e)→ p (lp([0, t ];e)), given in the following way pf (x) = {f ∈ lp([0, t ];e) : f(t) ∈ f (t, xt, x(t)) a.e. t ∈ [0, t ]} is well defined (see, for example, [24], [25]). it is clear that the multioperator pf is causal. example 2.2: let f : [0, t ]× bc → kv(e) be a multimap satisfying conditions (f1)− (f3) from example 2.1. suppose that {k(t, s) : 0 ≤ s ≤ t ≤ t} is a continuous (with respect to the corresponding norm) family of bounded linear operators in e and m ∈ l1([0, t ];e) is a given function. consider the volterra integral multioperator v : c ((−∞, t ];e) ( l1 ([0, t ];e) defined as v(u)(t) = m(t) + ∫ t 0 k(t, s)f (s, us)ds, i.e., v(u) = {y ∈ l1 ([0, t ];e) : y(t) = m(t) + ∫ t 0 k(t, s)f(s)ds : f ∈ pf (u)}. it is also clear that the multioperator v is causal. copyright © 2021 assa. adv syst sci appl (2021) 46 m. afanasova, v. obukhovskii, g. petrosyan 3. the controllability problem for functional inclusions with the causal operators we will assume that the causal operator q : c ((−∞, t ];e)→ c (lp ([0, t ];e)) satisfies the following conditions: (q1) q is weakly closed in the following sense: conditions {un}∞n=1 ⊂ c ((−∞, t ];e) , {fn}∞n=1 ⊂ lp ([0, t ];e) , 1 ≤ p ≤ ∞, fn ∈ q(un), n ≥ 1, un → u0, fn l1 ⇀ f0 implies f0 ∈ q(u0); (q2) there exists a function α ∈ l∞+ ([0, t ]) such that ‖q (u) (t) ‖e ≤ α (t) (1 + ‖u‖c) , for a.e. t ∈ [0, t ], for all u ∈ c((−∞, t ];e); (q3) there exists a function ω : [0, t ]× r+ → r+ such that (ω1) for all x ∈ r+ : ω(·, x) ∈ lp+([0, t ]), 1 ≤ p ≤ ∞, ; (ω2) for a.e. t ∈ [0, t ] a function ω(t, ·) : r+ → r+ is continuous, nondecreasing and quasihomogeneous in the sense that ω(t, λx) ≤ λω(t, x) for all x ∈ r+ and λ ≥ 0; (ω3) for each bounded set ∆ ⊂ c ((−∞, t ];e) we have χ (q (∆) (t)) ≤ ω ( t, sup s∈[0,t] ϕ (∆s) ) for a.e. t ∈ [0, t ], where the set ∆s = {ys : y ∈ ∆} ⊂ bc and ϕ is the modulus of fiber noncompactness in bc. note that the condition (ω2) means that ω(t, 0) = 0 for a.e.t ∈ [0, t ] and as an example of such a function we can consider ω(t, x) = k(t) · x, where k(·) ∈ lp+([0, t ]). consider a linear operator s : lp([0, t ];e)→ c([0, t ];e), which is causal in the sense that for every τ ∈ (0, t ] and f, g ∈ lp([0, t ];e) condition f(t) = g(t) for a.e. t ∈ [0, τ ] implies (sf) (t) = (sg) (t) for all t ∈ [0, τ ]. following [24], we impose the next conditions on operator s : (s1) for 1 ≤ p <∞ there exist d ≥ 0 such that ‖sf(t)− sg(t)‖pe ≤ d ∫ t 0 ‖f(s)− g(s)‖peds for all f, g ∈ lp([0, t ];e) and 0 ≤ t ≤ t ; if p =∞ there exist d1 ≥ 0 such that ‖sf(t)− sg(t)‖e ≤ d1 ∫ t 0 ‖f(s)− g(s)‖eds for all f, g ∈ l∞([0, t ];e) and 0 ≤ t ≤ t. (s2) for an arbitrary compact set k ⊂ e and a sequence {fn}∞n=1 ⊂ lp ([0, t ];e) , 1 ≤ p ≤ ∞, such that {fn(t)}∞n=1 ⊂ k for all t ∈ [0, t ] the weak convergence fn l1 ⇀ f0 implies sfn → sf0 in c([0, t ];e). also we suppose that s satisfies the relation: (s3) (sf) (0) = 0 for each function f ∈ lp([0, t ];e). copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay47 notice, that condition (s1) implies that the operator s satisfies the lipschitz condition: (s1′) ‖sf − sg‖c ≤ d‖f − g‖l1 . consider following important examples. (i) let a closed linear operator a : d (a) ⊂ e → e be the infinitesimal generata of a c0semigroup {eat}t≥0. the operator l : l1([0, t ];e)→ c([0, t ];e) defined as lf(t) = ∫ t 0 ea(t−s)f(s)ds is a special case of the operator s. note that taking a = 0 we obtain, as a special case, the usual integral operatorli : l1([0, t ];e)→ c([0, t ];e), lif(t) = ∫ t 0 f(s)ds; (ii) let a : d(a)→ e be a closed linear operator e generating a c0semigroup {u(t)}t≥0 . the operator g : lp([0, t ];e)→ c([0, t ];e), p > 1/q, defined as gf(t) = ∫ t 0 (t− s)q−1t (t− s)f(s)ds, 0 < q < 1, where t (t) = q ∫ ∞ 0 θξq(θ)u(tqθ)dθ, ξq(θ) = 1 q θ−1− 1 qψq(θ −1/q), ψq(θ) = 1 π ∞∑ n=1 (−1)n−1θ−qn−1 γ(nq + 1) n! sin(nπq), θ ∈ r+, is a special case of the operator s. lemma 3.1: ( [24], lemma 4.2.1, [12], lemma 3.4). the operators l and g satisfy conditions (s1)− (s3). consider a control system governed by a functional inclusion with causal operatorsq and s, of the following form: y(t) ∈ g(t) ( ψ(0) + σtk 0 such that ‖ikx‖ ≤ n for all x ∈ e. lemma 3.2: (see [12]) the operator functions g and t possess the following properties: 1) for each t ∈ [0, t ], g(t) and t (t) are linear bounded operators, more precisely, for each x ∈ e we have ‖g(t)x‖e ≤m ‖x‖e , ‖t (t)x‖e ≤ qm γ(1 + q) ‖x‖e , where m = supt≥0 ‖u(t)‖ 2) the operator functions g(·) and t (·) are strongly continuous, i.e., functions t ∈ [0, t ]→ g(t)x and t ∈ [0, t ]→ t (t)x are continuous for each x ∈ e. suppose ψ ∈ bc is a given function. for a function y ∈ pc([0, t ];e) such that y(0) = ψ(0) we define the function y[ψ] ∈ c((−∞, t ];e) as y[ψ](t) = { ψ(t), −∞ ≤ t < 0, y(t), 0 ≤ t ≤ t . we denote by d the closed convex subset of pc([0, t ];e), consisting of all functions y satisfying the condition y(0) = ψ(0). definition 3.1: a function y ∈ c((−∞, t ];e) is called a mild solution of problem (3.1)-(3.3), if it satisfies conditions: (1) the function y|[0,t ] ∈ d and satisfies inclusion (3.1); (2) y(t) = ψ(t), for t ∈ (−∞, 0]; (3) iky(tk) = y(t+k )− y(tk), k = 1, ...,m. now, the controllability problem which we solve in this paper may be formulated in the following way: for a given initial function ψ ∈ bc and a given x1 ∈ e we consider the existence of a mild solution y of problem (3.1)-(3.3) and a control u such that y(t) = ψ(t), t ∈ (−∞, 0], and y(t ) = x1. (3.4) copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay49 the pair (y, u) consisting of an integral solution y of problem (3.1)-(3.3) and a control u ∈ l∞([0, t ];u) will be called the solution of controllability problem (3.1)-(3.4). consider the multioperator γ : d( d defined as γ(y) = {x ∈ d : x(t) = g(t) ( ψ(0) + σtk 0 we take δ ∈ (0, ε), such that for every m ⊂ [0, t ], with the measure meas(m) < δ, we have: ∫ m |υ(s)|p < ε, for 1 ≤ p <∞, and respectively for p =∞ : ∫ m |υ(s)| < ε. taking mδ and bn corresponding to {fn} from lemma 2.2, we, by using property (s1), obtain that the sequence {s(bn)} is relatively compact in c([0, t ];e). let 1 ≤ p <∞, then the following estimates hold: ‖s(fn)(t)− s(bn)(t)‖pe ≤ d ∫ t 0 ‖fn(s)− bn(s)‖pe ds ≤ d ∫ [0,t]\mδ ‖fn(s)− bn(s)‖pe ds+d ∫ [0,t]∩mδ ‖fn(s)‖pe ds ≤ d ∫ [0,t]\mδ [2υ(s)− δ]p ds+d ∫ mδ |υ(s)|p ds ≤ d ∫ t 0 |2υ(s) + ε|p ds+ εd ≤ d ∫ t 0 2p |2υ(s)|p ds+d ∫ t 0 2pεpds+ εd ≤ 4pd ∫ t 0 υp(s)ds+ 2pεptd + εd. copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay51 therefore, the relatively compact set sgδ(t) is a ( 4pd ∫ t 0 υp(s)ds+ 2pεptd + εd )1/p net for the set {s(fn)(t)} . since ε > 0 is arbitrary, we obtain the conclusion of the lemma for the case 1 ≤ p <∞. if p =∞, then the following estimates hold: ‖s(fn)(t)− s(bn)(t)‖e ≤ d1 ∫ t 0 ‖fn(s)− bn(s)‖e ds ≤ d1 ∫ [0,t]\mδ ‖fn(s)− bn(s)‖e ds+d1 ∫ [0,t]∩mδ ‖fn(s)‖e ds ≤ d1 ∫ [0,t]\mδ |2υ(s)− δ|ds+d1 ∫ mδ |υ(s)| ds ≤ 2d1 ∫ t 0 υ(s)ds+ εtd1 + εd1. thus, the relatively compact set sgδ(t) is a 2d1 ∫ t 0 υ(s)ds+ εd1(t + 1) net for the set {s(fn)(t)} . since ε > 0 is arbitrary, we obtain the conclusion of the lemma also for the case p =∞. let m1, m2 be positive constants, such that ‖b‖ ≤m1, ∥∥w−1 ∥∥ ≤m2. (3.5) consider the measure of noncompactness ν in the space pc([0, t ];e) with values in the cone r2 +. on a bounded subset of ω ⊂ pc([0, t ];e) we define the values of ν as follows: ν(ω) = (γ (ω) ,modc (ω)) , where modc is the modulus of equicontinuity, γ is the fading modulus of fiber noncompactness γ(ω) = sup t∈[0,t ] e−ltχ(ω(t)). the constant l > 0 is chosen so that max{q1, q2} < 1, where q1 = sup t∈[0,t ] ( 4d1/p ( 1 + 4m1σd 1/pt 1/p ) ∫ t 0 e−lp(t−s)ωp (s, 1) ds )1/p , q2 = sup t∈[0,t ] ( 2d1 (1 + 2m1σd1t ) ∫ t 0 e−l(t−s)ω (s, 1) ds ) , where the constants d,d1 are from condition (s1), ω is a function from condition (q3). it is easy to see that the mnc ν is monotone, nonsingular, and algebraically semi-additive. it follows from the arzela–ascoli theorem that it is also regular. theorem 3.2: let a causal multioperatorq : c((−∞, t ];e) ( lp ([0, t ];e) satisfy conditions (q2) and (q3) and for a causal operator s : lp ([0, t ];e)→ c ([0, t ];e) the conditions (s1)–(s3) be satisfied. then, under conditions (i1), (i2), (w ) the multioperator γ is ν-condensing. copyright © 2021 assa. adv syst sci appl (2021) 52 m. afanasova, v. obukhovskii, g. petrosyan proof by lemma 3.2 and conditions (i1), (i2), it is suffices to prove the assertion of the theorem for the multioperator s ◦ q+ s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ s ◦ q ) . let ω ⊂ d be a bounded set such that ν ( s ◦ q (ω[ψ]) + s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ s ◦ q(ω[ψ]) )) ≥ ν (ω) . (3.6) let us show that the set ω is relatively compact. inequality (3.6) means that γ({s ◦ q (ω[ψ]) + s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ s ◦ q(ω[ψ]) ) }) ≥ γ(ω). (3.7) applying the condition (q3) and by using the properties of the function ω, we obtain for a.e. t ∈ [0, t ] χ ({f(t) : f ∈ q (ω[ψ])}) ≤ ω ( t, sup s∈[0,t] ϕ ({y[ψ]s : y ∈ ω}) ) = ω ( t, ϕ ( {y|[0,t] : y ∈ ω} )) = ω ( t, elte−ltϕ ( {y|[0,t] : y ∈ ω} )) ≤ ω ( t, eltγ ( {y|[0,t] : y ∈ ω} )) ≤ ω ( t, eltγ (ω) ) ≤ ω ( t, elt ) · γ (ω) . at first, we consider the case 1 ≤ p <∞. by lemma 3.5 we have for each t ∈ [0, t ] : χ ({sf(t) : f ∈ q (ω[ψ])}) ≤ ( 4pd ∫ t 0 ωp ( s, els ) ds · γp (ω) )1/p ≤ 4d1/p (∫ t 0 eplsωp (s, 1) ds )1/p · γ (ω) . further, χ ( {bw−1 ( x1 − g(t )ψ(0)− ζ ◦ sf(t) : f ∈ q (ω[ψ])} ) ≤ m1σχ ({ζ ◦ sf(t) : f ∈ q (ω[ψ])}) = m1σχ ({sf(t ) : f ∈ q (ω[ψ])}) ≤m1σ ( 4pd ∫ t 0 eplsωp (s, 1) ds · γp (ω) )1/p = 4m1σd 1/p (∫ t 0 eplsωp (s, 1) ds )1/p · γ (ω) , where σ = supt∈[0,t ] σ(t). using lemma 3.5 again, we have for each t ∈ [0, t ] : χ ( s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ sf(t) : f ∈ q (ω[ψ]) ) ≤( 4pd ∫ t 0 4pmp 1σ pd (∫ t 0 eplsωp (s, 1) ds ) dτ · γp (ω) )1/p ≤ copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay53 16d2/pm1σt 1/p (∫ t 0 eplsωp (s, 1) ds )1/p · γ (ω) . inequality (3.7) and the last inequality imply the following γ(ω) ≤ sup t∈[0,t ] ( 4d1/p ( 1 + 4m1σd 1/pt 1/p ) ∫ t 0 e−lp(t−s)ωp (s, 1) ds )1/p γ (ω) = q1 · γ (ω) , therefore γ (ω) = 0, thus ϕ (ω[ψ]t) = 0 for all t ∈ [0, t ]. let us turn to the case p =∞. by lemma 3.5 we have for each t ∈ [0, t ] : χ ({sf(t) : f ∈ q (ω[ψ])}) ≤ 2d1 ∫ t 0 ω ( s, els ) ds · γ (ω) ≤ 2d1 ∫ t 0 elsω (s, 1) ds · γ (ω) ; χ ( {bw−1 ( x1 − g(t )ψ(0)− ζ ◦ sf(t) : f ∈ q (ω[ψ])} ) ≤ m1σχ ({ζ ◦ sf(t) : f ∈ q (ω[ψ])}) = m1σχ ({sf(t ) : f ∈ q (ω[ψ])}) ≤m1σ2d1 ∫ t 0 elsω (s, 1) ds · γ (ω) = 2m1σd1 ∫ t 0 elsω (s, 1) ds · γ (ω) , where σ = supt∈[0,t ] σ(t). using lemma 3.5, we have for each t ∈ [0, t ] : χ ( s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ sf(t) : f ∈ q (ω[ψ]) ) ≤ 2d1 ∫ t 0 2m1σd1 (∫ t 0 elsω (s, 1) ds ) dτ · γ (ω) ≤ 4d1m1σt ∫ t 0 elsω (s, 1) ds · γ (ω) . inequality (3.7) and the last inequality imply the following γ(ω) ≤ sup t∈[0,t ] ( 2d1 (1 + 2m1σd1t ) ∫ t 0 e−l(t−s)ω (s, 1) ds ) γ (ω) = q2 · γ (ω) , therefore γ (ω) = 0, thus ϕ (ω[ψ]t) = 0 for each t ∈ [0, t ]. now we will show that the set ω is equicontinuous. we take sequences {yn}∞n=1 ⊂ ω, n ≥ 1 and {fn}∞n=1, fn ∈ q(yn[ψ]). from conditions (q2) and (q3) it follows that the sequence {fn}∞n=1 is lp-semicompact, and therefore by lemma 3.4 the sequence {sfn}∞n=1 copyright © 2021 assa. adv syst sci appl (2021) 54 m. afanasova, v. obukhovskii, g. petrosyan is relatively compact. hence modc({sfn}∞n=1) = 0. from the conditions that the operatorsb,w−1, ζ are bounded and linear, we can conclude that modc ( s ◦bw−1(x1 − g(t )ψ(0)− ζ{sfn}∞n=1) ) = 0. thus ν ( {s ◦ q (ω[ψ]) + s ◦bw−1 ( x1 − g(t )ψ(0)− ζ ◦ s ◦ q(ω[ψ]) ) } ) = (0, 0), but then it follows from the inequality (3.6) that ν(ω) = (0, 0), and the last expression yields that the set ω is relatively compact. to prove the main theorem of the paper, we need the following statements, known as the gronwall bellman lemma and the generalized gronwall bellman lemma. lemma 3.6: let v(t) and f(t) be nonnegative continuous functions on the segment [a, b], moreover v(t) ≤ c+ ∫ t a f(s)v(s)ds, t ∈ [a, b], where c is a positive constant. then for each t ∈ [a, b] the inequality v(t) ≤ ce ∫ t a f(s)ds, holds. lemma 3.7: let h(t), u(t) and v(t) be nonnegative functions integrable on [a, b] satisfying the inequality: v(t) ≤ u(t) + ∫ t a h(s)v(s)ds, t ∈ [a, b]. then the following inequality holds: v(t) ≤ u(t) + ∫ t a e ∫ t a h(θ)dθh(s)u(s)ds, t ∈ [a, b]. theorem 3.3: let a causal multioperator q : c((−∞, t ];e)→ cv(lp([0, t ];e)), 1 ≤ p ≤ ∞, satisfy conditions (q1)–(q3) and a linear causal operator s : lp([0, t ];e)→ c([0, t ];e) satisfy conditions (s1)–(s3). then, under conditions (i1), (i2), (w ) the set σψ of all solutions to problem (3.1)-(3.4) is a non-empty compact set. proof let us show that the set of all solutions y ∈ d of a one-parameter inclusion y ∈ λγ(y), λ ∈ [0, 1], (3.8) is a priori bounded. we divide the proof into three cases: p = 1, 1 < p <∞, p =∞. copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay55 let p = 1, if y ∈ d satisfies condition (3.8), then for each t ∈ [0, t ], using assumptions (b0), (s1), (i1), (i2) and (3.5), we have the following estimates: ‖y(t)‖e ≤ ‖g(t) ( ψ(0) + σtk 0 such that ‖ikx‖ ≤ n for all x ∈ e. notice that when q = 1 : g(t) = eat, t (t) = eat, therefore, in accordance with [24], a function y ∈ c((−∞, t ];e), is a mild solution of problem (4.10)-(4.12), if it can be represented in the form: y(t) =  eat ( ψ(0) + σtk 1/q satisfies condition (q1) can be verified as in the paper [28]. conditions (q2) and (q3) for pf follow from (f3) and (f4), respectively. taking into account lemma 3.1, we can consider the relation (4.13) as a special case of functional inclusion (3.1) with q = pf , and s = g is the cauchy type operator. as a direct consequence of theorem 3.3, we obtain the following result. theorem 4.2: suppose that conditions (a), (f1)–(f4), (i1), (i2), (w ) hold. then the set of solutions to problem (4.13)-(4.15), (3.4) is a non-empty compact subset of the space c((−∞, t ];e). acknowledgements the work of the first author was supported by the state contract of the russian ministry of education as part of the state task (contract fzgf-2020-0009). the work of the second author was supported by rfbr (project number 20-51-15003). the work of the third author was supported by rfbr (project number 19-31-60011). references 1. lupulescu, v. (2008). causal functional differential equations in banach spaces, nonlinear anal., 69(12), 4787-4795. 2. obukhovskii, v., zecca, p. (2011). on certain classes of functional inclusions with causal operators in banach spaces, nonlinear anal., 74(8), 2765-2777. 3. kulmanakova, m.m., ulyanova, e.l. (2019). o razreshimosti kauzalnih funkcionalnih vkluchenii s beskonechnim zapazdivaniem [on the solvability of causal functional inclusions with infinite delay]. tambov university reports. series: natural and technical sciences, 24(127), 293-315, [in russian]. copyright © 2021 assa. adv syst sci appl (2021) a controllability problem for causal functional inclusions with an infinite delay61 4. tikhonov, a.n. (1938). o funkcionalnih uravneniyah tipa volterra i ih prilozheniya v nekotorih zadachah matematicheskoi fiziki [on functional equations of volterra type and their applications to some problems of mathematical physics]. bulletin of moscow university, 1(8), 1-25, [in russian]. 5. corduneanu, c. (2002). functional equations with causal operators. stability and control: theory, methods and applications. london, taylor and francis. 6. bulgakov, a.i., maximov, v.p. (1981). funkcionalnie i funkcionalno-differencialnie vkluchenia s operatorami volterra [functional and functional differential inclusions with volterra operators]. differential equations, 17(8), 1362-1374, [in russian]. 7. drici, z., mcrae, f.a., vasundhara devi, j. (2005). differential equations with causal operators in a banach space, nonlinear anal., 62(2), 301-313. 8. drici, z., mcrae, f.a., vasundhara devi, j. (2006). monotone iterative technique for periodic boundary value problems with causal operators, nonlinear anal., 64(6), 12711277. 9. jankowski, t. (2008). boundary value problems with causal operators, nonlinear anal., 68(12), 3625-3632. 10. kilbas, a.a., srivastava, h.m., trujillo, j.j. (2006). theory and applications of fractional differential equations. amsterdam, north-holland mathematics studies, elsevier science b.v. 11. podlubny, i. (1999). fractional differential equations. san diego, academic press. 12. afanasova, m., liou, y. ch., obukhoskii, v., petrosyan, g. (2019). on controllability for a system governed by a fractional-order semilinear functional differential inclusion in a banach space, journal of nonlinear and convex analysis, 20(9), 1919-1935. 13. appell, j., lopez, b., sadarangani, k. (2018). existence and uniqueness of solutions for a nonlinear fractional initial value problem involving caputo derivatives, j. nonlinear var. anal., 2, 25-33. 14. gomoyunov, m.i. (2018). fractional derivatives of convex lyapunov functions and control problems in fractional order systems. fract. calc. appl. anal., 21, 1238-1261. 15. kamenskii, m., obukhoskii, v., petrosyan, g., yao, j.-c. (2019). existence and approximation of solutions to nonlocal boundary value problems for fractional differential inclusions, fixed point theory and applications, 30(2). 16. mainardi, f., rionero, s., ruggeri, t. (1994). on the initial value problem for the fractional diffusion-wave equation. in waves and stability in continuous media (pp. 246-251). singapore, world scientific. 17. agarwal, r.p., ahmad, b. (2011). existence theory for anti-periodic boundary value problems of fractional differential equations and inclusions, comput. math. appl., 62, 1200-1214. 18. gomoyunov, m.i. (2019). approximation of fractional order conflict-controlled systems, prog. fract. differ. appl., 5, 143-155. 19. kamenskii, m., obukhoskii, v., petrosyan, g., yao, j.-c. (2019). on a periodic boundary value problem for a fractional order semilinear functional differential inclusions in a banach space, mathematics, 7(12), 5-19. 20. kamenskii, m., obukhoskii, v., petrosyan, g., yao, j.-c. (2021). on the existence of a unique solution for a class of fractional differential inclusions in a hilbert space, mathematics, 9(2), 136-154. 21. kamenskii, m.i., petrosyan, g.g., wen, c.-f. (2021). an existence result for a periodic boundary value problem of fractional semilinear differential equations in a banach space, journal of nonlinear and variational analysis, 5(1), 155-177. 22. petrosyan, g. (2021). antiperiodic boundary value problem for a semilinear differential equation of fractional order, the bulletin of irkutsk state university. series: mathematics, 34, 51-66. 23. petrosyan, g. (2020). on antiperiodic boundary value problem for a semilinear differential inclusion of fractional order with a deviating argument in a banach space, copyright © 2021 assa. adv syst sci appl (2021) 62 m. afanasova, v. obukhovskii, g. petrosyan ufa mathematical journal, 12(3), 69-80. 24. kamenskii, m., obukhovskii, v., zecca, p. (2001). condensing multivalued maps and semilinear differential inclusions in banach spaces. berlin-new york, walter de gruyter. 25. obukhovskii, v., gelman, b. (2020). multivalued maps and differential inclusions. elements of theory and applications. singapore, world scientific. 26. hale, j.k., kato, j. (1978). phase space for retarded equations with infinite delay, funkc. ekvac., 21, 11-41. 27. hito, y., murakami, s., naito, t. (1991). functional differential equations with infinite delay. lecture notes in mathematics. berlin-heidelberg-new york, springer-verlag. 28. petrosyan, g.g. (2015). teorema o slaboi zamknutosti superpozicionnogo multioperatora [a theorem on the weak closure of superposition multioperator], tambov university reports. series: natural and technical sciences, 20(5), 1355-1358, [in russian]. copyright © 2021 assa. adv syst sci appl (2021) introduction preliminaries multivalued maps and measures of noncompactness phase space causal multioperators with infinite delay the controllability problem for functional inclusions with the causal operators controllability problems for semilinear differential inclusions with a delay and impulse effects controllability problems for a first order semilinear functional differential inclusions controllability problems for fractional semilinear functional differential inclusions microsoft word 20 donglin li--preliminary design on automatic production line of high frequency induction welding light gauge advances in systems science and applications (2011), vol.11, no.3-4 355-361 issn 1078-6236 international institute for general systems studies, inc preliminary design on automatic production line of high-frequency induction welding light gauge h-beam donglin li school of mechanical engineering, hubei univ. of technology, wuhan 430068, china abstract:the automatic production line of high-frequency induction welding h-beam is preliminarily designed, which is applied to produce the h-beam of the web plate with the height less than 100mm. in the first place, the technological process for the whole production line is framed based on the principle of the hf induction welding. subsequently, the scheme design is conducted by dividing the production line into pre-welding processing, hf induction welding and post-welding processing. finally, the software solidworks is applied to design the virtual prototype which consists of the feature modeling of the parts and the virtual assembly of the machine, and carry out the motion simulation in order to check the interference and movement coordination of mechanisms. process provides an effective basis and good guidance for the succeeding design of the production line. keywords: high-frequency (hf) induction welding h-beam quality control web plate wing plate virtual prototype 1.introduction with the rapid development of steel structure, the demand for h-beam is growing, especially for the hf induction welding h-beam of the web plate with the height less than 100mm which is now being widely applied because of its economic and rational cross-section and its highly accurate size in the construction, bridges, machinery, shipbuilding, light industry and other industries. the small size of the cross-section of this type of hf welding h-beam, however, usually leads to series of problems such as difficulties in adjusting the welding contact and controlling the welding quality as well as the large welding deformation[1,2]. so the automated production line technology for this specification has always remained the focus of the research and development at home and abroad. this paper will probe into the preliminary design of automatic production line of high-frequency induction welding h-beam of the web plate with the height less than 100mm. in the first place, the technological process for the whole production line is framed based on the principle of the hf induction welding. and then, the scheme design is conducted with the view of improving the welding quality. finally, the software solidworks is applied to design the virtual prototype of the whole production line and carry out the motion simulation, calculate out the smallest traction of the automatic production line. this process provides important data reference for the practical design. 2.the method of hf induction welding h-beam hf induction welding is the welding achieved by rapidly heating the steel through the use of skin and proximity effects of the hf current. when welding the h-beam, first of all, the upper and lower wings will form a v-angle[3,4]. and then, supply power to the wings and the web through electrode to form the back-and-forth loops and liquid beam at the joining point. with the forward movements of the work piece, the beam will continuously extrude the liquid metal and oxide from the edge under the compression of the rolling. the remaining metal will keep in close contact with each other, producing plastic deformation and re-crystallization and resulting 356 li: preliminary design on automatic production line of high-frequency induction welding light gauge h-beam in the solid welding line. hf induction welding method of the h-beam diagram is shown in figure 1. fig.1 the schematic diagram of hf induction welding method 3.the technological process the technological process of the entire automatic production lines is shown in figure 2. production line machines include the flattening machine, angle-controlled machine, hf induction welding machine, deburring machine, correcting machine, traction engine and others, totaling to 28 sets of unit constructions. by the application of these unit constructions, many functions such as sheet flattening, upsetting web plate, hf welding, deburring, post-welding cooling, profiles correcting are expected to be achieved. 4.the scheme design of the production line in order to effectively control the product quality, the design of the entire production line divides into pre-welding processing, hf induction welding, post-welding processing. 4.1 pre-welding processing 4.1.1 loading material, decoiling and clipping the raw materials of hf induction welding light gauge h-beam are hot-rolled strip steel coil[5], so the equipment on the production line in the first place is the loading and decoiling machine of strip steel. the decoiled strip steel will then enter into the process of clipping during which the strip steel will undergo cutting according to the required size for wings and web plate. 4.1.2 loop storage, cutting and butt-welding for uninterrupted production, loops and the equipment of cutting and butt-welding are required on the production line. loops are used to store strip steel, the equipment of cutting and butt-welding to cut the ends of the strip steel. and then the two volumes of steel will be welded together smoothly, in order to ensure the non-stop production. 4.1.3 sheet flattening the decoiled strip steel usually has certain deficiencies such as buckling, bending arc, wave shape and so on. therefore, the sheet flattening is very necessary. flattening is being carried out on the roll flattening machine equipped with the interlaced top and bottom working rollers. during the process of flattening, the clearance between the top and bottom working rollers should be adjusted a little less than the thickness of the steel plate being flattened. in addition, the intake side must be ensured to undergo complete plastic deformation, and the delivery side must undergo complete elastic deformation. the steel will undergo a couple of times of forward and backward alternating bending and finally become flat and smooth. 4.1.4 upsetting of the web plate relying on the self-melting welding, the deposition rate of the hf induction welding h-beam can only reach 85 %~90 %, and the width of the weld seam can only be the 85% of the web plate. the strength of the weld seam cannot exceed the strength of the base metal. in order advances in systems science and applications (2011), vol.11, no.3-4 357 to expand the width of the weld seam equal to the width of the web plate, it is necessary to upset both sides of the web plate by the use of rolling mill before welding in order to expand the area of the ends by 30%. by the use of the u-shaped groove within the rollers, the rolling mill upsets the web plate through the extruding force produced from between the top and bottom rollers. fig.2 the technological process 4.2. hf induction welding 4.2.1 control of the vee angle in order to make an effective use of the skin and proximity effects of the hf current, it is required to set an angle control machine in the front of a hf induction welding machine to form a vee angle by the upper and lower wings and web plate. the most crucial point in the design of the angle control machine is to control the size of the angle because the size has a large impact on the stability of the hf flashing process and the welding quality as well as the welding efficiency. in many cases, the vee angle is required to be less 10°. the size of the vee angle needs the succeeding study. 4.2.2 the best location of the welding contactor during the process of welding, the two contactors respectively belonging to the upper and lower hf induction welding machine heat the web plate and wings separately, as shown in figure 1. the position of the contactors determines the length of the heating period, and loading material decoiling and clipping cutting and butt-welding loop storage sheet flattening web plate upsetting controlling the vee angle hf induction welding deburring post-welding cooling correcting the welding deformation sizing and shearing stacking upper and lower wing plate web plate loading material decoiling and clipping cutting and butt-welding loop storage sheet flattening 358 li: preliminary design on automatic production line of high-frequency induction welding light gauge h-beam produces a large impact on the temperature field and stress field of the h-beam. in order to raise the efficiency of welding, contactors should be located close to the squeeze roller as much as possible. but as for the h-beam of the web plate with the height less than 100mm, the design and location of welding contactor have always remained a difficulty due to the limitation of spacing between the web plates. 4.2.3 extrusion between the welded joints since welding is completed under the extrusion pressure when the sheet material is heated up to the fusing temperature, it is required to set up a squeeze roller above the wings where the sheet material is welded. extrusion force has a strong impact on the quality of the weld seam. in production, the extrusion force is replaced by the extrusion quantity of the plates for welding which is adjusted and measured by alternating the gap of the squeeze rollers. on the premise of the welding quality to be guaranteed, the extrusion quantity should be reduced as much as possible in order to reduce the metal consumption. 4.3 post-welding processing 4.3.1 deburring it is inevitable to produce the unevenly scars in the four weld seams. deburring is a requirement because the weld scars not merely influence the appearance of the weldment, but also leads to the uneven scatter of the stress in the weld seam. two sets of deburring machine are arranged in symmetry on the production line, with a milling cutter and planning tool loaded respectively on each of the deburring machine. the first step is to mill the two weld seams in order to get rid of some larger scars. and then, some minor welding scars and burr after milling will be removed by the use of planning tool. in this way, the four weld seams will become even and smooth, which improves the quality of the h-beam’s appearance to a large extent. 4.3.2 post-welding cooling post-welding cooling is not only for decreasing temperature before correcting profiles, but also for the consideration of the changes in the structure property of the weld seams and the production of the residual stress. the post-welding processing of the hf induction welding h-beam generally adopts the normal water-cooling, with the control of the cooling rate and cooling temperature at the ordinary level. in order to produce the high quality h-beam, our design adopts the air-cooling craft. during the process of water-cooling, we increase the spray equipment and lengthen the period of air-cooling properly. 4.3.3 correcting profiles due to the uneven scatter of the temperature field, the steel in the post-welding will endure large deformation, especially regarding the light gauge h-beam, the condition is even worsening. failure in the control and correction of the deformation will have a direct impact on the erection and connection, operational life span as well as its bearing capacity. the welding deformation of the h-beam lies mainly in the deformation of the angle of the wing plate. this condition requires the symmetric arrangement of two sets of correction machine on the production line. each correction machine presses both sides of the wings through two sets of upper and lower rollers with conicity in order to produce the plastic deformation of the reversed direction during the process of continuous feed-in, completing the continuous correction of the h-beam steel wing plates. 4.4 design of the traction machine the traction machine is the main driving force on the production line because it mainly depends on the traction force of the traction machine to facilitate the whole continuous and stable process from flattening, upsetting, vee angle controlling, to hf induction welding, deburring, correction and so on. in the production line are set up many direct-current and adjustable speed traction machine, which, using the principle of loops control, through the differential speed and the swing of loops, conducts the micro tension control over the steel advances in systems science and applications (2011), vol.11, no.3-4 359 material and then facilitate the smooth going of the whole process, providing guarantee to the welding quality, size and shape tolerance of the h-beam. 5.design of the virtual prototype after determining the technological process and the design scheme of the production line, we take advantage of the software solidworks’ strong three-dimensional modeling function to carry out the virtual prototype of the production line, and conduct the motion simulation to test and verity the coordination of the movements of all the separate parts[6,7], providing the important design parameters for the concrete design of the framework. 5.1 modeling of the parts’ parametric features the design of the parts by the solidworks mainly depends on the basic operations such as drawing, rotation, scanning and so on to set up the three-dimensional solid model. as for the complicated parts, it is essential to choose correct way to generate because the incorrect choices are not only inefficient, but also cannot generate solid model at all in certain cases. take the modeling of the deburring machine as an example, due to its symmetric structure, half of the model can be built at first, and then another half can be completed through mirror. when constructing the half modeling, first, choose the sketch plane to enter the sketch mode and draw out the sketch of the framework. then, generate the three-dimensional solid entities using the drawing method. the completed framework model of the deburring machine is shown in figure 3. fig.3 the framework model of the deburring machine 5.2 virtual assembly after modeling the three-dimensional solid entities of the separate part, the virtual assembly is required for each of these parts. in the assembly module of the solidworks, it is workable to assemble these separate three-dimensional solid entities into a working machine, and check whether or not there is interference between different parts and whether or not the assembling entities move according to the requirements of design[8,9]. take the deburring machine as an example, first, insert the machine framework which the system defaults as a built-in fitment. then, insert the model of other parts and choose the location of each separate part. finally, select appropriate assembly constraint types to complete the location of the parts model. the deburring machine model after virtual assembly is shown in figure 4 (in order to achieve the true-to-life visual effect, h-beam is fitted together into the deburring machine). fig.4 the assembling model of the deburring machine 5.3 the motion simulation of the machine this paper chooses the cosmos motion in solidworks to carry out motion simulation. before simulation, it is necessary to enter the assembly environment to input some necessary 360 li: preliminary design on automatic production line of high-frequency induction welding light gauge h-beam parameters beforehand in order to confirm the moving parts, the fixing parts, the kinematics pair, and install the motion environment and simulation options. the motion simulation of the deburring machine is shown in figure 5. the highlighted parts shown in the picture are the kinematics pair between different machines. according to the requirements of the h-beam production line, with the given horizontal velocity, certain parameters such as the mean momentum and the average torque and others will be attained after the automated simulation calculation. fig.5 the motion simulation of the deburring machine 5.4 the motion simulation of the production line in order to get a clearer and more direct-viewing glimpse of the whole production motion process, we carry out the motion simulation to the entire production line. considering the fact that the entire production line is too bulky for the software to conduct a one-off calculation, we divide the production line into two stages: the pre-welded process of the hf induction welding which consists of the sheet material flattening machine, web plate upsetting machine and the angle control machine; the post-welding process of the hf induction welding which consists of the deburring machine and the correction machine. fig.6 the simulation model of the pre-welding processing. fig.7 the simulation model of the post-welding processing. the simulation steps for the entire production line are almost the same as that of the single machine. the simulation model of the pre-welding processing is shown in figure 6. the simulation model of the post-welding processing is shown in figure 7. after the process of simulation, the momentum and the torque of each moving parts can be accessible from the data-process module. for example, if the given horizontal velocity of the h-beam is 30m/min, the average traction torque of the h-beam will be 450nm based on the calculated mean momentum and average torque. 6.conclusion the automatic production line of the hf induction welding h-beam has the advantage of high productive efficiency, stable quality and the convenience to alternate the varieties and advances in systems science and applications (2011), vol.11, no.3-4 361 specifications and so on. this paper has delved into the preliminary design of automatic production line of hf induction welding h-beam with the web plate’s height less than 100mm. in the first place, the technological process of the production system is framed based on the principle of hf induction welding. subsequently, the scheme design is conducted with the view of improving the welding quality. finally, the software solidworks is applied to design the virtual prototype of the whole production line and carry out the motion simulation. all the preliminary design will provide a significant guidance for the succeeding engineering design in practice. how to promote the welding quality and control deformation, however, still remains the largest problem in the automatic production line of hf induction welding light gauge h-beam. in this paper, such measures as steel material flattening, web plate upsetting, profiles correcting are adopted to guarantee the quality; nevertheless, the best location of the welding contactor and the size of the welding vee angle still remain the focus of the succeeding research and design. 7.acknowledgements the emphasis scientific and technological key project of wu han city fund the project. (contract no. 030877). references [1] scott p. f. and smith w. “a study of the key parameters of high frequency welding”. tube china ’95. (1995) 161-181. [2] scott p. f. “the effects of frequency in high frequency welding”. transactions of tube 2000 toronto, ita conference. (2000) 74-80. [3] asperheim j. i., markegaard l., induksjon e. and lombard p. “temperature distribution in the weld vee cross section”. tube international. 17(70) (1998) 563-567. [4] asperheim j. i. and grande b. “temperature evaluation of weld vee geometry and performance”. tube international. 19(110) (2000) 497-502. [5] wesling v., schram a. and rekersdrees t. “high-frequency welding of martensitic hot-rolled strip”. welding research abroad. (1998) 145-162. [6] mao-peng chen and wei-fang. “automatic modeling for modular fixture components based on solidworks”. journal of south china university of technology (natural science). (2005) 56-59. [7] c.k.chua, s.h.teh and r.k.l.gay. “rapid prototyping versus virtual prototyping in product design and manufacturing”. the international journal of advanced manufacturing technology. 7(1999) 597-603. [8] yang fuqin, chang degong and wang xianlun. “dynamic simulation of the tripod sliding universal joint based on virtual prototype”. 2008 asia simulation conference 7th international conference on system simulation and scientific computing, icsc 2008. (2008) 711-714. [9] li feng, li shuguang, zheng lei and lou li. “dynamic virtual prototype of soldier system”. proceedings of the 27th chinese control conference. (2008) 329-333. adv syst sci appl 2019; 01; 61-74 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/632 identification of nonlinear systems having hard function adil brouri1*, tarik rabyi1, abdelmalek ouannou1 1) ensam, imsm, l2mc lab, moulay ismail university sciences, meknes, morocco e-mail: a.brouri@ensam-umi.ac.ma received august 14, 2018; revised march 18, 2019; published april 15, 2019 abstract: in this work, an identification approach of nonlinear systems is studied. presently, the nonlinear system can be described by wiener-hammerstein model. this latter is composed of a nonlinearity surrounded by two linear blocks. the linear blocks can be nonparametric. the nonlinear element is allowed to be discontinuous, specifically it can be of hard element (e.g. preload, coulomb friction, dead-zone). roughly, it is very difficult to model these types of nonlinearities by orthogonal decompositions (e.g. polynomial). keywords: nonlinear systems, linear system, static and dynamic systems, control system, system identification, discontinuous nonlinearity, hard element. 1. introduction nonlinear systems exist widely in industry and science applications [2,5-7,13], among which the wiener-hammerstein model (fig. 1.1) is one of the most typical cases [1,3,9,15-16]. the wiener-hammerstein models consist of a nonlinear block surrounded by two linear elements (fig.1.1). the identification problem have been paid considerable attention due to their benefits such as control [7,11-12]. the wiener-hammerstein like systems are used in a wide range of applications such as identification of skeletal muscle [1]. note that, wiener and hammerstein nonlinear systems can be modeled by wiener-hammerstein systems. the solution of this problem identification can be dealt using several method. the available methods have been developed following three main approaches i.e. iterative nonlinear optimization procedures [19]; stochastic methods e.g. [3,21]; in [27], an approach based on the standard svm (support vector machines) for regression was presented. the quite poor results obtained in that work highlighted some of the limitations of the method. in particular, only a nfir (nonlinear finite impulse response) model structure was taken into account, which did not perform well since the considered system has a long impulse response. another problem was given by the high computational time and memory usage, which made it difficult to work with a large amount of data. several svm-like approaches [18], based on the least squares svm (ls-svm), are characterized by a very high number of parameters. many approaches use the bla, or a similar correlation analysis, as a starting point for the algorithm (e.g. [24-26]). then, the user does not have to take order decisions needed to parametrize the bla (or the qbla). in [23], a nonparametric approach to separate the front and back dynamics starting from the best linear approximation (bla) is proposed. in [16], a recursive identification method of wiener-hammerstein system with internal noises is developed. in [17], a recursive leastsquares identification method of wiener-hammerstein system is proposed. this approach is treated in the case where the system nonlinearity is dead-zone. then, a wide multiplicity of approaches is currently under study. these contain parametric and non-parametric methods. there exist different class of non-parametric * corresponding author: a. brouri, a.brouri@ensam-umi.ac.ma mailto:a.brouri@ensam-umi.ac.ma 62 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) solutions. among these concepts the frequency modeling, for which control and output are relying using periodic signals [4,8,10,14,20,22]. depending on the nature of the control (the input excitation) to the system, several estimators have been developed. in this paper, a solution based on analytical geometry (frequency study) is proposed. presently, the linear elements are allowed to be nonparametric. the nonlinear block can be discontinuous or of hard nature. examples of popular hard functions are shown by figs. 1.2ab. for convenience, fig. 1.2a illustrates coulomb friction nonlinearity with dead-zone and fig. 1.2b shows example of viscous friction function. a hard function is often known as being a nonlinearity having several discontinuities and it is an affine nonlinearity elsewhere [9,14]. on the other hand, the modeling of hard function using orthogonal polynomials decomposition often remains a challenge. it is very complicated to approximate these types of nonlinearities using any orthogonal basis approximation (e.g. polynomial decomposition). for convenience, fig. 1.3 illustrates example of orthogonal polynomial decomposition of hard function. this example shows that remarkable modeling errors are thus induced especially around discontinuity points even though the troncature degree m is too big. presently, the nonlinear element is allowed to be discontinuous and not necessarily affine function between two consecutive discontinuities. except in a small interval where it is supposed to be parametric function (e.g. polynomial nonlinearity). then, recall that the identification method is based only on control signal u(t) and observed output signal y(t) (i.e. all inner signals are not accessible). in this work, the linear blocks are not necessarily parametric. accordingly, because of these last difficulties, it is not surprising to notice that they are very rare papers dealing the identification of wienerhammerstein models having hard function. fig. 1.1. wiener-hammerstein model the paper is organized as follows: the identification problem is formulated in section 2, which also introduces the problem of multiplicity of identification solutions; section 3 is devoted to the determination of nonlinear system parameters (linear block and nonlinear element); the performances of the identification method are illustrated by simulation in section 4. -0.4 -0.3 -0.2 -0.1 0 0.1 0.2 0.3 0.4 -1 -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 0.8 1 example of hard nonlinearity fig. 1.2a. coulomb friction nonlinearity with dead-zone identification of nonlinear systems having hard function 63 copyright ©2019 assa. adv. in systems science and appl. (2019) -1 -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 0.8 1 -1 -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 0.8 1 nonlinearity with preload fig. 1.2b. viscous friction function -2 -1.5 -1 -0.5 0 0.5 1 1.5 -1 -0.8 -0.6 -0.4 -0.2 0 0.2 0.4 0.6 0.8 true n.l m=5 m=10 m=20 fig. 1.3. decomposition of hard nl with series of polynomial basis 2. problem formulation presently, the problem of determination of nonlinear system parameters is discussed. the considered nonlinear system is structured by wiener-hammerstein model (fig. 1.1). let gi(s) and go(s) denote the transfer function of linear blocks. the associated impulse responses (or the inverse laplace transforms) are respectively denoted  1 )( ()i ig l g st  and  1 )( ()o og l g st  . the control signal ( )u t and the inner signal v(t) are thus related by the following relation: ( ) ( )* ( ) i v t g t u t (2.1) where the notation “ * ” designates the convolution product. then, the inner signals v(t) and w(t) undergo the following equation:    ( ) ( ) ( )* ( )iw t f v t f g t u t  (2.2) as far as that goes, the inner signal w(t) and the undisturbed output x(t) are related as follows: ( ) ( )* ( ) o x t g t w t (2.3) finally, it follows from (2.3) and fig. 1.1 that, the system output can be expressed as:   ( ) ( )* ( ) ( ) ( )* ( ) ( ) o o y t g t w t t g t f v t t       (2.4) where the extra input ( )t accounts for measurement noise and other modelling effects. on the other hand, note that this identification problem does not have a unique solution [6,12]. indeed, if the triplet  ( ), ( ), ( )i og s f v g s is solution of this wiener-hammerstein 64 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) identification problem, then any set of form    ( ) , ( / ) / , ( ) ; 0 and 0i i i o o o i ok g s f v k k k g s for k k  is also solution of this problem (by distributing nonzero constants between the blocks of system). in this respect, the question that arises is how to choose a solution to this problem? the answer to this will be dealt with in the next section. the considered wiener-hammerstein nonlinear system is completed by the following assumptions: assumptions 2.1:  the nonlinear element is hard function having multiple discontinuities. except in a small interval where it is supposed to be parametric or polynomial function of degree increased by a finite integer n.  the linear elements have a nonzero static gains, i.e. (0) 0ig  and (0) 0og  .  the extra input ( )t is supposed to be zero-mean ergodic and uncorrelated with the control input u(t). □ except of these assumptions, the system is arbitrary. in particular, the nonlinear element is allowed to be discontinuous and it is not necessarily an affine function between two consecutive discontinuities. then, the linear elements are nonparametric. 3. system parameters estimation 3.1. nonlinear element estimation the goal presently is to develop a solution allowing to give the estimate of a set of points belonging to nonlinear function f(.). the question: from the plurality problem discussed in section 2, what is the system to be estimate? a key idea is to get benefit from this model plurality to make the identification problem more tractable. in this respect, the following selection of the free scalars ( , )i ok k will prove to be judicious: 1 (0)i i k g  and 1 (0)o o k g  (3.1) without loss of generality, it is readily follows from (3.1) and model plurality that, the linear elements of system to be determined check the following property: (0) (0) 1i og g  (3.2) on the other hand, let excite the system by a constant value: 1( )u t u ,  0t t (3.3) where t is more superior to the system rise time rt . then, it follows from (2.1), (3.2) and (3.3) that, the inner signal v(t) boils down (in steady state) to: 1( )v t u (3.4) one immediately gets from (2.2) and (3.4) that (after transient regime):  1( )w t f u (3.5) accordingly, in view of (2.3), (3.2) and (3.5), it follows that x(t) is written as follows (after transient regime): identification of nonlinear systems having hard function 65 copyright ©2019 assa. adv. in systems science and appl. (2019)  11( )x t x f u  (3.6) then, it is readily seen from (2.4), (3.2) and (3.6) that the output y(t) undergoes (after transient regime) the following expressions:  1( ) ( )y t f u t  (3.7) at this point, it is worth emphasizing that, the signal y(t) is constant up to noise. furthermore, it readily seen (in steady state) that:      1 1 1 1, , , ( )u x u x u f u  (3.8) is a point belonging to the nonlinearity f(.). one difficulty with the considered identification problem is that, the output of the system y(t) is infected by the disturbance ( )t whose stochastic law is not known. then, it follows from the assumption on the noise ( )t , just as suggested in [6,12] the following estimator for 1( )f u is proposed: 1 1 1ˆ ˆ( ) ( ) r r t t t f u x y t t t      (3.9) indeed, one has from (3.7) and (3.9) that:  11 1ˆ ( ) ( ) r r t t t uf u f t t t       (3.10) bearing in mind that, the noise ( )t is zero-mean ergodic stochastic sequence, the last term in (3.10) converges (with probability 1) to zero. this implies that:  11 ˆ ( ) t uf u f   w.p.1 (3.11) this result shows that, an accurate estimate of a point belonging to the function f(.) can be obtained. accordingly, the estimate of set of points belonging to f(.) can be achieved using the same procedure, i.e. applying the control (input) sequence: ( ) ku t u , ( 1) t k t kt     for 1k n (3.12) then, ˆ( )kf u ( 1k n ) can be obtained using the estimator (3.9): ( 1) 1ˆ ( ) ( ) k r r kt t k t t f u y t t t       for 1k n (3.13) remark 3.1: consider the nonlinear system (wiener-hammerstein) described by (2.1)-(2.4). then, the system nonlinearity estimator, given by (3.9), enjoys the consistency property for any value .ku specifically, one has the following result:  ˆ ( ) kk t uf u f   w.p.1 (3.14) the proof of this can be found by combining (3.9)-(3.10) and the property of ( )t (zeromean ergodic stochastic). □ 66 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) remark 3.2:  the identification method can be easily applied in the case where f(.) is parametric (e.g. polynomial function of degree n). then, a number of points 1n n  (arbitrarily chosen by the user) is largely sufficient to determine the nonlinearity over the entire working range.  in the case where f(.) is hard element, it is thus not necessarily a parametric function and can be discontinuous. furthermore, in a small interval (e.g. between two consecutive discontinuities), f(.) is supposed to be parametric or polynomial function of degree increased by an integer n. then, using the set of estimated points   ˆ, ( ) 1; k k u f u k n , choose a candidate interval and verify if the nonlinear element can be decomposed with orthogonal polynomials approximation. to this end, we can excite the system by other inputs in this interval. then, find out if we can approximate the set of points  ˆ, ( ) k k u f u with a polynomial function of degree n.  the last statement can be checked using a sine control in the chosen interval. accordingly, by observing the spectrum of the output signal if it does not contain harmonics of higher rank than n. for further information, please see the following subsection « the linear blocks determination ». in this interval, the output of nonlinear block can thus be expressed as: 0 ( ) ( )( ) k k n k w t c vf v t     (3.15) where  0 ... t nc c c is the parameters vector corresponding to the nonlinear element.□ 3.2. the linear blocks determination presently, the aim is to present an identification approach permitting to provide the estimate of linear blocks parameters. for convenience, the nonlinear system (2.1)-(2.4) is excited by the following control signal:  ( ) cos( )u t u t   (3.16) where the parameter  is adjusted such that u(t) belongs to the chosen interval using remark 3.2 for any small amplitude u. this result can be practically depicted by observing the spectrum of the output signal. let ( )i  and ( )o  designate the phases (argument) of linear elements ( ) i g j and ( ) o g j , for any frequency  , respectively. then, it readily follows from (2.1), (3.2) and (3.16) that, the inner signal v(t) (in the steady state) is given by:  ( ) ( ) cos( ( )) i i v t u g j t       (3.17) one has then using (3.15) and (3.17) that:   0 ( ) ( ) cos( ( )) i i kk k n k w t c u g j t         (3.18) bearing in mind that:     0 ( ) cos( ( )) ( ) cos( ( )) i i i i k k lk l l k l g j t c g j t                (3.19) where: identification of nonlinear systems having hard function 67 copyright ©2019 assa. adv. in systems science and appl. (2019)   ! ! ! k l k c l k l   (3.20) then, it immediately follows from (3.18)-(3.20) that:   0 0 ( ) ( ) cos( ( )) i i k lk k l k l n k k l w t c u c g j t           (3.21) for convenience, the formulas of power identities of types 2cos m  and 2 1cos m  can also be given analytically as:  2 2 2 2 2 1 1 0 1 1 cos cos 2( ) 2 2 m m m m pm m m p c c m p        (3.22a)  2 1 2 1 0 1 cos cos (2 1 2 ) 4 m m pm m p c m p       (3.22a) in view of (3.22a-b) and grouping the components having the same harmonics, (3.21) becomes (for any limited integer n):       0 1 ( ) ( ) ( ) cos ( ) i i ik k n k w t a g j a g j k t          (3.23) where the unknown variables in the amplitude  ( ) ika g j ( 0 )k n and the phases  ( ) ik   ( 1 )k n are the parameters of linear element ( ) i g j (i.e. the modulus gain ( ) i g j and the phase ( ) i   ). therefore, one immediately gets using (2.3), (3.2) and (3.23):        0 1 ( ) ( ) ( ) ( ) cos ( ) ( ) i i o i ok k n k x t a g j a g j g jk k t k               (3.24) it readily seen that, the unknown parameters in the expression of undisturbed output signal are  ( ) , ( ) i i g j   and   ( ) , ( ) ; 1 o o g jk k k n    . in this respect, note that the inner signal x(t) is equivalent to sum of n sinusoidal signals and dc component. specifically, having the spectrum of x(t), the equation (3.24) leads to a system of (2 1)n equations and (2 2)n unknowns, i.e. more unknowns than equations. this problem can be overcome by repeating the same experiment using the input signal (3.19) having the frequency 2 , 3 , …. the difficulty that arises is that, the signal x(t) is not accessible to measurement. fortunately, an estimate of x(t) can be established using the fact that this latter is periodic of the same period 2 /t   of control signal u(t). this property suggests the following estimator: 1 1 ˆ( ) ( ) l p x t y t pt l    for [0, )t t (3.25a) ˆ ˆ( ) ( )x t kt x t  for any integer k (3.25b) where l is any integer preferably large. indeed, one immediately gets from (2.4) and (3.25ab) that: 68 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) 1 1 1 1 1 ˆ( ) ( ) ( ) 1 ( ) ( ) l l p p l p x t x t pt t pt l l x t t pt l                (3.26) the periodic stationarity of ( )t means that  ( ) ( ( ))e t pt e t   , for any p and t. then, the zero-mean ergodicity of ( )t means that: 1 1 ( ) 0 l l p t pt l      (3.27) this implies that (using (3.26)-(3.27)): ˆ( ) ( ) l x t x t   (3.28) 4. simulation presently, the system (2.1)-(2.4) is characterized by the linear elements of transfer functions: 0.1 ( ) (0.4 )(0.1 ) i g s s s    (4.1a) 1 ( ) (0.2 )(0.5 ) o g s s s    (4.1b) the curve of nonlinear block (.)f , used in simulation, is shown by fig. 4.1. the noise signal ( )t is a sequence of random numbers, with zero-mean and standard deviation 0.5  . in the first stage, we apply to the input of system the sequence plotted in fig. 4.2. the collected system output is illustrated in fig. 4.3a. then, using the estimator (3.9) or (3.13), the estimate values ˆ( )kf u ( 1k n ) are also given by fig. 4.3a. for convenience, a zoom of these results is given by fig. 4.3b. -2 -1 0 1 2 3 -1.5 -1 -0.5 0 0.5 1 1.5 2 2.5 the n.l considered in simulation fig. 4.1. shape of the function f(.) considered in simulation identification of nonlinear systems having hard function 69 copyright ©2019 assa. adv. in systems science and appl. (2019) 200 400 600 800 1000 1200 1400 1600 1800 2000 2200 -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 2.5 3 time (s) u the input sequence u(t) fig. 4.2. the applied control sequence then, to compare the estimated and the true nonlinearities, fig. 4.4 shows the set of points   ˆ, ( ) 1 16; k k u f u k  and the true function (.)f (of rescaled nonlinear system using (3.1)-(3.2)). these results show that the estimated points   ˆ, ( ) 1 16; k k u f u k  are very close to their true values. it follows that the function does not seem to be of hard type in the interval  1 3 . in this respect, the system can be excited with other inputs within this interval and comparing the interpolation of these points with a polynomial of degree n. 200 400 600 800 1000 1200 1400 1600 1800 2000 2200 0 50 100 150 200 250 300 time (s) y(t) and f(u ) the system output y(t) the estimates of f(u ) k ^ k fig. 4.3a. the system output signal and the estimate of x(t) 600 700 800 900 1000 1100 1200 1300 1400 1500 -10 -8 -6 -4 -2 0 2 4 6 8 time (s) y(t) and f(u ) fig. 4.3b. zoom of the signal y(t) and the estimate of x(t) in the second stage, we apply to the system input the sine signal (3.16) where  is chosen such that u(t) belongs to the interval  1 3 . taking e.g. 2  , the resulting system output signal y(t) for an amplitude 1u  and a frequency 0.02 ( / )rd s  is illustrated by fig. 4.5. 70 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) the measured signal y(t) is collected on a sufficiently large interval. then, the collected sample is used to generate the undisturbed output ( )ˆ tx using (3.25a-b). the obtained estimate is plotted in fig. 4.6 over one period of time. -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 2.5 3 0 50 100 150 200 250 300 the true nonlinearity f(.) the set of estimated points fig. 4.4. comparison between the true and estimated nonlinearities accordingly, it readily follows from the results of first stage and using an orthogonal polynomial decomposition of degree 3n  that, the obtained expression of ˆ(.)f within the interval  1 3 is given as: 3 2ˆ( ) 0.003 +31.24 0.02 24.8f v v v v    (4.2a) while the expression of the true nonlinearity is as follows: 2( ) 31.25 25f v v  (4.2b) let us consider: 3 ˆ 0.003c  ; 2 ˆ 31.24c  ; 1̂ 0.02c  ; 0 ˆ 24.8c   ; (4.3) denote the coefficients of the estimated nonlinearity. on the other hand, it readily follows from (3.18)-(3.23) and (4.2a) ( 1)u  that, the inner signal w(t) can be rewritten as:       0 3 1 ( ) ( ) ( ) cos ( ) i i ik k k w t a g j a g j k t          (4.4a) where:               2 2 0 0 1 2 3 2 3 2 2 1 1 2 3 2 2 2 3 3 3 3 ( ) ˆ ˆ ˆ ˆ ˆ ˆ( ) 3 ; 2 ( ) ˆ ˆ ˆ( ) 2 3 ( ) ; 4 1 ˆ ˆ( ) 3 ( ) ; 2 1 ˆ( ) ( ) ; 4 i i i i i i i i i g j a g j c c c c c c g j a g j c c c g j a g j c c g j a g j c g j                                          (4.4b) and:      1 2 3( ) ( ); ( ) 2 ( ); ( ) 3 ( ); i i i i i i                  (4.4c) then, one immediately gets using (3.2), (3.24) and (4.4a-c): identification of nonlinear systems having hard function 71 copyright ©2019 assa. adv. in systems science and appl. (2019)               0 1 2 3 ( ) ( ) ( ) ( ) cos ( ) ( ) ( ) ( 2 ) cos 2 2 ( ) (2 ) ( ) ( 3 ) cos 3 3 ( ) (3 ) i i o i o i o i o i o i o x t a g j a g j g j t a g j g j t a g j g j t                                 (4.5) where des parameters  ( ) ika g j ( 0 ... 3k  ) are given by (4.4b). further, note that the inner signal ( )tx is periodic of the same period of ( )tu (i.e. 2 /t   ). accordingly, ( )tx can be expanded in series fourier, where its parameters (the amplitudes and arguments) can be easily generated using the estimate ( )ˆ tx . then, it readily follows from (4.5) and ( )ˆ tx that, 7 equations are provided. specifically, using the estimate of dc component 0a , the value of harmonic amplitudes ka ( 1... 3k  ), and the estimate of harmonic phases. accordingly, it readily seen from (4.4b) and (4.5) that, the modulus gain ( ) i g j can be determined using the dc component of ˆ( )x t . then, the modulus gains ( ) o g jk , for 1... 3k  , can be immediately estimated using the amplitude of the first 3 harmonics (three unknowns and three equations). furthermore, using the argument of the fundamental component and that of the first two harmonics, one has 4 unknowns (i.e. ( ) i   , ( ) o   , (2 ) o   , and (3 ) o   ) and 3 equations (see (4.5)). we have more unknowns than equations. the nonlinear system is thus excited with the sine input (3.16) with the frequencies 2 and 3 . it readily follows that, these experiments generate 9 equations (from the phases) involving 9 unknowns ( ( ), i   (2 ) i   , (3 ) i   , ( ) o   , (2 ) o   , (3 ) o   , (4 ) o   , (6 ) o   and (9 ) o   ). finally, an estimate of ( ) i l  and ( ) o kl  , for 1... 3l  and 1... 3k  , can be obtained. 0 100 200 300 400 500 600 700 0 50 100 150 200 250 time (s) y true system output fig. 4.5. the system output y(t) then, it follows from these three experiments that, for any frequency  , the estimate of modulus gains ( ) i g jl and ( ) o g jkl ( 1... 3l  and 1... 3k  ) can also be given. on the other hand, in the case where the linear elements are parametric, the coefficients of transfer functions (numerators and denominators) can easily be determined using these estimates [6,12]. among the advantages of the proposed method, the estimate of the gains ( ) i g jl ( ( ) i g jl and ( ) i l  ) and ( ) o g jkl ( ( ) o g jkl and ( ) o kl  ), for 1... 3l  and 1... 3k  , can be obtained using only 3 experiments. then, repeating this steps for 0.04 ( / )rd s  and 0.06 ( / )rd s  , i.e.: excite the system by the control (3.16) with these frequencies and estimate the corresponding ( )ˆ tx . this latter allows us to get the fourier parameters. finally, the gains ( ) i g jl and ( ) o g jkl ( 1... 3l  and 1... 3k  ) can easily be determined using (4.4b) 72 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) and (4.5). table 4.1 gives the estimated numerical values. the obtained results show that the estimated parameter values are very close to their true values. 0 50 100 150 200 250 300 50 100 150 200 250 time (s) the filtered output fig. 4.6. the stimated signal ˆ( )x t over one period of time table 2.1. estimate of linear elements and their true values ( / )rd s 0.02 0.04 0.06  ( ) , ( )i ig j   (0.97, -0.25) (0.92, -0.48) (0.85, -0.69)  ( ) , ( )o og j   (0.99, -0.14) (0.98, -0.27) (0.95, -0.41)  ( 2 ) , (2 )o og j    (0.98, -0.27) (0.91, -0.54) (0.83, -0.77)  ( 3 ) , (3 )o og j    (0.95, -0.41) (0.83, -0.77) (0.7, -1.08)  ˆ ˆ( ) , ( )i ig j   (0.98, -0.24) (0.9, -0.46) (0.87, -0.67)  ˆ ˆ( ) , ( )o og j   (1.01, -0.15) (0.96, -0.25) (0.98, -0.39)  ˆ ˆ( 2 ) , (2 )o og j    (0.96, -0.25) (0.93, -0.57) (0.80, -0.81)  ˆ ˆ( 3 ) , (3 )o og j    (0.98, -0.39) (0.80, -0.81) (0.74, -1.04) 5. conclusion in this paper, an identification method of nonlinear systems is proposed. presently, the nonlinear system can be described by wiener-hammerstein model. the nonlinear element is allowed to be discontinuous or of hard shape. it is interesting to point that, it is very complicate to decompose or approximate these types of nonlinearities using polynomial decomposition. the estimation of system parameters is done using two stage. firstly, the nonlinear block is determined using a simple sequence of constant controls. in the second stage, the linear elements parameters are estimated using sine signal. the identification method also features the fact that the linear elements identification is made decoupled from the nonlinear element identification. to the author's knowledge very few previous studies have been dealt with discontinuous or hard nonlinearity and nonparametric linear elements. identification of nonlinear systems having hard function 73 copyright ©2019 assa. adv. in systems science and appl. (2019) references [1] bai, e.w., cai, z., dudley-javorosk, s., & shields r.k. (2009). identification of a modified wiener-hammerstein system and its application in electrically stimulated paralyzed skeletal muscle modeling. automat. contr., 45(3), 736-743. [2] benyassi, m. & brouri, a. (2017). identification of nonlinear systems having nonlinearities at input and output. european conf. on elec. eng. and comp. sc. (eecs), bern, 311-313. [3] bershad, n.j., celka, p., & mclaughlin, s. (2001). analysis of stochastic gradient identification of wiener-hammerstein systems for nonlinearities with hermite polynomial expansions. ieee trans. signal process., 49, 1060-1071. [4] brouri, a. (2016). frequency identification of nonlinear systems. lap lambert academic publishing, isbn (978-3-659-94991-3). [5] brouri, a. (2016). wiener-hammerstein models identification. int. journal of math. mod. & meth. in applied sc., 10, 244-250. [6] brouri, a. (2017). frequency identification of hammerstein-wiener systems with backlash input nonlinearity. w. trans. on syst. & contr., 12, 82-94. [7] brouri, a. (2017). identification of nonlinear systems. aip conference proceeding, icamcs, rome, 1836(1), https://doi.org/10.1063/1.4981971. [8] brouri, a., chaoui, f.z., amdouri, o., & giri, f. (2014). frequency identification of hammerstein-wiener systems with piecewise affine input nonlinearity. 19th ifac world congress, cape town, 10030-10035, https://doi.org/10.3182/20140824-6-za-1003.00303. [9] brouri, a. & giri, f. (2012). identification d’un modèle de wiener-hammerstein comprenant une non-linéarité affine par morceaux. cifa, grenoble, france. [10] brouri, a., giri, f., ikhouane, f., chaoui, f.z., & amdouri, o. (2014). identification of hammerstein-wiener systems with backlash input nonlinearity bordered by straight lines. 19th ifac world congress, cape town, 475-480, https://doi.org/10.3182/20140824-6-za-1003.00678. [11] brouri, a. & kadi, l. (2018). contribution on the identification of nonlinear systems. codit'18 conference, thessaloniki, 605-610. [12] brouri, a., kadi, l., & slassi, s. (2017). frequency identification of hammerstein-wiener systems with backlash input nonlinearity. int. j. of control, automation & systems, 15(5), 2222-2232. [13] brouri, a., kadi, l., & slassi, s. (2017). identification of nonlinear systems. european conf. on elec. eng. and comp. sc. (eecs), bern, 286-288. [14] brouri, a., rabyi, t., & ouannou, a. (2018). identification of nonlinear systems with hard nonlinearity. codit'18 conference, thessaloniki, 506-511, https://ieeexplore.ieee.org/document/8394834/. [15] brouri, a. & slassi, s. (2016). identification of nonlinear systems structured by wiener-hammerstein model. intern. j. of elec. and comp. eng., 6(1), 167-176. [16] falck, t., pelckmans, k., suykens, j., & de moor, b. (2009). identification of wiener-hammerstein systems using ls-svms. 15th ifac symposium on system identification, saint-malo, https://doi.org/10.3182/20090706-3-fr-2004.00136. https://doi.org/10.1063/1.4981971 https://doi.org/10.3182/20140824-6-za-1003.00303 https://controls.papercept.net/conferences/conferences/cifa12/program/cifa12_contentlistweb_1.html https://controls.papercept.net/conferences/conferences/cifa12/program/cifa12_contentlistweb_1.html https://doi.org/10.3182/20140824-6-za-1003.00678 https://www.scopus.com/record/display.uri?eid=2-s2.0-85050194294&origin=resultslist&sort=plf-f&src=s&sid=76dfaa66776891fe92b39fedca761532&sot=autdocs&sdt=autdocs&sl=18&s=au-id%2823567866000%29&relpos=3&citecnt=0&searchterm= https://www.scopus.com/record/display.uri?eid=2-s2.0-85050194294&origin=resultslist&sort=plf-f&src=s&sid=76dfaa66776891fe92b39fedca761532&sot=autdocs&sdt=autdocs&sl=18&s=au-id%2823567866000%29&relpos=3&citecnt=0&searchterm= https://doi.org/10.3182/20090706-3-fr-2004.00136 74 a. brouri, t. rabyi, a. ouannou copyright ©2019 assa adv. in systems science and appl. (2019) [17] li, l. & ren, x. (2017). decomposition-based recursive least-squares parameter estimation algorithm for wiener-hammerstein systems with dead-zone nonlinearity. inter. jour. of syst. sc., 48(11), 2405-2414. [18] marconato, a. & schoukens, j. (2009). identification of wiener-hammerstein benchmark data by means of support vector machines. 15th ifac symposium on system identification, saint-malo, https://doi.org/10.3182/20090706-3-fr2004.00135. [19] marconato, a., sjoberg, j., & schoukens, j. (2012). initialization of nonlinear state-space models applied to the wiener-hammerstein benchmark. control engineering practice, 20, 1126-1132. [20] mu, b.q. & chen, h.f. (2016). recursive identification of wienerhammerstein systems. siam j. contr. opt., 50(5), 2621–2658. [21] pillonetto, g., chiuso, a., & nicolao, g.d. (2011). prediction error identification of linear systems: a nonparametric gaussian regression approach. automatica, 47(2), 291-305. [22] pintelon, r., guillaume, p., rolain, y., schoukens, j., & van hamme, h. (1994). parametric identification of transfer functions in the frequency domaina survey. ieee trans. on aut. cont., 39(11), 2245-2260. [23] schoukens, m., pintelon, r., & rolain, y. (2014). identification of wienerhammerstein systems by a nonparametric separation of the best linear approximation. automatica, 50(2), 628-634. [24] sjöberg, j. & schoukens, j. (2012). initializing wiener-hammerstein models based on partitioning of the best linear approximation. automatica, 48(2), 353359. [25] sjöberg, j., lauwers, l., & schoukens, j. (2012). identification of wienerhammerstein models: two algorithms based on the best split of a linear model applied to the sysid’09 benchmark problem. control engineering practice, 20, 1119-1125. [26] westwick, d.t. & schoukens, j. (2012). initial estimates of the linear subsystems of wiener-hammerstein models, automatica, 48, 2931-2936. [27] wills, a. & ninness, b. (2009). estimation of generalised wienerhammerstein systems. 15th ifac symposium on system identification, saintmalo, 1104-1109, https://doi.org/10.3182/20090706-3-fr-2004.00183. https://doi.org/10.3182/20090706-3-fr-2004.00135 https://doi.org/10.3182/20090706-3-fr-2004.00135 https://doi.org/10.3182/20090706-3-fr-2004.00183 adv syst sci appl 2018; 02:84–92 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/542. segregation model for dynamic frequency allocation alexander kuznetsov1∗, 1voronezh state university, voronezh, russia abstract: we apply the schelling type ii segregation model to the dynamic frequency allocation. an algorithm is introduced for agents segregation over initially unknown radio channels. we relate the number of algorithm iterations until complete agents’ segregation to the number of agents, networks, and the other parameters via numerical experiment. also, there are some parallels with the continuous-time model of segregation used in sociology. keywords: segregation model, dynamic frequency allocation, cellular automaton 1. introduction nowadays, the number of robotic systems consisting of a large number of autonomous mobile agents (terrestrial robots, uavs, etc.) with wireless radio communication network constantly increases. agents interact with each other to perform a common task, for example, for emergency response, intrusion detection to a protected area, reconnaissance or networking in areas where it is difficult to organize a communication network in the usual ways [4] and so on. human-mediated deployment or restoration of the communication system of such agents can be extremely difficult – for example, in an emergency situation it can be too dangerous for human operators. for this reason, it is necessary to provide the agents themselves with all the logic necessary to maintain the communication network, initially setting only the basic principles of agents’ interaction. this problem is investigated in detail in [5], and below we introduce a simple model of such self-organization. briefly, agent periodically scans radio channels to discover other agents or stops at one of the channels and transmits a beacons sequence in order to be discovered by other agents. the problem we consider is one of cognitive radio ad hoc networks problems. the article [8] introduces dynamic spectrum access (dsa) which uses spectrum policy reasoning to determine allowed frequencies, requesting sensing periods on those frequencies, classifying the results from sensing events, then providing the list of allowed frequencies for use in frequency assignment. the papers [3, 10] describe different physical and informational aspects of the dynamic spectrum access. according to these works, two main approaches to dsa exists: dynamic spectrum allocation and opportunistic spectrum access. dynamic spectrum allocation exploits temporal and spatial traffic statistics and aims at improving spectrum efficiency through timeand space-dependent spectrum sharing among coexisting radio services. different from dynamic spectrum allocation, which uses the statistics of spectrum occupancy, opportunistic spectrum access uses the instantaneous spectrum availability by opening the licensed spectrum to secondary users. the idea is to allow ∗corresponding author: avkuz@bk.ru segregation model for dynamic frequency allocation 85 secondary users to identify available spectrum resources and communicate opportunistically in a manner that limits the level of interference perceived by primary users. this article is motivated by the question, whether it is possible to create a completely autonomous cognitive radio network, without any base station or initial spectrum allocations, or coordination from a base station is necessary for agents to distribute over radio channels in a finite number of iterations? the purpose of the present article is the attempt to disconnect network self-organization from its radio-physics ground and to study it as a purely mathematical segregation problem. the dynamical network organization is, in its essence, a segregation model related to the schelling segregation model [9]. the schelling cellular automaton (fig. 1.1) is a grid where each cell corresponds to a person. a person can have one of several “colours” that symbolizes her belonging to a particular social group. colours are distributed randomly on the grid, and there is also some amount of unoccupied cells. every person wants a certain part of people around him to be like him. if the number of adjacent cells of the same colour is below the specified threshold, then the person goes to the next free cell, otherwise, it remains in place. however, unlike the two-dimensional classical schelling model, the proposed automaton is practically one-dimensional: we need the coordinate n of an agent only to simplify the description of the dynamics of the automaton. fig. 1.1. the schelling segregation model 2. the cellular automaton model denote as id ⊂ z, 0 ∈ id the set of possible agent’s identifiers, asq ⊂ n the set of channel qualities. definition 2.1: let us define the set of channels as the simply connected domain c ⊆ z2 where each (n, f) ∈ c associated with the cell (qnf , idnf ) ∈ q× id and qkf = qlf for every (k, f), (l, f) ∈ c. we assume that if idnf = 0 there is no agent in the cell (n, f). therefore, the f th column of cells {(i, f)} corresponds to the f th channel. the number nf of cells in the f th column corresponds to the maximum possible number of agents using the f th channel. definition 2.2: let us call the agent the following vector ag = (idag, colag, fself , fold, nself ; ftarget, rag,qag; sensed; t), copyright c© 2018 assa. adv syst sci appl (2018) 86 a. kuznetsov where idag ∈ id is the unique agent’s identifier, fself is the current channel number, nself is the position of the agent on the current channel fself , ftarget is the number of a target channel, fold is the auxiliary variable. further, colorag ∈ col is the agent’s colour, rag : col × id → {0, 1} is the grouping predicate, qag : q → [0, 1] is the agent preferences function, t is an agent’s overall functioning discrete time. the set of all agents we denote as ag. the sensed is the agent’s knowledge of environment (see below). define the predicate r : ag ×ag → {0, 1} as following r(ag1, ag2) = rag1(colag2 , idag2), ag1, ag2 ∈ ag. if agents ag1, ag2 ∈ ag need to occupy the same channel, then r(ag1, ag2) = 1, otherwise r(ag1, ag2) = 0. it is clear, rag1 should be defined that rag1(ag2) = rag2(ag1) for any ag1, ag2 ∈ ag. note that r determines the communication graph for ag. definition 2.3: let p, x are functions. we shall write p ∼ x, if it exists such monotonically non-negative increasing function ϕ that p(x) = ϕ(x). for simplification, assume that c = {(f, n)|f = 1, fmax, n = 1, nmax}. also, we denote qnf = qf , agf ⊆ ag is the set of agents on the channel f (i.e. with fself = f ), agag = {ãg ∈ ag|r(ag, ãg) = 1} is the set of agents from the same network as the agent ag, and agag f = {ãg ∈ agf |r(ag, ãg) = 1} is the set of agents on the channel f from the same network as the agent ag. discovering of agents can be unsuccessful. by this reason, introduce the function d : 2ag → 2ag . the agent ag ∈ agf falls into the d(agf ) with probability p1 ∼ qf . our automaton will function in a discrete time. an agent’s tact behaviour can be described by the following algorithm. 0. initialization. the agent ag scans all channels in c and randomly selects (n, f) with probability p0(n, f) ∼ qag(qnf ) and such that the cell (n, f) is not already occupied. 1. sensing. the agent ag senses the channel f . the agent sets fself := f, nself := n, t := t+ 1, sensed[f ] := {qf ,d(agf )}. 2. if fself = fmax, then the agent chooses a channel f̃ from sensed with probability p0(n, f). for example, p0 can be defined so that the agent selects the channel with the best quality. the agent sets ftarget := f̃ . 3. decision. if |d(agag f )| = |ag ag|, then go to the step 6, else if |d(agag f )| > |agag| α0 ∧ |d(agag f )| > |d(agf )| − |d(ag ag f )|, α0 > 1, then go to 5, else go to 4. 4. channel change. the agent ag memorizes the current frequency fold := fself copyright c© 2018 assa. adv syst sci appl (2018) segregation model for dynamic frequency allocation 87 and chooses a new frequency fself := fself + dir, 1 ≤ fself + dir ≤ fmax, fmax, fself + dir = 0, 1, fself + dir > fmax. if an arbitrary coordinate n such that the cell (fself , n) is unoccupied exists, the agent chooses the cell (fself , n) and occupies it. if there is no such cell, then the agent does not occupy any cell on this turn and remains on its old cell. in this case, the agent sets fself := fold, go to 1. if fself = ftarget, then go to 5 else go to 1. 5. waiting for agents. the agent waits twait = α1fmax + α2τ(|d(agag f )|, |ag ag|, |d(agf )|), α1 ≥ 0, α2 ≥ 0, (2.1) τ(x, y, z) ∼ min { x y , x z } , (2.2) turns and selects dir randomly from the set {−1, 1}. go to 1. 6. wait. the main problem is to find model’s parameters αi, i = 1, 3 and a function τ such that the segregation process would complete in a reasonable time. we should note that an unlucky choice of parameters will result in the algorithm not converging at all. physically, agents get information about other agents and about a state of channels, scanning the spectrum and exchanging beacons, as described in [5]. one tact of the cellular automaton corresponds to the full cycle of the pilot and beacons exchange. a detailed technical solution corresponding to the proposed model is described in the patent “telecommunication network data transmission means and telecommunication network” # ru 2 549 120. fig. 2.1. final configuration of the automaton the example of the automaton’s final state is shown on fig. 2.1. the lighter tone of cells in the picture corresponds to the better quality of corresponding channels. we can see the network net1 = {agi|i = 1, 4} formed on the channel f1, and net2 = {agi|i = 5, 6} formed on the channel f2. copyright c© 2018 assa. adv syst sci appl (2018) 88 a. kuznetsov 3. numerical experiment and discussion we performed simulation of the proposed algorithm with the “psychohod” simulator [7]. for these purposes, we introduced the new operating mode of the program (see fig. 3.1). fig. 3.1. “psychohod” program in the process of agents segregation we defined p1 = 1, the grouping predicate as the following r(ag1, ag2) = { 1, colag1 = colag2 , 0, colag1 6= colag2 . fig. 3.2 contains simulation results for 5 networks, 50 agents in each network, and for α1 = fmax + 1, α2 = (fmax + 1)β, τ(x, y, z) = min { x y , x z } , where α1, α2, τ are parameters from (2.1), (2.2). ordinates correspond to the average discrete time t , abscissas correspond to the α0. channels qualities q = 1, 9 were distributed uniformly over fmax = 100 channels, 100 experiments for each point were performed. fig. 3.2. average time to complete network segregation (50 agents per net, 5 nets) copyright c© 2018 assa. adv syst sci appl (2018) segregation model for dynamic frequency allocation 89 the simulation stopped when the segregation process was completed or when t > 30000. the points encircled correspond a case when for all 100 experiments the segregation was completed in less than 30000 turns. for other points, the segregation process was completed in 79–99% of experiments. fig. 3.3 represents histograms of segregation’s completion times for different values of α0 and β. we used function t (α0) = a+ b α0 − c , c > 0, b > 0 for approximations of data points. fig. 3.3. histograms of segregation’s times next, we fixed α0 = 2.2, β = 30 and performed experiments with different numbers of nets and agents in each net as described in table 3.1. note that segregation considered table 3.1. unsuccessful segregation percentage agents in net 10 20 30 40 50 60 70 80 90 100 unsuccesses number, five nets 37 15 3 4 1 0 0 5 9 30 unsuccesses number, eight nets 97 75 59 33 22 2 1 18 29 100 unsuccessful if it takes more than 30000 turns. therefore, we should change α0 and β proportional to the number of agents. for example, if we have 25 agents in each of five nets, we need to reduce β approximately in four times and slightly increase α0 (see fig. 3.4) to minimize the average segregation time. results of the numerical experiment immediately suggest an improvement of the previously proposed algorithm. we should add the following step: 2.5. the elapsed time check. if t > tr, then change αi, i = 0, 2 on arbitrary small values. also, we can memorize all changes of algorithm’s parameters and gradually choose the optimal ones. the further improvement can be dividing of the set of channels into subsets so that agents from one net would tend to search channels in the channel subset associated with their net number. copyright c© 2018 assa. adv syst sci appl (2018) 90 a. kuznetsov fig. 3.4. average time to complete network segregation (25 agents per net, 5 nets) the proposed cellular automaton is close to the so-called schelling type ii model. works [1, 2], which introduce continuous-time schelling dynamical system, characterize schelling ii model as the following. let the population is partitioned into disjoint sets, akin to the different districts in a city. the population is divided into “types”. let x and y denote the two types, inhabiting a space which is partitioned into m areas, denoted ai, i = 1, 2, ...,m. also, let xi and yi denote the total x-type population and y -type population in ai, |x| and |y | the total of each type in the population, and n = |x|+ |y | the total population. tolerances are allocated to a given type in a given area via a “tolerance schedule”. this function describes the maximum x-type population that would tolerate up to r(x)xi members of type y in the same area, where r(x) is necessarily monotone decreasing. it is also assumed that there is no lower bound on tolerance, i.e. no population insists on the presence of the opposing type. the simplest, single-area continuous-time schelling dynamical system has the following form: dx dt = [xrx(x)− y]x, (3.1) dy dt = [yry (y)− x]y, (3.2) where x, y are densities of populations x , y , |x| = k|y |, p > 0, and rx , ry are monotone decreasing tolerance schedules’ functions. here, for example rx(x) = a(1− x)p, (3.3) ry (y) = b(1− ky)p, (3.4) a > 0, b > 0, p > 0, k = |x|/|y |. we have found, that our automaton’s behaviour in some cases can be described by a model like (3.1), (3.2). denote as x(t) = xf (t)/|x(t)| the density of agents of the net 1 on the network channel f . let’s also denote y(t) = yf (t)/|y (t)| the density of agents of all other networks on the network channel f . as we have 50 agents per net and 5 nets, k = |x|/|y | = 1/4, x(18) = 1/50, y(18) = 1/200. at first, select p = 1.7, a = 1, b = 1/2. copyright c© 2018 assa. adv syst sci appl (2018) segregation model for dynamic frequency allocation 91 we can modify (3.1), (3.2) in the way described in [1], using the exponential schedule rx(x) = e−6x − e−6 1− e−6 , (3.5) ry (y) = e−6y/k − e−6/k 2(1− e−6/k) , (3.6) to obtain better fitting. we can see the comparison of experimental points with the solutions of schelling dynamical system on fig. 3.5, left. circular and diamond-shaped markers correspond to experimental values of x and y. the black dotted line corresponds to the linear schedule (3.3), (3.4), solid lines correspond to the exponential schedule (3.5), (3.6). interestingly enough, that such initial condition exists which can provide oscillatory behaviour for xf . this is the case of approximately equal initial numbers of agents of different nets at the channel f (see fig 3.5, right). unfortunately, the model (3.1), (3.2) can not explain this case at whole, and we should use more complex multi-area segregation model, but it is possible to model oscillations by multiplying (3.5), (3.6) with a periodic function. the aforesaid situation entails the potential impossibility of the network’s self-organization in a reasonable time. fig. 3.5. graphs of numerical experiment’s results in comparison with the solutions of schelling dynamical system 4. connection with the agents’ motion model previously, the author had developed the model of agents’ motion and conflict and related simulator “psychohod”. groups of agents move through a rough terrain and can destroy other groups of agents. this simulator can export data with coordinates of communicating agents, the list of destroyed agents, and landscape obstacles to communication. the data exported used by the communication simulator which contains described above automaton as well as by other telecommunication models (fig. 4.1). optionally, the communication simulator can export the message exchange data to a network simulator like ns-3 or omnet++. 5. conclusions we proposed the simulation of the automatic initial channels distribution in a wireless network. also, we studied the dependence of the channels distribution time on various parameters like the nets’ number, the agents’ number, the channels’ scanning speed. at last, we compared the simulation results with the solution of schelling dynamical system. we found that the proposed network self-organization is similar to social segregation schelling type ii models. copyright c© 2018 assa. adv syst sci appl (2018) 92 a. kuznetsov fig. 4.1. conjugation of the agents’ motion simulator, the communication simulator, and (optional) a network simulator in the future, we are planning to design an adaptive algorithm providing minimization of the channels distribution time, a model with relay agents based on the “psychohod” simulator, and to obtain numerical characteristics of aforesaid models. it is also planned to provide agents with additional characteristics, such as a memory of channels visited, and link the proposed model of self-organization of the communication system with the model of the movement of agents [6], previously developed by the author. it seems promising to compare the proposed cellular automaton with more complex continuous-time segregation models to explain the occurrence of the chaotic or oscillatory behaviour of the system. this work was supported by rfbr grant 18-07-01240 a. references 1. haw, d. (2016). measuring and understanding segregation (doctoral dissertation, bristol centre for complexity sciences). 2. haw, d. j., & hogan, j. (2018). a dynamical systems model of unorganized segregation. the journal of mathematical sociology, 42(3), 1-15. doi: 10.1080/0022250x.2018.1427091 3. horne, w. d. (2003). adaptive spectrum access: using the full spectrum space. in proc. of the international symposium on advanced radio technologies. 4. jones, m., & dudenhoeffer, d. (2000). a formation behavior for largescale micro-robot force deployment. in winter simulation conference (vol. 01, p. 972-982). los alamitos, ca, usa: ieee computer society. doi: doi.ieeecomputersociety.org/10.1109/wsc.2000.899900 5. kuznetsov, a. (2017). allocation of limited resources in a system with a stable hierarchy (on the example of prospective military communications system). large-scale systems control, 66, 68–93. 6. kuznetsov, a. v. (2017). a simplified combat model based on a cellular automaton. journal of computer and systems sciences international, 56(3), 397–409. doi: 10.1134/s106423071703011x 7. kuznetsov, a. v. (2018). model of the motion of agents with memory based on the cellular automaton. international journal of parallel, emergent and distributed systems, 33(3), 290-306. doi: 10.1080/17445760.2017.1410819 8. redi, j., & ramanathan, r. (2011, nov). the darpa wnan network architecture. in 2011 milcom 2011 military communications conference (p. 2258-2263). doi: 10.1109/milcom.2011.6127657 9. schelling, t. c. (1971). dynamic models of segregation. the journal of mathematical sociology, 1(2), 143-186. doi: 10.1080/0022250x.1971.9989794 10. zhao, q., tong, l., & swami, a. (2005). decentralized cognitive mac for dynamic spectrum access. in first ieee international symposium on new frontiers in dynamic spectrum access networks, 2005. dyspan 2005. (p. 224-232). doi:10.1109/dyspan.2005.1542638 copyright c© 2018 assa. adv syst sci appl (2018) advances in systems science and application (2015) vol.15 no.4 326-337 biodiversity of island ecosystems of the northern and middle caspian and a new outlook at the islands age and the caspian sea level regime g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev dagestan state university, institute of applied ecology, makhachkala, russia. abstract species diversity of some biota taxonomic groups of the larger islands (chechen, tyuleniy, nordovyy, kulaly) of the middle and northern has been established by the research expeditions of 2009-2013. endemic taxa of species and subspecies rank have been found in the flora and fauna of this island group. the results of the gis simulation (of the map), showing the water surface configuration under the sea level lowering have been presented and different approaches to explain the fluctuations in the caspian sea level and the mechanisms of islands and evolution and depositional bench lands of the northern and middle caspian have been briefly described in the article. the supposition on the presence of island firm land throughout the holocoen (even during the highest sea level) has been made on the basis of analysis of the island biota composition and the causal explanation of the endemic taxa presence in the water area of the middle and northern caspian. keywords the caspian sea; biodiversity; coastal and island ecosystems; level fluctuation,autochthonous speciation. 1 introduction level regime of the caspian sea is an issue to be much discussed in the scientific printed matter in the theoretical and applied perspective. there were a variety of views on the causes of the fluctuation in this sea-lake [1–4] and forecast expectations of its level regime [5-17]. in these and other papers different hypotheses on the causes of the caspian sea level regime instability are grouped together in the geological, hydrogeological, climatic and technological conceptions. we do not set a goal to give a detailed analysis of the causes and forecasts of the caspian sea level fluctuations. it should be noted that the climate hypotheses, in our opinion, are of greater importance. our interest to the caspian sea level fluctuations is connected with the necessity to give a causal explanation of the identified taxonomic characteristics of the caspian sea island biota. 2 materials and methods during the 2009-2013 large-scale comprehensive studies of flora and fauna of the coastal and island ecosystems of the northern and middle caspian were made jointly by ecology and geography department of dagestan state university and advances in systems science and application (2015) vol.15 no.4 327 institute of applied ecology of the republic of dagestan. the research covered the western and eastern coasts, as well as the large islands of this part of the sea. the main focus of the expedition was to study the specific diversity of flora and fauna. faunal material processing was performed in the zoological institute of the russian academy of sciences (saint petersburg). table 1 summary table of species diversity of the island middle and northern caspian no. taxa number of genera number of species new for the science new for russia 1 darkling beetles (coleoptera, tenebrionidae) 127 341 1 species 2 owletmoths (lepidoptera, noctuidae) 279 902 1subfamily 1genus 3 carabid beetles (coleoptera, carabidae) 98 608 1 genus 1species 4 spiders (aranei) 131 290 2 species 5 click beetles (coleoptera, elateridae) 6 12 6 oribatid mites (acariformes, oribatida) 39 49 2 species 12 species 8 snout beetles (coleoptera, curculionidae) 127 318 9 orthopterous insects (orthoptera) 24 30 10 dung beetles (coleoptera, scarabaeidae) 133 363 1 species 1 subspecies 11 higher plants (cormophyta) 186 269 1 species 2 species total: 1150 3182 7 20 3 results as a result of this processing work of floristic (higher plants) and fauna (some groups of invertebrates) field data, the general picture of species diversity of large islands (chechen, tyuleniy, nordovyy, kulaly) of the north and middle caspian was determined. firstly, it should be emphasized that the biota of the islands, in general, consists of taxa which are widely spread in the eastern and, to a lesser extent, in the western coast of the sea. there are also species which extend far beyond the caspian sea region. however, the fauna and flora data given in table 1 show that new species currently unknown for the science in 328 g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev:biodiversity of island... the continental part of the caspian are proven to have appeared in the modern biota of the caspian sea islands. these are taxa of species or subspecies rank. considering the above said, the modern floral or faunal status of the given species group is to be recognized as endemic. formation and evolution stabilisation of these taxa requires time sufficient for the diagnostically important features to emerge. undoubtedly, the present-day configuration of island biota taxa natural habitat of the middle caspian depends on the scale and time duration of transgressive and regressive sea cycles and on the causes predetermining those events. in line with the nature of these cycles changes in the species composition and structural organisation of the island biocenoses were taking place. the need for a causal interpretation of the present-day configuration of the natural habitat of the biota species composition of the islands investigated and the endemic taxa presence made us develop a hydrodynamic gis model of the caspian. to solve this problem, a three-dimensional model of the caspian sea has been made in map 2011 gis. matrices keeping the hydrodynamic gis models functioning are based on approximately forty thousand depth soundings. the base level of the caspian sea in this simulation is -28 m point. below (fig.1-10) is the configuration of the sea water surface at different points of sea level lowering. the analysis of these materials allows us to make some important points: 1. the catastrophic reduction of the sea water plane square takes place with a gradual lowering of its level up to -38 m (fig.6). 2. at the level of -30 m by the islands of chechen, nordovyy, kulaly overland connection with coastal land appears (fig.2). island tyuleniy has the same connection between -31 and -32 m (fig.3). 3. between -32 and -33 m within the ural furrow an independent pool is differentiated (fig.4). 4. from the level of -34 m the baring of underwater bench land chain from the area of chechen island in the direction of kulaly island begin to show (fig.4,5). 5. at the level of -38 m a separate pool of ural furrow disappears (fig.6). 6. at the level of -40 m the northern caspian is completely drained (fig.7). 7. a further drop in the level up to -45 and -50 m does not result in a large reduction of the water plane (fig.8). 8. a further drop in the level up to -100 m or to -150 m, does not result in a critical reduction of the water plane (fig.9,10). advances in systems science and application (2015) vol.15 no.4 329 4 discussion as it was previously shown [18], a causal interpretation of autochthonous profiles of the water biota of the caspian sea and, as a consequence, high levels of endemism of its taxa does not present a problem. a different situation arises with understanding and explaining the endemism pattern of the coastal and especially island taxa. the modern faunal or floral status of the new species for science is to be considered as endemic. the species and subspecies rank taxa were identified in this group of species. for those taxa to be formed and evolutionally stabilized sufficient time is required for diagnostically important features to appear. it is obvious that the modern natural habitat configuration of the island biota taxa of the middle caspian depends on the scale and the duration of transgressiveregressive sea cycles and the causes that predetermine the course of the events. according to the nature of these cycles the change in the species composition and structural organisation of the coastal island cenoses was taking place. the shelf area of the north caspian has a variety of numerous accumulative bench lands, shoals and islands. yet our attention is drawn to another feature in their distribution: they are found and grouped exclusively in the field of sea zone extensions of prikumskaya uplift zone and tersko-caspian trough. the same is observed in the eastern part of the northern caspian. according to o.k. leontyev [19], the concentration of alluvia is associated with the local uplifts. n.a. kasyanova [20] believes that concentration of alluvia is confined to the tectonically active currently local uplifts, proving this by the fact that along the sea extension of anticlinal zones of karpinski range there also was found a series of local uplifts and that accumulative bench lands here are not present (or their forms are few). this can be explained taking into account the features of modern geodynamics of those areas: prikumskaya uplift zone has a higher current geodynamic activity compared with that of karpinski range [20]. however, the high latest tectonic activity of the local uplifts, located within the sea extension of karpinski range proven by the power reduction of the upper levels in almost all uplifts should be noted [21]. while not rejecting the genetic connection between the accumulative bench lands and islands ye.n. badyukova et al. propose a different interpretation of the previous data and, accordingly, a new version of those accumulative forms origin. [22] she believes that the zonal arrangement of bench lands and islands is conditioned by the location of the ancient coastlines of the various stages of the new caspian transgression. some scholars proposed to reconstruct the pleistocene caspian basin based on summarizing the results of the malacofaunal analysis and the data of the comprehensive study of the caspian region sediments.t.a. yanina proposes another 330 g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev:biodiversity of island... sequence of the pleistocene events on the basis of the factual data collected over the years of field, and laboratory investigations of the key sections of the pleistocene sediments and the caspian malacofauna localities. [23]the main attention is paid to the leading for the caspian sea and endemic for the ponto-caspian brackish-water shell-fish of didacna eichw genus, which features rapid evolutionary development at the species and subspecies level played the essential role of the genus for stratification of the sea pleistocene of the caspian sea and paleogeographic reconstructions of its basins. to monitor the results of the malacofauna analysis they used a conjugated method of the latest sediments investigation and events reconstruction. the possibility of islands surface absolute altitude increase following the banked up water level [19] suggests that even during the periods of high water levels in the north caspian sea there were island lands where the process of new species formation took place. during the holocene (the last 10 thousand years), the level of the caspian sea did not rise above -20 m [23]. if it was so tyuleniy, chechen and possibly kulaly islands and the adjacent islands existed all this time since the highest points of the island rise above the present sea level by 5-8 metres and the emergence of the continental links of the western and eastern shores secured a rich representation of turanian species (up to 30-60%) in northern dagestan and the variety of reliquiae of sarykum sand dune with its typical central asian flora and fauna (to the level of subendemics). fig. 1 sea surface configuration at the level of 28,6 m. advances in systems science and application (2015) vol.15 no.4 331 fig. 2 sea surface configuration at the level of 30 m. fig. 3 sea surface configuration at the level of 32 m. 332 g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev:biodiversity of island... fig. 4 sea surface configuration at the level of 34 m. fig. 5 sea surface configuration at the level of 36 m. advances in systems science and application (2015) vol.15 no.4 333 fig. 6 sea surface configuration at the level of 38 m. fig. 7 sea surface configuration at the level of 40 m. 334 g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev:biodiversity of island... fig. 8 sea surface configuration at the level of 50 m. fig. 9 sea surface configuration at the level of 100 m. advances in systems science and application (2015) vol.15 no.4 335 fig. 10 sea surface configuration at the level of 150 m. acknowledgements the study was carried out with support of the ministry of education and science of the russian federation, agreement no. 14.574.21.0109 (a unique identifier for applied scientific researches (project) rfmefi57414x0032). references [1] berg l.s. (1934),“the level of the caspian sea in the historical time”, issues of physical geography, no.1, pp.11-64. [2] lilienberg d.a.(1994),“new approaches to the modern endodynamics evaluation of the caspian region and the issues of its monitoring ”, news of the russian academy of sciences. series geographer, no.2, pp.16-36 [3] rychagov g.i. (1997), the pleistocene history of the caspian sea , publishing house of mgu, pp.267. [4] varushenko s.i. varushenko a.n. and klige r.k.(1987),“ change in the caspian sea and stagnant basins in palaeotime”, nauka, pp.240. [5] apollov b.a. (1935),“water balance of the caspian sea and its possible changes” tr. tsiegm , issue vol.2, no.44, pp.11-18. 336 g. m.abdurakhmanov, a.a.teymurov and a.a. gadzhiyev:biodiversity of island... [6] bregman g.r. mikhalevsky a.i. (1935,“water balance of the caspian sea in connection with the great volga ”, news of the academy of sciences of azssr. physics-chemistry series.,vol.19, pp.7-138. [7] kritskiy s.n. korenistov d.v. and ratkovich d.ya, (1975),fluctuations of the caspian sea level, nauka, pp.160. [8] shlyamin b.a., the caspian sea, geografgiz, pp.128. [9] kislov a.v., panin a., toropov p.(2014),“current status and palaeostages of the caspian sea as potential evaluation tool for climate model simulations ”, quaternary international, pp.1-8 [10] bolgov m.v. filimonova m.k. (2005),“on the sources of uncertainty in forecasting the caspian sea level and evaluation of flooding risk of the coastal areas” , water resources, no.6, pp.664-669. [11] budyko m.i., yefimova i.a., v.v. lobanov (1988),“the future level of the caspian sea”, meteorology and hydrology , no.5, pp.86-94. [12] kroonenberg s.b., abdurakhmanov g.m., badyukova e.n., van der borg k., kalashnikov a., kasimov n.s., rychagov g.i., svitoch a.a., vonhof h.b.and wesselingh f.p., (2007), “solar-forced 2600 bp and little ice age highstands of the caspian sea ”, quaternary international , vol.173, pp.137143 [13] kroonenberg s.b., kasimov n.s. and lychagin, m.yu, (2008),“ the caspian sea: a natural laboratory for sea-level change”, geography, environment, sustainabilility ,vol.1, pp.22-37 [14] mikhailov v.n. “mysteries of the caspian sea”. http://journal.issep.rssi.ru/articles/pdf/0004s 063.pdf soros educational journal. [15] ratkovich d.ya., bolgov m.v. (1994),“ study the probable regularities of the long-term fluctuations in the caspian sea level”, water resources, no.6, pp.389-404. [16] shlyamin b.a.(1962),“super long-term forecast the caspian sea level”, news of vgo ,vol.94, no.1 ,pp.26. [17] abdurahmanov g.m., teymurov g.a., shokhin i.v., nabozhenko m.v., alieva s.v., eskendarova s.n. and elderkhanova z.m., (2010), “the caspian region: environmental consequences of the climate change”, moscow: faculty of geography, pp.60-64. advances in systems science and application (2015) vol.15 no.4 337 [18] abdurakhmanov g.m., karpyuk m.i., morozov b.n. and puzachenko yu.g. (2002),“current status and the factors influencing the biological and landscape diversity of the volga-caspian region of russia”, nauka,pp.416. [19] leontyev d.c.(1957),“ on the origin of some islands of the northern part of the caspian sea”, tr. oceanographer, commission of the ussr academy of sciences,vol. 2, pp.147-158 [20] kasyanova n.a.(1998), “new data on the structure and oil and gas content of the north-west caspian sea”, oil and gas geology,no.4,pp.10-16. [21] ulitsky yu.a. and turaev i.a. et al.(1967),“the main features of the structure of the upper-pliocene quaternary sediments of the north-west caspian region in connection with the finding the mesozoic structural plan features ”, structural-geomorphological investigations when studying oil and gas basins. l. pp.180-186. [22] badyukova ye.n., varushchenko a.n. solovyova g.d. (1996), “on genesis of the bottom topography of the northern caspian”, bulletin of moip. geology department. ,vol.71, no.5, pp.80-88. [23] yanina t.a., “paleogeography of the caspian sea in the pleistocene ”, geology of the oceans and seas: proceedings of the 18th international scientific conference (school) on sea geology, m: geos, vol.3, pp.355-360. corresponding author tatiana can be contacted at: yal05@mail.ru adv syst sci appl 2020; 04:125–131 published online at https://ijassa.ipu.ru. on some global properties of multivalued simple waves dmitry v. tunitsky1 1v.a. trapeznikov institute of control sciences of ras, profsoyuznaya 65, 117342 moscow, russia abstract: the cauchy problem for one-dimensional quasilinear wave equation is considered. in the case that its solutions are multivalued simple waves, we derive explicit expressions in quadratures by introducing global characteristic coordinates. the main result of the paper is a theorem on some global properties of multivalued simple waves. keywords: hyperbolic quasilinear wave equation; multivalued solution; cauchy problem; simple wave introduction a number of models in continuum mechanics concern the wave equation zxx − g2(zy)zyy = 0, (0.1) where g = g(q) is a given positive coefficient, g(q) > 0, (0.2) cf. [1], [2], and [3], and the monge notations p = zx(x, y), q = zy(x, y) (0.3) are used. the variable x designates time, y is the spatial coordinate, and z is the displacement of continuum. in particular, vibration of a string, cf. [3], [1]; wave propagation in a bar of elastic-plastic material, cf. [4]; and isentropic flows of a compressible gas with plane symmetry, cf. [4], [5], are described by equations of the type (0.1). equation (0.1) is a quasilinear second-order partial differential equation with two independent variables x and y and unknown function z = z(x, y). (0.4) due to condition (0.2) it is hyperbolic. it is common knowledge what is a classical solution of equation (0.1). classical solutions have a rather serious drawback: in finite time singularities, a so called gradient catastrophes, can develop in them, cf. [2] and [1]. the latter means that there is a point (x, y) such that in some its vicinity a classical solution z (0.4) itself and its first derivatives p and q (0.3) are bounded but at least one of second derivatives is unbounded. this is a motivation and reason to generalize a notion of classical solution and define the notion of a multivalued solution, which is a relatively well-known one, cf. [6], [7], [8], and [9]. ∗corresponding author: dtunitsky@yahoo.com 126 d.v. tunitsky 1. multivalued solutions the differential 2-form ω2 = dp ∧ dy − g(q)dx ∧ dq, (1.5) is in an obvious way associated with the left hand side of the equation (0.1), cf. [8], and the linear differential form and its exterior derivative ω0 = dz − pdx− qdy, ω1 = dω0 = dx ∧ dp+ dy ∧ dq (1.6) are associated with equalities (0.3). an immersion σ : s −→ r5 (1.7) of a two-dimensional hausdorff paracompact manifold s is a multivalued solution of equation (0.1) if it satisfies the following system of exterior differential equations σ∗ω0 = 0, σ∗ω1 = 0, σ∗ω2 = 0, (1.8) cf. [7] and [8]. obviously, the graphic of a classical solution is a multivalued solution but not vice versa. and it can be proven that any classical solution is a part of maximal multivalued solution, cf. [10]. if a gradient catastrophe takes place at a point s ∈ s, then (dσ∗x ∧ dσ∗y) (s) = 0. multivalued solutions have a serious advantage over classical solutions because gradient catastrophes do not happen to them; cf. [10]. the linear differential forms ωj,1 = dp− (−1)jg(q)dq, ωj,2 = dy − (−1)jg(q)dx (1.9) will be called characteristic. it is not hard to see that the equalities ω2 − g(p)ω1 = ω1,1 ∧ ω1,2, ω2 + g(p)ω1 = ω2,1 ∧ ω2,2, ω2 = 1 2 (ω1,1 ∧ ω1,2 + ω2,1 ∧ ω2,2), ω1 = 1 2g(p) (ω1,1 ∧ ω1,2 − ω2,1 ∧ ω2,2) are through for characteristic forms ωj,1, ωj,2 (1.9) and 2-forms ω2 (1.5) and ω1 (1.6). thence, an immersion σ (1.7) is a multivalued solution of the equation (0.1) iff σ∗ω0 = 0, σ∗(ω1,1 ∧ ω1,2) = 0, σ∗(ω2,1 ∧ ω2,2) = 0. (1.10) a curve γ : γ −→ r5, where γ is a one-dimensional connected manifold, will be called a characteristic curve of the equation (0.1) that belongs to the j-th family, j = 1, 2, if γ∗ω0 = 0, γ∗ωj,1 = 0, γ∗ωj,2 = 0. (1.11) suppose an immersion σ (1.7) is a multivalued solution of equation (0.1). then equations (1.10) are through for this immersion. hence the pull-backs σ∗ωj,1 and σ∗ωj,2 of the copyright c© 2020 assa. adv syst sci appl (2020) multivalued simple waves 127 characteristic forms ωj,1 and ωj,1 (1.9) are linearly dependent and the linear algebraic equations σ∗ωj,1 = 0, σ∗ωj,2 = 0 (1.12) uniquely define the one-dimensional subbundle of the tangent bundle ts for j = 1, 2. therefore, by the frobenius theorem [11], for any point s ∈ s and number j = 1, 2 there exists the maximal integral manifold γj,s : γj,s −→ s, (1.13) of the system γ∗j,sωj,1 = 0, γ∗j,sωj,2 = 0, (1.14) containing the point s, where γj,s is a connected one-dimensional manifold. it follows from the equations (1.11), (1.12), and (1.14) that the composition σ ◦ γj,s of the maps (1.7) and (1.13) is a characteristic curve of the equation (0.1), belonging to the j-th family. of cause, for classical solutions conventional characteristic curves are obtained this way. 2. cauchy problem an initial curve or initial value for equation (0.1) is a smooth curve l : (−∞,+∞) −→ r5 (2.15) that satisfies the conditions ω0(l̇(τ)) = 0, ω2 1,1(l̇(τ)) + ω2 1,2(l̇(τ)) 6= 0, ω2 2,1(l̇(τ)) + ω2 2,2(l̇(τ)) 6= 0 (2.16) for −∞ < τ < +∞. the last two conditions mean that the initial curve l (2.15) is not characteristic, i.e. it is free, see (1.11). a classical-type initial value problem at x = 0 for equation (0.1) is defined by conditions z(0, y) = z0(y), zx(0, y) = q0(y), (2.17) where z0 : (−∞,+∞) −→ r, q0 : (−∞,+∞) −→ r (2.18) are given functions. this initial values determine initial curve l (2.15) with the coordinates l∗x(τ) = 0, l∗y(τ) = τ, l∗p(τ) = p0(τ), l∗q(τ) = q0(τ), l∗z(τ) = z0(τ). (2.19) obviously, this curve is an immersion and meets conditions (2.16) since ω2 1,1(l̇(τ)) = ω2 2,2(l̇(τ)) = 1 copyright c© 2020 assa. adv syst sci appl (2020) 128 d.v. tunitsky according to expressions (1.9). if for a multivalued solution (1.7) of equation (0.1) there exists an imbedding l : (−∞,+∞) −→ s (2.20) such that l = σ ◦ l, (2.21) then σ is called a solution of cauchy problem (0.1), (2.15) and l – an initial embedding. thus a solution of the cauchy problem (0.1), (2.15) is a pair (σ, l) of an immersion σ (1.7), satisfying the system (1.10), and imbedding l (2.20), satisfying (2.21). a multivalued solution (σ, l) of the initial value problem (0.1), (2.15) is said to be definite if for any point s ∈ s and number j = 1, 2 the intersection γj,s(γj,s) ∩ l(−∞,+∞) of the image of initial imbedding l (2.20) and characteristic curve γj,s (1.13) consists of exactly one point. theorem 2.1: (characteristic uniformization, [10]) let (σ, l) be a definite solution (1.7), (2.20) of the cauchy problem (0.1), (2.15)–(2.16). then there exists a unique diffeomorphism φ : s −→ φ(s) (2.22) such that the following three properties hold. (a) the image φ(s) is a subset of r2 and contains its diagonal δ = {(τ, τ) ∈ r2| −∞ < τ < +∞}. (b) for −∞ < τ < +∞ the initial imbedding l (2.20) satisfies the equality φ ◦ l(τ) = (τ, τ). (c) images φ ◦ γj,s(τ) of characteristic curves γj,s (1.13), where s ∈ s, lie in coordinate straight lines u = const if j = 1 and v = const if j = 2 for τ ∈ γj,s. a diffeomorphism φ (2.22), which is uniquely defined by theorem 2.1, is called a characteristic uniformization of the definite solution (σ, l). the coordinate plane of the parameters u, v and the same coordinates in it are called characteristic as well. the image φ(s) of the uniformization φ (2.22) is a uniformized domain and the composition ψ = σ ◦ φ−1 (2.23) is a uniformized solution. 3. simple waves a solution σ (1.7) is said to be a simple wave, if the equality dσ∗p ∧ dσ∗q = 0 (3.24) holds for it, cf. [4]. it is possible to show that a definite solution (σ, l) of the cauchy problem (0.1), (2.17) is a simple wave iff either p0(y) = −g(z′0(y)), (3.25) or p0(y) = g(z′0(y)), (3.26) where g = g(q) is a primitive function of g = g(q). copyright c© 2020 assa. adv syst sci appl (2020) multivalued simple waves 129 by definition (3.24), in case of simple waves linearization of equation (0.1) by means of hodograph transformation is not applicable. but it is possible to apply theorem 2.1 and use characteristic uniformization φ (2.22) in this case instead. indeed, by theorem 2.1, definition (1.13)–(1.14) of characteristic curve, and initial values (3.25) and (3.26) we get for coordinates x, y, p, q, and z of uniformized solution (2.23) the following equations ∂ ∂v y ψ∗ω1,1 = pv + g(p)qv = 0, ∂ ∂v y ψ∗ω1,2 = yv + g(p)xv = 0, ∂ ∂v y ψ∗ω0 = zv − pxv − qyv, ∂ ∂u y ψ∗ω2,1 = pu − g(q)qu = 0, ∂ ∂u y ψ∗ω2,2 = yu − g(p)xu = 0, ∂ ∂u y ψ∗ω0 = zu − pxu − qyu (3.27) where the symbol y denotes the interior multiplication, see [11] (section 2.11), and the initial values z0(τ, τ) = z0(τ), q0(τ, τ) = z′0(τ), p0(τ, τ) = −g(z′0(τ)) (3.28) in the case of (3.25) and z0(τ, τ) = z0(τ)), q0(τ, τ) = z′0(τ)), p0(τ, τ) = g(z′0(τ)) (3.29) in the case of (3.26). integration of cauchy problems (3.27), (3.28) and (3.27), (3.29) allows to represent characteristic uniformization (2.23) of a simple wave in quadratures. proposition 3.1: in characteristic coordinates u and v the coordinates x, y, p, q, and z of uniformization (2.23) of any definite multivalued solution (σ, l) of the cauchy problem (0.1), (2.17) that is a maximal simple wave are representated in the following way: x(u, v) = h(u, v) 2 √ g0(v) , y(u, v) = v + √ g0(v) 2 h(u, v), z(u, v) = z0(v) + p0(v) + g0(v)z′0(v) 2 √ g0(v) h(u, v), p(u, v) = p0(v), q(u, v) = z′0(v) (3.30) copyright c© 2020 assa. adv syst sci appl (2020) 130 d.v. tunitsky in case (3.25) and x(u, v) = h(u, v) 2 √ g0(u) , y(u, v) = u− √ g0(u) 2 h(u, v), z(u, v) = z0(u) + p0(u) + g0(u)z′0(u) 2 √ g0(u) h(u, v), p(u, v) = p0(u), q(u, v) = z′0(u) (3.31) in case (3.26). here (u, v) ∈ r2, g0(τ) = g(z′0(τ)), h(u, v) = u∫ v dτ√ g0(τ) . put π : r5 3 (x, y, p, q, z) 7−→ (x, y) ∈ r2. the following statement is an immediate corollary of representations (3.30) and (3.31) from proposition 3.1. it gives sufficient conditions for projection π ◦ σ of a maximal simple wave (σ, l) to be a proper map of degree 1. theorem 3.1: if lim τ→∞ g0(τ) τ = 0, lim τ→∞ |h(τ, 0)| = +∞, then the projection π ◦ σ of the maximal simple wave (σ, l) is a proper map of degree 1 and, therefore, π ◦ σ(r2) = r2. acknowledgements the reported study was funded by rfbr and jsps (project number 1951-50005) and rfbr (project number 20-01-00610). references 1. n.j. zabusky, ”exact solution for the vibrations of a nonlinear continuous model string”, j. math. phys. vol. 3, no. 5, pp.1028–1039, 1962. 2. p.d. lax, ”development of singularities of solutions of nonlinear hyperbolic partial differential equations”, j. math. phys. vol. 5, no. 5, pp.611–613, 1964. 3. m. taylor, partial differential equations, vol. 1, basic theory, springer-verlag, new york, 1996. 4. r. courant and k.o. friedrichs, supersonic flow and shock waves, interscience publishers, inc., new york, 1948. 5. b. l. rozhdestvenskii and n. n. janenko, systems of quasilinear equations and their applications to gas dynamics. transl. math. monogr., vol. 55, amer. math. soc., providence, ri, 1983. copyright c© 2020 assa. adv syst sci appl (2020) multivalued simple waves 131 6. é. cartan, les systèmes différentiels extérieurs et leur applications géométriques, hermann, paris, 1945. 7. a. m. vinogradov, ”multivalued solutions and a principle of classification of nonlinear differential equations”, soviet math. dokl., vol. 14, pp.661–665, 1973. 8. a. kushner, v. lychagin, and v. rubtsov, contact geometry and nonlinear differential equations, encyclopedia math. appl., 101, cambridge univ. press, cambridge, 2007. 9. r.l. bryant, s.s. chern, r.b. gardner, h.l. goldschmidt, and p.a. griffiths, exterior differential systems, springer-verlag, new york–berlin–heidelberg–london–paris– tokyo–hong kong–barcelona, 1991. 10. d.v. tunitsky, ”on hyperbolic systems of monge-ampère equations”, sbornik: mathematics, vol. 200, no. 11, 2009, pp. 1681–1714. 11. f. w. warner, foundations of differentiable manifolds and lie groups, springer-verlag, new york–berlin–heidelberg–tokyo,1983. copyright c© 2020 assa. adv syst sci appl (2020) multivalued solutions cauchy problem simple waves microsoft word article final.doc adv syst sci appl 2025; 2; 1-10 published online at https://ijassa.ipu.ru. computer simulation of a magnetic separator for iron ore dressing under conditions of uncontrolled disturbances nina osipova* university of science and technology misis, moscow, russia financial university under the government of the russian federation, moscow, russia abstract: the article presents a theoretical material about the principle of operation, design and advantages of the magnetic iron ore separator. a review of researches is given, where special attention is paid to the study of factors that affect the quality of concentrate and the loss of a valuable component in the tailings. the specification of the pdm-sc-120/300 magnetic separator is given. regression models are obtained that reflect the relationship between the content of the class -0.074 mm and the solid phase flow rate in the separator feed with a high coefficient of determination. the statistical significance of this coefficient was verified using the fisher criterion. a mathematical description of the object elements such as governing valve, motor, and magnetic separator is given. these models are simplified by dropping small time constants and linearization in the vicinity of nominal modes using taylor series expansion. the main disturbances that cause the deviation of the dressing indicators from the optimal values are highlighted. a computer model of the magnetic separator was constructed using the simulink matlab application. the simulation results showed that with the specified control actions, an increase in the content of the class -0.074 mm in the feed to 94-95 % brings the mass fraction of iron in the concentrate to the acceptable limits but reduces the productivity of the separator. keywords: magnetic separator, concentrate, tails, matlab, ms excel, regression model, fisher criterion, correlation, coefficient of determination, taylor series. 1. introduction over the past century, the constant growth of human needs for iron has led to the development and improvement of new technologies for ore dressing. the most widespread process is the magnetic separation. its main purpose is to separate ore material particles into two phases: concentrate with a high iron content and tailings with a small fraction of the valuable component. by design the magnetic separator is made in the form of a drum that rotates at the given speed and a stationary magnetic system located in its inner part. separation of the ore stream is based on the magnetic susceptibility of its constituent particles. if this parameter has a high value, then the particles stick to the drum and enter the concentrate compartment, and the rest, without attracting, are immediately washed off into the tailings compartment. the advantage of the magnetic dressing method is the ability to create a high force of attraction to the drum, which is hundreds of times higher than the particles weight, as well as safety during maintenance of the separator and harmlessness to the environment [5]. there are dry and wet magnetic dressing. in the first case, the ore is fed to the separator after preliminary crushing and screening. in the second case, after grinding in the mill and classification by size, the ore enters the magnetic separator in a mixture with a liquid (pulp). the above processes are multi-stage. * corresponding author: nvo86@mail.ru 2 n. v. osipova copyright ©2025 assa adv. in systems science and appl. (2025) the iron ore raw material coming to the metallurgical plant in the form of agglomerate or pellets must be of the specified quality with permissible deviations from the norm established by technical conditions and regulations. otherwise, you have to adjust the modes of melting units or spend more additional materials, which increases the cost of steel production. it is also necessary to ensure that iron losses in the tailings that do not exceed the permissible value. 2. review of the research the main purpose of the research in the study of mineral processing is to identify the main factors that affect the iron content in the products of the magnetic separator, including controlling and disturbing influences. control actions can be changed by a human operator or by an automatic system. disturbing influences are random changes in the physical and mechanical properties of the ore or pulp. several researches have been devoted to the study of the influence of controlling and disturbing influences on the magnetic separation process. in [1], multiple regression equations for a specific type of ore are obtained. one of them relates the iron content in the concentrate and the control variables: the filling level of the mill at the second stage of grinding, the density of the hydrocyclone discharge. another is the association of loss of iron in the tails of the first stage of dressing and fill level of the mill, the water flow into the mill the first stage of grinding, drain density classifier. the coursebook [10] describes the principle of creating a regression model based on experimental planning, which can predict the mass fraction of iron in the concentrate for a predetermined step forward in time. this model is based on the calculated average iron content for a certain period, the deviation of the current value of the iron content from the average, and changes in the load on the ore section. the paper [4] presents a regression model, in which the dependent variable is the iron mass fraction in the enrichment products, and the factors are the load on the industrial product of dry magnetic separation, water flow into the classifying apparatus and magnetic separators at the 1st, 2nd, 4th stages of dressing. the disadvantage of these methods is a large delay between obtaining of input and output variables, which makes it difficult to quickly update the coefficients of the regression model and calculate the control variables. the above papers do not specify what specific disturbances can cause deviations of the dressing indicators from the set ones. in [6], static characteristics of magnetic separators are given, reflecting the dependence of the iron content in the concentrate and its losses in the tailings on the drum rotation speed and pulp density, which can be controlled by supplying additional water to the separator bath. however, the graphs do not indicate the numerical values on the axes and the equations that they were based on. it is only known that in a wide range of parameters, the dependencies are nonlinear, close to the second-order polynomial. therefore, it is necessary to solve the following problems: to select the control and disturbing effects that affect the content of the useful component in the concentrate and in the tailings, so that there is no lag between the input and output of the object; to develop a computer model of the magnetic separator, which allows to study its operation of the separator under nominal control effects under conditions of uncontrolled disturbances. 3. object of research the object of research is the drum semi-countercurrent magnetic separator pdm-sc-120/300, which is used in many iron ore mining and processing plants to separate particles of less than 1 mm in size at the final stages of dressing. it’s required for simulation specification is given in table 3.1 [2]. computer simulation of a magnetic separator for iron ore enrichment 3 copyright ©2025 assa. adv. in systems science and appl. (2025) table 3.1. specification of the pdm-sc-120/300 magnetic separator feed option drum rotation speed, min-1 permissible productivity for solid phase, tons/h content in the feed separator, % class -0.074 mm solid phase of the pulp 1 60-70 30 140-180 2 20 80-120 3 75-85 30 100-140 4 20 70-100 5 94-96 30 60-80 6 20 40-60 as you can see from the table, there are six separator feed options. in this case, the disturbing effects are the content in the feed of the class -0.074 mm and the of solid phase flow rate (productivity) at two limit values of the solid phase content in the pulp. before finding the degree of influence of factors on the dressing indicators, it is necessary to check their correlation. for this purpose, ms excel generated samples of random variables with a normal distribution law that characterize the content of the class -0.074 mm and the corresponding solid phase flow rate for two values of the solid phase content in the separator feed (table 3.2). table 3.2. characteristics of random values of disturbances for simulation in ms excel solid phase content in pulp, % mathematical expectation maximum deviation (3σ) standard deviation (σ) performance, tons/h the content of the class -0.074 mm, % performance, tons/h the content of the class -0.074 mm, % performance, tons/h the content of the class -0.074 mm, % 20 100 65 20 5 20/3 5/3 30 160 20 85 80 15 15/3 30 120 20 20/3 20 50 95 10 1 10/3 1/3 30 70 for each of the above six options, the sample consists of 10 values, i.e. 30 positions for the solid phase content in the feed, equal to 20 % and 30 %. the diagram with the trend line based on the linear dependence is presented and the coefficient of determination is calculated (fig. 3.1). fig. 3.1. correlation diagram of solid phase flow rate and -0.074 mm class content in the separator feed the high determination coefficients r1 2 = 0,9134 и r2 2 = 0,9767 indicate a strong dependence of the solid phase flow rate on the content of the class -0.074 mm in the 4 n. v. osipova copyright ©2025 assa adv. in systems science and appl. (2025) separator feed. let’s evaluate the statistical significance of the coefficients using the fischer criterion. in the first case, we find the observed f-value for two dependencies: 2 1 1 2 1 2 2 2 2 2 0.9134 28 38.55, 1 1 0.9134 1 0.9767 28 1173.72, 1 1 0.9767 1 r f f r m r f f r m           (3.1) where n = 30 is the number of experimental points; m = 1 is the number of factors; f = n-m-1 is the number of degrees of freedom; α = 0.05 is the level of significance. it turns out that the found value is greater than the critical value indicated in the table fcr = 4.2 (f1 > 4.2; f2 > 4.2), which confirms the statistical significance of r2 with a probability p =1 α = 0.95. therefore, if there is a correlation in the mathematical model, it is sufficient to use just one of the disturbances: the solid flow rate or the content of the class -0.074 mm in the feed of the separator. 4. computer model of a magnetic separator 4.1 nonlinear dynamic model of a magnetic separator as control variables, we will use the water flow rate into the separator bath and the rotation speed of its drum. this choice is justified by the absence of a lag between these impacts and the output indicators, as well as the fact that there is no need to change the separator design. the model block diagram is shown in figure 4.1. when controlling the degree of the valve opening and the motor shaft rotation speed by setting different values εs and ωs, the water flow rate into the bath wb and the rotation speed of the drum ω, respectively, are regulated. it causes a change in the content of magnetite iron in the concentrate β and tailings ν monitored by analyzers. the parameter ξ is a disturbance that characterizes the instability of the physical and mechanical properties of the pulp, which also causes fluctuations in the dressing indicators β and ν. fig. 4.1. block diagram of the magnetic separator model let’s look at the individual components of the model in detail. the model of the governing valve is described by the following differential equation: v v v s ( ) ( ) , d t t t dt      (4.1) computer simulation of a magnetic separator for iron ore enrichment 5 copyright ©2025 assa. adv. in systems science and appl. (2025) εs is the set value for the degree of opening of the valve, %; εv is the current value of the valve opening degree, %; tv is the time constant of the controlled valve, sec. for the most valves, tv is approximately 0.3 seс [7]. the water flow rate in the separator bath is related to the value εv and the ratio [9]: b v( ) 11 ( ).w t t  (4.2) the model of an asynchronous drive motor has a very complex mathematical description in the form of a system of nonlinear equations. however, in the case of operation in a small neighborhood of the point that corresponds to the rated mode, it is allowed to consider a linearized system equivalent to the dc motor model: m m m s ( ) ( ) , d t t t dt    (4.3) ωs is the set value of the motor drive shaft rotation speed, min-1; ωm is the current value of the motor drive shaft rotation speed, min-1; tm is the time constant of the drive motor, s. the model of the gearbox is determined based on the ratio of the nominal rotation speed of the separator drum ωs.nom = 19 min–1 and the motor ωm.nom = 1000 min –1: s.nom m.nom ω 19 0.02. ω 1000g k    (4.4) the rotation speed of the separator drum ωg using a reducer is equal to: mgω ( ) 0.0 .2ω ( )t t (4.5) according to the simulation of asynchronous motors the value of tm can be assumed to be equal to 0.03 seс [9]. the models must have "dead zones" δv and δm for the valve and drive, respectively. this means that if the input signal is |εv|≤ δv or |ωg|≤ δm, then the output signal is ε = 0, ω = 0, otherwise ε = εv δv or ω = ωg – δm. let’s take δv = 3 %, δm = 0,1 %. the dynamics of the content of magnetite iron in the concentrate and tailings is governed by the following system of differential equations: β β β β β ν ν ξ ξ v v ν β 1 1 1 β ,ω ( ) (ξ) ( ) ( , ν 1 1 1 ν ,ω ξ), d f w f dt t t t d f w f dt t t t      (4.6) where w = wl + wb is the sum of the liquid phase consumption in the pulp and the water consumption in the separator bath, fβ(w, ω) и fν(w, ω) describe the relationship of the dressing indices β, ν and the control variables w, ω and the control variables w, ω in steady state when βd dt = 0 и ν d dt = 0: 2 2 β 2 2 ν , ω – 0.0000042 0.0082 – 0.023ω 1.21ω 46.5, , ω – 0.0000025 0.0087 0.035ω – 1.3ω 10.2, ( ) ( ) f w w w f w w w         (4.7) where tβ, tv are the time constants of magnetic separator (sec), tβ = 1…10 sec, tv = 1...10 seс [7]; fβξ(ξ), fνξ(ξ) are the functions on uncontrolled disturbances ξ. to simplify the simulation, we put tβ = tv = 1 sec. the system (4.7) is obtained in [9] based on the reference data of the average concentration of iron ore mining and processing plants, provided that fβξ(ξ) = const, fνξ(ξ) = const. the units of measurement for w and ω are m3/h and min-1 respectively. 6 n. v. osipova copyright ©2025 assa adv. in systems science and appl. (2025) 4.2 linearization of the magnetic separator model for the convenience of research, simulation and automation of the magnetic separator, we linearize equations (4.6). the most of automatic systems are designed and configured under the assumption that the object under control is linear. using the rule for calculating the function increment δy = y'δx, we perform the system linearization by expanding the taylor series of equations (4.6) in the vicinity of the nominal modes (wnom, ωnom) with respect to δβ and δν with the rejection of nonlinear terms. given that tβ = tv = 1 seс and that fβ(w, ω), fν(w, ω) are found by equation (4.7), we obtain: nom nom nom nom nom nom nom nom β β – 2 0.0000042 0.0082 – 2 0.023ω ω 1.21 ω= = β – 0.0000084 0.0082 – 0.046ω ω 1.21 ω, ν ν – 2 0.0000025 0.0087 2 0.035ω ω-1.3 ω ν – 0.000005 0.0087 0.07ω ω – w w w w w w w w w w w w                                        1.3 ω. (4.8) for simulation, we assume that the required content of magnetite iron in the concentrate βs = 63.8 %, which is the average for many iron ore mining and processing plants [2]. substituting ωnom = 19 min-1 in the first equation in (4.7) for fβ(wnom, ωnom) = βs, we get wnom = 400 m3/h. then the loss of magnetite iron in the tailings at wnom, ωnom is fν(wnom, ωnom) = νs = 1.22 %. assuming the time constants of the valve and motor are negligible in compare with the time constants of the magnetic separator, the water flow rate into the separator bath is associated with a given degree of opening of the valve as w = 11εs, and the rotation speed of the drum is determined by the equality ω = ωs. we introduce new notation for the variables δβ = x1, δν = x2, δw = u1, δω = u2. we assume that the change δw is due to the control of the valve. is due to the control of the valve. taking into account the above, substituting wnom, ωnom in equation (4.8), we get: 1 1 1 2 2 2 1 2 0.00484 0.336 , 0.0067 0.03 x u u x u x x u          (4.9) or in matrix form: ( ) ,( ) ( )t x tx ta bu  (4.10) where: 1 0 0.0048 0.336 , . 0 1 0.0067 0.03 a b             (4.11) the linear representation of the system (4.7) in the interval (wnom, ωnom) has the form:         β nom nom βlin β nom nom nom β nom nom nom ν nom nom νlin ν nom nom nom ν nom nom nom , ω ) , ω , ω ) , ω ) ω ω 55.48 0.00484 0.336ω, ω , ω ) , ω , ω ) , ω ) ω ω 2.035 0.0067 0. ( ( ) ( ( ( ( ) ( 03ω. ( ω f w f w f w w w w f w w f w f w f w w w w f w w                            (4.12) as a disturbance, we will consider the change in the percentage q of solid and the solid phase flow rate in the feed q. as can be seen from table 4.1, the separator maximum performance is q = 180 tons/h with a percentage of solid q = 30 %, and the minimum q = 40 tons/h with a solid q = 20 %. let’s calculate the minimum and maximum values of the liquid phase flow rate in the pulp, taking into account that the water density ρw = 1 ton/m3: computer simulation of a magnetic separator for iron ore enrichment 7 copyright ©2025 assa. adv. in systems science and appl. (2025)       l w 3 l.min 3 3 l.max 3 100 % , 100 % 30 % 40 tons/h 93.3 m /h, 1 ton/m 30 % 100 % 20 % 180 tons/h 720 m /h. 1 ton/m 20 % q w q q w w                (4.13) substituting w = wb + wl in (4.12) and replacing the liquid phase flow rate of the with the expression from equation (4.13), we get:     ξ ξ βlin b β νlin b ν β l ν ξ w ξ w l ( ) ( , ) ( ) ( , ω 55.48 0.00484 0.336ω+ , , ω 2.035 0.0067 0.03ω+ 0.00484 , ), 100 ( , ) , 100 ( , ) 0.00484 0.0067 0.0067 . f w w f f w w f f q q q q q q q q q q q q q q w f w                    (4.14) according to the specifications, the permissible deviation for the content of magnetite in the concentrate x1,2max ≤ 1 %. then βmin = 63.8 % 1 % = 62.8 %, βmax = 63.8 % + 1 % = 64.8 %. obviously, the minimum losses in the tailings are 0. their limit is νlim = 1.22 % + 1 % = 2.22 %. then the deviation x2max = 1 %. consider the procedure for determining the upper limits on control actions. the main requirement for the linear representation of the model (4.7) is that at the boundary values of ±u1,2max, the absolute value of the difference in the dressing indicators calculated from the nonlinear and linearized models should be equal to the error of their measurement. according to russian standard, it is 0.9 % for magnetite iron in concentrate and 0.3 % for losses in tailings during laboratory research by chemical analysis [3]. let the error of the analyzers also be equal to these values, then restrictions can be found from the solution of the system: β nom 1max nom 2max βlin nom 1max nom 2max ν nom 1max nom 2max νlin nom 1max nom 2max ( ) ( ) ( , ω + , ω + 0.9, , ω + , ω + 0) ) 3.( . f w u u f w u u f w u u f w u u          (4.15) the expressions fβlin(wnom + u1max, ωnom + u2max), fνlin (wnom + u1max, ωnom + u2max) are found by substitution in (4.12) wnom + u1max, ωnom + u2max instead of w, ω: β nom nom β nom nom βlin nom 1max nom 2max β nom nom 1max 2max ν nom nom ν nom nom νlin nom 1max nom 2max nom nom 1max 2max , ω ) , ω ) , ω + , ω ) , ω , ω ) , ω ) , ω + ( ( ( ) ( ( ( , ω ) . ω ( ) ( f w f w f w u u f w u u w f w f w f w u u f w u u w                 (4.16) the system (4.16) is easily solved using the matlab computing package of the optimization toolbox application. in this case, among the real roots, u1max = ± 346.83 m3/h, u2max = ± 4.14 min-1. so wnmin = 400 m3/h -346.83 m3/h = 53.2 m3/h, wnmax = 400 m3/h + 346.83 m3/h = 746.8 m3/h, ωnmin = 19 min-1 4.14 min-1 = 14.86 min-1, ωnmax = 19 min-1 + 4.14 min-1 = 23.14 min-1, the following restrictions must be met: wnmin ≤ w ≤ wnmax and ωnmin ≤ ω ≤ ωnmax. the nominal value of degree of opening of the valve εnom = 36.36 %, which corresponds to w = 400 m3/h. we will assume that the data from the analyzers is transmitted with a very small delay, the voltage at the output of their microchip is related to the physical value through a coefficient equal to one. therefore, we will not take them into account in the model. 8 n. v. osipova copyright ©2025 assa adv. in systems science and appl. (2025) the diagram of the computer model of the magnetic separator is shown in fig. 4.2. it was also built using the matlab environment of the simulink application based on the block diagram from fig. 4.1. fig. 4.2. diagram of a computer model of a magnetic separator in matlab simulink 4.3 simulation result the simulation was performed as follows. disturbances were periodically fed to the input of the object at an interval equal to the end time of transients of the magnetic separator tend = 5 seс [8]: the change in the percentage of solid in the feed and the solid phase flow rate, which was set using a random number generator in the range according to table 4.1, i.e., six variants of pulp properties were modeled. let’s take the first option as an example. when the solid content in the feed is 30 %, the solid phase flow rate increases to 120 tons/h. the model calculates the water consumption in the pulp using the equation (4.13), which becomes equal to wl = 280 m3/h, which is within the range of wlmin ≤ wl ≤ wlmin. the change in this signal is fed to the magnetic separator model (4.14) and used in the functions fβξ(q, q), fνξ(q, q) (fig. 4.3). fig. 4.3. changing the liquid and solid phases flow rate in the pulp computer simulation of a magnetic separator for iron ore enrichment 9 copyright ©2025 assa. adv. in systems science and appl. (2025) fig. 4.4. the content of magnetite iron in the concentrate and tailings when the properties of the pulp are unstable in the absence of automatic control of the separator, constant nominal values of water flow rate wb = 400 m3/h and rotation speed ω =19 min-1. the following conditions are met: wnmin ≤ w ≤ wnmax. the content of magnetite iron in the concentrate in the time interval of 520 seс is greater than the value of βmax. the same can be said for tailings that are superior to the νlim. the largest of the bursts is β = 65.34 % и ν = 3.35 % (fig. 4.4). this corresponds to the content of the class -0.074 mm in the feed of 60-70 %. only after t = 20 sec, the iron content falls below βmax, when the pulp is 94-95 % of the class -0.074 mm, but there is a sharp decrease in the productivity of the separator, which becomes below 80 tons/h. 5. conclusion as control actions were selected the rotation speed of the magnetic separator drum and the water flow rate into its bath. the nominal operating modes corresponding to the specified dressing parameters in the absence of disturbances are determined by the content and the solid phase flow rate in the feed, which correlates with the content of the class -0.074 mm in the feed, are taken as the parameters. a computer linear dynamic model of the separator in the matlab simulink package is constructed. based on the results of the simulation, the following conclusions can be drawn. setting control actions with constant nominal values is not enough to stabilize the quality of iron ore concentrate. when changing the content of the class -0.074 mm in the feed in the range from 60 % to 85 %, the separator operates with a sufficiently high performance, but the iron content in the concentrate is higher than normal. the same can be said about tailings. increasing the content of the class -0.074 mm in the feed to 94-95 % brings the mass fraction of iron in the concentrate to the acceptable limits but reduces the productivity of the separator. therefore, controlling the grinding and classification processes before magnetic separation in order to stabilize the granulometric composition is not promising. 10 n. v. osipova copyright ©2025 assa adv. in systems science and appl. (2025) a more rational approach is to create an optimal control system. it should be based on the formulated criterion for minimizing deviations of the iron content in the concentrate and tailings relative to the set values with restrictions on the control variables selected based on the permissible linearization error at the boundaries of their change intervals. in the future, the model of optimal control of the magnetic separator will allow evaluating and predicting the possibility of maintaining the quality of iron ore concentrate with acceptable deviations. this will reduce the number of experiments in laboratory and industrial installations and thus increase their service life and reduce the number of consumables. references 1. anikin, a. i. (1984). research and development of an optimal control system for a section of a self-grinding magneto-concentrating plant as a subsystem of the automated control system of the lebedinsky mpp kma. ph.d. thesis, leningrad order of lenin, order of the october revolution and order of the red banner of labor plekhanov mining institute 2. bogdanov, o. s. (1984) spravochnik po obogashcheniyu rud. obogatitel'nye fabriki (tom 4) [handbook of ore dressing. processing plants (volume 4)]. moscow, ussr: nedra, [in russian]. 3. gost 16589-86 (1987) rudy zheleznye tipa zhelezistyh kvarcitov. metod opredeleniya zheleza magnetita [iron ores of the ferruginous quartzite type. method for determining iron magnetite]. moscow, ussr [in russian]. 4. kazakov y. m. (1994) control of the wet magnetic dressing process line. ph.d. thesis, ural state mining and geological academy 5. karmazin, v. v., (2013). problemy i perspektivy magnitnogo obogashcheniya [problems and prospects of magnetic dressing]. gornyj informacionno-analiticheskij byulleten', s1, 560–575, [in russian] 6. maryuta a. n., kachan yu. g., bun'ko v. a. (1983) avtomaticheskoe upravlenie tekhnologicheskimi processami obogatitel'nyh fabrik [automatic management of technological processes in concentrating plants] moscow, ussr: nedra, [in russian]. 7. osipova n. v. (2018). ispol'zovanie fil'tra kalmana pri avtomaticheskom kontrole pokazatelej magnitnogo obogashcheniya zheleznyh rud [the use of kalman filter in automatic control of indicators of iron ores magnetic concentration], izvestiya visshikh uchebnykh zavedenii. chernaya metallurgiya = izvestiya. ferrous metallurgy, vol. 61, iss. 5, 372-377. doi: https://doi.org/10.17073/0368-0797-2018-5-372-377 8. osipova n. v. (2018). sintez asimptoticheskogo nablyudatelya dlya sistemy upravleniya magnitnym separatorom pri obogashchenii zheleznoj rudy [synthesis of asymptotic observer for magnetic separator control in iron ore processing] gornyj informacionnoanaliticheskij byulleten', no 6, 153-160, [in russian] doi: 10.25018/0236-1493-2018-60-153-160 9. osipova n. v. (2018). model of stabilization of the quality of iron-ore concentrate in the process of magnetic separation with the use of extreme regulation, metallurgist, vol. 62, nos. 3-4, 303-309. doi 10.1007/s11015-018-0660-8 10. vinogradov, s. v. (1984) avtomatizaciya tekhnologicheskih processov gornogo proizvodstva [automation of technological processes of mining production]. moscow, ussr: nedra, [in russian]. microsoft word 9-linhai zhao.doc advances in systems science and applications (2010), vol.10, no.1 55-60  issn 1078-6236 international institute for general systems studies, inc. study on determinants of the abnormal returns of a-shares on the first listed day linhai zhao and ridong hu academy of quantitative economics, huaqiao university, quanzhou 362021, china email: longmen168@126.com abstract on the background of the reform of split share structure, applying weighted least squares method of multiple linear regression model, 206 a-shares issued and listed from january 1st,2004 to december 31st,2007 is empirically studied, and it is found that turnover rate on the first listed day, size of money raised, and market index,have significant impact on the abnormal returns of a-shares on the first listed day, whereas, the influence of issuing price, issuing p/e ratio, net asset per share before issuance, real negotiable share ratio after issuance, method of issuance and the reform of split share structure, is not significant. keywords abnormal returns ipo underpricing reform of split share structure 1. introduction in the stock markets throughout the world there is an obvious phenomenon of the ipo underpricing, that is, shares after an initial public offering will have high abnormal rate of returns on the first listed day[1]. in spite of the complexity of china's capital market[2], the regulators have carried out drastic reform and the market-oriented objective is approaching step by step. the method of issuance and the mode of pricing are changing, meanwhile the split share structure reform are putting forward and the regime of new stock issues is about to change. on the background of the great institutional innovation, the study of the determinants of the abnormal returns of a-shares on the first listed day will help understand the determinants of stock prices and the mechanism of changes in the prices, promote china's reform of the regime of the issue and the pricing of the new stocks, and is of practical significance to china’s stock market to play the reasonable and adequate role. 2. research method 2.1 research assumptions in view of the big going-up margin in the prices of china's stocks during their initial public offering, the paper made the following assumptions: first, the subscription costs of new shares are ignored. the subscription costs include the opportunity costs of the funds used for subscription and the subscription expenses incurred. due to the inaccuracy of the measurement of the opportunity costs, and the very small ratio of the subscription expenses to the funds used for subscription, they are all ignored here. second, transaction costs are not included. the transaction cost of stock mainly comprises the commission charged by stock brokers and the stamp duty. though the rate of stamp duty was increased to 3 ‰ on may 30th, 2007, the transaction costs accounted for the amount of the investment funds only about one percent. third, the won rate of the issuance is not considered. the won rate is often closely linked with the method of issuance. because of the seriously low won rates under each of the methods of issuance, they are not considered hereby. 2.2 sample selected this paper examined the a-shares listed during the period from january 1st, 2004 to zhao: study on determinants of the abnormal returns of a-shares on the first listed day 56 december 31st, 2007 in shenzhen stock exchange and shanghai stock exchange. all the data came from the "ccer capital market research database". the shares with incomplete data or not enough observations are excluded. in the sample period, there are 300 stocks newly listed in shenzhen and shanghai. 206 stocks in the 300 stocks meet the above requirement, of which 59 are in shanghai and 147 are in shenzhen. this study used a relatively large sample carefully selected, experienced the important process of the market-oriented evolution of china's securities market, and can more comprehensively and accurately reflect actual level of the ipo abnormal returns of the china's a-share market and the changes in the process. excel and spss 11.5 are applied to conduct the statistical analysis [3,4]. 2.3 the measurement of the ipo abnormal returns to measure it, the following formula is used: 1 0 1 0(ln ln ) (ln ln )j j j j jr p p m m= − − − (1) where, jr stands for the ipo abnormal return of the j-th stock adjusted by the market return, 1 jp stands for the closing price of the j-th stock on the first listed day, 0 jp stands for the issuing price of the j-th stock, and 1 jm and 0 jm are the closing market indices on the first listed day of the j-th stock and on the issuing day of this stock respectively. 2.4 the hypotheses to be examined through analysis, there are 8 hypotheses to be examined: hypothesis 1: the methods of issuance influence the abnormal returns on the first listed day. the methods of issuance of the sampled stocks are categorized as following: issuance through the trading network with fixed price, issuance off the trading network by inquiry, the secondary market placement, corporate placement (including the strategic placement, placement off the trading network and private placement). hereby, issuance through the trading network with fixed price and issuance off the trading network by inquiry are classified by the methods of pricing; the secondary market placement and corporate placement are classified by the types of investors. these methods of issuance are often combined. different method of issuance may affect the ipo abnormal returns differently. hypothesis 2: the issuing price is negatively related to the abnormal returns on the first listed day. for the stocks with lower issuing price, the rising intervals of their prices are larger and the magnitudes of the rise are larger. hypothesis 3: the issuing p/e ratio is related to the abnormal returns on the first listed day. hypothesis 4: the negotiable share ratio positively correlates with the abnormal returns on the first listed day[5,6]. for firms with lower negotiable share ratio, their risks of corporate governance tend to be larger, and the probability of mergers and acquisitions is bigger after the issuance. this may lead to the smaller magnitude of the rise of the price of the stock on the first listed day. hypothesis 5: the turnover rate on the first listed day and net asset per share before issuance positively correlates with the abnormal returns on the first listed day. the larger the turnover rate and the more the net asset per share of the newly issued stock, the more investors are in the secondary market for this stock. that too many investors rush for the stock will push the price upward. hypothesis 6: the size of money raised negatively correlates with the abnormal returns on the first listed day. the smaller size of money raised, the more speculators buy the stock and the higher of the abnormal returns on the first listed day. hypothesis 7: the market index positively correlates with the abnormal returns on the first listed day. the higher market index shows that the market is in better condition and investors are more confident. this will heighten the trading price and lead to higher abnormal returns on the first listed day. hypothesis 8: listed before or after the reform of split share structure is related to the advances in systems science and applications (2010), vol.10, no.1  57 abnormal returns on the first listed day. the reform of split share structure will change the expectations of the investors and cause the fluctuations of stock prices. so it is thought that listed before or after the reform of split share structure is related to the abnormal returns on the first listed day. 3. empirical studies 3.1 multivariate linear regression model ⅰ 3.1.1 variables according to the above-mentioned hypotheses, this model includes following variables, r (abnormal returns)、t (turnover rates)、p (issuing price)、n (net asset per share before issuance)、e (p/e ratio)、l (logarithm of the size of money raised)、i (market index, the sum of the comprehensive index of shanghai stock exchange and the composite index of shenzhen stock exchange is adopted hereby)、c (change of policy, a dummy variable, its value is 1 for the stock issued before the implementation of the reform of split share structure, otherwise 0)、r (negotiable share ratio)、m (methods of issuance, dummy variables. three kinds of them, m1、 m2 and m3 are involved hereby, where m1 is the combination of issuance through the trading network with fixed price and issuance off the trading network by inquiry; m2 is the combination of issuance off the trading network by inquiry and the secondary market placement; m3 is the secondary market placement). 3.1.2 model ⅰ 3 1 2 3 4 5 6 7 8 9 1 j j j r p t n e r l i c mα β β β β β β β β β ε = = + + + + + + + + + +∑ (2) where, r is dependent variable, p 、t 、 n 、 e 、 r 、 l 、 i 、 jm ( 1, 2,3j = )、 c are independent variables, iβ ( 1, 2, ,8i = )and 9 jβ ( 1, 2,3j = ) are coefficients, α is constant, ε is stochastic error. 3.1.3 empirical analysis process first, exclude multicollinearity. this can be done in two ways. one way is to examine tolerance and variance inflating factor. when the tolerance is not bigger than or near 0.1 and variance inflating factor is bigger than 10, multicollinearity may exist. the variable correlated more with other variables should be deleted. the other way is to examine the condition index and variance ratio. when the condition index of some row in the collinearity diagnose table is bigger than 15 and at least two variables have variance ratios bigger than 0.90, multicollinearity may exist. the variable correlated more with other variables should be deleted. second, exclude abnormal points. abnormal points are points with extremely large or extremely small values. these points should be removed in order to keep robust. third, examine the goodness-of-fit. if the regression equation passes the f-test, the equation established. the more the determinable coefficient r2is, the better the goodness-of-fit is. the determinable coefficient used hereby is the adjusted determinable coefficient, adjusted r2. fourth, carry out t-test of the regression coefficients. *** stands for having passed t-test at the 1% significant level, ** stands for having passed t-test at the 5% significant level, and * stands for having passed t-test at the 10% significant level. fifth, select appropriate regression method. weighted least squares method with total shares is adopted. 3.1.4 analysis of the regression results the regression results are listed in table 1 and table 2. m1 should be removed due to its tolerance less than 0.0001. α 、t、r、l、i are significant at the 1% level, but n、p、e、m2、 m3、c are not significant at the 10% level. the condition index in the 11th row in table 2 is 17.020, bigger than 15, and the variance ratios of c and m3 are both larger than 0.90. this shows zhao: study on determinants of the abnormal returns of a-shares on the first listed day 58 that multicollinearity may exist and the model should be modified. table 1 regression result of model i coefficients non-normalized coefficients normalized coefficients statistics of collinearity model b std. error beta t significance tolerance variance inflation factor α 4.001 0.404 9.907 0.000*** t 0.270 0.115 0.123 2.339 0.020*** 0.615 1.625 n 0.005 0.026 0.010 0.187 0.852 0.644 1.552 p -0.007 0.005 -0.077 -1.401 0.163 0.559 1.789 e 0.002 0.002 0.058 1.031 0.304 0.533 1.877 r -0.740 0.186 -0.483 -3.975 0.000*** 0.116 8.650 l -0.164 0.015 -0.924 -11.184 0.000*** 0.250 3.999 i 0.000 0.000 0.888 9.531 0.000*** 0.196 5.091 m2 -0.179 0.152 -0.066 -1.175 0.241 0.548 1.824 m3 0.060 0.243 0.069 0.245 0.806 0.022 45.915 i c 0.094 0.229 0.110 0.411 0.682 0.024 41.786 r2 adjusted r2 f prob>f d.w. 0.667 0.650 39.122 0.000 1.785 variables removed statistics of collinearity model beta in t significance partial correlation tolerance variance inflation factor minimum tolerance i m1 . . . . 0.000 . 0.000 table 2 diagnosis of collinearity of model i variance ratio mod el numbers of dimensi on eigenval ue condit ion index α t n p e r l i m2 m3 c 1 3.343 1.000 0.00 0.01 0.01 0.00 0.01 0.01 0.01 0.00 0.00 0.00 0.00 2 2.071 1.270 0.00 0.00 0.00 0.08 0.06 0.01 0.00 0.03 0.00 0.00 0.00 3 1.342 1.578 0.00 0.15 0.10 0.00 0.00 0.00 0.02 0.01 0.03 0.00 0.00 4 1.102 1.742 0.00 0.06 0.28 0.05 0.06 0.00 0.00 0.00 0.09 0.00 0.00 5 1.000 1.828 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 6 0.963 1.864 0.00 0.12 0.01 0.02 0.02 0.00 0.00 0.00 0.40 0.00 0.00 7 0.452 2.720 0.00 0.03 0.02 0.19 0.20 0.03 0.13 0.06 0.01 0.00 0.00 8 0.357 3.058 0.00 0.16 0.38 0.57 0.45 0.00 0.00 0.00 0.00 0.00 0.00 9 0.293 3.376 0.00 0.43 0.09 0.09 0.14 0.01 0.27 0.16 0.03 0.00 0.00 10 0.064 7.213 0.00 0.03 0.10 0.01 0.06 0.90 0.56 0.71 0.00 0.01 0.03 i 11 0.012 17.020 0.00 0.01 0.02 0.00 0.00 0.03 0.01 0.03 0.43 0.98 0.96 3.1.5 the modification of model ⅰ to avoid multicollinearity, correlation coefficients are examined. if the absolute value of correlation coefficient of two variables is bigger than 0.5, they can be thought of as having strong correlation. this may cause serious multicollinearity. so stepwise method is adopted, variables not significant are removed and variables significant are left. (1) i and e:because the correlation coefficient between i and e is almost 0.6 and the p-value of e in the t-test is 0.304, e is deleted. (2) r:r is strongly related with l、c、m1 and m3. (3) m:m2 and m3 are not significant at 10% level. (4) c:c is not significant at 10% level. by the means of stepwise, advances in systems science and applications (2010), vol.10, no.1  59 model ⅰ can be modified as model ⅱ. 3.2 model ⅱ 1 2 3r i t lγ εγ γ γ= + + + + (3) 3.2.1 analysis of the regression coefficients from table 3, it can be found that γ、t、l、i are all significant at 1% level. multicollinearity does not exist in model ⅱ. 3.2.2 analysis of variance as it is shown in table 3, this model has passed the f-test and is significant. 3.2.3 autocorrelation test through the durbin-watson test, autocorrelation is examined. as it is shown in table 3, there does not exist autocorrelation in model ⅱ. table 3 regression result of model ii coefficients non-normalized coefficients normalized coefficients statistics of collinearity model b std. error beta t significa nce tolerance variance inflation factor γ 2.212 0.226 9.768 0.000** * i 2.729e-0 5 0.000 0.661 14.215 0.000** * 0.898 1.113 t 0.337 0.112 0.154 3.005 0.003** * 0.743 1.345 ii l -0.101 0.009 -0.566 -11.111 0.000** * 0.748 1.336 r2 adjusted r2 f prob>f d.w. 0.607 0.602 104.166 0.000 1.717 3.2.4 the economic explanation of the model based on the regression equation, it can be seen that r positively correlates with i and t , and negatively correlates with l .the regression coefficients γ 、 1γ 、 2γ 、 3γ are 2.212、 2.729e-5、0.337、-0.101 respectively, and their t-test values are 9.768、14.215、3.005 and -11.111. so they passed statistical tests. the concrete economic meanings can be explained as following: (1) t: the positive correlation between r and t shows that almost all the winners of the subscription wish to sell stocks subscribed in order to gain higher yield. meanwhile the losers of the subscription want to buy these stocks eagerly. because the quantity of the losers is more than that of the winners, supply cannot meet demand in the market. this pushes stock prices upward continually and the closing price on the first listed day is much higher than the issuing price. (2) i: the partial regression coefficient of the market index is 2.729e-05. this result presents that investors are more confident when the market is in good condition. this will pull the trading price further on the first listed day. (3) l: the logarithm of the size of the money raised negatively correlates with the abnormal returns of newly issued stocks. this shows that the stocks with smaller size of the money raised are favored by the short-term speculators and their abnormal returns tend to be higher. but the stocks issued by some large-sized companies supported by the state are not favored by the short-term speculators. from another view, this tells that asymmetric information theory of underpricing sponsored by some western economists[7-9] is not suitable for china’s stock market. zhao: study on determinants of the abnormal returns of a-shares on the first listed day 60 4. conclusion on the background of the reform of split share structure, through the empirical analysis of 206 a-shares issued and listed from january 1st,2004 to december 31st,2007, it is found that turnover rate on the first listed day, size of money raised, and market index,have significant impact on the abnormal returns of a-shares on the first listed day, whereas, the influence of issuing price, issuing p/e ratio, net asset per share before issuance, real negotiable share ratio after issuance, method of issuance and the reform of split share structure, is not significant. references [1] yunhui jin, wen yang. an empirical analysis of the phenomenon of underpricing of newly issued stocks. statistical research, 2003, 3: 49-53. [2] bowen sun, xiaoguang luo, baiyu sun and tianli liu. the study on the topology structure of china stock market and the self-organized criticality. advances in systems science and applications, 2007, 7(1): 127-131. [3] wei xue. data analysis based on spss. beijing: renmin university of china press, 2006. [4] shancun liu. application of excel in financial modeling analysis. beijing: people’s postal and telecommunication press, 2004. [5] huayang yin, xinping xia and pengcheng zhang. factors resulting in illiquidity premium of illiquid share in china. advances in systems science and applications, 2006, 6(1): 161-170. [6] bo qu. an empirical analysis on liquidity premium value of non-floating shares and influencing factors based on artificial neural network. advances in systems science and applications, 2007, 7(2): 306-312. [7] allen, f., faulhaber, g. r. signaling by underpricing in the ipo market. journal of financial economics, 1989, 23(2): 303-324. [8] grinblatt, m., hwang, c. y. signaling and the pricing of unseasoned new issue. journal of finance, 1989, 44(2): 393-420. [9] welch, i. seasoned offerings, imitation costs, and the underpricing of initial public offerings. journal of finance, 1989, 44(2): 421-450. advances in systems science and applications (2014) vol.14 no.1 84-100 optimized resource-to-task scheduler for cloud service r. a. karthika and a. shajin nargunam computer science & engineering department, noorul islam university, tamilnadu, india-629180 abstract a key concern in cloud computing setting is to capitalize on turnover by accommodating all arriving needs and to diminish bad consequences for cloud providers. attaining these purposes extremely depends on optimal usage of accessible resources in datacenters. to diminish the power and time consumption in cloud computing environment, the previous work evaluate the process of identifying level of each server peer consuming power and task execution time to perform web requests from client peers. this work explains inadequacies caused by the absence of resource organization method and propose a geometric representation using scheduler-based optimal resource allocation (s-ora). the proposed s-ora scheme investigates the suitability of commercial cloud service using amazon ec2 to hierarchical data exchange between multiple cloud instances with minimal resource utilization. the experimental evaluation shows that proposed performance of cloud computing services including amazon ec2, with suitable management of resources result in an increment of profit by reducing rejected requests of the cloud. keywords cloud instance, distributed networks, resource utilization, rejected requests. 1 introduction with the fast development of computing, storage space and system tools, dispersed calculating paradigms have undergone intense amends in the precedent decade. such a mutiny allows request service providers (asps) to organize treacle or even pet scale requests for manufacture reason. for instance, employing cluster or grid calculating amenities, scientists are capable to sprint large-scale climate anticipate representations with a huge quantity of data engendered from the most superior scientific tools, and regularly issue the estimate outcome to the universal public. nevertheless, owing to the vast computational intricacy and elevated quantity of data, such requests need demanding resource practice and place a serious financial weight on the association who systems, upholds and functions these possessions. in the kingdom of large-scale dispersed computing, the rising cloud computing notion, with its almost infifinite possessions and suppleness, is being accepted with the assurance to release asps from the upfront system and preservation cost of the communications. as an efficient means of giving subtracting resources in a shape of usefulness, advances in systems science and applications (2014) vol.14 no.1 85 cloud computing has lately concerned a considerable quantity of concentration for both manufacturing and academic world. the example move to cloud computing is determined by strapping command, particularly from endeavors, to progress the general efficiency by means of and running computing possessions. cloud repair providers employ datacenters to stipulation a common pool of calculation, storage space and bandwidth possessions, to be exercised by requests when the requirement arises. as resources at datacenters are common by using virtualization, requests are permissible to statistically complex such possessions in the structure of practical machines. completion of cloud arrangement has been probable with release source and profitable software letters. in the middle of them opennebula, eucalyptus, and openqrm are finest option for making iaas cloud communications as a release basis service. bringing services in iaas stage has been probable by means of datacenters as possessions with the assist of virtualization knowledge. virtualization is equipment that intangibles absent the particulars of corporeal hardware and gives virtualized resources for high-level requests. virtualization is a leading expertise in the countenance of resource operation sorting in huge kind of resources from mainframe, recall to storage and system. leveraging virtualization creates it probable to combine, disconnect and evolutes resource operation. spanning and dimension of the cloud possessions demand arrangements for resource administration and transmitting requests. this subject is an inspiration to suggest algorithms and machines for allocation, harmonizing and organization of resources in cloud surroundings. current years there have been planned different algorithms for organization and qos provisioning in the cloud resources. the objective of such a cloud is to supply competent good method questioning on the back-end data at a short cost with clever mode while being reasonably practical, and in addition, best advantageous and also receiving decrease of preparation cost on command changes. a worth over the commission charge for every structure can make sure proceeds for the cloud. and also inside cloud is occasionally inform his caches in turn as necessary on diverse cloud and other command will almost runs a diverse enthusiastic server at an instance on cloud. 2 related work cloud computing is a hopeful commercial framework that promises to do away with the obligation for supporting pricey computing amenities. nonetheless, the current cost-effective clouds created to continue web and minute database workloads tremendously vary from characteristic computing workloads. while services related to current cloud computing are insufficient for scientific computing, many-task computing [1] employing loosely coupled applications which did not provide with resources, required instantly. our work differs in 86 r. a. karthika: optimized resource-to-task scheduler for cloud service the problem that the address is towards the minimization of resource utilization obtained from the resource pool instantly. subsequently, vm migration algorithm based on nash equilibrium [2] solves resource utilization maximization problem in virtualized data centers. cloud computing is a deep uprising method in its calculation ability. the major purpose currently is to decrease the charge of organizing a service in the cloud and containing correct coordinative in among the models. public, private, and mixture cloud [3] surroundings all facade the presentation confines intrinsic in todays requests and networks. in turn for endeavors to exploit the plasticity and cost investments of the public, private, and cross cloud they have to conquer the similar latency and bandwidth restraints that confront dispersed it communications environments [4]. the optimization trouble of reducing resource cost in cloud for meeting service necessities is analyzed in literature[5]. social clouds provide the possibility to share resources among clients. the benefit provided to clients within a social network community using different economic patterns and evolved technical metrics, through simulation [6]. a novel framework called, network flow based resource allocation nfra, for reducing the energy consumption and increasing the profit was introduced in [7]. the ocrp [8] gave provisioning for computing resources to be used in multiple provisioning stages as well as a long-term plan in which cloud consumer minimize cost of provisioning in cloud environments. in literature[9], a novel highly decentralized information accountability framework to keep track of the behavior of the users in the cloud was maintained to provide the accountability. the problem for assigning with a set of clients for certain demands towards a set of servers with capacities and degree constraints are presented in literature[10]. the goal of heterogeneous resource allocation was to find an allocation, in such a way that the number of clients allocated with a server is smaller than the degree of assigning to the server. at the same time, their overall demand of the client should be smaller than the server’s capacity, while maximizing the overall throughput. however, achieving security in cloud was another major concern. the author presented fade [11], a mechanism to achieve full security goals upon a set of cryptographic key functions supported by a quorum of key managers. in literature[12], the task of assigning a third party auditor was the main focus that works on behalf of the cloud and client to verify the integrity of data stored in client. consequently, the choosing nodes for implementing a job in the cloud computing must be measured to develop the efficiency of the resources [13]. the author provided a method called, nephele, a first data processing framework that uses dynamic resource allocation [14] for both the tasks, scheduling and execution. with cpu and memory as constraints, a protocol called gossip advances in systems science and applications (2014) vol.14 no.1 87 [15] was designed that minimizes the cost for adapting an allocation that executes on dynamic input and does not require global synchronization. scheduling algorithms for parallel jobs [16] make efficient use of the two tier vms to improve the responsiveness which significantly outperforms commonly used algorithms such as extensible argonne scheduling system in a data center setting. a method for the efficient mapping of resource requests with a heuristic methodology [17] was addressed. routing and scheduling algorithm [18] for cloud architecture target minimal total energy consumption by switching off unused network and/or information technology (it) resources. the author provided maximized resource utilization with optimal execution efficiency [19] using the proportional share model. various workflow scheduling algorithms are listed and compared their characteristics and applicability for cloud scheduling and concluded that hcoc schedulers superior performance may partially be due to its multicore awareness, which is clearly a characteristic requiring consideration in hybrid cloud computing [20]. a mechanism for scheduling single tasks considering two objectives: monetary cost and completion time and dynamic scheduling of scientific workflows [21] was proposed. a cloud scheduler [22] considers both user requirements and infrastructure properties assured users that their virtual resources are hosted using physical resources that match their requirements without getting users about the details of the cloud infrastructure. all the works mentioned above provide mechanisms for resource management in cloud environment. certain other problems related to resource allocation for cloud in the existing p2p infrastructure has to be addressed. so, the problem specification is provided in depth, followed by the method adopted to solve the issue. 3 problem specification normally, cloud consists of set of resources. to manage the set of resources in the cloud, it is necessary to provide a resource manager (rm) to control and set up. the request sent by the clients (c) for the requisition of resources is processed based on dispatcher algorithm. once the resources are allocated and released, resource monitoring unit (rmu) set up a list of resources which are free from the data centers on it. the resource monitoring unit also provides information about the unavailability of the resources in the data centers. the sequence diagram described in fig.1 uses web service as an interface to enhance the communication between the clients and clouds. the workflow of managing the resources proceeds with the clients’ requests. the foremost step is that the client sends a resource request to the cloud, through the web interface (wi). the wi then passes onto the authorization unit 88 r. a. karthika: optimized resource-to-task scheduler for cloud service fig.1 resource allocation (without scheduler) (au). the au check for authorization and proceeds with the next process upon successful completion of the authorization. it will pass the clients requests to the resource manager (rm). upon analyzing the set of resources required by the client, the rm passes the requests to the resource dispatcher unit (rdu). before dispatching the set of resources to the client, the rdu checks the clients authorization information from the authorization unit (au). then the au allocates the resources to the clients who send requests. by this way, the resources are allocated to the clients from the clouds. 4 scheduler-based optimal resource allocation in cloud environment the previous section discussed about the process of managing the resources in the clouds in conventional manner. if allocation of resources is done improperly, then the resources which are free are simply assigned to the task accomplished. this results in the wastage of resources in the resource pool. to enhance and to further optimize the resource allocation in the cloud, in this section we describe a mathematical model using scheduler-based model with a set of constraints. as a solution for resource utilization in hierarchical distributed peer networks, inward requests to cloud primarily are sent to an essential component. this vital component called the scheduler, judge quantity of present resources in cloud instances and employs an improvised dispatcher algorithm to choose which cloud instance swarm this request and propels essential instructions to its virtual machine formation. at first we design the architecture, further explain the scheduler-based optimal resource allocation using improvised dispatcher algorithm. advances in systems science and applications (2014) vol.14 no.1 89 fig.2 architecture of s-ora 4.1 system architecture of scheduler-based optimal resource allocation the scheduler-based optimal resource allocation is designed based on sensible clouds like amazon ec2. the architecture of scheduler-based resource allocation is shown in fig.2. the process of allocating the resources is described individually based on the four components namely, the client, the allocator, the scheduler and finally the optimizer. the process starts with the (i) client (c), (ii) allocator comprises of web interface (wi) and authorization unit (au) (iii) scheduler schedules the availability of resources according to the availability and finally (iv) optimizer comprises of resource manager (rm), resource monitoring unit (rmu) and resource dispatcher unit (rdu). the notations will be used throughout the work. 4.2 process of scheduler-based optimal resource allocation to balance the set of utilized resources and free resources in the cloud, in this section, the resources are distributed among the clients to increase the availability with the help of scheduler. the scheduling of resources in the clouds are focused oriented towards the management of memory and cpu usage of every clients required. a mathematical model is presented here to introduce a parametric environment for better allocation of resources in cloud computing environment. before assigning the resources to the clients, the current stature of the client is to be noted. the stature includes the requirement of resources, memory and cpu usage. for this purpose, a parametric setting is initiated in a method to think about proximity and resource consumption on cloud instances as illustrated below 90 r. a. karthika: optimized resource-to-task scheduler for cloud service using the following constraints. let us consider the scheduler-based optimal resource allocation problem for processing and memory usage denoted by the equations (1) and (2) respectively. α = n∑ i=1 (procc[avail]− proci[cons] total procc (1) β = n∑ i=1 (memc[avail]−memi[cons] total memc (2) where i = 1, 2, ..., n requests placed by the clients to n cloud instances denoted by c = 1, 2, ...,m. the scheduler-based optimal resource allocation using processing and memory are further denoted by objective function as given below: x = α ∈ procnsymbolizes x =  x1 x2 ... xn  , xi ∈ proc (3) similarly, y = β ∈ memnsymbolizes y =  y1 y2 ... yn  , yi ∈ mem (4) where procn denotes set of n representation or n requests placed by the clients. procc[avail] and memc[avail] represent the current processing and memory availability in cloud instances denoted by ‘c’ for a particular request with total processing and memory available in cloud ‘c’ represented by total procc and total memc respectively. proci[cons] and memi[cons] represents the processing and memory consumed for the corresponding ith request. this design creates the model selfsufficient of quantity of resources in each cloud instance and creates a reliable representation for all cloud instances. the two parameters, proci and memi show the ratio of resource utilization for every machine. the two parameters are evaluated for every request and for every clients present in the cloud. the process of scheduler-based optimal resource allocation in cloud computing environment is performed using (a) cloud interface generation and (b) allocating optimal set of resources which will be discussed in the forthcoming section. 4.3 cloud interface generation the cloud interface generation (cig) forms the foremost step in schedulerbased optimal resource allocation. the cig generate, retrieve and modify the advances in systems science and applications (2014) vol.14 no.1 91 fig.3 model for scheduler-based cloud interface generation data items from the cloud instances. with the interface generated, the clients in the clouds derive the capability of the cloud to manage, link and data placed in it. the cloud interface send and process the data items through the set of cloud instances. with the cloud instances, the client sent the request for resources. most of the existing cloud process through these interfaces. a sample model for the cloud interface generation using scheduler-based model is shown below in fig.3. for the purpose of data storage operations, the client requires to know only about the link objects and data objects. as illustrated in fig.3, the client performs a put to the link url and creates a new link with the specified name. once the link for the client is created, the scheduler does the job of scheduling the processing and memory of the respective client and finally performs a put to form a new data object url. the subsequent get then fetches the actual data object and its corresponding value section. 4.4 allocating optimal set of resources once the cloud interfaces has been generated, the clients select the cloud instances from the set of cloud instances available in the network. next, the optimal resource allocation is taken place for each client who sends requests to the scheduler as illustrated in fig.4. the scheduler allocates the resources based on the resource availability. the optimal resource allocation is done based on the process of data exchange between the available cloud instances and assigning/releasing the resources based on the users tasks. the processes is based on two steps, one is task 92 r. a. karthika: optimized resource-to-task scheduler for cloud service selection and node selection. for the optimal set of allocating resources, the scheduler maintains resource pool and task pool. the resource pool consists of set of freed resources and the task pool consists of set of tasks which are not assigned to any resources. at first, the scheduler checks these two pools before assignment of resources. the scheduler identifies the rate of resource utilization and further identifies the task which is to be performed. once the resource from the resource pool is selected, the particular resource is locked by the scheduler and it cannot be assigned to any other tasks requested by the clients. once the resource is locked, the corresponding node which holds its task is ready to send the task to finish. once the task is received by the node, the resource starts its process to accomplish the task. fig.4 scheduler-based optimal resource utilization using improvised dispatcher algorithm for distributed hierarchical peer networks if any new task enters into the network, the scheduler checks the status of the task. if the task is unassigned, the scheduler determines the resource requirements of the corresponding task and further assigns it. based on the utilization of resources and task, the nodes will adjust its corresponding resource allocation advances in systems science and applications (2014) vol.14 no.1 93 processes. the performance of the scheduler-based optimal resource allocation is analyzed in the forthcoming section. 5 experimental evaluation the scheduler-based optimal resource allocation (s-ora) for hp2p networks with efficient data exchange among cloud instances are implemented in java using amazon ec2 a web service that provides resizable computing capacity in cloud. the s-ora uses the amazon ec2’s simple to use web service interface to obtain and configure that reduces the time taken to obtain and boot new server instances to minute. the performance evaluation tests aimed at comparing the direct invocation for cloud computing services using hierarchical distribution process with challenging interactions through traditional scientific computing. the cloud computing services at first recognizes the source and destination node for data transfer from one cloud instance to the other. the source nodes effectively chose the destination node based on hierarchical distribution process. so, the data transfer is successfully performed in cloud computing environment with amazon relational database service. then the fitness of commercial cloud service to hierarchical data exchange between multiple cloud instances is measured. presenting an efficient commercial cloud structure for service provider with optimal resource utilization and fast and accurate data exchange between the cloud service users. the performance of scheduler-based optimal resource allocation is evaluated by, number of rejected requests, time consumption to exchange data between cloud instances and resource utilization by clients from different cloud instances. in this section the experimental setup for designing s-ora that is used in our experiments is explained. the experiments were conducted on amazon’s ec2 infrastructure due to the popularity, feature rich, and stable commercial cloud available. it offers distinct resource configurations for virtual machine instances. amazon ec2’s interface minimizes the time required for various instances according to the changes observed in computing requirements. s-ora experiments with c1.medium, a compute optimized instance type, a 32-bit processor, 1.7 gb ram and 350 gb local disk storage. 6 results and discussion in s-ora, hp2p networks are designed for cloud computing service composition to capture the minimal resource allocation for service performance to other systems written in mainstream languages such as java using amazon ec2 web service. independent tests are run with growing number of applications, and constant number of service requests sent by each user. the performance graph and table describes the evaluation of the performance of s-ora. 94 r. a. karthika: optimized resource-to-task scheduler for cloud service 6.1 measure of number of rejected requests the number of rejected requests [clientrejreq] measures the number of request rejected as made by the client to the cloud instance through scheduler using amazon ec2 with windows server. the rejected request for a client is evaluated based on the difference between the requests made by the client and rejected request made by the cloud for the specific client at a particular time period. a rejected request for a client is given as: clientrejreq = clientreq − cloudrejreq (5) table 1 described the performance of s-ora and compared the results with existing works like bargaining towards maximized resource utilization in video streaming datacenters [1], ccs using scientific computing [2] based on requests to be rejected. table 1 number of requests vs. rejected requests number of requests (task/minute) number of rejected requests (task/minute) proposed s-ora existing nash bargaining solutions existing ccs using scientific computing 5 1 3 4 10 3 5 7 15 4 8 13 20 7 12 16 25 10 14 17 30 12 15 22 fig.5 shows the ratio of requests made by the client to cloud instances, to the number of rejected requests from the cloud instances according to the availability of resources. sending the client requests to cloud, either result in accept of the particular request, if the cloud has resource to be allocated from pool of resources or reject the clients request if the resource is not available. with the help of improved dispatcher algorithm, the number of rejected request is measured. the algorithms effectiveness is measured using the rejection rate. higher the number of rejection rate, lower the performance of algorithm. since the resource is allocated with the help of scheduler in the proposed model, the performance of s-ora is improved when compared to other algorithms. the number of rejected request using s-ora is 40%, 50% using nash bargaining solutions and 70% with ccs using scientific computing. when compared with the three algorithms, the rejected rate using s-ora is comparatively less as the scheduler schedules the advances in systems science and applications (2014) vol.14 no.1 95 fig.5 number of requests vs. rejected requests resource using the resource pool whereas the rejected rate using nash bargaining solutions is 50% which is less when compared to ccs as it uses the nash bargaining solution. 6.2 measure of time consumption for data exchange the time consumption [timecon] in s-ora measures the time consumed for data exchange between the cloud instances. the time consumption is evaluated on the basis of distance for particular data exchange by the cloud instances divided by the speed with which the operation is performed. the formula to evaluate time consumption is given below: timecon = distancedataexc/speed (6) the time consumed to exchange the data by the cloud instances are illustrated in table 2 for the s-ora with the existing works like bargaining towards maximized resource utilization in video streaming datacenters, ccs using scientific computing. fig.6 describes the consumption of time taken to exchange the data among the cloud instances in the distributed peer networks. in the proposed s-ora, the deployment of commercial cloud computing services is efficiently achieved using the scheduler-based cloud interface generation. the is due to clear separation between resource manager (rm) and resource dispatcher unit (rdu), where the scheduler maintains the set of resources and at the same time checks through the pool of resources and allocates accordingly. this in turn results in minimum 96 r. a. karthika: optimized resource-to-task scheduler for cloud service table 2 data size vs. time consumption data size (kb) time consumption for data exchange (seconds) proposed s-ora existing nash bargaining solutions existing ccs using scientific computing 10 10 15 20 20 18 24 26 30 22 33 34 40 30 42 45 50 38 56 58 60 43 68 70 fig.6 data size vs. time consumption time consumption to the hierarchical peer users. rather than existing works, the s-ora consume less time to exchange the data among the set of cloud instances. the variance in time consumption is 5-10% low in the s-ora whereas using nash bargaining solutions the time consumption is comparatively higher due to the use of vm migration algorithm and finally higher in ccs with many task computing. 6.3 measure of resource utilization by clients the resource utilization [resutil] in s-ora measures the utilization of resources by the clients from different cloud instances through scheduler. resource utilization is the summation of resource consumption by the total resource availability. advances in systems science and applications (2014) vol.14 no.1 97 table 3 algorithms vs. resource utilization. algorithms resource utilization (%) proposed s-ora 50 existing nash bargaining towards maximized resource utilization in video streaming datacenters 70 existing ccs using scientific computing 80 fig.7 algorithms vs. resource utilization resource utilization is evaluated using the formula given below: resutil = actualresutil/totalresutil = proci +memi/total proci + total memi (7) the comparison of utilization of resources for the purpose of exchange of data using three different algorithms is depicted in table 3. fig.7 describes the utilization of resources for the efficient data exchange between the cloud instances in the hierarchical distributed peer networks. when a network resource, for instance a cpu or a meticulous disk, is engaged by an operation or query, it is occupied for handing out further requests. awaiting requests have to stay for the resources to turn out until the resources have been freed by other clients. existing woks such as bargaining towards maximized resource utilization in video streaming datacenters and ccs using scientific computing has higher the percentage of time that the resource is occupied, the longer each operation must wait for its turn. so, rather than using existing schemes, the s-ora utilizes minimal resources for transfer of data which is handled efficiently 98 r. a. karthika: optimized resource-to-task scheduler for cloud service by the scheduler with the set of cloud instances. further, from the table 3 it is evident that s-ora uses minimum resources (50%) whereas the nash bargaining solutions uses (70%) that is minimum when compared to the ccs (80%) as the nash bargaining solutions equips the pivot point which is determined by the user, but comparatively higher to the s-ora as it uses the scheduler-based cig resulting in minimum resource utilization. finally, it is being observed that the scheduler-based optimal resource allocation technique efficiently manage the resources and tasks of the nodes in the cloud instances. by generating the scheduler-based cloud interface, the allocation of resources is done in an optimal manner. 7 conclusion this paper efficiently analyzes and resolves the issue of resource management in cloud environment by adapting scheduler-based optimal resource allocation technique. to manage the resources in the clouds, the cloud interface is generated primarily with the help of scheduler. with the cloud interface, the resources and the unassigned tasks are balanced efficiently using the scheduler. then the improved dispatcher algorithm is presented to manage the resources and tasks efficiently in the cloud instances. the scheduler-based optimal resource allocation using improved dispatcher algorithm is implemented on amazon ec2 as a practical cloud. the obtained results are compared with the existing works like bargaining towards maximized resource utilization in video streaming datacenters, ccs using scientific computing. the results had shown that improper resource allocation in the existing works leads to decline in resource utilization. experimental results showed that using the optimal resource allocation algorithm provides at most 5-10% enhancement in resource consumption than unique works in commercial cloud instances. references [1] a. iosup, s. ostermann, m. n. yigitbasi, r. prodan, t. fahringer and d. h. j. epema. (2011), “performance analysis of cloud computing services for many-tasks scientific computing”, ieee transactions on parallel and distributed systems, vol.22, pp.931-945. [2] yuan feng, baochun li and bo li. (2012), “bargaining towards maximized resource utilization in video streaming datacenters”, ieee infocom, no.25-30, pp.1134-1142. [3] v. khoshdel, s. a. motamedi, s. sharifian and m. farhadi (2011), “a new approach for optimum resource utilization in cloud computing environadvances in systems science and applications (2014) vol.14 no.1 99 ments”, international econference on computer & knowledge engineering, no.13-14, pp.314-321. [4] deepak mishra and manish shrivastava (2012), “optimal service pricing for cloud based services”, journal of soft computing & engineering, vol.2, pp.531-540. [5] han zhao, miao pan, xinxin liu, xiaolin li, yuguang fang. (2012), “optimal resource rental planning for elastic applications in cloud market”, ieee 26th ipdps, pp.808-819. [6] m. punceva, i. rodero, m. parashar, o. f. rana, i. petri. (2012), “incentivising resource sharing in social clouds”, ieee international workshop on infrastructure for collaborative enterprises, pp.185-190. [7] k. patel, m. annavaram, m. pedram, nfra. (2013), “generalized network flow based resource allocation for hosting centers”, ieee transactions on computers, vol.62, pp.1772-1785. [8] s. chaisiri, bu-sung lee, d. niyato. (2012), “optimization of resource provisioning cost in cloud computing”, ieee transactions on services computing, vol.5, pp.164-177. [9] s. sundareswaran, a. c. squicciarini, d. lin. (2012), “ensuring distributed accountability for data sharing in the cloud”, ieee transactions on dependable & secure computing, vol.9, pp.556-568. [10] beaumont olivier, eyraud-dubois lionel, thraves caro christopher, rejeb hejer. (2013), “heterogeneous resource allocation under degree constraints” ieee transactions on parallel and distributed systems, vol.24, pp.926-937. [11] yang tang, p. p. c. lee, j. c. s. lui, r. perlman. (2012), “secure overlay cloud storage with access control and assured deletion”, ieee transactions on dependable & secure computing, vol.9, pp.903-916. [12] qian wang, cong wang, kui ren, wenjing lou, jin li. (2011), “enabling public auditability and data dynamics for storage security in cloud computing”, ieee transactions on parallel and distributed systems, vol.22, pp.847-859. [13] shu-ching wang, kuo-qin yan, shun-sheng wang, ching-wei chen. (2011), “a three-phases scheduling in a hierarchical cloud computing network”, third international conference on communications & mobile computing, pp.114-117. 100 r. a. karthika: optimized resource-to-task scheduler for cloud service [14] d. warneke, odej kao. (2011), “exploiting dynamic resource allocation for efficient parallel data processing in the cloud”, ieee transactions on parallel and distributed systems, vol.22, pp.985-997. [15] f. wuhib, r. stadler, m. spreitzer. (2012), “a gossip protocol for dynamic resource management in large cloud environments”, ieee transactions on network & service management, vol.9, pp.213-225. [16] xiaocheng liu, chen wang, bing bing zhou, junliang chen, ting yang, a. y. zomaya. (2013), “priority-based consolidation of parallel workloads in the cloud”, ieee transactions on parallel and distributed systems, vol.24, pp.1874-1883. [17] c. papagianni, a. leivadeas, s. papavassiliou, v. maglaris, c. cervellopastor, a. monje. (2013), “on the optimal allocation of virtual resources in cloud computing networks”, ieee transactions on computers, vol.62, pp.1060-1071. [18] j. buysse, k. georgakilas, a. tzanakaki, m. de leenheer, b. dhoedt, c. develder. (2013), “energy-efficient resource-provisioning algorithms for optical clouds”, ieee journal of optical communications & networking, vol.5, pp.226-239. [19] sheng di, cho-li wang. (2013), “dynamic optimization of multiattribute resource allocation in self-organizing clouds”, ieee transactions on parallel and distributed systems, vol.24, pp.462-478. [20] l. f. bittencourt, e. r. m. madeira, n. l. s. da fonseca. (2012), “scheduling in hybrid clouds”, ieee communicatons magazine, vol.50, pp.42-47. [21] h. m. fard, r. prodan, t. fahringer. (2013), “a truthful dynamic workflow scheduling mechanism for commercial multicloud environments”, ieee transactions on parallel and distributed systems, vol.24, pp.1203-1212. [22] i. m. abbadi, anbang ruan. (2013), “towards trustworthy resource scheduling in clouds”, ieee transactions on information forensics & security, vol.8, pp.973-984. corresponding author r. a. karthika can be contacted at: karthika78@gmail.com adv syst sci appl 2017; 3:34–41 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/498 dynamic fracture tests data analysis based on the randomized approach marina v. volkova1∗, oleg n. granichin1,2, yuri v. petrov1,2, grigory a. volkov1,2 1saint petersburg state university, st. petersburg, russia 2institute of problems of mechanical engineering russian academy of sciences, st. petersburg, russia abstract: methods of measurement and prediction of materials dynamic strength are complicated and not standardized. usually dynamic tests are labour-consuming and each assay demands a lot of time in contrast to static experiments. the way of experimental data treatment based on incubation time criterion and randomized approach of sign-perturbed sums is considered. it is shown that a few experimental points are enough to determine strength parameter with accuracy proper for engineering. the lack of real experimental data is compensated in spsalgorithm by carrying out a large number of virtual random tests in which the bernulli signs are randomly generated in this method. keywords: sign-perturbed sums, incubation time, randomized approach, data analysis 1. introduction one of the most important problems in mechanical engineering is measuring the strength and rheological parameters of materials. determination of the static strength or young modulus is a common procedure, because these values result from direct observation. there are standard measurement techniques to obtain the average value of a required parameter with a certain degree of accuracy, for example σc ∈ [σc−;σc +]. also direct observations stipulate that the physical sense of the measured parameter is established and well-known, and all researchers have a unified conception about the final result. the more complicated situation is observed in dynamics, where material strength cannot be characterized by one parameter of the critical stress. experimental test results demonstrate that under intensive impacts, specimens can resist to stress level significantly higher than their static strength σc and stress value at fracture moment σ∗ depends on the rate of loading and the shape of breaking pulse [1–7]. these strain-rate dependencies are interpreted as passport specifications of materials in certain theories. the main difficulty of these approaches is that the variety of strain-rate curves is infinite due to strong influence of impact conditions to σd value. also it is impossible to describe material strength in the same way when fracture is initiated by threshold impacts. in this case the stress level at the breaking moment can be less even than the static strength, this phenomenon is called the fracture delay effect. the structure-temporal approach based on the incubation time criterion of fracture [8,9] is considered at the present work. the main idea of this criterion is that fracture does not occur instantly and there is some characteristic time for every transient process. implementation of only one additional strength parameter as the incubation time τ permits to predict fracture ∗corresponding author: m.volkova@spbu.ru http://ijassa.ipu.ru/ojs/ijassa/article/view/498 dynamic fracture tests data analysis based on the randomized approach 35 stress value and calculate strain-rate dependencies for all types of impacts. this structuretemporal approach was successfully applied for many different problems of dynamic strength estimation for different materials and condensed matters e.g. dynamic fracture of rocks and concretes, dynamic yielding of metals, acoustic ultrasonic cavitation of liquids, etc. [9–12]. the tests measuring incubation time directly are not realized, and now its value could be obtained only implicitly. the simplest way is to select an appropriate value of τ ensuring good correspondence between a model curve and a scatter of the experimental data. any method could be applied for fitting the model curve, e.g. least mean squares (lms) method, but it gives only one value of the incubation time and there is no any estimation of inaccuracy when we have an absence of any adequate variability in observations. standard estimation algorithms use typically the condition of persistent excitation in data. however, this condition is difficult to ensure in the considered problem since dynamic tests are very complicated and labour-consuming, so usually there are no many experimental points to treat. this often results in a degenerate observation data and complicated identification problems. the wellknown set-membership approach to identification uses no statistical properties of the noise, but assumes some known upper bounds on uncertain system components instead. the purpose of the approach is typically to compute some upper or lower set estimates on the set of dataconsistent parameters and no convergence of this set estimates to the true unknown parameter vector can be achieved without any significant additional assumptions (see, e.g. [13, 14]). in the context of the considered problem the robust estimates of min & max methods give a very wide interval for τ with high inaccuracy and this result is very conservative. randomness of the experimental values stipulates to use stochastic methods for treatment algorithms [15,16]. new sps (sign-perturbed sums) procedure proposed in [17] provides rigorously guaranteed non-asymptotic confidence regions for the unknown parameters τ of a linear dynamical control plant in the small-sample setting. in this paper we adopt this approach to the problem of the incubation time evaluation. this method is applicable to the problem because it can provide a confidence interval of τ with the rate of accuracy admissible for engineering. approximately ten data points are enough to determine an average value of incubation time with ε = 20− 35%. the applicability of the new algorithm to the incubation time approach will be illustrated on few dynamic fracture tests of different types of rocks [18]. 2. incubation time approach the main features of the structure-temporal approach are the following. the general form of the incubation time criterion is 1 τ ∫ t t−τ ( σ(t′) σc )α dt′ ≤ 1 (2.1) where σ(t′) is the loading stress function, τ is the incubation time of the fracture, α is a dimensionless parameter, for most brittle materials α = 1. according to (2.1) the fracture does not occur while left part of the criterion is less than one and the fracture moment t∗ corresponds to equality. experimental dynamic tests are often carried out in overthreshold regimes hence stresses grow linearly until the fracture. then the shape function of the loading impacts can be determined by strain-rate of load ε̇ and elastic modulus k: σ(t) = h(t)kε̇t, (2.2) where h(t) is heaviside step function. substitution function (2.2) to criterion (2.1) leads to the following equation for fracture time t∗ h(t∗) ( t∗ τ )α+1 − h(t∗ − τ) ( t∗ τ − 1 )α+1 = s (2.3) copyright c© 2017 assa. adv syst sci appl (2017) 36 marina v. volkova, oleg n. granichin, yuri.v. petrov, grigory.a. volkov fig. 1. dynamic fracture of gabbro-diabase [18]. blue points are experimental data. lines are theoretical curves plotted by the incubation time fracture criterion: solid line least mean squares method; dashed lines min & max robust method where s = (α + 1) ( σc/(kε̇τ) )α is a dimensionless parameter, which value depends on strainrate of load impacts. fracture time t∗ could not be negative, hence h(t∗) = 1 and there are two cases: when t∗ > τ or t∗ ≤ τ . from expression (2.3) it follows that s = 1 for t∗ = τ , hence the final equation set determining fracture moment t∗ is{ (t∗/τ) α+1 = s, s < 1, (t∗/τ) α+1 − (t∗/τ − 1)α+1 = s, s ≥ 1. (2.4) thus, incubation time approach is able to predict strain-rate dependence for the dynamic threshold of fracture σ∗(ε̇) = kε̇t∗. in most cases parameter α equals to 1 during calculation of dynamic strength of brittle materials such as rocks. then roots of equation (2.4) could be expressed in an explicit form and the critical fracture stress is the following: σ∗(ε̇) = ϕ(τ, ε̇) = { σc + τ 2 kε̇, ε̇ ≤ 2σc/kτ,√ 2σcτkε̇, ε̇ > 2σc/kτ. (2.5) an application of the incubation time criterion is demonstrated on the example of an impact fracture test on gabbro-diabase and marble. the theoretical curve in comparison with experimental [18] points is shown in fig. 1. the value of τ = 0.61 µs (solid line) is calculated by lms method, as it is mentioned above this result does not provide any information about inaccuracy of this value. another min & max method gives very wide interval for incubation time value τ ∈ [1.31; 2.21] µs, and there is no information about properties, that new experimental points would lie in this interval. hence, it is necessary to apply some method of data analysis, which gives, at first, a more accurate interval for possible values of τ and, at second, the degree of confidence in it. 3. problem description in dynamic test experiments we can choose acting factor ε̇ and to get the correspondent observation σ∗. dynamic test data are fitted using the following model of n noisy copyright c© 2017 assa. adv syst sci appl (2017) dynamic fracture tests data analysis based on the randomized approach 37 observations: σ∗i = ϕ(τ, ε̇i) + vi, i = 1, 2, . . . , n (3.6) where vi is a random noise (an inaccuracy) with symmetrical distribution. if the strain-rate dependence satisfies the principles of structure-temporal approach, then function ϕ(τ, ε̇i) is ϕ(τ, ε̇i) = kε̇t∗(τ), (3.7) where t∗ is the fracture time predicted by the incubation time criterion by solving equation (2.4). equation (3.7) allows to calculate fracture stress for different τ , then the lms method gives the best fitted value of the incubation time, for that follow sum has a minimum value n∑ i=1 ( ϕ(τ, ε̇i)− σ∗i )2 → min τ . however, we are not able to get a sufficiently good confidence interval for unknown τ without significant restrictions for the noise vi when n is small. the objective is to construct confidence regions for unknown τ that have guaranteed userchosen confidence probabilities for finite, and possibly small number of data points. it must be defined by the observations of outputs {σ∗i}ni=1 and known acting factors {ε̇i}ni=1 which may be chosen. the constructed regions are almost distribution-free, as the only assumption is the noise to have a property of symmetry. this is important since in practice the knowledge about the noise distribution is limited. additionally, the confidence regions should contain the least-squares point estimate. 4. sps procedure for constructing of confidence regions for a finite number of observations we can use the following procedure which is similar to sps procedure from [17]: the lms estimate is obtained as the solution of the equation h0(τ) = n∑ i=1 (σ∗i − ϕ(τ, ε̇i)) dϕ(τ, ε̇i) dτ = 0, where dϕ(τ, ε̇) dτ = { 1 2 kε̇, ε̇ ≤ 2σc/kτ, 1√ 2τ √ σckε̇, ε̇ > 2σc/kτ. we will try to exploit the information in the data as much as possible while assuming minimal prior statistical knowledge about the noise. our core assumption is the symmetry of the noise. for some m > 0 we generate n(m − 1) bernoulli random values βij = ±1 with probability of 1/2, and introduce m − 1 sign-perturbed sums hj(τ) = n∑ i=1 βij(σ∗i − ϕ(τ, ε̇i)) dϕ(τ, ε̇i) dτ , j = 1, 2, . . . ,m − 1. if τ ? is a nominal value of τ then h0(τ ?) and hj(τ ?) have the same distribution since {vi} are symmetric. therefore, there is no reason why a particular |hj(τ ?)| should be bigger or smaller than another |hj′(τ ?)| and the probability that a particular |hj(τ ?)| is the m-th copyright c© 2017 assa. adv syst sci appl (2017) 38 marina v. volkova, oleg n. granichin, yuri.v. petrov, grigory.a. volkov largest one in the ordering of {|hj(τ ?)|}m−1j=0 will be the same for all j, including j = 0 (the case where there are no sign-perturbations). as it can take different values, this probability is exactly 1/m . algorithm: 1. given a (rational) confidence probability p ∈ (0, 1), set integers m > q > 0 such that p = 1− q/m . 2. generate n(m − 1) i.i.d. random signs {βij} with prob{βij = 1} = = prob{βij = −1} = 1/2 for i = 1, 2, . . . , n and j = 1, 2, . . . ,m − 1. 3. set t := {τ : sps indicator(τ) = 1}. procedure: sps indicator(τ) 1. for the given τ compute the prediction error for i = 1, 2, . . . , n δi(τ) = σ∗i − ϕ(τ, ε̇i). 2. evaluate h0(τ) = n∑ i=1 δi(τ) dϕ(τ, ε̇i) dτ , hj(τ) = n∑ i=1 βijδi(τ) dϕ(τ, ε̇i) dτ , for j = 1, 2, . . . ,m − 1. 3. order scalars |hj(τ)| from smallest to biggest. 4. compute the rank r(τ) of |h0(τ)| in the ordering, where r(τ) = 1 if |h0(τ)| is the smallest in the ordering,r(τ) = 2 if |h0(τ)| is the second smallest, and so on. 5. return 1 ifr(τ) ≤m − q, otherwise return 0. note that the lms estimate τ̂ has by definition the property that h0(τ̂) = 0. therefore, the lms estimate τ̂ is included in the sps confidence region. the probability that τ ? belongs to t is given in the following theorem. theorem 1: if the observation noise in (3.6) is independent and it has a property of symmetry then for nominal value τ ? of the incubation time we have prob{τ ? ∈ t } = 1− q/m, where m , q, t from steps 1 3 of the algorithm described above. the proof of theorem 1 is similar to the correspondent one in [17] and it is provided at the expanded version of this paper [19]. the main difference is the nonlinearity of function ϕ(·, ·) from (2.5), but it is a convex function with monotonic property. 5. experiments results of sps procedure applied to experimental data of dynamic fracture tests for different types of rocks are shown in fig. 2-4. there are different amount of data points for each material, but it does not influence the applicability of sps algorithm. it provides values of the incubation time with the proper degree of accuracy approximately 20− 40%. copyright c© 2017 assa. adv syst sci appl (2017) dynamic fracture tests data analysis based on the randomized approach 39 fig. 2. dynamic fracture of gabbro-diabase [18]. red line is plotted by lms method; dotted lines are plotted by sps procedure for τ0.9 ∈ [0.39; 0.89] µs; dashed lines τ0.75 ∈ [0.48; 0.78] µs in the case of gabbro-diabase (see fig. 2) two values of confidence probabilities 90% and 75% were achieved by the following values of sps procedure parameters m = 200, q = 20 and correspondingly m = 200, q = 50. this gives us two intervals of the incubation fracture time τ0.9 ∈ [0.39; 0.89] µs and τ0.75 ∈ [0.48; 0.78] µs, while lms method provides τ = 0.61 µs. thus, sps procedure allows us to calculate the incubation time value with relatively small inaccuracy 25% and with sufficiently high degree of confidence 75%. fig. 3. dynamic fracture of coelga marble [18]. red line is plotted by lms method; dotted lines are plotted by sps procedure for τ0.9 ∈ [0.39; 0.89] µs; dashed lines τ0.75 ∈ [0.48; 0.78] µs in the case of coelga marble (see fig. 3) we have similar situation for following values of the sps procedure parameters: m = 100, q = 10 for 90% and m = 100, q = 25 for 75%. the inaccuracy does not exceed acceptable for engineering 30% for both confident intervals of the incubation time τ0.9 ∈ [0.72; 1.26] µs and τ0.75 ∈ [0.78; 1.1] µs, since lms method gives us τ = 0.91 µs. sps algorithm also provides good results for pervouralsky marble (see fig. 4) despite there are only five data points and lms method provides τ = 1.81µs. copyright c© 2017 assa. adv syst sci appl (2017) 40 marina v. volkova, oleg n. granichin, yuri.v. petrov, grigory.a. volkov fig. 4. dynamic fracture of pervouralsky marble [18]. red line is plotted by lms method; dotted lines are plotted by sps procedure for τ0.9 ∈ [0.39; 0.89] µs; dashed lines τ0.75 ∈ [0.48; 0.78] µs this number of points leads to relatively small number of iterations m = 20, so proper confidences are achieved by follows q = 2 and q = 5. these parameters give us incubation time intervals τ0.9 ∈ [1.35; 2.05] µs and τ0.75 ∈ [1.66; 1.91] µs with inaccuracy less than 20%. it should be noted, that 90%-confident and min & max intervals for pervouralsky marble are approximately the same. the advantage of the sps procedure in comparison with robust estimation is that it gives additional information about inaccuracy of the determined interval. 6. conclusion the incubation time approach proved itself as a good instrument to estimate and predict the limit parameters of loading impacts. the main problem is that there is no standard procedure to determine the incubation time value τ with estimation of its inaccuracy ε. the main difficulty is that the incubation time could be measured only implicitly and standard methods of data analysis do not work. the proposed application of sps procedure demonstrates how to calculate the incubation time with certain inaccuracy ε under limited number of experimental points. the advantage of this method is that it allows us to predetermine the confidence of the target intervals for parameter τ . the incubation time approach complemented by sps procedure for data analysis can become a universal instrument of engineering to determine material strength in dynamics. the further researches are aimed to improve developed method in case of arbitrary noises like in [20, 21] acknowledgements this work was supported by russian science foundation (project 16-19-00057). references 1. campbell j.d. & ferguson w.g. (1970) the temperature and strain-rate dependence of shear strength of mild steel. the philosophical magazine, 21 (1), 63–82. copyright c© 2017 assa. adv syst sci appl (2017) dynamic fracture tests data analysis based on the randomized approach 41 2. rosakis a.j., duffy j. & freund l.b. (1984) the determination of dynamic fracture toughness of alsi 4340 steel by the shadow spot method. j. mech. phys. solids, 32 (4), 443–460. 3. shockey d.a., seaman l. & curran d.r. (1983) advanced mathematical tools for automatic control engineers: deterministic techniques, new york, us: springer. 4. nikiforovsky v.s. & shemyakin e.i. (1979) dynamic fracture of solids [dinamicheskoe razrushenie tverdyh tel], novosibirsk, ussr: nauka [in russian]. 5. ravi-chandar k. (2001) experimental challenges in the investigation of dynamic fracture of brittle materials. physical aspects of fracture. nato science series, ser. ii : mathematics, physics and chemistry, 32, 323–342. 6. kalthoff j.f. & wincler s. (1987) failure mode transition at high rates of shear loading. international conference on impact loading and dynamic behavior of materials (impact87), bremen, west germany, 161–176. 7. kanel g.i., razorenov s.v., baumung k. & singer j. (2001) dynamic yield and tensile strength of aluminum single crystals at temperatures up to the melting point. journal of applied physics, 90 (1), 136–143. 8. petrov y.v. & utkin a.a. (1989) dependence of the dynamic strength on loading rate. mater. science, 25 (2), 153–156. 9. petrov y.v. (2004) incubation time criterion and the pulsed strength of continua: fracture, cavitation, and electrical breakdown. doklady physics, 49 (4), 246–249. 10. gruzdkov a.a. & petrov y.v. (2008) cavitation breakup of lowand high-viscosity liquids. technical physics, 53 (3), 291–295. 11. pugno n.m. (2006) dynamic quantized fracture mechanics. international journal of fracture, 140 (1-4), 159–168. 12. wang q.z., zhang s., & xie h.p. (2010) rock dynamic fracture toughness tested with holed-cracked flattened brazilian discs diametrically impacted by sphb and its size effect. experimental mechanics, 50, 877–885. 13. bai e.-w., nagpal k.m. & tempo r. (1996) bounded-error parameter estimation: noise models and recursive algorithms. automatica, 32 (7), 985–999. 14. blanchini f. & sznaier m. (2012) a convex optimization approach to synthesizing bounded complexity hinf filters. ieee transactions on automatic control, 57 (1), 219– 224. 15. kolbin v.v. (2014) generalized mathematical programming as a decision model. applied mathematical sciences, 8 (69-72), 3469-3476. 16. volkova m.v. (2017) one method to solve a problem of nonlinear two-stage perspective stochastic planning. international journal of pure and applied mathematics, 112 (4), 817-826. 17. csaji b., campi m.c. & weyer e. (2015) sign-perturbed sums: a new system identificaiton approach for contructing exact non-asymptotic confidence regions in linear regression models. ieee trans. on signal processing, 63 (1), 169–181. 18. bragov a.m., konstantinov a.y., petrov y.v. & evstifeev a. (2015) structuraltemporal approach for dynamic strength characterization of rock. materials physics and mechanics, 23, 61–65. 19. volkova m., granichin o., petrov y. & volkov g. (2017) sign-perturbed sums approach for data treatment of dynamic fracture tests. proc.of the 56th ieee conference on decision and control(cdc), melbourne, australia, 1652-1656. 20. senov a., amelin k., amelina n. & granichin o. (2014) exact confidence regions for linear regression parameter under external arbitrary noise. proc.of the 2014 american control conference (acc), portland, usa, 5097–5102. 21. senov a.a. & granichin o.n. (2014) identifikacija parametrov linejnoj regressii pri proizvol’nyh vneshnih pomehah v nabljudenijah. xii vserossijskoe soveshhanie po problemam upravlenija (vspu-2014), moscow, russia, 2708–2719 [in russian]. copyright c© 2017 assa. adv syst sci appl (2017) introduction incubation time approach problem description sps procedure for constructing of confidence regions experiments conclusion advances in systems science and application (2016) vol.16 no.3 11-20 regions of russia economic alignment: trends and tasks for the policy of regional development albert r. bakhtizin1 ,evgeniy m. bukhvald2 and anastasiya kolchugina2 1 central economics and mathematics institute of russian academy of sciences, moscow; 2 institute of economics of russian academy of sciences, moscow. abstract the article provides an analysis of the various programs and concepts, anyway focused at the economic alignment of the regions of russia. the paper discusses the reasons for the stability of interregional economic disparities, its prospects in connection with the priority focus on the innovative modernization of russian economy. the authors provide suggestions concerning the tasks of positive regions economic alignment in russia within the emerging system of strategic planning and its documents, concentrating basic orienteer of spatial development of the russian economy and the federal regional policy. keywords strategic planning; subjects of federation; economic differentiation; policy of regional development 1 introduction during recent years, was clearly delineated the need for qualitatively new approaches to the system of goals and instruments of the federal policy of regional development in the russian federation, including the “classic” or traditional for this policy task of equalizing the levels of socio-economic development of regions of the country. necessary legal and institutional framework for the realization of task by the moment has been formed by the recently adopted federal law on strategic planning in the russian federation no.172[1]. as one of the main priorities of such strategizing we see the need to maintain a high of integrity or integration of the economic space of the country. however, the most significant obstacle to providing such integrated economic space in the country at the present is concerned primarily with continuing deep rupture in the level of social-economic development of regions in russia. these gaps are due both to the structural deformations of the national economy and as well to unsuccessful options of economic reforms in 1990-s and later to low effectiveness of regional policy of the federal center. long-term maintenance in the country regions with qualitatively different levels of socio-economic development has a powerful disintegration effect over the national economy. recognizing the importance of this problem, russian government repeatedly attempted to ensure positive, (i.e. oriented at pulling up lagging regions, and not at the deterrence of the leading regions) economic alignment of subjects of federation as a priority in 12 albert r. bakhtizin et al:regions of russia economic alignment: trends and tasks for ... the regional policy of the federal state. however, serious practical results in this direction havent been achieved. so, in the early 2000-s this range of problems in the sphere of regional development was address to the federal target program[2], which initially was planned to be completed by 2015. but since 2006 this program has been terminated. however, this program worth to be briefly considered since now we are just at the “time point”, when this program was designed to end its action. the program quite objectively estimated the situation with the spatial parameters of the russian economy development. the government document noted: “at the present time the differences in expanding regions of the russian federation in the basic socio-economic indicators has reached a critical level. sharp interregional differentiation is the inevitable consequence of the increase in the number of lagging regions, the weakening of mechanisms of interregional economic interaction and augmentation of interregional discrepancies, which greatly complicate the conduction of national policy of socio-economic transformation”. in this context, the program set very ambitious objectives, namely to reduce the differences in socio-economic development of regions of the russian federation due to the reduction of the gap in main indicators between the developed and lagging regions by 2010 to 1,5 times and by 2015 to 2 times. the reason, why this program was stopped and, accordingly, failed to meet its goals was not only in the insufficient volumes of its financing, but in the apparent isolation of the program from other key directions and priorities of the state economic policy. another reason institutional and instrumental “poverty” of the program, which in fact operated with only one target institution of the “economic alignment policy”, namely with so called regional development fund (doesnt exist now). today we are fully entrenched in the understanding that the economic alignment of the regions of russia can not be a result of one or more funds or other specialized financial institutions, although their role can be very substantial. this alignment in significant parameters is reachable only as a result of the deep changes in the driving factors of the russian economy development, via reducing its dependence on natural resources exploitation in favor of the more rapid growth of high-tech industries. however, as concerned the functions of the above regional development fund, the program actually contained some interesting ideas, which can be used today. according to the program, the fund was intended to act not as an independent financial institution with its “own” budget, but as a “regulating and control center” for various (by purpose and by government affiliation) financial flows, anyway affecting economic and social development of russian regions. in this program it was noted that the fund should be a supervisor for the combined set of relevant parts of federal and regional programs as well as programs and advances in systems science and application (2016) vol.16 no.3 13 projects of industry funding. of course, such multilateral functionality doesnt fully fit with the traditional concept of “fund”, however, the idea of the state institution, aimed to provide unified management, coordination and control of all resources of the regional policy of the federal center seems to us relevant and still in demand at modern time. later adopted conceptual documents of the rf government (for example, the famous “concept-2020”, strategy of innovative development of the russian federation for the period till 2020, etc.) contained certain goals regarding the spatial aspects of the socio-economic development of the country. however, these documents didn’t record either clear-cut priorities of positive economic alignment of the regions or, respectively, special institutions and interventions for achieving this goal. finally we have to admit: for more than 20 years of sovereign russia, several attempts have been made in order to prepare a long-term concept (strategy) of regional (spatial) development of the national economy; to define its priorities, the objectives of this policy and the tools, needed to achieve them. attempts to prepare such a document were numerously undertaken by the former ministry of regional development of russia, by profile committees of the state duma and council of federation of the federal assembly (parliament) of the russian federation, by various expert institutions, etc. however, none of such documents has been brought to the stage of legal approval and contained cost-based objectives, aimed to overcome the extreme parameters of interregional economic disparities in the country. noteworthy that all these papers on regional development policy issues were very similar. all of them operated with approximately same set of instruments, namely: intergovernmental fiscal relations (including inter-budgetary distribution of tax revenues); state programs of territorial development, localization of federal investment projects, etc. besides, the developers of these documents in determining the strategic direction of the federal regional development policy invariably “rotated” around one important contradiction. one position in this dispute insisted on the need to continue to maintain economic alignment as the leading priority for the federal regional development policy. the other position seemed it possible to overcome the absolutzation the idea of economic alignment of the regions towards optimally balancing it with the primary focus of the policy of spatial development on the group of regions-leaders, which can play the role of locomotives for the russian economy as a whole. none of these positions revealed obvious evidence of its benefits and up to now has taken the place of a clearly expressed priority in the economic policy of the federal center. 14 albert r. bakhtizin et al:regions of russia economic alignment: trends and tasks for ... 2 can economic alignment remain a priority of the federal regional development policy? as regards above mentioned dispute, today economic science in russia meets several important issues. first, what is the current degree of economic differentiation of regions of russia; secondly, what is the actual trend of this differentiation at present and what we can expect in the future; thirdly, what is the extent to which economic alignment of the regions of russia its necessary or simply possible to keep among the priorities of economic policy of the federal state. if positive socio-economic alignment of territories still remains among the most important goals of the federal regional development policy, this task should be considered as unsolved. according to our estimates, based on the latest published data on gross regional product grp (2013), the gap between economically extreme regions of russia (tyumen region and the republic of ingushetia) in per capita grp amounted to 15.7 times in comparison with 17.7 times in 1995. however, its hardly an irrefutable argument in favor of an apparent reduction of interregional economic disparities in the country; at best, we can resume only about the stabilization of the situation. this is due to the fact that these year-by-year calculations, based on the two “extreme” subjects of the federation have a high degree of conditionality, particularly in the context of the ranking of the economically less developed regions. the implementation in one of these regions a large, co-financed by the federal government investment project for some time can remove the region from the last positions in the ranking. then the situation changes in favor of some other region, etc. in this regard, more reliable estimates of economic differentiation of regions should be based not on individual points of two regions, but as the relation of upper and lower decile groups of russian regions (see table 1). table 1. the decile coefficients of interregional economic differentiation for the whole set of regions of the russian federation . 1995 r. 1996 r. 1997 r. 1998 r. 1999 r. 2000 r. 2001 r. 3.46 3.13 3.36 3.41 4.07 4.15 3.68 2002 r. 2003 r. 2004 r. 2005 r. 2006 r. 2007 r. 2008 r. 3.18 3.21 3.86 .91 3.78 3.55 3.27 2009 r. 2010 r. 2011 r. 2012 r. 2013 r. 3.33 3.61 3.52 3.86 3.73 source: russian statistical agency data. according to the data, presented in table 1 and based on the so-called ”decile advances in systems science and application (2016) vol.16 no.3 15 ratio of differentiation” (in this case calculated as the ratio for the relevant years between the lowest value of grp per capita among the 10% of regions with highest per capita grp to the maximum rate for the 10% of regions with the smallest per capita grp) within the period 1995-2013 there was no substantial increase in inter-regional economic differentiation in russia. this figure showed rather a wavy trend. the highest ratio took place in 1999-2000 (4.1 4.2) and at present (2012-2013) these figures are located at the level of 3.7 3.8. in economically developed countries similar differentiation coefficients are somewhat lower: they are within 1.5 times (france, usa) or about 2.0 times (germany, italy), reflecting more integrated spatial economic development of these countries and, consequently, a higher level of integration of the regional segments of their economies. this level of differentiation is typical as well for developed countries with a high share of extractive industries in the economy, such as canada, where the decile ratio also does not exceed twice level. at the same time, more pronounced inter-territorial differences can be found in developing countries, for example in china, where the degree of differentiation of regions by the decile ratio varies from 2.8 to 3.5 times[3]. similar is the situation in the investment sphere. in recent years up to half or more of all fixed capital investment in the russian economy was accumulated in 10 leading regions (in 2014 51.8%). another evidence of a high degree of economic differentiation of subjects of the russian federation is also a limited number of regions-donors, the number of which in the current system of intergovernmental fiscal relations still amounts to 14 (2015), although the substantial increase in the number of such regions (up to half of the total number of subjects of federation) numerously has been declared in policy documents on regional and budgetary policy in the russian federation. most likely, the current trend towards relative stabilization of the economic differentiation of russian regions will continue. first of all, it is connected with absence at the moment and in the near perspective evident regions-locomotives of the rapid growth of the economy as a whole. to day the potency of maintaining rapid growth for the former leading regions (regions rich in raw materials and regions, performing the role of leading transactional centers in the russian economy), has largely exhausted. as concerns regions potential “dots” of innovative development, they are still too weak to take on entirely the role of “locomotives of economic growth”. thus, at present its too early to remove economic alignment of the regions from the position of one of the main priorities of the state economic policy. sometimes, the hypothesize is proved that its not necessary to focus this policy specifically on the task of “alignment”; and its only necessary to create for all regions of russia equal favorable conditions for socio-economic development and the prob16 albert r. bakhtizin et al:regions of russia economic alignment: trends and tasks for ... lem of regional and municipal disparities will be resolved “automatically”. but both our and foreign experience of regional development doesn’t confirm such an opportunity. this is true as well as the fact that any isolated measures aimed to align regions will obviously be unproductive, if these measures are carried out in isolation from other positive changes in the russian economy. for positive alignment of the regions the country needs, first of all, to change key drivers of economic growth. it largely co-operates the priorities of regional development policy, including the positive alignment of the regions, with the implementation of structural reforms in the national economy. in this sense, the policy of innovative modernization represents an opportunity not only to achieve structural changes in the economy and, accordingly, radical increase of its competitiveness, but also to create the basis for a “breakthrough” in solving the problems of regional development, including the task of positive alignment of subjects of the russian federation. however, it is difficult to hope for the automaticity of market regulators of comparative regional development, especially in view of the fact that according to their innovative potential, regions of russia differ even in greater degree than in such indicators as grp per capita[4, pp. 35-52]. 3 differentiation of regions and the rates of economic growth in russia is it possible to prove that changes in the degree of economic differentiation of russia’s regions are subject to certain regularities? basing on the analysis of the available data its principally possible to state a hypothesis concerning the fact that in periods of faster economic growth (for example, at the beginning of the 2000s) the rate of differentiation of regions of russia slightly increased, and during the periods of slow growth and stagnation of the economy differentiation is reducing. this hypothesis can be illustrated with the data given in table 2. the data of table 2 illustrate the regularity, proving that a higher economic growth in russia rate usually corresponds to a higher economic differentiation of regions and, conversely, during the period of slowdown of the growth rates, ratio of differentiation is comparatively lower. of course, this is only a preliminary statement, which can not be a reason to think that increasing trend of this differentiation is a obligatory condition or prerequisite of high economic growth rates. accordingly, any slowdown in growth rates is not the very trend of the alignment regions, which could satisfy the country from the point of view of achieving greater integration of its economic space. here we meet one of the most difficult questions for economic theory and practice of regional development policy: is it possible to combine high rates of economic development of the country with a consistent reduction of the gaps which still exist between the regions and, if so, via what means of economic policy this result can be achieved? to our mind, there is a real chance to combine economic alignment of regions advances in systems science and application (2016) vol.16 no.3 17 table 2. decile inter-regional differentiation ratio and general growth rates of the russian federation economy the period of unstable development and crisis decile differentiation ratio for rf regions 1995 1996 1997 1998 3.46 3.13 3.36 3.41 mean rate for 1995-1998 3,34 gdp growth rates ( %) +3.3 -3.6 +1.4 5.3 mean rate for 1995-1998 1.05 the period of relative economic growth decile differentiation ratio for rf regions 1999 2000 2001 2002 2003 2004 2005 2006 2007 4.07 4.15 3.68 3.18 3.12 3.86 3.91 3.78 3.55 mean rate for 1999-2007 3.71 gdp growth rates ( %) 6.4 10.0 5.1 4.7 7.3 7.2 6.4 8.2 8.5 mean rate for 1999-2007 7.09 the period of declining of growth rates and entering stagnation decile differentiation ratio for rf regions 2008 2009 2010 2011 2012 2013 3.27 3.33 3.61 3.52 3.86 3.73 mean rate for 2008-2013 3.55 gdp growth rates ( %) 5.2 -7.8 4.5 4.3 3.4 1.3 mean rate for 1999-2007 7.09 source: russian statistical agency data; authors calculations. with sufficiently high and sustainable rates of the russian economy growth. first, this chance demand for changes in the structure of the economy, which should be clearly associated with the identification and implementation of reasonable measure to reduce economic differentiation of regions and macro-regions of russia. the national development strategy should not just declare this task in general, but should contain economically relevant indicators, quantitatively showing the necessary progress in the reduction of such a differentiation. but this progress is possible only via targeted government economic policy, based on the alignment as a strong priority of national economy spatial development. there is another important condition, which could contribute to the reduction 18 albert r. bakhtizin et al:regions of russia economic alignment: trends and tasks for ... of regional economic disparities simultaneously with the sustainable growth of the national economy. that is setting the entire consistency of “spatial activities of all institutions and instruments of the federal economic policy. by the moment the total picture of ”spatial activities of all these institutions and policy instruments isn’t anyway regulated by the rf government. the main tool of the well-known “social alignment” (i.e. alignment of provision of public goods of a social nature for the russian regions population) is the system of so-called “fiscal equalization” transfers from the federal budget to regional budgets. however, as has shown already the practice of several decades, the current model of financial equalization, partially offsetting the social effects of spatial disparities in the russian economy, actually leads to the conservation of economic background of inter-regional disparities and doesn’t ensure trend to growing financial self-sufficiency of subjects of federation. the grant component of these transfers has almost nothing to do with the intensification of investment processes and growth rates in the regions and for the country in total [5]; doesn’t stimulate sufficient efforts of the regional and local authorities, aimed to support entrepreneurial activity and attraction of investors, including on the basis of public-private partnership. in other words, the financial equalization mechanism acts, but real economic alignment, as shown above, doesn’t occur. practical implementation of the all-national vertical of strategic planning needs for a new model of fiscal federalism for russia with active stimulating mechanisms for the stabilization of sub-federal (regional and local) budgets [6]. however, in the current situation the sphere of inter-budget relations already isn’t the only or even the predominant instrument of the federal impact over socio-economic development of the regions. this impact at present is multilink, and, unfortunately, unsystematic. this situation puts the task of strategic coordination of all the channels (instruments, links) of federal funding for socioeconomic development of regions, in particular, resources going through state targeted territorial development programs, through the operation of institutions for territorial development, as well as in the framework of the “standard” system of inter-budgetary interactions. this practically means a binding of spatial strategic planning and monitoring the final picture of the spatial distribution of all expenditures of the federal budget. here, in fact, it would be possible to return to the idea, proposed in the early 2000s, in the framework of the target state program. that is the idea of setting special “regional development fund”. but now this idea should be implemented at a qualitatively higher level of institutional arrangement, i.e. not simply in the form of a financial institution, but as a supervising institution with the functions (authority) for coordination of the spatial distribution of all types (channels) of federal resources, directed for socio-economic development of regions of the country. advances in systems science and application (2016) vol.16 no.3 19 respectively, should be offered a system of clear and transparent criteria for the selection of regions (macro-regions), to which federal funds are implemented through state target territorial development programs or are established special territorial development funds and corporations, etc. according to the experience of some countries, may be approved a special status the regions, which on the basis of statutory criteria can be declared “economic disaster areas”. these regions will receive in such a case additional forms of government assistance while taking measures to conserve and rehabilitate the regional finances. for example, in germany such territories for a long time are the states (lands) of bremen and saar [7]. similarly, should be formalized principles of spatial localization and concrete functions of federal “development institutions” (“science-towns, special zones and other territories of “advanced” development, etc.), as well as conditions and forms of federal support to various regional development institutions (regional economic zones, industrial parks, etc.). it is important that each institution has a transparent, targeted adjusted spatial pattern (program) of its activities, i.e. among other functions, operated as a policy instrument for regional development. these institutions should be objective-oriented on the spatial distribution of “pulses” of industrial-innovative development and, thus, contribute to the integration of the economic space of the country on the basis of “economy of innovations. 4 conclusion to summarize the research, the following results can be formulated: 1) most likely, the current trend towards relative stabilization of the economic differentiation of russian regions will continue, but letter cannot serve as satisfactory solution of the problem of positive alignment of the total economic space of the country. 2) at present its too early to remove economic alignment of the regions from the position of one of the main priorities of the state economic policy. 3) national development strategy should not just declare this task in general, but should contain economically relevant indicators, quantitatively showing the necessary progress in the reduction of such a differentiation. but this progress is possible only via targeted government economic policy, based on the alignment as a strong priority of national economy spatial development. 4) the russian federation government must ensure the entire consistency of “spatial activities of all institutions and instruments of the federal economic policy and establish a special state institution, aimed to provide unified management, coordination and control of all resources of the regional policy of the federal center. 5) for the state regional policy should be approved a system of transparent criteria for the selection of regions (macro-regions), to which federal funds are 20 albert r. bakhtizin et al:regions of russia economic alignment: trends and tasks for ... implemented through state target territorial development programs; besides following the experience of some countries, may be approved special status the regions, which on the basis of statutory criteria can be declared as “economic disaster areas”. references [1] law principle. (2014),“on the strategic planning in the russian federation” federal law of june, the 28th, russian . [2] law principle. (2001),“on the federal target program :reduction of distinctions in socially-economic development of regions of the russian federation (2002-2010 and till 2015)”, the decree of the russian federation government of october, the 11s, no. 717, russian. [3] nikolaev i. and tochilkina o. (2011),“economic differentiation of regions: estimations, dynamics, comparisons.” analytical report: fbk. [4] valentey s.d. and bakhtizin a.r., bukhvald, e.m. and kolchugina a.v. (2014),“trends of development of russian regions”, economy of the region. ekaterinburg, pp. 9-22. [5] cosgrove michael and marsh daniel. (2013),“ why the slow u.s. economic growth”, journal of academy of business and economics , vol. 13,no. 4. [6] dahlberg matiz. (2001), “fiscal federalism and state-local finance”, regional science and urban economics, vol.31, no.1. [7] bukhvald e.m. (2003), special modes of intergovernmental relations: the realities of russia and the german experience., finance and statistics, moscow. corresponding author albert r. bakhtizin can be contacted at: albert.bakhtizin@gmail.com advances in systems science and applications (2012) vol.12 no.2 162-173 parametric modeling and simulation of orthogonal milling process based finite element method jing sheng1 and liping yao2 1dpt. of mechanical engineering, hubei automotive industries institute, hubei shiyan 442002, china 2dpt. of foreign languages, hubei automotive industries institute, hubei shiyan 442002, china abstract the key techniques of 2d modeling with msc.marc software and the whole modeling procedure of metal oblique cutting process was presented. the rule based on the modeling process was investigated. the finite element simulation of metal machining is a complex process. it is essential to exploit a system to construct a model of simulation so as to obtain simulation data more conveniently and rapidly. this research is significant for the development of parametric modeling. the system’s interface, designed using c++ builder, can access data which includes the geometrical angles and dimensions of tool, the sizes of work, the relative position between tool and work, properties of tool and work, cutting conditions, etc.. the procedure file that is able to model in the msc.marc environment automatically is generated by the program. the parametrical modeling of simulation is completed by the system which calls the procedure file. finally, case studies were performed to study simulation model, interface and simulation. so the parametric modeling is a kind of effective avenue for metal machining process simulation. keywords orthogonal milling machining, parametric modeling, fea, interface design 1 introduction the experimentation, analytic and numerical method are frequently applied on the research on the metal cutting process. the disadvantages of experimentation include the high cost and labor-intensive process, and the analytic methods are very difficult to understand and analyze the machining process in detail. nowadays numerical approaches have been growing acceptance in industries and academia as a method for characterizing machining process [1-10]. it is known that modeling is one of some key techniques, in which geometrical angle and structural parameters of cutting tool are important factors. the modeling procedure, however, is relatively complex and professional. very little research on parametric modeling in metal machining has been reported. numerical simulation will be taken as a tool by consumer, and the parametric modeling is the first problem to be solved because of the practicability. the advances in systems science and applications (2012) vol.12 no.2 163 paper described the general modeling procedure of metal cutting simulation, and presented some key techniques, especially parametric modeling. 2 the parametric modeling of milling process 2.1 geometrical modeling of milling system geometric modeling is very important for the simulation of milling. it directly affects the operation and results of the simulation system. the geometrical model of cutting system is presented in fig.1. for confirming the milling tool’s rotating center conveniently, the origin of the coordinate is placed on the milling tool’s rotating center. in fig.1, ∆x represents the transverse distance between the outer circle of the milling tool and the lateral surface of the workpiece, and ∆y is the height between the outer circle of the tool and the machining surface of the workpiece. fig.1 the mode of two-dimension milling 2.2 geometric modeling of milling tool fig.2 shows the polygon line of a milling tool during upmilling. created the shape of single tooth, the 2-d model of milling tool was derived from duplicating the shape of single tooth along origin o. computing formulas of the coordinates of the points on a tooth are seen in table 1. where θ1 is the angle between face and back of tooth,θ2 is the angle of chip flute; while γ, α01, α02 represent the rake angle, the first relief angle and the second relief angle respectively. the tooth-spacing angle, the width, the height of teeth and the diameter of cutting tool are denoted by ε, ba1, h and d0. the computing formulas of the coordinates of down milling can be obtained by changing the abscissa in the formula into negative number, while keeping the value of ordinate the same. 164 jing sheng:parametric modeling and simulation of orthogonal milling process... fig.2 the polygon line of milling tool table 1 computing formulas of the coordinates of the points on a tooth point coordinates of the points p1 x1 = −d0 2 + h; y1 = h× tgγ0 p2 x2 = −d0 2 ; y2 = 0 p3 x3 = −d0 2 + ba1 × tgα01; y3 = ba1 p4 x4 = y5−y3+ctgα02×x3−tg(θ2−ε+γ0)×x5 ctgα02−tg(θ2−ε+γ0) y4 = y5ctgα02−y3tg(θ2−ε+γ0)+tg(θ2−ε+γ0)(x3−x5)ctgα02 ctgα02−tg(θ2−ε+γ0) p5 x5 = − √( −d0 2 + h )2 + (h× tgγ0) 2 × cos ( ε+ arctg h×tgγ0 d0 2 −h ) y5 = − √( −d0 2 + h )2 + (h× tgγ0) 2 × sin ( ε+ arctg h×tgγ0 d0 2 −h ) p6 x6 = −d0 2 × cos ε; y6 = d0 2 × sin ε 2.3 geometrical modeling of workpiece the computing formulas of the coordinates of the points on workpiece was derived from fig.1. the formulas are shown in table 2. advances in systems science and applications (2012) vol.12 no.2 165 table 2 computing formulas of the coordinates of the points on workpiece point x y w1 d0/2 + ∆x d0/2−∆y w2 d0/2 + ∆x d0/2−∆y + b w3 a+ d0/2 + ∆x d0/2−∆y + b w4 a+ d0/2 + ∆x d0/2−∆y where a, b are length and height of workpiece respectively. 2.4 geometrical modeling of rigid walls two lines,r1r2 and r3r4, were used to represent two rigid walls.the two rigid walls are used to restrict the movement of workpiece on x-direction and ydirection. the coordinates of the points on two rigid walls are shown in table 3. table 3 the coordinates of the points on two rigid walls point x y r1 a+ d0/2 + ∆x d0/2−∆y + 1 r2 a+ d0/2 + ∆x d0/2−∆y + b− 1 r3 d0/2 + ∆x− 1 d0/2−∆y − b r4 a+ d0/2 + ∆x+ 1 d0/2−∆y − b 2.5 material modeling it is necessary to configure properties of material. the workpiece material during machining generates elastic-plastic deform under high temperature, large deformation and large deformation rate. considering the effect of strain hardening, strain-rate hardening, and thermal softening on the stresses, johnson-cook’s empirical model (see (1)) was adopted. the jcmaterial law parameters are obtained by shpb equipment (see table 4). σ̄ = [a+b(ε̄)n] [ 1 + c ln ( ˙̄ε ε̇0 )][ 1− ( t − troom tmelt − troom )m] (1) where ε̄ is the equivalent plastic strain, ˙̄ε the equivalent plastic strain rate, t the temperature; while a, b, n, c, m and ε̇0 is the parameters determined by a material itself; tmelt and troom represent the melting temperature and the room temperature respectively. 166 jing sheng:parametric modeling and simulation of orthogonal milling process... table 4 jc material law parameters parameter a/mpa b/mpa n c m value 626 3614 0.82 0.0268 1 meshing, configures of contact, boundary conditions, remeshing, analysis conditions, etc. are no longer mentioned here. 2.6 friction modeling between tool and chip there are two explicit areas on the rake surface: slip region and glue region. on the basis of research, constant coefficient friction is applied in slip region and constant friction stress is used in glue one. the friction stress is written as [2]. f = { µσn σf = µσn k σf = k (2) where σn is normal stress. µ is friction coefficient and k is shear stress. 2.7 the criterion of chip separation during the simulation, there are criterions that make the chip separate from workpiece and rake face. they are divided into geometric criterion and physical criterion. the geometric criterion decides the separation through the changes of geometric dimension of deformable body. the physical one is used to identify whether magnitude of physical quantity causes critical value or not. in fact, chips are separated by setting a minimum force or stress of the nodes as threshold. 2.8 equation of heat conduction because the system consists of workpiece, chip and tool generates heat continuously, the first and the second deformation zone of the workpiece go through plastic and elastic deformation. besides, the rake surface of the tool has severe friction [4-5]. equation of the heat conduction in unsteady-state temperature field (take variable thermal conductivity into account) is defined as follows: ρc ∂t ∂t = k ( ∂2t ∂x2 + ∂2t ∂y2 ) + dk dt [( ∂t ∂x )2 + ( ∂t ∂y )2 ] −ρc ( wx ∂t ∂x + wy ∂t ∂y ) +q∗ (3) where k represents the thermo conductivity coefficient and t is temperature.ρ is the material density and c is thermal capacity. x and y are cartesian coordinate advances in systems science and applications (2012) vol.12 no.2 167 system. wx and wy represent velocity component of kinetic heat-source in x and y axis respectively. q∗ is heat generation rate per unit volume. q∗ = whσ ˙̄ε/j (4) where wh is the ratio that plastic deformation workpiece turn into heat energy. σ̄ is equivalence stress. ˙̄ε is equivalence strain ratio. j is coefficient of thermal equivalent of workpiece. because the amount of radiant heat is little, it is ignored. 3 key techniques of parametric modeling of simulation of milling process 3.1 interface design 3.1.1 the interface design between c++ builder and database exploiting database, database module and database engine provided by c++ builder or ado(active data object)were employed to access database. the tables whose type is dbf and frequent bde engine were used, while parameters about bde were set, such as path, type and language drive. 3.1.2 the interface file of parametric modeling in msc.marc governed msc.marc software characters, the system’s knowledge base rules were established. thus, procedure files were written using c++ builder code according to the rule. based on the model of the tool and the workpiece, the topology and geometric information of the tool, such as points, line and surface, the procedure of their modeling as well as meshing was written in procedure file line-by-line. then the modeling of the rigid walls was done, too. while the relative position between tool and workpiece, material model, friction model between tool and chip, properties of tool and workpiece, cutting conditions, the configures about finite element simulation and so on, were written into the file in same way. therefore the created procedure file can be operated according to specified manner. so the modeling process becomes easy and rapid. fig.3 shows the block diagram of modeling process. the structure of procedure file is seen in fig.4. 3.2 creating the parametrical modeling file parametric modeling file (procedure file) can finish the scheduled task in finite element software msc.marc to model and simulate the process of milling. so through explanation facility of a process file, the geometric information of cutting system’s points and lines can be performed. 168 jing sheng:parametric modeling and simulation of orthogonal milling process... fig.3 the block diagram of modeling process fig.4 structure of procedure file 3.3 parametric setting of workpiece and property of milling tool. before the simulating of machining process, geometric properties and material property of elements, contact relationship of bodies, mechanics and thermal conductivity between milling tool and workpiece need to be defined and evaluated. because the dimensions of workpiece, structural sizes and geometrical angles of tool have influence on the number of elements, dynamically meshing workpiece and tool are important. here element sets were employed to store the workpiece and tool elements respectively. the system implemented the method that setting units invisible or visible instead of calculating the number of elements developed before. by programming, all of configures of the model parameters were carried out automatically. advances in systems science and applications (2012) vol.12 no.2 169 4 execution of example the user interface of modeling system is presented in fig.5. the interface has five functions that include file management, adding modeling database, browsing modeling data, parameters of tool and work and contact parameters between tool and work. in the example, the diameter of tool is 30.2 mm, the height of teeth is 4 mm, and other parameters are given. the model of down milling is shown in fig.1a and the model of up-milling is shown in fig.1b.when the amount of feed is 46 mm·min-1,the rotating rate of milling is 275 rpm, and the width of cutting 3 mm, the situations of down milling and up milling are shown in fig.6 and fig.7. fig.8 and fig.9 show the predicted forces of down milling and upmilling. fig.5 the user interface of modeling system fig.6 the simulation of down milling 170 jing sheng:parametric modeling and simulation of orthogonal milling process... fig.7 the simulation of up-milling fig.8 the milling force of down milling fig.9 the milling force of up-milling 5 experimental verification for the purpose of verifying the simulation result, a milling experiment was conducted. the milling equipment is xh719 manufactured by qingdao first machine advances in systems science and applications (2012) vol.12 no.2 171 tools factory, and its power is 28 kw. the material of a milling tool is ys2t. the diameter of milling tool with four teeth is d=30.2 mm. the measure equipment is kistler9237a. down milling was used. cutting conditions were the same as the value used in simulation. the history of cutting force is shown in fig.10. it is shown that the experimental result agree with the simulation data. then orthogonal table was designed to measure cutting temperature(see table fig.10 the simulation of cutting force of down milling table 5 orthogonal table axial radial amount cutting cutting cutting of feed speed width /mm width /mm /mm·min−1 /m·min−1 2.00 5.0 37.5 23.6 2.00 7.1 47.5 29.5 2.00 10.0 60.0 37.3 2.45 5.0 47.5 37.3 2.45 7.1 60.0 23.6 2.45 10.0 37.5 29.5 3.00 5.0 60.0 29.5 3.00 7.1 37.5 37.3 3.00 10.0 47.5 23.6 5). fig.11 shows experiment value and simulation date. 6 conclusions following the expatiations of the whole process of parametric modeling in msc.marc, the paper discussed the key techniques. it has been proven that the ways and means are effective. it is helpful to simulate under different cutting parameters, various dimensions and geometric angles of a tool. therefore parametric modeling will provide good foundation for creating further cutting databases 172 jing sheng:parametric modeling and simulation of orthogonal milling process... fig.11 the comparison of cutting temperature and designing tool. acknowledgements the project is funded by the grants from the doctor fundation of hubei automotive industries institute. references [1] domenico u. (2008), “finite element simulation of conventional and high speed machining of ti6al4v alloy”, journal of materials processing technology, vol.196, pp.79-87. [2] pantalé, bacaria j l, dalverny o, rakotomalala r, s caperaa. (2004), “2d and 3d numerical models of metal cutting with damage effects”, comput. methods appl. mech. engrg, vol.193, pp.4383-4399. [3] vernaza-pena k m, mason j j, li m. (2002), “experimental study of the temperature field generated during orthogonal machining of an aluminum alloy”, experimental mechanics, vol.42, pp.221-229. [4] calamaz m, coupard d, girot f. (2008), “a new materialmodel for 2d numerical simulation of serrated chip formation when machining titanium alloy ti6al4v ”, international journal of machine tools & manufacture, vol.48, pp.275-288. [5] raviraj shetty, laxmikant k r. (2008), “pai finite element modeling of stress distribution in the cutting path in machining of discontinuously reinadvances in systems science and applications (2012) vol.12 no.2 173 forced aluminium composites”, journal of engineering and applied sciences, vol.3, pp.25-31. [6] limido j, espinosa c, salaun, m. (2007), “sph method applied to high speed cutting modeling”, international journal of mechanical sciences, vol.49, pp.898-908. [7] mamalis a g, horváth m, branis a s, manolakos d e. (2001), “finite element simulation of chip formation in orthogonal metal cutting”, journal of materials processing technology, vol.110, pp.19-27. [8] wen q, guo y b, todd b a. (2006), “an adaptive fea method to predict surface quality in hard machining”, journal of materials processing technology, vol.173, pp.21-28. [9] simoneau a, elbestawi m a. (2005), “surface defects during microcutting”, international journal of machine tools & manufacture, vol.10, pp.1-10. [10] xie l j, schmidt j, schmidt c, biesinger f. (2005), “2d fem estimate of tool wear in turning operation”, wear, vol.258, no.10, pp.1479-1490. adv syst sci appl 2017; 3:58–66 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/499 regression tree control of multidimensional static object ekaterina mangalova1∗ 1system analysis and control department, siberian state aerospace university, krasnoyarsk, russia abstract: in this paper the control problem of a static system under incomplete information is discussed. a nonparametric control algorithm based on classification and regression tree is proposed and evaluated for different task formulations: controlled inputs – one-dimensional output, controlled and observed uncontrolled inputs – one-dimensional output, controlled and observed uncontrolled inputs – multi-dimensional outputs. keywords: control, regression tree, machine learning, optimization 1. introduction the modeling of technical systems and processes is the essentially important tool used in engineering to design, improve, optimize and control systems. using simulation modeling is generally cheaper, safer and faster than conducting real-world experiments. this allows to use modeling for different alternatives analyses. there are two different ways to create simulation models: • manually built models. a specialist builds the simulation model manually based on his theoretical knowledge and technical experience. this method is the most useful and imminent for systems that do not yet exist. • statistical modeling. statistical modeling is based on generating models by observations. this method can be useful for complicated systems when input-output relationships are not evident or are unknown, internals are not essential for the technical problem. statistical models are used to detect a deviation of the normal behavior of a system and to control. model predictive control is a family of controllers in which there is a direct use of an explicit and separately identifiable model. control algorithms based on model predictive control concept have found wide acceptance in practical applications and have been studied by researchers. in [1] one of the first successful applications of model predictive heuristic control is described. there are two commonly used and well studied approaches to model predictive control: based on artificial neural networks and based on fuzzy logic. the use of artificial neural networks in model based control, both as process models and as controllers, is investigated by d. c. psichogios and l. h. ungar [2]. a. draeger, s. engell, h. ranke [3] implemented a feed-forward neural network as the nonlinear prediction model in an extended dmc-algorithm to control. h. sarimveis and g. bafas [4] introduceed the method based on a dynamic fuzzy model of the process to be controlled, which is used for predicting the future behavior of the output variables. in the paper [5] j. m da costa sousa and u. kaymak investigated the use of fuzzy decision making in model predictive control. ∗corresponding author: e.s.mangalova@hotmail.com http://ijassa.ipu.ru/ojs/ijassa/article/view/499 regression tree control of multidimensional static object 59 in contrast to neural network and fuzzy logic approaches, in [6] a control algorithm based on the nadaraya-watson estimator [7] (as a predictive model) and their sequence is proposed. however, there are major problems with this approach in the case of a multidimensional control task. on the one hand, it is connected with observations distribution in a highdimensional feature space (especially in the case of small number of observations). on the other hand, the nadaraya-watson estimator has the high computational complexity. the bigger feature space dimension, the harder to optimize the vector of bandwidths. decision trees are used for solving such regression tasks [8]. to avoid these problems it is suggested to use a decision tree instead of the nadaraya-watson estimator. in this paper an approach based on classification and regression tree (cart) is proposed to solve control tasks. the paper is organized as follows: in the second section, statements of the modeling and control problems are introduced; in the third section, the nonparametric modeling and control algorithm based on the nadaraya-watson estimator is presented; the fourth section is devoted to classification and regression tree (cart) and the control algorithm based on cart with its variants for different control task statements. 2. modeling and control tasks the block scheme of the control process is shown in figure 1. the following designations are taken: x̄ is the vector of output variables, x̄∗ is the vector of the desired output x̄ values, ū is the vector of controlled inputs, v̄ is the vector of observed uncontrolled inputs, ξ is the unobserved input (noise). fig.1 the block scheme of the control process it is evident from figure 1 that the output variables x̄ depend on the inputs ū, v̄, ξ. the control task is to build a control unit which generates ū such that e (x̄, x̄∗) is minimized, where e is some error measurement. the modeling (regression) and control tasks are adjacent. if a system reaction to an input (prediction) is known, it is possible to get the desired outputs by identifying the controlled inputs [9]. the forward model (regression model) x̂ (ū, v̄) is fit using the training set (ūi, v̄i, x̄i, i = 1, ..., n), where n is the observations number (figure 2). using the forward model x̂ (ū, v̄) the reactions to a given input can be predicted. for control task solving, it is necessary to invert the model x̂ (ū, v̄). it means that we should find controlled inputs ū∗ that could result in the desired outputs x̄∗ in the case of estimated copyright c© 2017 assa. adv syst sci appl (2017) 60 e. mangalova fig.2 forward model scheme (regression task) uncontrolled inputs v̄predicted (figure 3). in contrast to the forward model, the backward model û ( x̄∗, v̄predicted ) is fit using the training set (ūi, v̄i, x̄i, i = 1, ..., n) such that the outputs are given, uncontrolled inputs are predicted and we need to determine an appropriate controlled input. fig.3 backward model scheme (control task) 3. nadaraya-watson estimator approach the nadaraya-watson estimator approach was proposed to solve control tasks in such forward-backward models statement [6]. 3.1. modeling the nadaraya-watson estimator has been widely applied for the nonparametric regression, using the weighted average observations output in the neighborhood around (ū, v̄): x̂k (ū, v̄) = ∑n i=1 xi ∏mu j=1k ( uj, uji , c j u )∏mv j=1k ( vj, vji , c j v )∑n i=1 ∏mu j=1k ( uj, uji , c j u )∏mv j=1k ( vj, vji , c j v ) , k = 1, ...,mx, (3.1) where the kernel function k is a non-negative function that integrates to one and has zero mean, c̄u, c̄v are the vectors of bandwidths, mx is the number of output variables. to identify the forward models x̂k (ū, v̄), the bandwidths c̄u, c̄v should be optimized according to the selected accuracy measurement. 3.2. control backward models of nadaraya-watson estimators have the following form: copyright c© 2017 assa. adv syst sci appl (2017) regression tree control of multidimensional static object 61 ûm ( u1, ..., um−1, x̄, v̄ ) =∑n i=1 u m i ∏m−1 j=1 k ( uj, uji , c j u )∏mv j=1k ( vj, vji , c j v )∏mx j=1k ( xj∗, xji , c j x )∑n i=1 ∏m−1 j=1 k ( uj, uji , c j u )∏mv j=1k ( vj, vji , c j v )∏mx j=1k ( xj∗, xji , c j x ) , m = 1, ...,mu, (3.2) the bandwidths c̄u, c̄v in the equations (3.1) and (3.2) are not identical. there are some important limitations on the nadaraya-watson algorithm implementation: • ranking input features. the controlled input variables ū should be ranked by their importance. the process (3.2) starts from the most important controlled input feature and continues with less and less important ones. • bijective relationship. the forward models (the true relationships between input features and output features) should be bijective functions. suppose there are two significantly different controlled inputs ū∗1 and ū∗2 that produce the desired output. the result of the algorithm is a controlled input values between ū∗1 and ū∗2 (this control does not lead to the desired output) instead of choosing one of them. • curse of dimensionality. it is connected with observations distribution in a high dimensional feature space (especially in the case of a small number of observations). suppose, there are n = 1000 points uniformly distributed over the ten dimensional unit cube [0, 1]10. an average over the neighborhood of diameter 0.25 (in each coordinate) results in the volume of 0.2510 ≈ 0.00000095 for the corresponding ten-dimensional cube. hence, the expected number of observations in this cube will be 0.00095 and any averaging can not be expected. if we fix the count k = 1 of observations over which to average, the diameter of the typical neighborhood will be larger than 0.5. it means that the average is calculated over at least one-half of the range along each coordinate [10]. 4. decision tree approach decision trees are commonly used for solving regression tasks in multidimensional cases and allow preventing the nadaraya-watson algorithm limitations described in section 3. 4.1. modeling decision tree is a piecewise constant nonparametric model. the most popular decision tree model is classification and regression tree (cart) [8]. cart is a binary tree where each root node represents an input variable ujr and a split point br. let us aggregate controlled and observed uncontrolled inputs to make the description easier: ū = {u1, ..., umu , v1, ..., vmv}. each leaf node contains values of the output variable x which is used to make a prediction. cart fitting involves input variables and split points selection until a suitable tree is constructed. input variables and split points are chosen using a greedy algorithm to minimize the: min j min b ∑ i:uj ib l ( xi, x̂ + (ūi, j, b) ) (4.3) where x̂− (ūi, j, b) or x̂+ (ūi, j, b) is the prediction in the point ūi based on the subspaces after splitting: copyright c© 2017 assa. adv syst sci appl (2017) 62 e. mangalova x̂− (ūi, j, b) = ∑ i:uj ib xi∑ i:uj i>b 1 . (4.4) the tree construction ends using a predefined stopping criterion, such as the minimum number of observations assigned to each leaf node of the tree or the maximum depth of the tree. each node in the tree corresponds to a rectangular region of the predictor space sk, a subset of the observations lying in the region sk, a constant x̂k which is the average response of the observations in k-th rectangular region. thus, the binary tree model can be formalized as follows: x̂ (ū) = { x̂k : ū ∈ sk, k = 1, 2, ..., k } . (4.5) 4.2. control control tasks can be solved using the regression trees in different formulations depending on the number of input and output values, presence of observed uncontrolled input variables. consider the basic formulation of the problem. 4.2.1. controlled inputs one-dimensional output. assume there are only controlled variables ū and one output variable x. decision tree is fit using the training dataset containing simultaneous observations of ū and x. the control algorithm: 1. set the control target x∗. 2. search for such leaf node that k∗ = arg min k d1 (x̂k, x ∗) , (4.6) where d1 is a one-dimensional distance measurement. call k∗-th node ”target node”. 3. find control ū∗ contained in the region sk∗ . there are different ways to choose this control: • the center of the rectangle guaranties the most accurate forward model prediction (decision tree prediction). • points on the rectangle boundary allow to explore object and collect more information. • also the distance between the previous control and the current control can be minimized to reduce interference in the system. the experimental result. suppose the object is represented by the equation: x = sin (2πu1) + √ u2 − u3 + ξ . this equation is used in the computational experiment to simulate a real process. input variables u1, u2, u3, u4 are bounded by [0, 1]. the initial dataset contains 10 points distributed uniformly. control is selected as the center of the target node. figure 4 illustrates the desired outputs and the reactions to a generated control. figure 5 shows the initial dataset and the observations obtained during the simulation process. 4.2.2. controlled and observed uncontrolled inputs one-dimensional output. assume there are controlled and observed uncontrolled variables ū and one output variable x. the decision tree is fit using the training dataset containing simultaneous observations of ū and x. the control algorithm: 1. set the control target x∗. copyright c© 2017 assa. adv syst sci appl (2017) regression tree control of multidimensional static object 63 fig.4 simulation of a process. desired outputs and reactions to a generated control fig.5 simulation of a process. initial dataset and generated dataset 2. predict the uncontrolled input variables uj,predicted, j = mu + 1, ...,mu +mv. 3. cut nodes which can not be achieved with the predicted uncontrolled input variables. 4. search for a leaf node such that k∗ = arg min k,ū′∈sk d1 (x̂k, x ∗) , (4.7) where ū′ is the input values such that u′j ,j = 1, ...,mu can take any value and u′j = uj,predicted,j = mu + 1, ...,mu +mv. 5. find a control ū∗ contained in the region sk∗ . the experimental result. suppose now that u2 is uncontrolled and changed as follows u2 i = sin (πi/25). the previous value of u2 i−1 is used as a prediction for the next step u2,prediction. figure 6 illustrates the desired outputs and the reactions to a generated control. figure 7 shows the initial dataset and the observations obtained during the simulation process. 4.3. multi-dimensional output, controlled and observed uncontrolled inputs assume there are controlled and observed uncontrolled variables ū and the vector of output variables x̄. decision trees are fit using the training dataset containing simultaneous observations of ū and x̄. the control algorithm: 1. set the control targets x̄∗. 2. predict the uncontrolled input variables uj,predicted, j = mu + 1, ...,mu +mv. 3. find intersections of nodes which contain the predicted uncontrolled variables sk 1,...,mu , k = 1, ..., k ′ and the output values associated with these intersections x̂1 k,..., x̂mx k . copyright c© 2017 assa. adv syst sci appl (2017) 64 e. mangalova fig.6 simulation of a process with uncontrolled input. the desired outputs and the reactions to the generated control fig.7 simulation of a process with uncontrolled input. the initial dataset and the generated dataset 4. search for a leaf node such that k∗ = arg min k dmx ( x̂1 k, ...x̂ mx k , x̄∗ ) , (4.8) where dmx is a multi-dimensional distance measurement. 5. find a control ū∗ contained in the region sk∗ 1,...,mu . the experimental result. suppose now there are two outputs x1 = sin (2πu1) + √ u2 − u3 + ξ and x2 = √ u2 − u3 + ξ. figure 8 illustrates the desired outputs and the reactions to a generated control. the desired outputs for x1 and x2 are the same. figure 9 shows the initial dataset and the observations obtained during the simulation process. fig.8 simulation of a multidimensional process. the desired outputs and the reactions to a generated control copyright c© 2017 assa. adv syst sci appl (2017) regression tree control of multidimensional static object 65 fig.9 simulation of a multidimensional process. the initial dataset and the generated dataset) conclusion in this paper a control task of a multidimensional static system is discussed. the paper presents the advantages and disadvantages of the nonparametric algorithm based on the nadaraya-watson estimator and their sequence. in order to avoid the disadvantages of this nonparametric algorithm, the modification with piecewise constant approximation of the nonparametric estimator is put forward. in the proposed algorithm classification and regression trees are used as a predictive model instead of the nadaraya-watson estimator. the control algorithm based on cart is tested for different task formulations: controlled inputs and one-dimensional output, controlled and observed uncontrolled inputs and onedimensional output, controlled and observed uncontrolled inputs and multi-dimensional outputs. references 1. richalet j., rault a., testud j.l. & papon j. (1978). model predictive heuristic control: applications to industrial processes, automatica, 14 (5), 413–428. 2. psichogios d. c. & ungar l. h. (1991). direct and indirect model based control using artificial neural networks, industrial & engineering chemistry research, 30 (12), 2564– 2573. 3. draeger a., engell s. & ranke h. (1995). model predictive control using neural networks, ieee control systems magazine, 15 (5), 61–66. 4. sarimveis h. & bafas g. (2003). fuzzy model predictive control of non-linear processes using genetic algorithms, fuzzy sets and systems, 139 (1), 59–80. 5. da costa sousa j.m. & kaymak u. (2001). model predictive control using fuzzy decision functions, ieee transactions on systems, man, and cybernetics, part b (cybernetics), 31 (1), 54–65. 6. medvedev a.v., & raskina a.v. (2017). on the nonparametric identification and dual adaptive control of dynamic processes, journal of siberian federal university. mathematics & physics, 10 (1), 96–107. copyright c© 2017 assa. adv syst sci appl (2017) 66 e. mangalova 7. nadaraya e. a. (1964). on estimating regression, theory of probability and its applications, 9, 141–142. 8. breiman l., friedman j., stone c. j. & olshen, r. a. (1984). classification and regression trees. crc press. 9. kunh s. & guhmann c. (2008) modeling and control with local linearizing nadaraya watson regression, arxiv:0809.3690. 10. härdle w. (1990) applied nonparametric regression. cambridge university press. copyright c© 2017 assa. adv syst sci appl (2017) https://arxiv.org/abs/0809.3690 introduction modeling and control tasks nadaraya-watson estimator approach modeling control decision tree approach modeling control controlled inputs one-dimensional output. controlled and observed uncontrolled inputs one-dimensional output. multi-dimensional output, controlled and observed uncontrolled inputs bibliography paper title (use style: paper title) adv syst sci appl 2019; 04; 79-86 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/829 an optimal itinerary generation in a configuration space of large intellectual agent groups with linear logic dmitry maximov* institute of control science russian academy of science, moscow, russia e-mail: dmmax@inbox.ru received may 16, 2019; revised december 12, 2019; published december 31, 2019 abstract: a group of intelligent agents which fulfill a set of tasks in parallel is represented by the tensor multiplication of corresponding processes in a linear logic game category. an optimal itinerary in the configuration space of the group states is defined as a play with maximal total reward in the category. the reward is represented as a degree of certainty (visibility) of an agent goal, and the system goals are chosen by the greatest value corresponding to these processes in the system goal lattice. keywords: intellectual agent; itinerary choice; goal lattice; game semantics. 1. introduction the artificial intelligence is represented in the artificial general intelligence (agi) approach as an information processor which consumes and gives out information. investigations in the field are focused on systems which act rationally. a formal description of the most intelligent agent (aixi) behavior, in the sense of some intelligence measure, is suggested in agi framework [1]. the model is based on probabilistic modeling of the environment, and on the next system move determination based on previous experience, and on a numerical estimation of the system position reward and on the maximization of the expected reward along the trajectory. however, the method to obtain this numerical estimation is absent. also, there are no models to describe such agent groups’ behavior. it has been demonstrated in [2-4] that the structure existence (a lattice structure or else a monoid structure, i.e. the linear logic structure) in the system task [3, 4] or goal [2] set is sufficient for the system to behave quite reasonable. the behavior looks even like ants’ behavior in something [4]. but it is not supposed here the environment modeling unlike [1]. in this paper, the topic develops based on the idea that it is possible to represent different aim parallel achievement processes fulfilling by various agents in some environment as a tensor multiplication in linear logic. the logic is modeled in some game category [5]. thus, it is possible to describe goals achievement process by the intelligent agent system in some environment as a game. position rewards in the game are represented by sets which describe the information about goals or their distinctness degree. thus, the rewards are provided by the environment, and they are lattice elements but not numbers as in [1]. these lattice and linear logic structures are provided by the environment. but similar structures are also given by our ideas about the system and agent purposes. these are used to determine the priority of different parallel processes in such game category. the project was partly supported by rfbr grant 16-08-00832a * corresponding author dmmax@inbox.ru mailto:dmmax@inbox.ru 80 d. maximov copyright ©2019 assa adv. in systems science and appl. (2019) 2. mathematical backgrounds 2.1. lattices [6] definition 2.1: a partially-ordered set p is the set with such a binary relation x ≤ y for elements in it, that for all 𝑥, 𝑦, 𝑧 ∈ 𝑃 the following relationships are performed:  x ≤ x (reflexivity);  if x ≤ y and y ≤ x, then x = y (anti-symmetry);  if x ≤ y and y ≤ z, then x ≤ z (transitivity). the definition means that in the partially-ordered set not all elements are compared with each other. this property distinguishes these sets from linear-ordered ones, i.e., from numeric sets which are ordered by a norm. thus, the elements of the partially-ordered set are the objects of more general nature than numbers. in the partially-ordered set diagram, the greater the element (i.e., vertex, node) is the higher it lies, and the elements are compared with each other lie in the same path from the minor element to the greater one. two examples of a partially-ordered set diagram are represented in fig. 2.1which are also lattice diagrams. definition 2.2: the upper bound of a subset x in a partially-ordered set p is the element 𝑎 ∈ 𝑃, containing all 𝑥 ∈ 𝑋.the supremum or join is the smallest subset x upper bound.the infimum or meet defines dually as the greatest element𝑎 ∈ 𝑃 containing in all 𝑥 ∈ 𝑋. definition 2.3: a lattice is a partially-ordered set, in which every two elements have their meet, denoting as 𝑥 ⋀𝑦, andjoin, denoting as 𝑥 ⋁ 𝑦. in the lattice diagram the elements join is the nearest upper element to both of them, and the meet is the nearest lower one to both. the elements generating by joins and meets all other elements are called generators. they refer to the lattice as complete lattice if its arbitrary subset has the join and the meet. thus, any complete lattice has the greatest element “1”, and the smallest one “0” and every finite lattice is complete. fig. 2.1. examples of a goal lattice (a) and of an agent desire lattice (b) an optimal itinerary generation with linear logic 81 copyright ©2019 assa adv. in systems science and appl. (2019) 2.2. linear logic[7] if a multiplication operation is additionally defined at the lattice elements, then the operations of linear logic also exist at the lattice. we use the phase semantic of linear logic from [7]. definition 2.4: a phase space is a pare (𝑀, ⊥), where mis a multiplicative monoid (i.e., a triple (𝑀0,⋅, 𝑒) with 𝑀0 isa set and ⋅ is a multiplication with the unit e), which is also a lattice and the element false of the lattice ⊥ ⊂ 𝑀 is anarbitrary subset of the monoid. in linear logic, the element false differs from 0 (the minimal lattice element) in general in contrast to classical logic or intuitionistic one. the multiplication 𝑋 ⋅ 𝑌 = {𝑥 ⋅ 𝑦|𝑥 ∈ 𝑋; 𝑦 ∈ 𝑌 } is defined for arbitrary monoid subsets (i.e., the lattice elements) 𝑋, 𝑌 ⊂ 𝑀. the linear implication 𝑋 ⇒ 𝑌 = {𝑧|𝑥 ⋅ 𝑧 ∈ 𝑌, ∀𝑥 ∈ 𝑋} is also defined. for𝑋 ⊂ 𝑀 its dual is defined as 𝑋⊥ ⇒⊥. the dual element isa generalization of the negation in the case of linear logic. definition 2.5: facts are such subsets 𝑋 ⊂ 𝑀 that 𝑋⊥⊥ = 𝑋or equivalently 𝑋 = 𝑌⊥ for some 𝑌 ⊂ 𝑀. thus, facts are lattice elements coinciding with their double negations. e.g. ⊥⊥ = 𝐼 = {𝑒}⊥⊥; 1 = 𝑀 = ∅⊥; 0 = 1⊥ = 𝑀⊥ = ∅⊥⊥. here 1 is the maximal element of the lattice m, 0is its minimal element, e is the monoidal unit, and i is the neutral element of the multiplicative conjunction (see after this). it is easy to get the next properties: 𝑋⊥𝑋 ⊂ ⊥; 𝑋 ⊂ 𝑋⊥⊥; 𝑋⊥⊥⊥ = 𝑋; 𝑋 ⇒ 𝑌 ⊥ = (𝑋 ⋅ 𝑌 )⊥; (𝑋 ∨ 𝑌 )⊥ = 𝑋⊥ ∧ 𝑌 ⊥. from here we get only facts may be the values and the consequents of the implication. at facts the lattice operations of the additive conjunction& and the additive disjunction + are defined in the following way: 𝑋 & 𝑌 = 𝑋 ∧ 𝑌 = (𝑋⊥ ∨ 𝑌 ⊥)⊥; 𝑋 + 𝑌 = (𝑋⊥&𝑌 ⊥)⊥ = (𝑋⊥ ∧ 𝑌 ⊥)⊥ = (𝑋 ∨ 𝑌)⊥⊥ . the duality of the operationsunderstands here as in the set theory: ∨ ⊥ =∧ ∧⊥ = ∨ inwhich the duality means the negation. at facts, multiplicative operations are also defined. these are the multiplicative conjunction ⊗ and the multiplicativedisjunction ð: 𝑋 ⊗ 𝑌 = (𝑋 ∙ 𝑌 )⊥⊥ = (𝑋 ⇒ 𝑌 ⊥)⊥ = (𝑋⊥ ð 𝑌 ⊥)⊥; 𝑋 ð􀀀 𝑌 = (𝑋⊥ ∙ 𝑌 ⊥)⊥ = 𝑋⊥ ⇒ 𝑌 . theneutral element of the operation & is 1, the dual to it (neutralelement of the operation +) is 0. the neutral element of theoperation ð is ⊥, the dual to it, the neutral element of the operation ⊗ is i. the set of facts is divided into two classes dual to each other: the class of open facts op and the class of closed facts cl. the set op is closed by operations + and ⊗. its maximalelement by inclusion is i, and the minimal one is 0. the set cl is closed correspondingly by operations &and ð, and its maximal element is 1, and the minimal one is ⊥. in fig.2.1a, the class of open facts is encircled as op with 𝐼 = 𝑚1𝑒, and the class of closed facts is encircled as cl with ⊥= 𝑋3 2.3.game semantics in a linear logic category [5] definition 2.6: a conway game is defined as a rooted graph with vertices v as the game positions and edges 𝐸 ⊂ 𝑉 × 𝑉 as the game moves. each edge has a polarity ±1 which depends on whether it is the proponent or the opponent move. definition 2.7: 82 d. maximov copyright ©2019 assa adv. in systems science and appl. (2019) a trajectory or a play is some path from the graph root ∗. the path is alternated if the adjacent edges are of different polarities. definition 2.8: a strategy is defined as a non-empty set of alternated paths of even length, which are started from the opponent move, closed up to the prefix of even length and determined. determinism means that two paths with the common prefix, which differ in two moves, should coincide. definition 2.9: a dual play 𝑋⊥is obtained from the play x by reversing the polarity of moves. definition 2.10: the tensor product 𝑋⨂yof two conway games x and y is defined in such a way: positions 𝑥⨂y are 𝑉𝑋⨂𝑌 = 𝑉𝑋 × 𝑉𝑌with the root ∗𝑋⨂𝑌=∗𝑋×∗𝑌, moves are 𝑥⨂y ↦ { 𝑧⨂y, x ↦ z 𝑖𝑛 𝑋 𝑥⨂z, y ↦ z 𝑖𝑛 𝑌, and the polarity of a move in 𝑋⨂yis inheritedfrom the polarity of the underlying move in x or y. a generalized linear logic is modeled in a category of such games. the category objects are conway games and morphisms 𝑋 → 𝑌 are strategies in 𝑋⊥ðy.these multiplications are linear implications𝑋⊥ðy = 𝑋 ⇒ 𝑌. it should be mentioned that on game graphs the operations ð and ⨂ are the same so in [5] they are not even distinguished. a conway game with a payoff is the play with an additional weight 1, 1/2 or0 in each vertex. the weight depends on whether the position is winning (i.e., of the weight 1 or 1/2) or not. in the tensor product, these weights obey rules of boolean conjunction and implication. a strategy is winning if it terminates in the winning position. in the category of conway games with a payoff, morphisms are winning strategies now. it is possible to prove [8] that the categorical construction is conserved if the weights’ numbers are replaced with some sets form a lattice, and the boolean operations are replaced with the lattice operations. the greater set is connected with a position, the more advantageous it is, and it is winning if its weight is not 0. we suppose the existence of a universal set containing all the others. thus, all such estimation sets form a complete brouwer lattice. 3. behavior determination of an intellectual agent system an example of the goal lattice ms matches a system of l agents (fig. 2.1,a). in the lattice vertices, xi and e are the generators and denote the system goals. vertices ji are the generators joins and denote combined goals achieving. these combined goals, as well as generators, may be associated with the correspondent task fulfilling thus the vertices may have meets, i.e., some subtask included in different tasks (or correspondent goals). the higher goal lies (hence the more tasks it contains), the more important this behavior variant is. it supposed that the preferable behavior of the system is to achieve all its goals. this variant corresponds to the top lattice element 1. and the bottom element 0 corresponds to complete inactivity and to the least essential behavior variant. all the estimations may be considered as partially true truth values. thus, we can say that the more important the behavior is, the truer it is. the agent lattice mi (fig. 2.1, b) has another meaning. the same generators and their joins mean the agents’ desires. one of the vertices is marked as an intention for every agent. in the same manner, the more desires are included in the intention, the more essential the intention is. a variant of the open (op) and the closed (cl) classes definition is indicated in the system lattice (fig. 2.1, a). this definition (as well as the multiplication definition) defines the structure of linear logic in the lattice. the element multiplication is obtained (usually ambiguously, [4]) from the demand to implement the linear logic operations properties. in this case, it is possible to consider an optimal itinerary generation with linear logic 83 copyright ©2019 assa adv. in systems science and appl. (2019) parallel processes of combined goals achieving as the tensor product of corresponding lattice elements in the logic. and the priority of different processes is obtained from the demand of the greatest correspondent tensor product estimation in the lattice (the product is an element of the lattice; hence the higher it lies, the more important it is, i.e., the greater its truth-value is). it is supposed some initial agent goal distribution. for such a system, we consider the process of the interaction of the system with the environment as a conway game. in the game, the environment is the proponent which provides the system (the opponent) with the information about the environment objects. the opponent moves from one position in the environment to the other by the use of the information to achieve his goals. for the system of l agents, we, in fact, have l parallel processes and, therefore, the resulting game is their tensor product. thus, agents are placed initially in the configuration space (environment) in the root ∗=∗1 ⨂ … ⨂ ∗𝑙 of the system game 𝐴 = 𝐴1 ⨂ … ⨂𝐴𝑙 with the system goal lattice ms and agent intention lattices mi, 𝑖 = 1. . 𝑙. the game aj represents possible agent moves in the environment. but the real trajectory or the play is chosen from the demand of the maximal total position reward along the projected path. the agent move in the environment is estimated corresponding to an optimality criterion with the reward 𝑘(𝑝𝑖, 𝑏𝑗) in the position 𝑝𝑖of the goal 𝑏𝑗achieving process. it may be that the system does not see any goal initially and moves according to a criterion of an optimal move in goals absence. thus, the system has additionally l goals (tasks) 𝑎𝑗of movement in the environment, which are included as in the generator set of the lattice ms, as in the generator sets of lattices mi. the optimality criterion may represent the highest degree of the correspondence to the demanding system configuration, or the highest freedom in future moves or the most excellent visibility from a position or so on in the case of the free movement in the goals absence. and in the case of the system goals achieving we suppose the better a goal object is visible, the higher the reward is. it is supposed that agents see not the whole environment, but up to some horizon which may change in each direction fig. 3.1. moreover, the agent sees the environment as in the fog – the nearer the object, the better it is perceived†. therefore, we can predict a play only up to some finite step number with the increasing uncertainty along the path. let us n goals b1…bn are discovered in the environment by the system with information about them (or some other reward) 𝑘(𝑝𝑖, 𝑏𝑗) in positions 𝑝𝑖of the play 𝐴 = 𝐴1 ⨂ … ⨂𝐴𝑙of l agents. the game a corresponds to l parallel processes of achieving l movement goals 𝑎𝑗. then, a winning strategy of the game𝐴′ = 𝐴1 ⨂ … ⨂𝐴𝑙 ⇒ 𝐵1 ⨂ … ⨂𝐵𝑘defines a transition (morphism) to this new game 𝐴′. this game corresponds to l parallel processes of moving and achieving those k goals from discovered n ones, which may be better achieved in the next sense (that means that these k goals are the system intentions). the reward 𝑘(𝑝𝑖 , 𝑏𝑗) is some set. in the case of, e.g., unmanned vehicles, it may be an image of the object. the better the object is visible, the better the coincidence of the image with the original and the greater the reward. † it is possible to consider different horizons and different uncertainty degree for different goals. we do not do it here. 84 d. maximov copyright ©2019 assa adv. in systems science and appl. (2019) fig. 3.1.the visibility horizon it is reasonable to choose the trajectory (play) from the demand to maximize the reward along the path within the visibility horizon: 𝑘𝑝𝑙𝑎𝑦 (𝐴1 ⨂…⨂𝐴𝑙)⊥ð𝐵1 ⨂…⨂𝐵𝑘 = (3.1) = 𝑚𝑎𝑥𝑝𝑙𝑎𝑦𝑠 [ ⋃ 𝑘(𝐴1 ⨂…⨂𝐴𝑙)⊥ 𝑝𝑙𝑎𝑦 & ⋃ 𝑘𝐵1 𝑝𝑙𝑎𝑦 & … & ⋃ 𝑘𝐵𝑘 𝑝𝑙𝑎𝑦 here the reward 𝑘𝑝𝑙𝑎𝑦 (𝐴1 ⨂…⨂𝐴𝑙)⊥ð𝐵1 ⨂…⨂𝐵𝑘 is maximized in the game 𝐴′ and corresponds to that process of k goals achieving that has the highest priority (𝑎1 ⨂ … ⨂𝑎𝑙) ⊥ð𝑏1 ⨂ … ⨂𝑏𝑘 in the system goal lattice. the priority is maximal among all possible parallel processes of achieving n discovered goals. the maximum is taken among all possible plays and it joins the rewards along these plays (i.e., trajectories) in (𝐴1 ⨂ … ⨂𝐴𝑙)⊥, and in 𝐵1, and so on and in 𝐵𝑘. thus, the sign & means conjunction. if there are parallel processes which are not compared by priority, i.e., if there are several incompatible priorities(𝑎1 ⨂ … ⨂𝑎𝑙)⊥ð𝑏1 ⨂ … ⨂𝑏𝑘, it is possible to reorder the goal lattice in the manner that some lattice vertex has used as an additional priority [3]. in the reordered lattice, initially incompatible elements may become compatible, so we can choose the processes’ priority. such a lexicographical rule follows from the teleological system assignment, i.e., our vision of the system purpose. it may be an occasion that the lattice is so symmetric that does not allow make such a choice. in this case, the ambiguity of the linear logic structure allows the necessary variant choosing [8]. fig. 3.2. an example of a game (depicted only the opponent moves) an optimal itinerary generation with linear logic 85 copyright ©2019 assa adv. in systems science and appl. (2019) also, if we are not able to determine the trajectory in (3.1) uniquely, it is possible to use the goal lattices of different agents to determine the agent (and the correspondent trajectory) priority as in [2]. in this case, we assign a definite agent to achieve a particular goal from agents’ intention estimations. specifically, let us there are three agents’ plays a1, a2 and a3 in which the agents in positions p1 have discovered two goals b1 and b2 from possible three ones (fig. 3.2). let us also the rewards in the positions p1 satisfy the next relations: 𝑘𝐴1 (𝑝1, 𝑏1) > 𝑘𝐴2 (𝑝1, 𝑏1)< >𝑘𝐴3 (𝑝1, 𝑏1) (3.2) 𝑘𝐴3 (𝑝1, 𝑏2) > 𝑘𝐴2 (𝑝1, 𝑏2)< >𝑘𝐴1 (𝑝1, 𝑏2) (3.3) where the sign < > means the rewards’ incompatibility. this means that though the visibility degrees in positions p1 of a2 and a3 does not differ, the foreshortenings are different thus the images are different and, therefore, incompatible. therefore, we cannot define the system trajectory uniquely. it is clear that agent-1 should achieve the goal b1because it sees it better than the others. for the others let us, in this case, the desires’ lattices for these two agents have the form of fig. 3.3. thus, the agent-2 may reach goals b1 and b2, and the agent-3 has an additional desire b3. then, we may evaluate the lattice vertices’ weights by the formula [2]: vertex weight = joined in the vertex desires /total desires’ number (3.4) the formula evaluates vertices’ weights by two parameters: the number of desires joined in the intention and the vicinity of the vertex to the top element, i.e., to the most desirable behavior variant. the play of the two parameters allows weights’ comparing in different lattices. thus, e.g. the u12 weight is 𝑉𝑈12 = 2/3. since 𝑉𝑏2 = 1/2 in the agent-2 lattice and 𝑉𝑏2 = 1/3 in the agent3 lattice, we should choose the bigger value and correspondent agent-2 to achieve the goal b2.this is reasonable because the agent-3 is left free and can look for the goal b3. finally, the obtained play is depicted in fig. 3.3 by bold lines. fig. 3.3. the agent-2 lattice (a) and the agent-3 lattice (b) the additional subtle point is that the graphs of the games a and b do not always coincide, i.e., a goal may be seen in b, but the way to the goal may not exist in a (i.e., in the environment). thus, we have a deal with the game a of l parallel processes of agent moves, on the positions of which the rewards of k games b are possible but are not obligatory. it should be pointed out that the structures of tensorial multiplications ⨂ and ð in (𝑎1 ⨂ … ⨂𝑎𝑙)⊥ð𝑏1 ⨂ … ⨂𝑏𝑘 and in (𝐴1 ⨂ … ⨂𝐴𝑙)⊥ð𝐵1 ⨂ … ⨂𝐵𝑘 are different. in (𝑎1 ⨂ … ⨂𝑎𝑙)⊥ð𝑏1 ⨂ … ⨂𝑏𝑘the tensor ⨂ is a monoid in the goal lattice. these structures of ðand ⨂ are pre-existent and does not depend on the environment. it may be chosen from very general considerations [4] and determines the system purpose. but the tensor ⨂ and the co-tensor ðin(𝐴1 ⨂ … ⨂𝐴𝑙)⊥ð𝐵1 ⨂ … ⨂𝐵𝑘are determined by the environment, by the goals and obstacles 86 d. maximov copyright ©2019 assa adv. in systems science and appl. (2019) distributions, by visibility and so on. thus, there are two different linear logic structures in the approach. 4. conclusion in the paper, a movement of an agent group with its goal lattice was considered in some environment. each agent move was represented as a game in which the environment corresponds to the proponent which provides the opponent (the agent) with some information to achieve the agent goal. different agent moves are considered as parallel processes which are represented as the tensor product of correspondent games and form a comprehensive game. the game has rewards on its positions, which estimates the quality of the information provided by the environment. the reward may be very different: it may represent the highest degree of the correspondence to the demanding system configuration, or the highest freedom in future moves, or the most excellent visibility from a position, or the degree of the coincidence of the goal object image with the original or some else. we demand the greatest total reward along all agents’ plays to choose the group trajectory. and we choose those goal achieving processes (i.e., agent plays) from all possible, which have the highest estimation in the agent group goal lattice. it is so because every goal and the corresponding process of its achieving has a definite correspondent truth value in the lattice. the higher the value lies in the lattice diagram; the higher priority of the process is. thus, we consider two types of estimations: the goal lattice value defines the choice of the goal achieving processes from all possible ones, and the position rewards of the game determine the optimal trajectory of these chosen processes in the environment. when it is not possible to determine the most significant estimations in both these cases, some additional methods may be used to select the optimal variant. references 1. hutter m. (2012). one decade of universal artificial intelligence, in theoretical foundations of artificial general intelligence (pp. 66–88), atlantis press. 2. legovich y. s. & maximov d. y. (2017). selecting executor in a group of intellectual agents, automation and remote control, 78 (7), 1341–1349. 3. maximov d. y. (2016). reconfiguring system hierarchies with multi-valued logic, automation and remote control, 77(3), 462–472. 4. maximov d. y., legovich y. s. & ryvkin s. e. (2017). how the structure of system problems influences system behavior, automation and remote control, 78(4), 689–699. 5. mellies p. -a. & tabareau n. (2010). resource modalities in tensor logic, ann. pure. appl. logic, 161(5), 632–653. 6. birkhoff g. (1967). lattice theory. rhode island, providence. 7. girar j. -y. (1987 ). linear logic, theoretical computer science, 50, 1–102. 8. maximov d. y. (2019). formirovanie optimal’nogo marshruta bol’shih grupp intellectual’nyh agentov [an optimal itinerary generation of large intelligent agent groups], large-scale systems control, 78, 46–70. adv syst sci appl 2018; 02; 93-106 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/531 risk analysis in seawater desalination sector: a case study of beni saf water company “bwc” mohammed bouamri1, hassiba bouabdesselam1 1) lte research laboratory, national polytechnic school of oran, algeria e-mail:bouamrimohammed25@gmail.com, has.bouab@yahoo.fr abstract: in this present paper, a risk analysis approach is applied to an algerian reverse osmosis seawater desalination plant using the mads mosar method. mads mosar method is a stepwise risk analysis approach containing many phases. our work begins with analyzing a review of past accidents triggered by the ben in saf water company (bwc) seawater desalination plant locating in the algerian coast in ain temouchent region and analyzing their similar seawater desalination plants (or plants that using similar and potential equipment). then, the mads mosar method will apply essentially for the macroscopic vision (module a). the macroscopic vision corresponds to a main risk analysis. in the current case study, we were able to identify eight subsystems where sources and scenarios of hazards are identified, accident scenarios are assessed, recognized and ranked by "severity×probability" grid. we found twenty-six scenarios whose we were able to assess them in function of their probability and severity using "probability x severity" grid criteria. at the end of the analysis, we were able to define and suggest the most appropriate prevention and protection barriers for the potential elements in the studied seawater desalination plant, including pipelines, transformers, compressors, high pressure pumps, pressure vessels and energy recovery devices. keywords: risk analysis; beni saf water company (bwc); seawater desalination plant; mads mosar. 1.introduction a hazard is a situation or substance that has the potential to cause harm while risk is the likelihood or probability of a certain undesired event to occur within a certain period of time or under specified circumstances [1]. although the bwc seawater desalination plant has become a priority for the algerian freshwater production (provides 200.000 m3/d of water to meet the needs of ain temouchent and oran regions), there are also potential downsides. apart hazards inherent of any seawater desalination plant (chemicals occur in the desalination process from numerous origins. these include the source water and chemicals that are used in the treatment process to aid its efficient functioning, to ensure microbiological safety, to stabilize the water before it enters the distribution system, and to control corrosion from contact surfaces during storage and distribution to consumers [2] and the presence of electrical equipments), the plant studied using the reverse osmosis ro membrane type that operates with high pressures than what is normally found in conventional plants assured by high pressure equipments. membrane desalination processes use semi-permeable membranes with pumping pressure as the driving force-to separate a saline feedwater into two streams: a low-salinity product (permeate) and a high-salinity stream (concentrate or reject) [2]. according to [3] 40 % of total cost of water turns to seawater ro system (percentage could vary depending on project-specific factors). reliability, service time and safety of these 94 m. bouamri, h. bouabdesselam copyright ©2018 assa. adv. in systems science and appl. (2018) equipments must be considered. as regards death and injuries experienced in such desalination unit, one could refer to two deaths caused primarily by not observing the exiting safety and hse instructions [4]. safety precautions should be taken in the plant and shutdown-startup procedures were developed [5] and process monitoring devices are provided to prevent the damage of components due to process upsets [6]. also, employers should be aware of these hazards and even workplaces must be appropriately designed to ensure proper operations and better health and safety conditions (workers, the public, property and the environment safety). in chemical industries, risk assessment is an approach for better and more efficient management of the process safety, whereby risk is estimated and considering various factors involved, decisions are made on the appropriate tolerability required [6, 7]. 1.1 case study: beni saf water copany seawater desalination plant the first step of our risk analysis study is to describe study area and desalination process used. based on a-7-years history (during 2010) of the bwc plant operation, maintenance records and accident reports, the second step is getting past internal/external accident review (the return of experience ) for enhanced risk prevention. then, the mads mosar method is applied and recommended preventives and protective barriers are proposed. 2.study area description before hazard identification can begin, it is necessary to understand the system operations so that the risk analysis covers all activities of interest [7]. some though be given to grasping how the system works and how the hardware, software, people, and environment all interact [8]. clearly, we need to define the system and theirs functions to bind our analysis. this case study concerns beni saf water company bwc seawater desalination plant located on the mediterranean in the algerian coast in ain temouchent region, which has an area of 65,700m2, with a 200 000 m3/d production capacity. bwc has worked since march 2010. the location of the bwc plant is presented in the following map (fig. 1). fig.1. location of the bwc seawater desalination plant (source: author) risk analysis in seawater desalination sector 95 copyright ©2018 assa. adv. in systems science and appl. (2018) 3. process description bwc uses the reverse osmosis’s method for the seawater desalination. in general, the main treatment unit includes the following stages: seawater intake and pumping, pretreatment, membrane ro separation unit and post-treatment. seawater intake system allows seawater to flow to a desalination unit where minerals are then removed from the saline water through a desalination process. at present, practically all ro desalination plants incorporate two key treatment steps designed to sequentially remove suspended and dissolved solids from the source water. the purpose of the first step-source water pretreatment-is to remove the suspended solids and prevent some of the naturally occurring soluble solids from turning into solid form and precipitating on the ro membranes during the salt separation process [9]. pressures applied in reverse osmosis applications vary between 15 bar (brackish water desalination) and 60 to 80 bar in seawater desalination [10]. which is 65 bar in the bwc case. the ro system consists of a number (usually 2-18) of individual ro trains, each of which is capable of independently producing desalinated water from pretreated source water [11]. once the desalination process is complete, the freshwater produced by the ro system is further treated for corrosion and health protection and disinfected prior to distribution for final use. this third step of the desalination plant treatment process is referred to as post-treatment [9]. fig. 2 presents a general schematic process of the bwc plant. fig.2. schematic of bwc plant (source: author) 4. implementation of the mads mosar medthod: macroscopic vision risk management: a continuous management process with the objective to identify, analyze, and assess potential hazards in a system or related to an activity, and to identify and introduce risk control measures to eliminate or reduce potential harms to people, the environment, or other assets [12]. two main purposes of the risk management are to ensure that adequate measures are taken to protect people, the environment and assets from undesirable consequences of the activities being undertaken, and to balance different concerns, for 96 m. bouamri, h. bouabdesselam copyright ©2018 assa. adv. in systems science and appl. (2018) example safety and costs [13]. the risk analysis discussed here is one that anticipates and prevents undesired events during the implementation of the combustion pilot, using mads– mosar. mads refers to “the analysis method of dysfunctional systems”, and mosar refers to “the organized and systemically method of risk analysis”. mads proposes a general model of hazard, mosar builds a global methodology for the risk analysis [14,15]. the retained approach is the "mads-mosar" methodology was developed by a team of specialists from the cea (french atomic energy authority) [16]. it makes it possible to model our industrial field by a "man-installation-environment" system [17, 18]. it is an integrated approach that allows progressive analysis of risks an industrial site [19]. mosar method consists of two modules (module a and b) which can be used more or less independently. module a corresponds to a macroscopic analysis of the risks of an industrial site and requires a preliminary risk analysis. module b is used for a more detailed analysis of the scenarios identified in module a, realized with specific instruments of safe operation. the two modules are almost the same structure [20]. mads mosar structure process is shown in fig.3. fig.3. mads mosar process steps (adapted from [21]). as shown in the figure above, the first step of the mosar method consists of modeling the bwc plant by means of functional division into subsystems. so that for each subsystem (ssi) the type of hazard source is identified and the hazardous processes are defined (short and long scenarios). hazard identification is an iterative process. to identify each hazard, we need to decompose the system into many subsystems. 4.1. systems identification and modeling the main system constituting the study context is the industrial installation on which the risk analysis is carried out (the bwc desalination plant in our case). there are several possible divisions, the aim being to make the complex entity to be studied in sub-units (subsystems) simpler. from the previous description of the process, we can distinguish eight subsystems: ss1: seawater intake and pumping section; ss2: seawater pretreatment section; risk analysis in seawater desalination sector 97 copyright ©2018 assa. adv. in systems science and appl. (2018) ss3: reverse osmosis section; ss4: post-treatment section; ss5: electrical substation; ss6: administrative building; ss7: human; ss8: environment. our system can be broken down into following subsystems (fig.4). fig.4. system decomposition into subsystems. 4.2. identify the sources of hazards the source of hazard should not be ignored. the first step is to identify the sources of hazard for each subsystem or to identify how each subsystem can be a source of hazard. by making this identification for all subsystems, a list of the hazards of the installation is obtained. table 1 illustrates the areas, subsystems decomposition and hazard sources associate to each subsystem. table 1. sources of hazards presented the bwc plant area subsystems (ssi) hazard sources seawater intake and pumping area ss1: seawater intake and pumping -seawater intake pipelines -intake screens -feed pumps (11 pumps) -seawater storage reservoir -electric substation -sodium hypochlorite tank (120 m3) -2 compressors ss2 : seawater pretreatment -ferric chloride tanks (34 m3) -sodium metabisulfite tanks (14 m3) -antiscalant storage tank (7 m3) -sulfuric acid storage tanks (2x 100 m3) -sand filters (48 reservoirs) -antracite filters (28 reservoirs) -cartridge filter (20 filters) ss3 :ro unit -10 high pressure pumps (65 bar) -ro membrane trains racks (10 98 m. bouamri, h. bouabdesselam copyright ©2018 assa. adv. in systems science and appl. (2018) production area racks) :  1792 membranes -energy recovery devices (22 units)  pressure exchanger  booster pumps and valves -256 pressure vessels -caustic soda tank (1 m3) ss4 : post-treatment  -chemical dosing system: neutralisation or chlorination  -sodium hypochlorite tank  -produced water tank  -brine disposal  -10 pumps  -chemical cleaning tanks  -sand and anthracite filters cleaning  -03 membranes cleaning pumps.  -produced water tanks : 11 pumps & pipelines ss5 : electrical substation -transformer power station ss6 : administrative building -offices -laboratory -control room ss 7 : human -operators trainees ss 8 : environment climatic and natural conditions material environment to identify hazards, our work begins with analyzing a review of past accidents triggered by the bwc plant and their similar facilities. based on the analysis of the accident balance sheets, an example of the repartition per month by accident triggered by the bwc plant during 2010 is shown in the following diagram (fig.5). fig.5. accident distribution per month during 2010 (source: bwc) 0 0,5 1 1,5 2 2,5 3 3,5 number of accident number of accident risk analysis in seawater desalination sector 99 copyright ©2018 assa. adv. in systems science and appl. (2018) then, and according to accident reports and accident balance sheets, we analyzed accidents triggered by major desalination plants in algeria. as is apparent from the table 2, 37.5% of accidents are related to the transformer accidents while 25% of the accidents caused pipeline leak. also, 25% of the accidents present an environmental pollution while 12,5% of them are due to an electrical cause (short-circuit generated fire in the storage area). table 2. accident triggered by the algeria's desalination plants (source: author) desalination plant years accidents arzew (oran) 17-10-2005 transformer fire arzew (oran) 12-02-2006 hydrocarbon pollution skikda 10-02-2008 fire in the storage area ain temouchent 14-04-2010 explosion of water hp pipeline mers el hadjadj (oran) 28-07-2011 transformer fire kahrama (oran) 19-05-2012 hydrocarbon pollution honaine (tlemcen) 25-08-2012 transformer explosion honaine (tlemcen) 26-09-2016 water leak results transformer accidents 37,5 % pipeline leaks 25 % environmental pollution 25 % electrical causes 12.5 % 4.3. identify the scenarios of hazards mosar method insist on hazard chaining flow between component systems of industrial plants and is specially adapted for studying the effects of simultaneous accidents or "domino“ effects [19]. each subsystem in table a (from module a) is characterized by inputs (initiating events) and outputs (principal events). we can also make the relation to find the danger flux to compete the models. the table a is the best tool to establish that. as a part of our risk analysis study, table 3 presents initiating events, initial events and main events (flux of danger) for the ss5 for the element transformer. table 3. identify the source of hazard for the transformer source of danger phase of process initiating events initial events main events (flow of danger) ss5 electrical substation external internal container contonent transformer exp -transformer overheating -thermal radiation -no compliance with the safety instruction -maintenance operations -as a consequence of the lightning -hazardous discharge of static electricity -humidity -effect of defective joint -low impedance faults -electrical malfunction -corroded cables -internal heat -locking of the disc -electric arc -short-circuit -over-voltage -degradation of insulation like decay of the transformer oil due to moisture or ageing or decomposition -loss of containment -corrosion of the tank -crack in the wall of the oil conservator corrosion and rupture of the tank due to aging or magnetic shock insufficient maintenance or too infrequent inspection. -oil overheating -oil leak -accidental oil spill inflammation of the transformer oil contained inside the metal envelope the presence of flammable gas -explosion -fire -pollution human damages -production shutdown -atmospheric pollution due to smokes -loss of cooling oil in transformer -materials loss or damage (premises destruction, plant damage) 100 m. bouamri, h. bouabdesselam copyright ©2018 assa. adv. in systems science and appl. (2018) using the black-box concept by subsystems such as presented the fig.6, we can make the reduction of the variety. short scenarios of undesired events can be structured from the links between them. from table a and short scenarios, we can easily define long undesired scenarios (as an example, fig.7 presents long scenarios for transformer explosion). 4.4. assess the scenarios of risks probability assessment can be made, at the option of the analyst, qualitative, semiquantitative or quantitative, using classical instruments (trees of failures, trees of events). mosar method explicitly provides for identify and assess security measures and distinguish between technical measures and usage measures (called in other methods, human protective measures”) [22]. in this study, 26 accident scenarios were identified as those most relevant to our system and analyzed. the figure bellow (fig.8) shows the combination of risk severity and probability for each scenario. it allows scenarios to be classified and decided the relative priority. fig.8. severity x probability combination. as mentioned in the figure above, scenarios 5, 10, 11, 19 and 20 have a high “severity x probability” scores whereas scenarios as 16, 18 and 26 present the low “severity x probability” scores. the rest scenarios have medium “severity x probability” scores. sc 1 sc 2 sc 3 sc 4 sc 5 sc 6 sc 7 sc 8 sc 9 sc 10 sc 11 sc 12 sc 13 sc 14 sc 15 sc 16 sc 17 sc 18 sc 19 sc 20 sc 21 sc 22 sc 23 sc 24 sc 25 sc 26 probability 4 4 2 2 3 2 2 2 2 3 1 3 2 3 2 1 1 1 3 3 2 3 3 3 3 1 severity 2 2 2 1 4 2 2 2 2 4 3 4 2 2 2 1 3 1 4 4 2 2 1 2 2 1 0 1 2 3 4 5 6 7 8 severity x probability evaluation 101 ss1 ss3 ss5 ss8 ss7 -earthquake -mechanical shock -bad welding -breaking fiberglass reinforced polyster seawater collector -external/internal corrosion -design, operation or manufacture defects -loss of seal overpressure -pressure controller failure -compressor components failure -sodium hypochlorite injection pump failures -defective valve -flood -total/partial shutdown -explosion -disrupt operation or damage the apparatus -fragment projections -materials damage -pollution -insufficient maintenance or too infrequent inspection -stress and fatigue -lack of adequate training -human error -non-compliance with the safety instructions -non-compliance with service parameters -overpressure -suction pressure too high -suction and/or discharge valves closed or clogged -excessive air entrapped in liquid -pump is run dry -pump run off design point air leak in suction line -cavitation -inadequate lubrication inadequate lubricant cooling -internal/external corrosion -overheating lack of seal flush at seal faces chemicals in liquid other than specified -mechanical failure -bursting of simple pressure vessels -total/partial shut-down -explosion -fire -noises -fragment projections -materials damage -internal heat -locking of the disc due to wear or sand -advanced pollution -imprudence of a smoker -thermal radiation -transformer overheating -insufficient maintenance -mechanical/electrical malfunction -corrosion and rupture of the tank due to aging or magnetic shock -oil leak -hazardous discharge of static electricity -short-circuit -total/partial shut-down -explosion -fire -fragment projections -materials damage -dust storm -earthquake -mechanical degradation -pollution -atmospheric pollution -flood -mechanical or thermal shock lightning humidity -advanced electrical substation pollution -materials damage -overheating -overpressure -control instruments failures -explosion -fire -short-circuit -injuries fig.6. examples of short and long accident scenarios 102 copyright ©2018 assa. adv. in systems science and appl. (2018) dielectric oil lack neighbor equipment bringing of a thermal radiation nearby fire corroded cables maintenance operations crack in the wall of the oil conservator corrosion of the tank short-circuit imprudence of a smoker locking of the disc due to wear or sand faulty seal causing an oil leak lightning consequences internal heat accumulation of dust and grease deposits explosion of the transformer electric arc overheating transformer oil leak and gate or gate no smoking leak detector periodic inspection preventive maintenance leak detector lightning protection electrical grounding disc checking insulating oil retention pond periodic inspection corrosion protection emergency plan fig.7. long scenarios for ss5: transformer explosion risk analysis in seawater desalination sector 103 copyright ©2018 assa. adv. in systems science and appl. (2018) risk management plan is a step in the overall risk management procedure following the risk assessment step. after all risks are identified in the risk assessment step, risks that are not acceptable must be selected. the main task in the risk management plan step is to treat each selected unacceptable risk. to perform risk quantification, a risk matrix must be developed for each accident initiator from the corresponding accident matrix [7]. once scenarios have been identified, they must be assessed and grouped in a grid by their severity and probability as shown in the figure 9. fig.9. “severity x probability” grid the “severity x probability” grid (fig.9) shows that only 42.3% of the scenarios are located in the acceptable zone (green zone). however, 57.7% of the scenarios are located in the unacceptable zone (red zone). 4.5. define the means of prevention and qualify the barriers the last step of the mads mosar approach is to suggest/define and to quantify barriers. in order to improve safety approach, preventive and protective appropriate barriers are cited in table 6. the table below presents some proposed preventive and other protective barriers that should be available for potential elements (elements that may present a significant hazard) in the bwc seawater desalination plant. these elements are respectively: pipelines, transformers, high pressure pump, compressors, pressure vessels and energy recovery devices. 104 copyright ©2018 assa. adv. in systems science and appl. (2018) table 6. recommendations component recommendations pipelines 1the implementations of an accidental spill prevention plan for preventing and controlling accidental spills or discharges (especially brine discharges). 2damaged or leaking containers will be isolated, when possible, in a containment area or repackaged to prevent loss, exposure or hazards. 3leak detection of water pipeline. 4periodic inspection of the water pipeline systems (routine maintenance of piping system). 5the replacement of old pipes with new pipelines. 6prevention of corrosion. 7preventive and corrective maintenance of the fiberglass reinforced polyster seawater collector (estimate the lifetime of collectors and change them periodically). transformer 1transformer status check/control. 2preventive maintenance of the various safety and protection devices of the transformer (relays, fuses, circuit breakers, powder and carbon dioxide extinguishers). 3dielectric oil transformer quality control (transformer oil aging) and replacement. 4compliance with safety instructions. 5design to the appropriate electrical standards. 6periodically temperature oil checks. 7gas emission control and periodic oil analysis. high pressure pumps 1preventive and corrective maintenances are recommended. 2install flow indicators. 3increase pressure indication and alarms. 4preventive and corrective maintenance can reduce the physical hazard of noise. 5recommend anti-vibration pump core. compressors 1installing a pressure sensor. 2carbon dioxide extinguishers. 3emergency response for projections accidents. 4fire fighting equipment, procedures and alarms for emergency response. 5personal protective equipment (ppe) uses. pressure vessels 1periodic pressure vessels inspection. 2operating service pressure respect. 3respect of the safety distance (explosion prevention). 4pressure control instrument. erd (energy recovery devices) 1periodic verification is recommended. 2leak and seal detectors. 3ensure the erd support. 4never exceed permitted flow limits. 5parameters service respect (flow and pressure: erd recommends au minimum 1.0 bar). 6erd cavitation prevention (most recognized problem for erd). 5. conclusion in this study, we carried out the bwc plant risk analysis throughout its life cycle, including seawater intake and pumping, pretreatment, reverse osmosis desalination, posttreatment, electrical substation, environment and human. to control or mitigate risks those appear to be the most critical, we were applied the mads mosar method. the most important problem in the bwc plant or in the similar plants lies in ro membranes in several pressure vessels, which are placed in a certain order. this process requires that a higher pressure (more than 65 bars) be exerted on the high concentration side of the membrane which would render the unit being out of service and pipes being burst. https://www.travelers.com/boilerre/boiler-reinsurance-services/risk-control/risk-control-guides.aspx risk analysis in seawater desalination sector 105 copyright ©2018 assa. adv. in systems science and appl. (2018) many hazards can be related to the bwc plant equipments such as the transformers, compressors, high pressure pumps and the energy recovery devices. they can involve fires, explosions, high-voltage electrical arcs, fragments projections, oil ignition and dispersions, and potential injuries or may up to death. they make noises and they also give off heat during operation. there is also a wide range of chemicals in the workplace (for the seawater desalination process or laboratory analysis). hazardous chemicals can be a health, physicochemical or environmental hazards if not handled or stored correctly. the safe clean up of a chemical spill requires knowledge of the properties and hazards posed by the chemical to also protect the environment. this study offered valuable results in terms of safety as indeed the recommendations and improvements yielded are expected to lead directly/indirectly to help the bwc plant in its safety approach. the risk management plan is a continuous risk controlling and monitoring process. risk assessment is only one step in the overall approach to risk management. it belongs to a dynamic approach. in addition, it is recommended that a new risk assessment be carried out on a regular basis to determine whether the risks have been eliminated definitively or whether other risks have arisen since the last assessment. the actions resulting from the evaluations must be performed, formalized and documented. references [1] pollard, s.j.t. (2007). risk management for water and wastewater utilities. water and wastewater technologies, united kingdom: iwa publishing. [2] cotruvo, j., voutchkov, n., fawell, j., payment, p., curnliffe, d et al. (2010). desalination technology health and environmental impacts. chemical aspects of desalinated water. boca raton, florida: crc press. [3] voutchkov, n. (2018, august 21). seawater reverse osmosis design and optimization. [online]. available https://web.stanford.edu/group/ees/rows/presentations/voutchkov.pdf [4] rezaei, k & aziznezhad, v & chenani, a & tabibian, m & mozdianfard.m.r. (2016).risk analysis of the sea desalination plant at the 5th refinery of south pars gas company using hazop procedures. journal of fundamental and applied sciences. 8(2s) 602-613, doi: http://dx.doi.org/10.4314/jfas.8vi2s.75. [5] abdel-fatah, m. a & el-gendi, a & ashour. f. (2016). performance evaluation and design of ro desalination plant: case study. journal of geoscience and environment protection. 4, 53-63. [6] mannan, s. (2005). lee’s loss prevention in the process industries, 3rd edition, vol. 1, oxford, uk: elsevier butterworth heinemann,. [7] bley, d.c., droppo j.g., eremenko v.a., lundgren r. e. (2003). risk methodologies for technological legacies: proceedings of the nato advanced study institute, bourgas, bulgaria. [8] bahr, n.j. (2015). system safety engineering and risk assessment: a practical approach, 2nd edition, new york: taylor & francis. [9] voutchkov, n. (2013). desalination engineering: planning and design. mcgraw hill professional. [10] fritzmann, c., lowenberg, j., wintgens, t., melin t. (2007). state-of-the art of reverse osmosis desalination, desalination, 216, 1-76. https://web.stanford.edu/group/ees/rows/presentations/voutchkov.pdf http://dx.doi.org/10.4314/jfas.8vi2s.75 106 copyright ©2018 assa. adv. in systems science and appl. (2018) [11] voutchkov, n. (2014). desalination engineering: operation and maintenance. mcgrawhill professional publishing. [12] rausand, m. (2010). risk assessment: theory, methods, and applications, hoboken, n.j.: j. wiley & sons. [13] aven, t. (2011). quantitative risk assessment: the scientific platform, cambridge university press, cambridge. [14] périlhon, p. (2000). analyse des risques, élément méthodiques, phoebus, la revue de la sureté de fonctionnement, l’analyse de risques, 12, 31–49. [15] bultel, y., aurousseau, m., ozil, p., & perrin, l. (2007). risk analysis on a fuel cell in electric vehicle using the mads/mosar methodology, process safety and environmental protection, 85(3), 241-250. [16] ayrault, n. (2005). evaluation des dispositifs de prévention et de protection utilisés pour réduire les risques d’accidents majeurs (dra-039), rapport oméga 10 – evaluation des barrières techniques de sécurité, ineris, france. [17] marref, a.s., bahmed, b.l. & benoudjit, a.c. (2014).multi-field analysis of the systems sociotechnic in the industrial developing countries: case of the industry of cotton in algeria. international journal of innovation, management and technology. 5(4), 287-293. [18] périlhon, p. (1998). risk analysis risk: development of a method mosar organized and systemic risk analysis method, school of mines saint-etienne. [19] felegeanu, d –c., nedeff, v. & panainte, m. (2013).analysis of technological risk assessment methods in order to identify definitory elements for a new combined/complete risk assessment method, journal of engineering studies and research, 19 (3), 32-43. [20] perilhon, p. (2000). mosar, techniques de l'ingénieur, traité sécurité et gestion des risques, se, 4(060), 1-16. [21] laurent, a. (2013). sécurité des procédés chimiques connaissances de base et méthodes d’analyse des risques, lavoisier tec & doc. [22] perrin, l & felipe, m.g & dufaud, o & laurent, a. (2012). normative barriers improvement through the mads/ mosar methodology, safety science. 50(7), 1502 1512. http://eu.wiley.com/wileycda/section/id-302479.html?query=marvin+rausand adv syst sci appl 2017; 4; 78-92 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/521 copyright ©2017 assa. adv. in systems science and appl. (2017) developing a strategy of environmental management for electric generating companies using dea-methodology svetlana ratner, pavel ratner v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: lanaratner@gmail.com abstract: this paper investigates the possibility of utilization of the data envelopment analysis (dea), which is a non-parametric method of optimization, to solving problems of environmental management in electric generating companies. an advantage of dea is the possibility to work with dmus without any knowledge of the actual functional relation between inputs and outputs. we analyze methods of incorporating the negative ecologic effects into a model and propose an algorithm for applying the basic dea ccr input-oriented model twice in succession for the purpose of developing an optimal (ecologically and economically) strategy for environmental management in electricity energy generating companies. the developed method consists of sequentially solving several dea models: the first-stage model determines the effectiveness of dmus from an ecologic perspective and calculates target values for decreasing negative ecologic effects of non-effective dmus. the second stage requires solving one input-oriented ccr model for each non-effective object, using economic and social characteristics of projects meant to reduce negative environmental influence, and using the target values calculated in the first stage as outputs. besides the problem of evaluation the comparative efficiency of dmus, ecologically oriented studies also often needs to evaluate the changes in a dmu’s efficiency dynamically. for this, the malmquist productivity index (mpi) is used. mpi is a non-parametric method for analyzing time series that allows to track changes in dmu efficiency over time by means of dea models. we test this algorithm on the statistical data provided by russian electric companies for the period 2009-2011, and discuss methods for its practical application. the statistical data used in our calculations is averaged, and the results do not reflect the entire picture and should not be used to judge the quality of ecologic management in these companies. nevertheless, the calculations can be used to evaluate the progress of completion and the practicality of investment projects of companies, from an environmental viewpoint. they also may be used to help develop state programs for support of modernization in electric energy industry, ecologic standards or energy-saving programs. keywords: data envelopment analysis, non-parametric optimization, ecologic effects, environmental management, ecology management, electric companies. 1. introduction currently, electric generating objects that operate on hydrocarbon fuels are one of the largest emitters of greenhouse gases and other atmospheric pollutants (16% of total emissions from stationary sources in russia), as well as consumers of fresh water (35% of total water use in russia), pollutants of soil, underground and surface waters. increasing the ecologic efficiency of electric generating companies is one of the primary conditions for sustainable development of both this industry and the country itself. one of the most important problems for ecologic optimization of the development of electric energy generation industry is decreasing negative environmental influence as much as possible, via a variety of environmental protection measures (both technologic and organizational), while maintaining the existing volumes of electricity production [20]. investment priorities for energy companies directed towards decreasing negative ecologic effects are determined primarily by conventional system of ecologic penalties for overhttp://ijassa.ipu.ru/ojs/ijassa/article/view/521 79 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) standard pollution and rarely correspond to the actual ecologic situation of the region. many sources (e.g. [5, 19]) consider the existing russian economic mechanisms and incentives for minimizing negative environmental influence ineffective. at the same time, there are currently no examples of successful transition to ecologically safe energy generation technologies at regional or national scales anywhere in the world. the well-known brazil energy crisis of 2001 that happened due to the droughts and lack of water for hydro-power plants [9], the recent chinese ecologic crisis due to the increase of coal use for power generation, the non-proportionally intensive growth of solar power plants in czech republic [10] show the complexity of the problem of optimizing the energy system configuration. a lack of complete understanding of the way certain energy-generating processes influence ecosystems as well as a lack of attention to ecologic aspects of energy systems on their planning stage can lead to unexpected and undesired results, where the decrease of negative influence from one parameter (or group) is completely overshadowed by the increase in another parameter (or group). for instance, attempting to capture со2 to decrease greenhouse gas emissions leads to a significant increase in water use by power plants [11]. increasing the amount of optimization criteria for incorporation while planning the structure of a region energy system (decrease of negative environmental influence, decrease of energy production price, maximization of useful social and economic effects) served as another reason for using non-parametric methods of data envelopment analysis (dea) for solving this class of problems. among the multitude of approaches to modelling energetic and ecologic problems in foreign literature, dea has attained a leading position. one of the main reasons is the possibility to model comparatively the efficiency of various energy sectors in a variety of countries, which has become especially relevant with the liberalization of energy markets [24]. currently, dea is a well-developed methodology for comparing the efficiency of various homogenous economic agents, operating in production or elsewhere, via a variety of mathematical programming models. the agents the efficiency of which is evaluated by dea are usually known as “decision-making units” (dmus). all dmus perform the same production function that transforms a set of inputs into a set of outputs. an advantage of dea is the possibility to work with dmus without any knowledge of the actual functional relation between inputs and outputs. russian economists conventionally employ dea to analyze the efficiency of the budgetary system, regional authorities or banking structures, etc. [17, 18]; however, during the recent years, studies that use dea to analyze ecologic aspects of economic activity, including electric energy [12], started appearing. the classic dea model known as ccr (named after its developers: charnes a., cooper w.w., and rhodes e. [4]) involves solving a fractional linear programming problem that maximizes the ratio of the linear combination of weighted outputs to the linear combination of weighted inputs:    k kjk i iji vu xv yu h , max (1.1) this ratio is known as the efficiency coefficient, and its value lies between zero and one. any dmus with their efficiency coefficient equal to one are considered efficient, and all the others are, inversely, deemed ineffective. if the efficiency coefficient is defined in the form of (1.1), the dea problem itself is considered to be “input-oriented”. the other form for defining the efficiency coefficient (a ratio of the linear combination of weighted inputs to the linear combination of weighted outputs) is known as “output-oriented”. in our case, we’re dealing with an output-oriented problem. developing a strategy of environmental management for electric generating 80 copyright ©2017 assa. adv. in systems science and appl. (2017) calculating the projections to the efficiency frontier for inefficient dmus in the input/output space allows us to determine the estimated targets for decreasing inputs or increasing outputs. achieving these projected values will allow the dmu to become efficient. a peculiarity of using dea for optimizing energy systems is the presence of so-called undesirable outputs, that is to say, the negative ecologic effects. for solving ecological problems, a special class of dea models was developed: these are known as “environmental dea (edea)”. the goal of this paper is to review methods and approaches of accounting for undesirable outputs in environmental dea models and developing an algorithm for applying the basic input-oriented dea model for performing a comparative analysis of the ecologic efficiency of large electric generating companies of russia. we tested the capabilities of the developed two-stage algorithm on 24 dmus on its first stage and on detailed data of 11 power plants (that are part of ogk-2) on its second stage. 2. methogology overview: dea models used for optimizing energy systems based on ecologic criteria in its coefficient form, the classic input-oriented ccr dea model is as follows: 0 1 , max m m m m vu yu  (2.1) s.t. ;,2,1,2,10, ,1 ,,2,10 0 1 11 nnmmvu xv kkxvyu nm n n n n nk n n n m m mkm          where 0 – index of the dmu being optimized, x – input vector of size n, y – output vector of size m, к – amount of dmus. or, in dual form:   min (2.2) s.t: kk mmyy nnxx k mok m m mk nok n n nk    ,2,1,0 ,2,1, ,2,1, 1 1           this model searches for the possibility of proportionally decreasing inputs without a decrease in outputs. the ccr production set is the following set of vectors (x, y):              n j n j jjjjj njyyxxyxт 1 1 ,1,0,,),(  (2.3) 81 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) classic dea models, including ccr, assume their inputs and outputs are strictly monotonous, in other words, the production set follows this rule: if tyx );( and xx  or yy  , then tyx  );( or tyx );( (2.4) however, this property does not always describe the real production situation. for instance, energy generation via hydrocarbon fuel will always be linked to the production of sulfur dioxide, and decreasing this output without a decrease in the corresponding input is technologically impossible. thus, using a production set that follows (2.4) will lead to incorrect modelling results. a literature review allows one to conclude, that a significant number of attempts to consider undesirable outputs in dea models have been made. these fall into two main categories: i) recalculating (modifying) original data and using a traditional dea model [22]; ii) using original data with models based on the concept of weak disposability [6-8]. when using the first approach, the overall efficiency of a company may be divided into technical/economic efficiency, defined as a ratio of the weighted sum of wanted outputs to the weighted sum of inputs, and ecologic efficiency, which is defined as the ratio of weighted sums of wanted and unwanted outputs. let the first k of m outputs of the model (2.1) be desirable, and the others undesirable. then, the economic efficiency of dmu with the index of zero can be represented as:     m i ii k r rr economy xv y h 1 0 1 0 (2.5) and the ecologic efficiency as:     p ks ss k r rr yeco y y h 1 0 1 0 log   (2.6) to incorporate both efficiency measures in the basic ccr model, we have to somehow combine them in a way that corresponds to the general logic of the problem: maximization of desirable outputs and minimization of undesirable outputs and inputs. the following option (a) fits these conditions rather well:       m i ii p ks ss k r rr a xv yy h 1 0 1 01 0  besides that, undesirable outputs can be treated as inputs of the model (option b), which transforms the efficiency measure as follows:       p ks ss m i ii k r rr b yxv y h 1 01 0 1 0   . in this case, the decrease in undesirable outputs happens simultaneously with the decrease of inputs. paper [17] provides the proof that basic ccr models that use option a are analogous to those which employ option b. the production set corresponding to the property of weak disposability is defined thusly: developing a strategy of environmental management for electric generating 82 copyright ©2017 assa. adv. in systems science and appl. (2017)                                  ;,2,1,0 ,,2,1 ,,2,1 ,,2,1|),,( 1 1 1 kkz jjuuz mmyyz nnxxzuyx t k k k jjkk k k mmkk k k nnkk e     where u is a vector of undesirable outputs, e is an index with the meaning of “environmental”, since these types of production sets are used in environmental dea. besides that, et also fulfills the following condition: if tuyx ),,( and u=0, then y=0 (2.7) this indicates that a complete elimination of undesirable outputs is only possible via a complete halt of the production process. this approach can be used for undesirable inputs as well: models with weakly disposability of inputs are introduced in [7, 16]. besides disposability, operational characteristics of inputs and outputs (their unit measures) may also serve as a distinct characteristic of environmental dea. two models [12] often mentioned in the literature use categories for input and output variables. these models describe production processes during implementation of environmental laws rather well, and allow to consider the influence from other external factors that dmus themselves have no control of. return to scale is also an important characteristic of the dea production set. the basic ccr model uses constant return-to-scale. if we add the following condition to the set t:    k i i 1 1 , we’ll end up with the bcc model with variable returns to scale (while production grows, the scale effect changes from increasing to decreasing). another popular approach is the use of non-increasing return to scale, defined as:    k i i 1 1 . besides the orientation of the efficiency measure (input, output or undesirable output), an important property of environmental dea problems is the method for reducing inputs and increasing outputs, that is to say, the direction of movement towards the efficiency frontier. radial efficiency measures are the ones most frequently encountered in any dea models. in this case, the inputs decreasing proportionally by the same value of oc co  (radial movement from source point to efficiency frontier) (see fig. 1). 83 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 2.1 graphical illustration of the radial efficiency measure in a basic ccr model (2d case) combining radial efficiency measures with different production sets allows us to obtain a variety of dea models, including the basics: ccr and bcc. for example, the papers [7, 22, 23] use the production set et with a radial efficiency measure: etuyx ),,(:min 000  non-radial efficiency measures allow to decrease inputs and increase outputs nonproportionally, and usually have a better discriminative power than radial measures. a particularly well-known example is the non-radial russell measure:         n n n tyx n 1 00 ),(: 1 min θ , where θ is a diagonal matrix containing n ,1 . if n ,1 have different weights, a weighted non-radial efficiency measure may be used. such measures reflect the preferences of specific dmus [25]. a hyperbolic efficiency measure, also known as the graphical measure, decreases inputs and increases outputs by the same value at the same time (moving towards the efficiency frontier hyperbolically):       t y x ),(: 0 0 θ θθ this efficiency measure works best when both desirable and undesirable outputs are present. the directional distance function (ddf) is an efficiency measure that allows one to simultaneously increase desirable outputs and decrease inputs using a directed vector. this is a generalized form of the traditional radial efficiency measure [3]. 3. results 3.1. analyzing the comparative efficiency of energy generating companies in russia with ecologic indicators let’s consider the problem of evaluating the efficiency of the russian energy generating companies based on a set of ecologic parameters. to calculate the ecologic efficiency with the basic input-oriented ccr model (2.1) with a radial efficiency measure, we’ll use freely available statistical information [15] on ecologic aspects of the activity of the main players on the electric energy market: five wholesale generating companies (ogk) that unify the largest heat power-plants, and some territorial generating companies (tgk) that unify the developing a strategy of environmental management for electric generating 84 copyright ©2017 assa. adv. in systems science and appl. (2017) power plants of several neighboring regions that did not become parts of ogk and work as isolated energy systems (24 dmus total). we view atmospheric emissions (in thousands of tons), solid waste (in thousands of tons) and freshwater consumption (in millions of cubic meters) as our unwanted outputs. we consider the generated electricity as our sole wanted output. the results of our calculations, done in the maxdea software, using a radial and a nonradial efficiency measure, represented in table 3.1. table 3.1. scores of ecologic efficiency of generating companies in 2011 company name radial efficiency non-radial efficiency ogk-1 0.374 0.190 ogk-2 0.168 0.115 ogk-3 0.142 0.101 ogk-4 ojsc «e.on rossiya» 0.611 0.543 ogk-5 «enel ogk-5» 0.132 0.093 tgk-1 0.300 0.265 tgk-2 0.112 0.111 ojsc «mosenergo» (tgk-3) 1.000 1.000 tgk-4 ojsc «kvadra» 0.558 0.477 tgk-5 0.428 0.402 tgk-6 0.591 0.412 ojsc «volzhskaya tgk» (tgk-7) 1.000 1.000 tgk-9 0.122 0.096 ojsc «fortum» (tgk-10) 0.366 0.299 tgk-11 0.532 0.206 ojsc «kuzbassenergo» (tgk-12) 0.115 0.081 ojsc «eniseyskaya tgk» (tgk-13) 0.089 0.075 tgk-14 0.717 0.475 generiruyushchiye kompanii «lukoyl» 1.000 1.000 ojsc «dal'nevostochnaya gk» 0.177 0.091 ojsc «irkutskenergo» 0.528 0.284 ojsc «tatenergo» 1.000 1.000 ojsc «bashkirenergo» 0.678 0.616 ojsc «sibeko» 0.136 0.101 the dmu efficiency score herein is to be interpreted as a ratio of the minimal possible negative ecologic effects to the real ones. that is to say, the effective dmus are those who use the best available technologies and the cleanest fuel (from the ecologic point of view). the efficiency coefficient for the effective dmus is equal to 1, and highlighted in bold. it is easy to note that the efficiency coefficient of non-effective dmus is greater when calculated with the radial measure than with the non-radial one. the target indicators (for 2011) that have to be reached by non-efficient companies (calculated under a radial efficiency measure) to become efficient are presented in table 3.2. table 3.2. values of target indicators for inputs that need to be reached in 2011 company name target indicators emissions waste water consumption ogk-1 34.256 84.108 332.035 ogk-2 63.444 155.770 614.933 ogk-3 26.513 65.097 256.985 85 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) ogk-4 ojsc «e.on rossiya» 55.864 147.106 412.725 ogk-5 «enel ogk-5» 43.741 121.476 241.682 tgk-1 16.385 29.442 158.045 tgk-2 9.305 26.654 40.894 tgk-4 ojsc «kvadra» 11.685 19.870 112.633 tgk-5 10.929 31.453 46.131 tgk-6 15.532 18.012 270.476 tgk-9 14.927 42.799 65.070 ojsc «fortum» (tgk-10) 18.787 51.954 106.641 tgk-11 9.231 26.567 38.964 ojsc «kuzbassenergo» 19.991 52.513 149.342 ojsc «eniseyskaya tgk» 11.266 31.170 63.775 tgk-14 26.226 75.477 110.700 ojsc «dal'nevostochnaya gk» 23.341 67.172 98.520 ojsc «irkutskenergo» 63.063 181.488 266.183 ojsc «bashkirenergo» 20.362 53.769 148.491 ojsc «sibeko» 11.243 31.417 59.607 when the indicators given in table 3.2 are reached, each of these companies can match its reference point on the efficiency frontier in the multi-dimensional input/output space. the process of calculation of target indicators under non-radial efficiency measure has shown that their values change only in case then dmu has several benchmarks (see table 3.3.) table 3.3. changes in target indicators depending from the type of efficiency measure company name reference points δemissions δwaste δwater ogk-4 «mosenergo»; «tatenergo» 5.93 24.50 -71.27 ogk-5 «mosenergo»; «tatenergo» -3.51 -14.52 42.22 tgk-1 «mosenergo»; «volzhskaya tgk» 3.39 -2.45 32.13 tgk-2 «mosenergo»; «tatenergo» -0.10 -0.41 1.20 tgk-4 «mosenergo»; «volzhskaya tgk» 2.77 -2.01 26.27 tgk-6 «mosenergo»; 6.25 -4.77 180.53 tgk-9 gk lukoyl -0.13 -0.52 1.53 tgk-10 «mosenergo»; «tatenergo» -1.68 -6.96 20.24 tgk-12 «mosenergo»; «tatenergo» -4.00 -16.53 48.08 tgk-13 «mosenergo»; «tatenergo» -0.99 -4.13 12.01 ojsc «bashkirenergo» «mosenergo»; «tatenergo» 2.25 9.30 -27.05 ojsc «sibeko» «mosenergo»; «tatenergo» -0.75 -3.09 8.99 moving towards the efficiency frontier non-radially leads to some target parameters being greater than in the radial efficiency measure (such differences are indicated as negative developing a strategy of environmental management for electric generating 86 copyright ©2017 assa. adv. in systems science and appl. (2017) numbers), and in some cases, they will be smaller (which is indicated by a positive number). which of these two efficiency measures is the best depends on the expenditure of the specific company trying to reach these target indicators. 3.2. dynamic of ecology efficiency of generating companies in this section we investigate how ecologic efficiency of generating companies changes through time. table 3.4 shows efficiency scores of generating companies throughout 20092011, calculated according an input-oriented ccr model with a radial efficiency measure. table 3.4. ecology efficiency of generating companies in 2009-2011 company name 2009 2010 2011 ogk-1 0.374 0.387 0.374 ogk-2 0.173 0.165 0.168 ogk-3 0.171 0.158 0.142 ogk-4 ojsc «e.on rossiya» 0.624 0.577 0.611 ogk-5 «enel ogk-5» 0.147 0.126 0.132 tgk-1 0.534 0.378 0.300 tgk-2 0.130 0.115 0.112 ojsc «mosenergo» (tgk-3) 1.000 0.993 1.000 tgk-4 ojsc «kvadra» 0.588 0.400 0.558 tgk-5 0.505 0.464 0.428 tgk-6 0.338 0.366 0.591 ojsc «volzhskaya tgk» (tgk-7) 0.731 0.627 1.000 tgk-9 0.136 0.118 0.122 ojsc «fortum» (tgk-10) 0.438 0.308 0.366 tgk-11 0.667 0.714 0.532 ojsc «kuzbassenergo» (tgk-12) 0.123 0.107 0.115 ojsc «eniseyskaya tgk» (tgk-13) 0.097 0.078 0.089 tgk-14 0.099 0.079 0.717 generiruyushchiye kompanii lukoyl (before 2010: «yugk tgk-8») 1.000 1.000 1.000 ojsc «dal'nevostochnaya gk» 0.208 0.189 0.177 ojsc «irkutskenergo» 0.719 0.597 0.528 ojsc «tatenergo» 1.000 1.000 1.000 ojsc «bashkirenergo» 0.527 0.510 0.678 ojsc «sibeko» 0.142 0.126 0.136 in 2009 and 2010, out of all analyzed dmus, only 3 companies: ojsc “mosenergo” and the generating companies of the “lukoil” and “tatneft” groups, were ecologically effective. another company joined this list in 2011: ojsc “volzhskaya ttk”. however, it is not correct to judge changes in actual ecologic efficiency merely by these coefficients, since the efficiency frontier itself may shift from year to year. to evaluate the dynamic changes of dmu efficiency in dea problems, the malmquist productivity index (mpi) is used. this is a non-parametric method for analyzing time series [13]. in its general from, the mpi can be defined with a distance function as its base, however, it can also be presented as a ratio of efficiency measures. 87 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) let ),( 00 ttt yx and ),( 00 1 ttt yx be the input-oriented efficiency measures for dmu0, calculated using inputs and outputs from the solution of (2.2) during the moment t and the production set t, in the moments t and t+1. let ),( 1 0 1 0  ttt yx и ),( 1 0 1 0 1  ttt yx be the inputoriented efficiency measures for dmu0, calculated using inputs and outputs from the solution of (2.2) during the moment t+1 and the production set t, in the moments t and t+1. then, the input-oriented mpi can be defined as: 2 1 00 1 1 0 1 0 1 00 1 0 1 0 0 ),( ),( ),( ),(          ttt ttt ttt ttt yx yx yx yx mpi     (3.1) values of mpi0<1, mpi0=1 and mpi0>1 signify, respectively, a decrease, an increase or constancy in the efficiency of dmu0 throughout the research period [13]. various sources often employ the following form for the malmquist index: ),( ),( ),( ),( ),( ),( 00 1 0 1 0 12 1 1 0 1 0 1 1 0 1 0 00 1 00 0 ttt ttt ttt ttt ttt ttt yx yx yx yx yx yx mpi                  (3.2) the form (3.2) defines changes in efficiency in a decomposed format, with the first part representing the frontier shift effect, and the second part – the catch-up effect. the calculated malmquist indices are shown in table 3.5. table 3.5. malmquist index values (output-oriented) throughout 2009-2010 and 2010-2011 company name mpi 20092010 mpi 20102011 ogk-1 0.9522 0.9426 ogk-2 0.8574 0.9784 ogk-3 0.9522 0.9426 ogk-4 ojsc «e.on rossiya» 0.6700 1.0242 ogk-5 «enel ogk-5» 0.8805 1.0170 tgk-1 0.5365 0.6175 tgk-2 0.8805 0.9484 ojsc «mosenergo» (tgk-3) 0.9723 0.9865 tgk-4 ojsc «kvadra» 0.6705 1.3079 tgk-5 0.8805 0.9200 tgk-6 0.8469 1.2539 ojsc «volzhskaya tgk» (tgk-7) 0.7708 1.2641 tgk-9 0.8805 0.9520 ojsc «fortum» (tgk-10) 0.8891 1.1160 tgk-11 0.8805 0.9200 ojsc «kuzbassenergo» (tgk12) 0.8556 1.0659 ojsc «eniseyskaya tgk» (tgk-13) 0.8157 1.1212 tgk-14 0.8805 1.2143 generiruyushchiye kompanii lukoyl 1.0171 1.0569 ojsc «dal'nevostochnaya gk» 0.8805 0.9200 ojsc «irkutskenergo» 0.8805 0.9200 ojsc «tatenergo» 0.9383 1.0000 ojsc «bashkirenergo» 0.7630 1.3424 developing a strategy of environmental management for electric generating 88 copyright ©2017 assa. adv. in systems science and appl. (2017) ojsc «sibeko» 0.8550 1.0862 thus, a real increase in environmental efficiency has only been observed within the “lukoyl” generating companies in 2009-2010, and it has decreased for the other companies. during 2010-2011, already 12 companies have exhibited an increase in ecology efficiency. malmquist index values signifying an increase in efficiency are highlighted in bold in the above table. 4. policy application: elaboration of an optimal environmental management strategy for generating companies it is necessary to mention that all of the calculations done so far only considered ecologic efficiency of dmus, without taking into account any economic parameters of power plant operating. in fact, incorporation of economic parameters in analysis is possible with the use of methods described in section 2. however, we believe that elaboration of proactive environmental investment strategy for a generating company can be done with a far simpler approach, which involves solving the two basic dea problem consistenly. on the first step, we solve the problem of evaluating ecologic efficiency of the company’s activity using primary ecologic indicators, and calculate target values if the company ends up ineffective. on the second stage, we build another dea model for each ineffective dmu, which uses economic parameters (cost, implementation time, etc.) of environmental projects (meant to reduce negative ecologic effects down to the calculated target value, or a value sufficiently close to that) as inputs. solving these models allows us to choose projects that are most effective at achieving positive ecologic results, from an investor’s point of view. sequentially processing data for each of the power plants that is part of an ogk or a tgk, we obtain a problem for optimizing the development of the entire energy system. let us consider the problem of evaluating the efficiency of specific power plants that are part of ogk-2 (which is recognized as ineffective in 2011). table 4.1 shows the results of a solving ccr dea model with a radial efficiency measure, calculated in maxdea, using data from the company’s official 2013 reports (http://www.ogk2.ru). table 4.1. ecologic efficiency of ogk-2 power plants (ccr model) power plant efficiency coefficient input target indicators emissions, thousands of tons waste, thousands of tons water, millions of cubic meters adlerskaya 1 743.25 3.09 642.33 kirishskaya 1 2415.57 3.40 574645.2 krasnoyarskaya-2 0.066 1572.01 1.26 34978.45 novocherkasskaya 0.057 3262.63 2.57 52722.81 pskovskaya 0.876 488.38 0.69 116182.1 ryazanskaya 0.422 3507.38 14.60 3031.13 serovskaya 0.039 645.04 0.49 3281.98 stavropol'skaya 0.744 2489.71 3.51 592281.4 surgutskaya -1 1 7432.52 5.57 21622.2 troitskaya 0.237 1818.42 7.57 1571.507 cherepovetskaya 0.041 932.33 0.72 11876.88 89 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) it can be seen that three power plants that belong to ogk-2 (adlerskaya, kirishskaya and surgutskaya-1) are effective from an ecologic point of view. all three of these power plants use gas as their primary source of fuel; however, the pskovskaya and stavropolskaya power plants use gas also and are, nevertheless, ineffective. this means that the fuel type is not the only factor of ecologic efficiency. the “closest” statistical indicator (meaning-wise) to our calculated efficiency coefficient, that is used in the practice of energy management in generating companies is the ratio of fuel consumption to generated energy, which characterizes the efficiency of different generating equipment (the linear pearson correlation coefficient, calculated with ogk-2 power plant data, equals -0.81). improving ecologic indicators for generating companies is possible thanks to the implementation of the following projects, unrelated to changing fuel or the technology used for energy generation [12, 15, 19-20]: 1) reduction of atmospheric emissions by implementation of low-toxicity burners for highly-concentrated dust, thereby decreasing nitrogen oxides, installation (or repairing) ventilation technologies, energy filters and ash collectors; 2) reduction of water consumption by cleaning and discarding mineralized waste waters, installation (or repairing) treatment facilities for industrial wastewater, implementation of oil collectors; 3) reduction of waste generation by recycling ashes, and/or developing technologies to increase the reliability of their storage; 4) increasing overall environmental efficiency by implementation iso 14001 compliant environmental management systems (currently implemented on the stavropolskaya, pskovskaya, surgutskaya, serovskaya and troitskaya power plants), decrease fuel expenditure for energy generation. besides that, one also needs to consider the advanced technological capabilities for replacing existing fuel with cleaner alternatives or switching to completely different energy generation technologies (options of innovative development). the least expensive fuel option is currently to use natural gas for energy generation, and the least expensive options for new energy technologies, according to russian experts, are photovoltaics and wind energy [15]. atom energy and “clean coal” technologies that minimize atmospheric pollution are midrange in terms of expense. cost and time indicators of the aforementioned projects (as well as some indicators of their social efficiency) can be used as input parameters for a next-level ccr dea model. same model can be used to optimize economic and social parameters of investment project meant to decrease negative environmental influence of the non-efficient power plants. as outputs, we can use the target indicators of ecologic effects calculated during the last stage. the overall algorithm with which the model can be built is shown on figure 4.1. we cannot provide an example calculation for this case due to a lack of available statistical data on economic values of investment projects. developing a strategy of environmental management for electric generating 90 copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 4.1. algorithm of second-level dea model creation 6. conclusion the main result of this study is the proof of possibility for using a basic input-oriented ccr dea model for solving the problem of choosing an optimal environmental management strategy for a generating company. it is important to point out that each of the large generating companies analyzed in this research consists of several separate dmus – power plants that work using different equipment and different generating technologies. therefore, the statistical data used in our calculations is averaged, and the results do not reflect the entire picture and should not be used to judge the quality of ecologic management in these companies. nevertheless, the calculations can be used to evaluate the progress of completion and the practicality of investment projects of companies, from an environmental viewpoint. they also may be used to help develop state programs for support of modernization in electric energy industry, ecologic standards or energy-saving programs. the developed method consists of sequentially solving several dea models: the firststage model determines the effectiveness of dmus from an ecologic perspective and calculates target values for decreasing negative ecologic effects of non-effective dmus. the second stage requires solving one input-oriented ccr model for each non-effective object, using economic and social characteristics of projects meant to reduce negative environmental influence, and using the target values calculated in the first stage as outputs. the choice of an efficiency measure (radial or non-radial) depends on the cost of the projects, and can only be determined while solving the second-stage dea problems (determining economic efficiency). besides the cost, the choice of an efficiency measure can be made by judging the importance of a specific ecologic effect for the specific region or territory. besides the problem of evaluation of the comparative efficiency of dmus, ecologically oriented studies also often need to evaluate the changes in a dmu’s efficiency dynamically. for this, the malmquist productivity index (mpi) is used. mpi is a non-parametric method for analyzing time series that allows to track changes in dmu efficiency over time by means of dea models. 91 s. ratner, p. ratner copyright ©2017 assa. adv. in systems science and appl. (2017) acknowledgements this study is partly supported financially by the russian foundation for basic research (rfbr), project no. 16-06-00147 “development of data envelopment analysis models for optimizing development of regional economic systems based on ecologic parameters”. references [1] banker, r.d. & morey, r.c. (1986) efficiency analysis for exogenously fixed inputs and outputs, operations research, 34, 513–521 [2] banker, r.d. & morey, r.c. (1986) the use of categorical variables in data envelopment analysis, management science, 32, 1613–1627 [3] chung, y.h., fare, r. & grosskopf, s. (1997) nproductivity and undesirable outputs: a directional distance function approach, journal of environmental management, 51, 229–240 [4] cooper, w.w., seiford, l.m. & tone, t. (2006) introduction to data envelopment analysis and its uses: with dea-solver software and references. springer, new york. [5] fadeyeva, a.v. (2007) protivorechiya v ekologo-ekonomicheskoy sisteme sovremennogo rossiyskogo obshchestva kak faktor aktivizatsii investitsiy v chelovecheskiy kapital [contradictions in the ecological and economic system of modern russian society as a factor of activation of investments in human capital] ekonomicheskiye nauki, 2(27), 5-11 [in russian] [6] fare, r., grosskopf, s. & hernandez-sancho, f. (2004) environmental performance: an index number approach, resource and energy economics, 26, 343–352 [7] fare, r., grosskopf, s. & lovell, c.a.k. (1994) production frontiers. cambridge university press, cambridge. [8] fare, r., grosskopf, s. & pasurka jr., c.a. (2001) accounting for air pollution emissions in measures of state manufacturing productivity growth, journal of regional science, 41, 381– 409 [9] global wind energy council. (2007). global wind 2006 report. [online]. available http://gwec.net/publications/global-wind-report-2/global-wind-2006-report/ [10] han, l. (2015, april). the brief and wondrous life of solar energy development. [online] available http://www.pritomnost.cz/en/economics [11] international energy agency. (2013). world energy outlook 2012. special topics. [online] available https://www.iea.org/publications/freepublications/publication/world-energy-outlook2012.html [12] khrustalev, ye.yu. & ratner, p.d. (2015) analiz ekologicheskoy effektivnosti elektroenergeticheskikh kompaniy rossii na osnove metodologii analiza sredy funktsioniro-vaniya [analysis of ecological effectiveness of electric energy companies of russia based on dea-methodology]. ekonomicheskiy analiz: teoriya i praktika, 35, 33-42 [in russian] [13] korhonen p.j. & luptacik m. (2004). eco-efficiency analysis of power plants: an extension of data envelopment analysis, european journal of operational research, 154, 437-446 [14] lo k. a. (2014). critical review of china rapidly developing renewable energy and energy policies, renewable and sustainable energy reviews, 29, 508-516 developing a strategy of environmental management for electric generating 92 copyright ©2017 assa. adv. in systems science and appl. (2017) [15] ministry of energy of russian federation. (2012) funktsionirovaniye i razvitiye elektroenergetiki v 2011 godu. informatsionno-analiticheskiy doklad. [online]. available at https://www.minenergo.gov.ru/system/download-pdf/3399/3196 [16] oudelansink, a. & bezlepkin, i. (2003). the effect of heating technologies on co2 and energy efficiency of dutch greenhouse firms, journal of environmental management, 68, 73–82 [17] piskunov, a.a., ivanyuk, i.i., danilina, ye.p., lychev, a.v. & krivonozhko, v.ye. (2008) sistema reytingovaniya regionov s ispol'zovaniyem metodologii asf [rating regions with dea]. vestnik aksor, 4, 24-30 [in russian] [18] piskunov, a.a., ivanyuk, i.i., lychev, a.v. & krivonozhko, v.ye. (2009) ispol'zovaniye metodologii asf dlya otsenki effektivnosti raskhodovaniya byudzhetnykh sredstv na gosudarstvennoye upravleniye v sub"yektakh rossiyskoy federatsii [utilizing dea-methodology for evaluation of budget spending on state management effectiveness in the regions of russian federation]. vestnik aksor, 2, 28-36 [in russian] [19] ratner, s.v. & almastyan, n.a. (2014) ekologicheskiy menedzhment v rossiyskoy federatsii: problemy i perspek-tivy razvitiya [environmental management in russian federation: the problems and perspectives]. natsional'nyye interesy: prioritety i bezopasnost', 17, 37-45, [in russian] [20] ratner, s.v. & almastyan, n.a. (2015) rynochnyye i admini-strativnyye metody regulirovaniya negativnym vozdey-stviyem ob"yektov elektroenergetiki na okruzhayushchuyu sredu [market and administrative methods of regulation of negative environmental impacts of electric generation on environment]. ekonomicheskiy analiz: teoriya i praktika, 16, 2-15, [in russian] [21] seiford, l.m. & zhu, j. (2002). modelling undesirable factors in efficiency evaluation, european journal of operational research, 142, 16–20 [22] tyteca, d. (1996). on the measurement of the environmental performance of firms – a literature review and a productive efficiency perspective, journal of environmental management, 46, 281–308 [23] tyteca, d. (1997) linear programming models for the measurement of environmental performance of firms – concepts and empirical results, journal of productivity analysis, 8,183–197 [24] zhou p., ang b.w. & poh k.l. (2008). a survey of data envelopment analysis in energy and environmental studies, european journal of operational research, 189, 1-18 [25] zhu, j. (1996) data envelopment analysis with preference structure, journal of the operational research society, 47, 136–150. java based distributed learning platform adv syst sci appl 2019; 01; 116-140 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/591 models of continuous dynamics on the 2-simplex and applications in economics denis stijepic*1 1) university of hagen, department of macroeconomics, hagen, germany e-mail: denis@stijepic.com received may 16, 2018; revised january 21, 2019; published april 15, 2019 abstract: in this paper, we discuss the models of continuous dynamics on the 2-simplex that arise when different qualitative restrictions are imposed on the (continuous) functions that generate the dynamics on the 2-simplex. we consider three types of qualitative restrictions: inequality (or set-theoretical) conditions, monotonicity/curvature (or differential-geometrical) conditions, and topological conditions (referring to (transversal) non-(self-)intersection of trajectories). we discuss the implications of these restrictions for transitional and limit dynamics on the 2-simplex and the wide range of potential and existing applications of the resulting system-theoretical models in economics and, in particular, in economic growth and development theory. keywords: dynamics, trajectory, 2-simplex, continuous, monotonous, intersection, self-intersection, poincaré-bendixson theory, economics 1. introduction in this paper, we discuss the models of continuous dynamics on the 2-simplex that arise when different qualitative restrictions are imposed on the (continuous) vector function x(t) ≡ (x1(t), x2(t), x3(t)) that generates the dynamics on the 2-simplex (where t represents time). in particular, there are three major types of qualitative conditions that can be imposed on this function: (1.) inequality conditions of the type ∀t ∈ a ∀i ∈ b xi(t) ≶ ai = const., which can be treated by using set-theoretical concepts (referring to the points or segments of the corresponding trajectory and the partitions of the 2-simplex); (2.) (strict) monotonicity conditions referring to all or some of the functions xi(t), which can be treated by (differential) geometrical concepts of tangential vector angles and curvature; and * corresponding author: denis@stijepic.com mailto:denis@stijepic.com models of continuous dynamics on the 2-simplex and applications in economics 117 copyright ©2019 assa. adv. in systems science and appl. (2019) (3.) conditions regarding (transversal) trajectory non-(self-)intersections, which can be treated by using topological concepts (e.g., homeomorphisms). we discuss the implications of these restrictions for the transitional and limit dynamics on the 2-simplex (among others, fixed points, waves, or (limit) cycles may arise). the models that result from this discussion are relatively simple from the mathematical point of view, yet they seem widely applicable in economic growth and development theory and, thus, may be regarded as powerful system-theoretical constructs. as discussed in section 6, many core topics of long-run economic theory (e.g., savings rate dynamics, functional income distribution dynamics, and sector dynamics) can be modeled by continuous trajectories on simplexes, to which our system-theoretical results apply. in general, the system-theoretical analysis of economic dynamics seems a highly valuable complement to the standard approaches of economic modeling, which rely on quantitative theoretical micro-foundations: system-theoretical models can be relatively crude and do not necessarily require detailed/quantitative assumptions about the nature of economic phenomena. this can be an advantage, since, often, detailed/quantitative assumptions about economic agents and economic environments cannot be supported by empirical evidence and, thus, may open the door for speculation, ideology, and prediction errors. for these reasons, system-theoretical results are not only transferable across many economic topics, but can also be relatively robust in comparison to the results of quantitative micro-founded economic models, and this fact is particularly useful in prediction of long-run economic dynamics (see [20] and section 6 for a detailed discussion). moreover, the dynamics on the 2-simplex merit a detailed consideration from the (application oriented) mathematical point of view: although the 2-simplex can be regarded as a bounded subset of a plane (in ℝ3), the description of the dynamics on the 2-simplex requires a greater variety of analytical concepts in comparison to the description of the dynamics in ℝ2 (see, e.g., the discussion of the monotonicity concepts in sections 2.3 and 3). the rest of the paper is organized as follows. in section 2, we discuss the characterization of trajectory families on the 2-simplex via geometrical and topological concepts. sections 3-5 discuss the implications of these concepts for transitional and limit dynamics on the 2-simplex. this discussion yields system-theoretical models. the potential and existing applications of these models in economics and, in particular, in growth and development theory are discussed in section 6. concluding remarks are provided in section 7. 2. characterization of the trajectories on the 2-simplex in section 2, we summarize the concepts that can be used to characterize continuous dynamics on the 2-simplex as applied by [13-15] in structural change modeling. while there are different mathematical notational conventions, we choose the following notation for reasons of simplicity: small letters (e.g., x), bold small letters (e.g., x), capital letters (e.g., x), and greek letters (e.g., α) denote scalars, vectors/points, sets, and vector angles, respectively. cl(a) denotes the closure of the set a. if i denotes an open interval (e.g., (a, b)), then [i], [i), and (i] denote the corresponding closed (e.g., [a, b]), left-closed (e.g., [a, b)), and right-closed (e.g., (a, b]) interval, respectively. 118 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) 2.1 trajectories on the 2-simplex the (standard) 2-simplex (s), which is defined by (1), is a triangle in ℝ3, as depicted by fig. 1. the cartesian coordinates of the simplex vertices v1, v2, and v3 are stated by (2). s := {(x1, x2, x3) ∈ ℝ3: x1 + x2 + x3 = 1 ∧ ∀i ∈ {1, 2, 3} 0 ≤ xi ≤ 1} (1) v1 := (1, 0, 0) (2a) v2 := (0, 1, 0) (2b) v3 := (0, 0, 1) (2c) we define the vector function x(t, j) as follows: x(t, j) ≡ (x1(t, j), x2(t, j), x3(t, j)): t × j → s (3a) 0 ∈ t ⊆ ℝ (3b) j ⊆ s (3c) the trajectory x(t, j) and the trajectory segment x(t.+, j) are defined by (4). ∀j ∈ j x(t, j) := {x(t, j) ∈ s: t ∈ t} (4a) ∀j ∈ j x(t +, j) := {t ∈ t: t ≥ 0} (4b) x3 v3 v2 v1 x2x1 fig. 1. the standard simplex s in ℝ3. in fact, (4a) defines a trajectory family indexed by the set j, where each trajectory x(t, j) describes a path on s that is traversed over the period t. x(t +, j) is the segment of this path that is traversed over t ≥ 0. 2.2 set-theoretical trajectory classification equations (5) introduce a partitioning of s, which can be used for describing the location of relevant trajectory points or segments (e.g., initial segment/state, empirically observed segment, or segment representing the future dynamics), as we will see later. ∀i ∈ {1, 2, 3} svi := {(x1, x2, x3) ∈ s: xi > 1/2} (5a) sv0 := s \ (sv1 ∪ sv2 ∪ sv3) (5b) models of continuous dynamics on the 2-simplex and applications in economics 119 copyright ©2019 assa. adv. in systems science and appl. (2019) (5a) and (1) imply that the partition svi contains all the points of s that are dominated by xi; i.e., if a point (x1, x2, x3) is located in partition svi, then ∀j ∈ {1, 2, 3}\i xi > xj. the geometrical interpretation of the partitioning (5) is depicted in fig. 2. as we can see, for i ∈ {1, 2, 3}, the partition svi contains all the points of s that are closer to the vertex vi than to the other vertices (vj, j ≠ i). the following (set-theoretical) definitions allow us to assess the prediction range of monotonous models, as we will see later. let a(k) denote the area function, which assigns to a set k ⊆ s the (real number indicating the) area of k. b(t, f) := ⋃j∈f x(t, j) is the image of the family f of trajectories x(t, j), j ∈ f ⊆ j (cf. (4)). among all the path-connected and closed subsets of s that cover b(t, f), let m(t, f) denote one of the sets that cover the smallest area of s. a*(t, f) := a(m(t, f)) is the family image size of the family f. 23113 12332 31221 vwvv vwvv vwvv    sv1 sv2 sv3 sv0 v3 v2v1 w31 w23 w12 fig. 2. the partitioning of s. 2.3 differential-geometrical trajectory classification while the previous discussion can be used for a set-theoretical characterization of trajectories, we focus now on a differential-geometrical characterization of trajectories referring to the angles of the tangential vectors and expressing the monotonicity characteristics and the curvature of a trajectory. we say that the trajectory x(t, j) is continuous if for the given j, x(t, j) is continuous in t on the time interval t (cf. (4a)). moreover, a trajectory family is continuous if all the trajectories belonging to this family are continuous. let (a) d(t, j) be the directional (or tangential) vector associated with the point x(t, j), (b) ℓ be a line through the point x(t, j) that is parallel to the simplex edge v1-v2, and (c) δ(t, j) := ∡(d(t, j), ℓ) ∈ [0°, 360°] be the angle between the directional vector d(t, j) and the line ℓ (cf. fig. 3). 120 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) x(t, j) ℓ|| x(t, j) d(t, j) δ(t, j) v2v1 v3 fig. 3. the vector angle δ(t, j). moreover, we define the angles αi and the angle intervals ii by relying on the ‘saw tooth’ function ϕ: ℝ → ℝ as follows (cf. fig. 4): ϕ(z) = [(z – 1)/6 – floor((z – 1)/6)]360° (6a) ∀i ∈ ℕ ii ≡ (αi, αi+1) := (ϕ(i), ϕ(i + 1)) (6b) ∀i ∈ ℕ ∀j ∈ {n ∈ ℕ: i < n} [ii~j] := ⋃ j k = i [ik] ∧ [ii~j) := ⋃ j k = i [ik] \ αj+1 ∧ (ii~j]:= ⋃ j k = i [ik] \ αi ∧ ii~j := ⋃ j k = i [ik] \ αi \ αj+1 (6c) by using our definition of the vector angle δ(t, j) and the vector angles and intervals (6), we can formulate the properties 1-3 reflecting the relation between the tangential vector angles and the dynamics of x(t, j) in the case that ẋ(t, j) ≠ 0 (cf. fig. 1, 3, and 4). property 1: if ẋ(t, j) ≠ 0, then (a) δ(t, j) ∈ i3~5 ⟺ ẋ1(t, j) > 0, (b) δ(t, j) ∈ i0~2 ⟺ ẋ1(t, j) < 0, and (c) δ(t, j) ∈ {α3, α6} ⟺ ẋ1(t, j) = 0. property 2: if ẋ(t, j) ≠ 0, then (a) δ(t, j) ∈ i5~7 ⟺ ẋ2(t, j) > 0, (b) δ(t, j) ∈ i2~4 ⟺ ẋ2(t, j) < 0, and (c) δ(t, j) ∈ {α2, α5} ⟺ ẋ2(t, j) = 0. property 3: if ẋ(t, j) ≠ 0, then (a) δ(t, j) ∈ i1~3 ⟺ ẋ3(t, j) > 0 (b) δ(t, j) ∈ i4~6 ⟺ ẋ3(t, j) < 0, and (c) δ(t, j) ∈ {α1, α4} ⟺ ẋ3(t, j) = 0. we rely on the following definitions of monotonicity. first, xi(t, j) is monotonous (in t) if (∀t ∈ t ẋi(t, j) ≥ 0) or (∀t ∈ t ẋi(t, j) ≤ 0). second, xi(t, j) is strictly monotonous (in t) if either (∀t ∈ t ẋi(t, j) > 0) or (∀t ∈ t ẋi(t, j) < 0) but not both. third, the trajectory x(t, j) (associated with the function x(t, j)) is (strictly) monotonous in one dimension if (a) there exists an i ∈ {1, 2, 3} such that xi(t, j) is (strictly) monotonous and (b) for all k ∈ {1, 2, 3}\i xk(t, j) is not (strictly) monotonous. fourth, the trajectory x(t, j) is (strictly) monotonous in two models of continuous dynamics on the 2-simplex and applications in economics 121 copyright ©2019 assa. adv. in systems science and appl. (2019) dimensions if (a) there exist an i ∈ {1, 2, 3} and a k ∈ {1, 2, 3}\i such that xi(t, j) and xk(t, j) are (strictly) monotonous and (b) xl(t, j) is not (strictly) monotonous with l ∈ {1, 2, 3}\{i, k}. fifth, the trajectory x(t, j) is (strictly) monotonous (in three dimensions) if ∀i ∈ {1, 2, 3} xi(t, j) is (strictly) monotonous. ϕ(z)/100 z α1 = 0 α2 = 60 = α8 α3 = 120 α4 = 180 α5 = 240 α6 = 300 = α0 i6 = (300 , 360 ) = i0 i1 = (0 , 60 ) = i7 i2 = (60 , 120 ) i3 = (120 , 180 ) i4 = (180 , 240 ) i5 = (240 , 360 ) . s5(x0) s6(x0) = s0(x0) s4(x0) s3(x0) s2(x0) s1(x0) = s7(x0) v1 v2 v3 x0 ≡ (x01, x02, x03) fe d c b l1(x0) = l2(x0) = l3(x0) = l4(x0) = l5(x0) = l6(x0) = a fig. 4. the function ϕ(z), the angle intervals ii, the line segments li(x0), and the sets si(x0). 122 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) instead of using the curvature definition that is widespread in differential geometry (and which is difficult to apply in the proofs of our theorems), we use the following definition of curvature relying on vector angles: let β(q, r, j) denote the angle between the two tangential vectors d(q, j) and d(r, j) associated with the (monotonous) trajectory x(t, j) on s (cf. (4)), where q, r ∈ t. among all the tangential vector pairs (d(t, j), d(s, j)) associated with the (monotonous) trajectory x(t, j), where t, s ∈ t, let d(t*, j) and d(s*, j) be among the ones that are characterized by the largest angle β, i.e., β(t*, s*, j) =: κ(t, j) is the maximal tangential vector angle β associated with the trajectory x(t, j). the greater κ(t, j) is, the greater the curvature of the trajectory x(t, j) is. obviously, a linear trajectory has a curvature of 0. we say that a strictly monotonous trajectory x(t, j) (or a trajectory segment) that describes a clockwise (counterclockwise) movement on s has a positive (negative) signed curvature and write κ(t, j) > 0 (κ(t, j) < 0). let f be a family of trajectories x(t, j), j ∈ f ⊆ j (cf. (4)). then, κ*(t, f) := max(cl({κ(t, j): j ∈ f})) is the maximum curvature of the family f on the time interval t. 2.4 topological trajectory classification here, the topological characterization refers to the question whether a trajectory family is non-(self-)intersecting, which is a characteristic that can be expressed by homeomorphisms. moreover, it is deciding for the applicability of the poincaré-bendixson theory (cf. section 5) that the 2-simplex is homeomorphic to a bounded (and closed) subset of a plane. two trajectories x(t, j) and x(u, k) are non-intersecting if x(t, j) ∩ x(u, k) = ∅, where u ⊆ ℝ. otherwise they are intersecting. a trajectory x(t, j) is self-intersecting if ∃(r, s, t) ∈ t 3 r < s < t ∧ x(r, j) = x(t, j) ≠ x(s, j). otherwise, the trajectory is non-self-intersecting. a trajectory x(t, j) is transversally self-intersecting if ∃(t, s) ∈ t 2 t ≠ s ∧ x(t, j) = x(s, j) ∧ δ(t, j) ≠ δ(s, j). otherwise, the trajectory is not transversally self-intersecting. according to these definitions, a closed trajectory corresponding to a jordan curve is self-intersecting but not transversally self-intersecting. 3. implications of monotonicity in contrast to monotonous and bounded trajectories in ℝ2, monotonous trajectories on the 2-simplex can have (a) a wide range of different shapes and (b) omega limit sets consisting of more than only one (fixed) point. in this section, we discuss the geometrical aspects of the transitional and limit dynamics associated with continuous trajectories that are monotonous in one, two, or three dimensions. as we will see, these geometrical properties have interesting applications in economic dynamics modeling. 3.1 general properties of monotonous trajectories on the 2-simplex in this section, we show that continuous trajectories that are monotonous in three dimensions (two dimensions) are characterized by relatively low curvatures, allow for relatively weak waves, and are placed in relatively small subsets of the 2-simplex in comparison to the ‘related’ trajectories that are monotonous in two dimensions (one dimension). propositions 1-3 and corollary 1 summarize these results formally. the models of continuous dynamics on the 2-simplex and applications in economics 123 copyright ©2019 assa. adv. in systems science and appl. (2019) readers who are less interested in this formal discussion can also go directly to the discussion of figure 4 (see the paragraphs below proposition 3), which elaborates on the intuitive/graphical interpretation of these geometrical properties. given a point x0 ≡ (x01, x02, x03) ∈ s, (7) defines different subsets of s. as we will see (in proposition 3), each of the subsets s1-s6 defined by (7) corresponds to the closure of one of the vector angle intervals i1-i6 defined by (6), and each of the line segments l1-l6 defined by (7) corresponds to one of the angles α1-α6 defined by (6) (cf. fig. 4). l1(x0) := {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x3 = x03} (7a) s1(x0) := {(x1, x2, x3) ∈ s: x2 ≥ x02 ∧ x3 ≥ x03} (7b) l2(x0) := {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 = x02} (7c) s2(x0) := {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 ≤ x02} (7d) l3(x0) := {(x1, x2, x3) ∈ s: x1 = x01 ∧ x2 ≤ x02} (7e) s3(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x3 ≥ x03} (7f) l4(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x3 = x03} (7g) s4(x0) := {(x1, x2, x3) ∈ s: x2 ≤ x02 ∧ x3 ≤ x03} (7h) l5(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x2 = x02} (7i) s5(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x2 ≥ x02} (7j) l6(x0) := {(x1, x2, x3) ∈ s: x1 = x01 ∧ x2 ≥ x02} (7k) s6(x0) := {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x3 ≤ x03} (7l) s0(x0) := s6(x0) s7(x0) := s1(x0) (7m) ∀i ∈ ℕ ∀j ∈ {n ∈ ℕ: i < n} si~j(x0) := ⋃j k = i sk(x0) (7n) in propositions 1-3, we derive the characteristics of monotonous trajectories and, in particular, the ranges of tangential vector angles δ(t, j) and the sets in which the trajectory segments x(t +, j) are located. later, we will discuss the geometrical/graphical interpretation of these characteristics and their applications in economic dynamics analysis. proposition 1 focuses on the trajectories that are continuous and monotonous in one dimension on s and shows that there are only nine trajectory types belonging to this monotonicity class, where each type is identified by one of the condition-sets p11-p19. proposition 1: assume that (a) the trajectory x(t, j) defined by (4a) is continuous and monotonous in one dimension on s, (b) x(0, j) = x0 ≡ (x01, x02, x03) ∈ s, and (c) ∀t ∈ t ẋ(t, j) ≠ 0. then, x(t, j) and x(t +, j) (cf. (4b)) satisfy one and only one of the condition sets p11-p19, which are defined as follows (cf. (6) and (7)): a) for i ∈{1, 2, …, 6}, condition set p1i is: (∀t ∈ t δ(t, j) ∈ [i(i–1)~(i+1)]) ∧ (∃t ∈ t δ(t, j) ∈ i(i–1)~(i+1)) ∧ (∃(r, s) ∈ t 2 δ(r, j) ∈ [ii–1) ∧ δ(s, j) ∈ (ii+1]) ∧ x(t +, j) ⊂ s(i–1)~(i+1)(x0); b) for i ∈ {7, 8, 9}, condition set p1i is: (∀t ∈ t δ(t, j) ∈ {αi–6, αi–3}) ∧ (∃(p, q) ∈ t 2 δ(p, j) ∈ {αi–6} ∧ δ(q, j) ∈ {αi–6}) ∧ x(t +, j) ⊆ li–6(x0) ∪ li–3(x0). proof. as defined in section 2, x(t, j) is monotonous in one dimension if for one and only one i ∈ {1, 2, 3}, the function xi(t, j) is monotonous while for all other i, xi(t, j) is non-monotonous. thus, for proving proposition 1, we have to consider only three alternative scenarios of monotonicity in one dimension: (a) x1(t, j) is monotonous, (b) x2(t, j) is 124 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) monotonous, and (c) x3(t, j) is monotonous. moreover, since a monotonous function can be monotonously increasing or monotonously decreasing (or both), we have three alternative sub-scenarios for each of the three scenarios (a)-(c): (a) monotonously increasing, (b) monotonously decreasing, and (c) both, monotonously increasing and monotonously decreasing (which means constant). thus, overall, we have nine sub-scenarios: (aa)-(ac), (ba)-(bc), and (ca)-(cc). according to properties 1-3, each of the condition sets p11-p19, to which proposition 1 refers, represents one of the nine sub-scenarios (aa)-(cc). for example, condition sets p12, p15, and p17 represent the sub-scenarios (ca), (cb), and (cc), respectively. consider first the sub-scenario (ca), i.e., assume that x3(t, j) increases monotonously. property 3 and (6b) imply that (8) is valid in sub-scenario (ca). (∀t ∈ t δ(t, j) ∈ [i1~3]) ∧ (∃t ∈ t δ(t, j) ∈ i1~3) (8) moreover, according to the definition of ‘monotonicity in one dimension’, to which proposition 1 refers, (9) is valid in sub-scenario (ca). x1(t, j) and x2(t, j) are non-monotonous. (9) the interval [i1~3], to which (8) refers, can be partitioned into three subintervals [i1), [i2], and (i3]. properties 1 and 2 and (6b) imply (10). ∀t ∈ t δ(t, j) ∈ (i3] ⇒ x1(t, j) is monotonous. (10a) ∀t ∈ t δ(t, j) ∈ [i1) ⇒ x2(t, j) is monotonous. (10b) ∀t ∈ t δ(t, j) ∈ [i2] ⇒ x1(t, j) and x2(t, j) are monotonous. (10c) the statements (10) imply statement (11). x1(t, j) or x2(t, j) is monotonous if for all t ∈ t, δ(t, j) is within one and only one of the sub-intervals [i1), [i2], and (i3]. (11) (9) and (11) imply that over the period t, the tangential vectors δ(t, j) cannot stay within one and the same subinterval, i.e., at least one subinterval switch must occur over the period t. given the three subintervals [i1), [i2], and (i3], the set of all possible subinterval switches is: (i) switch from [i1) to [i2], (ii) switch from [i1) to (i3], (iii) switch from [i2] to [i1), (iv) switch from [i2] to (i3], (v) switch from (i3] to [i1), and (vi) switch from (i3] to [i2]. we analyze now these interval switches. in case (i), i.e., if (1.) initially, the tangential vector angles are within the interval [i1) and (2.) at some later time point, the tangential vector angles switch to the interval [i2], x1(t, j) is monotonous (cf. property 1). this contradicts (9). analogously, it can be shown that cases (iii), (iv), and (vi) contradict (9), since: in case (iii), x1(t, j) is monotonous; in case (iv), x2(t, j) is monotonous; in case (vi), x2(t, j) is monotonous. only, in cases (ii) and (v), both, x2(t, j) and x1(t, j), are non-monotonous, which is consistent with (9). in each of the cases (ii) and (v), (12) is true. ∃(r, s) ∈ t 2 δ(r, j) ∈ [i1) ∧ δ(s, j) ∈ (i3] (12) the fact that x3(t, j) increases monotonously in sub-scenario (ca) implies that ∀t ≥ 0 x3(t, j) ≥ x3(0, j), where x3(0, j) = x03 according to the assumptions made in proposition 1. in other words, in sub-scenario (ca), x(t +, j) ⊂ {(x1, x2, x3) ∈ s: x3 ≥ x03} =: sca(x0) (cf. proposition 1). if x(t +, j) ⊂ sca(x0) ⇒ x(t +, j) ⊂ s1~3(x0), then (13) is valid in sub-scenario (ca). x(t +, j) ⊂ s1~3(x0) (13) models of continuous dynamics on the 2-simplex and applications in economics 125 copyright ©2019 assa. adv. in systems science and appl. (2019) we prove now that x(t +, j) ⊂ sca(x0) ⇒ x(t +, j) ⊂ s1~3(x0). given the point x0 ≡ (x01, x02, x03) ∈ s (cf. proposition 1), the definition of sca(x0) (and (1)) implies that (14)-(16) are true if x(t, j) ∈ sca(x0). either x3(t, j) > x03 or x3(t, j) = x03 but not both. (14) x2(t, j) < x02 or x2(t, j) > x02 (or x2(t, j) = x02 ). (15) x1(t, j) < x01 or x1(t, j) > x01 (or x1(t, j) = x01). (16) the statement (16) can be divided into the two (disjunctive) cases (17a) and (17b). either x1(t, j) > x01 or x1(t, j) = x01 but not both. (17a) x1(t, j) < x01 (17b) if (14), (15), and (17a) are true and x(t, j) ∈ s, then x(t, j) ∈ {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x3 ≥ x03} and, thus, x(t, j) ∈ s3(x0) (cf. (7f)). we consider now the cases in which (14), (15), and (17b) are true. these cases are: x1(t, j) < x01 ∧ x3(t, j) > x03 ∧ x2(t, j) > x02 (18a) x1(t, j) < x01 ∧ x3(t, j) > x03 ∧ x2(t, j) < x02 (18b) x1(t, j) < x01 ∧ x3(t, j) > x03 ∧ x2(t, j) = x02 (18c) x1(t, j) < x01 ∧ x3(t, j) = x03 ∧ x2(t, j) > x02 (18d) x1(t, j) < x01 ∧ x3(t, j) = x03 ∧ x2(t, j) < x02 (18e) x1(t, j) < x01 ∧ x3(t, j) = x03 ∧ x2(t, j) = x02 (18f) obviously, the cases (18e) and (18f) violate (1). thus, if (18e) or (18f) is true, then x(t, j) ∉ s. if (18b) or (18c) is true and x(t, j) ∈ s, then x(t, j) ∈ {(x1, x2, x3) ∈ s: x1 < x01 ∧ x2 ≤ x02 ∧ x3 > x03} =: sbc(x0). if x(t, j) ∈ s2(x0), then x3(t, j) ≥ x03, since, otherwise, (1) is violated (cf. (7d)). in other words, s2(x0) = {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 ≤ x02 ∧ x3 ≥ x03}. obviously, sbc(x0) ⊂ s2(x0). thus, if (18b) or (18c) is true and x(t, j) ∈ s, then x(t, j) ∈ s2(x0). analogously, if (18a), (18c), or (18d) is true and x(t, j) ∈ s, then x(t, j) ∈ {(x1, x2, x3) ∈ s: x1 < x01 ∧ x2 ≥ x02 ∧ x3 ≥ x03} =: sacd(x0). moreover, if x(t, j) ∈ s1(x0), then x1(t, j) ≤ x01, since, otherwise, (1) is violated (cf. (7b)). in other words, s1(x0) = {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 ≥ x02 ∧ x3 ≥ x03}. obviously, sacd(x0) ⊂ s1(x0). thus, if (18a), (18c), or (18d) is true and x(t, j) ∈ s, then x(t, j) ∈ s1(x0). overall, we have shown that if x(t, j) ∈ sca(x0) ⊆ s, then the statements (14)-(17) are valid, which imply several feasible cases. in each of these cases, x(t, j) is in one of the sets s1(x0), s2(x0), and s3(x0), i.e., x(t, j) ∈ sca(x0) ⇒ x(t, j) ∈ s1(x0) ∪ s2(x0) ∪ s3(x0). this implies that x(t +, j) ⊂ sca(x0) ⇒ x(t +, j) ⊂ s1(x0) ∪ s2(x0) ∪ s3(x0), since x(t +, j) is the union of the points x(t, j) ∈ s for which the statements (14)-(17) (and (1)) hold (cf. proposition 1). according to (7n), s1~3(x0) = s1(x0) ∪ s2(x0) ∪ s3(x0). this completes the proof that x(t +, j) ⊂ sca(x0) ⇒ x(t +, j) ⊂ s1~3(x0). by now, we have shown that in the sub-scenario (ca), the statements (8), (12), and (13) must be true. these three statements reduce to condition set p12. it can be shown in the same way that (1.) the sub-scenarios (cb) and (cc) correspond to condition sets p15 and p17, respectively, and (2.) each of the sub-scenarios (ba)-(cc) corresponds to one and only one of the condition sets p11, p13, p14, p16, p18, and p19. 126 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) this completes the proof that each of the alternative (sub-)scenarios of monotonicity in one dimension (i.e., each of the sub-scenarios (aa)-(cc)) corresponds to one and only one of the condition sets p11-p19. □ now, we turn to proposition 2, which is similar to proposition 1 except for the fact that it focuses on the trajectories that are continuous and monotonous in two dimensions on s. in particular, proposition 2 derives the corresponding ranges of tangential vector angles δ(t, j) and the sets in which the trajectory segments x(t +, j) are located and shows that there are only six types of trajectories that are continuous and monotonous in two dimensions on s, where each type is identified by one of the condition-sets p21-p26. proposition 2: assume that (a) the trajectory x(t, j) defined by (4a) is continuous and monotonous in two dimensions on s, (b) x(0, j) = x0 ≡ (x01, x02, x03) ∈ s, and (c) ∀t ∈ t ẋ(t, j) ≠ 0. then, x(t, j) and x(t +, j) (cf. (4b)) satisfy one and only one of the condition sets p21-p26, where for i ∈ {1, 2, …, 6}, condition set p2i is: (∀t ∈ t δ(t, j) ∈ [ii~(i+1)]) ∧ (∃(r, s) ∈ t 2 δ(r, j) ∈ [ii) ∧ δ(s, j) ∈ (i(i+1)]) ∧ x(t +, j) ⊂ si~(i+1)(x0) (cf. (6)/(7)). proof. according to the definition of monotonicity in two dimensions, two of the functions x1(t, j), x2(t, j), and x3(t, j) must be monotonous, while the remaining one must be non-monotonous. thus, we have to consider only three cases: (a) x1(t, j) and x2(t, j) are monotonous (while x3(t, j) is non-monotonous), (b) x1(t, j) and x3(t, j) are monotonous (while x2(t, j) is non-monotonous), and (c) x2(t, j) and x3(t, j) are monotonous (while x1(t, j) is non-monotonous). for each of these cases, we must distinguish between four subcases. for example, in case (a), we can distinguish between the following subcases: (a) x1(t, j) and x2(t, j) are monotonously increasing, (b) x1(t, j) is monotonously increasing, while x2(t, j) is monotonously decreasing, (c) x1(t, j) and x2(t, j) are monotonously decreasing, and (d) x1(t, j) is monotonously decreasing, while x2(t, j) is monotonously increasing. subcases (a) and (c) are infeasible, since they violate (1): for example, if x1(t, j) and x2(t, j) are monotonously increasing, then x3(t, j) must be monotonously decreasing (instead of being non-monotonous), since x1(t, j) + x2(t, j) + x3(t, j) must be equal to 1 for all t. thus, we must consider only the subcases (b) and (d) of case (a). properties 1 and 2 imply that in subcase (b) of case (a), the statement (19) is valid (cf. (6)). ∀t ∈ t δ(t, j) ∈ [i3~4] (19) moreover, since case (a) requires that x3(t, j) is non-monotonous, property 3 implies that (20) is valid in case (a). ∃(r, s) ∈ t 2 ẋ3(r, j) > 0 ∧ ẋ3(s, j) < 0 (20) according to (6), the interval [i3~4], to which (19) refers, can be partitioned into the following partitions: [i3), α4, and (i4]. property 3 and (6) imply (21). ∀t ∈ t δ(t, j) ∈ [i3) ∨ δ(t, j) ∈ [i3) ∪ α4 ⇒ ∀t ∈ t ẋ3(t, j) ≥ 0 (21a) ∀t ∈ t δ(t, j) ∈ (i4] ∨ δ(t, j) ∈ α4 ∪ (i4] ⇒ ∀t ∈ t ẋ3(t, j) ≤ 0 (21b) ∀t ∈ t δ(t, j) ∈ α4 ⇒ ∀t ∈ t ẋ3(t, j) = 0 (21c) models of continuous dynamics on the 2-simplex and applications in economics 127 copyright ©2019 assa. adv. in systems science and appl. (2019) the statements (21) imply that if (19) and (20) are true, the tangential vectors δ(t, j) cannot stay within one and only one of the subintervals [i3), [i3) ∪ α4, α4, α4 ∪ (i4], and (i4] for all t ∈ t, and there must occur a switch from subinterval [i3) to subinterval (i4] or from subinterval (i4] to subinterval [i3) over the period t. thus, (22) is valid. ∃(r, s) ∈ t 2 δ(r, j) ∈ [i3) ∧ δ(s, j) ∈ (i4] (22) since in subcase (b) of case (a), x1(t, j) increases monotonously and x2(t, j) decreases monotonously, the assumptions made in proposition 2 and (4b) imply that x(t +, j) ⊂ {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x2 ≤ x02} =: sab(x0). sab(x0) can be partitioned as follows: sab(x0) = sab1(x0) ∪ sab2(x0), where sab1(x0) ∩ sab2(x0) = ∅ and sab1(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x2 ≤ x02 ∧ x3 ≥ x03} and sab2(x0) := {(x1, x2, x3) ∈ s: x1 ≥ x01 ∧ x2 ≤ x02 ∧ x3 < x03}. we can see immediately that sab1(x0) ⊂ s3(x0) (cf. (7f)) and sab2(x0) ⊂ s4(x0) (cf. (7h)). thus, sab(x0) ⊂ s3(x0) ∪ s4(x0). this result, (7n), and the previously shown fact that x(t +, j) ⊂ sab(x0) imply (23). x(t +, j) ⊂ s3~4(x0) (23) overall, we have shown that in the subcase (b) of case (a), condition set p23 must be true (cf. (19), (22), and (23)). analogously, it can be shown that in all the feasible subcases of cases (a)-(c), one and only one of the statements p21, p22, p24, p25, and p26 is true, which proves proposition 2. □ finally, we focus on the trajectories that are continuous and monotonous in three dimensions on s. in particular, proposition 3 derives the corresponding ranges of tangential vector angles δ(t, j) and the sets in which the trajectory segments x(t +, j) are located and shows that there are twelve types of trajectories that are continuous and monotonous in three dimensions on s, where each type is identified by one of the condition-sets p31-p312. proposition 3: assume that (a) the trajectory x(t, j) defined by (4a) is continuous and monotonous (in three dimensions) on s, (b) x(0, j) = x0 ≡ (x01, x02, x03) ∈ s, and (c) ∀t ∈ t ẋ(t, j) ≠ 0. then, x(t, j) and x(t +, j) (cf. (4b)) satisfy one and only one of the condition sets p31-p312, where (cf. (6) and (7)): a) for i ∈ {1, 2, …, 6}, condition set p3i is: (∀t ∈ t δ(t, j) ∈ [ii]) ∧ (∃s ∈ t δ(s, j) ∈ ii) ∧ x(t +, j) ⊂ si(x0); b) for i ∈ {7, 8, …, 12}, condition set p3i is: ∀t ∈ t δ(t, j) ∈ {αi–6} ∧ x(t +, j) ⊆ li–6(x0). proof. according to our definition of monotonicity (in three dimensions), x1(t, j), x2(t, j), and x3(t, j) must be monotonous if x(t, j) is monotonous (in three dimensions) on s. since a monotonous function can be (a) monotonously increasing, (b) monotonously decreasing, or (c) both (monotonously increasing and monotonously decreasing and, thus, constant), we have per function xi(t, j) three cases ((a)-(c)). moreover, we have three functions x1(t, j), x2(t, j), and x3(t, j). thus, overall, there are 33 possible combinations. this set of 27 combinations contains the combination (a) ∀i ẋi(t, j) ≤ 0, the combination (b) ∀i ẋi(t, j) ≥ 0, three times the combination (c) ẋi(t, j) ≥ 0 ∧ ẋk(t, j) ≥ 0 ∧ ẋl(t, j) = 0 ∧ i ≠ k ≠ l, three times the combination (d) ẋi(t, j) ≤ 0 ∧ ẋk(t, j) ≤ 0 ∧ ẋl(t, j) = 0 ∧ i ≠ k ≠ l, six times the combination (e) ẋi(t, j) = ẋk(t, 128 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) j) = 0 ∧ ẋl(t, j) ≠ 0 ∧ i ≠ k ≠ l, and the combination (f) ∀i ẋi(t, j) = 0. the combinations (a)-(e) are infeasible, since they violate (1) unless they reduce to combination (f). the combination (f) represents a fixed point (ẋ(t, j) = 0) and is excluded by the assumptions made in proposition 3. in the rest of the proof, we have to consider the remaining 12 combinations.1 each of these 12 combinations is covered by one of the conditions sets p31-p312. we leave it to the reader to prove the validity of proposition 3 in all these 12 cases; we prove the validity in only two representative cases. consider the case ∀t ∈ t ẋ1(t, j) ≤ 0 ∧ ẋ2(t, j) ≤ 0 ∧ ẋ3(t, j) ≥ 0, where ∃(r, s, p) ∈ t 3 ẋ1(r, j) < 0 ∧ ẋ2(s, j) < 0 ∧ ẋ3(p, j) > 0. then, (a) properties 1-3 imply almost directly that the tangential vector angles δ(t, j) satisfy the condition set p32, and (b) the assumptions made in proposition 3 and (4b) imply that x(t +, j) ⊂ {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 ≤ x02 ∧ x3 ≥ x03} =: sz(x0), and, thus, (7d) implies that sz(x0) ⊂ s2(x0); thus, x(t +, j) ⊂ s2(x0) as stated by the condition set p32. alternatively, consider the case ∀t ∈ t ẋ1(t, j) ≤ 0 ∧ ẋ2(t, j) ≥ 0 ∧ ẋ3(t, j) = 0, where ∃r ∈ t ẋ1(r, j) < 0 ∧ ẋ2(r, j) > 0.2 properties 1-3 imply almost directly that in this case, the tangential vector angles δ(t, j) satisfy the condition set p37. moreover, the assumptions made in proposition 3, (4b), and (7a) imply that x(t +, j) ⊆ {(x1, x2, x3) ∈ s: x1 ≤ x01 ∧ x2 ≥ x02 ∧ x3 = x03} ⊆ l1(x0). thus, x(t +, j) ⊆ l1(x0) as stated by condition set p37. □ we discuss now the geometrical interpretation of propositions 1, 2, and, 3, as depicted by fig. 4. to construct fig. 4, we choose an arbitrary point (x0) in the interior of s. then, we draw three line segments going through x0 and each being parallel to one of the simplex edges v1-v2, v2-v3, and v3-v1. the intersections of the line segments with the simplex edges are denoted by the points a-f. we can see that the line segments that connect x0 with one of the points a-f are the line segments l1(x0)-l6(x0), which are defined by (7) and which localize the six (closed) subsets s1(x0)-s6(x0) defined by (7). the angles between the line segments l1(x0)-l6(x0) and the simplex edge v1-v2 (according to the definition of tangential vector angles and intervals (6)) are depicted in the middle panel of fig. 4. (7n) and fig. 4 imply almost directly that (1.) each of the six sets si~(i+1)(x0), to which proposition 2 refers, is simply the union of two neighboring sets sj(x0) and sk(x0), (2.) each of the six sets s(i–1)~(i+1)(x0), to which proposition 1 refers, is simply the union of three neighboring sets sj(x0), sk(x0), and sm(x0). in particular, propositions 1-3 can be interpreted easily by using fig. 4: 1.) proposition 3 implies three geometrical properties of a trajectory segment x(t +, j) that is monotonous in three dimensions. first, x(t +, j) is located in one of the line segments l1(x0)-l6(x0) or in one of the sets s1(x0)-s6(x0). second, if x(t +, j) is in li(x0), then for all t ≥ 1 these feasible combinations are: (1.) ẋ1 ≤ 0 ∧ ẋ2 ≤ 0 ∧ ẋ3 ≥ 0, (2.) ẋ1 ≤ 0 ∧ ẋ2 ≥ 0 ∧ ẋ3 ≤ 0, (3.) ẋ1 ≤ 0 ∧ ẋ2 ≥ 0 ∧ ẋ3 ≥ 0, (4.) ẋ1 ≤ 0 ∧ ẋ2 ≥ 0 ∧ ẋ3 = 0, (5.) ẋ1 ≤ 0 ∧ ẋ2 = 0 ∧ ẋ3 ≥ 0, (6.) ẋ1 ≥ 0 ∧ ẋ2 ≤ 0 ∧ ẋ3 ≤ 0, (7.) ẋ1 ≥ 0 ∧ ẋ2 ≤ 0 ∧ ẋ3 ≥ 0, (8.) ẋ1 ≥ 0 ∧ ẋ2 ≤ 0 ∧ ẋ3 = 0, (9.) ẋ1 ≥ 0 ∧ ẋ2 ≥ 0 ∧ ẋ3 ≤ 0, (10.) ẋ1 ≥ 0 ∧ ẋ2 = 0 ∧ ẋ3 ≤ 0, (11.) ẋ1 = 0 ∧ ẋ2 ≤ 0 ∧ ẋ3 ≥ 0, and (12.) ẋ1 = 0 ∧ ẋ2 ≥ 0 ∧ ẋ3 ≤ 0, where ∃t ∈ t ẋi(t) < 0 if it is stated that ẋi ≤ 0, and, analogously, ∃t ∈ t ẋi(t) > 0 if it is stated that ẋi ≥ 0. 2 note that the cases ẋ1(r, j) < 0 ∧ ẋ2(r, j) = ẋ3(r, j) = 0 and ẋ2(r, j) > 0 ∧ ẋ1(r, j) = ẋ3(r, j) = 0 are infeasible (see the discussion of combinations (a)-(e)). models of continuous dynamics on the 2-simplex and applications in economics 129 copyright ©2019 assa. adv. in systems science and appl. (2019) 0, the tangential vector angles δ(t, j) associated with x(t +, j) are equal to the angle that is associated to the line segment li(x0) in fig. 4 +/–180°. for example, if x(t +, j) is in l3(x0), then δ(t, j) ∈ {120°, 300°} for t ≥ 0. third, if x(t +, j) is located in one of the sets si(x0), then for t ≥ 0, the tangential vector angles δ(t, j) associated with x(t +, j) are within the angle range indicated by the angles associated to the line segments li(x0) and li+1(x0) that bound the set si(x0) in fig. 4. for example, if the trajectory segment x(t +, j) that is monotonous in three dimensions is in s3(x0), then δ(t, j) is within the angle range [120°, 180°] for t ≥ 0 (cf. proposition 3 and condition set p33). 2.) the geometrical interpretation of proposition 2 is analogous. in particular, the trajectory segment x(t +, j) that is monotonous in two dimensions is located in two neighboring sets sj(x0) and sk(x0), and for all t ≥ 0, the tangential vector angles δ(t, j) of x(t+, j) are within the angle range indicated by the angles associated to the two line segments lj(x0) and lk+1(x0) that bound the union of the sets sj(x0) and sk(x0) in fig. 4. for example, if the trajectory segment x(t +, j) that is monotonous in two dimensions is in s3~4(x0) = s3(x0) ∪ s4(x0), then δ(t, j) is within the angle range [120°, 240°] for t ≥ 0 (cf. proposition 2 and condition set p23). 3.) analogously, proposition 1 implies that the trajectory segment x(t +, j) that is monotonous in one dimension is located in three neighboring sets sj(x0), sk(x0), and sm(x0). moreover, for all t ≥ 0, the tangential vector angles δ(t, j) associated with this trajectory segment are within the angle range indicated by the angles associated to the two line segments lj(x0) and lm+1(x0) that bound the union of the sets sj(x0), sk(x0), and sm(x0) in fig. 4. for example, if the trajectory segment x(t +, j) that is monotonous in one dimension is in s3~5(x0) = s3(x0) ∪ s4(x0) ∪ s4(x0), then δ(t, j) is within the angle range [120°, 300°] for t ≥ 0 (cf. proposition 1 and condition set p14). this graphical interpretation highlights important implications of propositions 1-3: first, a trajectory that is monotonous (in three dimensions) is captured in a smaller subset of s than a related trajectory that is monotonous in two dimensions. second, a trajectory that is monotonous in two dimensions is captured in a smaller subset of s than a related trajectory that is monotonous in one dimension. moreover, the maximum curvature κ* of trajectories that are monotonous in three dimensions (two dimensions) is greater than the maximum curvature of related trajectories that are monotonous in two dimensions (one dimension). this intuitive discussion does not explicitly define the meaning of the term ‘related’. thus, we define the meaning of this term and then formulate corollary 1 (which is implied by propositions 1-3) on the basis of this definition such that the discussion becomes more precise. let f(x0) ⊆ j be a family of continuous trajectory segments x(t +, j) ⊂ s, j ∈ f(x0), satisfying ∀j ∈ f(x0) x(0, j) = x0 ∈ s and ∀t ∈ t + ∀j ∈ f(x0) ẋ(t, j) ≠ 0 (cf. (4)). moreover, let p11(x0), p12(x0), …, p16(x0), p21(x0), p22(x0), …, p26(x0), p31(x0), p32(x0), …, and p36(x0) denote the subfamilies of f(x0) satisfying the conditions sets p11, p12, …, p16, p21, p22, …, p26, p31, p32, …, and p36, respectively. that is, j ∈ pcd(x0) ⊂ f(x0) implies that x(t +, j) satisfies the condition set pcd, where c ∈ {1, 2, 3} and d ∈ {1, 2, …, 6}. for (h, k) ∈ {1, 2, …, 6}2, we say that the families p1h(x0) and p2k(x0) are related if ∃i ∈ {1, 2, 3} ∀j ∈ p1h(x0) ∪ p2k(x0) (∀t ∈ t + ẋi(t, j) ≥ 0) ∨ (∀t ∈ t + ẋi(t, j) ≤ 0) ∧ (∃tj ∈ t + ẋi(tj, j) ≠ 0). that is, a family defined by proposition 1 is related to a family defined by proposition 2 if there exists an i for 130 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) which the monotonicity characteristics of xi(t, j) are identical in both families. for example, the families p12(x0) and p21(x0) are characterized by a monotonously increasing x3(t, j), i.e., ∀j ∈ p12(x0) ∪ p21(x0) (∀t ∈ t + ẋ3(t, j) ≥ 0) ∧ (∃tj ∈ t + ẋ3(tj, j) > 0); thus, p12(x0) and p21(x0) are related. we define the relations between the families defined by propositions 2 and 3 analogously: for (p, q) ∈ {1, 2, …, 6}2, we say that the families p2p(x0) and p3q(x0) are related if ∃(v, w) ∈ {1, 2, 3}2 ∀j ∈ p2p(x0) ∪ p3q(x0) (∀t ∈ t + ẋv(t, j) ≥ 0) ∨ (∀t ∈ t + ẋv(t, j) ≤ 0) ∧ (∀t ∈ t + ẋw(t, j) ≥ 0) ∨ (∀t ∈ t + ẋw(t, j) ≤ 0) ∧ (∃tj ∈ t + ẋv(tj, j) ≠ 0) ∧ (∃sj ∈ t + ẋi(sj, j) ≠ 0) ∧ v ≠ w. that is, a family defined by proposition 2 is related to a family defined by proposition 3 if (a) there exists a v for which the monotonicity characteristics of xv(t, j) are identical in both families and (b) there exists a w ≠ v for which the monotonicity characteristics of xw(t, j) are identical in both families. for example, as implied by (6), properties 1 and 3, and propositions 2 and 3, the families p21(x0) and p31(x0) are characterized by (a) a monotonously decreasing x1(t, j), i.e., ∀j ∈ p21(x0) ∪ p31(x0) (∀t ∈ t + ẋ1(t, j) ≤ 0) ∧ (∃tj ∈ t + ẋ1(tj, j) < 0), and (b) a monotonously increasing x3(t, j), i.e., ∀j ∈ p21(x0) ∪ p31(x0) (∀t ∈ t + ẋ3(t, j) ≥ 0) ∧ (∃sj ∈ t + ẋ3(sj, j) > 0). thus, p21(x0) and p31(x0) are related. corollary 1: a) consider the trajectory family p1h(x0), where h ∈ {1, 2, …, 6} and x0 ≡ (x01, x02, x03) ∈ int(s). there exist two trajectory families p2k(x0) and p2m(x0), (k, m) ∈ {1, 2, …, 6}2, k ≠ m, that are related to p1h(x0) and satisfy the following condition: ∀n ∈ {k, m} a*(t +, p1h(x0)) > a*(t +, p2n(x0)) ∧ κ*(t +, p1h(x0)) > κ*(t +, p2n(x0)) (cf. sections 2.2. and 2.3). b) consider the trajectory family p2p(x0), where p ∈ {1, 2, …, 6} and x0 ≡ (x01, x02, x03) ∈ int(s). there exist two trajectory families p3q(x0) and p3r(x0), (q, r) ∈ {1, 2, …, 6}2, q ≠ r, that are related to p2p(x0) and satisfy the following condition: ∀u ∈ {q, r} a*(t +, p2p(x0)) > a*(t +, p3u(x0)) ∧ κ*(t +, p2p(x0)) > κ*(t +, p3u(x0)). proof. we only sketch here the proof. starting with corollary 1a, assume that h = 2, i.e., consider the family p12(x0). according to our definition of relatedness, p12(x0) is related to p21(x0) and p22(x0), since (6), property 3, and propositions 1 and 2 imply that p12(x0), p21(x0), and p22(x0) are characterized by a monotonously increasing x3(t, j). according to proposition 1, (7b), (7d), (7f), (7n), and the definitions of a and m given in section 2.2, the following statements are true: m(t +, p12(x0)) ⊆ s1~3(x0) = s1(x0) ∪ s2(x0) ∪ s3(x0) = {(x1, x2, x3) ∈ s: x3 ≥ x03} (24) m(t +, p21(x0)) ⊆ s1~2(x0) = s1(x0) ∪ s2(x0) ⊂ s1~3(x0) (25) m(t +, p22(x0)) ⊆ s2~3(x0) = s2(x0) ∪ s3(x0) ⊂ s1~3(x0) (26) as implied by (24), all the trajectories belonging to the family p12(x0) are located in s1~3(x0), where the latter is a triangle obtained by constructing a line on s going through x0 and being parallel to the simplex edge v1-v2 (cf. property 3a, (6), and fig. 3 and 5). according to proposition 1 and (6), all the trajectories belonging to the family p12(x0) satisfy the following vector angle condition: ∀j ∈ p12(x0) (∀t ∈ t + δ(t, j) ∈ [0°, 180°]) ∧ (∃(r, s) ∈ t + × t + δ(r, j) ∈ [0°, 60°) ∧ δ(s, j) ∈ (120°, 180°]) (27) models of continuous dynamics on the 2-simplex and applications in economics 131 copyright ©2019 assa. adv. in systems science and appl. (2019) if we allow for non-smooth trajectories and, in particular, trajectories that are unions of line segments, it is easy to show geometrically by referring to fig. 5 that such trajectories can be constructed to any point on s1~3(x0) while satisfying the condition (27).3 thus, m(t +, p12(x0)) = s1~3(x0) and, thus, a*(t +, p12(x0)) = a(s1~3(x0)) (cf. section 2.2). moreover, (25) and (26) imply that a*(t +, p21(x0)) ≤ a(s1~2(x0)) < a(s1~3(x0)) and a*(t +, p22(x0))) ≤ a(s2~3(x0)) < a(s1~3(x0)). thus, (28) is true. a*(t +, p12(x0)) > a*(t +, p21(x0)) ∧ a*(t +, p12(x0)) > a*(t +, p22(x0)) (28) if we require that the trajectories belonging to the family p12(x0) are smooth (i.e., ∀j ∈ p12(x0) ∀t ∈ t + x(t, j) is differentiable with respect to t), then it is not possible to construct a trajectory that obeys (27) and goes through the points/vertices (x01, 0, x03) ∈ s1~3(x0) and (0, x02, x03) ∈ s1~3(x0), which can be easily proven by referring to fig. 4 and 5. that is, the smooth trajectories belonging to the family p12(x0) cannot cover two infinitesimally small areas of s1~3(x0). however, even in this case, it is still ensured that a*(t +, p12(x0)) > a*(t +, p21(x0)), since (a) s1~3(x0) = s1~2(x0) ∪ s3(x0) (cf. (24) and (25)), (b) s3(x0) (cf. (24)) is not infinitesimally small (in generic cases), and (c) a*(t +, p21(x0)) ≤ a(s1~2(x0)). the definition of κ* (cf. section 2.3) and (27) imply that κ*(t +, p12(x0)) = 180°. analogously, proposition 2, the definition of κ*, and (6) imply that κ*(t +, p21(x0)) = 120°. thus, κ*(t +, p12(x0)) > κ*(t +, p21(x0)). overall, by now, we have (heuristically) proven corollary 1a for h = 2. the proof is analogous for h ∈ {1, 3, 4, 5, 6}. the proof of corollary 1b is very similar to the proof of corollary 1a. thus, we omit it here. □ overall, among related trajectories and trajectory families the following is true: the higher the dimension of monotonicity, the smaller is (a) the family image (and, thus, the set of predicted states) and (b) the maximal curvature (and, thus, the potential strength of waves). if two families are unrelated, then a higher degree of monotonicity does not necessarily imply a smaller family image and a smaller maximum curvature. we can see that corollary 1 does not categorize all the condition sets postulated by propositions 1 and 3 and, in particular, not the condition sets p17-p9 and p37-p312. these condition sets imply that one of the xi(t, j) is constant for all t ∈ t + and, thus, the trajectory segments x(t +, j) are located on line segments. obviously, the constancy requirement is much stronger than a monotonicity requirement; thus, in the cases represented by the condition sets p17-p9 and p37-p312, the size of the family image m(.) is relatively small and the curvature is zero. 3 exactly speaking, (a) each of the line segments constituting such a trajectory is characterized by an angle to the v1-v2-edge of s in the range of [0°, 180°], (b) each trajectory contains a line segment that has an angle in the range of [0°, 60°), and (c) each trajectory contains a line segment that has an angle in the range (120°, 180°] (cf. (27)). 132 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) . s1~3(x0) v1 v2 v3 x0 ≡ (x01, x02, x03). .(x01, 0, x03) (0, x02, x03) fig. 5. the set s1~3(x0). 3.2 implications for prediction of transitional dynamics if we regard t = 0 as now and t > 0 as the future (and, thus, t + as the predicted trajectory segment), corollary 1 implies that the trajectories that are monotonous in one dimension (two dimensions) are harder to predict than the trajectories that are monotonous in two dimensions (three dimensions), ceteris paribus, since (a) the set of all possible future states is relatively great and (b) relatively stronger curvatures/waves may arise in the former case (in comparison to the latter case). however, besides the monotonicity characteristics, the location of the initial state x0 is decisive for the predictability of the future dynamics. in particular, the family image m and its size a* depend on x0 (cf. corollary 1 and its proof). in general, monotonicity implies that the system moves from x0 along the trajectory segment t + towards a vertex or an edge of the 2-simplex. thus, if the initial state x0 is relatively close to this vertex/edge, t + is captured in a relatively small set, i.e., the set of potential future states of the system is relatively small. this is almost a direct implication of the boundedness of the 2-simplex. as implied by propositions 1-3, (6), and section 2.3 (cf. proof of corollary 1), the maximum curvature κ* of trajectories that are monotonous in three dimensions, two dimensions, and one dimension is 60°, 120°, and 180°, respectively. thus, the cyclical behaviors corresponding to a transversally self-intersecting trajectory or a closed trajectory (jordan curve) are prohibited in all cases of monotonicity, since these types of cyclical behavior require a curvature greater than 180°. yet, monotonous trajectories allow for a cyclical behavior corresponding to waves on the simplex (see fig. 6). the angle range of 180° associated with monotonicity in one dimension allows for waves of high amplitude and short wavelength on the 2-simplex. in contrast, monotonicity in three dimensions allows only for relatively low-amplitude/long-wavelength waves (cf. fig. 6). models of continuous dynamics on the 2-simplex and applications in economics 133 copyright ©2019 assa. adv. in systems science and appl. (2019) . s3(x0) s2(x0) s1(x0) v1 v2 v3 x0 d c b l1(x0) = || l2(x0) = || l3(x0) = || l4(x0) = || a x(t +, j) monotonous in one dimension (condition set p12) . s2(x0) s1(x0) v1 v2 v3 x0 c b a x(t +, j) monotonous in two dimensions (condition set p21) . s1(x0) v1 v2 v3 x0 b monotonous in three dimensions (condition set p31) a x(t +, j) fig. 6. examples of monotonous waves. obviously, all three types of monotonicity allow for curved trajectories on the 2-simplex. however, only monotonicity in three dimensions allows for unidirectional linear trajectories, while monotonicity in one dimension allows for linear (non-smooth) trajectories yet requires at least one direction change. 134 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) 3.3 limit dynamics in all three cases of monotonicity (monotonicity in one, two, and three dimensions), the following two facts are true for continuous dynamics. first, the cyclical limit dynamics where the omega limit set is a jordan curve are excluded, since such cycles require that x1(t), x2(t), and x3(t) are non-monotonous in the limit. second, the system may converge to a fixed point or reach the fixed point in finite time (and stay there). the proof of these facts is obvious. in the case of monotonicity in three dimensions, only the fixed point outcome is possible, as implied by the monotone convergence theorem: since each of the functions x1(t), x2(t), and x3(t) is monotonous and restricted by an upper/lower limit of 0 and 1, each of the x1(t), x2(t), and x3(t) converges to its fixed point (x1 *, x2 *, and x3 *, respectively) or reaches it in finite time (and stays there). thus, x(t) converges to a fixed point x* ≡ (x1 *, x2 *, x3 *) ∈ s or reaches it in finite time. additionally, in the case of monotonicity in one dimension, the system may converge to a line segment if the trajectory is a wave. in this case, the wavelength decreases and the vector-angle range converges to the range of 180° as the system converges to the line segment (see the first part of fig. 6). note that in the case of monotonicity in two dimensions, the convergence to a line segment is not possible, as explained in the following. if the trajectory converges to a line segment, the tangential vector angle range must increase to a range of 180° which is prohibited by the definition of monotonicity in two dimensions, which allows only for a vector angle range of 120°. the former fact follows from the definition of the omega limit set, where for each of the points on the line segment (constituting the omega limit set), a sequence of points on the wave must be found that converges to it. 4. implications of non-self-intersection for transitional dynamics while non-intersecting trajectories and limit dynamics are treated in section 5, we focus, now, on the implications of (transversal) non-self-intersection for transitional dynamics. the class of transversally non-self-intersecting continuous trajectories on the 2-simplex is a subclass of the class of continuous trajectories on the 2-simplex. moreover, the class of non-self-intersecting continuous trajectories on the 2-simplex is a subclass of the class of transversally non-self-intersecting trajectories on the 2-simplex, since the former does not allow for jordan-curves in contrast to the latter. thus, by imposing the condition of (transversal) non-self-intersection, we can reduce the set of feasible trajectories on the 2-simplex, which can be exploited in prediction of dynamics, as explained in the following. obviously, the non-self-intersection is an important constraint in systems of continuous trajectories on two-dimensional domains. in the case of discontinuous trajectories, non-self-intersection still may reduce the class of feasible trajectories significantly depending on the type of discontinuity and the physical/social system being analyzed. however, in extreme cases and, in particular, in the case of point sequences on the models of continuous dynamics on the 2-simplex and applications in economics 135 copyright ©2019 assa. adv. in systems science and appl. (2019) 2-simplex (e.g., discrete-time paths), non-self-intersection becomes obsolete as a restraint (in natural and social sciences where the exact position of a system on the simplex is not measurable). for the same reasons, non-self-intersection is an obsolete restraint in threeor higher-dimensional dynamical system domains (see, e.g., [13]). 4.1 qualitative simulation the (transversal) non-self-intersection constraint on continuous trajectories on the 2-simplex can be understood as a dynamic constraint: at each point of time t ∈ t +, we have a restriction on system dynamics x(t, j) preventing certain type of dynamics (namely the dynamics that correspond to a self-intersection of the trajectory). the constraint is dynamic in the sense that it changes over time. in particular, it depends on the current position of the system on the 2-simplex and the form of the trajectory segment x(t –, j) representing the dynamics over the past time period t – (e.g., the longer the latter segment, the stronger is the constraint on the current dynamics), i.e., the constraint is updated continuously. this fact can be used in qualitative simulation, as discussed in detail by [7-8]. 4.2 non-self-intersection in combination with a determined trajectory-segment the concept of (transversal) non-self-intersection can be very useful even if we do not assume the dynamic constraint view discussed in section 4.1. in particular, assume that the trajectory segment (x(t –, j)) representing past dynamics is given by empirical data on past dynamics or by empirical laws. then, in general, x(t –, j) can be used as a basis for a partitioning of the 2-simplex. for example, since the lines that are parallel to the 2-simplex edges have a clear intuitive interpretation, x(t –, j) and such lines can constitute an intuitively meaningful partitioning of the simplex (see [13]). then, paths on the 2-simplex can be understood as sequences of partition switches, and the non-self-intersection constraint as an exclusion of certain switches, as demonstrated in fig. 7, where (immediate) switches between the partitions a and c are prohibited by the non-self-intersection constraint. if we interpret t = 0 as present, t – as past, and t + as future, this prohibition corresponds to infeasible future scenarios, i.e., non-self-intersection can be used in prediction of future dynamics (cf. [13]). obviously, depending on the (natural/social sciences) topic analyzed by these concepts, a certain length and positioning of x(t –, j) may be necessary to derive significant predictions. in particular, if x(t –, j) is relatively short or located in a relatively small or peripheral subset of the 2-simplex, it may not be possible to establish a relevant partitioning inducing a prohibition of paths that is of significant relevance for the topic/theory being analyzed (cf. [13]). these requirements are well known in statistics, where the length of the past time series and avoidance of outliers is important for the (statistical) significance of the predictions based on empirical (time-series) data. 136 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) x(t –, j) x(0, j) x(t +, j) v2v1 v3 d c a b fig. 7. paths as partition switches. 5. poincaré-bendixson theory the poincaré-bendixson theory, which is one of the fundaments of the dynamical systems theory, can be used to predict the qualitative properties of the limit dynamics of a smooth dynamical system in the plane. it applies to continuous systems, yet requires additional restrictions on the system (to ensure a sufficient degree of smoothness). we discuss here these requirements from a rather topological point of view applying the concepts discussed in section 2.4. for a general, discussion of the requirements and predictions/statements of the poincaré-bendixson theory, see, e.g., [1, pp.362f.], [2], [3, p.45], [4, p.55], and [21, chapter 7.3]. assume that the dynamics on the 2-simplex are representable by a (relatively) smooth autonomous differential equation system in terms of the coordinates (y1, y2) of a two-dimensional coordinate system that is parallel to the 2-simplex (see fig. 8). then, the poincaré-bendixson theory states that the limit dynamics of this system are either cyclical or transitory. in particular, the omega limit set of a trajectory generated by such a system consists of a fixed point, a jordan curve, or a homo-/heteroclinic union (of curves and fixed points). the geometrical interpretation of the requirement of the representability by a smooth differential equation system in y1-y2-coordinates is that the trajectories of the dynamical system on 2-simplex constitute a simple covering (of a connected subset) of the 2-simplex. in particular, such a simple covering consists of non-intersecting and transversally non-self-intersecting trajectories, where the union of these trajectories is a connected subset of the 2-simplex (see [16,20] for a detailed discussion and literature references). 6. applications in economics the mathematical theories of continuous dynamics on the standard 2-simplex developed in the previous sections have almost direct applications in the analysis of economic dynamics. in particular, they can be used in prediction of economic structural change, discussion of models of continuous dynamics on the 2-simplex and applications in economics 137 copyright ©2019 assa. adv. in systems science and appl. (2019) sectoral production functions, assessment of structural change costs, and design of cost-minimal structural change policies, as discussed in the following sections. v3 v2v1 y2 y1 ẏ(t) = φ(y(t)) y(t) ≡ (y1(t), y2(t)) φ: u → u' u, u' ⊂ ℝ2 fig. 8. representation of a dynamical system on s by a two-dimensional differential equation system. 6.1 economic topics covered by the models of continuous dynamics on the 2-simplex a major pillar of economics is the study of long-run economic dynamics, where short-run fluctuations are neglected and the dynamic patterns that persist over long periods of time (e.g., 100 years) are studied. in this context, the concept of structural change is essential, where not only aggregate economic indices (e.g., gross domestic product, trade volume, and economy-wide employment) are studied but also their structure. in particular, the aggregate indices are subdivided into components and the significance of these components for the aggregate index is indicated by the components’ shares in the aggregate index. many of these ‘shares’, such as savings rate, investment rate, and sectoral employment shares, are well known even in public debates. in general, structural change refers to the dynamics of these ‘shares’, where the shares satisfy the conditions stated by (1). in other words, economic structural change, i.e., the long-run dynamics of the ‘shares’ can be depicted by trajectories on standard simplexes. moreover, the assumption of continuous-time frameworks and continuous functional forms is a general convention in long-run economic dynamics modeling (although there are exceptions from this convention), which, in general, yields continuous dynamics of the shares on standard simplexes. for an overview of the topics that are covered by the system-theoretical models of continuous trajectories on standard simplexes and for corresponding references from the economics literature, see [15]. to provide some details and references on the economic applications of the system-theoretical models derived in the previous sections, we focus on a specific sort of economic structural change, namely long-run labor allocation dynamics in the three-sector framework, in section 6.2. 138 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) 6.2 example: long-run labor allocation dynamics the three-sector framework is one of the major concepts for studying economic structural change (for an overview of the literature, see, e.g., [5-6,9-12]. it is based on the assumption that economic activities can be divided into three categories or ‘sectors’: agriculture, manufacturing, and services. one of the major indices for studying the importance of these sectors are the shares of each of the sectors in aggregate employment (abbr. ‘employment shares’). these shares and their dynamics can be depicted by continuous trajectories on the 2-simplex (cf. [13,14]). thus, we can use the concepts discussed in sections 2-5 in the analysis of labor allocation dynamics as discussed in the following. while it is difficult to derive any consensus statements on the quantitative characteristics of labor allocation dynamics, the literature implies that there seem to be some empirically observable qualitative laws of labor allocation dynamics, which can be expressed by using the geometrical concepts discussed in our paper. in particular, [19] shows that the typical long-run labor allocation dynamics of a nowadays highly developed country over the last two centuries can be described by a trajectory that has the following characteristics: (1.) it is monotonous in two dimensions (cf. sections 2.3 and 3) and (thus) non-self-intersecting (cf. sections 2.4 and 4); (2.) it has a negative signed curvature κ (cf. section 2.3); (3.) its initial segment is located in the simplex partition sv1 (cf. (5) and fig. 2); and (4.) its final segment is located in partition sv3 (cf. (5) and fig. 2). as discussed by [14], these empirical observations can be interpreted as “natural” laws of structural change (since, among others, they are supported by the theoretical literature consensus) and, thus, can be exploited for predictions of structural change. for example, proposition 2 and corollary 1 (cf. [14]) and the approach discussed in section 4.2 (cf. [13]) can be used to predict the future (transitional) dynamics of labor allocation in developing and developed economies. moreover, [16,20] uses the empirical findings of [19] and the topological approach discussed in section 5 for a discussion of the applicability of the poincaré-bendixson theory in the prediction of limit dynamics of labor allocation. beside these applications, which focus on prediction of structural change, the models of continuous dynamics discussed in our paper have further applications in in structural change modeling: [17] shows that labor allocation trajectories that are monotonous in three dimensions minimize the structural change costs (e.g., unemployment, geographical relocation costs, and environmental pollution) and uses this result to elaborate a development policy minimizing the labor reallocation costs in a developing economy. [18] uses the model of monotonous and continuous trajectories and the concept of curvature (cf. section 2.3) to discuss a widespread assumption in theoretical structural change modeling (namely the assumption of cobb-douglas production functions) by applying an axiomatic-geometrical approach. finally, [15] discusses how the topological concepts discussed in section 2.4 can be interpreted and applied in the context of labor allocation dynamics modeling. models of continuous dynamics on the 2-simplex and applications in economics 139 copyright ©2019 assa. adv. in systems science and appl. (2019) 7. conclusions in this paper, we have studied models of continuous dynamics on the 2-simplex that arise when set-theoretical, differential-geometrical, or topological restrictions are imposed on the trajectories of the model. we focused on the qualitative properties of transitional and limit dynamics of these models and discussed their applications in long-run economic dynamics modeling. many of our results (in general, the results of section 3 and their applications in economics) can be extended to discrete or discontinuous systems or higher-dimensional simplexes. yet, the rather topological concepts discussed in sections 4 and 5 (e.g., the poincaré-bendixson theory) are, in general, not applicable or not useful in discrete or higher-dimensional systems and their applications (cf. [13,15-16,20]). in the latter systems, the concept of chaos as well as existence theorems on fixed points are of interest. thus, further research could focus on them and, in particular, their system-theoretical significance for long-run economic dynamics. empirical evidence implies that there are fluctuations of the labor allocation shares that correspond to the waves on the 2-simplex discussed in sections 3.2 and 3.3. further research could focus on a detailed discussion of waves on the 2-simplex and the application of the resulting system-theoretical models in the explanation of the empirically observed waves arising in labor-allocation dynamics. acknowledgement no funding has been received for this research. references [1] andronov, a.a, vitt, a.a. & khaikin, s.e. (1987). theory of oscillators. dover publications, mineola, new york. [2] ciesielski, k. (2012). the poincaré–bendixson theorem: from poincaré to the xxist century, central european journal of mathematics, 10, 2110–2128. [3] guckenheimer, j. & holmes, p. (1990). nonlinear oscillations, dynamical systems, and bifurcations of vector fields. springer, new york. [4] hale, j.h. (2009). ordinary differential equations. dover publications, mineola, new york. [5] herrendorf, b., rogerson, r. & valentinyi, á. (2014). growth and structural transformation, in p. aghion & s.n. durlauf (eds.), handbook of economic growth, volume 2b. elsevier b.v., amsterdam. [6] krüger, j.j. (2008). productivity and structural change: a review of the literature, journal of economic surveys, 22, 330–363. [7] kuipers, b. (1986). qualitative simulation, artificial intelligence, 29, 289–338. 140 d.stijepic copyright ©2019 assa. adv. in systems science and appl. (2019) [8] lee, w.w. & kuipers, b.j. (1988). non-intersection of trajectories in qualitative phase space: a global constraint for qualitative simulation, proceedings of the seventh national conference on artificial intelligence (aaai-88), morgan kaufmann, los altos, ca. [9] van neuss, l. (2018). the drivers of structural change, forthcoming in journal of economic surveys, 32. [10] schettkat, r. & yocarini, l. (2006). the shift to services employment: a review of the literature, structural change and economic dynamics, 17, 127–147. [11] silva, e.g. & teixeira, a.a.c. (2008). surveying structural change: seminal contributions and a bibliometric account, structural change and economic dynamics, 19, 273–300. [12] stijepic, d. (2011). structural change and economic growth: analysis within the partially balanced growth-framework. südwestdeutscher verlag für hochschulschriften, saarbrücken. an older version is online: http://deposit.fernuni-hagen.de/2763/ [13] stijepic, d. (2015). a geometrical approach to structural change modelling, structural change and economic dynamics, 33, 71–85. [14] stijepic, d. (2017a). positivistic models of structural change, journal of economic structures, 6, 1–30. [15] stijepic, d. (2017b). empirical evidence on the topological properties of structural paths and some notes on its theoretical explanation, mpra working paper no. 82473. [online]. available https://mpra.ub.uni-muenchen.de/82473/1/mpra_paper_82473.pdf [16] stijepic, d. (2017c). on the predictability of economic structural change by the poincaré-bendixson theory. [online]. available https://ssrn.com/abstract=3015889. [17] stijepic, d. (2017d). on development paths minimizing the structural change costs in the three-sector framework and an application to structural policy. [online]. available https://ssrn.com/abstract=2919806 [18] stijepic, d. (2017e). an argument against cobb-douglas production functions (in multi-sector growth modeling), economics bulletin, 37, 1143–1150. [19] stijepic, d. (2018). empirical evidence on the geometrical properties of structural change trajectories, research journal of economics, 1, 1–10. [20] stijepic, d. (2019). on the predictability of economic structural change by the poincaré-bendixson theory, forthcoming in foresight. [21] teschl, g. (2012). ordinary differential equations and dynamical systems. american mathematical society, providence, rhode island. http://deposit.fernuni-hagen.de/2763/ https://mpra.ub.uni-muenchen.de/82473/1/mpra_paper_82473.pdf https://ssrn.com/abstract=3015889 https://ssrn.com/abstract=2919806 advances in systems science and applications (2013) vol.13 no.1 53-67 soft probability of large deviations d.a. molodtsov dorodnitsyn computing center, russian academy of sciences, moscow, russia abstract a concept of soft probability is presented. an analogue of chebyshev’s inequality for soft probability is proved. soft large deviation probabilities for a nonnegative random variable under a mean hypothesis is calculated. keywords soft probability, hypothesis replacement, soft chebyshev inequality, large deviations. 1 what is soft probability? in modern textbooks of probability theory, the exposition usually begins with a discussion of the subject matter of this science. all phenomena are divided into three types. phenomena of the first type are those characterized by deterministic regularity; this means that a given set of circumstances always leads to the same outcome. phenomena of the second and the third type are not deterministic regular. the second type consists of statistically regular phenomena, and the third type, of the remaining phenomena. by statistical regularity the statistical stability of outcome frequencies is usually understood. as a rule, this notion is associated with the example of coin tossing. some authors outline more constructive ways of interpretation [1], but a final formalization of statistical stability has never been proposed, and its verification is always left to the reader’s judgment and intuition. thus, the most important question of the applicability of the theory to real phenomena remains essentially unanswered. soft probability [2-7] is merely the logical completion of the construction of statistical regularity, whose approximate description is usually contained in probability theory textbooks. our approach is based on examine the conclusions to which the logical development of the notion of statistical regularity leads. first, we mention at once that the verification of statistical regularity suggested here requires a certain set of trial outcomes, that is, a statistical database. if there are no trial results, then there is no object of examination. 2 statistical database first, we introduce the base outcome space ω. each trial is associated with an element of the set ω, namely, the outcome of this trial. a statistical database is merely a finite sequence of outcomes: base = {ω1, ..., ωn}, ωi ∈ ω. 54 d.a molodtsov: soft probability of large deviations by an event a we mean a subset of ω : a ⊆ ω. we say that an event occurs under a certain trial if the outcome of this trial belongs to the given event. we have to define the statistical regularity of the occurrence of an event a. in effect, statistical regularity means the closeness of frequencies for a given set of samples. thus, in formalizing this concept, we must begin with specifying a set of samples for a statistical database. we identify a sample i with the positions in the sequence base = {ω1, ..., ωn} occupied by its elements; in other words, a sample is a subset i ⊆ ind(base) of the set ind(base) = {1, ..., n}. in specifying a set of samples, it is natural to first constrain their size; thus, we introduce a parameter determining the size of a sample. as admissible samples we consider any samples of size m which consists of consecutive elements of the set ind(base) and are not too “old”. we denote the set of such samples by s(base,m, τ): s(base,m, τ) = {(i, i+ 1, ..., i+m− 1), i = τ, ..., n −m+ 1}. apparently, the set s(base,m, τ) is a minimal set of samples which are natural to consider in defining the notion of statistical regularity. of course, other definitions of admissible samples are also possible, which lead to different definitions of statistical regularity; it is important that the set of admissible samples be precisely specified. we define the frequency of occurrence of an event a in a sample i as µ(base, χ(a, ·), i) = 1 |i| ∑ i∈i χ(a,ωi). here |i| denotes the cardinality of the set i and χ(a,ω) = {1,ω∈a0,ω /∈a. 3 statistical regularity now, it is natural to understand the statistical regularity of the occurrence of an event a as the closeness of the frequencies of occurrence of a in any admissible samples from the set s(base,m, τ). below we give a more formal definition of this notion. definition 1. an event a is said to be statistically (m, τ, δ) − regular on a database base if |µ(base, χ(a, ·), i)− µ(base, χ(a, ·), j)| ≤ δ. for any samples i, j ∈ s(base,m, τ). we see that the definition of statistical regularity involves several parameters. apparently, it is for this reason that this notion has not been used, because the advances in systems science and applications (2013) vol.13 no.1 55 dependence of statistical regularity on parameters inevitably makes probability dependent on parameters as well, while probability is traditionally thought of as a single real number in the unit interval. however, the parameterization of probability only means the detailing of the description of the situation under consideration. in different situations, probabilities corresponding to different parameters are important. the description using the same number as probability in all situations is coarser than that using parameterized probability. parameterized families are used very extensively in theory and practice. thus, different methods of measuring physical quantities yield different results, which naturally leads to parameterized families. a striking example of such a family is the description of mine deposits in geology. in computational mathematics, it often happens that only approximate solutions can be found numerically; such solutions form parameterized families too [2,8]. for dealing with such objects, the notion of a soft set was introduced and the theory of soft sets was developed, which has found numerous applications in various areas of mathematics [2-7, 9-24]. for this reason, the alternative to classical probability considered in this paper is called soft probability. let us contemplate definition 1. the statistical database consists of the outcomes of events which have already occurred, while the main purpose of the theory is to produce informative statements concerning future events. thus, the only interesting aspect of the notion of statistical regularity is its use as a hypothesis on the future behavior of trial outcomes. accordingly, the application of soft probability is divided into two processes: • verifying hypotheses at each step. • given a set of accepted hypotheses, constructing estimates, predictions, etc. note that, in classical probability theory, the former process is virtually absent; the statistical regularity hypothesis is accepted only once, before launching the apparatus of probability theory, after which this hypothesis is neither controlled nor verified anew. moreover, in classical probability theory, the statistical regularity hypothesis is accepted for all possible events together. it remains unclear how to handle situations in which some of the events are statistically regular and the other events are not. the definition of statistical regularity is easy to generalize to a random function. by a random function f we mean any real-valued function defined on the set ω, that is, any function f : ω → e, where e is the set of real numbers. an example of such a function is the characteristic function χ(a, ·) of a set. 56 d.a molodtsov: soft probability of large deviations the mean value of a random function f on a sample i is defined as µ(base, f, i) = 1 |i| ∑ i∈i f(ωi). definition 2. a random function f is said to be statistically (m, τ, δ) − regular on a database base if |µ(base, f, i)− µ(base, f, j)| ≤ δ. for any samples i, j ∈ s(base,m, τ). thus, the statistical regularity of an event means simply that the characteristic function of this event is statistically regular. the condition that a random function is statistically regular can be written in a different equivalent form as follows. we set 2a = max i∈s(base,m,τ) µ(base, f, i) + min i∈s(base,m,τ) µ(base, f, i), and 2b = max i∈s(base,m,τ) µ(base, f, i)− min i∈s(base,m,τ) µ(base, f, i). it is easy to see that the statistical regularity of f is equivalent to the inequality 2b ≤ δ. for any i ∈ s(base,m, τ), we have |µ(base, f, i)−a| ≤ b. this readily implies the equivalence of the statistical regularity of f to the fulfillment of the inequality |µ(base, f, i) − a| ≤ δ/2, or the inclusion µ(base, f, i) ∈ [a − δ/2, a + δ/2], for any i ∈ s(base,m, τ). the inclusion is also equivalent to [ min i∈s(base,m,τ) µ(base, f, i), max i∈s(base,m,τ) µ(base, f, i)] ⊆ [a− δ/2, a+ δ/2]. this suggests the following natural definition. definition 3. the interval λ(f,base,m, τ) = [ min i∈s(base,m,τ) µ(base, f, i), max i∈s(base,m,τ) µ(base, f, i)] is called the (m, τ)− approximate mean value of the random function f on the database base. we denote the left and right endpoints of this interval by an underscore and an overscore, respectively: λ(f,base,m, τ) = [λ(f,base,m, τ), λ(f,base,m, τ)]. advances in systems science and applications (2013) vol.13 no.1 57 4 hypotheses on the behavior of a random function on the basis of the notion of the statistical regularity of a random function, we can formulate hypotheses of two types on the future behavior of the values of a random function. in fact, these are hypotheses on the future values of the statistical database. definition 4. a databasebase is said to be statistically (m, δ)− regular with respect to a random function f if |µ(base, f, i)− µ(base, f, j)| ≤ δ. for any samples i, j ∈ s(base,m, 1). definition 5. a databasebase is said to be statistically significantly (m, δ, a)− regular with respect to a random function f if |µ(base, f, i)− a| ≤ δ. for any samples i ∈ s(base,m, 1). the difference between definitions 4 and 5 is in that definition 4 supposes only the closeness of mean values on any admissible samples, while definition 5 specifies the number to which these mean values must be close with a given accuracy. in other words, the mean values must belong to the interval [a− δ.a+ δ], or the approximate mean values of the random function under consideration must belong to [a− δ.a+ δ]. these definitions can also be regarded as hypotheses on certain properties of the approximate means of a random function. definition 4 specifies only the length of an interval containing the approximate mean, and definition 5 specifies the boundaries of the entire range of this mean. an analysis shows that these hypotheses are often insufficient for obtaining instructive results. in addition to hypotheses on approximate mathematical expectation, hypotheses describing the deviation of the random function under consideration from its mathematical expectation (that is, hypotheses similar to that of the existence of variance) are very useful. such hypotheses can be formulated in various forms. we give only one version. together with a random function f , consider the random function equal to the absolute value of the deviation of f from the interval [a− δ.a+ δ], that is, defined by g(ω) = max{|f(ω)− a| − δ, 0}. definition 6. we say that a database base satisfies the (m, δ, a,△) − variance hypothesis with respect to a random function f if base is statistically significantly (m,△, 0) − regular with respect to the random function g, i.e. µ(base, g, i) ≤ △. 58 d.a molodtsov: soft probability of large deviations for any samples i ∈ s(base,m, 1). the approximate mathematical expectation and approximate variance hypotheses describe a variable oscillating about a certain interval. it is also natural to consider other hypotheses, which describe the growth, decline, periodicity, and other properties of random variables [6-7]. after classes of hypotheses are chosen, dealing with hypotheses is a dynamical step process. at each step, new realization from the base space (trial outcomes) appear. thus, at each step, it is required to form a list of accepted hypotheses and perform calculations on the basis of this list. 5 properties of the approximate mean value of a random function properties of an approximate mean are similar to those of mathematical expectation. inequalities involving intervals are assumed to hold componentwise, that is, at each endpoint of the interval. 1). λ(c,base,m, τ) = [c, c], c ∈ e. 2). if f(ω) ≥ g(ω) for any ω ∈ ω, then λ(f,base,m, τ) ≥ λ(g,base,m, τ). 3). λ(cf,base,m, τ) = cλ(f,base,m, τ), c ∈ e, c ≥ 0. 4). λ(f + c,base,m, τ) = λ(f,base,m, τ) + c, c ∈ e. 5). λ(−f,base,m, τ) = −λ(f,base,m, τ) = [−λ(f,base,m, τ),−λ(f,base,m, τ)]. 6). λ(f + g,base,m, τ) ⊆ λ(f,base,m, τ) + λ(g,base,m, τ). 7). λ(f,base,m, τ) ⊆ λ(f,base,m, t), τ ≥ t. 6 an approximate variance of a random variable in classical probability theory, the variance of a random variable is defined as the mathematical expectation of the squared difference between this variable and its expectation. the mathematical expectation of a random function is merely a number. in the case under consideration, the approximate mean is a parametric family of intervals which characterizes the random variable on a certain database. hypotheses involving the notion of statistically significant regularity impose constraints on approximate means. the intervals describing the constraints are not required to equal the corresponding approximate means. thus, it is convenient to define approximate variance as the measure of deviation of a random variable from a certain interval rather than from the corresponding approximate mean. definition 7. the approximate (m, τ, a, δ) − variance of a random function f on a database base is the (m, τ)− approximate mean of the random function max{|f(ω)− a| − δ, 0}, that is, the interval d(f,base,m, τ, a, δ) = λ(max{|f(·)− a| − δ, 0}, base,m, τ). advances in systems science and applications (2013) vol.13 no.1 59 the simplest properties of approximate variance are as follows. 1). d(f,base,m, τ, a, δ) ≥ 0. 2). d(cf,base,m, τ, ca, cδ) = cd(f,base,m, τ, a, δ), c ∈ e, c ≥ 0. 3). d(−f,base,m, τ,−a, δ) = d(f,base,m, τ, a, δ). 4). d(f+g,base,m, τ, a+b, δ+γ) ≤ d(f,base,m, τ, a, δ)+d(g,base,m, τ, b, γ). 5). d(f+g,base,m, τ, a+b, δ+γ) ≤ d(f,base,m, τ, a, δ)+d(g,base,m, τ, b, γ). 6). d(f,base,m, τ, a, δ) ≥ d(f,base,m, τ, a, γ), δ ≤ γ. 7). d(f,base,m, τ, a, δ) ⊇ d(f,base,m, γ, a, δ), τ ≤ γ. 7 chebyshev’s inequality for soft probability statement 1 (chebyshev’s inequality). if a random function f is nonnegative everywhere on ω and ε > 0, then λ(χ({ω′ ∈ ω|f(ω′) ≥ ε}, ·), base,m, τ) ≤ 1 ε λ(f,base,m, τ). proof. if a random function f is nonnegative everywhere on ω, then, as is easy to see, we have f(ω) ≥ εχ({ω′ ∈ ω|f(ω′) ≥ ε}, ω) for any ε > 0 and any ω ∈ ω. property 2 of approximate mean implies λ(f,base,m, τ) ≥ ελ(χ({ω′ ∈ ω|f(ω′) ≥ ε}, ω), base,m, τ). the approximate means of the characteristic function of a set equals the soft probability of this set; therefore, the soft probability of the set {ω′ ∈ ω|f(ω′) ≥ ε} is estimated as λ(χ({ω′ ∈ ω|f(ω′) ≥ ε}, ·), base,m, τ) ≤ 1 ε λ(f,base,m, τ) (recall that the inequality is interval). this completes the proof of the statement. now, let f be an arbitrary random function. for the function |f |, we have λ(χ({ω′ ∈ ω| |f(ω′)| ≥ ε}, ·), base,m, τ) ≤ 1 ε λ(|f |, base,m, τ). since the inequality |f(ω′)| ≥ ε is equivalent to |f s(ω′)| ≥ εs, where s > 0, it follows that λ(χ({ω′ ∈ ω| |f(ω′)| ≥ ε}, ·), base,m, τ) ≤ 1 εs λ(|f |s, base,m, τ). therefore, λ(χ({ω′ ∈ ω| |f(ω′)| ≥ ε}, ·), base,m, τ) ≤ inf s>0 1 εs λ(|f |s, base,m, τ). 60 d.a molodtsov: soft probability of large deviations now, consider the nonnegative random function g(ω) = max{|f(ω)−a|−δ, 0}, where f is an arbitrary random function. applying chebyshev’s inequality to g, we obtain λ(χ({ω′ ∈ ω|max{|f(ω′)−a|−δ, 0} ≥ ε}, ·), base,m, τ) ≤ 1 ε d(f,base,m, τ, a, δ). elementary transformations yield λ(χ({ω′ ∈ ω||f(ω′)− a| ≥ δ + ε}, ·), base,m, τ) ≤ 1 ε d(f,base,m, τ, a, δ). this allows us to write the inequality in the equivalent form λ(χ({ω′ ∈ ω||f(ω′)− a| ≥ δ}, ·), base,m, τ) ≤ inf 0<ε<δ 1 ε d(f,base,m, τ, a, δ − ε). thus, chebyshev’s inequality gives an estimate of the soft probability of “large” deviations in terms of approximate mean or approximate variance. if the database is known, then this information is of little value, because it is easy to directly calculate the exact values of any probabilistic characteristics, including the probabilities of large deviations. apparently, it is of more interest to apply this inequality to estimating the probability of a random function on a future database. naturally, this requires hypotheses on the approximate mean or the approximate variance of the function under consideration. however, in the presence of hypotheses, of interest are sharp bounds for the probability of large deviations. 8 soft probability of large deviations for a nonnegative random function under an approximate mean hypothesis suppose that a database base is statistically significantly (m, δ, a) − regular with respect to a nonnegative random function f , i.e., given any sample i ∈ s(base,m, 1), we have |µ(base, f, i)− a| ≤ δ we assume that a ≥ 0. we define a large deviation as the event a(f, ε) = {ω ∈ ω|f(ω) ≥ ε}, where ε > 0. we are interested in the range of values of the soft probability of this event under the above hypothesis. clearly, the solution essentially depends on the range of the function f . consider the case where f(ω) = e+ = {x ∈ e|x ≥ 0}. let f(ωi) = xi ∈ e+. then the constraints on the database can be written as constraints on the vector x = {x1, ...xn} ∈ x(n,m, a, δ), where x(n,m, a, δ) = {x ∈ en +| | j+m−1∑ i=j xi − am| ≤ δm, j = 1, ..., n−m+ 1}. advances in systems science and applications (2013) vol.13 no.1 61 the boundaries of the range of the soft probability are u∗(n,m, a, k, τ, δ, ε) = sup x∈x(n,m,a,δ) max τ≤j≤n−k+1 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi)}. and u∗(n,m, a, k, τ, δ, ε) = inf x∈x(n,m,a,δ) min τ≤j≤n−k+1 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi)}. note that, at n = m, the set x(n,m, a, δ) takes the form x(m,m, a, δ) = {x ∈ em + | | m∑ i=1 xi − am| ≤ δm}. in the case k ≤ m, the evaluation of u∗ and u∗ is based on the following assertion. statement 2. if x ∈ em + and (xj , xj+1, ..., xj+m−1) ∈ x(m,m, a, δ) , then there exists a vector y ∈ x(n,m, a, δ) such that yi = xi for i = j, j+1, ..., j+m−1. proof. let yi = xj+(i−j)modm for i = 1, ..., n. then ∑l+m−1 i=l yi = ∑j+m−1 i=j xi for any l = 1, ..., n−m+1. therefore, y ∈ x(n,m, a, δ). this proves the required assertion. the u∗ and u∗ problems can be formulated as u∗(n,m, a, k, τ, δ, ε) = max τ≤j≤n−k+1 sup x∈x(n,m,a,δ) 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi)}. and u∗(n,m, a, k, τ, δ, ε) = min τ≤j≤n−k+1 inf x∈x(n,m,a,δ) 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi)}. at k ≤ m, in the problems sup x∈x(n,m,a,δ) 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi) and inf x∈x(n,m,a,δ) 1 k j+k−1∑ i=j χ({y|y ≥ ε}, xi) the function to be optimized depends only on the variables xj = (xj , xj+1, xj+m−1); hence, we can perform optimization over the projection of the set x(n,m, a, δ) 62 d.a molodtsov: soft probability of large deviations on the corresponding coordinates rather over the entire set. it follows from statement 2 that this projection coincides with the set x(m,m, a, δ); thus, at k ≤ m, the u∗ and u∗ problems take the forms u∗(n,m, a, k, τ, δ, ε) = sup x∈x(m,m,a,δ) 1 k k∑ i=1 χ({y|y ≥ ε}, xi)}. and u∗(n,m, a, k, τ, δ, ε) = inf x∈x(m,m,a,δ) 1 k k∑ i=1 χ({y|y ≥ ε}, xi)}. for x ∈ em + , we set π(x,m, ε) = |{i ∈ {1, ...,m}|xi ≥ ε}|; this is the number of components greater than or equal to ε. we have u∗(n,m, a, k, τ, δ, ε) = sup x∈x(m,m,a,δ) min{π(x,m, ε), k} k = 1 k min{ sup x∈x(m,m,a,δ) π(x,m, ε), k}. and u∗(n,m, a, k, τ, δ, ε) = 1 k max{k + inf x∈x(m,m,a,δ) π(x,m, ε)−m, 0}. thus, it is required to find the maximum and the minimum value of the function π(x,m, ε) on x(m,m, a, δ). let us introduce the set π(m, ε, p) = {x ∈ em + |π(x,m, ε) = p}. it is easy to see that the image of the function ∑m i=1 xi on the set π(m, ε, p) equals { [εp,+∞), p > 0 [0,mε), p = 0, therefore, the function π(x,m, ε) takes the value p on the set x(m,m, a, δ) if and only if π(m, ε, p) ∩x(m,m, a, δ) ̸= ∅. that is, 1 ≤ p ≤ ma+mδ ε or a− δ < ε at p = 0. let [x] denote the largest integer not exceeding x. then the condition on those positive values p which the function π(x,m, ε) can take on x(m,m, a, δ) can be written in the form 1 ≤ p ≤ min{m, [ m(a+ δ) ε ]}. advances in systems science and applications (2013) vol.13 no.1 63 the condition that π(x,m, ε) vanishes on x(m,m, a, δ) has the form a < δ + ε. thus, we have proved the following assertion. statement 3. let k ≤ m. 1). if 0 ≤ a < δ + ε, then u∗(n,m, a, k, τ, δ, ε) = 0. 2). if a > δ + ε, then u∗(n,m, a, k, τ, δ, ε) = max{k+1−m,0} k = {1/m,k=m 0,k m. the function ∑j+k−1 i=j χ({y|y ≥ ε}, xi) depends only on those components of the vector x whose numbers belong to {j, j+1, ..., j+ k − 1}. thus, we introduce the set x(k,m, a, δ) = {x ∈ ek +| | j+m−1∑ i=j xi − am| ≤ δm, j = 1, ..., k −m+ 1}. for this set, an assertion similar to statement 2 is valid. statement 4. if x ∈ en + and (xj , xj+1, ..., xj+k−1) ∈ x(k,m, a, δ), then there exists a vector y ∈ x(n,m, a, δ) such that yi = xi for i = j, j + 1, ..., j + k − 1. proof. we set • yi = xj+(i−j)mod m for i = 1, ..., j − 1, • yi = xi for i = j, j + 1, ..., j + k − 1, • yi = xj+k−m+(i−j−k+m)mod m for i = j + k, ..., n. we have • ∑l+m−1 i=l yi = ∑j+m−1 i=j xi for l = 1, ..., j − 1, • ∑l+m−1 i=l yi = ∑l+m−1 i=l xi for l = j, j + 1, ..., j + k −m, • ∑l+m−1 i=l yi = ∑j+k−1 i=j+k−m xi for l = j + k −m, ..., n−m+ 1. therefore, y ∈ x(n,m, a, δ). this completes the proof of the statement. now, the problems for u∗ and u∗ with k > m take the forms u∗(n,m, a, k, τ, δ, ε) = 1 k sup x∈x(k,m,a,δ) π(x, k, ε). and u∗(n,m, a, k, τ, δ, ε) = 1 k inf x∈x(k,m,a,δ) π(x, k, ε). let us divide k by m with a remainder, that is, write k = mq + r,m > r ≥ 0. take an arbitrary vector x ∈ x(k,m, a, δ) and consider its decomposition into the parts x0 = (x1, ..., xm) ∈ em and xj = (xr+(j−1)m+1, xr+(j−1)m+m) ∈ em, j = 1, ..., q. 64 d.a molodtsov: soft probability of large deviations it is assumed that r > 0; if r = 0, then the vector x0 is absent. obviously, π(x, k, ε) = π(x0, r, ε) + q∑ j=1 π(xj ,m, ε) ≤ min{π(x0,m, ε), r}+ q∑ j=1 π(xj ,m, ε). note that the condition x ∈ x(k,m, a, δ) implies that xj ∈ x(m,m, a, δ) for any j = 0, ..., q. hence, we have sup x∈x(k,m,a,δ) π(x, k, ε) ≤ min{ max y∈x(m,m,a,δ) π(y,m, ε), r}+ q max y∈x(m,m,a,δ) π(y,m, ε). let us show that this inequality is, in fact, an equality. we take a vector y∗ ∈ x(m,m, a, δ) at which max y∈x(m,m,a,δ) π(y,m, ε) is attained and place all components of this vector greater than or equal to ε at the first positions. let x∗ ∈ ek + be the vector formed by as many copies of y∗ written one after another as needed to achieve the required dimension. it is easy to see that x∗ ∈ x(k,m, a, δ) and sup x∈x(k,m,a,δ) π(x, k, ε) ≥ π(x∗, k, ε) = min{ max y∈x(m,m,a,δ) π(y,m, ε), r} + q max y∈x(m,m,a,δ) π(y,m, ε). thus, we have proved the following assertion. statement 5. suppose that n ≥ k > m and k = mq + r, m > r ≥ 0 . 1). ifm(a+δ) ≥ ε, then u∗(n,m, a, k, τ, δ, ε) = min{[m(a+δ) ε ],r}+q min{[m(a+δ) ε ],m} k . 2). if m(a+ δ) ≤ ε, then u∗(n,m, a, k, τ, δ, ε) = 0. it is easy to see that statement 5 is also valid for n ≥ m ≥ k > 0. now, consider the u∗ problem. for the arbitrary vector x ∈ x(k,m, a, δ) under consideration and its partition constructed above, we have π(x, k, ε) = π(x0, r, ε)+ q∑ j=1 π(xj ,m, ε) ≥ max{k−m+π(x0,m, ε), 0}+ q∑ j=1 π(xj ,m, ε). since xj ∈ x(m,m, a, δ) for any j = 0, ..., q, it follows that inf x∈x(k,m,a,δ) π(x, k, ε) ≥ max{r−m+ inf y∈x(m,m,a,δ) π(y,m, ε), 0}+q inf y∈x(m,m,a,δ) π(y,m, ε). as above, this inequality is, in fact, an equality. to show this, we take a vector y∗ ∈ x(m,m, a, δ) at which min y∈x(m,m,a,δ) π(y,m, ε) is attained and place all components of this vector which are greater than or equal to ε at the last positions. advances in systems science and applications (2013) vol.13 no.1 65 consider the vector x∗ ∈ em + consisting of copies of y∗ written one after another. it is easy to see that x∗ ∈ x(k,m, a, δ) and inf x∈x(k,m,a,δ) π(x, k, ε) ≤ π(x∗, k, ε) = max{r −m+ min y∈x(m,m,a,δ) π(y,m, ε), 0}+ q min y∈x(m,m,a,δ) π(y,m, ε). thus, we have proved the following assertion. statement 6. suppose that n ≥ k > m and k = mq + r,m > r ≥ 0 . 1). if 0 ≤ a < δ + ε, then u∗(n,m, a, k, τ, δ, ε) = 0. 2). if a ≥ δ + ε, then u∗(n,m, a, k, τ, δ, ε) = q k . 9 conclusion the exact boundaries of the range of the soft probability of large deviations under a single mean hypothesis, which were found in this paper, show (although, for a very simple example) that it is quite possible to deal with soft probabilities, in spite of the presence of parameters and the interval form of soft probability. the next goal is to solve more complicated problems on evaluating various probabilities and other characteristics in the presence of several hypotheses, preferably of different types. of special interest is the application of the ideas and results presented in this paper to a real statistical problem, which would make it possible to verify the effectiveness of the approach for real data. all readers interested in such a practical experiment are kindly requested to send their suggestions to the author at dmitri molodtsov@mail.ru. references [1] v. n. tutubalin. (1972), probability theory: a short course and scientificmethodological notes, izd. moskov. univ, moscow, russian. [2] d. a. molodtsov. (2004), theory of soft sets, urss, moscow, russian. [3] d. a. molodtsov. (2007), “portfolio control using soft probability”, vestn. nats. assots. uchastnikov fond. rynka, no.7-8. [4] d. a. molodtsov. (2007), “portfolio control using soft probability (short positions)”, vestn. nats. assots. uchastnikov fond. rynka, no.10. [5] d. a. molodtsov. (2011), “soft sets and prediction”, nechetkie sist. myagkie vychisl, vol.6, no.1. [6] d. a. molodtsov. (2010), “finite frequency stability and probability”, nechetkie sist. myagkie vychisl, vol.5, no.1. 66 d.a molodtsov: soft probability of large deviations [7] d. a. molodtsov. (2008), “soft uncertainty and probability”, nechetkie sist. myagkie vychisl, vol.3, no.1. [8] i. s. berezin and n. p. zhidkov. (1962), computing methods, pergamon press, vol.1/2. [9] h. aktas and n. cagman n. (2007), “soft sets and soft groups”, inf. sci, vol.177, pp.2726c2735. [10] m. i. ali et al. (2009), “on some new operations in soft set theory”, computers and math. with appl, vol.57, pp.1547-1553. [11] d. chen. ((2005)), “the parameterization reduction of soft sets and its applications”, comput. math. appl, vol.49, pp.757-763. [12] f. feng et al. (2008), “soft semirings”, computers and math. with appl, vol.56, pp.2621-2628. [13] f. feng et al. (2010), “soft sets combined with fuzzy sets and rough sets: a tentative approach”, soft computing, vol.14, pp.899-911. [14] y. b. jun. (2008), “soft bck/bci-algebras”, computers and math. with appl, vol.56, pp.1408-1413. [15] y. b. jun and c. h. park. (2008), “applications of soft sets in ideal theory of bck/bci-algebras”, inf. sci, vol.178, pp.2466-2475. [16] z. kong et al. (2008), “the normal parameter reduction of soft sets and its algorithm”, j.comp. appl. math, vol.56, pp.3029-3037. [17] p. k. maji et al. (2002), “an application of soft sets in a decision making problem”, comput. math. appl, vol.44, pp.1077-1083. [18] p. k. maji et al. (2003), “soft set theory”, comput. math. appl, vol.45, pp.555-562. [19] p. majumdar and s. k. samanta. (2010), “on soft mappings”, computers and math. with appl, vol.60, pp.2666-2672. [20] d. molodtsov. (1999), “soft set theory first results”, comput. math. appl, vol.37, pp.19-31. [21] d. pie and d. miao. (2005), “from soft sets to information systems, granular computing”, ieee iinter. conf, no.2, pp.617-621. advances in systems science and applications (2013) vol.13 no.1 67 [22] m. shabir et al. (2009), “soft ideals and generalized fuzzy ideals in semigroups”, new math. nat. comput, no.5, pp.599-615. [23] m. shabir et al. (2011), “on soft topological spaces”, computers and math. with appl, vol.61, pp.1786-1799. [24] zou yan et al. (2008), “data analysis approaches of soft sets under incomplete information”, knowl.-based syst, vol.21, pp.941-945. corresponding author d.a molodtsov can be contacted at: dmitri molodtsov@mail.ru advances in systems science and applications (2013) vol.13 no.4 299-316 currency wars and a possible self-defense (ii) : a plan of self protection jeffrey forrest1, zack hopkins2 and sifeng liu3 1department of mathematics, slippery rock university, slippery rock, pa 16127, usa 211934 route 89, wattsburg, pa 16442, usa 3school of economics and management, nanjing university of aeronautics and astronautics, nanjing 210016, pr china abstract continuing what is presented in part one (forrest,hopkinsand liu, 2013)[1], and based on how a currency war could be potentially raged against a nation, this sequel makes use of the results of feedback systems to develop a self-defense mechanism that could conceivably protect the nation under siege. keywords purchasing power, supply and demand of money, feedback system, economic sector, monetary policy 1 introduction when combining the previous theoretical analysis in (forrest and hopkins, 2013)[1] with the recent cases of speculative attacks in the arena of international finance, we surely see the following predicament. when a nation tries to develop economically, due to its loosening economic and monetary policies, large amounts of foreign investments would be welcomed ; and at the same time, a lot of such foreign investments would strategically rush into the nation in order to ride along with the forthcoming economic boom. now what we have shown earlier is that if a large amount of foreign investments leaves suddenly, then the nation would most likely suffer from a burst of the economic bubble with a large percentage of economic activities interrupted either temporarily or indefinitely. so, a natural question at this junction is : how could we possibly design a measure to counter such sudden leaves of foreign investments in order to avoid the undesirable disastrous consequences ? we will address this problem in the rest of this paper. 2 a model for categorized purchasing power let us look at the following model that relates the purchasing power of money with the demand and supply of the money of a national economy : dp dt = k(d − s) (1) where d stands for the demand for money,sthe money supply, p the purchasing power of money, k > 0 is a constant,and t represents time. 300 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection what this model says is that the rate of change in purchasing power is directly proportional to the difference between the demand and supply of money. in particular, the model says that with all other variables staying constant, if the money supply s increases by a large amount and satisfies s > d, then the purchasing power of the money p decreases so that more money is needed to buy essentials of living and inflation will increase due to the increase in the money supply s. now, let us divide the overall national economy into three sectors e1, e2 and e3 as follows : e1 stands for the goods, services, and relevant production of these goods and services that are needed for maintaining the basic living standard, e2 he goods, services, and relevant productions that are used to acquire desired living conditions, and e2 the goods, services, and relevant productions that are used for the enjoyment of luxurious living conditions. the reason why we divide the economy in such a way is that according to (allen and goldsmith, 1972)[2],the following are four base requirements that allow for a stable society to be achieved and maintained : (1). minimum disruption of ecological processes ; (2). maximum conservation of materials and energy, or an economy of stock rather than flow ; (3). a population in which recruitment equals loss ; (4). a social system in which individuals can enjoy rather than feel restricted by the first three conditions. so, in terms of economics, to maintain a stable society, a relative stability of the economic sector e1 has to be achieved first. next, let us accordingly divide the overall demand d of money into three corresponding categories as follows : d1 = the demand of money for meeting the minimum requirement to maintain the basic living standard ; d2 = the demand of money for acquiring desired living conditions ; d3 = he demand of money for enjoying luxurious living conditions. assume that in a stable economy, we have the following allocation of the money demand : d = d1 +d2 +d3 = α1d + α2d + α3d (2) where the weights αi, i=1,2,3, stands for the average allocation of the citizens of the economy over the three categories as described above, satisfying α1+α2+α3 = 1, and di = αid, i=1,2,3. for instance, in the stable economy, an average family allocates half of its monthly income on necessities of living, such as food, utilities, etc., 4 tenth of the income on acquiring the desired quality of life, and one tenth of the income on luxurious items, then α1 = 0.5, α2 = 0.4 and α3 = 0.1. if the money supply s increases drastically, along with the decreasing purchasing power of money, all goods will cost more. if somehow the goods in the category advances in systems science and applications (2013) vol.13 no.4 301 of living necessities rise more rapidly, then a re-allocation of household income will appear. for instance, due to rumors about potential interruptions in the supply of food and clean water accompanying a substantial increase in the money supply, the average family has to reallocate its income as follows :α1 = 0.625, α2 = 0.3 and α3 = 0.075. when such a reallocation of income of the average family is forced to take place, the stability of the economy would very likely be in trouble. so, to stabilize the economy, the purchasing power of money in category d1 should stay relatively constant, while in d2 increases somehow slightly, and in d3 it should be allowed to increase in order to attract and trap the additional money supply away from category d1. so, let us assume p1 = the purchasing power of money in category d1 ; p2 = he purchasing power of money in category d2 ; p3 = the purchasing power of money in category d3. similarly, let us define s1 = he money supply that goes into category d1 ; s2 = the money supply that goes into category d2 ; s3 = he money supply that goes into category d3. so, equ.(1) would look as follows : dp1 dt = k11(d1 − s1) + k12(d2 − s2) + k13(d3 − s3) + ∑n j=1 q1jxj dp2 dt = k21(d1 − s1) + k22(d2 − s2) + k23(d3 − s3) + ∑n j=1 q2jxj dp3 dt = k31(d1 − s1) + k32(d2 − s2) + k33(d3 − s3) + ∑n j=1 q3jxj (3) where (k)(ij), and (q)(ij) are constants, and (x)(j) stands for monetary policies, i=1,2,3, and j=1,2, . . . ,n,by using matrix notations, equ.(3) can be rewritten as follows : ṗ = kz +qx (4) where p = [p1 p2 p3] t , ṗ is newton’s original notation for derivatives such that ṗ = [dp1dt dp2 dt dp3 dt ] t ,k = [kij ]3×3 the coefficient matrix of the variables (di − si), i=1,2,3, q= [qij ]3×n the coefficient matrix of the variables xj , j=1,2, . . . ,n. in terms of systems research, the mode in equ.(2) can be seen as a feedback system as depicted in fig.1, where s represents the initial state of the economy. after the monetary policies x1, x2, . . . , xn are introduced, the participants of the economy introduce either consciously or unconsciously a feedback component system sf so that the overall system with the added feedback produces the desired output p1,p2 and p3. what is shown in fig.1 is the fact that each and every market economy, where the participants are allowed to design their own methods (without violating the 302 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection fig.1 the feedback loop between monetary policies and categorized purchasing powers established laws) to achieve their individually defined financial successes, then the economy constitutes a “rotational” field. here, the word of rotation means that as soon as a monetary policy is introduced, the market participants will find ways to take advantage of the policy so that the policy and the individually designed methods jointly produce the individually desired outputs. 3 the functional relationship between p and (d − s) in this section, we investigate the relationship between the purchasing power vector [p1 p2 p3] t of money and the difference vector [d1 − s1 d2 − s2 d3 − s3] t of demand and supply of money. in particular, we discuss why the trends found in purchasing power of money could righteously be described by using linear models, although the purchasing power is clearly not linear in terms of the difference of the demand and the supply of money. to this end, let usintroduce the following three models, the first oneis linear, the second one quadratic, and the third one cubic, to demonstrate the effects of supply and demand of money on purchasing power of money in each case : p (t) = a(d(t)− s(t)) + ε (5) p (t) = a(d(t)− s(t))2 + b(d(t)− s(t)) + ε (6) and p (t) = a(d(t)− s(t))3 + b(d(t)− s(t))2 + c(d(t)− s(t)) + ε (7) where p (t) stands for the purchasing power of money, d(t) the demand of money, s(t) the supply of money,ε a random variable with mean c ̸= 0,and a, b and c are constant. more specifically, the random variable is the error term in the sense that it compensates for any unpredicted event or factor that could impact the purchasing power and that is not taken into account in the model. when the nonlinearity in the trend of purchasing power is considered, such as advances in systems science and applications (2013) vol.13 no.4 303 in the case of japanese yen (fig.2), the quadratic and cubic models in equs.(6) and (7) seem to model adequate. in particular, the non-linear graph of the purchasing power of japanese yen has two distinguishable patterns : one is parabolic and the other cubic. these two types of general patterns are respectively depicted by the nonlinear models in equs.(6) and (7). fig.3 depicts both the purchasing power and the amount of money in circufig.2 urchasing power and currency in circulation, from www.dollardaze.com as accessed on april 4, 2012 lation of the u.s. dollars over time. here a clear inverse relationship between the currency in circulation, supply of money, and the purchasing power of money can be seen. that is to say, as the amount of money in circulation increases the purchasing power of the us dollars decreases. the graph of purchasing power is fairly linear, especially if broken up into two segments : from 1971 to approximately 1981 and from 1981 until the present day. this trend in the purchasing power of u.s. dollars suggests that the linear model in equ.(5) seems to adequately reflect what happened to the purchasing power over time. in particular, if we employ the linear model in equ.(5) to describe the relationship between the purchasing power of the u.s. dollars and the difference between the demand and supply of the money with respect to time, then this model well explains how the u.s. dollar declines in purchasing power. for instance, the initial high purchasing power was due to the fact that the currency in circulation was fairly low and the value of the u.s. dollar was fixed at $35 an ounce of the gold. however, starting in may 1971 the u.s. dollar suffered from a major crisis, and consequently began to depreciate against the gold from the initial $ 35 an ounce to $38 an ounce on august 15, 1971, then to $42.22 an ounce at the start of 1973, and then to as high as $96 an ounce in march of the same year (wang and hu, 304 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection 2005, p.12)[3]. in 1976 a new international agreement was reached in the capital city kingston of jamaica ; with several rounds of modifications the international financial system entered the jamaica system in 1978. within this new system, gold is no longer considered as a form of money. that is, when the demand for the u.s. dollars is brought into the equation, it can justify why there are more severe or gradual drops in purchasing power as the amount of money in circulation increases. for example, if the change in the supply of money is equal to the increase in demand of money, then the purchasing power of the money should remain constant. to cause the severe drop in purchasing power as seen from 1971 until approximately 1981, there was an increase in the money supply accompanied by a decrease in the demand for money. this decrease in the demand for money was caused by the transition from the u.s. dollar being backed by gold to that not having gold backing (dollardaze.org). this also makes sense intuitively. if we already know that a rise in the money supply decreases the purchasing power of money, then people wanting the money less would further exacerbate that decrease in purchasing power. from that point on, the demand of money must have increased but still not greater than the supply of money because the decline in the purchasing power flattens out instead of becoming more severely negative. as the graph of the stock prices of the dow jones industrial average indicates, fig.4, the stock market began to rise in the early 1980s and had a sharp incline until the start of the 21st century. this rise is an indicator that supports the claim that the demand for money increased during this time fig.3 the feedback loop between monetary policies and categorized purchasing powers period and, in conjunction with the supply of money, influenced the purchasing power to decrease less severely than during the time period from 1971 to 1981. advances in systems science and applications (2013) vol.13 no.4 305 fig.4 historical data of djia (from yahoo finance, accessed on april 17, 2012) looking at the same graph created for the japanese economy, fig.2, we see a similar but different story. by breaking the graph up into more intervals, fig.5 we fig.5 the fluctuation in the purchasing power of japanese ye are able to form various linear patterns. from 1971 until around 1986 there was a general decline in the purchasing power of japanese yen. then once it came to a peak again in 1996 there again was an overall decline in the purchasing power. these are clearly not linear graphs ; but looking at the general motion of the purchasing power in the graph they are fairly linear in two segments, like the situation with the u.s. dollar. therefore, it is also reasonable for us to use the linear model in equ.(5) to explain the evolutionary trends in the purchasing power of japanese yen. under the assumption that the supply and purchasing power of money are inversely related when all other factors remain constant, we can see that the purchasing power is a function of the demand and supply with respect to time. by using the linear model in equ.(5) and by breaking up the graph in fig.5 into two 306 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection fig.6 the money supply and the nikkei 225 stock market (okina, shirakawa and shiratsuka, 2001)[4] intervals, we see a similar relationship that was discussed with the supply and demand of money in the u.s. dollar. if both the supply and demand increase at the same rate, then the purchasing power will remain constant. if the supply increases with the demand increasing at a faster rate than before but still less than the increase in he supply, then a less severe decline in the purchasing power would be observed. however, if the demand increases at a rate greater than that of the supply of money, then this would lead to an increase in the purchasing power. furthermore, an increase in the supply matched with a decreasing rate in demand would result in a more severe decrease in the purchasing power. these explanations help explain varying linear slopes that are reflected in the purchasing power of money in the japanese economy. for example, notice that for japan 1986 was the first year of the asset pricing bubble. in the graph of the money supply and the nikkei 225 stock market,fig.6, for details,see (okina, shirakawaand shiratsuka, 2001)[4], during this time there was a small increase in the money supply while the nikkei 225 increased drastically. with the stock market increasing during this time, the demand for money also increased because people wanted to take advantage of the market rise. as we mentioned earlier, a relatively low increase in the money supply accompanied by a large increase in the demand for money would result in an increase in the purchasing power. this explanation corresponds with the depiction of the purchasing power of japanese yen until 1990 when it burst reflecting the sharp decline in the purchasing power as reflected in the graph and by our model. advances in systems science and applications (2013) vol.13 no.4 307 the sharp rise in the purchasing power that appeared at around the year of 1973, fig.3, could have been due to the oil crisis which caused a shift in japanese economy toward huge investments in the electronic industries. from (hutchison, 1986)[5], we can see that japan’s money supply was fairly constant in the year of 1973 but the demand of money from the oil crisis increased drastically. that made the purchasing power of the money to increase drastically during the early 1970s ; this conclusion is further validated by the discussions of(okina, shirakawa and shiratsuka, 2001)[4]. fig.7 is a graph that reflects the japanese inflation rate over the time period from april 1971 until april 2012, where we can notice the continuing patterns and similarities with those of the purchasing power. fig.7 japanese inflation rate from april 1971 to april 2012. the original source was accessed on april 4, 2012 as a matter of fact, the similarities between the graphic patterns of the purchasing power of money and the inflation rate, as described above, are also seen in the current movements of oil prices. on the ed show of the msnbcat 11 :00 p.m., barny frank referenced certain people who bought crude oil to drive up the prices only to sell it at a later time for a windfall of profits. the market system unconsciously allows the oil prices to riseand to consequently manipulate the gasoline prices, fig.8. that situation works in conjunction with the demand and supply of money. summarizing what is discussed above we conclude that it is theoretically reasonable for us to analyze the relationship between the purchasing power of money and the difference of the demand and supply of money by using the linear model in equ.(5), where the random variable accounts for all the unexpected factors that are not included in the model. by combining what is obtained in section3 with the linear model in equ.(5), we have the following relationship between categorized purchasing power [p1 p2 p3] t of money and categorized difference [d1 − s1 d2 − s2 d3 − s3] t of demand and 308 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection supply of money : p1 p2 p3  = r3×3  d1(t)− s1(t) d2(t)− s2(t) d3(t)− s3(t)  +  ε1 ε2 ε3  (8) where r3 × 3 is a constant square matrix with real entries, and [ε1 ε2 ε3] t a random vector with a none zero mean. fig.8 the prices of crude and gasoline move in concert, the original source was accessed on april 17, 2012 4 separating economic categories using feedback component systems if we consider the mathematical expectations of the variables in equ.(8), we have p1 p2 p3  = a3×3  d1(t)− s1(t) d2(t)− s2(t) d3(t)− s3(t)  +  c1 c2 c3  (9) whereeεi = ci ̸= 0, i = 1, 2, 3. by substituting equ.(9) into equ.(4), we have r3×3ż = kz +qx (10) without loss of generality, we assume that r3×3 is invertible. that is, we assume in general the categorized purchasing power of money is completely determined by the categorized differences of demand and supply of money. then, equ.(10) can be rewritten as follows : ż = az +bx (11) advances in systems science and applications (2013) vol.13 no.4 309 wherea = r−1k and b = r−1q. to make the model in equ.(11) technically manageable, we assume without loss of generality that b is a 3×3 matrix, meaning that the monetary policies x1, x2, . . . , xn are accordingly categorized into three groups : x1=the set of all those monetary policies that deal with the population meeting the minimum need to maintain the basic living standard ; x2=the set of all those monetary policies that deal with the populationąŕs need for acquiring desired living conditions ; x3=the set of all those monetary policies that deal with the populationąŕs need for enjoying luxurious living conditions. without loss of generality, we will still use the same symbol x to represent the vector [x1 x2 x3] t of categorized monetary policies. similar to the concept of consumer price index (cpi), let us introduce an economic index vector y = [y1 y2 y3] t such that yi measures the state of the economic sector i, i = 1, 2, 3. then from equ.(11) we can establish the following model for the national economy of our concern : ż = az +bx y = cz +dx z(0) = 0 (12) where z is the 3 × 1 matrix [d1 − s1 d2 − s2 d3 − s3] t of the categorized difference of demand and supply of money, referred to as the state of the economic system, a,b,c, and d are respectively constant 3 × 3 matrices, such that d is non-singular (meaning that each introduction of monetary policies does have direct, either positive or negative, effect on the performance of the economy), and the input space x and output space y are following : x = y = {r : [0,+∞) → r3 : r is a piecewise continuous function} (13) where r stands for the set of all real numbers and r3 the nth dimensional euclidean space. what is described by equ.(12) is that the state of the national economy is representable through the use of the state variable z that helps the economy to absorb the positive and negative effects of the monetary policies x1, x2 and x3.then, both the internal mechanism z of the economy and the monetary policies x jointly have a direct effect on the overall performance y of the economy. the condition, as imposed on the input space x and the output space y means that monetary policies, which form the input space x, are introduced based on the effects of the previously implemented policies, while the overall performance, indexes of which constitutes the output space y , of the economy evolves from previous 310 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection states mostly continuous. such assumptions historically speaking are not always true, while failures occur rarely in terms of frequencies. for instance, in modern china, the so-called cultural revolution occurred abruptly (macfarquhar and schoenhals, 2008)[6] ; and several times in the history of russia, the tsars had introduced rushed social reforms creating a torn country (huntington, 1996)[7]. geometrically, the systemic model in equ(12) of the national economy is depicted in fig.9. the monetary policy input x splits into two portions f1f2f3f4 and f5f6f7f8.the portion labeled f1f2f3f4 directly affects the performance y of the economy. the other portion, labeled f5f6f7f8,is fed into the initial state of the economy s, leading to the introduction of the feedback component systemsf ,which stands for the market reactions to the introduced monetary policies, and the formation of a feedback loop within the economy. this feedback loop in fact constitutes the main body of the economy, while the overall performance vector y of indexes is merely an artificially designed measure. what needs to be noted is that within each of our three economic sectors e1, e2 and e3,the market reactions, which constitute parts of the feedback component system sf ,to monetary policies in general are unique and economic sector specific. that is, the market reactions in one economic sector are different from those of another sector. the inner most loop f1f2f3f4 is the second stage of the economy after the introduction of new monetary policies. they by the joint effect of the marfig.9 the geometry of the systemic model of the national economy ket reactions the introduction of new monetary policies. they by the joint effect of the market reactions (sf ) and the non-reactionary aspects (f5f6f7f8) of the monetary policies, the final numerical readings y of the economy are produced. to see how monetary policy input x could have aspects, one is reactionary and the other non-reactionary, let us assume that x stands for such a monetary policy that allows the inflation to inch higher. therefore, the price of crude oil is expected to rise accordingly. now, the reactionary aspects of the policy x make advances in systems science and applications (2013) vol.13 no.4 311 the gradual and calculated increase in the price of crude oil more or less random. on the other hand, the expected outcome of the policy is the theoretical, non-reactionary aspect of the policy. when the actual outcome deviates from the expectation, the difference is caused by the market reactions to the policy. according to (lin, 1994)[8], the 3-dimensional system in equ.(12), meaning that both the input x and the output y are elements from r3, can be decoupled into three independent systems of the same kind with one-dimensional input and out. specifically, if we let s = {(x, y) ∈ x × y : ∃z ∈ zsuch that x, y, z satisfy equ.(12)} = the system of all the ordered pairs (x,y) satisfying equ.(12) where the state space z = {r : [0,+∞) → r3 : r is a piecewise continuous functi -on}, and for each i = 1, 2, 3, define a system si as follows : ż = az +bixi y = cizi +dixi z(0) = 0 (14) where bi is the ith column of b,ci the ith raw of c,di a non-zero constant, and the input space xiand the output space yi are given as follows : xi = yi = {r : [0,+∞) → r3 : r is a piecewise continuous function} in particular, the system s can be decoupled through feedback into the factor systems si, i = 1, 2, 3, as follows. let α = [ a 0 0 0 a 0 0 0 a ] β = [ b1 0 0 0 b2 0 0 0 b3 ] γ = [ c1 0 0 0 c2 0 0 0 c3 ] δ = [ d1 0 0 0 d2 0 0 0 d3 ] then, the cartesian product system sd = s1s2s3 is represented by the set of all ordered pairs (x, y) satisfying  ż = αz + βx y = γz + δx z(0) = 0 (15) because δ is non-singular, the inverse system sd −1 is obtained as follows : sd −1 = {(y, x) : (y, x) satisfies equ.(16)} 312 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection where equ.(16) is given as follows : ż = (α− βδ−1γ)z + βδ−1y x = δ−1γz + δ−1y z(0) = 0 (16) the particular feedback component system sf : y → x used in this decoupling is given as follows : [ ż ż′ ] = [ a−bd−1c 0 0 α− βδ−1γ ] [ z z ′ ] + [ bd−1 βδ−1 ] y x = [−d−1c δ−1γ] [ z z ′ ] + (d−1 − δ−1)y[ z z ′ ] (0) = 0 (17) in terms of economics, what the concept of decoupling the 3-dimensional system s into component systems si,i = 1, 2, 3,as discussed above, implies is that when monetary policies are established individually and respectively for each of the three economic sectors e1, e2, and e3, although these economic-sector specific policies will most definitely have joint effects on the economy, there is at least one way to design a feedback component system sf so that the overall feedback system f (s, sf ), which represents the whole system as depicted in fig.1, can be controlled through adjusting individually each of the economic sectors e1, e2and e3. 5 a strategy for national defense the significance of the previous discussion is that we can now propose based on the sound analytical reasoning presented above a strategy for national defense against currency warfare in case that initial‘friendly’ foreign investments turn out to be an aggressive act of war by suddenly withdrawing all or significant amount of the investments. in particular, let us assume that at time period t, the total money supply of the specific nation of our concern is $1,000, $100 of which is from foreign investments. that is, the foreign investments amount to 10% of the total domestic money supply. with the additional money supply (due to the foreign investments) the speed of money circulation increases, signaled by increased spending and elevated levels of economic activities. if these foreign investments leave the nation suddenly, then it is reasonable to expect that more than 10% of the economic activities from around the nation will be more or less affected adversely due to the sudden exhaustion of the money flow, because accompanying the foreign investments there advances in systems science and applications (2013) vol.13 no.4 313 tends to be domestic investments attached too, making the total investments on the related economic activities more than 10% of the national economy. now, if before the foreign investments depart while leaving behind disastrous aftermath, the national government has been keeping the exchange rate (to this end not all nations from around the world are currently able to do this successfully) the same while gradually and strategically increased its money supply, say, to the level of $10,000, then in the entire monetary circulation around the nation the proportion of the foreign investments shrinks to about 1% from the original 10%. and if the money supply had been increased to the level of $1,000,000, then the proportion of the foreign investments would have shrunk from the original 10% to about 0.01%, which is nearly zero. so, if at this moment of time the foreign investments are suddenly withdrawn as an aggressive act of war, only around 1% or 0.01% of the consumption and economic activities of the nation will be affected adversely and the overall economic health of the nation will be relatively stable. on the other hand, along with the drastically increased money supply, all prices in the nation will most certainly go through the roof, placing a large portion of the national population in financial crises due to the run-away inflation. for instance, a certain commodity a, which is piece of living necessity, used to cost $1.00 a unit ; now it requires $10.00 or even $1,000 to purchase. such dramatic increase in prices will surely cause hardships for a good number of citizens of the nation. to this end, the national government needs to work on how to redistribute the additional money supply. in fact, as long as the distribution of the increased money supply does not cause social upheaval, then to this nation, nothing disastrous will really happen and the potentially damaging impacts of sudden departure of the foreign investments will be under control, too. however in reality, the distribution of the extra money supply could easily lead to major societal instabilities to the nation due to the increased unevenness in the economic scene : the rich become richer while the poor become poorer. the situation here is similar to that of the normal inflation, which has been employed to make the economic structure more uneven than before so that the yoyo structure of the economy spins with more strength. in other words, with additional money supply injected into the economy, the redistribution of the wealth that is represented by the increased money supply is surely uneven, where some people receive more than their share, some simply keep pace with the decrease of the purchasing power of their income, while others fall behind or further behind with their financial status. so, if our suggested measure is adopted to counter the damaging effects of sudden departure of foreign investments, the nation needs to develop a practical plan to distribute the extra money supply in order to : (1). keep societal peace and national stability so that a normal and operational economy can be maintained. 314 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection (2). increase the nation’s economic prosperity by taking advantage of the foreign investments even though they leave sooner or later either slowly or suddenly. specifically, as suggested by the systemic model in the previous section, to help protect the innocent citizens of the nation from suffering from the potential economic turmoil, the national government could purposely divide the economy into three sectors e1, e2, and e3 as described earlier to meet the following goals. in sector e1, which consists of living necessities, the sector performance, such as the sector specific cpi, evolves as normally as possible ; in sector e2,which consists of such goods, services, and relevant productions that are used by citizens to acquire desired living conditions, the sector performance index, say the particular cpi of e2, could outpace that of sector e1 by a large amount ; and critically, the national government needs to manage to trap most of the additional money supply in the economic sector e3, which consists of such goods, services, and relevant productions that are used by the citizens for their enjoyment of luxurious living. the previous discussion on the systemic model of the national economy indicates that by managing the market reactions correctly, that is the design of the feedback component system sf , these three economic sectors can be well separated from each other. when the economic sector e1 evolves normally as expected based on the history of the nation, the citizens would not need to worry about their basic living and survival. that naturally leads to the desired societal stability and peace. although the prices in the economic sector e2 increase drastically when compared to those in e1, that in general should not affect the mood or the happiness of the population much, because desired living conditions change from time to time and vary from one family to another. in other words, isolated, temporary personally desires,which are not yet satisfied and which are generally inconsistent with each other and even contradictory to each other, could not amalgamate into a political force to cause turmoil. now, the key is how to keep most of the additional money supply in the economic sector e3.to address this problem, let us pay attention to our earlier assumption that a large amount of foreign investments entered the nation. that means commercial products and services that are new to the people in the specific nation are expected to and will become available soon. because these products and services are brand new to the region, high prices can of course be charged for these commercial goods and services. the situation is similar to the scenarios with developed nations, where the focus has been placed on innovations, design and production of unprecedented products and services. it is these never-seen products and services that generate the most profits. at this junction, let us address the problem of how to classify a product or a service into one of the three economic sectors e1, e2, and e3. first of all, goods and services that should be classified into e1 are clear.they represent such commodiadvances in systems science and applications (2013) vol.13 no.4 315 ties that directly relate to the survival of any human being, such as foods, utilities, and shelters of the basic quality. now, if a family lives in a rented two-bedroom apartment and wants to own a house, then houses to that particular family will be in the economic sector e2.f the people in a community go to work, shopping, and entertainment mostly on foot, then individually owned transportation tools, such as bicycles, motorcycles, cars, private jets, etc., will belong to the economic sector e3.that is, as time evolves forward, what belongs to the economic sector e3 gradually lowers itself into the economic sector e2, and what used to be in e2 also gradually moves to e1.that is, the classification of one particular product or service is a function of time and space. just by looking around shopping centers, one can see vividly that such classification of commercial goods and services have been successfully done by merchants throughout the world without much trouble. now, let us consider the scenario that when the nation purposefully and strategically increases its money supply in order to prevent disastrous aftermaths that could be potentially left behind by sudden withdrawals of foreign investments, additional foreign investments could continue to pour in. in this case, if the nation could not pick up its corresponding speed of economic development, then it will be taken over by the run-way high rate of inflation. in this case, if the foreign investments depart strategically, then the nation will be in real major political, societal, and economic crises. to prevent such disastrous consequences, the continued inflow of foreign money has to been managed so that they cannot depart quickly. another practical scenario that is different of what has been discussed above is that at a high speed develops the nation itself into a super economic power such that there is no longer such a single amount of foreign investments that can amount to an influencing percentage in the nation’s domestic consumption and/or economic activities. in this case the nation will be safe in the face of any potentially sudden withdrawal of foreign investments. 6 a few final words similar to how many different methods are there to invest money in life, there should be that many different ways to launch a currency war. what is presented in this work of two parts simply considers only one possible way to launch a currency war and one possible way to protect oneself against the disastrous consequences of such an attack. in other words, more scientific efforts should be devoted to the studies of financial attacks (or wars) along the line as outlined in this work in the spirit of how conventional warfare is conducted, analyzed, and planned strategically. the purpose of doing so is not to launch financial attacks on the innocent people who simply work hard to make a living ; instead, one should always be prepared for the worst that might be imposed upon him by those who always look for 316 jeffrey forrest :currency wars and a possible self-defense (ii) : a plan of self protection ways to take advantage of others. in this regard, we hope that our presentation in this work of two parts will play the role of a brick that has been thrown out there to attract beautiful and practically meaningful gemstones. acknowledgements we like to use this opportunity to express our appreciation to all the participants of our weekly meetings. our specific thanks go to professor yirong ying of shanghai university and professor xiangdong li of jiangsu normal university of technology. their insightful comments and input helped to enrich this work greatly. références [1] forrest j., hopkins z. and liu s. f.(2013), “currency wars and apossible selfdefense (i) : how currency wars take place”, advances in systems science and application, vol.13, no.3, pp.197-216. [2] allen r., and goldsmith e. (1972), “towards the stable society : strategy for change.the ecologist archive”, the ecologist, located at http ://www. theecolo-gist.info/page32.html, accessed on may 22, 2012. [3] wang r. x., and hu g. h. (2005), international finance, wuhan, hubei : press of wuhan university of science and technology. [4] okina k., shirakawa m., and shiratsuka s. (2001), “the asset price bubble and monetary policy : japan’s experience in the late 1980s and the lessons”, monetary and economic studies (special edition), february, pp.395-450. [5] hutchison m. m. (1986), “japan’s ‘money focused’ monetary policy”, economic review(federal reserve bank of san francisco), summer, no.3, pp.3346. [6] macfarquhar r., and schoenhals m. (2008), mao’s last revolution. cambridge, ma : harvard university press. [7] huntington s. p. (1996), the clash of civilizations and the remaking of world order, new york : simon & schuster. [8] lin y. (1994), “feedback transformation and its application”, journal of systems engineering, vol.1, pp.32-38. corresponding author jeffrey forrest can be contacted at : e-mail : jeffrey.forrest@sru.edu. advances in systems science and application (2015) vol.15 no.4 299-315 variable precision rough set model based on set pair homeopathic similarity relationship yong liu1 and yi lin2 1school of business, jiangnan university, wuxi 214122, china. 2school of business,slippery rock university, slippery rock, u.s.a. abstract due to the complexity and uncertainty of the objective world, as well as the limitation of human understanding, it has been difficult for the classical rough set method to deal with incomplete information system with noise data, ambiguity and other informational uncertainties. in view of this, the thought and method of the set pair theory and variable precision rough set are used in this paper to construct a novel variable precision rough set model. after scrutinizing the advantages and limitations of the models previously established based on the set pair similarity relationship, we define the set pair homeopathic similarity relationship based on threshold α connection degree. then we employ this new concept to substitute the equivalence relationship of the variable precision rough set so that a novel variable precision rough set model based on the set pair homeopathic similarity relationship is established. upon developing the properties of this model, we construct an example to show the feasibility and effectiveness of our novel model. keywords incomplete information system; set pair analysis; variable precision rough set; equivalence relationship. 1 introduction as an useful mathematical tool to deal with information that involves inaccuracy, uncertainty and fuzziness, rough set theory was proposed by pawlak [1]. this theory has been widely applied in such areas as the knowledge discovery, data mining, decision analysis, pattern recognition, and other fields [2–4]. the classical rough set theory is developed to deal with complete information systems with discrete attribute, where all the attribute values are known, thus the upper and lower approximation sets of the object sets based on equivalence relations, which satisfy the reflexive, symmetric and transitive properties, are defined. however, due to the complexity of the real world, uncertainty and limitation of human knowledge, not all attribute values acquired are known. so it is difficult for the traditional rough set theory to deal with such information systems and to materialize the desired knowledge discovery and extraction of relevant rules in order to make rough set theory effectively deal with systems of incomplete information, many scholars have been involved in the relevant research, leading to some important results. in summary, there exist two main classes of relat300 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... ed works. on one hand, each incomplete information system is pre-processed in order to make up the missing data; then the given system is investigated as a system with complete information. on the other hand, the theoretical equivalence relation among the objects of the system of incomplete information is directly weakened so that some new kind of relation, such as tolerance relationship [5, 6], similarity relationship [7], dominance relationship [8], etc., is introduced. for the former class of research, due to the information loss caused by pre-processing of the information system, the related works will not be discussed in this paper. for the latter class of works, when all the unknown attribute values of the system of incomplete information are of the omission type, kryszkiewicz builds the tolerance relationship, which satisfies the properties of reflexivity and symmetry.[9] when the unknown attributes values of the incomplete information are of the lost type, stefanowski and tsoukias establish the similarity relationship, which satisfies the properties of reflexivity and transitivity.[10] in view of the characteristics of incomplete information systems, wang puts forward the limited tolerance relationship, which satisfies the properties of reflexivity and symmetry.[11] corresponding to the systems of incomplete information that contain preference information,greco, matarazzo and slowiski[12, 13],shao and zhang [14] respectively put forward rough set model based on tolerance dominance relationship, which satisfies the property of reflexivity. for systems with incomplete order information, yang, yang, chen, et al [15],xie, song, chen, et al [16] construct a rough set model based on the similar dominance relationship, which satisfies the property of reflexivity. they then analyze the shortcomings of the tolerance dominance relationship and the similar dominance relationship, leading to the establishment of a rough set model based on the limited tolerance dominance relationship, which satisfies the property of reflexivity. however, when these rough set models are used to classify the domain with many condition attributes, the definition of rough sets still seems too loose. in order to deal with this problem, the set pair analysis method is introduced into the study of incomplete information systems [17], and additional rough set models, which are based on set pair analysis, are constructed [18–26]. by carefully analyzing these models, it can be seen that most of them are based on a local aspect and consider the classification effect of some specific objects without an overall classification performance of all objects by taking into account of unified approach, relevant uncertainty, oppositions, etc. it has been difficult for each of these established models to deal with incomplete information systems with preference information. in view of the shortcomings of the existing literature on incomplete information systems and the fact that there exist noise data, fuzzy and grey information in the incomplete information of systems, in this paper we will employ the set pair analysis method, combined with the previous studies, to construct a novel advances in systems science and application (2015) vol.15 no.4 301 variable precision rough set model based on the set pair homeopathic similarity relationship. the organization of this paper is as follows. section 2 compares the advantages and disadvantages of the currently available set pair similarity relationships. section 3 introduces the concept of our set pair homeopathic similarity relationship. section 4 establishes the variable precision rough set model based on set pair homeopathic similarity relationship. a case study is given in section 5. then this presentation is concluded in section 6. 2 the set pair analysis as a useful tool of system analysis to characterize and to study a variety of certainty and uncertainty and their laws of transformation in systems, the methodology of set pair analysis is proposed by chinese scholar keqin zhao in 1989. its core idea is to holistically analyze and deal with the certainty relationship and the uncertainty relationship of the objective matter of concern. this methodology has been widely applied to the study of such areas as mathematics, physics, systems science, management science, decision-making science, theory of predictions, computer science, social science, artificial intelligence, and many others. it has shown theoretical significance and potential for success for applications[17]. definition 1[17].assume that two given setsa andb from a set pairh(a,b),and that under the particular background (w ) of the problem of concern, h have n attributes, where s attributes are shared by both of a and b ,p attributes are about the opposition of a and b,and f attributes are neither shared by a and b nor about the opposition of a and b.then the ratio s/n is called the identity degree of a and b under the background (w ) ; f/n the discrepancy degree of a and b under the background (w ) ;and p/n the contrary degree of a and b under the background (w ) ; and the connection degree of and under the background is defined as follows: µw (a,b) = s n + f n i+ p n j (1) the connection degree reflects the relationship between a and b under the background (w ),denoted as = a + bi + cj , 0 ≤ a, b, c ≤ 1 and a + b + c = 1 , where the parameters a, b, c stand for some relationship between the sets a and b under the background (w ) . definition 2[17].for the given connection degree = a + bi + cj ,the rate of the identity degree a and the contrary degree c under some specific background (w ) is referred to as the set pair situation, denoted as follows: s(h) = a c (2) 302 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... if c = 0 ,then s(h) → ∞ ,which is referred to as the infinite situation with respect to the specific background (w ) . 3 set pair homeopathic similarity relationship 3.1 incomplete information system definition 3. suppose that s = (u,a, v, f) is an information system, where u = {u1, u2, ..., un} is a finite and nonempty set, known as the universe, a = c ∪ d a finite and nonempty attribute set, c = {a1, a2, ..., am} the condition attribute set, d = {d1, d2, ...., dp} the decision attribute set, and v = ∪ vij the value range of u on a, where vij is the attribute value of ui on aj ∈ a, f : u × a → v is an information function. for ∀aj ∈ a,∀xi, xk ∈ u, f(xi, aj) ∈ vij ,∀b ∈ b,b ⊆ a ,the equivalence relationship of xi, xk with respect to the attritube b is defined as follows: r = {(xi, xk) ∈ u × u : f(xi, b) = f(xk, b)} (3) definition 4. suppose that s = (u,a, v, f) is an information system. if there is at least one object ui and one attribute aj ∈ a such that vij takes the null value, denoted by *, then information system s = (u,a,a, f) is called incomplete information system, denoted as s∗ = (u,a, v, f); otherwise, it is called a complete information system. 3.2 comparison of different kinds of set pair similarity relationships due to the incompleteness of information of an incomplete information system, it is difficult for the traditional method of rough sets based on equivalence relationship to solve problems involving such systems. in view of this fact, some scholars suggest to use the methodology of the set pair analysis to define the similarity relationship to investigate incomplete information systems. in short, there exist the following four categories of similarity relationships [18–26]. definition 5.in the incomplete information system s∗ = (u,a, v, f), the general similarity relationship is defined as follows: sim(x, y) = {(x, y) ∈ u × u | ∀a ∈ a, a(x) ≥ a(y)or a(x) = ∗or a(y) = ∗} (4) for ∀a ∈ a,∀x, y ∈ u,b ⊆ a, |b| = n, teh connection degree based on the incomplete information system is denoted as µb(x, y) = s n + f n i+ p n j where s = |{a ∈ b | a(x) = a(y) and a(x) ̸= ∗ and a(y) ̸= ∗}| stands for the number of the attributes that are clear and equal for x and y with respect to the attribute set b; while p = |{a ∈ b | a(x) ̸= a(y) and a(x) ̸= ∗ and a(y) ̸= ∗}| stands for the number of the attributes that are clear but not equal for x and y advances in systems science and application (2015) vol.15 no.4 303 with respect to the attributes set b; f = |{a ∈ b | a(x) ̸= ∗ or a(y) ̸= ∗}| stands for the number of the attributes that are not clear and equal for x and y with respect to the attributes set b. write µb(x, y) = s n + f n i+ p nj = a+ bi+ cj where a = s n , b = f n , and c = p n (1)set pair similarity relationship and similarity classes based on the connection degree definition 6.[18]. for the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a,the set pair similarity relationship of x and y is defined as follows: smh(b) = {(x, y) ∈ u × u | ub(x, y) = a+ bi+ cj} (5) definition 7.[18].for the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a,the set pair similarity classes of x and y with respect to the attribute set b are defined as follows: smh(b) = {y ∈ u | ub(x, y) = a+ bi+ cj} (6) obviously, this set pair similarity relationship satisfies the property of symmetry, but does not satisfy the properties of reflexivity and transitivity. in particular, in practical applications, it can be found that when the information system of concern contains a large amount of missing values, the result of model analysis tends to be not satisfactory, and the operational properties of the lower and upper approximations of rough sets are imperfect. (2) set pair similarity relationship and similarity classes based on a threshold value α definition 8. [19].for the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0 ≤ α ≤ 1,the set pair similarity relationship of x and y based on the threshold value α is defined as follows: smhα(b) = {(x, y) ∈ u × u | ub(x, y) = a+ bi+ cj, a+ b ≥ α} (7) definition 9. in the incomplete information system s∗ = (u,a, v, f),for ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0 ≤ α ≤ 1,the set pair similarity classes of x and y with respect to the attribuye set b based on the threshold value α are defined as follows: smhα b(b) = {y ∈ u | ub(x, y) = a+ bi+ cj, a+ b ≥ α} (8) evidently, this set pair similarity relationship satisfies the property of symmetry, but not that of transitivity. the rough set model based on the set pair similarity relationship only considers and limits the known identity degree and the proportion of the uncertain 304 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... attributes among the objects. it ignores the influence of the discrepancy degree on the similarity, which will affect the performance of classification. in addition, identifying the influence of the identity degree and that of the discrepancy degree on the similarity will also to a certain extent increase the classification error. for the rough set based on the set pair similarity relationship, there exist two types of errors. on one hand, it ignores the difference between uncertain attributes and plays too much emphasis on their similarity. on the other hand, it ignores the interfering effect of the discrepancy degree on the decision-making system in which requirements of relatively higher accuracy are imposed, leading to major differences in the decision-making rules. (3) set pair similarity relationship and similarity classes based on identity and contrary degrees (α, λ) definition 10. [22].in the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0.5 ≤ α ≤ 1, 0 ≤ λ ≤ 0.5,the set pair similarity relationship of x and y based on identity degree α and contrary degree λ is defined as follows: smhα(b) = {(x, y) ∈ u × u | ub(x, y) = a+ bi+ cj, a ≥ α, c ≤ λ} (9) obviously, this set pair similarity relationship satisfies the properties of reflexivity and symmetry, but not that of transitivity. definition 11.in the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0.5 ≤ α ≤ 1, 0 ≤ λ ≤ 0.5,the set pair similarity relationship of x and y with respect to the attribute set b based on identity degree α and contrary degree λ are defined as follows: smhα,λ b = {y ∈ u | ub(x, y) = a+bi+cj, a ≥ α, c ≤ λ, 0.5 ≤ α ≤ 1, 0 ≤ λ ≤ 0.5} (10) (4) set pair similarity relationship and similarity classes based on threshold values (α, λ) definition 12[23].in the incomplete information system s∗ = (u,a, v, f), ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0 ≤ α ≤ 1, 0 ≤ λ ≤ 1,the set pair similarity relationship of x and y based on the threshold values (α, λ) is defined as follows: smhα(b) = {(x, y) ∈ u × u | ub(x, y) = a+ bi+ cj, a+ λb− c ≥ α} (11) definition 13[23].in the incomplete information system s∗ = (u,a, v, f), for∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0 ≤ α ≤ 1, 0 ≤ λ ≤ 1,the set pair similarity classes of x and y with respect to the attribute set b based on the threshold values (α, λ) are defined as follows: smhα b(x) = {y ∈ u | ub(x, y) = a+ bi+ cj, a+ bλ− c ≥ α} (12) advances in systems science and application (2015) vol.15 no.4 305 the concept of set pair similarity classes reflects the object set that the similarity degree of x with respect to the attribute set b is greater or equal to(α, λ) . in view of the existing two kinds of errors the above set pair similarity relationship suffers from, tao, dai and zhang put forward the rough set model based on the threshold value (α, λ) .[23] here the threshold value α is introduced into the model by separating the contributions of both uncertain and certainty attributes to the relationship of similarity and by paying much more attention to the disturbance of the contrary degree. the relationship of similarities between things is considered and characterized in three aspects: the identity degree, discrepancy degree and contrary degree. due to the fact that the discrepancy degree possesses both positive and negative incentives on the connection degree, while the threshold value satisfies 0 ≤ λ ≤ 1 , the model only considers the positive incentives of the uncertain attributes on the similarity while neglecting their punitive effects. at the same time, because of the introduction of the threshold value λ , the subjectivity of the model is increased. additionally, the model does not satisfy the property of reflexive, which affects the quality of the operation of the model. 3.3 set pair homeopathic similarity relationship and similarity classes in view of the existing problems in connection degrees, we will use the homeopathic value method to make improvement on the effect of connection degrees. the idea of the homeopathic value method is that along with the identity degree, the value is determined while keeping the same potential class status. for example, for the connection degree µ = ai+ bj+ ck , the homeopathic value method is used to acquire the novel connection degree µ = a(1+ b)i+ b2j+ c(1+ b)k. from comparing the two connection degrees, we can see that by using the homeopathic value method the uncertainty of the original connection degree µ = ai+bj+ck is divided into three parts in the novel connection degree ab. one part is classified into the identity degree b2 ; one part into the discrepancy degree cb; and the third part into the contrary degree . at the same time, the set pair potentials of these two connection degrees are the same as a c = a(1+b) c(1+b) . for the connection degree a+ λb− c based on the fourth set pair similarity relationship, the homeopathic value method can be used to deal with the threshold λ so that a novel connection degree a + (a − c)b − c + b2λ is obtained. because the impact from the part of the uncertainty that is classified into the discrepancy degree b2 on the relationship of similarity is very little, b2 can be ignored. because of this reason, we define the set pair homeopathic similarity relationship as follows. definition 14.suppose that s∗ = (u,a, v, f)is an incomplete information system.for ∀a ∈ a,∀x ∈ u,b ⊆ a, 0 ≤ α ≤ 1, the set pair homeopathic similarity relationship of x and y based on the threshold value α is defined as follows: 306 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... smhα(b) = {(x, y) ∈ u ×u | ub(x, y) = a+ bi+ cj, a+(a− c)b− c ≥ α} (13) the set pair homeopathic similarity relationship satisfies the properties of reflexivity and symmetry, but not that of transitivity. definition 15.in the incomplete information system s∗ = (u,a, v, f).for ∀a ∈ a,∀x, y ∈ u,b ⊆ a, 0 ≤ α ≤ 1, the set pair homeopathic similarity classes of x and y with respect to the attribute set b based on the threshold value α are defined as follows: smhα b(x) = {y ∈ u | ub(x, y) = a+ bi+ cj, a+ (a− c)b− c ≥ α} (14) the set pair homeopathic similarity classes reflect such object set that the similarity degree of with respect to the attribute set b is greater or equal to α. 4 variable precision rough set based on set pair homeopathic similar relationship 4.1 construction of the model in any real-life information system, there always exist a variety of noise data, which is difficult for the traditional rough set method to deal with. in order to deal with such problem, ziarko [27–29] put forward the variable precision rough set model by introducing the threshold value β and approximation space to reflect this kind of restrictions. based on the thought and methodology of variable precision rough set model, we introduce the set pair homeopathic similarity relationship to substitute the equivalence relationship of rough set so that a novel variable precision rough set model can be constructed. accordingly, the lower and upper approximations of the variable precision rough set model based on the set pair homeopathic similarity relationship can be defined as follows. definition 16.suppose that s∗ = (u,a, v, f) is an incomplete information sysytem.for x ⊆ u,b ⊆ c, 0 ≤ α ≤ 10.5 ≤ β ≤ 1, the β-lower and β-upper approximations of xbased on the set pair homeopathic similarity relationship are defined respectively by: aprα,βb (x) = ∪{x ∈ u | |smhα(x) ∩x| |smhα(x)| ≥ β} (15) aprα,β b (x) = ∪{x ∈ u | |smhα(x) ∩x| |smhα(x)| ≥ 1− β} (16) the β-lower approximation aprα,β b (x) of the set x based on the set pair homeopathic similarity relationship can also be called as the positive region of the variable precision rough set model, denoted as posα,βb (x).in other words, for given confidence threshold values α and β,aprα,β b (x) is the set in which the universe advances in systems science and application (2015) vol.15 no.4 307 u can be classified definitely into all the element sets of the set x based on the set pair homeopathic similarity relationship. the β-upper approximation aprα,βb (x) of the set x reflects that for a given threshold value β ,the universe u can probably be classified into all the element sets of x definition 17.suppose that s∗ = (u,a, v, f) is an incomplete information sysytem.for 0 ≤ α ≤ 1, 0.5 ≤ β ≤ 1,the β-negative region and β-boundary of a subset x based on the set pair homeopathic similarity relationship are defined as follows: negα,βb (x) = ∪{x ∈ u : |smhα(x) ∩x| |smhα(x)| ≤ 1− β} (17) bndα,βb (x) = ∪{x ∈ u : 1− β < |smhα(x) ∩x| |smhα(x)| < β} (18) the β-negative region negα,βb (x) of x based on the set pair homeopathic similarity relationship reflects that for a given confidence threshold β , the universe certainly cannot be classified into all the elements of the collection set x. the boundary bndα,βb (x) of x based on the set pair homeopathic similarity relationship reflects that for the given confidence threshold β , the universe u certainly cannot be classified into all the element sets of either the set x or the set -x . definition 18.suppose that s∗ = (u,a, v, f) is an incomplete information sysytem.for 0 ≤ α ≤ 1, 0.5 ≤ β ≤ 1,the β classification quality of x based on the set pair homeopathic similarity relationship is defined as follows: rα,β(b,d) = ∣∣∣∪{x ∈ u : |smhα(x)∩r(x)| |smhα(x)| ≥ β} ∣∣∣ |u | (19) the quantity γα,β(b,d) measures the proportion that the possible correct classification knowledge is in the existing knowledge for the given values α, β in the universe. 4.2 properties of the model theorem 1.for the given incomplete information system s∗ = (u,a, v, f), let x ⊆ u, y ⊆ u,∀b ⊆ c, 0 ≤ α ≤ 1, 0.5 < β ≤ 1, the following hold true: (1)aprα,βb (x ∪ y ) ⊇ aprα,βb (x) ∪ aprα,βb (y ); (2)aprα,β b (x ∩ y ) ⊆ aprα,β b (x) ∩ aprα,β b (y ); (3)aprα,β b (x ∪ y ) ⊇ aprα,β b (x) ∪ aprα,β b (y );and (4)aprα,βb (x ∩ y ) ⊆ aprα,βb (x) ∩ aprα,βb (y ); proof.(1)for x ⊆ u, y ⊆ u, 0 ≤ α ≤ 1, 0.5 < β ≤ 1, we have |smhα(x) ∩ (x ∪ y )| |smhα(x)| ≥ |smhα(x) ∩x| |smhα(x)| 308 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... and |smhα(x) ∩ (x ∪ y )| |smhα(x)| ≥ |smhα(x) ∩ y | |smhα(x)| therefore, we obtain aprα,βb (x ∪ y ) ⊇ aprα,βb (x) ∪ aprα,βb (y ); (2)for x ⊆ u, y ⊆ u, 0 ≤ α ≤ 1, 0.5 < β ≤ 1, we have |smhα(x) ∩ (x ∩ y )| |smhα(x)| ≤ |smhα(x) ∩x| |smhα(x)| and |smhα(x) ∩ (x ∩ y )| |smhα(x)| ≤ |smhα(x) ∩ y | |smhα(x)| therefore, we obtain aprα,β b (x ∩ y ) ⊆ aprα,β b (x) ∩ aprα,β b (y ); similar to the argument of statement (1), statement (3) can be proved; and statement (4) can be proved in the same way as that of statement (2). qed theorem 2.for the given incomplete information system s∗ = (u,a, v, f), let x ⊆ y ⊆ u,∀b ⊆ c, 0 ≤ α ≤ 1, 0.5 < β ≤ 1, the following hold true: (1)aprα,β b (x) ⊆ aprα,β b (y );and (2)aprα,βb (x) ⊆ aprα,βb (y ). proof.(1)due to x ⊆ y ,it follows that x ∩y = x and x ∪y = y . so we obtain aprα,β b (x) = aprα,β b (x ∩ y ) = aprα,β b (x) ∩ aprα,β b (y ) therefore,we have aprα,β b (x) ⊆ aprα,β b (y ). (2)from x ⊆ y ,it follows that x ∪ y = y so that we have aprα,βb (y ) = aprα,βb (x ∪ y ) = aprα,βb (x) ∪ aprα,βb (y ) therefore,we obtain aprα,βb (x) ⊆ aprα,βb (y ).qed theorem 3.for the given incomplete information system s∗ = (u,a, v, f), let x ⊆ u,b1 ⊆ b2 ⊆ c, 0 ≤ α ≤ 1, 0.5 < β ≤ 1, the following hold true: (1)aprα,β b1 (x) ⊆ aprα,β b2 (x);and (2)aprα,βb1 (x) ⊇ aprα,βb2 (x). proof.(1)for ∀x ∈ aprα,β b1 (x),from b1 ⊆ b2 ⊆ a it follows that [x]α,βb1 ⊆ x such that [x]α,βb2 ⊆ [x]α,βb1 .so we obtain [x]α,βb1 ⊆ x such that x ∈ aprα,β b2 (x) therefore,we have aprα,β b1 (x) ⊆ aprα,β b2 (x) (2)for ∀y ∈ aprα,βb2 (x),we have [y]α,βb2 ∩x ̸= ∅,satisfying that [y]α,βb1 ∩x ̸= ∅.so, we get y ∈ aprα,βb1 (x).therefore,aprα,βb1 ⊇ aprα,βb2 (x) follows.qed advances in systems science and application (2015) vol.15 no.4 309 theorem 4.for the given incomplete information system s∗ = (u,a, v, f), let x ⊆ u,b ⊆ c, 0 ≤ α1 ≤ α2 ≤ 1, and 0.5 < β1 ≤ β2 ≤ 1. then following hold true: (1)aprα,β1 b (x) ⊆ aprα,β2 b (x); (2)aprα,β1 b (x) ⊇ aprα,β2 b (x). (3)aprα1,β b (x) ⊆ aprα2,β b (x); and (4)aprα1,β b (x) ⊇ aprα2,β b (x). proof.(1)for ∀x ∈ aprα,β1 b (x),from 0.5 < β1 ≤ β2 ≤ 2 it follows that [x]α,β1 b ⊆ x. so, we have [x]α,β2 b ⊆ [x]α,β1 b from which we get x ∈ aprα,β2 b (x). and consequently we obtain aprα,β1 b (x) ⊆ aprα,β2 b (x). (2)for ∀y ∈ aprα,β2 b (x), [y]α,β2 p ∩x ̸= ∅ holds true so that [y]α,β1 b ∩x ̸= ∅ follows. hence, we get y ∈ aprα,β1 b (x), and consequently aprα,β1 b (x) ⊇ aprα,β2 b (x) similar to the proof of statement (1), statement (3) follows; and statement (4) can be shown in the same way as that of statement (2). qed the results in theorem 4 indicate that β value is negatively related to the classification quality. in particular, when β value increases, the classification quality decreases and the positive and negative regions of the set x based on the variable precision rough set model, as proposed in this paper, will become narrower, while the boundary region of the set x will become wider. it means that only a small number of objects are classified. on the other hand, as the value of β decreases, the classification precision increases, and the positive and negative regions of the set x based on the variable precision rough set will widen, while the boundary region narrows. that means that most of the objects are classified, but possibly misclassified. 4.3 attribute reduction algorithm according to the similarity degree of the positive region, as constructed based on the proposed variable precision rough set model, the following attribute reduction algorithm is established for solving the reduction problem. the special steps are given as follows. input: the incomplete information system s∗ = (u,a, v, f) , the set pair homeopathic threshold value α and confidence threshold value β . output: a reductionb of the incomplete information system s∗ = (u,a, v, f). step 1: suppose b = c; step 2: compute the system classification quality γα,β(c,d); step 3: for any condition attribute a ∈ b, compute γα,β(b − a,d) and the dependence degree sigα,β(a) of the attribute a; step 4: if γα,β(b − a,d) ≥ γα,β(c,d) and sigγ(a) is the least, then b = b − {a} . if for some a ∈ b, γα,β(b − a,d) ≤ γα,β(c,d) and all sigγ(a) are equal, then the attribute a with the missing values for all objects is selected and 310 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... let b = b − {a} ; step 5: if for some a ∈ b, γα,β(b − a,d) ≤ γα,β(c,d) and all sigγ(a) are equal, then go to step 6; otherwise go to step 3; and step6:output a reduction of the initial incomplete information system. 5 case analysis suppose that the decision-making information system s = (u, v,a, f) is about the initial accessories of the military aircraft engine system, where u = {u1, u2, ..., u10} is the set of initial accessories, c = {a1, a2, a3, a4, a5} the set of condition attributes, where ,a1, a2, a3, a4, and a5 stand respectively for importance, consumable, convertibility, procurability, and economical efficiency, and d = {d}. for va1 = {2, 1, 0} , 2 represents key, 1 represents more important, 0 represents general; for va2 = {2, 1, 0} , 2 represents consumable, 1 represents vulnerable, 0 represents general; for va3 = {1, 0} , 1 represents substitution, 0 represents cannot be replaced;, for va4 = {2, 1, 0} , 2 represents longer for the order cycle, 1 represents general for the order cycle, 0 represents immediate purchase; for va5 = {2, 1, 0} , 2 represents excellent, 1 represents general, 0 represents bad; for vd = {1, 0} , 1 represents yes, 0 represents no. due to the fact that there exist missing attribute values, the decision-making information system s = (u,a, v, f) is an incomplete decision-making information system s∗ = (u,a, v, f), the details of shown are given in table 1. (1)when α = 0.2, based on the set pair homeopathic similarity relationship, table 1 the decision-making information system for the initial accessories selection of the military aircraft engine system order# a1 a2 a3 a4 a5 d 1 2 1 0 * 1 1 2 1 0 1 2 2 1 3 2 * 0 1 1 1 4 0 2 * 2 0 0 5 * 1 0 1 1 0 6 2 * 1 0 2 1 7 0 2 1 1 * 0 8 0 1 0 0 1 0 9 2 0 1 2 2 1 10 2 * 1 1 2 0 the universe on the condition attribute set c can be divided into: u/c = {x0.2 1 , x0.2 2 , x0.2 3 , x0.2 4 } advances in systems science and application (2015) vol.15 no.4 311 where x0.2 1 = {u1, u3, u5, u8}, x0.2 2 = {u2, u6, u9, u10} and x0.2 3 = {u4, u7}. according to the decision attribute set d , the universe can be divided into: u/d = {y1, y2} where y1 = {u1, u2, u3, u6, u9} and y2 = {u4, u5, u7, u8, u10}. when β = 1 , the lower and upper approximations of the decision-making classes y1 and y2 can be computed as follows: aprα,β b (y1) = ∅, aprα,βb (y1) = {u1, u2, u3, u5, u6, u8, u9, u10}, aprα,β b (y2) = {u4, u7}, aprα,βb (y2) = {u1, u2, u3, u4, u5, u6, u7, u8, u9, u10}, and γα,β(b,d) = |{u4,u7}| |{u1,u2,u3,u4,u5,u6,u7,u8,u9,u10}| = 2 10 = 0.2 when β = 0.7 , the lower and upper approximations of the decision-making classes y1 and y2 can be produced as follows: aprα,β b (y1) = {u2, u6, u9, u10}, aprα,βb (y1) = {u1, u2, u3, u5, u6, u8, u9, u10}, aprα,β b (y2) = {u4, u7}, aprα,βb (y2) = {u1, u3, u4, u5, u7, u8}, and γα,β(b,d) = ||{u4,u7}|∪|{u2,u6,u9,u10}|| |{u1,u2,u3,u4,u5,u6,u7,u8,u9,u10}| = 6 10 = 0.6 (2) when α = 0.4 , based on the set pair homeopathic similarity relationship, the universe on the condition attribute set c can be divided into: u/c = {x0.4 1 , x0.4 2 , x0.4 3 , x0.4 4 , x0.4 5 , x0.4 6 } where x0.4 1 = {u1}, x0.4 2 = {u2, u9, u10}, x0.4 3 = {u3, u5}, x0.4 4 = {u4}, x0.4 4 = {u6}, x0.4 5 = {u7} and x0.4 6 = {u8}. according to the decision attribute set d , the universe can be divided into: u/d = {y1, y2} where y1 = {u1, u2, u3, u6, u9} and y2 = {u4, u5, u7, u8, u10}. when β = 1 , the lower and upper approximations of the decision-making classes y1 and y2 can be produced as follows: aprα,β b (y1) = {u1, u6}, aprα,βb (y1) = {u1, u2, u3, u5, u6, u9, u10}, aprα,β b (y2) = {u4, u7, u8}, aprα,βb (y2) = {u2, u3, u4, u5, u7, u8, u9, u10}, and γα,β(b,d) = ||{u1,u6}|∪|{u4,u7,u8}|| |{u1,u2,u3,u4,u5,u6,u7,u8,u9,u10}| = 5 10 = 0.5 when β = 0.65 , the lower and upper approximations of the decision-making classes y1 and y2 are given as follows: aprα,β b (y1) = {u1, u2, u6, u9, u10}, aprα,βb (y1) = {u1, u2, u3, u5, u6, u9, u10}, aprα,β b (y2) = {u4, u7, u8}, aprα,βb (y2) = {u3, u4, u5, u7, u8}, and γα,β(b,d) = ||{u1,u2,u6,u9,u10}|∪|{u4,u7,u8}|| |{u1,u2,u3,u4,u5,u6,u7,u8,u9,u10}| = 8 10 = 0.8 based on the attribute reduction algorithm in section 4.3, when α = 0.4 and β = 0.65, a reduction {a1, a5} can be acquired. from the reduction {a1, a5}, some probabilistic decision-making rules set can be generated as shown in table2. 312 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... table 2 probabilistic decision-making rules generated by the reduction {a1, a5} with respect to the threshold values (α = 0.4, β = 0.65) rules# supprot number confidence(%) a1 = 2 and a5 = 2 66.7%−−−→ d = 1 3 66.7 a1 = 1 and a5 = 2 100%−−−→ d = 1 1 100 a1 = 0 and a5 = 0 100%−−−→ d = 1 1 100 a1 = 0 and a5 = 1 100%−−−→ d = 1 1 100 a1 = 0 and a5 = 2 100%−−−→ d = 1 1 100 according to table 2, we can see that eight objects in the object set can be correctly classified. that is, the classification quality is 80%. from the reduction {a1, a5} , it follows that the company, which manufactures military aircraft engine systems, pays more attention to importance and economical efficiency of the initial accessories. from the set of probabilistic decision-making rules, it follows that if the importance of the initial accessories is general, then the initial accessories should not be selected; if the importance is important and the economical efficiency is excellent, then the initial accessories should be selected. based on the previous calculation and analysis, by adjusting the set pair homeopathic parameter α and the confidence threshold parameter β , the classification ability of the model cam be improved; and the model can be made to possess some fault tolerant ability so that it can materialize the correct classification, and effectively extract the rules for decision making. 6 conclusions due to the complexity and uncertainty of the objective world, as well as the limitation of human understanding, available information systems often contain incomplete information, and the omission of attribute values in the systems increases the uncertainty and the noise of the system. it has been difficult for the traditional rough set method, developed for handling complete information systems, to deal with incomplete information systems. by taking advantage of the set pair analysis methodology, in this paper we constructed the variable precision rough set based on set pair homeopathic similar relationship. by adjusting the confidence threshold parameters (α, β) , we can make the model possess the fault tolerant ability so that it can effectively deal with information systems with noise data, ambiguity, and other kinds of uncertain information. it is shown that our model can materialize the extraction of decision-making rules and knowledge discovery out of a given incomplete information system. for any incomplete information system that contains preference information, grey information, noise data, advances in systems science and application (2015) vol.15 no.4 313 and other kinds of uncertain information, how to construct the grey dominance variable precision rough set model based on the set pair homeopathic similar relationship needs to be investigated, while its potential scope of applications needs to be explored. acknowledgement this work is partially funded by a marie curie international incoming fellowship within the 7th european community framework programme (grant no. fp7-piif-ga-2013-629051); national natural science foundation of china (71301061; 71503103), national social science fund project (12azd111); ministry of education humanities and social sciences youth fund(13yjc630120); natural science foundation of jiangsu province (bk20150157); jiangsu province social science fund project(14glc008); the research base of chinese iot development strategy(133930), the fundamental research funds for the central universities(jusrp11583; jusrp1507zd). references [1] pawlak, z. (1998), “rough sets”, international journal of information and computer sciences, vol. 49, no. 5, pp. 415-422. [2] pawlak, z., and skowron, a. (2007a), “rudiments of rough sets”, information sciences, vol. 1, pp. 3-27. [3] pawlak, z., and skowron, a. (2007b), “rough sets: some extensions”, information sciences, vol. 1, pp. 28-40. [4] pawlak, z., and skowron, a. (2007c), “rough sets and boolean reasoning”, information sciences, vol. 1, pp. 41-73. [5] zhao s., zhang l., xu x. s.(2014). “hierarchical description of uncertain information”, information sciences, vol. 261, no. 1, pp. 133-146. [6] xu w. h., wang q. r. zhang x. t. (2014). “multi-granulation rough sets based on tolerance relations”, soft computing, vol. 17, no. 7, pp. 1241-1252. [7] millo s., reinier g.c., deborah c.c..(2014), “aggregation of similarity measures for ortholog detection: validation with measures based on rough set theory”, computation systems, vol. 18, no. 1, pp. 19-35. [8] zhang h. y., leung y, zhou l..(2013). “variable-precision-dominance-based rough set approach to interval-valued information systems”, information sciences, vol. 244, no. 2, pp. 75-91. 314 yong liu and yi lin:variable precision rough set model based on set pair homeopathic ... [9] kryszkiewicz, m. (1998),“rough set approach to incomplete information systems” , information sciences, vol. 112, pp. 39-49. [10] stefanowski, j., and tsoukias, a.. (2001), “incomplete information tables and rough classification”, computational intelligence, vol. 17, pp. 545-566. [11] wang g. y. (2002). “the extensions of rough set theory in incomplete information system”, researcher and development of computer, vol. 39, no. 10, pp. 1238-1243. [12] greco, s., matarazzo, b., and slowiski, r. (2002a), “rough approximation by dominance relations”, international journal of intelligent systems, vol. 17, pp. 153-171. [13] greco, s., matarazzo, b., and slowiski, r. (2002b), “rough set s theory f or multi-criteria decision analysis”, european journal of operational research, vol. 129, pp. 1-47. [14] shao, m. w., and zhang, w. x.. (2005), “dominance relation and rules in an incomplete ordered information system”, international journal of intelligent systems, vol. 20, pp. 13-27. [15] yang, x. b., yang, j. y, and chen, w., et al. (2008).“ dominancebased rough set approach and knowledge reductions in incomplete ordered information system”, information sciences, 2008, vol. 178, no. 4, pp. 1219-1234. [16] xie, j., song, y. q., chen, j. m., et al. (2008).“ extensions of rough set model and set pair analysis in incomplete ordered decision system”, computer science, vol. 35, no. 12, pp. 154-157. [17] zhao k. q.. (2000). “set pair analysis and preliminary application”, hangzhou: zhejiang science and technology press, pp. 68-91. [18] huang, b., and zhou, x. z. (2002), “rough set model based on set pair analysis in incomplete system” computer science, vol. 29, pp. 1-3. [19] huang, b., and zhou, x z (2004), “the extensions of rough set theory in incomplete information system based on connection degree”, system engineering theory and practice, vol. 1, pp. 88-92. [20] xu, y., li, l. s., and li, x. j.. (2008). “generalized rough set model based on set pair situation”, journal of system simulation, vol. 20, no. 6, pp. 1515-1518. advances in systems science and application (2015) vol.15 no.4 315 [21] xu, y., and li, l. s.. (2010). “variable precision rough set model based on set pair situation”, control and decision, vol. 25, no. 11, pp. 1732-1736. [22] xu y., and li, l. s.. (2011). “variable precision rough set model based on (α, β) connection degree tolerance relation”, acta automatica sinica, vol. 37, no. 3, pp. 303-308. [23] tao, z., dai, h. j., and zhang, y.. (2008). “set pair rough set in incomplete system”, computer application, vol. 28, no. 7, pp. 1684-1685,1691. [24] liu, f. c.. (2005), “set pair rough set model based on limited tolerance relations”, computer science, vol. 32, no. 6, pp. 124-128. [25] liu, f. c.. (2006), “an algorithm for attributes reduction in variable precision rough set model based on set pair analysis”, computer science, vol. 33, no. 3, pp. 185-187. [26] zhou l., and shu, l.. (2006).“ rough set model based on new set pair analysis”, fuzzy systems and mathematics, vol. 20, no. 4, pp. 111-116. [27] ziarko, w. (1993a). “variable precision rough set model”, journal of computer and system sciences, vol. 46, no. 1, pp. 39-59. [28] ziarko, w. (1993b).“ analysis of uncertain information in the framework of variable precision rough sets”, foundations of computing and decision sciences, vol. 18, pp. 381-396. [29] ziarko, w. (1999). “rough sets, fuzzy sets and knowledge discovery”, singapore: springer, pp. 1-98. corresponding author yong liu can be contacted at:clly1985528@163.com adv syst sci appl 2018; 03; 111-122 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/581 maximum power point tracking technique for pv system by using golden section search method sonvir singh1 1) department of computer science & engineering, visvesvaraya national institute of technology, nagpur, india e-mail: sonvirsingh036@gmail.com abstract: importance of non-conventional source of energy (ncse) is upsurging day by day to tackle the increased load, energy consumption, and environmental pollution related problems. out of non-conventional source of energy (ncse), solar energy, mainly solar photovoltaic (spv) energy, is one of the most promising renewable source of energy (res) because of pollution-free, no rotating part, an abundance of solar energy, low-cost energy, and high efficiency. the main thing is how to extract maximum power from the solar pv panel. power extract from the pv system depends upon the solar radiation (isolation), ambient temperature and terminal voltage. regulation of terminal voltage is required to extract the maximum power at the operating point. the technique of extracting maximum power by changing its terminal voltage by the various methods is called as mppt. this paper takes up the problem of finding the mppt power point under changing environmental conditions, developing efficient mmpt and changing the operating point of the system very fast according to the environmental condition. in this paper, a new mppt technique is developed for tracking maximum power point which is efficient to existing techniques with the help of the golden section search optimization method. this paper is intended to serve an efficient method and reliable to find the maximum power point at which maximum power is extracted from the pv system. keywords: maximum power point tracking (mppt) techniques, non-conventional source of energy (ncse), solar photovoltaic (spv) energy, photovoltaic (pv) system 1. introduction to extract maximum power from the photovoltaic (pv) systems under different environmental conditions, the maximum power point techniques are used. this principle is generally used for those sources which have variable power outcome, for example, optical power transmission, thermophotovoltaic etc. pv system has a different configuration in respect to their relationship to the inverter, grid system, batteries or electrical load. but the ultimate destination is that to transfer more power from the pv cell, which depends upon the three factors-falling of sunlight on the solar panels, the temperature of the pv panel and electrical load characteristic. temperature remains constant by using heat sink material at the backside of the panel. as the load characteristic, when sunlight intensity changes the maximum power transfer changes accordingly then the load characteristic that gives maximum transfer efficiency also changes. maximum power point technique (mppt) is a technique, which includes the electronics controller used for the extracting maximum available power from the pv panel under various weather conditions. the maximum power extraction from the pv panel varies with solar radiation, ambient temperature, solar cell temperature. mmpt helps in finding the point at which power transfer from the panel to load is maximum. pv cells have a complicated relationship between temperature and resistance that produce non-linear characteristic. which can be analyzed based on the i-v plot. the main function of mppt is to change the output mailto:sonvirsingh036@gmail.com 112 s. singh copyright ©2018 assa. adv. in systems science and appl. (2018) voltage of the pv cells and provide a proper load to obtain maximum power for any given condition [1]. fig. 1 block diagram of the proposed model. an mppt is used for adjusting the voltage level according to the varying output of pv panel for extracting the maximum power for the pv and transferring that power to the variable or fixed the load [2] [3]. a converter and controlling unit as an interface between the load and pv module. by changing the load voltage, the load impedance seen by the source is varied and matched at the point of peak power with the source so as to transfer the maximum power[3]. therefore mppt technique is needed to operate the pv system at the maximum power point. many mppt techniques have been proposed in the literature, examples p&o methods, incremental conductance(ic) methods etc [4] [5]. the main advantage of the p&o mppt algorithm is its low cost and easy implementation but, it may fail to track the maximum operating point under varying weather conditions [6]. the step size of the duty cycle change required at the initial stage and final stage to find the maximum power point [7]. value of step size of the duty cycle should be taken large for points far away from the maximum power point and small for points near the mppt [8]. different methods of finding the maximum power point have different outcomes. some are better in finding the maximum power point at varying environmental condition and some are fast in the respect of time, accuracy and efficiency [9,10]. in this research paper, the system performance is optimized by golden section search method. by changing the interval or selection of section without using a buck-boost converter, the source impedance can be matched to adjust the load impedance to improve the efficiency of the system. the proposed model is shown in fig. 1. 2. photovoltanic system. the system which converts solar energy into electrical energy is called a photovoltaic (pv) system. the basic or smallest unit of pv system is pv cell. these cells are arranged in parallel, series or series-parallel as per requirement. the voltage and currently available at the terminal of pv system, it is directly fed to small loads like lighting, dc motors or connected to the distribution grid by using energy conversion devices [2]. main parts of pv system are pv module, charger, battery, inverter, and load. 2.1. equivalent model a photovoltaic cell is a device used to convert solar radiation into electrical energy. the equivalent circuit diagram of the photovoltaic cell shown in fig. 2 pv cell is a current source, it is produced by breaking of bonds and generation of electron and hole which causes the flow of current in the cell. the equivalent circuit consists current source, diode, shunt resistance, series resistance and load. shunt resistance represents the electron-hole combination before its maximum power point tracking technique 113 copyright ©2018 assa. adv. in systems science and appl. (2018) reach to load. the generation currently in the pv panel depends upon the physical and chemical properties of the material, the age of the solar cell, irradiation per unit area and temperature of the cell, environmental conditions, spectral characteristics of sunlight, dirt, and shadow and so on [11,12]. fig. 2 equivalent single pv cell model the diode current is given by the equation l l lq(v -i r ) nkt d oi = i e -1       (1) the load current which flows in the load is a difference of the photocurrent, diode current and shunt resistance current is given by. l ph d shi i i i   (2) shunt resistance (rsh) represent the electron-hole combination before it reaches to load (for simplifying ignore rsh), ish current becomes zero. ( ) 1 l l lq v i r nkt ol phi i ei           (3) where,  il is the load current or cell current (a).  n is the ideality factor.  q is the charge of the electron (coulomb).  k is the boltzmann's constant (j/k).  t is the temperature of the cell.  iph is the photocurrent (a).  rs ,rsh are series and shunt resistance of the cell (ohms).  vl is the cell output voltage (v). after solving equation no.-3, the load voltage is ln 1 ph l l l s o i inkt v i r q i         (4) in the case, open circuit, il=0 than the corresponding voltage is open circuit voltage (voc). ln 1 ph oc o inkt v q i        (5) 114 s. singh copyright ©2018 assa. adv. in systems science and appl. (2018) this is the order of 10-5. it should be noted that voc will not reduce as much as iph change. in the case of short circuit, vl=0 and impedance are low tags corresponding current is short circuit current which can be approximated as:  sc n pi qg l l  (6) where g is the rate of generation of electrons and holes, ln and lp are the diffusion length of electrons and holes respectively. above equation is not true for the most of cell because of several assumptions. the above equation shows that the short circuit current depends upon the generation rate and diffusion rate. fig.3 i-v and p-v characteristics of pv panel the current-voltage characteristic curve of a pv cell for a certain irradiance at fixed cell temperature is shown in the fig no.-3. the current of pv cell depends upon the external applied voltage and amount of sunlight sticking the surface of the cell. the power vs voltage curve showed in fig no.-3.here p is the power extracted from the pv cells. the curve mainly depends upon current insolation and temperature when insolation increases the power output of pv panel increases whereas when the temperature of the cell increases the power output of pv panel is decreases [13]. constant power generated by the pv panel may also be used to provide the energy to the main grid and micro-grid. energy distribution priority is decided by the priority factor, load and prize forecasting [14]. 2.2. temperature effect. in this paper, consider only two factor falling sunlight on the solar panel and electrical load characteristics. make temperature as a constant throughout the process by using temperature controller like heat sink material system [15]. heat sink material put in between glass cover of solar panel and cells. increased in temperature of solar panel decreased the output of the solar panel. to improve the efficiency of the solar panel, consider the temperature as constant in this paper with the help temperature control system. for controlling the temperature of the solar panel the temperature sensor is used to detect the temperature and heat sink material sink the heat if the temperature is more than the pre-set value and this heat is transfer outside by maximum power point tracking technique 115 copyright ©2018 assa. adv. in systems science and appl. (2018) the help of small cooling fans. so that temperature remains constant throughout the day during processing. systematic diagram is shown in figure no.-4. fig.4 side view of solar panel 3. golden section search technique. the golden search technique is a technique to find out the minimum and maximum (extremum) of strictly increasing and decreasing curve for particular interval or say for unimodal function by successively narrowing the range of values inside in which extremum point know to exist. here, this technique is used to find out the maximum power point in the pv curve [16]. 3.1 unimodal function. a function f(x) is unimodal on [a,b] interval if for some point x* on [a,b],f(x) is strictly increasing on [a,x*] and strictly decreasing on [x*,b].concave and convex shape type functions are unimodal. unimodal type function make the search easier to find maximum point in the given interval.start with an interval of uncertainty (the interval in which maximum point must lie) equal to [a,b],whose length is difference of end point of interval i.e b-a. consider two points in the interval [a,b] and evaluate the function at these points a b x1 x2 if f(x1) < f(x2) , function is increasing in the range [x1,x2]. therefore, the function value must be greater than f(x1). since the function is unimodal, then maximum cannot be lying in the range of (a, x1]. thus concluding that the maximum is in the range of ( x1, b]. if f(x1) > f(x2), since function is unimodal, the function’s maximum must be greater than x2 .therefore the maximum must lie in the range [a, x2) .if f(x1) > f(x2), then the maximum must lie in the range of (x1,x2). since the points, x1 and x2 have to be on either side of the maximum. according to fibonacci method 21   nnn fff (7) 116 s. singh copyright ©2018 assa. adv. in systems science and appl. (2018) dividing the above equation by fn-1. 1 2 1 1     n n n n f f f f (8) golden ratio ( ) 1 1 2 ( 1) lim lim lim n jn n n n nn n n j ff f f f f            (9) take limit on both side of equation (2) and replace with the golden ratio. 1 1    , 2 1 0    (10) after solving the equation no.-4 1 5 2    , consider the only positive root of  1.618  (golden ratio value) [17]. r = 1 0.618   r is conjugate root or conjugate golden ratio. 3.2 iterative process for finding maxima. step-1: given initial of uncertainty = [a, b],whose length is (b-a), stopping tolerance ( ). step-2: )(* 1 abrl  , * 1 1 0lim n n n f l l f     2 1 0 1 *lim n n n n n f f l f f     now, * 1 1x a l  and * 2 1x b l  a x2 bx1 l2 * l2 * f(x2) f(x1) points on x-axis are x1 and x2, their corresponding values of function f(x1) and f(x2) respectively if f(x1) > f(x2), the interval of uncertainty is [a, b] but now the new interval become [a, x2) l1=r(b-a) , check the condition, if l1<  than stop the iteration process and maxima must lie in [a,x2). if not, then goes to next step step-3: generate x3. maximum power point tracking technique 117 copyright ©2018 assa. adv. in systems science and appl. (2018) a x1 x3 x2 f(x3) f(x1) now, x2=b )( 2 * 2 axrl  now, x3= x2l2 * and x1= a+ l2 * points on x-axis are x1 and x3, their corresponding values of function f(x1) and f(x3) respectively if f(x1) > f(x3) , discard (x3, x2] and new interval become [a,x3]. l2=r2 (b-a), check the condition, if l2 <  than stop the iteration process and maxima must lie in [a, x3]. if not, then goes to the next step this process is stopped when the distance between the outer points is smaller than stopping tolerance. if lk <  maxima are found, stop the iteration. 4. the golden search method. consider the voltage range in which the possibility of finding the maximum power point is (v, v). normally in the case of pv curve v equal to zero and v is open circuit voltage. after each iteration, the voltage range is changing accordingly the range length also reducing after execution of each iteration. searching will stop when range length less than or equal to the tolerance. length of uncertainty interval. )( oo k k vvrl  (11) here k represents no. of iteration to find out the voltage at which power transfer is more. solve above education by taking natural logarithmic both side to find the value of k r vvl k ook ln ))/(ln(   (12) algorithm: golden section search algorithm 1: define the interval of uncertainty as (vo – vo) =(v,v) 2: calculate v1k= vk r(vk -vk) 3: calculate v2k = vk + r(vk vk) 118 s. singh copyright ©2018 assa. adv. in systems science and appl. (2018) 4: if (p(v1k) < p(v2k)) then. 5: p* must be in (v1k, vk) 6: vk+1 = v1k 7: vk+1 = vk 8: else 9: p* must be in (vk, v2k) 10: vk+1 = vk 11: vk+1 = v2k 12: if lk = (vk vk) < then 13: stop 14: else 15: k = k + 1 16: go to line 2 k shows the no. 12 of iteration which required to find the voltage at which maximum power will be extracted from the pv panel. firstly define the interval of uncertainty, vo normally zero and vo equal to open circuit voltage of the pv panel than calculated v1k and v2k by the above mention formula of the golden search method. after the calculation of v1k and v2k , calculate power at v1k and v2k and compare these power if (p(v1k) < p(v2k)) than maximum power point must be lie in (v1k, vk). now voltage range will be a change from vk+1 = v1k and vk+1 = vk otherwise maximum power point must be lie in (vk, v2k) and voltage range will change from vk+1 = vk and vk+1 = v2k. check the stopping condition, lk = (vk vk) <  if it is true then stop the process otherwise increase the value of k by 1 and repeat the whole process till the finding of the voltage at which maximum power is extracted from the pv panel. 5. simulation result in this paper, the model is developed with matlab/simulink. the proposed model is shown in fig.-4. the pv module parameter considers for the proposed model are: voc (open circuit voltage) =21.8 v. isc (short circuit current) =3.11 a. vmp (voltage at mpp) =17.44. imp (current at mpp) =2.86 a. pmp (power at mpp) =50 w. the outputs of the pv panel are current and voltage and it depends upon the solar radiance and temperature of the panel. the temperature of pv panel is maintained constant at 25 degree celsius and solar intensity is varied up to the rated value of 1000 w/m2. the current fed to the multiplier to obtain the power in respect to the maximum power point tracking technique 119 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.5 i-v and p-v characteristics of pv panel table 1. iteration process chart. iteration vo vo v1 v2 p(v1) p(v2) p* in interval lk=(vk-vk) 1. 0 21.8 8.33 13.47 25.91 41.76 (8.33,21.8) 21.8 2. 8.33 21.80 13.48 16.65 41.79 49.60 (13.48,21.8) 13.47 3. 13.48 21.80 16.66 18.62 49.61 47.26 (13.48,18.62) 8.32 4. 13.48 18.62 15.44 16.66 47.25 49.61 (15.44,18.62) 5.14 5. 15.44 18.62 16.65 17.41 49.6 50.02 (16.65,18.62) 3.18 6. 16.65 18.62 17.41 17.86 50.02 49.99 (16.65,17.86) 1.97 7. 16.65 17.86 17.11 17.40 49.99 50.02 (17.11,17.86) 1.21 8. 17.11 17.86 17.40 17.57 50.02 49.94 (17.11,17.57) 0.75 9. 17.11 17.57 17.29 17.39 50.03 50.02 (17.11,17.39) 0.46 10. 17.11 17.39 17.21 17.28 50.02 50.03 (17.21,17.39) 0.28 11. 17.21 17.39 17.28 17.32 50.03 50.03 018 output voltages of voltage interval changing system and output of pv panel is fed to the voltage interval changing system.v1 and v2 voltages are changed in each iteration. powers are calculated in respect to voltages v1 and v2 respectively by the multiplier. side by side length is compared by tolerance value if its length equal to or less than the tolerance value than iteration is stopped. at voltage, the iteration stop is voltage at which maximum power will be extracted. no. of iteration requires to finding the maximum power point at different weather condition show in table 1. no. of iteration is more compare to other methods but the time required in each iteration is very less. therefore this method has high speed to trace the maximum power point compare to others. the simulation circuit is shown in the fig-4. in this proposed circuit is without the buck-boost converter, the output voltage level is changed according to the load as well as radiation of sunlight and it is decided or change by golden section search method. during weather changing condition, the maximum power is not extracted so adjustment of the output voltage is required to extract maximum power. changing of the load is not in hand because it changes according to demand of consumers. in fig 5 the various parameter of the proposed method is shown and in fig-6 and fig-7 comparison of the proposed method is done with existing methods in all respect. the units of power and length are watts and voltage respectively. 120 s. singh copyright ©2018 assa. adv. in systems science and appl. (2018) iteration iteration fig.6 changing of parameters with an iteration of the proposed method. fig.7 comparisons of the proposed method with others table 2. iteration process chart. 6. conclusion in the present work, the maximum power point tracking is successfully carried out by using a golden search method. golden search method improves the efficiency and optimized the photovoltanic system without using a buck-boost converter compare to others methods of maximum power point technique. by changing the interval of voltage after each iteration, the length is decreasing by discarding the intervals in where the possibility of obtaining maximum power point is nil. this method is good in respect of accuracy and performance. in this paper, the temperature is removed by making it constant throughout the process by using the heat sink material layer at the back side of the pv panel. the performance is studied by the matlab/simulink. in future, the maximum power point tracking carried out by golden search method without using the buck-boost converter in order to reduce cost, improve response and efficiency, reduce complexity in making and complication of hardware can be removed. references [1] reisi, a. r., moradi, m. h. & jamasb, s. (2013.) classification and comparison of maximum power point tracking techniques for photovoltaic system: a review, renewable and sustainable energy reviews, 19, 433–443. [2] maxwell, j. c. (1881). a treatise on electricity and magnetism. clarendon press. mppt technique classification complexity tracking accuracy tracking speed efficient for partial shading application voltage-based mppt offline simple low slow no off grid current based mppt offline simple low slow no off grid perturb and observe online simple medium medium no both inc online complex very high fast yes both inr online complex high fast yes both rcc online complex high fast yes grid intelligent based hybrid complex very high fast yes both parasitic capacitance online complex very high medium yes off grid gradient descent hybrid medium high medium no off grid golden section search method (proposed) hybrid simple very high fast yes both maximum power point tracking technique 121 copyright ©2018 assa. adv. in systems science and appl. (2018) [3] chin, s., gadson, j., & nordstrom, k. (2003). maximum power point tracker. tufts university department of electrical engineering and computer science, 1-66. [4] ahmed, j., & salam, z. (2015). an improved perturb and observe (p&o) maximum power point tracking (mppt) algorithm for higher efficiency. applied energy, 150, 97-108. [5] subudhi, b., & pradhan, r. (2013). a comparative study on maximum power point tracking techniques for photovoltaic power systems. ieee transactions on sustainable energy, 4(1), 89-98. [6] ilyas, a., ayyub, m., khan, m. r., husain, m. a., & jain, a. (2018). hardware implementation of perturb and observe maximum power point tracking algorithm for solar photovoltaic system. transactions on electrical and electronic materials, 19(3), 222-229. [7] husain, m. a., & tariq, a. (2018). transient analysis and selection of perturbation parameter for pv-mppt implementation. international journal of ambient energy, (just-accepted), 1-8. [8] husain, m. a., jain, a., & tariq, a. (2016). a novel fast mutable duty (fmd) mppt technique for solar pv system with reduced searching area. journal of renewable and sustainable energy, 8(5), 054703. [9] husain, m. a., khan, a., tariq, a., khan, z. a., & jain, a. (2018). aspects involved in the modeling of pv system, comparison of mppt schemes, and study of different ambient conditions using p&o method. in system and architecture (pp. 285-303). springer, singapore. [10] ilyas, a., ayyub, m., khan, m. r., jain, a., & husain, m. a. (2017). realisation of incremental conductance the mppt algorithm for a solar photovoltaic system. international journal of ambient energy, 1-12. [11] agrawal, j., & aware, m. (2012, december). golden section search (gss) algorithm for maximum power point tracking in photovoltaic system. in power electronics (iicpe), 2012 ieee 5th india international conference on , 1-6. [12] ishaque, k., & salam, z. (2013), a review of maximum power point tracking techniques of pv system for uniform insolation and partial shading condition, renewable and sustainable energy reviews, (19), 475–488. [13] suresh, k. s., & prasad, m. v. pv cell based five level inverter using multicarrier pwm. international journal of modern engineering research, 1(2), 545-551. [14] funde, n., dhabu, m., deshpande, p., & patne, n. r. (2018). sf-oeap: starvationfree optimal energy allocation policy in a smart distributed multi-microgrid system. ieee transactions on industrial informatics. [15] husain, m. a., khan, z. a., & tariq, a. (2017). a novel solar pv mppt scheme utilizing the difference between panel and atmospheric temperature. renewable energy focus, 19, 11-22. [16] patrón, r. s. f., botez, r. m., & labour, d. (2012). vertical profile optimization for the flight management system cma-9000 using the golden section search method. in iecon 2012-38th annual conference on ieee industrial electronics society (54825488). [17] koupaei, j. a., hosseini, s. m. m., & ghaini, f. m. (2016). a new optimization algorithm based on chaotic maps and golden section search method. engineering applications of artificial intelligence, 50, 201-214. adv syst sci appl 2018; 4:13–38 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/624 a unified approach to reliability, availability, performability analysis based on markov processes with rewards valentina s. viktorova1∗, nikolai v. lubkov1 armen s. stepanyants1 1institute of control sciences, russian academy of sciences, moscow, russia received july 29, 2018; revised december 17, 2018; published december 31, 2018 abstract: this paper discusses a unified approach to reliability, availability and performability analysis of complex engineering systems. theoretical basis of this approach is continuous-time discrete state markov processes with rewards. from reliability modeling point of view complex systems are the systems with static and dynamic redundancy, imperfect fault coverage, various recovery strategies, multilevel operation and varying severity of failure states. we propose a unified method of calculating the reliability, availability and performability indices based on the definition of special forms of reward matrix. this method is proved to be effective in calculating both cumulative and instantaneous measures in steady-state and transient cases. we describe special analytical software which implements suggested method. we demonstrate the flexibility of the proposed method and software by analyzing multilevel process unit with protection and demand-based warm standby system. keywords: availability, markov process, markov reward model, multilevel engineering systems, performability, reliability. 1. introduction at the present state of development of reliability theory the models of analysis can be subdivided into two categories (classes): the dynamic models and the static models. all changes of system states in the class of dynamic models are considered as the processes developing in time. in the class of the static models, states of the system are determined by the sets of states of system elements at the time moment t. markov [1–4], semi-markov processes, asymptotic methods of the renewal theory and regenerative processes [5–7], monte carlo simulation techniques [8] are used within the framework of dynamic models. dynamic models allow to calculate all the main dependability measures both for repairable and non-repairable systems. these measures are: instantaneous indices (e.g. availability at the time instant); interval indices (e.g. reliability during the time interval); time-independent stationary indicators (e.g. mean time between failures). known drawbacks inherent in markov models are the size and stiffness of transient solution of a system of kolmogorov-chapman equations (or ill-conditionality for stationary case). possible ways to solve these problems are discussed in [9–11]. monte carlo simulation of modern high reliable systems may require large amount of simulations to obtain results with desired accuracy. these shortcomings can be eliminated with the help of special acceleration techniques. there is no limitation to type of distribution of waiting times between the changes in the class of monte carlo models. but the problem of creating a universal monte carlo ∗corresponding author: vsviktorova@gmail.com 14 v.s. viktorova, n.v. lubkov, a.s. stepanyants model describing the complex reliability behavior of different kind of systems remains unsolved. static models use two main groups of methods – combinatorial-probabilistic [3] and logical-probabilistic [12]. combinatorial-probabilistic methods use the basic formulas of combinatorics and probability theory (probability of the sum and product of events, the formula of total probability). these formulas are used mainly for serial-parallel and ”m out of n” redundant schemes. logical-probabilistic methods, which are based on the construction of a boolean function relating the state of the system with the states of its elements. the resulting boolean function is transformed to a form that allows to replace logical variables by corresponding probabilities. classic failure trees and reliability block diagrams are the main methodology of logical-probabilistic methods [13, 14]. classic static models for the case of repairable systems allow us to calculate only the differential (instantaneous) reliability indicators determined at the time instant t (e.g. availability, failure frequency, average efficiency at the time instant t). in the last decade, combined approaches to reliability models construction are successfully developing. these approaches are based on the implementation of dynamic properties into static models. dynamic fault trees are the most popular practical implementation of these approaches [15–19]. modern complex engineering systems with high demands on reliability are characterized by various features, which include: • multiple levels of operation efficiency (e.g., performance) and the graceful degradation of performance in the event of failures [20–23]; • a variety of redundancy implementation techniques (hardware, software, functional, temporal redundancy) [24–26], types of redundant schemes (parallel, standby, hybrid), modes of reserve components (cold and warm standby, shared) [3, 4]; • smart recovery strategies, restricted repair resources [27–29]; • imperfect built-in test equipment leading to the presence of undetected failures and false alarms [30, 31]; • implementation of special multiphase error handling procedures with classification on permanent and transient faults (mainly for fault tolerant computers) [32–34]; • the possibility of the occurrence of mutually exclusive failures that, with a certain multiplicity and sequence, can lead to different consequences at the system level [13,14]. markov modeling allows taking into account the above-mentioned features of the reliability behavior of systems and calculating most part indicators of reliability, safety, efficiency. at the design stage, we don’t have, as a rule, objective information about the distribution functions of random time variables of the model. therefore, we believe that the assumption of exponentiality is entirely permissible. the problem of large size of the markov model can be solved by decomposition and aggregation of the model parts and by automation of the model construction. in this paper, we consider the markov reward models as common tools for analyzing complex reliability behavior of technical systems. we propose the extension of markov reward modeling for the reliability, abailability, performability (rap) analysis of the systems at the design stages. we suggest the method of calculating reliability, availability and performability indicators based on a special definition of the reward matrix. this method is quite efficient and well suited for software implementation. it makes calculation both cumulative and point indices possible only by solving systems of differential or algebraic equations and does not require the usage of additional numerical integration operations. the remainder of this paper is organized as follows. a brief summary of the mathematical foundations of markov reward model is given in section ii. in section iii, firstly we give definitions of interval, point, steady-state rap measures. next, we present the method for for calculation of these measures in the framework of markov reward model. section iv presents analytical software implementing the proposed method. the flexibility of described copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 15 techniques and software is illustrated by investigations of two practical examples in section v. in section vi, we give some concluding remarks. 2. mathematical foundations a system state at the time moment t can be described by stochastic process z(t) of transitions in system state space ω. we assume that the state space ω is discrete, finite and of sizen . ω is defined by states of system components (failed or operational). we define a quality function φ on z(t). changing system states leads to a change in the quality function φ. z(t) is stochastic process, therefore, φ(z(t)) = φ̃(t) are also stochastic functions. due to the discrete nature of ω, the functions φ̃(t) are the step functions, whose values are quality levels (for example, system performability) corresponding to the states. g are indices of system performability. g can be defined as a measure on the trajectories of the φ̃(t) . general representation of g takes the form g = m{f [φ(z(t)]}, (2.1) where m represents mathematical expectation; f is a functional, which is determined by the type of performability index. thus, the model of performability contains three components ω, z(t),φ. it is natural and obvious to present the model of performability by means of transition graph (state graph). the elements of the state space correspond to the vertices. possible transitions in the state space correspond to the arcs of the graph. φ is defined as a function of the system state si. this function takes the value of wi corresponding to the performance level of the system in this state. that is φ(si) = wi. (2.2) in addition, this model can display effects occurring at the transition of the system from one state to another. this is accomplished by determining reward matrix w = [wij] n n , (2.3) where wij – is an effect on the system that arises at the transition from state i to state j; wii = wi. random transition process z(t) is most simply set in the case where the value of time of stay in each state have an exponential distribution, i.e. when the process is a continuous-time markov chain (ctmc). the process z(t) is determined by transition rates, which correspond to arcs of the markov graph. similarly to the case of matrix w , we can create a matrix of transition rates λ = [λij] n n , (2.4) where λij – is the rate of transition from state i to state j; λii = − ∑ j,j 6=i λij . 2.1. construction of rap indices thus, the markov model of performability is defined by two matrices λ and w , i.e. {r, z(t),φ} ∼ {λ,w}. model {λ,w} is called markov reward model (mrm). markov reward model describes the behavior of systems with exponential distribution of elements failure and repair time. states of the model represent the states of the system, which correspond to different sets of failed and operable items. if criterion of system failure is specified, then expression (2.1) defines also the indices of reliability and availability. copyright c© 2018 assa. adv syst sci appl (2018) 16 v.s. viktorova, n.v. lubkov, a.s. stepanyants let failure criterion is {φ̃(t) < φcr} ⇒failure, where φ̃(t) is the current value of the quality function; φcr is limit value of the quality, which is allowed by operating requirements. then, by setting a particular type of the functional f the following indices can be determined from equation (2.1): • reliability over the time interval (0, t)r(t): r(t) = p{φ̃(τ) > φcr, ∀τ ∈ (0, t)} under f [·] = { 0, if ∃τ ∈ (0, t) : φ̃(τ) < φcr 1, otherwise (2.5) • mean time to first failure (mttf) tff : tff = ∞∫ 0 r(t)dt underf [·] = min{τ : φ̃(τ) < φcr} (2.6) • availability a(t) at the time instant t: a(t) = p{φ̃(τ) ≥ φcr} under f [·] = { 0, φ̃(τ) < φcr 1, φ̃(τ) ≥ φcr (2.7) • probability of being at the ith level of performance at the time instant t: p (t) = p{φ̃(t) = φi} under f [·] = { 1, φ̃(t) = φi 0, otherwise (2.8) • average accumulated time spent at the ith level of performance: under f [·] = ∫ t:φ̃(t)=φi dt (2.9) • average number of transitions from the ith to the jth performance level during the time interval (0, t): under f [·] = ∑ l : φ̃(tl−) = φi φ̃(tl+) = φj 1, (2.10) where tl the time moment of transition from ith level to jth performance level. • time-averaged accumulated reward for the time interval (0, t): f [·] = 1 t t∫ 0 φ̃(τ)dτ (2.11) 2.2. description of mrm basic relations the consept of reward in markov reward model is a generalized concept. this can be any effect, loss, costs. according to the provisions of previous section we define the model state space as a discrete set {si}, i = 1, n . also we define reward matrix. the elements of the reward matrix are interpreted as follows: wii is a reward at time unit while the system is copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 17 in state si (reward rate); wij is impulse reward received in the system under the transition from state i into state j; n the number of states of the system. we define the total (or accumulated [20]) reward obtained by the system for a given time t as the functional f [φ(z(t))]: g̃(t ) = n∑ i=1 τiwii + ∑ i,j nijwij, (2.12) where τi is the residence time in state si during the time interval (0, t ); nij is the number of system transitions from state i into state j during the time interval (0, t ). accumulated reward g̃(t ) is a random variable, since τi and nij are random values. average accumulated reward is defined as the mathematical expectation of g̃(t ) y (t ) = m{g̃(t )}. (2.13) according to (2.1), y (t ) is an estimate of the performability index value. if we consider that members of (2.12) τi and nij depend only on the specific implementation of process z(t), it is obvious that the type of index is completely determined by the type of matrix w . the system of howard differential equations describes the behavior of the average accumulated reward: dyi(t)/dt = wii + ∑ j,j 6=i λijwij + ∑ j λijyj(t); i, j = 1, n, yi(0) = 0, (2.14) where yi(t) is the average accumulated reward for the time interval (0, t), given that the initial state is si(z(0) = si); λij the rate of transition from state si into state sj . the matrix form of equation (2.14) is dy (t)/dt = λy (t) +r; y (0) = 0, (2.15) where λ = ‖λij‖ matrix of transition rates; r is a column vector of constant terms: ri = wii + ∑ j,j 6=i λijwij; i, j = 1, n. (2.16) we assume that for absorbing states sr of the markov process valuewrr = 0. the absorbing state is a state that once entered, can not be left. this assumption is quite natural, since the absorbing states are usually identified with the non-operational states of the system. 3. method of calculating rap indices based on markov reward model the expression (2.1) of expected reward will determine the rap indices in accordance with (2.5) – (2.11), if the value of the reward for a particular process realization (2.12) coincides with the functional f[.] in (2.5) – (2.11). it requires to select the appropriate reward matrix w and optionally to correct the transition rate matrix λ. correction of matrix λ is necessary in case when the index reflect operation of the system till some event e, for example, till the system failure. since all performability indices in the model are interpreted through the reward, it is obvious that reward should be dismissed after event e has occurred. formally, this can be achieved by identification of the event e with a transition to an absorbing state. the reward associated with this state is assumed to be zero. the correction of matrix λ is that we treat the states corresponding to e as absorbing states. to do this correction, it is enough copyright c© 2018 assa. adv syst sci appl (2018) 18 v.s. viktorova, n.v. lubkov, a.s. stepanyants to equate to zero the elements of rows of matrix λ corresponding to these states. to carry out further calculations, we partition the set of model states ω into ωg, the set of operational system states, and ωf , the set of failed states. all rap measures can be subdivided into three groups: 1. interval measures these metrics are dependent on time and are determined on finite time interval [0, t ], for example, reliability r(0, t ). 2. point measures these metrics are dependent on time and are determined in time point t, for example, availability a(t). 3. steady-state measures these metrics are independent on time and are determined on an unlimited time interval (0,∞), for example, mttf. calculation of steady-state measures is reduced to finding the stationary solutions of the system (2.15), i.e., to solving a system of linear algebraic equations: λy +r = 0. (3.1) we mention that all indices calculated by the expressions (2.15) and (3.1) are vectors. the ith element of this vector is the estimated value of the index for the initial state si. 3.1. unreliability system unreliability over the time interval (0, t)(q(t)) can be calculated if we define elements of matrix w as wij = 0 ∀ i, j ∈ ωg;wij = 1 ∀ i ∈ ωg and j ∈ ωf : wij = { 1, if i ∈ ωg, j ∈ ωf 0, otherwise . (3.2) besides we must treat states in ωf as absorbing states. namely we should to delete in markov graph all arcs leading from states in ωf to states in ωg. the calculation is performed by (2.15). complementary index is reliability: r(t) = 1−q(t). 3.2. mean time to first failure mean time to first failure (tff ) can be calculated if we define diagonal elements of matrix w corresponding to the operational states as wii = 1 ∀ i ∈ ωg. wij = { 1, if i = j, i, j ∈ ωg 0, otherwise . (3.3) besides we must treat states in ωf as absorbing states. calculation is performed by (3.1). for the given type of reward matrix expression (3.1) is reduced to λ∗tff =  −1 ... −1  , (3.4) where λ∗ − k × k matrix derived from matrix λ by deleting rows and columns corresponding to nonoperational states; k total number of operational states. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 19 3.3. standard deviation of random time to first failure standard deviation (σ) and variance (σ2) of mean time to first failure can be calculated in accordance with the expression σ = √ m(t2)−m2(t) = √ t (2) ff − t 2 ff . to calculate vector t (2) ff , we have to find the solution of algebraic equations λ∗t (2) ff =  −2tff1 ... −2tffk  , (3.5) where tffi is mean time to first failure given the process starts from the ith state. 3.4. average accumulated time spent in the operational states average accumulated time spent in the operational states during the time interval (0, t) (t∑(t)) can be calculated if we define diagonal elements of matrix w corresponding to the operational states as wii = 1 ∀ i ∈ ωg. wij = { 1, if i = j; i, j ∈ ωg 0, otherwise . (3.6) here, we should not artificially create absorbing states. calculation is performed by the equation (2.15). we can also use this index for analysis of multilevel systems. in this case we should partition the set of system’s states ω into m sets ω1, . . .ωr, · · · ,ωm, where ωr the set of operational system states with rth performance level. average accumulated time spent in ωr during the time interval (0, t) (t r∑) can be calculated at wii = 1 ∀ i ∈ ωr. 3.5. availability availability a(t) is the probability that the system is operational at the time instant t. a(t) is the point measure. taking into account that a(t) = dtς(t)/dt we can calculate a(t) after calculation of tς(t) as follows: a(t) = λtς +r. (3.7) the formation of reward matrix is carried out in accordance with (3.6). unavailability u(t) = 1− a(t). 3.6. time-averaged availability general expression of time-averaged availability over the time interval (0, t) (aav(t)) has the form aav(t) = 1 t t∫ 0 a(τ) dτ. (3.8) for calculation of this indicator, it is necessary to take the reward matrix, in which the diagonal elements of the operational system states, are equal to 1/t and all other elements are zero: wij = { 1/t, if i = j; i, j ∈ ωg 0, otherwise . (3.9) artificial creation of absorbing states is not required. calculations are carried out by (2.15). copyright c© 2018 assa. adv syst sci appl (2018) 20 v.s. viktorova, n.v. lubkov, a.s. stepanyants 3.7. average failures number average (expected) failures number during the time interval (0, t)(n(t)) can be calculated if we define elements of matrixw columns corresponding to the failed states aswij = 1 ∀ i 6= j. values of other elements of matrix w are zero: wij = { 1, if i ∈ ωg, j ∈ ωf 0, otherwise . (3.10) n(0, t) is failure measure for repairable systems so we should not artificially create absorbing states. calculation is performed by (2.15). 3.8. failure frequency failure frequency (ω(t)) is a differential index with respect to n(t). after calculating of vector n(t) for each initial state, we put it into the right side of (2.15). the resulting vector dy (t)/dt will be the vector (ω(t)): ω(t) = λn(t) +  λ1 σ ... λnς  , (3.11) where λiς is total transition rate from ith operational state to failure states. for the failure state, it is zero. 3.9. average accumulated reward average accumulated reward over the time interval (0, t)(et (t)) integrates reward (loss) of the system during the time interval (0, t) proportionally to the time of stay in the states and the quantity of transitions between the states. reward matrix w has the most general form. the elements wii, wij are real rewards (losses) measured in terms of units of system performance. calculations are carried out by (2.15). 3.10. average reward average reward at the time instant t(eav(t)) is a point measure. this indicator characterizes the multi-level systems. therefore, system states set is subdivided into classes, in accordance with level of system performance. states of the ith class are characterized by performance (ei) per time unit (for example, throughput). the first step of eav(t) calculation is performed by the equation (2.15). as a result of the first calculation step, we get the vector of the average accumulated reward et (t). eav(t) is a differential index with respect to et (t). so the second calculation step is to put the vector et (t) into right side of equations (2.15): eav(t) = λet (t) +r. (3.12) 3.11. performability ratio this index is the ratio of the average accumulated reward to the nominal reward during the time interval (0, t): kp(t) = et (t) en(t) , (3.13) where en(t) nominal (maximum) accumulated reward generated by the system in the absence of failures. en(t) = emaxt. usually emax = e1. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 21 4. analytical software based on markov reward model the above uniform approach to calculating performance, availability and reliability measures was used by us in the development of analytical software. the software is operated under windows os and is written in c] programing language. it has an advanced graphical user interface (fig.4.1) and uses ole automation for generation reports in ms word formats. this software allows you to take into account all features of reliability behavior of complex technical systems, mentioned in sectoion 1: • multiple levels of operation performance and the graceful degradation of performance in the event of failures; • different types of redundancy implementation techniques; • variety of redundant schemes and backup components load levels; • recovery strategies and restrictions on repair resources; • imperfect built-in test; • possibility of transient faults occurrence; • common cause failures; • and so on. the user of this software can calculate all the rap indicators described in section 3: • reliability (unreliability) for a given time interval (0, t) ; • mean and standard deviation of random time to first failure; • average and accumulated time spent in the selected subset of states during the time interval (0, t); • availability (unavailability) at given time instant t; • average failures number for a given time interval (0, t); • failure frequency at given time instant t; • average reward at given time instant t; • average accumulated reward for a given time interval (0, t); • performability ratio for a given time interval (0, t). the software execution process is divided into three steps: creation of the mrm, setting up of the mrm, numerical solution of the system of howard equations. the composition and interrelationships of these processes are shown in fig.4.2-fig.4.4. mrm creation is executed by the user in interactive mode with the help of built-in graphical editor. the created model can be saved in an external xml file. a saved model can be reloaded from the xml file (fig.4.2). the model’s setting up is performed in the matrix configuration module in accordance with equations of section 3. setting up operations are performed automatically after the user chooses the metric for calculation (fig.4.3). calculation of values of user-selected rap indices is made in the numerical solution module (fig.4.4). in this module, numerical procedures for solving systems of differential and algebraic howard equations are realized. different approaches and methods of computational mathematics with regard to their software implementation outlined in [35 37]. detailed consideration of numerical evaluation of markov models transient behavior was presented in [9], [10]. three numerical techniques for finding the transient solution of large markov models were examined. they are uniformization, an explicit differential equation solution method (runge-kutta) and special stable implicit method (tr-bdf2). for the software implementation of numerical solution of the mrm we used efficient method based on evaluation of matrix exponential at small step. this method had been proposed and described in [38]. for a stationary case, a numerical solution is obtained by solving the system of algebraic equations by the gauss method or the rotation method [35,36]. the rotation method works well for bad conditioned systems, which are often generated by rap analysis models. copyright c© 2018 assa. adv syst sci appl (2018) 22 v.s. viktorova, n.v. lubkov, a.s. stepanyants fig. 4.1. screenshot of main form of the analytical software. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 23 fig. 4.2. execution process of the analytical software (step 1). copyright c© 2018 assa. adv syst sci appl (2018) 24 v.s. viktorova, n.v. lubkov, a.s. stepanyants fig. 4.3. execution process of the analytical software (step 2). copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 25 fig. 4.4. execution process of the analytical software (step 3).copyright c© 2018 assa. adv syst sci appl (2018) 26 v.s. viktorova, n.v. lubkov, a.s. stepanyants 5. case studies to demonstrate the effectiveness of the software implemented mrm, we present two examples of reliability, availability, performability evaluation of complex systems. in first example we investigate technological object with protection system. the second one is demand-based warm standby system (db-wss), described in article [23]. 5.1. markov reward model of multilevel process unit we consider process unit consisting of technological object (to) with protection system (ps). the protection system includes a diagnostic unit (du) and the actuator (a) (fig.5.5). fig. 5.5. technological object with protection. protection systems are designed to generate control actions to the protected object in order to prevent transition from process equipment failures to accident. control actions may be different, such as change of the operating mode, reduction in productivity, emergency shutdown of faulty elements and elements belonging to the same processing chain. these actions prevent the development of accident. here, we consider the case of technological process shutdown. however, proposed approach is also suitable for the control actions that lead to poor performance or mode change. the main function of protection system is performed sporadically at the time instants of the occurrence of the object failures, so that the protection system operates in standby temporary mode. let us define rewards of operational states and losses of failed states. in this example, all possible system states can be subdivided into four classes. the first class includes such technical states in which the object is functioning normally and earns specific rewards per unit of time of stay in these states (e.g., these states correspond to the nominal performance of technological object). let us note that this group includes states with latent failures such as failure of diagnostic unit. the second class consists of the states of accident-free shutdown of the technological object. the losses associated with these states are only related to the downtime of the object. the reward in these states is either zero or negative, if the idle leads to additional losses per unit time. the third and fourth classes of states are catastrophic failures of the object. the transition to this class of states from the states of the first class brings a onetime damage (negative reward) associated with the occurrence of an accident (death of people, equipment breakdowns, emissions into the atmosphere, etc.). in this model, we consider two types of catastrophic failures differing in the severity of the consequences – accident i and accident ii. thus, the normal functioning of the technological object is accompanied by a linear increase in the accumulated reward in proportion to the time spent in the first class states. idle time leads to the preservation of the achieved level of accumulated reward (at zero values of the reward rate in each state) or to its descent (at negative values of the reward rates) in proportion to the time spent in states of second, third and fourth classes. when impulse rewards associated with the transitions between the states are not zero, there is an abrupt change (more often a decrease) in the accumulated reward. negative impulse rewards are due to the costs of recovery from failures or accidents, purchase of equipment, payment of fines or insurance, etc. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 27 we suppose that all accidents occur only due to failures of the technological equipment. let protection system in the event of critical situation instantly shut down the technological process by making the necessary equipment control (for example, switching off). let the protection system processes critical situation with coverage β(0 ≤ β ≤ 1). failures of the protection system which occur during interval of normal to operation can lead to different consequences. let us single out ps failures of two types: latent failures (no operation) and explicit failures (false alarm). latent failures of ps are manifested in the form of absence of protection actions in the event of critical situation which entails an accident. false alarms lead to the undesired protection actions in the absence of to failures. parameters of the markov reward model of the technological object with protection system are: wij impulse losses (negative impulse reward) under the transition from state i into state j; wii reward rate (negative or positive) in state i; β coverage probability for to, determined as conditional probability of processing of to catastrophic failure by a protection system, provided that the failure has occurred; αdu fraction of the protection system latent failures such as non operation of du; αa fraction of the protection system latent failures such as non triggering of actuator; 1− αdu fraction of the protection system explicit failures such as du false alarm; 1− αa fraction of the protection system explicit failures such as actuator false triggering; ηi fraction of technological object catastrophic failures of the first kind (accident i); ηii fraction of technological object catastrophic failures of the second kind (accident ii); λto, λdu , λa – failure rates of technological object, diagnostic unit, actuator respectively; µ repair rate of technological object after shutdown; µa repair rate of technological object after falling into accident ii state. the transition graph of markov reward model of the technological object with protection is shown in 5.6. we partition the set of the mrm states into four subsets: • normal operation (op) (states 1,3,4) • accident-free shutdown (sd) (states 2,6) • accident i (ai) (state 5) • accident ii (aii) (state 7) transition rates matrix λ of the model is λ =  λ11 λ12 λ13 λ14 λ15 0 λ17 λ21 λ22 0 0 0 0 0 0 λ32 λ33 0 λ35 0 λ37 0 0 0 λ44 λ45 λ46 λ47 λ51 0 0 0 λ55 0 0 λ61 0 0 0 0 λ66 0 0 0 0 0 0 0 λ77  . (5.1) the diagonal elements of the matrix λii are given by − ∑ j,j 6=i λij . transition rates between states of the graph are determined based on the model parameters as follows: λ12 = [1− (1− β)(ηi + ηii)]λto + (1− αdu)λdu + (1− αa)λa; λ13 = αduλdu ; λ14 = αaλa; λ15 = (1− β)ηiλto; λ17 = (1− β)ηiiλto; λ21 = λ61 = µ; λ32 = (1− ηi − ηii)λto + (1− αa)λa; copyright c© 2018 assa. adv syst sci appl (2018) 28 v.s. viktorova, n.v. lubkov, a.s. stepanyants λ35 = λ45 = ηiλto; λ37 = λ47 = ηiiλto; λ46 = (1− ηi − ηii)λto; λ51 = µa; fig. 5.6. markov graph of to with protection. by specifying the form of reward matrix, it is possible to obtain equations for calculating various indices. we want to determine the cumulative effect of operation of technological object taking into account the positive and negative impact of the protection system. therefore, it is advisable to calculate the following indices: 1. q(t) unreliability over (0, t). q(t) can be calculated from equation (2.15). to calculate q(t), we must equate to zero λ21, λ51, λ61 and define w as w =  0 1 0 0 1 0 1 0 0 0 0 0 0 0 0 1 0 0 1 0 1 0 0 0 0 1 1 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0  . (5.2) 2. tς(t) -average accumulated time spent in the operational states over the time interval (0, t). tς(t) is determined by (2.15), setting r as copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 29 r = [ 1 0 1 1 0 0 0 ]t . (5.3) 3. a(t) system availability at the time instant t. a(t) is calculated by (3.7) after calculation of vector tς(t). 4. tff mean time to first failure. tff is determined by (3.4), where matrix λ∗ is λ∗ = [ λ11 λ13 λ14 0 λ33 0 0 0 λ44 ] . (5.4) 5. σ standard deviation of random time to first failure. σ is obtained from equation (3.5), where vector of free terms is [ −2tff1 −2tff2 −2tff3 ]t . 6. ta mean time to first catastrophic failure (accident i or ii). ta is determined by equation (3.4), where matrix λ∗ is 5× 5 matrix derived from matrix (5.1) by deleting rows and columns corresponding to accident i and accident ii states (states 5 and 7). 7. n(t) expected failures number. n(t) can be calculated from (2.15) where matrix w takes the form (5.2). 8. ω(t) failure frequency. ω(t) can be calculated from (3.11) after n(t) calculation and setting free terms vector as [ λ12 + λ15 + λ17, 0, λ32 + λ35 + λ37, λ45 + λ46 + λ47, 0, 0, 0 ]t . 9. na(t) expected catastrophic failures number. na(t) can be calculated from (2.15) where matrix w has unit elements only in the transitions to catastrophic failures (states 5 and 7). similarly, we can find separately nai(t) and naii(t). 10. tσsd(t) –average accumulated time spent in the states of technological object shutdown over time interval (0, t). tσsd(t) is determined by equation (2.15), setting r as [ 0, 1, 0, 0, 0, 1, 0 ]t . 11. nsd(t) average number of transitions to states of shutdown. nsd(t) can be calculated from equation (2.15) where matrix w has unit elements only in the transitions to shutdown states (states 2 and 6). 12. et (t) average accumulated reward over the time interval (0, t). calculation of et (t) is performed by (2.15). matrix of reward w is w =  w11 w12 w13 w14 w15 0 w17 w21 w22 0 0 0 0 0 0 w32 w33 0 w35 0 w37 0 0 0 w44 w45 w46 w47 w51 0 0 0 w55 0 0 w61 0 0 0 0 w66 0 0 0 0 0 0 0 w77  . (5.5) 13. eav(t) average reward at the time instant t. in accordance with (3.12) eav(t) is determined by solving the system of algebraic equations  eav1(t) ... eav7(t)  = λ  et1(t) ... et7(t) +  w11 ... w77(t) +  ∑ j 6=1 λ1jw1j ...∑ j 6=7 λ7jw7j  . (5.6) copyright c© 2018 assa. adv syst sci appl (2018) 30 v.s. viktorova, n.v. lubkov, a.s. stepanyants the numerical results of above indices calculation for different values of parameter β are summarized in table 5.1. all time-dependent metrics were calculated on one year period (t = 8760hours). we note once again that the solution, obtained by the mrm, is a vector whose ith component corresponds to the start of the markov process from state i. the table and diagrams include the first components of solution vectors, i.e the system starts from completely operational state s1. the functions of et (t), eav(t) for different values of catastrophic failure coverage parameter β are shown in fig.5.7 and fig.5.8 respectively. calculation of indices and plotting are performed with the following system parameters: λto = 0.0001(/hour);λdu = 0.00001(/hour);λa = 0.00005(/hour). µ = 0.1(/hour);µa = 0.004(/hour). αdu = 0.5;αa = 0.9. ηi = 0.4; ηii = 0.05. numerical values of impulse rewards are w15 = −105;w17 = −108;w12 = w13 = w14 = 0; w21 = 0;w32 = 0;w35 = −105;w37 = −108; w45 = −105;w46 = 0;w47 = −108; w51 = w61 = 0. numerical values of reward rates are w11 = 15(/hour);w33 = 10(/hour);w44 = 10(/hour); w22 = w55 = w66 = w77 = 0(/hour). fig. 5.7. to average accumulated reward et (t) versus time t for different coverage values β. the first family of curves (fig.5.7) shows the significant dependence of the average accumulated reward on the catastrophic failure coverage β. for the model parameters under consideration β should be greater than 0,9. the second family of curves(fig.5.8) shows a sharp decline in the growth of the average accumulated reward in time for all coverage values. it occurs due to the high rate of ps latent failures. to prevent such a recession, it is advisable copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 31 fig. 5.8. to average reward eav(t) versus time t for different coverage values β. table 5.1. calculation results over one year period. catastrophic failure coverage β index 0.8 0.85 0.9 0.95 1 q(t) 0.612 0.612 0.612 0.612 0.612 tς(t) 8670.6 8682.4 8694.3 8706.2 8718.1 a(t) 0.9814 0.9836 0.9858 0.988 0.99 tff 9360.1 9360.1 9360.1 9360.2 9360.1 σ 9555.6 9555.6 9555.6 9555.6 9555.6 ta 4774.7 5143.4 5573.8 6082.8 6694.1 n(t) 0.942 0.943 0.944 0.946 0.947 na(t) 0.124 0.108 0.091 0.075 0.058 nai(t) 0.11 0.096 0.081 0.0663 0.0516 naii(t) 0.014 0.012 0.01 0.0083 0.0064 tσsd(t) 8.16 8.34 8.52 8.7 8.88 nsd(t) 0.817 0.835 0.853 0.871 0.889 et (t) -15489 3146 21834 40575 59368 eav(t) -6.068 -4.202 -2.33 -0.443 1.45 to carry out preventive maintenance of the process unit (to+ps), during which latent failures of the protection system are detected. 5.2. reliability and performability investigation of db-wss we take demand-based warm standby system (db-wss), described in section iv.a of article [23] as the second object of study. this db-wss is composed of two components a1 and a3 (we keep the numbering made by the authors of [23]). the components a1 and a3 can be in four degradation states, which differ in capacity. the components capacity levels are 5, 4, 2, 0 and 5,4,1,0, respectively. component a1 is online. component a3 can be online or in warm standby state depends on the system demand. we copyright c© 2018 assa. adv syst sci appl (2018) 32 v.s. viktorova, n.v. lubkov, a.s. stepanyants consider scenario b from [23], when demand determined as 5. so initial state of a3 is warm standby. the transition time distributions for the degradation processes for each component are exponential distributions with parameters λ1 = 1/100, λw3 = 1/300, λ3 = λo3 = 1/150 (/day), where (w) indicates warm standby state and (o) – operating online state. first, we will construct a markov model completely corresponding to the degradation process described in [23]. namely, we will construct a model of a binary-state system consisting of multistate components. the transition markov graph of this db-wss is shown in fig.5.9. the states s1 and s10 are operational states of the system that differ in the capacity of components. the state (j, k) corresponds to the jth capacity level of element a1(ca1 j ) and the kth capacity level of element a3(ca3 k ). state s11 is failed system state. ca1 j + ca3 k ≥ 5 for operational states; ca1 j + ca3 k < 5 for failed state. fig. 5.9. markov graph of binary-state db-wss. for the first case, we calculate the following reliability indicators: 1. r(t) reliability during (0, t). unreliability q(t) can be calculated from equation (2.15), where vector r = [ 0 0 0 0 0 0 λ3 λ1 + λ3 λ1 + λ3 λ1 0 ]t . reliability r(t) = 1−q(t). table 5.2 shows the results of our calculation of r(t) and values of r(t) taken from [23]. for the model under study (fig.5.9), an exact analytical solution was obtained (see appendix). the results of r(t) calculations using (a.2), (a.3) are given in the last column of table 5.2. 2. tff -mean time to first failure. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 33 table 5.2. calculation results of system reliability. r(t) time [day] mrm method proposed in [23] monte carlo from [23] exact analytical solution 50 0.99317 0.9917 0.9931 0.993170 100 0.93879 0.9333 0.9385 0.938794 200 0.67104 0.6610 0.6730 0.671043 300 0.38090 0.380899 400 0.18958 0.189584 500 0.08858 0.0864 0.0884 0.088584 the tff is determined in accordance with (3.4) tff=284.61 [day]. 3. σ standard deviation of random time to first failure. the σ is obtained from (3.5), where vector of free terms is [ −2t1 −2t2 −2t3 −2t4 −2t5 −2t6 −2t7 −2t8 −2t9 −2t10 ]t . σ = 154.05 [day]. next, we assume the possibility of a multi-level system operation. we assume that the states of the system with a capacity less than 5 can be subdivided into groups (levels) with capacity {4, 3, 2, 1, 0}. the model of the multistate system consisting of multistate components is shown in fig.(5.10). the properties of the levels are given in table 5.3. table 5.3. the groups properties and measures for multistate model levels the group capacity t i σ 1 s1,s2,s3,s4,s5,s6,s7,s8,s9,s10 5 273.10 2 s11,s13 4 63.48 3 s12 3 17.75 4 s15 2 16.38 5 s14 1 52.43 6 s16 0 76.86 for the second case, we calculate the following reliability and performability indicators: 1. t iς average total residence time of the system in the states of the ith group during (0, t). the t iς is determined in accordance with section 3.4 . for example, t 2 σ is calculated by (2.15) with vector r = [ 0 0 0 0 0 0 0 0 0 0 1 0 1 0 0 0 ]t . last column of table 5.3 shows the results of calculation of t iς for t = 500[day]. 2. kp(t) performability ratio. the kp(t) is calculated in accordance with (3.13). to calculate performability ratio, we should determine the reward matrix. for this system, the reward matrix will consist only of diagonal elements. we will define reward rates through the capacities, corresponding to the states of operation and degradation as wii = ci/(cmaxt). for cmax = 5 and t = 500 [day] we have w1 = · · · = w10 = 0.002;w11 = w13 = 0.0016;w12 = 0.0012;w14 = 0.0008;w15 = 0.0004;w16 = 0. as a result of solving the system of equations (2.15), we obtain kp(500) = 0.70315. this performability index is useful in choosing the best project version of a multi-level system, provided that the reliability requirements are met. in this case, the reliability of the system during the time interval (0,500) [days] is extremely small (see last raw of table 5.2). however, the value of the performability index is quite satisfactory, which confirms the correctness of the choice of the redundancy scheme. copyright c© 2018 assa. adv syst sci appl (2018) 34 v.s. viktorova, n.v. lubkov, a.s. stepanyants fig. 5.10. markov graph of multistate db-wss. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 35 6. conclusion this paper presents a unified approach for the estimation of reliability, availability, performability of engineering systems. the approach uses markov process with rewards to model transient and steady-state behavior of the systems. by using markov reward models, we can derive various instantaneous and cumulative measures and estimate wide range of indices. we present straightforward calculation technique for evaluation of reliability, availability and performability indices based on special definition of reward matrix. the technique implementation requires only solution of system of differential equations describing the behavior of the average reward. any additional calculations, such as numerical integration, are not required. based on this approach a software for reliability, availability, performability analysis was created. the software is operated under windows os and is written in c] programing language. it has an advanced graphical user interface, uses effective numerical method for solving mrm and applies ole automation for generation reports in ms word formats. the efficiency of the proposed approach and accuracy of numerical solution have been shown by the case studies of the technological system with protection and demand-based warm standby system. 7. appendix in this section, we present analytical solution of the model of a binary-state system consisting of multistate components (fig.5.9). the column vector p(t) of the state probabilities is calculated by solving a system of kolmogorov-chapman equations: p′1(t) = −p1(t) · (λ1 + λw3 ); p′2(t) = p1(t) · λ1 − p2(t) · (λ1 + λ3); p′3(t) = p1(t) · λw3 − p3(t) · (λ1 + λw3 ); p′4(t) = p2(t) · λ1 − p4(t) · (λ1 + λ3); p′5(t) = p2(t) · λ3 + p3(t) · λ1 − p5(t) · (λ1 + λ3); p′6(t) = p3(t) · λw3 − p6(t) · (λ1 + λw3 ); p′7(t) = p4(t) · λ1 − p7(t) · λ3; p′8(t) = p4(t) · λ3 + p5(t) · λ1 − p8(t) · (λ1 + λ3); p′9(t) = p5(t) · λ3 + p6(t) · λ1 − p9(t) · (λ1 + λ3); p′10(t) = p6(t) · λw3 − p10(t) · λ1; (7.1) the solution is copyright c© 2018 assa. adv syst sci appl (2018) 36 v.s. viktorova, n.v. lubkov, a.s. stepanyants p1(t) = −e−(λ1+λw3 )t; p2(t) = λ1 λ3 − λw3 (e−(λ1+λw3 )t − e−(λ1+λ3)t); p3(t) = λw3 te −(λ1+λw3 )t; p4(t) = ( λ1 λ3 − λw3 ) 2 (e−(λ1+λw3 )t − e−(λ1+λ3)t)− λ2 1 λ3 − λw3 e−(λ1+λ3)t; p5(t) = λ1 λ3 − λw3 (e−(λ1+λw3 )t − e−(λ1+λ3)t)− λ1λ3t λ3 − λw3 e−(λ1+λ3)t + λ1λ w 3 t λ3 − λw3 e−(λ1+λw3 )t; p6(t) = (λw3 t) 2 2 e−(λ1+λw3 )t; p7(t) = λ1 λ1 + λw3 − λ3 e−λ3t + λ2 1 + λ1(λ3 − λw3 ) (λ3 − λw3 )2 e−(λ1+λ3)t − λ3 1 (λ3 − λw3 )2(λ1 + λw3 − λ3) e−(λ1+λw3 )t + λ2 1t λ3 − λw3 e−(λ1+λ3)t; p8(t) = 2λ2 1 (λw3 − λ3)2 e −(λ1+λw3 )t + λ2 1(λw3 − 2λ3)t (λw3 − λ3)2 e−(λ1+λ3)t − (λ1t) 2λ3 λ3 − λw3 e−(λ1+λ3)t + λ2 1λ w 3 t (λw3 − λ3)2 e −(λ1+λw3 )t − 2λ2 1 (λw3 − λ3)2 e −(λ1+λ3)t; p9(t) = λ1λ w 3 t λw3 − λ3 e−(λ1+λ3)t − λ1(λ3t) 2 2(λ3 − λw3 ) e−(λ1+λ3)t − λ1 λw3 − λ3 e−(λ1+λw3 )t − λ1λ w 3 t λw3 − λ3 e−(λ1+λw3 )t − λ1(λw3 t) 2 2(λw3 − λ3) e−(λ1+λw3 )t + λ1 λw3 − λ3 e−(λ1+λ3)t; p10(t) = e−λ1t − (λw3 t) 2 + 2λw3 t+ 2 2 e−(λ1+λw3 )t. (7.2) sought reliability is found as r(t) = 10∑ i=1 pi(t). (7.3) references 1. barlow, r. e. & proschan, f. (1965) mathematical theory of reliability. ny, wiley. 2. gnedenko, b. v., belyaev, y. k. & solovyev, a. d. (1969) mathematical methods in reliability theory. new york, academic press. 3. kozlov, b. a. & ushakov, i. a., (1970) reliability handbook. new york, n y, usa: holt, rinehartand winston. 4. viktorova, v. s. & stepanyants, a. s. (2016) models and methods of reliability analysis of technical systems. 2th ed. moscow: lenand publ.. 5. smith, w. l., (1958) renewal theory and its ramifications. j. roy. statist. soc., ser. b, vol. 20, no. 2, 243–302. copyright c© 2018 assa. adv syst sci appl (2018) a unified approach to rap analysis based on markov processes with rewards 37 6. cox, d. r. (1962) renewal theory. new york, wiley. 7. kalashnikov, v. v. (1994) topics on regenerative processes. boca raton: crc press. 8. malinowski, j. (2013) “a monte carlo method for estimating reliability parameters of a complex repairable technical system with inter-component dependencies,” ieee trans. reliability, vol. 62, no.1, 256–266. 9. reibman, a. l. & trivedi, k. s. (1988) “numerical transient analysis of markov models,” computers&operations research, vol. 15, no. 1, 19–36. 10. reibman, a. l., smith, r. & trivedi, k. s. (1989) “markov and markov reward model transient analysis: an overview of numerical approaches,” european journal of operational research, vol. 40, 257–267. 11. lindemann, c., malhotra, m. & trivedi, k. s. (1995) “numerical methods for reliability evaluation of markov closed fault-tolerant systems,” ieee trans. reliability, vol. 44, no. 4, 694–704. 12. ryabinin, i. a. (1976) reliability of engineering systems. m., mir pbl.. 13. henley, e. j. & kumamoto, h. (1996) fault tree construction in probabilistic risk assessment and management for engineers and scientists, (2nd ed.) ieeepress, ch. 4, sec. 4.3, 166–172. 14. twigg, d. w., ramesh, d. w., sandadi, u. r. & sharma, t. c. (2000) “modeling mutually exclusive events in fault trees,” annual reliability and maintainability symposium. 2000 proceedings. international symposium on product quality and integrity, los angeles, ca, 8–13. 15. meshkat, l., dugan, j. b. & andrews, j. d. (2002) “dependability analysis of systems with on-demand and active failure modes, using dynamic fault trees,” ieee trans. reliability, vol.51, no. 2, 240–251. 16. tang, z. & dugan, j. b. (2006) “bdd-based reliability analysis of phased-mission systems with multimode failures,” ieee trans. reliability, vol. 55, no. 2, 350–360. 17. xing, l., morrissette, b. a. & dugan, j. b. (2014) “combinatorial reliability analysis of imperfect coverage systems subject to functional dependence,” ieee trans. reliability, vol. 63, no. 1, 367–382. 18. zhu, p., han, j., liu, l. & lombardi, f. (2015) “a stochastic approach for the analysis of dynamic fault trees with spare gates under probabilistic common cause failures,” ieee trans. reliability, vol. 64, no. 3, 878–892. 19. li, z., gu, j., x, j., fu, l., an, j. & dong, q. (2017) “reliability analysis of complex system based on dynamic fault tree and dynamic bayesian network,” 2017 second international conference on reliability systems engineering (icrse), beijing, china, 1–6. 20. smith, r. m., trivedi, k. s. & ramesh, a. v. (1988) “performability analysis: measures, an algorithm, and a case study,” ieee trans. computers, vol. 37, no. 4, 406–417. 21. amari, s. v., xing, l., shrestha, a., akers, j. & trivedi, k. s. (2010) “performability analysis of multistate computing systems using multivalued decision diagrams,” ieee trans. computers, vol. 59, no. 10, 1419–1433. 22. viktorova, v. s., sverdlik, y. m. & stepanyants, a. s. (2010) “analyzing reliability for systems with complex structure on multilevel models,” automation and remote control, vol. 71, no. 7, 1410–1414. 23. jia, h., ding, y., peng, r. & song, y. (2017) “reliability evaluation for demand-based warm standby systems considering degradation process,” ieee trans. reliability, vol. 66, no. 3, 795–805. 24. hazra, n. k. & nanda, a. k., (2014) “component redundancy versus system redundancy in different stochastic orderings,” ieee trans. reliability, vol. 63, no. 2, 567–582. 25. dolega, b., kopecki, g. & tomczyk, a. (2016) “possibilities of using software redundancy in low cost aeronautical control systems,” 2016 ieee metrology for aerospace (metroaerospace), padua, italy, 33–37. copyright c© 2018 assa. adv syst sci appl (2018) 38 v.s. viktorova, n.v. lubkov, a.s. stepanyants 26. bolvashenkov, i., kammermann, j. & herzog, h.-g. (2017) “methodology for quantitative assessment of fault tolerance of the multi-state safety-critical systems with functional redundancy,” 2017 international conference on information and digital technologies (idt), zilina – slovakia, 74–83. 27. pan, x., yang, y., zhang, g. & zhang, b. (2017) “resilience-based optimization of recovery strategies for network systems,” 2017 second international conference on reliability systems engineering (icrse), beijing, china, 1–6. 28. lim, t.-j. & lie, c. h. (2000) “analysis of system reliability with dependent repair modes ,” ieee trans. reliability, vol. 49, no. 2, 153–162. 29. abhilash, chakka, r. & challa, r.-k. (2013) “numerical performance evaluation of heterogeneous multi-server models with breakdowns and fcfs, lcfs-pr, lcfs-npr repair strategies,” 3rd ieee international advance computing conference (iacc),” ghaziabad, india, 566–570. 30. amari, s. v., myers, a. f., rauzy, a. & trivedi, k. s. (2008) “imperfect coverage models: status and trends,” in misra k.b. (eds) handbook of performability engineering, london, springer. 31. spiridonov, i. b., stepanyants, a. s. & victorova, v. s. (2012), design testability analysis of avionic systems. reliability: theory&applications. [online]. vol. 7, no. 3 (26). 66–73. available: http://www.gnedenko-forum.org/journal/2012/032012/rta 3 201208.pdf. 32. ng, y. w. & avizienis, a. (1980) “a unified reliability model for fault tolerant computers,” ieee trans. computers, vol. c-29, no.11, 1002–1011. 33. makam, s. v. & avizienis, a. (1982) ”aries 81: a reliability and life-cycle evaluation tool for fault-tolerant systems,” ieee 12-th fault-tolerant computing symposium (ftcs-12), santa monica, ca, 267–274. 34. bavuso, s. j., dugan, j. b. , trivedi, k. s., rothmann, e. m. & smith, w. e. (1987) “analysis of typical fault-tolerant architectures using harp,” ieee trans. reliability, vol. r-36, no. 2, 176–185. 35. forsythe, g. e., malcolm, m. a. & moler, c. b. (1977) computer methods for mathematical computations, nj, prentice hall. 36. ikramov, k.d. (1987) numerical methods of linear algebra. moscow: znanie publ.. 37. ustinov, s. m. & zimnitskiy, v. a. (2009) computational mathematics. spb., bhvpetersburg publ.. 38. rakitskiy, y. v., ustinov, s. m. & chernorutskiy, i. g. (1979) numerical methods for solving stiff systems. moscow: nauka publ., 81–112. copyright c© 2018 assa. adv syst sci appl (2018) adv syst sci appl 2022; 01:15–34 published online at https://ijassa.ipu.ru. direct torque control-fuzzy type 2 for direct current link voltages balancing of the five-level cascade converters used in a wind energy conversion system abdelhafidh moualdia1*, saleh boulkhrachef1, patrice wira2 1lrea laboratory, faculty of technology, university of medea, ain d’heb medea, algeria 2irimas laboratory, university of haute alsace, mulhouse, france abstract: the available power of a wind system depends mainly on the wind speed. in addition, the wind system will give a power output that varies according to the speed of its generator which is a double feed asynchronous machine in our case. in other words, there is an optimal operating point that makes the most of the power available. in this work an innovative technique of capturing maximum power based on type-2 fuzzy systems. the principle of this maximum power point tracking algorithm is to look for an optimal operating relationship at maximum power and then track the maximum power based on this relationship. as part of the variable speed conversion of wind energy, this article proposes a simplified power electronics for injecting the energy produced in the network, the conversion chain includes a variable speed double feed asynchronous generator, two (back-to-back) converters five-level neutral point clamped type operating in grid side rectifier and rotor-side inverter mode. the main objective of this article is to develop a new stabilization strategy of direct control compatible with voltage inverters at five levels, more particularly neutral point clamped structure. this strategy allows flow and torque control of the double feed asynchronous machine and stabilizes the input capacitor voltages of the inverter. the response of the system obtained with this algorithm makes it possible to validate the soft solution proposed and show, during variation of the wind speed, a fast and precise adaptation of the speed of the double fed induction generator. keywords: double fed induction generator, direct current link voltage, fuzzy type 2, direct torque control, wind energy, five level converter cascade 1. introduction from a general point of view, regardless of the topology, multilevel conversion structures offer huge advantages over a conventional solution, based on a two-level converter [1, 2]. these advantages are visible, on the one hand from a technological point of view and on the other hand from a functional point of view. first of all, the quality of the output signal of the inverter can be improved thanks to the additional degree of freedom which is the number of voltage level [3–6]. the switched voltage is of reduced amplitude and the switching is therefore easier to manage. despite the advantages of multi-level inverters, the instability of capacitor voltages on the dc side remains the major disadvantage of multi-level npc (neutral point clamped) inverters [7, 8]. as a result, the imbalance of these voltages leads to the failure of power components and deformation of the output voltage. for this, several solutions are proposed. these solutions include methods based on vector modulation techniques, where the concept of redundant ∗corresponding author: amoualdia@gmail.com 16 a. moualdia, s. boulkhrachef, p. wira voltage vectors has been applied to balance the electrical charge between capacitors [9–11]. however, for high levels, the number of voltage vectors increases considerably and thus the control becomes complex. other solutions, based on the addition of auxiliary circuits for the balancing of these dc voltages of the inverter, are proposed in the literature . to be able to replace the dc drive and enjoy the advantages of the asynchronous motor, the control must be more and more efficient. dtc (direct torque control) control strategy has emerged as competitive with vector control techniques. this dtc command was invented by i. takahashi in themid− 1980s. it is based on the separate regulation of the rotor flux and the electromagnetic torque of the double-feed asynchronous generator. in order to provide better control of npc-structured five-level inverter voltages and achieve better performance, the wind energy conversion system is connected to the power grid using rectifiers controlled by the width modulation. pulse (pwm). these rectifiers can provide a low harmonic distortion in the input currents, a grid-side adjustable power factor, and a constant dc output voltage. control of rectifiers based on fuzzy systems is also possible [12]. in particular, the fuzzy controller methodology appears useful when information sources are considered unclear or uncertain. in an ordinary fuzzy system, the membership functions, once determined, are completely precise, and therefore unable to take into account the uncertainty of the linguistic terms used in the premises and consequences of the rules. to solve this limitation, the fuzzy set type-2 was introduced as an extension of the fuzzy set type-1, where each degree of membership of each element is itself a fuzzy set in [0, 1]. [13–15]. first, we first present a description of the dfig-based wind energy conversion system. the second part of this work will be dedicated to the synthesis of innovative technique of maximum power capture based on type-2 fuzzy systems. the principle of this maximum power point tracking (mppt) algorithm is to look for an optimal operating relationship at maximum power and then track the maximum power based on this relationship. in the third step, the development of a stabilizing dtc control, to fig. 1.1. wind energy conversion system based on five-level npc cascaded converters solve the problem of the imbalance of the input voltages of the 5-level converter with npc structure on the rotor side. taking advantage of the redundancies of the state of the rotor side converter, producing the same voltage vector, but opposite effects on the voltages of the capacitors, an algorithm is developed. this allows us to balance the capacitor voltages in addition to controlling the torque and flux of the double feed asynchronous generator. the copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 17 wind energy conversion system description based on five-level npc cascaded converters is shown in figure 1.1. nomenclature β blade pitch angle; v wind speed (m/s); ρ air density; λ speed ratio; cp power coefficient; pm mechanical power (kw); tm mechanical torque ; vsdq, vrdq dq axis stator and rotor voltages ; isdq, irdq, dq axis stator and rotor current ; ωs, ωr, stator and rotor pulsation (rd/s) ; φsdq,φrdq, dq axis stator and rotor flux ; m, mutual inductance ; ps, qs generator active and reactive powers ; p1,p2 logical function ; j, f moment of inertia and coefficient of friction ; e1,e2,e3 load conditions ; k phase number (k=1, 2, 3); c1; c2; c3; c4 dc-link capacitors; uc1; uc2; uc3; uc4 dc bus voltages; ic1; ic2; ic3; ic4 capacitors currents; irec1; irec2; irec3; irec4 output rectifier currents; iref netk reference network phase current; iref rec reference network phase current; vsk; isk stator phase voltage and current; vsαβ, isαβ stator voltage and current in the stationary α− β plane ω,ωn rotor speed and speed nominal value; tem,tref electromagnetic torque and reference value; φs stator flux magnitude; φsα,φsβ stator flux magnitude in α− β plane fnet network frequency; vdi discrete voltage level of vector vs; cfl output hysteresis flux controller; ctr output hysteresis torque controller; abbreviations used npc : neutral point clamped hvdc : high voltage direct current dc : direct current dtc : direct torque control pwm : pulse width modulation csr : current source rectifier im : induction machine svm : space-vector modulation flc : fuzzy logic controller t1fls : type-1 fuzzy sets t2fls : type-2 fuzzy sets pi : proportional integral controller dfig : double fed induction generator mppt : maximum power point tracking rsc : rotor-side converter gsc : grid-side converter copyright © 2022 assa. adv syst sci appl (2022) 18 a. moualdia, s. boulkhrachef, p. wira 2. description of the wind conversion system 2.1. wind turbine model the mechanical power available on the shaft of a wind turbine can be expressed as [22, 26]: pm = 0.5cp (λ) πρr2v 3 1 , (2.1) for the variables speed wind turbines, approximate expression of the power coefficient can be described by the following expression: cpf (λ, β) = c1 ( c2 λi − c3 − c4 ) exp ( −c5 λi ) +c6λ, (2.2) where: 1 λi = 1 λ+ 0.08β − 0.035 β3 + 1 (2.3) where, c1 = 0.5176, c2 = 116, c3 = 0.4, c4 = 5, c5 = 21, c6 = 0.0068.the torque produced by the turbine is expressed in the following way: tm = pm ωm = ( 0.5πρr3v 2 1 ) cp(λ, β)) λ (2.4) no wind turbine could convert more than 59 of the kinetic energy of the wind into mechanical energy turning the rotor [16]. this is known as the betz limit and it’s the cpmax theoretical maximum coefficient of power (for any wind turbine): cpmax = 16/27 ≈ 0.593. in practice, for the good turbines it’s in the range of 0.45 to about 0.50. the tip speed ratio (λ) for wind (a) (b) fig. 2.2. (a): for pitch angle effect on the aerodynamic coefficient of power and (b): for mechanical power output. turbines is the ratio between the rotational speed of the tip of a blade and the actual velocity of the wind, see figure 2.2. copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 19 table 2.1. inference matrix de\e ng nm np ez pp pm pg ng ng ng nm nm np np ez nm ng nm nm nm np ez pp np ng nm np np ez pp pm ez ng nm np ez pp pm pg pp nm np ez pp pp pm pg pm np ez pp pm pm pm pg pg ez pp pp pm pg pg pg fig. 2.3. tip speed ratio control (mppt). 2.2. maximum power tracking via fuzzy-type 2 systems at a given wind speed, the maximum turbine energy conversion efficiency occurs at an optimal tsr figure 2.3. therefore, as wind speed changes the turbine’s rotor speed needs to change accordingly in order to maintain the optimal tip speed ratio tsr and thus to extract the maximum power from the available wind resources. the structure of the fuzzy type-2 regulator is shown in figure 2.4. in order to have the desired performances, the normalization gains at the input and at the output of the regulator are determined by adjustment. for input and output variables consisting of seven fuzzy sets type-2 interval, the conventional anti-diagonal inference matrix of a fuzzy system is given in table 2.1. copyright © 2022 assa. adv syst sci appl (2022) 20 a. moualdia, s. boulkhrachef, p. wira fig. 2.4. structure of the fuzzy type-2 controller. (a) (b) fig. 2.5. membership functions.(a):membership functions of input variables,(b):membership functions of output variables 2.3. modelling of the double fed induction generator in the rotating field reference frame of park, the model of the dfig is given by the following equations [30, 31]: vsd = rsisd + d dt φsd − ωsφsq vsq = rsisq + d dt φsq + ωsφsd vrd = rrird + d dt φrd − ωrφrq vrq = rrirq + d dt φrq + ωrφrd, (2.5) stator and rotor voltages components: φsd = lsisd +mird φsq = lsisq +mirq φrd = lrird +misd φrq = lrirq +misq. (2.6) copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 21 double fed induction generator electromagnetic torque: cem = cr + j dω dt + fω, (2.7) generator active and reactive powers at the grid side are:{ ps = vsdisd + vsqisq qs = vsqisd − vsdisq. (2.8) 2.4. five-level npc converter multilevel converters are power-conversion systems composed by an array of power semiconductors and capacitive voltage sources that, when properly connected and controlled, can generate a multiple-step voltage waveform with variable and controllable frequency, phase, and amplitude. the stepped waveform is synthesized by selecting different voltage levels. the numbers of levels of a converter is defined as the number of steps that can be generated by the converter between the output terminal and any reference node within the converter, is usually denoted by n and called neutral. to be called a multilevel converter, each phase of the converter has to generate at least three different voltage levels. this differentiates the classic tw o-level voltage source converter (2l− v sc) from the multilevel family. the neutral clamped inverter, also known as diode clamped inverter. the basic architecture of this inverter discussed in references [17, 18]. the neutral clamped inverter obtained the fig. 2.6. the five-level npc converter. staircase output voltage. if m is the number of level, then the number of capacitors required on the dc bus are (m -1), the number of power electronic switches per phase are 2(m -1) and the number of diodes per phase are 2(m− 2). the dc bus voltage has three levels using two capacitors c1 and c2, for five levels using four capacitors c1, c2, c3 and c4 as shown in figure 2.6. table 2.2 lists the voltage levels and their corresponding switch states. state 1 means that the switch is on, and 0 means that the switch is off. we suppose that all the dc voltage sources are the same and equal to nominal value . the output voltage vector is defined copyright © 2022 assa. adv syst sci appl (2022) 22 a. moualdia, s. boulkhrachef, p. wira table 2.2. switching states and output voltage of the first leg of the five-level npc inverter state v1m s11 s12 s13 s14 s15 s16 s17 s18 +2 +2uc 1 1 1 0 0 0 0 0 +1 +uc 1 1 0 0 0 1 1 0 0 0 1 0 0 1 0 0 0 0 -1 −uc 0 0 1 1 1 0 0 1 -2 −2uc 0 0 0 1 1 1 0 0 as: vs = vame j0 + vbme −j2π/3 + vcme −j4π/3 = vα + jvβ. (2.9) it can be observed that 24 vectors can be generated by a unique switching state, 18 vectors can be generated using two switching states each (2 redundancy), 12 vectors can be generated using three switching states each (3 redundancies), 6 vectors can be generated using four switching states each (4redundancies), and one vector can be generated using five switching states (5redundancies). [15, 16] 3. dtc for five-level rotor side converter (rsc) in recent years, a new control strategy based on direct control of flux and torque has been proposed. this technique, known as dtc, enables the induction motor to deliver a very quick and accurate torque response.the instantaneous values of flux and torque are calculated from measured variables (voltages and currents) and then controlled directly by selecting optimum inverter switching modes.the objective of this section is to present the dtc of induction machine fed by a five–level npc inverter. the schematic diagram of the proposed dtc system is shown in figure 3.7. as in the original dtc principle the α− β plane will be divided into several sectors [26]. in a five-level inverter the number of discrete voltage vectors is more important than those obtained with a two-level inverter. thus, the α− β plane will be divided into 12 sectors rather than six. at each one of these sectors, an appropriate voltage vector will be assigned to keep flux and torque references as needed.the speed of the stator flux vector is given by the modulus of the applied space voltage vector. thus, the space voltage vectors will be chosen according to the rotor speed. voltage vectors with low amplitude will be chosen for low speeds, and vectors with greater amplitude will be chosen for higher speeds. in the five-level inverter, four tables have been used, according to a specific speed range, as shown in table 2.2 . three different states are used to identify torque status and two states for flux. 4. control strategy of the balancing capacitor voltage we will present the proposed dtc algorithm to balance dc bus voltages using redundant configurations of the five-level inverter. the, each discrete voltage level can be obtained by more than one switching state. as the voltage evolution for a given capacitor will be different for each state, this redundancy let’s to control the capacitors voltages while the requested space vector voltage is supplied. based on this property, a control strategy will be copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 23 fig. 3.7. schematic diagram of the proposed dtc system. table 4.3. six groups of redundant vectors group 1 vp1 vp4 vp7 vp10 vp13 vp16 group 2 vp2 vp6 vp8 vp12 vp14 vp18 group 3 vp3 vp5 vp9 vp11 vp15 vp17 group 4 vp19 vp21 vp23 vp25 vp27 vp29 group 5 vp20 vp22 vp24 vp26 vp28 vp30 group 6 vp31 vp32 vp33 vp34 vp35 vp36 presented and applied to a five-level inverter. to do so, we firstly study the effect of different redundant vectors on capacitor voltages. in table 5.4, all the redundant vectors of the space vector diagram and the corresponding capacitors currents according to the load (i1, i2, i3) are presented. depending on the forms of relationships (equation 2.9, 4.10 and 4.11 in table 4.3), we distinguish six groups of redundant vectors: 4.1. effect of redundant vectors on capacitor voltages redundant vectors of each group can increase or decrease capacitor voltages, depending on load conditions sign of (equation 2.9, 4.10 and 4.11. for groups with one equation 2.9 (groups 1, 4 and 6), we have two possibilities of load conditions, each one is associated with a logical function: p1 = 1 if e1 > 0 else p1 = 0 p2 = 1 if e1 ≤ 0 else p2 = 0. (4.10) copyright © 2022 assa. adv syst sci appl (2022) 24 a. moualdia, s. boulkhrachef, p. wira for groups with three (equation 2.9, 4.10 (groups 2, 3 and 5), we have six possibilities of load conditions, associated with six logical functions: p1 = 1 if e1 < 0, e2 < 0, and e3 > 0, else p1 = 0; p2 = 1 if e1 < 0, e2 > 0, and e3 < 0, else p2 = 0; p3 = 1 if e1 < 0, e2 > 0, and e3 > 0, else p3 = 0; p4 = 1 if e1 > 0, e2 < 0, and e3 < 0, else p4 = 0; p5 = 1 if e1 > 0, e2 < 0, and e3 > 0, else p5 = 0; p6 = 1 if e1 > 0, e2 > 0, and e3 < 0, else p6 = 0. (4.11) 4.2. choice of redundancies for each case of redundancy, the vector which reduces the fluctuation voltages in capacitors will be selected. the diagram of the control algorithm is shown in figure 4.8. we select the vector which charge the undercharged capacitors, and discharge the overcharged ones. to do so, we must measure capacitor voltages and calculate their deviation case. each deviation case is characterized by a logical function cj (table 5.4.). for example, the first case : uc1 k, then set direfr ′′ = direfr . • step 4 :if direfr ′′ 6= direfr ′, then go to step 5. if direfr ′′ = direfr ′ then set direfr = direfr ′′ and go to step 6. • step 5 :set direfr ′ = direfr ′′ and return to step 2. • step 6 : end the procedure for computing direfl is very similar, only two changes need to be made: in step 2, we need to find k (1 ≤ k′ ≤m − 1) such that uml ≤ direfl ′ ≤ uml , m=k’+1 and in step 3, let f i l = f i for i ≤ k′ and f i l = f i for i > k′ . from the type-reduction stage, we have for each output a type-reduced set. the crisp output of the type-2 fuzzy controller can be obtained by using the average value of direfr and direfl . hence, the defuzzified crisp output becomes: direfk = direfkl + direfkr 2 (5.18) fig. 5.11. wind speed profile 6. simulation results in order to demonstrate the feasibility of the proposed control method, simulation testing are undertaken on the structure shown in figure 1.1. the most commonly encountered disturbances in drive applications are changes in the load torque or changes in the speed. it is important thereby to show if the proposed cascade is able to handle aforementioned transients and ensure the stability of dc-bus voltages. the wind speed profile applied to energy conversion system is shown with in figure 5.11. the multilevel dtc strategy has been tested by simulation in different level of speed control loop.the reference speed is obtained by the mppt block. the control objective in this section is to show the performance in both interest operation regions (i and iii) of the turbine characteristic. in the region iii, one can see that the stator power is limited at its copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 29 fig. 6.12. mppt result maximum value by the pitch angle control,as shows figure 6.12. this is approved by the waveforms illustrated in figure 6.17 of the tip speed ratio (λ) and the power coefficient (cp). the mechanical speed is kept constant, at its limit value, as plotted in figure 6.12. fig. 6.13. response of electromagnetic torque and rotor flux however, in region i the power is maximized by the mppt algorithm. in this case, as expected, tip speed ratio and the power coefficient are maintained at constant values see figure 6.12 (β = 0,λ = λopt = 8.1, cp = cp−max = 0.49). the reference tracking reflecting the robustness of the proposed control under random behavior of wind speed. figure 6.13 shows the rotor flux waveform, it is circular and kept constant at 1.6 wb. the random evolution of the rotor current magnitude is related to the electromagnetic torque variations and their pulsation is depending on the slip variations. the operation as a generator is illustrated by the negative sign of the electromagnetic torque cem < 0, with a variable amplitude, which depends on the wind speed evolution, as presented in figure 6.13 where we can see that the operation with constant power is clear within the over-speed zone. in figure 6.14, the copyright © 2022 assa. adv syst sci appl (2022) 30 a. moualdia, s. boulkhrachef, p. wira capacitor voltages with the proposed dtc balancing strategy are shown . it can be seen that the balancing of the capacitor voltages is achieved sufficiently over the full speed state and for all load torque of induction machine, which prove that the stability of dc voltages with proposed strategy is independent of machine operating points. in the studystate condition, the maximum of each capacitor voltage ripple is less than 4v (2%) , which shows the effectiveness of the dtc balancing strategy. consequently, the different between the voltages capacitor tends to zero, as shown in figure 6.15. fig. 6.14. response of dc-link voltages with proposed dtc control strategy. fig. 6.15. error dc-link voltages. 7. conclusion in this work, the application of direct torque control in a wind energy conversion system supplied by a five level npc back-to-back converter has been presented. the choice of this copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 31 fig. 6.16. response of stator and rotor currents structure allows a bidirectional flow of power, and provides regeneration capability.another advantage of the back-to-back system is that it can control easily the input power factor. the effectiveness of the proposed dtc-based voltage balancing strategy it demonstrated under various operating conditions of double fed induction generator. is may therefore be concluded that the proposed multilevel direct torque control not only has the advantage of reducing the undesirable torque ripple, but also has the additional advantage of preventing the voltage drift phenomen on of the dc-link capacitors of the back-to-back system, longer life of dc-link capacitors can be achieved. furthermore, the employment of multilevel topology improves the stator voltage quality, reducing electromagnetic interference and isolation stress problems of windings. this results in the reduction of radiated emissions, which makes this kind of drives a less polluting system than the traditional two-level drive. in addition, this study has successfully demonstrated the application of type-2 fuzzy systems to control the dc-link voltage, and the maximization of power tracking. it found that the type-2 fuzzy control scheme could achieve good performance in terms of overshoot, steady-state error, torque disturbance, and variable speed tracking. copyright © 2022 assa. adv syst sci appl (2022) 32 a. moualdia, s. boulkhrachef, p. wira fig. 6.17. response of stator active and reactive power acknowledgements the research is part of a project prfu’2020, realized respectively in the laboratory of electrical engineering and automatic lrea research, university of medea, and irimas laboratory, university of haute alsace, mulhouse, france. references 1. polinder, h., f. f. a. van der pijl, g. j.de vilder and p. j. (2006) tavner.comparison of direct-drive and geared generator concepts for wind turbines. ieee transactions on energy conversion., 21 (3), 725–733. 2. p.liutanakul , a.awan, s.pierfederici, b.nahid-mobarakeh , f. m.tabar.(2010) linear stabilization of a dc bus supplying a constant power load: a general design approach. ieee transactions on power electronics, 22 (2),475–488. 3. x.sun , y.tian , z.chen. (2014) adaptive decoupled power control method for inverter connected dg. iet renewable power generation, 8 (2),171–182. 4. brabandere, k, bolsens, b., keybus, j. woyte, a. driesen, j. belmans, r.(2007) a voltage and frequency droop control method for parallel inverters. ieee trans. power electron, 22 (4),1107–1115. 5. wang, y., x. wang and d. gerling.a.(2018) precise voltage distortion compensation strategy for voltage source inverters. ieee transactions on industrial electronics., 65 (1) 59–66. 6. vovos, n. a., p. n. vovos and k. g.georgakas.(2014) harmonic reduction method for a single-phase dc-ac converter without an outputfilter. ieee transactions on power electronics., 29 (9), 4624–4632. 7. kapoor, p. and m. renge.(2017) improved performance of modular multilevel converter for induction motor drive. energy procedia., 117 (1),361–368. 8. g.c. diyoke, c.u.ogbuka, c.m.nwosu. (2019) a novel control dc-dc-ac buck converter for single phase capacitor-start-run induction motor drives. power engineering and electrical engineering,17 (2),87–95. 9. ajami, a., h. ardi and a. farakhor.(2014) design,analysis and implementation of a buck-boost dc/dc converter. iet power electronics., 7 (12),2902–2913. 10. bartman, j.(2017) analysis of output voltage distortion of inverter for frequency lower than nominal. journal of electrical engineering. 68 (3),194–199. copyright © 2022 assa. adv syst sci appl (2022) dtc-fuzzy type 2 for dc link voltages balancing of the five-level cascade 33 11. hanafy, h. (2014) dynamic analysis of a new connection for dual voltage operation of single phase capacitor run motor. journal of electrical engineering., 13 (1),1–8. 12. ojaghi, m. and s. daliri.(2017) analytic model for performance study and computeraided design of single-phase shaded-pole induction motors.ieee transactions on energy conversion., 32 (2),649–657. 13. leon, j. i., s. vazquez and l. g. franquelo. (2017) multilevel converters: control and modulation techniques for their operation, industrial applications. proceedings of the ieee., 105 (11),2066–2081. 14. franquelo, l., j. rodriguez, j. leon, s. kouro, r. portillo and m. prats.(2007) the age of multilevel converters arrives. ieee industrial electronics magazine., 2 (2) ,28–39. 15. babaei, e., s. laali and s. bahravar.(2015) a new cascaded multi-level inverter topology with reduced number of components and charge balance control methods capabilities.electric power components and systems, 43 (19) ,2116–2130. 16. kimball, j. and m. zawodniok. (2011) reducing common-mode voltage in threephase sine-triangle pwm with interleaved carriers. ieee transactions on power electronics. 26 (8),2229–2236. 17. mohamad r. banaei, m. r. jannati oskuee and f. mohajel kazemi.na. (2014) new advanced topology of stacked multicell inverter. international journal of emerging electric power systems, 15 (4),327–333. 18. sadigh ak, hosseini sh, sabahi m, gharehpetian gb.(2010) double flying capacitor multicell converter based on modified phase-shifted pulsewidth modulation.ieee trans power electron, 25 (6),1517–1526. 19. w. fengxiang, c. zhe, p. stolze, j.-f. stumper, j. rodriguez, r. kennel. (2014) encoderless finite-state predictive torque control for induction machine with a compensated mras.ieee trans. ind. informat, 10 (2),1097–1106. 20. wang, l., m. zhao, x. wu, x. gong and l. yang. fully. (2017) integrated highefficiency high step-down ratio dc–dc buck converter with predictive over-current protection scheme. iet power electronics, 10 (14),1959–1965. 21. h.kouki, m.fredj, b. rehaoulia. (2016) harmonic analysis of svpwm control strategy on vsi-fed double-star induction machine performances, electrical engineering,98 (2),133–143. 22. a.ziaei, r.ghazi, r.z.davarani.(2018) linear modal analysis of doubly-fed induction generator (dfig) torsional interaction: effect of dfig controllers and system parameters,advances in electrical and electronic engineering ,16 (4),388 – 401. 23. zhu, c., l. fan and m. hu. (2010) control and analysis of dfig-based wind turbines in a series compensated network for ssr damping.in: ieee pes general meeting. michigan: ieee, 1–6. 24. van, t. l., t. d. ngyen, t. t. tran and h. d. nguyen.(2015) advanced control strategy of back-to-back pwm converters in pmsg wind power system. advances in electrical and electronic engineering, 12 (2),81–95. 25. leon, a. e. and j. a. solsona. (2015) sub-synchronous interaction damping control for dfig wind turbines. ieee transactions on power systems, 30 (1),419–428. 26. t.vaimann, o.kudrjavtsev, a.kilk, a.kallaste, a.rassolkin. (2018) design and prototyping of directly driven outer rotor permanent magnet generator for small scale wind turbines.advances in electrical and electronic engineering, 16 (3),271– 278. 27. boulkhrachef, s., berkouk, e. m., barkat, s. benkhoris, m. f. (2012) self voltage balancing for the five level back to back converter using multilevel dtc and type2 fuzzy logic controller. the mediterranean journal of measurement and control,8 (4),491–506. 28. shukla, r. d., tripathi, r. k. thakur, p. (2017) dc grid/bus tied dfig based wind energy system .renewable energy, 108,179–193. 29. copyright © 2022 assa. adv syst sci appl (2022) 34 a. moualdia, s. boulkhrachef, p. wira 29. lalili, d., berkouk, e. m., boudjema, f. , lourci, n. (2008) self-balancing of dc-link capacitor voltages using redundant vectors for svpwm controlled five-level inverter, the fifth international multi-conference on systems, signals devices, ieee ssd08, amman, jordan, paper reference, ssd08-15691–17326. 30. boulkhrachef, s., moualdia, a., boudana, dj. , wira, p. (2019) higher order sliding mode controler of a wind energy conversion system, nonlinear dynamics and systems theory, 19 (04), 486–496. 31. beltran1, b., benbouzid, m.e.h., ahmed-ali, t. (2009) high-order sliding mode control of a dfig-based wind turbine for power maximization and grid fault tolerance, ieee international electric machines and drives conference, miami,fl,usa, 183—189. copyright © 2022 assa. adv syst sci appl (2022) introduction description of the wind conversion system wind turbine model maximum power tracking via fuzzy-type 2 systems modelling of the double fed induction generator five-level npc converter dtc for five-level rotor side converter (rsc) control strategy of the balancing capacitor voltage effect of redundant vectors on capacitor voltages choice of redundancies control of five-level grid side converter (gsc) type-2 fuzzy logic controller design simulation results conclusion microsoft word 892 article source, copyedited.doc adv syst sci appl 2020; 02; 82-97 published online at https://ijassa.ipu.ru. original russian text © yu. v. mitrishkin, n. m. kartsev, a. e. konkov, m. i. patrov, 2019, published in problemy upravleniya / control sciences, 2019, no. 3, pp. 3–15. plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter yuri mitrishkin1,2*, nikolay kartsev2, artem konkov1,2, mikhail patrov3 1) m.v. lomonosov moscow state university, faculty of physics, moscow, russia 2) v.a. trapeznikov institute of control sciences of russian academy of sciences, moscow, russia 3) a.f. ioffe institute of russian academy of sciences, st. petersburg, russia e-mail: yvm@mail.ru e-mail: n.kartsev@yandex.ru e-mail: konkov@phys.msu.ru e-mail: michael.patrov@mail.ioffe.ru abstract: the plasma magnetic control systems for iter (international thermonuclear experimental reactor) are presented. the systems comprise original engineering solutions for plasma position, current and shape control for the two versions of iter: iter-1 and iter-2, including those designed in v. a. trapeznikov institute of control sciences of the ras. it is noted that in iter-1 the plasma position and shape were controlled by all pf-coils and robust h∞-controllers, while to decrease peaks of control power for suppressing minor disruptions the additional nonlinear circuit was used without significant changes in displacements of the gaps between the plasma separatrix and the first wall. in iter-2 the special circuit with a fast voltage rectifier was used for plasma vertical speed stabilization about zero, while for plasma current and shape control the special cascade control systems were designed with and without the control channels decoupling, with robust h∞-controllers and predictive model, and with adaptive stabilization of the plasma vertical position as well. to increase the plasma controllability region in the vertical direction the additional horizontal field coils were introduced into the iter-2 vacuum vessel and the capabilities of the system with the new circuits to control the plasma vertical position at the noise presence were investigated. keywords: tokamak, plasma, plasma magnetic control, iter introduction currently, the flagship in solving the problem of controlled thermonuclear fusion is iter – international thermonuclear experimental reactor [1-3]. tokamak reactor iter is being built in france (cadarache) by an international consortium consisting of the european union, the usa, japan, russia, china, india, and south korea. all the most advanced tokamaks work in support of iter to provide a basic physical understanding of its reliable operation capabilities, including through magnetic and kinetic plasma control systems. on the russian side, a number of magnetic control systems for the plasma position, current, and shape have been developed for iter, which have been modelled on plasmaphysical codes dina (triniti, troitsk, russia) and pet (d.v. efremov research institute of electrophysical equipment, st. petersburg, russia). due to the peculiarity of the iter poloidal system associated with superconductivity of the poloidal field coils, in v.a. trapeznikov institute of control sciences of ras (moscow, russia) the robust magnetic * corresponding author: yvm@mail.ru plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 83 copyright ©2020 assa. adv. in systems science and appl. (2020) plasma control systems were developed. these systems are capable of working in conditions of uncertainty caused by the unmodeled dynamics of the plasma as a complex not fully studied plant. the developed systems work in the general hierarchical structural scheme of the plasma control system in iter, taking into account the structural scheme of controlling the position, current and gaps between the plasma separatrix and the first wall of the vacuum vessel (fig. i.1) [4]. fig. i.1. block diagram of the hierarchical system of magnetic plasma control in iter in the block diagram of magnetic plasma control in iter (fig. i.1) the poloidal field control device (pf-controller, pf – poloidal fields) will presumably consist of three separate units, namely: a supervisor located at the upper (adaptive) level of control, a feedforward controller and a feedback controller of the lower (basic) level. the basic level (control loop) of the poloidal field control system includes a controlled plant consisting of a plasma, coils of a poloidal magnetic field, passive metal structures, sensors, a measuring (diagnostic) unit, an actuator device representing power sources for the central solenoid and poloidal field coils; a feedback controller and a feedforward controller. the supervisor (fig. i.1), interacting with the plasma control system and the interlock system, controls the operation of the basic level. it generates program signals at the rate of observation and performs adaptive correction of the feedback controller about 10 times slower than the basic circuit. at present, magnetic plasma control systems are also being developed for the first thermonuclear power plant demo (demonstration power plant) [5]. it is important to develop a poloidal system, which would not be forced to introduce additional plasma position control coils inside the vacuum vessel, as was done in iter-2 which reduces the reliability of the thermonuclear reactor in stationary mode. coils for effective control of the plasma position, providing a sufficient controllability region in vertical direction with a given limit on the supply voltage at the vertical instability of the plasma, can be installed between the vacuum vessel and the coil of the toroidal magnetic field, as is done in the asdex upgrade tokamak and the t-15md tokamak project [2]. in the japanese variant of demo cs & pf coil power supplies disturbance diagnostics actuator plant measuring unit supervisor feedforward controller feedback controller ref y v ff pf controller interlock system plasma control system (codac) y v fb + + + adaptive level basic level i coils y b coils metalllic structures plasma sensors 84 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) [5] such magnetic coils are not provided, which requires further development and research of the demo poloidal system. 1. the iter project variants and scenarios in the iter project the iter-1 version with self-sustaining thermonuclear reaction and q = ∞ initially was developed but since 1998 the iter-2 version with q = 10 was begun (fig. 1.1, table. 1.1) [4], where q = pout/pin is the ratio of the output power of thermonuclear fusion pout to the input power pin. in iter, it was proposed to control the plasma boundary in the divertor configuration (lower x-point) by controlling the distances (called gaps) between the separatrix and the first wall at six different points, including the intersection points of the separatrix with the divertor plates. the plasma cross section, the location of the nine poloidal field coils (pf1 – pf9) in iter-1 and six pf coils (pf1 – pf6) in iter-2, the solid central solenoid (cs – central solenoid) in iter-1 and the six-section cs in iter-2, and the locations of the controlled gaps between the plasma boundary and the surface of the surrounding elements are shown in fig. 1.1. table 1.1. the plasma parameters in two versions of iter plasma parameters iter-1 iter-2 plasma current, ma 21 15 major radius, m 8.1 6.2 minor radius, m 2.8 2.0 elongation 1.6 1.85 thermonuclear fusion power, gw 1.5 0.5 q ¥ >10 burning time, с 1000 400 vertical instability time constant, sec 1.1 0.1 separatrix deviations from the given location must be within ±10 cm, which is less than 5% of the plasma major radius. the gap locations shown in the figures are chosen to provide a reasonably good overall performance of plasma shape control and at the same time for precise control at key points such as divertor strike points (g1, g2) and the plasma boundary in front of the ion-cyclotron radio frequency heating antenna (g3). all poloidal field coils are superconducting with a relatively small cross-sectional area. an ac/dc converter controls each coil. in iter-1, all the pf coils and the central solenoid control the unstable vertical position of the plasma. in iter-2, a special circuit is provided to control the plasma vertical position or speed, and in the subsequent version for this purpose, additional horizontal field coils are introduced into the vacuum vessel. plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 85 copyright ©2020 assa. adv. in systems science and appl. (2020) a b fig. 1.1. poloidal systems and arrangements of controlled gaps g1 g6 between the separatrix and the first wall: iter-1 (a) and iter-2 (b) the iter-2 scenario includes the following sequence of phases [3]: plasma start and growth of the plasma current, divertor configuration (x-point) formation, heating to ignition, controlled continuous burning, controlled reduction of fusion power and plasma current, termination of the x-point, plasma shutdown (the end of a discharge). the required evolution of the plasma configuration is created by changing of the poloidal currents (feedforward components) and adding of the feedback control. two main phases that can be identified are the most different in terms of control requirements: • limiter phase characterized by the plasma restricted by a limiter with a relatively low thermal and magnetic energy content: the control requirements are not very strict during this phase, and it is necessary basically only to suppress the vertical instability of the plasma; • divertor phase, during which the high amount of thermal and magnetic energy is a potential danger to the preservation of the outer layer of the plasma and its surrounding elements; this phase requires precise control of the plasma shape and current. 2. synthesis and modeling of plasma control systems in iter-1 the pf-controller (fig. i.1) should be able to control with given performance the gaps between the plasma and the first wall under the action of specified disturbances of the type of minor disruption (md). in doing so, the controller must guarantee to stabilize the vertical unstable position of the plasma in iter-1 or to stabilize the vertical velocity of the plasma around zero value in iter-2. the h∞ robust plasma current, position and shape control system was developed in iter-1 at q = ∞ [4, 6, 7] with an ncf-controller in the feedback (ncf – normalized coprime factorization). the robust system is weakly sensitive to errors and uncertainties of the tokamak plasma model. it remains stable at a large set of iter magnetic plasma configurations in comparison to classical systems, for example linear-quadratic control. for the first time by the method of mathematical modelling on the plasma-physical code dina it was shown that a larger robust stability margin of the multivariable control system provides a greater number of magnetic configurations from the iter database. for these configurations the control system stabilized the position of the separatrix under the action of mdperturbations on the divertor phase of the discharge [6, 7]. the study of the effect of perturbations of the elm type (edge located modes) showed a slight increase in the 86 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) amplitude of the currents in the poloidal field coils compared to their natural amplitude without feedback. in addition to the linear h∞ ncf robust controller in the feedback system, a nonlinear correction unit of the total control power (power management system – pms) was developed, which allowed to reduce power peaks by 30-40% in the simulation of the system in the presence of md-type perturbations without significant changes in the behavior of gaps [4, 6, 7]. fig. 2.1 shows the block diagram of closed-loop control system with the linear h∞ ncf controller in the feedback and the nonlinear correction unit. the principle of operation of the correction unit is as follows. the total control power signal, which is the sum of the products of the currents in and the voltages of the control coils un where ieqn is the current in the poloidal field coil, which specifies the plasma equilibrium (see fig. 2.1). the vector of corrective (additional) voltages is created as a product of the power module and voltages on the control coils in accordance with δu = –|p|u. this vector is introduced into the feedback to reduce the input voltages at a strong increase in power, which reduces the power peaks under the action of perturbations. fig. 2.1. control system with the correction unit of the full control power the developed controllers were modeled in closed-loop control systems on linear (create-l, pet-l, tsps-l, corsica-l) and non-linear (tsc, dina, corsica, pet, maxfea, tsps) models of the plant under the action of specified md-type and elm-type disturbances. modeling has shown that systems with developed controllers met technical requirements. as an example in fig. 2.2 for the system with the h∞ ncf controller at the point of the sob (start of burn) scenario on the pet-l linear model, graphs of transients after the md are given for variations of gaps, voltages, poloidal coil currents, total control power, its derivative, variations of plasma current. 10 1 ,n n n eqn n n p i u i i i = = = + då saturation -st e (1+st) delays planт t cs & pf coil power supplies feedback controller plasma current variation cs&pf coil current variations gap displacements hinfncf scalar product modulus product power management system power i nom + + _ _ u i plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 87 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 2.2. transient responses with h¥ ncf controller in case of md-type disturbance (pet linear model, sob) 3. synthesis and modeling of plasma control systems in iter-2 the new principles and plasma control systems for iter-2, which were developed in solving of problems of modeling magnetic plasma control systems in iter were proposed and applied. the focus of the development of the systems is associated with an increase in their reliability, survivability and performance of control (speed and accuracy). (1) synthesis of h∞ control system based on the scheme of external disturbance rejection (fig. 3.1–3.3) [4, 8, 9]. the block diagram of the control system is shown in fig. 3.1a [3]. it consists of two control loops: the first scalar loop serves to stabilize the vertical plasma velocity relative to zero, and the second mimo loop is designed to control the plasma current and shape: the six gaps between the plasma separatrix and the first wall (fig. 1.1, b). to stabilize the vertical plasma velocity, a special scheme is applied (see fig. 3.1, b), consisting of parallel connection of pf2 pf5 coils with appropriate directivity and slow 88 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) rectifiers, to which a fast rectifier vs (vertical stability) is connected, directly acting on vertical plasma velocity. first, the loop shaping of the open-loop system [10] and the mixed sensitivity [11] were used to synthesize robust controllers and by means of μ-analysis the controllers were compared with each other, as well as with lqg-, lead-lagand p-controllers. the loopshaping controller showed the smallest peak of µ, that is, the greatest robust stability margin [8]. a b fig. 3.1. iter double-loop magnetic control system: a) vsc is a vertical stabilization converter, mc is a main converter, kbd is a block-diagonal controller, fd is a differentiating filter, f is a filter; b) connection diagram of a fast voltage converter for suppressing the vertical plasma velocity, vs is a vertical plasma stabilization thyristor rectifier, m2 – m5 are main rectifiers for pf2 – pf5 coils then two block-diagonal controllers were synthesized to control the speed of plasma vertical movement, current, and shape with a cascade of currents in the coils of the poloidal field (fig. 3.2, a) and without this cascade (fig. 3.2, b) [4, 9]. the systems were tested at various points of the iter scenario on linear models obtained from a plasma-physical pet code under the action of the md perturbation. a typical transient with such testing is shown in fig. 3.3 for a controller without a cascade with pf currents, when, after a splash of gaps, they enter the specified tube, and the plasma position in the vertical and horizontal reach certain levels consistent with the plasma shape, when the vertical plasma velocity is stabilized about zero. the currents in the central solenoid and in the coils of the poloidal field come to some new level, corresponding to the compensation of disturbances. this controller showed the best robust properties when comparing it with other controllers for which the lqg approach and decoupling of control channels were used [4, 9]. a b fig. 3.2. block diagrams of controller synthesis: a) currents in cs & pf coils, b) plasma current and shape plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 89 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 3.3. variations of the output signals of the control system in case of md-type disturbance at the point of the sof (start of flattop) scenario: a gaps, vertical and horizontal plasma positions; b cs & pf currents (2) minimization of the h∞-norm of the mixed sensitivity function for the linearized plasma model while stabilizing the vertical plasma velocity about zero and simultaneously controlling the plasma shape and current for synthesizing robust scalar and multivariable controllers (see fig. 3.1) [12]. at the next stage, a siso and mimo controllers were synthesized while minimizing the h∞-norm of mixed sensitivity , where s(s) is the sensitivity transfer function of the closedloop system, k(s) is the controller. the block-diagonal controller was tested both on a linear model obtained by linearizing a plasma-physical dina code that implemented partial differential plasma equations, and on the nonlinear code itself, which distinguishes this work from the previous results. the plots of the processes obtained are similar to those shown in fig. 3.3. (3) cascade control (fig. 3.4 –3.6) [13–15]. further progress is related to the transition to cascade control and the solution of the problem of tracking the scenario values of gaps. fig. 3.4. structural scheme of the internal cascade (the lower level) for cs/pf coil current control ( ) ( ) ( ) ( ) ( ) ( ) 1 2 min k s w s s s w s k s s s ¥ ® 90 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) fig. 3.5. structural scheme of the external cascade (the upper level) for plasma shape and current control fig. 3.6. tracking of gaps g1–g6 here, in the inner loop, the original decoupling of control currents is carried out (see fig. 3.4). in the outer loop, the moore-penrose pseudo-inverse matrix establishes a connection between the displacements of the gaps between the first wall and the separatrix and the variation of the plasma current with control currents through the diagonal pii controller. at the same time, in order to avoid saturation of control currents, a nonlinear quadratic programming problem was solved (increasing system survivability) (see fig. 3.5). fig. 3.6 shows the operation of the cascade control system for the position, current and shape of the plasma with a decoupling of channels in the tracking mode for gaps on the dina code with a sufficiently high accuracy. (4) model predictive control (mpc). this control is attractive because the synthesis of the controller takes into account the restrictions on the input and output signals, which eliminates the development of specialized approaches, in particular, to prevent the saturation of plasma control currents. therefore, the proposed approach has no separation on two control loops of the plasma vertical velocity and plasma current and shape (fig. 3.7) [16]. this configuration takes into account limitations on the control currents. numerical modeling of the system with a predictive model on the non-linear dina code showed that, due to the natural consideration of constraints and the solution of the optimal problem at each control step, the system with a predictive model gives smaller deviations in the gaps when a md disturbance occurs [16]. plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 91 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 3.7. block diagram of a magnetic plasma control system in iter with a predictive model the predictive model approach for plasma control in tokamaks continues to evolve. thus, a scalar system for stabilizing the vertical position of the plasma in iter is developed and simulated relative to zero with prediction [17]. the method with a predictive model involving the singular decomposition of matrices [18] is used to control the current and shape of the plasma in iter. prediction is used to control the profile of safety factor q in iter with variable constraints on the raptor code [19], as well as to control the vertical position of the plasma on the t-15md tokamak model with a variable parameter [20]. (5) a cascade control system with decoupling of channels in the inner cascade for controlling currents in the poloidal field coils and robust controller in the outer cascade (fig. 3.8) [21]. cascade control of plasma shape and current leads to a trend towards increasing margines of robust system stability, which is associated with an increase in system reliability. to achieve this goal, the outer cascade was synthesized as a robust when solving the minimization problem of the h∞-norm of mixed sensitivity. then, on a linear model, the synthesized system was compared with the system of issure (3) according to the criteria of margins of robust stability and robust performance. the result showed that the robust system surpasses the system from issue (3) by about five times according to robust criteria. simulation of the robust system on the dina code in the tracking mode for gaps showed the result on tracking accuracy close to the system of issue (3) (see fig. 3.6). fig. 3.8. block diagram of cascade control system of plasma shape and current in tokamak: the keys provide the ability to switch contrtollers in the outer cascade (6) hierarchical plasma control with mimo robust control loop for plasma current and shape, synthesized by macfarley-glover method, and adaptive control circuit for the vertical position of the plasma with predictive model (fig. 3.9) [22, 23]. in the 92 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) previous approaches, the speed of the plasma vertical position was stabilized around zero, and not the position itself. this approach has two drawbacks. one of them is connected with the fact that this approach does not allow the plasma to maintain its position in a limited region of vertically controllability, since the limitation of this region arises due to the instability of the plasma and the restriction on the control voltage. another disadvantage is associated with the fact that with such control the system is not strictly stable, since the speed and not the position of the plasma in the vertical direction is stabilized. for these reasons, a system was developed to control the position, current, and shape of the plasma, in which the vertical unstable position of the plasma is stabilized. to avoid contradiction between the plasma shape and its position, an adaptive contour with a controller based on a predictive model was added, which adapts the position of the plasma to its shape (see fig. 3.9). рис. 3.9. block diagram of hierarchical system of magnetic plasma control with adaptation in iter: mpc controller with a predictive model a brief comparison of the applied approaches. iter used various principles and structures of control systems for magnetic control of plasma (1) (6). this is because the plasma in iter is a complex plant with distributed and time-varying parameters, having parametric and structural uncertainties, which makes it necessary to search for the most effective methods of multivariable hierarchical plasma control. first, to counter uncertainties of plasma models, robust control systems based on the h∞optimization theory (1), (2), (5) and (6) were developed. various system structures were proposed: (1) reflection of external disturbances; (2) minimization of the h∞-norm of the mixed sensitivity function; (5) decoupling of current control channels and h∞-controller for the shape of the plasma. all these schemes were used with the stabilization of the vertical plasma speed relative to zero, in order to eliminate contradiction between the position of the plasma and its shape. these schemes led to approximately the same perormance at control of current and plasma shape, providing certain robust stability margins due to h∞-controllers. systems with control of the vertical speed of the plasma, rather than its vertical position, are flawed: they are not strictly stable, and the vertical position can change uncontrollably in transients. for this reason, it was proposed to directly control the vertical position of the plasma in order to overcome this drawback in the vertical stabilization circuit (6). therefore, an adaptation of the vertical position of the plasma to its shape (6) was developed. at the same time, the systems with decoupling of control channels both in the inner cascade of currents control in the coils of the poloidal field and in the outer cascade of plasma current and shape (3) were developed. this is due to the fact that when decoupling the channels, it is easier to tune each control channel, while in comparison with the h∞theory of optimization a multivariable controller is synthesized while minimizing one h∞criterion, which makes it difficult to understand how the whole system works, although it gives large robust stability margins. a current, position, and shape control system with a predictive model has been applied (4). this approach is promising in that with it it is possible to take into account the restrictions on the input and output values of the controlled plant, since these restrictions are laid into the algorithm with prediction at the input and output horizons. in addition, a method plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 93 copyright ©2020 assa. adv. in systems science and appl. (2020) with a predictive model can be applied to create a multivariable adaptive plasma current and shape control system, if there is an opportunity to identify a linear plasma model at each time instant. 4. synthesis and modeling of plasma control systems in iter-2 with internal coil for controlling the vertical plasma position it has been established in [24] that the plasma controllability region on the vertical in iter is limited to a relatively small value namely 4 cm, which is of 2% of the minor iter radius. the idea of controllability region estimation is presented in fig. 4.1. it consists in the fact that under the same conditions and the same saturation voltage supplied to the coils of the horizontal iter field according to the scheme shown in fig. 4.1, b, the plasma was released from different initial conditions vertically, and the voltage sign was chosen such that it counteracted the increase in vertical displacement. if the initial condition was within the controllability region, then the voltage after a certain time interval caused the vertical displacement to change in the opposite direction. this study was conducted, in particular, on the corsica code (usa). four proposals were put forward and investigated [25] on how to increase the region of vertical plasma controllability: increase in voltage vs of the power supply from 6 to 9 kv; insertion of a second vs-circuit using two cs sections: cs2l and cs2u; adding stabilizing rings inside the vacuum vessel; insertion of special control coils inside the vacuum vessel. these proposals were studied and compared with each other in [25]. the most effective means turned out to be coils inside the vessel. for this reason, at the end of 2013, special coils were installed inside the vacuum vessel in the iter project in order to bring them closer to the plasma as much as possible, thereby expanding the controllability region vertically and increasing the reliability of the plasma magnetic control system (see fig. 4.1, b) [24]. a b fig. 4.1. investigation of the vertical controllability region in iter-2 and its expansion: a) experimental method of estimation of the controllability region dzmax [24]; b) coils of the horizontal field which are located inside the vacuume vessel to increase the controllability region; the power supply vs3 is connected to these coils that is controlled by the controller in the feedback [25]. two systems for stabilizing the plasma vertical speed with respect to zero with external and internal coils were investigated [26]. the authors used the jin-trac (jet) transport 94 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) code and the equilibrium create-nl code with the free boundary, which were connected together for modeling. in the first case, the control law was chosen as follows where vs1(s) is the laplace transform of the voltage function of the external coil, zdot(s) is the laplace transform of the function of the vertical speed movement of the plasma current center, and in the second case the control law in the laplace transform domain was chosen as follows: where vs3(s) is the laplace transform of the voltage function on the inner coil, is3(s) is the laplace transform of the current function in the inner coil. the case of scenario 1 was chosen, l-mode, sof (start of flattop: the beginning of the flat phase), ip = 15 ма, li = 0,88, bpol = 0,06. the simulation results are summarized in table 4.1 where vdemax (vertical displacement event: the phenomenon of vertical displacement) is the maximum possible initial position of the plasma in the vertical direction, which can be stabilized (controllability region), mj is the phase stability margin, wt is the cut-off frequency, mgu and mgl are the upper and lower limits of the stability margin in amplitude. the results showed that the area of control vdemax significantly more for the case of an inner coil, about 5-6 times. table 4.2 shows the values of vdemax in the presence of noise with a limited bandwidth of 1 khz, zero mean, standard deviation of σ = 210 mm/s and σ = 430 mm/s. the machine configuration is rather insensitive to noise: at the noise level in question, there is a negligible degradation in performance in terms of vdemax. table 4.1. characteristics of the control system obtained in the analysis vdemax, mm mj wt rad/s mgu, db mgl, db outer coil |34 22o 13 5.6 –4.0 inner coil > 200 65o 78 13 –18 table 4.2. the best performance in the presence of noise obtained in the analysis configuration vdemax, mm noise level s, mm/s 210 430 vs1 (6 kv) 28 24 vs3 (2,3 kv) 148 141 in iter scenarios are continued to be developed and refined using plasma magnetic control systems on plasma-physical codes, taking into account only external coils for controlling the vertical position of the plasma [27], and taking into account the internal coil [26, 28]. fig. 4.2 shows the block diagram of the plasma shape control system, which was used in [26] to simulate transients during the transition from mode l (15 ma) to mode h (15 ma) and vice versa on the linear model create-l. the paper [26] presents the evolution of the internal g6 and external g3 gaps (see fig. 1.1, b) for the cases: the direct link is updated at t = 1 s; direct link is updated at t = 5 s; the direct link is updated at t = 1 s exactly after the transition and then at t = 9 s; the direct link is updated at t = 1 s and a transport delay of 0.2 s; direct link is calculated as a function of bpol. ( ) ( ) ( )1 2 1 /1815000 1 / 60 dot svs s z s s + = + ( ) ( ) ( )5 3 3 1 / 408000 1,2 10 1 / 6 dot s svs s z s i s s -+ é ù= ´ ´ë û+ plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 95 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 4.2. block diagram of the plasma shape control system, similar to the extreme shape controller scheme in the jet tokamak (uk) studies of the plasma vertical controllability region for iter according to the method in [24] were continued in [27] on the pet code in more detail for various plasma variants. the results are summarized in table. 4.3, which shows that when using only external coils and vs1 power supplies, the range of the controllability region is in the range of 1.5–13.7 cm. if both external vs1 coils and internal vs3 coils are involved, then the range increases noticeably up to 16–25.5 cm. table 4.3. performance indicator max (z0) and instability time constant [26] plasma 15 ма plasma 1 li(3) = 1.2, βp = 0.1 plasma 2 li(3) = 1.0, βp = 0.1 plasma 3 li(3) = 0.73, βp = 0,6 maxz0, vs1 15 30 137 vs3+vs1 160 175 255 instability time constant, ms 56 78 171 the systems of magnetic plasma control have been investigated taking into account noise and various combinations of coils for stabilizing the vertical position of the plasma in iter on the dina code [27]: external coils vs1; vs3 internal coils; simultaneous use of external and internal coils. the studies were based on limited data analysis from c-mod, jet, and asdex-u tokamaks, which made it possible to specify for plasma activity in iter a model of vertical plasma speed noise dz/dt as a random signal evenly distributed over the frequency interval [0, 1] khz with a standard deviation of 0.6 m/s. the simulation was performed with low-frequency noise injection into the diagnostic signal dz/dt, which was used in the feedback of the vertical plasma stabilization. at the same time, the standard deviation of the uniformly distributed noise was increased until the vertical stabilization system lost stability or one of the engineering parameters reached its design limit. the best result was given by the combination of the windings of the internal and external coils. in this case, two signals were applied to the input of the vertical stabilization controller: dz/dt and ivs3. the maximum allowed standard deviation for noise in this case was 3 m/s, the standard deviation of the ivs3 current was about 10 ka, the standard deviation of the vertical plasma displacement during the flat phase of the plasma current was about 44 mm. the plasma control system was synthesized using the standard lqg method [29]. conclusion in this part of the review, attention is paid to the development and modeling of plasma position, current, and shape control systems in iter, and the contribution of v.a. trapeznikov institute of control sciences of russian academy of sciences to this work is highlighted. the results show a trend in the development of these systems, associated with an increase in accuracy and speed when tracking scenario signals and rejecting disturbances such as minor disruption. the trend reflects increase in the robust stability margin, which in turn leads to increase in reliability and survivability of plasma control systems in iter. 96 y.v. mitrishkin, n.m. kartsev, a.e. konkov, m.i. patrov copyright ©2020 assa adv. in systems science and appl. (2020) the introduction of horizontal field coils inside the iter vacuum vessel expands the controllability region in the vertical direction of plasma movement, which, at the given the control coil voltage limits, specifications for minor disruption disturbances and vertical plasma instability, significantly moves away the closed-loop system of magnetic plasma control from the boundaries of stability loss. magnetic plasma control systems in operating tokamaks, as well as for iter, continue to evolve in different ways. one of these areas is related to the development of highprecision methods for solving inverse plasma diagnostic problems, the data of which serve as input data for plasma control systems for internal plasma parameters, and approbation of control systems based on these data, rather than direct data, when all plasma parameters are known accurately enough [30, 31]. in section 43.2, experimental processing of scenarios for iter on diii-d (usa) and west (france) tokamaks, approaches in modeling and implementing plasma control systems in iter, preparation of the plasma control system for launch and operation in iter will be presented. road maps of the development and creation of the first demo fusion power station (following the iter step) will be shown, which indicate two demo development directions: on traditional tokamaks with relatively large aspect ratios and modular-type spherical tokamaks that will significantly reduce the creation time of demo. it will be also shown the main trends in the development of demo poloidal systems, as well as the initial version of the magnetic plasma control system in demo. acknowledgements this work was supported by the russian science foundation (rsf), grant no 17-1901022 (sections 2-5) and the russian foundation for basic research (rfbr), grant no 1708-00293 (introduction, section 1). references 1. mitrishkin, y. v., korenev, p. s., prokhorov, a. a., et al. (2018). plasma control in tokamaks. part 1. controlled thermonuclear fusion problem. tokamaks. components of control systems, advances in systems science and applications, 18 (2), 26-52. 2. mitrishkin, y. v., kartsev, n. m., pavlova, e. a., et al. (2018). plasma control in tokamaks. part. 2. magnetic plasma control systems, advances in systems science and applications, 18(3), 39-78. 3. iter. – url: https://www.iter.org. 4. mitrishkin, y. v. (2006). plasma control in experimental thermonclear devices: adaptive self-oscillations and robust control systems. (upravlenie plazmoy v ehksperimentalnyh termoyadernyh ustanovkah: adaptivnye avtokolebatelnye i robastnye sistemy upravleniya). – moscow: krasand, 400 p. [in russian]. 5. takase, h., utohab, h., sakamoto, y., et al. (2016). analysis of plasma position control for demo reactor, fusion engineering and design, 109–111, 1386-1391. 6. mitrishkin, y. v., lukash, v. e., khayrutdinov, r. r. (2005). plasma current, position and shape robust control system in iter, problems of atomic science and technology. series: thermonuclear fusion, 1, 61-81 [in russian]. 7. mitrishkin, y. v. (2004). comprehensive design and implementation of plasma adaptive selfoscillations and robust control systems in thermonuclear installations, proc. of 8th world multi-conference on systemics, cybernetics and informatics, orlando, fl, usa, xv, 247-252. 8. mitrishkin, y. v. & kimura, h. (2001). plasma vertical speed robust control in fusion energy advanced tokamak, proc. of 40th ieee conf. on decision and control, florida, usa, 1292– 1297. 9. mitrishkin, y. v., kurachi, k., kimura h. (2003). plasma multivariable robust control system design and simulation for a thermonuclear tokamak-reactor, international journal of control, 76 (13), 1358-1374. 10. mcfarlane, d., glover, k. (1990). robust controller design using normalized coprime factor plant descriptions. lecture notes in control and information sciences, springer-verlag, 212 p. 11. safonov, m. (1999). robust control / encyclopedia of electrical and electronics engineering (ed. j.g. webster), vol. 18. ny: wiley, 592-602. plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter 97 copyright ©2020 assa. adv. in systems science and appl. (2020) 12. dokuka, v. n., kadurin, a. v., mitrishkin, y. v., khayrutdinov, p. p. (2007). synthesis and modeling of the h∞-system of magnetic control of the plasma in the tokamak-reactor, automation and remote control, 68(8), 1410-1428. 13. mitrishkin y. v., korostelev a. y., sushin i. s., et al. (2009). plasma shape and current tracking control system for tokamak, proc. of 13th ifac symposium on information control problems in manufacturing, c7.1, 2133–2138. 14. mitrishkin, y. v., korostelev, a. y. (2010). a cascade system of tracking plasma current and shape of a tokamak with decoupling control channels (kaskadnaya sistema slezheniya za tokom i formoy plazmy v tokamake s razvyazkoy kanalov upravleniya), vestnik bmstu, 79(2), 21-38 [in russian]. 15. mitrishkin, y. v., korostelev, a. y., dokuka, v. n., khayrutdinov, r. r. (2011). synthesis and simulation of a two-level magnetic control system for tokamak-reactor plasma. plasma physics reports, 37(4), 279-320. 16. mitrishkin, y. v., korostelev, a. y. (2008). system with predictive model for plasma shape and current control in a tokamak, control sciences, 5, 19-25. [in russian]. 17. veremei, e. i., sotnikova, m. v. (2011). plasma stabilization on the base of model predictive control with the linear closed-loop system stability, vestnik s.-petersburg univ. ser. 10. prikl. mat. inform. prots. upr, 1, 116-133. [in russian]. 18. gerkšič, s., pregelj, b., perne, m., et al. (2018). model predictive control of iter plasma current and shape using singular value decomposition, fusion engineering and design, 129, 158-163. 19. maljaars, e., felici, f., baar, m. r., et al. (2015). control of the tokamak safety factor profile with time-varying constraints using mpc, nuclear fusion, 55(2). 20. mitrishkin, y. v., golubtsov, m. p. (2018). a hybrid control system for an unstable nonstationary plant with a predictive model, automation and remote control, 79(11), 2005-2017. 21. kadurin, a. v., mitrishkin, y. v. (2011). multidimensional system of cascaded control of plasma form and current in tokamak with channel decoupling and h∞-controller, automation and remote control, 72, 20-53. 22. mitrishkin, y. v., kartsev, n. m. (2011). hierarchical plasma shape, position, and current control system for iter, proc. of the 50th ieee conf. on decision and control and european control conf., orlando, fl, usa, tua13.2, 2620-2625. 23. kartsev, n. m., mitrishkin, y. v., patrov, m. i. (2017). hierarchical robust systems for magnetic plasma control in tokamaks with adaptation, automation and remote control, 78(4), 700-713. 24. humphreys, d. a., casper, t. a., eidietis, n., et al. (2009). experimental vertical stability studies for iter performance and design guidance, nuclear fusion, 49(11). 25. albanese, r., ambrosino, g., ariola, m., et al. (2009). iter vertical stabilization system, fusion engineering and design, 84, 394-397. 26. cavinato, m., ambrosino, g., figini, l., et al. (2014). preparation for the operation of iter: eu study on the plasma control system, fusion engineering and design, 89, 2430-2434. 27. casper, t., gribov, y., kavin, a., et al. (2014). development of the iter baseline inductive scenario, nuclear fusion, 54(1). 28. gribov, y., kavin, a., lukash, v., et al. (2015). plasma vertical stabilisation in iter, nuclear fusion, 55(7). 29. belyakov, v., kavin, a., kharitonov, v., et al. (1999). linear quadratic gaussian controller design for plasma current, position and shape control system in iter, fusion engineering and design, 45, 55-64. 30. zaitsev, f. s., shishkin, a. g., lukianitsa, а. а., et al. (2018). the basic components of software-hardware system for modeling and control of the toroidal plasma by epsilon-nets on heterogeneous mini-supercomputers, commun. comput. phys., 24 (1), 1-26. 31. shishkin, a. g., stepanov, s. v., suchkov, e. p., et al. (2018). distributed access to the resources of plasma modelling and control complex hasp cs, 2018 ieee international conference on computer and communication engineering technology (ccet), beijing, china, august 18-20, 193-197. microsoft word 761 enhancing hadoop performance in homogenous and heterogeneous big data environment adv syst sci appl2020; 01; 13-26 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/761 enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration ekhlas k. hamza 1) control and systems engineering department, university of technology, iraq e-mail: 100374@uotechnology.edu.iq, drekhlaskadhum@gmail.com, abstract: hadoop is one of the most famous platform solutions for processing large volume and scale of data in parallel processing in cloud computing. a hadoop system can be characterized based on three main factors: cluster, workload and user. each of these factors can be described as either heterogeneous or homogenous, which reflects the heterogeneity degree of the hadoop systemthe objective of this proposed research work is to investigate the degree of influence of heterogeneity for each of these factors on the performance of hadoop based on different schedulers. three schedulers are considered with different levels of hadoop heterogeneity and are tested and analyzed: the first algorithm considered is the fifo (first in first out), the second is the fair sharing, and the final is the coshh (classification and optimization based scheduler for heterogeneous hadoop). performance issues are related to hadoop schedulers and comparative performance analysis between different cases of jobs submission. these jobs are processed in different homogenous or heterogeneous data environments and under fixed or reconfigurable slot between map and reduce tasks for hadoop mapreduce java programming clustering model. the results showed that when assigning tunable knob between map and reduce tasks under certain schedulers like fifo algorithm, the performance enhanced significantly especially in cases of heterogeneity environment where the workload decreased significantly and the utilization of computational resources increase was obvious. keywords: hadoop, mapreduce, big data, slot configuration, scheduling algorithms, makespan, resources utilization. 1. introduction recently, big and parallel data processing has become a main concern for modern analysis and knowledge extraction for feasible prediction and decision making in many fields of businesses. data has recently become more complicated when generated in huge volumes and in many cases mixed between structure and unstructured data. thus, analyzing it exposes many challenges. moreover, modern data is acquired from different server sources and from different users in real-time, which add more complexity to the analysis processes [1],[2]. analyzing these big complicated and parallel data is always beneficial in terms of mining and extraction of useful information and for knowledge discovery purposes [3]. recently, hadoop has been considered as the most promising open source platform for the big and parallel data processing in real-time [4]. understanding the hadoop architecture and how it works is a challenge by itself. the target of utilizing hadoop is always to process the big and parallel data thereby reducing the computational costs of the usage of the processors and memory resources [5] as well as clustering efficiently the data in homogeneity and heterogeneity classifications. the accuracy of the data classification is aimed at researching 14 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) this topic, as well as reducing the latency time when answering the end users’ enquiries for the purposes of data mining or discovery of certain knowledge [6]. from the beginning, this research has been motivated from certain challenging factors, which are as follows: firstly, simulating the real-time hadoop homogenous and heterogeneous environments in one desktop. secondly, to simulate different users’ enquiries and even if they are heterogeneous users it may mean loading the data to hadoop from different distributed file systems and servers. besides that, the makespan or alternatively the workload should be determined especially for job enquiries submitted to hadoop from different distributed file systems and the data obtained from different heterogeneous clusters [7]. the focus here is to analyze the volume of workload assigned to each job, which is assigned to the hadoop system by determining the consumption of the processors and memory sources as well as the time taken for completing the job. most important of all is how to reduce the makespan or the workload in hadoop mapreduce architecture, especially for obtaining data from heterogeneous clusters loaded from different distributed file systems (servers). mainly fixing the number and the ratio between map and reduce processes by using traditional scheduling algorithms such as fifo and fair sharing algorithms seems to have some difficulties as some jobs maybe stuck sometimes, especially in case of fifo scheduling [8]. further, the job convergence is not optimized especially in the case of dealing with intensive heterogeneous data environments. this will lead to increased workload (make span), which consumes more computational resources and resulting in more time processing. the proposed system implements a mechanism to dynamically allocate slots for map and reduce tasks or alternatively it means optimally adjusting the ratio between map and reduce processes by using more efficient scheduling algorithms for optimized classification in a heterogeneous data environment. this, as the results show, will decrease the make span and more efficiently utilize the processors and memory resources. 2. research background and literature review in this section, the main concepts of this research field are defined and explained to understand the system architecture of the proposal and how it works in processing and analyzing big and medical data in real time. at this point it is important to understand from the beginning the characteristics and definition of the modern big data. big data can be identified mainly in three parameters, which are complexity, volume and dynamicity. as regards the complexity parameter, this will make sense when it is understood that the generated data in recent modern scenarios are coming from different sources and formats. this can be text, audio, videos, symbols or numbers, which are different extensions and formats. it is clear that dealing with such unstructured data will impose enormous complexity for purpose of analysis and data mining. there should be efforts during pre-processing to make such data more uniformed in order to make pattern recognition feasible. with reference to the volume parameter, the volume and the size of the data are increased rapidly, dynamically and spontaneously in real-time. this can become clear if we imagine that in modern scenarios of medical data collection to support personal health records and monitoring health status remotely, the data are collected by different sensors for different health parameters such as heart rate, blood pressure, body temperature and glucose levels. all these data are collected simultaneously and in parallel in one database center for further analysis so as to send certain pattern recognition or trend analysis as results for end users to allow them to monitor their health conditions. in this case the data is growing exponentially in size, volume and dynamically in real-time with different formats and structures. in the centralized management efforts of the big data analysis, another added complexity factor is represented as the data come from different distributed sources and in many cases enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 15 copyright ©2020 assa adv. in systems science and appl. (2020) from heterogeneous sources. hadoop recently appeared as the most interesting open-source software platform utilized for distributed storage and parallel processing of big data by using the mapreduce programming model [9]. the main parts of the hadoop architecture consist of the storage part, which is known as hadoop distributed file system (hdfs) and the processing part, which is the mapreduce programming model. hadoop’s main job is to split large blocks of data and distribute them across nodes in different clusters in parallel and simultaneously [10]. the simplest cluster contains one master node and multiple slave nodes. the master node keeps tracking the tasks acomplished in the slave nodes by the job tracker in each master node.some researchers have discussed scheduling techniques in [ 12-15]. they have provided scheduling-aware data, prefetching and eviction mechanisms based on spark, alluxio and hadoop. the main advantage of the mapreduce programming model in hadoop is scaling the data processing over multiple computing nodes. first, the map process takes a set of data and converts it to another form by simply breaking the data into (key/value pairs) format. thereafter, the reduce process takes the output from the map task as input and then shuffles the data combined, as shown in figure 1. the jobs for assigned jobs of hadoop mapreduce are scheduled from the beginning based on certain scheduling algorithms such as first input first output (fifo), fair sharing and others. fig. 1. hadoop map and reduce processes [10] it is always desirable to reduce the number of maps and processes so as to decrease the workload in hadoop processing and that depends on the power of the scheduling algorithms as well as the ratio between maps, which reduces implementation. this research work focuses in reducing the workload of the hadoop processing jobs (referred to as same meaning as make span) by adjusting the ratio between maps and reducing the optimal minimum number to achieve the different tasks even in homogenous or heterogeneous data clustering to answer different end-users’ enquiries. the incoming jobs submitted to the hadoop system can be considered heterogeneous with respect to certain factors such as the number of tasks, data volume, required computation, arrival rates and the execution times. it is worthy to notice that hadoop resources are different in capabilities of data storage as well as the processing units. this leads to different priorities and minimum requirements for sharing between different users taking into consideration that the number and type of jobs assigned to each user can be different as well. the mean execution time for certain jobs is considered as indicator of the heterogeneity degree for both the workload and clusters in the hadoop system. mainly fixing the number and the ratio between map and reducing processes by using the traditional scheduling algorithms such as fifo and fair sharing algorithms, seem to have some difficulties as some jobs maybe stuck sometimes, especially in case of fifo scheduling. moreover, the job convergence is not optimized especially in case of dealing with intensive heterogeneous data environment. this will lead to increase in workload 16 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) (makespan), which consumes more computational resources and resulting in more time processing. table 1 shows the comparisons of various scheduling techniques. table 1. show the comparative a of various scheduling techniques disadvantages advantages environment heterogeneous or homogenous idea of implementation preemp tion working taxonomy scheduler 1-considered only for singlekind of work. 2. low performance when run manykinds of works. 3. poor response periods for short worksrelated to large works. .1-rate of entire cluster arrangement method is less 2. simple to implement andcapable homogenous schedule works based on their significances in first-come first out. no better work with small clusters nonadaptive fifo .1-does not consider thework weight of each nodule. 1-less difficult. 2. mechanisms well when both lesser and large clusters. 3. it can offerfast responseperi ods for small jobs mixed with larger works. homogenous do a equal spreadingof calculatepropertie s among the users/jobs in the system yes better work with small clusters adaptive fair scheduling 1 . improve the overall system per formance. 2. addresses the fairness and the minimum share requirements. heterogeneous planned to improve the mean completion time of jobs yes better work with large clusters adaptive coshh 3. architecture of the proposed system in both the homogeneous and heterogeneous hadoop environments, the two modules are initially added to the job tracker, which is described below.  workload monitoring (wm)  slot assignment (sa) the function of wm is to collect previous workload information asexecution times are needed for completed tasks. based on this collected information, the workload of the current map and reduced tasks are estimated. thereafter, the role of sa is to adjust the slot ratio between map and reduce tasks by using estimated information received from wm. this adjustment is performed on each slave node. based on adjusting the slot ratio as well as the current slot status, the job tracker is used to assign tasks to slave nodes by both the workload monitor and slot assigner functions. in addition, the task tracker module along with the task manager function for checking the number of individual maps, reduce tasks running on the node. the performance of checking is done based on the new slot ratio received from the job tracker by running the task on each slave node.fig. 2 shows the architecture of the proposed method a batch of jobs is submitted from a user to the job tracker. the three modules involved in the job tracker are workload monitor, slot assigner and scheduler. the status of slot enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 17 copyright ©2020 assa adv. in systems science and appl. (2020) assignment and task assignment is communicated between job tracker and task tracker in order to perform the data communication with large-scale information. the task is split into several parts, which are functioned by different slots. thus, various slot configurations are performed and each slot contains separate task tracker and task manager. the four functions involved in the proposed architecture are:  the present workloads and their functions are estimated.  the best slot assignment of each node is decided in order to communicate between the job tracker and task tracker.  the task is assigned to each slave node.  constant monitoring of the task execution as well as the slot occupation situation. fig. 2. architecture of the proposed method [11] the modules involved in the proposed method to minimize the makespan are as follows:  homogenous environment  heterogeneous environment  makespan analysis  slot configuration prediction  evaluation 3.1. homogenous environment the homogenous system is defined as the homogenous characteristics of both workload and cluster. based on the job sizes,the homogenous system can be classified into either homogenous-small or homogenous-large. if all the jobs are small, then the system is defined as homogenous-small. if all the jobs are large, then the system is classified as homogenouslarge. mostly, the job size parameters affect the performance of the hadoop schedulers. the average completion time of all schedulers is almost equal. as both the cluster and workload are homogenous, the coshh algorithm suggests all resources as the best choice for all job classes. the two conditions used to develop homogenous computing environment are as follows:  the same storage representations along with the same results are considered as correct in hardware and software module of processor for operations on floating point numbers. 18 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020)  the correct floatingpoint value is transmitted and guarantee is provided to the communication layer during the process of floating point number between processors. the module involved in the homogenous phase is described as a flow diagram, as shown in figure 3. the five representations of data-analyzing hadoop benchmarks are described in the heterogeneous phase as a module for makespan and slot configuration prediction. the benchmarks are derived from mapreduce benchmarks suite, which are described below.  inverting the index task: generate certain list of the words as document indexing when entering a text file as an input.  rating the frequency to generate the histogram: as example, calculating the histogram from certain input of movie rating data.  counter for certain word tasks: this job can be described as counting the occurrence of each word when entering input text-file.  classifier: when inputting the movie rating data then, classify these movies into predefined clusters.  grepping task: to find certain patterns on the files from certain input text files. fig. 3. flow diagram of homogeneous environment the implementation of the proposed system architecture is achieved by simulating the medical unstructured big data in cloud environment developed in a single desktop using java programming in netbeans ide and wamp server services all working together, as depicted in figure4. enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 19 copyright ©2020 assa adv. in systems science and appl. (2020) fig. 4. developed cloud computing platform for medical big data of homogenous tasks 3.2. heterogeneous environment based on the workload and cluster factors, the user is either classified as homogenous or heterogeneous. here, these factors are considered as heterogeneous, thus, the user is also selected from a heterogeneous system.the challenging issue for the schedulers in this system is the arrival rate of different size jobs. based on the challenges,the system is classified into three such as heterogeneous-small, heterogeneous-equal and heterogeneous-large. if the arrival rate is high for small jobs, then the system is classified as heterogeneous-small. if the arrival rate is equal for all job sizes, then the system is classified as heterogeneous-equal. if the arrival rate is high for large jobs, then the system is classified as heterogeneous-large. due to the resolving resource and job mismatch problem schedulers algorithm, the average completion time is intended to be improved in all three heterogeneous systems using the proposed architecture. in the proposed method, the modules are divided into three sectionsas follows:  query process  three different queries  sub-query the processes involved in the proposed heterogeneous platform of medical big data is illustrated in the figure below. 20 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) fig. 5. processes flow diagram of heterogeneous environment enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 21 copyright ©2020 assa adv. in systems science and appl. (2020) the implementation of the heterogeneous part of the proposed architecture methodology is depicted from the clients’ enquiry side as well as the distributed servers’ side, as shown infigures 6 and 7. fig. 6. client query from heterogenous distributed server in medical big cloud environment fig. 7. processed data from heterogenous distributed servers in medical big cloud environment 3.3. makespan prediction the proposed algorithm tunable slot configurations for reduced makespan (tuscrd) is used to improve the performance of a batch of mapreduce jobs by adjusting a basic system parameter with the objective of the proposed system. the significant challenge available in the cloud is a prediction of makespan. the three modules involved in the makespan prediction are described below:  launch a hadoop cluster  tuning the slot assignment  tunable slot configurations for reduced makespan (tuscrd) under fifo a classic hadoop cluster consists of multiple slave nodes and a single master node. the master node executes the jobtracker section, which acts for the scheduling jobs along with 22 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) coordination between each job’s executions of tasks. each slave node executes the tasktracker part from the architecture for hosting the execution of mapreduce jobs. the task is stored as input data in the data node, which is derived as local whereas the remote is denoted by storing all the tasks in another node. if the data nodes are unknown, then the reduce tasks are considered as remote on all nodes and map tasks are considered as scheduler on the input node. every task is non-preemptable and shares the same deadline and release time as its job. initially, the system model is considered without failure and with pipeline/speculative execution. then, jobs are submitted continuously to the cluster and the master scheduler needs to schedule their tasks on the slaves without prior knowledge of future job arrivals. fig. 8. flow diagram of slot configurations 4. performance analysis of homogenous hadoop submitted jobs makespan and resource utilization are calculated once any homogeneity tasks such as line indexer, word counting and classification are clicked or the heterogeneity task of enquiring about medical data from different distributed servers. figure 9 shows an example of the makespan and cpu processing time as well as the memory space consumed during processing. fig. 9. evaluation of the inverted index homogenous task in term of makepsan and processing time enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 23 copyright ©2020 assa adv. in systems science and appl. (2020) table1explains thecomparative values of makespan between the inverted index under tunable slot configuration and the existing system. this shows that the inverted index is slightly higher compared tothe existing system as the medical big data are loaded from the same file. table 1. performance of tunable slot solution in hadoop homogenous inverted task on the other hand,tables 2 and 3 show significant enhancements of reducing the makespan workload where the medical big data are loaded from different distributed files. table 2. performance of tunable slot solution in hadoop classification table 3. performance of tunable slot solution for word count it is clear even in the homogenous hadoop big data environment but when loading the data from different files the makespan load is reduced significantly. in the second part of the implementation of the tunable slot configuration of the mapreduce hadoop, the environment considered is heterogeneous medical big data environment. in these cases, the end-clients can have different specific enquiries and the data are loaded and processed in different distributed servers.thereafter, the makespan workload and resources utilization in terms of cpu processing and memory space are determined, as shown in the following figure. 24 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) fig. 10. client query from heterogenous distributed server workload and resources utilization table4 below shows that the workload significantly reduces, in case of loading the data from heterogeneous distributed files from distributed servers with the proposed technique of tunable slot configuration for the hadoop mapreduce platform of processing the big data. table 4. makespan work load analysis of answering end-client enquires from heterogeneous hadoop big data environment. in brief from the presented data analysis, it clearly shows that allowing the slot between map and reduce processes to be tunable under certain scheduling algorithm in hadoop big data environment, will result in significant workload reduction and computational resources utilization, especially when the heterogeneity level increased when loading the data from different distributed files and servers. recent researchers [16-21] have been presented relating to the big data mobile network. they have discussed the ultimate research of wireless classifications for big data in the structure of a smart factory, which are a) mobile nodes and b) animatedly changing heavy traffic load. these can cause differences in traffic and energy consumptions of cluster head nodes in categorized networks. however, supporting the program big data transmission over 5g wireless systems forces many new and open challenges because the program big data services are both time-subtle and bandwidth-intensive over time-varying wireless channels with controlled wireless resources. the new generations of mobile devices have high enhancing hadoop performance in homogenous and heterogeneous big data environments by dynamic slot configuration 25 copyright ©2020 assa adv. in systems science and appl. (2020) processing power and storage but they lag behind in terms of software systems for big data storage and processing. hadoop is a accessible platform that arrange for spread storage and computational capabilities on clusters of commodity hardware. 5. conclusion and future works the problem of high workload and therefore, increased computational resources in hadoop mapreduce platform for big data processing, is an attempt to mitigate the optimized hadoop solution by allowing the slot configuration between maps and reduce the tunable according to certain job priority or homogeneity as well as the available resource spontaneously. the simulation of the proposed system shows that a high percentage of reduced workload and increased utilization of the computational resources can simply be observed by the lesser number of map and reduce tasks needed to accomplish certain jobs. this is shown in the proposed solution for three jobs submitted in a homogenous hadoop environment where the data in some cases are loaded from different distributed files in different servers as well as for different job queries in heterogeneous hadoop clustering environment as the data is loaded as well from distributed files. it is recommended to test as well the other scheduling algorithms such as fair sharing and classification and optimizationbased scheduler for heterogeneous hadoop (coshh). it is expected that the performance will be more optimized specially for the coshh scheduling algorithm with dynamic slot configuration. references 1. bhosale, h. s., gadekar, d. p. (2014). a review paper on big data and hadoop. international journal of scientific and research publications, 4(10), 1. 2. big data. (2015). new delhi: springer, india, private ltd. 3. hilbert, m. (2016). big data for development: a review of promises and challenges. development policy review, 34(1), 135-174. 4. che, d., safran, m., peng, z. (2013, april). from big data to big data mining: challenges, issues, and opportunities. in international conference on database systems for advanced applications (pp. 1-15). springer berlin heidelberg. 5. o’driscoll, a., daugelaite, j., sleator, r. d. (2013). ‘big data’, hadoop and cloud computing in genomics. journal of biomedical informatics, 46(5), 774-781. 6. suthaharan, s. (2014). big data classification: problems and challenges in network intrusion prediction with machine learning. acm sigmetrics performance evaluation review, 41(4), 70-73. 7. raju, r., amudhavel, j., pavithra, m., anuja, s., abinaya, b. (2014, march). a heuristic fault tolerant mapreduce framework for minimizing makespan in hybrid cloud environment. in green computing communication and electrical engineering (icgccee), 2014 international conference on (pp. 1-4). ieee. 8. rasooli, a., down, d. g. (2014). coshh: a classification and optimization based scheduler for heterogeneous hadoop systems. future generation computer systems, 36, 1-15. 9. shvachko, k., kuang, h., radia, s., chansler, r. (2010, may). the hadoop distributed file system. in mass storage systems and technologies (msst), 2010 ieee 26th symposium on (pp. 1-10). ieee. 10. mctaggart, c. (2014). hadoop/mapreduce. object-oriented framework presentation, csci5448. 26 ekhlas k. hamza copyright ©2020 assa adv. in systems science and appl. (2020) 11. dwivedi, k., & dubey, s. k. (2014, september). analytical review on hadoop distributed file system. in confluence the next generation information technology summit (confluence), 2014 5th international conference(pp. 174-181). ieee. 12. frederic nzanywayingoma, yang yang (2018) task scheduling and virtual resource optimisingin hadoop yarn-based cloud computing environment int. j. cloud computing, vol. 7, no. 2, 2018 13. ibrahim abaker, targio hashem, et al (2018) mapreduce scheduling algorithms: a review. the journal of supercomputing 10 december 2018 14. sa'ed abed, duha s. shubair (2018) enhancement of task scheduling technique of big data cloud computing 2018 international conference on advances in big data, computing and data communication systems (icabcd) 15. xiaomin li, di li, song li, shiyong wang, chengliang liu (2017) exploiting industrial big data strategy for load balancing in industrial wireless mobile networks ieee access (volume: 6) 29 december 2017 16. xi zhang, qixuan zhu (2019) information-centric virtualization for software-defined statistical qos provisioning over 5g multimedia big data wireless networks ieee journal on selected areas in communications (volume: 37, issue: 8, aug. 2019) 17. johnu george, chien-an chen, radu stoleru, geoffrey g. xie (2019) hadoop mapreduce for mobile clouds. ieee transactions on cloud computing (volume: 7, issue: 1, jan.-march 1 2019) 18. sudip misra, ayan mondal, swetha khajjayam (2019) dynamic big-data broadcast in fat-tree data center networks with mobile iot devices ieee systems journal ( volume: 13, issue: 3, sept. 2019) 19. muhammad rehan raza, ahmad rostami, lena wosinska, paolo monti (2019) a slice admission policy based on big data analytics for multi-tenant 5g networks journal of lightwave technology (volume: 37, issue: 7, april1, 1 2019) 20. cristo suarez-rodriguez, ying he, eryk dutkiewicz (2019) theoretical analysis of rem-based handover algorithm for heterogeneous networks ieee access (volume 7) 17 july 2019 advances in systems science and application (2015) vol.15 no.2 148-163 living and banking systems comparison: prisoners’ dilemma “win-win” is not the solution pierre bricage international academy for systems and cybernetic sciences, iascys, vienna, austria abstract to survive that is ‘to eat and not to be eaten’ so as to live on. whatever its spatial and temporal level of organization, every living system owns 7 invariant qualitative degrees of freedom. any living system is formed by embedments and juxtapositions of pre-existing systems. the same goes for man banking systems! how are the local quantitative laws of the spatial-temporal structuring and functioning of banking systems associated with the basic law of survival of living systems? how do the local actors become mutually integrated into their global whole? and reversely (systemic constructal law), why and how is the global whole reciprocally integrating the local parceners? is victory a strategic success? what are the roots of interdependence, conflicts and strategic order challenges?how is emerging a new power balance? can banking systems survive as parasitic systems? like a “food chain” is a “money chain” a way of violence escalade? the evolution of living systems is often seen as a “cooperative evolution” resulting from altruist behaviours. it could be modelled and simulated using games such as the prisoners’ dilemma game, a game that shows why 2 individuals might not cooperate, even though it appears to be in their best interests to do so. is the prisoners’ dilemma game justifying extortion? what can we learn from reinforcement learning dynamics in social dilemmas? in reality, humans display a systematic bias towards cooperative behaviour, much more so than predicted by models of “rational” self-interested action. models based on different kinds of payoffs and driving forces (where people forecast how the game would be played if they formed coalitions to maximize their forecasts) are shown to make better predictions which resemble reality. keywords parasitism, pareto equilibrium, ponzi pyramid, prisoners’ dilemma game. 1 introduction for living systems to survive that is first ‘to eat and not to be eaten’ [1]. in a predator-prey relationship to survive the predator must eat preys but not too much. the predator’s survival is limited by its prey’s limitations of survival [2]. soon or late the preys are to be eaten. but sometimes preys can win and even eat predators. that is the case of bacterial preys which eat predator amoebas, mycobacteria or cancer cells which destroy their predator immune cells. whatever its level of organization, to live on, any living system has ‘to be lucky’ so advances in systems science and application (2015) vol.15 no.2 149 as ‘to be at the right place at the right time’ [3]. whatever is its spatial and temporal level of organization, to survive and live on it owns 7 invariant qualitative characteristics [1]: figure 1. to survive every living system is a particular place for mobilization of matter and energy. entering flows (inputs) are used to make products and wastes which are then used, stored (throughputs) or excreted (outputs). if the outputs/inputs balance allows an internal accumulation of matter and energy, the system may grow in mass. growth (quantitative increase) is not a goal in itself but it is always a prerequisite phase for development (acquisition of new qualitative capabilities) and reproduction (figure 1). formed by embedments and juxtapositions of pre-existing systems in a new whole (endophysiotope), a living system is always part of a food chain; it eats and is eaten, within an ecoexotope of survival it shares with other living systems (fig.1). ‘soon or late it is impossible not to be eaten.’ man is not an exception [2, 4-5]. man species is a champion for enhancing growth of domestic plants fig. 1 living systems structuring and functioning: interaction is construction [1, 6, 10] and animals but for its own growth [4-6]! man species is ensuring its survival through the increase of the hosting capacity of its ecoexotope of survival (figure 1). the same goes for man societies which are often guided by quantitative economic considerations such as “saving more and more money” [7-8] rather than 150 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... by qualitative ethic considerations: “enough food for everyone everywhere”. living systems are food producing systems. we do know how they work [5, 9]. the modularity of living systems allows both a partial location and a global recycling of matter and energy [1]. the pleiotropy of the structures and functions, allowing ‘to kill two birds with one stone’, is a mechanism of exaptation. within any ecoexotope, the agoantagonistic relations balance soon or late ends in the disappearance of predators and a reduction of biodiversity [6]. only the merging into associations for the reciprocal and mutual sharing of advantages and disadvantages (armsada) allows the emergence of a new biodiversity [6,10]. banking systems are money producing systems. we do know how they work [11]. can their comparison with living systems permit to understand the origin of the current economic and living crises? and thus suggest adequate solutions [4, 12]? 2 money flows modelling workers exchange their labour against money [7]. this money input allows a person to buy food (to survive that is first to eat) and goods (in order not to be eaten by infectious diseases or to enhance her/his capacity of reproduction): figure 2.the two goals for our society to make money are health and beauty! if money inputs are exceeding their needs working consumers (or consuming workers) may put their money into a bank (or not). money flows can be analysed in terms of interactions, but biologically sounding! in food chains, into an ecosystem, in a commensalism situation only one partner receives benefits. the same goes in a parasitic situation but here the other partner is harmed! only a mutualistic situation allows both partners to have some benefits: table 1. if the bank takes your money and if you can get it back as you want, when you want and freely (without charge), it is a commensalism situation. you and the bank are “eating” the same money at the same table [7]. but, without your money, the bank has no money and cannot make money with your money. what does that mean in terms of advantages and disadvantages? if the bank has enough money -because a lot of people are storing their money in itthe bank can lend you some money. it is an advantage for you. you can buy something you could not otherwise. but there are never advantages without disadvantages. you must pay off your loan with interests. it is a disadvantage for you but an advantage for the bank. all that is an advantage for you is a disadvantage for the bank and reciprocally. how is that situation balanced? if enough deposits the bank has enough money to make loans, to give consumers credits that allow the bank to make money with your money:fig.3. it is a store-take-make situation (fig.2 & 3) [5-7]. to increase cash-flow [13], to have more money to loan, a bank can give you interests for the money you will put in its stock for a while. obviously, the inadvances in systems science and application (2015) vol.15 no.2 151 terest you will earn will be less than the money the bank will store with its loans interests. it is a mutualistic situation. you are at the same table, in the same food chain (the same money chain). you are nourishing the bank and the bank is nourishing you. of course the bank is eating more than you are. of course you can earn money only if other people in the same bank (the same food chain) have not enough money and if you have too much (figure 2). with increasing money stocks, the bank can keep on making more money and is growing. depending on the situation you are, either your money is growing too (if you are loaning money to the bank) or your debt is if you are a borrowing money from the bank. only those associated with the bank growth see their money increase too [13]. whatever its spatial and temporal level of organisation, to live on, any living system,owns 7 invariant qualitative capabilities (top left scheme): c1 mobilisation of matter and energy flows ,c2 growth (=accumulation), c3 reaction to stimulations, c4 organisation into space and through time, c5 integration (from integer: “to make one”) into an ecoexotope (exo: external, tope: spacetime, eco: of inhabitation) of survival, shared with other forms of life, c6 reproduction of its self, and c7 movement. these 7 capacities allow its endophysiotope (endo: internal, tope: space-time, physio: of functioning) to be hosted by an ecoexotope which furnishes the endophysiotope a hosting capacity, but only if the endophysiotope owns an adapted capacity to be hosted in it (integration). a system is always made of 3 kinds of entities: actors, interactions and the whole (top right model). it is always more and less than the sum of its parts. but whatever the system, actors are interacting and each action is a cause of an effect which may also be a cause (feedback) [14]: interaction is construction and construction is interaction (systemic constructal law). a living system is always a system of systems (down left scheme) made by embedments and juxtapositions of previous systems. an endophysiotope at a i level is an ecoexotope of survival for a i-1 level endophysiotope. it is an iteration process (a “matryoshka race”). formed by embedments and juxtapositions of pre-existing systems in a new whole (endophysiotope), a living system is always a part of a food chain -it eats and is eaten, within an ecoexotope of survival shared with other living systems-. growth (quantitative mass increase) is not a goal in itself but it is always a prerequisite phase (larval stage) before the endophysiotope development (acquisition of new qualitative capacities) and its reproduction (number growth). during the mass growth eco-phase, a mass threshold must be passed to reach the minimum volume so as to acquire the adult stage.“the capacity of reproduction has a cost paid by growth.” but a minimum duration (generation time) is required in order to gain it. to survive every living system is a particular place for mobilisation of matter and energy (down right scheme). entering flows (inputs) are used to make products and wastes which 152 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... are then used, stored (throughputs) or excreted (outputs). if the outputs/inputs balance allows an internal accumulation of matter and energy, the system’s mass may grow [1,9]. there are different kinds of description of interactions between living systems. fig. 2 banking structuring and functioning: from synergism to commensalism and mutualism. wether a species/a partner (species a or species b) receives (or not) benefits from another one, we usually distinguish 3 patterns of interactions: commensalism (e.g. “to eat at the same table”, but unequally!), mutualism (e.g. “to share benefits”, though maybe unequally in quality or quantity), parasitism (e.g. “to eat another system” which is harmed). depending on the kinds of interactions (+ benefit, no benefit but harmed effects, 0 no positive or negative effect) we can add up neutralism, competition or injury situations. some people even add up the notions of amensalism or proto-cooperation [15]. these interactions build up a network (which can be graphed) between the different actors of a system of systems (figure 1). yet whatever the interaction (for example here: mutualism or competition, each actor’s action among every couple of actors is both a cause and an effect: systemic constructal law. usually, and particularly in economadvances in systems science and application (2015) vol.15 no.2 153 table 1. ago-antagonistic interactions between 2 locally isolated species. ic models, only a simplified point of view is taken into account with 2 actors (prisoners’ dilemma “game”), which results in 3 situations: “win-win” (the 2 actors cooperate.), “lose-lose” (the 2 actors defect.), “lose-win” or “win-lose” (one cooperates and loses, the other one defects and wins) [8]. of course, due to feed-backs [11, 14], reality is far more complex [16]! 3 the current crisis situation today in france, after the 2009 economic crisis (due to a lack of growth from states and banks), by the law, you must put your money in banks. you can use cash only for small payments. you cannot be paid with cash. your salary must be deposited into a bank. you must use credit cards, checks or money orders. you must pay for all of that! when you trust banks with your money you pay a bank service management to be able to use your money or to give your money back. you will never get back all the money you shall deposit! and when you get your money back you have to pay again some money to the bank. the bank is always making money with every money deposit. you get no advantage in putting your money into the bank. the bank always has advantages when you put your money into the bank. the bank takes a part of your money for you to use your money. and the bank freely uses your money to make loans and money. services banks were freely giving to you a century ago in exchange for them to 154 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... fig. 3 banking systems structuring and functioning as parasites[6-7]. use your money, you must now pay for! it is a parasitic situation. the bank is eating your money! in less than a century, we ran from a synergetic situation to a parasitic one (table i). every day you are gaining and using money, the bank grows and you have less money. the body (money mass) of the bank grows allowing the bank to have more and more money and to create new shelters— in order to store more money—just like parasitic living systems are laying more and more eggs! more money entering the bank, more growth for the bank (figure 2) and increased lifetime for the bank [7]. that is exactly what we see in a ponzi scheme (figure 3), and that is exactly how modern states are growing and extending their area of growth and lifetime. looking back to 2007, the 2009 and 2012 crises were crises for consumers not for banks. for banks, it was an extraordinary growth increase [13]. their cashflow is now bigger than ever— with a 3 fold increase in 5 years. “you cannot eat the pie and have the pie” but banks can! not only banks but also states are growing, adding more and more “layers” to a ponzi pyramid [7-8]. but, soon or late, and faster and faster, a limit comes up! and the limitation for restoring growth is harder and harder to break. crises were exploding when the u.s.a. local growth stopped and the u.s.a. government had to extend its market outside, and worldwide, to a more and more global world market. china did the same but better. the european union did the same but with less efficiency [13]. in a ponzi pyramid the higher you are the more money you get. and the lower advances in systems science and application (2015) vol.15 no.2 155 you are the less you get. it is a pareto situation: 20% people get 80% of money or goods (and power), the remaining 80% get only 20%. you take money from the poorest to make money for the richest. it is typically a prisoners’ dilemma game situation [7-8, 17]. 4 is crisis a problem or a solution? ecosystems are often graphed as a pyramid. and just like in a ponzi pyramid the survival of the highest level depends on the survival of the largest down one [5, 10]. but in a food chain, the living pyramid is built from the bottom up. the ponzi one is built from top to bottom. indeed, food chains are not pyramids but mixed networks: “all eggs are not in the same basket”, “diversity is the rule”, “reciprocal exchanges are the law”. only reciprocal rewards can stabilize cooperation [18-19] and, soon or late, allow the merging of all the actors into an armsada [5-6,9]. that situation of reciprocal and mutual sharing of advantages and disadvantages can be modeled: figure 5. economic systems are neither machineries nor mechanistic linear systems but complex non-linear dynamic systems of systems as are ecosystems. thus, quality (rather than quantity), development (rather than) growth, creation and variety (rather than accumulation) are the keys. sharing limited resources could be equally done but parts amounts are decreasing with number in an hyperbolic way (figure 4) and “you cannot eat the pie and have the pie” [4,8]. the 2009 crisis is the cause (or the effect?) of a double increase of households debts [13]. in the loads of interests, commissions, financial services and administrative costs are voluntarily and artificially excluded, resulting in a false double reduction of these interests. through debt negotiation loans durations could be twice as long resulting in at least a double increase of loans costs, and interests are paid first! more debts are more money [7-8]! recession periods intervals (figure 3) are shorter and shorter: 1993-2002 (9 years), 2003-2009 (6 years), 2009-2012 (3 years), and each year now? that is the same sign of extension as for a pandemic infectious disease like flu! flu intervals were shorter and shorter: 1918-1957 (39 years), 1957-1968 (11 years), 1968-1977 (9 years), before -as in a ponzi pyramida new host was invaded (as a new food chain basis for the pyramid) [7]! and, as in a ponzi pyramid, regrowth resulted in a calm of 20 years (1977-1997). but then, flu crisis intervals bursted again shorter and shorter: 1997-2003 (6 years), 20032005 (2 years); in 2014, 4 different influenza viruses simultaneously were at the origin of epidemics in the popular republic of china. (-from top left to down right-) works people do allow them to gain money they use to buy foods and goods (economic exchange). if some money is not spent it can be stored (accumulation) through bank deposit (or not!). storing 156 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... more money allows banks to have enough money flow control (take) and thus to loan money to people who need it to buy foods or goods. but the bank takes “guarantees” to ensure that the money returns and with interests. money flow is always balanced in favour of the bank (money creation), but without any cost for other people (local commensalism). to have a credit is an advantage for people if they can pay the interests and give the money back. money is neither their property nor that of the bank but the deposits/property of other people (global synergism). to have more money deposits, bigger amounts and for longer durations, banks can give (or not...) money interests paid by the bank to “money-sharers” (“i can put your money to work for you. don’t ask me how. just let me show you”). but the interests paid by banks are always smaller that the interests the banks earn through money loans. there are never advantages without disadvantages. for consumers, disadvantages are greater than advantages. but for banks, advantages are greater than disadvantages. of course banks have taxes and salaries to pay, but they do so with money they make from debts and with money they get from people who do not have debts but “stored” money (pareto situation [16]: figure 4). nowadays money is virtual. only banks know how much money there is! the first upturn was an advantage for asia and latin america, but there was no second upturn. after each break, when the upturns came, only countries from the o.p.e.c. (organization of petroleum exporting countries) still had the same growth they had before. in europa, from the first to the second relative upturn, in 6 years unemployment doubled and no job was created. using a ponzi pyramid process at the world scale, only the u.s.a. stabilized their re-growth, but that was an advantage only for u.s.a. and a disadvantage for all the other countries. only japan enters into a no-inflation new developmental stage [7, 13]. inputs (money deposits) are always greater than outputs (take: figure 2) because the interests paid by banks are always smaller that the interests the banks earn through money loans (accumulation flow). but if you must always pay a fee to put your money into a bank and pay again to get your money back from the bank, you will never have the “availability” of your money— the bank is eating your money! it is a parasitic situation! and when states say “you must put all your money into a bank”, banks can easily make money with your money and take the percentage they want from your money, when they want! it is a take-make situation (figure 1) in favour of banks. for each money flow a part is “rapt” by the bank, banks are growing but not your money! if we model this situation, with t time, with a money amount of a bank, or debt of a consumer (debt makes money), we typically have a ponzi pyramid graph. crisis situations, r (shaded areas), are promoting the slope increase of money accumulation and crises are nearer and nearer [13]. this graph is exactly that of the advances in systems science and application (2015) vol.15 no.2 157 evolution of the us federal debt (as published by the us department of treasury financial management service, from 1965 to 2013, with debt increasing from a factor 1 to 20, in 50 years -from research.stlouisfed.org-). more debts make more money only for banks: “banks can eat the pie and have the pie.” hyperbolic causes and consequences correlations are common laws in living systems (figure 4), as it is for example with the metabolic rate of the nitrogen/phosphorus ratio (n/p). for living systems, development x growth = constant, quality x quantity = constant. what about economic systems? -figure 4just like in economic supply/demand models, hyperbolic graphs can be linearized, by changing the variables or the representation used (logarithmic scale), but the dilemma remains the same... growth or development? quantity or quality? an hyperbolic graph typically is the mark of a pareto phenomenon (20% of actors having 80% of rewards, the remaining 80% getting only 20%), but is this an optimum or an equilibrium [20-22]? in a pareto optimum [16], in the case of the prisoners’ dilemma game [21], gamers/prisoners may cooperate and have equal rewards. in this case any other outcome gives a worse outcome for at least one player. one is rewarded: the bank, the other is threatened: the consumer. and the reward is the biggest banking systems can have. in a nash equilibrium, each player’s strategy is the best response to all other players’ strategies. but it has a cost: rewards or outcomes are lesser and equally shared! that is never the case... except in the benevolent voluntary and united sector of social and solidarity economy [8]! armsadas are emerging from such processes [4-5]: the maximum value of the whole is superior to the sum of the maximum and minimum values for an individual actor [21]. -leftthe winners-losers prisoners’ dilemma game (table 1) fig.4banking/living systems structuring and functioning: limits and limitations. [6,9] 158 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... as a pareto equilibrium example [16]: simplified example of the greek bailout. -rightexample of biological hyperbolic laws xy=k of growth for 2 different legumes species, with y=n/p (where n is a limiting factor for both protein and nucleic acid syntheses, with protein synthesis limiting both mass and number growths, and p a limiting factor for both nucleic acid synthesis and energy storage, with energy storage limiting both growth and development) and x=relative growth rate. for living systems, development x growth = constant, quality x quantity = constant [4-7,9]. the higher the growth x, the smaller the developmental step [4]. -middlethe systemic constructal law: “interaction is construction, construction is interaction”. actors from adjacent levels of organization are interacting, causes getting effects and effects being new causes getting other effects and so on. growth is a prerequisite for development [9]. but just like in a predator-prey relationship, banks or states tell us only what they want us to hear! it is a sort of take-make-waste phenomenon like there is in dying living systems1. the hosting capacity, of any market or ecoexotope, is limited because of a supply and production interaction between the hosting capacity (of an ecoexotope) and the capacity to be hosted (of an endophysiotope). for living systems, growth is durable only if it is sustainable and sustained by each partner, for a benefit for their whole, as is it for an armsada. for a partner to survive all the partners must survive first (and their whole too). that is not the case in hidden banking networks2.the loss of resources among a trophic chain causes big and small animals to be less common and less productive higher up in the food chain. but medium ones are preserved with an increase in biodiversity. in ponzi pyramids the opposite happens. biggest ones will be bigger, poorest ones poorer and medium ones will disappear... the pareto 20%-80% situation! only reciprocal rewards may stabilise a system of systems, particularly in ecoexotopes of hard survival. capitalism must be redefined in terms of an e1if a living system does not stop its functioning because the concentration threshold of a toxic waste is under that of substrates for minimal activity, its growth shall stop quickly. the low hosting capacity of its ecoexotope lowers its endophysiotope growth. but if its endophysiotope has a high capacity to be hosted because of a low threshold of demand, its growth can be durable, even in presence of toxic wastes (figure 5). 2in the u.s.a., in 1950, a unit of tax paid for 1 social security recipient was supported by 16.5 workers but only by 3.0 in 2009, e.g. a load increase of 5.5 fold in 60 years, a duration which equals the time before retirement, and with an increase of payment durations of 10 years (from the age of 68 to 78!). this is a hidden ponzi pyramid sponsored by banks, insurances and states (source:social security administration, the 2010 annual report of the board of trustees, cdc, us life tables. if you are 55 today, current law will pay you 75 cents on the dollar: you paid 100 at 45 and will get 75 at 65 (source:the us social security trust fund, may 2011 report to the congress. payable benefits as percent of scheduled benefits). where are the missing 25? advances in systems science and application (2015) vol.15 no.2 159 cosystem. growth (quantitative increase) is only a prerequisite and a tool for development (qualitative acquisition). development is quality creation, step by step, using growth [1,5-6]. “quantity or growth is the problem” and “quality or development is the solution” [4]. modelling with biological concepts [9,15-16] supports input-output-recycling processes [4, 6,21-23]. -top, from left to righthyperbolic law qq=k [4]: the higher the number of parts q, the smaller the equal amount of each q (not a pareto situation here), just like in an economic situation (with the demand and production relationship); in an ecological situation the bigger the amount of food consumed, the bigger the wastes. but food is limited, under a minimal food threshold or above a maximal supported wastes threshold growth stops. -down, from left to rightecoexotope hosting capacity limitation and endophysiotope capacity to be hosted [6]. if a living system does not stop its functioning due to a concentration threshold of toxic waste (q toxicity) being lower than the concentration threshold of substrates for its minimal activity (q activity), its growth will stop very rapidly and eventually it shall die (down left). the low hosting capacity of its ecoexotope lowers the duration of its endophysiotope growth (figure 1). but if its endophysiotope has a high capacity to be hosted, “thanks for” a low threshold of demand (q activity is low) the system growth can last, even in presence of toxic wastes (q toxicity above q activity) [4]. armsada emergence, as a whole, results (down right) from complementing effects of the different local capacities to be hosted of the various partners (k and k). globally they share the same ecoexotope of survival but using different local parts of its global hosting capacity [6,9]. depending on the local mutual changes of the hosting capacity all that is an advantage for a partner is a disadvantage for another one and reciprocally [5]. 5 discussion-conclusion for banking systems to grow and survive that is ’to eat money and have their money not to be eaten’ just like for living systems it is to eat food and not to be eaten [1]. but with living systems we know that in a predator-prey relationship to survive the predator must eat the matter of preys but not too much [2, 4]. the survival of the predator is limited by the limitations of survival of its preys! isn’t the same ecologic law valuable for economic systems [7]? in a century, banking systems have gone from a mutualistic functioning to a parasitic one [7]. inputs are used only for banks profits, to make products that earn money for people having money. waste products, that is to say all that has a cost for banks, are paid by consumers (figure 3). services are servitudes now! the outputs/inputs balance allows an internal accumulation of money but only for bank shareholders. banking systems are growing and reproducing! and growth 160 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... fig. 5 living systems structuring and functioning: limitations and limits [1,4,67] (quantitative increase) is their goal! is this functioning a cooperative one [15, 22-23] or an extortive one [8]? with living systems we know that growth is only, and always, a prerequisite phase for acquisition of new qualitative capabilities (development) [1, 9]. we also know that growth, development and survival are durable only if they are sustainable for and sustained by all the partners [1, 14]. banking system has to learn that to survive all the consumers must survive first! that is not the case nowadays—if banking systems evolution obeys the same law living systems evolution does, only systems with a high capacity to be hosted -because of a low threshold of demand (no cost of services, no extra fees)will last, particularly in crisis and even in presence of toxic products such as toxic sub-primes. if win-win situations may exist locally (chinese growth was 8.1% in 2011 but 7.5% in 2012 and is decreasing, taiwan growth was only 3.6% in 2013, 2.4% is expected for great-britain in 2014, but uk is playing his local game not the european global one!), with globalization no win-win situation can persist at a world scale. you can never always be a winner, soon or late you will be a loser. only armsadas are lasting! like ecology, economy obeys a cyclic organization of life [16], with growth and differentiation phases, with entrepreneurial and managerial behavior phases that are replacing one another all the time. each phase change is an emergency situation, as is metamorphosis for living systems [9]. macroeconomic disasters fit a power-law model [24] as fit living systems evolutionary changes due to ecological disasters [4,7,10]. you can never have too much money! it is only by coupling insights from ecology and economy [22] that we can begin to model and understand the complex dynamics which underlie the advances in systems science and application (2015) vol.15 no.2 161 generation of poverty [12] and bank profits [8, 13], through growth but not for development [4]! references [1] p.bricage. (2000), “survival of living systems” ,afscet systémiqu & biologie, fac. médecine st pères, paris, france, pp.33. [2] p.bricage. (2008), “cancer is a breaking of the cells armsada through an aggression that results in a lack of non-autonomy”, ues-eus congress, res. systemica, lisboa, portugal, no.7, pp.8, see http://www.afscet.asso.fr/ ressystemica/lisboa08/bricage1.pdf. [3] p.bricage. (2013), “time management by living systems: time modularity, rhythms and conics running calendars. methodology”, theory and applications. systems research and behavioral science, no.30, pp.677-692. [4] p.bricage. (2011), “the social and environmental responsibility of mankind. 1. about man interventions in the living networks”,ues-eus congress, iascys workshop, bruxelles, belgique, pp.25 see http:// www.armsada.eu/files/pbmanserqash.pdf. [5] p.bricage. (2014), “survival management by living systems. a general system theory of the space-time modularity and evolution of living systems: armsada”, world conference on complex systems, pp.16, see https://hal.archives-ouvertes.fr/hal-01065974. [6] p.bricage. (2011), “thinking and teaching systemics: bio-systemics inhigher education”, iascys g.a., chengdu, p.r.china, pp.1-14, see http://tinyurl.com/biosystemics. [7] p.bricage. (2014), “alive and banking systems comparison: prisoners dilemma”, in globalization and crisis. systems complexity and governance, methodologies and tools for governance of world system, no.7, pp.301303. [8] j.-p.delahaye. (2014), “prisoners’ dilemma game and the illusion of extortion.” (the original is in french:“le dilemme du prisonnier et l’illusion de l’extorsion”), pour la science, no.435, pp.78-83. [9] p.bricage. (2005), “metamorphoses of living sstems: associations for the reciprocal and mutual sharing of sdvantages and disadvantages”, ueseus congress, res. systemica, paris,france, no.5, pp.12, see http:// www.afscet.asso.fr/ressystemica/paris05/bricage.pdf . 162 pierre bricage: living and banking systems comparison. prisoners dilemma “win-win” is ... [10] p.bricage. (2010), “cyber-systemics of living systems evolution” ,teilhard aujourdhui, no.33, pp.31-39, see http://hal.archives-ouvertes.fr/ docs/00/42/37/30/pdf/phylotagmotaphologie.pdf. [11] m. mulej,m. rebernik and s. kajzer. (1996), “ systems thinking, entrepreneurship and management”, ifsr newsletter, vol.15(1), pp.2-3. [12] c.n. ngonghala et al. (2014), “poverty, disease, and the ecology of complex systems”, plos biol, vol.12(4), pp.e1001827. [13] m.draghi(dir.). (2014), eurosystem. 2013 annual report of the european central bank , c.e.e., bruxelles, belgique, pp.289. [14] y.lin. (1994), “feedback transformation and its applications”. journal of systems engineering, vol.21(3), pp.32-18. [15] r.axelrod. (2006), the evolution of cooperation(revised edition), perseus books group, new york, usa, pp.241. [16] m.e.j. newman. (2006), “power laws, pareto distributions and zipf’s law”,arxiv: cond-mat/0412004v3, no.2066, pp.28. [17] j. maynard smith. (1976), “evolution and the theory of games”,american scientist, no.61, pp.41-45. [18] s.s. izquierdo. l.r. izquierdo and n.m. gotts (2008), “ reinforcement learning dynamics in social dilemmas”, journal of artificial societies and social simulation, vol.11(2), pp.1, see http://jasss.soc.surrey.ac.uk/ 11/2/1.html. [19] e.t. kiers et al. (2011),“reciprocal rewards stabilize cooperation in the mycorrhizal symbiosis”,science vol.333(6044), pp.880-882. [20] r.l. trivers. (1971), “the evolution of reciprocal altruism”, quarterly review of biology, no.46, pp.35-57. [21] william poundstone. (1993), prisoner’s dilemma, anchor, see http:// en.wikipedia.org/wiki/prisoner’s dilemma. [22] z. wang, a. szolnoki and m. perc. (2013), “optimal interdependence between networks for the evolution of cooperation”, scientific reports, no.3, pp.2470. [23] z. wang,a. szolnoki and m. perc. (2012), “if players are sparse social dilemmas are too: importance of percolation for evolution of cooperation”, scientific reports, no.2, pp.369. advances in systems science and application (2015) vol.15 no.2 163 [24] r.j.barro and t. jin. (2011), “on the size distribution of macroeconomic disasters”, econometrica, vol.79(3), pp.434-455. corresponding author pierre bricage can be contacted at: pierre.bricage@univ-pau.fr adv syst sci appl 2019; 02; 63-79 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/667 study and characterization of thermal comfort in a desert climate oudrane abdellatif1*, aour benaoumeur1,2 1) faculty of science and technology, university el-wancharissi of tissemsilt (cut), road bougra ben hamouda, 38004, tissemsilt, (algeria). e-mail: abdellatif.habadat@gmail.com 2) laboratory of applied biomechanics and biomaterials (labab), polytechnic of oranmaurice audin (enpo-ma), bp 1523 el mnaour, 31000, oran, (algeria). e-mail: ben_aour@yahoo.fr received february 12, 2018; revised december 9, 2018; published december 31, 2018 abstract: the objective major of this numerical study is the characterization of thermal comfort in new habitable architectures located in a completely desert area. this numerical characterization is intended to determine the parameters that affect the thermal comfort for the occupants of these architectures in this region. to achieve this objective, a numerical model describing the thermal exchanges taking place in a model of a habitable envelope has been developed. this model is based on thermal balances established at the level each wall of the habitat. the numerical models developed have been validated using climatic data recently measured in the renewable energy research unit of the saharan medium at adrar ‘’urer'ms‘’. a well detailed analysis of the some parameters that influence the thermal comfort in this architecture was raised and discussed. the fundamental equations governing thermal exchanges have been concretized by an implicit finite difference method, based on the nodal procedure. the system of algebraic equations obtained was solved by the iterative gaussian method. the results of the numerical simulation have shown that the material currently used in the construction of this architecture of adrar region, as well as the current climatological conditions, are the main causes of the thermal discomfort. keywords: finite difference, numerical models, adrar region, characterization, thermal discomfort, heat exchange. 1. introduction environmental issues and energy consumption are increasingly alarming global concerns today. it is essential to adopt solutions to obtain more energy-efficient and sustainable buildings, both new and existing. the use of solar energy is a means of improving the use of natural energy, which can reduce energy consumption [1]. the energy performance of a habitat can be estimated by performing thermodynamic simulations that take into account different conventional assumptions such as: the weather, occupancy, temperature set points, and uses of the window and shutter by the occupants and even the architecture of the habitats. the experience feedbacks in the habitable envelope with high energy performance highlight significant differences in energy consumption between forecasts and summer overheating [2]. ryzhov and al [3] in his recent study, he approached the importance of renewable energy exploitation, more specifically the solar energy in buildings. he attributed the strategy that promotes reduced energy consumption and environmentally friendly techniques. the development of a good energy concept for a house involves the search for a balance between reinforced insulation, compactness, passive solar gains, thermal inertia and comfort, * corresponding author: abdellatif.habadat@gmail.com mailto:ben_aour@yahoo.fr 64 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) indoor air quality and investment cost. this has been the subject of specific attention since the pre-project phase and most often requires the use of premature studies [4]. the essential function of a house is to ensure an interior ambiance well suited to our needs and our comfort. the inhabitant often places his comfort before saving energy. it is therefore necessary to plan the construction and installations to consume less energy while ensuring adequate comfort [5]. saheli and al [6] studied the influence of climatic conditions on the local heating efficiency of a building installed in the south algeria. furthermore, benhammou and al [7,8] analyzed the thermal behavior of a housing submitted to periodic solicitations under hot and dry climate. this analysis was made for the hottest july in the summer season, in order to examine the effect combined of the heat insulation and the passive cooling by eahe on the thermal performances of the livable envelopes for desert regions in the south algeria. this numerical study consists of visualizing not only the impact of climatic conditions and building materials used in adrar region's architecture on thermal comfort, but also its effects on the main heat transfer mechanisms related to the thermal comfort process. in addition, from the results of digital investigation of different variables such as: the internal temperature, external temperature and the density of the solar flux, we show ourselves such an approach which allows improving the thermal comfort in this architecture. 2. domestic architecture the raw earth is a material available everywhere in the planet. it is the universal material the privileged and the easiest handled by the human being to make a shelter while several millennia. globally, earthen houses today house more than a third of the world's population [9]. this construction material, respects the man and the environment, is perfectly recyclable. it ensures thermal comfort by its characteristics and offers advantageous economic convenience. algeria, this vast territory has experienced diversified land architecture according to the diversity of climatic zones. the domestic conception of adrar region, south-west of the sahara, differs from that of algiers, or that of m'zab valley at ghardaïa, at the lower sahara. the model of earth allows a real diversity of architectural language [9]. unfortunately, since the floods of 2009, the use of this material is banned, and replaced by cement (reinforced concrete). this material seen as modernity symbol and social promotion. in figure 2.1, we present the new architecture of adrar region which is based on reinforced concrete without total control of thermal insulation. fig. 2.1. new reinforced concrete architecture in adrar region. study and characterization of thermal comfort in a desert climate 65 copyright ©2019 assa. adv. in systems science and appl. (2019) on the other hand, in figure 2.2 we present the old architecture of this region. this architecture allows a thermal insulation based on the raw earth. the latter played the role of an insulating material against thermal discomfort. fig. 2.2. the old architecture with raw earth in adrar region. 3. climatological characteristics adrar has a typical hot desert climate of the hyper-dry saharan zone. it's the heart of sahara, with a hot, very long summer and a hot short, moderate winter. the annual average of precipitation in this region reaches hardly 14 15 mm, falling essentially in autumn or to spring [10]. the maximal average temperatures are 46 48 °c in july (the hottest month), what makes of adrar one of the world hottest cities. as an example: the peak of record temperature was established on monday, july 9th, 2018 with a temperature of 65 °c [11,12,13]. 3.1. solar data of the typical day chosen in adrar region table 3.1 presents the astronomical data for the designated typical day. noting that these data will be used to calculate the different densities of solar flux influencing heat exchange in habitable envelope. these data were recently measured by a radiometric station to in the renewable energy research unit in the saharan medium of adrar ‘’urer'ms’’. this autonomous radiometric measuring station was realized at the end of 2010. furthermore, it measures the global radiation parameters on three plans (horizontal, inclined at the latitude of the place and optimal monthly) and the ambient temperature. figure 3.1 present a local image of this radiometric station in adrar site. 66 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) fig. 3.1. local picture of adrar radiometric station. table 3.1. astronomical data of the typical day [12]. typical day (july 17, 2014) maximum ambient temperature in (°c) 47,70 minimum ambient temperature in (°c) 32,50 average ambient temperature in (°c) 40,70 maximum flux in (w/m2) 1051,00 average flow in (w/m2) 323,00 average wind speed in (m/s) 5,80 duration of the day in (h) 14,00 the time sunrise in (h) 5,00 the time sunset in (h) 19,00 the sun declination in (°) -13°,12’ the time correction in (minute) 17’,97’’ the hourly distribution of the wind speeds represents an indicator for the wind potential. his knowledge allows the estimation of the wind energy available on the site. the figure 3.2 represents the frequency distribution of average speeds measured for day given in percentage term for adrar site. the analysis of these curves shows that site of adrar has important wind energy potential and which is more favorable to exploitation of this energy type for electricity production. the reached average speed is from 5 to 6m/s in july. study and characterization of thermal comfort in a desert climate 67 copyright ©2019 assa. adv. in systems science and appl. (2019) fig.3.2. hourly evolution of the average wind speed [12]. 4. description of physical model the physical model used in this study is a parallelepiped-type habitable envelope composed of a single four-facade room, see figure 4.1. the construction is located on a surface of 20m2. this room is built according to the following standards:  the walls of the room are constructed of a light structure usually, 15 cm of block full of cement.  the floor slab is placed on a flat earth. it is cast directly onto a 4cm layer of thermal insulation.  the roof consists of a slab of reinforced concrete. fig.4.1. sketch of the piece studied. 5. implementation of model and resolution in this numerical study, we have opted for the nodal method use, which makes it possible to establish a thermal network equivalent to the modeling element. conductive, convective and radiative processes are considered in global form (conservation of the heat flux). according to on the complexity of the network adopted, one can to incorporate all or part of the heat transfer phenomena, which allows predicting, with more precision, the thermal behavior of 68 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) the studied part [14]. the phenomenological equation of heat relative to an infinitesimal volume makes it possible to write [15,16,17]: p.dτ).ddiv(dτ t t ρ.cp     (5.1) with; t t   : is the rate of variation of stored energy in dτ ; )div( : is the outgoing flow in dτ ; p.dτ : is the energy flow generated. by integrating this expression on a volume vi at temperature ti, the energy balance is written [15-18,19,20]: iradianceconvectionairconduction i i pφφφφ t t .ρ.cp    (5.2) or; cpi : is the calorific capacity of the volume vi pi : is the source of flux dissipated in the volume vi conductionφ , airφ , convectionφ and conductionφ : are the flows exchanged between vi and its environment by conduction, fluidic transport, convection and radiation. figure 5.1 shows the different modes of heat transfer in the study room. fig.5.1. description of the different modes of heat transfer for the studied room. the analog schematization, we allow to propose an electric circuit amounting to the thermal system to be able to apply the law of ohm [21,22,23,24]. study and characterization of thermal comfort in a desert climate 69 copyright ©2019 assa. adv. in systems science and appl. (2019) fig.5.2. electrical map of the studied room. and: extt : external ambient temperature in (° c); skyt : temperature of the celestial vault in (° c); soilt : soil temperature in (° c). as a condition the limits of the room studied, consideration was given to the ground temperature equal to the external ambient temperature. in addition, the temperature of the celestial vault is calculated by the following formula [25,26]: 5,10552,0 ambsky tt  (5.3) indeed, the coefficient of external convection is given by [25,26]: speed wind8,37,5 vh extc  (5.4) the law enforcement of ohm, in every knot of the studied room, gives the system of following equations [17-20]:  thermal balances of external facades of the studied room                                            absorpeesolpeevc b ext pee pe b pe absorpfpevc b extc pfpe pfp bpfp absorpoesolpoevc b extc poe po bpo absorpnevc b exctc pne pn bps absorpsesolrpsevc b extc pse ps b ps pt.hrt.hrt. e λ t.hr δt dt . s .cpm pt.hrt. e λ t.h δt dt . s .cpm pt.hrt.hrt. e λ t.h δt dt . s .cpm pt.hrt.hrt. e λ t.h δt t . s .cpm pt.ht.hrt. e λ t.h δt dt . s .cpm pnesol (5.5) 70 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019)  thermal balance of interior facades of the studied room                                                                               t..sht..sht..sh t..sht..sht..sh δt dt ..cpm t.ht.ht.h t.ht.ht. e λ t.h δt dt . s .cpm t.ht.ht.h t.ht.h.δδ e λ t.h δt dt . s p .cm t.ht.ht.h t.ht.ht. e λ t.h δt dt . s .cpm t.ht.ht.h t.ht.ht. e λ t.h δt dt . s .cpm t.ht.ht.h t.ht.ht. e λ t.h δt dt . s .cpm pipc6pfp5cpe4c po3cpn2cps1c int a a pippeirpfpipeirpoipeir pnipeirpsipeir b intc pei pe bpe pippfpirpeipfpirpoipfpir pnipfpirpsipfpir b intc pfpi pfp bpfp pippoirpfpipoirpeipoir pnipoirpsipoir b intc poi po bpo pippnirpfpipnirpoipnir peipnirpsipnir b intc pni pn b pn pippsirpfpipsirpeipsir poipsirpnipsir b intc psi ps bps (5.6) the algebraic equations systems resulting from the discretization of transfer equations in both habitat environments (external and internal) are expressed in the form of matrices that can be written as: a . x = b. this equations system is solved by using the iterative numerical method of gauss.                                                                                                                    11 10 9 8 7 6 5 4 3 2 1 211111110118116114112 10111010109108106104102 91099 8118108887868482 7877 6116106866656462 5655 4114104846444342 3433 2112102826242221 1211 00000 0000 000000000 0000 000000000 0000 000000000 0000 000000000 0000 000000000 b b b b b b b b b b b t t t t t t t t t t t aaaaaa aaaaaaa aa aaaaaaa aa aaaaaaa aa aaaaaaa aa aaaaaaa aa tt a tt pei tt pee tt pfpi tt pfpe tt poi tt poe tt pni tt pne tt psi tt pse (5.7) or:            psersolpservc b cext ps pspstt pse hh e h ts cpm ta 1 . . 11 (5.8) study and characterization of thermal comfort in a desert climate 71 copyright ©2019 assa. adv. in systems science and appl. (2019)         e ta btt psi  12 (5.9)           ps t solpsersol t vcpservc t extcext t pse ps bps fsgththth t t s cpm b .... . 1 (5.10)            piprpsipfpirpsipeirpsipoirpsipnirpsi b c ps bpstt psi hhhhh e h ts cpm ta  int22 1 . . (5.11)  pnirpsi tt pni hta   24 (5.12) the temperature of the thermal comfort (tc) for a habitable one makes it possible to estimate the heating needs [27]. however, based on the bibliographic synthesis that was conducted in order to find a suitable model for the calculation of the comfort temperature, we opted for a relationship that was established between the comfort temperature (tc), internal temperature (tint) and mean radiant wall temperature (tpi), as follows [28,29]: 2 tt t 6n 1n piint c     (5.13) ppipfpipoipeipnipsi 6n 1n pi ttttttt    (5.14) taking into account the air as a transmission fluid between the ambiance and the room facades. table 2 shows the physico-thermal properties of the air that were taken during the numerical simulation. table 5.1. the air thermal properties [29]. λ thermal conductivity in (w.m-1.k-1) 0,0262 cp specific heat in (j.kg-1.k-1) 1006 ρ density in (kg.m-3) 1,177 μ dynamic viscosity in (kg.m-1.s-1) 1,85.10-5 υ kinematic viscosity in (m2.s-1) 1,75.10-5 pr number of prandtl 0,708 6. flow chart of numerical modeling the calculation of the different heat transfer parameters within the habitable envelope under consideration is based on the following steps:  introduction of climatic data of the considered region for the selected typical day.  calculation of the geometrical and astronomical parameters of the sun in the considered typical day.  introduction of climate data such as: the average amount of the ambient solar flux and the ambient temperature of the considered day.  determination of solar radiation incidence angles for each facade of the habitat as a function of the hour angle and the position of the sun in the sky. 72 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019)  computation of the amount of solar flux density for of all the house facades. the numerical tests carried out during the development of the computer code in fortran language led to the retention of a time step of 300 seconds. in figure 6.1, we show the computational steps developed in the numerical simulation code. fig.6.1. flow chart of calculation steps in the habitat performed by the developed fortran code. 7. results and discussion 7.1. internal evolution of temperature of the room figure 7.1 shows the hourly variation of the temperature of different internal facades of the room during the typical day. it can be seen that the thermal inertia of the concrete plays a very important role in the heat transfer at the walls. indeed, the ceiling temperature (tpfp) is greater compared to the temperatures of the other internal faces of the south walls (tpsi), west (tpoi), north (tpni) and east (tpei). this is due to the conjugation of two essential factors that led to this increase in temperature: in the first place, the walls thickness of the room which involves the elevation of the thermal inertia and secondly, the inclination angle of the roof facade which is equal to 0 °. study and characterization of thermal comfort in a desert climate 73 copyright ©2019 assa. adv. in systems science and appl. (2019) fig.7.1. internal evolution of air temperature in the room, according to the local time. 7.2. internal evolution of temperature according to building materials to use construction materials wisely, it is essential to know their thermal properties. the thermo-physical properties of the materials used in this study are shown in table 7.1. the thermal capacity of a wall is especially useful if it is placed in the interior of the room and isolated from the external climatic conditions. to build in strong inertia, it is therefore to use heavy materials inside the room in order to store the solar heat and to attenuate the internal temperature variations [30]. conversely, the temperature of a room with low inertia increases rapidly at the slightest ray of sun without the possibility of storing solar heat. internal temperature differences will be significant and the risk of overheating will be high. strong inertia is especially useful in case of permanent occupation. low inertia can be interesting for habitable environments with intermittent use [30,31]. table.7.1. thermophysical properties of building materials [30,31]. heavy concrete lightweight concrete concrete stone light wood heavy wood thermal conductivity  (w.m-1. k-1) 1,75 1,00 1,40 0,14 0,20 density  (kg.m-3) 2200 1500 1895 540 800 specific heat cp (j.kg-1. k-1) 1000 1000 1000 2400 2700 emissivity  0,54 0,54 0,54 0,86 0,86 absorption coefficient  0,6 0,6 0,6 0,07 0,07 figure 7.2 illustrates the thermal capacity effect of solar energy storage in the room wall of for different building materials (light concrete, heavy concrete, stone concrete, light wood and heavy wood). it can be noted that for all habitable rooms built with high thermal capacity materials such as lightweight concrete (see fig. 7.2(a)), heavy concrete (see fig. 7.2(b)) and concrete stone (see fig 7.2(c)), the temperature of its internal space is optimal in the month of july because of the thermal inertia of these materials. indeed, it is well observed in figures (see fig. 7.2(a), fig.7.2(b) and fig.7.2(c)) that the temperature of the internal ambiance varies between 37°c and 38,5°c. in addition, if the construction of habitable rooms is based on low heat capacity materials, the temperature of the internal space varies around 34°c and 35°c because of the thermal inertia of this building materials type (heavy wood (see fig. 7.2(d)) and light wood (see fig. 7.2(e)). furthermore, it is noted that the temperature of the internal facades for building 74 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) materials with high thermal capacity is very high compared to the internal facades of buildings with low heat capacity materials. study and characterization of thermal comfort in a desert climate 75 copyright ©2019 assa. adv. in systems science and appl. (2019) fig.7.2. evolution of temperature of the room internal facades for different building materials. 7.3. evolution of the thermal comfort temperature figure 7.3 shows the hourly change in temperature of thermal comfort of different building materials as a function of local time. from this illustration, it can be seen that the choice of construction material has a significant influence on the hourly evolution of the thermal comfort temperature. indeed, when one uses for example the concretes selected categories in this study as building material with these dry climatological conditions, and without control of thermal insulation, the comfort temperature will exceed the conventional norms. that is to say, undesirable overheating of the habitable 76 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) environment with a temperature of 39°c at 15h00 in the afternoon as illustrated in figure 8 for the three categories of concrete (heavy, light and stone). on the other hand, the exploitation of other building materials such as the categories of wood (heavy and light) in this region with the same climatic conditions will help to get closer to the conventional values of the comfortable temperature in the environment. habitable, since the maximum temperature obtained using the wood does not exceed 37°c (value obtained at 15h00 pm). fig.7.3. evolution of the thermal comfort temperature in the room for different building materials in a typical day. 7.4. overlay on the psychrometric chart figure 7.4 shows the superimposition on the psychometric chart of temperature of thermal comfort within the room for different building materials in the typical day (july 17th, 2014). it can be seen that the thermal comfort temperature of concrete structures (heavy, light and stone) is positioned in the area of thermal discomfort and the need for evaporation. while for both types of wood (heavy and light) is positioned in the same area as that of the concrete categories, except that the comfort temperature for a concrete construction varies between 38°c and 39°c and the temperature of a wood construction varies between 34°c and 36°c. noting that these thermal comfort temperatures were calculated for the typical day of july 17th, 2014, which is the hottest day of the summer season in adrar region. fig.7.4. super-positioning temperatures of thermal comfort of the room for different building materials during the typical day. study and characterization of thermal comfort in a desert climate 77 copyright ©2019 assa. adv. in systems science and appl. (2019) 8. conclusion thermal comfort is an essential element for the well-being of the occupant in his built environment. taking it into account that the habitat implies there different aspects. however, after the numerical examination of the temperatures of different internal facades of the habitable room, the change effect of the building materials on the evolution of the internal temperature and the superposition of the thermal comfort temperatures on the psychometric chart, it has been found that:  the use of bare concrete as a building material in the current architecture of adrar region, involves undesirable overheating. in addition, the same effect has been noted for wood in this region, but with temperatures a little closer to conventional temperatures of thermal comfort.  the severe climatic conditions of this region contribute in a direct way to the thermal discomfort. on the basis of this study, it is recommended that before beginning the construction of a habitable envelope in this desert region, it is imperative to undertake an overall revision of conventional standards for thermal comfort, in particular the respect for bioclimatic concepts. on the other hand, we suggest:  ventilation and evaporative cooling during the summer season.  insulation and shading of the walls the most requested for excess solar radiation. nomenclatures abbreviation e the material thickness m pe east wall hc coefficient of internal heat exchange by convection -1-2 .kw.m pn north wall mps the south wall mass kg po west wall pabsor amount of heat absorbed w ps south wall spl south wall surface m2 pfp the ceiling wall tamb ambient temperature °c pni internal north wall tsol temperature of soil °c psi internal south wall tair air temperature °c pfpi internal wall of the ceiling t temperature difference °c pse external south wall tint full air temperature in the habitat °c poe external west wall text outside air temperature °c pip internal floor wall tb temperature of the concrete slab °c pee externe east wall tc thermal comfort temperature °c pfpe external wall of the ceiling fsg global solar flux w/m2 m'zab region north of sahara algerian references [1] dengjia w., yanfeng l., jing j., and jiaping l., (2015). the optimized matching of passive solar energy supply and classroom thermal demand of rural primary and secondary school in northwest china, procedia engineering vol. 121, 1089 – 1095. [2] batier c., (2016). confort thermique et énergie dans l’habitat social en milieu méditerranéen d’un modèle comportemental de l’occupant vers des stratégies architecturales, ph.d. thesis, université of montpellier, la france. 78 a. oudrane, b. aour copyright ©2019 assa adv. in systems science and appl. (2019) [3] ryzhov a., ouerdane h., gryazina e., bischi a. and turitsyn k., (2019). model predictive control of indoor microclimate: existing building stock comfort improvement, energy conversion and management vol.179, 219–228. [4] alexis s., (2014). etudes de systèmes énergétiques et optimisation énergétique de bâtiments basse énergie, master thesis, université of lorraine, la france. [5] errebai f., derradji l., maoudj y. & amara m., (2012). confort thermique d’un local d’habitation : simulation thermoaéraulique pour différents systèmes de chauffage, revue des energies renouvelables vol. 15 n°1, 91–102. [6] sehli a., hasni a. & m.tamali, (2012). the potential of earth-air heat exchangers for low energy cooling of buildings in south algeria, energy procedia vol.18, 496–506. [7] benhammou m., draoui b. & hamouda m., (2017). improvement of the summer cooling induced by an earth-to-air heat exchanger integrated in a residential building under hot and arid climate, applied energy vol. 208, 428–445. [8] benhammou m., draoui b., zerrouki m. & marif y., (2015). performance analysis of an earth-to-air heat exchanger assisted by a wind tower for passive cooling of buildings in arid and hot climate, energy conversion and management vol. 91, 1–11. [9] boutabba h., mili m., boutabba s., (2016). l’architecture domestique en terre entre préservation et modernité: cas d'une ville oasienne d'algérie (aoulef) domestic architecture in the ground between conservation and modernity: if an oasis city in algeria, j. mater. environ., sci. n°7 (10), 3558-3570. [10] chellali f., khellaf a., belouchrani a. & recioui a., (2011). a contribution in the actualization of wind map of algeria, renewable and sustainable energy reviews vol. 15, 993–1002. [11] oudrane a., (2011). modélisation et caractérisation d’un système solaire destiné pour la production du chauffage des locaux, magister thesis, l’école polytechniques of oran maurice audin, algeria. [12] radiometric station, (enermena), technical report, (2014). high precision meteorological station of research unit for renewable energies in the saharan environment, in adrar, algeria. [13] peel m., finlayson b. & macmahon t. (2007) updated world map of the köppengeiger climate classification, in hydrology and earth system sciences no 11, p. 1633, copernicus publications pour european geosciences union, göttingen, issn 10275606. [14] smaïl m., (2004). modélisation électromagnétique et thermique des moteurs à induction, en tenant compte des harmoniques d’espace, ph.d. thesis, université of lorraine, nancy, la france. [15] lutun j., (2012). modélisation thermique des alternateurs automobiles, ph.d. thesis, université of grenoble, la france. [16] jabbar a., khalifa n. & abbas e., (2009). a comparative performance study of some thermal storage materials used for solar space heating, energy and buildings vol. 41, pp. 407–415. [17] hernández-pérez i., álvarez g., gilbert h., xamán j., chávez y. and shah b., (2014). thermal performance of a concrete cool roof under different climatic conditions of mexico, energy procedia vol. 57, 1753 – 1762 [18] bekkouche s. & benouaz t., (2014). thermal resistances of local building materials and their effect upon the interior temperatures case of a building located in ghardaïa region , construction and building materials vol. 52, 59–70. [19] mezrhab a. & bouzidi m., (2006). computation of thermal comfort inside a passenger car compartment, applied thermal engineering vol. 26, 1697–1704. [20] boukli m. a., amara s., chabane n. e., (2011). thermal requirements and temperatures evolution in an ecological house, energy procedia vol. 6, 110–121. https://fr.wikipedia.org/wiki/european_geosciences_union study and characterization of thermal comfort in a desert climate 79 copyright ©2019 assa. adv. in systems science and appl. (2019) [21] peuser f., remmers & schnauss k., (2005). installation solaires thermiques, conception et mise en œuvre, edité par systèmes solaires, solar praxis et le moniteur. [22] bekkouche m. e. a., (2009). modélisation du comportement thermique de quelques dispositifs solaires, ph.d. thesis, université of tlemcen, algeria. [23] amar b., (2010). contribution à l'étude de séchage solaire de produits agricoles locaux, magister thesis, université of mentouri – constantine, algeria. [24] virginia g., valentina m., phillip b., clifford a., (2016). inferring the thermal resistance and effective thermal mass distribution of a wall from in situ measurements to characterise heat transfer at both the interior and exterior surfaces, energy and buildings, doi.org/10.1016/j.enbuild.2016.10.043. [25] oudjedi s., boubghal a., braham chaouch w., chergui t. et belhamri a., (2008). etude paramétrique d’un capteur solaire plan à air destiné au séchage (partie: 2), revue des energies renouvelables smsts’08 alger, 255 – 266. [26] basile k., cossi n., (1999). evaluation of the amount of the atmospheric humidity condensed naturally, renewable energy vol. 18, 223-247. [27] chandel s., aggarwal r., (2012). thermal comfort temperature standards for cold regions, ashdin publishing innovative energy policies, vol. 2, article id e110201, pp. 5. [28] humphreys m. a., (1978). outdoor temperatures and comfort indoors, batiment international, building research and practice vol. 6, pp. 92. [29] sam f., (2012). réhabilitation thermique d’un local dans une zone aride-cas de ghardaïa, magister thesis, université mouloud mammeri of tizi ouzou, algeria. [30] tixier n., (2012). environnement thermique et maîtrise énergétique, ensag. [31] mavromatidis l. e., m. el mankibi, (2012). numerical estimation of time lags and decrement factors for wall complexes including multilayer thermal insulation, in two different climatic zones, applied energy vol. 92, 480–491. adv syst sci appl 2019; 02; 120-133 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/593 optimal rotor design of claw-pole alternator for performance improving n.brahimi*1, s.tahi2, e. boudissa1, m.bounekhla1 1) laboratoire de système electrique et télécommande, université blida 1, blida, algeria. e-mail: nouzha_m74@yahoo.fr 2) electrical and industrial systems laboratory, usthb/fei, algiers, algeria. e-mail: stahi@usthb.dz received may 17, 2018; revised june 21, 2019; published july 10, 2019 abstract: the demand in electric energy, in vehicles, has experienced a remarkably rapid growth for a long time and this trend is expected to remain during the incoming years. the claw pole alternator is designed to provide the necessary power for all the electrical accessory loads over the entire operating speed, ranging from an idle speed of 1800rpm to a top speed of 6000 rpm. because of the relationship between the speed and the output power, producing enough power at idle speed becomes a challenging problem. the aim of this paper is to optimize the rotor design of conventional claw pole alternator in order to improve its performance during the idle speed. to derive the optimal rotor structure, cyclic coordinate method linked with magnetic equivalent circuit (mec) model is used, including magnetic saturation effect and leakage flux. an increase of up to 70% in output power at idle speed is demonstrated and significant improvements in performance over the whole speed range were observed. keywords: claw pole alternator, reluctance network, mec model, electric generator, shape optimization, cyclic coordinate method. 1. introduction automotive alternators are responsible for the generation of electricity in almost all vehicles using an internal combustion engine. they must be capable of powering all the relevant systems the vehicle requires during operation. until now, the claw pole alternator constitutes the unique device responsible for the generation of electricity within a vehicle [4]. the main advantages of the claw pole alternator are its low cost and construction simplicity. therefore, it is well adapted for industrial and mass production. however it suffers from low efficiency caused by high leakage flux [19] and poor output power [22]. performance and efficiency vary widely depending on the alternator speed and load conditions [12]. commonly, during idle speed (1800 rpm), the claw pole alternator isn’t able to sustain active electrical loads for a reasonable amount of time without undue drain on the battery. therefore, obtaining sufficient output power at idle speed is required. recently there has been a considerable amount of work into developing higher power alternators with improved performance. some of this work has focused on designing a new model of alternators such as the new claw pole alternator where the dc-excited winding is located in the stator [10]. another example is the doubly excited brushless claw pole alternator, where the hybrid excitation is achieved by ferrite permanent magnets suitably mounted on both the rotor and the stator in one hand, and a field winding in the stator on the other hand [5, 18]. most of these new designs of alternators need a lot more development before they become a * corresponding author: nouzha_m74@yahoo.fr mailto:stahi@usthb.dz optimal rotor design of claw-pole alternator 121 mature technology with possible integration in the vehicle power generating system. improved performances are also obtained by, simply, adding surface or interior permanent magnet to the conventional claw pole alternator. several studies have been carried out on hybrid excitation claw pole alternators [13, 24, 25]. in hybrid excitation, at least two magnetization sources are present in the rotor. in addition to the dc excited wound field, permanent magnets are placed into the rotor to compensate leakage flux and to boost rotor flux. permanent magnet increases the power density and efficiency [21], however their higher price undermines their usage [16]. furthermore, current regulations and industry marketing policies push the development process cycle energetically less expensive and environment friendly. this tight context implies finding new technical solutions to keep the same efficiency without permanent magnets. for these reasons a lot of efforts are being spent to improve performance and efficiency by optimizing the design of an existing traditional (conventional) claw-pole alternator [8, 23]. this paper describes the rotor design optimization of an existing conventional claw pole alternator while keeping same stator design and the same alternator footprint, in an attempt to increase the output power under different load conditions at the idle speed. the design improvement is achieved by using the magnetic equivalent circuit (mec), which is widely considered in literature [2, 7, 9]. it has been shown in [15] that mec model helps to predict performance of claw pole alternator with high accuracy and minimum cpu computation-time. optimal rotor is designed by using cyclic coordinate method. the optimization results which provide the optimized rotor geometrical dimensions are discussed. the performance and the output power of the optimal claw pole alternator over the entire speed range from idle speed of 1800rpm to top speed of 6000 rpm are presented. 2. 3d nonlinear magnetic equivalent circuit model claw-pole alternator mathematical modeling is complicated because the paths of the magnetic flux are 3d. therefore, the analytical method of magnetic equivalent circuit (mec) based on reluctance networks theory is used [6, 7]. the claw pole mec model employed in this work has been developed by rakotovao [17] and used for optimization by albert et al [2]. this mec model is less time consuming compared with numerical model based on finite element theory [9]. thus, the mec model is particularly suitable for optimization purpose [15]. in the mec method, each flux path section is characterized by magnetic reluctance. the good estimate of reluctances is essential because it guarantees the precision of the model and its coherence during optimization. accuracy of the mec model also depends on the topology of the network [14]. 2.1 expression of reluctances when material saturation is taken into account, the reluctance is non-linear function of the flux, as shown in equation (2.1) [9]: 𝑅 = 𝐿 ф 𝐻 ( ф 𝑆 ) (2.1) where, 𝐿 is the average flux path length, 𝑆 is the cross-section area and ф is the magnetic flux. h(b) is the analytical expression of the magnetic field versus the flux density 𝐵, which is expressed as follows [9]: 𝐵 = ф 𝑆 (2.2) 122 n.brahimi, s.tahi, e. boudissa, m.bounekhla 2.2 resolution method of the mec referring to the literature [9, 11], it has been reported that numerical resolution procedure of the claw pole mec is based on newton raphson algorithm and reluctance network theory. 2.2.1 mec resolution in no load condition mec resolution enables the evaluation of the fluxes circulating in the different parts of the magnetic circuit. in the case of no load condition, alternator’s back electromotive force (backemf) is a function of the direct flux component only. the computation of the no-load backemf is expressed as follows [9]: 𝐸 = 1 √2 𝑁𝑎 𝜔 ф𝑚 (2.3) where, 𝑁𝑎 is the number of turns per armature phase, 𝜔 is the angular frequency and ф𝑚 is the maximum flux crossing a phase which is evaluated using the mec resolution. no-load back-emf is plotted in fig. 2.1, where it can be noticed that the magnetic saturation occurs when the field current is greater than 1a. fig. 2.1. no-load back-emf emf versus field current 2.2.2 mec resolution under load condition claw pole alternator is a salient-pole synchronous generator with high saturated areas. therefore, including the effect of the armature magnetic reaction in both q-and d-axis is required. within the q-axis, magnetic saturation is low, resulting in a constant stator transverse inductance, while in the d-axis, saturation effect is important [2, 4]. then, mec accounting for the d-axis armature reaction is consider in order to carry out the flux linkage, especially, the one which allows the determination of the d-axis component of the back-emf ( 𝐸 𝑑 ), given by equation (2.3). under load condition, the back-emf phase is expressed in terms of its d and q components, as follows [9, 11]: 𝐸 = 𝐸 𝑑 + 𝐸 𝑞 (2.4) with { ‖𝐸 𝑑‖ = 𝐸𝑑(𝐼𝑓 , 𝐼𝑑) ‖𝐸 𝑞‖ = 𝑋𝑞. 𝐼𝑞 (2.5) where 𝐼𝑑 and 𝐼𝑞 are the d and q components of the armature current. 𝐼𝑓 is the field current. 𝑋𝑞 is the q-axis reactance. moreover, for resistive load and adopting blondel model illustrated in fig.2.2, the backemf: 𝐸 is given by equation (2.6) [2, 4]: 𝐸 = 𝑉 + 𝑅𝑠𝐼 + 𝑗𝑋𝜎𝐼 (2.6) optimal rotor design of claw-pole alternator 123 where: 𝑅𝑠 and 𝑋𝜎 are the phase resistance and leakage reactance respectively. rl is the resistive load. 𝑉 and 𝐼 are voltage phasors and the armature current respectively. (a) (b) fig.2.2. (a) blondel phase diagram in the case of resistive load. (b) claw pole alternator phase equivalent circuit. procedure for an operating point calculation is illustrated in fig 2.3. at each step, the change in the solution is monitored. the process continues either until the change is less than the newton raphson tolerance of 0.1%. the solution is obtained, in several seconds, after maximum ten iterations. fig.2.3. operating point calculation algorithm. 3. alternator performance in the following section the alternator performances are evaluated across the whole speed range. the 1800 rpm speed corresponds to idle speed and 6000 rpm corresponds to cruising speed. figure 3.1 (a) shows the calculated alternator output voltage versus the field current, parameterized by the load current. the alternator generated voltage increases along with 124 n.brahimi, s.tahi, e. boudissa, m.bounekhla alternator speed and the field current. at idle speed, the maximum current drawn from the alternator is 20.62 a, with field current and output voltage are equal to 5a and 12vac respectively. at this operating point, the saturation is well considered. it can be observed that when the load current increases the output voltage drops. 3.1 armature current characteristic the calculated armature current versus alternator speed is illustrated in fig.3.1 (b). the field current is set to 5 a. the armature current curve is characterized by three operating point. the first one is the generation starting speed (pgs) or 0-ampere-speed at which the alternator reaches its rated voltage without delivering power. the second one is the maximum output current at cruising speed (pc), which corresponds approximately to the shortcircuit current of the alternator. the third one is the output current at idle speed (pi)[12]. (a) (b) (c) fig. 3.1. (a) stator phase voltage versus field current. (b) armature current versus speed and the three operating points. (c) output power versus speed at fixed output voltage 3.2 alternator output power at fixed output voltage: particular attention is paid to the output power requirement at idle speed, at which, the alternator must deliver at least the power needed for long-term consumers. no output power is required below the idle speed. the calculated output power curve versus alternator speed is shown in fig.3.1. (c). 3.3 maximum output power of the alternator. the calculated maximum output power of the alternator is given in fig. 3.2. the field current is set to 5a. figure.3.2 (a) shows the load current that corresponds to the maximum power delivered from the alternator. figure.3.2 (b) shows the output voltage that corresponds to the maximum power. these curves show the amount of maximum available output power of the alternator. optimal rotor design of claw-pole alternator 125 (a) (b) fig.3.2.calculated alternator output power parameterized by the alternator speed, versus (a) armature current, (b) stator phase voltage. 4. optimization design procedure the aim is to optimize rotor structure of claw pole alternator at idle speed, including the magnetic saturation effect, while maintaining the same initial alternator footprint. the stator design parameters remain unchanged. table 4.1 reports the alternator parameters, with assigned constant values that do not change during the optimization process. two optimization procedures are described, namely: first optimization procedure in the first optimization procedure, the objective function is to maximize the output power at idle speed while maintaining the output voltage at 12vac. in this case, maximizing the output power leads to maximizing the delivered armature current. second optimization procedure in the second optimization procedure, the objective function is to maximize the output power without imposing the output voltage or armature current. table 4.1. main fixed parameters of the cpa. fixed parameter value number of phases 3 number of pairs of poles 6 number of stator slots 36 outer stator diameter (mm) 127 inner stator diameter (mm) 88.6 shaft diameter (mm) 18.1 height of plateau (mm) 11.95 length of rotor (mm): 52 4.1 rotor parameters optimization  first optimization procedure five crucial geometrical optimization parameters are selected, as shown in table 4.2 and illustrated in fig 4.1. some of these geometrical parameters determine the shape of the claw pole. determining the shape of the claw pole is the most important obstacle in the claw pole alternator designing process. the claw pole has an important role in closing the magnetic field lines because it ensures the route from the rotor excitation towards the air gap. the parameter, related to rotor core (core length) is, also, very important, especially as it conditions the space booked for the field winding.  second optimization procedure in this case, besides the five geometrical optimization parameters, the armature current comes to be added as the sixth optimization parameter. 126 n.brahimi, s.tahi, e. boudissa, m.bounekhla table 4.2. design of the sixth optimization parameters of the claw pole rotor optimization variable optimization parameter base value (mm) geometrical optimization parameters x1 width of the base of the claw 24.7 x2 height of claw tip 4.7 x3 core length 29 x4 air gap length 0.65 x5 width of claw tip 6 — x6 armature current — in addition to these optimization parameters, there are other five geometrical implicit parameters, listed in table 4.3 and illustrated in fig 4.1. to ensure the geometrical coherence of the rotor, these parameters are subject to change according to optimization parameters evolution during the optimization process. fig.4.1. geometrical parameters of rotor pole structure table 4.3.implicit optimization parameters of the claw pole rotor implicit variable implicit parameter base value (mm) x7 distance between claw 6.6 x8 claw pole length 29 x9 claw side plate thickness 11.5 x10 core radius 41.4 x11 outer rotor radius 87.3 the optimization of the rotor claw pole is a multivariable non-linear problem. in this investigation, the objective functions cannot be expressed as closed forms thus a derivative free optimization method is used for the objective function evaluation. therefore, cyclic coordinate method is applied. 4.2 cyclic coordinate method a cyclic coordinate method [3, 20, 26] using the unknown parameters as the search directions, is applied to maximize the objective function successively along each coordinate. more specifically, the method searches along the directions , where is a vector of zeros except for a one at the j th position. thus, along the search direction , the design variable is changed while all other variables are kept fixed. this technique is described in fig. 4.2 for the first iteration and in the case of two coordinate axes x1 and x2. starting from p0 (x10, x20), the maximization is performed successively along x1 and x2 and leads, respectively, to p1 (x11, x20) and p2 (x11, x21). the iterative process is repeated until the error test is satisfied. the algorithm shown in fig. 4.3, ndd ,.......,1 jd jd jx optimal rotor design of claw-pole alternator 127 estimates the vector parameters p that maximizes the objective function which increases at each iteration. fig.4.2. 2-d illustration of the cyclic algorithm fig.4.3. flowchart of the cyclic coordinate algorithm in this study, the alternator output power is chosen as objective function, which is expressed as follows [22]: 𝑃 = 3 ∗ 𝑉0 ∗ 𝐼𝑠 (4.1) where, 𝑃 is the output power, 𝑉0 and 𝐼𝑠 are the rms value of output voltage alternator and armature current respectively. both speed and field current are kept constant and equal to 1800 rpm and 5a respectively. orientation of the optimal solution toward a technically correct solution requires us to limit the search space. therefore, the optimization variables are constrained to vary in a range, between a lower and an upper limit, as shown in table4.4. the armature current is introduced, only in the second optimization procedure, as a sixth variable. table 4.4. optimization variables (mm or a). optimization variable lower limit upper limit x1 0.3 0.65 x2 1 5.5 x3 25 29 x4 6 38.02 x5 3 10 x6 10 80 128 n.brahimi, s.tahi, e. boudissa, m.bounekhla 5. optimization results and commentaries as expected, using mec coupled with cyclic coordinate method yields results, a few iterations (about 1h30) are required to achieve the maximum of the objective function. cyclic coordinate is a deterministic method, thus it is appropriate to secure the obtained results. this is done by using different sets of initial points and to check if the same maximum output power is obtained. figure.5.1 (a) and fig.5.1 (b) show the evolution of the objective function. the results of the optimal geometrical dimension of the rotor parameters are summarized in table 5.1 and table 5.2.  first optimization procedure : the output power, at idle speed, increases from 0.742 kw to around 1.27 kw which constitutes 71.15% improvement. the armature current increases from 20.62 a to 35.27 a.  second optimization procedure: at idle speed, the maximum output power reaches 1.3kw which constitutes 53.68% improvement. table 5.1.first optimization procedure, base and optimized values of design variables optimization parameter value before optimization (mm) value after optimization (mm) air gap length 0.65 0.3 height of claw tip 4.7 1 core length 29 25 width of the base of the claw 24.7 23.13 width of claw tip 6 4.35 implicit parameter distance between claw 6.6 8.7 claw pole length 29 25 claw side plate thickness 11.5 13.5 core radius 41.4 42.1 outer rotor radius 87.3 88 table 5.2.second optimization procedure, base and optimized values of design variables optimization parameter value before optimization (mm or a) value after optimization (mm or a) difference in parameter value (between 1st and 2nd optimization) (mm) air gap length 0.65 0.3 0 height of claw tip 4.7 1 0 core length 29 25 0 width of the base of the claw 24.7 23.94 -0.81 width of claw tip 6 4.17 0.18 armature current 29.08 32.09 — implicit parameter distance between claw 6.6 8.4 0.3 claw pole length 29 25 0 claw side plate thickness 11.5 13.5 0 core radius 41.4 42.2 0.1 outer rotor radius 87.3 88 0 optimal rotor design of claw-pole alternator 129 (a) (b) fig.5.1. evolution of output power versus number of iteration for the (a) first optimization procedure. (b) second optimization procedure.  impact of the air gap length: the air gap length can affect the output power considerably. in theory, to maximize the output power of an alternator, the air gap length should be designed as small as possible. for a typical alternator, the nominal air gap is approximately 0.4 mm but it can be larger or smaller. when the air gap is reduced, the output power is substantially improved [1]. however, the air gap of a claw-pole alternator should not be designed too small. when running at very high speeds, the alternator rotor poles will deflect due to centrifugal forces and the pole tip will touch the stator. in this investigation, the optimization process, can give solutions for very small values of the air gap, but the technically feasible solution selected is the one which corresponds to the air gap lower limit given above. the air gap flux is increased from 2.627 10-4 wb to 3.333 10-4 wb.  impact of the width of the base of the claw and the width of claw tip: these two parameters, determine the exchange surface between rotor and stator through the air gap. it is important to recall that, the magnetic radial flux is more concentrated at the basis of the claw pole than at its extremity. hence, the radial fluxes increases when the width of the base of the claw pole increases. however, the leakage fluxes between two adjacent claw poles, which are the most important leakage fluxes, increase when the width of the base of claw increases. the optimization process gives the optimal value of these two variables which lead to the maximum output power. in table 5.3, one can observe the reduction of the leakage flux between two adjacent claw poles and leakage flux between claw pole tip and opposite plate.  impact of the height of claw tip: the leakage fluxes between the claw pole and rotor winding can be affected by this parameter. to reduce these leakage fluxes, the distance between rotor core and the bottom of the claw should be designed as far as possible. this explains the fact that the optimal value of this variable tends to the lower limit given above. table 5.3 shows the reduction of the leakage flux between the claw pole and rotor winding.  impact of the core length: the optimization process assigns the lower limit to this variable. however, what is lost in length will be gained in claw side plate thickness. as a result, the junction who binds the core to the claw side plate is larger which enables more axial flux to be captured. table 5.3. leakages fluxes before and after optimization leakage fluxes base value (*10-4 wb) value after optimization (*10-4 wb) between adjacent claw pole fingers 1.1286 0.6855 130 n.brahimi, s.tahi, e. boudissa, m.bounekhla 6. alternator performance after optimization after optimization process, the output performances of the optimal claw pole alternator are evaluated. in fig. 6.1(a) and fig. 6.1 (b) the armature current and output power, are respectively plotted versus the rotational speed. it can be seen that, at idle speed of 1800rpm, optimized rotor design leads a higher armature current and hence to a higher output power. figure 6.1 (c) illustrates that only 1a field current is needed to deliver current at idle. in fig.6.2, (a) and (b) the maximum output power reaches 1.3 kw at idle speed. it can be observed that output performances of the optimazed alternator has been improved in the whole speed range from idle speed of 1800rpm to the cruising speed of 6000rpm. at the cruising speed of 6000 rpm the maximum output power reaches 4.5 kw. (a) (b) (c) between claw pole and rotor winding 0.04016 0.02449 between claw pole tip and opposite plate 0.02933 4.5953 10-6 optimal rotor design of claw-pole alternator 131 fig. 6.1. (a) stator phase current versus alternator speed. (b) alternator output power versus alternator speed. (c) stator phase voltage versus field current. (a) (b) fig. 6.2. (a) output power versus armature current. (b) output power versus stator phase voltage. 7. conclusion in this paper, the optimization method based on cyclic coordinate method coupled with magnetic equivalent circuit (mec) model, has been applied successfully for the design of rotor claw pole alternator. the main objective was to maximize the alternator output power for given stator dimensions and footprint at the idle speed of 1800rpm. the output power increased significantly at idle speed, it was demonstrated that output power and armature current of the optimized claw-pole alternator increased from 0.742 kw to 1.27 kw and from 20.62 a to 35.27a, which constitute 71.15% and 71.04% improvement, respectively. this is predominantly due to change in reluctance of geometric parameters, consequently affecting the flux throughout the alternator. at idle, the maximum output power delivered from the optimized alternator achieved 1.3 kw which constitutes 53.68% improvement. the obtained results indicate significant performance improvement of the optimized alternator over the whole speed range (1800 rpm to 6000 rpm), whereas the geometry optimization process has been fully concentrated at the idle speed. at high speed of 6000 rpm, the maximum available output power increased from 2.9 kw to 4.5 kw, a 50.90% improvement future investigation would consider the use of the mec model in the design optimization of various generators taking advantage of the claw-pole structure in harvesting energy systems. references [1] alaeddini,a., darab, a., & tahanian, h., (2015). influence of various structural factors of claw-pole transverse flux permanent magnet machines on internal voltage using finite element analysis. serbian journal of electrical engineering, 2 (12), 129-143. doi: 10.2298/sjee1502129a [2] albert, l., chillet, c., jarosz, a. & wurtz, f., (2003). analytic modeling of automotive claw-pole alternator for design and constrained optimization. 10th eur. conf. power electronics and applications, in processing, toulouse, france. [3] bazaraa m.s., sherall h.d. & shetty c.m., . (1993). nonlinear programming theory and algorithms, john wiley & sons, inc. 132 n.brahimi, s.tahi, e. boudissa, m.bounekhla [4] boldea, i., (2006). variable speed generators. crc press, taylor & francis group, second edition, boca raton. [5] dajaku, g., lehner, b., dajaku, xh., pretzer, a. & gerling, d. (2016). hybrid excited claw pole rotor for high power density automotive alternators. electrical machines (icem), xxii international conference of the ieee, lausanne, switzerland. [6] delale, a., albert, l., gerbaud, l. & wurtz, f. (2004). automatic generation of sizing models for the optimization of electromagnetic devices using reluctance networks. ieee transactions on magnetics, 2 (40), 830-833. [7] hecquet,m. & brochet,p., (1995) modeling of a claw-pole alternator using permanence network coupled with electric circuits, ieee transactions on magnetics, vol 31, n°3,2131-2134. [8] hidaka,y. & igarashi, h. (2018). three dimensional shape optimization of claw pole motors. j. adv. simulat. sci. eng., 1 (4), 64-77. [9] ibala, a., rabhi, r., & masmoudi, a. (2011). mec-based modeling of claw pole machines: application to automotive and wind generating systems. international journal of renewable energy research, ijrer, 3 (1), 1-8. [10] ibala, a. & masmoudi, a. (2010). accounting for the armature magnetic reaction and saturation effects in the reluctance model of a new concept of claw-pole alternator. ieee transactions on magnetics. 11 (46), 3955 – 3961. doi: 10.1109/tmag.2010.2055882 [11] ibala, a. & masmoudi, a. (2015). mec-based prediction to the loss of claw pole alternators, the inter. journal for computation and mathematics in electrical and electronic compel, 6 (34), 1719 – 1730. http://dx.doi.org/10.1108/compel-062015-0216. [12] ivankovic, r., cros, j., et al., (2012). power electronic solutions to improve the performance of lundell automotive alternators. new advances in vehicular technology and automotive engineering. intechop isbn: 9535106982. [13] jurca, f.n. & martis, c. (2012). theoretical and experimental analysis of a three-phase permanent magnet claw pole synchronous generator. iet electric power applications, 8 (6), 491-503. [14] moallem, m. & dawson, g.e., (1998), an improved magnetic equivalent circuit method for predicting the characteristic of highly saturated electromagnetic devices. ieee transactions on magnetics, 5 (34) , 3632-3635. [15] ostovic, v., miller, j.m., k.garg, v., schultz, r.d. and shawn, h. (1999). a magnetic equivalent circuit based performance computation of a lundell alternator, ieee transaction on industry application, 4 (35) [16] perez, s., (2013), contribution au dimensionnement optimal d’alternateur à griffe sans aimants apport des alliages feco [contribution to optimal sizing of claw-pole alternator without magnetcontribution of feco alloys], phd thesis, edeeats, grenoble, france. [in french]. [17] rakotovao, m.,(1993), un modèle operationnel complet pour l’alternateur à griffe [a complete operational model for automotive claw-pole alternator], phd thesis, cachan france, [in french] https://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=7592391 https://www.researchgate.net/journal/0018-9464_ieee_transactions_on_magnetics https://www.researchgate.net/journal/0018-9464_ieee_transactions_on_magnetics https://www.researchgate.net/journal/0018-9464_ieee_transactions_on_magnetics https://doi.org/10.1109/tmag.2010.2055882 optimal rotor design of claw-pole alternator 133 [18] rebhi, r., ibala, a. & masmoudi, a. (2013). on the modeling of a doublyexcited brushless claw pole alternator: application to a solar-wind-wave hybrid energy conversion system, eighth international conference and exhibition on ecological vehicles and renewable energies, ever, monaco, doi: 10.1109/ever.2013.6521547 [19] sharma,a., (2013),use of anisotropic materials in claw-pole alternator to reduce leahage flux, proc. annual ieee india conference (indicon),india doi: 10.1109/indcon.2013.6725997 [20] tahi, s., ibtiouen, r. & bounekhla, m. (2011). design optimization of two synchronous reluctance machines structures with maximized torque and power factor, progress in electromagnetic research b, 35, 369–387. [21] upadhayay,p., lebouc, a.k., garbuio, l., mipo, j.c. & dubus, j.m., (2017), design & comparison of conventional & permanent magnet based claw-pole machine for automotive application, proc. of the 15h int. conf. on electrical machine (elma), sofia, bulgaria, doi: 10.1109/elma.2017.7955390 [22] whaley, d.m., soong, w.l. & ertugrul, n. (2004). extracting more power from the lundell car alternator, australasian universities power engineering conference, aupec, brisbane, australia. [23] xiao-hual, bao & liu mou-zhi, (2012). parameter analysis and optimal design for mobile claw-pole alternator, applied mechanics and materials vols. 130-134, 658-661, trans tech publications, switzerland. [24] xiaohua bao, qingling he, qunjing wang & youyuan ni, (2008). research and optimal design on hybrid excitation claw-pole alternator for automobile application, electrical machines and systems, ieee international conference, 34933496, icems, wuhan, china. [25] zhenyang zhang, huijuan liu & tengfei song, (2016), optimization design and performance analysis of a pm brushless rotor claw pole motor with fem, machines, mdpi, 3 (4). [26] zaïm m. e., dakhouche k. & bounekhla m., (2002). design for torque ripple reduction of a three-phase switched reluctance machine, ieee transactions on magnetics, 38(2): 1189–1192. https://doi.org/10.1109/ever.2013.6521547 https://doi.org/10.1109/indcon.2013.6725997 https://doi.org/10.1109/elma.2017.7955390 https://ieeexplore.ieee.org/xpl/mostrecentissue.jsp?punumber=4753968 adv syst sci appl 2021; 03:63–74 published online at https://ijassa.ipu.ru. the model of the production side of the russian economy nikolay pilnik1, stanislav radionov1* 1financial research institute of the ministry of finance of the russian federation, moscow, russia abstract: we consider the model of the production side of the russian economy. this model is derived as the solution of the nonlinear dynamic optimization problem of the macroeconomic agent we call producer. this agent maximizes his discounted profit flow under technologic, demographic and financial constraints. we use the method of relaxation of complementary slackness conditions to transform the model to the more regular one and evaluate its parameters on the russian macroeconomic data. we show that this model can successfully replicate the large set of russian macroeconomic indicators such as gross domestic product, loans of producers, volume of fixed assets etc. keywords: dynamic optimization, complementary slackness conditions, production of gdp 1. introduction and literature review we present the model of production side of the economy, which consists of all economic agents who produce the added value. in the context of general equilibrium models these agents can be considered as elements of macroeconomic agent producer, similarly to agents consumer, bank, government and others. this agent generates a supply of good and demand for labor and capital, which is typical for macroeconomic models, but also has connections with banking system in terms of loans both in national and foreign currency. to our knowledge, this is the first model which describes the process of production and interactions of producers with the banking system on such a detailed level. we consider the maximization problem of discounted profit flow of this macroeconomic agent under technologic, demographic and financial constraints. it leads to the nonlinear dynamic optimization problem. we solve it using lagrange method and transform the solution by relaxation of the complementary slackness conditions (see [22]). this transformation leads to the nonlinear dynamic system. we estimate its parameters on the russian macroeconomic data and find a set of parameters that allows to reproduce it with rather high accuracy. in the modern literature, the works which are closest to ours are devoted to the estimation of production functions. this topic is considered as one of the most important in the empiric economic studies, because it deals with the fundamental economic process of formation of goods. the description of this process in terms of mathematical models is thus essential for the understanding of the functioning of the economy. apart from theoretical interest, it also has crucial policy implications. a large body of econometric literature is devoted to the estimation of production functions on the microdata (cobb-douglas function is usually used). the problem of obtaining reliable estimations turns out to be rather challenging because of simultaneity – if firm knows its productivity level when choosing inputs, it leads to the correlation between errors and ∗corresponding author: saradionov@edu.hse.ru 64 n. pilnik, s. radionov regressors. another problem is that firms with low productivity levels are harder to observe because they cease to exist sooner. several methods were developed to overcome these difficulties. probably the most popular ones were proposed in [20] and [17], but they were questioned in [2]. other methods were proposed, among others, in [7] and [15]. modelling of gross domestic product in terms of production functions is also widely used in the literature, especially in the works dealing with output gaps and potential growth rates, see for example [11], [12], [14], [16]. these works adopt a rather straightforward methodology which could potentially also suffer from simultaneity problem. in the context of gdp estimation this problem was addressed, for example, in [23]. description of gdp production in terms of production function is also ubiquitous in the modern macroeconomic dynamic stochastic general equilibrium (dsge) models, however they usually use a very simple cobb-douglas production functions. it should be noted that over the last years, a ces production function becomes more and more popular tool of production side of the economy, because empirical studies reject the hypothesis of elasticities of substitution between capital and labor being equal to one (the case of cobb-douglas function). the constant elasticity of substitution production function was derived in [4], and the numerous articles with estimation of ces production function on us data soon followed. in the first estimates the elasticity of substitution was indeed close to one (e.g., [10]), but the methodology used for these estimations was criticized. use of more rigorous econometric methods led to estimates well below unity (e.g., [18], [19], [9], [13]). new arguments in favor of cobb-douglas function were presented in [6], but they were again challenged in [3]. it should also be noted that more flexible production function than cobbdouglas is required in many growth models to generate the plausible dynamics (e.g. [1]), and in policy analysis (e.g. [8], [5]). a thorough discussion of ces production function is presented in the special issue of journal of macroeconomics in june 2008. 2. the model: statement we consider the model of the whole sector of the economy, which generates the added value, as a single agent. we will reduce it to the usual dynamic model that determines the supply of produced product, the demand of producer for loans, investments and labor, taking prices, wages, interest rates and other external factors on the market into account. there are several non-standard features in the proposed model. first, we divide investments into two parts: investments in the maintenance of fixed assets and investments in building up fixed assets. this approach allows to explain the fluctuations of output with higher accuracy, perceiving these investments as transaction costs and capital costs respectively. secondly, the production function depends on the volume of used fixed assets, adjusted on the level of investment in the maintenance of fixed assets. thirdly, we use a similar scheme to describe the contribution of labor in output – similarly to the fixed assets, we include human capital in the production function, which can be increased due to special wage investments. the formation of production capital (fixed assets) by the producer is described by the following relations: d dt m (t) = jm (t) − δam (t)m (t) , (2.1) 0 ≤ jm (t) . (2.2) in this equation m (t) are the fixed assets, jm (t) are the investments in building up fixed assets, δam (t) is the depreciation coefficient. this equation allows to calculate the series jm (t) based on statistics on the level of fixed assets and depreciation. the balance of investments j (t) = jm (t) + ju (t) , (2.3) copyright © 2021 assa. adv syst sci appl (2021) the model of the production side of the russian economy 65 where j (t) is the overall level of investment (gross fixed capital formation) in base year prices, allows to calculate the investments in the maintenance of fixed assets ju (t). further we assume that the prices of these investments may be different, therefore the balance of investments in nominal volumes can be written as pj (t) = pm (t) jm (t) + pu (t) ju (t) , (2.4) where pj (t) is the overall level of investments (gross fixed capital formation) in the current prices, pm (t) is the deflator of investments in building up fixed assets, pu (t) is the deflator of investments in the maintenance of fixed assets. the formation of human capital is described in a manner similar to the description of fixed assets: d dt h (t) = jh (t) rmax (t) − δah (t)h (t) , (2.5) 0 ≤ jh (t) , (2.6) where h (t) is the human capital of the economically active (but not necessarily employed) population per capita, jh (t) is the investments in building up human capital, δah (t) is the depreciation coefficient. everywhere further in the model we will distinguish the total economically active population rmax (t) and the number of employed r (t). we assume that human capital is distributed among all participants of the economically active population, but only the part corresponding to the employed workers is used. the statistical data provides information on the total amount of payments made by the producer to its employees – wage and mixed income w (t). further, we assume that this flow (as well as investment costs) consists of two parts: payments to employees wr (t)r (t) and investments in human capital wh (t) jh (t). hence, investment in human capital actually leads to the productivity increase of the economically active population in the economy, but does not lead to its growth. the most important element of the producer’s problem is the production function, which describes the relationship between the use of factors of production and the resulting output. when modeling the russian economy, linear production function can be used. the procedure its estimation is quite simple and can be is carried out by standard econometric tools. moreover, when instead of the indicator of fixed capital, the indicator of investments in fixed capital is used, it demonstrates a fairly high accuracy on the russian data. however, it cannot be used in the optimization problem because investments will not enter the first order conditions of the producer and it is impossible to obtain an expression for it. another approach is using the cobb-douglas function, which is popular, for example, in dsge models. but another problem arises here: on the russian data, the use of capital in such a production function leads to a poor model accuracy, and using of investments instead of capital is debatable from a theoretical point of view. that is why the presented producer model separates the long-term process of capital replacement and the short-term process of changing the load of this capital. such a technique, as will be shown below, allows to obtain a sufficiently high accuracy of the data reproduction without restrictions of the linear production function and with explicit use of capital in the model. thereby we use the following production function: y (t) = a (um (t)m (t))α (uh (t)h (t)rmax (t))β , (2.7) where um (t) and uh (t) are the utilization rates, respectively, of fixed assets and human capital. we define these indicators in the following non-linear way: um (t) = ( ju (t) m (t) )b , uh (t) = ( r (t) rmax (t) )a . (2.8) copyright © 2021 assa. adv syst sci appl (2021) 66 n. pilnik, s. radionov we substitute (2.8) into (2.7) and also substitute a with several normalization coefficients in order to avoid the dimensionality problem. as a result, the production function takes the form y (t) = y0 ( ju (t) j0 )bα( m (t) m0 )α (1−b)( h (t) h0 )β ( r (t) r0 )aβ ( rmax (t) rmax0 )(1−a)β . (2.9) depending on the market conditions, however, the volume of production 0 ≤ y (t) and the volume of products sold 0 ≤ yp (t) may vary. the producer can thus form the stock of the product in size 0 ≤ z (t), which can be calculated as d dt z (t) = y (t) − yp (t) . (2.10) denote the loans attracted by the producer by l (t). the average terms for which loans are attracted will be denoted by (βl (t)) −1. the variable βl (t) will be referred below as the inverse duration and interpreted as the average frequency of the loans return. then the dynamics of loans is described by the equation d dt l (t) = k (t) − βl (t)l (t) , (2.11) where 0 ≤ k (t) is the flow of newly attracted loans. the producer pays interest on borrowed funds rl (t)l (t), where rl (t) is the effective interest rate on loans. in addition to loans in national currency (rubles), the producer also attracts loans vl (t) in dollars. the dynamics of foreign currency loans and deposits is described by the equation d dt vl (t) = vk (t) − βvl (t) vl (t) . (2.12) here βvl (t) is the reverse durations of currency loans. newly attracted foreign currency loans are denoted by 0 ≤ vk (t). the effective interest rate on foreign currency loans is denoted by rvl (t). during its activities, the producer pays four types of taxes: value added tax, labor tax (social contributions), property tax and income tax. the rates of these taxes are denoted by τy (t), τr (t), τtm (t), τpr (t). for servicing operations related to loans and investments, the producer uses current account n (t). it is assumed that its volumes in the balance sheet are proportional to ruble and foreign currency loans, fixed assets and human capital. these proportions are defined by the coefficients νl, νvl, νm, νh: n (t) ≥ νll (t) + νvlwvl (t) vl (t) + νmpm (t)m (t) + + νhwh (t) (1 + τh (t))rmax (t)h (t) , (2.13) where wvl (t) denotes the exchange rate. other expenses of the producer are denoted by oc o (t) and are considered as the exogenous variable. thus, the financial balance of the producer can be written as d dt n (t) = k (t) − βl (t)l (t) − rl (t)l (t) − (1 + τpr (t)) pr (t) + wvl (t) (vk (t) − βvl (t) vl (t) − rvl (t) vl (t)) −oc o (t) + (1 − τy (t)) py (t)yp (t) − pj (t) ju (t) − pm (t) jm (t)− τtm (t) ptm (t)m (t) − (1 + τr (t))wr (t)r (t) − (1 + τh (t))wh (t) jh (t) , (2.14) copyright © 2021 assa. adv syst sci appl (2021) the model of the production side of the russian economy 67 where pr (t) is the profit after taxation. the above relationships represent the limitations imposed on the producer’s ability to choose the values of its planned variables (controls): h (t) , jh (t) , jm (t) , ju (t) , k (t) , l (t) ,m (t) , n (t) ,pr (t) , r (t) ,yp (t) , z (t) , vk (t) , vl (t) , y (t) . (2.15) according to the principle of rational expectations, when planning its control variables, the producer can rely on an accurate forecast of information variables: oc o (t) , rmax (t) , βl (t) , βvl (t) , δah (t) , δam (t) , pj (t) , pm (t) , ptm (t) , py (t) , rl (t) , rvl (t) , τh (t) , τpr (t) , τr (t) , τtm (t) , τy (t) , wh (t) , wr (t) , wvl (t) . the planned variables of the producer are thus a function of current and future values of information variables. the goal of the producer in the model is to maximize the total discounted utility from the undistributed profit after taxation with discount rate being equal to the deflator: ∫ t t0 e−∆ t 1 − β ( pr (t) py (t) )1−β dt. (2.16) the problem is supplemented with terminal condition, which can be interpreted as a growth condition for some linear combination of the phase variables: ω (t0 ) γ ≤ ω (t ) , (2.17) where ω (t) = ah (t)h (t) + al (t)l (t) + am (t)m (t) + an (t)n (t) + az (t)z (t) + avl (t) vl (t). this is an analogue of no ponzi condition written in terms of the producer’s own capital which is a difference between its assets and liabilities. to derive the solution of the maximization problem of the function (2.16) under constraints (2.1) – (2.14) and the terminal condition (2.17) with respect to the variables (2.15), we write down a lagrange functional and find its saddle point. the dual variables are denoted by φ2 (t) ,φ4 (t) ,φ7 (t) ,φ9 (t). the obtained system, which is a set of sufficient conditions for optimality, is presented below divided into several groups. the first group are equations for primal variables, including equations from the original problem statement: d dt h (t) = jh (t) rmax (t) − δah (t)h (t) , d dt m (t) = jm (t) − δam (t)m (t) , d dt l (t) = k (t) − βl (t)l (t) , d dt vl (t) = vk (t) − βvl (t) vl (t) , d dt pr (t) = ( ρ (t) − ∆ η − d dt τpr (t) (1 + τpr (t)) η + d dt py (t) py (t) ( 1 − η−1 )) pr (t) , n (t) = νll (t) + νvlwvl (t) vl (t) + νmpm (t)m (t) + νhwh (t) (1 + τh (t))rmax (t)h (t) , copyright © 2021 assa. adv syst sci appl (2021) 68 n. pilnik, s. radionov the second group are differential equations for dual variables d dt φ4 (t) = ( βvl (t) + ρ (t) − d dt wvl (t) wvl (t) ) φ4 (t) + (1 − νvl) ρ (t) − rvl (t) − d dt wvl (t) wvl (t) , (1 + τr (t))wr (t)r (t) = aβ (1 − τy (t)) py (t)y (t) , d dt φ7 (t) = ( δam (t) + ρ (t) − d dt pm (t) pm (t) ) φ7 (t) − (1 + νm) ρ (t) δam (t) − τtm (t) + d dt pm (t) pm (t) + α (1 − b) (1 − τy (t)) py (t)y (t) pm (t)m (t) , d dt φ2 (t) = (βl (t) + ρ (t)) φ2 (t) + (1 − νl) ρ (t) − rl (t) , pj (t) ju (t) = bα (1 − τy (t)) py (t)y (t) , d dt φ9 (t) = ( δah (t) + ρ (t) − d dt τh (t) 1 + τh (t) − d dt wh (t) wh (t) − d dt rmax (t) rmax (t) ) φ9 (t) − δah (t) − ( νh + νh d dt τh (t) + 1 1 + τh (t) ) ρ (t) + d dt wh (t) wh (t) + d dt rmax (t) rmax (t) + β (1 − τy (t)) py (t)y (t) (1 + τh (t))wh (t)rmax (t)h (t) , third group are complementary slackness conditions: [φ9 (t)][jh (t)], [φ7 (t)][jm (t)], [φ4 (t)][vk (t)], [φ2 (t)][k (t)], [ ( −τy (t) ρ (t) + ρ (t) + d dt τy (t) ) py (t) + ( d dt py (t) ) τy (t) − d dt py (t)][z (t)], where [a][b] means 0 ≤ a, 0 ≤ b, ab = 0. the last condition is the transversality condition that defines terminal condition for conjugate differential equations: ω (t0 ) γ = ω (t ) , where ω (t) = ah (t)h (t) + al (t)l (t) + am (t)m (t) + an (t)n (t) + + az (t)z (t) + avl (t) vl (t) . we transform this system according to the method described in detail in [22]. first, we transform it to the discrete time by replacing derivatives with increments. note that backward increments are used for direct variables and forward increments are used for dual variables. second, some direct variables are replaced by their values in the previous period. third, differential equations for the dual variables are replaced by expressions of corresponding separatrices. forth, the complementary slackness conditions are replaced by their more regular approximations (the relaxation of complementary slackness conditions). after all these transformations, we proceed to the dynamic system which can also be presented by several groups of expressions. copyright © 2021 assa. adv syst sci appl (2021) the model of the production side of the russian economy 69 first group consists of several expressions defining growth rates of different exogenous variables (in different forms) and one more auxiliary variable: gpm (t) = pm (t) − pm (t− 1) pm (t− 1) , gwh (t) = wh (t) − wh (t− 1) wh (t− 1) , gpy (t) = py (t) − py (t− 1) py (t− 1) , gtaupr (t) = τpr (t) − τpr (t− 1) τpr (t− 1) + 1 , gtauh (t) = τh (t) − τh (t− 1) 1 + τh (t− 1) , grmax (t) = rmax (t) −rmax (t− 1) rmax (t− 1) , gttauy (t) = −τy (t) − τy (t− 1) −τy (t− 1) + 1 , b (t) = −oc o (t) + y (t) (1 − τy (t)) py (t) − pr (t) (1 + τpr (t)) . second group consists of expressions defining gdp and its components included in the model – different kinds of investments, change of stocks and profit: y (t)(−aβ−bα+1) = y0 ( bα (1 − τy (t)) py (t) pj (t) j0 )bα( m (t− 1) m0 )α (1−b) × × ( h (t− 1) h0 )β ( aβ (1 − τy (t)) py (t) wr (t) (1 + τr (t))r0 )aβ ( rmax (t) rmax0 )(1−a)β , (1 + τr (t))wr (t)r (t) = aβ (1 − τy (t)) py (t)y (t) , pj (t) ju (t) = bα (1 − τy (t)) py (t)y (t) , jh (t) = b1 (grmax (t) + gtauh (t) + gwh (t) − δah (t) − ρ (t))rmax(t)h(t− 1)− − a1 ((−νh − 1) ρ (t) + grmax (t) + gtauh (t) + gwh (t) − δah (t))rmax (t)h (t− 1)− a1 β (τy (t) − 1) (gtauh (t) + 1) py (t)y (t) (1 + τh (t))wh (t) + cc1b (t) wh (t) (1 + τh (t)) , jm (t) = b2 (δam (t) − gpm (t) + ρ (t))m (t− 1)− a2 ((νm + 1) ρ (t) + δam (t) + τtm (t) − gpm (t))m (t− 1)− α (b− 1) (τy (t) − 1) py (t)y (t) pm (t) + cc2b (t) pm (t) , z (t) = z (t− 1) + (b5 − a5 (−gpy (t) − gttauy (t) + ρ (t)))z (t− 1) + cc5b (t) py (t) (1 − τy (t)) , pr (t) = ( −∆ − gtaupr (t) + ρ (t) η + ( −η−1 + 1 ) gpy (t) + 1 ) pr (t− 1) . hereinafter the expressions without time dependence such as y 0, a, α etc. are the parameters which will be estimated in the next section. next group are the variables defining stocks of physical and human capital: m (t) = jm (t) − δam (t)m (t− 1) +m (t− 1) , h (t) = jh (t) rmax (t) − δah (t)h (t− 1) +h (t− 1) . copyright © 2021 assa. adv syst sci appl (2021) 70 n. pilnik, s. radionov the last group are the variables defining financial variables, both stocks and flows. vk (t) = b3 (βvl (t) − gwvl (t) + ρ (t)) vl (t− 1) − a3 ((νvl − 1) ρ (t) + rvl (t) + gwvl (t)) vl (t− 1) − cc3b (t) wvl (t) , k (t) = (b4 (βl (t) + ρ (t)) − a4 ((νl − 1) ρ (t) + rl (t)))l (t− 1) + + (−1 + cc1 + cc2 + cc3 + cc5)b (t) , n (t) = νll (t− 1) + νvlwvl (t) vl (t− 1) + νmpm (t)m (t− 1) + + νhwh (t) (1 + τh (t))rmax (t)h (t− 1) , l (t) = k (t) − βl (t)l (t− 1) + l (t− 1) , vl (t) = vk (t) − βvl (t) vl (t− 1) + vl (t− 1) . 3. the model: identification and results the model is identified on the russian macroeconomic data from 2005q1 to 2021q1. the data used in the model can be divided into four groups. first, it is gdp series in the current and fixed prices and its elements by use (gross fixed capital formation and changes of stocks) and incomes (wages and profits). this data is published by rosstat, but it is currently presented by several partially overlapping series. we compose unified time series of gdp and its components using the assumption of preserving growth rates. gdp and its components have significant seasonalality, so the seasonality elimination procedure described in [21] is used. this procedure is chosen because it performs rather well for the short time series. second, several indicators from the balance of the russian banking system published by the bank of russia are used. these indicators include volumes of loans to producers in rubles and in foreign currency, its interest rates and durations. moreover, we use current accounts of non-financial organizations as the indicator of liquidity available to producer. all these indicators do not have a significant seasonal component, so no additional transformation is required. third, the tax data published by the federal treasury is used. the taxes related to producer include value added tax, property tax, profit tax and social payments included in the consolidated budget of russia and budgets of non-budgetary funds. the treatment of this data is the most complicated because it has not only seasonal component but also in-year irregularity – the taxes not collected in the current quarter are paid in the next one. the algorithm similar to the one described in [21] is used for the effective tax rates calculated as the ratio of the collected tax to its base. finally, forth group consists of all the other indicators such as dollar to ruble exchange rate, volume of fixed assets and amortization rate. the last two indicators are published on a yearly basis, so they were decomposed into a quarterly data using the assumption of the fixed growth rate during the year. the parameters of the model are estimated using the following approach. we consider growth rates of the following eight variables: y ,r,m , l, vl,n , pr,w to the corresponding quarter of the previous year. for example, in case of gdp we have the following growth rates from the model and from data: gry mod(t) = y (t) − y stat(t− 4) y stat(t− 4) , gry stat(t) = y stat(t) − y stat(t− 4) y stat(t− 4) , (3.18) where y stat is the statistical value of the real gdp, y is the value calculated via the model. the function of errors is the sum of squares of gry mod(t) − gry stat(t) for every t and copyright © 2021 assa. adv syst sci appl (2021) the model of the production side of the russian economy 71 similar expressions for other variables. it should be noted that flow variables such as k and vk are not considered because of their exceptionally high volatility. this leads to the problem of unconstrained optimization, because every parameter of the model can be an arbitrary real number, which is solved using lsqnonlin command from matlab optimization package. we use the monte carlo method – initial points are generated in a rather wide region based on our a priori assumptions on the values of the model parameters and, after large number of iterations, we find the set of parameters which provides a good fit of the data and is also economically meaningful. mean absolute percent errors of the model variables are presented in the table 3.1. table 3.1. mean absolute percent errors of the model variables growth rates y r m l vl n pr w 2.46 0.82 1.62 2.96 6.43 6.86 3.03 2.85 moreover, the accuracy of data fit is shown on the plots below. (a) gdp growth rate (b) number of employed growth rate (c) fixed assets growth rate (d) ruble loans growth rate copyright © 2021 assa. adv syst sci appl (2021) 72 n. pilnik, s. radionov (e) foreign currency loans growth rate (f) current accounts growth rate (g) profit growth rate (h) wages growth rate as we can see in the plots above and the table 3.1, the model performs rather well for all model variables. the percent error for vl is relatively high because the values of this variable are very small, so small errors in absolute terms may lead to high relative errors. the percent error for n is relatively high because of the volatility of this variable. 4. conclusion the model of the macroeconomic agent producer is presented. to our knowledge, this is the first model which describes not only the technologic and demographic, but also financial constraints of firms and other agents who produce the added value. application of the relaxation of complementary slackness conditions method allows to transform the model to a nonlinear dynamic system. the model is estimated on the russian macroeconomic data, and the set of parameters was found which allows to reproduce the dynamics of russian macroeconomic indicators with rather high accuracy. acknowledgements copyright © 2021 assa. adv syst sci appl (2021) the model of the production side of the russian economy 73 this work is supported by the russian science foundation under grant 21-18-00482. references [1] d. acemoglu, s. johnson, j. robinson, and y. thaicharoen. institutional causes, macroeconomic symptoms: volatility, crises and growth. journal of monetary economics, 50(1):49–123, 2003. [2] d. a. ackerberg, k. caves, and g. frazer. identification properties of recent production function estimators. econometrica, 83(6):2411–2451, 2015. [3] p. antras. is the us aggregate production function cobb-douglas? new estimates of the elasticity of substitution. contributions in macroeconomics, 4(1). [4] k. j. arrow, h. b. chenery, b. s. minhas, and r. m. solow. capital-labor substitution and economic efficiency. the review of economics and statistics, pages 225–250, 1961. [5] d. backus, e. henriksen, and k. storesletten. taxes and the global allocation of capital. journal of monetary economics, 55(1):48–61, 2008. [6] e. r. berndt. reconciling alternative estimates of the elasticity of substitution. the review of economics and statistics, pages 59–68, 1976. [7] r. blundell and s. bond. gmm estimation with persistent panel data: an application to production functions. econometric reviews, 19(3):321–340, 2000. [8] r. s. chirinko. corporate taxation, capital formation, and the substitution elasticity between labor and capital. national tax journal, 55(2):339–355, 2002. [9] r. m. coen. tax policy and investment behavior: comment. the american economic review, 59(3):370–379, 1969. [10] p. j. dhrymes and p. zarembka. elasticities of substitution for two-digit manufacturing industries: a correction. the review of economics and statistics, pages 115–117, 1970. [11] b. dunaev. calculating gross domestic product as a function of labor and capital. cybernetics and systems analysis, 40(1):86–96, 2004. [12] d. ecfin. the production function approach to calculating potential growth and output gaps estimates for eu member states and the us. 2006. [13] r. eisner and m. i. nadiri. investment behavior and neo-classical theory. the review of economics and statistics, pages 369–382, 1968. [14] n. p. epstein and c. macchiarelli. estimating poland’s potential output: a production function approach. number 10-15. international monetary fund, 2010. [15] a. gandhi, s. navarro, and d. rivers. on the identification of production functions: how heterogeneous is productivity? technical report, cibc working paper, 2011. [16] t. kawamoto, t. ozaki, n. kato, k. maehashi, et al. methodology for estimating output gap and potential growth rate: an update. 2017. [17] j. levinsohn and a. petrin. estimating production functions using inputs to control for unobservables. the review of economic studies, 70(2):317–341, 2003. [18] r. e. lucas jr and l. a. rapping. real wages, employment, and inflation. journal of political economy, 77(5):721–754, 1969. copyright © 2021 assa. adv syst sci appl (2021) 74 n. pilnik, s. radionov [19] g. maddala. productivity and technological change in the bituminous coal industry, 1919-54. journal of political economy, 73(4):352–365, 1965. [20] g. s. olley and a. pakes. the dynamics of productivity in the telecommunications equipment industry. econometrica: journal of the econometric society, pages 1263– 1297, 1996. [21] n. pilnik, i. pospelov, and i. stankevich. ob ispolzovanii fiktivnikh peremennikh dlya resheniya problemy sezonnosti v modelyah obshego ravnovesiya (on the use of dummy variables for the solution of the seasonality problem in the general equilibrium models). hse economic journal, 19(2):249–270, 2015. [22] n. pilnik, s. radionov, and a. yazikov. the model of the russian banking system with indicators nominated in rubles and in foreign currency. in y. evtushenko, m. jaćimović, m. khachay, y. kochetov, v. malkova, and m. posypkin, editors, optimization and applications, pages 427–438, cham, 2019. springer international publishing. [23] a. swamy and b. fikkert. estimating the contributions of capital and labor to gdp: an instrumental variable approach. economic development and cultural change, 50:693– 708, 02 2002. copyright © 2021 assa. adv syst sci appl (2021) introduction and literature review the model: statement the model: identification and results conclusion adv syst sci appl 2019; 04; 58-65 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/763 resource allocation with a probabilistic amount of resource vladimir burkov*1, irina burkova1,2, larisa rossikhina2,3 1) trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: vlab17@bk.ru 2) academy of management, the ministry of internal affairs of the russian federation, moscow, russia 3) voronezh institute of the federal penitentiary service of the russian federation, voronezh, russia received july 7, 2019; revised december 14, 2019; published december 31, 2019 abstract: the allocation problem of a limited resource under the probabilistic uncertainty over its amount is considered. the principal has a certain amount of resources to be allocated by him/her among consumers (agents). each agent submits the request for the resource to the principal. the principal allocates the resource in accordance with a specified resource allocation mechanism. in the theory of active systems, the priority-based resource allocation mechanisms were proposed and investigated. with these mechanisms, the resource is allocated proportionally to the values of the agent’s priority functions. three types of the priority-based mechanisms were identified, namely, the mechanism of absolute priorities, the mechanism of straight priorities and the mechanism of reverse priorities. previously, the priority-based mechanisms were considered under the assumption that the amount of available resource to be allocated by the principal is known. however, in many real resource allocation problems arising in practice this amount is often unknown. in this paper, the priority-based mechanisms are studied for the case in which the agents know the principal’s resource allocation function. the mechanisms of resource allocation based on the principle of reverse priorities are studied. keywords: resource allocation; mechanisms of reverse priorities; probabilistic uncertainty; nash equilibrium. 1. introduction the allocation problems of limited resources are widespread in practice [1, 2, 3]. the classical resource allocation scheme is as follows [4]. the principal has a certain amount of resources to be allocated by him/her among consumers (agents). each agent submits the request for the resource to the principal. the principal allocates the resource in accordance with a specified resource allocation mechanism [5, 6, 7]. in the theory of active systems, the priority-based resource allocation mechanisms were proposed and investigated [8, 9]. with these mechanisms, the resource is allocated proportionally to the values of the agent’s priority functions. three types of the priority-based mechanisms were identified, namely, the mechanism of absolute priorities, the mechanism of straight priorities and the mechanism of reverse priorities. in absolute priority mechanisms, resource allocation is proportional to the priorities set by the principal. these mechanisms are non-manipulated. it means that submit a request that * corresponding author: vlab17@bk.ru resource allocation with a probabilistic amount of resource 59 copyright ©2019 assa adv. in systems science and appl. (2019) reflects the true needs of the resource is beneficial for agents. their disadvantage is the fact that agents do not actually affect resource allocation. in the mechanisms of direct priority resources are allocated proportionally to the priorities that grow with the request on the principle of "more you ask – more you get". this principle leads to a trend of increasing of the request. it is significant disadvantage of the principle of direct priorities. however, it is still widely used in practice. in the mechanisms of reverse priority resource allocation is also proportional to functions of priority. however, these functions are decreasing functions of the request value on the principle of "more you ask – less you get". the principle of reverse priorities encourages resource savings. this is its significant advantage. it is experimentally tested in the allocation of water resources [2]. of the three described mechanisms of priority, the reverse priority principle is the most preferred. previously, the priority-based mechanisms were considered under the assumption that the amount of available resource to be allocated by the principal is known. however, in many real resource allocation problems arising in practice this amount is often unknown [10, 11]. in this paper, the mechanism of reverse priorities is investigated for the case in when agents know the principal’s resource allocation function. 2. problem statement consider an active system consisting of the principal and agents. the principal has a resource to be allocated among the agents in accordance with their requests. while submitting the requests, the agents have information about the principal’s allocation function f(r) only, where r is the amount of available resource to be allocated. the goal function of each agent i is an increasing function of the received resource xi. denote by si the request for the resource submitted by agent i. the principal allocates the resource using the mechanism of reverse priorities. the problem is to determine the nash equilibrium situation for three cases. in the first case the amount of resource available to the principal is known to the agents. in the second case, the agents know the distribution function f(r) of the resource available to the principal. in the third case, the distribution function of the resource available to the principal is discrete (more precisely, the resource takes two values with corresponding probabilities). 3. mechanism of reverse priorities: deterministic case consider the deterministic case in which the amount of resource available to the principal is known to the agents. define the priority function of agent i by   is ia isi  , ni ,1 , where ai is a parameter that restricts the agent’s priority. the resource allocation mechanism has the form   min , a rix s si i s yi           , (3.1) where aj y sj j   . assume the goal functions of the agents are increasing in xi. the deterministic statement of the problem yields the nash equilibrium 60 v. burkov, i. burkova, l. rossikhina copyright ©2019 assa adv. in systems science and appl. (2019) 2 1 y aj r j         , a ri si a j j    . in this case, isix  for all ni ,1 . now, consider another modification of the priority function given by   isiaisi  , ni ,1 . first, analyze the deterministic case:           r y isia issix ,min , where    j jsjay . in the nash equilibrium, r y isia is   , which gives ry ira is   and ry yia isia   . from the condition      j j ja yr y jsjay it follows that ray  * , r j ja ia is   * . the nash equilibrium exists only if the condition a > r is satisfied. 4. mechanism of reverse priorities: probabilistic case consider the probabilistic case in which the agents choose their requests for the resource using the information about the distribution f(r) of the available amount r only. consider the priority function   is ia isi  , ni ,1 , and the resource allocation mechanism   min , a rix s si i s yi           . denote by f(r) a distribution function of the available amount of the resource and let it be continuously differentiable with respect to r. the expected amount of resource requested by agent i is resource allocation with a probabilistic amount of resource 61 copyright ©2019 assa adv. in systems science and appl. (2019)      1 0 ri a rim s df r s f ri i s yi        , where 2 s yiri ai   , ni ,1 . find the maximum of this value over si under the hypothesis of weak contagion (the influence of si on y is small). in other words, while choosing their requests for the resource, the agents neglect the influence of si on y, considering y to be just a parameter. note that        ir ir drrfirfirrrdf 0 0 . calculate     2 0 riadm i r f r f r dri i ds s yi i                 ' ' 1 a dr dri i ir f r f r s f ri i i i i s y ds dsi i i      . a series of trivial transformations finally give     1 1 2 0 ridm f r f r dri ds ri i     . the resulting expression is independent of ai. assume the function м(s) is convex and therefore has a maximum point. the equilibrium value ri is the same for all agents, i.e., ri=r* for all i. the maximum point r* can be determined from the first-order optimality condition     * 1* 2 1 0 r f r f r dr ri   , where y ria is * *   . further, calculate  j ja r y j js ja y ** * , * * r j ja y   . finally, the nash equilibrium is ** r j ia ia is    . now, consider the priority function   isiaisi  , ni ,1 and the resource allocation mechanism           r y isia issix ,min . by analogy with the previous case, 62 v. burkov, i. burkova, l. rossikhina copyright ©2019 assa adv. in systems science and appl. (2019)             ir irfisrrdf y isia ism 0 1 , where s yiri a si i   ,  2isia iya ids dr   . calculate               irf ir drrfirfir yids dm 1 0 1           ir drrf y irf y ir 0 1 11 ,             irf y irf y ir irf ydr md 1' 1 1 2   0 ' 1        irf y ir . hence, the function  ism is concave, and the maximum point can be determined from the first-order optimality condition (the same for all agents)        r ydxxfrfyr 0 . 5. mechanism of reverse priorities: discrete case suppose the available resource r is q1 with a probability p1 and q2 with the probability 2 11p p  . adopt the mechanism of reverse priorities with the priority function   is ia isi  , ni ,1 . establish how the expected amount of the agent’s resource depends his/her request (for simplicity, the agent’s number i is omitted). calculate y aq d 1  and y aq d 2  . then three situations are possible as follows: 1) ds  . in this case , with the mechanism of reverse priorities the minimum will be achieved on the request, if 1qr  and if 2qr  . therefore,   ssm  . 2) dsd  . in this case, the minimum point ys aq x   1 will be achieved if 1qr  , and the minimum point sx  will be achieved if 2qr  . the expected amount of the resource will be   sp ys aq psm 2 1 1    . 3) ds  . in this case, the minimum point ys aq x   1 will be achieved if 1qr  , and the minimum point ys aq x   2 will be achieved if 2qr  . the expected amount of the resource will be     r sy a sy a qpqpsm     2211 . resource allocation with a probabilistic amount of resource 63 copyright ©2019 assa adv. in systems science and appl. (2019) the graph of the function  sm can be seen in fig 4.1. fig. 4.1 graph of function  sm . on the interval  dd , ,  sm is a convex function of s. therefore, it achieves a maximum at the points d or d. note that, if 1ds  , then dm  ; with 1qq  or 2qq  , the agent will receive the same amount dx  . if ds  , then dpp yd aq m 21 1    . the maximum of m is given by                 221 2 1 ;1max2 1 1;max pqp q q q y a dp yd q pd if the maximum is achieved on 1q , the agent will choose the strategy d; if on 22 2 1 1 qp q q p  , then the strategy d. define * 1p from the equation   2121 2 1 1 pqp q q q  . after simple calculations, 21 2* 1 qq q p   . so, if * 1 pp  , then each agent will choose the strategy d; if * 1 pp  , then the strategy d. in both cases, in the nash equilibrium the resources will be allocated as follows: q b ia ix  , where  i iab . next, consider the priority function   isiaisi  , ni ,1 . by analogy with the previous analysis,           y qsa sx ;min , where    j jsjay . using the equation   y qsa s   , find 64 v. burkov, i. burkova, l. rossikhina copyright ©2019 assa adv. in systems science and appl. (2019) 1 1 qy aq d   and 2 2 qy aq d   . if ds  , then   dxm  . if dsd  , then              2 11 2 1 1 p y qp s y a rsp y qsa pxm . if ds  , then       2211 qpqp y sa xm    . obviously, the maximum of  xm is achieved either at the point d or at the point d. hence,              dp y qda pdxm 2 1 1;             2 2 2 1 2 1 ; 1 1 max p yq q p yq q yq q a . the boundary value * 1p can be defined from the equation 2 2 2 1 2 1 1 1 p yq q p yq q yq q      . trivial calculations yield yq y p   1 * 1 . in contrast to the previous case, the value * 1p depends on y. if jdjs  , then    j bjdjay , 1qby  , b q p 1 1 * 1  . if idis  , ni ,1 , then b q p 2 1 * 1  . again, three cases are possible as follows. 1) b q p 1 1 . in this case, all agents will choose the strategy id , ni ,1 . 2) b q p 2 1 . in this case, all agent will choose the strategy id , ni ,1 . 3) b q p b q 1 1 2 1  . this situation has uncertainty. if all agents choose the strategy id , then b q p 1 1 * 1  ; therefore, * 1pp  and they will benefit more from the strategy id . if they choose the strategy id , then b q p 2 1 * 1  ; therefore, * 1pp  and the agent will benefit more from the strategy id . the situation is difficult to predict. resource allocation with a probabilistic amount of resource 65 copyright ©2019 assa adv. in systems science and appl. (2019) 6. conclusions in this paper, a study of the mechanism of inverse priorities for resource allocation in the deterministic and probabilistic case has been presented. for different priority functions the nash equilibrium has been obtained under the assumption that the goal functions of the agents are monotonically increasing in their amounts of resource. references 1. burkov v. n., gorgidze i. i., novikov d.a & yusupov b. s. (1997). models and mechanisms of distribution of costs and income in the market economy. moscow: institute of control sciences 2. burkov v. n., danev b. d. & enaev, a. k. (1989). large systems: modeling of organizational mechanisms. moscow: science. 3. burkov v. n., korgin n. a. & novikov d. a. (2009). introduction to the theory of control of organizational systems. moscow: librocom. 4. novikov, d. (2013), theory of control in organizations, new york: nova science publishers. 5. control mechanism (2013). organization management: planning, organization, stimulation, control, ed. by d. a. novikov. moscow: lenand. 6. mas-colell, a., whinston, m. d., & green, j. r. (1995). microeconomic theory (vol. 1). new york: oxford university press. 7. maschler, m., solan, e., & zamir, s. (2013). game theory (translated from the hebrew by ziv hellman and edited by mike borns). cambridge university press. 8. salanié, b. (2005). the economics of contracts: a primer. mit press. 9. goubko, m., burkov, v., kondrat’ev, v., korgin, n., & novikov, d. (2013). mechanism design and management: mathematical methods for smart organizations. new york: nova science publishers 10. burkov v. n., & kondratyev v. v. (1981) mechanisms of functioning of organizational systems. moscow: science. 11. ponomarev, v. a. & rossikhina , l.v. (2018) applied modeling of the mechanism of inverse priorities in the tasks of resource allocation using active experiment. bulletin of the voronezh institute of the federal penitentiary service of russia, 2, 94–99. 12. ponomarev v. a. (2018) a game-theoretic models of resource allocation. bulletin of the voronezh institute of the federal penitentiary service of russia, 4, 98–105. adv syst sci appl 2018; 1; 1-11 published online at http://ijassa.ipu.ru. copyright ©2018 assa. adv. in systems science and appl. (2018) algorithms for prosodic discourse feature interpretation in case of its processing using low-speed codecs. maxim a. bessonov 1 , natalia a. bessonova 2 , mais p. farkhadov 3 1) peoples' friendship university of russia, moscow, russia e-mail: bessonovma@gmail.com 2) peoples' friendship university of russia, moscow, russia e-mail: nbesson@yandex.ru 3) v.a. trapeznikov institute for control sciences of russian academy of science, moscow, russia e-mail: mais.farhadov@gmail.com abstract: in this article we propose two algorithms for discourse prosodic feature interpretation. the first algorithm based on wide phonetic categories and second algorithm based on audio signal melodic cross-correlation functions and short-timed energy series – as well as methodical recommendations for their use are proposed as a part of the problem of audio signal language identification based on a prosodic approach. an experimental evaluation of both algorithms is proposed. neural networks are used as a decision rule. wide phonetic categories were pause, pitch, noise. we have expanded wide phonetic categories to pause, pitch, noise, five levels of pitch, sites of decreasing energy, main maximum, adverse maximum. the total number of categories was 14. these algorithms can be applied for language identification or speaker identification. at the same time there is no requirement to restore the speech signal after processing it by low-speed codec. certainly, frames of the speech codec must contain such parameters as pitch, tone-noise parameter, energy. the base of speech signals consists of 10 languages 10 speakers per language. total time of the speech per speaker is 100 minutes. this time takes into account statistical regularities of languages. tests for evaluation of the algorithms were carried out with a multilayer perceptron. keywords: language identification, neural networks, discourse prosodic feature, wide phonetic categories. 1. introduction as numerous on-line human-machine interfaces are being developed, the problem of language identification stays unsolved. moreover, those systems are often required to support numerous languages. there are four methods for language identification: acoustic, phonotactic, lexical and prosodic. in one or more of their aspects, the first three methods are based on the discourse signal parameters: acoustic, mel-frequency cepstral coefficients, mixed mel-frequency cepstral coefficients and others. the prosodic approach uses parameters such as discourse melody, rhythm, tone and others [5–9]. prosodic parameters are difficult to describe and to interpret. therefore, in the present article two algorithms are proposed for discourse prosodic feature description in order to use them in automatic audio signal language assessment systems. the first algorithm is based on wide phonetic categories [4]. the second is based on discourse melody cross-correllation function and short-timed energy series. the main difference between the proposed algorithms and the ones described in the literature lays in their utilization for language assessment of audio signals that have been processed using low-speed codecs. this is based on the fact that low-speed codecs transmit mailto:nbesson@yandex.ru 2 m. bessonov, n. bessonova, m. farkhadov copyright ©2018 assa. adv. in systems science and appl. (2018) in communication channels such parameters as basal tone frequency, tone-noise signal and amplification of quasi-periodic fragments. 2. algorithms for prosodic features interpretation 2.1. algorithm based on wide phonetic categories let l={l1,l2,…,ln} be an ensemble of languages on which the language assessment procedure is performed, with n being the overall language number. let every language be represented by an ensemble li={l1,l2,…,lmi} of audio recordings from different speakers of this specific language, with mi being the overall number of audio recordings for a given language li. an audio recording is divided in quasi-stationary segments si(m) , each one of which having a duration of k samples, with i being the segment number of a given discourse signal: i = 1, 2, …, p, and p being the overall number of segments for a given discourse audio signal recording m=1,…,k-1. features are computed on each segment i, depending on its nature: vocal, non vocal or break    1 2i ia t s m ,i , , ..., p  (2.1) t being the operation allowing to define the type of segment. the evaluation of the segment's short-timed energy is denoted as    1 2 ik ie e s m ,i , , ..., p  (2.2) with e being the computation of the segment's short-timed energy. if the algorithm is used without reconstructing the source vocal signal waveform, then the parameters аi and ike are computed from the vocal transmission. thus series  paaaa ...,,, 21 and   pkkkk eeee ...,,, 21  are formed. if the segment is classified as a break, then ai=0. if the segment is classified as nonvocal, then ai=1. for each vocal segment the basal tone frequency (btf) is computed as   0 1 2i if f s m ,i , , ..., p  (2.3) with f being the operation of basal tone computation. afterwards the series  pffff 0...,,0,00 21 are formed. if the algorithm is used without reconstructing the source vocal signal waveform, the parameter f0i is computed from the vocal transmission. the basal tone frequency's range of variations is then divided into 5 intervals. each vocal segment is attributed a number, depending on which btf interval its frequency corresponds to  00 1 2 iuf uf ,i , , ...,f p  (2.4) f0ui being the btf level, uf the operation of btf change computation and segment encoding with numerical values  pufufufuf 0...,,0,00 21 . this allows the formation of btf values series for audio signal segments. afterwards segments during which the discourse short-timed energy increases or decreases are computed as   1 2 iu ke ue e ,i , , ..., p  (2.5) the encoding eui=(+/–)1 depends on whether the energy variation is increasing or decreasing correspondingly, ue being the operation of audio signal short-timed energy computation. the series  peueueueu ...,,, 21 are formed. if the short-timed energy decreases for a given segment, its btf value is multiplied by (-1). principal and lateral btf maxima on a segment between two breaks are used in order to algorithms for prosodic discourse feature interpretation 3 copyright ©2018 assa. adv. in systems science and appl. (2018) assess the principal and lateral accents. if the btf and short-timed energy maxima correspond in time and are maximum for a given fragment, then this segment is considered a principal maximum. if the maxima do not correspond in time, then the fragment is considered a lateral maximum ),0( euufmax i  with θ being the operation of principal and lateral maxima computation for the btf and short-timed energy series. this allows the constitution of the series  1 2 pmax max ,max ,...,max (2.6) therefore the final series  pxxxx ,...,, 21 of wide phonetic categories for a given audio record is constituted by elements xi, where 0, if break, 1, if non-vocal, 2,if interval1, 2, if interval1, 1, 3, if interval 2, 3, if interval 2, 1, 4, if interval 3, 4, if interval 3, 1, 5, if interval 4, 5, i 0 0 0 0 0 0 0 0f i i i i i i i i i i i u u u u u u u u u i u i u i f f e f f e f f a x e f f a                      interval 4, 1, 6, if interval 5, 6, if interval 5, 1, 7 0 0 lateral maximum principal maximum , if , 8, if . i i i i u u u u i i e f f e max max                                (2.7) figures 2.1 and 2.2 show the algorithm diagram that implements the discourse signal encoding. the autocorrelation function )(xr   is then computed on the series of wide phonetic categories x , with ψ being the operation of autocorrelation function computation. if the algorithm is used without reconstructing the source vocal signal waveform, the btf values are computed from the vocal transmission. if the algorithm is used with the source vocal signal waveform being reconstructed, then an algorithm for btf evaluation is required. there are numerous algorithms for btf evaluation [2]. this article presents the comparison of already implemented algorithms: the algorithm sift, based on the autocorrelation function; the algorithm amdf, based on short-time average difference function; and the algorithm for btf evaluation of the melp language encoding algorithm. table 2.1 shows the percentage values for correct btf computation p(от), erroneous assumption that a vocal fragment is non-vocal р(нв/в) and erroneous assumption that a non-vocal fragment is vocal р(в/нв). 4 m. bessonov, n. bessonova, m. farkhadov copyright ©2018 assa. adv. in systems science and appl. (2018) table 2.1. evaluation of btf correct evaluation algorithm sift amdf melp р(от), % 87±1 89±1 95±1,5 р(нв/в), % 7±1 6±1 3±0,5 р(в/нв), % 0,5 0,5 0,5 the algorithm melp was used for btf computation as its experimental evaluation proved it to be the most effective. fig. 2.1. algorithm diagram for vocal signal segments encoding algorithms for prosodic discourse feature interpretation 5 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 2.2. algorithm diagram for vocal signal segments encoding (continued) 2.2. algorithm based on basal tone melody cross-correlation functions and short-timed energy series the prosodic classification may be realized using basal tone melody cross-correlation function and short-timed energy of audio recordings. each audio recording is divided in quasi-stationary segments si(m) of k samples where i is the vocal signal number: i = 1, 2, …, p, р is the overall number of segments in a vocal signal m=1,…,k-1. for each segment i features are computed in accordance to the segment's nature: vocal, non vocal or break    1 2i ia t s m ,i , , ..., p  (2.8) t being the operation of segment type computation and the computation of the short-timed energy for one segment being    1 2 ik ie e s m ,i , , ..., p  (2.9) e being the operation of short-timed energy computation. the series  paaaa ...,,, 21 and  pekekekke ...,,, 21 are formed correspondingly. if a segment is classified as a break then ai=0. if a segment is classified as non-vocal, then ai=1. for each vocal segment the btf is computed   0 1 2i if f s m ,i , , ..., p  (2.10) f being the operation of btf computation. afterwards series  1 20 0 , 0 , ..., 0pf f f f are formed. if the algorithm is used without reconstructing the source vocal signal waveform, the 6 m. bessonov, n. bessonova, m. farkhadov copyright ©2018 assa. adv. in systems science and appl. (2018) parameter аi, eki, and f0i is computed from the vocal transmission. the cross-correlation function is computed on the btf series and the short-timed energy series ),0( ekfb  (2.11) ф being the operation of btf melody cross-correlation function computation and shorttimed energy series. the vector formed by the cross-correlation function values and series of wide phonetic categories is then given at the first layer of the neural network, that is used to infer the language group to which the presented vector corresponds. the feature computation algorithm is presented in figure 2.3. fig. 2.3. diagram of the algorithm of vocal segment encoding 3. methodical recommendations for the use of algorithms for prosodic feature interpretation algorithms as a part of the audio signal language identification methodical recommendations were developed in order to apply the aforementioned algorithm. they contain a succession of phases. phase 1. discourse dataset formation for training. the training dataset must fulfill the following conditions: if n is the overall number of languages, dm i the number of male speakers for a given language i, dm i the number of female speakers for a given language i, then vi(dm i , df i ) = vj(dm j , df j ), where i, j are the numbers of languages i, j = 1, …, n, meaning that all age groups must be represented in equal proportion among male and female speakers or that the volume of voice data for all age groups must be equal. the volume of voice data must be sufficient from a statistical standpoint in order for all the pronunciation variations to be described. the overall data volumes must be equal for every languages. step 1. reception from the source of a digital signal under the form st(fd, m, p, fr) , having the following characteristics: format "wav", sampling frequency fd = 8khz, regime m = mono, datadepth p = 16 bits, t being the audio signal number. step 2. filtering of the audio signal st(fd, m, p, fr) for unwanted noise suppression. this allows to receive the filtered signal st f (fd, m, p, fr) = ρ[st(fd, m, p, fr)], where p is the filtering operation. algorithms for prosodic discourse feature interpretation 7 copyright ©2018 assa. adv. in systems science and appl. (2018) step 3. training and control dataset formation. a set of audio signals zli{s1 f li(fd, m, p, fr), s2 f li(fd, m, p, fr), …, smi f li(fd, m, p, fr)} is formed for each language li, where mi is the overall number of audio signals for a given language li. the complete set of audio signals is then denoted by z = {zl1, zl2, …, zln}. step 4. treatment of all audio signals in every language made available by the vocal transmission. z vok = vok(z), vok where denotes the processing of the voice transmission dataset , z vok = {z vok l1, z vok l2, …, z vok ln}. step 5. audio signal parameter computation using the presented algorithms. afterwards a set of parameters z vok li mod1 = mod1(z vok li), z vok li mod2 = mod2(z vok li) is created, where mod1 and mod2 are the operation of parameter computation using the presented algorithms for prosodic parameters description. phase 2. neural network training. the neural network training operations allow the fine tuning of its parameters. neural networks with different topologies are described by different mathematical models, therefore in every specific situation a neural network is described by a specific formula. neural networks are defined for language groups. their number is equal to the number of combinations of 2 elements out of n. phase 3. neural network performance evaluation step 1. reception from the source of a digital signal under the form st(fd, m, p, fr) , having the following characteristics: format "wav", sampling frequency fd = 8khz, regime m = mono, data depth p = 16 bits, t being the audio signal number. step 2. filtering of the audio signal st(fd, m, p, fr) for unwanted noise suppression. this allows to receive the filtered signal st f (fd, m, p, fr) = ρ[st(fd, m, p, fr)], where p is the filtration operation. step 3. neural network’s testing. at the neural network's input, for each language pair li and lj, audio signals in the language i and j are given. the neural network's output gives an evaluation of the audio signal’s language identity, given t: the audio signal’s number:   f t t d rl net s fˆ ,m, p, f (3.1) step 4. evaluation of the number of correctly assessed audio signal for each language pair. this forms the vector d = (d12, d21, d13, d31, …, dn(n–1), d(n–1)n), where dij is the number of correctly assessed audio signal for a given language pair lilj, i ≠ j. step 5. hierarchical language tree construction based on the agglomerative hierarchical algorithm     k i l j min i j l x ,x ω k ω ρ ω ,ω min d xx ,    (3.2) where ωi and ωj are the languages li and lj, and ρ(ωi,ωj) is the distance between li and lj. the hierarchical language tree is the base upon which language groups are formed. 4. discourse database formation an audio signal dataset was formed in order to conduct test according to the presented methodical recommendation. its content is summarized in the table 4.1. audio records were taken from internet translation resources: television and radio, which implies that the discourse were processed using different codecs. in order to exclude the influence of the dataset's constitution on the experiment results, the number of speakers in each language was chosen equal. the overall duration of the audio signals was as well equal. the dataset was split equally in training and validation subsets. samples from the validation set were not present in the training set. for experimental purposes, all audio records both in the training and validation datasets were divided in 10seconds long fragments. the speakers' age repartition was approximated: men and women aging from 20 to 50 years old. 80% of each speaker's audio signal time was used for training, 20% for validation. the separation into training and validation was performed randomly. 8 m. bessonov, n. bessonova, m. farkhadov copyright ©2018 assa. adv. in systems science and appl. (2018) table 4.1. contents of the dataset used for prosodic feature models experimental validation language number of speakers overall audio signal duration for each speaker, min speaker sex (m-male, f-female) repartition in the training/testing sets, % chinese 10 100 5m/5f 80/20 english 10 100 5m/5f 80/20 finnish 10 100 5m/5f 80/20 french 10 100 5m/5f 80/20 german 10 100 5m/5f 80/20 japanese 10 100 5m/5f 80/20 farsi 10 100 5m/5f 80/20 portuguese 10 100 5m/5f 80/20 russian 10 100 5m/5f 80/20 spanish 10 100 5m/5f 80/20 5. neural network definition and tuning pattern recognition tasks are in most cases solved using statistical methods. however in the case of vocal data in different languages, it proves difficult to build the repartition function of the considered parameters. therefore in the present article neural networks were used for vocal segments classification. as stated in literature, for classification-type problems, the number of neurons in the network's first layer is equal to the number of elements in the feature vector that is given at the entrance [3]. the number of output neurons depends on the type of problem and the output number interpretation rules [3]. the number of neuron in the intermediate layers is given by the formula [3]     2 1 1 1 y p p w y x y y xp n n n n n n n n nlog n             (3.3) where ny is the neural network's (nn) output vector's number of element, np is the number of elements in the test dataset, nх is the number of element in the input vector and nw is the overall number of neurons. the choice of the nn's class and architecture is a non-trivial problem for which exact solutions do not exist [3]. in order to choose the number of neurons, one can highlight two methods: the more neurons, the more reliable the network will be and the more neurons, the worse the network will approximate the transfer function. the neural networks were implemented in matlab, using the environment's built-in functions. the following architectures were experimentally evaluated: kohonen maps, cascadeforward nn, elman networks, multilayer perceptron, hopefield networks, probabilistic networks, networks with radial basis functions (rbf), counter-propagation network with learning vector quantization. the networks were trained with built-in matlab functions [1]: quasi-newton algorithm, levenberg-marquardt algorithm with bayes regularization, fletcher-reeves conjugate gradient method, polak-ribiére conjugate gradient method, powell-beale conjugate gradient method, gradient descent, gradient descent with variable learning rate, levenberg-marquards algorithm, scaled conjugate gradient method, gradient descent with momentum, gradient descent with momentum and variable learning rate, one step secant method, random increment method and elastic error backpropagation algorithm. for the first phase, in order to build the limited groups of 10 languages, experiments were performed with each separated language pair, accounting for 45 neural networks in total. the best results have been obtained using a multilayer perceptron. therefore this architecture has been selected for fine tuning. since the language given at the nn's entrance is a priori unknown, it was decided to use a algorithms for prosodic discourse feature interpretation 9 copyright ©2018 assa. adv. in systems science and appl. (2018) unified architecture for each language pair. 6. evaluation of the wide phonetic categories algorithm according to the formula and the starting conditions for nn testing ny = 2, np = 600, the number of neurons in the hidden layers is then 117 ≤ nw ≤ 2015 for the wide phonetic categories model. since nw is between 117 and 2015, at the moment of the nn architecture definition the number of neuron was chosen from 100 to 2000, correspondingly to the number of neurons from 1 (one layer from 100 to 2000 neurons) to 20 (20 layers of 100 neurons) in the following configurations: from 100 to 1000 neurons with a 10 neuron step in a given layer or from 1000 to 2000 with a 100 neuron step. the maximal number of neurons in one given layer was 800. in order to build the different multilayer perceptron architectures for the 45 language pairs, a vector d=(d1,2, d2,1, d1,3, d3,1, di,j, dj,i, dn,n–1, dn–1,n) of goal indicators for assessment confidence was built, with n being the overall number of languages in the automatic language assessment system. therefore the length of the vector is d=90. each element di,j, dj,i = 100. the vector dk=(d k 1,2, d k 2,1, d k 1,3, d k 3,1, d k i,j, d k j,i, d k n,n–1, d k n–1,n) of goal indicators for assessment confidence for the current architecture has as well 90 elements. the distance between d and dk is defined as       2 2 2 1 2 1 2 2 1 2 1 k k k r , , , , i , j i , jd d d d d d d             2 2 2 i i 1 1 1 1 k k k j , j , n ,n n ,n n ,n n ,n... d d d d d d         (3.3) thus, the lower the distance dr, the better the nn tuning. at the end it was found that dr lays in the interval from 59.1861 to 532.4106. the best value, dr = 72.5358, was obtained for a nn with 1400 neurons overall, organized in one layer of 800 neurons and 2 layers of 600 neurons. the results of language assessment are presented in the table 6.1. table 6.1. average confidence values for language identification chinese english finnish french german japanese persian portugu ese russian spanish chinese 94.5 95.1 96.2 95.9 97.5 96.6 95.2 94.4 97.9 english 93.8 97.4 92.8 93.8 93.6 98.1 94.5 94.0 97.8 finnish 93.8 93.7 93.2 93.4 93.9 93.9 96.1 93.7 94.3 french 94.2 93.6 93.2 93.9 93.4 94.0 94.8 93.8 94.4 german 94.5 92.6 93.7 92.5 94.6 94.0 97.5 96.3 93.9 japanese 83.6 94.1 74.0 98.3 93.3 94.0 84.9 94.4 98.0 persian 84.4 94.0 74.6 93.3 93.8 83.6 92.7 84.3 93.2 portuguese 94.2 93.6 93.5 93.9 94.2 94.5 93.5 93.9 98.4 russian 94.4 95.1 94.1 95.3 93.4 94.0 94.4 94.3 94.5 spanish 93.9 94.3 93.4 93.2 94.2 93.8 94.1 94.5 93.2 the languages were used to form groups using the agglomerative algorithm. language pairs were used in quality of patterns to be recognized. the average confidence value for fixed first and second order error rates was used as a measurement of the distance between the two languages of one pair. the distance between classes is defined by the distance to the nearest neighbor:     k i l j min i j l x ω ,x ω kρ ω ,ω min d xx ,    (3.5) 10 m. bessonov, n. bessonova, m. farkhadov copyright ©2018 assa. adv. in systems science and appl. (2018) where ωi,ωj are the languages li and lj, and ρ(ωi,ωj) is the distance between li and lj. this allows the creation of a graph of hierarchical classification. this graph can then be used to assess groups of languages close to each other. 7. evaluation of the basal tone melody cross-correlation function and short-timed energy series algorithm according to the formula and the nn test starting conditions, ny = 2, np = 600, nx = 797, the number of neurons in the hidden layer is 117 ≤ nw ≤ 2806 for the cross-correlation function and btf and short-timed energy series model. since nw is comprised between 117 and 2086, at the time of the nn architecture definition the number of neurons in a layer was varied from 100 to 3000, correspondingly layers having from 1 (1 layer from 100 to 3000 neurons) to 20 (30 layers of 100 neurons each). the following configurations were used: from 100 to 1000 neurons with a 10 neuron step by layer, from 1000 to 3000 with a 100 neuron step. the maximum number of neurons was 800, dr = 89.1449. the results of language identification are presented in the table 7.1. table 7.1. average confidence values for language chinese english finnish french german japanese persian portuguese russian spanish chinese 97.7 94.7 92.8 97.8 97.9 93.8 91.7 93.1 92.1 english 91.2 91.4 92.3 92.9 94.8 92.7 97.7 90.3 92.0 finnish 90.9 91.5 95.8 94.7 94.6 95.4 90.9 93.6 95.9 french 92.1 92.9 92.4 93.9 96.7 97.5 92.1 91.8 91.8 german 92.5 90.2 91.4 90.4 91.8 92.2 92.4 93.0 95.4 japanese 80.6 91.8 90.1 82.3 71.9 90.7 90.5 94.7 97.2 persian 71.1 91.5 82.3 91.6 82.6 78.2 97.5 92.5 91.5 portuguese 90.7 91.0 92.0 92.0 93.2 93.4 92.1 94.5 92.5 russian 91.0 91.7 90.6 92.6 92.3 92.4 91.7 91.6 96.2 spanish 90.5 92.9 90.9 92.8 91.2 93.1 91.4 92.1 93.6 8. conclusion the algorithms presented in the article aim at a complex description of discourse prosodic feature for their usage in special data processing tasks, in particular audio signal language assessment. as seen in the presented tables, prosodic feature description using wide phonetic categories allows for high-confidence language identification. however this performance insignificantly surpasses cross-correlation function. distance indicators for current results of language assessment with respect to the goal indicator dr scored at dr = 72.5358 for the autocorrelation model from wide phonetic categories and dr = 89.1449 for the signal's crosscorrelation function model from basal tone value and short-timed energy series. the presented algorithms demark themselves from others in that they are used for audio signal language identification, after vocal transmission, but without reconstructing the source vocal signal waveform. references 1. diakonov v. & kruglov v. (2006) matlab 6.5 sp1/7/7 sp1/7 sp2 + simulink 5/6.instrumenti iskusstvennogo intellekta i bioinformatiki [matlab 6.5 sp1/7/7 sp1/7 algorithms for prosodic discourse feature interpretation 11 copyright ©2018 assa. adv. in systems science and appl. (2018) sp2 + simulink 5/6. tools of artificial intelligence and bioinformatics]. moscow, russia: solon-press [in russian]. 2. imamverdiev y. & suhostat l. (2014) podhodi dlia ozenki perioda osnovnogo tona rechevogo signala v zashumlennoy srede [the approaches for pitch evaluation in noisy environment]. speech technology., 4: 84-102. 3. komartsova l. & maksimov a. (2004) neirokimpiuteri: uchebnoe posobie [neurocomputers: tutorial]. moscow, russia: bmstu [in russian]. 4. miloshenko a.a. (2010) razrabotka metodiki ispolzovania shirokih foneticheskih kategoriy v zadachah verifikazii diktora [development of method of using broad phonetic categories in speaker identification task]. moscow, russia: miit [in russian]. 5. ambikairajah e., li h., wang l. & yin b. (2011) language identification: a tutorial. ieee circuits and systems magazine. 11(2), 82–108. 6. bhattacharjee u. & sarmah k. (2013) language identification system using mfcc and prosodic features. int. conference on intelligent systems and signal processing (issp), gujarat, 194–197. 7. lee r., leung c.-c. & ma b. (2013) spoken language recognition with prosodic features. ieee trans. on audio, speech, and language processing. 21(9), 1841–1853. 8. martinez d., jeida e. & ortega a. (2013) prosodic features and formant modeling for an ivector-based language recognition system. ieee int. conference on acoustics, speech and signal processing (icassp), vancouver, canada, 6847–6851. 9. martínez d., burget l., ferrer l. & scheffer n. (2012) ivector-based prosodic system for language identification. ieee int. conference on acoustics, speech and signal processing (icassp), kyoto, japan, 4861–4864. adv syst sci appl 2019; 02; 101-119 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/673 modeling the management of the economies of developing countries† s. bayzakov*1, jeffrey yi-lin forrest3, n.a. baizakov2 1) institute of mathematics and mathematical modelling, almaty,kazakhstan; e-mail: baizakov37@mail.ru; 2) institute of mathematics and mathematical modelling, almaty,kazakhstan; e-mail:nauryz.b@mail.ru; 3) school of business, slippery rock university, slippery rock, pa 16057, usa; e-mail: jeffrey.forrest@sru.edu received november 28, 2018; revised march 4, 2019; published july 10, 2019 abstract: in the present, the issues of algorithmization of the tasks of sustainable development of the economies of developing countries of the world are considered. the final result of the work is represented by a system of models for analyzing the development of the country's economy and the development of the well-being of its population, which is compared with fao models to ensure the productive security of the countries of the world proposed by the un world bank, as well as the oecd management criteria system. the basis for constructing the proposed system of models was actually the achieved level of human capital development. as a result, a new base has been created for assessing the contribution of the country's scientific and technological potential to the development of the productive power of labor, which is determined by the difference between labor productivity by income and labor productivity by the total labor input. and the total financial and economic capacity of a country is measured by the product of the number of man-hours spent in a calendar year by the achieved level of growth of the productive power of labor in the country's economy. keywords: potential, productivity, strength, labor, income, fao , model, technology. 1. a new eco-economic system of commodity-money management strategic planning for the development of market forces for developing countries relies on the old interpretation of the direct costing system. [1] thus, the process of determining net present value (npv) is based on the direct costing system, which was designated for accounting direct production expenditure by american scholar d. harris. according to this system, direct production expenditure is divided into two parts: that which is fixed and that which is variable. [2] in kazakhstan, with the help of the system developed here fixed expenditure includes expenses related to administrative and management activities aimed at the sale of output products. here, market research, commercial and general-administrative expenses are † this study was financially supported by grant no. ap05131044 from the scientific and technical programs and projects of the ministry of education and science of the republic of kazakhstan. * corresponding author: baizakov37@mail.ru mailto:baizakov37@mail.ru mailto:nauryz.b@mail.ru mailto:jeffrey.forrest@sru.edu modeling the management of the economies of developing countries 102 copyright ©2019 assa. adv. in systems science and appl. (2019) included. the variable expenditure includes expenses related to changes in the output volume of production: materials, energy and fuel, salaries for employees and engineers. as a result, the sum of direct and fixed costs represents the prime operating costs of the output. the amount of prime operating costs and depreciation of fixed assets of the enterprise, including taxes, determine the total key cost of production. as a result, the cash flow in terms of enterprise development years is calculated according to the formula: 𝑁𝑃𝑉 = 𝑅𝐸𝑉 − (𝑂𝑃𝐸𝑅 + ∑ 𝑇𝐴𝑋 + ∑ 𝐶𝑅), (1) where 𝑅𝐸𝑉 stands for the revenue from the sale of commodity ore in the reference year; 𝑂𝑃𝐸𝑅 the operating costs (filled by the prime cost) in the reference year; ∑ 𝑇𝐴𝑋 the tax paid in the reference year; and ∑ 𝐶𝑅 the credit repayment in the reference year (including the interest on credit). hence, the net present value 𝑁𝑃𝑉о(𝑖) is calculated according to the next formula: 𝑁𝑃𝑉о(𝑖) = 𝑁𝑃𝑉(𝑖) (1+ii1)∗(1+ii2)∗(1+ii3)∗…∗(1+iiн+1) , (2) where 𝑁𝑃𝑉(𝑖) stands for the cash flow of the reporting year i in million dollars, ii𝑖 the inflation index of year 𝑖 in %. finally, the present value, considering the discount coefficient, is calculated according to the formula: 𝑁𝑃𝑉(𝑖) = 𝑁𝑃𝑉𝑂1 ∙ (1 + 𝑟)1 + 𝑁𝑃𝑉𝑂2 ∙ (1 + 𝑟)2 + ⋯ + 𝑁𝑃𝑉𝑂𝑛 ∙ (1 + 𝑟)𝑛, (3) where r is the discount rate (money devaluation rate), indices 1, 2, 3, ..., n the discount intervals: years, quarters, months, days. in real-life applications, accounting and analysis based on the direct costing system revealed great opportunities to set tight connections between production volumes, gross and net profits, prime costs and gross revenues. they allowed practical use in the works of mathematical economists, production function methods, econometrics and regression analysis. however, any initiative becomes “a dead dogma”, if it does not continuously develop and does not find new arguments for the development of its viability. this point of view, focused on the continuous development of human capital, objectively determines the strict correspondence of the methods of market relations management with the level of development of labor and capital productive forces in the market-oriented economy of any country in the world. hence, in determining the break-even point of development according to the direct costing system, variable costs include salaries for workers and engineers, and fixed costs include salary costs for personnel employed in marketing, general and administrative divisions of an enterprise. it means that this system, which was developed in the 1940s under the conditions of tough competition between the economies of the soviet union and other developed countries of the world, did not take into account the development of human capital. such theoretical deficit hampers opportunities for the development of the working man himself, his creative, scientific and technological potential: the rich grow richer, the poor become poorer. in general, the direct costing system specifically limits the creative initiatives of employees' participation in the production of goods and services. it determines employment, 103 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) income and pricing indicators by including in the fixed and variable costs of production, the remunerations of people employed by an enterprise in the economy. however, the main drawback of this system is that it contradicts the macroeconomic approach of the quantitative theory of money, which describes the development of the market forces of the modern world economy and the globalization. that is, there is an objective need for us to develop management methods appropriate for the ecological and economic systems of developing countries of the world. such methods need to be adequate for the current level of forces underneath the productivity development. they need to ensure a mutual correspondence between monetarism macroeconomic indicators and the microeconomic indicators that measure the use efficiency of the local ecological and economic resources. in addition, the research results of the post-harris period, especially those related to the development of input-output tables by l. kantorovich and t. ch. koopmans, nobel laureates in 1975 [3], as well as the fundamental work of michael porter “competition” in 2001 [4], allow for consistency and assess the competitiveness of firms and enterprises on the productivity of local environmental and economic resources. both these works are unidirectional, are consistent with milton friedman’s [5] quantitative theory of money, and ensure that microeconomic indicators are comparable with macroeconomic ones. additionally, more related works should be added. for example, u. nordhaus and p. romer, who made a step back towards accounting of the use effectiveness of local ecological and economic resources, initiate an evaluation of scientific achievements[6]‡. in recent years, thanks to the initiatives of astana economic forum, large-scale global researches have been carried out in kazakhstan. among them, the works of s. baizakov, ye. utembayev, n. akimov, n. sagadiyev, m. amrenov, and k. berentayev [7] are considered important. these works are supported and responded by analysts from the developed western countries. ragnar bentzel was right in his presentation of the works of kantorovich and koopmans, as a representative of the royal swedish academy of sciences, by noting the obvious that “the main economic problems can be studied in a purely scientific sense, regardless of the political organization of the society in which they are studied” [3]. in general, new horizons are opened up for the development of the scientific and technological bases that will be useful for recording and analyzing production, employment, incomes and prices. therefore, the time is ripe for us to overcome the shortcomings of the direct-costing system in terms of the development of human capital and the initiative of the busiest person in the economy. to this end, firstly, it is necessary to ensure that the methods developed for managing the evolution of human capital, as a key component within the advance of the productive forces of developing countries of the world, are in line with the conditions for the development of human capital in the developed countries. secondly, it is recommended that in any strategic perspective the objectives of an operational planning are consistent with the long-term, sustainable development goals. thirdly, there is a need to identify key operators that allow the integration of commodity and financial markets through the evolutionary development of keynesian and monetarist theories. and fourthly, there is a need for a new multiplier developed for scientific and technological progress so that capital development models in their commodity and monetary forms can be integrated into the system of market relations. ‡ the swedish academy, in its message on the laureates, explains that u. nordhaus and p. romer created economic models which explain how the merket economy interacts with nature and knowledge. thus, p. romer in his works showed how knowledge could serve as an engine for long-term economic growth. while the works of u. nordhaus describe the interaction between society and nature. modeling the management of the economies of developing countries 104 copyright ©2019 assa. adv. in systems science and appl. (2019) 2. setting the task of digitization of sustainable development goals and the main operators of the algorithm for its implementation at the last astana forum (astana, june 2017), president of the state n.a. nazarbayev stated “the need for new methods for calculating indicators that measure the country's wealth and the well-being of citizens, as the frequently used indicator of gdp has a number of significant flaws. gdp does not reflect the long-term nature of economic activity, does not take into account damage to the environment, including the depletion of natural resources. in addition, it does not reflect the quality of life in a particular country, gdp per capita does not show the real well-being of citizens, does not consider the income stratification of the population. what i say is, the world community can adopt an updated methodology for calculating gdp on the basis of “green” gdp and such indices as the human development index, the oecd better life index. it should adequately reflect the need for a balanced development of countries” [8]. back in 2009, the president of the state wrote about this in the articles “keys to the crisis” and “the fifth way” [9-10], and in his address to the people of kazakhstan in 2018, he indicated specific ways to implement his project on a new model of economic growth [11]. algorithmizing the task of sustainable development of developing countries at the economic research institute is carried out by a limited number of operators. it allows policy makers to effectively manage the mutually agreed movements of indicators of economic and financial markets up to the implementation of their goals. setting up the task of achieving the goals of sustainable development and the operator tool for its algorithmization are aimed at implementing a transfer of national moneys of developing countries to their dollar equivalents. it uses, on a parity basis, the whole arsenal of analysis and forecasting methodology developed by international organizations, such as un and oecd. such a formulation of the task and the chosen tool for its implementation, according to the developers, proceeds from the objective need of ensuring parity of the national moneys of developing countries with the national currencies of developed countries that effectively interact with the world reserve currency, represented by the u.s. dollar. the main tool chosen for solving this problem allows developing countries to ensure the commensurability of their moneys with their dollar equivalents. this warrants the alignment of the development methodology of the managerial economics of all countries in the world. that is, each country on a parity basis can use innovative technologies proposed by international organizations and other tools for the effective and sustainable management of its national economies and finances. at the same time, a condition is created for the mutual coordination of the economic interests of local populations, businesses and regions of each country based on the development of human capital and effective use of local ecological and economic resources. in such a case, high rates of economic growth and intended goals of increasing the well-being of citizens are ensured in the most efficient way. 3.key operators that integrate product and financial markets through developing theories of keynesianism and monetarism our proposed operators are based on the qualitative theory of money developed by s. baizakov [12]. this theory was built by the evolutionary development of economic laws, identified by supporters of the macroeconomic theory of j. keynes [13] and the quantity theory of money of m. friedman [5]. and the corresponding system of models of the qualitative theory of money 105 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) is built on the basis of the conceptual idea of the president of kazakhstan n.a. nazarbayev, set out in “the fifth way”[10]. the algorithm for solving the problem of the sustainable economic development of developing countries ensures the commensurability of national moneys of developing countries with those of developed countries by establishing their true dollar equivalences. this algorithm consists of the following operators, accompanied by the functions they perform in accordance with their economic purposes. the microeconomic operator for determining the multiplier of the productivity of resource costs at the enterprise level is 𝜇 = 𝑁𝐺𝐷𝑃/𝑄𝑃, (4) where 𝑁𝐺𝐷𝑃 is nominal gdp, 𝑄𝑃 the local ecological and economic resources consumed for the production of nominal gdp. the macroeconomic operator for determining the multiplier of scientific and technological progress at the country level is 𝑐 = 𝜇/(1 + 𝜇). (5) the macroeconomic operator for determining the gdp deflator (inflation multiplier) is 𝑝𝑏. according to the monetary policy methodology, 𝑝𝑏 = 𝑁𝐺𝐷𝑃/𝑅𝐺𝐷𝑃, (6) where 𝑅𝐺𝐷𝑃 is real gdp determined by the same monetarism methodology. the microeconomic operator for determining the value multiplier of the u.s. dollar (𝑝𝑝) as a reserve currency is 𝑝𝑝 = 𝑐 𝑝𝑏⁄ . (7) the macroeconomic operator for determining the purchasing power of the national money, which is expressed as a multiplier of the dollar equivalent of the aggregate of the true prices of goods and services of the developing country is 𝑝𝑐 = 1 𝑝𝑝⁄ . (8) the inter-sectoral operator for determining the true price of each product or service in a developing country is given by calculating the savings in working time, determined by the labor equivalent 𝑡𝑖, the direct costs of local ecological and economic resources for producing a product 𝑖, and labor equivalents 𝑇𝑗, and total production costs of the same product. the macroeconomic operator for determining the goods and services actually consumed in a developing country for consumption and accumulation in physical units is calculated on the dollar equivalent of their money 𝐹𝐺𝐷𝑃 = 𝑐 × 𝑅𝐺𝐷𝑃, (9) where according to the current monetarism model 𝑅𝐺𝐷𝑃 = 𝑁𝐺𝐷𝑃/𝑝𝑏. the macroeconomic operator for determining the goods and services actually consumed in a developing countries for consumption and accumulation is calculated in national money in their dollar equivalent 𝐹𝐺𝐷𝑃 = 𝑝𝑝 × 𝑁𝐺𝐷𝑃, (10) where according to the current model of monetarism 𝑁𝐺𝐷𝑃 = 1 𝑝𝑝 × 𝑀, the mass of national money in circulation. modeling the management of the economies of developing countries 106 copyright ©2019 assa. adv. in systems science and appl. (2019) 4. system of operators of developing countries as a tool for further liberalization of the world economy based on the above system of operators, a developing country can fully establish the dollar equivalent of its money. as a result, it can apply the existing international experience and positive practices of other countries on a par with the developed countries in managing its economy. moreover, to analyze and forecast the development of any developing country, such international organizations as the un and the oecd countries can use analogues to assess key indicators determined by the multipliers of scientific and technological progress and the productivity of local ecological and economic resources. thus, the first operator for determining the productivity of local ecological and economic resources was used by representatives of the un’s world bank and consultants working in kazakhstan through the tacis of the economic commission for europe in 1999-2004. the ginsim models (united kingdom) for analyzing the economics of enterprises were particularly valuable in kazakhstan [14]. the results of these works were published in kazakhstan in 2002 under the scientific editorship of s. baizakov, and are used in the practical work of the economic research institute [15]. according to the results of research presented by kateryna schroeder (world bank), the world bank is still currently using the multiplier of the productivity of local ecological and economic resources. so, according to k. schroeder, this multiplier is developed in its original form by assessing the effectiveness of domestic spending on the resources of developing countries in central asia when exporting agricultural products to china (drc) [16]: 𝐷𝑅𝐶𝑖𝑗 = 𝑐𝑖𝑗 𝑑 (𝑝𝑖𝑗−𝑐 𝑖𝑗 𝑓 ) , (11) where 𝑐𝑖𝑗 𝑑 and 𝑐𝑖𝑗 𝑓 represent, respectively, the domestic and foreign costs of input products “𝑖” in the country. in the case of export of meat, for example, from the countries of central asia to a specific country (china), if the multiplier 𝐷𝑅𝐶𝑖𝑗 < 1, this exporting country “𝑖” has a “comparative advantage” in the production of good “𝑗” compared to china; that is, its good (meat) costs less than the same good in china (in the production of one dollar). accordingly, the smaller the indicator 𝐷𝑅𝐶𝑖𝑗 in central asian countries, the greater the comparative advantage of these countries will be. as can be seen, the constant for evaluating the efficiency of exports of developed countries is the dollar, not the dollar equivalent of the national moneys of developing countries. in the economies of developing countries, not all goods and services aim at fulfilling the purpose of export. and, not all of their goods consumed domestically are imported from other countries. for example, the entire consumption fund and gross accumulation of these countries turn around by means of national moneys, which are the same goods as other goods – in our case, the tenge (not an international reserve currency) is itself a commodity. this means that national moneys of developing countries become lightly liquid commodities and determine the value of the nominal volume of production of the final product. this nominal volume is the nominal gross domestic product (ngdp) in the current prices of national money, determined in terms of the dollar. at the same time, its physical volume, that is, the real volume of the final product, called the real gross domestic product (rgdp), is determined by means of the gdp deflator (inflation). according to the common practice of the world, in analyzing the development of economies of developing countries, world bank experts use this indicator of real gdp (rgdp). however, real gdp (rgdp) is not able to assess continuously the productivity of 107 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) local ecological and economic resources used to produce the nominal gdp (ngdp). without such an assessment, as shown by the multiplier of k. schroeder, a representative of the world bank, it is not possible to determine the effectiveness or impact of any innovative solution. so, the reciprocal of the multiplier of k. schroeder, represents the value (productivity) of a country's local ecological and economic resources used to produce one-dollar equivalent of this commodity (1/𝐷𝑅𝐶𝑖𝑗): 1/𝐷𝑅𝐶𝑖𝑗 = ( 𝑝𝑖𝑗−𝑐𝑖𝑗 𝑓 ) (𝑐𝑖𝑗 𝑑 ) = 𝑝𝑖𝑗 𝑐𝑖𝑗 𝑑 − 𝑐𝑖𝑗 𝑓 𝑐𝑖𝑗 𝑑 . (12) this indicator 1/𝐷𝑅𝐶𝑖𝑗 is in fact an indicator “𝜇” – that is, an indicator of the productivity of ecological and economic resources in the regions of this country. for this reason, it suffices, for example, to replace in the last multiplier 𝑐𝑖𝑗 𝑓 𝑡𝑜 𝑐𝑖𝑗 𝑑 . as a result of this replacement, we get an analogue of the multiplier "𝜇" for any country in the world that can be employed to conduct a comparative assessment of the effectiveness of its exports: 𝜇 = ( 𝑝𝑖𝑗−𝑐𝑖𝑗 𝑓 ) (𝑐𝑖𝑗 𝑑 ) = 𝑝𝑖𝑗−𝑐𝑖𝑗 𝑑 𝑐𝑖𝑗 𝑑 . (13) the above model of the multiplier of the un representative k. schroeder is a very valuable tool for analysis, as it expresses the production and economic relationship between producers of goods and ecological and economic factors of its production. the most important thing about it is that it reflects the relationship between microeconomic indicators and macroeconomic indicators. moreover, this model has been used in assessing the comparative advantages of more than 30 commodity positions of enterprises in kazakhstan. this work had been done by a group of kazakhstan experts under the methodological guidance of european expert jean michel, now an employee of the world bank of the un. the results of this work are also presented in the book [14, p. 41-78]. along with j. michel, the methodological basis of s. baizakov’s model can be seen in the ginsim model (uk), the software product of which was developed by the maxwel stamp company [ibid., p. 111-144]. in particular, the nominal rate of protection (nrp), determined by using both the schroeder multiplier and the ginsim project model, shows how much the price of local goods (pd) increases compared to the world price (pw) for the same product due to establishing custom duties [ibid, p. 133]: 𝑁𝑅𝑃 = [(𝑃𝑑 − 𝑃𝑤)/𝑃𝑤] × 100. (14) despite the possibility of integrating microeconomic indicators with macroeconomic ones, individual representatives of the world bank, when analyzing the sustainability of transition countries, limit themselves to assessing fluctuations of real gdp without considering studies, in particular, k. schroeder’s work on the cost intensity of the obtained final product or the productivity of ecological and economic resources. to this end, in table 1, the analysis of kazakhstan’s data was conducted by non-authors of the article having cited this table, the authors refer to the work of the representative of the world bank in kazakhstan, who studied the state of kazakhstan's economy in 2016. footnote to table 1 indicates this source. modeling the management of the economies of developing countries 108 copyright ©2019 assa. adv. in systems science and appl. (2019) table 1. contribution to real gdp growth, 2013-2016 (in percentage points, unless otherwise indicated) 2013 2014 2015 2016 estimate real gdp growth (in percent) 6.0 4.2 1.2 1.0 domestic demand 6.9 3.8 2.7 1.3 private consumption 5.1 0.7 1.0 -0.3 public consumption 0.2 1.0 0.3 0.3 gross capital formation 1.6 2.1 1.5 1.3 net export -1.0 0.1 -1.2 -0.2 exports of goods and services 1.1 -1.1 -1.2 -1.0 imports of goods and services -2.1 1.1 0.0 0.8 statistical discrepancy 0.1 0.3 -0.3 -0.1 source: world bank calculations based on the data published by the committee on statistics (some figures may not be accurate due to rounding) table 1 reflects only the abstract contribution of all participants, first of all, the real and monetary and financial sectors of the economy of kazakhstan, without detailing the contribution of each of them. that is, there is a need for such details on the basis of identifying the causes. according to w. leibniz, whose name was used in the name of the famous institute of hannover, “reality can be understood for its reasons”. therefore, the establishment of causes will be the basis for making new managerial decisions that improve the health of the economy. the second operator, according to s. baizakov, seems to be a multiplier of scientific and technological progress. it is essentially macroeconomic; and by its nature, scientific and technological progress is the main engine of innovation in managing not only the development of the real and monetary and financial sectors, but also in the managerial economic sector [17]. in the western economic literature, the stp multiplier is known by the name of “tfp” only as an indicator of the level of technological progress. according to k. marx, “the first most important of the innate properties of matter is motion, not only as a mechanical and mathematical motion, but even more as aspiration, life spirit, tension or, using the expression of jacob boehme, suffering [qual] of matter” [18]. the driving force – the vital spirit of an economy – is any desire of the three aforementioned sectors of the economy. it is called scientific and technological progress. however, the most commonly used models of econometrics and the concept of production functions do not fully reflect the scientific and technological progress (stp). such progress is scientific and technological because technology is being improved not only in the real sector of economy, but also in all three managerial sectors: real, financial and managerial. a piece of research work is considered fruitful only if all three sectors as the ten-year work of the astana economic forum (aef) showed, such is, for example, the concept of the “fifth way” of the president of kazakhstan n. a. nazarbayev [9-10], which, by the level of debatability, relates to world problems. to assess the applicability of this concept of a three-economy, you can refer to the work of the frenchman lionel guy stoléru, which serves as its counterpart and is connected with a specific example of cars, where the market value of the car production increases 90 times, the price of one car increases 3 times, and the overall index prices increased 5 times. then the final results of the three prices differ so much that one cannot ignore these differences and not recognize their objectivity[19]: 1. the growth index of the nominal value of products 90 00 1  jj jtjt qp qp i ; 109 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) 2. product growth index in physical terms 30 3 90 000 0 2  j jt jj jtj q q qp qp i ; 3. price increase index in constant french francs at relative value: 18 5 3 *30 00 3  jj jtjt q q i   , where jtp and jtq accordingly, the price and quantity of goods j during the period t , and j all the many benefits, the dynamics of which are studied in time 0 before t . t good j relative price. indeed, l. stoléru is right, it would be a gross mistake to measure and assess the level of “real” growth on the basis of two, nominal and real growth indices. it is the third index of balanced economic growth, which expresses the inter-branch nature of the sphere of production, and is the carrier of feedback between the sectors of the real economy. thanks to the third index of assessing the real value of goods and services, it will be possible to assess the gaps in the development of the real and financial sectors of the national economy. in fact, this third index is the main key to relieve economic and financial shocks. the book by p. sraffi “production of goods through goods” (1999) also explored the triple connection between individual product prices and market prices for goods and services. to explore the relationship between the two, sraffa introduces a new measure of the value “standard output” and “standard product” [20]. thus, taking the produced annual cumulative product for a standard equal to one, it determines the share of wages in the value of the “standard output”: c+v+m=1. (15) and, having accepted the produced national income as a standard product equal to one, he studies in it the ratio of wages and profits: v+m=1. (16) the first of these expresses the basis of the cost method, defined by the labor theory of value (x = c + v + m), and the second, the income method, determined by the quantity theory of money (y = v + m). the cost method, except for wages (v) and profits (m), takes into account the total costs of labor, representing the real costs of production (c). another standard in his theory is determined by equality according to work: la + lb + ... + lk= 1, (17) where la, lb, ...., lk the annual amount of labor, respectively, employed in industries producing products a, b, ..., k important is the result of the work of sraffa in determining the productivity of intermediate consumption goods by dividing the annual income generated (y = v + m) by the total labor input used for its production (x = c + v + m). in this context, alisher tleubayev, a doctoral candidate at the leibniz institute of agricultural development in transition economies (germany), presented the denominator of the tfp model, not in the former two-factor form of the production function, but in a qualitatively new form, namely in the form of the stp multiplier [21]: 𝑇𝐹𝑃 = 𝑌 𝐴(𝑡)×𝐾𝛼×𝐿𝛽×𝑀𝛾. (18) modeling the management of the economies of developing countries 110 copyright ©2019 assa. adv. in systems science and appl. (2019) this multiplier of technological progress (in his other article, jointly with i. bobojonov, one of the leaders of the leibniz institute [21]), in fact, is the multiplier of scientific and technological progress, given in the second operator of the qualitative theory of money of s. baizakov. now, the multiplier of scientific and technological progress, which is supported by such an authoritative institute named after w. leibniz (represented as the tfp model in terms of the interpretation of tleubayev-bobojonov), confirms the emergence of s. baizakov's new qualitative theory of money. in fact, this new qualitative theory of money has been developed through the evolutionary development of the quantitative theory of money of m. friedman, that is, by complementing it instead of denying it. thus, in calculations and theoretical studies of bobojonov and tleubayev, the numerator of the tfp technological progress indicator is expressed by the “𝑌” indicator, which is the nominal gdp of the second baizakov operator (𝑌 = 𝑁𝐺𝐷𝑃), and the denominator of the tfp indicator is determined not by two, but by a three-factor production function in the approach of bobojonov-tleubayev. in this case, the information base for the forecast of the tfp indicator itself is the dynamics of the relationship of the final product 𝑌 to the total cost of ecological and economic resources 𝑋, and the integrator of three factors (𝐿-labor, 𝐾-capital and 𝑀current cost of local ecological and economic resources). in general, the tfp technological progress indicator can be predicted based on the release of the system of national accounts. respectively, parameters 𝛼, 𝛽 and 𝛾 represent the elasticities of the above three factors. thus, the bobojonov-tleubayev’s model can be used in predicting the system of the un sustainable development goals [21]. the advantage of their model is that technological progress is expressed by the production attitude not of two factors, but of three factors: labor, capital and natural matter. due to the lack of such a three-factor approach, which is systemic, kazakhstan and other developing countries suffer a shortage of analytical and forecasting tools for long-term forecasting. as far as short-term forecasting is concerned, kazakhstan’s interpretation of the multipliers of scientific and technological progress was previously described in the article by baizakov and others (2014) [22]. unlike the multiplier 𝑇𝐹𝑃, by baizakov and others forecast model (2014) is designed for short-term planning of up to three-five years. and it is implemented according to the reported data of the statistical authorities of each country. therefore, for the model of forecasting the economic performance for the long-term perspective up to 2025–2050, it is proposed, mainly, to be based on the dynamics of the stp multiplier and other related progress multipliers. in general, this baizakov’s and others forecast model, as well as s. baizakov’s multiplier of scientific and technological progress, are being developed in the name of the bobojonovtleubayev’s (germany) long-term model and other models of developed countries and international organizations such as the human development index and oecd indicators. 5. stp multiplier: the key to integrate development patterns of capital in their commodity and monetary forms into a system of market relations as is known, according the definition of the value of a dollar using a macroeconomic textbook (reprinted in russia more than 13 times), “value of money is the amount of goods and services that can be exchanged per unit of money (dollar); and the purchasing power of a monetary unit is the reciprocal of the price level” [21]. this means that the purchasing power of the us dollar (pp), as the price of a reserve currency is determined by the formula: 111 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) 𝑝𝑝 = 𝑐 𝑝𝑏 , (19) where 𝑝𝑏 is the growth rate of the gdp deflator, 𝑐 is the growth rate of the multiplier of scientific and technological progress. this relation shows that the growth rate of the gdp deflator is decomposed into two indicators: 𝑝𝑏 = 𝑐 𝑝𝑝 = 𝑐 × 𝑝𝑐, (20) which implies 𝑝𝑐 = 1 𝑝𝑝 , (21) where 𝑝𝑐 is the true price of a national money in circulation, the market price of which is determined in the currency market of the country (m). but how is the supply of money in circulation determined? according to the basic economic law of the quantity theory of money, the following equality takes place: 𝑣 = 𝑁𝐺𝐷𝑃/𝑀, (22) where 𝑣 is the velocity of money circulation. according to this theory, not only the cost of nominal gdp (ngdp) is circulated, but also the cost of the volume of total costs for its production (𝑋): 𝑣𝑥 = 𝑋/𝑀, (23) where 𝑣𝑥 is the velocity of money circulation. that is, we have [21]: 𝑣 = 𝑁𝐺𝐷𝑃 𝑀 = 𝜇 1+𝜇 × 𝑋 𝑀 = 𝑐 × 𝑣х, (24) where 𝜇/(1 + 𝜇) = 𝑐 is the multiplier of scientific and technological progress, and 𝑣х is the multiplier of the total costs of producing the nominal gdp. since 𝑣 = 𝑐 × 𝑣х, the money supply 𝑀 is determined by the formula: 𝑁𝐺𝐷𝑃 = (𝑐∗𝑣х) × 𝑀, (25) where 𝑝𝑐 = 1/𝑝𝑝 = (𝑐 × 𝑣х) is the true price of deficiency of the dollar equivalent of a national money, as a liquid commodity, in relation to reserve currencies. that is, in the case of kazakhstan, the policy of the national bank should not focus on inflation targeting, but on reducing the deficit of the supply of money in circulation. it is this multiplier “𝑐∗𝑣х” that should become a guideline for setting the discount rates of the central banks of developing countries around the world. modeling the management of the economies of developing countries 112 copyright ©2019 assa. adv. in systems science and appl. (2019) as a result, the nominal gdp in the current national money will grow indefinitely, and the national money itself will depreciate indefinitely. this conclusion in relation to kazakhstan, as an intensively developing country, with the intention to enter the 30 best economic performing countries in the world, according to its economic development, makes it possible to argue that the current focus on the nominal gdp indicator is necessary but not sufficient. and, therefore, relying directly on the real gdp indicator obtained from the nominal one distorts the true rates of economic growth. the world bank report (2017) mentions that the development of the political system of kazakhstan indicates the transfer of activity on economic policy management from the administration of president to the structures of the government and parliament and the expansion of powers of local executive bodies. kazakhstan has announced an increase in the competitiveness of its economy – modernization 3.0 – and the implementation of five priorities for technical modernization, the business environment, macroeconomic stability, improving the quality of human capital, strengthening security institutions and measures to combat corruption [23]. however, the implementation of the super-program of modernization and the five priorities faces an imbalance on the part of the readiness of state (heads of the economic bloc of the government, the national bank) and market institutions to synchronize their actions with the problem of human capital development. this imbalance has led to a rather unexpected for outsiders’ view, but an absolutely adequate response on the part of the authorities to the demands of society, namely focusing on one of the five priorities the development of human capital. this was announced by the president of the state this year, noting the need to implement the five social presidential initiatives (mortgage “7-20-25”, student housing, etc.) [11]. at the same time, it should be assumed that the potential of these five initiatives will be difficult to assess and maximize in all spheres of the country's activities without an appropriate methodology. such a methodology should be based both on the existing traditional indicators and on the new three multipliers associated with the implementation of science and technology in the life of every citizen of the country. as a result, they are aimed at improving citizens’ well-being by increasing human capital and harmonizing relations in the following three chains: “man-science” (stp/tfp multiplier), “man-man” (multiplier of socio-economic progress), as well as “man-society” (multiplier of socio-political progress). according to the comparative analysis, the baizakov’s model is not aiming at increasing the money capital, which, unfortunately, is happening now in kazakhstan. but it is aiming at the development of human capital, which is the most important among all tasks related to the well-being of humanity. it also is confirmed by the un human development index, and is being transferred to the oecd countries. therefore, the use of the baizakov’s model means that in developing countries the passion for speculative capital ceases to exist, and it is not the subjective propositions of the quantitative theory of money that is working at present, but the objective laws of the qualitative theory of money. 6. experimental calculations for the analysis of the comparative effectiveness of innovative projects initial conditions of the comparative analysis. the initial base was the information of the consolidated balance of ssgpo for 2016, as an example for carrying out experimental calculations. it is supplemented with a forecast for 2017 and an expected estimate of data for 2018. for the forecast of the development of ssgpo for 2019–2030, a database of project data for its operational and investment development, exclusively defined in conventional numbers, has been used. consequently, the analysis presented below has only an experimental character for testing the algorithm for solving the problem according to the model of s. baizakov [24]. 113 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) the analysis was carried out by identifying the causes of risks and risk management in the real sector. the level of innovativeness of technology in each version of the real sector is determined by the cost effectiveness of current material resources (which are used to produce the final product). and the function of scientific and technological progress is defined as the result of the productivity of material resources. thus, table 2 shows the calculation scheme for determining the function of scientific and technological progress in one of two options by year of the analytical period, where its last columns express the indicators of the second option for 2018 (idle) and 2028 (forecast) years. according to a given database of design data, the baizakov’s calculation work algorithm is able to assess any prevailing macroeconomic or microeconomic situation. table 2. evaluation of the scientific technological potential of ssgpo algorithm the name of indicators unit. 2016 (fact) 2028 (forecast) nx = ngdp + qp scientific and technological potential of ssgpo in value terms thousan d.tg 183 963 219,0 302 054 905,9 µ = ngdp / qp productivity of variable costs excluding payroll components tg/tg 2,85 0,98 c = ngdp / nx = µ/(1+µ) the multiplier of scientific and technological progress on the productivity of variable costs without wages tg/tg 0,74 0,49 same percentage % 100,00 66,74 ngdp=µ/(1+µ)* nx the algorithm for calculating the ntp multiplier: 1% nx provides µ / (1 + µ)% ngdp thousan d.tg 136 140 105,9 149 191 474,1 nx=tw+tr+qp total consumption of resources for the production of rational consumption in ssgpo thousan d.tg 183 963 219,0 302 054 905,9 ngdp=tw+tr ngdp – gva ssgpo, analogue of the nominal gdp of the country thousan d.tg 136 140 105,9 149 191 474,1 ngdp=µ*qp ngdp – gva ssgpo, analogue of the country's nominal gdp thousan d.tg 136 140 105,9 149 191 474,1 ngdp= =µ/(1+µ)*nx ngdp – gva ssgpo, analogue of the nominal gdp of the country nominal gdp thousan d.tg 136 140 105,9 149 191 474,1 7. analysis of the innovativeness of projects (and operational plans) in the financial sector and the assessment of the cost effectiveness of their implementation. the analysis was performed by identifying the causes, as sources of risk and risk management in the financial sector of the economy. the innovativeness of technology in the economy of the financial sector is determined by the productivity of investment costs and the development of human capital (tw) in the form of normal profits (tr). as you know, the normal profit is determined relative to the wage fund, and it serves as one of the main sources of savings and, consequently, gross accumulation. in a civilized world focused on the development of human capital, the spiritual and creative potential of the person himself, all other conditions being the same, the level of the human development index should constantly increase and, at least, should not decrease in general, the function of socio-economic progress is defined as the result of the productivity of expenses for normal profit, defined as an indicator of rational savings. thus, table 3 shows the calculated scheme for determining the function of socio-economic progress by the years of the analytical period in our conditional example, in ssgpo, where the last modeling the management of the economies of developing countries 114 copyright ©2019 assa. adv. in systems science and appl. (2019) columns express changes in the indicators for 2018-2028 according to the same option 2. table 3 as table 2 is compiled by method of determining the multiplier s. baizakov, and the sources of the data used are the same. table 3. assessment of the socio-economic potential of ssgpo algorithm the name of indicators unit. 2016 (fact) 2028 (forecast) η = tw/tr productivity of expenses in ssgpo on normal profit in the short-term period tg/tg 0,39 0,23 q=η/(1+η) the multiplier of socio-economic potential in ssgpo, tg/tg 0,28 0,19 q=η/(1+η) same percentage % 100,00 67,33 tw=η/(1+η)*ngdp algorithm for calculating the effect of the multiplier of socio-economic potential in ssgpo thousand .tg 38 044 699,7 28 070 475,5 ngdp/l=φ average annual labor productivity per worker in ssgpo thousand .tg 7 318,9 11 387,7 8. analysis and assessment of socio-political progress in the development of the management economy sector the innovation of technology in the economy of the management sector is determined by the product of the multipliers of scientific and technological progress and socio-economic progress (c * q). innovation in the present means both the innovation itself, the innovation, and the process, that is, the potential capable of introducing, say, the economy of the real, financial and managerial sectors of something new, spiritually enriched with the desire not only of a human entrepreneur, but also as simple worker to progress. for example, new technology, or new technologies, patent inventions, or ideas for making management decisions in all the above areas of their activities. but as a potential, not every innovation contributes to an increase in the efficiency of production of goods and services, which are currently in demand and are supplied by demand from consumers. that is, innovation does not necessarily provide a comparative or absolute increase in the efficiency of production of the final results of intellectual activity, its imagination, creative process, discoveries, inventions and rationalization. the fruit of this aspiration of a person of labor and management is the synergistic effect associated with the productivity of the country's local ecological and economic resources and the multiplier of scientific and technological progress, determined on its basis. but the socio-political progress in the development of the managerial economy sector is connected not only with the harmonization of the development of commodity and financial capital. it is also connected with the development of the main component of the development of the productive forces of the country, the spiritual and material development of its human potential. such a three-level system means that the corresponding function of socio-political progress is harmonized with the scientific, technological and socio-economic progress. thus, the following table shows the calculated scheme for determining the function of socio-political progress over the years of the analytical period, and the last columns express the indicators for 2016-1028. table 4, as tables 3-4 compiled by the algorithm for determining the multiplier s. baizakov, on the basis of the previous information of option 2. table 4. assessment of the social and political potential of the ssgpo algorithm the name of indicators unit. 2016 (fact) 2028 (forecast) 115 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) tw/l= w the basic level of the average annual wage per worker in ssgpo thousand tenge 2 045,30 2 142,62 ngdp/l=φ basic annual average labor productivity per worker in ssgpo. thousand tenge 7 318,97 11 387,79 w=μ/(1+μ)*η/(1+η)*ѱ basic level of average annual wage per one worker at full costs in ssgpo thousand tenge 2 045,30 2 142,62 ѱ basic annual average labor productivity at full cost, per worker in ssgpo thousand tenge 9 889,96 23 055,87 с= φ/ ѱ=µ/(1+µ) the basic level of the multiplier ntp and ssgpo tg/tg 0,74 0,49 same percentage % 100,00 66,74 c*q=μ/(1+μ)*η/(1+η)* multiplier of social and political potential in ssgpo tg/tg 0,21 0,09 comparative analysis of two project options for innovative development of ssgpo (with all other conditions being the same). table 5. comparative analysis of innovative development projects of ssgpo algorithm the name of indicators unit. 2028 (option 1) 2028(option 2) nx = ngdp + qp scientific and technological potential of ssgpo in value terms thousand.t g 302 054 905,9 302 054 905,9 µ = ngdp / qp productivity of variable costs excluding payroll components tg/tg 0,98 0,98 c = ngdp / nx = µ/(1+µ) the multiplier of scientific and technological progress on the productivity of variable costs without wages tg/tg 0,49 0,49 same percentage % 68,39 66,74 ngdp=µ/(1+µ)* nx the algorithm for calculating the ntp multiplier: 1% nx provides µ / (1 + µ)% ngdp thousand.t g 136 140 105,9 149 191 474,1 l increase in the number of employed person -53 ѱ=nx/l productivity gains at full cost t/person +95,5 npv gain net present value thousand.t g + 926939.0 npv gain net present value thousand. dollars + 1806.4 as can be seen from the comparative analysis, the baizakov model is not aimed at increasing the money capital, which, unfortunately, is happening now in kazakhstan. and on the development of human capital, which is the most important of the tasks of the well-being of humanity, which is confirmed by the un human development index, and which are being transferred to the oecd countries. therefore, the application of the baizakov model means that in developing countries the passion for speculative capital ceases to exist, and it is not the subjective propositions of the quantitative theory of money that begin to work, but the objective laws of the qualitative theory of money. in conclusion, it should be noted that the theory of preserving the quality of money without applying the gold standard and in the context of globalization and digitalization is called the qualitative theory of money by analogy to the quantitative theory of money by m. friedman. the obtained practical results on the qualitative theory of money make it possible to bridge the gaps between: modeling the management of the economies of developing countries 116 copyright ©2019 assa. adv. in systems science and appl. (2019) micro and macroeconomic indicators; the theory of marginal utility and labor theory of value, short-term budget management plan and long-term strategic planned management tool; tools for assessing the effectiveness of investments in the accounting system of cost accounting of the “direct costing” resources of microeconomics and the productivity of current environmental and economic resources in applied models of trade and investment policy; using cobb-douglas-type macroeconomic production functions and estimating payback period based on net present value (npv) at the microeconomic level; criteria for managing human development in the real and financial sectors of the economy. 9. conclusions on the quality of national money and on the assessment of their impact on economic growth 1. the proposed methodology for the qualitative theory of money, which further develops the theory of keynesianism and the quantitative theory of m. friedman and is successfully applied in the practice of analyzing macroeconomic indicators of production, exchange, distribution and consumption. hence, the basic equation of a balanced economic growth assessment of the real final product fgdp fully has the form of a single cube: 1*fgdp =pp*ngdp= c*rgdp, (26) where ngdp and rgdp are respectively nominal and real gdp, pp and c are respectively indicators of the quality of national money (ngdp), which are in circulation and real gdp (rgdp). thus, a comparison of the growth rates of russia and kazakhstan with the corresponding indicators of four developed countries of the world is given in table 6. table 6. dynamics of changes in the parameters of balanced growth and equilibrium in the economy of russia and kazakhstan compared to the developed countries of the world in 2010-2015. ngdp rgdp fgdp fgdp/ ngdp fgdp/ rgdp 1*fgdp=0.76*ngdp=1.02*rgdp russian 526.9 107 294.5 0.56 2.75 1*294.5=0.56*526.9=2.75*107=294.5 kazakhstan 1572.5 92.7 224.2 0.14 2.42 1*224.2=0.14*1572.5=2.42*92.7=224. 2 germany 173.1 160 183 1.06 1.14 1*183=1.06*173.1=1.14*160=183 france 179 167 168 0.94 1.01 1*168=0.94*179=1.01*167=168 usa 176.2 160 149 0.85 0.93 1*149=0.85*176.2=0.93*160=149 great britain 175.1 130 133 0.76 1.02 1*133=0.76*175.1=1.02*130= 133 source: world development bank database, the analysis was specially conducted over the three five-year plans (2000-2015), on the basis of which, according to the technical task of the project program of the science committee of the republic of kazakhstan, a forecast for 2020 will be made. 2. taking into account the coefficients of bringing them into a single standard form of fgdp, these six countries of the world will be ranked in the order indicated in figure 1. russia 2000-2015 (2000=100%) kazakhstan 2000-2015 (2000=100%) 117 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 1. rating assessment of the growth rate of the actually used final product (fgdp) in six countries of the world. 3. based on the calculations made for these six countries of the world, the following conclusions can be made: 3.1. three-dimensional development figures of six countries of the world are constructed by reducing the three-dimensional measurement parameters to a one-dimensional standard measurement (fgdp). in terms of the growth rate of this key parameter of the economy, over the past five years, russia (294%), kazakhstan (224%), germany (183%), france (168%) ranked second, and the united states (149%) ranked second. ), united kingdom (133%). 3.2. the growth rates of the scientific and technological potential of each country in the world allow deciphering the growth rates of the socio-economic potential in the simplest form, as is the case for calculating tw = q * ngdp. but the actual final consumption product (ftw) created in the country is ftw = c * tw, and the actually created accumulation fund is equal to ftr = c * tr. 3.3. finally, the final result of the development of the economy of each country in the world is actually created by its socio-political potential, which, thanks to its system of sustainable development models, allows us to build a human development index. references 1. rasskazova-nikolaev, s. (2008). direkt-kosting: pravdivaya sebestoimost [direct costing: true cost]; business economics no. 50 (9264); https://www.egonline.ru/article/51932/ [in russian] germany 2000-2015 (2000=100%) france 2000-2015 (2000=100%) usa 2000-2015 (2000=100%) uk 2000-2015 (2000=100%) https://www.eg-online.ru/article/51932/ https://www.eg-online.ru/article/51932/ modeling the management of the economies of developing countries 118 copyright ©2019 assa. adv. in systems science and appl. (2019) 2. adamov n.a. & adamova g.a., management accounting. m.: finance and statistics, 2008 3. sagintayeva c. & baizakov c., (2012) v poiskah konsensusnoy valuti regionalnoy integracii [in search of a consensus currency of regional integration] astana, kazakhstan, aina-print [in russia]. 4. porter, m. e. (2008). competitive strategy: techniques for analyzing industries and competitors. simon and schuster. 5. friedman, m. (2009). capitalism and freedom. university of chicago press. 6. zamulin, o. a & sonin, k.i (2018) economic growth: the nobel prize of 2018 and lessons for russia, questions of the economy, 1, 11 -36. 7. baizakova, s., utembaeva, e., akimova, n. & oinarov, a. (2018) osnovy kachestvennoy teorii deneg [basics of quantitative theories of money], lambert academic publishing [in russian]. 8. anonymous, (2017, june 16) nazarbayev proposed to develop a new method for calculating gdp, economy, [online]: https://dixinews.kz/articles/ekonomika/28238/ [in russian] 9. nazarbayev, n. a. (2009), kliuchi of krizisa [keys to the crisis], rossiyskaya gazeta, 4839 [in russian] 10. nazarbayev, n. a. (2009, september 22) piatiy put’ [the fifth way]. [online]: https://www.zakon.kz/147878-statja-prezidenta-rk-pjatyjj-put.html. [in russian] 11. message from the president of the republic of kazakhstan n.nazarbayev to the people of kazakhstan. october 5, 2018 [online]: http://www.akorda.kz/ru/addresses/addresses_of_president/poslanie-prezidenta-respublikikazahstan-nnazarbaeva-narodu-kazahstana-5-oktyabrya-2018-g 12. baizakov, s. & baizakov, n. (2018) the task of a three-economy economy and the use of its model for comparing the development rates of countries strategic partners of kazakhstan economics and statistics. 2, .25-29. 13. keynes, j. m. (2018). the general theory of employment, interest, and money. springer. 14. baizakov s. et al. (2002), kazakhstan: analysis of trade and investment policy: the reader of the tacis project in kazakhstan (edkz9902) koment. groups of local consultants, under. scientific ed. baizakov s., astana: arkaim. 15. materials of the international conference agricultural trade and foreign investments for sustainable regional integration in the south caucasus and central asia, september 6-7, 2018, baku, through the world food organization (fao). 16. schröder, k. (2018, august 27) china 2030 implications for agriculture in central asia, [online]: http://documents.worldbank.org/curated/en/563161535404143460/china-andrussia-2030-implications-for-agriculture-in-central-asia 17. baye, m. r., prince, j., & squalli, j. (2006). managerial economics and business strategy (vol. 5). new york, ny: mcgraw-hill. 18. marx, k., & engels, f. (1975). marx & engels collected works vol 02: frederick engels:1838-1842. lawrence & wishart. 19. stoléru, l. (1975). economic equilibrium and growth (vol. 1). amsterdam: northholland publishing company; new york: american elsevier. https://dixinews.kz/articles/ekonomika/28238/ https://www.zakon.kz/147878-statja-prezidenta-rk-pjatyjj-put.html 119 s. bayzakov, jeffrey yi-lin forrest, n.a. baizakov copyright ©2019 assa. adv. in systems science and appl. (2019) 21. tleubaev, a., bobozhonov, i., gets, l., hockmann, h. & glauben, t. (2017) determining the productivity and efficiency of grain production in kazakhstan: an approach to stochastic boundaries, discussion paper no. 160, leibniz institute for agricultural development in transition economies. http://hdl.handle.net/10419/155329 22. baizakov, s., dzhunusova, d., shuneev, s. & baizakov, n. (2014), towards a methodology of macroeconomic analysis and economic expertise. economics and statistics, 1, 4-13. 23. world bank annual report (2017, june 30) [online]: http://documents.worldbank.org/curated/en/143021506909711004/world-bank-annualreport-2017 24. baizakov, s., khambar, b., baizakov, n (2018) economic and mathematical foundations of digitization of kazakhstan, economics and statistics, 3, 18-26. adv syst sci appl 2020; 02:56–70 published online at https://ijassa.ipu.ru. numerical methods for constructing solutions of functional differential equations of pointwise type armen beklaryan1* 1national research university higher school of economics, moscow, russia abstract: the article discusses construction of traveling wave type solutions for the frenkelkontorova model on the propagation of longitudinal waves. for the first time, based on the existence and uniqueness theorem of traveling wave type solutions, as well as the approximation theorem, a complete family of traveling wave type solutions is constructed in the form of subfamilies of bounded solutions (horizontals) and unbounded solutions (verticals). keywords: traveling waves, functional differential equations, equations of mathematical physics, splines, forward-backward differential equation 1. introduction the theory of differential equations with delay, which has developed rapidly in recent decades [1–3], has in many ways acquired a complete form and is now actively used in modeling various objects. numerical algorithms for solving equations with delay of various types were also developed [4, 5]. at the same time, numerical methods for equations with advanced arguments (and, moreover, with mixed type deviations) are practically not studied, although references to them have been encountered for a long time, mainly in connection with the classification of equations with deviating argument [6]. as a rule, there is a parsing of equations of a particular kind with further obtaining the existence and uniqueness theorems based on the use of the properties of the right-hand side [7] and application of methods, such as the study of the roots of the characteristic quasi-polynomial [8], collocation methods and finite element scheme (expansion of a solution in terms of basis functions of some finitedimensional space) [9–11] or through the construction of a hilbert space of the reproducing kernels on the basis of boundary conditions [12]. for equations of a delayed type, as a rule, a well-known initial problem is considered when the initial moment of time coincides with the left end of the interval of equation’s domain. such a statement is characterized by the fact that the initial-boundary conditions are local in nature, and the equation itself can be integrated by the step method. in all other cases, and even more so in the case of functional differential equations of pointwise type (fdept) with mixed-type deviations, the initial-boundary conditions are nonlocal in nature. one of the approaches to the construction of numerical methods for fdept is the variational method based on solving the induced optimal control problem (ocp) as an unconditional optimization problem for the residual functional. as noted earlier, the initialboundary conditions for fdept are not local, therefore, the solution of the variational problem for the residual functional cannot be obtained by simple local improvements and requires global optimization. note that for problems of finite-dimensional optimization there ∗corresponding author: abeklaryan@hse.ru numerical methods for functional differential equations 57 are many approaches to finding a global extremum, sufficiently efficient algorithms are built, there are a large number of publications that present precedents for successfully solved problems [13, 14, for example]. at the same time, the task of finding global extremum for ocp remains one of the most acute tasks in the extreme problems theory. among classical approaches to this task one cannot fail to mention the works [15] and [16]. some works that apply search and genetic algorithms are also well-known [17, 18]. but unfortunately these efforts have not yet brought about effective algorithms that would be able to solve a wide spectrum of practical problems [19]. one of the ideas in this area is the idea of reducing the optimal control problem to a low-dimensional problem on the reachability set of a controlled system [20]. unfortunately, the problem of approximating the reachable set in the numerical solution is not much simpler than the problem of finding a global extremum. nevertheless, the properties of the reachable set known from theoretical studies can be used to construct specialized optimization algorithms for ocp [21]. in particular, the connectivity property of the reachability set allows one to construct a scheme of continuous control variation that leads to the solution of the problem. the idea of using the connectivity property of the reachability set when constructing numerical optimization algorithms belongs to a.g. chentsov [22]. in the majority of well-known approaches to the construction of non-convex optimization methods, the solution of the problem is divided into two stages: “global”, where a wide scan of the variable space is performed, and “local”, aimed on local refinement of the obtained solution [14]. the combination of different methods at each stage, as well as the order of the alternation of stages, determines the specific computational algorithm. for mentioned class of problems, we consider a idea of combining both stages in one algorithm package and creating a method that allows one to select and change both the “global scan” procedure and local improvement of the obtained approximate solution at each iteration [23]. 2. statement of the initial boundary-value problem the most important goal in fdept is the study of the basic initial-boundary value problem ẋ(t) = f(t, x(q1(t)), . . . , x(qs(t))), t ∈ br, (2.1) ẋ(t) = ϕ(t), t ∈ r\br, ϕ(·) ∈ l∞(r,rn), (2.2) x(t̄) = x̄, t̄ ∈ r, x̄ ∈ rn. (2.3) where f : r× rns −→ rn is a mapping of the c(0) class, qj(·), j = 1, . . . , s, represent diffeomorphisms of the line preserving orientation, andbr is either the closed interval [t0, t1], or closed half-line [t0,+∞), or the whole line r. in the case of a delayed-type equation, for br = [t0, t1], [t0,+∞), t̄ = t0, the initialboundary-value problem becomes a well-known initial problem with the initial function ζ(t) = x̄− ∫ t̄ t ϕ(τ)dτ, t ≤ t̄. the statement of the initial-boundary value problem (2.1)-(2.3) is correct without any restrictions on the type of argument deviations. definition 2.1: the solution to the basic initial-boundary value problem is any absolutely continuous solution of the equation (2.1) satisfying the boundary condition (2.2) and the initial condition (2.3). the aim of the fde research is to study the solution space of the boundary value problem (2.1)-(2.2) for each given boundary function ϕ, as well as to describe the obstacles due to which the solutions of the boundary value problem (2.1)-(2.2) do not inherit the properties of solutions of ordinary differential equations. copyright © 2020 assa. adv syst sci appl (2020) 58 a. beklaryan if we introduce notation for a fixed function ϕ(·) fϕ(t, z1, . . . , zs) = { f(t, z1, . . . , zs), t ∈ br, ϕ(t), t /∈ br, the boundary value problem (2.1)-(2.3) becomes a cauchy problem ẋ(t) = fϕ(t, x(q1(t)), . . . , x(qs(t))), t ∈ r (2.4) x(t̄) = x̄, t̄ ∈ r, x̄ ∈ rn. (2.5) obviously, in the case br = r, the equality fϕ = f holds. in the study of the cauchy problem (2.4)-(2.5), the role of the group < q1, . . . , qs > is very important. often, instead of such a group, it is useful to consider some wider finitely generated group q of diffeomorphisms of the line, that is, < q1, . . . , qs >⊆ q. moreover, we assume that the condition hq = supt∈r |q(t)− t| < +∞ is satisfied for all elements q of the group q. let’s formulate a system of restrictions on the right-hand side f : r× rn·s 7−→ rn of fdept (2.1), and on diffeomorphisms qj(·), j = 1, . . . , s: (a) f(·) ∈ c(0)(r× rn·s,rn); (b) for any t, zj, z̄j, j = 1, s, there is a quasilinear growth ‖f (t, z1, . . . , zs) ‖rn ≤m0(t) +m1 s∑ j=1 ‖zj‖rn , m0(·) ∈ c(0)(r,r), and a lipschitz condition ‖f (t, z1, . . . , zs)− f (t, z̄1, . . . , z̄s) ‖rn ≤ lf s∑ j=1 ‖zj − z̄j‖rn ; (c) there exists µ∗ ∈ r+ such that m0(·) ∈ lnµ∗c(0)(r); (d) the values hqj = sup t∈r |t− qj(t)|, j = 1, . . . , s, are finite; (e) with the constant µ∗ ∈ r+ from condition (c) the family of functions f̃q,z1,...,zs(t) = f(q(t), z1, . . . , zs)(µ ?)hq , q ∈ q, z1, . . . , zs ∈ rn, is equicontinuous on any finite interval. the continuity condition, growth conditions with respect to the phase variables and time variable, and the lipschitz condition (conditions (a)− (b)) are standard in the theory of ordinary differential equations. in fact, in item (b), the first inequality in the form of the growth condition with respect to the phase variables and time variable is a consequence of the second condition in the form of the lipschitz condition. under the lipschitz constant lf , the minimum value among the possible values of these constants should be understood. accordingly, we can assume that m1 = lf . moreover, we wrote the first inequality out separately in order to formulate condition (c) for functionm0(·). condition (c) for function f is related to the study of solutions on the half-line and line, which requires certain restrictions copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 59 on the time growth of the right-hand side. it should be noted that we can always satisfy condition (d) by making a time change, but we may break condition (c). the last condition (e) is necessary for the right-hand side of the induced infinite-dimensional ordinary differential equation, with the phase space in a suitable banach space, to easily establish the fact of its bochner integrability. in fact, condition (e) can be removed, but this leads to further technical complications, as the right-hand side of the induced infinite-dimensional ordinary differential equation will only be measurable, and it will be necessary to establish the fact of its bochner integrability [24]. the right-hand side f(·) of fdept will be considered as an element of the banach space vµ∗(r× rns,rn) with lipschitz norm vµ∗(r× rns,rn) = { f(·) : f(·) satisfies the conditions (a)-(d) } , ‖f(·)‖lip = sup t∈r ‖f(t, 0, . . . , 0)(µ∗)|t|‖rn + + sup (t,z1,...,zs,z̄1,...,z̄s)∈r1+2ns ‖f(t, z1, . . . , zs)− f(t, z̄1, . . . , z̄s)‖rn∑s j=1 ‖zj − z̄j‖rn , where the parameter µ∗ ∈ r+ coincides with the corresponding constant from the condition (c). obviously, for the function f(·) ∈ vµ∗(r× rns,rn) the smallest value of the constant lf from the lipschitz condition (the condition (b)) coincides with the value of the second summand in the definition of the norm f(·). in what follows, speaking of the lipschitz condition, by the constant lf we mean exactly its smallest value. the right-hand side of the fdept (2.1) uniquely defines a pair (f(·), h), where f(·) is an element of the banach space vµ∗(r× rns,rn), and h = (hq1 , . . . , hqs), hqj ≥ 0, j = 1, . . . , s, are maximum deviation of the argument from the condition (d), which we will consider as parameters. therefore, such a right-hand side of the equation will uniquely determine the pair (lf ;h). we will seek a solution to the fdept with a quasilinear right-hand side in a oneparameter family of banach spaces of functions with a given exponential growth. the exponent is the parameter of the selected family of functions, which is defined as follows lnµc(k)(r) = { x(·) : x(·) ∈ c(k) (r,rn) , max 0≤r≤k sup t∈r ‖x(r)(t)µ|t|‖rn < +∞ } , ‖x(·)‖(k) µ = max 0≤r≤k sup t∈r ‖x(r)(t)µ|t|‖rn , k = 0, 1, . . . , µ ∈ (0,+∞). 2.1. existence and uniqueness theorem for the basic initial-boundary value problem in terms of the parameter µ ∈ (0, 1) and the space lnµc(0)(r) we formulate a condition guaranteeing the existence and uniqueness of a solution for the initially-boundary problem. theorem 2.1 ( [25]): let function f and diffeomorphisms qj(·), j = 1, . . . , s, satisfy the conditions (a)− (e) from the section 2. if for some µ ∈ (0, µ∗) ∩ (0, 1) the inequality lf s∑ j=1 µ−hqj < lnµ−1, (2.6) is satisfied then for any fixed initial-boundary conditions x̄ ∈ rn, ϕ(·) ∈ l∞(r,rn) there exists a solution (absolutely continuous) x(·) ∈ lnµc(0)(r) of the basic initialboundary value problem (2.1)-(2.3). such a solution is unique and, as an element of the copyright © 2020 assa. adv syst sci appl (2020) 60 a. beklaryan space lnµc(0)(r), depends continuously on the initial-boundary conditions x̄ ∈ rn, ϕ(·) ∈ l∞(r,rn) and the right-hand side of the equation (function f(·)). the condition (2.6) is exact and not improvable: there are differential equations for which violation of this condition leads to a violation of either the existence of the solution or uniqueness. note that if the inequality (2.6) has a solution, then there are µ1(lf ;h), µ2(lf ;h) such that for each solution µ the following condition holds µ1(lf ;h) < µ < µ2(lf ;h). 3. variational approach to constructing solutions to the basic initial-boundary value problem we study absolutely continuous solutions of the system fi(t, ẋ(t), x(t+ τ1), . . . , x(t+ τs)) = 0, i = 1, k, t ∈ [tl, tr], (3.7) under boundary conditions outside the system definition interval ẋ(t) = ϕl(t), t ∈ [tll, tl], (3.8) ẋ(t) = ϕr(t), t ∈ [tr, trr], (3.9) as well as initial-boundary conditions on the system definition interval km(ẋ(ξ), x(ξ1), . . . , x(ξp)) = 0, (3.10) where fi : r× rn × rns −→ rn, i = 1, k, is a mapping of the c(0) class; τj ∈ z, j = 1, s; tl, tr ∈ r; tll = tl + min{0, τ1, . . . , τs}, trr = tr + max{0, τ1, . . . , τs}; ϕl, ϕr : rn −→ rn are mappings of thec(0) class, defining fixed boundary vector functions;km : rn × rnp −→ rn,m = 1, q, is a mapping of the c(0) class; ξ, ξi ∈ [tl, tr], i = 1, p, is a fixed set of points on the system definition interval. we give a formal statement of the optimization problem induced by the initial-boundaryvalue problem. problem a. find the trajectory x̂(t), which provides a minimum of residual functional i(x̂(t)) = v(n) ( k∑ i=1 ∫ tr tl f 2 i (t, ˙̂x(t), x̂(t+ τ1), . . . , x̂(t+ τs))dt+∫ tl tll [ ˙̂x(t)− ϕl(t)]2dt+ ∫ trr tr [ ˙̂x(t)− ϕr(t)]2dt ) + v(k) q∑ m=1 k2 m( ˙̂x(ξ), x̂(ξ1), . . . , x̂(ξp)), (3.11) where v(n), v(k) ∈ r+ are weighting factors. the following statement holds. proposition 3.1: each solution to the initial-boundary value problem (3.7)-(3.10) is a solution to the optimization problem a. in particular, under the conditions of theorem 2.1, a solution to the initial-boundary value problem (2.1)-(2.3) exists and is a solution to the corresponding problem a. in the general copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 61 case, the converse statement to proposition 3.1 is false, but, nevertheless, taking into account the uniqueness theorem and also the stability theorem of the solution, in the case of successful construction of a numerical solution, we can speak with high confidence about constructing a solution close to analytical one. the proposed variational approach to constructing solutions of the basic initial-boundary value problem lies at the basis of the numerical implementation of such solutions. the numerical methods themselves are based on the ritz method and spline collocation constructions and were implemented in [23, 26]. in order to solve the problem of the class under consideration, the trajectories of the system are discretized on a grid with a constant step, and a generalized residual functional is formulated that includes both the weighted residual of the original differential equation and the residual of the boundary conditions. a spline differentiation technique is used to evaluate the derivatives of the desired trajectories of the system, based on two spline approximation designs: using cubic natural splines and using a special type of spline whose second derivatives at the edges are also controlled using optimized parameters. 3.1. software package optcon-f a set of algorithms for local and global optimization was implemented for solving the stated finite-dimensional problems. the used technology includes: an algorithm for sequentially increasing the accuracy of approximation by multiplying the number of nodes in the grid of the discretization; algorithms for the difference evaluation of the derivatives of the functional from the first to the sixth degree of accuracy inclusive; method of successively increasing the precision of spline differentiation. the corresponding software complex (sc) optcon-f was implemented in the language c under the control of operating systems os windows, os linux and mac os using compilers bcc 5.5 and gcc. sc was designed to obtain a numerical solution of boundary value problems, parametric identification problems and optimal control for dynamical systems described by fdept [27]. among the local algorithms included in the sc optcon-f there are: 1) partan method; 2) powell-brent’s method; 3) gradient method of confidence intervals; 4) barzilaiborwein method; 5) newton’s method with the difference estimate of the hessian matrix; 6) generalized quasi-newtonian method; 7) direct-dual method of gradient descent; 8) differential euler optimization method of the 2nd order; 9) differential adams optimization method of the 4th order and others. as algorithms of “closers” there are: 1) adaptive modification of the hooke-jeeves method; 2) stochastic search methods in random subspaces of the indicated (2, 3, 4, or 5) dimension; 3) local version of the curvilinear search method. non-local algorithms that form the basis of the sc are: 1) “parabolic” method – a combination of coordinate-wise descent with a periodic multistart and a non-local onedimensional search by the parabola algorithm; 2) non-local method of curvilinear search; 3) luus-jaakola method; 4) “forest” method – multivariant adaptive method of random multistart and others. as part of the work on the sc, taking into account the specifics related to the qualitative properties of fdept, the following key results were obtained: • algorithm of construction of “controllable splines” has been developed. the problem of obtaining a high-precision approximation to the derivative of a function of one variable over a set of values of the function itself, defined on a fixed grid, was investigated. the performed computational experiments showed that in order to achieve good accuracy in estimating derived trajectories for fdept systems, it is not possible to use known types of splines (“natural”, with additional boundary conditions, akima splines, etc.). the proposed algorithm is based on the use of derivatives of cubic spline functions. • a technology has been developed for approximation of general fdept systems using a finite-dimensional unconditional minimization problem. by analyzing the copyright © 2020 assa. adv syst sci appl (2020) 62 a. beklaryan behavior of homeomorphisms on the initial (main) time interval, an extended time interval is calculated, for which the discretized grid is constructed. to approximate the initial continuous problem on a fixed grid over time, the ritz method is used: the trajectories are approximated using controlled spline functions, the coefficients of which are selected by searching for the minimum of the residual functional. • a specialized global optimization algorithm was developed, based on the idea of curvilinear search. to search for the minimum of the non-convex residual functional, a specialized global optimization algorithm has been developed, based on the idea of curvilinear search on the reachable set of a control system produced using a pair of randomly generated “support” controls. using the property of connectedness, variations in the space of control functions that do not violate the existing constraints are constructed at iterations of the algorithm. this is achieved by direct projection on a parallelepiped. to improve the current approximation, a modification of a non-local onedimensional search algorithm based on the “parabolic” method is applied [28]. also, as a globalizing mechanism, a non-local search in random directions is used, repeated many times at each iteration of the algorithm. to solve auxiliary non-convex problems of one-dimensional search, a modification of the stochastic p -algorithm, proposed by a. zhigljavsky and a. žilinskas [14], is implemented [29]. to enhance reliability of the proposed method, a periodic random multi-start was provided in the algorithm’s construction. • an algorithm for the numerical integration of fdept systems based on the sequential discretization technique has been developed. to improve the efficiency of calculations, tools have been implemented to build a sequence of approximative finitedimensional optimization problems with a growing number of variables. at the same time, solutions obtained at the previous stages of calculations performed on the current discretization grid are projected onto a new grid with an increased number of nodes while preserving the qualitative and quantitative characteristics of the trajectories. • the proposed numerical algorithms were tested on a wide range of tasks [27] with using the principle of “the best of known solutions” [30]. in all considered problems, the proposed algorithm allowed us to find the best known solution. the calculation experiments that have been carried out demonstrate considerably high fidelity of the proposed algorithms. the above heuristic search algorithm for the solution x̂(t) can be justified on the basis of the existence and uniqueness theorems for initial-boundary value problems for the investigated fdept, as well as theorems on approximating solutions of such equations on the whole line by solutions of the initial-boundary value problem on a sequence of expanding intervals. a description of such equations and corresponding results are presented in the following sections. 4. existence and uniqueness theorem for soliton solutions in the problem of longitudinal vibrations of an infinite homogeneous rod we consider a problem from the theory of plastic deformation, in which solutions of the traveling wave type for the following system are studied mÿi(t) = yi+1(t)− 2yi(t) + yi−1(t) + φ(yi(t)), i ∈ z, yi ∈ r, t ∈ r, (4.12) yi(t+ τ) = yi+1(t), τ ≥ 0. (4.13) we will study such a system under the most weak conditions on the potential φ(·) in the form of the lipschitz condition. the lipschitz constant will be denoted by lφ. we formulate a theorem on the existence and uniqueness of a traveling wave type solution (soliton solution), copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 63 τ µ 1 µ̂ τ̂ µ1(τ ) µ2(τ ) 0 fig. 4.1. graphs of functions µ1(τ), µ2(τ) based on the correspondence of traveling-wave type solutions of the infinite-dimensional system (4.12)-(4.13) to the solutions of the induced fdept [24], reduced to a first-order system ẋ1(t) = x2(t), x ∈ r2, x = (x1, x2)′, t ∈ r, ẋ2(t) = m−1 ( x1(t+ τ)− 2x1(t) + x1(t− τ) + φ(x1(t)) ) , (4.14) as well as the existence and uniqueness theorem for solutions to such equations (theorem 2.1). to do this, we consider a transcendental equation with respect to two variables τ ∈ [0,+∞) and µ ∈ (0, 1) cτ ( 2µ−1 + 1 ) = lnµ−1, (4.15) where c = max{1; 2m−1 √ l2 φ + 2}. the set of solutions of the equation (4.15) is described by functions µ1(τ), µ2(τ) given in fig. 4.1. if we denote the right-hand side of the system (4.14) by f , then relations µ1(lf ;h) = τ √ µ1(τ) and µ2(lf ;h) = τ √ µ2(τ) will be valid, where lf = c, h = (τ, τ). for each µ ∈ (0, 1) we define the phase space of the infinite-dimensional equation (4.12)as a hilbert space of sequences with the corresponding norm k2 z2µ = {κ : κ = {zi}+∞ −∞, zi ∈ r2, i ∈ z, ∑ i∈z ‖zi‖2 r2µ2|i| < +∞}. theorem 4.1 ( [24]): let the potential φ satisfy the lipschitz condition with constant lφ. then, for any initial values ī ∈ z, a, b ∈ r, t̄ ∈ r, and characteristics τ > 0 satisfying the condition 0 < τ < τ̂ , copyright © 2020 assa. adv syst sci appl (2020) 64 a. beklaryan for the initial system of differential equations (4.12) there exists a unique solution of the traveling wave type {yi(·)}+∞ −∞ with characteristic τ such that it satisfies the initial conditions yī(t̄) = a, ẏī(t̄) = b. for any parameter µ ∈ (µ1(τ), µ2(τ)) the vector function ω(t) = {(yi(t), ẏi(t)) ′}+∞ −∞ belongs to the space k2 z2µ for any t ∈ r, and the function ρ(t) = ‖ω(t)‖2µ belongs to the space l1 τ √ µc (1)(r). such a solution depends continuously on the initial values a, b ∈ r, as well as on the mass m. theorem 4.1 not only guarantees the existence of a solution but also determines the limitation of its possible growth both in time t and in coordinates i ∈ z (over space). it is obvious that for each 0 < τ < τ̂ the space k2 z2(µ2(τ)−ε), for small ε > 0, is much narrower than the space k2 z2(µ1(τ)+ε). the theorem guarantees the existence of a solution in narrower spaces and uniqueness in wider spaces. 5. approximation of the solution of a functional differential equation defined on the line by solutions of an initialboundary value problem with expanding intervals of the definition next, we consider soliton solutions of the system, which will be implemented as solutions of the induced functional differential equations defined on the whole line. for the numerical integration of such an equation, one should be able to approximate the solutions of such an equation by the solutions of initial-boundary value problems defined on an expanding family of finite intervals. this section is devoted to this problem. let’s consider the initial boundary value problem ẋ(t) = f(t, x(t+ τ1), . . . , x(t+ τs)), t ∈ br, (5.16) ẋ(t) = ϕ(t), t ∈ r\br, ϕ(·) ∈ l∞(r,rn), (5.17) x(t̄) = x̄, t̄ ∈ r, x̄ ∈ rn, (5.18) where τj ∈ r. we formulate again the theorem on the continuous dependence on the initialboundary conditions. we define a banach space of functions with weight lnµl∞(r) = { ϕ(·) : ϕ(·) ∈ l∞ (r,rn) , sup t∈r vrai‖ϕ(t)µ|t|‖rn < +∞ } , µ ∈ (0, 1) and norm ‖ϕ(·)‖µ = sup t∈r vrai‖ϕ(t)µ|t|‖rn . if in the boundary condition (5.17), starting with some sufficiently large k ∈ z+ with the condition br ⊂ [−k, k], the boundary condition ϕ(·) ∈ l∞(r) replace outside the interval br ⊂ [−k, k] so that the new boundary function ϕ̃(·) satisfies the condition ϕ̃(·) ∈ lnµl∞(r), then the corresponding solution x̃(·) of the boundary value problem on br ⊂ [−k, k] will match the original solution x(·). we are going to formulate a proposition on the approximation of solutions of an initialboundary value problem defined on the whole line by solutions of the initial-boundary value copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 65 problem defined on the interval [−k, k] as k → +∞. we consider the initial-boundary value problem on the whole line br = r ẋ(t) = f(t, x(t+ τ1), . . . , x(t+ τs)), t ∈ r, (5.19) x(t̄) = x̄, t̄ ∈ r, x̄ ∈ rn, (5.20) where τj ∈ r, j = 1, . . . , s, j = 1, . . . , s. for each k ∈ z we consider the initial-boundary value problem on a finite interval br = [−k, k] ẋ(t) = f(t, x(t+ τ1), . . . , x(t+ τs)), t ∈ [−k, k], (5.21) ẋ(t) = ϕ(t), t ∈ r\[−k, k], ϕ(·) ∈ ln1l∞(r), (5.22) x(t̄) = x̄, t̄ ∈ r, x̄ ∈ rn. (5.23) theorem 5.1 ( [31]): let the map f(·) satisfies the conditions (a)− (d). if for µ ∈ (0, µ?) ∩ (0, 1) the inequality lf s∑ j=1 µ−|τj | < lnµ−1 (5.24) holds, and (µ1(lf ;h), µ2(lf ;h)) is the maximum solution interval for the inequality (5.24), then for any x̄ ∈ rn, ϕ(·) ∈ ln1l∞(r), the solution x̂(·) of the initial-boundary value problem (5.19)-(5.20), as an element of the space lnµc(0)(r), is approximated by solutions x̂k(·) of the initial-boundary value problem (5.21)–(5.23) as k → +∞. moreover, for any arbitrarily small ε, 0 < ε < µ2 − µ1 there is cfϕε such that the following estimate is true ‖x̂(·)− x̂k(·)‖(0) µ ≤ cfϕε ( µ1(lf ;h) µ2(lf ;h)− ε )k . from the estimate it follows that the convergence rate will be the higher, the more the µ1(lf ;h) and µ2(lf ;h) values will differ from each other, the difference between them is inversely proportional to the value of lf max1≤j≤s |τj|. 6. construction of the complete family of traveling wave type solutions in the frenkel-kontorova model considering periodic functionals for the model from the section 4, we obtain the frenkelkontorova model from the theory of plastic deformation. in this case, taking the functional φ(y) = a sin(by), a,b ∈ r mÿi(t) = yi+1(t)− 2yi(t) + yi−1(t) + asin(byi(t)), i ∈ z, yi ∈ r, t ∈ r, (6.25) yi(t+ τ) = yi+1(t), τ > 0, (6.26) we construct solutions of the traveling wave type with characteristic τ . as noted earlier, the space of traveling wave type solutions for such a system coincides with the space of solutions of the fdept system ẋ1(t) = τx2(t), t ∈ r, ẋ2(t) = τm−1[x1(t+ 1)− 2x1(t) + x1(t− 1) + a sin(bx1(t))]. (6.27) moreover, there is a correspondence of solutions according to the rule yi(t) = x(τ−1t+ i). we will construct numerical solutions of the system (6.27) using the approximation theorem 5.1. copyright © 2020 assa. adv syst sci appl (2020) 66 a. beklaryan 6.1. numerical experiments next, the results of the computational experiments on the study of initial-boundary value problems for systems of fdept using optcon-f software will be presented. before demonstrating the examples, it is necessary to make a number of significant observations: (a) in the sc optcon-f, the possibility of satisfying the condition of uniform boundedness of the derivative is realized. the maximum deviation from zero is determined by the lipschitz constant of the right side of the equation, by the parameter µ, and also by the deviations of the argument; (b) numerical integration on an interval with initial-boundary conditions is realized as an integration process with given boundary conditions at the left end and procedures for minimizing deviations from the boundary conditions at the right end for the solution constructed with observance of the restrictions from the preceding item; (c) in the obtained theorem on approximating solutions of the original equation on the whole line by solutions on expanding finite intervals [−k, k], the only restriction on the boundary conditions themselves is the condition for their uniform boundedness. in particular, we choose the zero boundary condition; (d) we note that the obtained condition for the existence of a solution of the traveling wave type is just a sufficient condition. therefore, solutions of the traveling wave type can also be numerically constructed for τ > τ̂ . nevertheless, many of the central conditions of the presented theory are “exact”, that is, there are examples of equations for which violation of the indicated conditions leads to the lack of solutions; (e) in the sc optcon-f there is the possibility of sequential application of various algorithms within the framework of constructing a solution for one task. thus, the constructed intermediate solution in the previous step becomes the starting solution (“baseline”) for the following algorithm. in this case, such an implementation does not prevent the global search algorithms from “popping out” of the local solution. separately, we note the presence of a programming module that allows predetermining the order of application of algorithms, as well as the construction of complex chain of steps (conditional statements, loops, etc.) depending on the current or historical values of a number parameters (for example, error estimation or number of iterations). for the examples presented below, in addition to the basic global optimization algorithm described in section 1, the following scheme was used in cycle: the generalized quasinewtonian and powell-brent’s methods (with bi-directional line search along each dimension) were used alternately, and after the error changed by less than 10−p the adaptive modification of the hooke-jeeves method was used l times (p and l are computable functions on the basis of the loop iteration number, as well as the lipschitz constants of the equation itself and a number of other technical characteristics). the stopping criterion depended on the number of iterations in the first part of the cycle, as well as the current error estimate and its dynamics; (f) in view of all the points listed above, as well as stochastic elements in the applied algorithms, the presented value of the residual functional (rf, i.e. error of a numerical solution) can not be used to estimate the theoretical rate of convergence. we consider dynamical system in the following form:{ ẋ1(t) = 0.15x2(t), ẋ2(t) = 0.15× 100−1(x1(t+ 1)− 2x1(t) + x1(t− 1) + 500 sin(0.1x1(t))), t ∈ r, initial conditions{ x1(0) = c, x2(0) = 0. (6.28) copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 67 (a) trajectories of x1(t) for the system (6.29) (b) trajectories of x1(t) for the system (6.28) fig. 6.2. trajectories of x1(t) at different c here, with respect to the system (6.27), we have a = 500, b = 0.1, τ = 0.15,m = 100, and the equality (4.15) takes the form 0.003 √ 2502(2µ−1 + 1) = lnµ−1 and has on the interval (0, 1) two solutions with approximate values µ1(0.15) = 0, 22 and µ2(0.15) = 0, 424191 (the exact values are expressed in terms of the lambert w -function and can’t be written out in quadratures). taking into account the impossibility of considering the numerical solution of the system on an infinite interval, we introduce the parameter k and the corresponding family of expanding initial-boundary value problems{ ẋ1(t) = 0.15x2(t), ẋ2(t) = 0.15× 100−1(x1(t+ 1)− 2x1(t) + x1(t− 1) + 500 sin(0.1x1(t))), t ∈ [−k, k], boundary conditions{ ẋ1(t) = 0, ẋ2(t) = 0, t ∈ (−∞,−k] ∪ [k,+∞), initial conditions{ x1(0) = c, x2(0) = 0. (6.29) according to the theorem 5.1, convergence is guaranteed under any given essentially bounded boundary conditions. in the system (6.29), the simplest boundary function is chosen, the identity zero. thus, the solution of the system (6.29) converges (according to the metric of the space lnµc(0)(r) for µ ∈ (µ1(0.15), µ2(0.15))) to the solution of the system (6.28) as k →∞. since the equation (6.27) is autonomous, the solution space of such equation is invariant with respect to time-variable shifts. on the other hand, from the periodicity of the right-hand side with respect to the x1(t) it follows that the solution space of such equation is invariant copyright © 2020 assa. adv syst sci appl (2020) 68 a. beklaryan (a) trajectories of x1(t) at different k (b) trajectories of x1(t) at different c fig. 6.3. unbounded trajectories of x1(t) with respect to a shift in x1(t) for a period equal to 2π b . therefore, it suffices to consider a family of solutions of the initial system (6.28) with a value of initial conditions x1(0) from zero to the value of the period. nevertheless, the stationary solutions x1(t) are repeated each half-period π b . since the right-hand side of the equation is an odd function with respect to the x1(t), the solution space of such equation can withstand the reflection transformation with respect to the axis t. hence, it is sufficient to construct trajectories x1(t) in the strip from zero to the half-period. figure 6.2 shows the graphs x1(t) for different values of the parameter c = x1(0) for both the system (6.29) and the initial system (6.28) (the values of c are reduced to half-period). note that fig. 6.2 shows the complete family of bounded graphs x1(t) (up to the above transformations) from the space lnµc(0)(r) for µ ∈ (µ1(0.15), µ2(0.15)). at the same time, for some initial conditions, there are only unbounded x1(t). figure 6.3a demonstrates the evolution of such an unbounded graph x1(t) with increasing k for the initial conditions x1(0) = π b , x2(0) = 50. figure 6.3b, in turn, shows a one-parameter family of unbounded graphs x1(t) for x1(0) = π b and different values of x2(0) = c. 7. conclusion the construction of numerical solutions of the traveling wave type for the frenkel-kontorova model on the propagation of longitudinal waves using the developed software package was demonstrated. for this model, a complete family of traveling wave solutions was constructed. acknowledgements the reported study was funded by rfbr according to the research project 19-01-00147. references 1. bellman, r. & cooke, k. l. (1963) differential-difference equations. new york: academic press. 2. hale, j. k. (1977) theory of functional differential equations. new york: springer. copyright © 2020 assa. adv syst sci appl (2020) numerical methods for functional differential equations 69 3. myshkis, a. d. (1972) linear differential equations with retarded arguments. moscow: nauka, (in russian). 4. bellen, a. & zennaro, m. (2003) numerical methods for delay differential equations. usa: oxford university press. 5. maset, s. (2003) numerical solution of retarded functional differential equations as abstract cauchy problems. journal of computational and applied mathematics, 161, 259–282, https://dx.doi.org/10.1016/j.cam.2003.03.001. 6. el’sgol’ts, l. e. & norkin, s. b. (1973) introduction to the theory and application of differential equations with deviating arguments. new york: academic press. 7. baotong, c. (1995) functional differential equations mixed type in banach spaces. rendiconti del seminario matematico della università di padova, 94, 47–54. 8. ford, n. j. & lumb, p. m. (2009) mixed-type functional differential equations: a numerical approach. journal of computational and applied mathematics, 229, 471– 479, https://dx.doi.org/10.1016/j.cam.2008.04.016. 9. abell, k. a., elmer, c. e., humphries, a. r., & van vleck, e. s. (2005) computation of mixed type functional differential boundary value problems. siam journal on applied dynamical systems, 4, 755–781, https://dx.doi.org/10.1137/040603425. 10. lima, p. m., teodoro, m. f., ford, n. j., & lumb, p. m. (2010) analytical and numerical investigation of mixed-type functional differential equations. journal of computational and applied mathematics, 234, 2826–2837, https://dx.doi.org/10.1016/j.cam.2010.01.028. 11. lima, p. m., teodoro, m. f., ford, n. j., & lumb, p. m. (2010) finite element solution of a linear mixed-type functional differential equation. numerical algorithms, 55, 301– 320, https://dx.doi.org/10.1007/s11075-010-9412-y. 12. li, x. y. & wu, b. y. (2014) a continuous method for nonlocal functional differential equations with delayed or advanced arguments. journal of mathematical analysis and applications, 409, 485–493, https://dx.doi.org/10.1016/j.jmaa.2013.07.039. 13. evtushenko, y., posypkin, m., & sigal, i. (2009) a framework for parallel large-scale global optimization. computer science – research and development, 23, 211–215, https://dx.doi.org/10.1007/s00450-009-0083-7. 14. zhigljavsky, a. & žilinskas, a. (2008) stochastic global optimization. us: springer, https://dx.doi.org/10.1007/978-0-387-74740-8. 15. bellman, r. (1957) dynamic programming. princeton, new jersey: princeton university press. 16. krotov, v. f. (1996) global methods in optimal control theory. new york: marcel dekker inc. 17. chachuat, b., singer, a. b., & barton, p. i. (2006) global methods for dynamic optimization and mixed-integer dynamic optimization. industrial & engineering chemistry research, 45, 8373–8392, https://dx.doi.org/10.1021/ie0601605. 18. lopez cruz, i. l., van willigenburg, l. g., & van straten, g. (2003) efficient differential evolution algorithms for multimodal optimal control problem. applied soft computing, 3, 97–122, https://dx.doi.org/10.1016/s1568-4946(03)00007-3. 19. floudas, c. a. & gounaris, c. e. (2009) a review of recent advances in global optimization. journal of global optimization, 45, 3–38, https://dx.doi.org/10.1007/s10898-008-9332-8. 20. khrustalev, m. m. (1988) an accurate description of feasibility sets and conditions for the global optimality of dynamic systems. ii. conditions for global optimality. automation and remote control, 49, 874–881. 21. tolstonogov, a. (2000) differential inclusions in a banach space. netherlands: springer, https://dx.doi.org/10.1007/978-94-015-9490-5. 22. subbotin, a. i. & chentsov, a. g. (1981) guarantee optimization in control problems. moscow: nauka, (in russian). copyright © 2020 assa. adv syst sci appl (2020) https://dx.doi.org/10.1016/j.cam.2003.03.001 https://dx.doi.org/10.1016/j.cam.2008.04.016 https://dx.doi.org/10.1137/040603425 https://dx.doi.org/10.1016/j.cam.2010.01.028 https://dx.doi.org/10.1007/s11075-010-9412-y https://dx.doi.org/10.1016/j.jmaa.2013.07.039 https://dx.doi.org/10.1007/s00450-009-0083-7 https://dx.doi.org/10.1007/978-0-387-74740-8 https://dx.doi.org/10.1021/ie0601605 https://dx.doi.org/10.1016/s1568-4946(03)00007-3 https://dx.doi.org/10.1007/s10898-008-9332-8 https://dx.doi.org/10.1007/978-94-015-9490-5 70 a. beklaryan 23. zarodnyuk, t. s., anikin, a. s., finkelshtein, e. a., beklaryan, a. l., & belousov, f. a. (2016) the technology for solving the boundary value problems for nonlinear systems of functional differential equations of pointwise type. modern technologies. system analysis. modeling, 49, 19–26, (in russian). 24. beklaryan, l. a. (2007) introduction to the theory of functional differential equations. group approach. moscow: factorial press, (in russian). 25. beklaryan, l. a. (1991) a method for the regularization of boundary value problems for differential equations with deviating argument. soviet math. dokl., 43, 567–571. 26. zarodnyuk, t. s., gornov, a. y., anikin, a. s., & finkelstein, e. a. (2017) computational technique for investigating boundary value problems for functionaldifferential equations of pointwise type. in proc. of the viii international conference on optimization and applications (optima-2017), petrovac, montenegro, october 27, ceur-ws.org. 27. gornov, a. y., zarodnyuk, t. s., madzhara, t. i., daneeva, a. v., & veyalko, i. a. (2013) a collection of test multiextremal optimal control problems, 257–274. springer. 28. gornov, a. y., zarodnyuk, t. s., finkelstein, e. a., & anikin, a. s. (2016) the method of uniform monotonous approximation of the reachable set border for a controllable system. journal of global optimization, 66, 53–64, https://dx.doi.org/10.1007/s10898-015-0346-8. 29. gornov, a. y. & finkelstein, e. a. (2015) algorithm for piecewise-linear approximation of the reachable set boundary. automation and remote control, 76, 385–393, https://dx.doi.org/10.1134/s0005117915030030. 30. floudas, c. a. & pardalos, p. m. (1990) a collection of test problems for constrained global optimization algorithms. berlin, heidelberg: springer-verlag, https://dx.doi.org/10.1007/3-540-53032-0. 31. beklaryan, l. a. & beklaryan, a. l. (2017) traveling waves and functional differential equations of pointwise type. what is common? in proc. of the viii international conference on optimization and applications (optima-2017), petrovac, montenegro, october 2-7, ceur-ws.org. copyright © 2020 assa. adv syst sci appl (2020) https://dx.doi.org/10.1007/s10898-015-0346-8 https://dx.doi.org/10.1134/s0005117915030030 https://dx.doi.org/10.1007/3-540-53032-0 introduction statement of the initial boundary-value problem existence and uniqueness theorem for the basic initial-boundary value problem variational approach to constructing solutions to the basic initial-boundary value problem software package optcon-f existence and uniqueness theorem for soliton solutions in the problem of longitudinal vibrations of an infinite homogeneous rod approximation of the solution of a functional differential equation defined on the line by solutions of an initial-boundary value problem with expanding intervals of the definition construction of the complete family of traveling wave type solutions in the frenkel-kontorova model numerical experiments conclusion мягкая вероятность ошибки при передаче информации adv syst sci appl 2017; 4; 14-21 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/482 selective attention compensates detrimental effect of auditory noise on performance in posner experiment borzou alipourfard1, kamyar arbabifard2, fatemeh bakouie3*, somayeh sadat hashemi kamangar3, shahriar gharibzadeh4 1 department of bioengineering university of texas at arlington, arlington, united states 2 department of computer engineering, amirkabir university of technology, tehran, iran 3 institute for cognitive and brain sciences, shahid beheshti university, tehran, iran 4 department of biomedical engineering, amirkabir university of technology, tehran, iran abstract: environmental noise has a degrading effect on individual’s capacity to pay attention. in this paper, we revealed that the individual’s impaired performance under noise can be counterpoised by recollecting and recalling his attention. we studied the effect of auditory white noise on individual’s performance in posner experiment; and examined if actively reorienting the individual’s attention to stimuli would mitigate the effects of auditory noise. our subjects, ten undergraduate students, participated in the classical posner experiment, while auditory white noise was being played for them through earphone. we varied the intensity of the auditory noise and studied subject’s performance in three conditions: neutral, valid cued, and invalid cued trials. our results showed that the presence of auditory noise degraded subject’s performance in neutral trials. however, subject’s performance was not affected by auditory noise in valid cued trials. keywords: auditory white noise, attention, posner experiment. 1. introduction most people prefer a quiet place for work; they commonly believe that in the presence of environmental noise, their performance is negatively impacted; and they have had experiences that seem to prove this point. thus, studying why environmental noise may affect performance on a task and how we can mitigate its effects is not only of theoretical importance, but is also of practical application and importance. there have been many studies concerned with various effects of auditory noise. some have considered the effects of noise in residential areas. for instance, auditory noise, such as road traffic, has been proved to have an unhealthy effect on residents sleeping patterns or children learning capabilities [1]. there are also studies concerning the detrimental effects of noise in working environments and the physiological and psychological consequences of long exposure to noise on workers [2]. at the physiological level, noise has been seen to cause increased heart rate and blood pressure [3]. furthermore, the neural correlation of the effects of background noise have been reported as increased direct-coupled potential, as induced decrease in brain alpha wave, or as increased dopamine turn over in pfc [4,5]. in experimental settings, the immediate effects that auditory noise may have on an individual’s performance have been examined. in most of these studies, subject’s performance under noise condition was compared with those recorded in standard condition; and impacts of noise, at a psychological and behavioral level, have been described in terms of an overload framework, arousal model, inverted-u hypothesis and stochastic resonance [6-10]. interestingly, the effects of noise on performance have been seen to be consistent with that of performance under induced arousal [8, 9]. that is, under the effect of noise, a subject’s * corresponding author: f_bakouie@sbu.ac.ir http://ijassa.ipu.ru/ojs/ijassa/article/view/482 mailto:f_bakouie@sbu.ac.ir 15 b. alipourfard, k. arbabifard, f. bakouie, s. sadat, h. kamangar, s. gharibzadeh copyright ©2017 assa. adv. in systems science and appl. (2017) performance in low priority task components is degraded while his performance in high priority task components is unaffected; and the subject also shows increased selectivity in utilizing cues [6-9]. although most studies have observed that noise induces arousal, it is to mention that not all the studies have observed such effect: there has been cases in which noise has induced simply a vegetative stress state, or has even improved a subject’s accuracy and sensitivity to stimuli [10, 11]. in this paper, we aimed to study why environmental noise may affect an individual’s performance on a task and how we can decrease this effect. to find an answer for the proposed question, we noticed that many of the previous studies have attributed changes in subject’s performance (under noise) to changes in attentional mechanisms (such as narrowing of attention) [7]. an intuitive thinking suggests that this conjecture is consistent with our personal experience: most of us find it more difficult to focus in a noisy environment. if such a hypothesis is true and that noise disrupts subject’s attentional mechanisms then we expect that with the aid of some other mechanism that captures, recalls and reorients individual’s attention, the impaired performance under noise can be recovered. a suitable experimental setting to determine if recollecting an individual’s attention would mitigate the effects of auditory noise on his performance is posner experiment. in posner task, the experimenter is able to orient and manipulate subject’s attention with the aid of visual cues of his own design. in such an experimental setting, we can study if reorienting an individual’s attention, with the aid of posner cues, compensates the detrimental effect of noise on his performance. but before that, we first have to show that noise does have a detrimental effect on a subject’s performance in posner experiment. posner experiment was one of the early demonstrations of selective attention. in a visual detection task, posner, prior to the onset of the detection stimuli (a black dot), used a cue (an arrow) to indicate the stimuli location to subjects [12]. posner showed that orienting subject’s attention to stimuli location (pre-cueing) decreased subject’s response time, facilitating his perceptual processing. furthermore, posner task has since provided a strong framework for many computational models of attention [13]. dosher and lu studied a subject’s performance in posner experiment, when the detection stimulus (visual stimulus, a dot) was embedded in different levels of background visual noise. through this paradigm they showed that visual noise had a detrimental effect on a subject’s performance and orienting subject’s attention to stimuli location mitigated this detrimental effect; hence they discovered a new attentional mechanism, which they called external noise exclusion. it is to point out that in such an experimental setting we would expect the noise to be harmful to a subject’s performance as it affects the visual information he requires to complete the task. on the other hand, it seems that auditory noise does not affect any sensory information an individual requires to complete the posner task. now we are encountered with two questions: 1) does the auditory noise affect subject’s performance in posner experiment detrimentally? 2) does recalling and reorienting individual’s attention mitigate the detrimental effect of auditory noise on subject’s performance in posner experiment? in this study we design an experiment to answer these two questions that leads the two novelties in comparison with dosher and lu study, i.e. i) the presence of auditory noise degraded subject’s performance; ii) orienting of attention disrupts the detrimental effect of auditory noise on subject. it should be mentioned that in the study of dosher and lu, both the stimulus and noise are visual; while in this study, we designed an experiment with visual stimulus and auditory noise. in this study, our subjects participated in the classical posner experimental setting while different levels of auditory white noise were played through an apple earpods earphone. we varied the intensity of the noise and studied subject’s performance in three conditions; namely: neutral, valid cued and invalid cued trials. through this methodology, we examined interactions between auditory noise and oriented attention, and we sought to learn whether actively reorienting the individual’s attention to stimuli location would mitigate the effects of auditory noise or not. with selective attention compensates detrimental effect of auditory noise 16 the intuitive thinking we mentioned before, we expected to observe an antagonistic interaction between noise and oriented attention in subject’s performance. 2. methods ten undergraduate students, all male, aged between 19 and 22, and with no observable medical conditions that could affect their performance in the task, participated voluntarily in this experiment. the subjects each sat on a chair placed in front of a lcd monitor and at a distance that each found to be comfortable. each subject was instructed to fix his gaze at the center of the fourteen-inch monitor, which had a white background. we asked each subject to press a key as quickly as possible upon detecting the stimulus. the detection stimulus, a black dot, occurred at either the left or the right of the fixation point. the detection stimulus was offset with respect to the fixation point; the offset was set at 600 pixels. our experiment consisted of 5 blocks of trials for five different levels of auditory noise intensity (specifically, no noise, 70 db, 85 db, 100 db, 115 db). white noises of different intensities were generated using matlabr2010a; the noise intensity levels were measured by an hm-303-6 voltage meter. in the study of arnsten & goldman-rakic, it is mentioned that the background noise intensity is 60-70 db; the loud noise intensity is more than 95 db and the noise intensities more than 115 db is annoying. as we wanted to cover all the noise intensities, from background noise to loud and annoying one, we chose the intensity levels of 70 db, 85 db, 100 db, 115 db. the generated noise was then fed to an apple earpods earphone to produce the desired level of auditory noise intensity for each block of trials. to prevent noise intensity adjustments from affecting the subject’s decisional criteria, the order in which these blocks took place in the experiment was designed to be random. each block of trial then consisted of ten cued trials and ten neutral trials. so, the total number of collected dataset for each subject is 5×10×2=100. at the beginning of each trial, subjects were presented with either a plus sign (neutral trials) or an arrow pointing to the right or left (cued trials). as shown in fig. 1, if the plus sign was presented for the trial, then the detection stimuli occurred with equal probability to the left or right of the fixation point. for the cued trials, the detection stimuli occurred at the indicated location with the probability of 0.8 (valid trials) and occurred at the opposite side with the probability of 0.2 (invalid trials). detection stimuli occurred 400 ms after the onset of either the plus sign or the arrow. the time between trials was set at three seconds. the subject’s performance was measured in each trial as indicated by his response time. the described framework was developed virtually in java environment. before the official experiment commenced, subjects were instructed on the paradigms and methods of the experiment; and each subject was allowed to practice the experiment for as many times as he desired, until he found himself comfortable with the process. finally, in the actual testing process, care was taken to provide each subject with a comfortable and a stress-free environment. all subjects were tested during the period between 14:00 and 18:00. 3. results the results of our experiment are shown in fig. 2. the upper plots show the statistical distribution of subjects’ response time (rt). in the lower plots of this figure, the average response time is shown as a function of noise intensity. the leftmost plots in this figure show the results for neutral trials; the middle plots are for the valid cued trials and the rightmost plots show the response time for invalid cued trials. data is plotted using logarithmic scale for noise intensity (x axis) and linear scale for response time (y axis). to check the accuracy of the collected data, we measured the difference in the average response time when the stimuli occurred at the left of the fixation point and when it occurred at its right; in an accurate data set, we would expect this difference to be very close to zero. in our experiment, this difference accumulated to only 2 ms. 17 b. alipourfard, k. arbabifard, f. bakouie, s. sadat, h. kamangar, s. gharibzadeh copyright ©2017 assa. adv. in systems science and appl. (2017) fig. 1. organization of posner task. if the plus sign was presented for the trial, then the detection stimuli occurred with equal probability to the left or right of the fixation point. for the cued trials, the detection stimuli occurred at the indicated location with probability of 0.8 and occurred at the opposite side with the probability of 0.2. a preliminary examination of the data in fig.2 shows that valid cue facilitates performance in all noise conditions while invalid cue does not facilitate performance. this observation is more prominent at higher noise intensities: in 115 db intensity condition, invalid cueing degrades performance while accurate cueing improves performance. as the data in fig. 2 shows, effects of noise on performance in neutral trials and valid cued trials are unalike. on the one hand, for neutral trials, the response time is an increasing function of noise intensity; in neutral trials, as the leftmost plots in fig. 2 show, the response time increases from 721 ms at no noise condition to 738 ms at noise intensity of 115 db. f-test was conducted by r3.1.1 software to compare subject’s response time in these conditions. there was a significant effect of noise on subject’s response time, f(1,9) = 5.64, mse = 1896, p <0.05. on the other hand, for the valid cued trials, as shown in the middle plot of fig. 2, the response time is seemingly independent of noise intensity, exhibiting a total variation of only 9ms. interestingly, in valid cued trials, by conducting f-test to compare subject’s response time in no noise condition and 115db noise condition, we see that there was no significant effect of noise on subject’s response time, f(1,9) = 0.33, mse = 2938 ns. as demonstrated in the right most plot of fig. 2, the response time for invalid cued trials is also an increasing function of noise intensity. specifically, in these instances, the increase in response time is greater as compared with neutral trials. the ample increase in subject’s response time in the 115db noise condition shows that the degrading effect that noise may have on subject’s performance can be as detrimental as the effect of removing valid cues: in the no-noise condition, we see that valid cueing improves performance by 34.2 ms and 115 db noise condition increases subject’s response time by 29.8 ms. in addition, the detailed data of fig. 2 is presented in table 1. in this table, we have provided the exact values of means and sds (std) of cued and neutral trials. on average, the difference between the means of cued and neutral trials as calculated from the table below is 48.50 ms. to further expand upon our hypothesis, we calculated differences in subject’s rt in neutral and valid cued trials and plotted them against noise intensity. the result is presented in fig. 3. selective attention compensates detrimental effect of auditory noise 18 fig.2 recorded response time is plotted against noise intensity level. statistical distributions of the collected rts are shown in the upper plots of this figure. the average rt is plotted in the lower figure. plotted data in the leftmost figures corresponds to the data collected in the neutral trials, while the middle figures and the rightmost figures show the results for the valid cued trials and invalid cued trials respectively. the data is plotted using logarithmic scale for noise intensity (x axis) and linear scale for response time (y axis) table 1. the exact values of means and sds (std) of cued and neutral trials noise level 1 2 3 4 5 mean rt in cued trials (std) 686.82(51.4) 684.77(48.8) 684.86(45.5) 678.68(44.6) 681.28(56.7) mean rt in neutral trials (std) 721.83(39.5) 724.05(46.5) 734.79(50.5) 734.66(49.8) 743.60(51.4 ) we then built a simple linear regression model (least square regression) using r3.1.1 software to study the relation between noise and differences in subjects' performance in valid cued and neutral trails. the value of r2 was calculated to be 0.0177. the results are shown in table 2. note the significant p-value for the noise level coefficient. the “intercept” of the calculated linear model shows that cueing, on average, facilitates performance by 31.09 ms. table 2 the coefficient for the regression model. coefficients estimate std. error tvalue pr(>|t|) intercept 31.09 6.32 4.898 1.56e-06 noise level 6.68 2.56 2.606 0.009 19 b. alipourfard, k. arbabifard, f. bakouie, s. sadat, h. kamangar, s. gharibzadeh copyright ©2017 assa. adv. in systems science and appl. (2017) fig.3 the difference in subject’s rt between valid cued and neutral trials as function of noise intensity. the blue line is a least square linear regression model that shows the dependency of the aforementioned data on the intensity level of the auditory noise 4. discussion in this study, our goal was to examine whether the underlying mechanisms by which environmental noise impacts performance, were related to human attentional mechanisms; we examined if the individual’s impaired performance under noise could be counterpoised by recollecting and recalling his attention. our subjects participated in the classical posner experiment, with the addition that different levels of auditory white noise were played through an apple earpods earphone. we varied the intensity of the noise, studied the changes in subject’s response time, and examined interactions between environmental noise and oriented attention. we found that the presence of noise was degrading to subjects’ performances in neutral trials. however, subjects’ performances were not affected by noise in valid cued trials. we have specifically compared our design with that of the dosher & lu 2000 [13]. in order to better capture our experimental setting, we examined a similar experimental setting proposed by dosher and lu. they studied a subject’s performance in posner experiment, when the detection stimulus (visual stimulus, a dot) was embedded in different levels of background visual noise. through this paradigm they showed that visual noise had a detrimental effect on a subject’s performance and that orienting subject’s attention to stimuli location mitigated this detrimental effect. hence, they discovered a new attentional mechanism, which they called external noise exclusion. it is to point out that in such an experimental setting we would expect the noise to be harmful to a subject’s performance as it affects the visual information he requires to complete the task. on the other hand, in our experimental setting, auditory noise does not affect any sensory information an individual requires to complete the visual posner task. hence, we would not expect it to be harmful to his performance. however, presence of noise has been shown to have a degrading effect. the main question that remains is that why noise may affect a subject’s performance. in neutral trials, our results show that noise impacts a subject’s performance negatively; this degrading effect of noise has also been seen in previous researches [8, 9]. on the other hand, for selective attention compensates detrimental effect of auditory noise 20 the valid cued trials we found that a subject’s response time was independent of noise intensity. this suggests that orienting attention to the stimuli location somehow counteracted the detrimental effects of noise. interestingly, we see that this interaction is present at all noise levels. and irrespective of the intensity, orienting the subject’s attention fully counteracts the effect of noise. this persistent antagonistic relation not only suggests that these two events are counteracting each other at some cognitive level [2], but is also an evidence to our main hypothesis since recalling subject’s attention fully negated the detrimental effect that noise had had on his performance. to further expand upon our hypothesis, we calculated differences in subject’s rt in neutral and valid cued trials and plotted them against noise intensity (see figure.3). the “noise level” coefficient can be interpreted as the average increase in a subject’s rt per noise level in neutral trials that was compensated by cueing in valid cued trials. if there was no interaction between noise and oriented attention on a cognitive level, we would expect the “noise level” coefficient to be zero and for cueing to decrease response time on average by 31.09 ms regardless of noise intensity (see table 1). however, our results suggest a very strong interaction (pr(>|t|)=0.009 ) between noise and oriented attention. this result is a strong evidence for the cognitive level presence of auditory noise that is fully nullified when orienting attention. another interesting aspect of the collected data set is concerning a subjects’ performance in invalid cued trials; however, we want to note that the results here are more qualitative than quantitative, since the small collected data for the invalid cued trials does not allow for a statistical inference. our results show that the presence of noise was the most hurtful to a subject’s performance in these trials (see fig. 2). in cued trials, we then have a task component in which performance is continuously worsened when increasing noise intensity (invalid cued trials) and a task component that is not affected by the presence of noise (valid cued trials). previous studies recognized that under noise conditions, a subject’s performance is continuously degraded in low priority task components, while the performance in high priority task components is not affected [8, 9]. in other words, the amount of noise effect on subject’s performance depends on the priority of the task; if the task has high priority, negative noise effect on subject’s attention would be less [8, 9]. taking into consideration the results of such studies, we may conclude that in our experiment, valid cueing attracts attention of the subject and make a priority for the task. subjects’ performance in valid cued trials represents their performance in the high priority component of the task, and in invalid cued trials represents their performance in the low priority task component. hence, we hypothesize that, pre-cueing had a prioritizing effect on subject’s surrounding sensory events. therefore, visual events on the cued side became his sensory events with high priority, while the visual events on the opposite side were given a lower priority. this prioritizing effect favors the hypothesis that pre-cueing orients subjects' attention and the decrease in his response time in valid cued trials is a result of the orientation of attention rather than a change in response criteria. 5. conclusion in this paper we used posner task to study why environmental noise may affect an individual’s performance on a task. we studied the accuracy of the common belief that the presence of noise makes it more difficult to focus and to pay attention; and examined if recollecting and recalling an individual’s attention would counterpoise his impaired performance under noise. our results supported our main hypothesis. they showed that orienting a subject’s attention to stimuli location mostly negated the detrimental effect that noise had on his performance. our study suggests that, it would be possible to decrease the negative effect of environmental noise even if its modality differs from the stimuli's. this would certainly raise the efficiency of individuals and helps them work well and accurately in noisy workplaces. 21 b. alipourfard, k. arbabifard, f. bakouie, s. sadat, h. kamangar, s. gharibzadeh copyright ©2017 assa. adv. in systems science and appl. (2017) references [1] hygge, s., boman, e., & enmarker, i. (2003). the effects of road traffic noise and meaningful irrelevant speech on different memory systems. scandinavian journal of psychology, 44(1), 13-21. [2] broadbent, d. e. (1971). decision and stress. london, uk: academic press. [3] hsu, s.-m., ko, w.-j., liao, w.-c., huang, s.-j., chen, r. j., li, c.-y., & hwang, s.-l. (2010). associations of exposure to noise with physiological and psychological outcomes among post-cardiac surgery patients in icus. clinics, 65(10), 985-989. [4] arnsten, a. f., & goldman-rakic, p. s. (1998). noise stress impairs prefrontal cortical cognitive function in monkeys: evidence for a hyperdopaminergic mechanism. archives of general psychiatry, 55(4), 362-368. [5] briggs, f., mangun, g. r., & usrey, w. m. (2013). attention enhances synaptic efficacy and the signal-to-noise ratio in neural circuits. nature, 499, 476–480. [6] bell, p. a. (1978). effects of noise and heat stress on primary and subsidiary task performance. human factors: the journal of the human factors and ergonomics society, 20(6), 749-752. [7] woodhead, m. m. (1966). an effect of noise on the distribution of attention. journal of applied psychology, 50(4), 296. [8] o'malley, j. j., & gallas, j. (1977). noise and attention span. perceptual and motor skills, 44(3), 919-922. [9] hygge, s., & knez, i. (2001). effects of noise, heat and indoor lighting on cognitive performance and self-reported affect. journal of environmental psychology, 21(3), 291-299. [10] manjarrez, e., mendez, i., martinez, l., flores, a., & mirasso, c. r. (2007). effects of auditory noise on the psychophysical detection of visual signals: cross-modal stochastic resonance. neuroscience letters, 415(3), 231-236. [11] hasenfratz, m., michel, c., nil, r., & bättig, k. (1989). can smoking increase attention in rapid information processing during noise? electrocortical, physiological and behavioral effects. psychopharmacology, 98(1), 75-80. [12] posner, m. i. (1980). orienting of attention. quarterly journal of experimental psychology, 32(1), 3-25. [13] dosher, b. a., & lu, z.-l. (2000). noise exclusion in spatial attention. psychological science, 11(2), 139-146. adv syst sci appl 2018; 03; 39-78 published online at https://ijassa.ipu.ru/index.php/ijassa/article/view/645 original russian text © yu. v. mitrishkin, p. s. korenev, a. a. prokhorov, n. m. kartsev, m. i. patrov, e.a. pavlova, 2018, published in problemy upravleniya / control sciences, 2018, no. 2, pp. 2–30 plasma control in tokamaks. part. 2. plasma magnetic control systems yuri v. mitrishkin 1,2 *, nikolay m. kartsev2, evgeniya a. pavlova1, artem a. prohorov1,2, pavel s. korenev1,2 , m.i. patrov 3 1) m. v. lomonosov moscow state university, faculty of physics, moscow, russia 2) v. a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia 3) ioffe institute, russian academy of sciences, st. petersburg, russia abstract: various systems of magnetic control of plasma position, current, and shape are considered in vertically elongated tokamaks including spherical tokamaks being in operation. the systems of resistive wall modes suppression are described. tokamaks constructions and crosssections, structural schemes of plasma control systems are given; heed is paid to operating principles of plasma control systems; experimental results of plasma control in tokamaks are given. various engineering realizations of plasma magnetic control systems in tokamaks are presented. keywords: tokamak, plasma, plasma magnetic control, plasma position control, plasma current and shape control, resistive wall modes suppression, plasma control systems realization. introduction the first part [1] of the survey was devoted to the general problem of controllable thermonuclear fusion. it covers the key features of tokamaks and components of plasma control systems, describes the constructions of tokamaks. in particular, the diagnostics system of the globus-m tokamak is considered. the experimental data obtained from this tokamak were used in m.v. lomonosov moscow state university and v.a. trapeznikov institute of control sciences to develop original plasma position, current, and shape control systems. the second part of the survey considers plasma magnetic control systems (plasma is a completely ionized gas) in tokamaks, which are distributed dynamic plants with a complex structure and uncertainties subjected to uncontrollable disturbances. plasma in a magnetic field is not in thermodynamic equilibrium, and, as a consequence, is liable to various instabilities, which are the main cause of a relatively slow approaching of plasma parameters to the lawson criterion. the scientific direction related to research, development, and improvement of tokamaks has been significantly advanced in former soviet union under the leadership of academician l. a. artsimovich [2] and then has spread across the world [3, 4]. the first tokamaks had the circular cross section and were meant to wide-ranging high temperature plasma physics research. a general trend toward enlargement of such tokamaks was observed. these tokamaks included devices which were located in the i.v. kurchatov institute of atomic energy (moscow, russia): т-3, т-4, т-7, то-1, т-10 и т-15; in the a.f. ioffe institute (st. petersburg, russia): tuman-3 (toroidal device with plasma magnetic adiabatic heating); and also * corresponding author: yvm@mail.ru 40 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) foreign devices: plt, tftr (test fusion tokamak reactor) (usa), tore-supra (france), textor-94 (germany), ft-u (italy), etc. a specific feature of the next generation of tokamaks consists in the vertical elongation of the cross section. this feature enables to increase the plasma pressure and to raise the heating by the own plasma current [5]. a “price to pay” for these advantages is vertical instability of the plasma caused by the magnetic fields required for elongation. nevertheless, the vertically elongated tokamaks are now the main experimental base for researches on the problem of controlled thermonuclear fusion. these are jet (joint european torus), united kingdom, jt60u, japan, asdex upgrade, germany (max planck institute for plasma physics), diiid, c-mod, usa, tcv, switzerland, compass (compact assembly), czech republic; east (experimental advanced superconducting tokamak), china, kstar (korean superconducting tokamak reactor), south korea. the spherical tokamaks with a small aspect ratio have also appeared: mast (mega-amp spherical tokamak), united kingdom; nstx (national spherical torus experiment), usa; globus-m/m2, ioffe institute, russia. these tokamaks allow to even further increase in the gas kinetic plasma pressure at a specified magnetic field and can lead to an additional reduction in cost of a reactor as a result. plasma as an automatically controlled plant has the following typical characteristics, which create difficulties of both fundamental and technical nature: plasma is a distributed system with an infinite number of degrees of freedom; imperfection of the theoretical models and the fact that processes proceeding in a plasma have not been sufficiently explored lead to considerable uncertainties in the structure and parameters of plasma models; plasma is a non-stationary plant which means that plasma parameters may change by several orders of magnitude for a short period of time when plasma is produced and heated at one operational cycle or experiment; plasma can be non-minimal-phase plant, since the transfer functions on certain control channels, under the assumption that the parameters are replaced by the fixed values, may have both poles and zeros with a positive real part; plasma is also subjected to uncontrollable disturbances, which in some cases can be estimated in re al time by observing the inputs and outputs of the plant; plasma is a source of broadband noises which are not sufficiently explored. this leads to difficulties in plasma parameters identification; plasma is a nonlinear dynamic plant by its nature; large values of natural frequencies of plasma oscillations require high speed of response and significant power of control systems; actuators, which are generating the control signals sent to the plasma, may contain energy converters with nonlinear characteristics (often discontinuous), dead-zones, and transport delays. this makes it significantly difficult to design and analyze closed loop plasma control systems. the complexity of dynamics and nonlinearity of the actuators serve as an additional source of uncertainties when building a controlled plant (plasma in a tokamak) model; diagnostic units in thermonuclear devices, in many cases, have uncertainties when identifying plasma, which also provide contribution to the total uncertainty of plasma models. despite the presence of peculiarities of plasma, which characterize it as one of the most complex controlled plants in nature, the automatic feedback control systems started being used in the 1960s for plasma confinement in the magnetic traps, and then started to play the significant role in controlled thermonuclear fusion. the investigations along this direction were begun in 1967-1968 by doctors in physics and mathematics v.v. arsenin and v.a. chuyanov in experiments on the trap with magnetic mirrors ogra-2 in the i.v. kurchatov institute of atomic energy. in the experimental device ogra-2 the flute and ion-cyclotron (kinetic) plasma instabilities have been suppressed and then the control systems have obtained 41 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) a spread to suppress other instabilities: drift, ionization, helical instabilities in tokamaks, to stabilize -pinches etc. in [6] a detailed overview of the results of researches and experiments on suppressing plasma instabilities is presented. initially, the main control problem in tokamaks was to stabilize the position of a plasma ring along the major radius by poloidal (lying in a meridional plane) magnetic field. the first experiments aimed to solve this task were carried out on the to-1 tokamak using an impedance controller in 1971 jointly by the staff of the i.v. kurchatov institute of atomic energy (l.n. artemenkov, i.n. golovin, et al.) and the v.m. glushkov institute of cybernetics of academy of sciences of the ukraine soviet socialist republic (yu.i. samoilenko, v.f. gubarev et al.) [7]. combined plasma equilibrium control is used in modern tokamaks, namely: preprogrammed reference control provides a scenario and correction of the plasma location is performed by a feedback control system. these systems became standard components of tokamaks and found application for the joint stabilization of orthogonally decoupled stable horizontal and unstable vertical plasma position in tokamaks. then this approach has started to be used for plasma shape control by means of poloidal field coils and in fact the controlled plant started to belong to a class of multivariable systems, these include the plasma in iter [8, 9]. the use of automatic control methods to ensure the stability and equilibrium of plasmas in thermonuclear devises with magnetic confinement became a generally recognized necessity. experiments on tokamaks have shown that the main parameters of plasma, which directly provide creation of the conditions for initiation of thermonuclear reaction, are solely sensitive even to slight displacements of the external magnetic surface of plasma column in relation to the vacuum vessel or diaphragm limiting the column. therefore, the accuracy of the plasma equilibrium control enables to reduce the rate of harmful impurities inflow into the plasma as well as loss of particles that makes it possible to increase attainable values of plasma density, temperature, and energetic lifetime [2, 7-9]. the position of the plasma boundary is stabilized as close as possible to the first wall in order to ensure the efficient use of the internal space of the vacuum vessel and also to reduce the increments of unstable displacements of the vertically elongated plasma. the smallest failures in the plasma control system can cause the melting of the vacuum vessel. uncontrollable contacts of the plasma with the vacuum vessel lead to the plasma powerful energy release outward, and also result in excessive mechanical loads and damage of the thermonuclear device. these failures are totally unacceptable for the thermonuclear reactor. apart from the plasma magnetic control systems [8, 9], plasma kinetic control systems are also being developed [10, 11] that enables to control the profiles of plasma parameters: plasma current, safety factor, temperature, density, and pressure, and also the burning power during thermonuclear reaction. these systems are necessary for obtaining the most profitable (optimal) operation modes of future thermonuclear reactors. the research work aimed at developing and implementing of plasma control systems in the soviet union was commenced in the laboratory of professor, doctor of engineering sciences lev n. fitsner in v.a. trapeznikov institute of control sciences since 1973. the introduction of developed plasma control systems to physical experiments and their numerical investigation on the distributed parameter models of the plasma were done in cooperation with collaborators of the i.v. kurchatov institute of atomic energy (moscow), troitsk institute for innovation & fusion research (triniti) (troitsk, moscow region), a.f. ioffe physicaltechnical institute (st. petersburg), the d.v. efremov institute of electro-physical apparatus (st. petersburg). general requirements for plasma control systems in tokamaks and general design methodologies of new plasma control systems have not been developed so far. each scientific group who works on operating tokamak designs plasma control systems depending on device functionalities and the type of tasks to be solved. since tokamaks have different configurations of poloidal systems and energy resources of power supplies, different plasma control systems 42 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) are obtained. presently, at the international conference on decision and control (cdc) the sections devoted to plasma control systems have been organized on the initiative of american specialists. there are also the papers on that subject presented at the ifac world congress in 2014 (references are in the survey). the papers have been published by the v.a. trapeznikov institute of control sciences of the russian academy of sciences at the cdcs [12–18] and ifac congresses [19–21]. 1.control systems of plasma position in all operating tokamaks there are control systems of a plasma position, namely, stabilization systems of the plasma magnetic axis or the center of the plasma current in various implementations. several russian and foreign developments will be showed in greater depth. 1.1. the t-14, tuman-3 and tvd tokamaks (russia) in the early national projects a number of plasma position control systems were developed: t14 (tokamak with a strong field, triniti, troitsk), tuman-3 (toroidal installation with magnetic adiabatic heating, ioffe institute, st. petersburg), tvd (elongated tokamak with a divertor, kurchatov institute, moscow). control systems were developed for the t-14 and tuman-3 devices, modeled and implemented in practice of physical experiments on tuman-3: control system that evaluates and compensates the external disturbances during stabilization of a major radius of the plasma column [22-25], as well as an adaptive self-oscillating system for stabilization of the horizontal plasma position, that minimizes the amplitude of self-oscillations at each quasi-period under variable parameters of the controlled plant [26-30]. the additive perturbation and two variable parameters of the model were estimated on line by the adaptive kalman filter [31]. a two-loop orthogonally decoupled self-oscillating system for stabilization of the horizontal and vertical plasma positions with thyristor voltage inverters as actuators was developed and applied in experiments [32-35]. 1.2. the globus-m tokamak (russia) the block diagram of the stabilization system for the plasma vertical position of the globusm tokamak is shown in fig. 1 [36]. the control system includes an actuator based on a current inverter [37], loaded on the control coil of the horizontal field. the control coil generates a radial magnetic field that is proportional to the current passing through it, and has an impact on the plasma column (controlled plant) and shifts it in the vertical direction z. the feedback loop of the system is closed through a controller that implements a proportionally-differentiate control algorithm (see fig. 1 for designations) α(ε ε)contr du t   , where (t)=z(t)–zref(t) is the error, zref(t) is the vertical displacement of the plasma. the error  is formed as  ε p ref p r ref pzi z i k t z i     , where r(t) is the radial flux, ip is the plasma current, k is the proportional gain. the results of the work of this version of the stabilization system of the plasma vertical position are shown in fig. 2 (the case of plasma stabilization in the equatorial plane of the tokamak at zref = 0). it can be seen from the oscillogram r(t) that plasma vertical stabilization system holds the plasma in the equatorial plane (r=0) with good accuracy until the end of the discharge pulse. the calculated elongation k of the plasma, that is determined by the efit code, at the end of the discharge pulse equals to two: k = 2. 43 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) br ÷ u + ccci actuator pdcontroller ucontr plasma in tokamak × kψr z zref ip ɛ fig. 1. the block diagram of the stabilization system for the plasma vertical position of the globus-m tokamak: cc – control coil; ci – current inverter; ψr – radial flux [36]; br – induction of a radial magnetic field fig. 2. oscillograms ip(t), i(t), r(t): discharge pulse № 10446 [36] 1.3. the jet tokamak (united kindom) the jet (joint european torus) tokamak [38] is one of the largest operating machines in the world. its magnetic configuration is close to the architecture of the iter tokamak project. in the jet tokamak magnetic control system there is a plasma adaptive vertical stabilization system, which stabilizes the vertical plasma velocity about zero. it consists of three main subsystems: the vertical stabilization controller, the controller of the actuator current and the adaptive controller. four poloidal coils connected to a fast frfa amplifier, that are capable of switching between nine output voltage modes during time of the order of 200 μs according to the non-linear law with hysteresis zones, are used as an actuator. the measuring system (speed observer) calculates the plasma vertical speed on the basis of the ampere's law and the approximation of the vertical moment of the total current as a weighted sum of measurements of magnetic field outside the plasma [38–40]. the controller of vertical stabilization is a proportional unit. however, the behavior of the system closed by such controller will be similar to the case of a relay controller due to the presence of dead zones and hysteresis in the model of the actuator: as soon as the plasma velocity becomes greater than the threshold value, a voltage appears at the output of the actuator, a current creating a force acting on the plasma in the opposite direction increases rapidly in the coil. when the plasma velocity passes through zero, the voltage at the output of the actuator disappears. further, when the opposite threshold value is reached, the reverse voltage appears. thus, an oscillating process arises. the use of a single closed stabilization loop of the plasma vertical velocity around zero cannot guarantee the stability of the plasma in the vertical direction [40]. the process will oscillate around zero, however, a drift of the average values of the plasma vertical position and current in the control coils is possible, since the control is carried out by speed, and not directly ε 44 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) by position. the problem of vertical stabilization is solved jointly with the plasma shape control system, which gives a guaranteed stable closed control system. fig. 3. stabilization system of the vertical plasma velocity around zero on the jet tokamak: frfa is fast radial field amplifier [38] it is also necessary to avoid exceeding the current limit value in the control coil and the output of the frfa amplifier in a saturation mode. a slow pi controller, which stabilizes the current in the control coil around zero, is used to solve this problem. due to the non-minimal phase nature of the vertical plasma motion model [40] the gain in this loop is negative. the risk of overheating of the frfa amplifier and the release of heat on it depend on the switching frequency, these determine the limit of the permissible gain in the stabilization loop of speed around zero. this limitation is further enhanced by the fact that the vertical speed is calculated from the magnetic measurements, i.e. the sensors are connected to the field of the control coil current, which creates additional oscillations. the model of the vertical plasma velocity varies substantially during the discharge. in addition, a significant uncertainty of the model arises when linear approximation of the actuator, working on complex nonlinear principles, is done. thus, it is impossible to use one set of stationary parameters of the controllers throughout the whole discharge, so the adaptation of the parameters of the controller of vertical stabilization is used on the jet tokamak. there is an approximation of the dependence of the frfa switching frequency fsw on the gain kvs in the stabilization loop of the vertical velocity and on the module  of a single unstable pole of its linear model [38-40]: fsw=γf(kvs), where f is a monotone function. an adaptive controller operates on the basis of this ratio, leading away the frfa mode from saturation. the switching frequency of the frfa in the jet tokamak is about 500 hz, which ensures that there is no overheating and provides small amplitude of the plasma oscillations over the entire range of the module of the unstable pole γ of the linear vertical plasma motion model. 1.4. the east tokamak (china) stabilization of the vertical position and plasma velocity around zero on the east tokamak (experimental advanced superconducting tokamak) is given serious attention, since vertical instability is a risk for the operation of tokamaks. in the east, an algorithm called rzip [41] is used to control the current and position of the plasma. the block diagram of the plasma position and current control system on the limiter phase is presented in fig. 4. the value of the plasma current is measured by the rogowski loop, and the estimations of the vertical and horizontal positions of the magnetic axis are reconstructed from the signals of the magnetic diagnostics outside the plasma. the feedback is closed via pid controllers and matrices m_matrix(rp, zp) and m_matrix(ip) [41]. the control is carried out by means of currents in poloidal field coils, for which the target values consist of the sum of the scenario signals and signals of the closed-loop control system. the coils of the central solenoid are used 45 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) to control the plasma current, the vertical and horizontal positions of the plasma are controlled by the coils of the poloidal field. additional coils connected to high-speed power supplies and located inside the vacuum vessel are used to suppress the vertical plasma instability, as well as in the iter project. pf error east command to pf power supply pf current control pf measured current pf current trajectories м-matrix ip м-matrix rp, zp δpf1~ δpf14 δpf11~ δpf14 pid pid ip error + + ip target ip measured current command to ic power supplypidfilter rp, zp error zp error rp, zp target e-matrix rp, zp magnetic diagnostic data fast z control ic current + + fig. 4. the block diagram of the plasma current and position control system on the east tokamak on the limiter phase of the discharge [41] (ic is an active feedback coil inside the vacuum vessel) the results of tracking the vertical and horizontal position of the plasma in the east for the target in the form of triangular pulses in the real experiment, discharges 10112 and 10113 respectively, are shown in fig. 5. the control error for the vertical position does not exceed 1 mm. the control contour of the horizontal position tracks the target noticeably more slowly due to the effect of the field penetration into the conducting structures of the vessel during horizontal displacement. the control error does not exceed 2 mm after the end of the transient process. fig. 5. tracking the vertical and horizontal position of the plasma on the east tokamak [41] the italian create group and chinese specialists during the experimental campaign of 2016 on east developed and applied the stabilization system of the plasma vertical velocity around zero, that was frequency decoupled with the plasma current and shape control system [42]. the control directly of the plasma position was conducted using the plasma shape control system. the controller scheme of this system is shown in fig. 6, it is similar to the plasma control system used on the jet tokamak [38-40]. the vertical plasma position control system in the east tokamak is modeled on the tsc (tokamak simulation code, usa) plasma-physical code. it is claimed in [43] that the results of modeling on the main parameters, such as plasma current, plasma shape and position, flux contours and magnetic measurements, coincided well with the experimental data. 46 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) plasma shape and position control isoflux control mode rzip control mode pf coils current controllerplasma current controller 12 siso pidcontrollers ipf + iff m(·, ip)pid ipf, ref1 ipf, ref ip, ref ip + m(·, zp; rp) 2 siso pidcontrollers rp, ref zp, ref rp, zp + m(·, seg1÷13) 13 siso pidcontrollers δψcp m(·, zx1; rx1) 2 siso pidcontrollers rx1, ref zx1, ref rx1, zx1 + m(·, zx2; rx2) 2 siso pidcontrollers rx2, ref zx2, ref rx2, zx2 + m(·, sym)pidsymref sym + + ipf, ref2 control mode selection fig. 6. the simplified scheme of the position, current and plasma shape controller on the east tokamak: m matrices are used to distribute signals between 12 current control loops in poloidal field coils; the target currents are monitored by the current controller in pfc coils [42] 2.plasma position, current, and shape control systems all tokamaks given in table 2 of the first part of this survey [1] have a common similarity consisting in the vertical elongation, but all of them differ from each other by the poloidal systems, that are the magnetic coils systems creating poloidal fields. the jet tokamak has an iron core that other tokamaks do not have. all the rest of tokamaks have an air central solenoid. the iron core makes additional problems associated with its nonlinear magnetization curve. tokamaks diii-d, nstx, jt-60u, and tcv have the poloidal field coils inside the toroidal field coil; other tokamaks on the contrary have the poloidal field coils outside the toroidal field coil. the east and iter tokamaks have all superconductive coils. these tokamaks have the horizontal field coils inside the vacuum vessel in order to stabilize the unstable plasma vertical position, which significantly expands the area of controllability and stability on the vertical coordinate with limited power supply sources of these coils. in all tokamaks, the coils of the poloidal field are located differently in the space around the vacuum vessel. such difference of the poloidal systems of known tokamaks leads to the different configurations of the plasma position, current, and shape control systems. 2.1. diii-d tokamak (usa) a general view of the diii-d tokamak is shown in fig 7, a and the magnetic configuration of the plasma in it is shown in fig. 7, b [44]. on the diii-d tokamak the magnetic flux is controlled at 13 points on the separatrix (fig. 7, b). the block diagram of the magnetic plasma control in the diii-d tokamak is shown in fig. 8, a, fig. 8, b shows the result of tracking the upper and lower gaps between the separatrix and the first wall, as well as the x-point coordinates. 47 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) a b fig. 7. diii-d tokamak: а – general view (internet resource), б – poloidal flux contour lines (configuration with lower х-point) [43] a b fig. 8. plasma control on the diii-d tokamak [45]: a – block diagram of the isoflux-control at the 13 separatrix points (see fig. 7, b); rt – real time; b – system response to control of the upper and left gaps between the separatrix and the first wall and x-point position coordinates. 48 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) on the diii-d tokamak, the following ideology of developing plasma position, current, and shape control system is adopted [44]. the vertical plasma instability is suppressed through a separate circuit, which stabilizes the plasma vertical velocity around zero (see fig. 8, a). a set of coils is used to control the plasma current. the 18 poloidal field f-coils are used to control the plasma shape. the isoflux-control method is applied: magnetic flux at 13 points on the plasma separatrix, as well as vertical and horizontal position of the x-point are calculated in real time using signals of magnetic diagnostics of currents, fluxes and fields by the efit algorithm. the error signals between the specified (reference) values of magnetic flux on the separatrix and the x-point coordinates are fed to the multivariable isoflux controller (mimo controller). the isoflux controller also receives signals of currents in the control coils. the output signals of the controller come to the inputs of the actuators, thereby closing the multivariable feedback loop. the multivariable plasma shape controller itself is designed on the basis of linear models, which are generated by a special software package named toksys. the set of points is given, which determines the desired location of the plasma separatrix and in which control will be carried out. currents in the poloidal field coils are tuned so as to keep equal magnetic poloidal flux at the x-point and at all the other boundary points. let the difference between the flux at the controlled points and the given flux references at these points be defined as δψ. the relation between the change of currents in the control coils δi and δψ can be represented as i = m–1. here m–1 is a control matrix, which is the inverse matrix to the matrix m composed of the values of the green's function for the grad–shafranov equation. the elements of the matrix m are poloidal flux values at each control point while control currents equal to unit. 2.2.asdex upgrade tokamak (germany) a general view of the asdex upgrade tokamak (modernized axially symmetric divertor experiment) is shown in fig. 9, a, and its cross section is shown in fig. 9, b. the control of the vertical and horizontal position of the plasma, as well as the coordinates of the divertor strike-points is carried out on this tokamak (fig. 10, 11). 2.3.jet (uk) tokamak on the jet tokamak (fig. 12) various control modes are used in conjunction with the vertical stabilization of the plasma, they are the horizontal position of the plasma, the currents in the poloidal fields coils, the plasma current, the coordinates of the divertor strike-points control modes, and the control mode of the plasma shape by the gaps between the first wall and the separatrix [38]. the structure of the controller in jet is founded on physical principles and based on the kirchhoff’s vector equation pf pf pf di v m ri dt   , where vpf and ipf are voltage vectors on the coils and the measured currents in them, m and r are matrices of mutual inductance and resistance, respectively, control of currents in the poloidal field coils is carried out in absolute values. not all coils are used for different control modes, but the most efficient individual coils or sets of coils. control modes can be combined: for example, most discharges can be performed with simultaneous use of modes of the currents in the poloidal field coils control, the plasma current control, the divertor strike-points coordinates control, and plasma shape control by values of the gaps. the control system has a cascade structure [49] shown in fig. 13, where there are controllers for plasma shape, current, and position of divertor strike-points in the block "shape and plasma current controller". 49 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) a b fig. 9. tokamak asdex upgrade: a – general view [46]; b –cross section with poloidal field coils [47]; the control coil for the vertical position of the plasma is placed between the vacuum vessel and the toroidal winding 50 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) references position shape current controllercontroller controller ffc ffc asdex upgrade experiment plasma equilibrium determination ffc icol iv ioh load compensation fig. 10. block diagram of the magnetic plasma control system of the asdex upgrade tokamak [48] fig. 11. plasma position and divertor strike-points location control response on the asdex upgrade tokamak [48] the general control law for controllers [38] is vpf = rest ipf + k (yref – y), where rest is the estimation of the resistance matrix, k is the matrix gain, yref and y are the vectors of the reference inputs and the measured outputs. in the control law structure there is the compensation element rest ipf of the voltage on the resistance rest that makes the behavior of the closed-loop control system of currents in jet coils close to the devices with superconducting coils and eliminates the need to introduce integral components for increasing the degree of astatism and achieve the required tracking error 1-2% [38]. the matrix gain is synthesized in the form of k = h*(mest t –1c –1), where h is the matrix of inputs/outputs choice, mest is the estimation of the mutual inductance matrix used for decoupling of the currents in the poloidal field coils, t is the static gain matrix linking the variations of the currents in the coils with variations of output variables: y = t ipf, and c is the matrix of the desired values of time constants in the closed-loop system. the plasma model in the vessel is taken into account as a coil with a distributed current. in doing so, the plasma resistance and mutual inductance of the plasma with all coils [50], except the coil p1, are neglected, that is justified by the slow dynamics of the plasma current. in practice, only during disruption the plasma can induce a significant voltage on the poloidal field coils. the values of magnetic fluxes at the points of the equatorial plane on the outer and inner circumference are used for the plasma horizontal position control mode. since the plasma separatrix is a line of the equal flux, it is enough to make the difference of fluxes in these points equal to zero by the coil p4 in order to stabilize the horizontal position (fig. 12, b, 14, a). 51 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) а b fig. 12. jet tokamak: a – general view (internet resource), b – cross-sectional view with control coils and central solenoid [38] shape controller system pf current controller ipf (t) + iff (t) shape and plasma current controller iref (t) plasma current and shape references + + plasma / circuits system vpf (t) vertical stabilization system + + zp (t), irfa (t) . ip (t), shape(t) fig. 13. block diagram of the plasma position, shape, and current on jet [49] the xloc algorithm [51] is used to calculate the gaps values between the vessel first wall and the plasma separatrix in the mode of plasma shape control. the gaps response to the test signals in the closed-loop control system on the jet tokamak is shown in fig 14, b when a discharge duration of at least 2 seconds. the poloidal field coils d2 and d3 are used to control the locations of the divertor strikepoints (see fig. 12, b). the signals of location of the intersection points of the separatrix with two vertical or horizontal axes (the specific option is set by the operator) are used as feedback signals. an oscillating mode of control for the location of the divertor strike-points is used for energy scattering, which allows to distribute the energy over a larger area of the divertor. it is implemented by introducing the reference saw-edged signal at frequency of 4 hz while the control loop bandwidth has a value of 10 hz. 52 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) a b fig. 14. tokamak jet: a – vertical cross-section, b – gaps response to the test signals [38] 2.4.tcv tokamak (switzerland) the control method of the plasma shape in the tcv tokamak (variable configuration tokamak) [52] by the values of the magnetic flux and field at certain points on the plasma separatrix is illustrated in fig. 15 [53]. the method is based on the fact that the separatrix is a line of a constant magnetic flux, and the x-point in the divertor configuration is the point of zero field. thus, having chosen the plasma shape scenario in advance, it is necessary to align the flux values at the selected points, reducing the field to zero at the expected x-point. a b fig. 15. tokamak tcv: a – cross-section; b – block diagram of the plasma shape control by flux and field values on the separatrix of the tcv tokamak [53] 53 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) the control scheme (fig. 15, b) consists of four main elements. one of them is the element k, it serves to calculate the deviations of the magnetic flux and field on the separatrix from the scenario values as a linear combination of the measurements of the sensors outside the plasma. another element is a block of pid controllers, one for each deviation signal, which reduces the values of controlled signals to zero. the next element 1m  calculates the derivatives of currents in the poloidal field coils, in which the current deviations will be reduced to zero in the finite time. this calculation is carried out for the model of magnetic configuration without plasma and currents in the passive structures of the vessel. finally, the element l calculates on the basis of the kirchhoff’s equations the required values of the voltages on the poloidal field coils as a function of the currents and their derivatives, as well as the voltage on the plasma ohmic heating coil. the given algorithm does not take into account the model of the plasma in the vessel of the tokamak and, therefore, can be used throughout the discharge; however, it does not allow achieving a high degree of decoupling of control channels. this approach was applied to the kstar tokamak [54], and its modified versions are being used in diii-d, east and alcator c-mod [55]. a b fig. 16. tokamak east: a – construction; b – cross-section [42] 2.5.east tokamak (china) the complete reduced copy of the iter is the chinese east tokamak with superconducting coils (sc) (fig. 16), in which there are control coils inside the tokamak vessel. the system of plasma shape magnetic control for the east, that was developed jointly by the chinese and american specialists, is one of the most advanced. the two algorithms for the plasma shape control on the isoflux-control principle (control of the magnetic flux and field values at certain points on the plasma separatrix) are used. these 54 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) algorithms are iso-elong and iso-dnull to control the elongated limiter configuration and the divertor configuration with two x-points, respectively. fig. 17 shows the basic block diagram of the control system for these algorithms. pf error eastcommand send pf current control pf measured current pf current trajectories pid pid error + + control points x-point error ip, zp ip, zp target e-matrix magnetic diagnostic data ic current м-matrix magnetic diagnostic data rtefit parameter calculation shape trajectories fig. 17. block diagram of the plasma current, shape, and vertical position control systems on the east tokamak for the divertor discharge phase (isoflux control): rtefit (real time equilibrium fitting) [41] a b fig. 18. plasma shape and position stabilization on the east tokamak for discharge 10618: a – magnetic configuration of east at the moment of 4.958 s; b – tracking errors at the reference points and the x-points [41] the positions of characteristic points on the separatrix, such as x-points or points of contact with the limiter, are calculated during the plasma shape reconstruction. the maximum value of the magnetic flux from the characteristic points is selected as the reference for all points. then, the differences of the magnetic flux values are used to close the feedback through the pid controllers (see fig. 17). one or two nearest poloidal field coils control the value at each 55 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) reference point, the plasma current is controlled by the central solenoid coils. the reference currents in the coils are calculated using the matrices of the control channel decoupling, scenario signals are added to them, and they are fed to the multivariable control loop of the currents in the poloidal field coils. the typical value of the unstable pole module on the east tokamak is 200-300 s–1; it can exceed 1000 s–1 for configurations with greater elongation or non-standard plasma parameters. it is necessary to quickly calculate the vertical position of the plasma and use high-speed power supplies as actuators in order to suppress vertical instability with such dynamics. an analogue of the rzip algorithm is used in the control system to calculate the estimation of the plasma vertical position, which is controlled by individual coils located inside the vessel. fig. 18, b shows the signals for discharge 10618 [41] with an elongation of 2.0 and a plasma current of up to 250 ka. algorithm iso-dnull starts working at the moment of 2.7 s and by 3.0 s the differences of values of the magnetic flux are less than 0.001 v·s/rad, and the error of xpoints position control is less than 1 cm. 3.spherical tokamaks 3.1.advantages of spherical tokamaks for successful operation of a thermonuclear reactor the plasma parameters must satisfy the numerous constraints imposed by the magnetohydrodynamics theory. one of the most important constraints is the maximum achievable β which characterizes the efficiency of plasma confinement and is defined as the ratio of the plasma pressure to the magnetic field pressure. numerical calculations [56] show that as the aspect ratio a decreases, the maximum acceptable β increases as 1/ a for conventional tokamaks ( 3a  ) and faster for spherical tokamaks ( 1.5a  ). thus, the best plasma confinement can be achieved on spherical tokamaks. another advantage of spherical tokamaks is the high value of the safety factor q at the plasma boundary. the safety factor increases as 2 3/21/ (1 )a a with decreasing aspect ratio, allowing to suppress the kink instability and enabling higher plasma current than on conventional tokamaks with the same magnetic field strength and small radius of the plasma. finally, relatively small size of the spherical tokamaks makes their creation and operation less costly, and allows for higher values of magnetic and electric fields to be attained with the same currents as in conventional tokamaks. these advantages mark the spherical tokamaks as promising candidates to be the future commercial thermonuclear power plants, which gives more relevance and significance to the plasma control problems on spherical tokamaks. in this context, the results of plasma control for the operating spherical tokamaks mast, nstx, and globus-m are given below. 3.2.mast-u tokamak (uk) the cross-section and structure of the mast-u tokamak (mega ampere spherical tokamak) are shown in fig. 19. the vertical position control on mast-u is implemented via the pd controller [57], while not the plasma vertical coordinate is taken as controlled variable but the product of plasma vertical coordinate and its current (fig. 20, a). like on east and diii-d tokamaks, the isoflux control method is applied for the shape control on the mast-u tokamak, with the real-time rtefit plasma equilibrium reconstruction code calculating flux values at control points [59]. 56 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) a b fig. 19. mast-u tokamak: a – tokamak structure, b – cross-section [57] a b fig. 20. plasma shape and position stabilization on mast-u: a – vertical position control [56]; b – inner and outer plasma radii control, x-point position control [59] 57 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) a b fig. 21. nstx-u tokamak: a – structure; b – cross-section [58] 3.3.nstx-u tokamak (usa) the structure and cross-section of the nstx-u tokamak (national spherical torus experiment) are shown in fig. 21. the isoflux control method [58] is used to control the plasma shape on the nstx-u tokamak. magnetic fluxes in the set of control points are calculated in real time using the rtefit code and via pid controllers are equated to a given flux that is determined by the flux at the x-point or the flux at the contact point of the plasma with the vessel. a b fig. 22. plasma shape stabilization on nstx-u [58]: a – resulting gaps on the outer side of the plasma; b –xpoint position control 58 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) pid controllers are also used to control the position of x-points and strike points. a control contour that minimizes the horizontal distance between the flux isolines of the upper and lower x-points is used for configurations with two x-points. fig. 22a shows the resulting outputs of the control system for gaps between control points on the outer side of the plasma and the walls of the vessel, the result of the x-point position control is shown in fig. 22, b. 3.4.globus-m tokamak (russia) spherical tokamak globus-m (a.f. ioffe physical-technical institute of the russian academy of sciences, st. petersburg) [1] is the only tokamak with a vertically elongated cross-section in the russian federation, fig. 23 [60, 61]. feedback systems are applied on the globus-m tokamak to control the horizontal and vertical position of the plasma with high-speed thyristor current inverters as actuators [37]. there is a set of poloidal field coils on the tokamak that is incorporated into the control loops of currents in these coils with multiphase thyristor rectifiers [72] and pd-controllers. so these loops enable to use program control to control plasma magnetic surfaces in each discharge. that has made it possible to collect a database of plasma discharges, which was used to develop and simulate hierarchical control systems for the position, current, and shape of the plasma with the equilibrium reconstruction codes in the feedback. two codes for the plasma equilibrium reconstruction from magnetic measurements outside the plasma have been developed in the matlab environment: one using picard iterations to solve the grad-shafranov equation by means of the green's functions [18, 62] and one using non-iterative moving filaments method to approximate the plasma current distribution [21]. the plasma equilibria reconstructed from the experimental data were used to construct arrays of linear plasma models from which linear models with variable parameters were created by means of linear interpolation [18, 21]. a b c fig. 23. globus-m tokamak: a – cross-section: – limiter; – magnetic loops; – vacuum vessel; – pf coils; – central solenoid; b – magnetic configuration reconstructed by the moving filaments method [21]; c – magnetic configuration reconstructed by the picard iterations method with the measuring displacements directions of gaps between the tokamak first wall and separatrix[18, 58] systems with time-varying robust h∞ controllers with switching for isoflux control (fig. 24) [21] and linear interpolation for controlling gaps between the first wall and separatrix (fig. 25) [18] were designed for these time-varying plasma models and modeled, using a new technique with simultaneous application of the linear models, experimental scenario signals, and reconstruction codes in the feedback [18, 21, 63, 64]. the combination of the moving filaments reconstruction code with the flux control on the separatrix showed the highest 59 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) computational speed of the system with the possibility of implementing it in real time on the globus-m tokamak using the speedgoat industrial computers (https://www.speedgoat.com) with the simulinkrt operating system from mathworks [21]. lpv model of plasma in globus-m tokamak current inverter current inverter thyristor rectifiers analog controller of the plasma vertical position analog controller of the plasma horizontal position pidcontrollers δz δr ihfc ivfc uhfc uvfc δupf δipf, δics δip switching robust mimo controller kmimo(s) br, bz, δψb ipf, ics, ip,ψ 0 plasma equilibrium reconstruction algorithm ipf, ics, ip, ψ δψ matrix exchanging at tk moment x(tk) vector state matching x(t) controller matrix database hierarchical level of controller switching pid-controllers of plasma current 0 0 0 a b fig. 24. time-varying system with controllers switching and moving filaments equilibrium reconstruction code for poloidal flux on separatrix and magnetic field in x-point control on the globus-m tokamak: a – block-diagram with lpv (linear parameter varying) model, b – deviations δb of magnetic field in x-point and deviations δψ of the poloidal flux difference between points on the separatrix during the shift from limiter to divertor phase of the discharge [21]. lpv model of plasma in globus-m tokamak δz δr ihfc ivfc uhfc uvfc δupf δipf, δics δip h∞ lpv gap g controller δg ipf, ics, ip, ψ δψ δgref _ δg switch s kmimo pid-controller of plasma current thyristor rectifiers mimo pidcontroller current inverter current inverter analog controller of the plasma vertical position analog controller of the plasma horizontal position plasma equilibrium reconstruction code ipf, ics, ip,ψ 00 0 0 a b fig. 25. system with the time-varying controller based on linear interpolation of controllers array and picard iterations equilibrium reconstruction code for gaps control on the globus-m tokamak: a – block diagram, b – deviations of gaps between the plasma separatrix and vacuum vessel [18] it is worth noting that the poloidal flux control in vacuum vessel without plasma was studied prior to the development of shape control system for the globus-m tokamak [64]. 60 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) stabilization of the plasma vertical position with fast contour is considered in [16], where the method for the plasma vertical position adaptation to its shape is proposed. this method was applied to the globus-m tokamak [65]. the block diagram of this application is shown in fig. 26, and the operation of the plasma position, shape and current control system is shown in fig. 27. fig. 26. block diagram of the hierarchical control system for the plasma position, current and shape with the adaptation of the vertical position of the magnetic axis on the globus-m tokamak [61] fig. 27. modeling results for plasma control system on globus-m tokamak: a – plasma vertical position deviation; b – gaps deviation; c – plasma current variation [61] 3.5.t-15m tokamak (russia) during the developing of plasma control system for the t-15m tokamak (national research center “kurchatov institute”) a part of the work was carried out at the v.a. trapeznikov institute of control sciences of the russian academy of sciences [66]. an important result of the conducted investigations is the transfer of the horizontal magnetic field coil, intended for control of an unstable vertical position of the plasma, from the position outside of the toroidal field coil to the position between the vacuum vessel and the toroidal field coil [17]. this is due to the fact that the horizontal field coil in its initial position was shielded by the neighboring poloidal field coils needed for the plasma shape control. because of that the control system acquires the internal instability property. this meant that the plasma position control system required unlimited increase of the current in the horizontal field coil to stabilize the vertical position of the plasma under the minor disruptions. this instability made the system, and, consequently, the entire plant, unworkable. this property of the system was discovered during its development and modeling [67, 68]. while the coil was transferred close to the vacuum 61 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) vessel, but inside the toroidal field coil, the control system became able to achieve internal stability [69] and operability with the unstable plant under control. the model of the vertical plasma motion in the t-15 tokamak for the poloidal system with the transferred horizontal field coil [71] was obtained by the identification of the plasmaphysical code dina (src rf triniti) [70]. after that the plasma vertical position control systems with different actuators have been developed: a multiphase thyristor rectifier and a transistor voltage inverter [72]. in the first case, a modal control system was synthesized with pole placement of the closed-loop system at single point in negative real part of the complex plane for maximum convenience of system tuning, in the second case the system was put into a sliding mode, and revealed its weaker robust properties. a system with an adaptive predictive model for the variable parameter of the plasma model was developed [73], as well as the system with the placement of the poles of the system in the lmi (linear matrix inequalities) -regions to reject the minor disruptions perturbations was designed. the latter system showed the advantages over the modal system in terms of rejection of the external disturbance and accuracy [74]. it has been shown that the plasma shape is controlled in the t-15 tokamak without the use of the horizontal field coil. with the multivariable controller with a state estimator [75] the shape control requires only poloidal field coils and the sections of the central solenoid, like in the iter version of 1995-1997 [9]. for the t-15 tokamak, the plasma shape control system was modeled on the dina code with stabilization of the vertical plasma velocity around zero and the lqg controller in the feedback [76]. 4.magnetic control systems of resistive wall modes let us consider the resistive wall modes (rwm) and methods of their suppression [77]. the general trend seems to be toward an increase of the plasma parameter β = 2μ0

/ b0 2 (

is the average pressure, b0 is the toroidal magnetic field, μ0 is the vacuum permeability) and normalized parameter βn = aβ / (b0 ip) (iр is the plasma current, a is the minor radius) in modern tokamaks. the instability related to the rwm is one of the factors, which limits the growth of  and often leads to disruption of plasma discharge. thus, the low toroidal modes n = 1, n is a toroidal wave number, which might arise with the growing pressure, are of particular interest from the viewpoint of control systems design (the modes of the form ( )( ) i t m nr e     , m is the a poloidal wave number, are considered when examining the rwm, see § 5 of part 1 of this survey [1]). the tokamak conductive wall located close enough to the plasma could provide rwm modes suppression. however, it may only merely slow down the growth of the modes, since the wall with a finite conductivity could suppress the mode only for a certain time roughly the same as the time of magnetic field penetration into the wall. therefore, these modes are called the resistive wall modes. the following system of equations describes the dynamics of rwm using a simple cylindrical model approximation and the assumption of a rigid mode structure and ignoring the effect of plasma rotation [79]: 0, 0, , eff p pw w pc c wp p w w wc c w w cp p cw w c c c c c l i m i m i m i l i m i r i m i m i l i r i v            where , , p w ci i i are the current on the plasma surface, the current in surrounding passive (wall) structures, and the control coil current, respectively, abm represents the mutual inductance between conductors a and b, ar is the resistance in conductor a, al is the self inductance of conductor a, , { , , }a b p w c , effl is the effective self inductance, cv is the 62 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) voltage applied to the control coil. the model has been obtained for a tokamak where the control coils are inside the vacuum vessel (fig. 28). the degree of interactions between currents on the plasma surface, in the surrounding wall, and control coil are characterized by mutual inductances abm . plasma vessel wallcontrol current in coil mpc mwc mpw fig. 28. cross section of a cylindrical model of the rwm dynamics [78] the interconnection of the plasma rotation and the appearance of rwm is not well studied, however, the model for describing this phenomenon was derived in [80] using an argument based on the exchange of energy between the plasma mode and external conductors:   ,iw p pw ww i d b c b    where , p wb b are changes in the field at the plasma surface and at the vessel wall when the plasma is rotating;  is the plasma toroidal rotation frequency; 1 pw pwс m  ; iww is the value of the coupling of rwm energy transferred through the field pb to the toroidally in phase component of the field wb ; φd represents the energy coupled to the component wb that is 90o toroidally advanced, d represents dissipation mechanism. there are two primary approaches to suppressing the rwm. the first one is that of applying feedback magnetic control systems in order to suppress the system’s instabilities. the growth rate of oscillations decreases by an order of magnitude because of the tokamak vessel conductive wall, making it possible virtually to realize the feedback control by the control coils. the second approach is that of rotation stabilization of the plasma. in modern tokamaks a beam of neutral atoms is injected into the system, which transmits an angular momentum that is sufficient for keeping the plasma rotation, which in turn results in rwm suppression. the combination of these two approaches to control the rwm is considered as the most prospective for future thermonuclear reactors. the rwm control systems [81] for the diii-d tokamak (fig. 29) [82] were created using different design methods of controllers. the feedback control underlies in most of the proposed methods. the structural scheme of the closed loop control system is presented in fig. 30. the pd [83], lqg [83], h∞ncf [84], dk [84], etc. controllers are used as controllers in the feedback loop. the functional   0 ( ) t tj u x qx u ru dt    , where 0, 0 q r  are weighting matrices, is minimized when designing an lqg optimal controller (u = –kx is the control law, x is the state). the kalman filter (the observer) was applied to estimate the state [83]. dk-iteration also can be applied to synthesize the controller that is a combination of the н∞-synthesis and μ-analysis. in this case the controller k meets the criterion  1min min ( ) k d dn k d  , where 63 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) n(k) is the transfer matrix of the closed loop system,  is the set of matrices commuting with  ( i   , ( ) 1   ), d d   , where  is the maximum singular value of the matrix. the results of numerical simulation when the rwm growth rate changes are presented in fig. 31. it is shown that the lqg controller effectively suppresses the rwm and allows to increase the stability margin of the system [83, 84]. fig. 29. diii-d tokamak [78] plasma model reference outputfilter controller actuator fig. 30. structural scheme of the rwm control system fig. 31. the outputs of the rwm control system of the diii-d tokamak when the rwm growth rate changes [83,84] the tokamak diii-d construction provides the neutral beam injectors (fig. 32) which allow giving additional energy and momentum to the plasma in the direction opposite to its movement, or transmitting up to 10 mw of power without giving any additional momentum to the plasma [85]. the possibility of decoupling of the channels of transmission of power and momentum makes it possible to suppress the rwm and to slow down the plasma rotation considerably at .n n no wall   , where .n no wall is the value of the n coefficient without the wall. the high values of n are also achieved by two series of control coils, located inside (icoils) and outside (c-coils) of the vacuum vessel. fig. 33, a-e shows that the rwm increase without the i-coils feedback (i-coils are switched off), which results in the plasma disruption. however, the rwm are successfully suppressed in case of the i-coils feedback (i-coils are switched on) (fig. 33, f-j). 64 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) fig. 32. the location of the neutral beam injectors in the diii-d tokamak [85] fig. 33. a comparison of systems: а – e –without feedback control system (i-coils are switched off); f – j – with feedback control system (icoils are switched on) [85] on the other tokamaks, such as nstx [82], iter [83, 84] et al., similar principles are being applied to control the rwm taking into account their specific design features. 5.implementation of plasma control systems 5.1.real time test beds for tokamaks currently, the real time test beds are applied in various fields of engineering, industry, and science that allow to model control systems in real time, enable them to be tuned and debugged and then to be switched to a real controlled plant [89]. such an approach is also applied to the tokamaks as real controlled plants (fig. 34) [90]. for instance, this approach was applied to the tuman-3 tokamak and was based on the analog-digital controllers and analog plant model [9, p. 204, fig. 5.16]. a controlling computer on the diii-d [91] and east [92] tokamaks enables switching from the plant model to the tokamak or vice versa through a corresponding switch. this allows saving time during physical experiment and thoroughly tuning plasma control systems in real time on the tokamak plasma models. tokamak model of the tokamak controller switch k1 k21 2 1 2 fig. 34. the concept of the real time computer test bed: к1 и к2 are switches from the plant model to the tokamak and vice versa [9] http://context.reverso.net/%d0%bf%d0%b5%d1%80%d0%b5%d0%b2%d0%be%d0%b4/%d0%b0%d0%bd%d0%b3%d0%bb%d0%b8%d0%b9%d1%81%d0%ba%d0%b8%d0%b9-%d1%80%d1%83%d1%81%d1%81%d0%ba%d0%b8%d0%b9/design+features 65 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 5.2.software realization of plasma shape control system in the jet tokamak the interface of the plasma shape control system in the jet tokamak [38] enables the user to divide the time of the experiment into segments, referred to as time windows, where it is possible to specify different combinations of the control regimes. for each time window the user may activate the specific controllers and set the programmed signals to track them. the user also can choose a shape control (sc) scenario containing the prepared set of controllers and programmed signals, the interface will adjust the signals shape for the specified initial and final values. the xsc scenarios became available when the extreme shape controller [93] system had been implemented, allowing to specify the cross section of the plasma visually and interactively update the programmed currents and gaps signals. the difference between the defined values for adjacent windows may cause substantial spikes of signals during the transition. to exclude this effect, the interface automatically introduces the additional smoothing time windows for the smooth transition from measured values at the end of the first window to defined values at the beginning of the next window. measurement => physics real-time diagnostic (time synchronization) evaluator ip, pos, shape, β, q, w, prad, regimedirect input · scaling · offset correction 1 2 control 3 i monitor: · operation limits, · generator speed feedback controller: · plasma current; · position; · shape; · confinement; · radiation; · profiles iv 4 5 reference generator: · waveforms interpolator; · intelligent waveform generator iiipulse scheduler: · waveforms; · transition rules segment scheduler: · watchdog; · conditional branching ii v command => actuator actuator adaptor: · load; · distribution; · modulation command output: · clipping; · scaling; · offsetting; · packaging vi diagnostic systems i actuator systems vi coils heating fuelling protection time event generators fig. 35. general scheme of the digital plasma control system for the asdex upgrade tokamak [94] 5.3.plasma control system for the asdex upgrade tokamak the digital dcs (discharge control system) for the asdex upgrade tokamak [94] (fig. 35) comprises the following elements: diagnostic system i, which includes the system of direct input 1, the real-time diagnostic 2 for time synchronization of the system, and evaluator for all measuring outputs 3, the control unit iv, that contains monitor, checking the operation limits and generator speed 4, and feedback control algorithms 5, which allow controlling the outputs of the system through actuators vi, pulse scheduler ii, reference generator iii, and segment scheduler v, containing watchdog and conditional branching. the plasma current, position, and shape control systems, plasma magnetic confinement system, radiation and plasma profiles control systems are integrated into the dcs. the control actions are formed on the 66 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) basis of the data of the diagnostic system. the control signals are fed to the actuators (control coils, additional heating system, fueling system, protection system, timing system of the events, generators), furthermore, a preliminary load, distribution and modulation and subsequently clipping, scaling, offsetting, and packaging of signals are carried out. the dcs provide the functional capabilities that allow one to guide the plasma discharges, coordinate measuring and actuating devices, and optimize the plasma parameters in the tokamak. 5.4.plasma control system of the tcv tokamak the closed-loop control system for the tcv (tokamak à configuration variable) [95] tokamak is presented in fig 36. the output signals from tcv tokamak (currents in coils, plasma density, magnetic characteristics, x-ray, etc.) come to the diagnostic system 1, and then are archived in the tcv database, and come to the input of the control system. the control signals are formed in blocks scd (système de contrôle distribué) 2 and hybrid control system 3, and are outputted to the programmable adder/switch 4, and then through actuators 5 (electron cyclotron resonance heating (ecrh), electron cyclotron current drive (eccd), toroidal (tf) and poloidal (pf) field coils, gas valve) are fed to the tokamak. the development of control systems for the tcv tokamak is carried out in a special interface of the scd host computer, that has access to the tcv database, can receive and transmit signals to the control unit, and realizes cooperation among all blocks of the digital system through the tcvpc computer. scd hybrid control system + programmable adder / switch actuators: ecrh/eccd, tf coils, pf coils, gas valve. diagnostics: coil currents, magnetics, density, x-ray, etc. archived in tcv database timer cards: trigger signals, clock signals. signals to program hardware tcvpc computer plant control & supervision shot-prepare parameters scd state, timer parameters scd host computer shot-prepare parameters post-shot data тсv database post-shot data algorithm code + waveforms signals to program hardware from tcvpc 5 1 3 4 2 fig 36. general scheme of the plasma control system for the tcv tokamak [95]. 67 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 6.approaches to development of magnetic control systems for plasma in tokamaks the review of magnetic control systems that have been undertaken makes it possible to classify them into different groups depending on approaches to the plasma vertical position and shape control. there are two approaches to plasma vertical movement control. the stabilization of the plasma vertical speed around zero is the first one. this approach is being applied to many devices such as jet, diii-d, east, kstar, nstx, iter project, where there are a wide range of plasma shape control systems. the stabilization of plasma vertical speed around zero makes it possible to avoid a contradiction with the plasma shape control problem, however, at the same time the system is not strictly stable and may have low stability margins. this approach also leads to the necessity of switching the controller in the stabilization loop of an unstable plant, since the plasma position stabilization system have to be applied before turning on the plasma shape control system (the stabilization of the plasma vertical speed around zero can be used only in combination with the plasma shape control system). the stabilization of plasma vertical position is the second one. the plasma vertical position control is applied directly during the limiter phase of the discharge on the asdex upgrade, jt-60sa, globus-m, etc. in some cases, there is a possibility to use control of the plasma shape during the divertor phase of the discharge without using the additional control loops for an unstable vertical position of plasma or plasma vertical speed around zero. such technical solution was used in an earlier draft version of the iter-1998 project [9] and during the analysis of such opportunity using mathematical modeling for the t-15 tokamak [75] that is under construction. the values of controllable parameters in real time, which are impossible to obtain by direct measurement, are necessary to solve the plasma shape control problem. the parameters of plasma shape may be reconstructed from the magnetic measurements outside the plasma by indirect methods, which are, in fact, a complex computational task [1]. today, there are a variety of approaches, using the values of different parameters of plasma, which are applied for plasma shape control in modern devices with a vertically elongated magnetic configuration. the plasma shape control using the values of gaps between the separatrix and the first wall (gap control). this approach is being used in the jet, asdex upgrade tokamaks, iter project. the control of the gaps between the separatrix and the first wall (vacuum vessel or blanket) has direct physical significance since indeed the safety parameters, namely the distances between the plasma boundary and the first wall, are controlled. the disadvantages of this approach include the computational complexity of the algorithms of plasma equilibrium reconstruction and the difficulty of their implementation in real time. the plasma shape control using the values of the magnetic flux in a set of points on the separatrix (isoflux control) is used in the diii-d, tcv, east, kstar tokamaks. this approach lessens the computational complexity compared to the previous one since it applies the control on the basis of indirect parameters and the full plasma shape reconstruction is not required. conclusion control of plasma dynamics is one of the central fundamental problems of the theoretical and experimental study of thermonuclear fusion and the project of transition to thermonuclear energy. however, the control methods, especially with respect to the internal plasma parameters (see later the fourth part of this survey), remain insufficiently developed due to the complexity of the construction and the variety of tokamaks, the need to use complex mathematical models, and the solution of complicated ill-posed problems in plasma diagnostics (see the first part [1] of the review), the development of voluminous science68 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) intensive software and the use of high-performance computing. in practice, this leads to a long and costly work on the experimental selection of control system parameters and a large number of premature discharge disruptions during research and development campaigns. therefore, the survey, systematization and classification of the real control systems for toroidal plasma are important and relevant. the given overview of the magnetic control systems of the position, shape, and current of the plasma on modern vertically elongated tokamaks, also including spherical tokamaks, shows that the problem of selecting (developing) an efficient and reliable structure of the plasma magnetic control system has not been finally solved. in the world practice, there are competing approaches to vertical stabilization of the plasma and various approaches to controlling the shape of the plasma, which have their advantages and disadvantages that cause their use on specific devices. hence, in particular, there is a lack of standards for the development of plasma control systems in tokamaks. nevertheless, it is possible to note the main trends both in the development of the tokamaks themselves and in the systems of magnetic control of the plasma in them. the vertically elongated tokamaks develop in the direction of decreasing the aspect ratio, which leads to spherical tokamaks, that have better physical and technical characteristics for the purposes of creation of future thermonuclear power plants based on tokamaks. in this regard, the work on the development, research, optimization, and simplification of the engineering implementation of plasma control systems is especially relevant for spherical tokamaks, in particular, primarily for the russian spherical globus-m2 tokamak, since it exceeds in the plasma parameters the known foreign analogs like mast and nstx (see subsection 3.4). on the other hand, the analysis of plasma magnetic control systems done allows to mark out the general features of the configurations of such systems, namely: multi input multi output (mimo);; multiple-loop; cascade control; hierarchy; robustness; adaptability. а tendency has been outlined for the multivariable control of the plasma shape and current to decouple the control channels and to use the simplest pid controllers in them (see subsections 1.3, 1.4, 2.3-2.5). it should be noted that in addition to control systems for the plasma position, current, and shape, the development of control systems for the resistive wall modes takes place as well, which must be suppressed when the plasma density increases by means of special additional coils and feedback control. similar properties and trends may appear in these systems. acknowledgements this work was supported by the russian science foundation (rsf), grant no 17-1901022 (sections 2-6) and the russian foundation for basic research (rfbr), grant no 17-0800293 (introduction, section 1). references 1. mitrishkin, y.v., korenev, p.s., prokhorov, a.a., kartsev, n.m., et. al. (2018). plasma control in tokamaks. part 1. controlled thermonuclear fusion problem. tokamaks. components of control systems. advances in systems science and applications. vol. 18, no. 2, pp. 26-52 . 69 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 2. artsimovich, l.a. (1961). upravlyaemye termoyadernye reaktsii [controlled thermonuclear reactions]. moscow, fismatgiz, [in russian]. 3. wesson, j. (2004). tokamaks: 3rd ed. oxford: clarendon press. 4. balshaw, n. all-the-world's tokamaks. [online]. available http://www.tokamak.info 5. kadomtsev, b.b. & shafranov, v.d. (1983). magnitnoe uderzhanie plazmy [magnetic confinement of plasma]. physics-uspekhi, 139(3), 399–434, [in russian]. 6. arsenin, v.v. & chuyanov, v.a. (1977). podavlenie neustojchivostej plazmy metodom obratnoj svyazi (obzor) [suppression of plasma instabilities by the feedback method (survey)]. physics-uspekhi, 123(1), 83–129, [in russian]. 7. samoilenko, y.i., gubarev, v.f. & krivonos, y.g. (1988). upravlenie bystroprotekayushchimi protsessami v termoyadernykh ustanovkah [rapid processes control in the thermonuclear devices]. kiev, naukova dumka, [in russian]. 8. ariola, m. & pironti, a. (2016). magnetic control of tokamak plasmas. berlin, springer, 2nd ed. 9. mitrishkin, y.v. (2016). upravlenie plazmoi v eksperimentalnykh termoyadernykh ustanovkakh: adaptivnye avtokolebatelnye i robastnye sistemy upravleniya [plasma control in the experimental thermonuclear devices: adaptive self-oscillating and robust control systems]. moscow, urss-krasand, [in russian]. 10. shi, w., wehner, w., barton, j., et al. (2012). a two-time-scale model-based combined magnetic and kinetic control system for advanced tokamak scenarios on diii-d. proc. 51st ieee conference on decision and control. maui, hawaii, usa, 4347–4352. 11. mitrishkin, y.v. (2012). multivariable plasma magnetic and kinetic control systems in tokamaks. proc. of iii joint symp. of taiwan-russia research cooperation on advanced problems in intelligent mechatronics, mechanics and control. moscow, russia, 169–177. 12. ambrosino, g., ariola, m., mitrishkin, y. & portone, a. (1997). plasma current and shape control in tokamaks using h∞ and µ-synthesis. proc. of the 36 ieee conf. on decision and control. san diego, usa, 3697–3702. 13. mitrishkin, y.v., kimura, h. (2001). plasma vertical speed robust control in fusion energy advanced tokamak. proc. of the 40th ieee conf. on decision and control. florida, usa, 1292–1297. 14. mitrishkin, y.v., dokuka, v.n., khayrutdinov, r.r. & kadurin, a.v. (2006). plasma magnetic robust control in tokamak-reactor. proc. of 45th ieee conf. on decision and control. san diego, usa, 2207–2212. 15. mitrishkin, y.v., korostelev, a.y., dokuka, v.n. & khayrutdinov, r.r. (2009). design and modeling of iter plasma magnetic control system in plasma current ramp-up phase on dina code. proc. of the 48th ieee conf. on decision and control. shanghai, china, 1354–1359. 16. mitrishkin, y.v. & kartsev, n.m. (2011). hierarchical plasma shape, position, and current control system for iter. proc. of the 50th ieee conf. on decision and control and european control conf. orlando, usa, 2620–2625. 17. mitrishkin, y.v., zenckov, s.m., kartsev, n.m., efremov, a. a., et. al. (2012). linear and impulse control systems for plasma unstable vertical position in elongated tokamak. proc. of the 51st ieee conf. on decision and control, maui, hawaii, usa, 1697–1702, https://doi.org/10.1109/cdc.2012.6426971. 18. mitrishkin, y.v., korenev, p.s., prohorov, a.a. & patrov, m.i. (2017). tokamak plasma magnetic control system simulation with reconstruction code in feedback based on http://www.tokamak.info/ 70 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) experimental data, proc. of ieee 56th annual conf. on decision and control. melbourne, australia, 2360–2365. 19. mitrishkin, y.v., kadurin, a.v. & korostelev, a.y. (2011). tokamak plasma shape and current h controller design in multivariable cascade system. proc. of 18th ifac world congress. milan, italy, 3722–3727. 20. mitrishkin, y.v., ivanov, v.a. (2011). combined nonlinear tokamak plasma current profile control system design with input constraints. proc. of ifac world congress. milan, italy, 3728–3733. 21. mitrishkin, y.v., korenev, p.s., prohorov & a.a., patrov, m.i. (2017). robust h∞ switching mimo control for a plasma time-varying parameter model with a variable structure in a tokamak. proc. of ifac 2017 world congress. toulouse, france, 11883– 11888. 22. mitrishkin, y.v. (1985). upravlenie dinamicheskimi obektami s primeneniem avtomaticheskoj nastrojki [control of dynamic plants with application of automatic adjustment]. moscow, nauka, [in russian]. 23. gribov, y.v., chuyanov, v.a. & mitrishkin, y.v. (1985). sposob stabilizacii polozheniya plazmennogo shnura v tokamake [method for stabilizing the plasma position in a tokamak]. certificate of authorship, 1119490, ussr, 19, 243, [in russian]. 24. koscov, y.a., gribov, y.v. & mitrishkin, y.v. (1985). ustrojstvo dlya stabilizacii ravnovesnogo polozheniya plazmennogo shnura v tokamake [device for stabilizing the equilibrium position of a plasma in a tokamak]. certificate of authorship, 1153698, ussr, 37, 258, [in russian]. 25. gribov, y.v., kuznecov, e.a., mitrishkin, y.v. & chuyanov, v.a. (1986). relejnaya sistema stabilizacii polozheniya plazmy tokamaka [relay system for stabilizing the position of the tokamak plasma]. problems of atomic science and technology, ser. thermonuclear fusion, 4, 51–57, [in russian]. 26. kuznecov, e.a., mitrishkin, y.v. & savkina, i.s. (1987). adaptivnaya sistema minimizacii amplitudy avtokolebanij plazmennogo shnura otnositel'no zadannogo polozheniya v tokamake [adaptive system for minimizing the amplitude of autooscillations of a plasma with respect to a given position in a tokamak]. upravlenie slozhnymi tekhnicheskimi sistemami, moscow, ipu ras, 30–33, [in russian]. 27. gribov, y.v., kuznecov, e.a., mitrishkin, y.v., savkina i.s., & chuyanov, v.a. (1988). adaptivnaya optimal'naya sistema upravleniya gorizontal'nymi smeshcheniyami plazmennogo shnura v tokamake [adaptive optimal control system for horizontal displacements of a plasma in a tokamak], problems of atomic science and technology, ser. thermonuclear fusion, 4, 28–32, [in russian]. 28. gribov, y.v., mitrishkin, y.v., chuyanov, v.a. & shahovec, k.g. (1988). sposob stabilizacii polozheniya plazmennogo shnura v tokamake [method for stabilizing the plasma position in a tokamak]. certificate of authorship, 1399824, ussr, 20, 231, [in russian]. 29. gribov, y.v., koscov, y.a., kuznecov, e.a., mitrishkin y.v.& chuyanov, v.a. (1988). ustrojstvo dlya stabilizacii polozheniya plazmennogo shnura v tokamake [device for stabilizing the equilibrium position of a plasma in a tokamak]. certificate of authorship, 1418817, ussr, 31, 244, [in russian]. 71 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 30. mitrishkin, y.v. (1989). minimizaciya amplitudy avtokolebanij v relejnoj sisteme upravleniya s ustojchivoj linejnoj dinamicheskoj chast'yu [minimization of the amplitude of self-oscillations in a relay control system with a stable linear dynamic part]. avtomatika i telemekhanika, 9, 91–102, [in russian]. 31. mitrishkin, y.v. & kuznetsov, e.a. (1993). estimation of parameters of stabilized plasma, plasma devices and operations, 2(3), 277–286. 32. bortnikov, a.v., gerasimov, s.n., mitrishkin, y.v., & polianchik, k.d. (1990). relejnoe avtomaticheskoe upravlenie polozheniem plazmennogo shnura po gorizontali v tokamake tvd s pomoshch'yu stacionarnyh regulyatorov [relay automatic control of the horizontal plasma position in the tvd tokamak using stationary regulators]. preprint iae, 5218/7, 1-51, [in russian]. 33. bortnikov, a.v., gerasimov, s.n., mitrishkin, y.v., & cypakin, i.a. (1990). sistema avtomaticheskogo upravleniya polozheniem plazmy po gorizontali i vertikali v tokamake tvd [the system of automatic control of the horizontal and vertical plasma position in the tvd tokamak]. preprint iae, 5096/7, 1-12, [in russian]. 34. bortnikov, a.v., gerasimov, s.n., mitrishkin, y.v., and cypakin, i.a. (1990). invertor napryazheniya avtomaticheskoj sistemy upravleniya polozheniem plazmennogo shnura v tokamake tvd [the voltage inverter of the automatic control system of the plasma position in the tvd tokamak]. preprint iae, 5068/7, 1-36, [in russian]. 35. abramov, a.v., bortnikov, a.v., mitrishkin y. v., brevnov, n.n., et. al. (1991). shaping, vertical stability and control elongated plasmas on the tvd, preprint iae, 5301/7, 1-41. 36. kuznecov, e.a. & mitrishkin, y.v. (2005). avtokolebatel'naya sistema stabilizacii neustojchivogo vertikal'nogo polozheniya plazmy sfericheskogo tokamaka globus-m [self-oscillatory system for stabilizing the plasma unstable vertical position of the spherical tokamak globus-m]. moscow, scientific edition, v.a. trapeznikov institute of control sciences of ras, [in russian]. 37. kuznetsov, e.a., yagnov, v.a., mitrishkin, y.v., shcherbitsky, v.n. (2017). current inverter as actuator for plasma position control systems in tokamaks. proc. of the 11th ieee intern. conf. on application of information and communication technologies (aict2017), moscow, russia, 485–489. 38. sartori, f., tommasi, g.d. & piccolo, f. (2006). the joint european torus. plasma position and shape control in the world’s largest tokamak. ieee control syst. magazine, 26(2), 64–78. 39. neto, a., albanese, r., ambrosino, g., ariola, m., et al. (2011). exploitation of modularity in the jet tokamak vertical stabilization system, proc. of the 50th ieee conf. on decision and control and european control conference, orlando, usa, 2644– 2649, https://doi.org/10.1109/cdc.2011.6160510. 40. bellizio, t., albanese, r., ambrosino, g., ariola, m., et al. (2011). control of elongated plasma in presence of elms in the jet tokamak, ieee trans. on nuclear science, 58(4), 1497–1502, https://doi.org/10.1109/tns.2011.2157524. 41. yuan, q.p., xiao, b.j., luo, z.p., walker, m. l., et. al. (2013). plasma current, position and shape feedback control on east. nuclear fusion, 53(4), 043009, https://doi.org/10.1088/0029-5515/53/4/043009. 42. albanese, r., ambrosino, r., castaldo, a., de tommasi, g., et. al. (2017). iter-like vertical stabilization system for the east tokamak, nuclear fusion, 57(8), https://doi.org/10.1088/1741-4326/aa7a78. https://doi.org/10.1088/1741-4326/aa7a78 72 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) 43. qiu, q., xiao, b., guo, y., liu, l., et al. (2016). simulation of east vertical displacement events by tokamak simulation code, nuclear fusion, 56(10), https://doi.org/10.1088/00295515/56/10/106029. 44. walker, m.l., johnson, r.d., leuer, j.a. & penaflor, b.g. (2005). on-line calculation of feedforward trajectories for tokamak plasma shape control, proc. of the 44th ieee conf. on decision and control, and the european control conference, seville, spain, 8233–8239, https://doi.org/10.1109/cdc.2005.1583495. 45. walker, m.l., humphreys, d.a., leuer, j.a., ferron, j.r., et. al. (2000). implementation of model-based multivariable control on diii–d. ga–a23468. [online]. available https://fusion.gat.com/pubs-ext/soft00/a23468.pdf. 46. herrmann, a. & gruber, o. (2003). chapter 1: asdex upgrade – introduction and overview, fusion science and technology, 44(3), 569–577. 47. streibl, d., lang, p.t., leuterer, f., et. al. (2003). chapter 2: machine design, fueling, and heating in asdex upgrade, fusion science and technology, 44(3), 578–592. 48. mertens, v., raupp, g. & treutterer, w. (2003). chapter 3: plasma control in asdex upgrade, fusion science and technology, 44(3), 593–604. 49. de tommasi, g., maviglia, f., neto, a.c., et. al. (2014). jet-efda contributors plasma position and current control system enhancements for the jet iter-like wall, fusion engineering and design, 89(3), 233–242. 50. albanese, r., cocorrese, e. & rubinacci, g. (1989). plasma modeling for the control of vertical instabilities in tokamaks, nuclear fusion, 29(6), 1013–1022, https://doi.org/10.1088/0029-5515/29/6/011. 51. beghi, a. & cenedese, a. (2005). advances in real-time plasma boundary reconstruction, ieee control syst. mag., 25(5), 44–64, https://doi.org/10.1109/mcs.2005.1512795. 52. lister, j.b., hofmann, f., moret, j.m., et. al. (1997). the control of tokamak configuration variable plasmas, fusion technology, 32(3), 321–373, https://doi.org/10.13182/fst97a1. 53. ambrosino, g. & albanese, r. (2005). magnetic control of plasma current, position, and shape in tokamaks, ieee control syst. mag., 25(5), 76–92, https://doi.org/10.1109/mcs.2005.1512797. 54. jhang, h., kessel, c.e., pomphrey, n., kim, j.y., et. al. (2001). simulation studies of plasma identification and control in korea superconducting tokamak advanced research, fusion engineering and design, 54(1), 117–134, https://doi.org/10.1016/s09203796(00)00431-2. 55. hutchinson, i.h., horne, s.f., tinios, g., stephen, m. w., et. al. (1996). plasma shape control: a general approach and its application to alcator c-mod, fusion technology, 30(2), 137–150, https://doi.org/10.13182/fst96-a30746. 56. freidberg, j. (2007). plasma physics and fusion energy. cambridge: cambridge university press. 57. cunningham, g. (2013). high performance plasma vertical position control system for upgraded mast, fusion engineering and design, 88(12), 3238–3247, https://doi.org/10.1016/j.fusengdes.2013.10.001. 58. boyer, m.d., battaglia, d.j., mueller, d., eidietis, n., et. al. (2018). plasma boundary shape control and real-time equilibrium reconstruction on nstx-u, nuclear fusion, 58(3), https://doi.org/10.1088/1741-4326/aaa4d0. https://doi.org/10.1088/0029-5515/56/10/106029 https://doi.org/10.1088/0029-5515/56/10/106029 https://fusion.gat.com/pubs-ext/soft00/a23468.pdf https://doi.org/10.1088/0029-5515/29/6/011 https://doi.org/10.13182/fst97-a1 https://doi.org/10.13182/fst97-a1 https://doi.org/10.1109/mcs.2005.1512797 https://doi.org/10.1016/s0920-3796(00)00431-2 https://doi.org/10.1016/s0920-3796(00)00431-2 https://doi.org/10.13182/fst96-a30746 https://doi.org/10.1088/1741-4326/aaa4d0 73 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 59. pangione, l., mcardle, g. & storrs, j. (2013). new magnetic real time shape control for mast, fusion engineering and design, 88, 1087–1090, https://doi.org/10.1016/j.fusengdes.2013.01.048. 60. sakharov, n.v. (2001). spherical tokamak globus-m construction and operation. plasma devices and operations, 9(1-2), 25–38. 61. gusev, v.k., e.a. azizov, e.a., alekseev, a.b., et. al. (2013). globus-m results as the basis for a compact spherical tokamak with enhanced parameters globus-m2. nuclear fusion, 53(9), 093013. 62. korenev, p.s., mitrishkin, y.v. & patrov, m.i. (2016). rekonstrukciya ravnovesnogo raspredeleniya parametrov plazmy tokamaka po vneshnim magnitnym izmereniyam i postroenie lineinykh plazmennykh modelei [reconstruction of equilibrium distribution of tokamak plasma parameters by external magnetic measurements and construction of linear plasma models]. mechatronics, automation, control, 17(4), 254–265, [in russian]. 63. mitrishkin, y.v., prohorov, a.a., korenev, p.s. & patrov, m.i. (2017). sposob modelirovaniya sistem magnitnogo upravleniya formoj i tokom plazmy s obratnoj svyaz'yu v tokamake [a method for modeling the magnetic control of the shape and current of a plasma with a feedback in a tokamak]. application no. 2017115081 of the rf patent, [in russian]. 64. mitrishkin, y.v., prohorov, a.a., korenev, p.s. & patrov, m.i. (2017). metod modelirovaniya sistem magnitnogo upravleniya plazmoj v tokamake s kodom vosstanovleniya ravnovesiya plazmy v obratnoj svyazi na osnove eksperimental'nyh dannyh [the method of simulation of plasma magnetic control systems in a tokamak with a code for the plasma equilibrium reconstruction in the feedback on the basis of experimental data]. lomonosovskie chteniya-2017. sekciya fiziki, moscow, russia, 181– 184, [in russian]. 65. karcev, n.m., mitrishkin, yu.v. & patrov, m.i. (2017). ierarhicheskie robastnye sistemy magnitnogo upravleniya plazmoj v tokamakah s adaptaciej [hierarchical robust systems of magnetic plasma control in tokamaks with adaptation]. avtomatika i telemekhanika, 4, 149–165, [in russian]. 66. azizov, e., velikhov, e., belyakov & v., filatov, o. (2010). status of project of engineering-physical tokamak, 23rd iaea fusion energy conf. daejeon. republic of korea, ftp/p6-01. 67. mitrishkin, y.v., karcev, n.s. & zenkov, s.m. (2014). stabilizaciya neustojchivogo vertikal'nogo polozheniya plazmy v tokamake t-15. chast' i [stabilization of the unstable vertical position of the plasma in the t-15 tokamak. part i]. avtomatika i telemekhanika, 2, 129–145, [in russian]. 68. mitrishkin, y.v., karcev, n.s. & zenkov, s.m. (2014). stabilizaciya neustojchivogo vertikal'nogo polozheniya plazmy v tokamake t-15. chast' ii [stabilization of the unstable vertical position of the plasma in the t-15 tokamak. part ii]. avtomatika i telemekhanika, 9, 31–44, [in russian]. 69. zhou, k. & doyle, j.c. (1998). essentials of robust control. prentice hall, available http://www.dl.offdownload.ir/ali/essentials%20of%20robust%20control.pdf 70. lukash, v.e., dokuka, v.n. & hajrutdinov, r.r. (2004). programmno-vychislitel'nyj kompleks dina v sisteme matlab dlya resheniya zadach upravleniya plazmoj tokamaka [the dina software-computing complex in the matlab system for the https://inis.iaea.org/search/search.aspx?orig_q=author:%22velikhov,%20e.%22 https://inis.iaea.org/search/search.aspx?orig_q=author:%22belyakov,%20v.%22 https://inis.iaea.org/search/search.aspx?orig_q=author:%22filatov,%20o.%22 http://www.dl.offdownload.ir/ali/essentials%20of%20robust%20control.pdf 74 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) solution of tokamak plasma control problems]. problems of atomic science and technology, ser. thermonuclear fusion, 1, 40–49, [in russian]. 71. mitrishkin, y.v., kartsev, n.m. & zenkov, s.m. (2013). plasma vertical position, shape, and current control in t-15 tokamak, proc. of the ifac conference on manufacturing modelling, management and control. saint petersburg, russia, 1820–1825, https://doi.org/10.3182/20130619-3-ru-3018.00250. 72. mitrishkin, y.v., pavlova, e.a., kuznetsov, e.a. & gaydamaka, k.i. (2016). continuous, saturation, and discontinuous tokamak plasma vertical position control systems. fusion engineering and design, 108, 35–47, https://doi.org/10.1016/j.fusengdes.2016.04.026. 73. golubtcov, m.p., mitrishkin, y.v. & sokolov, m.m. (2016). adaptive model predictive control of tokamak plasma unstable vertical position, proc. 2016 intern. conf. stability and oscillations of nonlinear control systems (pyatnitskiy's conf.), https://doi.org/10.1109/stab.2016.7541185. 74. pavlova, e.a., mitrishkin, y.v. & khlebnikov, m.v. (2017). control system design for plasma unstable vertical position in a tokamak by linear matrix inequalities, proc. of the 11 th ieee intern. conf. on application of information and communication technologies (aict2017), moscow, russia, 458–462. 75. zenkov, s.m., mitrishkin, y.v. & fokina, e.k. (2013). mnogosvyaznye sistemy upravleniya polozheniem, tokom i formoj plazmy v tokamake t-15 [multivariable control systems for the position, current and shape of the plasma in the t-15 tokamak]. problemy upravleniya, 4, 2–10, [in russian]. 76. dokuka, v.n., kavin, a.a. & lukash, v.e. (2014). chislennoe modelirovanie upravleniya plazmoj v modernizirovannom tokamake t-15 [numerical simulation of plasma control in the modernized t-15 tokamak]. problems of atomic science and technology, ser. thermonuclear fusion, 37(3), 56–70, [in russian]. 77. pustovitov, v.d. (2003). usilenie rezonansnogo polya na granice ustojchivosti rwm v tokamake [amplification of the resonance field at the stability boundary of the rwm in a tokamak]. jetp letters, 78(5), 727–730, [in russian]. 78. walker, m.l., humphreys, d.a., mazon, d., moreau, d., et. al. (2006). emerging applications in tokamak plasma control, ieee control systems magazine, 2, 35–63, https://doi.org/10.1109/mcs.2006.1615272. 79. okabayashi, m., pomphrey, n., hatcher, r.e. (1998). circuit equation formulation of resistive wall mode feedback stabilization schemes, nuclear fusion, 38(11), 1607–627. https://doi.org/10.1088/0029-5515/38/11/302. 80. liu, y.q. & bondeson, a. (2000). active feedback stabilization of toroidal external modes in tokamaks, physical review letters, 84(5), 907–910, https://doi.org/10.1103/physrevlett.84.907. 81. in, y., kim, j.-s., humphreys, d.a. & walker, m.l. (2006). model-based static and dynamic filter application to resistive wall mode identification and feedback control in diii-d, proc. of 45th ieee conf. on decision and control. san diego, usa, 2250–2256. 82. garofalo, a.m. (2006, november 6-8) resistive wall mode control in diii-d. workshop on active control of mhd stability: active mhd control in iter, princeton plasma physics laboratory, princeton, new jersey. http://ieeexplore.ieee.org/search/searchresult.jsp?searchwithin=%22authors%22:.qt.m.%20p.%20golubtcov.qt.&newsearch=true http://ieeexplore.ieee.org/search/searchresult.jsp?searchwithin=%22authors%22:.qt.y.%20v.%20mitrishkin.qt.&newsearch=true http://ieeexplore.ieee.org/search/searchresult.jsp?searchwithin=%22authors%22:.qt.m.%20m.%20sokolov.qt.&newsearch=true http://ieeexplore.ieee.org/document/7541185/ http://ieeexplore.ieee.org/document/7541185/ http://dx.doi.org/10.1109/stab.2016.7541185 75 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 83. dalessio, j., schuster, e., humphreys, d., walker, m., et. al. (2008). extending the rwm stability region by optimal feedback control, proc. of 47th ieee conf. on decision and control, cancun, mexico, 3145–3150, https://doi.org/10.1109/cdc.2008.4739359. 84. schuster, e. (2008, may 7-8). integrated plasma control. workshop “control for nuclear fusion”, eindhoven university of technology, netherlands. 85. garofalo, a.m., jackson, g.l., la haye, r.j., okabayashi, m., et. al. (2007). stability and control of resistive wall modes in high beta, low rotation diii-d plasmas, nuclear fusion, 47(9), 1121–1130, https://doi.org/10.1088/0029-5515/47/9/008. 86. katsuro-hopkins, o., sabbagh, s.a. & bialek, j.m. (2009). analysis of resistive wall mode lqg control in nstx with mode rotation. proc. joint 48th ieee conference on decision and control and 28th chinese control conference. shanghai, p.r. china, 309– 314. https://doi.org/10.1109/cdc.2009.5400543. 87. katsuro-hopkins, o., bialek, j., maurer, d.a. & navratil, g.a. (2007). enhanced iter resistive wall mode feedback performance using optimal control techniques, nuclear fusion, 47, 1157–1165, https://doi.org/10.1088/0029-5515/47/9/012. 88. maurer, d.a., bialek, j., navratil, g.a., mauel, m.e., et. al. (2006). controllability and reduced state space models for feedback control of the resistive wall kink mode, proc. of 45th ieee conf. on decision and control. san diego, usa, 2271–2275, https://doi.org/10.1109/cdc.2006.377492. 89. applied dynamics international. (2018, march 1). solutions in real time. [online]. available http://www.adi.com/. 90. mitrishkin, y.v., efremov, a.a. & zenkov, s.m. (2013). experimental test bed for real time simulations of tokamak plasma control systems, journal of control engineering and technology, 3(3), 121–130. 91. walker, m.l., humphreys, d.a., leuer, j., et. al. (2003). practical control issues on diiid and their relevance for iter, general atomics, engineering physics memo, epm111803a. 92. wang, s., yuan, q. & xiao, b. (2017). development of the simulation platform between east plasma control system and the tokamak simulation code based on simulink, plasma science and technology, 19(3), 035601, https://doi.org/10.1088/2058-6272/19/3/035601. 93. ariola, m. & pironti, a. (2005). plasma shape control for the jet tokamak: an optimal output regulation approach, ieee control syst. mag, 25(5), 65–75, https://doi.org/10.1109/mcs.2005.1512796. 94. treutterer, w., cole, r., lüddecke, k., neu, g., et. al. (2014). asdex upgrade discharge control system – a real-time plasma control framework, fusion engineering and design, 89(3), 146–154. https://doi.org/10.1016/j.fusengdes.2014.01.001. 95. le, h.b., felici, f., paley, j.i., duval, b.p., et. al. (2014). distributed digital real-time control system for tcv tokamak, fusion engineering and design, 89(3), 155–164, https://doi.org/10.1016/j.fusengdes.2013.11.001. https://doi.org/10.1109/cdc.2008.4739359 https://doi.org/10.1088/0029-5515/47/9/008 https://doi.org/10.1109/cdc.2009.5400543 https://doi.org/10.1109/cdc.2006.377492 http://www.adi.com/ https://doi.org/10.1016/j.fusengdes.2013.11.001 http://www.sciencedirect.com/science/journal/09203796 http://www.sciencedirect.com/science/journal/09203796/89/3 https://doi.org/10.1016/j.fusengdes.2013.11.001 76 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) 77 plasma control in tokamaks. part 2. copyright ©2018 assa adv. in systems science and appl. (2018) 78 y.v. mitrishkin, p.s korenev, a.a. prokhorov, n.m. kartsev, m.i. patrov et al. copyright ©2018 assa adv. in systems science and appl. (2018) advances in systems science and applications (2014) vol.14 no.4 361-377 evaluation of users’ behavior using adaptive fuzzy prediction technique joan m., john, a. and shajin nargunam computer science & engineering department, noorul islam university, tamilnadu, india-629180 abstract user behavior is a considerable determinant of a products ecological blow as manufacturing advances authorize enlarged effectiveness of product purpose, the consumers alternatives and habits ultimately enclose a major outcome on possessions utilized by the product. a variety of design methods developed in diverse contexts propose chances for engineers and other stakeholders functioning in the ground of sustainable improvement to involve consumers actions at communication with the result or system, in result building the user more competent. approaches to varying users actions from varied amount of fields are evaluated and discussed. but it fails to predict the user behaviors accurately. in this work, adaptive fuzzy prediction technique is presented to interpret and predict the users behavior towards product demand, quantity of products, features of the product/services, annual sales with the multi-clustered web usage data. the fuzzy model works on the feedback pattern to arrive at predictive feature of the web usage data value moments, across different time periods, and adapt to the recent consumer behavior and liking of the products and service. experimentation conduct with real data sets extracted from web log files collected from websites. performance evaluation is made to show the effectiveness of the proposed technique. keywords users behavior, product, behavior-patterns, feedback pattern, multiclustered data, adaptive fuzzy prediction technique. 1 introduction web data mining is a type of approaches that proficiently processes the tasks of providing the required information from the internet, enhancing the web site framework to offer better internet service quality and identifying the informative expertise from the internet for enhanced web appliances. usually, web data mining are of three types, web content, web structure and, web usage mining. in addition to identifying the requirements of the customers, organization also require to recognize what motivates them to buy, and how can processes the purchasing procedures to make sure that the products or services are on the purchasing list. understanding the customers will assist, to enhance and provide the product, in addition to obtaining the right price crossing points and enhancing successful promotional behaviors. the psychology of the purchasing procedures has been greatly noted and no matter what size organization business, proce362 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique dures that assist company turn into be more successful. both organizations and consumers provide patterns of purchasing actions. the business representation is less process to debate as the business consumers will almost dominantly encompass some generalized procedures of buying in place. the organization task is to recognize the procedure and match the marketing behaviors to the diverse stages of the procedure. this means that the consumer will obtain the right type of contact at the precise time. to process insights into areas for instance indicators of consumer defection, crossing points sensitivity, clustering, and consumer needs examination, to process a few. most marketers identify the value of sensing consumer data, but also act as the endeavors of leveraging this expertise to produce intelligent, proactive routes rear to the consumer. data mining approaches and techniques for identifying and monitoring patterns among data assists businesses sift during layers of seemingly unspecified data for purposeful relationships, where they could processes, rather than merely route to, consumer requirements. defined and precise predictive representations are very significant in viewing schemes. the requirement for novel techniques and attitudes in representing calculation and vulnerability are prejudiced by the current proceeds in soft computing also the problematic accurateness and inapplicability to person calculation of formerly required after geometric analysis methods. thus founding accurate analytical representations become gradually harder for multivariable projecting representations. conventionally, such troubles have been lectured by geometric logistic deterioration methods for binary reliant variables. therefore, it can be accomplished that the web is an assorted, energetic and amorphous data depository, which gives huge quantity of valuable information and also demonstrates the difficulty of managing enormous amount of information from the diverse perspectives, users, web repair providers, commerce analysts. the efficient search tools are very greatly essential to recognize significant and practical information accurately. the growth has enthused the web service suppliers to predict the users web usage behaviors so that, they can • personalize the information offered to them • construct the websites more users sociable • decrease the traffic load • make or adapt their website to outfit diverse group of people in this work, predict the users behaviors based on their selection over product demand, quantity of products, features of the product/services, annual sales, etc., by employing the proposed adaptive fuzzy prediction technique. advances in systems science and applications (2014) vol.14 no.4 363 2 related work world wide web is an enormous depository of web pages and associations. it presents great quantity information for the internet users. data mining has developed as a ground of essential and functional study in computer science. in [1], an adaptive technique for predicting cancer susceptibility was designed by combining fuzzy concept with statistical logistic regression which was tested on cancer dataset. as to the authors suggestion regarding the generalization of the algorithm into prediction problem involving other types of intrinsic linear functions, and adaptive fuzzy prediction technique is used to evaluate users behavior. the development of web is unbelievable as it can be observed in current days. users admittance is traced in web logs. from the users viewpoint, it is very complicated to haul out practical knowledge from the enormous quantity of information and secondly, it is also hard to take out for the users to admit pertinent information professionally. web logs extracted and purpose of this exposition is to assess, suggest and progress the exercise of a quantity of the current approaches, architectures and web mining methods (gathering individual information from consumers) are the earnings of operating data mining techniques to persuade and take out helpful information from web information and service where data mining has been practiced in the pastures of e-commerce and e-business (that funds users actions) [2]. one method to solve such problem is the application of web usage mining as in [3]. the evaluation of users behavior using adaptive fuzzy prediction technique deploys the multi-clustered web usage data in order to derive the behavior pattern. a developing clustering algorithm is planned for cluster production. clusters are recognized and customized supported on restraint measure of mapping consistence and companionable dimension. auspiciously, online social networks and social standard have completed it simple for users to point to whom they reliant and whom they achieve not. nevertheless, this does not resolve the crisis since every user is only possible to recognize an insignificant portion of other users; the author must encompass techniques for deducing reliant and disbelieve among users who do not recognize one another. in [4], the authors present a novel technique for calculating both reliant and non-reliant (i.e., optimistic and pessimistic trust). the authors supply to this region a novel algorithm for efficiently forecasting reliance and mistrust in web-based communal systems. the authors in [5] unite a path possibility reliance inference algorithm with a new method employing spring-embedding to deduce system distance. the product of the classifiers is evaluated with offered techniques which exercise considerably diverse techniques. when a user needs admin contact, a communal conversation page is associate for users to converse and cast 364 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique your vote on whether to disclose the mediator. optimistic and pessimistic votes are calculated as optimistic and pessimistic trust ratings [6]. the methods in [7] have exposed an enormous potential for humanizing classification accuracy. this revision is concerned with the assessment of training data allotment and its strength on the arrangement of frequent classifier schemes. temporal data clustering presents establishment techniques for shaping the intrinsic association and concentrating information in excess of chronological data. the author in [8] calculated techniques for communication seek out in associated, high-dimensional data sets, which are resultant surrounded by a clustering organization. the author takes you back that indexing by vector processes (va-file), which was anticipated as a way to match the disfigurement of dimensionality, uses scalar qualification, and so fundamentally ignores dependences obliquely magnitude, which represents an origin of sub-optimality [9]. researches on position-based service (lbs) have been promising in modern years remaining to a large cycle of probable applications [10]. one of the vital subjects is the removal and computation of mobile actions and associated communication. feature clustering is a scheming method to reduce the dimensionality of distinctive vectors for content tagging. in this paper, the author in [11] suggests fuzzy similarity-based self-constructing practices for characteristic grouping [12]. the paper [13] urbanized a novel resolution structure for adaptive production response fuzzy direct schemes aimed at efficiently selling with nonlinear schemes with numerous input-multiple production (mimo) setback and in the attendance of dynamics and normal function suspicions. a multiple delay fuzzy scheme calculation representation is consequent and its scheme possessions are elucidated. such a calculation representation allows the exercise of a model-based technique for fuzzy control. the paper [14] investigated the possibility of concerning a comparatively new neural system method, i.e., great learning mechanism (elm), to understand a neuro-fuzzy takagi-sugeno-kang (tsk) fuzzy supposition scheme. the tsk technique is an enhanced report of the standard neuro-fuzzy tsk fuzzy deduction scheme. at the equivalent time, the consequential fraction of the fuzzy rules is attained by numerous elms. finally, the estimated calculation value is dogged by a load calculation system. in [15], the author are disturbed with a technique for building quantum-based adaptive neuro-fuzzy systems (qanfns) by means of a takagi-sugeno-kang (tsk) fuzzy category supported on the fuzzy granulation from a specified effort production data set. for this reason, the author urbanized a methodical technique in creating habitual fuzzy rules supported on fuzzy subtractive quantum grouping. the clustering method is not only an addition of thoughts intrinsic to scale-space and support-vector grouping but also symbolizes an efficient sample that shows advances in systems science and applications (2014) vol.14 no.4 365 definite uniqueness of the objective scheme to be represented from the fuzzy subtractive technique. through the precedent few years, the primary deduction for corporeal repossession is being anticipated employing regression methods from sounder comments. the current system employs fuzzy reason and data clustering to institute an association among replicated sounder comments and impressive outlines. this association is more reinforced employing the adaptive neuro-fuzzy supposition scheme (anfis) by fine-tuning the offered fuzzy-rule stand [16]. the paper [17] presents an appliance of fuzzy-logic methods to the reversible firmness of grayscale images. with suggestion to a spatial degree of difference pulse code intonation (dpcm) system, calculation might be skilled in a space-varying manner moreover as adaptive, i.e., with predictors recalculated at every pixel, or as confidential, in which image chunks or pixels are tagged in a number of lessons. a novel method to the devise and employ of inferential sensors in the procedure trade is proposed in [18], which is supported on the lately introduced idea of developing fuzzy models (efms). they speak to the confront that the current procedure industry faces nowadays, that is, to expand such adaptive and selfcalibrating online inferential sensors that decrease the preservation costs while charging the elevated accuracy and interpretability/transparency. kalman filter (kf) is the mainly regularly employed assessment method for incorporating signals from short-term elevated presentation schemes, similar to inertial directionfinding schemes (inss), with position systems showing long-term constancy, similar to the comprehensive positioning scheme (gps) [19]. the paper [20] presents completion consequences using lately pioneered discrete-time adaptive calculation and direct methods employing online task approximators. a fuzzy time series [21] has been functional to the calculation of employment, hotness, stock directories, and other areas. associated studies mostly center on three issues, specifically, the separation of conversation, the pleased of forecasting system, and the techniques of defuzzification, all of which really control the calculation accurateness of forecasting representations. the presentation of an adaptive neurofuzzy assumption scheme (anfis) considerably falls when improbability lives in the data or scheme process. prediction period (pis) can enumerate the indecision connected with anfis point forecasts. the paper [22] first nears a method to acclimatize the delta method for the creation of pis for conclusions of the anfis representations. in [23], a developing fuzzy system (efs) is urbanized for scheme condition forecasting. 366 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique 3 proposed adaptive fuzzy prediction technique to evaluate users behavior the process of identifying the users behavior based on their demands over the products and its quality is done by accomplishing the adaptive fuzzy prediction technique. the adaptive fuzzy prediction model is used to predict the users behaviors efficiently across different time periods and adapt to the recent user behavior and liking of the products and service. here the adaptive fuzzy prediction model is combined with the representation of fuzzy assumption with user behavior predictor. the fuzzy prediction technique is applied to the multi-clustered web usage data. the architecture diagram of the proposed evaluation of users behavior using adaptive fuzzy prediction technique [eubafp] is shown in fig. 1. from the fig.1, it is being observed that the web log file data is given as input to the system to estimate the performance. before fuzzification, the web log file data is being pre-processed with the corresponding method and normalized to the self-organizing maps. with these, web usage data was multi-clustered in an appropriate manner based on the similarity based clustering. with the web usage multi-clustered data, fuzzy prediction model is applied to classify the users data based on their products demands over the market. the fuzzy inference model will form a set of rules at first as a base for classifying the rules. the adaptive fuzzy prediction model of the website could be processed in the constructing phase under which it collects web browsing log. the browsing behaviors of the users are modeled and compared with the representation of the prediction model to improve the browsing performance of the users. fig.1 architecture diagram of the proposed eubafp 3.1 design of adpative fuzzy prediction technique fig.2 shows the block diagram of an adaptive fuzzy prediction technique. it comprises of two subsystems: a fuzzy prediction system, and behavior predictor. advances in systems science and applications (2014) vol.14 no.4 367 at first, in the fuzzy prediction model, the users in the web log data file are clustered based on their behaviors. the fuzzy prediction system assumes the possibility of the active time of the user in the cluster based on the behaviors of the user. next, in order to identify the active time of the user in the group, behavior predictor is used. the behavior predictor identifies the possibility of number of users visited the web page. the expected possibility information is employed in web usage data management to handle users’ behavior. fig.2 adaptive fuzzy prediction technique the browsing behaviors of users are predicted based on the prediction model. since the browsing behaviors are identified, the predictable patterns should be combined or the prediction pattern should be built, correspondingly. the browsing behavior of a user is compared with all matching visited web pages in prediction model for expecting the future traveling path. the design of the subsystems is given below. normally, there is a greater chance of the people to process the inaccurate information obtained from the web usage data. at first, before deriving the set of fuzzy rules, the clustering process is done based on users, session and the web pages they viewed. after forming the user groups, fuzzy inference system is presented. the fuzzy prediction follows the fuzzy ifthen rules for identifying the user behaviors. the fuzzy prediction model to identify the user behavior is processed based on subsequent steps. • choose appropriate input and output variables. • form a group based on the web pages, users and sessions they entered • select a appropriate kind of fuzzy prediction method • design a set of fuzzy if-then rules (information base) • predict the behavior of the user 368 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique fuzzy prediction-based user behavior. the fuzzy prediction subsystem is a unique authority scheme. it utilizes an information support, articulated in terms of fuzzy implication rules, and a suitable implication engine to identify the active time of the user in the respective group. the facts based on the identification of the users behavior is considered as, • identify the relation between the user visited web pages and their behaviors • predict the rules based on their behaviors the primary step is to obtain the inputs and decide the amount to which they fit in to each of the suitable fuzzy sets by means of association functions. after the efforts are fuzzified, we recognize the extent to which every part of the visited web pages are used for every rule. the fuzzy rules are then formed once the groups are generated, based on user, web pages visited and sessions. in each set of fuzzy rule, the subsequent steps have to be followed. • assign each rule contain a binary weight (0 and 1). • consider weight of each rule is 0 which has no outcome of the behavior of the user. • assign appropriate weight to each of the rule formed to identify the behavior of the user. • apply inference technique to the corresponding rules design process of fuzzy prediction-based user behavior.the design process of fuzzy prediction-based user behavior comprises of two steps. • fuzzification • fuzzy assumption engine: a. fuzzy implication rules; b. fuzzy implication engine. in the first step of fuzzy prediction-based user behavior, fuzzification is applied where numerical value is converted to qualitative value based on the user input called as linguistic variables. these linguistic variables are explained using membership function which has a value between zero and one. let linguistic variable lm,j be the received input from the user at time tm with probability pm,i. for our study, three linguistic variables called triangular membership function are selected and stored in ulm,j which consists of elements, short (s), normal (n), and large (l) given as follows. triangle(y, s,n,l) =  0, y < s y − s/n − s, s ≤ y ≤ n l− y/l−n,n ≤ y ≤ l 0, l ≤ y (1) after the fuzzification process is completed the results are fed to fuzzy assumption engine which comprises of fuzzy implication rules and fuzzy implication engine. the fuzzy implication rules consist of a set of if-then statements known as linguistic rules which describes the behavior of the user for particular set of inputs advances in systems science and applications (2014) vol.14 no.4 369 at a given time. the rule set for deriving user behavior using fuzzy assumption engine for ith rule is given below: ri = { f (lm,0isa0,i) and (lm,1isa1,i) and (lm,2isa2,i) then (pm,0ispo,i) and (pm,2isp1,i) and (pm,2isp2,i) (2) where i = 1,2,...,i, is the total number of fuzzy rules for predicting user behavior using fuzzy adaptive preference techqnique. (lm,0, lm,1, lm,2) ∈ ulm,0, ..., and (pm,0, pm,1, pm,2 ∈ upm,0,.... denote linguistic variables, (a0,ia1,ia2,i, po,ip1,ip2,i) denotes the fuzzy set present in ulm,0 and upm,0 respectively. the fuzzy implication engine then evaluates the set of rules to compute a qualitative output result and matches with the visited web page which generates nine rules for analyzing the user behavior. the procedure below describes the process of the proposed adaptive fuzzy prediction technique shown below: input:set of parameters like product demand, quantity of products, features of products step 1: recognize the parameters that finest suits the fuzzy system for inputs, positions, and the outputs step 2: divide the discourse or the period extent by every variable step 3: significance into a number of fuzzy subsets step 4: conveying all a linguistic label step 5: assign or decide an association utility for every fuzzy subset step 6: determine the fuzzy relationships among the inputs, states fuzzy subsets step 7: determine the fuzzy subsets for outputs step 8: form the rule base step 9: standardize the constraint variables step 10: select suitable scaling features for the input and output variables step 11: fuzzify the inputs to the checker step 12: use fuzzy logic to deduce the output contributed from every rule step 13: amassed the fuzzy outputs suggested by every rule step 14: concern defuzzification to shape a crisp output the above procedure determined the process of the fuzzy sets to identify the set of users behavior based on their activities over the demand and qualities of products. the next section will describe about the experimental evaluation of the proposed scheme. 4 experimental evaluation experimentations are conduct at all the web usage data obtained from web server logs with real data sets extracted from websites. performance evaluation is 370 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique made to show the effectiveness of the proposed evaluation of users behavior using adaptive fuzzy prediction technique [eubafp] compared to the existing one referred from different contributors of web usage mining. a sample set of data has been taken for experimental evaluation by taking five sets of attributes. the attributes used here are product demand, quantity of products, features of the product/services, annual sales, etc., to identify the behavior of the user based on their actions. table 1 web server logs dataset details the attributes are analyzed with the multi-clustered web usage data generated in the first phase. the predictive feature of the web usage data value is generated based on the fuzzy model across different time periods. the performance of the proposed evaluation of users behavior using adaptive fuzzy prediction technique [eubafp] is measured in terms of • implementation cost • prediction accuracy • execution time implementation cost.implementation cost for adaptive fuzzy prediction technique refers to the cost incurred building the procedure based on the number of users in clustered parts, pattern, and running, testing, and building essential changes. the implementation cost is evaluated using the formula with simulated time (st) and the simulation time (snt), to predict the user behavior for t interval of time is derived below: advances in systems science and applications (2014) vol.14 no.4 371 ic (n) = snt (n) /st (3) prediction accuracy.prediction accuracy defines the accuracy rate for predicting the users behavior based on adaptive fuzzy prediction technique towards product demand, quantity of products, features of the product/services, annual sales. the prediction accuracy is evaluated using the formula given below: pm,j ∗ = n ∗ pm,j 1 + (n − 1) ∗ pm,j (4) where n denotes the total number of test conducted for a specific user at varied interval of time and pm,j denotes the reliability of adaptive fuzzy prediction technique. execution time.the parameter execution time using adaptive fuzzy prediction technique refers to the time taken to execute, towards predicting the behavior of users product demand in network environment. exect ime = (insncount)i ∗ (clockcyclet ime)i (5) where execution is evaluated by the products of number of instructions to be executed for n users (i=1,2,3..,n) and clock cycle time for the specific ith user (i = 1,2,3...,n). 5 results and discussion in this work, we have seen that the proposed technique identified the behavior of the user based on their activities present in the network environment. the predictive feature of the web usage data are identified based on the fuzzy prediction model across different time periods. the below table and graph describes the performance of the proposed technique and compared the results with the existing adaptive fuzzy regression model (afr) [1] and user navigation pattern discovery using fast adaptive neuro-fuzzy inference system (unpd) for mining user profiles referred from different contributor s of web usage mining. fig. 4 describes the implementation cost required to handle the fuzzy inference systems based on the number of users in the clustered parts. the implementation cost for the fuzzy inference system in the proposed eubafp is low in the sense that adaptive fuzzy prediction technique is denoted by binary weighted model that results in less implementation cost than compared to the other two models afr and unpd. when the number of rules generated is small, its implementation cost is also less. as a result, the eubafp proposed system is realistic even when the number of users is large with a variance of 50-60% when compared to afr. 372 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique table 2 number of users vs. implementation cost. the implementation cost required to build the fuzzy inference system based on the number of users present no. of users implementation cost proposed eubafp existing afr existing unpd 5 7 25 28 10 11 30 34 15 15 38 40 20 19 40 43 25 22 46 49 30 26 50 53 35 30 54 58 fig.4 no. of users vs. implementation cost table 3 number of users vs. prediction accuracy. the behaviors of users are identified by implementing the fuzzy inference systems by diverse set of approaches no. of users prediction accuracy (%) proposed eubafp existing afr existing unpd 5 65 55 40 10 70 62 45 15 72 64 51 20 78 68 54 25 80 70 59 30 82 71 62 35 84 74 64 advances in systems science and applications (2014) vol.14 no.4 373 fig.5 describes the prediction accuracy of users behavior based on the number of users present in it. from the figure it is evident that an increase in number of users also results in increased rate of prediction accuracy. when compared to the other two techniques afr and unpd, the proposed adaptive fuzzy prediction technique as shown in figure 5, achieves higher prediction accuracy. the higher prediction accuracy in adaptive fuzzy prediction technique is due to combining of both fuzzy prediction system and behavior predictor model which results in variance of 10-15% when compared to afr. fig.5 number of users vs. prediction accuracy the proposed eubafp construction presents the benefit of execution time. compared to the existing works like adaptive fuzzy regression model (afr) [1] and user navigation pattern discovery using fast adaptive neuro-fuzzy inference system (unpd), the proposed eubafp provides an improvement in execution time in predicting users behavior. fig. 6 describes the efficiency of fuzzy prediction systems in terms of execution to identify the behavior of users in the network environment. rather than using other existing works for user behaviors, the proposed eubafp achieves better efficiency in predicting the users behavior in terms of time with a variance of 1015% when compared to two other methods, using defuzzificaiton, center of area. as eubafp consumes less execution time and proceed with a set of rules, the efficiency in predicting the user product demands is also being high. finally it is being observed that the proposed eubafp performs well in terms of prediction accuracy and efficiency with less implementation cost. 374 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique table 4 number of users vs. efficiency. the efficiency of the fuzzy prediction in terms of execution time for diverse set of approaches is illustrated no. of users execution time (sec) proposed eubafp existing afr existing unpd 5 32 40 50 10 40 43 55 15 46 52 65 20 50 58 70 25 55 65 75 30 60 70 80 35 62 75 85 fig.6 number of users vs. execution time 6 conclusion an adaptive fuzzy prediction system is developed to foresee the possibilities that the behavior user will be dynamic in the close by clusters at prospect instants employing the real-time dimension data established at the user. the merits of the adaptive fuzzy prediction system recline in • simplicity it is a one-way build-up method that does not need timeconsuming on-line accessing, • usefulness the possibilities are serious for balancing proficient operation of the web usage data of users, and • low cost the expected possibilities are attained supported on the appropriate demands of the user profiles. experimental results reveal that the performance of the adaptive fuzzy predicadvances in systems science and applications (2014) vol.14 no.4 375 tion system depends on the extent of behaviors of the user based on their product demands, and on the accessibility of information which confines the number of probable behaviors of the user. taking into report that the overall inference accuracy can be appreciably increased with value multiplexing, the fuzzy prediction scheme presents a fine resolution to getting users information. references [1] pankaj nagar and sumit srivastava, (2008), “adaptive fuzzy regression model for the prediction of dichotomous response variables using cancer data: a case study”, journal on applied mathematics, statistics and informatics (jamsi), 4. [2] belsare satish and patil sunil, (2012), “study and evaluation of user’s behavior in e-commerce using data mining”, research journal of recent science, vol.1(isc-2011), pp.375-387. [3] j vellingiri, dr. s.chenthur pandian, (2011), “user navigation pattern discovery using fast adaptive neuro-fuzzy inference system”, ijcsi international journal of computer science issues, vol.8, issue 6, no.1, november. [4] t. dubois, j. golbeck, and a. srinivasan, (2009), “rigorous probabilistic trust inference with applications to clustering”, in proceedings of the 2009 ieee/wic/acm international joint conference on web intelligence and intelligent agent technology-volume 01. ieee computer society, pp.655c658. [5] j. leskovec, d. huttenlocher, and j. kleinberg, (2010), “predicting positive and negative links in online social networks”, in proceedings of the 19th international conference on world wide web. acm, pp.641c650. [6] u. kuter and j. golbeck, (2010), “using probabilistic confidence models for trust inference in web-based social networks”, acm trans. internet technol., vol.10, no.2, pp.1c23. [7] dara, r.a. , makrehchi, m. , kamel, m.s. (2010), “filter-based data partitioning for training multiple classifier systems”, ieee transactions on knowledge and data eng, vol.22, no.4, april. [8] sharadh ramaswamy and kenneth rose, (2011), “adaptive cluster distance bounding for high-dimensional indexing”, ieee computer society, july. 376 joan m.: evaluation of users’ behavior using adaptive fuzzy prediction technique [9] lu, e.h.-c., tseng, v.s. , yu, p., (2011), “mining cluster-based temporal mobile sequential patterns in location-based service environments”, ieee transactions on knowledge and data eenineering, vol.23, no.6, june. [10] jung-yi jiang , kaohsiung, taiwan , ren-jia liou , shie-jue lee, (2011), “a fuzzy self-constructing feature clustering algorithm for text classification”, ieee transactions on knowledge and data eenineering, vol.23, no.3, march. [11] pierrakos, d. , paliouras, g, (2010), “personalizing web directories with the aid of web usage data”, ieee transactions on knowledge and data engineering. [12] nasraoui, o. ,soliman, m. ,saka, e. ; badia, a., (2008), “a web usage mining framework for mining evolving user profiles in dynamic web sites”, ieee transactions on knowledge and data engineering. [13] ruiyun qi, nanjing, gang tao , bin jiang , chang tan, (2012), “adaptive control schemes for discrete-time tcs fuzzy systems with unknown parameters and actuator failures”, ieee transactions on fuzzy systems. [14] sun zl, au kf, choi tm , (2007), “a neuro-fuzzy inference system through integration of fuzzy logic and extreme learning machines”, ieee transactions on systems, man, and cybernetics, part b: cybernetics. [15] kim ss, kwak kc, (2010), “development of quantum-based adaptive neuro-fuzzy networks”, ieee transactions on systems, man, and cybernetics, part b: cybernetics. [16] ajil, k.s. , lulea, sweden ,shukla, m.v. , pal, p.k., (2010), “a new technique for temperature and humidity profile retrieval from infraredsounder observations using the adaptive neuro-fuzzy inference system”, ieee transactions on geoscience and remote sensing. [17] aiazzi, b. , baronti, s. , (2002), “fuzzy logic-based matching pursuits for lossless predictive coding of still images”, ieee transactions on fuzzy systems. [18] angelov, p. , kordon, a., (2010), “adaptive inferential sensors based on evolving fuzzy models”, ieee transactions on systems, man, and cybernetics, part b: cybernetics. [19] abdel-hamid, w. , noureldin, (2007), “adaptive fuzzy prediction of lowcost inertial-based positioning errors”, ieee transactions on fuzzy systems. advances in systems science and applications (2014) vol.14 no.4 377 [20] ordonez, r. , spooner, j.t. ,passino, k.m., (2006), “experimental studies in nonlinear discrete-time adaptive prediction and control”, ieee transactions on fuzzy systems. [21] wai-keung wong , enjian bai , chu, a.w., (2010), “adaptive time-variant models for fuzzy-time-series forecasting”, ieee transactions on systems, man, and cybernetics, part b: cybernetics. [22] khosravi, a. , nahavandi, s. , creighton, d.,(2011), “prediction interval construction and optimization for adaptive neurofuzzy inference systems”, ieee transactions on fuzzy systems. [23] wang, w. , vrbanek, j., (2008), “an evolving fuzzy predictor for industrial applications”, ieee transactions on fuzzy systems. corresponding author author can be contacted at: joanmjohn123@yahoo.co.in adv syst sci appl 2020; 04; 27-35 published online at https://ijassa.ipu.ru. determining the vertical force when steering nguyen tuan anh1*, hoang thang binh2 1) automotive engineering department, thuyloi university, 175 tay son, dong da, hanoi, vietnam e-mail: anhngtu@tlu.edu.vn 2) school of transportation engineering, hanoi university of science and technology, 1 dai co viet, hai ba trung, hanoi, vietnam e-mail: binh.hoangthang@hust.edu.vn abstract: when the vehicles move at high velocity and quickly steer, the vehicles can be rollover. the first sign of this phenomenon is that two wheels on the same side are completely separated from the road surface. the typical value for this sign is the vertical force fz at each wheel. if this value gradually decreases to zero, the wheel runs the risk of separating from the road surface that may lead to rollover. therefore, the value of the vertical force fz in different conditions needs to be determined to detect the imminent limit of the rollover phenomenon. this research established the spatial dynamic model of the vehicle to determine the vertical force fz based on the simulation method. besides, the equation describing the relationship between vertical force fz, velocity v, and acceleration steering  corresponding to the value of steering angle  is also established by the calculation process. from this equation, the value of the vertical force fz can be simply calculated with relatively high accuracy based on determining conditions. the results of this research are the basis for determining and establishing the vehicle's rollover limit. keywords: dynamic vehicle model, vertical force, function, steering acceleration, rollover 1. introduction 1.1. the instability problem of the vehicle automobiles are a popular vehicle in everyday life. however, the number of vehicle accidents is also very large. when the vehicles move at high velocity and quickly steer, the vehicles often encounter the situation: side slip or rollover. the side slip phenomenon usually occurs when the vehicles go on slippery roads, wheels are not able to contact the road surface. when the side slip occurs, the driver could not control the trajectory and direction of the motion of the vehicles. the sideslip problem is usually less dangerous than the rollover problem. the rollover phenomenon occurs when two wheels on the same side are completely separated from the road surface. if there is only one wheel separated from the road surface, the vehicle will be in an unstable state and at risk of rollover [15, 17]. the main cause of this phenomenon is due to the lateral acceleration ay appeared suddenly when quickly steering, especially in the case of fishhook steering [11, 18]. the parameter that warns the risk of rollover is the vertical force at wheels fz. if this value is sufficiently close to zero, the wheel runs the risk of separating from the road surface. therefore, if it is possible to determine the vertical force value fz, the timing of the rollover of the vehicle can be predicted. some solutions have been proposed to improve this situation such as the use of the active stabilizer bar, the electronic power steering, the active suspension, the electronic stability * corresponding author: anhngtu@tlu.edu.vn 28 n.t. anh, h.t. binh copyright ©2020 assa. adv. in systems science and appl. (2020) program, [etc.] [1, 16, 19]. the above systems use parameters of input signals such as the lateral acceleration ay, the roll angle , the vertical force at wheels fz [5, 20]. 1.2. literature review the topic of "rollover vehicle" has been researched by many authors. several authors have introduced a rollover index r [6, 8, 9, 10, 14]. this index is determined based on the difference in the vertical force of the wheels. if |r| <1, the car operates stably. in the case of |r| = 1, both wheels on one side are separated from the road surface, the vehicle rollovers completely. however, this index only really makes sense in case both wheels of the same side are separated from the road surface. if only one of the wheels is lifted off the road surface, the rollover index r cannot be determined. besides, the value of lateral acceleration ay is often used to determine the rollover threshold of the vehicles [3, 7, 13]. the value of the lateral acceleration ay can be calculated and simulated through link equations or experiments. however, this value is only really meaningful in case of ignoring the influence of the dimensions of the vehicles. some other studies use the roll angle of the vehicle  to determine the vehicle's rollover limited [2, 4, 12]. the value of the roll angle of the vehicles  is calculated by the lateral acceleration ay, which includes the influence of the dimensions. to accurately determine the limits of the vehicle instability, it is necessary to find the time when the wheels are separated from the road surface (fz = 0). this research focuses on identifying the limits at which the wheels are separated from the road surface. at the same time, the research also established the equation showing the relationship between vertical force fz, steering angle , and steering acceleration . therefore, it is easy to determine the vertical force fz at different times and conditions. 2. methodology 2.1. nomenclature : pitch angle, rad : roll angle, rad : steering angle, rad : yaw angle, rad ay: lateral acceleration, m/s2 b: half of the base width, m cij: damping coefficient, ns/m fcij: force of the damper, n fkij: force of the spring, n fktij: force of the tire, n fxij: longitudinal force, n fyij: lateral force, n fzij: vertical force, n h: distance from center to roll axis, m iz: moment of inertia of the z-axis, kgm2 ix: moment of inertia of the x-axis, kgm2 kij: stiffness of the spring, n/m ktij: stiffness of the tire, n/m l1: distance from the center to the front axle, m l2: distance from the center to the rear axle, m m: sprung mass, kg mij: unsprung mass, kg determining the vertical force when steering 29 copyright ©2020 assa. adv. in systems science and appl. (2020) uij: bump on the road, m vx: longitudinal velocity, m/s vy: lateral velocity, m/s z: displacement of the sprung mass, m zij: displacement of the unsprung mass, m 2.2. double-track dynamic vehicle model the motion of the vehicle is set based on the model of 10 degrees of freedom. the doubletrack model is set as below (fig. 2.1). assuming that the vehicle is moving at a constant velocity, the steering angle is small. ignore the influence of other factors. the motion determination matrix is described as follows [2]: 1 2 z ij i; j = 1 y xy1 y2 2 2 z i;j = 1 l1 im + m (v ψ) = (f f ) v (ψ 0) l1 im + m ij                 (2.1) where: y1 y11 y12f = f + f y2 y21 y22f = f + f fig. 2.1. the double-track model 2.3. spatial dynamic model 7 dof to determine the oscillation of the vehicle, it is necessary to establish the spatial dynamic model with 7 degrees of freedom as fig. 2.2. 30 n.t. anh, h.t. binh copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 2.2. dynamic vehicle model 7 dof the matrix of vehicle oscillation is given as follows [2]: 22 yx 1 2 2 x 2 x 1 b m (a cos φ + gsinφ)mhi + mh (z φ) = (f f ) + 0 1 b i + mh m i + mh                    (2.2) where: 1 c11 k11 c21 k21f = f + f + f + f 2 c12 k12 c22 k22f = f + f + f + f to be able to identify the above matrix, use the corresponding link equations. the vertical displacement of the unsprung mass: ij ij ktij cij kij m z = f f f (2.3) the elastic force of the suspension system: kij ij ij f = k (z z ± bφ) (2.4) the damping force of the suspension system: cij ij ij f = c (z z ± bφ) (2.5) the elastic force of the tire: ktij tij ij ij f = k (u z ) (2.6) the vertical force fz is calculated by using the equation below: zij ktij cij kij f = f f f (2.7) when the vertical force fz becomes close to zero, the wheel tends to separate from the road surface, the vehicle is in an unstable state. determining the vertical force when steering 31 copyright ©2020 assa. adv. in systems science and appl. (2020) 3. result and discussion 3.1. simulation conditions the established model is used for most of the common vehicles. to conduct simulation and review, the specifications of the reference vehicles are given in table 3.1. table 3.1. vehicle specifications symbol description value unit m sprung mass 2000 kg mij unsprung mass 45 kg l1 distance from center of gravity to front axle 1250 mm l2 distance from center of gravity to rear axle 1550 mm b half of the base width 750 mm ix moment of inertia of the x–axis 800 kgm2 iz moment of inertia of the z–axis 2400 kgm2 h distance from the center to roll axis 650 mm at the velocity v = 70 km/h, the steering angle  = 0–50 (linear increase in 1 second), the graph in fig. 3.1 shows the value of the vertical force fz at the wheels at the same time. fig. 3.1. the value of the vertical force fz from fig. 3.1 it can be seen that the value of the fz21 is the smallest, this wheel tends to separate from the road surface first. therefore, simulation is concentrated at this position to determine the time of separating the wheels from the road surface. the research will simulate the oscillation of the vehicle with two cases: case 1: steering angle  = 0–40 (linear increase), corresponding to the different values of steering acceleration and velocity of the vehicle. case 2: steering angle  = 0–50 (linear increase), corresponding to the different values of steering acceleration and velocity of the vehicle. 32 n.t. anh, h.t. binh copyright ©2020 assa. adv. in systems science and appl. (2020) 3.2. results using the parameters of the reference vehicle and the two proposed cases, the graph in fig. 3.2 shows the relationship between steering acceleration , velocity v, and vertical force fz21 of the wheel at position (21). the graph in fig. 3.2 indicates that: + at the same value of the velocity, if the steering acceleration increases, the vertical force at the wheel will decrease. in case the velocity value is small, this decrease is not much. otherwise, if the velocity value is large, this attenuation will be significant. + at the same value of steering acceleration, if the velocity increases, the vertical force will drop sharply. this decrease is almost linear (in case the value of steering acceleration is large). + at the same value of velocity and steering acceleration, if the steering angle is larger, the value of vertical force will be smaller. the relationship between the parameters in fig. 3.2 has the form of irregular planes. therefore, it is possible to set up many plane equations to determine the vertical force value fz when the velocity v and steering acceleration  are known. the equation for determining the value of the vertical force fz is as follows: z 1 2 3f = a + a ε + a v (3.1) where: a1, a2, and a3 are the coefficients of the equation for each specific determination interval. based on the results obtained from the simulation process, the above coefficients are given in table 3.2. the coefficients a1, a2, and a3 correspond to different values of steering angle and steering acceleration. fig. 3.2. the relationship between steering acceleration, velocity, and vertical force table 3.2. the coefficients of the equation steering angle (deg) steering acceleration (deg/s2) the coefficients of the equation a1 a2 a3 40   0.14 4989.07 – 71.94 – 43.73 40 0.07   < 0.14 5064.91 – 913.04 – 43.57 40  < 0.07 5099.01 – 2571.43 – 42.80 50   0.14 5023.14 – 136.69 – 54.07 50 0.07   < 0.14 5127.43 – 1434.78 – 53.63 50  < 0.07 4704.57 – 2714.29 – 48.03 determining the vertical force when steering 33 copyright ©2020 assa. adv. in systems science and appl. (2020) the vertical force value fz can easily be determined through the coefficients of the equation. table 3.3 shows the difference between the results of calculating vertical force fz by two different methods (simulation and equation). table 3.3. the value of the vertical force fz velocity (km/h) steering acceleration (deg/s2) steering angle  = 0–40 steering angle  = 0–50 fz simulation (n) fz equation (n) tolerance (%) fz simulation (n) fz equation (n) tolerance (%) 60 0.279 2345 2345 0.0 1741 1741 0.0 60 0.140 2354 2355 0.0 1757 1760 0.2 60 0.093 2366 2366 0.0 1776 1776 0.0 60 0.070 2377 2387 0.4 1784 1810 1.4 60 0.056 2387 2387 0.0 1787 1771 0.9 60 0.047 2388 2411 1.0 1791 1797 0.3 60 0.040 2389 2428 1.6 1795 1815 1.1 60 0.035 2391 2441 2.0 1796 1828 1.8 70 0.279 1877 1908 1.6 1165 1200 2.9 70 0.140 1884 1918 1.8 1179 1219 3.3 70 0.093 1896 1930 1.8 1199 1240 3.3 70 0.070 1909 1951 2.2 1216 1273 4.5 70 0.056 1919 1959 2.0 1224 1191 2.8 70 0.047 1926 1983 2.9 1231 1216 1.2 70 0.040 1931 2000 3.5 1237 1234 0.2 70 0.035 1939 2013 3.7 1241 1248 0.6 80 0.279 1440 1471 2.1 635 659 3.6 80 0.140 1448 1481 2.2 649 678 4.3 80 0.093 1461 1494 2.2 674 704 4.3 80 0.070 1477 1516 2.6 705 737 4.3 80 0.056 1493 1531 2.5 722 711 1.5 80 0.047 1506 1555 3.2 736 736 0.0 80 0.040 1516 1572 3.6 750 754 0.5 80 0.035 1525 1585 3.8 760 767 0.9 90 0.279 1033 1033 0.0 119 119 0.0 90 0.140 1043 1043 0.0 138 138 0.0 90 0.093 1059 1059 0.0 167 167 0.0 90 0.070 1080 1080 0.0 200 201 0.5 90 0.056 1103 1103 0.0 230 230 0.0 90 0.047 1123 1127 0.4 253 256 1.2 90 0.040 1142 1144 0.2 273 274 0.4 90 0.035 1157 1157 0.0 287 287 0.0 34 n.t. anh, h.t. binh copyright ©2020 assa. adv. in systems science and appl. (2020) the data in the table show that there is a difference between the two methods, but the tolerance is quite small. in the case of steering angle  = 0–40, the biggest tolerance is only 3.8%. in the case of steering angle  = 0–50, this value does not exceed 4.5%. in general, this difference is not much and this equation can be used in many different cases. 4. conclusion the stability and safety of the vehicle are expressed through the value of vertical force at the wheel fz. the velocity, steering angle, and steering acceleration have a great influence on this problem. when the steering angle increases or the steering acceleration increases or both factors increase, the value of the vertical force fz will decrease significantly. if the value of fz gets close to zero, the wheels tend to separate from the road surface. at this time, the vehicle is in an unstable situation and very dangerous. to determine the value of the vertical force fz at different velocities, steering angle, and steering acceleration, the established dynamic vehicle model can be used. however, the setup and simulation process is complicated and inconvenient. this research gave the function of determining the value of the vertical force fz at specific conditions. therefore, it is easy to identify fz without using the dynamic vehicle model to simulate. the function is not perfectly accurate, there is a difference in the vertical force value fz when determined by this function and simulation process. however, this difference is quite small (less than or equal to 3.8% when the steering angle is from 00 to 40 and less than or equal to 4.5% when the steering angle is from 00 to 50). the results of the research can be used in many different survey cases. this research has only focused on the theoretical calculation process and establish corresponding equations, it is necessary to experiment to accurately evaluate the results. the results of this research will be the basis for the development of other research in the future. references 1. anh n.t. & binh h.t. (2019). research on dynamic vehicle model equipped active stabilizer bar, advances in science, technology and engineering systems journal, 4(4), 271–275. doi: 10.25046/aj040434. 2. anh n.t. & binh h.t. (2020). research on determining the limited roll angle of vehicle, lecture notes in networks and systems, 104, 613–619. doi: 10.1007/978-3-03037497-6_70. 3. cruz p., echaveguren t. & gonzalez p. (2017). estimation of heavy vehicle rollover potential using reliability principles, revista ingenieria de construction, 32(1), 5– 14. doi: 10.4067/s0718-50732017000100001. 4. ding e.l. & massel t. (2005). estimation of vehicle roll angle, proceedings of the 16th triennial world congress. prague, czech republic, 122–127. doi: 10.3182/200507036-cz-1902.01908. 5. eisele d.d. & peng h. vehicle dynamics control with rollover prevention for articulated heavy trucks, proceedings of the 5th symposium on advanced vehicle control, michigan, usa, 1–8. 6. jin z., li j., huang y. & khajepour a. (2019). study on rollover index and stability for a triaxle bus, chinese journal of mechanical engineering, 32(2), 1–15, 2019. doi: 10.1186/s10033-019-0376-0. 7. kazemian a.h., fooladi m. & darijani h. (2017). rollover index for the diagnosis of tripped and un-tripped rollovers, latin american journal of solids and structures, 14(11), 1979–1999. doi: 10.1590/1679-78253576. determining the vertical force when steering 35 copyright ©2020 assa. adv. in systems science and appl. (2020) 8. li b. & bei s. (2019). research method of vehicle rollover mechanism under critical instability condition, advances in mechanical engineering, 11(1), 1–11. doi: 10.1177/1687814018821218. 9. licea m.a.r., rodriguez e.a.v, pinal f.j.p. & olivares j.p. (2018). the rollover risk in delta tricycles: a new rollover index and its robust mitigation by rear differential braking, mathematical problems in engineering, 29(1), 1–14. doi: 10.1155/2018/4972419. 10. lin z., li b., li j. & khajepour a. (2019). dynamic stability and control of tripped and un-tripped vehicle rollover. san rafael. doi: 10.2200/s00916ed1v01y201904aat006. 11. moshchuk n., mousseau c., norman k. & motors g. (2008). simulation study of oscillatory vehicle roll behavior during fishhook maneuvers. proceedings of the 2008 american control conference. seattle, washington, usa, 3933–3940. doi: 10.1109/acc.2008.4587107. 12. parczewski k. & wnek h. (2017). the influence of vehicle body roll angle on the motion stability and maneuverability of the vehicle, combustion engines, 168(1), 133–139. doi: 10.19206/ce-2017-121. 13. phanomchoeng g. & rajamani r. (2011). new rollover index for detection of tripped and un-tripped rollovers, proceedings of the 50th ieee conference on decision and control and european control conference. orlando, florida, usa, 7440–7445. doi: 10.1109/cdc.2011.6160823. 14. rajamani r., piyabongkam d., tsourapas v. & lew j. y. (2011). parameter and state estimation in vehicle roll dynamics, ieee transactions on intelligent transportation systems, 12(4), 1558–1567. doi: 10.1109/tits.2011.2164246. 15. sakurai t. (2015). overview of research of rollovers, journal of traffic and transportation engineering, 3(1), 1–6. doi: 10.17265/2328-2142/2015.01.001. 16. schofield b., hagglund t. & rantzer a. (2006). vehicle dynamics control and controller allocation for rollover prevention, proceedings of the 2006 ieee conference on computer-aided control system design, 2006 ieee international conference on control applications, 2006 ieee international symposium on intelligent control. munich, germany, 149–154. doi: 10.1109/cacsd-cca-isic.2006.4776639. 17. wu x., ge x., luo s. & huang h. (2010). research on vehicle rollover and control, proceedings of the 2nd international conference on advanced computer control. shenyang, china, 510–514. doi: 10.1109/icacc.2010.5486695. 18. xiong f., lan f., chen j. & xiong y. z. (2015). the study for anti-rollover performance based on fishhook and j turn simulation, proceedings of the 3rd international conference on material, mechanical and manufacturing engineering. guangzhou, china, 2084–2093. doi: 10.2991/ic3me-15.2015.401. 19. yu f., li d.f. & crolla d.a. (2008). integrated vehicle dynamics control, proceedings of the 2008 ieee vehicle power and propulsion conference. harbin, china, 1– 6. doi: 10.1109/vppc.2008.4677809. 20. zulkarnain n., imaduddin f., zamzuri h. & and mazlan s.a. (2012). application of an active anti-roll bar system for enhancing vehicle ride and handling, proceedings of the 2012 ieee colloquium on humanities, science and engineering, kota kinabalu, malaysia, 260–265. doi: 10.1109/chuser.2012.6504321. 1. introduction 1.1. the instability problem of the vehicle 1.2. literature review 2. methodology 2.1. nomenclature 2.2. double-track dynamic vehicle model 2.3. spatial dynamic model 7 dof 3. result and discussion 3.1. simulation conditions 3.2. results 4. conclusion adv syst sci appl 2021; 02:58–70 published online at https://ijassa.ipu.ru. case study: influence of muscle fatigue and perspiration on the recognition of the emg signal narek unanyan1*, alexey belov1* 1v.a. trapeznikov institute of control sciences of russian academy of sciences, moscow 117997, russia abstract: emg data processing and muscle activity recognition has become the most popular method for upper limb prosthetics. the high sensitivity of emg sensors with respect to external disturbances and other factors prevent from accurate muscle activity recognition. the aim of the paper is to investigate robustness of window recognition method with respect to muscle fatigue and perspiration of the forearm skin. the current experiment was carried out using arduino nano microcontroller connected to emg sensors. the subject under study is a healthy man of 26 years old with an average build. the subject was asked to do physical exercises, thereby loading the muscles of the fingers of the hand to achieve partial or complete fatigue and perspiration. during the whole process, emg sensors have installed on the subject and transmitted the signal to the computer using arduino. all signal processing is done directly on the computer with a pre-recorded signal. experimental results have been shown that with the appearance of external factors during prosthesis operation recognition accuracy may degrade to unsatisfactory. false positives occur with perspiration of skin surface and complete muscle fatigue. an algorithm for automatic self-correction of the boundaries of motion detection zones has been introduced. instead of identification of causes that leads to performance degradation, we use correction scheduling started by timer. experimental results have shown that proposed automatic adaptive correction is effective. despite higher recognition delay, proposed auto-tuning method provides satisfactory muscle activity identification and feature extraction in real-time. keywords: sensors, electromyogram, myoelectric control, prosthetics, pattern recognition 1. introduction in recent decades, and especially in recent years, much effort has been made to implement efficient control algorithms based on processing of electromyographic (emg) signals [1]. emg signals provide easy and non-invasive access to the physiological processes that cause muscle contraction. that is why emg technique is extensively used in control of upper limb prostheses. since the first attempts in the late 1940s [2], several emg-based algorithms have been developed to improve the functionality and ease of use of hand prostheses [3]. currently, emg signal processing is the most common approach used to manipulate prosthetic hands. limitations in prosthetic mechanics and emg data processing are currently the main restriction to the development and implementation of a complete bioelectric prosthesis [4–7]. moreover, when it comes to recognizing myoelectric patterns, the big problem is that statistical properties of emg signal change over time. this leads to the fact that control systems become unstable or difficult to use after a certain period of time [8]. in practice, it is very important to maximize emg data efficiency. to do this, it is necessary to assess the condition of the end user of the prosthesis, observing the condition of ∗corresponding author: n.unanyan@mail.ru case study: influence of muscle fatigue and perspiration 59 the skin, tissues, skeletal anatomy, muscle strength and activity, as well as the range of motion. also, an important factor is monitoring the shape, amplitude, frequency of the signal. only after a complete study of the user it is possible design the system for correct receiving and recognizing of emg signal [9]. the high sensitivity of emg sensors with respect to external disturbances and other factors prevent from accurate muscle activity recognition. such factors are investigated in [5]. among them: 1. muscle fatigue during work [10, 11]. 2. perspiration of the skin. 3. displacement of sensitive elements [12–14]. 4. physiological characteristics of a person [15, 16]. factors such as muscle fatigue and perspiration can lead to incorrect muscle activity recognition and feature extraction. the researchers argue that the accumulation of sweat under the sensor leads to a decrease in amplitude and filtering of high frequency components [17–20]. however, little is known about the amount of sweat above which a number of problems can arise. many authors also study the effect of muscle fatigue on recognition. [21] argue that with muscle fatigue, both the amplitude and frequency of the emg signal change. it was found that muscle fatigue is usually quantified as a decrease in maximum muscle strength or power, resulting in different recordings of signals from emg electrodes over time [22]. various traits were examined to assess muscle fatigue, such as the wavelet transform [23,24], the number of zero crossings [25], and autoregressive coefficients [26]. it is important to note that muscle fatigue identification commonly is carried out in frequency domain. this is cased by emg spectrum shift towards lower frequencies [5]. in this paper, we investigate influence of muscle fatigue and perspiration at feature selection accuracy performed by window method developed and presented in [27]. additionally, fail-safe correction to prevent false recognition is introduced. in the proposed method fatigue or perspiration identification technique is not used. this is caused by low computational requirements and real-time implementation of the algorithm. instead of identification of causes that leads to performance degradation, we use correction scheduling started by timer. the rest of the paper is organized as follows. sect. 2 describes experimental details and some necessary information on hardware and software equipment. sect. 3 presents main result of the paper. some conclusive remarks and future works are emphasized in sect. 4. 2. preliminaries 2.1. hardware description the current experiment was carried out using arduino nano microcontroller connected to emg sensors. arduino nano was selected due to its low cost and low power consumptions. the microcontroller itself is connected to the computer. all signal processing is done directly on the computer with a pre-recorded signal. matlab software is used for processing and analysis received data. to read the muscle activity we use emg sensors from df robotics. these sensors combine the filter circuit and the amplifier circuit. emg sensor amplifies the minimum electron diffraction signal within 1.5 mv by a factor of 1000 and suppress noise (especially frequency interference) using a differential input and an analog filter. the output signal is analog, which takes 1.5 v as the reference voltage. output voltage range 0–3 v. the signal level depends on the intensity of muscle activity. the output signal indicates muscle activity and contributes to the analysis and study of the emg signal. for laboratory experiment, sensor surface has been cleaned and degreased at each sensor. after that, 5 electrodes are placed on the upper and lower parts of the forearm in the correct order. most of the electrodes were located near the forearm (see fig. 2.1), since the flexors copyright © 2021 assa. adv syst sci appl (2021) 60 n. unanyan, a. belov fig. 2.1. location of emg sensors. and extensors of all fingers except the thumb are mainly located in this area. these muscles are oriented mainly parallel to the axis between the elbow and the wrist. unlike wearable band or sleeve such kind of electrode location provides better emg signal recording. 2.2. software tools before evaluating the influence of various factors on recognition, it is worth to briefly explain an emg signal recognition window method software. this algorithm is given and described in [27]. we try to recognize finger movements as close to its physiological behavior as possible. to this end we selected three levels of finger states. they are defined as follows: g = 1, the muscle is relaxed, 2, the muscle is in half tense, 3, the muscle is in full tense. in last expression, value g = 1 corresponds to straightened finger while values g = 2 and g = 3 correspond to half and fully bent finger respectively. briefly, this algorithm can be summarized as follows: 1. preparation of the sensor and signal reading. 2. signal preprocessing and normalization. preprocessing consists of sliding mav calculation within the window of length n . here n is chosen experimentally to provide satisfactory accuracy. normalized signal is calculated as absolute difference between sliding mav and current measurement. 3. boundary values ai and bi identification. 4. depending on the area in which the processed signal is located, a decision is made on the value of g as: g(xmax) = 1 if a1 < xmax ≤ b1, 2 if a2 < xmax ≤ b2, 3 if a3 < xmax ≤ b3, (2.1) copyright © 2021 assa. adv syst sci appl (2021) case study: influence of muscle fatigue and perspiration 61 fig. 2.2. emg data processing: a.) raw signal. b.) results of muscle activity recognition using window method. where xmax is a maximum value of normalized signal at the window of length n counts. the result of muscle activity recognition using window method is depicted in figure 2.2. window algorithm has been tested for quality and correct operation under normal conditions. also, the window method was compared with artificial neural networks, which made it possible to conclude about benefits of this method [27]. 2.3. methods identification algorithm, described in [27], was tested at 28 healthy people both males and females from 20 to 55 years old. identification error for both hands is depicted in figure 2.3. analyzing the data obtained during the experiment, it was found that there are no substantial differences in identification depending on age. female muscle activity is defined a little bit better than male. in view of this, it was decided to select a healthy man of 26 years old with an average build for further experiments. the subject was asked to do physical exercises, thereby loading the muscles of the fingers of the hand. during the whole process, emg sensors have installed on the subject and transmitted the signal to the computer using arduino. we note that emg sensor data for different sensors have the same waveform. so, the further investigation and illustrations are carried out at one selected sensor dataset. the subject, due to physical exertion, led the muscles of the finger to partial fatigue. after that, a signal was recorded with three types of finger states. when the subject led to complete fatigue signal values for three types of finger states was also recorded. after that these signals were analyzed. another experiment was to test the effect of perspiration on the emg signal. for this, the subject was placed in a special room with a high temperature where he also performed physical exercises. after that emg data recording was carried out similar to previous experiment. all the results of these experiments are discussed below. 3. main result copyright © 2021 assa. adv syst sci appl (2021) 62 n. unanyan, a. belov fig. 2.3. identification error statistics. 3.1. muscle fatigue the term muscle fatigue is used to describe a temporary decrease in one’s physical capacity of performing motions. several researchers investigate the impact of muscle fatigue using time domain or frequency domain characteristics of emg signal such that maximum amplitude or power spectrum density [23]. many efforts have been made to identify the level of fatigue. in order to mitigate for the effects of fatigue research focuses on collecting data from multiple levels of fatigue and utilizes the abundance of information. this requires recording more data than a simple classification case and creates different computational requirements to the system. window method described in previous section also may suffer from muscle fatigue. further we investigate effect of muscle fatigue on feature extraction by window method. in this experiment, the subject performed physical exercises with chest expander to load the muscles with work until muscle would be tired. experiment is stopped when muscle has reached a complete fatigue. during all of the process, emg recording was carried out. we are interested in states of partially tired and completely tired muscles. emg pattern of all three finger positions is depicted in figure 3.4. figure 3.5 shows the raw emg signal on an enlarged scale. one can see the difference between two states, namely untired hand and completely tired one. both signals indicate relaxed muscles. it is noticeable that noises appear in the channel with accumulation of fatigue. however, it can also be concluded that noises are not significant and can be neglected. the root mean square (rms) value is also not much different. rms for ideal conditions is equal 1.53 while partially tired and tired muscles gives rms value equal to 1.51 and 1.53 respectively. the number of crossings of the mean square value is also the same. however, the number of changes in the sign of the function and its amplitude are significantly different due to the appearance of noise. we should note that the best recognition accuracy is reached only when muscle is untired. partially tired muscles also provides satisfactory recognition. however, activity of tired muscle cannot be recognized correctly. in figure 3.9 one can see that half tensed muscle recognized with incorrectly almost in 50% cases. 3.2. perspiration now let’s check how the system will behave if, during operation, the subject in the places where the sensing element is placed is wet due to perspiration. for this, emg data were collected from the subject, after which he performed a series of physical exercises for perspiration. the measurement result can be seen in figure 3.7. copyright © 2021 assa. adv syst sci appl (2021) case study: influence of muscle fatigue and perspiration 63 fig. 3.4. comparison of muscle fatigue charts: a.) untired muscle; b.) partially tired muscle; c.) tired muscle. fig. 3.5. comparison of emg shapes of untired and tired muscles. it is shown that emg signal has sufficiently distorted. in particular, the amplitude decreased and the number of changes in the signal sign increased. the rms has not changed much. it is equal to 1.52. however, the number of rms crossings has been increased. with such strong interference, there is a possibility of false triggering of the recognition algorithm. figure 3.8 illustrates recognition results after perspiration. one can see that algorithm provides false recognition during operation if parameters ai and bi are not corrected. copyright © 2021 assa. adv syst sci appl (2021) 64 n. unanyan, a. belov fig. 3.6. results of muscle activity recognition for muscle fatigue: a.) untired muscle; b.) partially tired muscle; c.) tired muscle. fig. 3.7. comparison of waveforms with perspiration. results of muscle activity recognition under sweat skin surface has shown high probability of false recognition. figure 3.8 clearly shows that the boundaries that were copyright © 2021 assa. adv syst sci appl (2021) case study: influence of muscle fatigue and perspiration 65 fig. 3.8. results of muscle activity recognition operation for perspiration. calculated before perspiration are no longer suitable and during operation, if the boundaries are not automatically corrected, the user will lose control of the prosthesis. 3.3. malfunction correction experimental results have been shown that with the appearance of external factors during prosthesis operation recognition accuracy may degrade to unsatisfactory. therefore, it is required to develop correction algorithm that performs correction of boundaries ai and bi to provide fail-safe recognition. the most interesting case is automatic recalculation to perform fail-safe recognition without manual mode even during active operating of the prothesis. we propose the following boundary values correction strategy. the program code for recalculation of new boundaries starts every 5 minutes in automatic mode without user interaction. the 5 minutes time interval has been selected empirically after several experiments. time interval can be changed individually for each patient. the correction process is a cycle in which the system inspects the user’s activity. after initialization of the boundary correction, the system is switched to waiting mode and wait for muscle activities. after activity detected, the correction algorithm collects about 150 measurements for each type of muscle activity. when data for three types of muscle activity are collected, correction step is carried out between contiguous movements, for example, between relaxed and half tensed muscle. new boundaries are defined by the following relations: copyright © 2021 assa. adv syst sci appl (2021) 66 n. unanyan, a. belov a1 = 0, a2 = b1 = n∑ i=1 xrelaxed i + n∑ i=1 xhalf i 2n , a3 = b2 = n∑ i=1 xhalf i + n∑ i=1 x full i 2n , b3 = n∑ i=1 x full i n + 2.5, (3.2) where xrelaxed i , xhalf i , and x full i are maximum window value for relaxed, partially tensed and fully tensed muscle respectively, n is a number of counts for calculating the average value (n = 150). table 3.1. recognition accuracy of window method. muscle activity relaxed half tensed tensed fatigue no correction, % 12 52 28.8 fatigue with correction, % 3.5 8.2 2.5 perspiration no correction, % 24 19.5 6 perspiration with correction, % 12 9.8 4.5 results of muscle activity recognition before and after correction are given in table 3.1. it is easy to see that correction algorithm allows to improve accuracy of window method. also we should note errors correspond to the transient process between states of muscle activity. it has been experimentally obtained that overall delay in recognition between two states of muscle activity may be increased up to 100 ms. this is satisfactory value for realtime operations. additionally, we derive boundaries for ideal conditions (untired muscle and dry skin surface) and corrected values for each type of deflection. boundaries for ideal conditions are defined in (3.3), corrected values for tired muscle are defined by (3.4), and corrected values for sweaty skin are given by (3.5). g = 1 if 0 < xmax 6 0.11, 2 if 0.11 < xmax 6 0.42, 3 if 0.42 < xmax 6 3.3. (3.3) g∗ fatig. = 1 if 0 < xmax 6 0.03, 2 if 0.03 < xmax 6 0.3, 3 if 0.3 < xmax 6 3. (3.4) g∗ persp. = 1 if 0 < xmax 6 0.12, 2 if 0.12 < xmax 6 0.45, 3 if 0.45 < xmax 6 3.5. (3.5) copyright © 2021 assa. adv syst sci appl (2021) case study: influence of muscle fatigue and perspiration 67 fig. 3.9. algorithm operation after border correction for muscle fatigue. fig. 3.10. algorithm operation after border correction for perspiration. one can see that for tired muscle corresponding boundary values are less than ideal ones while corrected boundaries for sweaty skin surface are larger than ideal. the operation of the algorithm after border correction can be seen in figures 3.9 and 3.10. it is shown that proposed correction procedure provides improved recognition without substantial fails during operation. copyright © 2021 assa. adv syst sci appl (2021) 68 n. unanyan, a. belov 4. conclusion the paper presents an experimental investigation of the influence of muscle fatigue and perspiration on the recognition of muscle activity using emg signal. a number of experiments have been carried out. during study of the perspiration and muscle fatigue, it was found that complete fatigue may increase identification error to unsatisfactory (over 50%). perspiration as well as fatigue leads to performance degradation. to improve identification accuracy, automatic correction method has been proposed. experimental results has shown that proposed solution allows to improve working precision of the amplitude-based window identification under perturbing factors. it has been shown that when deflection from ideal conditions occurs, the algorithm adapts to changes. the future works is focused on: 1. study of other factors for emg signal recognition. 2. complex tests on several subjects with different parameters of external disturbance. 3. development of rehabilitation strategy for amputees to provide prosthetic hand operation. compliance with ethical standards conflict of interest the authors declare that there is no conflict of interest. ethical standard the participant signed an informed consent before the experiment. the study was conducted in accordance with the declaration of helsinki. the protocol was approved by the v.a. trapeznikov ics ras institutional review board. acknowledgements the reported study was funded by rfbr, project number 19-38-90293. references 1. prakash, a., sharma, s. & sharma, n. (2019) a compact-sized surface emg sensor for myoelectric hand prosthesis. biomed. eng. lett., 9, 467479. https://doi.org/10.1007/s13534-019-00130-y 2. reiter,r. (1948) eine neue elecktrokunsthand. grenzgebiete der medizin, 4–183. 3. englehart, k., hudgins, b. & parker, p.a. (2001) a wavelet-based continuos classification scheme for multifunction myoelectric control. ieee trans biomed eng vol. 48, 302–311. 4. tenore, f.v.g., ramos, a., fahmy, a., acharya, s., etienne-cummings, r. & thakor, n.v. (2009) decoding of individuated finger movements using surface electromyography. ieee trans biomed eng, vol. 56, no. 5, 1427–1434. 5. kyranou, k., vijayakumar, s. & erden, m.s. (2018) causes of performance degradation in non-invasive electromyographic pattern recognition in upper limb prostheses. front neurorobot, 12–58. 6. scott, r.n. & parker, p.a. (1988) myoelectric prostheses: state of the art. j med eng technol, vol. 12, 143–151. copyright © 2021 assa. adv syst sci appl (2021) case study: influence of muscle fatigue and perspiration 69 7. vujaklija, i., roche, a.d., hasenoehrl, t., sturma, a., amsuess, s., farina, d. & et al. (2017) translating research on myoelectric control into clinics are the performance assessment methods adequate. front. neurorobot, vol. 10, 11–17. 8. park, k.h., suk, h.i. & lee, s.w. (2016) position-independent decoding of movement intention for proportional myoelectric interfaces. ieee trans neural syst rehabil eng, vol. 24, no. 9, 928–939. 9. clancy, e.a. & farry, k.a. (2000) adaptive whitening of the electromyogram to improve amplitude estimation. ieee trans biomed eng, vol. 47, no. 6, 709–719. 10. cifrek, m., medved, v., tonkovic, s. & ostojic, s. (2009) surface emg based muscle fatigue evaluation in biomechanics. clin biomech (bristol, avon), vol. 24, 327–340. doi: 10.1016/j.clinbiomech.2009.01.010. epub 2009 mar 13. 11. venugopal, g. & ramakrishnan, s. (2014) analysis of progressive changes associated with muscle fatigue in dynamic contraction of biceps brachii muscle using surface emg signals and bispectrum features. biomed. eng. lett., 4, 269-276. https://doi.org/10.1007/s13534-014-0135-1 12. boschmann, a. & platzner, m. (2014) towards robust hd emg pattern recognition: reducing electrode displacement effect using structural similarity. annu int conf ieee eng med biol soc., vol. 2014, 4547–4550. doi: 10.1109/embc.2014.6944635. pmid: 25571003. 13. unanyan, n. & belov, a. (2019) signal-based approach to emg-sensor fault detection in upper limb prosthetics, proceedings of the 20th international carpathian control conference (iccc 2019, krakow-wieliczka, poland), 1–6. 14. raj, r., ramakrishna, r. & sivanandan, k.s. (2016) a real time surface electromyography signal driven prosthetic hand model using pid controlled dc motor. biomed. eng. lett., 6, 276-286. https://doi.org/10.1007/s13534-016-0240-4 15. farina, d., cescon, c. & merletti, r. (2002) influence of anatomical, physical, and detection-system parameters on surface emg. biol cybern, 86, 445-456. 16. enoka, r.m.(2019) physiological validation of the decomposition of surface emg signals, journal of electromyography and kinesiology, vol. 46, 70–83. 17. kenny, g.p. & jay, o. (2007) evidence of a greater onset threshold for sweating in females following intense exercise. eur j appl physiol, vol. 101, no. 4, 487-493. 18. de luca, c.j. (1997) the use of surface electromyography in biomechanics. j appl biomech volume 13, 135-136. 19. ray, g.c. & guha, s.k. (1983) equivalent electrical representation of the sweat layer and gain compensation of the emg amplifier, ieee trans biomed eng, vol. 30, no. 2, 130-132. 20. taniguchi, y., sugenoya, j., nishimura, n., iwase, s., matsumoto, t., shimizu, y. & et al. (2011) contribution of central versus sweat gland mechanisms to the seasonal change of sweating function in young sedentary males and females. int j biometeorol vol. 55, no. 2, 203-212. 21. song, j.h., jung, j.w. & bien, z. (2006) robust emg pattern recognition to muscular fatigue effect for human-machine interaction. in mexican international conference on artificial intelligence, 1190-1199. 22. enoka, r.m. & duchateau, j., muscle fatigue: what, why and how it influences muscle function, j physiol, vol. 586, no. 1, 11–23. 23. bartuzi, p. & roman-liu, d. (2014) assessment of muscle load and fatigue with the usage of frequency and time-frequency analysis of the emg signal. acta bioeng biomech, vol. 16, no. 2, 31–39. 24. camata, t.v., dantas, j.l., abrao, t., brunetto, m.a., moraes, a.c. & altimari, l.r. (2010) fourier and wavelet spectral analysis of emg signals in supramaximal constant load dynamic exercise. annu int conf ieee eng med biol soc, vol. 2010, 1364–1367. 25. masuda, t., miyano, h. & sadoyama, t. (1982) the measurement of muscle fiber conduction velocity using a gradient threshold zero-crossing method. ieee trans copyright © 2021 assa. adv syst sci appl (2021) 70 n. unanyan, a. belov biomed eng, vol. 29, no. 10, 673–678. 26. al-mulla, m.r. sepulveda, f., colley, m. & kattan, a. (2009) classification of localized muscle fatigue with genetic programming on semg during isometric contraction. annu int conf ieee eng med biol soc, vol. 2009, 2633–2638. 27. unanyan, n. & belov, a. (2020) a real-time fail-safe algorithm for decoding of myoelectric signals to control a prosthetic arm. proceedings of the 21th international carpathian control conference iccc, 1–6. copyright © 2021 assa. adv syst sci appl (2021) introduction preliminaries hardware description software tools methods main result muscle fatigue perspiration malfunction correction conclusion  2017 г adv syst sci appl 2020; 03; 1-23 published online at https://ijassa.ipu.ru. mathematical modeling of antitumor viral vaccine therapy: from the experiment to the clinic nina a. babushkina1*, ekaterina a. kuzina1 1) v.a. trapeznikov institute of control sciences of russian academy of sciences, 117997, profsoyuznaya street, 65, moscow, russia e-mail: babushkina_na@mail.ru, kate_k93@mail.ru abstract: the paper presents the model developed to identify efficient strategies of antitumor viral vaccine introduction. these strategies are able to produce complete suppression of the tumor growth. the model was developed in matlab-simulink. three efficient strategies of viral vaccine introduction were produced. it was found that the choice of the strategy depends on the tumor size at the start of the treatment, and the range of the tumor sizes for each of the strategies was identified. for the small tumors, elimination of the tumor can be achieved through single-shot vaccine administration in dosages that lead to the death of tumor cells caused directly by the virus. for the big tumors that are within the threshold size, elimination of the tumor can be achieved through repeated vaccine administrations with stepwise reduction of time periods between them. for the tumors of any size, the strategy of repeated administration of the virusbased vaccine that allows stabilizing the tumor size as per the start of the treatment was defined. keywords: mathematical model, experimental oncology, tumor cells, kinetic growth curves, vaccine therapy, virus, immune response, antibodies, allometric proportions 1. introduction effective treatment of cancer remains one of the key challenges of contemporary healthcare. addressing this challenge requires extensive research based on experimental studies that are further translated into clinic. one of the promising directions for combating oncological diseases is the development of different treatment methods within the domain of immunotherapy. cancer immunotherapy research by j. allison and t. honjo was awarded 2018 nobel prize in medicine. significant progress in cancer treatment was produced by breakthroughs in molecular biology and immunology that made it possible to understand the reasons behind tumor degeneration and the development of tumor process [1-3]. immune system can resist the emergence and development of the tumor process, and the aim of immunotherapy is to stimulate the immune system to fight against the tumor cells. there are two ways to stimulate the immune system: specific and non-specific. specific antitumor vaccines are based on dendritic cells that carry the information on the antigens specific for each type of tumor [4-8]. non-specific virus-based antitumor vaccines can induce the immune response against the tumor by producing new protein formations on the surface of the tumor cells [913]. such vaccines are based on the viruses that are not dangerous for the humans but are able to identify and destroy the tumor cells [14-17]. the virus is able to induce the process of tumor cells’ death that has two stages. the first stage is produced by the virus itself. settling down on the tumor cell, the virus penetrates * corresponding author: babushkina_na@mail.ru mailto:babushkina_na@mail.ru mailto:kate_k93@mail.ru n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 2 inside, multiplies itself and destroys the cell. meanwhile the immune system reacts at the virus intrusion and destroys it together with the tumor cell. the second stage of the tumor cells’ death evolves later, due to the immune system’s reaction to the population of infected tumor cells. when settling down on the surface of the cells, the virus leaves specific protein formations on the membrane surface. this produces a population of infected tumor cells that are detected by immune system as foreign. in response, immune system produces antibodies that destroy the infected tumor cells [17-19]. thus the virus acts as a specific marker of the tumor cells and helps to overcome the irresponsiveness of the immune system against uninfected tumor cells. one of the key advantages of immunotherapy is the absence of toxic effects associated with chemotherapy and radiation therapy. since there is no need for additional treatment to mitigate the negative side effects, this makes immunotherapy a more harmless and costeffective method as compared to other approaches to cancer treatment. low toxicity makes it possible to expand the range of dosages – however, this requires additional experimental studies and therefore implies higher research costs. computational experiments based on the mathematical model of vaccine therapy make it possible to define the optimal dosages, assess the effectiveness of immune response depending on the vaccine dosage, and describe the dynamics of the tumor process for different treatment strategies. data obtained through computational experiments may be used to define optimal treatment strategies that produce complete tumor regression. given the high costs of experimental and clinical studies, computational experiments based on mathematical models can produce a lot of useful information based on the limited amount of experimental data. modelling makes it possible to expand the study of the range of applicable dosages and treatment strategies when exploring the new methods of antitumor therapy, thus reducing time and labor costs and the number of animals required for conducting an experiment in vivo. moreover, research in molecular biology, cell biology, biophysics and immunology has produced a large amount of data that also require processing and analysis with the use of specialized computer technologies based on mathematical modelling [8-11]. however, despite its advantages, modelling is still not widely used in experimental and clinical oncology. this paper aims to demonstrate that computational experiments based on mathematical modelling may produce meaningful results for defining cancer treatment strategies, complementing in vivo experimental approaches and enabling more efficient translation of experimental results to the clinic, building on the study published in mathematical biology & bioinformatics earlier in 2019 [19]. 2. problem statement this study aims to develop an algorithm for finding strategies for the use of antitumor viral vaccine to completely suppress tumor growth. to achieve this aim, a software package was developed in the matlab-simulink system, which includes several mathematical models: mathematical model of vaccine therapy, describing the mechanism of tumor cell death as a result of the immune response to the injection of the virus [10,18,19], mathematical model of infectious disease by g. i. marchuk, describing the dynamics of the formation of antibodies against the virus by the immune system [20-22], mathematical model by h. f. skipper, describing how the proportion of the fast proliferating tumor cells decreases in the growing tumor [10,23], mathematical model of antitumor therapy with discontinuous trajectories, evaluating the effectiveness of therapeutic effects on the experimental trajectories of tumor growth after the injection of a viral vaccine [10,18,19]. the values of the parameters of the complex mathematical models are given in the table 1 [19, p. 44]. the complexity and nonlinearity of differential equations of mathematical models does mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 3 not allow to obtain the solution analytically without the use of computational software. on the basis of the developed software package, a computational experiment was conducted, which allowed to investigate the effectiveness of various strategies for the injection of viral vaccines depending on the dose and size of the tumor at the time of its injection. 3. mathematical model of vaccine therapy with virus val the mathematical model of vaccine therapy describes the mechanism of tumor cells’ death under the influence of two factors – the virus itself and antibodies against infected tumor cells, on the basis of experimental data on the growth of erlich tumor in experimental animals after a single administration of the vaccine with venezuelan equine encephalitis (vee) virus [17]. the experiments were carried out on 2-month-old female mice of the line balb/c, с57в1/6 and dba/2. in total, 690 animals were used in the experiments [17]. the process of tumor cell growth without treatment (control) is described by a simple differential equation [19, p. 38]: ( ) ( ) ( ) dn t t n t dt =   , with 0 0( )n t n= , (1) where n(t) is the number of tumor cells, t denotes time,  (t) is the parameter characterizing the growth rate of tumor cells, n0 is the initial number of tumor cells at time t = 0. the type of function describing tumor growth without treatment was determined by experimental curves of erlich adenocarcinoma growth by regression analysis in matlab (fig. 1). the experimental curve of tumor growth without treatment is most accurately described by the gompertz function, which is the solution of the differential equation (1) for  (t) = n n exp(– n t). fig. 1. approximation of experimental data on the growth of ehrlich adenocarcinoma in control by the gompertz function the gompertz function is: 0( ) exp( exp( )) exp( (1 exp( ))n n n n nn t n t n t= −  − =  − − , (2) where 0 exp( )nn n =  is the maximum tumor size for t →  . calculated values of the parameters of the gompertz function and the sum of squared deviations (sко = 0.21) are shown in table 1 [19, p. 44]. analysis of experimental data on tumor growth after the introduction of a viral vaccine (fig. 2, lower curve) allows us to distinguish two periods of intensive death of tumor cells. n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 4 the first period lasts from 6 to 8 days, and the second from 13 to 16 days. we can assume that the first stage is associated with the reaction of the immune system to the virus and the formation of antibodies specific to this virus. and the second stage is associated with the appearance of dead tumor cells infected with a virus, and the formation of antibodies specific to these tumor cells [15-17]. fig. 2. experimental data on the growth of ehrlich adenocarcinoma without treatment and after a single injection of the vaccine [17-19] as a result of the interaction of the virus with tumor cells, three populations of tumor cells emerge (fig. 3). fig. 3. interaction of the virus with tumor cells and immune system the population of tumor cells carrying the virus is part of the tumor cells on which the virus is absorbed. penetrating inside the cell and multiplying there, the virus makes the tumor cell die. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 5 the population of infected tumor cells has specific protein formations on the cell membrane which are perceived as alien by the immune system. thus this population dies under the action of the antibodies produced by the immune system. the population of uninfected tumor cells remains alive and is capable of multiplying and resuming tumor growth. a mathematical model of vaccine therapy describing the dynamics of two-stage tumor cell death is presented in the form of a system of differential equations (1–10) [18–20]. the dynamics of tumor growth before the introduction of the vaccine is described by a differential equation of the form:  0, , ( ) ( ) ( ), vt dn t t n t t dt =   with 0 0( )n n t= , (3) where ( )n t is the population of tumor cells before administration of the vaccine, 1v cvt z=  + is the beginning of the immune response against viruses. 1 is the moment of injection of the vaccine, cvz is the period of the delay of immune response against the virus. the dynamics of growth and death of infected tumor cells after the introduction of the virus is described by the differential equation of the form: ( , , ( ) [ ( ) ( ) ( )] ( ), v v av v v v n dn t t k v t k a t n t t t t dt =  − −   with 0 ( )v vn n t= (4) where ( )vn t is the population of infected tumor cells following injection of the vaccine, ( )v t is the number of viruses, ( )va t is the number of antibodies against the virus, vk is the coefficient of the rate of reproduction of the virus in a tumor cell, avk is the coefficient of the rate of death of viruses as a result of interaction with antibodies ( )va t . the dynamics of growth and death of infected tumor cells under the action of antibodies at the second stage is described by the differential equation of the form:  , , ( ) [ ( ) ( )] ( ), v an n v n l dn t t k a t n t t t t dt =  −   with 0 ( )v nn n t= , (5) where 1n cnt z=  + is the moment of the immune response against infected tumor cells, cnz is the delay time of the immune response against infected tumor cells, ( )na t is the number of antibodies against infected tumor cells, ank is the dimensional coefficient. the death of tumor cells at each of the two stages occurs as a result of the development of the body's immune response to the introduction of a viral vaccine. the first stage is associated with the formation of antibodies against the virus ( )va t , and the second stage is associated with the formation of antibodies against infected tumor cells ( )na t . graphs of antibody formation dynamics ( )va t against the virus and against infected tumor cells ( )na t are shown in fig. 4 [19, p. 39]. n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 6 fig. 4. graph of the dynamics of the number of antibodies against the virus ( )va t (the first stage of the immune response) and against infected tumor cells ( )na t (the second stage of the immune response)) for the experimental dose 0 0.015v = on τ1 = 1day the mechanism of their formation is calculated on the basis of the model of infectious disease by g. i. marchuk [20-22]. the parameters of this model are adapted for experimental curves of tumor growth after the introduction of viral vaccines. the dynamics of the number of viruses according to the mathematical model of an infectious disease [20, 22] is described by the equation of the form: ( ) ( ) ( ) ( ),v v v dv t v t a t v t dt =  − (6) where 0 1( )v v=  is the initial dose of virus vaccine, 1 is the moment of the first introduction of a viral vaccine, v is the rate of reproduction of the virus inside the cell, v is the rate of death of viruses in their interaction with antibodies ( )va t [19, p. 39]. the initial condition for the solution of equation (4) is taken as a control parameter characterizing the introduced dose of the viral vaccine 0 1( )v v=  . the first stage of the body's immune response to the introduction of the virus is determined by the number of antibodies ( )va t , which is calculated from the following equations: ( ) ( ) ( ) ( ) ( ),v a v v av v v v da t c t t a t v t a t dt =  − − − (7) where a is the rate of formation of antibodies from one plasma cell, av is the rate of loss of antibodies due to the interaction with viruses ( )va t , v is the rate of reduction of the number of antibodies due to natural destruction, 1v cvt z=  + is the moment of the beginning of immune response against viruses [19, p. 40]. due to the fact that the time of virus reproduction in the experimental tumor was not recorded in the available experimental data [15-17], when constructing the model, it was assumed that the period of virus reproduction inside the tumor cell, leading to its death, can be considered with sufficient accuracy equal to the time of delay of the immune response against the virus cvz . then in equations (7) and (8) the time of introduction of the virus was taken into account as a parameter tv = 1 +zcv, which records the moment of the beginning of the immune response against viruses. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 7 number of plasma cells ( )vc t is determined from the equation: ( ) ( ) ( ) ( )v vc v v cv v vn dc t v t t a t t c t c dt      =  − − − − , with ,( )v v vnc t c= (8) where с is the rate of formation of plasma cells, cv is the constant coefficient, cvz is the period of delay in the development of antibodies against the virus (first stage of immune response). the second term of this equation reflects the maintenance of the initial number of plasma cells in the norm vnc [19, p. 40]. the second stage of the immune response of the body to the formed infected tumor cells was determined by the number of antibodies, which was calculated from the following equations: ( ) ( ) ( ) ( ) ( ),n an n n an n v nn n da t c t t a t n t a t dt =  − − − (9) where an is the rate of formation of antibodies, an is the coefficient that describes the rate of reduction of the number of antibodies ( )na t due to their interaction with infected tumor cells ( )vn t , nn is the rate of reduction of the number of antibodies due to natural destruction [19, p. 40]. number of plasma cells ( )nc t is determined from the equation: , ( ) ( ) ( ) ( )n cn v n n n cn n nn dc t n t t a t t c t c dt      =  − − − − with ,( )n n nnc t c= (10) where cn is the rate of formation of plasma cells, cn is the constant coefficient, сnz is the period of delay in the development of antibodies against the infected tumor cells (second stage of immune response), nnc is the initial number of plasma cells in the norm [19, p. 40]. to adequately describe the effectiveness of viral vaccines, it is necessary to take into account that one of the factors of high selectivity of viruses in relation to tumor cells is the high rate of their division in comparison with normal body tissues [4-7]. the measured tumor volume contains fractions of rapidly and slowly proliferating tumor cells, which is described in detail in the mathematical model of tumor growth by skipper [23]. as the size of the tumor increases, the fraction of rapidly dividing tumor cells decreases, while the fraction of slowly dividing cells and temporarily non-dividing cells increases. to describe how the tumor size is related to the efficiency of the vaccine, ( )np t function is included in the model. this function describes the dynamics of the decline in the share of rapidly proliferating cells with increasing size of the tumor [19]: 2 2 21 ( ) 1 ( ) , 1 p p p p t p t arctg k t     = −    −    (11) where p and pk are constant parameters, t denotes time (in days), *1/p t = , where *t is the moment when numbers of the rapidly and slowly proliferating cells are equal [19, p. 41]. then the number of infected cells in the measured tumor volume is calculated as ( ) ( ) ( )v n n nn t n t p t= , where 1n cnt z=  + is the moment when the immune response against infected tumor cells begins. as a result of the interaction of the virus with tumor cells, three populations of tumor n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 8 cells develop in the tumor volume ( )vn t (fig. 5) [19, p. 41]. the structure of the equations of the model is such that the existence, uniqueness and nonnegativity of solutions in this work under nonnegative initial conditions are satisfied, which was proved and stated in the work of n. v. pertsev [21, p. 156-157]. fig. 5. calculation curves of the dynamics of infected cells ( )vn t (dotted line) and the kinetic trajectory of tumor growth ( )n t after dose injection 0 0.015v = at τ1 = 1 day, v and n are duration of the delay tumor growth after each stage of tumor cells’ death. experimental data is denoted by + figure 5 shows two calculated curves. the dotted line describes the dynamics of growth and death of infected tumor cells in accordance with the model of vaccine therapy (1) – (11). the solid line shows the dynamics of the number of surviving tumor cells after the first stage and the second stage of immune response in accordance with the mathematical model of anticancer therapy with discontinuous trajectories [19]. 4. a mathematical model of anticancer therapy with discontinuous trajectories to predict the dynamics of tumor growth when constructing a mathematical model of antitumor therapy with discontinuous trajectories, several assumptions were made to describe the death and subsequent growth of the experimental tumor [10, 18]. 1. tumor growth without the introduction of a viral vaccine (control) and after its introduction is described by the gompertz function with the preservation of the values of the parameters of the function. 2. the death of tumor cells occurs instantly, causing a sudden decrease in the size of the tumor at the beginning of the immune response at each stage of cell death. 3. the trajectory of tumor growth after the death of tumor cells is shifted in time for the duration of the delay in tumor growth 0( )v after each stage of their death. 4. duration of tumor growth delay 0( )v is the time interval between the start of immune response (i.e. when the number of tumor cells starts to decline affected first by the virus at vt and then by the antibodies at nt ) and the moment when the tumor, resuming its growth, reaches the same size as at vt and nt respectively [10, 18]. the mathematical model of anticancer therapy with discontinuous trajectories is used to construct dynamic trajectories of tumor growth after two-stage death of tumor cells after the mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 9 introduction of a viral vaccine (fig. 5 – ( )n t ) [19, p. 41]. the dynamics of tumor growth before the injection of the vaccine is described by a differential equation (12), which is similar to the equation of tumor growth in the control (1):  0, , ( ) ( ) ( ), vt dn t t n t t dt =   with 0 0( )n n t= , (12) where ( )n t is the size of the tumor before vaccine injection,  (t) = n n exp(– n t) is tumor growth rate in control, 1v cvt z=  + is the moment when tumor cells start to die in response to the virus (1st stage of immune response), 1 is the moment of vaccine injection, cvz is the period of delay of the immune response against the virus. the dynamics of tumor growth after the first stage of tumor death is described by a differential equation of the form:  , , ( ) ( ) ( ), v v v nt t t dn t t n t t dt = − −  with ) )( ( ) ( ) (r v v v v vtn n t s t t t=  − (13) where ( )vt t − is pulsed dirac function, describing the instantaneous death of tumor cells in the first stage of the immune response, ( )v vs t denotes the proportion of dying tumor cells infected with the virus at the first stage of the immune response, which is determined from the equation: ( ) ( ) ( ) v v v v v v n t s t n t  = (14) where ( )v vn t denotes the population of infected tumor cells at the start of the immune response, ( )v vn t is the number of dying tumor cells in the first stage of the immune response, which is calculated as the difference between the maximum and minimum number of infected cells in the period from the beginning to the end of the 1st stage of immune response according to the equation: 1 2( ) ( ) ( )v v v vv vn t n t n t = − (15) where 1 vt and 2 vt are the moments when the number of infected cells during the 1st stage of the immune response against the virus reaches maximal and minimal level (fig. 5, dotted line), ( )r vn t is the number of remaining tumor cells that continue to produce tumor growth, which is calculated from the equation: ) ( )( ) (r v v v vn tn t n t −= (16) the trajectory of tumor growth after the first stage of tumor cell death is described by the gompertz equations with a time shift for the duration of tumor growth delay (fig. 5): 0 0( ) exp( (1 exp( ( ( ))))),v n n v vn t t n t v− =  − − − (17) where 0( , )v vv t is the delay in tumor growth after the first stage of the immune response. the dynamics of tumor growth after the second stage of tumor cells’ death is described n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 10 by the differential equation of the form:  , , ( ) ( ) ( ),n n n lt dn t t t n t t t dt = − −  with ) )( ( ) ( ) (r n v n n ntn n t s t t t=  − (18) where 1n cnt z=  + is the moment when the immune response against virus-infected tumor cells begins, cnz is the period of delay of immune response against virus-infected tumor cells, ( )nt t − is the pulsed dirac function, describing the instantaneous death of tumor cells in the second stage of the immune response, ( )n ns t is the proportion of dying tumor cells infected with the virus at the second stage of the immune response, which is determined from the equation: ( ) ( ) , ( ) v n n n n n t s t n t  = (19) ( )v nn t denotes the number of dying tumor cells in the second stage of the immune response. it is calculated as the difference between the maximum and minimum number of infected cells in the period from the beginning to the end of the 2nd stage of immune response according to the equation: 1 2( ) ( ) ( ),n n v vv nn t n t n t = − (20) where 1 nt and 2 nt are the moments when the number of infected cells during the 2nd stage of immune response against the virus reaches maximal and minimal level (fig. 5, dotted line), ( )r nn t is the number of remaining tumor cells that continue to produce tumor growth, which is calculated from the equation: ) ( )( ) (r n n v nn tn t n t −= (21) then the trajectory of tumor growth after the second stage of tumor cell death is described by the gompertz equation with a time shift for the duration of tumor growth delay (fig. 5): 0 0( ) exp( (1 exp( ( ( )))))n n n n nn t t n t v− =  − − − (22) where 0( )n v is the delay in tumor growth after cells’ death in the second stage of the immune response. the duration of the delay of tumor growth 0( )v v and 0( )n v was determined as the time interval between the start of immune response (i.e. when the number of tumor cells starts to decline affected first by the virus at vt and then by the antibodies at nt ) and the moment when the tumor, resuming its growth, reaches the same size as at vt and nt respectively. therefore: 0( ) ( ( ))v v vn t n t v= +  , (23) 0( ) ( ( ))n n nn t n t v= +  . (24) the values of the parameters of the vaccine therapy model and the model of anticancer therapy with discontinuous trajectories are given in appendix 1. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 11 5. finding effective strategies of viral vaccine introduction evaluation of the effectiveness of various strategies of tumor growth after the injection of the viral vaccine was carried out on the basis of a mathematical model of anticancer therapy with discontinuous trajectories according to the equations (12) (24). the aim of efficient strategy is to reduce the number of tumor cells below the lethality threshold. in this case tumor will not be able to resume its growth. this threshold is denoted let 0n . if the number of surviving cells is below the threshold after the first stage of immune response ( 0( )r let vn t n ) or after the second stage of the immune response ( 0( )r let nn t n ), this is interpreted in the model as the complete destruction of tumor cells. in this case, the life expectancy of treated animals will be equal to the average life expectancy of experimental animals without tumors tl = 3 years. 5.1. strategy of stabilization of tumor growth with multiple injection of the viral vaccine modeling of multiple injections of a viral vaccine within the framework of the constructed model allowed us to determine the algorithm for finding a strategy for stabilizing tumor growth with multiple injections of a constant dose of a viral vaccine v0 stab (τ1). results indicate that the strategy of stabilization is possible for a narrow range of dosage. this dosage should be able to produce an equal number of antibodies at each of the two stages of the immune response stab stab 0 1 0 1( , ) ( , )v na v a v =  (fig. 6, a,b) and (fig. 7, a,b). а b fig. 6. strategy to stabilize the tumor size with the injection of a viral vaccine on τ1 = 3 days with an interval between injections ∆t = 18.4 days: а shows the trajectories of growth and death of tumor cells; b shows the dynamics of antibody formation 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of immune response n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 12 а b fig. 7. strategy to stabilize the tumor size with the injection of a viral vaccine on τ1 = 20 days with an interval between injections ∆t = 25.7 days: а shows the trajectories of growth and death of tumor cells; b shows the dynamics of antibody formation 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of immune response the strategy of tumor growth stabilization is performed in the range of doses from stab 0v =0.04 to stab 0v =0.06 (fig. 8,b) at a constant value of the duration of the interval between administration of the vaccine. duration of intervals between repeated introductions of the vaccine ∆t increases with the size of the tumor (fig. 8,a). а b fig. 8. the dependence of the values of the parameters under the strategy of stabilization of tumor growth on the size of the tumor at the beginning of treatment: а shows the intervals between vaccine introductions stab 0 1( , )t v  , b shows the value of the introduced dose stab 0 1( )v  this strategy requires sufficiently large dosage which will be able to produce equally big number of antibodies at the 1st and the 2nd stages of immune response. thus, this strategy allows to transfer the course of the disease into a chronic state by restraining the growth of the tumor. within the framework of the constructed model, it is shown that it is possible to restrain the growth of a tumor of any size for an unlimited time. however, if vaccine is not introduced any longer, the tumor growth could resume. the strategy of tumor growth stabilization is similar to the effect of inoculation from own tumor cells in the period between two successive injections of the vaccine. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 13 5.2. strategy of tumor growth suppression with multiple injection of the viral vaccine to solve the problem of reducing the number of tumor cells below the lethality threshold with repeated injection of the vaccine for all tumor sizes, agorhythm was developed, based on the strategy of stabilizing tumor growth, in which the duration of intervals between repeated injections of the viral vaccine was reduced step by step. figures 9,а and 10,а show the trajectories of tumor growth when the vaccine is injected 3 and 4 times, and the time intervals between consecutive injections are reduced. this allows to drive the number of tumor cells below the lethality threshold. a b fig. 9. the strategy for multiple injection of the vaccine by reducing the interval between the doses 0v = 0,056 at the beginning of the treatment on τ1 = 10 day with the initial interval between injections ∆tstab = 20.1 days: а shows the trajectories of growth and death of tumor cells; b shows the dynamics of formation of antibodies 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of the immune response (solid curve denotes av(t), dashed curve denotes an(t)) а b fig. 10. the strategy for multiple injection of the vaccine by reducing the interval between the doses 0v = 0,06 at the beginning of the treatment on τ1 = 20 day with the initial interval between injections ∆tstab = 25.7 days: а shows the trajectories of growth and death of tumor cells; b shows the dynamics of formation of antibodies 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of the immune response (solid curve denotes av(t), dashed curve denotes an(t)) the dose of the viral vaccine was determined depending on the size of the tumor to implement the strategy of stabilizing the tumor size at the beginning of treatment v0 stab( τ1) (fig.8 b). the interval between the first and second introduction of the vaccine ∆t1 stab (τ1) was also determined depending on the size of the tumor at the beginning of treatment based n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 14 on the condition of stabilization of the tumor size (fig. 8,а). the moment of the second injection of the vaccine was defined as τ2 = τ1 + ∆t1 stab. beginning from the third moment of the vaccine injection τ3 and further τi, the moment of the subsequent injection of the vaccine was defined as τi = τi-1 + ∆ti-1. the duration of the interval between injections was reduced by the step value and was calculated as ∆ti-1 = ∆ti-2 – step. in fact, the step value is the control parameter for the implementation of the strategy of gradual reduction of tumor size with each subsequent injections of a viral vaccine. the number of repeated injections of the vaccine to achieve regression of tumor growth depends on the step value. in the framework of the constructed model, this occurs when the number of tumor cells decreases below the lethality threshold (fig. 9,a-10,а). this strategy allows to achieve a reduction in the number of tumor cells below the lethality threshold only if the duration of tumor growth does not exceed 20 days (fig. 10,a). for large tumors it is necessary to carry out surgery, or use a strategy to stabilize the growth of the tumor (fig. 11,a). а b fig. 11. the strategy for multiple injection of the vaccine by reducing the interval between the doses 0v = 0,06 at the beginning of the treatment on τ1 = 25 day with the initial interval between injections ∆tstab = 28.9 days: а shows the trajectories of growth and death of tumor cells; b shows the dynamics of formation of antibodies 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of the immune response (solid curve denotes av(t), dashed curve denotes an(t)) 5.3. strategy of single injection of the viral vaccines testing the effectiveness of various doses of viral vaccine in a single injection at different points in tumor growth was carried out on the developed program module. it was determined that even with a single injection of a viral vaccine in certain doses, it is possible to obtain a reduction in the number of tumor cells below the v0let threshold, i.e. achieve complete destruction (fig.12a, 13a). at the same time, the number of antibodies produced at the first stage of the immune response exceeds their number at the second stage (fig. 12b, 13b). mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 15 а b fig. 12. the strategy of a single lethal dose injection on let 0v = 0.048 on τ1 = 3 day; а shows the trajectories of growth and death of tumor cells; b shows the dynamics of formation of antibodies 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of the immune response [19, p. 48, fig. 9] а b fig. 13. the strategy of a single lethal dose injection on let 0v = 0.075 on τ1 = 8 day; а shows the trajectories of growth and death of tumor cells; b shows the dynamics of formation of antibodies 0 1( , )va v  at the 1st and 0 1( , )na v  at the 2nd stage of the immune response thus, the death of tumor cells with a single injection of lethal doses occurs immediately in the first stage of the immune response as a result of the formation of large amounts of antibodies against the virus 0( )va v and a slight formation of antibodies against infected tumor cells 0( )na v . as a result of the analysis of the obtained results, a dependence was constructed, indicating that the magnitude of lethal doses leading to the destruction of all tumor cells at the first stage of the immune response depends on the size of the tumor at the time of its introduction (fig. 14). however, to achieve the complete destruction of tumor cells is possible only for small tumors, the growth of which does not exceed 8 days. n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 16 fig. 14. the graph of lethal doses from the beginning of treatment ( )let 0 1v  with a single injection of a viral vaccine [19, p. 49, fig.11] 6. predicting the duration of the periods between the repeated injections of virus vaccine for the clinic to predict the duration of the periods between repeated injections of the viral vaccine for the clinic according to the results obtained in the experiment on mice, it is proposed to use the well-known in biodynamics allometric ratio, which connects the rate of metabolic processes in the body with the body weight in mammals [24-26]: bm a s=  , (25) where m is the body weight of the mammal, s is the rate of metabolic processes in mammals, a and b are constant parameters. allometric relations connect the body weight of animals and humans with a variety of other biological parameters that reflect the temporal characteristics of the processes of vital activity of the organism, including the duration of its life [24-26]. then the ratio (25) can connect the body weight of the animal with any other time parameter. in this case, such a time parameter is considered the duration of the periods between successive injections of the viral vaccine in the implementation of the strategy of stabilizing the tumor size ∆t : ( )bm a t=   . (26) for two species of mammals, which are human and experimental mouse, we can write the allometric ratio of the form: ( )human human bm a t=   and ( )mouse mouse bm a t=   , (27) where humant and mouset denote the duration of the periods between the injections of viral vaccines for human being and for mouse. to define the duration of the period between injections for the human being we can use the ratio: ( ) human human b mouse mouse m t m t  =  based on the fact that the average body weight values are known for a human being and a mouse, and the interval between successive injections of the vaccine is determined, the mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 17 duration of the period between injections of the vaccine for a person can be calculated from the equation: 1/( ) human human mouse b mouse m t t m  =  (28) the value of the parameter b was calculated as the slope of the straight line formed by the logarithmation of the allometric ratio, which relates the average body weight of different mammals and humans with the average life expectancy lifet (fig. 15) [10]: ln ln ln lnlife lifem b t a b t c=  + =  + (29) as a result of the regression analysis, the calculated value of the parameter b was obtained as b = 2,4981, which characterizes the slope of the regression line. fig. 15. dependence of the life expectancy of animals on their body weight for ○ mouse, □ rats, + guinea pigs, х rabbits, ◊ dogs, and * humans [10] taking in equation (28) the mass of the mouse equal to mousem = 0,025 kg and the weight of a human being is equal to humanm = 70 kg, with b = 2,4981 = 2,5 and 1/b = 0,39 we can calculate the time conversion factor from mouse to human. 0.470 ( ) 23.9 0.025 k = = then the duration of the period between injections of the vaccine for a human being can be calculated from the expression: 1/( ) human human mouse b mouse mouse m t t t k m  =  =   (30) the results of the evaluation of the duration of the periods between repeated injections of viral vaccines in the implementation of the strategy of tumor size stabilization for humans and mouse are shown in the table 1. table 1. intervals between repeated injections of viral vaccines for humans and mouse obtained on the basis of allometric ratios dose of viral vaccine 0v and when the vaccine was first interval between injections to stabilize tumor size in the experiment on mouse 0 1( , )mouset v  interval between injections to stabilize the size of the tumor in the clinic for humans 0 1( , )humant v  n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 18 introduced 1 0v = 0.04 on 1 =1 day 0 1( , )mouset v  = 17.2 day 0 1( , )humant v  =1,14 year 0v = 0.047 on 1 =3 days 0 1( , )mouset v  = 18.4 day 0 1( , )humant v  =1,22 year 0v = 0.05 on 1 =5 days 0 1( , )mouset v  = 19.0 day 0 1( , )humant v  =1,26 year 0v = 0.056 on 1 =10 days 0 1( , )mouset v  = 20.1 day 0 1( , )humant v  =1,33 year 0v = 0.06 on 1 =15 days 0 1( , )mouset v  = 22.8 day 0 1( , )humant v  =1,5 year 0v = 0.06 on 1 =20 days 0 1( , )mouset v  = 25.7 day 0 1( , )humant v  =1,7 year 0v = 0.06 on 1 =25 days 0 1( , )mouset v  = 28.9 day 0 1( , )humant v  =2 year the method of allometric ratios is indirect and cannot guarantee 100% accuracy and reliability of the assessment of the duration of the periods between repeated injections of viral vaccines for their use in the clinic. however, this period of repeated injections of the viral vaccine may be a recommendation for the patient and for the doctor about the timing of the control examination. according to the results of the examination, it is necessary to decide on the condition of stabilization of the tumor size, which should not exceed the size of the tumor at the time of treatment. the results of the medical inspection should clarify the dose and adjust the strategy of repeated injections of the viral vaccine. 7. conclusion the results of the computational experiment conducted within this study indicate that the choice of cancer treatment strategy depends on the tumor size at the start of the treatment. the model also indicates that it is possible to reduce the number of tumor cells beyond the lethality threshold which guarantees that the tumor does not resume its growth. for the small tumors (less than 8 days’ growth as per our model), elimination of tumor cells can be achieved through single-shot virus-based vaccine administration. in this case, the death of tumor cells is caused by their destruction by the virus. however, in these conditions the number of antibodies produced against tumor cells is limited, which can result in recurrence of the tumor growth in future. for bigger tumors (within 20 days’ growth as per our model) elimination of tumor cells is possible if the vaccine is administered repeatedly with stepwise reduction of time periods between administrations. the model developed within this study also makes it possible to define treatment strategy that may restrain the tumor development, i.e. lead to cancer chronification. dosages and time intervals between vaccine administrations for this stabilization strategy were defined in the present study. an advantage of this strategy is that it can be used for the tumors of any size. thus it can be beneficial for the patients who cannot be subjected to surgical treatment, and creates new treatment possibilities for far-advanced cancer. this study also proposes a method of transferring the results of the computational experiment to the treatment of humans, based on allometric relationships between human and animal body weights, enabling clinicians to calculate the appropriate intervals between repeated vaccine administrations. this method makes it possible to plan the return visit to the therapist to make a decision on subsequent vaccine administration. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 19 finally, results of the computational experiment based on mathematical modelling may substantially and meaningfully complement in vivo experiments for defining efficient cancer treatment strategies. development of advanced methods of cancer treatment is highly dependent on the costs of conducting experiments. active implementation of mathematical modelling approach on different stages of experimental research and clinical studies may contribute to the reduction of these costs as well. references 1. baryshnikov, a. yu. (2003). vzaimootnoshenie opukholi i immunnoy sistemy organizma [interaction between tumor and immune system]. prakticheskaya onkologiya, 4(3), 128–130 [in russian]. 2. grinevich, yu. a., khranovskaya, n. n. (2007). vaktsiny na osnove antigenprezentiruyushchikh dendridnykh kletok v immunoterapii bol'nykh so zlokachestvennymi opukholyami [vaccines on the base of antigen presenting dendritic cells for immunotherapy of cancer patients]. prakticheskaya onkologiya, 9(4), 365–370 [in russian]. 3. kose, e., moore, s., ofodile, c., radunskaya, a., ellen, r. (2017). immuno-kinetics of immunotherapy: dosing with dcs. letters in biomathematics, 4(1), 39–58. 4. loktev, v. b., ivankina, t. yu., netesov, s. v., chumakov, p. m. (2012). onkoliticheskie parvovirusy. novye podhody k lecheniju rakovyh zabolevanij [oncolytic parvoviruses. new approaches to the treatment of cancer diseases]. vestnik rossijskoj akademii medicinskih nauk [bulletin of the russian academy of medical sciences], 67(2), 42–47 [in russian]. 5. lezhnin, yu. n., kravchenko, yu. e., frolova, e. i., chumakov, p. m., chumakov, s. p. (2015). onkotoksicheskie belki v protivorakovoj terapii: mehanizmy dejstvija [oncotoxic proteins in anticancer therapy: mechanisms of action]. molekuljarnaja biologija [molecular biology], 49(2), 264–278 [in russian]. 6. rommelaere, j., geletneky, k., angelova, a.l., daeffler, l., dinsart, c., kiprianova, i., schlehofer, j. r., raykov, z. (2010). oncolytic parvoviruses as cancer therapeutics. cytokine & growth factor reviews, 21(2), 185–195. 7. raykov, z., grekova, s., galabov, a. s., balboni, g., koch, u., aprahamian, m., rommelaere, j. (2007). combined oncolytic and vaccination activities of parvovirus h-1 in a metastatic tumor model. oncology reports, 17(6), 1493–1500. 8. de pillis, l., gallegos, a., radunskaya, a. (2013). a model of dendritic cell therapy for melanoma. front. oncol., 3(56), 1-14. 9. kim, r., woods, t., radunskaya, a. (2018). mathematical modeling of tumor immune interactions: a closer look at a pd-l1 inhibitor in cancer immunotherapy. spora: a journal of biomathematics, 4(1), 25–41. 10. babushkina, n., kuzina, e. (2018). analytical study of the antitumor viral vaccine introduction regimens based on mathematical modeling. advances in systems science and applications, 18(1), 59–84. doi: 10.25728/assa.2018.18.1.493 11. kogan, y., halevi–tobias, k., elishmereni, m., vuk-pavlovic, s., agur, z. (2012). reconsidering the paradigm of cancer immunotherapy by computationally aided real-time personalization. cancer res.,72(9), 2218–2227. https://doi.org/10.25728/assa.2018.18.1.493 n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 20 12. lai, x., friedman, a. (2017). combination therapy of cancer vaccine and immune checkpoint inhibitors: a mathematical model. plos one, 12(5): e0178479. doi: https://dx.doi.org/10.1371%2fjournal.pone.0178479 13. enderling, h., chaplain, m. a. j. (2014). mathematical modeling of tumor growth end treatment. current pharmaceutical design, 20(30), 4934–4940. 14. mutsenietse, a. ya. (1972) onkotropizm virusov i problema viroterapii zlokachestvennyh opuholej [oncotropism of viruses and the problem of viral therapy of malignant tumors]. riga: zinatne, [in russian]. 15. gromova, a. yu. (1999) protivoopuholevye svojstva vakcinnogo shtamma virusa venesujel'skogo jencefalomielita i ego onkolizata [antitumor properties of the vaccine strain of the venezuelan encephalomyelitis virus and its oncolyte]. thesis of candidate of science (biology). st. petersburg, [in russian]. 16. urazova, l. n. (2003) jeffektivnost' i mehanizmy protivoopuholevogo dejstvija virusnyh vakcin pri jeksperimental'nom onkogeneze [efficiency and mechanisms of the antitumor action of viral vaccines in experimental oncogenesis]. thesis of candidate of science (biology). st. petersburg, [in russian]. 17. vidyaeva, i. g. (2005) virusnye vakciny i ih onkolizaty v terapii jeksperimental'nyh opuholej [viral vaccines and their oncolytes in the therapy of experimental tumors]. thesis of candidate of science (medicine). tomsk, [in russian]. 18. babushkina, n. a., kuzina, e. a., loos, а. а., belyaeva, е. v. (2018) rezul'taty issledovaniya rezhimov primeneniya protivoopuholevyh virusnyh vakcin na osnove matematicheskogo modelirovaniya [the results of the study of the modes of application of anticancer virus vaccines based on mathematical modeling] // problemy upravlenija [control sciences], 4, 61–70, [in russian]. 19. babushkina, n. a., kuzina, e. a., loos, a. a., belyaeva, e.v. (2019). otsenka effektivnykh strategiy primeneniya protivoopukholevoy vaktsinoterapii na osnove matematicheskogo modelirovaniya [assessment of the efficient strategies for applying atitumor viral vaccine therapy based on mathematical modeling]. matematicheskaya biologiya i bioinformatika [mathematical biology and bioinformatics], 14(1), 34–54. doi: 10.17537/2019.14.34 [in russian]. 20. marchuk, g. i. (1991) matematicheskie modeli v immunologii. vychislitel'nye metody i jeksperimenty [mathematical models in immunology. computational methods and experiments]. moscow: nauka, [in russian]. 21. pertsev, n. v. (2018). global'naya razreshimost' i otsenki resheniy zadachi koshi dlya funktsional'no-differentsial'nykh uravneniy s zapazdyvaniem, ispol'zuemykh v modelyakh zhivykh system [global solvability and estimates for solutions to the cauchy problem for the retarded functional differential equations that are used to the model living systems]. sibirskiy matematicheskiy zhurnal [siberian mathematical journal], 59(1), 143–157 [in russian]. 22. romanyukha, a. a. (2011) matematicheskie modeli v immunologii i jepidemiologii infekcionnyh zabolevanij [mathematical models in immunology and epidemiology of infectious diseases]. moscow: binom. laboratory of knowledge [laboratorija znanij], [in russian]. 23. skipper, h. f. (1971). kinetics of mammary tumor cell-growth and implications for therapy. cancer, 28(6), 1479–1499. 24. monichev, a. yu. (1984) dinamika krovetvorenija. [dynamics of hematopoiesis]. moscow: medizina [medicine], [in russian]. mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 21 25. monichev, a. yu. (1987). a mathematical model of the spatial structure of bone marrow in hemopoietic dynamics. cybernetics and systems analysis, 23(2), 274–280. 26. kuzina, e.a., babushkina, n.a. (2016). programmnaja realizacija metoda prognozirovanija jeffektivnosti vakcinoterapii ot jeksperimenta v kliniku [software implementation of the method for predicting the efficacy of vaccine therapy from experiment to the clinic]. proc. of the 12th int. conf. “physics and radioelectronics in medicine and ecology” phreme-2016. vladimir: vgu, 125–130, [in russian]. n.a. babushkina, e.a. kuzina copyright ©2020 assa. adv. in systems science and appl. (2020) 22 appendix 1. values of model parameters equation parameter description ( )dn t dt – (1-2) n = 3.3613, n = 0.0332 n = 23 0n = 0.75 the parameters of the function of gompertz approximating the experimental curves of growth of population of tumor cells without vaccination (control) ( )dn t dt – (3) 1 = 1 1v cvt z=  + cvz = 4.5 the moment of the first injection of a viral vaccine the moment the beginning of the immune response against the virus delay time of the immune response against the virus ( )vdn t dt – (4-5) 1n cnt z=  + the moment the beginning of the immune response against infected tumor cells сnz = 10.5 delay time of immune response against infected tumor cells vk = 0.25 avk = 0.8 ank = 0.8 the constant coefficients of the equation of dynamics of the number of cells after a single injection of the vaccine ( )dv t dt – (6) v = 0.1 the reproductive rate of the virus v = 15 the rate of death of viruses in their interaction with antibodies 0v = 0.015 the initial condition of equation (2), which reflects the dose of the viral vaccine ( )vda t dt – (7) a = 100 the rate of formation of antibodies from the single plasma cells av = 70 the rate of loss of antibody due to the interaction with viruses v = 5 the rate of reduction of the number of antibodies due to natural destruction max va = 1.05 the maximum calculated number of the antibodies ( )vdc t dt – (8) с = 100 the rate of formation of plasma cells cv = 4.5 the size factor vnc = 0.001 the initial number of plasma cells ( )nda t dt – (9) an = 30 the rate of formation of antibodies from the single plasma cells an = 6.2 the rate of decline of antibodies due to the interaction with the tumor cells nn = 6.3 the rate of the reduction of the number of antibodies due to the natural destruction max na = 4.64 the maximum calculated number of antibodies ( )ndc t dt – (10) cn = 76.677 the rate of formation of plasma cells cn = 38 the size factor nnc = 0.0001 the initial number of plasma cells ( )p t – (11) p = 0.3 pk = 0.95 *1/p t = the parameters of the p(t) function describing the dynamics of the decrease in the proportion of rapidly proliferating cells as the tumor size increases the moment at which the number of fractions of rapidly and mathematical modeling of antitumor viral vaccine therapy… copyright ©2020 assa. adv. in systems science and appl. (2020) 23 *t = 35 суток slowly proliferating cells are equal ( )dn t dt – (13) v = 4.5 сут. the delay in the growth of tumor cells as a result of their death under the action of the virus v = 9.8 сут. the delay in the growth of tumor cells as a result of their death under the action of antibodies against infected tumor cells 1 vt = 5.82 1 nt = 12.32 the moments of reaching the maximum number of infected tumor cells by the beginning of the immune response 2 vt = 7.13 2 nt = 14.29 the moments of reaching the minimum number of infected tumor cells by the end of the immune response ( )1 v v vn t = 1.31, ( )1 n n vn t = 1.15 the maximum values of the number of infected tumor cells before the immune response at each of the two stages of their death ( )2 v v vn t = 0.84, ( )2 n n vn t = 0.1 the minimum values of the number of infected tumor cells at the end of the immune response at each of the two stages of their death ( )1 v v vn t = 0.47, ( )1 n n vn t = 1.04 the number of infected cells killed in the first and second stages of immune system stimulation 1. introduction 2. problem statement adv syst sci appl 2021; 01:86–94 published online at https://ijassa.ipu.ru. enhanced results on stability criteria for linear time delay systems with distributed delay via relaxed double integral inequality r. jeetendra1∗, b. jeevanandan1,2 1assistant professor,department of mathematics, kongu engineering college, erode-638 060, india. 2assistant professor,department of mathematics, kongu engineering college, erode-638 060, india. abstract: this paper investigates the matter of stability criteria for linear time delay systems with distributed delay. firstly, a relaxed double integral inequality is established to estimate the double integral terms appearing within the derivative of lyapunov-krasovskii functionals (lkfs) with a triple integral term. unlike the recently introduced jensen’s inequalities, wirtinger based integral inequalities, refined jensen’s inequalities and therefore the auxiliary function based integral inequalities the proposed relaxed integral inequality provides large feasible solution region and fewer conservative results. secondly, by constructing an augmented lyapunov-krasovskii functional with a triple integral term, the robust stability criteria for linear time delay systems with distributed delay are given in terms of linear matrix inequalities (lmis), which may be easily computed by the lmi toolbox of matlab. finally, two numerical examples are performed to indicate the effectiveness of the proposed criterion. keywords: robust stability, delay-dependent stability, lyapunov functional, linear matrix inequalities, time-delay systems, distributed delay 1. introduction many dynamic systems in the real world inevitably have time delays, and such delays often cause poor performance, oscillation or even instability of the system. consequently, the stability issue of time-delay systems has attracted researchers for many years. the main attention is paid to determine the admissible delay region, for which the systems remain stable, by developing effective delay-dependent stability criteria via the lyapunov-krasovskii stability theory. the issue relies on the handling of the integral terms arising within the derivative of the lyapunov-krasovskii functionals (lkfs). the development of new methods for this problem has always been a very important consideration. so as to scale back conservatism of stability criteria, variety of techniques are presented, including for example, jensen’s integral inequality [1], the wirtinger-based integral inequality [2], the various forms of wirtinger-based double integral inequality [3-5], bessel-legendre inequality [6], auxiliary function-based integral inequalities [7,8], the free-weighting matrix approach [4,9-11], reciprocally convex method [12,13] and free matrix based multiple integral inequality [14]. the well-known jensen’s inequality is often adopted because it could lead on to a stability test with fewer matrix variables. recently, wirtinger integral inequality introduced in [2] is shown more powerful than jensen’s inequality. later, another forms of integral inequalities are reported in [6,7,15,16-19] to further reduce the conservatism. ∗corresponding author: jee4maths@gmail.com enhanced results on stability criteria for linear time delay systems... 87 this paper presents a new relaxed double integral inequality to estimate the double integral term within the derivative of lyapunov-krasovskii functionals(lkfs). a replacement stability criterion is established by applying the newly proposed inequality which provides less conservatism and fewer number of decision variables. to indicate the effectiveness of the proposed criterion, two numerical examples are provided. notations: throughout this paper, rn is the n-dimensional euclidean space and rn×n is the set of all n× n real matrices. x > 0(x ≥ 0) means that the matrix x is a real symmetric positive definite matrix (positive semi definite). i denote the identity matrix with appropriate dimensions; col {·} means a column vector. ∗ in a matrix represents the elements below the main diagonal of a symmetric matrix. the superscript t denotes the transpose of the matrix. 2. problem formulation consider the following system with state and distributed delays: ẋ(t) = ax(t) + a1x(t− h) + a2 ∫ t t−h x(s)ds, (2.1) x(t) = φ(t), t ∈ [−h, 0] where, x(t) ∈rn is the state vector, a,a1, a2 ∈rn×n are constant matrices, h is a constant time delay satisfying h > 0 and φ(t) is a continuous vector-valued initial condition. 2.1. lemma[20]: for a given matrix m > 0 , the following inequality holds for all continuously differentiable functions x : [a, b]→ rn: (b− a) ∫ b a ẋ(s)tmẋ(s)ds ≥ ωt 1mω1 + 3ωt 2mω2 + 5ωt 3mω3 + 7ωt 4mω4 where ω1 = x(b)− x(a) ω2 = x(b) + x(a)− 2 b− a ∫ b a x(s)ds ω3 = x(b)− x(a) + 6 b− a ∫ b a x(s)ds− 12 (b− a)2 ∫ b a ∫ b θ x(s)dsdθ ω4 = x(b) + x(a)− 12 b− a ∫ b a x(s)ds+ 60 (b− a)2 ∫ b a ∫ b θ x(s)dsdθ − 120 (b− a)3 ∫ b a ∫ b σ ∫ b θ x(s)dsdθdσ relaxed double integral inequality 2.2. lemma: for symmetric positive definite matrix z ∈ rn×n, scalars α < β, and vector φ : [α, β]→ rn such that the integration concerned is well defined, the following inequality holds:∫ β α ∫ β u ϕt (s)zϕ(s)dsdu ≥ 2 (β − α)2 ωt 5zω5 + 16 (β − α)2 ωt 6zω6 (2.2) copyright c© 2021 assa. adv syst sci appl (2021) 88 r. jeetendra, b. jeevanandan where ω5 = ∫ β α ∫ β u ϕ(s)dsdu ω6 = − ∫ β α ∫ β u ϕ(s)dsdu+ 3 β − α ∫ β α ∫ β u ∫ β θ ϕ(s)dsdθdu proof: for a function λ(s) = k1 + k2s, integration by parts, we have∫ β α ∫ β u λ(s)ϕ(s) dsdu = λ(a) ∫ β α ∫ β u ϕ(s) dsdu+ 2k2 ∫ β α ∫ β u ∫ β θ ϕ(s)dsdθdu by setting λ(a) = −1, 2k2 = 3 β−α , the above equality is rewritten as∫ β α ∫ β u λ(s)ϕ(s) dsdu = ω6 then the following equality is obtained for any vector ω0 and any matrix m > 0:∫ β α ∫ β u λ(s)ωt 0mϕ(s) dsdu = ωt 0mω6 similarly, the following equalities are derived:∫ β α ∫ β u ωt 0lϕ(s) dsdu = ωt 0lω5∫ β α ∫ β u ωt 0lr −1ltω0 dsdu = (β − α)2 2 ωt 0lr −1ltω0∫ β α ∫ β u ωt 0lr −1mtλ(s)ω0 dsdu = 0∫ β α ∫ β u λ2(s)ωt 0mr−1mtλ(s)ω0 dsdu = (β − α)2 16 ωt 0mr−1mtω0 using the above equalities and the schur complement derives the following equality: ∫ β α ∫ β u [ ω0 λ(s)ω0 ϕ(s) ]t [ lz−1lt lz−1mt l ∗ mz−1mt m ∗ ∗ z ][ ω0 λ(s)ω0 ϕ(s) ] dsdu = ∫ β α ∫ β u ϕt (s)rϕ(s)dsdu+ sym { ωt 0lω5 + ωt 0mω6 } + (β − α)2 2 ωt 0 { 8lz−1lt +mz−1mt 8 } ω0 ≥ 0. where ωt 0 = [ ωt 5 ωt 6 ] , l = −2 (β − α)2 [z 0]t and m = −16 (β − α)2 [0 z] copyright c© 2021 assa. adv syst sci appl (2021) enhanced results on stability criteria for linear time delay systems... 89 that is, ωt 0l = −2 (β − α)2 ωt 5z and ωt 0m = −16 (β − α)2 ωt 6z which leads to (2.2). this completes the proof. 2.3. remark: the proposed relaxed double integral inequality provides the tightest estimation value of the double integral term ∫ b a ∫ b θ xt (s)zx(s)dsdθ > 0, compared with the widely used jensen’s integral inequality and wirtinger-based integral inequality. moreover, the additional positive term 16 (β−α)2 ωt 6zω6 reduces the estimation gap. therefore, the proposed relaxed double integral inequality will cause less conservative than the prevailing ones within the literature. by setting ϕ(s) = ẋ(s), the subsequent lemma are often obtained from the above lemma 2.2. 2.4. lemma: for symmetric positive-definite matrix z ∈ rn×n, scalars α < β, and vector ẋ : [α, β]→ rn such that the integration concerned is well defined, the following inequality holds:∫ β α ∫ β u ẋt (s)zẋ(s)dsdu ≥ 2χt1zχ1 + 16χt2zχ2 where χ1 = x(β)− 1 β − α ∫ β α x(s)ds χ2 = −1 2 x(β)− 1 β − α ∫ β α x(s)ds+ 3 (β − α)2 ∫ β α ∫ β u x(s)dsdu 3. main results in this section, delay dependent stability criteria for the system with distributed delays are derived interms of lmi as follows: 3.1. theorem: given h > 0, the system (2.1) is assymptotically stable if there exists positive definite matrices p ∈ r4n×4n, q, s, z ∈ rn×n, such that the following lmi holds: ξ = γpυt + υpγt + ψ < 0 (3.3) where γ = [e1 e3 e4 e5] υ = [ e0 e1 − e2 he1 − e3 h2 2 e1 − e4 ] copyright c© 2021 assa. adv syst sci appl (2021) 90 r. jeetendra, b. jeevanandan ψ = e1qe t 1 − e2qet2 + h2e0se t 0 + h2 2 e0ze t 0 − (e1 − e2)s(e1 − e2)t − 3 ( e1 + e2 − 2 h e3 ) s ( e1 + e2 − 2 h e3 )t − 5 ( e1 − e2 + 6 h e3 − 12 h2 e4 ) s ( e1 − e2 + 6 h e3 − 12 h2 e4 )t − 7 ( e1 + e2 − 12 h e3 + 60 h2 e4 − 120 h3 e5 ) s ( e1 + e2 − 12 h e3 + 60 h2 e4 − 120 h3 e5 )t − 2 ( e1 − 1 h e3 ) z ( e1 − 1 h e3 )t − 16 ( −1 2 e1 − 1 h e3 + 3 h2 e4 ) z ( −1 2 e1 − 1 h e3 + 3 h2 e4 )t e0 = ae1 + a1e2 + a2e3 and ei ∈ r5n×n are elementary matrices, for example et2 = [0 i 0 0 0]. proof: consider a lyapunov-krasvoskii canditate as v (t) = v1(t) + v2(t) + v3(t) + v4(t) where v1(t) = ηt (t)pη(t), v2(t) = ∫ t t−h xt (α)qx(α)dα v3(t) = h ∫ t t−h ∫ t β ẋ(α)tsẋ(α)dαdβ and v4(t) = ∫ t t−h ∫ t β ∫ t σ ẋ(α)tzẋ(α)dαdσdβ. where η(t) = col [ x(t), ∫ t t−h x(α)dα, ∫ t t−h ∫ t β x(α)dαdβ, ∫ t t−h ∫ t β ∫ t σ x(α)dαdσdβ ] the time derivative v (t) along the trajectories of system can be computed as follows: v̇1(t) = 2ηt (t)p η̇(t) = 2ξt (t)γpυt ξ(t) v̇2(t) = xt (t)qx(t)− xt (t− h)qx(t− h) v̇3(t) = h2ẋt (t)sẋ(t)− h ∫ t t−h ẋ(α)tsẋ(α)dα v̇4(t) = h2 2 ẋt (t)zẋ(t)− ∫ t t−h ∫ t β ẋ(α)tzẋ(α)dαdβ where ξ(t) = col [ x(t), x(t− h), ∫ t t−h x(α)dα, ∫ t t−h ∫ t β x(α)dαdβ, ∫ t t−h ∫ t β ∫ t σ x(α)dαdσdβ ] copyright c© 2021 assa. adv syst sci appl (2021) enhanced results on stability criteria for linear time delay systems... 91 and it can be rewritten as v̇ (t) = ξt (t) { γpυt + υpγt + e1qe t 1 − e2qet2 + h2e0se t 0 + h2 2 e0ze t 0 } ξ(t) − h ∫ t t−h ẋ(α)tsẋ(α)dα− ∫ t t−h ∫ t β ẋ(α)tzẋ(α)dαdβ. applying lemma 2.1 and 2.4 to the above integrals leads to −h ∫ t t−h ẋ(α)tsẋ(α)dα ≤ −ξt (t) { (e1 − e2)s (e1 − e2)t + 3 ( e1 + e2 − 2 h e3 ) s ( e1 + e2 − 2 h e3 )t + 5 ( e1 − e2 + 6 h e3 − 12 h2 e4 ) s ( e1 − e2 + 6 h e3 − 12 h2 e4 )t + 7 ( e1 + e2 − 12 h e3 + 60 h2 e4 − 120 h3 e5 ) s ( e1 + e2 − 12 h e3 + 60 h2 e4 − 120 h3 e5 )t } ξ(t) − ∫ t t−h ∫ t β ẋ(α)tzẋ(α)dαdβ ≤ −ξt (t) { 2 ( e1 − 1 h e3 ) z ( e1 − 1 h e3 )t + 16 ( −1 2 e1 − 1 h e3 + 3 h2 e4 ) z ( −1 2 e1 − 1 h e3 + 3 h2 e4 )t } ξ(t) hence, we have v̇ (t) ≤ ξt (t) { γpυt + υpγt + ψ } ξ(t) v̇ (t) ≤ ξt (t)ξξ(t). this completes the proof. 4. numerical examples in this section, two examples are used to illustrate the effectiveness of the proposed method. 4.1. example: consider the following system with distributed delay: ẋ(t) = [ 0.2 0 0.2 0.1 ] x(t) + [ 0 0 0 0 ] x(t− h) + [ −1 0 −1 −1 ] ∫ t t−h x(s)ds copyright c© 2021 assa. adv syst sci appl (2021) 92 r. jeetendra, b. jeevanandan the purpose is to match the utmost allowable upper bounds of h that guarantees the asymptotic stability of the above system. table 4.1 lists the computed maximum allowable upper bounds and also the number of decision variables which keep the system stability by different methods. from table 4.1, it’s clear that the proposed approaches can provide higher upper bounds than those within the existing results. it should be noted that our method provides maximum allowable boundary which is adequate to the analytical bound with fewer number of decision variables. table 4.1. upper bounds on h obtained for example 4.1 methods maximum h allowed nodv chen and zheng 2007 1.6339 85 seuret and gouaisbaut 2013 1.877 16 park et al. 2015 1.9504 59 zeng et al. 2015 2.0395 75 trinch 2015 2.0395 27 zhao et al. 2017 2.0402 45 park et al. 2018 2.0412 42 3.1 theorem 2.0412 39 analytical bound 2.0412 4.2. example: consider the following system with distributed delay: ẋ(t) = [ −2 0 0 −0.9 ] x(t) + [ −1 0 −1 −1 ] x(t− h) + [ 0 0 0 0 ] ∫ t t−h x(s)ds table 4.2 lists the computed upper bounds by different methods and it shows that our method provides an upper bound which is quite close to the analytical bound. table 4.2. upper bounds on h obtained for example 4.2 methods maximum h allowed nodv zhao et al. 2017 6.1663 45 zeng et al. 2015 6.1664 75 gu et al. 2003 (n=3) 6.171 67 chen et al. 2016 6.1719 106 park et al. 2018 6.1719 42 3.1 theorem 6.1719 39 analytical bound 6.1725 5. conclusion in this article, delay dependent stability criteria for linear time-delay system with distributed delay are proposed by the employment of the lyapunov method . by the development of augmented lyapunov functional and relaxed double integral inequality, the delay dependent stability criterion has been proposed. compared to the recently proposed integral inequalities the obtained ones could provide more accurate estimations on the handling of the integral terms arising within the derivative of the lyapunov-krasovskii functionals(lkfs) the obtained stability condition provides larger feasible solution region and fewer conservatism with fewer number of decision variables than the present ones within the literature. two numerical examples are presented to indicate the effectiveness of the proposed approach. copyright c© 2021 assa. adv syst sci appl (2021) enhanced results on stability criteria for linear time delay systems... 93 acknowledgements the authors sincerely thank the editors and anonymous reviewers for their careful reading, constructive comments and suggestions to improve the quality of the manuscript. references 1. gu, k., chen, j. & kharitonov, v.l. (2003) stability of time-delay systems. springer science bussiness media. 2. seuret, a., & gouaisbaut, f. (2013) wirtinger-based integral inequality: application to time-delay systems. automatica 2013, 49(9), 2860–2866. 3. park, m., kwon, o., park, j.h., lee, s., & cha, e. (2015) stability of time-delay systems via wirtinger-based double integral inequality. automatica 2015, 55, 204–208. 4. zeng, h.-b., he, y., wu, m., & she, j. (2015) new results on stability analysis for systems with discrete distributed delay. automatica 2015, 60, 189–192. 5. zhao, n., lin, c., chen, b., wang, & q.-g. (2017) a new double integral inequality and application to stability test for time-delay systems. applied mathematics letters 2017, 65, 26–31. 6. seuret, a., & gouaisbaut, f. (2015) hierarchy of lmi conditions for the stability analysis of time-delay systems. systems control letters 2015, 81, 1–7. 7. park, p., lee, w.i., lee, & s.y. (2015) auxiliary function-based integral inequalities for quadratic functions and their applications to time-delay systems. journal of the franklin institute 2015, 352(4), 1378–1396. 8. park, p., lee, w.i., lee, & s.y. (2016) auxiliary function-based integral/summation inequalities: application to continuous/discrete time-delay systems. international journal of control, automation and systems 2016, 14(1), 3–11. 9. he, y., wang, q.-g., xie, l., & lin, c. (2007) further improvement of freeweighting matrices technique for systems with time-varying delay. ieee transactions on automatic control 2007, 52(2), 293–299. 10. chen, w.-h., zheng, & w.x. (2007) delay-dependent robust stabilization for uncertain neutral systems with distributed delays. automatica 2007, 43(1), 95–104. 11. zeng, h.-b., he, y., wu, m.,& she, j. (2015) free-matrix-based integral inequality for stability analysis of systems with time-varying delay. ieee transactions on automatic control 2015, 60(10), 2768–2772. 12. park, p., ko, j.w., & jeong, c. (2011) reciprocally convex approach to stability of systems with time-varying delays. automatica 2011, 47(1), 235–238. 13. lee, w.i.,& park, p. (2014) second-order reciprocally convex approach to stability of systems with interval time-varying delays. applied mathematics and computation 2014, 229, 245–253. 14. chen, j., xu, s., & zhang, b. (2016) single/multiple integral inequalities with applications to stability analysis of time-delay systems. ieee transactions on automatic control 2016, 62(7), 3488–3493. 15. trinh, h. (2015) refined jensen-based inequality approach to stability analysis of timedelay systems. iet control theory applications 2015, 9(14), 2188–2194. 16. kim, j.-h. (2016) further improvement of jensen inequality and application to stability of time-delayed systems. automatica 2016, 64, 121–125. 17. lee, t.h., park, j.h., park, m.-j., kwon, o.-m., jung,& h.-y. (2015) on stability criteria for neural networks with time-varying delay using wirtinger-based multiple integral inequality. journal of the franklin institute 2015, 352(12), 5627–5645. 18. zhang, c.-k., he, y., jiang, l., wu, m., zeng,& h.-b. (2016) stability analysis of systems with time-varying delay via relaxed integral inequalities. systems control letters 2016, 92, 52–61. copyright c© 2021 assa. adv syst sci appl (2021) 94 r. jeetendra, b. jeevanandan 19. park, m.-j., kwon, o., & ryu, j. (2018) generalized integral inequality: application to time-delay systems. applied mathematics letters 2018, 77, 6–12. 20. zhang, s., & qi, x (2017) improved integral inequalities for stability analysis of interval time-delay systems. algorithms 2017, 10(4), 134. copyright c© 2021 assa. adv syst sci appl (2021) introduction problem formulation lemma[20]: lemma: remark: lemma: main results theorem: numerical examples example: example: conclusion adv syst sci appl 2020; 03:153–165 published online at https://ijassa.ipu.ru. control model of the floodplain territories structure alexander a. voronin1*, alexander v. khoperskov1, inessa i. isaeva1, anna yu. klikunova1 1volgograd state university, volgograd 400062, russia abstract: we have developed a model of an integrated hydrological, environmental and socioeconomic structure of the floodplain area, which is the subject of strategic control based on the observational data, the results of numerical hydrodynamic and geoinformation modeling. the state assessment of each structural element is determined by the degree of correspondence between the hydrological type and the socio-economic type of the floodplain area. the goal of control is to implement a territorial structure that maximizes the value of the aggregated criterion for a stable state of the system. the control mechanism is a complex of hydraulic engineering projects (multiproject) in various channels in the floodplain area. the simulations results of the implementation of the sustainable development strategy of the environmental and socio-economic system of the volga-akhtuba floodplain northern part are presented. keywords: simulations, development management, hydrological regime, environmental and socio-economic systems, hydraulic engineering projects, volga-akhtuba floodplain 1. introduction the problem of sustainable development (sd), which until recently was in the focus of management/control theory and practice, has largely lost its popularity in the last decade. in our opinion, the reason for this is the existence of a difficult gap between existing conceptual and formal models, on the one hand, and the development management practice of real environmental and socio-economic systems (eses), on the other. for example, in [1, 2] the authors developed the concept of sd control based on game-theoretic models and the active systems theory. the main parameters of these models are the “system compatibility indices” calculated on the basis of the objective functions of its actors and ecosystem sustainability criteria. however, justifying these values that underlie these and many other sd models is an almost insurmountable difficulty in most cases. indeed, the shortest periods of degradation and, moreover, the restoration of regional ecosystems are decades, so the “compatibility indices” should characterize conflicts between generations to a greater extent than conflicts between its actors. in the absolute majority of cases, the development of modern eses occurs within the framework of the “weak sustainability” concept, according to which natural capital is exchanged for human and productive capital with varying degrees of efficiency. a possible transition to sd in a state of degrading ecosystems requires both a study of the conditions for their restoration and a readiness for the necessary self-restraint in the context of economic competition. in addition, the level of modern theoretical and empirical knowledge about real ∗corresponding author: a.voronin@volsu.ru 154 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova ecosystems at the regional level does not allow us to confidently identify the their stability thresholds and predict their dynamics. the floodplains of regulated large rivers with a developed eses are among the territories in which, on the one hand, this gap manifests itself in the most acute form, on the other, it can potentially be bridged more easily than on other areas. the high conditionality of their territorial structure by the unstable nature of floods is an objective factor of weak environmental sustainability, which, as a rule, is destroyed in the process of their socioeconomic development. the development of high-performance supercomputers and effective methods of computational fluid dynamics in recent years makes it possible to effectively simulate a realistic flood regime, which is a system-forming factor in floodplain areas, allowing to significantly reduce the complexity of the constructed sd models. an example of a model successful construction for sustainable development of a regional environmental system is the work [3] in which the authors created the model of ecological and economic control of the biological resources extraction of the azov sea with sustainable reproduction. their complex numerical simulation model successfully combines models of hierarchical differential games, biological kinetics, hydrodynamics using high performance computational algorithms for supercomputers. however, we note that the ecological criterion in this model (the stability of the biosystem) essentially coincides with the long-term economic criterion (sustainable extraction of biological resources). the work [4] is aimed at creating computer technology for effectively assessing the reliability of meeting the demands of volga water users and, in particular, for the lower volga including projects to improve environmental conditions. the model is based on the multicriteria analysis methods and the compromises theory. integrated decision support systems (dss) allow solving the problems of planning water resources management in river basins, scarcity of river water resources for socio-economic needs, rational balance between socio-economic and environmental needs. let us point out some good examples of dss for the haihe river basin [5, 6], the northern part of the volga-akhtuba floodplain [7], the illinois river [8], the pinhao river basin [9], the biosphere reserve elbe river landscape in lower saxony [10], the pecos river [11]. this work and our research use similar methods and technologies, such as interdisciplinary modeling, gis technologies, multi-criteria assessments, peer review. our research is distinguished by a more active use of direct numerical hydrodynamic modeling. the objective of this study is to solve the problem of identification and realizability of the values of the territorial structure parameters, necessary for sd. our model is based on a simplified version of the “indices of system compatibility” (see [1]), implemented as a vector for assessing the self-compatibility of the territorial structure of the eses, the components of which are calculated by territorial aggregation of expert consistency indices between the hydrological (h) type and the environmental and socio-economic (ese) type for each site of the territory. the practical implementation of the proposed research concept is the identification and study of the attainability of the target h-structure of the northern part of the volga-akhtuba floodplain within the volgograd region (hereinafter called vaf) in the context of the sustainable development of its eses through hydrotechnical multiproject, based on the construction of hydraulic structures (canals and dams) with calculated characteristics. these characteristics should stabilize the h-structure of the floodplain area and thereby make it possible to implement the mechanisms of ecological and economic control by its economic entities, which is described in [12]. the set of software tools presented in this paper includes the decision support system (dss), which is an extension of the dss for river water resources distribution. this set of tools for control sustainable development of floodplain areas is an implementation of the general model of the sd information-analytical system proposed by the authors of works [1, 2]. copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 155 2. research methods, tools and technologies the hydrological regime of the floodplain area is determined by three main factors: the river hydrograph (q(t) is the volume of water flowing per unit of time or the variation of discharge), the channel structure and the area topography [13–15]. natural changes in the last two factors, as a rule, are insignificant for several tens of years, and the first factor has high variability at short time intervals. if the distribution function of flooding volumes for the period t is determined, then the frequencies of flooding of interfluve fragments calculated for this period remain constant over time t (t � t ). in other words, the h-structure of the floodplain area z(h) exists stably during the time interval t . this structure is described as a set of maps territories 〈m(h) i (i = 1, ..., n)〉, flooded by water with frequencies from the specified ranges, and the distribution function of the relative flooded area at the floods peak ψ(x) : p (s(h)/s < x) gives the aggregated state of the system (s is the total area of the vaf). the time interval t is determined experimentally from the observations data of the parameters of flood river hydrographs over a long period t � t . the choice of the value of t affects the existence conditions, the type and accuracy of determining the stable hydrological structure z(h). in turn, the stability of the hydrological structure is the main factor in the formation of the territorial environmental and socio-economic (ese) systems of the floodplain area z(ese), characterizing the spatial sites distribution by functionally homogeneous ese-types. the term “territorial natural structure” in this work means a biotopes system of the main structural elements of a floodplain ecosystem. ese-structure can be built on the basis of the corresponding cadastral maps. hand esestructures define the complex structure z(c) = z(h) ∩ z(ese), which is given by flood maps of functionally homogeneous territories or ese maps of hydrologically homogeneous territories m(c) = ∪ijm(c) i,j , wherem(c) i,j =m(h) i ∩m(ese) j (i = 1, ..., n; j = 1, ...,m), (2.1) where ‖ s(c) ij ‖, s (c) ij = s(m(c) ij ) is the land area. the complex structure z(c) is also determined by the distribution functions of the relative area of the flooded terrain at the floods peak ψj(x) : p (s (h) j /s (c) j < x), s (c) j = n∑ i=1 s (c) ij , (2.2) where the number of frequency intervals n and their sizes, the number of ese-types m are determined by the purpose and accuracy of modeling. the construction of a hydroelectric dam and / or abrupt climatic changes are often the reasons for the formation of a new h-structure and an adaptive transformation of the esestructure of the floodplain. but these factors can also be the reason for the activation of slow changes in rivers channel topography, disrupting the stability of the h-structure, and, as a result, causing a decrease in the degree of self-compatibility between the hand esestructures, which leads to ese-degradation. ese-degradation counteraction requires periodic fine tuning of hand ese-structures. within the framework of the weak stability concept, the control task is the implementation of a constant or variable h-structure, which ensures high efficiency and safety of a constant or variable territorial ese-structure. the objective h-structure should first of all ensure the sustainability of the ecosystem in the case of a sustainable development concept. this in turn requires the stabilization of the h-structure itself with parameters that ensure stable flooding of its biotope. in the absence of a floodplain eses model, we estimate the potential efficiency of z(c) based on the expertly constructed vector of “disharmony” δ = (δ1, ..., δm), δj = 1− copyright © 2020 assa. adv syst sci appl (2020) 156 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova kj (j = 1, ...,m), in which the functions kj = k(ψj) ∈ [0; 1] characterize the degree of correspondence between hydrological regime of functionally homogeneous territory and its ese-type j. the complete uncertainty of the kj functions for some j, their exact coincidence for different j, the insensitivity of the functions k(ψj) for all j to certain ranges of flooding frequencies allow us to reduce the dimension of the actual structures z(h),z(ese),z(c) and the dimension of the problem (2.3). we can rewrite the general control problem for the structure of the floodplain area as a multi-criteria optimization problem: δτ+θ(u) −→ min u∈u , z(c) τ+θ∈ωz , (θ = 1, 2, ...,θ) , (2.3) where ωz is the set of admissible structures, depending on the chosen concept of development stability (weak or strong stability), u is the set of realizations of a hydraulic engineering multiproject, θ is planning horizon, the τ index corresponds to the calculation of hydrological structures for the period [τ − t , τ ]. the lack of effective methods for solving inverse problems of hydrodynamics allows using only heuristic branch-and-bound methods to solve the problem (2.3), while the exact methods can be used at some stages in (2.3). special particular problems of (2.3) and methods for their solution are described in the works [7, 16] using hydrodynamic, geoinformation and game-theoretic simulations. 3. tools and simulations results for volga-akhtuba floodplain vaf is located in the lower reaches of the volga river (the area of the northern part of the floodplain is approximately 870 km2, the number of small natural channels exceeds 200 with total length of about 1000 km). figure 3.1 shows a diagram of the general hydrological system of vaf, which can be conditionally divided into the western part, which is flooded mainly from the volga river. the hydrological regime in the eastern part is determined by the water dynamics in the akhtuba river and in the large middle trunk water system of natural channels of the floodplain. the structure of our software package for solving the problem (2.3) for eses of vaf is shown in the figure 3.2. individual parts of this software are described in detail in [7,16,17]. the key component is a numerical hydrodynamic model of the area flooding, which takes into account all the main physical and meteorological factors [13, 17, 18]. this mathematical model is based on 2d shallow water equations, which has been successfully applied to similar problems, see, for example, [19, 20]. the existence, type and accuracy of determining the stable hydrological structure of the vaf under the condition of the stability of the channel system and topography in the period 1962–2018 are established as a result of computational experiments. we have generated sequences of all possible samples of fragments of the time series v (t) (t = 5, 6, ..., 57), containing t data units, using the observations array of normalized volumes of v for flood hydrograph through the hydroelectric dam. for these fragments, we have constructed a set of experimental distribution functions ψ(t)(x) : p (v (t)/vmax < x) (vmax is the maximum volume for the entire observation period). the performed analysis indicates that the error in determining the magnitude of the flood volume realized in the vaf with frequency of nflood ≥ 0.85 for t ≥ 20 does not exceed 0.05. initial hydrological data are based on the official primary statistical information on the operation of the volga hydroelectric power station (see also [21]). the regression model for describing the flow of flood waters into the akhtuba river in the period 1962-2018 was built based on the results of measurements of flood water levels at the gauging station in the volgograd city [7]. the modeling results and the constructed retrospective hydrological copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 157 fig. 3.1. scheme of the complex of hydraulic engineering projects. the blue lines show the system formed by the erik gniloj and the erik pahotnyj in the northern part and by the kashirskij water tract in the eastern part of the vaf. the yellow lines show only some of the small channels associated with this system. light purple channels are filled directly from the volga river. red lines indicate canals from the akhtuba river. magenta diamonds show possible positions of dams of different types. magenta dashed line delimits the conditional “western” zone from the “eastern” zone. structure z(h) 1982 show that the main factor of the loss of stability of this hydrological structure and the drying up of the interfluve is a gradual decrease in the average volume of low-flow and flood waters in the akhtuba river due to the progressive degradation of the volga riverbed after the construction of the volga hydroelectric power station in 1962. figure 3.3 shows the distribution functions ψ1982(x) and ψ2018(x). comparative analysis of curves 1 and 2 shows both the general reduction of the flooded territory and the qualitative change in its structure. the share of sustainably flooded territory is reduced from 85 percent to 47 percent. the northern territory of the vaf cadastral map contains 607 cadastral parcels that belong to 35 cadastral types. moreover, cadastral types have not been established for 35 percent of the vaf. such areas and recreational areas are not evaluated and the k(ψj) functions are undefined for them. we combine the remaining types into four groups based on the proximity of the estimated coefficients. the first two are wetlands (j = 1) and flood-meadows (type of farmland) (j = 2) and they collectively represent the floodplain ecosystem biotope for which non-flooding is damage. the other two are areas of residence of the population (j = 3) and forests + economic areas (j = 4), for which flooding is a damage. very small floods copyright © 2020 assa. adv syst sci appl (2020) 158 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova fig. 3.2. the structure of the software package for our model of the floodplain territory control. on the volga river were observed every 10–15 years in the period 1880–1930 before the construction of the volga system of hydroelectric power plants [21]. this allows us to relate the stability boundary of the floodplain ecosystem with the value of the threshold frequency of flooding of the floodplain biotope nflood = 0.85. table 3.1. evaluative coefficients of the efficiency of typical elements of the volga-akhtuba floodplain. nflood 0 (0; 0.25] (0.25; 0.50] (0.50; 0.85) [0.85; 1.0] ki,1 0 0 0 0.25 1 ki,2 0 0 0.25 0.5 1 ki,3 1 0 0 0 0 ki,4 1 0.1 0 0 0 copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 159 fig. 3.3. model integral distribution functions of the flood area are shown by the following curves: 1 — ψ1982(x); 2 — ψ2018(x); 1e is eastern zone for 1962–1982; 1w is western zone for 1962–1982; 2e is eastern zone for 1998–2018; 2w is western zone for 1998–2018. the symbols 2wp and 2ep denote the corresponding curves, taking into account the implementation of hydraulic projects. the horizontal dashed line corresponds to flooding of the floodplain with a frequency of at least 0.85. evaluation of the effectiveness of the territorial structure of vaf is carried out using special functions of the form k(ψj) = 1 sj n∑ i=1 kij(nflood)sij (j = 1, ...,m) , (3.4) where kij(nflood) are the dependencies of the coefficients on the flooding frequency (see table 3.1). thus, the current complex (integrated) structure of vaf is characterized by n = 5, m = 4. the presence of a significant part of the territory with undefined evaluation functions k(ψj) is a source of great uncertainty in the task (2.3). the expert estimation of the threshold value of the area of the floodplain ecosystem biotope is approximately 0.5s, which noticeably exceeds both the area of modern stable flooding (0.3s) and the value of s1 + s2. to remove this uncertainty, we included in the first group of ese structures (j = 1) a part of the territory with indefinite evaluation functions k(ψj), which was steadily flooded in the period 1962– 1982. we estimate the error in determining the this territory area as 0.02, while it is 0.07 for the period 1998-2018. the reason for this is the stable flooding in the past of all the plains, while the modern border of the stable flooded area is within the plains. figure 3.4 shows both the target areas of stable flooding and areas where flooding causes damage. for the ese structure z(ese) 2018 changed in this way, we calculated the “disharmony” vectors for the complex and virtual structures using a series of computational experiments. z(c) 2018 = z(h) 2018 ∩ z (ese) 2018 : δ2018 = (0.45; 0.53; 0; 0.11) , z(c) virt = z(h) 1982 ∩ z (ese) 2018 : δvirt = (0; 0.06; 0.02; 0.17) . the first vector shows a high level of disharmony of the vaf eses due to catastrophic drying up of the interfluve, the second vector is the objective vector for solving the problem (2.3). the main part of the multiproject is the channel from the volgograd reservoir to the akhtuba river bypassing the volga river and a dam that closes the entrance to akhtuba from the volga river (see “barrier dam on the akhtuba river” on the figure 3.1). the parameter of the flood hydrograph is the maximum discharge through the canal (qch). the product of qch by the duration of the flooding τf (which is equal to 15 days in all experiments) gives copyright © 2020 assa. adv syst sci appl (2020) 160 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova fig. 3.4. upper panel: the target area for stable flooding of the vaf. bottom panel: areas where flooding causes damage. the volume of flood waters that determine the flooded territory map and its area. since the amount of flood water entering the akhtuba river is less than 10 percent of the total spring water (v (flood) tot ), the project allows for stable flooding of almost any target area in the eastern part of the vaf, limited by the condition of non-flooding of socio-economic territories (curve 2ep in figure 3.3) even in conditions of high natural variability of v (flood) tot . considering that qch is a small value of the total water discharge in the river volga during the flood period, we restrict ourselves to the case of a fixed qch. we determine the location of the dams by heuristic analysis of the directions of water flows across the boundary between the east and west zones (see dashed line in figure 3.1) during the flooding stage. the designed dams in the channels should prevent the flow of water from the western zone to the eastern zone. features of the real topography limit the number of copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 161 fig. 3.5. maps of the stably flooded area of the vaf in the period 1998-2018. up panel: model without bypass channel with qch = 0. bottom panel: model with bypass channel qch = 2000 m3 sec−1. positions of such dams. we have found the dam locations that have a noticeable effect using hydrodynamic simulations. the results of hydrodynamic modeling and optimization of projects for deepening small channels and installing additional dams on the vaf territory (see works [22–24]) showed the possibility of reducing the discussed negative effect by 8 percent. figure 3.1 demonstrates the hydrological system of the floodplain and the complex of hydraulic projects. we distinguish copyright © 2020 assa. adv syst sci appl (2020) 162 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova the eastern zone (“e”), adjacent to the akhtuba river, and the western zone (“w”), for which the hydrological regime is determined directly by the volga river. we have highlighted in red small channels in the western part of the vaf, for which deepening of the channels is also advisable. the optimal locations of dams on small canals are marked by rhombuses in figure 3.1. fig. 3.6. dependencies of the vector components of “disharmony” δm on qch for the eses. figure 3.3 shows the distribution functions of flooded areas for two time intervals (1962– 1982 and 1996–2018) separately for the western zone (curves 1w, 2w, 2wp) and for the eastern zone (curves 1e, 2e, 2ep), which are highlighted in figure 3.1. the horizontal dashed line corresponds to stable flooding of the interfluve. the post-design distribution function for the western zone is denoted as 2wp. a comparison of the 2w and 2wp curves shows an approximately 25 percent reduction in stable flooded area. distribution function x = const = 0.47 (see vertical line 2ep) corresponds to the annual flooding with qch = 2000 m3 sec−1. the arrow in figure 3.1 indicates the position of a possible large bypass channel from the volgograd reservoir to the akhtuba river, the project of which is being discussed. such construction should include a barrier dam on the akhtuba river. figure 3.5 shows the results of our spring flood simulations with and without such a bypass channel. the bypass canal can significantly improve the hydrological regime for the eastern area of the vaf. small projects are based on the work to clear and deepen small channels (they are indicated in yellow, purple and red in figure 3.1), as well as the use of special dams to regulate the water flow in the middle channels. figure 3.6 shows the curves δm(qch), which determine the dependence of the objective functions on the main project parameter qch. the values δ(1) 1 and δ (1) 2 correspond to the project with a bypass canal and a barrier dam, and the values δ(2) 1 and δ (2) 2 additionally include hydraulic engineering projects on the territory of the vaf (deepening of channels, dams on small channels). the case qch = 0 corresponds to the pre-design state and δ2018 = (0.45; 0.53; 0; 0.11). the increasing sequence of qch values determines the sequence of nested flood maps of the vaf at the peak of floods. there are two critical values of the parameter qch that define the set of feasible solutions to problem (2.3). the value qch1 determines the largest value of qch, at which the equalities δ3 = δ4 = 0 are satisfied. this value corresponds to the largest flood map for the peak of high water included in the stable flood target map. the value qch2 defines the largest flood map by inclusion, at which the equality δ3 = 0 is fulfilled. thus, the segment [qch1;qch2] is the “compromise zone” that contains the value of qch. a decrease in copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 163 the total disharmony index δ1 + δ2 + δ4 (δ1 = δ1e + δ1w, δ2 = δ2e + δ2w) in this segment is possible if some economic zones are added to the multiproject or if the target area of stable flooding is reduced. 4. conclusion we present an approach to the control of the spatial structure of the floodplain territory, which is the basis for creating strategies for their development. the consistency functions between the hydrological type of each spatial structural element and its ese-type are used as objective functions. the hydrotechnical multiproject for the interfluve channels is a control tool. the implementation of this approach for the northern part of the volga-akhtuba floodplain makes it possible to determine the parameters of the hydraulic engineering multiproject, which provide the conditions for its sustainable development. the analysis of expenses and ese-effects of the implementation of the investigated multiproject are beyond the scope of our task. we also do not discuss here game-theoretic models of the environmental and economic control mechanisms of the economic entities of the vaf. three such mechanisms (the control of the economic entity of the volga hydroelectric power station, the control of the urbanization process, the mechanism of co-financing for the project of deepening small channels) were investigated in our previous works [16, 24, 25]. modeling a complex of projects and mechanisms for sustainable development of the vaf territory, as well as a comparative analysis with a similar complex that implements the strategy of weak sustainability, are the subject of our future research. undoubtedly, each floodplain territory has a unique structure, therefore, the combination of hydraulic projects and mechanisms of ecological and economic control in each of them is also individual. however, we believe that the proposed approach to the control of the territorial structure of floodplain environmental and socio-economic systems will be useful both for theoretical research and for the practical implementation of strategies for their sustainable development. acknowledgements the authors were supported by the ministry of science and higher education of the russian federation (government task no. 0633-2020-0003). references 1. ougolnitsky, g. a. (2017). a system approach to the regional sustainable management. advances in systems science and applications, 17(2), 52–62. 2. ougolnitsky, g. a., anopchenko, t. y., gorbaneva, o. i., lazareva, e. i., & murzin, a. d. (2018). systems methodology and model tools for territorial sustainable management. advances in systems science and applications, 18(4), 136–150. 3. sukhinov, a. i., chistyakov, a. e., ugol’nitskii, g. a., usov, a. b., nikitina, a. v., puchkin, m. v., & semenov, i. s. (2017). game-theoretic regulations for control mechanisms of sustainable development for shallow water ecosystems. automation and remote control, 78(6), 1059–1071. 4. bolgov, m. v., buber, a. l., komarovskii, a. a., & lotov, a. v. (2018). searching for compromise solution in the planning and managing of releases into the lower pool of the volgograd hydropower system. 1. strategic planning. water resources, 45(5), 819–826. 5. ge, y., li, x., huang, c., & nan, z. (2013). a decision support system for irrigation water allocation along the middle reaches of the heihe river basin, northwest china. environmental modelling & software, 47, 182–192. copyright © 2020 assa. adv syst sci appl (2020) 164 a.a. voronin, a.v. khoperskov, i.i. isaeva, a.yu. klikunova 6. weng, s. q., huang, g. h., & li, y. p. (2010). an integrated scenario-based multicriteria decision support system for water resources management and planning — a case study in the haihe river basin. expert systems with applications, 37(12), 8242– 8254. 7. isaeva, i. i., voronin, a. a., khoperskov, a. v., dubinko, k. e., & klikunova, a. y. (2019). decision support system for the socio-economic development of the northern part of the volga-akhtuba floodplain. communications in computer and information science, 1083, 63–77. 8. misganaw, d., guo, y., knapp, h.v. & bhowmik, n.g. (1999). the illinois river decision support system (ilrdss), report prepared for the: illinois department of natural resources, 48. 9. fernandes, l. f. s., marques, m. j., oliveira, p. c., & moura, j. p. (2014). decision support systems in water resources in the demarcated region of douro-case study in pinhao river basin, portugal. water and environment journal, 28(3), 350–357. 10. wriggers, p., kultsova, m., kapysh, a., kultsov, a., & zhukova, i. (2014). intelligent decision support system for river floodplain management. communications in comp. and inform. sci, 466, 195–213. 11. mccord, j., carron, j.c., liu, b., rhoton, s., rocha, m. & stockton, t. (2004). pecos river decision support system: application for adjudication settlement and river operations eis. southern illinois university carbondale opensiuc, 95. 12. burkov, v. n., novikov, d. a. & shchepkin, a.v. (2015). simulation models for control mechanisms in ecological-economic systems studies in systems, decision and control, 10, 117–154. 13. khrapov, s. s., pisarev, a. v., kobelev, i. a., zhumaliev, a. g., agafonnikova, e. o., losev, a. g., & khoperskov, a. v. (2013). the numerical simulation of shallow water: estimation of the roughness coefficient on the flood stage. advances in mechanical engineering, 5, 787016. 14. gorelits, o. v., ermakova, g. s., & terskii, p. n. (2018). hydrological regime of the lower volga river under modern conditions. russian meteorology and hydrology, 43(10), 646–654. 15. agafonnikova, e. o., klikunova, a. yu. & khoperskov, a. v. (2017). a computer simulation of the volga river hydrological regime: a problem of water-retaining dam optimal location bulletin of the south ural state university, series: mathematical modelling, programming and computer software, 10(3), 148–155. 16. voronin a., isaeva i., khoperskov a. & grebenuk s. (2017). decision support system for urbanization of the northern part of the volga-akhtuba floodplain (russia) on the basis of interdisciplinary computer modeling. communications in computer and information science, 754, 419–429. 17. dyakonova t., khoperskov a. & khrapov s. (2016). numerical model of shallow water: the use of nvidia cuda graphics processors. communications in computer and information science, 687, 132–145. 18. khrapov, s. s. & khoperskov, a.v. (2020). application of graphics processing units for self-consistent modelling of shallow water dynamics and sediment transport lobachevskii journal of mathematics, 41(8), 1475–1484. 19. lebedeva, s. v., zemlyanov, i. v., & lomonosov, a. a. (2018). flow distribution, flooding and water engineering measures in the volga-akhtuba floodplain region: 2 modeling. russian meteorology and hydrology, 43(10), 655–663. 20. elizarova, t. g. e., & ivanov, a. v. (2018). regularized equations for numerical simulation of flows in the two-layer shallow water approximation. computational mathematics and mathematical physics, 58(5), 714–734. 21. bolgov, m. v., shatalova, k. yu., gorelits, o. v. & zemlyanov, i. v. (2017) ekosistemy: ekologiya i dinamika [water and ecology problems of the volga-akhtuba floodplain] , 1(3), 15–37, [in russian] copyright © 2020 assa. adv syst sci appl (2020) control model of the floodplain territories structure 165 22. voronin a., vasilchenko a. & khoperskov a. (2018). a project optimization for small watercourses restoration in the northern part of the volga-akhtuba floodplain by the geoinformation and hydrodynamic modeling. journal of physics: conference series, 973, 012064. 23. vasilchenko, a. a., voronin, a. a., dubinko, k. e., & isaeva, i. i. (2018). program complex for simulation modelling of hydrotechnical projects in floodplain terrains. mathematical physics and computer modeling, 21(2), 59–74. 24. vasilchenko, a., voronin, a., svetlov, a., antonyan, n. (2016). assessment of the impact of riverbeds depth in the northern part of the volga-akhtuba floodplain on the dynamics of its flooding. international journal of pure and applied mathematics, 110(1), 183–192. 25. voronin, a. a., eliseeva, m. v., khrapov, s. s., pisarev, a. v., & khoperskov, a. v. (2012). the regimen control task in the eco-economic system ”volzhskaya hydroelectric power station–the volga-akhtuba floodplain’”. ii. synthesis of control system. problemy upravleniya, 6, 19–25. copyright © 2020 assa. adv syst sci appl (2020) introduction research methods, tools and technologies tools and simulations results for volga-akhtuba floodplain conclusion advances in systems science and applications (2014) vol.14 no.3 199-229 the theory of parametric control of macroeconomic systems and its applications(iii) a. ashimov, zh. adilov, r. alshanov, yu. borovskiy and b. sultanov kazakh national technical university named after k. satpayev, 22 satpaev street, almaty, 050013, kazakhstan abstract this work consists of three parts and presents the recent results of development of the theory of parametric control of macroeconomic systems and some its applications for solving a number of concrete problems. keywords mathematical model, structural stability, parametrical identification, parametric control part 3. applications of the theory of parametric control of macroeconomic systems 3.1 macroeconomic analysis and parametric control of macroeconomic system based on an econometric model of small open economy 3.1.1 building an econometric model of a small open economy general view of the model for small open economy of kazakhstan, that describes equilibrium conditions in macroeconomic markets of goods, money, labor and of capital, taking into account its interaction with the russian federation and the rest of the world, is presented by the following relations [1-2]. equilibrium in the goods market of the republic of kazakhstan is presented by the formula : y d = y s (1) where y s real supply of goods in the republic of kazakhstan, in billions of tenge (national currency of the republic of kazakhstan) ; y d = c + i + g + nefull-real demand for goods in the republic of kazakhstan, in billions of tenge ; nefull = ne + neru real volume of goods net export from the republic of kazakhstan, in billions of tenge ; neru = qru ex − erq ru im real net export of goods from the republic of kazakhstan to the russian federation, in billions of tenge (indicator, that takes into account the terms of cooperation of the republic of kazakhstan within the regional customs union) ; qex real volume of exports of goods from the republic of kazakhstan to the russian federation, in billions of tenge ; qru im real volume of imports of goods to the republic of kazakhstan from russian federation, in billions of tenge ; er = epz/p real exchange rate in the republic of kazakhstan, tenge/us dollar ; e exchange rate of national currency in the republic of kazakshtan, tenge/us dollar ; pz the general price level in the outside world ; p the general price level in the republic of kazakstan ; ne = qw ex − erq w im real net export of goods from the republic 200 a. ashimov : the theory of parametric control of macroeconomic systems and ... of kazakshtan to the rest of the world, in billions of tenge (indicator that takes into account the terms of cooperation between the republic of kazakhstan and the rest of the world) ; qw ex real volume of exports of goods from the republic of kazakhstan to the rest of the world, in billions of tenge ; qw im real volume of imports of goods to the republic of kazakhstan from the rest of the world, in billions of us dollars ; g real volume of government expenditures in the republic of kazakhstan, in billions of tenge ; i real volume of the republic of kazakshtan investments to the basic capital, in billions of tenge ; c real volume of consumption by households in the republic of kazakhstan, in billions of tenge. all real data are presented for the year of 2000. equilibrium in the money market of the republic of kazakhstan is presented by the following relation : m/p = l (2) where l is real cash balances in the republic of kazakhstan, in billions of tenge ; m nominal money supply in the republic of kazakhstan (in billions of tenge). equilibrium in the labor market of the republic of kazakhstan : pdy/dn = w (3) where w nominal wage rate in the republic of kazakhstan, in thousands of tenge ; dy/dn marginal productivity of labor in the republic of kazakhstan ; y real gross domestic product (hereinafter gdp) in the republic of kazakhstan, in billions of tenge ; n number of employed people in the republic of kazakhstan, in thousands of persons. equilibrium in the capital market : pnefull = nke (4) where nke net export nominal volume from the republic of kazakhstan, in billions of tenge. let’s introduce additional notations for economic indicators used in development of a model : mru real money supply in russian federation (in billions of rubles) and gru real volume of government expenditures in russian federation (in billions of rubles), indicators allowing for conditions of operating of a country as part of regional customs union ; i the average interest rate of banks for loans in the republic of kazakhstan ; iz interest rate of the outside world (market yield on us, treasury securities at 1-year constant maturity, quoted on investment basis) ; pavg = 0.6p + 0.4epz/e2000 weighted average price level in the republic of kazakhstan ; expected exchange rate in the republic of kazakhstan (tenge/us dollar) ; êe = (ee− e)/e expected growth rate of the exchange rate in the republic of kazakhstan ; p oil average oil price (in thousands of tenge for a advances in systems science and applications (2014) vol.14 no.3 201 barrel) ; ∆ operator of first difference for the series : ∆x = x −x−1;x−1lag variable. preliminary econometric analysis showed the possibility for evaluation of macroeconomic indicators c, l, w, i, y, qru ex , qru im , qw ex , qw im, nke of equilibrium conditions in macroeconomic markets as a function of the regression on the basis of the following set of time series : y, c, l, i, nefull, w, n, pavg, i, qru ex , p oil, mru , gru , qru im , qw ex , er, qw im, nke, iz , êe for the years of 2000-2011 according to statistical data of national economies of the republic of kazakhstan and russian federation. in order to build non-spurious regression functions the considered sets were checked for stationary with the help of augmented dickey-fuller method (adf)and were decomposed [3]. according to the results of a check-up for stationarity of time series : regression functions of the following type c = c(y ), l = l(y, i,nefull), w = w (n,pavg), nke = nke(iz , êe, i), i = i(i), y = y (n), qru im = qru im (y,gru ), qw im = qw im(y, er) were built by the method of ordinary least squares on the basis of the following non-stationary series and the results of their analysis on the statistical significance are presented respectively in table 1 : table 1 regression functions for non-stationary series consumption of domestic products in the : c= 558.3 + 0.38y republic of kazakhstan(r2=0.99) (0.00) (0.00) demand for real cash balances in the l=0.7y -82.6 i -0.8 nefull republic of kazakhstan(r2=0.92) : (0.00) (0.02) (0.09) price of labor supply in the republic w =-0.15 n + 1174.8 pm+0.4ez/2000 of kazakhstan, where pm = 0.6p (r2 = 0.98) : (0.00) (0.00) net capital export from the republic of nke = 291iz − 6090ẽe − 25.6i kazakhstan (r2 = 0.65) : (0.03) (0.02) (0.35) investments of the republic of kazakhstan i = 5885.5-291.5 i to the basic capital (r2 = 0.33) : (0.01) (0.05) production function in the republic of y =-25255+4.27n kazakhstan (r2=0,95) : (0.00) (0.00) import of goods from the russian qru im = 0.07y + 0.89∆gru federation (r2=0.89) : (0.00) (0.00) import of goods from the rest of the world qw im = 0.22y + 3.75er (r2 = 0.79) : (0.00) (0.01) regression functions of the type qru ex = qru ex (er, p oil,mru , gru ), qw ex = qw ex(er, p oil) were built by the method of least squares on the basis of corresponding stationary series or stationary with respect to determined trends and the results of the analysis for statistical significance are presented respectively in table 2 : 202 a. ashimov : the theory of parametric control of macroeconomic systems and ... table 2 regression functions for stationary series import of goods from the republic of : qru ex = 4.4∆er + 31.4p oil kazakhstaninto the russian federation −0.066mru + 0.128gru (r2=0.72) : (0.03) (0.04) (0.02) (0.00) exportof goods from the republic of kazakhstan to the rest of the world qw ex = −11.2er + 278.0p oil + 1830.4 (r2=0.99) : (0.01) (0.00) (0.01) regressions functions (where corresponding time series are stationary with respect to the determined trends) were checked on spuriousness by the t-test [4]. spurious regressions functions (where time series were non-stationary) were checked on spuriousness by the engle-granger cointegration test [5]. the model of small open economy of the republic of kazakhstan based on the equilibrium conditions in macroeconomic markets of goods, money, labor and capital (1)-(4) and based on the built regression functions (tables 1, 2) has the following form :  y d = 3201.8 + 0.84 m p − 17.63 epz p + 4.01 e−1p z −1 p−1 + 281.8p oil − 0.06mru − 0.694gru + 0.811gru −1 y s = −25525 + 19967.4p + 13392.7pz e e2000 y zb0 = 6311.72 + 1066.9p oil − 66.72 epz p + 15.17 e−1p z −1 p−1 − 0.2276mru + 3.07gru −1 − 1003.45 iz p + 21000 ee e − 19369.8 1 p − 0.221 m p 2 − 1.04 epz p 2 + 0.205 e−1p z −1 pp−1 + 19.77 p oil p − 0.034 mru p − 0.04 gru p + 0.0478 gru −1 p y s = y d = y zb0 i = 18.5− 0.0025 m p − 0.0117 ep 2 p + 0.0023 e−1p z −1 pp−1 + 0.19p oil − 0.00003mru + 0.0005gru −1 (5) here y zb0 is the function of zero balance of payments of the republic of kazakhstan. advances in systems science and applications (2014) vol.14 no.3 203 3.1.2 estimating stability indicators of the model for a small open economy the quality of the researched econometric model of the small open economy is estimated by the stability indicator β (section 1.3.2), which characterizes a change in the equilibrium model solutions by small deviations of the input parameters used. if the stability indicator of the model takes a small value for a small deviation of input variables used, it is considered that the model is qualitative in the sense of stability indicator. estimation of stability indicators is made by the algorithm 5, where the vector x = {m,g,pz , iz , ee, p oil,mru} has been considered as the vector of input parameters, and the vector z = {y, e, i} has been considered as the vector of output variables. conducted computing experiments show that deviations of equilibrium solutions up to 1 % correspond to the deviations of input factors within 1%. this confirms the fact that the considered econometric model (5) is a qualitative model in sense of stability indicator β. 3.1.3 parametric control of the country’s export depending on uncontrollable factors based on the fact of dependence of the solution of algebraic equations on its coefficients, we propose an approach to parametric control of national economy evolution taking into account the requirements for equilibrium on macroeconomic markets, which comes down to making recommendations based on the optimal values of economic tools in the form of solutions of mathematical programming problems based on the econometric model of a small open economy. let us consider the possibility of estimating the optimal values of m and g tools of economic policy for the given values of the uncontrollable input parameters p oil, iz , pz , ee, mru and gru that represent the values of these factors in the framework of the customs union by example of one country and the rest of the world in 2011 within the framework of the model is-lmzb0 (built on statistics for 2000-2011) in sense of the maximum criterion : qex = qw ex +qru ex → max (6) here qex is the function of total exports of goods. the stated estimate can be obtained by solving the following problem of mathematical programming. problem 3.1. based on the mathematical model (5) find values (m, g), that provide maximum to the criterion (6) under the constraints (7) here m * and g* are accepted values of money supply and government expenditures respectively, for the years of 2008-2011 ; y ∗, p ∗, e∗, i∗ basic equilibrium solutions of the system (5) ; y , p , e, i optimal equilibrium solutions of the system (5). 204 a. ashimov : the theory of parametric control of macroeconomic systems and ...  |m −m∗| ≤ 0.1m∗, |g−g∗| ≤ 0.1g∗, |p − p ∗| ≤ 0.1p ∗, |e− e∗| ≤ 0.1e∗, |i− i∗| ≤ 0.1i∗, |y − y ∗| ≤ 0.1y ∗, (7) the proposed approach to the parametric control of national economy evolution consists in realization of the following algorithm : 1. choice of mathematical model based on statistical analysis of the regression functions and estimation of stability indicator of the econometric model of economic general equilibrium for the open economy of the republic of kazakhstan ; 2. statement of the mathematical programming problem ; 3. prediction of uncontrollable factors pz , iz , p oil, mru , gru and ee for the period of choosing the recommendations on economic policy ; 4. solution to the mathematical programming problem based on the selected mathematical model to forecast values of uncontrollable factors ; 5. making recommendations on values of the m and g tools based on the analysis of the results of the mathematical programming problem for predicted values of the uncontrollable factors and possible additional information on the economic conjuncture. below we present an illustration of the proposed approach of parametric control of the national economy evolution for 2012. 1. let the model of a small open economy be a mathematical model selected on the basis of estimation of stability indicators (less than 1%) in 2011. 2. as a statement of the optimization problem for the model of a small open economy in 2011 we take the statement of the problem 3.1. 3. forecasted values of uncontrollable factors, obtained on the basis of the models built taking into consideration the results of time series decomposition into components, took the following values for 2012 : pz = 1.32 ; p oil = 14.75 thousandtenge per barrel ; iz = 0.17% ; ee = 148.68 tenge for one us dollar ; mru = 27949.1 billion rubles and gru =10898.3 billion rubles. 4. solution to the mathematical programming problem on the basis of the model for a small open economy by example of the republic of kazakhstan and the predicted values of uncontrollable factors for 2012 are : m = 6733.0 billion tenge ; g = 1444.4 billion tenge ; the value of the criterion max(qex w + qex ru ) = 3465.1 + 335.6 = 3800.7 billion tenge ; 5. the following can be proposed as a recommendation : solutions obtained during the experiment m = 6733.0 billion tenge and g = 1444.4 billion tenge or advances in systems science and applications (2014) vol.14 no.3 205 some correcting values, those can be obtained on the basis of the additional data analysis on economic conjuncture. 3.2 macroeconomic analysis and parametric control of cyclical dynamics the major section of modern macroeconomic theory is propositions on market cycles, in which the factors generating them are considered and different mathematical models for their analysis are proposed [1,6-8]. suppression of market cycles is the major field of stabilization policy [9-10]. 3.2.1 macroeconomic analysis and parametric control of cyclical dynamics based on kondratiev cycle model model description this model combines descriptions of non-equilibrium economic growth and nonuniform scientific and technological advancement [11]. the model is described by the following system of equations, including two differential and one algebraic equation :  n(t) = ay(t)a, dx/dt = x(t)(x(t)− 1)(y0n0 − y(t)n(t)), dy/dt = n(t)(1− n(t))y(t)2(x(t)− 2 + µ+ l0 n0y0 ), n0 = ayα0 . (8) here t is the time (in months) ; x is the efficiency of innovations ; y is the capital productivity ratio ; y0 is the capital productivity ratio corresponding to the equilibrium trajectory ; n is the rate of savings ; n0 is the rate of saving corresponding to the equilibrium trajectory ; µ is the coefficient of withdrawal of funds ; l0 is the job growth rate corresponding to the equilibrium trajectory ; a and a are some model constants. estimation of the model parameters is carried out based on statistical information from the republic of kazakhstan for the years 2001-2005 [12]. the deviations in the observed statistical data and the calculated data do not exceed 1.9% within the considered period. as a result of solving the problem of parametric identification, the following values of the exogenous parameters are obtained : α = −0.0046235, y0 = 0.081173, n0 = 0.29317, µ = 0.00070886, l0 = 0.00032161, x(0) = 1.911144. a retrospective prediction for 2006 and 2007 are characterized by errors equal to 6.1% and 12.1%, respectively, for the capital productivity ratio, and 2.3% and 11%, respectively, for the rate of savings. the respective cyclic phase trajectory of the kondratiev cycle model is presented in fig.1. the period of cyclic trajectory corresponding to the statistical information of the republic of kazakhstan for the given years is estimated to be 206 a. ashimov : the theory of parametric control of macroeconomic systems and ... 232 months. fig.1 cyclic phase trajectory of the kondratiev cycle model fig.2 chain-recurrent set for the kondratiev cycle model estimating the robustness of the kondratiev cycle model without parametric control the estimation of structural stability (robustness) of the mathematical model advances in systems science and applications (2014) vol.14 no.3 207 is carried out according to the 4th component of the parametric control theory (section 1.1) in the chosen compact set of the model phase space. fig.2 presents an estimate of the chain-recurrent set r(f,n) obtained by the application of the chain-recurrent set estimation algorithm for the region n = [1.7; 2.3]×[0.066; 0.098] of the phase plane oxy of system (8). since the set r(f,n) is not empty, one can draw no conclusion about the weak structural stability of the kondratiev cycle model in n on the basis of robinson’s theorem. however, since there is a non-hyperbolic singular point in n , namely, the center (x0 = 2− µ+l0 n0y0 , y0), then system (8) is not weakly structurally stable in n. parametric control of the evolution of economic system based on the kondratiev cycle model choosing the optimal parametric control laws is carried out in the environment of the following four relations : 1) n0(t) = n0 ∗ + k1 y(t)− y(0) y(0) ; 2) n0(t) = n0 ∗ − k2 y(t)− y(0) y(0) ; 3) n0(t) = n0 ∗ + k3 x(t)− x(0) x(0) ; 4) n0(t) = n0 ∗ − k4 x(t)− x(0) x(0) . (9) here ki is the scenario coefficient ; n∗ 0 is the value of the exogenous parameter n0 obtained as a result of the estimation of parameters. the problem of choosing the optimal law of parametric control at the level of the econometric parameter n0 can be formulated as follows. on the basis of mathematical model (8), find the optimal parametric control law in the environment of the set of algorithms (9), ensuring reach of optimal values of the following criterion : k = 1 t t∑ t=1 (( x(t)− x0 x0 )2 + ( y(t)− y0 y0 )2 ) → min (10) (here t = 232 is the period of the cycle) under the constraints 0 ≤ y(t) ≤ 1, 0 ≤ n(t) ≤ 1, 0 ≤ x(t) (11) the base value of the criterion (without parametric control) is as follows : k = 0.0307. the value of criterion k = 0.007273 for the control law, that is optimal in 208 a. ashimov : the theory of parametric control of macroeconomic systems and ... fig.3 capital productivity ratio without parametric control and with use of law 4, optimal in the sense of criterion k fig.4 efficiency of innovations without parametric control and with use of law 4, optimal in the sense of criterion k the sense of the criterion (10) of the 4th law, from the set (9) represented before is obtained by solving the problem formulated above through application of the parametric control approach to the evolution of the economic system. corresponding value of adjustable coefficient of this law is . the values of the model’s endogenous variables without applying parametric control and with use of the optimal parametric control law for criterion k are presented below in graphic form (fig.3 and fig.4). estimating the structural stability of the kondratiev cycle mathematical model with parametric control advances in systems science and applications (2014) vol.14 no.3 209 to carry out this analysis, the expressions for optimal parametric control laws (11) with the obtained values of the adjustable coefficients are substituted into the right-hand side of the second and third equations of system (1) for the parameter n0. then, by using a numerical algorithm for estimating the weak structural stability of the discrete-time dynamical system for the chosen compact set n determined by the inequalities 1.7 ≤ x ≤ 2.3, in the state space of the variables (x, y), the estimation of the chain-recurrent set r(f,n) as the empty (or onepoint) set is obtained. this means that the kondratiev cycle mathematical model with optimal parametric control law is estimated as weakly structurally stable in the compact set n . analysis of the dependence of the optimal value of criterion k on the parameter for the variational calculus problem based on the kondratiev cycle mathematical model let us analyze the dependence of the optimal value of criterion k on the exogenous parameters µ (share of withdrawal of capital production assets per month) and a for parametric control laws (11) with the obtained optimal values of the adjusted coefficients ki, where the values of the parameters (µ, a) belong to the rectangle a = [0.00063; 0.00147]× [−0.01; 0.71] in the plane. plots of dependencies of the optimal value of criterion k (for parametric control laws 0 and 2, yielding the maximum criterion values) on the uncontrollable parameters (see fig.5) were obtained by computational experimentation. the projection of the intersection line of the two surfaces in the plane (µ, α) consists of the bifurcation points of the extremals of the given variational calculus problem. 3.2.2 macroeconomic analysis and parametric control of cyclical dynamics based on dynamic stochastic general equilibrium model for the economy of kazakhstan in nonlinear dynamical stochastic general equilibrium (dsge) model is presented on the base of given composition and behavior of agents, their inter action in stochastic conditions and of taking the principle of rational expectations [13]. this nonlinear dsge model of the economy consists of : first-order aggregate conditions of optimization problems of agents (house hold sand intermediate product producers) [14] ; description of government activity rules ; and rules of shocks specifying in terms of either first-order auto regression s or gaussian white noises. first-order aggregate conditions involve equilibrium conditions in market of labor, capital, intermediate and final goods. the nonlinear dsge model that was built involves both the model of actual economy, and the model of potential economy. the model of potential economy is similar to the model of actual economy by its composition, except that potential economy functions under flexible prices and wages (in the model of actual economy the se prices are not flexible), and also 210 a. ashimov : the theory of parametric control of macroeconomic systems and ... fig.5 plots of the dependencies of the optimal value of criterion k on exogenous parameters µ, a in absence of “extra charge” shocks (those are in the model of actual economy). the nonlinear dsge model in question has the following vector form : etf θ(xt−1, xt, xt+1,h σh t ) = 0 (12) here et is sign of conditional mathematical expectation given information available at the point of time t (t = 1, 2, ...) ; f θ is known vector function ; θ is parameters set, consisting of structural parameters of the model and auto regression parameters of shocks ; xt is vector, consisting of endogenous variables and shocks, determined by first-order auto regressions ; x0 is given ; hσh t is vector, consisting of gaussian white noises, σh is corresponding diagonal covariance matrix. according to the technique taken for dsge model, linear approximation of nonlinear dsge model (12) was built in neighborhood of its stationary point x [13]. mentioned point is found by solving vector equation (15) obtained from (14) by dropping time subscripts and nulling white noises : f θ(x,x,x, 0) = 0 (13) log-linearization of dsge model of f . s mets and r. wouters in neighborhood of its estimated stationary point gives linear dsge model of the following form : aθx̂t−1 +bθx̂t + cθetx̂t+1 +dθhσh t = 0, t = 1, 2, 3... (14) advances in systems science and applications (2014) vol.14 no.3 211 here the sign ≪∧≫ corresponds to linearized variable, aθ, bθ, cθ, dθ are matrices of corresponding dimensions. in this example we consider the case of implementing state economic policy by the taylor rule [15], describing behavior of national bank insetting interest rate, and rules for determining the size of government spending as well. in the framework of linear model (14) the taylor rule for determining governmental bonds yield is presented in the following form : r̂t = ρr̂t−1 + (1− ρ) ( π̄t + rπ(π̂t−1 − π̄t−1) + ry (ŷt − ŷ p t ) + r∆π(π̂t−1 − π̂t−1) + r∆y ( ŷt − ŷ p t − ( ŷt−1 − ŷ p t−1 )) + ηrt (15) and the rule of government spending in the following form : ĝt = ρgĝt−1 + ηgt (16) here r̂t, ŷt, ŷ p t , ĝt, π̂t are variables, corresponding to : governmental bondsyield (1 + interest rate), output, potential output, government spending and inflation. ηrt is interest rate shock, given in terms of gaussian white noise ; π̄t is inflation shock ; ηgt is government spending shock, ρ, rπ, r∆π, r∆y , ρg are the parameters of the equations (15), (16). estimating parameters of linear dsge model on the basis of statistical data of the economy of the republic of kazakhstan model (16) solution was obtained by the blanchard-kahn algorithm [16-17]. this solutionis presented in the form of first-order vector auto regression : x̂t r̂t ĝt  = qθ  x̂t−1 r̂t−1 ĝt−1 + f θhσh t , t = 1, 2, 3... (17) hereinafter x̂t is vector-column consisting of all endogenous variables in the model (including shocks determined in terms of auto regression), excluding state policy tools of governmental bonds yield r̂t, and the size of government spending ĝt. vectors x̂0, r̂0, ĝ0 are given ; qθ, f θ are matrices of corresponding dimensions. estimating parameters of the model in question (14) (using (17)) was made by the bayesian estimation method (the metropolis-hastings algorithm with the number of simulations of 4 000 000) using the kalman filter [18]. as observations were used quarterly data for seven macroeconomic indicators of kazakhstan (gdp, investments, consumption, employment, average wage, refinancing rate, and inflation) from 2002i till 2010iii. we found the logarithm of mentioned statistical indicators and linearly detrended them. for using the kalman filter within 212 a. ashimov : the theory of parametric control of macroeconomic systems and ... the bayesian approach the model (17) was supplemented with vector equation of the dimension : ŝt = m  x̂t r̂t ĝt  (18) here m is matrix, each row of which contains one unity, all of the rest its elements are equal to 0. as the results of measuring of observed variables were taken log-deviations from its linear trends of macroeconomic indicators values (consumption, investments, gdp, inflation, average wage, employment, refinancing rate), corresponding to observed variables. statistical data for the republic of kazakhstan from 2002i till 2011iii was used in this study. for using the bayesian approach there were given a priori density distribution p = p0(θ ′,σh) of parameters θ′,σh . the form and probabilistic characteristics of this distribution from were used in the research [13], with the exception of mathematical expectations of a priori distribution of parameters σh . mentioned mathematical expectations were increased 2.5 times relative to corresponding values from s mets f . and wouters r. in connection with large sampled standard deviations of economic indicators of kazakhstanin comparison with eurozone. according to the bayesian approach method [18], using like lihood function, obtained on the basis of the model(17), (18) using the kalman filter, and a priori distribution of parameters p0(θ′,σh) as well,was found posterior joint density distribution of initial estimates of parameters : p = p1(θ ′,σh). then using the metropolis-hastings algorithm with density was p1(θ ′,σh) generated a sample, consisting of 4 000 000 sets of parameters θ′,σh . finally, as required estimates of parameters were taken corresponding sampled averages. quality of applied method for finding the parameters estimates was tested by retro prognosis. for this purpose there were made predictions for mentioned observed economic indicators for four periods from 2010 iv till 2011 iii. root mean square deviations of obtained expected predicted values of economic indicators from corresponding statistical data were about 3%. analysis of shocks effectson gdp and inflation using estimating of impulse responses on disturbances within the framework of internal shocks of the economy of the republic of kazakhstan fig.6 and fig.7 present relatively impulse responses of real gdp and inflation on (unit positive) shocks. each diagram presented in figures is obtained by calculation of linear model(17) for initial zero values of all endogenous variables of the model and the value of chosen shock equal to its standard deviation for zero period. under this all of the values of this shock for non-zero time values, and also the values of all other shocks of the model for all of time values were taken as the null. advances in systems science and applications (2014) vol.14 no.3 213 fig.6 impulse responses of real gdp on singular positive shocks fig.7 impulse responses of inflation on singular positive shocks analysis of the diagrams of impulse response of real gdp presented in fig.6 shows the following : • given positive shocks of productivity (εat ), labor supply (εlt ) and investments (εit ) gdp increases. • positive shocks of preferences (εbt ) and government spending (ηgt ) also in214 a. ashimov : the theory of parametric control of macroeconomic systems and ... crease gdp (since these shocks increase,respectively, consumption and government spending) • positive shock of extra charge for goods (ηpt ) decreases gdp, and shock of extra charge for wage (ηwt ) increases gdp. • positive monetary shock (ηrt ) results in production decline (because of interest rate growth). analysis of the diagrams of impulse response of inflation presented in the fig.7 shows the following : • given positive shocks of extra charge for goods and extra charge for wage inflation increases • given positive shock of productivity inflation negligible decreases. the rest of shocks do not practically have an effect on inflation. these responses of the model to shocks correspond with theoretical propositions. decomposing indicators evolution to shocks effect parts in retrospective period indicators decomposition in retrospective period, presented in fig.8 and 9,illustrates the contribution (in percentage) of each shock effect of the estimated model (17) to deviations of actual values of gdp and inflation indicators from corresponding equilibrium values for the period 2002i-2012i. presented diagrams show how deviations of gdp and inflation from their corresponding trends in retrospective period (from 2002 i till 2012 iii) emerged according to positive and negative effects of the shocks in question. for instance, deviation of gdp from trend in 2009iiiequal to -4.93% is the sum of positive summands : 1. shock effects of labor supply (εlt ) equal to 3.02% deviation of gdp from trend, 2. shock effects of interest rates (ηrt ) equal to 1.51% deviation of gdp from trend, 3. shock effects of extra charge for goods (ηpt ) equal to 0.85% deviation of gdp from trend and negative summands : 4. shock effects of preferences (εbt ) equal to -3.88% deviation of gdp from trend, 5. shock effects of extra charge for wage (ηwt ) -3.63% deviation of gdp from trend, 6. shock effects of extra charge for capital (ηqt ) -1.16% deviation of gdp from trend, 7. shock effects of government spending (ηgt ) -1.12% deviation of gdp from trend, 8. shock effects of productivity (εat ) -0.69% deviation of gdp from trend (rest advances in systems science and applications (2014) vol.14 no.3 215 shock effects of investments (εit ) and shock effects of inflation (π̄t) are negligible). that is deviation of gdp from trend equal to -4.93% is the sum of all shock effects of the model for mentioned period. for instance, deviation of inflation from trend in 2007iv equal to 6.32% is the sum of only positive summands (all of the rest effects are negligible) : 1. shock effects of extra charge for goods equal to 4.87% deviation of inflation from trend ; 2. shock effects of extra charge for wage 1.43% deviation of inflation from trend. analysis of the diagram in fig.8 shows also that break-neck growth of gdp during the period 2004-2007 was mainly because of extra charge shocks (for wage, good, capital), lessening of gdp growth in 2008-2011 was mainly because of negative shock effects of preferences, extra charge for capital and wage. analysis of the diagram in fig.9 shows that inflation deviation from the equilibrium level in the period 2002-2011 was almost fully because of extra charge shock for good and wage. in other words, all of the rest shocks of the model do not practically have an effect on inflation values in mentioned period. fig.8 shocks effects on gdp rate fig.9 shocks effects on the quarterly inflation 216 a. ashimov : the theory of parametric control of macroeconomic systems and ... prediction of shocks effects on economic indicators and suppression of their effects based on dsge model of smets-wouters in the paper by estimated model (17) were obtained predicted values of macroeconomic indicators (gdp and inflation) for 1, 4, 10, 20, 30 and 40 quarters (i.e. correspondingly for 2011iv, 2012iii, 2014i, 2016iii, 2019i, 2021iii). for estimating shock effects on error variances of economic indicators predictions were used the standard technique for defining decomposition of error variance of predictions for the models of vector auto regressions [19]. the results obtained by the software dynare matlab toolbox [http ://www.dynare.org] are presented in tables 1 and 2. there are no shocks, in these tables, which effects on variance less than by 0.01%. analysis of the tables 3 and 4 shows the following. error variances of prognosis of gdp generally are determined by preference shocks, government spending shock, and extra charge shocks on the cost of capital. error variances of prognosis of inflation generally are determined by extra charge shocks on goods and extra charge shocks on wages. table 3 prognosis (billion tenge in average prices of 1994) and decomposition of variance of the quarterly gdp prognosis time prognosis decomposition of variance (in %) hori mathematical standard ηat ηbt ηgt ηlt ηit ηrt ηqt ηpt ηwt-zon expectation deviation 1 271.413 8.17 2.25 22.90 17.57 1.47 0.09 1.07 45.32 6.13 3.19 4 292.166 10.960 4.04 21.73 15.30 3.25 0.15 1.56 41.99 8.27 3.71 10 327.034 13.229 4.04 21.73 15.30 3.25 0.15 1.56 41.99 8.27 3.71 20 385.810 16.067 4.09 20.69 14.20 3.70 0.14 1.51 39.84 7.90 7.92 30 456.450 19.136 4.03 20.39 13.98 3.67 0.14 1.49 39.47 8.01 8.82 40 542.395 22.751 4.02 20.37 13.97 3.67 0.14 1.49 39.45 8.03 8.86 table 4 prognosis (in %) and decomposition of variance of the quarterly inflation prognosis time prognosis decomposition of variance (in %) horizon mathematical standard ηat ηqt ηpt ηwtexpectation deviation 1 1.22% 0.75% 0.29 0.01 92.24 7.47 4 1.36% 0.78% 0.30 0.06 72.31 27.33 10 1.62% 0.80% 0.28 0.15 62.65 36.92 20 1.78% 0.81% 0.27 0.16 60.71 38.85 30 1.83% 0.81% 0.28 0.17 60.55 39.00 40 1.85% 0.81% 0.28 0.17 60.53 39.01 advances in systems science and applications (2014) vol.14 no.3 217 in realization of the state policy for minimizing shock effects on economic indicators in the capacity of its tools we choose additive summands ηrt , η g t in the expressions (14), (15), desired values of which are searched in terms of deterministic values instead of respective shocks. the parametric control approach for minimizing shock effects consists in the following. let t be number of quarter, starting from which the state realizes parametric control policy for minimizing shock effects. at each time point t = t, t + 1, t + 2, ... on the basis of estimated model (17), written in the form x̂t+i r̂t+i ĝt+i  = qθ  x̂t+i−1 r̂t+i−1 ĝt+i−1 + f θhσh t+i , i = 1, 2, 3, ..., 40 (19) such deterministic values of tools ηrt+1, ..., ηrt+40, η g t+1,..., η g t+40 are defined, those give the minimum for criterion lt = et 40∑ i=1 βi(π̂2 t+i + λy ŷ 2 t+i),min(ηrt+1, ..., η r t+40, η g t+1, ..., η g t+40)lt (20) (characterizing expected discounted total deviation of gdp and inflation values from respective equilibrium values (trend)) under the following constraints on endogenous variables of the model. mathematical expectations for inflationand bond yield in this time horizon should not deviate from respective equilibrium values more than by 0.5% : |etπ̂t+i| ≤ 0.005, i = 1, 2, 3, ..., 40 (21) |etr̂t+i| ≤ 0.005, i = 1, 2, 3, ..., 40 (22) and mathematical expectation for the size of government spending by 5.0% from their trend values : |etĝt+i| ≤ 0.05, i = 1, 2, 3, ..., 40 (23) moreover, at each specified time point t, the value of current state of economy for the time point t−(x̂t, r̂t, ĝt) is known. after receiving a new information (the values of variables (x̂t+1, r̂t+1, ĝt+1) at the next time point (t + 1)), the values of tools ηrt+2, ..., ηrt+41, ηgt+2,..., ηgt+41 are calculated again by solving the problem (19)-(23) for respective period. here discount factor, is some weight coefficient. introduce a new minimization criterion : l̃t = 40∑ i=0 βi ( (etπ̂t+i) 2 + λy (etŷt+i) 2 ) ,min(ηrt+1, ..., η r t+40, η g t+1, ..., η g t+40)l̂t (24) 218 a. ashimov : the theory of parametric control of macroeconomic systems and ... it is not difficult to check, that this criterion differs from by the value, independent from variables ηrt+1, ..., η r t+40, η g t+1, ..., η g t+40. from the relation (19), by taking mathematical expectations for both of its parts, we get et  x̂t+i r̂t+i ĝt+i  = qθet  x̂t+i−1 r̂t+i−1 ĝt+i−1 + f θ′′ [ ηrt+i ηgt+i ] , i = 1, 2, 3, ..., 40 (25) here f θ ′′ is matrix, comprised by two corresponding columns of the matrix f θ consequently, optimal values of variables of the problems (19)-(23)and (21)-(25) coincide between each other. derived optimization problem (21)-(25) applies to classical (deterministic) type of variational calculus problems, which for each time point t is solved by the linear and quadratic programming method (with matlab application). below in the paper it is estimated the effectiveness of application of formulated above parametric control in assumption that the state will implement this policy during 30 years (120 periods). that is, assume that at each time point t (t = t, t +1, t +2, ..., t +119, t is number of quarter, corresponding to 2011iii) the stated etermines the values of parameters ηrt+1, η g t+1 by solving above mentioned optimization problem(19)-(23) (or, that is the same, (21)-(25)). in computing experiment it is assumed that the economy is precisely described by estimated model :  x̂t+i r̂t+i ĝt+i  = qθ  x̂t+i−1 r̂t+i−1 ĝt+i−1 + f θhσh t+i , i = 1, 2, 3, ..., 160 where x̂t is known. in the paper, for estimating the effectiveness of application of formulated above parametric control approach the monte-carlo method was used with estimate of 100 development scenarios of economy. let us present aggregative algorithm for estimating application of the parametric control approach. 1. generation of the sample, consisting of 100 elements-sets of values of vector gaussian random values (white noises) {h ′σh t+1,h ′σh t+2, ..., h ′σh t+120}j where j = 1, ..., 100, hσh t+i = [h ′σh t+i , η r t+i, η g t+i] t with known probabilistic characteristics of noises σh . 2.for each element of the sample {h ′σh t+1,h ′σh t+2, ...,h ′σh t+120}1, {h ′σh t+1,h ′σh t+2,..., h ′σh t+120}2,..., {h ′σh t+1,h ′σh t+2, ..., h ′σh t+120}120 : 2.1 the calculation of the model with parametric control : we solve optimization problem (21)-(25) for period t = t i.e. we find respective advances in systems science and applications (2014) vol.14 no.3 219 ηrt+1, ..., η r t+40, η g t+1, ..., η g t+40 ; from obtained set of valueswe take ηrt+1, η g t+1 (we drop rest values ηrt+2, ..., η r t+40, η g t+2, ..., η g t+40. 2.2. the model is calculated for 1 step with shocks values hσh t+1, which consist of tools values ηrt+1, η g t+1 determined by the state and shocks values h ′σh t+1,which were realized by economy independently (exogenously) from the state policy. x̂t+1 r̂t+1 ĝt+1  = qθ  x̂t r̂t ĝt + f θhσh t+1 2.3 the steps 2.1 and 2.2 are iteratedfor values t = t+1, t+2, t+3, ..., t+120. 3. on the bases of obtained 100 trajectories of gdp and inflation is built average trajectory (expected prognosis value) and standard deviations of prognosis for the period t = t, t +1, t +2, t +3, ..., t +120. obtained values are compared with basic prognosis. realization results of formulated algorithm show that for used sample of shocks the parametric control of suppressing shocks effects provides diminishing predicted standard deviations of gdp by 58.3% at the average in prognosis horizon from 2011iii till 2021iii (see fig.10). the parametric control of suppressing shocks effects provides diminishing predicted standard deviations of inflation by 32.0% at the average in prognosis horizon from 2011iii till 2021iii and diminishing samples tan dard deviation of inflation by 47.8% in comparison with actual data during the period2002itill 2011iii (see fig.11) 3.3 macroeconomic analysis and parametric control of the economic growth based on computable general equilibrium model for the economic sectors presentation of computable general equilibrium model non-autonomous computable general equilibrium model (cge model) in general form is presented by the following system of relations [2], [20]. 1) subsystem of differential equations, connecting endogenous variables values for two successive years : x1(t+ 1) = f1 (x1(t), x2(t), x3(t), µ(t), a(t)) (26) here t = 0, 1, ..., n−1 is number of year, discrete time ; x(t) = (x1(t), x2(t), x3(t)) ∈ rm is vector of endogenous variables of the system ; xi(t) ∈ xi(t) ⊂ rmi , i = 1, 2, 3 (27) here the variables x1(t) involve the values of capital assets of the sectors-producers, budgets of 220 a. ashimov : the theory of parametric control of macroeconomic systems and ... fig.10 prognostic values of real gdp for the basic scenario and the parametric control approach fig.11 prognostic values of inflation for the basic scenario and the parametric control approach (in %) economic agents and so on ; x2(t) involve the values of demand and supply of agents in different markets and so on ; x3(t) are different kinds of market prices and budget parts in markets with state-set prices for various economic agents ; m1 +m2 +m3 = m ; advances in systems science and applications (2014) vol.14 no.3 221 u(t) ∈ u(t) ⊂ rq is vector function of controllable (adjustable) parameters. coordinate values of this vector correspond to various state economic policy tools, for instance, such as state budget parts and budget parts of economic agents, various tax rates, governmental bonds yield and so on ; a(t) ∈ a ⊂ rs is vector function of uncontrollable parameters (factors). coordinate values of this vector characterize various external and internal social and economic factors depending on time : export and import goods prices, population size of the country, production functions parameters and so on ; x1(t), x2(t), x3(t), u(t) are compact sets with non-empty interiors ; xi =∪n t=1xi(t), i = 1, 2, 3 ; x = ∪3 t=1xi ; u = ∪n−1 t=0 u(t), is open connected set ; f1 : x × u ×a → rm1 , is continuous mapping. 2) subsystem of algebraic equations, describing behavior and interaction of agents in various markets within sampled year, these equations allow expressing the variables by exogenous parameters and rest endogenous parameters : x2(t+ 1) = f2 (x1(t), x3(t), x3(t), u(t), a(t)) (28) here f2 : x1 ×x3 × u ×a → rm2 is continuous mapping. 3) subsystem of recurrence relations for iterative calculations of equilibrium values of market prices in various markets and budget parts in markets with state-set prices for various economic agents : x3[q+ 1] = f3 (x2(t)[q], x3(t)[q], l, u(t), a(t)) (29) here q = 0, 1, ... is number of iteration ; l is the set of positive numbers (adjustable constants of iterations, when their values decrease,economic system comes faster to its equilibrium condition, however,at the same time the risk of the case when prices go to negative range increases ; f3 : x2×x3× (0,+∞)ms ×u ×a → rm2 is continuous mapping (that is compressing at fixed t ; x1(t) ∈ x1(t) ; u(t) ∈ u(t) ; a(t) ∈ a and some fixed l. in this case the mapping f3 has the only fixed point, to which converges the iterative process (28), (29). computable model (26), (28), (29) under fixed values of functions u(t) and a(t) for each time point t determine the value of exogenous variables x(t), corresponding to price equilibrium of demand and supply in the markets of goods and services of agents in the framework of the following algorithm. 1) in the first step it is assumed that t=0 and it is determined the initial values of variables x1(0). 2) in the second step for current the initial values of variables x3(0)[0] are determined in various markets and for various agents ; the values x2(t)[0] = f2 (x1(t), x3(t)[0], x3(t), µ(t), a(t)) (the initial values of demand and supply of agents in the markets of goods and services) are calculated by (28). 3) in the third step for current t it is run iterative process (28), (29). in this, 222 a. ashimov : the theory of parametric control of macroeconomic systems and ... for each value q current values of demand and supply are found from (29) : x2(t) = f2 (x1(t), x3(t)[q], x3(t), u(t), a(t)) by improvement of market prices and budget parts of economic agents. condition for stopping iterative process is equality of demand and supply values in various markets accurate within 0.01%. consequently,there are determined equilibrium values of market prices in each market and budget parts in markets with state-set prices for various economic agents. we omit the index q for such equilibrium values of endogenous variables. 4) in the following step on the basis of obtained equilibrium solution for the time point using differential equations (26) we define the values of variables x1(t+ 1). the value increases by unity. transition to the step 2. quantity of iterations of steps 2, 3, and 4 are determined in accordance with the parametric identification problems, prognosis and control in chosen in advance periods. considered cge model can be presented in the form of continuous mapping f : x × u × a → rm, determining transformation of the values of endogenous variables of the system for zero year to respective values of the next year according to presented above algorithm. here the compacts x(t) = x1(t)×x2(t)×x3(t), determining the compact x in the space of endogenous variables are defined by the set of possible values of variables x1 and respective equilibrium values of variables x2 and x3 calculated by relations (30)-(31). we will assume that for chosen point x1(0) ∈ int(x1) and corresponding, calculated by (28), (29) points x(0) = (x1(0), x2(0), x3(0)), the inclusion x(t) = f t(x(0)) ∈ int(x(t)) is true under some fixed u(t) ∈ int(u(t)), a(t) ∈ a for t = 0, ..., n. (n-fixed positive integer). this mapping f defines discrete dynamical system in the set x, on the trajectory of which imposed appropriate initial condition : {f t, t = 0, 1, ...}, x|t=0 = x0 (30) based on this conception specific cge model of economic sectors is considered below. parametric identification of cge model of economic sectors the model in question by statistical data of the republic of kazakhstan is presented by the following 19 economic agents. economic agent no.1. agriculture, hunt and forestry ; economic agent no.2. fishery, fish breeding ; economic agent no.3. mining ; economic agent no.4. manufacturing ; economic agent no.5. production and distribution of electricity, gas and water ; economic agent no.6. construction ; economic agent no.7. trade ; automobile and house articles maintenance ; advances in systems science and applications (2014) vol.14 no.3 223 economic agent no.8. hotels and restaurants ; economic agent no.9. transportation and communication ; economic agent no.10. financial activities ; economic agent no.11. transactions with real estates, lease and services to enterprises ; economic agent no.12. public administration ; economic agent no.13. education ; economic agent no.14. public health and social services ; economic agent no.15. other municipal, social and personal services ; economic agent no.16. housekeeping services ; economic agent no.17. aggregate consumer, combining households ; economic agent no.18. government, presented by the sum of central, regional and local governments, and non-budget funds as well. government determines tax rates and amount of subsidies for agents-producers and the size of social transfers for households. moreover, this sector includes non-commercial organizations, serving households (political parties, labor unions, social associations, etc.) ; economic agent no.19. bank sector, involving national bank and commercial banks. economic agent no.20. outside world. here,economic sectors no.1-16 are agents-producers. the considered model is presented as general expressions of relations of (26), (28), (29) respectively m1 = 67, m2 = 597, m3 = 34 by expressions, which help to calculate values of its 698 endogenous variables. this model contains also 2045 estimated exogenous parameters. in the result of combined solution of the problems a and b as per the constructed algorithm of parametric identification (section 1.2) using statistical data on evolution of the economy of the republic of kazakhstan. the relative value of deviations of estimated values of variables used mostly as criteria of corresponding observed values was equal to less than 0.63%. further the calculation of the model outside the period of parametric identification (forecast calculation) using extrapolated for the forecast period values of functions u(t), a(t) will be called as basic calculation. the results of calculation and retrospective basic calculation of the model for 2008 partially presented in the table 5 show estimated values, observed values and deviations of estimated values of main output variables of the model from corresponding observed values. here, the period 2000-2007corresponds to the period of parametric identification of the model ; 2008 is the retroprognosis period ; y is gross output (in prices of 2000) ; yg is gdp (in prices of 2000) ; p is consumer price index in percentages relative to the previous year ; sign ≪ ∗ ≫ corresponds to observed values, sign ≪ a ≫ corresponds to deviations (in percentages)of esti224 a. ashimov : the theory of parametric control of macroeconomic systems and ... table 5 observed, calculated values of output variables of the model and corresponding deviations indicator year 2000 2001 2002 2003 2004 2005 2006 2007 2008 y ∗(t) 5.44 6.32 6.47 6.86 7.72 8.52 9.25 9.69 9.84 y (t) 5.38 6.32 6.47 6.86 7.72 8.52 9.27 9.64 9.82 ∆y (t) -1.22 -0.02 0.00 0.00 0.05 0.08 0.21 -0.51 -0.26 y ∗ g (t) 2.45 2.78 3.05 3.36 3.72 4.09 4.55 5.01 5.18 yg(t) 2.47 2.78 3.05 3.35 3.72 4.09 4.55 5.01 5.20 ∆yg(t) 0.88 0.07 -0.04 -0.02 -0.02 -0.02 -0.04 -0.15 0.38 p ∗(t) 106.4 106.6 106.8 106.7 107.5 108.4 118.8 109.5 p (t) 107.6 106.8 106.9 106.7 107.3 108.2 118.6 109.4 ∆p (t) 1.13 0.18 0.08 -0.05 -0.23 -0.22 -0.24 -0.05 mated values from respective observed values. analysis of sources for economic growth based on computable general equilibrium model of economic sectors now, let us analyze sources for economic growth of economic sectors based on cge model of economic sectors and based on retrospective data for 2000-2009. for this, using expressions of production functions of economic agents of the model, we estimate the effect of change of arguments of these functions on the rates of growth of gav of sectors v y _i[t + 1] in assumption about constancy of coefficients ca_z_ji[t] under consumed by sector intermediate products v d_pj_iz[t], coefficients ca_k_i[t] under capital (v k_i[t] + v k_i[t + 1])/2 and coefficients ca_l_i[t] under labor v d_pi_il[t]. here i, j = 1, ..., 16 are numbers of economic agents ; t is number of year ; power (x,y ) corresponds to xy , exp(x) corresponds to ex ; ca_r_i is coefficient, characterizing technical progress in i-th sector. v y _i[t+ 1] = ca_r_i× exp(v d_p1_iz[t]× ca_z_1i)× exp(v d_p2_iz[t] × ca_z_2i])× exp(v d_p3_iz[t]× ca_z_3i)× exp(v d_p4_iz[t]× ca_z_4i) × exp(v d_p5_iz[t]× ca_z_5i)× exp(v d_p6_iz[t]× ca_z_6i) × exp(v d_p7_iz[t]× ca_z_7i)× exp(v d_p8_iz[t]× ca_z_8i) × exp(v d_p9_iz[t]× ca_z_9i)× exp(v d_p10_iz[t]× ca_z_10i) × exp(v d_p11_iz[t]× ca_z_11i)× exp(v d_p12_iz[t]× ca_z_12i) × exp(v d_p13_iz[t]× ca_z_13i)× exp(v d_p14_iz[t]× ca_z_14i) × exp(v d_p15_iz[t]× ca_z_15i)× exp(v d_p16_iz[t]× ca_z_16i) × power(((v k_i[t] + v k_i[t+ 1])/2), ca_k_i)× × power(v d_pi_il[t], ca_l_i) (31) advances in systems science and applications (2014) vol.14 no.3 225 after finding the logarithms of both parts (31) and total increment of the function ln (v y _i) and dropping members of the highest order infinitesimal, we get the following estimate of the growth rate yi of real gav of i-th sector depending on growth of arguments of production function : ca_r_i, v d_pj_iz, kim = (v k_i[t] + v k_i[t+ 1])/2, and v d_pi_il[t]. yi = ∆v y _i v y _i = ∆ca_r_i ca_r_i + 16∑ j=1 (ca_z_ji× v d_pj_iz) ∆v d_pj_iz v d_pj_iz + ca_k_i ∆kim kim + ca_l_i ∆v d_pi_il v d_pi_il (32) denote by ai = ∆ca_r_i ca_r_i the rate of technical progress in i-th sector ; zij = ∆v d_pj_iz v d_pj_iz is the rate of intermediate products consumed by i-th sector, and produced by j-th sector ; ki = ∆kim kim is the rate of capital accumulation in i-th sector ; li = ∆v d_pi_il v d_pi_il is the growth rate of labor inputs in i−th sector, where the sign ”∆” means change of variable ; time values in (32) were dropped for short. coefficients on the right-hand-side of formula (32) at stated above rates and characterize degree of effect of the factors in question on economic growth and allow comparing their effect with effect of technical progress, coefficient at which is equal to 1. by denoting these coefficients by aij = ca_z_ji × v d_pj_iz, βi = ca_k_i, γi = ca_l_i, from (32) we get its reduced writing : yi = ai + 16∑ j=1 αijzij + βiki + γili (33) let us present the values of coefficients, defining contributions of sources of economic growth of sectors based on the model in question for 2008. (see table 6). coefficients in the table show how much the growth rate of gav would increase given the one percent increase of growth factors (fixed assets, labor or demand for intermediate goods of economic agents). analysis of coefficients table βi, γi, αs i = ∑16 j=1 αij of the table 4 shows that if we exclude the rate of technical progress, effect of which on the rate of growth of all sectors in given model is the same, then from rest three factor rates of economic growth, the largest effect on the rate of real output of sectors 1, 5, 7, 12, 13, 16 of the economy has the rate of labor inputs ; for sectors 4, 6, 8, 10, 15-the rate of capital accumulation ; and for other sectors 2, 3, 9, 11, 14 -the rate of consumed by the sector intermediate products, produced by all of the sectors. also note that for sectors 5, 7 12, 16 the rates of capital accumulation practically do not have effect on corresponding rate of output growth ; the rates of labor inputs haven on-zero effect on the rates of output growth of all sectors ; for sectors 226 a. ashimov : the theory of parametric control of macroeconomic systems and ... 6, 16 the rates of consumed intermediate good shave no effect on corresponding rate of output growth. the results of analysis allow choosing the following budget parts of 16 economic sectors in the capacity of tools for solving the problems of economic growth. oij-budget part of i-th sector, which is for payment for goods and services, bought from j-th sector ; ol i-budget part of i-th sector, which is for payment for labor ; on i -budget part of i-th sector, which is for payment for investment goods. table 6 coefficients, characterizing an impact of economic growth factors number of βi γi αi1 αi2 αi3 αi4 αi5 αi6 αi7sector i 1 0.3089 0.9051 1.345· 9.480· 2.897· 2.171· 1.602· 2.028· 1.345· 10−18 10−02 10−14 10−13 10−14 10−14 10−12 2 0.2426 2.4964 1.590· 7.308· 7.087· 9.884· 1.390· 1.120· 1.590· 10−16 10−01 10−16 10−15 10−15 10−15 10−16 3 0.9650 0.6886 2.886· 0.000 2.970· 8.269· 9.894· 1.478 2.886· 10−15 10−12 10−13 10−14 10−15 4 1.2900 0.0805 2.227· 1.343· 8.634· 9.989· 9.029· 2.641· 2.227· 10−13 10−1 10−13 10−13 10−03 10−16 10−13 5 1.0· 2.4083 3.199· 5.537· 8.913· 6.720· 5.406· 5.300· 3.199· 10−10 10−17 10−17 10−14 10−4 10−3 10−4 10−17 6 0.9343 0.7721 2.078· 9.186· 3.313· 3.940· 2.467· 1.324· 2.078· 10−16 10−18 10−14 10−13 10−15 10−14 10−16 7 1.0· 1.8792 8.133· 1.061· 3.115· 2.660· 4.124· 2.508· 8.133· 10−10 10−16 10−15 10−14 10−13 10−14 10−13 10−16 8 1.0691 0.4706 2.003· 1.498· 4.343· 6.171· 3.375· 7.288· 2.003· 10−02 10−15 10−17 10−14 10−15 10−15 10−02 9 0.8660 0.2153 4.289· 2.265· 1.666· 8.753· 3.057· 4.005· 4.289· 10−16 10−13 10−15 10−14 10−14 10−16 10 0.6702 0.5492 0.000 0.000 0.000 4.397· 5.602· 8.811· 0.000 10−15 10−15 10−15 11 1.2022 0.1006 2.929· 2.267 5.064· 1.030· 1.664· 5.308· 2.929· 10−05 10−14 10−12 10−04 10−13 10−05 12 1.0· 2.5822 5.415· 0.000 0.000 4.255· 2.573· 0.000 5.415· 10−10 10−14 10−13 10−14 10−14 13 0.2635 1.7177 2.370· 0.000 0.000 8.847· 1.142· 5.739· 2.370· 10−14 10−13 10−13 10−15 10−14 14 0.0227 1.7814 6.736· 2.018· 6.280· 1.288· 1.919· 4.809· 6.736· 10−14 10−1 10−15 10−12 10−13 10−13 10−14 15 0.9304 0.2173 1.041· 1.408· 7.681· 3.712· 3.926· 1.579· 1.041· 10−15 10−3 10−15 10−13 10−14 10−13 10−15 16 0 1.9372 0 0 0 0 0 0 0 continuation of table 6. advances in systems science and applications (2014) vol.14 no.3 227 number of αi8 αi9 αi10 αi11 αi12 αi13 αi14 αi15 αi16sector i 1 4.325· 4.595· 1.807· 1.177· 0 1.681· 1.178· 9.304· 0 10−15 10−2 10−14 10−14 10−16 10−15 10−18 2 8.054· 1.087· 2.616· 2.280 0 0· 0 0 0 10−17 10−14 10−17 3 8.054 2.221· 1.436· 4.626· 0 1.581· 4.991· 1.571· 0 10−14 10−2 10−13 10−12 10−15 10−4 10−14 4 8.038· 1.444· 4.353· 6.338· 0 1.247· 3.379· 5.445· 0 10−15 10−13 10−14 10−16 10−16 10−03 10−16 5 1.642· 2.760· 6.132· 2.512· 0 3.273· 2.026· 5.494· 0 10−15 10−1 10−3 10−14 10−5 10−17 10−4 6 7.977· 1.117· 6.525· 2.014· 0 8.436· 1.554· 1.089· 0 10−15 10−2 10−16 10−14 10−17 10−17 10−16 7 6.067· 4.548· 3.407· 1.505· 0 5.394· 8.576· 4.404· 0 10−14 10−1 10−13 10−12 10−16 10−5 10−16 8 1.528· 8.267· 3.053· 1.507· 0 2.339· 1.338· 2.244· 0 10−15 10−15 10−14 10−14 10−17 10−17 10−16 9 3.346· 1.002· 3.203· 4.339· 0 3.657· 4.803· 1.091· 0 10−14 10−12 10−13 10−13 10−15 10−16 10−15 10 1.991· 3.270· 2.862· 5.408· 0 1.347· 0 8.805· 0 10−14 10−01 10−13 10−15 10−14 10−18 11 6.664· 6.596· 2.965· 2.646· 0 9.260· 1.661· 9.599· 0 10−14 10−13 10−13 10−12 10−14 10−03 10−06 12 4.185· 1.548· 9.506· 7.964· 0 2.126· 1.408· 2.886· 0 10−1 10−1 10−13 10−13 10−14 10−15 10−15 0 13 4.820· 5.807· 1.794· 1.331· 0 3.939· 7.738· 1.836· 0 10−14 10−1 10−14 10−14 10−15 10−17 10−14 14 1.891 2.118· 4.749· 2.743· 0 2.219· 1.782· 6.646· 0 10−13 10−17 10−13 10−16 10−14 10−14 15 6.285· 2.139· 9.739· 8.436· 0 9.538· 1.416· 1.222· 0 10−14 10−1 10−14 10−14 10−17 10−17 10−13 16 0 0 0 0 0 0 0 0 0 finding the optimal parametric control laws on the basis of cge model of economic sectors in computational experiments with cge model of economic sectors,the criterion below was used as maximization criterion k = 1 6 2015∑ t=2010 v y [t] (34) here k is average value of country’s gross output for 2010-2015 in prices of 2000. in experiments with optimization criterion (34), the constraints on growth of consumer price index of the following form were used : v pr[t] ≤ 1.0 228 a. ashimov: the theory of parametric control of macroeconomic systems and ... here is calculated consumer price index of the model without parametric control, is consumer price index with parametric control. in computational experiments control was performed for 1536 exogenous parameters j -th agent-producer’s budget parts for purchasing goods and services, produced by i -th agent-producer for 2000-2015 : oj i [t] ; t=2010,...,2015 ; i,j=1,...,16. here ∑16 i=1o j i (t) ≤ 1 for mentioned values of t. basic values of mentioned parts, obtained by solving the parametric identification problem of the model on data of 2000-2008, we will denote by ōj i ; i,j=1,...,16. the following problem of finding optimal values of adjustable parameters vectors was considered. on the basis of cge model of economic sectors to find mentioned values of budget parts of agents-producers oj i [t], which provide the maximum of criterion k under additional constraints on these parts of the following form : 0.5 ≤ oj i [t]/ō j i ≤ 2; i, j = 1, ..., 16; t = 2010, ..., 2015. solutions of these optimization problems were made using the nedler-mead algorithm. after using parametric control of budget parts of the model, criterion value turned out k = 1.6283 ·1013, criterion value increased by 33.14% in relation to the basic variant. references [1] tarasevich l.s., grebennikovp.i., leusskiy a.i. (2006), “macroeconomics”, moscow: vysshee obrazovanie, in russian. [2] ashimov a.a., sultanov b.t., adilov zh.m., borovskiy yu.v., novikov d.a., alshanov r.a. ashimov as.a. (2013), “ macroeconomic analysis and parametrical control of a national economy”, new york: springer. [3] e.s. said, d.a. dickey. (1984), “testing for unit roots in autoregressivemoving average models of unknown order”, biometrika, vol.71, no.3, pp.599607. [4] r.c. fair. (2004), “estimating how the macroeconomy works”, cambridge, mass.: harvard university press. [5] r.f. engle, c.w.j. granger. (1987) “co-integration and error correction: representation, estimation, and testing”, econometric, vol.55, no.2, pp.251276. [6] turnovsky s.j. (2000), “methods of macroeconomic dynamic”, cambridge: mit press. [7] kolemaev v.a.matematnqeskaya. (2002), “konomnka”, moskva: yuniti. [8] kydland f.e., prescott e.c. (1977), rules rather than discretion: the inconsistency of optimal planes, springer, journal of political economy, vol.87. advances in systems science and applications (2014) vol.14 no.3 229 [9] dornbusch r., fischer s. (1990), “macroeconomics”, new york: mcgraw-hill. [10] tumanova e.s., shagas n.a. (2004), “macroeconomics”, moscow: infra-m. [11] dubovskiy s.v. (1995), “the kondratiev cycle as the simulation object”, mathematical modelling, vol.7, no.6, pp.65-74, russian. [12] k.s. abdiev. (2001-2009), “statistical yearbook of kazakhstan”, astana: agency on statistics of the republic of kazakhstan. [13] smets f., wouters r (2003), “an estimated dynamic stochastic general equilibrium model of the euro area”, journal of the european economic association, vol.1, no.5, pp.1123-1175. [14] sims c. (2002), “random lagrange multipliers and transversality”, eco 504, spring, pp.1-7. [15] orphanides a. (2007) , taylor rules, federal eeserve board, working paper, no.18, washington, d.c.. [16] fernández-de-córdoba g., torres j.l. (2009), “forecasting the spanish economy with an augmented var-dsge model”, malaga economic theory research center working papers, no.1. [17] blanchard o.j., kahn c.m. (1980), “the solution of linear difference models under rational expectation, econometrica”, comput. math. math. phys, vol.24, pp.1812. [18] fernández-villaverde j. (2010), “the econometrics of dsge models, series”, vol.48, no.5, pp.1305-1312. [19] hamilton j.d. (1994), “time series analysis”, new jersey: princeton. [20] makarov v.l., bahtizin a.r., and sulakshin s.s. (2007), “application of computable models in state administration”, moscow: nauchniy expert. corresponding author a. ashimov can be contacted at: ashimov37@mail.ru advances in systems science and application(2016) vol.16 no.3 33-51 depth reduction factor assessment for evaluation of cyclic stress ratio based on site response analysis farzad farrokhzad department of civil engineering, babol university of technology abstract earthquake properties that affect the liquefaction of a soil are described with one parameter known as the cyclic stress ratio (csr). the stress reduction coefficient parameter accounts for the flexibility of the soil profile. the previous proposed equations could be used in routine engineering practice so long as the simplified (rd) was used to assess csr, but could be unconservative if used in sites with complex soil layers or in design of vital facilities such as dams, hospitals, bridges and . . . . this research presents the results of studies in babol city to develop, depth reduction factor (rd) based on equivalent linear analysis of study area for evaluation of cyclic stress ratio. as a part of microzonation study for the babol city a total of 35 boreholes have been drilled in 35 km2 of the research area. the depths of these boreholes ranged about 25 to 35 m. spt blow counts were taken in each 2 m depth. many geophysical investigations, generally 35 downhole logging surveys, are carried out in 35 mentioned boreholes for generation and measurement of shear wave velocity. based on these analyses, the (rd) diagrams and equations are developed for each mentioned zone. keywords depth reduction factor (rd), liquefaction, site response analysis, cyclic shear stress, earthquake return period, microzonation 1 introduction in geology, a fault is a planar fracture or discontinuity in a volume of rock, across which there has been significant displacement along the fractures as a result of earth movement [1]. an earthquake is caused by a sudden slip on a fault [2]. stresses in the earth’s outer layer push the sides of the fault together [3]. stress builds up and the rocks slips suddenly, releasing energy in waves that travel through the earth’s crust and cause the shaking that we feel during an earthquake [4]. the site safety during earthquakes is related with geotechnical phenomena such as amplification, liquefaction, landsliding and fault movements [5]. loose sand and silt that is saturated with water can behave like a liquid when shaken by an earthquake [6]. soil liquefaction describes a phenomenon whereby a saturated or partially saturated soil substantially loses strength and stiffness in response to an applied stress, usually earthquake shaking or other sudden change in stress condition, causing it to behave like a liquid [7]. evaluation of liquefaction potential requires comparison of the anticipated level of loading imposed on a 34 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting soil profile with the inherent resistance of the soil profile to liquefaction [8]. that procedure essentially compares the cyclic resistance ratio (crr) [the cyclic stress ratio required to induce liquefaction for a cohesionless soil stratum at a given depth] with the earthquake-induced cyclic stress ratio (csr) at that depth from a specified design earthquake [defined by a peak ground surface acceleration and an associated earthquake moment magnitude] [9]. during an earthquake, the soils will be subject to cyclic shear stresses induced by the ground shaking. the average cyclic stress ratio (csr) during an earthquake may be estimated by the following equation 1 [10-14]: csr = τave σ0′ = 0.65( amax g )( σ0 σ0′ ).rd (1) where amax = maximum acceleration at the ground surface, σ0 = total overburden pressure at depth under consideration, σ0 ′ = effective overburden pressure at depth under consideration and rd = stress reduction coefficient. seed and idriss considered a soil column as a rigid body [15]. in reality, soil behaves as a deformable body instead of as a rigid body. hence, the rigid body shear stress should be reduced with a correction factor to give the deformable body shear stress (τmax)d [16-18]. this correction factor is called the stress reduction coefficient (rd) and can be computed as follows: values of rd are commonly estimated from the diagram which is proposed by seed and idriss in 1971 (fig. 1). this chart was determined analytically using a variety of earthquake motions and soil conditions. average rd values in the diagram can be estimated using the following functions (equation 2) [19]. rd = 1− 0.00765z for z ≤ 9.15(m) rd = 1.174− 0.0267z for 9.15(m) < z ≤ 23(m) rd = 0.774− 0.008z for 23(m) < z ≤ 30(m) (2) also, equation 3 shows the suggestion of blake [20] rd = 1− 0.4113z0.5 + 0.04052z + 0.001753z1.5 1− 0.4177z0.5 + 0.05729z − 0.006205z1.5 + 0.00121z2 (3) regarding to the research of golesorkhi , statistically based rd curves were developed for different earthquake magnitude ranges [21]. the initial study of golesorkhi has been further extended by idriss and golesorkhi. advances in systems science and application(2016) vol.16 no.3 35 fig. 1 rd results from response analyses for 2153 combinations of site conditions and ground motions, superimposed with heavier lines (seed and idriss 1971) the proposed correlation for estimation of rd as a function of depth, magnitude, intensity of shaking, and site stiffness is presented in equation 4. in(rd) = α(z) + β(z).mw α(z) = −1.012− 1.126.sin( z 38.5 + 5.133) β(z) = 0.106 + 0.118.sin( z 37.0 + 5.142) (z = depth in feet) (4) rd(d,mw , amax, v ∗ s,12m) = for d < 65(ft)[ 1 + −23.013−2.949amax+0.999mw+0.016v ∗ s,40′ 16.258+0.201e 0.104(−d+0.0785v ∗ s,40′ +24.888) ] [ 1 + −23.013−2.949amax+0.999mw+0.016v ∗ s,12m 16.258+0.201e 0.104(0.0785v ∗ s,40′ +24.888) ] ± σεrd rd(d,mw , amax, v ∗ s,12m) = for d ≥ 65(ft)[ 1 + −23.013−2.949amax+0.999mw+0.016v ∗ s,40′ 16.258+0.201e 0.104(−d+0.0785v ∗ s,40′ +24.888) ] [ 1 + −23.013−2.949amax+0.999mw+0.016v ∗ s,12m 16.258+0.201e 0.104(0.0785v ∗ s,40′ +24.888) ] − 0.0014(d− 65)± σεrd σεrd = { 0.0072d0.850 d < 40ft 0.007240d0.850 d ≥ 40ft (5) 36 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting cetin et al recognized that rd is a function of site response, and developed a new correlation for estimation of rd. the proposed correlation is presented in equation 5 [22]. numerous seismic events around study area, such as manjil-rudbar earthquake on june 21, 1990, mazandaran earthquake on may 28, 2004, tabas earthquake on september 16, 1978 and bam earthquake on december 26, 2003 have demonstrated the relevance of subsurface soil layers and geotechnical conditions on seismic ground response [23-25]. equivalent-linear ground response modeling is by far the most commonly utilized procedure in practice [26]. equivalent-linear soil material modeling is widely used in practice to simulate true nonlinear soil behavior for applications such as ground response analyses [27-28]. the advantages of equivalent-linear modeling include small computational effort and few input parameters. in equivalent linear model, the soil stiffness g is modified in response to computed strains and g and ξ degradation curves. ground motion is computed for selected g and ξ pair at each layer; in particular, strain histories are calculated; from above effective shear strain γeff is calculated; from obtained effective shear strain, new pair of g(γ) and ξ(γ) are selected using the degradation curves; these steps are repeated until the maximum difference between computed shear modulus and damping ratio values in two successive iterations be less than ∼5% [29-31]. in the research of ishibashi and zhang, equivalent shear moduli and damping ratios for sandy soils were collected and equation 6 was proposed to best fit data points [32]. g gmax = k(γ, ip)σ̄ m(γ,ip)−m0 0 k(γ.ip) = 0.5 { 1 + tanh [ ln( 0.000102+n(ip) γ )0.492 ]} m(γ, ip)−m0 = 0.272 { 1− tanh [ ln(0.000556γ )0.4 ]} e−0.0145i13p n(ip) =  0 for ip = 0 sandy soils 3.37× 10−6i1.404p for 0 < ip ≤ 15 low plastic soils 7× 10−7i1.976p for 15 < ip ≤ 70 medium plastic soils 2.7× 10−5i1.115p for ip > 70 high plastic soils (6) 2 study area babol is located in the in the north of iran, between the northern slopes of the alborz mountains and approximately 20 kilometers south of caspian sea on the west bank of babolrud river and receives abundant annual rainfall. iranian tectonic plate in middle east affects the seism tectonic condition of babol (figure 2) [33-34]. the tectonic environment near babol city is unusually complicated. khazar and north alborz faults are the most significant faults around the study advances in systems science and application(2016) vol.16 no.3 37 area (figure 3). these faults are directed e-ne and w-sw. the khazar fault is the boundary between the caspian plain and alborz mountain (figure 4, figure 5). table 1 presents a list of active faults affecting the babol city [35]. regarding to seismicity of mazandaran, it can be assumed that for large earthquakes, the faulting process primarily involves repeated breaking of the same fault segment rather than creation of a new fault surface. table 2 shows recent major earthquakes that are felt in the zone (648584 e, 4049660 n) and (652032 e, 4043901 n). this zone closely bounds babol city. also, magnitude distribution of earthquakes in the range of 200 km around the babol city are presented in figure 6. fig. 2 tectonic of study area fig. 3 alborz and khazar faults fig. 4 surface section of khazar fault zone fig. 5 satelite map of north alborz fault 38 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting table 1 active faults around the study area (babol city) fault name type of fault distance from fault length babol city (km) (km) north alborz thrust fault 44 300 khazar thrust fault 16 550 firouzabad thrust fault 85 112 atari thrust fault 91 85 astaneh thrust fault 93 75 kandovan thrust fault 100 64 mosha thrust fault 91 400 north of tehran thrust fault 115 108 firouzkooh thrust fault 84 40 bayejan thrust fault 60 45 damghan thrust fault 136 100 orim thrust fault 72 44 table 2 recent major earthquakes around the study area (babol city)) earthquake magnitude depth (km) year chalus, mazandaran 6.3 17 2004 babol, mazandaran 5.1 15 2012 damghan, semnan 5.7 7 2010 roudbar and manjil 7.4 10 1990 (a) earthquakes of babol from 800 to 1900 (b) earthquakes of babol from 1900 to 2014 fig. 6 magnitude distribution of earthquakes in the range of 200 km around the babol city babolrud river originates in the alborz mountains and is one of the major rivers in iran. it is located on the left side of babol city. the study area is conadvances in systems science and application(2016) vol.16 no.3 39 stantly filled with the new alluvial sediments of the babolrud river. following paragraphs provide an overview of the site investigation conducted at 35 km2 of babol city and present the results of both geotechnical and geophysical investigations. also, supplemental activities related to babol seismic microzonation project are illustrated too. regarding to mentioned project, a total of 35 new boreholes have been drilled in study area. also, the results of previous investigations (consist of 60 borholes) were collected too (figure 7). the depth of boreholes ranged from 25 to 40 m. downhole logging surveys and spt tests carried out through new 35 boreholes, were performed at every 1.5 m and 2 m of depth. the location of drilled boreholes over babol city, a sample of geotechnical test results and shear wave velocity profile obtained by downhole test are shown in figure 8 and figure 9. the general topographic gradient in the study area is constant. during the investigation, the site soils were observed to consist of silty clay, silty sand and sandy clay. the depth of the water table as measured during drilling should be fig. 7 location of drilled boreholes over babol city 40 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting (a) sample of borehole log report (0-15 m) (b) sample of borehole log report (15-30 m) fig. 8 results of laboratory and in-situ tests (geotechnical investigations) advances in systems science and application(2016) vol.16 no.3 41 (a) sample of shear wave velocity profile (b) sample of s wave data obtained during downhole test fig. 9 results of geophysical (downhole) surveys fig. 10 location of drilled boreholes over babol city 42 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting carefully evaluated. it is always necessary to wait for at least 24 hours to check on the stabilized water table for the final measurement. groundwater levels in areas of the babol city is presented figure 10 [36]. 3 results and discussion in reality, soil behaves as a deformable body instead of as a rigid body. as a result, the actual peak shear stress induced at each depth is less than that predicted in surface. therefore, the rigid body shear stress should be reduced with depth reduction factor (rd), to give the deformable body shear stress. the depth reduction factor (rd) can be computed as follows (equation 7) [37-38]: rd = (τmax)depth (τmax)rigidboady = (amax)depth (amax)surface (7) amplification refers to the increase in the amplitudes of seismic waves as they propagate through the soil layers near the surface of the earth. this is because the ground under these districts is relatively soft. soft soils usually, amplify ground shaking. influence of the soil response on the seismic motion at the ground surface, is considered through the equivalent linear or nonlinear one-dimensional response of a soil column. these methods of analysis require selection and scaling of ground motions appropriate to design hazard levels. it is necessary to select empirical recordings of ground motion and scale these ground motions to the level of the design spectrum. in order to develop the depth reduction factor (rd), the max acceleration at surface and each depth should be calculated [39]. in this research, site response analyses, performed by the geostudio software, were carried out for 35 boreholes using 4 scaled earthquakes with 109, 475, 975 and 2500-year return period. the details of mentioned scaled accelerograms are presented in table 3 and figure 11. finally the (rd) diagrams for each zone of study area is drown and a new empirical correlation is proposed. table 3 details of scaled earthquakes earthquake name a max (g) at bedrock return period (year) kareh-bas 0.108 109 bam 0.215 475 northridge 0.302 950 san fernando 0.364 2500 advances in systems science and application(2016) vol.16 no.3 43 (a) scaled time history records of the kareh-bas earthquake (109-year return period) (b) scaled time history records of the bam earthquake (475-year return period) (c) scaled time history records of the northridge earthquake (950-year return period) (d) scaled time history records of the san fernando earthquake (2500year return period) fig. 11 time history records of scaled earthquake for performing site response analyses fig. 12 35 zones of study area (site response analyses are performed for each zone) 44 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting advances in systems science and application(2016) vol.16 no.3 45 46 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting fig. 13 variations of the stress reduction coefficient (rd) with depth for various zones of babol city advances in systems science and application(2016) vol.16 no.3 47 in order to define the depth reduction factor for babol city, 35 zones were each analyzed using scaled input motions (figure 12). then the rd diagrams were developed for each zone. figure 13 shows the depth reduction factor for each zone based on different input motions. most researchers today agree that there is no correlation for assessment of depth reduction factor which can be applicable for different soil types and all regions. different site conditions such as soil and sediment erosion, geological situation, soil structure and fabric, seismicity of area and ... are the main factors of mentioned uncertainty. so it can be said that the best empirical correlations for a specific region should be assessed regarding to in-situ tests and site response analyses which are performed in that region. in this research, an empirical correlation is suggested based on more than 200 site response analyses for different zones of babol city, using 4 scaled earthquake motions. the proposed new correlation for estimation of rd as a function of depth is presented in equation 8. table 4 details of scaled earthquakes rd = −0.0425z + 1.0333 0 ≤ z ≤ 6 equation 8 rd = 0.0006z2 + 0.0048z + 0.7865 6 < z < 12 rd = −0.02z + 1.06 12 ≤ z ≤ 14 rd = 0.0025z2 + 0.072z + 0.261 14 < z ≤ 20 4 conclusion liquefaction in soil is one of the major problems in geotechnical earthquake engineering. the simplified method of seed and idriss (1971) is the most common procedure used for evaluation of liquefaction potential. in mentioned simplified method, an empirical correlation was developed for estimation of rd based on the 2153 site response analyses. previous developed rd functions can be unconservative in comparison to actual case-specific seismic response analysis. therefore, in places with high potential of liquefaction, the csr is recommended to be calculated using a modified dynamic response analysis. this research proposed the values of depth reduction factor (rd) based on the results of several extensive site response analyses of babol city at 35 zones. firstly, 4 modified and scaled acceleration records (kareh-bas earthquake with 109-year return period, bam earthquake with 475-year return period, northridge earthquake with 975-year return period and san fernando earthquake with 2500year return period) were used as input motions at the bedrock of the 35 zones. soil layers were idealized as horizontal layers. engineering properties of each 48 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting layer, ground water level and other necessary data collected during geotechnical investigations and geophysical surveys were used to develop the soil column extending form the ground surface to bedrock. then, local site effects of study area were evaluated numerically with equivalent linear method and finally the depth reduction factor (rd) diagrams and correlation were prepared for study area. in some cases, depending on geotechnical and geological situation of a site or seismicity of a study area, the csr calculated by the simplified method could be incapable, overestimated or underestimated. to summarize it can be said that, the variation of csr obtained by site response analysis could be considered more reliable. also in complex and unusual sites or during strong shaking levels, it can be recommended that the more comprehensive alternative is to develop depth reduction factor based on site response analysis. references [1] nguyen kh and gatmiri b. (2007), “evaluation of seismic ground motion induced by topographic irregularity”, soil dynamics and earthquake engineering, vol. 27, no. 2, pp. 183-188. [2] chavez-garcia fj, raptakis d, makra, k. and pitilakis, k. (2000), “site effects at euroseistest-ii. results from 2d numerical modeling and comparison with observations”, soil dynamics and earthquake engineering, vol. 19, no. 1, pp. 23-39. [3] chopra , a k (2001), dynamic of structures theory and application to earthquake engineering, prentice hall. [4] frankel a. (1993), “three-dimensional simulations of ground motions in the san bernardino valley, california, for hypothetical earthquakes on the san andreas fault”, bulletin of the seismological society of america, vol.83, pp.1020-1041. [5] gueguen p, chatelain jl, guillier b et al. (1998), “site effect and damage distribution in pujili (ecuador) after the 28 march 1996 earthquake”, soil dynamics and earthquake engineering, vol.15, no.5, pp.329-334. [6] youd tl, idriss im, andrus rd et al. (2001), “liquefaction resistance of soils. summary report from the 1996 nceer and 1998 nceer/nsf workshops on evaluation of liquefaction resistance of soils.”, journal of geotechnical and geoenvironmental engineering , vol. 127, no. 10, pp. 817-833. [7] seed hb and peacock wh (1971), “the procedure for measuring soil liquefaction characteristics”, soil mechanics and foundation division (asce) , vol. 97, pp. 1099-1119. advances in systems science and application(2016) vol.16 no.3 49 [8] kramer sl (2008), evaluation of liquefaction hazards in washington state, ground settlement. [9] zhou yg, chen ym and ke h (2005), “correlation of liquefaction resistance with shear wave velocity based on laboratory study using bender element.”, journal of zhejiang university science , vol. 6, no. 8, pp. 805-812. [10] youd tl and perkins dm (1978) , “mapping of liquefaction induced ground failure potential. ”, geotechnical engineering division (asce), vol.104, no. 4, pp. 433-446. [11] youd tl and hoose sn (1977) , “liquefaction susceptibility and geologic setting ”, proceeding, 6th world conference on earthquake engineering, prentice-hall, englewood cliffs, vol.3, no. 4, pp. 2189-2194. [12] wang w (1979), “some findings in soil liquefaction”, water conservancy and hydroelectric power scientific research institute,beijing. [13] robertson pk, woeller dj and finn wd (1992), “seismic cone penetration test for evaluating liquefaction potential under cyclic loading”, canadian geotechnical journal, vol. 29, no. 3, pp. 686-695. [14] olsen rs, youd tl and idriss im (1997), “proceeding, nceer workshop on evaluation of liquefaction resistance of soils”, national center for earthquake engineering research,state university of new york, buffalo, pp. 225-276. [15] seed hb and idriss im (1971) , “simplified procedure for evaluating soil liquefaction potential.”, geotechnical engineering division (asce) ,vol. 97, no. 9, pp. 1249-1273. [16] seed rb, cetin ko, moss re et al. (2003), recent advances in soil liquefaction engineering: a unified and consistent framework, university of california, berkeley. [17] seed hb (1979), “soil liquefaction and cyclic mobility evaluation for level ground during earthquake”, geotechnical engineering division (asce),vol. 105, no. 2, pp. 201-255. [18] liao ss and lum ky (1998), “statistical analysis and application of the magnitude scaling factor in liquefaction analysis.”, geotechnical earthquake engineering and soil dynamics, vol.1, pp. 410-421. [19] robertson pk and wride ce (1998), “evaluating cyclic liquefaction potential using the cone penetration test”, canadian geotechnical journal, vol. 35, no. 3, pp. 442-459. 50 h. santhi, n. jaisankar: load-aware congestion adaptive multipath multicasting [20] youd, t. l., idriss, i. m. et al.(2001), “liquefaction resistance of soils: summary report from the 1996 nceer and 1998 nceer/nsf workshops on evaluation of liquefaction resistance of soils”,journal of geotechnical and geoenvironmental engineering, vol. 127, no. 10, pp. 817-833. [21] golesorkhi r (1989), factors influencing the computational determination of earthquake-induced shear stresses in sandy soils, phd dissertation, university of california at berkeley. [22] cetin ko, seed rb, kiureghian ad et al. (2004), “standard penetration testbased probabilistic and deterministic assessment of seismic soil liquefaction potential.”, journal of geotechnical and geoenvironmental engineering, vol. 12, pp.1314-1340. [23] farrokhzad f, choobbasti aj, barari a et al. (2011), “assessing landslide hazard using artificial neural network: case study of mazandaran, iran.”, carpathian journal of earth and environmental sciences, vol. 6, no. 1, pp. 251-261. [24] farrokhzad f, choobbasti aj and barari a (2012), “liquefaction microzonation of babol city using artificial neural network.”, journal of king saud university (elsevier), vol. 24, no. 1, pp. 89-100. [25] farrokhzad f, barari a, ibsen lb et al. (2011), “predicting subsurface soil layering and landslide risk with artificial neural networks: a case study from iran”, geologica carpathica , vol. 62, no. 5, pp. 23-30. [26] kramer sl, paulsen sb (2004), proceedings, international workshop on uncertainties in nonlinear soil properties and their impact on modeling dynamic soil response, university of california, berkeley. [27] kawano m, asano k, dohi h et al. (2010), “verification of predicted nonlinear site response during the 1995 hyogo-ken nanbu earthquake.”, soil dynamics and earthquake engineering, vol. 20, no. 5, pp. 493-507. [28] godano c and oliveri f (1998), “nonlinear seismic waves: a model for site effects”, international journal of non-linear mechanics,vol. 34,pp. 457-468. [29] assimakia d and kausel e (2002), “an equivalent linear algorithm with frequencyand pressure-dependent moduli and damping for the seismic analysis of deep sites”, soil dynamics and earthquake engineering, vol. 22, pp. 959-965. advances in systems science and application(2016) vol.16 no.3 51 [30] choobbasti aj, rezaei s and farrokhzad f (2013), “evaluation of site response characteristic using microtremors”, gradevinar, vol. 65, no. 8, pp. 731-741. [31] idriss im and seed hb (1968), “seismic response of horizontal soil layers”, journal of the soil mechanics and foundations division, vol. 94,pp. 10031031. [32] ishibashi i, zhang x (1993), “unified dynamic shear moduli and damping ratios of sand and clay”, soils and foundations, vol. 33, no. 1, pp182-191. [33] choobbasti aj, shooshpash e, farrokhzad f (2012), “3-d modeling of soils unsaturated depth using artificial neural network (case study of babol)”,unsaturated soils: research and applications (speringer), vol. 2, pp. 317-324. [34] farrokhzad f, barari a, choobbasti aj et al. (2011), “neural network-based model for landslide susceptibility and soil longitudinal profile analyses: two case studies”,journal of african earth sciences (elsevier), vol. 61, pp. 349-357. [35] choobbasti aj, barari a, safaie m et al. (2008), “mitigation of flourd (pol e safid) landslide, northern iran, using non-woven geotextiles”, review of the bulgarian geological society, vol. 69, no. 1, pp. 49-56. [36] choobbasti aj, shooshpasha e and farrokhzad f (2013), “3-d modeling of groundwater table using artificial neural network-case study of babol”, indian journal of geo-marine science,vol. 42, no. 7, pp. 903-906. [37] seed hb, idriss im and arango i. (1983), “evaluation of liquefaction potential using field performance data”,journal of geotechnical engineering (asce), vol. 109, no. 3, pp. 458-482. [38] seed hb and lee kl (1966), “liquefaction of saturated sands during cyclic loading”, soil mechanics and foundations division (asce), vol. 92, pp. 105-134. [39] farrokhzad f, choobbasti aj and barari a (2011), “determination of liquefaction potential using artificial neural networks”, gradevinar,vol. 63, pp. 837-845. corresponding author farzad farokhzad can be contacted at: farzadfarokhzad@mit.ac.ir adv syst sci appl 2019; 01; 141-149 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/594 copyright ©0000 assa. adv. in systems science and appl. (0000) the application of graph decomposition to development of large-scale agent-based economic models v.l. makarov1, a.r. bakhtizin1, e.d. sushko1, g.b. sushko1 1) central economics and mathematics institute cemi ras, moscow, russia e-mail: albert.bakhtizin@gmail.com received may 17, 2018; revised march 13, 2019; published april 15, 2019 abstract: in this work we describe the application of the graph decomposition algorithms for the development of a scalable high-performance agent-based model of population of russia described in terms of demography, migration and transport flows. the simulated system consists of agents representing individuals and sets of links to other agents, which represent the social interactions of individual. individual agents in the model participate in several independent processes, for which different sets of social links is important such as family and neighbors. to perform a load balancing of agents between cluster computer nodes the metis graph decomposition algorithm was used. these algorithms allow to split the graph of agents and links into parts of similar size with least possible number of links between them. a number of numerical experiments was carried out for test model to estimate the influence of the parameters of the model on scalability. keywords: agent-based modelling, numerical modelling, parallel computing, graph decomposition. 1. introduction the simulation of complex economic processes by means of agent-based approach requires a large number of agents to be used. one can expect the number of agents to be comparable to the population of the modelled society such as the city or a country i.e. up to 109 agents. high number of agents requires the use of supercomputer clusters to perform the simulation and requires a special programming technique to be applied in order to be able to run the model on such computer. a typical supercomputer nowadays is a cluster of multiprocessor multicore nodes with the shared memory inside the node and distributed memory between the nodes of the cluster. that distributed memory model requires the data describing agents to be evenly split between the nodes of the cluster. the paper describes the application of the metis[1,2,3] spatial decomposition algorithm for conducting multi-agent simulations of the behavior of society using supercomputers. there is a number of technologies and frameworks for development of abm (agentbased modeling) simulation software such as microsoft axum [4] and repast hpc [5] and swages [6]. these frameworks implement mechanisms of exchange of messages between agents but the problem of optimal distribution of agents between processes should be solved by the simulation code taking into account all types of interactions in the model. for example, a similar approach using metis algorithm for a large-scale epidemiologic abm using repast hpc and metis was implemented in paper [7]. another example is the application of the metis decomposition for modeling of railway scheduling in paper [8]. e ach of these codes used its own implementation of the decomposition code to transform the objects in the model into graph representation for metis. mailto:albert.bakhtizin@gmail.com 142 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©2019 assa. adv. in systems science and appl. (2019) the simulated system consists of agents which represent individuals and a set of links to other agents, which represent the social interactions of individual. the individual agents in the model participate in several independent processes, for which different sets of social links is important such as family and neighbors. the agents of the system participate in two processes: 1) the process of reproduction of the population, and 2) the process of migration. in the first process they use messages exchange to search for the partner to form a family. in the second process the message exchange mechanism is used to obtain information about available jobs in different regions to determine the direction of migration. to conduct agent modeling of demographic processes with a large number of agents, it is necessary to create an initial state of the population of agents with a given distribution by region, age, sex, income and other relevant parameters. in the test case, reading of the data by region from the text file and the geometry of regions from the image file was realized. the data is used to construct a rectangular grid, each cell of which is connected to a certain region and parameters of agents in the cell are determined by the parameters of the region. a random distribution of the number of resident agents between grid cells (uniform distribution of population within each region with the fixed total number of agents in the system) and calculation of grid cells decomposition by processors was implemented. 2. the model description the most common way of writing simulation programs for cluster computers is to use c++/fortran language and mpi library which are available on all supercomputers. to simplify the development of the model the distributed abm framework [9] was developed for use with the high-level microsoft .net platform. the use of the high-level programming language for model description requires the implementation of wrapper code to system-level functions for inter-process exchange using native mpi library. as most of modern supercomputers run on linux operating system, we have decided to use microsoft .net core and mono [10] implementations of .net platform available for this os. the choice of these technologies was determined by the following criteria: 1. the system has to be scalable across multiple computational cluster nodes (i.e. use resource of multiple nodes for speedup) therefore the multithreading calculation model was not suitable as it is limited to single cluster node. 2. the model should be easy to develop and maintain and therefore the high-level programming language c# was used. 3. the system should be efficient and therefore the native mpi library was used instead of tcp/ip sockets or .net libraries like windows communication foundation as these technologies are not optimal for supercomputers and hpc applications. mpi libraries installed on each cluster computer are usually tuned for particular proprietary network system which is used on the cluster such as infiniband. 4. the program should be portable and should be compatible with any operating system i.e. running on both developer workstations and on cluster computers. 5. the results of the simulation should not depend significantly on the number of processors used for the calculation. all interactions between agents on different processes on each simulation step should be taken into account. the simulation solution calculated on different number of processors can still be different due to the difference in the order of operations (due to the limited float point numbers precision) and the difference in the random number sequence (each process has its own independent random number sequence). an efficient mechanism of message exchange between agents was implemented by means of message queue and native mpi collective operations. the message queue accumulates a buffer of messages to different processes and then uses mpi alltoall exchange operation to deliver contents of messages. this operation delivers the buffers of the implementation of the scalable modelling framework 143 copyright ©2019 assa. adv. in systems science and appl. (2019) arbitrary size from each mpi process to all other processes in most efficient way by splitting the buffer into chunks of optimal size for network transfer and hiding the latency of network operations by performing simultaneous several send and receive operations. to use native operations with managed c# objects operations of binary serialization and deserialization of objects were implemented and c# wrappers for native functions were written. 2.1 the system decomposition algorithm the modeled system consists of a large number of agents connected with each other by the number of social links. to perform an efficient simulation of the system it is necessary to split the set of agents between mpi processes taking into account their links with other agents. this distribution of agents between processes determines the parallel scalability of the program. the quality of the decomposition is determined by the balance of the agents count on different processes and the number of links between agents assigned to different processes. the decomposition of a graph of agents with links can be performed by the application of graph decomposition algorithms such as metis [1], scotch [11] or jostle [12]. these algorithms provide a computationally efficient way of decomposition of large graphs into parts of equal size with minimal borders between them. such approach leads to a best possible system decomposition with a fine distribution of agents but it also leads to a number of drawbacks: 1) it requires the formation and a decomposition of a sparse matrix of size n, where n is the total number of agents, 2) each change of a number of agent due to agent’s birth or death leads to a recalculation of the decomposition. in order to calculate the decomposition in a more efficient way the system was split by a rectangular grid covering all regions. the decomposition of cells is much faster due to smaller matrix and doesn’t require recalculation of the decomposition after each time step as the distribution of population changes slowly with time. due to the small size of the input matrix the decomposition can be performed using sequential version of metis algorithm without sensible impact on the parallel performance of the simulation. to calculate the decomposition of the grid, a graph metis algorithm with a weighting was used (metis_partgraphrecursive). the metis algorithm takes an input on a graph specified through the constraint matrix and an array of weights of the nodes of the graph and returns the optimal distribution of the graph by a given number of parts with minimizing the links between the parts. the link matrix is given in csr [13] format i.e. by the number of rows (n) and two arrays ia and ja. the array ia contains a sequence of partial sums of the number of nonzero elements in the matrix. this array is used to separate the elements of the ja array by strings. the ja array contains index lists of non-zero elements in all rows. in the test example, the size of the matrix corresponded to the number of non-empty grid cells, the adjacent grid cells (not diagonally) were considered to be related. as the weight of the cell, the number of agents that should be created in the corresponding cell was used. the calculation of the most optimal decomposition of such matrix is very computationally expensive. the metis library implements a multilevel recursive coarsegrained algorithm for calculation of the reasonably good decomposition. on each level the original graph is transformed to a smaller graph using the coarse-graining algorithm, the decomposition of the smaller graph is performed and then a special refinement of the decomposition is performed. as a result of calling the metis algorithm for a given coupling matrix and weights and a given number of processors, an array of numbers is created that describes the optimal binding of the grid cells to the processors. these data are used to distribute cells and agents 144 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©2019 assa. adv. in systems science and appl. (2019) by processors during the calculation, but do not affect the aggregation procedure of the received calculation results. the results of the calculation are aggregated by region and should not depend on the decomposition used. to test the decomposition algorithm a sample system was constructed using the data on the population of russia. russia can be a good example of a country with very different regions in terms of population. the total population of russia is 144 million people and the average density of population is 8.58 person per square kilometer ranging from 4626 people/km2 in moscow to 0.07 people/km2 in chukotka region. large cities such as moscow and st. petersburg were considered as separate regions in order to distinguish the parameters of the region and parameters of the enclosing region because of the principal difference in population density, income and other parameters. the figures below show the results of automatic grid decomposition using the metis algorithm for 8 processors in two variants: without taking into account the number of agents in cells (the weight of each cell = 1) and taking into account the number of agents (cell weight equals to number of agents). the implementation of the scalable modelling framework 145 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 1. the decomposition of a territory of russia by 8 processors taking into account only the number of cells (top) and the number of agents in cells (bottom). one can notice much smaller regions in the western part of russia on a second plot due to much higher population density. different shades of red show areas assigned to different processors. in the first case, the algorithm equalizes the number of grid cells on different processors with the minimum length of the boundary between them. in the second case, the number of agents attributed to the cells of a specific processor is equalized, while minimizing the length of the border. one can see that in the second case the regions have different area because of the difference in population density. the difference in density between european part of the country (27 people per km2) and asian part (3 people per km2) leads to much smaller regions in european part in terms of area and cell count. another effect associated with the formulation of criteria for optimizing distribution is that the cells located on islands and exclaves are isolated and not directly connected to the nearest regions and in terms of decomposition can be linked to any processor. 146 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©2019 assa. adv. in systems science and appl. (2019) the use of the sparse matrix of links between cells allows the further development of decomposition criteria which should in principle include the connectivity of cells which have no common border. these connections may arise due to air communication, railways and other ways of traveling which have to be taken into account for correct simulation of migration. in the test example presented above, the total number of agents was set at 8400000. at the same time, the share of agents in each region was determined by the share of this region in the population of the country. within the region, agents were distributed uniformly across cells. such a calculation scheme allows us to change the total number of agents for test purposes, keeping their correct geographical distribution. using the algorithm described above, the distribution of the source system by the processors was obtained. in the implementation, it is important that such a distribution must be done before the agents are created, since the agents created must by initially correctly distributed among the nodes of the cluster to meet the limitations of the system memory. to analyze the obtained distribution of agents by processors, one can use amdahl's law, which connects the maximum achievable acceleration in parallel computations (sp) with the number of processors (p) and the fraction of consecutive computations (a). fig. 2. the illustration of amdahl’s law for description of parallel speedup on the number of processors and the value of parallel portion. in our case, sequential computations arise due to an imbalance in the number of agents on the nodes, i.e. if we have two nodes and the first 400 agents, and on the second 300, then we can say that the first 300 agents are processed by each node in parallel, and the second 100 agents are processed sequentially by the second one, because the second node can’t perform calculations at this time. the figure below shows the dependence of the parallel speedup on the number of mpi processes. the comparison of amdahl’s law estimates for decomposition with and without cell weights show that decomposition with cell weights leads to a very balanced distribution of agents and potentially to a very good parallel speedup. the real parallel speedup of the test simulations is significantly lower than amdahl’s law estimates but it still shows rather reasonable speedup. the implementation of the scalable modelling framework 147 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 3. the dependence of the parallel speedup on the number of mpi processes. the scalability of the test simulation is compared to the ideal speedup curve, and to the estimate speedup based on the amdahl’s law for the cases of decomposition with and without cell weights. for the test calculation, a system consisting of 8400000 agents was selected. for each agent the number of neighbor agents was selected randomly (1..10), the neighbor agents were selected randomly within the same cell on the map and it’s direct neighbors. the table shows the time of calculation of the step and the average number of agents on one processor. table. 1. the results of the simulation of 8.4 million of agents for different number of processors. number of cores time step simulation time, s parallel speedup number of agents per core fraction of remote message sends, % 1 55.683 1 8417590 0 2 28.872 1.92 4208795 0.15 4 16.114 3.45 2104397 0.39 8 10.262 5.42 1052198 0.68 12 7.006 7.94 701465 0.91 16 5.268 10.56 526099 1.10 24 3.796 14.66 350732 1.72 48 2.161 25.75 175366 2.26 96 1.378 40.38 87683 3.41 192 0.846 65.78 43841 5.56 148 v.l. makarov, a.r. bakhtizin, e.d. sushko, g.b. sushko copyright ©2019 assa. adv. in systems science and appl. (2019) the results shown in the table above show that the distribution of agents by processors as a result of the decomposition was fairly uniform, but the overall parallel speedup in this version is limited to a several dozen times. the table also shows the dependence of the fraction of remote message sends which were delivered to agent on another processor. the increase of the number of processors leads to increase of the number of borders between areas attributed to different processors, the number of remote sends and the ratio between network exchanges and local calculations within one processor which is another factor which affects the scalability of the simulation. uniform distribution of the number of agents by processors allows you to evenly distribute the load of ram of each node of the cluster. as a test case, a multi-agent system was calculated with 148 million agents on 192 cluster processors. the main difficulty in this case was that with the total amount of ram of the allocated nodes in 192gb on each node of the cluster there was only 8gb of ram. thus, it was necessary to distribute the agents evenly, so that on none of the cluster nodes the simulation processes exceed the memory consumption limit. fig. 4. decomposition of the computational grid into 192 processors taking into account population density. the figure above shows the distribution of the calculated grid for 192 processors, built taking into account the population density. in the distribution, the average number of agents per processor was 772000. the number of agents per processor ranged from 743711 to 787266, i.e. the spread of the number of agents on one processor was ~ 6% of the average. the calculation of one step in this simulation took about 15 seconds. in the future, the spread of the number of agents can be reduced by using a smaller calculation grid, which will improve the accuracy of the decomposition. 3.conclusion in this work a new framework for parallel calculations of agent-based models was presented and tested. the framework uses metis graph decomposition algorithm to perform natural spatial distribution of agents in abm-simulations to achieve high level of balance and parallel scalability. the decomposition algorithm based on cell distribution allows us to compute a balanced distribution of agents efficiently. which can be used for both initial distribution of agents and redistribution of agents during simulation. the provided results of the implementation of the scalable modelling framework 149 copyright ©2019 assa. adv. in systems science and appl. (2019) the test simulations show good scalability of the program across multiple computational nodes and possibility to perform abm simulations with number of agents comparable to population of large countries and regions. acknowledgements this work was supported by the russian science foundation (grant # 14-18-01968). the possibility to perform computer simulations at the mvs-100k joint supercomputer center and tianhe 2 supercomputer is gratefully acknowledged. references [1] karypis, g. & kumar, v. (1995). metis-unstructured graph partitioning and sparse matrix ordering system, version 2.0. [2] karypis, g. & kumar, v. (1999) parallel multilevel k-way partitioning scheme for irregular graphs. // siam review, vol. 41, no. 2, pp. 278 300 [3] lasalle, d. & karypis, g. (2013) multi-threaded graph partitioning // 27th ieee international parallel & distributed processing symposium [4] https://en.wikipedia.org/wiki/axum_(programming_language) [5] https://www.bsc.es/computer-applications/pandora-hpc-agent-based-modellingframework [6] scheutz, m., connaughton, r., dingler, a., & schermerhorn, p. (2006). swages an extendable distributed experimentation system for large-scale agent-based alife simulations. in proc. of artificial life x, pp. 412-419. [7] collier, n., ozik, j., & macal, c. m. (2015). large-scale agent-based modeling with repast hpc: a case study in parallelizing an agent-based model. european conference on parallel processing, pp. 454-465 [8] salidol, m. a., abril, m., barber, f., ingolotti, l., tormos, p. et. al. (2006). domain dependent distributed models for railway scheduling. in international conference on innovative techniques and applications of artificial intelligence, pp. 163-176 [9] bakhtizin a.r. et al. (2017) the development of the agent-based demography and migration model of eurasia and its supercomputer implementation. // advances in systems science and applications, [s.l.], v. 17, n. 4, p. 34-45 [10] http://www.mono-project.com [11] chevalier, c. & pellegrini, f. (2008). pt-scotch: a tool for efficient parallel graph ordering. // parallel computing. 34 (6): 318–331. doi:10.1016/j.parco.2007.12.001 [12] walshaw, c. & cross, m. (2000). mesh partitioning: a multilevel balancing and refinement algorithm. // journal on scientific computing. 22 (1): 63– 80. doi:10.1137/s1064827598337373. [13] tinney w. & walker j. (1967) direct solutions of sparse network equations by optimally ordered triangular factorization. // proceedings of the ieee, 55(11):1801–1809 https://en.wikipedia.org/wiki/axum_(programming_language https://www.bsc.es/computer-applications/pandora-hpc-agent-based-modelling-framework https://www.bsc.es/computer-applications/pandora-hpc-agent-based-modelling-framework http://www.mono-project.com/ https://en.wikipedia.org/wiki/digital_object_identifier https://doi.org/10.1137%2fs1064827598337373 adv syst sci appl 2018; 03; 29-38 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/654 situational awareness formation for large network-systems viacheslav abrosimov1* 1) smart solutions, samara, russia e-mail: avk787@ yandex.ru abstract: an approach to situational awareness formation as a necessary condition for the operation of large network-centric systems is considered. the basis is information acquisition from all system elements according to the “initial data-parameter-event-situation-solution” scheme. the situational awareness model is constructed as a neural network with an ensemble organization and associative-projective structure. the model is learned using precedent examples. keywords: network, situation, awareness, neuron, network, ensemble 1. introduction large systems include a large number of elements and have a complex hierarchical structure and, therefore, a complex character of connections. it is necessary to take into account the existence of random factors from the environment and also the elements of the system itself. the characteristic feature of modern large systems is network centricity, which consists in the network nature of systems and use of various information for decision making. situational awareness is a necessary condition for the operation of large network-centric systems. large systems differ from complex systems, as the operator in the control loop is included in system. the situational awareness systems are actively developed for ensuring the efficiency and safety of large systems in manned aircraft, transport and other fields (e.g., see [1, 2]). the basis of such systems is the high operational efficiency of information interaction. in defence applications, situational awareness models are developed as organizational models for knowledge management on the battlefield and also for the accumulation of incoming information about the current situation in special-purpose databases. these data come to situational awareness system from intelligence agents, advanced units and from higher-level governmental structures including a combat command control information network. managing the processes of data collection allows to perform a single visualization of the combat space, depending on the level of management and the access rights of government officials [3]. 2. situational awareness problems for large systems according to the endsley model [4], the situational awareness concept distinguishes between three basic elements: a) information about the spatial and temporal environment of the situation b) situation analysis and prediction in terms of scenarios and c) the actions of other participants. obviously, to ensure a high degree of situational awareness, all elements of a large system should be considered as information providers. analysis shows that the main problem of modern large-scale economic systems is the lack of properly organized information and communication network between the elements of the system that implements the principles of situational awareness, i.e., a common information environment and its continuous * corresponding author: avk787@ yandex.ru 30 v. abrosimov copyright ©2018 assa adv. in systems science and appl. (2018) updating using the largest possible information sources (large system inputs). but the major unresolved problems in the situational awareness fields is decision uncertainty caused by the abundance of redundant information. events and situations objectively occur in large systems. an event in a large system is understood as the realization of a random process associated with predicted or previously neglected factor demonstration. it is clear that events in large systems are not independent. moreover, in large systems, the emergence of an event from outside (the influence of the external environment) and the event generated by the system are both observed. often the source of an event remains undetermined, and its impact on large system functioning is critical. a situation in a large system is understood as a group of events that lead to the result of system functioning and requires decision making. a precedent in a large system is understood as a situation that arose as a result of the onset of an events aggregate caused by external factors or as a result of setting the aggregate parameters of a large system, according to which a) appropriate decisions were made and b) the consequences of such decisions are known. consider an example. in the large economic and ecological system "a forest production and extraction" described by various parameters (total area, volume of felling, percentage of deadwood, etc.), the event of abandoned fire occurs, leading to the situation of fire and causing a precedent (annual fires on the territory of the russian federation). it is impossible to study the influence of all factors determining events and situations in large systems. in this case, it is necessary to apply analytical methods allowing, on the one hand, to carry out estimation with maximum complete consideration of all features/ each element of a large system determines and accumulates information, assesses the situation in accordance with its limited capabilities, thereby being able to act only in its narrow area of responsibility. but in a properly constructed large network-centric system, there is the gradual accumulation of experience and knowledge. all system elements participate in this process with the course of time. however, such phenomena as "the of situational awareness impulse " may arise. this reflects the fact that previously missing information is suddenly detected by observing any of system elements. such impulses are associated with events occurring in large organizational and technical systems. in practice, also there may be a “gap in situational awareness”. it occurs as a result of limited or even lost communication and, accordingly, interaction between the elements of the system. with a high degree of situational awareness, the elements of the system may function as autonomous ones. however, in future system life, the accumulated underdevelopment can lead to dangerous errors. the elements of the system must know, predict and respond to potentially possible situations related to the performance of their own functions. they should also know the strategies of the other elements that are a part of the larger system to coordinate their actions. this means that some information resources should be constantly filled by current information: initial data, events and situations that arise during large system functioning. large system situational awareness formation is the process of collective accumulation of information that describes the events and situations generated by the environment and system initial conditions, which is carried out in parallel with system functioning and implemented by all large system elements. based on the mathematical model of neural networks, this paper suggests a methodological approach for investigating the problems of situational awareness formation. 3. basic principles let us formulate some basic principles. the principle of recurrence. it means that situations that occur during large system functioning in general way are repetitive, similar. however, situations could never be the same. but, according to our approach, a difference comes from a number of inessential background situation awareness formation for large network systems 31 copyright ©2018 assa. adv. in systems science and appl. (2018) details that do not affect the situation, while the situation’s similarity appears due to the coincidence of a few main features. the application of methods based on the principle of analogies for the study of events and situations is exposed to certain criticism in scientific papers. however, this principle is widely and successfully used by many researchers. the principle of hierarchy. the situational awareness formation can be conditionally described as a certain hierarchical structure that allows to design situational awareness components interacting with each other. the feasibility of such formalizations is obvious. the problem consists not so much in decomposition, but in the representation of constituent elements, the definition of appropriate interconnections between different-level elements of the hierarchy and also in the completeness of these connections and the degree of their mutual influence. the principle of comparability. the situational awareness entities can be compared using various criteria. in practice, such a comparison is made frequently and, depending on criteria, has either the character of compared integrated numbers or fuzzy statements. the principle of importance. the situational awareness information can be described by a limited set of most important parameters, events and situations. other parameters can be considered as a certain background. 4. approach to situational awareness formalization among the rich variety of possible mathematical approaches to the formalization of situational awareness, we suggest to use neural networks. application of neural networks for solving different problems was mentioned in literature. artificial neural networks are currently actively used in forecasting, pattern recognition, and control. first of all, the architecture of neural networks is very easy to redesign; this is extremely convenient for adding elements for structure when new influencing factors appear. secondly, the idea of future events and situations is always objectively incomplete and fuzzy. an advantage of neural networks is that they have ability to restore information from its separate parts. thirdly, neural networks have learning ability, which allows to make decisions by analogy in case of situations’ similarity. fourthly, neural networks enable to consider complex situations formed under the influence of factors of various nature and dimension. they can use heterogeneous information that is also essential for analysis models. with such networks, statistical data and expert estimations can be used together and the information of various probability can be processed. a neural representation is based on the concept of a neuron and neural ensemble [6]. as usually a neuron is a formal element whose output activity ranges from 0 to 1 and is determined by the level of input excitation. the excitement level is determined by the sum of input excitations transmitted from other similar neurons taking into account the appropriate weight functions between the neurons. a neural ensemble is a set of neurons united by mutual exciting connections so that the whole ensemble becomes excited if a given number of neurons in it are excited. neural ensembles have unique properties, which allow working with them as with a single unit and when necessary to split them up on small parts and to create new compound structures. the ensembles are intended for the formalization of neural image concepts with various degree of complexity. our model is based on the hypothesis about the separation of a neural network into individual hierarchical levels. at every hierarchical level, neurons are connected to each other and create neural ensembles. there are special operations over ensembles, which allow to add, subtract, create, and destroy ensembles, as well as to establish (destroy) connections between them, and so on. the ensembles can be united into groups (fields). the opportunity to obtain various fields of ensembles is provided, such as: associative (for creation of ensembles), buffer (for storage of ensembles) and difference (for comparison of ensembles among themselves). ensembles and 32 v. abrosimov copyright ©2018 assa adv. in systems science and appl. (2018) ensemble groups are linked to each other by the so-called “connections” transferring the excitation image from one level of hierarchy to another. through such “connections,” the excitement of ensembles, e.g., at a lower level of hierarchy causes the adjustable excitation of the higher-level ensembles. high reliability is an important feature of neural models. practical experiments have shown that the failure of 10% and even more elements of a neural network does not lead to significantly different results. it happens because each element in neural networks with the ensemble structure is responsible for many functions, not only for one. the failure of such an element deteriorates some functions, but owing to the presence of many elements and their interrelations this deterioration is so small that it becomes insignificant for the whole network. in conventional approaches, the factor of uncertain input information (which can be incomplete, fuzzy, doubtful etc.) may cause errors in decision making, as the elimination of initial uncertainty in prediction models can appear impossible in principle. in neural models, an image can be created using partially uncertain information. this image does not completely fit the reality but nevertheless is very "close" to it. by defining an excitement edge, the most significant factors can be studied. in other words, the influence of uncertain (and very often doubtful) input factors on the result can be reduced. 5. hierarchical neural network for describing large system situational awareness the basic idea of using neural networks to describe the situational awareness formation processes is to employ an excited ensemble as a semantic image of a parameter, characteristic, object, factor, event, situation, precedent, solution and other concepts of large systems operation [7]. from the author’s point of view, the idea to use a neural network for describing situational awareness has certain novelty. let us construct a multilayer hierarchical neural network and ensembles exciting by such neurons. the first level will be composed of neurons and ensembles of “input parameters,” corresponding to the initial data on various factors of the large system and its elements environment. the second level consists of neurons and ensembles of “events” that will semantically describe the events of the influencing environment or the facts occurring in the course of large system functioning. the third level consists of neurons and ensembles of “situations” that simulate situations arising in a large system. and, finally, at the fourth level, we will place the ensembles of “precedents,” which are associated with the decisions on such precedents. between the levels, we create projective connections. they will transmit excitation from one level to another. the excitation of lower-level ensembles in the hierarchy by projective connections leads to the controlled excitation of the upper-level hierarchy ensembles. the connections between neurons and neural ensembles will be established and changed according to the results of learning. they can be either exciting (increasing the activity of neurons and ensembles) or inhibitory (reducing activity) as a result, each concept is formalized by means of neural ensemble. at each level of hierarchy, the ensembles are identical in terms of structure (neurons which are linked by mutual exciting connections), yet different in terms of their meaning. the level of excitement of each ensemble depends on the importance of its influence on the ensembles in each consecutive level of hierarchy. connections between neurons and neural ensembles can be established and changed according to the results of training. they can be stimulating (increasing the activity of neurons and ensembles) or breaking (reducing the activity of neurons). the training procedure is as follows. “teacher” forms a variety of source data for neurons at the first hierarchical level. the excitation of the inputs of neurons and neural ensembles (the “parameters” of the first layer of situation awareness formation for large network systems 33 copyright ©2018 assa. adv. in systems science and appl. (2018) the hierarchy) through the ensemble and projective connections initiates the process of excitation for the neural network. the excitement of the projective links of the first and second layers is transmitted to the “event” ensembles. due to the connections within the ensembles and between the ensembles, the ensembles of the second layer are excited, and then, through the projective connections of the second and third layers, the “situation” ensembles of layer are excited. to each such ensemble at the third layer, the “precedent” ensemble is aligned at the fourth layer of the model hierarchy. the “precedent” ensembles are put in line with the decisions taken under given conditions and the consequences of such decisions. initiating the excitation of the corresponding neurons and ensembles in the course of training, “teacher” can set new non-standard events that form previously unknown situations, create the necessary precedents and point out the right solutions in such situations with the aim of shaping the behavior of a large system in the required effective direction. 6. situational awareness: filling of information the neural model is located and functions in a certain "internet-cloud". all elements of a large network-centric system have access to this source. situational awareness formation here is the excitation process of a multilayer neural network based on the “parameter-characteristic-eventsituation-precedent” principle, which runs sequentially as soon as information is received. each element of a large system at each time transmits its current data to the model. the "current situation" ensemble, which is excited in the way described above under the "parameter" "event" "situation" scheme, is compared with ensembles that are in the "precedents" field and built on the learning outcomes or mark during the experience of system development. the type and name of the ensemble that is closest to the emerged situation are then determined. the decisions that were made in similar cases and the consequences of such decisions are also determined. the analysis of these consequences makes it possible to estimate the occurrence of problems for a given realization of the set of input parameters and initial data. this information is used for making the decision. the neural model of situational awareness formation allows us to seek a solution in a neighborhood of the already known effective ones. for this purpose, an ensemble of "precedents" and the most effective solutions and satisfying consequences obtained in the process of learning is selected. a comparison between the characteristics of the input data vectors applied to the inputs of 1-st neurons level is made. the process of excitation again propagates through the neural network but with the modified input vector. a new “situation” ensemble is formed, and the ensembles are again compared. it is assumed that the necessary knowledge of precedents and effective solutions will be developed as the knowledge base of situational awareness during large system functioning. 7. conclusion the paper has reflected only the general principles of neural model functioning for the analysis of relations between different states. unfortunately, the limited format of a conference paper does not allow to give more details on many other important aspects of modeling. the problems of model training and other difficulties in its realization, such as the principles of coding for input information and decoding for output information, the issues of interpretation of modeling results, etc., have not been touched. the idea of neural models application for situational awareness formation has become an efficient tool to describe the situational awareness processes, to choose most effective strategies and to predict the consequences of decisions based on situational awareness information. 34 v. abrosimov copyright ©2018 assa adv. in systems science and appl. (2018) acknowledgements this work was supported in part by the russian foundation for basic research, project no. 16-08-00832-a. references [1] fedunov, b.e. (2002). bortoviye operativno sovetuyuschie ekspertnie sistemy takticheskih samoletov piatogo pokoleniya (obzor po materialam zarubezhnoy pechati) [onboard advisory expert systems of tactical aircraft of the fifth generation (review of foreign press materials)] moscow, sic gosniias, [in russian]. [2] ermakov, a.n.,merkulov, а.а., panfilov, s.a., raikov a.n. (2015). podderzhka reshenij v avarijnyh situacijah na zheleznoj doroge na osnove setevoj jekspertno-analiticheskoj sistemy [support of solutions in emergency situations on the railway road on the basis of the network expert-analytical system], intellektual'nye tehnologii na transporte, 2, 5-9 [in russian]. [3] rubtsov, i.v., boslyakov a.a.,lapshov,v.s., mashkov k.y, et.al. (2015). problemy i perspektivy razvitija mobil''noj robototehniki voennogo naznachenija [problems and prospects for the development of mobile robotics for military use]. inzhenernyj zhurnal: nauka i innovacii, 5(41), 26-31 [in russian]. [4] endsley, m.r. (1995) toward a theory of situation in dynamic systems. human factors, 37 (1), 32-64. [5] callan, r. (1999). the essence of neural networks. the essence of computing series. prentice hall europe, london. [6] kussul e.m. (1992). associativnye nejropodobnye struktury [associative neural structures]. kiev, naukova dumka [in russian]. [7] abrosimov, v.k. (2014). nejronnaja prostranstvenno-vremennaja model' dvizhenija ob"ektov upravlenija [neural space-time model of the motion of objects], nejrokomp''jutery: razrabotka, primenenie ,3, 26-35 [in russian]. adv syst sci appl 2021; 03:31–39 published online at https://ijassa.ipu.ru. convolution neural network based covid-19 screening model ashish nainwal*, gorav kumar malik, amrish jangra dept. of electronics and communication engineering, fet gurukul kangri (deemed to be) university, haridwar, uttarkhand, india abstract: coronavirus disease 2019 (covid-19) is a high death rate respiratory condition that requires easy-to-reach markers for prediction. the electrocardiograph (ecg) alterations that may occur after covid-19 hospitalization have not been fully studied yet. covid-19 also affects heart function, which can be seen on an ecg. as a result, ecg can be used to detect virusinfected individuals. the database consists of ecg images. in this scenario, a convolution neural network (cnn) is utilized to classify covid-19 ecg. the model is made up of eight layers, including a convolution layer, a max-pooling layer and a dense layer. the ecg image is fed into a cnn model, which classifies the covid-19 ecg. the model provides us with 98.11% accuracy, 98.6% sensitivity and 96.40% specificity. although 100.00% of the categorization of normal images and covid-19 ecgs were not accurately determined by the proposed cnn model, this is the first cnn model to categorize ecg images into normal and covid-19 classes from the ecg database and provide additional diagnostic to medical experts. keywords: ecg, convolution neural network, covid-19 1. introduction coronavirus disease 2019 (covid-19) is the clinical sign of contamination with the severe acute respiratory syndrome coronavirus-2 (sars-cov-2). more than 170,000,000 persons in more than 180 nations or regions worldwide suffered as a result of the 2019 coronavirus disease (covid-19) global pandemic [1]. the infection’s clinical course is characterized by respiratory symptoms (fever, cough, and tiredness), which can progress to pneumonia, acute respiratory distress syndrome (ards) and shock [2]. the modern world is in the midst of a never-before-seen health disaster. the covid-19 outbreak is wreaking havoc on hospitals and medical professionals worldwide. according to epidemiological statistics, coronavirus patients with prior cardiovascular irregularities are at a higher danger of pre and post covid19 related issues and mortality [3]. covid-19 has the potential to affect the heart and lungs. according to several articles, covid-19 has been related to myocarditis, acute coronary syndromes and decompensated heart failure [4] [5] [6]. the pandemic also emphasizes the importance of preparing cardiac healthcare for covid-19 as well as acute and chronic cardiac therapy in individuals where patient and medical delays could be hazardous. lecun et al. introduced cnn in 1990 and in recent years, it has become one of the most potent methods of machine learning [7]. cnn offers benefits both in accuracy and performance in imaging recognition, sound classification and semantic identification with the support of fast-growing graphics process unit (gpu) technology [8] [9] [10]. in addition, recent research has demonstrated the considerable promise of cnn with biological applications including categorization of animal behavior, the prediction of protein structures ∗corresponding author: ashishnainwal86@gmail.com 32 a. nainwal, g k malik, amrish and pattern recognition (emg) [11] [12]. recent research has also identified intriguing uses of cnn in bio-signals such as ecg for time series [13]. the cnn framework benefits from using the huge training data set to overall improvements of classification parameters. degerli et al. have developed a unique methodology to the composite location, gradation and detection of covid-19 from 15495 cxr images with the creation of the ”infection mappings” capable of precisely detecting and grading covid-19 seriousness with a precision of 98.69% [14]. kesim et al. suggested a new cnn model for chest radiography data (xray images) categorization [15]. several in-depth learning models, including chest x-ray pictures and computer ct scans have been proposed in recent research to detect anomalies of covid-19 in patient healthcare [16]. in this article, the covid-19 patient ecg and normal ecg images were put to the proposed convolution neural network. the proposed cnn model has eight layers, comprising three convolution layers, three pooling layers and two dense layers. the approaches yielded good results for classifying covid-19 patients based on their ecg. 2. method and material 2.1. database the collection of data is in image data form [17]. ch. pervaiz elahi institute of cardiology, nishtar medical university and punjab institute of cardiology are the data sources, all of which are situated in pakistan. this image data collection is collected with 12 lead-based edan series equipment. the rate of sampling is 500hz. several medical professionals have marked all ecg images. figure 2.1 and 2.2 shows the image of normal ecg and covid-19 patient ecg respectively. s. no type of ecg no of images sampling rate 1 covid-19 250 500 2 normal 859 500 table 2.1. data set details all ecg devices used to collect data have been set up as ”on” for important warnings. ecg technicians are taught to respond to all forms of edan ecg device alerts so that ecg technicians are able to carry out all precautionary actions during ecg operations. this step is crucial to more precisely capture ecg images, the data created by highly qualified professionals actively involved for many years, collected ecg images were manually reviewed by medical experts and supervised by senior physicians with ecg interpretation expertise. 2.2. preprocessing the ecg interpretation is done by image, so we need to improve image quality; we used a non-linear method gamma correction for enhancement of images its improve the interpretability or view of data and give ’better’ contribution to input preparation for image processing techniques. for normalization of images several operations are performed like scalar multiplication, addition and subtraction gamma correction alludes to the image enhancement on contrast by changing the unique range of pixel intensity distribution and conveys a non-linear technique on the source image pixels and can cause immersion of the picture is changed. moreover, if the gamma value is too large or small shows a low contrast image, from the outset, gamma revision appears to either obscure or light up a picture, however, this is a gross misrepresentation. figure 2.3 shows gamma-corrected image on two different gamma values 0.47 and 1.83. we would already be able to change the normal brightness of a picture by some adjusted standardization calculation, the easiest of which copyright © 2021 assa. adv syst sci appl (2021) convolution neural network based covid-19 screening model 33 fig. 2.1. normal ecg fig. 2.2. ecg of covid-19 patient would be basically enhancing every pixel force esteem, viably ”moving” the mean pixel power esteems across the whole picture. pixel (p ) of the image defined in a range between 0-255, δ is a representation of angle value, grayscale image of ecg signal represented by k of the pixel (kεp ) choose a midpoint km of the range [0, 255]. the linear map from pixel group p to group δ represents as: φ : p → δ = {ω|ω = φ(k)},φ(k) = πk/2km (2.1) the δ is mapped to γ (gamma value symbol) h : δ → γ,γ = {γ|γ = h(k)} (2.2) h(k) = 1 + f1(k) (2.3) f1(k) = a cos(φ(k)) (2.4) where a lies between 0 to 1 as a weighted factor and h is the histogram intensity level. according to map, group p and gamma group pixel values are interrelated. gamma number helps to find the arbitrary pixel value. let γ(k) = h(k), and gamma correction function is defined as: t (k) = 255 ( k 255 )1/γ(x) (2.5) copyright © 2021 assa. adv syst sci appl (2021) 34 a. nainwal, g k malik, amrish fig. 2.3. : input image: (i)ecg trace image original (ii)gamma corrected image when γ=.47 (iii) gamma corrected image when γ=1.83 where the output pixel correction value is represented by t(k) in grayscale. after gamma correction, the dataset is treated so that the ecg images scale to meet the cnn network input picture size needs. the z-score normalization was performed with the average and standard deviation of images. 2.3. convolution neural network cnn is the popular sort of neural artificial network that includes algorithms for supervised learning tasks. it is also an essential tool for deep learning with many hidden layers and parameters. in several areas like image processing, design recognition and other cognitive tasks, the cnn was widely used. a standard cnn is made up of three-layer types: the convolutions layer, the fully connected layer and the pooling layer. the convolution layer conducts overlapping operations throughout the entire image spatially to create characteristics. for downsampling maps, the pooling layer is responsible. the fully connected layer classifies depending on the characteristics learned. one-dimensional signal evolution is referred to as 1d convolution or simply convolution. if the convolution occurs between two signals covering two perpendicular dimensions, the convolution shall be called a 2d convolution. figure 2.4 shows the 2d convolution operation. this notion may be expanded to include multi-dimensional signals via which multi-dimensional convergence is possible. polling is used to simplifying or reducing the information acquired from feature maps spatial dimensions. the most frequent usage of pooling is the maximum of pooling because of its speed and increased convergence. this is what makes a filter (usually the size 2x2) and a step the length is the same. figure 2.5 shows the polling operation. each input is connected to each output and hence has the phrase ”fully connected. ”it is the last layer, usually following cnn’s last pooling layer. fully linked layers act like a typical neural network and contain around 90% of cnn parameters. this copyright © 2021 assa. adv syst sci appl (2021) convolution neural network based covid-19 screening model 35 fig. 2.4. 2d convolution operation fig. 2.5. operation of max polling layer essentially enters the last pooling layer output and produces a n-dimensional vector where n is the number of classes to pick from. as the function for neuron activation, non-linear transfer functions are employed in ann. for example, the most frequent activation functions are sigmoid f(x) = 1/(1 + exp(−x)) and hyperbolic tangent f(x) = tanh(x). sigmoid and hyperbolic tangents are both non-linear saturating functions, which with an increasing input decrease to almost zero. a recent study has shown the growth in cnn applications’ rate of speed and rating performance, as well as the classification performance of such a linear rectified function f(x) = maximal(0,x)(relu). the activation function relu is in the dense layer of our cnn model. the suggested 2-dimensional (2d) cnn model is shown in figure 2.6. the eight layers of the proposed model consist of three convolution layers, three max-pooling layers, and two dense layers(fully connected). the layers are converged to the corresponding kernel size for each convolution layer (64, 32 and 64). the max-pooling layer also called a downsample layer, should be utilized to the features maps after every two convolution layers. it was used to copyright © 2021 assa. adv syst sci appl (2021) 36 a. nainwal, g k malik, amrish fig. 2.6. cnn model decrease computer complexity and override controls. the stride is set to 1 and 2, respectively, for the convolution and the max-pooling layer. two activation functions relu and softmax are used in dense layer to improve the performance of the model. 3. result and discussion this section discusses the performance of the ecg image classification network. we suggested an efficient approach to extract and detect characteristics from each ecg image to examine the performance of the proposed system. the cnn model is built in python by using the keras library in colab.research (open source) google platform. the filtered image is sent into the cnn model, which classifies it. the image is filtered during preprocessing to increase system accuracy. different measures have been used to evaluate the system’s performance. three key indexes are used to assess the classifier’s performance, namely accuracy(acc), sensitivity(se) and specificity(sp). se = (tp/tp + fn))× 100 sp = (tn/(tn + fp )× 100 acc = ((tp + tn)/(tp + tn + fp + fn))× 100 predicted normal covid-19 accuracy sensitivity specificity original normal 847 12 98.11 98.60 96.40 covid-19 9 241 98.11 96.40 98.60 table 3.2. confusion matrix for complete data set predicted normal covid-19 accuracy sensitivity specificity original normal 243 7 97 97.2 96.8 covid-19 8 242 97 96.8 97.2 table 3.3. confusion matrix for balanced data set copyright © 2021 assa. adv syst sci appl (2021) convolution neural network based covid-19 screening model 37 fig. 3.7. training and validation accuracy and loss plot for complete data set and balanced data set where the number of impostor acceptations is false positive (fp), true negative (tn) is an impostor refusal number; false negative (fn) is a valid denial number and true positive (tp) is a valid acceptance. table 3.2 shows the confusion matrix for complete data set with 859 samples of normal ecg and 250 samples of covid-19 patient ecg.proposed convolution neural network used 60% of data for training and 40% data for testing purposes. it demonstrates that 98.60% of normal ecg segments are accurately classified in the normal class, while 96.40% of covid19 images are accurately classified in the covid-19 class. only 1.40% and 3.6%of ecg images are misclassified as covid-19 and normal class, respectively. the model incorrectly detects 12 normal ecg images as covid-19 ecg and correctly detects 847 images. out of 250 covid -19 ecg images, the network incorrectly classifies 9 of them, while the rest are correctly identified. with the entire data set, the accuracy is 98.11%. table 3.3 illustrates the confusion matrix for a balanced data set with 250 samples of each class. out of 250 images, 7 normal ecg images and 8 covid-19 ecg images are classified copyright © 2021 assa. adv syst sci appl (2021) 38 a. nainwal, g k malik, amrish wrong. in this data set 2.8 % of normal ecg images are incorrectly labeled as covid-19. furthermore, the misclassification rate of the covid-19 ecg image is approximately 3.2 %. the accuracy of this scenario is 97%. figure 3.7 shows the graph between accuracy and loss versus no of epochs. the first two graphs show the training and validation accuracy vs. epochs and loss vs. epochs, respectively for the complete data set. the following two plots are for a balanced data set. it can also be shown that after a few epochs, the networks achieve and stable with the highest accuracy and lowest loss. 4. conclusion this article presents an approach for the screening of covid-19 based on the convolution neural networks using data-set features. the suggested approach provides links between multiple layers of the original cnn architecture using convolution blocks, which produce dynamic layer combinations of different layers. the suggested approach is used to examine two scenarios of the classification task. the accuracy, sensitivity and specificity are used to analyze outcomes. the model gives 98.11% accuracy for complete data set and 97% accuracy for balanced data set. since the suggested approach leverages the ecg trace image that smartphone acquired and widely accessible facilities in low resource nations, this study helps to diagnose covid-19 and other heart defects as a second opinion by computer-aided methods. references 1. dong, e., du, h., & gardner, l. (2020) an interactive web-based dashboard to track covid-19 in real time. the lancet infectious diseases, 20, 533–534. 2. grasselli, g., et al. (2020) baseline characteristics and outcomes of 1591 patients infected with sars-cov-2 admitted to icus of the lombardy region, italy. jama, 323, 1574– 1581. 3. bergamaschi, l., et al. (2021) the value of ecg changes in risk stratification of covid-19 patients. annals of noninvasive electrocardiology, p. e12815. 4. doyen, d., moceri, p., ducreux, d., & dellamonica, j. (2020) myocarditis in a patient with covid-19: a cause of raised troponin and ecg changes. the lancet, 395, 1516. 5. zheng, y.-y., ma, y.-t., zhang, j.-y., & xie, x. (2020) covid-19 and the cardiovascular system. nature reviews cardiology, 17, 259–260. 6. angeli, f., spanevello, a., de ponti, r., visca, d., marazzato, j., palmiotto, g., feci, d., reboldi, g., fabbri, l. m., & verdecchia, p. (2020) electrocardiographic features of patients with covid-19 pneumonia. european journal of internal medicine, 78, 101–106. 7. lecun, y., bottou, l., bengio, y., & haffner, p. (1998) gradient-based learning applied to document recognition. proceedings of the ieee, 86, 2278–2324. 8. lee, h. & kwon, h. (2017) going deeper with contextual cnn for hyperspectral image classification. ieee transactions on image processing, 26, 4843–4855. 9. hershey, s., et al. (2017) cnn architectures for large-scale audio classification. in 2017 ieee international conference on acoustics, speech and signal processing (icassp), 131– 135, ieee. 10. pei, z., li, c., qin, x., chen, x., & wei, g. (2019) iteration time prediction for cnn in multi-gpu platform: modeling and analysis. ieee access, 7, 64788–64797. 11. zhai, x., jelfs, b., chan, r. h., & tin, c. (2017) self-recalibrating surface emg pattern recognition for neuroprosthesis control based on convolutional neural network. frontiers in neuroscience, 11, 379. 12. cheng, j., liu, y., & ma, y. (2020) protein secondary structure prediction based on integration of cnn and lstm model. journal of visual communication and image copyright © 2021 assa. adv syst sci appl (2021) convolution neural network based covid-19 screening model 39 representation, 71, 102844. 13. li, y., pang, y., wang, j., & li, x. (2018) patient-specific ecg classification by deeper cnn from generic to dedicated. neurocomputing, 314, 336–346. 14. degerli, a., ahishali, m., yamac, m., kiranyaz, s., chowdhury, m. e., hameed, k., hamid, t., mazhar, r., & gabbouj, m. (2021) covid-19 infection map generation and detection from chest x-ray images. health information science and systems, 9, 1–16. 15. kesim, e., dokur, z., & olmez, t. (2019) x-ray chest image classification by a smallsized convolutional neural network. in 2019 scientific meeting on electrical-electronics & biomedical engineering and computer science (ebbt), 1–5, ieee. 16. qiblawey, y., et al. (2021) detection and severity classification of covid-19 in ct images using deep learning. diagnostics, 11, 893. 17. khan, a. h., hussain, m., & malik, m. k. (2021) ecg images dataset of cardiac and covid-19 patients. data in brief , 34, 106762. copyright © 2021 assa. adv syst sci appl (2021) introduction method and material database preprocessing convolution neural network result and discussion conclusion adv syst sci appl 2019; 01; 31-43 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/649 activity of russian companies of the agri-food sector in regional industrial value-added chains veronika yu. chernova 1*, inna v. andronova 2, ekaterina a. degtereva 1,3,4, alexander m. zobov 1, vasily s. starostin 5 1) department of marketing, peoples' friendship university of russia, russian federation, moscow, russia 2) department of international economic relations, peoples' friendship university of russia, russian federation, moscow, russia 3) institute of europe of the russian academy of sciences (ras), moscow, russia 4) mgimo-university, moscow, russia 5) institute of marketing, state university of management, moscow, russia received september 24, 2018; revised march 12, 2019; published april 15, 2019 abstract: this article is concerned with the study of the factors, underlying the decision of russian companies to relocate production units to foreign countries and form global value-added chains (gvacs). it is proposed by the authors to consider the emergence of economic interests in relocation from the standpoint of synthesis of the concept of integration based on the natural advantages and the concept of gvac. a high potential for the development of cooperation by russian companies is noted herein, in particular, in the agro-industrial complex, and at the same time, there is a limited interest of russian business in the integration processes in the eaeu format. the factors that hinder the expansion of russian business to the partner countries in the eaeu, as well as the risks of deepening the cooperation ties, are singled out. keywords: industrial value-added chains, eaeu, industrial integration, cooperation ties. 1. introduction the problem of the prospects for the development of economies of the eaeu member states in the format of their participation in the eurasian integration project [1] to create four freedoms: the freedom of movement of goods, services, finance, and labor, is one of the most controversial nowadays [2-3]. the eaeu is a young integration project, which has been operating in the customs union environments since 2011, and in the economic union environments – since 2015. however, the key theme, from the point of view of enhancing macroeconomic stability and implementation of the competitive potential, is the formation of a coherent industrial, transport, energy and agrarian policy, and the deepening of industrial cooperation [4]. in the context of increasing global risks, the deepening of integration and the development of russia's cooperation with the eaeu countries are becoming ever more urgent. the eaeu was established to realize more completely the economic potential within the regional economic ties of the former single union state in the creation of the conditions for improving the competitiveness of the participating countries. currently, russia is aimed at modernization of the economy based on the technological breakthrough, the diversification, and expansion of the exports to world markets, the integration into global value-added chains along with the * corresponding author: veronica.urievna@mail.ru 32 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) partners in the eaeu. the obvious advantages of the international fragmentation of production, which are pushing the enterprises of the developed countries to participate in gvacs, carry quite significant risks in themselves. in this regard, it seems relevant to analyze the sources of the emergence of the economic interests of russian enterprises of the agri-food sector in the formation of and participation in regional value-added chains in the long term. 2. literature review the problem of the prospects for the development of economies of the eaeu member states in the format of their participation in the eurasian integration project [1] to create four freedoms: the freedom of movement of goods, services, finance, and labor, is one of the most controversial nowadays [2-3]. the eaeu is a young integration project, which has been operating in the customs union environments since 2011, and in the economic union environments – since 2015. however, the key theme, from the point of view of enhancing macroeconomic stability and implementation of the competitive potential, is the formation of a coherent industrial, transport, energy and agrarian policy, and the deepening of industrial cooperation [4]. according to the calculations presented [5], two-thirds of the cumulative integration effect in the form of gdp growth are potentially due to direct production effects resulting from the expansion of cooperation, technical and technological ties between the subjects of the eaeu, and the formation of intraregional value-added chains [6-7]. therefore, according to tkachuk, the coordinated economic policies of the eaeu member countries should be directed to the creation of such structures. under the value chain is understood the full range of operations for the creation and consumption of a product – from developing ideas to after-sales service. in the paper [8], the author has distinguished between producer-oriented value chains and buyer-oriented value chains. in later studies of value chains in agribusiness [9], it is stated that retailers and branded marketers are not the only buyers, and international traders and processors play a similar role in the value chains. buyers in different market segments have different requirements and organize different types of chains. in addition, not all chains have explicit coordinators. as gereffi [10] puts it, the type of value chain is determined by various combinations of three factors: the difficulty of transferring information and knowledge, especially with regard to product specifications and the process for making an inter-firm transaction; the extent to which this complexity can be reduced, and information and knowledge can be effectively transferred without making significant investment; the ability of actual and potential suppliers to fulfill customer requirements. value chain types include a wide range of variations from low levels of explicit coordination and power asymmetry between buyers and suppliers in case of market relations to high levels of explicit coordination and power asymmetry between buyers and suppliers in case of hierarchy, and intermediate chain types of relational, modular and subordinate types. humphrey [11] notes that in agribusiness, coordination through market relations is increasingly being replaced by coordination through the direct exchange of information between firms involved in the creation of the final product. the leading coordinating company should be able to provide instructions and monitor their implementation, make key decisions on the structure of manufacturing, inclusion in the chain and exclusion from it of particular suppliers, distribution of specific types of activities between the various participants in the chain. the fulfillment of such functions is connected with the presence of a certain market power in the leading firm. the asymmetry of market power in value chains identifies the links of the profit concentration [12] and, consequently, the concentration of resources for innovation and growth. activity of russian companies of the agri-food sector 33 copyright ©2019 assa adv. in systems science and appl. (2019) increasing concentration in individual links in the value chain and increasing quality and safety requirements for food, according to humphrey [11], are the basic trends in value chains in agribusiness. the participation in gvacs is facilitated by such factors as favorable business climate and the level of economic development. the international organizations define the tariffs and other trade restrictions as the barriers to participation in gvacs. in addition to the tariffs, the wto and the oecd identify three additional important barriers to participation in value-added chains: the inadequate infrastructure, the limited access to financing and non-compliance with the world standards. the costs associated with the introduction of new standards and the tightening of existing requirements have a significant impact on small firms included in the value chain. ollinger and moore [13] show that for small enterprises, this means a decrease in profitability and the need to either leave the industry or switch to other products. thus, food safety and quality requirements contribute to concentration in the value chains in agribusiness. concentration growth is marked in all links of the value chain. thus, in the paper [14], the concentration in the agrochemical and seed industries of the agro-industrial complex was studied, which according to the author is related to the protection of intellectual property. there are some works [15] which show concentration at the stage of making agricultural products and the stage of their processing, the main reasons for which are achieving the economy of scale, stricter food safety and quality requirements, the need for continuous development and implementation of innovation. moreover, the institutional environment, the business environment, the transport infrastructure and the qualification of the employees are the significant constraints. a key role in this is played by transportation costs. until recently, it was believed that distance was an important factor in the effective functioning of gvacs and the choice of suppliers. stephenson argues that the distance factor can be overcome by greater efficiency of the transport and logistics system [16]. preigerman [17] puts that the development of transport infrastructure and technologies, but not the geography of the country as before, determine the opportunities and the trajectory of the socio-economic development of countries. the benefits of participation in value chains have been studied by many scholars. thus, the research conducted by the uk companies [18] confirms the influence of participation in gvacs on the increase in productivity: a 10% increase in the external presence in the industry increases by 0.5% the overall factor productivity of the national producer of this industry. the studies of the influence of the integration of the companies in gvacs [19] showed that in the first year after integration in gvacs, the analyzed companies achieved +5% of the productivity advantage and +9% in four years. such significant differences in productivity growth can be associated with both the modernization of the company with its integration into the value chain and differences in the location of a link in the chain, ceteris paribus. at the same time, the companies that had left the global chains lost 1% of productivity during the first year, and the cumulative loss of performance over four years was a loss of productivity of 8%. according to baier & bergstrand [20], for international fragmentation of production, a significant obstacle is transport, tariff and non-tariff barriers. they showed that a 7.5% reduction in the tariff rates combined with a 5% reduction in the transportation costs contributes to vertical specialization (offshoring) by almost one-third. based on the study of german companies making decisions on offshoring, marin [21] noted the importance of not only the low wages and closeness to germany but also the importance of a lower level of corruption, better conditions for contracting. a study of more than 16,000 german companies [22] made it possible to identify the following factors (in the descending order) as very significant and important: linguistic and cultural barriers; institutional and administrative barriers; the cost-benefit ratio; distance to production sites; budgetary issues (taxes, etc.); interests of employees; business ethics; uncertainty regarding international 34 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) standards; risk of non-compliance with patent law; distance to major markets; the lack of local suppliers that meet the requirements of the company. according to heifets [23], the mainstreaming factors influencing the decision of russian companies to expand into foreign countries are the benefit and convenience of development. tkachuk [4] justifies the creation of sectoral regional clusters with export potential by the identification of “convergence zones” on the basis of analysis of sectoral priorities of the national economies, which are of interest at least for eaeu, which have competitive advantages in these sectors and correspond to the trends of the world economy. natural advantages and the available development potential can allow the eaeu agro-industrial sector to compete for the consumer in the world market [24] while maintaining the principle of security [25]. 3. materials and methods the aim of this paper is to study the possibility of developing cooperative ties between enterprises of the eaeu countries in the agri-food sector from the standpoint of the formation of regional value chains. in the study, the author relies on the concept of value chains in the agri-food business. the author distinguishes the main types of value chains: hierarchy, chains of subordinate, modular relational type and chains with market relations. the trends of their development in the agroindustrial sector detail the growing concentration in the individual links of the value chain and the increasing requirements for the quality and safety of food products, as well as the high significance of the efficiency of transport and logistics. the author wishes to show which types of chains are predominantly formed in the eaeu, who is the initiator of the formation of chains, how much the eaeu countries are involved in value chains and what obstacles exist for the development and deepening of cooperative ties between enterprises of different eaeu countries. the study is based on relevant empirical data, statistics and publications of the eec, edb cis, and the world bank for the analysis of which traditional methods and techniques of economic analysis were used in the form of the comparison of absolute, relative and average values, index method, and grouping methods. the use of mathematical analysis tools for research is difficult due to the lack of statistical information over a long period of time, the inability to level the impact of political factors on the integration processes and the development of cooperative ties, as well as the presence of factors distorting statistical data, such as illegal imports from third countries to the eaeu (for example, illegal imports from china to kazakhstan in 2017 amounted to more than $6 billion, in 2018 – $4 billion and exceeded legal imports [26]) and in the intra-union trade, the use of various tax evasion schemes by agri-food organizations. in this connection, in the study of the concentration of markets, the author used expert estimates, which in this situation are the most accurate. 4. results the agro-industrial complex is one of the most important strategic sectors of the economy of the eaeu countries. functioning in the eaeu format provides each of the five member states a number of advantages of a general economic nature [27]. 4.1. eaeu food safety requirements there is a unified regulatory and legal framework in the territory of the eaeu. the main legal instruments that establish mandatory requirements for products and processes of their life cycle are technical regulations [28]. in total, 47 technical regulations of the union were adopted, most of which have already entered into force, and unified requirements for safety and product quality cover already about 85% of all products traded in activity of russian companies of the agri-food sector 35 copyright ©2019 assa adv. in systems science and appl. (2019) the eaeu market. the requirements of the technical regulations of the eaeu are in many ways wider than those laid down in european standards. they can combine several dozen of such standards or european directives. technical regulations of the eaeu are applied not only to protect human life and/or health or to protect the environment. given the degree of risk of harm, they may contain special requirements for products, as well as related processes. currently, as shown by the global food security index [29], calculated by economist intelligence unit, of the three eaeu countries, for which the food safety index was calculated, the highest quality and food safety index (75.2 points and the 25th place) is in russia [29]. belarus and kazakhstan are far behind in terms of quality and safety, as well as in terms of accessibility, availability of resources and efficiency (fig. 1). source: [29] fig. 1. global food security index of russia, kazakhstan and belarus 4.2. concentration russia, being the largest economy of the union, produces the majority (about 85%) of food products of the eaeu, but the share of belarus (about 10%) and kazakhstan (4.27%) is also noticeable, the republic of armenia and kyrgyzstan together produce less than one percent (table 1). table 1. share of food production in the eaeu in 2017, % indicator armenia belarus kazakhstan kyrgyzstan russia eaeu food production, including: 0.58 9.98 4.27 0.31 84.86 100 processing and preserving of meat and making of meat products 0.21 10.42 2.29 0.13 86.95 100 processing and canning of fish, crustaceans and mollusks 0.13 7.58 0.88 0.02 91.38 100 processing and canning of fruits and vegetables 1.10 7.21 11.52 0.21 79.96 100 manufacture of vegetable and animal oils and fats 0.02 3.06 5.45 0.01 91.45 100 dairy production 0.55 18.83 4.02 0.43 76.17 100 making of bakery and flour products 2.01 5.32 5.54 0.74 86.39 100 beverages 2.13 6.82 6.60 0.67 83.79 100 source: compiled and calculated by the author according to [30]. 67 71 61 75 73 66 68 63 67 63 58 66 51 58 68 40 45 50 55 60 65 70 75 80 total accessibility availability quality and safety natural resources and efficiency russia belarus kazakhstan 36 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) in the eaeu countries in agriculture, the share of private farms of the population, which are small producers, is traditionally high. in kyrgyzstan and armenia, the share of commercial organizations is less than 5%, in kazakhstan – less than ¼, in the republic of belarus in agriculture the share of the public sector in the form of collective farms is very high (about 80%) and the sector of commercial organizations is completely absent (fig. 2). source: compiled by the author according to [30]. fig. 2. the structure of agricultural production by business types (as a percentage of the total, at current prices), 2017 taking into account the above, the concentration on the eaeu market is considered in the case study of russia. the russian meat market is characterized by a high concentration of production facilities mainly in the european part of russia; the meat industry is mainly represented by diversified holdings that produce poultry, pork, and beef and build a vertically integrated value chain. in the meat market, the share of the largest players continues to increase, which together produced 3 million tons of poultry meat, 1.7 million tons of pork and over 100 thousand tons of beef, which is 60% of the total national agricultural production and 46% of the total national meat production in the country. at the same time, more than a quarter of the national production of all kinds of meat and about 37% of poultry accounted for the five largest meat producers [31]. the highest concentration of production facilities and the highest competition is observed in the poultry industry, although, according to experts, compared to brazil, the united states or europe, the share of leading players in the russian poultry market is not high enough being 10-12% [31]. in the pork market, the 20 largest companies account for just over 60% of the total russian industrial output. according to experts, competition in the pork market will intensify in the coming years due to the implementation of large projects in this area. the beef market is significantly different from the pork and poultry markets. the largest producer accounts for about 8% of the total national output, and more than 60% of this type of meat is produced by private farms. at the moment, beef cattle are not attractive to investors, even if there is a demand for this type of meat due to the long payback periods over 12 years. vertical integration along the value chain from manufacturing to sales is a general development trend in vegetable production, especially greenhouse vegetables. high concentration is noted in other markets of the agro-industrial sector. thus, in the sugar market, a quarter of sugar beet processing and sugar production accounts for the 10 largest sugar factories in terms of daily output. high concentration is also observed in those parts of the value chain that are close to the buyer. for the past 15 years, large modern retail chains have steadily increased their share in 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% armenia belarus kazakhstan kyrgyzstan russia agricultural organizations commercial organizations private farms activity of russian companies of the agri-food sector 37 copyright ©2019 assa adv. in systems science and appl. (2019) the russian market, gradually crowding out small business and open markets, especially in large cities in the european part of the country. in 2016, retail chains accounted for 27.7% of the total turnover in the sector, in the food category – 33.1%. currently, large russian retail chains are beginning to develop value chains and invest in food production and agriculture. 4.3. transport and logistics of the eaeu the formation and development of value chains and the fragmentation of production at the level of regional specialization objectively lead to the need for the uninterrupted flow of raw stock, semi-finished products and services through customs territories along the gvacs on the way to the creation of finished products, increasing the requirements for the transport and logistics system. an important characteristic of the state of development of transport and logistics is the length of the existing communication lines. according to the eec, at the end of 2017, the total length of the existing eaeu communication lines is 1,712.8 thousand km of public roads (about 2.5% of the world total), 109.7 thousand km of railway tracks (about 7.8% of the world total) and 287.8 thousand km of pipelines. at the same time, the increase in the length of communication lines in 2013-2017 was provided mainly by the increase in the length of the communication lines of the russian federation. for example, the increase in the length of roads in russia in 2013-2017 (111.3 thousand km) exceeded the current length of roads in kazakhstan (102 thousand km) and belarus (95.4 thousand km). a generalized indicator of the effectiveness of the national transport and logistics systems is the logistics performance index (lpi). from 2007 to 2018, all the countries of the eaeu improved the position of their transport and logistics in the international lpi global ranking [32-33]. in general, the eaeu transport and logistics are characterized by the low lpi index, which testifies to low competitiveness and insufficient efficiency in comparison with other countries (fig. 3). the lack of a unified transport and logistics market of the union as such [34] is the result of the effect of the emerging consolidation of transport and logistics markets in these countries. source: compiled by the author based on the world bank data [33]. 2,0 2,5 3,0 3,5 4,0 4,5 5,0 customs infrastructure international shipments logistics competence tracking & tracing timeliness kazakhstan russian federation armenia belarus kyrgyz republic china germany 38 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) fig. 3. the logistics performance index of the eaeu countries, china and germany in 2018 currently, the issue of modernization of transport and logistics is a priority for the development of the eaeu. in russia, for example, the volume of investment in “transportation and warehousing” and “information and communications” in 2017 amounted to 18.3% and 3.5% of the total investment in the economy, in kazakhstan – 14.4% and 0.9%, in belarus – 10% and 3.2%, in armenia – 12.2% and 3.5%, in kyrgyzstan – 12.6% and 2.5%, respectively [35]. in general, the business conditions in the eaeu member countries are approximately the same, with the exception of the kyrgyz republic. so, according to the world bank, in the eastern europe-central asia region (24 countries), in 2018 russia ranks 4th, being inferior to georgia, latvia, macedonia. kazakhstan and belarus are the 5th and the 6th, respectively, armenia is the 12th and the kyrgyz republic is the 21st (fig. 4). source: [36] fig. 4. business conditions in the eaeu member countries and the rf in comparison 4.4. investment russia and kazakhstan are the main investors of mutual direct investments in the eaeu ($1,162 and $375 million, respectively), including in transport and logistics. at the end of 2017, russia accounted for about 75.6% of the exported volume of mutual investments in the eaeu countries. during the period from the 2nd quarter of 2017 to the 2nd quarter of 2018, the volume of mutual investments in the eaeu decreased by $1.2 billion, and the share of russia increased to 95.34%. mutual investments of the eaeu countries in the agro-industrial sector remain low. the agri-food sector of the eaeu countries is not interesting for russian investors. so, mutual (mainly russian) investments in the agro-industrial sector of kazakhstan make up 1.5% of the total investments, in the agro-industrial sector of belarus – 0.8%, while 0 5 10 15 20 25 30 st ar ti n g a b u si n es s d ea li n g w it h c o n st ru ct io n p er m it s g et ti n g e le ct ri ci ty r eg is te ri n g p ro p er ty g et ti n g c re d it p ro te ct in g m in o ri ty i n v es to rs p ay in g t ax es t ra d in g ac ro ss b o rd er s e n fo rc in g c o n tr ac ts r es o lv in g in so lv en cy kazakhstan belarus armenia kyrgyz republic russian federation activity of russian companies of the agri-food sector 39 copyright ©2019 assa adv. in systems science and appl. (2019) investments of the eaeu countries in the russian agro-industrial sector exceed 15% of the total investments (fig. 5). source: [37] fig. 5. the structure of mutual investments in the eaeu countries in 2017 such a result is explained not only by the large volume of the russian market and its unsaturation but also by the policy of import substitution and support for the agro-industrial sector in russia. the agri-food sector of russia is in the second place (15.8%) in terms of the attractiveness for the eaeu investors after the chemical sector (35.1%). the main volume of investments was provided by kazakhstan companies specializing in crop and dairy production [37]. 5. discussion as a result of the policy of import substitution and significant governmental support for domestic manufacturers, the russian food markets are currently saturated with a number of products (for example, poultry, pork). further production growth of the largest manufacturers is possible only at the expense of exports growth, since the volume of production already exceeds consumption. in addition, the growth of production is constrained by the increase in production cost against the background of higher prices for imported fodder components [31]. increasing competition, according to experts, will encourage manufacturers to look for ways to reduce costs, including via transferring part of their production facilities to the eaeu countries with cheaper labor and more suitable climatic conditions for agriculture. in addition, for russian manufacturers already interested in exports development, it may be interesting to place enterprises closer to end users – the countries of southeast asia. the development of vegetable markets, especially greenhouse vegetables using photoculture technology in the winter, is unprofitable for a number of vegetable crops, especially in the central, northwestern and eastern regions of russia due to high heating and lighting costs for greenhouses. cultivation of vegetable crops in the eaeu countries with a warmer climate and a greater number of sunny days would be economically justified subject to the development of the eaeu transport and logistics infrastructure. according to pak [34], 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% fdi to russia from the eaeu fdi to belarus from the eaeu fdi to kazakhstan from the eaeu fdi to armenia from the eaeu fdi to kyrgyzstan from the eaeu agri-food sector fuel ferrous metallurgy nonferrous metallurgy machinery and equipment infrastructure chemistry construction transport communications and it wholesale and retail finance tourism other industries 40 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) the low efficiency of the infrastructure component of eurasian integration, including transport, warehousing, customs, management, telecommunications, etc., in the long run can have a no less divisive role in eurasian integration than the preservation of non-tariff barriers in mutual trade. in the analysis of the domestic market of the eaeu on the basis of the market research of meat products [38], it was found that veterinary, sanitary and phytosanitary measures are the main obstacles to the access of belarusian, kazakh and kyrgyz exporters to the russian market. at the same time, the population of the eaeu countries is quite loyal to the import of goods of russian origin, especially in kazakhstan, tajikistan, and kyrgyzstan, while the republic of belarus is the most interested in foreign investment. in general, more than a third of the population of the eaeu countries supports the policy of rapprochement with russian business. russia is also the most attractive country on the part of the eaeu countries for cooperation in the field of science and technology, in joint development and partnership with which about half of the population of tajikistan, belarus and kyrgyzstan are interested [39]. the edb analysts believe that the deepening of eurasian integration in the eaeu and the establishment of common markets will gradually change the situation of corporate interaction. in the meantime, the benefits of membership in the eaeu are mainly used by large investors, for which cross-border barriers are less painful and their resources allow them to effectively overcome these barriers. the medium business tends not to go abroad so far [40]. in the opinion of the edb cis, the general factors preventing the integration and development of cooperation between the russian business and the partners in the eaeu is the presence of non-tariff barriers in the mutual trade of the eaeu countries, the excessive administrative and tax burden on business, the low level of financial support for small and medium-sized businesses, the interest rates on loans, as well as underdevelopment of cooperation between large business and small and medium-sized enterprises. for russian enterprises, the main risks related to the formation of regional value-added chains in the eaeu format are the following: national protectionism, trade wars between the individual member countries, currency conflicts, the macroeconomic instability and the existing system of economic and political relations both within the eaeu and between the individual countries [23]. the value chains within the eaeu format are currently not sufficiently studied. there is a very small amount of studies in this area. in particular, pobyvayev [41] in his research came to a conclusion that value chains, being a tool for increasing the degree of economic integration of regional economies, would benefit russian-belarusian integration, and the most effective way to participate in global value chains at this stage would be to embed small and medium businesses of the two countries into them. 6. conclusion the use of the eaeu format for the formation of regional value-added chains seems to be economically profitable for russian business. the agribusiness sector has a high potential for the development of cooperation ties, in a number of which the cooperation is a necessity and a condition for further development. despite this, russian business does not currently demonstrate an increased interest in the development of joint projects and cooperation ties, which is hampered by a number of factors. this study showed that regardless of some criticism of the very possibility of forming value chains in a national territory, they have already been formed. the predominant type of value chains is a hierarchical type (vertically integrated agri-holdings), which is associated with the high cost or impossibility of coordinating the manufacturing, monitoring, and control with a different organization of the production process in conditions of stricter safety requirements activity of russian companies of the agri-food sector 41 copyright ©2019 assa adv. in systems science and appl. (2019) and product quality. however, the author has not revealed widespread value chains with any enterprises of the eaeu countries. from the position of value chains, the factors hindering the development of cooperative ties and the formation of joint value chains in the agri-food sector of the eaeu are:  discrepancy between the quality and level of food safety of belarus, kazakhstan, kyrgyzstan and armenia with the requirements of technical regulations of the eaeu and the russian federation, which, if enterprises of these countries are included in the value chains of russian agricultural holdings, will significantly increase their costs on monitoring and controlling the quality of products and manufacturing processes;  mainly private farms’ economy in the structure of agricultural production in armenia, kyrgyzstan, kazakhstan – very small commodity producers and an extremely low share of commercial agricultural organizations, which, from the standpoint of the value chain management concept, also causes the growth of costs of agricultural holdings on monitoring and control;  missing sector of commercial organizations in the structure of agricultural production in belarus with a high proportion (about 80%) of state organizations (collective farms of the old soviet type) with an inefficient management system;  insufficient efficiency of the existing transport and logistics of the eaeu, which could ensure timely delivery of perishable products from the place of manufacturing to the processing sites or end users. acknowlegements the publication has been prepared with the support of the “rudn university program 5-100” in the frame of the project “improvement of marketing tools to support and expand the import of consumer goods in the real sector of the russian economy”. references [1] zhiltsov, s.s. (2016). eurasian integration: problems and prospects. journal of the peoples' friendship university of russia. series: political science, 1, 7-15. [2] vinokurov, e.yu., & tsukarev, t.v. (2015). economy of the eaeu: the agenda. journal of eurasian economic integration, 4, 7-21. [3] pereboyev, v.s. (2017). eaeu for business: opportunities and tools. cis edb. retrieved september 21, 2018, from http://eurasian-studies.org/archives/6372 [4] tkachuk, s.p. (2016). from the cu – eeu to the eaeu itself: on the harmonization of the industrial policies of the member countries (conceptual considerations). russian economic journal, 1, 54-65. [5] glazyev, s.yu. (2011). the real core of post-soviet economic integration: the results of the creation and prospects for the development of the customs union of belarus, kazakhstan and russia. russian economic journal, 6, 78-81. [6] fomina, a.v., berduygina, o.n., & shatsky, a.a. (2018). industrial cooperation and its influence on sustainable economic growth. entrepreneurship and sustainability issues, 5(3), 467-479. https://doi.org/10.9770/jesi.2018.5.3(4) [7] mikhaylov, a.s. (2018). socio-spatial dynamics, networks and modelling of regional milieu. entrepreneurship and sustainability issues, 5(4), 1020-1030. https://doi.org/10.9770/jesi.2018.5.4(22) [8] gereffi, g. (1994). the organisation of buyer-driven global commodity chains: how u.s. retailers shape overseas production networks. in g. gereffi, & m. korzeniewicz (eds.), commodity chains and global capitalism (pp. 95-122). westport: praeger. [9] ponte, s., & gibbon, p. (2005). quality standards, conventions and the governance of global value chains. economy and society, 34(1), 1-31. [10] gereffi, g., humphrey, j., & sturgeon, t. (2005). the governance of global value chains. review of international political economy, 12(1), 78-104. 42 v.yu. chernova, i.v. andronova, e.a. degtereva, a.m. zobov, v.s. starostin copyright ©2019 assa adv. in systems science and appl. (2019) [11] humphrey, j., & memedovic, o. (2006). global value chains in the agrifood sector. retrieved february 17, 2019, from https://www.unido.org/sites/default/files/200905/global_value_chains_in_the_agrifood_sector_0.pdf [12] milberg, w. (2004). the changing structure of trade linked to global production systems: what are the policy implications? international labour review, 143(1-2), 4590. retrieved february 17, 2019, from https://static1.squarespace.com/static/53ce7840e4b01d2bd01192ee/t/53e8f8a5e4b0b05 3addadd1f/1407776933063/changing-structure-of-trade-linked-to-globalproduction-systems.pdf [13] ollinger, m., & moore, d. (2006). the economic forces driving the costs of food safety regulation. retrieved february 17, 2019, from https://www.researchgate.net/publication/23506496_the_economic_forces_driving_th e_costs_of_food_safety_regulation [14] srinivasan, c.s. (2003). concentration in ownership of plant variety rights: some implications for developing countries. food policy, 28(5-6), 519-546. [15] humphrey, j., & schmitz, h. (2004). governance in global value chains. in h. schmitz (ed.), local enterprises in the global economy (pp. 95-109). cheltenham: edward elgar. [16] stephenson, s. (2013). global value chains: the new reality of international trade. geneva, switzerland: ictsd & wef. retrieved february 17, 2019, from http://e15initiative.org/wp-content/uploads/2015/09/e15-gvcs-stephenson-final.pdf [17] preigerman, ye. (2018). infrastructure connectivity and political stability in eurasia. valdai notes, 85, 16. [18] haskel, j.e., pereira, s.c., & slaughter, m.j. (2007). does inward foreign direct investment boost the productivity of domestic firms? review of economics and statistics, 89(3), 482-496. [19] baldwin, j., & yan, b. (2014). global value chains and the productivity of canadian manufacturing firms. economic analysis research paper series. ottawa: statistics canada. [20] baier, s., & bergstrand, j. (2000). the growth of world trade and outsourcing. notre dame university. [21] marin, d. (2006). a new international division of labor in europe: outsourcing and offshoring to eastern europe. journal of the european economic association, 4, 612622. https://doi.org/10.1162/jeea.2006.4.2-3.612 [22] godard, o., & görg, h. (2011). the role of global value chains for german manufacturing. retrieved september 21, 2018, from https://ssrn.com/abstract=2179948 or http://dx.doi.org/10.2139/ssrn.2179948 [23] heifets, b.a. (2015). eurasian economic union: new challenges for business. society and economy, 6, 5-22. [24] kalenova, s.a. (2017). on the joint research project of kazakhstan and russian scientists "integration effects of economic interaction within the framework of the eurasian economic union. in russia: trends and prospects of development (pp. 286290). moscow: institute of scientific information in social sciences, russian academy of sciences. [25] fyodorov, m.v., & kuzmin, e.a. (2013). agriculture and economic security of russia: retrospective research. journal of international scientific researches, 5(1-2), рр. 42-45. [26] mendkovich, n. (2019). “gray import” to the eaeu countries: how to deal with smuggling. eec. eurasian studies. retrieved february 17, 2019, from http://eurasianstudies.org/archives/11251 [27] mikhailovskaya, s. (2018). cross-border cooperation of agrarians. belorusskaya dumka, 1, 24-31. activity of russian companies of the agri-food sector 43 copyright ©2019 assa adv. in systems science and appl. (2019) [28] nurgaliyeva, m.t., smagulov, a.k., & iskakova, zh.a. (2016). issues of regulating the quality and safety of food products within the framework of the european and eurasian economic union. science and world, 3(31), 86-91. [29] global food security index. (2018). the economist. retrieved february 17, 2019, from https://foodsecurityindex.eiu.com/country/details#russia [30] eurasian economic commission. (2018). agro-industrial sector. statistics of the eurasian economic union: statistical bulletin. moscow. (p. 132). [31] kulistikova, t. (2018, august 7). top 25 largest russian meat producers. market leaders will continue to consolidate. agroinvestor. retrieved february 17, 2019, from https://www.agroinvestor.ru/rating/article/30208-top-25-rossiyskikh-proizvoditeleymyasa/ [32] international lpi global ranking 2007. (2007). retrieved february 17, 2019, from https://lpi.worldbank.org/international/global/2007 [33] international lpi global ranking 2018. (2018). retrieved february 17, 2019, from https://lpi.worldbank.org/international/global/2018 [34] pak, y. (2014). prospects and challenges of transport and logistics cooperation in the eurasian economic union. journal of strategic and international studies, 6, 73-79. [35] eurasian economic commission. (2019). socio-economic statistics. investments. retrieved february 17, 2019, from http://www.eurasiancommission.org/ru/act/integr_i_makroec/dep_stat/econstat/p ages/investments.aspx [36] the world bank. (2018). doing business. measuring business regulations. retrieved september 21, 2018, from http://www.doingbusiness.org/rankings?region=europe-andcentral-asi [37] edb center for integration studies. (2017). monitoring mutual investments in the cis countries. st. petersburg: cis edb. (p. 60). [38] eurasian economic commission. (2017). the conjugation of the eaeu and the sreb takes real shape: the list of infrastructure projects has been agreed. retrieved february 17, 2019, from http://www.eurasiancommission.org/ru/nae/news/pages/2-032017-1.aspx [39] edb centre for integration studies. (2017). integration barometer of the edb. st. petersburg: cis edb. (p. 108). [40] edovina, t. (2017). capital returned to its neighbors. mutual investment in the eaeu proceeded to growth. retrieved february 17, 2019, from https://www.kommersant.ru/doc/3385720 [41] pobyvaev, s.a. (2016). global value chains and their potential role in the development of russian-belarusian integration. the world of new economy, 4. retrieved february 17, 2019, from https://cyberleninka.ru/article/n/globalnyetsepochki-stoimosti-i-ih-potentsialnaya-rol-v-razvitii-rossiysko-belorusskoy-integratsii на рассматриваемом множестве приоритетных проектов после нахождения решения выпуклой задачи распределения ресурсов в соответствии с общим объёмом ассигнований , достигнутой вероятностной оценкой реализуемости вместе с экспертными приоритетами adv syst sci appl 2019; 04; 66-78 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/827 a planning and scheduling method for large-scale innovation projects vladimir topka*, anatoliy d. tsvirkun trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: topka3@mail.ru, tsvirkun@ipu.ru received may 16, 2019; revised december 12, 2019; published december 31, 2019 abstract. we consider the problem of uniform allocation according to the minimax criterion of a non-renewable resource of an innovative project's activities. a lexicographic method is used for ordering these minimax criteria to solve the problem one objective function developed a greedy heuristic for finding the maximum path on an acyclic digraph with double weights on arcs. the greedy heuristic gives approximate solution of the problem mentioned. the obtained solution allows building a project resources plan and schedule with four types of temporary slacks with inaccurate initial data. the proposed approach allows performing an analysis of the project risk of performance parameter, for which has not developed the apparatus for its analysis. these include the risks that the project when complete fails to perform as intended or fails to meet the mission or business requirements that generated the justification for the project. performance risks are estimated on the probability distribution function of a favorable/unfavorable outcome. proposed lexicographic disintegration is a decomposition procedure and is suitable for planning and scheduling large-scale projects. practical calculations for the convex problem in the initial statement may be performed with the help of ready-made software, and the obtained greedy solution of the problem-consequence has theoretical significance in graph theory. keywords: large-scale innovative project, planning, scheduling, reliability index of project activities, greedy algorithm, the maximum path with double weights on arcs. 1. introduction in existing project management software systems, some modules perform risk of cost and risk of schedule assessment. on the other hand, an approach based on the concepts of the theory of reliability can be proposed for modeling uncertainty and risk of innovative projects. in the theory of reliability, the indicator of network reliability is determined by the probability of its connectivity. this requires that at least one path is found connecting the beginning and the end of the network. however, in managing projects in the pert network, in order to ensure reliable execution of the project, it is necessary to perform all activities included in the project. in this case, the task is to assess not only the connectivity of the beginning and the end of the network, but also the completeness of the whole project. in present project management software, a quantitative assessment of risk of cost and risk of schedule is made on the basis of monte carlo simulation. the proposed approach allows one the analysis of the risk of performance [1] of a project for which the apparatus of its analysis has not been developed. the execution risk, upon completion, accomplish the required mission to achieve the specified technical characteristics, is based on the probability distribution function of a favorable/unfavorable outcome. for her: the probability of success of the project is the probability that the execution of this project will not fail, and the probability of reliable execution is taken as an index of the reliabil * corresponding author: topka3@mail.ru mailto:tsvirkun@ipu.ru a planning and scheduling method for large-scale innovation projects 67 copyright ©2019 assa adv. in systems science and appl. (2019) ity of implementation. the advantage of the proposed approach is that it provides a tool for working in an unexplored area. for activities of innovative projects with a high degree of uncertainty, to obtain the characteristics of random variables of their implementation the likelihood of achieving the goal or probability of technical success, that is, to build the distribution function of such a random variable, you can use empirical data or data obtained as a result of an active experiment in which the monte carlo method obtained the probability distribution function of the technical success of the activity, the degree of achievement of the specified performance characteristics of this development, in the form of a monotonously increasing with the saturation of the function of expended resources [2]. the likelihood of performing the activity by a particular time characterizes the possibility of attaining specific technical indicators of the result of the activity, provided that the previous activity, ensuring its beginning, has been completed. this probability depends on the amount of those or other resources spent, which have a value expression as cost. the approach developed in the article is aimed at developing quantitative methods for assessing and optimizing the project parameters, taking into account the criterion of uniform allocation of a non-renewable (stored) resource with a restriction on the assessment of project reliability index. 2. indicators of the reliability of activities and the project as a whole previously, integer, linear and some non-linear models were used to describe innovative projects execution [3]. we will model [4] the dependence of the probability of technical success (reliability index) of the project’s activities on the homogeneous non-renewable resource spent by 2-parameter power-concave functions of the form     α 0 1/ α 0 ( ) ε, 1 ,ε 0, ε, , 1,..., , j j j j j j j j u p u u u u j n          (2.1) where 0 α 1j  is the form parameter, 0 ju > 0 is the scale parameter. which at the late phases of an innovation project, when it is advisable to use quantitative methods to assess the uncertainty of the project, quite visually and precisely approximates the practical [2; 5] functions of the activities reliability of the innovation project. let an acyclic directed graph g (e, γ) be given, where e is the set of vertices corresponding to the project’s activities (activity-on-node representation), and г is the set arcs defining partial order relations of direct technological activity precedence. the network g (e, γ) has one final and one initial vertex; n final vertex the last number among all the vertices of the project and the dummy vertex 0, which means the launch of the project. thus, we have a network g (e, γ), at the vertices of which a development reliability index is given that satisfies the relation (2.1). we will set the reliability of the project by the probability of its technical success (the degree of achievement of the specified performance characteristics of the project) by a certain allowable period, the assessment of the reliability of the project in the most unfavorable case. in the theory of reliability, such an assessment is determined by the weakest link, i.e. the worst of the technological chains from the vertex of the lower level to the final vertex. 68 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019)   α 0μ ( , ) μ 0 0 ( ) min ε, 1 , ε 0, α (0,1), [ε , 1], 0; μ , j i i j n jk j k j j jk j i u p u w u w u          (2.2) where 0 0 [ε , 1], ε 0jkw   is arc (j, k) reliability coefficient, (from the matrix of paths along the arcs of the network g (e, г)) transfer the result of the j-th activity by the arc (j, k) to perform the k-th activity (like the transfer function of connecting links in theory of automatic control). we denoteμ i the path with the number 1,...,i m on the network from the initial vertex to the final vertex n. when solving problems of project planning in terms of a power model [6], the time variable is not taken into account. in this article we will consider a deterministic network model of the project with disjunctive input arcs or, or with conjunctive input arcs and having the property that for the project reliability indicator is performed assessment of lower bound (2) [6]. thus, we will speak about pi from (2.1) as a reliability index of activity, and about pn from (2.2) as an assessment of indicator of guaranteed reliability of project performance. in the model for representing network constraints, we will use the matrix of paths (along arcs) of the network g (e, г): which is based on the list of arcs of the network as the characteristic function of the constraint system. in the network model of the project g (e, г), each path from the initial vertex to the final one is uniquely represented by a line of a specially constructed matrix of paths (along arcs)    , , μj k i l j k  , in which its elements  arcs stand in the i-th line only if (j, k) arc belongs to i -th path. all other cases when  , μ i j k  there are empty cells. 3. the problem of successive uniform allocation of nonrenewable resource by chebyshev criterion let  1 2u= , , ..., nu u u be the vector of a non-renewable resource allocated for the execution of all activities of the project g (e, г). as an objective function for the allocation of resources, we will use uniform, or chebyshev criterion in the form u max inf , 0, 1,..., .j j uj e u u j n     (3.1) a problem of the form (3.1) is usually called a discrete minimax problem. therefore, this discrete minimax problem is considered in the form of the following problem of smooth conditional minimization 0 0 ( , ) inf , ju u u u       0 α 00 ( , ) μ 1/ α 0 0 , 1,..., , ε, 1 , μ , 1,..., , ε, , 1,..., . j i j j j jk i j k j j j u u j n u u w p i m u u u u j n                       (3.2) by logarithmizing the second group of constraints of the problem (3.2) and taking into account the first one, we get a planning and scheduling method for large-scale innovation projects 69 copyright ©2019 assa adv. in systems science and appl. (2019) 0 0 ( , ) min , ju u u u   (3.3) 0 0 ( , ) μ ( , ) μ0 ( , ) μ 0 ln ln ln ln , α . i i i j jk j k j k j j k j u w p u u u u                 (3.4) let the constraints of the problem be consistent, both here and in the future .u  due to inequality (3.4), the consequence of the original formulation, taking into account the network specifics of the problem, we will seek its solution in the form of the maximum path with double weights, when the optimal solution of the problem (3.3), (3.4) is realized on a path (possibly a set of paths) with a value 0λ such that 0 0 ( , ) μ ( , ) μ0 μ ( , ) μ ln ln ln λ max α i i i i j jk j k j k j j k u w p          ( , ) μ μ ( , ) μ σ max . α i i i jk j k j j k      wherein, 0 0 0 ( , ) min expλ ju u u u   . therefore, it is necessary to find (possibly, a set) of maximal paths with double weights on arcs in an acyclic directed graph g (e, γ).     0* μ * 0 0 0* 0* μ max λ μ , exp λ exp λ μ , μ i i j arg u u j      we assume that the parameters of the problem are such that here and in the future 0 u u it is satisfied. received resource critical path, with the maximum consumption of a non-renewable resource. let  0* μ μ max λ μ i iarg are found. the optimal solution of (3.3), (3.4) we fix:   * 0 0* 0 0* fixed, μ , μ fixed. ju u j e j j       and we consider the problem of allocating a non-renewable resource according to a successively applied minimax criterion, varying uncommitted variables. in this case, the task of the second stage is: 0 u max inf , 0, 1,..., .j j uj e e u u j n      or     1 1 ( , ) 1 α α* 00 ( , ) μ 0 inf , , 1,..., , ε, 1 , , μ , 1,..., . j j j i u u u j j j jk j k j j i u u u j n u u u w p u u u i m                    70 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019) by logarithmizing the second group of constraints of the problem and taking into account the first one, we get 0* 0 * 0 ( , ) μ ( , ) μ μ1 ( , ) μ ln ln ln α ln ln α i i i j jk j j j k j k j j j k u w p u u             0* 0 * 0 ( , ) μ ( , ) μ μ1 μ ( , ) μ ln ln ln α ln λ max α i i i i j jk j j j k j k j j j k u w p u                1 ( , ) μ 1* 1 μ μ ( , ) μ σ max , μ max λ μ . α i i i i jk j k i j j k arg       we also fix the optimal solution to the problem of the second stage:  1* 1 1 1* 1* 0* 1 exp λ μ fixed, μ \ μ = fixed. ju u j e      at the same time, we gradually expand the viewer set  1 0j e e e    . next, we consider the minimax problem of the third stage, the optimal solution of the previous stages is fixed, and the unfixed variables vary.  1 0 u max inf , 0, 1,..., .j j uj e e e u u j n       we continue to allocate resources according to the minimax objective function successively until at some stage q (the set indicator) we get the domain of definition of the problem  0... ... q re e e e     and   * * * 1 * 0 exp λ μ fixed, μ \ μ fixed, q q q q j q q q r u u e j j               0 q re e . (3.5) the overall problem is solved by successive optimization of the remaining sub graph by the minimax objective function at each stage, as in the lexicographic method of organizing the solution of multi-criteria optimization problems. when the whole set of vertices is covered in optimal ways, this will solve the problem of a successive, monotonically expanding the scanned area, uniform allocation of a non-renewable resource according to a minimax criterion with a restriction on the assessment of reliability index of the project in the worst case. this procedure converges in a finite number of stages. such a lexicographic disintegration, in essence, is a decomposition procedure and is suitable for the planning of large-scale projects. to find the path of maximum efficiency in [7, p. 13], an algorithm of a search type was proposed (with exponential complexity) that reduces to finding the maximum path in the network. in [8, p. 192], to find in the graph with double weights of the cycle with the minia planning and scheduling method for large-scale innovation projects 71 copyright ©2019 assa adv. in systems science and appl. (2019) mum value, the detection algorithm in the graph of the negative weight cycle is used (provided that for all cycles the sum of weights in the denominator is positive). the solution of this problem with double weights using the proposed algorithm requires 3 1 log η o e       operations, where η the magnitude of the error (weakly polynomial algorithm), e is the number of vertices in the network. in [9], a general method was proposed for solving the average problem in linear spaces. from it, as a special case, follows a strong polynomial algorithm for finding the minimal average contour in a weakly connected oriented graph, which complexity is  3 o e г . to solve the problem of finding the maximum path with double weights in the acyclic digraph g (e, г), we propose a greedy heuristic, for which, as will be shown below, the estimate  2 o г of its computational complexity is valid. the advantage of which, compared with the above, is lower computational complexity. 4. greedy algorithm of finding the maximum path in digraph with double weights on arcs construction of a set of paths  sm  , with weights ji and j on arcs lji , (j,i)  г of the digraph g (e, γ) covering all its vertices e. the initial step. watched: arcs :l  , vertices :e  , path number s: = 1. 1). selection of an arc on a set of unvisited arcs г =г\l. in the acyclic digraph g (e, γ), we choose an arc jil , where  ,j i г is optimal by the local criterion ( , ) σ max . α ji ji j i г j l arg   in the array l we bring the arc jil : : jil l . 2). build a path. 2.1). adding new arcs. construct a path s from the initial dummy vertex j: 1 jг   to the final vertex j = n: jг  by local criterion, i.e. for an arc jil , adjacent arcs kjl and itl are selected from conditions 1 ( , ) ( , ) σ σ max , α α σ σ max , α α j i kj ji j i l kj k г k j j l ji it j i l it t г j i j l l arg l arg                  which are added to the existing arcs : jil l , forming a path  ... ...s kj itl l l    with double weights. where ( , ) σ ji j i l  and α j j l  values for the path found in the previous step. 72 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019) putting μ : 1s  , for the constructed in such a way the track μ s , we define the value  λ μs s , s = 1,2, ...  ( , ) μ μ μ ( , ) μ σ max λ μ α ji j i s s s j j i       . list of viewed arcs  : ... ...kj itl l l l   ; for viewed vertices – es= {j|j  s}; e' := e'  es. 2.2). avoiding algorithm loops on scanned paths. in order to prevent the algorithm from looping along the scanned path μ s , in the future we will calculate the indicator value 1 μ s . if for non-viewed vertices \e e  , then let \k e e we put 1, if , δ 0, if \ . k k e k e e     then μ μ 0 0 δ 1, ; δ 0, if μ : \ . k k k k k e k k e e             in this case, we set μ1 μ μ , if δ 1, 1, if δ 0, k ks k k           and then in the future, we will sequentially determine  ( , ) μ 1 1 μ 1 μ ( , ) μ σ max λ μ α ji j i s s s j j i           1 1λ μs s     only for new paths 1μ s that do not coincide with μs , s = 1,2, ... 3). cycle if for unvisited arcs г   , then cycle through : 1s s  and go to step 1. otherwise: viewing the constructed set of paths  1μs and determining the optimal path(s) with double weights on arcs 0*μ argmaxλ (μ )s s s  , s = 1,2, ... for the first stage (3.3), (3.4) of the problem:  0 0*λ λ μ , and for the subsequent ones  *λ λ μr r , r = 1, ..., q. in the algorithm, the review of arcs for constructing a set of paths is carried out from г to  , which allows you to avoid a complete search of the paths, and the introduction of a counter μ s   for paths that have at least one unwatched arc eliminates looping the algorithm on the same old path. described algorithm is, in essence, greedy heuristic and so it gives approximate solution of the above problem. the complexity of the algorithm described is estimated as follows. the selection of the initial locally optimal arc lji requires г operations, building a path  from the initial vertex to the final vertex is also requires the order of d г operations, where d is the maximum degree of the vertices of the graph g (e, γ). the construction of all such paths for each of the г arcs is repeated about  о г d г г   time plus the successive selection of the λ-optimal path, which will require order г operations. therefore, the complexity makes a planning and scheduling method for large-scale innovation projects 73 copyright ©2019 assa adv. in systems science and appl. (2019) up    2 о ог d г г г г    of operations. and for a complete vertex coverage, you will need to repeat such an algorithm one more г time, then the total will be required  3 о г operations. for clarity, we illustrate (table 1) finding the first maximal path with double weights on arcs by calculating all the available routes and choosing the maximal one. table 1. finding the first maximal values with double weights on the paths for 0lnu path number 0lnu 1 4.0628 2 2.3828 3 2.7416 4 2.3891 5 2.9082 6 2.1519 7 2.3016 8 2.1389 9 2.8276 10 2.5899 11 2.2212 12 2.2322 13 2.2010 14 2.5093 15 2.3978 16 2.2619 17 2.7578 18 3.1007 19 2.3034 20 2.2053 a test network consisting of 32 vertex-activities, from the well-known library of network models of projects from http://www.om-db.wi.tum.de/psplib/ provide precedence relations of the selected network project as in the table 2 below. table 2. precedence relations of the test network project. jobnr. activity number; #modes the number of execution modes (does not matter); #successors the number of vertices at which the edges of the relations of precedence go out of the vertex of this work; successors numbers of activities that go to the edges of the relations of the preceding from this vertex. table 2. precedence relations of the test network project. jobnr. #modes #successors successors 1 1 3 2 3 4 2 1 3 6 11 15 3 1 3 7 8 13 4 1 3 5 9 10 5 1 1 20 6 1 1 30 7 1 1 27 http://www.om-db.wi.tum.de/psplib/ 74 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019) 8 1 3 12 19 27 9 1 1 14 10 1 2 16 25 11 1 2 20 26 12 1 1 14 13 1 2 17 18 14 1 1 17 15 1 1 25 16 1 2 21 22 17 1 1 22 18 1 2 20 22 19 1 2 24 29 20 1 2 23 25 21 1 1 28 22 1 1 23 23 1 1 24 24 1 1 30 25 1 1 30 26 1 1 31 27 1 1 28 28 1 1 31 29 1 1 32 30 1 1 32 31 1 1 32 32 1 0 above information from table 2 about relations of technological precedence of selected network project will be organized then with the help of matrix "columns vertices" and "rows paths by vertices". which specifies the characteristic function of the constraint system, and the parameters are specified in the last two columns of table 3 and table 4. table 3. upper the border ub for 0lnu № j= upper the border ub for 0lnu ub = form parameter j  scale parameter 0 ju  0 0.3000 0.5100 1 4.3944 0.2500 0.6000 2 4.1650 0.2200 0.5000 3 0 0.2400 0.5000 4 0 0.2500 0.5000 5 0 0.2800 0.4800 6 4.3565 0.2300 0.4780 7 0 0.2330 0.4750 8 0 0.2350 0.4600 9 0 0.2380 0.4550 10 0 0.2400 0.4500 11 0 0.3000 0.4480 a planning and scheduling method for large-scale innovation projects 75 copyright ©2019 assa adv. in systems science and appl. (2019) 12 0 0.3200 0.4460 13 0 0.3500 0.4450 14 0 0.3800 0.4420 15 0 0.4000 0.4400 16 0 0.4200 0.4300 17 0 0.4300 0.4200 18 0 0.4400 0.4000 19 0 0.4420 0.3800 20 0 0.4450 0.3500 21 0 0.4460 0.3200 22 0 0.4480 0.3000 23 0 0.4500 0.2400 24 0 0.4550 0.2380 25 0 0.4600 0.2350 26 0 0.4750 0.2330 27 0 0.4780 0.2300 28 0 0.4800 0.2800 29 0 0.5000 0.2500 30 4.4742 0.2800 0.7000 31 0 0.5000 0.3200 32 4.1759 0.3000 0.7000 table 4. numerical solution to the problem (3.2) as a whole № j= ln ju = restrictions c = upper the border ub = form parameter j  scale parameter 0 ju  0 3.1203 25.000 3,1203 0,300 0,510 1 4.3944 4.1021 4,3944 0,250 0,600 2 4.1650 3.7607 4,1650 0,220 0,500 3 3.8179 3.4393 3,8179 0,240 0,500 4 3.6652 3.2914 3,6652 0,250 0,500 5 0.2924 0.0444 3,1267 0,280 0,480 6 0.2924 0.0444 4,3565 0,230 0,478 7 0.2924 0.0394 3,7124 0,233 0,475 8 0.2924 0.0444 3,5443 0,235 0,460 9 0.2924 0.1276 3,4537 0,238 0,455 10 0.2924 0.0444 3,3789 0,240 0,450 11 0.2924 0.0444 2,6883 0,300 0,448 12 0.2924 0.0444 2,5063 0,320 0,446 13 1.7583 1.4512 2,2850 0,350 0,445 14 0.2924 0.0444 2,0868 0,380 0,442 15 0.2924 0.0444 1,9711 0,400 0,440 16 0.2924 0.0444 1,8225 0,420 0,430 17 1.2483 0.9752 1,7254 0,430 0,420 18 1.5753 1.2756 1,5753 0,440 0,400 76 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019) 19 0.2924 0.0444 1,4522 0,442 0,380 20 1.2576 0.9652 1,2576 0,445 0,350 21 0.2924 0.0307 1,0538 0,446 0,320 22 0.9051 0.6127 0,9051 0,448 0,300 23 0.4052 0.1128 0,4052 0,450 0,240 24 0.3823 0.0899 0,3823 0,455 0,238 25 0.3506 0.0582 0,3506 0,460 0,235 26 0.3215 0.0291 0,3215 0,475 0,233 27 0.2924 0 0,2924 0,478 0,230 28 0.7010 0.4086 0,7010 0,480 0,280 29 0.4463 0.1539 0,4463 0,500 0,250 30 4.4742 4.1818 4,4742 0,280 0,700 31 0.9400 0.6476 0,9400 0,500 0,320 32 4.1759 3.8835 4,1759 0,300 0,700 0u 1.3396. if we solve the problem in question (3.2) as a whole, as the problem of mathematical programming using ready-made software, we obtain the numerical solution presented in table 4 above. for the 0 0 [ , 1], 0jkw    – arc reliability coefficient (j, k), (from the matrix of paths along the arcs of the network g (e, г):) transfer the result of the j-th activity through the arc (j, k) to perform the k-th activity (like the transfer function of the connection links in the theory of automatic control) we have the following table 5 values. table 5. arcs reliability coefficient 0г arcs value jkw 1 (5,20) 0.95 2 (11,20) 0.95 3 (18,20) 0.95 4 (16,22) 0.95 5 (17,22) 0.95 6 (18,22) 0.95 7 (10,25) 0.95 8 (15,25) 0.95 9 (20,25) 0.95 for all other arcs of the network 1 01 \ .jk jkw w г г г    5. determination of the schedule with inaccurate source data since the values of the scheduling model parameters 0 0 0 (0,1), [ , 1], 0, 0, , 1,..., , ( , ) ( , ) j jk jw u j k n j k г g е г          are known with a certain error, the resource allocation algorithm, as well as the physical volume of work v [usd  hour] and the amount of a non-renewable resource, found as a result of solving the considered problem, which has the value expression u [usd], are approximate, the project’s time parameters the duration of the activity * */ r j j jt v u , r = 0,1, ..., a planning and scheduling method for large-scale innovation projects 77 copyright ©2019 assa adv. in systems science and appl. (2019) q 1,...,j n ; and the activities slacks are determined by processing the numbers known with an error. the obtained solution of problem (3.1), (3.5) together with reliability constraints, vector       * * *u , 0, 1,..., , 0,1,..., , , r r j ju u j n r q r j q n      such that (omitting the index r)  * * * * * 0 0( ) ( )j j j j ju u u u u          . (5.1) by denoting inaccurately set time values as  δt , using actions on approximate numbers [5], one can obtain approximate parameters of the optimal schedule. to do this, we calculate the deterministic project schedule for one pass in the forward and reverse directions based on the obtained optimal (under given error) duration of the project * 0( )it  , starting from the initial project activities. the critical path method, with values, specified inaccurately (5.1), using actions on approximate numbers [5], also allows you to calculate late deadlines for performing project activities in a back pass through the network, starting from the project completion date (calculated by direct passing through the network). and also to calculate slacks of activities [5]: total (full) slack, free (local) slack, safe slack, independent slack. the total (full) slack tsl of the project the period of time by which you can postpone the activity without violating the limitations and deadlines for the project t ( ) ( ) ( ) ( ) ( ).lb eb lc ec i i i i isl t t t t         free (local) slack fsl of the project the period of time by which you can postpone the activity without violating the deadline for the subsequent work – f ( ) min( ( ) ( )) ( ). i eb ec i j ij i j г sl t d t         the safe slack ssl of the project activity is the period of time for which activity j can be prolonged when all its predecessors 1 ji г  started working at the latest date so that the shortest possible time to complete the project does not increase – 1 s ( ) ( ) max ( ). j lb lc j j i i г sl t t       an independent slack i jsl of the project activity j is equal to the time for which the duration of activity j can be prolonged regardless of the time of completion of its predecessors 1 ji г  and the time of the beginning of its followers jk г  1 i *( ) max 0;min ( ) max ( ) ( ) . j j eb lc j k i j k г i г sl t t t               the following relations are valid (under 0  ): t f i t s i( ) ( ) ( ); ( ) ( ) ( ). j j j j j jsl sl sl sl sl sl          conclusion the article discusses a deterministic network model of a project in which, for a project's overall reliability indicator, it is fair its assessment from below in the form of the weakest link – in the worst technological chain from the initial vertex of the project to the final one. 78 v. topka, a. d. tsvirkun copyright ©2019 assa adv. in systems science and appl. (2019) for the indicator of the reliability of the activity-vertex of the project, a power-concave model is used depending on the spent non-renewable (stored) resource. the peculiarity of the model is weighted network arcs, which set the arc reliability factor (j, k), transferring the result of the j-th activity by the arc (j, k) to perform the k-th activity. under these conditions, the problem of planning and scheduling projects is considered on the basis of a successively applied minimax criterion according to the type of lexicographic ordering of criteria in multi-criteria optimization problems. such a lexicographic dissociation, in essence, is a decomposition procedure and is suitable for planning large-scale projects. to solve the problem of uniform allocation of a non-renewable resource, a greedy heuristic has been developed for finding the maximum path with double weights on arcs of an acyclic digraph, computational complexity that is quadratic in the number of arcs. the greedy heuristic gives approximate solution of the problem under consideration. which allows you to get a resource critical path, with the maximum consumption of a non-renewable resource. after that, the project schedule is determined, including its temporary critical path, with inaccurate data. practical calculations for the considered convex problem (3.2) can be performed using ready-made software, and the resulting greedy solution of the problemconsequence (3.3), (3.4) has theoretical significance in graph theory. references 1. the owner’s role in project risk management. (2005) washington, d.c.: the national academies press, www.nap.edu. 2. elkjaer m. (2000). stochastic budget simulation, int. j. project management. 18(2), 139-147. 3. rabbani m., tavakkoli moghaddam r., jolai f. & ghorbani h.r. (2006). a comprehensive model for r&d project portfolio selection with zero-one linear goalprogramming, ije transactions a. basics.19(1), 55-66. 4. tsvirkun a.d., akinfiyev v.k. & konovalov ye.n. (1991). modelirovaniye i upravleniye innovatsionnymi programmami v krupnomasshtabnykh tekhnicheskikh sistemakh. [modeling and management of innovative programs in large-scale technical systems] – moscow, russia: (working paper / ics). [in russian]. 5. topka v.v. (2014). lexicographic solution of two-objective project planning problem under constrained reliability index, j. computer & systems sci. int. 53(6), 877-895. 6. topka v.v. (2012). minimization of project time and cost under constrained reliability index in the disjunctive project model, automation and remote control. 73(7), 11731180. 7. burkov, v.n., zalozhnev, a.yu. & novikov, d.a. (2001). teoriya grafov v upravlenii organizatsionnymi sistemami. [graph theory in the management of organizational systems.]. moscow, russia: sinteg, [in russian]. 8. christofides n. (1975). graph theory: an algorithmic approach (computer science and applied mathematics). n.y.: academic press. 9. karzanov a.v. (1985). o minimal'nykh po srednemu vesu razrezakh i tsiklakh oriyentirovannogo grafa [on minimal cuts and cycles with respect to average weight of an oriented graph] / in: kachestvennyye i priblizhennyye metody issledovaniya operatornykh uravneniy. (pp. 72-83). [qualitative and approximate methods for the investigation of operator equations.] yaroslavl', ussr: yar su, [in russian]. http://www.nap.edu/ https://dl.acm.org/author_page.cfm?id=81100484367&coll=dl&dl=acm&trk=0 adv syst sci appl 2018; 3:1–16 published online at http://ijassa.ipu.ru. person re-identification based on sad and histogram in a non-overlapping camera network abdullah salem baquhaizel ∗, belal alshaqaqi, meriem boumehed, mokhtar keche laboratoire signaux et images, département d’électronique, université des sciences et de la technologie d’oran mohamed boudiaf, usto-mb. oran, bp 1505, el m’naouer. 31000, algeria abstract: in this paper, we present the conception and implementation of a system for person re-identification in a camera network, based on the appearance. this system aims to associate an identifier to each detected person, which keeps this identifier in the same camera and in other cameras even if he or she disappears and then reappears again. our system comprises an improved moving objects detection step that is implemented by combining the mixture of gaussians method (mog) and a proposed difference method, to improve the detection results. then localization and verification step to eliminate false detection. it also comprises a tracking step that is implemented using the sum of absolute differences algorithm (sad), with an acceleration strategy to reduce calculation complexity. the re-identification stage is realized using three steps: the tracking for the temporal association, the histogram and the new smart exploitation of the interpolation technique for comparison. our proposed system build an online database that contains the history of every person that enters the field of view of the cameras, it does not require any previously collected training data, it integrates both the detection and the tracking so it can be used in real applications, and another advantage is the simplicity of the used techniques in term of implementation and calculation rapidity, without affecting the quality of results. the global system was tested on a real data set collected by three cameras. the experimental results show that our approach gives very courageous results. keywords: person re-identification, mog, sad, histogram, interpolation 1. introduction in recent times, video surveillance has grown more and more. this resulted in an increase of cameras installed in different places (public or private), making their exploitation and monitoring extremely difficult for a human being. consequently, much research has been done to create intelligent vision systems that can help the human being, in interpreting scenes and reacting with alarms in case of any anomaly. currently, there are several types of video surveillance systems (people recognition, access control in sensitive locations, control of traffic congestion, ...etc.). in this paper, we are particularly interested in the problem of person re-identification in a camera network. person re-identification in computer vision systems aims to follow a person, associate an identifier to him, and store it in a database. when the person leaves the scene then reappears in the field of view of any camera, it will be assigned the same identifier. in a crowded and uncontrolled environment observed by cameras from unknown distances, person re-identification relying upon conventional biometrics, such as face recognition, is neither feasible nor reliable, due to insufficiently constrained conditions and insufficient ∗corresponding author: abdullah.baquhaizel@{univ-usto.dz ; gmail.com} 2 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche image details for extracting robust biometrics [31]. alternatively, visual features based on the appearance of people, determined by their clothing and objects carried or associated with them, can be exploited more reliably for person re-identification. the main contributions of this paper are: our proposed system is able to build a fully automated online database that contains the history of every person captured by the cameras, based on their appearance. the proposed system does not require any previously collected training data, so no specific reference for a person is stored a priori. compared to most of the other state-of-the-art approaches, our proposed system is a self-contained system that is ready to use in real applications since it integrates both the detection and the tracking phases. the simplicity of used techniques. furthermore, the improvement in each block, object detection by combining the mog method and the difference method, the tracking and the re-identification. the rest of this paper is organized as follows: section 2 presents some related works from the literature. in section 3, we describe the proposed system and give the details for its blocks. the experimental results and their discussion are presented in section 4. finally, some conclusions are drawn in section 5. 2. related works in the literature, the approaches of person re-identification can be grouped into several classes, according to several criteria [20]: 1. the number of images per person: this class comprises two families. the first family is the family of mono-sample methods, where the signature of a person is extracted from a single image as in [4, 9, 12, 26, 30, 39]. for example, in[4], authors project the gallery and the probe data into the regularized canonical correlation analysis (rcca) subspace then the reference descriptors (rds) of the gallery and probe are constructed by measuring the similarity between them and the reference data. in [12], a metric learning framework is used to obtain a robust metric for large margin nearest neighbor classification with rejection. the second family is the family of multi-sample methods, where multiple images are used to calculate the signature of a person as in [10, 11, 13, 19, 22, 29, 34, 36]. for example, authors in [10] propose custom pictorial structure (cps) for re-identification. in [22] they use the implicit shape model (ism) and sift features for the person reidentification. in [29] spatiotemporal person features are extracted using multi-frame twin-channel descriptor based on a gabor filter, then mahalanobis distance metric learning algorithm is used for matching. in [36] the discriminating nature of the sparse representation is exploited in order to perform people re-identification task. 2. the type of representation: the first family in this class is the family of global approaches, where the whole information in the image is exploited for calculating the person’s signature, as in [4, 5, 7, 17, 21, 27]. for example, in [21], the jensen-shannon kernel is proposed to learn nonlinear distance metrics. in [5] non-articulated 3d body models are exploited to spatially map appearance descriptors (color and gradient histograms) into the vertices of a regularly sampled 3d body surface. in[17], the appearance by a set of region covariance descriptors is modeled, then a discriminative model is learned using boosting for feature selection. in [27], authors extract kdes (kernel descriptor) from human roi for person classification. the second family is that of local approaches, which represent the image by several feature vectors, each vector describes a region or an interest point locally detected, such copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network3 as in [11, 14, 15, 42]. for example, in [11] interest points are detected with fast-hessian detector, then they are described in color and surf description. the approach in [15] uses matching of signatures based on interest points descriptors detected by a variant inspired from surf called camellia key-points. in [42], authors use the face to video retrieval, they analyzed four local features for describing face images: harris corner operators, sift descriptors, surf descriptors, and eigenfaces. 3. the existence of a set of images mapped a priori: this class includes supervised approaches like in [7, 9, 22, 25, 28]. for example, two structured learning based approaches are proposed in [25], they explore the adaptive effects of multiple low-level visual features with an optimal ensemble of their metrics. in [28], a data driven approach is proposed for learning color patterns, they model color feature generation by jointly learning a linear transformation and a dictionary to encode pixel values. in the other hand, we have the unsupervised approaches as in [18, 23, 38, 40]. for example, a video representation, called spatio-temporal pyramid sequence (stps) is developed by [23] to encode space-time information, and they formulate a novel time shift dynamic time warping (ts-dtw) model and its multi-dimensional extension named mdts-dtw for selective matching between sequences. in [40] a dynamic graph matching (dgm) method is proposed; a graph for samples in each camera is constructed, and then graph matching scheme is introduced for cross-camera labeling association. a very nice survey of people re-identification approaches is presented in [37]. they are therein grouped as a multidimensional taxonomy according to camera setting, sample set cardinality, adoption of a body model, signature, machine learning techniques, and application scenario. 3. description of the proposed system we present the conception and implementation of a system for person re-identification in a camera network, based on the appearance. in this section, we describe the different blocks of the proposed system and how all these blocks are connected. these blocks are: the detection of moving objects, their localization and verification, their tracking and their re-identification. the detailed flowchart of the proposed system is shown in figure 3.1. in the beginning, we initialize the number of found objects to zero. then for each frame, we start detecting moving objects, next we perform their localization, then a verification and confirmation stage is applied to improve the results by eliminating the false detection. in parallel to the detection and for each frame also, we have the tracking process of any found objects from the previous frame. after that, the obtained results from the tracking are fused with those of the detection and verification stage, by using the technique of intersecting and overlapping. then comes the feature extraction step where the histogram for each found object is calculated. for the re-identification, we use the tracking result for the temporal association, the histogram and the trajectory interpolation for the comparison. the re-identification and association block tries to associate the found objects with objects already seen or add them as new objects if they appear for the first time, thus, building an online database which contains the history of every person that appears in front of the camera. in the coming subsections, we will discuss in detail the techniques and the tools used in each block. 3.1. detection of moving objects this initial phase is performed by combining the mixture of gaussians (mog) method [35] and a proposed method we called the difference method. copyright c© 2018 assa. adv syst sci appl (2018) 4 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche 11 initiation: n° found objects=0 image capture intersection differencemog or morphological operations update db with new id *1 trajectory interpolation *3 histogram *3 update db with matched id *2 if found objects from update db with id of tracking *2 detection only intersection or tracking only localization verification yes no retrieved detected objects n° found objects >0 tracking yes yes retrieved tracked objects no noyes yes no no update db database db test db *3:comparison *1: new id *2: existed id update db with interpolated id *2 histogram fig. 3.1. the detailed flowchart of the proposed system the mog is one of the most used and successful methods in surveillance systems, because it is adaptive, and can handle multimodal backgrounds [8]. it is the most common method to build a background[32]. furthermore, it is robust to slow lighting changes, periodic motions from a cluttered background, slow moving objects, long term scene changes, and camera noises [41]. in this method, the recent history of each pixel, {x1, ..., xt}, is modeled by a mixture of k gaussian distributions. the probability of observing the current pixel value is given by: copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network5 p (xt) = k∑ i=1 ωi,t ∗ η(xt, µi,t,σi,t) (3.1) where ωi,t is the portion of the data accounted by the ith gaussian in the mixture at time t and η(xt, µi,t,σi,t) is a gaussian probability density function with mean value µi,t and covariance matrix σi,t of the ith gaussian in the mixture at time t , where η is given by: η(xt, µ,σ) = 1 (2π) n 2 |σ| 1 2 e− 1 2 (xt−µt)t σ−1(xt−µt) (3.2) k is the maximum number of gaussian components, depending on the available memory. currently, this parameter ranges from 3 to 5. at each pixel xt in a new frame, the following update equations for the k distributions are performed: ωk,t = (1− β)ωk,t + β (3.3) β is the learning rate of adaptation. this value ranges from 0.0 to 1.0, determines the importance of the previous observation. the first b distributions are chosen as the background model, where: bt = argminb ( b∑ k=1 ωk > t ) (3.4) the value of t determines the minimum portion of the data that should be accounted by the background. when higher values for t are chosen, a multi-model background model is formed which can handle repetitive background motion [35]. in the other hand, we have the difference method, we first take the difference between two successive images in gray-scale ig(t) and ig(t−1), as in eq. (3.5), and then we compare the resulting difference image idiff to a threshold to detect pixels in movement. idiff = ig(t) − ig(t−1) (3.5) the combination of the detection resulting from the mog and the difference methods is performed using the logical or operation. after that, we apply morphological operations as erosion, dilatation, and fill the holes. the holes of a binary image corresponding to the set of its regional minima, which are not connected to the image border [33]. 3.2. localization and verification to localize the detected objects, we use the labeling technique [16]. it consists in separating the areas in the mask obtained from the detection step. we associate with each area an integer value (label) by using an 8-connected neighborhood. then we propose a verification phase to eliminate false detection. we first calculate some proprieties for each localized area, e.g. x and y coordinates, height, width and sum of foreground pixels. then to be validated, each object has to verify the following three conditions: • the height to width ratio has to lie between min and max thresholds. • the surface of the rectangle containing the object (surface = height x width) has to lie between min and max thresholds. this is to reject very small and very big objects due to false detection. • the ratio of the sum of foreground pixels to the surface also has to be bounded. copyright c© 2018 assa. adv syst sci appl (2018) 6 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche 3.3. tracking the tracking process is done by template matching using the sum of absolute differences algorithm (sad), like in our previous proposed driver drowsiness detection systems[2, 3]. this algorithm is widely used for image compressing and object tracking in real-time application [1], it can give high localization rate when the image is with high illumination variation [24]. in digital image processing, the sad is a measure of the similarity between image blocks. it is calculated by taking the absolute difference between each pixel in the original block x (a portion from the current frame) and the corresponding pixel in the y block being used for comparison (model from the previous detection). these differences are summed to create a simple metric of block similarity as in eq. (3.6), zero means that the two blocks are identical. we sweep all the positions in the frame, then the block with the smallest metric is the tracked block. the sad value for two blocks x and y is calculated by: sad = m∑ i=1 n∑ j=1 |x(i, j)− y (i, j)| (3.6) for a given y model, the most similar block x is the one that minimizes the sad. 3.4. re-identification after the stages of detection, localization, verification, and tracking, we have the stage of reidentification and online construction of database db containing the history of each person that appeared in the view field of the cameras. the lower part of figure 3.1 presents a detailed flowchart of this stage. this stage deals with the moving objects obtained from the detection and tracking stages, which are called ’found objects’. first, we proposed to perform the intersection between the found objects resulting from the detection and tracking, then we calculate the percentage of the intersection if it exceeded a predefined threshold we merge the objects from detection and tracking into one object, otherwise, they stay separated objects. the intersection (a ∩b) of two rectangles a and b is the rectangle that contains all elements of a that also belong to b. now, we test each found object, if it resulted from the detection only, the tracking only or from both (intersection). if the found object comes from the intersection or tracking only, we update the database with the identifier of tracked object. on the other hand, if that found object comes from detection only, then we calculate its histogram. an image histogram is a type of histogram that acts as a graphical representation of the tonal distribution in a digital image. it plots the number of pixels for each tonal value. the histogram of the found object is compared to the histograms of identified objects stored in the database. if there is a match, then we update the database by associating that object with the matched identifier, otherwise, we make a trajectory interpolation. the interpolation is a method of constructing new data points within the range of a discrete set of known data points. if there is an interpolation, we update the database with the interpolated identifier, otherwise, we consider this object as a new one and assign to it a new identifier that is added to the database. 4. experimental results in this section, we present the material and the database used, the experimental results, and their discussion. copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network7 fig. 4.2. layout of the cameras (a) (b) (c) fig. 4.3. fields of view of the 3 cameras: (a) camera 1, (b) camera 2, (c) camera 3 4.1. system development environment the material that was used for the development of our application is: 1. a laptop with: • processor: intel r©coretm i7 4702mq cpu @ 2.20 ghz 2.20 ghz. • ram memory: 8.00 go. • operating system: windows 8.1, 64-bits • hard drive: 1 tb. 2. digital video recorder dvr. 3. camera with characteristics: • 1 /3 sony hr ccd • 420 tv lines • 0.2 lux • adjustable focal between (3mm and 8 mm). to test our system we build our own database [6], composed of sequences of images recorded on the third floor of the department of electronics at usto university. three cameras, set to a height of (2.30m) and with an angle of ( −30◦), were used to take these images. each sequence contains from one to three people who walk in the fields of view of the three cameras. the cameras were placed as shown in the layout presented in figure 4.2. copyright c© 2018 assa. adv syst sci appl (2018) 8 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche to fulfill the condition of a non-overlapping camera network, the database was realized so that a person lies in the field of view of only one camera, at a given instant. figure 4.3 shows the fields of view of the three cameras. 4.2. experimental results, and discussion in this section, we will present and discuss the results of each step of the proposed system. 4.2.1. results of the detection the mixture of gaussian gives us raw results of detection from each camera, after having defined suitable settings according to some criteria, like: indoor or outdoor environment, people movement speed and lighting changes. figure 4.4 (b) presents an example of these results. to improve these raw results, we combine them with the results of the difference method (figure 4.4 (c)), which allows for the detection of the edges of moving objects, then we apply morphological operations as erosion, dilatation and proceed to a holes filling of the resulting image to minimize false detection and obtain better results as illustrated in figure 4.4 (d). (a) (b) (c) (d) fig. 4.4. example of moving objects detection. (a) original image, (b) results of detection by mog, (c) results of detection by difference, (d) results of the holes filling of the or between b and c 4.2.2. results of the verification to evaluate the verification phase, we compare the results before and after the verification. we consider the true detection from the detection before verification as positive and the false detection as negative. table 4.1 presents for videos from camera 1, camera 2 and camera 3, the obtained evaluation measures, namely the tp (true positive), the fn(false negative), the fp (false positive), and the tn (true negative). to understand these measures we present figure 4.5, let us assume that we have four objects; copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network9 1 2 3 4     (a) 1 2 3 4     (b) fig. 4.5. before and after verification evaluation, (a) detection results (before the verification), (b) results after the verification table 4.1. evaluation results of the proposed verification technique camera 1 camera 2 camera 3 tp 411 399 341 fn 23 8 7 fp 2 2 3 tn 19 30 73 objects 1 and 2 are not people, however, objects 3 and 4 are people. figure 4.5 (a) is the results before the verification, so objects 1 and 2 are false detection, however objects 3 and 4 are true detection. in figure4.5 (b), we have the results after the verification, objects 1 and 3 are kept, but objects 2 and 4 are removed. now we can assign each measure of evaluation to the corresponding object: • object 1 is a false positive (fp). • object 2 is a true negative (tn). • object 3 is a true positive (tp). • object 4 is a false negative (fn). from these measures, we can define the following rates: tp tpfn % = tp/ (tp + fn): % of true person kept. fn tpfn % = fn/ (tp + fn): % of true person removed. fp tnfp % = fp/ (tn + fp): % of false person kept. tn tnfp % = tn/ (tn + fp): % of false person removed. the obtained evaluation rates are given in table. 4.2, for videos from camera 1, camera 2 and camera 3. from this table one can observe, that the proposed verification technique performs well, since it manages to remove 90.48%, 93.75% and 96.05% of false detection in camera 1, camera 2 and camera 3, respectively, while removing only 5.30%, 1.97% and 2.01% of true detection, in camera 1, camera 2 and camera 3, respectively. table 4.2. evaluation rates of the proposed verification technique camera 1 camera 2 camera 3 tp tpfn% 94.70 98.03 97,99 fn tpfn% 5.30 1.97 2.01 fp tnfp% 9.52 6.25 3.95 tn tnfp% 90.48 93.75 96.05 copyright c© 2018 assa. adv syst sci appl (2018) 10 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche in figure 4.6, we give the localization and verification results. after the localization by the labeling technique, we apply the verification procedure to each object. in figure 4.6 (a), only the objects that pass the conditions of verification are kept (the person in green rectangle), the others in red are ignored. figure 4.6 (b) shows the results of detection . (a) (b) fig. 4.6. example of localization and verification, (a) localization and verification results on the original image, (b) results of detection 4.2.3. results of the tracking the tracking step is run in parallel with the detection step and it is realized by the sad. to accelerate its execution we decided to apply it only to a limited region of interest, instead of searching in the whole frame. this region is determined by the coordinates of the model to track. the obtained results of tracking are satisfactory. figure 4.7 shows the tracking results, figure 4.7 (a) is a detection in frame 137 and figure 4.7 (b) is its tracking in frame 194, it means that we still maintain this object even after about 57 frames. (a) (b) fig. 4.7. tracking results, (a) detection in frame 137, (b) the tracking of detection of frame 137 in frame 194 4.2.4. results of the re-identification the re-identification stage is realized using three techniques, the tracking for the temporal association, the histogram and the interpolation technique for comparison. in figure 4.8, we present the multiplication of the detected object with its mask to extract the silhouette only. then, we calculate histograms of red, green, blue channels and gray-scales as shown in figure 4.9. we use the histogram of the silhouette to avoid the effect of the background. these different histograms are used for comparison with the models stored in the database. if there is a match, we associate the matched identifier to copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network11 the actual object, otherwise, we use the trajectory interpolation technique, as shown in figure 4.10 (a), to determine if this object is in the trajectory of an existing object. if so, we associate the actual object to this object, otherwise, we consider that the actual object is new and add it to the database with a new identifier. from figure 4.10 (b), we can deduce that in the period between the frames 220 and 270, mentioned by two verticals lines in red, the two person appear together and figure 4.10 (c) is a sample from that period. note that the appearance of person 1 starts from the frame 120, and that of person 2 from the frame 180. the miss in the appearance of the person are due to the elimination of certain items by the verification stage. the performance of our re-identification system is evaluated using two criteria: 1. number of created ids, 2. number of appearances and associations of each id, evaluation based on the number of created ids: we can evaluate our proposed reidentification system by comparing the number of assigned ids that created by the system to the number of real ids. the ideal is when these two numbers are equal. in table. 4.3, these numbers are compared, using videos from the different cameras. the number of real ids is obtained by the ground truth of used dataset, and their appearance is how many time they appear. however, the number of created ids is obtained by the system, and their appearance is how many time they appear. it can be observed that the number of created ids is larger than the number of real ids; this is due to the false detection. it can also be observed that by eliminating most of the false detection using the verification stage, we reduce the number of created ids to a level close to the number of real ids. table 4.3. comparison of the number of created ids (nb c ids) to the number of real ids (nb r ids), and their number of appearance (nb a) real created nb r ids nb a nb c ids nb a camera 1 3 434 7 515 camera 2 3 407 5 541 camera 3 3 348 8 494 table 4.4. number of appearances (nb a) and person-ids associations of each id(p) in camera 1, 2 and 3 id camera 1 camera 2 camera 3 nb a p nb a p nb a p 1 175 p1 181 p1 98 p1 2 36 p2 27 131 p2 3 31 172 p2 29 4 41 p3 57 40 p1 5 117 p1 104 p3 30 6 27 33 7 88 p3 102 p3 8 31 p3 evaluation based on the number of appearances and associations of each id: another way to evaluate our proposed system is the number of appearances of each id,(see table. 4.4). from this table, we can determine the number of ids for each camera and the number of appearances of each id. for example, we can see that for camera 1, seven copyright c© 2018 assa. adv syst sci appl (2018) 12 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche ids were created, and id 5 was associated 117 times. in addition, from this table, we can determine the ids associated with each person. for example, it can be seen from table. 4.4, that with camera 3 the ids 1 and 4 were associated with the person 1, named p1. (a) (b) (c) fig. 4.8. the multiplication of detected object with its mask, (a) detected object, (b) mask of detected object and (c) results of multiplication (a) (b) (c) (d) fig. 4.9. different histograms of the silhouette, (a), (b), (c) and (d) are the histograms of red, green, blue channels and gray-scales respectively 5. conclusion in this paper, we proposed the conception and implementation of a system for person reidentification in a camera network, based on the appearance. this system is able to assign an identifier to each detected person, that it keeps everywhere in the fields of view of the cameras and even if he or she disappears and then appears again. our system implements an improved detection technique that combines the mixture of gaussian method with the difference method. the sad algorithm with an acceleration copyright c© 2018 assa. adv syst sci appl (2018) person re-identification based on sad and histogram in a non-overlapping camera network13 time ( frames ) (a) time ( frames ) (b) (c) fig. 4.10. trajectory interpolation, (a) the trajectories of the two objects through time, (b) the trajectories of the two objects along the y-axis (c) the two objects that appear in this sequence strategy is used for the tracking step, whereas the re-identification stage is realized using three techniques: the tracking for the temporal association, the histogram and the trajectory interpolation technique for comparison. our proposed system does not require any previously collected training data to build an online database that contains the history of every person that enters the field of view of the cameras. it can be used in real applications because it integrates both the detection and the tracking. another advantage is the simplicity of the used techniques in term of implementation and calculation rapidity, without affecting the quality of results. a real data set is collected by three cameras to test the global system. the experimental results show that our approach leads to very courageous results with an opportunity for improvement in the re-identification stage, by using local histograms instead of using the global one. also as a future work, we plan to evaluate our method quantitatively and compare it with other methods. references [1] alshaqaqi, b., baquhaizel, a. s., ouis, m., & boumehed, m. (2016). real-time implementation of face detection and tracking. algerian journal of research and technology, 1(1), 25–30. [2] alshaqaqi, b. [belal], baquhaizel, a. s., amine ouis, m. e., boumehed, m., ouamri, a., & keche, m. (2013). driver drowsiness detection system. in 2013 8th international copyright c© 2018 assa. adv syst sci appl (2018) 14 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche workshop on systems, signal processing and their applications (wosspa) (pp. 151– 155). doi:10.1109/wosspa.2013.6602353 [3] alshaqaqi, b. [belal], baquhaizel, a. s., ouis, m. e. a., boumehed, m., ouamri, a., & keche, m. (2013). vision based system for driver drowsiness detection. in 2013 11th international symposium on programming and systems (isps) (pp. 103–108). doi:10.1109/isps.2013.6581501 [4] an, l., kafai, m., yang, s., & bhanu, b. (2013). reference-based person reidentification. in 2013 10th ieee international conference on advanced video and signal based surveillance (pp. 244–249). doi:10.1109/avss.2013.6636647 [5] baltieri, d., vezzani, r., & cucchiara, r. (2015). mapping appearance descriptors on 3d body models for people re-identification. international journal of computer vision, 111(3), 345–364. doi:10.1007/s11263-014-0747-z [6] baquhaizel, a. s., kholkhal, s., alshaqaqi, b., & keche, m. (2018). ssd and histogram for person re-identification system. in ifip advances in information and communication technology (vol. 522, pp. 585–596). doi:10.1007/978-3-319-897431 50 [7] bauml, m., & stiefelhagen, r. (2011). evaluation of local features for person reidentification in image sequences. in 2011 8th ieee international conference on advanced video and signal based surveillance (avss) (pp. 291–296). doi:10 . 1109 / avss.2011.6027339 [8] boumehed, m., alshaqaqi, b., ouamri, a., & keche, m. (2013). moving objects localization by local regions based level set: application on urban traffic. journal of mathematical imaging and vision, 46(2), 258–274. doi:10.1007/s10851-012-0400-9 [9] cai, y., huang, k., & tan, t. (2008). human appearance matching across multiple nonoverlapping cameras. in 19th international conference on pattern recognition (pp. 1–4). doi:10.1109/icpr.2008.4761704 [10] cheng, d. s., cristani, m., stoppa, m., bazzani, l., & murino, v. (2011). custom pictorial structures for re-identification. in proceedings of the british machine vision conference (pp. 68.1–68.11). doi:10.5244/c.25.68 [11] de oliveira, i. o., & pio, j. l. d. s. (2009). people reidentification in a camera network. in 2009 eighth ieee international conference on dependable, autonomic and secure computing (pp. 461–466). doi:10.1109/dasc.2009.33 [12] dikmen, m., akbas, e., huang, t. s., & ahuja, n. (2011). pedestrian recognition with a learned metric. in k. ron, k. reinhard, & s. akihiro (eds.), computer vision accv 2010. lecture notes in computer science lncs (vol. 6495, part 4, pp. 501–512). doi:10.1007/978-3-642-19282-1 40 [13] farenzena, m., bazzani, l., perina, a., murino, v., & cristani, m. (2010). person re-identification by symmetry-driven accumulation of local features. in 2010 ieee computer society conference on computer vision and pattern recognition (pp. 2360– 2367). doi:10.1109/cvpr.2010.5539926 [14] gheissari, n., sebastian, t., & hartley, r. (2006). person reidentification using spatiotemporal appearance. in 2006 ieee computer society conference on computer vision and pattern recognition volume 2 (cvpr’06) (vol. 2, pp. 1528–1535). doi:10. 1109/cvpr.2006.223 [15] hamdoun, o., moutarde, f., stanciulescu, b., & steux, b. (2008). person reidentification in multi-camera system by signature based on interest point descriptors collected on short video sequences. in 2008 second acm/ieee international conference on distributed smart cameras (pp. 1–6). doi:10 . 1109 / icdsc . 2008 . 4635689 [16] haralick, r. m., & shapiro, l. g. (1992). computer and robot vision, vol 1. (pp. 28–48). addison-wesley. [17] hirzer, m., beleznai, c., roth, p. m., & bischof, h. (2011). person re-identification by descriptive and discriminative classification. in h. kalviainen, j. parkkinen, & copyright c© 2018 assa. adv syst sci appl (2018) https://dx.doi.org/10.1109/wosspa.2013.6602353 https://dx.doi.org/10.1109/isps.2013.6581501 https://dx.doi.org/10.1109/avss.2013.6636647 https://dx.doi.org/10.1007/s11263-014-0747-z https://dx.doi.org/10.1007/978-3-319-89743-1_50 https://dx.doi.org/10.1007/978-3-319-89743-1_50 https://dx.doi.org/10.1109/avss.2011.6027339 https://dx.doi.org/10.1109/avss.2011.6027339 https://dx.doi.org/10.1007/s10851-012-0400-9 https://dx.doi.org/10.1109/icpr.2008.4761704 https://dx.doi.org/10.5244/c.25.68 https://dx.doi.org/10.1109/dasc.2009.33 https://dx.doi.org/10.1007/978-3-642-19282-1_40 https://dx.doi.org/10.1109/cvpr.2010.5539926 https://dx.doi.org/10.1109/cvpr.2006.223 https://dx.doi.org/10.1109/cvpr.2006.223 https://dx.doi.org/10.1109/icdsc.2008.4635689 https://dx.doi.org/10.1109/icdsc.2008.4635689 person re-identification based on sad and histogram in a non-overlapping camera network15 a. kaarna (eds.), image analysis: 17th scandinavian conference, scia 2011, ystad, sweden, may 2011. proceedings (vol. 3540, march 2016, pp. 91–102). lecture notes in computer science. doi:10.1007/978-3-642-21227-7 9 [18] hirzer, m., roth, p. m., & bischof, h. (2012). person re-identification by efficient impostor-based metric learning. in 2012 ieee ninth international conference on advanced video and signal-based surveillance (pp. 203–208). doi:10 .1109/avss. 2012.55 [19] huang, c.-h., wu, y.-t., & shih, m.-y. (2009). unsupervised pedestrian reidentification for loitering detection. in advances in image and video technology (pp. 771–783). doi:10.1007/978-3-540-92957-4 67 [20] ibn khedher, m. (2014). ré-identification de personnes à partir des séquences vidéo (dissertations, institut national des télécommunications, paris). [21] ijiri, y., & lao, s. (2012). human re-identification through distance metric learning based on jensen-shannon kernel. in visapp: international conference on computer vision theory and applications (pp. 603–612). retrieved from http://www.murase.nuie. nagoya-u.ac.jp/$%5csim$murase/pdf/951-pdf.pdf [22] jungling, k., & arens, m. (2011). view-invariant person re-identification with an implicit shape model. in 2011 8th ieee international conference on advanced video and signal based surveillance (avss) (pp. 197–202). doi:10 . 1109 / avss . 2011 . 6027319 [23] ma, x., zhu, x., gong, s., xie, x., hu, j., lam, k.-m., & zhong, y. (2017). person re-identification by unsupervised video matching. pattern recognition, 65, 197–210. doi:10.1016/j.patcog.2016.11.018. arxiv: 1611.08512 [24] nourain dawoud, n., belhaouari samir, b., & janier, j. (2011). fast template matching method based optimized sum of absolute difference algorithm for face localization. international journal of computer applications, 18(8), 30–34. doi:10. 5120/2302-2912 [25] paisitkriangkrai, s., wu, l., shen, c., & van den hengel, a. (2017). structured learning of metric ensembles with application to person re-identification. computer vision and image understanding, 156, 51–65. doi:10.1016/j.cviu.2016.10.015. arxiv: 1511.08531 [26] park, u., jain, a., kitahara, i., kogure, k., & hagita, n. (2006). vise: visual search engine using multiple networked cameras. in 18th international conference on pattern recognition (icpr’06) (pp. 1204–1207). doi:10.1109/icpr.2006.1176 [27] pham, t. t. t., le, t. l., vu, h., dao, t. k., & nguyen, v. t. (2017). fullyautomated person re-identification in multi-camera surveillance system with a robust kernel descriptor and effective shadow removal method. image and vision computing, 59, 44–62. doi:10.1016/j.imavis.2016.10.010 [28] rama varior, r., wang, g., lu, j., & liu, t. (2016). learning invariant color features for person reidentification. ieee transactions on image processing, 25(7), 3395–3410. doi:10.1109/tip.2016.2531280. arxiv: 1410.1035 [29] sathish, p. k., & balaji, s. (2017). multi-frame twin-channel descriptor for person re-identification in real-time surveillance videos. international journal of multimedia information retrieval, 6(4), 289–294. doi:10.1007/s13735-017-0136-9 [30] schwartz, w. r., & davis, l. s. (2009). learning discriminative appearance-based models using partial least squares. in 2009 xxii brazilian symposium on computer graphics and image processing (pp. 322–329). doi:10.1109/sibgrapi.2009.42 [31] shaogang gong, marco cristani, c. c. l., & hospedales, t. (2014). the reidentification challenge. in s. gong, m. cristani, s. yan, & c. c. loy (eds.), person re-identification (chap. 1). doi:10.1007/978-1-4471-6296-4 1 [32] shi, y., cheng, s., quan, s., chen, j., & chen, d. (2011). moving objects detection by gaussian mixture model: a comparative analysis. in 2011 international conference on electrical and control engineering (20090143110004, pp. 1121–1124). doi:10.1109/ iceceng.2011.6058008 copyright c© 2018 assa. adv syst sci appl (2018) https://dx.doi.org/10.1007/978-3-642-21227-7_9 https://dx.doi.org/10.1109/avss.2012.55 https://dx.doi.org/10.1109/avss.2012.55 https://dx.doi.org/10.1007/978-3-540-92957-4_67 http://www.murase.nuie.nagoya-u.ac.jp/$%5csim$murase/pdf/951-pdf.pdf http://www.murase.nuie.nagoya-u.ac.jp/$%5csim$murase/pdf/951-pdf.pdf https://dx.doi.org/10.1109/avss.2011.6027319 https://dx.doi.org/10.1109/avss.2011.6027319 https://dx.doi.org/10.1016/j.patcog.2016.11.018 https://arxiv.org/abs/1611.08512 https://dx.doi.org/10.5120/2302-2912 https://dx.doi.org/10.5120/2302-2912 https://dx.doi.org/10.1016/j.cviu.2016.10.015 https://arxiv.org/abs/1511.08531 https://dx.doi.org/10.1109/icpr.2006.1176 https://dx.doi.org/10.1016/j.imavis.2016.10.010 https://dx.doi.org/10.1109/tip.2016.2531280 https://arxiv.org/abs/1410.1035 https://dx.doi.org/10.1007/s13735-017-0136-9 https://dx.doi.org/10.1109/sibgrapi.2009.42 https://dx.doi.org/10.1007/978-1-4471-6296-4_1 https://dx.doi.org/10.1109/iceceng.2011.6058008 https://dx.doi.org/10.1109/iceceng.2011.6058008 16 a.s.baquhaizel, b.alshaqaqi, m. boumehed, m.keche [33] soille, p. (2004). morphological image analysis. in morphological image analysis principles and applications (p. 208). doi:10.1007/978-3-662-05088-0. arxiv: 1011. 1669 [34] souded, m. (2013). people detection, tracking and re-identification through a video camera network (doctoral dissertation, signal and image processing, institut national de recherche en informatique et en automatique, universite de nice sophia antipolis ecole, nice). [35] stauffer, c., & grimson, w. (1999). adaptive background mixture models for realtime tracking. in ieee computer society conference on computer vision and pattern recognition (vol. 2, 100, pp. 246–252). doi:10 . 1109 / cvpr . 1999 . 784637. arxiv: cvpr.1999.784637 [10.1109] [36] truong cong, d.-n., achard, c., & khoudour, l. (2010). people re-identification by classification of silhouettes based on sparse representation. in 2010 2nd international conference on image processing theory, tools and applications (pp. 60–65). doi:10. 1109/ipta.2010.5586809 [37] vezzani, r., baltieri, d., & cucchiara, r. (2013). people reidentification in surveillance and forensics: a survey. acm computing surveys, 46(2), 1–37. doi:10.1145/2543581. 2543596 [38] wang, t., gong, s., zhu, x., & wang, s. (2014). person re-identification by video ranking. in d. fleet, t. pajdla, b. schiele, & t. tuytelaars (eds.), computer vision – eccv 2014. lecture notes in computer science lncs (vol. 8692, part 4, pp. 688–703). doi:10.1007/978-3-319-10593-2 45 [39] wang, x., doretto, g., sebastian, t., rittscher, j., & tu, p. (2007). shape and appearance context modeling. in 2007 ieee 11th international conference on computer vision (pp. 1–8). doi:10.1109/iccv.2007.4409019 [40] ye, m., ma, a. j., zheng, l., li, j., & yuen, p. c. (2017). dynamic label graph matching for unsupervised video re-identification. proceedings of the ieee international conference on computer vision, 2017-octob, 5152–5160. doi:10.1109/ iccv.2017.550. arxiv: 1709.09297 [41] ying-li tian, lu, m., & hampapur, a. (2005). robust and efficient foreground analysis for real-time video surveillance. in 2005 ieee computer society conference on computer vision and pattern recognition (cvpr’05) (vol. 1, pp. 1182– 1187). doi:10.1109/cvpr.2005.304 [42] zhou, j., & huang, l. (2017). a comparative study of local features in face-based video retrieval. journal of computing science and engineering, 11(1), 24–31. doi:10. 5626/jcse.2017.11.1.24 copyright c© 2018 assa. adv syst sci appl (2018) https://dx.doi.org/10.1007/978-3-662-05088-0 https://arxiv.org/abs/1011.1669 https://arxiv.org/abs/1011.1669 https://dx.doi.org/10.1109/cvpr.1999.784637 https://arxiv.org/abs/cvpr.1999.784637 https://dx.doi.org/10.1109/ipta.2010.5586809 https://dx.doi.org/10.1109/ipta.2010.5586809 https://dx.doi.org/10.1145/2543581.2543596 https://dx.doi.org/10.1145/2543581.2543596 https://dx.doi.org/10.1007/978-3-319-10593-2_45 https://dx.doi.org/10.1109/iccv.2007.4409019 https://dx.doi.org/10.1109/iccv.2017.550 https://dx.doi.org/10.1109/iccv.2017.550 https://arxiv.org/abs/1709.09297 https://dx.doi.org/10.1109/cvpr.2005.304 https://dx.doi.org/10.5626/jcse.2017.11.1.24 https://dx.doi.org/10.5626/jcse.2017.11.1.24 introduction related works description of the proposed system detection of moving objects localization and verification tracking re-identification experimental results system development environment experimental results, and discussion results of the detection results of the verification results of the tracking results of the re-identification conclusion paper title (use style: paper title) adv syst sci appl 2018; 02; 59-62 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/578. specifics of long-term forecasting for global gas markets vyacheslav kulagin1,*, anna galkina1 1) department of world and russian energy complex research energy research institute of the russian academy of sciences moscow, russia abstract: the article presents a methodology for developing long-term forecasts for global gas markets using optimization modeling. this approach can be effectively used for an integrated analysis of the global gas market conjuncture, and assessment of management decisions in the gas industry in both the short and long term. keywords: gas market; modeling; forecast; demand; production; prices 1. introduction the article covers the specifics of forecasting for global markets employing optimization models. these models are based on the principle of the monge-kantorovich problem of large dimension, in which the world’s total costs of satisfying the global gas demand are minimized: 𝑓(𝑥) → 𝑚𝑖𝑛 where 𝑓(𝑥) stands for the total global costs of gas production, including taxes, global transportation costs, gas storage, as well as liquefaction and regasification in the case of lng transportation. 2. input and output parameters in forecasting of world gas markets development the main results of calculations in regarded optimization models are the optimal volumes of global gas trade by route, production by countries and gas prices. the key modeling input parameters and assumptions are as follows:  volumes of gas supply with corresponding production costs;  gas demand volumes;  capacity of the existing and planned infrastructure for piped gas and lng, including corresponding transportation costs;  volumes and pricing conditions of the concluded contracts for piped gas and lng supply;  prices of alternative energy resources, taking into account prices of co2 emissions and the extent of demand switching for such energy resources. the level of detail in the global supply curve presentation, which contains information on the opportunities for gas production by fields and groups of fields per each forecast year, including corresponding costs, largely determines whether the model reflects the real market with due accuracy. a number of market factors and assumptions about future production volumes and costs (for instance, the rate of decline in gas production in europe, the cost of extracting shale gas and coalbed methane in asia, the prospects of gas production in east * corresponding author: vakulagin@yandex.ru 60 v. kulagin, a. galkina copyright ©2018 assa adv. in systems science and appl. (2018) africa, etc.) greatly affect the relative competitiveness of gas producers. thus, when the development of individual national or regional gas markets, it is necessary to incorporate changes in the balance of supply and demand across all the global markets, even without any need for receiving such a detailed forecast. gas production costs are not a fixed parameter, either: in particular, they are influenced by the development of technologies for the production of conventional and unconventional gas, devaluation/revaluation of currencies, changes in oil and gas market prices and tax burden. thus, since h2 2014, the costs of hydrocarbon production have significantly decreased (fig. 1). fig. 1. indices of operational (uoci) and capital (ucci) costs in the oil and gas production sector. source: ihs cera [1]. in this regard, an important assumption in plotting the gas supply forecast curve, which directly influences the forecast level of gas prices, is the rate of production cost escalation. the global gas demand drivers have an impact on modeling results that is no less important than that of the assumptions for gas supply. the demand for gas, as part of energy demand, shall be determined by both the level of domestic energy consumption (as a whole and by sectors) and the conditions of inter-fuel competition. forecasting energy consumption of countries (nodes) by types of energy resources shall be carried out separately and shall not be deemed the subject-matter of the study described. but the demand for gas in the model shall be refined further in order to consider inter-fuel competition, the opportunities for switching the demand for alternative energy resources are added when certain threshold levels of prices are reached in the process of calculations. here, the coefficients reflecting the assumptions for co2 emission prices are added to the prices of carbon fuels. countries (more precisely, nodes, since in the detailed analysis of certain countries, especially big ones, it is expedient to allocate several nodes within a single country) with the given production capacities and the volume of gas demand in the models under consideration are connected with one another by gas transportation and storage infrastructure. this includes gas pipelines, gas storage facilities, lng plants and lng regasification terminals. the main parameters of gas transportation infrastructure are capacity, transportation costs, their escalation rates, expected service life, and directions of routes (dispatch nodes and destination nodes). thus, the possibilities of gas transportation between the nodes with surpluses and deficit of supply are formed. in view of a large share of transportation costs in the market price of gas, manipulating the possibilities can significantly distort simulated conditions of competition. in particular, in the short run, when analyzing a volatile market environment with price wars, scenarios can become realistic in which only operational costs for gas production and transportation are considered while the correct option for long-term forecasts is to incorporate long-term marginal costs. specifics of long-term forecasting for global gas markets 61 copyright ©2018 assa adv. in systems science and appl. (2018) a part of supply and demand for gas is withdrawn from the free competition environment through the introduction of gas supply contracts (piped gas and lng). gas supplies under contracts can differ from optimum competition alternatives by volumes, routes and price. in the real market swap deals take place, and it is possible to overestimate future demand when signing the long-term contracts. however, long-term contracts, first and foremost, ensure the security of supply through the guarantee of supply and demand and reduction of transaction costs. the increased flexibility of gas supply conditions under long-term contracts also contributes to their role in gas markets. however, on the background of market development the volume of gas supplies under short and medium-term contracts has markedly increased: in 2016, almost 30 % of global lng supplies were made under contracts valid for up to 4 years [2]. in the models, the input parameters characterizing gas supplies under long-term contracts include annual volumes of supplies, contract duration, gas price formula, level of quarterly and annual minimum volumes. gas prices in contracts can be directly indexed to oil prices, the price of the japanese crude cocktail (jcc), prices of petroleum products (gasoil, fuel oil), as well as to spot prices (for example, henry hub, gaspool, nbp, ttf). accordingly, the projected oil prices and spot gas prices to be indexed in contracts are exogenous variables. these prices are also used to calculate the cost of the fuel required to transport lng. the optimization results are the volumes of gas supplies in all directions (can be grouped into import and export forecasts by countries), the gas production required for these supplies (in the same level of detail as gas supply input in the model), gas storage volumes, gas prices as long-term marginal costs of supplying gas to each node and refinement of gas consumption volumes, with consideration of the switching of part of the demand for alternative fuels. based on the calculation results per each node, the following condition is fulfilled: gas production + piped gas and lng import + withdrawal from gas storage gas consumption piped gas and lng export injection into gas storage = 0. in general, the scheme of input parameters and assumptions and output data appears as follows (fig. 2). fig. 2. inputs and outputs in optimization models of global gas market development. source: eri ras. 3. sensitivity analysis and results as the studies and results of the calculations show [3, 4], currently, gas prices are more influenced by oil prices in the asian and european markets, but in the long term, the role of gas-on-gas competition and hence the role of gas production and transportation costs will increase in all key global markets. in the us market, gas production costs already directly affect market gas prices.the prospects for changing directions and the level of global gas trade in the next 25 years will largely be determined by the demand in asia. russian exports to the east will mainly depend on the opportunities for production, the projects implemented and contracts signed. in the european and the cis market the dynamics of demand and the possibilities of competition will 62 v. kulagin, a. galkina copyright ©2018 assa adv. in systems science and appl. (2018) play a key role, while the production and transportation capacities will have sufficient reserves for increasing supplies when needed. the considered approach allows for the shaping of complex scenarios, exploring interactions between market participants and analyzing the sensitivity of national gas markets to various local and global factors (for example, to changes in global or regional demand, the costs and volumes of supply of conventional and unconventional gas, changes in taxation, terms of supply contracts), as well as assessing the competitiveness of projects for gas production and transportation under various conditions. references [1] indices of operational (uoci) and capital (ucci) costs in the oil and gas production sector. ihs indexes [online] available https://www.ihs.com/info/cera/ihsindexes/ [2] the lng industry annual report 2017, giignl. [3] makarov , a. grigoriev, t. & mitrova t. (2016) global and russian energy outlook. moscow, russia: eri ras – acrf. [4] kulagin, v., & mitrova, n. (2015). gazovyj rynok evropy: utrachennye illyuzii i robkie nadezhdy [the european gas market: lost illusions and glimmers of hope]. moscow: national research university higher school of economics. moscow, russia. energy research institute of the russian academy of sciences. [in russian]. adv syst sci appl 2018; 04; 64-73 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/555 solution of one global problem by approach of the parametric control theory abdykappar ashimov1,*, yuriy borovskiy1 1)kazakh national research technical university named after k. satpayev, almaty, kazakhstan e-mail: ashimov37@mail.ru, yuborovskiy@gmail.com received february 12, 2018; revised december 9, 2018; published december 31, 2018 abstract: purpose of this paper is to show the potential of parametric control theory approach for the optimal solution of global problems (in the case of the food problem) and to substantiate the benefits of joint action to solve them on the basis of the developed global multi-country computable general equilibrium model (hereinafter model). the main results of the paper: model verification (testing the possibility of its practical application) of the constructed and calibrated model was successfully carried out by methods of parametric control theory: using sustainability assessments and assessments of sustainability indicators of smooth mappings, which are defined by the model; as well as through conducting of series of counterfactual and forecast scenarios. based on the model a number of parametric control problems were solved in finding the optimal values of the instruments of public policy, which provide economic growth, reducing the gap in economic development between rich and poor regions, as well as increasing the production of agricultural products in the regions. value consists in the developed model and in showing approaches of parametric control theory on providing guidance to address global problems and to substantiate the benefits of joint action within the framework of their decision. keywords: global problem, parametric control theory, computable general equilibrium (cge) model, model verification. 1. introduction the formed world economy (after independence of former colonies, collapse of the soviet union, emergence of new states and economic unions) for many years keeps the imbalances in the economic development of individual countries and regions of the planet. these imbalances formulate global and regional problems (including individual countries problems), without their solution further humanity progressive advance along the path of economic progress is impossible. the most urgent problems are the "north-south" problem the gap in economic and social development between rich and poor countries and the associated food problem, which lies in inability of poor countries to support them completely with even vital food. different mathematical models are widely used to assess the scenarios of food products output in various regions of the world for different periods of time. for example, the econometric model which has been proposed in [1] is used for short-term assessment of grain yield in different parts of the us. the paper [2] contains a description of the international food security assessment model, which is designed to assess food consumption, food access, and food gaps (previously called food needs) in low-income countries through 2025. let us * corresponding author: ashimov37@mail.ru mailto:ashimov37@mail.ru mailto:yuborovskiy@gmail.com solution of one global problem by approach of the parametric control theory 65 copyright ©2018 assa. adv. in systems science and appl. (2018) give as example the basic characteristics of the most ambitious impact model [3]. it is partial equilibrium, multi-commodity, multi-country model which generates projections of global food supply, demand, trade, and prices. impact covers over 46 crops and livestock commodities and it includes 115 countries/regions where each country is linked to the rest of the world through international trade and 281 food producing units. the forecast data until 2050 for the three scenarios (baseline, yield increase, energy shock) are presented in the mentioned paper. it should be noted, that the above and other discussed models in the available literature are mainly used to obtain projections, and they have not been used for evaluation of optimal agreed measures of economic policies of different regions to address the food problem in these regions. however, it is generally recognized that for solution of these global problems it is required to have all humanity coordinated efforts and, in particular, efforts of all countries concerned states. this interest in coordinated efforts can be improved if based on the model tested for the possibility of its practical application (that is, tested for transferring the computational experiments results into the subject area of macroeconomic analysis and the choice of the values of economic policy tools), will: 1) reasonable optimal steps will be offered (steps which will be defined by economic policy instruments) which should be taken in order to solve these global challenges; 2) the justified assumptions that coordinated actions of all states (or at least states of the countries within regional economic union) will give a greater effect for each country than optimal uncoordinated action of this country. in this paper, based on the macroeconomic model developed and tested on the possibility of its practical application, the possibilities of the parametric control theory [4] approach were demonstrated to solve a number of formulated problems of parametric control; and the benefits of joint action to address them are justified. the considered parametric control problems are aimed to evaluate the optimal values of economic instruments that provide economic growth, the convergence of macroeconomic indicators and the growth of regions’ agricultural output. note that the main difference between the parametric control theory and other well-known macroeconomic theories is in using only such mathematical models that have been tested for the possibility of their practical application through a series of original and classical methods. at the same time model verification is carried out not only for its basic calculation, but for optimal scenarios resulting from solving these parametric control problems. 2. the model 2.1. general features of the model in this paper, a static computable general equilibrium model globe 1 [5] is developed to describe the dynamics of functioning of interacting economies of 9 regions: five countries (kazakhstan, russia, belarus, armenia, kyrgyzstan) forming the eurasian economic union (eaeu); european union (eu, as one country); usa; china and the rest of the world (as one country). the economy of each region of the model includes the following 16 industries: 1 – agriculture, forestry and fisheries; 2 – production of crude oil and gas; 3 – metal-working production and machine-building; 4 – metallurgical industry; 5 – education, health, and public administration; 6 – production and transmission of electricity, gas, and hot water; 7 – production of food, beverages, and tobacco; 8 – professional, scientific and technical activities; 9 – other industries; 66 a. ashimov, yu. borovskiy copyright ©2018 assa. adv. in systems science and appl. (2018) 10 – other services; 11 – mining (except for oil and gas production); 12 – construction; 13 – production of textiles, apparel, and leather and related products; 14 – financial services; 15 chemical and petrochemical production; 16 – transport. in addition to the specified producers, present in each region of the model, there are such consumer agents as the households and the state. in the output of products, the industries utilize two production factors which are owned by the households in their region: the labor and the capital. in the model there is also yet another special agent (a.k.a. region) the globe, importing transport services from the rest of the regions and providing these services to other regions for the import of goods. the industries, the households, and the government (within the framework of productions consumption) in each region (except the globe) are composed of a large number of representative price-taker agents, annually solving the appropriate optimization problems. the model, as compared to the base case of the globe 1, was developed as follows. the static globe 1 model was developed to become a dynamic model by describing the following variables by using dynamic equations: technological factors of production functions for the gva of all the industries in the regions, and the supply of factors by the households in the regions. 2.2. conceptual description of the economy and the mathematical model it is assumed that producer agents, household agents and government agents are assumed to be perfect rationality agents. here we list their basic functions. producer agent in its activities annually: • produces one, corresponding to the name of the industry, type of production (subject to the condition of minimizing the costs); • produces gva gross value added (based on the use of factors as the labor and the capital of the households); • exports a part of produces production (subject to the condition of maximizing the profits); • imports intermediate productions from other regions and consumes intermediate production; • pays net tax payments to their government. producer agents solve following two pairs of the nested optimization problems: • minimization costs for the purchase of the intermediate productions and gva costs for a given production output; • minimization costs for the purchase of the production factors at the given output of the final production; • profit maximization from sales within the region and beyond, at the given production output; • profit maximization from exports to various regions at the given level of production exports. household agent in its activities annually: • gains income from the producers in their region on the basis of demand for the belonging to them factors; • consumes the productions from all the regions (according to the solution of the problem of maximizing their utility functions under the budget constraints); • perform savings in the form of investment products based on their income and consumption; solution of one global problem by approach of the parametric control theory 67 copyright ©2018 assa. adv. in systems science and appl. (2018) • pays net tax payments to their government. government of each region (except for globe) in its activities annually: • determines the effective tax rates and receives income in the form of net tax revenue (including revenue from customs duties); • consumes the final productions (government spending); • performs savings in the form of investment products based on its income and expenses. industry, household, and government agents each year jointly solve the following optimization problems: • determination of the optimal share of imports in the consumption of each type of production subject to minimum of costs of domestic and imported components of this production; • determination of the optimal regional structure of each type of imported production subject to minimum of costs of this type of imported production. the conceptual description of the model economy contains the statements of the above optimization problems with the corresponding first order conditions, other equations describing the functions of the agents, balance ratios for prices and quantities (indicators, measured in the prices of the producer), the internal balances in the accounts of the government and the external balances of trade accounts. in the model is used the system of endogenous prices for all the 16 types of productions of each region, including the buyer's and seller's prices, the prices of the exporter and importer, and etc. the calculated values of the prices ensure the performance of the annual balance sheet ratios, providing for: • the equilibrium in factor (labor and capital) markets; • the equilibrium in the markets for each type of production; • bilateral current account balances for each pair of the regions; • the equilibrium of savings (households, governments) and their investments in the regions’ industries. the mathematical model (which has been built on the basis of its conceptual description) was obtained as the result of combining into a single set of equations describing: the first order conditions of all the agents’ optimization tasks and other rules of the agents’ activity, including the equations describing the dynamics of technological factors of production functions and the factors supply. the balance and the auxiliary equations are also included to the mathematical model. 2.3. source database, calibration, and calculation of the model the core of the database of the model is the set of concordant social accounting matrices (sam) of the regions for each year under review (2004-2021), which are demonstrating how the productions flows (in monetary terms) and financial flows are distributed among the producers, households, governments, importers and exporters. the mentioned sam sets for 2004, 2007, and 2011 were extracted with a special converter from the [6]. note that in the database there are no data for the remaining years of the considered historical 2004-2014 period. therefore, for 2005, 2006, 2008-2010, and 2012-2014 the desired sets of sam were calculated using the developed algorithm (algorithm 1) on the basis of the available statistical sources, containing input-output tables (see, e.g., [7], indicators of mutual trade [8], using the basic ratios calculated using the known sam of the nearest last year (2004, 2007, or 2011). for the forecast period (2015-2021), algorithm 2 was used, which allows to calculate the mentioned sam sets on the basis of the following lookahead indicators of the regions provided by the [9]: gdp, total investment, imports of goods, imports of services; exports of goods, exports of services, total government revenues, and total government expenditures. at the same time were used the basic ratios, calculated using the obtained sams for 2014. 68 a. ashimov, yu. borovskiy copyright ©2018 assa. adv. in systems science and appl. (2018) in addition to the obtained sam sets, the source database of the model includes the values of replacement coefficients of various factors in the production function of producers; replacement coefficients of various types of productions in the output functions of producers, utility functions of households and aggregation functions describing the consumption by the agents as well as the initial values of the dynamic equations of the corresponding model. these replacement coefficients were obtained from the gtap 9 database for 2004, 2007, and 2011 and then extrapolated for the remaining years of calculation period of the model (20042021). the calibration is carried out at each start-up of the model. at this stage, on the basis of the formed database using special expressions, computation is performed of all the values of the exogenous variables in the model for the time from 2004 to 2021. the solution of the calibrated equation set of the model (model calculation) is performed with the help of the software which was implemented in the integrated design environment of the gams [10] with the use of the built-in solver path [11]. the results of calculation of the considered baseline scenario of the calibrated model precisely reproduce the statistical and forecast data used in the building of sams for the baseline database of the model. 3. verification (testing the possibility for practical application) of the model verification of the calibrated model was performed by three methods, the first two of which were proposed by the authors within the framework of the development of the parametric control theory. 3.1. assessing the stability of smooth mappings defined by the model the availability of the stability properties of mapping baf : , transferring the values of exogenous parameters ap into the solutions (the values of endogenous variables) suggests preserving the qualitative properties of such mapping at small (in some sense) changes of the mapping. to assess the stability of smooth mapping of the specified type, based on the numerical evaluation of the fulfillment of its conditions, the authors proposed the corresponding set of numerical algorithms for the cases of immersion, submersion, and submersion with a fold [12]. in the experiments, basic mappings of types with 5dim a and 9dim b were considered, whereas the arguments of f were taken the value-added tax rates of the five eaeu countries (kazakhstan, russia, belarus, armenia, and kyrgyzstan) for 2015. as the output variables of f were taken of values gdp in all the 9 regions of the model for 2021. a five-dimensional a box (with the center at the point  51 ,, ppp  , corresponding to the baseline values of the specified tax rates) boundaries are distanced from the values ip to the value of ip5.0 . it should be noted that the time to calculate the implemented mappings stability algorithms will increase approximately exponentially with increasing the dimension of a box ( adim ). this limits the use of this approach, so to obtain a reasonable calculation time the set (of the most important factors used in the solution on the basis of the model of specific problems for macroeconomic analysis and parametric control) was selected. the results of the specified numerical experiments demonstrated the absence of singular points of the f mapping in a box and the stability of this immersion. 3.2. estimation of stability indicators of smooth mappings defined by the model the ),(  pf stability indicator of (defined by the model) mapping baf : at the ap point and for the selected positive  number is the diameter of the image (when f mapping) of the ball with its radius  and with its center at the point p (in relative terms). if solution of one global problem by approach of the parametric control theory 69 copyright ©2018 assa. adv. in systems science and appl. (2018) for all the ap points the numerical assessment of the ),(lim)( 0   pp ff value is uniformly close to zero, then the f mapping, defined by the tested model is assessed on the a set as continuously dependent on exogenous values. in the experiments with the model, as a set, a ball was considered centered at p point corresponding to the baseline values of all the tax rates in all the regions in 2004, whereas endogenous variables tb set – gdp, exports, and imports of all the regions of the model for the fixed computational year t (from 2004 to 2021). as an example, table 1 shows the calculated values of the model stability indicators (in percent) for the base point p and 01.0 . table 1. values of stability indicators for the basic calculation of the model year 2004 2005 2006 2007 2008-2021 )01.0,(pf 0.4536 0.0813 0.0097 0.0060 0.0000 all the specified in table 1 assessments of stability indicators do not exceed 0.4536, which characterizes the stability of the model (in terms of the considered stability indicators) for calculations up to 2021 as sufficiently high. in particular, the acquired for the year 2004 value of the )01.0,(pf indicator means that the image of a sphere centered at the point p (corresponding to the basic values of all the 2004 tax rates) and having the radius of 0.01 (in relative terms) in the calculation of the model is transformed into a set with the diameter of 0.4536 (in relative terms) for the output variables values (gdp, exports, and imports of all the regions in 2021). various options of the calculated values of the limit indicators )(pf were close enough to zero at 0001.0 , which evaluates the tested mapping as a continuous one in a box. 3.3. implementation of counterfactual and forecast scenarios according to the well-known macroeconomic theory, the reduction of taxes levied on producers and consumers, as well as the increased state demand for consumer products will increase the country’s output and gdp. as part of its testing, based on the model counterfactual and forecast scenarios were calculated to assess the implementation of the provision of the theory. in particular, the scenario was performed featuring a 10% decrease in the effective tax rates of value added tax, tax on the producers' income and 10% increase in government consumption in each eaeu country. the results of the calculation of this scenario demonstrated an changes in gva in each sector in the relevant country (ranging from -3.85% to 6.16%) and increases in gdp in each region ranging from 0.0279% in 2009 for global economy to 0.7715% in 2012 for eaeu, compared with the observed data. the above results of the three test methods make it possible to draw the conclusion of a successful verification of the tested model. 4. solutions to parametric control problems on the basis of the model as part of the proposal of optimal measures to reduce the gap in economic development among the regions, as well as of the action to increase agricultural output, a series of parametric control problems were solved on the basis of the model. in those problems, the values of all the uncontrolled exogenous variables of the model match the baseline forecast of the variables. hereinafter indices correspond to the region number ( 9,,1i – the regions of the model): 1 – kazakhstan, 2 – russia, 3 – belarus, 4 – armenia, 5 – kyrgyzstan, 6 – european union, 7 – usa, 8 – china, 9 – the rest of the world; 10 – eurasian economic union, and 11 – world economy. 70 a. ashimov, yu. borovskiy copyright ©2018 assa. adv. in systems science and appl. (2018) statement of the ip problem ( 11,,1i ): to identify on the basis of model the values of the control parameters u(t) (effective tax rates for producers revenues, sales and customs duty taxes differentiated by production and industry sector: shares of public spending for consumption) that provide the maximum value of the ik criterion (2) (4) with constraints on control instruments of )()( tutu  type and with restrictions (1) for some of the endogenous variables. here, for 2021,,2015t ; u(t) – is a box with the center at the point of base values u(t) and boundaries spaced at ± 10% of the baseline values. for ip , ( 9,,1i ) problems the u(t) control parameters are the specified public policy tools only for i-th region, for the 10p problem – those of the five eaeu countries, and for the 11p problem – those of all of the nine regions of the model in the aggregate. the restrictions for the endogenous variables of the model in all the ip problems look like: )()( tqvaptqvap rr  , 9,,1r , 2021,,2015t . (1) here: )(tqvapr is gdp per capita in the region r, )(tqvapr are the basic values of this indicator. meeting the conditions (1) will guarantee that subject to the solving by any region (or by the 5 countries making up the eaeu, or by all the regions of the model) of their ip problems, the values )(tqvapr will not decrease (compared to the baseline forecasts), both in their specific regions and also in all the other regions of the model. in each of the first formulated ip problems, ( 9,,1i ), pursuance of the optimal fiscal policy aimed at the economic growth is supposed and the growth of agricultural output during 2015-2021 in their specific regions is supposed, whereas the behavior of the rest of the regions of the model is characterized by the baseline values of the exogenous variables. therefore, in these problems the ik criterion is recorded as:     2021 2015 )()( t iii ttqxattqvak , 9,,1i (2) where )(ttqvai is gdp per capita rate, )(ttqxai is the rate of output of industry (agriculture, forestry, and fisheries) per capita in the region i in year t (in current usd) and  is adjusted factor ( 3.0 ). it is assumed that in the 10p problem carrying out a coordinated optimal fiscal policy within the framework of the five eaeu countries in 2015-2021 as aimed at the economic growth, the growth of eaeu agricultural output, and the approximation of the per capita gdp values in each of the five eaeu countries to the respective values in the united states. (usa, (r = 7) is the region having the highest value of gdp per capita of all the regions according to the calculation in the model until 2021). therefore, the problem criterion 10k was selected as follows:                      5 1 2021 2015 7 7 5 1 2021 2015 101010 )( )()( )()( r t r r r r t tqvap tqvaptqvap ttqxattqvak . (3) here )(10 ttqva is the rate of gdp per capita of eaeu in the year t; is the rate of output of agriculture, forestry, and fisheries per capita in the eaeu in year t; )(tqvapr is gdp per solution of one global problem by approach of the parametric control theory 71 copyright ©2018 assa. adv. in systems science and appl. (2018) capita in the region r in year t; r is the weighting factor, its value equal to 1.0r for medium-developed regions (kazakhstan and russia), and 1r for all other regions eaeu members;  is adjusted factor ( 1 ). in the 11p problem it is considered the hypothetical possibility of coordinated optimal fiscal policy within the framework of all regions of the model in 2015-2021, aimed at the global economic growth, the growth of world agricultural output, and the approximation of the gdp values per capita in each region to the respective values in the united states. therefore, the problem 11k criterion was selected as follows:                      9 1 2021 2015 7 7 9 1 2021 2015 111111 )( )()( )()( r t r r r r t tqvap tqvaptqvap ttqxattqvak . (4) here )(11 ttqva is the rate of the world gdp per capita in the year t; )(11 ttqxa is the rate of the global output of agriculture, forestry, and fisheries per capita in year t; the value of factor 1.0r for  8,6,2,1r , 07  and 1r for all other r values. the formulated problems were solved numerically using the optimization algorithm provided by gams [13]. the results of their resolve as the increments of the average gdp value of and the increments of the average output of agriculture, forestry and fisheries in the regions for 2015-2021 (per capita, in percent relative to the baseline scenario) are shown in tables 2 and 3. table 2. the average value increments of the output in agriculture, forestry, and fisheries in the regions following the solutions of the 11 ip problems problem number of region r = 1 r = 2 r = 3 r = 4 r = 5 r = 6 r = 7 r = 8 r = 9 1p 3.8219 0.0026 0.0224 0.0007 0.7509 0.0009 0.0012 0.0028 0.0007 2p 0.0353 6.7404 0.5553 0.8582 0.8786 0.0059 0.0078 -0.0008 -0.0056 3p 0.0107 0.0121 6.8270 0.0056 0.7580 0.0014 0.0011 0.0010 0.0013 4p 0.0047 0.0011 0.0018 2.8709 0.0076 0.0000 0.0001 0.0001 0.0000 5p 0.0095 0.0011 0.0065 0.0002 1.3197 0.0000 0.0001 0.0008 0.0001 6p 0.4617 0.4226 0.8165 0.9414 0.6130 10.1658 0.3555 0.1951 0.4326 7p 0.0582 0.0856 0.0863 0.0586 0.7361 0.1427 5.7370 0.0741 0.1958 8p 0.0767 0.0818 0.0397 0.0370 0.6741 0.0415 0.2032 6.7602 0.1166 9p 0.2160 0.2357 0.1609 0.2441 0.0277 0.1398 0.3250 0.0760 6.3946 10p 4.0196 6.8185 11.6397 3.7374 1.3891 0.0084 0.0082 0.0002 -0.0054 11p 4.9816 7.8776 13.1665 5.3487 4.5170 10.9244 7.0752 7.6066 6.8962 table 3. the average values of increments of gdp of the regions following the solutions of the 11 ip problems problem number of region r = 1 r = 2 r = 3 r = 4 r = 5 r = 6 r = 7 r = 8 r = 9 1p 0.5949 0.0008 0.0089 0.0005 0.0000 0.0005 0.0000 0.0007 0.0003 2p 0.0298 1.4875 0.4111 0.6758 0.0574 0.0000 0.0000 0.0000 0.0000 3p 0.0073 0.0073 0.7963 0.0042 0.0000 0.0005 0.0000 0.0000 0.0009 4p 0.0002 0.0001 0.0001 1.4413 0.0000 0.0000 0.0000 0.0000 0.0000 5p 0.0002 0.0001 0.0000 0.0000 1.5198 0.0000 0.0000 0.0000 0.0000 6p 0.3606 0.4084 0.4073 0.9281 0.8454 1.9219 0.1035 0.1062 0.2595 7p 0.0310 0.0498 0.0448 0.0656 0.0000 0.0611 0.7446 0.0265 0.1107 8p 0.0678 0.0256 0.0210 0.0374 0.8326 0.0390 0.0402 1.1463 0.0467 9p 0.1206 0.1123 0.0630 0.2134 0.2164 0.0885 0.0558 0.0939 0.9169 72 a. ashimov, yu. borovskiy copyright ©2018 assa. adv. in systems science and appl. (2018) 10p 0.7659 1.4433 4.3875 2.1344 3.0802 0.0000 0.0000 0.0003 0.0000 11p 1.4083 2.0763 5.2096 3.7368 5.9613 2.1219 0.8123 1.4071 1.1225 the obtained optimal values of economic policy tools of the discussed above parametric control problems were tested for the possibility of their practical implementation as follows. the scenario options of the calibrated model for the specified optimum values of the tools were tested by the three methods referred to in section 3. in all cases, the results were obtained similar to the above ones specified in section 3 for the baseline scenario, namely: • permissible values of assessments of stability indicators; • absence of singular points of the considered mappings in their appropriate domains and the stability of these mappings and • coincidence of the results of the 2015-2021 forecast scenarios to the main provisions of the macroeconomic theory. analysis of bold type numbers in each column of tables 2 and 3 shows that within problems ip , 11,,1i , the approach of parametric control at the level of all regions (the problem 11p ), as well as at the level of the five countries of the eaeu (the problem 10p ), for the most part gives greater effects for separate region in comparison with parametric control at the level of separate region (problems ip , 9,,1i ). additionally, as a result of solving the problem 11p was reached the alignment of regions economic development, which is characterized by a decrease ratio of the maximum to minimum per capita gdp of all regions to 4.86% in 2021 as compared to the option without control; 4.17% to 2021 compared to 2015. the growth rate of per capita gdp in 2021 was also achieved in comparison with 2015 at 29.32%, 30.86% and 20.10% respectively for kyrgyzstan, armenia, and rest of the world (less developed regions). generally, across the world, as a result of solving the problem 11p the growth of gdp per capita in 2021 is 1.31%, while industry output growth in agriculture, forestry and fisheries per capita is 7.89% compared to the base case. at that growth of global gdp per capita in 2021 is 22.38%, and industry global output growth in agriculture, forestry and fisheries per capita is 38.41% compared to 2015. analysis of the problems ip solving results and the results of relevant tests shows the high potential of parametric control approach to make recommendations for optimal coordinated government economic policy at the global level and at the level of regional economic union. 5. conclusion 1. the results of the development of the dynamic global multi-country computable general equilibrium model are presented. the model is calibrated using the sams, as extracted from gtap data base, and obtained by extrapolation of extracted sams. 2. the results of the model verification with using of the three approaches are presented. 3. there are presented problem statements and the solution results of the 11 parametric control problems which are aimed at economic growth, output growth of agricultural production and reducing the gap in economic development between the rich and poor regions. these results demonstrate the effectiveness of the approach of the parametric control theory to develop recommendations to address global problems. solution of one global problem by approach of the parametric control theory 73 copyright ©2018 assa. adv. in systems science and appl. (2018) references [1] sakamoto, t., gitelson, a., & arkebauer, t. (2013). modis-based corn grain yield estimation model incorporating crop phenology information, remote sensing of environment, 131, 215-231. [2] rosen, s., meade, b., & murray, a. (2015) international food security assessment, 2015-2025. united states department of agriculture: economic research service. [3] rosegrant, m., tokgoz, s., & bhandary, p. (2012). the new normal? a tighter global agricultural supply and demand relation and its implications for food security. american journal of agricultural economics, april, 1-7. [4] ashimov, a., sultanov, b., adilov, zh., borovskiy, yu., novikov, d., alshanov, a., & ashimov, as. (2013) macroeconomic analysis and parametrical control of a national economy. new york, ny: springer. [5] globe 1. (2016). applied general equilibrium modelling. [online]. available www.cgemod.org.uk/globe1.html [6] gtap 9 data base. (2017). gtap data bases. [online]. available https://www.gtap.agecon.purdue.edu/databases/v9/default.asp [7] world input-output database. (2017). [online]. available http://www.wiod.org/home [8] world integrated trade solution. (2017). [online]. available http://wits.worldbank.org [9] world economic outlook databases. (2017). international monetary fund. [online]. available http://www.imf.org/external/ns/cs.aspx?id=28 [10] gams. (2017). [online]. available www.gams.com [11] ferris, m. & munson, t. (2015). path 4.7. [online]. available https://www.gams.com/latest/docs/s_path.html [12] ashimov, a., adilov, zh., alshanov, r., borovskiy, yu., & sultanov, b. (2014). the theory of parametric control of macroeconomic systems and its applications (i), advances in systems science and application, 14(1), 1-21. [13] nlpec. (2015). gams. [online]. available https://www.gams.com/latest/docs/s_nlpec.html http://www.cgemod.org.uk/globe1.html http://www.wiod.org/home http://wits.worldbank.org/ http://www.imf.org/external/ns/cs.aspx?id=28 http://www.gams.com/ adv syst sci appl 2020; 02:119–130 published online at https://ijassa.ipu.ru. structural change in multisector monopolistic competition model igor pospelov1,2, stanislav radionov1,3,4* 1frc csc of the ras, moscow, russia 2national research university higher school of economic, moscow, russia 3lebedev physical institute of the ras, moscow, russia 4financial research institute of the ministry of finance of russia, moscow, russia abstract: we present a natural generalization of the dixit-stiglitz monopolistic competition model (dsm) — we assume that there is a continuum of industries, each of them described as in dsm, and each characterized with its own elasticity of substitution. although firms in all industries share the same level of productivity and costs, exogenous technological progress leads to non-trivial reallocations of labor and production to industries with lower elasticities of substitution. thus the model, despite is simplicity and the absence of additional assumptions about industry structure, generates the structural changes described in the economic growth literature. keywords: dixit-stiglitz model, monopolistic competition, structural change, market reallocations. 1. introduction in modern economic growth, simon kuznets wrote: “we identify the economic growth of nations as a sustanined increase in per capita or per worker product, most often accompanied by an increase in population and usually by sweeping structutal changes. in modern times these were changes in the industrial structure within which product was turned out and resources employed — away from agriculture toward nonagricultural activities, the process of industrialization...” existence of these structural changes was considered by kuznets as one of the main stylized facts of development. indeed, the loss of relative importance of agricultucal sector in favor of industrial sector and then to services sector is well documented, see for example [3], [5], [10]. a number of economic models were developed to describe these structural changes and their relation to economic growth. the simplest mechanism, which generates structural changes, was proposed in [4] — different economic sectors grow at different rates because they have different rates of technological progress. this idea was developed in [13], [15] and [2] among many others. our work contributes to another strand of literature based on the idea that structural changes are driven not by differences in production technologies but by demand factors, namely by non-homotheticity of consumer preferences. for example, in [12] authors assume that household consumption consists of agricultural, manufacturing and services goods. a special form of utility function generates reallocations of consumption: after the consumption of certain amount of agricultural good, the household starts to demand other goods, moreover, it consumes service goods only after some level of manufacturing ∗corresponding author: saradionov@edu.hse.ru 120 i.g. pospelov, s.a. radionov consumption is reached. the model in [8] is based on the idea that new goods are introduced in economy as a luxury (a good with high income elasticity), but with time income elasticity decreases and good becomes necessity. there are also several models incorporating, for example, financial development, demographic transition, urbanization, migration, production organization, human capital, fertility, see [1] for a survey. as we can see, all these models use rather special assumptions about industry structure. the model we present in this paper, on the other hand, is very simple and free of special assumptions, but nevertheless is able to generate the reallocations described by kuznets. the key idea of our work is the same as in [8] — the main difference between goods is income elasticity, and structural change is a decline of industries producing goods with low income elasticities and rise of the ones producing goods with high income elasticities. but the model in [8] is highly stylized — there is not much said about the production side of the economy, so it remains unclear what kind of industry structure generates new goods with high income elasticities and why income elasticities of old goods decrease with time. we, on the other hand, describe both consumption and production sides of economy, thus making our model more microfounded. we present a simple monopolistic competition model, based on the constant elasticity of substitution utility function, proposed in [7]. we assume there is a continuum of industries each of them described as in [7]. the only difference between industries is intra-industry elasticity of substitution. to the best of our knowledge, this description of production side of economy, although rather straigthforward, is new in the literature. similarly to [8], we assume that “simpler” the commodity, higher the elasticity of intra-industry substitution. this assumption is in line with empirical studies on product differentiation, see for example [11] for the calculations and also [9] for the discussion. as in [8], industries with high elasticities of substitution may be interpreted as agricultural, with middle elasticities — as manufacturing, with low elasticities — as services. technological progress in our framework is modeled naturally as a decrease of variable costs of a firm (an increase of workers’ productivity) at the expense of an increase of fixed costs (which may be interpreted as investments). although firms in all industries share the same levels of productivity and costs, labor and production flows from less differentiated (higher elasticity of substitution) to more differentiated (lower elasticity of substitution) goods. thus our model, despite its simplicity and with no additional assumptions on industry structure, generates kuznets structural changes. 2. model 2.1. model setup there is a continuum of industries indexed with ρ ∈ (0, 1). for every industry, preferences of aggregate consumer are given by a ces function: v (ρ) = n(ρ)∑ i=1 ci (ρ) ρ  1 ρ , (2.1) where n(ρ) is the number of firms working in this industry, ci(ρ) is the quantity of consumed good produced by i-th firm. elasticity of substitution between goods in a given industry is σ = 1/(1− ρ). so, if ρ is high then the elasticity of substitution in this industry is high and goods in this industry are weakly differentiated and vice versa. preferences of the aggregate consumer over the goods in the continuum of industries are given by the function copyright © 2020 assa. adv syst sci appl (2020) structural change 121 u = 1 ν ∫ 1 0 a (ρ)v (ρ)ν dρ, (2.2) where a(ρ) ≥ 0 may be interpreted as consumer’s preference for the products of industry ρ. this is also needed for the dimension correctness, because goods in different industries can have different units of measurement. firms are assumed to have both fixed and variable costs, so in order to produce ci(ρ) units of good, i-th firm has to use li (ρ) = α ci (ρ) + f (2.3) units of labor. productivity level 1/α and fixed costs f are assumed to be equal across firms. the model becomes highly complicated without this assumption (see, for example [6], who deal with heterogeneous markets). full stock of labor in the economy is denoted by l, the wage, the same for all workers, is w. so the aggregate consumer has wl units of income. 2.2. market equilibrium and social welfare consider the aggregate consumer’s problem of maximizing utility given budget constraint: u = 1 ν ∫ 1 0 a (ρ)  n(ρ)∑ i=1 ci (ρ) ρ  1 ρ  ν dρ→ max, ∫ 1 0 n(ρ)∑ i=1 ci (ρ) pi (ρ) dρ ≤ wl. to solve it, we form a lagrange function in the following form: l = 1 ν ∫ 1 0 a (ρ)  n(ρ)∑ i=1 (ci (ρ)) ρ  1 ρ  ν dρ+ ξν−1 wl− ∫ 1 0 n(ρ)∑ i=1 ci (ρ) pi (ρ) dρ  . first of all, we should check if the problem has a solution, for this the integrand must be concave. applying sylvester’s criterion, we conclude that the function is concave if the two following conditions are valid: ν < 1 and ν < ρ. so we have two possibilities: if ν ∈ (0, 1), then the set of possible values of ρ must be restricted to (ν, 1), or we can assume that ν < 0. we will consider only the second case, because, as we will see later, the first case leads to some very unpleasant degeneracies at the point ν = ρ. from an economic point of view, the first case is problematic because we have to assume a relationship between parameters which does not have an economic meaning. so, from now on, ν is assumed to be negative. fixing some j ∈ {1 . . . n} and equalizing the partial derivative of lagrange function with respect to cj(ρ) to zero, we get the optimum level of consumption of cj(ρ): cj (ρ) = ( pj (ρ) ξ ν−1 a (ρ)v (ρ)ν−ρ ) 1 ρ−1 . (2.4) we involute both sides of this equality to the power ρ and find a sum over j = 1 . . . n(ρ), and after some calculation we get copyright © 2020 assa. adv syst sci appl (2020) 122 i.g. pospelov, s.a. radionov v (ρ) = ξ a (ρ)− 1 ν−1 p (ρ) 1 ν−1 , (2.5) where p (ρ) ρ ρ−1 = n(ρ)∑ i=1 pi (ρ) ρ ρ−1 (2.6) is the price index associated with the goods index v (ρ) as in [7] or [14]. taking (2.5) into account, (2.4) may be rewritten in the following form: cj (ρ) = pj (ρ) 1 ρ−1 v (ρ)p (ρ)− 1 ρ−1 . (2.7) now consider the behavior of a firm. first of all, it is obvious that due to the fact that all firms have the same levels of fixed and variable costs, and the symmetry of consumer’s preferences over the goods in a fixed industry, all firms in the industry face the same consumer’s demand and set the same price. taking (2.3) into account, the problem of firm in the industry ρ has the form π (ρ) = p (ρ) c (ρ)− (c (ρ)α + f)w → max . (2.8) substituting (2.7) and finding maximum with respect to p(ρ), we find the price set by firm: p (ρ) = αw ρ , (2.9) like, again, in [14]. now we will impose a free entry condition and demand firm profit to be equal to zero. substituting (2.5), (2.6), (2.9) into (2.8), after some calculations we get the following expression for the number of firms in the industry ρ: nm(ρ) = ( (1− ρ)α ν ν−1 ξ w 1 ν−1 ρ ν ν−1fa(ρ) 1 ν−1 ) (ν−1)ρ ν−ρ (2.10) substituting (2.10) to consumer’s budget constraint, we get the equation defining ξ: ∫ 1 0 a (x)− x ν−x w x ν−xf− ν (x−1) ν−x α ν x ν−x ξ (ν−1)x ν−x ( x− ν ν−1 − x− 1 ν−1 ) ν (x−1) ν−x x− ν ν−1dx = l. (2.11) unfortunately, ξ can be found from (2.11) only numerically. substituting (2.9) into (2.8) and setting firm profit to zero, we find that firm output may be rewritten in the following form: cm (ρ) = fρ α (1− ρ) , (2.12) which is again in line with [7]. now consider the problem of the benevolent social planner, who maximizes consumer utility with respect to technological limitation. copyright © 2020 assa. adv syst sci appl (2020) structural change 123 u = 1 ν ∫ 1 0 a (ρ)  n(ρ)∑ i=1 ci (ρ) ρ  1 ρ  ν dρ, α ∫ 1 0 n(ρ)∑ i=1 ci (ρ) dρ+ f ∫ 1 0 n (ρ) dρ ≤ l. form the lagrange function in the following form: l = 1 ν ∫ 1 0 a (ρ)  n(ρ)∑ i=1 ci (ρ) ρ  1 ρ  ν dρ+ λ l− α ∫ 1 0 n(ρ)∑ i=1 ci (ρ) dρ+ f ∫ 1 0 n (ρ) dρ  . differentiating with respect to c(ρ) and n(ρ), after some calculations we get nw (ρ) = ( (1− ρ)α ν ν−1λ ρfa(ρ) 1 ν−1 ) (ν−1)ρ ν−ρ , (2.13) cw (ρ) = fρ α (1− ρ) . substituting (2.12) into the planner’s technological constraint, we get the formula defining λ:∫ 1 0 α ν x ν−xa (x)− x ν−x λ (ν−1)x ν−x f− ν (x−1) ν−x x− (ν−1)x ν−x (1− x) ν (x−1) ν−x dx = l. (2.14) similarly to (2.11), it can be solved only numerically. as we can see, outputs in market equilibrium and in the social welfare problem are the same, but the numbers of firms are different. this proves the following proposition 1. for any ν ∈ (−∞, 0) and any parameters of the economy l, f, α, w market equilibrium is inefficient. thus, in the market equilibrium consumer cannot optimally distribute her expenses across industries, but can do it within an industry. recall that in one-sector dixit-stiglitz model the equilibrium is efficient, but in the presence of second market of homogenous product (as it was proposed in the original paper), it is not. this result is disappointing, but rather expectable — market efficiency is rare thing in the presence of monopolists. it is also worth noting that this result cannot be considered as a trivial corollary of general theorems about the efficiency in monopolistic competition models proven, for example, in [6] and [16], because utility function (2.1)–(2.2) does not belong to the class of variable elasticity of substitution utility functions analyzed in these papers. figure 1 shows the distribution of the number of firms in the market equilibrium and in social welfare state for arbitrarily chosen parameters of the economy α = 0.01, f = 0.5, l = 30, a(ρ) = 1, w = 1 and consumer’s preferences characterized with ν = −1. figure 2 shows the distribution of firms in the same economy and for consumer’s preferences characterized with ν = −5. copyright © 2020 assa. adv syst sci appl (2020) 124 i.g. pospelov, s.a. radionov fig. 2.1. ν = −1. fig. 2.2. ν = −5. as we see, smaller ν, closer the market equilibrium to the social welfare state. note that as ν tends to −∞, utility function (2.2) converges to leontieff-like minimum function: u−∞ = min ρ∈(0,1) v (ρ). (2.15) this observation leads us to proposition 2. for consumer’s preferences defined by (2.15), market equilibrium is efficient. proof. as ν tends to −∞ in (2.10) and (2.11), we get n∗m (ρ) = ( fρ ξ (1− ρ)α )−ρ , n∗w (ρ) = ( fρ λ (1− ρ)α )−ρ ,∫ 1 0 ξxf−x+1αxx−x (−x+ 1)x−1 dx = l,∫ 1 0 λxf−x+1αxx−x (−x+ 1)x−1 dx = l. obviously, λ = ξ and n∗m(ρ) = n∗w (ρ). � 2.3. effects of technological progress and population growth consider the effects of technological progress in this model. we assume that in our framework technological progress means an increase of workers’ productivity 1/α. this, however, doesn’t come without cost — we assume that fixed costs f also increase. thus, progress is due to the increase in capital expenditures. figures 3–6 show the effect of technological progress in economy with parameters α = 0.01, f = 0.5, l = 30, w = 1 and consumer’s preferences defined by parameters a(ρ) ≡ 1, ν = −∞, on the distribution of the number of firms, output and labor. technological progress is modeled by decreasing α 3 times and increasing f 3 times. the solid line is before progress, dashed — after. conclusions are summed up in the following copyright © 2020 assa. adv syst sci appl (2020) structural change 125 fig. 2.3. number of firms. fig. 2.4. number of workers. fig. 2.5. output. fig. 2.6. ratio of outputs. proposition 3. with decrease of variable costs and increase of fixed costs: α′ = 1 k α, f ′ = kµf, k > 1, µ > 0, for any parameters of economy and consumer’s preferences of the type (2.1)-(2.2) with a(ρ) = 1 for any ρ, 1. for µ ≤ 1 output increases in all industries ρ ∈ (0, 1), 2. number of workers increases in industries with ρ ∈ (0, ρ∗) and decreases for industries with ρ ∈ (ρ∗, 1) for some ρ∗ ∈ (0, 1). 3. if µ+ ν ≤ 0, consumer’s utility increases. proof. 1. obviously, with the change of costs, ξ will change as well: ξ′ = θξ, where θ can depend on µ. we need to prove that c (ρ)n (ρ) < c′ (ρ)n′ (ρ). substituting the expressions from (2.11) and (2.12), we get ρ f ′ α′ (1− ρ)w ( (1− ρ)α ν ν−1 ξ w 1 ν−1 ρ ν ν−1fa(ρ) 1 ν−1 ) (ν−1)ρ ν−ρ = k− ν(µ+1)(ρ−1)+ρ ν−ρ θ (ν−1)ρ ν−ρ × × c (ρ)n (ρ) > c (ρ)n (ρ) . (2.16) copyright © 2020 assa. adv syst sci appl (2020) 126 i.g. pospelov, s.a. radionov note that −ν(µ+1)(ρ−1)+ρ ν−ρ > 0 and (ν−1)ρ ν−ρ > 0 for all ν < 0 and ρ ∈ (0, 1). so if we prove that θ > 1, inequality (2.16) will be proven. in order to prove that θ > 1, consider budget constraint (2.11) of economy after technological progress, which may be rewritten in the following form: ∫ 1 0 k −ν(µx−µ+x) ν−x θ (ν−1)ρ ν−ρ (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx = l f , (2.17) where a is a combination of parameters of the model. budget constraint for economy before technological progress is then∫ 1 0 (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx = l f . (2.18) note that if we prove that∫ 1 0 k −ν(µx−µ+x) ν−x (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx < l f , (2.19) that will mean exactly that for (2.17) to be true, θ must be greater than 1. due to (2.18), inequality (2.19) may be rewritten in the following form:∫ 1 0 ( 1− k− ν(µx−µ+x) ν−x ) (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx > 0. (2.20) denote the integrand in (2.20) as g(x). figure 6 shows the graph of g(x). this function is positive for ρ > µ 1+µ and negative otherwise. fig. 2.7. g(x) for µ = 0.7. we will prove that if µ ≤ 1, than g(x) > −g(1− x) for all x ∈ ( µ 1 + µ , 2 µ 1 + µ ) , (2.21) which proves (2.20). note that for µ > 1 it doesn’t have to be true. consider the function h(x) = − g(x) g(1−x) . analysis of this function shows that limx→ µ 1+µ h(x) = 1 and it is copyright © 2020 assa. adv syst sci appl (2020) structural change 127 monotonically increasing. so h(x) > 1 for ρ ∈ (0.5, 1) and hence (2.21) holds, hence (2.20) holds, hence θ > 1, which concludes the proof. 2. we will prove that l′(1)n′(1) < l(1)n(1), l′(0)n′(0) < l(0)n(0) and l · n is a monotonic function of ρ. limρ→1 l(ρ)n(ρ) = α ν ν−1 ξ w 1 ν−1 , so l′(ρ)n′(ρ) l(ρ)n(ρ) = k− ν ν−1 θ. we need to prove that θ < k ν ν−1 . (2.22) assume otherwise: θ ≥ k(ν/(ν−1)) and substitute it to the budget constraint in the form l = ∫ 1 0 k− ν(µx−µ+x) ν−x θ (ν−1)ρ ν−ρ (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx ≥ ≥ ∫ 1 0 k νµ(1−x) ν−x (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx >∫ 1 0 (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx = l this contradiction proves (2.21) and the fact that l′(1)n′(1) < l(1)n(1). limρ→0 l(ρ) = f , so obviously l′(0) > l(0). finally, analysis of the derivative of l shows that is indeed monotonically increasing. 3. consumer’s utility may be written in the following form: u ′ = 1 ν ∫ 1 0 ( c′ (ρ)n′ (ρ) 1 ρ )ν dρ = 1 ν ∫ 1 0 ( k− µρ−µ+ρ ν−ρ θ ν−1 ν−ρ c (ρ)n (ρ) 1 ρ )ν dρ > > 1 ν ∫ 1 0 k ν(µρ−µ−ν+ρ) ν−ρ ( c (ρ)n (ρ) 1 ρ )ν dρ > 1 ν ∫ 1 0 ( c (ρ)n (ρ) 1 ρ )ν dρ = u. here we used (21) in the first inequality and the assumption ν ≤ −1 in the second one (µρ− µ− ν + ρ > −µ− ν > 0). � there are few things regarding these results we should point out. first, the decrease of the number of firms in all industries, which can be seen in figure 3, is expectable result of increasing fixed costs and may be interpreted as creation of big corporations. the second part of proposition 3 is the most important to us: it means that because of technological progress, labor flows from less differentiated to more differentiated industries. this reallocation is exactly of the type described in the citation in the beginning of the article and the main feature of this model. figure 6 indicates another good property of output: the rate of output growth is higher in more differentiated industries, which is in line with what we observe in the modern economy. we do not include a proof because it can be done trivially by analyzing the derivative of the ratio. note that another important indicator, output per worker, does not show nontrivial dynamics in our framework: it can be easily calculated that c(ρ)/l(ρ) = ρ/α, so with technological progress this ratio increases in all industries equally. now consider the effect of population growth. figures 7, 8, 9 show the effects of population growth of 150% in the same economy as in the previous figures, but everything is now per capita. proposition 4. with the population growth, for any parameters of economy and consumer’s preferences of the type (2.1)-(2.2) with a(ρ) = 1 for any ρ ∈ (0, 1), 1. output per capita decreases in industries with ρ ∈ (0, ρ1) and increases in industries with ρ ∈ (ρ1, 1) for some ρ1 ∈ (0, 1). 2. labor per capita decreases in industries with ρ ∈ (0, ρ2) and increases in industries with ρ ∈ (ρ2, 1) for some ρ2 ∈ (0, 1). copyright © 2020 assa. adv syst sci appl (2020) 128 i.g. pospelov, s.a. radionov fig. 2.8. number of firms per capita. fig. 2.9. output per capita. fig. 2.10. number of workers in industry per capita. proof. 1. denote l′ = kl, k > 1, ξ′ = δξ. first, we shall prove that δ > 1. assume otherwise, δ ≤ 1, then l′ = kl = ∫ 1 0 δ (ν−1)x ν−x (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx ≤ ≤ ∫ 1 0 (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx = l, which contradicts to the assumption k > 1. next, we prove that δ > k. to do so, consider budget constraint in the form (2.17): l′ = kl = ∫ 1 0 δ (ν−1)x ν−x (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx < < δ ∫ 1 0 (1− x) ν (x−1) ν−x x− ν x ν−1a (ν−1)x ν−x dx = δl. copyright © 2020 assa. adv syst sci appl (2020) structural change 129 limρ→0 c(ρ)n(ρ) l = 0, so we have to analyze derivatives: limρ→0 ( c(ρ)n(ρ) l )′ = f l and f l′ < f l , hence c(ρ)n′(ρ) l′(x) < c(ρ)n(ρ) l(x) for small ρ’s. limρ→1 c(ρ)n(ρ) l = α 1 ν−1 ξ w 1 ν−1 l and xi′ l′ = δξ kl > ξ l , hence c(ρ)n′(ρ) l′(x) > c(ρ)n(ρ) l(x) for big ρ’s. analysis of derivative of the function c(ρ)n(ρ) l(x) shows that it is monotonically increasing. 2. similarly to 1: limρ→0 c(ρ)n(ρ) l = f l , limρ→1 c(ρ)n(ρ) l = α ν ν−1 ξ w 1 ν−1 l . � as we can see, and as it might be expected, population growth leads the economy to, in a sense, opposite direction compared to economic growth, consumer now cares more about “simple” homogenous goods. this behavior may be interpreted in the following way: as the population increases, it becomes harder and harder to “feed” them, so people start to care less about luxury (differentiated goods) and more about simple material goods. 3. conclusion we developed a simple and natural model, which generalizes the dixit-stiglitz monopolistic competition model. it provides a natural way to describe technological progress, which leads to non-trivial labor and product reallocations. these reallocations may be interpreted as transfers from “simple” homogeneous to “more complicated” differentiated goods (from agricultural to manufacturing goods and then to services). under some assumptions, it leads to an increase in consumer utility. so, despite its simplicity, our model is able to reproduce essential features of modern economies, described in economic growth literature. references 1. d. acemoglu. introduction to modern economic growth. princeton university press, 2007. 2. d. acemoglu and v. guerrieri. capital deepening and nonbalanced economic growth. journal of political economy, 116(3):467–498, june 2008. 3. r. c. allen. economic structure and agricultural productivity in europe, 1300–1800. european review of economic history, 4(1):1–25, 2000. 4. w. j. baumol. macroeconomics of unbalanced growth: the anatomy of urban crisis. economic journal, 57(3):415–426, june 1967. 5. t. boppart. structural change and the kaldor facts in a growth model with relative price effects and non-gorman preferences. econometrica, 82(6):2167–2196, 2014. 6. s. dhingra and j. morrow. monopolistic competition and optimum product diversity under firm heterogeneity. technical report, london school of economics, 2012. 7. a. k. dixit and j. e. stiglitz. monopolistic competition and optimum product diversity. american economic review, 67(3):297–308, 1977. 8. r. foellmi and j. zweimüller. structural change, engel’s consumption cycles and kaldor’s facts of economic growth. journal of monetary economics, 55(7):1317–1328, 2008. 9. h. gray and j. martin. on the meaning and measurement of product differentiation in international trade: a reply. review of world economics (weltwirtschaftliches archiv), 118(2):335–337, june 1982. 10. b. herrendorf, r. rogerson, and a. valentinyi. growth and structural transformation. in handbook of economic growth, volume 2, pages 855–941. elsevier, 2014. 11. g. hufbauer. the impact of national characteristics & technology on the commodity composition of trade in manufactured goods. in the technology factor in international copyright © 2020 assa. adv syst sci appl (2020) 130 i.g. pospelov, s.a. radionov trade, nber chapters, pages 143–232. national bureau of economic research, inc, august 1970. 12. p. kongsamut, s. rebelo, and d. xie. beyond balanced growth. review of economic studies, 68(4):869–82, october 2001. 13. k. matsuyama. agricultural productivity, comparative advantage, and economic growth. journal of economic theory, 58(2):317–334, december 1992. 14. m. j. melitz. the impact of trade on intra-industry reallocations and aggregate industry productivity. econometrica, 71(6):1695–1725, 2003. 15. l. r. ngai and c. a. pissarides. structural change in a multisector model of growth. american economic review, 97(1):429–443, 2007. 16. e. zhelobodko, s. kokovin, m. parenti, and j.-f. thisse. monopolistic competition: beyond the constant elasticity of substitution. econometrica, 80(6):2765–2784, 2012. copyright © 2020 assa. adv syst sci appl (2020) introduction model model setup market equilibrium and social welfare effects of technological progress and population growth conclusion microsoft word 977 article text, copyedited.doc adv syst sci appl 2020; 03; 166-177 published online at https://ijassa.ipu.ru. a differential stackelberg game theoretic model of the promotion of innovations in universities mukharbek malsagov1*, gennady ougolnitsky2, anatoly usov3 1) ingush state university, nazran, russia e-mail: mmm1956@bk.ru 2) southern federal university, rostov-on-don, russia e-mail: gaugolnickiy@sfedu.ru 3) southern federal university, rostov-on-don, russia e-mail: abusov@sfedu.ru abstract: a two-level dynamical game theoretic model "federal state universities" in openloop strategies is built and investigated. a refinement of the electronic learning courses, and their differentiation using modern methods of information technologies made by competing a la cournot universities (agents) are treated as innovative investments. the algorithms of building nash (in the non-coalitional game of agents) and stackelberg (in the hierarchical game of the principal with the grand coalition of agents) equilibria are proposed and implemented. for the solution of the respective dynamic control problems the pontryagin maximum principle and simulation modeling are used. a comprehensive analysis of the received results is given. keywords: nash equilibrium, stackelberg equilibrium, stackelberg differential games, innovations, method of qualitatively representative scenarios in simulation modeling 1. introduction a review of applications of the game theoretic models to the analysis of innovations is presented in [5]. the authors divide three levels of these applications: (1) intra-organizational level within a firm, where players are innovators, project managers, and resource administrators; (2) inter-organizational level where the players are competitive firms; (3) meta-organizational level, where players are a social planner (innovation policy maker, government, a social or government institution, e.g., a research foundation) and an aggregate innovative entrepreneur. a traditional approach to the building and investigation of the differential game models is exposed in [1,4]. an original method of solution of the stackelberg differential games based on building of a mutually benefit program of actions and punishment in the case of deviations, is described in [9]. in our papers [13,16] the approach of [9] is extended for the case of several followers with consideration of the requirements of sustainable development of the controlled dynamic system [12]. ougolnitsky and usov [14] have proposed a method of qualitatively representative scenarios in simulation modeling. this method allows for a good enough qualitative forecast of the dynamics of a controlled system by means of few scenarios of the computer simulation. an interesting dynamic game theoretic model of the oligopolistic investment to the product differentiation is studied in [2, 3]. the paper [2] considers open-loop strategies, and the paper [3] analyzes closed-loop ones. cellini and lambertini [2, 3] have found that in the case of open-loop strategies the nash stable volume of investments to the product differentiation (innovations) increases with a number of firms, i.e. the level of market * corresponding author: mmm1956@bk.ru a differential stackelberg game theoretic model… 167 copyright ©2020 assa. adv. in systems science and appl. (2020) competition (arrow hypothesis). in contrary, in the case of closed-loop strategies the dependence is inverse (schumpeter hypothesis). the authors notice that innovation activity in the product differentiation can be considered as a contribution of private investors to the production of a public good. this fact embeds their paper to the context of public goods economics. interesting models of stackelberg oligopoly are proposed and investigated in [68]. this paper develops the papers [2, 3] with consideration of other mentioned sources in the following directions. first, its problem domain is a development of the electronic learning courses in universities, where a modification of the courses is treated as an innovative activity. there is no doubt in the actuality of this problem. on march 25, 2020 the it holding talenttech and on-line universities "netology" and edmarket have presented the research results about the on-line education market†. as the authors notice, a volume of the russian market of on-line education in the b2c segment was 38,5 billion rubles in the end of 2019. in the end of the year 2023, according to that forecast, its value will be equal to 60 billion rubles a year. the global market of on-line education up to the year 2023 tends to the amount $282,62 billion. an even higher estimate is given by the interfax academy. according to its report, the russian market of the on-line education after 2019 equals to 45–50 billion rubles, and in 2020 will be equal to the amount 55–60 billion rubles, the annual growth is 20–25%. the global on-line education market is estimated by the value $74 billion (about 4,8 trillion rubles) after the year 2019, so the potential for russia is great‡. it is evident that the growth of the distant forms of education due to covid-19 pandemics will essentially enforce the noticed trends. second, we propose a hierarchical setup of the problem with the federal state as a leader (principal), and competitive universities as followers (agents). a description of a university as active system [11] is given in [10]. the hierarchical impact of the principal to the agents may be administrative (compulsion) or economic one (impulsion). in the former case the principal impacts to the sets of feasible strategies of the agents, and in the latter case to their payoff functionals [12]. thus, the contribution of the paper has four aspects. first, we consider a combination of the aggregative non-cooperative oligopolistic game of the agents with the stackelberg game of the type "principal-agents". second, the parameter of the demand function varies in time, and the character of this variation depends on the agents' actions (strategies) in the form of a differential equation. third, the agents choose both outputs and investments, i.e. an agent's strategy includes a parameter of its cost function called the constant cost. fourth, the following approach is used. from the point of view of the agents, their interaction is modeled as a game in normal form where the nash equilibrium is built. from the point of view of the principal it is supposed that the agents form the grand coalition, and respectively a stackelberg two-person game "principal coalition of agents" arises. this permits to avoid a challenging question about what should be considered as a best response of several agents to the principal's strategy. the rest of the paper is organized as follows. in the section 1 the setup of a dynamic problem of hierarchical control is given. in the section 2 the nash equilibrium for two symmetrical agents with administrative and economic impact of the principal in open-loop strategies is built. algorithms of building the stackelberg equilibrium by means of the method of qualitatively representative scenarios in simulation modeling are presented. the section 4 describes a numerical simulation in the problem of building the stackelberg equilibrium. the section 5 contains a comparative analysis of the received results, and the conclusions are formulated. †http://neorusedu.ru/news/rossijskij-rynok-onlajn-obrazovaniya-ozhidaet-burnyj-rost ‡https://www.kommersant.ru/doc/4275439 168 m. malsagov, g. ougolnitsky, a. usov copyright ©2020 assa. adv. in systems science and appl. (2020) 2. the problem setup let us consider the following hierarchical modification of the model proposed by cellini and lambertini [2, 3]. several universities (agents) competing a la cournot develop electronic learning courses for sale. a refinement of the courses, and its differentiation using modern methods of learning and information technologies are treated as innovative investments. the differentiation of the courses may be considered as a public good, and the respective investments as a private production of the public good [3]. on the higher (relative to the agents) control level a principal (federal state or its authorized bodies) is situated. the principal tends to increase a public good (with possible additional consideration of its own interests) by means of administrative or economic control methods. in the case of administrative impact (compulsion) the principal bounds from below the contributions of the agents to the innovative development (production of the public good) that incurs control cost. in the case of economic impact (impulsion) the principal grants the agents on the base of an available budget. in both cases the principal and the agents use open-loop strategies. the game is played on a finite interval of time in the case of compulsion the model for n agents has the form: the principal's payoff functional (2.1) the principal's control constraints ; (2.2) the agents' payoff functionals (2.3) the agents' control constraints ; ; (2.4) the equation of system dynamics [9, 10] (2.5) the current payoff function of the i-th agent (2.6) in the principal's payoff function the profit πi reflects an approximate estimate of the positive externality from the universities' activity, namely, the gnp growth due to increasing of the educational level of the society; the inverse demand function [2,3]; (2.7) the payoffs of the principal and the agents in the moment of time t respectively the total payoff functionals of the principal and the agents; a symmetrical degree of substitutability between any pair of courses. if then the courses are completely homogeneous. if then the courses are unique and each agent is a monopolist [2,3]; an output level of the i-th agent at constant returns to scale, then total operative cost per period are ; individual investment of ].,0[ t max)()())()(( 0 0 1 1 0 ®+ú û ù ê ë é ÷ ø ö ç è æ--= ò å å = = tgdttyztstej t n i n i iii t pr max)(0 ktyi ££ ;max)())()(( 0 ò ®++= t iii t i tgdttstej pr max)()( ktkty ii ££ max)(0 qtqi ££ );( )(1 )( td tk tk dt dd + -= ;)0();()( 1 bdtktk n i i ==å = )()(])([)( tktqctpt iiiii --=p å ¹ --= ij jii tqtdtbqatp )()()()( ;,...,2,1),(])()()([)( nitqctqtdtbqatg ii ij jii =---= å ¹ .)(])()()([)()( 11 0 å åå = ¹= ---== n i ii ij ji n i i tqctqtdtbqatgtg ijj ,0 ],0[)( btd î btd =)( 0)( =td )(tqi ),0(),()( iiiii actqctc î= )(tki a differential stackelberg game theoretic model… 169 copyright ©2020 assa. adv. in systems science and appl. (2020) the i-th agent to the innovative development, the overall industry expenditure, an upper bound for any ; an upper bound for any ; functions are the principal's grants to the i-th agent (in the case of compulsion they are given); t – the length of the game. then, are the principal's lower bounds established for a discount factor; demand parameters; a convex increasing administrative control cost function, the function is supposed to be linear the equation of dynamics (5) is treated as a production function with the output created by the input . this technology can be shown to exhibit decreasing returns to scale w.r.t. . thus is non-increasing function of time which tends to zero when [3]. in the case of impulsion the model has the form: the principal's payoff functional (2.8) the principal's control constraints ; (2.9) the agents' payoff functionals (2.10) the agents' control constraints ; . (2.11) here are the principal's control to be determined; is a total volume of the principal's grants (with consideration of possible savings). all input functions of the model are supposed to be continuous, and the controls of the agents and the principal belong to the class of piecewise continuous functions. we investigate the model (2.1)–(2.7) for compulsion and (2.5)-(2.11) for impulsion from the point of view of different control agents. from the point of view of the agents there is a noncooperative n person game which solution is supposed to be a nash equilibrium. from the point of view of the principal there is a stackelberg game. let us assume in this case that the agents cooperate (create the grand coalition) and have the summary payoff functional in the form (2.12) other model relations do not change. thus, we have a stackelberg game between the principal and the grand coalition of agents. this game has the following information structure: 1. the principal chooses its open-loop strategies for impulsion or for compulsion, . 2. given these strategies for impulsion or for compulsion, the agents within the grand coalition choose their open-loop strategies , using (2.12). as an integrand function in (2.12) depends continuously on its arguments, and the domains of feasible controls of the agents (2.4) or (2.11) are non-empty closed sets, the problem of determination of the agents' best response to any principal's strategy is resolvable. )(tk constmax =k )(tki constmax =q )(tqi )(tsi )(tyi );(tki )1,0(îr 0,0 >> ba )(×z ;0)0( =z )(×z .cons;)( teexxz == )(/ tdd!)(tk )(tk )(td ¥®)(tk max)()]()([ 0 0 1 0 ®+-= ò å = tgdttstej t n i ii t pr å = £³ n j ji ststs 1 )(;0)( ò ®++= t iii t i tgdttstej 0 max)())()((pr max)(0 ktki ££ max)(0 qtqi ££ )(tsi s .max)())()(( 1 01 ®÷÷ ø ö çç è æ ++== å òå = = n i i t ii t n i ia tgdttstejj pr )(tss ii = )(tyy ii = nitt ,...,2,1],,0[ =î )(tss ii = )(tyy ii = ),(),( tqqtkk iiii == nitt ,...,2,1],,0[ =î 170 m. malsagov, g. ougolnitsky, a. usov copyright ©2020 assa. adv. in systems science and appl. (2020) 3. the principal maximizes its payoff functional (2.1) or (2.8) for the worst best response of the coalition of agents to its strategy or . 4. the received set of strategies for impulsion or for compulsion is the stackelberg equilibrium. 3. building the nash equilibrium let us first investigate the model from the point of view of the agents when they use openloop strategies. the principal's strategies are supposed to be given. this interpretation corresponds to the case of an indifferent principal without its own objectives. then we receive a differential n-person game (2.3)-(2.7) in the case of compulsion and (2.5) – (2.7), (2.10), (2.11) in the case of impulsion. its solution is assumed to be a nash equilibrium. to build it we use the pontryagin maximum principle [15]. the hamilton function of the -th agent both for compulsion and impulsion has the form: where is a conjugate variable (as function of time). from the necessary condition of extremum in the case of symmetrical agents ( ) we receive the system of equation for determination of their control variables (3.1) thus . (3.2) besides, we have the system of differential equations (3.3) from (3.1) we receive therefore, the following proposition is proved. proposition (3.1): the formulas (3.2), (3.3) determine the point of maximum of the hamilton function for some value if the system (3.2), (3.3) has a solution and the values (3.2) belong to the domains of feasible controls (2.4) for compulsion or (2.11) for impulsion. if these conditions are not satisfied for some t then the hamilton function attains its maximum for this at one of the bounds of the segments (2.4) or (2.11). 0j ),...,( 1 nssr ),...,( 1 nyyr ),...,,,...,,,...,( ** 1 ** 1 ** 1 nnn qqkkss ),...,,,...,,,...,( ** 1 ** 1 ** 1 nnn qqkkyy i ,))(1 ,)( )()()()()()),(),(),(),(( 1 1 1 å å å = = ¹ + ++÷ ÷ ÷ ø ö ç ç ç è æ ---= n j j n i i iiiii n j ij jiiiii tk tkd ttsktqctqdtbqattdttqtkh ll )(til ï ï î ïï í ì = = ¶ ¶ = ¶ ¶ ni q h k h i i i i ,...,2,1, 0 0 ;cci º niiggqqkksshh iiiiii ,...,2,1;;;,;;; ===ººººº ll ;0 )1( 1 2 =+ --= ¶ ¶ nk d k h l .0)()1()(2 =----= ¶ ¶ ctqndtbqa q h ))1(2 )(;1)( -+ = -+= ndb сatq n dtk l ;11 2 2 ÷ ø ö ç è æ -+÷ ø ö ç è æ + -== ¶ ¶ ddb ca dt d d h l ll ; )1)((2 ))(1( )( )( -+ -= ¶ ¶ = ntdb aсn td gtl .)0(;11 bd ddt dd = +-= l .0;0)1(2; )1( 2 2 2 2 32 2 = ¶¶ ¶ <---= ¶ ¶ + = ¶ ¶ qk hndb q h nk dn k h l t t a differential stackelberg game theoretic model… 171 copyright ©2020 assa. adv. in systems science and appl. (2020) an analytical investigation of the system of equations (3.3) is impossible due to the form of the equation of dynamics (2.5) and presence of the third summand in (2.7) that is principal for the considered problem setup. thus the system (3.3) was analyzed numerically by the method of shooting. the system of equations (3.3) has a solution not for all values of the model input parameters. the picard theorem of existence and uniqueness of solution of the system of differential equations is not satisfied, the right hand sides are defined only for the negative values of λ(t)d(t). besides, in the neighborhood of zero the lipschitz conditions on λ and d are not satisfied. therefore, the range of input model parameters for which the system (3.3) has the unique solution was determined numerically. a numerical identification of the model has a testing character and provides a reasonable relation of the model parameters and variables that allows for acceptable qualitative conclusions about the comparative analysis of the results of numerical modeling. thus, the professors of southern federal university develop annually about 100 new learning courses, therefore the maximal possible value is taken an average cost of the development is equal to 50-80 thousand rubles, so (thousand rubles per year). the value of parameter b varied in the range (thousand rubles per year). we divided the cases of small investment of the principal ( ), middle investment ( ), and considerable investment ( ), as well as different cases of the agents' control by the principal: namely, the value of z varied from 2 thousand rubles per year (soft control) till 10 thousand rubles per year (hard control). the discount factor was estimated as according to the annual inflation rate. the period of modeling was equal to 3 years, i.e. days for two universities ( ). other parameters were evaluated by the experts (table 3.1). table 3.1. test values of the model parameters parameter value 100 50 900 5 3000 0.04 600 2 dimension thousand rubles per year thousand rubles per year thousand rubles per year thousand rubles per year thousand rubles per year for determination of the range of model parameters in which the system (3.3) has the only solution we have implemented about 80 numerical simulations with variation of the values: (thousand rubles per year); n from 2 till 100. as a result it was established that for any input data there is a moment of time exists a solution of (3.3): . the moment belongs to the range from 950 till 1000 days and depends not essentially on the parameters . its value slightly increases with n or b, decreases with a, and does not depend on c. some of the results are presented in table 3.2. table 3.2. dependence of the moment of time on model parameters 2 110 100 10 997 2 110 100 1 999 2 110 100 25 989 2 110 100 250 980 2 200 100 10 999 .100max =q ,3000max =k ,900=а 50=c 151-=b 10/maxkns = 5/maxkns = maxkns = 04.0=r 1095=t 2=n maxq cci º a b maxk r ssi º z ]200,10[îc ]250,1[; îb ]1000,100[; îa :01 >t ),[0)(;0)( 1 tttanyfortdt î³|| )( 0 )( 0 ji jj qrsuv l ï)(),( qrsuv j î)(),( d£|| )( 0 )( 0 jl jj )( 000 ,, )()( ljjj ji ),...,,,...,( )( 1 )( 1 )()( 10 )( 0 sss n ss uuvvjj = 0;,, >d= ljis d ),( ),(),( max 000 0000 ),(),( ),( 00 vuj vujvuj vuvu qrsvu ¹ î }, 2 ,0{)( maxmax n k n kyty ii îº 0=iy n kyi 2 max= n kyi max= s },,0{)( )2()1( iiii sssts îº åå == == n i i n i i ssss 1 )2( 1 )1( ; 2 0=is )1( ii ss = )2( ii ss = }, 2 ,0{ max max qqqi î 174 m. malsagov, g. ougolnitsky, a. usov copyright ©2020 assa. adv. in systems science and appl. (2020) . for simplicity it is assumed that the strategies do not vary with time. then we receive for small values the computer simulations are quite implementable. with consideration of the proposed strategies from qrs set the state variable is found analytically as 5. numerical calculation of the stackelberg equilibrium the computer simulations were implemented on personal computer with microprocessor a10 intel pentium g4620 with operative memory 4 gb using an object oriented programming language c# according to the proposed algorithm. an average time of one simulation for the construction a qrs set was less than three seconds. the received results were analyzed by the following criteria: (a) summary discounted payoff of the principal calculated by formulas (2.1) or (2.8); (b) index of system compatibility [12]: , where is the principal's payoff in the stackelberg equilibrium. this index demonstrates how necessary is principal's presence in a control system. the closer is the value of to one, the more the system is compatible, and the less is a need in the hierarchical control by the principal. for an initial qrs set the conditions (a) and (b) from the definition 1 are checked. the value of is selected so that a difference between the principal's payoffs for two any strategies from the initial qrs set does not exceed 10%. then given the model parameters the second condition from the definition 1 is checked. if necessary, the initial qrs set is extended or narrowed. in the case of extension the initial qrs set is added by new strategies. they correspond to the values situated between the previous ones (the strategies are supposed to be constant in time). then the computer simulations are implemented. we have realized about 150 numerical calculations for three agents that form the grand coalition. the model parameters varied in the following range: for compulsion; for impulsion. in all cases the index of system compatibility is equal to one, and the system is completely compatible. besides, in the case of impulsion on a short period of modeling (up to 4 years) both for the principal and for the agents it is not profitable to invest to the development of innovative technologies (learning courses). below the results of computer simulations for several sets of input data are presented. let the grand coalition includes three agents, and days (4 years). example 5.1: the values of model parameters are (thousand rubles per year). we varied the values the results for impulsion in the case are given in table 5.1. in the majority of calculations . table 5.1. results of numerical simulations (example 2) for impulsion }, 2 ,0{ maxmax n k n kki î .27nqrsm == 61 corresponds to the indicator reaching its threshold value, i.e. in this case, the indicator is in a stable zone. one of the main indicators of the operational stability of the industrial sector is considered to be the market component, namely the growth of production output. to this end, a hypothesis was put forward that there is a direct relationship between the volume of production in the industry and the level of development of its it infrastructure. that is, the created information and communication infrastructure in the industry becomes a determining factor in its transformation into the region’s industry of specialization. to test the hypothesis, the linear regression modeling method is applied, the significance of the obtained equations is established using the coefficient of determination (r2). the equation of the linear regression has the form: (3) where y is the average value of the effective indicator; x is the influencing factor; a and b is regression coefficients. the coefficients of the regression equation can be interpreted in this case as follows. if a> 0, then with increasing the coefficient x, fiscal security increases, and if a <0, then with increasing the coefficient x it decreases. to check the adequacy of the equation, the coefficient of determination is calculated: 2 2 2 )( 1 ( ) ii i a by x r yy − − = −  − (4) where yi is the level of socio-economic sustainability in the i-th year; xi is factor assessment in the i-th year; is the average level of socio-economic sustainability over several years. to verify the adequacy, the fisher test was used: 2 2 ( 2) (1 ) r f n r = − − (5) where r2 is the coefficient of determination; n is the sample size. if the value of f exceeds the critical value in absolute value, then the equation is adequate and can describe existing patterns. the obtained regression models are used to predict possible prospects for the development of the industrial sector depending on the expected values of sectoral technological development. in this forecast, the point type of the forecast was conventionally applied to visually indicate the trend of the digital transformation of industry. 4. results as part of the study, development indicators for the industrial sector of russia, including enterprises of processing and construction industries (appendix a), were selected. the indicators are defined from the perspective of the industry’s characteristics for five development potentials – market, production, innovation, financial, and hr. from appendix a, it is seen that for many indicators there is a negative trend of change. to calculate the economic sustainability of industrial enterprises, it is necessary to bring all indicators to a single normalized value. the calculations on normalizing the indicators for assessing the economic sustainability of the industrial sector make it possible to form curves transformation of industries 41 copyright ©2020 assa. adv. in systems science and appl. (2020) of indicators of changes in the economic stability of the industrial sector of russia for the five considered potentials as shown in fig. 2. 0,250 0,300 0,350 0,400 0,450 0,500 0,550 0,600 0,650 0,700 2010 2011 2012 2013 2014 2015 2016 2017 2018 market potential production potential innovation potential financial potential hr potential economic security (integral indicator) fig. 2. the dynamics of indicators of economic sustainability of the industrial sector of russia for 2010-2018 fig. 2 shows that the sustainability of the industrial sector is most determined by the market potential of economic security. it is important to note that the production potential of the industry is more stable and less affected by the crisis of 2014, in contrast to indicators in the context of finance and innovation (innovation and financial potential). as is known, the introduction of innovations and modern technologies is one of the main components of economic growth [6,20,26]. one of the challenges of technological development at the moment is digitalization [2,4,5]. to assess the level of digitalization in the industrial sector, the authors consider the dynamics of a number of indicators characterizing this process. appendix b provides information on the use of information and communication technologies. in 2017, the share of organizations using personal computers in construction is 88.9%, and in processing – 95.5%. currently, digitalization processes affect the mineral wealth sector of the economy to the greatest extent and the construction sector to the lowest. it is important to note that in terms of the use of personal computers, there is a negative trend in all considered sectors. in this study, the authors will identify the trend component that describes the impact of the digitalization level of industries on the development of the industrial sector. for this purpose, paired linear regression equations are constructed that describe the time series of the main growing trend under consideration (figs. 3-6). as this level, the authors determined the average value of the indicator “the use of information and communication technologies in organizations” as a percentage of the total number of examined organizations of the corresponding type of activity, which is an exogenous (influencing) variable. the endogenous (dependent) variable (y) is the value of the volume of own-produced shipped goods, own works and services provided. 42 o.p. smirnova, l.m. averina, a.o. ponomareva copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 3. the dynamics of the digitalization trend for the industrial sector as a whole, 2011-2018 fig. 4. the dynamics of the digitalization trend for the mineral wealth industry, 2011-2018 fig. 5. the dynamics of the digitalization trend for the processing industry, 2011-2018 fig. 6. the dynamics of the digitalization trend for the construction industry, 2011-2018 analyzing the obtained equations, one can conclude that the relationship between the level of production and the level of digitalization of the industry is close. the highest value of the coefficient of determination was found in the processing industry of the economy, a weaker value is observed in the construction industry. this conclusion can also be drawn by examining the values of the coefficients of the dependent variable: β = 1,437.7 in processing, β = 522.2 in mineral wealth mining, to the lowest degree the level of digitalization affects construction (β = 166.8). the variability of the development of the industrial sector of russia from the level of digitalization is shown in table 1. table 1. variability of the digitalization status of the industrial sector of russia indicator regression equation determination coefficient fisher test predictive critical total y=2,731x – 130,987 0.9739 2.1629 2.0048 mineral wealth mining y=522.09x – 26,503 0.7043 0.8758 2.7764 processing y=1,437.7x – 78,052 0.9289 1.6503 4.3026 construction y=166.81x – 6,480 0.5519 0.6211 3.1824 the significance of the obtained equations was checked using the fisher test. based on the calculated critical value of the fisher test, the authors accept an alternative hypothesis and conclude that there are statistically significant differences in the frequency of outcome depending on the impact of the factor. according to the established pattern, there is a rather strong dependence of the volumes of own-made products on equipping the production environment of it infrastructure. transformation of industries 43 copyright ©2020 assa. adv. in systems science and appl. (2020) 5. discussion the advent of new technologies leads to the modernization of existing industries and the emergence of new ones. traditional industries of regional specialization of the industrial sector are supplemented by new ones such as ict, industries using biotechnology, photonics, robotics, artificial intelligence, etc. the authors present the results of a rating of the costs of russian regions for ict. in 2013, the cost of ict in russia as a whole amounted to 69.5 billion rubles. by 2019, this figure increased by 2.3 times. note that in industrialized regions the growth rate of this indicator is even higher, for example, in the sverdlovsk region over the same period, ict expenses increased by 2.5 times (from 1.06 to 2.73 billion rubles); in the chelyabinsk region, ict expenses increased by 4.5 times (from 0.26 to 1.14 billion rubles). 40000 60000 80000 100000 120000 140000 160000 180000 2013 2014 2015 2016 2017 2018 2019 million rubles fig. 7. dynamics of costs for information and communication technologies in russia these trends suggest support for the continued use of ict. the presence of an information barrier in the growth of companies has been repeatedly noted in studies [39]. the dependencies obtained earlier allow comparing the growth of the real sector of the economy with the growth rate of using ict in these sectors. here is the forecast for the development of the russian industry as a whole, as well as in mineral wealth mining, processing and construction (figs. 8-11). this forecast is based on the obtained relationships of the influence of the digitalization level on the development of the industrial sector. the digitalization trend, as mentioned above, in this study is described by the average of indicators such as the share of enterprises in the relevant industry using personal computers, servers, local area networks, global information networks, and the share of organizations that have a web page. the forecast for the development of the manufacturing industry as a whole shows that, while maintaining the average growth rate of the digitalization in the industry (2.2% per year), the predicted volume of industrial goods and services by 2023 may amount to more than 76.9 billion rubles, and the average growth of this indicator will be more than 8% a year. 44 o.p. smirnova, l.m. averina, a.o. ponomareva copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 8. forecast of the impact of the digitalization trend on the industrial sector as a whole for the period of 2019-2023 fig. 9. forecast of the impact of the digitalization trend for the mineral wealth industry for the period of 2019-2023 fig. 10. forecast of the impact of digitalization trend for the processing industry for the period of 2019-2023 fig. 11. forecast of the impact of the digitalization trend for the construction industry for the period of 2019-2023 the digitalization trend of the mineral wealth industry has somewhat slower dynamics of changes and a lower connection with the resulting volume indicator. while maintaining the existing average growth rate (about 1.5% per year), by 2023 the volume of output in the mineral wealth mining industry will amount to more than 14.24 billion rubles per year, and the average growth rate in the industry is about 7% per year. the highest degree of digitalization is noted in the processing industry, and here the predicted growth values will be about 2.2% per year. this will allow the industry to reach a production level of over 52 billion rubles per year by 2023, provided that 90% of organizations in this industry will use digital technology (under this condition, the industry will show average growth rates at almost 9%). the construction industry has shown the least degree of digitalization. the predicted share of organizations using digital technologies in the industry by 2023 will be less than 70%. the output volume will approach 5 billion rubles, and the average growth rate in the construction industry will be about 2.8% per year. the obtained data on the dynamics of industrial development make it possible to formulate a forecast for the structure of the real sector of the national economy (fig. 12). transformation of industries 45 copyright ©2020 assa. adv. in systems science and appl. (2020) 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 mineral wealth mining processing construction forecast fig. 12. real and forecast structure of the real sector of the economy fig. 12 shows a gradual decrease in the share of the construction industry, as well as of mineral wealth mining, which is caused by lower development rates of these industries compared to the processing industry. that is, the influence of new technologies and the digital transformation of the manufacturing industry contributes to positive structural changes in the economy. an increasing degree of influence of new technologies, including ict, will inevitably contribute to the transformation of the manufacturing industry. industries of regional specialization will also be transformed. the study showed that existing technological trends must be taken into account, as they have a significant impact on the development of the economy. the results of this study suggest that the feature of new industries is rapid growth, accumulation and combination of various types of resources, of great importance for their development is scientific and industrial cooperation, interregional interaction and scientific collaboration [3]. one of the popular tools for the development of new technologies is such a kind of strategic planning and management of the regional economy as smart specialization. this concept makes it possible to single out individual specialized activities that do not correspond to the profile type of the region, but represent a unique highly specialized niche in the markets of high-tech products, services and technologies, i.e. creation of effective promising regional specializations [28]. 6. conclusion this paper identifies the factors that influence the transformation of regional specializations, taking into account the global challenges of technological development. the role of digitalization in increasing the competitiveness of industries with promising economic specialization of the region was identified, which is due, inter alia, to cooperative interindustry interaction and the level of development of innovative infrastructure. the results obtained confirm the hypothesis that the created information and communication infrastructure in the industry is becoming a determining factor in its transformation into the region’s industry of specialization. that is, digitalization in the aspect of regional development is a mechanism that allows certain sectors to achieve competitiveness. therefore, there is an objective need to take into account the level of digitalization in the methods for identifying regional specialization. these methods should account for not only traditional approaches based on the inter-regional division of labor and the rational 46 o.p. smirnova, l.m. averina, a.o. ponomareva copyright ©2020 assa. adv. in systems science and appl. (2020) distribution of productive forces, but also the priority of high-tech industries and their digitalization level. thus, the modern methodological approach to identifying the specialization of the region should be guided to a greater extent by the qualitative indicators of development, their technological transformation of the economy and the current market situation. acknowledgements the paper was prepared in accordance with the research plan for the laboratory for spatial development of the territory at the institute of economics of the ural branch of the russian academy of sciences for 2019-2021. references 1. aglietta, m., & boyer, r. (1982). poles de competitivite, strategie industrielle et politique macro-economique. paris. couverture orange cepremap, no. 8223. 2. akberdina, v., kalinina, a., & vlasov, a. (2018). transformation stages of the russian industrial complex in the context of economy digitization. problems and perspectives in management, 16(4), 201-211. 3. aldieri, l., kotsemir, m.n., & vinci, c.p. (2017). the impact of research collaboration on academic performance: an empirical analysis for russian universities. journal of eurasian social dialogue, 2(1), 17-33. 4. batkovsky, a.m., kalachykhin, p.a., telnov, yu.f., & fomina, a.v. (2019). assessment of the level of requirements to key competences of enterprises in the conditions of the digital economy. radio industry (russia), 29(3), 91-99. 5. blatova, t.a., makarov, v.v., & shuval-sergeeva, n.s. (2019). quantitative and qualitative aspects of measuring the digital economy. radio industry (russia), 29(4), 63-72. 6. dellis, k., karkalakos, s., & kottaridi, c. (2016). entrepreneurship targeting policies, technological growth, and unemployment. journal of eurasian economic dialogue, 1(6), 19-39. 7. edmondson, g., mccollam, s., & kelly, e. (2014). 5 steps to smarter specialization. science business publishing. 8. foray, d. (2013) smart specialisation and the new industrial policy agenda. paper presented at the 2013 erac mutual learning seminar, 20th march 2013. retrieved february 23, 2020, from https://era.gv.at/object/document/360/attach/industrial_policy_agenda.ppt 9. foray, d., david, p., & hall, b. (2009). smart specialization – the concept. knowledge economists policy brief, 9(85), 1-5. retrieved february 23, 2020, fromhttp://ec.europa.eu/invest-in-research/pdf/download_en/kfg_policy_brief_no9.pdf 10. frenken, k., yan oort, f., & verburg, t. (2007). related variety, unrelated variety and regional economic growth. regional studies, 41(5), 665-697. 11. gardiner, b., martin, r., & tyler, p. (2004). competitiveness, productivity and economic growth across the european regions. regional studies, 38, 1045-1067. 12. hausmann, r., & klinger, b. (2006). structural transformation and patterns of comparative advantage in the product space. cid working papers, no. 128. 13. heckscher, e. (2006). the effect of foreign trade on income distribution (milestones of economic thought. vol. 6). moscow: ekonomicheskaya shkola. (p. 720). 14. hidalgo, c., & hausmann, r. (2009). the building blocks of economic complexity. proceedings of the national academy of sciences, 106, 10570-10575. 15. karpenko, s.v., silina, t.a., & ordinskaya, m.e. (2011). a systematic approach to the study of the meso-level of economic relations in regional economic systems. https://era.gv.at/object/document/360/attach/industrial_policy_agenda.ppt transformation of industries 47 copyright ©2020 assa. adv. in systems science and appl. (2020) retrieved february 23, 2020, from https://www.sworld.com.ua/index.php/ru/efficienteconomy-today-c112/11951-c112-088 16. kazantsev, s.v. (2017). models for calculating the security indicators of a country and its regions. region: economics and sociology, 2(94), 32-51. 17. korovin, g.b., & averina, l.m. (2018). prospects for the development of the green economy on the principles of “smart specialization” in the old industrial region. in v.v. akberdina, & o.a. romanova (eds.), multisubject industrial policy: joint monograph (pp. 84-93). yekaterinburg: institute of economics, ural branch of the russian academy of sciences. 18. kroll, h. (2015) efforts to implement smart specialization in practice – leading unlike horses to the water. european planning studies, 10(23), 2079-2098. 19. kutsenko, e.s., abashkin, v.l., & islankina, e.a. (2019). focusing regional industrial policy through industry specialization. issues of economics, 5, 65-89. 20. lambert, t.e. (2018). monopoly capital and innovation: an exploratory assessment of r&d effectiveness. journal of eurasian economic dialogue, 3(6), 20-35. 21. larsen, p. (2011). cross-sectoral analysis of the impact of international industrial policy on key enabling technologies. brussels: european commission. retrieved june 10, 2020, from https://publications.europa.eu/en/publication-detail//publication/713f63c6-9d8a-4680-99f3-de63d489e79e. 22. makedon, v., drobyazko, s., shevtsova, h., maslosh, o., & kasatkina, m. (2019). providing security for the development of high-technology organizations. journal of security and sustainability issues, 8(4), 759-774. https://doi.org/10.9770/jssi.2019.8.4(18) 23. mccann, p., & ortega-argilés, r. (2015). smart specialization, regional growth and applications to european union cohesion policy. regional studies, 8(49), 1291-1302. 24. mccann, p., & ortega-argilés, r. (2016). the early experience of smart specialization implementation in eu cohesion policy. european planning studies, 8(24), 1407-1427. 25. mityakov, s.n. (2014). development of a system of indicators of economic security of russian regions. in economic security of russia: problems and prospects: proceedings of the ii international scientific and practical conference (pp. 70-78). nizhniy novgorod. 26. molina, j.a., velilla, j., & ortega, r. (2016). entrepreneurial activity in the oecd: pooled and cross-country evidence. journal of eurasian economic dialogue, 1(6), 118. 27. oecd. (2013). innovation-driven growth in regions: the role of smart specialization. paris: oecd publishing. 28. olaniyi, e.o., & reidolf, m. (2015). organisational innovation strategies in the context of smart specialization. journal of security and sustainability issues, 5(2), 213-227. 29. pilyasov, a.n. (2018). regional investment policy: how to overcome “dependence on the path?” region: economics and sociology, 4, 134-167. 30. porter, m. (1993). international competition. moscow: mezhdunarodnye otnosheniya. 31. ricardo, d. (1935). selected works. vol. 2: on the principles of political economy and taxation. moscow. (p. 350). 32. romanova, o.a., korovin, g.b., & kuzmin, e.a. (2017). analysis of the development prospects for the high-tech sector of the economy in the context of new industrialization. espacios, 38(59), 25. 33. smirnova, o.p., & averina, l.m. (2019). a study of the features of promising economic specialization of an industrial region. regional economy: theory and practice, 6(17), 1006-1019. 34. smith, a. (2007). an inquiry into the nature and causes of the wealth of nations. moscow: eksmo. (p. 960). https://www.sworld.com.ua/index.php/ru/efficient-economy-today-c112/11951-c112-088 https://www.sworld.com.ua/index.php/ru/efficient-economy-today-c112/11951-c112-088 48 o.p. smirnova, l.m. averina, a.o. ponomareva copyright ©2020 assa. adv. in systems science and appl. (2020) 35. stryabkova, e.a. & lyshchikova, yu.v. (2019). development of methodological approaches to identifying the priorities of “smart specialization” of territories. economics: yesterday, today, tomorrow, 12a(9), 73-82. doi: 10.34670/ar.2020.92.12.037. 36. toomsalu, l., tolmacheva, s., vlasov, a., & chernova, v. (2019). determinants of innovations in small and medium enterprises: a european and international experience. terra economicus, 17(2), 112-123. 37. veselovsky, m.y. pogodina, t.v., ilyukhina, r.v., sigunova, t.a., & kuzovleva, n.f. (2018). financial and economic mechanisms of promoting innovative activity in the context of the digital economy formation. entrepreneurship and sustainability issues, 5(3), 672-681. 38. zyablitskaya, t.s. (2012). formation and development of the competitiveness of a region with tourist specialization (phd thesis). altai state university. 39. litau, e. (2018). the information problem on the way to becoming a “gazelle.” in proceedings of the european conference on innovation and entrepreneurship, ecie (vol. 2018-september, pp. 394–401). appendix а table. indicators of the potential of the industrial complex of russia (for enterprises of processing and construction industries) indicator 2010 2012 2014 2016 2018 market potential dynamics of income, % 100.00 111.11 108.30 104.15 105.00 ratio of the production index and the index of changes in numbers (in % to the previous year) 107.6 104.2 103.5 101.3 102.9 index of changes in labor intensity and index of changes in the working time in % of the production index (in % to the previous year) 110.5 96.4 96.7 99.4 99.2 costs of production and sale of products per 1 ruble of manufactured products, kopecks 75.3 77.2 80 80.1 80 production potential share of investments aimed at reconstruction and modernization in the total volume of investments in fixed assets 18.8 19.5 17.4 16.3 15.5 index of physical volume of investments in fixed assets aimed at reconstruction and modernization, % 103.5 107.1 92.5 95.1 101.6 average age of machinery and equipment available at the end of the year 11.1 11.5 11.2 11.3 12.2 index of changes in capital productivity, % 100.9 104.2 88.7 101.2 97.0 innovation potential volume of innovative goods, works, services, billion rubles 1.244 2.873 3.580 4.364 4.516 index of physical volume of investments in fixed assets aimed at reconstruction and modernization, % 103.5 107.1 92.5 95.1 92 number of advanced manufacturing technologies used, units 236 350 439 548 524 number of new technologies (technical achievements) acquired by organizations, units 119 63 28 14 48 financial potential return on sales, % 10 8.6 7.3 7.6 12.3 return on assets % 6.7 6.1 2.5 5.9 6.4 equity to total assets ratio 0.52 0.48 0.4 0.42 0.48 current ratio 1.34 1.25 1.21 1.24 1.01 hr potential average annual number of employees in organizations, million people 16.901 17.166 17.111 16.941 17.06 labor productivity index 105.2 104.8 102.5 100.2 95.9 index of changes in capital-labor ratio, % 102.2 99.3 113.5 100.9 103.00 average monthly nominal accrued wages per employee, thousand rubles 9.078 24.512 29.511 34.592 40.722 source: rosstat. transformation of industries 49 copyright ©2020 assa. adv. in systems science and appl. (2020) appendix b table. digitalization indicators of the industrial sector of russia (at the beginning of the year) indicator 2012 2014 2016 2018 pcs total 94.1 94.0 92.3 92.1 mineral wealth mining 94.6 95.6 93.1 90.7 processing 97.3 97.2 97.1 95.5 construction 96.0 94.3 92.9 88.9 servers total 19.7 19.7 47.7 50.6 mineral wealth mining 30.0 30.4 69.9 69.1 processing 25.9 25.2 67.4 74.5 construction 20.7 19.4 61.2 58.0 local area networks total 71.3 73.4 63.5 61.1 mineral wealth mining 85.1 86.3 78.3 73.3 processing 84.2 85.2 76.6 76.2 construction 82.7 81.6 68.3 59.9 global information networks total 85.6 88.7 89.0 89.7 mineral wealth mining 91.8 92.9 91.8 89.0 processing 94.3 95.2 96.0 94.5 construction 92.5 92.3 91.4 87.1 web pages total 33.0 41.3 42.6 47.4 mineral wealth mining 30.0 36.8 37.2 39.7 processing 53.3 57.9 57.5 63.8 construction 34.3 38.7 40.1 38.7 source: rosstat. 1. introduction 2. literature review 3. materials and methods 4. results 5. discussion 6. conclusion acknowledgements appendix а appendix b paper title (use style: paper title) adv syst sci appl 2019; 04; 87-93 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/830 exchange rate of the ruble modeling: macroeconomic level anton kuzmin department of data analysis, decision-making and financial technologies the financial university under the government of the russian federation moscow, russian federation e-mail: a_kuzmin@rambler.ru received may 16, 2019; revised december 12, 2019; published december 31, 2019 abstract: as underlying determinants of the author’s model are used capital flows, import and export operations, elasticities of cross-country trade, terms of foreign trade, indexes of real gdp, indexes of domestic and export prices, solutions of microagents. the model was verified on data from the crisis period in russia 2013-2015. the simulation results are applied to the analysis of contemporary monetary policy of russia. keywords: exchange rate of the ruble, modelling, balance of payments, currency crisis. 1. introduction the monetary approach to exchange rate modeling is important: the models of frenkel [1], mussa [2, 3], hooper-morton [4] and others. model with sticky prices was created by dornbush [5, 6]. the two-country model with maximization of individual preferences was created by stockman [7]. representative of another class of dynamic stochastic general equilibrium (dsge) models is the monopolistic competition model obstfeld-rogoff [8] of exchange rate of two open economies. it should be noted that economic fundamentals takes an important place both in these models and in the model clark-macdonald [9, 10]. a number of researches was devoted to the dynamics of the ruble: strizhkova [11], mirkin [12] and others. however, to construct a mathematical fundamental determinants model of the russian ruble new developments are necessary. we have to base on the study of the fund movements of the russian balance of payments. we also need to engage in modeling the terms of foreign trade as one of the most important fundamental determinants. as a result, it's necessary to investigate the equilibrium medium-term dynamics of the ruble in the context of the currency crisis in 2013-15. in this paper the proposed dynamic model of the equilibrium exchange rate of russian ruble is a development of previous author’s research [13, 14]. it takes into account in the medium-term period capital flows, import and export operations, elasticity of cross-country trade, terms of foreign trade, indexes of real gdp, indexes of export and domestic prices, solutions of micro agents. the constructed exchange rate will balance the balance of payments. as a consequence, it will balance the offer of national and foreign currencies on the cross-country monetary market. the main objective of the modeling was to identify the dynamics of the ruble exchange rate from the system of major macroeconomic factors. this system must necessarily include the factors of the monetary sector, as well as the factors of the real sector. modeling will be mailto:a_kuzmin@rambler.ru 88 a. kuzmin copyright ©2019 assa adv. in systems science and appl. (2019) based on the principles of the international flows equilibrium exchange rate (ifeer), developed in the works of the author. but in this article we will present a new model. 2. a dynamic model of the exchange rate of the ruble the exchange rate policy of the bank of russia is the regime of independent floating according to the classification of the international monetary fund. russia is also characterized by special public attention to the dynamics of the nominal exchange rate of the national currency. as a result, the importance of the bank of russia's exchange rate policy is typical in modern conditions. initially, we consider all real market transactions at nominal exchange rates , (1, )ie i n on the domestic currency market that occurred in a certain period of time. then iii rde ,, in i-th transaction are: the nominal exchange rate, the amount in the determined foreign currency and the amount in the national currency respectively. these variables are linked by ratios: iii rde  , and therefore, iii dre / . in the russian foreign exchange market, a significant part of transactions takes place in us dollars. transactions in other currencies, such as the euro, the british pound are directly linked to these prevailing rates through the system of cross-rates in the international and local market. thus, we further consider the us dollar as a foreign currency and the direct quotes of us dollar to russian ruble. it should be noted that the contribution of each market transaction is different. it depends on the volume of the transaction. therefore, we define the synthetic value of the exchange rate for the certain period as the sum of exchange rates, weighted by amounts in determined foreign currencies: 1 1 n i in i j j d e e d     , (2.1) using the definition of the exchange rate in i-th transaction, we get from (2.1):          n j j n i in i i i n j j i d r d r d d e 1 1 1 1 so, on a conceptual level, we can disaggregate the country's balance of payments flows:          cbkca cbkca ddd rrr e (2.2) in this formula (2.2) we consider flows in the respective currencies that came to the domestic currency market: as the supply of the foreign and national currencies in accordance with current account operations amount of funds with index ca, as the supply of the foreign and national currencies in accordance with capital operations amount of funds with index к, as a result of currency interventions of the bank of russia amount of funds with index св. this formula takes into account such important components as export revenue, demand for imports, demand for foreign assets and so on. exchange rate of the ruble modeling 89 copyright ©2019 assa adv. in systems science and appl. (2019) conceptually, formula (2.2) can be applied not necessarily to analyze the dynamics of the russian ruble, but in other cases as well. in the work of the author [6], its adapted version was used to model the dynamics of equilibrium exchange rates under the equality of economies. however, in this paper we will conduct a modeling of the ruble exchange rate in the conditions “open economy rest of the world”. further two-period model will be considered in periods t, t-1. it should be noted that russian exports affect the economic situation in the country. of course, the dynamics of actual export prices will directly determine the state of the domestic foreign exchange market. amount of foreign currency in the market at period t will be determined by the volume of exports decisions of exporters at period t-1. and the most important factor in increasing the physical volume of exports is the real terms of international trade. we define these conditions in this situation as the division of exports and domestic prices adjusted to the nominal exchange rate: 1 * 1 11     t t tt r p p ee , where pt the internal price index (e.g., cpi), * 1tp the actual export price index. here the terms of international trade can be structurally correlated with the real exchange rate, used in the author's work [13]. thus, the amount in foreign currency at period t that came to the domestic foreign exchange market as part ek of export earnings in actual export prices  tp is equal to zr t x x t x tet ca eqqkpd )()( 1 1 1 1 1 1        (2.3) where te qconstk , the index of total real output (e.g., real gdp). let’s give some explanations to the formula (2.3). index z values response to the state of the terms of international trade at period t-1. the part 1 11 1 1( ) x x x e t tk q q    reflects that the physical export is a part of averaged total real output not only in current period, but also in the previous time. at the same time averaging method in 1 11 1 1( ) x x x e t tk q q    shall not have a significant influence on the final result from economic positions. this is due to the insignificance in the real output changes, compared to possible fluctuations of price indices in the medium term. rate  shows a "slightly higher" non-negative growth of national exports in relation to imports as a function of real output. the main reason is the restriction of domestic demand and thus the need to realize the growing domestic output due to the growth of exports. the parameter x=(z-y) associates parameters y and z in the dependencies (2.3) and (2.4). next, make an assumption about the import dependencies: the consumption of imports at time t by residents with proportion ik of their income )( 1 1 1 1    x x t x tt qqp is represented both current income and income in the previous period as follows yr t x x t x tti ca eqqpkr ))(( 1 1 1 1 1     (2.4) the parameter у also values response to the state of the real terms of international trade. 90 a. kuzmin copyright ©2019 assa adv. in systems science and appl. (2019) of course, as it can be seen from formula (2.2), capital flows directly determine the dynamics of the exchange rate. at the same time, a significant number of exchange rate models ignore this problem, while remaining at the level of the trade balance. however, within the framework of the proposed ifeer modeling the capital flows can be involved in the process. we accept the hypothesis for the functional dependence of capital outflows: this is the demand for international assets of russian residents for savings purposes and it is part of total income in domestic prices: 1 1 1 1 1 ( ) ( )( ) k cb x r yx x t t t tk r r p k q q e          , (2.5) where constk k  . it will also assumes that in (2.5) and further in (2.6) transactions of the bank of russia are considered. for the functional dependence of capital inflows, we accept that it reflects the inflows of portfolio and direct investments. according to this international speculators and investors want to buy a part of the real national product in international prices. another important factor is the terms of international trade because the changes in them will certainly affect the investment climate: 1 11 1 1 1 ( ) ( ) ( ) k cb x r zx x t t t tk d d p k q q e            (2.6) where constk k  , 0 . dependence (2.6) suggests that the impact of national product growth is increased more than proportionally due to psychological factors and the implemented policy of import substitution, which is carried out by the government of the russian federation in recent years. in this dependence it is expressed by the parameter  in part     1 1 1 1 1 )( x x t x tk qqk . we will directly substitute dependencies (2.3) (2.6) into the formula (2.2):                                        1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 )()( )()( ))(( ))(( zr t x x t x ttk zr t x x t x tte yr t x x t x ttk yr t x x t x ttit eqqpk eqqpk eqqpk eqqpke   exchange rate of the ruble modeling 91 copyright ©2019 assa adv. in systems science and appl. (2019) 1 1 1 1 1 1 11 1 1 1 1 1 1 1 1 ( )( ) ( ) ( ) ( ) ( ( ) ) . x r yx x t t t t i k x r zx x t t t t x x x e t t k p q q e k k p q q e k q q k                                   (2.7) based on economic logic, we can assume that ( ) 0   and the part )( 1 1 1   õ t õ õ t qq is stable, comparing to other model variables in the medium term. this makes it possible to consider the component as a constant in (2.7): constk kqqk kk x k x x t x te ki           1/ 1 1 1 1 )( ))(( )(  . thus as you can see, this component k' is the more stable, the smaller the contribution of the current balance flows compared to the capital flows. formula (2.7): / 1 1 1 1 1 1 / 1 1 * 1 1 1 1 1 1 ( ) ( ) ( ) ( ) . ( ) ( ) x t t x r õx x t t t t x t x õx x t t t t t t k p e p q q e k p p p q q e p                    then: x x t t tx t t tx tt q p p kq p p kee                         1 1 1 1/1/ 1)(  (2.8) and then after intertemporary separation of variables and notation 1 x  : ./1/        t t tx t t t t q p p kq p p ke (2.9) 3. verification of the model of course, to verify the model most appropriate would be the crisis period 2013 2015, associated primarily with the significant depreciation of the ruble. the starting point of the research period is december 2013. at the moment, both the domestic macroeconomic situation and the international environment were characterized by stability. the exchange rate of the ruble according to the bank of russia at that time was 32,73 to the us dollar. 92 a. kuzmin copyright ©2019 assa adv. in systems science and appl. (2019) we used the cpi (consumer price index) as the model variable p [15], the prices of blend crude oil as the model variable p* (us dollars per brent-mix barrel, intercontinental exchange, [16]), the index of the real gdp as the model variable q [17]. our setting used the least squares method with normalization of the investigated variables. thus, parametric minimization problem            t t tt e ee 2 )nominal( )nominal()( min   (3.1) was solved here under the conditions 0 . the exchange rate e(nominal) is the nominal exchange rate usd/rur [17], according to the bank of russia calculations monthly at the end of the each period. solution of parametric optimization (3.1) in accordance with numerical simulation is 5,0 . here we use the updated data compared with the data of work [14]. in figure 3.1 (author's calculations, monthly updated data) presents the dynamics of the exchange rate of the russian ruble against the us dollar in accordance with the result formula of our research (2.9) e(theor) and the nominal exchange rate e(nominal). fig. 3.1. the nominal and calculated exchange rates of russian ruble usd/rur (2013 2015, monthly updated data). as results, the average of absolute normalized deviations of the calculated and nominal exchange rates is 0.3% and average of normalized deviations is 3%. this allows us to speak about the high quality of the developed mathematical model. 4. conclusion the results of the modeling generally confirm the conclusions made in the media about the causes of the currency crisis in this period. a serious drop in oil prices and closely correlated the prices of natural gas and processed products was the main reason for the actual double devaluation of the ruble. however, it should be noted that the internal situation in the form of accelerated inflation and gdp decline has made a serious contribution to this process. of course, a statement of this fact cannot say about the adequacy of the situation. in the transition to a floating exchange rate such a rigid dependence of the ruble exchange rate from the foreign economic environment requires increasing degrees of intervention by the bank of russia on the domestic forex market. references 1. frenkel j. (1976). a monetary approach to the exchange rate: doctrinal aspects and empirical evidence, the scandinavian journal of economics, 78(2), 200 – 224. exchange rate of the ruble modeling 93 copyright ©2019 assa adv. in systems science and appl. (2019) 2. mussa m. (1976). the exchange rate, the balance of payments and monetary and fiscal policy under regime of controlled floating, scandinavian journal of economics, 78(2), 229 – 248. 3. mussa m. (1984). the theory of exchange rate determination, in exchange rate theory and practice. in j.e.o. bilson and r.c. marston, eds. (pр. 13 – 78), chicago: university of chicago press, nber. 4. hooper p. & morton j. (1970). fluctuations in the dollar: a model of nominal and real exchange rate determination, international finance discussion paper (board of governors of the federal reserve system), october, №168. 5. dornbush r. (1976). expectations and exchange rates dynamics, journal of political economy, 84, 1161-1176. 6. dornbush r. (1976). capital mobility, flexible exchange rate and macroeconomic equilibrium, in recent developments in international monetary economics, e. claasen and p. salin, eds., (pр. 261-278). north-holland 7. stockman a.c. (1980). a theory of exchange rate determination, journal of political economy, 88, 673-698. 8. obstfeld m. & rogoff k. (1995). exchange rate dynamics redux, journal of political economy, 103(3), 624-660. 9. clark p.b. & macdonald r. (2000). filtering the beer a permanent and transitory decomposition / imf working paper 00/144 washington: international monetary fund. 10. clark p.b. & macdonald r. (1998). exchange rates and economic fundamentals: a methodological comparison of beers and feers, imf working paper 98/67, washington: international monetary fund, march. 11. strizhkova l.a. (2017). the relationship between inflation, exchange rate and parameters of economic policy (on example of russia), bulletin of the institute of economics of the russian academy of sciences, 5, 156-176. [in russian] 12. mirkin y.м. (2018). future dynamics of russian ruble exchange rate, finance, money, investment, 67(3), 3-7. [in russian] 13. kuzmin a. (2018). equilibrium exchange rate modeling, in eleventh international conference on management of large-scale system development (mlsd 2018), https://ieeexplore.ieee.org/document/8551843/metrics#metrics, publisher: ieee, electronic isbn: 978-1-5386-4924-4, doi: 10.1109/mlsd.2018.8551843. 14. kuzmin a. (2015). exchange rate modeling: the case of ruble, review of business and economics studies, 3(3), 39-48. 15. federal service of state statistics of russia [www.gks.ru]. 16. information agency *bloomberg*. 17. bank of russia [www.cbr.ru]. https://ieeexplore.ieee.org/document/8551843/metrics#metrics microsoft word 897 article text, copyedited.doc adv syst sci appl 2020; 03; 91-104 published online at https://ijassa.ipu.ru. classifying amorphous polymers for membrane technology basing on accessible surface area of their conformations oleg miloserdov1* 1) institute of control sciences, russian academy of sciences, moscow, russia e-mail: oleg_milos@mail.ru abstract: almost 400 amorphous polymer materials used in membrane gas separation technology are clustered on the basis of the shape of their polymer chains conformations. obtained clusters, which rely solely on the geometry of polymer chains and not on the chemical class (polyamides, polyacetylenes, etc.), are shown to discriminate polymers with respect to their transport properties, in particular, the coefficient of diffusion. the method proposed consists of several steps. firstly, realistic conformations of large polymer macromolecules are constructed using the program code developed in the rdkit environment for python. then, polymer conformations are characterized by a curve that relates the “accessible surface area” (i.e., the contact surface between the spherical model of a macromolecule and a spherical “probe”) to the radius of this probe, and also seven similar curves, which relate the polarized (neutral, positively or negatively charged, etc.) accessible surface area to the radius of the spherical probe that represents the variety of penetrant gases. an improved algorithm for surface area calculation maps out the outer surface of the macromolecule to eliminate its influence. the curves are averaged between ten polymer conformations to obtain more robust figures. finally, agglomerative clustering is used to separate different polymers in the space of these curves that align their accessible-surface-area-related quantities against the probe radius. the proposed classification of polymers can be used to develop more precise predictive models of polymers’ transport properties for the theory-guided and computer-aided materials design. keywords: machine learning, agglomerative clustering, qspr, molecular modeling, polymer membranes, gas separation 1. introduction membrane technologies of gas separation are extensively used all over the world in hydrogen and oxygen production, natural gas purification and carbon dioxide separation. despite the great variety of materials used in liquid, metal, ceramic and polymer membranes, there is still a high need for a scientifically-based search for new materials, especially for polymer membranes. the development of mathematical models to predict properties of polymers from their structure allows saving financial and time resources during the development of new membrane materials. unfortunately, the mathematical models for transport properties of polymeric materials so far have either insufficient predictive power or are limited to specific polymers or their classes. primarily, this is due to the fact that transport properties of polymers used for membrane gas separation are determined by the geometry of their polymer chains, characterized by a wide variety of chemical structures and methods of their organization (spatial isomerism, statistical sequence of polymer units, etc.). several approaches are known from the literature to predict transport properties of polymers. those most popular and well-founded are based on computer simulation of atoms and their interactions [19,8]. first of all, they include molecular dynamics (md) and the grand canonical monte carlo method (gcmc) [13] that are used to predict solubility and * corresponding author: oleg_milos@mail.ru 92 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) diffusion coefficients. while md is used to predict diffusion, gcmc models have shown their effectiveness in predicting solubility coefficients. one of the common shortcomings of both md and gcmc is their high computational complexity. they require a lot of computing resources and machine time. therefore, these methods seem to be too timeconsuming for mass prediction of the properties (screening) of hypothetical (previously not synthesized) polymeric materials. at the same time, there is an alternative approach to the prediction of sorption properties of polymers and the prediction of their transport parameters. this is a typical task for qspr methods [17,21]. their goal is to establish statistically significant correlations between the numerical values of the physicochemical properties of materials and molecular descriptors. earlier, in [11], the method was proposed and then improved in [10] for predicting the sorption parameters of polymers (gas solubility coefficient at infinite dilution s and the solubility constant of henry's law kd in the model of double sorption) for a number of light gases in glassy polymers. the approach [11,10] is based on calculating the dependence of a number of geometric indices on the effective penetrant radius for conformations of a polymer chain segment (200-600 atoms length), obtained from computer modeling of a polymer chain segment. in this article this approach to molecular modeling and data analysis is further developed and applied to building a classification of glassy polymers used in membrane technology. the isolated clusters turn out to be closely related to the physicochemical properties of substances that are important for the membrane technology. in the future, it can be used to construct refined regressions that can predict the main transport properties of polymeric materials — solubility, diffusion, and permeability coefficients. 2. molecular modeling 2.1. molecular mechanics modeling problems of quantitative property forecasting from statistics and those of massive polymer classification impose some requirements on the underlying approach to molecular mechanicsmodeling (hereinafter mmm). 1. conformations (macromolecule atoms’ positions on a scene) obtained from mmm should be as realistic as possible (conform to real positions of atoms in a polymer membrane). 2. manual tuning of the mmm algorithm should be avoided, and the algorithm should be as automated as possible. 3. the mmm algorithm should be able to parallelize on a server and have an acceptable calculation time for one polymer. 4. the method should ensure the stability of the results and their reproducibility. 5. the method should be applicable to specific polymers used in membrane gas separation. 6. the method should use the free software. in this paper, we further improve the mmm method from [11,10], which is based on the modeling of conformations of the polymer macromolecule of about several hundred atoms size. in [11], for each polymer, an oligomer (200–600 atoms length) consisting of several polymer monomer units, was created in the instantjchem molecular simulation environment [4]. for this chain, using conformer plugin [6] of the instantjchem chemaxon package, several conformations of a short chain segment of the polymer under consideration (typically including several monomer units and from 200 to 600 atoms) are generated by relaxation in the empirical dreiding field from random initial positions of atoms. further calculation of the indices necessary for constructing the regression model was also performed using instantjchem tools. despite the fact that, using the obtained conformations, one of the transport parameters of polymer membranes, s solubility, was successfully predicted, this approach did not meet many criteria, listed above. the most important drawback is the classifying amorphous polymers for membrane technology … 93 copyright ©2020 assa. adv. in systems science and appl. (2020) unrealistic nature of the many conformations obtained, as well as the difficulty in expanding this solution to server platforms. in [10], the problem of the unreality of the obtained conformations was solved by performing their modeling in a program in the perkinelmer chem3d package [5] using a similar technique. to assign a random initial position of atoms to the simulated polymer chain, molecular mechanics simulation was performed at a temperature of 300 k to 3000 k for 1000 iterations, after which the resulting structure was optimized by free energy in the empirical field mm2. this approach showed good results in terms of realistic conformations, but made it impossible to automate the mmm process and did not allow the efficient use of multiprocessor systems. during the search, many different mmm solutions were tested, for example, lammps [12], openmm [7], hoomd-blue [3,9] and others. some solutions did not allow automation of the process (lammps), the others (e.g., openmm) were developed for biopolymers and had no empirical force fields relevant for glassy polymers used in membrane gas separation. rdkit package [16] an open-source toolkit – is widely used by the scientific community to solve various problems in chemoinformatics and machine learning. the main data structures and rdkit algorithms are written in c++, which ensures high performance. rdkit also has shells in python, java, and c #, which makes it easy to use. to arrange the molecule in 3d space in the rdkit environment, the flexible embedmolecule procedure is used [15]. the initial coordinates of atoms positions in molecule can be determined both through the eigenvalues of the distance matrix, and through a random positioning of atoms, and the value of the random number generator can be fixed, which will later make it possible to obtain the same molecule. then mmm is performed in an empirical force field uff (universal force field [14]). it is important that rdkit tools allow to fix the part of a molecule during mmm. the polymer chain is built by a simple algorithm, which uses the embedmolecule function of rdkit and roughly mimics the real process of polymerization. at the beginning, the monomer unit of the polymer is located in the space, then a similarly processed unit is added to it, but with other random coordinates. all atomic coordinates, except the atomic coordinates of the last two added monomer units, are fixed, and two free units undergo energy minimization in the uff field, which simulates the process of sequential polymerization of a macromolecule. in this case, the algorithm takes into account the experimental values of the torsion angles and uses the base of universal rules for the mutual arrangement of atoms, for example, the mandatory presence in the same plane of the atoms of the benzene ring. in less than half of the samples, due to the peculiarities in their structure, attempts to construct a macromolecule in this way failed after several iterations. in this case, the procedure of placing the atoms of a molecule in space was performed with the use of random initial coordinates for the entire macromolecule, and not its individual monomer units. examples of the obtained conformations are presented in fig. 2.1. 94 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 2.1. examples of three polymer repeating units are shown above, and the corresponding macromolecules of about 700-1200 atoms size obtained with mmm are shown below 2.2. calculation of geometric indices as in [10], the curves that represent the value of various accessible-surface-area-based geometric indices of polymers’ conformations as functions of the radius of a spherical “probe” are used to characterize polymers, since the main hypothesis of this method is that transport properties of a polymer (e.g., solubility coefficient and the similar measures) to some extent depend on the characteristics of the contact surface between the polymer macromolecule and the penetrant gas molecule. the basic and the simplest index of this family is the accessible surface area (asa). its calculation is illustrated by fig. 2.2. asa is the surface area circumscribed by the center of the spherical "probe" of a given radius in all possible positions of its contact with the van der waals surface of this molecule. fig. 2.2. examples of calculating the accessible surface area of a molecule the whole family of the employed indices is presented in table 2.1. classifying amorphous polymers for membrane technology … 95 copyright ©2020 assa. adv. in systems science and appl. (2020) table 2.1. geometric indices index index description asa standardized (per cm3 of the sample) accessible surface area (å2⋅mol/cm3) asa+ accessible surface area where contact occurs at a point on the surface with a partial positive (q > 0) charge* (å2⋅mol/cm3) asa– accessible surface area where contact occurs at a point on the surface with a partial negative (q < 0) charge* (å2⋅mol/cm3) asah accessible area of hydrophobic (with low, |q| < 0.125, level of partial charge*) surface (å2⋅mol/cm3) asap accessible area of polarized (with high, |q| ≥ 0.125, level of partial charge*) surface (å2⋅mol/cm3) dpsa3 dpsa3 = ∑i asai⋅qi, where asai – is the contribution of the i-th atom to the specific accessible surface area of the molecule and qi is the partial charge* of i-th atom (å2 e моль/см3) * ppsa3 ppsa3 = ∑i asai⋅qi, where the summation is limited to atoms with a positive partial charge*: asai⋅ qi > 0 (å2e mol/cm3) pnsa3 pnsa3 = ∑i asai⋅qi where the summation is limited to atoms with a negative partial charge: asai⋅qi < 0 (å2e mol/cm3) * partial charges according to gasteiger–marsili were calculated using the rdkit.chem.rdpartialcharges module. ** here “e” means the electron charge unit. compared to the previous versions of this method, the new tools of mmm presented in this article allowed to build bigger molecules and more complex conformations, which address not only the interaction of adjacent atoms in a polymer chain but also that of atoms came close to each other due to the bending of a macromolecule chain (see fig. 2.1). also, a new improved index calculation algorithm was implemented, which excludes from consideration the outer surface of the oligomer, and thus takes into account interactions of a penetrant with the polymer macromolecule. the macromolecules in fig. 2.1 are rather short polymer chains, and their outer shell has considerable surface area. realistic polymer matrices are much bigger and have no such outer shell. the contribution of atoms that form these outer shells distorts the value of surface-area-based indices from table 2.1, and their contributions must be excluded from index calculation. therefore, based on the center-of-mass position of the molecule and the value of the current “probe” radius r, the indices from table 2.1 were calculated only for atoms of the macromolecule located at no more than d = 10 – r angstrom from its center of mass. hence, the atoms that belong to the outer shell of the conformation were excluded. to draw the dependences asa(r), asa+ (r), etc., the index value was calculated for each r from 0å to 0.2å with a step of 0.05å, from 0.2å to 1å with a step of 0.1å and from 1å to 2å with a step of 0.2å. typical dependencies of a number of geometric indices from table 2.1 from the “probe” radius for one of the polymers, calculated for different d, are presented in fig. 2.3. one can see from fig. 2.3 that curves for d = 7å – r resemble those for d = 10å – r, so the adopted approach is robust with respect to small deviations of the selection rule for “internal” atoms. to justify the sufficient chain length and the number of generated conformations for the stability of the results obtained below, 6 conformations of different lengths (200, 300, 400, 600, 800, atoms without hydrogen) were generated for 40 different polymers. for them, indices from the entire set were calculated for a number of r values from the range from 1 å to 3 å. in fig. 2.4 presents, as an example, the obtained dependences of three indices (at different r) on the macromolecule size for 16 random polymers from 40 calculated. it can be seen from the figure that the indices begin to stabilize when the length of the macromolecule exceeds 600 atoms, which was taken as the base length. 96 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) the same logic shows that averaging over 10 conformations is enough to obtain stable index values (see fig. 2.5). fig. 2.3. typical dependencies of a number of geometric indices from table 2.1 from the “probe” radius for one of the polymers, calculated for various d fig. 2.4. asa, asa+ and dpsa3 indices vs the macromolecule size for 16 random polymers classifying amorphous polymers for membrane technology … 97 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 2.5. asa, asa+ and dpsa3 indices averaged over conformations vs the number of conformations being averaged for 16 random polymers 3. classification of conformational structures in machine learning problems correctly collected and pre-processed data play a key role in obtaining a qualitative result. the database “gas separation parameters of glassy polymers” (hereinafter referred to as the database) was created in 1998 at the laboratory of membrane gas separation at the topchiev institute of petrochemical synthesis, russian academy of sciences. the database is a unique source of information and a tool for predicting gas separation properties of polymers. since its inception, the database is being continuously updated. the main transport characteristics of polymer gas separation membranes are coefficients of permeability, diffusion, and solubility. the permeability coefficient p can be written as a product of the solubility coefficient s and the diffusion coefficient d. the first (s) describes the driving force of the process of transporting gas molecules, while the second (d) corresponds to the kinetic component of the process. to solve the clustering problem, 397 unique polymers were selected from the database, each having several records containing experimental data on the interaction of the gaspolymer pair under different conditions (2629 records in total). while the database covers 98 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) more unique polymer structures, for the clustering problem it is necessary to maintain a balance in the data on their chemical classes. therefore, the structures were selected according to the rule: • if more than 42 structures of a certain chemical class exist in the database, then records that have experimental data for p, d, s were selected, • otherwise, all structures of this chemical class were selected. as justified in section 2 above, for each polymer in a sample ten conformations were generated of size at least 600 atoms (not counting hydrogen atoms). see fig. 2.1 for the examples of conformations built. then, for each conformation and for each r in the range from 0å to 3å, the indices from table 2.1 are calculated using the selection rule d = 10å – r, which rules off the atoms of the outer hull of the conformation. typical dependences of geometric indices asa, asa+ and dpsa3 on the “probe” radius for one of the polymers are shown in fig. 3. in contrast to [10], where the optimal range of [r–, r+] linearization of the obtained dependencies was searched for, the obtained slope and bias coefficients were used as an explanatory regression variables, in this article the entire curve is used for classification. for the clustering of polymers, the agglomerative method (a sort of hierarchical classification) was used [1]. the agglomerative clustering method was launched for the number of clusters from k = 2, ..., 30, since the optimal number of clusters was unknown. according to the values of the silhouette coefficient, the calinski–harabasz index (c-h) and the davies–bouldin index (dbi) for the clustering constructed, the value k = 15 has been chosen. the data supplied to the input of the algorithm was previously standardized. to illustrate the quality of the resulting clustering using the tsne algorithm [20], the placement of points in 2d space was constructed. it is clearly seen that the tsne algorithm, being different from agglomerative clustering in its nature, distributes the data points in accordance with the obtained clustering. the embedding also allows to evaluate the distance between clusters (see fig. 3.1). table 3.1. the distribution of polymers of different chemical classes over clusters chemical class cluster number sum by class 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 polyacrylates 1 3 1 9 1 15 polyesters 2 18 1 1 4 4 4 8 42 polyethers 1 8 12 4 2 2 5 3 37 polyphosphasenes 3 3 polyacetylenes 1 4 3 11 19 polynorbornenes 5 7 2 2 2 9 1 6 3 37 polysulphones 19 1 18 2 40 other n-containing 4 1 5 3 4 9 26 polyamides 11 1 16 3 2 4 37 polystyrenes 2 3 1 2 1 13 22 vinylic polymers 1 3 2 1 5 1 1 14 polycarbonates 14 1 1 3 19 polyimides 3 3 5 6 9 16 42 polyamidoimides 2 2 1 1 11 23 40 other carbo-chained 1 2 3 other hetero-chained 1 1 sum by cluster 4 15 64 34 63 20 6 34 6 31 28 11 64 4 13 397 classifying amorphous polymers for membrane technology … 99 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 3.1. data embedding in 2d with tsne algorithm (clusters are shown with different signs and colors) 4. clustering and transport properties the clusters constructed turn out to be closely related to the transport properties of polymers that are important for membrane technology, first of all, permeability, diffusion, and solubility coefficients of a polymer material with respect to different penetrant gases. the, so-called, robeson diagram [18] is a conventional representation of transport properties for a collection of polymers. the robeson diagram is constructed for a pair of penetrant gases (for example, oxygen-nitrogen) in the coordinates αp = po2/pn2 — separation selectivity of a pair of gases (oxygen-nitrogen) depending on po2 — the permeability coefficient for a more permeable gas (oxygen) and in coordinates αd = do2/dn2 — diffusion selectivity of a pair of gases (oxygen-nitrogen) depending on do2 — the diffusion coefficient (see fig. 4.1 and fig. 4.2, respectively). 100 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) robeson diagrams are often used by membrane gas separation specialists for comparative analysis of different groups of polymers. hence, it is quite natural to represent the clusters of polymers constructed in the previous section as the points on a robeson diagram to analyze the relation between the obtained classification of polymers and their transport properties. the experimental data on the coefficient of permeability and diffusion were taken from the database. since the experimental data were obtained at various temperatures, they were brought to a single standard value of 308k using the algorithm from [2]. then, the diagrams were constructed using the experimental data on the permeability (see fig. 4.1) and diffusion (see fig. 4.2) coefficients available from the database. for the convenience of the analysis, along with the points of an individual cluster the center of this cluster and the center of the entire data sample (calculated without the explicit outliers) are added to the robeson diagram. the clusters turn out to be closely related to the transport properties of polymers. the location of the center of mass of the cluster relative to the center of mass of the entire data sample, the shape of the cluster and its composition demonstrate these relationships. the first thing worth noting is that the points of most clusters are rather crowded on the robeson diagrams. consequently, materials with similar transport properties are collected in one cluster. it is important that clustering was built only on the basis of the shape and geometry of conformations of polymer molecules without using any hint like the chemical class of a polymer. now consider the clusters separately. the most revealing relationships can be found on the robeson diagram for the diffusion selectivity in figure 4.2. so, the centers of mass of large clusters 3, 5, 13 are shifted to the upper left quadrant of the general diagram. this indicates a high selectivity for d with a low diffusivity. moreover, in combination with the information from fig. 4.1, polymers from these clusters are a good combination of permeability selectivity and permeability itself. also, cluster 11 can be attributed to them if there were more experimental data on the diffusion coefficient. for the most part, polymers from clusters 6, 8, 12 turned out to be low-selective in terms of permeability and diffusion. however, they were highly permeable and highly diffuse. from table 3.1 it can be noted that, in general, clustering does not depend on the chemical class of polymers. so, the above clusters 3, 5, 6, 8, 13 consist of polymers of various chemical classes. however, there are exceptions. for example, all highly diffuse and highly permeable polymers with a low selectivity from 12 cluster are polyacetylenes. all polyphosphasenes entered cluster 8, most polycarbonates entered cluster 3, and polysulphones were divided into two clusters 5 and 10. polyamidoimides and polyimides, being often similar in structure and properties, were mainly divided into clusters 11 and 13. let us turn to the structure of polymers themselves. cluster 6 consists only of polymers with bulky fluorine substituents, which in this case belong to different chemical classes. in combination with the information that the polymers of this cluster are highly permeable and highly diffuse and, at the same time, low selective, one can judge the consequences of adding bulky fluorine substituents. moreover, cluster 7, which contains mainly polymers with a pentafluorophenyl group, is rather the opposite of low selectivity and low diffusion. this suggests that not always the presence of a large number of fluorine molecules leads to properties similar to polymers of 6. it is also interesting to compare cluster 6 and cluster 10. cluster 10 also has bulky fluorine substituents; however, monomer units are often several times longer and most polymers belong to polyamidoimides and polyimides, which are never met in cluster 6. classifying amorphous polymers for membrane technology … 101 copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 4.1. polymer clusters on the robeson chart for permeability. the center of the whole data sample is depicted with the red circle, while the center of the cluster is depicted with the green triangle 102 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 4.2. polymer clusters on the robeson chart for diffusion. the center of the whole data sample is depicted with the red circle, while the center of the cluster is depicted with the green triangle classifying amorphous polymers for membrane technology … 103 copyright ©2020 assa. adv. in systems science and appl. (2020) the above analysis reveals the deep relation between the transport properties of polymers and the shape and geometry of conformations of their macromolecules. the resulting clustering allows one to identify the signs and characteristics of polymer molecules with various extreme properties, which will undoubtedly be useful in the search and synthesis of new promising polymers. 5. conclusion based on a large sample consisting of almost 400 amorphous polymers from 16 different chemical classes used in membrane gas separation, clustering was constructed only on the basis of data on the shape and geometry of the conformations of polymer molecules. obtained 15 clusters are closely related to transport parameters important for membrane gas separation, such as permeability coefficient p and diffusion coefficient d. the method proposed consists of several steps. at the beginning, 10 realistic conformations of polymer macromolecules are constructed. for these purposes, a program code is written in the rdkit environment for python. it satisfies all the requirements listed in chapter 2: the realism and reproducibility of the resulting structures, fast automated calculation, the possibility of parallelization, and free distribution of software. then, the dependencies were calculated of 8 accessible-surface-area-based indices (see table 2.1) on the radius of a spherical probe that represents the variety of penetrant gases. an improved algorithm for calculating indices eliminates the influence of the outer shell of the macromolecule, which allows focusing on the processes occurring inside the polymer membrane. the obtained 8 curves are averaged between ten polymer conformations to obtain more robust figures. then, these dependences characterizing the polymer are used as predictors for the agglomerative clustering method. the plans for further research include constructing separate regressions for the resulting clusters in order to more accurately predict transport properties (p, d, and s coefficients). the methods developed allow us to approach the solution of the problem of obtaining substances with predetermined properties. for example, the prediction of transport parameters for a collection of hypothetic polymers, followed by the manual selection of several most promising polymers, could guide the synthesis experiments of novel polymers. acknowledgements the reported study was funded by rfbr according to research projects 18-37-00265 and 1937-90004. references 1. agglomerative clustering. recursively merges the pair of clusters that minimally increases a given linkage distance [online]. available https://scikitlearn.org/stable/modules/generated/sklearn.cluster.agglomerativeclustering.html 2. alent'ev a.ju. (2003) prognozirovanie transportnyh svojstv stekloobraznyh polimerov: rol' himicheskoj struktury i svobodnogo ob'ema. doctor of science thesis, moscow [in russian]. 3. anderson j. a., lorenz c. d., travesset a. (2008) general purpose molecular dynamics simulations fully implemented on graphics processing units, journal of computational physics, 227(10), 5342-5359. 4. chemaxon. software solutions and services for chemistry & biology [online]. available https://www.chemaxon.com 5. chemoffice professional. chemical communications software [online]. available http://www.perkinelmer.com/product/chemoffice-professional-chemofficepro 6. conformer plugin. walk-through manual on how to use the conformer plugin [online]. available https://docs.chemaxon.com/display/docs/conformer+plugin 104 o. miloserdov copyright ©2020 assa. adv. in systems science and appl. (2020) 7. eastman p., et al. (2017) openmm 7: rapid development of high performance algorithms for molecular dynamics, plos comp. biol, 13(7), e1005659. 8. fried, j. (2006) materials science of membranes for gas and vapor separation, ed. by yu. yampolskii, i. pinnau, b.d. freeman, wiley, chichester, p. 95. 9. glaser j., et.al. (2015) strong scaling of general-purpose molecular dynamics simulations on gpus, computer physics communications, 192, 97-107, https://doi.org/10.1016/j.cpc.2015.02.028 10. goubko m., miloserdov o., yampolskii yu., alentiev a., ryzhikh v. (2019) prediction of solubility parameters of light gases in glassy polymers on the basis of simulation of a short segment of a polymer chain, polymer science, 61(5), 718– 732. 11. goubko m., miloserdov o., yampolskii yu., alentiev a., ryzhikh v. (2016) a novel model to predict infinite dilution solubility coefficients in glassy polymers, journal of polymer science part b: polymer physics, 55(3), 228-244. 12. lammps. molecular dynamics simulator [online]. available https://lammps.sandia.gov/ 13. norman, g.e.; filinov, v.s. (1969) high temp, 7, 216. 14. rappe a. k. et.al. (1992) uff, a full periodic table force field for molecular mechanics and molecular dynamics simulations, journal of the american chemical society, 114 (25), 10024-10035. 15. rddistgeom module. module containing functions to compute atomic coordinates in 3d using distance geometry [online]. available https://www.rdkit.org/docs/source/rdkit.chem.rddistgeom.html 16. rdkit. open-source cheminformatics software [online]. available http://www.rdkit.org 17. robeson l.m. et. al. membr. sci., 1997, 132, 33. 18. robeson, l.m. (1991) correlation of separation fac-tor versus permeability for polymeric membranes, journal of membrane science, 62 (165), 165–185. 19. theodorou, d. (2006) materials science of membranes for gas and vapor separation, ed. by yu. yampol-skii, i. pinnau, b.d. freeman, wiley, chichester, p. 49. 20. tsne. t-distributed stochastic neighbor embedding [online]. available https://scikit-learn.org/stable/modules/generated/sklearn.manifold.tsne.html 21. yampolskii yu., shishatskii s., alentiev a., loza k. (1998), j. membr. sci. 148, 59. adv syst sci appl 2018; 1; 110-131 published online at http://ijassa.ipu.ru. models of adaptation in dynamical contracts under stochastic uncertainty mikhail. v. belov1, dmitry. a. novikov2 1) ibs, moscow, russia e-mail: mbelov59@mail.ru 2) trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: novikov@tushino.com abstract: in this paper, we establish new sufficient conditions for the optimality of lump-sum and compensatory incentive schemes (contracts) under stochastic uncertainty. in addition, we suggest and analyze dynamical models of principal’s and agents’ adaptation to changes in statistical characteristics of an external environment. keywords: contract theory, incentive problem, stochastic uncertainty, adaptive behavior. 1. introduction incentive problems in organizational systems, i.e., stimulation of controlled subjects to choose certain actions for the benefit of a control subject, are studied in organizational control [12, 13] and contract theory [21, 28]. within the latter, much attention is paid to the situations in which the results of activity of controlled subjects depend on their own actions and also on external random factors. if the participants of an organizational system have a repeated interaction with the course of time, then dynamical contracts arise naturally. these contracts are described using the methods of repeated games (in discrete time [10, 24, 27] or continuous time [22, 29]), see the surveys in [11, 23, 27]. however, in some situations the characteristics of an external uncertainty change with time, and hence it is necessary to develop and analyze models for a proper consideration of such effects by organizational system participants, including their efficient response. this paper is organized as follows. in section 2, we introduce a classification system for the models of contracts in organizational systems and consider the static models of contracts, particularly the model under additive uncertainty (subsection 2.4) and the simple agent model (subsection 2.5). for these models, we establish new sufficient conditions for the optimality of lump-sum and compensatory incentive schemes (contracts). section 3 is dedicated to contracts in multiagent systems while sections 4 and 5 to the dynamical models of earned value and adaptation processes of organizational system participants to changeable characteristics of an external environment. 2. static model consider an organizational system (os) [[13]] that consists of a single control subject (principal) and a single controlled subject (agent). the agent chooses a nonnegative action y ≥ 0. in combination with a realized value of an uncertain parameter (known as the state of nature   [0, ∆]), the action uniquely defines the result z = y –  of his activity. this setup is called the additive uncertainty model, as the uncertainty distorts the agent’s action models of adaptation in dynamical contracts 111 copyright ©2018 assa. adv. in systems science and appl. (2018) additively. assume the agent’s cost c(y, r) depends on his action y and type r > 0 (a parameter that reflects the efficiency of the agent’s activity). let c(y, r) be a strictly monotonically increasing smooth convex function in y that vanishes in the origin and also a strictly monotonically decreasing function in r. the agent’s cost representation as a function of two variables is standard in organizational control (e.g., see [5–13]). in most cases considered below, the agent’s type r is fixed and/or conceptually insignificant. for the sake of compactness, we will omit r in the cost function, using the notation c(y) whenever no confusion occurs. the principal offers the agent to conclude a contract σ(z) that determines a nonnegative reward (and conditions to obtain it) depending on the achieved result. the agent’s goal function is the difference between the incentive and cost functions, i.e., ).,()(),),(( ryczzyf −=  (1) the principal obtains an income h(z) from the agent’s activity (where h(∙) is a continuous nondecreasing function) and also bears incentive cost. in other words, the principal’s goal function is the difference between the income and incentive functions: ).()()),(( zzhz  −= (2) accept the following sequence of moves in this system, which is traditional for the incentive problems in organizational control [[2], [12], [13]] and contract theory [[21], [26], [28], [30]]. first, the principal offers the agent a contract; second, the agent chooses his action in response to this offer and the payments are made accordingly. note that the agent may reject the contract, choosing the zero action. in this case, he obtains no reward but bears no cost. assume both participants of this organizational system seek to maximize their goal functions. since the result of the agent’s activity depends on his action and also on the state of nature, the principal and agent have to use all available information in order to eliminate the external uncertainty. there are the following types of uncertainty depending on the awareness of the system participants. – no uncertainty (the deterministic case in which ∆ ≡ 0, and this is common knowledge for the principal and agent); – complete awareness (both subjects know the true value of the state of nature); – interval uncertainty (both subjects know merely a set of admissible values for the state of nature, i.e., a range [0, ∆]); – probabilistic uncertainty (both subjects know a probability distribution over the set of admissible values for the state of nature or for the result of the agent’s activity); – fuzzy uncertainty (both subjects know a membership function for an uncertain parameter defined on the set of admissible values). denote by and <φ(σ(∙), y)> the deterministic goal functions of the agent and principal, respectively; these functions are obtained after elimination of the existing uncertainty for the state of nature. as a rule, an interval uncertainty is eliminated using the principle of maximum guaranteed result (see a survey of uncertainty elimination methods for incentive problems in [[12]]); a probabilistic uncertainty, using the principle of expected utility; a fuzzy uncertainty, the principle of maximally undominated alternatives. designate as p(σ(∙)) the set of agent’s optimal actions that are implemented by the principal within a contract σ(∙), i.e., )}.),(({maxarg))(( 0 yfp y =   similarly, designate as m the set of admissible contracts. accept the hypothesis of benevolence [[12], [13]], stating that the agent chooses the most profitable action for the principal from the set of implementable actions. then the incentive problem is to find an optimal contract σ*(∙), i.e., an admissible contract that maximizes the principal’s goal function: 112 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) )}}.),(({max{maxarg)( )( * y pym     (3) if the hypothesis of benevolence is not satisfied, the problem is to find an incentive scheme (contract) that yields the maximum guaranteed efficiency: )}}.),(({min{maxarg)( )( * y pym g     even for the model with one uncertain parameter, there are many possible combinations of the principal’s and agent’s awareness (more specifically, the number of such combinations is 17 = 1 + 42, as the cases of nontrivial mutual awareness [[14]] are beyond the scope of this paper). let us study some of them. 2.1. deterministic case here z ≡ y, and the solution of problem (3) is a lump-sum incentive scheme based on the principle of cost compensation for the agent [[1], [3], [13]]. it has the form 0 0 0 0 ( ) , , , ( ) 0, , c c x y x y x x  y       =    (4) where the optimal plan (the agent’s action desired by the principal) is )}.()({maxarg 00 zczhx z −=  (5) substituting (4) and (5) into (2) gives the principal’s optimal payoff )}.()({max 00 zczhk z −=  (6) the corresponding value of the agent’s goal function is 0. if the hypothesis of benevolence fails, the solution is the ε-optimal incentive scheme 0 0 0 ( ) , , ( , ) 0, ;gc c x y x x y y x   +  =   (7) where ε > 0 represents an arbitrarily small constant. 2.2. complete awareness of principal and agent assume that the agent and principal know the realized value of the state of nature before choosing the action and incentive scheme. then the principal can use the so-called flexible planning mechanism [[9]] in which the optimal plan x*() = arg 0 max y [h(y – ) – c(y)] (8) and the incentive scheme * * * * ( ( )), ( ) , ( ( ), ) 0, ( ) , c c x z x x z z x          − =   − (9) both explicitly depend on the state of nature . the value of the agent’s goal function is 0 while the principal’s optimal payoff takes the form k() = 0 max y [h(y – ) – c(y)]. (10) obviously, as the agent’s cost function is nondecreasing, we obtain k() ≤ k0 0  . this means that the uncertainty has a negative effect on the principal’s payoff. 2.3. interval uncertainty if there is incomplete awareness in the system, we should discriminate between the cases of symmetrical (identical) and asymmetrical awareness of the principal and agent. (a standard hypothesis is that the agent has at least the same awareness of the uncertain parameters as the principal [[13]]. hence, considering the case of asymmetrical awareness below, we assume that the agent knows the realized value of the state of nature while the principal makes his decision under uncertainty). note that the agent may report information to the principal [[2], [13]], this setup is not studied here. models of adaptation in dynamical contracts 113 copyright ©2018 assa. adv. in systems science and appl. (2018) asymmetrical awareness. in this case, the principal knows merely the admissible range [0, ∆] for the state of nature and hence is forced to guarantee cost compensation for the agent. xmgr = arg 0 max y [0 ] min ;  [h(y – ) – c(y)] = (11) = arg 0 max y [h(y – ∆) – c(y)]. the principal uses the incentive scheme σc(xmgr, z) = mgr mgr mgr ( ), , 0 . c x z x , z x  −    −  (12) the agent knows the realized value for the state of nature  and chooses the action y*() = xmgr – ∆ + , (13) which leads to the expected result (xmgr – ∆) of his activity. the agent benefits from plan fulfilment, which is easy to check directly by comparing his payoffs. for any state of nature, the principal’s optimal payoff k∆ = 0 max y [h(y – ∆) – c(y)] (14) under asymmetrical awareness is not higher than under complete awareness (see (10) and (14)). the agent’s goal function takes the value f(σc(xmgr, xmgr), y*(), xmgr) = c(xmgr) – c(xmgr – ∆ + ) ≥ 0 . (15) value (15) is called informational rent, i.e., the payoff obtained by a subject (in our model, agent) owing to better awareness in comparison with other subjects (the principal). symmetrical awareness. in this case, neither the agent nor the principal know the realized value of the state of nature at the moment of decision-making. the only available information is the admissible range [0, ∆]. thus, the principal uses the incentive scheme (12) and obtains payoff(14), whereas the agent is forced to choose the action guaranteeing a nonzero reward to him: ymgr = xmgr . (16) as a result, the agent has zero payoff. note that the organizational system suffers from the overproduction (∆ – ) ≥ 0. 2.4. probabilistic uncertainty: additive model let the uncertain state of nature  be a random variable with a continuous distribution function f̂ (∙): [0; ∆] → [0; 1], which has a density function p(∙). in the sequel, we will use a distribution function fθ(∙): (–∞; +∞) → [0, 1] of the form fθ(ξ) = 0, 0, ( ), [0, ], 1, . f̂            assume there is asymmetrical awareness of this uncertain parameter in the system. for a given agent’s action y, the result of activity z = y –  is a random variable with the distribution function fz(∙, y): [y – ∆, y] → [0, 1] defined by fz(q, y) = 1 – f(y – q) . (17) let the principal’s choice be limited to the parametric class of lump-sum incentive schemes σc(π, z) = , , 0, , z z       (18) where π ≥ 0 and λ ≥ 0 denote some parameters, π is a planned result (the result of the agent’s activity desired by the principal). (as demonstrated in [[6], [12]], in the case under study the lump-sum incentive schemes can be not optimal, see the discussion below; nevertheless, they are simple and widespread in applications). under an action y ≥ π chosen by the agent, the expected value of his reward (18) is 114 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) ez σc(π, z) = λf(y – π) (19). suppose the agent seeks to maximize his expected utility [[12], [13], [21]]. by the firstorder optimality conditions, the agent’s action y*(π, λ) ≥ π satisfies the equation λ p(y *(π, λ) – π) = c(y*(π, λ)). (20) the principal’s problem is to choose the parameters π ≥ 0, λ ≥ 0 of the incentive scheme (contract) in order to maximize his expected utility: .max)),(()()),(( 0,0 * 0 *   →−−−    yfdpyh (21) example 1. assume the principal has a linear income function h(z) = γz, where γ > 0 is a given constant; the agent has a quadratic cost function c(y, r) = y2/2r, where r > 0 is his type [[13]], which describes the efficiency of the agent’s activity; the distribution f(∙) is uniform, i.e., f(v) = v/∆, v  [0, ∆]. according to formula (19), the expected reward of the agent is λ(y – π)/∆. hence, for any action y ≥ π the agent receives the expected payoff ez f(σc(π, z), y, z) = λ (y – π)/∆ – y2/2r. (22) maximizing his expected payoff (22), the agent chooses the optimal action (also, see expression(20)) * for , 2 ( , ) 0 for . 2 r r y r            =      (23) the expected value of the principal’s goal function is ez φ(σc(π, z), z) = γ(y – ∆/2) – λ(y – π)/∆. (24) substituting (23) into (24), we obtain the following optimization problem for the parameters of the incentive scheme (18) (see (21)): γ (λr/∆ – ∆/2) – λ(λr/∆ – π)/∆ → 0, /(2 ) max r     . (25) the solution of problem (25) has the form λ* = γ∆, π* = γr/2. then the expected payoff of the principal is γ(γr – ∆)/2 while the expected payoff of the agent is 0. •* let us design the optimal incentive scheme for the probabilistic additive uncertainty model. the general solution procedure of probabilistic incentive problems is as follows [[12]]. first, for each agent’s action x, find a minimal incentive scheme min(x, ) in terms of the expected incentive cost of the principal that implements this action, i.e., the scheme stimulating the agent to choose x  p(min(x, )). second, find the action that is most beneficial to implement in terms of the principal’s goal function, i.e., the one maximizing his expected payoff (also see expression (3)): x* = arg 0 max x eφ(min(x, ), x), (26) where e means the expectation operator. denote x** = arg max x [ 0 ( ) ( )h x p d    − – c(x)]. lemma 1: in the probabilistic incentive problem, for any agent’s action x ≥ 0 there does not exist an incentive scheme implementing this action with the principal’s expected incentive cost strictly smaller than the agent’s cost, i.e., min(x, ) ≥ c(x). * throughout the paper, the symbol “•” indicates the end of an example or proof. models of adaptation in dynamical contracts 115 copyright ©2018 assa. adv. in systems science and appl. (2018) proof. assume on the contrary that there exists an agent’s action 0x  and an incentive scheme ( )z such that ez ( )|z x = 0 ( ) ( )zz p v, x dv +  < ( )c x , (27) and the incentive scheme ( )z implements the action 0x  , i.e., for any y  0, ez ( )|z x – ( )c x ≥ ez ( )|z y – c(y). (28) for y = 0, by c(0) = 0 inequality (28) takes the form ez ( )|z x ≥ ( )c x + ez ( )|0z . this result contradicts (27), as the incentive and its expected value are nonnegative. • we will establish sufficient conditions for the optimality of an incentive scheme of form (18), namely, the contract σc(x, z) = ( ), , 0, c x z x z x .  −    −  (29) proposition 1: for all x  [x** – , x**], let ** ** ( ) ( ) ( ) c x p x x c x   −  . (30) then in the probabilistic additive uncertainty model the incentive scheme (29) implements the agent's action x** with the minimum expected incentive cost c(x**) of the principal and is optimal. proof. calculate the expected value of the agent’s reward from choosing an action y given a plan x: ez σc(x, z)|y = 0 c ( ) ( ), px  y d    −  − . it follows from (19) that ez σc(x, z)|y = c(x) f(y – x + ∆). the agent’s expected utility is ez σc(x, z)|y – c(y) = ( ), , ( )  ( ) ( ), [ , ], ( ) ( ), c y y x c y y x x c x c c x f y  x y y x     .  − −  −   −  −  −  +   given the plan x = x**, by condition (30) the maximum value of this expected utility is achieved at the zero action or at the action coinciding with the plan x (condition (30) guarantees that the agent’s expected utility is a nondecreasing function in his action on the interval [x** – , x**]). on the strength of the hypothesis of benevolence, the agent chooses the action x**, which makes the expected incentive cost of the principal equal to the agent’s cost. hence, by lemma 1 the incentive scheme (29) is optimal. • obviously, if the hypothesis of benevolence fails, then under condition (30) the ε-optimal incentive scheme is σcg(x, z) = ( ) , , 0, c x z x z x . +  −    −  example 2. for the data of example 1, condition (30) takes the form γr ≥ 2∆. • to proceed, study sufficient conditions for the optimality of a compensatory incentive scheme. to this end, find an incentive scheme ( )ˆ z that nullifies the agent’s expected utility for any actions, i.e., ( ) ( ) ( ) y y ˆ z p y z dz c y − − = . (31) proposition 2: if there exists a contract ( ) 0ˆ z  satisfying relationship (31) then it is optimal in the probabilistic additive uncertainty model. 116 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) the validity of proposition 2 follows from property (31) of the incentive scheme ( )ˆ z and lemma 1. therefore, an incentive scheme ( )ˆ z is optimal if there exists a positive solution of the integral equation (31). let us formulate the existence conditions of such a solution. proposition 3: if the functions ( )yc y and ( )p v  are continuous on the domains of definition, then the integral equation (31) has a unique solution, which can be calculated using the sequential approximations        = − −  =  = + ....,2,1, )0( )( )(ˆ )0( )( )(ˆ , )0( )( )(ˆ 0 1 0 idu p uzp u p zc z p zc z z ii       (32) for this solution to be positive, a necessary condition is ....,2,1,)()(ˆ)( 0 =−  iduuypuyc y i  (33) proof. differentiation of the integral equation (31) yields ).()()(ˆ)(ˆ)()(ˆ)0( zcduuzpuzpzp z z =−+−−  −   (34) first, solve (34) for z  : ).()()(ˆ)(ˆ)0( 0 zcduuzpuzp z =−+    write this expression as . )0( )( )(ˆ )0( )( )(ˆ 0 du p uzp u p zc z z     − −  =  (35) equation (35) is a volterra integral equation of the second kind. as is well-known (e.g., see [[17]]), under the accepted hypotheses it has a unique solution, which can be calculated using the sequential approximations (32). condition (33) is necessary for the positive solution. thus, we have completely solved the problem for z  . based on (34), it is possible to obtain a similar solution for z  [j; (j + 1)], j = 1, 2, …: 0 ˆ{ ( ) ( ) ( )} ( ) ˆ ( ) ( ) . (0) (0) z c z p z p z u z u du p p         +  − − = −  note that it suffices to prove the positivity of ( )ˆ z on the first interval only, as c'(z) + p()(z – )  c'(z) for the subsequent intervals. example 3. consider the uniform density function )(p defined by 0 1 0 1 1 for [ , ], ( ) 0 for [ ; ], a a p a a       =    where 0 < a0 and a1 = a0 + . then (31) can be written as  − − = 0 1 ).,()( az az rzcdzu differentiation of both sides leads to the functional equation (z – a0) – (z – a1) = с'(z). since c(y) = 0 and c (y) = 0 for y  0, the solution of this equation is given by the functional series   = −=− 00 ).()( i izcaz models of adaptation in dynamical contracts 117 copyright ©2018 assa. adv. in systems science and appl. (2018) recall that )(c is continuously differentiable and convex. hence, the sum of the functional series is positive and increasing, which agrees with the main requirement to the incentive scheme (). particularly, for the data of example 1, c(y, r) = y2/2r and c(y, r) = y/r for y > 0; then   = −−=− 00 ),)(()( i iziz r az  where )( is the heaviside step function 1 for 0, ( ) 0 for 0. u u u   =   • 2.5. probabilistic uncertainty: the simple agent model an alternative to the additive uncertainty model considered in subsection 1.4 is the socalled simple agent model [[3], [5], [6], [12]], in which the distribution function of activity results has the form ( ), , ( , ) 1, . z g q q y f q y q y  =   (36) here g(∙): [0; +∞) → [0; 1] is a given distribution function such that g(0) = 0, with a density g(∙). like in the additive model, the agent’s action defines the maximum possible result while the distribution g(∙) does not explicitly depend on the action. denote by pz(q, y) the density function associated with distribution (36). if the principal uses the lump-sum incentive scheme (18) and the agent chooses an action y ≥ π, the expected reward of the agent is λ(g(y) – g(π)). for the simple agent model, an analog of the first-order optimality condition (20) has the form λ g(y*(π, λ)) = с(y*(π, λ)). as proved in the book [[12]], compensatory incentive schemes are optimal in the simple agent model with a risk-neural agent. for this class of models, the following optimal incentive schemes were constructed in [[6]]: – for a risk-averse agent, the compensatory incentive schemes σk(z) = 0 ( ) 1 ( ) z c v dv g v  − ; (37) – for a risk-seeking agent, the lump-sum incentive schemes σc(x, z) = ( ) , , 1 ( ) 0, . c x z x g x z x   −   (38) clearly, the incentive scheme (37) is nonnegative, increasing and convex. by analogy with [[6], lemma 1], we may establish the following properties of the incentive schemes (37) and (38). proposition 4: in the simple agent model, for any y  0 the incentive schemes satisfy the relationships 1) 0 ( ) ( , ) ( )k zq p q y dq c y + = , 2) 0 ( ) ( , ) ( )c zq p q y dq c x + = . at the conceptual level, relationship 1) of proposition 4 means that for any action of the agent his expected reward (37) coincides with his cost to choose this action. hence, under the incentive scheme (37) used by the principal, any actions of the agent yield zero utility for him and make his choice indifferent. by the hypothesis of benevolence, the agent will choose the action that is most profitable for the principal. 118 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) if the principal adopts the incentive scheme (38), then the agent is indifferent between the zero action (reject of the contract) and plan fulfilment. for the agent’s expected utility to reach a unique maximum as the result of plan fulfilment, the principal has to increase payments for plan fulfilment by an arbitrarily small positive value ε. note that this incentive scheme is not optimal but ε-optimal. the next property can be verified directly. proposition 5: in the simple agent model, for any x  0 the incentive scheme c  (x, z) = ( , ) , , 1 ( ) 0, , c x r z x g x z x +  −   (39) is ε-optimal, i.e., it implements the agent’s action x with the minimum expected incentive cost of the principal. regardless of the scheme used by the principal (compensatory or lump-sum), his optimal plan for the agent is given by x* = arg 0 max y [ 0 ( ) ( ) y h z g z dz + (1 – g(y)) h(y) – c(y, r)]. (40) the first-order optimality condition for (40) has the form h(x*) + (1 – g(x*)) h'(x*) = c(x*, r). (41) example 4. for the data of example 1, let g(z) = z/(1 + z). then it follows from (37) that 2 2 ( ) 1 2 3 k z z z r    = +    ; using formula (41) we calculate         − − + = 1 1 31 2 1* r r x   . • this paper considers the simple agent model with a risk-neural agent. hence, choosing between the incentive schemes (37) and (38), we should give preference to the lump-sum incentive scheme, as (a) it is simpler and (b) its ε-optimal analog stimulates the agent to fulfill the plan (see proposition 5). the main results of section 1 dedicated to static problems of contract theory are the analytical relationships (31) and(37), which allow to formulate and solve complex problems (particularly, the dynamical ones with changeable characteristics of agents and/or state of nature, e.g., the parameters of distribution). first, we will extend the model with one agent and additive uncertainty to the multiagent case (section 3). then, we will generalize the static simple agent model to the case of several sequential action periods (section 4). 3. multiagent model consider an organizational system composed of the principal and n subordinate agents with simultaneous and independent decision-making. denote by n = {1, …, n} the agent set and by ci(yi, ri) = c(yi, ri) the cost function of agent i; as before, yi ≥ 0 specifies the action of agent i and ri > 0 is his type. designate as y = i i n y   the total action of all agents. assume the principal is interested in a total result x ≥ 0 of the activity of all agents with a probability not smaller than a given threshold α [0; 1]. the value α is called contract reliability [[4]]. for the additive uncertainty model, this condition takes the form y ≥ x + n 1f − (α). (42) the value n 1f − (α) can be treated as payment for uncertainty in terms of agents’ activity. consider the following problem. what are the optimal plans for actions? using proposition 1, we obtain that for each agent the expected incentive cost of the principal coincide with the agent’s cost to choose a corresponding action (under the incentive scheme models of adaptation in dynamical contracts 119 copyright ©2018 assa. adv. in systems science and appl. (2018) (31) used by the principal, each agent receives a constant expected payoff regardless of his action; hence, by the hypothesis of benevolence each agent prefers plan fulfilment). since the cost functions of the agents are nondecreasing, in the optimal solution condition (42) holds as equality. hence, optimal plan calculation is reduced to the constrained optimization problem      += →    −   ni i ni x ii nfxx rxc i ).( min),( 1 }0{  (43) using the lagrange method of multipliers, we easily establish the following result. proposition 6: in the additive uncertainty model, the optimal plans {xi *} in the contract yielding a total result x ≥ 0 with a reliability α are given by * ix = c'–1(, ri), i  n, (44) where μ > 0 is the solution of the equation   −− += ni ii nfxxrc ).(),( 11   (45) as the distribution function is strictly monotonic, a direct analysis of problem (43) leads to proposition 7: in the additive probabilistic uncertainty model, the principal’s minimum cost to implement a given total result of agents’ activity does not decrease for higher contract reliability. example 5. consider the cobb–douglas cost functions for the agents, i.e., c(y, r) = 1  yμr1–μ, μ > 1. it follows from (44) and (45) that * ix = i j j n r r   (x + n 1f − (α)), i  n . (46) using the optimal plans (46), calculate the optimal value of the goal function: ( )    −  −         + = ni nj j ii r nfx rxc . )( ),( 1 1 *      (47) the right-hand side of (47) is not decreasing in α (see proposition 5). the payment for uncertainty (the difference between (47) and the value of the goal function in the corresponding deterministic problem) constitutes 1 1 ( ( )) ( )j j n x nf x r       − −  + −  , not decreasing for higher contract reliability. • 4. earned value model as a matter of fact, the earned value model is widespread in project management, both in theory and applications. in this section, we show a connection between the suggested models of contracts and the earned value model. consider the interaction of the principal and one agent within the scope of a certain project (a sequence of discrete periods). by a period t0 (called the project completion time), the principal has to implement a given total result x0 ≥ 0 of activity. let the states of nature { t}t=1,2,… in different periods be independent random variables obeying the same distribution f(∙). assume the principal concludes the optimal contract ( )tˆ z with the agent 120 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) (see proposition 2) that satisfies (31) and specifies the agent’s reward depending on the result of his activity in period t, where t = 1, 2, … . the agent’s type and cost function are independent of periods. hence, under a given reliability α of each single-period contract, in each period the principal has to allocate the same plan to the agent given by (48)x0 = x0/t0 + 1f − (α) (compare with formula (42)). by (31), the agent benefits from plan fulfilment. the total plan for the agent’s activity by period t makes up 0 tx = t x0 – t 1f − (α) = tx0/t0. (49) in terms of the earned value approach of modern project management [[7], [16]], sequence (49) is called the budgeted quantity of work scheduled (bqws). since the result z = x0 –   of the agent’s activity in period  is a random variable, the total result xt achieved by a period t is also a random variable of the form xt = tx0 – 1 t    =  = t(x0/t0 + 1f − (α)) – 1 t    =  = = 0 tx + t 1f − (α) – 1 t    =  .(50) sequence (50) is called the actual quantity of work performed (aqwp). now, introduce other indices of the earned value approach subject to the additive uncertainty model (t = 1, 2, …, t) [[7]] as follows. – the expected planned cost of the principal, or the budgeted cost of work scheduled (bcws), defined by 0 tc = tс(x0/t0 + 1f − (α), r); (51) – the actual cost of the principal, or the actual cost of work performed (acwp), defined by ct = 1 0 )( t ˆ x        = − ; (52) – work underrun (in terms of time, positive or negative), defined by δ(t) = min {δ | 0 t tx x− = }; (53) – earned value (ev), or the budgeted cost of work performed (bcwp) as the planned cost of actually performed work, defined by t ec = ( ) 0 t tc − ; (54) – the current forecast t(t) of project completion time defined by t(t) = t0 + ε(t); (55) – total planned cost, also called the budget at completion (bac) or the budget cost (bc), defined by c0 = t0 с(x0/t0 + 1f − (α), r); (56) – the current linear estimate of total cost, defined by c(t) = t(t) ct / t; (57) – the actual project completion time, defined by t’ = min {t ≥ 0 | xt ≥ x0}; (58) – the difference between the actual and budgeted cost, or cost overrun (co), defined by ∆сe(t) = ct – ct e; (59) – schedule performance index (spi), defined by at = t ec / 0 tc ; (60) – cost performance index (cpi), defined by bt = t ec /ct. (61) models of adaptation in dynamical contracts 121 copyright ©2018 assa. adv. in systems science and appl. (2018) the budgeted cost indices (48)–(61), which are traditionally divided into primary (48)– (52) and derived (53)–(61), are efficient tools for project management, at the stages of planning and implementation. example 6. for the data of example 1, let r = 1, t0 = 100, x0 = 100, ∆ = 1, and α = 0.2. the trajectories of the total plan (49), the total result (50) and the expected result xt = 0 tx + t( 1f − (α) – e θ)) are illustrated fig. 1. simulation was performed in rds software complex [[15]]. fig. 1. dynamics of results in example 6 the dynamics of the planned and actual cost and also of earned value are shown in fig. 2. the calculations yielded t’ = 145, bac = 72, and ∆сe(t’) ≈ 60%. • fig. 2. dynamics of cost and earned value in example 5 5. adaptation model incentive problems in dynamical organization systems can be classified using different bases such as the relationship between periods, the foresight of system participants, the mode of decision-making and others, see [[11]]. in this section, we first introduce a classification of dynamical incentive problems with a single-time change of a model parameter, e.g., the principal’s goal function, the agent’s cost function or the distribution function of activity results. such a change occurs at a time td and will be further called a discord. by assumption, the system participants are short-sighted: in each period, they make decisions for this period only and do not consider the consequences in future periods. then we study a model in which a discord occurs for the distribution function while the time of discord is unknown to the principal and agent. the changes in the behavior of the system participants after total plan (black), expected result (blue), and total result (red) periods planned cost (black), actual cost (red), and budgeted cost (blue) periods 122 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) detection of the discord can be treated as their adaptation to the new operating conditions [[8], [18]]. 5.1. a classification of dynamical incentive problems in table 1 we present all possible cases for the awareness of the principal and agent about the new functions (income, cost, distribution) after a discord occurs. these functions will be denoted using appropriate symbols with tilde. assume that, before choosing their decisions for a period t, the principal knows the history zt–1 = (z1, …, zt–1) while the agent the histories zt–1 and yt–1 = (y1, …, yt–1). the awareness of the principal and agent are shown in columns 3 and 4 of table 1. 1-2. a change of the principal’s income function ( )( ~ )( zhzh → ) is considered in rows 1 and 2 of table 1. if the principal knows the new income function and the discord time td (row 1 of table 1), then the problem is reduced to a set of typical static incentive problems studied in section 1, which are solved for each period independently. however, if the new income function )( ~ zh or the discord time td are not available for the principal, the problem makes no sense: the principal does not have enough information for decision-making. 3-9. a change of the agent’s cost function ( )(~)( ycyc → ) is considered in rows 3–9 of table 1. if the agent and principal both know the new cost function and the discord time (row 3 of table 1), then the problem again is reduced to a set of typical static incentive problems, which are solved for each period independently. assume the principal is aware of the new cost function of the agent but has no information about the discord time (row 4 of table 1). in this case, his rational behavior is to offer the agent a menu of optimal contracts for a set of possible cost functions. this is the screening principle used under asymmetric awareness [[13], [21]]). if the agent knows the new cost function and the discord time while the principal neither of them (row 5 of table 1), the problem makes no sense due to the following. for obtaining a positive payoff himself, the principal has to stimulate the agent’s activity by offering a contract with a nonnegative payoff of the latter. yet, being unaware of the agent’s cost function, the principal cannot form such a contract. for the same reasons, the problems with the awareness described by rows 8, 12, and 15 of table 1 are ill-posed: the principal does not know the new cost function )(~ yc or the distribution function ),( ~ yzf . models of adaptation in dynamical contracts 123 copyright ©2018 assa. adv. in systems science and appl. (2018) table 1. classification of dynamical incentive problems n o. discord agent knows principal knows problem 1 )( ~ )( zhzh → no matter td; )( ~ zh typical 2 nothing or td ill-posed 3 )(~)( ycyc → td; )(~ yc td; )(~ yc typical 4 )(~ yc solvable by screening 5 nothing ill-posed 6 )(~ yc td; )(~ yc typical 7 )(~ yc typical 8 nothing ill-posed 9 nothing no matter ill-posed 1 0 ),( ~ ),( yzfyzf  → td; ),( ~ yzf td; ),( ~ yzf typical 1 1 ),( ~ yzf solvable by screening 1 2 nothing ill-posed 1 3 ),( ~ yzf td; ),( ~ yzf typical 1 4 ),( ~ yzf d1 1 5 nothing ill-posed 1 6 nothing no matter ill-posed let the principal know the new cost function of the agent and the discord time while the agent merely this function (row 6 of 124 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) table 1). this situation leads to a set of typical problem. really, using complete information the principal will offer a contract with c(y) before the discord and with )(~ yc after it; in turn, the agent can identify the discord time by the changing offers of the principal and then respond optimally. again, we obtain a set of typical static incentive problems. now, consider the case in which the principal and agent are both aware of the new cost function but the discord time is uncertain (row 7 of table 1 models of adaptation in dynamical contracts 125 copyright ©2018 assa. adv. in systems science and appl. (2018) table 1). recall that the agent’s cost function is continuously differentiable and strictly monotonic. hence, observing his actual cost, the agent can detect the fact of discord (if any); more specifically, the agent reliably detects the change of the cost function at the end of the period after the discord, in which the agent chooses some action y such that )(~)( ycyc  . however, in this period the agent chooses his action without any knowledge of the cost function. dealing with the short-sighted agent, the principal has to offer a contract associated with the worst-case cost function for the agent, i.e., with the function ( ) max{ ( ); ( )}с̂ y c y c y= , before the discord and after it including detection by the agent. thus, we arrive at a typical problem with additional cost of the principal, which can be assessed by both players. if the agent does not know the new cost function, the problem is ill-posed regardless of the principal’s awareness (row 9 of table 1). in this case, he cannot estimate possible loss in several periods until the new cost function is identified. the agent prefers zero action accordingly. 10-16. a change of the distribution function of activity results or of the distribution function of the state of nature ( ),( ~ ),( yzfyzf  → ) is considered in rows 10–16 of table 1. assume the principal knows the new distribution function ),( ~ yzf and the discord time while the agent at least ),( ~ yzf (rows 10 and 13 of table 1). then the problem again is reduced to a set of typical static incentive problems. let the agent be aware of both the new distribution function ),( ~ yzf and the discord time and let the principal be aware of ),( ~ yzf only. here a sequential screening problem arises naturally (row 11 of table 1). if the agent does not know the new distribution function ),( ~ yzf , the problem is ill-posed regardless of the principal’s awareness (row 16 of table 1). in this case, he cannot estimate possible loss or payoff in several periods until the new distribution function ),( ~ yzf is identified. the agent prefers zero action accordingly. model d1 (row 14 of table ) is a multiperiod contract model with a change of the distribution function ),( yzf at some time. in this model, the principal and agent both know the new distribution function ),( ~ yzf and also expect a single change of this function. but they are a priori unaware of the discord time. before period 1 of their interaction, the agent and principal have no information about the discord except the prior. hence, they have to act under the hypothesis that in period 1 the result corresponds to ),( yzf or ),( ~ yzf . dealing with the short-sighted agent, the principal has to offer a contract associated with the worst-case cost function for the agent (otherwise, the agent rejects the contract and both players do not receive new information about the state of nature, which makes further interaction unreasonable). if the principal forms such a contract and the agent is rational, then the former can predict the action y of the latter. the result z is observed by the principal, and hence the principal has the same posterior information as the agent. this fact can be generalized as the principle of transparent stimulating contract and formulated in the following way. assume the principal can form a stimulating contract while the agent is rational and does not have strategic behavior; then the principal can reliably predict the agent’s actions and use the same complete information as the agent. hence, after period 1 the principal can design a contract for period 2 using the prior knowledge about the functions ),( yzf , ),( ~ yzf and also using the observations z1, y1. in subsequent periods t, the principal will act in the same way, forming contracts t(z) based on the available functions ),( yzf , ),( ~ yzf and the observations zt–1, yt–1. 126 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) for the short-sighted principal, an alternative is to use an optimal contract given the distribution f(z, y) until he detects the discord (this situation is studied in detail below for model d1). note that, owing to identical awareness, the principal and agent detect the discord “simultaneously.” depending on the availability of additional prior information for players, there may exist several contract design statements. if both players have prior information about the distribution of the discord times, it is possible to obtain a sequential optimal bayesian algorithm to design the contract t(z). finally, if both participants possess no additional information about the possible discord times and their interaction can be terminated in any period, then the minimax approach is optimal. 5.2. discord problem (problem d1) consider the multiperiod simple agent model with unconnected periods [[10]] and let the principal use the optimal incentive scheme (38) with the optimal plan (40). assume that initially the principal and agent have the same information about the distribution function g(∙). at some time td > 0 a discord occurs, which changes the distribution g(∙) to ( )g  . the new distribution ( )g  is a priori known to the principal and agent, but none of them have information about the discord time. as the result of this discord, in a single period the expected utility of the agent varies by the value ∆f(g(∙), ( )g  ) = c(x*) * * * ( ) ( ) 1 ( ) g x g x g x − − . (62) choosing in each period the action x*, the agent (as well as the principal) observes the sequence of results zt (also the agent observes the sequence yt = (x*, …, x*)). both players have to decide whether a discord occurs or not. the sequential problem is therefore decomposed into single-period problems with additional information (owing to independent periods). for period t, the agent and principal renegotiate the contract using information about the possible distributions )(g , )( ~ g and the additional observations zt–1. define the value lt = ln( ( , ) ln( ( , ))t t t tg z y g z y− to formulate the optimal sequential maximal likelihood rule for discord detection as follows. in each period t > 0, calculate lt (l0 = 0) by 1 1 1 0, if 0, , if 0. t t t t t t t l l l l l l l − − − +  =  + +  (63) if for some period lt > , then the discord takes place. here the value  describes the error characteristics of the first and second kind. as is well-known [[19]], in comparison with other statistics the maximal likelihood statistic yields the most efficient decision rule in terms of the following criterion. one of the errors is detected with a probability not smaller than a given threshold while the other error is optimized. the threshold δ is chosen using the error characteristics of the first and second kind. example 7. for the data of examples 1 and 4, let g(z) = z/(β1 + z) and ( )g z = z/(β2 + z). then ez(y) = βi ln (1 + y/βi). using condition (41), we find x* = 4 1 1 2 ( -1) r r r        − − −     . expression (38) takes the form models of adaptation in dynamical contracts 127 copyright ©2018 assa. adv. in systems science and appl. (2018) σc(x*, z) = * 2 * * * ( ) ( ) , , 2 0, . x x z x r z x    +     (64) according to the principle of transparent stimulating contract, the agent always chooses the action         − − −−== 1 )1( 4 1 2 1 1*     yr r rxy . before the discord time, the result of the agent’s actions obeys the distribution 2 1 1* * 1 1* )( )(),( z xz x xzg + +− + =      for z  [0, x*]; after the discord time, the distribution 2 2 2* * 2 2* )( )(),(~ z xz x xzg + +− + =      for z  [0, x*], where () denotes the delta-function. then *12 1 2 * *2 1 * 1 2 ln 2ln for [0, ), ln 2ln for . t t t t z z x z l x z x x            + +      +     =     + + =    +    take t = 500, td = 200, r = 1, γ = 10, β1 = 100, and β2 = 60. the dynamics of the planned, expected and actual results (in cumulative sums) are shown in fig. 3 (dashed lines correspond to the dynamics without discord detection). the dynamics of the cumulative cost of the agent and the principal’s incentive cost are illustrated in fig. 4 (dashed lines correspond to the dynamics without discord detection). fig. 3. dynamics of cumulative results in example 7 planned result(black), expected result (blue), and actual result (red) periods 128 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) fig. 4. dynamics of cumulative cost (agent, principal) in example 7 in the cumulative sum method, we may use the discord indicators defined by 1 1 ( ) t *ts t ezz x  = = − (65) or * * 2 1 ( , ) ( , ) t t cs x z tc x r   = = − . (66) their dynamics in our example can be observed in fig. 5. let us apply the maximal likelihood rule (63). the corresponding dynamics are presented in fig. 6. the means of the statistic lt before and after the discord are –0.04 and +0.04; the root-mean-square deviations are –0.29 and 0.27, respectively. note that, before the discord, the statistic lt takes the value * 2 1 * 1 2 ln ln 0.29 x x        + + = −    +    with the probability 1 * 2 0.5 x   = + ; after the discord, with the probability 2 * 2 0.38. x   = + fig. 5. dynamics of discord indicators s1 and s2 in example 7 in our example, for δ = 2 the discord is detected after 10 periods since its occurrence. • agent’s actual cost (black) and principal’s incentive cost (red) periods discord indicators: s1 (red) and s2/100 (blue) periods discord indicator models of adaptation in dynamical contracts 129 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 6. dynamics of discord indicator in example 7 6.conclusions in this paper, we have considered contracts between the short-sighted principal and agents that operate under external probabilistic uncertainty (knight’s measurable uncertainty [[25]]) with changeable characteristics (a reflection of knight’s true uncertainty). a proper response to true uncertainty is a basic function of control subjects for the adaptive behavior of subordinate structural elements of activity [[1], [20]]. among the promising lines of future research, we mention other descriptive methods for the influence of external uncertainty on activity results, contract renegotiation conditions for the longsighted principal and agents, and discord problems in multiagent dynamical organizational systems. references [1] belov, m. & novikov, d. (2017). strukturnye modeli kompleksnoi deyatel’nosti [structural models of complex activity]. moscow, russia: fizmatlit [in russian]. [2] novikov, d. (ed.) (2013). mechanism design and management: mathematical methods for smart organizations. new york, usa: nova science publishers. [3] burkov, v., gorgidze, i. & ivanovskii, a. (1974). prostoi aktivnyi element. realizatsiya plana i pereotsenka busushchego sostoyaniya. sintez funktsii dokhoda [simple active element. plan implementation and future state assessment. income function design], in aktivnye sistemy, ser. problemy i metody upravleniya v aktivnykh sistemakh, vol. 2. moscow: institute of automation and remote control, 52–62. [4] burkov, v. & novikov, d. (1997). kak upravlyat’ proektami [how to manage projects]. moscow, russia: sinteg, [in russian]. [5] burkov, v. (1977). osnovy matematicheskoi teorii aktivnykh sistem [foundations of mathematical theory of active systems]. moscow, russia: nauka, [in russian]. [6] goubko, m. (2000). zadacha teorii kontraktov dlya modeli prostogo agents [a problem of contract theory for simple agent model], upravlenie bol'shimi sistemami, 2, 22–27. periods 130 m.v. belov, d.a. novikov copyright ©2018 assa adv. in systems science and appl. (2018) [7] kolosova, e., novikov, d. & tsvetkov, a. (2000). metodika osvoennogo ob’ema v operativnom upravlenii [earned value approach in operational management]. moscow, russia: apostrof [in russian]. [8] novikov, d. (2008). matematicheskie modeli formirovaniya i funktsionirovaniya komand [mathematical models of team formation and operation]. moscow, russia: fizmatlit [in russian]. [9] novikov, d. (1997). mekhanizmy gibkogo planirovaniya v aktivnykh sistemakh s neopredelennost’yu [flexible planning mechanisms in active systems under uncertainty], avtomatika i telemekhanika, 5, 118–125. [10] novikov, d. (1997). mekhanizmy stimulirovaniya v dinamicheskikh i mnogoelementnykh sotsial’no-ekonomicheskikh sistemakh [incentive mechanisms in dynamical and multiagent socioeconomic systems], avtomatika i telemekhanika, 6, 3–26. [11] novikov, d., smirnov, i. & shokhina, t. (2002). mekhanizmy upravleniya dinamicheskimi aktivnymi sistemami [control mechanisms for dynamical active systems]. moscow, russia: institute of control sciences [in russian]. [12] novikov, d. (1998). stimulirovanie v sotsial’no-ekonomicheskikh sistemakh [stimulation in socioeconomic systems]. moscow, russia: institute of control sciences [in russian]. [13] novikov, d. (2013). theory of control in organizations. new york, usa: nova science publishers, 2013. [14] novikov d. & chkhartishvili a. (2014). reflexion and control: mathematical models. london, uk: crc press. [15] roshchin, a. (2012). raschet dinamicheskikh system (rds). rukovodstvo dlya programmistov. prilozhenie: opisanie funktsii i strukter. prilozhenie k rukovodstvu dlya programmistov [calculation of dynamical systems (rds). manual for programmers. appendix: description of functions and structures. appendix to the manual for programmers]. moscow, russia: institute of control sciences [in russian]. [16] a guide to the project management body of knowledge (pmbok® guide), 5th ed. (2013). project management institute. [17] tricomi, f. (1957). integral equations. new york, usa: interscience publishers. [18] tsyganov, v. (1991). additivnye mekhanizmy v otraslevom upravlenii [additive mechanisms in sectoral management]. moscow, russia: nauka [in russian]. [19] shiryaev, a. (1976). statisticheskii posledovatel’nyi analiz. optimal’nye pravila ostanovki [statistical sequantial analiz. optimal stopping rules]. moscow, russia: fizmatlit [in russian]. [20] belov m. & novikov d. (2017). reflexive models of complex activity. proc. of wosc world congress. [21] bolton, p. & dewatripont, m. (2005). contract theory. cambridge: mit press. [22] cvitanic, j. & zhang, j. (2012). contract theory in continuous-time models. heidelberg: springer. models of adaptation in dynamical contracts 131 copyright ©2018 assa. adv. in systems science and appl. (2018) [23] horne, m. (2016). essays on dynamic contract theory. ann arbor: the university of north carolina. [24] ljungqvist, l. & sargent, t. (2004). recursive macroeconomic theory (2nd ed.). cambridge: mit press. [25] knight, f. (1921). risk, uncertainty and profit, in hart, schaffner, and marx prize essays, no. 31. boston and new york: houghton mifflin. [26] menard, c. (2000). institutions, contracts and organizations: perspectives from new institutional economics. northampton: edward elgar pub. [27] renner, p. & schmedders, k. (2016). dynamic principal-agent models, swiss finance institute research paper no. 16–26. zurich: university of zurich. [28] salanie, b. (2005). the economics of contracts (2nd ed.). massachusetts: mit press. [29] sannikov, y. (2008). a continuous-time version of the principal-agent problem, review of economic studies, 75(3), 957–984. [30] stole, l. (1997). lectures on the theory of contracts and organizations. chicago: univ. of chicago. http://restud.oxfordjournals.org/ adv syst sci appl 2018; 03:17–28 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/583 visual odometry approaches to autonomous navigation for multicopter model in virtual indoor environment alexander v. abdulov1, alexander n. abramenkov1 ,∗andrey a. shevlyakov1 1 institute of control sciences of the russian academy of sciences, moscow, russia abstract: using simulators and virtual environments can simplify the development and testing of various control, behavior and navigation scenarios significantly. in this article we describe our experience using the ros environment and the gazebo simulator to develop an autonomous multicopter uav that can navigate an unknown building and locate an exit using computer vision techniques. keywords: virtual simulation, visual odometry, uav, autonomous navigation 1. introduction as more powerful embedded computers appear, along with smaller cameras, autonomous indoor navigation for mobile robots becomes a sufficiently interesting and tractable problem, inciting new research [8, 13]. since global positioning signal is usually unavailable indoors, mobile robots face hard problems of determining their position in reference to some external beacons, route planning and actuating movement. the essential difficulty is not navigation per se, but locating beacons using available sensors. autonomous mobile robots may face additional challenges in dynamic environments. today’s technology allows for relatively easy development of autonomous mobile robots that can navigate using only onboard sensors [1, 9]. however, there are no universal algorithms for navigation in unknown scenes with dynamic obstacles. simulation software exists to facilitate the development and lower the entry threshold of knowledge (e.g. diy robotics). well-designed simulators (see gazebo [11]) make rapid testing possible, as well as robot model building and their deployment in realistic scenarios. in this article we share our experience if solving a problem of autonomous navigation in virtual environment using widespread visual odometry methods. we proceed with the problem and scenario description, which is considered as a test ground for two different approaches. we discuss strong points and limitations of the methods, as well as evaluation of their applicability and overall difficulties of this type of research. ∗corresponding author: aash29@gmail.com 18 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov 2. problem statement as an example of a navigation problem, we chose the following mission: a multicopter of a specified model must autonomously locate a window in an unknown building and leave via this window, without colliding with obstacles or using any kind of external positioning system. onboard sensors include imu (accelerometers and gyroscopes) and computer vision sensors (lidar and video cameras). building in question is an enclosed virtual hangar with one open window (see fig. 2.1). the hangar has a rectangular plan. texture, size and color of walls may vary. obstacles are represented by cylindrical pillars of varying height, diameter and color. pillars are static, but their position is random. they may also be combined to form more complicated structures. the window is rectangular, although its position and dimensions are not known. as the mathematical models of sensors and whole uavs are usually available, we shall not pay too much attention to the hardware. for our research we chose a quadrotor scheme (fig. 2.2) as the most widespread and small enough (less than 0.5 m in dia) to operate indoors. a multitude of virtual sensors can be combined and installed onboard. virtual multicopter movement is controlled by feeding an embedded controller the reference values of linear and angular velocity. the make of embedded controller is fixed, thus we implement the autonomous movement with higher-level pid controllers that have to be tuned separately. fig. 2.1. top view of the virtual hangar fig. 2.2. multicopter model copyright c© 2018 assa. adv syst sci appl (2018) visual odometry approaches to multicopter autonomous navigation 19 3. proposed solution 3.1. technology simulation is run on ubuntu 16.04 os with ros kinetic installed. ros (robot operating system) is a framework for robot programming. it is based on graph architecture, where data processing is performed in the nodes that can receive and transmit messages to other nodes. there are many useful packages for ros, which can also be written from scratch using standard interface. c++ and python are supported. we should mention separately the mavros package that provides the data exchange interface between ros and autopilot firmware using mavlink (micro air vehicle link) protocol, which is designed for uav communication. as a simulator we chose the gazebo software due to the availability of necessary uav and sensor models, as well as ros integration. additionally, we use opencv library for image processing and dlib library for numerical optimization. 3.2. flight strategy high-level strategy for a chosen uav model is described as a finite state machine and contains the following steps (see fig. 3.3a). 1. climb to a given altitude. 2. survey the area. on this step the multicopter rotates (up to 360 degrees). two tasks are accomplished – searching for window and building a local obstacle map. 3. if a window is found, multicopter changes heading to face the window, and next step is performed; otherwise, multicopter flies forward in the direction chosen according to the algorithm and surveys again (step 2). 4. head to the window avoiding obstacles and exit the building. local obstacles are mapped using frontal laser lidar (see fig. 3.3b). for each heading we save the distance to the nearest obstacle and add a safety margin which allows to avoid collisions. a new direction of movement is chosen as the clearest one and closest to the course set. on a lower level, we must control and stabilize the uav model by flight altitude, yaw angle and linear velocity. (a) flowchart of states change (b) obstacles are grey, safety margins are pink fig. 3.3. flight strategy and local obstacles map copyright c© 2018 assa. adv syst sci appl (2018) 20 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov 3.3. angle stabilization orientation estimate obtained via mems gyroscopes can accumulate error due to “zero drift” problem. algorithms exist for fusing different sensors (accelerometers, gyroscopes and magnetometers) to refine estimates. for example, madgwick ahrs [12] algorithm or extended kalman filter [14]. we should note that the model firmware used imu readings to interpret the control signals. in case of yaw angle error this feature did not allow to control the uav correctly and had to be accounted for while specifying control signal. since in case of simulation the magnetometer readings were relatively accurate, they were used to rectify the yaw estimate. 3.4. altitude stabilization it is reasonable to use a downward-facing range-finder (altimeter) to stabilize altitude. an ultrasound range-finder can be used, but a more narrow-bear laser range-finder is preferable. additionally, orientation of the aircraft must be accounted for. on the basis of this estimate, it is relatively simple to tune the coefficients for a pid controller which can adequately stabilize the reference altitude (see fig. 3.4). here and below, we denote: target – reference value; real – ground truth, obtained from simulator; pid – sensor estimate, fed to the pid controller. 3.5. velocity stabilization in gps-denied indoor spaces we need to use other information to obtain current velocity estimate. in principle, imu can be used, but due to large errors it needs to be amended by other means. visual odometry is a promising approach [3]. a sequence of images captured by camera can be a main or auxiliary source of position information. depending on number of cameras and their configuration, various methods can be used. we’ll consider two approaches, mono and stereo visual odometry, that necessitate installing one or two calibrated cameras respectively. fig. 3.4. flight altitude stabilization by pid copyright c© 2018 assa. adv syst sci appl (2018) visual odometry approaches to multicopter autonomous navigation 21 fig. 3.5. perspective projection camera model 3.6. camera model visual odometry operates on a basis of a camera model. often this is a relatively simple perspective projection model (see fig. 3.5), represented by a matrix (3.1) of intrinsic parameters (focal length and principal point). in more complicated cases, distortion transformations are also accounted for. c = ∣∣∣∣∣∣ fx 0 cx 0 fy cy 0 0 1 ∣∣∣∣∣∣ (3.1) usually the intrinsic parameters are determined by camera calibration. for a virtual camera we chose image center as the principal point (cx, cy) and obtain focal length f = fx = fy from equation (3.2), knowing resolution s and field of view (fov) angle α. α = 2 · arctan s 2 · f (3.2) equation (3.3) gives us the relation of point (x, y, z) in 3d scene and its projection (u, v) on image plane up to a scale coefficient w. this is a composition of rotation matrix r and translation vector t . w · ∣∣∣∣∣∣ u v 1 ∣∣∣∣∣∣ = c · |r t | · ∣∣∣∣∣∣∣∣ x y z 1 ∣∣∣∣∣∣∣∣ (3.3) 3.7. monocular visual odometry consider a case of one camera, laser altimeter and 6-axis imu. if a camera is downward-facing, then we can assume that most of the observed feature points belong to one plane. this is a reasonable assumption if the flight altitude is low, and it was almost always valid in our simulation. knowing altitude h and camera orientation quaternion qimu , we can compute floor plane equation a · x+ b · y + c · z + d = 0 coefficients, as indicated in (3.4). first three coefficients are obtained by rotation operation ◦ a unit vector pointing along an axis z. (a, b, c) = q−1 imu ◦ (0, 0, 1); d = −h · c (3.4) copyright c© 2018 assa. adv syst sci appl (2018) 22 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov (a) keypoints optical flow with homographic filtering between two consequent images of floor; matched points in green, filtered points in red (b) keypoints flow projection on the floor plane fig. 3.6. monocular odometry approach to velocity estimation taking coordinates (u, v) of floor image feature points and substituting (3.5) in (3.6), we reconstruct their position (x, y, z) in camera frame. since frame used in aircraft firmware does not match the camera frame, some axes and signs are inverted in equations for x and y. û = u− cx fx ; v̂ = v − cy fy ; (3.5) z = d a · û+ b · v̂ + c ; x = −z · v̂; y = −z · û (3.6) as a result, linear velocity estimate is carried out using feature keypoints position (fast [5], good features to track [7]) from to consequent measurements. correspondence between feature points can be found, for example, by optical flow or by comparing descriptors [4]. we used a well-established lucas-kanade algorithm [2], due to its modest computational requirements and adequate performance in case of small movements. the implementation of this algorithm is available in opencv library. it is useful to apply additional filters to reject outliers: mean velocity, direction, etc. an assumption of all points belonging to one plane allows us to successfully use the filter based on homography matrix (see fig. 3.6a), obtained on ransac [10] method. this approach is schematically depicted on fig. 3.6b. in this case, multicopter speed is towards negative direction of axis y . thus, projections of feature points are moving in the opposite direction. red line is normal distance to the floor, blue line is its estimate by altimeter and imu readings, magenta plane is the observed floor plane, cyan markers are feature points and their movement. in general, monocular approach allows us to estimate velocity while flying at low altitude (less than 1 m), and is computationally inexpensive. flying higher leads to poor performance. this is due to two factors: 1. aircraft rotation at higher altitude leads to more displacement for optical flow to handle, which leads to more errors and outliers. 2. flying higher increases the probability of obstacles in view field, which invalidates the assumption of all points belonging to the floor plane. here the task of odometry is not fully solved. only velocity is estimated, and then stabilized by a suitably tuned pid controller, without trajectory construction. copyright c© 2018 assa. adv syst sci appl (2018) visual odometry approaches to multicopter autonomous navigation 23 fig. 3.7. velocity stabilization via monocular odometry approach therefore, the accuracy of this method makes it possible to successfully use it. speed estimates are adequate, as can be seen on fig. 3.7. stabilization time is acceptable, although overshooting is present. 3.8. stereo visual odometry stereo visual odometry is an alternative approach to estimating velocity. in this work we used a stereoscopic method based on reprojection error minimization [6]. it consists of several steps: 1. for each feature point p 3d need to find correspondence between its projections on current left and right frames of a stereo camera (pl t, p r t ) and on one of the previous frames (pl t−1). this can be seen in fig. 3.8. as in monocular case, we use lucas-kanade optical flow. 2. using the established correspondence between the projections of all feature points on the left (ul, vl) and right (ur, vr) frames of a stereo camera with base b, the triangulation problem (3.7) is solved to restore their position (x, y, z) in 3d space (see fig. 3.9). (x, y, z) = b ul − ur · (ul − cx, vl − cy, f) (3.7) 3. knowing current position on each feature point p3d i,t in 3d space, and its projection pl i,t−1 at a previous time step, linear displacement t and rotation r are estimated by minimizing (3.8) reprojection error. argmin r,t ∑ i ‖pl i,t−1 − ĉ(r · p3d i,t + t )‖ (3.8) equation (3.8) contains a projection operator ĉ that maps a point from 3d space to the image plane according (3.3). we used the euclidian distance ‖·‖ as a norm. this approach is more versatile, and does not demand additional hardware apart from two properly fixed cameras (downward or forward facing). the main limitation is the necessity of multiple observed feature points. their allowed distance depends on stereo camera base. copyright c© 2018 assa. adv syst sci appl (2018) 24 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov fig. 3.8. schematic representation of 3d point projection on stereo camera matrix during movement. cl t, cr t – left and right cameras at time t. cl t−1, cr t−1 – left and right cameras at time t− 1 fig. 3.9. triangulating p for a stereo camera with base b stereo visual odometry approach is also more computationally expensive – for the same video parameters (648x486, 36fps) processing was 2-3 times longer. in our case, monocular odometry allowed for real time computation, while stereo odometry had to skip a half of available frames. it is possible, however, to decrease the precision in order to speed up the computations. in the examples considered, forward facing camera did not always perform well due to lack of texture on pillars which yielded few feature points. downward facing camera was more stable, although there was a minimum necessary altitude that allowed both cameras to see the same objects. copyright c© 2018 assa. adv syst sci appl (2018) visual odometry approaches to multicopter autonomous navigation 25 fig. 3.10. velocity stabilization via stereo odometry approach fig. 3.11. errors of calculated velocities for monocular and stereo approaches, error = |pid −real| figure 3.10 shows the velocity stabilized with tuned pid controllers via the estimate obtained by this approach. velocity estimation is somewhat better than in monocular case – the plots of real and pid nearly match (see fig. 3.11). 4. search for window to detect a window, we used color segmentation. assuming that the background visible through the window was a different color than the rest of the scene, we needed to find a suitably colored shape (see fig. 4.12a). furthermore, shape filtering, e.g. rectangle detection (see fig. 4.12b), can be used. at a last stage, when the aircraft approaches a window, it can become disoriented as number of feature points decreases or they disappear entirely. in this case, we solve a ballistics problem – aim for the window and fix a exit speed, disabling regular velocity controllers. copyright c© 2018 assa. adv syst sci appl (2018) 26 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov (a) color segmentation and threshold binarization (b) specifically-colored rectangle searching fig. 4.12. searching for window (a) simple (b) difficult fig. 4.13. examples of the scenes 5. simulation results as a result of several simulation trials, we established that the described approach solves a problem of exiting the building through the window. two scenes of different layouts were explored. simple scene dimensions are 40x50x10 m (see fig.4.13a). obstacles are represented by cylindrical pillars of different radii. aircraft could detect the window directly from the starting point and approached it while avoiding the obstacles. difficult scene dimensions are 50x90x15 m, obstacles are simple pillars and composite objects of several different-sized cylinders, imitating trees. (see fig. 2.1, 4.13b). in this case, obstacles blocked the window completely, thus search strategy had to be activated. the parameters chosen resulted in a fairly conservative behavior – there was no collisions, but search required a sufficiently long time. thus, a relatively simple search algorithm, that didn’t map the scene or localize of the aircraft, could accomplish the task of exiting the building. 6. technical difficulties we encountered some complications while using the gazebo simulator, which should be accounted for. notably, high computational requirements, which resulted in unstable work. sometimes, a kind of “numerical wind” was apparent, a seemingly random perturbations that forced the aircraft to crash. copyright c© 2018 assa. adv syst sci appl (2018) visual odometry approaches to multicopter autonomous navigation 27 the effect could be reproduced by running both client and server parts of the simulator on the same medium-powered computer. for all our experiments we used an intel core i7-4510u cpu @ 2.00ghz, 8gb ram pc. additional disturbances were created by other applications running in the background. practically, gazebo is not the most comfortable software to use in case of multiple test runs. each run required restarting the whole virtual environment and took a lot of time, which combined with the “numerical wind” made our work more difficult – especially the tuning of pid controllers. we encountered various other difficulties, including the necessity to convert between every coordinate frame involved (world, aircraft, sensor). this, however, is anticipated by mavlink protocol which offers the possibility of specifying control in several coordinate frames (ned, enu, etc.) 7. results numeric simulation using the ros framework, the gazebo simulator and the rest of the corresponding ecosystem allowed us to test the applicability of the chosen algorithms in specified conditions. use of this class of simulators and frameworks can be of much help while developing real hardware and software platforms for uavs. references [1] a. weinstein, a. cho, g. loianno, & v. kumar. (2018). visual inertial odometry swarm: an autonomous swarm of vision-based quadrotors. ieee robotics and automation letters, 3 (3), 1801–1807. [2] b.d. lucas, & t. kanade. (1981). an iterative image registration technique with an application to stereo vision. international joint conference on artificial intelligence, 647–679. [3] c. cadena, l. carlone, h. carrillo, y. latif, d. scaramuzza, j. neira, . . . j.j. leonard. (2016). past, present, and future of simultaneous localization and mapping: towards the robust-perception age. ieee transactions on robotics, 32 (6), 1309–1332. [4] d. bekele, m. teutsch, & t. schuchert. (2013). evaluation of binary keypoint descriptors. in 2013 ieee international conference on image processing (pp. 3652–3656). doi:10.1109/icip.2013.6738753 [5] e. rosten, & t. drummond. (2006). machine learning for high speed corner detection. 9th european conference on computer vision, 1, 430–443. [6] h. badino, & t. kanade. (2011). a head-wearable short-baseline stereo system for the simultaneous estimation of structure and motion. iapr conference on machine vision applications (mva),nara, japan. [7] j. shi, & c. tomasi. (1994). good features to track. proceedings of ieee conference on computer vision and pattern recognition, 593–600. [8] k. mo, h. li, zh. lin, & j. lee. (2018). the adobeindoornav dataset: towards deep reinforcement learning based real-world indoor robot visual navigation. arxiv preprint arxiv:1802.08824. [9] l. silvestri, l. pallottino, & s. nardi. (2017). design of an indoor autonomous robot navigation system for unknown environments. in international conference on modelling and simulation for autonomous systems (pp. 153–169). springer. [10] m.a. fischler, & r.c. bolles. (1981). random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. commun. acm, 24 (6), 381–395. doi:10.1145/358669.358692 copyright c© 2018 assa. adv syst sci appl (2018) https://dx.doi.org/10.1109/icip.2013.6738753 https://dx.doi.org/10.1145/358669.358692 28 a.v. abdulov, a.n. abramenkov, a.a. shevlyakov [11] n. koenig, & a. howard. (2004). design and use paradigms for gazebo, an open-source multi-robot simulator. proceedings of 2004 ieee/rsj international conference on intelligent robots and systems, 2149–2154. [12] s. madgwick. (2010). an efficient orientation filter for inertial and inertial/magnetic sensor arrays. retrieved from http : //x io . co .uk/ res/doc/ madgwick internal report.pdf [13] sh. yang, s.a. scherer, x. yi, & a. zell. (2017). multi-camera visual slam for autonomous navigation of micro aerial vehicles. robotics and autonomous systems, 93, 116–134. [14] x. li, m. chen, & l. zhang. (2016). quaternion-based robust extended kalman filter for attitude estimation of micro quadrotors using low-cost mems. control conference (ccc), 2016 35th chinese, 10712–10717. copyright c© 2018 assa. adv syst sci appl (2018) http://x-io.co.uk/res/doc/madgwick_internal_report.pdf http://x-io.co.uk/res/doc/madgwick_internal_report.pdf introduction problem statement proposed solution technology flight strategy angle stabilization altitude stabilization velocity stabilization camera model monocular visual odometry stereo visual odometry search for window simulation results technical difficulties results adv syst sci appl 2018; 03; 79-89 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/653 simulation modelling as a tool for design and development in large-scale automated systems: smart city application in terms of lack of statistical information dmitri razumov1*, prof., dr. (eng.) vladimir aleshin2 1) dept. school of it management russian academy of national economy (ranepa) moscow, russia e-mail: dmitrirazumov@yandex.ru 2) dept. school of it management russian academy of national economy (ranepa) moscow, russia e-mail: aleshin_vladimir@mail.ru abstract: design and development of large-scale automated systems (as) is always associated with overcoming a large number of uncertainties, which structuring often through the using of practice standards and modeling at the systems & software engineering. methods of analysis of large systems (ls) involve their decomposition in the context of structural and/or functional paradigms, which allow us to consider the obtained components as objects for modelling: functional modelling, simulation etc. design the automated control system (acs) of region (smartcity project) need to optimize the decisions. in this paper we propose the new methods of determining the parameters of project in terms of lack statistical information based on the simulation model. keywords: acs automated control system, life cycle, modeling, smart city. 1. introduction already at the initial stages of systems engineering, different models and methods of mathematical science are used. these are: mathematical statistics, discrete mathematics, game theory etc. advantages in area of information technology has allowed to bring to the practice of creating ls such powerful tools as simulation. in many cases simulation is the only one way to get any notion of the behavior of a complex system and to conduct its analysis. the acs in the crisis and emergency situations is a big complex system (bs), in the nodes of which are located decision-making centers, related with the aid of information flows with operating forces and resources from the state and municipal institutions. the complexity is due to the fact that the organization of such systems does not have a straightforward structural implementation and is determined by many factors (availability of communication channels, level of automation, duty services, interaction between departments, the level of financial support, etc.). therefore the design of such bs is connected with the decomposition of the general problem at the level of paradigm life cycle (lc) [1] in the context of structural-functional component hierarchy, which have a nodes (fig. 1) as the duty dispatcher service (dds) or command control centers(ccc). the problem of simulation of the overall structure of the system is the lack of sufficient statistical * corresponding author: dmitrirazumov@yandex.ru 80 d. razumov, v. aleshin copyright ©2018 assa adv. in systems science and appl. (2018) information for all dds. the authors propose a method, which ensures a sufficiently adequate model in these limits. 2. functional and structural analysis for designing acs of region the function "to manage the life cycle stage" on the stage of the lifecycle "design” solves a non-trivial problem of building a hierarchy of command control centers and decision-support centers. the situation with organization of this structure differs from region to region. as a rule the hierarchy of divisions in the area is formed from the newly created command control centers (ccc) at the level of municipal districts from the unified duty dispatcher service (udds) and existing ccc (utilities and other emergency services, fig. 1). fig. 1. structural analysis of system often due to lack of technical capacity information can be transmitted not to the main samples but in any of the duty services. the task of designing the structures contains many elements of uncertainty. the requirements of operators, resources, emergency response number, power and performance of computer systems and throughput of communication channels in the project document do not have a clear quantitative and scientific basis and developed at best on the basis of comparative assessments of the operational histories of the existing call centers (cc). however it is not a sufficient basis for parameters definition of the designed system. underestimation of requirements for operators, resources, the quantity and quality of technical means cannot lead to lower response time, but rather its can increase it. such approaches may lead to waste of millions of dollars during the construction of such systems. 3. task definition for simulation modelling we need to offer a science-based method of assessment requirements for the resources ccc/udds, guaranteeing the fulfillment of key performance indicators (kpi) the system functioning in framework the main operational process, under limited amount of statistically reliable data and uncertain structure that represents a hierarchy of decision-support centers, communicate links and the management links. simulation modelling as a tool for design and development 81 copyright ©2018 assa. adv. in systems science and appl. (2018) 3.1. building the base model class of ccc as already noted, ccc is defined as the basic element of the system in context of structuralfunctional decomposition. the ccc model is used also to model the main operational process which uses a discrete event simulation model on the basis of libraries of standard classes of the system modeling anylogic. anylogic allows presenting the ccc functionality in an object-oriented paradigm. discrete event simulation model on the basis of libraries of standard classes anylogic is used for modelling the main process, which implements the basic blocks at the level of mechanisms described in the theory of mass service: event sources, the queue, processing servers, etc. you can pick up an optimal set of characteristics for the compilation of parameters for the specifications a particular system by manipulating the parameters of the objects based on the implementation of these classes, as well as analyzing the statistical dependencies. the designed system, as already noted, is a hierarchical structure of control centers interacting with each other to coordinate control of forces and means of response. 4. methods for evaluation of the stochastic characteristics of the model as the main stochastic characteristics of the basic simulation model of the considered functional unit (ccc) allocated: • intensity of input flow requests; • the length of the intervals between user requests of call centers automated dispatch control system (adcs) one of the ccc; • the temporal characteristics of requests processing; • actions of emergency response services. for each of these parameters were applied a method involving the following steps: • pre-selection of a family of theoretical distribution; • estimation of parameters (position, scale, and shape) of distributions; • testing of hypotheses about the correctness of the choice one; • verification of the results obtained by comparing the estimated and real data, using statistics on one of the nodes (ccc). 4.1. estimation the intensity of input flow request at the fig. 2 the first result obtained in the evaluation of the intensity the input flow of requests ccc. it needs the separation the original time series from non-random components. fig. 2. lengths of the intervals between the requests ones of the ccc 82 d. razumov, v. aleshin copyright ©2018 assa adv. in systems science and appl. (2018) at first we accept for an initial hypothesis about the probability distribution of the lengths of the intervals between the receipts of requests in the ccc in accordance with the theoretical recommendations. this assumption is the exponential character of this distribution. the program package statistica was used for the test and analyze these hypotheses based on the data for 2013 year. the result being analyzed with the kolmogorov-smirnov criterion histogram of distribution density of time intervals between receipts of bids (fig. 3). fig. 3. the gistogramm analyze in the programm statistica thus we obtain the average value of the intervals between receipt of bids is equal to 7.8 min., the dispersion of the intervals of the arrivals of 10.12, the value of the kolmogorov-smirnov 0,05335. hypothesis of an exponential distribution is confirmed. 4.2. estimation the temporal characteristics request processing the gamma distribution is the most "suspicious" law from the perspective of better applicability for estimating parameters of a stochastic variable of type "duration of service". scale parameter b and shape parameter a of this distribution are connected by a system of equations. simulation modelling as a tool for design and development 83 copyright ©2018 assa. adv. in systems science and appl. (2018) fig.4. the histogram analysis of duration time bids service (4.1), thus m(x) – expected value ( ); d(x) – distribution (s2); the solution of system equations (4.1) allows you to determine the most plausible values of the parameters of the gamma distribution. accordingly, the parameter estimates have the form: , (4.2) the data for 2013 and the criterion of kolmogorov-smirnov have been used to test the hypothesis about the applicability of the gamma distribution for estimating the processing time of applications in the executive division of ccc (fig.4). when we checked these parameters on the model (fig.5) we get the deviation from the actual number of filings for the same period in 2014 year of no more than 12%. fig.5 clearly shows that the model graphs of the probability density the service requests and interval requests are almost identical with the statistics (fig.3, fig.4). 5. development the base ccc model on the data of the input streams queueing system modeling each of the call centers should be provide for: classification of bids according degree of service complexity (which service responds, how many services respond etc.); possibility of information exchange according applications between ccc, both vertically and horizontally (fig.1). 84 d. razumov, v. aleshin copyright ©2018 assa adv. in systems science and appl. (2018) fig. 5. the test of ccc simulation model 5.1. partitioning the original input stream requests on the substreams in order to be able to draw conclusions about the all participating ccc, for which we have no statistics, and the capacity of inter-node relationships, you have to divide the available input stream on certain types: • 01 – only required response of forces and means of dds 01, • 02 – ccc 02, • 03 – ccc 03, • 0103 – ccc 01 and 03, respectively etc. on this base being made the conclusion about the streams of 01, 03 and then created the simulation model of common ccc. thus "traces" of the joint activity is possible to estimate the parameters of the "missing statistics". fig. 6 shows the identification of distribution laws for each of the "substream". way above is also the identity of the distribution laws for the service interval for each of the selected types substream. the experiment confirms that the splitting of the flow of requests into classes with the corresponding parameters of the distribution does not change the parameters of the distribution of initial flow. thus the basic universal template (class complexsource) simulates a typical ccc. simulation modelling as a tool for design and development 85 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 6. identification of distribution laws for each of the "substream" fig. 7. the common universal simulation model integrated dds/ccc with distributed input streams in terms of anylogic 5.2. development the base simulation model ccc standard model of ccc (base class) with split input and output, generating a request to ccc 01, 02, 03 (fig. 7), allows testing the hypothesis that the model command control center by splitting the input stream will be adequate in terms of the formation the final steam model without splitting (fig. 8). 86 d. razumov, v. aleshin copyright ©2018 assa adv. in systems science and appl. (2018) fig.8. experiment 6. development the model of regional government system based on the base (universal) model for local ccc as node objects being used the entities which classes are inherited from the base class of the model local ccc/dds (fig.7). fig. 9. graphic interpretation the common simulation model on base universal ccc(complexsource) class simulation modelling as a tool for design and development 87 copyright ©2018 assa. adv. in systems science and appl. (2018) fig. 10. the parameters of local model ccc 03. fig. 11. parameters of the model ccc 02 service in figures 10 and 11 it is seen that the input streams of an object that models behavior of some local ccc 03 differ from the object of the same class modeling of local ccc 02. the specificity of 03 is reflected by the fact that is dominated by flows associated with the receipt of events including type 03 (for more details see above) which is reflected by the intensity of the exponential distribution. streams which do not fit the specifics of ccc 03 are ignored because their intensity is negligible. such techniques are used to describe the specifics of 02 ccc too (see the parameters of the object complexsource_02 fig. 9). here you can see that are significant the streams that have type of applications including reaction units 02. the highest intensity has a flow of "clean" applications 02. graphics system allows you to choose the optimal resource values that manipulating it is easy to identify the processes mutual influence in the work of each of the nodes on other subsystems (fig.12 ). 88 d. razumov, v. aleshin copyright ©2018 assa adv. in systems science and appl. (2018) fig. 12. the identification of optimal resources for each local ccc through simulation model. 5. conclusion thus the authors proposed the formulation of the problem, which served as the basis for the construction of a simulation model. this allows you to determine the required amount of smartcity system resources to perform key performance indicators within the existing constraints. the presented material can be useful for designers and developers of large automated systems which include the acs of the smartcity, as well as the acs of the regional and municipal level, large sports events of international and state scale, situation centers of regional and municipal governments including such a specific area as public authorities and management. the following articles will discuss the problems of system design from the point of view of functional modeling in the life cycle including obtaining an adequate structure of the model as well as approaches related to multivariate analysis. acknowledgements we thank the members of the organizational committee of mlsd’2017 for their kind invitation to contribute the paper presented at the conference to the journal and andrey shevlyakov, the secretary of advances in systems science and applications, for technical assistance. references [1] blanchard, benjamin s. (2004) system engineering management. john wiley & sons. [2] gost r iso/mehk 12207 (1999) «informacionnaya tekhnologiya. processy zhiznennogo cikla programmnyh sredstv», identichnyj [iso/iec 12207(1995) system and software engineering — software life cycle processes]. simulation modelling as a tool for design and development 89 copyright ©2018 assa. adv. in systems science and appl. (2018) [3] malyshev v.v. (2010). metody optimizacii v zadachah sistemnogo analiza i upravleniya. [optimization methods in system analysis and management problems] moscow: maiprint [in russian]. [4] razumov d.a., aljoshin v.d. (2014). modelirovanie v zhiznennom cikle avtomatizirovannyh sistem upravleniya v krizisnyh i chrezvychajnyh situaciyah. prikladnaya informatika. [modeling in the life cycle of automated control systems in crisis and emergency situations. applied informatics] 6(54), 102-116. [in russian]. advances in systems science and applications (2013) vol.13 no.4 355-368 technique of constructing predictive inferences for future outcomes with applications to inventory management problems konstantin n. nechval1 and nicholas a. nechval2 1applied mathematics department, transport and telecommunication institute lomonosov street 1, lv-1019, riga, latvia 2statistics department, evf research institute, university of latvia raina blvd 19, lv-1050, riga, latvia abstract predictive inferences (predictive distributions, prediction and tolerance limits) for future outcomes on the basis of the past and present knowledge represent a fundamental problem of statistics, arising in many contexts and producing varied solutions. in this paper, new-sample prediction based on a previous sample (i.e., when for predicting the future outcomes in a new sample there are available the observed data only from a previous sample), within-sample prediction based on the early data from a current experiment (i.e., when for predicting the future outcomes in a sample there are available the early data only from that sample), and new-within-sample prediction based on both the early data from that sample and the data from a previous sample (i.e., when for predicting the future outcomes in a new sample there are available both the early data from that sample and the data from a previous sample) are considered. it is assumed that only the functional form of the underlying distributions is specified, but some or all of its parameters are unspecified. in such cases ancillary statistics and pivotal quantities, whose distribution does not depend on the unknown parameters, are used. in order to construct predictive inferences for future outcomes, the invariant embedding technique representing the exact pivotal-based method is proposed. furthermore, this technique can be used for optimization of inventory management problems. an illustrative example is given. keywords future outcomes, inventory management, predictive inferences 1 introduction prediction of future outcomes based on the past and current data is the most prevalent form of statistical inference. predictive inferences for future outcomes are widely used in risk management, finance, insurance, economics, hydrology, material sciences, telecommunications, and many other industries.practical problems often require the computation of predictive distributions, prediction and tolerance limits for future values of random quantities. consider the following examples : 1) a consumer purchasing a refrigerator would like to have a lower limit for the failure time of the unit to be purchased (with less interest in distribution of the population of units purchased by other consumers) ; 2) financial managers in 356 konstantin n. nechval : technique of constructing predictive inferences for future... manufacturing companies need upper prediction limits on future warranty costs. a large number of problems in inventory management, production planning and scheduling, location, transportation, finance, and engineering design require that decisions be made in the presence of uncertainty. most of the inventory management literature assumes that demand distributions are specified explicitly. however, in many practical situations, the true demand distributions are not known, and the only information available may be a time-series of historic demand data. when the demand distribution is unknown, one may either use a parametric approach (where it is assumed that the demand distribution belongs to a parametric family of distributions) or a non-parametric approach (where no assumption regarding the parametric form of the unknown demand distribution is made). under the parametric approach, one may choose to estimate the unknown parameters or choose a prior distribution for the unknown parameters and apply the bayesian approach to incorporating the demand data available. scarf [1] and karlin [2] consider a bayesian framework for the unknown demand distribution. specifically, assuming that the demand distribution belongs to the family of exponential distributions, the demand process is characterized by the prior distribution on the unknown parameter. further extension of this approach is presented in [3]. parameter estimation is considered in [4]. liyanage and shanthikumar [5] propose the concept of operational statistics and apply it to a single period newsvendor inventory control problem.in this paper we consider the case, where it is known that the demand distribution function belongs to a parametric family of distribution functions. however, unlike in the bayesian approach, we do not assume any prior knowledge on the parameter values. conceptually, it is useful to distinguish between “new-sample” prediction, “within-sample” prediction, and “new-within-sample” prediction. for new-sample prediction, data from a past sample are used to make predictions of future outcomes from the same process or population. for within-sample prediction, the problem is to predict future outcomes in a sample or process based on early data from that sample or process. for new-within-sample prediction, the problem is to predict future outcomes in a sample or process based on early data from that sample or process as well as on a past sample data from the same process or population. various solutions have been proposed for the prediction problem, that is, the problem of making inferences on a random sample {yj ; j = 1, ...,m} given independent observations {xi; i = 1, ..., n} drawn from the same distribution. the y ′ j s and the x ′ is are commonly featured as “future outcomes” and “past outcomes” respectively. inferences usually bear on some reduction z of the y ′ j s-possibly a minimal sufficient statistic and consist of either prediction intervals or likelihood or predictive distribution for z, depending on different authors. kaminsky and nelson [6] discussed point and interval prediction of order statistics. best linear unbiased predictors advances in systems science and applications (2013) vol.13 no.4 357 based on location-scale family of distributions are reviewed. prediction intervals based on such predictors as well as those based on pivotals are studied. a brief discussion of bayesian prediction is also given. predictive distributions are found in the bayesian framework (see aitchison and sculthorpe [7]). lawless [8] applied the conditional method, which was first suggested by fisher [9] and promoted further by a number of others (nechval et al. [10] ; murthy et al. [11]), to different problems relating to the weibull and extreme value distributions. in practice the proposed methods have limited applications and it is the purpose of this paper to obtain predictive inferences concerning z via the simple invariant embedding technique [12-16]. the obtained results are given below. 2 constructing predictive inferences for future outcomes let us assume that the random variable x follows the exponential distribution with the probability density function fσ(x) = σ−1 exp(−x/σ), σ > 0, (1) and the cumulative distribution function fσ(x) = 1− exp(−x/σ), (2) where σ is the scale parameter (σ >0). 2.1 new-sample prediction theorem 1. let x1 ≤ ... ≤ xk be the first k ordered observations from the previous sample of size n from the exponential distribution (1) and yr be the rth order statistic in a set of m future ordered observations y1 ≤ ... ≤ ym also from the distribution (1).then the probability distribution functionof the ancillary statistic yr/sk is given by pr { yr sk ≤ η } = m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i 1 [1 + (i+m− j)η]k , (3) where sk = ∑k i=1xi + (n− k)xk. proof. the joint density of x1 ≤ ... ≤ xk is given by fσ(x1, ...,xk) = n! (n− k)! k∏ i=1 1 σ exp ( −xi σ ) exp ( −(n− k) xk σ ) (4) it is known that sk = k∑ i=1 xi + (n− k)xk (5) 358 konstantin n. nechval : technique of constructing predictive inferences for future... is the sufficient statistic for σ. then vk = sk/σ (6) is the pivotal quantity, the probability density function of which is given by f(vk) = 1 γ(k) vk−1 k exp(−vk), vk ≥ 0 (7) using the invariant embedding technique [12-16], we reduce pr {yr ≤ yr} = m∑ j=r ( m j ) [fσ(yr)] j [1− fσ(yr)] m−j = m∑ j=r ( m j )[ 1− exp ( −yr σ )]j[ exp ( −yr σ )]m−j (8) to pr { yr sk ≤ η|vk } = m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i exp (−(i+m− j)ηvk) (9) where η = yr/sk (10) it follows from (9) that pr { yr sk ≤ η } = e  m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i exp (−(i+m− j)ηvk)  (11) = ∫ ∞ 0 m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i exp (−(i+m− j)ηvk) f(vk)dvk = m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i 1 [1 + (i+m− j)η]k (12) this ends the proof. corollary 1.1. a lower one-sided new-sample α prediction limit h on the rth order statistic yr in a set of m future ordered observations y1 ≤ ... ≤ ym is given by h = ηsk, (13) advances in systems science and applications (2013) vol.13 no.4 359 where η satisfies the equation m∑ j=r ( m j ) j∑ i=0 ( j i ) (−1)i 1 [1 + (i+m− j)η]k = α. (14) (observe that an upperone-sided 1− α prediction limit h may be obtained from a lower one-sided α prediction limit by replacing α by 1− α.) 2.2 within-sample prediction theorem 2. let y1 ≤ ... ≤ yl be the first lordered observations (order statistics) in a sample of size m from a continuous distribution with some probability density function fθ(x) and distribution function fθ(x), where θ is a parameter (in general, vector). then the joint probability density function of y1 ≤ ... ≤ yl and the rth order statistics yr(1 ≤ l < r ≤ m) is given by gθ(y1, ..., yl, yr) = gθ(y1, ..., yl)gθ(yr|yl), (15) where gθ(y1, ..., yl) = m! (m− l)! l∏ i=1 fθ(yi)[1− fθ(yl)] m−l, (16) gθ(yr|yl) = (m− l)! (r − l − 1)!(m− r)! [ fθ(yr)− fθ(yl) 1− fθ(yl) ]r−l−1[ 1− fθ(yr)− fθ(yl) 1− fθ(yl) ]m−r · fθ(yr) 1− fθ(yl) = (m− l)! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) (−1)j [ 1− fθ(yr) 1− fθ(yl) ]m−r+j fθ(yr) 1− fθ(yl) (17) represents the conditional probability density function of yr given yl = yl. proof. the joint density of y1 ≤ ... ≤ yl and yr is given by gθ(y1, ...,yl, yr) = m! (r − l − 1)!(m− r)! l∏ i=1 fθ(yi)[fθ(yr)− fθ(yl)] r−l−1· fθ(yr)[1− fθ(yr)] m−r = gθ(y1, ..., yl)gθ(yr|yl) (18) it follows from (18) that gθ(yr|y1, ..., yl) = gθ(y1, ...,yl, yr)/gθ(y1, ..., yl) = gθ(yr|yl), (19) 360 konstantin n. nechval : technique of constructing predictive inferences for future... i.e., the conditional distribution of yr, given yi = yi for all i = 1,...,l, is the same as the conditional distribution of yr, given only yl = yl, which is given by (17). this ends the proof. corollary 2.1. the conditional probability distribution function of yr given yl = yl is pθ {yr ≤ yr|yl = yl} = 1− (m− l)! (r − l − 1)!(m− r)! · r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [ 1− fθ(yr) 1− fθ(yl) ]m−r+1+j (20) corollary 2.2. let y1 ≤ ... ≤ yl be the first lorder statistics in a sample of size mfrom the exponential distribution (1). then the conditional probability distribution function of yr given yl = yl is pθ {yr ≤ yr|yl = yl} = 1− (m− l)! (r − l − 1)!(m− r)! · r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [ exp ( −yr − yl σ )]m−r+1+j (21) theorem 3. let y1 ≤ ... ≤ yl be the first lordered observations (order statistics) in a sample of size m from the exponential distribution (1), where the parameter α is unknown.then the probability distribution function of the ancillary statistic yr/yl is given by pr { yr yl ≤ yr yl } = 1− m! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j · (22)( l−1∏ i=0 [( yr yl − 1 ) (m− r + 1 + j) +m− l + 1 + i ])−1 proof. we reduce (20) to pr{yr ≤ yr |yl = yl} = 1− (m− l)! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) · (−1)j m− r + 1 + j [ exp ( −yl σ ( yr yl − 1 ))]m−r+1+j , (23) where yr/yl is the ancillary statistic whose distribution does not depend on the parameter σ, yl/σ is the pivotal quantity. using the probability density function advances in systems science and applications (2013) vol.13 no.4 361 of yl/σ, we eliminate the unknown parameter σ from the problem as pr { yr yl ≤ yr yl } = ∫ ∞ 0 pr{yr ≤ yr |yl = yl}g(yl/σ)dyl/σ = 1− m! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j · ( l−1∏ i=0 [( yr yl − 1 ) (m− r + 1 + j) +m− l + 1 + i ])−1 (24) where g(yl/σ) = m! (l − 1)!(m− l)! [ 1− exp ( −yl σ )]l−1 exp ( −yl σ (m− l) ) exp ( −yl σ ) = m! (l − 1)!(m− l)! l−1∑ i=0 ( l − 1 i ) (−1)i exp ( −yl σ (m− l + 1 + i) ) , (yl/σ) ∈ (0,∞) (25) represents the probability density function of the pivotal quantity yl/σ. this ends the proof. corollary 3.1. a lower one-sided within-sample α prediction limit h on the rth order statistic yr in a set of m future ordered observations y1 ≤ ... ≤ ym is given by h = ηyl, (26) where η satisfies the equation 1− m! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j( l−1∏ i=0 [(η − 1)(m− r + 1 + j) +m− l + 1 + i] )−1 = α. (27) theorem 4. let y1 ≤ ... ≤ yl be the first lordered observations (order statistics) in a sample of size m from the exponential distribution (1), where the parameter α is unknown.then the probability distribution function of the ancillary statistic (yr − yl)/sl is given by pr { yr − yl sl ≤ yr − yl sl } = 1− 1 b(r − l, (m− r + 1) r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [ 1 + (m− r + 1 + j) yr − yl sl ]−l (28) 362 konstantin n. nechval : technique of constructing predictive inferences for future... where sl = l∑ i=1 xi + (m− l)yl (29) is the sufficient statistic for σ. proof. using the technique of invariant embedding [12-16], we reduce (20) to pr {yr ≤ yr|yl = yl} = 1− (m− l)! (r − l − 1)!(m− r)! r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [ exp ( −sl σ yr − yl sl )]m−r+1+j (30) where (yr−yl)/sl is the ancillary statistic whose distribution does not depend on the parameter σ ; sl/σ is the pivotal quantity. since the probability density function of sl/σ is known,we eliminate the unknown parameter σ from the problem as pr { yr − yl sl ≤ yr − yl sl } = ∫ ∞ 0 pr{yr ≤ yr |yl = yl}g(sl/σ)dsk/σ = 1− 1 b(r − l,m− r + 1) r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [ 1 + (m− r + 1 + j) yr − yl sl ]−l (31) where f(sl/σ) = (γ(l))−1(sl/σ) l−1 exp(−sl/σ), (sl/σ) ≥ 0. (32) this ends the proof. corollary 4.1. a lower one-sided within-sample α prediction limit h on the rth order statistic yr in a set of m future ordered observations y1 ≤ ... ≤ ym is given by h = yl + ηsl (33) where η satisfies the equation 1− 1 b(r − l,m− r + 1) r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [1 + (m− r + 1 + j)η]−l = α. (34) advances in systems science and applications (2013) vol.13 no.4 363 2.3 new-within-sample prediction theorem 5. let x1 ≤ ... ≤ xkbe the first kordered observations from a previous sample of size n from the exponential distribution (1) and y1 ≤ ... ≤ yk be the first lordered early observations from a new sample of size m also from the distribution (1), where the parameter α is unknown. then the probability distribution function of the ancillary statistic (yr − yl)/(sk + sl) is given by pr { yr − yl sk + sl ≤ yr − yl sk + sl } = 1− 1 b(r − l, (m− r + 1)) r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j × [ 1 + (m− r + 1 + j) yr − yl sk + sl ]−(k+l) (35) proof. for the proof we refer to theorems 1 and 4. corollary 5.1. a lower one-sided within-sample α prediction limit h on the rth order statistic yr in a set of m future ordered observations y1 ≤ ... ≤ yk is given by h = yl + η(sk + sl), (36) where η satisfies the equation 1− 1 b(r − l,m− r + 1) r−l−1∑ j=0 ( r − l − 1 j ) (−1)j m− r + 1 + j [1 + (m− r + 1 + j)η]−(k+l) = α. (37) 3 inventory management froblem and predictive inference in this section, we consider a single-period newsvendor model. single-period stocking decisions often occur in practice ; these require the decision maker to choose the stocking level of an item for which demand exists for only a single period. several factors affect this decision : the distribution of demand, the cost and price of the item, the salvage value of the item, and the loss of customer goodwill due to stockouts. the newsvendor model addresses this problem and develops a formula for what is usually called the ‘critical-fractile’. the optimum order quantity is calculated using the critical-fractile of the distribution of the demand for the period. underlying the mathematical simplicity of the critical-fractile formula is a powerful and intuitively appealing insight for the determination of the order quantity. the order quantity depends only on the optimum balance between two types of costs. the first is the cost per unit associated with the unavailability of stock to meet the manifest demand (i.e., underage cost). the second is the unit cost associated with excess inventory at the end of the period for which there is no demand (i.e., overage cost). following hadley and whitin [17], we review 364 konstantin n. nechval : technique of constructing predictive inferences for future... the single-period newsvendor model and provide a broader interpretation to the structure of its solution. the notation, we use for the newsvendor model, is given below. y random variable for single-period demand fθ(y) probability density function of single-period demand fθ(y) probability distribution function of single-period demand θ parameter (in general, vector) c1 unit selling price c unit procurement cost, which is independent of the procured amount g unit salvage value for unsold items remaining at the end of the period c2 unit stockout penalty cost (over and above any lost profit) u variable representing the order quantity u∗ optimal order quantity q(u) expected profit as a function of the order quantity co unit overage cost cu unit underage cost different versions of the problem may equivalently consider expected opportunity cost minimization or expected profit maximization. we examine the latter and write the expected profit as q(u) = c1 ∫ u 0 yfθ(y)dy + c1u ∫ ∞ u fθ(y)dy − cu+ g ∫ u 0 (u− y)fθ(y)dy − c2 ∫ ∞ u (y − u)fθ(y)dy (38) assuming that c1 + c2 > g (which is generally true in most situations), we can show that the expected profit function q(u) is concave. we, therefore, set the first derivative equal to zero to find a maximizing solution. the value of u that maximizes (38) is the one that satisfies fθ(u ∗) = (c1 − c+ c2)/(c1 − g + c2). (39) clearly the lost contribution margin (c1-c) plus the stockout penalty (c2) represents the ‘underage cost’. similarly the item cost (c) minus the salvage value (g) represents the ‘overage cost’. if we refer to the underage cost as cu and the overage cost as co, we may rewrite (39) as fθ(u ∗) = cu/(cu + co). (40) we should choose the order quantity u∗ such that the cumulative distribution function (cdf ) of u∗ equals the ratio of the underage cost to the sum of the advances in systems science and applications (2013) vol.13 no.4 365 underage and overage costs. a relatively high underage cost results in a higher order quantity, whereas a relatively high overage cost leads to a lower order quantity, as one would expect. if the single-period demand y follows the exponential distribution (1) then q(u) = σ [ (c1 − g)− (c− g) u σ − (c1 − g + c2) exp ( −u σ )] (41) u∗ = σ ln ( 1 + c1 − c+ c2 c− g ) (42) and q(u∗) = σ [ c1 − c− (c− g) ln ( 1 + c1 − c+ c2 c− g )] (42) parametric uncertainty. consider the case when the parameter σ is unknown. let x1 ≤ ... ≤ xn be the past observations (order statistics) of single-period demands from the exponential distribution (1).then s = n∑ i=1 xi, (44) is a sufficient statistic for σ ; sis distributed with gσ(s) = [γ(n)σn]−1sn−1 exp(−s/σ)(s > 0) (45) to find the best invariant decision rule µbi , we use the invariant embedding technique [12-16] to transform (41) to the form, which is depended only on the pivotal quantity v = s/σ and the ancillary factor η = u/s. in statistics, a pivotal quantity or pivot is a function of observations and unobservable parameters whose probability distribution does not depend on unknown parameters. note that a pivotal quantity need not be a statisticąłthe function and its value can depend on parameters of the model, but its distribution must not. if it is a statistic, then it is known as an ancillary statistic. transformation of q(u) is given by q(η|v) = σ [(c1 − g)− (c− g)ηv − (c1 − g + c2) exp (−ηv)] (46) we find the expected profit for the statistical decision u = ηs as e{q(η)} = ∫ ∞ 0 q(η|v)g(v)dv =σ [ (c1 − g)− (c− g)ηn− (c1 − g + c2)(η + 1)−n] (47) where g(v) = [γ(n)]−1vn−1 exp(−v)(v > 0). (48) 366 konstantin n. nechval : technique of constructing predictive inferences for future... the value of η that maximizes (47) is given by η∗ = [1 + (c1 − c+ c2)/(c− g)]1/(n+1) − 1. (49) thus, ubi = η∗s = [ [1 + (c1 − c+ c2)/(c− g)]1/(n+1) − 1 ] s. (50) comparison of decision rules. for comparison, consider the maximum likelihood decision rule that may be obtained from (42) as uml = σ̂ ln [1 + (c1 − c+ c2)/(c− g)] = ηmls, (51) where σ̂ = s/n is the maximum likelihood estimator of σ, ηml = ln [1 + (c1 − c+ c2)/(c− g)]1/n. (52) since µbi and µml belong to the same class c = {u : u = ηs}, (53) it follows from the above that µml is inadmissible in relation to µbi . if, say, c= 2, c1=490, c2=2, g=1 (in terms of money), and n=1, we have that rel.eff.e{q(η)}{uml, ubi, σ}=e{q(ηml)} / e{q(η∗) = (c1 − g)− (c− g)ηmln− (c1 − g + c2) 1 (ηml+1)n (c1 − g)− (c− g)η∗n− (c1 − g + c2) 1 (η∗+1)n = 0.93. (54) thus, in this case, the use of µbi leads to a growth in the expected profit of about 7% as compared with µml. the absolute expected profit will be proportional to σ and may be considerable. predictive inference.it will be noted that the predictive probability density function of the single-period demand y, which is compatible with (38), is given by f(y|s) = n+ 1 s ( 1 + y s )−(n+2) (y > 0). (55) using (55), the predictive profit is determined as q(p)(u|s) = c1 ∫ u 0 yf(y|s)dy + c1u ∫ ∞ u f(y|s)dy − cu+ g ∫ u 0 (u− y)f(y|s)dy − c2 ∫ ∞ u (y − u)f(y|s)dy = s n [ (c1 − g)− (c− g) u s n− (c1 − g + c2) (u s + 1 )−n ] , (56) advances in systems science and applications (2013) vol.13 no.4 367 which can be reduced to q(p)(η|s) = s n [ (c1 − g)− (c− g)ηn− (c1 − g + c2) 1 (η + 1)n ] . (57) it follows from (57) that eσ{q(p)(η|s)} = ∫ ∞ 0 q(p)(η|s)gσ(s)ds = σ [ (c1 − g)− (c− g)ηn− (c1 − g + c2)(η + 1)−n] = e{q(η)} (58) thus, ubi can be found immediately from (56) as ubi = argmax u q(p)(u|s). (59) 4 conclusion the technique proposed in this paper represents a simple and computationally attractive statistical method based on the constructive use of the invariance principle in mathematical statistics. the main advantage of this technique consists in that it allows one to eliminate unknown parameters from the problem and to use the previous and current data of observations for obtaining predictive inferences as completely as possible. we have illustrated the technique for the exponential distribution. applications to other log-location-scale distributions could follow directly. références [1] scarf h. (1959), “bayes solutions of statistical inventory problem”, ann. math. statist, vol.30, pp.490-508. [2] karlin s. (1960), “dynamic inventory policy with varying stochastic demands”, management sci, vol.6, pp.231-258. [3] azoury k.s. (1985), “bayes solution to dynamic inventory models under unknown demand distribution”, management sci, vol.31, pp.1150-1160. [4] conrad s.a. (1976), “sales data and the estimation of demand”, oper. res. quart, vol.27, pp.123-127. [5] liyanage l.h., shanthikumar j.g. (2005), “a practical inventory control policy using operational statistics”, oper. res. lett, vol.33, pp.341-348. [6] kaminsky k.s., nelson p.i. (1998), “prediction of order statistics”, in : handbook of statistics-17 : order statistic : applications, n. balakrishnan and c. r. rao (eds.), elsevier science, pp.431-450. 368 konstantin n. nechval : technique of constructing predictive inferences for future... [7] aitchison j., sculthorpe d. (1965), “some problems of statistical prediction”, biometrika, vol.52, pp.469-483. [8] lawless j.f. (1982), statistical models and methods for lifetime data, new york : john wiley. [9] fisher r.a. (1934), “two new properties of mathematical likelihood”, proc. roy. statist. soc, a, vol.144, pp.285-307. [10] nechval n.a., nechval k.n., vasermanis e.k. (2003), “statistical models for prediction of the fatigue crack growth in aircraft service”, in : fatigue damage of materials 2003, a. varvani-farahani and c. a. brebbia (eds.). southampton, boston : wit press, pp.435-445. [11] murthy d. n. p., xie m., jiang y. (2004), weibull models, new york : john wiley and sons inc. [12] nechval n.a., nechval k.n., vasermanis e.k. (2003), “effective state estimation of stochastic systems”, kybernetes (an international journal of systems & cybernetics), vol.32, pp.666-678. [13] nechval n.a., berzins g., purgailis m., nechval k.n. (2008), “improved estimation of state of stochastic systems via invariant embedding technique”, wseas transactions on mathematics, vol.7, pp.141-159. [14] nechval n.a., nechval k.n., purgailis m. (2011), “prediction of future values of random quantities based on previously observed data”, engineering letters, vol.9, pp.346-359 [15] nechval n.a., purgailis m., nechval k.n., strelchonok v.f. (2012), “optimal predictive inferences for future order statistics via a specific loss function”, iaeng international journal of applied mathematics, vol.42, pp.40-51. [16] nechval n.a., purgailis m., cikste k, berzins g., nechval k.n. (2010), “optimization of statistical decisions via an invariant embedding technique”, in : lecture notes in engineering and computer science : proceedings of the world congress on engineering 2010, wce 2010, london, 30 june-2 july 2010, pp.1776-1782. [17] hadley g., whitin t.m. (1963), analysis of inventory systems, new jersey : prentice-hall inc. corresponding author nicholas a. nechval can be contacted at : nechval@junik.lv microsoft word 1111-article text-7257-1-6-20220131 adv syst sci appl 2024; 03; 104-113 published online at https://ijassa.ipu.ru. the approximation matrix method and its comparison with the analytical hierarchy process by t. saaty viktor p. korneenko v.a. trapeznikov institute of control sciences of the russian academy of sciences, moscow, russia e-mail: vkorn@ipu.ru abstract: the article presents an optimization method for the formation of quantitative weights of objects (importance of criteria, priorities of alternatives) according to the initial expert judgment matrix in multi-criteria selection problems. since the matrix of pairwise comparisons can be considered as some perturbation of the multiplicative matrix, the proposed method is based on the approximation of the original matrix of pairwise comparisons by the multiplicative matrix according to the matrix criterion of minimum distances between matrices. there is a one-to-one mapping between the elements of the weight vectors and the elements of the multiplicative matrix. for the first time, using a specific example using the matrix criterion, a relative estimate of the approximate solution of the analytical hierarchy process by t. saaty concerning the optimal solution obtained by the approximation matrix method is given. on account of the approximation matrix method being mathematically justified and due to the simplicity of finding optimal solutions, it can be recommended instead of the analytical hierarchy process by t. saaty. keywords: multi-criteria choice, normalized object weights, expert judgment matrix, multiplicative matrix, matrix criterion. 1. introduction when solving applied problems of multicriteria choice on a set of objects (alternatives, management decisions, options) presented in the form of preference relations, the problem of their expert measurement in the quantitative scale of relations arises. to date, many approaches and methods have been proposed to solve this problem, based on the resulting preference relations to narrow the set of non-dominant alternatives, as well as on paired comparisons of objects (solutions, criteria) [1, 2], which do not always allow us to identify a single alternative or management solution without attracting additional information. however, when solving applied problems related to the measurement of objects in expert scales, as well as the formation of local weights of criteria presented in the form of a hierarchical tree of the importance of criteria, expert methods of evaluating and ranking objects are usually used [3]. direct methods of expert evaluation of criteria weights have found application in the planning methodology through relative indicators of technical evaluation (pattern). experts are asked to evaluate the normalized local weights of criteria at each level of the hierarchy on a quantitative scale, and then the global weights are found by multiplying local weights along the branches of a multi-level criteria tree [4]. another expert approach based on the matrix of paired comparisons to assign "weights" to a finite set of compared objects is the analytical hierarchy process (ahp) by t. saaty which is now firmly established in the theory and practice of multicriteria selection problems [5–7]. following the analytical hierarchy process, experts form a so-called matrix of paired comparisons (judgments) of objects 𝑉 = [𝑣 ], 𝑖, 𝑗 = 1, 𝑛, in the scale of relations, and then find the approximation matrix method and its comparison… 105 copyright ©2024 assa adv. in systems science and appl. (2024) the right eigenvector �⃗� = (𝑤 , . . . , 𝑤 ) of this matrix, corresponding to the maximum eigenvalue. the desired weight vector is a vector whose elements are normalized by the sum of the elements of the right eigenvector. since the calculation of the vectors of weights of objects (criteria and alternatives) is performed by a numerical (approximate) method, t. saaty aware of this, introduced a special numerical indicator: consistency index the compatibility index of the judgment matrix 𝑉 and the multiplicative matrix 𝑊 obtained based on the values of the eigenvector, in the form of the hadamard product [9]: 𝑆. 𝐼. = 𝑒 𝑉°𝑊 𝑒, 𝑒 = (1,1, … ,1), where v is the original judgment matrix, and 𝑊 = [𝑤 ] = is a multiplicative square matrix whose elements are determined from the normalized elements of the right eigenvector �⃗� of the judgment matrix v. in this case, the ratio takes place: 𝑤 = 𝑤 𝑤 = 𝑤 / ∑ 𝑤 𝑤 / ∑ 𝑤 = 𝑤 𝑤 , 𝑖, 𝑗 = 1, 𝑛. the s.i. index it characterizes the degree of confidence in the results obtained with the help of ahp and is interpreted as a kind of measure of the deviation of the initial perturbed judgment matrix 𝑉 from the multiplicative one w. in the work of t. saaty [8, p. 76], it is shown that if we perform the hadamard matrix product, then the compatibility index takes the form 𝑆. 𝐼. = > 1, where λ is the maximum eigenvalue of the matrix 𝑉. with a sufficiently close approximation to the unit value of the index, the matrix of paired comparisons 𝑉 is "close" to the multiplicative matrix w. if the consistency index exceeds a certain "threshold" value, then it is impossible to conclude the proximity of these matrices, and therefore it is not recommended to use ai in such cases. the hierarchy analysis method and its applications are described in many reviews, monographs, scientific articles, as well as works popularizing this method [10–15]. however, it can be clearly stated that the method of hierarchy analysis by t. saaty is approximate since the method of determining the weight vector is based on numerical (approximate) methods for calculating the roots of a polynomial. the problem of finding the roots of polynomials of degree 𝑛 ≥ 5 is unsolvable in radicals. therefore, in t. saaty's method, for the number of objects at least five, the procedure for finding the eigenvalues of the matrix of degrees of the superiority of the importance of criteria or preferences of alternatives is carried out using numerical approximate methods for finding the roots of a polynomial implemented in the package expert choice [16]. v.d. noghin states that "the value of the compatibility index can only indirectly judge the magnitude of the final "model" error: it can never be precisely determined by anyone. this is the specificity of this heuristic approach" [10, p. 1194]. the article suggests a more efficient method of the approximation matrix (mam) for forming optimal object weights based on the matrix criterion of distance to the original matrix of judgments than the method of analyzing hierarchies of t. saaty. 2. statement of the problem of approximation of matrices of judgments in multi-criteria applied problems the aggregation mechanism is usually represented in the form of an additive candle of object ratings (alternatives, variants) 𝑎 ∈ 𝐴 = {𝑎 |𝑙 = 1, 𝑛 }, according to criteria with weights of importance in the form [1]: 𝐹(𝑎 , 𝑤 , … , 𝑤 ) = ∑ 𝑤 𝑓 (𝑎 ), ∑ 𝑤 = 1, 106 v.p. korneenko copyright ©2024 assa adv. in systems science and appl. (2024) where 𝑓 (𝑎 ) is the object score 𝑎 in the resulting scale according to the criterion 𝑓 , 𝑗 = 1, 𝑛; 𝑤 = 𝑤( 𝑓 ) is the quantitative (normalized) weight 𝑓 of the criterion. let us consider the formulation of the formation of object weights by the criterion of proximity to the original matrix of paired comparisons of a multiplicative matrix. let us have as initial data: 𝑉 = [𝑣 ] – the initial expert matrix of judgments about the relative importance of objects (criteria) 𝑓 ∈ 𝐹, 𝑖, 𝑗 = 1, 𝑛; 𝑊 – the set of multiplicative square matrices n of the nth order over the field of real numbers. as a measure of proximity 𝑑(𝑉, 𝑊) between the original 𝑉 = [𝑣 ] and the multiplicative 𝑊 = 𝑤 matrix, where 𝑊 ∈ 𝑊 , we take the square of the euclidean l2-norm equal to the difference of these matrices [9]: 𝑑(𝑉, 𝑊) = ‖𝑉 − 𝑊‖ = (𝑣 − 𝑤 ) . (2.1) then the mathematical formulation of the problem of choosing a multiplicative matrix w that approximates the original matrix 𝑉 = [𝑣 ], i.e., the closest approach to the original expert judgment matrix, is reduced to minimizing the indicator (2.1) in the form (v − w ) → 𝑚𝑖𝑛 ∈ , (2.2) provided that the matrix elements are multiplicative 𝑊 = [𝑤 ]: w = w w , ∀ i, j, k = 1, n. (2.3) due to condition (2.3), it is not analytically possible to obtain a solution to the original problem using one of the classical optimization methods, for example, the lagrange multiplier method. let us use the properties of the multiplicative matrix and solve this problem by reducing the original problem to an equivalent problem. 3. special linear properties of a multiplicative matrix to solve the original problem (2.2)–(2.3), it is necessary to find elements of the multiplicative 𝑊 = 𝑤 , ∀𝑖, 𝑗 = 1 , 𝑛, matrix, that provides a minimum for the quadratic criterion 𝑑(𝑉, 𝑊) (2.1) under condition (2.3). it turns out that for a multiplicative matrix, the relationship between the elements of the columns of the matrix and the elements of the right eigenvector is valid. to do this, consider two statements. the work of b.g. mirkin [17, p. 183–184] provided that the matrix 𝐵 = 𝑏 , ∀ 𝑖, 𝑗 = 1, 𝑛 is over traditional if there exists a positive vector 𝑥 = (𝑥 , . . . , 𝑥 ) such that that 𝑏 = and the vector 𝑥 is a point of equilibrium process: 𝑞 = 𝐵𝑞 , 𝑡 = 1, 2, . . ., which in the limit leads to a private vector 𝑞 = lim ⟶ 𝑞 . we show that the multiplicative matrix has several other properties that will be useful in reducing the original, analytically unsolvable problem (2.2)–(2.3) in the framework of classical optimization methods to an equivalent one. 3.1. the relationship between normalized column elements of a multiplicative matrix and elements of the right eigenvector theorem 3.1: 1. between the elements 𝑤 of a multiplicative matrix 𝑊 = 𝑤 , ∀𝑖, 𝑗 = 1, 𝑛, and any pair (𝑤 , 𝑤 ) component of the right eigenvector �⃗� = (𝑤 , . . . , 𝑤 ) true bijective mapping ↦ 𝑤 , whose every attitude to one mapping element 𝑤 of the matrix w and back the approximation matrix method and its comparison… 107 copyright ©2024 assa adv. in systems science and appl. (2024) 𝑤 ↦ 𝑤 𝑤 , provided this is true equality: 𝑤 = 𝑤 𝑤 , ∀ 𝑖, 𝑗 = 1, 𝑛. (3.4) 2. the normalized elements 𝑤 of the columns �⃗� = (𝑤 , … , 𝑤 ) , 𝑗 = 1, 𝑛, coincide with each other and are equal to the normalized right eigenvector, i.e. �⃗� = 𝑤 … 𝑤 = 𝑤 … 𝑤 = �⃗�, ∀ 𝑗 = 1, 𝑛, (3.5) where 𝑤 = ∑ is the normalized element 𝑗 of the j-th column of �⃗� ; 𝑤 = ∑ is the normalized element of the i-th row of the right eigenvector �⃗�. 3. the multiplicative matrix has one basis row. proof. 1. since by hypothesis the matrix 𝑊 = 𝑤 multiplicative then will provide that, if for any pair of numbers from {𝑤 , . . . , 𝑤 } the validity of the equation (3.4), then in this case the condition of multiplicative between the elements of the matrix 𝑊 = , i.e. thus there exists a one-to-one mapping between the elements: 𝑤 ↔ . indeed, if 𝑤 = and 𝑤 = , then we have 𝑤 𝑤 = × = = 𝑤 , that is, for all 𝑖, 𝑗, 𝑘 = 1, 𝑛 , the multiplicativity condition is satisfied. 2. let the elements of the vector �⃗� = (𝑤 , … , 𝑤 ) satisfy equality (3.4). we show that �⃗� = (𝑤1,. . . , 𝑤 𝑤𝑛) is the right eigenvector of the matrix 𝑊 = , ∀𝑖, 𝑗 = 1, 𝑛. let us verify that 𝑊�⃗� = λ�⃗� is valid: 𝑊�⃗� = 𝑤 /𝑤 𝑤 /𝑤 … 𝑤 /𝑤 𝑤 /𝑤 … 𝑤 /𝑤 … … … 𝑤 /𝑤 … 𝑤 /𝑤 𝑤 /𝑤 … 𝑤 /𝑤 𝑤 𝑤 … 𝑤 = 𝑛𝑤 𝑛𝑤 … 𝑛𝑤 = 𝑛 𝑤 𝑤 … 𝑤 = 𝑛�⃗�, where λ = 𝑛 is the maximum eigenvalue. for arbitrary columns �⃗� и �⃗� 𝑤𝑘 𝑎𝑛𝑑 𝑤𝑞 (1 ≤ 𝑘, 𝑞 ≤ 𝑛) of the matrice w, whose elements satisfy the multiplicativity condition 𝑤 (2.3), we normalize the elements. since 𝑤 = 𝑤 𝑤 , then 𝑤 = 𝑤 /𝑤 , whence for any column numbers 𝑘, 𝑞 we have: 𝑤 = ∑ = / ∑ = ∑ = ∑ = 𝑤 , ∀ 𝑖 = 1, 𝑛, i.e., the normalized components of the columns of the matrix coincide with each other. on the other hand: 𝑤 = 𝑤 ∑ 𝑤 = 𝑤 /𝑤 ∑ 𝑤 /𝑤 = 𝑤 ∑ 𝑤 = 𝑤 , ∀ 𝑖 = 1, 𝑛. thus, the normalized components of the matrix columns are also equal to the normalized elements of the right vector of the multiplicative matrix. 3. since the matrix 𝑊 = by linear transformations is reduced to the form: 108 v.p. korneenko copyright ©2024 assa adv. in systems science and appl. (2024) 𝑤 /𝑤 𝑤 /𝑤 . . . 𝑤 /𝑤 𝑤 /𝑤 . . . 𝑤 /𝑤 . . . . . . . . . 𝑤 /𝑤 . . . 𝑤 /𝑤 𝑤 /𝑤 . . . 𝑤 /𝑤 ~ ∏ 𝑤 1/𝑤 1/𝑤 . . . 1/𝑤 1/𝑤 . . . 1/𝑤 . . . . . . . . . 1/𝑤 . . . 1/𝑤 1/𝑤 . . . 1/𝑤 ~ ∏ 𝑤 1/𝑤 0 . . . 0 1/𝑤 . . . 0. . . . . . . . . 0. . . 1/𝑤 0 . . . 0 , the rank of the multiplicative matrix is equal to one: rg 𝑊 = 1, i.e., the multiplicative matrix has one basic row. the theorem is proved. ∎ 3.2. reducing the elements of the matrix columns to integer values and equality to the right eigenvector theorem 3.2: let the elements 𝑞 , 𝑖, 𝑘 = 1, 𝑛, of the multiplicative matrix 𝑊 = [𝑞 ], ∀ 𝑖, 𝑘 = 1, 𝑛, be represented in general by integers and rational numbers in the form of irregular fractions. then, if the elements of 𝑞 any column �⃗� , 𝑘 = 1, 𝑛, of the matrix is multiplied by the least common multiple 𝑛 denominators of the rational elements of the column, obtained in the result of this procedure, the integer elements 𝑧 = 𝑛 𝑞 , ∀ 𝑘 = 1, 𝑛, the columns of the matrix coincide with each other on any line number, and will be equal to the integer elements 𝑤 right eigenvector �⃗� = (𝑤 , . . . , 𝑤 , … 𝑤 ) corresponding to its own maximum eigenvalue 𝜆 = 𝑛: z … z = 𝑤 … 𝑤 , 1 ≤ 𝑘 ≤ 𝑛. (3.6) proof. by theorem 1, the elements of any multiplicative inverse-symmetric matrix can be brought to mind (3.4), namely: 𝑞 = , where 𝑤 are positive integer numbers – the elements of the right eigenvector, the least common multiple of positive integers to rational denominators of the elements of such a multiplicative matrix is equal to 𝑛 = 𝑤 , ∀ 𝑘 = 1, 𝑛. as a result, for k-th column �⃗� = , . . . , , we find 𝑛 �⃗� = 𝑛 × 𝑤 /𝑛 … 𝑤 /𝑛 = 𝑤 × 𝑤 /𝑤 … 𝑤 /𝑤 = 𝑤 … 𝑤 = �⃗�, ∀ 𝑘 = 1, 𝑛, that is, we proved the correctness of (3.6). it is easy to verify that �⃗� is the right eigenvector of the matrix 𝑊, which was required to prove. the theorem is proved. ∎ example 1. consider a multiplicative matrix 𝑊 = ⎝ ⎜ ⎜ ⎜ ⎛ 1 1 2 1 3 1 5 2 1 2 3 2 5 3 3 2 1 3 5 5 5 2 5 3 1⎠ ⎟ ⎟ ⎟ ⎞ (3.7) for the columns of the matrix, the lowest common multiples are: 𝑛 = 𝑤 = 2 × 3 × 5 = 30, 𝑛 = 𝑤 = 3 × 5 = 15, 𝑛 = 𝑤 = 2 × 5 = 10, 𝑛 = 2 × 3 = 6. from here we come to the matrix: 𝑊 = 30 15 10 6 30 15 10 6 30 15 10 6 30 15 10 6 . it is easy to verify that the right the approximation matrix method and its comparison… 109 copyright ©2024 assa adv. in systems science and appl. (2024) eigenvector �⃗� = (30, 15, 10, 6) of the matrix w (3.7), corresponding to the eigenvalue 𝜆 = 4 and the normalized elements of the columns of the matrix coincide with each other and are equal to the elements of the normalized right vector of the matrix: �⃗� = 30 61 , 15 61 , 10 61 , 6 61 . the original matrix can also be represented as 𝑊 = : 𝑊 = 30/30 15/30 10/30 6/30 30/15 15/15 10/15 6/15 30/10 15/10 10/10 6/10 30/6 15/6 10/6 6/6 . 4. reducing the original problem to an equivalent one since, following (3.5), the elements of the 𝑖-th rows of the normalized vector columns of the multiplier and vector matrix are equal to the 𝑖-th normalized component of the right proper vector, namely: 𝑤 = ⋯ = 𝑤 = ⋯ = 𝑤 = 𝑤 , (4.8) where ∑ 𝑤 = 1, to we have: 𝑤 = = / ∑ / ∑ = , ∀𝑖, 𝑗 = 1, 𝑛. we show that the multiplicativity condition (2.3) for the normalized row elements (4.8) of the matrix is satisfied: 𝑤 𝑤 = 𝑤 𝑤 × 𝑤 𝑤 = 𝑤 𝑤 = 𝑤 , 𝑖, 𝑗, 𝑘 = 1, 𝑛. thus, finding the elements of the approximating matrix 𝑊 = [𝑤 ] of the problem (2.2)– (2.3) is equivalent to finding the elements of a normalized column vector: �⃗� = (𝑤 , … , 𝑤 , … , 𝑤 ) . as the target indicator of the approximation problem, we take the square of the difference between the normalized elements of the original and multiplicative matrices in the form: 𝜌 𝑉, 𝑊 = 𝑣 − 𝑤 (4.9) then the mathematical formulation of the original problem is reduced to finding the normalized right column vector of the approximating matrix and providing the minimum criterion (4.9): 𝑣 − 𝑤 → min ( ,…, ) , (4.10) where 𝑣 = ∑ are the normalized elements of the matrix 𝑉 = [𝑣 ] of judgments, 𝑗 = 1, 𝑛. the reasonableness of the transition from the original perturbed matrix 𝑉 = [𝑣 ] to the normalized 𝑉 = [𝑣 ] is based on the fact that the rank and magnitude of the relation between the elements of the same vector column are preserved for the original and normalized matrix. 110 v.p. korneenko copyright ©2024 assa adv. in systems science and appl. (2024) 4.1. theorem on the optimal solution: theorem 4.3: the optimal solution to the problem (4.10) is a normalized vector �⃗�∗ = 𝑤∗ … 𝑤∗ , the components of which are taken as the coefficients of the importance the criteria 𝑤∗ = 𝑤(𝑓 ) and calculated by the formula: 𝑤∗ = 1 𝑛 𝑣 , ∀ 𝑖 = 1, 𝑛, while delivering a minimum of the indicator (4.9). the elements of the optimal approximation matrix 𝑊∗ = 𝑤∗ are determined by the formula: 𝑤∗ = 𝑤∗ 𝑤∗. as an estimate of the approximation matrix to the original judgment matrix, we take the euclidean matrix norm [9]: 𝑑(𝑉, 𝑊∗) = ‖𝑉 − 𝑊∗‖ = 𝑣 − 𝑤∗ (4.11) the proof is trivial and is based on necessary and sufficient conditions for the existence of an optimal solution of a function of many variables. 4.2. algorithm for generating optimal weights for criteria step 1. normalization of elements of the original matrix of pairwise comparisons of the importance of criteria by the sum ∑ 𝑣 of elements columns of the original matrix of judgments: 𝑣 = ∑ are the normalized elements of the judgment matrix 𝑉 = [𝑣 ]. step 2. the calculation for each row of the normalized matrix 𝑉 = [𝑣 ] of judgments of its average value, we take it as the normalized weights of the criteria, namely: 𝑤∗ = 1 𝑛 𝑣 , 𝑤∗ = 𝑤(𝑓 ) ∀ 𝑖 = 1, 𝑛. step 3. recovery of the elements of the multiplicative matrix by the optimal normalized weights of the criteria: 𝑊 = ∗ ∗ . step 4. estimation of the proximity between the elements of the original matrix of pairwise comparisons and the optimal multiplicative matrix by the formula 𝑑(𝑉, 𝑊∗) (4.11). 5. comparison of the effectiveness of the method with the analytical hierarchy process by t. saaty let us compare the error of calculating the weights of the importance of objects (criteria, alternatives) by t. saaty's method with the optimization method of the approximation matrix for the approximation matrix method and its comparison… 111 copyright ©2024 assa adv. in systems science and appl. (2024) the formation of object weights in multi-criteria problems. as multiplicative matrices consider the square matrix 𝑊 = , , . and 𝑊 = ∗ ∗ , , . , in which the elements are determined by the vector of priorities, found by the method of analysis of hierarchies and using approximate matrices. to do this, we will use the data from the example of buying a house by a family with average incomes, given in the work of t. saaty (see [6, p. 41–44]). the problem consists in choosing one house from three available alternatives {a, b, c} based on eight factors that serve as criteria for the multi-criteria selection problem. table 5.1 shows expert estimates of pairwise comparison of the importance of criteria and the priority vector calculated from the maximum eigenvalue following the method of hierarchy analysis by t. saaty (𝜆 = 8,811, compatibility index 𝑆. 𝐼. = , ≈ 1,101). the summed elements of the judgment matrix and the vector of weights of the importance of criteria found from these data using the approximation matrix method are presented in table 5.2. table 5.1. initial matrix 𝑉 of judgments and priority vector according to ahp [6, p. 42] factor (criteria) 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 priority vector, 𝑤 size, 𝑓 1 5 3 7 6 6 1/3 1/4 0,175 transport, 𝑓 1/5 1 1/3 5 3 3 1/5 1/7 0,062 environment, 𝑓 1/3 3 1 6 3 4 1/2 1/5 0,103 age, 𝑓 1/7 1/5 1/6 1 1/3 1/4 1/7 1/8 0,019 yard, 𝑓 1/6 1/3 1/3 3 1 1/2 1/5 1/6 0,034 facilities, 𝑓 1/6 1/3 1/4 4 2 1 1/5 1/6 0,041 state, 𝑓 3 5 2 7 5 5 1 1/2 0,221 finance, 𝑓 4 7 5 8 6 6 2 1 0,348 table 5.2. normalized matrix 𝑉 = ∑ of judgments and a vector of priorities by mam factor (criteria) 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 priority vector, 𝑤 size, 𝑓 0,11 0,23 0,25 0,17 0,23 0,23 0,07 0,10 0,174 transport, 𝑓 0,02 0,05 0,03 0,12 0,11 0,12 0,04 0,06 0,068 environment, 𝑓 0,04 0,14 0,08 0,15 0,11 0,16 0,11 0,08 0,108 age, 𝑓 0,02 0,01 0,01 0,02 0,01 0,01 0,03 0,05 0,021 yard, 𝑓 0,02 0,02 0,03 0,07 0,04 0,02 0,04 0,07 0,038 facilities, 𝑓 0,02 0,02 0,02 0,10 0,08 0,04 0,04 0,07 0,047 state, 𝑓 0,33 0,23 0,17 0,17 0,19 0,19 0,22 0,20 0,212 finance, 𝑓 0,44 0,32 0,41 0,20 0,23 0,23 0,44 0,39 0,333 to evaluate the accuracy, we restore multiplicative matrices of pairwise relations by priority vectors: �⃗� = (0,175; 0,062; 0,103; 0,019; 0,034; 0,041; 0,221; 0,348), (5.12) �⃗� = (0,174; 0,068; 0,108; 0,021; 0,038; 0,047; 0,212; 0,333), (5.13) obtained by the method of calculating the eigenvector and the method of approximating the matrix of pairwise comparisons using the minimum distance criterion, and compare the results. tables 5.3 and 5.4 present multiplicative matrices �⃗� , �⃗� , with priority vectors �⃗� (5.12) and �⃗� (5.13) formed from normalized weights. table 5.3. multiplicative matrix w , formed from normalized weights of the priority vector 𝑤 factor (criteria) 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 size, 𝑓 1,00 2,82 1,70 9,21 5,15 4,27 0,79 0,50 transport, 𝑓 0,35 1,00 0,60 3,26 1,2 1,51 0,28 0,18 environment, 𝑓 0,59 1,66 1,00 5,42 3,03 2,51 0,47 0,30 age, 𝑓 0,11 0,31 0,18 1,00 0,56 0,46 0,09 0,05 yard, 𝑓 0,19 0,55 0,33 1,79 1,00 0,83 0,15 0,10 112 v.p. korneenko copyright ©2024 assa adv. in systems science and appl. (2024) facilities, 𝑓 0,23 0,66 0,40 2,16 1,21 1,00 0,19 0,12 state, 𝑓 1,26 3,56 2,15 11,63 6,50 5,39 1,00 0,64 finance, 𝑓 1,99 5,61 3,38 18,32 10,24 8,49 1,57 1,00 table 5.4. multiplicative matrix 𝑊 formed from normalized weights of the priority vector 𝑤 factor (criteria) 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 𝑓 size, 𝑓 1,00 2,54 1,62 8,37 4,63 3,70 0,82 0,52 transport, 𝑓 0,39 1,00 0,64 3,30 1,83 1,46 0,32 0,21 environment, 𝑓 0,62 1,57 1,00 5,17 2,86 2,29 0,51 0,32 age, 𝑓 0,12 0,30 0,19 1,00 0,55 0,44 0,10 0,06 yard, 𝑓 0,22 0,55 0,35 1,81 1,00 0,80 0,18 0,11 facilities, 𝑓 0,27 0,69 0,44 2,26 1,25 1,00 0,22 0,14 state, 𝑓 1,22 3,10 1,98 10,21 5,65 4,52 1,00 0,64 finance, 𝑓 1,91 4,86 3,10 16,02 8,86 7,08 1,57 1,00 we find the values of the norm of the difference between the original matrices and the multiplicative matrices of relations formed from the values of the priority vector (the importance of objects): 𝑑 = ‖𝑉 − w ‖ = √200,66 ≈ 14,2; 𝑑 = ‖𝑉 − 𝑊 ‖ = √139,61 ≈ 11,8. let us determine the accuracy of the solution 𝑑 by the method of t. saaty concerning the optimal 𝑑 obtained in the framework of the optimization problem by the criterion (4.11). using the ε-approximation formula, we find that: 𝜀 = |𝑑 − 𝑑 | 𝑑 × 100 % = 14,2 − 11,8 11,8 × 100 % ≈ 20,3 %. comparison of methods for the accuracy of obtaining the weights of objects for the original perturbed matrix of the 8th order is shown in fig. 5.1. fig 5.1. comparison of methods based on the accuracy of obtaining object weights it follows that the error estimate is 20,3 % and the ahp solution that differs from the optimal one by this value cannot be considered satisfactory. thus, t. saaty's method of finding priorities for the importance of criteria and objects based on the eigenvector of the matrix of pairwise comparisons should be attributed to approximate, and not to exact, as the author of ahp declares. 7. conclusion the paper explores the problem of forming quantitative weights of objects based on the matrix of pairwise comparisons in the relationship scale. since in the hierarchy analysis process of t. saaty, the method of determining the weight vector is carried out using polynomials, the problem of finding the roots of polynomials of degree 𝑛 ≥ 5 is unsolvable in radicals. therefore, in t. saaty's method, for the number of objects at least five, the procedure for finding the eigenvalues of the matrix of degrees of the superiority of the importance of criteria 11,8 14,2 10,0 11,0 12,0 13,0 14,0 15,0 the approximation matrix method the analytical hierarchy process the distance of multiplicative matrices to the original matrix of judgments the approximation matrix method and its comparison… 113 copyright ©2024 assa adv. in systems science and appl. (2024) or preferences of alternatives is carried out using numerical methods for finding the roots of a polynomial implemented in the package expert choice [17]. on the other hand, and because of the inverse symmetry of the elements of the matrix of judgments 𝑉 = [𝑣 ], 𝑖, 𝑗 = 1, 𝑛, obtained by sequential comparison of all pairs of objects, the expert has to answer ( ) questions about the values 𝑣 . in this regard, the t. saaty's method is justified only for a small number of criteria and objects. in this article, using a concrete example, it is shown that the analytical hierarchy method is approximate and at the same time, its error is estimated relative to the optimal solution obtained by the method of the approximation matrix for the formation of object weights in multi-criteria problems. since the presented method is mathematically justified, and also because of the computational simplicity of forming object weights, it can be recommended instead of the method t. saaty in solving applied problems. references 1. figueira, j., greco, s., & ehrgott, m. (2005). multiple criteria decision analysis: state of the art surveys multiple criteria decision analysis: state of the art surveys. berlin, germany: springer. 2. noghin, v. d. (2018). reduction of the pareto set: an axiomatic approach. berlin, germany: springer. 3. korneenko, v. p. (2019). optimizatsionniy metod vybora rezuliruyushego ranzhirovaniya obiektov, predstablennyh v rangovoy shkale izmereniya [optimization method for selecting the resulting ranking of objects represented in the rank scale of measurement], large-scale systems control, 82, 44–60, [in russian], doi: 10.25728/ubs.2019.82.3 4. sigford, s. v., & parvin, r. h. (1995). project pattern a methodology for delernining relevance in complex decision-making, ieee trans., 12(1), 9–13. 5. saaty, t. l., & vargas, l. g. (1984). comparison of eigenvalue, logarithmic least squares and least squares method in estimating ratios, mathematical modeling, 5, 309–324. 6. saaty, t. l. (1986). axiomatic foundation of the analytic hierarchy process, management science, 32(7), 841–855. 7. saaty, t. l. (1990). the analytic hierarchy process: planning, priority setting, resource allocation. pittsburgh, pa: rws. 8. saaty, t. (2008). prinyatie resheniy pri zavisimostyah i obratnyh svyazyah: analiticheskiye seti [decision-making under dependencies and feedback: analytical networks]. moscow, lki, 360 p. [in russian]. 9. horn, r. a., & johnson, c. r. (2013). matrix analysis. cambridge, uk: cambridge university press. 10. noghin, v. d. (2004). a simplified variant of the analytic hierarchy processes based on a nonlinear scalarizing function, computational mathematics and mathematical physics, 44(7), 1194–1202. 11. zahedi, f. (1986). the analytic hierarchy process – a survey of the method and its applications, interfaces, 16, 96–108. 12. takeda, e. (1990). the analytic hierarchy process: an overview, systems, control and information, 34, 669–675. 13. forman, t. y., & gass, s. i. (2001). the analytic hierarchy process – an exposition, operations research, 21, 469–486. 14. bodin, l., & gass, s. i. (2003). on teaching the analytic hierarchy process, computer and operations research, 30, 487–1497. 15. vaidia, j. s., & kumar, s. (2006). analytic hierarchy process: an overview of applications, european journal of operational research, 168, 1–29. 16. expert choice. (2021). [online]. available: www.expertchoice.com. 17. mirkin, b. g. (1974). problema gruppovovo vybora [the problem of group selection]. moscow, ussr: nauka [in russian]. adv syst sci appl 2019; 03; 80-92 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/689 influence of information sharing, partnership, and collaboration in supply chain performance; study on apples agroindustry alfredo tutuhatunewa*1,2, surachman1, purnomo b. santoso1, imam santoso1 1) brawijaya university, mt haryono street 167 malang, indonesia e-mail: alfredo.tutuhatunewa@hotmail.com 2) pattimura university, ir. m. putuhena street poka, ambon, indonesia received january 11, 2019; revised august 26, 2019; published october 1, 2019 abstract: supply chain plays an essential role in the development of business organizations to achieve competitive advantage. supply chain studies are continuing mainly to improve supply chain performance. many of the literature write about the role of information sharing, partnership processes and collaboration in improving the performance of the supply chain. this research was conducted to find the relationship and influence of information sharing, supply chain partnership, and collaboration on supply chain performance on smes. the object of research is apple agroindustry, in east java, indonesia. the analysis was conducted with partial least square (pls). the results indicate only an indirect effect of information sharing on supply chain collaboration. also, indirect effect of information sharing on supply chain performance of smes. keywords: information sharing, partnerships, collaboration, partial least square, small and medium enterprises, supply chain performance. 1. introduction apple is an annual fruit plant originating from the west asian region with a sub-tropical climate. in indonesia, apples can grow and bear good fruit in the highlands. the centers of apple production in east java are malang (batu, pujon, and poncokusumo) and pasuruan (nongkojajar). in addition to apple cultivation, the processed industries of apples continue to be developed. processed industries are conducted to increase the added value of apples into various food and beverage products. the majority of the processed apple industry is a micro and small business unit (smes), which is a home industry. the apple agroindustry supply chain is dynamic because it involves the flow activity, among others: raw materials, finished products, ordering, shipping, payment, and information among the parties involved. it causes the overall management of these streams to be challenging to perform effectively. it is hard to meet the production targets due to the absence of raw materials at certain times, due to failure to share information and establish coordination and collaboration between actors in the supply chain of apple agroindustry. the supply chain (sc) is defined as the management of upstream and downstream relationships with suppliers and customers in order to deliver superior customer value at less cost to the supply chain as a whole [10]. sc has become an essential focus for business organizations to enhance competitive advantage. companies must implement the right supply chain management strategy to compete at the sc level. this strategy needs to be integrated and coordinated throughout the sc to produce the performance of sc members [14,18]. supply chain management (scm) studies emphasize how to maximize the overall value of a * corresponding author: alfredo.tutuhatunewa@hotmail.com influence of information sharing, partnership 81 copyright ©2019 assa. adv. in systems science and appl. (2019) company by using and sharing resources across the company better. its become apparent for many organizations that are assessing their performance is essential to succeeding efficient and effective sc [1]. due to globalization, outsourcing, customization, time to market, and pricing pressure have compelled enterprises to adopt efficient and effective scm [40]. therefore, it is crucial to coordinate decisions and actions among partners in sc to improve the performance of the sc [47]. a suitable coordination mechanism between actors in sc through an online information network plays an essential role in increasing the effectiveness of material flow, information, and money [34]. that is, if companies want to improve collaborative capabilities, then companies need to prepare themselves by building a network of information technology to support the ability to share information first. furthermore, the ability to share information and collaborative capabilities jointly affects supply chain performance [46]. the benefits of information sharing include improving inventory management, increasing sales, and knowing the demand better [19]. failure to share information in the supply chain causes a bullwhip effect [11], i.e., amplification of the demand flow variance, which is flowing in the entire sc, from customer to factory. sc collaboration is also influenced by the practice of partnerships among actors taking place in it [50]. 2. materials and methods information sharing information sharing refers to the extent to which critical information is communicated to other supply chain partners [27]. many researchers have emphasized the importance of sharing information in scm practice. moreover, yu et al. [49] suggest that the adverse effects of the bullwhip effect on the supply chain can be reduced or eliminated by sharing information with trading partners. the empirical findings of childerhouse & towill [9] reveal that a simplified flow of materials, including streamlining and making all the information flows through the chains, is an integrated and effective supply chain. currently, the companies do not operate alone; they are now connected to many other partners [29]. information sharing can help supply chain members establish partnerships for better supply chain system performance [49]. information sharing is found to impact operational performance [13]. meanwhile, information sharing is one of the characteristics of collaboration [35]. information sharing also plays an important role in supporting collaborative capabilities [46]. h1a: information sharing has a positive impact on sc partnership. h1b: information sharing has a positive impact on sc collaboration. h1c: information sharing has a positive impact on sc performance. partnership the partnership is defined as purposive strategic relationships between independent firms having common goals, striving for mutual benefit, and recognizing a high degree of interdependence [25]. in the context of cooperation, the relationship between the company and its suppliers can take many forms. the first, such as joint ventures or strategic alliances, involves negotiating and maintaining explicit contracts that explain expectations and deliveries and sometimes revenue-sharing [7] and they have a legal structure that sets the boundaries [45]. supply chain partnerships, on the other hand, tend to operate without formal contracts [22]. the type of relationship between buyers and suppliers might vary from hostilities to cooperatives [6]. research from zhang et al. [50] has concluded that partnership management 82 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) had a significantly positive influence on supply chain collaboration. moreover, the partnership has an importance role to supply chain performance [20,33,49]. h2a: partnership has a positive impact on sc collaboration. h2b: partnership has a positive impact on sc performance. collaboration collaboration is a process of participation of a group of people or organizations, working together to achieve mutually desired outcomes, and building interconnected systems to address problems and opportunities. collaboration is defined by several researchers among others, [26], which refer to long-term conditions, win-win conditions, and open information exchange agreements, in which both parties are engaged in joint efforts to improve performance and commit to quality, cooperation, and conflict resolution. meanwhile, [2] defines collaboration as two or more companies sharing responsibility in planning, management, execution, and performance measurement information. from the above understanding, it can be concluded that collaboration is a form of cooperation, interaction, and compromise of some individuals, or organizations, involved directly or indirectly, over a long period, receiving consequences and benefits. collaboration is based on values among others, the common goal, the similarity of perception, the willingness to process, and mutual benefit. collaboration involves multiple shared resources and responsibilities, in planning, implementing, and evaluating all activities to achieve common goals. all parties involved must be willing to share their vision, mission, resources, and strengths. in supply chains, collaboration involves designing a set of strategies where two or more different actors, with complementary capabilities, achieve shared aspirations and goals in a competitive environment, which cannot be achieved individually [21]. from this perspective, collaboration in the supply chain becomes an essential strategy for achieving competitive advantage. the escalation of competition, the flow of globalization, and the increasing demands of customers led to the idea that companies cannot compete on their own in the marketplace. therefore, companies look beyond their boundaries and establish cooperation with other complementary firms to minimize potential risks [32]. the literature suggests that collaboration is associated with increased performance, regarding increased visibility, increased service levels, increased flexibility, improved end-customer satisfaction, and reduced cycle time [12,37,43]. h3a: collaboration has a positive impact on sc performance. supply chain performance some experts and practitioners recommend several methods that accommodate all the dimensions of supply chain performance [39], namely: • total supply chain cost. fulfillment costs as a percentage of revenues or fulfillment costs per order case. • service level. it includes the level (availability the ratio of the number of items ordered by the customer and the number of items sent to the customer). • asset management. it focuses on capital utilization of investments in facilities and equipment and working capital invested in inventories. • customer accommodation. it aims to capture the size of the request flawlessly. • cash-to-cash cycle time. it is time it takes to convert the costs spent on inventory into profits that are collected from the proceeds of the sale. • benchmarking. it makes management aware of state-of-the-art business practices. it includes: internal benchmarking, competitor benchmarking, and benchmarking limited. influence of information sharing, partnership 83 copyright ©2019 assa. adv. in systems science and appl. (2019) the global performance of the supply chain can be enhanced by exchanging information between its members at different decision levels [28]. the performance of the supply chain is strongly influenced by two things, namely information sharing and collaboration capabilities [28,38,46,48]. the research framework is shown in fig. 1. fig. 2.1. research framework methodology the research was conducted in the supply chain of apple agroindustry, from apple farmers, suppliers of food additives, suppliers of plastic cups and bottles, packaging suppliers, apple processing agroindustry, distribution, and retailers, in batu city, east java. data collection was done by distribution the questionnaire, which was compiled to know how the process of share the information, partnership and collaboration in the supply chain of apple agroindustry, as well as supply chain performance. research variables and measuring instruments have described in table 2.1. data were collected from 36 smes, with product classification as in table 2.2. table 2.1. variable and its measuring instruments variable measuring instrument ref. information sharing (x1) • smes only shares inventory data with supply chain partners (x1-1); • smes shares inventory and demand data with other supply chain partners. (x1-2) • smes shares inventory, demand, and capacity or production data with other supply chain partners. (x1-3). [8] partnership (y1) • trust (y1-1) • commitment (y1-2) [5] [16] collaboration (y2) • resource sharing (y2-1) • decision synchronization (y2-2) • incentive alignment (y2-3) [4] supply chain performance (y3) • supply chain flexibility (y3-1) • the extent of co-operation (y3-2) • customer responsiveness (y3-3) [42] [17] [51] information sharing partnership collaboration sc performance h1a h1b h1c h2a h2b h3a 84 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) data analysis and structural equation modeling using partial least square (pls), are as follows: 1. measurement model (outer model). the measurement model defines how each indicator block corresponds to its latent variable. the measurement model design determines the indicator properties of each latent variable, whether reflexive or formative, based on the operational definition of the variable. 2. structural model (inner model). the structural model describes the relationship between latent variables based on substantive theory. the latent variables in this study are information sharing, partnership, collaboration, and supply chain performance. 3. path diagram. 4. evaluation of the goodness of fit q-square can be calculated by the equation: 𝑄2 = 1 − (1 − 𝑅1 2)(1 − 𝑅2 2). . (1 − 𝑅𝑝 2) (2.1) where: 𝑅1 2, 𝑅2 2 ... 𝑅𝑝 2 are r-square of the endogenous variable in the model. 5. hypothesis testing (resampling bootstrapping) table 2.2. product classification of respondents product number of respondents apple cider 14 apple chips 9 apple vinegar 2 another apple processed 11 3. results and discussion evaluation of measurement model convergent validity is used to measure the validity of a reflexive indicator as a latent variable measure, which can be seen from the outer loading of each variable indicator. an indicator is said to have excellent reliability if the value of outer loading above 0.70. results of outer loadings can be seen in table 3.1. for information sharing variables (x1), the three indicators show values greater than 0.7. thus, it can be concluded that the three indicators are capable of measuring information sharing variables well. in partnership variable (y1), two indicators, also have outer loading value more than 0.7, which means that both indicators capable of measuring partnership variables well. for collaboration variable (y2), which is an endogenous latent variable, capable explained by three indicators. so is the performance variable (y3), which has outer loading value greater than 0.7. table 3.1. outer loadings variables x1 y1 y2 y3 x1-1 0.901 x1-2 0.935 x1-3 0.923 y1-1 0.914 y1-2 0.916 influence of information sharing, partnership 85 copyright ©2019 assa. adv. in systems science and appl. (2019) y2-1 0.919 y2-2 0.946 y2-3 0.910 y3-1 0.918 y3-2 0.897 y3-3 0.822 the criteria for measuring discriminant validity can be seen on the cross-loading between the indicator and the construct. the result of cross-loading is shown in table 3.2. table 3.2. cross loading variables x1 y1 y2 y3 x1-1 0.901 0.587 0.528 0.559 x1-2 0.935 0.641 0.711 0.646 x1-3 0.923 0.586 0.599 0.577 y1-1 0.633 0.914 0.674 0.637 y1-2 0.574 0.916 0.715 0.676 y2-1 0.563 0.648 0.703 0.919 y2-2 0.636 0.703 0.835 0.946 y2-3 0.598 0.637 0.702 0.910 y3-1 0.661 0.753 0.918 0.759 y3-2 0.589 0.622 0.897 0.831 y3-3 0.509 0.625 0.822 0.509 the result of cross-loading shows that the correlation of the construct between each variable with the indicator has a higher value than the correlation with other indicators. it is concluded that the latent variables predict the indicator on it block better than the indicator on the other block. construct reliability can be measured by cronbach's alpha. this value reflects the reliability of all indicators in the model. the minimum value of 0.7 is ideally 0.8 or 0.9. results of data processing can be seen in table 3.3. table 3.3. construct reliability and validity cronbach's alpha rho_a composite reliability information 0.909 0.917 0.943 partnership 0.806 0.806 0.912 performance 0.855 0.876 0.911 collaboration 0.916 0.923 0.947 from table 3.3, it can be seen that all variables have cronbach's alpha more than 0.7 with the lowest value in the partnership variable. thus, it can be concluded that no reliability or unidimensionality problems were found in the established model. evaluation of the structural model the structural model analysis was conducted to examine the effect of information sharing, partnership, and collaboration on supply chain performance of apple agro-industrial. the analysis is done by using r-square. 86 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) r-square indicates the extent to which a construct can describe the model as a whole, or by indicating the magnitude of a particular influence of latent variables on the latent dependent variable. r-square value can be seen in table 3.4. table 3.4. r-square original sample (o) sample mean (m) standard deviation (stdev) t statistics p values partnership 0.434 0.441 0.153 2,835 0.005 performance 0.733 0.754 0.075 9,758 0,000 collaboration 0.569 0.594 0.096 5.904 0,000 from table 3.4., it can be summarized as follows:  r-square of partnership variables indicates that information sharing gives 43.4% influence on the partnership.  r-square of performance variable indicates that information sharing (x1), partnership (y1) and collaboration (y2) give the effect of 73.3% to performance.  r-square of the collaboration variable indicates that the information sharing (x1) and partnership (y1) give 56.9% influence to the collaboration variables (y2). path diagram fig. 3.1 shows the path diagram obtained. the diagram shows the relationships among information sharing, partnerships, collaboration and supply chain performance of apple agroindustry. fig. 3.1. path diagram evaluation of the goodness of fit evaluation of the goodness of fit is done by calculating a q-square value for the constructive model. q-square measures how well the model and its parameter estimation generate the observation value. the q-square value > 0 indicates that the model has predictive relevance, otherwise if the q-square value ≤ 0 indicates the model lacks predictive relevance. q-square can be calculated by equation (1): 𝑄2 = 1 − (1 − 𝑅1 2)(1 − 𝑅2 2)(1 − 𝑅3 2) information sharing collaboration performance partnership influence of information sharing, partnership 87 copyright ©2019 assa. adv. in systems science and appl. (2019) where: 𝑅1 2 = 0.434; 𝑅2 2 = 0.569; and 𝑅3 2 = 0.733 (see table 6) so, 𝑄2 = 0,935 from the calculation of q-square, it can be concluded that the resulting model has a very good predictive relevance of 93.5%. hypothesis testing the impact of information sharing on supply chain partnership. path diagram results in fig. 3.1 shows that at 5 percent significant level, information sharing affects sc partnership, with path coefficient 0.659 (p-value = 0.000). thus, this result support h1a is information sharing has a positive impact on sc partnership. the impact of information sharing on supply chain collaboration. path diagram results in fig. 3.1 shows that at 5 percent significant level, information sharing has no impact on sc collaboration, with path coefficient 0.311 (p-value = 0.077). thus, this result does not support h1b is information sharing has a positive impact on sc collaboration. the impact of information sharing on supply chain performance path diagram results in fig. 3.1 shows that at 5 percent significant level, information sharing has no impact on sc performance, with path coefficient 0.150 (p-value = 0.188). thus, these results not support h1b is information sharing has a positive impact on sc performance. the impact of partnership on supply chain collaboration path diagram results in fig. 3.1 shows that at 5 percent significant level, the partnership has a positive impact on sc collaboration, with path coefficient 0.513 (p-value = 0.005). thus, this result support h2a is partnership has a positive impact on sc collaboration. the impact of the partnership on supply chain performance path diagram results in fig. 3.1 shows that at 5 percent significant level, the partnership does not affect sc performance, with path coefficient 0.306 (p-value = 0.110). thus, these results do not support h2b is partnership has a positive impact on sc performance. the impact of collaboration on supply chain performance path diagram results in fig. 3.1 shows that at 5 percent significant level, collaboration has a positive effect on sc performance, with path coefficient 0.494 (p-value = 0.003). thus, this result support h3a is collaboration has a positive impact on sc performance. indirect effect testing table 3.5 shows the indirect effects of the research variables. indirect effect on table 7 show that at 5 per cent significant level, information sharing has indirect effect to the sc performance (t-statistic = 4.636; p-value = 0.000), information sharing has indirectly effect to the sc collaboration (t-statistics = 2.643; p-value = 0.008) and partnership has indirectly effect to the sc performance (t-statistics = 2.025; p-value = 0.043) 88 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) table 3.5. indirect effects original sample (o) sample mean (m) standard deviation (stdev) t statistics p values information -> partnership information -> performance 0.522 0.522 0.113 4,636 0,000 information -> collaboration 0.338 0.316 0.323 2,643 0.008 partnerships -> performance 0.253 0.237 0.125 2,025 0.043 partnerships -> collaboration collaboration -> performance 4. discussion this study uses pls to analyze the correlation between information sharing, partnerships, and collaboration on sc performance. the research was conducted at the apple agroindustry supply chain, batu city, indonesia. the results of the study did not find the direct effect of information sharing on sc performance, nor the direct effect of partnerships on sc performance. likewise, the direct effect of information sharing on sc collaboration was not found. the results of the study only indicate that there is an indirect effect of information sharing on sc collaboration, information sharing on sc performance, and partnerships on sc performance. this result is in line with the study from baihaqi and sohal [3], which states that information sharing does not have a direct relationship with organizational performance. this relationship is mediated by the practice of collaboration between supply chain actors. the absence of such correlations is typical for energy consumption since power grids possess constructive constraints and, hence, other tools have to be developed to improve performance of energy supply chains [30]. the present result is in line with the findings of zhao et al. [51], which states that information sharing significantly affects the supply chain performance, both in total costs and service levels. zhou & benton [52] also state that supply chain practices will be effective when the level of information sharing increases, which in turn will also improve supply chain performance. this result also supports the results of research by lin et al. [23], which states that information sharing can reduce demand uncertainty, which will improve supply chain performance. information sharing provides a large number of benefits, which in turn will increase the efficiency of supply chain performance in the manufacturing sector [24]. the effect of partnership on sc performance was following the results proposed by khan et al. [20] that buyer-supplier partnerships in the supply chain have a positive effect on supply chain performance. results of gallear et al. [15] also show that there is a positive relationship between partnership management and supply chain performance. wibowo and sholeh [44] also state that supplier partnerships are one of the supporting factors for supply chain performance in construction projects. the effect of collaboration on sc performance was following the findings of vereecke & muylle [41], who found empirically the relationship between sc collaboration and improved performance obtained. these empirical findings support the statement that proper collaboration with suppliers and customers provides benefits for improving performance. singhry et al. [36] found a significant relationship between scc and sc performance, which has been tested through covariance structural equation modeling. ramanathan and gunasekaran [31] also found that collaborative alliances improve sc performance, however, there are no results of research that show an indirect relationship of the effect of information sharing on sc performance. this condition can be caused by the population, which is mostly smes with limited resources. the process of information sharing is not carried out as done by large companies. influence of information sharing, partnership 89 copyright ©2019 assa. adv. in systems science and appl. (2019) this present study is subject to a few limitations that should be addressed in future research. first, although the number of samples is considered satisfactory for the research model using pls techniques, there are still opportunities to complete this study. second, with extensive sampling and covering more smes, the relationship model between variables can be further clarified. references [1] afonso, h., cabrita, m. do r. (2015) developing a lean supply chain performance framework in a sme: a perspective based on the balanced scorecard. procedia engineering 131, 270–9, doi: 10.1016/j.proeng.2015.12.389. [2] anthony, t. (2000) supply chain collaboration: success in the new internet economy. achieving supply chain excellence through technology, pp. 41–4. [3] baihaqi, i., sohal, a.s. (2013) the impact of information sharing in supply chains on organisational performance: an empirical study. production planning & control 24(8– 9), 743–58, doi: 10.1080/09537287.2012.666865. [4] cao, m., zhang, q. (2011) supply chain collaboration: impact on collaborative advantage and firm performance. journal of operations management 29(3), 163–80, doi: 10.1016/j.jom.2010.12.008. [5] capaldo, a., giannoccaro, i. (2015) how does trust affect performance in the supply chain? the moderating role of interdependence. international journal of production economics 166, 36–49, doi: 10.1016/j.ijpe.2015.04.008. [6] carr, a.s., pearson, j.n. (1999) strategically managed buyer–supplier relationships and performance outcomes. journal of operations management 17(5), 497–519, doi: 10.1016/s0272-6963(99)00007-8. [7] chauhan, s.s., proth, j.-m. (2005) analysis of a supply chain partnership with revenue sharing. international journal of production economics 97(1), 44–51, doi: 10.1016/j.ijpe.2004.05.006. [8] chen, m.-c., yang, t., yen, c.-t. (2007) investigating the value of information sharing in multi-echelon supply chains. quality & quantity 41(3), 497–511, doi: 10.1007/s11135-007-9086-2. [9] childerhouse, p., towill, d.r. (2003) simplified material flow holds the key to supply chain integration. omega 31(1), 17–27, doi: 10.1016/s0305-0483(02)00062-2. [10] christopher, m. (2011) logistics & supply chain management. 4th ed., pearson education, great britain: [11] costantino, f., di gravio, g., shaban, a., tronci, m. (2014) the impact of information sharing and inventory control coordination on supply chain performances. computers & industrial engineering 76, 292–306, doi: 10.1016/j.cie.2014.08.006. [12] daugherty, p.j., richey, r.g., roath, a.s., min, s., chen, h., arndt, a.d., genchev, s.e. (2006) is collaboration paying off for firms? business horizons 49(1), 61–70, doi: 10.1016/j.bushor.2005.06.002. [13] fawcett, s.e., osterhaus, p., magnan, g.m., brau, j.c., mccarter, m.w. (2007) information sharing and supply chain performance: the role of connectivity and willingness. supp chain mnagmnt 12(5), 358–68, doi: 10.1108/13598540710776935. [14] flynn, b.b., huo, b., zhao, x. (2010) the impact of supply chain integration on performance: a contingency and configuration approach. journal of operations management 28(1), 58–71, doi: 10.1016/j.jom.2009.06.001. [15] gallear, d., ghobadian, a., chen, w. (2012) corporate responsibility, supply chain partnership and performance: an empirical examination. international journal of production economics 140(1), 83–91, doi: 10.1016/j.ijpe.2012.01.016. [16] ghijsen, p.w.th., semeijn, j., ernstson, s. (2010) supplier satisfaction and commitment: the role of influence strategies and supplier development. journal of purchasing and supply management 16(1), 17–26, doi: 10.1016/j.pursup.2009.06.002. 90 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) [17] graham, t.s., daugherty, p.j., dudley, w.n. (1994) the long-term strategic impact of purchasing partnerships. international journal of purchasing and materials management (fall), 12–8, doi: https://doi.org/10.1111/j.1745-493x.1994.tb00269.x. [18] green, k.w., whitten, d., inman, r.a. (2008) the impact of logistics performance on organizational performance in a supply chain context. supp chain mnagmnt 13(4), 317–27, doi: 10.1108/13598540810882206. [19] kaipia, r., hartiala, h. (2006) information‐sharing in supply chains: five proposals on how to proceed. int jrnl logistics management 17(3), 377–93, doi: 10.1108/09574090610717536. [20] khan, s.a., liang, y., sumaira, s. (2015) the effect of buyer-supplier partnership and information integration on supply chain performance: an experience from chinese manufacturing industry. international journal of supply chain management 4(2). [21] kumar, g., banerjee, r.n. (2012) an implementation strategy for collaboration in supply chain: an investigation and suggestions. international journal of services and operations management 11(4), 407–27, doi: 10.1504/ijsom.2012.046077. [22] lambert, d.m., emmelhainz, m.a., gardner, j.t. (1996) developing and implementing supply chain partnerships. int jrnl logistics management 7(2), 1–18, doi: 10.1108/09574099610805485. [23] lin, f., huang, s., lin, s. (2002) effects of information sharing on supply chain performance in electronic commerce. ieee transactions on engineering management 49(3), 258–68, doi: 10.1109/tem.2002.803388. [24] lotfi, z., mukhtar, m., sahran, s., zadeh, a.t. (2013) information sharing in supply chain management. procedia technology 11, 298–304, doi: 10.1016/j.protcy.2013.12.194. [25] mohr, j., spekman, r. (1994) characteristics of partnership success: partnership attributes, communication behavior, and conflict resolution techniques. strategic management journal 15(2), 135–52, doi: 10.1002/smj.4250150205. [26] monczka, r.m., handfield, r.b., giunipero, l.c., patterson, j.l. (2008) purchasing and supply chain management. 4th ed., cengage learning. [27] monczka, r.m., petersen, k.j., handfield, r.b., ragatz, g.l. (1998) success factors in strategic supplier alliances: the buying company perspective*. decision sciences 29(3), 553–77, doi: 10.1111/j.1540-5915.1998.tb01354.x. [28] montoya-torres, j.r., ortiz-vargas, d.a. (2014) collaboration and information sharing in dyadic supply chains: a literature review over the period 2000–2012. estudios gerenciales 30(133), 343–54, doi: 10.1016/j.estger.2014.05.006. [29] mourtzis, d. (2011) internet based collaboration in the manufacturing supply chain. cirp journal of manufacturing science and technology 4(3), 296–304, doi: 10.1016/j.cirpj.2011.06.005. [30] popov, i., krylatov, a., zakharov, v., ivanov, d. (2017) competitive energy consumption under transmission constraints in a multi-supplier power grid system. international journal of systems science 48(5), 994–1001, doi: 10.1080/00207721.2016.1226986. [31] ramanathan, u., gunasekaran, a. (2014) supply chain collaboration: impact of success in long-term partnerships. international journal of production economics 147, 252–9, doi: 10.1016/j.ijpe.2012.06.002. [32] rokkan, a.i., heide, j.b., wathne, k.h. (2003) specific investments in marketing relationships: expropriation and bonding effects. journal of marketing research 40(2), 210–24. [33] ryu, i., so, s., koo, c. (2009) the role of partnership in supply chain performance. industr mngmnt & data systems 109(4), 496–514, doi: 10.1108/02635570910948632. influence of information sharing, partnership 91 copyright ©2019 assa. adv. in systems science and appl. (2019) [34] sahin, f., robinson, e.p. (2007) flow coordination and information sharing in supply chains: review, implications, and directions for future research. decision sciences 33(4), 505–36, doi: 10.1111/j.1540-5915.2002.tb01654.x. [35] simatupang, t.m., sridharan, r. (2005) the collaboration index: a measure for supply chain collaboration. int jnl phys dist & log manage 35(1), 44–62, doi: 10.1108/09600030510577421. [36] singhry, h.b., rahman, a.a., imm, n.s. (2015) measurement for supply chain collaboration and supply chain performance of manufacturing companies. international journal of economics and management 9, 1–22. [37] smirnova, m., henneberg, s.c., ashnai, b., naudé, p., mouzas, s. (2011) understanding the role of marketing–purchasing collaboration in industrial markets: the case of russia. industrial marketing management 40(1), 54–64, doi: 10.1016/j.indmarman.2010.09.010. [38] smith, g.e., watson, k.j., baker, w.h., pokorski ii, j.a. (2007) a critical balance: collaboration and security in the it-enabled supply chain. international journal of production research 45(11), 2595–613, doi: 10.1080/00207540601020544. [39] thakkar, j., kanda, a., deshmukh, s.g. (2009) supply chain performance measurement framework for small and medium scale enterprises. benchmarking 16(5), 702–23, doi: 10.1108/14635770910987878. [40] varma, t.n., khan, d.a. (2014) information technology in supply chain management. journal of supply chain management systems volume 3(3), 35–46. [41] vereecke, a., muylle, s. (2006) performance improvement through supply chain collaboration in europe. int jrnl of op & prod mnagemnt 26(11), 1176–98, doi: 10.1108/01443570610705818. [42] voudouris, v.t., consulting, a. (1996) mathematical programming techniques to debottleneck the supply chain of fine chemical industries. computers & chemical engineering 20, s1269–74, doi: 10.1016/0098-1354(96)00219-0. [43] whipple, j.m., lynch, d.f., nyaga, g.n. (2010) a buyer’s perspective on collaborative versus transactional relationships. industrial marketing management 39(3), 507–18, doi: 10.1016/j.indmarman.2008.11.008. [44] wibowo, m.a., sholeh, m.n. (2015) the analysis of supply chain performance measurement at construction project. procedia engineering 125, 25 – 31, doi: doi: 10.1016/j.proeng.2015.11.005. [45] wilson, d.t. (1995) an integrated model of buyer-seller relationships. journal of the academy of marketing science 23(4), 335, doi: 10.1177/009207039502300414. [46] wu, i.-l., chuang, c.-h., hsu, c.-h. (2014) information sharing and collaborative behaviors in enabling supply chain performance: a social exchange perspective. international journal of production economics 148, 122–32, doi: 10.1016/j.ijpe.2013.09.016. [47] xu, x., meng, z. (2014) coordination between a supplier and a retailer in terms of profit concession for a two-stage supply chain. international journal of production research 52(7), 2122–33, doi: 10.1080/00207543.2013.854940. [48] yang, j., wang, j., wong, c.w.y., lai, k.-h. (2008) relational stability and alliance performance in supply chain. omega 36(4), 600–8, doi: 10.1016/j.omega.2007.01.008. [49] yu, z., yan, h., cheng, e. (2001) benefits of information sharing with supply chain partnerships. industr mngmnt & data systems 101(3), 114–21, doi: 10.1108/02635570110386625. [50] zhang, h., wang, h.-c., zhou, m.-f. (2015) partnership management, supply chain collaboration, and firm innovation performance: an empirical examination. innovation science 7(2), 127–38, doi: 10.1260/1757-2223.7.2.127. 92 a. tutuhatunewa, surachman, p. b. santoso, i. santoso copyright ©2019 assa adv. in systems science and appl. (2019) [51] zhao, x., xie, j., zhang, w.j. (2002) the impact of information sharing and ordering co‐ordination on supply chain performance. supp chain mnagmnt 7(1), 24–40, doi: 10.1108/13598540210414364. [52] zhou, h., benton, w.c. (2007) supply chain practice and information sharing. journal of operations management 25(6), 1348–65, doi: 10.1016/j.jom.2007.01.009. microsoft word 992 article text, copyedited.doc adv syst sci appl 2021; 02; 29-41 published online at https://ijassa.ipu.ru. cooperation between smes and large industrial enterprises: russian case olga a. romanova1, evgeny a. kuzmin1,2*, marina v. vinogradova3, olga s. kulyamina3 1) institute of economics of the ural branch of russian academy of sciences, yekaterinburg, russia 2) ural state university of economics, yekaterinburg, russia 3) russian state social university, moscow, russia abstract: in the modern global economy, the system of functional cooperative relations that lead to the formation of hybrid structures – production networks has become actively spread. however, the strength and scale of cooperation are not uniform. using the example of russia, the authors consider the effectiveness of the production network of cooperation between small and mediumsized enterprises with large companies. we tested a number of hypotheses based on common ideas about the effects of cooperation. empirical results make it possible to clarify the mechanism of formation and features of interfirm production chains in russia. the “anchor” role of large enterprises with state participation as centers of cooperation formation is noted. in the course of the study, 14 enterprises were selected, distributed across key sectors of the russian economy. statistical and correlation analysis methods were used to evaluate the effects of cooperation. the results showed that the orders placed by large manufacturing enterprises with small and mediumsized enterprises increased over the period of 2015–2019. “anchor” enterprises, as a rule, reduce the production localization degree. however, this does not have a significant impact on improving the profitability of their activities, and also does not depend on the share of state participation. besides, placing orders with small and medium-sized enterprises does not allow them to reduce the number of employees. many of the expected internalities that are characteristic of cooperative relations in developed countries are not reflected in the specifics of the russian economy, or their manifestation is limited. the russian experience clearly demonstrates the weakness of cooperative partnership, although with positive trends of change. there is a need to further improve the mechanisms for supporting small and medium-sized enterprises in the production sector, aimed at creating sustainable networks. the proposed approach can be applied to assess inter-firm production chains in other countries. a comparative study will determine the strength of the formation of production networks across countries, which will expand the understanding of the economic processes of networkization. keywords: cooperation, globalization, small and medium-sized businesses, production networks, subcontracting, specialization, production chains 1. introduction transformations in the economy that occur under the influence of globalization processes encourage enterprises to look for effective forms of organization of production activities. such forms of integration as subcontracting, franchising, leasing, venture financing, technology parks, joint ventures, tolling, etc. are becoming more and more popular [1]. the system of functional production of cooperative relations has become widely used. bringing a small number of specialized companies in private manufacturing businesses provides the flexibility to respond to market conditions. one of the main factors determining cooperation is the creation of useful communication mechanisms, which are usually carried out in the form of subcontracting processes. * corresponding author: kuzmin.ea@uiec.ru 30 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) the closeness of ties and the depth of cooperation between industrial enterprises are not uniform. in russia, this is facilitated by the procurement system of state-owned companies, which obliges large enterprises to purchase from small and medium-sized enterprises (smes), which in turn can be carried out by transferring part of the production cycle to a subcontractor (subcontract). the volume of purchases from smes, including purchases in which the contractor must engage the sme as a subcontractor, must be at least 20% of the total annual value of contracts concluded by customers based on the results of purchases. at the same time, at least 18% of the total annual value of contracts should be allocated to procurement involving only smes [2]. however, these measures have not yet significantly changed the role of smes in the economy as a whole. according to the analytical center for the government of the russian federation, in comparison with foreign countries, the share of public procurement by smes is low is about 1–5% against 20%. the reasons for these discrepancies are both institutional and structural [3]. it is obvious that for sustainable economic development, it is necessary to search for optimal mechanisms for expanding cooperation between industrial enterprises of various levels. therefore, the purpose of the study is to assess the effects of cooperation between smes and large industrial enterprises in russia. to do this, the authors analyzed the purchasing activities of large enterprises with state participation in key russian sectors: oil, gas, minerals and related activities (petroleum products transportation); industrial production, and electric-power supply industry. preliminary data indicate that the degree of cooperation between enterprises is relatively low. in this regard, the authors put forward several assumptions that should explain this feature. the identified impact factors will allow understanding better the mechanism of forming production chains through cooperation and assessing its potential in the russian economy. 2. literature review the relationship between competition and cooperation is one of the key issues in the strategic management theory. this is reflected in the variety of approaches that explain the pros and cons of various strategies for organizing the production process of enterprises in their interaction with the external environment [4-7]. cooperation develops along the entire value chain. preference is given to all possible contractual forms in comparison with intra-company integration. the intensive growth of industrial cooperation raises questions about the blurring of lines of the economic agent, the formation of hybrid structures, which are increasingly referred to as networks. the phenomenon of intercompany network relations, which has become widespread in recent years, attracts researchers who are trying to explain the reasons for its occurrence [8]. in the most general terms, intercompany networks are perceived as a way to regulate the interdependence between companies. it should be taken into account that initially, the definitions of intercompany networks differ both in the terminology used and in the emphasis [9-11]; the objective and research direction are the decisive factors. the development of network production cooperation with the participation of smes is presented in the works of many scientists [12-14]. mins and schneider [15] define the transformation of the world economy and the principles of business introduction as metacapitalism. justifying this idea, among the reasons, along with globalization, integration of global capital markets, the spread of information and communication technologies and ebusiness, they pointed to the fundamental restructuring of companies, which led not only to their reengineering of business processes but also contributed to the creation of transnational production networks [15]. as the modern world practice shows, it is necessary to use production cooperation in its newest forms for the sustainable development of industrial enterprises. petrishcheva [16] cooperation between smes and large industrial enterprises 31 copyright ©2021 assa. adv. in systems science and appl. (2021) attempts to set the concept of industrial cooperation and points out the potential for its development. it is generally agreed [17-21] that one of the most promising organizational forms of integration of small, medium, and large enterprises is subcontracting. this form of cooperation is designed to use a wide network of suppliers [16, 22-23]. more generally, subcontracting refers to a specific aspect of the organization of industrial production, in which large and small firms coexist (with a high degree of specialization) in production, and sometimes in making investment decisions [24]. in fact, a large enterprise transfers part of its production functions, which ultimately reduces inventory and optimizes the production process, focusing on the assembly of the final product and quality control [25]. the use of this mechanism leads to a reduction in capital investment in the means of production and a reduction in the number of people employed in production [26]. as a result, subcontracting helps to diversify business risks and regulate production levels more flexibly [27-28]. some scientists [29-31] have suggested that cost minimization is the main explanation for subcontracting production. at the same time, tijun et al. [32] stated that the main idea of subcontracting went beyond cost minimization. they justified the idea that the key factors influencing the use of this production strategy were the focus on the core business, access to the professional and technological capabilities of partners, and the release of internal resources. according to handfield and nichols [33], manufacturers can match future product needs with existing resources through close collaboration with key suppliers. in developed countries, industrial cooperation is a tool for improving the efficiency of industrial production and ensuring overall economic growth [34-36]. since smes are the initiators of many innovations and provide the basis for sustainable economic development [37], they must be adequately protected in order to survive in the industrial market [38-39]. this can be achieved through government policies that encourage industrial sectors to increase the pace of production cooperation with smes. one of the mechanisms is the regulation of public procurement. the topic of supporting small and medium-sized businesses in the procurement system is very popular [40]. some scientists [41-43] consider the mechanism of granting preferences through the introduction of special procedures involving smes to be economically unjustified. support methods include several mechanisms that differ from state to state. in the united states, information and consulting support are provided, procurement quotas are set, some contracts are subcontracted to small and medium-sized businesses, and innovative developments are supported. many eu countries use a simplified procedure for purchasing goods, works, and services at the expense of budget funds, and provide for the transfer of part of contracts to small businesses, but do not carry out procedures in which only small and medium-sized businesses participate. the question of the effectiveness of public procurement procedures, in which only smes can participate, attracts the attention of many scientists, but their conclusions are ambiguous [42, 44]. the review has shown that the economy is no longer contrasting small and medium-sized businesses with large businesses, and their relations are transformed, moving to a new stage of development. modern forms of cooperation are gradually changing the philosophy of inter-company communications. the target function of cooperation is to expand external economic relations – to create a production network. 3. materials and methods the initial analysis of trends in industrial cooperation and the expected effects of its implementation in practice allowed forming the following main hypotheses, which will be tested in the course of the study: h1. despite the legal requirements for the share of public procurement from smes, the development of cooperation between small enterprises and large state-owned companies is 32 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) inefficient. in this regard, it is assumed that the share of orders transferred to the sme sector within the framework of production cooperation does not depend on the share of state participation in the capital of large industrial enterprises. h2. “anchor” industrial enterprises that are more actively involved in industrial cooperation have higher indicators of production efficiency, in particular, the indicator of gross profitability and profitability of outlay. h3. placing orders with smes allows large industrial enterprises with state participation to reduce the number of employees. to assess the degree of development of cooperation between industrial enterprises and smes in russia, “anchor” large enterprises were selected, divided into key industries: oil, gas, minerals and related activities (petroleum products transportation); industrial production, and electric power (see appendix a). the list includes such companies as bashneft, vankorneft, gazprom, rosneft, russian helicopters, ngo almaz, united aircraft corporation (uac), united engine corporation (uec), eastern energy company (eec), mosenergo, rosseti moscow region (moesk), rosseti, rushydro, and transneft. the choice of “anchor” companies is determined by their contribution to the russian economy. all these companies are among the top 100 largest companies in russia, their revenue for the last report in 2019 varies from 39 billion rubles to 4.8 trillion rubles. besides, all industrial enterprises have a share of state participation. this feature is also an area of research restrictions. in addition to the selection of industrial enterprises with state participation, the analysis of the degree of cooperation with smes is carried out only among legal entities (companies), individual entrepreneurs and individuals are ignored (although they perform work, provide services and produce products for “anchor” companies). the study used the following indicators of enterprises: revenue, cost of production, works (services); the amount and share of revenue of large industrial enterprises attributable to smes; the amount and share of the cost of production, works (services) of large industrial enterprises attributable to smes. the data panel was supplemented with information on the volume of purchases from smes, the share of state participation in the capital of industrial enterprises, and the number of employees. the sources of information were the news agencies spark-interfax and interfax corporate information disclosure center [45]. statistical and correlation analysis methods were used for a comprehensive assessment of the effects of sme cooperation with large industrial enterprises in russia. the production localization degree (a1) was calculated as part of the statistical approach: , (1) where is the orders of large industrial enterprises placed with smes; is the cost of products, works (services) of large industrial enterprises. the share of revenue (a2) and the share of cost (a3) of large industrial enterprises attributable to smes were determined using the formulas: , (2) , (3) where is the cost of large industrial enterprises attributable to smes; is the revenue of large industrial enterprises; is the revenue of large industrial enterprises attributable to smes. in order to compare the performance of enterprises, the authors used the profitability indicators – gross profit margin (gpm) and outlay (po): smes 1 orders lie a pc = smesorders liepc smes 2 re relie a = smes 3 lie pca pс = smespc relie smesre cooperation between smes and large industrial enterprises 33 copyright ©2021 assa. adv. in systems science and appl. (2021) , (4) , (5) where gp is the gross profit; re is the revenue; oi is the operating profit; pc is the production, works (services) cost. the pearson correlation ratio was calculated to determine the strength of the statistical relationship between the indicators: , (6) where n is the sample size, , are the mean values of parameters; , are the variances of parameters; , are the root-mean-square (standard) deviations of parameters. 4. results and discussion during the observation period from 2015 to 2019, all selected large industrial enterprises with state participation increased the volume of orders placed with smes by several times from 2.15 (rosneft) to 210.8 (transneft). data on order dynamics is provided in appendix b. at the same time, the level of internal production localization (autonomy) for all the considered enterprises decreased (table 1), and enterprises increased the share of orders placed with smes in the cost price. however, the degree of localization varied heterogeneously over the period under review. table 1. production localization degree of large industrial enterprises with state participation in 2015–2019, % enterprise 2015 2016 2017 2018 2019 bashneft 8.42 9.38 6.89 2.92 79.53 vankorneft 4.09 61.35 0.92 0.35 74.85 russian helicopters 10.97 3.13 8.52 33.22 134.75* gazprom 0.34 0.57 0.31 0.03 5.50 eec 0.55 0.65 0.63 0.78 36.01 mosenergo 1.68 1.56 1.72 0.30 75.15 moesk 13.63 7.639 8.99 23.76 78.22 ngo almaz 1.94 1.13 1.13 0.11 73.15 uac 8.76 3.60 5.07 0.60 67.96 uec 1.85 1.80 3.99 0.59 51.99 rosneft 5.43 1.55 0.26 0.02 6.40 rosseti 16.41 9.86 10.26 1.62 735.63* rushydro 17.61 15.44 8.37 4.27 157.79* transneft 0.04 0.71 0.32 0.25 7.01 note: * – for holding companies, the cost of production, works (services) has a heterogeneous distribution within the group. source: [45]. based on the collected data, the share of revenue and the share of the cost of “anchor” large industrial enterprises accounted for by smes was calculated – separately for mediumsized, small, and micro enterprises according to the classification adopted in russia [46]. the criteria for a medium-sized enterprise are that the average number of employees is not more gpgpm re = oipo pc = 1 ( )( ) ( 1) n i i i xy x y x x y y r n s s = = å x y 2 xs 2 ys xs ys 34 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) than 250, and the annual income is not more than 2 billion rubles. the share of organizations in the capital of medium-sized enterprises that are not related to smes should not exceed 49%, the share of the state, regions, or non-profit organizations shall not exceed 25%. the small business criteria are that the average number of employees is not more than 100, and the income is not more than 800 million rubles. the micro-enterprise criteria are the average number of employees no more than 15 and annual income no more than 120 million rubles. restrictions on the structure of the authorized capital are similar. the largest share of orders for medium-sized enterprises in the revenue of large industrial enterprises in revenue in 2019 was observed in rosseti (48.48%), uac (39.6%), rushydro (39.69%). the largest share of orders for medium-sized enterprises in the cost of large industrial enterprises was observed in russian helicopters (36.75%), rushydro (31.38%), and moesk. gazprom, transneft, and rosneft placed the smallest share of orders with medium-sized businesses (fig. 1). fig. 1. shares of revenue and cost of large industrial enterprises attributable to medium-sized enterprises in 2019, % source: [45, 47]. a similar situation is observed for small businesses. the largest share of orders for small businesses in the revenue of large industrial enterprises in 2019 was recorded in russian helicopters (39.12%), rushydro (44.79%), rosseti (33.28%). mosenergo (36.95%), rushydro (32.33%), and bashneft (26.16%) accounted for the largest share of orders for small enterprises in the cost of large industrial enterprises. gazprom, rosneft, and transneft demonstrated the worst work with small enterprises (fig. 2). cooperation between smes and large industrial enterprises 35 copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 2. share of revenue and cost of large industrial enterprises attributable to small enterprises in 2019, % source: [45, 47]. gazprom (0.30%), rosneft (0.19%), and transneft (0.32%) showed a low share of the revenue from large industrial enterprises attributable to micro-enterprises. such a low share of micro-enterprise participation is also observed in the cost of these companies (fig. 3). rushydro (10.99%), ngo almaz (7.04%), and eec (6.16%) show the best positions in working with micro-enterprises in terms of the revenue share. in terms of cost, the degree of participation of these companies is close – rushydro (7.32%), mosenergo (6.66%), and ngo almaz (5.11%). fig. 3. shares of revenue and cost of large industrial enterprises attributable to micro-enterprises in 2019, % source: [45, 47]. 36 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) in general, the leaders in placing orders for the sme sector in 2019 were such enterprises as rushydro (95.46% of revenue), rosseti (86.90% of revenue), russian helicopters (80.76% of revenue), mosenergo (68.21% of revenue), moesk (68.22% of revenue), ngo almaz (66.64% of revenue), and uac (65.91% of revenue). the smallest share of orders attributable to smes in revenue is observed in the “anchor” companies in the oil and gas sector. the average order amount placed with one small enterprise in 2019 was 249,114.9 thousand rubles, with one medium-sized enterprise – 943,719.6 thousand rubles, and with one micro-enterprise – 30,243.6 thousand rubles. thus, the structure of the distribution of orders for smes is dominated by medium-sized enterprises. recall that medium-sized enterprises are characterized by the presence of an average number of employees up to 250 people and an annual income (revenue) of no more than 2 billion rubles. the calculated correlation ratio between the parameter of the share of state participation in large industrial enterprises and the share of orders attributable to the sme sector is -0.04. on a chedoke scale [48], the reverse relationship is weak; therefore, the hypothesis h1 was confirmed. this means that the largest industrial state-owned companies prefer not to place orders with smes, which is due to both the low level of development of the sme sector in russia and the industry specifics of large businesses (the largest russian production companies belong to the oil and gas industry), while smes are mainly concentrated in the services and trade sectors. the correlation ratio between the indicator of the share of orders from smes of large industrial enterprises in the cost price and the gross profit margin is 0.1469. the data on profitability and the share of cost attributable to smes are given in appendix c. the results obtained allow concluding that the correlation is direct and weak. the first part of the h2 hypothesis about the manifestation of greater activity in production cooperation on the part of “anchor” enterprises was not confirmed. this means that large industrial enterprises do not sufficiently use the advantages of production cooperation with smes to increase efficiency by reducing production costs. the correlation ratio between the indicator of the share of orders from smes of large industrial enterprises in the cost and profitability of outlay is 0.2924. the correlation is also direct and weak, and the second part of the h2 hypothesis is not confirmed. the h3 hypothesis implements the assumption that placing orders with smes allows “anchor” industrial enterprises with state participation to reduce the number of employees. the analysis of the growth rate of the average number of employees of the majority of large industrial enterprises under consideration for 2015/2019 leads to the conclusion that despite the increase in the level of placing orders with smes, the average number of employees has increased. the largest growth rate in the number of employees is observed in the uec (40 times), which is due to the reorganization of the enterprise. taking the uec observation as a statistical outlier, the average rate of growth in the number of employees of large enterprises is 107.61% (appendix c). the correlation ratio between the indicator of the share of orders from smes of large industrial enterprise in the cost price and the growth rate of the average number of employees is -0.2023. the relationship is inverse and weak, so the h3 hypothesis is not confirmed. based on the study, it can be concluded that production cooperation in russia between “anchor” large industrial enterprises with state participation through the use of subcontracting with the sme sector is not carried out effectively. many of the expected internalities that are characteristic of cooperative relations in developed countries, both for contractors and subcontractors, are not reflected in the specifics of the russian economy or their manifestation is limited. cooperation between smes and large industrial enterprises 37 copyright ©2021 assa. adv. in systems science and appl. (2021) 5. conclusion the development of specialization and cooperation of small, medium and large enterprises in the modern conditions of the global market is becoming an economic necessity and is a consequence of the new competitiveness paradigm [49]. this statement finds convincing arguments in world practice. production cooperation is formed along the entire value chain and leads to the emergence of a new phenomenon of intercompany relations – production networks. however, the strength and scale of cooperation are not uniform. the russian experience clearly demonstrates the weakness of cooperative partnership, although with positive trends of change. the results indicate that the level of internal production localization (autonomy) of large industrial enterprises with state participation in russia in 2015–2019 decreased; all these enterprises increased the share of orders placed with smes. therefore, there is an expansion of subcontracting as a form of industrial cooperation. the largest share of orders placed with smes is observed in large industrial enterprises in the electric power supply industry, and the smallest – in the oil and gas industry. in terms of the volume of orders placed by “anchor” large industrial enterprises, the leaders are medium-sized enterprises; the number of orders placed with such enterprises is 3.8 times higher than the number of orders placed with small businesses, and 31.2 times – with micro-enterprises. at the same time, the hypothesis that placing orders with smes allows large industrial enterprises with state participation to reduce the number of employees has not been confirmed, as well as the hypothesis that “anchor” industrial enterprises, which are more actively involved in industrial cooperation, have higher production efficiency indicators. this allowed concluding that cooperation is inefficient among large companies with state participation in russia. references 1. asaul, a.n., skumatov, e.g., & lokteeva, g.e. (2004). methodological aspects of formation and development of business networks: study guide. st. petersburg: gumanistika. (p. 55). 2. decree of the government of the russian federation no. 1352 “on features of participation of subjects of small and medium business in procurement of goods, works, services by separate types of legal entities”. (2014, december 11). retrieved september 24, 2020, from www.pravo.gov.ru 3. chernova, v.y., golodova, z.g., degtereva, e.a., zobov, a.m., & starostin, v.s. (2018). russian industry in global value-added chains. european research studies journal, 21(3), 165–178. https://doi.org/10.35808/ersj/1051 4. golovanova, s.v., avdasheva, s.b., & kadochnikov, s.m. (2010). intercompany cooperation: analysis of cluster development in russia. russian journal of management, 8(1), 41–66. 5. katkalo, v.s. (2006). evolution of the strategic management theory. st. petersburg: publishing house of st. petersburg state university. 6. uhlig, c. (1980). industrial cooperation as an instrument of development policy. intereconomics, 15, 188–193. https://doi.org/10.1007/bf02930851 7. zobov, a.m., degtereva, e.a., starostin, v.s., & chernova, v.y. (2016). innovative strategies of transnational companies and synergy effect of technologization. indian journal of science and technology, 9(39). https://doi.org/10.17485/ijst/2016/v9i39/96521 8. tretyak, o.a., & rumyantseva, m.n. (2003). network forms of intercompany cooperation: approaches to explaining the phenomenon. russian journal of management, 1(2), 25–50. 9. gerlach, m.l. (1992). the japanese corporate network: a block model analysis. administrative science quarterly, 37(1), 105–139. https://doi.org/10.2307/2393535 10. granovetter, m. (1985). economic action and social structure: the problem of embeddedness. american journal of sociology, 91, 481–510. https://doi.org/10.1086/228311 38 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) 11. miles, r.e., snow, c.c., mathews, j.a., & coleman, h.j. (1997). organizing in the knowledge area: anticipating the cellular form. academy of management executive, 11(4), 7–20. https://doi.org/10.5465/ame.1997.9712024836 12. galaso, p., & kovářı́k, j. (2018). collaboration networks and innovation: how to define network boundaries. journal of eurasian economic dialogue, 3(2), 1–17. 13. lambert, t.e. (2017). monopoly capital and entrepreneurism: whither small business? journal of eurasian economic dialogue, 2(3), 1–12. 14. lambert, t.e. (2018). monopoly capital and innovation: an exploratory assessment of r&d effectiveness. journal of eurasian economic dialogue, 3(6), 20-35. 15. mins, gr., & schneider, d. (2001). metacapitalism and revolution in electronic business: what will companies and markets be like in the 21st century. moscow: alpina business books. (p. 280). 16. petrishcheva, i.v. (2011). industrial cooperation in the context of interaction between small and large enterprises: essence and forms. almanac of modern science and education, 1(44), 168–170. 17. chen, l. (2011). the governance and evolution of local production networks in a cluster: the case of taiwan’s machine tool industry. geojournal, 76, 605–622. https://doi.org/10.1007/s10708-009-9317-2 18. kimura, f. (2002). subcontracting and the performance of small and medium firms in japan. small business economics, 18, 163–175. https://doi.org/10.1023/a:1015187507379 19. kuriyama, n. (2017). the formation of industrial subcontracting in the japanese manufacturing industry. in japanese human resource management (palgrave macmillan asian business series). cham: palgrave macmillan. https://doi.org/10.1007/978-3-31943053-9_4 20. park, y., shin, j., & kim, t. (2010). firm size, age, industrial networking, and growth: a case of the korean manufacturing industry. small business economics, 35, 153– 168. https://doi.org/10.1007/s11187-009-9177-7 21. sprietsma, h.b. (1978). international subcontracting and developing countries. de economist, 126, 220–242. https://doi.org/10.1007/bf01770341 22. aydinliyim, t., & vairaktarakis, g.l. (2011). sequencing strategies and coordination issues in outsourcing and subcontracting operations. in k. kempf, p. keskinocak, & r. uzsoy (eds.), planning production and inventories in the extended enterprise. international series in operations research & management science, vol. 151. new york, ny: springer. https://doi.org/10.1007/978-1-4419-6485-4_12 23. sidorenko, e.e., & chizhova, e.n. (2009). production outsourcing as a form of interaction between large and small industrial enterprises: monograph. belgorod: bstu publishing house. (p. 126). 24. nagaraj, r. (1984). sub-contracting in indian manufacturing industries: analysis, evidence and issues. economic and political weekly, 19(31/33), annual number: focus on industry, 1435–1453. 25. analytical center for the government of the russian federation. (2018). subcontracting as a mechanism for small business development. bulletin on competition development, 22. retrieved july 3, 2020, from https://ac.gov.ru/files/publication/a/17284.pdf 26. litovchenko, i.l. (2016). economic efficiency of subcontracting in production cooperation. economic studies, 5(13), 169–176. 27. bamberg, g. (1987). risk sharing and subcontracting. in g. bamberg, & k. spremann (eds.), agency theory, information, and incentives. berlin, heidelberg: springer. https://doi.org/10.1007/978-3-642-75060-1_4 28. kashbraziev, r.v. (2014). quantitative approach to risk assessment of international cooperation. finance and credit, 46(622), 45–58. 29. abraham, k.g., & taylor, s.k. (1996). firms’ use of outside contractors: theory and evidence. journal of labor economics, 14(3), 394–424. https://doi.org/10.1086/209816 cooperation between smes and large industrial enterprises 39 copyright ©2021 assa. adv. in systems science and appl. (2021) 30. holmes, j. (1986). the organization and locational structure of production subcontracting. in a.j. scott, & m. storper (eds.), production, works, and territory: the geographical anatomy of industrial capitalism (pp. 80–106). boston; london and sydney: allen and unwin. 31. mcmillan, j., (1995). reorganizing vertical supply relationships. in h. siebert (ed.), trends in business organization: do participation and cooperation increase competitiveness. university of michigan. 32. tijun, f., sandal, l., jiehong, k, & dandan, l. (2009). a survey and analysis of outsourcing in east china. norwegian school of economics and business administration. 33. handfield, r.b., & nichols, e.l (2003). reorganizing supply chains. forming integrated systems of value creation (trans. from english, pp. 59–82). moscow: williams. 34. litau, e. (2018). entrepreneurship and economic growth: a look from the perspective of cognitive economics. in iceme 2018: proceedings of the 2018 9th international conference on e-business, management and economics (pp. 143–147). https://doi.org/10.1145/3271972.3271978 35. lowry, a. (2020). industrial policy and technological cooperation in the eaeu: the case of eurasian technology platforms. in m. broad, & s. kansikas (eds), european integration beyond brussels. security, conflict and cooperation in the contemporary world. cham: palgrave macmillan. https://doi.org/10.1007/978-3-030-45445-6_9 36. pochukaeva, o.v. (2006). the interindustry cooperation of production complexes as an economic growth factor. studies on russian economic development, 17, 182–186. https://doi.org/10.1134/s1075700706020067 37. toomsalu, l., tolmacheva, s., vlasov, a., & chernova, v. (2019). determinants of innovations in small and medium enterprises: a european and international experience. terra economicus, 17(2), 112–123. https://doi.org/10.23683/2073-6606-2019-17-2-112-123 38. litau, e. (2017). ‘evolution of species’ in business: from mice to elephants. the question of small enterprise development. journal of advanced research in law and economics, 8(6), 1812–1824. https://doi.org/10.14505/jarle.v8.6(28).16 39. litau, e. (2018). the information problem on the way to becoming a “gazelle”. in proceedings of the european conference on innovation and entrepreneurship, ecie, 2018september (pp. 394–401). 40. eom, s., kim, s., & jang, w. (2015). paradigm shift in main contractorsubcontractor partnerships with an e-procurement framework. ksce journal of civil engineering, 19, 1951–1961. https://doi.org/10.1007/s12205-015-0179-5 41. kidalov, m., & snider, k. (2011). us and european public procurement policies for small and medium-sized enterprises (sme): a comparative perspective. business and politics, 13(4), 1–41. https://doi.org/10.2202/1469-3569.1367 42. nakabayashi, j. (2013). small business set-asides in procurement auctions: an empirical analysis. journal of public economics, 100, 28–44. https://doi.org/10.1016/j.jpubeco.2013.01.003 43. nicholas, c., & fruhmann, m. (2014). small and medium-sized enterprises policies in public procurement: time for a rethink? journal of public procurement, 14(3), 328–360. https://doi.org/10.1108/jopp-14-03-2014-b002 44. krasnokutskaya, e., & seim, k. (2011). bid preference programs and participation in highway procurement auctions. american economic review, 101(6), 2653–2686. https://doi.org/10.1257/aer.101.6.2653 45. spark-interfax. (2020). retrieved september 24, 2020, from https://www.sparkinterfax.ru/ 46. federal law no. 209-fz “on the development of small and medium-sized businesses in the russian federation”. (2007, july 24). retrieved september 24, 2020, from http://www.pravo.gov.ru 40 o.a. romanova, e.a. kuzmin, m.v. vinogradova, o.s. kulyamina copyright ©2021 assa. adv. in systems science and appl. (2021) 47. center for corporate information disclosure interfax. (2020). retrieved september 24, 2020, from https://www.e-disclosure.ru/ 48. ishkhanyan, m.v., & karpenko, n.v. (2016). econometrics. part 1. dual regression. moscow: msru (miit). (p. 117). 49. heywood, j.b. (2002). outsourcing: in search of competitive advantages (trans. from english). moscow: williams. (p. 174). appendix a selected large industrial enterprises of russia for the analysis of industrial cooperation with smes enterprise activity type by okved* share of state participation in the capital, % average number of employees, people revenue in 2019, thousand rubles cost of sales in 2019, thousand rubles 1. production of oil, gas, and minerals bashneft 06.10.1 60.5 9,183 703,150,528 514,466,674 vankorneft 06.10.1 0.01 1,600 383,329,128 308,750,775 gazprom 46.71 50.0 26,691 4,758,711,459 2,657,654,354 rosneft 06.10.1 40.4 4,553 6 827 526 407 4 782 222 071 2. production russian helicopters 30.30.3 85.71 427 39,853,657 23,884,842 ngo almaz 26.30.17 1.16 11,387 101,586,138 92,535,194 united aircraft corporation (uac) 30.30.3 8.99 661 54,734,083 53,083,289 united engine corporation (uec) 30.30.13 87.45 14,297 94,038,976 63,187,526 3. electric-power supply industry eastern energy company (eec) 35.14 0.08 3,962 97,746,207 90,981,227 mosenergo 35.11.1 26.4 7,922 189,781,589 172,256,268 rosseti moscow region (moesk) 35.12 88.4 14,377 160,375,521 139,860,598 rosseti 35.12 88.4 642 39,434,924 4,658,385 rushydro 35.11.2 62.2 5,396 155,180,091 93,884,445 4. other activities transneft 49.50.1 78.55 1,257 960,811,881 787,367,559 note: * is the russian national classifier of economic activities (okved); 06.10.1 is the extraction of crude oil; 46.71 is the wholesale of solid, liquid and gaseous fuels and related products; 30.30.3 is the production of helicopters, planes and other flying vehicles; 26.30.17 is the production of radio and television transmitting equipment; 30.30.13 is the manufacture of jet engines and their parts; 35.14 is the power-supply trade; 35.11.1 is the electricity production by thermal power plants, including activities for ensuring the operability of power plants; 35.12 is the transmission of electricity and technological connection to distribution power grids; 35.11.2 is the production of electricity by hydroelectric power stations, including activities for ensuring the operability of power plants; 49.50.1 is the transportation via pipelines of crude oil and petroleum products. source: [45, 47]. appendix b volume of orders of large industrial enterprises with state participation from smes in 2015–2019, thousand rubles enterprise 2015 2016 2017 2018 2019 change 2015/2019 bashneft 26,710,063 29,465,010 26,231,483 14,708,667 409,149,836 +1,532% vankorneft 8,113,407 114,701,623 2,325,058 1,127,552 231,096,848 +2,848% russian helicopters 668,099 531,842 1,154,521 6,176,591 32,185,779 +4,818% gazprom 7,646,833 12,769,846 7,773,306 816,985 146,174,290 +1,912% cooperation between smes and large industrial enterprises 41 copyright ©2021 assa. adv. in systems science and appl. (2021) enterprise 2015 2016 2017 2018 2019 change 2015/2019 eec 249,589 311,905 331,445 620,313 32,758,637 +13,125% mosenergo 2,509,753 2,586,024 2,800,891 509,229 129,458,616 +5,158% moesk 15,489,675 9,472,757 12,216,046 33,163,019 109,405,277 +706% ngo almaz 495,657 555,598 534,157 59,048 67,692,533 +13,657% uac 4,262,882 1,795,264 2,608,787 310,220 36,075,422 +846% uec 492,931 531,626 546,298 191,144 32,849,933 +6,664% rosneft 141,847,000 44,426,939 8,977,148 1,021,700 306,224,474 +216% rosseti 712,467 424,912 461,972 75,382 34,268,268 +4,810% rushydro 11,263,353 8,701,764 7,017,781 4,135,669 148,137,273 +1,315% transneft 261,718 4,993,648 2,317,868 2,022,262 55,168,592 +21,079% source: [45]. appendix c data on gross profit margin, the profitability of outlay and the share of cost attributable to smes for selected large industrial enterprises in russia for 2019, % enterprise profitability share of cost attributable to smes growth rate of the average number of employees 2015/2019 gross profit outlay bashneft 26.69 11.03 55.96 121.12 vankorneft 19.46 19.06 50.08 35.56 russian helicopters 40.07 33.84 65.15 66.51 gazprom 44.15 15.23 2.24 107.45 eec 6.92 2.61 25.41 116.09 mosenergo 9.23 10.17 63.39 99.99 moesk 12.79 4.71 34.18 95.48 ngo almaz 8.91 9.59 54.12 248.57 uac 3.02 -4.66 45.14 105.42 uec 32.81 20.85 22.72 4016.01 rosneft 29.86 12.51 3.29 111.51 rosseti 88.19 746.54 62.09 98.17 rushydro 39.5 65.29 71.03 95.76 transneft 18.05 12.24 4.19 97.29 source: [45]. adv syst sci appl 2021; 03:75–90 published online at https://ijassa.ipu.ru. applications of sine-cosine wavelets method for solving drinfel’d–sokolov–wilson system naser azizi1, reza pourgholi1* 1school of mathematics and computer science, damghan university, damghan, iran abstract: in this article, we use the sine-cosine wavelets (scws) method to numerically solve the drinfel’d–sokolov–wilson (dsw) system. for this purpose, we use an approximation of functions with the help of scws, and we approximate spatial derivatives using this method. the operational matrix based on scws has a large number of zero components, which ensures good system performance and provides acceptable accuracy even with fewer collocation points. in the end, to show the effectiveness and accuracy of the method in solving this system one numerical example is provided. keywords: drinfel’d–sokolov–wilson system, numerical method, sine-cosine wavelets method, operational matrix 1. introduction nonlinear coupled partial differential equations (pdes) are very significant in a type of scientific field, especially in fluid mechanics, solid-state physics, plasma waves, plasma physics, and chemical physics. since many nonlinear physical phenomena can be explained by the exact and numerical solutions of nonlinear equations, the attempt for finding the exact and numerical solutions to these phenomena is important. in this article, our main goal is to solve numerically the dsw system. a generalized form of the dsw system is given by:{ ψt + αφφx = 0, φt + βφxxx + γψφx + δψxφ = 0, (1.1) where α, β, γ, and δ are some nonzero parameters. system (1.1) plays an important role in fluid dynamics [8,12] and is originally introduced by drinfel’d and sokolov [7] and wilson [24] as a model of dispersive water waves. many researchers have devoted considerable efforts by successfully implementing various methods to extract solitary wave solutions and other solutions of dsw system [1, 10, 17, 23, 25]. one way to solve equations numerically is to use wavelets. the basic idea of wavelets goes back to the early 1960s [4,5]. there are developments concerning the multiresolution analysis algorithm based on wavelets [6] and the construction of compactly supported orthonormal wavelet bases [16]. so far, several problems have been solved numerically using different wavelets, for example, we can refer to references [2, 3, 9, 13, 18, 19, 26, 27]. in this paper, we consider system (1.1) by using the scws method to find numerical solutions. scw has been used and showed efficiency to solve various problems. to indicate this, we can refer to some ∗corresponding author: pourgholi@du.ac.ir 76 n. azizi, r. pourgholi works. razzaghi and yousefi in [20] have employed a scw to solve variational problems. tavassoli kajani et al. [15] for solving integro-differential equations have presented a method based on scws. a numerical evaluation of hankel transform for seismology has been given in [14] using the scws approach. amir and umer saeed in [22] have used scws to solve the fractional nonlinear oscillator equations. in the present article, we intend to use the scws method to numerically solve the dsw system (1.1) with the initial conditions ψ(x, 0) = f1(x), φ(x, 0) = f2(x), x ∈ [0, 1], (1.2) and the boundary conditions ψ(0, t) = g1(t), φ(0, t) = g2(t), t ∈ [0, tfin], φ(1, t) = k2(t), φx(0, t) = w2(t), t ∈ [0, tfin], (1.3) where tfin represents the final time. the differentiable functions fi(x), gi(t), for i = 1, 2, k2(t), and w2(t) are known. the structure of this article is as follows: in section 2, we describe the properties of the scws. in the following, expanding functions into the scws series and operational matrix of them for the numerical solutions are discussed. in section 3, the procedure of implementation of the scws method, for system (1.1) with specified initial and boundary conditions (1.2) and (1.3) is presented. the numerical performance of the method is made in section 4, and finally, concluding remarks are given in section 5. 2. properties of scws wavelets are useful mathematical functions constructed from the dilation and translation of a single function called the mother wavelet, which can be denoted by ω. assuming that the expansion parameter η and the translation parameter ν are considered, we have the continuous wavelets family as follows [11]: wη,ν(x) = |η|− 1 2ω (x− ν η ) , η, ν ∈ r, µ 6= 0. if the parameters η and ν are restricted to take values η = η−κ0 and ν = rν0a −κ 0 , a family of discrete wavelets is obtained as: wκ,r(x) = |η0| κ 2ω(ηκ0x− rν0), (2.4) where η0 > 1, ν0 > 0, and r and κ are positive integers. the set {wκ,r(x)} in (2.4), forms a wavelet basis for l2(r). especially, if η0 = 2 and ν0 = 1, the set {wκ,r(x)} forms an orthonormal basis. scws are defined on interval x ∈ [0, 1) as [14]: wr,s(x) = 2 κ+1 2 gs(2 κx− r)χ[ r 2κ , r+1 2κ ), (2.5) where κ = {0} ∪ n, r = 0, 1, 2, . . . , 2κ − 1, and χ[ r 2κ , r+1 2κ ) denotes the characteristic function given as χ[ r 2κ , r+1 2κ ) = { 1, x ∈ [ r 2κ , r+1 2κ ), 0, elsewhere. (2.6) also, gs(x) =  1√ 2 , s = 0, cos(2sπx), s = 1, 2, . . . , `, sin(2(s− `)πx), s = `+ 1, `+ 2, . . . , 2`, (2.7) copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 77 where ` is any positive integer. scws have compact support and are an orthonormal basis forl2([0, 1)). the orthonormal basis functions for scws by assuming κ = 1 and ` = 1 are obtained as follows: for 0 ≤ x < 1 2 =⇒ w0,0(x) = √ 2, w0,1(x) = 2 cos(4πx), w0,2(x) = 2 sin(4πx), for 1 2 ≤ x < 1 =⇒ w1,0(x) = √ 2, w1,1(x) = 2 cos(2π(2x− 1)), w1,2(x) = 2 sin(2π(2x− 1)). (2.8) so, with the collocation points xm = 2m− 1 2n , m = 1, 2, . . . ,n = 2κ(2`+ 1), (2.9) the graphs of wr,s(x) for κ = ` = 1, are shown in fig. 2.1. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 w0,0(x) w0,1(x) w0,2(x) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 w1,0(x) w1,1(x) w1,2(x) fig. 2.1. the graphs of wr,s(x) for κ = ` = 1. 2.1. expanding functions into the scws using the set of scws, any function υ(x) ∈ l2([0, 1)) can be approximated as an infinite series of these functions as follows: υ(x) = ∞∑ r=0 2∑̀ s=0 ar,swr,s(x), (2.10) where ar,s =< υ,wr,s >= ∫ 1 0 υ(x)wr,s(x) dx. by truncating the infinite series (2.10) at levels r = 2κ − 1 and s = 2`, we obtain an approximate representation for υ(x) as υ(x) ' 2κ−1∑ r=0 2∑̀ s=0 ar,swr,s(x) = atγ(x), (2.11) where a and γ are (n × 1)-vectors and are introduced as follows: a = [ a0,0, a0,1, . . . , a0,2`, a1,0, a1,1, . . . , a1,2`, . . . . . . , a2κ−1,0, a2κ−1,1 . . . , a2κ−1,2` ]t , γ = [ w0,0,w0,1, . . . ,w0,2`,w1,0,w1,1, . . . ,w1,2`, . . . . . . ,w2κ−1,0(x),w2κ−1,1 . . . ,w2κ−1,2` ]t . (2.12) copyright © 2021 assa. adv syst sci appl (2021) 78 n. azizi, r. pourgholi the scws matrix γn×n at the collocation points (2.9), is given as follows: γn×n = [ γ ( 1 2n ) ,γ ( 3 2n ) , . . . ,γ (2n − 1 2n )] , in other words γn×n =  w0,0( 1 2n ) w0,0( 3 2n ) . . . w0,0(2n−1 2n ) w0,1( 1 2n ) w0,1( 3 2n ) . . . w0,1(2n−1 2n ) ... ... . . . ... w0,2`( 1 2n ) w0,2`( 3 2n ) . . . w0,2`( 2n−1 2n ) w1,0( 1 2n ) w1,0( 3 2n ) . . . w1,0(2n−1 2n ) w1,1( 1 2n ) w1,1( 3 2n ) . . . w1,1(2n−1 2n ) ... ... . . . ... w1,2`( 1 2n ) w1,2`( 3 2n ) . . . w1,2`( 2n−1 2n ) ... ... . . . ... ... ... . . . ... w2κ−1,0( 1 2n ) w2κ−1,0( 3 2n ) . . . w2κ−1,0(2n−1 2n ) w2κ−1,1( 1 2n ) w2κ−1,1( 3 2n ) . . . w2κ−1,1(2n−1 2n ) ... ... . . . ... w2κ−1,2`( 1 2n ) w2κ−1,2`( 3 2n ) . . . w2κ−1,2`( 2n−1 2n )  . in particular, for κ = ` = 1, the scws matrix γ6×6 is given as follows: γ6×6 =  √ 2 √ 2 √ 2 0 0 0 1 −2 1 0 0 0√ 3 0 − √ 3 0 0 0 0 0 0 √ 2 √ 2 √ 2 0 0 0 1 −2 1 0 0 0 √ 3 0 − √ 3  . 2.2. the operational matrix of scws due to the vector form (2.12), the integration of γ(x) can be calculated as follows:∫ x 0 γ(t) dt = qγ(x), where q is n ×n operational matrix given by q = 1 2κ+ 1 2  f s · · · s 0 f · · · s ... ... . . . ... 0 0 · · · f  , where s and f are (2`+ 1)× (2`+ 1) matrices (see [15]). copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 79 3. application of the method in this section, we use the scws method for finding the approximate solutions of system (1.1) with specified initial and boundary conditions (1.2) and (1.3). for this, dividing the interval [0, tfin] into m equal parts of length ~t = tfin m and denoting tn = (n− 1)~t, n = 1, 2, . . . , (m+ 1), we expand ψ̇′ and φ̇′′′ in terms of scws as, ψ̇′(x, t) ∼= 2κ−1∑ r=0 2∑̀ s=0 ar,swr,s(x) = atγ(x), (3.13) φ̇′′′(x, t) ∼= 2κ−1∑ r=0 2∑̀ s=0 br,swr,s(x) = btγ(x), (3.14) where prime and dot mean differentiation concerning x and t, respectively. now, we consider the following two cases: case 1: considering the equation (3.13): by integrating this equation once concerning t from tn to t and once concerning x from 0 to x, we have ψ′(x, t) = (t− tn)atγ(x) + ψ′(x, tn), (3.15) ψ̇(x, t) = atqγ(x) + g′1(t), (3.16) now, integrating equation (3.16) once concerning t from tn to t, we obtain ψ(x, t) = (t− tn)atqγ(x) + [g1(t)− g1(tn)] + ψ(x, tn). (3.17) case 2: considering the equation (3.14): by integrating this equation once concerning t from tn to t and three times concerning x from 0 to x, we obtain φ′′′(x, t) = (t− tn)btγ(x) + φ′′′(x, tn), (3.18) φ′(x, t) = (t− tn)btq2γ(x) + φ′(x, tn) + [w2(t)− w2(tn)] + x[φ̇′′(0, t)− φ̇′′(0, tn)], (3.19) φ(x, t) = (t− tn)btq3γ(x) + φ(x, tn) + [g2(t)− g2(tn)] + x[w2(t)− w2(tn)] + x2 2 [φ̇′′(0, t)− φ̇′′(0, tn)], (3.20) φ̇(x, t) = btq3γ(x) + g′2(t) + xw′2(t) + x2 2 φ̇′′(0, t). (3.21) by using the boundary condition φ(1, t) = k2(t) equations (3.19)-(3.21) are changed as follows: φ′(x, t) = (t− tn)btq2γ(x) + φ′(x, tn) + (1− 2x)[w2(t)− w2(tn)] + 2x[k2(t)− k2(tn)] − 2x[g2(t)− g2(tn)], (3.22) φ(x, t) = (t− tn)btq3γ(x) + φ(x, tn) + (1− x2)[g2(t)− g2(tn)] + x(1− x)[w2(t)− w2(tn)] + x2[k2(t)− k2(tn)], (3.23) φ̇(x, t) = btq3γ(x) + (1− x2)g′2(t) + x(1− x)w′2(t) + x2k′2(t). (3.24) copyright © 2021 assa. adv syst sci appl (2021) 80 n. azizi, r. pourgholi discretizing the results (3.15)-(3.17) and (3.18) and (3.22)-(3.24), by assuming x→ xm and t→ tn+1, we have ψ′(xm, tn+1) = ~tatγ(xm) + ψ′(xm, tn), (3.25) ψ̇(xm, tn+1) = atqγ(xm) + g′1(tn+1), (3.26) ψ(xm, tn+1) = ~tatqγ(x) + [g1(tn+1)− g1(tn)] + ψ(xm, tn), (3.27) φ′′′(xm, tn+1) = ~tbtγ(xm) + φ′′′(xm, tn), (3.28) φ′(xm, tn+1) = ~tbtq2γ(xm) + φ′(xm, tn) + (1− 2xm)[w2(tn+1)− w2(tn)] + 2xm[k2(tn+1)− k2(tn)]− 2xm[g2(tn+1)− g2(tn)], (3.29) φ(xm, tn+1) = ~tbtq3γ(xm) + φ(xm, tn) + (1− x2 m)[g2(tn+1)− g2(tn)] + xm(1− xm)[w2(tn+1)− w2(tn)] + x2 m[k2(tn+1)− k2(tn)], (3.30) φ̇(xm, tn+1) = btq3γ(xm) + (1− x2 m)g′2(tn+1) + xm(1− xm)w′2(tn+1) + x2 mk ′ 2(tn+1), (3.31) where, xm’s are the collocation points that are introduced in (2.9). to linearized the nonlinear terms φφx, ψxφ, and ψφx in system (1.1), we use the linearization form given by rubin and graves [21] as follows: φφx = φx(x, tn)φ(x, tn+1)− φx(x, tn)φ(x, tn) + φ(x, tn)φx(x, tn+1), (3.32) ψxφ = φ(x, tn)ψx(x, tn+1)− φ(x, tn)ψx(x, tn) + ψx(x, tn)φ(x, tn+1), (3.33) ψφx = φx(x, tn)ψ(x, tn+1)− φx(x, tn)ψ(x, tn) + ψ(x, tn)φx(x, tn+1). (3.34) using linear expressions (3.32)-(3.34), the discrete form of system (1.1) considering xm and tn+1 is as follows: ψ̇(xm, tn+1) + αφ′(xm, tn)φ(xm, tn+1) + αφ(xm, tn)φ′(xm, tn+1) = αφ′(xm, tn)φ(xm, tn), φ̇(xm, tn+1) + βφ′′′(xm, tn+1) + δφ(xm, tn)ψ′(xm, tn+1) + δψ′(xm, tn)φ(xm, tn+1) +γφ′(xm, tn)ψ(xm, tn+1) + γψ(xm, tn)φ′(xm, tn+1) = δφ(xm, tn)ψ′(xm, tn) +γφ′(xm, tn)ψ(xm, tn). (3.35) now, by using equations (3.25)-(3.31), system (3.35) leads to{ at θ1 + bt θ2 = h1(xm, tn), at θ3 + bt θ4 = h2(xm, tn), (3.36) where the matrices θi, i = 1, 2, 3, 4 are matrices with dimensions n ×n as follows: θ1 = qγ(xm), θ2 = [ αφ′(xm, tn)~tq3γ(xm) + αφ(xm, tn)~tq2γ(xm) ] , θ3 = [ δφ(xm, tn)~tγ(xm) + γφ′(xm, tn)~tqγ(xm) ] , θ4 = [ q3γ(xm) + β~tγ(xm) + δψ′(xm, tn)~tq3γ(xm) + γψ(xm, tn)~tq2γ(xm) ] , copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 81 and the matriceshi, i = 1, 2 are matrices with dimensions n × 1 as follows: h1(xm, tn) =− αφ(xm, tn)φ′(xm, tn)− g′1(tn+1)− αφ′(xm, tn) [ (1− x2 m)[g2(tn+1)− g2(tn)] + xm(1− xm)[w2(tn+1)− w2(tn)] + x2 m[k2(tn+1)− k2(tn)] ] − αφ(xm, tn) [ (1− 2xm)[w2(tn+1)− w2(tn)] + 2xm[k2(tn+1)− k2(tn)] − 2xm[g2(tn+1)− g2(tn)] ] , h2(xm, tn) =− βφ′′′(xm, tn)− δφ(xm, tn)ψ′(xm, tn) − γφ′(xm, tn)[g1(tn+1)− g1(tn) + ψ(xm, tn)] − [ (1− x2 m)g′2(tn+1) + xm(1− xm)w′2(tn+1) + x2 mk ′ 2(tn+1) ] − δψ′(xm, tn) [ (1− x2 m)[g2(tn+1)− g2(tn)] + xm(1− xm)[w2(tn+1)− w2(tn)] + x2 m[k2(tn+1)− k2(tn)] ] − γψ(xm, tn) [ (1− 2xm)[w2(tn+1)− w2(tn)] + 2xm[k2(tn+1)− k2(tn)]− 2xm[g2(tn+1)− g2(tn)] ] . the matrix-vector form of system (3.36) is as follows:[ θ1 θ2 θ3 θ4 ] 2n×2n [ a b ] 2n×1 = [ h1 h2 ] 2n×1 (3.37) from (3.37), the coefficients vector a and b can be calculated. with these coefficients and using the equations (3.27) and (3.30), the approximate solutions are successively obtained. 4. numerical experiments in this section, we apply the scws method to obtain the numerical solutions of the dsw system (1.1). to compare the obtained numerical results, we use the following solutions that obtained by arnous et al. ( [1]): ψ(x, t) = 6c γ+2δ sech2 (√ c βk2 (k(x− ct)− ξ0) ) , φ(x, t) = ± √ 12c2 α(γ+2δ) sech (√ c βk2 (k(x− ct)− ξ0) ) , where c and k are arbitrary real constants. to show the effectiveness and accuracy of the proposed method, we considered an example with α = β = γ = δ = 1, tfin = 1, ~t = 0.01. remark 4.1: for describing the error, we introduce the infinity-norm of absolute error and the root mean square (rms) error norm as follows: lψ ∞ = ||ψ(xm, t)−ψ∗(xm, t)||∞ = max 1≤m≤n |ψ(xm, t)−ψ∗(xm, t)|, rmsψ = [ 1 2n 2n∑ m=1 ( ψ(xm, t)−ψ∗(xm, t) )2 ] 1 2 , (4.38) copyright © 2021 assa. adv syst sci appl (2021) 82 n. azizi, r. pourgholi where ψ∗ is the approximate solution of ψ. similarly, the lφ ∞ and rmsφ are obtained according to formulas (4.38). the numerical results for ψ(x, t) and φ(x, t) at time t = 1 when ` = κ = 1 are reported in table 4.1. for different values of κ, table 4.2 stated the calculated errors (4.38) at time t = 0.5, also, the execution times for these values are given in table 4.3. difference between exact and numerical solutions ψ and φ at time t = 1 are shown in figs. 4.2-4.4 and 4.5-4.7, respectively. table 4.1. the numerical results for ψ(x, t) and φ(x, t) at t = 1 when ` = κ = 1. xm ψ(xm, 1) ψ∗(xm, 1) |ψ(xm, 1)−ψ∗(xm, 1)| φ(xm, 1) φ∗(xm, 1) |φ(xm, 1)− φ∗(xm, 1)| 0.0833333 0.153514 0.153517 2.358647e− 06 0.159955 0.159953 2.434473e− 06 0.25 0.157382 0.157384 2.526247e− 06 0.161958 0.161949 8.477983e− 06 0.416667 0.160643 0.160646 2.873385e− 06 0.163627 0.163619 7.684178e− 06 0.583333 0.163242 0.163244 2.279712e− 06 0.164945 0.164933 1.248603e− 05 0.75 0.165133 0.165128 5.211826e− 06 0.165898 0.165906 7.514665e− 06 0.916667 0.166281 0.166280 1.685403e− 06 0.166474 0.166457 1.740170e− 05 table 4.2. the calculated errors (4.38) at time t = 0.5 with ` = 1. ψ(x, 0.5) φ(x, 0.5) l∞ κ = 1 1.333830e− 06 8.630715e− 06 κ = 2 6.055654e− 08 1.078673e− 06 κ = 3 2.810132e− 09 7.847945e− 08 κ = 4 8.968863e− 11 5.304455e− 09 κ = 5 3.105035e− 12 3.172111e− 10 κ = 6 9.736049e− 14 2.018670e− 11 rms κ = 1 7.661924e− 07 5.199718e− 06 κ = 2 3.315228e− 08 5.115186e− 07 κ = 3 1.000311e− 09 2.933696e− 08 κ = 4 3.271928e− 11 2.026272e− 09 κ = 5 9.981293e− 13 1.217716e− 10 κ = 6 3.154537e− 14 7.797156e− 12 table 4.3. the execution times for different values of κ with ` = 1. κ = 1 κ = 2 κ = 3 κ = 4 κ = 5 κ = 6 cpu time (s) 102.835546 189.641444 373.585049 753.232887 1519.095320 3163.873643 copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 83 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -6 -5 -4 -3 -2 -1 0 1 2 3 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -6 (ℓ = 1, κ = 1) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2.5 -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 2.5 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -7 (ℓ = 1, κ = 2) fig. 4.2. difference between exact and numerical solutions ψ at time t = 1, when ` = 1 and κ = 1, 2. copyright © 2021 assa. adv syst sci appl (2021) 84 n. azizi, r. pourgholi 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -12 -10 -8 -6 -4 -2 0 2 4 6 8 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -9 (ℓ = 1, κ = 3) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -4 -3 -2 -1 0 1 2 3 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -10 (ℓ = 1, κ = 4) fig. 4.3. difference between exact and numerical solutions ψ at time t = 1, when ` = 1 and κ = 3, 4. copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 85 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -1.5 -1 -0.5 0 0.5 1 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -11 (ℓ = 1, κ = 5) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -4 -3 -2 -1 0 1 2 3 ψ (x ,1 ) − ψ ∗ (x ,1 ) ×10 -13 (ℓ = 1, κ = 6) fig. 4.4. difference between exact and numerical solutions ψ at time t = 1, when ` = 1 and κ = 5, 6. copyright © 2021 assa. adv syst sci appl (2021) 86 n. azizi, r. pourgholi 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2 -1.5 -1 -0.5 0 0.5 1 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -5 (ℓ = 1, κ = 1) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2.5 -2 -1.5 -1 -0.5 0 0.5 1 1.5 2 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -6 (ℓ = 1, κ = 2) fig. 4.5. difference between exact and numerical solutions φ at time t = 1, when ` = 1 and κ = 1, 2. copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 87 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -2 -1.5 -1 -0.5 0 0.5 1 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -7 (ℓ = 1, κ = 3) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -12 -10 -8 -6 -4 -2 0 2 4 6 8 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -9 (ℓ = 1, κ = 4) fig. 4.6. difference between exact and numerical solutions φ at time t = 1, when ` = 1 and κ = 3, 4. copyright © 2021 assa. adv syst sci appl (2021) 88 n. azizi, r. pourgholi 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -8 -6 -4 -2 0 2 4 6 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -10 (ℓ = 1, κ = 5) 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 x -5 -4 -3 -2 -1 0 1 2 3 4 φ (x ,1 ) − φ ∗ (x ,1 ) ×10 -11 (ℓ = 1, κ = 6) fig. 4.7. difference between exact and numerical solutions φ at time t = 1, when ` = 1 and κ = 5, 6. given the approximation function υ(x) expressed in section 2, the solutions of the system (1.1), can be expanded as: υ(x) = ∞∑ r=0 2∑̀ s=0 ar,swr,s(x). (4.39) copyright © 2021 assa. adv syst sci appl (2021) applications of sine-cosine wavelets method for solving drinfeld-sokolov-wilson 89 in the investigated example, we approximate the solution of this equation as follows: υ(x) ' 2κ−1∑ r=0 2∑̀ s=0 ar,swr,s(x), (4.40) which is the truncating the infinite series (4.39). by substituting the solutions ar,s in (4.40), we get the error function e (x) as follows: e (x) = ∣∣∣∣∣υ(x)− 2κ−1∑ r=0 2∑̀ s=0 ar,swr,s(x) ∣∣∣∣∣. therefore, as κ increases, the series (4.40) becomes larger and closer to the series (4.39), in other words, e (x) approaches zero. the obtained numerical results for ψ and φ confirm this. 5. conclusion in this article, using the scws method and using the initial and boundary conditions (1.2) and (1.3), we solved the dsw system (1.1) numerically. considering the obtained numerical results in tables 4.1 and 4.2, figs. 4.2-4.7, and also comparing these results with the exact solutions, it can be concluded that the presented method for solving the dsw system (1.1) is an efficient and high accuracy method. the strength of this method is the simplicity of calculations with low storage space. references 1. arnous, a., mirzazadeh, m., & eslami, m. (2016). exact solutions of the drinfel’d– sokolov–wilson equation using bäcklund transformation of riccati equation and trial function approach. pramana, 86(6), 1153–1160. 2. aziz, i., khan, f., et al. (2014). a new method based on haar wavelet for the numerical solution of two-dimensional nonlinear integral equations. journal of computational and applied mathematics, 272, 70–80. 3. chen, c., & hsiao, c. (1997). haar wavelet method for solving lumped and distributedparameter systems. iee proceedings-control theory and applications, 144(1), 87–94. 4. chui, c. k. (2016). an introduction to wavelets. elsevier. 5. daubechies, i. (1992). ten lectures on wavelets. siam. 6. daubechies, i., & lagarias, j. c. (1992). two-scale difference equations ii. local regularity, infinite products of matrices and fractals. siam journal on mathematical analysis, 23(4), 1031–1079. 7. drinfeld, v. g., & sokolov, v. v. (1981). equations of korteweg–de vries type, and simple lie algebras. in doklady akademii nauk, 258, 11–16. 8. drinfel’d, v. g., & sokolov, v. v. (1985). lie algebras and equations of korteweg-de vries type. journal of soviet mathematics, 30(2), 1975–2036. 9. foadian, s., pourgholi, r., tabasi, s. h., & damirchi, j. (2019). the inverse solution of the coupled nonlinear reaction–diffusion equations by the haar wavelets. international journal of computer mathematics, 96(1), 105–125. 10. fu, z., yuan, n., chen, z., mao, j., & liu, s. (2009). multi-order exact solutions to the drinfel’d–sokolov–wilson equations. physics letters a, 373(41), 3710–3714. 11. heil, c. (1993). ten lectures on wavelets (ingrid daubechies). siam review, 35(4), 666–669. 12. hirota, r., grammaticos, b., & ramani, a. (1986). soliton structure of the drinfel’d– sokolov–wilson equation. journal of mathematical physics, 27(6), 1499–1505. copyright © 2021 assa. adv syst sci appl (2021) 90 n. azizi, r. pourgholi 13. hosseininia, m., heydari, m., & avazzadeh, z. (2020). numerical study of the variableorder fractional version of the nonlinear fourth-order 2d diffusion-wave equation via 2d chebyshev wavelets. engineering with computers, 1–10. 14. irfan, n., & siddiqi, a. (2016). sine-cosine wavelets approach in numerical evaluation of hankel transform for seismology. applied mathematical modelling, 40(7-8), 4900– 4907. 15. kajani, m. t., ghasemi, m., & babolian, e. (2006). numerical solution of linear integro-differential equation by using sine–cosine wavelets. applied mathematics and computation, 180(2), 569–574. 16. mallat, s. g. (2009). a theory for multiresolution signal decomposition: the wavelet representation. in fundamental papers in wavelet theory, 494–513. 17. misirli, e., & gurefe, y. (2010). exact solutions of the drinfel’d–sokolov–wilson equation using the exp-function method. applied mathematics and computation, 216(9), 2623–2627. 18. pourgholi, r., esfahani, a., foadian, s., & parehkar, s. (2013). resolution of an inverse problem by haar basis and legendre wavelet methods. international journal of wavelets, multiresolution and information processing, 11(05), 1350034. 19. pourgholi, r., tavallaie, n., & foadian, s. (2012). applications of haar basis method for solving some ill-posed inverse problems. journal of mathematical chemistry, 50(8), 2317–2337. 20. razzaghi, m., & yousefi, s. (2002). sine-cosine wavelets operational matrix of integration and its applications in the calculus of variations. international journal of systems science, 33(10), 805–810. 21. rubin, s. g., & graves jr, r. a. (1975). a cubic spline approximation for problems in fluid mechanics. nasa sti/recon technical report n, 75, 33345. 22. saeed, a., & saeed, u. (2019). sine-cosine wavelet method for fractional oscillator equations. mathematical methods in the applied sciences, 42(18), 6960–6971. 23. wen, z., liu, z., & song, m. (2009). new exact solutions for the classical drinfel’d– sokolov–wilson equation. applied mathematics and computation, 215(6), 2349–2358. 24. wilson, g. (1982). the affine lie algebra c (1) 2 and an equation of hirota and satsuma. physics letters a, 89(7), 332–334. 25. xiao-xing, n., & qing-ping, l. (2015). darboux transformation for drinfel’d– sokolov–wilson equation. communications in theoretical physics, 64(5), 491. 26. yang, y., heydari, m., avazzadeh, z., & atangana, a. (2020). chebyshev wavelets operational matrices for solving nonlinear variable-order fractional integral equations. advances in difference equations, 2020(1), 1–24. 27. yuttanan, b., & razzaghi, m. (2019). legendre wavelets approach for numerical solutions of distributed order fractional differential equations. applied mathematical modelling, 70, 350–364. copyright © 2021 assa. adv syst sci appl (2021) introduction properties of scws expanding functions into the scws the operational matrix of scws application of the method numerical experiments conclusion microsoft word _done_1116-article text-5520-1-6-20220524 adv syst sci appl 2023; 01; 91-98 published online at https://ijassa.ipu.ru. nowcasting gdp of major economies during the crisis: does energy matter? ivan a. kopytin1, nikolay p. pilnik1, ivan p. stankevich1* 1) primakov national research institute of world economy and international relations, russian academy of sciences, moscow, russia abstract: in this article we compare the accuracy on nowcasts obtained with different models and different sets of indicators used as predictors for a set of 19 major economies. we compare the performance of mixed-frequency bayesian var models, dynamic factor models and unrestricted midas models with l1 regularization. we test different groups of commodity prices as possible predictors: energy indicators, agricultural commodities, precious metals and industrial metals prices. we find that among all the indicator groups tested energy commodities prices yield the highest average nowcasting accuracy, even though the accuracy of models utilizing all the indicators available remains slightly higher. among all the models tested, the highest quality is yielded by mixed-frequency bayesian var models. we also emphasize the importance of manual selection of predictors for non-diversified economies, where it can significantly improve the accuracy of nowcasts compared to the models with a wide set of predictors keywords: nowcasting, bayesian vector auto regression, gdp, oil price. 1. introduction one of the significant problems with macroeconomic statistics is the publication lag: the most important macroeconomic data is often published with several month lag (or longer). this problem becomes especially crucial in circumstances such as the recent crisis, caused by the new coronavirus infection: the necessity of stimulus is evident, but the exact amount is hard to determine without all the relevant statistics. one of the possible solutions for this problem is nowcasting of macroeconomic indicators. nowcasting is an estimation of the current level of an indicator that is published with a significant lag and is not observable at the moment the calculations are made, using a set of indicators that are published faster. the most typical nowcasting problem is gdp nowcasting as the gdp data, on the one hand, is usually published with a several month lag after the end of the corresponding period, and on the other hand, gdp is crucial in policy making and decision making of many agents. in this paper we investigate gdp nowcasting based on a set of commodity prices and pmis. from a technical point of view, there are several main approaches to gdp nowcasting. first, these are coupling equations based on predicting a high-frequency series using standard methods, then aggregating it into lower frequency data and using the resulting indicator as an explanatory variable in the low-frequency equation for the indicator of interest. this approach was presented, for example, in [11], and still retains considerable popularity, see [17, 18]. another very common class of econometric models is the so-called midas (mixed data sampling) mixed data frequency). this approach is based on using in the equation for a lower frequency series (for example, quarterly) as explanatory variables several values of a higher frequency series related to the current period (for example, three monthly values of another indicator related to this quarter). depending on the specification, a number of * corresponding author: ivanstankevich0@gmail.com 92 i. kopytin, n. pilnik, i. stankevich copyright ©2023 assa adv. in systems science and appl. (2023) restrictions on the values of the coefficients of the model can be introduced: models without restrictions are usually called unrestricted midas models (u-midas). presented in [8, 9], models of this class are widely used for the nowcasting of macroeconomic indicators, see, for example, [2, 7, 14]. from technically more complex approaches, one can single out mixed frequency vector autoregression (mfvar), based on the description of the joint dynamics of high-frequency indicators and the unobserved decomposition of low-frequency indicators into higher frequency data, see, for example, [12] or [16] in the application to nowcasting. there are also bayesian generalizations of this approach [15]. when building a nowcasting system, one has to choose between different model types and different indicators that can be used as explanatory variables in a nowcasting model. this choice becomes especially important for tasks of building a nowcasting system for a set of countries with different structures of the economy. in this paper, we test a set of widely used nowcasting models (midas, mixed frequency bayesian vars, dfm) on data for 19 major economies. we also compare different sets of operative indicators: a set of pmi indicators and several sets of world commodity prices: prices for energy commodities, metals (both precious and industrial) and agricultural commodities. pmis are a common choice in nowcasting and are often found to increase the nowcasting accuracy in real-time settings (see, e.g., [13] for the us and [5] for the euro area). the most popular choice among commodity prices is oil price, see, e.g., [2,4], and our main hypothesis is in line with this choice: among all the commodity groups, prices of energy indicators (including oil and gas prices) can yield the highest nowcasting accuracy. 2. methods we employ a relatively standard for nowcasting literature set of models: midas (mixed data sampling) models, dynamic factor models and mixed-frequency bayesian vector autoregressions. midas models are based on the utilization of higher-frequency indicators as explanatory variables. this class of models was presented [8,9] and is widely used in nowcasting of macroeconomic indicators (see [2, 7, 14]). in the general case, a midas model assumes that ( ) ( ) 1 0 0 i i p k i i t j t j j tm j t i m j j y y x u          , where ty is lower-frequency data (in our case – quarterly series of gdp growth rate); ( )i tx factors of higher frequency (in our case – monthly series of pmi of energy commodities prices); im defines the number of ( )ix observations in one observation of dependent variable (3 months in a quarter in our case). midas models also often imply restrictions on ( )i j , but in our case, we focus on unrestricted midas models without autoregressive component. instead, we use lasso regularization so that our resulting model can be written as: ( ) ( 1 0 ) 0 i i k i i t j tm j t i j m y x u        with target function  ( ) 2 ( ) 1 1 0 ( ) min i i j t k i t t j t i j m y y           . regularization for u-midas models can prevent overfitting while allowing for more complex dependencies and yielding higher forecasting accuracy (see [19]). mixed-frequency bvars describe the joint dynamics of variables of different frequencies in a single var model. they are widely used in nowcasting literature: see [15, 16] for one of the first applications, or more recent [3] shows that nowcasting performance of nowcasting gdp of major economies during the crisis: does energy matter? 93 copyright ©2023 assa. adv. in systems science and appl. (2023) mfbvars matches the performance of state-of-the-art dfm while having a more general structure and allowing for a greater flexibility. mfbvar assumes that all the processes (of different observed frequencies) evolve at higher frequency (monthly in our case), but are observed at lower frequency. for the variables observed in a lower frequency we assume that observed values ,q ty are obtained from the original monthly process ,q tx with intra-quarterly averaging (here and further when describing mfbvar model we follow the notation of [1]) , , 1 , 2 , 1 ( ), { , , , } 3 , q t q t q t q t x x x t march june september december y else         the same scheme is used in [16] for a very similar task of short-term forecasting of macroeconomic variables, including gdp. for the tx vector process we assume a standard var(p) model 1 1 , ~ n(0, )t t p t p t tx x x        т т , where 1,..., p  are coefficient matrices. by combining ,...,t t px x  into a single vector tz , we can get a companion form of var(p) model 1 , ~ n(0, )t t t tz z u u     . the observation equation for the original ty process is t t ty m z  , where tm is a selection matrix (that determines at which periods quarterly indicators are observed),  is an aggregation matrix based on weighting scheme employed (intra-quarterly averaging in our case). both these matrices are known and determined by the data structure and assumptions made.  and  are coefficients of companion form var model. we also utilize a minnesota-style prior which assumes for a model written in matrix form x w e  , where 1( , , )tw w w  • , 1( , , ,1)t t t pw x x  • • • ,  is the coefficients matrix and e errors such that vec( ) | ~ n(vec( ), )    ,  [( 1) 1]( ) diag( ) 0n p     • , 3 2 1 2 2 4 , lag l of variable r, i =(l-1)n+r ( ) , 1 ri l s i np           , where ,  parameters are prior beliefs specified by the researcher, i are diagonal elements of  , all  parameters control the “tightness” of the prior and are tuned when estimating a mfbvar model, 2 is are the residual variances from auxiliary ar(4) models. in our case, we utilize widely-used minnesota-style prior, where prior means for all the parameters except ar(1) parameters are chosen to be 0 and prior means for ar(1) parameters are  . hence, we imply that processes for all the series analyzed are ar(1) processes until the evidence (taken from the data) is enough for the estimation procedure to make the model more complex. priors for  in our case are taken to be 0 in order not to restrict them to any given value because our target variable (gdp growth rate) is stationary. another popular choice 1  is usually applied in non-stationary dependent variable cases. for the error covariance matrix we assume the quite common inverse wishart prior: 94 i. kopytin, n. pilnik, i. stankevich copyright ©2023 assa adv. in systems science and appl. (2023) 2 2 1 ~ ( , ) ( 1) ( ,..., ) 2 n iw s s n diag s s n          where 2 is are the residual variances from auxiliary ar(4) models. for the calculations we use the r package mfbvar of [1]. dynamic factor model implies that all the series observed are a combination of unobserved common factors (the number of factors is lower than the number of variables studied), that can be identified and used to make forecasts and nowcasts of indicators analyzed. the general specification of a dynamic factor model can be given by t t ty f    , 1 , ~ (0, ) p t i t i t t q i f a f bu u iid n i    , where ty is a vector of variables investigated, vector  and matrix  are parameters to be estimated, unobserved common factors tf are assumed to follow a var(p) process. we follow two-stage estimation approach as in [10], where parameters of  matrix and unobserved factors tf are estimated with principal components based on a standardized and balanced panel of explanatory variables and then re-estimated on an unbalanced set of regressors using kalman smoothing. the resulting factors are used in the model. estimation is performed in r using nowcasting package, [6] 3. data and research methodology we use the data on quarterly gdp growth rates for the period of 2001 – 2020 for the following countries: usa, russia, china, india, argentina, australia, brazil, britain, canada, france, germany, indonesia, italy, japan, mexico, saudi arabia, turkey, korea, south africa. the countries selected represent a significant part of the world gdp and in the same time belong to different geographical regions and have different structures of the economy – it can be important to test the stability of results obtained. we use 4 groups of monthly commodity prices (obtained from world bank’s pink sheet – [20]) as explanatory variables:  energy commodities (crude oil brent, dubai, wti; coal, australian, south african, natural gas us, europe; liquefied natural gas, japan),  agricultural commodities (wheat, us srw; rice, thai 5%; maize, palm oil, soybeans),  industrial metals (aluminum, iron ore, copper, nickel),  precious metals (gold, platinum, silver) we also use a set of monthly pmi series for the countries presented in the sample as a benchmark we compare the nowcasting accuracies over the last 10 points (quarters) for different groups of indicators. during the testing phase for each data point tested we restrict the sample to the information that replicates the actual nowcasting practice: all the explanatory variables for the current quarter are known (and are not known for the subsequent quarter), the indicator nowcasted is not known. we measure the accuracies using mean absolute errors (mae) for quarterly gdp growth rates nowcasting gdp of major economies during the crisis: does energy matter? 95 copyright ©2023 assa. adv. in systems science and appl. (2023) 4. results and discussion for each country we choose the model with the lowest mae. tables 4.1-4.3 present the types, factor groups and mae for the best model for each of 19 countries and the best model with mae for the case when all the indicators were used as a reference. accuracies in the table 4.1 are estimated on the whole sample, in the table 4.2 – without 2020, to measure the quality during a relatively stable period. we see that not in all the cases the best nowcast is obtained using the bigger sample of indicators. for the countries with economies focused on a particular group of products, such as saudi arabia, the accuracy of nowcast using only the most appropriate group of prices is almost twofold higher than for the model with all the indicators (even though this set still includes energy commodities prices). it allows us to conclude that in at least some cases the more is not always the better and manual selection of predictors can boost the performance of the model table 4.1. types of best models and their nowcasting mae for models with one groupof factors and models with all factors, % gdp one group all indicators type factor mae type mae usa midas_l1 indmetal 1.56 mfbvar 1.77 russia dfm pmi 1.74 midas_l1 1.60 china midas_l1 precmetal 1.86 mfbvar 1.71 india midas_l1 energy 3.89 dfm 4.34 argentina mfbvar energy 3.85 mfbvar 2.64 australia midas_l1 pmi 1.44 midas_l1 1.35 brazil dfm pmi 2.04 mfbvar 1.48 britain midas_l1 pmi 3.17 midas_l1 2.96 canada mfbvar pmi 1.13 mfbvar 1.49 france midas_l1 indmetal 2.82 mfbvar 3.04 germany dfm pmi 1.79 mfbvar 1.15 indonesia midas_l1 precmetal 1.30 midas_l1 1.45 italy dfm pmi 2.60 midas_l1 2.79 japan midas_l1 indmetal 1.63 mfbvar 1.55 mexico mfbvar energy 3.16 mfbvar 2.79 saudi arabia mfbvar energy 1.19 dfm 2.26 turkey mfbvar energy 2.00 mfbvar 2.97 korea mfbvar indmetal 1.31 mfbvar 0.96 south africa midas_l1 agriculture 0.57 midas_l1 0.60 the first table answers the question of the best way to predict the economic growth of each of the countries, especially in the specific conditions of 2020. the first thing to pay attention to is the question of whether one group of factors can be distinguished, or whether models with the entire set of variables work best. moreover, the group of those countries for which models with the entire set of indicators work better includes, for example, russia, china, germany and the united kingdom. on the other hand, for example, for countries such as the united states, india and france, it is possible to single out one group of variables that work best. we also see that, among all the indicator groups, the most frequent “winners” are energy commodity prices and pmi. these results hold even for the more stable sample without 2020. 96 i. kopytin, n. pilnik, i. stankevich copyright ©2023 assa adv. in systems science and appl. (2023) table 4.2. types of best models and their nowcasting mae for models with one group of factors and models with all factors without 2020, % gdp one group all indicators type factor mae type mae usa midas_l1 indmetal 0.39 mfbvar 0.40 russia mfbvar agriculture 0.58 midas_l1 0.69 china midas_l1 precmetal 0.37 midas_l1 1.40 india mfbvar agriculture 0.57 mfbvar 0.66 argentina mfbvar energy 2.40 mfbvar 2.46 australia mfbvar pmi 0.41 mfbvar 0.34 brazil mfbvar indmetal 0.49 midas_l1 0.47 britain midas_l1 pmi 0.52 midas_l1 0.49 canada midas_l1 pmi 0.44 mfbvar 0.63 france midas_l1 indmetal 0.34 midas_l1 0.53 germany dfm pmi 0.68 mfbvar 0.68 indonesia midas_l1 precmetal 0.21 midas_l1 0.23 italy midas_l1 energy 0.52 midas_l1 0.41 japan midas_l1 indmetal 0.63 midas_l1 0.68 mexico mfbvar pmi 0.84 mfbvar 1.15 saudi arabia mfbvar energy 0.88 dfm 1.38 turkey mfbvar energy 2.03 mfbvar 2.49 korea mfbvar pmi 0.42 mfbvar 0.46 south africa midas_l1 agriculture 0.43 midas_l1 0.70 comparing the results from the two previous tables is the most interesting. this approach actually allows us to draw meaningful conclusions about how radically the mechanisms that determine a significant part of the economic growth of the world's largest economies have changed. all countries can be divided into three groups depending on what composition of factors is used, and whether this composition of factors changes in models with and without 2020. the first group of countries includes australia, brazil, the uk and germany. the most working models for these countries both without and taking into account 2020 include all variables. the second group of countries includes the usa, canada, france, indonesia, saudi arabia, turkey and south africa. the most working models for these countries contain only one group of variables, regardless of whether 2020 is included in the models. for the us and france, industrial metal prices are the best leading indicators. for saudi arabia and turkey, prices for energy products. that is, a factor that is a good leading indicator can be both the main export product and the product on whose import the economy is critically dependent. finally, the third group includes russia, china, india, argentina, italy, japan, mexico and korea. for them, models with the addition of data for 2020 differ significantly from models without this period. almost all of these countries (the only exceptions are italy and india, which at different times experienced a very strong shock from the spread of the coronavirus) have the same trend. with the addition of the 2020 data, the model that used only one group of indicators changes to a model that includes all groups of factors. it seems to us that this is a rather significant effect associated with the crisis events of 2020. during this period, the interrelationships in the economy become more complex, and it is impossible to say in advance which variable will make the greatest contribution. therefore, models that take into account all the analyzed factors begin to work more accurately. nowcasting gdp of major economies during the crisis: does energy matter? 97 copyright ©2023 assa. adv. in systems science and appl. (2023) table 4.3 present the mean mae for all the models and predictor groups with and without 2020. in both cases the best nowcasting quality is obtained using mfbvar and in both cases the best indicator group is energy commodities prices. however, for the whole sample mfbvar with all the indicators performs better. table 4.3. average errors (mae) across all country models, % gdp energy pmi precmetal indmetal agriculture all indicators all time mfbvar 2.66 3.34 3.13 2.76 4.38 2.37 dfm 2.96 2.74 3.10 3.03 3.07 3.02 midas_l1 3.03 3.30 3.23 3.36 3.23 2.55 without 2020 mfbvar 1.14 1.26 1.29 1.16 1.23 1.17 dfm 1.63 1.54 1.66 1.61 1.64 1.62 midas_l1 1.63 1.94 2.12 1.95 1.70 1.17 5. conclusion in this article we compare the accuracy on nowcasts obtained with different models and different sets of indicators used as predictors for a set of 19 major economies. we compare the performance of mixed-frequency bayesian var models, dynamic factor models and unrestricted midas models with l1 regularization. we test different groups of commodity prices as possible predictors: energy indicators, agricultural commodities, precious metals and industrial metals prices. we found that energy really matters in nowcasting tasks: especially for non-diversified economies and during stable periods of time. energy commodities prices used as predictors for nowcasting models yield lower mean nowcasting errors than any other group of indicators, including widely-used pmis. we also found that mixed-frequency bayesian vars tend to have the highest accuracy in most of the cases analyzed. acknowledgements the article was prepared within the project "post-crisis world order: challenges and technologies, competition and cooperation" supported by the grant from ministry of science and higher education of the russian federation program for research projects in priority areas of scientific and technological development (agreement № 075-15-2020-783). references 1. ankargren, s., & yang, y. (2019). mixed-frequency bayesian var models in r: the mfbvar package. [online] availabe https://www.diva-portal.org/smash/ record.jsf?pid=diva2%3a1345141&dswid=8425 2. chernis, t., & sekkel, r. (2017). a dynamic factor model for nowcasting canadian gdp growth, empirical economics, 53(1), 217–234. 3. cimadomo, j., giannone, d., lenza, m., sokol, a., & monti, f. (2020). nowcasting with large bayesian vector autoregressions, working paper series, 2453, 1–29. 4. dahlhaus, t., guénette, j. d., & vasishtha, g. (2017). nowcasting bric+ m in real time, international journal of forecasting, 33(4), 915–935. 5. de bondt, g. j. (2019). a pmi-based real gdp tracker for the euro area, journal of business cycle research, 15(2), 147–170. 98 i. kopytin, n. pilnik, i. stankevich copyright ©2023 assa adv. in systems science and appl. (2023) 6. de valk, s., de mattos, d., & ferreira, p. (2019). nowcasting: an r package for predicting economic variables using dynamic factor models, the r journal, 11(1), 230–244. 7. ferrara, l., & marsilli, c. (2019). nowcasting global economic growth: a factor‐augmented mixed‐frequency approach, the world economy, 42(3), 846–875. 8. ghysels, e., santa-clara, p., & valkanov, r. (2006). predicting volatility: getting the most out of return data sampled at different frequencies, journal of econometrics, 131(1-2), 59–95. 9. ghysels, e., sinko, a., & valkanov, r. (2007). midas regressions: further results and new directions, econometric reviews, 26(1), 53–90. 10. giannone, d., reichlin, l., & small, d. (2008). nowcasting: the real-time informational content of macroeconomic data, journal of monetary economics, 55(4), 665–676. 11. ingenito r., trehan b. (1996). using monthly data to predict quarterly output, economic review, 3, 3–11. 12. kuzin v., marcellino m. & schumacher c. (2011). midas vs. mixed-frequency var: nowcasting gdp in the euro area, international journal of forecasting, 27(2), 529–542. 13. lahiri, k., & monokroussos, g. (2013). nowcasting us gdp: the role of ism business surveys, international journal of forecasting, 29(4), 644–658. 14. marcellino, m., & schumacher, c. (2010). factor midas for nowcasting and forecasting with ragged‐edge data: a model comparison for german gdp, oxford bulletin of economics and statistics, 72(4), 518–550. 15. mccracken, m. w., owyang, m., & sekhposyan, t. (2015). real-time forecasting with a large, mixed frequency, bayesian var, frb st. louis working paper, 2015. 16. schorfheide, f., & song, d. (2015). real-time forecasting with a mixed-frequency var, journal of business & economic statistics, 33(3), 366–380. 17. schumacher c. (2016). a comparison of midas and bridge equations, international journal of forecasting, 32(2), 257–270. 18. soybilgen b. & yazgan e. (2018). evaluating nowcasts of bridge equations with advanced combination schemes for the turkish unemployment rate, economic modelling, 72, 99–108. 19. stankevich, i. (2020). comparison of macroeconomic indicators nowcasting methods: russian gdp case, applied econometrics, 59, 113–127. 20. world bank commodity price data (the pink sheet) (2021, june). [online] available https://www.worldbank.org/en/research/commodity-markets microsoft word 792-article text-3045-1-18-20200630.doc adv syst sci appl 2020; 02; 20-31 published online at https://ijassa.ipu.ru. global stability analysis of typhoid fever model olumuyiwa james peter1*, adebisi ajimot folasade2, michael oyelami ajisope3, fidelis odedishemi ajibade4, adesoye idowu abioye5, festus abiodun oguntolu6 1,5 department of mathematics, university of ilorin, ilorin, nigeria 2 department of mathematics, osun state university, oshogbo, osun state, nigeria 3department of mathematics, federal university oye-ekiti, oye-ekiti, nigeria 4department of civil engineering, federal university of technology, akure, nigeria 6department of mathematics, federal university of technology, minna, niger state, nigeria e-mail: peterjames4real@gmail.com abstract: we analyze with four compartments a deterministic nonlinear mathematical model of typhoid fever transmission dynamics. using the lipchitz condition, we verified the existence and uniqueness of the model solutions to establish the validity of the model and derive the equilibria states of the model that is, disease-free equilibrium (dfe) and endemic equilibrium (ee). the computed basic reproductive number r0 was used to establish that the disease-free equilibrium is globally asymptotically stable when its numerical value is less than one, the disease will be under control. in addition, the lyapunov function was applied to investigate the stability property for the (dfe). the model was numerically simulated to validate the results of the analysis. keywords: typhoid fever, equilibria, stability, nonlinear mathematical model 1. introduction typhoid fever is an infection caused by salmonella typhi bacteria. typhoid is usually triggered by the ingestion of food or water contaminated with feces or urine of infected individuals and is, therefore, a typical illness in regions with poor sanitation. (brooks [1]; roumagnac et al., [2]). in developing economies, typhoid fever outbreaks occur from time to time in overwhelmed areas and refugee camps with high population density. the disease causes high morbidity in children below ten years of age, with at least seventeen million new cases globally and nearly 600 000 deaths annually. the disease is not rampant in north america: in the us, an estimated four hundred cases are reported every year; seventy percent of the cases are traced to those who return from endemic regions. the mortality of typhoid fever is ten percent, but with adequate treatment, it can be limited to one percent. (lin et al., [3]; sinha et al., [4]; hyman, [5]). intestinal fever treatment is anchored on the blood culture condition of the patients. if the species is sensitive, the oral antibiotic is used. typhoid fever is becoming an increasingly common illness worldwide, making antibiotic treatment more expensive and harder. enteric fever signs vary and are similar to the signs of different microbial infections. symptoms of typhoid fever are as follows: variable level of high fever in 75 cases, body aches and muscle pain, chills, shriveled loss of appetite, abdominal pain in 20 to 40 cases, nose bleeds, headache, dizziness, rashes on the skin, bowel constipation or looseness, weakness and fatigue, sore throat and cough (lifshitz [6]). some mathematical models have been formulated on the transmission of typhoid fever. (adetunde [7]; lauria et al., [8]; * corresponding author: peterjames4real@gmail.com global stability analysis of typhoid fever model copyright ©2020 assa. adv. in systems science and appl. (2020) 21 kalajdzievska [9]; mushayabasa [10]; cvjetanovic et al.,[11]; moffact [12]; pitzer et al., [13]; date et al.,[14]; muhammad, et al.,[15]; watson & edmunds [16]; nthiiri [17]; moatlhod & gosaamang [18]; mushayabasa [19]; peter & ibrahim [20]; peter et al.,[21]; peter and ibrahim [22]). the aim of this study is to extend and complement previous works by formulating a model that captures the following controls; vaccination and education. 2. materials and methods the model comprises four compartments: susceptible class; represents the proportion of those who are prone to typhoid. infected class; represents the population of individuals who have been infected with typhoid fever and capable of transmitting the infection to susceptible populations through interaction. carrier-class; represents the population of individuals who have been infected with typhoid fever but without signs of infectiousness in them. recovered compartment; represent the number of individuals infected with typhoid fever but are now cured of typhoid as a result of treatment. recruitment into susceptible populations is by birth or immigration at the rate . it is assumed that a certain proportion in the susceptible class moved to the carrier infectious class at rate while the complement moved to the infectious compartment. we also assume that the rate of disease transmission of carrier individuals will be higher than the disease transmission rate of infected individuals this is because they are more likely to be unaware of their infection. carrier's disease symptom is noticeable at the rate . contagious individuals may receive treatment and recuperate at the rate . susceptible individuals receive vaccination to protect themselves from contracting the disease at the rate . is an educational parameter that serves as the limitation for carriers and indicative persons from transmitting typhoid fever. this parameter lies within the range . when it implies that educational campaign not in position so that the vulnerable population are oblivious of typhoid fever and when it denotes that all vulnerable persons are well informed of the causes of typhoid fever, that is, they explicitly understand what brings about the diseases, how the disease is being transmitted and how to safeguard themselves from being infected by the disease. figure 1. pictorial description of the model q r r-1 b g a y f-1 10 ££ f 0=f 1=f o. j. peter, a. f. adebisi, m. o. ajisope, f.o. ajibade, a. i. abioye copyright ©2020 assa adv. in systems science and appl. (2020) 22 (1) where (2) table 1. variables and parameters interpretation variables description vulnerable populations carrier contagious populations contagious populations recovered populations parameters interpretation recruitment rate in vulnerable class mortality rate of vulnerable class mortality rate for carrier infected class mortality rate for infected class mortality rate for recovered class the rate at which individual carriers develops symptoms parameter for education rate of vaccination the probability of newly infected persons becoming asympt omatic or carrier rate of transmission of infection for carrier individuals rate of transmission of infection for infectious individuals force of disease infection rate at which individual in the infectious class recovered ï ï ï ï þ ïï ï ï ý ü -+ +--+-------ris dt dr iis dt di iis dt dic sss dt ds c cc 4 3 2 1 = )()(1)(1)(1= )(1)(1= )(1= µdy dµfaflr faµfrl yflµq iic gbl += ï ï ï ï þ ïï ï ï ý ü -+ +--++-----+ --+-ris dt dr iiiis dt di iiiis dt dic siiss dt ds cc ccc c 4 3 2 1 = )()(1)())(1(1= )(1))(1(= ))(1(= µdy dµfagbfr faµfgbr yfgbµq ( )s t ( )ic t )(ti )(tr q 1µ 2µ 3µ 4µ a f y r b g l d global stability analysis of typhoid fever model copyright ©2020 assa. adv. in systems science and appl. (2020) 23 3. solution of the model 3.1. existence and uniqueness of the model the validity of any mathematical model is a function of the solution of the model, provided the solution is unique. we verify the singularity of the solution of the model in equation (2) via the popular lipchitz condition. given the system of equation (2) be as follows (3) theorem 3.1 assuming the region is denoted by (4) and suppose that r’ meets the lipchitz condition on each occasion, the pairs and , where is a positive constant. hence, there is a constant such that there exists a unique continuous vector solution of the system in the interval . it is vital to mention that the requirement is fulfilled by the condition that is continuous and bounded in . considering the model equation (2), we are of the interest in the region we look for the solution that is bounded in the region and whose partial derivatives satisfy where and are positive constants. theorem 3.2 let j represent the region then equation (2) has a unique solution if it is established that are continuous and bounded in r for , for , ï ï ï þ ïï ï ý ü -+ +--++-----+ --+-risj iiiisj iiiisj siissj cc ccc c 44 33 22 11 = )()(1)())(1(1= )(1))(1(= ))(1(= µdy dµfagbfr faµfgbr yfgbµq â ),........,(=),........,,(=1,, 20102100 noon yyyyyyyyyyacc £-£2121 ),(),( yykycfycf -£( )1, yc ( ) jyc î2, 0³d )(cy d£0cc ..1,2,=,, ji x f j i ¶ ¶ â .0 ⣣a 0££ af a d 1,2,3,4=,, ji x j j i ¶ ¶ 1j ¥--+-¶ ¶ <))(1(= 1 1 yfgbµ ii s j c ¥-¶ ¶ <)(1=1 fbs i j c ¥-¶ ¶ <)(1=1 fgs i j ¥ ¶ ¶ <0=1 r j 2j ¥-+ ¶ ¶ <))(1(=2 fgbr ii s j c ¥---¶ ¶ <)(1)(1= 2 2 faµfbrs i j c o. j. peter, a. f. adebisi, m. o. ajisope, f.o. ajibade, a. i. abioye copyright ©2020 assa adv. in systems science and appl. (2020) 24 , for , , for , , , the partial derivatives exist, continuous and bounded, so the model has a unique solution that completes the proof of theorem 3.1. 3.2. feasible region and equilibrium from equation (2) we have that and thus, along with each solution. also from (2), we see that where . therefore, we have omitted r because r does not appear in other equations. this reveals that the model can be studied in the feasible region; . is positively invariant in relation to (2). once the dynamics are understood, those of r can then be determined from the equation . the first stage of our analysis is to find the disease-free equilibrium states from the equations by setting the right-hand side to zero i.e., the model always has a disease-freeequilibrium . and the endemic equilibrium satisfies from the equilibrium equations we can show that a unique exist with where ¥¶ ¶ <)(1=2 fgrs i j ¥ ¶ ¶ <0=2 r j 3j ¥+-¶ ¶ <))()(1(1=3 ii s j c gbfr ¥-+-¶ ¶ <)(1))(1(1=3 fagfr s i j c ¥+--¶ ¶ <)())(1(1= 3 3 dµgfr s i j ¥ ¶ ¶ <0=3 r j 4j ¥ ¶ ¶ <=4 y s j ¥ ¶ ¶ <0=4 ci j ¥ ¶ ¶ <=4 d i j ¥¶ ¶ <= 4 4 µ r j s dt ds )( 1µyq +-£ 1 )(suplim µy q + £ ¥® ts t ncrcicicsc dt dn c -£----£ qq 4321 }{ 4321 ,,.min ccccc = c tn t q £ ¥® )(suplim î í ì þ ý ü £++ + £âî=g + c iissiis cc q µy q ,:),,( 1 3 g ),,( iis c risr 4 ' µdy --= ),,( iis c 0=))(1(1 siiss c yfgbµq --+-0=)(1))(1( 2 ccc iiiis faµfgbr ----+ 0=)()(1)())(1(1 3 iiiis cc dµfagbfr +--++-ú û ù ê ë é + ,0,0= 1 0 µy qa ),,( **** iisa c= .0,, *** >iis c *a )( ))(( = 223 23* gµrgµrbµbdrag aµdµ +-++ ++ kk k s )(1= f-k global stability analysis of typhoid fever model copyright ©2020 assa. adv. in systems science and appl. (2020) 25 for to exist in the feasible region , the necessary and sufficient condition requires , or equivalently, define (4) then a threshold parameter determines the actual number of equilibria. in the next section, we shall explain the basic reproduction number. proposal 3.1 if then is the only equilibrium in ; if so, two equilibria exist, , and a unique endemic equilibrium, 3.3. basic reproduction number the basic reproductive number measures the number of secondary cases that the bacterium is able to introduce into the whole population of fully susceptible individuals in a stable demographic state (kalajdzievska [9]). the basic number of reproduction is a significant quantity in epidemiology since it sets the threshold for the analysis of an outbreak of disease and the evaluation of its control strategies. the value of the reproductive number, therefore, indicates whether a disease becomes endemic or dies out in a population. if , it indicates that each contagious person will not cause an infection and therefore, the disease will always disappear but when the basic reproduction number is greater than one i.e., each contagious person will cause at least one secondary infection which will subject the entire population to the attack of the disease. it is obtained by taking the largest eigenvalue(spectral radius) of the matrix where is the rate of new infections in infectious class, is the transfer of persons out of the compartment by another means, is the disease-free equilibrium. applying the next generation matrix approach, the basic reproduction number for the model is given as (5) in equation (5), the expression in the big square brackets is the per capita mean number of secondary infections. this expression is multiplied by , the number of susceptible individuals in the absence of infection to come about the basic reproduction number 3.4. disease-free equilibrium’s stability in order to check the local stability in the absence of infection a0, the jacobian matrix will be evaluated at *a g 1 *0 µy q + £< s .1 )( * 1 ³ + sµy q [ ] )])([( )1()(1 231 23 1 *0 k kk s r aµdµyµ rgµdµbragq µy q +++ -+++ = + = 10 r 0a *a 10 r 1 0 )()(= ú ú û ù ê ê ë é ¶ ¶ ú ú û ù ê ê ë é ¶ ¶ j oi j oi x xv x xfr if iv ox ú û ù ê ë é ++ + + + +++ = ))(( )1( )())(( 23 2 2231 0 kkk kkr aµdµ rgµ aµ br aµdµ ag yµ q yµ q +1 or ú û ù ê ë é + ,0,0= 1 0 µy qa o. j. peter, a. f. adebisi, m. o. ajisope, f.o. ajibade, a. i. abioye copyright ©2020 assa adv. in systems science and appl. (2020) 26 the following stability results indicate that is a sharp threshold. proposal 3.2 a0 locally asymptotically stable if and is unstable if . proof one eigenvalue of is < 0. the other two eigenvalues are matrix. we want to show that whenever then the routh-hurwitz conditions hold, that is, and . going by the assumption that and therefore, and this shows that and the whenever . this establishes the proposition. ú ú ú ú ú ú ú ú ú ú û ù ê ê ê ê ê ê ê ê ê ê ë é +-÷÷ ø ö çç è æ + ---+÷÷ ø ö çç è æ + -÷÷ ø ö çç è æ + --+-÷÷ ø ö çç è æ + ÷÷ ø ö çç è æ + --÷÷ ø ö çç è æ + --+)()1)(1()1()1)(1(0 )1())1(()1(0 )1()1()( =)( 3 11 1 2 1 11 1 0 dµ yµ q gfrfa yµ q bfr yµ q frgfaµ yµ q frb yµ q fg yµ q fbyµ aj 10 r )( 0aj )( 11 yµl +-= 22´ ú ú ú ú ú ú ú ú ú û ù ê ê ê ê ê ê ê ê ê ë é +-÷÷ ø ö çç è æ + ---+÷÷ ø ö çç è æ + -÷÷ ø ö çç è æ + --+-÷÷ ø ö çç è æ + )()1)(1()1()1)(1( )1())1(()1( =)( 3 11 1 2 10 dµ yµ qgfrfa yµ qbfr yµ qfrgfaµ yµ qfrb dj 10 a ú ú ú ú ú û ù ê ê ê ê ê ë é + ÷÷ ø ö çç è æ + -++ ú ú ú ú ú û ù ê ê ê ê ê ë é -+ ÷÷ ø ö çç è æ + -+= 1 )( )1(1( )(1 )1( )1( ))1(()( 3 1 3 2 1 2 dµ yµ qgfr dµ faµ yµ qfrb faµatra 1 ))(( )1( )())(( 23 2 2231 0 <ú û ù ê ë é ++ + + + +++ = kkk kkr aµdµ rgµ aµ br aµdµ ag yµ q )1( )1( 2 1 faµ yµ qfrb -+ ÷÷ ø ö çç è æ + 1 )( )1(1( 3 1 < + ÷÷ ø ö çç è æ + -dµ yµ qgfr )1( )1( 2 1 faµ yµ qfrb -+ ÷÷ ø ö çç è æ + 01 )( )1(1( 3 1 <+ ÷÷ ø ö çç è æ + -dµ yµ qgfr 0)( a 10 = threshold (threshold = 0,7 according to the initial conditions); when a tectonic plate collides, another random variable wave acceleration in seismometer alert event is generated (random variable generation and event generation modules are applied). if wave acceleration > 0,05 m/s2 (minimum value of risk) then seismometer alert is activated, also the time is stopped by using timestamp = “stop” else seismometer alert = “inactive.” such an event triggers the measurement station starts event when seismometer alert = “active” for changing the seismograph state to “on” and automatically inserting the measurement data (local time, slip distance, fault area); in this { } { } { } { } code code initial conditionsearth zone has earthquake arrives { } = rand >= yes threshold = “collision” random value = “position” no random value station earthquare date epicenter focus seismologist code name has magnitude emerges has inserts earthquake local time = time name measurement local time has location slip distance manages inserts updates deletes searchs plate tectonic state plate tectonic state time = 0 seconds threshold = 0,7 random value = 0,0 has moment time starts station measurement wave acceleration > 0,05 m/s2 seismometer alert = “active” seismometer alert = “inactive” wave acceleration time simulation = 28800 seconds yes no rock rigidity = * * plate tectonic has code name seismometer alert latitude longitude name altitude fault area seismograph state “on” “active” seismometer alert = seismograph state “off” ends station measurement { } earthquake magnitude / = 3 log10 * 10.7 2 final slip distance final fault area seismic moment intensity { } = + 1 seconds time time passes time earthquake seismic moment earthquake seismic moment earthquake slip distance earthquake fault area wave acceleration = 0,0 m/s2 seismometer alert = “inactive” rock rigidity = “position” plate tectonic state = seismograph state “off” = = 7 x 107 kg/m2 = micro light moderate strong -very strong violent -extreme wave acceleration * = “next”timestamp <= time time simulation = rand measurement timestamp “next” = and <= time time simulation timestamp “next” = and = “collision” plate tectonic state = “next”timestamp = “stop”timestamp >= time time simulation seismometer alert “inactive” = or simulating events in requirements engineering by using pre-conceptual-schema… 9 copyright ©2021 assa. adv sys sci appl (2021) case, we use the implication link for obligatorily relating to the cause of other event. thus, timestamp = “next” for increasing the time and returning to seismometer alert emerges, comparing the wave acceleration value, and saving measurement to the end of the earthquake. if seismometer alert = “inactive” or time >= simulation time, then station measurement ends. when measurement ends, then seismologist manages—updates, inserts, searches, deletes, which are dynamic relationships—the earthquake class. such a class has leaf concepts time moment, date, epicenter, focus, final slip distance, final fault area, magnitude, intensity, and seismic moment. magnitude and seismic moment are derived attributes (values calculated by applying mathematical equations with leaf concepts inserted, numbers, and parameters; magnitude = 2/3 * (log 10 seismic moment – 10,7) and seismic moment = rock rigidity * slip distance * fault area * wave acceleration; the intensity is micro, light, moderate, strong, very strong, violent, and extreme according to the magnitude.) simulating events from ssd by using pcs. the complete logic allow performing the event simulation with data required in the system and the database. thus, both scientists and business analysts should validate the requirements with scientific stakeholders with representation, then they can understand the complete system, simulating events and its data previously to the development phase of the scientific software. example, we simulate events in an earthquake occurred in santander, colombia. such a phenomenon is analyzed by using the pcs of the fig. 3. we show the epicenter and station in fig. 4, where the earthquake was measured and detected. we use the real information of variables, parameters, classes, and leaf concepts for simulating the earthquake. fig. 3. earthquake occurred in santander, colombia [4] values (leaf concepts) of the tectonic plate, earth zone, seismologist, and station classes are stored in the tables of the database tectonic plate, earth zone, seismologist, and station (see tables 2, 3, and 4). values (local time, slip distance, fault area, and station code) of the 10 p.a. noreña, c.m. zapata copyright ©2021 assa. adv sys sci appl (2021) measurement class are automatically inserted by using the event station measurement starts (see fig. 3) and stored in the table measurement of the database (see table 6; wave acceleration random value assigned is ‘1,2 m/s2’ for applying the mathematical equation of the seismic moment, its unit is newton (n) meters. values (date, moment time, earth zone code, station, focus, epicenter, latitude, longitude, final slip distance, final fault area, seismic moment, magnitude, and intensity) of the earthquake class are inserted by the seismologist when the station measurement ends event happens (see fig. 3) and stored in the table earthquake of the database (see table 7). earthquake had an intensity “strong” and a magnitude ‘5,5 n meters’, value summarized in ‘5,5’ by excluding units in richter scale. some fictional information is added to the real data in order to complete the simulation, e.g., the seismologist. table 2. tectonic plate. the authors code name earth zone code ptc1 cocos 101 ptc2 south american 101 ptc3 nazca 101 ptc4 caribbean 101 table 3. earth zone the authors code name 101 colombia table 4. earth zone the authors code name 104521452 juan manuel gaviria table 5. station. the authors code name location latitude longitude altitude earth zone code seismologist code 62 barranca santander 7,101o -73,7o 132 101 104521451 table 6. measurement. the authors local time (s) slip distance (m) fault area (m2) station code 30000 0,5 33 62 31000 1,5 34,1 62 32000 2 36,3 62 33000 2,8 38,5 62 34000 3 39 62 35000 3,35 40,2 62 simulating events in requirements engineering by using pre-conceptual-schema… 11 copyright ©2021 assa. adv sys sci appl (2021) table 7. earthquake. the authors d at e m om en t ti m e( s) ea rth z on e co de st at i o n co de fo cu s ( m ) ep ic en te r la tit ud e lo ng itu de fi na l s lip d ist an ce (m ) fi na l f au lt ar ea (m 2 ) se ism ic m om en t ( m ) m ag ni tu de in te ns ity 30 /0 5/ 20 18 30 00 10 1 62 14 0. 00 0 sa nt an de r 6, 83 ° 7 3, 11 ° 3, 35 40 ,2 94 26 9x 1 05 5, 5 st ro ng 5. validation we evaluate the understanding and the usage the pcs components by the expert opinion. we collect the data with a survey as instrument. we select two independent variables: current role and area. the sample size is 28 experts (10 scientists: 6 in geology, 1 in energy, 1 in petroleum, and 2 in simulation; 11 developers, 3 professor, and 4 analysts in re and software engineering) and four dependent variables: pcs representation. first, experts observed the pcs in fig.3. after, they answered the question q1. describe with your words, what is being represented in the model? for analyzing the pcs representation variable. we observe 61% of experts described phrases related to an earthquake early warning system, simulation, detection, and automated measure of an earthquake, i.e., a system to measure earthquakes and tectonic activity, tectonic plate movement simulation, a system to detect and collect information on earthquakes. in addition, the 28% of experts indicated phrases related to an earthquake, a tectonic activity, a seismic event, and events of an earthquake and a seismological station, i.e., a seismologist takes measurements of an earthquake, a conceptual model of earthquakes, their properties, and measurements. only 11% of experts described different phrases as: relational model, a system, and remembers. the results indicate the 89% (25) of experts perform a near description to the system domain (see fig. 4). fig. 4. pcs representation component integration. experts answered the question: q2. does the pcs integrate scientific, software, and simulation components? for analyzing of the component integration variable. the 86% (24) of experts and affirm the model integrate such components and the 11% (3) indicates perhaps integrate the components. the results confirm the pcs include 12 p.a. noreña, c.m. zapata copyright ©2021 assa. adv sys sci appl (2021) scientific, software and simulation components, which allow represent a ssd and simulating events (see fig. 5). fig. 5. component integration pcs understand. experts answered the question: q3. is the pcs understandable? for analyzing the pcs understand variable. from the point of view of the experts, 20 experts are strongly agreed and agreed with the pcs is understandable, 7 experts are neutrals, and 1 expert is disagreed (see fig. 7). the majority opinion of the results in show acceptation in the understanding of the schema. fig. 6. pcs understand pcs usage. experts answered the question: q4. would you use the pcs in scientific software development? e.g., in modeling, simulation and requirements elicitation for analyzing the pcs usage variable. the 72% (20) experts affirm they could use the pcs in scientific software development processes, the 21% (6) experts indicate perhaps would use it and the 7% (2) answer they would not use it (see fig. 7). the results present a positive position in the pcs usage for modeling ssd, event simulation, and requirements elicitation in scientific software development processes. simulating events in requirements engineering by using pre-conceptual-schema… 13 copyright ©2021 assa. adv sys sci appl (2021) fig. 7. pcs usage 6. conclusion initial conditions, timers, class concepts, leaf concepts, events, dynamic relationships, variables, and parameters are identified in pre-conceptual schemas as components (elements) for simulating events. such a representation was refined by defining initialization, clock generation, random variable generation, and event generation modules; such modules were proposed for completing components of event simulation in requirements engineering. pcs allow for modeling and simulating events in ssd by integrating simulation and software components in the pcs notation. such components are applied in a lab study of an early earthquake alert generation system in the ssd of geology. event simulation by using the pcs is a new approach for integrating scientific and software components in the same model. thus, scientists and software analysts can use the pcs as a computing model for modeling and simulating events in re for the scientific software development process: they can validate the requirements with scientific stakeholders by using pcs, understand the elements of the complete system, and simulate events and data previously to the development phase of the scientific software. we also evaluate the understanding and the usage the pcs components by the expert opinion was favorable and both scientists and software experts understanded the model. this approach offers a better understanding of real-world events in any domain and consistency among the modeling, simulation, and implementation phases reducing the gap between science and software engineering. we suggest as future work the automation of pre-conceptual schemas for simulating events, which allows for automatically generating results and reports from a simulation software. acknowledgements this paper is a product of the ph.d. research project an extension to pre-conceptual schemas for refining event representation and mathematical notation with hermes code 39886 at universidad nacional de colombia. this work is sponsored by minciencias (ministry of science, technology, and innovation) from colombia in the program ph.d. students in colombia, bid 727 of 2015. 14 p.a. noreña, c.m. zapata copyright ©2021 assa. adv sys sci appl (2021) references 1. bonaventura, m. a., and castro, r. d. (2015). ingeniería guiada por modelado y simulación de eventos discretos: metodología y caso de estudio en la red de datos del experimento tech. rep., atl-daq-proc-2015-024. 2. brosig, f., meier, p., becker, s., koziolek, a., koziolek, h., and kounev, s. (2014). quantitative evaluation of model-driven performance analysis and simulation of component-based architectures. ieee transactions on software engineering, 41(2), 157–175. 3. camejo, i. m., sailema, g. l. a., carrillo, k. m. g., and verdecia, j. a. m. (2018). computational simulation model of milk production process, case study: dairy plant fcp-espoch. kne engineering, 179–191. 4. colombian geological services journals, https://www2.sgc.gov.co/paginas/serviciogeologico-colombiano.aspx. 5. durango, c.e., noreña, p.a., zapata, c.m. (2018). representación de eventos de ruido ambiental a partir de esquemas preconceptuales y buenas prácticas de educción geoespacial de requisitos. research in computing science, 147(2), 327–341. 6. garcía, a., ortega, m., izquierdo, d. (2012). elementos de simulación un enfoque práctico con witness. universidad politécnica de madrid: madrid. 7. gonzález, a., luna, c., abella, r. (2016). uml state machine as modeling language for devs formalism. xlii latin american computing conference (clei), ieee, 1–12. 8. gonzález, a., luna, c., cuello, r., perez, m., daniele, m. (20140). metamodel-based transformation from uml state machines to devs models. xl latin american computing conference (clei), ieee, 1–12. 9. johanson, a. & hasselbring, w. (2018). software engineering for computational science: past, present, future. computing in science & engineering, 20(2), 90–109. 10. kanewala, u.e., bieman, j. (2014). testing scientific software: a systematic literature review. information and software technology, 56(10), 1219–232. 11. kelly, d. (2015). scientific software development viewed as knowledge acquisition: towards understanding the development of risk-averse scientific software. journal of systems and software, 109, 50–61. 12. laurindo, m. peixoto, t., de assis, j. (2019). communication mechanism of the discrete event simulation and the mechanical project software for manufacturing systems. journal of computational design and engineering, 2019, 6(1), 70–80. 13. li, y., guzman, e., tsiamoura, k., schneider, f. bruegge, b. (2015). automated requirements extraction for scientific software. procedia computer science, 51, 582–591. 14. luckham, d. (2011). event processing for business: organizing the real-time enterprise. new jersey: john wiley & sons. 15. luckham, d. (2002). the power of events. an introduction to complex event processing in distributed enterprise systems. boston: addison-wesley, 2002. 16. noreña, p.a., zapata, c.m. (2018). una representación basada en esquemas preconceptuales de eventos determinísticos y aleatorios tipo señal desde dominios de software científico. research in computing science, 147(2), 207–220. 17. noreña, p.a., zapata, c.m. (2018). a pre-conceptual-schema-based representation of time events coming from scientific software domain. 22nd world multi-conference on systemics, cybernetics and informatics (wmsci), 53–58. simulating events in requirements engineering by using pre-conceptual-schema… 15 copyright ©2021 assa. adv sys sci appl (2021) 18. noreña, p.a., zapata, c.m., villamizar, a.e. (2019). representing chemical events by using mathematical notation from pre-conceptual schemas. ieee latin america transactions, 17(01), 46–53. 19. noreña, p.a., zapata, c.m. (2019). business simulation by using events from preconceptual schemas. developments in business simulation and experiential learning 2018, 46, 258–263. 20. noreña, p.a. (2020). an extension to pre-conceptual schemas for refining event representation and mathematical notation. ph.d. thesis, universidad nacional de colombia, medellín campus, colombia. 21. rodríguez, j.m. serrano, d., monleón, t., caro, j. (2008). los modelos de simulación de eventos discretos en la evaluación económica de tecnologías y productos sanitarios.’ gaceta sanitaria 2008, 22, 151–161. 22. rus, i., neu, h., münch, j. (2014). a systematic methodology for developing discrete event simulation models of software development processes. arxiv:14033559 [csse], https://arxiv.org/abs/1403.3559.4 23. song, c., guo, l., wang, n. ma, l. (2014). a grey box testing method for availability simulation software based on event tree model. prognostics and system health management conference (phm-2014 hunan), ieee, 368–372. 24. velásquez, s.: ‘un modelo ejecutable para la simulación multi-física de procesos de recobro mejorado en yacimientos de petróleo basado en esquemas preconceptuales.’ m.sc. thesis, universidad nacional de colombia, medellín campus, colombia, 2019. 25. zapata, c.m. (2012). the unc-method revisited: elements of the new approach. saarbrücken: lambert academic publishing, 2012. 26. zapata-jaramillo, c. m., zapata-tamayo, j. s., & noreña, p. a. (2020). conversión de eventos desde esquemas preconceptuales en código pl/pgsql: simulación de software en la cuarta revolución industrial. revista ibérica de sistemas e tecnologias de informação, (39), 18–34. microsoft word 3. 872_done adv syst sci appl 2023; 02; 11-23 published online at https://ijassa.ipu.ru. __________________________________ *corresponding author: ssg.mech.official@gmail.com an overview of multiple criteria decision making techniques in the selection of best laptop model shankha s. goswami*, dhiren k. behera indira gandhi institute of technology, dhenkanal, sarang, india, 759146 abstract: this research article presents a detailed study of a laptop selection problem by analyzing through three (multi-criteria decision making) mcdm techniques namely, (weighted sum model) wsm, (weighted product model) wpm and (weighted aggregated sum product assessment model) waspas. the main goal of this study is to propose the best laptop model among a group of seven alternative models based on 10 selection criteria. for this purpose, (analytic hierarchy process) ahp is used to determine the criteria weightages whereas, wsm, wpm, waspas techniques are applied individually to determine the best model and to rank the alternatives. the best model is proposed based on the output results of the three mcdm techniques which suggest that alternative 3 and alternative 2 are the optimum and the foulest choice respectively. all the rankings given by the three methodologies are also compared with the previous result which shows that all the three methods are giving the same outcome and the rankings are more or less same with very minor alterations. keywords: laptop selection, multi-criteria decision making, analytic hierarchy process, alternative ranking, weighted sum and product model 1. introduction mcdm have become a trending topic for the decision makers to utilized it as a judgmental tool that helps to make proper selection of products or processes. for any sector, it is very important to make right decision at the right time. the primary problem associated in decision making process is the selection criteria that creates confusion among the buyers while purchasing some items like mobile [15, 16, 21], refrigerator, car [42], air-conditioner [1] etc. let us take an example of a cell phone [21] to explain it clearly, suppose someone wants to buy a cell phone having descent specifications and obviously within the budget, now there are lots of phones available in the market from different brands that create confusion among the buyers which one is the appropriate model to buy and the buyers also have some choices and preferences regarding the phone specifications (i.e. criteria’s) like processor, camera, battery capacity, screen size, etc. ideally, it is not possible that the buyers may get all of their preferable specifications in one model but still they have to select one best model so that most of their requirements can be fulfilled [21]. to get rid from such type of hard situations that arises in our everyday life while taking any decisions, mcdm can guide us to take the right path and provide solution to these types of problems. proper selection of laptop model [14, 17, 23, 27, 35, 44] is such type of problem that came under the category of mcdm, since its selection is based on various criteria like screen resolution, ram, cost, hard disk capacity, brand, color, etc. so, it is quite difficult for the buyers to choose the appropriate laptop model among thousands of available models based on their budget and specifications that perfectly suits their personality. therefore, mcdm would be the suitable tool for investigating the present problem having various conflicting factors. although, 12 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) many other techniques like artificial neural network (ann), linear programming problem (lpp), goal programming etc. are present, but none of the mentioned optimization tools can ever reach the level of mcdm in terms of simplicity and easiness. mcdm techniques are very simple and easy to understand, it requires less computational time, and mcdm necessitates some simple mathematical steps for its operation. on the other hand, alternative optimization process like ann is quite difficult to grab and full of complex calculations. a detailed analysis of the best laptop model selection is presented in this article from which the buyers can develop some idea about the laptop model while purchasing it. not only that, but also the laptop manufacturing companies will be benefitted from the present market scenario that helps to reconstruct their future production according to the present market demand. the ongoing problem is adopted from the paper presented by adali and isik [2], where the selection of laptop is carried out by applying multi-objective optimization on the basis of ratio analysis (moora) [7], full multiplicative form of moora (multimoora) [8] and multiobjective optimization on the basis of simple ratio analysis (moosra) [22,26] methods. a panel of five expert members having high market experience of more than 10 years was formed. the panel board members are higher officials and executives who are associated with some of the reputed electronic firms and laptop manufacturing companies. after going through a massive brainstorming session, they [2] considered 10 selection parameters as follows: processor clock speed (pcs) (in ghz), cache memory (cm) (in mb), storage space (ss) (in gb), display card storage (dcs) (in gb), ram (r) (in gb), screen resolution (sr) (in pixel), screen dimension (sd) (in inch), brand trustworthiness (bt), weight (w) (in kg), and price (p) (in turkish liars). according to the experts, these are the most important factors for a laptop that a customer usually looked for before purchasing it. the seven laptop models were chosen after studying the present market demand and rigorous research of the past laptop sales history. from the records and statistics of some electronic stores and verbal communication with some of the laptop users, the board members sorted out the seven most suitable laptop models for this analysis. the number of alternative laptop models is restricted to 7 because, with the increasing of alternatives and selection criteria, the order of the decision matrix increases; this makes the calculation portion time-consuming and complicated. the selection process was carried out on the basis of 10 selection parameters and the best model was proposed among the seven chosen alternative models. the criteria weightages were determined by adali and isik [2] using ahp [34,38] and the same weightages values are reused in this analysis. the decision makers put their own views and opinions for constructing the ahp pair-wise comparison analysis matrix followed by the criteria weightage’s determination which are as follows: wc1 = 0.297, wc2 = 0.025, wc3 = 0.035, wc4 = 0.076, wc5 = 0.154, wc6 = 0.053, wc7 = 0.104, wc8 = 0.017, wc9 = 0.025, wc10 = 0.214. table 4.1 presents the seven chosen alternative laptop models and their respective specifications (criteria’s). the screen resolution is measured in 1-3 judgment scale which denotes that: 1 – worst (1366 × 768 pixels), 2 – medium (1600 × 900 pixels), 3 – best (1920 × 1080 pixels) and similarly, brand reliability is also measured in 1-10 judgment scale which denotes that: 1 worst and 10 best [2]. a problem similar to the stated above is further extended by implementing three other mcdm techniques, i.e. wsm [11, 30, 33, 45], wpm [32, 45] and waspas [48–50] which are discussed in this article. the output results are also compared with the previous results to check whether the outcomes from all the methods are same or not. an overview of multiple criteria decision making techniques… 13 copyright ©2023 assa adv. in systems science and appl. (2023) 2. literature review many potential researchers have on various mcdm techniques applied in wide variety of areas including industrial sector [9,13,18,19,24], banking and finance sector [25,39], energy sector [20], health and educational sector [6,36], environmental management [37,43], etc. afshari et al. [3] applied simple average weighting (saw or wsm) method to solve the personnel selection problem. men et al. [31] used technique for order preference by similarity to ideal solution (topsis) method to analyze the harmonius society development of east china. zhang et al. [51] established a technology project credit evaluation index system model, where ahp is used to determine the weights of each evaluation index and the evaluation results was analyzed through fuzzy comprehensive evaluation method (fcem). gong et al. [12] applied fuzzy analytic hierarchy process (fahp) for determining the customer requirements in software quality function deployment, and the vitality degree was analyzed using trapezoidal fuzzy function. wang et al. [47] used anp and fuzzy assessment method to assess and rank the risk factor strategies using triangular fuzzy numbers. zavadskas et al. [50] measured the accuracy of wsm and wpm methods and proposed a joint waspas methodology for evaluation of accuracy, based on initial criteria values. bagocius et al. [5] developed a deep-water sea port in the klaipeda region by a combination of waspas and entropy methods. dejus and antucheviciene [10] applied the entropy and waspas method for the assessment and selection of appropriate solutions for occupational safety. siozinyte and antucheviciene [40] proposed a rational solution to improve daylighting in vernacular buildings by applying ahp to determine the weightages and copras, waspas, topsis to rank the alternatives. staniunas et al. [43] evaluated the greenhouse emissions and aimed at decreasing emissions significantly by applying complex proportional assessment (copras), waspas and topsis methodologies. zavadskas et al. [48] selected the most preferable façade’s alternative for public or commercial building by waspas and the results was also validated by moora and multimoora. zolfani et al. [52] used stepwise weight assessment ratio analysis (swara) method to calculate the relative weightages of the criteria and waspas method to evaluate the alternatives for selecting the perfect location to establish a shopping mall. chakraborty and zavadskas [9] explored eight manufacturing decision making problems such as arc welding process, machinability of materials, cutting fluid selection, electro-discharge micro machining process parameters, industrial robot, electroplating system, milling and forging condition by the application of waspas method. liu and lin [29] assessed audit risk judgment by constructing a dominant if rough set based mcdm methodology. karande et al. [24] used couple of mcdm methods, i.e. multimoora, waspas, wpm, moora and wsm to investigate two real time industrial robot selection problems. li et al. [28] proposed a new decision method to address the challenge on criteria reduction in mcdm problems with numerical values based on the rough set theory and the relation of criteria values. venkateswarlu and sarma [46] selected the best supplier for implementing the spring manufacturing industry by incorporating saw and višekriterijumsko kompromisno rangiranje (vikor) methods. apart from this, many researchers also adopted different mcdm methods to choose the best alternatives used for domestic purposes like, srichetta and thurachon [41] evaluated the notebook selection problem by applying fuzzy-ahp. srikrishna et al. [42] used topsis technique for a new car selection problem. lakshmi et al. [27] and tampi et al. [44] selected the best laptop model by using topsis and ahp methodology. adali and isik [1] executed an airconditioner selection problem with copras and additive ratio assessment (aras) methods. kalyani et al. [23] implemented extended topsis method for selecting the best laptop model. 14 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) agajie [4] adopted ahp to select the best personal computer on the basis of three criteria i.e. cost, user-friendliness and software availability. from the above scenario of the literatures it is noticeable that, there are very few research works have been recorded till date where wsm, wpm and waspas methods are used for solving household decision making problems rather than the industrial applications. so, an initiative is taken in this article to utilize these three methods as a decision-making tool for solving a laptop selection problem. moreover, mcdm is a quite straightforward 3. theoretical analysis all the three methods are explained clearly in details step by step under this section. there is not much difference between these methods and the initial steps of all the three methods starts with a decision matrix which is shown below by equation (3.1) followed by the normalization as shown in equations (3.2) and (3.3) [24]. [𝐴 ] × = ⎣ ⎢ ⎢ ⎢ ⎡ 𝑎 𝑎 𝑎 … 𝑎 𝑎 𝑎 𝑎 … 𝑎 𝑎 𝑎 𝑎 … 𝑎 … … … … … 𝑎 𝑎 𝑎 … 𝑎 ⎦ ⎥ ⎥ ⎥ ⎤ , (3.1) where ‘m’ is the number of alternatives and ‘n’ is the number of criteria (i = 1, 2, 3…, m; j = 1, 2, 3…, n). now, the normalization of the decision matrix is done by using equation (3.2) or (3.3) according to the criteria nature. for beneficial criteria (maximum criteria) whose higher values are expected, nij = (3.2) for non-beneficial criteria (minimum or cost criteria) whose smaller values are expected, nij = (3.3) 3.1. weighted sum model (wsm) wsm [11,30] is the most widely used mcdm tool [24] due to its simplicity and less complex calculations. the best option and the ranking are evaluated based on the overall weighted sum of the alternatives which is determined by using equation (3.4) shown below. 𝑊 = ∑ 𝑁 𝑊 (3.4) 3.2. weighted product model (wpm) wpm [32,45] is very similar to wsm [24] but instead of addition, multiplication is done and the criteria weightages are used as the power to the normalized values. overall weighted product is determined for every alternative and the ranking is done based on these values. the weighted product values are determined by using equation (3.5). 𝑊 = ∏ 𝑁 , (3.5) where ‘𝑊 ’ are the ahp computed weightages of criterion in equation (3.4) and (3.5). an overview of multiple criteria decision making techniques… 15 copyright ©2023 assa adv. in systems science and appl. (2023) 3.3. weighted aggregated sum product assessment model (waspas) waspas [48-50] is the combination of wpm and wsm where the aggregated sum product weightage is determined to rank the alternatives. a joint generalized criterion of the weighted summation and the multiplication methods is proposed by zavadskas et al. [48,49] which is shown by equation (3.6) [24]. qi = 0.5𝑊 + 0.5𝑊 = 0.5∑ 𝑁 𝑊 + 0.5∏ 𝑁 , (3.6) where, where qi is the aggregated sum product weightage of the ith alternative (i = 1, 2, 3…, m; j = 1, 2, 3…, n). in order to increase the effectiveness and the ranking accuracy, a more generalized equation was developed by zavadskas et al. [50] which is given by equation (3.7) [24]. qi = λ𝑊 + (1-λ)𝑊 = λ∑ 𝑁 𝑊 + (1-λ)∏ 𝑁 , (3.7) where, where λ = 0, 0.1, 0.2, 0.3…, 1. from equation (3.7), if λ = 0 then it is converted to wpm and if λ = 1 then it is converted to wsm. 4. research methodology adali and isik [2] solved a laptop selection problem by using moora, multimoora and moosra method and the same problem is adopted in this present research work and further extended by analyzing through wsm, wpm and waspas method. in [2], they determined the criteria weightages using ahp which are kept constant in this analysis. the three methods are applied to the above stated problem and the output ranking of the alternatives are compared to the previous results. all the calculation details are covered in this section and the decision matrix as created by adali and isik [2] according to equation (3.1) is shown in table 4.1 which is the initial step of the analysis. table 4.1. selected alternative models and their specifications criteria nature maxi maxi maxi maxi maxi maxi maxi maxi mini mini models pcs cm ss dcs r sr sd bt w p lm1 3.5 6 1256 4 16 3 17.3 8 2.82 4100 lm2 3.1 4 1000 2 8 1 15.6 5 3.08 3800 lm3 3.6 6 2000 4 16 3 17.3 5 2.9 4000 lm4 3 4 1000 2 8 2 17.3 5 2.6 3500 lm5 3.3 6 1008 4 12 3 15.6 8 2.3 3800 lm6 3.6 6 1000 2 16 3 15.6 5 2.8 4000 lm7 3.5 6 1256 2 16 1 15.6 6 2.9 4000 max or min value 3.6 6 2000 4 16 3 17.3 8 2.3 3500 (source: adali and isik, 2017) now the normalization of the decision matrix shown in table 4.1 is done by using equations 3.2 and 3.3 based on the nature of the criteria. the maximum criteria (i.e. beneficial criteria) is normalized using equation 3.2 whereas, the minimum criteria (i.e. non-beneficial) is normalized using equation 3.3. the normalized decision matrix is shown in table 4.2. table 4.2. normalized matrix weights 0.297 0.025 0.035 0.076 0.154 0.053 0.104 0.017 0.025 0.214 16 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) pcs cm ss dcs r sr sd bt w p lm1 0.972 1 0.628 1 1 1 1 1 0.816 0.854 lm2 0.861 0.667 0.5 0.5 0.5 0.333 0.902 0.625 0.747 0.921 lm3 1 1 1 1 1 1 1 0.625 0.793 0.875 lm4 0.833 0.667 0.5 0.5 0.5 0.667 1 0.625 0.885 1 lm5 0.917 1 0.504 1 0.75 1 0.902 1 1 0.921 lm6 1 1 0.5 0.5 1 1 0.902 0.625 0.821 0.875 lm7 0.972 1 0.628 0.5 1 0.333 0.902 0.75 0.793 0.875 (source: author’s own elaboration) the overall weighted sum (𝑊 ) for each alternative is calculated by using equation (3.4). the calculated weighted sum is shown in table 4.3 below. table 4.3. overall weighted sum of the alternatives pcs cm ss dcs r sr sd bt w p weighted sum lm1 0.289 0.025 0.022 0.076 0.154 0.053 0.104 0.017 0.020 0.183 0.943 lm2 0.256 0.017 0.018 0.038 0.077 0.018 0.094 0.011 0.019 0.197 0.743 lm3 0.297 0.025 0.035 0.076 0.154 0.053 0.104 0.011 0.020 0.187 0.962 lm4 0.248 0.017 0.018 0.038 0.077 0.035 0.104 0.011 0.022 0.214 0.783 lm5 0.272 0.025 0.018 0.076 0.116 0.053 0.094 0.017 0.025 0.197 0.892 lm6 0.297 0.025 0.018 0.038 0.154 0.053 0.094 0.011 0.021 0.187 0.897 lm7 0.289 0.025 0.022 0.038 0.154 0.018 0.094 0.013 0.020 0.187 0.859 (source: author’s own elaboration) the overall weighted product (𝑊 ) for each alternative is calculated by using equation (3.5). the calculated weighted sum is shown in table 4.4 below. table 4.4. overall weighted product of the alternatives pcs cm ss dcs r sr sd bt w p weighted product lm1 0.992 1 0.984 1 1 1 1 1 0.995 0.967 0.938 lm2 0.957 0.990 0.976 0.949 0.899 0.943 0.989 0.992 0.993 0.983 0.712 lm3 1 1 1 1 1 1 1 0.992 0.994 0.972 0.959 lm4 0.947 0.990 0.976 0.949 0.899 0.979 1 0.992 0.997 1 0.755 lm5 0.974 1 0.976 1 0.957 1 0.989 1 1 0.983 0.885 lm6 1 1 0.976 0.949 1 1 0.989 0.992 0.995 0.972 0.879 lm7 0.992 1 0.984 0.949 1 0.943 0.989 0.995 0.994 0.972 0.831 (source: author’s own elaboration) by using equation (3.7) a joint weightage of wsm and wpm is determined for every alternative. now, λ value ranges from 0 to 1 at an interval of 0.1. so, eleven joint weightages are proposed for each and every alternative based on eleven λ values i.e. λ = 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 in order to observe that, whether there will be some differences in the final output results and ranking for different λ values. the joint generalized criterion [24,50] weightages for different λ values are calculated and shown in table 4.5. table 4.5. joint generalized criterion weightages for different λ values λ = 0 λ = 0.1 λ = 0.2 λ = 0.3 λ = 0.4 λ = 0.5 λ = 0.6 λ = 0.7 λ = 0.8 λ = 0.9 λ = 1 an overview of multiple criteria decision making techniques… 17 copyright ©2023 assa adv. in systems science and appl. (2023) lm1 0.938 0.939 0.939 0.940 0.940 0.941 0.941 0.941 0.942 0.942 0.943 lm2 0.712 0.715 0.718 0.721 0.724 0.727 0.730 0.733 0.737 0.740 0.743 lm3 0.959 0.959 0.959 0.959 0.960 0.960 0.960 0.961 0.961 0.961 0.962 lm4 0.755 0.758 0.761 0.764 0.766 0.769 0.772 0.775 0.777 0.780 0.783 lm5 0.885 0.885 0.886 0.887 0.888 0.888 0.889 0.890 0.891 0.892 0.892 lm6 0.879 0.881 0.882 0.884 0.886 0.888 0.890 0.891 0.893 0.895 0.897 lm7 0.831 0.833 0.836 0.839 0.842 0.845 0.848 0.850 0.853 0.856 0.859 (source: author’s own elaboration) from the above table 4.5 it can be clearly seen that, for λ = 0, the joint weightages are exactly same as the weighted product shown in table 4.4 since the first part of the equation (3.7) is eliminated in this case and on the other hand, when λ = 1, the second part of the equation (3.7) got eliminated so the joint weightages exactly matches with the weighted sum shown in table 4.3. 5. results and discussions the weighted sum, weighted product and the joint generalized weightages are determined for every alternative by applying wsm, wpm and waspas method which are presented in table 4.3, table 4.4 and table 4.5 respectively. now according to the decreasing magnitude of these weightages, three ranking order of the alternatives can be done for each and individual methods which are presented in table 5.1 (wsm and wpm) and table 5.2 (waspas). table 5.1. ranking of the alternatives by wsm and wpm wsm wpm alternatives weighted sum rank weighted product rank lm1 0.943 2 0.938 2 lm2 0.743 7 0.712 7 lm3 0.962 1 0.959 1 lm4 0.783 6 0.755 6 lm5 0.892 4 0.885 3 lm6 0.897 3 0.879 4 lm7 0.859 5 0.831 5 (source: author’s own elaboration) table 5.2. ranking of alternatives by waspas alternatives λ = 0 rank λ = 0.1 rank λ = 0.2 rank λ = 0.3 rank λ = 0.4 rank λ = 0.5 rank lm1 0.938 2 0.939 2 0.939 2 0.940 2 0.940 2 0.941 2 lm2 0.712 7 0.715 7 0.718 7 0.721 7 0.724 7 0.727 7 lm3 0.959 1 0.959 1 0.959 1 0.959 1 0.960 1 0.960 1 lm4 0.755 6 0.758 6 0.761 6 0.764 6 0.766 6 0.769 6 lm5 0.885 3 0.885 3 0.886 3 0.887 3 0.888 3 0.888 3 lm6 0.879 4 0.881 4 0.882 4 0.884 4 0.886 4 0.888 4 lm7 0.831 5 0.833 5 0.836 5 0.839 5 0.842 5 0.845 5 alternatives λ = 0.6 rank λ = 0.7 rank λ = 0.8 rank λ = 0.9 rank λ = 1 rank lm1 0.941 2 0.941 2 0.942 2 0.942 2 0.943 2 lm2 0.730 7 0.733 7 0.737 7 0.740 7 0.743 7 lm3 0.960 1 0.961 1 0.961 1 0.961 1 0.962 1 18 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) lm4 0.772 6 0.775 6 0.777 6 0.780 6 0.783 6 lm5 0.889 4 0.890 4 0.891 4 0.892 4 0.892 4 lm6 0.890 3 0.891 3 0.893 3 0.895 3 0.897 3 lm7 0.848 5 0.850 5 0.853 5 0.856 5 0.859 5 (source: author’s own elaboration) from the above two tables 5.1 and 5.2, it can be clearly seen that all the three mcdm techniques are giving the same output results and suggest that model 3 is the best laptop model among the group followed by model 1 in the 2nd position whereas, model 2 is the worst choice. all the rankings are more or less same with very little alteration. there is a very tough competition between model 5 and model 6 for the 3rd and 4th place. from table 5.1, wsm is suggesting that model 6 should come in the 3rd place whereas, wpm is suggesting that model 5 should occupy that place. to remove this confusion waspas is applied and the variations in ranking for all the values of λ is observed. it is noticeable in table 5.2 that the ranking matches with the wpm ranking as the λ value varies from 0 to 0.4, since the second part of the equation (3.7) is given more importance in the favor of wpm. similarly, as the λ value varies from 0.6 to 1 the ranking shifted towards the wsm, since the second part of the equation (3.7) slowly tends to zero and took the favor of wsm. so, in waspas method equal importance should be given to each wsm and wpm method by keeping the value of λ as 0.5 and it is suggested to use equation (3.6) while calculating the joint generalized weights. hence, the ranking given by waspas method when λ = 0.5 can be accepted as the final ranking. it should be noted, that the joint weights for model 5 and model 6 are given as same i.e. 0.888 in table 5.2 under the column λ = 0.5 is due to round off error. if up to five decimal places are considered then for model 5 and model 6 the joint weights would be 0.88850 and 0.88775 which clearly shows that model 5 should occupy the 3rd position and model 6 should come in the 4th place. the present three rankings are also compared with the previous researchers results in table 5.3 and graphically in fig. 5.1 as proposed by adali and isik [2] using moora, multimoora and moosra methods. table 5.3. ranking comparisons of different mcdm methods alternatives wsm wpm waspas moora multimoora moosra final ranking lm1 2 2 2 2 2 2 2 lm2 7 7 7 7 7 7 7 lm3 1 1 1 1 1 1 1 lm4 6 6 6 6 6 6 6 lm5 4 3 3 3 3 3 3 lm6 3 4 4 4 4 4 4 lm7 5 5 5 5 5 5 5 (source: adali and isik, 2017; author’s own elaboration) an overview of multiple criteria decision making techniques… 19 copyright ©2023 assa adv. in systems science and appl. (2023) fig. 5.1. graphical comparisons of the proposed rankings (source: author’s own elaboration; created using ms word chart option) 6. conclusion from the above research study, it can be concluded that model 3 is the best laptop model followed by model 1 among these chosen seven alternatives and model 2 is the worst model to buy for the customers. the best and the worst choice of the alternatives suggested by all the six methods are exactly the same but, there is a very minor alternation in the wsm ranking regarding the 3rd and 4th places. since, the ranking of the rest five mcdm methods shown in table 5.3 exactly matches with each other so, the final ranking of the laptop models can be proposed as follows: model 3 > model 1 > model 5 > model 6 > model 7 > model 4 > model 2 limitations: there are lots of mcdm tools available whose implementation to the same problem may generate different output results and ranking. moreover, if different subjective or objective weightage calculation methods are used other than ahp like, criteria importance through inter criteria correlation (critic), entropy, best worst method (bwm) etc. it may also alter the final results. future scope: different mcdm methods can be applied to the above problem and the ranking can be compared with these proposed rankings. the criteria weightages can also be determined by other methods and the variation in results can be noted. not only these but also, mcdm tools can also be utilized as a decision tool in the selection of other electronic gadgets and household appliances e.g. television, washing machine, camera etc. acknowledgement we are grateful to igit sarang for providing us such a wonderful research environment. we are thankful to bput rourkela for their immense cooperation and support. we are also obliged to aicte india for their monthly financial assistance. 0 1 2 3 4 5 6 7 8 lm1 lm2 lm3 lm4 lm5 lm6 lm7 wsm wpm waspas moora multimoora moosra final ranking 20 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) references [1] adali, e. a., & isik, a. t. (2016). air conditioner selection problem with copras and aras methods, manas journal of social studies, 5(2), 124-138. [2] adali, e. a., & isik, a. t. (2017). the multi-objective decision making methods based on multimoora and moosra for the laptop selection problem, journal of industrial engineering international, 13, 229–237, https://doi.org/10.1007/s40092-016-0175-5. [3] afshari, a., mojahed, m., & yusuff, r. m. (2010). simple additive weighting approach to personnel selection problem, international journal of innovation, management and technology, 1(5), 511–515. [4] agajie, t. m. (2017). application of analytic hierarchy process in the case of purchasing, international journal of engineering science and computing, 7(4), 10357–10364. [5] bagočius, v., zavadskas, e. k., & turskis, z. (2013). multi-criteria selection of a deepwater port in klaipeda, procedia engineering, 57, 144–148, https://doi.org/10.1016/ j.proeng.2013.04.021. [6] bhattacharyya, a., & chakraborty, s. (2014). a dea-topsis-based approach for performance evaluation of indian technical institutes, decision science letters, 3(3), 397–410. [7] brauers, w. k. m., & zavadskas, e. k. (2006). the moora method and its application to privatization in a transition economy, control and cybernetics, 35(2), 445–469. [8] brauers, w. k. m., & zavadskas, e. k. (2010). project management by multimoora as an instrument for transition economies, technological and economic development of economy, 16(1), 5–24. [9] chakraborty, s., & zavadskas, e. k. (2014). applications of waspas method in manufacturing decision making, informatica, 25(1), 1–20. [10] dėjus, t., & antuchevičienė, j. (2013). assessment of health and safety solutions at a construction site, journal of civil engineering and management, 19(5), 728–737. [11] fishburn, p. c. (1967). additive utilities with incomplete product set: applications to priorities and assignments. baltimore, md: operations research society of america. [12] gong, l., yang, m., guo, s., & wang, q. (2012). a method for determining importance degree of customer requirements in software quality function deployment, advances in systems science and applications, 12(1), 46–53. [13] goswami, s. s., & behera, d.k. (2020). implementation of entropy-aras decision making methodology in the selection of best engineering materials, materials today: proceedings, 38, 2256–2262, https://doi.org/10.1016/j.matpr.2020.06.320. [14] goswami, s. s., behera, d. k., & mitra, s. (2020). a comprehensive study of weighted product model for selecting the best laptop model available in the market, brazilian journal of operations & production management, 17(2), e2020875, https://doi.org/10.14488/ bjopm.2020.017. [15] goswami, s. s., & behera, d. k. (2021). evaluation of best smartphone model in the market by integrating fuzzy-ahp and promethee decision making approach, decision, 41(1), 71–96, https://doi.org/10.1007/s40622-020-00260-8. [16] goswami, s. s., & behera, d. k. (2021). an analysis for selecting best smartphone model by ahp-topsis decision making methodology, international journal of service science, management, engineering, and technology, 12(3), 116–137, http://doi.org/10.4018/ ijssmet.2021050107. [17] goswami, s. s., & behera, d. k. (2021). best laptop model selection by applying integrated ahp-topsis methodology, international journal of project management and productivity assessment, 9(2), 29–47, http://doi.org/10.4018/ijpmpa.2021070102. an overview of multiple criteria decision making techniques… 21 copyright ©2023 assa adv. in systems science and appl. (2023) [18] goswami, s. s., & behera, d. k. (2021). solving material handling equipment selection problems in an industry with the help of entropy integrated copras and aras mcdm techniques, process integration and optimization for sustainability, 5, 947–973, https://doi.org/ 10.1007/s41660-021-00192-5. [19] goswami, s. s., et al. (2021). analysis of a robot selection problem using two newly developed hybrid mcdm models of topsis-aras and copras-aras, symmetry, 13, 1331, https://doi.org/10.3390/sym13081331. [20] goswami, s. s., mohanty, s. k., & behera, d. k. (2021). selection of a green renewable energy source in india with the help of merec integrated piv mcdm tool, materials today: proceedings, 52, 1153–1160, https://doi.org/10.1016/j.matpr.2021.11.019. [21] goswami, s. s., & mitra, s. (2020). selecting the best mobile model by applying ahpcopras and ahp-aras decision making methodology, international journal of data and network science, 4(1), 27–42. [22] jagadish, & ray, a. (2014). green cutting fluid selection using moosra method, international journal of research in engineering and technology, 3(3), 559–563. [23] kalyani, k. s., nagarani, s., maragatham, l., & kumar, n. d. (2016). multi-criteria decision making for selecting the best laptop, international journal of control theory and applications, 9(36), 437–441. [24] karande, p., zavadskas, e. k., & chakraborty, s. (2016). a study on the ranking performance of some mcdm methods for industrial robot selection problems, international journal of industrial engineering computations, 7(3), 399–422. [25] korhonen, p., koskinen, l., & voutilainen, r. (2006). a financial alliance compromise between executives and supervisory authorities, european journal of operational research, 175(2), 1300–1310, https://doi.org/10.1016/j.ejor.2005.06.033. [26] kumar, r., & ray, a. (2015). selection of material under conflicting situation using simple ratio optimization technique. in k. das, k. deep, m. pant, j. bansal, & a. nagar (eds.), advances in intelligent systems and computing: vol. 335. proceedings of fourth international conference on soft computing for problem solving (pp. 513–519). springer. [27] lakshmi, t. m., venkatesan, v. p., & martin, a. (2015). identification of a better laptop with conflicting criteria using topsis, i.j. information engineering and electronic business, 6, 28–36, https://doi.org/10.5815/ijieeb.2015.06.05. [28] li, z., liu, j., zhang, z., & liu, s.-feng (2016). a new decision method for multicriteria decision making with numerical values based on criteria reduction, advances in systems science and applications, 16(4), 29–42. [29] liu, y., & lin, y. (2015). a novel multi-attribute decision making methodology and application, advances in systems science and applications, 15(3), 202–219. [30] maccrimon, k. r. (1968). decision making among multiple attribute alternatives: a survey and consolidated approach. santa monica, ca: rand corporation. [31] men, k., jiang, l., & liu, j. (2010). research on comprehensive evaluation of harmonious society in east china, advances in systems science and applications, 10(1), 48–54. [32] miller, d. w., & starr, m. k. (1969). executive decisions and operations research (2nd ed.). englewood cliffs, nj: prentice-hall. [33] mitra, s., & goswami, s. s. (2019). application of simple average weighting optimization method in the selection of best desktop computer model, advanced journal of graduate research, 6(1), 60–68, https://doi.org/10.21467/ajgr.6.1.60-68. [34] mitra, s., & goswami, s. s. (2019). selection of the desktop computer model by ahptopsis hybrid mcdm methodology, international journal of research and analytical reviews, 6(1), 784–790, http://doi.one/10.1729/journal.19551. 22 s.s. goswami, d.k. behera copyright ©2023 assa adv. in systems science and appl. (2023) [35] mitra, s., & goswami, s. s. (2019). application of integrated mcdm technique (ahpsaw) for the selection of best laptop computer model, international journal for research in engineering application & management, 4(12), 1–6, https://doi.org/10.18231/24549150.2019.0091. [36] podgórski, d. (2015). measuring operational performance of osh management system a demonstration of ahp-based selection of leading key performance indicators, safety science, 73, 146–166, https://doi.org/10.1016/j.ssci.2014.11.018. [37] qin, x. s., huang, g. h., chakma, a., nie, x. h., & lin, q. g. (2008). a mcdmbased expert system for climate change impact assessment and adaptation planning. a case study for the georgia basin, canada, expert systems with applications, 34(3), 2164–2179, https://doi.org/10.1016/j.eswa.2007.02.024. [38] saaty, t. l. (1980). the analytic hierarchy process. new york, ny: mcgraw-hill. [39] seçme, n. y., bayrakdaroglu, a., & kahraman, c. (2009). fuzzy performance evaluation in turkish banking sector using analytic hierarchy process and topsis, expert system applications, 36(9), 11699–11709, https://doi.org/10.1016/j.eswa.2009.03.013 [40] šiožinytė, e., & antuchevičienė, j. (2013). solving the problems of daylighting and tradition continuity in a reconstructed vernacular building, journal of civil engineering and management, 19(6), 873–882. [41] srichetta, p., & thurachon, w. (2012). applying fuzzy analytic hierarchy process to evaluate and select product of notebook computers, international journal of modeling and optimization, 2(2), 168–173, http://dx.doi.org/10.7763/ijmo.2012.v2.105. [42] srikrishna, s., reddy, a. s., & vani, s. (2014). a new car selection in the market using topsis technique, international journal of engineering research and general science, 2(4), 177–181. [43] staniūnas, m., medineckienė, m., zavadskas, e. k., & kalibatas, d. (2013). to modernize or not: ecological–economical assessment of multi-dwelling houses modernization, archives of civil and mechanical engineering, 13(1), 88–98, https://doi.org/10.1016/ j.acme.2012.11.003. [44] tampi, y. a. n., pangemanan, s. s., & tumewu, f. j. (2016). consumer decision making in selecting laptop using analytical hierarchy process (ahp) method (study: hp, asus and toshiba), jurnal riset ekonomi, manajemen, bisnis dan akuntansi, 4(1), 315–322. [45] triantaphyllou, e., & mann, s. h. (1989). an examination of the effectiveness of multidimensional decision-making methods: a decision-making paradox, decision support systems, 5(3), 303–312, https://doi.org/10.1016/0167-9236(89)90037-7. [46] venkateswarlu, p., & sarma, d. b. (2016). selection of supplier by using saw and vikor methods, international journal of engineering research and application, 6(9), 80–88. [47] wang, y., song, j., & zhang, j. (2012). research on chinese airline strategy risk based on anp and the fuzzy assessment method, advances in systems science and applications, 12(2), 122–132. [48] zavadskas, e. k., antucheviciene, j., šaparauskas, j., & turskis, z. (2013a). multicriteria assessment of facades’ alternatives: peculiarities of ranking methodology, procedia engineering, 57, 107–112, https://doi.org/10.1016/j.proeng.2013.04.016. [49] zavadskas, e. k., antucheviciene, j., saparauskas, j., & turskis, z. (2013b). mcdm methods waspas and multimoora: verification of robustness of methods when assessing alternative solutions, economic computation and economic cybernetics studies and research, 47(2), 5–20. an overview of multiple criteria decision making techniques… 23 copyright ©2023 assa adv. in systems science and appl. (2023) [50] zavadskas, e. k., turskis, z., antucheviciene, j., & zakarevicius, a. (2012). optimization of weighted aggregated sum product assessment, elektronika ir elektrotechnika, 122(6), 3–6, https://doi.org/10.5755/j01.eee.122.6.1810. [51] zhang, b., wang, y., & chen, d. (2011). a research on technology project credit evaluation model based on ahp and fcem, advances in systems science and applications, 11(3–4), 249–256. [52] zolfani, s. h., aghdaie, m. h., derakhti, a., zavadskas, e. k., & varzandeh, m. h. m. (2013). decision making on business issues with foresight perspective; an application of new hybrid mcdm model in shopping mall locating, expert systems with applications, 40(17), 7111–7121, https://doi.org/10.1016/j.eswa.2013.06.040. adv syst sci appl 2020; 04:36–44 published online at https://ijassa.ipu.ru. statistical properties of vanet-based information spreading imre varga1∗, gergely kocsis1 1university of debrecen, department of it systems and networks, debrecen, hungary abstract: in this work, we present a robust model of vehicular ad hoc networks (vanet) in order to study information spreading on such topologies. vehicles are moving along the fastest routes between their starting points and their destinations on a map derived from a real urban topology. vehicles can exchange information through the use of short-range wireless communication devices. the source of information is a roadside unit which provides packets for the nearby vehicles. these agents can carry and forward messages for others. as a result of the data dissemination, the major part of the system becomes informed quite quickly. presented results include the investigation of information spreading in the system, e.g. the time evolution of the average awareness, the age distribution of information owned by separate vehicles and the statistical properties of time intervals between information exchange events. it has been shown how the effectiveness of this complex system depends on the density of intelligent devices. scalefree behavior was found by time series analysis. our computer simulation results can help to design smartcity applications of the future. keywords: agent-based simulation, vanet, information diffusion, traffic simulation, time series analysis 1. introduction the spreading of information in vehicular networks plays a key role in many smart city services. because of this, the topic is in the focus of research in the last decade. the aim of these applications is to make urban traffic safer and comfortable. previous studies have analyzed the topological properties of urban road maps [1, 2] and the traffic flow was measured and studied [3] in some other works in order to increase the efficiency of these intelligent transportation systems. several different algorithms and methods were developed to simulate the motion of vehicles and generate traffic in urban or in highway environment [4–6]. the communication of moving wireless devices (possibly carried by members of these vehicular networks) may be described by using standardized communication protocols, like dedicated short range communication (dsrc) or ieee 802.11p standard [7–9]. in vanets (vehicular ad hoc network) both the routing [10–12] and the broadcasting [4,13] are actively investigated fields. distribution of information was also studied in different wireless systems, e.g. in mobile peer-to-peer systems [14] or self-organized sensor networks [15]. some questions related to the statistical properties of the general spreading processes in vanets are however still open. the goal of our research is to introduce and analyze a new framework in order to be able to answer some of the following questions. what are the limits of the information spreading? can we reach all actors of the traffic system based only on self-organization? do all vehicles own up-to-date information? similar questions have been appeared and already answered in ∗corresponding author: varga.imre@inf.unideb.hu statistical properties of vanet-based information spreading 37 social networks [16,17], but due to the continuously changing topology, the characteristics of spreading can be very different. in section 2 the construction of the realistic urban topology is shown. section 3 presents the details of the simulation of vehicular motion. information spreading based on carry-andforward and multihop broadcast dissemination schemes is presented in section 4 and then the first results of our studies are shown in section 5. at the end we close by some conclusions. 2. underlying map topology of simulations we applied a real city map, namely the map of our home city as the underlying network topology to reach a realistic simulation environment. the dataset describing the map was gained and is available at the page of the openstreetmap project [18]. since in this work we would like to use this map for agent-based simulation a much more simplified topology is needed. because of this the source was reduced keeping the topology of the crossroad network and the distances between junctions, but losing the real geographical locations of road sections. according to the original osm format any crooked road can be built up from shorter straight segments and the geographical coordinates of their endpoints. in this way, a road section between crossroads can be described by a list of internal nodes with degree 2. in our approach, the shape of a road section is negligible and only the length of the section is important. this was the base of our topology simplifying method. in case of any two road segments between nodes a−b and nodes b − c, node b was eliminated if it has no other neighbors than a and c, merging the segments to only one longer segment between nodes a− c with a distance equal to the sum of lengths of the previous two segments. after the above reduction process the resulting network can be analyzed from two different points of view. (i) taking a look to the network as an abstract topology (undirected graph) we can find that there are 3422 nodes (junctions) connected by 4812 links (road sections). it was found that 84% of these nodes have a degree of 3 or 4. in the unit of link number (ignoring road length) the diameter of the network is 96. (ii) from the geographical aspect however our network still has some spatial properties. taking into account the distances, the average distance between two crossroads is 121.5m, however, the distribution is quite wide, there are almost 3 orders of magnitude difference between the shortest and the longest road section. the average distance between two randomly chosen nodes is 4.1± 1.9km. (more details are available in [19, 20].) 3. motion of vehicles on the above described map vehicles are moving from their randomly chosen starting node toward their randomly chosen destination node along the fastest path. even though today many different navigation options are available to be considered, practice shows that drivers usually use a route with the shortest travel time instead of the shortest distance route, or the smallest number of left turns [21]. the original data set contains information about the rank of all road segments (e.g.: primary, secondary, residential, living street, etc.). the average speeds of cars depend on the rank of the road. based on the speed prediction/offer of the google maps [22], different average speed is applied in case of different road ranks. thus the shortest and the fastest route can be different. vehicles move with constant speed between two neighboring nodes, at a crossroad they turn according to their route (and perhaps change speed). traffic jams, traffic lights or the finite size of vehicles are not taken into account during the simulation because from the point of the later spreading process the short-term fluctuations of the speed of cars are negligible. thus, in this kind of mean-field approach, the velocity of vehicles is not influenced by the copyright c© 2020 assa. adv syst sci appl (2020) 38 i. varga, g. kocsis actual traffic situations. even though, the source and destination nodes are random the density of the traffic is really diverse due to the topology (connectivity, ranks). we assumed that the number of moving cars in the system at a given time can be constant since the simulated (typically few 10 minutes long) time interval is small compared to the daily life cycle of a city or the duration of rush-hour traffic. in this way, different scenarios (e.g. rush-hours or off-peak time) can be simulated separately using a distinct number of cars. at the beginning of the simulation, there are n vehicles in the system. later, when a vehicle arrives to its destination, it is removed and immediately a new one is initialized and started. at the beginning of the simulation, all the cars are just departed. it is easy to understand that in order to avoid artificial transient effects the measurement related to spreading is started only later (t = 0) when the system becomes randomized, however, the simulation of the traffic is started at t = −t0. the length of the randomization time interval (−t0 ≤ t < 0) is longer then most of the trips (t0 = 750s, average travel time is 459± 261s), so when the scientific observation is started all the initial cars have been arrived and others are launched in different time moments. the simulation is stopped at t = tmax. the system evolves in discrete time steps. the time step ∆t is small enough to move only a few meters, so it is tiny compared to the whole simulation time ∆t� t0 + tmax. the time interval of the analysis (0 ≤ t ≤ tmax) is long enough to cover several generations of vehicles. we found that in order to get reliable results the total number of simulated cars (nt) has to be at least five times greater than number of cars at a given moment (nt > 5n ). this approach of macroscopic traffic is quite simplified. the motion of vehicles is much slower than, spreading of information (detailed in the next section) due to the consecutive quick message forwarding. this fact allows us to neglect more details of the micro-level traffic. in addition to the approximate spatial distribution of agents, motion is only required to regularly change the communication topology. 4. spreading of information in the system, smart vehicles are represented by agents able to interact by short-range communication. if the distance of two vehicles at a given time moment is less than the range r of the wireless communication, they can exchange information. if agent i can receive information from agent j, communication to the opposite direction is also possible technically. based on this, in our model the agents can have two different states. on the one hand agent i can be uninformed, so it has not received any data (denoted by si = 0). on the other hand, it can be informed, so it has already got some data (denoted by si = 1). beside this vehicle-to-vehicle (v2v) communication, there is infrastructure-to-vehicle communication (i2v) as well. in the latter case the on board units (obu) of smart vehicles can receive information from road side units (rsu). in our model initially all agents are in uninformed state and only one rsu is present, playing the role of an information source. when an agent passes close enough to the rsu it receives new up-to-date public information (e.g. traffic or weather alert). the actual content of messages is negligible. the agent stores it together with the actual timestamp and later it shares with others within the communication range. if one of these neighboring agents is uninformed it becomes informed. if both agents that are in contact have been already informed, the agent with an older timestamp will update its knowledge storing the newer information with its given timestamp. thus information can spread in this dynamically changing network from the rsu to any vehicle even if they have never passed by the rsu. in order to characterize agent i in detail we introduce the quantity ti which is the latest/newest timestamp of information owned by the informed agent i or ti = −1 if agent i is uninformed. (so ti > 0 is the simulation time when the given information unit carried by agent i entered into the system by the rsu.) the behavior of the system is shown in fig. 4.1. a vehicle (agent i) proceeds from node a to node d. it goes by the rsu in node b receiving new information at t = t . an other vehicle (agent j) moves from node e towards node f . both of them are copyright c© 2020 assa. adv syst sci appl (2020) statistical properties of vanet-based information spreading 39 fig. 4.1. the behavior of the system. in the vicinity of the node c at the same time. since they are within the range r, agent i can transmits the information to agent j. between nodes a and b agent i is uninformed, but between b and d it is in an informed state, having timestamp t . agent j becomes informed at node c and possesses also timestamp t between nodes c and f . at simulation time t, an informed agent i has information with age ai = t− ti. the average age of information 〈a〉 owned by agents in a given time moment can be written as 〈a〉 = ∑ i tisi n inf , (4.1) where n inf is the number of informed agents, defined as n inf = ∑ i si. a large value of n inf/n indicates extensive information spreading. when the average age of information 〈a〉 is low, it means that our smart traffic system is in an up-to-date phase. thus the number of informed agents n inf and the average age of information 〈a〉 are good measures of the effectiveness of information spreading in vanet. 5. results since during the simulation an si (susceptible-infected) model [23] is applied, initially more and more agents become informed. however, the system never reaches a fully informed state, because during the simulation new, uninformed agents appear in the system, while informed ones disappear as they reach their destination. investigating the time evolution of the agents it was found that the system reaches a steady state described by saturating functions. in fig. 5.1 a, one can observe that at t = 0 (when the rsu is just activated) there are no informed agents in the system, but soon some agents pass by the information source of the infrastructure. then the vehicles carry the information during their motion to different places of the city meanwhile they behave as secondary information sources speeding up the spreading of information so leading to increasingn inf (t)/n function with a significant slope. after a quite short time period, spreading slows down resulting in saturation of the number of informed agents. the average movement of vehicles during a simulation step ∆t is the half of the applied range of communication r. (of course, increasing range r speeds up the spreading.) the reason why the system reaches this almost steady state is the fact that the propagation of information can be faster than the motion of vehicles. the saturation level depends on the number of agents (the density of smart vehicles in the city) and in most cases the n inf (t)/n curves never reach 1.0. this is shown in fig. 5.1 b. as one can observe the information coverage of vanet can be effective only if the number of smart vehicles copyright c© 2020 assa. adv syst sci appl (2020) 40 i. varga, g. kocsis 0.0 0.2 0.4 0.6 0.8 1.0 n in f /n 0 500 1000 1500 2000 t n=100 n=1000 n=10000 a) 0.0 0.2 0.4 0.6 0.8 1.0 n in f (t m a x )/ n 10 10 2 10 3 10 4 n b) fig. 5.1. a) number of informed agents (vehicles) as a function of time for different numbers of agents. after a short time period a saturation is achieved at a quite high value. b) the previous saturation level depends on the number of vehicles in the system (of course more smart vehicle leads to higher level of awareness). exceeds a given threshold (about few hundreds of vehicles in the case of the medium-sized city debrecen). even though in most cases the number of informed agents n inf is proved to be relatively high in the system, the really important questions are the following ones. how old is the average information? is the system in an up-to-date phase continuously? the average information age as a function of time 〈a〉 (t) can give the answers. as it is shown on fig. 5.2 a, most of the agents have relatively young information. recent information from rsu overwrites the system very quickly without any outer control. of course the saturation level of 〈a〉 (t) (far from the opening time period) is determined by the number of agents. more smart vehicles lead to a more up-to-date system. (see the fig. 5.2 b.) the average age of information is even less than the length of the time period needed to reach the saturation of the number of informed agents. 0 100 200 300 400 500 600 < a > 0 500 1000 1500 2000 t n=1000 n=10000 a) 0 200 400 600 < a > (t m a x ) 10 10 2 10 3 10 4 n b) fig. 5.2. a) the average age of information owned by the vehicles as a function of time. it shows saturation for different system size. b) the average age of information in the saturation phase decreases logarithmically with the number of vehicles, so a denser vehicle park in the city results in a more up-to-date system. the above averages describe the system on a macro scale. for microscale characterization, we analyzed the time intervals between information exchanges for all agents. the average information age can be low only if the time interval ∆trecv between two subsequent information receive events is short for each vehicle. as one can see in fig. 5.3 a, the copyright c© 2020 assa. adv syst sci appl (2020) statistical properties of vanet-based information spreading 41 1 10 -1 10 -2 10 -3 10 -4 p ( t r e cv ) 10 10 2 10 3 trecv n=5000 n=10000 n=15000 a) 1 10 -1 10 -2 10 -3 10 -4 p ( t f o rw ) 10 10 2 10 3 tforw n=5000 n=10000 n=15000 b) fig. 5.3. the distribution of the average message receiving (a) and forwarding (b) time interval of vehicles. they obey a power law with exponents around 3.4 (a) and 2.5 (b) respectively, independently of the vehicle density. distribution of ∆trecv has a power-law form: p(∆trecv) ∼ ∆trecv −γrecv . this means that most of the agents frequently receive new information packets, but some of them keep old information for relatively long time. the density of vehicles has no effect on the value of the exponent which is close to γrecv ≈ 3.4. the distribution of the time intervals between message forwards ∆tforw also proved to be scale-free, however its exponent γforw is definitely lower, it is around 2.5. (see fig. 5.3 b.) one can ask why this difference in the distributions appear even though the number of receive and forward events are the same. this can be explained however by conditions of information exchange. when vehicles approach each other within the communication range, only the up-to-date agent can forward its recent information. in this way, an agent, who has a quite old information packet has to wait a long time to get a chance to forward it. contrarily old information can be overwritten by almost any other message, so the high receiving chance leads to very few old information to overwrite. on micro scale, one can follow the spreading of each information holding any given timestamp t . agents receive new information from the rsu at time t = t . then it is transmitted to other vehicles, so the number of agents nt having the same timestamp t is increasing. sooner or later they will be updated so nt (t) is decreasing, finally all of these information units will disappear at t = t + tl, where tl denotes the lifetime of the given information. note that tl is not the lifetime of a given information packet, but the time interval when the given information (with a given timestamp) is present somewhere in the system. the time evolution of nt (t) is qualitatively similar for all timestamps, however quantitatively they are very different. the general form of thent (t) function can be obtained by rescaling and averaging all the separate curves. the result is illustrated in fig. 5.4 a. it shows that at the beginning information spreads very quickly but after te time the expansion reaches its maximum, where m = nt (t + te) is the maximal number of agents carrying the information originated at time t . after the expansion, we found a dying out phase with duration td = tl − te . the whole average lifetime 〈tl〉 of information in a system depends on the number of vehicles proceeding in the town. see fig. 5.4 b. the curve has a minimum. in case of small systems, more agents result in lower lifetime tl due to the increasing number of updates (younger information) related to fig. 5.2 b. in order to understand the increasing regime, we have to study the maximum of nt (t− t ) denoted by m . figure 5.4 c illustrates that for large system 〈m〉/n increases linearly with n , so an increasing proportion of agents is carrying the given timestamp in the whole population. if there are more smart vehicles in the city due to the fast spreading a few information timestamps dominate. there can be hundreds of different information in the system, however a dominant part of agents carries copyright c© 2020 assa. adv syst sci appl (2020) 42 i. varga, g. kocsis 0 20 40 60 80 n t 0 40 80 120 t-t te tl m a) 0 100 200 300 400 500 600 < t l > 0 3000 6000 9000 12000 15000 n 0 3000 6000 9000 12000 15000 n b) 0.0 0.02 0.04 0.06 < m > /n 0 7000 14000 n c) 0.0 0.2 0.4 0.6 0.8 < t e > / < t l > 0 3000 6000 9000 12000 15000 n d) fig. 5.4. a) the average time evolution of the number of vehicles carrying information with a given timestamp (t ) for n = 5000. a shorter expansion period te is followed by a longer dying out period separated by a maximum m at time t = t + te . b) the average lifetime of information with a given timestamp as a function of the number of cars. increasing vehicle density first leads to decreasing lifetime, but above a certain number of vehicles, given information can present in the system for a longer time interval. c) the maximal ratio of vehicles carrying given information with timestamp t at t = t + te is increasing linearly for large systems. solid, gray line just guides the eyes indicating linear dependence. d) the ratio of the expansion period and the lifetime as a function of the number of agents n . the expansion phase is shrinking by increasing the number of agents (indicating speeding up of the spreading since m is increasing). only a few timestamps. in case of a given timestamp the number of agents carrying it can be a few percent of the population, so hundreds of vehicles. the time needed to overwrite all of them is long, thus the average lifetime of timestamps 〈tl〉 can increase with n . the expansion phase is usually shorter than the dying out phase, but their ratio is not fixed. the 〈te〉/〈tl〉 ratio as a function of the n number of smart vehicles can be characterized by an almost linear decreasing curve in the studied systems. (see fig. 5.4 d.) in case of n = 103 agents, more than 40% of the lifetime of the average timestamp is in the expansion phase, while in case of n = 104 only 13% of the lifetime is spent in the expansion period. a crowded traffic system results in a short and really intensive spread of new information and it is followed by a long disappearing section with the presence of a few vehicles having old (maybe no longer valid) information. copyright c© 2020 assa. adv syst sci appl (2020) statistical properties of vanet-based information spreading 43 6. summary in this work, we presented an agent-based model of information spreading on vehicular ad hoc networks. the time-dependent network topology of agents was based on the motion of communicating smart vehicles. vehicles are moving based on shortest travel time paths between the randomly selected starting and destination points of a real city map. due to the short-range communication moving vehicles can receive public information from each other or from a fixed road side unit. in this ad hoc network, the statistical properties of information spreading were investigated. above a certain number of smart vehicles in the system information spreads very fast, and a dominant part of the system can be in an almost homogeneous informed state. this efficient dissemination of information does not require considerable computational power of devices because there is no addressing or routing process and the devices store only the latest packages. however, we found that some agents can stay in an out-of-date state for a quite long time. the number of smart vehicles has a huge effect on spreading, only a large self-organized system can be effective. on micro scale, we analyzed the time intervals between information exchanges for all agents. we found that the distribution of these intervals has a power-law form where the density of vehicles has no effect on the exponents. the evolution of the spreading of each information holding any given timestamps was also analyzed. we found out that the lifetime of information can be separated into two periods: an expansion period and a dying out period. we found that the number of agents in these complex systems may affect the information spreading and the ratio of these two periods in significant and well describable ways. all these results claim appropriate treating of out-of-date vehicles. the presented results have shown that the system has many interesting features, although real-life applications require more realistic simulations. in our further research, we try to find answers to essential, practical questions. what happens if the rsu is removed (turned off) or more than one such units are placed to the system? how to avoid the presence of old (out-of-date, fake) information? what is the effect of the introduction of a susceptible-infected-susceptible (sis) model [23] (forgetting old information)? how to optimize spreading reducing the number of information exchanges (for energy efficiency), but keeping the system in an up-to-date phase? what is the topology of this ad hoc communication network? acknowledgements the publication is supported by the efop-3.6.1-16-2016-00022 project. the project is co-financed by the european union and the european social fund. this work was supported by the construction efop-3.6.3-vekop-16-2017-00002. the project was supported by the european union, co-financed by the european social fund. map data copyrighted openstreetmap contributors and available from https://www.openstreetmap.org. references 1. porta, s., crucitti, p., latora, v. (2006). the network analysis of urban streets: a dual approach, physica a, 369, pp. 853–866. 2. jiang, b. (2007). a topological pattern of urban street networks: universality and peculiarity, physica a, 384, pp. 647–655. 3. yan, y., zhang, s., tang, j., wang x. (2017). understanding characteristics in multivariate traffic flow time series from complex network structure, physica a 477, pp. 149–160. copyright c© 2020 assa. adv syst sci appl (2020) 44 i. varga, g. kocsis 4. zeadally, s., hunt, r., chen, y. s., irwin, a., hassan, a. (2012). vehicular ad hoc networks (vanets): status, results, and challenges, telecommunication systems, 50, pp. 217–241. 5. fiore, m., härri, j., filali, f., bonnet, c. (2007). vehicular mobility simulation for vanets, 40th annual simulation symposium (anss ’07), norfolk, pp. 301–309. 6. bátfai, n., besenczi, r., mamenyák, a., ispány, m. (2015). traffic simulation based on the robocar world championship initiative, infocommunications journal, 7, pp. 50–58. 7. salvo, p., felice, de m., baiocchi, a., cuomo, f., rubin, i. (2013). timer-based distributed dissemination protocols for vanets and their interaction with mac layer, ieee 77th vehicular technology conference, dresden, pp. 1–6. 8. malla, a. m., sahu, r. k. (2013). a review on vehicle to vehicle communication protocols in vanets, international journal of advanced research in computer science and software engineering, 3, pp. 409–414. 9. xu, q., sengupta, r., mak, t., ko, j. (2004). vehicle-to-vehicle safety messaging in dsrc, proceedings of the 1st acm international workshop on vehicular ad hoc networks, pp. 19–28. 10. nishtha, d. m. (2016). vehicular ad hoc networks (vanet), international journal of advanced research in electronics and communication engineering, 5, pp. 1003–1008. 11. gong, j., xu, c. z., holle, j. (2007). predictive directional greedy routing in vehicular ad hoc networks, 27th international conference on distributed computing systems workshops (icdcsw ’07), toronto, pp. 2–2. 12. ramakrishna, m. (2012). dbr: distance based routing protocol for vanets, international journal of information and electronics engineering, 2, pp. 228–232. 13. sanguesa, j. a., fogue, m., garrido, p., martinez, f. j., cano, j. c., calafate, c. t. (2016). a survey and comparative study of broadcast warning message dissemination schemes for vanets, mobile information systems, 8714142, pp. 1–18. 14. busetta, p., bouquet, p., adami, g., bonifacio, m., palmieri, f. (2003). k-trek: a peerto-peer approach to distribute knowledge in large environments, international workshop on agents and p2p computing, pp. 174–185. 15. bicocchi, n., mamei, m., zambonelli, f. (2012). self-organizing virtual macro sensors, acm transactions on autonomous and adaptive systems, 7, pp. 1–28. 16. varga, i. (2017). comparison of network topologies by simulation of advertising, in: gusikhin, o., méndez muñoz v., firouzi, f., mønster, d., chang, c. (eds.) proceedings of the 2nd international conference on complexity, future information systems and risk 2017, porto, pp. 17–22. 17. kocsis, g., varga, i. (2014). agent based simulation of spreading in social-systems of temporarily active actors, in: was, j., sirakoulis, g. ch., bandini, s. (eds.) cellular automata for research and industry 2014. lncs vol. 8751, springer, pp. 330–338. 18. openstreetmap (2018). [online]. available https://www.openstreetmap.org. 19. varga, i., némethy, a., kocsis, g. (2018). agent-based simulation of information spreading in vanet, in: mauri g., el yacoubi s., dennunzio a., nishinari k., manzoni l. (eds.) cellular automata for research and industry 2018, lncs vol. 11115, springer, pp. 166–174. 20. varga, i. (2020). a complex sis spreading model in ad hoc networks with reduced communication efforts, advances in complex systems (accepted). 21. stephan winter (2002). modeling costs of turns in route planning, geoinformatica, 6 (4), pp. 345-361. 22. google maps (2018). [online]. available https://maps.google.com. 23. m. e. j. newman, (2010). networks an introduction, oxford university press, pp. 627–675. copyright c© 2020 assa. adv syst sci appl (2020) introduction underlying map topology of simulations motion of vehicles spreading of information results summary adv syst sci appl 2021; 04:65–86 published online at https://ijassa.ipu.ru. dynamics of covid-19 outbreak and optimal control strategies: a model-based analysis naba kumar goswami1*, b. shanmukha2 1department of mathematics, pet research foundation, mysuru, india 2department of mathematics, pes college of engineering, mandya, india abstract: covid-19 is an infectious disease caused by the sars-cov-2 virus, which spreads so fast in the inhabitants. the virus is transmitted through direct contact with respiratory droplets of an infected individuals through coughing and sneezing or indirect contact through contaminated objects or surface. in this article, a non-linear mathematical model is proposed and analyzed to manifest the impact of transmission dynamics of the covid-19 pandemic based on indian condition by considering asymptomatic and symptomatic infections. it is assumed that the transmission rates due to asymptomatic and symptomatic individuals are different. the basic reproduction number of the model is computed and studied the stability of different equilibria of the model in detail. the sensitivity analysis is presented to identify the key parameters that influence the basic reproduction number, which can be regulated to control the transmission dynamics of the disease. also, this model is extended to the optimal control model and is analyzed by using the pontryagin’s maximum principal and solved numerically. it has been observed that the optimal control model gives better result as compacted to the model without optimal control model as it reduces the number of infectives significantly in a desired interval of time. keywords: covid-19, basic reproduction number, stability analysis, sensitivity analysis, optimal control 1. introduction the ongoing covid-19 outbreak has put mathematical models in the limelight. in 1960, the human coronavirus was first identified and in 21st century, a large number of people in this world have been affected by the three outbreaks such as sars, mers, and 2019-ncov. in 2003, the severe acute respiratory syndrome(sars) outbreaks, especially in the chinese mainland, hong kong, taiwan, and canada of the world [1] and in 2012 and 2015, the middle east respiratory syndrome(mers) outbreak in saudi arabin [2] and south korea [3] respectively. coronaviruses belonging to the family of coronaviridae and order of nidovirales, enveloped, non-segmented, single-stranded positive-sense rna viruses [4]. all coronavirus are zoonotic. they start in animals and can then, following mutation, recombination, and adaptation, be passed on to humans. in the human-to-human transmission of covid-19 can occurs via respiratory droplets directly (through droplets from coughing or sneezing) or indirectly (touching surfaces or objects contaminated with virus and touching their mouth, nose, eyes) [5]. the incubation period for the new coronavirus from 2 to 14 days in human to human transmission [6]. the common symptoms of covid-19 are fever, cough, difficulty of breathing, and fatigue. at present, there are no specific drugs for the disease to protect the people, only hygiene measures can reduce the ∗corresponding author: nabakrgoswami@gmail.com 66 n.k. goswami, b. shanmukha rate of transmission. according cdc covid-19 vaccination is a safer way to help build protection measure. covid vaccines can stop or reduce most of the people from getting sick, but not everyone. despite taking all recommended doses of vaccines and waits a few weeks for immunity to build up, still there is a chance that people can get infected. so far there is no evidence, however, that human coronaviruses can be transmitted by animals. in december 2019, a new outbreak of pneumonia of unknown cause has been identified in wuhan city, the capital of china’s hubei province [9]. this outbreak has some other potential causes like influenza, avian influenza, adenovirus, and sars, but there are no symptoms like mers [7]. later on 7th january, 2020, the causative pathogen was identified as a novel coronavirus (2019-ncov) [8]. as the virus is very closely related to sars and mers, so the name 2019-ncov can distinguish the virus from both [9]. on 30th january 2020, the world health organization (who) declared that the outbreak is a public health emergency of international concern(pheic) [10] as it is worldwide spread. as the number infective is increasing rapidly and the epidemiological evidence of human-to-human (doorplates, direct contact, etc.) transmission suggests that 2019-ncov is more contagious than both sars and mers [11–13]. on 11th february 2020, the world health organization announced a new name of coronavirus disease as covid-19 [14]. on 11th march 2020, the world health organization officially declares the covid-19 outbreak as a pandemic as it spread worldwide. the most affected countries in this world are the usa, india, brazil, italy, spain, france, uk, germany, iran, china, belgium, netherlands, etc. in this ongoing outbreak, in the usa more than 43,107,628 people are infected and more than 694,619 people have died. presently more than 229,382,253 infected cases have been reported across 223 territories, where more than 206,052,168 infected cases are recovered and more than 4,707,336 infected cases have been died till 20th sept 2021. in india, the first novel coronavirus case was reported on 30th january 2020 [15] in the state of kerala. due to the crises of coronavirus pandemic, prime minister of india declared 21 + 19 = 40 days nationwide lockdown from 25th of march to 3rd of may 2020 as a preventive measure for the covid-19 [16]. the ministry of health and family welfare of india has suggested various precautionary measures to prevent the spread of viruses such as washing hands frequently, physical/social distancing, wearing mask, avoiding touching face, nose, and eyes [17, 18]. apart from lockdown the government of india performing many awareness programs about preventative measures through media and social networks (tv, radio, newspaper, facebook, twitter, etc.). the effect of lockdown and social distancing play an important role to reduce the coronavirus infection. in 2020, total number of infected cases are recorded in india are 10,266,674 and january to september 2021, number infected cases are 23,211,745. in 2021, also statewide or locality wise lockdown declared according to basis of active cases in that particular state or region. on 20th september, 2021, more than 33,478,419 infected covid-19 cases have been reported in 32 states in india, while 32,715,105 cases are recovered and 445,165 many cases have been died. most affected states in india are maharashtra, karnataka, kerla, delhi, gujarat, rajasthan, madhya pradesh, uttar pradesh, tamil nadu, andhra pradesh, telangana, west bengal, jammu and kashmir. among the cities in india, mumbai, bangalore and delhi are the badly effected by covid-19 outbreak. the primary aimed to study this model to study indian conditions and will forecast future pandemic by using information available. in [1] authors presented a deterministic model and simplified from the seijr model, which is adapted to analyze the important parameters of the model of sars epidemic; in [19] authors constructed a seqijr model of epidemic disease transmission which includes immunization and varying population size is studied; in [20] authors studied a mathematical about early transmission dynamics of the infection and evaluating the effectiveness of control measures; in [21] authours proposed and studied a copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 67 model on the various impact of the intervention on the spread of covid-19 in india; in [22] authors discussed a data-driven analysis in the early phase of the outbreak. they estimated the basic reproduction number of novel coronavirus (2019-ncov) in china; in [23] authors develop a mathematical model for the spread of the coronavirus disease 2019 (covid-19); in [24] authors developed a bats-hosts-reservoir-people transmission network model for simulating the potential transmission from the infection source to the human infection. in [31] authors proposed a mathematical model and they focuses on the impacts of face mask, hospitalization of symptomatic individuals and quarantine of asymptomatic individuals on the transmission dynamics of covid-19 pandemic in india. this paper is organized as follows: in section 2 presents the model; in section 3 discussed existence of equilibria and basic reproduction number; in section 4 presents the stability analysis of the model; in section 5 deal with sensitivity analysis of basic reproduction number; in section 6 illustrate the effects of parameters on disease outbreak ; in section 7 we extend the model to optimal control model and analysis it. demonstrates the numerical simulation results of the optimal control model; finally, section 7 we conclude the paper. 2. the model in this section, a dynamic model for covid-19 pandemic is presented and discussed based on india condition. the model divides the total population n(t) = s + e + ia + is +h +r into six different compartments according to the nature of the disease such as susceptible individuals (s), exposed individuals (e), asymptomatic infective individuals (ia), symptomatic infective individuals (is), hospitalized individuals (h) and recovered individuals (r). it is assumed that the total population is varying and homogeneously mixed i.e., all people are equally likely to be infected by the infectious individuals if they come into contact. the individuals are employed in the province at a constant rate λ and join the susceptible class. the natural birth and deaths in the population are also considered in the model. it is assumed that susceptible individuals after being exposed to the covid19 infection can progress to asymptotic infective and symptomatic infective at the rates β. assume that an asymptomatic individual joins the symptomatic populations class at the rate ρ. further, both asymptomatic and symptomatic infectious individuals will progress to the hospitalized or quarantine compartment with clinical symptoms of covid-19 at the rate δ1 and δ2 respectively. also, some asymptotic and symptomatic individuals may recover without hospitalized or quarantine at the rates γ1 and γ2 respectively. hospitalized individuals may recover and after recovery, it progresses to recovered class at the rate γ3. however, the rates of recovery may vary from one compartment to another. the natural mortality rate of each individuals class is µ. due to critical illness of covid-19 disease some of the symptomatic and hospitalized individuals class have additional mortality rates µ1 and µ2, respectively. the flow diagram of the model is given in figure 1, and biological interpretations of parameters are shown in table 1. keeping the above facts/assumptions in mind, a mathematical model covid-19 is proposed as follows: ds dt = λ− βasia − βssis − µs de dt = βasia + βssis − (κ+ µ)e dia dt = ξκe − (ρ+ γ1 + µ+ δ1)ia dis dt = (1− ξ)κe + ρia − (γ2 + µ1 + µ+ δ2)is (1) copyright © 2021 assa. adv syst sci appl (2021) 68 n.k. goswami, b. shanmukha dh dt = δ1ia + δ2is − (γ3 + µ2 + µ)h dr dt = γ1ia + γ2is + γ3h − µr fig. 1. flow diagram of the model. table 1. biological interpretations of parameters parameter biological interpretations λ : rate of recruitment in the susceptible class, βa : rate of infection of susceptible with asymptomatic individuals βs : rate of infection of susceptible with symptomatic individuals κ : rate of incubation ρ : rate of progression from asymptomatic to symptomatic ξ : fraction of exposed individuals not showing symptoms δ1 : hospitalized/quarantine rate of asymptomatic individuals δ2 : rate of hospitalized symptomatic individuals γ1 : recovery rate of asymptomatic individuals γ2 : recovery rate of symptomatic individuals γ3 : recovery rate of hospitalized individuals µ : natural mortality rate of human µ1 : disease related mortality rate for symptomatic individuals µ2 : disease related mortality rate for hospitalized individuals copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 69 3. analysis of the model system from the system (1), we have ds dt ∣∣∣ s=0 = λ > 0, de dt ∣∣∣ e=0 = βasia + βssis ≥ 0, dia dt ∣∣∣ ia=0 = ξκe ≥ 0 dis dt ∣∣∣ is=0 = (1− ξ)κe + ρia ≥ 0, dh dt ∣∣∣ h=0 = δ1ia + δ2is ≥ 0, dr dt ∣∣∣ r=0 = γ1ia + γ2is + γ3h ≥ 0 here, all the rates are non-negative on the bounding planes. so, if we start in the interior of the 6-dimensional closed hyperoctant r6 +, we will always remain there, in view of the fact that the direction of the vector field is inward on all the bounding planes. thus, non-negativity of all the solutions of the model system (1) is guaranteed. further, from the model system (1), we note that the total human population n = s1 + s2 + i +h +q+r satisfies, dn dt = λ− µn − µ1is − µ2h this gives lim sup t→∞ n ≤ λ µ therefore, all the solutions s(t), e(t), ia(t), is(t), h(t), r(t) are bounded by λ µ . hence, the biologically feasible region for the system (1) is given by the following positively invariant set: ω = {(s,e, ia, is, h,r) ∈ r6 + : 0 ≤ s + e + ia + is +h +r ≤ λ µ } 3.1. basic reproduction number the disease-free equilibrium for the model (1) as e0 = (s0, e0, i0 a , i 0 s , h 0, r0) = ( λ µ , 0, 0, 0, 0, 0). we find the basic reproduction number r0 by using the next generation matrix method [26]. the new infection terms of the matrix f and the transition terms of the matrix v of the system (1) are respectively, as follows: f =  βasia + βssis 0 0 0  , v =  (κ+ µ)e −ξκe + (ρ+ γ1 + δ1 + µ)ia −(1− ξ)κe − ρia + (γ2 + δ2 + µ1 + µ)is −δ1ia − δ2is + (γ3 + µ2 + µ)h  now, we find the matrices f (of new infection terms) and v (of the transition terms) as copyright © 2021 assa. adv syst sci appl (2021) 70 n.k. goswami, b. shanmukha f =  0 βas βss 0 0 0 0 0 0 0 0 0 0 0 0 0  , v =  κ+ µ 0 0 0 −ξκ ρ+ γ1 + δ1 + µ 0 0 −(1− ξ)κ −ρ γ2 + δ2 + µ1 + µ 0 0 −δ1 −δ2 γ3 + µ2 + µ  it follows that fv −1 =  m11 m12 m13 0 0 0 0 0 0 0 0 0 0 0 0 0  where, m11 = κξβaλ µ(κ+ µ)(ρ+ γ1 + δ1 + µ) + βsλκ(ξρ+ (1− ξ)(ρ+ γ1 + δ1 + µ)) µ(κ+ µ)(ρ+ γ1 + δ1 + µ)(γ2 + δ2 + µ1 + µ) m12 = βaλ µ(ρ+ γ1 + δ1 + µ) − ρβsλ µ(ρ+ γ1 + δ1 + µ)(γ2 + δ2 + µ1 + µ) , m13 = βsλ µ(γ2 + δ2 + µ1 + µ) the basic reproduction number is same as the spectral radius of the next-generation matrix fv −1. thus, from above, we obtain the expression for r0 as r0 = κλ µ(κ+ µ)(ρ+ γ1 + δa + µ) [ ξβa + βs{ξρ+ (1− ξ)(ρ+ γ1 + δ1 + µ)} γ2 + δ2 + µ1 + µ ] the quantity r0 is known as basic reproduction number, the expected number of secondary cases produced in completely susceptible population, by a typical infective individual for the system (1). 3.2. existence of endemic equilibrium point the endemic equilibrium of the system (1) satisfies the following algebraic equations such as ds dt = 0, de dt = 0, dia dt = 0, dis dt = 0, dh dt = 0, dr dt = 0 the system (1) realize a unique positive solution e1 = (s∗, e∗, i∗a , i ∗ s , h ∗, r∗) s∗ = κ+ µ (βad1 + βsd2) , e∗ = λ(βad1 + βsd2)− µ(κ+ µ) (βad1 + βsd2)(κ+ µ) , i∗a = ξκe∗ ρ+ γ1 + δ1 + µ = d1e ∗, i∗s = (1− ξ)κe∗ + ρd1e ∗ γ2 + δ2 + µ1 + µ = d2e ∗, h∗ = (δ1d1 + δ2d2)e∗ γ3 + µ2 + µ = d3e ∗, r∗ = (γ1d1 + γ2d2 + γ3d3)e∗ µ where, d1 = κξ ρ+ γ1 + δ1 + µ , d2 = (1− ξ)κ+ ρd1 γ2 + δ2 + µ1 + µ , d3 = (δ1d1 + δ2d2) γ3 + µ2 + µ copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 71 4. stability analysis of the model theorem 4.1: for model system (1), the disease-free equilibrium e0 is locally asymptotically stable if r0 < 1 and unstable if r0 > 1. 4.1. global stability of disease-free equilibrium to prove the global stability of disease-free equilibrium, we are using the theorem by castillochavez et al. [25] theorem 4.2: if the given mathematical model can be written in the form: dx dt = f (x, y ), and dy dt = g(x, y ), g(x, 0) = 0 (2) where x = s , y = (e, ia, is, h)t , denoting the number of uninfected and denoting the number of covid-19 infected people respectively. then the disease-free equilibrium is represented here by e0 = (x0, 0) = ( λ µ , 0) for the global asymptotically stable, the condition (h1) and (h2) given below must be satisfied. h1 : for dx dt = f (x0, 0), h2 : g(x, y ) = ay − ĝ(x, y ), ĝ(x, y ) ≥ 0, here a = dyg(x0, 0) is m-matrix (in m-matrix, all the off diagonal element of matrix are non-negative) . if the given system of differential equation in mathematical model satisfies the given condition in (2) then the point e0 = (x0, 0) is a global asymptotically stable equilibrium of given mathematical model provided r0 < 1. and for the given mathematical model, the result is shown in the next theorem, as given below. theorem 4.3: the point e0 = (x0, 0) of the system (1) is global asymptotically stable (g.a.s.), provided r0 < 1. and the condition given in (2) are satisfied. proof by using theorem (2.1) to our model system (1), we get f (x0, 0) = λ− µs, g(x, y ) = ay − ĝ(x, y ) where, a = f − v =  −(κ+ µ) βas βss 0 ξκ −(ρ+ γ1 + δ1 + µ) 0 0 (1− ξ)κ ρ −(γ2 + δ2 + µ1 + µ) 0 0 δ1 δ2 −(γ3 + µ2 + µ)  copyright © 2021 assa. adv syst sci appl (2021) 72 n.k. goswami, b. shanmukha then g(x, y ) = ay − ĝ(x, y ) = ay −  ĝ1(x, y ) ĝ2(x, y ) ĝ3(x, y ) ĝ4(x, y )  = (s0 − s)(βaia + βsis) 0 0 0  where ay =  −(κ+ µ)e + s0(βaia + βsis) ξκe − (ρ+ γ1 + δ1 + µ)ia (1− ξ)κe + ρia + (γ2 + δ2 + µ1 + µ)is δ1ia + δ2is + (γ3 + µ2 + µ)h  and ĝ(x, y ) =  (κ+ µ)e + s(βaia + βsis) ξκe − (ρ+ γ1 + δ1 + µ)ia (1− ξ)κe + ρia + (γ2 + δ2 + µ1 + µ)is δ1ia + δ2is + (γ3 + µ2 + µ)h  here, we can easily see s0 ≥ s, hence g(x, y ) ≥ 0 for all (x, y ). also by the defination of m matrix we can say that the matrix a is m matrix. hence, disease-free equilibrium (e0) is global asymptotically stable. 4.2. global stability of endemic equilibrium theorem 4.4: the endemic equilibrium e1 = (s∗, e∗, i∗a , i ∗ s , h ∗, r∗) of the given mathematical model is globally asymptotically stable. proof for the global stability result, we will use the method discussed in korobeinikov and wake [27], li and muldowney [28, 29]. here we consider the following lyapunov function: l = c1 ( s − s∗ − s∗ln s s∗ ) + c2 ( e − e∗ − e∗ln e e∗ ) + c3 ( ia − i∗a − i∗a ln ia i∗a ) +c4 ( is − i∗s − i∗s ln is i∗s ) then the time derivative of l is given by dl dt = c1 ( 1− s∗ s ) ds dt + c2 ( 1− e∗ e ) de dt + c3 ( 1− i∗1 i1 ) di1 dt + c4 ( 1− i∗2 i2 ) di2 dt now from the mathematical model we put the expressions for ds dt , de dt , dia dt , dis dt , in the above equation, which gives dl dt = c1 ( 1− s∗ s ) {λ− βaias − βsiss − µs} + c2 ( 1− e∗ e ) {βaias + βsiss − (κ+ µ)e} + c3 ( 1− i∗a ia ) {ξκe − ρia − (γ1 + µ)ia − δ1ia} copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 73 + c4 ( 1− i∗s is ) {(1− ξ)κe + ρia − (γ2 + µ1 + µ)is − δ2is} (3) the mathematical model system satisfies the following relation at the equilibrium point. λ = βai ∗ as ∗ + βsi ∗ ss ∗ + µs∗, κ+ µ = βai ∗ as ∗ + βsi ∗ ss ∗ e∗ , ρ+ γ1 + δ1 + µ = ξκe∗ i∗a , γ2 + δ2 + µ1 + µ = (1− ξ)κe∗ + ρi∗a i∗s putting all the above expressions in (3) we get, dl dt = c1 ( 1− s∗ s ) [βai ∗ as ∗ + βsi ∗ ss ∗ + µs∗ − βaias − βsiss − µs] + c2 ( 1− e∗ e )[ βaias + βsiss − ( βai ∗ as ∗ + βsi ∗ ss ∗ e∗ ) e ] + c3 ( 1− i∗a ia )[ ξκe − ( ξκe∗ i∗a ) ia ] + c4 ( 1− i∗s is )[ (1− ξ)κe + ρia − ( (1− ξ)κe∗ + ρi∗a i∗s ) is ] then, dl dt = −c1 ( (s∗ − s)2 s ) µ+ c1 [( 1− s∗ s ) {βai∗as∗ + βsi ∗ ss ∗ − β1ias − β2iss} ] + c2 ( 1− e∗ e )[ βaias + βsiss − ( βai ∗ as ∗ + βsi ∗ ss ∗ e∗ ) e ] + c3 ( 1− i∗a ia )[ ξκe − ( ξκe∗ i∗a ) ia ] + c4 ( 1− i∗s is )[ (1− ξ)κe + ρia − ( (1− ξ)κe∗ + ρi∗a i∗s ) is ] dl dt = −c1 ( (s∗ − s)2 s ) µ+ g(x1, x2, x3, x4) where, s s∗ = x1, e e∗ = x2, ia i∗a = x3, is i∗s = x4, βas ∗i∗a = p, βss ∗i∗s = q, ξκe∗ = r, (1− ξ)κe∗ = s now, g(x1, x2, x3, x4) = c1(p+ q − px1x3 − qx1x4)− c1p 1 x1 − c1q 1 x1 + c1px3 + c1qx4 + c2(px1x2 + qx1x4 − ax2 − qx2)− c2p ( x1x3 x2 ) − c2q ( x1x4 x2 ) copyright © 2021 assa. adv syst sci appl (2021) 74 n.k. goswami, b. shanmukha + c2(p+ q) + c3(rx2 − rx3 − r x2 x3 + r) + c4(sx2 − sx4 − s x2 x4 + s) = x2(−c2p+ c3r + c4s) + x3(c1p− c3r) + x4(c1q − c4s) + x1x3(−c1p+ c2p) + x1x4(c1q + c2q) + c1(2p+ q) + c2(p+ q) + c3r − c3r ( x2 x3 ) + c4s− c4s ( x2 x4 ) − c1(p+ q) ( 1 x1 ) − c2(px1x3 + qx1x4) ( 1 x2 ) to get the values of c1, c2, c3, c4 we take the coefficients of x1x3, x1x4, x2, x3, x4 equal to zero and solve the algebraic equations in c1, c2, c3, c4. this gives c1 = c2; c3 = c1p r ; c4 = c1q s choosing c1 = c2 = 1 and n = 1, we get g(x1, x2, x3, x4) = p ( 3− 1 x1 − x1x3 x2 − x2 x3 ) + q ( 3− 1 x1 − x1x4 x2 − x2 x4 ) since the arithmetic mean is greater than or equal to geometric mean, we have 1 x1 + x1x3 x2 + x2 x3 ≥ 3 and 1 x1 + x1x4 x2 + x2 x4 ≥ 3 hence, dl dt = − ( (s∗ − s)2 s ) µ+ p ( 3− 1 x1 − x1x3 x2 − x2 x3 ) + q ( 3− 1 x1 − x1x4 x2 − x2 x4 ) thus it is easy to observe that dl dt ≤ 0 and the equality dl dt = 0 hold on;y for x1 = x2 = x3 = x4 = 1 for which s = s∗, e = e∗, ia = i∗a , is = i∗s . from the lasalle’s invariance principle [30], the equilibrium e1 of the given system is globally asymptotically stable for r0 > 1. 5. sensitivity analysis of reproduction number in this section, we also perform sensitivity analysis for the parameters involved in reproduction number r0, which reflects that increase or decrease in these parameter causes increase or decrease in r0. the sensitivity of r0 to different parameters is shown in figure 3. it is used to discover the parameters that have a high impact on r0 and should be targeted by intervention strategies. sensitivity indices allows to measure the relative change in a variable when parameter changes. for that we use the forward sensitivity index of a variable, with copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 75 respect to a given parameter, which is defined as the ratio of the relative change in the variable to the relative change in the parameter. if such variable is differentiable with respect to the parameter, then the sensitivity index is defined using partial derivatives. the normalized forward sensitivity index of r0, which is differentiable with respect to a given parameter α, is defined by y r0 α = α r0 ∂r0 ∂α the above formula can be used to compute the analytical expression for the sensitivity of r0 to each parameter that it includes. accordingly, the sensitivity indexes of the model (1) are illustrate in figure 2. consequently, the value of r0 increases with increase in the values of all positive indices parameters λ, βa, βs, κ, and ξ with r0. also, the parameters γ1, γ2, δ1, δ2, ρ, µ, and µ1 have negative index with r0. it is clearly observed that the effect of the parameter λ is the maximum and hence it is the most sensitive parameter of r0. it means that small change (increase or decrease) in the parameters (λ) will significant change in the value of r0 by 100%. it is obvious that phenomenon of a lower value of r0 will boost to prevent the disease prevalence. thus, to control the disease from the population, we have to control the increase of parameters having positive indices with r0, whereas parameters with negative indices should be maintained. 6. effects of parameters on disease outbreak for the numerical simulation of the model, we consider all the parameters are in per day basis. first we consider the following set of parameters which corresponds to disease-free equilibrium. λ = 2; βa = 0.00085; βs = 0.00029; γ1 = 1/14; γ2 = 0.00245; γ3 = 0.002;κ = 1.19; µ1 = 0.1;µ2 = 0.015;µ = 0.00425; ρ = 0.001; ξ = 0.0015; δ1 = 0.16; δ2 = 0.01; for the above set of parameters we get r0 = 0.0932 < 1 and the disease-free equilibrium point e0(462.52, 0, 0, 0, 0, 0) is stable. this fact is demonstrated in figure 4 (a). later, we did few changes for endemic equilibrium as follows: λ = 40; βa = 0.00085; βs = 0.00029; γ1 = 1/14; γ2 = 0.000245; γ3 = 0.002; κ = 0.89; µ1 = 1.05; µ2 = 0.015; µ = 0.0125; ρ = 0.69; ξ = 0.95; δ1 = 0.012; δ2 = 0.019; from the above set of parameter we get r0 = 2.0075 > 1, and the endemic equilibrium e1(801.56, 26.82, 25.04, 24, 43, 38.75, 208.6, ) is stable. the stability of the equilibrium point e1 is shown in figure 4(b). the effect of different values of recovery rate from asymptomatic (γ1), symptomatic (γ2) and hospitalized (γ3), which corresponds to infective human is demonstrated in figure 4(c), 4(d), and 4(e) respectively. it is clear that the parameter of recovery rate γ1, γ2, and γ3 increases, simultaneously the infected population decreases. the effect of s, h, and r also shown in the figure 4(f) with respect to endemic equilibrium point. copyright © 2021 assa. adv syst sci appl (2021) 76 n.k. goswami, b. shanmukha a s 1 2 1 1 2 parameters -1.5 -1 -0.5 0 0.5 1 f or w ar d s en si tiv ity o f r 0 fig. 2. forward sensitivity of r0 7. the optimal control model in this section, the model (1) is extended to formulate optimal control problem by incorporating two time-dependent optimal control parameters, namely u1(t), and u2(t). if u1, and u2 equal zero, no effort is placed in these controls at time t, and if they equal one, maximum effort is applied. thus, optimal control variables are given, as follows: the control variable u1(t) represents the reduction in the transmission between humanto-human via using surgical face masks, social distancing, self-isolation, sensitization and awareness of transmission of the disease. the control variable u2(t) represents the increase in the testing facility and treatment, which can lead to fast detection of covid infected cases and recovery and add additional time-dependent parameter ηu2(t) in the rate of direction δ2. keeping in view of the above assumptions, the optimal control model is formulated as follows: ds dt = λ− (1− u1)βaias − (1− u1)βsiss − µs de dt = (1− u1)βaias + (1− u1)βsiss − (κ+ µ)e dia dt = ξκe − (ρ+ γ1 + µ+ δ1)ia dis dt = (1− ξ)κe + ρia − (γ2 + µ1 + µ)is − (δ2 + ηu2)is (4) dh dt = δ1ia + (δ2 + ηu2)is − (γ3 + µ2 + µ)h dr dt = γ1ia + γ2is + γ3h − µr copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 77 0 1 1 1 r 0 0.8 s 0.5 0.6 a 2 0.4 0.20 0 0 0.5 1 1.5 (a) 0 1 1 1 r 0 0.80.5 0.6 a 2 0.4 0.20 0 0.5 1 1.5 (b) 0 1 1 1 r 0 0.80.5 0.6 s 2 0.4 0.20 0 0 2 4 6 8 10 12 (c) 0 1 1 1 r 0 0.80.5 0.6 a 2 0.4 0.20 0 0.5 1 1.5 (d) 0 1 1 1 r 0 0.80.5 0.6 2 0.4 0.20 0 0 0.5 1 1.5 (e) 0 1 1 1 r 0 0.80.5 0.6 3 2 0.4 0.20 0 0 0.5 1 1.5 2 (f) fig. 3. (a) influence of βa and βs on r0, (b) influence of βa and ρ on r0, (c) influence of βs and ξ on r0 (d) influence of βa and ξ on r0, (e) influence of λ and ξ on r0, (f) influence of γ3 and κ on r0 copyright © 2021 assa. adv syst sci appl (2021) 78 n.k. goswami, b. shanmukha 0 200 400 600 800 1000 time(in days) 0 200 400 600 800 1000 p op ul at io ns s e i a i s h r fig. 4. variation of s, e, ia, is, h , r performing the stability of disease-free equilibrium point with r0 = 0.0932 < 1. 0 200 400 600 800 1000 time(in days) 0 200 400 600 800 1000 1200 p op ul at io ns s e i a i s h r fig. 5. variation of s, e, ia, is, h , r performing the stability of endemic equilibrium point with r0 = 2.0075 > 1. 0 50 100 150 200 250 300 350 400 450 500 time(in days) 0 50 100 150 in fe ct ed p op ul at io ns 1 =0.526 1 =0.0526 1 =0.00526 fig. 6. variation of ia with time performing different values of γ1 copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 79 0 50 100 150 200 250 300 350 400 450 500 time(in days) 0 50 100 150 in fe ct ed p op ul at io ns 2 =1.5 2 =0.5 2 =.005 fig. 7. variation of is with time performing different values of γ2 0 50 100 150 200 250 300 350 400 450 500 time(in days) 0 50 100 150 in fe ct ed p op ul at io ns 3 =0.2 3 =0.02 3 =0.0002 fig. 8. variation of h with time performing different values of γ3 7.1. the optimal control problem in this section, we study the behavior of the proposed model by using optimal control theory. the objective functional for fixed time tf if given by j(u1, u2) = ∫ tf 0 [ a1ia + a2is + 1 2 (b1u 2 1 +b2u 2 2) ] dt (5) subject to the model system (4). the parameter a1 ≥ 0, a2 ≥ 0, b1 ≥ 0, b2 ≥ 0 are the weight and balancing constants, which measure the respective cost involvement over the interval [0, tf ]. in order to find an optimal control, u1 ∗, and u2 ∗ such that j(u1 ∗, u2 ∗) = min (u1,u2)∈ω j(u1, u2), (6) where ω is the control set and is defined as ω = {(u1, u2) : 0 ≤ u1, u2 ≤ 1, t ∈ [0, tf ]} here, all the controls are bounded and measurable. 7.1.1. existence and characterization of optimal controls here, we shall first establish the existence of such control functions that minimizes the cost functional j . the lagrangian l copyright © 2021 assa. adv syst sci appl (2021) 80 n.k. goswami, b. shanmukha of this problem is defined as: l(ia, is, u1, u2) = a1ia + a2is + 1 2 b1u1 2 + 1 2 b2u 2 2 now, we shall use pontryagin’s maximum principle [32, 33] for necessary conditions for optimal controls system (4). for that by choosing x = (s,e, ia, is, h,r) ,ω = (u1, u2) and λ = (λ1, λ2, λ3, λ4, λ5, λ6), the associated hamiltonianh can be written as h(x,ω, λ) = l(ia, is, u1, u2) + λ1 ds dt + λ2 de dt + λ3 dia dt + λ4 dis dt + λ5 dh dt + λ6 dr dt (7) since u∗1, and u∗2 are solutions to the control problem (4), there exists the adjoint variables λ1, λ2, λ3, λ4, λ5, λ6 satisfying the following conditions. dx dt = ∂h(t, x, u∗1, u ∗ 2, λ1, λ2, λ3.λ4, λ5, λ6) ∂λ 0 = ∂h(t, x, u∗1, u ∗ 2, λ1, λ2, λ3.λ4, λ5, λ6) ∂u dλ dt = −∂h(t, x, u∗1, u ∗ 2, λ1, λ2, λ3.λ4, λ5, λ6) ∂x (8) theorem 7.1: for the objective functional (5) and the control set (8) subject to control system (4) there exists an optimal control u∗ = (u1 ∗, u2 ∗) ∈ ω such that j(u1 ∗, u2 ∗) = min ω j(u1, u2). proof to establish this result, we follow the theorem 4.1 mentioned in [40] for the existence of optimal controls. as, we have discussed above that all the state variables (population) are bounded for each bounded controls coming from the control set ω. furthermore, lipschitz condition with respect to state variables is satisfied by the right hand part of the model system (4). the control variable set ω is also convex and closed by the definition and the model system (4) is linear in control variables. thus, all the conditions for the existence of controls are fulfilled (for more details one can follow [28, 29]). hence the result. theorem 7.2: for optimal controls measures u∗1, u ∗ 2 and the state solutions s∗, e∗, i∗a , i ∗ s , h ∗, r∗ of the state system (4), there exists adjoint variables λ = (λi) tf ∈ r6, i = 1, 2, 3, 4, 5, 6 such that dλ1 dt = (1− u1)βaia(λ1 − λ2) + (1− u1)βsis(λ1 − λ2) + µλ1 dλ2 dt = (κ+ µ)λ2 + ξκ(λ4 − λ3)− κλ4 dλ3 dt = −a1 + (1− u1)βas(λ1 − λ2) + ρ(λ3 − λ4) + γ1(λ3 − λ6) + (µ+ δ1)λ3 dλ4 dt = −a2 + (1− u1)βss(λ1 − λ2) + γ2(λ4 − λ6) + (µ1 + µ)λ4 + (δ2 + ηu2)(λ4 − λ5) dλ5 dt = γ3(λ5 − λ6) + (µ2 + µ)λ5 (9) dλ6 dt = µλ6 copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 81 with transversality conditions λ1(tf ) = λ2(tf ) = λ3(tf ) = λ4(tf ) = λ5(tf ) = λ6(tf ) = 0 (10) proof let u∗1, u ∗ 2 be the optimal control functions and s∗, e∗, i∗a , i ∗ s , h ∗, r∗ are the corresponding state variables. then, pontryagin’s maximum principle ensures the existence of the following adjoint variable λi(i = 1, 2, 3, 4, 5, 6) ∈ r6, which satisfies the following canonical equations: dλ1 dt = −∂h ∂s , dλ2 dt = −∂h ∂e , dλ3 dt = −∂h ∂ia , dλ4 dt = −∂h ∂is , dλ5 dt = −∂h ∂h , dλ6 dt = −∂h ∂r with transversality conditions (10) and the hamiltonian (7). the adjoint system (9) can be obtained. in the following result, we shall state the analytical forms of the optimal controls. theorem 7.3: the optimal controls (u∗1, u ∗ 2) which minimizes j over the region ω are given by u∗1 = min{1, max(0, ũ1)} u∗2 = min{1, max(0, ũ2)} where ũ1 = (β1ia + β2is)s(λ2 − λ1) b1 , ũ2 = ηis(λ4 − λ5) b2 , proof using optimally condition, we have ∂h ∂u1 = 0, ∂h ∂u2 = 0 we have ∂h ∂u1 = b1u1 + (βaia + βsis)s(λ1 − λ2) = 0 this gives u1 = ((βaia + βsis)s)(λ2 − λ1) b1 := ũ1 similarly, ∂h ∂u2 = b2u2 + ηis(λ5 − λ4) = 0 this implies u2 = ηis(λ4 − λ5) b2 := ũ2 moreover, lower and upper bounds of these control are 0 and 1 respectively. thus, if ũ1 > 1, ũ2 > 1, then u1 = u2 = 1. copyright © 2021 assa. adv syst sci appl (2021) 82 n.k. goswami, b. shanmukha also, ũ1 < 0, ũ2 < 0 then u1 = u2 = 0. otherwise, we have u1 = ũ1, and u2 = ũ2 hence, for these controls u∗1, u ∗ 2 we get optimum value of the function j . 8. simulation of optimal control problem in this section, we simulate our optimal control model using matlab. the parameter values are keeping same the parameters corresponding to stability of endemic equilibrium point e1 of the model (1). the weight constants for the optimal control problem are taken as a1 = 1, a2 = 1, b1 = 45, b2 = 65. we solve the optimality system by iterative method with the help of forward and backward difference approximations [32, 34–36]. we consider the time interval as [0,180]. first we solve the state equations by the forward difference approximation method then we use the backward difference approximation method to solve the adjoint equations. we consider different types of strategies to see the impact of optimal control in the total number of human infectives. 8.1. strategy a: employing hygiene promotion, social distancing and self-isolation (u1), only. here, only control measure u1(t) is used to optimize the objective function j , while control intervention u2(t) = 0, were not employed. the influence of u1(t) is demonstrated in figure 5(a), to minimize the objective function, the optimal control u1(t) is maintained at the maximum level. a single preventive measure can influence the spread of the covid-19 in the population. maintaining social distancing and self-isolation leads to control significant number reduction in asymptomatic and symptomatic cases in the population. from the figures, it is clear that the optimal control u1(t) is a little more effective compared to other types of controls but we need to maintain it to one for a longer period which is not easy to achieve. this control strategy is for using surgical face masks, social distancing, selfisolation, awareness of the transmission of disease, and sensitization. 8.2. strategy b: increase testing facility and treatment of the symptomatic individuals (u2) only. here, only control measure u2(t) is used to optimize the objective function j , while control intervention u1(t) = 0, were not employed. in figure 5(b), we present the plots of population and the effects of the increase in testing facility and treatment are demonstrated to minimizing the cost and reducing the number of coronavirus infections in the population. 8.3. strategy c: employing both the control interventions (u1, u2). here both the control interventions (u1(t), u2(t)) are used to optimize the objective function j. from figure 5(c), it is easy to say that by combining both optimal controls u1(t) and u2(t), the total number of infectious individuals decreases significantly. the simulation result indicates the effectiveness of optimal control strategies in reducing the number of infectives. it is observed that combined controls are more useful in reducing the number of infected cases significantly. finally, it observed that from figure 5(d), the optimal control model gives a better result as compacted to the model without the optimal control model as it reduces the number of infectives significantly in a desired interval of time. copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 83 0 20 40 60 80 100 120 140 160 180 time(in days) 0 0.2 0.4 0.6 0.8 1 u 1(c o n tr o l p ro fi le ) u 1 , when u 2 = 0 fig. 9. influence of u1, when u2 = 0 0 20 40 60 80 100 120 140 160 180 time(in days) 0 0.2 0.4 0.6 0.8 1 u 2(c o n tr o l p ro fi le ) u 2 , when u 1 = 0 fig. 10. influence of u2, when u1 = 0 0 20 40 60 80 100 120 140 160 180 time(in days) 0 0.2 0.4 0.6 0.8 1 u 1(c o n tr o l p ro fi le ) u 1 u 2 fig. 11. influence of u1, u2 0 20 40 60 80 100 120 140 160 180 time(in days) 0 50 100 150 200 in fe ct ed p op ul at io ns without control with control fig. 12. variation of infected human populations against time with and without optimal control copyright © 2021 assa. adv syst sci appl (2021) 84 n.k. goswami, b. shanmukha 9. conclusion dynamics of covid-19 pandemic and its optimal control strategies are discussed in this present study through a mathematical model. the host population of the model is subdividing into six different compartment according to nature of the disease such as susceptible, exposed, asymptomatic infective,symptomatic infective, hospitalized and recovery. this non-linear mathematical model is proposed based on indian condition by considering asymptomatic and symptomatic infections populations. it is assumed that the transmission rates due to asymptomatic and symptomatic individuals are different. the epidemiological threshold (basic reproduction number) is computed by using next generation matrix method and discussed different equilibria of the model in details. the model is globally asymptotically stable for the global stability of disease-free equilibrium when basic reproduction number is less than unity using the castillo-chavez theorem. also the theoretical analysis is carried out the global stability of endemic equilibrium is globally asymptotically stable using lyapunov method when basic reproduction number is greater than one. sensitivity analysis of the model performed to identify the key parameter that influence the basic reproduction number, which will regulate to control transmission dynamics of the disease. it emphasized that λ and µ are the most sensitive parameter, followed by asymptomatic transmission coefficient β1 and ξ. furthermore, the mathematical model is extended to optimal control problem by incorporating two time-dependent optimal control parameters to reducing the burden due to covid-19 using pontryagin’s maximum principal. introduced control parameter u1 in the model as surgical face masks, social distancing, self-isolation, sensitization and awareness of transmission of the disease in asymptomatic and symptomatic invective individuals, which effectively reduce transmission rate. social distancing should always implement at a higher percentage than self-isolation. the optimal control drives a significant reduction in the asymptomatic and symptomatic infected populations. the control parameter u2 included reducing the cost infrastructure like testing facility, which can lead to fast detection of covid infected cases. the optimal control suggests that unless the cost is very high, the social distancing should implement at the maximum level throughout the time examined. the optimal control model provides a more reliable result as compacted to the model without the optimal control model. this control strategy reduces the number of infectives significantly in a desired interval of time. acknowledgements the authors would like to thank the editor and anonymous referees for their valuable comments and suggestions which led to an improvement of our original manuscript. references 1. guanghong, d.,chang,l.,jianqiu,g.,ling, w., cheng, k., and zhang, d. (2004) sars epidemical forecast research in mathematical model,chinese science bulletin, 49 (21), 2332—2338. 2. killerby,m., h. biggs, h., midgleys, c., gerber, s., watson, j. (2020) middle esat respiratory syndrome coronavirus transmission, emerg.infect.dis. 26 (4),191-198. 3. willman, m., kobasa, d., kindrachuk, j. (2019) a comparative analysis of factors influencing two outbreaks of middle eastern respiratory syndrome (mers) in saudi arabia and south korea, viruses 11(12):1119.doi :10.3390/v11121119.pmid 31817037. 4. kolifarhood, g., aghaali, m., saadati, h. m., taherpour, n., rahimi, s., izadi, n.,and nazari, s.s.h. (2020) epidemiological and clinical aspects of covid-19: a narative review,archives of academic emergency medicine, 8(1). 5. chan, j.f.w. ,yuan, s., kok , k.h., chu,h., yang, j. (2019) afamilial cluster of pneumonia associated with the 2019 novel coronavirus indicating person-to-person copyright © 2021 assa. adv syst sci appl (2021) dynamics of covid-19 outbreak and optimal control strategies: a model-based... 85 transmission: a study of a family cluster. lancet. http://dx.doi.org/10.1016/s01406736(20) 30154-9. 6. bai, y., yao, l., wei, t., tian, f., jin, d. y., chen, l. (2020) presumed asymptomatic carrier transmission of covid-19. jama. 2020 feb 21. pubmed pmid: 32083643. pubmed central, pmcid: pmc7042844. epub 2020/02/23. eng. 7. world health organization. novel coronavirus – china. geneva,switzerland: world health organization.https://www.who.int/csr/https://www.who.int/csr/don/12-january2020-novel-coronavirus-china/en/ 8. hui, d. s., azhar, e. i. , madani, t. a., ntoumi, f., kock, r., dar , o. (2020) the continuing 2019− ncov epidemic threat of novel coronaviruses to global health the latest 2019 novel coronavirus outbreak in wuhan, china. int. j. infect. dis. 9. n. zhu, d. zhang, w. wang,x. w. li, b. yang, j. d. song , et al. a novel coronavirus from patients with pneumonia in china, 2019, n engl jmed.http: //dx.doi.org/10.1056/nejmoa2001017. [2020-01-24]. 10. world health organization. statement on the second meeting of the international health regulations (2005) emergency committee regarding the outbreak of novel coronavirus (2019-ncov). geneva, switzerland: world health organization.https://www.who.int/newsroom/detail/30-01-2020-statement-on-thesecond-meeting-of-theinternational-health-regulations-(2005)-emergency-committeeregar ding-the-outbreak-of-novel-coronavirus-(2019-ncov). [2020-01-30]. 11. wang, c., hornby, p. w., hayden, f. w., gao, g. f. (2020) a novel coronavirus outbreak of global health concern. lancet. http://dx.doi.org/10. 1016/s0140-6736(20)30185-9.. 12. munster, v.j., koopmans, m., doremalen , n. v., riel, d. v., wit, e. d. (2020) a novel coronavirus emerging in china – key questions for impact assessment. n engl j med. http://dx.doi.org/10.1056/nejmp2000929. 13. huang, c., wang, y., li, x., ren, l., zhao, j., hu, y. (2020) clinical features of patients infected with 2019 novel coronavirus in wuhan, china. lancet. http://dx.doi.org/10.1016/s0140-6736(20)30183-5. 14. who, (2020) naming the coronavirus disease (covid-19) and the virus that causes it. available online:,https://www.who.int/emergencies/diseases/novel-coronavirus2019/technical-guidance/[retrieved: 25/03/2020]. 15. who, (2020) coronavirus disease 2019 (covid-19), situation report -10. available online:, https://www.who.int/emergencies/diseases/novel-coronavirus-2019/situationreports/[retrieved: 25/03/2020]. 16. pulla, p. (2020) covid-19: india imposes lockdown for 21 days and cases rise, bmj 368, doi: 10.1136/bmj.m1251. 17. ministry of health and family welfare (mohfw), (2020) coronavirus disease 2019 (covid-19). available online:, https://www.mohfw.gov.in/, [retrieved: 25/03/2020] . 18. khanna, r.c., honavar, s.g., (2020) all eyes on coronavirus what do we need to know as ophthalmologists, indian journal of ophthalmology, 68(4) 549–553. 19. siriprapaiwana, .s , moorea, e. j., koonpraserta, s. (2018) generalized reproduction numbers, sensitivity analysis and critical immunity levels of an seqijr disease model with immunization and varying total population size, mathematics and computers in simulation,146,70-89. 20. kucharski, a.j., russell, t. w., diamond, c., liu, y., edmunds, j., funk, s.,eggo, r. m. (2020) early dynamics of transmission and control covid-19: a mathematical modelling study, lancet infect. dis. doi:10.1016/s1473–3099(20)30144–4. 21. senapati, a., rana1b, s., dasb, t., chattopadhyaya, j. (2021)impact of intervention on the spread of covid-19 in india: a model based study, j. theor biol., doi:10.1016/j.jtbi.2021.110711, pmid:33862090. 22. zhao, s., lin, q., ran, j. (2020) preliminary estimation of the basic reproduction number of novel coronavirus (2019-ncov) in china, from 2019 to 2020: a data-driven analysis in the early phase of the outbreak, int. j. infect. dis. 92, 214–217. copyright © 2021 assa. adv syst sci appl (2021) 86 n.k. goswami, b. shanmukha 23. ivorra, b., ferrández, m. r., vela-pérez, m.,and ramos.a. m. (2020) mathematical modeling of the spread of the coronavirus disease 2019 (covid-19) considering its particular characteristics. the case of china. communications in nonlinear science and numerical simulation. accepted . preprint version in researchgate , doi link: http://www.doi.org/10.13140/rg.2.2.21543.29604. 24. chen, t., rui, j.,wang, q., zhao1, z., cui, j,.and yin, l. (2020) a mathematical model for simulating the phase-based transmissibility of a novel coronavirus, infectious diseases of poverty 9:24, https://doi.org/10.1186/s40249-020-00640-3 25. castillo-chavez, c., feng, z., and huang, w. (2002) on the computation of r0 and its role on global stability, in: mathematical approaches for for emerging and reemerging infectious diseases. springer-verlag, 229-250. 26. driessche, p.v. and watmough, j. (2002) reproduction numbers and sub-threshold endemic equilibria for compartmental models of disease transmission mathematical biosciences, 180, 29-48. 27. korobeinikov, a., wake, g. c. (2002) lyapunov function and global stability for sir, sirs, and sis epidemiological models. appl. math. lett. 15:955–960. 28. li my, muldowney js (1996) a geometric approach to global stability problems. siam j math anal 27(4):1070–1083. 29. li, m. y., muldowney, j. s. (1995) global stability for the seir model in epidemiology. math biosci 125:155–164. 30. lasalle, j.p. (1976) the stability of dynamical systems. regional conference series in applied mathematics. siam, philadelphia. 31. srivastav, a.k., tiwari, p.k., srivastava, p.k., ghosh, m., kang, y. (2020) a mathematical model for the impacts of face mask, hospitalization and quarantine on the dynamics of covid-19 in india: deterministic vs. stochastic, mathematical biosciences and engineering, 18 (1), 182–213. 32. lenhart, s., workman, j. t. (2007) optimal control applied to biological models. crc press, boca raton. 33. pontryagin, l.s., boltyanskii, v. g., gamkrelidze, r.v., mishchenko, e. f. (1962) the mathematical theory of optimal processes, inter science publishers, geneva. 34. goswami, n. k. and shanmukha, b. (2020)‘modeling and analysis of symptomatic and asymptomatic infections of zika virus disease with non-monotonic incidence rate’, applied mathematics and information sciences, 14(4), 655-671. 35. goswami , n. k. (2021)’modeling analysis of zika virus with saturated incidence using optimal control theory’, international journal of dynamical systems and differential equations, 11 (3/4), 287–301. 36. olaniyi, s., obabiyi, o.s., okosun, k.o., adewale, a.t., adewale, s.o. (2020)mathematical modeling and optimal cost-effecting control of covid-19 transmission dynamics. eur phys j plus.135(11). copyright © 2021 assa. adv syst sci appl (2021) introduction the model analysis of the model system basic reproduction number existence of endemic equilibrium point stability analysis of the model global stability of disease-free equilibrium global stability of endemic equilibrium sensitivity analysis of reproduction number effects of parameters on disease outbreak the optimal control model the optimal control problem existence and characterization of optimal controls simulation of optimal control problem strategy a: employing hygiene promotion, social distancing and self-isolation (u1), only. strategy b: increase testing facility and treatment of the symptomatic individuals (u2) only. strategy c: employing both the control interventions (u1,u2). conclusion microsoft word 641 article text, copyedited.doc adv syst sci appl 2021; 02; 1-7 published online at https://ijassa.ipu.ru. smart control mechanisms in digital technologies of decision-making vladimir n. burkov1, irina v. burkova2*, viktor g. zaskanov3 1) trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: vlab17@bk.ru 2) trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: irbur27@mail.ru 3) samara national research university, samara, russia e-mail: zaskanov@mail.ru abstract: this paper considers the problems of adopting digital technologies in decision-making. two types of digital decision technologies are described, namely, the direct (or traditional) technology with human decision-making based on computer-aided advisor systems and also the inverse technology with independent computer decision-making in which a human merely monitors the decision process. these technologies are compared with each other, and some applications are discussed. keywords: digital economy, correct mechanisms, decision support system, the kantorovich– glushkov conditions, direct and inverse technologies 1. introduction the development of digital economy as well as a wide adoption of digital technologies, not only in data processing and electronic documentation but also in solid decision-making, require a proper review of decision methodology. the attempts to use digital technologies in decision-making were undertaken long ago [1]. in this context, note the role of soviet academicians l.v. kantorovich and v.m. glushkov. in the wake of the 1970s large-scale automation in the ussr, glushkov suggested the idea of a computer-aided planning system for the national economy. kantorovich designed a mechanism to coordinate national economic plans with the interests of enterprises based on his original methodology of duality. the optimal dual estimates of associated planning problems were considered as a tool of such coordination. unfortunately, the attempt failed because the real interests of economic agents poorly fit the strict framework of duality. kantorovich and glushkov formulated two prerequisites for an efficient use of digital technologies in decision-making as follows: (1) all information necessary for decisionmaking must be reliable and (2) the economic agents must benefit by implementing given decisions (fulfilling their plans). in the sequel, these prerequisites will be called the kantorovich–glushkov conditions. in the 1970s, they were difficult to satisfy [1] but the situation has dramatically changed in the recent time, following the design of smart control mechanisms: contract theory [2, 3, 5, 8], the theory of mechanism design, where the incentive compatibility conditions and the revelation principle [6, 7], as well as the work [4, 7] and many others. * corresponding author: irbur27@mail.ru 2 v.n. burkov, i.v. burkova, v.g. zaskanov copyright ©2021 assa. adv. in systems science and appl. (2021) note that the incentive compatibility conditions, and the revelation principle were originally proposed in 1971 in article [9], where the revelation principle was called the "fairplay principle". let us discuss this class of mechanisms in detail. any changes in social being, particularly, control mechanisms, lead to fast consequences in human behavior. smart control mechanisms are the control mechanisms that change human behavior for public interests, e.g., stimulate truth-telling (the provision of reliable information), decision implementation, efficient development, etc. of crucial importance for digital technologies are the smart mechanisms that guarantee truth-telling and decision implementation: decision-making with these mechanisms satisfy the kantorovich–glushkov conditions. it is possible to identify two types of digital decision technologies as follows. the first type, known as the direct technology (traditional for modern society), relies on human decision-making with computer-aided decision support systems. the development of such technologies is connected with the design of active advisor systems: after the implementation process, the decision of a human decision-maker (dm) is compared with the one suggested by an advisor system using the so-called recalculation models. the technology of the second type (inverse) is based on independent computer decisionmaking while a human merely monitors the decision process without any interference (except for critical situations and force majeur). in such technologies, a decision support system (dss) becomes a decision-making system (dms) while a dm a person who assesses decisions (decision analyzer, da). however, control mechanisms are designed by managers and other interested persons who are responsible for the results of functioning. 2. direct (traditional) decision technology fig. 1. direct decision technology the structural diagram of the direct decision technology is shown in fig. 1. as mentioned above, the power of decision belongs to a decision-maker (dm), who uses information about an object and external environment as well as his/her experience and dss recommendations. figure 1 has the following notations: u as the principal’s decision (in this case, the principal is the dm); v as the dss recommendation; j as the available information about the object; finally, x(u) as the result of decision implementation. this technology demonstrates high efficiency if the principal well knows the object, seeks for its successful functioning and has operative tools to guarantee decision implementation. otherwise, the dm encounters unreliable information from the object, the non-implementation of required decisions and a possible occurrence of corruption ties. consider the direct technology on elementary examples of order allocation. dss object dm x(u) j v u smart control mechanisms in digital technologies of decision-making 3 copyright ©2021 assa. adv. in systems science and appl. (2021) example 1. the principal needs products in an amount r. there are n suppliers of these products. denote by xi the production plan of supplier i. let the product cost be given by , where a parameter ri describes the production efficiency of supplier i. the principal’s problem is to find the plans xi so that (1) (the supply satisfies the demand) and the total cost ф(х) are minimized, i.e., . (2) for the exact values r = (ri), the optimal plan has the form , , (3) where h is the total efficiency of all suppliers, . however, the principal often knows only approximate values of the parameters ri. in this case, the information s = (si) about these parameters is reported by the suppliers. the interest of supplier i is determined by the profit , (4) where λ denotes the product price established by the principal. he/she makes some decision using a control mechanism x = π(s) and λ(s), where π(s) are λ(s) are the functional relationships between the plan (x), control (λ) and the information s reported by the suppliers. the revelation mechanism [2] suggested in the theory of active systems is correct, i.e., guarantees truth-telling and plan fulfilment. in example 1, the revelation mechanism has the form (5) the direct technology suffers from several drawbacks as follows. first, it is difficult to construct an adequate model of an object, and even more difficult to design an adequate correct decision mechanism. second, this technology is susceptible to corruption––possible collusions between the principal and suppliers and also between suppliers themselves. so, the direct technology was further refined to the two-channel control mechanisms (active advisor systems) based on recalculation models (rms) [3]. an rm yields a more accurate estimate of the object’s parameters using additional information about the results of decision implementation. this approach allows comparing the economic effects from the decisions made by the principal and the ones suggested by a dss. the structural diagram of decision-making processes in two-channel mechanisms is presented in fig. 2. 21 2 ii irz x= i i x r=å ( ) 21min 2 = å ix i i ф x x r i i rx r h = 1, ,= !i n j j h r=å ( ) 21, 2i i i i i y x x x r l l= , 1, , , . = = = å !i i i i x s i n r s l l 4 v.n. burkov, i.v. burkova, v.g. zaskanov copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 2. decision-making in two-channel mechanisms figure 2 has the following additional notations: rm as a recalculation model; δ as the difference between the economic effects from the dm’s decision and dss recommendation. if δ > 0, the dm obtains an incentive; if δ < 0, a penalty. the two-channel mechanisms stimulate the principal’s interest in correct decisions and also reduce the corruption component. example 2. for the order allocation mechanism of example 1, a recalculation model can be easily constructed in the following way. let zi be the real cost of supplier i ( ). then , , and hence , (6) where wi denotes the plan recommended by the dss. consider a collusion between the dm and some supplier (without loss of generality, supplier 1). using his/her power of decision, the dm overrates the price λ1 for supplier 1. to this end, supplier 1 pays the dm a bribe––α% of his/her profit. the profit of supplier 1 makes up , and his/her beneficial plan is x1 = (1 – α)λ1 r1. without corruption, supplier 1 obtains the profit . the collusion makes sense for him/her if . so the beneficial corruption condition takes the form , or . let , where k > 1 is some parameter adjusted by the dm if he/she enters collusion with supplier 1. then the beneficial corruption condition can be written as 21 2 ii irz x= 2 2 i i i xr z = 1, ,= !i n 2 2 i i i i i w z z x æ ö d = -ç ÷ è ø å ( ) 2 1 1 1 1 1 2 xx r a l( )21 12 r h r ( ) 2 2 2 1 1 1 1 11 2 2 rr r h a l æ ö> ç ÷ è ø ( ) 11 r h a l> ( ) 1 1 1r r r h a l> > 1 r hkl = dss object dm y(u) j v u rm δ smart control mechanisms in digital technologies of decision-making 5 copyright ©2021 assa. adv. in systems science and appl. (2021) . note that, the higher is the supplier’s efficiency, the smaller is his/her gain from collusion. moreover, under a sufficiently large penalty for any deviations from the recalculation model-based cost, corruption becomes unbeneficial to the supplier. the major difficulty for a practical use of two-channel mechanisms is the design of adequate recalculation models. 3. inverse decision technology consider the inverse decision technology in which the principal’s role is played by a computer (digital decision support system, ddss) while the principal’s supervisor monitors the decision process without any interference (except for critical situations and force majeure). actually the supervisor and ddss change their places, which explains the term “inverse technology”. for making this technology efficient, the mechanism should be discussed with all interested persons and then enacted as a law. in this sense, the ddss performs a purely computational function. the whole responsibility for decisions lies on the designers of the mechanism (in the first place, the principal’s supervisor), not on the ddss. if he/she concludes that the mechanism is inefficient (e.g., in new conditions of functioning), then the supervisor considers possible corrections. of course, as noted earlier, the mechanism must satisfy the kantorovich–glushkov conditions, i.e., be correct. this forms the main problem for adopting digital technologies in decision-making and management: the correctness of mechanisms is not so trivial to guarantee. let us describe a possible design of correct mechanisms. consider a system of a single principal and n agents (enterprises, organizations, etc.). the interests of each agent i, , are determined by his/her goal function yi(l, xi), where l denotes control and xi is the plan established by the ddss. the sequence of functioning is as follows. 1. the principal announces the set l of admissible controls. 2. for each λ ∈ l, the agents report their beneficial plans to the principal: . 3. the principal chooses λ ∈ l and the corresponding plans xi(λ), , that maximize his/her goal function ф(x, λ). this mechanism is correct under the hypothesis of weak contagion [2], stating that all agents do not consider the influence of their messages on the control l. indeed, for any values l, the agents obtain beneficial plans and hence are interested in truth-telling and plan fulfillment. remark. if the set of admissible controls is large, the principal does not need to announce all its elements. he/she may organize an iterative procedure, adding new controls at each step to increase the value of his/her goal function. such a procedure resembles decomposition methods. example 3. consider the order allocation problem as described in example 1. assume the principal announces the set of admissible prices l (step 1) while the agents report their beneficial plans . the principal chooses l to minimize the total cost subject to the constraint . for this problem, the hypothesis of weak contagion was justified in the theory of active systems in the case of sufficiently many agents [2]. ( ) 1 1 1h k r a> > 1, ,= !i n ( ), 1, , ,= î!ix i n ll l 1, ,= !i n ( ) , 1, ,= = !i ix r i nl l ( )i i x rl =å 6 v.n. burkov, i.v. burkova, v.g. zaskanov copyright ©2021 assa. adv. in systems science and appl. (2021) an important advantage is that the inverse technology considerably reduces corruption effects. really, the principal’s supervisor has no right to interfere into the ddss and hence cannot affect decision-making. so collusions between the principal and agents take no place. however, corruption may occur when the agents report information to the principal. for struggling against this corruption component, the inverse technology should be augmented by a recalculation model. the main stages of adopting digital technologies in economy include the following. 1) choosing a situation that can be formalized as a model; 2) designing a smart control mechanism that guarantees truth-telling and decision implementation as well as coordinates the interests of all persons involved; 3) developing a software product, rules and regulations in accordance with this mechanism; 4) implementing this mechanism, analyzing its efficiency and making periodic corrections subject to new conditions of functioning. 4. some applications of digital technologies consider some applications of digital technologies [3]. 1. the automated quantitative complex assessment of activity results (akkord). this system was designed at trapeznikov institute of control sciences in the 1980s for assessing the efficiency of all industrial enterprises within the ministry of instrumentation, automation and control systems of the ussr (minpribor), following the initiative of minister m.s. shkabardnya. the general assessment method and procedure were discussed at the scientific and engineering council of the ministry and also at the board of the ministry and then approved by shkabardnya. calculations were performed using a special software complex without human interference. 2. cost-effective taxation systems. in 1990–1991, trapeznikov institute of control sciences participated in an experiment on new taxation systems for research activities organized by the state committee for science and technology of the ussr. the experts of the institute suggested the cost-effective taxation system––a smart incentive mechanism that stimulates any organizations (even monopolies) to reduce cost and prices. the corresponding rules, regulations and software were developed after debates at the scientific council of the institute. for two years, the institute was functioning in experimental conditions, and the results confirmed theoretical expectations—it was not beneficial to overrate the cost of contract work. 3. the reverse priority mechanism for limited resource allocation. this smart mechanism stimulates customers to submit objective requests for a certain resource (truthtelling). the mechanism was adopted in the water resource management system of bulgaria and highly appreciated later (awarded by the golden mark for high technological advance). 5. conclusion the article discusses two smart decision-making mechanisms based on digital technologies – direct and inverse. the question arises why such mechanisms are few in practice. we give four reasons: 1. the complexity of the smart mechanisms design. it is not by chance that several nobel prizes in economics were received for the development and implementation of one or another smart mechanism. 2. weak interest of officials and managers. indeed, the introduction of managerial innovations in an organization, region or industry is a considerable risk. drawing an analogy with medicine, we can say that this is an operation on a living organism, which is carried out smart control mechanisms in digital technologies of decision-making 7 copyright ©2021 assa. adv. in systems science and appl. (2021) by the patient himself. not every leader will dare to do it. moreover, an effective system of estimation and motivation of managers and officials for the result has not yet been created. 3. the problem of compatibility, the essence of which is that an excellent mechanism is often rejected by the organization because it is incompatible with its mentality, traditions, etc. 4. lack of trained personnel capable of developing and implementing smart mechanisms. the theory of smart mechanisms design and the theory of active systems are studied in a small number of universities. nevertheless, the authors are convinced that the future belongs to smart mechanisms. note the most promising directions to adopt smart digital technologies. 1. assessment systems for the activity of managers and officials at any level, with stimulation based on activity results. 2. allocation systems for limited resources (finances, etc.). 3. efficiency improvement programs for organizations, municipalities, regions, etc. acknowledgements this work was partially supported by the russian foundation for basic research, project no. 18-07-01258. references 1. ivanov, v. & malinetskii, g. (2017). strategicheskie prioritety tsifrovoi ekonomiki [strategic priorities of digital economy], strategicheskie prioritety, 3, 54–95. 2. burkov, v. (1977). osnovy matematicheskoi teorii aktivnykh sistem [fundamentals of mathematical theory of active systems]. moscow, russia: nauka. 3. goubko, m., burkov, v., kondrat’ev, v., korgin, n. & novikov, d. (2013). mechanism design and management: mathematical methods for smart organizations. new york, usa: nova science publishers. 4. mas-colell, andreu, michael dennis whinston, and jerry r. green (1995) microeconomic theory. vol. 1. new york: oxford university press. 5. bolton p., dewatripont m. contract theory. – cambridge: mit press, 2005. – 740 p. 6. harsanyi j. (1967, 1968) games with incomplete information played by "bayesian" players // management science. part i: 1967. vol. 14. № 3. p. 159 – 182. part ii: 1968. vol. 14. № 5. p. 320 – 334. part iii: 1968. vol. 14. № 7. 486–502. 7. maschler m., solan e, zamir s. (2013) game theory, cambridge: cambridge university press, 1008 p. 8. salanie b. (2005) the economics of contracts. 2nd edition, massachusetts: mit press. – 224 p. 9. burkov v.n., lerner a.j. (1971) fairplay in control of active systems. differential games and related topics, amsterdam: north-holland publishing company, 325–344. adv syst sci appl 2021; 03:1–11 published online at https://ijassa.ipu.ru. simultaneously robust chaos synchronization and identification of unknown parameters for nonlinear gyro system abolghasem daeichian1,2*, shahram aghaei3 1department of electrical engineering, faculty of engineering, arak university, arak, 38156-8-8349, iran 2institute of renewable energy, arak university, arak, 38156-8-8349, iran 3department of electrical engineering, yazd university, yazd, iran abstract: this paper concerns robust synchronization and parameter identification for nonlinear gyroscope systems. gyros are widely utilized in navigational applications where synchronization plays a vital role. a system of nonlinear dynamical equations with some parameters presents a model of gyro systems. the parameters of gyro can vary in time, which can lead to desynchronization of the gyro systems. in this paper, we assume that a gyro system has bounded time-varying unknown parameters and the synchronization problem is considered in two situations. first, the synchronization of two gyroscopes with identical dynamical model and, second, the synchronization of a gyroscope with the rössler system. the lyapunov stability theory with control terms is employed to cope with the problem. also, the identification of timevarying unknown parameters is the side goal of the paper. the proposed scheme synchronizes chaotic nonlinear systems in both situations appropriately. in addition, the slave parameters converge to the nominal values of master parameters despite uncertainty. simulation results illustrate the superiority of the proposed method. keywords: robust synchronization, chaos, parameter identification, gyroscope, rössler system 1. introduction numerous studies confirm that complex and chaotic behaviors are observed in physics, mechanics, engineering, etc. [1–4]. the chaotic behavior of dynamical systems is interested in many researches [5,6]. also, chaos synchronization in applied systems is studied [7]. the intent of chaos synchronization is to synchronize the states of at least two chaotic systems despite the difference in initial conditions, the difference in model parameters, and even the difference in dynamical models. how to reach synchronization under such distinctions is a challenging crux [8]. the identification of unknown system parameters is also the side goal of many studies. many schemes have been proposed to synchronize two chaotic systems, such as adaptive control [9–11], sliding mode control [12, 13], robust methods [14, 15], intelligent base methods [16–19] and lyapunov base methods [20–23]. the gyroscopes are widely employed in navigational, aeronautical, and space engineering. those have complex dynamical motion including periodic, sub-harmonic, quasi-periodic, and chaotic [24]. gyro synchronization for a couple of chaotic systems with one-way linking is considered in [25] and it is extended in [26] by applying an active control. fuzzy sliding mode control is employed to synchronize uncertain chaotic nonlinear gyros in [27]. a variable structure control approach using neural network with multi-quadratic radial basis ∗corresponding author: a-daeichian@araku.ac.ir, a.daeichian@gmail.com 2 a. daeichian, s. aghaei function is introduced in [28] where gyros with known and unknown system parameters are synchronized. the existence of external disturbance is investigated in [29], where fast nonsingular terminal sliding mode surface is employed to derive an adaptive finite-time control in order to synchronize the gyro systems. synchronization and parameter identification for the symmetric gyroscope system with constant parameters have been done in [30]. it is not robust with respect to parameter’s variation and does not converge to the nominal value of parameter. this paper concerns developing robust synchronization based on the lyapunov theorem with control terms for nonlinear gyroscope systems with unknown time-varying parameters. in addition, unknown parameter identification is expected. the idea is to find the control input u such that the parameters’ uncertainty eliminates in the lyapunov function. two situations are considered for the slave system. first, a similar gyro dynamical system, and second, the rössler system which has another dynamical equation of master gyro. the parameters of slave system or control input converge to the nominal value of master system parameters by a parameter updating law in spite of the uncertainty. it is proved by the lyapunov stability theorem that the proposed controller can synchronize chaotic systems with time-varying unknown parameters. simulation results show that method in [30] has a larger margin of error than the proposed robust method here. section 2 states the problem. controller and the parameter identifier are designed in section 3. simulation results verify achieving synchronization in section 4. finally, section 5 concludes the paper. 2. problem statements fig. 2.1 shows a gyro which is established on a vibrating base. the dynamical equations of the system are expressed by the euler’s angles θ, φ and ψ and a multiple harmonic equation∑n k=1ak sinωkt which represents base vibration. if we define x1 = θ, x2 = θ̇ and x3 = φ̇, the dynamical equations of the system in the state space framework are [30, 31]: ẋ1 = x2 ẋ2 = −(βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 − c1(t) i x2 + mgl i sinx1 − mg i n∑ k=1 ak sin (ωkt) sinx1 ẋ3 = −2 cosx1 sinx1 x2x3 + βϕx2 i sinx1 (2.1) where i , mg, and l denote the polar moments of gyro inertia, the gravity force, and the distance between origin and the center of gravity. βφ and βϕ are constants of the motion. c1(t) is a parameter vector which is unknown, time varying, and uncertain. it worth noting that x1 = θ, which is known as nutation angle, is the angle between thexy z (fixed) axis and xyz (body) axis. it is shown in fig.(11-16) of [31] that θ(t) is limited by two extreme values 0 < θ1 < θ(t) < θ2 < π which correspond to the turning points of the central-force. this fact prevents eq. (2.1) to become singular. for more details on the derivation of this model we refer to chapter 11 of [31]. however, c(t) has the nominal value c with uncertainty or disturbance ∆c(t), that is to say c1(t) = c + ∆c(t). it is assumed that the uncertainty is bounded, namely: assumption 2.1: (uncertainty boundedness) ‖∆c‖ ≤ k ∀t ∈ r+ (2.2) copyright © 2021 assa. adv syst sci appl (2021) simultaneously robust chaos synchronization and identification of unknown 3 where k is a known, positive, and real number. fig. 2.1. schematic of a gyroscope on a vibrating base eq. (2.1) is the master (drive) system model. the major goal is that the states of a slave (response) system synchronize with the states of the master system. here, we consider two different dynamical systems as slave system. first, a gyro with state space representation similar to the master system: ẋs1 = xs2 + u1 ẋs2 = −(βφ − βϕ cosxs1)(βϕ − βφ cosxs1) i2 sin3 xs1 − c2(t) i xs2 + mgl i sinxs1 − mg i n∑ k=1 ak sin (ωkt) sinxs1 + u2 ẋs3 = −2 cosxs1 sinxs1 xs2x s 3 + βϕx s 2 i sinxs1 + u3, (2.3) second, the rössler system as slave system: ẋs1 = −xs2 − xs3 + u1 ẋs2 = xs1 + b1x s 2 + u2 ẋs3 = a1 + xs3 (xs1 − c1) + u3 (2.4) where a1, b1, and c1 are the rössler system parameters. three control inputs u1, u2 and u3 are augmented to the slave dynamical equations to control the synchronization of the systems. c2(t) is the parameter vector of the slave system. the aim is to not only synchronize the states of two gyros but also c2(t) converges to c (nominal value of the unknown parameter), simultaneously. considering assumption 2.1, any robust controller has to satisfy the following conditions: 1. limt→∞ ‖e(t)‖ = 0 2. limt→∞ ‖c̃(t)‖ = 0 where e(t) = xs(t)− x(t) and c̃(t) = c2(t)− c. copyright © 2021 assa. adv syst sci appl (2021) 4 a. daeichian, s. aghaei 3. design of controller and identifier 3.1. gyro system as slave if gyro dynamical equations (2.3) are considered as a slave system, then the error dynamic can be written in the form: ė1 = e2 + u1 ė2 = −(βφ − βϕ cosxs1)(βϕ − βφ cosxs1) i2 sin3 xs1 + (βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 −c2 i xs2 + c1 i x2 + mgl i (sinxs1 − sinx1) −mg i n∑ k=1 ak sin (ωkt)(sinx s 1 − sinx1) + u2 ė3 = −2 cosxs1 sinxs1 xs2x s 3 + 2 cosx1 sinx1 x2x3 + βϕx s 2 i sinxs1 − βϕx2 i sinx1 + u3 (3.5) where e1 = xs1 − x1, e2 = xs2 − x2, and e3 = xs3 − x3. the following theorem introduces control inputs u1, u2, u3, and parameter identification updating law such that limt→∞ ‖e(t)‖ = 0 and limt→∞ ‖c̃(t)‖ = 0. theorem 3.1: consider the master system (2.1) with unknown parameter c1(t) which satisfies assumption 2.1 and the slave system (2.3). then, there exist the control laws u1 = −e2 − e1 u2 = (βφ − βϕ cosxs1)(βϕ − βφ cosxs1) i2 sin3 xs1 − (βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 +( c2 i − 1)e2 − mgl i (sinxs1 − sinx1)− ke2x2 i‖e2x2‖ x2 + mg i n∑ k=1 ak sin (ωkt)(sinx s 1 − sinx1) u3 = 2 cosxs1 sinxs1 xs2x s 3 − 2 cosx1 sinx1 x2x3 − βϕx s 2 i sinxs1 + βϕx2 i sinx1 − e3 (3.6) and the parameter updating law ċ2 = e2x2 i (3.7) which leads to master-slave synchronization. moreover, the unknown parameter of the slave system converges to the nominal value of the master one. proof consider the following lyapunov function: v (e1, e2, e3, c̃) = 1 2 (e21 + e22 + e23 + c̃2) (3.8) copyright © 2021 assa. adv syst sci appl (2021) simultaneously robust chaos synchronization and identification of unknown 5 where c̃(t) = c2(t)− c. the first derivation of the lyapunov function along the error dynamic eq.(3.5) is: v̇ = e1(e2 + u1) + e2(− (βφ − βϕ cosxs1)(βϕ − βφ cosxs1) i2 sin3 xs1 + (βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 − c2 i xs2 + c1 i x2 + mgl i (sinxs1 − sinx1)− mg i n∑ k=1 ak sinωkt(sinx s 1 − sinx1) + u2) +e3(− 2 cosxs1 sinxs1 xs2x s 3 + 2 cosx1 sinx1 x2x3 + βϕx s 2 i sinxs1 − βϕx2 i sinx1 + u3) + c̃ċ2 = −e21 − e22 − e23 + ∆c i x2e2 − ke22x2 2 i‖e2x2‖ (3.9) taking into account the inequality ‖a+b‖ ≤ ‖a‖+ ‖b‖, we obtain the inequality: v̇ ≤ −‖e‖2 + ‖∆c i ‖‖e2x2‖ − ‖ k i ‖‖e2x2‖ (3.10) finally, assumption 2.1, gives: v̇ ≤ −‖e‖2 (3.11) so, v̇ < 0 for ‖e‖ 6= 0. this lyapunov function and its derivation demonstrate that the error dynamic eq.(3.5) is asymptotically stable, that is to say, the master-slave synchronization is achieved. it also guarantees the convergence of the slave system parameter c2(t) to the nominal value of the master system parameter c. 3.2. rössler system as slave now, the rössler system is considered as a slave. the error dynamic in this situation is: ė1 = −xs2 − xs3 − x2 + u1 ė2 = xs1 + b1x s 2 + c1 i x2 − mgl i sinx1 + (βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 + mg i n∑ k=1 ak sin (ωkt)(sinx1) + u2 ė3 = a1 + xs3(x s 1 − c1) + 2 cosx1 sinx1 x2x3 − βϕx2 i sinx1 + u3 (3.12) the control inputs and parameter identification updating law are given as the following theorem. theorem 3.2: consider the master system (2.1) with unknown parameter c1(t) which satisfies assumption copyright © 2021 assa. adv syst sci appl (2021) 6 a. daeichian, s. aghaei 2.1 and the slave system (2.4). then, there exist the control laws u1 = xs2 + xs3 + x2 − e1 u2 = −xs1 − b1xs2 − c2 i x2 − ke2x2 i‖e2x2‖ x2 + mgl i sinx1 −(βφ − βϕ cosx1)(βϕ − βφ cosx1) i2 sin3 x1 − mg i n∑ k=1 ak sin (ωkt)(sinx1)− e2 u3 = −a1 − xs3(xs1 − c1)− 2 cosx1 sinx1 x2x3 + βϕx2 i sinx1 − e3 (3.13) and the parameter updating law, ċ2 = e2x2 i (3.14) which leads to master-slave synchronization. moreover, the unknown parameter of the slave system converges to the nominal value of the master one. proof the proof can be written down in the same way as the proof of theorem 1 by considering the error dynamic eq.(3.12). so, the two gyro systems synchronize despite initial state and parameter mismatches. the next section shows the simulation results of the proposed method. furthermore, the outcomes of the proposed method are compared with the results of [30]. 4. simulation consider the dynamical model parameters βφ = 2, βϕ = 5, i = 1,mg = 4, l = 0.25, ω1 = 1, and a1 = 12.1. the rössler system is chaotic where a1 = 0.2, b1 = 0.2, and c1 = 5.7. also, we consider different initial conditions for the master gyro as x1(0) = −0.5, x2(0) = −1.2, x3(0) = 10, slave gyro as xs1(0) = xs2(0) = xs3(0) = 0.1, and slave rössler system as xs1(0) = −0.1, xs2(0) = 0.2, xs3(0) = 0.5. assuming the unknown, time varying, and uncertain parameter c1(t) = 0.5 + 0.7 sin t where the nominal value is c = 0.5 and ∆c = 0.7 sin t. thus, the maximum norm of ∆c can be set as k = 0.7. the initial value of slave system parameter is also assumed as c2 = 0. moreover, the simulation time step is chosen as 0.0001 (sec). the states of two master and slave gyros as well as the error of their synchronization by the method of [30] are depicted in fig. 4.2. the mean square errors (mse) of the states synchronization are 6.617× 10−4, 1.151× 10−2, and 1.802× 10−1 for the e1, e2, and e3, respectively. since the unknown parameter is in the dynamical equation of x2, the steady state synchronization error is observed in e2. this can be quantified by neglecting transient time, i.e. considering mse for data obtained after 10 seconds from the beginning which are 1.721× 10−12, 7.749× 10−03, and 4.686× 10−10 for e1 to e3, respectively. it can be seen in fig. 4.2 that e2 does not converge to zero and it has fluctuations between −0.120 and 0.265. this fact is due to the steady state error in the parameter identification algorithm which is shown in fig.4.5. the mean value of the identified parameter by the method of [30] after initial transient time is 0.447 which suffers approximately 10.6% steady state error. theorems 1 and 2 assert that control inputs (3.6) and (3.13) with the parameter-updating laws (3.7) and (3.14) not only synchronize the master and slave systems but also guarantee the identification of the unknown parameter of the master system correctly. figs. 4.3 and 4.4 illustrate the states and synchronization error for gyro-gyro and gyro-rössler systems by the proposed method, respectively. the simulation results show that synchronization copyright © 2021 assa. adv syst sci appl (2021) simultaneously robust chaos synchronization and identification of unknown 7 0 50 100 150 -3 -2 -1 0 x 1 0 50 100 150 -5 0 5 x 2 0 50 100 150 0 10 20 x 3 0 50 100 150 -3 -2 -1 0 x 1s 0 50 100 150 -5 0 5 x 2s 0 50 100 150 0 10 20 x 3s 0 50 100 150 time (sec) -0.5 0 0.5 1 e 1 0 50 100 150 time (sec) -0.5 0 0.5 1 1.5 e 2 0 50 100 150 time (sec) -10 -5 0 e 3 fig. 4.2. method of [30]: row 1: master system states, row 2: slave gyro system states, row 3: synchronization error 0 50 100 150 -3 -2 -1 0 x 1 0 50 100 150 -5 0 5 x 2 0 50 100 150 0 10 20 x 3 0 50 100 150 -3 -2 -1 0 x 1s 0 50 100 150 -5 0 5 x 2s 0 50 100 150 0 10 20 x 3s 0 50 100 150 time (sec) -0.5 0 0.5 1 e 1 0 50 100 150 time (sec) -0.5 0 0.5 1 1.5 e 2 0 50 100 150 time (sec) -10 -5 0 e 3 fig. 4.3. proposed method (gyro-gyro): row 1: master system states, row 2: slave gyro system states, row 3: synchronization error error asymptotically tends to zero for all states regardless of uncertainty in parameters. in the case of gyro-gyro synchronization, the mse of the synchronized x2 reduces to 1.019× 10−3 which shows a decrease of more than 91%. omitting initial transient time give mse equal to 3.45× 10−5, which is almost 99.5% more accurate than the previous method. synchronization mses for the gyro-rössler situation are 2.941× 10−4, 1.202× 10−3, and 1.659× 10−1 related to e1, e2, and e3, respectively. the mses decrease to 0.765× 10−12, 0.399× 10−4, and 4.315× 10−10 if neglecting the initial transient time. the gyro-rössler copyright © 2021 assa. adv syst sci appl (2021) 8 a. daeichian, s. aghaei 0 50 100 150 -3 -2 -1 0 x 1 0 50 100 150 -5 0 5 x 2 0 50 100 150 0 10 20 x 3 0 50 100 150 -3 -2 -1 0 x 1s 0 50 100 150 -5 0 5 x 2s 0 50 100 150 0 10 20 x 3s 0 50 100 150 time (sec) -0.5 0 0.5 1 e 1 0 50 100 150 time (sec) -0.5 0 0.5 1 1.5 e 2 0 50 100 150 time (sec) -10 -5 0 e 3 fig. 4.4. proposed method (gyro-rössler): row 1: master system states, row 2: slave gyro system states, row 3: synchronization error 0 100 200 300 time (sec) -0.2 0 0.2 0.4 id e n ti fi e d v a lu e 0 100 200 300 time (sec) 0 0.2 0.4 0.6 id e n ti fi c a ti o n e rr o r fig. 4.5. gyro-gyro synchronization: left: parameter identification, right: identification error; dashed-dotted line (black): nominal value, dashed line (blue): proposed method, solid line (red): method of [30] 0 100 200 300 time (sec) -0.2 0 0.2 0.4 id e n ti fi e d v a lu e 0 100 200 300 time (sec) 0 0.2 0.4 0.6 id e n ti fi c a ti o n e rr o r fig. 4.6. gyro-rössler synchronization: left: parameter identification, right: identification error; dasheddotted line (black): nominal value, dashed line (blue): proposed method, solid line (red): method of [30] synchronization also ensures convergence to zero of all errors. the mses for all situations are written in table 4.1. the unknown parameter is also identified perfectly by the developed method. figs. 4.5 and 4.6 indicate the convergence of the unknown parameter to the nominal value c as well as the error of estimation for the proposed method and that of [30] in the situation of gyro-gyro and gyro-rössler systems, respectively. it can be seen that the fluctuations in the e2 decrease as time goes by and the parameter identification error converges to zero in the proposed method while the method of [30] suffers lasting fluctuations as well as steady state error. copyright © 2021 assa. adv syst sci appl (2021) simultaneously robust chaos synchronization and identification of unknown 9 table 4.1. mse of synchronization mse mse for time>10 (sec) method e1(×10−4) e2(×10−3) e3(×10−1) e1(×10−12) e2(×10−4) e3(×10−10) method of [30] 6.617 11.51 1.802 1.721 77.49 4.686 proposed gyro-gyro 6.617 1.019 1.802 1.721 0.345 4.686 method gyro-rössler 2.941 1.202 1.659 0.765 0.399 4.315 5. conclusion the proposed synchronization of gyro system has been founded based on the elimination of dynamical equations’ nonlinearity by the controller. simultaneously, a parameter identification algorithm estimates the unknown, time-varying, and uncertain parameters. the proposed method not only guarantees convergence of synchronization error to zero but also ensures the convergence of unknown parameter estimation to the nominal value. considering a typical gyroscope, the synchronization error in order of magnitude 10−12, 10−5, and 10−10 for e1, e2, and e3, respectively, are obtained in the situation of gyro-gyro and gyro-rössler synchronization. moreover, the unknown parameter is estimated at time 300 second with less than 2.7% of error. however, the proposed algorithm requires the upper bound of the parameter uncertainty to be determined. another point is that the control inputs are a complex nonlinear function of states that impose high computation load. nomenclature parameter definition x, y, z inertia base vectors (fixed) x, y, z body axis θ, φ, ψ the euler’s angles i the polar moments of gyro inertia l the distance between origin and the center of gravity m the mass of gyro g gravitational acceleration constant ωk k-th harmony of the base vibration frequency ak amplitude of base vibration with frequency ωk βφ, βϕ the constants of the motion a1, b1, c1 the rössler system parameters xi i-th state of the master system (i=1,2,3) xsi i-th state of the slave system (i=1,2,3) references 1. rakkiyappan, r., latha, v. p., zhu, q., & yao, z. (2017) exponential synchronization of markovian jumping chaotic neural networks with sampled-data and saturating actuators. nonlinear analysis: hybrid systems, 24, 28–44. 2. aghaei, s., daeichian, a., & puig, v. (2020) hierarchical decentralized reference governor using dynamic constraint tightening for constrained cascade systems. journal of the franklin institute, 357, 12495–12517. copyright © 2021 assa. adv syst sci appl (2021) 10 a. daeichian, s. aghaei 3. daeichian, a. & honarvar, e. (2018) modified covariance intersection for data fusion in distributed nonhomogeneous monitoring systems network. international journal of robust and nonlinear control, 28, 1413–1424. 4. zhu, q. & wang, h. (2018) output feedback stabilization of stochastic feedforward systems with unknown control coefficients and unknown output function. automatica, 87, 166–175. 5. vaidyanathan, s. & rasappan, s. (2014) global chaos synchronization of n-scroll chua circuit and lur’e system using backstepping control design with recursive feedback. arabian journal for science and engineering, 39, 3351–3364. 6. aghaei, s. & daeichian, a. (2017) a step-by-step algorithm for plotting local bifurcation diagram. journal of nonlinear analysis and application, 2017, 12–18. 7. vaidyanathan, s. (2015) global chaos synchronization of the lotka-volterra biological systems with four competitive species via active control. international journal of pharmtech research, 8, 206–217. 8. azar, a. t. & vaidyanathan, s. (2016) advances in chaos theory and intelligent control, vol. 337. springer. 9. tan, m., pan, q., & zhou, x. (2016) adaptive stabilization and synchronization of nondiffusively coupled complex networks with nonidentical nodes of different dimensions. nonlinear dynamics, 85, 303–316. 10. fotsin, h. & woafo, p. (2005) adaptive synchronization of a modified and uncertain chaotic van der pol-duffing oscillator based on parameter identification. chaos, solitons & fractals, 24, 1363–1371. 11. park, j. h. (2005) adaptive synchronization of rossler system with uncertain parameters. chaos, solitons & fractals, 25, 333–338. 12. li, z. & shi, s. (2003) robust adaptive synchronization of rossler and chen chaotic systems via slide technique. physics letters a, 311, 389–395. 13. wang, c.-c. & su, j.-p. (2004) a new adaptive variable structure control for chaotic synchronization and secure communication. chaos, solitons & fractals, 20, 967–977. 14. suykens, j. a., curran, p. f., vandewalle, j., & chua, l. o. (1997) robust nonlinear h∞ synchronization of chaotic lur’e systems. ieee transactions on circuits and systems i: fundamental theory and applications, 44, 891–904. 15. liu, b., chen, g., teo, k. l., & liu, x. (2005) robust global exponential synchronization of general lur’e chaotic systems subject to impulsive disturbances and time delays. chaos, solitons & fractals, 23, 1629–1641. 16. alasty, a. & salarieh, h. (2008) identification and control of chaos using fuzzy clustering and sliding mode control in unmodeled affine dynamical systems. journal of dynamic systems, measurement, and control, 130. 17. yau, h.-t. & shieh, c.-s. (2008) chaos synchronization using fuzzy logic controller. nonlinear analysis: real world applications, 9, 1800–1810. 18. iwasaki, m., miwa, m., & matsui, n. (2005) ga-based evolutionary identification algorithm for unknown structured mechatronic systems. ieee transactions on industrial electronics, 52, 300–305. 19. poznyak, a. s., yu, w., & sanchez, e. n. (1999) identification and control of unknown chaotic systems via dynamic neural networks. ieee transactions on circuits and systems i: fundamental theory and applications, 46, 1491–1495. 20. faieghi, m., baleanu, d., et al. (2014) sampled-data nonlinear observer design for chaos synchronization: a lyapunov-based approach. communications in nonlinear science and numerical simulation, 19, 2444–2453. 21. ouannas, a., odibat, z., shawagfeh, n., alsaedi, a., & ahmad, b. (2017) universal chaos synchronization control laws for general quadratic discrete systems. applied mathematical modelling, 45, 636–641. 22. shen, l. & wang, m. (2008) robust synchronization and parameter identification on a class of uncertain chaotic systems. chaos, solitons & fractals, 38, 106–111. copyright © 2021 assa. adv syst sci appl (2021) simultaneously robust chaos synchronization and identification of unknown 11 23. zhi, l. & chong-zhao, h. (2001) adaptive control and identification of chaotic systems. chinese physics, 10, 494. 24. wang, c.-c. & yau, h.-t. (2011) nonlinear dynamic analysis and sliding mode control for a gyroscope system. nonlinear dynamics, 66, 53–65. 25. chen, h.-k. (2002) chaos and chaos synchronization of a symmetric gyro with linearplus-cubic damping. journal of sound and vibration, 255, 719–740. 26. lei, y., xu, w., & zheng, h. (2005) synchronization of two chaotic nonlinear gyros using active control. physics letters a, 343, 153–158. 27. yau, h.-t. (2008) chaos synchronization of two uncertain chaotic nonlinear gyros using fuzzy sliding mode control. mechanical systems and signal processing, 22, 408–418. 28. deori, p. b. & kandali, a. b. (2016) synchronization and chaos control of heavy symmetric chaotic unknown gyroscope using mqrbvsc. in 2016 international conference on intelligent control power and instrumentation (icicpi), 12–16, ieee. 29. yin, l., deng, z., huo, b., & xia, y. (2017) finite-time synchronization for chaotic gyros systems with terminal sliding mode control. ieee transactions on systems, man, and cybernetics: systems, 49, 1131–1140. 30. ge, z.-m. & lee, j.-k. (2005) chaos synchronization and parameter identification for gyroscope system. applied mathematics and computation, 163, 667–682. 31. thornton, s. t. & marion, j. b. (2004) classical dynamics of particles and systems. brooks/cole—thomson learning. copyright © 2021 assa. adv syst sci appl (2021) introduction problem statements design of controller and identifier gyro system as slave rössler system as slave simulation conclusion adv syst sci appl 2020; 04; 11-26 published online at https://ijassa.ipu.ru. frequency-geometric identification of magnetization characteristics of switched reluctance machine laila kadi1*, adil brouri1, abdelmalek ouannou1 1) imsm team, l2mc laboratory, ensam, moulay ismail university, meknes, morocco e-mail: lailakadigi@gmail.com, a.brouri@ensam-umi.ac.ma, abdelmalekouannou@gmail.com abstract: accurate modeling of electrical drives for online testing is a relevant problem. switched reluctance machine (srm) has lately attracted significant attention because it has several advantages compared to conventional engines. it is simple, free of rare-earth and faulttolerant machine. an analytical model of a srm has not been reported yet due to the dynamic and strong nonlinearity of srm. the most srm control and applications are based on several assumptions and simplifications. therefore, it is convenient to develop an accurate approach to identify the srm characteristics. in this paper, an analytical modeling and identification method of magnetization characteristics of switched reluctance machine (srm) is proposed. presently, an exact mathematical model of srm is established. unlike several previous studies, in this approach the system nonlinearities of srm are allowed to be hysteresis (i.e. the hysteresis effect is considered) and taking account the inherent magnetic nonlinearity. indeed, the srm is considered as highly nonlinear which makes the modeling of these machines difficult to achieve. then, it is convenient to develop an accuracy model of srm because it is always operated in the magnetically saturated mode to maximize the energy transfer. the developed model can be used in control, simulation and design development. furthermore, an identification method, at standstill test, based on frequency technics is developed allowing the identification of srm nonlinearities (considering the saturation and the hysteresis effects). in this respect, it is shown that the nonlinear behavior of srm can be exactly described by a block-oriented nonlinear structure. specifically, the srm can be described by a wiener nonlinear model. compared to the existing methods, the proposed study gives good accuracy of flux-linkage value characteristics and enjoys the simplicity of implementation. keywords: switched reluctance machine, hysteresis effect, saturation effect, nonlinear identification, nonlinear modeling, wiener model, frequency-geometric method 1. introduction the world is now focused on how to reduce the cost of equipment used in renewable energy, electrical vehicles (ev), as well as in other applications. for wind energy conversion, many research has been done to evaluate the generator that is characterized by low cost, easy to maintain, high performance, and able to work under variable speeds [1, 2]. moreover, the choice of motor drives for ev is very critical. in this context, the srm is typically selected as a generator or motor because of its low cost, simplicity, robust structure, small moment of inertia rotor and fault-tolerant capability. then, it is easy to maintain. note that the majority of losses show up in the stator because the srm has no rotor winding, no brushes, and does not have a magnet permanent [3] (see fig. 1), there are also others benefits of srm. in view of these significant remarks, the srm is relatively easy to cool and is insensitive to high temperatures [4]. the srm can be a suitable choice for wind energy and ev applications. in this respect, the srm becomes more competitive compared to the classical machines such as induction machines. however, its major defects are torque ripples and acoustic noise [4]. note that several studies are still ongoing to mitigate these effects [5, 6]. * corresponding author: lailakadigi@gmail.com 12 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) even though the several advantages of the srm, the strong nonlinearity makes the modeling of srm not an easy task. despite all efforts in modeling and identification of srm, finding an analytical model, which can accurately represent the nonlinear behavior of srm, is widely open for scientific research. as commonly used in the literature, the modeling and identification methods of srm flux linkage or self-inductance characteristics are classified into 4 categories: experimental measurements [7, 8], numerical solutions [9, 10], intelligent methods [11, 12], and analytical approaches [13, 14]. roughly, the experimental methods are costly, require considerable time, can damage the machine and possibly create a safety hazard. therefore, this method can be used to validate the obtained results. previous research has shown that two magnetization characteristics have been generally looked for srm modeling. the first method is established using the inductance characteristics based model. the second solution is based on the flux-linkage [4]. at this stage, it is worth emphasizing the difficulty to establish a model that can accurately describe the behavior of the machine. this study proposes the development of an exact analytical model for the srm, which is featured by a simple structure and reduced computational time. then, an identification method is proposed for determining the flux linkage magnetization characteristics. lately, researchers have shown an increased interest in using block-oriented nonlinear models based on the input and output systems [16–18]. in this paper, it is proved that the srm behavior can be accurately designed using block-oriented nonlinear model. specifically, it is shown that the srm at standstill can be exactly described by the wiener model (fig. 2). this later is given by the cascade connection of a linear dynamic system followed by a static or dynamic nonlinearity block [17, 18] as shown in fig. 2. to the author’s knowledge, for the first time, it has been shown that the srm can be exactly described by the wiener model. as soon as a model of srm is obtained, the identification method of the system parameters using this model constitutes also an important problem. in this article, a new approach is suggested for the identification of the magnetization characteristics often called flux linkage characteristics are much dependent on current and rotor position of the motor. note that the problem of identifying the srm nonlinearities will be carried out including saturation and hysteresis effects. presently, a frequency-domain solution is designed to determine the srm nonlinearities. the frequency identification methods have been widely used to identify the nonlinear systems parameters [16, 19, 20]. it is shown that the srm can be analytically described by a nonlinear system of wiener model. the proposed frequency approach consists of applying a sine input voltage ( ) sin( )v t v t= to the phase winding, for any given rotor position. then, by recording the current measurement at the winding terminal, the system nonlinearities of the model describing the srm can be identified using a simple frequency-geometric method. then, the srm can be modeled analytically using the data acquisition obtained by the finite element method (fem). note that this study is quite different from previous works [21–24]. the first key step in the development of this identification approach is the determination of an accurate model describing the nonlinear behavior of srm. another key step consists of designing an identification method using an appropriate excitation signal to estimate the parameters and the characteristics of srm. then, the system nonlinearities of srm are allowed to be static or hysteretic operator. furthermore, the model and the identification solution take into account the inherent magnetic nonlinearity. then, using the developed nonlinear model of srm, an identification method at standstill test is developed. the identification approach is based on frequency technics using simple sine input voltage. it is shown that the inner signal of the srm model can be analytically determined at any time t. in the first stage, the hysteresis effect is not considered. a filtering algorithm of the current at the winding terminal is suggested. accordingly, the system nonlinearities of the model describing the srm can be easily identified using the recording of the phase current https://www.linguee.fr/anglais-francais/traduction/commonly+used.html 13 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) (the filtered signal) over one half-period of time. then, the flux-linkage characteristics can be easily obtained. furthermore, these results are compared with those obtained using magnetostatic method. in the second stage, the saturation and the hysteresis effects of srm nonlinearities are considered. for any given rotor position, the srm phase winding is excited using also a sine input voltage. then, using the proposed filtering algorithm of the phase current, the system nonlinearities of the model describing the srm can be accurately estimated using the recorded data over one period of time. accordingly, the flux-linkage characteristics can be provided without any other experiment (using uniquely the phase voltage and the recorded current). compared to the magnetostatic method, it is shown that the proposed study gives good accuracy of flux-linkage value characteristics, enjoys the simplicity of implementation and characterized by a reduced computational time. to the authors’ knowledge, no previous study has found an exact model simulating the behavior of srm. in addition, very few published works have attempted to identify the hysteresis operator. the paper is organized as follows: the principle of operation and the equations describing the srm are stated in section 2. then, the mathematical model and the description of the identification method of srm nonlinearity are dealt within section 3. section 4 presents examples of obtained results using the proposed study. fig. 1.1. example of a 8/6 srm ξ v x i y fig.1.2. wiener model 2. operation principle and equations describing srm srm is a double saliency device in which the stator comprises a salient magnetic circuit provided with coils supplied by a current i, and the rotor is also a salient ferromagnetic circuit but without any conductor or magnet [24], [25] as shown in fig.1.1. in such machine, torque is produced by the tendency of its rotor to move to a position where the reluctance of the excited winding is minimized (aligned position) [24], [25]. then, the machine rotates the magnetic field by synchronizing continuously and consecutively the supply of different phases with the rotor position by means of the firing angles (turn-on and turn-off). srm can operate as a motor or generator. in order to obtain motoring mode, a stator phase is excited when the rotor is moving from an unaligned position towards the aligned position. likewise, rotor (does not contain any coil) stator coils ( )ig s h(.) 14 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) by exciting a stator phase when the rotor is moving from aligned towards the unaligned position, a generating mode will be attained [26]. to feed each stator phase of the srm, we generally use an asymmetric half-bridge converter [27], [28] (fig.2.3). the magnetic flux created by the ampere-turns ni oscillates between two extreme values correspond to [29]: unaligned position in which the magnetic circuit has a maximum reluctance (minimum inductance) as illustrated by fig. 2.4a. aligned position in which the magnetic circuit has a minimum reluctance (maximum inductance) as explained by fig. 2.4b. the structure of the studied machine is of four phase 8/6 srm, where 8 designates the number of poles in the stator (ns) and 6 is the number of poles in the rotor (nr). the choice of ns and nr is important since they have significant implications on the torque. the material of the srm stator and rotor is steel m36 with the nonlinear b-h characteristic. the mutual coupling between adjacent phases is very small and can be neglected[25], [29]. let i, r, l, ω, θ, and v denote the phase current, phase resistance, inductance, rotor speed, rotor position angle, and phase voltage respectively. then, the voltage equation of a phase winding can be given as: ( ) ( ) ( , )d i i dt v t r t   = + () where λ is the flux-linkage per phase. under the hypothesis of nonmagnetic coupling between phases (at any given time, only one phase is excited), the flux-linkage can be expressed as follows: ( , ) ( , )i l i i  = (2.2) due to the nonlinear behavior of the ferromagnetic material, the flux in the stator phases varies nonlinearly according to the rotor position θ and the phase current i. accordingly, through direct derivation of (2.2) in (2.1), one immediately gets the following expression of the phase voltage: ( ) ( ) ( ) ( ) ( ), , ,i i l i l idi di v t r t l i dt i dt      = + +    +     () the last term in (2.3) is the back electromotive force (emf) voltage that is expressed by: ( ), e i l i   =   () fig.2.3. asymmetric half bridge converter 15 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) (a) (b) (c) fig.2.4 (a) unaligned position; (b) aligned position; (c) midway position 3. modeling and identification approach of srm at standstill test, the back electromotive force (emf) voltage is nil because it is proportional to the shaft speed. accordingly, the expression of phase voltage in (2.3) becomes: ( ) ( ) ( ) ( ), ,i i l idi di v t r t l i dt i dt  = + +   () the rotor of srm is blocked at specific positions relative to the phase to be tested, e.g. at aligned, unaligned, and a midway between the above two positions. at blocked rotor (for any given rotor position  ), the nonlinearity ( )l i  (i.e. the inductance according to the current phase) can be described using a finite orthogonal function, e.g. polynomial decomposition [30], [31]: 0 ( ) n k k k θ a il i = =  () where n is the polynomial function degree and ka its coefficients. then, it is readily follows from (3.1) and (3.2) that the phase voltage can be expressed as: ( )k+1 k=0 ( ) n k d b i dt v t =  (3.3a) where: 1 1 2 1 k k r a for k b a for k + = =      (3.3b) this result means that the phase voltage can be obtained by deriving a polynomial function according to the phase current. then, let introduce the following polynomial function: k+1 k=0 ( ) n kf i b i=  (a) then, in view of (3.4a), it follows that the phase voltage ( )v t in (3.3a) can be rewritten as follows: ( ) ( ) df i dt v t = (3 .4b) remark 3.1: 16 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) it follows from the formulae (3.2)-(3.4) show that the nonlinear system having, the phase current i(t) as input and the phase voltage v(t) as output, can be described by a hammerstein model, i.e., nonlinear element (f(i)) followed by a linear block (derivative term). at this point, it is worth emphasizing that the derivative action can be modeled by a linear block. it is readily follows that the analytical nonlinear system (3.4a-b) (the phase voltage v(t) according to the phase current i(t) for any given rotor position) can be modeled by a nonlinear subsystem followed by a linear dynamical element as shown by fig.3.1. this nonlinear model is called hammerstein system. fig. 3.1. hammerstein nonlinear model the aim presently is to develop an identification approach allowing to estimate the inductance nonlinearity ( )l i  , for any given rotor position  , and the flux-linkage ( , ) ( , )i l i i  = . furthermore, the hysteretic shape of the flux linkage can be seen. then, it is worth emphasizing that, from practical point of view, the input signal of srm is not a current excitation, but rather a phase voltage and the output signal of srm is the phase current. indeed, the srm is generally excited by a phase voltage and the phase current can be viewed as the output measurement. then, the output electrical signal of srm (i.e. the phase current i(t)) according to the input excitation (i.e. the phase voltage v(t)) can be modeled by a nonlinear system of wiener model (fig.1.2), where h(.) is the system nonlinearity, ( )ig s is the linear subsystem and ξ(t) is the measurement noise. specifically, the linear element is an integrator block ( ( ) 1/ig s s= ). the above model is analytically described by the following equations: ( ) ( )* ( )ix t g t v t= () ( )( ) ( )i t h x t= () ( ) ( ) ( )y t i t t= + () where: 0 ( ) ( ) st i ig t g s e ds + =  (the inverse laplace transform of ( )ig s ); the symbol '*' refers to the convolution operation; v(t) and y(t) represent the (measurable) input and output of the system, respectively. in this respect, note that the internal signal x(t) is not accessible to measurement. at this stage, an identification approach is proposed based on frequency method. then, the srm characteristics and the parameters determination can be dealt using frequency-geometric tools. accordingly, the following sine signal is being applied to the system input (phase voltage): ( ) sin( )v t v t= () in the steady state, the internal signal x(t) is also sinusoidal (of the same period of v(t)) and can be written as [15], [16], [18]: ( ) ( / ) cos( )x t v t = − () f(.) g(s) i v f ( . ) 17 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) commonly, the main complexity of wiener model identification problem is due to the fact that the internal variable x(t) is not accessible to measurement. fortunately, in our nonlinear model, it follows (3.9) the inner signal x(t) can be easily estimated at any time t. presently, a frequency domain method is proposed to identify the system nonlinearity h(.) using the data obtained from finite element method. in this respect, the sinusoidal voltage ( ) sin( )v t v t= is being applied to the phase winding and the resulting phase current is recorded. then, the srm nonlinearities characteristics can be accurately determined using only the system input (i.e. the phase voltage v(t)) and the measurable output signal (i.e. the phase current i(t)). it is interesting to bear in mind that, if the input excitation v(t) is t=2π/ω-periodic, then the steady state of inner signal x(t) and output system i(t) are also periodic of the same period t as v(t) [15], [18]. this result is a quite interesting achievement as it shows that, the curve of i(t) according to x(t) can be easily obtained using uniquely the set of points ( ) ( )( ), ( ) -( / ) cos( ), ( )x t i t v t i t = , in the steady state, over one period t. the determination of the functions ( ) ( )( ), ( ) ( / ) cos( ), ( )x t i t v t i t = − constitutes another key challenge of this work, because it makes possible the estimation of nonlinearity h(.) of srm model (fig. 1.2) at any given rotor position  . practically speaking, the useful information (the phase current i(t)) is not directly accessible to the measurement. indeed, the measurement signal y(t) is the phase current with additive noise (please see (3.7)). let ˆ ( ) m i t denotes the phase current estimate. in this respect, thanks to periodicity of i(t) (t=2π/ω), an accurate estimate of the phase current can be obtained using the following estimator: ( )  ) 1 0 1ˆ ( ) 2 / f 0 2 /or m mdef k i t y t k t m     − = = +  (a) ˆ ˆ( 2 / ) ( ) m m i t k i other iset w + = (b) where m is an integer, preferably of large value. then, the identification of the system nonlinearity h(.) can easily be achieved using the collected data over the interval  )0 2 /m  . for convenience, let us summarize the identification steps of system nonlinearity h(.): firstly, using the set of recorded data  ) ( ); 0y t for t mt , the estimate of the current phase i(t) can be obtained using the estimator (3.10a-b). then, we generate the fictive internal signal ( ) ( / ) cos( )x t v t = − over one period t. therefore, the estimate ˆ (.) m h of nonlinear element h(.) can be simply performed by constructing the set of points: ( )  ) ˆ( ), ( ) ; 0 m x t i t for t t . finally, an accurate estimate of srm nonlinearity, for different rotor positions, can be obtained using the above frequency-geometric method. the above results lead to the following proposition. proposition 3.1: let consider the identification problem of srm model at standstill (fig. 1.2), where the noise ( )t is supposed to be a zero-mean ergodic stochastic process. this latter is excited by the sine input ( ) sin( )v t v t= . accordingly, one has with probability 1 (w.p.1): 1. the current estimate ˆ ( ) m i t converges to the true phase current signal i(t). 18 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) 2. the system nonlinearity estimate ˆ (.) m h converges to the true function h(.). proof. to prove the first part of the proposition, let us combining (3.7) and (3.10a-b), one thus gets: ( ) 1 1 0 0 1 1 ˆ ( ) 2 / 2 /( ) m m m k k i t i t k t k m m     − − = = = + + +  () bearing in mind that the phase current i(t) is t=2π/ω -periodic, (3.11) can thus be rewritten as: 1 0 1 ( ) ( ) ( 2 / )ˆ m m k t i t t k m i    − = = + + () using the fact that the noise is a zero-mean ergodic stochastic process, the last term in (3.12) converges to zero as m tends to infinity (w.p.1). this completes the first part of proposition. to prove the second part of proposition, let us recall that for the input ( ) sin( )v t v t= (fig.1.2), the inner signal x(t) belongs to the interval / /v v   − . accordingly, the system nonlinearity h(.) can be described, within the interval / /v v   − , as the set of all points ( ) ( )( )  ) ( )( ), ( ) ( ), ; 0x tx t i t x t h t t=  . therefore, the nonlinearity estimate ˆ (.) m h can be featured by plotting the set of points ( )  ) ˆ( ), ( ) ; 0 m x t i t t t . it follows from the first part of proposition that: ( ) ( )  )ˆ( ), ( ) ( ), ( ( )) 0lim   m m x t i t x t h x t t tfor any → =  () this means that the nonlinearity estimate ˆ (.) m h converges to the true function h(.) w.p.1 for any  )0t t . this completes the proof of proposition. this result is a quite interesting achievement as it shows that, it is possible to obtain an accurate estimate of srm characteristics for any given rotor position . at this stage, other significant results can be observed. on one hand, if the system nonlinearity h(.) is static, the estimate ˆ (.) m h of h(.) can be done by using simply the set of points ( )ˆ( ), ( ) m x t i t over one-half period  )0 / 2t t . on the other hand, in the case where the nonlinear element h(.) of srm is hysteresis, the system nonlinearity estimate ˆ (.) m h can be featured using the set of points ( )ˆ( ), ( ) m x t i t over one period of time (i.e. ( )  ) ˆ( ), ( ) ; 0 m x t i t for t t ). these results lead to the following proposition. proposition 3.2: let us reconsider the above identification problem of srm model (fig. 1.2). this machine is excited by the phase voltage (3.8), where the amplitude v and the frequency ω (rd/s) of v(t) are kept constant. let’s assume that, in steady state, ( ) ( )( ) (( 1) )h hx kt x k t= + . then, the following results are obtained: 1. if the hysteresis effect of srm magnetization characteristics is not considered, for / /x v v    − the nonlinear element h(.) (in the steady state) can be constructed by fitting the set of points (x(t), i(t)) over one half-period (  )0 / 2for t t ). 19 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) 2. now, the hysteresis effect of srm magnetization characteristics is considered. then, for any fixed v and  , the hysteretic nonlinearity h(.) (in the steady state) can be obtained by plotting the set of points ( )  ) ( ), ( ) ; 0x t i t t t (i.e. over one period of time). proof. first, one immediately gets using (3.6) -(3.9) for any  )0t t : ( ) ( )( ) ( ) ( / ) cos( )i t h x t h v t = = − () then, if the srm system nonlinearity h(.) is static (does not contain the hysteresis effect), then for any  )0 / 2t t there exists 't t t= − such that: ( ) ( )  )' / 2( ) ( ') ( )   t t tani t h t h x t dx = = () accordingly, the static nonlinear function h(.) can be described uniquely by the set of points ( )  ) ( ), ( ) ; 0 / 2x t t t ti  . this completes the first part of proposition. at this stage, the function h(.) is hysteretic nonlinear operator. accordingly, using the assumption ( ) ( )( ) (( 1) )h x kt h x k t= + one immediately gets that all points of coordinates ( )  ) ( ), ( ) ; 0x t t t ti  build a closed locus [16], [18]. furthermore, for any given amplitude v and frequency ω, a single closed locus is obtained. finally, the shape of hysteretic nonlinearity h(.), for fixed values of (v, ω), can be described by plotting the points of coordinates (x(t), i(t)) for  )0t t . presently, the aim is to present a method of modeling and identification of the fluxlinkage ( , )i  . the above obtained achievement is useful to determine the flux-linkage features. accordingly, it is readily follows from (2.1) that, the flux-linkage ( , )i  can be expressed as: ( )( )( )( , ) tv t ti r i d   −= (3.16) then, it follows from (3.9) and (3.16) that, the flux-linkage can be rewritten as: ( )( , ) cos( ) t v i t r i dt    = − −  (3.17) ( )( )( )( , ) x t h x ti r dt  = −  (3.18) where the first term part on the right side corresponds to x(t), given by (3.9). therefore, it readily follows from (3.17) that, the flux-linkage characteristics can be achieved using the resistor measurement r, the values of x(t) and the phase current measurement ˆ ( ) m i t . in this respect, it is worth emphasizing that no other information (experiment) is required to determine the flux-linkage characteristics. indeed, the latter can be estimated using only the available information’s in the first stage (i.e. ˆ ( ) m i t and x(t)). 20 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) 4. simulation results & discussion to show the efficiency of this approach, the machine 8/6 srm is considered in simulations. let us choose the rotor position 0 =  as the aligned position of phase a (fig. 2.4a). accordingly, 30 =  corresponds to the unaligned position of phase a (fig. 2.4b) and 15 =  corresponds to the midway position. simulations were carried out using a fem software by collecting data acquisition at several specific positions (e.g. at aligned, midway and unaligned positions). firstly, the srm is being excited by the sinusoidal phase voltage ( ) sin( )v t v t= , where the amplitude 400v v= and the frequency 314 / srd = . according to the proposition 3.2, note that the amplitude v and frequency  (rd/s) of the input signal are kept constant. accordingly, the steady state of phase current i(t) is also periodic of the same period of v(t). furthermore, note that a noise signal ( )t is added to the useful information (i.e. i(t)). this noise is supposed to be a zero-mean ergodic stochastic process. then, the system output y(t), i.e. the undisturbed phase current added to the measurement noise, is collected on the interval  )0t mt , where m=10. fig.4.1 illustrates y(t) for three rotor positions. the plot confirms that the phase current measurement y(t) is also periodic of the same period of v(t). then, using the estimator (3.10a-b), the phase current estimate ˆ ( ) m i t can be obtained for any given rotor position. this latter is plotted in fig. 4.2 over one period of time. it is clearly seen that, as the rotor gets closer to the aligned position, the saturation effect becomes more significant and notably distorts the shape of the phase current (the black curve in fig. 4.2). bearing in mind that x(t) is not accessible to measurement, but it can be easily determined at any time t using the expression (3.9) (fig. 1.2). then, the nonlinearity estimate ˆ (.) m h can be featured by the set of points ( )  ) ˆ( ), ( ) ; 0 m x t i t t t (using the phase current estimate ˆ ( ) m i t ). at this stage, it is worth emphasizing that two cases can be observed according to the system nonlinearity h(.) (proposition 3.2). fig. 4.1. output signal ( )y t fig. 4.2. phase current estimate signal ˆ ( )mi t 4.1. srm magnetization characteristics without considering the hysteresis effect (static curves) presently, we suppose that h(.) is static nonlinearity. then, the shape of h(.) can be described using the phase current estimate and the inner signal x(t) over one half-period (proposition 21 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) 3.2). accordingly, an accurate estimate ˆ (.) m h of the system nonlinearity h(.) can be obtained using the set of points ( )  ) ˆ( ), ( ) ; 0 / 2 m x t i t t t (i.e. by plotting the phase current estimate ˆ ( ) m i t in respect to x(t)). then, the nonlinearity estimate ˆ (.) m h for different rotor positions, e.g. aligned, midway and unaligned, can be achieved by repeating the same experiment. in this respect, the obtained nonlinearity estimate ˆ (.) m h for aligned, midway, and unaligned positions are shown by fig.4.3. the plots confirm that the function ˆ (.) m h presents a severe nonlinearity at aligned position due to the saturation effect. then, another important feature of srm is the flux-linkage, which is very interesting to be estimated. for convenience, let us consider the expression of flux-linkage ( , )i  in (3.17). accordingly, this latter can be determined by integrating the estimated current and using the signal x(t) (given by (3.9)). in this respect, it is worth emphasizing that the determination of flux-linkage features can be obtained using uniquely the results of experiments carried out in the first stage. in the case where the phase current is unidirectional, the first quadrant of (λ, i) plane is considered. then, the obtained flux-linkage characteristics are shown in fig.4.4. these results show that, using the proposed approach, the flux-linkage characteristics are acquired with good accuracy compared to those obtained using magnetostatic model. fig. 4.3. the estimate of the function h(.) without taking into account the hysteresis effect 22 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) fig. 4.4 flux-linkage characteristics without considering hysteresis effect 4.2. srm magnetization characteristics considering the hysteresis effect presently, the nonlinear elements are allowed to be hysteretic. to this end, a frequencygeometric tool is used as shown in section 3. in this respect, the determination of electrical nonlinearities and flux-linkage characteristics of srm can be found using the same experiment suggested in the first case (without considering the hysteresis effect). then, it follows proposition 3.2 and using phase current estimate ˆ ( ) m i t , the hysteretic nonlinearity estimate ˆ (.) m h can be featured by the set of points ( )ˆ( ), ( ) m x t i t for  )0 2 /t   . so, the curves of system nonlinearity estimate ˆ (.) m h for three key positions (aligned, midway and unaligned positions) are shown by fig.4.5. as can be seen, the obtained curves of ˆ (.) m h are not static and are closer to closed locus (in steady state). these locus will be referred to the hysteretic behaviors of srm. furthermore, note that the surface of hysteretic curve increases according to the position rotor. accordingly, the hysteretic losses are small at the aligned position and become high at the unaligned position. . fig. 4.5. the estimate of the function h(.) considering the hysteresis effect 23 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) in the case where the phase current is unidirectional, the first quadrant of ( , )i plane is considered. accordingly, the obtained flux-linkage ( , )i  characteristics are shown in fig.4.6. the achieved flux-linkage characteristics show that ( , )i  is not static, but it is close to a closed locus. moreover, the hysteretic surface is considerably narrow at the aligned position and increases according to the position rotor. note that, the validation of the identification method was performed using magnetostatic technic as can be seen in fig .4.4. let us recall that, the magnetostatic method is a pure numerical solution which requires a full set of data obtained from fem [25], [32]. this method differs substantially from our approach, which is based only upon the input-output measurement data. as can be seen in fig.4.4, the achieved flux-linkage shapes using the proposed method and those obtained using magnetostatic give very close results. in the case where the hysteresis effect is considered, the present analysis confirms the existence of hysteresis effect, which is not too discussed in the literature. fig. 4.6. flux-linkage characteristics with hysteresis effect 5. conclusion this paper proposes an analytical modeling and identification method of magnetization characteristics of srm. one key step in the design of such method is the construction of the nonlinear block-oriented model modelling the behavior of the srm. specifically, it is shown that the srm can be described by a wiener model. to our knowledge, no previous work has been dealt with this technic. then, an analytical identification method is developed to obtain the magnetization characteristics of the srm. the identification method is dealt based on a simple frequency-geometric method by applying a sine input voltage to the phase winding, for any given rotor position. then, the inner signal x(t) of the nonlinear wiener system (the srm model) can be easily estimated at any time t. accordingly, all the electrical parameters of srm can be accurately estimated using the phase voltage and the recording of the current measurement at the winding terminal. on one hand, if the hysteresis effect of srm nonlinearities is not considered, the magnetization characteristics can be estimated using data-acquisition over one-half period. the obtained results are compared with other achieved data using the fem method. on the other hand, the identification of srm magnetization features can be performed even if the hysteresis effect is considered using only data-acquisition over one period. in this respect, it 24 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) is worth emphasizing that very few available works have attempted to identify the hysteresis effects of srm. furthermore, it is shown that the hysteretic losses are generally small at the aligned position and can be of high value at the unaligned position. this result is confirmed by the area of the obtained hysteresis curve. another originality of the present study lies in the fact that the machine model can be analytically represented using block-oriented systems. the proposed wiener model still be very simple in use and the magnetization characteristics can be easily achieved using the frequency-geometric method. another feature of this work is the fact that the modeling design and the identification scheme require less computational time, unlike several previous studies. the accuracy of the model has been verified by comparing the obtained characteristics using this approach with those using the magnetostatic method. despite all efforts in modeling and identification of srm, finding an accurate analytical model, which represents the nonlinear behavior of srm, is wide open for scientific research. in terms of future work, it would be interesting to use the same model in real time. then, this model can used to obtain the dynamical model of srm. the hysteresis measurement (surface of the hysteretic curve) can offer some suggestions to handle the losses. this approach could be tested on some applications of srm using the obtained model. references 1. r. cardenas, r. pena, m. perez, j. clare, g. asher, and p. wheeler, “control of a switched reluctance generator for variable-speed wind energy applications,” ieee transactions on energy conversion, vol. 20, no. 4, 781–791, 2005. 2. d. a. torrey, “switched reluctance generators and their control,” ieee transactions on industrial electronics, vol. 49, no. 1, 3–14, 2002. 3. t. j. e. miller, switched reluctance motors and their control. magna physics , oxford : clarendon press , new york : oxford university press., 1993. 4. a. emadi, handbook of automotive power electronics and motor drives. taylor & francis., 2017. 5. n. saha, a. k. panda, and s. panda, “speed control with torque ripple reduction of switched reluctance motor by many optimizing liaison technique,” journal of electrical systems and information technology, vol. 5, no. 3, 829–842, dec. 2018. 6. d. marcsa and m. kuczmann, “design and control for torque ripple reduction of a 3phase switched reluctance motor,” computers & mathematics with applications, vol. 74, no. 1, 89–95, jul. 2017. 7. t. a. s. barros, p. j. s. neto, m. v. paula, a. b. moreira, p. s. n. filho, and e. r. filho, “automatic characterization system of switched reluctance machines and nonlinear modeling by interpolation using smoothing splines,” ieee access, 1–1, 2018. 8. m. v. terzic, h. li, b. bilgin, and a. emadi, “comparison of experimental methods for electromagnetic characterization of switched reluctance motors,” in 2018 xiii international conference on electrical machines (icem), 2018, 1881-1888. 9. b. parreira, s. rafael, a. j. pires, and p. j. costabranco, “obtaining the magnetic characteristics of an 8/6 switched reluctance machine: from fem analysis to the experimental tests,” ieee transactions on industrial electronics, vol. 52, no. 6, 1635–1643, dec. 2005. 25 frequency-geometric identification of magnetization characteristics copyright ©2020 assa. adv. in systems science and appl. (2020) 10. b. ganji, m. heidarian, and j. faiz, “modeling and analysis of switched reluctance generator using finite element method,” ain shams engineering journal, vol. 6, no. 1, 85–93, mar. 2015. 11. o. ustun, “a nonlinear full model of switched reluctance motor with artificial neural network,” energy conversion and management, vol. 50, no. 9, 2413–2421, sep. 2009. 12. x. yao and y. yang, “online modeling for switched reluctance motors using adaptive network based fuzzy inference system,” in 2017 29th chinese control and decision conference (ccdc), 2017, 1568-1573. 13. a. khalil and i. husain, “a fourier series generalized geometry based analytical model of switched reluctance machines,” in ieee international conference on electric machines and drives, 2005., 2005, 490-497. 14. s. li, s. zhang, t. habetler, and r. harley, “fast and accurate analytical calculation of the unsaturated phase inductance profile of 6/4 switched reluctance machines,” in 2016 ieee energy conversion congress and exposition (ecce), 2016, 1-8. 15. a. brouri, l. kadi, and s. slassi, “frequency identification of hammerstein-wiener systems with backlash input nonlinearity.,” international journal of control, automation and systems, vol. 15, no. 5, 2222–2232, oct. 2017. 16. f. giri, a. radouane, a. brouri, and f. chaoui, “combined frequency-prediction error identification approach for wiener systems with backlash and backlash-inverse operators,” automatica, vol. 50, no. 3, 768–783, mar. 2014. 17. f. giri, y. rochdi, a. radouane, a. brouri, and f. z. chaoui, “frequency identification of nonparametric wiener systems containing backlash nonlinearities,” automatica, vol. 49, no. 1, 124–137, jan. 2013. 18. a. brouri, f. giri, f. ikhouane, f. z. chaoui, and o. amdouri, “identification of hammerstein-wiener systems with backlash input nonlinearity bordered by straight lines,” ifac proceedings volumes, vol. 47, no. 3, 475–480, jan. 2014. 19. a. brouri, o. amdouri, f. z. chaoui, and f. giri, “frequency identification of hammerstein-wiener systems with piecewise affine input nonlinearity,” ifac proceedings volumes, vol. 47, no. 3, 10030–10035, jan. 2014. 20. m. aguado-rojas, p. maya-ortiz, and g. espinosa-pérez, “on-line estimation of switched reluctance motor parameters,” international journal of adaptive control and signal processing, vol. 32, no. 6, 950–966, 2018. 21. j. hur, c. c. kim, and d. s. hyun, “modeling of switched reluctance motor using fourier series for performance analysis,” journal of applied physics, vol. 93, no. 10, 8781–8783, may 2003. 22. z. peng, p. a. cassani, and s. s. williamson, “an accurate inductance profile measurement technique for switched reluctance machines,” ieee transactions on industrial electronics, vol. 57, no. 9, 2972-2979., sep. 2010. 23. j. e. stephen, s. k. kumar, and j. jayakumar, “nonlinear modeling of a switched reluctance motor using lssvm abc,” acta polytechnica hungarica, vol. 11, no. 6, 143–158, 2014. 24. t.j.e. miller, electronic control of switched reluctance machine. new york : oxford university press, 1993. 25. k. laila and b. adil, “nonlinear numerical study of mutual inductances for 26 l. kadi, a. brouri, a. ouannou copyright ©2020 assa. adv. in systems science and appl. (2020) switched reluctance machine,” 1–6, 2019. 26. a. arifin, “switched reluctance generator drive in the low and medium speed operation : modelling and analysis,” massey university, 2013. 27. e. darie and c. cepişcǎ, “the use of switched reluctance generator in wind energy applications,” 2008 13th international power electronics and motion control conference, epe-pemc 2008, no. october 2008, 1963–1966, 2008. 28. s. vukosavic and v. stefanovic, “srm inverter topologies: a comparative evaluation,” ieee transactions on industry applications, vol. 27, no. 6, 1034–1047, 1991. 29. r. krishnan, switched reluctance motor drives : modeling, simulation, analysis, design, and applications. crc press, 2001. 30. y. sofiane, a. tounzi, and f. piriou, “a non linear analytical model of switched reluctance machines,” applied physics, vol. 172, 163-172., 2002. 31. s. w. zhao, n. c. cheung, c. k. lee, x. y. yang, and z. g. sun, “survey of modeling methods for flux linkage of switched reluctance motor,” in 2011 4th international conference on power electronics systems and applications, 2011, 1-4. 32. l. kadi and a. brouri, “numerical modeling of a nonlinear four-phase switched reluctance machine,” in 2017 international renewable and sustainable energy conference (irsec), 2017, 1-6. frequency-geometric identification of magnetization characteristics of switched reluctance machine 1. introduction 2. operation principle and equations describing srm 3. modeling and identification approach of srm 4. simulation results & discussion 5. conclusion microsoft word 1165 article text, copyedited.doc adv syst sci appl 2021; 04; 100-114 published online at https://ijassa.ipu.ru. assessment of impact of trade wars on production and exports of the russian federation using the agent-based model aleksandra mashkova 1,2*, albert bakhtizin2 1) orel state university named after i.s. turgenev, orel, russia e-mail: aleks.savina@gmail.com 2) central economics and mathematics institute russian academy of sciences, moscow, russia e-mail: albert.bakhtizin@gmail.com abstract: in this article we consider methods of simulating consequences of trade wars using the agent-based model. the presented model simulates dynamics of trade relations between russia, the united states, china, the european union, and the rest of the world. we present event structure of the model which reflects interaction among different types of agents in the model: countries, organizations and residents. international trade is simulated at micro-level, as a set of supplies (purchases and sales) of organizations in different countries. volume and structure of trade flows is changed annually, based on the algorithm that takes into account existing and newly imposed trade restrictions and the expected change in final demand. for information support of the model we use official statistical sources of the countries included in the model. before loading these data to the model we process it to the similar sectoral structure and time period. for scenario calculations we consider optimistic and pessimistic scenarios for the world economy dynamics in the context of epidemiological risks and three possible options for world trade policy: preservation, cancellation or imposition of new trade restrictions. simulation results for the russian federation show that current sanctions against it affect exports in a number of industries, but do not have a significant impact on the dynamics of value added, while imposition of new restrictions could slow down the economic growth rate by 0.3-0.5% annually. keywords: agent-based model, trade wars, trade restrictions, import, export, scenario calculations 1. introduction various sanctions measures have increasingly become an instrument of world economic policy in recent years. imposed restrictions affect various sectors of the economy, access to financial markets, imports and exports of countries. despite the fact that a number of studies have shown that all participants suffer to a greater or lesser extent from participation in a trade war [4,5,18], restrictions continue to operate and be updated. in particular, since 2014, the russian federation has been under sanctions from the united states and the european union in relation to the financial sector, mining and transportation of minerals. in this regard, the question arises of assessing how much damage the imposed sanctions cause to the industries that fall under their influence, how they affect the output of these industries and the volume of export of their products to various countries. the subject of this study is trade flows among countries involved in trade conflicts, as well as changes in their volumes and structure under the influence of imposed trade restrictions. this task implies taking into account a large number of factors influencing trade flows, such as the final demand for products, availability of raw materials, utilization of fixed * corresponding author: aleks.savina@gmail.com assessment of impact of trade wars on production and exports of the russian… 101 copyright ©2021 assa. adv. in systems science and appl. (2021) assets, access to investments and fluctuations in prices for goods, therefore, to solve it, it is proposed to develop of a computer model for assessing consequences of trade wars. agent-based modeling was chosen as the main research method, which makes it possible to evaluate the dynamics of the global system as a result of the interaction of different agents: countries, organizations and residents. in the proposed model interaction among these agents determines the direction and structure of the trade flows among countries and their change under the influence of demand and government regulation. the developed model is based on arrays of real data on the economy and population of its member countries and trade exchange among them, which makes it possible to make reasonable forecasts of trade relations dynamics in various political conditions for each participating country. 2. literature review in the early 2000s most of the scientific publications devoted to analysis of trade wars referred to reviews of events and econometric analysis of retrospective data series [3,11]. later in 2010s the researcers turned to developing models including sets of abstract countries that do not reflect the characteristics of real countries [2,8,18]. in the last decade the task of forecasting the economic consequences of trade wars for the participating countries is most often solved using computable general equilibrium (cge) models, in which statistical data for different countries are loaded.the most famous project in the field of developing tools for quantifying trade wars between real countries is the global trade analysis project (gtap), which brings together scientific researchers from all over the world [1,7]. gtap was initiated in 1992 and proposed a unified methodology based on the cge approach. the model complexes developed using the gtap methodology include groups of countries (sometimes the whole world) and all sectors of the economy.within gtap a set of simulation complexes was developed: 1. worldscan, which consists of general equilibrium models of 29 products, trading in 30 countries, having the largest weight in world gdp, and consolidated regions that include several countries. in the model, supply and demand for a specific product in a given country are formed taking into account supply and demand of this product in other countries, and the price depends on substitution possibilities, transport costs, trade barriers, and other factors [4]. 2. globe multisector model [16], developed by specialists of the hohenheim university and the us naval academy, was used to assess the consequences of trade wars between the countries of the north american free trade zone, regulated by north american free trade agreement. 3. miragrodep multicountry multisector model, developed at the international research institute for food policy (washington, usa), based, in addition to the gtap methodology, on the more general mirage model (modeling international relationships under applied general equilibrium). the focus of this model is on goods trade between the us, china and mexico [5]. 4. the center for international trade and economics and the institute of world economy and politics of the chinese academy of social sciences have developed a global model to assess the consequences of the trade war between the united states and china [13]. the largest international organizations have developed their multi-country models: organization for economic co-operation and development (oecd) и international monetary fund (imf). the oecd’s new global model considers 6 groups of oecd countries and 3 groups of non-oecd countries. oecd [12]. the imf’s global macrofinancial model belongs to the class of dynamic stochastic general equilibrium models (dsge) and considers more than 40 of the largest economies in the world [20]. the listed models are aimed at considering the world's largest players, most often the usa and china, sometimes also eu countries. even though in some of the pre-sented model 102 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) complexes russia stands out as a separate participant, the available publications do not provide an assessment of the consequences of economic sanctions for russia, which have been regularly introduced against it since 2014. thus, developing a model of trade wars, which would include russia as one of the key participants, remains an urgent task. 3. materials and methods 3.1. structure of the model of trade wars within the agent-based model of trade wars, we study dynamics of trade relations between russia, the united states, china, the european union, and the rest of the world (fig. 3.1). the model is programmed as a software simulator which represents interactions among three types of agents: organizations, residents and states. fig. 3.1. concept of the model resident agents act as employees of organizations, on the one hand, and consumers of goods and services, on the other. agents' income goes to the accumulator, from which purchases of the final product are then paid. trade relations in the model are simulated at the level of individual organizations, and states regulate the processes of commodity exchange by introducing and canceling restrictions and changing tariffs on imports and exports of products. for agent-organizations, the current volume of output, the price in national currency, the number of workers and the volume of fixed assets are set. organizations interact with each other through supplies that reflect purchases and sales, including international ones. the suppliers in the model are agent-organizations, and the buyers can be other agent-organizations, the state, or the agentsresidents. supplies are divided into three types: 1. intermediate. intermediate supplies include supplies of raw materials, components and services to other organizations. for the convenience of further calculations, intermediate assessment of impact of trade wars on production and exports of the russian… 103 copyright ©2021 assa. adv. in systems science and appl. (2021) supplies are divided into basic (directly dependent on the dynamics of the organization's output) and additional (general business needs and accompanying services). 2. investment supplies, which include various types of fixed assets. investment supplies are also divided into basic (equipment and buildings) and additional (other supplies of products that are put on the balance sheet of organizations as fixed assets). 3. final products supplies purchased by resident agents. for each supply, the identifier of the supplier and the buyer is specified, as well as volume of supply in standard units, selling and purchase prices. the selling price is set in the currency of the state in which the supplier agent is located, excluding taxes. the purchase price consists of the sales price, sales taxes, export and import taxes (for international shipments), and is set in the currency of the country in which the agent-buyer is located. the states in the model regulate tax rates, benefits to the population, amount of financing of public sector organizations and subsidies, and also adjust the existing sanctions on commodity exchange with other countries. trade restrictions on import and export are set as the share of the turnover of industry j that falls under their influence, while ; . interaction of various agents in the model within its main events is presented on fig. 3.2. in the event “benefits payment”, the state transfers benefits to resident agents, while the corresponding accumulators (state budget and residents' funds) are changed. the event "public sector financing" implements transfer of funds from the budget to public sector organizations to pay their expenses. also, the state makes payments on the federal debt and may impose sanctions to other countries in the form of trade restrictions. fig. 3.2. event structure of the model the event "purchase of materials" simulates supply of intermediate products among production organizations, including import and export of products. when the product is sold, the payment is credited to the account of the agent-supplier at the selling price in the currency of sale and debited from the account of the agent-buyer in the currency of purchase. 104 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) trade organizations act in the same way in the event "wholesale purchase", resident agents in the event "retail purchase", and public sector organizations in the event "state procurement". after the completion of sales, taxes are transferred to the budgets of countries: internal taxes are charged on supplies where the supplier and the buyer were from the same country, and import / export taxes are charged on the international supplies. the event "wage payment" is implemented by organizations of all types, these funds are transferred to the accumulators of resident agents. for commercial organizations, the financial result is calculated. profit is defined as the difference between sales and costs of intermediate supplies, salaries and taxes. the profit is divided into two parts: the first is distributed as personal income from entrepreneurial activities, the second is used to pay for investments, along with subsidies received from the state. the production and sales forecast is carried out taking into account the existing and newly imposed trade restrictions and the expected change in final demand, which, in turn, entails a change in the output of raw materials and components. 3.2. algorithms the actions of organizations in countries, which became subject to trade restrictions, are based on a number of assumptions: supplies from countries that have imposed restrictions may be replaced by domestic supplies or supplies from other countries of products of a similar industry. purchasing organizations of various industries have equal priority in the procurement of raw materials, that is, the resulting shortage of some materials and components would be proportionally distributed among all buyers. substitution of supplies that have fallen under trade restrictions for supplies from available sources is carried out in proportion to the share of available sources in the initial distribution of supplies. that is, if the supply of products of industry j from country b accounted for 60% of the supply of this industry to country a, then in replacing the missing supplies from other countries, suppliers from country b would also account for 60% of the deficit in industry j. the algorithm below shows the sequence of recalculation of supplies under the influence of trade restrictions in accordance with the introduced assumptions: 1. calculation of permitted supplies for organizations in industry j, for products of which trade restrictions has been imposed: (3.1) where – volume of supply of organization-buyer b available from organization-supplier s of industry j in the next modeling year (y+1), – volume of supply of organization-buyer b delivered from organization-supplier s in the current modeling year y; – trade restriction on import of products of industry j, imposed in the modeling year (y+1); , – trade restriction on export of products of industry j, imposed in the modeling year (y+1). 2. calculation of lacking supplies for organizations: (3.2) where – volume of lacking supply of organization-buyer b from organizationsupplier s in the next modeling year (y+1); – volume of supply of organization-buyer b available from organization-supplier s in the next modeling year (y+1), – volume of 𝑉𝑏𝑠 𝑦+1 = 𝑉𝑏𝑠 𝑦 ∗ )1 − 𝑟𝑗 𝑖𝑚𝑝 0 ∗ (1 − 𝑟𝑗 𝑒𝑥𝑝 ) 𝑉𝑏𝑠 𝑦+1(𝑙𝑎𝑐𝑘) = 𝑉𝑏𝑠 𝑦 − 𝑉𝑏𝑠 𝑦+1 assessment of impact of trade wars on production and exports of the russian… 105 copyright ©2021 assa. adv. in systems science and appl. (2021) supply of organization-buyer b delivered from organization-supplier s in the current modeling year y. 3. grouping of lacking supplies by industry j across all organizations: (3.3) where – volume of lacking supply of organization-buyer b from organizationsupplier s in the next modeling year (y+1); – total volume of lacking supply of organizations-buyers from organizations-suppliers of industry j in the next modeling year (y+1). 4. choice of organizations for increase of supplies of industry j from countries that have not imposed trade restrictions. 5. calculation of increase of supplies of organizations from countries that have not imposed trade restrictions: (3.4) where – volume of additional supply of organization-buyer b available from organization-supplier s of industry j in the next modeling year (y+1), – total volume of lacking supply by industry j in the next modeling year (y+1); – share of organization s in supplies of products of industry j; – share of organization b in purchase of products of industry j. 6. recalculation of supplies of organizations from countries that have not imposed trade restrictions: (3.5) where – volume of supply of organization-buyer b delivered from organizationsupplier s of industry j in the next modeling year (y+1), – volume of supply of organization-buyer b delivered from organization-supplier s in the current modeling year y; – volume of additional supply. the change in trade flows between countries is also greatly influenced by the change in demand for the products of industries from organizations and consumers. to model this process, the following assumptions are made in the model: the starting point for changing the output and supplies of organizations is the final consumer demand, and the intermediate demand for materials and components is considered as a derivative of the final one. the industry structure of supply organizations is considered unchanged throughout the modeling period. intermediate supplies of organizations are divided into basic and additional, and within the framework of the sales adjustment algorithm, only basic intermediate supplies are changed. supplies are divided in such a way so that in the terminal sectors of the algorithm (agriculture and mining), all intermediate supplies would be additional. this assumption makes it possible to avoid the infinite recursion of the algorithm, since the organizations of industries that do not have basic supplies are processed at last. the algorithm below shows the sequence of recalculation of supplies due to change in final demand is carried out in accordance with the introduced assumptions: the work of the algorithm begins with assessing of the expected change in final demand on the wholesale purchases of the final products (fig. 3.3). for this, the demand for the final product is recalculated: 𝑉𝑗 𝑦+1(𝑙𝑎𝑐𝑘) = --𝑉𝑏𝑠 𝑦+1(𝑙𝑎𝑐𝑘) 𝑚 𝑠=1 𝑛 𝑏=1 𝑉𝑏𝑠 𝑦+1(𝑎𝑑𝑑) = 𝑉𝑗 𝑦+1(𝑙𝑎𝑐𝑘) ∗ 𝑃𝑠 𝑗 ∗ 𝑃𝑏 𝑗 𝑉𝑏𝑠 𝑦+1 = 𝑉𝑏𝑠 𝑦 + 𝑉𝑏𝑠 𝑦+1(𝑎𝑑𝑑) 106 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) (3.6) where – final demand for products of industry i in the next modeling year (y+1); – final demand for products of industry i in the current modeling year y; – yearly dynamics of final demand. fig. 3.3. algorithm for changes in sales due to the dynamics of final demand then there is a sequential processing of organizations supplying the final product, within which the values of their sales are adjusted due to changes in final demand for their product: (3.7) where – total sales of products of industry i in the next modeling year (y+1); – final demand for products of industry i in the next modeling year (y+1); – intermediate demand for products of industry i in the current modeling year y; – investment demand for products of industry i in the current modeling year y. the coefficient of change of intermediate supplies is calculated: (3.8) where – yearly dynamics of sales of products of industry i; – total sales of products of industry i in the next modeling year (y+1); – total sales of products of industry i in the current modeling year y. 𝐹𝐷𝑖 𝑦+1 = 𝐹𝐷𝑖 𝑦 ∗ (1 + ∆𝑑) 𝑆𝑖 𝑦+1 = 𝐹𝐷𝑖 𝑦+1 + 𝐼𝑛𝑡𝐷𝑖 𝑦 + 𝐼𝑛𝑣𝐷𝑖 𝑦 ∆𝑠𝑖 = 𝑆𝑖 𝑦+1 𝑆𝑖 𝑦 − 1 assessment of impact of trade wars on production and exports of the russian… 107 copyright ©2021 assa. adv. in systems science and appl. (2021) the values of intermediate deliveries of organizations-suppliers of final products are recalculated: (3.9) where – intermediate demand for products of industry j from organizations of industry i in the next modeling year (y+1); – intermediate demand for products of industry j from organizations of industry i in the current modeling year y; – yearly dynamics of sales of products of industry i. for each organization-supplier of intermediate products, the values of their sales are adjusted, due to the cumulative change in final and intermediate demand for their product: (3.10) where – total sales of products of industry i in the next modeling year (y+1); – final demand for products of industry i in the next modeling year (y+1); – intermediate demand for products of industry i in the next modeling year (y+1); – investment demand for products of industry i in the current modeling year y. the corresponding coefficients of change in intermediate supplies and their new values are calculated by the formulas (3.8) and (3.9). the procedure for checking supplier organizations is determined by their industry affiliation: first, organizations that produce final products (light industry), then intermediate (production of fuel, materials and chemical products), and at the end raw materials (agriculture and mining). 3.3. information support of the model initial modeling data is based on official sources presented on the portals of the russian federation federal state statistics service [19], bureau of economic analysis of the united states department of commerce [6], national bureau of statistics of china [17] and eurostat [9]. from these sources tables of gdp structure, input-output tables, structure of import and export between different countries is used. sectoral classifiers of the economy differ significantly in different countries, therefore, in the model, aggregated industries were identified, representing commodity groups that are significant for international trade: 1. agriculture & food production. 2. mining. 3. fuel production. 4. public sector. 5. chemical production. 6. production of materials. 7. production of equipment and transport. 8. light industry. 9. service. 10. trade. 11. construction table 3.1 presents industries in different countries included in the aggregated industry “production of materials”. the time periods for which production information is available also differ. the most relevant input-output tables for the eu and the united states are presented for 2019, for china and russia for 2017, therefore, for the last two countries, the tables were recalculated in proportion to the change in value added in the respective industries in 2018 and 2019. 𝐼𝑛𝑡𝐷𝑖𝑗 𝑦+1 = 𝐼𝑛𝑡𝐷𝑖𝑗 𝑦 ∗ (1 + ∆𝑠𝑖) 𝑆𝑖 𝑦+1 = 𝐹𝐷𝑖 𝑦+1 + 𝐼𝑛𝑡𝐷𝑖 𝑦+1 + 𝐼𝑛𝑣𝐷𝑖 𝑦 108 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) table 3.1. structure of aggregated industry “production of materials” aggregated industry in the model industries in the eu industries in russia industries in the usa industries in china production of materials manufacture of wood and of products of wood and cork, except furniture; manufacture of wood and of products of wood and cork, except furniture; wood products manufacture of wood and wood products, paper and paper products; publishing and printing manufacture of pulp, paper and paper products manufacture of paper and paper products paper products publishing, printing and reproduction of recorded media publishing, printing and reproduction of recorded media printing and related support activities manufacture of basic metals and fabricated metal products manufacture of basic metals manufacture of basic metals primary metals manufacture of fabricated metal products, except machinery and equipment manufacture of fabricated metal products, except machinery and equipment fabricated metal products initial data was loaded into the agent-based model of trade wars in the form of excel spreadsheets. 3.4. program realization the presented model of trade wars was programmed in microsoft visual studio using c# and postgresql as the database management system. this instrument was chosen instead of specified multi-agent simulation platform, since the model works with large arrays of data and quite a large number of agents (about 800 thousand), and programming in c# allows to reach high-performance computing. fig. 3.4 shows the software architecture of the trade wars model, which includes the model database and two main procedures: generation of model objects based on source data and simulation procedure that reproduces the dynamics of trade flows between countries in accordance with the algorithms described above. fig. 3.4. software architecture of the model generation procedure is the link from aggregated statistical data to the objects of the agent-based model, stored in the model database. within generation procedure population and organizations are created and interconnections among them are set. population in both models is reproduced taking into account gender-age structure, as discussed in [15]. one agent in the model stands for 10000 of residents in the real world, which results in 779400 agents living in 5 countries in the model (table 3.2). assessment of impact of trade wars on production and exports of the russian… 109 copyright ©2021 assa. adv. in systems science and appl. (2021) table 3.2. count of agents in countries of the model country resident count in 2020, mln agents count russia 146,75 14675 the usa 331,45 33145 the european union 447 44700 china 1411,78 141178 the rest world 5457,1 545710 total 7794 779400 organizations in the model of trade wars are aggregated: one organization in the model responds to a set of organizations of one aggregated industry in one country. thus, in the developed model, there are 55 organizations of 11 industries in 5 countries, which create and maintain more than 2000 interconnected trade flows. agents of working age are assigned to organizations through workplaces. these issues are discussed in more detail in [14]. the generated objects are stored in the model database and used for multivariate calculations using the procedure for simulation of the dynamics of trade flows between countries under the conditions of imposed restrictions, which are set as the control parameters of the model. the obtained results are also stored in the database, and are available for further analysis via the postgresql database management system interface. 4. results and discussion based on the developed software model, a series of calculations was carried out to assess the impact of trade restrictions on the output and export of organizations in countries under sanctions restrictions. in the first year of modeling, corresponding to 2020, the fall in the countries' gdp was reproduced, corresponding to the official estimates of their statistical offices. in particular, for russia, the federal state statistics service estimates the decline in gdp at 3.1% [19]. based on the algorithms of the model, a corresponding recalculation of supplies among organizations was made. thus, the forecast simulation period starts from 2021. to make forecasts of dynamics of production and commodity exchange, a medium-term period of 5 years was chosen, which is associated, first of all, with the significant uncertainty of the global epidemiological situation. in the context of this uncertainty, two scenarios are considered: an optimistic one, assuming a global economic recovery in 2021-2022, and a pessimistic one, in which a repetition of waves of coronavirus infection would lead to slowdown in the recovery of production and demand. within each scenario, three options for trade relations among countries are considered: 1. cancellation of all existing restrictions on the import and export of products. under these conditions, the dynamics of commodity exchange would be determined only by speed of economic recovery. 2. preservation of existing restrictions. currently, there are eu and us restrictions packages against russia, russian restrictions against the eu and the united states, and reciprocal sanctions between the united states and china. for russia, the imposed restrictions affect sectors of agriculture and food production (aggregated industry #1 in the model), aircraft and rocketry (aggregated industry #7) [10]. the eu, and especially the united states, restrict the construction of the nord stream 2, which affects industries of mechanical engineering, mining and fuel production (aggregated industries #2, #3 and #7 in the model). 3. imposition of new restrictions since 2022 affecting about 5% of trade volume in key industries. within the framework of scenario calculations, the model was used to compare the dynamics of russia's gdp and export volumes under various variants of world trade policy. consider the results obtained under the optimistic scenario. as the data in fig. 4.1 show, the cancellation or preservation of existing trade restrictions does not have a significant 110 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) effect on the dynamics of gdp, while imposition of new restrictions leads to slowdown in annual economic growth rates by 0.5%, and in the first two years of after imposition of restrictions, their influence is more noticeable. fig. 4.1. dynamics of russian gdp within alternative trade policy (optimistic scenario) fig. 4.2 shows the forecast for the dynamics of export volumes under the optimistic scenario. here, the comparatively weak influence of the existing restrictions is also noticeable, however, the effect of imposition of new ones is more noticeable: from 3-3.5% growth, the dynamics of exports falls to 1-2% after imposition of new restrictions. fig. 4.2. dynamics of russian exports within alternative trade policy (optimistic scenario)) in general, the results presented in fig. 4.1 and 4.2 indicate a fairly high degree of stability of the russian economy and a relatively weak impact of the current sanctions on its development. the results obtained as a result of computer calculations can also be presented in a more detailed form, for separate exporting industries (table 4.1). the obtained data demonstrates that, while maintaining the existing restrictions, the export of the industry "production of machinery and equipment" suffers insignificantly, while growth rates of products of other industries stay at the same level. after imposition of new restrictions in 2022, the export of mining industry changes the growth rate from positive (3-4%) to negative (about -1%), while fuel sales continue to grow at a slower pace (12% instead of 4%). also the export of the "manufacturing of machinery and equipment" sector shows a slowdown in growth rates by about 0.5% annually. assessment of impact of trade wars on production and exports of the russian… 111 copyright ©2021 assa. adv. in systems science and appl. (2021) table 4.1. dynamics of export by industry, percents relative to the previous year (optimistic scenario) industry year 2020 2021 2022 2023 2024 2025 cancellation of restrictions agriculture & food production -5,3 5,1 4,6 5 5,6 5,9 mining -3,9 4,7 3,5 3,6 3,6 3,7 fuel production -4,3 4,9 3,9 4,1 4,3 4,5 chemical production -3,4 3,9 3,6 3,7 3,9 4 production of materials -1,1 3 2,9 3,1 3,2 3,4 production of machinery and equipment 0 3 3,1 3,3 3,4 3,5 miscellaneous manufacturing -8,8 7,4 6 6,2 6,3 6,5 preservation of restrictions agriculture & food production -5,3 5,1 5,3 4,9 5,9 5,6 mining -3,9 4,7 3,5 3,6 3,6 3,7 fuel production -4,3 4,9 3,9 4,1 4,3 4,5 chemical production -3,4 3,9 3,6 3,7 3,9 4 production of materials -1,1 3 2,9 3,1 3,2 3,4 production of machinery and equipment 0 2,7 2,9 3 3,2 3,4 miscellaneous manufacturing -8,8 7,4 6 6,2 6,3 6,5 imposition of restrictions agriculture & food production -5,3 5,1 4,6 5 5,4 5,7 mining -3,9 4,7 -1,2 -1 -0,7 -0,3 fuel production -4,3 4,9 1,1 1,5 1,9 2,2 chemical production -3,4 3,9 3,6 3,7 3,9 4 production of materials -1,1 3 3 3,1 3,3 3,4 production of machinery and equipment 0 2,7 2,4 2,6 2,8 3,1 miscellaneous manufacturing -8,8 7,4 6,1 6,2 6,4 6,5 in the context of the pessimistic scenario, there is also no significant effect of cancellation of trade restrictions on the dynamics of gdp compared to their preservation, while imposition of new ones slows down the economic growth rate by 0.3% annually (fig. 4.3). fig. 4.3. dynamics of russian gdp within alternative trade policy (pessimistic scenario) in 2022, the preservation of restrictions leads to even greater growth in russian export volumes compared to their cancellation, which is associated with restrictions to other countries resulting in rise of orders to suppliers from russia (fig. 4.4). imposition of new restrictions in the context of the pessimistic scenario leads to a drop in export growth rates to 0.8-1.2%. 112 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4.4. dynamics of russian gdp within alternative trade policy (pessimistic scenario) calculations based on the presented agent model of trade wars for the russian federation show that the imposed sanctions affect agriculture, mining and fuel production to a greater extent, however, only the mining industry has a negative dynamic with the introduction of new restrictions (up to -1,2%). the rest of the industries maintain a positive growth rate, losing only a little in its speed (from 4,5% to -2,2% for the fuel production). as for the general economic situation, sanctions do not have a significant impact on the dynamics of value added, and the income of the population. the results obtained as a result of modeling reflect possible changes in trade turnover between countries in the sectoral context, as well as the impact of these changes on the overall economic indicators of countries, in particular, gdp. this structure of the modeling output data coincides with the structure of the results of the models based on the gtap methodology. the results obtained are directly related to the assumptions of the model, the most important of which are the following: we consider goods from one industry are interchangeable and do not take into account preferences of the customers to goods from a particular country. such assumptions are widespread in economic models [3,4]. investment in fixed assets is considered elastic, directly proportional to the change in output, which makes it possible to neglect the delays associated with expansion of production. in the calculations performed, the change in trade volume in prices of the base year is estimated, excluding inflation and fluctuations in the exchange rates. direct bans on the import of certain categories of products are considered as sanctions measures, and the increase in customs tariffs that affect the price of products for buyers is not taken into account. the assumptions made generally exclude price competition as a factor influencing international trade, and shift the main emphasis to the fundamental availability of supplies of products of certain industries in the required quantity. the assumption of the interchangeability of goods from different countries makes it possible not to take into account such a problem as the complete unavailability of raw materials and materials of a certain type. thus, the model does not assess the effects of trade wars, such as reduced output or production interruptions due to lack of required materials. the assumption that there are no delays associated with the expansion of production is supposed to be removed in subsequent works to take into account the impact of investment processes in different countries on international trade. assessment of impact of trade wars on production and exports of the russian… 113 copyright ©2021 assa. adv. in systems science and appl. (2021) 5. conclusion this article examined issues of developing a model of trade wars, including as the main participants russia, the european union, the united states and china, as well as the united rest of the world. for the development, an agent-based approach was chosen, which distinguishes the presented model from known analogs based on the principles of computational general equilibrium. the chosen method makes it possible to study international trade relations at the micro level, in the context of supplies of individual organizations, purchases of residents and state procurements. the article presents a number of results obtained, among which the most significant are: 1. the event structure of the model of trade wars, reflecting interaction of state, public sector organizations, manufacturers, trade agents and residents. 2. the algorithm reflecting annual change in the volume of sales and supplies of organizations, which, in turn, determines change in the volume of imports, exports and added value of industries. 3. the sectoral structure of the model, leading to a general view of sectoral classifiers in different countries. 4. parametrization of scenario calculations, including two scenarios for the restoration of the world economy in the context of epidemiological risks (optimistic and pessimistic) and three options for world trade policy (preservation, cancellation or imposition of new trade restrictions). 5. forecasts of the dynamics of gdp, volumes and sectoral structure of russian exports in the context of various scenarios and a comparative assessment of the consequences of trade restrictions. as a direction for further research in this area, a more detailed analysis of the impact of investing in expanding production on the consequences of trade restrictions for different countries was chosen. acknowledgements the reported study was funded by russian science foundation according to the research project № 21-18-00136 “development of a software and analytical complex for assessing the consequences of intercountry trade wars with an application for functioning in the system of distributed situational centers in russia”. references 1. aguiar, a., narayanan, b., mcdougall, r. (2016). an overview of the gtap 9 data base. journal of global economic analysis, 1, 181-208. https://doi.org/10.21642/jgea.010103af 2. alvarez, f. & lucas, r. (2005). general equilibrium analysis of the eaton–kortum model of international trade. journal of monetary economics, 54(6), 1726-1768. https://doi.org/10.1016/j.jmoneco.2006.07.006 3. baier, s.l. & bergstrand, j.h. (2001). the growth of world trade: tariffs, transport costs, and income similarity. journal of international economics, 53 (1), 1–27. https://doi.org/10.1016/s0022-1996(00)00060-x 4. bollen, j. & rojas-romagosa h. (2018, july). trade wars: economic impacts of us tariff increases and retaliation: an international perspective. cpb background document. cpb netherlands bureau for economic policy analysis. [online]. available https://www.cpb.nl/sites/default/files/omnidownload/cpb-background-documentnovember2018-trade-wars-update.pdf 114 a. mashkova, a. bakhtizin copyright ©2021 assa. adv. in systems science and appl. (2021) 5. bouët, a. & laborde, d. (2017, august). us trade wars with emerging countries in the 21st century. make america and its partners lose again. ifpri discussion paper. [online]. available http://ebrary.ifpri.org/cdm/ref/collection/p15738coll2/id/131368. 6. bureau of economic analysis of the united states department of commerce. [online]. available https://www.bea.gov/. 7. corong, e., hertel, t., mcdougall, r., tsigas, m., van der mensbrugghe, d. (2017). the standard gtap model, version 7. journal of global economic analysis 2, 1-119. https://doi.org/10.21642/jgea.030101af 8. dixon, p., jerie, m., rimmer, m. (2016). modern trade theory for cge modelling: the armington, krugman and melitz models. journal of global economic analysis 1(1),1– 110. https://doi.org/10.21642/jgea.010101af 9. eurostat. [online]. available https://ec.europa.eu/eurostat. 10. federal law no. 127-fz. (2018, june 4) on measures of influence (counteraction) on unfriendly actions of the united states of america and other foreign states. [online]. available http://publication.pravo.gov.ru/document/view/0001201806040032. 11. glick, r. & taylor, a. (2005). collateral damage: trade disruption and the economic impact of war. nber working paper series. working paper no. 11565. [online]. available http://www.nber.org/papers/w11565 12. hervé, k., pain, n., richardson, p., sédillot, f., beffy, p. (2010). the oecd’s new global model. economic modelling, 28, 589-601 (2011). https://doi.org/10.1016/j.econmod.2010.06.012 13. li, c., he, c., lin, c. (2018). economic impacts of the possible china-us trade war. emerging markets finance and trade, 54, 1557-1577. https://doi.org/10.1080/1540496x.2018.1446131 14. mashkova, a.l., nevolin, i.v., savina, o.a., burilina, m.a. & mashkov, e.a. (2020). generating social environment for agent-based models of computational economy. communications in computer and information science, 1349, 291-305. https://doi.org/10.1007/978-3-030-67238-6_21 15. mashkova, a.l., novikova, e.v., savina, o.a. & mashkov, e.a. (2021). generating synthetic population for the agent-based model of the russian federation spatial development. advances in social simulation (essa 2019). springer proceedings in complexity. 1349, 183-187. springer nature switzerland. https://doi.org/10.1007/978-3030-61503-1_17 16. mcdonald, s., thierfelder, k.: globe v2: a sam based global cge model using gtap data. [online]. available http://cgemod.org.uk/globev2_2014.pdf 17. national bureau of statistics of china. [online]. available http://www.stats.gov.cn/english/. 18. ossa, r. (2014). trade wars and trade talks with data. american economic review 104(12), 4104–4146. doi: 10.1257/aer.104.12.4104. 19. russian federation federal state statistics service. [online]. available http://www.gks.ru. 20. vitek, f. (2018, april). the global macrofinancial model. international monetary fund working paper, 18/81. [online]. available https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3182506. adv syst sci appl 2020; 01; 114-118 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/864 a data management model for proactive risk management in healthcare dmitry meshkov, lyudmila bezmelnitsyna, sergey cherkasov v.a. trapeznikov institute of control sciences, russian academy of sciences, russia e-mail: dmitrymeshkov@mail.ru, blyu18@gmail.com, cherkasovsn@mail.ru abstract: the delphi investigation made possible to create and present a model of data flows and expertise for decision making in healthcare. the model structures and visualizes data flows including population health risks assessment and healthcare interventions value. most elements of this model are used in practice and are regulated by international or local documents. the investigation also indicates the need for the creation and utilization of additional level of expertise including integrated analysis and modeling of the results of management decisions. structuring of data management in accordance with this model provides an opportunity to implement proactive risk management in healthcare there to increase healthcare effectiveness significantly. keywords: model, proactive, healthcare, management, data, health technology assessment, guidelines, clinical. 1. introduction decision making and management in healthcare always involve the redistribution of budgets and existing major healthcare models are traditionally based on the description of financial flows [4]. the emergence of innovative technologies and their costs not only affected the quantitative changes in financial flows, but also the change in information quality, quantity and distribution in healthcare decision-making. creation in majority of countries health technology assessment agencies responsible for expertise of clinical economic characteristics of healthcare interventions indicates the need for novel supportive models for healthcare decision making based on development and rational use of appropriate methods of prophylactics, diagnostics and treatment [7]. it is evident now that achieving both the expected results in population health as well as in rational spending of funds requires appropriate conditions for use of the intervention within healthcare system such as decrease of accessibility barriers (technological; administrative; and psychological); availability and affordability. our early research demonstrated correlations between the number of new drugs (inn) available in the market and the reduction of patient mortality in the example of cancer. this effect was more probably induced by additional opportunities and proper healthcare management rather than the increase of clinical efficacy of these drugs [5]. these results indicate the need for a novel model of healthcare management formed around population health risks and medical technologies, which meets modern requirements and would make it possible to increase the efficiency of decision-making in healthcare independently or when used in conjunction with other models. 2. objectives the objective of the study was to identify the structural model for proactive healthcare management corresponding to the modern challenges of population health and health mailto:dmitrymeshkov@mail.ru mailto:blyu18@gmail.com mailto:cherkasovsn@mail.ru a data management model for proactive risk management in healthcare 115 copyright ©2020 assa adv. in systems science and appl. (2020) technologies and applicable to any healthcare system regardless of funding sources and health care model as well. oncology was chosen for the case study since this area is of great importance for population health and a significant number of new technologies allowed us to clearly identify the challenges and the processes occurring in this area [5]. 3. method modified delphi investigation was used to determine the nodal elements of the model and their hierarchical relationships [3]. the group of analysts included 7 people. all of them had experience of practical work as doctors and healthcare management as well. two of them had also an economic high education and experience in solving economic issues in healthcare. the other two experts had experience of data generation of medicinal products starting from laboratory investigations to healtheconomic assessments and inclusion of these products in reimbursement lists. the group of experts included 64 physicians specialized in oncology who had also practical experience in management of oncology units and hospitals (finance/management/staff load). they were presented in databases of clinicians participating in international industry sponsored clinical trials and thus were well aware about worldwide experience of oncology diseases treatment. the study was conducted in several stages: 1. analyze of the source data and creation a questionnaire for the expert group to fill in. 2. conducting a survey of the expert group. 3. analysis of the responses received and development of proactive risk-management model in healthcare. the questionnaire consists of 33 questions covering different aspects of changes during last 10 years (to link with assessment of statistical data provided by national oncology register) of factors influencing accessibility, efficacy, efficiency, effectiveness and safety of medical care and interventions to oncology patients and to express the impact of each factor in points from 1 to 5. overall, each expert had to respond 462 questions covering 14 groups of oncology diseases according to icd-10 who fic. 4. results the proposed structural model for assessing and proactively risks managing in healthcare is presented at fig.1. it consists of node elements or blocks most of which have already been developed and applied in practice for expertise and decision making in healthcare. the model is based on world health organization (who) recommendations and country commitments to use the who family of international classifications (who fic). the who fic contains three basic classifications making possible to form a global terminology base for discrete description of the presence or absence of diseases (international classification of diseases (icd), degrees of functional body or organ impairment (international classification of functioning (icf) and international classification of healthcare interventions (ichi). these classifications allow to achieve comparability of data between different geographical regions and countries, as well as to observe the comparability of data over time, determine trends and make a forecast of population health. fic makes possible to monitor population healthcare and healthcare technologies as well as making forecast in both areas [5]. these classifications allow us to achieve uniform representation of diverse data and to build predictive models according to format acceptable for regulatory decision making. next level of model is the assessment of disease burden summarizing the complex of medical, social and economic consequences of a particular disease. it is the basis for decision 116 d. meshkov, l. bezmelnitsyna, s. cherkasov copyright ©2020 assa adv. in systems science and appl. (2020) on the need for medical technologies to reduce the impact of the disease and for the successive data generation on the value of this medical technology in the treatment of the disease [1,6]. next sequential levels include data generation on the value of healthcare intervention in laboratory conditions (“preclinical studies”) and studies with patients participation (“clinical studies”) assessing the likelihood of an outcome, systematic reviews, and clinical guidelines (in accordance with who recommendations) that summarize all available information on diseases and their treatment options. clinical guidelines are systematically developed provisions designed to help a doctor make decisions about medical tactics and the use of appropriate medical technologies in certain clinical situations. in clinical recommendations, there is a link between each statement and scientific data, and scientific facts take precedence over expert opinion. these documents do not have formal legal force, but are a tool that helps doctors make optimal therapeutic choices. the use of clinical recommendations developed on the basis of international clinical studies in specific economic conditions and the health system is the task of the government and health managers, and is carried out with the help of other formalized documents that have legal force. the development of a prognosis focused on the needs of practical health care will also provide a timely evidence base for the creation of clinical recommendations that meet the requirements of who and serve as the basis for treatment standards and procedures for providing medical care. [2]. the obtained information is used for the “integrated assessment and scenario models” developing with the purpose of the issue resolution, as well as for predicting of medical technologies development (“horizon scanning”) [8]. it is obvious that market authorization of the new intervention and overcoming barriers requires time which delay the availability of medicines n market and decrease the effectiveness of healthcare system. the earlier the preparation for introduction of new interventions begins the better result can be achieved on population level. horizon scanning deliver data for clinical and economic evaluation of drugs (“health technology assessment”) and for the “integrated assessment and scenario models” block which summarizes and present information for administrative solutions. it turned out that this last element of the model is practically not represented in the process of forming an expert assessment and making a decision while that the requirements for the final expert report depend on the requests of the regulatory authorities and the tasks they face. local legislation and formal regulatory procedures in each country form the landscape influencing on regulatory solutions which must lead to appropriate proactive risk management aimed at eliminating the predicted population health issues by accelerated development of appropriate healthcare structure and regulations as well as introduction of medicines into healthcare practice. proper healthcare management is also linked to timely provision of information about population health threats and healthcare interventions value by physicians education at high medical schools and within in the framework of continuous medical training and professional development. providing information to specialists and relationships with medical schools will allow them to prepare specialists beforehand, taking into account the forecast of the dynamics of public health indicators and making them familiar with innovative healthcare interventions aimed at improving these indicators. a data management model for proactive risk management in healthcare 117 copyright ©2020 assa adv. in systems science and appl. (2020) fig.1. node elements and data flow of the proactive risk management healthcare model. regulatory control according to international requirements world health organization family of international classification (who fic) international statistical classification of diseases and related health problems (icd) international classification of functioning disability and health (icf) international classification of healthcare interventions (ichi) population health monitoring healthcare interventions monitoring preclinical studies clinical studies systematic reviews and clinical guidelines health technology assessment regulatory control according to local legislation medical education and training regulations and procedures healthcare organization and management population health forecast horizon scanning disease burden integrated assessment & scenario models 118 d. meshkov, l. bezmelnitsyna, s. cherkasov copyright ©2020 assa adv. in systems science and appl. (2020) 5. discussion and conclusions the model of data generation and distribution for proactive healthcare management is presented. a comprehensive evidence-based approach linked to the population health forecast provides an opportunity to improve the quality of healthcare management as well as the effectiveness of medical care. most elements of the presented model are already existing and are used in practice. there are international or local documents regulating data proceeding for each blocks and data submission to the regulatory authorities for making management decisions. the survey indicated that the only block “integrated assessment & scenario models” is not used in practice. regulatory decisions are made on separate assessment of clinical efficacy and safety (market authorization), clinical guidelines (education and training, regulatory procedures) and economic effectiveness (inclusion into reimbursement lists) while a comprehensive assessment of these data and modeling of outcomes under different scenarios would be the most effective solution of expert assessment and regulatory decisions in healthcare continuity. references 1. cutler rl, fernandez-llimos f, frommer m, benrimoj c, garcia-cardenas v (2018) economic impact of medication non-adherence by disease groups: a systematic review. bmj open. 8(1), 1-13. doi: 10.1136/bmjopen-2017-016982. 2. guidelines for who guidelines. who practice guidelines: recommended process. global programme on evidence for health policy (2003). world health organization. geneva, switzerland eip/gpe/eqc/2003,1. 1-24 available: http://apps.who.int/iris/bitstream/10665/75146/1/9789241548441_eng.pdf 3. jandhyala r. (2020) a novel method for observing proportional group awareness and consensus of items arising from list-generating questioning. curr med res opin. mar 11:1-11. doi: 10.1080/03007995.2020.1734920. available: https://www.tandfonline.com/doi/full/10.1080/03007995.2020.1734920 4. katarina bohm, achim schmid, ralf gotze, claudia landwehr, heinz rothgang (2013) five types of oecd healthcare systems: empirical results of deductive classification/ health policy 113(2013), 258-269 5. khabriev r.u., isaeva a.v., bezmenitsyna l.yu., meshkov d.o., berseneva e.a., cherkasov s.n. (2016) dostupnost innovacionny`x lekarstvenny`x preparatov i snizhenie pokazatelej smertnosti ot onkologicheskix zabolevanij [the availability of innovative drugs and a reduction in cancer mortality rates] bulletin of the national public health research institute, 6, 5-19 [in russian] 6. mcdonald sa, nijsten d, bollaerts k, bauwens j, praet n, et. al. (2018) methodology for computing the burden of disease of adverse events following immunization. pharmacoepidemiol drug saf. 2018 jul; 27(7):724-730. doi: 10.1002/pds.4419 7. proposal for a regulation of the european parliament and of the council on health technology assessment and amending directive (2011) 2011/24/eu. brussels, 31.1.2018 com (2018) 51 final 2018/0018 (cod); 1-48 8. the international information network on new or emerging, appropriate use and reassessment needed health technologies (2020) [online] available: https://www.euroscan.org 9. world health organization. classifications (2020) [online] available: https://www.who.int/classifications/en/ https://www.tandfonline.com/doi/full/10.1080/03007995.2020.1734920 https://www.euroscan.org/ https://www.who.int/classifications/en/ adv syst sci appl 2020; 04:83–104 published online at https://ijassa.ipu.ru. on optimal therapy protocols in the mathematical model of prostate cancer progression marie-christin litzinger1, y. todorov1, miriam föller-nord1, manitjaiswal k. chaudhary2, alexander s. bratus3,4∗ 1university of applied sciences, paul-wittsack-str. 10, 68163, mannheim, germany 2nikolsky mathematical institute, rudn university, miklukho-maklaya str.6, moscow, 117198, russia 3russian university of transport, obraztsova 15, moscow, 127994, russia 4moscow center of fundamental and applied mathematics, msu, gsp-1, leninskie gory, moscow, 119991, russia abstract: a mathematical model for the treatment of prostate cancer regarding androgen ablation is considered. the effect of the therapy on healthy cells is taken into account and modelled as a second order state constraint. a mathematical overview about possible therapy strategies is presented. keywords: prostate cancer, androgen ablation, therapy strategies, optimal control, state constraint 1. introduction regarding prostate cancer progression and development testosterone and its role in the biological mechanisms of epithelial prostate cells are of particular interest for developing therapy strategies. therefore the main steps of testosterone procession are now outlined briefly. after entering the cells testosterone is either able to bind to androgen receptors (ar) straight away or it is first converted to dihydrotestosterone (dht). in contrast to testosterone, dht and ar form a more stable complex. after being phosphorylated and dimerized that complex binds to the dna, where it causes a higher transcription rate of with proliferation, survival and differentiation associated genes [11][10]. consequently, there are several points in the mechanism where the therapy can intervene, for example through inhibiting the testosterone production or blocking androgen receptors [2]. after some time prostate cancer cells are able to grow even in the absence of androgen. due to the following therapy resistance an intermittent androgen therapy can be applied [9, 19]. it is important that healthy and cancer cells both produce prostate-specific antigene (psa). in noncancerous cells only a limited amount of psa enters into the bloodstream due to natural barriers. in the presence of cancer these barriers break down and more psa leaks into the blood. the psa-level in the blood serum therefore acts as an indicator for the stadium and the presence of prostate cancer [11]. due to the complex mechanisms of cancer the cooperation of different scientific areas is necessary and mentioned in [3] and [6]. since the last fifty years the application of ∗corresponding author: alexander.bratus@yandex.ru 84 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus mathematical approaches in order to analyze the dynamics of cancer is therefore common and discussed in [3]. furthermore mathematical models and methods for the analysis of the dynamics of tumor-immune interaction are considered to evolve research perspectives [4]. however, not only the competition between cancer and healthy cells, but also the toxicity of the therapy, which damages healthy cells, are negative side effects. these are essential for developing effective medicine and therapy strategies. consequently, the mathematical theory of optimal control processes is needed to develop successful therapy protocols [17][22]. due to the adverse effects caused by the ablation of androgen, for example the decrease of quality of life [5], it is convenient to introduce a limitation of the healthy epithelial prostate cells during androgen ablation therapy. in the present work three approaches – an alternative, suboptimal and optimal one – in the mathematical model for prostate cancer progression in response to androgen ablation therapy are introduced. in all three approaches the side effects of the androgen ablation therapy are considered as a second order pure state constraint on the amount of the healthy epithelial prostate cells. the goal is to minimize the amount of cancer cells while keeping the amount of healthy cells above a given critical limit. the resistance of prostate cancer cells during the androgen ablation therapy is also taken into account. in the following a brief description of each therapy strategy is given. the alternative therapy strategy is based on the dynamical analysis of the system and its critical points. this approach circumvents the classical optimal control theory. in combination with a numerical improvement algorithm it provides feasible results which are similar to the optimal one. pontraygin’s maximum principle without state constraints is used to develop a therapy protocol at which the conditions of the states on the constraint boundary are analyzed. the considered approach always guarantees a feasible therapy process if there exists such possible one. the results are used to obtain the so called suboptimal therapy strategy. an analytical determination of an optimal therapy strategy in the presence of state constraints of order two is considered by means of analytical and numerical techniques. therefore the software tool bocop is used. the results of the presented approaches are compared and discussed. 2. mathematical model of prostate cancer in this article a mathematical model of prostate cancer based on [11] is considered. copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 85 de(t) dt = µee(t) ( 1− e(t) ηe ) − εe(t)(m(t) +n(t)) ηe − δee(t) dn(t) dt = (1− αm)µnn(t) ( 1− n(t) +m(t) η ) − δnn(t) dm(t) dt = (µmm(t) + αmµmn(t)) ( 1− n(t) +m(t) η ) − δmm(t) dr(t) dt = αr − λrr(t)− kdf r(t)d(t) + kdr a(t)− ktf r(t)t (t) + ktr a(t) dd(t) dt = βt t (t) kt + t (t) − λdd(t)− kdf r(t)d(t) + kdr a(t) dt (t) dt = f(ts)− λtt (t)− βt (t) t (t) kt + t (t) − ktf r(t)t (t) + ktr at(t) da(t) dt = −λaa(t) + kdf r(t)d(t)− kdr a(t) dat(t) dt = −λatat(t) + ktf r(t)t (t)− ktr at(t) dp (t) dt = βp(a(t) + at(t) + ϕ)− αpp (t)efrac − γpp (t) (m(t) +n(t))2 kp + (m(t) +n(t)) − λpp (t) t ∈ [0, t ], e(0), n(0), m(0), r(0), d(0), t (0), a(0), at(0), p (0) ∈ r+, (2.1) where efrac = vc e pv ol , pv ol = αvt dv bdvv + tdv . in (2.1) it is assumed that the healthy cells e have the proliferation rate µe , the death rate δe and a limit ηe . the magnitude of the competition between healthy and cancer cells is modelled by means of the parameter ε. the androgen-dependent cancer cells n proliferate at rate µn , decease with rate δn and have the limit ηn . with the probability αm ∈ (0, 1) cancer cells n mutate to castrationresistant cancer cells m where the parameters µm and δm represent the proliferation and dead rate respectively. because of the character of this article the remaining state variables of the dynamic system (2.1) which represent specific hormone levels in the organism are only mentioned: t testosterone concentration, p tissue psa concentration, r free androgen receptor concentration, d dihydrotestosterone concentration, a dihydrotestosterone-activated androgen receptor concentration,at testosterone-activated androgen receptor concentration. in this article all cell types are estimated in billions. detailed information, the biological background of the variables and the associated parameters are given in [?]. assuming µn = µm = µ, δn = δm = δ, n(t) and m(t) can be substituted in order to represent the total number of the prostate cancer cells. from the addition n(t) +m(t) = z(t) follows dz(t) dt = µz(t) ( 1− z(t) η ) − δz(t). (2.2) the following figures illustrate the behaviour of healthy and cancer cells as well as the psa-level in the absence of therapy. the parameters of the model which were used in this copyright c© 2020 assa. adv syst sci appl (2020) 86 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus article describe a scenario in which the amount of healthy cells is not strongly influenced by cancer cells. the set of parameters and their description is given (see a.1). 0 50 100 150 non-dimensional time 0 50 100 150 200 250 300 fig. 2.1. time response of healthy e(t) and cancer cells z(t) 0 50 100 150 non-dimensional time 0 1 2 3 4 5 6 7 8 9 10 fig. 2.2. time response of the psa function p (t) copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 87 as mentioned before intermittent therapy is one way to slow down the development of therapy resistance. the following figures show the dynamics of the system when such therapy is applied. 0 50 100 150 non-dimensional time 0 1 2 3 4 5 fig. 2.3. amount of the administered drug u(t) 0 50 100 150 non-dimensional time 0 50 100 150 200 250 300 fig. 2.4. time response of healthy e(t) and cancer cells z(t) copyright c© 2020 assa. adv syst sci appl (2020) 88 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus 0 50 100 150 non-dimensional time 0 1 2 3 4 5 6 7 8 9 10 fig. 2.5. time response of the psa function p (t) in the following sections different therapy strategies are determined and used to minimize cancer cells without violating the constraint boundary on healthy cells. 3. an alternative therapy strategy consider the optimal control problem after the substitution (2.2) min u j = ∫ t 0 z(t)dt→ min (3.3) copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 89 de(t) dt = µee(t) ( 1− e(t) ηe ) − εe(t)(m(t) +n(t)) ηe − δee(t) dz(t) dt = µ(h)z(t) ( 1− z(t) η ) − δz(t) dr(t) dt = αr − λrr(t)− kdf r(t)d(t) + kdr a(t)− ktf r(t)t (t) + ktr a(t) dd(t) dt = βt t (t) kt + t (t) − λdd(t)− kdf r(t)d(t) + kdr a(t) dt (t) dt = f(ts)− λtt (t)− βt (t) t (t) kt + t (t) − ktf r(t)t (t) + ktr at(t) da(t) dt = −λaa(t) + kdf r(t)d(t)− kdr a(t) dat(t) dt = −λatat(t) + ktf r(t)t (t)− ktr at(t) dp (t) dt = βp(a(t) + at(t) + ϕ)− αpp (t)efrac − γpp (t) (m(t) +n(t))2 kp + (m(t) +n(t)) − λpp (t) dh(t) dt = −γh(t) + u(t) 0 ≤ u(t) ≤ umax, t ∈ [0, t ], t -fixed e(0), z(0), r(0), d(0), t (0), a(0), at(0), p (0), h(0) ∈ r+, (3.4) and the state constraint on the amount of healthy cells g(t) = e(t)− ec ≥ 0 ∀t ∈ [0, t ]. (3.5) the last equation in (3.4) describes the pharmacokinetic and the functions µ(h) = µ0 ( 1− kh h+ 1 ) and µe(h) = µ0 ( 1− keh h+ 1 ) the pharamcodynamic of the administered medicine. the parameters γ, k, ke, µ0, µ0 e ∈ r+, where γ stands for the dissipation rate of the administered drug respectively, k and ke describe the effect of the drug on cancer and healthy cells and µ0 and µ0 e are the replication rates of the two cell populations. the equation (3.5) has to guarantee that the amount of the healthy cells does not undercut a given critical limit ec during the therapy process. because of the mathematical complexity of the optimal control problem (3.3)-(3.5) an alternative approach, which will further be called alternative therapy strategy, is considered at first. in order to analyze whether small changes of the initial conditions of the considered system lead to a different behaviour of the corresponding trajectories its critical points are analyzed (see a.2). due to the practical irrelevance of the trivial critical point (ē, z̄, h̄) = (0, 0) only z̄ = ( 1− δ µ(h) ) η, ē = ηe − ε(µ(h)− δ)η µe(h)µ(h) − δeηe µe(h) , h̄ = u γ (3.6) is taken into account. from the dynamical analysis of the considered system (see a.3) it is obvious, that for realistic parameters of the model the critical point (3.6) is a stable node. copyright c© 2020 assa. adv syst sci appl (2020) 90 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus assuming now that u is a constant parameter such that 0 ≤ u ≤ umax we consider the objective function j̃(u) := j̃(ē, z̄) = z̄ + κ(ē − ec)2 → min, (3.7) where with the scale of the weighting factor κ ∈ r+ the state constraint on the amount of healthy cells is taken into account. it follows from the implicit function theorem that ē and z̄ can be considered as functions of u and consequently the minimization of the functional (3.7) can be approached as a mathematical programming problem. due to the theorem of weierstrass a solution of this problem always exists. assuming that u = ū ∈ [0, umax] is the value for which the minimum of the functional (3.7) is achieved. now consider the function u(t) = umax, 0 ≤ t ≤ t̂. here umax and t̂ are two constants such that the solution of the last equation in (3.4) with initial condition h(0) = 0 reaches the value h = ū γ at the moment t = t̂. since h(t) = umax γh (1− e−γt) (3.8) the condition h(t̂) = ū γ is fulfilled iff t̂ = 1 γ ln ( 1− ū umax ) . (3.9) the last equality provides the required time of the function h(t) to reach the value of ū γ if the control function u(t) = umax in the time interval 0 ≤ t ≤ t̂ , which is called intensive therapy time. in the remaining time interval t̂ < t ≤ t the control function u(t) has to keep the value ū γ which minimizes the objective function j̃(u). this therapy interval time is called relaxation therapy time. the determined strategy is called alternative therapy strategy and can be defined as u(t) = umax, 0 ≤ t ≤ t̂, ū γ , t̂ < t ≤ t. (3.10) according to the considerations above the alternative control strategy consists out of two stages: the stage of the intensive therapy and the stage of relaxation. in order to avoid disadvantages in the approach considered above, the following modification, presented in [22], will be used. this requires the introduction of the shifting variable s ∈ r+ in the penalty term of the objective function (3.7) j̃(u) := j̃(ē, z̄) = z̄ + κ(ē − (ec + s))2 → min. (3.11) in the numerical procedure some iterations are usually necessary to obtain a good result and the new value of the shifting variable after every iteration has to be calculated as follows: s(q+1) = s(q) + max{ec − e(t)| e(t) < ec, t ∈ (0, t )}, where q is the number of the last iteration. if the maximal number of iteration qmax or the desired accuracy θ ≥ |s(q) − s(q−1)| is reached the procedure has to be terminated. the following numerical results illustrate the time response of healthy and cancer cells as well as the psa-level and the control function using the alternative therapy strategy. copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 91 0 50 100 150 non-dimensional time 44 46 48 50 52 54 56 58 fig. 3.6. time response of healthy cells e(t) using the alternative therapy strategy 0 50 100 150 non-dimensional time 115 120 125 130 135 140 145 150 fig. 3.7. time response of cancer cells z(t) using the alternative therapy strategy cancer and healthy cells decrease (fig. 3.6), (fig. 3.7). healthy cells reach the constraint boundary without violating it. copyright c© 2020 assa. adv syst sci appl (2020) 92 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus 0 50 100 150 non-dimensional time 0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 fig. 3.8. amount of the administered drug u(t) using the alternative therapy strategy 0 50 100 150 non-dimensional time 0 1 2 3 4 5 6 7 8 9 10 fig. 3.9. time response of the psa function p (t) using the alternative therapy strategy the maximal dose of therapy is used for a short amount of time and then a constant amount is administered (fig. 3.8). due to the reduction of cancer cells less psa is produced in the tissue and can consequently be found in the blood serum (fig. 3.9). copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 93 4. suboptimal therapy strategy in this section a suboptimal therapy strategy for the problem (3.3)-(3.5) is introduced. in order to minimize the objective function (3.3) it is natural to apply the maximal admissible amount of the drug so long as possible. however, it has to be guaranteed that the amount of healthy cells does not violate the constraint (3.5). i.e. that the time derivation ė on the constraint boundary has to be non negative. from ė(tj) ≥ 0 follows µe(h)ec ( 1− ec ηe ) − εecz(tj) ηe − δeec ≥ 0 ⇒ µ0 e ( 1− keh(tj) h(tj) + 1 ) ≥ εz(tj) + δeηe ηe − ec ⇒ keh(tj) ≤ ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) (h(tj) + 1) (4.12) ⇒ h(tj) ( ke − ( 1− εz(tj) + δeηe (ηe − ec)µ0 e )) ≤ ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) ⇒ h(tj) ≤ ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) ke − ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) , where tj := {t ∈ (0, t )| e(t) = ec}. due to the continuity of e the strong inequality in the last expression can not appear. the last expression in (4.12) gives the necessary condition for h on the constraint boundary, which guarantees an admissible process. therefore the following analysis is needed. from the last equation of the system (3.4) holds for u(t) = umax h(tj) = e−γtj ∫ tj 0 eγtumaxdt = umax γ ( 1− e−γtj ) . (4.13) thus if the maximal amount of the drug, u(t) = umax, t ∈ [0, tj), is administered until the critical boundary of the healthy cells is reached, the following relation has to be fulfilled ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) ke − ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) = umax γ (1− e−γtj). (4.14) note, that this is a special case which can only appear with a given set of parameters of the model and initial conditions. in this case the therapy strategy can be changed when the amount of the healthy cells reaches the critical value without violating the therapy process. the therapy strategy stated below is in this case optimal (see a.4). copyright c© 2020 assa. adv syst sci appl (2020) 94 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus u(t) =  umax, 0 ≤ t ≤ tj, γ 1− εz(t) + δeηe (ηe − ec)µ0 e  ke− 1− εz(t) + δeηe (ηe − ec)µ0 e  , tj < t ≤ t. (4.15) in general the change of the administered amount of the drug can not induce the change of the slope of the development of the healthy cells immediately. this is caused by the inertia of the dynamic system. therefore investigations of the more realistic cases have to be done. the case in which the maximal admissible amount of the drug has to be changed before the critical boundary of the healthy cells is reached, is considered. i.e.,( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) ke − ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) < umax γ (1− e−γtj). (4.16) the admissible process can be guaranteed if the maximal admissible amount of the drug is reduced at the point of time t′ < tj so that h(t′) = ( 1− εz(tj(umax)) + δeηe (ηe − ec)µ0 e ) ke − ( 1− εz(tj(umax)) + δeηe (ηe − ec)µ0 e ) holds, where z(tj(umax)) is the amount of the cancer cells on the constraint boundary using the maximal dose of medicine (u(t) = umax, t ∈ [0, tj]). thus the suboptimal therapy strategy can be given by u(t) =  umax, 0 ≤ t ≤ t′, γ 1− εz(tj(umax)) + δeηe (ηe − e(t))µ0 e  ke− 1− εz(tj(umax)) + δeηe (ηe − e(t))µ0 e  , t′ < t ≤ t. (4.17) the following figures present the results of the suboptimal control approach. cancer and healthy cells decrease and the healthy cells do not reach the constraint boundary (fig. 4.10), (fig. 4.11). since cancer cells are decreasing continuously the quantity of psa in the tissue and therefore in the blood serum also decreases (fig. 4.13). 5. optimal therapy strategy in this section the therapy strategy related to the optimal control problem (3.3)-(3.5) is investigated with pontryagin’s maximum principle, which is the necessary condition for an optimal process. the restriction on the minimal admissible amount of healthy cells (3.5) has to be interpreted as a second order state constraint. because of the high complexity of optimal copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 95 0 50 100 150 non-dimensional time 44 46 48 50 52 54 56 58 fig. 4.10. time response of healthy cells e(t) using the suboptimal therapy strategy 0 50 100 150 non-dimensional time 120 125 130 135 140 145 150 fig. 4.11. time response of cancer cells z(t) using the suboptimal therapy strategy control problems with higher order state constraints the introduced approach is a composition of analytical investigations and numerical results. copyright c© 2020 assa. adv syst sci appl (2020) 96 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus 0 50 100 150 non-dimensional time 0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 fig. 4.12. time response of the quantity of the u(t) using the suboptimal therapy strategy 0 50 100 150 non-dimensional time 0 1 2 3 4 5 6 7 8 9 10 fig. 4.13. time response of the psa function p (t) using the suboptimal therapy strategy because of the nature of the maximum principle it is essential to primarily investigate the optimal therapy strategy without state constraints. the optimality of this process till the state variable e reaches the constraint boundary is determined. copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 97 the investigations given in a.3 prove the following theorem. theorem 5.1: the optimal therapy strategy in the absence of the state constraint (3.5) is given as u∗(t) = umax ∀t ∈ [0, t ]. from the practical point of view this result is equivalent to the situation in which the test results of a patient are acceptable during the therapy. therefore the therapy has to be executed with the maximal admissible amount of the drug in order to eradicate the cancer cells. in the opposite case the amount of healthy cells e reaches the critical value ec. because of the complexity the strict mathematical investigation of the constraint problem is not considered. instead of that numerical results using the software tool bocop are presented. the numerical results show the development of healthy cells, cancer cells and the optimal control. 0 50 100 150 non-dimensional time 40 42 44 46 48 50 52 54 56 58 60 fig. 5.14. time response of healthy cells e(t) using the optimal therapy strategy healthy and cancer cells decrease (fig. 5.14), (fig. 5.15). healthy cellse(t) then progress along the boundary. after a period where the maximal amount of drug is administered follows a short relaxation phase. then the maximal admissible amount of medicine which minimizes the objective function is applied. the following table presents the values of the objective function, the healthy cells and cancer cells at the end of the therapy process (t = 150). note that in the results achieved by the alternative therapy strategy the modification (3.11) is used. therefore the results of the alternative strategy are better than the ones obtained with the suboptimal one. copyright c© 2020 assa. adv syst sci appl (2020) 98 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus 0 50 100 150 non-dimensional time 100 110 120 130 140 150 160 fig. 5.15. time response of the cancer cells z(t) using the optimal therapy strategy 0 50 100 150 non-dimensional time 0 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 fig. 5.16. amount of the administered drug u(t) using the optimal therapy strategy 6. summary a simplified mathematical model for the therapy of prostate cancer regarding androgen ablation is considered. the effect of the therapy on cancer and healthy cells is assumed to copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression 99 therapy strategy j(t ) e(t ) z(t ) alternative therapy strategy 1.988 · 104 46.222 117.88 suboptimal therapy strategy 2.029 · 104 47.414 123.14 optimal therapy strategy 1.195 · 104 46.000 114.98 table 5.1. comparison of the terminal values of the three considered approaches reduce the birth rate of the populations of the considered cell types. an optimization problem with a constraint on the amount of healthy cells during the therapy process is introduced. three therapy strategy approaches of the problem are illustrated and solved by combinating analytical and numerical procedures. the presented results of the therapy strategies using the suboptimal and alternative control demonstrate that the mathematical challenges of the classical optimal control problem with state constraints can be circumvented and the approaches can be used for medicine. moreover the suggested mathematical model produces possibilities for further investigations regarding the physical condition of a patient which can be modelled by means of different parameters of the model. from the mathematical point of view the considered problem covers theoretical investigations on the necessary and sufficient conditions of optimal control problems with state constraints. acknowledgements this research is supported by the russian scientific foundation grant 19-11-00008 and grant rfbi 20-04-60157 and ministry of science of higher education grant 075-15-20191621. references 1. agus, d.b., et al. (1990), prostate cancer cell cycle regulators: response to androgen withdrawal and development of androgen independence, j natl cancer inst, 91, 1869– 1876. 2. balk, s.p. (2002), androgen receptor as a target in androgen-independent prostate cancer, urology, 60, 132–138. 3. byrne, h.m. (2010), dissecting cancer through mathematics: from the cell to the animal model, nature reviews cancer, 10, 221–230. 4. bellomo, n. & preziosi, l. (2000), modelling and mathematical problems related to tumor evolution and its interaction with the immune system, mathematical and computer modelling, 32, 413–452. 5. donovan, k.a., walker, l.m., wassersug, r.j., thompson, l.m. & robinson, j.w. (2015), psychological effects of androgen-deprivation therapy on men with prostate cancer and their partners, cancer, 121, 4286–4299. 6. gatenby, r.a. and maini, p.k. (2003), mathematical oncology: cancer summed up, nature, 421, page 321. 7. gann, p.h., hennekens c.h., ma j., longcope, c., stampfer, m.j. (1996), prospective study of sex hormone levels and risk of prostate cancer, j natl cancer inst, 88, 1118– 1126. 8. heinlein, c.a., chang, c. (2004), androgen receptor in prostate cancer, endocr rev, 25, 276–308. 9. hirata, y., tanaka, g., bruchovsky, n. and aihara, k. (2012), mathematically modelling and controlling prostate cancer under intermittent hormone therapy, asian journal of copyright c© 2020 assa. adv syst sci appl (2020) 100 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus andrology, 270–277. 10. hussain, m., et al. (2009), prostate-specific antigen progression predicts overall survival in patients with metastatic prostate cancer: data from the southwest oncology group trials 9346 (intergroup study 0162) and 9916, j clin oncol, 27, 2450–2456. 11. jain, h.v., clinton, s.k., bhinder, a. & friedman, a. (2011), mathematical modelling of prostate cancer progression in response to androgen ablation therapy, proc. natl acad. sci. usa, 108, 19701–19706. 12. ledzewicz, u., schättler, h. (2006), application of optimal control to a system describing tumor anti-angiogenesis, proceedings of the 17th international symposium on mathematical theory of networks and systems (mtns), kyoto, japan, 478–484. 13. ledzewicz, u., schättler, h. (2005), a synthesis of optimal controls for a model of tumor growth under angiogenic inhibitors, proceedings of the 44th ieee conference on decision and control, 945–950. 14. miyamoto, h., messing, e.m., chang, c. (2004), androgen deprivation therapy for prostate cancer: current status and future prospects, prostate, 61, 332–353. 15. murray, j.m. (1990), optimal control for a cancer chemotherapy problem with general growth and loss functions, math biosci., 98(2), 273–287. 16. pontryagin, l.s., boltyanskii, v.g., gamkrelidze, r.v. & mishchenko ef (1962), the mathematical theory of optimal processes, interscience publishers. 17. rubinow s.i., lebowitz j.l. (1976), a mathematical model of the chemotherapeutic treatment of acute myeloblastic leukemia, biophys. j., 1257–1271. 18. seidenfeld j., et al. (2000), single-therapy androgen suppression in men with advanced prostate cancer: a systematic review and meta-analysis, ann intern med, 132, 566–577. 19. suzuki, y., sakai, d., nomura, t., hirata, y. and aihara, k. (2014), a new protocol for intermittent androgen suppression therapy of prostate cancer with unstable saddle-point dynamics, journal of theoretical biology, 350, 1–16. 20. swanson, k.r., true, l.d., lin, d.w., buhler, k.r., vessella, r. and murray, j.d. (2001), a quantitative model for the dynamics of serum prostate-specific antigen as a marker for cancerous growth: an explanation for a medical anomaly, the american journal of pathology, 158(6), 2195–2199. 21. swan, g.w. (1990), role of optimal control theory in cancer chemotherapy, math biosci., 101(2), 237–284. 22. todorov, y., nuernberg, f. (2014), optimal therapy protocols in the mathematical model of acute leukemia with several phase constraints, ocam, 35(5), 559–574. copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression101 a. appendixes a.1. model parameter values the following values were used for the model parameters (taken from [11]). parameter value parameter value parameter value ηe 110.24 δe 0.0081 δ 0.00405 γ 0.1 µe0 0.22 µ0 0.1 ke 0.889 k 0.99 ats 1.00 bts 22.23 cts 0.23 ktf 0.14 ktr 0.07 αp 0.001515 ε 0.009 λd log(2) 9 λp log(2) 12.3 λat log(2) 3 βm 0.003 βn 0.002 βt 4.569 pn 10 pm 10 kp 7 γp 0.6 104 kdf 0.018 kdr 0.053 100 kt 0.104 aet 1.62 λr log(2) 3 λa log(2) 3 λt log(2) 3 av 18.397 bv 67.753 cv 28.4 dv 12 vc 5.56 106 r0 e 90 0.002622 β table a.2. parameters of the model more detailed information is given in [11]. a.2. dynamical analysis the motivation for the construction of the alternative control strategy rests upon the asymptotically behaviour of the considered system (3.4). a relocation of an asymptotic stable critical point in order to minimize the functional (3.3) provides the alternative therapy strategy. therefore the type of the nontrivial critical point given in (3.6) is analyzed. define the jacobian j = ∂f ∂x =  ∂e(t) ∂e ∂e(t) ∂z ∂z(t) ∂e ∂z(t) ∂z  = µe(h)− 2eµe(h) ηe − εz ηe − δe −εe ηe 0 µ(h)− 2zµ(h) η − δ  at the considered critical point (3.6) holds j |(ē,z̄) = −µe(h̄) + ēµe(h̄) ηe + δe −εē ηe 0 −µ(h̄) + δ  . copyright c© 2020 assa. adv syst sci appl (2020) 102 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus in order to analyze the behaviour of the considered critical point it is necessary to analyze the sign of the eigenvalues of the jacobian. i.e., to find the roots of the polynomial defined by det(j − λē) = ( − µe(h̄) + ēµe(h̄) ηe + δe − λ )( − µ(h̄) + δ − λ ) . the eigenvalues λ1, λ2 of j are λ1 = −µ(h̄) + δ < 0 and λ2 = −µe(h̄) + ēµe(h̄) ηe + δe. for the given set of parameters (see a.1) the critical point (ē, z̄) is always stable. a.3. optimal control without state constraint constructing the pontryagin-hamilton function [16] to the problem (3), (4) h := h(e,z,r,d, t,a,at, p, h, ψ0, ψ1, ψ2, ψ3, ψ4, ψ5, ψ6, ψ7, ψ8, ψ9, u) = + ψ0(z) + ψ1 ( µee ( 1− e ηe ) − εez ηe − δee ) + ψ2 ( µz(t) ( 1− z(t) η ) − δz(t) ) + + ψ3 ( αr − λrr− kdf rd + kdr a− ktf rt + ktr a ) + + ψ4 ( βt t kt + t − λdd − kdf rd + kdr a ) + + ψ5 ( f(ts)− λtt − βt t kt + t − ktf rt + ktr at ) + + ψ6 ( − λaa+ kdf rd − kdr a) + ψ7(−λtat + ktf rt − ktr at ) + + ψ7 ( −λatat(t) + ktf r(t)t (t)− ktr at(t) ) + + ψ8 ( βp(a+ at + ϕ)− αppefrac − γpp (n +m)2 kp +n +m − λpp ) + ψ9 ( − γh+ u ) the optimal control law is given as u∗(t) = umax, ψ9(t) > 0, 0, ψ9(t) < 0, us ∈ (0, umax), ψ9(t) = 0 ∀t ∈ (t1, t2), (t1, t2) ⊂ [0, t ], where ψ9(t) is the so called switching function. note, that the optimal control function u∗(t) maximizes h in all points t ∈ [0, t ]. copyright c© 2020 assa. adv syst sci appl (2020) on optimal therapy protocols in the math. model of prostate cancer progression103 from pontryagin’s maximum principle the system of adjoint variables is given by ψ̇1(t) = −∂h ∂e = −ψ1 ( +µe(h(t))− 2e(t)µe(h(t)) ηe − εz(t) ηe − δe ) , (a.18) ψ̇2(t) = −∂h ∂z = −ψ0 + ψ1 e(t)ε ηe + ψ2 ( −µ(h(t)) + 2µz(t) η + δ ) , (a.19) ψ̇9(t) = −∂h ∂h = −ψ1 dµe(h(t)) dh(t) e(t) ( 1− e(t) ηe ) − ψ2 dµ(h(t)) dh(t) µz(t) ( 1− z(t) η ) + +ψ9γ, ψ1(t ) = ψ2(t ) = ψ9(t ) = 0. lemma a.1: in the optimal control problem (3.4) holds ψ0 = −1. proof from ψi(t ) = 0, 1 ≤ i ≤ 9, and the condition (ψ0, ψ1, ..., ψ9) 6= 0 follows that ψ0 = −1. from lemma a.1 follows ψ1(t) = 0 ∀t ∈ [0, t ], ψ2(t) = ew2(t) (∫ t 0 e−w2(s)ds− ∫ t 0 e−w2(t)dt ) , ψ9(t) = eγt[− ∫ t 0 e−γsψ2(s) dµ(h(s)) dh(s) z(s) ( 1− z(s) η ) ds+ + ∫ t 0 e−γtψ2(t) dµ(h(t)) dh(t) z(t) ( 1− z(t) η ) dt], where w2(t) = ∫ t 0 ( −µ(h(s)) + 2µ(h(s))z(s) η + δ ) ds. ψ9(t) depends on ψ2(t) and therefore this function is analysed for the mathematical solution of this problem. note, that ψ9(t) also depends on ψ1(t) but as shown before ψ1(t) ≡ 0. proof of theorem 5.1 lemma a.2: ψ2(t) is a negative monotone increasing function on [0, t ]. proof in order to analyse the behaviour of ψ2(t) the auxiliary function ψ̃2 = e−w2(t)ψ2(t) which has the same sign and zeros as ψ2(t) is used. it holds ψ̃2(t) < 0, t ∈ [0, t ) due to t∫ 0 e−w2(t)dt > t∫ 0 e−w2(s)ds ∀t ∈ [0, t ) and ψ̃2(t ) = 0. copyright c© 2020 assa. adv syst sci appl (2020) 104 m.-c. litzinger, y. todorov, m. föller-nord, m.k. chaudhary, a.s. bratus because t∫ 0 ew2(s)ds is a strictly increasing function ψ̃2(t) has no zeros on [0, t ) and therefore ψ2(t) is a negative monotone increasing function with ψ2(t) ≤ 0 ∀t ∈ [0, t ]. lemma a.3: ψ9(t) is a positive decreasing function on [0, t ]. proof from ψ9(t) = eγt ( ψ90 − ∫ t 0 e−γsψ2(s) µ(h(s)) dh(s) z(s) ( 1− z(s) η ) ds ) and ψ9(t ) = 0 ψ90 = ∫ t 0 e−γtψ2(t) µ(h(t)) dh(t) z(t) ( 1− z(t) η ) dt from the definition of the given problem follows z(t) ≥ 0 ∀t ∈ [0, t ] and( 1− z(t) η ) ∈ (0, 1). moreover it holds ψ2 ≤ 0 and dµ(h) dh < 0 ∀t ∈ [0, t ]. therefore it follows∫ t 0 e−γtψ2(t) dµ(h(t)) dh(t) z(t) ( 1− z(t) η ) dt > ∫ t 0 e−γsψ2(s) dµ(h(s)) dh(s) z(s) ( 1− z(s) η ) ds ∀t ∈ [0, t ]. thus ψ9(t) > 0 ∀t ∈ [0, t ) and u∗(t) = umax. a.4. proof of the statement proof from theorem 5.1 follows u∗(t) = umax for t ∈ [0, tj). because of the minimization of the objective function (3.3) the maximal admissible value of u on the constraint boundary has to be chosen. the limit of this value is given by u(t) = ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) ke − ( 1− εz(tj) + δeηe (ηe − ec)µ0 e ) , t ∈ [tj, t ]. copyright c© 2020 assa. adv syst sci appl (2020) introduction mathematical model of prostate cancer an alternative therapy strategy suboptimal therapy strategy optimal therapy strategy summary appendixes model parameter values dynamical analysis optimal control without state constraint proof of the statement adv syst sci appl 2019; 02; 33-43 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/709 potential scales of structural transformation of the russian electric power system based on low-carbon distributed generation fedor veselov1, irina erokhina1, andrey khorshev1* 1) the energy research institute of the russian academy of sciences, moscow, russia received march 3, 2019; revised may 5, 2019; published july 10, 2019 abstract: this article considers conditions and scenarios of transformation of the traditional structure of generating capacity in the russian electric power system in the next 10-15 years considering the mass decommissioning of operating thermal power plants (tpps) and active development of low-carbon distributed generation (dg) technologies. for this purpose, the ultimate development potential of russia's key distributed cogeneration technology was assessed, taking into account the forecast of capacity balance situation in the national electric power system. in addition to it, the possible scales of the involvement of renewable power (res) plants, as well as the potential impact of consumer’s demand response (dr) programs to replace large-scale generating facilities are considered. keywords: electric power system, distributed generation, low-carbon technologies, cogeneration, renewable power plants, demand response, capacity balance. 1. introduction for many decades, the development of the electric power industry in most countries has mainly relied on large power plants that produce power to high voltage transmission grid. however, in recent years, the development of small power sources† connected to the distribution network has been more and more intensive, and this is becoming one of the key trends in the transformation of the electric power industry of the 21st century worldwide. distributed generation (dg) becomes significant, and in some countries, such as denmark, [1] – it is also the dominant component of the national electric power system. the rapid growth of this segment is facilitated (to varying degrees in different countries) by several drivers that are of regulatory, technological and market nature. the most significant regulatory drivers are formed by purposeful public actions to implement the priorities of the national energy policy. many countries, having raised the level of energy security, increase in the energy efficiency and environmental acceptability of the electric power industry as political goals, have chosen to maximize the involvement of local (primarily renewable) energy resources and the intensive development of other low-carbon cogeneration technologies using gas, biomass and biofuel, as one of the main ways to achieve these goals. the regulatory framework and economic (tariff-based, tax-related or credit-based) mechanisms for supporting such projects have made the new sector attractive for large-scale investments. technological drivers that have provided for a significant improvement in the dg competitiveness in recent years, based on the fold reduction in the cost of wind and solar power as a result of technological improvements and the scale effect due to the multifold equipment production growth [2-4]. according to recent estimates, in a number of countries, renewable * corresponding author: epos@eriras.ru † in this article, distributed generation includes all types of power plants with an installed capacity of up to 25 mw and that are connected to a distribution grid 34 f. veselov, i. erokhina, a. khorshev copyright ©2019 assa adv. in systems science and appl. (2019) energy technologies have equalized or even surpassed traditional technologies for levelized cost of electricity (lcoe), even without special incentive measures provided by the state (for example, [5]) or a higher price of carbon; for other countries, this parity is expected in the next decade. finally, a significant role in dg development is played by market drivers associated with the development of competitive electricity markets, due to which distributed energy supply often becomes more attractive to consumers. the opportunity to choose the supplier, the differentiation of the final price of electricity with the allocation of competitive and tariff component, the transition to hourly and even shorter pricing periods stimulate consumers to a wider, more active influence on market conditions. the demand curve for electricity is becoming more elastic, and it is distributed power that offers consumers additional options for choosing power supply conditions, based on a combination of price, reliability and quality of services. often, consumers' own generation sources are combined with storage devices, as well as other mechanisms for price-dependent demand response (dr). it is necessary to note one more important feature of the dg, making it a more attractive investment solution in the context of slowly growing demand for electricity. unlike power plants with large power units, distributed sources due to small power gains can more accurately balance supply and demand for electricity and capacity both in amount and territory. this significantly reduces the risks of over-investment and excess capacity. finally, in conditions where competitive market mechanisms are limited mainly by the wholesale level, the public regulatory actions in the retail market, the rules and validity of approval of network tariffs, the practice of cross-subsidizing in the interests of certain consumer groups (usually, population) is another incentive for consumers to the abandonment of the traditional format of energy supply through a common grid and full or partial switching to own energy sources. this driver is, for example, the main incentive for the development of the dg in the russian energy system, since distortions in prices introduced by non-economic factors (in particular, keeping energy prices low for the population while having higher prices for other consumers) are the main economic reason for industrial and commercial consumers to choose an alternative model of energy supply [6]. combinations of causes and incentives that contribute to the dg development in different countries are unique and affect both the technological priorities and the overall scale of this new segment of the electric power industry. one of the countries, where the influence of regulatory and technological drivers is most significant, is germany. the active policy on the structural transformation of the electric power industry introducing renewable energy sources (res) led to the fact that in 2015 the installed capacity of the latter exceeded 90 gw. and the majority of res plants with a capacity below 25 mw connected to the distribution grid approached 79 gw [7]. together with 6 gw of small thermal power plants, this amounts to more than 40% of the total installed capacity of the national electric power system. a significant dg segment was created in the uk (25 gw, or 25%, of the installed capacity) [8]. unlike these countries, for example, in the usa, the active development of renewable energy is mostly linked with large wind farms and solar farms [9]. with a total res plants capacity of more than 100 gw, only 18 gw have a capacity below 25 mw. most of the dg sources in the us are thermal, cogeneration plants (17.5 gw), and the total share of the dg in the installed capacity of the national electric power system is still insignificant and does not exceed 5%. as we shall show below, compared to other countries, the dg role in the energy system of russia is still negligible; however, the growth rate of the small-scale power sources is noticeably higher than that of large power plants. the purpose of this article is to analyze key drivers, priorities and possible scales of the dg development for the russian electric power industry, based on current forecasts of growth in demand for electricity, the status of existing capacities, price factors in the wholesale and retail markets. potential scales of structural transformation 35 copyright ©2019 assa. adv. in systems science and appl. (2019) 2. distributed generation in the electric power system of russia and the potential for its development the unified energy system (ues) of russia is one of the largest in the world, occupying the fifth place in the world in terms of the installed capacity and the volume of electricity production in 2016, according to eia [10]. staged ues formation in the 20th century took a traditional way of generating capacities concentration, construction of large power plants (up to thousands of megawatts) and development of a high voltage network (330-750 kv). at present, thermal power plants (67.7%), nuclear power plants (11.7%) and large hydro (20.3%) play the main role in the structure of installed capacity. one feature of the ues of russia is the significant spread of cogeneration technologies (chp), which constitute about 50% of the total capacity of all thermal power plants (tpp). another feature is the extremely low level of renewable energy – the total capacity of wind, solar, geothermal and small hydro power plants in 2016 was about 0.5 gw or 0.3% of the total volume in the ues. at present, russia is pursuing a rather cautious policy to stimulate renewable energy. in accordance with the current decisions of the regulator, until 2024, about 5 gw of wind and solar power will receive tariff support, but subsequently the volume of res plants that receive support will be reduced due to the tariff growth limitations for the final consumers. according to state statistics, in 2016, russia had 36,000 power plants with a capacity below 25 mw (that are considered as dg in this article), and their total capacity was about 13.0 gw. over the past 10 years, the total installed capacity of dg power plants in russia increased by 30% (table 1), which is twice as much as the increase in installed capacity of large power plants (15.5%). however, to date, the dg is only a small, albeit an actively growing segment of the electric power industry: for 10 years, its share in the total capacity of russia's power plants increased from 4.5 to 4.9%, and in electricity production – from 2.0 to 2.6%. table 1. assessment of the scale of the dg development in russia in 2006 and 2016 2006 2016 installed capacity, gw production of electricity, twh quantity, pcs. installed capacity, gw production of electricity, twh quantity, pcs. dg – total, including: 9.97 19.76 23625 12.96 27.85 36002 % of total values for russia 4.5 2.0 4.9 2.6 in the zone of centralized power supply 2.76 8.38 261 4.47 13.65 445 tpp 2.39 7.05 196 3.38 11.18 256 res 0.37 1.33 65 1.09 2.48 189 in the zone of autonomous power supply (tpp) 7.21 11.38 23364 8.50 14.19 35557 the share of total indicators for the russian electric power industry 4.5% 2.0% 4.9% 2.6% source: russian state statistical service in the dg structure of russia, the dominant role is played by thermal power plants used for autonomous power supply of settlements isolated from the national electric power system (ues). approximately 8.5 gw (that is, about 2/3 of the entire dg capacity) is operated outside any connection to the ues of russia. the remaining 4.5 gw are used by consumers connected to the ues, to reserve supplies from the grid and to provide for themselves with electrical energy and heating. the most common types of dg are diesel, gas turbine, gas piston units, which are used as backup or peak, but mostly as cogeneration sources. in remote and isolated from power system areas distributed generation is a natural and nonalternative investment solution, whereas within the ues boundaries its development takes place in conditions of economic competition with the traditional schemes of electricity (and also heating) supply from the grid. in recent years, the growth of the dg in russia in the 36 f. veselov, i. erokhina, a. khorshev copyright ©2019 assa adv. in systems science and appl. (2019) absence of special measures of their regulatory support is due to the insufficiency of market mechanisms in the retail market, because of which the cost of supplying electricity from the grid (and for new consumers – also the level of connection fee) is higher than the cost of their own production. an important role is played by the imperfection of regulation in the heating market, where heat tariffs are less attractive for many consumers compared to the cost of heating produced from an own boiler house or a small chp. the planned transition to the heating price formation model based on long-term marginal costs envisages that the cost of heating will be determined by a competing technology – a new boiler house. such a change in the heating market will also contribute to the growth of its price and commercial attractiveness of distributed cogeneration power plants. this will reduce the risks of unprofitable sales of heating and will reduce the cost of electricity for consumers, based on the lcoe for chp based on the calculation of total discounted costs for electricity and heat production. thus, market drivers are key factors in the dg growth in russia, and the pace of this development depends on government regulatory actions in the retail electricity and heating markets that can make distributed generation an effective alternative to the traditional supply chain. however, speaking about the restructuring of the electric power system, it is necessary to assess the technically possible potential for changing the proportions between large and distributed sources in the context of a possible change in the balance of supply and demand in the ues. at present, the ues of russia has a capacity surplus, due to a suboptimal configuration of incentive mechanisms for investment into new power plants in conditions of stagnant demand. at present, it is estimated at 32 gw, and the actual power reserve in the power system is 38% and more than 2 times higher than the normative (17%). despite the extremely moderate forecast of demand growth, already in the period until 2025 this excess capacity may disappear, and in the subsequent period the capacity deficit in the ues will rapidly increase. the reason for this is the deterioration of a significant proportion of the existing thermal power plants, which equipment has already reached or will soon reach the service life limit. accordingly, investment decisions on reconstruction, replacement of equipment with more efficient or final dismantling and replacement with new capacities, should be made for these objects. currently, the wholesale market prices of electricity and capacity are insufficient to pay back investments in the reconstruction of existing large thermal power plants. the special mechanism to support investments in such projects will be launched in 2019 but any case it will not cover the entire volume of capacities that require replacement. under these conditions, there are high risks that significant volumes of the tpp’s existing capacity will leave the market. this opens up good competitive prospects for a distributed generation to replace the retiring capacity of large power plants. first of all, this refers to the replacement of large chps, many of which, as a result of a sharp decline in industrial demand for heat, are forced to contain excess heat capacity. the substitution of such power plants with smaller units, which heating capacities are adequate to actual heating demand, is one of the important directions for the dg development (the quantitative estimates of such substitution are given below). in order to assess the growing capacity deficit in conditions when the existing tpps are shutting down, it is necessary to compare the capacity requirements of the ues of russia with the volume of the guaranteed installed capacity of all types of power plants (table 2). with the macroeconomic expectations of the ministry of economic development of russia‡ the demand for electricity in the ues of russia will increase by an average of 1.2% per year and by 2035 it will amount to more than 1300 billion kwh. taking into account the normative ‡ the assessment of the change in demand for electricity and the demand for installed capacity was performed on the basis of the basic version of the long-term forecast of the socioeconomic development of the russian economy for the period until 2035, submitted to the government of russia by the ministry of economy in may 2017. potential scales of structural transformation 37 copyright ©2019 assa. adv. in systems science and appl. (2019) reserve margin, as well as various restrictions on the use of installed power plant capacity, the demand capacity requirement by 2035 corresponding to this forecast will exceed 250 gw. the value of the guaranteed installed capacity of all types is determined by the following components: the capacity of the existing tpps that have not yet reached the service life limits and do not require investment decisions (table 2), which by 2035 will decrease by 71 gw and will amount to about 55% of the 2016 level. the main reduction in the size of this segment will have to occur already in the period until 2025; the capacity of the tpps already under construction, the commissioning of which will be implemented in the coming years (until 2020) in the amount of 9 gw, based on the mid-term ues development plans for 2017-2023, approved by the system operator and by the ministry of energy of the russian federation. [11]; capacity of all types of non-carbon power plants (hydro, nuclear and res), adopted in accordance with the long-term ues development plans approved by the government of the russian federation in 2016 [12]. the capacity of nuclear power plants in the ues will increase by 7.3 gw by 2035 (by 26% from the level of 2016), the growth of large hydro plants will be about 4 gw (8% from 2016), mainly in siberia and the far east. the growth in the res capacity is expected to be about 6.5 gw, mainly in the european part of the country. table 2. characteristics of the forecast balance situation in the ues of russia in the period until 2035, gw 2016 2020 2025 2030 2035 demand for electricity 1048.5 1096.4 1165.7 1234.8 1308.0 demand for centralized heating 1235 1260 1285 1242 1310 demand for installed capacity 205.0 216.1 226.8 239.6 252.0 guaranteed installed capacity of power plants – total 236.6 245.8 188.4 186.8 189.5 operating tpp 157.2 153.1 93.6 89.2 86.2 tpps under construction 2.9 9.0 9.0 9.0 9.0 nuclear 27.9 30.6 29.7 30.5 35.2 hydro 48.1 50.7 51.2 52.1 52.1 res plants 0.5 2.4 5.0 6.0 7.0 excess capacity 31.6 29.7 capacity deficit after 2020 – total -38.6 -52.7 -62.5 including due to the growth in demand -10.7 -23.5 -35.9 thus, the amount of the guaranteed generating capacity in the ues of russia by 2025 will be reduced by almost 50 gw, remaining approximately at this level until 2035. taking into account the expected in demand for electricity and capacity requirements, the need for additional (newly commissioned) capacity, which at least partially can be provided by dg sources, can reach 39 gw by 2025, and 63 gw by 2035. based on these estimations of the potential needs for the dg, several scenarios for the mix of distributed energy technologies will be further considered, including: i. distributed cogeneration technologies; ii. res based microgeneration iii. demand management technologies. 3. scenarios for the development of the dg based on cogeneration sources from the perspective of the end-users, the problem of efficient heating supply in russian climatic conditions is no less (and perhaps more) relevant than reliable and high-quality power supply. that is why the development of distributed energy now and in the future will, first of all, focus on new small-scale cogeneration (and in the southern regions – also trigeneration) 38 f. veselov, i. erokhina, a. khorshev copyright ©2019 assa adv. in systems science and appl. (2019) sources. in this study, when assessing the development potential of the dg based on cogeneration, three groups were identified. the first group consists of the dg units, which substitute the heating supply in the service area of the existing chp plants, while the latter's capacity is reduced due to the decommissioning of the deteriorated equipment. with the introduction of market mechanisms supporting the reconstruction of existing tpp that do not take into account the low efficiency of their operation in the heating market, the reduction in their capacity as a whole in the ues of russia by 2025 may be about 28 gw, and by 2035 – more than 30 gw. accordingly, heating supply to consumers will decrease relative to 2016 by 30% by 2035. if this reduction in supplies is compensated for by dg cogeneration sources with a full load in the heating mode, the capacity of such plants can be about 20 gw by 2025-2030. the second group consists of the dg units which can respond to the additional demand for heating. given the continuing decline in the heat intensity of the russian economy, the increase in demand for heating relative to 2016 is estimated to be a modest amount – only 2% by 2025 and about 6% by 2035. taking into account the fact that cogeneration, as the most energy efficient way of electricity and heat supply, is a priority for the national energy policy, its share in the heat balance will increase, replacing traditional boiler houses. at the same time, heating output from chp will grow faster than the total demand for heating, and will increase by 7% by 2025 and by 26% by 2035. in the event that all this increase in heat supply from the chp will be provided exclusively by small sources of combined generation of electricity and heating, this will make it possible to additionally use about 5 gw of cogeneration plants by 2025 and about 18 gw by 2035. in total, two groups of dg considered above can provide input of more than 23 gw of cogeneration plants by 2025 and 39 gw by 2035. however, this value is less than the potential balance capacity deficit that may arise in the ues during this period (table 2). one of the ways to eliminate this deficit is an even more intensive development of distributed cogeneration. in this connection, the third group of dg is identified in this assessment, which will have to substitute the existing boiler houses in the heat balance, instead of their capital intensive reconstruction. one of the options for this replacement of boiler houses is their reconstruction into small chp plants with the installation of gas turbine or gas piston units. according to russian experts, the economically feasible potential of such a reconstruction of boiler houses is estimated at 60-70 gw [13]. however, based on the capacity requirements, the reasonable substitution scales should be 2-3 times smaller and will result to 17 and 25 gw of additional small chp capacity in 2025 and 2035 (table 3, scenario 1), respectively. this will lead to a serious structural rearrangement of the centralized heat balance with a sharp increase in the share of cogeneration (totally large and distributed chp) from 45% in 2016 to 70% by 2035. taking into account the high fuel utilization factor at chp plants (up to 8090%), even partial implementation of such scenario and structural reorganization of not only the electric power industry, but also the district heating system is attractive from the point of view of energy efficiency and ecological compatibility. another concomitant effect with such an active development of the chp plants may occur in the electricity market. in general, according to the ues, the share of distributed generation can reach 21% in 2025 and more than 29% in 2035. the high efficiency of cogeneration sources, mainly based on the heating load profile, will result in their annual capacity factor (50-57%) being significantly lower than the annual load factor of the ues (about 75%). to bridge this gap in the electricity balance, it will be necessary to increase the capacity factor of the most efficient large tpps with the lowest fuel costs, from 48% in 2016 to 64% by 2035 (table 3, scenario 1). in addition to the increased volumes of electricity generation at the chp in the cogeneration mode with minimal fuel costs, this change in the electricity balance structure will help reduce the price of the spot market, reshaping the profile of the electricity supply curve [14]. potential scales of structural transformation 39 copyright ©2019 assa. adv. in systems science and appl. (2019) table 3. changes in the structure of installed capacity and electricity production in the ues under various scenarios for the dg development dg development scenarios 2016 2020 scenario 1 (chp) scenario 2 (chp and res) scenario 3 (chp and demand management) 2025 2035 2025 2035 2025 2035 capacity required – total, gw 204.1 214.5 227.0 252.0 227.0 252.0 227.0 252.0 guaranteed installed capacity total, including: 236.4 246.0 188.4 189.5 188.4 189.5 188.4 189.5 existing dg 4.5 4.5 4.5 4.5 4.5 4.5 4.5 4.5 excess (+) / deficit (-) 32.3 31.5 -38.6 -62.5 -38.6 -62.5 -38.6 -62.5 dg capacity additions i. cogeneration, incl. 40.3 64.0 23.3 38.8 23.3 38.8 substitution of existing chps 18.6 20.7 18.6 20.7 18.6 20.7 ensuring the increase in demand for heating 4.8 18.1 4.8 18.1 4.8 18.1 substitution of boiler houses 16.9 25.1 ii. micro-res 16.9 25.1 additional capacity for res reservation 13.5 / 0§ 20.1 / 0 iii. demand management 16.9 25.1 electricity production – total, twh, incl. 1048.5 1096.4 1165.7 1308.0 1165.7 1308.0 1165.7 1308.0 nuclear 196.4 205.8 222.7 245.3 222.7 245.3 224.1 274.4 hydro 176.5 186.5 188.8 195.1 188.8 195.1 188.8 195.1 res plants – total 1.8 6.2 12.0 18.0 62.8 93.5 12.0 18.0 incl. res microgeneration 50.8 75.4 large tpps 662.6 686.7 529.6 518.6 563.5 568.9 604.2 598.5 capacity factor, % 48 49 61 64 57 / 65 58 / 71 69 74 distributed cogeneration 11.2 11.2 212.5 331.0 127.9 205.3 138.0 222.0 capacity factor, % 38 38 55 56 55 56 60 60 decrease in annual со2 emissions relative to the case without dg, mln tons -26 -44 -40 -64 -18 -44 share of dg in total installed capacity,% 2 3 21 29 20 / 21 27 / 29 15 22 share of dg in electricity generation,% 1 2 19 27 16 23 13 18 4. combined scenarios for the dg development with res and dr components the development of the distributed cogeneration plants can be combined with other dg technologies, including wind and solar microgeneration, as well as demand response options. each of these scenarios has its positive sides and risks and requires more detailed analysis, which is presented below. in the scenario of dg development as a mix of cogeneration and res units (table 3, scenario 2) there are additional risks associated with ensuring the reliability of energy supply. § the first value when reserving additional res with traditional tpps, the second – for the case of combination of res with accumulators 40 f. veselov, i. erokhina, a. khorshev copyright ©2019 assa adv. in systems science and appl. (2019) one risk factor is the stochastic operation of res plants and the need to reserve their capacity. in conditions of expected capacity deficit, this will require, along with commissioning of new res plants, to put in operation additional peaking (gas turbine or gas piston) units or to develop a lot of power storage systems that are still expensive for such mass use, including the pumped hydro storage plants. another risk factor is the low capacity factor of res plants (about 20-25% for onshore wind and 10-15% for solar plants on average for the conditions in russia). therefore, in addition to their reservation with peaking capacities to balance the growing electricity demand in the ues, it will be necessary to increase the electricity production of thermal power plants and their capacity factor to 58% (table 3, scenario 2). in this scenario, the replacement of large plants by a combination of distributed cogeneration and renewable sources can potentially provide for the whole ues, about 16% of the total electricity production in 2025 and about 23% by 2035. an alternative option is a combination of res plants with storage systems to provide controlled output of generated (and stored) electricity to the power system for a long time (4 to 6 hours). such a solution, as in the case of the reserve capacity additions, will require more capital and maintenance costs (system integration costs), and also involves losses of about 10% of the electricity in the storage systems. as a result [15-16], the levelized cost of electricity (lcoe) at renewable energy plants in russia, already exceeding these indicators of traditional generation (gas and coal-fired tpp and nuclear), increases by another 1.5 times (table 4). table 4. relative competitiveness of thermal, nuclear and res plants in russia (relative to lcoe of ccgt – 100%) 2015 2035 coal-fired tpp 137% 126% nuclear 171% 103% wind (capacity factor 23%) without system integration costs 327% 199% with cost of peaking plants reservation 412% 258% combined with storage systems 605% 322% solar (capacity factor 17%) without system integration costs 501% 248% with cost of peaking plants reservation 615% 346% combined with storage systems 879% 442% the results for scenario 2 provided in table 3 show that integration of res plants with storage systems will not require additional capacity reserves in the ues, but for the compensation of losses in storage process, the tpps capacity factor will increase to 65% in 2025 and up to 71% by 2035. another scenario, in contrast to the two previously discussed, provides that the remaining capacity deficit is eliminated not by capacity additions, but by reducing the peak load of consumers through demand response programs. at this it is assumed that in spite of the load profile changes the forecasted electricity demand will remain the same as in other scenarios, i.e. the impact of the energy saving programs is not considered. in this scenario, to balance the electricity demand it will be required to strongly increase the utilization of practically all types of generating capacities (table 3, scenario 3), including not only large tpps (capacity factor will grow to 74%), but also nuclear power plants (capacity factor growth to 89%) and the sources of distributed cogeneration commissioned to compensate the retirement of large chp plants and the increase in demand for heating (capacity factor growth from 56 to 60%). these estimates show that intensive demand response measures can in extremis lead to a tense balance situation in terms of marginal (and effective) modes for the operation of generating capacities and, if possible, should be accompanied by energy saving measures, resulted to the demand reduction, but not only reshaping of its hourly and seasonal profile. potential scales of structural transformation 41 copyright ©2019 assa. adv. in systems science and appl. (2019) otherwise, especially given the large share of chp in russian electric power industry, this will lead to an increase in the marginal costs of electricity generation and, as a result, an increase in electricity prices in the spot market. taking into account the possible implementation of dr programs, the contribution of dg to the required capacity will be the same as in previous versions, but due to the smaller volume of electricity production, its share in the total electricity production in the ues of russia will be slightly lower: up to 15% in 2025 and about 22% in 2035. 5. conclusion the estimates of the potential for transformation the russian electric power industry based on the use of dg sources, provided in the article, can be considered as the upper limit of its impact on the structure of the installed capacity and electricity production in the ues. clarification of these estimates, determination of cost-effective potential will require more detailed analysis of costs and benefits, comparison of additional costs for the construction and operation of different types of dg sources, costs savings of developing large power plants, also taking into account the possible increase in environmental charges. despite the significant differences in the mix of technological solutions, all three scenarios considered for restructuring the ues on the basis of distributed generation sources (low-carbon gas power plants, wind and solar power plants) are united by the fact that they provide a significant reduction in greenhouse gas (ghg) emissions compared with the "case as usual" option, within which the existing technological structure of generating capacities on the basis of large thermal power plants is preserved. the main reason for this is a gain in the fuel savings of cogeneration plants having a higher efficiency in the contrast with existing steam turbine thermal power plants (on average 38.5% for gas and 34% for coal-fired tpp), which in the "case as usual" scenario will remain in operation after equipment replacement with the same one. fuel savings makes it possible to reduce ghg emissions by 2035 by more than 40 million tons of со2, and the growth of microgeneration based on res will increase this effect by about 1.5 times (table 3). given the possible development of economic mechanisms for constraining ghg emissions to meet russia's national obligations under the paris protocol, this effect creates an additional competitive advantage for distributed generation, which can significantly facilitate the implementation of the scenarios considered. the economic effects in the grid will be ambivalent: on the one hand, the dg development allows reducing the costs of strengthening the transmission and distribution network within the traditional “vertically-oriented” (or “up-down”) scheme of energy supply. on the other hand, the rapid growth of the dg will require additional investments into the strengthening of “horizontal” electrical connections in the distribution grid, as well as the adaptation of the electric power system to bi-directional electricity flows. another important fact is that a significant part of the dg is the investment projects of electricity consumers, mainly industrial and commercial. their willingness and financial ability to invest in electricity generation assets that are not related to their main type of economic activity is a separate research task that requires microand macroeconomic modeling. at the same time, the state itself, as a regulator of the electricity market, can both stimulate and restrain the growth of the dg, primarily by liberalizing the rules of the retail market and reducing (albeit gradual) the amount of cross-subsidies that distort the economically feasible cost of electricity for consumers, influencing their investment choices. on the other hand, the state can ensure effective competition between the dg and large generation in the framework of the new economic mechanism for renewing large thermal power plants, which will enable them to identify the most effective locations for their replacement by new technologies. 42 f. veselov, i. erokhina, a. khorshev copyright ©2019 assa adv. in systems science and appl. (2019) acknowledgements the study was carried out with the financial support of the russian science foundation (project no. 17-79-20354). references 1. ropenus, s., & klinge jacobsen, h. (2015). a snapshot of the danish energy transition: objectives, markets, grid, support schemes and acceptance. study. berlin, germany: agora energiewende. [online]. available https://www.agoraenergiewende.de/fileadmin/projekte/2015/integration-variabler-erneuerbarer-energiendaenemark/agora_snapshot_of_the_danish_energy_transition_web.pdf 2. irena (2016). the power to change: solar and wind cost reduction potential to 2025. [online]. available https://www.irena.org/publications/2016/jun/the-power-to-changesolar-and-wind-cost-reduction-potential-to-2025 3. iea (2014). world energy investment outlook 2014. [online]. available https://www.iea.org/publications/freepublications/publication/weio2014.pdf 4. frankfurt school-unep centre/bnef (2018). global trends in renewable energy investment. [online]. available http://www.iberglobal.com/files/2018/renewable_trends.pdf 5. eu/irena (2018). renewable energy prospects for the european union. [online]. available https://www.irena.org//media/files/irena/agency/publication/2018/feb/irena_remap_eu_2018.pdf 6. pankrushina, t.g., solyanik, a.i., & zolotova, i.y. (2018). impact of grid tariffs on the competitiveness of distributed generating sources in the regions of russia. proceedings of the international scientific conference "far east con" (iscfec 2018). vladivostok, russia, 633-637, https://doi.org/10.2991/iscfec-18.2019.155 7. bundesnetzagentur (2016). kraftwerksliste der bundesnetzagentur. stand 10.05.2016 [power plant list of federal network agency as of 10.05.2016], [in german]. [online]. available https://www.bundesnetzagentur.de/shareddocs/downloads/de/sachgebiete/energie/untern ehmen_institutionen/versorgungssicherheit/erzeugungskapazitaeten/kraftwerksliste/kraftw erksliste_2015.xlsx 8. beis (2016). digest of united kingdom energy statistics 2016. [online]. available https://www.gov.uk/government/uploads/system/uploads/attachment_data/file/577712/duk es_2016_final.pdf 9. eia (2017). eia form 860. [online]. available https://www.eia.gov/electricity/data/eia860/ 10. eia (2017). international energy statistics. [online]. available https://www.eia.gov/beta/international/rankings/#?cy=2015&aid=7&pid=2&tl_id=2-a 11. prikaz minenergo rossii №143 ot 01.03.2017 "ob utverzhdenii skhemy i programmy razvitiya ees rossii na 2017 – 2023 gody [the order of the ministry of energy of the russian federation dated march 1, 2017 no. 143 "on the approval of the scheme and program for the development of the unified energy system of russia for 2017-2023"]. [online]. available https://minenergo.gov.ru/system/download-pdf/8170/72659 12. generanaya skhema razmeshcheniya obyektov elektroenergetiki do 2035 goda. utverzhdena rasporyazheniyem pravitelstva rossiyskoy federatsii ot 12.06.2017 № 1209-r [general scheme of the placement of electric power facilities in russia until 2035. approved https://www.agora-energiewende.de/fileadmin/projekte/2015/integration-variabler-erneuerbarer-energien-daenemark/agora_snapshot_of_the_danish_energy_transition_web.pdf https://www.agora-energiewende.de/fileadmin/projekte/2015/integration-variabler-erneuerbarer-energien-daenemark/agora_snapshot_of_the_danish_energy_transition_web.pdf https://www.agora-energiewende.de/fileadmin/projekte/2015/integration-variabler-erneuerbarer-energien-daenemark/agora_snapshot_of_the_danish_energy_transition_web.pdf https://www.irena.org/publications/2016/jun/the-power-to-change-solar-and-wind-cost-reduction-potential-to-2025 https://www.irena.org/publications/2016/jun/the-power-to-change-solar-and-wind-cost-reduction-potential-to-2025 https://www.iea.org/publications/freepublications/publication/weio2014.pdf http://www.iberglobal.com/files/2018/renewable_trends.pdf https://www.irena.org/-/media/files/irena/agency/publication/2018/feb/irena_remap_eu_2018.pdf https://www.irena.org/-/media/files/irena/agency/publication/2018/feb/irena_remap_eu_2018.pdf https://doi.org/10.2991/iscfec-18.2019.155 https://www.bundesnetzagentur.de/shareddocs/downloads/de/sachgebiete/energie/unternehmen_institutionen/versorgungssicherheit/erzeugungskapazitaeten/kraftwerksliste/kraftwerksliste_2015.xlsx https://www.bundesnetzagentur.de/shareddocs/downloads/de/sachgebiete/energie/unternehmen_institutionen/versorgungssicherheit/erzeugungskapazitaeten/kraftwerksliste/kraftwerksliste_2015.xlsx https://www.bundesnetzagentur.de/shareddocs/downloads/de/sachgebiete/energie/unternehmen_institutionen/versorgungssicherheit/erzeugungskapazitaeten/kraftwerksliste/kraftwerksliste_2015.xlsx https://www.gov.uk/government/uploads/system/uploads/attachment_data/file/577712/dukes_2016_final.pdf https://www.gov.uk/government/uploads/system/uploads/attachment_data/file/577712/dukes_2016_final.pdf https://www.eia.gov/electricity/data/eia860/ https://www.eia.gov/beta/international/rankings/#?cy=2015&aid=7&pid=2&tl_id=2-a https://minenergo.gov.ru/system/download-pdf/8170/72659 potential scales of structural transformation 43 copyright ©2019 assa. adv. in systems science and appl. (2019) by the order of the government of the russian federation dated june 9, 2017, 1209-r]. [online]. available http://static.government.ru/media/files/zzvuuhfq2f3ojik8azkvsxrgibw8engp.pdf 13. filippov, s.p., & dilman, m.d. (2014). perspektivy ispolzovaniya kogeneratsionnykh ustanovok pri rekonstruktsii kotelnykh [prospects for the use of cogeneration plants in the reconstruction of boiler houses]. industrial power engineering, 4, 7-11, [in russian]. 14. veselov, f.v., novikova, t.v., & khorshev, a.a. (2015). technological renovation of thermal power plants as a long-term check factor of electricity price growth. thermal engineering, 62, 843-852. https://doi.org/10.1134/s0040601515120137 15. veselov, f.v., makarova, a.s., novikova, t.v., tolstoukhov, d.a., & atnyukova, p.v. (2017). konkurentnyye perspektivy aes v formirovanii nizkouglerodnogo profilya rossiyskoy elektroenergetiki [competitive prospects of nuclear power plants in building the low-carbon profile of the russian electric power industry], energeticheskaya politika, 3, 6877, [in russian]. 16. makarov, a.a., veselov, f.v., makarova, a.s., novikova, t.v., & pankrushina, t.g. (2017). strategic prospects of the electric power industry of russia. thermal engineering. 64, 817828. https://doi.org/10.1134/s0040601517110064 http://static.government.ru/media/files/zzvuuhfq2f3ojik8azkvsxrgibw8engp.pdf https://doi.org/10.1134/s0040601515120137 https://doi.org/10.1134/s0040601517110064 adv syst sci appl 2020; 04:70–82 published online at https://ijassa.ipu.ru. multi-valued neural networks ii: a robot group control dmitry maximov1∗ 1v.a. trapeznikov institute of control sciences of russian academy of sciences, moscow, russia abstract: for the new concept of a multi-valued neural network introduced earlier, an analogue of the t-norm in fuzzy mathematics is considered. in the multi-valued neural network, all variables are elements of the lattice of linguistic variables, i.e., they are all only partiallyordered. the lattice operations are used to build the network output by inputs. however, a lattice elements’ multiplication may also be used to determine such operations in the case when not all of them are allowed by the lattice construction. in this paper, the lattice is assumed to be residuated, and the residual construction gives the analogue of a t-norm. the lattice elements’ multiplication determines the implication which is used, together with other lattice operations, in output determining of the neural network. though, such a construction determines a multi-valued associative memory similar to the brouwer lattice case considered earlier, this variant is more natural to use in the kohonen-like networks that we demonstrate with the example of a robot group management. keywords: multi-valued neural networks, fuzzy neural networks, robot group management, associative memory, linguistic variable lattice, system tasks’ lattice 1. introduction the concept of multi-valued neural networks was introduced in [1], [2], and “multi-valued neural networks i” by d. maximov, v. i. goncharenko, and yu. s. legovich. it continues a line of research to assess the system state using elements of a lattice. the lattice may be a lattice of sets as in [3], [4] or, even, a lattice of graphs of system state configurations [5], [6]. such an approach is used in different tasks when the situation is estimated by a set of linguistic variables along with degrees of significance/certainty of the values of these variables. usually, in fuzzy systems, confidence levels take values in a numerical interval. however, the degree of fuzziness is determined by experts. in [7] it was demonstrated that it is not necessary to use numbers: linguistic variables (but already partially-ordered), which do not require a mandatory numerical evaluation, can again serve as an estimate. such assessments are determined by the situation itself and do not need an expert opinion or, at least, the experts assessment of the situation is greatly facilitated. in this case, to compare the valuations and obtain control solutions in [7], the concepts of multi-valued logic are used in which the scale of truth values is a brouwer lattice of a general form. such scales of truth values generalize a linearly ordered scale (in particular, in fuzzy logic) and naturally define the implication as a lattice operation. in the case of a complicated lattice, the calculation of implications becomes a laborious task. in [1] and [2], it is proposed to use a multi-valued neural network to calculate implications and a multi-valued associative memory to obtain a control solution quickly. the ∗corresponding author: jhanjaa@ipu.ru, dmmax@inbox.ru multi-valued neural networks ii 71 associative memory generalizes a fuzzy one of [8] to the case of using multi-valued logic operations instead of fuzzy ones. in “multi-valued neural networks i” the results of [8] on fuzzy associative memory with thresholds to the case of multi-valued one were expanded: the authors introduced the concept of such a network, investigated the properties of the associative memory and gave a generalization of the learning algorithm of fuzzy associative memory with thresholds to the multi-valued case. the inputs, outputs and connection weights of the network are linguistic variables (partially ordered), not numbers, which, however, made it possible to use such a network for processing control and diagnostic information of complex dynamic objects. however, a brouwer lattice was used as the scale of truth value in both these cases, which restricts application possibilities. in [8], a general fuzzy associative memory with a t-norm is also considered. in residuated lattices, properties in the definition of the t-norm coincide with the same as in the residuated construction definition. in this paper we use this fact to expand the results of [8] to associative memories with variables taking values in residuated lattices that may not be brouwerian in this case. we investigate the properties of such a multi-valued associative memory and give a generalization of the learning algorithm of fuzzy associative memory to the multi-valued case similarly to “multi-valued neural networks i” and [1]. however, the application of such a construction is natural in kohonen-like neural networks rather than in associative memories due to the residuals’ intuitive meaning. we demonstrate such an application — with the example of a robot group management. 2. backgrounds feedforward fuzzy neural networks in which internal operations are based on fuzzy operations of joins and meets ∨ − ∧, were proposed in [9]. such networks are called ∨ − ∧ fuzzy associative memory, and fuzzy information in them represented by the elements of the interval [0, 1]. we will use elements of a general lattice in multi-valued associative memory instead of elements of the interval [0, 1]. 2.1. lattices [10] definition 2.1: a lattice is a partially-ordered set having, for any two elements, their exact upper bound or join ∨ (sup, max) and the exact lower bound or meet ∧ (inf, min). definition 2.2: the exact upper bound of a subset x ⊆ p of a partially-ordered set p is the smallest p element a, larger than all the elements of x: min(a) ∈ p : a ≥ x, ∀x ∈ p . definition 2.3: the exact lower bound is dually defined as the largest p -element, smaller than all the elements of x . definition 2.4: a complete lattice is a lattice in which any two subsets have a join and a meet. this means that in a non-empty complete lattice there is the largest “>” and the smallest “0” elements. if we take such a lattice as a scale of truth values in a multi-valued logic, then the largest element will correspond to complete truth (true), the smallest to complete falsehood (false), and intermediate elements will correspond to partial truth in the same way as the elements of the segment [0,1] evaluate partial truth in fuzzy logic. in logics with such a scale of truth values, implication can be determined by multiplying lattice elements, or internally, only from lattice operations. copyright c© 2020 assa. adv syst sci appl (2020) 72 dmitry maximov definition 2.5: lattice elements, from which all the others are obtained by join and meet operations are called generators of the lattice. definition 2.6: a lattice is called atomic if every two of its generators have null meets. definition 2.7: a brower lattice is the lattice that has internal implications. definition 2.8: in such a lattice, the implication c = a⇒ b is defined as the largest c : a ∧ b = a ∧ c. definition 2.9: the implication ¬a = a⇒ 0 is called the pseudo-complement of a. distribution laws for join and meet are satisfied in brouwer lattices. the converse is true only for finite lattices. 2.2. residuated lattices [11] in non-distributive lattices, the implication cannot be defined. however, we may introduce a multiplication of the lattice elements and use it to define an external implication. definition 2.10: a residuated lattice is an algebra (l,∨,∧, ·, 1,→,←) satisfying the following conditions: • (l,∨,∧) is a lattice; • (l, ·, 1) is a monoid; • (→,←) is a pare of residuals of the operation ·, that means ∀x, y ∈ l : x · y 6 z ⇔ y 6 x→ z ⇔ x 6 z ← y in this case, the operation · is order preserving in each argument and for all a, b ∈ l both the sets {y ∈ l|a · y 6 b} and {x ∈ l|x · a 6 b} each contains a greatest element (a→ b and b← a respectively). the monoid multiplication · is distributive over ∨: x · (y ∨ z) = (x · y) ∨ (x · z). also, x · 0 = 0 · x = 0. a special case of residuated lattices is a heyting algebra, when the monoid multiplication coincides with ∧. in non-commutative monoids, residuals → and ← can be understood as having a temporal quality: x · y 6 z means “x then y entails z,” y 6 x→ z means “y estimates the transition had x then z,” and x 6 z ← y means “x estimates the opportunity if-ever y then z.” you may think about x, y, and z as bet, win, and rich correspondingly (wikipedia). though, the temporal quality refers to residuals, we refer the same meaning to any multiplication x · w 6 z. however, only the residual y estimates the transition x→ z. definition 2.11: a residuated lattice a is cancellative if it satisfies the equations x · y = x · z =⇒ y = z and y · x = z · x =⇒ y = z [12]. a residuated lattice is cancellative if, and only if, it satisfies the equations x→ xy = y and yx← x = y [12]. copyright c© 2020 assa. adv syst sci appl (2020) multi-valued neural networks ii 73 definition 2.12: a residuated lattice a is said to be integrally closed if it satisfies the equations x · y 6 x =⇒ y 6 1 and y · x 6 x =⇒ y 6 1, or equivalently, the equations x→ x = 1 and x← x = 1 [13]. every cancellative residuated lattice is an integrally closed one, and any upper or lower bounded integrally closed residuated latticel is integral, i.e., a 6 1, ∀a ∈ l [13]. therefore, a finite calcellative lattice is integral. we will assume that the lattices used in a multi-valued neural network are complete, finite and residuated. we can compare the definition of a residuated lattice and the definition of a fuzzy operator and a t-norm [14], [8]: definition 2.13: we call the mapping t : [0, 1]→ [0, 1] a fuzzy operator, if the following conditions hold: • t (0, 0) = 0, t (1, 1) = 1; • if a, b, c, d ∈ [0, 1], then a 6 c, b 6 d, ⇒ t (a, b) 6 t (c, d); • ∀a, b ∈ [0, 1], t (a, b) = t (b, a); • ∀a, b, c ∈ [0, 1], t (t (a, b), c) = t (a, t (b, c)). if t is a fuzzy operator, and ∀a ∈ [0, 1], t (a, 1) = a, we call t a t-norm. for a, b ∈ [0, 1], let us define a→ b ∈ [0, 1] : a→ b = sup{x ∈ [0, 1]|t (a, x) 6 b}. we see that the latter definition coincides with the residual construction definition in the case of a commutative monoid and l = [0, 1]. excluding the condition of commutativity and the requirement l = [0, 1], we obtain the construction of a complete and residuated lattice (we demand also its finiteness). 3. multi-valued ∨ − ∗ associative memory as in the case of fuzzy neural networks [8], a multiplication of variables is used together with lattice operations in ∨ − ∗ multi-valued associative memory, where ∗ stands for the multiplication. let us suppose, an input signal is x ∈ ln, and an output signal is y ∈ lm, where l is the residuated lattice used. in this case, the input-output relationship in a twolayer multi-valued associative memory can be written as y = x ~ w, where ~ stands for the ∨ − ∗ composition operation, and w = (wij)n×m ∈ µn×m is the n×mmatrix of the weights of connections with elements from l, i ∈ n = {1...n}, j ∈m = {1...m} (fig. 3.1). in [1], several of the simplest composition variants have been suggested for ∨ − ∧ multi-valued associative memory and the next two of them are considered: yj = ∨ i {xi ∧ wij}; (3.1) yj = ∧ i {xi ⇒ wij}. (3.2) in the theory of the simplest fuzzy associative memory, only option (3.1) is considered. more complex models use a combination of ∨ and t-norms [8], [15] and a combination with implication [15]. however, they all use a linear number segment as a set of variables, unlike this paper where a general non-number lattice with multiplying elements is used. copyright c© 2020 assa. adv syst sci appl (2020) 74 dmitry maximov fig. 3.1. bilayer associative memory thus, we consider the composition 3.1, but with the monoid multiplication ∗ instead of ∧ (fig. 3.1), as in [8], however with all variables taking values in the residuated lattice l: yj = ∨ i {xi ∗ wij}. (3.3) in the vector notation, it can be written as: y = x ~ w. (3.4) we denote (x, y ) = {xk, yk | k ∈ p}, p = {1...p} — the family of pairs multi-valued patterns with xk = (xk1, ..., x k n) and yk = (yk1 , ..., y k m). we establish the connection weight matrix w ∗ = (w∗ij)n×m: w∗ij = ∧ k∈p (xki → ykj ). (3.5) the definition generalizes the similar one from [8]. then, let us define the sets [8]: sij(w∗, x, y ) = {k ∈ p | ykj 6 xki ∗ w∗ij}. mw = {w ∈ µn×m | ∀k ∈ p, xk ~ w = yk}. this is the set of the pattern pair family and the connection matrix, which satisfy the equation 3.3. the next property follows from the residual operation definition and holds both in fuzzy [8] and residuated [11] cases: a ∗ (a→ b) 6 b. (3.6) also, the following properties hold in residuated lattices [11]: c→ ( ∧ y ) = ∧ {c→ y|y ∈ y }; (3.7) b→ (a→ c) = (a ∗ b)→ c. (3.8) the property (3.6) is used in the following theorem proof exactly in the same way as in [8] due to the fact that the t-norm is a special case of the residuated lattice construction. however, such a theorem in [8] contains a necessity condition in point (iii) which does not hold in the case of a general non-chain lattice. in the general case, it may be ∨ixi = y, though ∀xi < y that contradicts the necessity. copyright c© 2020 assa. adv syst sci appl (2020) multi-valued neural networks ii 75 theorem 1: let w∗ = (w∗ij)n×m ∈ µn×m ∈ lm. then: (i) ∀k ∈ p, xk ~ w∗ ⊂ yk, and if the matrix w satisfies: ∀k ∈ p, xk ~ w ⊂ yk, then, w ⊂ w∗; (ii) if mw 6= ∅, then, w∗ ∈mw, and ∀w ∈mw, w ⊂ w∗, i.e., ∀i ∈ n, j ∈m, wij 6 w∗ij; (iii) the set mwcd 6= ∅ if ∀j ∈m, ⋃ i∈n sij(w∗, x, y ) = p . theorem 1 tells us that w∗ is the largest solution of the equation (3.4), if such a solution exists, and we will use this fact in the next section when we go down from the largest lattice element to w in a learning algorithm. 3.1. learning algorithm the iteration scheme for learning the connection weight matrix w of the multi-valued associative memory is similar to [8] with changes as in [1] and “multi-valued neural networks i”. it generalizes the dynamic δ-learning algorithm of fuzzy associative memory, introduced in [16], [17], [18] for different fuzzy cases, to the ∨ − ∗ multi-valued case. however, we are bound by atomic cancellative (hence, integrally closed) residuated lattices in the following algorithm, since theorem 2 is proved only in this case. hence, 1 = > in the lattice used since it is finite. thus, we get an algorithm of wij iteration for i ∈ n and j ∈m : step 1. initialization: for i ∈ n , j ∈m let us put wij(0) = >, t = 0; step 2. let w(t) = (wij(t)); step 3. let us calculate the resulting output: y = x ~ w, i.e., ∀k ∈ p, j ∈m, ykj (t) =∨ i[x k i ∗ wij(t)]. step 4. weights selection. denotation 1: from this place, we denote by {ykj } the set of generators contained in the lattice element ykj : {gl|gl 6 ykj } and by {w}ij the matrix of such sets. the matrix elements are the sets of generators of the weight matrix elements. matrices wij of weights and {w}ij are one-to-one correspondent to each other in distributive lattices. in non-distributive lattices, not one set {w}ij may correspond to one wij . in this case, we take max{w}ij as corresponding set to wij .the set difference will be denoted by a minus sign. then, {w}ij(t+ 1) = ⋂ k  {w}ij(t)− sym({ykj (t)} − {ykj }), if k : xki ∗ wij(t) > ykj ; {w}ij(t), otherwise. (3.9) here operation sym means the symmetrized set difference: sym(a−b) = { a−b, a ∩b ⊂ a; b − a, a ∩b = a. (3.10) step 5. for i ∈ n , j ∈m , let us check {w}ij(t+ 1) = {w}ij(t)? if this is true, then the algorithm stops, otherwise t = t+ 1 and goes to step 2. copyright c© 2020 assa. adv syst sci appl (2020) 76 dmitry maximov theorem 2: let the matrix sequence {w(t) | t = 1, 2...} is obtained by the learning algorithm. then, (a) {w(t) | t = 1, 2...} is non-increasing sequence; (b) {w(t) | t = 1, 2...} converges; (c) {w(t) | t = 1, 2...} converges to w ⊆ w∗, where w∗ij is defined in (3.5). proof (a) for i ∈ n , j ∈m , k ∈ p we get from (3.9) that {w}ij(t+ 1) 6 {w}ij(t). therefore, for i ∈ n , j ∈m : w(t+ 1) ⊂w(t). thus, {w(t) | t = 1, 2...} is a non-increasing sequence. (b) {w(t) | t = 1, 2...} converges, since the sequence wij(t) is bounded below by the smallest lattice element 0 ∀t = 1, 2.... (c) let us prove ∀t ∈ {0, 1, ...}, wij(t) 6 w∗ij . if t = 0, then wij(0) = > ≡ 1. if wij(t) = > for t = 1, 2, ..., then xki ∗ wij(t) 6 ykj , or wij(0) = > is a decision, by (3.9). by the definition of the residual, ∀k ∈ p, xki → ykj = >, or wij(0) = w∗ij by theorem 1. therefore, w∗ij = > ≡ 1 = wij(t), ∀t ∈ {0, 1, ...}. let us suppose limt{wij}(t) 6 w∗ij . let t0 is the first step such that wij(t0) 6 w∗ij . then, xki ∗ wij(t0) 6 xki ∗ w∗ij = xki ∗ ∧ k′∈p (x k′ i → yk ′ j ) 6 xki ∗ (xki → ykj ) 6 ykj by (3.6). thus, by (3.9), we get wij(t0 + 1) = wij(t0) 6 w∗ij . let limtwij(t) = wij(t0) = lij 6 w∗ij . then, lim t ykj (t) = ∨ i∈n (xki ∗ lij) 6 ykj . (3.11) therefore, by (3.9) and (3.11), we get: lim t {wij}(t) = {lij} = = ⋂ k [lim t {wij}(t)− sym({lim t ykj (t)} − {ykj })] = = ⋂ k [{lij} − ({ykj } − {lim t ykj (t)})]. (3.12) thus, limt y k j (t) = ykj for ∀k ∈ p , and lij = wij(t0) = limtwij(t) 6 w∗ij is a decision of (3.4), or {lij} ⋂ ({ykj } − {limt y k j (t)}) = ∅. the latter case means, at least, that {lij} is incomparable with ({ykj } − {limt y k j (t)}). thus, {ykj } ⋂ {limt y k j (t)} ⊇ {ykj } ⋂ {lij}. then, {ykj } ⋂ {limt y k j (t)} = {ykj } ⋂ {lij}, since xki ∗ lij 6 lij in integral lattices, and, therefore, lim t ykj (t) = ∨ i∈n (xki ∗ lij) 6 lij. (3.13) then, let us suppose limt{wij}(t) = lij > xki → ykj for ∀k ∈ p , hence lij > w∗ij . then, we get by the definition of the residual: xki ∗ lij > ykj and, therefore, lim t ykj (t) = ∨ i∈n (xki ∗ lij) > ykj . (3.14) hence, we obtain similarly (3.12) with (3.13) that limt y k j (t) = ykj that contradicts (3.14). thus, limt{wij}(t) = lij ≯ xki → ykj for ∀k ∈ p that means limt{wij}(t) 6 {w∗ij}, or limt{wij}(t) is incomparable with {w∗ij}. copyright c© 2020 assa. adv syst sci appl (2020) multi-valued neural networks ii 77 however, limt{wij}(t) may not be incomparable with w∗ij . indeed, let t0 is the first step such that wij(t0) is incomparable with w∗ij . then, if xki ∗ wij(t0) is incomparable with ykj , we get limt{wij}(t) = lij = wij(t0) by (3.9), and limt y k j (t) = ∨ i∈n(x k i ∗ lij) is incomparable with ykj or limt y k j (t) = ∨ i∈n(x k i ∗ lij) > ykj . however, the latter case contradicts (3.12) with (3.10) similarly to (3.14) and (3.13). the same holds if xki ∗ wij(t0) > ykj , also similarly to (3.14) and (3.13). let limt y k j (t) = ∨ i∈n(x k i ∗ lij) is incomparable with ykj . however, (3.12) and (3.10) demand then {limt y k j (t)} − {ykj } = ∅ since (xki ∗ lij) 6 lij , and, hence, {limt y k j (t)} ⊆ {lij}. now, let us limt{wij}(t) = lij is incomparable with w∗ij , and xki ∗ lij 6 ykj . then, xki ∗ lij is also incomparable with xki ∗ w∗ij , since the lattice is cancellative (otherwise, e.g., xki ∗ lij 6 xki ∗ w∗ij =⇒ lij 6 xki → xki ∗ w∗ij = w∗ij). hence, it should be xki ∗ lij < ykj and xki ∗ w∗ij < ykj , otherwise, they are comparable. let us consider lij → w∗ij = lij → ∧ k x k i → ykj . thus, lij → w∗ij = ∧ k lij → (xki → ykj ) = ∧ k(x k i ∗ lij)→ ykj by (3.7) and (3.8). therefore, (xki ∗ lij)→ ykj = 1 in integral lattices, since (xki ∗ lij) < ykj [13]. hence, lij → w∗ij = 1, and lij 6 w∗ij which contradicts their incomparability. thus, we obtain the unique result: limtwij(t) 6 w∗ij . 4. example of a robot group management such a network may be used as an associative memory or pattern classifier in the same manner as ∨ − ∧ networks in [2] and “multi-valued neural networks i”. it may be useful in the case of a non-distributive lattice of linguistic variables. however, a systematic method to introduce the monoid multiplication in such a lattice is absent, and we might use heuristics in every particular case. moreover, the residual intuition treats the multiplication as action: we had something in the past, have something else now, that entails a third thing as a result of the operation. all these reasons lead us to the fact that the networks are more suitable for using, not with patterns (their storing or classifying), but in some leader output detecting. thus, we consider here the problem of task distribution in a group of janitor robots. the group is fulfilling a set of tasks. then, a new task arises. robots must decide which one (or several) of them will perform the new task. such a problem was considered in [5] based on linear logic. however, the structure of linear logic is rather intricate and difficult to implement programmatically . the residuated construction is much simpler and allows the task lattice to be determined more adequately. let us consider a group of robots that can perform the following tasks: x1 — trash search, e — taking out the trash, x3 — sawing of the found garbage when it cannot be taken out, x2 — the stupor: it is the state of a robot when it tries, but cannot perform a regular task, e.g., the robot cannot pick up the trash, or cannot saw it, or can not move. all these main tasks have subtasks (fig. 4.2): trash search x1 includes d1 — a move, and d4 — a video camera operation; trash removing e includes d1 and d3 — grabbing of the debris; sawing the garbage x3 includes d3 and d2 — sawing which is not depicted for simplicity. the stupor x2 includes only the grip: a robot has grabbed a thing, but can do nothing with it. elements x2, x3, d1, and d4 may be taken as generators. all others are their meets and joins. such tasks are, e.g., c42 — a robot in a stupor operates with a video camera, or c43 — the same during sawing. in this case, we consider the joined tasks as fulfilling in parallel. however, in other cases, we may consider joined tasks as performed sequentially. the lattice in fig. 4.2 differs from the similar one in [5]: now, we include all joined tasks copyright c© 2020 assa. adv syst sci appl (2020) 78 dmitry maximov fig. 4.2. a lattice of system task estimations which have meanings. however, including all of them may be unnecessary. then, the lattice may be non-distributive. thus, we consider, as in previous papers, the task lattice as the lattice of linguistic variables denoting tasks. though, the tasks may also be associated with the sets of operations required for a task to be performed. as usually, the lattice partial order orders lattice elements by their values: the higher an element lies in the lattice diagram fig. 4.2, the greater the value it has. thus, maximal active system behaviour has the top value >, and activity absent has the bottom value 0. these values estimate transitions a→ b = c: the value c estimates the transition from the task a to the task b. or, once more, “had a then c (arising) entails b.” 4.1. monoid definition there is no a general method to define the monoid operation in a residuated lattice. every time we should use some heuristics. in our case, we demand c42 = 1, since we want to change the stupor x2 into the grip d3 if we obtain the stupor in two consequent iterations: x2x2 = d3 †, though in other cases, e.g., d1d1 = d1 ‡. therefore, it may be x2 < 1. subtasks d3 and d4 arising should not usually change the current task. thus, we consider c42 = 1 as the only right neutral element of the monoid multiplication (correspondingly, we consider only the right multiplication and left residuals x→ y), since one can hardly give equally meaningful meaning to a left neutral element. then, we want ex2 = d3, i.e., if we get stupor during debris pick up because the robot cannot lift the trash, only the grab remains from the trash remove. therefore, d1x2 ∨ d3x2 = d3, and, hence, d1x2 = d3 or d1x2 = 0, and d3x2 = 0 or d3x2 = d3. we chose d1x2 = 0 (that is evident: the stupor during the move entails an activity cessation) and d3x2 = d3, because it should be for all lattice elements x1 = x: d11 = d1x2 ∨ d1d3 ∨ d1d4 = d1, and, therefore, †we omit the multiplication sign in this section for simplicity. ‡we may also demand x2x2 = 0 — stupor changes to activity absent. copyright c© 2020 assa. adv syst sci appl (2020) multi-valued neural networks ii 79 d1x2 6= d3 §. thus, d1d3 = 0 and x2d3 = d3 by monotonicity, and, hence, d1d4 = d1. also, we demand x1x2 = d4, i.e., if a robot cannot move during a debris searching, only camera functioning remains. then, d4x2 = d4 — the stupor does not influence camera functioning. similarly, x3x2 = d3 and x1d3 = d4d3 = d4. then, evidently, x1d1 = x1, d4d1 = x1 (since, a move does not cancel the previous camera operation), x1d4 = x1 (since, the robot may stop during the search), and d4x1 = x1x1 = x1. hence, x1e = x1, and ex1 = d3x1 = e since we suppose d3d1 = d1 and d3d4 = d3 — from d31 = d3. thus, e.g., c42 = 1 = d4 → d4 and c1e = x1 → x1. note, that the robot in search does not normally switch to trash removing, as in ants’ colonies, because x1e = x1 — “had x1, now e arises, that entails x1”. this result was obtained in [3] [3] from the linear logic structure definition, while here the conclusion is the consequence of our multiplication operation choice. now, let us x2d1 = d1. therefore, x2x1 = c2e, because x2d4 = x2 from x21 = x2. then, let it be x3d1 = d1 and x3d4 = x3 — from x31 = x3 and evidently x3d3 = d3. then, x3x1 = c3e. at least, let us consider x3 multiplication: d1x3 = x3, d3x3 = x3, d4x3 = c43. the first two are evident, in the latter case we establish that sawing does not cancel the camera operation. also, we suppose x2x3 = > since we do not know if the stupor can switch to sawing at this place, or transition to any task (or again to the stupor) is possible. then, u23e = x2 → >. thus, we obtain the following basic definitions from the consideration above, though, our choice was sometimes not unique and, therefore, may be changed: d1d1 = d1 d3d3 = d3 d4d4 = d4 x2x2 = d3 or 0 x3x3 = x3 d1d3 = 0 d3d1 = d1 d4d3 = d4 x2d1 = d1 x3d1 = d1 d1d4 = d1 d3d4 = d3 d4d1 = x1 x2d4 = x2 x3d4 = x3 d1x2 = 0 d3x2 = d3 or 0 d4x2 = d4 x2d3 = d3 x3x2 = d3 d1x3 = x3 d3x3 = x3 d4x3 = c43 x2x3 = > x3d3 = d3 it is easy to verify that the (right) monoid definition satisfies these equations. 4.2. neural network operation a variant of the kohonen network may be used to determine the leading output or outputs in the problem of task distribution in a robot group. the kohonen net [19] is a bilayer neural network in which the output layer consists of n neurons with m inputs in each neuron. the output of the j-th neuron is obtained according to the following formula: yj = dj + ∑ i∈[0;m] wijxi, where x = (x1, ..., xm) is an input vector, and dj is a threshold. the network operates by the rule “winner takes it all” which determines the leading output. residuated lattice operations may be used in the formula as in ∨ − ∗ ((3.3)) and ∨ − ∧ [2], [1] multi-valued associative memories. here, we consider only a variant without a threshold, with wij = 0 if i 6= j, and with a unique input x (this is the arising task in the robot group). the formula for neuron outputs is also different then: yi = ti ∨ x > ti(ti → (ti ∨ x)) (4.15) in the expression, ti are tasks which the i-th robot supposes to perform (its intentions), and ti ∨ x denotes the i-th robot intention to fulfil task x together with the other robot tasks. §if we chose d3x2 = 0, then ex2 = 0 that is also possible. copyright c© 2020 assa. adv syst sci appl (2020) 80 dmitry maximov fig. 4.3. the neural network topology for the system task choice the right part of (4.15) is a consequent of (3.3) and corresponds to the usual logic rule: [a ∧ (a⇒ b)]⇒ b. that means we obtain b if we have a, and a entails b. the formula has the same meaning in the residuated case: ac 6 b, where c 6 a→ b. hence, we take ti ∨ x as an output and we might have an extremal principal to choose the necessary output. we take here mini[ti ∨ x] as the principal, because it is naturally to chose a robot with a minimal set of tasks which it is supposed to perform. clearly, we take several such outputs in the absence of the minimum. thus, we consider the neural network depicted in fig. 4.3, in which the input variable x is an arising task in the robot group which is fulfilling tasks ti at the current time. an arising task may be here x1, e, and x3, since stupor x2 refers to only one concrete robot, but not to all of them. we might choose now the necessary output for different input cases and tasks performed. here, we will use an extended form for outputs instead of (4.15): tix 6 ti ∨ x (“there was ti, now x arises, entails ti ∨ x”), in the case of x 6 ti → (ti ∨ x). let us suppose a standard situation: robots are looking for a trash and remove it. then, if a search task x1 arises, some of the robots already performing x1 will include this new task into their intention set. indeed, x1x1 = x1 is the minimal intention of joined tasks. a robot fulfilling subtask d4 of x1 (the camera operation) may also include the new task into its intentions, since d4x1 = x1. such a state d4 may happen if a searching robot falls into a stupor: x1x2 = d4. a robot moving without any other task (subtask d1, e.g., which return after the trash remove) may also switch to this new one, since d1x1 = d1 < x1. if a task e of taking out the trash arises, then it will be performed by a robot already removing the debris: ee = e is the minimal value in the standard situation. though, a moving robot (in d1) or a grabbing one (in d3) may also switch to taking out: d1e = d1 < e, d3e = e. the latter case arises after sawing when only grabbing remains. thus, sawn-off pieces may be removed by the same robot that has sawn them off, or one robot can saw, and another remove the sawn-off pieces. if a task of garbage sawing x3 arises, then it will be performed by one of robots already fulfilling sawing (x3x3 = x3), or by a robot in grabbing: d3x3 = x3 — these values are minimal. the latter situation arises when the robot picking up the debris can not lift it: ex2 = d3 — only the grabbing remains. if the robot cannot saw the trash, it switches to grabbing: x3x2 = d3. however, in this situation, no regular task can be performed, and collective removing is required: d3e = e, where the task e is a part of the collective action. but this decision is external to the neural network application. also, such an external decision copyright c© 2020 assa. adv syst sci appl (2020) multi-valued neural networks ii 81 is needed if the robot cannot grab the trash: ex2 = d3, then, d3x2 = d3, and we obtain in two iterations x2x2 = d3 or x2x2 = 0. thus, we see that the system of robots behaves quite reasonable, and the behaviour is due only to our choice of basic multiplications in the residuated lattice. 5. conclusion a new concept of multi-valued neural networks with variables and weights taking values in a residuated lattice was introduced in this study. the concept continues a row of studies in which the state of a system is estimated not by numbers, but by elements of a partially ordered set, namely, the lattice. in our case the lattice is finite and residuated, and it may consist of some linguistic variables. such an approach can facilitate the assessment of the situation by experts in cases requiring their participation. we have expanded the results on fuzzy neural networks to such a ∨ − ∗multi-valued case in which variables may take values in a non-distributive lattice. we found out the conditions under which it is possible to store in the multi-valued associative memory given pairs of network variable patterns. we also gave the learning algorithm that generalizes the fuzzy one for the multi-valued case. however, the algorithm is suitable in the special case of the residuated lattice: it should be integral. such a network may be used as an associative memory or a pattern classifier like the ∨ − ∧multi-valued neural network. however, here, we gave the example of such a kohonenlike network using it in a janitor robot group management. references 1. maximov d. (2020) making a decision on the management of a group of unmanned aerial vehicles by using multi-valued networks, in russian materials of the 13th international conference ’management of large-scale system development’ (mlsd’2020), moscow, trapeznikov institute of control science russian academy of science. 2. maximov d. (2020) multi-valued neural networks and their use in decision making on the management of a group of unmanned vehicles, proceedings of the 13th international conference ”management of large-scale system development” (mlsd), ieee, 1–5. https://ieeexplore.ieee.org/document/9247800. 3. maximov d. yu., legovich yu. s., ryvkin s. (2017) how the structure of system problems influences system behavior, how the structure of system problems influences system behavior, automation and remote control, 78(4), 689–699. 4. maximov d. (2019) an optimal itinerary generation in a configuration space of large intellectual agent groups with linear logic, advances in systems science and applications, 19(4), 79–86. https://ijassa.ipu.ru/index.php/ijassa/article/view/829/513 5. maximov d., ryvkin s. (2017) systems smart effects as the consequence of the systems complexity, proc. 17th international conf. on smart technologies (ieee eurocon 2017, ohrid), ohrid, ieee, 576–582. 6. maximov d., ryvkin s. (2019) multi-valued logic in graph transformation theory and self-adaptive systems, annals of mathematics and artificial intelligence, 87(4), 395– 408. 7. maximov d. (2019) control in a group of unmanned aerial vehicles based on multi-valued logic, proc. of the 12th international conference ’management of large-scale system development’ (mlsd’2019), providence, ieee, 1–5. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=8911092 8. liu p., li h. (2004) fuzzy neural network theory and application, series in machine perception and artificial intelligence,59, london: world scientific publishing co. pte. copyright c© 2020 assa. adv syst sci appl (2020) 82 dmitry maximov ltd. 9. kosko b. (1987) fuzzy associative memories, fuzzy expert systems reading, ma: addison-weley. 10. birkhoff g. (1967) lattice theory, rhode island: providence. 11. blount k. , tsinakis c. (2003) the structure of residuated lattices, int. j. algebra comput. (13), 437–461. 12. bahls p., cole j., galatos n., jipsen p., tsinakis c. (2003) cancellative residuated lattices, algebr. univ., 50(1), 83–106. 13. gil-férez j., lauridsen f. m., metcalfe g. (2019) integrally closed residuated lattices, stud. logica, 1–24. https://doi.org/10.1007/s11225-019-09888-9 14. meneganti m., saviello f. s., tagliaferri r. (1998) fuzzy neural networks for classification and detection of anomalies, ieee trans. on neural networks, (9), 848– 861. 15. sussner p., valle m. e. (2006) implicative fuzzy associative memories, eee transactions on fuzzy systems, 14(6), 793–807. 16. li x. z., ruan d. (1997) novel neural algorithms based on fuzzy δ rules for solving fuzzy relation equations: part i, fuzzy sets and systems, 90, 11–23. 17. li x. z., ruan d. (1999) novel neural algorithms based on fuzzy δ rules for solving fuzzy relation equations: part i, fuzzy sets and systems, 103, 473–486. 18. li x. z., ruan d. (2000) novel neural algorithms based on fuzzy δ rules for solving fuzzy relation equations: part i, fuzzy sets and systems, 109, 355–362. 19. kohonen t. (1989) self-organizing maps, berlin-new-york: springer-verlag. copyright c© 2020 assa. adv syst sci appl (2020) introduction backgrounds lattices residuated lattices multi-valued associative memory learning algorithm example of a robot group management monoid definition neural network operation conclusion microsoft word 1115-source texts-6538-1-18-20231219 adv syst sci appl 2023; 04; 148-155 published online at https://ijassa.ipu.ru. comparison to the proposed hybrid model and machine learning techniques for survival prediction of corona, infected patients md. asadullah1*, md. murad hossain1, md. matiur rahman molla2, md. matiur rahaman1 1 bangabandhu sheikh mujibur rahman science and technology university, gopalganj, bangladesh 2 islamic university, kushtia, bangladesh abstract: sars-cov-2, a novel coronavirus discovered in wuhan, china is spreading quickly and has a high incidence rate around the globe. as a result, everyone on the planet is having difficulty adjusting to the effects of corona and is unable to foresee the devastation and disaster caused by covid-19. in this work, we predict the survival status of patients infected with coronavirus using three distinct machine learning (ml) techniques: random forest (rf), support vector machine (svm), and logistic regression (lr). we also assess the classification performances of these algorithms. here, we put out a hybrid model and evaluated it against the three previously discussed machine learning techniques. the outcomes demonstrated the 97.85% prediction accuracy of our suggested hybrid model. aside from our suggested hybrid model, the random forest machine learning method demonstrated the highest accuracy of 94.62% among the three. nonetheless, the prediction accuracy of the hybrid model outperforms that of the random forest and is significantly better than that of the other three ml techniques. the classification performances were assessed using the f-score, sensitivity, specificity, and precision metrics. using 10-fold cross-validation, roc assessments and confusion matrices produced by these machine learning algorithms were provided and examined. to assess the effectiveness of the classification. these machine learning algorithms' roc assessments and confusion matrices are shown and examined by 10-fold cross-validation. keywords: covid-19, machine learning, hybrid model, random forest, roc curve 1. introduction the entire world is struggling to adjust to corona's effects and financial losses. they were unable to foresee the disaster and loss brought on by covid-19. numerous research investigations have been carried out to characterize the various models, strategies, innovations, and levels of awareness about the transmission of covid-19. the categorization of verified covid-19 cases associated with the epidemic diseases identified a significant obstacle to sustained advancement. the neural network's group method of knowledge handling (gmdh) form, which is linked to fake intelligence techniques, provided an explanation for binary classification modeling [1]. a deep-learning neural network concept is used to provide a verification method. using the concept of a deep-learning neural network, this system employs long short-term memory (lstm) and a gated recurrent unit (gru) for the last step of dataset training [2]. the xgboost machine learning method was assisted by building a prognostic prediction model and emphasizing emphasis. as a result, 29 patients who were cleared after february 19th were tested [3]. ai-driven techniques assist in classifying covid-19 outbreaks in a way that makes sense for their global propagation. the original goal of the research is to identify active learning-based cross-population train/test * corresponding author: asadullahstat@gmail.com word template for assa manuscript 149 copyright ©2023 assa. adv. in systems science and appl. (2024) models that use multitudinal and multimodal data in order to detect covid-19, similar to other healthcare difficulties [4]. washkaro combines real sources of data with daily news and presents it in hindi using nlp techniques, machine learning, and m-health to provide an awareness-raising solution. additionally, the appliance provides community-focused audiovisual content in local languages that has been hand-picked and verified by humans [5]. when implementing a mobile phone-based web survey, recommend utilizing machine learning techniques to be prepared to improve potential covid-19 case identifications more quickly. additionally, this will lessen the spread among those who are vulnerable [6]. an early screening model to distinguish covid-19 pneumonia from influenza was created using deep learning techniques within a virus infection and healthy cases with lung ct pictures [7]. using chest x-ray radiographs, three distinct convolutional neural networkbased models are proposed for the identification of individuals afflicted with coronavirus pneumonia. using 5-fold cross-validation, roc assessments and confusion matrices using these explained models are provided and examined [8]. described the machine learningbased classification of the extracted deep feature on chest x-ray images using patients with covid-19 and pneumonia using resnet152. it is also possible to anticipate the spread of the novel coronavirus (covid-19) in likely patients using non-invasive and early forecasting of the virus by examining chest x-rays [9]. this study [10] reveals and uses machine learning-based ct radionics models to predict hospital stays for individuals with pneumonia associated to sars-cov-2 illness. uses the publicly available covid-19 chest x-ray dataset to illustrate how bayesian convolutional neural networks (bcnn) can quantify the uncertainty in deep learning solutions to improve the diagnostic performance of the human-machine combination and describe the uncertainty in prediction is strongly associated with the accuracy of the prediction [11]. this research assumes that artificial intelligence's deep learning methods could afford to extract covid-19's specific graphical features and supply a clinical diagnosis, thus saving critical time for disease decline, to do the pathogenic test, through the support of covid-19 radiographical changes in ct images [12]. utilizing the recently developed vaxign-ml machine learning technique in conjunction with vaxign reverse vaccinology, covid-19 vaccine candidates are predicted. six proteins, including the five non-structural proteins (nsp3, 3cl-pro, and nsp8-10) and s protein, were predicted to be adhesins after experiments with the entire proteome of sarscov-2 [13]. the previous study suggested three well-known techniques: enhancing the currently limited data, using a panel selection method to select the most straightforward forecasting model among multiple models, and optimizing a personal prediction model's parameters for optimal accuracy. however, this study uses a method that supports the three benefits of knowledge mining from a small, poor dataset [14]. to ascertain the scope, duration, and end date of covid-19 in china rather than using epidemiological models to analyze the dynamics of covid-19 transmission in china. this work suggests real-time covid-19 forecasting techniques influenced by fake intelligence (ai) [15]. try testing the model's forecasting power in this article by training it on data from the 24th of january to the 3rd of march window, matching the forecasts up to the 1st of april. a one-week buffer window (march 25–april 1) was provided even for italy and the republic of korea to validate the model's forecast. in the instance of the us, the model accurately depicts the progression of the infected curve and projects that the infection will stop by april 20, 2020 [16]. a case study utilizing fuzzy rule induction and cmc—which is improved by deep learning networks—is examined in order to gain more accurate stochastic insights on the evolution of the pandemic. fuzzy rule induction approaches are used in conjunction with a deep learning-based cmc, as opposed to relying solely on uniform and uncomplicated assumptions for an mc [17]. throughout this work, a novel bio-inspired metaheuristic simulates the transmission and infection of healthy individuals by the coronavirus. to replicate coronavirus behavior as closely as feasible, relevant parameters such as superspreading rate, traveling rate, and re-infection probability are added to the model [18]. 150 i. petrov, j.f. doe copyright ©2023 assa adv. in systems science and appl. (2023) incorporating city-to-city links, a mathematical framework is constructed to determine the total number of secondary cases caused by the imported cases, as well as the number of imported cases of the novel virus from epidemic resources. additionally, to replicate outbreaks in several cities, a meta-population compartmental model was constructed and backed by a classical sir technique [19]. this study compares the prediction outcomes of three distinct mathematical models based on many parameters and multiple locations. while the gompertz model's fitting effect is also superior to the bertalanffy model's, the logistic model's fitting effect is also the best [20]. three biomarkers were chosen by machine learning methods in order to show the survival of specific patients by identifying susceptible predictive biomarkers of the severity of the disease. in order to prioritize patients at the highest risk and maybe lower the rate, this research proposes a straightforward and workable method [21]. studies have shown that patients with cardiovascular and metabolic problems were more likely to contract covid-19 and that the infection worsened. this investigation aims to determine the relationship between metabolic and cardiovascular illnesses using covid-19 [22]. nonetheless, the aforementioned studies discuss diverse mathematical strategies, artificial intelligence models, neural network models, and verification and validation methods for the prediction of patients impacted by corona. some articles even used a statistical, stochastic, and mathematical model to describe how corona transmission occurs. however, a small number of studies—particularly those using hybrid models—used machine learning methods. no such noteworthy research paper attempted to use machine learning tools to create a survival model for covid-19 patients. in this work, we developed a hybrid model and focused on forecasting the survival of coronavirus-infected individuals utilizing currently available machine learning devices. we have made an effort to more closely evaluate our suggested hybrid model with the available machine learning tools. in this study, we also attempt to quantify the effectiveness of different machine learning technologies in terms of classification, validation, and accuracy. lastly, we try to demonstrate that the employed data set is well described by our suggested hybrid model, demonstrating strong performance in terms of accuracy, precision, sensitivity, specificity, f1 score, and auc. 2. materials and methods 2.1 dataset collection and processing we used metadata from "https://github.com/ieee8023/covid-chestxray-dataset" in this paper. there are 312 patients in our dataset, and each patient has additional data (referred to as a variable). the following unique characteristics are included in our dataset: patient id, finding/type of pneumonia, age (years of the patient), sex (male or female), survival status (yes/no), offset (number of days), went_icu (yes (y) if the patient was in the intensive care unit (icu) or critical care unit (ccu) at any point during this sickness or no (n) or blank if unknown), iintubated (yes (y) if the patient was intubated (or ventilated) at any time during this illness, etc. there are 109 female patients and 203 male patients among the 312 patient records in this collection. this is a list of all the metadata fields along with covid-19related explanations. we forecast the survival status of patients infected with corona based on the characteristics. version 3.6.3 of the r program is used for data processing and analysis. 2.2 proposed hybrid model the idea behind hybrid or ensemble techniques is that a combination of several single models can often produce effective discriminatory rules. we are going to build ensembles of machine learning algorithms. we integrate predictions from four caret models via stacking. this may indicate that the models are proficient in multiple domains, enabling a novel word template for assa manuscript 151 copyright ©2023 assa. adv. in systems science and appl. (2024) classifier to determine how to extract the best performance and accuracy from each model. in this instance, we will utilize the random forest method to combine the predictions, which is why the hybrid model shows the highest accuracy when compared to all or any machine learning software alone. 2.3 experimental setup figure 2.1 outlines the several stages that our investigation covered. we pre-processed the dataset for our predictive model in the first stage. the dataset is then divided into training and testing datasets as the next phase. three well-known classification methods as well as a hybrid model were used in our investigation. as a result, the most accurate prediction model gets approved to be used as a future model. fig.2.1. flowchart for proposed hybrid model 2.4. machine learning technique 2.4.1 support vector machines (svm) support-vector machines (svms, support vector networks) are supervised learning models with corresponding learning algorithmic programs used for regression analysis and compartmentalization through simulated data in machine learning techniques. in addition to performing linear variety, support vector machines (svms) can also do nonlinear classification by implicitly transforming their inputs into high-dimensional feature spaces, a technique known as the kernel trick. the support vector machine category was created by vladimir vapnik and alexey chervonenkis in an effort to transfer a linearly separable hyperplane and split the dataset into two groups. assuming perfect data separation, let's proceed. after that, we can maximize the subsequent: minimize 2, subject to: 152 i. petrov, j.f. doe copyright ©2023 assa adv. in systems science and appl. (2023) the last two constraints can be compacted to: a quadratic algorithm called linear svm works effectively with datasets that are easily divided into two sections by a hyper-plane. however, datasets can occasionally be complex and challenging to categorize with a linear kernel. 2.4.2 random forest model even without hyper-parameter adjustment, the random forest method effectively employs a machine-learning algorithm that produces excellent results most of the time. because of its ease of use and versatility, it's also one of the most used algorithms (used for classification and regression jobs). the random forest has a great advantage when used to classification and regression tasks, which make up the majority of machine learning systems in use today. the hyperparameters of random forests are identical to those of decision trees and bagging classifiers. we may also use the regressor algorithm to carry out regression operations for random forest. the irregular forest classifier is made up of a combination of tree classifiers. each tree casts a unit vote to identify the most prevalent input vector. each classifier is constructed using an arbitrary vector that may be freely examined from the input vector. 2.4.3 logistic regression there are significant similarities between logistic regression and linear regression. nevertheless, their intended application remains the primary distinction. while logistic regression is utilized for classification tasks, linear regression methods are employed for value prediction. let be a binary outcome from bernoulli with . the logistic regression model can be defined as: 2.5 performance evaluation we evaluate the machine learning methods for classification based on the following criteria. technique names formula accuracy = precision = sensitivity/recall = specificity = f1-score = 3. results and discussion to test the three machine learning classification algorithms and one hybrid model to predict the survival status of patients infected with coronavirus, we conducted several analyses. the prerequisites for machine learning algorithms' specialized performance evaluation are shown in table 3.1. table 3.1. performance measurements for classification technique algorithm accuracy sensitivity specificity precision f1 score auc rf 0.9462 0.6923 0.9875 0.9875 0.8432 0.9514 word template for assa manuscript 153 copyright ©2023 assa. adv. in systems science and appl. (2024) svm 0.9355 0.7692 0.9625 0.9625 0.8449 0.8683 lr 0.7634 0.1538 0.8625 0.8625 0.2611 0.8587 hybrid 0.9785 0.8462 1.0000 1.0000 0.9167 0.9793 the results of the performance evaluations for the three machine learning methods used to predict survival are shown in table 3.1. it also explains how a suggested hybrid model was developed. out of the three machine learning methods, random forest performs the best with an accuracy of 94.62%, while logistic regression has the lowest accuracy at 76.34%. we evaluate the performance of three machine learning algorithms using the current standards, and the random forest model consistently outperforms the other two. however, compared to random forest, our hybrid model produced superior accuracy (97.85%). out of all the machine learning techniques, our suggested hybrid model produced better results when compared to the other evaluation criteria. we also use the roc curve to illustrate the performance of the various ml methods. the discriminant power in four distinct algorithms is shown in graph 2 for various colors. the color black represents our suggested hybrid model. our suggested model has a strong ability to discriminate. the suggested model has an auc value of.9793, which indicates that the predictions of the proposed hybrid model are 97.93% accurate, according to table 3.1's auc value. if we take into account logistic regression as well, the auc=.8587 indicates that logistic regression accurately predicts 85.87% of patients with corona infection who do not survive. fig. 2. roc curve for different machine learning algorithm compare to hybrid model 4. conclusion after examining the trial data, we came to the conclusion that our suggested hybrid model achieves the best accuracy, 97.85%, and auc, 97.93%. as a result, our suggested approach can identify corona-infected patients who do not survive more precisely. in terms of classification accuracy or area under the roc curve, our hybrid model performed the best. the suggested hybrid model's auc score is close to 1, indicating a more accurate prediction. our suggested hybrid model appears to perform better across the board for all performance assessment factors, according to the experimental results. we observe that the suggested 154 i. petrov, j.f. doe copyright ©2023 assa adv. in systems science and appl. (2023) model with the lowest error has the highest accuracy. as the literature study has shown, we conclude that the criteria of a predictive model for corona-infected patient survival prediction has only been partially met. in conclusion, we can state that, when compared to the three machine learning approaches, our suggested hybrid model is the most predictive model for all machine learning methods and has a lower classification error. acknowledgement we thank all the co-authors who contributed to preparing this manuscript – especially the author md. asadullah has does analysis and contributed to interpretation. secondly md. murad hossain prepared an introduction and contributed to overall manuscript writing. remaining co-authors help us reviewing all write-ups. references [1] asadullah, m., hossain, m. m., rahaman, s., amin, m. s., sumy, m. s. a., et al. (2023). evaluation of machine learning techniques for hypertension risk prediction based on medical data in bangladesh, indonesian journal of electrical engineering and computer science, 31(3), 1794–1802. [2] bandyopadhyay, s. k. & dutta, s. (2020). machine learning approach for confirmation of covid-19 cases: positive, negative, death and release, iberoamerican journal of medicine, 03, 172–177. [3] dandekar, r. & barbastathis, g. (2020). quantifying the effect of quarantine control in covid-19 infectious spread using machine learning, medrxiv: 2020.04.03.20052084, [online]. available: https://www.medrxiv.org/content/10.1101/ 2020.04.03.20052084v1. [4] fong, s. j., li, g., dey, n., gonzalez-crespo, r. & herrera-viedma, e. (2020). finding an accurate early forecasting model from small dataset: a case of 2019ncov novel coronavirus outbreak, int. j. interact. multimed. artif. intell., 6(1) 132. [5] ghoshal, b. & tucker, a. (2020). estimating uncertainty and interpretability in deep learning for coronavirus (covid-19) detection, arxiv: 2003.10769, [online]. available: http://arxiv.org/abs/2003.10769. [6] hu, z., ge, q., li, s., jin, l. & xiong, m. (2020). artificial intelligence forecasting of covid-19 in china, arxiv:2002.07112, [online]. available: http://arxiv.org/abs/ 2002.07112. [7] hossain, m. m., asadullah, m., hossain, m. a., & amin, m. s. (2022). prediction of depression using machine learning tools taking consideration of oversampling. malaysian journal of public health medicine, 22(2), 244–253. [8] hossain, m. m., asadullah, m., rahaman, a., miah, m. s., hasan, m. z., et al. (2021). prediction on domestic violence in bangladesh during the covid-19 outbreak using machine learning methods, applied system innovation, 4(4), 77. [9] jia, l., li, k., jiang, y., guo, x. & zhao, t. (2020). prediction and analysis of coronavirus disease 2019, arxiv: 2003.05447, [online]. available: http://arxiv.org/abs/2003.05447. [10] kumar, r., arora, r., bansal, v. & sahayasheela, v. j. (2020). accurate prediction of covid-19 using chest x-ray images through deep feature learning model with smote and machine learning classifiers, medrxiv: 2020.04.13.20063461, [online]. available: https://www.medrxiv.org/content/10.1101/2020.04.13.20063461v1. [11] korkut, s., göl, s., & kilic, m. s. (2020). poly (pyrrole-co-pyrrole-2-carboxylic acid) / pyruvate oxidase based biosensor for phosphate: determination of the potential, and application in streams, electroanalysis, 32(2), 271–280. [12] li, b., yang, j., zhao, f., zhi, l., wang, x., et al. (2020). prevalence and impact of word template for assa manuscript 155 copyright ©2023 assa. adv. in systems science and appl. (2024) cardiovascular metabolic diseases on covid-19 in china, clin res cardiol, 109(5), 531–538. [13] mamani, s., nolan, d. a., shi, l. & alfano, r. r. (2020). special classes of optical vector vortex beams are majorana-like photons, opt. commun., 464, 125425. [14] martínez-álvarez, f., asencio-cortés, g., torres, j. f., gutiérrez-avilés, d., melgargarcía, l., et al. (2020). coronavirus optimization algorithm: a bioinspired metaheuristic based on the covid-19 propagation model, arxiv: 2003.13633, [online]. available: http://arxiv.org/abs/2003.13633. [15] ong, e., wong, m. u., huffman, a. & he, y. (2020). covid-19 coronavirus vaccine design using reverse vaccinology and machine learning, biorxiv: 2020.03.20.000141, [online]. available: https://www.biorxiv.org/content/10.1101/ 2020.03.20.000141v2. [16] pandey, r., gautam, v., bhagat, k. & sethi, t. (2020). a machine learning application for raising wash awareness in the times of covid-19 pandemic, arxiv: 2003.07074, [online]. available: http://arxiv.org/abs/2003.07074. [17] pirouz. b., haghshenas, s. s., & piro, p. (2020). investigating a serious challenge in the sustainable development process: analysis of confirmed cases of covid-19 (new type of coronavirus) through a binary classification using artificial intelligence and regression analysis, sustain, 12(6), 2427. [18] qi, x., jiang, z., yu, q., shao, c., zhang, h., et al. (2020). machine learning-based ct radiomics model for predicting hospital stay in patients with pneumonia associated with sars-cov-2 infection: a multicenter study, medrxiv: 2020.02.29.20029603, [online]. available: https://www.medrxiv.org/content/ 10.1101/2020.02.29.20029603v1. [19] rao, a. s. r. s. & vazquez, j. a. (2020). identification of covid-19 can be quicker through artificial intelligence framework using a mobile phone-based survey in the populations when cities/towns are under quarantine, infect. control hosp. epidemiol., 1400. [20] santosh, k. c. (2020). ai-driven tools for coronavirus outbreak: need of active learning and cross-population train/test models on multitudinal/multimodal data, j. med. syst., 44(5), 1–5. [21] wang, s., kang, b., ma, j., zeng, x., xiao, m., et al. (2020). a deep learning algorithm using ct images to screen for corona virus disease (covid-19), medrxiv: 2020.02.14.20023028, [online]. available: https://www.medrxiv.org/content/ 10.1101/2020.02.14.20023028v5. [22] xu, x., jiang, x., ma, c., du, p., li, x., et al., deep learning system to screen coronavirus disease 2019 pneumonia, arxiv: 2002.09334, [online]. available: http://arxiv.org/abs/2002.09334. [23] yuan, h.-y., hossain, m. p., tsegaye, m., zhu, x., jia, p., et al. (2020). estimating the risk on outbreak spreading of 2019-ncov in china using transportation data, medrxiv: 2020.02.01.20019984, [online]. available: https://www.medrxiv.org/content/10.1101/2020.02.01.20019984v1. [24] yan, l., zhang, h.-t., goncalves, j., xiao, y., wang, m., et al. (2020). a machine learning-based model for survival prediction in patients with severe covid-19 infection, medrxiv: 2020.02.27.20028027, [online]. available: https://www.medrxiv.org/content/10.1101/2020.02.27.20028027v3. [25] yan, l., zhang, h.-t., xiao, y., wang, m., guo, y., et al. (2020). prediction of criticality in patients with severe covid-19 infection using three clinical features: a machine learning-based prognostic model with clinical data in wuhan, medrxiv: 2020.02.27.20028027, [online]. available: https://www.medrxiv.org/content/ 10.1101/2020.02.27.20028027v2. microsoft word 1001-article text-5834-1-11-20221223 adv syst sci appl 2023; 01; 99-114 published online at https://ijassa.ipu.ru. numerical investigation of fluid-structure, thermal coupling for a heated slab oudrane a.1*, aour b.2, hamouda m.1, chesneau x.3, zeghmati b.3, balti j.4 1) laboratory of sustainable development and informatics (lddi), faculty of science and technology, ahmed draya university of adrar, algeria 2) laboratory of applied biomechanics and biomaterials (labab), bp 1523 el mnaouer, national polytechnic school of oran-maurice audin (enpo-ma), 31000, oran, algeria 3) laboratory of mathematics and physics groups of energy mechanics (lamps), university of perpignan via domitia, 52, avenue paul alduy, 66860 perpignan cedex france 4) department of physics, faculty of science, university of carthage, 7021 jarzouna, tunisia abstract: this work focuses on a numerical analysis of fluid-structure thermal coupling in a heating slab. the latter consists of a rectangular cross-section duct located in the concrete slab of a discredited habitat. we modeled the thermal transfers of fluid flow in the pipe. in fact, the navier-stokes equations that govern this flow have been solved numerically. these equations were by an implicit method of finite differences. the systems of algebraic equations thus obtained were solved by the algorithms of gauss and thomas. the equation of conduction in the concrete slab was solved using the same methodology as that of flow. in this work, we based on an algorithm that interacts non stationary solid medium with a fluid medium consisting of permanent a state by ensuring equal flows and temperatures on the common interface between the two mediums at every moment. the numerical simulation of heat transfers and the thermal behavior of the heating slab were analyzed for various parameters influencing thermal diffusion. the results obtained show that the numerical methodology adopted for the control of fluid-structure coupling is acceptable in comparison with the literature. keywords: heated slab; thermal coupling; fluid flow; heat transfer; not stationary processes 1. introduction one of the renewable energy sources is solar energy, which is the most used within cities in the mediterranean. passive solar is a process of generating thermal energy by converting solar radiation into heat. solar energy is the most developed compared to other renewable energies [1–3]. however, the behavior of conversion systems for this energy type is highly dependent on variations in climatic parameters such as temperature, solar irradiation and storage means [4]. in the context of thermal coupling, giles and al. [5] have studied the numerical stability and procedures of fluid-structure thermal coupling. the aim of this study is to analysis the thermal diffusion phenomenon with a continuity of temperature and a heat quantity at the interface. birken and al. [6] highlighted the importance of fluid-structure, thermal coupling in industrial cooling processes for the steel heat treatment [6]. this numerical study is devoted to the thermal interaction study between fluid and structure in the heating slab, also called conjugated heat transfer. monge and birken [7] have considered two areas with jumps in the material conductivity coefficient through the connection interface. heuzé and al. [8] * corresponding author: abdellatif.mebarek@univ-adrar.edu.dz 100 oudrane a. et al. copyright ©2023 assa adv. in systems science and appl. (2023) have developed a digital tool to simulate the thermomechanical coupling procedure based on a fluid-structure coupling to describe the state of the structure matter throughout. currently, numerical simulation of fluid-structure interaction problems is one of the biggest challenges for modern scientific computing. typical examples are found in aeronautics, where air flow around an elastic aircraft or air-sheet oscillations in air flow are well presented in the work of dowell and al. [9]. in addition, this interaction is important in the in-turbomachinery field, where the energy transfer will take place between a rotor and air [10]. furthermore, in the field of biomechanics, the elastic behavior of micro-pump or artificial membranes in blood flow is affected by this type of interaction by referring to the work of scotti and al. [11] and then tezduyar and al. [12]. underfloor heating is a technique that provides good comfort while minimizing energy consumption. in this context, the aim of this study is to characterize the variable heat exchange between a laminar flow of fluid in forced convection and a concrete slab of considerable thickness, the top of which is subjected to a constant ambient temperature of 28°c. the modelling is based on the thermal balance calculation at the level for system elements: fluid-structure. model validation was performed using the results obtained by andreo and al [13]. the latter used the same heating system with a heat supply provided by solar energy. the second step in this work is to test the parameters influencing the heat transfer within the heating slab, analysing the system thermal behavior. as part of this work, we propose to exploit this energy potential in the habitat through the use of a closed-circuit heating floor. studies on floor heating technology are numerous in the north african region. we are interested in the case of a solar heating system for a singlezone space in a dry climate similar to that studied by mokhtari and al [14]. the system is equipped with a concrete slab to store and produce heat from the floor within a habitable envelope. the principle of operation is to circulate directly into the concrete slab a fluid heated by solar collectors [15]. in fact, this slab will have the diffuser role of a soft and homogeneous heat throughout the house. 2. physical description of the slab studied the element of the numerical modelling is a hydronic heating floor which consists of three layers with a coil tube (figure 1). the insulation material is a plate of expanded polystyrene, which is the most used in this kind of habitable constructions. the insulation layer height is 5 cm. above it, we have a concrete screed 10 cm high, 1m long and 1m wide. in the latter are arranged the cross-linked polyethylene tubes [16,17]. they are very often used for the heated floors realization. in fact, these semi-rigid pipes are flexible, and they do not need welding to be carried out like those of copper. the tubes are arranged in (u) shape with a diameter (d) of 20 mm and a 10 cm for spacing [17]. a layer of concrete coating is superimposed on the heating grid. fig.1. physical description of the slab studied numerical investigation of fluid-structure… 101 copyright ©2023 assa. adv. in systems science and appl. (2023) 3. mathematical formulation 3.1. simplifying hypotheses a set of assumptions is retained in this study to simplify the mathematical modelling for thermal transfer model. these assumptions are derived from the physical properties of fluid flow in a horizontal pipe embedded in a concrete slab. the main assumptions taken into account in this study are as follows:  the fluid is newtonian, assumed to be viscous and incompressible;  the flow is transient with a laminar regime;  viscous dissipation is negligible: the diffusion of purely mechanical energy is neglected because the water speed and viscosity are low;  flow has only two velocity components: one longitudinal speed and the other transverse;  no internal heat sources 0s  ;  the physical properties (µ, cp, ρ, λ) are constant;  low pipe thickness is neglected in numerical calculations;  the fluid-structure interface is in thermodynamic equilibrium;  the solid medium to be isotropic. 3.2. fluid flow modelling in the pipe we initially opted for a flow study in a rectangular cross-section duct. this involves the flow of viscous fluid between two long plates (l), parallel and separated by a small distance (d). both plates are fixed, and the fluid is moved by a pressure gradient (figure 2). the solution governing a flow of poiseuille for maximum speed (u0) in the pipe medium is as follows [18]:              2 0 41)( d y uyu f (1) fig.2. physical description of fluid flow within the pipe it should be noted that a rigorous treatment of the boundary layer would require the complete solution to navier-stokes equations. their complexity prompted prandtl to simplify them to retain only the most important terms. the main idea is to neglect the axial gradients (∂/∂x) in front of the transverse gradients (∂/∂y). thus we obtain the prandtl equations for boundary layer which govern a laminar flow in the heating slab pipe as follows [19,20,21]:  equation the amount of movement 0 1 2 2                  y p and y u v x p y u v x u u t u ff f f f f f f f f  (2)  equation of mass conservation 0      y v x u ff (3) 102 oudrane a. et al. copyright ©2023 assa adv. in systems science and appl. (2023)  energy conservation equation 2 2 y t cy t v x t u t t f fpf f f f f f                (4) 3.3. mathematical model of thermal diffusion the floor slab is regarded as a homogeneous solid to which the classical equation of heat diffusion is applied [22]. heat diffusion is defined as a heat transmission mode in a solid caused by a temperature difference between two regions of this solid medium [23,24]. the numerical modelling is based on the two-dimensional study of heat conduction within the heating slab. the equation of thermal conduction is:                 2 2 2 2 y t x t ct t bb bb bb p  (5) this equation is discredited by the finite difference method using a three-point forward scheme in the contact interface. after the discretization, we get a tridiagonal algebraic system, the resolution of which is done by the thomas algorithm. 3.4. initial and boundary conditions  at the moment t = 0, the field of calculation under consideration is initialized by: 0)0(;0)0(;0)0(  fff tvu (6) 0)0(  b t (7)  the fluid velocity at the channel inlet is given by x = 0 and 0 ≤ y ≤ d;              2 0 41)( d y uyu f (8)  at the canal outlet we have x = l and 0≤ y ≤ d; 00;0          x t and x v x u fff (9)  the coupling interface (fluid-structure) is governed by the following expression 0f(1). in the next generation, plant no. 1, using cheaper hydraulic drives, will reduce the cost of manufacturing robots: the cost of a robot with cheaper hydraulic drives will be reduced by the amount of reduction in the cost of hydraulic drives and will be equal to c(1) talarm. the distribution parameters and talarm are estimated during calibration on the validation sample without anomalies, where tt has to take values close to 0 without growing trend. this approach allows to detect mean and variance changes of the residuals. an example of the cusum test is demonstrated in fig. 4.2. the test value increases for anomalies and falls when the anomaly disappears. to avoid the test values relaxation after the anomaly tt values can be reset right after the anomaly detected or by operator’s demand. 4.3. one-class classification one-class classification algorithms are used for anomalies and novelty detection in data. these algorithms learn boundaries of a normal class and everything that outside that boundary is considered as an anomaly. isolation forest [18] is one of the best one-class algorithms. suppose a sample with n objects is given. each object has d parameters. the idea of the isolation forest in the following: step 1 select a subsample with 0 < n < n objects to build a new isolation tree. step 2 randomly select a parameter. step 3 randomly select an object from the current tree node. use the selected parameter value of this object as a splitting rule for the node. copyright c© 2019 assa. adv syst sci appl (2019) 28 step 4 repeat steps 3 and 4 for the left and right children of the node. stop the process when the termination conditions are satisfied. step 5 repeat steps 1-4 to build forest of isolation trees. according to [18], anomalies have shorter decision paths in isolation trees than normal objects. for an object x anomaly score ŝ is estimated as: s(x) = 2− e(h(x)) c(n) (4.9) where h(x) is a length of s decision path for the object in an isolation tree; e(h(x)) is an average length of the decision path for the object in the forest; c(n) is an average length of the decision path for all n objects in the forest. the score takes values in [0, 1]. objects with the score close to 1 are anomalous. on the other hand, objects with small score are normal. in case, when for all objects in a sample s(x) ≈ 0.5, the sample has no anomalies. similar to the time series analysis approach, isolation forest uses vectors yt of measured parameters values to estimate the anomaly score ŝit for a measurement: ŝit = fi(yt, yt−1, ..., yt−k) (4.10) an example of working of isolation forest is demonstrated in fig. 4.3. the first 60000 seconds are used to train a classifier. the figure shows a slightly higher anomaly score for anomalies. all measurements with ŝit ≥ 0.5 can be considered as anomalies. results can be improved using cusum test similar to var models and will be considered further. 4.4. binary classification isolation forest binary classification fig. 4.3. (top) time series of percentage of cpu usage for one of the storage components of the storage system after aggregation. (middle) anomaly score predicted by isolation forest. (bottom) anomaly score predicted by binary classifier. binary classification is a powerful tool to detect known anomalies. normal and anomalous measurements are considered as two classes. the classifier learns the separation surface copyright c© 2019 assa. adv syst sci appl (2019) 29 between them and is using it for anomalies detection for new measurements. similarly to the var models, several binary classifiers [17] are considered: logistic regression, naive bayes classifier, shallow neural networks, random forest and gradient boosting over decision trees classifiers. they are used to predict anomaly scores: ŝit = fi(yt, yt−1, ..., yt−k) (4.11) where ŝit ∈ [0, 1] is a predicted anomaly score. parameters of the classifiers are optimized using grid search to minimize the logistic loss function: l = − ∑ t (sit log ŝit + (1− sit) log (1− ŝit)) (4.12) where sit ∈ {0, 1} is a true anomaly score for t-th measurement of i-th time series; the classifier with the lowest value of the loss function wins and is used for the anomalies detection. the example of the binary classification work is shown in fig. 4.3. the first 110000 seconds are used to train the classifiers. 4.5. models comparison 0.00 0.01 0.02 0.03 0.04 0.05 false alarm rate 0.95 0.96 0.97 0.98 0.99 1.00 fa ilu re d et ec tio n ra te binary classification var (a) 0.0 0.1 0.2 0.3 0.4 0.5 false alarm rate 0.5 0.6 0.7 0.8 0.9 1.0 fa ilu re d et ec tio n ra te isolation forest + cusum isolation forest (b) fig. 4.4. dependencies of failure detection rate from false alarm rate for different anomaly detection algorithms: (a) binary classification and var; (b) isolation forest and isolation forest with cusum methods. we collected data with induced hardware failures in different storage system components. a failure is considered as detected when at least one parameter of the system demonstrates anomalous behaviour. time series for all parameters corresponded to one component are united into a group and used to estimate anomaly scores as described in previous sections. a measurement is anomalous when its anomaly score ŝt > τ , where τ is a threshold value. for a set of τ values failure detection rate (fdr) and false alarm rate (far) are calculated to measure the failures detection quality: fdr(τ) = 1 nanomalies ∑ i i[ŝi ≥ τ |si = 1] (4.13) far(τ) = 1 nnormal ∑ i i[ŝi ≥ τ |si = 0] (4.14) wherenanomalies in the total number of anomalous measurements;nnormal is the total number of normal measurements. dependencies of failure detection rate from false alarm rate copyright c© 2019 assa. adv syst sci appl (2019) 30 roc auc fdr (far = 5%) fdr (far = 1%) binary classification 0.999 0.997 0.991 var 0.997 0.987 0.975 isolation forest 0.755 0.0 0.0 isolation forest + cusum 0.956 0.852 0.836 table 4.1. quality metrics for the anomalies detection algorithms. for different anomaly detection algorithms are shown in fig. 4.4. these dependencies are also known as roc-curves. area under the roc-curve (roc auc) is widely used in machine learning to measure quality of classification. roc aucs, failure detection rates that correspond to false alarm rate values of 5% and 1% for different anomaly detection algorithms are presented in tab. 4.1. the results demonstrate that the binary classification method has the best quality. it detects larger than 99% of all anomalous states with false alarm rate below 5%. these results can be explained by the learning procedure of the method. during this procedure it recognizes separation surface between normal and anomalous states in the system parameters space. then it uses this surface to distinguish the system states on new measurements. var model shows similar results. it recognizes about 98% of all anomalous states with false alarm rate below 5%. however, unlike the binary classification it does not use any information about anomalies. the model takes only the system parameters measurements for normal states to learn normal behaviour of the system. any significant deviation from this behaviour considered as anomaly. this allows var to detect unknown anomalies. isolation forest demonstrates high false alarm rates. due to the noise in the parameters measurements isolation forest detects some normal measurements as anomalous and generates false alarms. it does not use any information about anomalies during the training procedure, that makes it harder to detect anomalies compared with binary classification. cusum test helps to reduce a number of false alarms and improves quality of the anomalies detection in a small far region as it is shown in fig. 4.4. 5. conclusion several approaches for automatic anomalies detection for data storage systems were tested and compared. these approaches are based on binary and one-class classification algorithms as well as on var models of time series analysis. the binary classification method has demonstrated the best detection quality. it allows to detect about 99% of anomalous states with false alarm rate of about 1%. the main limitation of this method is that it requires information about anomalies to learn how to detect them. this provides the high detection quality but makes it impossible to use the method for detection previously unseen anomalies. to resolve this limitation two additional methods were considered. these methods do not use information about anomalies during their training procedures. they are able to detect previously unseen anomalies. isolation forest shows very small failures detection rate in a region of false alarm rate below 5%. however, var model of time series analysis in combination with cusum test demonstrates promising results. it detects about 97% of anomalous states with false alarm rate is about 1%. the method learns normal behaviour of the system and detects any changes of it. in our future works we are going to continue developing methods of anomaly detection for data storage systems based on time series analysis methods. the goal is to investigate the methods that are able to recognize any anomalies without training procedure that uses information about these anomalies. copyright c© 2019 assa. adv syst sci appl (2019) 31 6. acknowledgements the research was carried out with the financial support of the ministry of science and higher education of russian federation within the framework of the federal target program research and development in priority areas of the development of the scientific and technological complex of russia for 2014-2020. unique identifier rfmefi58117x0023, agreement 14.581.21.0023 on 03.10.2017. 7. reference references 1. a. patcha and j.-m. park, “an overview of anomaly detection techniques: existing solutions and latest technological trends,” computer networks, vol. 51, no. 12, pp. 3448 – 3470, 2007. 2. v. chandola, a. banerjee, and v. kumar, “anomaly detection: a survey,” acm comput. surv., vol. 41, pp. 15:1–15:58, july 2009. 3. m. aiello, m. mongelli, e. cambiaso, and g. papaleo, “profiling dns tunneling attacks with pca and mutual information,” logic journal of igpl, vol. 24, p. jzw056, 09 2016. 4. e. cambiaso, m. aiello, m. mongelli, and g. papaleo, “feature transformation and mutual information for dns tunneling analysis,” in 2016 eighth international conference on ubiquitous and future networks (icufn), pp. 957–959, july 2016. 5. m. mongelli, m. aiello, e. cambiaso, and g. papaleo, “detection of dos attacks through fourier transform and mutual information,” in 2015 ieee international conference on communications (icc), pp. 7204–7209, june 2015. 6. t. shon and j. moon, “a hybrid machine learning approach to network anomaly detection,” information sciences, vol. 177, no. 18, pp. 3799 – 3821, 2007. 7. k. limthong, “real-time computer network anomaly detection using machine learning techniques,” journal of advances in computer networks, pp. 1–5, 01 2013. 8. j. hamidzadeh, m. zabihimayvan, and r. sadeghi, “detection of web site visitors based on fuzzy rough sets,” soft computing, 01 2017. 9. m. zabihimayvan, r. sadeghi, h. n. rude, and d. doran, “a soft computing approach for benign and malicious web robot detection,” expert syst. appl., vol. 87, pp. 129–140, nov. 2017. 10. m. zabihi, m. v. jahan, and j. hamidzadeh, “a density based clustering approach for web robot detection,” in 2014 4th international conference on computer and knowledge engineering (iccke), pp. 23–28, oct 2014. 11. q. lin, h. zhang, j.-g. lou, y. zhang, and x. chen, “log clustering based problem identification for online service systems,” in proceedings of the 38th international conference on software engineering companion, icse ’16, (new york, ny, usa), pp. 102–111, acm, 2016. 12. s. he, j. zhu, p. he, and m. r. lyu, “experience report: system log analysis for anomaly detection,” in 2016 ieee 27th international symposium on software reliability engineering (issre), pp. 207–218, oct 2016. 13. w. xu, l. huang, a. fox, d. patterson, and m. i. jordan, “detecting large-scale system problems by mining console logs,” in proceedings of the acm sigops 22nd symposium on operating systems principles, sosp ’09, (new york, ny, usa), pp. 117–132, acm, 2009. 14. c. a. c. rincn, j. pris, r. vilalta, a. m. k. cheng, and d. d. e. long, “disk failure prediction in heterogeneous environments,” in 2017 international symposium copyright c© 2019 assa. adv syst sci appl (2019) 32 on performance evaluation of computer and telecommunication systems (spects), pp. 1–7, july 2017. 15. j. li, x. ji, y. jia, b. zhu, g. wang, z. li, and x. liu, “hard drive failure prediction using classification and regression trees,” in 2014 44th annual ieee/ifip international conference on dependable systems and networks, pp. 383–394, june 2014. 16. m. borisyak, f. ratnikov, d. derkach, and a. ustyuzhanin, “towards automation of data quality system for cern cms experiment,” journal of physics: conference series, vol. 898, p. 092041, oct 2017. 17. t. hastie, r. tibshirani, and j. friedman, the elements of statistical learning. springer series in statistics, new york, ny, usa: springer new york inc., 2001. 18. f. t. liu, k. m. ting, and z.-h. zhou, “isolation forest,” in proceedings of the 2008 eighth ieee international conference on data mining, icdm ’08, (washington, dc, usa), pp. 413–422, ieee computer society, 2008. 19. b. koo, b. shin, and t. krijnen, “employing outlier and novelty detection for checking the integrity of bim to ifc entity associations,” in isarc 2017 proceedings of the 34th international symposium on automation and robotics in construction, 28 june 1 july, taipe, taiwan, pp. 14–21, international association for automation and robotics in construction i.a.a.r.c), 2017. 20. b. schölkopf, j. c. platt, j. c. shawe-taylor, a. j. smola, and r. c. williamson, “estimating the support of a high-dimensional distribution,” neural comput., vol. 13, pp. 1443–1471, july 2001. 21. r. sadeghi and j. hamidzadeh, “automatic support vector data description,” soft comput., vol. 22, pp. 147–158, jan. 2018. 22. j. hamidzadeh, r. sadeghi, and n. namaei, “weighted support vector data description based on chaotic bat algorithm,” appl. soft comput., vol. 60, pp. 540–551, nov. 2017. 23. r. j. hyndman, g. athanasopoulos, and otexts.com, forecasting : principles and practice. otexts.com [heathmont?, victoria], print edition. ed., 2014 2014. 24. yadro, “tatlin storage description.” https://yadro.com/products/ tatlin, 2019. 25. yadro, “vesnin server description.” https://yadro.com/products/ vesnin-tech, 2019. 26. e. s. page, “continuous inspection schemes,” biometrika, vol. 41, pp. 100–115, 06 1954. 27. l. koepcke, g. ashida, and j. kretzberg, “single and multiple change point detection in spike trains: comparison of different cusum methods,” frontiers in systems neuroscience, vol. 10, p. 51, 2016. copyright c© 2019 assa. adv syst sci appl (2019) multiagent model of people evacuation from premises while emergency adv syst sci appl 2019; 01; 98-115 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/558 multiagent model of people evacuation from premises while emergency andrey samartsev1, vladimir ivaschenko1, alexander rezchikov*1, vadim kushnikov1 , leonid filimonyuk1, aleksey bogomolov1 1) institute of precision mechanics and control of ras, saratov, russia e-mail: iptmuran@san.ru received may 17, 2018; revised march 13, 2019; published april 15, 2019 abstract: the multi-agent model of emergency evacuation process in case of limited space conditions is offered. a distinctive feature of the proposed model is the systematic consideration of physical interactions between agents and their reasonable behavior in conditions of evacuation from premises of a complex configuration. the model allows to investigate the state and behavior of people in emergency and determine the time of their evacuation. the proposed model can be used as a basis for determining ways and developing plans for evacuation from premises in emergency situations. keywords: multiagent model, intelligent agent, evacuation of people, premises, emergencies. 1. introduction currently, the problem of combating emergencies is becoming especially urgent. the statistics for 2016 presented by the ministry of emergency situations of russia [37] indicate significant human and material losses caused by emergencies. emergency evacuation of people from the premises as well as pre-emptive solutions [10, 26, 29-32, 40-45] are the most effective ways to reduce damage in case of accidents, catastrophes, fires and terrorist acts. in many situations, evacuation occurs spontaneously. this, among other things, is largely due to the inadequate development of existing organizational, technical and software tools for managing the evacuation process, the need to develop adequate models and algorithms to support decision making to ensure the effective evacuation of people in emergencies. in cases of spontaneous people evacuation from the premises, each person independently chooses a route and, on the way, makes decisions that they consider beneficial for achieving their goals. at the same time people decision making is influenced by the state and behavior of other people, the environment (walls and other obstacles), as well as such factors as panic [7, 8, 25], the failure of warning systems, unfamiliarity with the evacuation plan, the instinct of self-preservation, etc. therefore, the research of spontaneous evacuation is significant. in addition, the investigation of spontaneous evacuation makes it possible to take into account the behavior of people and identify bottlenecks, such as the maximum intensity of the flow of people through the exits from the premises, the places of their maximum accumulation, the capacity of the premises, etc., to organize the optimal control of the evacuation process [33, 34]. *corresponding author: rezchikov1939@mail.ru mailto:iptmuran@san.ru multiagent model of people evacuation from premises while emergency 99 copyright ©2019 assa. adv. in systems science and appl. (2019) thus, it is very important to ensure quick and unhindered evacuation in case of emergency. the quality of the evacuation process is usually checked by computer simulation, since the natural experiment in this situation can be dangerous and difficult to reproduce. in connection with the above, it is not surprising that the dynamics of the behavior of large masses of people, in particular during evacuation, has long been actively studied. the model of social forces, proposed by helbing and molnar in 1995, is widely used. its modifications have been actively proposed and are being investigated so far [2, 14-17, 24, 48]. in the classical interpretation of this model, people in the crowd are acted upon by forces on the one hand that attract them to the goal, and on the other hand forces that keep them away from each other and from obstacles at a certain distance. in addition, people can be affected by randomly changing force, designed to take into account the deliberate deviation of pedestrians from a given route. in general, the model is considered quite successful, although it is criticized for the fact that the underlying assumptions simplify the behavior of pedestrians. pedestrians in the model tend to adjust the speed of motion in inverse proportion to the distance from obstacles and other pedestrians. in fact, pedestrians can use better escape strategies, for example, "squeeze" between two other pedestrians to find a calmer route. in addition, the model is rather complicated in terms of the computation speed [47]. there are also fluid dynamic models [13, 27, 35, 47, 49]. they are quite accurate when modeling human flows of high (but not extremely high) density in the absence of panic. disadvantages of these models are that they give little information about the behavior of specific people, and only simulate the situation in general. models of this class are often difficult to understand intuitively, in addition, in practice, there are serious differences between the patterns of fluid movement and the panicking crowd. at the moment, the most promising are the agent models [3, 4, 7-9 12, 21, 23, 25, 35, 36, 39, 47, 49], in which the system is defined as a set of agents interacting according to certain rules. such models a priori determine the states of agents, their characteristics, as well as the rules of their interaction and behavior in various situations. agent behavior is understood as a function (or algorithm) that links the information that an agent receives about its local environment and the actions that it takes. interaction of agents in the process of evacuation can consist in avoiding collisions with each other or walls, in physical contact, in the spread of panic sentiments. note that the process of spontaneous evacuation is a vivid example of a decentralized system, when management is not carried out from the outside, and each person acts in accordance with their personal interests, and the global behavior of the system arises as a result of the many evacuees actions. in this case, the global behavior of the system may differ from the behavior of the elements of the system (this phenomenon is called emergence). for example, if each evacuation participant begins to hurry, the overall evacuation rate may be reduced due to a crush at the exit (the effect "faster is slower"). agent models make it possible to take into account the emergence effect to the full [47]. among the agent models that are suitable for describing the evacuation process, there are continuous and discrete ones. discrete models are based on the apparatus of cellular automata [1, 18-20, 35, 38, 47, 49]. such models are conceptually simple and computationally effective. however, in most cases, discrete models are an oversimplification of reality: pedestrians move discretely from their cells to their neighboring cells, the trajectories of their movement are angular and resemble the movements of a chess rook, collisions with each other and with walls are described rather roughly. nevertheless, by using such models, it is possible to explain some aspects of crowd behavior, for example, crushes at the exits and the effect "faster is slower". the most natural and convincing realization of the agent approach are continuous agent models. in such models, the characteristics of agents can take any value on some continuous segment. this approach requires more computational resources than discrete, but it allows to 100 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) achieve better modeling accuracy. continuous agent models are presented in [4, 6, 21, 22, 40, 46]. the model presented in this article belongs to that class either. the existing works of these classes are not without some drawbacks. some models are too complex in terms of the quantity of parameters [4] or the computational speed, for instance [2, 14-17, 21, 24]. many works do not consider evacuation from premises of complex shape [4, 14-17, 20, 21, 24, 38]. some models do not take into account the physical interaction of agents, their inertia [4, 21]. many models do not take into account the reasonableness of evacuated people, in particular the possibility of overtaking one agent by another [14-17]. a lot of models do not take into account the specifics of the evacuation process, but only model the behavior of "walking people" [21-24]. commercial models often do not reveal many details of the crowd behavior modeling process they use. the purpose of this article is to describe a model without all problems mentioned in previous paragraph, to develop software package based on the model and to simulate the evacuation process with use of the software package. the aim of simulation is to validate model by comparing model simulation results with ones of other models and statistic data, visually observing agents behavior and estimating of “reasonability” of their actions. 2. problem formulation without loss of generality let us consider the premises, the floor plan of which is shown in fig. 1. 1 people, 2 walls, 3 – exit zones fig. 1. premises map and people’s distribution inside it. in the xoy plane parallel to the floor of the premises, the following components are specified: 1. a finite set of walls, each of which is specified by parameters: (xw, yw) – coordinates of the lower-left corner of the wall, xw w and yw w – its width and length. we assume that the walls are perpendicular to the coordinate axes ox и oy, since this is the case in majority of the floor plans. 2. a finite set of exits from the premises specified by the parameters: (xe, ye) – coordinates of the lower left corner of the exit; xw e and yw e – its width and length, respectively. after a person got into one of the exit zones, it can be assumed that they multiagent model of people evacuation from premises while emergency 101 copyright ©2019 assa. adv. in systems science and appl. (2019) successfully evacuated. note here that the exit zone should not be directly in the doorway, but at some distance from it, since after the person leaves the room, they still influences on the evacuation process [5]. 3. a finite set of evacuated people, which we will consider as a set of their circle projections onto the plane xoy. we consider the centers of these circles as coordinates of people. 4. a finite set of evacuation zones, each of which is specified by the parameters: (xz, yz) – coordinates of the lower left corner of the evacuation zone, xw z и yw z – its width and length. at the very beginning of the evacuation people are concentrated within such zones. the presence of evacuation zones allows to take into account the possible absence of people in some rooms and in the space outside the premises at the first moment of time. parameters of people: – radius vector (coordinates) of the center of the circle around the projection of the person on the plane xoy – x(t); – radius of projection – r; – mass of a person – m; – vectors of its velocity – v(t) and acceleration – a(t); – the maximum possible speed – vmax and acceleration – amax. the speed and acceleration of the person at the beginning of the evacuation are assumed to be equal to zero. it is necessary to perform a simulation of the process of spontaneous people evacuation from the premises and to develop activities aimed at managing and speeding up this process. the activities include the possible creation, movement and expansion of certain building openings, adjustment of evacuation plans, and personnel instructions. these events must help to speed up people's movement along evacuation routes and avoid crushing. 3. multi-agent model 3.1. model construction in the proposed model, as in [2, 14-17, 22, 24], the law of agents’ motion is determined by the relations: .)()()( ;)()()( ttattt tttxttx     (1) where δt is model time step. the agents’ velocities change after collisions. the description of their physical interaction with each other and obstacles can be presented in various ways. it is permissible to regard collisions of agents as partially elastic [28]. during partially elastic collision some of the kinetic energy passes into thermal or other forms of energy. to describe the partially elastic collision, a recovery coefficient 0 ≤ ε ≤ 1 is introduced, which determines the nature of the interaction of the colliding bodies. for ε = 1 the collision is absolutely elastic, if ε = 0 it is absolutely inelastic, and for 0 < ε < 1 it is partially elastic. the normal components of the velocity to the shared tangent plane to the surfaces of the colliding bodies at the point of their contact (the collision plane) after impact at a partially elastic collision are calculated by the formulas: .)1( ;)1( 21 2211 22 21 2211 11 mm vmvm vu mm vmvm vu nn nn nn nn         (2) here v1n и v2n are the normal projections of the agent velocities to the collision plane before impact, u1n and u2n are the normal projections of the agent velocities to the collision 102 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) plane after impact, m1 and m2 are the masses of the colliding agents. tangential projections of the agent velocities to the collision plane after impact do not change [28]. when the agent collides with a wall, the projection of its velocity of motion parallel to the wall does not change, while the other projection reverses its sign and decreases its value according to the coefficient ε. for example, when an agent collides with a wall parallel to the ox axis, its speed is recalculated according to the relations ,; yyxx uu   where v = (vx, vy), u = (ux, uy) is the agent's velocity before the collision with the wall and after, respectively. let the direction of the acceleration vector a(t) of each agent, provided that each of them tends to get to the nearest exit from the room as fast as possible. to do this, we introduce the concept of the optimal vector (from the agent's point of view) of the velocity vopt(t), which approximates the agent to the chosen exit and allows, if possible, not to interfere with obstacles other agents and walls. the actual velocity of the agent v(t) may differ from the optimal one due to its collisions with other agents and walls. then we can assume that the agent moves with the acceleration a(t), whose modulus is determined from the physiological capabilities of the person (|a(t)| = amax), and the direction coincides with vopt(t) – v(t). if vopt(t) = v(t), then the agent moves without acceleration. fig. 2 shows a graphical interpretation of the relationship between agent parameters. here we should pay attention to the fact that the vectors a и vopt – v are collinear. fig. 2. graphical interpretation of connection between agent parameters the optimal speed module vopt(t) is bounded from above by the physiological capabilities of the agent vmax. if there are no obstacles ahead of the agent along the route, for example other agents or walls, then usually |vopt(t)| = vmax. to avoid collision with other agents along the route, the vopt(t) module may decrease. let us consider the approach to the choice of the vector vopt(t) in more detail. let e be a vector specifying the direction to the nearest exit, taking into account the layout of the room. the choice of vector e is determined solely by the location of the agent, walls and exit zones, but not by the location of other agents. thus, it is possible to determine the vector field e for a given premises map. in this case, the calculations of this field during the simulation are carried out only once, which is an advantage from the point of view of the computational speed. the field of the vector e for the premises map under study is shown in fig. 4. let lα – be the distance from the center of the agent to the nearest obstacle (human or wall) in the direction at an angle α to the vector e, l is the specified critical distance, and r is the radius of the person projection on the xoy plane. multiagent model of people evacuation from premises while emergency 103 copyright ©2019 assa. adv. in systems science and appl. (2019) then the modulus of the agent optimal velocity vector vαopt(α) in the direction at an angle α to the vector e can be calculated from the expressions: .,0)( ;,/)()( ;,)( rl rllrlrl rll opt maxopt maxopt          (3) people have different physiological capabilities, therefore it is necessary that each agent has individual values of vmax and аmax. let function        2 ; 2 -);cos()()(   optf (4) then the angle γ between the vectors vopt and e is defined as the angle at which the value of the function f(α) maximal. that is, the direction of the vector vopt of the agent's movement is chosen according to the criterion        2 ; 2 -and ifmax)(  f (5) the value of the vector vopt is determined by substituting α = γ into expression (3): |vopt |=vαopt(γ) (6) the presented way of choosing the value of vopt allows the agent to maneuver between other agents, trying firstly to avoid collision with them as much as possible, and secondly, to approach the nearest exit as quickly as possible. here we are assuming a good familiarity with the floor plan of the premises. thus, the choice of the optimal speed vector vopt takes into account both the layout of the room and the location of other agents around it. the vopt vector must be recalculated for each agent at each step of the model time. note that in the described model there are also attributes used in known models [4, etc.], such as: the angle and range of the agent's view of the surrounding space, the moment of inertia and the angle of rotation of the agent's head: 1. the direction that approximates the agent to the exit (taking into account the presence of walls in the room) is determined by the field of the vector e. the viewing angle of the agent can be considered equal to π, since in expression (5)        2 ; 2   , i.e. the agent analyzes all possible alternatives for movement in the plane in front of him. 2. since the agent calculates the optimal speed values based on the situation analysis within the critical zone, we can assume that the agent's viewing range is not less than l. 3. part of the energy in the impact goes into a rotational motion. in the model, these energy losses are taken into account exclusively by the recovery coefficient (formula (2)). 4. the ability of the agent to rotate the head is accounted for by the large viewing angle mentioned in 1. thus, in the model some parameters are set implicitly, which simplifies the model and increases the computational speed. 3.2 approach to model implementation 3.2.1 calculation optimizations as follows from the model (1) (6), for each agent for each step of model time, it is required to calculate distances to all other agents and walls of the room. proceeding from the calculated values, it is necessary to calculate the maximum possible velocities for a certain set of directions of motion of a given agent by the ratio (3) (16 possible directions are accepted with step π/8), and then by criteria (4) (6), determine the direction and the magnitude of the optimal speed vector, and then the acceleration of the agent. 104 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) the algorithm (for one step of model time) in this case will have computational complexity o(n2+nm), where n is the number of agents, and m is the number of walls. the quadratic complexity is not acceptable for systems with a large number of agents, so optimization is needed here, aimed to improve the asymptotic of the computational speed. let с = {cij | 0 ≤ i < n, 0 ≤ j < m} is the network that divides the space into cells with side h; n, m – number of cells horizontally and vertically, respectively; cij is a cell with coordinates (i, j), and w ⊂ c is the set of cells to which the wall corresponds. in the software implementation, the room was divided by the network c into cells of width h = 10 cm. this cell size is small enough for having centers of two agents inside at once. the following data structure is being supported during calculations: – each agent has the coordinates of the cell in which he is located; – each cell cij contains a reference to an agent whose center is in the given cell (in the absence of such an agent, the reference is null). each cij cell has the following attributes: – belonging to a wall or floor. – distance to the nearest exit dij; – vector eij, which specifies the direction to the exit. therefore, for a given agent, it is necessary to test a limited set of neighboring cells for the presence of walls and other agents, which makes it possible to improve the computational complexity of the algorithm to o(n). 3.2.2 calculation of the distances scalar field to the exit the attribute of the distance between exit and cij cell dij was calculated using the graph g = (v, e) given in the following way. suppose that for each cell cij  w here exists a vertex of the graph g vij  v, and eijkl  e is the edge of the graph between the vertices vij and vkl. then the set of edges e is defined as follows: e = {eijkl |∀a,b (a ∈ [min(i, k); max(i, k)] & & b ∈ [min(j, l);max(j, l)] → cab ∉ w) & (7) & (i, j) ≠ (k, l) & (i–k) ⊥ (j–l) & |i–k |≤ 2 & |j–l| ≤ 2)}. in addition, for each edge of the graph g, its length is given, which is proportional to the real distance between the cell centers corresponding to the vertices of the edge. the length lijkl of the edge eijkl is defined by the relation 22 )()( ljkilijkl  (8) consider the formula for defining the set of edges in more detail. 1. ∀a,b (a ∈ [min(i, k); max(i, k)] & b ∈ [min(j, l); max(j, l)] → cab  w) means that among cells cij and ckl, and also cells between them in a rectangle with a diagonal connecting cij and ckl, there are no cells corresponding to the walls. 2. (i, j) ≠ (k, l) means that the graph is not reflexive. there is no sense for loops since they can not change the value of the distance from the exit to any vertex. 3. (i–k) ⊥ (j–l) means that numbers (i–k) and (j–l) are coprime, i.e. these numbers have no common divisors except number 1. this relation allows to simplify the graph, since edges for which this relation is not satisfied, can not change the value of the distance from the exit to any vertex. 4. |i–k |≤ 2 & |j–l| ≤ 2 means that the distance vertically and horizontally between the cells is not greater than 2. thus, two vertices are connected by an edge only in cases where the corresponding cells: a) directly touch each other (they are vertical or horizontal neighbors); b) touch each other at one point (they are diagonal neighbors); c) are one cell far from each other by one dimension and two cell far from each other by the other (here an analogy with the l-shaped move of the chess knight can be drawn). multiagent model of people evacuation from premises while emergency 105 copyright ©2019 assa. adv. in systems science and appl. (2019) additionally, the connected cells and the cells that are between them should not correspond to the cells with the wall. fig. 3 shows a fragment of graph g superimposed on the part of the room shown in fig. 1 near the point with coordinates (15, 8). for clarity, in the picture edges, incident to two vertices are highlighted in red. one of these vertices is far from the walls and therefore has 16 incident edges. the second vertex is near the walls, so the number of edges incident to this vertex is less, and, in this case, is 9. fig. 3. a fragment of the graph g near the point with coordinates (15, 8) on a given graph, it is necessary to implement the algorithm for finding the shortest paths for finding the distance from the exit to all other cells. thus, for each vertex vij for which cij w, the distance dij to the nearest exit is determined. for cells cij∈w we assume that dij = ∞. in fig. 4 graphically shows the scalar field d of the distances to the exit zone: the darker areas of the room are located further away from the exit. also, fig. 4 shows the field of the vector e, which is directed toward decreasing values of the field d. 3.2.3 calculation of the vector field setting the direction to the nearest exit let us dwell in more detail on the vector e field calculation procedure. we introduce the graph g’, which differs from the graph g by the set of edges. the set of edges e' of the graph g' is defined as follows e’ = { eijkl | (i, j) ≠ (k, l) & (i–k) ⊥ (j–l) & |i–k |≤ 2 & |j–l| ≤ 2)}. (9) thus, when choosing a set of edges the presence of walls in the room is ignored. for the cell cij there are 16 possible directions of the vector eij with step π/8. these directions correspond to all the incident edges of the vertex vij in the graph g’. we number these edges from 0 to 15 in counterclockwise order starting from the edge parallel to the ox and denote them respectively by e0 … e15. for each of these directions, we calculate the ratio mk = dk/lk, where lk is the length of the edge ek, and dk is the difference between the values of the distance attribute to the exit between vertices incident to ek. thus, we obtain a vector m of 16 elements (m1 … mk). we introduce the vector m0 = (m01 … m0k), obtained from m by averaging the neighboring elements as follows: .10/)(5/)(5/2 16mod)18(16mod)14(16mod)17(16mod)15(0   kkkkkk mmmmmm (10) let the element m0s be minimal, m0s = min m0. then the angle between the vector e and the ox axis is defined as s * π/8. the field of the vector e calculated for the premises is shown in fig. 4. 106 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 4. scalar field d and vector field e 3.2.4 multiparticle collisions when the model is implemented in practice, the question of multiparticle collisions arises. two-particle collisions in the model are resolved by recalculating the velocities of the collided agents according to formula (2), provided that the distance between the centers of the agents is less than or equal to the sum of their radii, and the agents are approaching closer to each other. multiparticle collisions are resolved by recalculating the velocities of colliding agents according to formula (2) in pairs for all colliding agents in the order determined by the indexes of the colliding agents on the list. it should be noted that many-particle collisions, as a rule, arise only in the case of a crush. in its absence, such collisions are extremely rare, in view of the fact that in the model the time step is sufficiently small, and the collision of agents takes only one step of model time. 4. model usage results 4.1 program structure the software for implementing the multi-agent model is written in the java programming language using the swing gui library. here is the structure of the program for implementing the model: 1. set the initial data: the room dimensions, the coordinates of the walls, the exits from the room, the number and parameters of the agents, and the type of simulation results. 2. initializing of the network c: constructing the graph g or the network c, defining the scalar field d and the vector field e. 3. generation of agents within the evacuation zones. 4. the calculation of vopt for each agent (3 6). 5. the calculation of a for each agent. 6. the calculation of v for each agent (1). 7. the calculation of x for each agent (1). 8. collision check for each pair of agents. if the collision takes place, then the speed should be recalculated. multiagent model of people evacuation from premises while emergency 107 copyright ©2019 assa. adv. in systems science and appl. (2019) 9. check of collisions with walls for each agent. if there is a collision, then the speed should be recalculated. 10. verification of the exit condition for each agent. if the agent has reached the exit, then he is excluded from the list. it is considered that the agent has reached the exit if his center is in the exit zone. 11. collection and storage of statistics. if there are no people left in the room, go to step 14. 12. display the position of agents and walls on the screen in the program window. 13. go to the next step of the model time. go to step 4. 14. if the required number of intermediate experiments is not completed, then proceed to step 3. 15. displaying the results of the experiment (graphics of functions). 16. go to step 1 (user request). 4.2 modelling results computational experiments conducted on the software and information complex have shown that during evacuation agents demonstrate "reasonable" behavior. most vivid example of such behavior is overtaking of one agent by another (fig. 5). fig. 5. illustration of overtaking one agent by another (three stage of overtaking from left to right) in the situation there is no physical contact between the agents. the agent behind and having the maximum speed greater than the one of the agent moving in front of him, analyzes the situation around and chooses a route for overtaking. eventually he gets to the exit, not waiting for the agent in front and does not spend his energy on physical contact with the other agent. the other examples of “reasonable” behavior of the agent are avoidance of collision with walls, deceleration when approaching obstacles (walls or congestion of other agents), maneuvering between other agents and choice of the nearest exit from premises during evacuation. however, agents demonstrate quite poor level of strategic “reasoning”, for example they never try to choose other exit when the one they approached occasioned to be flooded by other agents. all types of that behavior could be visually observed via software package ui. when studying the evacuation process with a large number of evacuated, there was an accumulation of agents in the exit zone and a crush which were also observed in [2, 14-17, 20, 22]. as an example, we present the results of numerical simulation of the evacuation process from a premises measuring 20 × 10 meters, the plan of which is shown in fig. 1. the room contains 34 walls, each of which is specified by the parameters (x, y, xw, yw) in meters: (3.0; 3.0; 0.2; 10.0); (5.9; 3.0; 0.2; 1.0); (5.9; 5.0; 0.2; 2.0); (5.9; 5.0; 0.2; 2.0); (5.9; 108 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) 8.0; 0.2; 3.0); (5.9; 12.0; 0.2; 1.0); (7.9; 3.0; 0.2; 5.0); (8.9; 10.0; 0.2; 3.0); (11.9; 10.0; 0.2; 3.0); (15.9; 10.0; 0.2; 1.0); (15.9; 12.0; 0.2; 1.0); (17.9; 3.0; 0.2; 5.0); (17.9; 10.0; 0.2; 3.0); (19.9; 3.0; 0.2; 1.0); (19.9; 5.0; 0.2; 2.0); (19.9; 8.0; 0.2; 2.0); (20.9; 10.0; 0.2; 1.0); (20.9; 12.0; 0.2; 1.0); (22.9; 3.0; 0.2; 10.0); (3.0; 3.0; 3.5; 0.2); (7.5; 3.0; 11.0; 0.2); (19.5; 3.0; 3.5; 0.2); (3.0; 6.0; 3.0; 0.2); (20.0; 6.0; 3.0; 0.2); (7.9; 7.8; 5.1; 0.2); (15.0; 7.8; 3.1; 0.2); (3.0; 10.0; 4.0; 0.2); (8.0; 10.0; 2.0; 0.2); (11.0; 10.0; 1.0; 0.2); (16.0; 10.0; 0.5; 0.2); (17.5; 10.0; 1.0; 0.2); (19.5; 10.0; 3.5; 0.2); (3.0; 12.8; 10.0; 0.2); (15.0; 12.8; 8.0; 0.2). on the map, in 3 meters from the walls of the room the exit zones are set and specified by the parameters (x, y, xw e, yw e) in meters: (0; 0; 0.2; 16); (0.2; 0; 25.8; 0.2); (25.8; 0.2; 0.2; 15.8); (0.2; 15.8; 25.8; 0.2). the evacuation zone is defined by one element, which is specified by the parameters (x, y, xw z, yw z) in meters: (3; 3; 20; 10). the following numerical values of the evacuation process parameters were used: – number of evacuees (agents): 100; – their maximum speed vmax: random uniformly distributed quantity on a segment 1 – 2 m/s; – their maximum acceleration amax: randomly uniformly distributed quantity on the interval 1 – 2 m/s2; – the initial coordinates of each agent were chosen according to the uniform distribution law inside evacuation zones, proceeding from the condition that initially the projection of a person should not intersect the projections of other people and the walls; – radius of projection r: random uniformly distributed quantity on the segment 0,22 – 0,29 m; – critical distance: l = 2 m; – mass m: random uniformly distributed quantity on the segment 60 100 kg; – model time step: δt = 0,004 s; – coefficient of recovery: ε = 0,4 [22, 28]. the results of modeling the evacuation process are shown in fig. 6 9. the function hn(t) shows the dependence of the people inside premises number on time, provided that at the initial moment there are n people in the room. the function m nh is the result of averaging the function hn(t) over m implementations obtained during the simulation. the chart of function )( 100 100 th is shown on fig. 6. fig. 6. chart of function )( 100 100 th multiagent model of people evacuation from premises while emergency 109 copyright ©2019 assa. adv. in systems science and appl. (2019) the results showed that in the first few seconds of the evacuation process very few people leave the premises, since most of them are only at the beginning of their way to the exits. then, at the exits, crowds of people begin to form, at which time the intensity of the flow of people through the doors increases. by about the 17th second, the intensity of the human flow through the exits is weakening. this is due to the fact that at this time, agents with lower speed of movement begin to evacuate, as well as with uneven distribution of the load over the exits from the premises (evacuation through some exits by this time could have already stopped). the long "tail" on the chart of the function )( 100 100 th in the range of the argument values from 35 to 70 s is explained by the formation of crushes at the exits from the room for some realizations of the experiment. thus, with some initial arrangements of agents and the distribution of their parameters, evacuation can be delayed. dependency resemble to )( 100 100 th was also analyzed in [2, 11] and graphs of a similar shape were obtained both in computer simulation and in the analysis of statistical data obtained during the actual evacuation. the quantitative differences between our research and results in [2, 11] are explained by different initial conditions (premises map and the number of people). a computational experiment was also performed to analyze the dependence of the average evacuation time tm(n) from the number of people n in the room: dtth n nt nm )( 1 )( 0 100    (11) the simulation results are shown in fig.7. fig. 7. chart of function tm(n) as it can be seen from fig. 7, the time of evacuation of people increases with the number of people increase. some deviations from the monotonicity of the function tm(n) can be explained by the instability of the evacuation process in relation to the initial conditions, in particular, to the location of the agents in the premises at the very beginning of evacuation, and by the frequent formation of crushes. let the function vh(n), which is defined for integer values of the interval [1; h], here h is the number of people at the beginning of the evacuation, n is the agent number in the agent list sorted by the order in which they reach the premises exit, and the value of the function is the maximum speed vmax of the agent with the number n in the specified list. 110 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) let the function )(tv m h as the result of averaging the function vh(n) over the m implementations obtained during the simulation. chart of function )( 1000 50 tv is shown in fig. 8. fig. 8. chart of function )( 1000 50 tv as it can be seen in the fig. 8, agents with a higher maximum speed are usually evacuates faster. the maximum speed of agents, which managed to evacuate among the very first ones, is especially great, because the crushes have not yet formed in the openings, and such agents can use their speed advantage to the full extent. after that, the maximum agent speed is reducing, and, finally, the last agents with the lowest maximum speed are evacuated, who did not managed to take advantageous positions in the evacuation process and had to wait for the disappearance of crushes in openings. also, the analysis of the premises was carried out for the presence of places most dangerous in terms of quantity of the agents collisions with each other. in fig. 9 the heat map of the premises is presented where the most dangerous places are marked with a more saturated red color. as it can be seen, the most dangerous are positions before the openings and slightly to the left and to the right of the opening. this can be explained by agents attempts to wedge into the main flow before the opening and the intensity of this flow. multiagent model of people evacuation from premises while emergency 111 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 9. heatmap for collision possibility this map was constructed after modelling of 50 people evacuation from a given premises and averaging the results after 50 experiments. 5. conclusion the model of the people evacuation process from premises in emergency situations in conditions of limited space of complex configuration is offered, which allows to obtain a more adequate assessment of the evacuation process due to the system considering of physical interaction of people in the evacuation process and the procedure for reasonable correction of goals pursued by intellectual agents in the evacuation process. the simulation of evacuation process in case of emergency on the basis of the proposed model makes it possible to assess the state and behavior of people in the occurrence of emergencies in the premises, as well as the number of people evacuated from the premises with different floor plan and occupation, depending on the maximum possible values of speed and acceleration. the model also allows to research impact of agent parameters on the order of their evacuation and find the most insecure places in the premises, where agents collide with each other and crushes appear. strengths of the model include decent computational performance, account of agents’ physical interaction and complex premises map, focus on spontaneous evacuation process and possibility of computation parallelism. experiments have shown results close to the ones of known evacuation models, during them agents have demonstrated patterns of “reasonable” behavior. weaknesses of the current model implementation include absence of consideration of some factors, for example, agents’ ignorance of floor plan, agents’ verbal interaction, agents’ mutual assistance, other weakness is poor strategic “reasoning” in some cases. integration with emergency models is of scientific and practical interest, for that emergency influence rules on agents must be considered. the model can be used for buildings evacuation safety validation or as a part of simulator for emercom personnel training. 112 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) references [1] alizadeh, r. (2011). a dynamic cellular automaton model for evacuation process with obstacles. safety science, 49(2), 315-323 [2] aptukov a.m., bratsun d.a. & liushnin a.v. (2013). modelirovanie povedeniia panikuiushchei tolpy v mnogourovnevom razvetvlennom pomeshchenii [modeling the behavior of a panicking crowd in a multilevel branched room], komp'iuternye issledovaniia i modelirovanie, vol. 5, № 3, 491– 508, [in russian]. [3] antonini g., bierlaire m. & weber m. (2006). discrete choice models of pedestrian walking behavior // transportation research part b. – 2006. – vol. 40, № 8. – p. 667–687. [4] beklarian a.l. & akopov a.s. (2016). agentnaia model' povedeniia tolpy pri chrezvychainykh situatsiiakh [agent model of crowd behavior in emergency situations], xvi aprel'skaia mezhdunarodnaia nauchnaia konferentsiia po problemam razvitiia ekonomiki i obshchestva, higher school of economics publishing house, 665-681, [in russian]. [5] beklarian a.l. (2015). front vykhoda v modeli povedeniia tolpy pri chrezvychainykh situatsiiakh [exit front in crowd behavior models in emergency situations], vestnik tambovskogo universiteta. seriia: estestvennye i tekhnicheskie nauki. vol. 20, № 4, 851-856, [in russian]. [6] bosse, t., hoogendoorn, m., klein, m.c.a., treur, j., wal, c.n. et al.: modelling collective decision making in groups and crowds: integrating social contagion and interacting emotions, beliefs and intentions. auton agnts mult-agnt syst j, vol. 27, 52-84 (2013) [7] breer v.v. & novikov d.a. (2012). modeli upravleniia tolpoi [crowd control models], problemy upravleniia, № 2, 38–44, [in russian]. [8] breer v.v., novikov d.a. & rogatkin a.d. (2014). stokhasticheskie modeli upravleniia tolpoi [stochastic crowd control models], upravlenie bol'shimi sistemami, № 52, 85–117, [in russian]. [9] cherif f. & chighoub r. (2010). crowd simulation influenced by agent’s sociopsychological state, journal of computing, vol. 2, № 4. 48–54. [10] filimonyuk l. the problem of critical events’ combinations in air transportation systems // advances in intelligent systems and computing, springer international publishing, 2017. vol. 573. pp. 384-392. [11] galea e., deere s. & filippidis, l. (2012). the safeguard validation data set— sgvds1 a guide to the data and validation procedures. fire safety engineering group, university of greenwich london. [12] gwynne, s., galea, e. r., owen, m., lawrence, p. j., & filippidis, l. (1999). a review of the methodologies used in the computer simulation of evacuation from the built environment. building and environment, 34(6), 741-749 [13] henderson l.f (1971). the statistics of crowd fluids, nature, № 229, 381–383 [14] helbing d., farkas i. & vicsek t. (2000). simulating dynamical features of escape panic, nature, № 407, 487–490. [15] helbing d. (2001). traffic and related self-driven many-particle systems, reviews of modern physics, vol. 73, № 4, 1067–1141. multiagent model of people evacuation from premises while emergency 113 copyright ©2019 assa. adv. in systems science and appl. (2019) [16] helbing d., farkas i., molnar т. & vicsek t. (2002). simulation of pedestrian crowds in normal and evacuation situations, pedestrian and evacuation dynamics, vol. 21, № 2, 21–58. [17] helbing d., johansson a. & al-abideen h.z. (2007). dynamics of crowd disasters: an empirical study, physical review, vol. 75, № 4, 0461091–0461097. [18] kirchner, a., klüpfel, h., nishinari, k., schadschneider, a. & schreckenberg, m. (2003). simulation of competitive egress behavior: comparison with aircraft evacuation data. physica a: statistical mechanics and its applications, 324(3-4), 689-697 [19] kirchner, a., nishinari, k., & schadschneider, a. (2003). friction effects and clogging in a cellular automaton model for pedestrian dynamics. physical review e, 67(5), 056122. [20] kirik e.s., kruglov d.v. & iurgel''ian t.b. (2008). o diskretnoi modeli dvizheniia liudei s elementom analiza okruzhaiushchei obstanovki [on a discrete model of people's movement with an element of analysis of the environment], zhurnal sfu. seriia «matematika i fizika», vol. 1, № 3, 266–276, [in russian]. [21] korepanov v.o. (2009). imitatsionnye modeli takticheskogo povedeniia agentov [simulation models of tactical behavior of agents], upravlenie bol'shimi sistemami. № 26, 66–76, [in russian]. [22] korhonen t. (2016). fire dynamics simulator with evacuation: fds+evac technical reference and user’s guide (fds 6.5.2, evac 2.5.2), vtt technical research centre of finland. [23] krasnoshchekov p.s. (1998). prosteishaia matematicheskaia model' povedeniia. psikhologiia konformizma [the simplest mathematical model of behavior. psychology of conformism], matematicheskoe modelirovanie, vol. 10, № 7, 76– 92, [in russian]. [24] moussaida m., helbing d. & theraulaza g. (2011). how simple rules determine pedestrian behavior and crowd disasters, pnas, vol. 108, № 17, 6884–6892. [25] novikov d.a. (2015). modeli informatsionnogo protivoborstva v upravlenii tolpoi [models of information confrontation in crowd management], problemy upravleniia, № 3, 29-39, [in russian]. [26] petrova m.s., petrov s.v. & vol''khin s.n. (2006). okhrana truda na proizvodstve i v uchebnom protsesse [labor protection at work and in the educational process], enas, 152 s, [in russian]. [27] predtechenskii v.m. & milinskii a.i. (1979). proektirovanie zdanii s uchetom organizatsii dvizheniia liudskikh potokov [designing of buildings taking into account the organization of movement of human flows], stroiizdat, 375 s, [in russian]. [28] putilov k.a. (1963). kurs obshchei fiziki. t. 1. mekhanika. akustika. molekuliarnaia fizika. termodinamika. izd. 11. [course of general physics. vol. 1. mechanics. acoustics. molecular physics. thermodynamics. ed. 11], gos. izd-vo fiz.-mat. lit-ry, 560 s, [in russian]. [29] rezchikov a., kushnikov v., ivaschenko v.a., bogomolov a., filimonyuk l., kachur k. (2016). control of the air transportation system with flight safety as 114 a. samartsev, v. ivaschenko, a.rezchikov, v.kushnikov et al. copyright ©2019 assa. adv. in systems science and appl. (2019) the criterion. automation control theory perspectives in intelligent systems, springer. [30] rezchikov a., kushnikov v., ivaschenko v.a., bogomolov a., filimonyuk l., et al. (2017) the approach to provide and support the aviation transportation system safety based on automation models, springer. [31] rezchikov a., dolinina o., kushnikov v., ivaschenko v., kachur et al. (2016). the problem of a human factor in aviation transport systems, indian journal of sciences & technology. [32] rezchikov, a.f., kushnikov, v.a. & ivashchenko, v.a. (2017). prevention of critical events combination in robotic welding. journal of machinery manufacture and reliability. №4 370-379. [33] samartsev a.a., rezchikov a.f. & ivashchenko v.a. (2016). ispol'zovanie dinamicheskikh planov evakuatsii pri pozhare na promyshlennykh predpriiatiiakh [use of dynamic evacuation plans in case of fire in industrial plants], matematicheskie metody v tekhnike i tekhnologiiakh mmtt-29 v.3, 177-181 [in russian]. [34] samartsev a.a., rezchikov a.f. & ivashchenko v.a. (2016). podkhod k minimizatsii vremeni evakuatsii liudei iz pomeshchenii promyshlennogo predpriiatiia [the approach to minimizing the time of evacuation of people from the premises of an industrial enterprise], komp'iuternye nauki i informatsionnye tekhnologii: materialy mezhdunar. nauch. konf, saratov: nauka, [in russian]. [35] santos, g. & aguirre, b. e. (2004). a critical review of emergency evacuation simulation models. [36] skobelev p.o. (2010). mul'tiagentnye tekhnologii v promyshlennykh primeneniiakh: k 20-letiiu osnovaniia samarskoi nauchnoi shkoly mul'tiagentnykh sistem [multiagent technologies in industrial applications: to the 20th anniversary of the founding of the samara scientific school of multi-agent systems], mekhatronika, avtomatizatsiia, upravlenie, № 12, 33-46, [in russian]. [37] sravnitel'naia kharakteristika chrezvychainykh situa-tsii, proisshedshikh na territorii rossiiskoi federatsii v 2016/2015 godakh [comparative characteristics of emergency situations that occurred on the territory of the russian federation in 2016/2015] [online]. available http:// http://www.mchs.gov.ru/folder/33160904, [in russian]. [38] stepantsov m.e. (2004). matematicheskaia model' napravlennogo dvizheniia gruppy liudei [mathematical model of directed movement of a group of people], matematicheskoe modelirovanie, vol. 16, №3, 43–49, [in russian]. [39] tarasov v.b. (1998). agenty, mnogoagentnye sistemy, virtual'nye soobshchestva: strategicheskoe napravlenie v informatike i iskusstvennom intellekte [agents, multi-agent systems, virtual communities: a strategic direction in in-formatics and artificial intelligence], novosti iskusstvennogo intellekta, № 2, 5-63, [in russian]. [40] tupikov d.v. & ivashchenko v.a. (2014). neirosetevoe prognozirovanie znachenii faktorov vozniknoveniia pozhara na proizvodstvennykh ob"ektakh [neural network forecasting of the values of the factors of occurrence of a fire at multiagent model of people evacuation from premises while emergency 115 copyright ©2019 assa. adv. in systems science and appl. (2019) production facilities], matematicheskie metody v tekhnike i tekhnologiiakh mmtt-27, v.3, 59-61 [in russian]. [41] tupikov d.v. & ivashchenko v.a. (2015). podkhod k obespeche-niiu pozharnoi bezopasnosti na promyshlennykh predpriiatiiakh [approach to ensuring fire safety in industrial premises], matematicheskie metody v tekhnike i tekhnologiiakh mmtt-28, vol. 3, 51-54, [in russian]. [42] tupikov d.v. & ivashchenko v.a. (2013). razrabotka bazy znanii dlia operativnogo upravleniia vzryvoi pozharo-opasnym proizvodstvom [development of a knowledge base for the operational control of explosive and fire hazardous production], vestnik saratovskogo gosudarstvennogo tekhnicheskogo universiteta, № 3 (72), 133-137, [in russian]. [43] tupikov d.v. & rezchikov a.f. (2014). podkhod k postroeniiu nechetkoi bazy znanii dlia opredeleniia pozharoopasnykh situatsii [approach to constructing a fuzzy knowledge base for determining fire hazard situations], matematicheskie metody v tekhnike i tekhnologiiakh mmtt-27, vol. 3, 125-127, [in russian]. [44] tupikov d.v., rezchikov a.f. & ivashchenko v.a. (2014). algoritm podderzhki priniatiia reshenii po ustraneniiu pozharoopasnykh situatsii na promyshlennykh predpriiatiiakh [algorithm for supporting decision-making on the elimination of fire-hazardous situations in industrial enterprises], upravlenie bol'shimi sistemami, №52, 148-163 [in russian]. [45] tupikov d.v. rezchikov a.f. & ivashchenko v.a. (2014). podkhod k prognozirovaniiu znachenii faktorov pozharo-opasnykh situatsii [approach to predicting the values of fire hazard factors], mekhatronika, avtomatizatsiia, upravlenie, № 7, 48-51, [in russian]. [46] van der wal, c. n., formolo, d., robinson, m. a., minkov, m., & bosse, t. (2017). simulating crowd evacuation with socio-cultural, cognitive, and emotional elements. in transactions on computational collective intelligence xxvii (pp. 139-177). springer, cham. [47] winter h. (2012). modelling crowd dynamics during evacuation situations using simulation, lancaster university. [48] yu w. & johansson a. (2017). modeling crowd turbulence by many-particle simulations, physical review, vol. 76, № 4, 046105. [49] zheng, x., zhong, t., & liu, m.: modeling crowd evacuation of a building based on seven methodological approaches. building and environment, 44(3), 437-445 (2009) adv syst sci appl 2019; 01; 12-30 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/651 evaluating environmental impacts of photovoltaic technologies using data envelopment analysis svetlana v. ratner*1, andrey v. lychev2 1) trapeznikov institute of control sciences of russian academy of sciences, moscow, russia, e-mail:lanarat@mail.ru 2) national university of science and technology misis, moscow, russia e-mail: lychev@misis.ru received september 29, 2018; revised february 20, 2019; published april 15, 2019 abstract. this study contributes to the literature by proposing a new method of complex evaluation of multiple life cycle environmental impacts of different pv technologies based on the data envelopment analysis (dea). the main advantage of dea as a non-parametric technique is that it does not require prior knowledge of underlying production functions. an empirical production technology frontier is estimated based on best-practice boundary of the input-output relationship. dea evaluates comparative or relative efficiency, which means the measurement with reference to some set of units we are comparing with each other. the proposed approach allows to aggregate disparate quantitative estimates of individual negative environmental effects from the literature and special databases in a transparent and easily understandable index or coefficient of ecology efficiency. the evaluation of environmental effects is performed on data from the ecoinvent database. the results of this study clearly show that from an environmental point of view it is more practical to prefer technologies, which are less resource and energy intensive in manufacturing and upstream activities. as of right now, this requirement is met by thin-film technologies: amorphous silicon (a-si), cadmium telluride (cdte), and copper-indiumdiselenide (cis); however, their ecologic efficiency evaluation may change as we obtain more data on the final stages of the lifecycle for pv modules of various types. our computational results show that 5 out of 19 pv technologies were identified as efficient compared to the other technologies against which they were assessed. cdte, a-si, and ribbon silicon laminated panels installed slanted-roof, cis panel mounted slanted-roof, and open ground installation m-si are demonstrated highest efficiency scores from an ecological point of view throughout the life cycle. based on results of this study a number of opportunities for improving existing government incentives and rationalizing the design of state programs under elaboration can be identified. keywords: environmental impact categories; solar power; photovoltaic system; life cycle assessment; data envelopment analysis. 1. introduction throughout the last decades, the global sales and installation of photovoltaic (pv) systems have grown rapidly. annual installations of pv systems reached a record 98 gw in 2017, while global total installations attained 402 gw [37]. although pv-technologies have very low environmental and human health impacts compared to conventional electricity generation, the processes of manufacturing, transportation, installation, and disposal of pv-modules are associated with significant energy consumption, usage of working fluids containing chlorates and nitrites, formation of sewage and other negative environmental effects that need to be considered. at present, several photovoltaic technologies are being developed in the world: monocrystalline silicon, polycrystalline (multicrystalline) silicon and thin-film. it is well known, * corresponding author: lanarat@mail.ru mailto:lychev@misis.ru s.v. ratner, a.v. lychev 13 copyright ©2019 assa adv. in systems science and appl. (2019) that silicon wafer manufacturing process is very energy intensive. mono-crystalline cells require up to 1000 kwh/kg-si and poly-crystalline cells’ production needs up to 700 kwh/kg-si [11]. manufacturing of silicon-based pv cells is accompanied by emissions of such hazardous materials as silica dust, silanes, diborane, phosphine, and solvents. heterojunction solar cells (shj) are produced from silicon wafers in a low-temperature process that does not exceed 200 °c, thus, the energy intensity of the manufacturing process and, consequently, its environmental impact should be lower. besides, a low-temperature coefficient of shj cells leads to higher energy yields compared to conventional c-si modules [29]. however, there is as of yet no precise data on the environmental effects of this technology over the whole life cycle due to its relative novelty. it is also known, that mining cadmium and manufacturing cdte solar cells can cause occupational health risks. some questions also remain regarding the toxicity of pv modules’ disposal [18,19]. therefore, the environmental issues of cadmium telluride pv cells and module production and utilization are still under discussion. tellurium is also considered as a rare earth element, so the problem of resource constraints must also be taken into consideration when choosing this technology [36]. the main negative environmental impacts of manufacture of copper indium gallium selenide (cigs) photovoltaic cells, as is reported in [2] are connected with up-stream activities including extraction and processing of primary materials. for example, the extraction and processing of silver, which is needed for stringer and screen printing processes, is accompanied with emissions of toxic heavy metals such as mercury, lead, and arsenic. the mining of copper (used in the cable for the balance of the system and in co-production of selenium, gallium, and indium) is potentially associated with the exposure of radioactive materials. such manufacturing processes as energy-intensive co-evaporation needed for to production of the cigs layer and water-intensive surface washing of the cell’s substrate also affect negatively on the environment. another issue that needs to be considered is the scarcity of indium, which can restrict the production process [34]. until recently, the choice between industrial developments of a particular technology was carried out, mainly, by two criteria: cost and efficiency (or energy conversion coefficient). however, with the growth of production capacities around the world, the ecological aspects, such as the emission of pollutants by production facilities into the air, water and soils, the consumption of rare earth metals, water, energy, etc., become more important when choosing a particular technology for development and support through multiple government policies. this problem become more relevant for russia with the booming developing of pv-industry in recent years [56, 33]. with the enactment in 2013 of new government’s renewable energy support scheme focusing on grid-connected power generation facilities more than 5 mw, the process of construction and connection of such facilities to the federal grid company has notably intensified. in 2013–2018 period, more than 100 solar generation projects of a total capacity of 1.852 gw were selected for a subsequent support (fig. 1). https://en.wikipedia.org/wiki/copper_indium_gallium_selenide_solar_cells https://en.wikipedia.org/wiki/copper_indium_gallium_selenide_solar_cells 14 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) source: own calculations fig. 1. planed pv-installations on the wholesale market supported by the state program in the framework of the government decree no. 449 (28 of may, 2013) an important feature of the renewable energy support programs introduced in russia is the high production localization index, which must be achieved in the project in order to be selected for financial support. such support conditions are aimed at creating in the country own full-cycle production of photovoltaic modules and other types of equipment for renewable energy. the active development of new pv-manufacturing technologies and the current and future increase in production volumes boost the relevance of the problem of estimating and predicting the possible negative environmental effects of photovoltaics at all stages of the life cycle: production, operation, and disposal. currently, life cycle assessment (lca) followed the international standards organization (iso) 14040 series is a mature methodology and a valuable tool for providing a comprehensive “cradle-to-grave” view of the environmental loads of a technology. it is often used for comparative analysis of manufacturing alternatives and optimization of product system design. due to the fact, that a full life cycle assessment is usually a time-, energy-, and data-intensive process requiring sophisticated methodology, many scholars prefer to use a simplified approach and quantify only greenhouse gas (ghg) performance of the technology or product under consideration, though for each product this specific type of environmental influence is the most significant. current papers on life cycle assessment of pv technologies can be divided into three groups. the papers of the first group focus on detailed analysis of one or several types of negative ecologic effects of specific types of pv-technologies. for example, hsu et al. [22] shows an analysis of ghg emissions from solar c-si pv lca. kim [25] evaluates the same type of environmental impact from thin-film pv lca. the research [20] analyzes and obtains estimates for сo2 emissions and embedded energy of a hybrid photovoltaic–thermal module. fthenakis [19] evaluates atmospheric cd emissions from the life-cycle of cdte models. the [28] research shows a very wide set of environmental impacts for new marine pv-technologies, including besides ghg emissions such categories as acidification potential, eutrophication, ecotoxicity, human toxicity, depletion of abiotic resources, terrestric ecotoxicity potential, and photochemical ozone creation potential. results of these papers improve upon the existing knowledge on the influence of new technologies on the environment. the second group compares multiple pv-technologies between each other or between other renewable energy technologies based on a specific category on environmental influence. most of the time the researched category is the ghg emissions. for example, the study [2] compares in detail the environmental influence from the production of two thin-film photoelements based on 35,2 140 199 255 285 308 373 132 125 0 50 100 150 200 250 300 350 400 2014 2015 2016 2017 2018 2019 2020 2021 2022 m w s.v. ratner, a.v. lychev 15 copyright ©2019 assa adv. in systems science and appl. (2019) cu(in,ga)se2: in particular, whether the use of zinc oxysulfide (zn(o,s)) or cadmium sulfide (cds) minimizes the ghg emissions during the production of photoelements through screen printing and stringer. paper [3] summarizes the results of over 75 researches performed via life cycle assessment (lca) and focusing on ghg emissions for electricity and heat generation from photovoltaic, solar thermal, onshore and offshore winds, hydropower, marine technologies (wave power and tidal energy), geothermal, biomass, waste, and heat pumps. the results of these papers may be used to improve the production chain, choose better designs for production systems, as well as make decisions on regional and national government levels to support one type of technology over another. however, for a full comparison of the ecologic influence of different technologies one must evaluate not just one (perhaps important) category of environmental influence, but the entire range of negative ecologic effects. hence, in modern literature the amount of works, which compare pv-technologies against each other or other renewable energy tech with multiple environmental influence categories in mind, is constantly increasing. furthermore, these comparisons can be performed both separately and using various aggregate indicators. for example, [27] low-concentration pv and conventional pv are compared separately in the climate change (ghg emissions), acidification potential, eutrophication potential, human toxicity, and ozone layer depletion categories. the same approach is used in [30], where mc-si, ingap and ingap/mc-si solar modules are compared in climate change, abiotic resource depletion, acidification, human toxicity, and fresh water ecotoxicity separately. in [41] one can find a comparative analysis of environmental impacts of organic and conventional pv-technologies completed using recipe v1.0.5 midpoint (h) impact categories. celik et al. [7] shows an aggregate indicator of toxicity (toxicity for humans, ecotoxicity of sea and freshwater, ecotoxicity of sea and freshwater bottom sediments), which is used to compare conventional si pv technologies and perovskite solar cells. paper [34] investigates ghg emissions, embedded energy and presents an aggregate indicator for environmental influence: the normalized environmental impact (eco-points) of several popular silicon-based pv technologies. in the case that several environmental effects are taken into consideration for comparison of designs of products or manufacturing processes, another important research question arises: which of these negative environmental impacts are more important and how should it be accounted for? this study contributes to the literature by proposing a new method of complex evaluation of multiple life cycle environmental impacts of different pv technologies based on the data envelopment analysis (dea). dea is a methodology for evaluating the relative performance of a set of homogeneous units, which transform multiple inputs into multiple outputs [9]. the main advantage of dea as a non-parametric technique is that it does not require prior knowledge of underlying production functions. an empirical production technology frontier is estimated based on best-practice boundary of the input-output relationship. dea evaluates comparative or relative efficiency, which means the measurement with reference to some set of units we are comparing with each other. the efficiency score has a clear economic meaning. it shows how much the dmu should reduce its resources or increase its outcome to become efficient. since dea was first introduced, there have been a large number of papers written on dea or applying this methodology on various sets of managerial, operational and economic problems. today dea is not only a tool for technical-efficiency analysis but also an increasingly popular performance management tool for benchmarking in many areas of research including energy sector [15]. the evaluation of environmental effects is performed on data from the ecoinvent database which is currently one of the most reliable and complete information resources for a significant number of industrial products. the paper’s added value as compared to above-mentioned research consists in the following: a) the work takes into account a much wider range of negative environmental effects; b) the developed method of their complex comparison allows to rank the technology according to the degree of aggregated environmental impact, as well as to determine the target parameters for reducing negative environmental effects. first attempts to apply dea to the problem of choosing the most ecologically efficient solar power technologies has been made https://www.sciencedirect.com/science/article/pii/s0927024816300605#! https://www.sciencedirect.com/science/article/pii/s2212827117308831#! 16 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) in an earlier work of one of the authors [35], and it is thanks to that we could prove the applicability of this approach to compare different photovoltaic technologies from the environmental perspective. however, during the last year data on photovoltaic technology has significantly expanded and updated thanks to progress both in r&d and in production. besides that, the aforementioned paper has shown a suboptimal choice of indicators of environmental impact: some impact categories haven’t been accounted for in the analysis. in [35] we also compared ecologic impact of photovoltaic technology without considering the method of solar panel installation which could have distorted the results of the analysis. this paper addresses and eliminates these issues, and considers the newly-attained results from a perspective of policy applications. the remainder of the paper is organized as follows: section 2 provides an overview of lca methods, used in the ecoinvent database, and explains the choice of datasets and impact categories, which we used in the evaluation of pv-technologies. it also provides insight into the methodology (dea). in section 3 we discuss the results of the dea-based evaluation of modern pv technologies and the possibilities to develop a proposed approach for management applications as well as its limitations. section 4 concludes and offers policy recommendations. 2. data and methodology in rest of the paper we use the following abbreviations for pv technologies and types of pvinstallations: pv technologies a-si amorphous silicon cdte cadmium telluride cis copper-indium-diselenide m-si multicrystalline silicon r-si ribbon silicon s-si singlecrystalline silicon pv-installations fi facade installation fri flat-roof installation li integrated laminate ogi open ground installations pm mounted panel sri slanted-roof installation in this study, we used the data from the ecoinvent database (a non-profit association of research organizations in switzerland) for assessment of the environmental impact of the photovoltaic life cycle. currently, ecoinvent is the world’s leading life cycle assessment (lca) database compliant with the iso 14040-14044 standards and contains life-cycle data sets of more than 12,800 products and services [31]. it is important to emphasize that the ecoinvent database is not simply a library of individual lca datasets. the data are interrelated in such a way that all semifinished products, input streams, electricity consumption, demand for raw materials and materials and equipment requirements depend on sub-processes of production and delivery of semi-finished products and services for product redistribution. thus, lca results are calculated in a matrix system, so that any update in one set of process data will affect the accumulated lca results in all other connected data sets. version 3.3 (2016) of the database contains 356 datasets on photovoltaic energy, from which 227 datasets are devoted to electricity production, 43 – to transportation of semi-finished products such as mounted systems, cells, panels, modules, wafers and single crystals of silicon, 3 – to cell and panel factory contraction, 32 – to installations, 10 – to cell/panel production, 12 – to laminate production, 2 – to single wafer production, 2 – to single crystal production, 10 – to panel production, 7 – to mounting system production and 4 – to treatment of waste form silicon cells and s.v. ratner, a.v. lychev 17 copyright ©2019 assa adv. in systems science and appl. (2019) panel production. each data set of environmental effects includes effects from all upstream activities under evaluation, therefore the consideration of last activity in a production chain is sufficient. there are no datasets in the database devoted to disposal and recycling of pv models and panels because even the earliest pv-installations considered in the database are still in operation. lifecycle ends with low voltage electricity produced with the 3 kwp module, assuming an average yield. the 3 kwp module has been chosen in most of the datasets devoted to electricity production as the basic module for building integrated pv electricity production, due to the fact that larger modules can easily be built with 3 kwp modules without producing a significant error in environmental impact calculations. the datasets represent several mature pv technologies: multicrystalline silicon (m-si), singlecrystalline silicon (s-si), amorphous silicon (a-si), cadmium telluride (cdte), copper-indium-diselenide (cis), and ribbon silicon (r-si). the modules made of m-si, s-si, a-si, and r-si can be installed of different parts of the building (facade installation (fi), flat-roof installation (fri) or slanted-roof installation (sri)) as an integrated laminate (li) of a mounted panel (pm). the modules made of cdte can be installed in the form of integrated laminate only while cis models only in the form of a mounted panel. besides this, the database contains several datasets for 570 kwp m-si open ground installations (ogi). thus, we have different production chains specified by the cells and module production technologies, form and place of installation. the referent product in all datasets on electricity production is low voltage 1 kwh1. the simplifying assumption was made in all datasets that all electricity is directly provided as low voltage electricity so that no transformation has to take place. the lifetime of all pv modules is supposed to be 30 years. in order to compare energy technologies, we need to choose the datasets, yield (or calculated by extrapolation of field data) for the same geographic location, which is characterized by the value of specific solar yield in kwh per kilowatt peak installed. since the most reliable statistical data in the database is collected for switzerland (the longest period of observation), we chose data on various options (19 total) for electricity production using photovoltaic installations in this geographic location (table 1). one should note that this dataset can be used for comparison of energy technologies only, but not for assessment of average production patterns in different geographic locations, for example, in russia. table 1. comparing electricity production options for switzerland option period of observation 3kwp fi m-si li 2005-2016 3kwp fi m-si pm 2005-2016 3kwp fi s-si li 2005-2016 3kwp fi s-si pm 2005-2016 3kwp fri m-si 2005-2016 3kwp fri s-si 2005-2016 3kwp sri a-si li 2005-2016 3kwp sri a-si pm 2005-2016 3kwp sri cdte li 2005-2016 3kwp sri cis pm 2005-2016 3kwp sri m-si li 2005-2016 3kwp sri m-si pm 2005-2016 3kwp sri m-si pm label-certified 2010-2016 3kwp sri r-si li 2005-2016 1 although it is known that large central photovoltaic power stations in the higher kilowatt to megawatt range can feed directly into the mediumor high voltage grid, in order to treat all photovoltaic installations the same way, low voltage electricity is assumed as a product in this dataset 18 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) 3kwp sri r-si pm 2005-2016 3kwp sri s-si li 2005-2016 3kwp sri s-si pm 2005-2016 3kwp sri s-si pm label-certified 2010-2016 570kwp ogi m-si 2008-2016 life cycle impact assessment can be carried out using several indicator methods, which differ from each other in the spectrum of considered environmental impact categories. the latest version of ecoinvent uses 13 basic methods, of which, according to our opinion, cml2001 is the most complete and informative. this method was developed by center of environmental science of leiden university and published in 2001 as a new operational guide for implementation of iso environmental management standards [21]. in addition to such basic categories of environmental impact as greenhouse gas emissions, this method takes into account many other negative environmental effects. according to [8] it gives more accurate estimates of chemical impacts on human health. main impacts on the environment, according to cml2001 method can also be considered at different time horizons. table 2 summaries impact categories, covered in cml2001, their indicators and measure units. table 2. cml2001 environmental impact categories and their indicators impact category group name of the impact category in the method unit acidification acidification potential average europe acidification potential – generic kg so2-eq climate change climate change gwp100 climate change gwp20 climate change gwp500 climate change lower limit of net gwp100 climate change upper limit of net gwp100 kg co2-eq depletion of abiotic resources depletion of abiotic resources elements, economic reserve depletion of abiotic resources elements, reserve base kg antinomy-eq ecotoxicity freshwater aquatic ecotoxicity faetp infinitive freshwater aquatic ecotoxicity faetp100 freshwater aquatic ecotoxicity faetp20 freshwater aquatic ecotoxicity faetp500 freshwater sedimental ecotoxicity fsetp infinitive freshwater sedimental ecotoxicity fsetp100 freshwater sedimental ecotoxicity fsetp20 freshwater sedimental ecotoxicity fsetp500 marine aquatic ecotoxicity maetp infinitive marine aquatic ecotoxicity maetp100 marine aquatic ecotoxicity maetp20 marine aquatic ecotoxicity maetp500 marine sedimental ecotoxicity msetp infinitive marine sedimental ecotoxicity msetp100 marine sedimental ecotoxicity msetp20 marine sedimental ecotoxicity msetp500 kg 1,4-dcb-eq s.v. ratner, a.v. lychev 19 copyright ©2019 assa adv. in systems science and appl. (2019) terrestrial ecotoxicity tetp infinitive terrestrial ecotoxicity tetp100 terrestrial ecotoxicity tetp20 terrestrial ecotoxicity tetp500 eutrophication eutrophication generic eutrophication average europe po4-eq nox-eq human toxicity human toxicity htp infinitive human toxicity htp100 human toxicity htp20 human toxicity htp500 kg 1,4-dcb-eq ionising radiation radiation dalys land use land use land competition m2a odour odour (malodours air) m3 air ozone layer depletion ozone layer depletion odp steady state ozone layer depletion odp10 ozone layer depletion odp15 ozone layer depletion odp20 ozone layer depletion odp25 ozone layer depletion – odp30 ozone layer depletion – odp40 ozone layer depletion odp5 kg cfc-11-eq photochemical oxidation (summer smog) photochemical oxidation ebir (low nox) photochemical oxidation high nox photochemical oxidation low nox photochemical oxidation mir (very high nox) photochemical oxidation moir (high nox) kg formed ozone kg ethylene-eq kg ethylene-eq kg formed ozone kg formed ozone some of the environmental effects of photovoltaics presented in table 2 (such as ionising radiation, ozone layer depletion, and photochemical oxidation) are negligibly small (less than 1×10-8), therefore, in order to compare the various variants of electricity generation with solar cells, we chose the following impact categories: acidification potential, climate change, eutrophication potential, ecotoxicity (all types), land use, malodorous air and depletion of abiotic resources. the duration of the impact/accumulation of each negative ecological effect was set to 100 years (medium-term period). all input data for comparative analysis of different pvtechnologies is presented in table a.1, see appendix a. for complex evaluation of multiple life cycle environmental impacts of different pv technologies we use data envelopment analysis (dea). based on the work initiated by farrell [16], charnes, cooper and rhodes [9] developed linear programming technique now called dea to estimate a production frontier. it is a tool to evaluate the relative efficiency of a set of homogeneous decision making units (dmu). in traditional dea models, each dmu is assumed to be a black-box whose internal structure is not considered, i.e. dmu is regarded as the entity responsible for converting inputs into outputs and whose performances are to be evaluated. in the energy sector, dmus can be, for instance, manufacturing plants, power plants, power distribution divisions etc. in this paper we consider a number of selected combinations of pv technology and installation method having common input and environmental outputs as distinct dmus. in dea model, the relative efficiency score of any dmu is determined as a measure of the relative improvements in inputs and outputs between the dmu and its assigned target. in order to provide relative comparisons, a collection of dmus is used to evaluate each dmu against each other. the efficiency score can be assessed by minimization of input keeping its output levels constant (input-oriented model), or by maximization of output keeping the inputs at the same rate (output-oriented model) [40]. 20 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) fig. 2. efficiency evaluation in dea approach fig. 2 explains the basic idea of dea approach for efficiency evaluation. points a-g are the set of dmus under assessment. units b-e are efficient, they form efficient frontier bcde against which other inefficient units are assessed. the efficiency score of inefficient unit a can be measured relative to the projection a’ obtained by simultaneous minimization of inputs. the ray from the origin shows the direction of radial projection of unit a onto efficient frontier. thus, the radial input efficiency score of unit a is equal to oa’/oa. selection of inputs and outputs plays important role in benchmarking studies using dea [12]. the inputs and the outputs should reflect the resources utilized by the dmus and its production, respectively. dea environmental assessment usually considers both desirable and undesirable outputs to assess the performance of dmus. see for example, sueyoshi and goto [39], zhou et al. [43], førsund [17]; wang et al. [42], etc. in this study, all environmental impact categories are treated as undesirable outputs, and desirable output is electricity production. we also assume that the maintenance costs per 1 kwh for all pv technologies are the same, and therefore input parameters are not considered. for this reason, we consider environmental impacts (undesirable outputs) as inputs in the efficiency evaluation process [39]. since all environmental impact indicators are provided per 1 kwh electricity production, no return to scale assumption is needed. therefore, the dea model chosen in this work to identify the efficient frontier is ccr model introduced by charnes et al. [9] based on constant return to scale assumption. the basic motivation of the paper is to compare the performance of pv technologies under minimization of environmental impacts, hence input oriented ccr model is chosen. the input-oriented ccr model is written as:               r i i m k ko ss 11 min  mksxx k n j jkjko ,,1,0 1      (1) riysy ioi n j jij ,,1, 1      njj ,,1,0  , where n is the number of dmus, s.v. ratner, a.v. lychev 21 copyright ©2019 assa adv. in systems science and appl. (2019) m is the number of inputs, r is the number of outputs, xj = (x1j,…,xmj) is the vector of inputs of dmuj, yj = (y1j,…,yrj) is the vector of outputs of dmuj, and (xo,yo) is the vector of inputs and outputs of dmuo under evaluation, θo, optimal objective function, is the relative efficiency score of dmuo, no 1 , j , nj ,,1 , are the coefficients of linear combination for assessing dmuo;  mss ,,1  and  rss ,,1  are slacks variables, ε is an infinitesimal constant (a non-archimedean quantity). we can avoid handling ε, but then problem (1) has to be solved in two stages [9]. next, we assume that each model is solved in this way. dea efficiency scores θo are evaluated as the distance from the dmus to the efficient frontier. from the model formulation, it follows that 0 < θo ≤ 1. if efficiency score θo = 1, and optimal slacks ),,( ** 1 *   msss  and ),,( ** 1 *   rsss  , then dmuo is efficient; if optimal solution satisfies θo = 1, then dmuo is weakly efficient; units with θo < 1 are inefficient. 3. results and discussion model (1) is solved in the present paper using frontiervision software [26] that calculates relative efficiency scores and target impacts. the efficiency scores reported in table 3 show that 5 out of 19 dmus were identified as efficient compared to the other dmus against which they were assessed. cdte, a-si, and r-si laminated panels installed slanted-roof, cis panel mounted slanted-roof, and open ground installation m-si are demonstrated highest efficiency scores from an ecological point of view throughout the life cycle. table 3. efficiency scores installation method, pv technology eff. score, % 3kwp fi m-si li 65.76 3kwp fi m-si pm 61.20 3kwp fi s-si li 64.20 3kwp fi s-si pm 59.19 3kwp fri m-si 92.77 3kwp fri s-si 90.61 3kwp sri a-si li 100.00 3kwp sri a-si pm 85.32 3kwp sri cdte li 100.00 3kwp sri cis pm 100.00 3kwp sri m-si li 98.75 3kwp sri m-si pm 80.79 3kwp sri m-si pm label-certified 90.31 3kwp sri r-si li 100.00 3kwp sri r-si pm 91.81 3kwp sri s-si li 97.49 3kwp sri s-si pm 77.81 3kwp sri s-si pm label-certified 87.42 570kwp ogi m-si 100.00 22 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) table 3 shows that slanted-roof installations are more efficient than flat-roof installations. facade installations have the lowest efficiency scores. this result is quite expected and can be explained by the fact that a tilt angle of a slanted roof is close to the local latitude in switzerland (46n average) and thus allows to harvest the maximum amount of solar radiation on an annual basis [10] and, therefore, produce more electricity for its lifespan. also, it should be noted that laminated integrated panels are more ecologically efficient in comparison with mounted panels. it can be explained by the absence of aluminum frame and some other construction details due to integration straight into the roof, which leads to exclusion from the life cycle of several production activities. the higher efficiency of 570 kwp open ground pv plant is also an expected result due to the fact that the system on the ground has more airflow to cool down the panels and, thus, to keep up an optimal temperature for pv modules’ operation and to extend its carrying capacity and increase its overall electricity production [38]. comparing the ecological effectiveness of s-si, m-si and thin-film pv technologies (cdte, cis, a-si), we can see that the mean of efficiency scores of single-si modules for different installation methods is 79.45%, the mean of multi-si is 81.6% and the mean of thin-film pv is 96.33% (see tables 4-6). table 4. efficiency scores for s-si technology installation method, pv technology eff. score, % 3kwp fi s-si li 64.20 3kwp fi s-si pm 59.19 3kwp fri s-si 90.61 3kwp sri s-si li 97.49 3kwp sri s-si pm 77.81 3kwp sri s-si pm label-certified 87.42 mean score 79.45% table 5. efficiency scores for m-si technology installation method, pv technology eff. score, % 3kwp fi m-si li 65.76 3kwp fi m-si pm 61.20 3kwp fri m-si 92.77 3kwp sri m-si li 98.75 3kwp sri m-si pm 80.79 3kwp sri m-si pm label-certified 90.31 mean score 81.60% table 6. efficiency scores for thin-film technologies installation method, pv technology eff. score, % 3kwp sri a-si li 100.00 3kwp sri a-si pm 85.32 3kwp sri cdte li 100.00 3kwp sri cis pm 100.00 mean score 96.33% s.v. ratner, a.v. lychev 23 copyright ©2019 assa adv. in systems science and appl. (2019) it is important to understand that these results cannot be explained by pv cells or module efficiency, traditionally measured as percentage of the incident solar energy that the pv cell converts into electricity under the standard rating conditions (table 7), and correspond mainly to resource and energy intensity of production processes [10-11]. therefore, the bigger amount of energy produced by silicon-based pv systems in the exploitation period comparing the amount of energy produced by thin-film pv systems does not compensate the negative ecological effects of upstream processes. table 7. cell/module efficiency for different pv technologies [13, 31-32] pv technology cell efficiency (%) module efficiency (%) s-si 15.3 14.0 m-si 14.4 13.2 r-si 13.1 12.0 a-si 6.5 6.5 cdte 10.9 10.9 cis 10.7 10.7 comparing the results obtained by dea-approach with the results of simple ranking of pvtechnologies by one of the indicators of environmental or energy efficiency (see table 8), it can be noted that the technology performance indicator calculated by the dea-model, although correlated with the ratings for energy pay back time, ghg emissions and human toxicity, overall gives a more consistent and balanced result. table 8. ranks of pv-technologies according to dea model (own calculation), energy pay-back-time [23], ghg-emissions (ecoinvent, version 3.3) and human toxicity (ecoinvent, version 3.3.) installation method, pv technology rank on eff.score rank on epbt rank on ghg rank on ht 3kwp fi m-si li 12 9 16 17 3kwp fi m-si pm 14 9 17 19 3kwp fi s-si li 13 10 18 16 3kwp fi s-si pm 15 10 19 18 3kwp fri m-si 4 6 7 6 3kwp fri s-si 6 8 13 5 3kwp sri a-si li 1 n/a 2 12 3kwp sri a-si pm 9 5 10 15 3kwp sri cdte li 1 2 1 11 3kwp sri cis pm 1 3 4 7 3kwp sri m-si li 2 n/a 5 3 3kwp sri m-si pm 10 4 12 14 3kwp sri m-si pm label-certified 7 n/a 8 10 3kwp sri r-si li 1 n/a 3 2 3kwp sri r-si pm 5 1 6 8 3kwp sri s-si li 3 n/a 11 4 3kwp sri s-si pm 11 7 15 13 3kwp sri s-si pm label-certified 8 n/a 14 9 570kwp ogi m-si 1 n/a 9 1 24 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) in addition, the dea solution also allows obtaining target values for each indicator of the environmental impact for each inefficient pv technology. the target impacts for each of the inefficient pv technology were calculated according to the following equation: * * 1 , 1, , n ko j kj j x x k m    , (3.1) where j * , j = 1,…,n are optimal -variables in problem (1). the values of the target impacts are presented in table a.2, see appendix. knowledge of the target parameters that each of the photovoltaic technologies needs to achieve in order to become efficient in environmental terms can significantly simplify the process of managing applied research aimed at improving these technologies, as well as developing new photovoltaic technologies, such as organic pv. table 9 presents a summary of peer frequency, i.e. number times each efficient pv technology is a peer for another. a sum of peer weights is calculated for each efficient dmuk as a sum of variables k * in optimal solutions of problem (1) for all inefficient dmus excluding dmuk. table 9. peer count summary installation method, pv technology peer frequency sum of peer weights 3kwp sri a-si li 6 1.96 3kwp sri cdte li 7 0.98 3kwp sri cis pm 0 0 3kwp sri r-si li 13 10.83 570kwp ogi m-si 13 0.23 note that dmus 3kwp sri r-si li and 570kwp ogi m-si are the peers for 13 inefficient technologies. however, first technology is more valuable because the sum of its weights is much greater. this means that this technology is closer to efficient targets of inefficient units than another. on the contrary, 3kwp sri cis pm technology is the self-evaluator peer according to the classification of edvardsen et al. [14], because it only references itself. since all types of studied pv technologies are still under research and development, it is reasonable to consider the different technical possibilities to reach target indicators and make prevalent technologies more environmentally efficient. it will require deeper analysis for each of the environmental impacts considered at various stages of the life cycle. it is also necessary to mention that all obtained results so far do not include the environmental impacts of pv-systems’ disposal due to the lack of primary data. the inclusion of such downstream stages of the life cycle as disposal and recycling can change the environmental efficiency estimates of pv-technologies in the future and proactive planning for a pv recycling infrastructure can contribute to the ecological efficiency of each technology. 4. conclusions and policy applications a number of opportunities for improving existing government incentives and rationalizing the design of state programs under elaboration can be identified based on results of this study. our evaluations of the complex ecologic efficiency of several commercially successful pv technologies may be used for developing various state programs for supporting pv equipment manufacturers in russia and in other countries, which are just starting their own production of pv modules. the currently existing approaches to stimulation do not differentiate between technologies, leaving a choice to the customer, and most customers prefer pv modules with a higher cell efficiency. a change is necessary in the current order for sustainable development. the results of this study clearly show that from an environmental point of view it is more practical to prefer technologies, which are less resource and energy intensive in manufacturing and upstream activities. as of right now, this requirement is met by thin-film technologies (a-si, cdte, and cis); s.v. ratner, a.v. lychev 25 copyright ©2019 assa adv. in systems science and appl. (2019) however, their ecologic efficiency evaluation may change as we obtain more data on the final stages of the lifecycle for pv modules of various types. considering the issue of the installation method of pv modules, open ground installations remain preferable. in this regard, the existing state support system for photovoltaics in russia, which focuses on large power plants with capacities above 5 mw, is logical and can be deemed successful. while developing new microgeneration state support programs, it is reasonable to prefer slanted-roof installations and laminated integrated panels. stimulating the population to use these specific installation types for pv systems can be done both via standardization and certification (for example, municipal standards for installation of solar panels), as well as via special educational events and training programs. r&d support programs in the field of solar energy can also be constructed in a way that stimulates not just research targeting the growth of cost-competitiveness and energy efficiency of photovoltaics, but also improvement of their environmental efficiency. in the near future, we can use target environmental impacts, obtained in this research as benchmarks for emerging each of the currently existing, commercially viable pv technologies, which are presented in this study. as technologies develop and new data on their environmental impacts is obtained, these target parameters may be re-calculated using the herein suggested dea approach. despite the fact that abovementioned practical conclusions have independent scientific value for environmental management, the main contribution of our research is the development of a method for the integrated assessment of the negative impact of renewable technology on the environment for the widest possible range of ecology effects taken into account. the combined lca and dea approach is proposed for a first time, therefore the paper expands the scope of practical applications of the popular decision-making methodology such as dea. the main advantage of proposed method of integrated assessment of the negative impact of pv-technologies on the environment during entire life cycle is its scalability: it is easy to integrate into the calculations both a greater number of new technologies and more identified environmental effects. even for a large number of studied technologies and a large number of categories of environmental impact, there is still the possibility of ranking alternatives. the main limitation of this study is the lack of empirical data on environmental impacts of new pv-technologies, which are only entering the stage of mass production: most importantly, heterojunction solar technologies, being developed by the biggest russian pv-manufacturer hevel group. this problem opens a further area of research, which, however, can be easily integrated into the proposed approach of relative complex environmental efficiency evaluation. another limitation is connected with the lack of discrimination capability for efficient dmus classical dea models. because of this, several technologies can have the highest rank in integrated environmental efficiency at the same time. as one can see in our study 5 technologies are classified as efficient ones. therefore, complete ranking is not possible with ccr model. in order to generate a complete ranking of dmus, many theoretical extensions of basic dea models have been proposed by various researchers. a recent survey on complete ranking methods in dea can be found in [1]. each approach has their own strength and weaknesses and detailed discussion on those methods is beyond the scope of the present paper. we assume that further choice between eco-efficient pv technologies can be made using one of dea ranking methods or with the help of traditional instruments of technical and economic analysis. acknowledgements dea calculations were conducted with the support of the russian science foundation (project no. 17-11-01353). 26 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) references [1] aldamak, a., & zolfaghari, s. (2017). review of efficiency ranking methods in data envelopment analysis. measurement, 106, 161–172. [2] amarakoon, s., vallet, c., curran, m., haldar, p., metacarpa, d. et al. (2017). life cycle assessment of photovoltaic manufacturing consortium (pvmc) copper indium gallium (di), selenide (cigs) modules. the international journal of life cycle assessment. https://www.springerprofessional.de/life-cycle-assessment-ofphotovoltaic-manufacturing-consortium-p/13322018 (accessed 12 february 2019). [3] amponsah, n.y., troldborg, m., kington, b., aalders, i. & hough, r.l. (2014). greenhouse gas emissions from renewable energy sources: a review of lifecycle considerations. renewable and sustainable energy reviews. 39, 461-475. [4] bracquene e., peeters j.r., dewulf w. & duflou j.r. (2018) taking evolution into account in parametric lca model for pv panels. procedia cirp. 69, 389-394. [5] butuzov v.a. (2018). solar heat supply: world statistics and peculiarities of the russian experience. thermal engineering. 65(10), 741–750. [6] butuzov v.a., amerkhanov r.a. & grigorash o.v. (2018). geothermal power supply systems around the world and in russia: state of the art and future prospects. thermal engineering. 65(5), 282–286. [7] celik i., song z., cimaroli a.j., yan y., heben m.j. et al. (2016) life cycle assessment (lca) of perovskite pv cells projected from lab to fab. solar energy materials and solar cells. 156, 157-169 [8] chevalier b., reyes-carrillo t. & laratte b. (2011) methodology for choosing life cycle impact assessment sector-specific indicators. international conference on engineering design, iced11, 15 18 august 2011, technical university of denmark. [9] cooper, w.w., seiford, l.m. & tone, k. (2007). data envelopment analysis: a comprehensive text with models, applications, references and dea-solver software. 2nd edition, springer, new york. [10] cucchiella, f. & d’adamo, i. (2012). estimation of the energetic and environmental impacts of a roof-mounted building-integrated photovoltaic systems. renewable and sustainable energy reviews. 16, 5245-5259. [11] dubey, s., jadhav, n.y. & zakirova, b. (2013). socio-economic and environmental impacts of silicon based photovoltaic (pv) technologies. energy procedia. 33, 322-334. [12] dyson, r.g., allen, r., camanho, a.s., podinovski, v.v., sarrico, c.s. et al. (2001). pitfalls and protocols in dea. european journal of operational research. 132(2), 245-59. [13] ecoinvent, 2011. report no.6, http://www.ecoinvent.ch (accessed 15 february 2018). [14] edvardsen, d.j., førsund, f.r. & kittelsen, s.a.c. (2008). far out or alone in the crowd: a taxonomy of peers in dea. journal of productivity analysis. 29, 201-210. [15] emrouznejad, a. & yang, g. (2018). a survey and analysis of the first 40 years of scholarly literature in dea: 1978–2016. socio-economic planning sciences. 61, 4-8. [16] farrell, m.j. (1957). the measurement of the productive efficiency. j of the royal statistical society. 120, 253-281. https://www.springerprofessional.de/the-international-journal-of-life-cycle-assessment/4936586 https://www.springerprofessional.de/the-international-journal-of-life-cycle-assessment/4936586 https://www.springerprofessional.de/life-cycle-assessment-of-photovoltaic-manufacturing-consortium-p/13322018 https://www.springerprofessional.de/life-cycle-assessment-of-photovoltaic-manufacturing-consortium-p/13322018 s.v. ratner, a.v. lychev 27 copyright ©2019 assa adv. in systems science and appl. (2019) [17] førsund, f. r. (2018). multi-equation modelling of desirable and undesirable outputs satisfying the materials balance. empirical economics. 54(1), 67–99. [18] fthenakis, v. & moskowitz, p. (2000). photovoltaics: environmental health and safety issues and perspectives. progress in photovoltaics. 8, 27-38. [19] fthenakis, v.m. (2004). life cycle impact analysis of cadmium in cdte pv production. renewable and sustainable energy reviews. 8(4), 303-334. [20] good c. (2016) environmental impact assessments of hybrid photovoltaic–thermal (pv/t) systems – a review. renewable and sustainable energy reviews. 55, 234–239 [21] guinèe, j.b., gorrée, m., heijungs, r., huppes, g., kleijn, r. et al. (2001). life cycle assessment; an operational guide to the iso standards; parts 1 and 2. ministry of housing, spatial planning and environment (vrom) and centre of environmental science (cml), den haag and leiden, the netherlands. [22] hsu, d., o’donoughue, p., fthenakis, v., heath, g., kim et al. (2012). life cycle greenhouse gas emissions of crystalline silicon photovoltaic electricity generation systematic review and harmonization. journal of industrial ecology. 16(s1), s122s135. [23] jungbluth n., stucki m, frischknecht r., & buesser, s. (2010) photovoltaics. in dones, r. (ed.) et al., sachbilanzen von energiesystemen: grundlagen für den ökologischen vergleich von energiesystemen und den einbezug von energiesystemen in ökobilanzen für die schweiz. ecoinvent report no. 6-xii, esu-services ltd, uster, ch, [in german]. [24] jungbluth n., tuchschmid m. & de wild-scholten m. j. (2008) life cycle assessment for photovoltaics: update for ecoinvent data v2.01 http://esuservices.ch/publications/energy/ [25] kim, h.c., fthenakis, v., choi, j.-k. & turney, d.e. (2012). life cycle greenhouse gas emissions of thin-film photovoltaic electricity generation. journal of industrial ecology. 16, s110-s121. [26] krivonozhko, v. e., lychev, a. v. & førsund, f. r. (2017). measurement of returns to scale in radial dea models. computational mathematics and mathematical physics. 57(1), 83–93. [27] li g., xuan q., pei g., su y., lu y. et al. (2018) life-cycle assessment of a lowconcentration pv module for building south wall integration in china. applied energy. 215, 174-185 [28] ling-chin j., heidrich o. & roskilly a.p. (2016) life cycle assessment (lca) – from analysing methodology development to introducing an lca framework for marine photovoltaic (pv) systems. renewable and sustainable energy reviews. 59, 352-378 [29] louwen, a., sark, w., schropp, r. & faaij, a. (2016). a cost roadmap for silicon heterojunction solar cells. solar energy materials and solar cells. 14, 295-314. [30] meijer, a., huijbregts, m.a.j., schermer, j.j., reijnders, l. (2003). life-cycle assessment of photovoltaic modules: comparison of mc-si, ingap and ingap/mc-si solar modules. progress photovoltaic res appl. 11, 275-287. [31] moreno ruiz e., lévová t., reinhard j., valsasina l., bourgault g. et al. (2016). documentation of changes implemented in ecoinvent database v3.3. ecoinvent, zürich, switzerland. http://esu-services.ch/publications/energy/ http://esu-services.ch/publications/energy/ https://www.sciencedirect.com/science/article/pii/s1364032115014410#! https://www.sciencedirect.com/science/article/pii/s1364032115014410#! https://www.sciencedirect.com/science/article/pii/s1364032115014410#! https://www.sciencedirect.com/science/journal/13640321 https://www.sciencedirect.com/science/journal/13640321/59/supp/c https://www.sciencedirect.com/science/article/pii/s0927024815006741#! https://www.sciencedirect.com/science/journal/09270248 28 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) [32] moreno ruiz e., valsasina l., fitzgerald d., brunner f., vadenbo c.et al. (2017). documentation of changes implemented in the ecoinvent database v3.4. ecoinvent, zürich, switzerland. [33] nizhegorodtsev, r.m. & ratner, s.v. (2016). trends in the development of industrially assimilated renewable energy: the problem of resource restrictions. thermal engineering. 63(3), 197-207. [34] pandey, a., tyagi, v., jeyraj, a., selvaraj, l., rahim, n. et al. (2016). recent advances in solar photovoltaic systems for emerging trends and advanced applications. renewable sustainable energy review. 53, 859-884. [35] ratner, s.v. & iosifov, v.v. (2017). strategizing for solar energy development in russia subject to environmental impact. economic analysis: theory and practice. 8, 1522-1540 [in russian] [36] ratner, s.v. & nizhegorodtsev, r.m. (2017). analysis of renewable energy projects’ implementation in russia. thermal engineering. 64(6), 429-436. doi:10.1134/s0040601517060052 [37] renewables, 2018. global status report. renewable energy policy network ren21 secretariat for the 21st century, paris, france, http://www.ren21.net/wpcontent/uploads/2018/06/17-8652_gsr2018_fullreport_web_final_.pdf; 2018 (accessed 10 february 2019). [38] ritzen, m.j., vroon, z.a.e.p., rovers, r., lupisek, a. & geurts, c.p.w. (2017). environmental impact comparison of a ventilated and a non-ventilated buildingintegrated photovoltaic rooftop design in the netherlands: electricity output, energy payback time, and land claim. solar energy. 155, 304-313. [39] sueyoshi, t., & goto, m. (2018). environmental assessment on energy and sustainability by data envelopment analysis. john wiley & sons, ltd. [40] sueyoshi, t., yuan, y., li, a. & wang, d. (2017). methodological comparison among radial, non-radial and intermediate approaches for dea environmental assessment. energy economics. 67, 439–453. [41] tsang m.p., sonnemann g.w. & bassani d.m. (2016) life-cycle assessment of cradle-to-grave opportunities and environmental impacts of organic photovoltaic solar panels compared to conventional technologies. solar energy materials and solar cells. 156, 37-48. [42] wang, d. d. & sueyoshi, t. (2017). assessment of large commercial rooftop photovoltaic system installations: evidence from california. applied energy. 188, 45– 55. [43] zhou p., poh k.l. & ang b.w. (2016). data envelopment analysis for measuring environmental performance. in: hwang sn., lee hs., zhu j. (eds) handbook of operations analytics using data envelopment analysis. international series in operations research & management science, vol 239. springer, boston, ma. http://www.ren21.net/wp-content/uploads/2018/06/17-8652_gsr2018_fullreport_web_final_.pdf http://www.ren21.net/wp-content/uploads/2018/06/17-8652_gsr2018_fullreport_web_final_.pdf s.v. ratner, a.v. lychev 29 copyright ©2019 assa adv. in systems science and appl. (2019) appendix a table a.1. input data for comparative analysis installation method, pv technology so2-eq (gen) gwp100a nox-eq pox-eq faetp100a fsetp100a htp100a land air maetp100a msetp100a resources 3kwp fi m-si li 0.00080479 0.11721 0.00036814 0.00040589 0.27186 0.63365 0.18929 0.009284 1388.1 0.91367 1.13230 0.0007772 3kwp fi m-si pm 0.00082222 0.12094 0.00037988 0.00041474 0.30161 0.70768 0.19433 0.00945 1439.8 1.00800 1.26120 0.0008035 3kwp fi s-si li 0.00091488 0.13939 0.00042352 0.00044450 0.27868 0.64846 0.18877 0.009791 1410.9 0.93778 1.15750 0.0009270 3kwp fi s-si pm 0.00093127 0.14289 0.00043456 0.00045282 0.30666 0.71806 0.19351 0.009947 1459.5 1.02650 1.27870 0.0009517 3kwp fri m-si 0.00548520 0.081018 0.00025345 0.00027563 0.20446 0.48003 0.12667 0.006288 956.25 0.68352 0.85491 0.0005488 3kwp fri s-si 0.00062211 0.095803 0.00029034 0.00030144 0.20776 0.48676 0.12636 0.006626 970.21 0.69560 0.86628 0.0006479 3kwp sri a-si li 0.00047230 0.064313 0.00020003 0.00023177 0.18401 0.43101 0.13757 0.003955 318.82 0.61342 0.76251 0.0004385 3kwp sri a-si pm 0.00059623 0.085838 0.00026219 0.00027205 0.24332 0.57601 0.18127 0.004633 469.1 0.80436 1.01540 0.0005568 3kwp sri cdte li 0.00043030 0.049012 0.00017623 0.00025864 0.18300 0.42532 0.13686 0.003735 318.36 0.61210 0.75626 0.0003297 3kwp sri cis pm 0.00051273 0.074027 0.00024123 0.00031670 0.20854 0.48391 0.12810 0.004859 454.52 0.69280 0.85354 0.0004850 3kwp sri m-si li 0.00051316 0.074603 0.00023597 0.00026671 0.18190 0.42439 0.11706 0.006198 923.45 0.61103 0.75818 0.0005064 3kwp sri m-si pm 0.00062838 0.092526 0.00029008 0.00031474 0.22285 0.53606 0.14868 0.007206 1091.2 0.76381 0.95532 0.0006138 3kwp sri m-si pm label-certified 0.00055929 0.082354 0.00025819 0.00028014 0.20338 0.47712 0.13233 0.006414 971.26 0.67983 0.85029 0.0005463 3kwp sri r-si li 0.00050074 0.070164 0.00023003 0.00026025 0.18086 0.42209 0.11463 0.005542 718.65 0.60785 0.75494 0.0004699 3kwp sri r-si pm 0.00055134 0.078666 0.00025441 0.00027498 0.20443 0.47994 0.13138 0.005779 771.09 0.68332 0.85597 0.0005136 3kwp sri s-si li 0.00058886 0.089770 0.0002739 0.00029305 0.18654 0.43445 0.11733 0.006542 939.36 0.62744 0.77533 0.0006081 3kwp sri s-si pm 0.00071034 0.109050 0.00033121 0.00034343 0.23228 0.54382 0.14795 0.007578 1105.9 0.77763 0.96840 0.0007254 3kwp sri s-si pm label-certified 0.00063224 0.097058 0.0002948 0.00030567 0.20674 0.48403 0.13169 0.006744 984.31 0.69214 0.86193 0.0006456 570kwp ogi m-si 0.00049902 0.083707 0.00025718 0.00020105 0.099367 0.22837 0.10382 0.042508 948.81 0.34096 0.40829 0.0005503 30 evaluating environmental impacts of photovoltaic technologies copyright ©2019 assa adv. in systems science and appl. (2019) table a.2. efficient targets for pv technologies installation method, pv technology so2-eq (gen) gwp100a nox-eq pox-eq faetp100a fsetp100a htp100a land air maetp100a msetp100a resources 3kwp fi m-si li 0.00052923 0.077077 0.00024209 0.00026691 0.17878 0.41669 0.12448 0.006105 912.81 0.60083 0.74460 0.0005111 3kwp fi m-si pm 0.00054069 0.079530 0.00024981 0.00027273 0.19834 0.46537 0.12779 0.006215 946.81 0.66286 0.82937 0.0005284 3kwp fi s-si li 0.00060163 0.091663 0.00027851 0.00029230 0.18326 0.42643 0.12414 0.006438 927.81 0.61668 0.76117 0.0006096 3kwp fi s-si pm 0.00061240 0.093964 0.00028577 0.00029777 0.20166 0.47220 0.12725 0.006541 959.77 0.67503 0.84087 0.0006258 3kwp fri m-si 0.00360707 0.053277 0.00016667 0.00018125 0.13445 0.31567 0.08330 0.004135 628.83 0.44948 0.56219 0.0003609 3kwp fri s-si 0.00040910 0.063000 0.00019093 0.00019823 0.13662 0.32009 0.08309 0.004357 638.01 0.45743 0.56967 0.0004261 3kwp sri a-si pm 0.00039208 0.056447 0.00017242 0.00017890 0.16001 0.37878 0.11920 0.003046 308.48 0.52895 0.66773 0.0003662 3kwp sri m-si li 0.00033745 0.049059 0.00015517 0.00017539 0.11962 0.27908 0.07698 0.004076 607.26 0.40181 0.49858 0.0003330 3kwp sri m-si pm 0.00041322 0.060845 0.00019076 0.00020697 0.14655 0.35251 0.09777 0.004739 717.57 0.50228 0.62822 0.0004036 3kwp sri m-si pm label-certified 0.00036779 0.054156 0.00016979 0.00018422 0.13374 0.31375 0.08702 0.004218 638.70 0.44706 0.55915 0.0003592 3kwp sri r-si pm 0.00036256 0.051731 0.00016730 0.00018083 0.13443 0.31561 0.08640 0.003800 507.07 0.44935 0.56289 0.0003377 3kwp sri s-si li 0.00038723 0.059033 0.00018012 0.00019271 0.12267 0.28569 0.07716 0.004302 617.72 0.41260 0.50986 0.0003999 3kwp sri s-si pm 0.00046712 0.071711 0.00021780 0.00022584 0.15275 0.35762 0.09729 0.004983 727.24 0.51137 0.63682 0.0004770 3kwp sri s-si pm label-certified 0.00041576 0.063825 0.00019386 0.00020101 0.13595 0.31830 0.08660 0.004435 647.28 0.45515 0.56681 0.0004246 adv syst sci appl 2019; 03; 23-37 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/747 study of operational efficiency of large-scale publicly owned companies in russia valentina e. guseva1, evgeny a. kuzmin2,3*, andrey i. vlasov4 1) tyumen industrial university, tyumen, russia 2) institute of economics of the ural branch of the russian academy of sciences (ie ub ras), ekaterinburg, russia 3) ural state university of economics, ekaterinburg, russia 4) bauman moscow state technical university, moscow, russia received february 12, 2018; revised december 9, 2018; published december 31, 2018 abstract: a sustainable economic growth is partly a consequence of the best possible balance between public and private sectors. reasons for changed government efforts in economy are highly diverse and lie in the area of strategic development goals. a review of evidence-based studies shows that public ownership in a capital of companies influences some characteristics of their operational activities. the specific components, elaborated in the paper, contribute to this. based on data provided in financial statements of leading publicly owned companies in russia, the research aims at identification of causal relationships in composition and ownership structure of companies, as well as their impact on economic ratios. researchers verified the assumptions, put forward in this respect, using the correlation-regression analysis. the basis for the operational efficiency comparison included the return on equity (roe), return on assets (roa), return on sales (ros), equity ratio and current liquidity ratio. findings show that there are sufficient arguments in favour of the assumption that publicly owned companies are consistently less efficient against the private ones. obtained data made it possible to conclude that a change in the state’s share in the capital of large-scale companies in russia influences the return on assets and return on equity only. other ratios under consideration express poor correlations. the heterogenic impact on adjacent ratios of operational efficiency prevents from concluding in a rigorous manner that publicly owned companies are a priori less efficient. one needs to keep in mind sector affiliation and a capital concentration ratio. keywords: operational efficiency, publicly owned companies, private companies, capital concentration, economic growth, sustainable development. 1. introduction efficiency of publicly owned and private companies has not become less of an issue in both the developed and developing countries. in this discussion, there is a main problem of a search for the best possible ratio between incorporation forms in economy [30]. steady trends towards a growth of the public sector in russian economy are a clear example saying that the issue is pressing. at the same time, this sector has historically concentrated production facilities and capital resources. the transition to market economy in the 1990-ies had been marked by a declined government share (to 38%) in gdp by 2006 [51]. however, later, this trend did not develop and a balance between sectors started recovering. by 2016, the share of the public sector in russia's gdp had reached 46% [39]. these data find their confirmation in ownership structure statistics. the national government is the largest shareholder in the russian market. the share of publicly owned companies that disclose their ownership structure, in the * corresponding author: kuzminea@gmail.com 24 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) moscow exchange index, is at least 45% [37]. 28 of 100 companies, ranking by ra expert in 2017, are the publicly owned companies ones (percentage holding is 25%) [55]. 5.5% of all the employees in the russian economy works for them in 2017. the share of proceeds of the top 100 public sector companies was 50% of their total revenue in 2017. according to riarating in 2018, the publicly owned companies account for more than 48% of top 100 largest companies capitalization [56]. it indicates that the bunching of publicly owned companies companies is significantly high. we have recently observed further strengthening of a role of the national government, not only as a participant in economic affairs, but also as their regulator. efficiency of the public sector is a key issue in definition and justification of its required existence. the component analysis might help to make aspects of this issue clearer. there are convincing evidence-based data saying that overall efficiency of factors of production at russian publicly owned companies is much lower than in the private sector [52]. labour productivity at publicly owned companies is more than 30% below the national average rate due to the fact that crucial decisions often serve political purposes instead of commercial ones. hence, efficiency criteria in state’s ownership management are not formal ratios (such as profitability, scope of tax deductions to a budget), but a quality of provided public goods and other institutional ratios [44]. however, followers of this approach exclusively view publicly owned companies one-sidedly, as an element of infrastructure support to the social and economic life. in the framework of the systematic view of the issue, researchers focus on an analysis of comparative efficiency of publicly owned companies as isolated agents. unsatisfactory management of public assets and an excessive size of the public sector indeed pose a serious risk to the economic growth [44]. in the countries where the public sector’s share is significant, they always serve for alignment of a business cycle and employment maintenance [1], often to the detriment of economic interests. the specifics available in the shift of interests manifests itself in the balance between the public and private capital in an ownership structure of companies. it makes sense in the discussion that this should cause a change in their performance, expressed in objective ratios. therefore, our research aims at identification of causal correlations in a composition and ownership structure of companies, their impact on economic ratios. the conclusions that we hope to make would allow us being closer to a search for the best possible ratio between public and private sectors in economy, which make it possible to achieve sustainable development. 2. review of literature boyett [6] believes that the process of public entrepreneurship appeared as a response to a transitional state in an economic system, as a new form of entrepreneurship in terms of an unstable external environment in response to uncertainty. there are various formal procedures, with which it is possible to calculate a share of the public sector in economy, its efficiency and various expert opinions, often contradictory. in the research, publicly owned companies refer to those companies where for various reasons, a national government exercises control being also a dominating owner. an ownership structure and management models might differ [46]. in any case, controlled parameters of companies' activities measure their financial standing, main characteristics of which affect competitiveness of an enterprise in a sector and in a market as a whole [13, 26]. they also affect a potential and opportunities of business cooperation with other companies [11, 19] and involvement of stakeholders [8, 17]. therefore, the analysis of corporate activities at runtime aims to show the extent, to which companies have properly consumed available resources [34]. the conventional view is that publicly owned companies are less efficient than private ones. some researchers explain this by the fact that, under market conditions, publicly owned operational efficiency of large-scale publicly owned companies 25 copyright ©2019 assa. adv. in systems science and appl. (2019) companies are usually established as a compensation for market failures [43]. if market failures are treated by joint efforts as ppp projects, the return on investment is maximized when the private capital prevails [3]. conceptual understanding of this phenomenon lies in the theory of public administration [54]. at the same time, in terms of institutional economy, a national government is unable to provide control over management in an efficient manner due to high information asymmetry (principal-agent theory). it is usual that for the companies with a high state’s share, external mechanisms of management monitoring do not actually perform their disciplining function [41]. borcherding, pommerehne and schneider [5] review findings from 50 evidence-based studies, in 40 of which authors concluded that private entrepreneurship had turned out to be much more efficient than public one. in 7 other cases, it was impossible to identify obvious advantages of one or another incorporation forms. in the similar comparison of functioning of private and publicly owned companies, boardman and vining [4] reviewed the results obtained earlier and compared them with own calculations. their main conclusion is as follows: when considering a wide range of factors that influence transaction efficiency, large-scale production companies with joint ownership and similar companies that are 100%owned by a national government are significantly less efficient against similar private companies. muller [31] comes to similar conclusions, summarising findings from 71 studies. in 56 cases, publicly owned enterprises were less efficient. specifics of publicly owned companies makes specific a number of components that affect operational and strategic efficiency [45, 49, 50]. they include politics-driven resource allocation, weak managerial mechanisms and limited managerial autonomy, poor coordination of activities, risk-evasive behaviour of managers, poor market discipline, strong social control, poor employment motivation, bureaucrats, lack of flexibility, etc. in general, the terms, in which publicly owned enterprises operate, are soft budget constraint conditions with no active incentives for development. at private companies with an insignificant state’s share, which has not yet become controlling, mentioned components act to a lesser extent. benefits from such cooperation come to the fore. there are reduction in a cost of investment attraction [9, 16, 29] and protectionism in commodity and resource markets [27]. private companies consider the presence of the national government as a factor to increase competitiveness in terms of growing globalization in economy. note that even the presence of a prevailing state’s share in a company's capital structure is not in itself a prerequisite of inefficiency. however, a circle of the papers that support this idea is narrow. it is noteworthy that such papers only cover a number of infrastructural sectors (4 out of the 5 cases, electric power production and distribution). there, for a number of reasons, competition is poor [39]. when analysing a wider market, there is lower efficiency of publicly owned enterprises against private enterprises operating in the same sector. in many cases, production costs at publicly owned enterprises are higher. a desire of managers at publicly owned enterprises to increase their budget/organizational size (niskanen model [33]) and bureaucrats largely explain inefficiency at publicly owned enterprises. findings from evidence-based research on an impact of public ownership on efficiency of russian companies are not categorical. kuznetsov and murav’yev think that the higher a proportion of outsiders is, the higher processing efficiency of companies is, measured as labour productivity. on the contrary, ownership of insiders and the national government causes lower productivity [25]. kapelyushnikov [20] shows that enterprises with a dominating state’s share in the authorized capital have the worst financial ratios. the higher school of economics’ review [53] includes offbeat findings in terms of the state's influence on efficiency of russian companies. authors of the review focused on the selective influence of the state on individual enterprises in conjunction with their modernization. 26 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) ruzhanskaya and krutikov [40] believe that, as an owner, the state does not demonstrate efficient management of companies. they have nevertheless found a positive influence of the state’s share in the capital on productivity, investment attractiveness and profitability ratios. at the same time, the state’s share in ownership is not statistically related to the companies’ growth. in absence of the obvious economic growth, the highly concentrated ownership of russian companies has became at the same time a condition and consequence of the developed integration processes in the corporate segment [2, 35]. findings from various evidence-based research of russian production enterprises [24, 42] support such a conclusion [24, 42]. in the completed review, we emphasized that efficiency of publicly owned and private companies is a subject of an extensive discussion. there are sufficient arguments in favour of the assumption that publicly owned companies are comparatively less efficient than the private ones. the question has remained open, to what extent this influence is apparent, whether it is of a selective nature by sector affiliation and concentration level of state ownership in the capital. in the narration below, we will try to make these points clearer. 3. materials and methods one might approach company’s efficiency in terms of various aspects [18, 32]. experts usually distinguish operational and strategic efficiency [38]. this research focuses operational efficiency of companies exactly. its important characteristics include return on equity (roe), return on assets (roa), return on sales (ros), equity ratio and current liquidity ratio. research source data included data from financial statements of companies as reported by the interfax centre for corporate information disclosure [10]. recent 7 years was a period for the analysis (2012-2018). all the publicly owned companies for the research were grouped by the international monetary fund (imf) classification [14]: (1) 100%-publicly owned. the group includes the entities that are 100-percent ownership of the national government (e.g. russian railways (rr), united shipbuilding corporation (usc)), (2) joint companies. the group includes the companies jointly established by governmental and non-governmental entities, as well as the companies established by governmental entities and non-resident entities, in which a share of a non-resident entity does not exceed 10%, (3) private companies with participation of the state. for the research, we selected 12 companies from various sectors of economy. see their general characteristics in table 1. companies represent different branches of economy, their gross output share is extremely large. according to spark-interfax, over 2018, these 12 companies account for 16.5 trillion rubles of revenue, which is about more than 12% of the turnover of all the russian companies. undoubtedly, the limited sample has its disadvantages. meanwhile we rely on the sample importance in terms of the scale and share of participation in economy. table 1. description of the companies that have been selected for the research name sector public ownership, % rr transportation 100.00 usc shipbuilding 100.00 gazprom oil and gas production 50.23 rosneft oil and gas production 50.00 uac aircraft industry 96.95 transneft transportation of oil and oil products 20.86 inter rao electric power sector 44.30 rushydro electric power sector 75.40 rosseti power sector 88.89 rostelecom telecommunications 54.90 tatneft oil and oil product production 34.00 alrosa diamond mining 66.00 operational efficiency of large-scale publicly owned companies 27 copyright ©2019 assa. adv. in systems science and appl. (2019) source: compiled as reported by spark-interfax [47]. to study specifics of operational efficiency at publicly owned companies in russia, we put forward a number of assumptions. assumption testing followed methods of the correlation-regression analysis. correlating makes it possible to estimate a strength of a component influence on the process under consideration. the linear correlation coefficient takes values from -1 to +1. correlations between attributes might be poor and strong (close). their criteria are estimated with the cheddok scale [21]. r2 sample determination coefficient is a measure of an overall quality of the regression equation (correspondence of the made equation with statistical data). by definition, r2 takes values in the range [0;1]. the closer r2 is to one, the better the regression approximates empirical data. statistical significance of individual regression ratios is checked with student’s t-test, and the equation as a whole – using fisher’s f-test. 4. results and discussion 4.1. overall analysis of economic standing of companies based on financial statements, we calculated operational efficiency ratios for each company under consideration (annex a). see table 2 for summarized estimates for 2018. the current state of publicly owned companies provides the first illustration of their operating performance at the moment that is regarded further over the preceding period. cumulative data provides with comparing, first of all, the differences in branch specific nature; secondly, study their heterogeneity. the demonstrative example is labour productivity, which is significantly diversified. the explanation of this consists in performing a range of executive functions by publicly owned companies as united shipbuilding corporation (usc). table 2. operational efficiency ratios for sample companies in 2018 company/ ratio roa roe ros equity ratio current liquidity ratio labour productivity, thousand rub/ human conventional operational efficiency rr 0.30 0.42 7.83 0.70 0.47 2391 low usc 0.04 0.12 -3.12 0.31 10.33 126265 low gazprom 6.8 9.88 18.02 0.66 1.96 11113 high rosneft 3.94 24.83 5.08 0.16 1.43 22624 low uаc -6.63 -9.81 -0.72 0.66 1.76 4375 low transneft 0.98 5.46 5.75 0.18 0.79 8147 low inter rao 4.4 4.92 1.64 0.86 2.14 1113 high rushydro 3.64 4.38 40.52 0.82 7.63 2336 high rosseti -2.93 -3.2 5.83 0.92 36.22 125 low rostelecom 0.92 2.11 3.99 0.41 0.58 2372 low tatneft 24.81 31.33 31.49 0.78 3.14 37551 high alrosa 5.29 8.15 20.13 0.63 1.58 6280 high let us consider in more detail a standing of each company at runtime in 2012-2018. at rr, roa and roe were low in the period under consideration. thus, in 2017, the sector-specific average return on assets at rail transport enterprises engaged in intercity and international passenger traffic was 1.1% [22], whereas at rr, it was 0.30%. at the same time, the return on sales was comparable to the sector-specific average of 8.3% (at rr, it was 8.25%). in 2012-2018, at rr, dynamics of profitability ratios was unstable. by 2014, there had been a clear decline in values. a recovery in values to values of 2012 had only been recorded by 2017-2018. in 2012-2018, at rr, the equity ratio was higher than the statutory value (0.5%), with an obvious decline of this ratio from 0.8% to 0.7%. the current liquidity ratio was significantly lower than the statutory value (1.5-2.5) and had the steady decreasing dynamics over the period under consideration. it is possible to conclude that rr has relatively low operational efficiency. 28 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) operational efficiency ratios of another publicly owned company, usc, assume that its profitability ratios are also low in spite of the positive dynamics in 2012-2015. as reported by the federal tax service of russia (fts), in 2017, the return on assets in production of other vehicles and equipment was 1.9% [22]. at usc, this ratio was at the level of 0.03%. the average sector-specific return on sales was 12.9%, while at usc, the value was negative (3.78%). at the same time, in 2018, the six-fold excess of the current liquidity ratio over the statutory value points out to an excess and inefficient use of funds. in 2012-2018 at gazprom, the return on assets and return on equity had a tendency towards an increase, while the return on sales significantly declined. according to the federal tax service, in the sector of crude oil and gas production, in 2017, the average return on assets was 11.1% [22]. at gazprom, this ratio was more than 2 times higher (25.2%). the company’s equity ratio is relatively high (share of equity capital is at the level of 66%), but there is a tendency towards its decline. the current liquidity ratio fully complies with statutory values. thus, at gazprom, operational efficiency ratios are quite high with the clear positive dynamics. next, let us analyse rosneft activities. the state’s share in the company’s capital was over 50%. calculated data assume the negative changes available in profitability ratios that are significantly lower than sector-specific average values. there are also low levels of the equity ratio. in 2018, the share of equity was only 16% of total liabilities. the current liquidity ratio follows this tendency being below the statutory ratio. it is obvious that at rosneft, operational efficiency is relatively low. with the state’s share of over 96%, united aircraft corporation (uac) was not distinctive in operational efficiency. as dynamics for 2012-2018 assumes (annex a), operational efficiency was low. negative profitability values were observed almost over the entire period under consideration. transneft was another company in the sample. in 2018, the company's return on assets, as well as return on equity sharply dropped, roa was 0.98%, and roe was 5.46%. at the same time, as early as in 2017, values were 5.40% and 31.53%, respectively. sector-specific average roa in the pipeline transportation sector is about 5.5% [22], while ros is 13.5% [22]. in 2018, the equity ratio was 18%. the current liquidity ratio also was below the normal level. at inter rao, the effective state’s share is 44.3%. its operational efficiency ratios show the positive dynamics, while the capital structure is highly sustainable. the current liquidity ratio was also within the standard statutory range. in 2012-2013, over the short period of time, profitability ratios had been negative. however, by 2016, they had already increased. the average return on sales in generation, transmission and distribution of electric power is at the level of 11.1% [22], whereas at inter rao, in 2017, this type of profitability was negative (0.19%). the return on assets was slightly lower than the sector-specific average value, 3.95% versus 4.8% in 2017. in general, inter rao typically has relatively high operational efficiency. rushydro is another representative of the electric power sector. minor fluctuations in roe and ros in 2012-2018 (annex a) confirm sustainability and rigidity of a company's business model. at the same time, values of profitability ratios are close to sector-specific average values of roa and much better in terms of ros. in 2018, the equity share in the capital structure was 82% and the current liquidity ratio was 5 times higher than the standard statutory rate. rosseti stands out from this list to some extent, which also operates in the electric power sector. its main activity includes the long distance electric power transmission and management of the energy system. in the structure of the shared capital at rosseti, the state’s share is 88.89%. the company has numerous subsidiaries, which, due to the high state’s share in the parent company, also have high dependence on government entities. operational efficiency of the company was low in terms of profitability ratios. at the same time, in 2018, operational efficiency of large-scale publicly owned companies 29 copyright ©2019 assa. adv. in systems science and appl. (2019) the equity share was 92%, the current liquidity ratio was more than 24 times more than the statutory value, which points out to excessive working assets. rostelecom is the largest telecoms provider in russia. in 2012-2018, its return on assets and profitability ratio (annex a) had negative dynamics being below the required threshold. thus, in 2017, in the information and communication sector, sector-specific average roa and ros were 9.1% and 14.2%, respectively [22], whereas, in 2017, rostelecom’s roa was at the level of 4.64%, ros was at the level of 4.37%. the equity ratio was also lower than the recommended rate (41% of equity). tatneft is another representative of publicly owned business. the state’s share at this company is 34%. the analysis completed showed that the company's operational efficiency is quite high as evidenced by the positive dynamics of profitability ratios (roa, roe and ros) (all ratios increased over the period and their values were higher than sector-specific average values). the equity ratio was 0.78 (against the statutory ratio of 0.50). in the period under review, the current liquidity ratio was also higher than the statutory value. alrosa was a final company in the research sample. as one can see from calculated data (annex a), all profitability ratios in 2012-2018 had tendencies towards a decline and their values were below average sector-specific values. in the mining industry, average roa is about 11.0% [22]. however, in 2017, the company achieved the profitability of almost 2 times below than this value is (4.18%). at the same time, average sector-specific ros is 25.9%. in 2012-2016, alrosa exceeded the average sector-specific ros, but in 2017-2018, it was significantly inferior in this aspect. the current liquidity ratio and equity ratio are in the statutory range. the analysis suggests that the company has relatively high operational efficiency. for the research, the dynamics of industry average of the russian companies operational efficiency is important to be noted. the federal tax service of russia regularly monitors average return on average assets (roa) and return on sales (ros). in 2018, roa for hc and gas production totaled 35.3; for railway sector, it totaled 7.6; for pipeline industry, it totaled 12.2; for electrical power distribution and transmission, it totaled 12.5; for information and communication service, it totaled 14.6; for mining operations if totaled 33.6; for other transport and equipment production it totaled 12.0. the average of ros for russia in 2018 totaled: for hc and gas production totaled 20.4; for railway sector it totaled 20.2; for pipeline industry it totaled 4.0; for electrical power distribution and transmission it totaled 5.0; for information and communication service it totaled 7.4; for mining operations if totaled 17.3; for other transport and equipment production it totaled 2.0. although, the industry average of operational efficiency resulted from both private and public companies. therefore, the comparison of the obtained values of efficiency with average ones will be incorrect and inconsistent. the research goal is to determine of efficiency dependency based on changing of government share in company’s assets. thus, observed heterogeneity in operational efficiency among publicly owned companies in the sample leads us to a need in a research of a component influence. 4.2. component influence on operational efficiency in order to study specifics of operational efficiency at publicly owned companies in russia, we put forward a number of assumptions. we tested assumptions using methods of the correlation-regression analysis. the study conceptual framework involves identification of a relationship between a value of the state's share in the authorized capital and a value of each operational efficiency ratio (profitability ratio, return on assets, and return on sales, equity ratio, and current liquidity ratio). numerous evidence-based studies and the findings obtained by us from the analysis of a financial standing of a number of publicly owned russian companies make it possible to put forward the idea that an ownership structure has an essential impact on operational efficiency of these companies. this assumption became the background of the research, which 30 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) presumes the correlation between operational efficiency values and the government share in company’s assets. this idea found its expression in five independent assumptions as follows. meanwhile, this correlation must be inexplicit, as operational efficiency is influenced by many other factors, which are not included in evaluation further. besides, the dependency of operational efficiency on government share should be observed not on the whole sample, but on its part, since differential peculiarities of public administration are common for not all managers, mechanism of business operations takes place to be objectively, socio-economic conditions of work are heterogenic, etc. hypothesis 1: roa depends on the state's share in the company’s capital. the higher the state’s share is, the lower roa is. the calculations completed show that the pair-correlation coefficient between the state's share in the company's ownership and return on assets was -0.57151. in this case, the correlation is inverse and significant against the parameters under consideration. the linear regression equation looks like y = -0.162 x + 14.038. the regression coefficient (b = -0.162) shows that the change in the resulting ratio is elastic. in this case, with the 1-point growth, roa has an average decline of -0.162. the free constant (a = 14.038) formally shows an expected initial rate of the resulting ratio with x (state’s share) = 0. to check whether the linear regression equation is significant, determination coefficient r2 is used and its value is 0.327. next, there is a calculation of the actual value of f-test (4.85). the tabular value of the criterion with degrees of freedom k1=1 and k2=10, ftable = 4.96. as far as the actual value of f > ftabl, the determination coefficient is statistically significant (estimated value of regression equation is statistically reliable). the assumption does not lack support, roa declines with the growth in the state’s share in the company's capital. although, cause-effect relationship is not well-defined. the influence of other factors that are not taken into account in the framework of this regression model is not excluded. hypothesis 2: roе depends on the state’s share in the company’s capital. the higher the state’s share is, the lower roe is. the correlation coefficient between roe and the state's share in the company’s capital was 0.64310. this suggests a measurable correlation between attributes. the linear regression equation looks like y = -0.2715 x + 24.231. the regression coefficient (b = -0.271) shows that the change in the resulting ratio is elastic. in this case, with the 1point growth in the state’s share, roe declines on average by -0.271. with х (state’s share) = 0, initial expected roe is 24.231. determination coefficient r2 is 0.4136. the actual value of f-test is 7.05. the tabular value of the criterion with degrees of freedom k1=1 and k2=10, f tabl = 4.96. as far as the actual value of f > ftabl, the determination coefficient is statistically significant (estimated value of regression equation is statistically reliable). the assumption does not lack support, roa declines with the growth in the state’s share in the company's capital. as in case of hypothesis 1, cause-effect relationship of observed values is questionable. the correlation can be confirmed with expanding of sample and the scopes of observations. hypothesis 3: ros depends on the state’s share in the company’s capital; the higher the state’s share is, the lower ros is. the pair correlation coefficient between the state's share in the company’s capital and the return on sales was -0,23513. according to the typology, the correlation between attributes is poor. operational efficiency of large-scale publicly owned companies 31 copyright ©2019 assa. adv. in systems science and appl. (2019) the linear regression equation looks like y = -0.117 x + 19.001. the regression coefficient (b = -0.117) shows that the change in the resulting ratio is elastic. in this case, with the 1-point growth in the state’s share in the capital, ros declines on average by -0.117. the free constant a = 19.001 while х (state’s share) is 0. determination coefficient r2 is 0.0553. next, there is a calculation of the actual value of f-test (0,59). a tabular value of the criterion with degrees of freedom k1=1 and k2=10, ftable = 4.96. as far as the actual value of f > ftabl, the determination coefficient is not statistically significant (estimated value of regression equation is not statistically reliable). the assumption lacks support. ros does not decline with a growth of the state’s share in the company's capital. hypothesis 4: the equity ratio depends on the state’s share in the company’s capital. the higher the state’s share is, the lower the er is. the correlation between the equity ratio and the state’s share was 0.26209 saying of the poor direct correlation. the linear regression equation is y = 0.00255 x + 0.425. the regression coefficient (b = 0.00255) shows that the change in the resulting ratio is elastic. in this case, with the 1-point growth of the state's share in the capital, the equity ratio has an average increase of 0.00255. if х (state’s share) = 0, an initial expected rate of the equity ratio is 0.425. determination coefficient r2 is 0.0687. the actual value of f-test is 0.704. as far as the actual value of f < ftabl, the determination coefficient is not statistically significant (estimated value of regression equation is not statistically reliable). the assumption lacks support. the equity ratio does not decline with an increase in the state’s share in the company’s capital. hypothesis 5: the current liquidity ratio depends on the state's share in a company's capital. the higher the state’s share is, the lower the current liquidity ratio is. the correlation between the equity ratio and the state’s share in the company’s equity was 0.38002. its level is an evidence of a direct moderate correlation in place. the linear regression equation is y = 0.142 x -3.565. the regression coefficient (b = 0.142) shows that the change of the resulting ratio is elastic. in this case, with a 1-point growth of the state’s share, the current liquidity ratio has an average increase of 0.142. at the same time, the free constant takes a negative value (a = -3.565) when x (state’s share) is 0. determination coefficient r2 is 0.1444. next, there is a calculation of an actual value of f-criterion (1.688). as the actual value of f < ftab, the determination coefficient is not statistically significant (found estimate of regression equation is not statistically reliable). the assumption lacks support. the current liquidity ratio does not decline as the share of public ownership in a company’s capital grows. thus, the component analysis of operational efficiency at publicly owned companies in russia has showed that only return on assets and profitability ratio are dependent on the ownership structure. with the 1%-increase in public ownership, the ratios have a decline by 0.162% and 0.271% respectively. the other operational efficiency ratios that we have considered (return on sales, equity ratio and current liquidity ratio) do not depend on the share of public ownership. 5. conclusion the public sector is a key player in economy of russia. over the past few years, a number of publicly owned companies has had a steady growth, while a share of the public sector in russia's gdp has increased up to 46%. the completed component analysis based on correlation-regression methods shows that the size of the publicly owned share has an overall 32 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) negative influence on operational efficiency ratios. this primarily refers to return on assets and return on equity. dynamics of the other efficiency ratios considered (return on sales, equity ratio and current liquidity ratio) has poor interconnections with the share of national government in a capital of a company. vague results of the research prevent us from concluding in a rigorous manner that publicly run companies are a priori less efficient. however, ceteris paribus, as compared to the private sector, it is more difficult for the national government to be an efficient owner. acknowledgments the russian foundation for basic research financially supported the research under research project no. 19-010-00994 “methods and models to evaluate and forecast public business efficiency in industrial complex”. references [1] albrecht, b. (1996). privatization, coordination and agency costs: the case for participation in eastern europe, international tax and public finance, 3(3), 351-368. https://doi.org/10.1007/bf00418949 [2] avdasheva, s. (2007). rossiyskiye kholdingi: novyye empiricheskiye svidetel'stva [russian holdings: new empirical evidence], voprosy ekonomiki, 1, 98-111 [in russian]. https://doi.org/10.32609/0042-8736-2007-1-98-111 [3] berduygina, o., vlasov, a. & kuzmin, e. (2017). investment capacity of the economy during the implementation of projects of public-private partnership, investment management and financial innovations, 14(3), 189-198. https://www.doi.org/10.21511/imfi.14(3-1).2017.03 [4] boardman. a. & vining a. (1989). ownership and performance in competitive environments: a comparison of the performance of private, mixed and state-owned enterprises, journal of law and economics, 32, 29. https://doi.org/10.1086/467167 [5] borcherding, t., pommerehne, w. & schneider, f. (1982). comparing the efficiency of private and public production: the evidence from five countries, zeitschrift fur nationaloekonomie, 89, 127-156. [6] boyett, i. (1997). the public sector entrepreneur — a definition, international journal of entrepreneurial behavior & research, 3(2), 77-92. [7] bucci, a. & del bo, c. (2012). on the interaction between public and private capital in economic growth, journal of economics, 106(2), 133-152. https://doi.org/10.1007/s00712-011-0239-3 [8] camilleri, m. (2015). valuing stakeholder engagement and sustainability reporting, corporate reputation review, 18(3), 210-222. https://doi.org/10.1057/crr.2015.9 [9] cavaliere, a., maggi, m. & stroffolini, f. (2017). investment-driven mixed firms: partial privatization by local governments, international tax and public finance, 24(3), 459-483. https://doi.org/10.1007/s10797-016-9426-z [10] centre for corporate information disclosure. [online]. available https://www.edisclosure.ru. [11] choi, y., kang, d., chae, h. & kim k. (2008). an enterprise architecture framework for collaboration of virtual enterprise chains, international journal of advanced manufacturing technology, 35(11-12), 1065-1078. https://doi.org/10.1007/s00170006-0789-7 [12] christ, k. & green, r. (2004). public capital and small firm performance, atlantic economic journal, 32(1), 28-38. https://doi.org/10.1007/bf02298616 https://doi.org/10.1007/bf00418949 https://doi.org/10.32609/0042-8736-2007-1-98-111 https://www.doi.org/10.21511/imfi.14(3-1).2017.03 https://doi.org/10.1086/467167 https://doi.org/10.1007/s00712-011-0239-3 https://doi.org/10.1057/crr.2015.9 https://doi.org/10.1007/s10797-016-9426-z https://doi.org/10.1007/s00170-006-0789-7 https://doi.org/10.1007/s00170-006-0789-7 https://doi.org/10.1007/bf02298616 operational efficiency of large-scale publicly owned companies 33 copyright ©2019 assa. adv. in systems science and appl. (2019) [13] chursin, a. & makarov, y. (2015). quantitative evaluation of the firm competitiveness. in management of competitiveness. cham: springer. https://doi.org/10.1007/978-3319-16244-7 [14] di bella, g., dynnikova, o. & slavov, s. (2019). the russian state’s size and its footprint: have they increased? international monetary fund, 2, 11. https://doi.org/10.5089/9781498302791.001 [15] dreyfus, p. (1980). the efficiency of public enterprise: lessons of the french experience. in w.j. baumol (ed.), public and private enterprise in a mixed economy. international economic association series. london: palgrave macmillan. https://doi.org/10.1007/978-1-349-16394-6_21 [16] fabuš, m. & csabay, m. (2018). state aid and investment: case of slovakia, entrepreneurship and sustainability issues, 6(2), 480-488. https://www.doi.org/10.9770/jesi.2018.6.2(1) [17] gnan, l., hinna, a., monteduro, f. & scarozza, d. (2013). corporate governance and management practices: stakeholder involvement, quality and sustainability tools adoption, journal of management & governance, 17(4), 907-937. https://doi.org/10.1007/s10997-011-9201-6 [18] hilkevics, s. & semakina, v. (2019). the classification and comparison of business ratios analysis methods, insights into regional development, 1(1), 48-57. https://www.doi.org/10.9770/ird.2019.1.1(4) [19] hua, h. & cong, p. (2011). information sharing of partnership in supply chain and enhancement of core competitiveness of enterprises. in m. ma (ed.), communication systems and information technology. lecture notes in electrical engineering, 100. berlin, heidelberg: springer. https://doi.org/10.1007/978-3-642-21762-3_115 [20] kapelyushnikov, r. (2001). sobstvennost' i kontrol' v rossiyskoy promyshlennosti [ownership and control in the russian industry], voprosy ekonomiki, 12, 103-124 [in russian]. [21] kashina, i., et al.(2012). avtomatizatsiya protsessov obrabotki informatsii v statistike [automation of information processing in statistics]. moscow: dm k press [in russian]. [22] kontseptsiya planirovaniya vyyezdnykh nalogovykh proverok (prikaz fns rossii ot 30.05.2007 № мм-3-06/333@) [conceptual framework for the on-site tax audit planning (order no. мм-3-06/333@ by fts of russia of may 30, 2007)]. [online]. available https://www.nalog.ru/rn77/about_fts/docs/3897151 [in russian]. [23] kortelainen, m. & leppänen, s. (2013). public and private capital productivity in russia: a non-parametric investigation, empirical economics, 45(1), 193-216. https://doi.org/10.1007/s00181-012-0615-z [24] kuzmin, e., volkova, e. & fomina, a. (2019). research on the concentration of companies in the electric power market of russia, international journal of energy economics and policy, 9(1), 130-136. https://doi.org/10.32479/ijeep.7169 [25] kuznetsov, p. & murav'yev, a. (2002). struktura aktsionernogo kapitala i rezul'taty deyatel'nosti firm v rossii: analiz “golubykh fishek” fondovogo rynka [share capital structure and performance of russian companies: the analysis of blue chips in the stock market]. moscow: rpei [in russian]. [26] li, j., et al. (2018). corporate level: new driving forces and models for enhancing the competitiveness of chinese enterprises. in china’s provincial economic competitiveness and policy outlook for the 13th five-year plan period (2016-2020). research series on the chinese dream and china’s development path. singapore: springer. https://doi.org/10.1007/978-981-13-2664-6_6 [27] li, m., cui, l. & lu, j. (2014). varieties in state capitalism: outward fdi strategies of central and local state-owned enterprises from emerging economy countries, journal of international business studies, 45(8), 980-1004. https://doi.org/10.1057/jibs.2014.14 https://doi.org/10.1007/978-3-319-16244-7 https://doi.org/10.1007/978-3-319-16244-7 https://doi.org/10.5089/9781498302791.001 https://doi.org/10.1007/978-1-349-16394-6_21 https://www.doi.org/10.9770/jesi.2018.6.2(1) https://doi.org/10.1007/s10997-011-9201-6 https://www.doi.org/10.9770/ird.2019.1.1(4) https://doi.org/10.1007/978-3-642-21762-3_115 https://doi.org/10.1007/s00181-012-0615-z https://doi.org/10.32479/ijeep.7169 https://doi.org/10.1007/978-981-13-2664-6_6 https://doi.org/10.1057/jibs.2014.14 34 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) [28] litau, e. (2018). entrepreneurship and economic growth: a look from the perspective of cognitive economics. acm international conference proceeding series, 143-147. https://doi.org/10.1145/3271972.3271978 [29] lópez-iturriaga, f. & rodríguez-sanz, j. (2001). ownership structure, corporate value and firm investment: a simultaneous equations analysis of spanish companies, journal of management and governance, 5(2), 179-20. https://doi.org/10.1023/a:1013078225905 [30] ming, x. (2008). optimal withdrawing path of state-owned capital in economic transition a dynamic model, frontiers of economics in china, 3(1), 38-62. https://doi.org/10.1007/s11459-008-0003-y [31] muller, d. (2007). public choice iii. moscow: su hse. https://doi.org/10.1007/s11127-007-9151-3 [32] narkunienė, j. & ulbinaitė, a. (2018). comparative analysis of company performance evaluation methods, entrepreneurship and sustainability issues, 6(1), 125-138. https://www.doi.org/10.9770/jesi.2018.6.1(10) [33] niskanen, w. (1971). bureaucracy and representative government. chicago: aldineatherton. [34] orekhova, s. & kuzmin, e. (2017). resource investment model in specifics of developing countries. proceedings of the international conference on trends of technologies and innovations in economic and social studies, 488-494. https://doi.org/10.2991/ttiess-17.2017.80 [35] pappe, ya. (2002). rossiyskiy krupnyy biznes kak ekonomicheskiy fenomen: spetsificheskiye cherty, modeli yego organizatsii [russian large-scale business as an economic phenomenon: specifics, models of its organization], problemy prognozirovaniya [forecasting problems], 2, 83-97 [in russian]. [36] pestieau, p., &tulkens, h. (2006) assessing and explaining the performance of public enterprises: some recent evidence from the productive efficiency viewpoint. in p. chander, j. drèze, c.k. lovell, & j. mintz (eds.), public goods, environmental externalities and fiscal competition. boston, ma: springer. [37] puchkarev, d. (2018, october 12). kompanii s gosudarstvennym uchastiyem, skol'ko ikh [publicly-owned companies, their number]. [online]. available https://bcsexpress.ru/novosti-i-analitika/kompanii-s-gosudarstvennym-uchastiem-skol-ko-ikh [in russian]. [38] purlik, v. (2016). operatsionnaya i strategicheskaya effektivnost' biznesa: sovremennaya traktovka [operational and strategic business performance: modern interpretation], risk: resursy, informatsiya, snabzheniye, konkurentsiya [risk: resources, information, supply, competition], 4, 229-232 [in russian]. [39] radygin a.d., et al. (2018). effektivnoye upravleniye gosudarstvennoy sobstvennost'yu v 2018–2024 gg. i do 2035 g. analiticheskiy doklad [effective management of state property in 2018–2024 and until 2035. the analytical report]. мoscow: economic development. centre for strategic development. [online]. available https://www.csr.ru/wpcontent/uploads/2018/02/doklad_effektivnoe_upravlenie_gossobstvennostyu_web.pdf [in russian]. [40] ruzhanskaya, l. & krutikov d. (2006). faktory povysheniya rynochnoy stoimosti ural'skikh kompaniy: problemy praktiki i politiki [multipliers of market value of the ural companies: application and strategy issues]. seriya “nauchnyye doklady: nezavisimyy ekonomicheskiy analiz” [series “scientific reports: independent economic analysis”], 183. moscow: moskovskiy obshchestvennyy nauchnyy fond [moscow public scientific foundation] [in russian]. [41] ruzhanskaya, l., ostanin, i., tychinskaya, t. & shcherbinina, a. (2009). regional businesses and state-owned companies: can the state-owned companies contribute to https://doi.org/10.1145/3271972.3271978 https://doi.org/10.1023/a:1013078225905 https://doi.org/10.1007/s11459-008-0003-y https://doi.org/10.1007/s11127-007-9151-3 https://www.doi.org/10.9770/jesi.2018.6.1(10) https://doi.org/10.2991/ttiess-17.2017.80 https://bcs-express.ru/novosti-i-analitika/kompanii-s-gosudarstvennym-uchastiem-skol-ko-ikh https://bcs-express.ru/novosti-i-analitika/kompanii-s-gosudarstvennym-uchastiem-skol-ko-ikh https://www.csr.ru/wp-content/uploads/2018/02/doklad_effektivnoe_upravlenie_gossobstvennostyu_web.pdf https://www.csr.ru/wp-content/uploads/2018/02/doklad_effektivnoe_upravlenie_gossobstvennostyu_web.pdf operational efficiency of large-scale publicly owned companies 35 copyright ©2019 assa. adv. in systems science and appl. (2019) the businesses' operational efficiency improvement if they have a share in their property and participate in managing the businesses? modern competition, 2(14), 4463. [42] seydalin, r. & musina, d. (2017). analiz tipov integrirovannykh kompaniy v neftegazovom komplekse [analysis of types of integrated companies in the oil and gas sector], problemy nauki [problems of science], 2 (5((18)), 40-41 [in russian]. [43] shapiro, c. & willig, r. (1990). economic rationales for the scope of privatization. the political economy of public sector reform and privatization. london: westview press. [44] shatin, a. (2012). efficiency of state ownership of the large companies in russia: modern aspect, messenger of the chelyabinsk state university, 9(263), 25-28. [45] sherstobitova, o. (2013). problema ekonomicheskoy effektivnosti gosudarstvennogo sektora ekonomiki [the problem of economic efficiency in public sector of economy], naukovedeniye [science studies], 3(16), 42 [in russian]. [46] state-owned enterprises in the eu: lessons learnt and ways forward in a post-crisis context. (2016). [online]. available https://ec.europa.eu/info/sites/info/files/file_import/ip031_en_2.pdf [47] system for professional analysis of markets and companies “spark-interfax”. [online]. available http://www.spark-interfax.ru [48] tatahi, m. (2006). public versus private ownership: theory and performance. in privatisation performance in major european countries since 1980. london: palgrave macmillan. https://doi.org/10.1057/9780230624955 [49] turova, e. (2015). sovremennyy mekhanizm povysheniya effektivnosti gosudarstvennogo predprinimatel'stva (opyt gosudarstvennykh kompaniy stran briks) [modern mechanism for improving efficiency of public sector entrepreneurship: the experience of brics state-owned enterprises], gosudarstvennoye upravleniye. elektronnyy vestnik [e-journal. public administration], 53, 237-255 [in russian]. [50] vickers, j. & yarrow, g. (1988). privatization: an economic analysis. cambridge: mit press. [51] world bank. (2016). russian federation systematic diagnostic: pathways to inclusive growth. [online]. available http://pubdocs.worldbank.org/en/184311484167004822/dec27-scd-paper-eng.pdf [52] world bank. (2016). tfp by sector and ownership, russian federation systematic diagnostic: pathways to inclusive growth. [online]. available https://www.worldbank.org/en/country/russia/publication/systematic-countrydiagnostic-for-the-russian-federation-pathways-to-inclusive-growth [53] yasin ye. (ed.). (2004). strukturnyye izmeneniya v rossiyskoy promyshlennosti [structural changes in the russian industry]. moscow: izd. dom gu vshe [publ. house of higher school of economics] [in russian]. [54] zerbinati, s. & souitaris v. (2005). entrepreneurship in the public sector: a framework of analysis in european local governments, entrepreneurship & regional development, 17(1), 43-64. https://doi.org/10.1080/0898562042000310723 [55] ra expert. (2018). [online]. available https://raexpert.ru/rankingtable/top_companies/2018/main/ [56] ria rating. (2019). [online]. available http://riarating.ru/infografika/20190129/630115992.html https://doi.org/10.1057/9780230624955 https://doi.org/10.1080/0898562042000310723 36 v.e. guseva, e.a. kuzmin, a.i. vlasov copyright ©2019 assa adv. in systems science and appl. (2019) annex a table. operational efficiency ratios at runtime at sample publicly owned companies in 2012-2018 ratio research period ( )s x ( )var x 2012 2013 2014 2015 2016 2017 2018 russian railways (rr) roa 0.33 0.02 -0.93 0.01 0.12 0.30 0.30 0.02 0.17 roe 0.41 0.02 -1.25 0.01 0.17 0.41 0.42 0.03 0.30 ros 4.94 4.27 4.17 5.53 7.43 8.25 7.83 6.06 2.59 er 0.80 0.77 0.73 0.70 0.74 0.72 0.70 0.74 0.00 cr 0.65 0.67 0.74 0.76 0.57 0.56 0.47 0.63 0.01 united shipbuilding corporation (usc) roa -0.37 -0.19 0.19 0.56 0.12 0.03 0.04 0.05 0.09 roe -0.46 -0.28 0.34 1.29 0.34 0.10 0.12 0.21 0.32 ros -275.72 0.12 -42.61 -6.49 -40.49 -3.78 -3.12 -53.16 9960.67 er 0.76 0.59 0.52 0.38 0.35 0.36 0.31 0.47 0.03 cr 1.88 2.58 4.44 10.55 6.15 5.65 10.33 5.94 11.78 gazprom roa 5.69 6.01 1.64 3.20 3.07 2.66 6.80 4.15 3.91 roe 7.21 7.73 2.16 4.38 4.04 3.51 9.88 5.56 7.59 ros 27.14 24.45 23.09 18.73 8.46 2.33 18.02 17.46 81.02 er 0.79 0.77 0.74 0.72 0.80 0.72 0.66 0.74 0.00 cr 2.21 2.41 2.28 2.35 2.02 1.63 1.96 2.12 0.07 rossneft roa 13.06 3.65 7.85 2.78 1.02 1.32 3.94 4.80 18.35 roe 25.34 10.34 36.54 17.16 6.69 8.64 24.83 18.51 119.56 ros 9.40 5.61 3.62 3.05 0.35 3.25 5.08 4.34 7.84 er 0.50 0.28 0.17 0.15 0.15 0.15 0.16 0.22 0.02 cr 4.36 1.24 1.26 2.08 1.33 1.31 1.43 1.86 1.30 united aircraft corporation (uаc) roa -0.03 0.26 2.75 -2.49 -0.70 0.04 -6.63 -0.97 8.61 roe -0.04 0.36 4.11 -3.59 -0.96 0.06 -9.81 -1.41 18.86 ros -2.64 -7.37 -3.06 -3.32 -0.47 -0.62 -0.72 -2.60 5.91 er 0.74 0.69 0.65 0.73 0.73 0.69 0.66 0.70 0.00 cr 6.97 11.84 2.80 3.42 2.82 2.52 1.76 4.59 13.04 transneft roa 1.17 1.22 1.14 1.07 2.58 5.40 0.98 1.94 2.64 roe 7.56 7.61 7.68 7.98 17.55 31.53 5.46 12.20 88.09 ros 2.68 1.71 1.71 2.76 4.40 5.78 5.75 3.54 3.11 er 0.16 0.16 0.14 0.13 0.16 0.18 0.18 0.16 0.00 cr 0.97 0.26 0.19 1.28 0.83 0.87 0.79 0.74 0.15 inter rao roa -3.58 -13.58 0.12 1.05 24.60 3.95 4.40 2.42 132.90 roe -4.11 -15.03 0.13 1.12 25.51 4.10 4.92 2.38 149.38 ros -2.68 -4.70 -1.43 6.39 0.60 -0.19 1.64 -0.05 12.54 er 0.86 0.95 0.94 0.94 0.98 0.94 0.86 0.92 0.00 cr 1.57 3.29 2.56 4.49 9.99 4.49 2.14 4.08 8.04 rushydro roa 2.08 4.50 3.68 3.43 4.65 3.82 3.64 3.69 0.71 roe 4.44 5.80 4.58 4.11 5.54 4.53 4.38 4.77 0.41 ros 40.57 43.49 37.17 40.27 37.85 42.08 40.52 40.28 4.88 er 0.79 0.76 0.84 0.83 0.85 0.84 0.82 0.82 0.00 cr 3.04 3.67 6.11 5.49 7.40 3.80 7.63 5.31 3.42 rosseti roa -2.11 -141.67 -30.77 -10.78 72.97 -3.24 -2.93 -16.93 4079.08 roe -2.36 -150.73 -31.01 -12.12 82.46 -3.53 -3.20 -17.21 4763.67 ros 44.41 46.59 38.12 33.34 5.37 5.49 5.83 25.59 369.26 er 0.87 1.00 0.99 0.82 0.92 0.92 0.92 0.92 0.00 cr 1.91 18.56 10.99 5.11 9.17 17.47 36.22 14.20 130.64 rostelecom roa 6.21 6.43 5.44 3.87 1.78 1.48 0.92 3.73 5.52 roe 11.45 13.28 11.80 7.99 5.38 4.64 2.11 8.09 17.83 operational efficiency of large-scale publicly owned companies 37 copyright ©2019 assa. adv. in systems science and appl. (2019) ratio research period ( )s x ( )var x 2012 2013 2014 2015 2016 2017 2018 ros 17.87 16.32 14.42 11.63 4.30 4.37 3.99 10.41 37.20 er 0.52 0.44 0.48 0.49 0.19 0.46 0.41 0.43 0.01 cr 0.62 1.34 0.61 0.48 0.49 0.66 0.58 0.68 0.09 tatneft roa 13.80 12.30 14.78 13.98 15.34 13.52 24.81 15.50 17.76 roe 18.86 15.96 18.10 16.50 17.92 16.26 31.33 19.28 29.42 ros 29.61 26.89 23.37 25.80 18.76 21.48 31.49 25.34 20.12 er 0.75 0.79 0.84 0.85 0.86 0.81 0.78 0.81 0.00 cr 5.63 4.50 4.02 3.51 3.31 3.75 3.14 3.98 0.74 alrosa roa 11.65 8.68 5.09 3.84 23.24 2.73 5.29 8.65 50.67 roe 19.96 15.39 9.67 7.63 40.18 4.18 8.15 15.02 150.92 ros 40.05 38.46 42.64 45.67 19.64 14.76 20.13 31.62 166.08 er 0.56 0.56 0.49 0.51 0.64 0.67 0.63 0.58 0.00 cr 2.26 1.50 3.64 2.45 4.26 2.03 1.58 2.53 1.09 note: er is equity ratio, cr is current ratio, ( )s x is time average value; ( )var x is dispersion. adv syst sci appl 2021; 01; 95-112 published online at https://ijassa.ipu.ru. improving dcp haze removal scheme by parameter setting and adaptive gamma correction cheng-hsiung hsieh1*, yi-hung chang2 1) department of computer science and information science, chaoyang university of technology, taichung, taiwan e-mail:chhsieh@cyut.edu.tw 2) department of computer science and information science, chaoyang university of technology, taichung, taiwan e-mail: s10827616@ cyut.edu.tw abstract: recently, single-image haze removal based on the dark channel prior (dcp), originally proposed by he et. al., has attracted much attention in the image restoration community. this dehazing algorithm, called the dcp scheme here, is well-known to have four main problems in its dehazed images: artifacts, hue distortion, color over-saturation, and halos. in this paper, an improved dcp (idcp) is proposed to deal with the four aforementioned problems, by setting the model parameters, i.e. scaling factors and window size and smoothing factor of a guided image filter in the dcp scheme. note that a dehazed image is generally dim and low in contrast. an adaptive gamma correction (agc) is introduced for dehazed image enhancement. the proposed idcp and agc are used to create the idcp/agc scheme, in which the idcp scheme performs haze removal and the agc enhances the dehazed image. the idcp/agc scheme was justified through extensive experiments and compared with the dcp scheme, an optimization-based scheme, and two learning-based schemes on two datasets. the results indicated that the proposed scheme is subjectively and objectively superior to the comparison schemes. keywords: dark channel prior, single image haze removal, adaptive gamma correction 1. introduction image haze removal by far is an active research field in the image restoration and image enhancement community. the haze mainly resulted from the adversary weather condition degrades the visual quality of an image. a hazy image is generally of low contrast and visibility, which may affects the following computer visionary applications, such as video-based surveillance systems, automatic driver’s assistance systems, and object tracking/recognition systems. since haze removal is in great need, much work has been done in this field. a popular class of haze removal schemes relies on statistical observations and thus assumptions, such as contrast assumption [1], un-correlation assumption [2], dark channel prior [3], and color line [4]. besides, in [5], a pioneer work based on machine learning for single image haze removal was proposed. in [6], a multi-scale convolutional neural networks was presented to estimate parameters in the haze image model. in [7], an end-to-end framework based on a deep learning neural network for single image haze removal was introduced. in [8], another end-to-end framework, called all-in-one dehazing network, was proposed. in [9-10], generative adversarial networks were applied to image haze removal. for more other schemes, one may consult survey papers in [11-13]. this paper will concentrate on the popular scheme based on dark channel prior, which was originated by he et al. in [3]. the scheme will be called the dcp scheme hereafter. with its * corresponding author: chhsieh@cyut.edu.tw 96 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) simplicity, the dcp scheme has attracted much attention recently. however, the original dcp scheme has several problems needed to be improved, i.e. computational cost, artifacts, color distortion, and halos. to reduce the computational cost, in [14] he et al. used a guided image filter (gif) to replace the soft matting algorithm in [3] for transmission map refinement. the artifacts and color distortion usually occur in the sky region of a dehazed image, while the problem of halos happens in large depth discontinuities. to improve the dcp scheme, many researchers have proposed a variety of solutions to relieve the aforementioned problems. some of them are listed below. in [15], with a soft segmentation, the transmission map estimation was divided into two part: one for the sky region and the other for the non-sky region. for the sky region, a luminance model was applied to estimate the transmission map, while the dark channel prior was used to estimate the transmission map in the non-sky region. then the two transmission maps were fused to obtain the final transmission map. in [16], a sky segmentation scheme based on quad-tree splitting was presented to deal with the problem of the dcp scheme in the sky region, where the transmission map was refined by an edge preserving gif. in [17], the sky region is segmented through a quad-tree decomposition and region-growing scheme. then the estimate of transmission map obtained from the dark channel was modified based on the segmented sky region. the atmospheric light was estimated through the segmented sky region. in [18], the transmission map was found through minimization of an energy function with a piecewise smooth assumption. in [19], a segmentation scheme based on the binary image from gradient image of input image was employed to separate sky regions, through which an estimate of atmospheric light was found. with the help of edge information, an initial transmission map was obtained and then refined by a gif. in [20], a light intensity reverse algorithm was proposed to deal with the artifact problem in the sky region or white objects. by a threshold, the region of interest was segmented, whose rgb pixel values were replaced with a constant manually specified by the user. then the dcp scheme was applied to find the dehazed image. in [21], the initial transmission map was found by the 5 × 5 minimum filter and then refined by the sobel filter and mean filter. to avoid color distortion, a pixel-based adaptive lower bound for the final transmission map was calculated through a constrained piece-wise linear function. in [22], a saliency detection was proposed to extract white objects according to superpixel intensity contrast. then the atmospheric light and transmission map were estimated with the preprocessed image. besides, an adaptive upper bound was given to avoid over-exposure in dehazed images. in [23], a two-stage transmission map estimation was employed. in the first stage, a dehazed image was found by the dcp scheme. in the second stage, the pixel-based transmission map was obtained from the dehazed image. then morphological operations were applied to find the final transmission map. finally, the image was restored by the final transmission map and atmospheric light estimated by the dcp scheme in the first stage. in the aforementioned schemes, most of them use a segmentation skill to separate sky and nonsky regions to avoid the artifacts and color distortion found in the sky region, since many researchers conjecture that the problem is due to the inappropriateness of dark channel prior for the sky region. moreover, many researchers attribute the halo problem to the large depth discontinuities in the input image. however, the two conjectures are arguable. though the dark channel prior is not suitable in the sky region or white objects, it may not suggest that the dcp scheme is inappropriate accordingly. in [3], it says that the sky and non-sky regions can be handled gracefully, even the dcp is not a good prior for sky regions. this paper will justify that the statement is correct, through the proposed improved dcp (idcp) scheme, where no segmentation is required for the sky and non-sky regions. in addition, problems of artifacts, hue distortion, color over-saturation, and halos in the dcp scheme will be relieved by the proposed idcp scheme through model parameter setting. section 3 will have the details. in general, a dehazed image is dim and of low contrast. conventional image enhancement methods, such as histogram equalization (he) and gamma correction (gc), are not suitable for dehazed image enhancement, since over enhancement and color distortion happen most of improving dcp haze removal scheme... 97 copyright ©2021 assa. adv. in systems science and appl. (2021) time. even so, some researchers have tried to apply he and gc to image haze removal. in [24], an adaptive gc was applied in the transmission map estimation. in [25], the haze image model and a gc were combined to form a model, called concise gamma-correction-based dehazing model. in their results, color distortion can be found in the given examples. in [26], a contrast limited adaptive he (clahe) was applied to visibility enhancement of the dehazed image. however, over enhancement is found, when sky region is large. this paper will propose an adaptive gamma correction, which can be applied to dehazed image enhancement without color distortion. this paper has four main contributions. first, an alternative solution is proposed to the problems of artifact and color distortion in sky regions, that is, by an adaptive scaling factor in the atmospheric light estimation. this is different from the sky and non-sky region segmentation schemes as described previously. second, artifacts in the sky region and color distortion in a dehazed image is relieved by an adaptive scaling factor in the estimation of initial transmission map. thus, no sky and non-sky region segmentation is required. this provides a way to mitigate this problem. third, the halo problem happened in large discontinuities is solved by a gif parameter setting in the transmission refinement. this doing gives a solution to the halo problem. fourth, an adaptive gamma correction (agc) is presented to alleviate the color distortion in the conventional gamma correction. then the proposed agc is employed to dehazed image enhancement. this brings for the possibility to apply gamma correction in the image haze removal field. this paper consists of two parts. the first part is to present an improved dcp (idcp) dehazing scheme and the second part is to introduce an adaptive gamma correction (agc) as a postprocessing to enhance the dehazed image. in section 2, the dcp scheme is briefly reviewed. in section 3, the proposed idcp dehazing scheme and the proposed agc for dehazed image enhancement are introduced. then the two schemes are combined as the idcp/agc scheme. next, the proposed idcp/agc scheme is extensively justified by two image datasets and compared with four dehazing schemes in section 4. finally, conclusion is made in section 5. 2. review of the dcp scheme the dcp scheme is based on the following haze image model, 𝑰(𝑥) = 𝑱(𝑥)𝑡(𝑥) + 𝑨[1 − 𝑡(𝑥)] (2.1) where 𝑰(𝑥) is the observed hazy image; 𝑱(𝑥) is the haze-free image; 𝑨 is the global atmospheric light or simply atmospheric light; 𝑡(𝑥) = 𝑒−𝛽𝑑(𝑥) is the transmission map which represents the portion of the non-scattered light to the camera; 𝛽 is the scattering coefficient of the atmosphere and 𝑑(𝑥) is the scene depth at position 𝑥. the dcp scheme is based on the following observation. in general, at least one of rgb components has very low intensity in the non-sky pixels of a haze-free image. this statistical observation is called dark channel prior in [3], which can be obtained through a block minimum filter. the result is called dark channel. fig. 2.1 is an example to justify the dark channel prior where the 15×15 minimum filter is employed. (a) (b) fig. 2.1. image forest (a) original (b) the corresponding dark channel the derivation to obtain the initial transmission map in [3] is briefly reviewed in the following. assume a hazy image 𝑰 is in the rgb color space and without sky regions or white objects. when considering one component of 𝑰, eq. (2.1) can be rewritten as 98 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) 𝐼𝑐(𝑥) = 𝐽𝑐(𝑥)𝑡(𝑥) + 𝐴𝑐 − 𝐴𝑐𝑡(𝑥) (2.2) where 𝑐 ∈ {𝑅, 𝐺, 𝐵}. next, eq. (2.1) is normalized by 𝐴𝑐 and is obtained as 𝐼𝑐(𝑥) 𝐴𝑐 = 𝐽𝑐(𝑥) 𝐴𝑐 𝑡(𝑥) + 1 − 𝑡(𝑥) (2.3) assume the transmission map within the block ω(𝑥) is a constant. through a block minimum filter, the dark channel in eq. (2.3) is obtained as min 𝑦∈ω(𝑥) min 𝑐 [ 𝐼𝑐(𝑥) 𝐴𝑐 ] = �̃�(𝑥) min 𝑦∈ω(𝑥) min 𝑐 [ 𝐽𝑐(𝑥) 𝐴𝑐 ] + 1 − �̃�(𝑥) (2.4) where �̃�(𝑥) denotes as the initial transmission map. by the property of dark channel prior, the dark channel for haze-free image 𝐽𝑐(𝑥) approaches to zero and thus min 𝑦∈ω(𝑥) min 𝑐 [ 𝐽𝑐(𝑥) 𝐴𝑐 ] = 0 is considered in eq. (2.4). consequently, by eq. (2.4) the initial transmission map can be estimated from the input image 𝑰 as �̃�(𝑥) = 1 − min 𝑦∈ω(𝑥) min 𝑐 [ 𝐼𝑐(𝑥) 𝐴𝑐 ] (2.5) to avoid the halo problem, the initial transmission map �̃�(𝑥) is further refined by the soft matting algorithm in [3] or a guided image filter (gif) in [14]. for more details, one may consult [3, 14]. given image 𝑰 in the rgb color space, the implementation steps of the dcp scheme are given as follows. step 1. find the initial block dark channel through a block minimum filter as 𝐼ω 𝑑𝑎𝑟𝑘(𝑥) = min 𝑦∈ω(𝑥) min 𝑐 [𝐼𝑐(𝑦)] (2.6) where ω(𝑥) is a 𝑁 × 𝑁 window centered at 𝑥 and 𝑐 ∈ {𝑅, 𝐺, 𝐵}. in [3], 𝑁 = 15 is employed. step 2. estimate the atmospheric light 𝑨 = [𝐴𝑅 𝐴𝐺 𝐴𝐵] by 𝐼ω 𝑑𝑎𝑟𝑘(𝑥). find the 0.1% pixels of the highest values in 𝐼ω 𝑑𝑎𝑟𝑘(𝑥); trace back to the corresponding pixels in image 𝑰; and find the pixel with the highest intensity as the estimate of 𝑨. step 3. calculate the normalized block dark channel as 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) = min 𝑦∈ω(𝑥) min 𝑐 [ 𝐼𝑐(𝑦) 𝐴𝑐 ] (2.7) step 4. obtain the initial transmission map as �̃�(𝑥) = 1 − 𝜔 × 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) (2.8) where 0 < 𝜔 ≤ 1 is a scaling factor. in [3], 𝜔 is set to 0.95. step 5. refine the initial transmission map �̃�(𝑥) by the soft matting algorithm in [3] or by the gif) in [14] to obtain the final transmission map 𝑡(𝑥). when the gif is used, the settings given in [14] are as follows: the input image 𝑰 as the guidance image; the window size 𝑁 = 20; the smoothing factor 𝜖 = 0.001. step 6. recover the scene radiance as 𝐽𝑐(𝑥) = 𝐼𝑐(𝑥)−𝐴𝑐 max[𝑡0,𝑡(𝑥)] + 𝐴𝑐 (2.9) where 𝑡0 is a user-defined lower bound of 𝑡(𝑥). it is set to 0.1 in [3]. 3. the proposed idcp scheme and agc this section will introduce the proposed idcp haze removal scheme and the agc for dehazed image enhancement in this study. in section 3.1, the idcp scheme is presented to improve the dcp performance. then, the post-processing agc scheme is proposed to enhance the visual quality of the dehazed image by the idcp scheme in section 3.2. 3.1. the idcp dehazing scheme it is well-known that for better performance, at least four problems should be appropriately addressed in the dcp scheme: estimation of 𝑨, estimation of 𝑡(𝑥), the sky/non-sky region handling, and halos. the solutions in the idcp scheme are given below. improving dcp haze removal scheme... 99 copyright ©2021 assa. adv. in systems science and appl. (2021) 3.1.1. atmospheric light estimation in the dcp scheme, the hue distortion is often found in the dehazed images, especially in the sky region. to see the problem, image village is served as an example, which is shown in fig. 3.1(a). the dehazed image by the dcp is shown in fig. 3.1(b) where the sky region is of hue distortion. in [27], it has proven that the hue distortion in a dehazed image results from the estimation error of 𝑨. in other words, the problem of hue distortion in the dcp scheme is from an inappropriate estimation of 𝑨. fortunately, the problem can be relieved by introducing a scaling factor 𝛼 on 𝑨. in the dcp scheme, the scaling factor of 𝑨 can be considered as 𝛼 = 1 as in eq. (2.6). when 𝛼 = 0.85, the dehazed image is shown in fig. 3.1(c). obviously, the hue distortion is relieved by scaling 𝛼 value. thus, an adaptive scaling factor 𝛼𝑎 will be employed in the idcp scheme. by experiments, 𝛼𝑎 = min[(𝜇1)0.0975, 0.975] works well most of time, where 𝜇1 = mean[𝐼1 𝑑𝑎𝑟𝑘(𝑥)] and 𝐼1 𝑑𝑎𝑟𝑘 the pixel-based dark channel. besides, 𝐼ω 𝑑𝑎𝑟𝑘(𝑥) in step 2 of the dcp scheme is replaced by 𝐼1 𝑑𝑎𝑟𝑘 to find 𝑨 for efficiency. (a) (b) (c) fig. 3.1. image village (a) original (b) 𝛼 = 1 (c) 𝛼 = 0.85 3.1.2. initial transmission map estimation in the dcp scheme, the fixed scaling factor 𝜔 = 0.95, as in eq. (2.8), is used to find the initial transmission map �̃�(𝑥) according to the aerial perspective phenomenon. however, we find that artifacts in the sky region and color distortion in a dehazed image are caused by the inappropriate fixed scaling factor 𝜔 = 0.95 . thus, it introduces the related distortions accordingly. to see the problem, image dogs is given as an example, which is shown in fig. 3.2(a). after the dcp scheme, the dehazed image is depicted in fig. 3.2(b), which shows the artificial contours and color distortion due to 𝜔 = 0.95. fortunately, an adaptive scaling factor 𝜔𝑎 can help the situation. with 𝜔 = 0.65, the result is given in fig. 3.2(c). in the idcp scheme, an adaptive scaling factor 𝜔𝑎 will be employed to estimate initial transmission map �̃�(𝑥) . as a rule of thumb, 𝜔𝑎 = min [(𝜇0.9)0.325, 0.95] is employed where 𝜇0.9 = mean[𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) ≤ 0.9] and 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) the normalized block-based dark channel. (a) (b) (c) fig. 3.2. image dogs (a) original (b) 𝜔 = 0.95 (c) 𝜔 = 0.65 3.1.3. sky and non-sky region handling in [3], it says that the sky and non-sky regions can be handled gracefully, even the dcp is not a good prior for sky regions. the statement can be verified as follows. in a bright sky region, the pixel intensity 𝐼𝑐(𝑥) → 1; the atmospheric light 𝐴𝑐 → 1; the normalized dark channel 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) → 1; the transmission map 𝑡(𝑥) → 1, and thus the dehazed scene can be found as 𝐽𝑐(𝑥) = 𝐼𝑐(𝑥)−𝐴𝑐 max[𝑡0,𝑡(𝑥)] + 𝐴𝑐 → 𝐴𝑐 (3.1) which is generally consistent with the real situation. consequently, by eq. (3.1) there is no need to handle sky and non-sky regions separately, though many published papers headed 100 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) toward the way to the sky/non-sky region segmentation. in other words, the dcp scheme can be applied to the sky and non-sky regions equally well, if appropriate parameters are employed. by our observations, the problems of color over-saturation distortion, hue distortion, artifacts, and over-exposure found in the sky region come from inappropriate estimations of 𝑨 and 𝑡(𝑥). this can be verified by fig. 3.1 and fig. 3.2. in the idcp scheme, these problems can be relieved by setting parameters appropriately. 3.1.4. initial transmission map refinement in general, halos occur in the large depth discontinuities of a dehazed image. it is observed that the halo problem in the dcp scheme comes from an inappropriate parameter setting in the gif for transmission map refinement. in the gif, three parameters are guidance image 𝑰𝑔, window size 𝑁, and smoothing factor 𝜖. in the dcp scheme, 𝑰𝑔 is the input image 𝑰; 𝑁 = 20; 𝜖 = 0.001. by observations, we find that appropriately setting parameters 𝑁 and 𝜖 in the gif is able to deal with the halo problem. a large 𝑁 in the gif reduces the halo effect in a dehazed image. however, it introduces minor color over-saturation. fortunately, it can be solved by using a large 𝜖. a. the effect of 𝑁 here, image house, shown in table 3.1, is used as an example to investigate the effect of the gif parameters, 𝑁 and 𝜖, on a dehazed image. in the simulation, the dcp scheme is employed with the gif setting 𝑁 = 20 and 𝜖 = 0.001, as suggested in [14]. to investigate the effect of parameter 𝑁, 𝜖 is fixed at 0.001. the dehazed images with different values of 𝑁, i.e. 20, 40, and 55, are given in table 3.1, where the corresponding refined transmission 𝑡(𝑥) are also shown for comparison. in the case of 𝑁 = 20, the dehazed house has a severe halo problem at large depth discontinuities. as 𝑁 increases to 40, the halos diminish significantly. when 𝑁 = 55, the halo is not visible at all. besides, one may observe that the significant edges in 𝑡(𝑥) become more clear and more details appear, as 𝑁 varies from 20 to 55. it implies that a large window size may avoid the halo problem, as expected. consequently, 𝑁 = 55 will be used in the proposed idcp scheme. though the window size 𝑁 = 55 is able to deal with the halo problem, minor color over-saturation is found in the dehazed house, as shown in the fourth row of table 3.1. fortunately, the introduced distortion can be avoided by setting smoothing factor 𝜖. the following subsection has the details. table 3.1. the effect of 𝑁 in the gif on the dehazed image house. gif setting 𝑡(𝑥) 𝐽(𝑥) original image 𝑁 = 20 𝜖 = 0.001 𝑁 = 40 𝜖 = 0.001 𝑁 = 55 𝜖 = 0.001 improving dcp haze removal scheme... 101 copyright ©2021 assa. adv. in systems science and appl. (2021) b. the effect of 𝜖 the effect of smoothing parameter 𝜖 on the dehazed house is investigated here. in the simulation, 𝜖 is set to 0.001, 0.05, and 0.1, where 𝑁 is fixed at 55. the corresponding dehazed images and their corresponding 𝑡(𝑥) are given in table 3.2, respectively. by table 3.2, one can observe that the details of 𝑡(𝑥) has been smoothened more and more, as 𝜖 increases from 0.001 to 0.1. besides, the problem of color over-saturation gradually relieves, as 𝜖 becomes larger. it suggests that the problem of color over-saturation can be dealt with a large 𝜖, i.e. 0.1. thus, the smoothing factor 𝜖 = 0.1 will be employed in the proposed idcp scheme. table 3.2. the effect of 𝜖 in the gif on the dehazed image house. gif setting 𝑡(𝑥) 𝐽(𝑥) original image 𝑁 = 55 𝜖 = 0.001 𝑁 = 55 𝜖 = 0.05 𝑁 = 55 𝜖 = 0.1 3.1.5. implementation of the idcp scheme with the above discussion, the proposed idcp scheme is summarized in the following implementation steps, where input image 𝑰 is assumed in the rgb color space. step 1. find the pixel-based dark channel as 𝐼1 𝑑𝑎𝑟𝑘(𝑥) = min 𝑐 [𝐼𝑐(𝑥)] (3.2) where 𝑐 ∈ {𝑅, 𝐺, 𝐵}. step 2. find the maximum in 𝐼1 𝑑𝑎𝑟𝑘(𝑥) and its corresponding pixel in 𝑰, 𝒑𝑚𝑎𝑥. then estimate the atmospheric light as 𝑨 = [𝐴𝑅 𝐴𝐺 𝐴𝐵] = 𝛼𝑎 × 𝒑𝑚𝑎𝑥 , where 𝛼𝑎 = min[(𝜇1)0.0975, 0.975] and 𝜇1 = mean[𝐼1 𝑑𝑎𝑟𝑘(𝑥)]. step 3. calculate the normalized block-based dark channel as 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) = min 𝑦∈ω(𝑥) min 𝑐 [ 𝐼𝑐(𝑦) 𝐴𝑐 ] (3.3) step 4. obtain the initial transmission map as �̃�(𝑥) = 1 − 𝜔𝑎 × 𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) (3.4) where 𝜔𝑎 = min [(𝜇0.9)0.325, 0.95] and 𝜇0.9 = mean[𝐼ω̅ 𝑑𝑎𝑟𝑘(𝑥) ≤ 0.9]. step 5. find the final transmission map 𝑡(𝑥) through refining �̃�(𝑥) by the gif with the guidance image 𝐼1 𝑑𝑎𝑟𝑘(𝑥), the window size 𝑁 = 55, and the smoothing parameter 𝜖 = 0.1. step 6. estimate the dehazed image as 𝐽𝑐(𝑥) = 𝐼𝑐(𝑥)−𝐴𝑐 max[𝑡0,𝑡(𝑥)] + 𝐴𝑐 (3.5) where 𝑡0 = 0.1 is a user-defined lower bound of 𝑡(𝑥). 102 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) in comparison with the dcp scheme described in section 2, the proposed idcp scheme has at least four main differences. first, the pixel-based dark channel 𝐼1 𝑑𝑎𝑟𝑘(𝑥) is used to estimate the atmospheric light 𝑨 with an adaptive scaling factor 𝛼𝑎. second, 𝐼1 𝑑𝑎𝑟𝑘(𝑥) is employed as the guidance image in the gif to refine the initial transmission map �̃�(𝑥). this helps improve the efficiency. third, an adaptive scaling factor 𝜔𝑎 is applied in the estimation of initial transmission map �̃�(𝑥) to avoid the artifacts and color distortion happened in the sky region. this doing makes the proposed idcp scheme is able to deal with the sky and non-sky regions equally well, without segmentation. fourth, the gif setting with large 𝑁 = 55 is employed to relieve the halo problem, while large 𝜖 = 0.1 is to handle the color over-saturation. with those modifications, the proposed idcp scheme shows much better performance than the dcp scheme. this will be verified in section 4. 3.2. dehazed image enhancement with an agc a dehazed image generally is dimmer than its original image. consequently, an adaptive gamma correction (agc) is used to improve visual quality of dehazed images. the agc is modified from a conventional gamma correction (cgc). given image 𝑰𝑖, whose elements are denoted as 𝐼𝑖 𝑐(𝑥) and 𝑐 ∈ {𝑅, 𝐺, 𝐵}, a cgc transforms 𝐼𝑖 𝑐(𝑥) to a new value 𝐼𝑜 𝑐(𝑥) as 𝐼𝑜 𝑐(𝑥) = [ 𝐼𝑖 𝑐(𝑥) 𝐼𝐻 𝑐 ] 𝑔 𝐼𝐻 𝑐 (3.6) where 𝐼𝐻 𝑐 = max 𝑥 [𝐼𝑖 𝑐(𝑥)] is the maximum value in component 𝑐. when put the dynamic ranges of input image 𝑰𝑖 and output image 𝑰𝑜 into account, eq. (3.6) can be rewritten in a more general form as 𝐼𝑜 𝑐(𝑥) = [ 𝐼𝑖 𝑐(𝑥)−𝐼𝐿 𝑐 𝐼𝐻 𝑐 −𝐼𝐿 𝑐 ] 𝑔 [𝐼𝑜,𝐻 𝑐 − 𝐼𝑜,𝐿 𝑐 ] + 𝐼𝑜,𝐿 𝑐 (3.7) where 𝐼𝐿 𝑐 = min 𝑥 [𝐼𝑖 𝑐(𝑥)] and 𝐼𝐻 𝑐 = max 𝑥 [𝐼𝑖 𝑐(𝑥)]. notations 𝐼𝑜,𝐿 𝑐 and 𝐼𝑜,𝐻 𝑐 are user-defined lower limit and upper limit of 𝐼𝑜 𝑐(𝑥) for the output image 𝑰𝑜, respectively. the superscript 𝑔 in eq. (3.7) is a user-defined factor. it is well-known that the cgc suffers from the problem of color distortion. therefore, it hinders the application to dehazed image enhancement. an example, image women, is given in fig. 3.3(a), where the dehazed image, shown in fig. 3.3(b), is obtained by the proposed idcp scheme. the dehazed image, post-processed by the cgc, is shown in fig. 3.3(c) which has a noticeable color distortion. (a) (b) (c) (d) fig. 3.3. image women (a) original (b) after the proposed idcp scheme (c) with the cgc (d) with the agc to relieve the color distortion in the cgc, an adaptive gamma correction (agc) is introduced here. by our observations, the color distortion in the cgc results from the inappropriate set of upper and lower limits, [𝐼𝐿 𝑐 𝐼𝐻 𝑐 ]. consequently, the proposed agc replaces the set [𝐼𝐿 𝑐 𝐼𝐻 𝑐 ] with set [𝐼𝐿 𝐼𝐻] where 𝐼𝐿 = min 𝑐 [𝐼𝐿 𝑐] and 𝐼𝐻 = max 𝑐 [𝐼𝐻 𝑐 ]. with the set [𝐼𝐿 𝐼𝐻], eq. (3.7) is modified as 𝐼𝑜 𝑐(𝑥) = [ 𝐼𝑖 𝑐(𝑥)−𝐼𝐿 𝑐 𝐼𝐻−𝐼𝐿 ] 𝑔𝑎 [𝐼𝑜,𝐻 𝑐 − 𝐼𝑜,𝐿 𝑐 ] + 𝐼𝑜,𝐿 𝑐 (3.8) where 𝑔𝑎 is an adaptive factor. the way to determine parameter 𝑔𝑎 in eq. (3.8) is discussed below. note that the parameter 𝜔𝑎 for initial transmission map estimation affects the strength of haze removal significantly. that is, a larger 𝜔𝑎 results in a stronger dehazing effect and vice versa. a stronger haze removal makes the dehazed image dimmer in general. consequently, 𝑔𝑎 should be inversely improving dcp haze removal scheme... 103 copyright ©2021 assa. adv. in systems science and appl. (2021) related to 𝜔𝑎. to increase the brightness, 𝑔𝑎 is restricted to be less than 1 and greater than 0.7 to avoid artifacts. by experiments, 𝑔𝑎 = max[(1 − 𝜔𝑎)0.095, 0.707] is employed in the agc. with this 𝑔𝑎 , the proposed agc is applied to fig. 3.3(b), whose result is shown in fig. 3.3(d). as expected, the color distortion is avoided and visual quality is enhanced. since the proposed agc is a post-processing scheme, it can be added to the proposed idcp scheme after step 6 in section 3.1.5. that is, step 7. enhance the dehazed image �̂� by the agc. the combination of the idcp scheme and the agc will be call the idcp/agc scheme. 4. results and discussions in this section, the proposed idcp/agc scheme is justified by the datasets reside in [13] and kedema in [28]. the reside dataset is a large-scale benchmark for single image dehazing algorithms, which consists of synthetic and natural images. in the following experiments, the indoor training set (its) [29] is employed, that has 10,000 clear images and 100,000 synthetic hazy images generated by eq. (2.1) with various 𝐴 and 𝛽 . besides, the outdoor training set (ots) [30] is also used, which consists of 3,981 clear images and 136,160 synthetically generated hazy images. by the large amount of images, the proposed idcp/agc scheme is justified and compared with four recently reported haze removal schemes in section 4.1, where both objective and subjective evaluation are considered. the four compared schemes are the conventional dcp scheme in [14], the optimization-based scheme in [31], which will be called rro scheme, the learning based scheme in [32], which will be called cap scheme, and the deep learning based scheme dehazenet (dnet for short) in [33]. in section 4.2, the kedema dataset is used, which consists of 25 natural hazy images with different scenarios, to further verify for the proposed idcp/agc scheme, objectively and subjectively, and compared with the dcp, rro, cap, and dnet schemes. in addition to the subjective comparison, the performances of the related schemes are also justified objectively. to compare the related schemes in different aspects, the objective assessments include full-reference methods: the peak signal-to-noise ratio (psnr) and the structural similarity index measure (ssim) in [34]; half-reference method: the dehazing quality index (dhqi) in [35]; no-reference methods: the blind/reference-less image spatial quality evaluator (brisque) in [36] and the integrated local natural image quality evaluator (il-niqe) in [37]. the compared results are given below. 4.1 results and comparisons with the reside dataset to justify the performance of the proposed idcp/agc scheme, an extensive experiment is conducted with the reside dataset in this section, where the results from the four compared schemes are also given for objective and subjective comparisons. first, the proposed idcp/agc scheme and the compared schemes are run with the its dataset in section 4.1.1 and then ots dataset in section 4.1.2. 4.1.1. results with the its dataset in this subsection, the idcp/agc scheme is verified by the its dataset. the results for the four compared schemes are also given for comparison. the first part shows the objective results, where the performance indices psnr, ssim, dhqi are better for a higher value, while brisque, il-niqe are the less the better. in the second part, subjective results will give for visual comparison. a. objective comparison with the its dataset the objective comparisons with psnr, ssim, dhqi, brisque, il-niqe are shown in table 4.1, for the proposed idcp/agc scheme and the compared schemes, where the ranking is shown in parentheses. by the results, the dcp scheme has the best and second best in 104 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) brisque and il-niqe, respectively. however, it has the worst ranking in psnr, ssim, and dhqi. that is, the dcp scheme lies on either the worst end or the best end with the average ranking 3.6. for the rro scheme, it ranks the second best in brisque; the third place in dhqi and il-niqe; the fourth place in psnr and ssim. it ends with the average ranking 3.2. as for the cap scheme, it has the second best psnr; the third place in ssim; but the fourth place in dhqi and brisque, and the worst ranking in il-niqe. the average ranking is 3.6. as for the dnet, it has the best psnr and ssim; the second best in dhqi; but the fourth place in il-niqe and the worst ranking in brisque. its average ranking is 2.6. the ranking of the proposed idcp/agc scheme falls within the top three among full-reference, semi-reference, and no-reference assessments. it suggests that the performance of the proposed idcp/agc scheme has a relatively stable performance among the compared schemes. overall, the proposed idcp/agc scheme has the best performance, then dnet, rro, cap, and dcp schemes, in terms of the average ranking (ar). table 4.1. objective comparison with the its dataset idcp/agc dcp rro cap dnet psnr 19.1669(3) 16.4216(5) 18.2363(4) 19.5537(2) 20.3097(1) ssim 0.8742(2) 0.8374(5) 0.84681(4) 0.8733(3) 0.8909(1) dhqi 59.8120(1) 55.6912(5) 57.3045(3) 56.9135(4) 59.2964(2) brisque 15.6748(3) 13.4386(1) 13.9044(2) 19.1409(4) 22.6275(5) il-niqe 32.6799(1) 33.4390(2) 33.5820(3) 33.9286(5) 35.3254(4) ar 2(1) 3.6(4) 3.2(3) 3.6(4) 2.6(2) b. subjective comparison with the its dataset in addition of objective comparison, the subjective comparison is also given for the its dataset. here, six images are selected to subjectively evaluate the proposed idcp/agc scheme and the compared schemes. the six images include two less hazy images (no. 1 and no. 2), two moderate hazy images (no. 3 and no. 4), and two heavy hazy images (no. 5 and no. 6), which are shown in table 4.2. for reference, the image filenames in its are also given in table 4.2 with the associated psnr for each selected image. by table 4.2, the dcp scheme tends to have hue distortion and color over-saturation in all six dehazed images. the rro scheme has hue distortion in images 2, 4, and 5; minor color over-saturation for images 1 and 3. the dehazing effect of the cap scheme seems not strong enough, especially for the heavy hazy images 5 and 6. basically, the dnet scheme performs well in the cases of less hazy and moderate images, but fails to remove haze in heavy hazy images 5 and 6. as for the proposed idcp/agc scheme, it provides better visual quality in the dehazed images. table 4.2. subjective comparison with the its dataset (6 selected images) no clear input idcp/agc dcp rro cap dnet 1 0025_08_0.7048 psnr=31.26 16.32 24.97 22.42 20.17 2 1075_03_0.7312 psnr=30.72 13.29 13.96 24.53 25.63 3 0499_05_0.8353 psnr=22.63 15.50 21.89 19.78 21.47 improving dcp haze removal scheme... 105 copyright ©2021 assa. adv. in systems science and appl. (2021) 4 0691_09_0.8315 psnr=24.53 19.06 18.08 20.01 23.91 5 0803_10_0.9055 psnr=22.97 13.28 13.83 19.42 19.44 6 0962_02_0.7819 psnr=22.56 18.43 15.69 14.01 14.77 next, three worse cases, shown in table 4.3, for the proposed idcp/agc scheme are given and discussed. for image 1, better visual quality is for the proposed idcp/agc scheme, even though the cap and dnet schemes have better psnr. it suggests that the psnr is not consistent with the subjective evaluation in some cases. similar results are found in images 2 and 3. the dehazed images by the proposed idcp/agc scheme are obviously better than the compared schemes, even some of their psnr are higher. the reason might be that the proposed idcp/agc scheme enhances the contrast and brightness in the dehazed images. thus, it increases the mean squared error between the dehazed images and their clear images, and lower psnr results. consequently, in this study five objective assessments are employed to have a more fair comparison among the proposed idcp/agc scheme and the compared schemes. table 4.3. subjective comparison with the its dataset (3 chosen cases) no clear input idcp/agc dcp rro cap dnet 1 0244_02_0.8649 psnr=25.28 16.14 15.09 33.10 34.54 2 0388_09_0.7129 psnr=20.84 16.78 30.50 25.11 29.11 3 0146_04_0.7433 psnr=15.72 16.07 27.77 24.16 21.88 4.1.2. results with the ots dataset to have more understanding the performance, the proposed imdcp/aglgc scheme is further justified by the ots dataset and compared with the dcp, rro, cap, and dnet schemes. the objective comparison is given first and then the subjective comparison follows. a. objective comparison with the ots dataset as in the its dataset, five objective assessments, psnr, ssim, dhqi, brisque, and ilniqe, for the proposed idcp/agc scheme, dcp, rro, cap, and dnet are given in table 4.4. the dcp scheme has the second best in il-niqe and the third place in brisque, but with the worst psnr, ssim, and dhqi. this ends up with average ranking 4. for the rro scheme, it has the second best in brisque, the third place in il-niqe, and the fourth place in psnr, ssim, and dhqi. its average ranking is 3.4. the cap scheme has the best ssim and the second best psnr, the third place dhqi, but the fourth place in brisque and ilniqe, which results in the average ranking 2.8. for the dnet, it has the best psnr and dhqi, the second best in ssim, but the worst brisque and il-niqe. the average ranking is 2.8. as for the proposed idcp/agc scheme, it has the best brisque and il-niqe, the second best dhqi, and the third place in psnr and ssim. its average ranking is 2. by the average 106 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) ranking, the best one is the proposed idcp/agc scheme, and then dnet or cap, rro, and dcp at final. by table 4.4, it implies that the proposed idcp/agc scheme has more stable performance than the compared schemes, among the five objective assessments. table 4.4. objective comparison with the ots dataset idcp/agc dcp rro cap dnet psnr 22.5973(3) 15.9340(5) 19.0970(4) 23.3067(2) 23.7265(1) ssim 0.9111(3) 0.8456(5) 0.8779(4) 0.9236(1) 0.9161(2) dhqi 50.6845(2) 44.9371(5) 47.7001(4) 49.9232(3) 51.6284(1) brisque 16.1972(1) 16.9098(3) 16.3732(2) 16.9410(4) 19.0443(5) il-niqe 21.2804(1) 21.3985(2) 21.4482(3) 21.4876(4) 23.1376(5) ar 2(1) 4(4) 3.4(3) 2.8(2) 2.8(2) b. subjective comparison with the ots dataset for subjective comparison, six images are selected from the ots dataset, which are shown in table 4.5. the six images are with different haziness. two images (no. 1 and no. 2) are less hazy, two images (no. 3 and no. 4) are of moderate haziness, and two images (no. 5 and no. 6) have more haziness. the corresponding psnr for each image is also given in table 7 for reference. in the given images, the dcp scheme has artifacts in image 2, color over-saturation in images 3 and 4 , and hue distortion in images 5 and 6. the rro scheme has less dehazing effect in image 1, artifacts in image 2, hue distortion in images 3 and 5, and minor color oversaturation in images 4 and 6. the cap scheme has a dark chiffon in images 5 and 6. similarly, the dnet scheme has similar, but minor, results in images 5 and 6. for the rest of images, the dnet generally works better than the cap scheme, in terms of visual quality of the dehazed images. as for the proposed idcp/agc scheme, it gives better subjective results than the four compared schemes, even images 3, 5, and 6 have less psnr, which is caused by brightness and contrast enhancement. table 4.5. subjective comparison with the ots dataset (6 selected images) no clear input idcp/agc dcp rro cap dnet 1 1178_0.8_0.04 psnr=34.09 26.94 22.88 25.56 24.14 2 0014_1_0.04 psnr=23.21 14.52 19.62 21.63 19.14 3 0255_1_0.1 psnr=26.11 18.55 18.09 25.65 27.06 4 0250_0.85_0.12 psnr=25.72 15.90 17.88 23.07 24.06 5 0027_0.8_0.2 psnr=21.02 14.38 23.01 20.68 25.01 improving dcp haze removal scheme... 107 copyright ©2021 assa. adv. in systems science and appl. (2021) 6 0004_0.8_0.2 psnr=19.29 12.77 17.00 21.47 14.77 for the ots dataset, three worse cases in the proposed idcp/agc scheme are given and discussed here. the three images are shown in table 4.6, where their psnr are given as well. obviously, the visual quality of the dehazed images by the proposed idcp/agc scheme is better than the four compared schemes, even though some of psnr are lower. for images 1 and 2, the ground true image, i.e. clear image, is somewhat hazy, not really clear. thus, it affects the psnr calculation. in other words, the schemes have less dehazing effect, like the rro, cap, and dehazenet, achieve better psnr. for image 3, it is a night shot. the proposed idcp/agc scheme not only removes haze but also enhances brightness and contrast. consequently, it increases the mean squared error and thus less psnr results. it explains why the proposed idcp/agc scheme takes the third place in psnr, as shown in table 4.6. that is why this study adopts five different objective assessments and the average ranking is used to evaluate the overall performance for a fair comparison. table 4.6. subjective comparison with the ots dataset (3 chosen cases) no clear input idcp/agc dcp rro cap dnet 1 0267_0.95_0.16 psnr=25.10 14.97 20.02 30.58 26.96 2 0150_0.85_0.2 psnr=16.36 16.38 16.69 32.36 19.74 3 0113_0.85_0.12 psnr=15.94 29.45 19.40 18.13 26.31 4.2 results and comparisons with the kedema dataset in this section, the proposed idcp/agc scheme is applied to the natural images, where no ground true images are available. in the experiments, the kedema dataset, which contains 25 natural hazy images, is employed to verify the proposed idcp/agc scheme further, whose results are compared with the dcp, rro, cap, and dnet schemes as previously. 4.2.1 objective comparison with the kedema dataset three objective metrics, dhqi, brisque, and il-niqe, are calculated with the dehazed images obtained by the proposed idcp/agc and the compared schemes. note that the full reference metrics, psnr and ssim, are not applicable, since ground true images are not available. for those three metrics do not require ground true images, the results are shown in table 4.7. by table 4.7, the dcp scheme ranks between 3 and 5; the rro scheme lies within top three ranking; the cap scheme takes the best in brisque and the fourth place in dhqi and il-niqe; the dnet scheme has the worst ranking in brisque and il-niqe, and the third place in dhqi; the proposed idcp/agc scheme has the best dhqi, and the second best both in brisque and il-niqe. by the average ranking, the proposed idcp/agc scheme outperforms the compared schemes, and the rro, cap, dcp, and dnet schemes follow. again, the proposed idcp/agc scheme provides a better and more stable performance than the compared schemes. 108 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) table 4.7. objective comparison with the kedema dataset idcp/agc dcp rro cap dnet dhqi 60.5390(1) 57.7240(5) 60.1320(2) 58.9870(4) 61.1010(3) brisque 11.8990(2) 12.5420(4) 12.3470(3) 10.9520(1) 13.3190(5) il-niqe 23.9100(2) 24.6400(3) 23.5640(1) 25.6000(4) 27.7350(5) ar 1.7(1) 4(4) 2(2) 3(3) 4.3(5) 4.2.2 subjective comparison with the kedema dataset for subjective comparison, the 25 dehazed images for the kedema dataset are shown in table 4.8, which are obtained from the proposed idcp/agc scheme and the four compared schemes. in general, the dcp scheme suffers from the problems of color over-saturation, halos, and artifacts throughout the given examples. the rro scheme has artificial contours in the sky region of image 1, hue distortion in image 5, color over-saturation in image 7, less dehazed results in images 9, 11. for the cap scheme, it has artificial contours in the sky region of image 1, and less dehazed results in images 3 to 6, 8, 18 to 23. generally, the cap scheme has a relatively weak dehazing performance. the dnet scheme has artificial contours in the sky region of image 1, less dehazed results in images 3, 4, 6, 11, 18, 20, 21, 22, and 23. as for the proposed idcp/agc scheme, it generally gives a stable and better result than the compared schemes in visual quality, which is demonstrated in the aspects of naturalness, brightness, and contrast, as shown in table 4.8. table 4.8. subjective comparison with the kedema dataset no input idcp/agc dcp rro cap dnet 1 2 3 4 5 6 7 improving dcp haze removal scheme... 109 copyright ©2021 assa. adv. in systems science and appl. (2021) 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 110 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) 24 25 5. conclusion in this paper, an improved dcp (idcp) scheme was presented to solve the four problems in the dcp scheme: artifacts, hue distortion, color over-saturation, and halos. the dehazed image by the idcp was further enhanced by an adaptive gamma correction (agc). the overall dehazing scheme is called idcp/agc. the proposed idcp/agc scheme was extensively justified by the synthetic hazy images in the its and ots datasets, and the natural images in the kedema dataset. besides, the proposed idcp/agc scheme was compared with four recently reported haze removal schemes objectively and subjectively. to have a balanced comparison, five objective assessments, including full-reference, half-reference, and noreference methods, were employed. the results indicated that the objective evaluation was for the proposed idcp/agc scheme in terms of the average ranking, while the subjective comparison showed that the proposed idcp/agc scheme was relatively stable and able to provide a better visual quality in the given image datasets. in the future, an optimization algorithm will be investigated for the heuristic parameter setting in the proposed idcp/agc scheme. references 1. tan, r.t. (2008). visibility in bad weather from a single image, proc. of ieee conference on computer vision and pattern recognition, anchorage, ak, 1-8, doi: 10.1109/cvpr.2008.4587643 2. fattal, r. (2008). single image dehazing, acm transactions on graphics, 27(3), 721 729, 2008, doi: 10.1145/1360612.1360671 3. he, k., sun, j., and tang, x. (2011). single image haze removal using dark channel prior, ieee transactions on pattern analysis and machine intelligence, 33(12), 2341-2353, doi: 10.1109/tpami.2010.168 4. fattal, r. (2014). dehazing using color-lines," acm transactions on graphics, article number: 13, doi: 10.1145/2651362 5. zhu, q., mai, j., and shao, l. (2015). a fast single image haze removal algorithm using color attenuation prior, ieee transactions on image processing, 24(11) 3522-3533, doi: 10.1109/tip.2015.2446191 6. ren, w., liu, s., zhang, h., pan, j., cao, x., et. al. (2016). single image dehazing via multi-scale convolutional neural networks, proc. of europe conference on computer vision, 154-169, 2016, doi: 10.1007/978-3-319-46475-6_10 7. cai, b., xu, x., jia, k., qing, c., and tao, d. (2016). dehazenet: an end-to-end system for single image haze removal, ieee transactions on image processing, 25(11), 51875198, doi: 10.1109/tip.2016.2598681 8. li, b., peng, x., wang, z., xu, j., and feng, d. (2017). aod-net: all-in-one dehazing network, proc. of ieee international conference on computer vision, venice, 47804788, doi: 10.1109/iccv.2017.511 9. chen, y., patel, a.k., and chen, c. (2019). image haze removal by adaptive cyclegan, proc. of asia-pacific signal and information processing association annual summit and conference, lanzhou, china, 1122-1127, doi: 10.1109/apsipaasc47483.2019.9023296 improving dcp haze removal scheme... 111 copyright ©2021 assa. adv. in systems science and appl. (2021) 10. liu, z., xiao, b., alrabeiah, m., wang, k., and chen, j. (2019). single image dehazing with a generic model-agnostic convolutional neural network, ieee signal processing letters, 26(6), 833-837, doi: 10.1109/lsp.2019.2910403 11. lee, s., yun, s., nam, j.-h., won, c.s., and jung, s.-w. (2016). a review on dark channel prior based image dehazing algorithms, eurasip journal on image and video processing, article number: 4, doi: 10.1186/s13640-016-0104-y 12. singh, d., and kumar, v. (2018). comprehensive survey on haze removal techniques, multimedia tools and applications, 77, 9595–9620, doi: 10.1007/s11042-017-5321-6 13. li, b., ren, w., fu, d., tao, d., feng, d., et. al. (2019). benchmarking single-image dehazing and beyond, ieee transactions on image processing, 28(1), 492-505, doi: 10.1109/tip.2018.2867951 14. he, k., sun, j., and tang, x. (2013). guided image filtering, ieee transactions on pattern analysis and machine intelligence, 35(6), 1397-1409, doi: 10.1109/tpami.2012.213 15. zhu, y., tang, g., zhang, x., jiang, j., and tian, q. (2017). haze removal method for natural restoration of images with sky, neurocomputing, 275, 499-510, doi: 10.1016/j.neucom.2017.08.055 16. yuan, h., liu, c., guo, z., and sun, z. (2017). a region-wised medium transmission based image dehazing method, ieee access, 5, 1735-1742, doi: 10.1109/access.2017.2660302 17. wang, w., yuan, x., wu, x., and liu, y. (2017). dehazing for images with large sky region, neurocomputing, 238, 365-376, doi: 10.1016/j.neucom.2017.01.075 18. zhu, m., he, b., and wu, q. (2018). single image dehazing based on dark channel prior and energy minimization, ieee signal processing letters, 25(2), 174-178, doi: 10.1109/lsp.2017.2780886 19. xiao, j., zhu, l., zhang, y., liu, e., and lei, j. (2017). scene-aware image dehazing based on sky-segmented dark channel prior, iet image processing, 11(12), 1163-1171, doi: 10.1049/iet-ipr.2017.0058 20. liu q. (2018). a light intensity reverse algorithm for improving dark channel prior dehazing, proc. of the 11th international congress on image and signal processing, biomedical engineering and informatics, beijing, china, 1-9, doi: 10.1109/cispbmei.2018.8633085 21. chen, y., li, z., bhanu, b., tang, d., peng, q., et. al. (2018). improve transmission by designing filters for image dehazing, proc. of ieee the 3rd international conference on image, vision and computing, chongqing, china, 374-378, doi: 10.1109/icivc.2018.8492834 22. zhang, l., wang, s., and wang, x. (2018). saliency-based dark channel prior model for single image haze removal, iet image processing, 12(6), 1049-1055, doi: 10.1049/ietipr.2017.0959 23. salazar-colores, s., cabal-yepez, e., ramos-arreguin, j.m., botella, g., ledesmacarrillo, l.m., et. al. (2019). a fast image dehazing algorithm using morphological reconstruction, ieee transactions on image processing, 28(5), 2357-2366, doi: 10.1109/tip.2018.2885490 24. putra, o.v., musthafa, a., and pradhana, f. r. (2019). ‘a hybrid approach on single image dehazing using adaptive gamma correction, j. phys.: conf. ser. 1381 012030, doi: 10.1088/1742-6596/1381/1/012030 25. ju, m., ding, c., zhang, d., and guo, y.j. (2018). gamma-correction-based visibility restoration for single hazy images, ieee signal processing letters, 25(7), 1084-1088, doi: 10.1109/lsp.2018.2839580 26. kapoor, r., gupta, r., son, l.h., kumar, r., and jha, s. (2019). fog removal in images using improved dark channel prior and contrast limited adaptive histogram equalization, multimedia tools and applications, 78, 23281-23307, doi: 10.1007/s11042-019-7574-8 112 c.-h. hsieh, y.-h. chang copyright ©2021 assa. adv. in systems science and appl. (2021) 27. shen, y., wu, x., and deng, x. (2015). analysis on spectral effects of dark-channel prior for haze removal, proc. of ieee international conference on image processing, quebec city, qc, 2945-2949, doi: 10.1109/icip.2015.7351342 28. ma, k., liu, w., and wang, z. (2015). perceptual evaluation of single image dehazing algorithms, proc. of ieee international conference on image processing, quebec city, qc, 3600-3604, doi: 10.1109/icip.2015.7351475 29. indoor training set, available https://www.dropbox.com/sh/mrtguzk1a1l11o6/aaarckahsqlz34psgm5muuza?dl=0. 30. outdoor training set, available https://www.dropbox.com/s/86qp410gen5u2uk/its.zip?dl=0. 31. shin, j., kim, m., paik, j., and lee, s. (2020). radiance–reflectance combined optimization and structure-guided l0-norm for single image dehazing, ieee transactions on multimedia, 22(1), 30-44, doi: 10.1109/tmm.2019.2922127 32. zhu, q., mai, j., and shao, l. (2015). a fast single image haze removal algorithm using color attenuation prior, ieee transactions on image processing, 24(11), 3522-3533, doi: 10.1109/tip.2015.2446191 33. cai, b., xu, x., jia, k., qing, c., and tao, d. (2016). dehazenet: an end-to-end system for single image haze removal, ieee transactions on image processing, 25(11), 51875198, doi: 10.1109/tip.2016.2598681 34. wang, z., bovik, a.c., sheikh, h.r., and simoncelli, e. p. (2004). image quality assessment: from error visibility to structural similarity, ieee transactions on image processing, 13(4), 600-612, doi: 10.1109/tip.2003.819861 35. min, x., zhai, g., gu, k., yang, x., and guan, x. (2019). objective quality evaluation of dehazed images, ieee transactions on intelligent transportation systems, 20(8), 28792892, doi: 10.1109/tits.2018.2868771 36. mittal, a., moorthy, a.k., and bovik, a.c. (2012). "no-reference image quality assessment in the spatial domain, ieee transactions on image processing, 21(12), 46954708, doi: 10.1109/tip.2012.2214050 37. zhang, l., zhang, l., and bovik, a.c. (2015). a feature-enriched completely blind image quality evaluator, ieee transactions on image processing, 24(8), 2579-2591, doi: 10.1109/tip.2015.2426416 https://www.dropbox.com/sh/mrtguzk1a1l11o6/aaarckahsq-lz34psgm5muuza?dl=0 https://www.dropbox.com/sh/mrtguzk1a1l11o6/aaarckahsq-lz34psgm5muuza?dl=0 https://www.dropbox.com/s/86qp410gen5u2uk/its.zip?dl=0 1. introduction 2. review of the dcp scheme 3. the proposed idcp scheme and agc 3.1. the idcp dehazing scheme 3.1.1. atmospheric light estimation 3.1.2. initial transmission map estimation 3.1.3. sky and non-sky region handling 3.1.4. initial transmission map refinement a. the effect of 𝑁 b. the effect of 𝜖 3.1.5. implementation of the idcp scheme 3.2. dehazed image enhancement with an agc 4. results and discussions 4.1 results and comparisons with the reside dataset 4.1.1. results with the its dataset a. objective comparison with the its dataset b. subjective comparison with the its dataset 4.1.2. results with the ots dataset a. objective comparison with the ots dataset b. subjective comparison with the ots dataset 4.2 results and comparisons with the kedema dataset 4.2.1 objective comparison with the kedema dataset 4.2.2 subjective comparison with the kedema dataset 5. conclusion microsoft word 1378-article text-6924-1-18-20240328 adv syst sci appl 2024; 01; 142-162 published online at https://ijassa.ipu.ru. a new face swap detection technique for digital images rasha thabit 1,2*, heba mohammed fadhil 3, hassan falah fakhruldeen 4,5, akram hatem shather6, mohanad a. al-askari7 1) dijlah university college, baghdad, iraq 2) al-iraqia university, baghdad, iraq 3) university of baghdad, baghdad, iraq 4) imam ja’afar al-sadiq university, baghdad, iraq 5) university of kufa, kufa, iraq 6) al kitab university, altun kopru, kirkuk, iraq 7) university of al-anbar, al-anbar, iraq abstract: in recent years, the rapid development of deep learning-based face image manipulation algorithms and applications became one of the challenges that are facing information forensics and information security systems. using these applications, one can easily swap the face in a digital image with another face for different intentions where most of them are malicious intentions. different face swap detection techniques have been presented in recent years to check the authenticity of the face in a digital image. most of the available techniques are machine-learning or deep-learning based which makes them vulnerable to false detection results in addition to the time-consuming training process. in this paper, a new technique for face swap detection (fsd) is presented based on the image watermarking process. the proposed technique consists of two main algorithms called embedding and authentication algorithms. several experiments have been conducted to evaluate the performance of the proposed technique and to prove its efficiency in detecting fake faces. the proposed technique outperforms various deep-learning-based techniques because no training is required and the detection accuracy is 100 %. the performance of the proposed fsd technique is promising therefore it is applicable in different practical applications. keywords: information security, information forensics, deep-fake detection, face swap detection, face image manipulation detection. 1. introduction biotechnology and biometric information have been utilized in different verification systems such as security and financial industries [1]. the biometric information can be generated using extrinsic sources such as iris, fingerprint, and face or using intrinsic sources such as hand-vein, finger-vein, and palm-vein [2,3]. over the years, the facial recognition systems [4] have been widely used to identify individuals or their feelings using images or videos processing techniques, however, identity theft and delivering fake information have been considered as challenges to these systems [5,6]. face swap is a type of deep-fakes that refers to the process of replacing the face of one person with another in a digital image [7,8]. the face swap has several advantages in the movies and games industry [9], however, it has been considered as one of the most dangerous attacks that are required to be detected effectively especially when it is applied with malicious intentions. the fake face images can be used for fake videos and news to destroy the reputation of people, threaten them, identity theft, and many other malicious intentions [10–13]. over the years, different face-swap algorithms and applications * corresponding author: rashathabit@yahoo.com a new face swap detection technique for digital images 143 copyright ©2024 assa. adv. in systems science and appl. (2024) have been presented consequently different face-swap detection techniques have been implemented to serve the data security and digital media forensics fields [14–17]. the research community in the last few years witnessed an increased interest in face swap detection because these techniques became requested by many institutions and companies [17]. in [18], a face-swap detection technique based on machine learning has been presented. in this work, 83 landmarks for each face are extracted and used for training different classifiers. different machine learning techniques have been tested such as support vector machine (svm) [19], random forest (rf) [20], and multi-layer perceptron’s (mlp) [21]. the images are classified into two classes that are innocent class and swapped images class. two types of classifiers (i.e., linear and non-linear) have been tested to find the best. the results of the nonlinear classifier were better than that of the linear classifier, however, both types of classifiers have recorded false detection results and the best accuracy that have been obtained was around 92 % for only a specific images dataset. in [22], another face-swap detection technique has been presented based on deep learning and the error level analysis (ela) process. the main principle of this technique is related to the errors that are generated in the manipulated face images. when a face region is cut and another face is pasted in its place, an error level will acquire between the manipulated area and the surrounding area. the residual network (resnet-18) has been used with custom dense layers for the classification process. to reduce the time for training which may reaches days or weeks, the authors suggested the use of transfer learning. the results of this work were promising; however, false detection results have been also recorded and the best accuracy that have been obtained was around 97 % for only a specific images dataset. in [23], a face-swap detection technique based on a convolution neural network (cnn) has been presented. the method of detection consists of two stages that are preprocessing stage and the classifier stage. in the preprocessing stage, the features of the face are extracted and the alignment process is applied [24, 25]. in the classifier stage, mobilenet-like cnn pretrained on an imagenet has been used [26, 27]. this technique obtained better accuracy results compared to previous techniques based on mesonet and xception networks [28, 29], however, false detection results have been obtained and the best accuracy that have been obtained was between 98 % to 99 % based on the images dataset. the main problems of this technique are the use of a specific images’ dataset and the time complexity for training and testing processes. as explained before, there are some limitations in the machine-learning and deep-learningbased face swap detection techniques such as the false detection results, the high time complexity for training and testing, the need for high-quality images for training, the need for large datasets, and others. to avoid these limitations, this paper presents a new face-swap detection (fsd) technique based on image watermarking. the main idea of the proposed technique is inspired by the region-of-interest (roi) based medical image authentication techniques [30–32]. in this paper, the face area will be considered as roi and its authentication data will be extracted and embedded in the area outside roi. the data generation, embedding, extraction, and authentication processes are all based on the watermarking techniques that have been presented in [33–35]. to detect the face area two face detection algorithms are tested to choose the one with the best performance. the proposed technique can distinguish between innocent and fake faces in digital images without the need for training and with 100 % accuracy thus it outperforms the machine learning and deep-learning-based techniques. the proposed technique can be applied for any image regardless its quality which increased its ability in practical applications. the rest of the paper is organized as follows: in section 2, the related works are presented; in section 3, the algorithms of the proposed fsd technique are explained in details; in section 4, the experiments and their discussion are presented; and finally, section 5 illustrates the conclusions that have been drawn from this research. 144 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) 2. related works the proposed fsd technique starts by detecting the face area in the digital image. to detect the face region two widely used face detection algorithms have been tested to choose the one with the best performance. the face detection algorithms are briefly explained in the following subsection. on the other hand, inspired by the medical image authentication technique presented in [34], the proposed fsd adopted slantlet transform (slt)-based watermarking. slt-based watermarking techniques [33], [35–39] proved their efficiency in various applications, therefore, this type of watermarking has been adopted as part of the implemented algorithms in the proposed fsd technique. the second subsection presents the slt-based embedding and extraction algorithms for a block of image. 2.1. face detection two well-known face detection algorithms are tested to choose the best one in order to be adopted as the first part of the proposed fsd technique. the adaboost-based algorithm [40] and multi-task cascaded cnn (mtcnn)-based algorithm [41] are tested for different face images. in [40], the detection algorithm starts by calculating the integral image for the input image followed by extracting the features using the harr-like filters. then a small number of the generated features is selected using the adaboost algorithm. to extract the promising regions of the image, the cascade structure is used for the complex classifiers. after several processing steps, the face regions are selected. in [41], a joint face detection and alignment algorithm is presented in which a shallow cnn algorithm is applied to generate the candidate windows. then a more complex cnn algorithm is used to reject the non-face windows. thereafter, another robust cnn algorithm is applied as the final touch-up to refine the result and generate the landmark positions. samples of the preliminary tests for the abovementioned algorithms are shown in fig. 2.1. the results proved that the mtcnn-based algorithm performs better in comparison with the adaboost-based algorithm where the latter has missed some of the faces. therefore, the mtcnn-based algorithm has been chosen to be applied as the first stage in the proposed fsd technique. 2.2. slt-based watermarking the slt-based watermarking algorithms have been applied in different applications and they proved their efficiency in comparison with different other algorithms in terms of visual quality, robustness, and time complexity [35], [42–44]. based on the previous studies, we suggested the used of slt-based watermarking in face swap detection algorithms. in the proposed fsd technique, the slt-based watermark embedding and extraction algorithms for a single image block have been adapted from [34]. the adopted slt-based embedding and extraction algorithms are explained in table 2.1 and table 2.2, respectively. a new face swap detection technique for digital images 145 copyright ©2024 assa. adv. in systems science and appl. (2024) fig. 2.1. preliminary test results for mtcnn-based and adaboost-based face detection algorithms. table 2.1. slt-based watermark embedding algorithm for a single block [34]. input: original image block (size 16×16 pixels) and binary sequence of 64 bits. output: watermarked image block (size 16×16 pixels). step 1 read the input image block 𝐵 and the binary sequence 𝐵𝑖𝑛 . step 2 transform 𝐵 using slt matrix as follows: 𝑇 = [𝑆𝐿𝑇 ] [𝐵] [𝑆𝐿𝑇 ] where 𝐵 and 𝑇 are the original and watermarked blocks, 𝑆𝐿𝑇 and 𝑆𝐿𝑇 are the slt matrix and its transpose both of size (16×16). step 3 divide the coefficients in 𝑇 into four subbands called (𝐿𝐿, 𝐻𝐿, 𝐿𝐻, 𝑎𝑛𝑑 𝐻𝐻). 𝐿𝐿 = 𝑇 (1: 8, 1: 8) 𝐻𝐿 = 𝑇 (1: 8, 9: 16) 𝐿𝐻 = 𝑇 (9: 16, 1: 8) 𝐻𝐻 = 𝑇 (9: 16, 9: 16) where 𝐿𝐿, 𝐻𝐿, 𝐿𝐻, 𝑎𝑛𝑑 𝐻𝐻 are low-low, high-low, low-high, and high-high subbands, respectively. step 4 for 𝑥 = 1 𝑡𝑜 64 𝑏 = 𝐵𝑖𝑛 (𝑥) for 𝑖 = 1 𝑡𝑜 8 for 𝑗 = 1 𝑡𝑜 8 𝐷1 = 𝐻𝐿(𝑖, 𝑗) − 𝐿𝐻(𝑖, 𝑗) if 𝑏 = 1 and 𝐷1 < 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 , then increase 𝐻𝐿(𝑖, 𝑗) and decrease 𝐿𝐻(𝑖, 𝑗) as follows: 𝑁𝑒𝑤𝐻𝐿(𝑖, 𝑗) = 𝐻𝐿(𝑖, 𝑗) + 𝑁𝑒𝑤𝐿𝐻(𝑖, 𝑗) = 𝐿𝐻(𝑖, 𝑗) − 146 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) if 𝑏 = 1 and 𝐷1 ≥ 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑, then do-nothing. 𝐷2 = 𝐿𝐻(𝑖, 𝑗) − 𝐻𝐿(𝑖, 𝑗) if 𝑏 = 0 and 𝐷2 < 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 , then increase 𝐿𝐻(𝑖, 𝑗) and decrease 𝐻𝐿(𝑖, 𝑗) as follows: 𝑁𝑒𝑤𝐻𝐿(𝑖, 𝑗) = 𝐻𝐿(𝑖, 𝑗) − 𝑁𝑒𝑤𝐿𝐻(𝑖, 𝑗) = 𝐿𝐻(𝑖, 𝑗) + if 𝑏 = 0 and 𝐷2 ≥ 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑, then do-nothing. note: the 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 variable is used for controlling the visual quality and the robustness of the embedded watermark. the 𝑇ℎ𝑟𝑒𝑠ℎ𝑜𝑙𝑑 value that has been adopted in [34] is 3 because it gives a good compromise between visual quality and robustness. step 5 replace the original 𝐻𝐿 and 𝐿𝐻 subbands in 𝑇 with 𝑁𝑒𝑤𝐻𝐿 and 𝑁𝑒𝑤𝐿𝐻 subbands. save the resultant matrix as 𝑁𝑒𝑤𝑇 . step 6 apply inverse slt on 𝑁𝑒𝑤𝑇 to obtain the watermarked image block as follows: 𝑊 = [𝑆𝐿𝑇 ] [ 𝑁𝑒𝑤𝑇 ] [𝑆𝐿𝑇 ] where 𝑊 is the watermarked image block that carries a binary sequence of 64 bits. table 2.2. slt-based watermark embedding algorithm for a single block [34]. input: original image block (size 16×16 pixels) and binary sequence of 64 bits.. output: watermarked image block (size 16×16 pixels). step 1 read the input watermarked image block 𝑊 . step 2 transform 𝑊 use the ng slt matrix as follows: 𝑇 = [𝑆𝐿𝑇 ] [𝑊 ] [𝑆𝐿𝑇 ] where 𝑇 is the transformed watermarked image block. step 3 divide the coefficients in 𝑇 into four subbands called (𝐿𝐿, 𝐻𝐿, 𝐿𝐻, 𝑎𝑛𝑑 𝐻𝐻). 𝐿𝐿 = 𝑇 (1: 8, 1: 8) 𝐻𝐿 = 𝑇 (1: 8, 9: 16) 𝐿𝐻 = 𝑇 (9: 16, 1: 8) 𝐻𝐻 = 𝑇 (9: 16, 9: 16) step 4 let 𝑥 = 1 for 𝑖 = 1 𝑡𝑜 8 for 𝑗 = 1 𝑡𝑜 8 𝑏(𝑥) = 1 𝑤ℎ𝑒𝑛 𝐻𝐿(𝑖, 𝑗) ≥ 𝐿𝐻(𝑖, 𝑗) 𝑏(𝑥) = 0 𝑤ℎ𝑒𝑛 𝐿𝐻(𝑖, 𝑗) > 𝐻𝐿(𝑖, 𝑗) 𝑥 = 𝑥 + 1 the loop continues until extracting a binary sequence of length 64 bits. 3. proposed fsd technique the proposed fsd technique consists of two main algorithms called embedding and authentication algorithms. the embedding algorithm is applied at the sender side to detect the face area, generate its authentication information and to hide this information in the region outside the face area. the authentication algorithm is applied at the receiver side in which the embedded information is extracted and compared with the information generated from the received face area to ensure the authenticity of the face in the digital image. when the compared information is not matched the face is classified as unauthentic and the manipulated region is localized in the received image. the following subsections illustrate the details of the proposed embedding and authentication algorithms. a new face swap detection technique for digital images 147 copyright ©2024 assa. adv. in systems science and appl. (2024) 3.1. proposed embedding algorithm the embedding algorithm of the proposed fsd technique can be summarized as illustrated in fig. 3.1. the algorithm starts by reading the original (red, green, and blue) rgb color face image 𝐼 (: , : , 𝑖) where 𝑖 = 1, 2, 3 which refers to the r, g, and b channels of the input image, respectively. in order to detect the face box, the mtcnn is applied to 𝐼 as explained in section 2.1. the output of the mtcnn algorithm is an array contains four readings (𝑦, 𝑥, 𝑤, ℎ) where (𝑥, 𝑦) is the top left corner of the detected face’s box, 𝑤 is the width of the box, and ℎ is the height of the box. to define the detected box in terms of pixels’ positions, the top left corner and the bottom right corner must be defined as integers. the readings (𝑦, 𝑥, 𝑤, ℎ) are rounded to their nearest integer numbers that are greater than or equal to the values. let the resultant array after round is (𝑦 , 𝑥 , 𝑤 , ℎ ), the top left corner of the face box is (𝑥 , 𝑦 ) and the bottom right corner of the face box (𝑥 + ℎ , 𝑦 + 𝑤 ). thus, the pixels’ positions of the face box can be defined as (𝑥 : 𝑥 + ℎ , 𝑦 : 𝑦 + 𝑤 ). a binary mask image 𝐼 is generated according to the obtained positions of pixels that are related to the face box. the following procedure is conducted to generate 𝐼 :  read the size of 𝐼 (𝑀 × 𝑁 × 𝐶) , where 𝑀 = ℎ𝑒𝑖𝑔ℎ𝑡 (𝐼 ) , 𝑁 = 𝑤𝑖𝑑𝑡ℎ (𝐼 ) , and 𝐶 = 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐ℎ𝑎𝑛𝑛𝑒𝑙𝑠 𝑖𝑛 (𝐼 );  generate a binary image of zeros with size (𝑀 × 𝑁);  convert the pixels at the positions (𝑥 : 𝑥 + ℎ , 𝑦 : 𝑦 + 𝑤 ) to ones;  save the resultant binary image as 𝐼 . after generating the mask image, each channel from 𝐼 can be watermarked using the following procedure:  read the channel image 𝐼 (: , : , 𝑖) where 𝑖 refers to the level of the channel.  divide the channel image and 𝐼 into non-overlapping blocks each of size (16×16) pixels.  classify the blocks of the channel image as shown in fig. 3.2 where the mean value of the pixels (𝜇) in each 𝐼 block is calculated to classify the blocks into two groups (i.e., face blocks and non-face blocks).  generate the authentication information from the ‘face blocks’ by calculating the mean value of the pixels for each block followed by rounding the result to the nearest integer value. the resultant values are converted to binary sequences and concatenated to generate one binary sequence which must be embedded in the ‘nonface blocks’.  to increase the robustness of the embedded sequence, apply bch (11,15) encoding. to embed binary sequence in ‘non-face blocks’, the sequence must be divided into non-overlapping subsequences each of length 64 bits. in order to prepare the binary sequence 𝑏𝑐ℎ for the embedding, the length of the sequence must be divisible by 64. the following steps are applied to prepare the binary subsequences for embedding: o 𝑅𝑒𝑚 = 𝑙𝑒𝑛𝑔𝑡ℎ 𝑏𝑐ℎ /64; o if 𝑅𝑒𝑚 = 0 then 𝐵𝑖𝑛 = 𝑏𝑐ℎ ; o else 𝐸𝑥𝑡𝑒𝑛𝑑 = 64 − 𝑅𝑒𝑚, 𝐸𝑥𝑡𝑒𝑛𝑑 = 𝑧𝑒𝑟𝑜𝑠(1: 𝐸𝑥𝑡𝑒𝑛𝑑); o 𝐵𝑖𝑛 = [𝑏𝑐ℎ , 𝐸𝑥𝑡𝑒𝑛𝑑 ].  apply slt watermarking algorithm (explained in section 2.2) to embed the binary subsequence of 𝐵𝑖𝑛 in ‘non-face blocks’. the resultant watermarked ‘non-face blocks’ and the original ‘face-blocks’ are used to construct the watermarked channel image. 148 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) fig. 3.1. the embedding algorithm of the proposed fsd technique. a new face swap detection technique for digital images 149 copyright ©2024 assa. adv. in systems science and appl. (2024) fig. 3.2. classification of channel image blocks into two groups. the procedure of watermarking one channel image is repeated to generate the watermarked channel images which are used to construct the resultant watermarked face image 𝐼 . the watermarked image is sent to the receiver side at which the authentication algorithm must be applied to ensure the safety of the face image and to reveal any manipulation in the face region if exists. the proposed fsd technique doesn’t need the original image or the authentication information at the receiver side which makes the technique completely blind and it can reveal manipulations using the received watermarked image only. 3.2. proposed authentication algorithm the authentication algorithm of the proposed fsd technique can be summarized as illustrated in fig. 3.3. the algorithm starts by reading the watermarked face image 𝐼 and applying the mtcnn algorithm to detect the face box as explained in the embedding procedure. the pixels specification, mask image 𝐼 generation, and blocks classification procedure are the same as those which have been explained at the embedding side. the following procedure is repeated to extract the embedded authentication data and to calculate the authentication data from the ‘face blocks’ in the received 𝐼 :  read channel image 𝐼 (: , ∶, 𝑖) from the received 𝐼 and read the generated 𝐼 .  divide 𝐼 (: , ∶, 𝑖) and 𝐼 into non-overlapping blocks of size (16×16).  classify the blocks into two groups ‘face blocks’ and ‘non-face blocks’ as explained in the embedding procedure.  calculate the new authentication data from the received ‘face blocks’.  extract the embedded authentication data using slt extraction algorithm (explained in section 2.2) from ‘non-face blocks’. as shown in fig. 3.3, the extracted authentication data and the calculated authentication data are compared to check the authenticity of the ‘face blocks’. the block is considered authentic when the compared data are identical. if the compared data are not identical, then the ‘face block’ is considered not authentic and a border is drawn on the block to localize it in the face image. the face image is considered authentic only when all the blocks in the face region are authentic. 150 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) fig. 3.3. authentication algorithm of the proposed fsd technique. 4. results and discussion to test the performance of the proposed fsd technique, the experiments have been conducted for color face images with different sizes which have been collected using google image search from different websites such as [45–48]. samples of the test images are shown in fig. 4.1. a new face swap detection technique for digital images 151 copyright ©2024 assa. adv. in systems science and appl. (2024) fig. 4.1. sample test images. the first experiment has been conducted to ensure the accuracy of the proposed fsd technique in detecting the face’s pixels and generating the mask image. to ensure the safety of the face images after watermarking, the visual quality of the watermarked face images has been tested. the ability of the proposed fsd technique in detecting fake faces has been evaluated. experiments have been conducted to test the capacity and its relationship with the number of ‘non-face blocks’. the experiments also include test of payload and its relationship with the number of ‘face blocks’. the following subsections present the experimental results followed by a general comparison between the proposed fsd technique and the previous deep-learning based face swap detection techniques. 4.1. accuracy of mask image generation the mask image generation process depends on the accuracy of selecting the pixels that are related to the face region as illustrated in section 3. the accuracy of generating the mask image has been tested for different test images before proceeding to other tests. the experimental results proved that the proposed technique can accurately select the positions of face’s pixels and generate the mask image without errors. samples of the mtcnn results and their related mask images are shown in fig. 4.2. 152 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) fig. 4.2. sample results of mask image generation test. 4.2. visual quality test the visual quality of the resultant watermarked images after watermark embedding using the proposed fsd technique must be tested to ensure that there are no visual artifacts in the image, no errors in rearranging the blocks, and no errors in constructing the watermarked images. the first test is conducted by displaying the original face image and its resultant watermarked image side by side to check differences. samples of this test are shown in fig. 4.3. the results proved that the watermarked images are unscathed and there are no visual artifacts in the image. the second test has been conducted to calculate the peak signal-to-noise ratio (psnr) in db and the mean squared error (mse) between the original image and the watermarked image. the experiments have been conducted for the test images shown in fig. 4.1 and the results are shown in table 4.1. the results proved that the visual quality of the watermarked image depends on the ratio of the face’s blocks to the non-face’s blocks and depends also on the contents of the image where some images require less changes to carry the authentication bits while others require more changes in the image contents. the psnr results shown in table 4.1, illustrate the efficiency of the proposed fsd technique in generating watermarked images with high visual quality. fig. 4.3. sample watermarked images. table 4.1. visual quality test results. image name image size size of face area mse psnr (db) image 1 3000×1987×3 668×533×3 0.313 +53.17 image 2 7952×5304×3 1452×1214×3 0.0181 +65.56 image 3 5472×3648×3 1206×944×3 0.0679 +59.81 a new face swap detection technique for digital images 153 copyright ©2024 assa. adv. in systems science and appl. (2024) image 4 3888×2592×3 1206×944×3 0.0496 +61.17 image 5 3456×5184×3 2029×1677×3 0.042 +61.90 image 6 5075×5760×3 2453×1955×3 0.0761 +59.32 image 7 3008×2008×3 409×338×3 0.003 +73.42 image 8 6240×4160×3 1037×885×3 0.0097 +68.28 image 9 2395×2395×3 1096×913×3 0.0763 +59.31 image 10 3027×2007×3 779×625×3 1.129 +47.60 image 11 5599×3733×3 1461×1277×3 0.4056 +52.05 image 12 3442×2295×3 1277×998×3 0.0426 +61.84 image 13 4090×7360×3 1291×985×3 0.1367 +56.77 image 14 6000×4000×3 1203×1024×3 0.0144 +66.56 image 15 2620×2096×3 1383×1076×3 0.4011 +52.10 4.3. fake faces detection test to test the ability of the proposed fsd technique in detecting fake faces in digital images, face swap attack has been imposed on the watermarked face images using ps adobe photoshop version (21.2.1). the results of this experiment proved the efficiency of the proposed fsd technique in revealing the fake face in the digital image and localizing the manipulated part in the face region. samples of the results are shown in fig. 4.4– fig. 4.9. the proposed fsd technique detects fake faces for all test images without error which makes the accuracy 100 % regardless the quality of the test images. fig. 4.4. fake face detection using the proposed fsd technique for ‘image 2’. 154 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) fig. 4.5. fake face detection using the proposed fsd technique for ‘image 4’. fig. 4.6. fake face detection using the proposed fsd technique for ‘image 6’. fig. 4.7. fake face detection using the proposed fsd technique for ‘image 7’. fig. 4.8. fake face detection using the proposed fsd technique for ‘image 13’. a new face swap detection technique for digital images 155 copyright ©2024 assa. adv. in systems science and appl. (2024) fig. 4.9. fake face detection using the proposed fsd technique for ‘image 15’. 4.4. embedding capacity test the embedding capacity of the proposed fsd technique depends on the size of the image and the size of the face region. as mentioned in the embedding procedure, each (16×16) block from the ‘non-face blocks’ can carry 64 bits which is based on the adopted slt watermarking technique (explained in section 2). the number of the ‘non-face blocks’ in one channel is multiplied by 3 to calculate total number of the ‘non-face blocks’ in the image. then the total number of the ‘non-face blocks’ is multiplied by 64 bits to calculate the total embedding capacity of the image. the results of this test are shown in table 4.2 and the relationship between the total embedding capacity and number of ‘non-face blocks’ is illustrated in fig. 4.10. the results proved that the larger the number of ‘non-face blocks’, the higher embedding capacity and vice versa. fig. 4.10. relationship between total capacity and number of ‘non-face blocks’. table 4.2. embedding capacity test results. image name image size size of face area no. of nfb 1 in single channel no. nfb 1 in three channels total capacity (bits) image 1 3000×1987×3 668×533×3 21683 65049 4163136 image 2 7952×5304×3 1452×1214×3 157500 472500 30240000 image 3 5472×3648×3 1206×944×3 73416 220248 14095872 image 4 3888×2592×3 1206×944×3 36720 110160 7050240 image 5 3456×5184×3 2029×1677×3 56416 169248 10831872 image 6 5075×5760×3 2453×1955×3 95178 285534 18274176 0 5000000 10000000 15000000 20000000 25000000 30000000 35000000 t ot al c ap ac it y (b it s) number of non-face blocks in single channel relationship between total capacity and number of non-face blocks 156 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) image 7 3008×2008×3 409×338×3 22928 68784 4402176 image 8 6240×4160×3 1037×885×3 97704 293112 18759168 image 9 2395×2395×3 1096×913×3 18141 54423 3483072 image 10 3027×2007×3 779×625×3 21625 64875 4152000 image 11 5599×3733×3 1461×1277×3 73957 221871 14199744 image 12 3442×2295×3 1277×998×3 25705 77115 4935360 image 13 4090×7360×3 1291×985×3 112278 336834 21557376 image 14 6000×4000×3 1203×1024×3 88810 266430 17051520 image 15 2620×2096×3 1383×1076×3 15437 46311 2963904 1 no. of nfb = number of ‘non-face blocks’. 4.5. payload test the payload in each channel from the input image refers to the total number of bits that are generated in 𝐵𝑖𝑛 as explained in (subsection 3.1). the generated payload depends on the size of the face region. as mentioned in the embedding procedure, the authentication information is generated from the ‘face blocks’, converted to binary, bch coded, and prepared for embedding in ‘non-face blocks’. the results of payload test are shown in table 4.3 and the relationship between the total payload and number of ‘face blocks’ is illustrated in fig. 4.11. the results proved that the larger the number of ‘face blocks’, the higher payload and vice versa. fig. 4.11. relationship between total payload and number of ‘face blocks’. table 4.3. payload test results. image name image size size of face area no. of fb in single channel length of 𝐵𝑖𝑛 in single channel total payload (bits) image 1 3000×1987×3 668×533×3 1505 16448 49344 image 2 7952×5304×3 1452×1214×3 7007 76480 229440 image 3 5472×3648×3 1206×944×3 4560 49792 149376 image 4 3888×2592×3 1206×944×3 2646 28928 86784 image 5 3456×5184×3 2029×1677×3 13568 148032 444096 image 6 5075×5760×3 2453×1955×3 18942 206656 619968 image 7 3008×2008×3 409×338×3 572 6272 18816 image 8 6240×4160×3 1037×885×3 3696 40384 120960 0 100000 200000 300000 400000 500000 600000 700000 t ot al p ay lo ad ( b it s) number of face blocks in single channel relationship between total payload and number of face blocks a new face swap detection technique for digital images 157 copyright ©2024 assa. adv. in systems science and appl. (2024) image 9 2395×2395×3 1096×913×3 4060 44352 133056 image 10 3027×2007×3 779×625×3 2000 21888 65664 image 11 5599×3733×3 1461×1277×3 7360 80320 240960 image 12 3442×2295×3 1277×998×3 5040 55040 165120 image 13 4090×7360×3 1291×985×3 5022 54848 164544 image 14 6000×4000×3 1203×1024×3 4940 53952 161856 image 15 2620×2096×3 1383×1076×3 5916 64576 193728 1 no. of fb = number of ‘face blocks’ 4.6. time complexity the execution time required for the embedding and authentication procedure is very important in practical applications, therefore, the run-time of the proposed fsd algorithms have been calculated for different test images. since the run-time is affected by the hardware and software used in the experimental tests it is useful to mention that the computer used in the experimental tests has 1.80 ghz intel® core tm i7 cpu and 16 gb memory. the software used is matlab (r2020a) and the commands used for this experiment are tic and toc commands. the results of this experiment are shown in table 4.4 which proved the efficiency of the proposed fsd technique in executing embedding and authentication algorithms in seconds. the execution time of the algorithms depends on the size of the face image and the total payload. the higher the payload, the longer the execution time. as shown in the results, the execution time for authentication is lower than that for the embedding which makes the proposed fsd technique suitable for the practical application that require face authentication before proceeding to other processing steps. the test for small size images gives much lower execution time, for instance an image of size (164×307×3) and payload (1024) required (0.31235) seconds for embedding and (0.2701) seconds for authentication. 4.7. comparison with previous techniques as mentioned in the introduction section, the false detection, high time complexity for training and testing, the need for high-quality images for training, and the need for large datasets are some of the limitations in the face swap detection methods that are based on machine learning and deep learning algorithms. since the proposed fsd techniques adopted digital watermarking, there is no need for training and it can be applied for any face image regardless its quality. the accuracy of the proposed technique is 100 % and there are no false detection results thus it outperforms the techniques in [18,22,23] which have recorded accuracies around [92 %, 97 %, 98 % to 99 %] for only specific datasets. the results of the techniques in [18,22,23] can be further degraded when the test images are different from the datasets used in the training process while the results of the proposed fsd technique are always accurate. table 4.4. time complexity test results. image name image size size of face area total payload (bits) embeddi ng time (sec.) extractio n time (sec.) image 1 3000×1987×3 668×533×3 49344 3.3191 1.8827 image 2 7952×5304×3 1452×1214×3 229440 24.6058 14.1453 image 3 5472×3648×3 1206×944×3 149376 12.1752 7.3814 image 4 3888×2592×3 1206×944×3 86784 6.4060 3.5111 image 5 3456×5184×3 2029×1677×3 444096 12.1816 8.7077 image 6 5075×5760×3 2453×1955×3 619968 15.9596 12.1469 image 7 3008×2008×3 409×338×3 18816 2.9516 1.0532 image 8 6240×4160×3 1037×885×3 120960 14.4640 7.1083 158 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) image 9 2395×2395×3 1096×913×3 133056 5.4332 3.2164 image 10 3027×2007×3 779×625×3 65664 3.8354 2.4048 image 11 5599×3733×3 1461×1277×3 240960 12.1705 11.3076 image 12 3442×2295×3 1277×998×3 165120 5.4606 4.0822 image 13 4090×7360×3 1291×985×3 164544 17.5745 10.5852 image 14 6000×4000×3 1203×1024×3 161856 14.5385 7.9509 image 15 2620×2096×3 1383×1076×3 193728 3.6695 2.3317 5. conclusions the rapid development of new technology and the spread of easy-to-use applications can be considered as double-edged sword where these applications can be used for innocent or malicious intentions. recently, several applications have been introduced to swap the faces in digital images which can be adopted for identity theft, threatening, destroying the reputation, spreading fake news, and many others. to authenticate the face image, the research community presented different face swap detection techniques based on machine-learning and deeplearning. there are some limitations in these techniques such as the false detections results, the long time required for training and testing, the need for large datasets for training, the need for high quality images to obtain better detection results, etc. in this paper, a new face swap detection (fsd) technique is presented based on images watermarking technology. the proposed fsd technique have two main algorithms each of them starts by the proposed steps to detect the face region in the image. the embedding algorithm is applied to generate and hide the authentication information while the authentication algorithm is applied to extract and authenticate the face information in the received image. experiments have been conducted to evaluate the performance of the proposed fsd technique for different test images. the results proved that capacity and payload depend on the size of the image and the size of the face region. the larger the size of the face region, the higher the payload. the embedding capacity increases with the increment in the ratio of the number of blocks outside face region to the number of blocks belong to face region. the subjective and objective evaluation of the visual quality proved the efficiency of the proposed fsd technique where high quality watermarked images are generated without any visible distortions. the proposed fsd technique can effectively detect fake face in the image and the accuracy is 100 %. the execution time of the algorithms is low even for large size images which makes the proposed technique suitable for practical applications. for the future work, different watermarking techniques can be applied and compared with the proposed fsd technique in the aim of further improving the performance. funding this research received no external funding. acknowledgements the authors would like to thank their institutions for encouraging and supporting their scientific researches. conflicts of interest the authors declare no conflict of interest. a new face swap detection technique for digital images 159 copyright ©2024 assa. adv. in systems science and appl. (2024) references 1. zhao, y., wu, m., zhang, l., wang, j. & wei, d. (2018). an effective feature segmentation algorithm for a hyper-spectral facial image, information, 9(10), 261. https://doi.org/10.3390/info9100261 2. qin, h. & wang, p. (2019). a template generation and improvement approach for finger-vein recognition. information, 10(4), 145. https://doi.org/10.3390/info10040145 3. ammar, s., bouwmans, t. & neji, m. (2022). face identification using data augmentation based on the combination of dcgans and basic manipulations, information, 13(8), 370, https://doi.org/10.3390/info13080370 4. zhi, j., song, t., yu, k., yuan, f., wang, h., et al. (2022). multi-attention module for dynamic facial emotion recognition, information, 13(5), 207, https://doi.org/10.3390/info13050207 5. semenkov, a., bragin, d., usoltsev, y., konev, a. & kostuchenko, e. (2021). generation of an eds key based on a graphic image of a subject’s face using the rc4 algorithm, information, 12(1), 19, https://doi.org/10.3390/info12010019 6. figueroa, a., peralta, b. & nicolis, o. (2021). coming to grips with age prediction on imbalanced multimodal community question answering data, information, 12(2), 48, https://doi.org/10.3390/info12020048 7. bitouk, d., kumar, n., dhillon, s., belhumeur, p. & nayar, s. k. (2008). face swapping: automatically replacing faces in photographs, acm transactions on graphics, 27(3), 1–8, https://doi.org/10.1145/1360612.1360638 8. korshunova, i., shi, w., dambre, j. & theis, l. (2017). fast face-swap using convolutional neural networks, proc. of 2017 ieee international conference on computer vision (iccv) (venice, italy), 3697–3705, https://doi.org/10.1109/iccv.2017.397 9. tolosana, r., vera-rodriguez, r., fierrez, j., morales, a. & ortega-garcia, j. (2022). an introduction to digital face manipulation. in: rathgeb, c., tolosana, r., vera-rodriguez, r. & busch, c. (eds.), handbook of digital face manipulation and detection (pp. 3–26). berlin, germany: springer international publishing, https://doi.org/10.1007/978-3-03087664-7_1 10. jiang, l., li, r., wu, w., qian, c. & loy, c. c. (2020). deeperforensics-1.0: a largescale dataset for real-world face forgery detection, proc. of the ieee/cvf conference on computer vision and pattern recognition (seattle, wa), 2889–2898, https://doi.org/10.1109/cvpr42600.2020.00296 11. li, l., bao, j., yang, h., chen, d. & wen, f. (2020). advancing high fidelity identity swapping for forgery detection, proc. of the ieee/cvf conference on computer vision and pattern recognition (seattle, wa), 5074–5083, https://doi.org/10.1109/cvpr42600.2020.00512 12. liu, k., perov, i., gao, d., chervoniy, n., zhou, w., et al. (2023). deepfacelab: integrated, flexible and extensible face-swapping framework, pattern recognition, 141, 109628, https://doi.org/https://doi.org/10.1016/j.patcog.2023.109628 13. goodfellow, i. j., pouget-abadie, j., mirza, m., xu, b., warde-farley, d., et al. (2014). generative adversarial nets, proc. of advances in neural information processing systems (montreal, quebec, canada), https://doi.org/10.48550/arxiv.1406.2661 14. ke, z., sun, j., li, k., yan, q. & lau, r. w. h. (2022). modnet: real-time trimapfree portrait matting via objective decomposition, proceedings of the aaai conference on artificial intelligence, 36(1), 1140–1147, https://doi.org/10.1609/aaai.v36i1.19999 160 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) 15. chang, h., lu, j., yu, f. & finkelstein, a. (2018). pairedcyclegan: asymmetric style transfer for applying and removing makeup, proceedings of the ieee conference on computer vision and pattern recognition (salt lake city, ut), 40–48, https://doi.ieeecomputersociety.org/10.1109/cvpr.2018.00012 16. rozsa, a., günther, m., rudd, e. m. & boult, t. e. (2019). facial attributes: accuracy and adversarial robustness, pattern recognition letters, 124, 100–108, https://doi.org/10.1016/j.patrec.2017.10.024 17. hassani, a., malik, h. & diedrich, j. (2022). efficiently mitigating face-swap-attacks: compressed-prnu verification with sub-zones. technologies, information, 10(2), 46, https://doi.org/10.3390/technologies10020046 18. zhang, y., zheng, l., & thing, v. l. l. (2017). automated face swapping and its detection, proc. of 2017 ieee 2nd international conference on signal and image processing (icsip) (singapore), 15–19, https://doi.org/10.1109/siprocess.2017.8124497 19. chang, c.-c. & lin, c.-j. (2011). libsvm: a library for support vector machines, acm transactions on intelligent systems and technology, 2(3), 1–27, https://doi.org/10.1145/1961189.1961199 20. kodovsky, j., fridrich, j. & holub, v. (2011). ensemble classifiers for steganalysis of digital media, ieee transactions on information forensics and security, 7(2), 432–444, https://doi.org/10.1109/tifs.2011.2175919 21. zheng, l., duffner, s., idrissi, k., garcia, c. & baskurt, a. (2016). siamese multi-layer perceptrons for dimensionality reduction and face identification, multimedia tools and applications, 75(9), 5055–5073, https://doi.org/10.1007/s11042-015-2847-3 22. zhang, w. & zhao, c. (2020). exposing face-swap images based on deep learning and ela detection, mdpi proceedings volumes, 46(1), 29, 1–8, https://doi.org/10.3390/ecea5-06684 23. volkova, s. s. & bogdanov, a. s. (2021). a deep learning approach to face swap detection, international journal of open information technologies, 9(10), 16–20. 24. zhang, s., zhu, x., lei, z., shi, h., wang, x., et al. (2017). s3fd: single shot scaleinvariant face detector, proc. of the ieee international conference on computer vision (venice, italy), 192–201. 25. tang, x., du, d. k., he, z., & liu, j. (2018). pyramidbox: a context-assisted single shot face detector. in: ferrari, v., hebert, m., sminchisescu, c., & weiss, y. (eds.), computer vision – eccv 2018 (pp. 812–828). berlin, germany: springer nature, https://doi.org/10.1007/978-3-030-01240-3_49 26. deng, j., dong, w., socher, r., li, l.-j., li, k., et al. (2009). imagenet: a large-scale hierarchical image database, proc. of the 2009 ieee conference on computer vision and pattern recognition, (miami, fl), 248–255, https://doi.org/10.1109/cvpr.2009.5206848 27. sandler, m., howard, a., zhu, m., zhmoginov, a. & chen, l. (2018). mobilenetv2: inverted residuals and linear bottlenecks, proc. of the 2018 ieee/cvf conference on computer vision and pattern recognition (salt lake city, ut), https://doi.org/10.48550/arxiv.1801.04381 28. a rössler, a., cozzolino, d., verdoliva, l., riess, c., thies, j., et al. (2019). faceforensics++: learning to detect manipulated facial images, proc. of the 2019 ieee/cvf international conference on computer vision (iccv) (seoul, korea), 1–11, https://doi.org/10.1109/iccv.2019.00009 a new face swap detection technique for digital images 161 copyright ©2024 assa. adv. in systems science and appl. (2024) 29. afchar, d., nozick, v., yamagishi, j. & echizen, i. (2018). mesonet: a compact facial video forgery detection network, proc. of 2018 ieee international workshop on information forensics and security (wifs) (hong kong, china), 1–7, https://doi.org/10.1109/wifs.2018.8630761 30. thabit, r. (2021). review of medical image authentication techniques and their recent trends, multimedia tools and applications, 80(9), 13439–13473, https://doi.org/10.1007/s11042-020-10421-7 31. khor, h. l., liew, s.-c. & zain, j. m. (2017). region of interest-based tamper detection and lossless recovery watermarking scheme (roi-dr) on ultrasound medical images, journal of digital imaging, 30(3), 328–349, https://doi.org/10.1007/s10278-0169930-9 32. eswaraiah, r. & reddy, e. s. (2014). roi-based fragile medical image watermarking technique for tamper detection and recovery using variance, proc. of the 2014 seventh international conference on contemporary computing (ic3) (noida, india), 553–558, https://doi.org/10.1109/ic3.2014.6897233 33. thabit, r., & khoo, b.e. (2014). a new robust reversible watermarking method in the transform domain. in: mat sakim, h., mustaffa, m. (eds.), the 8th international conference on robotic, vision, signal processing & power applications (pp. 161–168). singapore: springer, https://doi.org/10.1007/978-981-4585-42-2_19 34. thabit, r. & khoo, b. e. (2017). medical image authentication using slt and iwt schemes, multimedia tools and applications, 76(1), 309–332, https://doi.org/10.1007/s11042-015-3055-x 35. thabit, r., & khoo, b. e. (2014). robust reversible watermarking scheme using slantlet transform matrix, journal of systems and software, 88, 74–86, https://doi.org/https://doi.org/10.1016/j.jss.2013.09.033 36. thabit, r. & khoo, b. e. (2015). a new robust lossless data hiding scheme and its application to color medical images, digital signal processing, 38, 77–94, https://doi.org/10.1016/j.dsp.2014.12.005 37. thabit, r. (2019). multi-biometric watermarking scheme based on interactive segmentation process, periodica polytechnica electrical engineering and computer science, 63(4), 263–273, https://doi.org/10.3311/ppee.14219 38. liu, x., lou, j., fang, h., chen, y., ouyang, p., et al. (2019). a novel robust reversible watermarking scheme for protecting authenticity and integrity of medical images, ieee access, 7(1), 76580–76598, https://doi.org/10.1109/access.2019.2921894 39. ustubioglu, a. & ulutas, g. (2017). a new medical image watermarking technique with finer tamper localization, journal of digital imaging, 30(6), 665–680, https://doi.org/10.1007/s10278-017-9960-y 40. viola, p. & jones, m. (2001). rapid object detection using a boosted cascade of simple features, proc. of the 2001 ieee computer society conference on computer vision and pattern recognition (cvpr 2001), (kauai, hi), 1–9, https://doi.org/10.1109/cvpr.2001.990517 41. zhang, k., zhang, z., li, z. & qiao, y. (2016). joint face detection and alignment using multitask cascaded convolutional networks, ieee signal processing letters, 23(10), 1499– 1503, https://doi.org/10.1109/lsp.2016.2603342 162 r. thabit, h. m. fadhil, h. f. fakhruldeen, a. h. shather, m. a. al-askari copyright ©2024 assa adv. in systems science and appl. (2024) 42. rayachoti, e., tirumalasetty, s. & prathipati, s. (2020). slt based watermarking system for secure telemedicine, cluster computing, 23, https://doi.org/10.1007/s10586-02003078-2 43. mohammed, r. t., & khoo, b. e. (2012). image watermarking using slantlet transform, proc. of 2012 ieee symposium on industrial electronics and applications (bandung, indonesia), 281–286, https://doi.org/10.1109/isiea.2012.6496644 44. thabit, r. & khoo, b. e. (2022). robust reversible watermarking application for fingerprint image security, advances in systems science and applications, 22(1), 117–129, https://doi.org/10.25728/assa.2022.22.1.1176 45. creators, t. (2024). pexels. [online]. available https://www.pexels.com/photo/portraitphoto-of-man-3185944/ 46. buzz, p. (2024). here’s how to age multiple faces with faceapp’s old filter. [online]. available https://www.popbuzz.com/internet/viral/age-multiple-faces-faceapp-old-filter/ 47. website, f. (2024). freepik. [online]. available https://www.freepik.com/freephoto/family-having-great-weekend_857329.htm 48. images, i. g. (2024). family pictures, images and stock photos. [online]. available https://www.istockphoto.com/search/2/image?phrase=family adv syst sci appl 2021; 02:83–103 published online at https://ijassa.ipu.ru. stability analysis and optimal measure for controlling eco-epidemiological dynamics of prey-predator model jamiu a. ademosu, samson olaniyi*, sunday o. adewale department of pure and applied mathematics, ladoke akintola university of technology, ogbomoso, nigeria abstract: an eco-epidemiological model representing the interactions between prey and predator populations affected by a disease in an ecosystem is presented. the model is governed by a five-dimensional nonlinear system of ordinary differential equations coupling both ecological and epidemiological features of interacting populations. the well-posedness of the model is established with respect to positivity and boundedness of solutions. conditions for asymptotic stability of different equilibrium points are extensively investigated to determine the existence and coexistence of prey and predator species using local linearization and lyapunov functions techniques. additionally, the analysis of the model is extended to assess the effects of three timedependent control functions, such as disease prevention, treatment and alternative resource for predator, on the population dynamics of the prey-predator coexistence in the system. keywords: ecology, epidemiology, prey-predator model, stability, optimal control problem, simulations 1. introduction interaction between organisms of different species is a common phenomenon in an ecosystem. it is an intrinsic feature which cannot be downplayed in the study of ecological system. this interaction, especially between prey and predator species, often affects or changes the population sizes of organisms in the system [11]. since the advent of the classical lotka and volterra prey-predator models [20, 31], many ecosystem models, which can be defined as mathematical representations of ecological system of prey and predator interactions, have been studied in literature to better understand the real system (see, e.g. [5, 11, 13, 17, 27]). in [11], analysis of modified lotka-volterra model was carried out by taking into consideration an environmental case containing two related populations of prey and predator species having influence on the size of each other. kumar and kumari [17] studied the impact of fear on a chaotic model describing the interactions of the species in the food chain of one prey and two predators. prasad et al. [27] investigated the role of mutual interference between predators on the dynamics of additional food provided predator-prey system. in recent decades, a number of studies have considered ecosystems of interacting populations among which diseases spread. the resulting ecological systems are described by eco-epidemiological models [7–9, 21, 29]. das [9] studied the predator-prey model with parasitic infection transmitting among the predator population only. nandi et al. [21] developed a predator-prey system where prey species only are infected with a viral disease resulting into susceptible and two-stage infected classes. bera et al. [8] presented theoretical analysis supported by simulation of the dynamical behaviors of a prey-predator system where both populations are affected by diseases. ∗corresponding author: solaniyi@lautech.edu.ng 84 j.a. ademosu, s. olaniyi, s.o. adewale interest has been growing recently in the application of optimal control theory to the dynamics of eco-epidemiological models for prey-predator system. al-nassir [6] studied a two-dimensional continuous prey-predator system incorporating an optimal control variable with the aim of reducing the number of the predator density in the population. in [10], the authors extended a prey-predator mathematical model with infection and harvesting on prey to include one time-dependent control for preventing infection in the prey population only. in another development, hugo et al. [16] applied optimal control theory in minimizing the spread of newcastle disease among chickens (prey) and humans (predator) populations using three control variables, such as vaccination of prey, education campaign and treatment of infected humans. in this work, a five-dimensional nonlinear prey-predator model is studied to gain further insights into the dynamics of eco-epidemiological interactions between prey and predator populations. the analysis of the model is extended to investigate the impact of three optimal control variables, namely disease prevention in both prey and predator populations, treatment control for prey, and provision of alternative resources control for predator population. the rest of the work is organized as follows: the model is formulated and its well-posedness is established in section 2. in section 3, existence and stability of possible equilibrium points are investigated. in section 4, the model is extended to explore optimal control dynamics with numerical simulations. section 5 is devoted to concluding remarks. 2. model formulation the dynamic eco-epidemiological prey-predator system is formulated by sub-dividing the prey population at time, t, into susceptible xs(t), infected xi(t) and recovered xr(t). on the other hand, the predator population is sub-divided into two, namely susceptible ys(t) and infected yi(t). it is assumed that the susceptible prey population increases logistically with intrinsic growth rate r and carrying capacity k. following effective contact with the infected prey, the susceptible prey becomes infected at rate β1. to avoid extinction of the prey population due to disease, the infected prey is allowed to recover at per capita rate γ. the natural mortality rate for the prey population is given by µ1. further, it is assumed that the susceptible predator, being healthy, has the capacity to attack and consume both susceptible and infected prey, while the recovered prey is assumed to be under cover from predation. it is also assumed that predators depend on prey for sustenance. the susceptible predator contracts disease from infected predator at rate β2, while the infected predator, being unhealthy, has the capacity to attack and consume infected prey only as assumed in [8]. it has been assumed that only infected prey population recovers due to their accessibility to treatment. hence, recovered class for predators is not considered. this assumption is based on some eco-epidemiological cases, for examples: carnivorous animals (predator) and humans (prey), and hawks (predator) and chicks (prey). the predation and conversion rates for the predator population are given, respectively, by pi and ci, for i = 1, 2, 3. the natural mortality rate for predator population is given by µ2. concise description of the variables and the parameters of the prey-predator system is presented in table 2.1 and the nonlinear ordinary differential equations describing the interacting prey and predator populations is given by copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 85 dxs dt =rxs ( 1− xs k ) − β1xsxi − p1xsys − µ1xs dxi dt =β1xsxi − p2xiys − p3xiyi − (µ1 + γ)xi dxr dt =γxi − µ1xr dys dt =c1p1xsys + c2p2xiys − β2ysyi − µ2ys dyi dt =β2ysyi + c3p3xiyi − µ2yi. (2.1) table 2.1. the variables and parameters of the malaria model (2.1) variable description xs(t) susceptible (healthy) prey xi(t) infected (unhealthy) prey xr(t) recovered prey ys(t) susceptible (healthy) predator yi(t) infected (unhealthy) predator parameter description r intrinsic growth rate k carrying capacity β1 disease transmission rate in prey β2 disease transmission rate in predator pi, i = 1, 2, 3 predation coefficients ci, i = 1, 2, 3 conversion rates γ recovery rate of infected prey µ1 natural mortality rate of prey µ2 natural mortality rate of predator 2.1. well-posedness of the model the mathematical and eco-epidemiological relevance of the prey-predator system (2.1) depends on the well-posedness of the model. keeping in mind that all the parameters of the model are non-negative. here, the boundedness and positivity of solutions of the model are investigated to establish the well-posedness of the model. 2.1.1. boundedness of solutions the following result is required to establish the boundedness of solutions of the model (2.1). lemma 2.1: the susceptible prey population xs(t) is bounded. proof it is clear from the first equation of the model (2.1) that dxs dt ≤ rxs − x2 s k . (2.2) copyright © 2021 assa. adv syst sci appl (2021) 86 j.a. ademosu, s. olaniyi, s.o. adewale solving the nonlinear first order differential inequality of bernoulli type (2.2) yields xs(t) ≤ kxs(0) ke−rt +xs(0)(1− e−rt) . it follows that the lim supxs(t) < k as t→∞. theorem 2.1: the solutions of the system (2.1) are uniformly bounded. proof let the total prey and predator populations be represented, respectively, by x and y , so that the total population of both species be t = x + y . then it follows that dt dt ≤ rxs − µ1x − µ2y. (2.3) using lemma 2.1 and let µ = min{µ1, µ2}, then (2.3) becomes dt dt + µt ≤ rxs ≤ rk, which on integration gives t (t) ≤ rk µ (1− e−µt) + t (0)e−µt. consequently, lim supt (t) < (rk/µ) + ε, ∀ ε > 0 as t→∞. this ends the proof. 2.1.2. positivity of solutions theorem 2.2: the solution set {xs, xi, xr, ys, yi} of the eco-epidemiological model (2.1) with nonnegative initial conditions, xs(0), xi(0), xr(0), ys(0), yi(0) in ω, remain non-negative for all time t > 0. proof sincexs(t) < k as shown in lemma 2.1, then (1−xs/k) ≥ 0 with equality if the size of the susceptible prey is exactly its carrying capacity. hence, the first equation of the model (2.1) implies that dxs(t) dt + (β1xi + p1ys + µ1)xs(t) ≥ 0, (2.4) which on using integrating factor gives d dt [ xs(t) exp (∫ t 0 (β1xi(φ) + p1ys(φ))dφ+ µ1t )] ≥ 0. (2.5) further integration of (2.5) gives xs(t) ≥ xs(0) exp [ − (∫ t 0 (β1xi(φ) + p1ys(φ))dφ+ µ1t )] > 0, ∀t > 0. (2.6) all other variables can be proved to be positive in a similar approach. copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 87 hence, it is sufficient to consider the dynamics of the flow generated by the system (2.1) in a feasible region defined by ω = { (xs, xi, xr, ys, yi) ∈ r5 + : t < rk µ + ε } . in the region ω, the prey-predator model (2.1) can be said to be eco-epidemiologically wellposed. 3. analysis of the model in this section, the prey-predator model is analysed around the possible equilibrium points. 3.1. existence of equilibrium points the prey-predator model (2.1) has the following possible equilibrium points: 3.1.1. trivial equilibrium point (e0) this is a steady state in the absence of prey and predator populations. hence, no interactions exist. the equilibrium point is given by e0 = (0, 0, 0, 0, 0). (3.7) 3.1.2. axial equilibrium point (e1) this is a steady state where only the susceptible prey is present. in other words, it is the disease-free cum predator-free equilibrium point. it exists provided that the intrinsic growth rate r exceeds the natural mortality rate µ1. the equilibrium point is given by e1 = ( k [ 1− µ1 r ] , 0, 0, 0, 0 ) . (3.8) 3.1.3. disease-free equilibrium point (e2) this is a steady state where there is no disease in both prey and predator populations. it exists provided the inequality r ( 1− µ2 kc1p1 ) > µ1 holds. hence, it is obtained as e2 = ( µ2 c1p1 , 0, 0, 1 p1 [ r ( 1− µ2 kc1p1 ) − µ1 ] , 0 ) . (3.9) 3.1.4. predator-free equilibrium point (e3) this is a steady state where there is no predator in the system. it exists provided the inequality r ( 1− ( µ1+γ kβ1 )) > µ1 holds. the equilibrium is given as e3 = ( µ1 + γ β1 , 1 β1 [ r ( 1− ( µ1 + γ kβ1 )) − µ1 ] , γ β1µ1 [ r ( 1− ( µ1 + γ kβ1 )) − µ1 ] , 0, 0 ) . (3.10) 3.1.5. infected prey-free equilibrium (e4) this is a steady state where the infected prey does not exist. the equilibrium will exist if k > k r ( p1µ2 β2 + µ1 ) and c1p1 β2 [ k − k r ( p1µ2 β2 + µ1 )] > µ2. thus, it is given by e4 = ( k − k r ( p1µ2 β2 + µ1 ) , 0, 0, µ2 β2 , c1p1 β2 [ k − k r ( p1µ2 β2 + µ1 )] − µ2 ) . (3.11) copyright © 2021 assa. adv syst sci appl (2021) 88 j.a. ademosu, s. olaniyi, s.o. adewale 3.1.6. infected predator-free equilibrium (e5) this is a steady state where the infected predator is absent in the ecosystem. the equilibrium is given by e5 = (x∗ s , x ∗ i , x ∗ r , y ∗ s , 0), (3.12) such that x∗ s > 0, x∗ i > 0, x∗ r > 0, y ∗ s > 0, and where x∗ s = c2[p2(r − µ1)− p1(µ1 + γ)]− β1µ2 c2p2 ( 1 k + p1β1 p2 ) − c1p1β1 , x∗ i = µ2 c2p2 − c1p1 c2p2 c2[p2(r − µ1)− p1(µ1 + γ)]− β1µ2 c2p2 ( 1 k + p1β1 p2 ) − c1p1β1  , x∗ r = γx∗ i µ1 , y ∗ s = β1x ∗ s − (µ1 + γ) p2 . 3.1.7. interior equilibrium point (e6) this is a steady state where all the populations for prey and predator species are present in the system. the equilibrium is given by e6 = (x∗∗ s , x ∗∗ i , x ∗∗ r , y ∗∗ s , y ∗∗ i ), (3.13) such that x∗∗ s > 0, x∗∗ i > 0, x∗∗ r > 0, y ∗∗ s > 0, y ∗∗ i > 0, and where x∗∗ s = 1 k [ (r − µ1)−x∗∗ i ( β1 − c3p1c3 β2 )] , x∗∗ i = β2[(p2 − p3)µ2 + β2(µ1 + γ)− (r − µ1)(β1β2 − c1p1p3)] β2(p2p3c3 + c2p2p3) + (β1β2 − c1p1p3)(β1β2 − c3p1p3) , x∗∗ r = γx∗∗ i µ1 , y ∗∗ s = µ2 − c3p3x ∗∗ i β2 , y ∗∗ i = β1x ∗∗ s − p2y ∗∗ s − (µ1 + γ) p3 . 3.2. stability analysis the jacobian matrix of the prey-predator system (2.1) is given by j =  j1 −β1xs 0 −p1xs 0 β1xi j2 0 −p2xi −p3xi 0 γ −µ1 0 0 c1p1ys c2p2ys 0 j3 −β2ys 0 0 0 β2yi β2ys + c3p3xi − µ2  , (3.14) copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 89 where j1 = r − 2rxs k − β1xi − p1ys − µ1, j2 = β1xs − p2ys − p3yi − (µ1 + γ), j3 = c1p1xs + c2p2xi − β2yi − µ2. the stability of each of the equilibrium points is analysed by finding the eigenvalues of the jacobian matrix (3.14) evaluated at each point. 3.2.1. stability of e0 the jacobian matrix (3.14) is evaluated at the trivial equilibrium point (3.7). hence, solving the corresponding characteristic equation |j(e0)− λi5)| = 0, where λ is the eigenvalue and i5 is the identity matrix of order five, the following eigenvalues are obtained: λ1 = r − µ1, λ2 = −µ1, λ3 = −µ2, λ4 = −(µ1 + γ), λ5 = −µ2. it can be seen that all the eigenvalues are unconditionally negative except λ1. thus, e0 will be stable if r < µ1 and unstable if r > µ1. this result is theorized as follows. theorem 3.1: the trivial equilibrium point e0 of the prey-predator model (2.1), given by (3.7), is stable if r < µ1 and unstable if r > µ1. the implication of theorem 3.1 is that the coexistence of prey-predator population will not occur if the intrinsic growth rate r does not exceed the natural mortality rate µ1. in other words, stable equilibrium point e0 will result into extinction of prey and predator species in the ecosystem. 3.2.2. stability of e1 the characteristic equation of the matrix (3.14) evaluated at the axial equilibrium point e1 is given by |j(e1)− λi5)| = 0, which gives the following eigenvalues: λ1 = µ1 − r, λ2 = −µ1, λ3 = β1k ( 1− µ1 r ) − (µ1 + γ), λ4 = c1p1k ( 1− µ1 r ) − µ2, λ5 = −µ2. the stability result is theorized as follows. theorem 3.2: the axial equilibrium point e1 of the prey-predator model (2.1), given by (3.8), is stable if the inequalities: µ1 < r, β1k ( 1− µ1 r ) < (µ1 + γ) and c1p1k ( 1− µ1 r ) < µ2 hold. the implication of theorem 3.2 is that the stable e1 will support the existence of prey population only in the ecosystem provided these threshold conditions are satisfied: r01 < 1,r02 < 1 andr03 < 1, where r01 = µ1 r , r02 = β1k ( 1− µ1 r ) µ1 + γ , and r03 = c1p1k ( 1− µ1 r ) µ2 . now, suppose r0 = max{r01,r02,r03}, and considering the prey sub-model only such that the axial equilibrium point e1 becomes ẽ1 = (x0 s , 0, 0), where x0 s = k [ 1− µ1 r ] . the following global stability result is proved. theorem 3.3: the equilibrium point ẽ1 is globally asymptotically stable ifr0 < 1. proof since p1 = 0 in the absence of predation, thenr0 = max{r01,r02}. consider the lyapunov copyright © 2021 assa. adv syst sci appl (2021) 90 j.a. ademosu, s. olaniyi, s.o. adewale function (combination of goh-volterra [2, 3, 13, 15, 22] and linear types) given by l = xs −x0 s −x0 s ln xs x0 s +xi. (3.15) the time derivative of lyapunov function (3.15) along the trajectory of the prey only submodel is given by l̇ = ( 1− x0 s xs ) dxs dt + dxi dt . (3.16) using dxs dt = rxs ( 1− xs k ) − β1xsxi − µ1xs and dxi dt = β1xsxi − (µ1 + γ)xi in (3.16) with xs ≤ x0 s , it follows that l̇ = (xs −x0 s ) ( r [ 1− xs k ] − β1xi − µ1 ) + β1xsxi − (µ1 + γ)xi ≤ r [ 1− x0 s k ] − β1xi − µ1 + β1x 0 sxi − (µ1 + γ)xi. (3.17) since r [ 1− x0 s k ] = µ1 at the disease-free steady state of the prey only sub-model, then it follows from (3.17), by further simplifications, that l̇ ≤ −β1xi − (µ1 + γ) ( 1− β1x 0 s µ1 + γ ) xi. (3.18) sincer02 = β1x 0 s µ1 + γ , consequently (3.18) becomes l̇ ≤ −β1xi − (µ1 + γ) (1−r02)xi. therefore, l̇ ≤ 0 when r02 < 1 with equality, l̇ = 0, if and only if xi = 0. it follows from lasalle’s invariance principle [19] that the largest invariant set contained in {(xs, xi, xr) ∈ r3 + : l̇ = 0} is the singleton set {ẽ1}. hence, the equilibrium point ẽ1 = (k[1− (µ1/r)], 0, 0) is globally asymptotically stable. the implication of theorem 3.3 is that the presence of disease in the prey population can be eliminated if r02 < 1, irrespective of the initial sizes of the infected prey in the ecosystem. thus, every solution of the prey only sub-model will always converge to the point ẽ1 wheneverr02 is less than unity. this theoretical result is demonstrated in figure 3.1. 3.2.3. stability of e2 considering that there are only two disease classes in the prey-predator system (2.1), then using the notations in the next generation operator method [30], it follows that f = [ β1µ2 c1p1 0 0 β2 p1 m1 ] and v = [ p2 p1 m1 + µ1 + γ 0 0 µ2 ] , where m1 = r ( 1− µ2 kc1p1 ) − µ1. hence, the basic reproduction number of the full system (2.1), denoted by r∗ 0, is given by r∗ 0 = ρ(fv −1) = max  β1µ2 c1p1 ( p2 p1 m1 + µ1 + γ ) , β2m1 p1µ2  . (3.19) copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 91 0 5 10 15 0 10 20 30 40 50 60 70 time in fe ct ed p re y fig. 3.1. convergence of solution trajectories of prey sub-model to the equilibrium point ẽ1 at different initial sizes. the parameter values used are: r = 11.2, β1 = 0.04, µ1 = 0.85, γ = 0.75, k = 30, so thatr01 = 0.0759 andr02 = 0.6931, implying thatr0 = r02 < 1. therefore, by theorem 2 of [30], the following stability result is claimed. theorem 3.4: the disease-free equilibrium point e2 of the prey-predator model (2.1), given by (3.9), is stable provided that r∗ 0 < 1. the implication of theorem 3.4 is that the presence of few species around the equilibrium point e2 will support the coexistence of both prey and predator species in a disease-free ecosystem whenever the basic reproduction number, r∗ 0, is less than unity. the global dynamics of the prey-predator around the disease-free equilibrium point e2 is explored next. theorem 3.5: the disease-free equilibrium point of the prey-predator model (2.1), given by (3.9), is globally asymptotically stable if r∗ 0 > 1. proof the proof is based on using the lyapunov function defined by f = p1 p2m1 + p1(µ1 + γ) xi + 1 µ2 yi, (3.20) which when differentiated with respect to time, t, gives df dt = p1 p2m1 + p1(µ1 + γ) [β1xsxi − p2xiys − p3xiyi − (µ1 + γ)xi] + 1 µ2 (β2ysyi + c3p3xiyi − µ2yi). copyright © 2021 assa. adv syst sci appl (2021) 92 j.a. ademosu, s. olaniyi, s.o. adewale since xs ≤ β1µ2 c1p1 and ys ≤ 1 p1 [ r ( 1− µ2 kc1p1 ) − µ1 ] , noting that m1 = r ( 1− µ2 kc1p1 ) − µ1. therefore, it follows, after simplification, that df dt ≤  β1µ2 c1p1 ( p2 p1 m1 + µ1 + γ ) − 1 xi − p1p3 p2m1 + p1(µ1 + γ) xiyi + [ β2m1 p1µ2 − 1 ] yi + 1 µ2 (c3p3xiyi) ≤ max  β1µ2 c1p1 ( p2 p1 m1 + µ1 + γ ) , β2m1 p1µ2 − 1  (xi + yi) − [ p1 p2m1 + p1(µ1 + γ) − c3 µ2 ] p3xiyi = [r∗ 0 − 1](xi + yi)− [ p1 p2m1 + p1(µ1 + γ) − c3 µ2 ] p3xiyi. consequently, df/dt ≤ 0 if r∗ 0 < 1 and xi = 0 or yi = 0. thus, xs → µ2/(c1p1), xr → 0 and ys → m1/p1 as both xi → 0 and yi → 0. therefore, by lasalle’s invariance principle [19], the disease-free equilibrium point e2 is globally asymptotically stable. the implication of theorem 3.5 is that elimination of disease in the ecosystem is independent of initial sizes of infected prey and infected predator whenever r∗ 0 < 1. this result is demonstrated in figure 3.2(a) for infected prey and figure 3.2(b) for infected predator population, where every solution trajectory converges to the disease-free equilibrium. 0 5 10 15 20 25 30 35 40 0 1 2 3 4 5 6 7 8 9 10 time in fe c te d p re y (a) 0 5 10 15 0 5 10 15 20 25 time in fe c te d p re d a to r (b) fig. 3.2. convergence of solution trajectories of prey-predator model to the disease-free equilibrium point e2 at different initial sizes. the parameter values used are: mu1 = 0.8, r = 3, β1 = 0.4, p1 = 1.0, γ = 0.35, p2 = 0.2, p3 = 1.0, µ2 = 0.85, k = 30, β2 = 0.3, c1 = 0.25, c2 = 0.25, c3 = 0.15, so that r∗ 0 = 0.6565 < 1. copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 93 3.2.4. stability of e3 evaluating the jacobian matrix (2.1) at predator-free equilibriume3 leads to the characteristic equation |j(e3)− λi5| = 0. thus, the following eigenvalues are obtained: λ1 = −µ1 λ2 = −µ2, λ3 = c1p1 β1 (µ1 + γ)− c2p2 β1 (m2 + µ2), λ4 = −r 2 ( µ1 + γ kβ1 ) + 1 2 √ r2 ( µ1 + γ kβ1 )2 + 4m2(µ1 + γ), λ5 = −r 2 ( µ1 + γ kβ1 ) − 1 2 √ r2 ( µ1 + γ kβ1 )2 + 4m2(µ1 + γ), where m2 = r ( 1− µ1+γ kβ1 ) − µ1. the stability of e3 is theorized as follows theorem 3.6: the predator-free equilibrium point e3 of the prey-predator model (2.1), given by (3.10), is stable if the following conditions hold c1p1(µ1 + γ) < c2p2(m2 + µ2) and √ r2 ( µ1 + γ kβ1 )2 + 4m2(µ1 + γ) < r ( µ1 + γ kβ1 ) . this result implies that stable e3 will not be able to support the coexistence of both prey and predator populations. the presence of few species around the predator-free equilibrium point e3 will make predator population to go extinct in the ecosystem. 3.2.5. stability of e4 the jacobian matrix (3.14) is evaluated at the infected prey-free equilibrium point (3.11). hence, solving the corresponding characteristic equation |j(e4)− λi5)| = 0 gives one eigenvalue, λ1 = −µ1, and the remaining four eigenvalues are obtainable from λ4 +w1λ 3 +w2λ 2 +w3λ+w4 = 0, (3.21) where w1 = r − µ1 − p1 µ2 β2 , w2 = d1d2 − 2µ2(d1 + d2) + µ2 + c1p1d3 − c1p1 µ2 β2 with d1 = r − 2r d3 k − p1 µ2 β2 − µ1, d2 = β1d3 − p2 µ2 β2 − c1p1p3 d3 β2 − (µ1 + µ2 + γ) and d3 = k − k r ( p1µ2 β2 + µ2 ) . w3 = c1p1 [ (d2 − µ2) µ2 β2 − d3(d2µ2 + d1) ] − d2µ2(2d1 + µ2) and w4 = d2µ 2 2 ( d1 + p1 β2 ) . copyright © 2021 assa. adv syst sci appl (2021) 94 j.a. ademosu, s. olaniyi, s.o. adewale applying routh-hurwitz’s criterion, the roots of (3.21) will have negative real parts if w1 > 0, w1w2 > w3, w1w2w3 > w 2 1w4 +w 2 3 . the stability result is hereby theorized as follows. theorem 3.7: the infected prey-free equilibrium of the prey-predator model (2.1), given by (3.11), is stable if w1 > 0, w1w2 > w3 and w1w2w3 > w 2 1w4 +w 2 3 . the implication of theorem 3.7 is that the presence of few species around the stable infected prey-free equilibrium e4 will support the coexistence of both prey and predator populations in the ecosystem. 3.2.6. stability of e5 two of the eigenvalues of the jacobian matrix (3.14) evaluated at the infected predator-free equilibrium point e5 are given by λ1 = −µ1 and λ2 = −µ2, while the remaining can be obtained from the characteristic equation λ3 + e1λ 2 + e2λ+ e3 = 0, (3.22) where e1 = −(f1 + f2 + f3), with f1 = 2c1p1f̃1 − µ2(1 + c2p2) such that f̃1 = c2p2(r − µ1)− p1c2(µ1 + γ)− βµ2 c2p2 ( 1 k + p1β1 p2 ) − c1p1β1 , f2 = r − µ2 − [ 2r k + c2p2 + kp1β1(c2 − c1) k(c2p2 + kp1c2 − c1p1β1) ] f̃1 − β1 ( µ2 − c1p1 c2p2 f̃1 ) (1− p1) f3 = β1f̃1 − p2 [( c2p2 + kp1β1(c2 − c1) k(c2p2 + kp1c2 − c1p1β1) ) f̃1 − β1 ( µ2 − c1p1 c2p2 f̃1 )] − (µ1 + γ). e2 = f1f2 + f3(f1 + f2) + β2 1 f̃1 ( µ2 − c1p1 c2p2 f̃1 ) + p 2 1 c1 p2 f̃1(f3 + (µ1 + γ)− β1f̃1), e3 = 1 p2 (f3 + (µ1 + γ)− β1f̃1) [( µ2 − c1p1 c2p2 f̃1 ) (p2(c2[p1f̃1 − p2]− c1p1))− c1p 2 1 f̃1(f3 + f1f2) ] . by routh-hurwitz’s criterion, the roots of the characteristic equation (3.22) will have negative real parts if e1 > 0 and e1e2 > e3. this result is stated as follows. theorem 3.8: the infected predator-free equilibrium point of the prey-predator system (2.1), given by (3.12), is stable if e1 > 0 and e1e2 > e3. the result in theorem 3.8 implies that coexistence of prey and predator species in the ecosystem is possible in the absence of infected predator, provided that the given parametric conditions are valid. 3.2.7. stability of e6 recall from (3.19) that the basic reproduction number of the full system (2.1) is given by r∗ 0 = max  β1µ2 c1p1 ( p2 p1 m1 + µ1 + γ ) , β2m1 p1µ2  . the following global stability result is explored. copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 95 theorem 3.9: the interior equilibrium point e6 of the prey-predator model (2.1) is globally asymptotically stable if r∗ 0 > 1. proof let r∗ 0 > 1, such that the interior equilibrium point e6 exists. noting from (3.13) that e6 = (x∗∗ s , x ∗∗ i , x ∗∗ r , y ∗∗ s , y ∗∗ i ), then the following lyapunov function of quadratic type [28] is defined: v = 1 2 [(xs −x∗∗ s ) + (xi −x∗∗ i ) + (xr −x∗∗ r ) + (ys − y ∗∗ s ) + (yi − y ∗∗ i )]2 (3.23) the time derivative of v in (3.23) along the solution path of the full system (2.1) is given by v̇ = [(xs −x∗∗ s ) + (xi −x∗∗ i ) + (xr −x∗∗ r ) + (ys − y ∗∗ s ) + (yi − y ∗∗ i )] dt dt , (3.24) where t (t) = xs +xi +xr + ys + yi. recall from (2.3) of the boundedness result (theorem 2.1) that dt dt ≤ rk − µt, and since x∗∗ s +x∗∗ i +x∗∗ r + y ∗∗ s + y ∗∗ i = rk µ . it follows from (3.24) that v̇ ≤ [ t (t)− rk µ ] [rk − µt (t)] = µ [ t (t)− rk µ ] [ rk µ − t (t) ] = −µ [ t (t)− rk µ ]2 = −µ[(xs −x∗∗ s ) + (xi −x∗∗ i ) + (xr −x∗∗ r ) + (ys − y ∗∗ s ) + (yi − y ∗∗ i )]2. it can be seen that v̇ ≤ 0 with equality if and only if xs = x∗∗ s , xi = x∗∗ i , xr = x∗∗ r , ys = y ∗∗ s and yi = y ∗∗ i . hence, by lasalle’s invariance principle [19], the largest invariant set in {(xs, xi, xr, ys, yi) ∈ r5 + : v̇ = 0} is the only set {e6}. thus, the interior equilibrium point e6 = (x∗∗ s , x ∗∗ i , x ∗∗ r , y ∗∗ s , y ∗∗ i ) is globally asymptotically stable. this completes the proof. the implication of theorem 3.9 is that coexistence of prey and predator species in the ecosystem is possible while disease persists in the population. the presence of disease in the ecosystem, such that r∗ 0 > 1, is independent of initial population sizes of both infected prey and predator species. this global stability result is depicted in figure 3.3(a) for infected prey population and figure 3.3(b) for infected predator population. the phase portrait which illustrates the stable coexistence of prey and predator populations is presented in figure 3.4. 4. optimal control model in order to enhance the coexistence of both prey and predator species in the ecosystem, the prey-predator model (2.1) is extended by incorporating the following three time-dependent optimal control functions: copyright © 2021 assa. adv syst sci appl (2021) 96 j.a. ademosu, s. olaniyi, s.o. adewale 0 5 10 15 20 25 30 35 40 0 2 4 6 8 10 12 14 16 18 20 time in fe c te d p re y (a) 0 5 10 15 0 5 10 15 20 25 time in fe c te d p re d a to r (b) fig. 3.3. convergence of solution trajectories of prey-predator model to the interior equilibrium point e6 at different initial sizes. the parameter values used are:mu1 = 0.5, r = 11.2, β1 = 0.4, p1 = 1.0, γ = 0.35, p2 = 0.2, p3 = 1.0, µ2 = 0.65, k = 30, β2 = 0.65, c1 = 0.25, c2 = 0.25, c3 = 0.15, so that r∗ 0 = 9.7293 > 1. 0 1 2 3 4 5 6 7 8 0 1 2 3 4 5 6 7 prey p re da to r fig. 3.4. phase portrait of predator against prey showing stable coexistence of both populations around the interior (positive) equilibrium point. 1. the control function u1(t) ∈ [0, 1], which represents disease prevention measure for both prey and predator populations. 2. the function u2(t) ∈ [0, 1], which represents treatment control for prey population. 3. the control function u3(t) ∈ [0, 1] representing alternative source of food for predator survival. therefore, the prey-predator model (2.1) becomes a non-autonomous system of ordinary differential equations of the form: copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 97 dxs dt = rxs ( 1− xs k ) − (1− u1(t))β1xsxi − p1xsys − µ1xs dxi dt = β1xsxi − p2xiys − p3xiyi − (µ1 + γ)xi − u2(t)r0xi dxr dt = γxi + u2(t)r0xi − µ1xr dys dt = c1p1xsys + c2p2xiys − (1− u1(t))β2ysyi − µ2ys + u3(t)c0ys dyi dt = (1− u1(t))β2ysyi + c3p3xiyi − µ2yi, (4.25) where r0 and c0 are the rate constants for the optimal controls u2(t) and u3(t), respectively. since the goal is to reduce the risk of both prey and predator species going extinct, such that the populations of infected prey and predator are minimized at minimum costs possible, then the following objective functional is formed j = ∫ tf 0 ( a1xi + a2yi + 1 2 [b1u 2 1 +b2u 2 2 +b3u 2 3] ) dt, (4.26) where a1, a2, b1, b2 and b3 are the positive weight constants for balancing the terms in the objective functional over the time interval [0, tf ]. further, the cost function associated with the disease prevention measure for prey and predator populations is represented by the term 1/2b1u 2 1, while the treatment measure for prey population and the provision of alternative food source for predator are, respectively, represented by the terms 1/2b2u 2 2 and 1/2b3u 2 3. the objective functional is made of nonlinear cost on controls of quadratic type in line with the standards for optimal control problems (see, e.g. [1, 4, 14, 16, 24, 25]). therefore, the optimal control problem for the prey-predator system (4.25) is to seek a control triple u∗ = (u∗1, u ∗ 2, u ∗ 3, ), such that j(u∗) = min {j(u1, u2, u3) : u1, u2, u3 ∈ u} , (4.27) where u = {ui(t) : 0 ≤ ui(t) ≤ 1,lebesgue measurable, t ∈ [0, tf ], for i = 1, 2, 3} is a nonempty set of all admissible controls. 4.1. analysis of the control model the necessary conditions that must be satisfied by the optimal control triple u∗ for the minimization problem are derivable from pontryagin’s maximum principle [26]. this principle converts the problem in (4.27) with (4.26) subject to the state system (4.25) into a problem of minimizing pointwise a hamiltonian h, with respect to u1(t), u2(t) and u3(t). thus, the hamiltonian for the control problem is given by copyright © 2021 assa. adv syst sci appl (2021) 98 j.a. ademosu, s. olaniyi, s.o. adewale h = a1xi + a2yi + 1 2 [b1u 2 1 +b2u 2 2 +b3u 2 3] +η1 [ rxs ( 1− xs k ) − (1− u2(t))β1xsxi − p1xsys − µ1xs ] +η2 [β1xsxi − p2xiys − p3xiyi − (µ1 + γ)xi − u2(t)r0xi] +η3[γxi + u2(t)r0xi − µ1xr] +η4[c1p1xsys + c2p2xiys − (1− u1(t))β2ysyi − µ2ys + u3(t)c0ys] +η5[(1− u1(t))β2ysyi + c3p3xiyi − µ2yi], (4.28) where η1, η2, η3, η4 and η5 are the adjoint or co-state variables corresponding to the state variables xs, xi, xr, ys and yi, respectively. applying the existence results as given in [2, 12], the following result is obtained. theorem 4.1: for the optimal control triple u∗ that minimizes the objective functional (4.26) over u subject to (4.25), there exist adjoint variables η1, η2, η3, η4 and η5 satisfying the system governing the adjoint variables dη1 dt = 2rxs k η1 + (1− u1)β1xi(η1 − η2) + p1ys(η1 − c1η4) + (µ1 − r)η1 dη2 dt = (1− u1)β1xs(η1 − η2) + p2ys(η2 − c2η4) + p3yi(η2 − c3η5) +(γ + u2r0)(η2 − η3) + µ1η2 − a1 dη3 dt = µ1η3 dη4 dt = p1xs(η1 − c1η4) + p2xi(η2 − c2η4) + (1− u1)β2yi(η4 − η5) +(µ2 − u3c0)η4 dη5 dt = p3xi(η2 − c3η5) + (1− µ1)β2ys(η4 − η5) + µ2η5 − a2 (4.29) and with the transversality conditions η1(tf ) = η2(tf ) = η3(tf ) = η4(tf ) = η5(tf ) = 0. (4.30) moreover, the optimal control triple (u∗1, u ∗ 2, u ∗ 3) is characterized as follows u∗1 = max { 0,min { 1, 1 b1 [β1xsxi(η2 − η1) + β2ysyi(η5 − η4)] }} , u∗2 = max { 0,min { 1, 1 b2 [r0xi(η2 − η3)] }} , u∗3 = max { 0,min { 1, −c0ysη4 b3 }} . (4.31) proof the system (4.29) governing the adjoint variables is obtained by finding the partial derivative of the hamiltonian h (4.28) as follows: dη1 dt = − ∂h ∂xs , dη2 dt = − ∂h ∂xi , dη3 dt = − ∂h ∂xr , dη4 dt = − ∂h ∂ys , dη5 dt = − ∂h ∂yi copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 99 with the transversality conditions: η1(tf ) = η2(tf ) = η3(tf ) = η4(tf ) = η5(tf ) = 0. finally, differentiating the hamiltonian h partially with respect to each of the three controls u1, u2 and u3, so that ∂h ∂u1 ≡ b1u1 + β1xsxi(η1 − η2) + β2ysyi(η4 − η5) = 0, ∂h ∂u2 ≡ b2u2 + r0xi(η3 − η2) = 0, ∂h ∂u3 ≡ b3u3 + c0ysη4 = 0, which when solved on the interior of the control set u gives the required characterization (4.31). hence, the proof. 4.2. simulations of the control model here, the numerical simulations of the coupled system of state equations (4.25) with adjoint equations (4.29) are conducted. the standard iterative fourth order forward-backward rungekutta method is used to solve the optimality system due to difference in time orientations of the state and adjoint equations [18, 23]. the optimal control problem is simulated over the time interval [0, 15] in years using these parameter values: µ1 = 0.5, r = 11.2, β1 = 0.4, p1 = 1.0, γ = 0.35, p2 = 0.2, p3 = 1.0, µ2 = 0.65, k = 30, β2 = 0.65, c1 = 0.25, c2 = 0.25, c3 = 0.15. in addition, the balancing weight constants are given as a1 = 1, a2 = 0.1, b1 = 0.01, b2 = 0.0001 and b3 = 0.008. the rate constants for the controls are chosen as r0 = 0.4 and c0 = 0.15. thus, to minimize the objective functional (4.26) with disease prevention control only at low cost of implementation, the control u1 should be maintained at maximum (100%) effort for 8 years before being dropped to zero in final time as shown in the figure 4.5(a). the control profile shown in figure 4.5(b) indicates that single implementation of treatment control u2 requires maximum (100%) effort throughout the period of the control intervention. it is observed in figure 4.5(c) that alternative food source control u3 should be at the upper bound for 14 years before coming to the lower bound. the combined implementation of the three controls for minimizing the objective functional is depicted in figure 4.5(d), where it is can be observed that prevention control u1 should be maintained at the upper bound for 8 years before declining to the lower bound, treatment control u2 and provision of alternative food source u3 should be maintained at maximum (100%) for about 2 years and almost 1 year, respectively. the effects of combining all the three optimal controls on the dynamical behaviors of prey and predator populations are illustrated in figure 4.6. the size of susceptible prey population with control reduces due to predation when compared with the size without control, but the population does not go extinct due to the presence of the combined optimal control as shown in figure 4.6(a). this suggests the possibility of stable coexistence of predators with low density of prey population as indicated in [27]. the infected prey population with optimal control (figure 4.6(b)) decreases sharply when compared with the case without control. further, in figure 4.6(c), it can be seen that the size of susceptible predator population with combined optimal control is higher than the size without control. whereas, the presence of the combined optimal control reduces the population size of the infected predator in the ecosystem. copyright © 2021 assa. adv syst sci appl (2021) 100 j.a. ademosu, s. olaniyi, s.o. adewale 0 5 10 15 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 o p tim a l c o n tr o l u 1 time (a) 0 5 10 15 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 o p tim a l c o n tr o l u 2 time (b) 0 5 10 15 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 o p tim a l c o n tr o l u 3 time (c) 0 5 10 15 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 time o p ti m a l c o n tr o l p ro fi le s u 1 u 2 u 3 (d) fig. 4.5. control profiles for single and combined implementations of u1, u2 and u3. 5. conclusion a nonlinear mathematical model of prey-predator interacting populations in an ecological system with disease spread is studied. the five-dimensional eco-epidemiological model is analyzed with a view to providing insights into the behavior of ecosystem in the presence of disease. seven possible equilibrium points, namely trivial, axial, disease-free, predatorfree, infected prey-free, infected predator-free and interior equilibrium points are analytically determined. asymptotic stability of the eco-epidemiological system is analyzed to investigate the behavior of the system around each of the obtained equilibrium points. some conditions that guarantee the existence of each prey and predator species, as well as coexistence of both prey and predator populations are given using local linearization and lyapunov functions techniques. in addition, the eco-epidemiological model is extended to include three time-dependent optimal controls, such as disease prevention in both species, treatment strategy for prey and provision of alternative food source for predator survival. hence, the optimal control preypredator model is analyzed using pontryagin’s maximum principle in order to enhance the coexistence of both prey and predator species by reducing the risks of their extinction in the ecosystem, such that the populations of infected prey and predator species are minimized at minimum costs possible. copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 101 0 5 10 15 0 5 10 15 20 25 30 time s u s c e p ti b le p re y u 1 ≠ 0, u 2 ≠ 0, u 3 ≠ 0 u 1 =u 2 =u 3 =0 (a) 0 5 10 15 0 0.5 1 1.5 2 2.5 time in fe c te d p re y u 1 ≠ 0, u 2 ≠ 0, u 3 ≠ 0 u 1 =u 2 =u 3 =0 (b) 0 5 10 15 0 2 4 6 8 10 12 14 time s u s c e p ti b le p re d a to r u 1 ≠ 0, u 2 ≠ 0, u 3 ≠0 u 1 =u 2 =u 3 =0 (c) 0 5 10 15 0 2 4 6 8 10 12 time in fe c te d p re d a to r u 1 ≠ 0, u 2 ≠ 0, u 3 ≠0 u 1 =u 2 =u 3 =0 (d) fig. 4.6. combined effects of optimal controls u1, u2 and u3 on the dynamics of the prey-predator model. the analysis of the eco-epidemiological model presented in this study can be applied to a wide spectrum of interactions between different kinds of living beings in the environment. acknowledgement the authors are thankful for the constructive comments and suggestions of the editor and the anonymous reviewers. initial discussion on the original manuscript with professor mini ghosh of vellore institute of technology, chennai, india is also acknowledged with thanks. references 1. abidemi, a. and aziz, n.a.b. (2020). optimal control strategies for dengue fever spread in johor, malaysia, comput. meth. prog. bio., 196, 105585. https://doi.org/10.1016/j.cmpb.2020.105585 2. abimbade, s.f., olaniyi, s., ajala, o.a. and ibrahim, m.o. (2020). optimal control analysis of a tuberculosis model with exogenous re-infection and incomplete treatment, optim. control appl. meth., (41), 2349–2368. https://doi.org/10.1002/oca.2658 copyright © 2021 assa. adv syst sci appl (2021) 102 j.a. ademosu, s. olaniyi, s.o. adewale 3. adewale, s.o., podder, c.n. and gumel, a.b. (2009). mathematical analysis of a tb transmission model with dots, can. appl. math. q., 17, 1–36. 4. akanni, j.o., akinpelu, f.o., olaniyi, s., oladipo, a.t., and ogunsola, a.w. (2020). modelling financial crime population dynamics: optimal control and cost-effectiveness analysis, int. j. dynam. control, 8, 531–544. 5. al-darabsah, i., tang, x. and yaun, y. (2016). a prey-predator model with migrations and delays, discrete contin dyn syst ser b., 21(3), 737–761. 6. al-nassir, s. (2017). the dynamics and optimal control of a prey-predator system, global j. pure appl. math., 13, 5287–5298. 7. bakare, e.a., adekunle, y.a. and nwagwo, a. (2012). mathematical analysis of the control of the spread of infectious disease in a prey-predator ecosystem, int. j. comp. organ. trends, 2(1), 27–32. 8. bera, s.p., maiti, a. and samanta, g.p. (2015). a prey-predator model with infection in both prey and predator, filomat, 29(8), 1753–1767. 9. das, k.p. (2011). a mathematical study of a predator-prey dynamics with disease in predator, isrn appl. math., article id 807486, 1–16. 10. diva, a.r., fatmawati, windarto and didik, k.a. (2018). optimal control of predatorprey mathematical model with infection and harvesting on prey, journal of physics: conf. series, 974, 012050. 11. doust, m.h.r. and gholizade, s. (2014). an analysis of the modified lotka-volterra predator-prey model, gen. math. notes, 25(2), 1–5. 12. fleming, w.h. and richel, r.w. (1975). deterministic and stochastic optimal control. new york: springer. 13. ghosh, m. (2010). modeling prey-predator type fishery with reserve area, int. j. biomath., 3, 351–365. 14. ghosh, m., olaniyi, s. and obabiyi, o.s. (2020). mathematical analysis of reinfection and relapse in malaria dynamics, appl. math. comput., 373, 125044. https://doi.org/10.1016/j.amc.2020.125044 15. goswami, n.k. and shanmukha, b. (2021). a mathematical analysis of zika virus transmission with optimal control strategies, comput. meth. diff. equ., 9(1), 117–145. 16. hugo, a., makinde, o.d., kumar, s. and chibwana, f.f. (2017). optimal control and cost effectiveness analysis for newcastle disease eco-epidemiological model in tanzania, j. biol. dyn., 11, 190–209. 17. kumar, v. and kumari, n. (2019). controlling chaos in three species food chain model with fear effect, aims math., 5(2), 828–842. 18. lenhart, s. and workman, j.t. (2007). optimal control applied to biological models. london: chapman & hall. 19. lasalle, j.p. (1976). the stability of dynamical dystems, regional conference series in applied mathematics. philadelphia pa: siam. 20. lotka, a.j. (1925). elements of physical biology. baltimore: williams and wilkins. 21. nandi, s.k., mondal, p.k., jana, s., haldar, p. and kar, t.k. (2015). prey-predator model with two-stage infection in prey: concerning pest control, j. nonlinear dynam., article id 948728 (2015), 1–13. 22. obabiyi, o.s. and olaniyi, s. (2019). global stability analysis of malaria transmission dynamics with vigilant compartment, electron. j. differential equations., 2019(09), 1– 10. 23. okyere, e., olaniyi, s. and bonyah, e. (2020). analysis of zika virus dynamics with sexual transmission route using multiple optimal controls. scientific african, 9, e00532. https://doi.org/10.1016/j.sciaf.2020.e00532 24. olaniyi, s., okosun, k.o., adesanya, s.o. and lebelo, r.s. (2020). modelling malaria dynamics with partial immunity and protected travelers: optimal control and costeffectiveness analysis. j. biol. dyn., 14, 90–115. copyright © 2021 assa. adv syst sci appl (2021) stability analysis and optimal measure 103 25. olaniyi, s., obabiyi, o.s., okosun, k.o., oladipo, a.t. and adewale, s.o. (2020). mathematical modelling and optimal cost-effective control of covid-19 transmission dynamics, eur. phys. j. plus, 135, 938. https://doi.org/10.1140/epjp/s13360-020-00954z 26. pontryagin, l.s., boltyanskii, v.g., gamkrelidze, r.v. and mishchenko, e.f. (1962). the mathematical theory of optimal processes, new york: wiley. 27. prasad, b.s.r.v., banerjee, m. and srinivasu, p.d.n. (2013). dynamics of additional food provided predator-prey system with mutually interfering predators. math. biosci., 246, 176–190. 28. safi, m.a. and darassi, m.h. (2018). mathematical analysis of a model for ectoparasite-borne diseases, math. meth. appl. sci., 41(17), 8248–8257. 29. sani, a., cahyono, e., mukhsar and rahman, g.a. (2014). dynamics of disease spread in a predator-prey system, adv. stud. biol. 6, 169–179. 30. van den driessche, p. and watmough, j. (2002). reproduction numbers and subthreshold endemic equilibria for compartmental models of disease transmission, math. biosci. 180, 29–48. 31. volterra, v. (1926). variazioni e fluttaazioni of numero di individul in specie animali conviventi, mem. accd. linc., 2, 31–113. copyright © 2021 assa. adv syst sci appl (2021) introduction model formulation well-posedness of the model boundedness of solutions positivity of solutions analysis of the model existence of equilibrium points trivial equilibrium point (e0) axial equilibrium point (e1) disease-free equilibrium point (e2) predator-free equilibrium point (e3) infected prey-free equilibrium (e4) infected predator-free equilibrium (e5) interior equilibrium point (e6) stability analysis stability of e0 stability of e1 stability of e2 stability of e3 stability of e4 stability of e5 stability of e6 optimal control model analysis of the control model simulations of the control model conclusion adv syst sci appl 2022; 03:1–17 published online at https://ijassa.ipu.ru. stability and bifurcation of a delay cancer model in the polluted environment ahmed a. mohsen1*, raid k. naji2 1department of mathematics, college of education for pure science (ibn-al-haitham), university of baghdad. 2department of mathematics, college of science, university of baghdad. abstract: it is well known that the spread of cancer or tumor growth increases in polluted environments. in this paper, the dynamic behavior of the cancer model in the polluted environment is studied taking into consideration the delay in clearance of the environment from their contamination. the set of differential equations that simulates this epidemic model is formulated. the existence, uniqueness, and the bound of the solution are discussed. the local and global stability conditions of disease-free and endemic equilibrium points are investigated. the occurrence of the hopf bifurcation around the endemic equilibrium point is proved. the stability and direction of the periodic dynamics are studied. finally, the paper is ended with a numerical simulation in order to validate the analytical results. keywords: cancer, environment pollution, time delay, stability, the hopf bifurcation. 1. introduction cancer is an abnormal growth of cells that tend to proliferate in an uncontrolled way and in some cases to spread. in fact, cancer or tumor growth can be considered as not one disease but it is a group of more than one hundred different and distinctive diseases. cancer can involve any tissue of the body and have many different forms in all body areas [2]. according to the world health organization, the cancer is one of the global diseases that affect the human as much as 18 million new cases annually, and more than half of the mortalities in 2018, and the numbers are expected to nearly double in 2040. although there is many advanced experimental in developing interventional therapies for cancer such as immunotherapy, virotherapy, targeted drug therapies, and chemotherapy, in addition to the surgical resection, treatment options are still limited and the disease is considered fatal [5]. therefore, we can consider the mathematical models as one of the ways for studying this disease. for a brief review of mathematical models describing tumor-immune dynamics see tsygvintsev et al. [12]. lestari et al [6] proposed a mathematical model of the spread of cancer with chemotherapy and then studied their dynamical behavior. weerasinghe et al [14] discussed some models of plasticity, tumor progression, and metastasis using three broadly conceived mathematical modeling techniques: discrete, continuum, and hybrid, each with advantages and disadvantages. simmons et al. [9] presented a brief overview of breast cancer, focusing on its heterogeneity, and then explained the role of mathematical modeling and simulation in teasing apart the underlying biophysical processes. in [8], pillis and radunskaya proposed and investigated the immune response to tumor invasion. trisilowati et al. [11] proposed and analyzed the optimal control model of dendritic cell treatment of growing ∗corresponding author: aamuhseen@gmail.com 2 a.a. mohsen, r.k. naji tumors. moreover, wilson and levy [15] formulated a model to investigate the effect of immunotherapy on tumor growth. recently, abernathy et al [16], studied the long-term dynamics of a system of nonlinear differential equations that describes the role of virotherapy on tumors and the impact of immune response specific to fighting cancer. al-tuwairqi et al [1] presented a realistic mathematical model that describes the interaction between the innate immune system and uninfected tumor cells. this is based on the fact that both tumor and virus-infected cells are recognized by natural killer cells that are part of the innate immune system on the other hand, there have been several papers proposed to show the effect of the time delay of some epidemic disease models as seen in zuo et al. [17], which studied the impact of the media and delay on the spread of an epidemic. in [3], cooke et al. proposed an epidemic model with two delays. wang and wu [13], investigated an seir model with delay. recently, naji and majeed [7] studied the effect of delay on a stage structure prey-predator model. in this article, it is assumed that cancer start outbreak due to environmental pollution, and then the effect of delay in cleaning up this environment on the dynamic behavior of the proposed system is discussed. so that, the next section deals with the formulation of the model and presents its main properties. in section 3, the stability analysis of the diseasefree equilibrium point and endemic equilibrium point is discussed and the possibility of occurrence of the hopf bifurcation is also studied. section 4, discusses the stability and direction of the periodic dynamics that resulting from the hopf bifurcation as the delay starting increases. numerical simulations are used in section 5 to further understand the dynamics of the model. finally, the paper is ended with a discussion section as given in section 6. 2. the mathematical formulation of cancer model it is well known that one of the main reasons for the spread of cancer is the polluted environment. hence, in case of existence of such disease in a population of size n(t) at time t, then the population is divided into two compartments, namely susceptible compartment, which is denoted to their population’s size at time t by s(t) and cancer compartment, which contains all the individuals that infected by cancer and denoted to their population’s size at time t by c(t), such that n (t) = s (t) + c(t). moreover, it is assumed that the pollution level at time t in the environment is given by e(t). on the other hand, since the delay in clearance of the environment from their contamination has a vital effect on the spread of cancer. therefore, the effect of delay, with the amount of delay given by τ > 0, is considered in the formulation of the model. accordingly, the dynamics of the spread of cancer within such an environment can be represented in the following block diagram. fig. 2.1. the block diagram of cancer model. copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 3 and the corresponding dynamic model has been formulated by the following system of nonlinear delay differential equations ds dt = ψ − β (1−m)s (t)e (t)− µs (t) dc dt = β (1−m)s (t)e (t)− (µ+ α)c (t) de dt = p − θe (t− τ)  , (2.1) where s (r) = s (0) > 0, c (r) = c(0) ≥ 0 and e (r) = e (0) ≥ 0 for r ∈ [−τ, 0). it is assumed that all parameters in system (2.1) have the following description: ψ > 0 is the natural birth of susceptible; µ > 0 represents the death rate of all populations; β > 0 represents the contact rate of the susceptible with the contaminated environment. however, the constant m represents the body resistance rate due to chromosomal fixed against environment effect such that 0 ≤ m ≤ 1. while p ≥ 0 is the pollution rate in the environment and θ > 0 is the clearance rate in the environment. note that, all the interaction functions on the right-hand side of the system (2.1) are continuous and have continuous partial derivatives and hence the system has a unique solution in r3 +. furthermore, the boundedness of the solution is shown in the following theorem. theorem 2.1: all the solutions of system (2.1) with initial conditions belong to r3 +, are bounded. proof let (s (t) , c (t) , e (t)) be any solution of system (2.1) with initial conditions belongs to r3 +. hence, by adding the first two equations in system (2.1) to each other we get: ds dt + dc dt = ψ − µn − αc. therefore, it is obtained that: dn dt + µn ≤ ψ. then, solving the above inequality by using the gronwall lemma we obtain for t→ ∞ that: n ≤ ψ µ . similarly, from the 3rd equation in the system (2.1) we have: de dt = p − θe. this implies that: de dt + θe = p. hence for t→ ∞, we have e(t) ≤ p θ . this completes the proof. now, since the variable c in the second equation of system (2.1) dose not appear in the other two equations, then the following system, say system (2.3), can be solved independently. then by substituting the obtained solution, say (sυ, eυ) in second equation of system (2.1), we obtain the following solution: copyright © 2022 assa. adv syst sci appl (2022) 4 a.a. mohsen, r.k. naji c (t) = β (1−m)sυeυ µ+ α , (2.2) where (sυ, eυ) is a solution of the system: ds dt = ψ − β (1−m)s (t)e (t)− µs (t) de dt = p − θe (t− τ)  . (2.3) 3. analysis of system 2.3 in this section, the existence and their stability analysis are carried out. it is easy to verify that system (2.3) has two equilibrium points and can be described as follows: 1. the first point is called the disease-free equilibrium point dfep, which is denoted by e0 = (s0, 0), where: s0 = ψ µ . (3.4) clearly, e0 exists provided that: p = 0. (3.5) 2. the other point is called the endemic equilibrium point eep and that denoted by e1 = (s1, e1), where: s1 = ψθ β(1−m)p+µθ e1 = p θ  . (3.6) obviously, e1 exists uniquely in the interior of the positive quadrant of the se−plane provided that: p > 0. (3.7) now, to analyze the stability of the above equilibrium points the general jacobian matrix for the system (2.3) at any point, say (s,e), can be written as: j (s,e) = [ − [β (1−m)e + µ] −β (1−m)s 0 −θe−λτ ] . (3.8) accordingly, the jacobian matrix of system (2.3) at the dfep that given in eq. (3.4), becomes: j (e0) = [ −µ −β (1−m)s0 0 −θe−λτ ] . (3.9) then, the characteristic equation of the matrix (3.9) can be written as: λ2 + µλ+ (θλ+ µθ) e−λτ = 0. (3.10) obviously, when τ = 0, eq. (3.10) becomes: copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 5 λ2 + (µ+ θ)λ+ µθ = 0. (3.11) clearly, the eigenvalues in such a case can be written as: λ [τ=0] 1 = −µ, λ[τ=0] 2 = −θ. hence, e0 is always locally asymptotically stable for τ = 0. furthermore, by using the function h(s,e) = ( 1 se ) as the dulac function, which is a continuously differentiable unction in a simply connected ω ⊂ r2 + such that ∂(hf(s,e)/∂s + ∂(hg(s,e)/∂e dose not change the sign in ω and vanishes at most on a set of measure zero then the system (2.3) not have periodic orbits in ω. therefore, it is easy to verify that system (2.3) with τ = 0 has no periodic dynamics and hence e0 is always globally asymptotically stable. on the other hand, since the dfep exists under the condition (3.5), hence system (2.3) can be reduce to a single equation due to the fact é < 0. therefore, the system (2.3) has no periodic dynamics when τ > 0 and p = 0, and hence e0 is always globally absolutely stable [10] for all τ ≥ 0. now the stability of the eep of system 2.3 that given by (3.6), is investigated using the linearization technique. accordingly, the jacobian matrix that given in (3.8) of system (2.3) at the eep can be written as: j (e1) = [ − [β (1−m)e1 + µ] −β (1−m)s1 0 −θe−λτ ] . (3.12) hence, the characteristic equation of j(e1) can be written: λ2 +b1λ+ (b2λ+b3) e −λτ = 0, (3.13) where b1 = β (1−m)e1 + µ, b2 = θ, and b3 = θ [β (1−m)e1 + µ] . so for τ = 0, then eq. (3.13) becomes: λ2 + (b1 +b2)λ+b3 = 0. (3.14) it easy to see that the roots of the above equation have negative real parts. consequently, whenever the eep exists, it is always locally asymptotically stable (las). indeed, it is globally asymptotically stable as mentioned above with the help of the dulac criterion. now, for τ > 0, assume that eq. (3.13) has a pair of purely imaginary roots, say λ = ±iω with ω > 0, cross the imaginary axis. hence, by substituting λ = iω in eq. (3.13) we get: −ω2 + ib1ω + (ib2ω +b3) (cosωτ − isin ωτ ) = 0. therefore, by separating the real and imaginary parts we obtain: b1ω = b3sinωτ −b2ω cosωτ −ω2 = −b2ωsinωτ −b3cosωτ } . (3.15) clearly, squaring the equations in eq. (3.15) and adding to each other gives that: ω4 + ( b2 1 −b2 2 ) ω2 −b2 3 = 0, (3.16) set k = ω2, then eq. (3.14) becomes: k2 + ( b2 1 −b2 2 ) k −b2 3 = 0. (3.17) copyright © 2022 assa. adv syst sci appl (2022) 6 a.a. mohsen, r.k. naji clearly, by using descartes’ rule of signs, eq. (3.17) has a unique positive root, namely k = ω2 0 . hence, ω0 is a positive root of eq. (3.16) too. therefore, there is at least a pair of purely imaginary roots ±iω0 satisfying eq.(3.13). substituting ω0 in eq. (3.15) and then solving the resulting equation with respect to τ we get that: τj = 1 ω0 sin−1ω0 (b1b3 +b2ω0 2)( b2 2ω0 2 +b3 2 ) + 2jπ ω0 ; j = 0, 1, 2, · · · (3.18) now, define that τ0 = minj≥0 τj , then λ (τ) = γ (τ) + iω(τ) be a root of eq. 3.13, such that γ (τ0) = 0 and ω (τ0) = ω0. then we have the following theorem. theorem 3.1: assume that condition (3.7) holds. then the eep is conditionally stable. proof it is well known that the equilibrium point e1 is conditionally stable if it is asymptotically stable for τ ∈ [0, τ0). moreover, as shown above, the equilibrium point e1 is globally asymptotically stable for τ = 0 and the transcendental characteristic equation given by eq. (3.13) has at least a pair of purely imaginary roots ±iω0 at τ = τ0. consequently, if we assume that λ (τ) = γ (τ) + iω(τ) is an eigenvalue of eq. (3.13) with γ (τ0) = 0 and ω (τ0) = ω0 > 0, where τ0 is given by eq. (3.18) then the proof follows if we can show that the following transversality condition hold.[ d(reλ (τ)) dτ ] τ=τ0 ̸= 0. (3.19) by deriving eq. (3.13) with respect to τ we obtain that:[ 2λ+b1 +b2e −λτ − τ(b2λ+b3)e −λτ] dλ dτ = λ(b2λ+b3)e −λτ . (3.20) therefore, we obtain that:[ dλ dτ ]−1 = (2λ+b1) −λ(λ2 +b1λ) + b2 λ(b2λ+b3) − τ λ . (3.21) since λ = iω0 at τ = τ0, then eq. (3.21) can be rewritten as:[ dλ dτ ]−1 τ=τ0 = 2iω0 +b1 b1ω0 2 + iω0 3 + b2 −b2ω0 2 + iω0b3 − τ0 iω0 . now since sgn [ d(reλ) dτ ] τ=τ0 = sgn [ re ( dλ dτ )−1 ] λ=iω0 . (3.22) accordingly, from the fact: re [ 2iω0 +b1 b1ω0 2 + iω0 3 ] = 2ω0 2 +b1 2 ω0 2(b1 2 + ω0 2) , re [ b2 iω0b3 −b2ω0 2 ] = −b2 2 b2 2 ω0 2 +b3 2 , copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 7 re [ τ0 iω0 ] = 0. hence, we obtain that:[ re ( dλ dτ )]−1 τ=τ0 = 2ω0 2 +b1 2 ω0 2(b1 2 + ω0 2) − b2 2 b2 2 ω0 2 +b3 2 > 0. therefore, d(reλ(τ)) dτ at τ = τ0 is not equal zero and does not change sign. clearly, the obtained result shows that the eigenvalue of characteristic equation eq. (3.13) crosses the imaginary axis from left to right as τ passes through τ0. hence system (2.3) losses its stability and undergoes the hopf bifurcation at τ = τ0. thus the proof is complete. 4. stability and direction of the hopf bifurcation in the previous section, we have shown that system (2.3) undergoes the hopf bifurcation near the eep at τ = τj . in this section, however, the direction of the hopf bifurcation and their stability at τ = τ0 is investigated. the normal form theory and center manifold theorem are used in this investigation, see [4]. let x1 = s − s1, x2 = e − e1, xi(t) = xi(τt) and τ = τ0+u, then we drop the bars for simplification of notations. therefore system (2.3) is transformed into functional differential equations in c = c ([−1, 0] , r2) as: ẋ (t) = lu (xt) +h (u, xt) , (4.23) where x (t) = (x1, x2) t ∈ r2, lu : c → r2 and h:r×c → r2 are given as: lu (γ) = (τ0 + u) [ − [β (1−m)e1 + µ] −β(1−m)s1 0 0 ] [ γ1(0) γ2(0) ] +(τ0 + u) [ 0 0 0 −θ ] [ γ1(−1) γ2(−1) ]  , (4.24) and h (u,γ) = (τ0 + u) [ −β(1−m)γ1(0)γ2(0) 0 ] , (4.25) here γ(v) = (γ1(v),γ2(v)) t ∈ c. now, by the riesz representation theorem, there exists a function of bounded variation, say η(v, u) for v ∈ [−1, 0], such that: luγ = ∫ 0 −1 dη (v, u) γ (v),γ ∈ c. (4.26) here we can choose: η (v, u) = (τ0 + u) [ − [β (1−m)e1 + µ] −β(1−m)s1 0 0 ] δ (v) − (τ0 + u) [ 0 0 0 −θ ] δ (v + 1)  , (4.27) where δ(v) is the dirac function, which define as follows, then eq. (4.26) is satisfied. copyright © 2022 assa. adv syst sci appl (2022) 8 a.a. mohsen, r.k. naji δ (v) = { 1 v = 0. 0 v ̸= 0. (4.28) for γ ∈ c1 ([−1, 0] , r2) , define: a (u) γ(v) = { dγ(v) dv − 1 ≤ v < 0.∫ 0 −1 dη (ς, u) γ (ς) v = 0. (4.29) and r (u) γ(v) = { 0 − 1 ≤ v < 0. h (u,γ) v = 0. (4.30) hence, system (4.23) can be transformed into an operator differential equation of the form: x́t = a (u)xt +r (u)xt, (4.31) where xt = x (t+ v), v ∈ [−1, 0]. now, we defined the adjoint operator a∗ of a, where ψ ∈ c1 ( [−1, 0] , (r2) ∗), by: a∗ (0)ψ(s) = { −dψ(s) ds 0 < s ≤ 1,∫ 0 −1 dηt (ς, 0)ψ (−ς) s = 0, (4.32) with a bilinear inner product form: ⟨ψ (s) ,γ(v)⟩ = ψ(0)γ (0)− ∫ 0 v=−1 ∫ v ξ=0 ψ t (ξ − v) dη (v) γ (ξ) dξ, (4.33) where, η (v) = η (v, 0). clearly, a (0) and a∗(0) are adjoint operators. now since ±iτ0ω0 are the eigenvalues of a (0) as we show in the previous section then they are also eigenvalues of a∗(0). furthermore, we need to find the eigenvector q of a corresponding to iτ0ω0 and the eigenvector q∗ of a∗ corresponding to −iτ0ω0 so that they satisfying the normalization conditions ⟨q∗, q⟩ = 1 and ⟨q∗, q⟩ = 0, as determined in the following theorem. theorem 4.1: the eigenvectors of a (0) and a∗ corresponding to iτ0ω0 and −iτ0ω0 are given by q (v) = (1, r1) t eivτ0ω0 and q∗ (s) = d(1, r2) t eisτ0ω0 respectively, where: r1 = − [iω0 + β (1−m)e1 + µ] β (1−m)s1 , r2 = β (1−m)s1 iω0 − θeiω0τ0 , d = [ 1 + r1r2 − τ0θr1r2e −iω0τ0 ]−1 . proof since q (v) = (1, r1) t eivτ0ω0 , is the eigenvector of a(0) corresponding to iτ0ω0, then: a (0) q (v) = iτ0ω0 q (v) . so, we obtain: copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 9 a (0) q (0) eivτ0ω0 = iτ0ω0 q (0) e ivτ0ω0 . now according to the definition of a (0), we get τ0 ( iω0 + β (1−m)e1 + µ β (1−m)s1 0 iω0 + θe−iω0τ0 )( 1 r1 ) = ( 0 0 ) . consequently, we obtain that: r1 = − [iω0 + β (1−m)e1 + µ] β (1−m)s1 . similarly, since q∗ (s) = d(1, r2) t eisτ0ω0 is the eigenvector of a∗ corresponding to −iτ0ω0, hence: a∗q∗ (s) = −iτ0ω0q ∗ (s) . from the definition of a∗ we obtain: τ0 ( iω0 − [β (1−m)e1 + µ] 0 −β (1−m)s1 iω0 − θe−iω0τ0 )( 1 r2 ) = ( 0 0 ) . such that: r2 = β (1−m)s1 iω0 − θeiω0τ0 . to guarantee that ⟨ q∗ (s) , q (v) ⟩ = 1, then the parameter d should be determined, so by using eq. (4.33), we get that: ⟨q∗ (s) , q(v)⟩ = d (1, r2) ( 1 r1 ) − ∫ 0 v=−1 ∫ v ξ=0 d (1, r2) e −i(ξ−v)ω0τ0dη (v) ( 1 r1 ) eiξω0τ0dξ = d { 1 + r1r2 − ∫ 0 −1 (1, r2) ve ivω0τ0dη (v) ( 1 r1 )} = d { 1 + r1r2 + [ τ0 (1, r2) ( 0 −θr1 )] e−iω0τ0 } = d {1 + r1r2 − τ0θr1r2e −iω0τ0} . thus we can choose d so that: d = [ 1 + r1r2 − τ0θr1r2e −iω0τ0 ]−1 . in addition, it can be easily verify that ⟨q∗, q⟩ = 0 by applying the adjoint property ⟨φ,aγ⟩ = ⟨a∗φ,γ⟩, as follows: iτ0ω0 ⟨q∗, q⟩ = ⟨−iτ0ω0q ∗, q⟩ = ⟨a∗q∗, q⟩ = ⟨q∗, aq⟩ = ⟨q∗, iτ0ω0q⟩ = −iτ0ω0 ⟨q∗, q⟩ . then, we have ⟨q∗, q⟩ = 0. in the following, the technique given by hassard et al. [4] to compute the coordinates describing center manifold c0 at u = 0 is used. define copyright © 2022 assa. adv syst sci appl (2022) 10 a.a. mohsen, r.k. naji z (t) = ⟨q∗, xt⟩ w (t, v) = xt (v)− z (t) q (v)− z (t) q (v) = xt (v)− 2re {z (t) q (v)} } , (4.34) where xt (v) = x(t+ v) be the solution of eq. (4.24). also, on the center manifold c0, we have: w (t, v) = w (z (t) , z (t) , v) , (4.35) here w (z (t) , z (t) , v) = w20 (v) z2 2 + w11 (v) zz + w02 (v) z2 2 + . . . , such that, z(t) and z(t) are local coordinates of center manifold c0 in the direction of q∗ and q∗, respectively. clearly, we know that xt ∈ c0 of eq. (4.31), since u = 0, we know that ⟨φ,aγ⟩ = ⟨a∗φ,γ⟩, for ⟨γ, φ⟩ ∈ d(a)×d(a), then: ź(t) = ⟨q∗, x′(t)⟩ = ⟨q∗, a(0)xt +r(0)xt⟩ = ⟨q∗, a (0)xt⟩+ ⟨q∗, r (0)xt⟩ = ⟨a∗(0)q∗, xt⟩+ ⟨q∗, r (0)xt⟩ . thus ź (t) = iω0τ0z (t) + q∗ t (0)h(0, w (t, 0) + 2re {z (t) q(0)}) = iω0τ0z (t) + q∗ t (0)h0 (z, z) } , (4.36) here h0 (z, z) = h(0, z, z), now rewrite eq. (4.36) as: ź (t) = iω0τ0z (t) + g (z, z) , (4.37) where g (z, z) = 1 2 g20z 2 + g11zz + 1 2 g02z 2 + 1 2 g21z 2z + . . . so according to the eq. (4.34) with the eq. (4.31), we obtain: ẃ (t, v) = x′t (v)− ź (t) q (v)− ź (t) q (v) = a(0)xt +r(0)xt − iω0τ0z(t)q(v)− q∗ t (0)h0(z, z)q(v) +iω0τ0z (t) q (v)− q∗t (0)h0 (z, z) q (v) } = a(0)xt +r(0)xt − a(0)z(t)q (v)− a(0)z(t)q (v) −2re { q∗ t (0)h0(z, z)q(v) } } = a(0)w(t, v) +r(0)xt − 2re { q∗ t (0)h0(z, z)q(v) } . therefore, it is obtain that: ẃ (t, v) = a (0)w (t, v)− 2re { q∗ t (0)h0 (z, z) q (v) } for v ∈ [−1, 0) . a (0)w (t, 0) +h0 (z, z)− 2re { q∗ t (0)h0 (z, z) q (0) } for v = 0. (4.38) accordingly, we can rewrite eq. (4.38) for v ∈ [−1, 0) as follows: ẃ (t, v) = aw (t, v) +g (z, z, v) , (4.39) copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 11 where g (z, z, v) = g20 (v) z2 2 +g11zz +g02 z2 2 + . . . so, differentiate the two sides of (4.35), gives that: ẃ = wz ź + wz ź. (4.40) then, using the eqs. (4.35), (4.36) and (4.39) give that: (a− 2iτ0ω0)w20 (v) = −g20 (v) , aw11 (v) = −g11 (v) , (a+ 2iτ0ω0)w02 (v) = −g02 (v) . (4.41) according to eqs. (4.36) and (4.37), we have that: g (z, z) = q∗ t (0)h0(z, z) = 1 2 g20z 2 + g11zz + 1 2 g02z 2 + 1 2 g21z 2z + . . . , (4.42) where q∗t (0) = d (1, r2). also from eq. (4.34), it gets that: x (t+ v) = (x1 (t+ v) , x2(t+ v))t = zq (v) + zq (v) + w (t, v) , here q (v) = (1, r1) t eivτ0ω0 . hence by eq. (4.35), it is obtained that: x1 (t+ 0) = z + z + w(1) (t, 0) = z + z + w (1) 20 (0) z 2 2 + w (1) 11 (0) zz + w (1) 02 (0) z 2 2 + . . . , x2 (t+ 0) = zr1 + zr1 + w(2) (t, 0) = zr1 + zr1 + w (2) 20 (0) z 2 2 + w (2) 11 (0) zz + w (2) 02 (0) z 2 2 + . . . , x1 (t− 1) = ze−iτ0ω0 + zeiτ0ω0 + w(1) (t,−1) = ze−iτ0ω0 + zeiτ0ω0 + w (1) 20 (−1) z 2 2 +w (1) 11 (−1) zz + w (1) 02 (−1) z 2 2 + . . . , x2 (t− 1) = zr1e −iτ0ω0 + zr1e iτ0ω0 + w2 (t,−1) = zr1e −iτ0ω0 + zr1e iτ0ω0 + w (2) 20 (−1) z 2 2 +w (2) 11 (−1) zz + w (2) 02 (−1) z 2 2 + . . . . (4.43) now, substituting eq. (4.25), when u = 0 and γ(v) ≡ x(t+ v) in eq. (4.42), gives that: g (z, z) = dτ0 (1, r2) ( −β (1−m)x1 (t+ 0)x2 (t+ 0) 0 ) . therefore, the following is obtained: g (z, z) = −dτ0β (1−m) [ r1z 2 + (r1 + r1) zz + r1z 2 + ( 2w (2) 11 (0) + w (2) 20 (0) + 2r1w (1) 11 (0) + r1w (1) 20 (0) ) z2z 2 + ( 2w (2) 11 (0) + w (2) 20 (0) + 2r1w (1) 11 (0) + r1w (1) 02 (0) ) z z2 2 + · · · ]  . (4.44) copyright © 2022 assa. adv syst sci appl (2022) 12 a.a. mohsen, r.k. naji so comparing the coefficients in eq. (4.44) with those in eq. (4.42) gives that: g20 = −2dτ0β (1−m) r1, g11 = −dτ0β (1−m) (r1 + r1) , g02 = −2dτ0β (1−m) r1, g21 = −dτ0β (1−m) ( 2w (2) 11 (0) + w (2) 20 (0) + 2r1w (1) 11 (0) + r1w (1) 20 (0) ) . (4.45) since eq. (4.45) contains w11 and w20, hence it needs to compute them. now, from eqs. 4.38 and (4.39), for v ∈ [−1, 0), we have: g (z, z, v) = −q∗t (0)h0 (z, z) q (v)− q∗t (0)h0 (z, z) q (v) . then by using eqs. (4.36) and (4.37), it obtains that: g (z, z, v) = −g (z, z) q (v)− g (z, z) q (v) . (4.46) by comparing coefficients in both sides gives: g20 (v) = −g20q (v)− g02q (v) , g11 (v) = −g11q (v)− g11q (v) , (4.47) now, substituting eqs. (4.47) into eqs. (4.41) respectively gives that: w′ 20 (v) = 2iτ0ω0w20 (v) + g20q (v) + g02q (v) , w′ 11 (v) = g11q (v) + g11q (v) . (4.48) now by solving eqs. (4.48), it is easy to verify that the solutions are written respectively as: w20 (v) = −g20q(0) iτ0ω0 eiτ0ω0v − g02q(0) 3iτ0ω0 e−iτ0ω0v +k1e 2iτ0ω0v, w11 (v) = g11q(0) iτ0ω0 eiτ0ω0v − g11q(0) iτ0ω0 e−iτ0ω0v +k2. (4.49) wherek1 = ( k (1) 1 , k2 1 ) andk2 = ( k (1) 2 , k2 2 ) are constants vectors to be determined in the following. from the definition of a when v = 0, and eqs. (4.41), it is obtained that:∫ 0 −1 dη (v)w20 (v) = 2iτ0ω0w20 (0)−g20(0). (4.50) ∫ 0 −1 dη (v)w11 (v) = −g11(0). (4.51) also from eqs. (4.38) and (4.39), it is easy to verify that: g20 (0) = −g20q (0)− g02q (0) + 2τ0 [ r1β (1−m) 0 ] . (4.52) g11 (0) = −g11q (0)− g11q (0) + 2τ0 [ β (1−m) (r1 + r1) 0 ] . (4.53) therefore by substituting eqs. (4.49) and (4.52) into eq. (4.50), and then using the results: copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 13 ( iτ0ω0i − ∫ 0 −1 eiτ0ω0vdη (v) ) q (0) = 0,( −iτ0ω0i − ∫ 0 −1 e−iτ0ω0vdη (v) ) q (0) = 0. (4.54) it is obtained that:( 2iτ0ω0 i − ∫ 0 −1 e2iτ0ω0vdη (v) ) k1 = 2τ0 [ r1β(1−m) 0 ] . (4.55) similarly, by substituting eqs. (4.49) and (4.53) into eq. (4.51), and then using eqs. (4.54), it is obtained that: k2 ∫ 0 −1 dη (v) = −2τ0 [ β (1−m) (r1 + r1) 0 ] . (4.56) now, using eq. (4.28) with u = 0 into eqs. (4.55) and (4.56), gives that: [ 2iω0 + β (1−m)e1 + µ β (1−m)s1 0 2iω0 + θe−2iω0τ0 ] k1 = 2 [ r1β(1−m) 0 ] . (4.57) [ −β (1−m)e1 − µ −β (1−m)s1 0 −θ ] k2 = −2 [ β (1−m) (r1 + r1) 0 ] . (4.58) accordingly, solving the linear systems given by eqs. (4.57) and (4.58), gives that: k (1) 1 = 2r1β(1−m)[2iω0+θe−2iω0τ0 ] [2iω0+β(1−m)e1+µ][2iω0+θe−2iω0τ0 ] , k (2) 1 = 0. (4.59) and k (1) 2 = 2β(1−m)(r1+r1) [β(1−m)e1+µ] , k (2) 2 = 0. (4.60) consequently, all the values of gij; i = 0, 1, 2; j = 0, 1, 2 given in eqs. (4.45) can be determined by the parameters and delay. thus, it can calculate the following quantities: c1 (0) = i 2τ0ω0 ( g20g11 − 2|g11|2 − |g02|2 3 ) + g21 2 , µ2 = −re{c1(0)} re{λ′(τ0)} , β2 = 2re {c1(0)} , t2 = − im {c1(0)}+µ2im {λ′(τ0)} τ0ω0 . (4.61) these quantities are used to determine the direction of the hopf bifurcation and stability of bifurcated periodic solutions of system (2.3) at the critical value τ0 as shown in the following theorem. theorem 4.2: the direction of the hopf bifurcation is determined by the sign of µ2 at the critical value τ0, so that copyright © 2022 assa. adv syst sci appl (2022) 14 a.a. mohsen, r.k. naji 1. the the hopf bifurcation is supercritical if µ2 > 0 and the hopf bifurcation is subcritical if µ2 < 0. 2. the stability of bifurcated periodic solutions is determined by β2 so that the periodic solutions are stable if β2 < 0 and unstable if β2 > 0. 3. the period of bifurcated periodic solutions is determined by β2 so that the period increases if t2 > 0 and decreases if t2 < 0. 5. numerical simulation in this section, the obtained results are illustrated using numerical simulation. the following set of hypothetical parameters set is adopted throughout this section. ψ = 0.87, β = 0.010, m = 0.05, µ = 0.015, p = 0, α = 0.3, θ = 0.2, τ = 7.7 < τ0 ∼= 7.85. (5.62) all the obtained trajectories of the system (2.1) are drawn using the matlab of version 8. starting from different sets of initial data, system (2.1) is solved numerically using the parameters set given in eq. (5.62) and then the obtained trajectories are drawn in fig. (5.2). fig. 5.2. : the trajectories of the system (2.1) using data given by eq. (5.62) approach to dfep. (a) trajectories of susceptible. (b) trajectories of cancer. (c) trajectories of environment. (d) 3d phase plot for globally asymptotically stable dfep. clearly, fig. (5.2) illustrates that system (2.1) has a globally asymptotically stable dfep for that data (5.62). however, for the same data given by eq. (5.62) with p = 0.6, it is observed that, although that system (2.1) is solved with different sets of initial points, it has a globally asymptotically stable eep given by e1 = (17.9, 1.6, 3) as shown in fig. (5.3) below. copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 15 fig. 5.3. : the trajectories of the system (2.1) using data given by eq. (5.62) with p = 0.6 approach to eep. (a) trajectories of susceptible. (b) trajectories of cancer. (c) trajectories of environment. (d) 3d phase plot for globally asymptotically stable eep. now, to show the effect of varying the body resistance rate (m) due to chromosomal fixed against environment effect on the system behavior, the system is solved numerically for different values, say m = 0.1, 0.7 respectively, keeping other parameters fixed as given in eq. (5.62) with p = 0.6, and then the solution of system (2.1) is drawn in figs. (5.4a) and (5.4b). fig. 5.4. : time series of the trajectory of the system (2.1) using data given by eq. (5.62) with p = 0.6 and different values of m. (a) trajectory approaches to eep given by e1 = (18, 1.5, 3) when m = 0.1. (b) trajectory approaches to eep given by e1 = (32, 0.92, 3) when m = 0.7. clearly, in fig. (5.4) as m increases, the trajectory of the system (2.1) still approaches asymptotically to the eep point. in fact, it is observed that as m increases the number of individuals with cancer decreases, and the number of susceptible increases without any effect of environment. copyright © 2022 assa. adv syst sci appl (2022) 16 a.a. mohsen, r.k. naji now, the dynamical behavior of the system (2.1) near the eep point under the effect of increasing the time delay is investigated. the system (2.1) is solved numerically for the set of parameter values given by eq. (5.62) with p = 0.6 and τ = 7.86 then the trajectory of the system (2.1) is drawn in figs. (5.5a-5.5e). fig. 5.5. : the existence of periodic solution near eep of the system (2.1) for data given by equation (5.62) with p = 0.6 and τ = 7.86. (a) trajectories of susceptible. (b) trajectories of cancer. (c) trajectories of environment. (d) 3d periodic solution. 6. conclusion in this paper, a mathematical model that describes the spread of cancer in a polluted environment incorporating delay in cleaning up the environment from the contaminated has been proposed and studied. the properties of the solution are discussed. it is observed that the proposed model has two equilibrium points namely dfep and eep. the stability analysis of the model shows that the dfep is globally asymptotically stable for all τ ≥ 0. while the eep is conditionally stable so that it’s globally asymptotically stable for all τ ∈ [0, τ0) and there is the hopf bifurcation at τ0. however, it is an unstable point for τ > τ0 and periodic dynamics occurred. the stability and direction of the periodic dynamics are also investigated analytically by finding the normal form using the center manifold theory as well as numerically. it is observed that for the hypothetical set of data given by eq. (5.62) with p = 0.6 the eep is still globally asymptotically stable for 0 ≤ ô < τ0 ∼= 7.853. while the hopf bifurcation occurs and periodic solutions bifurcate near the eep point as τ passes through the above critical value τ0. on the other hand, all the quantities, which are given in eqs. (4.61), are determined for the periodic dynamics drawn in fig. 5.5 as c1 (0) = 7.360− 77.426i, β2 = 14.72 > 0, µ2 = −0.294 < 0 and t2 = 51.873 > 0. therefore, according to theorem (4.2), the periodic copyright © 2022 assa. adv syst sci appl (2022) stability and bifurcation of a delay cancer model ... 17 resulting from the hopf bifurcation is subcritical, unstable, and the size of the period increases. acknowledgements the authors are thankful to the referees for their valuable suggestions. references 1. al-tuwairqi, s. m., al-johani n., o., & eman a. s., (2020). modeling dynamics of cancer virotherapy with immune response. advances in difference equations, 438. 2. chaffer, c. l. & weinberg, r. a., (2011). a perspective on cancer cell metastasis. science journal, 331, 1559–1564. 3. cooke, k. l., & driessche, v. d., (1996). analysis of an seirs epidemic model with two delays, j. math. boil., 35, 240–260. 4. hassard, b., kazarino d., & wan, y., (1981). theory and application of the hopf bifurcation, cambridge university press. 5. international agency for research on cancer. (16 october 2019). [online] available: https://gco.iarc.fr/tomorrow/home. 6. lestari, d., sari, e. r. & arifah, h., (2019). dynamical of a mathematical model of cancer cells with chemotherapy. journal of physics conf. series, 1320. 7. naji, r. k. & majeed, s. j., (2020). the dynamical analysis of a delay pray-predator model with a refuge-stage structure prey population. iranian journal of mathematical sciences ad informatics, 15(1), pp 135–159. 8. pillis, l. c. & radunskaya, a., (2003). a mathematical model of immune response to tumor invasion, computational fluid and solid mechanics, pp. 1661–1668, elsevier science, amsterdam. 9. simmons, a., burrage, p. m., nicolau, d.v., lakhani s.r. & burrage, k., (2017). environmental factors in breast cancer invasion: a mathematical modeling review. pathology, 49(2), 172–180. 10. tipsri, s. & chinviriyasit w., (2015). the effect of time delay on the dynamics of an seir model with nonlinear incidence. chaos solitons & fractals 75, 153–172. 11. trisilowati, t., mccue, s., & mallet, d., (2013). numerical solution of an optimal control model of dendritic cell treatment of growing tumour, anziam journal, 54, 664– 680. 12. tsygvintsev, a., marino, s. & kirschner, d. e., (2012). a mathematical model of gene therapy for the treatment of cancer, in mathematical models and methods in biomedicine, 357–373. springer-verlag. 13. wang, l., & wu, x., (2018). stability and the hopf bifurcation for an seir epidemic model with delay, advances in the theory of nonlinear analysis and its applications, 113–127. 14. weerasinghe, h. n., burrage, p. m., burrage, k. & nicolau d.v., (2019). mathematical models of cancer call plasticity. journal of oncology, 2019, 2403483. 15. wilson, s. & levy, d., (2012). a mathematical model of enhancement of tumor vaccine efficacy by immunotherapy, bulletin of mathematical biology, 74(7), 1485–1500. 16. zachary, a., kristen a., & jessica s., (2020). a mathematical model for tumor growth and treatment using virotherapy. aims mathematics, 5(5), 4136–4150. 17. zuo, l., liu, m. & wang, j., (2015). the impact of awareness program with recruitment and delay on the spread of an epidemic, mathematical problems in engineering, 235935. copyright © 2022 assa. adv syst sci appl (2022) introduction the mathematical formulation of cancer model analysis of system 2.3 stability and direction of the hopf bifurcation numerical simulation conclusion adv syst sci appl 2017; 3:42–48 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/503 interval-valued data forecasting using dual-parametric neural networks yuliya. e. polozova1, pavel. v. saraev1 1lipetsk state technical university, lipetsk, russia abstract: this study is focused on the approach to modelling and forecasting interval-valued data using a dual-parametric neural network (dpnn). this concept was proposed as a subclass of interval neural networks that contains two types of parameters: real and interval ones. this approach makes it possible to get guaranteed inclusion of an exact (single value) solution into interval calculation results. in this paper we intend to give a theoretical overview of previous research on the subject and description of the new developed methods and algorithms for learning dpnn. the experiments demonstrate that interval calculation results obtained by using the proposed approach include of exact solution at least in 60% of cases. keywords: interval neural network, forecasting, dual-parametric neural network, intervalvalued data, parametric identification 1. introduction today neural networks are successfully applied to many different problems, particularly time series data forecasting. these applications usually use sets of single point values as input and output data. but in many applications it is more natural to use the input values and the predicted results in the form of intervals. an interval neural network (inn), which contains interval arithmetic, is used to calculate such interval-valued data. one of the most important features of interval analysis is the guaranteed inclusion of the exact solution into interval calculation results. previous studies show that this feature is not achieved in the case of inn. trying to eliminate this drawback the authors introduced dualparametric neural network (dpnn) as a new subclass of inn. models of this subclass contain two types of parameters: real and interval ones. in this paper, we generalize previously developed and new learning methods and algorithms of dpnn and apply them to forecasting interval-valued time series. 2. theoretical overview interval neural network (inn) is a neural network that contains at least one interval parameter, namely input, output, or weight. the inns can be used for the following reasons: • initial data are sometimes presented in the form of interval values rather than singlepoint ones; • training data sets are reduced in size due to the use of clustering analysis; • inns allow for making reliable accuracy estimates of calculation results. ∗corresponding author: julipolozova@yandex.ru http://ijassa.ipu.ru/ojs/ijassa/article/view/503 interval-valued data forecasting using dual-parametric neural networks 43 the existing training methods for interval neural networks are based on the backpropagation (bp) algorithm for interval data [1–5] with training error calculated as the quadratic function of quality: j(w) = 1 2 k∑ i=1 qi(w) = 1 2 k∑ i=1 { (yi(w)− ỹi)2 + (yi(w)− ỹi)2 } (2.1) where k is the number of examples in the training data set, w is the weight vector, qi(w) is the training error on example i of the training data set, yi(w) is the lower bound of the network‘s output interval on example i, ỹi is the lower bound of the training output interval on example i, yi(w) is the upper bound of the network‘s output interval on example i, ỹi is the upper bound of the training output interval on example i. the bounds of the output interval for neuron i (index i is already used for training samples. it would be more clear to use another letter, say, k or j for neurons) of layer l are computed as follows: y(l,i) = σ(net(l,i)) = σ ( nl−1∑ j=0,wj>0 w (l,i) j y(l−1,j) + nl−1∑ j=0,wj<0 w (l,i) j y(l−1,j) ) (2.2) y(l,i) = σ(net(l,i)) = σ ( nl−1∑ j=0,wj>0 w (l,i) j y(l−1,j) + nl−1∑ j=0,wj<0 w (l,i) j y(l−1,j) ) (2.3) where y(l,i) is the lower bound of the output interval for neuron i of layer l, σ(·) is the activation function for hidden layer neurons, net(l,i) the lower bound of the activation level for neuron i of layer l, nl−1 is the number of neurons in layer l − 1, w(l,i) j is weight j for neuron i of layer l, y(l−1,j) is the lower output bound for neuron j of layer l − 1, y(l−1,j) is the upper output bound for neuron j of layer l − 1, y(l,i) is the upper output bound for neuron i of layer l, net(l,i) is the upper bound of the activation level for neuron i of layer l. the gradient of the inn training quality function is calculated as follows: ∂qk(w) ∂w (l,i) j = { s(l,i)σ′net(net (l,i))y(l−1,j) + s(l,i)σ′net(net (l,i))y(l−1,j), w (l,i) j > 0, s(l,i)σ′net(net (l,i))y(l−1,j) + s(l,i)σ′net(net (l,i))y(l−1,j), w (l,i) j < 0. (2.4) where s(l,i), s(l,i) are the lower and upper bounds of the factor determined by the recurrent procedure s(l,i) = ( nm+1∑ j=1,wj>0 s(l+1,j)σ′net(net (l+1,j))w (l+1,j) i + nm+1∑ j=1,wj<0 s(l+1,j)σ′net(net (l+1,j))w (l+1,j) i ) s(l,i) = ( nm+1∑ j=1,wj>0 s(l+1,j)σ′net(net (l+1,j))w (l+1,j) i + nm+1∑ j=1,wj<0 s(l+1,j)σ′net(net (l+1,j))w (l+1,j) i ) (2.5) with the initial condition s(l,1) = y(w)− ỹ, s(l,1) = y(w)− ỹ, where l is the number of layers in the neural network model. copyright c© 2017 assa. adv syst sci appl (2017) 44 y.e. polozova, p.v. saraev 3. proposed approach 3.1. modification of error function one of the most important advantages of interval analysis methods is the possibility to get guaranteed estimates of single-point solution in the output interval. however, function (2.1) does not guarantee that output intervals of the training examples will be fully included in the output intervals of the network. to solve this problem, a training quality function for an inn model with one output was proposed in previous research: j([w]) = k∑ i=1 qi([w]) = k∑ i=1 pibi, (3.6) where bi = [ bi1 bi2 ] = [ ỹi − yi([w]) yi([w])− ỹi ] ; pi = [pi1 pi2] , piq = { 1, biq ≥ 0, − biq n , biq < 0; q = 1, 2. here k is the size of the training data set, [w] is the interval weight vector, qi([w]) is the training error on example i , pi is the row vector of weight coefficients for the deviation of the network‘s output interval bounds from the bounds of the example i of the training data set, bi is the vector of deviation between the network‘s output interval bounds and the bounds of example i, ỹi is the lower bound of the output interval on training example i, yi([w]) is the lower bound of the network’s output interval on sample i, yi([w]) is the upper bound of the network’s output interval on example i, ỹi is the upper bound of output interval of example i, piq is element q of row vector pi, biq is is element q of row vector bi, q is the number of the vector element, n is the tolerance level showing the admissible sequence of deviations of an interval bound (e.g. 0.1, 0.01, 0.001). to train the inn model, it was offered to use the following interval adaptive algorithm of global function optimization based on weight vector bisection [6] instead of interval bp algorithm. algorithm 1 input: the minimum width δ > 0 of the bar, required training error ε > 0. output: [w] is the weight vector, which is a minimum of function (3.6); j∗ is the minimum value of function (3.6). 1. initialization of the initial weights of [w]. 2. calculation of inn output y∗ and quality function j∗. 3. initialization of list l = [w], j∗. 4. initialization of bisection coordinate l = 0 and iteration counter c = 0 (which counts iterations needed for exiting the cycle). 5. cycle: as long as minwidni=1[wi] > δ and j∗ > ε. (a) calculation of bisection coordinate l = l + 1. if l exceeds the weight vector, then l = l. (b) bisection of [w] in coordinate l into [w′] and [w′′]. (c) calculation of inn outputs y∗ and quality functions j ′ and j ′′. (d) if j ′ > j∗ and j ′′ > j∗, then go to step 5e. otherwise, go to step 5f. (e) assignment c = c+ 1. if c is not equal to the weight vector value, then go to step 5a. otherwise, exit the cycle. (f) deletion of element ([w], j∗) from l. (g) addition of ([w′], j ′) and ([w′′], j ′′) to list l. (h) arrangement of list l in ascending order by the value of the second field. (i) the first record is denoted by ([w], j∗). thus, usage of the proposed error function and learning algorithm makes it possible to guarantee that output interval contains exact solution. copyright c© 2017 assa. adv syst sci appl (2017) interval-valued data forecasting using dual-parametric neural networks 45 it is known that a result of calculating a function in interval analysis depends on the way the variable is represented in the formal expression. to obtain a more accurate value the number of variables contained in the formal expression should be minimized. in the above formulation of the problem it means that the number of layers should be minimized too. this conclusion is confirmed by experimental results of previous research. thus, minimizing ofnn layers improves accuracy of training. but a too small number of layers reduces ability of inn to approximation. this conflict was overcome by introduction of the new subclass of inn. 3.2. dual-parametric neural network dual-parametric neuron network (dpnn) is a subclass of inn that contains two types of parameters: real and interval ones. a dpnn model with a single interval output contains: • n interval output neurons; • m hidden layers; • real neuron weights in all layers (from layer 1 to layer m− 1); • interval neuron weights in layer m. besides, input and output values are also interval. in the case when input or output values should be real it is possible to represent them as intervals whose left and right bounds are equal. real weights make it possible to get high quality of learning of the neural network model. interval weights guarantees that the intervals of the training examples will be fully included in the output intervals of the network. thus, it is possible to guarantee that output intervals contains single-point solutions. for the purpose of training the dpnn, it was proposed in previous research to use an algorithm combining the interval bp method and the global optimization algorithm for interval values. 1. at the first stage, the real weights are trained using the interval bp algorithm. 2. at the second stage, the resulting weights are only used to initialize the model, followed by the training of interval weights by means of the interval global optimization algorithm. besides, a structural identification algorithm was developed for dpnn models with one hidden layer. it allows for selecting the optimal number of neurons, which are consecutively added to the hidden layer until the training error is reduced. 3.3. learning algorithm based on the intervals extension procedure as we note earlier to train the second stage model of dpnn we use the interval global optimization algorithm based on bisection of weight vector. this approach has several drawbacks. one of them is that the initial intervals of the weight vector need to have enough width. this width should be enough to include training data outputs into the model outputs for all examples in a training set when the signal passes through the network at the first time. thus, there is a difficulty in generating initial weights to train the model at the second stage of dpnn learning algorithm. the second drawback is related to the first one. according to the traditional interval global optimization algorithm an efficient choise for the bisection is cut along a coordinate with the maximum width [7]. but if the widths of all intervals is the same there is no reason to consider any of them as the most appropriate. it is possible to generate random initial intervals with different widths. but the question remains how to achieve inclusion of training data outputs into model outputs at the first iteration if this is not satisfied. to eliminate these drawbacks we introduce a learning algorithm based on the intervals extension procedure instead of the bisection. copyright c© 2017 assa. adv syst sci appl (2017) 46 y.e. polozova, p.v. saraev algorithm 2 step 1. we initialize a model with one interval output, n input interval neurons, m hidden layers and weights assigned in real values to the neurons of all layers. step 2. we train the model using the bp algorithm. step 3. we initialize a model with one interval output, n input interval neurons, m hidden layers and weights assigned in real values to the neurons of all layers (from layer 1 to layer m− 1). the neuron weights in hidden layer m are assigned in interval values. step 4. we assign real values to the weights in the model trained at step 2, and the ones in the model defined at step 3 for all hidden layers. step 5. we train the model formed at step 4 in accordance with the algorithm of training inns with interval weights. the weights of neurons in all hidden layers (from layer 1 to layer m− 1) are regarded as constant values. only the weights of hidden layer m are subject to change. step 5.1. we select the increment value m of an interval. step 5.2. a loop starts along the coordinates of the weight vector. with each iteration m substracts from the lower bound of the weight vector interval. at each iteration, the network output and the value of quality function are calculated. if the received training error value is less than the error obtained in the previous step, the weights vector and training error are fixed as minimum ones. step 5.3. step 5.2 is repeated while the training error on each loop iteration is reducing the minimum one. step 5.4. steps 5.2 and 5.3 are executed in the same way for the upper bound of weight vector intervals. in this case, m is added to the upper bound of the interval at each loop iteration. 4. numerical experiments the goal of our experiments is to show a difference between forecasting results obtained using inn with the interval bp algorithm and the proposed approach with dpnn. in our experiments, for the learning of a dpnn we generate interval-valued time series: x(i) = 0.0, 0.1, 0.2, ...; y(i) = 0.2 sin(2πx(i)) + 0.1 ( x(i) )2 + 0.2 + 0, 005rnd[−1, 1]; y(i) = 0.2 sin(2πx(i)) + 0.2 ( x(i) )2 + 0.3 + 0, 005rnd[−1, 1]. here rnd[−1, 1] is a random real number in the closed interval [−1, 1]. training set includes 35 pairs ( y(i), y(i) ) . the next 5 pairs are used as a basis for projections. we use the proposed model of dpnn with a single output neuron, 9 input neurons and one hidden layer with 3 neurons in it. to compare proposed approach and interval bp algorithm we make projections using both of them. results are shown in fig. 4.1, 4.2 and in table 4.1. table 4.1. experimental results no. of the experiment 1 2 training time, sec 28 302 training error 0.07 4.13 average relative deviation of projected value bounds from actual ones 0.15 0.12 the experiments demonstrate that projected values obtained by using the proposed approach (experiment 2) include of actual values at least 3 of 5 values. thus, in 60% of cases we can guarantee that projected value includes the exact solution. projected results interval covers just a part of actual values interval because the forecasting error accumulates with copyright c© 2017 assa. adv syst sci appl (2017) interval-valued data forecasting using dual-parametric neural networks 47 fig. 4.1. forecasting using only interval bp algorithm (experiment 1). fig. 4.2. forecasting using the proposed approach with the learning algorithm based on the intervals extension procedure (experiment 2). each subsequent projected interval. for the same reason actual values lies out of the projected interval bounds in the last values and not in the first ones. but, nevertheless, projected values obtained in experiment 2 seem to be more informative for practical use than projected results in experiment 1. training time in table 4.1 is so different, because bp algorithm used in experiment 1 is a part of learning algorithm of experiment 2. training errors is so different, because algorithms in experiment 1 and 2 uses different error functions – (2.1) and (3.6) respectively. average relative deviation of projected value bounds from actual ones is calculated as follows: δ = 1 2n n∑ i=1 ( |yi − ỹi| |yi| + |yi − ỹi| |yi| ) , where n is the number of projected values; yi, yi are the lower and the upper projected value bounds; ỹi, ỹi are the lower and the upper actual value bounds. copyright c© 2017 assa. adv syst sci appl (2017) 48 y.e. polozova, p.v. saraev 5. conclusion this study was centered on the approach to modelling and forecasting addedof interval-valued data using dpnn models. in the article we provided a theoretical overview of the technology we use, and described the proposed approach that makes it possible to get guaranteed inclusion of exact (single value) solution into interval forecasting results. our experiments confirmed the efficiency of dpnn models in predicting intervalvalued data as at least in 60% of cases we can guarantee that projected values include the exact solutions. as a prospective investigation our research will focus on training of inn and dpnn models. we take the following training data set of interval values as the initial data: {[x̃i], [ỹi]}ki=1. the inn (dpnn) will be used to compute interval functions for interval arguments in the following manner: [y] = f([w], [x]), where [w] is the vector of interval parameters (weights) of the network. besides, we are going to consider the problem of computing such inn (dpnn) weight values that allow the model outputs to include all interval outputs comprised by the training data set. one more problem to be considered is the optimization of the training quality function q([w]) = max d([yi], [ỹi])→ min with the distance between intervals d calculated as follows: d([yi], [ỹi]) = { +∞, if yi < ỹi or yi < ỹi; max { |yi − ỹi|, |yi − ỹi| } , otherwise. references 1. patil r.b. (1995) interval neural networks. apic’95, el paso, extended abstracts, a supplement to the international journal of reliable computing, el paso, tx, 164. 2. belohlavek r. (1997) backpropagation for interval patterns. neural network world, 7(3), 335–346. 3. garczarczyk z.a. (2000) interval neural networks. 2000 ieee international symposium on circuits and systems. emerging technologies for the 21st century. proceedings (ieee cat no.00ch36353), 3, 567–570. 4. kim h.j. & ryu t.-w. (2005) time series prediction using an interval arithmetic fir network. neural information processing – letters and reviews, 8 (3), 39–47. 5. yang d., li z. & wu w. (2016) extreme learning machine for interval neural networks. neural computing and applications, 27(1), 3–8. 6. sharyi s.p. (2008) randomized algorithms in interval global optimization numerical analysis and applications, 1(4), 376389. https://doi.org/10.1134/s1995423908040083 7. jaulin l., kieffer m., didrit o. & walter e. (2001) applied interval analysis. london, uk: springer. copyright c© 2017 assa. adv syst sci appl (2017) https://doi.org/10.1134/s1995423908040083 introduction theoretical overview proposed approach modification of error function dual-parametric neural network learning algorithm based on the intervals extension procedure numerical experiments conclusion adv syst sci appl 2024; 01:69–81 published online at https://ijassa.ipu.ru. combined therapeutic strategies for cancer: integrating oncolytic viruses and inhibitors in a mathematical model majda el younoussi1*, khalid hattaf1,2, noura yousfi1 1laboratory of analysis, modeling and simulation (lams), faculty of sciences ben m’sick, hassan ii university of casablanca, p.o box 7955 sidi othman, casablanca, morocco 2equipe de recherche en modélisation et enseignement des mathématiques (ermem), centre régional des métiers de l’education et de la formation (crmef), 20340 derb ghalef, casablanca, morocco abstract: the greatest cause of death worldwide continues to be cancer, a complicated set of diseases characterized by uncontrolled cell development. although early identification and therapeutic approaches have improved, the incidence of the disease is still on the rise, demanding continued study into its underlying causes and cutting-edge treatment paradigms. for creative interventions and focused medicines, the variety of cancer kinds, which are influenced by genetics, way of life, and environmental variables, poses both difficulties and opportunities. in this paper, we present a mathematical model to treat cancer with combined therapies, oncolytic viruses and mitogen-activated protein kinase inhibitors. we demonstrate that our model is both biologically and mathematically well-posed through the existence, the non-negativity and the boundedness of solutions. furthermore, we study the equilibrium points as well as the stability of these equilibria. finally, we use numerical simulations to illustrate the effect of this combined therapy on tumor cells. keywords: mapk inhibitors, oncolytic viruses, mathematical modeling, stability, hopf bifurcation. 1. introduction a type of biological therapy called oncolytic viruses is made to target and eliminate cancer cells while sparing normal cells. these viruses are produced naturally or genetically modified in a lab to target only cancer cells. the virus replicates inside the cancer cells after it has infected them, leading to the cells’ bursting and dying. this process also releases new viral particles, which can infect nearby cancer cells and continue the cycle of destruction [2]. certain oncolytic adenoviruses are notably dependent on the coxsackie-adenovirus receptor (car), and variations in car expression levels within target cells could potentially impact the efficacy of viral infection and the resulting therapeutic advantages. car has been linked to numerous facets of cancer biology, including cell adhesion, signaling, and migration, in addition to its function in promoting viral entry, making it a viable therapeutic target [9]. additionally, mitogen-activated protein kinase, also referred to as mapk or mek, is a family of serine/threonine protein kinases that are important for cellular functions like cell growth, differentiation, proliferation, and promoting car expression. mapk has the potential to exploit the complex interplay between car and oncolytic viruses to enhance their cancer-killing capabilities and stimulate immune responses against tumor cells [1]. ∗corresponding author: majda.elyounoussi-etu@etu.univh2c.ma 70 the development of our knowledge of cancer biology and treatment, on the other hand, has been greatly aided by mathematical modeling, particularly the use of ordinary differential equations (odes). the intricate dynamics and interactions between various cellular processes, signaling pathways, and therapeutic interventions in cancer can be better understood using ode models. researchers can predict how different treatment modalities will affect cancer cells and their microenvironment by creating and analyzing ode-based models. they can also find potential targets for novel therapies. further, ode models can help integrate experimental and clinical data, allowing for the quantitative assessment of cancer progression and treatment outcomes. in 2007, zurakowski and wodarz [11] used an ode model to discuss the interactions between the populations of the average level of car expression on the surface of the cells, free virus populations, susceptible, uninfected and infected tumor cells. this model was generalized by youshan and qian in their work [10], the researchers constructed a mathematical model to simulate the impacts of both mek inhibitors and viruses on tumor cells. this model is a free boundary problem, which means that it takes into account the growth and shrinkage of the tumor as the therapies are applied. the researchers used the model to explore how the combined therapies could reduce the tumor size. recently, nono et al. [7] recently expanded upon a prior model, applying it to brain cancer and introducing optimal control techniques to optimize the combination of oncolytic virotherapy and mek inhibitors. incorporating delays in mathematical models for cancer treatment can play a crucial role in capturing the realistic behavior of biological systems and improving the accuracy of the model predictions. motivated by all that, we propose in this paper an ode model to study the dynamics of oncolytic viruses and their interaction with the coxsackie-adenovirus receptor and mapk inhibitors in the context of cancer therapy, taking into account the duration required by tumor cells that have been infected to generate fresh viruses following the entry of the virus. our paper is organized into several key sections to present our research cohesively. we introduce our mathematical model in section 2, and rigorously examine its well-posedness. moving to section 3, we delve into equilibrium points and their stability, including an exploration of the hopf bifurcation, shedding light on the system’s dynamic behavior. section 4 presents the results of our numerical simulations, providing practical insights into the behavior of the model under various conditions. finally, in section 5, we offer a comprehensive conclusion, summarizing our findings, discussing their implications, and highlighting the broader significance of our research. 2. presentation and well-posedness of the model within this section, we present the subsequent ordinary differential equation (ode) model: ds dt = r(1− u)s(t)(1− s(t)+i(t) k )− βw (t)s(t)v (t) 1+αv (t) − ds(t), di dt = βw (t−τ)s(t−τ)v (t−τ)e−mτ 1+αv (t−τ) − δ(1− u)i(t)− di(t), dv dt = nδ(1− u)i(t)− βw (t)s(t)v (t) 1+αv (t) − µv (t), dw dt = ηu(γ −w (t))− hw (t), (2.1) where s(t), i(t), v (t) and w (t) are the concentration of uninfected tumor cells, infected tumor cells, free oncolytic virus particles and the average level of car molecules on cell surfaces at the time t, respectively. the factor r is the rate of tumor growth per individual in a population, slowed down by the value (1− u), where u represents the intensity of mapk inhibitor and varies between 0 and 1. if u = 1, the mapk inhibitor has the maximum possible effect. if u = 0, every cell in the first phase continue to grow and the production of car copyright © 2024 assa. adv syst sci appl (2024) 71 molecule is stopped. moreover, for biological and mathematical reasons, we will suppose as in [11] that d < r(1− u), where d is the rate of natural death of cells. the parameter k signifies the maximal tumor size. the term βsvw 1+αv models the rate of tumor cells infection by the virus in presence of car receptor and the interaction between them on the uninfected cell, where α measures the saturation effect, and β is the rate of infection process. the factor δ represents the virus induced death rate while the rate µ represents the decay of the virus. the parameter denoted as n represents the quantity of newly released viruses following the lysis of an infected tumor cell. cells generate car molecules at a rate denoted as η and experience a loss of these molecules from their surface at a rate of h. the term γ − w characterizes the saturation of car expression. furthermore, τ represents the time required for the transition from tumor cell infection to new virus production, wherem denotes the death rate for infected cells prior to virus production, and e−mτ signifies the probability of survival during the time interval [t− τ, t]. to prove that the model is mathematically well-posed and biologically meaningful, it is important to demonstrate the existence, the boundedness and the non-negativity of solutions as time evolves. let c be the set of continuous functions from the interval [−τ, 0] to r4, with the supremum norm ||ϕ|| given by sup−τ≤ζ≤0 |ϕ(ζ)|, where ϕ ∈ c. applying the fundamental theory of functional differential equations [4], we conclude that a single solution exists (s(t), i(t), v (t),w (t)), where the initial condition (s0, i0, v0,w0) are in c and we suppose that: s0(ζ) ≥ 0, i0(ζ) ≥ 0, v0(ζ) ≥ 0,w0(ζ) ≥ 0, ζ ∈ [−τ, 0]. (2.2) theorem 2.1: let’s suppose that the initial conditions fulfill (2.2). then each solution of model (2.1) stays non-negative for all t ≥ 0. proof using (2.2), we derive the following: s(t) = s(0)e ∫ t 0 (1−u)(1−s(x)+i(x) k )−βw (x)v (x) 1+αv (x) −d dx, then for all t > 0, we get s(t) ≥ 0. the second equation of model (2.1) gives i(t) = i(0)e−αt + e−mτ−αt ∫ t 0 βs(x−τ)v (x−τ)w (x−τ) 1+αv (x−τ) eαxdx, where α = δ(1− u) + d. then i(t) ≥ 0 for every t ≥ 0. from the third equation of model (2.1), we obtain v (t) = (v (0)e− ∫ t 0 βs(x) 1+αv (x) dx +nδ(1− u) ∫ t 0 i(y)eµy− ∫ t y βs(x) 1+αv (x) dxdy)e−µt. thus, v (t) ≥ 0 for every t ≥ 0. utilizing the final equation from model (2.1), we acquire: w (t) = w (0)e ∫ t 0 ηuγ w (x) −ηu−h dx ≥ 0, for all t ≥ 0. hence, every solution of model (2.1) is non-negative for all t ≥ 0. theorem 2.2: each solution of the model (2.1), given non-negative initial conditions (2.2), remains bounded for all t ≥ 0. proof by the first equation of our model, we have copyright © 2024 assa. adv syst sci appl (2024) 72 ds dt ≤ r(1− u)s(t)(1− s(t)−i(t) k ). using the comparison principal, we get lim sup t→+∞ s(t) ≤ k. therefore, s(t) is bounded. let z(t) = s(t− τ)e−mτ + i(t), hence, we have dz dt = r(1− u)s(1− s + i k )e−mτ − dse−mτ − δ(1− u)i − di ≤ r(1− u)ke−mτ − (r(1− u) + d)se−mτ − (δ(1− u) + d)i ≤ r(1− u)ke−mτ − cz(t), where c = c ′ (1− u) + d and c′ = min{r, δ}. then, lim sup t→+∞ z(t) ≤ rk(1−u)e−mτ c , and we get lim sup t→+∞ i(t) ≤ rk(1−u)e−mτ c . hence, i(t) is bounded. from the third equation we deduce dv dt ≤ nδ(1− u)i − µv, by lim sup t→+∞ i(t) ≤ rk(1−u)e−mτ c , we obtain lim sup t→+∞ v (t) ≤ nδ(1−u)2e−mτ µc . thus, v (t) is bounded. by the fourth equation of (2.1) and u ∈ [0, 1], we get dw dt ≤ ηγ − (η + h)w, then lim sup t→+∞ w (t) ≤ ηγ η+h . thus, w (t) is bounded. 3. equilibria and stability analysis within this section, we explore the three equilibrium points of model (2.1) along with their stability characteristics. copyright © 2024 assa. adv syst sci appl (2024) 73 3.1. equilibrium points of the model when there is no virus, the model (2.1) admits two infection-free equilibrium. the equilibrium point e0 = (0, 0, 0,w0), which reflects the non-existence of cells and virus, and the equilibrium point e1 = (s1, 0, 0,w1) where s1 = k(1− d r(1−u) ) and w0 = w1 = ηuγ ηu+h . by r(1− u) > d, the equilibrium e1 exists. in the existence of the virus, there exists another equilibrium point called the endemic equilibrium e∗ = (s∗, i∗, v ∗,w ∗). we suppose that i∗ > 0, v ∗ > 0 and we put r0 = βnδw1s1(1− u)e−mτ (δ(1− u) + d)(µ+ βw1s1) . r0 is the reproduction number and represents the potential for the oncolytic virus to spread within the tumor cell population. by simple calculus, we show that the equilibrium e∗ exists if r0 > 1, this means that the virus can sustain its presence and spread within the tumor cells, leading to a persistent infection. (s∗, i∗, v ∗, w ∗) are the solution of this system: r(1− u)s∗(1− s∗ + i∗ k )− βw ∗s∗v ∗ 1 + αv ∗ − ds∗ = 0, (3.3) βw ∗s∗v ∗e−mτ 1 + αv ∗ − δ(1− u)i∗ − di∗ = 0, (3.4) nδ(1− u)i∗ − βw ∗s∗v ∗ 1 + αv ∗ − µv ∗ = 0, (3.5) ηu(γ −w ∗)− hw ∗ = 0. (3.6) by equation (3.6), we get w ∗ = ηuγ ηu+ h . (3.7) by adding the equation (3.4) to the equation (3.5), we get v ∗ = δ(1− u)(n − emτ )− demτ µ i∗. (3.8) by adding the equation (3.3) to the equation (3.4), we obtain i∗ = s∗(r(1− u)(k − s∗)− dk) r(1− u)s∗ +kemτ (δ(1− u) + d) . (3.9) obviously, by determining s∗ we will determine v ∗ and i∗, for that we will use (3.7), (3.8) and (3.9) to find s∗. for simplification, we put: p = δ(1−u)(n−emτ )−demτ µ and o = nδ(1− u). by equation (3.5), we get (1− αv ∗)oi∗ − βw ∗s∗v ∗ − µ(1 + αv ∗)v ∗ = 0, then, o − βw ∗s∗p − pµ+ (αpo − µαp 2)i∗ = 0. copyright © 2024 assa. adv syst sci appl (2024) 74 using (3.9), we obtain as2 + bs + c = 0, where, a = −βw ∗pr(1− u)− αopr(1− u) + αp 2µr(1− u), b = or(1− u)− pµr(1− u)− βw ∗p ∗k(1− u)emτ (δ + d) +αopk(r(1− u)− d)− αp 2µr(1− u)k + αp 2µdk, c = (oδ(1− u)− pµδ(1− u) +od− pµd)kemτ . then s∗ = −b− √ ∆ 2a , where ∆ = b2 − 4ac. thus s∗, i∗, v ∗ and w ∗ are defined. 3.2. stability analysis the characteristic equation at any equilibrium e = (s, i, v,w ) is given by∣∣∣∣∣∣∣∣∣∣∣∣∣∣ −r(1−u)s k − λ − r(1−u)s k − βws (1+αv )2 − βsv 1+αv βwv e−(m+λ)τ 1+αv −δ(1− u)− d− λ βwse−(m+λ)τ (1+αv )2 βv se−(m+λ)τ 1+αv − βwv 1+αv nδ(1− u) − βws (1+αv )2 − µ− λ − βsv 1+αv 0 0 0 −ηu− h− λ ∣∣∣∣∣∣∣∣∣∣∣∣∣∣ = 0. (3.10) theorem 3.1: the equilibrium state e0 = (0, 0, 0,w0) exhibits instability. proof at e0 = (0, 0, 0,w0) we get the following equation: (r(1− u)− d− λ)(δ(1− u) + d+ λ)(µ+ λ)(ηu+ h+ λ) = 0, (3.11) then the roots are λ1 = r(1− u)− d, λ2 = −δ(1− u)− d, λ3 = −µ and λ4 = −ηu− h. since we have r(1− u) > d, we get λ1 > 0. thus, e0 is unstable. theorem 3.2: the equilibrium point e1 = (s1, 0, 0,w1) is locally asymptotically stable for every τ ≥ 0 if r0 < 1 and unstable if r0 > 1. proof at e1, (3.10) becomes( 2r(1− u) s1 k + d+ λ )( ηu+ h+ λ )( λ2 + (δ(1− u) + d+ βw1s1 + µ)λ +(δ(1− u) + d)(βw1s1 + µ)(1−r0e −λτ ) ) = 0. (3.12) clearly, λ1 = −2r(1− u)s1 k − d and λ2 = −ηu− h represent two of the roots of the aforementioned equation, with the remaining roots arising from solutions to the subsequent copyright © 2024 assa. adv syst sci appl (2024) 75 equation: λ2 + (δ(1− u) + d+ βw1s1 + µ)λ+ (δ(1− u) + d)(βw1s1 + µ)(1−r0e −λτ ) = 0. (3.13) if r0 > 1, let f(λ) = λ2 + (δ(1− u) + d+ βw1s1 + µ)λ+ (δ(1− u) + d)(βw1s1 + µ)(1−r0e −λτ ). we have f(0) = (δ(1− u) + d)(βw1s1 + µ)(1−r0) < 0 and lim λ→+∞ f(λ) = +∞. in this case, the equation f(λ) = 0 possesses at least one positive root. hence, if r0 > 1 e1 is unstable. if r0 < 1, we discuss two cases: when τ = 0, we get δ(1− u) + d+ βw1s1 + µ > 0 and (δ(1− u) + d)(βw1s1 + µ)(1− r0) > 0. then all the roots of (3.12) have negative real parts for τ = 0 and r0 < 1. when τ > 0, let iω be a purely imaginary root of (3.13) where ω > 0. then,{ −ω2 + (δ(1− u) + d)(βw1s1 + µ) = (δ(1− u) + d)(βw1s1 + µ)r0cos(wτ), (δ(1− u) + d+ βw1s1 + µ)ω = −(δ(1− u) + d)(βw1s1 + µ)r0sin(ωτ), thus we get ω4 + ( (δ(1− u) + d)2 + (βw1s1 + µ)2 ) ω2 + (δ(1− u) + d)2(βw1s1 + µ)2(1−r2 0) = 0. let x = ω2, then we obtain x2 + ( (δ(1− u) + d)2 + (βw1s1 + µ)2 ) x+ (δ(1− u) + d)2(βw1s1 + µ)2(1−r2 0) = 0. hence if r0 < 1, there is no positive solution. thus, e1 is locally asymptotically stable for r0 < 1. theorem 3.3: if r0 < 1, the equilibrium e1 is globally asymptotically stable. proof take into consideration the presented lyapunov function: l(t) = δ(1− u)emτi(t) + δ(1− u) + d n emτv (t) + δ(1− u) ∫ t t−τ βw (ξ)s(ξ)v (ξ) 1 + αv (ξ) dξ, then, we obtain dl dt = δ(1− u) βwsv 1 + αv − (δ(1− u) + d)emτ n ( βwsv 1 + αv + µv ) . we have v 1+αv ≤ v , lim sup t→∞ s(t) ≤ s1 and lim sup t→∞ w (t) ≤ w1. then, we get dl dt ≤ (δ(1− u) + d)(βw1s1 + µ)(r0 − 1)v n . hence, if r0 < 1 we obtain dl dt ≤ 0. clearly, dl dt = 0 if and only if s = s0, i = 0, v = 0 and w = w0. then the largest invariant set contained in {(s, i, v,w )|dl dt = 0} is the singleton {e1}. by lasalle’s invariance principale [5], we deduce that e1 is globally asymptotically stable when r0 < 1. copyright © 2024 assa. adv syst sci appl (2024) 76 the characteristic equation at the equilibrium e∗ can be expressed in the following manner: (ηu+ h+ λ) ( λ3 + p1λ 2 + p2λ+ p3 + (q1λ+ q2)e −λτ ) = 0, (3.14) where, p1 = (1− u) ( rs∗ k + δ ) + βw ∗s∗ (1 + αv ∗)2 + d+ µ, p2 = (1− u) ( βw ∗s∗ (1 + αv ∗)2 + µ )( rs∗ k + δ + d 1− u ) − β2s∗v ∗w ∗2 (1 + αv ∗)3 + r(1− u)s∗ k ( δ(1− u) + d ) , p3 = (δ(1− u) + d)s∗ ( r(1− u) k ( βs∗w ∗ (1 + αv ∗)2 + µ ) − β2v ∗w ∗2 (1 + αv ∗)3 ) , q1 = (1− u)βs∗w ∗ 1 + αv ∗ ( rv ∗ k − δn 1 + αv ∗ ) e−mτ , q2 = (1− u)βs∗w ∗ 1 + αv ∗ ( µrv ∗ k + nδ 1 + αv ∗ ( βv ∗w ∗ 1 + αv ∗ − r(1− u)s∗ k )) e−mτ . since the root λ1 = −ηu− h is negative, it remains to determine the roots of the following equation: λ3 + p1λ 2 + p2λ+ p3 + (q1λ+ q2)e −λτ = 0, (3.15) when τ = 0, the equation (3.15) becomes λ3 + p1λ 2 + (p2 + q1)λ+ p3 + q2 = 0. (3.16) we have p1 > 0, and by simple calculation we can find that p3 + q2 > 0. by routh-hurwitz criterion, we conclude the result bellow: lemma 3.4: suppose that r0 > 1 and p1(p2 + q1)− (p3 + q2) > 0. then in the absence of delay (τ = 0), all the roots of (3.15) have negative real parts. hence, the equilibriume∗ = (s∗, i∗, v ∗,w ∗) is locally asymptotically stable. when τ > 0, let iω (ω > 0) be a purely imaginary root of the equation (3.15). thus,{ p1ω 2 − p3 = q1ω sin(ωτ) + q2 cos(ωτ), −ω3 + p2ω = −q1ω cos(ωτ) + q2 sin(ωτ), (3.17) then, ω6 + ( p21 − 2p2 ) ω4 + ( p22 − q21 − 2p1p3 ) ω2 + p23 − q22 = 0. (3.18) for x = ω2, the equation (3.18) is reduced to x3 + (p21 − 2p2)x 2 + (p22 − q21 − 2p1p3)x+ p23 − q22 = 0. (3.19) we consider the following function: f(x) = x3 + c2x 2 + c1x+ c0, (3.20) where c0 = p23 − q22 , c1 = p22 − q21 − 2p1p3 and c2 = p21 − 2p2. obviously f ′(x) = 3x2 + 2c2x+ c1, and ∆′ = 4(c22 − 3c1) its discriminant. hence, we get the following result: copyright © 2024 assa. adv syst sci appl (2024) 77 lemma 3.5: (i) if c0 < 0, then the equation f(x) = 0 has at least one positive root. (ii) if c0 ≥ 0 and ∆′ ≤ 0, then the equation f(x) = 0 has no positive roots. (iii) if c0 ≥ 0 and ∆′ > 0, then the equation f(x) = 0 has a positive root if x1 > 0 and f (x1) ≤ 0, where x1 = √ c22−3c1−c2 3 is a root of f ′(x) = 0. by lemma 3.4 and the previous lemma we deduce the following theorem: theorem 3.6: assume that r0 > 1, c0 ≥ 0 and p1 (p2 + q1)− (p3 + q2) > 0. if any of the subsequent conditions are met, • ∆′ ≤ 0, • ∆′ > 0 and x1 ≤ 0, • ∆′ > 0 and f (x1) > 0, then the equilibrium pointe∗ is locally asymptotically stable for any non-negative time delay. on the other hand, we study the hopf bifurcation of model (2.1) at the equilibrium point e∗. the stability of the equilibrium point e∗ changes when the equation (3.15) has purely imaginary roots. so, we assume that ω1, ω2 and ω3 are these positive roots. moreover, we consider τ as a parameter of bifurcation. substituting ω = ωϵ and τ = τ ϵ in (3.17) where ϵ = 1, 2, 3, we get q2(p1ω 2 ϵ − p3)− q1ωϵ(−ω3 ϵ + p2ωϵ) = q22cos(ωϵτ ϵ) + q21ωϵ 2cos(ωϵτ ϵ), thus, we obtain τ ϵn = 1 ωϵ ( arccos ( q2(p1ω 2 ϵ − p3) + q1ω 2 ϵ (ω 2 ϵ − p2) q22 + q21ω 2 ϵ ) + 2πn ) , (3.21) where n ∈ n. obviously, ±iωϵ are purely imaginary roots of (3.15) with τ = τ ϵn. let τ0 = min ϵ∈{1,2,3} {τ ϵ0} and λ(τ) = ψ(τ) + iω(τ) be the root of the equation (3.15) where ψ (τ ϵn) = 0 and ω (τ ϵn) = ωϵ. differentiating equation (3.15) with respect to τ , we get( dλ dτ )−1 = 3λ2 + 2p1λ+ p2 + q1e −λτ λ (q1λ+ q2) e−λτ − τ λ . therefore, it is straightforward to deduce re ( dλ dτ )−1 ∣∣∣∣∣ τ=τϵn = 3ω4 ϵ + 2 (p21 − 2p2)ω 2 ϵ + p22 − q21 − 2p1p3 q21ω 2 ϵ + q22 = f ′ (ω2 ϵ ) q21ω 2 ϵ + q22 . by f ′ (ω2 1) > 0, f ′ (ω2 2) < 0 and f ′ (ω2 3) > 0, the transversality condition is verified, and we suppose these conditions: (a) c0 < 0, (b) c0 ≥ 0,∆′ > 0, x1 > 0 and f (x1) ≤ 0, and we get the following result: theorem 3.7: suppose r0 > 1 and p1 (p2 + q1)− (p3 + q2) > 0. if one of conditions (a)-(b) is satisfied, the equilibrium point e∗ is locally asymptotically stable for all time delays τ ∈ [0, τ0). furthermore, e∗ becomes unstable when τ > τ0. in addition, when τ = τ ϵn model (2.1) undergoes a hopf bifurcation at e∗ where ϵ = 1, 2, 3 and n ∈ n. copyright © 2024 assa. adv syst sci appl (2024) 78 4. numerical simulations this section presents numerical simulations of our system to demonstrate its dynamics and behavior. the system is evaluated with various initial conditions satisfying s0, i0, v0,w0 > 0, and the time interval is set from t = 0 to t = 2000. the following set of parameters, selected based on previous works [3, 6, 8, 11], are used: r = 0.5, u = 0.5, d = 0.1, k = 2× 109, α = 1.95× 10−10, β = 1.2× 10−10, δ = 0.5,n = 1000, µ = 20, η = 0.17, h = 0.07, γ = 7, τ = 2, m = 1. using these values, we obtain r0 = 3.8481 > 1, p1(p2 + q1)− (p3 + q2) = 11.2223 > 0, c0 = 0.1873 > 0, ∆′ = 6.7108× 105 > 0 and f(x1) = 0.1224 > 0. according to theorem 3.6, the equilibrium point e∗ is locally asymptotically stable for any τ ≥ 0. this is illustrated and validated in figure 4.1. fig. 4.1. dynamical behavior of system (2.1) around the equilibrium point e∗ when τ = 2, µ = 20 and γ = 7. we then consider the same parameter values, except that γ is changed to 2 and µ to 0.65. in this case, we obtain c0 = −9.7104× 10−4 < 0, which means that the condition of theorem 3.6 is not satisfied. therefore, the equilibrium e∗ is unstable, as shown in figure 4.2. when m = 0 and n = 100, we obtain r0 = 19.0795 > 1, p1(p2 + q1)− (p3 + q2) = 0.0054 > 0, and c0 = −9.0776× 10−4 < 0. therefore, the conditions for theorem 3.7 are satisfied. using equation (3.19) and (3.21), we obtain ω = 0.4186 and τ0 = 0.8567. in figure 4.3, when τ = 0.7 < τ0, the equilibrium point e∗ is locally asymptotically stable. however, when τ = τ0 = 0.8567, model (2.1) undergoes a hopf bifurcation at the equilibrium e∗. furthermore, when the parameter τ exceeds the critical threshold τ0 associated with the hopf bifurcation, the equilibrium at the point e∗ transitions from stable steady state to stable limit cycle oscillation, as illustrated in figure 4.4. moreover, all these results confirm our theoretical results stated in theorem 3.7 copyright © 2024 assa. adv syst sci appl (2024) 79 fig. 4.2. dynamical behavior of system (2.1) around the equilibrium point e∗ when τ = 2, µ = 0.65 and γ = 2. fig. 4.3. the stability of the equilibrium point e∗ when τ = 0.7 < τ0. copyright © 2024 assa. adv syst sci appl (2024) 80 fig. 4.4. the instability of the equilibrium point e∗ when τ = 5 > τ0. 5. conclusion this paper has presented a novel contribution by introducing an ordinary differential equation model with delay designed to comprehensively capture and analyze the intricate dynamics of tumor cells following the administration of oncolytic viruses. our model has carefully considered the role of the coxsackie adenovirus receptor (car), as well as the influence of mitogen-activated protein kinase inhibitors, factors that have played pivotal roles in the interaction between the virus and the tumor microenvironment. by incorporating these crucial elements, our research seeks to provide a deeper understanding of the underlying mechanisms governing the response of tumor cells to oncolytic virus treatment. through this innovative model, our aim has been to provide valuable information that can inform the development of more effective therapeutic strategies to combat cancer. for that, we have proven that our proposed model is mathematically and biologically meaningful through its existence, non-negativity, and boundedness of solution. we have explored the equilibrium points of the model and have identified three possible steady states. the first, e0, represents a state where there are no cells (uninfected or infected) and no virus particles. this equilibrium may not have biological relevance, as it reflects the non-existence of tumor cells and virus particles. the second, e1, describes where there are uninfected tumor cells at a constant concentration s1, but no infected tumor cells or virus particles, and the average level of car molecules on the surface of the cells is at w1. this equilibrium could represent a state where the virus fails to infect and spread in the tumor. the third, e∗, was an endemic equilibrium in which uninfected tumor cells, infected tumor cells, virus particles, and car molecules on the surface of the cells reached constant concentrations over time: s∗, i∗, v ∗, and w ∗, respectively. this equilibrium could represent a coexistence between the tumor cells, infected cells, and the virus. furthermore, an examination of the local stability of these three equilibrium points has been carried out employing the characteristic equation. moreover, the global stability of the equilibrium point e1 has been established by employing copyright © 2024 assa. adv syst sci appl (2024) 81 an appropriate lyapunov function. in addition to the stability analysis of the equilibria, we have also investigated the effects of time delay on the dynamics of the system. our analysis has revealed that the introduction of a time delay parameter can lead to a hopf bifurcation at the equilibrium point e∗, inducing a shift in equilibrium from a stable steady state to a stable limit cycle oscillation. our numerical experiments additionally have validated the theoretical findings, demonstrating the impact of time delay on the system’s behavior and the emergence of the hopf bifurcation. overall, our study has highlighted the importance of considering the time delay in modeling the dynamics of oncolytic virus therapy and has provided valuable information on the long-term behavior of the system. acknowledgements we would like to extend our appreciation to everyone who contributed their knowledge and ideas to this research, regardless of how small a way. the editors and anonymous referees are thanked by the authors for their insightful criticism and recommendations, which significantly increased the caliber of this work. references 1. cuschieri, j. & maier, r. v. (2005) mitogen-activated protein kinase (mapk), critical care medicine, 33, s417–s419. 2. everts, b. & van der poel, h. g. (2005) replication-selective oncolytic viruses in the treatment of cancer, cancer gene therapy, 12, 141–161. 3. friedman, a., tian, j. p., fulci, g., chiocca, e. a. & wang, j. (2006) glioma virotherapy: the effects of innate immune suppression and increased viral replication capacity, cancer research, 66, 2314–2319. 4. hale, j. & verduyn lunel, s.m. (1993) introduction to functional differential equations. new york, ny: springer. 5. huo, h., zhao, h. & zhu, l. (2015) the effect of vaccines on backward bifurcation in a fractional order hiv model, nonlinear anal., real world appl., 26, 289–305. 6. linsenmann, t., jawork, a., westermaier, t., homola, g., monoranu, c. m. & vince, g. h. (2019) tumor growth under rhgm-csf application in an orthotopic rodent glioma model, oncolytic letter, 17(6), 4843–4850. 7. nono, m. k. & ngouonkadi, e. m. (2020) synergistic effects of oncolytic adenovirus and mek inhibitors on glioma treatment dynamics: analysis and optimal control, applied mathematical sciences, 14(16), 781–800. 8. okamoto, k. w., priyanga, a. i. & petty, t. d. (2014) modeling oncolytic virotherapy: is complete tumor-tropism too much of a good thing?, journal of theoretical biology, 358, 166–178. 9. wunder, t., schmid, k., wicklein, d., groitl, p., dobner, t., lange, t., anders, m. & schumacher, u. (2013) expression of the coxsackie adenovirus receptor in neuroendocrine lung cancers and its implications for oncolytic adenoviral infection, cancer gene therapy, 20, 25–32. 10. youshan, t. & guo, q. (2008) a mathematical model of combined therapies against cancer using viruses and inhibitors, science in china series a: mathematics, 51(12), 2315–2329. 11. zurakowski, r. & wodarz, d. (2007) model-driven approaches for in vitro combination therapy using onyx-015 replicating oncolytic adenovirus, j theoret biol, 245, 1–8. copyright © 2024 assa. adv syst sci appl (2024) introduction presentation and well-posedness of the model equilibria and stability analysis equilibrium points of the model stability analysis numerical simulations conclusion adv syst sci appl 2018; 1; 41-58 published online at http://ijassa.ipu.ru. copyright ©2018 assa. adv. in systems science and appl. (2018) review of recent trends in coarse grain reconfigurable architectures for signal processing applications raghavachari ramya 1 , sridharan moorthi 2 1) vlsi systems research laboratory, department of electrical and electronics engineering, national institute of technology, tiruchirapalli, india e-mail:407114003@nitt.edu 2) vlsi systems research laboratory, department of electrical and electronics engineering, national institute of technology, tiruchirapalli, india e-mail: srimoorthi@nitt.edu abstract: coarse grained reconfigurable architecture got the attention of researchers working in designing computing architectures for processing massive streaming data associated with the multimedia applications in portable entertainment and communication electronics. the algorithms for processing audio, video, and graphics are very complex in nature. these data intensive computation algorithms belong to the domain of signal processing. as the complexity of algorithms increases, a matching improvement in speed performance of the hardware becomes essential to maintain the quality of service. the observed growth of algorithmic complexity is much higher than the growth rate of integration density governed by moore’s law. also, the constraints on memory bandwidth in the traditional von neumann architectures along with the slow growth in the battery capacity demands a paradigm shift in computer architecture design. reconfigurable hardware architecture is proposed as a possible alternative in this regard. the reconfigurable architectures are designed to exploit the regular and repetitive structure of signal processing algorithms and the coarse grained processing elements are designed to match with the word level granularity of these complex algorithms. the research shows that the coarse grain reconfigurable architectures with heterogeneous processing elements are a better option for system design in dsp applications, which exploit granularity matching between the algorithms and the processing hardware, and the inherent parallelism of dsp algorithms for the realization of low power dsp systems. keywords: reconfigurable architectures, cgra, dsp, fpga, asic 1. introduction the fastest growing segment of the electronic industry is the battery driven products of entertainment electronics and wireless communication systems. this growth is heavily indebted to the recent advancements in digital signal processing (dsp). but the computationally intensive dsp applications limit the battery life in portable devices such as smart phones, mp3 players, hearing aids etc., the hardware platform chosen for the implementation of mobile wireless multimedia applications decides the speed, flexibility and cost. the major attraction of the microprocessor based system development is that its ram based structure offers large scale flexibility such that products for various applications can be software based and in turn avoids the necessity of costly application specific integrated circuits (asics) for each and every application. but the sequential nature of processing by microprocessors is the major performance limiting factor. the challenging design criteria are extremely low power, high performance, flexibility and low cost. the growing performance gap among application complexity, vlsi technology, and battery technology is the 42 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) constraint in delivering satisfactory performance at low power and high speed. the flexibility of a system is also an important aspect to accommodate the rapidly changing consumer needs. limiting the applications to a specific domain is expected to provide better energy efficiency by compromising flexibility. the speed and power performance of applications pertaining to dsp domain can be further improved by suitable architecture designs which exploit the regular and repetitive structure of dsp algorithms. reconfigurable architecture is a probable architecture suitable for the said purpose. the idea of restructurable computing put forward by gerald estrin in the early sixties [1, 2].the restructurable or reconfigurable architecture is a deviation from the von neumann architecture with adjustable hardware processing elements. the general purpose processors like microprocessors have fixed structure but estrin’s proposal has provision for adjustable hardware. the adjustable hardware is called as reconfigurable fabric (rf) and the process of adjusting the hardware for different applications is called reconfiguration. there are two classes of reconfigurable architectures: fine grain and coarse grain. field programmable gate array (fpga) is a fine grain general purpose reconfigurable device that supports a broad range of applications [3, 4]. the processor architecture that uses bit oriented simple processing elements like look up table (lut) of an fpga as the fundamental building block of the processor are called fine grained. the architecture that is composed of word oriented complex logical blocks like alu is termed as coarse grain. the main disadvantage of fpga like fine grain systems when compared with asics are low or medium performance, high configuration overhead, large silicon area, propagation delay, and high power consumption. coarse grained reconfigurable processors have gained more popularity in the recent past, because they offer a new method for a dynamic and programmable execution similar to fpga and tend to achieve the performance of application specific hardware. the paradigm shift to meet the computing compulsions of real-time communication and multimedia processing is from instruction stream driven systems to data driven systems and fixed hardware systems to reconfigurable hardware systems. the coarse grained reconfigurable architecture (cgra) is a convenient platform for data streaming processing in multimedia dsp applications. these reconfigurable computing architectures have been surveyed extensively in [5-12]. the review discusses the design goals and architectural features of cgra in section 2. section 3 enumerates the architectural adaptations of cgra which makes it very suitable for the domain specific applications of dsp. section 4 focuses on the development of energy efficient coarse grain architectures for dsp applications. finally conclusions are presented in the section 5. 2. coarse grain reconfigurable architectures the coarse grain reconfigurable architectures compromise on the flexibility of fpga to match with the performance of asic by limiting themselves to a particular application domain [10, 13-16]. the performance improvement over fpga is obtained by the inherent word level configurability which is designed to match with the instruction and data granularity of algorithms results in the reduced number of cycles per instruction and hence good performance in the aspects of speed of operation. the common architecture of cgra is a combination of a control processor and a rf which is either tightly coupled to act as a co-processor or loosely coupled to act as an augmenting independent unit for processing dedicated domain specific instructions. generally the control processors execute the nonloop sequential code, control the mapping of configurations to the rf grid and supervise the execution activities. the basic architecture of cgra is shown in fig. 2.1. the rf is a combination of coarse grain processing elements (pes) or functional units (fu), word level data paths, and fast interconnects. the presence of coarse grain processing elements reduces review of recent trends in coarse grain reconfigurable architectures 43 copyright ©2018 assa. adv. in systems science and appl. (2018) the configuration data and this makes the devices to do the reconfiguration faster and reduces the area, delay, and power consumption of circuits. fig. 2.1. basic architecture of cgra the data driven processing elements are configured by configuration words stored in a dedicated memory called configuration or context memory. the context words stored in the context memory decides the functionality of each processing element and also the flow of data between pes. reconfiguration is effected by choosing another context word in context memory. the different reconfigurable options available to the designer are: reconfiguration can be designed to take place at the beginning of the execution and retains the context for the rest of the time (static reconfiguration) or it may occur at one or more times during the execution (dynamic reconfiguration). reconfiguration can affect all elements in the architecture (total reconfiguration) or only some of them (partial reconfiguration). another flexibility that the cgra architecture designers can exercise is the option to select either a homogeneous set of pes or a heterogeneous set of pes for the rf. the recent approach is to give preference to have reconfigurable system on chip which includes even general purpose processor (gpp), dsp, and fpga integrated along with cgra in to the chip [17-19]. heterogeneous architecture for cgra is preferred since, some algorithms run more efficiently on bit level configurable architecture and some perform optimal on word level reconfigurable platforms. an understanding of the algorithmic domain is essential for the design of domain specific applications in order to achieve asic like performance. the major application of cgra is in the domain of dsp. table 2.1. illustrates the important coarse grain reconfigurable architectures listed in the literature. table 2.1. different coarse grain reconfigurable architectures listed in literature year of publication name of the processor research group processing element(pe) processing element configuration data path granularity target application 1990 paddi[20] university of california, berkeley cluster of 8 exu; exu: rfile,mux, nano store crossbar 16-bit dsp 1995 kress array[21] university of kaiserslautern rdpu 2d-mesh family: select path width to implement computational datapaths 1996 rapid [22] university of washington multiplier,3 alus, 6 dpr, 3 local memories 1d-array 16-bit pipelining applications 1996 matrix [23] mit alu, multiplier, 256 × 8-bit memory, control logic 2d-mesh 8-bit, multi granular general purpose co-processor 44 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) 1997 pleiades [24 ] university of california, berkeley satellite processors crossbar multigranular multimedia 1997 raw [25] mit alu, registers, instruction & data memories, config. logic,prog.switch 2d-mesh 8-bit, multigranular general purpose 1998 piperench [26] carnegie mellon university alu & pass register file 1d-array 128-bit co-processor for streaming multimedia acceleration 1998 remarc[27] stanford university 16-bit nano processor 2d-mesh 16-bit multimedia 1998 morphosys [28] university of california, irvine alumultiplier, a shift unit, two input mux, & 32-bit context register 2d-mesh 8-bit/16bit multimedia & dsp 1999 chess array[29] hewlett packard labs interleaved alus and switchboxes hexogonal array multigranular multimedia 2000 dream[30] university of darmstadt 8-bit alu, 2 shifters,cu,cm 2d-mesh 8-bit/16-bit next generation wireless 2001 pact-xpp[31] pact informations technologie ag 32-bit alu,adder, router,cm,cu 2d-array 16-bit& 32-bit multimedia, dsp,telecommuni cation 2002 montium [32] university of twente five alu, two rfiles, sequencer tile array 16-bit streaming dsp application 2002 imagine [33] stanford university alu clusters, multi port stream register file, scratch pad memory stream model 16-bit media processing 2002 dart[34] university of rennes reconfigurable data path cluster 16-bit mobile communication 2003 trips[35] university of texas, austin 32-bit alu, routers,cu,op-reg, configuration cache 2d-array 32-bit general purpose 2003 adres[36] mec , belgium 32-bit alu, registers,router, cm,cu array 8-bit multimedia and mobile communication 2004 nec-drp[37] nec electronics 8-bit alu,8-bit dmu,16 8-bit rfu,8-bit ffu tile 8-bit streaming multimedia application 2007 mora[38] university of calabria multipliers, adders,registers, muxes,3:2 compressors linear array 8-bit multimedia 2008 rica[39] university of edinburgh heterogeneous pe array 32-bit dsp,viterbi decoding 2008 smartcell[40] worcester polytechnic institute 16-bit alu, i/o registers tile array 8-bit dsp, multimedia 2009 flora[41] seoul national university alu,rfile,mux, configuration memory 2d-array 24-bit dsp, multimedia 2009 drra[42] kth royal institute of technology, sweden morphable dpuarithmetic partition, logic partition & post processing partition 2d-array 16-bit dsp 2011 syscore[43] university college, dublin configurable functional unit systolic 24-bit biomedical signal processing 2013 bilrc[44] bilkent university, ankara, turkey alu,memory, multiplier 2d-array 16-bit multichannel fir filters,viterbi & turbo decoder 2014 fpca[45] university of california ces, lmus, alus, on-chip buffers, registers 2dmesh 32-bit dsp,medical imaging, image processing 2017 hycube [46] national university of singapore alu,memory, switches 2d-mesh 32-bit general purpose review of recent trends in coarse grain reconfigurable architectures 45 copyright ©2018 assa. adv. in systems science and appl. (2018) alu: arithmetic logic unit; cu: control unit; cm: configuration manager; ce: computation elements; config. logicconfigurable logic; dmu: data management unit; dpu: data path unit; exu: execution unit; ffu: flip flop unit; lmu: local memory unit; mux: multiplexer; op-reg: operand registers; prog.switchprogrammable switch; rfile-register file; rdpu: reconfigurable data path unit; rfu: register file unit 3. coarse grain reconfigurable architectures for dsp the paradigm shift happened in the ic design space is that it changed from a two dimensional problem of area and speed to a three dimensional problem of speed, power and complexity. conventional processing architectures such as gpp, dsp, asic and even fpga finds it difficult to satisfy the design constraint of high speed, low power performance to meet the requirements of highly complex algorithms for the current and future applications. but the industry needs processors to fill the gap among algorithmic complexity, vlsi technology and battery technology. coarse grain computing paradigm which got a boost in the past two decades is one specific answer to the above defined problem. the cgra addresses these constraints by limiting the design to domain specific applications. the larger chunk of applications addressed by cgra for high performance is from the application domain of dsp [5, 19, 47-52]. the basic concept used in the design is that hardware adapts to algorithm instead of adapting the algorithm to the hardware. the dsp algorithms are characterized by the regular and repetitive structures present in them. these regular and repetitive computational parts of dsp algorithms that accounts for large fraction of the execution time and energy consumption are called dsp kernels. these kernels are suitable for spatially distributed computing with word level processing rather than the sequential computing at a bit oriented fashion. for example the dsp kernels of fir filter contains a multiply-accumulate operation and the kernel of a fast fourier transform (fft) contains the fft butterfly which are very much suitable for word oriented processing if suitable functional units are implemented for their processing at word level in a cgra fabric. algorithms belonging to the same algorithmic domain have similar kernels and operate on similar data structures. therefore the same processing elements can be reconfigured without much control overhead for execution of different algorithms. a thorough understanding of the algorithm domain is important in the design of a power efficient reconfigurable architecture. considering the application domain of dsp, there is an excessive demand for streaming communication and computation for wireless protocol processing and multimedia processing. also, the understanding of the underlying vlsi technology plays a vital part in the design of low power systems. the dominant vlsi technology is the cmos technology with the major component of energy consumption is dynamic power consumption. a first order approximation to dynamic power consumption in cmos circuitry is given by the formula where pd is the power in watts, is the effective switching capacitance in farads, v is the supply voltage in volts, α is the switching activity factor and f is the frequency of operation in hertz. the above equation suggests that the power can be reduced by, reducing the capacitive load ceff, reducing the supply voltage v, reducing the switching frequency f, or reducing the switching activity α. since the technology already touched the 1v wall [53,54] and increase of frequency of operation also causes undesirable effects, some techniques should be exploited to reduce the switching activity and capacitance. thus it is explored to see how the concept of locality of reference is exploited to reduce the capacitance and switching activity in order to achieve better power performance in low power systems. locality of reference [19, 52, 55] is a major concept widely exploited to 46 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) reduce the power consumption. the switching activity factor α is also reduced by the locality of reference. spatial locality of reference refers to the fact that once a specified location is referenced, there is often exists a chance to refer a nearby location in the near future. accessing a large distant memory is less energy-efficient than accessing small and local memory. energy efficiency can be substantially improved by exploiting the locality of reference principle. due to the locality of reference principle communication within a tile dominates communication between tiles. the functional-level reconfiguration offers opportunities to improve energy efficiency of flexible architectures. minimizing the reconfiguration data volume reduces the energy wastage. the coarse-grain paradigm of computation optimizes storage as well as computation resources from an energy point of view. another important characteristic for data streaming applications is computing parallelism, which means the computational task has the potential to be distributed across multiple processing components. the temporal and spatial parallelism inherent in dsp algorithms can be exploited to extract better performance. in temporal parallelism, the pipeline technology is adopted at either the instruction level or at the task level, in which the instruction code or computing task is separated into multiple stages. several instructions or tasks are overlapped in the same pipeline at different stages to improve system throughput. on the other hand, spatial parallelism distributes the data and tasks onto different computational nodes to process in parallel. in data streaming applications, data-level parallelism (dlp) can often be exploited since there are no data dependencies between different input data blocks. in this case, single instruction multiple data (simd) computational style is widely used to apply the same kernel functions to different data elements. similarly, task-level parallelism (tlp) is usually exploited to execute different application threads in data streaming applications. in many cases, the computing task involved in stream processing can be decomposed into multiple stages. these stages can be overlapped into multiple computing resources to concurrently process different data sets through the pipeline. given plentiful parallelism, it is a key requirement that the computing architecture design for data streaming applications should be able to efficiently exploit and map the parallelism onto available hardware resources. as mentioned earlier, the configurability of cgra is due the reconfigurability of the functionality of the pes according to context words stored in the configuration memory. recently the cgra developers focusing on configurability of the communication network in place of the configurability of the pe. in syscore [43] a conceptual interconnect structure named round about interconnect (rai) is introduced. the interconnect configuration registers associated with the rai supports distant neighbor data transfer without the cost of area and power of commonly used mesh network interconnect or a mesh variation structure. the hycube [46] architecture goes even further with a reconfigurable interconnect to provide single cycle communication between distant neighbors. the reconfigurable interconnect network reduces the density of the interconnect network, hence reduces the power dissipation in the interconnect network. 4. energy efficient architectures for dsp the dsp algorithms are characterized by the regular and repetitive structures called kernels present in them. since the algorithms belonging to the same algorithmic domain have similar kernels and operate on similar data structures, the same processing elements can be reconfigured without much control overhead for execution of different algorithms. by executing dominant kernels of a given domain of algorithms on optimized processing elements, significant energy savings with minimum of energy overhead is achieved. this review of recent trends in coarse grain reconfigurable architectures 47 copyright ©2018 assa. adv. in systems science and appl. (2018) section deals with important energy efficient domain specific low power cgra processors intended for signal processing and multimedia applications. 4.1 montium tile processor [19, 51, 56] the computer architecture design and test for embedded systems (cadtes) group at the university of twente designed the montium tile processor and the processing core has been further developed by recore systems. this coarse grain architecture targets for 16-bit dsp applications. the montium tile is characterized by its coarse-grained reconfigurability, high performance and low energy consumption. the montium achieves flexibility through reconfigurability. the block diagram of a single montium processing tile is shown in fig. 4.1. the lower part contains the communication and configuration unit (ccu) and the upper part shows the reconfigurable tile processor (tp). the tile processor is the computational part which can be configured to implement a specific algorithm. it consists of five alus connected to 10 memory banks through a circuit switched network and these five processing parts together called the processing part array. a sequencer controls the operation of the processing part array. an array of configuration memory is provided for memory, register, alu and interconnects. the sequencer with the help of respective decoders selects appropriate context words during execution of an algorithm. the ccu implements the interface for off-tile communication. fig. 4.1. block diagram of montium processing tile [56] high performance is achieved by parallelism, because the montium has several parallel processing elements. the principle of locality of reference is exploited to reduce the energy wastage. the memory hierarchy of montium tiles allows different levels of local register, local memory, global memory and global register. the two layer decoding employed in the montium tile design simplifies the instruction fetch part, which reduces energy consumption and cost. montium is designed for dsp algorithms found extensively in mobile applications. such algorithms are usually regular and have high computational density. it has also be noted that thread-level parallelism is addressed by the multi-core approach as different tiles can run different tasks, data-level parallelism (dlp) is achieved by the montium processing tiles, which employ parallelism in the data path and instructionlevel parallelism (ilp) is addressed by the montium processing tiles as multiple data path 48 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) instructions can be executed concurrently. the architecture found to be highly flexible in mapping dsp kernels and energy efficient too. the main characteristics required for a cgra to act as a dsp core are low power, high speed and flexibility. the architecture is found to be highly flexible in mapping dsp kernels. the energy performance characteristic of the montium is compared with the characteristics of other state-of-the-art reconfigurable architectures in [56]. the fft algorithm is used to compare the energy consumption of a montium tile processor with an asic, a xilinx virtex-ii pro fpga, a silicon hive avispa and an arm920t processor. the results presented in the paper shows that montium processor provide an alternative for mobile devices for which energy-efficiency is an important factor. the montium tile processor satisfies the requirements mentioned above for a dsp core and can be used as a single dsp accelerator core or clustered in a large reconfigurable subsystem. 4.2 dynamic reconfigurable resource array (drra) [57, 58] drra is a heterogeneous coarse grain reconfigurable architecture capable of hosting multiple, complete radio and multimedia applications. the design template was developed at kth royal institute of technology, sweden. a fragment of drra fabric which consists of drra cells arranged in rows and columns fashion is illustrated in fig. 4.2. the architecture relies on high bandwidth, distributed memory integrated in 3d within the logic tile. any signal processing applications such as fir, iir, fft etc., can be mapped efficiently in the drra fabric. the drra fabric is composed of morphable data path units (mdpu) and register files (rfile) organized in a 4×n matrix. bottom and top layers are of rfiles and the inner layers are of mdpus. a sequencer, switch box, mdpu and rfile form the basic drra cell. fig. 4.2. dynamically reconfigurable resource array fabric [57] reconfigurability of mdpus and the interconnect that combine the multiple mdpus and rfiles is central to creating algorithmic building blocks. drra cells are connected together through interconnects using a sliding window 3-hop communication scheme. every resource input is connected to every resource output within three-column range, making the boundaries of the architecture flexible enough to implement most of the algorithmic building blocks. extended connectivity beyond three columns can be achieved by using intermediate review of recent trends in coarse grain reconfigurable architectures 49 copyright ©2018 assa. adv. in systems science and appl. (2018) buffers and this seamless connectivity offers flexibility in creating on the fly partition to create new algorithmic building blocks as and when required. the drra architecture is characterized by the large pool of resources for computational, storage, and interconnects functionalities. the architecture relies on clustering these resources at run time to serve individual applications and once an application is over, it is reclaimed and reused for next application. drra being a fabric, the computation is distributed across the chip. multiple threads, algorithms and applications are intended to run in parallel in the fabric. the distributed memory architecture (dimarch) found in drra is partitionable and the architecture is designed to keep the cost of partitioning and re-partitioning low both in terms of cycles and energy. the distributed nature of memory architecture and the concept of private execution environments enable a short distance between storage and computation, which in turn satisfies the locality of reference. the drra reduces the granularity mismatch at three levels. at the building block level, the dpu and the register file have been endowed with custom modes like macs, butterflies, programmable burst address generation, bit-reverse addressing modes etc., to efficiently implement commonly occurring dsp operations in the signal processing algorithms. this coarse grain mode reduces the silicon and bit-width mismatches. the drra cell is the basic unit when composing a drra fabric, but its components dpu, register file and sequencer can be combined individually to create hierarchical fsmds (fsms + data-paths) of arbitrary complexity; the combining happens in the drra interconnect fabric. this is the third level at which drra reduces the instruction granularity mismatch. drra’s interconnect scheme enables dynamic creation of an arbitrary wide vliw, where each issue can be an arbitrary instruction. the power performance of drra is computed and compared with asic, fpga in detail and presented in [57]. the performance is tested by simulating fir, fft and sorting algorithms. the static and dynamic power analysis shows that the drra achieves performance comparable with that of an asic processor. for comparison, the dynamic power dissipated in the clock net is on an average over all the algorithms is 79.47x that of asic, but for an fpga it is 255.57x. similarly the dynamic power spend for computation is 2.76x compared to 60.87x for fpga compared to that of an asic. the drra consumes 24.64x of the static power consumption of asic and fpga consumes 403.19x that of an asic. the performance of drra is evaluated further by implementing a 1024 fft. for a 1024-point fft, in terms of fft operations per unit energy, drra-1 and drra-2 outperforms all cgras by at least 2x and is worse than asic by 3.45x. the number of operations per second achievable by drra equals that of the asic. also considering the good flexibility of drra, it possesses all the essential features of a dsp core for mobile platforms. 4.3 smartcell [40, 59,60] smartcell is a coarse grain reconfigurable architecture designed for stream based applications. the block diagram of smartcell architecture is shown in fig.4.3. the architecture is made up of three main components: cell unit, reconfigurable interconnect fabric and data i/o. 50 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) fig.4.3. block diagram of smartcell architecture [60] the reconfigurable cell units are the basic components in smartcell. they are organized in a tiled structure. there are four identical pes in each cell block. multiple pes can be chained together to implement more complex algorithms. a 4-stage pipeline structure is developed in each processor. the reconfigurable interconnect fabric are used for inter and intra cell communication. the data flow can be dynamically reconfigured for different applications. the deep pipeline and ilp at single pe and the tlp among multiple cells arranged in a tiled mesh are exploited to realize high computing capacity and energy efficiency. an instruction controller is provided to control the data-path and functionality of each pe. the number of pes involved in an application task, function of each pe and the inner and inter cell connectivity are reconfigurable in real time. also, the dynamic power consumption is significantly reduced by the use of gated clock technology to turn off the inactive pes. the features of the smart cell [59] are listed below: (i) coarse-grained granularity: smartcell is designed to generate coarse-grained configurable system targeted for computation intensive applications. the processing elements operate on 16-bit input signals and generate a 36-bit output signal, which avoids high overhead and ensures better performance compared with fine-grained architectures. (ii) flexibility: due to the rich communication resources, versatile computing styles can be easily mapped onto the smart cell architecture, including simd, mimd, and 1d or 2d systolic array structures. this also expands the range of applications to be implemented. (iii) dynamic reconfiguration: by loading new instruction codes into the configuration memory through the spi structure, new operations can be executed on the desired pes without any interruption with others. the number of pes involved in the application is also adjustable for different system requirements. (iv) fault tolerance: in the smart cell system, defective cells, caused by manufacturing fault or malfunctioned circuits, can be easily turned off and isolated from the functional ones. (v) deep pipeline and parallelism: two levels of pipeline are achieved—the instruction level pipeline in a single processor element and the task level pipeline among multiple cells. the data parallelism can also be explored to concurrently execute multiple data streams, which in combine ensures a high computing capacity. review of recent trends in coarse grain reconfigurable architectures 51 copyright ©2018 assa. adv. in systems science and appl. (2018) (vi) hardware virtualization: distributed context memories are used to store the configuration signals for each pe. the cycle-by-cycle instruction execution supports hardware virtualization that is able to map large applications onto limited computing resources. (vii) smartcell provides explicit synchronization that eases the exploration of computing parallelisms. (viii) unique system topology: the cell units are tiled in a 2d mesh structure with four pes inside each cell. this topology was designed to meet different computational requirements. with the help of the hierarchical on-chip connections, the smartcell architecture can be dynamically reconfigured to perform in variant operational styles. it is experimentally shown that the smartcell offers significant improvement in energy and speed performance [60] compared to fpga, dsp and morphosys cgra. the smart cell is found to be 1.5 times faster compared to xilinx’s virtex ii pro xc2vp20 fpga in executing an fft block. it is found to be 3.6 times energy efficient than the fpga. a comparison with dsp for the same experimental setup shows that smartcell is about 20.8 times faster and about 28.9 time more energy efficient than tms320c6713. the smartcell satisfies the basic requirement of low power, high speed, and flexibility required for portable dsp processors and hence a good substitute as a dsp core. 4.4 fully pipelined composable architecture (fpca) [45] the architecture of fpca is also that of a 2d-mesh with neighbor-to-neighbor(n2n) connectivity. each tile is a cluster of pes. the architecture of the pe cluster is shown in fig.4.4. the cluster is a set of heterogeneous pes including computation elements(ces), local memory units(lmus) and register chain to act as onchip buffers and registers respectively. the architecture may be considered as a two level architecture where processing elements are first connected by a permutation matrix with high connectivity within a cluster and then by a global n2n network for more scalable connectivity. fig.4.4. internal structure of a processing element (pe) cluster [45] mostly, a single application fit into a single processing element cluster and this alleviates the challenges of routability and dynamic composition. a small part of the resources is used for a single application and hence the unused resources can shared for other applications. multiple copies of a single application can also be mapped to the unused resources. this process is called dynamic composition which will improve the overall throughput. a runtime scheduler is used to map the idle resources as mentioned to the incoming applications. the challenge of complex routing in dynamic composition is overcome by the programmable 52 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) interconnects provided with the fpca architecture. the processing element cluster approach introduces heterogeneity which support memory, arithmetic and logic operations altogether in a tile. this enables one to offload execution control to the pe cluster and hence aggressive pipelining made possible based on a modulo scheduling scheme. the full pipelining and dynamic composition of the architecture improve the energy and speed performance of the device. it is observed that fpca prototype achieve 1.5-3.4x speedup compared to the dual-core arm. also shown that a 2x improvement in speed performance is achieved by duplication based on dynamic composition. experiments show that fpca architecture can achieve a >50x energy savings compared to arm processor. 4.5 hycube [46] hycube architecture follows the 2d-mesh topology as shown in fig.4.5. the functional units (fus) are connected in a 2d-mesh structure. the architecture comprises of two types of tiles. tiles of the first type are connected to the data memory and capable of performing memory operations. tiles of the second type compute only tiles and both have an alu, a configuration memory and cross bar switch. an additional load-store unit is present in the memory tile enable them to interface with data memory. fig.4.5. a 4 × 4 hycube cgra [46] by introducing a reconfigurable interconnect hycube provides 1-cycle communication between distant tiles. the main circuit element in the programmable interconnect is the crossbar switch that is driven by a clockless repeater. the repeaters can be configured to bypass data to successive hops asynchronously or to receive data. since data can be passed to distant tiles by bypassing the entry of data into the in between tiles, 1-cycle communication can be realized between tiles that are not nearest neighbors. the hycube has a compiler controlled noc and hence there is no routing and flow control logic. the input and output ports has only single registers and facilitates multihop multicast path between tiles scheduled completely at compile time. given the reconfigurable single-cycle multi-hop interconnect, hycube scales very well for 4×4 cgra and offers 3x performance and 4x performance-per-watt compared to arm. the experimental results also show that average power efficiency of hycube is 1.5x and 3x compared to a cgra with standard noc and a n2n cgra, respectively. the experiments gives a demonstration of the improvements made with regard to the interconnect of the hycube enables it to improve the energy performance. unlike the other processors discussed hycube is not a dedicated dsp cgra architecture. but the idea of reconfigurable interconnect and reduction of total hardware cost due to interconnects are also seems to be welcome features for the design of future cgras for improving power performance as demonstrated by the hycube architecture. review of recent trends in coarse grain reconfigurable architectures 53 copyright ©2018 assa. adv. in systems science and appl. (2018) 5. conclusions in the post dennard scaling era, the supply voltage scaling is limited by exponential increase in leakage current and it touched the 1v wall. the computer architecture designers are thus exploring alternatives to circumvent the limitations imposed by the vlsi technology. the speed and power performance can be improved by exploiting the inherent word level parallelism of dsp algorithms to design suitable parallel computer architectures matching with the structure and granularity of the algorithms. another concept well exploited in the design of low power dsp system is the judicious use of locality of reference by providing distributed memory. the flexibility of cgra were restricted by limiting the applications to a specific domain and suitable control mechanism are provided to turn off unutilized hardware in run-time to get better energy efficiency. also heterogeneous architectures ensures better granularity matching which can be exploited for realizing low operational frequencies which is directly proportional to the power dissipated. the study reveals that coarse grained reconfigurable heterogeneous architecture with distributed memory is the future solution for the electronic industry to assure speed, power and cost performance. recent research shows that reconfigurable interconnect networks reduces the density of the interconnect network, and hence reduces the power dissipation in the interconnect network. a clever design choice of reconfigurable processing elements and interconnects may result in delivering asic like energy and speed performance by cgras especially in the application domain of signal processing. references [1] estrin, g. (1960). organization of computer systems: the fixed-plus-variable structure computer. in papers presented at the may 3-5, 1960, western joint ireaiee-acm computer conference, new york, usa, 33-40, https://doi.org/10.1145/1460361.1460365 [2] estrin, g., bussell, b., turn, r., & bibb, j. (1963). parallel processing in a restructurable computer system. ieee transactions on electronic computers, ec12(6), 747-755, https://doi.org/10.1109/pgec.1963.263558 [3] trimberger, s. m. (2015). three ages of fpgas: a retrospective on the first thirty years of fpga technology. in proceedings of the ieee, 103(3), 318-331, https://doi.org/10.1109/jproc.2015.2392104 [4] tatas, k., siozios, k., & soudris, d. (2007). a survey of existing fine-grain reconfigurable architectures and cad tools. in s. vassiliades and d. soudris (eds), fine-and coarse-grain reconfigurable computing (pp. 3-87). springer netherlands. [5] hartenstein, r. (2001). a decade of reconfigurable computing: a visionary retrospective. in proceedings of design, automation and test in europe conference and exhibition 2001, munich, germany, 642-649, https://doi.org/10.1109/date.2001.915091 [6] krishnamurthy, r. b. (2001). a survey of next generation reconfigurable architectures for embedded computing, tech.rep., college of computing, georgia institute of technology. [7] compton, k., & hauck, s. (2002). reconfigurable computing: a survey of systems and software. acm computing surveys (csur), 34(2), 171-210, https://doi.org/ 10.1145/508352.508353 https://doi.org/10.1145/1460361.1460365 https://doi.org/10.1109/pgec.1963.263558 https://doi.org/10.1109/jproc.2015.2392104 https://doi.org/10.1145/508353 https://doi.org/10.1145/508353 54 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) [8] george, t., soudris, d., & vassiliadis, s. (2007). a survey of coarse-grain reconfigurable architectures and cad tools. in s. vassiliades and d. soudris (eds), fine-and coarse-grain reconfigurable computing (pp.89-149). springer netherlands. [9] hauck, s., & dehon, a. (2010). reconfigurable computing: the theory and practice of fpga-based computation, morgan kaufmann, amsterdam. [10] de sutter, b., raghavan, p., & lambrechts, a. (2013). coarse-grained reconfigurable array architectures. in handbook of signal processing systems (pp. 553-592). springer new york. [11] lyke, j. c., christodoulou, c. g., vera, g. a., & edwards, a. h. (2015). an introduction to reconfigurable systems. in proceedings of the ieee, 103(3), 291-317, https://doi.org/ 10.1109/jproc.2015.2397832 [12] chattopadhyay, a. (2013). ingredients of adaptability: a survey of reconfigurable processors. vlsi design, 2013,1-18, http://dx.doi.org/10.1155/2013/683615 [13] eguro, k., & hauck, s. (2003). issues and approaches to coarse-grain reconfigurable architecture development. in proceedings of the 11th annual ieee symposium on field-programmable custom computing machines,(fccm 2003), napa, ca, usa, 111-120, https://doi.org/10.1109/fpga.2003.1227247 [14] eguro, k., & hauck, s. a. (2005). resource allocation for coarse-grain fpga development. ieee transactions on computer-aided design of integrated circuits and systems, 24(10), 1572-1581, https://doi.org/10.1109/tcad.2005.852291 [15] verbauwhede, i., schaumont, p., piguet, c., & kienhuis, b. (2004). architectures and design techniques for energy efficient embedded dsp and multimedia processing. in proceedings of the design, automation and test in europe conference and exhibition, 2, 988-993, https://doi.org/10.1109/date.2004.1269022 [16] carroll, a., friedman, s., van essen, b., wood, a., ylvisaker, b., ebeling, c., & hauck, s. (2007). designing a coarse-grained reconfigurable architecture for power efficiency. in department of energy na-22 university information technical interchange review meeting. [17] smit, g. j., kokkeler, a. b., wolkotte, p. t., hölzenspies, p. k., van de burgwal, m. d., & heysters, p. m. (2007). the chameleon architecture for streaming dsp applications. eurasip journal on embedded systems, 2007:078082 https://doi.org/10.1155/2007/78082 [18] park, y., park, j. j. k., & mahlke, s. (2012). efficient performance scaling of future cgras for mobile applications. in 2012 international conference on field programmable technology(fpt’12), seoul, south korea, 335-342, https://doi.org/10.1109/fpt.2012.6412158 [19] smit, g. j., kokkeler, a. b., wolkotte, p. t., van de burgwal, m. d., & heysters, p. m. (2006). efficient architectures for streaming dsp applications. dynamically reconfigurable architectures, internationales begegnungs-und forschungszentrum fuer informatik (ibfi), schloss dagstuhl, germany. [online]. available http://drops.dagstuhl.de/opus/volltexte/2006/743 [20] chen, d., & rabaey, j. (1990). paddi: programmable arithmetic devices for digital signal processing. in vlsi signal processing iv, 240-249. https://doi.org/10.1109/jproc.2015.2397832 https://doi.org/10.1109/fpga.2003.1227247 https://doi.org/10.1109/date.2004.1269022 https://doi.org/10.1155/2007/78082 http://drops.dagstuhl.de/opus/volltexte/2006/743 review of recent trends in coarse grain reconfigurable architectures 55 copyright ©2018 assa. adv. in systems science and appl. (2018) [21] hartenstein, r. w., & kress, r. (1995). a datapath synthesis system for the reconfigurable datapath architecture. in proceedings of asia and south pacific design automation conference 1995, makuhari, chiba, japan, 479-484, https://doi.org/10.1109/aspdac.1995.486359 [22] ebeling c., cronquist d.c. & franklin p. (1996) rapidreconfigurable pipelined datapth. in: r.w. hartenstein, m. glesner (eds) field –programmable logic smart applications, new paradigms and compilers. fpl 1996. lecture notes in computer science, (pp. 126-135). springer, berlin, heidelberg, https://doi.org/10.10073/3-54061730-2_13 [23] mirsky, e. & dehon, a. (1996) matrix: a reconfigurable computing architecture with configurable instruction distribution and deployable resources. in proceedings of ieee symposium on fpgas for custom computing machines (fccm), napa valley, ca, usa, 157-166, https://doi.org/ 10.1109/fpga.1996.564808 [24] rabaey, j. m. (1997). reconfigurable processing: the solution to low-power programmable dsp. in 1997 ieee international conference on acoustics, speech, and signal processing, munich, germany, 1, 275-278, https://doi.org/10.1109/icassp.1997.599622 [25] waingold, e., taylor, m., srikrishna, d., sarkar, v., lee, w., lee, v., & babb, j. (1997). baring it all to software: raw machines. computer, 30(9), 86-93, https://doi.org/10.1109/2.612254 [26] copen, s., herman, g., matthew, s., mihai, m., cadambi, b. s., reed, r., & laufer,t.r(1999). piperench: a coprocessor for streaming multimedia acceleration. in proceedings of the 26 th international symposium on computer architecture, atlanta, ga, usa, 28–39, https://doi.org/10.1109/isca 1999.765934 [27] miyamori, t., & olukotun, k. (1999). remarc: reconfigurable multimedia array coprocessor. ieice transactions on information and systems, 82(2), 389-397. [28] singh, h., lee, m. h., lu, g., kurdahi, f. j., bagherzadeh, n., lang, t., & eliseu filho, m. c. (1998). morphosys: an integrated re-configurable architecture. in proceedings of the nato rto symp. on system concepts and integration, monterey, ca, usa , 1-11. [29] marshall, a., stansfield, t., kostarnov, i., vuillemin, j., & hutchings, b. (1999). a reconfigurable arithmetic array for multimedia applications. in proceedings of the 1999 acm/sigda seventh international symposium on field programmable gate arrays , monterey, california, usa, 135-143, https://doi.org/10.1145/296399.296444 [30] alsolaim, a., becker, j., glesner, m., & starzyk, j. (2000). architecture and application of a dynamically reconfigurable hardware array for future mobile communication systems. in proceedings of 2000 ieee symposium on fieldprogrammable custom computing machines, napa valley, ca, usa, 205-214, https://doi.org/10.1109/fpga.2000.903407 [31] baumgarte, v., ehlers, g., may, f., nückel, a., vorbach, m., & weinhardt, m. (2003). pact xpp-a self-reconfigurable data processing architecture. the journal of supercomputing, 26(2), 167-184, https://doi.org/10.1023/a:1024499601471 https://doi.org/10.1109/fpga.1996.564808 https://doi.org/10.1109/icassp.1997.599622 https://doi.org/10.1145/296399.296444 https://doi.org/10.1109/fpga.2000.903407 56 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) [32] heysters, p., smit, g., & molenkamp, e. (2003). a flexible and energy-efficient coarse-grained reconfigurable architecture for mobile systems. the journal of supercomputing, 26(3), 283-308, https://doi.org/10.1023/1025699015398 [33] kapasi, u. j., dally, w. j., rixner, s., owens, j. d., & khailany, b. (2002). the imagine stream processor. in proceedings of the 2002 ieee international conference on computer design: vlsi in computers and processors, freiberg, germany, 282-288, https://doi.org/10.1109/iccd.2002.1106783 [34] david, r., chillet, d., pillement, s., & sentieys, o. (2002). dart: a dynamically reconfigurable architecture dealing with future mobile telecommunications constraints. in proceedings of the 16 th international parallel and distributed processing symposium (ipdps 2002), fort lauderdale, fl,usa, https://doi.org/ 10.1109/ipdps.2002.1016554 [35] sankaralingam, k., nagarajan, r., mcdonald, r., desikan, r., drolia, s., govindan, m. s., ... & liu, h. (2006). distributed microarchitectural protocols in the trips prototype processor. in proceedings of the 39 th annual ieee/acm international symposium on microarchitecture, orlando,fl,usa, 480-491, https://doi.org/10.1109/micro.2006.19 [36] mei b., vernalde s., verkest d., de man h., lauwereins r. (2003) adres: an architecture with tightly coupled vliw processor and coarse-grained reconfigurable matrix. in: y. k. cheung p., constantinides g.a. (eds) field programmable logic and application. fpl 2003. lecture notes in computer science, (pp. 61-70). springer, berlin, heidelberg, https://doi.org/10.1007/978-3540-45234-8_7 [37] suzuki, m., hasegawa, y., yamada, y., kaneko, n., deguchi, k., amano, h., ... & awashima, t. (2004). stream applications on the dynamically reconfigurable processor. in proceedings of the 2004 ieee international conference on fieldprogrammable technology, brisbane, nsw, australia, 137-144, https://doi.org/ 10.1109/fpt.2004.1393261 [38] lanuzza, m., perri, s., corsonello, p., & margala, m. (2007). a new reconfigurable coarse-grain architecture for multimedia applications. in 2007 nasa/esa conference on adaptive hardware and systems, edinburgh, uk, 119-126, https://doi.org/10.1109/ahs.2007.10 [39] khawam, s., nousias, i., milward, m., yi, y., muir, `m., & arslan, t. (2008). the reconfigurable instruction cell array. ieee transactions on very large scale integration (vlsi) systems, 16(1), 75-85, https://doi.org/ 10.1109/tvlsi.2007.912133 [40] liang, c., & huang, x. (2008). smartcell: a power-efficient reconfigurable architecture for data streaming applications. in ieee workshop on signal processing systems, sips 2008, washington, dc, usa, 257-262, https://doi.org/ 10.1109/sips.2008.4671772 [41] lee, d., jo, m., han, k., & choi, k. (2009). flora: coarse-grained reconfigurable architecture with floating-point operation capability. in proceedings of the 2009 international conference on field-programmable technology( fpt 2009), sydney, nsw, australia, 376-379, https://doi.org/ 10.1109/fpt.2009.5377609 [42] shami, m. a., & hemani, a. (2009). morphable dpu: smart and efficient data path for signal processing applications. in proceedings of ieee workshop on signal https://doi.org/10.1109/iccd.2002.1106783 https://doi.org/10.1109/ipdps.2002.1016554 https://doi.org/10.1109/micro.2006.19 https://doi.org/10.1007/978-3-540-45234-8_7_ https://doi.org/10.1007/978-3-540-45234-8_7_ https://doi.org/10.1109/fpt.2004.1393261 https://doi.org/10.1109/ahs.2007.10 https://doi.org/10.1109/tvlsi.2007.912133 https://doi.org/10.1109/sips.2008.4671772 https://doi.org/10.1109/fpt.2009.5377609 review of recent trends in coarse grain reconfigurable architectures 57 copyright ©2018 assa. adv. in systems science and appl. (2018) processing systems, tampere, finland,167-172 , https://doi.org/10.1109/sips.2009.5336246 [43] patel, k., mcgettrick, s., & bleakley, c. j. (2011). syscore: a coarse grained reconfigurable array architecture for low energy biosignal processing. in 2011 ieee 19th annual international symposium on field-programmable custom computing machines(fccm),salt lake city, ut, usa, 109-112, https://doi.org/ 10.1109/fccm.2011.38 [44] atak, o., & atalar, a. (2013). bilrc: an execution triggered coarse grained reconfigurable architecture. ieee transactions on very large scale integration (vlsi) systems, 21(7), 1285-1298, https://doi.org/ 10.1109/tvlsi.2012.2207748 [45] cong, j., huang, h., ma, c. , xiao, b., & zhou, p. (2014). a fully pipelined and dynamically composable architecture of cgra. in 2014 ieee 22nd annual international symposium on field-programmable custom computing machines (fccm), boston, ma, usa, 9-16, https://doi.org/ 10.1109/fccm.2014.12 [46] karunaratne, m., mohite, a. k., mitra, t., & peh, l. s. (2017). hycube: a cgra with reconfigurable single-cycle multi-hop interconnect. in proceedings of the 54 th annual design automation conference 2017 , austin, tx, usa, article no.45, https:/doi.org/10.1145/3061639.3062262 [47] tessier, r., & burleson, w. (2001). reconfigurable computing for digital signal processing: a survey. the journal of vlsi signal processing systems for signal, image and video technology, 28(1-2), 7-27, https://doi.org/10.1023/a:1008155020711 [48] zhang, c., lenart, t., svensson, h., & öwall, v. (2009). design of coarse-grained dynamically reconfigurable architecture for dsp applications. in 2009 international conference on reconfigurable computing and fpgas,reconfig'09, quintana roo, mexico 338-343, https://doi.org/ 10.1109/reconfig.2009.49 [49] todman, t. j., constantinides, g. a., wilton, s. j., mencer, o., luk, w., & cheung, p. y. (2005). reconfigurable computing: architectures and design methods. iee proceedings-computers and digital techniques, 152(2), 193-207, https://doi.org/ 10.1049/ip-cdt:20045086 [50] abnous, a. (2001). low-power domain-specific architectures for digital signal processing (doctoral dissertation), university of california, berkeley, ca, [online] . available http://vada.skku.ac.kr/classinfo/dsp/sdr/thesis.pdf [51] heysters, p. m., & smit, g. j. (2003). mapping of dsp algorithms on the montium architecture. in proceedings of international parallel and distributed processing symposium, nice, france, p180.2, https://doi.org/10.1109/ipdps.2003.1213333 [52] galanis, m. d., dimitroulakos, g., & goutis, c. e. (2006). mapping dsp applications on processor/coarse-grain reconfigurable array architectures. in proceedings of 2006 ieee international symposium on circuits and systems (iscas), island of kos, greece, 3666-3669, https://doi.org/ 10.1109/iscas.2006.1693422 [53] itoh, k., yamaoka, m., & oshima, t. (2010). adaptive circuits for the 0.5-v nanoscale cmos era, ieice transactions on electronics, 93(3), 216-233, https://doi.org/10.1587/transele.e93.c.216 https://doi.org/10.1109/sips.2009.5336246 https://doi.org/10.1109/fccm.2011.38 https://doi.org/10.1109/tvlsi.2012.2207748 https://doi.org/10.1109/fccm.2014.12 https://doi.org/10.1145/3061639.3062262 https://doi.org/10.1109/reconfig.2009.49 https://doi.org/10.1049/ip-cdt:20045086 https://doi.org/10.1109/iscas.2006.1693422 https://doi.org/10.1587/transele.e93.c.216 58 r. ramya, s. moorthi copyright ©2018 assa. adv. in systems science and appl. (2018) [54] nowak, e. j. (2002). maintaining the benefits of cmos scaling when scaling bogs down. ibm journal of research and development, 46(2.3),169-180, https://doi.org/ 10.1147/rd.462.0169 [55] guo, y. (2006). mapping applications to a coarse-grained reconfigurable architecture (doctoral dissertation), university of twente, eindhoven, netherlands. [56] heysters, p. m., smit, g. j., & molenkamp, e. (2004). energy-efficiency of the montium reconfigurable tile processor. in proceedings of the international conference on engineering of reconfigurable systems and algorithms (ersa '04), las vegas, nev, usa, 38-44. [57] shami, m. a. (2012). dynamically reconfigurable resource array (doctoral thesis) stockholm: kth royal institute of technology, sweden. [58] tajammul, m. a., shami, m. a., hemani, a., & moorthi, s. (2011). noc based distributed partitionable memory system for a coarse grain reconfigurable architecture. in proceedings of the 2011 24 th international conference on vlsi design (vlsi design), chennai, india, 232-237, https://doi.org/ 10.1109/vlsid.2011.45 [59] liang, c., & huang, x. (2009). smartcell: an energy efficient coarse-grained reconfigurable architecture for stream-based applications. eurasip journal on embedded systems, 2009:518659, https://doi.org/10.1155/2009/518659 [60] liang, c. (2009). smartcell: an energy efficient reconfigurable architecture for stream processing (doctoral dissertation), wpi, worcester, ma. https://doi.org/10.1147/rd.462.0169 https://doi.org/10.1109/vlsid.2011.45 adv syst sci appl 2020; 01:119–127 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/866 local normal forms of autonomous quasi-linear constrained differential systems alexander m. kotyukov1*, stanislav o. nikanorov1, natalya g. pavlova1,2 1v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia 2moscow institute of physics and technology (state university), dolgoprudnyi, russia abstract: the paper presents a study of impasse (singular) points of autonomous quasi-linear constrained differential systems, also called differential-algebraic equations. the interest in such systems is motivated by their applications in various problems of pure and applied mathematics, including control theory, biology, and electric engineering. local normal forms of such systems in a neighborhood of their impasse points are established. keywords: impasse point, direction field, normal form, diffeomorphism, symmetry 1. introduction we consider systems of differential equations in the form a(ξ)ξ′ = b(ξ), ξ′ = dξ dt , ξ = (ξ1, . . . , ξn)∗ ∈ rn, (1.1) where a = (aij) is an n× n matrix and b = (b1, . . . , bn)∗ is a vector function. the components aij , bj are assumed to be c∞-smooth functions on ξ. there are many possible names of systems (1.1): differential systems of sobolev type, generalized vector fields, descriptor systems, quasi-linear constrained† differential systems, etc. the interest in systems (1.1) is motivated by their applications in various problems of pure and applied mathematics (including control theory and electric engineering). for instance, systems (1.1) describe the dynamic of electric circuit in nonlinear rlc-networks (networks consisting of a resistor, an inductor, and a capacitor). see [1]– [14] and the references therein. in a pioneer work [15], systems (1.1) appear as a tool for studying a systems of pdes of the mixed type, which describes the motion of a body filled with a viscous incompressible fluid. in a recent series of papers [16]– [20], systems (1.1) in dimension n = 3 describe geodesic lines in singular metrics. there exists several different approaches for studying systems (1.1) and, more generally, nonlinear systems of differential equations f (t, ξ, ξ′) = 0, ξ′ = dξ dt , ξ = (ξ1, . . . , ξn)∗ ∈ rn, (1.2) ∗corresponding author: amkotyukov@mail.ru †the reason for the this term will soon become clear. 120 a.m. kotyukov, s.o. nikanorov, n.g. pavlova which are often called implicit differential equations or differential-algebraic equations (daes). here f : r2n+1 → rn is assumed to be a c∞-smooth mapping. there exists several approaches to investigation of systems (1.1), (1.2). an analytical approach is based on the decoupling procedure. the idea is to rearrange terms within the given system so that it is decomposed in two subsystems of lower dimensions (separated as far as possible), where the first subsystem is equivalent to a standard ode of of maximum possible dimension and the second one is a dae having a special form. the highest goal is so-called complete decoupling, where the both subsystems are completely independent, that is, they do not have common unknowns. for linear daes with constant coefficients this method is comparable with the weierstrass–kronecker normal form of regular matrix pencils. this approach is presented in the book [21]. another geometric approach goes back to h. poincaré.‡ he used it for a single implicit differential equation f (x, y, y′) = 0, that is, the case n = 1. further development of this approach is given, e.g., in [22]– [34]. in the modern terminology, the main idea is the following. by jk denote the space of k-jets of vector functions ξ(t) : r→ rn. in the case of systems (1.2), consider a manifold mn+1 f ⊂ j1, j1 ' r2n+1, defined by the equation f = 0. moreover, dae (1.2) defines a direction field on mn+1 f , whose integral curves are 1-graphs (legendrian lifts) of integral curves of (1.2). in other words, we pass from the multivalued vector direction field given by dae (1.2) in j0 (i.e., the (t, ξ)-space) to a single-valued direction field on the manifold mn+1 f . the natural projection π : j1 → j0 sends integral curves of the field defined on mn+1 f to integral curves of dae (1.2). the restriction of the projection π : mn+1 f → j0 is a mapping of two (n+ 1)-dimensional manifolds. it has singularities at those points of mn+1 f where the matrix (∂f/∂ξ′) vanishes. integral curves of (1.2) have singularities (generically, cusps) at corresponding point of the (t, ξ)-space. in the case of systems (1.1), the above construction is essentially simplified: no need to consider the space j1. indeed, writing system (1.1) in the pfaffian form a11dξ1 + a12dξ2 + · · ·+ a1ndξn − b1dt = 0, . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . an1dξ1 + an2dξ2 + · · ·+ anndξn − bndt = 0, (1.3) one can see that system (1.1) defines a direction field in the (t, ξ)-space: ξ̇1 = ∆1(ξ), . . . , ξ̇n = ∆n(ξ), ṫ = ∆(ξ), (1.4) where ∆ is the determinant of the matrix a, and ∆i is the determinant of the matrix obtained from a by replacing of its ith column with b. here the dot over a symbol means the differentiation by a new parameter playing the role of time. a principal difference between system (1.1) and usual autonomous odes is that system (1.1) (and consequently, field (1.4)) possesses so-called degenerate hypersurface γ = {ξ : ∆(ξ) = 0}, which is also called the criminant of the system (1.1). generically, there are no integral curves of (1.1) passing through γ; see [29]. therefore, γ is also called impasse hypersurface of the system, and points of γ are called impasse points. it is worth observing that if system (1.1) describes a model of a natural process in real time, the existence of impasse points in the applicability domain often means that the model is not adequate. ‡mémoire sur les courbes définies par les équations différentielles (1885). copyright © 2020 assa. adv syst sci appl (2020) local normal forms of differential systems 121 the germs of two systems (1.1) are called ck-equivalent, if there exists a local ckdiffeomorphism of the ξ-space that conjugates the germs of the corresponding direction fields (1.3). our goal is to simplify field (1.3) (and consequently, the corresponding system (1.1)) in a neighborhood of its impasse points using an appropriate local ck-diffeomorphism of the ξ-space. here k ≥ 1 is integer or∞ or ω. cω is the standard symbol for the class of analytic maps.§ the germs of two systems (1.1) are called orbitally ck-equivalent, if there exist a local ck-diffeomorphism of the ξ-space and ck-mapping of the independent variable t 7→ τ(t, ξ) that conjugate the germs of the corresponding direction fields (1.3). the classification of systems (1.1) at all impasse points is not observable. according to the general ideology of singularity theory, a number of classes of impasse points that admits relatively simple normal forms is determined. for the cω-equivalence, many results of this sort appear in [28], where systems (1.1) are considered in real, complex and even infinite dimensional (banach) complex spaces. for the ck-orbital equivalence, many results appear in [32]. in this paper, we present some new results aboutck-normal forms of the germs of systems (1.1) at their impasse points of a special type. 2. main results there are several geometric objects naturally connected with system (1.1). the firts object of thid sort is the criminant γ. other important object are the family of linear operators a(ξ) : tξrn → tξrn, their images ima(ξ) and kernels kera(ξ). the criminant γ is the locus of points ξ where dim ima(ξ) < n or, equivalently, dim kera(ξ) > 0. the simplest type of impasse points is so-called non-singular impasse points ξ ∈ γ that satisfy three following conditions: 1. d∆(ξ) 6= 0, that is, γ is a regular hypersurface. 2. dim kera(ξ) = 1 and the direction kera(ξ) is transversal to γ. 3. b(ξ) /∈ ima(ξ), that is, ima(ξ)⊕ 〈b(ξ)〉 = rn. theorem 2.1: the germ of system (1.1) at every its non-singular impasse point is c∞-equivalent (in the analytic category, cω-equivalent) to ξ′1 = 0, . . . , ξ′n−1 = 0, ξn ξ ′ n = ±1. moreover, in the category of orbital equivalence, one can reduce ±1 to 1. the proof of theorem 2.1 can be found in [27, 28, 32]. now consider a point ξ◦ ∈ γ such that the 1st and 2nd conditions hold true, but the 3rd condition fails. from dim kera(ξ◦) = 1 (the 2nd condition) and the known identity dim kera(ξ) + dim ima(ξ) = n it follows that dim ima(ξ◦) = n− 1, i.e., rank of the matrix a(ξ◦) = n− 1. then there exists i ∈ {1, . . . , n} such that rank of the matrix obtained from a(ξ◦) by eliminating the ith column is n− 1, and the germs of ∆, ∆1, . . . ,∆n belong to the ideal i = 〈∆,∆i〉 in the ring of smooth functions, and the set of singular points (equilibriums) of the corresponding §in speaking about cω-equivalence, we assume that the systems (1.1) are also analytic. copyright © 2020 assa. adv syst sci appl (2020) 122 a.m. kotyukov, s.o. nikanorov, n.g. pavlova field (1.4) is given by two equations: ∆ = ∆i = 0. (the detailed proof of this fact can be found in [31].) therefore, singular point of the field (1.4) are not isolated, they fill a submanifold w ⊂ γ of codimension 2. at every point ξ ∈ w in a neighborhood of ξ◦, the spectrum of the linear part of the field (1.4) is (λ1, λ2, 0, . . . , 0). here λ1,2 are complex numbers, generically reλ1,2 6= 0 at almost all points of w . lemma 2.1: assume that reλ1,2(ξ◦) 6= 0. then in a neighborhood of ξ◦, the following statements hold true: (i) the set w is the center manifold of the field (1.4). (ii) there exist i1, i2 ∈ {1, . . . , n} such that the ideal i = 〈∆i1 ,∆i2〉. proof the first statement is trivial. for the second statement, remark that it is possible to select as generators of i any two elements v, w ∈ i such that dv, dw at the point ξ◦ are linearly independent. consider the jacobi matrices j0 = ∥∥∥∥∥∂(∆1, . . . ,∆n,∆) ∂(ξ1, . . . , ξn, t) ∥∥∥∥∥, j = ∥∥∥∥∥∂(∆1, . . . ,∆n) ∂(ξ1, . . . , ξn) ∥∥∥∥∥. at the point ξ◦. since the functions ∆1, . . . ,∆n,∆ do not depend on t, the spectrums of matrices j0 and j have the form (λ1, λ2, 0, . . . , 0) with the same λ1,2. therefore, among the functions ∆1, . . . ,∆n there exist a couple v = ∆i1 , w = ∆i2 whose differentials at ξ◦ are linearly independent. therefore, v, w are generators of the ideal i . lemma 2.1 shows that using an appropriated renaming of the variables (ξ1, . . . , ξn) 7→ (x, y, z), z = (z1, . . . , zn−2), one can write the corresponding field (1.4) in the form ẋ = v, ẏ = w, żi = αiv + βiw, i = 1, . . . , n− 2, ṫ = av + bw, (2.5) here a, b, α, β, v, w are c∞ (resp., cω) functions on x, y, z such that v, w vanish at ξ◦. the ideal i = 〈v, w〉, and the center manifold w consisting of singular points is given by the equations v = w = 0. since all components of the field (2.5) do not depend on t, one can consider the projection of (2.5) to the ξ-space: ẋ = v, ẏ = w, żi = αiv + βiw, i = 1, . . . , n− 2. (2.6) the classification of fields (2.6) much simpler than the classification of fields with isolated singular points with the same spectrum, see [17, 18, 31]. it is explained by the fact that the non-trivial dynamics is connected with the center manifold of the field. in the case of (2.6), the center manifolds consists of singular points of the fields, therefore the non-trivial dynamics is identically zero. lemma 2.2: assume that reλ1,2(ξ◦) 6= 0 and there are no resonances p1λ1(ξ◦) + p2λ2(ξ◦) = 0, p1,2 ∈ z+, p1 + p2 ≥ 1. (2.7) then for any integer k ≥ 1, the germ of (2.6) is ck-equivalent to the germ ẋ = ṽ(x, y, z), ẏ = w̃(x, y, z), żi = 0, i = 1, . . . , n− 2, (2.8) copyright © 2020 assa. adv syst sci appl (2020) local normal forms of differential systems 123 at the origin, with some ṽ, w̃ ∈ 〈x, y〉. moreover, if reλ1,2(ξ◦) have the same sign, the above equivalence holds true with k =∞. without loss of generality, further we assume that |λ1(ξ◦)| ≤ |λ2(ξ◦)|. put λ◦ = λ1(ξ◦) : λ2(ξ◦), µ◦ = λ2(ξ◦) : λ1(ξ◦). lemma 2.3: assume that the conditions of lemma 2.2 hold true. then the germ of the field (2.6) is orbitally ck-equivalent to the germ (2.8), where ṽ, w̃ have the form (a) ṽ = x, w̃ = λ(z)y, if λ1,2(ξ◦) are real and µ◦ /∈ n ∪q−, (b) ṽ = λ(z)x, w̃ = y + α(z)xµ◦ , if µ◦ ∈ n \ {1}, (c) ṽ = α(z)x+ β(z)y, w̃ = −β(z)x+ α(z)y, if λ1,2(ξ◦) are complex. note that λ(0) = λ◦, µ(0) = µ◦ in the cases (a), (b). in the case (c), α(0), β(0) are respectively the real and imaginary parts of λ1,2(ξ◦). lemmas 2.2, 2.3 can be found in [17, 18, 31]. in the cases (a), (b), (c) enumerated in lemma 2.3, consider three families of local diffeomorphisms (rn, 0)→ (rn, 0) given by the following changes of the variables x, y: (a) x 7→ xϕ, y 7→ yϕλ(z), (2.9) (b) x 7→ xϕ, y 7→ yϕµ(z) + α(z)µ(z)xµ◦ ϕµ(z) − ϕµ◦ µ(z)− µ◦ , z 6= 0, y 7→ yϕµ(z) + α(z)µ(z)xµ◦ϕµ◦ lnϕ, z = 0, (2.10) with all smooth strictly positive functions ϕ = ϕ(x, y, z), and (c) x 7→ eα(z)ϕ(x cos β(z)ϕ+ y sin β(z)ϕ), y 7→ eα(z)ϕ(y cos β(z)ϕ− x sin β(z)ϕ), (2.11) with all smooth functions ϕ = ϕ(x, y, z). lemma 2.4: the families of diffeomorphisms (2.9), (2.10), (2.11) consist of symmetries of the direction field (2.8) in the cases (a), (b), (c), respectively. in other words, every diffeomorphism from the families (2.9), (2.10), (2.11) sends vector field (2.8) into a parallel vector field¶ in the cases (a), (b), (c), respectively. it is worth observing that families (2.9), (2.10), (2.11) are subgroups of the whole group of symmetries of the direction field (2.8) in the cases (a), (b), (c), respectively, but non of them coincide with the whole group. for instance, the family (2.9) does not contain the linear changes x 7→ c1x, y 7→ c2y with arbitrary constants c1,2 6= 0. the proof of lemma 2.4 is by direct calculations and omitted. now we can get the main result of the paper. theorem 2.2: let ξ◦ ∈ γ be an impasse point of the system (1.1) satisfying the above conditions. in addition, in the cases (a), (b), we assume that the eigenvectors with λ1,2(ξ◦) are transversal to the ¶two vector fields are parallel, if one of them is obtained from another by multiplication by a non-vanishing smooth function. in other words, two vector fields are parallel, if the corresponding direction fields coincide. copyright © 2020 assa. adv syst sci appl (2020) 124 a.m. kotyukov, s.o. nikanorov, n.g. pavlova criminant γ. then for any integer k ≥ 1, the germ of system (1.1) is ck-equivalent to the germ (x+ y)ψx′ = ṽ, (x+ y)ψy′ = w̃, z′1 = 0, . . . , z′n−2 = 0, (2.12) at the origin, where ṽ, w̃ are defined in lemma 2.3 and ψ = ψ(x, y, z) is a c∞-smooth function, ψ(0) 6= 0. under the same conditions, the germ of system (1.1) is orbitally ckequivalent to (2.12) with the same ṽ, w̃ and ψ ≡ 1. moreover, if reλ1,2(ξ◦) have the same sign, the above equivalences hold true with k =∞. proof by lemma 2.3, there exist a ck-smooth local diffeomorphism f : (ξ1, . . . , ξn) 7→ (x, y, z), z = (z1, . . . , zn−2), that sends the point ξ◦ to 0 and brings the germ of the direction field (2.6) at ξ◦ to the orbital normal form (2.8). then f transforms the corresponding direction field (2.5) into the form ẋ = ṽ, ẏ = w̃, żi = 0, i = 1, . . . , n− 2, ṫ = aṽ + bw̃, (2.13) where ṽ, w̃ are defined in 2.3 and a, b are smooth functions on the variables x, y, z. the criminant of the system (1.1) corresponding to (2.13) is given by the equation aṽ + bw̃. the condition that the eigenvectors with λ1,2(ξ◦) are transversal to the criminant yields a(0) 6= 0, b(0) 6= 0. let us show that there exists a diffeomorphism of the families (2.9), (2.10), (2.11) in the cases (a), (b), (c), respectively, that convert the criminant γ into a hyperplane, for instance, the hyperplane x+ y = 0. consider that case (a). then ṽ = x, w̃ = λ(z)y. by lemma 2.4, any diffeomorphism of the family (2.9) transforms the direction field (2.13) into ẋ = x, ẏ = λ(z)y, żi = 0, i = 1, . . . , n− 2, ṫ = axϕ+ bλ(z)yϕλ(z) = aϕ (x+ a−1bλ(z)ϕλ(z)−1y). (2.14) without loss of generality, assume that a−1(0)b(0)λ(0) > 0. otherwise one can make the change variables y 7→ −y, which preserves the first n− 1 components of the field (2.14) and brings its last component to the desirable form. thus, it suffices to establish the existence of a positive smooth function ϕ satisfying the equation a−1(xϕ, yϕλ(z), z) b(xϕ, yϕλ(z), z)λ(z)ϕλ(z)−1 = 1. (2.15) the derivative of the left-hand side of (2.15) by ϕ is not equal to zero. by the implicit function theorem, a function ϕ with the desirable properties exists in a neighborhood of the origin. the diffeomorphism (2.9) with ϕ founded above preserves the first n− 1 components of the direction field (2.14) and transforms its last component into ṫ = ψ(x, y, z) (x+ y), where ψ is a smooth function non-vanishing at the origin. the corresponding system (1.1) can be written in the form (x+ y)ψx′ = x, (x+ y)ψy′ = λ(z)y, z′1 = 0, . . . , z′n−2 = 0. (2.16) finally, if we deal with the orbital equivalence, one can make the change of the independent variable: t 7→ τ , where the new and old variables are connected with the differential equation dt dτ = ψ(x, y, z). copyright © 2020 assa. adv syst sci appl (2020) local normal forms of differential systems 125 the geometric sense this equation is that we make a parameterization integral curves of system (2.16), which depends on a point on every curve, so that the function ψ in (2.16) becomes identically 1. the case (b) is similar to (a) and omitted. consider that case (c). it is easy to see that (α− β)a− (α + β)b, (α + β)a+ (α− β)b do not simultaneously vanish at the origin. without loss of generality, we assume that the latter functions does not vanish at the origin. by lemma 2.4, any diffeomorphism of the family (2.11) transforms the direction field (2.13) with ṽ = α(z)x+ β(z)y, w̃ = −β(z)x+ α(z)y into the direction field whose first n− 1 components coincides with those in (2.13) and the nth component is ṫ = xeαϕ((αa− βb) cos βϕ− (αb+ βa) sin βϕ)+ yeαϕ((αa− βb) sin βϕ+ (αb+ βa) cos βϕ). thus, it suffices to establish the existence of a smooth function ϕ such that (αa− βb) cos βϕ− (αb+ βa) sin βϕ = (αa− βb) sin βϕ+ (αb+ βa) cos βϕ, (2.17) where α, β are given functions on z, while a, b depend on ϕ: a = a(eαϕ(x cos βϕ+ y sin βϕ), eαϕ(−x sin βϕ+ y cos βϕ), z), b = b(eαϕ(x cos βϕ+ y sin βϕ), eαϕ(−x sin βϕ+ y cos βϕ), z). equation (2.17) can be transformed to the equivalent form tan(βϕ) = (α− β)a− (α + β)b (α + β)a+ (α− β)b (2.18) the existence of a function ϕ satisfying equation (2.18) follows from the implicit function theorem. the remaining part of the proof in the case (c) repeats the reasonings carried out in the case (a). acknowledgements this paper was written with the financial support of rfbr (grants no. 19-01-00080, no. 20-01-00610). theorem 2.2 was obtained by the first author with the financial support of the russian science foundation (project no. 20-11-20131). references 1. brayton, r. & moser, j. (1964) a theory of nonlinear networks i/ii, q. appl. math., 22 (1/2), 1–33 / 81–104. 2. campbell, s.l. (1982) singular systems of differential equations ii. res. notes math., san francisco–london–melbourne: pitman advanced publishing program. 3. chua, l.o. (1980) dynamic nonlinear networks: state-of-the-art, ieee trans. circuits and systems-i: fund. theory and appl., 27 (11), 1059–1087. 4. chua, l.o. & deng, an-c. (1989) impasse points. i: numerical aspects, int. j. circuit theory appl., 17 (2), 213–235. copyright © 2020 assa. adv syst sci appl (2020) 126 a.m. kotyukov, s.o. nikanorov, n.g. pavlova 5. gerdts, m. (2012) optimal control of odes and daes. berlin: de gruyter. 6. gerdts, m. (2015) a survey on optimal control problems with differential-algebraic equations, surveys in differential-algebraic equations, cham: springer, 103–161. 7. heiland, j. (2016) a differential-algebraic riccati equation for applications in flow control, siam j. control optim., 54 (2), 718–739. 8. kurina, g.a. & marz, r. (2007) feedback solutions of optimal control problems with dae constraints, siam j. control optim., 46 (4), 1277–1298. 9. rabier, p.j. & rheinboldt, w.c. (1994) on impassee points of quasilinear differentialalgebraic equations, j. math. anal. appl., 181 (2), 429–454. 10. rabier, p.j. & rheinboldt, w.c. (1994) a geometric treatment of implicit differentialalgebraic equations, j. differ. equations, 109 (1), 110–146. 11. reiszig, g. (1996) differential-algebraic equations and impassee points, ieee trans. circuits and systems-i: fund. theory and appl., 43 (2), 122–133. 12. reibig, g. & boche, h. (2003) on singularities of autonomous implicit ordinary differential equations, ieee trans. circuits and systems-i: fund. theory and appl., 50 (7), 922–931. 13. reich, s. (1991) on an existence and uniqueness theory for nonlinear daes, circuits, systems, signal proc., 10 (3), 344–359. 14. smale, s. (1972) on the mathematical foundations of electrical circuit theory, j. diff. geom., 7, 193–210. 15. sobolev, s.l. (1954) on a new problem of mathematical physics, izv. akad. nauk sssr, ser. mat. 18, 3–50. 16. remizov, a.o. (2009) geodesics on 2-surfaces with pseudo-riemannian metric: singularities of changes of signature, mat. sb., 200 (3), 75–94. 17. ghezzi, r. & remizov, a.o. (2012) on a class of vector fields with discontinuities of divide-by-zero type and its applications to geodesics in singular metrics, j. dyn. control syst., 18, 135–158. 18. remizov, a.o. & tari, f. (2016) singularities of the geodesic flow on surfaces with pseudo-riemannian metrics, geom. dedicata, 185 (1), 131–153. 19. pavlova, n.g. & remizov, a.o. (2018) a brief survey on singularities of geodesic flows in smooth signature changing metrics on 2-surfaces, springer proc. in math. & stat., 222, 135–155. 20. pavlova, n.g. & remizov, a.o. (2019) completion of the classification of generic singularities of geodesic flows in two classes of metrics, izv. math. 83 (1), 104–123. 21. lamour, r., marz, r., & tischendorf, c. (2013) differential-algebraic equations. a projector based analysis. berlin: springer. 22. arnold, v.i. (1988) geometrical methods in the theory of ordinary differential equations. new york: springer-verlag. 23. arnold, v.i. & il’yashenko, yu.s. (1988) ordinary differential equations. dynamical systems i. encycl. math. sci. 1, springer-verlag, 7–140. 24. bruce, j.w. & tari, f. (1996) implicit differential equations from the singularity theory viewpoint, singularities and differential equations, banach center publ. 33, 23–38. 25. davydov, a.a. (1994) qualitative theory of control systems. math. monogr., 141, providence, rhode island: amer. math. soc. 26. davydov, a.a., ishikawa, g., izumiya, s., & sun, w.-z. (2008) generic singularities of implicit systems of first order differential equations on the plane, jap. j. math., 3 (1), 93–119. 27. guzmán-gómez, a.m. (1997) constrained equations with impasse points, j. math. anal. appl., 214 (1), 292–306. 28. pazii, n.d. (1999) locally analytic classification of equations of sobolev type. ph. d. thesis, chelyabinsk: chel. state univ. press. 29. remizov, a.o. (2002) on proper singular points of ordinary differential equations unsolved for derivatives, differ. equ., 38 (5), 654–662. copyright © 2020 assa. adv syst sci appl (2020) local normal forms of differential systems 127 30. remizov, a.o. (2002) implicit differential equations and vector fields with non-isolated singular points, sb. math., 193 (11), 1671–1690. 31. remizov, a.o. (2008) multidimensional poincaré construction and singularities of lifted fields for implicit differential equations, j. math. sci., 151 (6), 3561–3602. 32. sotomayor, j. & zhitomirskii, m. (2001) impasse singularities of differential systems of the form a(x)x′ = f (x), j. differ. equations, 169 (2), 567–587. 33. takens, f. (1976) constrained equations; a study of implicit differential equations and their discontinuous solutions, lect. notes math., 525, 143–234. 34. zhitomirskii, m. (1992) typical singularities of differential 1-forms and pfaffian equations. math. monogr., 113, providence, rhode island: amer. math. soc. copyright © 2020 assa. adv syst sci appl (2020) introduction main results adv syst sci appl 2021; 01:139–149 published online at https://ijassa.ipu.ru. application of the minimum principle of a tikhonov smoothing functional in the problem of processing thermographic data eugeniy laneev1∗, natalia chernikova1, obaida baaj1 1peoples’ friendship university of russia, moscow, russia abstract: the paper considers a method for correcting thermographic images. mathematical processing of thermograms is based on the analytical continuation of the stationary temperature distribution as a harmonic function from the surface of the object under study to the heat sources. the continuation is performed by solving an ill posed mixed problem for the laplace equation in a cylindrical region of rectangular cross-section. the cylindrical area is bounded by an arbitrary surface and plane. the cauchy conditions are set on the surface-the boundary values of the desired function and its normal derivative. inhomogeneous conditions of the first kind are set on the side faces of the cylinder. the problem is the inverse of the corresponding mixed problem for the poisson equation. in this paper, an approximate solution of the problem is obtained that is stable with respect to the error in the cauchy data and inhomogeneity in the boundary conditions. in the course of constructing an approximate solution, the problem is reduced to the fredholm integral equation of the first kind, which is solved using the minimum smoothing functional principle. the convergence of the approximate solution of the problem is proved when the regularization parameter is matched to the error in the data. keywords: termogram, ill-posed problem, inverse problem, cauchy problem for the laplace equation, integral equation of the first kind, tikhonov regularization method 1. introduction digital technologies have penetrated into all branches of human activity and one of the urgent problems is to improve the quality and information content of representations of research results, in particular, the quality of images obtained from measurement data, through their mathematical (digital) processing. this applies, for example, to images obtained by thermal imaging methods using a thermal imager that registers thermal electromagnetic radiation from the surface of the object under study in the infrared range. in particular, in medicine, thermal imaging has become an effective means of early diagnostics [1]. the image on the thermogram, which is a map of the temperature distribution on the surface of the patient’s body, makes it possible to assess functional abnormalities in the state of his internal organs. at the same time, the image on the thermogram in some cases turns out to be somewhat distorted due to the processes of thermal conductivity and heat exchange. the paper proposes a method for correcting the image on a thermogram within a certain mathematical model. as a corrected thermogram, the image of the temperature distribution on the plane near the density of heat sources is considered as more accurately transmitting the image of heat sources. it is proposed to obtain this distribution as a result of the continuation (similar to the continuation of gravitational fields in geophysics problems [2]) of the temperature distribution from the surface from which the original thermogram is taken. the continuation is obtained by solving ∗corresponding author: elaneev@yandex.ru 140 e. laneev, n. chernikova, o. baaj the inverse problem to a certain mixed boundary value problem for the poisson equation. the considered inverse problem is ill-posed, since small errors in the initial data (the initial thermogram) may correspond to significant errors in the solution of the inverse problem. to construct its stable approximate solution, we use the tikhonov regularization method [3], based on optimization methods [4]. 2. statement of the problem let’s consider a physical and then a mathematical model, within which we will set the inverse problem. the physical model is a homogeneous heat-conducting body in the form of a rectangular cylinder, bounded by the surface s and containing heat sources with a time-independent density function that create a stationary temperature distribution in the body. we associate the density function of heat sources with the object under study. we assume that a given temperature distribution is maintained on the side faces of the cylinder, and on the surface s there is a convective heat exchange with the external environment of temperature u0, described by newton’s law, according to which the heat flux density at a point on the surface is directly proportional to the temperature difference inside and outside. let’s move on to the mathematical model. in the cylinder of rectangular cross section d∞ = {(x, y, z) : 0 < x < lx, 0 < y < ly, −∞ < z <∞} ⊂ r3 consider a cylindrical domain d(f,∞) = {(x, y, z) : 0 < x < lx, 0 < y < ly, f (x, y) < z <∞}, (2.1) bounded by the surface s = {(x, y, z) : 0 < x < lx, 0 < y < ly, z = f (x, y) < h}. (2.2) let γ be the sum of side faces of the domain d(f,∞). in the domain d(f,∞) consider the following mixed boundary value problem for the laplace equation ∆u(m) = ρ(m), m ∈ d(f,∞), ∂u ∂n ∣∣∣ s = h(u0 − u) ∣∣∣ s , u|γ = f1, u is bounded when z →∞. (2.3) the problem (2.3) corresponds to the steady-state temperature distribution created with heat sources of the distribution density function ρ, on the surface s – the third boundary condition is set, corresponding to convective heat exchange with the external environment of temperature u0 with the coefficient h, on the boundary γ the temperature is set as a function f1. we assume that the density carrier ρ is located in the domain z > h. we also assume that the functions ρ, f1 are such that the solution of the problem (2.3) exists in c2(d(f,∞)) ⋂ c1(d(f,∞)). in particular, the solution of the problem (2.3) gives the boundary value u|s . now let’s set the inverse problem. inverse problem 1. let within the model (2.3) be set the following functions f = u|s, f1 = u|γ. (2.4) we need to find a continuous function ρ. copyright c© 2021 assa. adv syst sci appl (2021) application of the minimum principle... 141 note that density recovery is associated with the same difficulties as solving the inverse potential problem [5], for which significant restrictions on uniqueness classes are known. therefore, to solve the inverse problem, we apply the approach [2] used in geophysics problems. source of information about the density of ρ we will consider the function u|z=h on the plane z = h, that is closer to the density carrier ρ than the surface s. since the carrier of the function ρ by the condition of the problem (2.3) is located in the domain z > h, then the solution of the problem (2.3) in the domain d(f,h) = {(x, y, z) : 0 < x < lx, 0 < y < ly, f (x, y) < z < h} (2.5) satisfies the laplace equation. the sum of side faces of the domain d(f,h) denote by γh . instead of the inverse problem 1, we will solve the following inverse problem inverse problem 2. let within the model (2.3) be set the following functions f = u|s, f1 = u|γh . (2.6) we need to find a solution u to the boundary value problem in the domain d(f,h) ∆u(m) = 0, m ∈ d(f,h), u|s = f, ∂u ∂n ∣∣∣ s = h(u0 − f) ∣∣∣ s , u|γh = f1. (2.7) we will consider the function u|z=h as a source of information about the density ρ. we assume that the functions f, f1 in (2.6), (2.7) are taken from the set of solutions of the direct problem (2.3), so the solution of the inverse problem exists in c2(d(f,h)) ⋂ c1(d(f,h)). we note that in the problem (2.7) on the surface s of the form (2.2), cauchy conditions are set, that is, the boundary values f of the desired function u and the values of its normal derivative are set, so the problem (2.7) has a unique solution. the boundary z = h of the domain d(f,h) of the form (2.5) is free and, thus, the problem (2.7) is unstable with respect to data errors, i.e. it is ill-posed. we will construct an explicit representation of the exact solution of the problem (2.7). 3. exact solution of the problem let’s construct an exact solution of the problem (2.7), following the scheme [6, 7]. consider the source function ϕ(m,p ) of the dirichlet problem in the cylinder d∞: ∆u(p ) = ρ(p ), p ∈ d∞, u|x=0,lx = 0, u|y=0,ly = 0, u→ 0 when |z| → ∞, (3.8) i.e., ϕ(m,p ) = 1 4πrmp +w (m,p ), (3.9) where rmp is the distance between points m and p and w (m,p ) is a harmonic function of point p . the source function can be obtained by the reflection method as a sum of point source functions with period 2lx in the variables x and period 2ly in the variable y, ϕ(m,p ) = 1 4π ∞∑ n,m=−∞ ( 1 r1,nm − 1 r2,nm − 1 r3,nm + 1 r4,nm ) , (3.10) copyright c© 2021 assa. adv syst sci appl (2021) 142 e. laneev, n. chernikova, o. baaj where r1,nm = [(xm − xp + 2lxn)2 + (ym − yp + 2lym)2 + (zm − zp )2]1/2, r2,nm = [(xm + xp + 2lxn)2 + (ym − yp + 2lym)2 + (zm − zp )2]1/2, r3,nm = [(xm − xp + 2lxn)2 + (ym + yp + 2lym)2 + (zm − zp )2]1/2, r4,nm = [(xm + xp + 2lxn)2 + (ym + yp + 2lym)2 + (zm − zp )2]1/2, and, in particular, r1,00 = rmp . let m ∈ d(f,h). then, applying the green formulas in the domain d(f,h) to the function u(p ), i.e., the solution of problem (2.7), and to the functions 1 4πrmp and w (m,p ) in (3.9), we obtain u(m) = ∫ ∂d(f,h) [∂u ∂n (p ) 1 4πrmp − u(p ) ∂ ∂np 1 4πrmp (m,p ) ] dσp , m ∈ d(f,h) (3.11) and 0 = ∫ ∂d(f,h) [∂u ∂n (p )w (m,p )− u(p ) ∂w ∂np (m,p ) ] dσp , m ∈ d(f,h). (3.12) summing (3.11) and (3.12) taking into account (3.9) we obtain u(m) = ∫ ∂d(f,h) [∂u ∂n (p )ϕ(m,p )− u(p ) ∂ϕ ∂np (m,p ) ] dσp , m ∈ d(f,h). (3.13) given homogeneous boundary conditions for ϕ and inhomogeneous ones for u on the side faces γh of the cylindrical domain d(f,h), we obtain u(m) = ∫ s [ h(u0 − f(p ))ϕ(m,p )− f(p ) ∂ϕ ∂np (m,p ) ] dσp− − ∫ γh [ f1(p ) ∂ϕ ∂np (m,p ) ] dσp + ∫ π(h) [∂u ∂n (p )ϕ(m,p )− u(p ) ∂ϕ ∂np (m,p ) ] dσp , where π(h) = {(x, y, z) : 0 < x < lx, 0 < y < ly, z = h}. (3.14) in the domain zm < h , we introduce the notation φ(m) = ∫ s [ h(u0 − f(p ))ϕ(m,p )− f(p ) ∂ϕ ∂np (m,p ) ] dσp − ∫ γh [ f1(p ) ∂ϕ ∂np (m,p ) ] dσp , (3.15) v(m) = ∫ π(h) [∂u ∂n (p )ϕ(m,p )− u(p ) ∂ϕ ∂np (m,p ) ] dσp , zm < h. (3.16) then we obtain the solution of the problem (2.7) in the form u(m) = v(m) + φ(m), m ∈ d(f,h), (3.17) copyright c© 2021 assa. adv syst sci appl (2021) application of the minimum principle... 143 where the function φ is calculated from known functions f and f1. if the solution of the problem (2.7) exists, then the function v of the form (3.16), harmonic in the domain d(−∞, h) = {(x, y, z) : 0 < x < lx, 0 < y < ly, −∞ < z < h}, can be represented in d(f,h) ⊂ d(−∞, h) according to (3.17) in the form v = u− φ and then it may be defined on the boundary of π(h) as a continuous function v |z=h= u |z=h −φ |z=h= vh . (3.18) thus, the function v can be viewed as a solution of the problem ∆v(m) = 0, m ∈ d(−∞, h), v|z=h = vh , v|x=0,lx = 0, v|y=0,ly = 0, v → 0 when z → −∞, (3.19) and the function v can be expressed in terms of the function vh by using the green function of problem (3.19) as follows: v(m) = − ∫ π(h) ∂g ∂np (m,p )vh(p )dxpdyp , m ∈ d(−∞, h), (3.20) where ∂g ∂np (m,p ) ∣∣∣ p∈π(h) = ∂g ∂zp (m,p ) ∣∣∣ p∈π(h) = = − 4 lxly ∞∑ n,m=1 exp { knm(−h + zm) } sin πnxm lx sin πmym ly sin πnxp lx sin πmyp ly , (3.21) knm = π ( n2 l2x + m2 l2y )1/2 . (3.22) it follows that if problem (2.7) has a solution, then (3.20) implies that the function v in the domain d(−∞, h) can be represented as the fourier series v(m) = v(x, y, z) = − ∞∑ n,m=1 (ṽh)nm exp { knm(z −h) } sin πnx lx sin πmy ly , (3.23) (ṽh)nm = 4 lxly lx∫ 0 ly∫ 0 vh(x′, y′) sin πnx′ lx sin πmy′ ly dx′dy′, (3.24) of a complete system of functions{ sin πnx lx sin πmy ly }∞ n,m=1 . (3.25) the series (3.23) uniformly converges in the domain d(−∞, h − ε) for any ε > 0, because∣∣(ṽh)nm exp { knm(z −h) } sin πnx lx sin πmy ly ∣∣ 6 ∣∣(ṽh)nm ∣∣ exp { − εknm } . copyright c© 2021 assa. adv syst sci appl (2021) 144 e. laneev, n. chernikova, o. baaj thus, it follows from the representation (3.17) of the solution of problem (2.7) and from (3.23) that, to obtain an explicit expression for the exact solution of problem (2.7), it suffices to express the function vh (3.18) in terms of the prescribed functions f and f1. let us show that the function vh satisfies a fredholm integral equation of the first kind. let m ∈ d(−∞, f ), where d(−∞, f ) = {(x, y, z) : 0 < x < lx, 0 < y < ly, −∞ < z < f (x, y)}. applying the green formula in the domain d(f,h) to the function u(p ), i.e., a solution of problem (2.7), and to a function ϕ(m,p ) of the form (3.9), we, by analogy with (3.11), (3.12), and (3.13), obtain the relation 0 = ∫ ∂d(f,h) [∂u ∂n (p )ϕ(m,p )− u(p ) ∂ϕ ∂np (m,p ) ] dσp , m ∈ d(−∞, f ). from this, with regard to the homogeneous boundary conditions for the function ϕ and to the inhomogeneous boundary conditions for u and notation (3.15) and (3.16), we obtain v(m) = −φ(m), m ∈ d(−∞, f ). (3.26) let a < min (x,y) f (x, y) andm ∈ π(a), where π(a) is a domain of the form (3.14) for z = a. then, by formulas (3.26) and (3.20), we obtain the integral equation of the first kind∫ π(h) ∂g ∂np (m,p )vh(p )dxpdyp = φ(m), m ∈ π(a). (3.27) from the equation (3.27) taking into account the decomposition (3.21) for zm = a we obtain the following relations between the fourier coefficients of the unique solution vh of this integral equation and the fourier coefficients of its right-hand side: − (ṽh)nm exp { − knm(h − a) } = φ̃nm(a), (3.28) where the φ̃nm(a) are the fourier coefficients of the function φ(m)|m∈π(a), φ̃nm(a) = 4 lxly ∫ π(a) φ(x, y, a) sin πnx lx sin πmy ly dxdy. (3.29) note that formula (3.28) characterizes the decrease in the fourier coefficients φ̃nm(a) with increasing n and m if, for the functions f and f1, there exists a solution of problem (2.7) and hence a function vh defined by (3.18). we express the fourier coefficients (ṽh)nm, substitute them into the series (3.23), and obtain the function v in the domain d(−∞, h) : v(m) = − ∞∑ n,m=1 φ̃nm(a) exp {knm(z − a)} sin πnx lx sin πmy ly , m(x, y, z) ∈ d(−∞, h). (3.30) the series (3.30), just as the series (3.23), uniformly converges in the domaind(−∞, h − ε) for any ε > 0 if there exists a solution of problem (2.7) for the given functions f and f1. formula (3.17), where the functions v and φ are given by (3.30) and (3.15), respectively, gives an explicit expression for the solution of problem (2.7). copyright c© 2021 assa. adv syst sci appl (2021) application of the minimum principle... 145 4. approximate solution of the problem let the functions f and f1 in problem (2.7) be given with an error; i.e., let, instead of them, functions f δ and f δ1 be given such that ‖f δ − f‖l2(s) 6 δ, ‖f δ1 − f1‖l2(γh) 6 δ. we construct an approximate solution of problem (2.7) converging to the exact solution as δ → 0. here the function φ defined by formula (3.15) can be obtained approximately as φδ(m) = ∫ s [ h(u0 − f δ(p ))ϕ(m,p )− f δ(p ) ∂ϕ ∂np (m,p ) ] dσp− − ∫ γh [ f δ1 (p ) ∂ϕ ∂np (m,p ) ] dσp . (4.31) we apply the cauchy-schwarz inequality to the difference of functions (4.31) and (3.15) for m ∈ π(a), a < min (x,y) f (x, y), and obtain an estimate of the right-hand side of integral equation (3.27), |φδ(m)− φ(m)| 6 h max m∈π(a) ( ∫ s ϕ2(m,p )dσp )1/2‖f δ − f‖l2(s)+ + max m∈π(a) ( ∫ s [ ∂ϕ ∂np (m,p ) ]2 dσp )1/2‖f δ − f‖l2(s)+ + max m∈π(a) ( ∫ γh [ ∂ϕ ∂np (m,p ) ]2 dσp )1/2‖f δ1 − f1‖l2(γh) 6 cδ. (4.32) for an approximate solution of eq. (3.27), we take the extremal of the tikhonov functional [3, p. 68] with zero-order stabilizer, mα[w] = ∥∥∫ s ∂g ∂n wdσ − φδ ∥∥2 l2(π(a)) + α‖w‖2 l2(π(h)), α > 0, (4.33) where π(a) and π(h) are domains defined by formula (3.14). the extremal can be obtained as a solution of the euler equation for the functional (4.33) which, in the fourier coefficients of the function w, has the form exp {−2knm(h − a)} w̃nm + αw̃nm = − exp {−knm(h − a)} φ̃δ nm(a), where φ̃δ nm(a) = 4 lxly ∫ π(a) φδ(x, y, a) sin πnx lx sin πmy ly dxdy (4.34) are the fourier coefficients of the function φδ(m)|m∈π(a). solving the equation for the fourier coefficients of the extremal and substituting the extremal wδα for vh into representation (3.23), we obtain an approximation vδα to the function copyright c© 2021 assa. adv syst sci appl (2021) 146 e. laneev, n. chernikova, o. baaj v in the domain d(−∞, h), vδα(m) = − ∞∑ n,m=1 φ̃δ nm(a) exp{knm(zm − a)} 1 + α exp{2knm(h − a)} sin πnxm lx sin πmym ly . (4.35) note that coefficients of the series (4.35) differ from corresponding coefficients of the series (3.30) in the factor (1 + α exp{2knm(h − a)})−1, and the series (4.35) converges uniformly. according to the representation (3.17), we obtain an approximate solution of problem (2.7) in the form uδα(m) = vδα(m) + φδ(m), m ∈ d(f,h), (4.36) where vδα and φδ are the functions defined by formulas (4.35) and (4.31) respectively. theorem. assume that there exists a solution of problem (2.7). then, for any α = α(δ) > 0 such that α(δ)→ 0 and δ/ √ α(δ)→ 0 as δ → 0, the function uα(δ) of the form (4.36) uniformly converges as δ → 0 to the exact solution of problem (2.7) on any compact k ⊂ d(f,h). proof. on any compactk ⊂ d(f,h), according to representations (4.36) and (3.17), we estimate the difference |uδα − u| 6 |vδα − v|+ |φδ − φ|. (4.37) obviously, there is ε > 0 such that k ⊂ d(−∞, h − ε). for the modul of the difference vδα − v in the domain d(−∞, h − ε) we obtain |vδα − v| 6 |vδα − vα|+ |vα − v|, (4.38) where vα is a function of the form (4.35) for exact functions f and f1, vα(m) = − ∞∑ n,m=1 φ̃nm(a) exp{knm(zm − a)} 1 + α exp{2knm(h − a)} sin πnxm lx sin πmym ly . (4.39) to estimate the difference vδα − vα on the right-hand side in inequality (4.38) for zm < h − ε, we use inequality (4.32) |vδα(m)− vα(m)| 6 ∣∣∣∣ ∞∑ n,m=1 exp{knm(zm − a)} 1 + α exp{2knm(h − a)} ∣∣∣∣ · 4 max p∈π(a) ∣∣φδ(p )− φ(p ) ∣∣ 6 6 c1δ ∞∑ n,m=1 exp{knm(h − ε− a)} 1 + α exp{2knm(h − a)} 6 6 c1δmax x [ ex 1 + αe2x ] ∞∑ n,m=1 exp{−knmε} 6 c2 δ√ α . (4.40) we estimate the difference vα − v in inequality (4.38) for zm < h − ε, |vα − v| 6 ∞∑ n,m=1 α exp{2knm(h − a)} exp{knm(h − ε− a)} 1 + α exp{2knm(h − a)} ∣∣φ̃nm(a) ∣∣. copyright c© 2021 assa. adv syst sci appl (2021) application of the minimum principle... 147 from this, using (3.28) and applying the cauchyschwarz inequality, we obtain |vα − v| = ∞∑ n,m=1 α exp{2knm(h − a)} exp{−knmε} 1 + α exp{2knm(h − a)} ∣∣ ˜(vh)nm ∣∣ 6 6 [ ∞∑ n,m=1 ( α exp{2knm(h − a)} 1 + α exp{2knm(h − a)} )2 exp{−2knmε} ]1/2 · 2√ lxly ||vh ||l2 . since the series depending on the parameter α is majorized by the converging numerical series with coefficients exp{−2εknm}, it is possible to pass to the limit in α, and hence |vα − v| → 0 when α→ 0. (4.41) it follows from (4.38), (4.40), (4.41), and the assumptions of the theorem that |vδα(δ) − v| → 0 when δ → 0. (4.42) the second difference on the right-hand side in inequality (4.37) can be estimated by analogy with (4.32). we apply the cauchy-schwarz inequality to this difference for m ∈ k, and obtain |φδ(m)− φ(m)| 6 hmax m∈k (∫ s ϕ2(m,p )dσp )1/2 ‖f δ − f‖l2(s)+ + max m∈k (∫ s [ ∂ϕ ∂np (m,p ) ]2 dσp )1/2 ‖f δ − f‖l2(s)+ + max m∈k ( ∫ γh [ ∂ϕ ∂np (m,p ) ]2 dσp )1/2 ‖f δ1 − f1‖l2(γh) 6 c3δ. from this relation, inequality (4.37), and formula (4.42) the assertion of the theorem follows. 5. numerical solution of the problem the effectiveness of the proposed method for solving the problem (2.7) is shown in the following model example. in the problem (2.3), let the surface s be the plane π(0), f1 = u0 = 24, h = 0.5, lx = 30, ly = 30, h = 1.4, and the function ρ corresponds to three point sources at points in the plane π(h) : (x1, y1) = (8.8), (x2, y2) = (10.8), (x3, y3) = (10.10). the boundary value of the solution of the model problem (2.3) in this case has the form f(x, y) = u0 + ∞∑ n,m=1 3∑ i=1 e−knmh knm + h sin πnxi lx sin πmyi ly sin πnx lx sin πmy ly , (5.43) where knm is calculated using the formula (3.22). to set the inverse problem (2.7), we consider that the function f, calculated by the formula (5.43), a known function. also f1 = u0 = 24, h = 0.5, lx = 30, ly = 30, h = 1.4 are known. copyright c© 2021 assa. adv syst sci appl (2021) 148 e. laneev, n. chernikova, o. baaj fig. 5.1. the initial data of the inverse problem (initial thermogram) fig. 5.2. the result of restoring the thermogram u|z=h to solve the inverse problem (2.7), we use the formulas (4.36), (4.35), (4.34), (4.31). in the formula (4.31) we use the representation for the fundamental solution ϕ(m,p ) = 2 lxly ∞∑ n,m=1 e−knm|zm−zp | knm sin πnxm lx sin πmym ly sin πnxp lx sin πmyp ly (5.44) when zm = a, zp = 0. the fourier coefficients in the formula (4.34) are calculated without calculating the function φ, similarly to [8]. when using the formula (4.34), integration is performed under the sign of the integral in (4.31) and under the sign of the sum in (5.44). taking into account the orthogonality of the system of functions (3.25), the calculation formulas for calculating the fourier coefficients φnm are significantly simplified. to obtain a numerical result, the problems (2.3), (2.7) are discretized. a uniform grid of 91x91 points is introduced on the rectangles π(a), a = −0.5 and π(h). the hamming algorithm [9, p.83] is used to sum discrete fourier series. the calculation results are shown in fig.5.1 and fig.5.2. fig.5.1 shows the initial data of the inverse problem – the function f calculated from the discrete analog of the formula (5.43). the relative magnitude of the added error is 0.28%. the three sources are perceived as a single whole. fig.5.2 shows the result of restoring the u|z=h function using the formulas (4.36), (4.35), (4.34), (4.31). three sources are clearly visible. regularization parameter α = 10−8. with the regularization parameter α = 0, the solution is destroyed. copyright c© 2021 assa. adv syst sci appl (2021) application of the minimum principle... 149 6. conclusion the inverse problem (2.7) and its stable solution can be used for mathematical processing of thermograms, in particular, in medicine [1], in order to correct the image. as already mentioned, a thermogram obtained using a thermal imager transmits an image of the structure of heat sources inside the body approximately. refinement of the image on the thermogram can be performed within the framework of the problem (2.7). in this case, the f function will be associated with the original thermogram, and the uh function will be considered as the result of processing the thermogram. since the function u|z=h represents the temperature distribution on a plane closer to the heat sources under study than the original surface s, we can expect a more accurate reproduction of the source image on the calculated thermogram u|z=h . the results of calculations, performed on the model example, show the effectiveness of the proposed method and algorithm based on the formulas (4.36), (4.35), (4.34), (4.31), which can be used for processing thermographic images. acknowledgements the research is supported by the russian science foundation, project n 21-11-00064. references 1. ivanitskii g.r. (2006). teplovideniye v meditsine [thermovision in medicine], vestnik ran, 76(1), 44-53, [in russian]. 2. tihonov a.n., glasko v.b., litvinenko o.k. & melihov v.r. (1968). o prodolzhenii potentsiala v storonu vozmushchayushchih mass na osnove metoda regulyarizatsii [on the continuation of the potential towards disturbing masses based on the regularization method], izvestiya an sssr. fizika zemli, 1, 30-48, [in russian]. 3. tihonov a.n. & arsenin v.ya. (1979). metody resheniya nekorrektnyh zadach [methods for solving ill-posed problems]. moscow, russia: nauka, [in russian]. 4. arutyunov a., jacimovic v.& pereira f. (2003). isecond order necessary conditions for optimal impulsive control problems ,journal on dynamical and control systems, 9(1), 131-153. 5. prilepko a.i. (1973). inverse problems of potential theory (elliptic, parabolic, hyperbolic, and transport equations), math. notes, 14(5), 990-996. 6. laneev, e.b. (2018). construction of a carleman function based on the tikhonov regularization method in an ill-posed problem for the laplace equation, differential equations, 54(4), 476–485, https://doi.org/10.1134/s0012266118040055 7. chernikova n.y., laneev e.b., muratov m.n.& ponomarenko e.y. (2020). on an inverse problem to a mixed problem for the poisson equation. in: pinelas s., kim a., vlasov v. (eds) mathematical analysis with applications. concord-90 2018. springer proceedings in mathematics & statistics, vol. 318, pp. 141–146. springer, cham. https://doi.org/10.1007/978-3-030-42176-2 14 8. laneev, e.b., mouratov, m.n. & zhidkov, e.p. (2008). discretization and its proof for numerical solution of a cauchy problem for laplace equation with inaccurately given cauchy conditions on an inaccurately defined arbitrary surface, physics of particles and nuclei letters, 5(3), 164–167. 9. hamming r.w. (1962). numerical methods for scientists and engineers. new york, usa: mcgraw-hill book company. copyright c© 2021 assa. adv syst sci appl (2021) introduction statement of the problem exact solution of the problem approximate solution of the problem numerical solution of the problem conclusion adv syst sci appl 2021; 02:104–116 published online at https://ijassa.ipu.ru. achieving angular superresolution of control and measurement systems in signal processing boris lagovsky1, evgeny rubinovich2* 1russian technological university (mirea), moscow, russia 2trapeznikov institute of control sciences, russian academy of sciences, moscow, russia abstract: algebraic methods of processing data obtained by control and measurement systems to achieve angular superresolution are presented. the efficiency of using the methods in the formation of approximate images of objects at low signal-to-noise ratios is shown. the results of numerical experiments demonstrate the possibility of obtaining images with a resolution exceeding the rayleigh criterion by 3-10 times. the robustness of the solutions obtained by the methods of algebraic exceeds many well-known approaches. the relative simplicity of the presented methods allows the use of inexpensive computing devices and perform real-time measurement processing. keywords: angular superresolution, inverse problem, integral equation, rayleigh criterion 1. introduction an important modern problem of improving control and measurement systems is increasing their information content based on new signal processing methods. one of the directions of its solution is to increase the angular resolution of angle-measuring systems, which makes it possible to detail the image of the object under study. in this regard, the problem of restoring the image of the object with an angular superresolution becomes highly significant. this paper is an extension of work in this direction originally presented in ieee 2020 7th international conference on control, decision and information technologies (codit-2020) [1]. due to the importance of the problem, hundreds of publications in many countries have already been devoted to the problems of achieving angular superresolution. we will highlight many articles that are general [2–5]. currently popular methods are music [6–8], esprit [9], the deconvolution method [10], the maximum entropy method [11, 12], the borgiottilagunas method [13], the capon method [14], the maximum likelihood method [15] and others [16]. all these methods are not universal and are not always effective. they begin to work successfully and allow us to increase the effective angular resolution at a signal to noise ratio (snr) not lower than 20-25 db. the developed algebraic methods [17–19] differ favorably from those mentioned above in that they have significantly higher noise immunity. in addition, they are relatively simple, which significantly reduces signal processing time and allows the use of relatively simple computing devices. as a result, unlike many other methods, it is possible to apply them in real-time. ∗corresponding author: rubinvch@hotmail.com achieving angular superresolution of control and measurement systems 105 2. problem statement let us first consider the one-dimensional case. the signal u(α) received by the goniometer system when scanning the area under study is an integral transformation u(α) = ∫ ω f (α − φ) i(φ) dφ, (2.1) where f (α) is the directional pattern (dp) of the system, i(α) is the unknown angular distribution of the amplitude of the signal reflected (or emitted) by the object, ω is the angular sector, in which the object under study is located. the problem of reconstructing the image of the source i(α) with the highest possible angular resolution exceeding the rayleigh criterion, i.e. with super-resolution, is posed. dp and measurement data in the form of u(α) are considered known. mathematically, the search for the angular distribution i(α) is reduced to the approximate solution of fredholm integral equations of convolution type (2.1) concerning the function i(·). it is known that of the three conditions for the correctness of hadamard problems (existence of the solution, uniqueness of the solution, stability of the solution), the fredholm equation (2.1) does not satisfy the second and third conditions. thus, the problem under consideration belongs to the class of inverses and is incorrect. because of this, attempts to increase the angular resolution and exceed the rayleigh criterion generally lead to instabilities in the solutions and, as a result, to significant errors. the main obstacle to obtaining solutions with super-resolution is the random components present in the received signal. when obtaining a resolution within the rayleigh criterion, their influence is usually negligible but increases dramatically (exponentially) when trying to obtain a higher resolution. achieving super-resolution, as studies show, is possible, but up to a certain limit, determined by the snr, measurement accuracy, and accuracy of the dp assignment. for a comparative assessment of the quality of the approximate solution of the problem (2.1), three indicators are usually used by different methods. the first one is the achieved degree of exceeding the rayleigh criterion. the second indicator is the value of the displacement of the found positions of objects from the true ones since the positions of the resolved objects and their elements found during signal processing do not always exactly correspond to the actual directions to the radiation sources. when observing a single radiation source, the bias is usually close to zero. in the case of two or more sources, the offset can be different from zero and increase as you increase the reach of effective solutions. the third indicator is the snr, at which the specified degree of excess of the rayleigh criterion is achieved. it is theoretically known and has been confirmed in the course of numerical experiments that the solutions of inverse problems are by their nature very sensitive to the presence of random components in the source data. because of this, the first and third indicators can be combined. this is the dependence of the degree of excess of the rayleigh criterion on the snr. this indicator was considered the main one in the development of methods and algorithms for achieving superresolution. 3. algebraic methods of solution to solve the inverse problem (2.1), we propose methods and algorithms for digital signal processing, which can be called algebraic. they consist in finding solutions in the form of decompositions over the given sequences of functions gm(α) [20], which are orthonormal in copyright© 2021 assa. adv syst sci appl (2021) 106 b. lagovsky, e. rubinovich the region of the source location ω. i(α) = ∞∑ m=1 bm gm(α) � n∑ m=1 bm gm(α), (3.2) where bm are unknown expansion coefficients. by using (3.2), instead of (2.1), we get a decomposition of the useful signal u(α) obtained by scanning over a system of non-orthogonal functions χm(α) u(α) � n∑ m=1 bm χm(α), (3.3) χm(α) � ∫ ω f (α − φ) gm(φ) dφ. (3.4) all the bm necessary to obtain the solution (3.2) are found from the condition of minimizing the root-mean-square deviation of the right-hand side (3.3) from u(α). as a result, the search for the vector b of unknown coefficients bm is reduced to the solution of slae: v = gb, (3.5) where the elements of the vector v and matrix g are: vm = ∫ θ u(α) χm(α) dα, gnm = ∫ θ χn(α) χm(α) dα, (3.6) n,m = 1, 2, . . . ,n. thus, the search for an approximate solution of i(α) in the form of a finite system expansion of the selected functions allows us to parametrize the inverse problem and reduce its solution to the slae solution [17–19]. the principal feature of the resulting systems is their ill-conditioned, which is a consequence of the attempt to solve the inverse problem. the larger the chosen area of integration θ, the higher the stability of the obtained approximate solutions. however, the received signal always contains random components. for this reason, to obtain an adequate solution to the inverse problem (2.1), the level of the useful signal must significantly exceed the level of the random components. with the growth of θ, the snr decreases in the areas close to the boundary, which limits the area of integration θ. therefore, the integration region should be defined as the sector of angles within which the snr value is sufficient to obtain stable solutions. it should be noted that the necessary snr for obtaining solutions should be provided not directly at the reception of the signal but after the primary signal processing. so, the presented algebraic method and its variants make it possible, by consistently increasing the number of functions used in (3.2), i.e., by increasing the effective angular resolution, to approach the limit resolution for each problem to be solved. the iterative procedure for finding an approximate solution continues as long as it is possible to obtain a stable solution. comparing the algebraic method with others, it should be noted that it potentially allows one to obtain an exact solution to the inverse problem (2.1) if the chosen system of functions gm(α) in (3.2) provides an exact representation of the solution i(α) using a finite number of terms. it is easy to prove that such a choice of the system gm(α) for a known distribution i(α) is always possible. thus, the algebraic method turns out to be at least as good as any other known method for solving the inverse problem (2.1). the question of choosing the optimal system of functions gm(α) for the desired i(α) remains, however, open. it is known that a significant improvement in the quality of solutions to inverse problems can be achieved by using a priori information about the solution. the presence of such copyright© 2021 assa. adv syst sci appl (2021) achieving angular superresolution of control and measurement systems 107 information allows, in particular, to use it when optimizing the choice of the system of functions gm(α) for constructing a solution. preliminary information about the solution, in addition, is realized by using algebraic methods in the form of: choosing the location, size, and shape of the studied area of localization of the signal source ω, introducing additional conditions in the form of equations and inequalities that connect the coefficients of expansion over sequences of functions and thereby regularize the problem. fig. 3.1 shows the solution to the problem of image reconstruction of two closely spaced smoothly inhomogeneous signal sources (dash line 1) in the form of a distribution of the reflected signal amplitude at a low noise level. for illustration, curve 3 shows a signal received by an angle-measuring system with a beamwidth of θ0.5 in the angular region of the source location. when constructing the solution (3.3)-(3.6), we used preliminary information fig. 3.1. the restoration of a smoothly inhomogeneous signal source about the smoothly inhomogeneous nature of the amplitude angular distribution of the signal reflected by the source. this predestined the use of trigonometric functions as a system of functions gm(α) in (3.2). the resulting approximate solution (solid curve 2) allowed us to resolve the sources and almost accurately determine their angular position and shape. due to the oscillating nature of the system of functions chosen to represent the solution, false sources with small amplitudes have emerged. fig. 3.2 shows the solution to the problem of image reconstruction of four closely located identical signal sources (dash line 1), close to the point at a very small snr of 13 db. the resulting approximate solution (solid line 2) allowed us to resolve all four sources and determine their angular position with good accuracy. for illustration, curve 3 shows the signal received and processed by the goniometer system. the usual methods of obtaining superresolution mentioned earlier do not allow us to obtain adequate solutions at such a low snr level. satisfactory solutions with the resolution of all four objects, but somewhat worse quality, can be obtained even at even lower snr values. copyright© 2021 assa. adv syst sci appl (2021) 108 b. lagovsky, e. rubinovich fig. 3.2. the restoration of the images of point sources at low snr as already noted, with an increase in the number of functions gm(α) used in the decomposition, the conditionality numbers of matrices g in (3.5) and similar ones, they increase sharply (according to the exponential law) and the solutions become less stable. the dimension of the matrices can be reduced by selecting the functions used. the purpose of selection is to select only those functions from the selected system that best represent the solution. selection can be carried out in several ways: 1. based on the analysis of the signal spectrum and function spectra, followed by their selection with the spectra closest to the spectrum of the signal under study. 2. based on the analysis of the mutual correlation functions of the signal and the functions from the selected family. 3. based on the selection of functions from the used family with the highest values of the expansion coefficients in the solution representation. fig. 3.3 shows the results of restoring the source image at a very high noise level, comparable to the useful signal (snr = 6-8 db). the dashed curve is the original intensity distribution, the solid bold curve is the reconstructed image using six wavelets. for comparison, a solid thin polyline is shown as a reconstructed image using the same six wavelets in the absence of noise. the obtained solutions provided a 2-3 fold excess of the rayleigh criterion, and accuracy of source localization θ 0.5/8, and very high noise immunity. most of the known methods are designed to obtain solutions to one-dimensional problems. their generalization to two-dimensional problems significantly complicates the algorithms and increases the instability of solutions. in addition, the signal processing time increases dramatically. to obtain satisfactory results, it is sometimes necessary to use parallel processors. generalization of algebraic methods for solving one-dimensional problems to two-dimensional ones does not lead to a serious complication of the algorithms. we present the desired angular two-dimensional distribution i(α, ϕ) as a decomposition over a finite system of orthogonal functions gnm(α, ϕ) in the two-dimensional domain ω with copyright© 2021 assa. adv syst sci appl (2021) achieving angular superresolution of control and measurement systems 109 fig. 3.3. restoring the source image using the selected images mhat wavelets at snr = 6 db unknown coefficients. it is convenient to use separable systems, which are represented as the product of one-dimensional systems gnm(α, ϕ) = gn(α)gm(ϕ). then: i(α, ϕ) � n∑ n,m=1 bnm gn(α) gm(ϕ). (3.7) and the received signal by (2.1) is obtained in the form: u(α, ϕ) � n∑ n,m=1 bnm ψnm(α, ϕ), ψnm(α, ϕ) = ∫ ω f (α − α′, ϕ − ϕ′) gn(α′) gm(ϕ′) dα′ dϕ′, n, m = 1, 2, . . . ,n. (3.8) the coefficients bnm, which provide the minimum root-mean-square deviation in the twodimensional region θ of the synthesized signal (3.8) from the received one, are found as slae solutions: ∫ θ u(α, ϕ) ψ jk(α, ϕ) dα dϕ = = n∑ n,m=1 bm ∫ θ ψ jk(α, ϕ) ψnm(α, ϕ) dα dϕ, j, k = 1, 2, . . . ,n. (3.9) numerical experiments have shown that the solution time and stability of solutions to twodimensional problems in comparison with one-dimensional problems at the same level of copyright© 2021 assa. adv syst sci appl (2021) 110 b. lagovsky, e. rubinovich superresolution vary slightly [21, 22]. thus, algebraic methods allow us to parametrize oneand two-dimensional inverse problems, and to reduce the solutions of integral equations to the slae solution, sometimes with additional conditions [23, 25, 27]. fig. 3.4 shows an approximate solution of problems (3.7)-(3.9). the original source was defined as four identical objects with small angular coordinates, located at the corners of a square with a side of 0.9θ 0.5, which were not resolved by direct observation. the width f (α, ϕ)θ 0.5 was assumed to be the same at both corners. when constructing the solution, we used preliminary information about the object under study as consisting of a group of small-sized sources. as gnm(α, ϕ), functions close to gauss functions were chosen to represent the solution, the maxima of which were located in the centers of squares with sides ∆α, ∆ϕ equal to θ 0.5/n, into which the region was successively divided. the best solution with minimal errors was obtained at n = 3, i.e. when dividing the area ω into 9 squares. step ∆α corresponded to the desired angular resolution. as a result, all small-sized sources were resolved and their position was found with good accuracy. since the inverse incorrect problem was solved, small errors appeared in the form of false sources with a small amplitude, which should be ignored when analyzing solutions. fig. 3.4. restored image of a two-dimensional source copyright© 2021 assa. adv syst sci appl (2021) achieving angular superresolution of control and measurement systems 111 4. two-beam method further increase of the degree of achieved super-resolution is possible based of of more complex measuring and signal processing. in [24] considered the use of interpolation to find solutions with superresolution, in [19] described the use of regression methods for the analysis of measurement data in [28] using extrapolation methods, in [26] described the application of the developed technique in the use of uwb signals. consider another possibility to increase the degree of superresolution when using active goniometer systems, i.e. systems that emit probing signals. let signal reception and radiation carried out by two independent scanning dptransmitting fe(β) and receiving dp fr(α). this, for example, can be achieved in the radio wave range using a smart antenna. then, the relationship of the values i, u, and the scanning dp is expressed as the convolution integral of two variables α and β, which correspond to the positions of the maxima of the transmitting and receiving rays: u(α, β) = ∫ ω fe(α − ϕ) fr(β − ϕ) i(ϕ) dϕ. (4.10) the amount of information obtained by scanning with two beams (4.10) significantly exceeds the amount obtained by single-beam scanning (2.1). this allows us to expect to obtain a more accurate solution of the inverse problem (3.2)-(3.5) with a higher achievable level of superresolution. one of the effective methods of signal processing (4.10) is the preliminary integration of the scanning angle β of the receiving or α emitting antenna. the larger the chosen area of integration λ, the higher the stability of the obtained approximate solutions. however as λ increases, the snr decreases in areas close to the boundary. therefore, the integration region should be defined as the sector of angles within which the snr value is sufficient to obtain stable solutions. usually, the area of λ is noticeably larger than the area of ω. integrating (4.10), for example, by the angle β we find v(α) = ∫ λ u(α, β) dβ = = ∫ ω fe(α − ϕ) i(ϕ) ∫ λ fr(β − ϕ) dβ dϕ = = ∫ ω fe(α − ϕ) i(ϕ) f(ϕ) dϕ = ∫ ω fe(α − ϕ) ĩ(ϕ) dϕ, (4.11) where f(ϕ) = ∫ λ fr(β − ϕ) dβ, ĩ(ϕ) = i(ϕ) f(ϕ). (4.12) the resulting integral equation (4.11) for the introduced function ĩ(ϕ) formally coincides with the original one (2.1). therefore, processing of the results may be carried out using all techniques and algorithms applicable to (2.1). the solution (4.11) in the form of ĩ(ϕ) allows using (4.12) to easily express the desired distribution i(α). numerical studies on models have shown that the conditionality numbers of matrices of type (3.4)-(3.5) obtained for the problem (4.11) are significantly smaller than those of matrices g from (3.4). this indicates increased stability of the resulting solutions. in addition, pre-integration (4.11) reduces the role of random rapidly oscillating noise components of the signal, and thus also increases the stability of solutions. increasing resiliency allows us to use more functions when presenting solutions (3.2) and, as a result, to get a higher resolution. copyright© 2021 assa. adv syst sci appl (2021) 112 b. lagovsky, e. rubinovich the disadvantage of pre-integration is the increased signal processing time. to increase the processing speed, instead of integrating by α and β, we can set a selection of angle values (αp, βq), p, q = 1, 2, . . . ,n, covering the entire scan area λ. then, instead of (4.11), using (3.2), we get v(α, β) = n∑ m=1 bm ∫ ω fe(αp − ϕ) fr(βq − ϕ) gm(ϕ) dϕ. (4.13) by entering the notation wpq = u(αp, βq), φpqm = ∫ ω fe(αp − ϕ) fr(βq − ϕ) gm(ϕ) dϕ, (4.14) we arrive at the slae with respect to the coefficients bm in the form wpq = m∑ m=1 φpqm bm, p, q = 1, 2, . . . ,n. (4.15) the system (4.15) is overridden. its solution is sought by standard algorithms in the rootmean-square approximation, which reduces the role of random components in the received signal. the solutions of the system (4.15), as shown by numerical experiments on a mathematical model, are not inferior in stability to solutions based on integration (4.11),(4.12). at the same time, the signal processing speed based on (4.13)-(4.15)is noticeably higher, which allows the algorithm to be used in real-time. fig. 4.5 illustrates the differences in the quality of solutions to inverse problems by the one-and two-beam method with a low snr of 11 db. the true source of the reflected signal was two separate small-sized sources with a distance between them of 0.7θ 0.5, which were not resolved by direct observation (the solid curve 1). step functions were chosen to represent the solution i(α). the solution based on the single-beam method is presented using the dashed curve 2. it should be considered inadequate: the image has little to do with the true sources. at the same time, the solution based on the two-beam method (the solid curve 3) allowed us to resolve the sources and correctly localize their location. the resulting false objects are characterized by a small amplitude of the reflected signal, and they should be ignored when solving such inverse problems. thus, the two-beam method of object image reconstruction reduces the role of random components, which ultimately manifests itself as an additional regularizing factor that increases the stability of solutions to the inverse problems under consideration. this is fully inherent in two-dimensional problems. when scanning the two-dimensional region θ with the transmitting and receiving beams at the angles (α, ϕ) and (γ, β), respectively, the received signal is obtained in the form: u(α, ϕ, γ, β) � n∑ n,m=1 bnm ψnm(α, ϕ, γ, β), ψnm(α, ϕ, γ, β) = = ∫ ω fr(α − α′, ϕ − ϕ′) fe(γ − α′, β − ϕ′)gn(α′)gm(ϕ′)dα′dϕ′. (4.16) copyright© 2021 assa. adv syst sci appl (2021) achieving angular superresolution of control and measurement systems 113 fig. 4.5. single and double-beam solutions at high noise levels using the representation (3.7), the coefficients bnm, which provide a minimum of the rootmean-square deviation of the sum in (4.16) from the received signal, are found as solutions of slae: ∫ θ u(α, ϕ, γ, β) ψnm(α, ϕ, γ, β) dα dϕ dγ dβ = = n∑ n,m=1 bnm ∫ θ ψ jk(α, ϕ, γ, β) ψnm(α, ϕ, γ, β) dα dϕ dγ dβ, j, k = 1, . . . ,n. to increase the processing speed, as previously for one-dimensional tasks, we can set a selection of angle values that cover the entire scan area instead of integrations. fig. 4.6,4.7 show the solution to the problem of restoring the image of a two-dimensional object. the rectangular coordinate system (α, ϕ) is used. a complex object consists of four identical small-sized objects located at the corners of a square with a side of 0.6θ 0.5. fig. 4.6 shows their location in the θ area. in the same figure, the u(α, ϕ) signal received during a single-beam scan is shown in the form of a grid. direct observation, as well as the processing u(α, ϕ), did not allow us to obtain a stable solution in the form of an image of an object with a super-resolution. the use of a two-beam method of measurement and signal processing provided an image of the object. fig. 4.7 shows the results of the mentioned signal processing in the region ω when using step functions. the decision based on of a two-beam method has allowed us to resolve all four sources. their angular position is correctly found and localized. the complex objects that arise when solving the inverse problem have a small amplitude and can be considered as angular noise. when solving problems of this kind, they should be filtered out, for example, using threshold devices. copyright© 2021 assa. adv syst sci appl (2021) 114 b. lagovsky, e. rubinovich fig. 4.6. the location of individual elements of a complex object and the angular distribution of the received signal 5. conclusion the developed algebraic methods of signal processing allow us to form approximate images of complex one-and two-dimensional objects with super-resolution. the achieved angular resolution increases when using a priori information about the solution reaches values 5-10 times higher than the rayleigh criterion. images are restored with relatively small errors in the values of the intensity and angular positions of individual elements of objects. the results of numerical studies have shown that the minimum required for obtaining a stable solution with a superresolution of snr is 12-16 db for simple algebraic methods, which is significantly less than for known methods. modifications of algebraic methods based on two-beam scanning, targeted selection of functions chosen to represent the solution, and extensive use of a priori information about the solution allow us to obtain a stable solution with a superresolution in some cases up to an snr of 7-9 db. in an alternative interpretation, the modified algebraic methods of constant snr allow increasing the level of the achieved superresolution. the created high-speed algorithms for algebraic methods allow to process of signals in a real-time mode. acknowledgment the reported study was partially supported by rfbr research project no. 20-07-00006. copyright© 2021 assa. adv syst sci appl (2021) achieving angular superresolution of control and measurement systems 115 fig. 4.7. restored image of the object in fig. 4.6 references 1. lagovsky b., rubinovich e. (2020). increasing the angular resolution of control and measurement systems in signal processing. proc. of the 7th int. conf. on control, decision and information technologies (codit), 496–500. 2. kasturiwala s., ladhake s. (2010). superresolution: a novel application to image restoration. j. comput. sci. inf. technol, (5), 1659–1664, . 3. uttam s., goodman n. (2010). superresolution of coherent sources in real-beam data. ieee trans. aerosp. electron. syst., 46(3), 1557–1566. 4. park s., park m., kang m. (2003). superresolution image reconstruction: a technical overview. ieee signal process. mag., 20(3), 21–36. 5. kulikov n., dung n. (2018). analysis of noise immunity of reception of signals with multiple phase shift keying under the influence of scanning interference. russian technological journal. 6(6), 5–12. 6. waweru n., konditi d., langat p. (2014). performance analysis of music root-music and esprit doa estimation algorithm. int. j. electr. comput. eng., 1(8), 209–216. 7. rao b., hari k. (1989). performance analysis of root-music. ieee trans. acoust. 37(12), 1939–1949. 8. kim k., seo d., kim h. (2002). ecient radar target recognition using the music algorithm and invariant features. ieee trans. antenn. propag. 50(3), 325–337. 9. lavate t., kokate v., sapkal a. (2010). performance analysis of music and esprit doa estimation algorithms for adaptive array smart antenna in mobile communication. proc. of the 2nd int. conf. on computer and network technology. 308–311. 10. almeida m., figueiredo m. (2013). deconvolving images with unknown boundaries using the alternating direction method of multipliers. ieee trans. image process. 22(8), 3074–3086. copyright© 2021 assa. adv syst sci appl (2021) 116 b. lagovsky, e. rubinovich 11. guan j., yang j., huang y., et. al. (2014). angular superresolution algorithm based on maximum entropy for scanning radar imaging. ieee int. geoscience and remote sensing symp. 3057–3060. 12. gamboa f., gassiat e. (1996). sets of superresolution and the maximum entropy method on the mean. siam j. math. anal. 27, 1129–1152. 13. borgiotti g., kaplan l. (1979). superresolution of uncorrected interference sources by using adaptive array technique. ieee trans. antenn. propag. 27(11), 842–845. 14. capon j. (1969). high-resolution frequency-wavenumber spectrum analysis. proc. ieee 1969. 27(1), 1408–1418, . 15. stoica p., sharman k. (1990). maximum likelihood methods for direction-of-arrival estimation. ieee trans. acoust. 38(7), 1132–1143, . 16. li l., hurtado m., feng x., et. al. (2018). a survey on the low-dimensional-model-based electromagnetic imaging. found. trends signal process. 12(2), 107–199, . 17. lagovsky b., samokhin a., shestopalov y. (2016). increasing effective angular resolution measuring systems based on antenna arrays. proc. 2016 ursi int. symp. electromagnetic theory (emts). 432–434, . 18. lagovsky b., chikina a. (2017). superresolution in signal processing using a priori information. piers proc. progress in electromagnetics research symposium (piers 2017). 944–947. 19. lagovsky b., samokhin a., shestopalov y. (2019). regression methods of obtaining angular superresolution. conference paper 2019 ursi asia-pacific radio science conference (ap-rasc). 236–239. 20. morse p., feshbach h. (1953). methods of theoretical physics. science/engineering/math. ny: mcgraw-hill. 21. lagovsky b., samokhin a. (2013). image restoration of two-dimensional signal sources with superresolution. piers proc. progress in electromagnetics research symposium (piers 2013). 315–319. 22. lagovsky b., samokhin a., shestopalov y. (2018). creating two-dimensional images of objects with high angular resolution. ieee asia-pacific conf. on antennas and propagation (apcap). 114–115. 23. lagovsky b. (2012). superresolution: data mining. piers proc. progress in electromagnetics research symp. (piers 2012-moscow). 993–996. 24. lagovsky b. (2012). image restoration of the objects with superresolution on the basis of spline interpolation. piers proc. progress in electromagnetics research symp. (piers 2012-moscow). 989–992. 25. lagovsky b. (2009). algorithm for the determination of targets coordinates in structure of the multiple targets with the increased effective resolution. piers proc. progress in electromagnetics research symp. (piers 2009). 1642–1645. 26. lagovsky b., samokhin a., shestopalov y. (2017). increasing accuracy of angular measurements using uwb signals. proc. of the 11th european conf. on antennas and propagation (eucap). 1083–1086. 27. lagovsky b., samokhin a. (2017). superresolution in signal processing using a priori information. proc. of the int. conf. electromagnetics in advanced applications (iceaa). 779–783. 28. lagovsky b., samokhin a., shestopalov y. (2015). superresolution based on the methods of extrapolation. piers proc. progress in electromagnetics research symp. (piers 2015-prague). 1548–1551. copyright© 2021 assa. adv syst sci appl (2021) introduction problem statement algebraic methods of solution two-beam method conclusion microsoft word 1167 adv syst sci appl 2022; 04; 233-241 published online at http://ijassa.ipu.ru. estimation of the posterior probabilities of classes by the approximation of the anderson discriminant function valery zenkov v.a. trapeznikov institute of control sciences of ras, moscow, russia email: zenkov-v@yandex.ru abstract: the approximation of the anderson discriminant function at a given point in the feature space of two classes using a supervised training sample allows estimating the posterior probabilities of classes as easily as converting fahrenheit degrees to celsius degrees. these probabilities make it possible to further solve the classification problem with subjectively specified costs of classification errors and criteria. the training sample is simply converted to a regression analysis sample by replacing the class numbers with differences in error costs. a nonparametric method for approximating the discriminant function at a point is proposed. it does not require the specification of an approximation function at a point more complex than a linear one. examples are given. keywords: machine learning, anderson discriminant function, approximation, class posterior probability, weighted least squares. 1. introduction the posterior probabilities (pop) of the classes are exhaustive information for solving the problem of object classification. they allow attributing a classification object to a particular class, taking into account different matrices of the costs of classification errors, set by the user not only for the entire task as a whole, but also for some categories of classified objects separately, for example, by the status of the classified object. we solve the problem in two steps. at the first stage, for the object of classification, we find the pop of the classes objective data. at the second stage, we assign the object to one or another class, taking into account the subjective preferences of the user: different the classification error costs, the method of making a decision taking into account some restrictions, etc. in connection with the importance of the pop of the classes for solving the classification problem and for interpreting the decisions made when classifying objects, add-ons to some already existing classification methods are used. so, for the support vector machine, the platt calibrator is used [1], which estimates the pop of the classes by the distance of a point in the class feature space to the boundary between classes. in this case, a hypothesis is put forward about the form of the dependence of the pop of the classes on the distance of the point to the boundary between the classes, and then the parameters of the hypothetical dependence are found using the maximum likelihood method using the training set. the anderson discriminant function (adf), according to the sign of which a point in the class feature space belongs to one of two classes, is a regression dependence. it is convenient in that, according to its approximation, obtained by the weighted least squares method from a supervised training set and by the cost of classification errors used to determine it, we calculate pop classes as simply as we get degrees celsius from degrees fahrenheit. we define adf from the bayesian solution of the classification problem [2], obtained by theodore wilburne anderson. adf is the difference between the two functions of average losses from classification errors in the space of attributes of the two classes at given costs of losses from classification errors. 234 v. zenkov copyright ©2022 assa adv. in systems science and appl. (2022) we approximate the adf using a supervised training set of the classification problem, transforming it into a regression analysis set by replacing the class numbers (labels) with the corresponding differences in the cost of classification errors. when approximating the adf at a point, there is no need to specify the form of the approximating dependence in the class feature space. it is enough to construct a linear approximation of the adf by the closest points to a given point, or by the nadaray-watson method [3] to find the adf value at this point. the pop estimates of the classes at a point based on the adf approximation at this point do not depend on the choice of error costs used to determine the adf. you can set the error costs equal, for example, 1 and 1, or not equal, for example, 2 and 5. the adf and approximations at the point will be different, but the pop estimates at the point will be the same. this is the main property of adf. the classification error costs are used to transform the supervised training set into a regression analysis set. class numbers in the first case are replaced by -1 and +1 in the first case and by -2 and +5 in the second case. due to the independence of the estimates of the pop of the classes on the choice of the costs of classification errors, the classification error can always be set to 0.5 and 0.5. the method for obtaining pop estimates from the adf approximation does not depend on the class imbalance in the training sample. class imbalance in the training set occurs when the number of points of one class is much less than the number of points of another class. class imbalance makes it difficult to detect objects of a small class if classification methods are used that are not based on pop estimations and costs of classification errors. the method for solving the classification problem at a point in the feature space using a supervised training set, when we use the entire training set for each classified point, is nonparametric, and the resulting model is a nonparametric model. 2. anderson discriminant function we define adf as the difference of two functions of average losses of classification errors in a problem with two classes:        12 1 2 1 2|, , , –k x k kf x c g x c g x c m c c   , (2.1) where         1 12 2 21, 1 1 , , 1| |g x c c p x g x c c p x   there are average losses at the point, if the point is classified as class 1 or, respectively, class 2; ijc is the cost of the error when a point from a class j is mistakenly referred to a class i. 0, # , 0ij iic i j c  . p(k|x) is the pop of the class k at the point x,    1 2 1| |p x p x  ,       ,/| |kp k x p p x k p x    | k kp x p p x k  , pk is the prior probabilities of classes, p(x|k) is the conditional distributions of class features; k is the class numbers 1 or 2;  |k xm  is the mathematical expectation of at a point. if  12 , 0f x c  , then the point referred to class 1, otherwise to class 2. if  12 , 0f x c  , then the point referred to class 1, otherwise to class 2. 3. adf properties 3.1. adf is by definition a regression function. to transform the training set of the supervised classification problem into a regression analysis set, due (2.1), on the set, the class number 1 should be replaced by -c21 and the class number 2 by c12. estimation of the posterior probabilities of classes… 235 copyright ©2022 assa. adv. in systems science and appl. (2022) 3.2. the first class pop and the adf defined for given c12 and c21 are identically related       12 12 12 211 – , /|p x c f x c c c  . (3.1) corollary 3.1: if we set the cost of errors under the condition 12 21 1c c  , that it does not lead to a loss of generality, then the equal (3.1) is simplified  121 * , * ( | )p x p f x p  (3.2) and on the boundary between the classes, where    12 12, * | 0, * 1f x p p p x с   and 211 *p c  . the pop of the first class on the class boundary is p*. corollary 3.2: the pop of the classes at the point x are not addicts on for which * 0p  , or c12 and c21, the adf is defined. for the formal proof, it suffices to substitute expression (2.1) for adf in (3.1) or (3.2). corollary 3.3: when choosing the error costs c for classifying a point by the pop of the classes obtained at the first stage of solving the classification problem, at the second stage there will be indistinguishability of classes in the feature space if     12 12, 0 v )( , 0x xmin f x c max f x c  . (3.3) in this case, all points will need to be assigned to the same class. 4. adf approximation methods the adf can be approximated in at least three ways: – to approximate the adf at a point in the class feature space to obtain estimates of the pop of the classes at this point according to (3.1) or (3.2). this is a nonparametric model for solving a problem. – to approximate the adf in the entire area of the training set, having specified the type of the approximating function and choosing the approximation method. for this parametric model, a method is proposed for approximating the adf in the initially unknown neighborhood of its zero values, to solve more accurately the classification problem for the selected type of approximating dependence and the given costs of classification errors and not pursue the goal of evaluating the pop of the classes. – to approximate several adf for the given pop values of the first class at the boundary between the classes, for example, 0.1, 0.3,…, 0.9 and for a given type of approximating dependence. this is a variant of a series of adf approximations. for a given point, the pop of the classes is found by the method of interor extrapolation over neighboring approximations of the adf, between which the point is located. the nonparametric model is slower it uses the entire training set in the test stage. parametric ones work faster, but they require the choice of the type of the approximating function for adf in the training stage. 4.1. adf approximation at point this version builds a nonparametric model. its distinctive feature is the need to use all or part of the supervised training sеt to estimate the pop of the classes at each given point. in [4], to estimate the adf at a point, we select such costs of classification errors for which the adf approximation at a given point takes zero value. but more preferable, in our opinion, 236 v. zenkov copyright ©2022 assa adv. in systems science and appl. (2022) is the adf approximation by a plane (at least by a second-order polynomial) in the vicinity of a given point. the criterion for approximating the adf at a given set point xj is the weighted root mean square error. at the training stage, for each set point xj used as a test one, there is a vector of coefficients λj approximating the adf in the form (1, xj) λj for given values of the parameters w and s of the weight function, for example, an exponent: 11, 2 2( ) [ ) ] ( || || ), = min – – (1, exp j n nj k k n j n n n s jn n j q x c c x w xx      , (4.1) where the row vector (1, xn), in addition to the row vector of features xn of the training set of components, contains a component one. n is the number of rows in the set without row xj at the training stage, as required by the loo overfitting method (leave-one-out). kn is the class number 1 or 2 in the set line n. xn is the feature vector in the set line n. 0w  is a weighting factor that sets the rate of decay of the weighting function (in this case, the exponent is from the distance of the training set point to a given point). s is an exponent to which the distance from the sampling point to a given point is raised. 1 2 – n n k k c c is the estimate of the adf in the set row n, obtained by replacing the class number kn, 1 or 2, in the row by the difference in error costs determined by the given p*. having obtained the vector of coefficients λj by minimizing criterion (4.1) for the given parameters of the method w and s, we calculate the estimate of the adf at the point xj, and, using (3.1), (3.2), we find the estimate of the pop of the first class at the point xj. if the number of classes 2k  , then similarly we obtain the pop of the class at this point, using the method one class against all the others. the best values of the parameters w and s are selected, for example, according to the minimum criterion 1 12 2 21 1( )r n c n c n  (4.2) where n1 is the number of points of the first class, erroneously assigned to the second class, n2 is the number of points of the second class, erroneously assigned to the first class. often s = 1 and the value of the only parameter w is selected. the type of the weighting function is not necessarily exponential. 4.2. adf approximation in the neighborhood of zero values in some cases (a large volume of the training set, the presence of a priori information about the conditional distributions of features of classes, or about the type of surface separating the classes), it is more expedient to build a parametric model. it can be in the form of a discriminant function of a given form with unknown parameters determined from the training set. to classify a point, we can use the bayes formula to recalculate the reconstructed conditional distributions of class features to the pop of the classes at the specified point and, taking into account the costs of classification errors, assign the point to a particular class. we propose in [5] a heuristic method for constructing a parametric model of the classification problem based on a supervised training set, used for the approximation of the discriminant function in the vicinity of its zero values for given costs of classification errors. in this case, we usually do not talk about evaluating the pop of the classes. the supervised training set does not contain information about the position of the adf zero points. for given values of classification errors, there may be no boundary between classes and all points should be assigned to the same class. to solve the problem, a heuristic method [5] of approaching the boundary between classes is used. first, based on the given costs of classification errors, the training set is transformed into a set of the regression analysis problem by replacing the class numbers (labels) with the corresponding differences of the costs of classification errors (2.1). then, several iterations of solving the problems of approximation of adf by the weighted least squares method are estimation of the posterior probabilities of classes… 237 copyright ©2022 assa. adv. in systems science and appl. (2022) started. the weight function from the distance of the point to zero on the previously obtained adf approximation is used as the weights of the set points. the measure of the distance of a point to the zero of the previous approximation is the module of the previous approximation of the adf at this point. solution criterion at each step with an exponential weighting function: 2 11 2 1 min – ]– ' exp ( ) {[ ( ) ( | ( ) | )}' i n n s i k k i n i n i n n nq c c x w x         , (4.3) where i is the iteration number, 1,2,...,{ }i i . at the first step, the problem is solved without a weighting function  0 0w  ; at subsequent steps, the weighting function gives more weight to points closer to the zero region of the previous adf approximation. i is a given number of iterations i  . wi is a given weighting factor at a step s is a given exponent, usually 1s  . the weighting factor in iterations can be constant or change, for example * , {0,1,..., }iw w i i i  . for each iteration, the losses are found from the obtained adf approximation (4.3). the best value of λ is the vector to which the lower losses correspond (4.2). 4.3. estimation of pop of the class at a point from a series of adf approximations in this parametric method [6], based on a supervised training set, we construct several approximations of the adf specified for multiple pop of the first class at the boundaries between classes. this completes the training process and the training set is no longer used. the constructed series of approximations can be used to estimate the pop of the first-class at a given point in the feature space by two methods: the interpolation method and the smoothing method. in the interpolation method for estimating the pop of the first class at a given point by the signs of the adf approximations in the series, at this point, there is a pair of neighboring approximations, between which the point is located. if a pair is found, then the interpolation method yields an estimate for pop of the first class. a measure of the distance to neighboring boundaries and known by the construction of pop of the first class on these boundaries is the module of the values of neighboring adf approximations at a given point. if a given point is outside the limits of the adf approximations, then the estimate of the pop of the class is found by extrapolation from the first or last adf approximation in the series, checking the result for non-negativity and not exceeding one. in the smoothing method, the weighted pj* values used to obtain the j-approximations of the adf are used to estimate the pop of the classes at a given point. as a measure of the distance of a point to the boundary j between classes, the modulus of the value of the japproximation of the adf at a given point   , * j jf x p is used. as a weighting function, for example, the exponent is used * 1 ( ) ( | ,( exp f , ) ) |j j j j jw x w x p    , (4.4) where w is the parameter setting the smoothing properties of the method. the weighted estimated pop at a point using a series of adf approximation and (4.4) will be * ( | ) ( ) ( )1 / j j j j jp x p w x w x   . (4.5) the best value of w can be selected according to (4.2), performing classifications according to the pop of the classes (4.5). 238 v. zenkov copyright ©2022 assa adv. in systems science and appl. (2022) 5. examples 5.1. approximation of the adf at a point we present the case of solving the problem by minimizing (4.3) and (4.2) in fig. 5.1 with normal conditional distributions of two-dimensional features of three classes. one of them, class a, is located between classes b and c. class a is designated by the number 1, classes b and c will be combined into one class and designated by the number 2. class middles:       0, 0 , 3, 0 , 3, 0a b cm m m    .the covariance matrices of the classes are the same. the prior probabilities of the classes 0.5ap  , 0.25b cp p  . the set generated by this parameter contains 120 points. the figure shows the theoretical densities of conditional distributions of class features, theoretical adf and pop of the first class, as well as estimates of pop at given points, obtained from 1s  and 3w  (4.2). fig. 5.1 the root-mean-square errors (rms) in the estimation of pop of the first-class about with concern to the theoretical ones were 0.09 and 0.18, respectively, for the linear and polynomial approximation of the second-order of the adf. the linear approximation of adf turned out to be more accurate in this set. the estimates of pop of the first class at the given points are located near the curve of pop of the first class, obtained from the estimates of the parameters of the normal laws of conditional distributions of features (the dashed curve next to the solid theoretical curve of pop of the first class). since there are two features, the flat figure shows a section by a plane passing along the axis of the values of the first feature and the ordinate. estimation of the posterior probabilities of classes… 239 copyright ©2022 assa. adv. in systems science and appl. (2022) fig. 5.2 fig. 5.2 shows an example with a one-dimensional feature distributed according to a uniform law of probability distribution. classes a, b and c have averages 0, 3, 3a b cm m m    . the limits of maximum deviations from the average in both directions are the same and equal to 2. the classes are intersected in the feature space. the prior probabilities of the classes are the same as in the previous case. there are 120 points in the training set, 1s  and 3w  . visually, the estimates of pop of the first class by a second-order polynomial in being more accurate. explanations for fig. 5.1 and fig. 5.2. the sampling points of the three classes a, b, and c are located at levels p* and 1 *p . the conditional distributions of features of classes 1 and 2 (one against combining the other two) are the dashed lines  |1p x and  | 2p x in the figures. the solid thin line is the distribution density of x, p(x). the inscription adf denotes the position of the theoretical adf (solid thick line with convexity downward).  |1p x is the theoretical pop of the first class (solid thick line with a convex upward). close to theoretical dashed lines is adf and pop constructed from the training set, for the known conditional distributions of class attributes and the given costs of classification errors, * 0.5p  . the adf constructed for the training set in fig. 5.2 is not shown. it is a mirror image of the pop constructed for the training set (dashed line). round points near pop of the first-class – estimates of pop by approximating adf at given points; crosses — estimates of pop using a second-order polynomial to approximate the adf. 5.2. approximation of the adf in the vicinity of zero values when choosing the type of adf approximating dependence, one can use the features, according to their correlation coefficients with the desired value. so in the example [5] with the dimension of the feature space of 217, the training set consisted of only 252 lines. three were selected with a correlation coefficient estimates of the desired value not lower than 0.64 and with a cross-correlation, not higher than 0.66. the classification error according to the adf approximation in the vicinity of the zeros according to the linear approximation was 6.8 %, according to the second-order polynomial – 4.8%. for comparison, the classification 240 v. zenkov copyright ©2022 assa adv. in systems science and appl. (2022) method based on the hypothesis of conditional normal distributions of the selected features of the classes, according to the polynomial and according to the linear model, was 10.3% and 9.9%, respectively. we made [7] a comparison on 15 examples of the quality of solving classification problems by the adf approximation method in the vicinity of zeros and the support vector machine method (svm) for linear adf approximation and approximation by a second-order polynomial at equal costs of classification errors. in 14 examples, the first method gave a lower classification error than the second. a few examples may not be a sufficient basis for judging the advantages of one method over another. 5.3. estimation of the pop of the first class at a point from a series of the adf approximations as a model example, fig. 5.3, [6] the classification problem with two classes with normal conditional distributions of class attributes in a two-dimensional space with different covariance matrices was considered. the average values of the features of the first class  1 0, 0m  . the covariance matrix is 1 ((1, 0); (0,1))s  . for the second class:  2 0, 2m  , 2 ((4,1); (1,1))s  . the prior probabilities of the classes are 1 0,6p  , 2 0,4p  . fig. 5.3 based on the initial data and the generated set of 500 points, the pop of the first class was calculated in the following ways: 1. exact according to the known distributions according to the bayes formula; 2. sample based on sample estimates of distribution parameters using bayes' formula for calculating pop of the first class; 3. inter-extrapolation for the series of the adf approximated by a second-order polynomial; 4. by the method of smoothing according to the series of adf approximations. a series of the adf approximations is built for p* values from 0.95 to 0.05 with a step of 0.1. circles are points of the first class, crosses are of the second. the root-mean-square error estimate of the pop of the first class for the training set by the interpolation method for a series of adf approximations was 0.055. the error in its estimation by the smoothing method was 0.06. the error in estimating pop of the first class based on sample estimates of feature distributions and using bayes' formulas was 0.02, i.e. three times less. but in justification of estimation of the posterior probabilities of classes… 241 copyright ©2022 assa. adv. in systems science and appl. (2022) other methods, it should be remembered that they do not require knowledge of the laws of distribution of the features of classes. 6. conclusion approximation of the anderson discriminant function based on a supervised training set allows you to simultaneously find estimates of the posterior probabilities of classes of classified objects. the posterior probabilities of the classes are comprehensive information for solving the problem of classification according to various criteria. the supervised training set of the classification problem is transformed into a training set of the regression analysis problem by replacing the class numbers in the training set with the corresponding differences with the costs of classification errors, set arbitrarily at the training stage to determine the approximations of the anderson discriminant function. at the second stage of solving the problem, according to the estimates of the objective posterior probabilities of classes, the actual classification of the object is performed according to the subjectively specified classification criteria. methods for solving the problem related to the use of nonparametric and parametric approaches are proposed. with the nonparametric approach, to approximate the anderson discriminant function at a point in the class feature space representing the classified object, a training set is used for each classified point, which may be unacceptable for larger set sizes and with certain performance requirements. for a non-parametric approach a rather linear approximating function. in parametric approaches, weighted least squares are used to obtain a more accurate approximation of a function in the vicinity of its zero values. in another parametric approach, a series of approximations of the anderson discriminant functions are constructed from a given set of a posteriori probabilities on the class boundaries. then, for the classified point, the posterior probabilities of the classes are found using the interpolation method or the smoothing method. methods are demonstrated by examples. references 1. platt, j. (2000). probabilistic outputs for support vector machines and comparisons to regularized likelihood methods, advances in large margin classifiers, 10(3), 61–74. 2. anderson, t.w. (2003) an introduction to multivariate statistical analysis, 3 rd. ed. new york, ny: john wiley & sons. 3. hardle, v. (1993) applied nonparametric regression. moscow, russia: mir. 4. zenkov, v.v. (2018). estimating the probability of a class at a point by the approximation of one discriminant function, autom remote control, 79(9), 1582–1592, https://doi.org/10.1134/s0005117918090047. 5. zenkov, v.v. (2017). using weighted least squares to approximate the discriminant function with a cylindrical surface in classification problems, autom. remote control, 78(9), 1662–1673, https://doi.org/10.1134/s0005117917090107. 6. zenkov, v.v. (2019). estimation of the posterior probability of a class from a series of discriminant anderson functions, autom remote control, 80(3), 447-458. https://doi.org/10.1134/s0005117919030056. 7. zenkov, v.v. (2020). applying an approximation of the anderson discriminant function and support vector machines for solving some classification tasks, autom remote control, 81(1), 118–129, https://doi.org/10.1134/s0005117920010105. microsoft word 988 article text, copyedited.docx adv syst sci appl 2021; 02; 20-28 published online at https://ijassa.ipu.ru. using patent landscapes for technology benchmarking: a case of 5g networks dmitry kochetkov 1*, marat almaganbetov 2 1) ministry of science and higher education of the russian federation, moscow, russia e-mail: kochetkovdm@hotmail.com 2) lexisnexis eastern europe, moscow, russia e-mail: marat.almaganbetov@lexisnexis.com abstract: patent landscape reports (plr) have long been used in business, science and r&d. they enable data-driven decision making and minimizing the risks of strategic technology choices. the authors analyzed the patent landscape for 5g wireless communications using the patentsight® software product. among the companies, the clear leader by the patent portfolio index is samsung electronics; among countries, it is usa and china. however, patents of high technological relevance and/or commercial value may appear in small tech companies, or as a “by-product” of other corporate activities. the research results are of interest to researchers and practitioners involved in the development and implementation of 5g technologies. keywords: patent landscape, 5g, patent assets index, technological relevance, market coverage, competitive impact 1. introduction the patent landscape is a systematic study of patent documents that identifies and visualizes trends in business, science and technology. patent landscape reports usually focus on one industry, technology, or geographic region. patent landscape reports (plrs) support informed decision making and are designed to effectively address high-stakes challenges in various areas of technology, thereby increasing trust. through patent analytics and plr, these critical decisions can be made using evidence that provides informed choices and mitigates decision-making risks [13]. this patent study focuses on 5g technology. 5g stands for fifth-generation cellular wireless communications, and the initial standards for it were established at the end of 2017. since 2019, mobile operators have begun rolling out 5g networks around the world. the main advantage of the new networks is that they will have higher bandwidth, providing higher data transfer rates. thanks to the increased network bandwidth, mobile operators will be able to compete with isps providing cable internet services. besides, latest generation wireless networks open up a whole range of new applications in the internet of things (iot), augmented/virtual reality, m2m systems [9]. the study used patent analytics provided by lexisnexis (founded in 1977, is part of the largest transnational information holding relx group). lexisnexis ip is a global leader in information solutions and services to meet the needs of the intellectual property market, government agencies, and academia. the company aggregates data from 111 patent offices around the world [1]. * corresponding author: kochetkovdm@hotmail.com using patent landscapes for technology benchmarking: a case of 5g networks 21 copyright ©2021 assa. adv. in systems science and appl. (2021) 2. theoretical background 5g publications began to appear in 2014, and since then their number has grown steadily. we used the web of science® database, which is one of the most widely used reference databases for scientific periodicals. for information retrieval, we used the '5g' query with the category constraint 'telecommunications.' the request returned 7093 publications. fig. 2.1 shows the dynamics of the number of publications from 2014 to 2019. fig. 2.1. dynamics of 5g-related publications. data source: clarivate analytics web of science®. we visualized the 5g scientific landscape using vosviewer® as in [7] (fig. 2.2). it is a software that enables creation, visualization and analysis of network maps based on bibliographic data. vosviewer® also offers text mining functionality that can be used to construct and visualize co-occurrence networks of important terms extracted from a body of scientific literature [18]. vosviewer® pays special attention to the visualization of bibliometric maps. vosviewer's functionality is especially useful for displaying large bibliometric maps in a simple and efficient way. it is especially useful for maps containing at least a moderately large number of items (e.g., at least 100 items) [15, 16]. the software is based on the idea of association strength: (1) where sij is the similarity measure, cij is the co-occurrence of i and j, and ciсj the expected number of co-occurrences of i and j under the assumption that co-occurrences of i and j are statistically independent. sij takes values in the range [0, 1]. vosviewer® applies the link strength [14] by default. then, the program uses the mapping technique detailed by [17]. finally, vosviewer® assigns the network nodes to clusters; the clustering technique is presented in [19]. vos determines the locations of items in a map by minimizing , (2) subject to: . (3) 22 d. kochetkov, m. almaganbetov copyright ©2021 assa. adv. in systems science and appl. (2021) xi and xj denote the location of i and j respectively. bibliometric maps have two dimensions and rely on the euclidean distance. for the purposes of this study, we mainly used text mining functionality for constructing co-occurrence networks of terms extracted from metadata. fig. 2.2. co-occurrence analysis of 5g-related publications. data source: clarivate analytics web of science®. built with vosviewer®. the red cluster covers general issues of 5g network management, including network architecture, green radiotechnological aspects of network functioning, yellow management of network power parameters and the blue one relates to optimization of resource allocation. however, the picture would be incomplete without the patent landscape because scientific inventions enter everyday life precisely through intellectual property objects and their commercialization. patent landscapes are a widespread tool for analyzing the evolution of technological fields and identifying hot topics [2]. patent landscapes (also patent landscape reports (plrs) enable informed business&technology decisions [13]. however, there is no single definition of the patent landscape; it could be an overview of patent information for a specific technology or geographic area. the landscape usually strives to answer specific practical questions and present complex information about patenting activities in a transparent and intelligible form. industry has long used the patent landscapes to make strategic decisions about investments, research and development (r&d), and competitors. högskolan i halmstad. & niamba (2016) provide a good overview of patent indicators, subdividing them into four categories [5]: • performance indicators deal with the patent output of the analyzed entities (inventors or applicants) that are used to monitor the technological performance of companies/institutions and inventors/researchers [10]. • technology indicators analyze patent classifications, classifying each patent into one or more classes according to its technology area. • patent value indicators give an idea of the economic value of a project, which is determined by various factors such as family size, citation, and geographic coverage. using patent landscapes for technology benchmarking: a case of 5g networks 23 copyright ©2021 assa. adv. in systems science and appl. (2021) • collaboration indicators provide information about collaboration patterns between entities. technology benchmarking is used to compare the performance of organizations in terms of technology applications. benchmarking focuses on analyzing your own processes and then comparing your own performance with the performance of others [8]. patents are reliable kpis for comparison [6]. at the same time, patents are a lagging indicator of the company's research, since it is at least 18 months old (most patent offices around the world have a secrecy period of 18 months). in 2011, a new technology benchmarking methodology based on patent information [3] was presented. the main indicator in this methodology is the patent asset index, which is constructed based on the 'portfolio size', 'market coverage' and 'technology relevance' metrics. 'market coverage' is a measure of the extent of patent protection in global markets. 'technology relevance' is a citation-based indicator to assess the technological impact of patents, which eliminates systematic distortions of existing citation-based patent indicators. authors argued that the patent asset index offered a more accurate assessment of a firm's patent than other similar methods. we used this technique in our research, so we describe it in more detail in the next section. 3. data and methods the study used data on 4069 patent families associated with 5g technology. the search was carried out using the keywords “5g” and “fifth generation” with a filter by domain (ipc) h04w (wireless communication networks [2009.01]). patent family is a set of patents obtained in different countries for the protection of one invention (when the first application in the country priority then extends to other patent offices) [11, p.60]. the data source is patentsight®(“patent research and analytics products | lexisnexis patentsight®,” n.d.) from lexisnexis. this tool provides objective metrics of global technological power and influence, developed by german scientists headed by n. omland and h. ernst [3] (fig. 3.1). fig. 3.1. the patent asset index structure. source: [12]. 24 d. kochetkov, m. almaganbetov copyright ©2021 assa. adv. in systems science and appl. (2021) the patent asset index ™ (pai) takes into account both the quantity of actively protected inventions and their quality. the method was developed and tested in scientific research and has been used by leading companies in many industries for several years. the patent assets index is the overall strength of the patent portfolio. it is based on the technology relevance ™ and market coverage ™ of the company's patents to analyze the competitive impact of a portfolio. technology relevance is the worldwide citation derived from later patents, taking into account age, patent office practice, and technology area. market coverage is the size of the market protected by active patents and pending patent applications for a particular invention. finally, competitive impact ™ measures the strength of an individual patent, i.e. the relative commercial value of the patent. 4. results fig. 4.1 presents the analysis of portfolio size, patent assets index, and competitive impact or relative business value. the undisputed market leader is samsung electronics, which has the largest 5g patent portfolio with the highest patent assets index. the other side of the coin is the low portfolio selectivity. at the same time, we observed two companies with small patent portfolios, but extremely high value (at&t and interdigital). fig. 4.1. portfolio size (horizontal axis), competitive impact (vertical axis), and patent assets index (bubble size). created with patentsight®. samsung leads not only in the area as a whole, but in all sub-areas (level 4 of the ipc classification) (fig. 4.2). using patent landscapes for technology benchmarking: a case of 5g networks 25 copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4.2. the size of the patent portfolio by level 4 of the ipc classification. created with patentsight®. however, if we look at the indicator of technological relevance, then we will not see companies with the largest patent portfolios in the list (fig. 4.3). often, the most cited patents appear in small tech companies or as a "by-product" of other businesses (eg toyota). fig. 4.3. technological relevance of the patent portfolio. created with patentsight®. finally, in terms of the origin of patent families, two jurisdictions are leading by a noticeable margin usa and the prc (fig. 4.4). 26 d. kochetkov, m. almaganbetov copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4.4. geographical breakdown of patent families by jurisdiction of origin. created with patentsight®. 5. conclusion patent landscape reports (plr) have long been used in business, science and r&d. they enable data-driven decision making and minimizing the risks of strategic technology choices. in this study, we analyzed the patent landscape for 5g wireless networks. among the companies, the clear leader by the patent portfolio index is samsung electronics; among countries, it is usa and china. however, patents of high technological relevance and/or commercial value may appear in small tech companies, or as a ‘by-product’ of other corporate activities. the analysis shows that the 5g family of technologies has entered a phase of maturity, i.e. further developments will be incremental. at this stage, the market is already formed, the emphasis is shifting to commercialization and implementation. the new generation of mobile communication systems is already here [4]. however, the new generation of communication networks (6g) will be built on the existing scientific groundwork related to understanding both target applications and the most promising candidate technologies. historical analysis of technology development based on scientometric analysis and patent landscapes enables making informed strategic decisions for technological breakthroughs in the future. acknowledgements the study was financially supported by the russian foundation for basic research within the framework of the research project no. 18-00-01040 komfi ‘the impact of emerging technologies on the urban environment and the quality of life of urban communities.’ using patent landscapes for technology benchmarking: a case of 5g networks 27 copyright ©2021 assa. adv. in systems science and appl. (2021) references 1. about us. (2020). retrieved july 14, 2020, from lexisnexis website: https://www.lexisnexisip.com/about-us/ 2. clarke, n. s., jürgens, b., & herrero-solana, v. (2020). blockchain patent landscaping: an expert based methodology and search query. world patent information, 61(march). https://doi.org/10.1016/j.wpi.2020.101964 3. ernst, h., & omland, n. (2011). the patent asset index – a new approach to benchmark patent portfolios. world patent information, 33(1), 34–41. https://doi.org/10.1016/j.wpi.2010.08.008 4. giordani, m., polese, m., mezzavilla, m., rangan, s., & zorzi, m. (2019). towards 6g networks: use cases and technologies. 5. högskolan i halmstad., f., & niamba, c. n. (2016). patent bibliometrics and its use for technology watch. journal of intelligence studies in business, 6(1), 17–33. retrieved from https://ojs.hh.se/index.php/jisib/article/view/152/0 6. jain, r., tripathi, m., agarwal, v., & murthy, j. (2020). patent data analytics for technology benchmarking: r-based implementation. world patent information, 60(april 2019), 101952. https://doi.org/10.1016/j.wpi.2020.101952 7. kochetkov, d., vuković, d., sadekov, n., & levkiv, h. (2019). smart cities and 5g networks: an emerging technological area? journal of the geographical institute jovan cvijic sasa, 69(3), 289–295. https://doi.org/10.2298/ijgi1903289k 8. lewis, w. p., & samuel, a. e. (1995). technological benchmarking: a case study in product development by a small manufacturer. in a. rolstadås (ed.), benchmarking — theory and practice. ifip advances in information and communication technology (pp. 110–119). https://doi.org/10.1007/978-0-387-34847-6_13 9. looper, c. de. (2020). what is 5g? the next-generation network fully explained |. retrieved july 14, 2020, from digital trends website: https://www.digitaltrends.com/mobile/what-is-5g/ 10. oecd patent statistics manual. (2009). https://doi.org/10.1787/9789264056442-en 11. oecd science, technology and industry scoreboard 2001. (2001). https://doi.org/10.1787/sti_scoreboard-2001-en 12. patent research and analytics products | lexisnexis patentsight®. (n.d.). retrieved july 14, 2020, from https://www.lexisnexisip.com/products/patent-sight/ 13. trippe, a. (2015). guidelines for preparing patent landscape reports. 14. van eck, n. j., & waltman, l. (2009). how to normalize cooccurrence data? an analysis of some well-known similarity measures. journal of the american society for information science and technology, 60(8), 1635–1651. https://doi.org/10.1002/asi.21075 15. van eck, n. j., & waltman, l. (2010). software survey: vosviewer, a computer program for bibliometric mapping. scientometrics, 84(2), 523–538. https://doi.org/10.1007/s11192-009-0146-3 16. van eck, n. j., & waltman, l. (2014). visualizing bibliometric networks. in measuring scholarly impact. https://doi.org/10.1007/978-3-319-10377-8_13 17. van eck, n. j., waltman, l., dekker, r., & van den berg, j. (2010). a comparison of two techniques for bibliometric mapping: multidimensional scaling and vos. journal of the 28 d. kochetkov, m. almaganbetov copyright ©2021 assa. adv. in systems science and appl. (2021) american society for information science and technology, 61(12), 2405–2416. https://doi.org/10.1002/asi.21421 18. vosviewer visualizing scientific landscapes. (n.d.). retrieved august 28, 2019, from https://www.vosviewer.com/ 19. waltman, l., van eck, n. j., & noyons, e. c. m. (2010). a unified approach to mapping and clustering of bibliometric networks. journal of informetrics, 4(4), 629–635. https://doi.org/10.1016/j.joi.2010.07.002 microsoft word 1130 article text, copyedited.doc adv syst sci appl 2021; 04; 31-44 published online at https://ijassa.ipu.ru. the role government policy and supports play to stimulate economic growth jeffrey yi-lin forrest1*, abdou k. jallow1, chao wen2, juehui shi3, huan guo4 1) slippery rock university, slippery rock, pa, usa e-mail: jeffrey.forrest@sru.edu; abdou.jallow@sru.edu 2) ohio northern university, ada, oh 45810, usa e-mail: c-wen@onu.edu 3) angelo state university, san angelo, tx, usa e-mail: juehui.shi@angelo.edu 4)jianghan university, wuhan, china e-mail: guohuan.2007@163.com abstract: this paper studies the reason why policy tools might work, if they do, what the functional mechanism is, and why government’s policies and supports are necessary for stimulating economic growth. by employing systems thinking and methods and the logical reasoning as that commonly used in mathematics and natural science, this paper establishes 6 formal and generally true propositions on these related issues. at the conclusion, we provide a whole list of recommendations for policy makers and government officers in terms of when and how their implemented policies will lead to their desired outcomes. at the conclusion, this paper provides directions and open problems for future research. keywords: economic agent, feedback, market, momentum, policy tool, system 1. introduction fast development of communication technology in recent decades has prompted a score of leading nations to introduce policies and provide supports in order to transform their industries. to maintain their leadership in the globalizing world, these nations challenge themselves to elevate their industries from automated manufacturing to intelligent manufacturing, such as industry 4.0 [1]. hence, the following questions arise naturally: why are policy tools expected to work in an economic system? why are government policy tools and supports fundamentally necessary for stimulating economic growth? answers to these questions are both theoretically and practically important for all nations from around the world. developed nations, at least some of them, attempt to preserve their leading positions in the increasingly globalizing world economy. developing and underdeveloped nations with their miserable failures of modernization experienced in the past one hundred plus years [2] face the likelihood of tumbling further behind the developed nations economically, socially, and politically. although many scholars have worked on these and other related questions since before the time of adam smith, the relevant debates have come and gone in waves without achieving anything definite other than producing inconsistent and uncompromising suggestions [3.4]. other than filling a gap existing in the literature on whether or not governmental policies and supports are necessary in economic development, by addressing these questions and related issues, this work also tempts to answer rostow’s [4] call for an * corresponding author: jeffrey.forrest@sru.edu 32 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) appropriate methodology for such investigations in order to actually produce scientifically sound and dependable results. in particular, in terms of how policies actually work, wieczorek and hekkert [5] consider how policy instruments can improve the function of an innovation system. by considering which channel is likely to amplify the effects of monetary policy, auclert [6] empirically finds that the following three channels work the earnings heterogeneity channel from unequal income gains, the fisher channel from unexpected inflation, and the interest rate exposure channel from real interest rate changes. boeckx et al. [7] examine the effectiveness and transmission mechanism of the euro-system’s credit support policies. boeckx et al. [7] find successes of policies in stimulating credit flows from banks to the private sector. rebei [8] empirically looks at the crowding-in effect of government spending on private consumption and finds that in early times government spending increasingly crowded in private consumption, while this relation was reverted in recent times. cúrdia and woodford [9] investigate consequences of a variable credit spread optimal policy responses and find that a simple target criterion provides a good approximation to optimal policy. jaehrling et al. [10] provides insights on negotiations and outcomes of labor clauses across different stages of policy process, demonstrating the importance of alliances among local politicians, unions and employers. enriching the literature, this study demonstrates why governmental policies are important, when a government is able to align business objectives of a large proportion of individual economic agents, and what policy tools can help accomplish the goal. as for why governmental policies and supports are necessary in economic development, the literature is naturally divided into two camps: (i) yes, they are necessary; (ii) no, they don’t work generally. for example, andreoni and chang [3] provide a framework for strategically coordinating packages of interactive industrial policy measures. howell et al. [11] argue that the future of us competitiveness in international markets depends on realistic federal, industry, and company policies with regard to microelectronics. conversely, jomo [12] studies eight asian economies (i.e., japan, south korea, taiwan, singapore, hong kong, malaysia, thailand, indonesia), which achieved a statistically unlikely rapid economic growth between 1965 and 1990. although government interventions played a role, he notes that the gains from industrial policies are ambiguous. regarding how well the advanced manufacturing initiative for america’s future, a concerted us government effort, would harmonize with the obama administration, hemphill [13] finds that the answer is not well. comparing to what is summarized above, this study contributes to the literature by establishing several generally true results. for example, it shows why a nation needs to provide policy supports to maintain an extant economic growth momentum, when an implemented policy will positively affect the economic performance of manufacturing enterprises, and why different policies are needed to promote economic growth in different geographical regions. epistemologically speaking, all results established in this paper do not suffer from the constraints of dataand anecdote-based approaches, as commonly employed in the literature, and avoid the methodological pitfall experienced by rounds of debate on whether or not governmental policies are necessary for economic development since the time even before adam smith [3] and by rounds of studies on what was really underneath the occurrence of the industrial revolution [4]. in particular, the former respectively produced inconclusive and noncompromising reasons for and against the use of policies based on anecdotes and statistics. and the latter led to questionable lists of factors that were underneath the occurrence of the industrial revolution [2,4]. in comparison, results developed in this work are scientifically definite. for the following reasoning to flow smoothly, assume that each business firm is established to satisfy a market niche so that the firm is operationally maintained by a positive cash flow as a consequence of its business conducts in the marketplace. and by consumers, it means end users of market offering; by customers those firms that employ their inputs to the role government policy and supports play to stimulate economic growth 33 copyright ©2021 assa. adv. in systems science and appl. (2021) produce their outputs. when both consumers and customers coexist, they are jointly known as customers. the rest of this paper is organized as follows. section 2 studies the functional mechanism of policy tools. section 3 investigates why governmental policies and supports are necessary in economic development. section 4 concludes the paper with practical recommendations and open questions for future research. 2. functional mechanism of policy tools this section investigates why policy tools are expected to work in the focal economic system? 2.1. a systemic modeling of the economy by system, it means such a notion that models an organization or a structure as a collection of components or objects and associations among the objects. the components have been conventionally treated as isolated in classical sciences, while the associations help make the components into an organic whole (or system) [14]. hence, the concept of systems really appears everywhere in life. for example, each family is a system; each business firm is a system; … as a matter of fact, a main characteristic of the world is the systemness of various kinds of organizations (and structures). this realization explains why the concept of systems can be appropriately applied in investigations of business-related matters and issues, see [4,15,16]. symbolically, a system is an ordered pair with containing all the isolated objects of system and relations that connect the objects in into a whole [14]. for example, each business organization consists of a set of components, such as employees, properties, equipment, etc. and, these objects are connected together through particular relations, because of which the whole is acknowledged as an organization. in this junction, one should note that each object in can also be a system again, just as in the case that a family consists of people as its objects, and each person is a biological system, made up of various organs; each organ is again a system of even smaller parts, … at the same time, a relation in does not have to be numerical. for example, it will not be appropriate to express the relationship that two people and are brothers, unless one only likes to capture this relationship partially. given system , another system is said to be a partial system of , if is a subset of and for any relation there is relation such that that is, the relation of the system is the restriction of a relation in the system . a system is multi-leveled, if there is at least one object such that is a system, and there is at least one object such that is a system, …, where is known as a first-level object system, a second-level object system, … a system is said to be centralized if each object in is a system and there exists a system such that and for any distinct objects x, y ∈ m, say and , and , where and . the system c is called a center of the centralized system s. speaking in non-mathematical terms, a system is centralized, if and only if there is another system that is a partial system of every object system in . to address the question stated above, let us model the economy of the focal nation as a system that consists of all economic agents as component parts, such as consumers, families, 34 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) business firms, economic sectors, and markets, and associations that connect the component parts into a functional economy. in this model, the associations reflect how money, information and knowledge flow within the economy, how the magnitudes and types of supplies and demands constantly change, how consumer’s preference and tastes evolve, etc. each family is a system that consists of family members; each business firm is a system that is made up of members of various families; each economic sector is composed of many firms; and each market connects individual consumers and firms through exchanges of goods, information, money, etc. that is, our focal economy can be well modeled as a multileveled system. within this systemic model of the economy, the government represents a formal structure that supports the stable existence and smooth operation of the economic system with the top level of the government serving as the center, as seen in the concept of centralized systems. any slight change in the center affects most areas of the economic system. the formal structure of the government can be imagined as the three-dimensional hierarchy of individuals, such as consumers on the bottom, whose consumption drives all other layers of the economy, business firms as the next layer upward, each of which serves a segment of consumers, … and markets on the top, where eventual exchanges of information, knowledge, money, goods, etc., take place. the government guarantees the economy operates smoothly through laws, regulations, and related reinforcements. 2.2. the importance of governments as a multi-leveled, large-scale complex system, a nation’s economy surely satisfies the following theoretical fact: theorem 1: under zfc let cardinality be infinite and regular, satisfying that for any , . if s = (m,r) is a system with each object also a system , satisfying • , • , for each object in s, and • there exists such an element g that belongs to at least objects in m, then there exists a partial system b of s satisfying that • the object set of b is of cardinality ≥ , • g belongs to each object of b, and • the partial system b forms a centralized system. (a) eddy motion of a general system (b) meridian field of the yoyo model (c) typical trajectory of how matters return fig. 1. the yoyo model for a general system for the proof of this result, see appendix. this theorem implies that within a nation’s economy, although there might be areas the government has either no or little influence, there is at least one portion of the economy that is of roughly the same scale of the entire economy, over which the government can influence. to better understand this explanation, the following yoyo model helps us intuitively imagine what a system is and how systems the role government policy and supports play to stimulate economic growth 35 copyright ©2021 assa. adv. in systems science and appl. (2021) evolve and possibly interact with each other. in particular, every system can be seen as an abstract yoyo in fig. 1 [14]: each system is a multi-dimensional entity that spins about its axis. if we fathom such an entity in our 3-dimensional space, such a structure, as shown in fig. 1(a), appears. the input side pulls in ‘things’ (e.g., materials, information, investment, and human talents). after funneling through the “neck”, things are spit out in the form of outputs (e.g., products, services, etc.). some things, spit out as outputs, never return to the other side and some will, fig. 1(b), where fig. 1(c) depicts the trajectory of how things return. due to its general shape, such a structure is known as a yoyo. from this systemic model, it can be seen that the government plays a non-negligible role in the nation’s economy as indicated by the center arrow in the yoyo body that provides the orientation for the system and influences component elements (industries/companies) of the abstract systemic model of the economy. speaking differently, when the government advocates and seeks after the goal of advancing the economy by increasing the degree of marketization, deepening political reform, and encouraging dramatic rise of private enterprises, its use of policy tools will be most likely successful. the reason why this approach practically works is that the government can effectively impose its will and desired outcome on a large segment of the economy through employing policy tools, as guaranteed by theorem 1 that characterizes centralized systems. so, the following conclusion follows: proposition 1: each government’s commitment stands for a process of social influence and support in which relevant governmental officials can gather the assistance and support of a large number of economic agents in the accomplishment of a determined national objective. such governmental commitment eventually creates a harmonious way for individual economic agents to work jointly and collectively to accomplish the desired outcome. a similar but more narrowly stated conclusion than this proposition has been empirically confirmed by [17] in the name of leadership and by [18]. by goal orientated system, it means such a system that focuses on a certain determined collection of tasks and consequences. the study on goal-oriented systems is voluminous. typical examples of goal-oriented systems are [19]: regulation, control, self-organization, learning, autopoiesis, self-reproduction, self-correction, adaptation, evolution, etc. by combining this concept and the systemic study of leadership [20], one has: proposition 2: the importance of the government is reflected in its ability to adjust the nation’s underlying organizational structure as a goal-oriented system so that most of individual economic agents will be able to adjust their orientations and operations without much difficulty. systemically speaking, there naturally exist development unevenness and imbalances among different regions of any nation, as shown by the evolution of flow patterns in the dishpan experiment [21]. in particular, a dishpan, filled with fluid, can model a nation’s economy with the “fluid” modeling the movement of such things as “money, information, goods, and others” within the economy. then the flow pattern, seen from either the input or output side, alternates between a uniform one, fig.1(a), and a chaotic one, fig. 2, demonstrating the fact that imbalances in economic development are commonly seen phenomena. fig. 2 depicts why this proposition holds true, where the entire economy is modelled as a spinning yoyo with fields around. the large dark-colored arrows through the middle stands for the orientation of the overall system, as desired by the government. if the orientation of a 36 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) local spin field, say a1, is not in agreement with that of the overall yoyo body, then a1 will be either dissolved under the joint effects of its adjacent fields or marginalized by forced to the periphery of the overall yoyo field, as indicated by theorem 1. so, a1 will not be any part of the centralized partial system b of the original economic system s. 2.3. conditions when policy tools will work proposition 3: if a government can employ its organizational influence by using policy tools to obtain aids and supports of other economic agents, be they large or small, in accomplishing a determined national objective that will benefit a large number of economic agents, there will appear a joint ambition that potentially involves a major portion of the economy. fig. 2. the structure of a goal-oriented system this is a restatement of theorem 1. the if-condition matches that the system s = (m,r) satisfies. therefore, the conclusion follows from the existence of the centralized partial system b. for a relevant but different study, see [17]. proposition 4: if a government is effective in terms of implementing its policies, then it will be able to align the business objectives of a large proportion of individual economic agents, even though they follow their respectively different and often opposing business strategies. hence, a united front of effort will be formed. this follows from theorem 1 and fig. 2, where the local fields located in the adjacent areas between a1, a2, a3, and a4, spin in directions opposite to those of a1, a2, a3, and a4. differences in directions reflect that businesses, even employing opposing strategies, can peacefully coexist. proposition 5: the initial concept of further advancing the nation’s economy can become an extraordinary reality through implementing appropriate policy tools. this result follows jointly from propositions 2 – 4 above. the importance of these propositions is that no matter what policy tool(s) are adopted and which implementation approaches are used, they need to be rooted in the outcome of benefiting a large economic segment for the policy and approach to actually work. such consistent, conceptual and practical specifications provide a common goal for a large proportion of economic agents to aim at and a cheering point for the citizens of the nation to feel excited about. the role government policy and supports play to stimulate economic growth 37 copyright ©2021 assa. adv. in systems science and appl. (2021) generalizing theorem 3 in [22] produces the following, which describes how a market signals its need for innovations and additional competition. the proof is given in appendix. theorem 2: in nash equilibrium, if the consumer surplus of the market described below is larger than the loyal-customer base of one incumbent company, then the market calls for new innovation and additional competition. if a company answers the call by entering the market with its version of substitute offer, then its expected profit can be potentially larger than that of at least one incumbent company. here, the market satisfies the following conditions: • it is served by a number of companies with their horizontally differentiated offers; • its operation is only affected by market forces, such as demand and supply, and consumers’ forever evolving preferences and tastes; • each incumbent company has a base of loyal customers if the price is not more than their reservation price. • there are switchers who make purchases based on whose price/value is lower; they are collectively known as consumer surplus; and • all companies’ pricing strategies are known to the incumbent companies who respond by playing nash equilibria through untainted self-analyses. for the market considered above, it generally means that the technology and management used in production have been standardized. hence, for a company to profitably enter the market, it must have introduced a more efficient technology and/or managerial routine that can significantly reduce its business expenditure in making its product of increased sophistication and functionality, as christensen et al. [23] confirms empirically. 3. the necessity of governmental policies and supports 3.1. linkage mappings within a supply-chain ecosystem to address why government policy tools and supports are fundamentally necessary for stimulating economic growth, theorem 2 indicates that when recognizing a market signal innovatively, a firm is in a competitive position over its competitors. however, in practice, the consequent enhanced performance of the firm really depends on other players in its supply-chain ecosystem [24]. in particular, to produce the offering to satisfy an innovatively recognized demand, a firm has to obtain its needed supplies from its suppliers. that can be a challenge to the supply-chain ecosystem. even after having met such challenge and created the expected value, the firm can still be hampered in its capturing of value. that might well depend on infrastructures necessary for the firm to reach consumers in the marketplace [24]. the supply-chain ecosystem of any firm consists of upstream components, such as suppliers, and/or downstream complements, such as customers, supporters and assistants, known as complementors, who help to make the firm’s product usable by consumers [24]. although outside the firm’s direct supply chain, complementors help connect the firm’s offering and consumers by constructing infrastructures. for example, governments, regulators, etc., are complementors that help build and maintain, for instance, roads necessary for transportation or specifications of new safety procedures, etc. the systemic structure of this supply-chain ecosystem is given in fig. 3. the focal firm utilizes inputs, called components, from n supplies and delivers its outputs to customers with the support and assistance of m complementators, m, n = 1, 2, 3, … other than the shown first-tier components and complements, this structure in real life needs to extend leftward and rightward along the value-creation chain to include other tiers, such as suppliers' suppliers, customers' customers, complementors’ complementors, etc. by including the full 38 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) range of different tiers, the systemic characteristics of this ecosystem can be vividly seen, where the focal firm internally focuses on providing organizational supports so that its employees can innovatively integrate components into market offerings. for similar systemic constructs, see, for example, [24,25]. fig. 3. the supply-chain ecosystem of a focal firm let denote the supply chain of a firm. model firm in this chain as an abstract system . assume the index set of these firm systems is , satisfying that for any , there is a unique firm system in the supply-chain ecosystem; and for any firm system in the supply-chain ecosystem, there is a unique index that labels the system . then, can be modeled as a hierarchy of firm systems , for . let , i = 1, 2, be systems and h: → a mapping. for each relation , define . without confusion, write h: → . a mapping h: → is said to be partial, if for some object , is not defined. the identity mapping is defined by , for . theorem 3: assume that the hierarchy does not include such systems and that some outputs of are used as inputs of , while some outputs of are employed by as inputs. then the supply-chain ecosystem can be made into a partially linked hierarchy of systems by a family of partial linkage mappings of . the proof is given in appendix. 3.2. maintenance of an existing momentum of economic growth to investigate how an existing momentum of economic growth can be maintained, let us first look at the concept of feedback systems [14]. assume that and are two linear spaces and and and and linear functions from to and from to , respectively. then the feedback system of s by is defined as the input-output system such that the following expression holds true in our current context, theorem 3 means that the existence of a large magnitude consumer surplus, as provided in the assumptions of theorem 2, represents the fact that the incumbent firms of the market can no longer satisfy the forever changing consumer preference and tastes. in real life, of course, this fact (or market call) can be understood in many different ways. if our focal firm innovatively deciphers the market call by introducing its original product, then that implies that each tree-like subset of the partially ordered set t in terms of the order relationship contains minimal elements, consisting of various market demands, and that every player in the supply-chain ecosystem of the focal firm will be the role government policy and supports play to stimulate economic growth 39 copyright ©2021 assa. adv. in systems science and appl. (2021) challenged to provide their correspondingly innovative supplies in order to jointly answer the market call. even though the initial market call might not be answered adequately by a single firm’s reply with its new product, the aggregate of all relevant original products offered by various entrepreneurial firms will definitely satisfy the market demand. so, this discussion combined with the concept of feedback systems leads to the following fact regarding the importance of government’s policies. proposition 6: to maintain the extant economic growth momentum, assuming that such a momentum already exists, a nation has to constantly provide “fuels” through policy supports for the prevailing feedback system, which reinforces the horse race between market demands and manufacturing production so that the race will continue to intensify, to continuously function. in fact, what is discussed in the previous paragraphs indicates that market exchange stimulates manufacturing production, while the production encourages consumers to alter their preferences and tastes. this cyclical evolution continues and helps both market exchange and manufacturing production enter into a horse race against each other. that is, an operational feedback system appears. in this feedback system, forever changing market demands provide stimuli for manufacturers to produce more products and make better products. and the growing magnitude of manufacturing production forces firms to hire more employees with rising salaries. such mutual reinforcements further strengthen the population’s purchasing power, which in turn pushes market demand onto a higher level; … to keep the prevailing feedback system functional smoothly, the economy (or the government) has to encourage as many entrepreneurial firms as possible to answer various market calls by providing their original products. however, many of these original products will not be readily usable by the consumers of the marketplace unless adequate infrastructures are constructed and maintained by forces supported either directly or indirectly by the government. speaking differently, the reason why policy supports are needed to fuel the prevailing feedback system in order for it to function smoothly is because the system is too colossal for any individual economic enterprise to handle, while the government can reach and mobilize a large proportion of the economic system (theorem 1). hence, the government can at least help coordinate many individual firms in very large scales to join their efforts, resources, etc., by utilizing policy tools and supports. for example, the literature shows that to stimulate economic growth the government can use policy tools to • promote the ideology of commercialization at various societal levels, especially in rural areas [2,26]; • encourage entrepreneurship and risk taking [27]; and • provide the freedom for people to move geographically and professionally and to make career choices [28], and can provide various supports [2,4,16,29,30], such as • development of the necessary domestic and international markets of sufficient depth and distribution networks of appropriate sophistication; • smooth operations of these markets and networks; • financial and politically stability; • construction of necessary infrastructures (such as roads, irrigation systems, energy supplies, etc.); • smooth flows of knowledge and information that promote coordination and specialization of labor; 40 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) • educational programs for necessary skill trainings; • resources and needed coordination of resources; • and others. moreover, the discussion in the previous sections elucidates why further development of the focal nation’s economy needs to pay equal attention to increasing degree of marketization, continuously deepening political reform, and dramatically raising the magnitude of market sector of private enterprises. in particular, by increasing the degree of marketization, market demands can be freely and adequately signaled through market phenomena and forces. by dramatically raising the magnitude of the market sector of private enterprises, entrepreneurs can individually and differently comprehend market signals and take their respectively appropriate actions by designing and offering their varied original products and services to the marketplace. by continuously deepening political reform, the government’s way of conducting its business can be brought closer to the current state of the accelerating horse race between the market demands and manufacturing productions. only by synchronizing the political system that is underneath the composition and operation of the government with the prevalent working state of the feedback system of the economy, which reinforces the horse race between market exchanges and manufacturing productions through raising purchasing power of the population, the overall organizational structure of the focal nation, seen as an entity that is simultaneously political, economic, social, etc., will stay as a stable, smoothly functional system. 4. conclusion because of the use of systems science and methodology, this work is able to scientifically address: (1) why are policy tools expected to work in an economic system? (2) why are government policy tools and supports fundamentally necessary for stimulating economic growth? and (3) several other relevant issues. historically, there had been three rounds, plus the current round for a total of four rounds, of debate that directly attempted to address these questions from different angles. however, due to the fact that only data/anecdote-based approaches were employed, what achieved turned out to be theoretically inconsistent and uncompromising suggestions that are practically unreliable [3]. the entire situation is similar to that reflected in the literature on the industrial revolution [2,4] but with such a major difference that for the latter rostow recognized the need for introducing an appropriate methodology in order for relevant studies to escape from the trap of not being able to produce consistent conclusions. by comparing these two lines of related but different studies, this paper develops generally-true conclusions by innovatively employing logical reasoning based on results established rigorously. in particular, among others, the following main results are established: • if a government is able to introduce and implement appropriate policy tools, then the determination of further advancing the nation’s economy can become an extraordinary reality. it is because (propositions 2 – 5) in this case the government is able to adjust the economy’s underlying organizational structure as a goal-oriented system within which most individual economic agents will be able to adjust their orientations and operations without much difficulty. • to maintain an existing momentum of economic growth, a nation needs to constantly provide policy supports for the prevailing feedback system, which reinforces the horse race between market demands and manufacturing production so that the race will continue to intensify, to continuously function (proposition 6). the role government policy and supports play to stimulate economic growth 41 copyright ©2021 assa. adv. in systems science and appl. (2021) because of the particular methodology employed here, the established theoretical conclusions provide policy makers with dependable bases to make their decisions. for example, the general recommendations below for a national government follow naturally: (1) it needs to constantly maintain its capability to make adjustment(s) to the underlying organizational structure of its economy. only with such capability, the government is able to fine-tune and to redirect the development and evolution of the economy (proposition 2). (2) it needs to frequently identify objectives of economic development that potentially benefit most economic agents. by accomplishing such a goal, the policy makers of the nation will be able to develop a national ambition supported by a majority of the nation (proposition 3). (3) if all possible, it needs to find effective ways to implement its policies in order to practically materialize the goal of further advancing the nation’s economy from the present state (propositions 4 and 5). (4) it needs to create and maintain a functional feedback mechanism that intensifies the horse race between market demands and manufacturing production (proposition 6). first, by creating such a feedback mechanism from scratch, an originally impoverished nation can potentially evolve into an economy of proto-industrialization [16,30], where local handicraft productions, alongside commercial agriculture, will be developed to such a level that, beyond local markets, can also satisfy external markets. secondly, by maintaining such an existing feedback mechanism, a developed economy will be carried to a higher level of economic development [2]. there are also some limitations to this work. first, although conclusions of this paper depend heavily on the methodology of systems science, we only applied one of the many tools available that are developed for analyzing organizations, their evolutions and interactions [31]. so, when other tools of systems science are employed one by one in studies of business-related issues and problems, we will be able to establish finer and reliably applicable conclusions. second, all reasonings used in this work assume why a firm exists – it attempts to satisfy a particular market niche with its operation financially maintained by a positive cash flow as a consequence of its business conducts. although this assumption is conventionally true in the past [32], it is no longer the case in the present business landscape. for example, some e-business operations have focused on pumping up their future promises and potentials to continuously attract sufficient venture capitals by placing emphasis on increasing their market shares, although they have been losing money year after year since their inception [33]. hence, any violation of this assumption can realistically stand for why applications of our general conclusions developed in this paper fail to work in practice, if the firms of concern exist for purposes other than what is assumed. these limitations, along with others not listed here, of this current work provide directions for future research. appendix: proofs of theorems for related terminologies and symbolic expressions used in these proofs, see [14]. proof of theorem 1. for convenience, assume that and that there is a common element in all the object systems in m, which implies (1) without loss of generality, for each object ∈ m, assume (2) so, for each , the order type < of is a subset of . because is regular and , there exists a so that = {z ∈ m: has order type } has cardinality . let us fix such a and deal only with the partial system of s, where is the restriction of the relation set r on . 42 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) for each , implies that less than objects of the partial system have object sets as subsets of and is cofinal in . if z ∈ and , let (ξ) be the ξth element of . because is regular, there is some ξ such that { (ξ): z ∈ } is cofinal in θ. let be the least such ξ. then the condition that there exists a common element in each system in implies that > 0. let (3) then < and < for all z ∈ and all . by transfinite induction on , pick so that and (4) let . then and whenever , are distinct objects. because for each , there exists an and a with and for each , , forms a centralized system, where is the restriction of the relation set r on b. proof of theorem 2. assume that the market is occupied by incumbent companies, the, individually different boundary conditions of companies can be normalized so that the production cost is 0, customer’s reservation price 1, price satisfies , loyal-customer bases of the incumbent companies are of the same percentage scale with being the scale of the consumer surplus. assume that the entering company randomizes its price between its cost 0 and consumer reservation price 1. to protect their established territories, each incumbent sets its price by considering all competitors. so, the equilibrium indifference condition of incumbent company is (5) where is the price distribution of company ; and in nash equilibrium, the incumbents do not have any pure pricing strategy [22]. hence, for the incumbents, their symmetric equilibrium pricing is (6) the assumption implies that equation (6) defines a mixed strategy for each incumbent for . because experiences a jump of at p = 1, the expected profits of the entrant are: (7) (8) the first term on right-hand side of equation (7) is equal to the entrant’s expected profits when it charges a price lower than the incumbents and captures all switchers, and the second term the entrant’s expected profits when it directly competes with the incumbent companies. the expected profits of an incumbent are the role government policy and supports play to stimulate economic growth 43 copyright ©2021 assa. adv. in systems science and appl. (2021) (9) because , and when , , there is so that when , . that is, the entrant can actually expect to make more profits in the said market than some incumbent. proof of theorem 3. in fact, the input-output relationship among the firms in implies that the index set can be ordered partially: for , if, and only if, some outputs of firm system are used as inputs of firm system . for , if , define a partial (linkage) mapping from firm into customer firm by (10) for any such that system ’s output is applied as a component in ’s product . because for any , satisfying , we have (11) this systemic modeling implies that we can treat the supply-chain ecosystem of any firm as a partially-linked hierarchy of firm systems . references 1. hermann, m., pentek, t., & otto, b. (2016) design principles for industry 4.0 scenarios. the 49th hawaii international conference on system sciences (hicss), koloa, hi, 3928–3937. 2. wen, y. (2016) the making of an economic superpower: unlocking china’s secret of rapid industrialization. singapore: world scientific. 3. andreoni, a., & chang, h.j. (2019) the political economy of industrial policy, structural change and economic dynamics, 48(march), 136–150. 4. rostow, w. w. (1960) the stages of economic growth: a non-communist manifesto. cambridge, uk: cambridge university press. 5. wieczorek, a.j., & hekkert, m.p. (2012) systemic instruments for systemic innovation problems, science and public policy, 39(1), 74–87. 6. auclert, a. (2019). monetary policy and the redistribution channel, american economic review, 109(6), 2333–2367. 7. boeckx, j., & de sola perea, m., & peersman, g. (2016) the transmission mechanism of credit support policies in the euro area. national bank of belgium working paper, no. 302. 8. rebei, n. (2017) evaluating changes in the transmission mechanism of government spending shocks, international monetary fund, wp/17/49. 9. cúrdia, v., & woodford, m. (2016) credit frictions and optimal monetary policy, journal of monetary economics, 84(december), 30–65. 10. jaehrling, k., johnson, m., larsen, t.p., refslund, b., & grimshaw, d. (2018) tackling precarious work in public supply chains, work, employment and society, 32(3), 546– 563. 11. howell, t.r., noellert, w.a., maclaughlin, j.h., & wolff, a.w. (2019) the microelectronics race: the impact of government policy on international competition. london: routledge 12. jomo, k.s. (2019) southeast asia’s misunderstood miracle: industrial policy and economic development in thailand, malaysia and indonesia. new york: routledge. 44 j.y-l. forrest, a.k. jallow, c. wen, j.h. shi, h. guo copyright ©2021 assa. adv. in systems science and appl. (2021) 13. hemphill, t.a. (2014) the us advanced manufacturing initiative: will it be implemented as an innovation – or industrial – policy? innovation: organization & management, 16(1), 67–70. 14. lin, y. (1999) general systems theory: a mathematical approach. new york: kluwer academic/plenum publishers. 15. porter, m.e. (1985) competitive advantage: creating and sustaining superior performance. new york, ny: free press. 16. forrest, j.y.l., zhao, h.c. & shao, l. (2018) engineering rapid industrial revolutions for impoverished agrarian nations, theoretical economics letters, 8, 2594–2630. 17. chemers, m.m. (2001) cognitive, social, and emotional intelligence of transformational leadership: efficacy and effectiveness. in r. e. riggio, s. e. murphy, f. j. pirozzolo (eds.), multiple intelligences and leadership (p.139–160). mahwah, nj: lawrence erlbaum. 18. kouzes, j., & posner, b. (2007). the leadership challenge. ca: jossey bass. 19. klir, g. (2001) facets of systems science. new york: springer. 20. lin, y., & forrest, b. (2011) systemic structure behind human organizations: from civilizations to individuals. new york: springer. 21. fultz, d., long, r. r., owens, g. v., bohan, w., kaylor, r., and weil, j. (1959) studies of thermal convection in a rotating cylinder with some implications for large-scale atmospheric motion. meteorol. monographs (american meteorological society) vol. 21, no. 4. 22. zhao, h.c., forrest, j.y.l. and jirasakuldech, b. (2018) a game analysis of trade dumping and antidumping, theoretical economics letters, 8, 2860–2881. 23. christensen, c.m., suárez, f.f., & utterback, j.m. (1998) strategies for survival in fastchanging industries, management science 44(12-part-2), s207-s220. 24. adner, r. (2006). match your innovation strategy to your innovation ecosystem, harvard business review, 84(4), 98–107. 25. iansiti, m., & levien, r. (2004) the keystone advantage. boston, ma: harvard business school press. 26. mendels, f.f. (1972) proto-industrialization: the first phase of the industrialization process, the journal of economic history, 32, 241–261. 27. petrovslaya, i., zaverskiy, s., & kiseleva, e. (2017) attribute to entrepreneurship in russia, advances in systems science and applications, 17(2), 29–43. 28. huntington, s.p. (1996) the clash of civilizations and the remaking of world order. new york: simon & schuster. 29. fleischman, r.k. & parker, l.d. (2017) what is past is prologue: costing accounting in the british industrial revolution, 1760-1850. london, uk, and new york, ny.: routledge. 30. coleman, d.c. (1983) proto-industrialization: a concept too many, the economic history review (new series), 36(3), 435–448. 31. forrest, j.y.l., duan, x.j., zhao, c.l., & xu, l.d. (2013) systems science: methodological approaches. new york: crc press, an imprint of taylor and francis. 32. sobel, r. (1999) when giants stumble: classic business blunders and how to avoid them. paramus, new jersey: prentice hall. 33. li, t., & ma, j.h. (2015) complexity analysis of dual-channel game model with different managers’ business objectives, communications in nonlinear science and numerical simulation, 20: 199–208. microsoft word 1049 article text, copyedited.doc adv syst sci appl 2022; 01; 1-14 published online at https://ijassa.ipu.ru. plasma magnetic control systems in d-shaped tokamaks and imitation digital computer platform in real time for controlling plasma current and shape yuri v. mitrishkin* lomonosov moscow state university, faculty of physics, moscow, russia v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: y_mitrishkin@hotmail.com abstract: the direction of developing plasma magnetic control systems in d-shaped tokamaks with its actuality and scientific novelty is presented. shortly the poloidal systems of iter and demo are characterized. the new imitation computer digital platform for modeling plasma magnetic control systems in real time for application in physical experiments is described. the preliminary results of testing the plasma magnetic control system for the globus-m2 tokamak (ioffe institute, s-petersburg, russia) on the target computer of the platform are given. the digital twin of the imitation platform is characterized. the methodology of risk minimization is presented. keywords: tokamak, plasma control, imitation digital platform, real time, globus-m2, iter, demo 1. introduction the world challenge of the controlled thermonuclear fusion (ctf), researches on which have been begun in the beginning of 50th years of the last century, is one of central problems in science and engineering. the solution of this problem will open a new, safe, almost inexhaustible source of energy from the synthesis of nuclei of light elements, as well as eliminate carbon dioxide emissions into the atmosphere from the combustion of natural resources and stop climate change on the earth in an unfavorable and dangerous direction for mankind. fig. 1. vertically elongated tokamak without an iron core: 1 is the vacuum vessel; 2 is the toroidal field coil; 3 is the poloidal field inner and outer coils; 4 is plasma and helical magnetic lines (© iter project center, russia) * corresponding author: y_mitrishkin@hotmail.com 2 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) the most perspective direction of the ctf is tokamaks with d-shaped cross-section (fig. 1), on the basis of which it is planned to create thermonuclear power plants. confinement and heating a plasma in tokamaks to provide a self-sustaining thermonuclear reaction are ensured by magnetic and kinetic plasma control systems with feedback. the plasma in the magnetic field of a tokamak vacuum vessel represents extremely difficult for diagnostics and control a nonstationary nonlinear dynamical plant with distributed parameters, with structural and parametric uncertainties, subject to the influence of uncontrolled disturbances in the presence of broadband noises. that's why one of the main tasks of the ctf has not been solved yet specifically prolonged confinement of the high-temperature plasma in tokamaks with set parameters and prevention of plasma major disruptions. the paper is structured as follows. section 2 highlights the significance of plasma magnetic control systems in tokamaks. in section 3 the scientific novelty of the direction presented in the paper is shown and concentrated on the imitation digital computer platform in real time for controlling plasma current and shape in tokamaks. section 4 is devoted to some important features of tokamaks in whole, iter, and demo. in section 5 the short introduction into practice of physical experiments of the globus-m2 tokamak of the digital plasma control system is given. section 6 presents the basic new idea of the paper specifically the imitation digital computer platform structure in real time having new design for application to plasma control in tokamaks. in section 7 the digital twin of the plasma magnetic control system is shown in the digital imitation platform. the results of the preliminary testing of one of the plasma current and shape control systems for the globus-m2 tokamak on the target computer is given in section 8. section 9 describes the methodology of risk minimization of application of the direction suggested. the conclusion summarizes the main new features of the direction of the plasma control systems development. 2. significance of plasma magnetic control systems in tokamaks in this situation the urgency of continuing the development and research of magnetic plasma control systems in tokamaks only increases [2, 3]. it is explained first of all by the fact that the systems of magnetic plasma control have not yet reached a proper level of reliability and survivability for their application in thermonuclear reactors, where increased reliability of systems operation in conditions of presence of the thermonuclear reaction in the plasma of reactors is required. it is especially important for round-the-clock operation of thermonuclear power plants. moreover, modern experiments on tokamaks require more and more advanced plasma control systems, which would allow deeper and more accurate study of plasma processes during generated discharges, could guarantee rejection of minor disruptions and also prevent major disruptions, so that they do not lead to emergency situations and destruction of fusion power plants. in doing so, at present nobody knows models of the plasma with the thermonuclear reaction and so nobody can guarantee the operability of the plasma magnetic control systems in thermonuclear reactors like iter (international thermonuclear experimental reactor) and demo (demonstration fusion power plant). design of plasma magnetic control systems in tokamaks should have the feature of automatic adaptation to fusion plasma new models. 3. scientific novelty of direction of control systems design further development of the plasma magnetic control systems in the stated direction in the continuation of the achievements in the grant of the russian science foundation № 17-19-01022 (2017-2019) “hierarchical automatic control systems of plasma position, current, and shape in toroidal axial-symmetric magnetic configurations in a wide range of aspect ratios” will be carried out in real time with their high scientific novelty by the team of [4]. plasma magnetic control systems in d-shaped tokamaks and imitation digital 3 copyright ©2022 assa. adv. in systems science and appl. (2022) the novelty will be provided primarily by the originality of the imitation digital computer platform created in real time. installation of new plasma control algorithms obtained on the digital twin of the control system is planned to be practiced over the internet. an application for a patent of the russian federation is being made for this digital imitation platform. the platform will allow setting up plasma control systems with an algorithm for plasma equilibrium reconstruction in real time. the russian federation patent for invention no. 2702137 was received in 2019 for the method of modeling such setting [5]. the industrial computer will allow replacing all analog controllers in the control loops of plasma position and currents in the coils of the central solenoid and poloidal field with digital controllers (8 controllers) on the globus-m2 tokamak itself, and will also allow applying the plasma shape control system with feedback in experiments. digitalization of the whole system of magnetic plasma control on the tokamak through the third platform computer will allow to set and solve a new task of optimization of joint coordinated operation of all plasma control circuits in real time. as it follows from the surveys on the systems of magnetic plasma control in tokamaks [6, 7], such a task was not set and was not solved by anyone. the solution of this problem is necessary for an optimal integration of the plasma magnetic control system with the kinetic control system of plasma parameters profiles (current density, safety factor q, density, temperature, pressure) in the long term. fig. 2. general configuration of three levels hierarchical control system [8] novelty of the plasma control systems will be provided by the fact that the control systems will be developed in the class of digital hierarchical control systems (fig. 2) for multivariable dynamic plants with time varying parameters [8]. this scientific direction allows to obtain various practically unlimited new combinations of different control and identification algorithms at different hierarchical levels. sensor networksactuators controlled process baseline controller adaptation algorithm selforganization algorithm disturbance noise i. basic control level ii. adaptive level iii. intelligent level general feedback loop decision-making self-learning, self-configuring self-optimizing level function adapting basic controller to time-varying plant level name _ r stabilizing of plant outputs y or tracking of references r yv e 4 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) the main directions in development of hierarchical plasma control systems will be as follows: ø identification (building models based on experimental data of observation of inputs and outputs of the controlled plant), ø robust control, ø adaptation, ø elements of artificial intelligence. in the control directions it is supposed to be developed: • control systems with a predictive model with adaptation of parameters to changing parameters of the controlled plant [9]; • robust systems with multivariable pid controllers adjusted by means of linear matrix inequalities [10]; • decoupling of control channels with subsequent adjustment of control channels by means of the method of quantitative feedback theory [11]; • decoupling of control channels by means of relative gain array approach and h∞optimization theory [12]; • robust and adaptive systems with coordination of control loops of position and shape of the plasma [13]; • methods of picard iterations, moving filaments, and neural networks for reconstruction of plasma equilibrium on the basis of the magnetic measurement outside the plasma [14-16]. at the same time, the development and application of new linear and non-linear plasma models will be continued. linear models are designed for synthesis of linear magnetic plasma control systems without taking into account slow plasma processes related to the transfer of matter and energy. non-linear models take into account the transfer processes and are aimed at future development and implementation of kinetic plasma control systems. novelty of plasma control systems development directions of our team is confirmed by three reports at the ifac world congress in 2020 (ifac wc 2020), which was held remotely in berlin (germany) [9, 10, 16]. 4. tokamaks, iter, and demo europe is a leader in solving the ctf problem. at the european tokamaks, in particular, jet (major radius r=3 m, england), asdex upgrade (r=1.65 m, germany), tcv (r=0.88 m, switzerland), the experimental researches of the high-temperature plasma are intensively conducted, new plasma control systems are developed, applied, and investigated. in europe, the first iter thermonuclear reactor with r=6.2 m is under construction in cadarache (france) [17, 18], where the assembly of the reactor itself has already been started. there are two roadmaps for the creation of demo on tokamaks: (a) with large aspects ratio [19] and (b) spherical tokamaks [20, 21]. the first roadmap leads to the giant demo having the very large major radius of about 9-10 m. the second roadmap leads to the modular demo on spherical tokamaks as modules with the lower aspect ratio and relatively small major radius specifically less than 2 m. the cost of electricity from giant tokamaks should be very high but the cost of electricity of the modular demo is very competitive and is of about 0.06 $/kwh [21]. plasma magnetic control systems in d-shaped tokamaks and imitation digital 5 copyright ©2022 assa. adv. in systems science and appl. (2022) a b c fig. 3. vertical cross-sections of (a) iter [22], (b) asdex upgrade [23], (c) globus-m2 [24] there are a set of poloidal systems (poloidal field coils around the vacuum vessel) for the giant demo as copies of the poloidal system of iter with the rough mistake in the design of this system [22]. this original system led to the nonsense: at the major radius of iter of 6.2 m the controllability region of the plasma unstable vertical position was about 3-4 cm. because of that the internal extra coils to control plasma unstable vertical position have been installed into the vacuum vessel of the iter tokamak (fig. 3a) [22]. the engineering solution of including the pf-coils inside the vacuum vessel of demo is not the best idea from the following points of view: • the thermonuclear plasma may burn these coils, • the space for the plasma became less inside the vacuum vessel and the plasma is located farther from the first wall in comparison with tokamaks which do not have the pf-coils inside the vacuum vessel. this essentially decreases the effectiveness of acting on the plasma of the eddy currents induced in the first wall. it is much effective and reliable to locate the pf-coils for the plasma vertical position control between the vacuum vessel and the toroidal field coil like for instance in the asdex upgrade tokamak (fig. 3b) [23] or design the whole poloidal system like in the spherical globus-m2 tokamak (fig. 3c) [24]. 5. introduction into practice of physical experiments of the globus-m2 tokamak of the digital plasma control system the digital simulation platform will provide an opportunity for the first time to implement the developed new plasma magnetic control systems in real time in the physical experiments of the operating globus-m2 spherical tokamak. solution of off-line identification tasks will be carried out on the existing globus-m2 tokamak equipment, which is planned in the direction for successful guaranteed application of new plasma control systems in the experiments. the multivariable plasma magnetic control system which is operating at present on the spherical globus-m2 tokamak in experiments is shown in fig. 4a [13]. it contains 2 closed control loops for plasma position, 1 control loop for the current in the central solenoid (cs), and 5 control loops for the currents in the poloidal field (pf) coils. to generate plasma discharge scenarios the program reference signals are fed to these loops. 6 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) a set of plasma magnetic control systems for the globus-m2 tokamak was developed and modelled in line with the new methodology presented in [25] for plasma current and shape control with feedback in the matlab/simulink environment. 6. description of the imitation digital platform there is the basic approach to model plasma control systems in real time by means of test bed containing two target computers connected by feedback [26]. for control of plasma current and shape with feedback in real time on the globus-m2 tokamak it is proposed to create and apply a new real-time imitation digital computer platform, which has a feedback bunch on two high-tech computers with high efficiency. these are target computers such as " plant simulator ̶ controller", where the switching in the “controller” of the control algorithm from the internal model to the “plant simulator” and back takes place to configure the control system with subsequent transfer of a control algorithm through remote access to the target computer with its internal model connected to the tokamak (fig. 4b). the real-time platform realizes real-time plasma simulation in the tokamak with diagnostics, actuators, control loops for the horizontal and vertical position of the plasma and currents in the pf-coils as well as plasma current and shape loops on three industrial computers (ic) with communication devices with the controlled plant, which is the plasma in the tokamak. the ic1 simulates only the plasma in the tokamak with acting control loops of plasma position and cs/pf-currents (fig. 4a), while the other ic2 and ic3 simulate the plasma in the tokamak with acting control loops, and the plasma current and shape controller connected to each other by internal feedback with the plasma equilibrium reconstruction code at the controller input. the ic2 and ic3 use two multivariate switches which switch the c controller of the ic2 from the plasma internal model (min) to the plasma external model (mext) or the tokamak and back to adjust the control system on the min and then apply it out on the mext or the tokamak. that allows to simulate by the c controller of the ic2 connected to the mext real physical experiments on the tokamak. host computers (hc) for the development and modeling of plasma control systems in computer time with a display d for visualization of modeling of plasma control processes are connected to each of ic1, ic2, and ic3. these hcs are used to carry out loading from them the developed controllers and plasma models into ic1, ic2, and ic3. as well as displays d are connected to ic1, ic2, and ic3 to visualize internal control processes in real time. a database server (s) is connected to each of the three hcs, through which the data are downloaded for all hcs to develop plasma control systems in computer time and further to be downloaded to ic1, ic2, and ic3 for the purpose of modelling plasma control systems in real time in line with the simulation method with plasma reconstruction code in the feedback [5] and future plasma control in real time. the ic3 with hc and d is installed and connected directly to the tokamak, the internal controller c with the reconstruction code at its input is set up on the min in real time, and then is switched to the tokamak. as a result, the controller c previously adjusted on the min controls the plasma current and shape in real time during plasma discharges. for the progress of plasma control systems, the internal controller c in the ic3 may be changed by means of remote access to a new internal controller. then this new controller is adjusted on the min and is switched to plasma control in the tokamak in real time by means of two multivariate switches. plasma magnetic control systems in d-shaped tokamaks and imitation digital 7 copyright ©2022 assa. adv. in systems science and appl. (2022) a b fig. 4. (a) block diagram of the plasma control system operating on the globus-m2 tokamak: cz, cr, ccs, cpf1, cpf2_top, cpf2_bottom, cpf3, and ccc are analogue controllers, ahfc, avfc, acs, apf1, apf2_top, apf2_bottom, apf3, and acc are actuators [13]. (b) structural scheme of the imitation computer platform: c is a controller, min, mout are internal and external plant models, d is a display, hc is a host computer, ic is an industrial computer, s is a server, tok is a tokamak 7. digital twin of real control sytem in imitation platform digital twins of control systems are being developed in industry when designing control systems because control processes can be simulated in real time on the digital twin, they can be optimized or new control algorithms can be synthesized and the optimal system results found are then applied to the real plant under control. this is much cheaper than solving the same problems at once on the real plant. at the same time, digital twins are also used to help to control complex dynamic plants, for example, in experimental physics, in order to find optimal control solutions for these plants independently of the physical experiments. in the imitation platform (fig. 4b) two industrial computers namely ic1 and ic2 constitute the digital twin and the third industrial computer ic3 is planned to be connected to the real plant specifically the globus-m2 tokamak to organize the real plasma digital control system fig. 5. digital twin of feedback control system of real dynamical plant [ifac 2020 wc] cz ahfc ― ccs acs tokamak ― ics reference magnetic flux (21 loops) ip z reference z ihfc ivv+p cr avfc ― r reference uhfc uvfc ivfc ics cpf1 apf1 ipf1 ipf1 reference cpf3 apf3 ipf3 reference ipf3 ccc acc icс reference icc ipf2_top ipf2_bottom r ucs upf1 upf3 ucc cpf2_top apf2_top ipf2_top reference cpf2_bottom apf2_bottom ipf2_bottom reference upf2_top upf2_bottom ― ― ― ― ― mest ― ic2 мin hc c ic1 dhc d ― ic3 мinc dhc d s msu/ics ras tok ― ic3 мinc dhc d ioffe inst. d plant controlled modelcontrol unit digital twin real time data virtual space feedback real controlled plantcontrol unit real control system real space feedbackrecommendations 8 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) 8. plasma control system tests in real time the preliminary tests of the control system of plasma current and shape as “isoflux control” (fig. 6) [15, 27] on the target ic of speedgoat type (www.speedgoat.com) with simulinkrt operation system were carried out in real time, which showed sufficient speed of response for physical experiments on the globus-m2 tokamak. from the mhd theory of tokamak plasma configurations it is known that the magnetic field at the x-point for any discharge is zero and the plasma boundary is an isolevel line of the poloidal flux. thus, by minimizing |br|, |bz| the x-point is placed at the desired location and by minimizing the plasma separatrix is placed at the points 1 and 2 (fig. 7). to test the operation of the control system in real time, the entire control system was transferred to discrete time and tested on the performance real-time computer of speedgoat (www.speedgoat.ch) with a processor intel(r) core(tm) i7-3770k cpu 3.50ghz. the main time step was 50 μs for all subsystems, except for the equilibrium reconstruction algorithm: for it the time step was 500 μs. as a result, a 50 μs task is performed on average in 24 μs and a 500 μs task is processed on average for 347 μs for 5 filaments. this shows that the system, taking into account computer power, is quite applicable in real experiments. the results of the simulation are shown in figs. 8 and 9. fig. 6. block diagram of multiloop and multivariable hierarchical plasma magnetic control system of the globus-m2 tokamak [15] plasma magnetic control systems in d-shaped tokamaks and imitation digital 9 copyright ©2022 assa. adv. in systems science and appl. (2022) fig. 7. the idea of the isoflux plasma shape control [15] fig. 8. deviations of the magnetic field in the location of x-point for the shot № 31648 in discrete time [15] fig. 9. deviations of the flux in the desired location of the plasma separatrix (fig. 4) for the shot № 31648 in discrete time [15] 10 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) 9. methodology of risk minimization of research any new scientific work is subject to different kinds of risks, which may not lead to the desired result, especially as in this project, when the research is associated with a complex technique, a large field of research of a complex controlled plant, such as the plasma in tokamaks, the development of new methods and systems of plasma control, the introduction of real-time platforms in the computer network, organizational difficulties in solving problems with foreign partners, etc. therefore, it is important to foresee in advance the possibility of appearance of different types of risks and methods of their minimization or complete elimination: ü when developing new tools, there may be the risk of their ineffectiveness or their application and refinement may take too long. this risk can be reduced by the creative abilities of team members. in this case, the most talented and persistent professionals get into the team who have abilities to minimize this risk. ü the risk that the purchased equipment in the form of speedgoat performance computers is not fast enough to solve the problems of magnetic plasma control in the globus-m2 tokamak. in order to minimize this risk, it is necessary to continue the work on studying these computers. for this purpose, it is necessary to consider variants, when delays in communication devices with the plant can have even less delay, than 8 µs. it can be fpga or use of the hdl coder. this issue is part of the study for the applicability of plasma position control channels of the imitation platform. for plasma current and plasma shape control, these delays are quite small and do not interfere with the control process, as our already conducted studies show. in order to reduce the risk of applying real-time plasma control platform to control the plasma in the globus-m2 tokamak, it is advisable to first purchase only two speedgoat performance computers, mount, configure, and test the platform on two "plant-controller" computers connected by the feedback, simulate on it the available plasma position, current, and shape control systems, to understand what needs to be changed in this platform (if necessary) in order to adapt it to solve plasma control tasks on the globus-m2 tokamak. after that, one can purchase one speedgoat performance computer with corrected characteristics to be applied on the tokamak directly. ü the risk from withdrawal from the team of some of its members. such risk exists, but it can be minimized by attracting the most talented young people to the team. for this purpose, the consortium has a leading higher education institution of russia namely lomonosov moscow state university, which can solve this problem because it is the source of young talents in their specialty. ü the risk from unacceptable input-output controllability of the controlled plant [28]. it means that it is impossible to achieve the set control goals on this plant because of its properties, which were obtained when creating this plant. for example, such situation arose with iter, when the poloidal system was already designed and the project development was in progress, it was found that at the major radius of tokamak of 6.2 meters, the vertical region of plasma controllability is only 3-4 cm. in this case it was necessary to change the controlled plant, which was done by introducing additional horizontal magnetic field coils inside the vacuum vessel. this led to expansion of the vertical controllability region by order of magnitude, which is enough for this machine. in the demo project such mistake should not be made: fusion electric power plant should work in twenty-four-hour mode and with increased reliability level, and presence of coils inside the vessel essentially reduces this reliability. therefore, at the next stage in the world project it is proposed for demo, first of all, to develop a poloidal tokamak system, so that the controlled plant has an acceptable level of input-output controllability for the design of a sufficiently reliable control system of the plasma position, current, and shape. ü the risk from feedback controller synthesis, which does not provide sufficient stability margin and does not provide the required performance. in this case, to minimize this risk, the developers of the control system should first analyze the sources of poor operability of the controller. such sources may be the following: plasma magnetic control systems in d-shaped tokamaks and imitation digital 11 copyright ©2022 assa. adv. in systems science and appl. (2022) o insufficient accuracy of the plant model on which the control system is synthesized; o insufficient accuracy of reduction of either the plant model or of the controller itself; o failure to take into account all the properties of the plant during synthesis of the controller, which affect its operability, for example, failure to take into account all disturbances, which affect the controlled plant, plant varying parameters, nonlinearities, saturated input and output signals, etc. in this case, the developers must identify the cause of the unsuccessful synthesis of the controller, redesign the controller, again simulate the entire system. synthesis of the system should be iterative and cover not only modeling, but also application to the controlled plant. fig. 10 shows the controller design cycle, which covers a number of stages [26]. in general, it should be noted that in thermonuclear tokamak reactors an explosion is impossible, because in case of emergency it is possible to close the valve of gas supply to the vacuum vessel and the process development will stop. therefore, fusion electric power plants are safe sources of energy in comparison with nuclear power plants. fig. 10. design cycle of dynamic plant control systems to achieve the required stability margins and performance of systems [26] 10. conclusion as follows from the indicated novelty of the direction and plenary reports at the ifac 2020 world congress, the magnetic plasma control systems in the globus-m2 tokamak will be developed in accordance with the world trends in the development of automatic control systems: • digitalization, automation, and elements of artificial intelligence lead to the autonomy of systems, which may give the human independent systems. the direction involves the use of artificial neural networks to reconstruct the plasma equilibrium, which are elements of artificial intelligence; • creation of digital twins of control systems to be created on the digital simulation platform; • the increasingly widespread use of machine learning, since the amount of information processed, is constantly increasing, for example, in iter, the codac information management system will process about 1 million signals. in this direction, similar methods will be used to 12 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) configure neural networks both on evolutionary plasma models and according to the data from real experiments. nobody knows the models of the plasma with thermonuclear reaction and so nobody can guarantee the operability of any known plasma magnetic control system. so, one should design plasma magnetic control systems for iter and demo having a feature of automatic adaptation to a fusion plasma new model. the design of such control systems is in line with our basic research direction which corresponds to the world development of control systems: combination of automation, digitalization, artificial intelligence giving autonomy. acknowledgements the author thanks the admins of v.a. trapeznikov institute of control sciences of the russian academy of sciences and lomonosov moscow state university for purchasing for two industrial speedgoat computers to design imitation digital computer platform in real time to develop plasma current and shape control system on the globus-m2 tokamak. this work was partially funded by the grant of the russian science foundation № 17-1901022 to develop plasma magnetic control systems in tokamaks. the plasma control systems developed in the framework of that grant (see section 3) are supposed to be installed into the digital imitation platform (see section 6) to operate in real time. references 1. wesson, j., campbell, d. (2004) tokamaks. international series of monographs on physics, clarendon press. 2. ariola, m., pironti, a. (2016) magnetic control of tokamak plasmas. 2nd ed. cham: springer international publishing switzerland. 3. mitrishkin, y.v., kartsev, n.m., kuznetsov, e.a., korostelev, a.y. (2020) metodi i sistemi magnitnogo upravleniya plazmoi v tokamakakh [methods and systems of plasma magnetic control in tokamaks] мoscow, russia, krasand [in russian]. 4. mitrishkin, y.v., kartsev, n.m., prokhorov, a.a., pavlova, e.a., korenev, p.s., konkov, a.e., kruzhkov, v.i., ivanova, s.l. (2020). tokamak plasma models development for plasma magnetic control systems design by first principle equations and identification approach. proc. of 14th int. symp. «intelligent systems», intels’20, moscow, russia. 5. mitrishkin, y.v., prohorov, a.a., korenev, p.s., patrov, m.i. (2017) sposob formirovfniya modeli magnitnogo upravleniya formoy i tokom plasmi s obratnoy svyazyu v tokamake. [method of formation of the model of magnetic control of plasma shape and current with feedback in a tokamak]. patent for invention of the russian federation no. 2702137. [in russian]. 6. mitrishkin, yu.v., korenev, p.s., prokhorov, a.a., kartsev, n.m., patrov, m.i. (2018) plasma control in tokamaks. part 1. controlled thermonuclear fusion problem. tokamaks. components of control systems, advances in systems science and applications, 18(2), 26–52, https://doi.org/10.25728/assa.2018.18.2.598 7. mitrishkin, y.v., kartsev, n.m., pavlova, e.a., prohorov, a.a., korenev, p.s., patrov, m.i. (2018). plasma control in tokamaks. part. 2. magnetic plasma control systems, advances in systems science and applications, 18(3), 39–78. 8. mitrishkin, y.v., and haber, r.e. (2009). hierarchical control system for complex dynamical plants // proc. of the 1st international workshop on networked embedded and control system technologies: european and russian r&d cooperation – nester 2009, milan, italy, 56-65. insticc press, portugal, a. pascal and v. ufnarovsky (eds). isbn 978-989-674-004-7. plasma magnetic control systems in d-shaped tokamaks and imitation digital 13 copyright ©2022 assa. adv. in systems science and appl. (2022) 9. mitrishkin, y., kartsev, n., korenev, p., patrov, m. (2020). model predictive control with time varying parameters for plasma shape and current in a tokamak. proc. international federation of automatic control 2020 world congress (ifac 2020 wc), berlin, germany, paper-id 343. 10. konkov, a., mitrishkin, y., korenev, p., patrov, m. (2020). robust cascade lmi design of mimo controllers for plasma position, current, and shape model with time-varying parameters in a tokamak. proc. ifac 2020 world congress, berlin, germany, paper-id 356. 11. garcia-sanz, m. (2017). robust control engineering: practical qft solutions. boca raton: crc press. 12. mitrishkin, y., pavlova, e., patrov, m. (2021) design and comparison of plasma h∞ loop shaping and rga-h∞ double decoupling multivariable cascade magnetic control systems for a spherical tokamak. advances in systems science and applications, 21(1), 22–45. 13. mitrishkin, y.v., prohorov, a.a., korenev, p.s., patrov, m.i. (2020). plasma magnetic timevarying nonlinear robust control system for the globus-m/m2 tokamak, control engineering practice (a journal of ifac, elsevier), 100, 104446. 14. korenev, p.s., mitrishkin, y.v., patrov, m.i. (2016). rekonstruktsiya ravnovesnogo raspredeleniya parametrov plasmi tokamaka po vneshnim magnitnim izmereniyam i postroyeniye lineynikh plazmennikh modeley. [reconstruction of equilibrium distribution of plasma parameters on the base of external magnetic measurements and construction of plasma linear models]. mechatronics, automatization, and control, 17 (4), 254–265, [in russian]. 15. mitrishkin, y.v., prokhorov, a.a., korenev, p.s., and patrov, m.i. (2019). hierarchical robust switching control method with the equilibrium reconstruction code based on improved moving filaments approach in the feedback for tokamak plasma shape, fusion engineering and design, 138, 138–150, https://doi.org/10.1016/j.fusengdes.2018.10.031 16. prokhorov, a., mitrishkin, y., korenev, p., patrov, m. (2020). the plasma shape control system in the tokamak with the neural network as a plasma equilibrium reconstruction algorithm, proc. ifac 2020 world congress, berlin, germany, paper-id: 1690. 17. mitrishkin, y., kartsev, n., konkov, a., & patrov, m. (2020). plasma control in tokamaks. part 3.1. plasma magnetic control systems in iter. advances in systems science and applications, 20(2), 82-–97. https://doi.org/10.25728/assa.2020.20.2.892 18. mitrishkin, y., kartsev, n., konkov, a., & patrov, m. (2020). plasma control in tokamaks. part 3.2. simulation and realization of plasma control systems in iter and constructions of demo. advances in systems science and applications, 20(3), 136–152. 19. eurofusion (2018). european research roadmap to the realization of fusion energy [online] – available https://www.euro-fusion.org/eurofusion/roadmap/ 20. gryaznevich, m. p., chuyanov, v.a., kingham, d., et al. (2015). advancing fusion by innovations: smaller, quicker, cheaper, journal of physics: conference series, 591, 012005, https://doi.org/10.1088/1742-6596/591/1/012005 21. chuyanov, v., & gryaznevich, m. (2017). modular fusion power plant. fusion engineering and design, 122, 238–252. 22. humphreys, d. a., casper, t. a., eidietis, n., et al. (2009). experimental vertical stability studies for iter performance and design guidance. nuclear fusion, 49(11). 23. mertens, v., raupp, g. & treutterer w. (2003) chapter 3: plasma control in asdex upgrade, fusion science and technology, 44(3), 593–604, doi: 10.13182/fst03-a401. 24. gusev, v.k., azizov, e.a., alekseev, a.b., et al. (2013). globus-m results as the basis for a compact spherical tokamak with enhanced parameters globus-m2, nuclear fusion, 53(9), 093013. 25. mitrishkin, y.v., korenev, p.s., kartsev, n.m., kuznetsov, e.a., prohorov, a.a., patrov, m.i. (2019). plasma magnetic cascade multiloop control system design methodology in a tokamak. control engineering practice, 87, 97–110. 14 y.v. mitrishkin copyright ©2022 assa. adv. in systems science and appl. (2022) 26. mitrishkin, y.v., efremov, a.a., zenkov, s.m. (2013) experimental test bed for real time simulations of tokamak plasma control systems. journal of control engineering and technology. 3(3), 121–130. 27. mitrishkin, y.v., korenev, p.s., prohorov, a.a., patrov, m.i. (2017) robust h∞ switching mimo control for a plasma time-varying parameter model with a variable structure in a tokamak. proc. ifac 2017 world congress, toulouse, france, ifac papersonline, 50(1), 11385–11390. 28. skogestad, s., postlethwaite, i. (2005) multivariable feedback control: analysis and design 2nd ed., john wiley & sons, ltd. adv syst sci appl 2021; 02:71–82 published online at https://ijassa.ipu.ru. algorithm for optimal two-link trajectory planning in evasion from detection problem of mobile vehicle with non-uniform radiation pattern andrey galyaev1, pavel lysenko1,2*and victor yakhno1 1institute of control sciences of ras, moscow, russia, 117997 2moscow institute of physics and technology, dolgoprudny, russia, 141701 abstract: the problem of an optimal pp/tp (path/trajectory planning) evasion of a mobile vehicle from the detection is considered. the objective is to minimize the risk of a moving vehicle being detected by a static sensor when moving between two specified points on the plane. detection is based on the primary acoustic field emitted by a vehicle with an inhomogeneous radiation pattern. an algorithm for finding a two-link optimal trajectory is proposed. the optimal trajectory and the law of speed of a mobile vehicle, as well as the value of the criterion, are found. keywords: uuv path/trajectory planning; non-detection probability; non-uniform radiation pattern, evasion from detection 1. introduction the problems of path/trajectory planning of autonomous and controlled vehicles are currently interesting due to their wide usage. for example, in the evasion problem, the moving object must remain hidden for the detector system [1]. these problems are extremely relevant for modern military applications. technological advances allow the onboard algorithms of unmanned mobile vehicles to be non-trivial and use solutions from broad scientific fields, such as optimal control and differential games. the problem of evading radar detection, in particular, is considered in [2] as a problem of automated planning of the trajectory of combat unmanned aerial vehicles (uavs) in the presence of radar-guided surface-to-air missiles. it turns out that the solution of the problem significantly depends on the assumption of the isotropic capabilities of the radar (i.e., the independence of the reflected signal from the spatial orientation of the uav). in case of isotropic characteristics, an analytical solution for the problem is possible, which facilitates the construction of an onboard control algorithm. homogeneous radiation capabilities of mobile vehicle are explored in papers [3–8]. the paper considers the problem of the covert infiltration of an unmanned underwater vehicle (uuv) into a given area under the supervision of a stationary passive sonar conducting a search on the primary hydro-acoustic field as in [9]. in contrast to previous works dealing with path planning problems, the acoustic signal emitted by a mobile vehicle has a non-uniform pattern. the task of the uuv optimal route planning from the starting point of space to the end point can be formulated as a problem of calculus of variations. the optimization criterion in this problem is the uuv non-detection probability, e.g. the chance of the sonar to make a decision about uuv presence in the region under study during ∗corresponding author: pashlys@yandex.ru 72 a. galyaev, p. lysenko uuv’s passing along the route. the decision rule in threshold statistics is examined to make the detection decision [10]. let the uuv trajectory be represented by a piecewise linear trajectory with the duration of a movement on each rectilinear section equal to the duration of the observation ”tact”. for each clock cycle, we denote the judgment ”uuv absent” by 0, and the judgment ”uuv present” by 1. then the results of making decisions on the uuv trajectory can be represented by a sequence of zeros and ones. uuv will not be detected on the path if there are no ones in the corresponding sequence of zeros and ones on the selected path. the optimal trajectory is the one for which the probability of such an event is maximum. articles [5, 11, 12] are devoted to the problems of optimal planning of the uuv route under threat conditions, in which other probabilistic and energy criteria are implemented. at codit 2020 conference results for optimal one-link trajectories of mobile vehicle, evading from one sonar, were presented [13]. the detection of sonar is based on primary hydro-acoustic field signals. but later in [14] it was shown, that sufficient optimal conditions can brake and then multi-link trajectory becomes optimal. in that article problem statements for two-link optimal trajectories were discussed. the current article considers the algorithm for solving of the supplementary problem, needed for two-link optimal trajectories constructing. firstly, in the article the non-detection probability of uuv on the route is derived. secondly, a path planning problem is formulated and analytically solved. later in next chapter the special case of two-link optimal trajectories is explored and the algorithm for finding such trajectories is obtained. finally, several examples are presented. 2. non-detection probability of uuv under passive surveillance the uuv covertness on the selected trajectory with a given speed law change can be characterized by the probability pnd that during the passage of the route it will not be detected at any tact. this probability is denoted pnd and will be called the non-detection probability of uuv on the trajectory. in the case of one sonar this probability is [3] pnd = ∏ j fn  χ2 1−α,n 1 + σ2 0(υj/υ0) µr20 σ2 nr 2 j γ  , (2.1) where fn is χ2 is a probability distribution with n degrees of freedom, α is a false alarm probability, γ is an attenuation coefficient, rj is the distance from uuv to sonar at j tact, starting from the moment of appearance of uuv on the trajectory, and υj is a constant speed of uuv at this tact, σ0, σn, µ, r0, υ0 are some model parameters. the number of factors in the product is equal to the number of tacts when moving uuv along a trajectory. in the case of several sonars, the product of expressions of the form (2.1) over all available sonars is taken [1]. thus the formalization of the uuv covertness concept is reduced to the formula (2.1). now the task is to construct the optimal trajectory and the optimal law of the uuv velocity variation, maximizing the probability pnd. based on actual algorithms for information processing, using decisive threshold rules, in the article [13] a formula was obtained for calculating the risk of mobile vehicle detection. it is represented as an integral, which we call the threat functional [6], [13] in the problem of uuv route planning for the case of passive sonar r = ∫ t0 0 σ2 s(t) σ2 n(t) dt. (2.2) copyright © 2021 assa. adv syst sci appl (2021) algorithm for optimal two-link trajectory planning in evasion 73 this conclusion coincides with the result given in [3] for σ2 s = σ2 0 (υj/υ0) µ r20 r2 γ, and σ2 n(t) = const. (2.3) moreover, for the path planning task, it is not required to know the exact values of the information processing parameters because they are not included in (2.2). 3. path planning problem 3.1. risk functional to formulate the path planning problem as the problem of calculus of variations let us rewrite the risk functional (2.2). due to (2.3), after omitting the constants the expression in integral can take a general form as a power model σ2 s(t) σ2 n(t) ∼ vµ rk . (3.4) as well as in [1] from the physical point of view this function is the instantaneous level of signal s, received by the sensor. it depends on the current distance between sensor and evading object r, for some types of physical fields – on absolute instant velocity of the object v too and can be represented as s = vµ rk . (3.5) the exponents k and µ characterize the physical field used for detection. depending on values of k and µ this can be magnetic, thermal, acoustic or electromagnetic fields. moreover, this signal depends on the receiving diagram of the sensor and the radiation pattern of the object s = vµ rk a0(ϕ)g(ψ, ϕ), (3.6) where multipliera0(ϕ) is responsible for diagram of sensor antenna and g(ψ, ϕ) is a radiation pattern of the moving object. the risk r is the integral value of this signal, so the criterion of optimization (2.2) is a function of phase coordinates of the object and qualities of the sensor and the object itself: r = ∫ t0 0 ( vµ rk a0(ϕ)g(ψ, ϕ) ) dt. (3.7) we consider sensor antenna diagram to be homogeneous, so a0(ϕ) ≡ 1. the geometric meaning of angles ψ and ϕ is as follows. here ψ is the angle of rotation of object’s velocity vector and ϕ is the angle of rotation of radius vector as shown in fig. 3.1. we study the case of k = 2 and µ = 2 for acoustic field. non-uniform radiation pattern of the object can be described as g(ψ, ϕ) = g(β). thus signal expression (3.6) has a form s = v2 r2 g(β). copyright © 2021 assa. adv syst sci appl (2021) 74 a. galyaev, p. lysenko fig. 3.1. the moving object in the cartesian coordinate system with sensor s. fig. 3.2. rectangular triangle of velocity vectors ~v,~vr and ~vϕ. 3.2. mathematical statement of the problem therefore we can formulate the optimal path planning problem. problem 3.1: it is required to find such trajectory (r(t), ϕ(t)), which minimizes the functional r = ∫ t0 0 v2 r2 g(β)dt, (3.8) where v is a velocity of the object, r is the distance between the sensor and the object. boundary conditions are r(0) = xa, r(t0) = rb, ϕ(0) = ϕa, ϕ(t0) = ϕb. (3.9) time t0 of moving on route from point a to b is fixed. the detection system consists of one static sensor placed in the origin of cartesian coordinate system with x axis passing point a. the polar coordinate system is used for solution simplicity. it is easy to see that ψ − ϕ = β is the angle in triangle built from radial vr and transversal vϕ velocities of the object, e.g. the angle between velocity of the object and its projection on radius vector as shown in fig. 3.2. next lemma is valid. lemma 3.1: substitution of variable ρ = ln r brings functional (3.8) to the form r(ρ(·), ϕ(·)) = ∫ t0 0 sdt = ∫ t0 0 (ρ̇2 + ϕ̇2)g ( arctan ϕ̇ ρ̇ ) dt. (3.10) copyright © 2021 assa. adv syst sci appl (2021) algorithm for optimal two-link trajectory planning in evasion 75 a proof for lemma 3.1 can be found in [13]. instead of solving the original problem 1 according to lemma 2 it is necessary to solve a two point boundary value variational problem on the minimum of the functional (3.10). problem 3.2: it is required to find the trajectory (ρ∗(t), ϕ∗(t)), which minimizes the functional r(ρ(·), ϕ(·)) = ∫ t 0 ( ρ̇2 + ϕ̇2 ) g ( arctan ϕ̇ ρ̇ ) dt→ min ρ(·),ϕ(·) . (3.11) with boundary conditions ρ(0) = ρa, ρ(t ) = ρb, ϕ(0) = ϕa, ϕ(t ) = ϕb. 4. solution of the path planning problem 4.1. the necessary optimality conditions theorem 4.1: suppose that 0 < g1 < g(β) < g2 for all β ∈ [0, 2π] is a twice continuously differentiated function of β, where g1, g2 are some constant values, and ρ̈(t), ϕ̈(t) exist and are continuous functions of t. then the extremal trajectory satisfies the following system of equations{ ρ̇ = const, ϕ̇ = const. (4.12) thus the extremal trajectory, which can be the solution for problem 3.2 can be represented in the parametric form of logarithmic spiral r(t) = ra exp ( t t0 ln rb ra ) , ϕ(t) = ϕa + ϕb − ϕa t0 t. (4.13) on the plane (r, ϕ) this line has a form r(ϕ) = ra exp ( ϕ− ϕa ϕb − ϕa ln rb ra ) . (4.14) next two lemmas define the velocity law of the vehicle on the extremal trajectory and the risk value on it. lemma 4.1: the velocity law on the extremal trajectory equation (4.14) is represented as follows v(t) = ra t exp ( t t ln rb ra )√ ln2 rb ra + (ϕb − ϕa)2. (4.15) lemma 4.2: the explicit dependence risk (3.11) from boundary conditions on the extremal trajectory equation (4.14) has the form r∗ = (ρb − ρa)2 + (ϕb − ϕa)2 t g(β0) = l2 t g(β0), (4.16) copyright © 2021 assa. adv syst sci appl (2021) 76 a. galyaev, p. lysenko where β0 = arctan ϕb − ϕa ρb − ρa , and l = √ (ρb − ρa)2 + (ϕb − ϕa)2 is the length of the straight line segment between points a and b in the space (ρ, ϕ). the proofs of lemmas 4.1 and 4.2 remain valid as shown in reference [9]. 4.2. the sufficient optimality conditions the sufficient optimality conditions for path planning problem 3.2 and the hessian matrix explicit form are presented in [9] as follows lemma 4.3: let s(ρ, ρ̇, ϕ, ϕ̇, t) = ( ρ̇2 + ϕ̇2 ) g (β), g(β) – thrice continuously differentiated function of β, then hessian matrix h equals h = ( h11 h12 h21 h22 ) where h11 = ∂2s ∂ρ̇2 = 2g (β)− 2ϕ̇ρ̇g′ (β)− g′′ (β) ϕ̇2 ϕ̇2 + ρ̇2 , h12 = ∂2s ∂ρ̇∂ϕ̇ = ( ρ̇2 − ϕ̇2 ) g′(β)− g′′(β) ϕ̇2 + ρ̇2 , h21 = ∂2s ∂ϕ̇∂ρ̇ = ( ρ̇2 − ϕ̇2 ) g′(β)− g′′(β) ϕ̇2 + ρ̇2 , h22 = ∂2s ∂ϕ̇2 = 2g (β) + 2ϕ̇ρ̇g′ (β) + g′′ (β) ρ̇2 ϕ̇2 + ρ̇2 , (4.17) and the hessian itself is the determinant of the matrix deth = 4g2(β) + 2g(β)g′′(β)− g′2(β). (4.18) theorem 4.2: assume that the conditions of theorem 4.1, lemma 4.3 are satisfied, and the inequality deth > 0 is valid for all values β. then the optimal trajectory given by equation (4.12) brings the strong minimum to the risk functional equation (3.11). 5. algorithm for finding two-link optimal trajectories if conditions of theorem 4.2 are not fulfilled, e.g. deth ≤ 0, then optimal trajectory can consist of many segments of logarithmic spirals. the current article explores the case of two segments, however it can be shown that any optimal multi-link trajectory is constructed from segments of two base directions. as explained above, these segments in (ρ, ϕ) space transform into straight lines. this fact is illustrated in figure 5.3. copyright © 2021 assa. adv syst sci appl (2021) algorithm for optimal two-link trajectory planning in evasion 77 0 ρ− ρ0 ϕ− ϕ0 l0 l1 l2 β0 β1 β2 fig. 5.3. the mobile vehicle in the (ρ, ϕ) coordinate system. next lemma allows to choose two optimal angles β1 and β2 from the whole set of such angles, fulfilling the boundary conditions of the problem. a corresponding to these angles optimal risk is denoted as r∗(β1, β2) and is calculated according to lemma 4.2. lemma 5.1: angles β∗1 , β ∗ 2 for optimal trajectory can be found from system cos(β2 − β1) = cos(ξ(β1)− ξ(β2)), g(β2) g(β1) = cos2 ξ(β2) cos2 ξ(β1) , (5.19) where function ξ(β) = arctan 1 2 g′(β) g(β) . we introduce two new functions χ(β) = β + ξ(β), η(β) = g 1 2 (β) cos ξ(β) (5.20) and rewrite the system (5.19) as { χ(β1) = χ(β2), η(β1) = η(β2). (5.21) notice that cos ξ(β) > 0 according to definition function ξ(β) in lemma 5.1. thereby the search of optimal two-link trajectories is reduced to the solution of the system (5.21). unfortunately, it can not be solved analytically, so we have developed a specific algorithm for finding solutions of system (5.21). first of all, let us investigate each of the functions χ(β) and η(β) by extreme’s and find their first derivatives dχ(β) dβ = 4g2(β)− (g′(β))2 + 2g′′(β)g(β) 4g2(β) + (g′(β))2 . the numerator of this fraction is a part of expression for deth, so dχ(β) dβ = 0 ⇐⇒ deth = 0. copyright © 2021 assa. adv syst sci appl (2021) 78 a. galyaev, p. lysenko as for η(β) function we have dη(β) dβ = g 1 2 (β)g′(β) 4g2(β) √ 4g2(β) + g′2(β) deth from last expression follows dη(β) dβ = 0 ⇐⇒ deth = 0 or g′(β) = 0. the illustration of the theorem 4.2 implementation is as follows. if deth > 0 for all β, then dχ(β) dβ > 0, i.e. the function χ(β) increases monotonically and therefore there is no solution for the first of the equations (5.21) and one-link trajectory is optimal. now we propose algorithm for finding system (5.21) solution. we consider β ∈ [0◦, 450◦] and denote all intervals β ∈ (a2i−1, a2i), i = 1, .., n , where deth > 0. algorithm 1: algorithm for finding optimal values of β1, β2 result: β∗1 , β∗2 initialization β1 = a1, β2 = a3, r∗ = r(β1, β2) while i < n , i = i+ 1 do k = i+ 1; while k < n , k = k + 1 do find system (5.21) single solution (β1, β2), β1 ∈ [a2i−1, a2i], β2 ∈ [a2k−1, a2k]; if such solution exists then calculate r(β1, β2); if r(β1, β2) < r∗ then r∗ = r(β1, β2); (β∗1 , β ∗ 2) = (β1, β2); else continue end else continue end end end the proposed algorithm has the following advantages. it looks for solutions on reduced β domains only where deth > 0 and additionally where solution is possible. all possible optimal directions are found, regardless of the initialization values of the algorithm, as will be shown in the example. the optimal value of the criterion depends on the directions β1 and β2 and the length of the corresponding trajectory segments. moreover, that is an efficient algorithm. the only other way to find optimal angles β1, β2 is a brute force approach of sorting through all possible two-link trajectories, which is a much more time-consuming process, compared to the algorithm above. 6. example this section demonstrates the use of proposed algorithm in one model case. the considered algorithm has been implemented in matlab. also matlab scripts have been developed to validate and illustrate obtained results. copyright © 2021 assa. adv syst sci appl (2021) algorithm for optimal two-link trajectory planning in evasion 79 figures 6.4 and 6.5 show a fairly simple radiation pattern of a mobile vehicle g(β) = 1 + 0.3 cos(4β) 1.3 . (6.22) fig. 6.4. radiation pattern of the mobile vehicle on cartesian plane fig. 6.5. radiation pattern of the mobile vehicle as a function of β figure 6.4 shows a radiation pattern on a cartesian plane. this pattern is symmetrical in respect to both cartesian axes. the solid line represents the function g(β) level line. the value g(β) is a radius-vector modulo directed from zero point to point lying on the shown level line. figure 6.5, on the other hand, presents a radiation pattern as a dependence from β. figures 6.6 and 6.7 demonstrate functions χ(β) and η(β) respectively. blue color solid line segments on those graphs are assigned to intervals β ∈ (a2i−1, a2i), i = 1, .., n , where deth > 0. we see that in the example n = 5. fig. 6.6. function χ(β) copyright © 2021 assa. adv syst sci appl (2021) 80 a. galyaev, p. lysenko fig. 6.7. function η(β) let us introduce notation a′i for further exploration. we declare that for β ∈ [a2i−1, a2i] and fixed k a′2i−1 = max(a2i−1, arg(χ(β) = χ(a2k−1))), a′2i = min(a2i, arg(χ(β) = χ(a2k))). this notation takes place because χ(β) is monotonically increasing function on β ∈ [a2i−1, a2i]. further we need to narrow down the exploration area on β of functions χ(β), η(β). figures 6.8 and 6.9 correspond to the first pair of narrowed intervals of hessian positivity [a′1, a ′ 2] and [a′3, a ′ 4]. black dots show found at this iteration of the algorithm 1 solutions β1, β2. fig. 6.8. enlarged segment of function χ(β). fig. 6.9. enlarged segment of function η(β). so the domain of values of function χ(β) when β belongs to the first interval [a1, a2] do not cross with the same one when β belongs to the third, forth and fifth interval [a5, a6], [a7, a8], [a9, a10], then the algorithm continues searching the solution comparing intervals two and three on β, then three and four, and so on. candidates of optimal pairs of (β1, β2) are approximately (59◦, 121◦), (149◦, 211◦), (239◦, 301◦), (329◦, 391◦). these directions and trajectories are presented in figure 6.10. if the direction β0 lies in the domain where deth > 0 then the optimal trajectory is one-link one. on the contrary if the direction β0 lies in the domain where deth < 0 then two-link trajectory is optimal. found by algorithm 1 optimal pairs (β1, β2) in every domain where deth < 0 are the same as shown on figure 6.10. copyright © 2021 assa. adv syst sci appl (2021) algorithm for optimal two-link trajectory planning in evasion 81 fig. 6.10. optimal trajectories for different β0. 7. conclusions the problem of minimizing the risk of uuv detection by a static sensor while moving between two given points on a plane is solved. the detection is based on the primary acoustic field radiated by the vehicle with a non-uniform radiation pattern, which differs this work form the previous researches in the area. in the first part of the article, the non-detection probability is derived. further it is used as an optimization criterion in the path planning problem. the optimal trajectory and velocity law of the uuv are found, as well as the risk value. a case of the two-link optimal trajectories is studied, when sufficient optimal conditions are not fulfilled. a special algorithm has been developed for finding optimal angles of logarithmic segments. it allows to solve the whole problem and find optimal angles β1, β2. in practical applications this algorithm allows to preprocess these angles for all initial conditions of β0 for a particular radiation pattern g(β) of the mobile vehicle, store it in memory and use later this information for onboard calculations. this approach saves a lot of computational power, which is very important for resource-efficient missions. acknowledgements the work of a.a.g. was partially supported by programm of basic research of ras; the work of p.v.l. was partially supported by russian foundation for basic research grant 2038-90215; the work of v.p.y. was partially supported by programm of basic research of ras. the authors declare that there is no conflict of interest. references 1. galyaev, a. a. & maslov, e. p. (2010) optimization of a mobile object evasion laws from detection. j. comput. syst. sci. int., 49, 560–569. 2. kabamba, p. t., meerkov, s. m., & zeitz, f. h. (2006) optimal path planning for unmanned combat aerial vehicles to defeat radar tracking. j guid control dynam, 29, 279–288. 3. sysoev, l. p. (2011) detection probability criterion on the path for mobile object control problem in conflict environment. autom remote control, 72, 65–72. copyright © 2021 assa. adv syst sci appl (2021) 82 a. galyaev, p. lysenko 4. dogan, a. & zengin, u. (2006) unmanned aerial vehicle dynamic-target pursuit by using probabilistic threat exposure map. j guid control dynam, 29, 944–954. 5. galyaev, a. a., lysenko, p. v., & yakhno, v. p. (2020) moving object evasion from single detector at given speed. probl. upr., 83–91. 6. zabarankin, m., uryasev, s., & pardalos, p. (2002) optimal risk path algorithms. in murphey, r. & pardalos, p. (eds.), cooperative control and optimization, (pp. 273– 298), dordrecht: kluwer acad. 7. sidhu h, m. g. & m, s. (2006) optimal trajectories in a threat environment. jbt , 9, 33–39. 8. pachter, l. s. & pachter, m. (2001) optimal paths for avoiding a radiating source. in proceedings of the 40th ieee conference on decision and control, vol. 4, 3581–3586. 9. galyaev, a. a., dobrovidov, a. v., lysenko, p. v., shaikin, m. e., & yakhno, v. p. (2020) path planning in threat environment for uuv with non-uniform radiation pattern. sensors, 20. 10. lehmann, e. l. & romano, j. p. (2005) testing statistical hypotheses. springer texts in statistics, new york: springer. 11. rubinovich, e. y. & andreev, k. v. (2016) moving observer trajectory control by angular measurements in tracking problem. autom remote control, 77, 106–129. 12. mercer, g. n. & sidhu, h. s. (2007) two continuous methods for determining a minimal risk path through a minefield. in read, w. & roberts, a. j. (eds.), proceedings of the 13th biennial computational techniques and applications conference, ctac-2006, vol. 48 of anziam j., 293–306. 13. galyaev, a. a., lysenko, p. v., & yakhno, v. p. (2020) evasion from detection of the moving object with non-uniform radiation pattern. in 2020 7th international conference on control, decision and information technologies (codit), vol. 1, 463–468. 14. galyaev, a. a., lysenko, p. v., & yakhno, v. p. (2021) 2d optimal trajectory planning problem in threat environment for uuv with non-uniform radiation pattern. sensors, 21. copyright © 2021 assa. adv syst sci appl (2021) introduction non-detection probability of uuv under passive surveillance path planning problem risk functional mathematical statement of the problem solution of the path planning problem the necessary optimality conditions the sufficient optimality conditions algorithm for finding two-link optimal trajectories example conclusions adv syst sci appl 2021; 01:113–138 published online at https://ijassa.ipu.ru. software for testability analysis of aviation systems valentina s. viktorova1∗, armen s. stepanyants1 1v. a. trapeznikov institute of control sciences of russian academy of sciences, moscow, russia abstract: this paper describes models, methods and software tool for testability analysis of aviation systems. it comprises analytical and programing aspects of calculations of main testability, reliability and availability indices. it presents general description of the software, xml schema of input data and technique of their mapping to the database structure. the procedure for generating the initial data for testability analysis based on the line replaceable units failure modes report is described. fault tree model for analysis of the built-in test conformity is suggested. markov models have been created for analyzing reliability and availability, taking into account the features of the built-in test and the specifics of the aircraft operation. an approach to the construction of trends in the operative availability of aviation systems in the inter-maintenance interval is proposed. keywords: testability, built-in-test, fault coverage, failure detection ratio, built-in-test equipment conformity, failure mode and effects analysis, markov reliability models, operative availability trend 1. introduction reliability, availability, testability, maintainability are the main properties of aviation items that determine the safety of an aircraft flight [1]. these properties are interrelated (chapter 6 in [2] ), an improvement (worsening) of one leads to a corresponding changes in the others. testability is a property of an item that characterizes its adaptability to be controlled by specified diagnostic systems. these systems include both internal and external diagnostic means. internal diagnostic means integrated on-line referred to as built-in-test (bit). in one of the first regulatory documents in this area [3], the targets of testability analysis are defined as assessing the ability to detect and isolate system faults to the faulty replaceable assembly level. our research focuses on design stage advanced testability analysis, including investigation of impact of bit performance on the reliability and availability of aircraft systems. analytical models that consider imperfect bit performance are known as imperfect fault coverage models. many articles are devoted to these models and associated reliability analysis techniques [4–8]. we have developed a reliability model considering imperfect fault coverage and aircraft-specific recovery strategies for the latent failures. when constructing the reliability models, a technical inspection and maintenance program [9, 10] typical for civil aviation was taken into account. at the aircraft design stage, it is mandatory to carrying out a qualitative and preliminary quantitative failure modes and effects analysis (fmea) [11]. if the fmea table records are supplemented with information on the failure detection methods [13–15], then they can be used as input data for calculating the testability indices and conducting reliability and availability analysis taking into account the bit characteristics. testability-oriented fmea ∗corresponding author: vsviktorova@gmail.com 114 v.s. viktorova, a.s. stepanyants analysis is performed for all functional aircraft systems on the level of the line replaceable units (lru) [16]. storage, selection, sorting and filtering of data of such complexity and volume as aircraft fmea reports can be carried out only with the involvement of modern database management systems. implementation of code of definition and solving reliability and testability analysis models is only possible through the advanced programming languages with high precision numeric data types and support of dynamic data structures. we chose oracle dbms and pl/sql programming language. in this paper, we present software for joint testability and reliability analysis of civil aircrafts. the models and analysis methods used in this software reflect the experience we gained during the project testability analysis of the sukhoi superjet and mc-21 aircraft family. this software provides the following options • testability indices calculation; • construction of distributions of testability indices by levels of criticality of failures; • construction of distributions of testability indices by fault detection methods and by impact on departure delay; • evaluation of built-in-test equipment conformity; • automatic construction of markov reliability models of redundant lru assemblies with imperfect fault coverage and aviation-specific recovery scenarios; • investigation of the impact of bit functional features on system reliability and availability; • creation of operative availability trends within the specified time frame between the aircraft maintenance periods. the rest of this paper is organized as follows. in section 2, we give the software description including meta definition of testability data and their mapping to the database structure. section 3 presents mathematical background of the testability&reliability analysis. in section 3.1, we derive the formulas for calculating single testability measures. in section 3.2, we define logical-probability model of bit conformity. in section 3.3, we suggest the parametric markov reliability model of redundant lru assemblies with imperfect fault coverage by bit. section 3.4 describes an approach to trending of aircraft operative availability. some results and graphical illustrations of the analysis are summarized in section 4. in the appendix, we provide the xml schema of testability data. 2. testability analysis software description 2.1. testability metadata structure the actual practice of database design is the preliminary creation of a platform-independent meta-definition of the data. in accordance with the recommendations of international reference books [17–19] on technical publications in aerospace industries, the xml language is used for the meta definition of information items of this software. special language referred to as xml schema definition (xsd) [20] are used for description of the structure, data types, elements and attributes of an xml document. we have developed the xml schema that defines the xml image of testability data testability data xml schema definition (txsd). the pilot version of the txsd was described in [21]. current txsd is presented in appendix. it defines four-level hierarchical structure with the help of four global complex types tprojects, tsystems, tlrus, tfmeas. each of them is an unbounded sequence of elements of type tproject, tsystem, tlru, tfmea, respectively. the complex type tproject defines project. project is a generic term. an aircraft as a whole, arbitrary functional systems, for example, air conditioning, fire protection, integrated modular avionics can be considered as a project. subsystems and assembles, for example, a landing gear steering or an engine interface unit of the central computer module of the indicating/recording system also may be represented as a project. thus, the level of copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 115 detail of the analyzed object is determined by the user. the tproject context is divided into three parts descriptors (codes and names of the project), time parameters (maintenance periods, recovery times...), set of functional systems (tsystems). the tsystem context also is divided into three parts descriptors (codes and names of the system), technical specifications for the system reliability and testability indices, set of the line replaceable units (tlrus). the tlru context is divided into six parts descriptors (codes and names of the lru), lru part and serial numbers, lru identifier in the centralized maintenance system (cms), redundancy schema, technical specifications for the reliability and maintainability indices, set of failure entries (tfmeas). the tfmea type describes output data of design stage analysis of failure modes and effects of the lru. tfmea has complex structure. the elements of this structure define various properties of failure, including failure descriptors in cms, failure rate, failure severity level, detection method, phase of flight, flag of departure delay, bit resolution features and so on. txsd contains global simple types which define • string type of ata code [22] (tcodata) of system and lru; • decimal types of time parameters and reliability indices (trindex), normalized testability indices (tratio); • enumeration types of lru properties lru redundancy schema (trschema), lru mode (tlrutype); • enumeration types of failure properties detection method (tfdm), severity level (tseverity), flight profile [23] (tphases), mode (tlrufailure), departure delay flag (tdelay). txsd includes a specialized complex datatype, namely resolution list (tlruchain). it is an unbounded sequence of tcodata elements. tlruchain defines a list of replaceable units that will be isolated by bit when a failure is detected. a global element named ”projects” specifies the root of any txsd-based xml document. the complete composition of the txsd is given in the appendix. 2.2. testability database structure the mapping of txsd to a relational database schema is carried out according to the following rules: • global complex types tproject, tsystem, tlru, tfmea define the presence and structure of four main database tables storing information about projects, functional systems, lru, fmea; • simple type elements define the data types, identity, check constraints of columns of these tables; • elements of the sequence type tsystems, tlrus, tfmeas contained in complex elements project, system, lru respectively define referential integrity within the database; • nesting order of txsd elements determines relations by foreign key between parents and childes tables of the database; • enumeration type elements are embodied in look-up, reference tables. fig.2.1 presents graphical interpretation of the mapping procedure. 2.3. testability software features testability analysis software is a web application that runs on an oracle database (fig. 2.2). this application is built using oracle database’s native low-code development platform apex. this choice enables to build high-performance, scalable, secure apps. oracle apex uses a 3-tier architecture where requests are sent from the browser, through a web server, to the database. within the database, the request is processed by oracle apex. once the copyright c© 2021 assa. adv syst sci appl (2021) 116 v.s. viktorova, a.s. stepanyants processing is complete, the result is sent back to the browser. processes of requesting and submitting pages are realized through a modern implementation of the oracle net listener called oracle rest data service (ords). user interface of the software is mobile friendly. pages have smart layout and include forms, charts, reports and other user interface components that can work across varying screen resolutions. database includes main tables (see subsection 2.2), reference tables and views. views contains data from join of fmea and reference tables. specifications and bodies of analytical procedures and functions are contained in four pl/sql packages testpckg, relpckg, apckg, numpckg. testpckg, relpckg, apckg includes procedures and functions of testability, reliability, availability analysis respectively. numpckg implements numerical methods for mathematical computations . reports of testability and reliability analysis can be download into external files of csv, xlsx, rtf, xml, pdf formats. fig. 2.1. txsd to database schema conversion (fk – foreign key, pk – primary key). . 3. mathematical background of computational algorithms the following subsections describe models, methods and calculation procedures applied in the software for quantification and analysis of various testability, reliability and availability indicators. we use the generally accepted in aviation reliability analysis assumption about an exponential distribution of random times to failure [12]. in this case, the reliability indicator called the failure rate is equal to the constant parameter of the distribution. 3.1. single testability indicators single testability indicators (sti) are used to specify bit functionality or reliability. these indicators are used to assess a single bit property. when calculating sti, a preliminary partition of the entire set of the lrus failure modes into disjoint subsets is carried out. sti can obviously be estimated as the ratio of the number of selected failures modes (ns) to the total number (nt ) of failure modes copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 117 fig. 2.2. testability analysis software structure. . sti = ns nt . (3.1) however, for joint reliability modeling of the object under test and bit, sti should be set as the ratio of some probabilistic measures. the expediency of such definition is explained by the fact that when modeling reliability behavior, it will be possible to split the total flow of the object failures into several components interesting for researcher. for example, detected failures and hidden (latent) failures. in this case, sti can be defined as the conditional probability of the failure modes of interest, provided that the failure of the object has occurred: sti = (1− exp−λ̃st) (1− exp−λ̃t t) , (3.2) where λ̃t = 1 t ∫ t 0 λt (τ)dτ – average failure rate of the object; λ̃s = 1 t ∫ t 0 λs(τ)dτ – average rate of selected failure modes of the object. if we consider the exponential distribution of random time to failure, then sti = (1− exp−λst) (1− exp−λt t) ≈ λt�1 λs λt . (3.3) if we assume the incompatibility of the object failure modes copyright c© 2021 assa. adv syst sci appl (2021) 118 v.s. viktorova, a.s. stepanyants sti = λs λt (1− exp−λt t) 1− exp−λt t = λs λt . (3.4) most testability indicators given in reference books [1, 3] are the ratios similar to (3.1), (3.4). the main single testability indices are: • fault detection ratio (fdr); • fault isolation ratio (fir); • bit efficiency; • bit unreliability factor (uf); • bit false alarm factor (faf). they are formed based on summing the failure rates (failure numbers) fetched from the fmea table: ∑ i∈ω1 λi/ ∑ j∈ω2 λj or ∑ i∈ω1 ni/ ∑ j∈ω2 nj . the software executes sql queries to fmea table and generates the following subsets of the set of failure modes: • ω – total set of the object failures; • ωb, ωm, ωc – disjoint subsets of the object failures detected by bit, maintenance, crew, respectively (see appendix, tfdm type); • ωn – subset of the object hidden (latent) failures; • ωb1 – subset of the object failures detected by bit and isolated to single faulty lru; • ωi , ωii , ωiii , ωiv , ωv – disjoint subsets of the object failures related to severity levels i, ii, iii, iv, v, respectively (see appendix, tseverity type); • ωbi , ωbii , ωbiii , ωbiv , ωbv – disjoint subsets of the object failures detected by bit and related to severity levels i, ii, iii, iv, v, respectively; • ωnop, ωfa – disjoint subsets of bit failures of type no operation and false alarm (see appendix, tfailuretype); • ωff– subset of functional failures; • ωd– subset of failures impacting on departure delay (see appendix, tdelay type). table 3.1 shows how sets of failure modes ω1,ω2 are formed when calculating single testability indices. table 3.1. definition of ω1, ω2 index ω1 ω2 fault detection ratio ωb ω fault isolation ratio ωb1 ωb fdr in flight ωb ∪ ωc ω bit efficiency ωb ωb ∪ ωm ∪ ωc bit unreliability factor (uf) ωnop ∪ ωfa ω bit false alarm factor (faf) ωfa ∪ ωfa ωnop ∪ ωfa fdr for severity level i(fdri) (i=i,ii,iii,iv,v) ωbi ωi construction of distribution of fault detection ratio by levels of failures criticality are based on fdri calculation. in the reliability theory, in particular, when solving maintenance problems, the index of average failures number (afn) is used. this indicator is the mathematical expectation of the number of the object failures during a given time interval (n(t)). at exponentially distributed with parameter λ random times to failures, afn is calculated by the formula [2, pp. 89–90]: n(t) = λt. (3.5) afn is not a regulated index of testability. but when it is calculated in relation to the failure modes covered by different detection methods, it becomes an obvious measure of testability. in this software, the afn is calculated over flight and different maintenance periods. copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 119 3.2. bit conformity for an integral assessment of bit quality we use a composed probabilistic indicator called the conformity of bit. bit conformity is a complex characteristic that takes into account both functional and reliability properties of bit. in this case, two incompatible types of bit failures are considered non-operation and false alarm. let us derive an expression for bit conformity under the following initial data and assumptions: • event a means the controlled object is in operable (good) state ; event a means the controlled object is in inoperable (failure) state; prob(a) = pa, prob(a) = qa are the probabilities of these events; • bit has two failure modes, the first mode is non-operation (i) and the second one is false alarm (ii); event k is a good (failure-free) bit state, event kno is occurrence of nonoperation bit failure, event kfa is occurrence of false alarm bit failure; prob(k) = pk , prob(kno) = qno, prob(kfa) = qfa are the probabilities of occurrence of corresponding events; • events k, kno, kfa are a complete group of mutually exclusive events, so pq + pno +qfa = 1; • b is the event that bit recognizes the object state as operable, b bit recognizes the object state as failed; • η is a fault detection ratio of bit. we define the event of bit conformity c as the correct recognition of the state of the object c = (a ∧b) ∨ (a ∧b). (3.6) to find the probability of a truth of this event prob{c = 1}, we use the fault tree model. one of the techniques for constructing fault trees is to apply the theorem of decomposition into inconsistent hypotheses in some logical variables, which ensures the decomposition of the logical model and simplifies the construction of individual tree branches [24, pp. 165– 172], [25, pp. 103–115]. to construct a fault tree with a top eventc, we use the decomposition with respect to the object states (events a and a ). and in the tree branches with a we apply the decomposition with respect to bit coverage, viz. insertion the inhibit gates with conditional events η and 1− η. the inhibit gate is used to indicate that the output occurs when the input events (bottom events) occur and the input condition (right side event) is satisfied. fig. 3.3 shows the resulting fault tree. the fault tree contains repeated basic events (event 1 (a), event 2 (k), event 4 (kfa)). repeated events are colored green. the logical expression corresponding to the constructed fault tree model is given below. c = a ∧ (k ∨kno) ∨ a ∧ η ∧ (k ∨kfa) ∨ a ∧ (1− η) ∧kfa = (3.7) a ∧ (k ∨kno) ∨ (a ∧kfa) ∨ (a ∧ η ∧k). let’s write down the probability function for a truth of the found logical expression. since the decomposition into inconsistent hypotheses was used when creating fault tree, we obtained logical expression in orthogonal form. therefore, one can directly substitute the corresponding probabilities instead of logical variables, and replace logical operations with arithmetic ones [26, pp. 39,278–336]. the final expression for bit conformity prob(c = 1) = pc = pa(pk +qno) +qaqfa +qaηpk . (3.8) copyright c© 2021 assa. adv syst sci appl (2021) 120 v.s. viktorova, a.s. stepanyants fig. 3.3. fault tree model of bit conformity. . 3.3. reliability measures analytical modeling of the systems reliability behavior taking into account bit fdr and redundancy is implemented on continuous-time markov chains with discrete finite state space of size n. markov reliability model (mrm) allows to model complex transient and steadystate reliability behavior and to estimate wide range of reliability and availability indices [25, pp. 147–199]. mrm is formalized by the system of kolmogorov differential equations that describes the behavior of the state probability vector p (t) as a function of time t: dp (t)/dt = p (t)λ; p (0) = p0, (3.9) where λ = ‖λij‖ – n× n infinitesimal matrix, λij, i 6= j is the rate of transition from state i to state j, λii = − ∑ i 6=j λij . the ith component pi(t) of the row vector p (t) is the probability the system is in state i at time t. the software performs automatic generation of the markov model. the mrm describes reliability behavior of the redundant assembly of line replaceable units (ral) between maintenance periods such as heavy checks. the model generation algorithm is based on the following provisions • a set of non-redundant line replaceable units are combined into assembly based on functional or design features; • both recovery and redundancy are implemented at the assembly level; • redundancy type is parallel operating k out of m, where m total quantity of identical assemblies, k quantity required for success operation (m ≥ 1, k ≤ m); copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 121 • incomplete coverage of lru failures occurs due to functional imperfections in the bit (0 ≤ η ≤ 1); • two ral recovery scenarios between checks are considered – the ral operability is restored in case of occurrence of the hidden failures of all assemblies (scenario i) (for example, engine control system); – the ral operability is not restored in case of occurrence of the hidden failures of all assemblies (scenario ii) (for example, a fire protection system). • the presence of at least one failure detected by the bit in the assembly leads to a complete assembly recovery, regardless of whether there are other (including hidden) failures in the assembly. when building the model, the states are aggregated. all detected (hidden) failures of the assembly elements are combined into one state of a detected (hidden) assembly failure. each mrm state is encoded with a m-character code d1d2. . .di. . .dm. the symbol di of the code is determined by the rule di =  1 if the assembly is operable d if a failure detected by the bit has occurred in the assembly h if a failure undetected by the bit (hidden) has occurred in the assembly hd(dh) if both detected and hidden failures have occurred in the assembly to ensure the correctness of the aggregated mrm, it is necessary to operate with the mean recovery times of the assembly mttrassembly = ∑ i λi µi∑ i λi , (3.10) where λi, µi are the failure and recovery rates of the ith lru of the assembly. the mrm state set is partitioning into two subsets of operational (good) states ωg and non operational (failure) states ωf . the formal condition of belonging of the model state si to one of these subsets is{ si ∈ ωg if total number of symbol ”1” in state code ≥ k si ∈ ωf if total number of symbol ”1” in state code < k the parameters of the procedure for generating the infinitesimal matrix of this mrm are m, k, η, recovery scenario number. values of parameters m, k, η are determined based on sql queries to the database. the markov graph corresponding to the generated matrix at m = 3, k = 1, scenario i is shown in fig. 3.4. the software calculates on the mrm the following reliability measures 1. interval measures reliability r(t) and unreliability u(t) on time interval (0, t ). 2. point measure availability a(t) in time instant t. 3. steady-state measure mean time to first failure mttff which is defined on unlimited time interval (0,∞). to calculate r(t), it is necessary to solve the system (3.9), having previously made all failure states absorbing. this requires equating to zero the elements of rows of matrix λ corresponding to failure states. r(t) = ∑ i∈ωg pi(t), u(t) = ∑ j∈ωf pj(t) = 1−r(t). (3.11) copyright c© 2021 assa. adv syst sci appl (2021) 122 v.s. viktorova, a.s. stepanyants fig. 3.4. graph representation of the mrm (m = 3, k = 1, scenario i). the calculation of a(t) is performed in a similar way, but in this case the failure states should not be forced to be absorbing. calculation of mttff is reduced to solving a system of linear algebraic equations: − p ∗(0) = tλ∗, (3.12) where λ∗ − r × r matrix derived from matrix λ by deleting rows and columns corresponding to nonoperational states; r – total number of operational states (r < n); p ∗(0) = [ p1(0) p2(0) . . . pr−1(0) pr(0) ]. provided that the system starts from a completely healthy state p ∗(0) = [ 1 0 . . . 0 0 ]. mttff = ∑r i=1 ti. the solutions of systems of differential (3.9) and algebraic (3.12) equations are implemented using numerical methods focused on the stiffness and bad conditioning [27], inherent in the reliability models of the repaired systems [28, pp. 48–59]. 3.4. operative availability model aviation systems are classified as hybrid recovery systems [29]. these systems have the cyclic mode of operation whose duty cycle consists of a target operation phases (top), repair phases (rp), inspection, maintenance and correction phases (imcp). recovery actions after failures are not possible on-line (top) and can be conducted only off-line during rp and imcp. when the system is online (top) information about failures detected by the bit is recorded in memory unit of cmcs (central maintenance computer system). cmcs collects and reports lru faults data in order to aid maintenance personnel in recovery and maintenance procedures. during rp, only faults detected by the bit can be eliminated. full recovery of system operation is carried out only at stage imcp where latent faults also eliminated. thus, due to the imperfection of the built-in test, between imcp stages there is a consistent decrease in the system availability at the beginning of each subsequent top. the main indicator that can be used to assess the reliability behavior of the cyclic mode systems, in particular aviation systems, is operative availability. the operative availability a(t, t+ t0) is index depending on two variables – the operating time t (usually counted from 0), when recovery from failure states is possible, and the operative time t0 (counted from t), when copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 123 recovery from failure states is impossible. in general, this index is calculated as a(t, t+ t0) = ∑ i∈ωg ai(t)r(t, t+ t0)/s(t)=i, (3.13) wherer(t, t+ t0)/s(t)=i – the reliability of the system on the interval (t, t+ t0), provided that the system is in state i at time t. this software builds a trend of operative availability for a given number of phases of target operation (qstep). the construction of the trend is carried out by a program that performs a step-by-step calculation of reliabilityr(ti) (ti is the duration of the ith top). the calculation is performed on the mrm shown in fig. 3.5. a step-by-step calculation of the reliability r(ti) is carried out each time with new initial conditions, which makes it possible to take into account the fact of a decrease in availability arising from incomplete coverage of failures at the previous i− 1 recovery phases (rp). the vector of initial conditions for the ith top c(i) are formed according to the following rule c(i) = [ c (i) 1 c (i) 2 . . . c (i) j . . . c (i) n ] , (3.14) where c (i) j = { p1(ti−1) + ∑ k∈ωb pk(ti−1) if j = 1 pj(ti−1) if j ∈ ω\ωb 0 if j ∈ ωb fig. 3.5. graph representation of the mrm over target operation phase (1 ≤ m ≤ 4, 1 ≤ k ≤ m). copyright c© 2021 assa. adv syst sci appl (2021) 124 v.s. viktorova, a.s. stepanyants 4. case studies of the software six cases below illustrates the software capability. the initial data when performing the calculations of testability and reliability indices are the output of the design stage analysis of modes and effects of lru failures. a constant failure rate is specified in fmea tables for each type of failures. the failure rates are calculated using the methods described in various reference books and standards [30] or/and processing statistics on failures. the calculation of interval indicators of testability and reliability are carried out for the periods of time regulated by the scheduling maintenance program referred as a b c d checks [10, pp. 75–84]. the complexity of service tasks grows from check a to check d. the actual periods of these inspections depend on the type of aircraft, but the lighter checks a and b are carried out more often than the heavier checks c and d. we use the following times between checks: tchecka = 600h, tcheckb = 3500h, tcheckc = 7000h, tcheckd = 70000h. flight duration tflight = 5h. we demonstrate software capability on sample system consisting of three lru. screenshot of the software page with the fmea report is in fig. 4.6. fig. 4.6. screenshot of the fmea report. copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 125 4.1. case 1. failure modes classification the software implements the classification of the lru failure modes by following criteria: • detection methods – bit – failure is detected by built-in-test – maint – failure is detected in ground maintenance – crew – failure is detected by crew – none – failure is not detected until check d • failure types – ff – functional failure – bit nop – no operation built-in-test failure – bit fa – false alarm built-in-test failure • impact on departure delay the percentage of each group is calculated. the percentage is calculated by summing failure rates (estimation by frequency) fig. 4.7, 4.8, 4.9 (a), or by the number of failures (quantification) fig. 4.7, 4.8, 4.9 (b). fig. 4.7. failure detection methods. a estimate by frequency, b quantification. 4.2. case 2. sti calculation the software implements calculation of the following single testability indices: • total fault detection ratio including bit, crew, maintenance (tfdr) • bit fault detection ratio (fdr) • bit fault isolation ratio (fir) • automated fault isolation capability (afic = fdr× fir) • bit efficiency calculations of tfdr, fdr, fir, afic, bit efficiency and creation of fdr distribution by failure mode severity levels are implemented in accordance with the definitions set out in section 3.1. fig. 4.10 presents the screenshot with results of the analysis. full coverage of high severity (catastrophic, critical) failures is desirable when designing bit. comparison of frequency (i) and quantity (ii) estimates can be useful in assessing the quality of testability solutions. for example, if the fdr i < fdr ii, then the design must be changed, since the most frequently failure modes are not covered by the bit. copyright c© 2021 assa. adv syst sci appl (2021) 126 v.s. viktorova, a.s. stepanyants fig. 4.8. failure mode types. a estimate by frequency, b quantification. fig. 4.9. failures impact on departure delay. a estimate by frequency, b quantification. 4.3. case 3. afn calculation the average failures number over time interval is calculated by (3.5) and the distribution of this indicator by failure severity is constructed. the calculation is performed for time intervals: flight, check a, check b, check c. distributions are constructed for the following types of failures: • failures detected by built-in test (afn bit); • failures detected during maintenance process (afn maint); • hiden (latent) failures not covered by bit, crew and maintenance (afn none). screenshots of the afn summary page and respective charts are presented in fig.4.11 and fig.4.12. 4.4. case 4. analysis of bit conformity in the framework of bit conformity analysis the following procedures are implemented: • calculation of bit conformity based on the technical specifications of bit reliability and functional properties (conformity i); copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 127 fig. 4.10. screenshot of the report of sti calculations. . • calculation of bit conformity based on the fmea data concerning rates and coverage of failures modes (conformity ii); • comparison of the values of conformity i and conformity ii and issuing recommendations for bit modernization if conformity ii < conformity i; • study of behavior of the conformity index in the space of bit reliability parameters. an approach to comparing identical indicators calculated on the basis of different type of initial data (technical specifications and fmea) is proposed in [12]. fig. 4.13 is a screenshot of the report of calculating bit conformity i and ii and their constituents. the conformity indicator is calculated in accordance with the definitions set out in section 3.2. the plots in fig. 4.14 (a) and (b) show the results of calculation of bit conformity as a function of bit unreliability factor uf (see table 3.1). uf characterizes unreliability of bit equipment relative to unreliability of object under test. the set of curves is formed by varying the parameter kfa. kfa determines the ratio of the failure modes of bit and is calculated as faf (see table 3.1). plot of fig. 4.14 (a) shows small decrease of conformity appears linear copyright c© 2021 assa. adv syst sci appl (2021) 128 v.s. viktorova, a.s. stepanyants fig. 4.11. screenshot of summary of the system average failures number. in uf on short time interval (flight). plot of fig. 4.14 (b) demonstrates limit behavior of conformity on long time interval (check a). (limt→∞c(t) = kfa). 4.5. case 5. reliability analysis. the calculation and study of reliability indicators are conducted on models described in section 3.3. the formation of the elements of matrix λ is carried out by executing sqlqueries to fmea table to select the failure modes rates and average recovery times of the lru. the curves of reliability (unreliability), availability (unavailability), mttff are plotted as a function of fdr. a family of curves is plotted for each indicator. a separate curve corresponds to one of three standard redundancy schemes duplicated (1 out of 2), tripled (1 out of 3), majority (2 out of 3). the fourth curve corresponds to a non-redundant assembly configuration (1 out of 1). the chart in fig. 4.15 (a) demonstrates a strong dependence of mttff on fault detection ratio. a sharp jump in the mttff is observed at fdr values close to 1. the chart in fig. 4.15 (b) demonstrates a weak dependence of the reliability on fdr at short time intervals. with an increase in the time interval, this dependence increases (fig. 4.15 (c). fig. 4.15 (c) also shows that a 2 out of 3 redundancy schema can only be used in conjunction with high fault detection ratio. otherwise, this option turns out to be worse than the non-redundant scheme. on long time intervals, high bit fdr and redundancy can dramatically increase the assembly reliability (fig. 4.15 (d)). fig. 4.16 (a) and (b) show availability as a function of fdr. the curves of this chart show that bit fdr is a factor affecting the efficiency of using redundancy at recovery systems. copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 129 fig. 4.12. distribution of the system average failures number by severity and detection methods. . copyright c© 2021 assa. adv syst sci appl (2021) 130 v.s. viktorova, a.s. stepanyants fig. 4.13. screenshot of the bit conformity analysis report (conformity i vs conformity ii) . fig. 4.14. bit conformity vs unreliability factor. a. t = tflight, b. t = tchecka . copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 131 fig. 4.15. reliability indices vs bit fault detection ratio a. mttff b. r(0, tflight) c. r(0, tchecka) d. r(0, tcheckb). . fig. 4.16. availability vs fault detection ratio a. a(tflight) b. a(tchecka) . . copyright c© 2021 assa. adv syst sci appl (2021) 132 v.s. viktorova, a.s. stepanyants 4.6. case 6. operative availability analysis. the software provides the ability to build an operative availability trend for an arbitrary time frame. the time frame is defined by the user by choosing the start time counted from the last maintenance cycle with 100% renewal and quantity of the tops. the trend construction is based on models of operative availability described in section 3.4. trend analysis enables the aircraft designer to define bit functionalities and maintenance periods during which latent failures are detected and eliminated. this avoids a decrease of the operative availability below the acceptable level. fig. 4.17 (a,b) shows the trends in operative availability at fdr values equal to 0.7 and 0.9. the curves clearly show that an increase in the fdr leads to increase a probability of starting the system from a failure-free state. the limiting value of this probability, equal to one, is achieved at 100% failure coverage by bit fig. 4.17 (c). fig. 4.17. trend of operative availability (qstep = 20) a). fdr=0.7 b). fdr=0.9 c). fdr=1.0. . copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 133 5. conclusion in this paper, we presented software implemented analytic techniques of testability evaluation and joint testability and reliability analysis of the aircraft systems [31, 32]. the approach to testability evaluation relies on parametric analysis of fmea data, including parameters as failure severity, fault detection method, failure type, impact on departure delay. we have presented two models for joint testability and reliability analysis. the first model is the fault-tree of bit conformity considering functional properties and unreliability of built-intest equipment. the second one is markov reliability model considering imperfection of bit fault coverage and cyclic operation mode specific for aviation systems. we have proposed a method for calculating operative availability taking into account the recovery scenario inherent in aviation. these models and methods are implemented in the web application of the oracle database. the architecture of the database of system characteristics and fmea reports was developed. extended preliminary effort of developing database architecture was devoted to creation of meta specification of testability data as xml schema. the codes of programs defining and solving the models were implemented in the pl sql language. the software pages with user interface components have smart grid layout adopted to varying screen resolutions both pc and mobile. we have demonstrated the software capability on the sample system analysis including testability evaluation, reliability analysis of redundant assemblies with imperfect fault coverage, modeling of operative availability trend on inter-maintenance interval. the software could be of practical value to designers of built-in-test systems and analysts of testability departments. references 1. gost r 56081 (2014) izdeliya aviatsionnoy tekhniki. bezopasnost’ poleta, nadezhnost’, kontroleprigodnost’, ekspluatatsionnaya i remontnaya tekhnologichnost’. poryadok normirovaniya i kontrolya pokazateley. [state standard r 56081-2014. aircraft items. flight safety, reliability, testability and maintainability indices. setting and control.] m.: standartinform, [in russian]. 2. beighelt, f., franken, p. (1988) nadezhnost’ i tekhnicheskoye obsluzhivaniye. matematicheskiy podkhod. [reliability and maintenance. mathematical approach.] m.:radio i svyaz’, [in russian]. 3. dod 3235.1-h (1982) test and evaluation of system reliability, availability and maintainability. third edition. ousdre/ddte, 3-3. 4. amari, s.v., myers, a.f., rauzy, a., trivedi, k.s. (2008) imperfect coverage models: status and trends. in: misra k.b. (eds) handbook of performability engineering. springer, london. https://doi.org/10.1007/978-1-84800-131-2 22 5. li, x., wan, h., gong, z., wang, z., huang, h. (2011) flight control system reliability study based on hidden markov model imperfect fault coverage model-hidden markov model. int.l conf. on quality, reliability, risk, maintenance, and safety engineering. xi’an, 126–131, doi: 10.1109/icqr2mse.2011.5976582. 6. xiong, x., zhang, p. (2012) reliability analysis of flight control system for large civil aircraft with imperfect fault coverage model. proc. ieee prognostics and system health management conf. (phm-2012 beijing). beijing, 1–5, doi: 10.1109/phm.2012.6228935. 7. lubkov, n.v., spiridonov, i.b., stepanyants, a.s. (2016) vliyaniye kharakteristik kontrolya na pokazateli nadezhnosti sistem. [influence of control characteristics on system reliability indicators.] // trudy mai, 85, [in russian]. [online]. available: http://www.mai.ru/science/trudy/published.php?id=67501. 8. rani, p., pahuja, g. (2018) reliability analysis of flight control system under perfect and imperfect fault coverage. 3rd ieee int. conf. on recent trends in electronics, copyright c© 2021 assa. adv syst sci appl (2021) 134 v.s. viktorova, a.s. stepanyants information & communication technology (rteict). 759-763. 9. a4a msg-3. (2018) operator/manufacturer scheduled maintenance development. revision 2018.1. 10. chekryzhev, n.v. (2015) osnovy tekhnicheskogo obsluzhivaniya vozdushnykh sudov: ucheb. posobiye. [aircraft maintenance fundamentals:study. manual]. samara: izd. sgau, [in russian]. 11. mil-std-1629a (1984) procedures for performing a failure mode, effects and criticality analysis. notice 2. 12. arp4761 (1996) guidelines and methods for conducting the safety assessment process on civil airborne systems and equipment. sae international. 13. bidokhti, n., loeser, m. (2012) does your system have sufficient diagnostics coverage? proc. an. reliability&maintainability symp.. reno, nv, 1–6, doi: 10.1109/rams.2012.6175510. 14. liu, d., zeng, z., huang, c., & li, f. (2012) the testability modeling and model conversion technology based on multi-signal flow graph. proc. ieee prognostics & system health management conf. (phm-2012 beijing). beijing, 1–8, doi: 10.1109/phm.2012.6228919. 15. long, w., yue, l., yanling, q. tengfei, q. & minhao, w. (2017) a method of testability analysis and design based on fmea extension. 13th ieee int. conf. on electronic measurement & instruments (icemi), yangzhou, 361–367, doi: 10.1109/icemi.2017.8265968. 16. viktorova, v.s., stepanyants, a.s. (2010) proyektnyy analiz kontroleprigodnosti tekhnicheskikh sistem (teoriya, metody rascheta, programmnoye obespecheniye). [design analysis of testability of technical systems (theory, calculation methods, software).] sc.pub. m.: ipu ran, [in russian]. 17. aviatsionnyy spravochnik as 1.1. s1000dr-2013 (2013). mezhdunarodnaya spetsifikatsiya na tekhnicheskiye publikatsii, vypolnyayemyye na osnove obshchey bazy dannykh [international specification for technical publications based on a common database]. m.: fgup ”niisu” [in russian]. 18. international procedure specification for logistic support analysis lsa (2014). s3000l-b6865-03000-00. issue no 1.1. asd, aia. 19. international specification for technical publications using a common source database (2019). s1000d-b6865-01000-00. issue no. 5.0. asd, aia. 20. w3school xml tutorial. xsd introduction. [online]. available: https://www.w3schools.com/xml/schema intro.asp 21. viktorova, v.s., spiridonov, i.b. (2015) universal’naya model’ dannykh kontroleprigodnosti [generic testability data model]. m.: ipu ran, [in russian]. 22. s-techenterprises, llc. (2005). ata 100 chapter and section headings. [online]. available: http://www.s-techent.com/ata100.htm 23. molina, m., carrasco, s., martin, j. (2014). agent-based modeling and simulation for the design of the future european air traffic management system: the experience of cassiopeia. j.m. corchado et al. (eds.): paams workshops, ccis 430, 22–33 https://doi.org/10.1007/978-3-319-077673 3. 24. henley, e.j., kumamoto, h. (1996). probabilistic risk assessment and management for engineers and scientists (2nd ed.) nj.: ieeepress. 25. viktorova, v.s., stepanyants, a.s. (2016) modeli i metody rascheta nadezhnosti tekhnicheskikh sistem [models and methods of reliability analysis of technical systems]. m.: lenand, [in russian]. 26. ryabinin, i. (1976) reliability of engineering systems. principles and analysis. english translation. m.: mir publishers. 27. ustinov, s.m. , zimnitskiy, v.a. (2009) vychislitel’naya matematika [computational mathematics]. spb.: bkhv-peterburg, [in russian]. copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 135 28. viktorova, v.s., stepanyants, a.s. (2020) analiz nadezhnosti i effektivnosti mnogourovnevykh tekhnicheskikh sistem [analysis of reliability and efficiency of multilevel technical systems]. m.: lenand [in russian]. 29. stepanyants, a.s., viktorova, v.s. (2019) reliability analysis of systems with hybrid recovery and imperfect built-in-test. proc. of the 22nd int. conf. on distributed computer and communication networks: control, computation, communications (dccn-2019, moscow). cham: springer, 413–423, [in russian]. 30. petrukhin, b.p. (2014) osobennosti prognozirovaniya nadozhnosti sistem i ustroystv po metodike 217 plus tm chast’ 1. modeli intensivnosti otkazov komponent [features of predicting the reliability of systems and devices according to the 217 plus tm method. part 1. models of component failure rate]. datchiki i sistemy, 6, 37–43 [in russian]. 31. viktorova, v.s., lubkov, n.v., stepanyants, a.s. (2017) sowtware for testability analysis of aircraft functional systems: certificate of state registration of computer programs 2017660269 rf; dated 20.09.2017. 32. viktorova,v.s., stepanyants, a.s. (2019) analysis of the trend of operational availability of aircraft systems: certificate of state registration of computer programs 2019618035 rf; dated 26.06.2019. 6. appendix. testability data xml schema definition testability software this schema defines a structure and properties of testability xml data copyright c© 2021 assa. adv syst sci appl (2021) 136 v.s. viktorova, a.s. stepanyants copyright c© 2021 assa. adv syst sci appl (2021) software for testability analysis of aviation systems 137 copyright c© 2021 assa. adv syst sci appl (2021) 138 v.s. viktorova, a.s. stepanyants < xs:element name=”fmea” type=”tfmea” minoccurs=”1” maxoccurs=”unbounded” /> < xs:element name=”lru” type=”lru” minoccurs=”1” maxoccurs=”unbounded” /> < xs:element name=”system” type=”tsystem” minoccurs=”1” maxoccurs=”unbounded” /> < xs:element name=”project” type=”tproject” minoccurs=”1” maxoccurs=”unbounded” /> copyright c© 2021 assa. adv syst sci appl (2021) introduction testability analysis software description testability metadata structure testability database structure testability software features mathematical background of computational algorithms single testability indicators bit conformity reliability measures operative availability model case studies of the software case 1. failure modes classification case 2. sti calculation case 3. afn calculation case 4. analysis of bit conformity case 5. reliability analysis. case 6. operative availability analysis. conclusion appendix. testability data xml schema definition adv syst sci appl 2023; 03:177–190 published online at https://ijassa.ipu.ru. on bi-laminar neural field models of electrical activity in the primary visual cortex evgenii burlakov1,2*, ivan malkov1 1tyumen state university, tyumen, russian federation 2 derzhavin tambov state university, tambov, russian federation abstract: we investigate the modelling framework for studying electrical activity in the primary visual cortex of the brain based on a bi-laminar neural field equation. the deep layer of the neural field models the orientation-independent electrical activity, whereas the orientation-dependent superficial layer captures the selectivity to spatially oriented stimuli of the orientation columns in the primary visual cortex. we verify the solvability of a cauchy problem for the bi-laminar neural field equation with both sigmoidal and heaviside-type neuronal activation. we also construct connections between the solutions that correspond to these types of neuronal activation, which justifies the use of the heaviside-type neuronal activation functions that is crucial in the problems of computer simulations involving vast ensembles on neurons. we prove the possibility of a correct approximation of the bi-laminar neural field model with a two-layer neuronal network. we also highlight some perspectives opened by the results of the present research related to the studies of travelling waves of evoked electrical activity in the visual cortex as well as the neural activity control problems in the framework of the neurofeedback paradigm. keywords: mathematical models of primary visual cortex, bi-laminar neural field models, twolayer neuronal nerwork models, well-posedness, heaviside activation function 1. introduction mathematical models of macroand mesoscopic neuronal activity of the human brain cortex involve the description of electrical activity of vast ensembles of neuronal elements, which can be registered using electroand magnetoencefalography (eeg and meg) [1, 2] and indirectly observed in fmri recordings [3], are usually presented in the form of neural field equations (see e.g. the pioneering work [4] and the review [5]). the most well-known neural field model is the amari neural field equation (see [4]) ∂tu(t, x) = −τu(t, x) + ∫ ω ω(x, y)g(u(t, y))dy. (1.1) here u(t, x) represents the level of electrical activity in the neural field ω at time t and position x, the so-called connectivity function ω defines the strengths of interneuronal connections in the neural field, the value g(u) determines the probability of activation (firing) of a neuron with electrical activity level u. in the mathematical neuroscience community, the connectivity ω is typically assumed to be an exponentially decaying function symmetric with respect to the vertical axis or a sum of such functions, and the activation function f is taken to be a continuous sigmoidal-shaped function. ∗corresponding author: eb @bk.ru 178 e. burlakov, i. malkov several extensions of (1.1) relying on the fact that the neural media is not spatially uniform formalize the heterogeneity of the brain cortex in general forms (see e.g. [6–8]). such mathematical models are rather well-studied, including the issues of well-posedness [8, 9], construction and justification of the schemes for numerical simulations [8–10], studies of special types of solutions that are physiologically relevant [6, 7, 11, 12]. however, none of these neural field models could capture the characteristical features of primary visual cortex, whose elements of microstructure are selective with respect to perceptions of visual stimuli of certain directions, before the introduction of the following model (see [13]): ∂tud(t, x) = −τdud(t, x) + ∫ ω ωd(x, y)gd(u(t, y))dy +νd π 2∫ −π 2 gs(us(t, x, ψ))dψ, ∂tus(t, x, φ) = −τsus(t, x, φ) + ∫ ω π 2∫ −π 2 ωs(x, φ, y, ψ)gs(u(t, y, ψ))dψdy +νsgd(ud(t, x)). (1.2) here ud(t, x) defines the level of orientation-independent activity in the deep layer and us(t, x, φ) is the orientation-dependent activity in the superficial layer. the connectivities in the two layers are denoted by ωd and ωs, respectively, and the corresponding time constants are given by τd and τs. we also include vertical inputs from the deep to the superficial layer with the strength νd and back from the superficial to the deep layer with the strength νs (the latter is averaged with respect to the orientation preference of neurons in the superficial layer). functions gd, gs are probabilistic functions of firing (neuronal activation functions) for the deep layer and the superficial layer, respectively. we refer the reader to the work [13] for more details on the biophysical justification of the modeling framework (1.2). to the best of our knowledge, the work [13] is the only published study capturing the orientation selectivity in the visual cortex in the framework of the neural field setting. however, several important mathematical issues were overlooked in this study. the aim of the present paper is to fill in these mathematical blank spots in the justification of the modelling framework (1.2). namely, in section 2 we verify the solvability of a cauchy problem for (1.2) and prove the possibility of correct spatially discretized approximation of (1.2). section 3 deals with the solvability problem for (1.2) in the case when the probabilistic activation functions are replaced with the instant-activation heaviside-type functions. in section 4 we establish connections between the modelling approaches of the two previous sections. in section 5 we highlight some close perspectives opened by the results of the present research. 2. bi-laminar neural field model with continuous activation functions for convenience of the presentation of the forthcoming material, we introduce the following notations. for any metric space λ, any λ0 ∈ λ, s ⊂ λ, and r > 0, we define bλ(λ0, r) to be the ball in the space λ of the radius r > 0 centered at λ0, s to be the closure of s in λ. we denote by rm the m-dimensional real vector space with the norm | · | and by ω – a compact subset of r2. denote by lp(ω× (−π 2 , π 2 ],r2) and cp(ω× (−π 2 , π 2 ],r2) the banach spaces of integrable and continuous, respectively, functions from ω× (−π 2 , π 2 to r2), which are πperiodic in the second variable. for all t > 0, we denote ξt = [0, t ]× ω× (−π 2 , π 2 ] and copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 179 definebcp(ξt ,r2) to be the banach space of bounded continuous functions ϑ : ξt → r2 πperiodic in the third variable with the norm ||ϑ||bcp(ξt ,r2) = max x∈ξt |ϑ(x)| and l∞,p(ξt ,r2×2) – to be the banach space of 2× 2 matrix functions from ξt to r with essentially bounded components π-periodic with respect to the third variable. in this section we consider the solvability of a cauchy problem for the bi-laminar neural field equation (1.2) in the case of continuous activation functions (i.e. gd = fd, gs = fs) and justify the possibility to approximate the bi-laminar model (1.2) with a two-layered neuronal network. we start out with the system ∂tud(t, x) = −τdud(t, x) + ∫ ω ωd(x, y)fd(ud(t, y))dy +νd π 2∫ −π 2 fs(us(t, x, ψ))dψ, ∂tus(t, x, φ) = −τsus(t, x, φ) + ∫ ω π 2∫ −π 2 ωs(x, φ, y, ψ)fs(us(t, y, ψ))dψdy +νsfd(ud(t, x)) (2.3) together with the initial condition ud(0, x) = ûd(x), us(0, x, φ) = ûs(x, φ) (2.4) where ûd : ω → r, ûs : ω× (−π 2 , π 2 ] → r are continuous and lim φ→−π 2 ûs(x, φ) = ûs(x, π 2 ). assume that (aω) the neuronal connectivity functions ωd, ωs : ω× ω → rn are continuous; (af ) the neuronal activation functions fd, fs : r2 → [0, 1] are lipschitz continuous; theorem 2.1: let assumptions (aω) and (af ) be satisfied. then there exists a unique solution of the problem (2.3), (2.4), which is a continuous function from [0,∞)× ω× (−π 2 , π 2 ] to r2. proof the system (2.3) can be written as ∂tu(t, x, φ) = −τu(t, x, φ) + ∫ ω π 2∫ −π 2 v (x, y, φ, ψ)f (u(t, x, φ))dψdy, (2.5) where f (u) = ( fd(ud) fs(us) ) , v (x, y, φ, ψ) = ( ωd(x, y)/π νsδ(x− y) νdδ(x− y)/π ωs(x, φ, y, ψ) ) . we will consider the problem of solvability of the system (2.3), (2.4) in a more general setting that reads as follows: u(t, x, φ) = û(x, ψ) + t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)f (u(s, y, ψ))µ1(dψ)µ2(dy)ds (2.6) copyright © 2023 assa. adv syst sci appl (2023) 180 e. burlakov, i. malkov where u(t, x, φ) = ( ud(t, x) us(t, x, φ) ) , û(x, φ) = ( ûd(x) ûs(x, φ) ) , w (t, s, x, y, φ, ψ) = ( exp(−τd(t−s)) 0 0 exp(−τs(t−s)) ) v (x, y, φ, ψ), and µ1 and µ2 are complete σ-additive measures defined on r and r2, respectively, and finite on bounded subsets of their ranges of definition. introduce the following two conditions: (aw ) for any t > 0 and any (t, x, φ) ∈ ξt , the function w (t, ·, x, ·, φ, ·) belongs to l∞,p(ξt ,r2×2), the function (t, x, φ) 7→ ∥w (t, ·, x, ·, φ, ·)∥l∞,p(ξt ,r2×2) is bounded, and for any measurable set i ⊂ ξt and any (t0, x0, φ0) ∈ ξt , it holds true that lim (t,x,φ)→(t0,x0,φ0) ∫ ∫ ∫ i∩ ( [0,t]×ω ) w (t, s, x, y, φ, ψ)µ1(dψ)µ2(dy)ds = ∫ ∫ ∫ i∩ ( [0,t0]×ω×(−π 2 ,π 2 ] ) w (t0, s, x0, y, φ0, ψ)µ 1(dψ)µ2(dy)ds. (af ) the function f : r2 → [0, 1]2 is lipschitz continuous (hereinafter, we denote [0, 1]2 = [0, 1]× [0, 1]). choose t > 0. we define ut ∈ bcp(ξt ,r2) to be a t -local solution to (2.6) if ut satisfies the equation (2.6) on the set ξt . we consider a continuous function u∞ : [0,∞)× ω× (−π 2 , π 2 ] → r2 to be a global solution to (2.6) if for any t > 0, the restriction of u∞ to ξt is a t -local solution to (2.6). let us formulate the statement on solvability of the integral equation (2.6) in terms of the definitions given above. lemma 2.1: let assumptions (aw ) and (af ) be satisfied. then for any t > 0, the equation (2.6) has a unique t -local solution that is the restriction to ξt of the unique global solution to (2.6). proof of lemma 2.1 we choose arbitrary t > 0 and take any two functions u1, u2 ∈ bcp(ξt ,r2). we estimate ∥itntu1 − itntu2∥bcp(ξt ,r2) = max (t,x,φ)∈ξt |(itntu1)(t, x, φ)− (itntu2)(t, x, φ)|, where for any ξ ∈ bcp(ξt ,r2), the operator it is defined by the relation (it ξ)(t, x, φ) = t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)ξ(s, y, ψ)µ1(dψ)µ2(dy)ds (2.7) and, due to condition (aw ), has the image in bcp(ξt ,r2) (see e.g. [14], chapter 3, § 5.5); the operator nt defined as (nt ξ)(t, x, φ) = f (ξ(t, x, φ)) acts from the space bcp(ξt ,r2) to itself due to (af ). by the virtue of (af ), we have max (t,x,φ)∈ξt |(itntu1)(t, x, φ)− (itntu2)(t, x, φ)| copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 181 = max (t,x,φ)∈ξt ∣∣∣∣∣ t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)f (u1(s, y, ψ))µ 1(dψ)µ2(dy)ds − t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)f (u2(s, y, ψ))µ 1(dψ)µ2(dy)ds ∣∣∣∣∣ ≤ max (t,x,φ)∈ξt t∫ 0 ∫ ω π 2∫ −π 2 |w (t, s, x, y, φ, ψ)||f (u1(s, y, ψ))− f (u2(s, y, ψ))|µ1(dψ)µ2(dy)ds ≤ max (t,x,φ)∈ξt t∫ 0 ∫ ω π 2∫ −π 2 |w (t, s, x, y, φ, ψ)|lf |u1(s, y, ψ)− u2(s, y, ψ)|µ1(dψ)µ2(dy)ds ≤ lf max (t,x,φ)∈ξt t∫ 0 ∫ ω π 2∫ −π 2 |w (t, s, x, y, φ, ψ)|µ1(dψ)µ2(dy)ds ∥u1(s, y, ψ)− u2(s, y, ψ)∥ where lf > 0 is the lipschitz constant of f . by choosing t = t1 > 0 in a way that lf max (t,x,φ)∈ξt1 t1∫ 0 ∫ ω π 2∫ −π 2 |w (t, s, x, y, φ, ψ)|µ1(dψ)µ2(dy)ds < 1 and applying banach fixed point theorem (see e.g. [15], chapter 2, § 14) to the problem u = it1nt1u (2.8) in the space bcp(ξt1 ,r2) we prove the existence of a unique solution ut1 ∈ bcp(ξt1 ,r2) of (2.8), which is a unique t1-local solution of the integral equation (2.6). we introduce a new time-variable t′ = t− t1. now we consider the problem (2.6) with t = t′. we will refer to this new problem as (2.6)′ (with the initial problem u′d(−t1, x) = û′d(x), u ′ s(−t1, x, φ) = û′s(x, φ)). we apply the procedure described above and prove the existence of a t ′ 1-local solution ut ′ 1 to (2.6)′ for some t ′ 1 > 0. we thus obtain a t2-local solution ut2 to (2.6) where t2 = t1 + t ′ 1, u = { ut1 t ∈ [0, t1] ut ′ 1 t ∈ [t1, t2] . the t2-local solution ut2 is continuous at (t1, x) for any x ∈ ω, that is, ut2 ∈ on the next step, we choose any t2-local solution ut2 ∈ to the equation (2.6) we introduce a new time-variable t′′ = t − t2 and repeat the procedure we thus obtain a strictly increasing sequence {ti}, i = 1, 2, . . . and the corresponding sequence of of local solutions uti , i = 1, 2, . . . such that for any i1 < i2, uti1 is the restriction of uti2 ∈ bcp (ξti2 ,r2) to the set ξti1 . we find lim i→∞ ti = t̂ . take any t∗ ∈ (0, t̂ ). for some number i, t ∈ (ti−1, ti) and ut∗ therefore is a t∗-local solution. we have constructed the mapping t∗ 7→ ut∗ . prove that {ti}, i = 1, 2, . . . is not bounded. indeed, assuming the contrary, we get ti < t ∗ for some t ∗ <∞ and all i = 1, 2, . . ., so that the norms of the copyright © 2023 assa. adv syst sci appl (2023) 182 e. burlakov, i. malkov corresponding solutions satisfy the relation lim t→t ∗−0 ∥ut∥bcp(ξt ,r2) = ∞ which contradicts to (aw ). we thus proved that {ti}, i = 1, 2, . . . is not bounded and, hence that the sequence of local solutions constructed in the proof has a unique limit that is a global solution to (2.6). thus, lemma 2.1 is proved. note that the validity of conditions (aω) and (af ) naturally provides the fulfillment of (aw ) and (af ). now, applying lemma 2.1 to (2.6) in the case when µ1(dψ) is the lebesgue measure on r µ2(dy) is the lebesgue measure on r2, we prove the theorem. the validity of the following statement is implied by lemma 2.1. remark 2.1: for any t > 0, the restriction of the unique (global) solution obtained in theorem 2.1 to the set ξt is the unique solution to the problem (2.3), (2.4) on the set ξt . consider now a two-layer neural network ∂tv i d(t, n)=−τdvid(t, n)+ n∑ j=1 ωij d(n)fd(v j d(t, n))+νd m∑ l=1 fs(v il s (t,m)), ∂tv ik s (t, n,m)=−τsviks (t, n,m)+ n∑ j=1 m∑ l=1 ωijkl s (n,m)fs(v jl s (t, n,m))+νsfd(v i d(t, n)) (2.9) having n ”spatial” elements in each of the layers and m ”directions” in the orientation columns layer, and parameterized by the dimensions n and m of its layers. for any natural n and m, the values vid(t, n), v ik s (t, n,m) correspond to the neuronal activity of the deep layer and the superficial layer in the network, the constants ωij d(n), ω ijkl s (n,m) define the strengths of connections inside each of the layers, and the constants νd, νs define the strengths of connections between the layers. the following statement establishes correspondence between the bi-laminar neural field model (2.3) and the two-layer neuronal network (2.9). proposition 2.1: let assumptions (aω), (af ) be satisfied. for each naturalm and n, let {δ1k (m), k = 1, ...,m} and {δ2i (n), i = 1, ..., n} be finite families of open subsets of r and r2, respectively, satisfying the conditions m⋃ k=1 δ1k (m) = (−π 2 , π 2 ], n⋃ i=1 δ2i (n) = ω, lim m→∞ max k=1,...,m mes(δ1k (m)) = 0, lim n→∞ max i=1,...,n mes(δ2i (n)) = 0, where mes(·) denotes the lebesgue measure. let ψk(m) and yi(n) (k = 1, ...,m, i = 1, ..., n) be arbitrary points in δ1k (m) and δ2i (n), respectively. then for each natural n and m and any α(n,m) ∈ rnm+n, there exists a unique continuous solution to the system (2.9) considered together with the initial conditions( vid(n) ) (0) = αi d(n), ( viks (n,m) ) (0) = αik s (n,m), (2.10) which is a function v(n,m) = (vd(n), vs(n,m)) from [0,∞) to rnm+n. copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 183 moreover, the sequence of solutions vik(n,m) = (vid(n), v ik s (n,m)) to the initial value problem (2.9), (2.9), where ωij d(n) = mes(δ2i (n))ωd(yi(n), yj(n)), ωijkl s (n,m) = mes(δ2i (n))mes(δ1k (m))ωs(yi(n), yj(n), ψk(m), ψl(m)), αik d(n) = ûd(0, yi(n)), α ik s (n,m) = ûs(0, yi(n), ψk(m)) (2.11) converges to the solution u(t, x, φ) (t ≥ 0, x ∈ ω, φ ∈ (−π 2 , π 2 ]) of the initial value problem (2.3), (2.4), as n,m→ ∞, in the following sense: lim n→∞ m→∞ max t∈[0,t ] ( sup 1≤i≤n 1≤k≤m ( sup x∈δ2i (n) φ∈δ1k (m) |u(t, x, φ)− ( vik(n,m) ) (t)| )) = 0 (2.12) for any t > 0. proof by a reasoning similar to the one applied in the proof of theorem 2.1, we can conclude that for any natural n and m, the problem (2.9), (2.10) is equivalent to the equation (2.6) with µ1 = µ1 1/m, µ2 = µ2 1/n, where µ1 1/m and µ2 1/n are the sums of m dirac point measures concentrated at the points ψk(m) ∈ (−π 2 , π 2 ] and n dirac point measures at the points yi(n) ∈ ω, respectively. we define µ1 0 and µ2 0 to be the lebesgue measures on r and r2, respectively. therefore, solvability of (2.9), (2.10) for each natural n and m follows from theorem 2.1. the proposition conditions imply that the measures µ1 (·) and µ2 (·) are weakly right-continuous at 0 on the sets (−π 2 , π 2 ] and ω, respectively. indeed, for any continuous function υ(x, φ) : (−π 2 , π 2 ]× ω → r2, ∫ ω π 2∫ −π 2 υ(y, ψ)µ1 1/m(dψ)µ 2 1/n(dy) = n∑ i=1 m∑ k=1 υ(yi(n), ψk(m))mes(δ1k (m))mes(δ2i (n)) = ∫ ω π 2∫ −π 2 υ(y, ψ)µ1 0(dψ)µ 2 0(dy) (2.13) as n,m→ ∞. finally, assumptions (aω), (af ), together with the relations (2.11) and the property (2.13) provide the convergence (2.12). copyright © 2023 assa. adv syst sci appl (2023) 184 e. burlakov, i. malkov 3. bi-laminar neural field model with heaviside-type activation functions in this section we derive conditions for solvability of (1.2) in the case of discontinuous heaviside-type activation functions (i.e., gd = hd, gs = hs): ∂tud(t, x) = −τdud(t, x) + ∫ ω ωd(x, y)hd(ud(t, y))dy +νd π 2∫ −π 2 hs(us(t, x, ψ))dψ, ∂tus(t, x, φ) = −τsus(t, x, φ) + ∫ ω π 2∫ −π 2 ωs(x, φ, y, ψ)hs(us(t, y, ψ))dψdy +νshd(ud(t, x)) (3.14) where (ah) the components of the neuronal activation function h = (hd, hs), h : r2 → {0, 1} are heaviside-type activation functions: hk(u) = { 0, u ≤ hk, 1, u > hk, where hk is the threshold of activation, k = d, s. following the procedure described in the proof of theorem 2.1 we rewrite the problem (3.14), (2.4) as follows: u(t, x, φ) = û(x, ψ) + t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)h(u(s, y, ψ))dψdyds (3.15) where h = ( hd hs ) . using the ideas of a.f. filippov (see [16], chapter 2, § 4), we address the problem of solvability of the integral equation (3.15) and, hence, the problem (3.14), (2.4), in the sense of the so-called generalized solutions. we define a generalized solution to the problem (3.14), (2.4) to be a continuous function from [0,∞)× ω× (−π 2 , π 2 ] to r2, which satisfies the inclusion u(t, x, φ) ∈ û(x, ψ) + t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)h(u(s, y, ψ))dψdyds (3.16) where (ah) the set-valued function h : r2 ⇒ [0, 1]2 is defined as h = (hd,hs), hk(u) = { 0, u < hk, [0, 1], u = hk, 1, u > hk, k = d, s (hk are the same as in (ah)). copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 185 theorem 3.1: let assumptions (aω) and (ah) be satisfied. then there exists a generalized solution to the problem (3.14), (2.4), which is a continuous function from [0,∞)× ω× (−π 2 , π 2 ] to r2, proof we start out with proving the solvability of the inclusion (3.16) in the following sense. for any t > 0, we define ut ∈ bcp(ξt ,r2) to be a t -local solution to (3.16) if ut satisfies the inclusion (3.16) on the set ξt . we say that a continuous function u∞ : [0,∞)× ω× (−π 2 , π 2 ] → r2 is a global solution to (3.16) if for any t > 0, the restriction of u∞ to ξt is a t -local solution to the inclusion (3.16). lemma 3.1: let assumptions (aw ) and (ah) be satisfied. then, for any t > 0, the inclusion (3.16) has a t -local solution. any t -local solution can be extended to a global solution to (3.16). proof of lemma 3.1 we choose arbitrary t > 0 and represent the mapping on the right-hand side of (3.16) as follows: (itntu)(t, x, φ) = t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)h(u(s, y, ψ))dψdyds where it : bcp(ξt ,r2) → bcp(ξt ,r2) is defined by (2.7) and for any ξ ∈ bcp(ξt ,r2), (nt ξ)(t, x, φ) = h(ξ(t, x, φ)). by the virtue of (aw ) we have that it is a continuous operator. the operator h is upper-semicontinuous due to (ah). as h is upper-semicontinuous and it is a linear continuous operator, for mξ r = bc(ω×(−π 2 ,π 2 ],r2)(ξ, r), the composition itnt ξ : mξ r → c(ω× (−π 2 , π 2 ],r2) is a convex valued mapping. show that the composition itnt ξ is uppersemicontinuous and closed. choose ui and ϑ such that ϑ ∈ inui and ∥ui − u0∥bcp(ξt ,r2) → 0, ∥ϑi − ϑ0∥bcp(ξt ,r2) → 0 where u0 ∈ bcp(ξt ,r2), ϑ0 ∈ bcp(ξt ,r2) are some limit points and vi → v0 ∈ bcp(ξt ,r2). choose also wi ∈ n vi such that ϑi = iwi. consider the sequence {wi} ⊂ l as the sequence of bochner integrable mappings wi : [0, t ] → l(ω× (−π 2 , π 2 ]). by the virtue of condition (ah) and convergence vi → v0 we can apply kolmogorov-riesz compactness theorem [17] and consequently proposition 4.2.1 [18], and conclude that wi → w0 weakly, w0 ∈ l. further, using mazur’s lemma (see [19], section 5.1, theorem 2) we have ŵi = ∞∑ j=i βijwj such that ∥ŵi − w0∥ → 0 (3.17) where the coefficients βij satisfy the following conditions: • ∞∑ j=i βij = 1 for all i = 1, 2, ...; • one can find a number j0 such that βij = 0 for all j > j0 and for each i = 1, 2, .... the relation (3.17) implies (see section 41, theorem 4 [15]) the existence of a subsequence of the sequence ŵi converging to w0 almost everywhere. copyright © 2023 assa. adv syst sci appl (2023) 186 e. burlakov, i. malkov due to upper semicontinuity of h, for almost all (t, x, φ) ∈ ξt and for any ε > 0 there exists a number i0 = i0(t, x, φ, ε) such that for all i > i0 it holds true that h(t, x, φ, vi) ⊂ b((h(t, x, φ0))(t, x, φ), ε) we thus have wi ∈ b((h(t, x, φ0))(t, x, φ), ε). as an ε-neighborhood of a convex set is convex, we obtain ŵi ∈ b((h(t, x, φ0))(t, x, φ), ε). by the virtue of the closedness of h, the latter relation implies that w0 ∈ b((h(t, x, φ0))(t, x, φ), ε) so that w0 ∈ nu0. putting ϑ̂i = it ŵi = it ∞∑ j=i βijwj = ∞∑ j=i βijitwj = ∞∑ j=i βijϑj we obtain ∥ϑ̂i − ϑ∥ → 0, which due to the continuity of it implies that ϑ = itw0 ∈ itntu0. thus, the closedness and, consequently, the upper semicontinuity of the composition itnt is proved. now we choose some sufficiently large t > 0 and put r = 2max(û). using (aw ) and (ah), find the maximal t1 ∈ (0, t ] such that max (t,x,φ)∈ξt1 ∣∣∣∣∣ t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ)ξr(s, x, ψ)− û(x, φ)dψdyds ∣∣∣∣∣ ≤ r. thus, we have it1nt1(mξ r) ⊂ mξ r, we apply bohnenblust–karlin theorem and prove the existence of a fixed point ut1 ∈ ξt1 that is a t1-local solution to the problem (3.16). choose any t1-local solution ut1 to problem (3.16). we introduce a new time-variable t1 = t− t1. now, we consider the problem (3.16) with t = t1. we apply the procedure described above and prove the existence of a t ′ 1-local solution u′t ′ 1 to (3.16) for some t ′ 1 > 0. we thus obtain a t2-local solution ut2 to (3.16), where t2 = t1 + t ′ 1. u ={ ut1 t ∈ [0, t1] u′t ′ 1 t ∈ [t1, t2] . the t2-local solution ut2 is continuous at (t2, x, φ). in the next step, we select any t2-local solution ut2 to the problem (3.15), (2.4). we introduce a new time-variable t2 = t− t2 and repeat the procedure above. we thus obtain a strictly increasing sequence {ti}, i = 1, 2, . . . and the corresponding sequence of of local solutions uti , i = 1, 2, . . . such that for any i1 < i2, uti1 is the restriction of uti2 ∈ bcp (ξti2 ,r2) to the set ξti1 . we find lim i→∞ ti = t̂ . take any t∗ ∈ (0, t̂ ). for some number i, t ∈ (ti−1, ti) and ut∗ therefore is a t∗-local solution. we have constructed the mapping t∗ 7→ ut∗ . prove that {ti}, i = 1, 2, . . . is not bounded. indeed, assuming the contrary, we get ti < t ∗ for some t ∗ <∞ and all i = 1, 2, . . ., so that the norms of the corresponding solutions satisfy the relation lim t→t ∗−0 ∥ut∥bcp(ξt ,r2) = ∞ which contradicts to (aw ). we thus proved that {ti}, i = 1, 2, . . . is not bounded and, hence that the sequence of local solutions constructed in the proof has a unique limit that is a global solution to (2.6). thus, lemma 3.1 is proved. lemma 3.1 implies that the problem (3.14), (2.4) possesses a generalized solution, which proves the theorem. copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 187 4. continuous dependence of solutions to bi-laminar neural field model under the transition from continuous to heaviside-type neuronal activation we introduce here the following parameterized version of (2.3) ∂tud(t, x) = −τdud(t, x) + ∫ ω ωd(x, y)f i d(ud(t, y))dy +νd π 2∫ −π 2 f i s(us(t, x, ψ))dψ, ∂tus(t, x, φ) = −τsus(t, x, φ) + ∫ ω π 2∫ −π 2 ωs(x, φ, y, ψ)f i s(us(t, y, ψ))dψdy +νsf i d(ud(t, x)) (4.18) with a natural parameter i. the next statement presents the main result of this section. theorem 4.1: let assumption (aω) be fulfilled and for all natural i, the functions f i d, f i s : r → [0, 1] satisfy assumption (af ). then for any natural i, there exists a unique solution, say ui, to the equations (4.18) with the initial condition (2.4), which is a continuous function from [0,∞)× ω× (−π 2 , π 2 ] to r2. moreover, if for any ε > 0, one can find a number iε such that |f i k(u)−hk(u)| < ε, u ∈ r \br(hk, ε), i > iε, k = d, s, (4.19) a continuous function u0 : [0,∞)× ω× (−π 2 , π 2 ] → r2, u0 = (u0 d, u 0 s ), obtained from the relation lim t→∞ lim i→∞ max (t,x,φ)∈ξt |ui(t, x, φ)− u0(t, x, φ)| = 0 in the case if it satisfies the condition mes ( {(t, x) ∈ [0,∞)× ω, u0 d(t, x) = hd, u 0 s (t, x, φ) = hs} ∪{(t, x, φ) ∈ [0,∞)× ω× (−π 2 , π 2 ], u0 d(t, x) = hd, u 0 s (t, x, φ) = hs} ) = 0, (4.20) is a generalized solution to the problem (3.14), (2.4). proof similarly to the previous sections we can rewrite (4.18) as u(t, x, φ) = û(x, ψ) + t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ) ( f(u, 1 i ) ) (s, y, ψ)dψdyds, (4.21) where f(u, 1 i ) = ( f i d(ud) f i s(us) ) . copyright © 2023 assa. adv syst sci appl (2023) 188 e. burlakov, i. malkov unique solvability of the problem (4.18), (2.4) for all natural i follows from theorem 2.1. choose arbitrary t > 0. assumptions (aω) and (af ) imply that the set ∞⋃ i=1 t∫ 0 ∫ ω π 2∫ −π 2 w (t, s, x, y, φ, ψ) ( ft (bcp(ξt ,r2), 1 i ) ) (s, y, ψ)dψdyds and, hence, the set of solutions to the problem (4.18), (2.4) defined on the set ξt , say ui t , (for all natural i) is relatively compact in bcp(ξt ,r2). choose any limit point, say u0 t ∈ bcp(ξt ,r2), of the set ∞⋃ i=1 ui t . denoting ft (u, 0) = ( hd(ud) hs(us) ) we notice that the relations (4.19), (4.20) and the properties of set-valued integral (see e.g. [20], §1.5.1) provide that for any ε > 0, there exists a number iε such that itft (u i, 1 i ) ∈ bbcp(ξt ,r2)(itf(u0, 0), ε) for all i > iε. the latter allows to get the following relation: bbcp(ξt ,r2)(u 0, ε/2) ∋ ui = itf(ui, 1 i ) ∈ bbcp(ξt ,r2)(itf(u0, 0), ε/2), which means that u0 is a limit point of itft (u 0, 0) = itntu 0 (see section 3). noticing that the values of itft are closed (see e.g. the proof of lemma 3.1), we prove that u0 ∈ bcp(ξt ,r2) is a generalized solution of the problem (3.14), (2.4) on the set ξt . taking into account the arbitrary choice of t > 0, we finalize the proof. 5. conclusions and outlook in this research we established fundamental properties of the mathematical model for the macroand meso-scale electrical activity in the primary visual cortex based on the neural field concept. this opens several perspectives of further applications of the modelling framework (1.2). in recent studies of the human brain travelling waves, which are the most well-identified phenomenon of the brain electrical activity, underlying the brain functioning in both normal regime and in various pathological states (such as e.g. epilepsy and parkinson’s disease), the neural field equation (1.1) was successfully used in the simulations of travelling waves in the sensorimotor cortex [21]. the model (1.1) was supplemented by an extra component formalizing a slow negative feedback in the neural media due to the presence of the inhibitory neurons (see e.g. [22]). a much more intriguing problem from the point of view of neurophysiology is to investigate and model the interaction of travelling waves in the visual cortex with the orientation columns under a presentation of spatially oriented stimuli. in a similar way to (1.1), the models (2.3) and (3.14) can be equipped with the corresponding negative feedback components for mathematical modelling in the framework of the aforementioned studies. moreover, the presentation of spatially oriented stimuli can be treated as an impulse control problem for the modelling system (1.2), which can be studied e.g. based on the ideas developed in [23]. copyright © 2023 assa. adv syst sci appl (2023) on bi-laminar neural field models 189 the models of the form (1.2) can be used in the framework of the neurofeedback paradigm, where the target characteristics of the brain activity are converted into humaninterpretable visual, auditory or tactile stimuli [24, 25]. for example, a person undergoing a neurofeedback session learns to regulate the activity of his own central nervous system by performing task of maintaining the stimulus in a certain state trying to keep the angle of a rotating arrow displayed on the monitor screen within certain limits, which corresponds to maintaining the target characteristic of brain activity in the required range. this paradigm is used both for the correction of the psycho-emotional state, and for training to improve the efficiency of cognitive functions, as well as the treatment of a wide range of neurodegenerative diseases, including epilepsy [26]. another paradigm involves stimulation and alteration (transcranial magnetic, using direct or alternating current) of brain activity, depending on the current state of the central nervous system [27, 28]. it is aimed at suppressing pathological activity or inducing a specific behavioral response, which is used e.g. to suppress tremor in patients with parkinsonism or reduce the probability of an epileptic seizure. the paradigms described can be related to minimization problems constructed based on the mathematical framework (3.14), as it is more suitable for computer simulations compared to the model (2.3). in these minimization problems, the discontinuity in the nonlinear activation functions implies difficulties in using the standard theory that relies on the smoothness of the mappings involved. however, we conjecture that due to the presense of the heaviside-type activation functions in (3.14), one can construct isotone operators acting in an appropriate ordered space and apply the results on minimization of functionals in partially ordered spaces developed in [29] to prove the solvability of the minimization problem. acknowledgements the present research is supported by the russian science foundation (grant no. 22-2100756). references 1. muller, l., chavane, f., reynolds, j. & sejnowski, t. (2018). cortical travelling waves: mechanisms and computational principles, nat. rev. neurosci., 5, 255–268. 2. verkhlyutov, v., sharaev, m., balaev, v., osadtchi, a., ushakov, v. et. al. (2018). towards localization of radial traveling waves in the evoked and spontaneous meg: a solution based on the intra-cortical propagation hypothesis, proc. comput. sci., 145, 617–622. 3. aquino, k.m., schira, m.m., robinson, p.a., drysdale, p.m. & breakspear, m. (2012). hemodynamic traveling waves in human visual cortex, plos comput biol., 8(3), e1002435. 4. amari, s. (1977). dynamics of pattern formation in lateral-inhibition type neural fields, biol. cybern., 27(2), 77–87. 5. bressloff, p.c. (2012). spatiotemporal dynamics of continuum neural fields, j. phys. a, 45, 033001. 6. coombes, s., laing, c.r., schmidt, h., svanstedt, n. & wyller, j.a. (2012). waves in random neural media, discrete and continuous dynamical systems series a, 32, 2951– 2970. 7. malyutina, e., wyller, j. & ponosov a. (2014). two bump solutions of a homogenized wilson–cowan model with periodic microstructure, physica d, 271(1), 19–31. 8. burlakov, e., zhukovskiy, e., ponosov, a. & wyller, j. (2015). existence, uniqueness and continuous dependence on parameters of solutions to neural field equations, copyright © 2023 assa. adv syst sci appl (2023) 190 e. burlakov, i. malkov memoirs on differential equations and mathematical physics, 65, 35–55. 9. burlakov, e. (2021). on inclusions arising in neural field modeling, differential equations and dynamical systems, 29, 765–787. 10. malyutina, e. ponossov, a. & wyller, j. (2015). numerical analysis of bump solutions for neural field equations with periodic microstructure, applied mathematics and computation, 260, 370–384. 11. kolodina, k., oleynik, a. & wyller, j. (2018). single bumps in a 2-population homogenized neuronal network model, physica d : non-linear phenomena, 370, 40– 53. 12. atmania, r., burlakov, e.o. & malkov i.n. (2022). on existence and stability of ring solutions to amari neural field equation with periodic microstructure and heaviside activation function, russian universities reports. mathematics, 27(140), 318–327. 13. bressloff, p.c. & carroll, s.r. (2015). laminar neural field model of laterally propagating waves of orientation selectivity, plos comput biol., 11(10), e1004545. 14. krein, s.g. (1962). funktsionalyi analiz [functional analysis]. moscow, russia: nauka, [in russian]. 15. kolmogorov, a.n. & fomin, s.v. (1957). elements of the theory of functions and functional analysis. dover publications inc. 16. filippov, a.f. (1988). differential equations with discontinuous right-hand sides. dordrecht, springer. 17. riesz, m. (1933). sur les ensembles compacts de fonctions sommables [on compact sets of summable functions]. acta sci math (szeged), 6(1), 136–142, [in french] 18. kamenskii, m.i., obukhovskii, v.v. & zecca, p. (2011). condensing multivalued maps and semilinear differential inclusions in banach spaces. berlin, de gruyter. 19. yosida, k. (1980). functional analysis. berlin, springer–verlag. 20. borisovich, yu.g, gelman, b.d., myshkis, a.d. & obukhovskii, v.v. (2011). vvedenie v teoriyu mnogozhachnykh otobrazheniy [introduction to the theory of multivalued maps]. moscow, russia: librokom, [in russian]. 21. burlakov, e., verkhlyutov, v. & ushakov, v. (2022). a simple human brain model reproducing evoked meg based on neural field theory. in b. kryzhanovsky, w. dunin-barkowski, v. redko, y. tiumentsev, v.v. klimov (eds.), advances in neural computation, machine learning, and cognitive research v. studies in computational intelligence, 1008, 109–116. 22. pinto, d.j. & ermentrout, g.b. (2001). spatially structured activity in synapticallyncoupled neuronal networks: i. traveling fronts and pulses, siam j. appl. math., 62(1), 206–225. 23. burlakov, e.o. & zhukovskii, e.s. (2016). on well-posedness of generalized neural field equations with impulsive control, russian mathematics (iz vuz), 60(5), 66–69. 24. evans, n., gale, s., schurger, a. & blanke, o. (2015). visual feedback dominates the sense of agency for brain-machine actions, plos one, 10(6), e0130019. 25. hady, a. el. (2016). closed loop neuroscience. elsevier. 26. sitaram, r., ros, t., stoeckel, l., haller, s., scharnowski, f. et. al. (2016). closed-loop brain training: the science of neurofeedback, nature reviews neuroscience, 18, 86–100. 27. mcintyre, c.c. & warren, m.g. (1999). excitation of central nervous system neurons by nonuniform electric fields, biophys. j., 76(2), 878–888. 28. zrenner, c., belardinelli,p. , mueller-dahlhaus, f. & ziemann, u. (2016). closedloop neuroscience and non-invasive brain stimulation: a tale of two loops, front cell neurosci., 00092. 29. arutyunov, a.v., zhukovskiy, e.s. & zhukovskiy, s.e. (2019). caristi-like condition and the existence of minima of mappings in partially ordered spaces, j. optimis. theory app., 180(1), 48–61. copyright © 2023 assa. adv syst sci appl (2023) introduction bi-laminar neural field model with continuous activation functions bi-laminar neural field model with heaviside-type activation functions continuous dependence of solutions to bi-laminar neural field model under the transition from continuous to heaviside-type neuronal activation conclusions and outlook microsoft word 803 a novel enriched lasso based compression technique for energy efficient wireless sensor networks adv syst sci appl 2020; 01; 66-82 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/803 a novel enriched lasso based compression technique for energy efficient wireless sensor networks anitha k1, jaison b2, nalini m3, shiny irene d4 1) department of computer science and engineering, saveetha school of engineering, chennai, india email: anithak.rameshr@gmail.com 2) department of computer science and engineering, r.m.k engineering college, chennai, india email: bjn.cse@rmkec.ac.in 3) department of computer science and engineering, saveetha school of engineering, chennai, india email: nalinitptwin@gmail.com 4) department of computer science and engineering, saveetha school of engineering, chennai, india email: dshinyirene@gmail.com abstract: the capability of transmission in wireless sensor networks (wsn) is circumscribed due to the precincts in energy utilization, controlled resources of transmission devices and network components. information compression is taken into account as the best option, because the major part of energy consumed is for transmission of information. habitually lossy compression is adopted, since wsn abides some error in the reconstructed signals subjected to some acceptable tolerance. lasso based models have been ascertained their capability to effectually compress both multivariate and univariate data. traditional lasso considers ℓ1-norm regularization for learning in multi-dimensional data sets and assumes sparsity as model parameters. lasso prominence on sparsity and deal with the correlation between the data points. however, model sparsity may be constricting and not essentially the foremost applicable assumption in several problem domains. to eliminate this limitation, an enriched lasso (mlasso) is proposed for compression bearing in mind both sparsity and correlation. in specific the strategy can select data that are having strong features to reconstruct the data and are less correlated between each other. furthermore, an efficient alternating direction method of multipliers (admm) is adopted to resolve the ensuing sparse non-convex optimization problem. extensive experiments on diverse datasets provides the proof that mlasso outperforms other similar algorithms for signal compression. thus the proposed method ensures less energy consumption, decreases power loss and improves the operational life and reliability of network components. keywords: multidimensional data, compression, energy efficiency, senor networks, lossy compression 1. introduction in current scenario, there is an immense development in usage of mobile technologies and remote sensing devices. progresses in compact hardware and sensing devices aided to impart them into every object leading to internet of things (iot) [1]. large scale wireless sensor networks (wsns) are employed to acquire and collect the data from different envia novel enriched lasso based compression technique for energy efficient wireless sensor networks 67 copyright ©2020 assa adv. in systems science and appl. (2020) ronments and objects. they are communicated to a data acquisition center for monitoring and surveillance. the application of wsns embrace environmental monitoring [2], remote patient monitoring in healthcare [3], animal behavior classification [4], smart grid [5] and structural health monitoring for infrastructures [6]. the efficiency and the results of the applications depend on the amount of data and the acquisition speed of sensors. more data at high speed will help the application to perform efficiently. the same can be realized by higher sampling speed that ends up with large volume of raw data from sensors. thus the network needs more energy for transmitting and huge memory for storage. each battery supplying energy to the sensor nodes have limited stored energy and it is challenging to recharge or replace. as per kimura and latifi [7] each remote sensor consumes 80% of battery capacity for data transmission. the results presented by barr and asanovic [8] provides the evidence that the energy consumed by the network for transmitting a single bit of information is almost same as that required by the processing unit for performing thousands of computing operations. also more energy consumption lead to more heating of sensor nodes and thus reduce the life. thus it becomes a big challenge to manage equipment life, operate and maintain the data acquisition and storage systems in application scenarios. the energy requirements can be regulated by adopting procedures at various sections of wsn protocol stack [9-12] by energy efficient routing [13] and battery saving media access control [14]. a packet forwarding protocol is presented by l. zhang and y. zhang [15] and proved as energy efficient. the other approaches involved in improving the life of networks includes sample data reduction, battery replacement and data compression [16]. the perspective of this work given in this paper is to reduce transmission of data by introducing effective data compression. in general there are two methods of data processing that will result in reduction [17] data aggregation and data approximation. aggregation methods use statistical data such as minimum, maximum and average and effectively reduces the data size but fails to capture the important trends and changes in data required for arriving results. data approximation is a model based method and will not disturb the actual information, if the data feed contains large amount of redundancy. based on the method and purpose of application, the approximation method are categorized as time series analysis [18-20], data mining model [21], probabilistic model [22, 23] and data compression [24-26]. data compression proves to be the more favourable technique that provides good data quality, best of system performance and significantly reduces the energy consumption of wsns. embracing data compression algorithm in data storage and transmission devours some energy to compute the compression but it will be much less than the energy required for storage and transmission. in literature numerous algorithms and techniques have been offered and their capability to compress the time series data is proven and established. however many of them are derived and tested for compression of single variate sensor data such as temperature, relative humidity, etc. in-spite of vast applications of multivariate signals only few literature works are available addressing the issues in systems used to acquire, analyse and store multivariate signals. the proposed work adopts an algorithm based on lasso approximation, enriched lasso (mlasso) that can compress both multivariate and univariate sensor data. lasso has proved itself very much suitable for multidimensional regression. lasso assumes that there is no correlation in input data. but in wsns, the data is captured using different sensor units in a node. these sensors will at the same time acquire many types of data, like sound intensity, acceleration, temperature, humidity, light intensity, and video. these data used to have certain correlation. lasso happens to select only one feature among the correlated features and ends up in system underperformance [27]. this issue is addressed in the enriched lasso (mlasso) technique. in addition to discovering the correlation between the data, it also better 68 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) discriminates similar data. this property of dlasso helps in achieving high reconstruction accuracy. the proposed algorithm in particular creates optimum predictions of multi variate data and thus achieves efficient energy utilization. to demonstrate the capability and suitability of the proposed algorithms in comparison with the competing algorithms available in literature, the system is evaluated and tested with benchmark data sets from different application domains. the experimental results of proposed system performed significantly better when tested with smooth multivariate data sets. this makes the system most suitable in applications such as behaviour monitoring. the work also includes development, optimizing and computing compression consuming less running time. in addition guidance for selection of parameters influencing the algorithm performance is also described. the organization of this paper is as follows: in section 2, a brief about the related work is given. the univariate and multivariate lasso compression algorithm and the methodology of implementation is proposed in section 3. the description of real world publicly available data sets and the comparison of several compression algorithms available in literature is done in section 4. finally, section 5 briefs few concluding remarks. 2. related work in sensors that are intricate in long term real time monitoring, efficient consumption of energy is a conspicuous feature for their satisfactory performance. several available compression algorithms are not directly malleable for sensor nodes. only explicitly formulated and derived compression algorithms are appropriate and implemented for sensor nodes applications [28]. multivariate data measured from one sensor node are correlated. each original time series is not linear, however deligiannakis [17] proposed an algorithm keeping base signal as an independent variable. then the other series data is approximated using regression model. the base signal values are directly extracted from the actual sensor raw data and replaced during changes occur. the algorithm is more suitable for data having larger multivariate correlation. on the other hand the algorithm doesn’t consider error constraints. race algorithm [24] is developed for data acquired from single sensor node. few of the other techniques developed includes distributed source coding (dsc) [29], compressed sensing (cs) [30], distributed source modelling (dsm) [31] and distributed transform coding (dtc) [32]. there are other compression techniques grouped as temporal compression techniques classified as: lossless and lossy compression algorithms. they are characterized based on different principles. data accuracy [8] is preserved in lossless compression. this accuracy is preserved by eliminating redundant data during compression and decompression process. on the other hand, lossy compression techniques are derived to achieve better compression ratio at the cost of sacrificing the accuracy. a number of classic lossless compression methods have been derived and analysed. few of the algorithms suitable and used in sensors is sensor lempel-ziv-welch (s-lzw) algorithm and a distinct variant of previous technique lsz [33] made suitable for sensors. it is specially developed to storage and energy constraints in sensor networks. lossless entropy compression (lec) algorithm [33] is a lossless compression approach proved very efficient is based on traditional information encoding. the lossless compression techniques are not well suited for sensor applications. in some cases they require large memory and in other cases their assumptions about entries of static dictionary are not appropriate. lossy compression relaxes the accuracy and accepts deviations from the original to reach the flexibility to achieve less energy consumption and reconstruction accuracy for higher compression ratio, in order this will increase the lifetime of wireless sensors. schoell hammer has presented a less complex lightweight temporal compression (ltc) technique. a a novel enriched lasso based compression technique for energy efficient wireless sensor networks 69 copyright ©2020 assa adv. in systems science and appl. (2020) small value of error is added with each reading, controlled by knob [34,35]. ltc technique is simple and less complex method and can be used for temporal data compression [36]. the performance of the algorithm decays for the fluctuating sensor readings. piecewise linear approximation based scheme with less number of line segments (plamlis) algorithm exploits and estimates minimum variety of segments to approximate the given statistic. the error limit [37] is fixed during the compression process and the difference between any approximation value and its actual value is maintained less than the fixed limits. the temporal lossy compression algorithms explained above are well suitable for slowly and gradually varying signals such as temperature and not for signals such as vibrations since they use piecewise linear representation of time series. also ltc and plamlis are compression methods that are not effective for compression of multi-dimensional or multivariate time series. the work presented in this paper presents adoption of enriched lasso approximation algorithm for data compression. the algorithm is best suitable to compress both univariate and multivariate signals. 3. the compression algorithm let the one dimensional signal be 1 2( , ,....... )ny y y y observed or sensed at times 1 2 1 2( , ,....... ) ...n nt t t t where t t t � � � . that is the signal iy is measured at a time it , by the sensor. we can say that the reading of accelerometer y used to measure the acceleration of an equipment during a seismic test in the time duration 1 0 30nt s and t s . it can be seen that the values measured by y will be varying during every time instant in a given time limit t, every values of y captured by the sensor in the time interval t should be acquired for monitoring and stored for further studies, analysis, modelling and development. furthermore, it can be for a sequence of instances the change in values of y can be of less significance. in this case, if the data read in during the time between p qt and t is such that 1 ......p p qy y y� for some p q� . in this scenario, we could eliminate the data observed during the time p and q and store the reading , 1(... ....)p qy y y � in the data base corresponding to instances , 1(... ....)p qt t t � . thus data sensed in the time frame p qt and t is not varying vastly, we can be able to store the space required to store p q� recorded sensor values and also the time stamps. thus reducing the storage requirements also reduces the energy required for transmitting the data. in addition a very significant economic benefit in developing the sensor network can also be achieved. in practical installations, the values read by the sensors may be with slight variations and will not be constant in a particular time period ( , )p qt t . in such cases, the size cannot be reduced just by eliminating the data, to compress the data that are varying continuously, an efficient algorithm derived from lasso regularization [38] is presented. previously lasso approximation has been applied to order the features in a meaningful sequence in applications such as comparative prostate cancer analysis [39], genomic hybridization [40] and timevarying networks [41]. from the literature it is evident that the basic lasso-based algorithms assume restricted freedom among the variables, and their idea is to perform regression separately for each response vector rather than performing it jointly for all the response vectors. as a result they perform only data approximation and representation. compression of signal y is attained by eradicating the small differences in the signal. the signal x a smoother approximation of signal y is computed by minimizing the differences between data at adjacent time intervals. many algorithms have been proposed and much improvement has been established in the basic lasso, but the achieved response may not be the optimum. the lasso based algorithms 70 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) considers the correlations between the data points and computes an approximate signal x of signal y . the lasso algorithm performs well if the sensor data has less variance and its performance decays with increased variance between the data points in sensor data set. to improve and achieve a generalized results, it is necessary to consider the correlation between the data points in the data set and correlation between data points and response. in view of above mentioned issues in existing lasso-type approximation methods, an enhanced lasso (referred to as mlasso) is proposed for data compression. this algorithm not only discovers the correlations between the data points and the response, in addition also discriminates similar data points. this characteristics of the algorithm differentiates the proposed algorithm from the rest of lasso based algorithms. the proposed mlasso method uses a novel graph regularizer on the data points which simultaneously considers the ‘response-data point correlation’ and the ‘data point-data point correlation’ in the data. in this work alternating direction method of multipliers (admm) is employed to optimize the mlasso algorithm method and does the data compression effectively. 3.1. lossy compression for sensor signal to achieve the requirement of compressing an original signal, it shall be sparse in some possibly transformed domain. compressive sampling theory or compressed sensing developed for data transmission has become an active investigation area. the basis of this theory is that the true data can be represented using reduced samples and the same shall be recoverable from this samples with an allowed error bound. on this basis the true sensor signal e can be recuperated from approximated fewer samples of signal derived by the solution of the problem: 1 min w e e (1) subject to yie where the original signal e is a 1pu matrix to be compressed and compression theorem w is a p pu known matrix. the n pu sensing matrix φ, contains in rows, n bases for the measurements needed to be sampled from the sensed signal. the general choice to initialize φ is n randomly generated data points. finally 1nu vector y will contain the sampled data points from the original signal e such that n p� . the data acquired by the sensor in the field may contain noise. considering the measurements in practical cases the problem given in (1) is generally rewritten as: 1 min w e e (2) 2 2 subject to y ie h� d where h represents the noise in the measured signal. as mentioned in (2) the 1l norm of e under the compression transform w is minimized, such that reconstructed signal *e from n measurements of y (compressed signal) of true signal shall be consistent. the optimization problem in eq. (2) with constraints shall be rewritten in the following form: 2 2 1 min y w e ie o e� � (3) a novel enriched lasso based compression technique for energy efficient wireless sensor networks 71 copyright ©2020 assa adv. in systems science and appl. (2020) the values of h in eq. (2) and o in eq. (3) can be determined empirically. the eq. (3) can be found that it is similar to the lasso eq. (4) 2 2 1, min 1y x d e d e o e� � � (4) a main difference between eq. (3) and eq. (4) is the compression transform w available in the second term of eq. (4) called penalty term. 3.2. enriched lasso for data compression since the introduction of lasso 1l regularization has been adopted for learning in highdimensional databases. extending lasso eq. (4), the learning compressible models can be formulated as follows: 1, min ( ,1 )l y x w d e d e o e� � (5) the proposed data compression method is motivated by the objective to support the selected samples to correlate more with the reconstructed signal due to low redundancy between them. therefore, equations eq.4 and eq.6, are combined and propose the mlasso method for data compression and formulated as: 2 1 22 1 min ty s e ie o e o e e� � � (6) where 1o , 2o ≥ 0 are tuning parameters. note that t se e is a non-convex constrain. our mlasso disparities with previous lasso-type data compression methods, which use convex methods and which may be suboptimal in terms of the accuracy. here, the proposed mlasso method enforces stricter non-convex constrictions, i.e., ‘data point-response correlations’ and ‘data point-data point correlations’, in locating the optimal regression e . once the solution e of eq.7 is obtained, the selected samples can be easily recuperated. 3.2.1 learning compressible models let x is a n pu matrix with constant p , y x e h � is the sensed data with noise. the lasso refers to: 2 2 min y xe� such that 1 le d (7) and in lagrangian it becomes 2 2 1min 2 n n y x n e o e� � (8) the proposed mlasso algorithm solve the non-convex problem in eq. (7) by using alternating direction method of multipliers (admm). the concept of admm approach is to bifurcate a very complicated problem into a set of simpler and solvable problems. the best features of augmented lagrangian methods and dual decomposition are combined to solve the constrained optimization problem [42]. by introducing an auxiliary variable j into the objective function eq.7, the problem solved by admm takes the following form: 2 1 212 1min ( ) ( ) 2 t t tf g y x s e e j e o j o e e� � � � (9) 72 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) 0subject to e j� which is found as same as the problem in eq.(7) and j in the eq. (8) shall be considered as a proxy for e . the augmented lagrangian associated with the constrained problem in eq. (8) is given by 22 2 2 1 1 2 1( , , ) ( ) ( , ) 2 2 t t tl z y x s z ue j e o e e o j e j e j � � � � � � � (10) here u is a positive penalty factor and z is a lagrange multiplier corresponding to the equality constraint e j . the objective function in eq. (7) is simplified by decoupling the function by introducing an additional variable j and a constraint 0e j� . the original eq. (6) is solved by iteratively minimizing ( , , )l ze j overe andj , and then z is updated. the learning rule adopted is as given below: 1 arg min ( , , ) d i k k r l z e e e j� � (11) 22 2 2 1 1 2 1 ( ) ( , ) 2 2 t t ty x s z ue o e e o j e j e j e e w w ª º � � � � � � �« »w w ¬ ¼ =0 2 ( ) 0t k kxy xx s ze o e u e j� � � � � 1 1 2 ( ) 0 2 ( 2 ) t k k k k t k t k k xy xx s z xy z xx s xx s i xy z e o e u e j uj e o e ue e o u uj� � � � � � � � � � � � � � � 1 1arg min ( , , ) d k k k r l z e j e j� � � (12) by applying 1k ie � and the multipliers , 1,2......k iz i d are fixed in the lagrangian, the minimization problem of 1, 1,2......k i i dj � is: � � � �21 1 1 1 1 min [ . ] 2i d d d k k i i i i i i i i i z j uo j j e j� � � �¦ ¦ ¦ (13) derivative of equation with respect to ij and equating it to zero, we get: 1 1 ( ) 0 ( ) i i k k i i i i i i k k i i i i z z o j u j e j o j u j e j � � w � � � w w � � w � � � � 1 1 1 1 1 1 1 1 0 [ , ] k k k k i i i i i i i k k k k i i i i i i k k i i i z if z z if z z ue o ue o u j ue o ue o u ue o o � � � � � � ­ � � � !° ° ° � � � � �® ° ° � � � ° ¯ (14) 1 1 1( )k k k k i i i iz z u e j� � � � � (15) a novel enriched lasso based compression technique for energy efficient wireless sensor networks 73 copyright ©2020 assa adv. in systems science and appl. (2020) the learning algorithm is considered converged and stops if the primal and dual residuals meets the predefined stopping criterion mentioned as absolute tolerance and relative tolerance. the penalty parameter u affects the primal and dual residuals, and hence affects the termination of the algorithm. a large ρ tends to produce small primal residuals, but increases the dual residuals. a fixed 0u is commonly used. but there are some schemes for varying the penalty parameter which achieve better convergence. 4. experimental results in this section, we brief and report the results to illustrate the performance of lasso approximation algorithm on both multivariate and univariate datasets. we implemented both algorithms in java with eclipse as front end. all experiments were performed on a hp with an intel processor with access to 5gb of ram. 4.1. experimental data the compression capability of mlasso algorithm is verified with datasets provided by sensor scope in collaboration with hes-so, clim-arbes project and microphone provided by cmu. sensor scope: the real time data is collected using a small sensor network [26]. the data collection consists of 1472 recordings of temperature and humidity sampled at 2 seconds. it is noticed that the recorded values of humidity is higher than the recorded temperature values. the recorded values are smooth i.e. there will not be any sudden changes. the data is represented as fn_t (temperature) and fn_h (humidity). microphone (cmu): the dataset contains 2887 data sampled at an interval of 4-9 seconds done at cmu room nsh 4622. the univariate data is discontinuous or non-smooth due to presence of lot of noise and surges. the dataset is denoted as sc_m. mobile health (mhealth): among the several multivariate data sets available for public, the mobile health (mhealth) three dimensional data set from uci machine learning repository [25] is the first of its kind. the dataset comprises of 483840 data with a sampling rate of 0.02 second. the collection contains body motion and vibrant signs from ten volunteers while performing twelve physical activities using shimmer wearable sensors. the multiple sensors used measures the motion experienced by the body parts namely acceleration, the rate of turn and the magnetic field orientation. from the mhealth data set, two-lead ecg (mh_ecg), three-axis acceleration of right wrist (mh_ar), three-axis acceleration of left ankle (mh_al), and three-axis magnet data (mh_mg) are selected for experimentation. crawdad: the second multivariate data set contains the three axis acceleration readings (cj_a) of vehicles. the data set is collected by the motion of different vehicles using the mobile phone of its drivers. the collection contains 16060 data sampled at 0.0625 second intervals. in addition to the above dataset the third multivariate data set is provided by samuel madden et al. (http://db.csail.mit.edu/labdata/labdata.html). the data set contains 2.30 million data and were collected by 54 mica2dot nodes at the same time. four attribute data humidity, temperature, voltage and light intensity were collected from each node. each real 74 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) number such as sampling data, wavelet coefficient, regression coefficient, which is stored in mica2nodes, needs 2 bytes for storage. 2 bytes are enough for storing an integer such as start number start and the count number length of data.. to understand the effectiveness of the proposed algorithm, the results achieved are compared with the below mentioned well-established algorithms available in literature: � lightweight temporal compression (ltc), is a low-complexity piecewise linear approximations lossy compression algorithm. it fits the consecutive measurements as a straight line within the desired error margin. the greater compression ratio is obtained, while the larger error bound is given. � piecewise linear approximation with minimum number of line segments (plamlis), takes an n-length sequence of measurements and finds the minimum number of line segments required to represent the sequence within an error bound. � lasso : compression of signal y is achieved by removing the small differences between the data points in the signal. the signal x a smoother approximation of signal y is computed by minimizing the differences between data at adjacent time intervals 4.2. performance evaluation 4.2.1. correlation length for a discrete time series � �x n with n = 1, 2. . . , n, xp is the mean and 2 xv is the variance the correlation length of � �x n can be defined as the smallest value *n subjected to the autocorrelation function of signal is lesser than a threshold w (predefined). the autocorrelation function is given by: 2 [( ( ) )( ( ) )( ) x x x x e x m x m nn p pu v � � � (16) and * 0 arg min{ ( ) }x n n nu g ! � (17) 4.2.1. compression ratio (k ) if the time series � �x n with n = 1, 2. . . , n, needs ( )bn x bits to store and its compressed version � �x̂ n needs ˆ( )bn x bits to store then k is estimated using: ˆ( ) ( ) b b n x n x k (18) 4.2.2. total energy consumption a novel enriched lasso based compression technique for energy efficient wireless sensor networks 75 copyright ©2020 assa adv. in systems science and appl. (2020) the total energy consumed is the sum of energy needed to compress the signal and to transmit the compressed signal expressed in joule. the number of operations (addition, subtractions, divisions, multiplications and comparisons) required for compression is estimated and mapped into number of clock cycles and then to energy consumption. energy required for transmission and receiving the data depends on the energy consumed by equipment to transmit a bit of data and number of bits. 4.2.3. total error if the compressed signal � �x̂ n is reconstructed to its original form and is given by � �y n . the error is the difference between the reconstructed and original signal. the total error is calculated using: � � 2 0 ˆ( ) ( ) n i error x i x i �¦ (19) 4.4. experimental results to evaluate the proposed algorithm, the synthetic signals of correlation length *n with lengths {1, 10, 20, 50 . . . 500} time slots are used to test the system. w = 0.05 have been adopted for the results presented in this paper. additionally, a gaussian noise with standard deviation noisev = 0.04 has been included in the signal. acceptable tolerance in error between actual signal and reconstructed signal has been set to noiseh [v where 0[ t . the same signal ( )x n has been used for all the compression methods for each simulation run and value of *n . the plot in fig. 1 between provides the compression ratio achieved for different correlation lengths *n for four compression methods. form these results, it can be found that compression performance of all the four methods are poor for small values of *n and improves with the increased correlation length. this ensures that the correlation length *n is a key factor that decides the performances of the algorithm. the energy required for compression is presented in fig. 2. we can achieve better compression values for higher values of *n . the same is not true in the case of energy consumption among all scheme. fig. 1. illustration of changes in compression ratio for various correlation length for univariate data. 76 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) fig. 2. compression ratio vs energy consumption for compression methods for fixed 4 noiseh v we shall look at the dependence of energy required for compression on correlation length *n . the energy consumption depends on the operational model of a particular algorithm adopted for compression. in case of ltc, the input signal ( )x n is incrementally, initially starts the first sample and including one sample at a time. thus correlation length *n has weak influence on the consumed energy. plamlis divides and reiterates based on the correlation length and thus the method performs lesser iterations if the correlation length increases. thus the number of operations required decreases, consequently decreasing the energy consumption. for lasso and mlasso, the energy requirements increases during the addition of new sample to the model. the error tolerances have to be verified after updating the model. thus energy consumption by these models rises with increased correlation length of the input time series. in general the reconstruction error increases with the increased compressed ratio. a method can be considered as superior if it compresses the dataset and if the reconstructed dataset from the compressed data is most similar(less reconstruction error) to the original dataset. in order to explore the capabilities of the methods in compressing the signals, error of the reconstructed signals from signals of different compression ratios are presented in fig.3. it can be found that the reconstructed signal from the compressed data using mlasso is more similar to the original signal(less error). also it can be observed that the proposed method for compression provides generalized performance for data of different types. the error remains almost same for a range of compression ratio. a novel enriched lasso based compression technique for energy efficient wireless sensor networks 77 copyright ©2020 assa adv. in systems science and appl. (2020) fig. 3. the change in error at different compression ratio for different datasets this performance of the proposed lasso based method mlasso allows to adopt the method to compress signals with more variance between the signal samples. also by adopting the method it is possible to compress the signal to very smaller sizes, since the original signal can be reconstructed using small error. the plots in fig.4 shows the total energy gain for different correlation lengths. the total energy gain, defined as the ratio between the energy spent for transmission in the case with no compression and the total energy spent for compression and transmission using the selected compression techniques. the proposed method mlasso, produces the highest energy gain. in wsn scenario the total energy is highly contributed by the energy required for computations required for compression. thus, lightweight methods, ltc and plamlis, also performs better. in the case of uwn, the energy required for computations of compression is negligible compared to transmission energy requirements. in this scenario, mlasso, whose performance is the best will support more for energy savings. in this case energy savings by methods such as ltc, is limited. it can be noted that till now the compression ratio has been considered as a parameter to evaluate the system performance. and the compression ratio depends on error tolerance noiseh [v , so the compression algorithm takes ξ as an input from the user. the relative error ξ can be related with the compression ratio. here we relate mathematically ξ as a function of compression ratio. in the same way cn can also be expressed with respect to compression ratio. these relationships have been derived. 78 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) fig. 4. energy gain vs correlation length through synthetic signals, extensive simulations and numerical fitting. the relative error tolerance can be related to compression ration through the following formulae: 2 * 4 3 2 1 2 3 1 ( , ) 1 2 3 4 5 1 p p p lasso q n p p p p p mlasso q k k k [ k k k k k k ­ � � ° �° ® � � � �° ° �¯ (20) the fitting parameters p1, p2, p3, p4, p5 and q1 depends on the correlation length. the fitting formulae are validated against the carwad dataset. the temperature and humidity values in the dataset are utilized for validation. the experimental relationships of eq. (14) for mlasso and lasso are provided in fig. 5(a) and 5(b). the plot with marks provide the performance obtained to compress temperature and humidity dataset using lasso and mlasso. from these plots, it can be noted that the numerical fittings obtained with the synthetic signals also suitable for representing the signals from datasets. it can be also observed for decreasing values of *n , the shape of the curves related to ξ and compression ratio remain similar but shifts towards the right. lastly, it is noticed that the influence of *n on the performance is mainly noticeable at lesser values of *n , and for correlation lengths larger than 100, the plots be likely to converge. it was also found that the computational cost cn depends linearly on compression ratio and the correlation length ( *n ) and can be expressed as in eq.(15): * *( , ) ( )cn n cr cr nd j e � � (21) from the experiments it is ensured that cn more depends on compression ratio and less on correlation length. hence eq.(15) takes a simple form as in eq.(16) ( ) ( )cn cr crd e � (22) fig. 6 provides the results of eq. (16). the results ensure results obtained for experimental signals against the results obtained on real world signals. a novel enriched lasso based compression technique for energy efficient wireless sensor networks 79 copyright ©2020 assa adv. in systems science and appl. (2020) fig. 5. relationship between compression ratio and fitting functions *( , )n[ k for mlasso and lasso fig. 6. relationship between compression ratio and no. of cycles per bit 5. conclusion in this paper we have systematically compared the lasso and a mlasso compression algorithms for saving the energy in sensor networking. the intension was to explore the better method for energy savings are possible depending on the statistics of sensor algorithm, performance of compression and the type of hardware used to construct the network. the results in the paper reveal that there is considerable energy savings when signal compression is done by using lasso based algorithms, since they need lesser computational cost. but it can be noticed that in the case of underwater transmission of data more energy shall be required and in these case algorithms that can produce higher compression performance will be more suitable. at the end formulas for parameters that are suitable for lasso based compression methods that relates computational requirements, reconstruction error and compression ratio are 80 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) derived and tested with real world datasets. the usefulness of proposed algorithm is assessed using these test results. references 1. fan wu, jean-michel redouté, mehmet rasit yuce, (2018). we-safe: a self-powered wearable iot sensornetwork for safety applications based on lora, ieee access, 6, pp. 40846 – 40853. 2. kozlovszky m., kovacs l., karoczkai k. (2015) cardiovascular and diabetes focused remote patient monitoring. in: braidot a., hadad a. (eds) vi latin american congress on biomedical engineering claib 2014, ifmbe proceedings, 49. 3. la gonzalez, gj bishop-hurley, rn handcock, and c crossman (2015). behavioral classification of data from collars containing motion sensors in grazing cattle, computers and electronics in agriculture, 110, pp. 91–102, 4. l.chhaya, p. sharma, g. bhagwatikar, and a. kumar (2017), wireless sensor network based smart grid communications: cyber attacks, intrusion detection system and topology control, electronics, 6(1), p. 5. 5. hong-nan li, ting-hua yi, liang ren, dong-sheng li, and lin-shen huo (2014). reviews on innovations and applications in structural health monitoring for infrastructures, structural monitoring and maintenance, 1(1), 1–45. 6. joseph azar, abdallah makhoul, mahmoud barhamgi, raphael couturier (2019), future generation computer systems, 96, pp. 168-175. 7. kenneth c barr and krste asanovic (2006). energy-aware lossless data compression, acm transactions on computer systems (tocs), 24(3):250–291. 8. c. x. wang, d. yuan, h. h. chen, and w. xu (2008). an improved deterministic sos channel simulator for multiple uncorrelated ravleigh fading channels, ieee transactions on wireless communications, 7(9), pp.3307–3311. 9. x. cheng, c. x. wang, h. wang et al. (2012). cooperative mimo channel modeling and multi-link spatial correlation properties, ieee journal on selected areas in communications, 30(2), pp.388–396. 10. c.x.wang, x.hong, h.h.chen, and j.thompson (2009). on capacity of cognitive radio networks with average interference power constraints, ieee transactions on wireless communications, 8 (4), pp.1620–1625. 11. q. ni and c. zarakovitis (2012). nash bargaining game theoretic scheduling for joint channel and power allocation in cognitive radio system, ieee journal on selected areas in communications,30 (1),pp.70–81. 12. s.bai, w.y.zhang, g.l.xue, j.tang, and c.g.wang (2012). dear: delay-bounded energy-constrained adaptive routing in wireless sensor networks, in proceedings of the 31st annual ieee international conference on computer communications (infocom ’12), pp.1593–1601. 13. j.kabaraand m.calle (2012), mac protocols used by wireless sensor networks and a general method of performance evaluation, international journal of distributed sensor networks, articleid 834784,11 pages. 14. l. zhang and y. zhang (2009), energy-efficient cross-layer protocol of channel-aware geographic-informed forwarding in wireless sensor networks, ieee transactions on vehicular technology, 58(6), pp.3041–3052. 15. chugh, amit; panda, supriya (2019), energy efficient techniques in wireless sensor networks, recent patents on engineering, 13 (1),pp. 13-19. 16. abhishek tiwari, tiago h. falk, (2019) lossless electrocardiogram signal compression: a review of existing methods, biomedical signal processing and control, 51 , pp. 338-346 a novel enriched lasso based compression technique for energy efficient wireless sensor networks 81 copyright ©2020 assa adv. in systems science and appl. (2020) 17. wenyu cai, meiyan zhang (2018), spatiotemporal correlation–based adaptive sampling algorithm for clustered wireless sensor network, international journal of distributed sensor networks, 14 (8). 18. mou wu, liansheng tan, naixue xiong (2016), data prediction, compression, and recovery in clustered wireless sensor networks for environmental monitoring applications, information sciences, 329(1), pp. 800-818. 19. y. l. borgne and g. bontempi (2012), time series prediction for energy-efficient wireless sensors: applications to environmental monitoring and videogames, in proceedings of the 4th international icst conference on sensor systems and software (scube ’12), 102, pp.63–72. 20. mohamed mostafa fouada, nour e. oweisb, tarek gaberb, maamoun ahmedd, vaclav snasel (2015), data mining and fusion techniques for wsns as a source of the big data, science direct,5, pp. 778 – 786. 21. fernando perez-cruz and sanjeev r kulkarni (2010). robust and low complexity distributed kernel least squares learning in sensor networks, signal processing letters, ieee, 17(4):355–358. 22. f. kazemeyni, e. b. johnsen, o. owe, and i. balasingham (2012), mule-based wireless sensor networks: probabilistic modeling and quantitative analysis, in proceedings of the 10th international conference on integrated formal methods, 7321, pp.143–157. 23. j. m. zhang, y. p. lin, s. w. zhou, and j. c. ouyang (2010). haar wavelet data compression algorithm with error bound for wireless sensor networks, journal of software, 21(6), pp. 1364–1377. 24. x.song, c.r.wang, j.gao, and x.hu (2012). dlrdg: distributed linear regressionbased hierarchical data gathering framework in wi reless sensor network, neural computing and applications, 15 pages. 25. m.sadler and m.martonosi (2006). data compression algorithms for energyconstrained devices in delay tolerant networks, in proceedings of the 4th international conference on embedded networked sensor systems(sensys’06),pp.265–278. 26. zou, h., hastie, t (2005), regularization and variable selection via the elastic net. journal of the royal statistical society: series b (statistical methodology) 67(2), pp. 301–320. 27. davide zordan, borja martinez, ignasi vilajosana and michele rossi (2014). on the performance of lossy compression schemes for energy constrained sensor networking, acm transactions on sensor networks (tosn), 11 (1):15. 28. sharadh ramaswamy, kumar viswanatha, ankur saxena, and kenneth rose (2010), towards large scale distributed coding, in proc. of acoustics speech and signal processing (icassp), 2010 ieee international conference, pages 1326–1329, 2010. 29. donoho d l., (2006). compressed sensing, ieee transactions on information theory, 52(4): 1289-1306. 30. fernando perez-cruz and sanjeev r kulkarni (2010), robust and low complexity distributed kernel least squares learning in sensor networks, signal processing letters, ieee, 17(4):355–358. 31. alon amar, amir leshem, and michael gastpar (2010), recursive implementation of the distributed karhunen-loeve transform, signal processing, ieee transactions,58(10):5320–5330. 32. francesco marcelloni and massimo vecchio (2009). an efficient lossless compression algorithm for tiny nodes of monitoring wireless sensor networks, the computer journal, 52(8):969–987. 33. tom schoell hammer, ben greenstein, eric osterweil, michael wimbrow, and deborah estrin(2004). lightweight temporal compression of microclimate datasets, center for embedded network sensing. 82 anitha k., jaison b., nalini m., shiny irene d. copyright ©2020 assa adv. in systems science and appl. (2020) 34. jialiang lu, fabrice valois, mischa dohler, and min-you wu (2010). optimized data aggregation in wsns using adaptive arma, in proc. of sensor technologies and applications (sensorcomm), 2010 fourth international conference, pages 115–120. 35. mohammad abu alsheikh, puay kai poh, shaowei lin, hwee-pink tan, and dusit niyato (2014). efficient data compression with error bound guarantee in wireless sensor networks, in proc. of the 17th acm international conference on modeling, analysis and simulation of wireless and mobile systems, pages 307–311. 36. ngoc duy pham, trong duc le, and hyunseung choo (2008). enhance exploring temporal correlation for data collection in wsns, in proc. of research, innovation and vision for the future, 2008, rivf 2008. ieee international conference, pages 204– 208. 37. yuan, m., lin, y (2006). model selection and estimation in regression with grouped variables. journal of the royal statistical society: series b (statistical methodology) 68(1), 49–67. 38. jacob, l., obozinski, g., vert, j.p.(2009). group lasso with overlap and graph lasso. in: proceedings of the 26th annual international conference on machine learning. pp. 433–440. 39. zhao, p., rocha, g., yu, b. (2009). the composite absolute penalties family for grouped and hierarchical variable selection, the annals of statistics, pp. 3468–3497. 40. amr ahmed and eric p xing (2009), recovering time-varying networks of dependencies in social and biological studies, in proc. of the national academy of sciences, 106(29):11878–11883. 41. boyd, s.,parikh, n., chu, e.,peleato, b., eckstein, j (2011). distributed optimization and statistical learning via the alternating direction method of multipliers, foundations and trends in machine learning 3(1), 1–122. microsoft word 859-article text-2988-1-6-20200617.doc adv syst sci appl 2020; 02; 44-55 published online at https://ijassa.ipu.ru. influence assessment of intelligent unmanned ground vehicles on the transport network state andranik akopov1,2, nerses khachatryan1,2, fedor belousov1,2* 1) national research university higher school of economics, moscow, russia e-mail: aakopov@hse.ru, nkhachatryan@hse.ru, fbelousov@hse.ru 2) central economics and mathematics institute ras, moscow, russia e-mail: nerses@cemi.rssi.ru abstract: this article is devoted to econometric analysis of the results of experiments conducted with two agent-based models, which describe the movement of ground vehicles. there are two types of road users in these models: manned ground vehicles (mgv) and unmanned ground vehicles (ugv). in the first model, the main difference between ugv and mgv is an ability to exchange massages between ugv for transmitting information about extreme situations, which allows them to adjust speed and direction of movement. in the second model, in addition to the above differences, ugv have an additional advantage, namely, the ability to intelligently assess density of traffic flow for efficient maneuvering. in these models, at a given roundabout, traffic characteristics such as output stream traffic and the number of traffic accidents are analyzed. the main task of the econometric analysis is to study dependence of these traffic characteristics on the model parameters such as average vehicle speed, input flow rate, message exchange rate between ugv, and the impact of the effect obtained from the implementation into ugv ability of intelligent estimation of traffic flow density. keywords: unmanned ground vehicles, manned ground vehicles, agent-based model, experimental results, econometric analysis 1. introduction currently, the task of developing new intelligent control systems for an ensemble of ground-based ugv is being actualized in order to ensure high-speed and safe traffic, maximize the capacity of the transport system and minimize the number of factors (for example, emerging traffic accidents) that threaten other road users. works [1-7] are devoted to development of ugv control systems. in particular, the paper [1] proposes an algorithm based on predictive tracking control, designed to reduce time for decision-making. a very promising direction in the design of ugv control systems is an approach based on the use of machine learning methods, in particular, reinforced learning [2]. such methods make it possible to ensure an effective response of ugv to the occurrence of certain situations on the road (for example, the sudden appearance of a pedestrian). at the same time, even the use of multilayer neural networks with a large number of neurons and synoptic connections (deep learning) [8, 9], the use of ultraprecise neural networks [10], etc. does not allow for zero learning error in pattern recognition and classification of moving objects with complex characteristics. therefore, the task of developing ugv control systems based on simple rules that take into account the most common road situations for making effective decisions about maneuvering, changing lanes [4, 7], adjusting the speed and direction of traffic, etc. is urgent. such a task can be solved using agent-based simulation methods [6, 7, 11, 12], fuzzy clustering [7, 13, 14, 15] and genetic optimization algorithms [16]. so, earlier we developed and implemented two agent-based models of mgv and ugv motion in anylogic simulation system [6, 7]. * corresponding author: fbelousov@hse.ru influence of unmanned vehicles on the state of the transport network 45 copyright ©2020 assa. adv. in systems science and appl. (2020) the first model, using the previously proposed phenomenological approach [11], is designed to assess the influence of various parameters (such as average initial speeds, input flow intensity, data exchange rate between ugv, etc.) on the behavior and condition of unmanned and manned ground vehicles in a dense flow [6]. the spatial dynamics of mgv and ugv is defined by a system of difference equations with a variable structure. the model takes into account the effects of "turbulence" and "traffic congestion" caused mainly by high vehicle density and occurrence of road accidents. at the same time, effective interaction is carried out between the ugv, in particular, periodic exchange of messages about the traffic situation (for example, about the places where accidents occur), in order to timely adjust speed and direction of movement. an important feature of this model is the proposed concept of the agent's personal space. the behavior of agents (mgv and ugv) in flow depends on density of the surrounding space. as the flow density increases, the radius of the agent's personal space is compressed (i.e., the transport flow is compacted). however, after reaching a certain threshold value of density, the radius of the agent's personal space increases significantly in manned ground vehicles, due to the occurrence of panic, and partially increases in ugv, due to a predetermined desire to avoid collision in dense traffic flow. fig. 1.1. configuration of road space for mgv and ugv the second model is designed for effective management of ugv, in particular, by maneuvering when changing lanes, the algorithm of which is based on the proposed fuzzy clustering algorithm [7]. the movement of the mgv and ugv ensemble in a certain twodimensional space is considered. the space consists of combination of circular motion, two entrances and two exits from the circle (fig. 1.1). at the same time, ugv makes individual decisions about trajectory adjustment based on simple rules. if there is less dense traffic in one of the adjacent lanes, this ugv is tuning to the corresponding lane. in this case, the flow density is estimated using fuzzy clustering methods for each of the available alternatives (lanes) across the entire ensemble of vehicles (both ugv and manned ground vehicles that do not have all the information about the density of the surrounding space). as a result, adaptive ugv management is supported, which minimizes the risks of accidents (accidents involving ugv) and maximizes traffic (total output stream) in conditions of heavy traffic flow. software implementation in anylogic of the developed simulation model was performed and numerical experiments were carried out. modes that ensure safe and highspeed movement of vehicles in dense traffic flow are found (fig. 1.2). the developed system (fig. 1.2) allows us to evaluate influence of multiple control parameters (for example, average speeds of mgv and ugv, intensities of input flows of mgv and ugv, etc.) on the total number of accidents and traffic of the output stream. 46 a. akopov, n. khachatryan, f. belousov copyright ©2020 assa adv. in systems science and appl. (2020) fig. 1.2. software implementation of the model in anylogic simulation system this work is devoted to the analysis of the results of experiments conducted with these models. the main task of this analysis is to study dependence of the output stream traffic and the number of accidents on model parameters such as average vehicle speed, input stream intensity, message exchange frequency between ugv, and the impact of the effect obtained from the implementation into ugv ability of intelligent estimation of traffic flow density. to do this, a pool of identical experiments was conducted within each model. a set of experiments is formed as a result of multiple runs of the model, in each of which the specified model parameters take values in a given range with a given step. finally, a database consisting of the results of more than one hundred thousand experiments is formed for each model. based on these results, an econometric analysis was carried out: the dependence of the output stream traffic and the number accidents on the parameters in both the first and second models was studied. also the effect obtained as a result of the introduction of ugv with the ability to intelligently assess the density of road traffic, consisting in increasing the output stream traffic and reducing the number of accidents, was estimated. 2. source data, primary data analysis this paper uses data obtained from a pool of experiments with each of the two agent-based models [6, 7] that describe the movement of vehicles, both manned and unmanned, in a given section of circular motion. in these experiments, such traffic characteristics as the output stream traffic and the number of traffic accidents over a certain period of time, which act as dependent variables, were calculated. the following parameters of agent-based models are explanatory variables: • intensityofunсrewedvehicles – intensity of entry of ugv to the specified section of circular motion (in units of model time); • intensityofusualagents – intensity of entry of mgv to the specified section of circular motion (in units of model time); • speedofuncrewedvehicles – average ugv speed (km/h); • speedofusualagents – average mgv speed (km/h); • intensityofconnections – intensity of messaging between ugv (in units of model time). as noted in the introduction, lot of experiments are formed as a result of multiple runs of the model implemented in anylogic using the special option “sensitivity analysis”. in this case, influence of unmanned vehicles on the state of the transport network 47 copyright ©2020 assa. adv. in systems science and appl. (2020) the specified model parameters take values in the specified range with the specified step (table 2.1). table 2.1. variation of model parameters № name of parameter range of variation variation step number of variations 1 intensityofuncrewedvehicles 0.005 – 0.05 0.005 10 2 intensityofusualagents 0.005 – 0.05 0.005 10 3 speedofuncrewedvehicles 40 – 140 10 11 4 speedofusualagents 40 – 140 10 11 5 intensityofconnections 0 – 1 0.1 11 it is easy to see that the total number of experiments in each pool is 133100. the table 2.2 shows the main descriptive statistics of dependent variables. table 2.2. descriptive statistics average median standart deviation minimum maximum output stream traffic (model 1) 56.5 56 25.4 3 149 output stream traffic (model 2) 67.8 65 35.3 3 207 number of accidents (model 1) 16 10 18.6 0 148 number of accidents (model 2) 5.7 1 12.7 0 145 according to table 2.2, introduction of intelligent ugv leads to an increase in output traffic by 20% and a decrease in the average number of accidents by 2.8 times. it is also worth noting that the median number of accidents is reduced from 10 to 1, i.e. in model with the ability to intelligently estimate the traffic density, approximately half of the experiments do not have an accident. in fig. 2.1 and fig. 2.2 histograms of the output stream traffic distribution are presented, and in fig. 2.3 and fig. 2.4 – histograms of the distribution of the number of accidents in model 1 and model 2, respectively. fig. 2.1. histogram of the output stream traffic distribution in model 1 0 200 400 600 800 1000 1200 1400 1600 1800 2000 0 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 10 0 10 5 11 0 11 5 12 0 12 5 13 0 13 5 14 0 14 5 15 0 n um be r o f e xp er im en ts output stream traffic 48 a. akopov, n. khachatryan, f. belousov copyright ©2020 assa adv. in systems science and appl. (2020) fig. 2.2. histogram of the output stream traffic distribution in model 2 fig. 2.3. histogram of the number of accident distribution in model 1 fig. 2.4. histogram of the number of accident distribution in model 2 from fig. 2.3 and fig. 2.4 it follows that the number of experiments without an accident in model 2 is three times less than in model 1. we finish the initial data analysis with the graph 0 10000 20000 30000 40000 50000 60000 70000 0 6 12 18 24 30 36 42 48 54 60 66 72 78 84 90 96 10 2 10 8 11 4 12 0 12 6 13 2 13 8 14 4 n um be r o f e xp er im en ts number of accidents 0 5000 10000 15000 20000 25000 0 6 12 18 24 30 36 42 48 54 60 66 72 78 84 90 96 10 2 10 8 11 4 12 0 12 6 13 2 13 8 14 4 n um be r o f e xp er im en ts number of accidents 0 200 400 600 800 1000 1200 1400 1600 0 7 14 21 28 35 42 49 56 63 70 77 84 91 98 10 5 11 2 11 9 12 6 13 3 14 0 14 7 15 4 16 1 16 8 17 5 18 2 18 9 19 6 20 3 21 0 n um be r o f e xp er im en ts output stream traffic influence of unmanned vehicles on the state of the transport network 49 copyright ©2020 assa. adv. in systems science and appl. (2020) below (fig. 2.5). it shows how the number of experiments in which the number of accidents is not less than the specified number (between the minimum and maximum number) decreases in both the first and second models. fig. 2.5. number of experiments with number of accidents not less than a given value 3. econometric model specification selection let's proceed to the study of the influence of the model parameters specified in table 2.1 on the output stream traffic and the number of accidents, i.e., to construct and evaluate the corresponding regression equations. at the initial stage of this study, we will try to determine the model specification. since the set of explanatory variables is defined in advance in the framework of building the above agent-based models, we are trying to determine the type of assumed dependency. we restrict ourselves to two types of functions that are most often encountered in econometric studies: linear and power-law. in this regard, we present matrices of pairwise correlations of all variables (both explanatory and dependent), as well as their logarithms for both models (tables 3.1, 3.2, 3.3, 3.4). the following notation is used in these tables: • x1 – intensityofuncrewedvehicles; • x2 – intensityofusualagents; • x3 – speedofuncrewedvehicles; • x4 – speedofusualagents; • x5 – intensityofconnections; • y1 – output stream traffic; • y2 – number of accidents. thus – x1, x2, x3, x4, x5 – explaining variables, а y1, y2 – dependent. table 3.1. correlation matrix of variables in model 1 x1 x2 x3 x4 x5 y1 y2 x1 1.00 -0.20 -0.008 0.03 -0.0002 0.07 0.47 x2 -0.20 1.00 0.002 0.03 0.0005 0.15 0.46 x3 -0.008 0.002 1.00 -0.02 -0.0006 0.40 0.11 x4 0.03 0.03 -0.02 1.00 0.002 0.74 -0.21 x5 -0.0002 0.0005 -0.0006 0.002 1.00 0.0003 -0.003 y1 0.07 0.15 0.40 0.74 0.0003 1.00 -0.16 y2 0.47 0.46 0.11 -0.21 -0.003 -0.16 1.00 0 20000 40000 60000 80000 100000 120000 1 7 13 19 25 31 37 43 49 55 61 67 73 79 85 91 97 10 3 10 9 11 5 12 1 12 7 13 3 13 9 14 5 n um be r o f e xp er im en ts number of accidents model 1 model 2 50 a. akopov, n. khachatryan, f. belousov copyright ©2020 assa adv. in systems science and appl. (2020) table 3.2. correlation matrix of logarithms of variables in model 1 log(x1) log(x2) log(x3) log(x4) log(x5) log(y1) log(y2) log(x1) 1.00 -0.22 -0.008 0.02 -0.0007 0.09 0.47 log(x2) -0.22 1.00 0.001 0.03 0.0005 0.08 0.44 log(x3) -0.008 0.001 1.00 -0.015 -0.0003 0.37 0.10 log(x4) 0.02 0.03 -0.015 1.00 0.002 0.75 -0.17 log(x5) -0.0007 0.0005 -0.0003 0.002 1.00 -0.0003 -0.001 log(y1) 0.09 0.08 0.37 0.75 -0.0003 1.00 -0.10 log(y2) 0.47 0.44 0.10 -0.17 -0.001 -0.10 1.00 table 3.3. correlation matrix of variables in model 2 x1 x2 x3 x4 x5 y1 y2 x1 1.00 -0.26 -0.02 0.12 -0.001 0.23 0.19 x2 -0.26 1.00 0.04 -0.008 -0.002 0.13 0.39 x3 -0.02 0.04 1.00 -0.05 -0.001 0.37 0.12 x4 0.12 -0.008 -0.05 1.00 -0.0009 0.74 -0.41 x5 -0.001 -0.002 -0.001 -0.0009 1.00 -0.003 0.001 y1 0.23 0.13 0.37 0.74 -0.003 1.00 -0.37 y2 0.19 0.39 0.12 -0.41 0.001 -0.37 1.00 table 3.4. correlation matrix of logarithms of variables in model 2 log(x1) log(x2) log(x3) log(x4) log(x5) log(y1) log(y2) log(x1) 1.00 -0.23 -0.02 0.10 -0.002 0.20 0.14 log(x2) -0.23 1.00 0.04 -0.03 -0.0006 0.07 0.45 log(x3) -0.02 0.04 1.00 -0.05 -0.001 0.35 0.10 log(x4) 0.10 -0.03 -0.05 1.00 -0.0006 0.77 -0.47 log(x5) -0.002 -0.0006 -0.001 -0.0006 1.00 -0.001 0.001 log(y1) 0.20 0.07 0.35 0.77 -0.001 1.00 -0.40 log(y2) 0.14 0.45 0.10 -0.47 0.001 -0.40 1.00 based on the data from these tables, we can draw the following conclusions: 1. the pairwise correlation both between explanatory variables and between their logarithms is low in both models, which, when evaluating the corresponding regression equations, cannot become a source of multicollinearity. this is important given that we are interested in the contribution of each explanatory variable to changes in dependent variables. 2. comparing pairwise correlations between dependent and explanatory variables does not allow us to determine which of the functions (linear or power-law) most adequately describes the dependence of output stream traffic and the number of accidents on the model parameters. therefore, in the future we will conduct econometric analysis using both types of functions. in connection with the use of power functions, which will be logarithmized before estimation, we discard the results of experiments in which a variable assumes a zero value. thus, to construct the regression equations for the output stream traffic, the data of 102448 observations were used, and for the number of accidents – the data of 65134 observations. 4. econometric analysis of traffic characteristics in the framework of linear regression equations we proceed to evaluate the dependences of the output stream traffic and the number of accidents in both models on the indicated parameters of the model in the framework of linear regression equations. so, we will build and study linear regression equations with the following dependent variables: • output stream traffic in model 1; • output stream traffic in model 2; • number of accidents in model 1; • number of accidents in model 2. the coefficients were estimated in the framework of the classical multiple regression model using the least squares method. since the white test confidently rejects the null hypothesis of influence of unmanned vehicles on the state of the transport network 51 copyright ©2020 assa. adv. in systems science and appl. (2020) homoskedasticity of regression residues, standard errors are estimated by using the white procedure. 4.1. results of an econometric analysis of the output stream traffic in the framework of linear regression equations here are the results of evaluating the dependence of the output stream traffic (table 4.1) on the above set of explanatory variables. table 4.1. dependent variable – output stream traffic, linear regression explanatory variables model 1 model 2 constant -37.76*** (0.27) -79.16*** (0.40) intensityofuncrewedvehicles 147.22*** (3.72) 499.05*** (5.29) intensityofusualagents 284.38*** (4.56) 446.01*** (6.70) speedofuncrewedvehicles 0.34*** (0.002) 0.48*** (0.002) speedofusualagents 0.60*** (0.002) 0.88*** (0.003) intensityofconnections 0.40 (0.28) -0.12 (0.25) r2 0.73 0.76 note. in parentheses are the values of standard errors. ***, **, * – significance at the 1, 5 and 10% level, respectively. as follows from table 4.1, all explanatory variables in both models except intensityofconnections, which describe the output stream traffic, are significant at a 1% level. we also note that in both models the values of the determination coefficients are quite high. comparing the coefficients for significant variables in both models, we can draw the following conclusions: 1. an increase in the ugv input stream intensity by 0.005 units of model time (variation step, table 2.1) leads to an increase in model 2 output stream traffic compared to model 1 output stream traffic by about 1.76 units. for mgv, the indicated difference is 0.81 units. 2. an increase in the average ugv speed by 10 km / h (variation step, table 2.1) leads to an increase in the output stream traffic of model 2 compared with the output stream traffic of model 1 by 1.4 units. for mgv, the indicated difference is 2.8 units. 4.2. results of an econometric analysis of the number of accidents in the framework of linear regression equations here are the results of evaluating the dependence of the number of accidents on the above set of explanatory variables (table 4.2). table 4.2. dependent variable – number of accidents, linear regression explanatory variables model 1 model 2 constant -22.46*** (0.20) -6.34** (0.24) intensityofuncrewedvehicles 797.77*** (2.98) 414.23*** (3.96) intensityofusualagents 810.13*** (3.14) 546.11*** (4.00) speedofuncrewedvehicles 0.07*** (0.001) 0.04*** (0.001) speedofusualagents -0.15*** (0.001) -0.22*** (0.002) intensityofconnections -0.16 (0.13) 0.12 (0.16) r2 0.62 0.45 note. in parentheses are the values of standard errors. ***, **, * – significance at the 1, 5 and 10% level, respectively. 52 a. akopov, n. khachatryan, f. belousov copyright ©2020 assa adv. in systems science and appl. (2020) from table 4.2 it follows that, as in the previous case, the intensityofconnections variable is not significant. all other explanatory variables describing the number of accidents in both models are significant at a 1% level. the positive values of the coefficients in the variables intensityofuncrewedvehicles and intensityofusualagents are quite natural. the average speed of ugv has a positive effect on the number of accidents, and the average speed of mgv is negative. this is explained as follows: an increase in the speed of mgv, in contrast to an increase in the speed of ugv, is a result of enough free traffic and leads to a decrease in road congestion, and an increase in the speed of ugv can lead to unpredictable actions by mgv. comparing coefficients for explanatory variables in both models, we can draw the following conclusions: 1. an increase in ugv input stream intensity by 0.005 units of model time (variation step, table 2.1) leads to a decrease in the number of accidents in model 2 compared to the number of accidents in model 1 by about 1.92 units. for mgv, the indicated difference is about 1.6 units. 2. an increase in the average ugv speed by 10 km / h (variation step, table 2.1) leads to a decrease in the number of accidents in model 2 compared with the number of accidents in model 1 by 0.3 units. for mgv, the indicated difference is 0.7 units. in conclusion, we note that evaluating the dependence of the output stream traffic and the number of accidents on the specified model parameters using the linear function gives generally good results. only an assessment of the dependence of the number of accidents on the studied parameters in model 2 can be called not very successful due to the rather low value of the determination coefficient equal to 0.45. we proceed to the estimation of these dependences using a nonlinear function. 5. econometric analysis of traffic characteristics in the framework of nonlinear regression equations let's proceed to estimate the dependences of the output stream traffic and the number of accidents in both models on the indicated parameters of the model in the framework of nonlinear regression equations. so, we will build and study power regression equations with the following dependent variables: • output stream traffic in model 1; • output stream traffic in model 2; • number of accidents in model 1; • number of accidents in model 2. after logarithming of the equation, the coefficients were estimated in the framework of the classical multiple regression model using the least squares method, and standard errors were estimated using the white procedure. here are the results of evaluating the dependence of the output stream traffic (table 5.1) and the number of accidents (table 5.2) on the indicated set of explanatory variables. table 5.1. dependent variable –. output stream traffic, nonlinear regression explanatory variables model 1 model 2 constant -2.65*** (0.03) -3.79*** (0.02) intensityofuncrewedvehicles 0.08*** (0.002) 0.17*** (0.002) intensityofusualagents 0.08*** (0.002) 0.13*** (0.002) speedofuncrewedvehicles 0.55*** (0.003) 0.68*** (0.003) speedofusualagents 1.07*** (0.004) 1.35*** (0.003) intensityofconnections 0.002 (0.002) -0.0000013 (0.002) r2 0.72 0.77 influence of unmanned vehicles on the state of the transport network 53 copyright ©2020 assa. adv. in systems science and appl. (2020) table 5.2. dependent variable – number of accidents, nonlinear regression explanatory variables model 1 model 2 constant 11.77*** (0.06) 13.86*** (0.07) intensityofuncrewedvehicles 1.11*** 0.005 0.66*** (0.006) intensityofusualagents 1.10*** (0.004) 1.12*** (0.006) speedofuncrewedvehicles 0.30*** (0.008) 0.20*** (0.009) speedofusualagents -0.61*** (0.008) -1.61*** (0.009) intensityofconnections -0.07*** (0.005) 0.004 (0.005) r2 0.58 0.52 note. in parentheses are the values of standard errors. ***, **, * – significance at the 1, 5 and 10% level, respectively. comparing table 5.1 and table 5.2 with table 4.1 and table 4.2, respectively, we can conclude that power regressions give qualitatively the same conclusions as linear ones. this is especially true for the description of these dependencies for output stream traffic. as for the description of the dependence for the number of accidents, here we can identify some advantage of power regression. firstly, the value of the determination coefficient for model 2 increases from 0.45 to 0.52, and secondly, in contrast to the linear dependence, the intensityofconnections variable is statistically significant in the first model at a 1% level. the corresponding coefficient value in table 5.2 (-0.07) allows us to state that an increase in message intensity between ugv by 14% allows reducing the number of accidents by 1%. moreover, in the second model, this variable is insignificant. this is explained by the fact that in the second model, ugv have ability to estimate density of traffic flow, which allows them to maneuver effectively. in such a situation, the exchange of messages about the presence of an accident becomes irrelevant for the ugv. 6. conclusion in [6, 7] two agent-based models of movement of ugv and mgv, which were developed and implemented in anylogic simulation system, are described. the main difference between the second model and the first is that ugv is endowed with an additional capability the ability to intelligently assess the density of traffic flow. the main purpose of this article is to evaluate the effect obtained from the implementation of these ugv, which is to increase output stream traffic and reduce the number of accidents. for this purpose a pool of identical experiments was carried out within each model, the results of which became the initial basis for conducting an econometric analysis. for both the first and second models, the dependence of the output stream traffic and the number of accidents on a given road section in a given period of time on a number of model parameters such as the average vehicle speed, input stream intensity, and the frequency of exchanges between ugv was studied. the study of these dependencies was carried out both in the framework of linear and power regressions. evaluation of these dependencies made it possible to determine an increase in the output stream traffic and a decrease in the number of accidents in the second model relative to the indicated characteristics in the first model depending on the model parameters. acknowledgements the reported study was funded by rfbr, project number 19-29-06003. 54 a. akopov, n. khachatryan, f. belousov copyright ©2020 assa adv. in systems science and appl. (2020) references 1. ziye, z., haiou, l., huiyan, c., shaohang, x., & wenli, l. (2019). tracking control of unmanned tracked vehicle in off-road conditions with large curvature. ieee intelligent transportation systems conference, itsc. auckland, new zealand, 3867–3873, https://doi.org/10.1109/itsc.2019.8917468 2. menda, k., chen, y.-c., grana, j., bono, j. w., tracey, b. d., kochenderfer, m. j., & wolpert, d. (2019). deep reinforcement learning for event-driven multi-agent decision processes. ieee trans. intell. transp. syst., 20(4), 1259–1268, https://doi.org/10.1109/tits.2018.2848264 3. chen, c., liu, x., chen, h.-h., li, m., & zhao, l. (2019). a rear-end collision risk evaluation and control scheme using a bayesian network model. ieee trans. intell. transp. syst., 20(1), 264–284, https://doi.org/10.1109/tits.2018.2813364 4. xiong, g., kang, z., li, h., song, w., jin, y., & gong, j. (2018). decision making of lane change behavior based on rcs for automated vehicles in the real environment. ieee intelligent vehicles symposium. changshu, china, 1400–1405, https://doi.org/10.1109/ivs.2018.8500651 5. cao, s., shen, l., zhang, r., yu, h., & wang, x. (2019). adaptive incremental nonlinear dynamic inversion control based on neural network for uav maneuver. ieee/asme international conference on advanced intelligent mechatronics. hong kong, china, 642–647, https://doi.org/10.1109/aim.2019.8868510 6. akopov, a. s., beklaryan, l. a., khachatryan, n. k., beklaryan, a. l., & kuznetsova, e. v. (2020). mnogoagentnaja sistema upravlenija nazemnymi bespilotnymi transportnymi sredstvami [multi-agent ground-based unmanned vehicle control system]. informacionnye tehnologii, (in print), [in russian]. 7. akopov, a. s., beklaryan, l. a., khachatryan, n. k., beklaryan, a. l., & fomin, a.v. (2020). sistema upravlenija bespilotnymi transportnymi sredstvami na osnove nechetkoj klasterizacii [unmanned vehicle control system based on fuzzy clustering]. vestnik komp'juternyh i informacionnyh tehnologij, (in print), [in russian]. 8. goodfellow, i., bengio, y., & courville, a. (2016). deep learning. mit press. 9. nikolenko, s. i., kadurin, a. a., & arhangel'skaja, e. o. (2018). glubokoe obuchenie [deep learning]. st. petersburg, piter, [in russian]. 10. lecun, y., boser, b., denker, j. s., henderson, d., howard, r. e., hubbard, w., & jackel, l. d. (1989). backpropagation applied to handwritten zip code recognition. neural computation, 1(4), 541–551, https://doi.org/10.1162/neco.1989.1.4.541 11. akopov, a. s., & beklaryan, l. a. (2015). an agent model of crowd behavior in emergencies. autom. and remote control, 76(10), 1817–1827, https://doi.org/10.1134/s0005117915100094 12. akopov, a. s., beklaryan, l. a., & saghatelyan, a. k. (2019). agent-based modelling of interactions between air pollutants and greenery using a case study of yerevan, armenia. environmental modelling and software, 116, 7–25, https://doi.org/10.1016/j.envsoft.2019.02.003 13. bezdek, c. j. (1974). cluster validity with fuzzy sets. journal of cybernetics, 3(3), 58–73, https://doi.org/10.1080/01969727308546047 14. bezdek, c. j. (1981). pattern recognition with fuzzy objective function algorithms. springer us, https://doi.org/10.1007/978-1-4757-0450-1 influence of unmanned vehicles on the state of the transport network 55 copyright ©2020 assa. adv. in systems science and appl. (2020) 15. beklaryan, a. l., & akopov, a. s. (2016). simulation of agent-rescuer behaviour in emergency based on modified fuzzy clustering. proceedings of the international joint conference on autonomous agents and multiagent systems, aamas. singapore, 1275–1276. 16. akopov, a. s., beklaryan, l. a., thakur, m., & verma, b. d. (2019). parallel multiagent real-coded genetic algorithm for large-scale black-box single-objective optimization. knowledge-based systems, 174, 103–122, https://doi.org/10.1016/j.knosys.2019.03.003 adv syst sci appl 2023; 01:8–21 published online at https://ijassa.ipu.ru. on codimension 3 singular points of first-order implicit differential equations yana s. agakhanova1, alexander s. nikachev1*, irina n. oblasova2 1moscow institute of physics and technology (state university), dolgoprudnyi, russia 2north-caucasus federal university, stavropol, russia abstract: we investigate local phase portraits of first-order implicit differential equations in a neighborhood of their singular points of codimension 3. namely, we consider singular points where the lifted field of the equation on a surface is non-singular, but the projection of the surface to the phase plane has a singularity of one of three types: lips, beaks, swallowtail. we also consider generic bifurcations in one-parameter families of such equations. keywords: implicit differential equations, vector fields, singularities of mappings, fold, cusp, swallowtail 1. introduction the work in this paper is a part of an ongoing research on understanding singular points of first-order implicit differential equations f (x, y, p) = 0, p = dy/dx, (1.1) that is, points where fp vanishes (see, for example, [1] – [19] and the references therein). here x ∈ r1, y ∈ r1, and f : r3 → r1 is a smooth (c∞) function. since a single point of the (x, y)-plane may correspond to several different values of p such that f (x, y, p) = 0, equation (1.1) determines a multi-valued direction field on the (x, y)-plane. for studying this equation, we shall use the legendrian lift of the equation to a surface in the 1-jet space, which goes back to poincaré. the idea of this method is quite similar to the riemann surface for multi-valued functions of the complex variable. it is used in the most of modern works devoted to implicit differential equations, including the works mentioned below. moreover, this method naturally generalizes the case of multi-dimensional systems of implicit equations. the latter leads to an interesting phenomenon: vector fields with non-isolated singular points, see [12, 16]. the recent paper [20] should be especially remarked. 1.1. lifting of equation let j1 be the 1-jet space of functions y(x), i.e., the space with coordinates (x, y, p). equation (1.1) defines a surface f in j1, which is regular at points where ∇f ̸= 0. further, we shall always assume that this condition holds true, and the tangent plane is defined at every point of ∗corresponding author: nikachev441@gmail.com singular points of implicit differential equations 9 f . the intersection of the field of contact planes pdx− dy = 0 with the fields of the tangent planes defines a direction field of f , which is called the lifted field of equation (1.1). integral curves of equation (1.1) are obtained by the projection π of trajectories of the lifted field from the surface f to the (x, y)-plane parallel to the p-direction. the p-direction in the space j1 we shall call vertical. thus, a study of the phase portrait of equation (1.1) is reduced to the projection of the phase portrait of the lifted field from the surface f to the (x, y)-plane. the projection π has singularity at point of the surface f where fp = 0. singularities of the germs of smooth mappings are very well studied, the first result in this direction belongs to h. whitney. see, for example, [21–23] and [24, 25]. the intersection of the contact planes pdx− dy = 0 with the fields of the tangent planes to f can be written in the pfaffian form: fxdx+ fydy + fpdp = 0, pdx− dy = 0. (1.2) which yields the field ẋ = fp, ẏ = pfp, ṗ = −g, g = fx + pfy, (1.3) where the dot stands for differentiation by a new parameter t, which plays a role of time. formula (1.3) presents the lifted field of equations (1.1). we remark that the lifted field is defined at all points of the surface f except points where the contact plane coincide with the tangent plane, that is, fp = g = 0. the equality fp = 0 defines a subset of f , which is called the criminant of equation (1.1). points of the criminant are called singular points of equation (1.1), they are critical points of the projection π. the equality g = 0 defines a subset of f , which is called the inflection curve of equation (1.1), since its projection on the (x, y)-plane is the locus of inflection points of integral curves of (1.1). generically, the criminant and the inflection curve are curves on the surface f intersecting at isolated points, which are singular points of the lifted field (1.3). 1.2. classification of singular points remind that we excluded from consideration points where ∇f = 0 and we call a point of f a singular points of equations (1.1), if fp = 0. here we shall use terminology from [16]: definition 1.1: a singular point of equation (1.1) is called proper, if g ̸= 0 and improper otherwise. for every proper singular point there exists a unique integral curve of the field (1.3). the projection of this integral curve form f to the (x, y)-plane has a singularity. for simplicity, assume that the proper singular point is the origin of the space j1. if the germ of f at this point has a finite multiplicity by p, that is, ∂f ∂p (0) = 0, . . . , ∂n−1f ∂pn−1 (0) = 0, ∂nf ∂pn (0) ̸= 0 (1.4) with an integer n ≥ 2, then the germ of the corresponding integral curve of equation (1.1) has the form x = tnφ(t), y = tn+1ψ(t), φ(0)ψ(0) ̸= 0, (1.5) with smooth functions φ, ψ, see [3]. from the geometric viewpoint, here there are two different cases: if n is even or odd. in fig. 1.1 we present integral curves (1.5) with n = 2 (left) and n = 3 (right). a more interesting and complicated problem is a study of local phase portraits of equation (1.1) in a neighborhood of its singular point. in a neighborhood of a proper singular point the copyright © 2023 assa. adv syst sci appl (2023) 10 y. s. agakhanova, a. s. nikachev, i. n. oblasova fig. 1.1. two types of integral curves passing trough a proper singular point phase portrait of the lifted field on f is diffeomorphic to a family of parallel lines (the flow box theorem), whence the problem is reduced to the study of its projection. the phase portraits of equation (1.1) near a proper singular point where the mapping π has a fold or a cusp (pleat), are well studied; see [9, 10] or survey in [16]. for the convenience of the reader, we shall briefly describe the related results below. the phase portraits of equation (1.1) near an improper singular point where the mapping π has a fold (codimension 2) are studied in [9,10]. the phase portraits of equation (1.1) near an improper singular point where the mapping π has a pleat (codimension 3) are studied in [7]. the cases mentioned above cover singularities of equation (1.1) of codimensions 1 and 2, but not all singularities of codimension 3, see the table below. the columns correspond to all types of singularities of the projection π : f → r2 up to codimension 3 called fold, pleat, lips, beaks, swallowtail. the first row presents the corresponding left-right normal forms (see, for example, [24, 25]). the second (third) row contains proper (respectively, improper) singular points of equation (1.1) with the corresponding singularity of the mapping π. the dash stands in the cases if the codimension of singularity is greater than 3. in other cases we write new if the given type of singularity is not studied before, otherwise we give the references to related sources. it is worth observing that the left-right normal forms mentioned above are obtained by independent c∞-diffeomorphisms (changes of variables) in the image and preimage. therefore, they cannot be directly applied to simplification of equation (1.1), since a differential equation is connected with another group of transformations. fold pleat lips beaks swallowtail (u, v) → (z, w) z = u2 z = u3 + uv z = u3 + uv2 z = u3 − uv2 z = u4 + uv w = v w = v w = v w = v w = v g ̸= 0 codim = 1 codim = 2 codim = 3 codim = 3 codim = 3 [1, 2] [2, 9, 10] new new new g = 0 codim = 2 codim = 3 codim = 4 codim = 4 codim = 4 [2, 9, 10] [7] − − − the paper presents a study of three cases marked as new in the table above. for convenience of the reader, we start with a brief description of results for proper singular points with the fold and pleat of π. after that, we consider proper singular points with the swallowtail, lips and beaks of π, which are not considered before. together with [7], the obtained results cover all singularities of equation (1.1) of codimensions 1, 2, 3. copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 11 2. preliminary results without loss of generality, we shall assume that the considered singular point of equation (1.1) is the origin of the space j1. then the condition g(0) ̸= 0 is equaivalent to fx(0) ̸= 0, and the germ of f is the graph of a smooth function x = f(y, p), where f(0) = 0 and fp(0) = 0. therefore, we shall further assume that f (x, y, p) = f(y, p)− x. by the hadamard lemma, we have f (x, y, p) = f(y, p)− x = f0(y) + f1(y)p+ f2(y, p)p 2 − x, where fi are smooth functions, f0(0) = 0 (follows from f (0) = 0) and f1(0) = 0 (from fp(0) = 0). after the change of variables x̃ = f0(y)− x, we get f0(y) ≡ 0, and we brought equation (1.1) to the form f (x, y, p) = f(y, p)− x = αp2 + pyg(y) + p3h(y, p)− x = 0, (2.6) where g, h are smooth functions, α = fpp(0)/2 is a real number. 3. fold and pleat from now on, we shall deal with equation (2.6). if α ̸= 0, that is, fpp(0) ̸= 0, then the germ of π has a fold. if α = 0, but g(0) ̸= 0 and h(0) ̸= 0, that is, fpp(0) = 0, fppp(0) ̸= 0, fpy(0) ̸= 0, then the germ of π has a pleat (cusp). in fig. 3.2, we present the surface f and the projections of integral curves of the lifted field to the (x, y)-plane for the both cases. fig. 3.2. two singularities of the mapping π : f → r2. fold (on the left) and pleat (on the right). theorem 3.1: the germ of equation (1.1) at every proper singular point where π has a fold, is equivalent to the simple equation p2 − x = 0. (3.7) in a neighborhood of such singular point, the phase portrait of equation (1.1) is a family of semicubic parabolas, whose cusps fill the discriminant curve of the equations, the projection of the criminant to the (x, y)-plane; fig. 3.2 (left). theorem (3.1) states that an appropriate local diffeomorphism of the (x, y)-plane brings this family to the standard form y = 2 3 x3/2 + const, which corresponds to equation (3.7). the proof can be found in [1]. copyright © 2023 assa. adv syst sci appl (2023) 12 y. s. agakhanova, a. s. nikachev, i. n. oblasova consider the germ of equation (2.6) at a proper singular point where π has a pleat; see fig. 3.2 (right). using topological arguments, in [9] it is proved that the classification of the germs of equations (2.6) with pleat has functional invariants. moreover, functional invariants do not disappear even if changes of variables are homeomorphisms. however, the geometric classification of possible phase portraits is quite simple. consider integral curves of the corresponding lifted field on the surface f , which has the form ẏ = pfp(y, p), ṗ = 1− pfy(y, p). (3.8) let us present f in the form f(y, p) = apy + bp3 + py2h1(y, p) + p4h2(y, p), ab ̸= 0. (3.9) then, substituting (3.9) into (3.8), one can see that the criminant of equation (2.6) is given by the equation y = kp2 + o(p2) with k = −3b/a. in a neighborhood of the origin, the criminant is similar to the parabola y = kp2, which is tangent to the direction (3.8) of zero. this point is the cusp of the discriminant curve, which has the form x = f(kp2 + o(p2), p) = −2bp3 + o(p3), y = kp2 + o(p2). this yields the following result: theorem 3.2: there exist only two geometrically different types of phase portraits of the field (3.8) corresponding to different signs of k. these phase portraits and their projections to the (x, y)plane are presented in fig. 3.3 (up and down, respectively). fig. 3.3. down: two geometrically different types of phase portraits of equation (2.6) at a proper singular point where π has a pleat. up: the corresponding lifted fields. in the both cases, there exists a narrow domain on the (x, y)-plane, where integral curves constitute a 3-web. 4. singularities of codimension 3 in this section, we present the main results of the paper: a study of the local phase portraits of equation (1.1) at its proper singular points, where the mapping π has a singularity of one of three following types: swallowtail, lips, beaks. copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 13 4.1. swallowtail consider the germ of equations (1.1) at its proper singular points, where the mapping π has a swallowtail (not to be confused with a surface with the same name). this means that fp(0) = fpp(0) = 0, fpy(0) ̸= 0, fppp(0) = 0, fpppp(0) ̸= 0. (4.10) under conditions (4.10), the germ of equation (2.6) has the form f (x, y, p) = f(y, p)− x = 0, f(y, p) = apy + bp4 + py2h1(y, p) + p2yh2(y, p) + p5h3(y, p), ab ̸= 0. (4.11) the surface of such equation is presented in fig. 4.4, together with the criminant of the equation and its discriminant curve. the lifted field on f has the form ẏ = pfp(y, p), ṗ = 1− pfy(y, p). (4.12) the above formulas show that the criminant of (4.11) has the form y = kp3 + o(p3), where k = −4b/a. in a neighborhood of the origin, the criminant is similar to the cubic parabola y = kp3, which is tangent to the field (4.12) at the origin. the tangency point correspond to a singular point of the discriminant curve: x = f(kp3 + o(p3), p) = −3bp4 + o(p4), y = kp3 + o(p3). (4.13) fig. 4.4. left: the surface f of equations (1.1) whose mapping π has a swallowtail; the bold curve on f is the criminant of the equation, and the bold curve on the plane is its discriminant curve. right: intersection of the surface s (blue) with the surface swallowtail (red) in the space of the coefficients. lemma 4.1: locally, the discriminant curve of equation (4.11) splits the (x, y)-plane into two open domains, in which equation (4.11) has two real roots p or zero real roots p, respectively. at point of the discriminant curve (except the origin) it has one double real root p, while at the origin the multiplicity of p is equal to 4. proof by the division theorem [22, 23], the germ of equation (2.6) satisfying conditions (4.10) is equivalent to f̃ (x, y, p) = p4 + 3∑ i=0 ai(x, y)p i = 0, p = dy/dx, (4.14) where ai(x, y) are smooth functions, ai(0) = 0, and da0, da1 at zero are linearly independent. using an appropriate change of variables (x, y), one can kill the monomial p3, that is, bring copyright © 2023 assa. adv syst sci appl (2023) 14 y. s. agakhanova, a. s. nikachev, i. n. oblasova (4.14) to the form p4 + a(x, y)p2 + b(x, y)p+ c(x, y) = 0, p = dy/dx, (4.15) with smooth functions a, b, c, a(0) = b(0) = c(0). form the independence of da0, da1 in (4.14) it follows the independence of db, dc in (4.15). then the statement of the lemma can be interpreted in the following way. in the space with cartesian coordinates a, b, c consider the discriminant surface of the polynomial m(p) = p4 + ap2 + bp+ c, (4.16) that is, a surface that consists of points (a, b, c) that the polynomial m(p) has multiple roots. this surface w defined by the equations m(p) =m ′(p) = 0 is called the swallowtail, it is presented in fig. 4.4 (right) in red. it separates the whole space into three open domains u4, u2, u0, where m(p) has 4, 2, 0 real roots, respectively. in the (a, b, c)-space equation (4.15) defines a surface s a = a(x, y), b = b(x, y), b = c(x, y), (4.17) it is presented in fig. 4.4 (right) in blue. since the differentials of the functions b, c at zero are linearly independent, in a neighborhood of the origin the surface s is regular and it intersects the swallowtail w as presented in fig. 4.4 (right). this proves all statements of the lemma. theorem 4.1: there exist only two geometrically different types of local phase portraits of equation (4.11) corresponding to different signs of k, which is equivalent to those of ab. these phase portraits are presented in fig. 4.5. fig. 4.5. down: two different types of local phase portraits of equation (1.1), whose mapping π has a swallowtail, which correspond to different signs of ab in (4.11). up: the corresponding phase portraits of the lifted field. the dotted lines present the criminant of the equation (up) and its discriminant curve (down). proof consider local phase portraits of the lifted field (4.12) for equation (4.11). in a neighborhood of the origin integral curves of the field (4.12) are similar to straight lines parallel to the paxis, and the criminant is similar to the cubic parabola y = kp3. this proves the statement of the theorem. copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 15 4.2. lips and beaks consider the germ of equations (1.1) at its proper singular points, where the mapping π has a singularity of the type lips or beaks. then, after an appropriate linear change y 7→ αy equation (2.6) can be brought to the form f (x, y, p) = f(y, p)− x = 0, f(y, p) = p3 + h1(y, p)yp 2 + h2(y, p)y 2p+ h3(y, p)p 4, (4.18) where the hessian of the function fp(y, p) at the origin is not zero: h[fp](0) = 4(3h2(0)− h21(0)) ̸= 0. (4.19) the mapping π has a singularity of the type lips (respectively, beaks) if h[fp] > 0 (respectively, h[fp] < 0). the corresponding surfaces of the equation (1.1) are presented in fig. 4.6. fig. 4.6. the surface f of equations (1.1) whose mapping π has a singularity of the type lips (left) or beaks (right). the bold curve on f is the criminant and the bold curve on the plane is its discriminant curve. by the division theorem, the germ of the function f from (4.18) has the form f (x, y, p) = φ(x, y, p) ( p3 + a(x, y)p2 + b(x, y)p+ c(x, y) ) , where a, b, c,φ are smooth functions, a(0) = b(0) = c(0) = 0, and φ(0) ̸= 0. therefore, the germ of equation (4.18) is equivalent to p3 + a(x, y)p2 + b(x, y)p+ c(x, y) = 0, p = dy/dx. (4.20) lemma 4.2: in (4.20), the coefficient c has the form c(x, y) = xc̃(x, y) with a smooth function c̃. the condition g ̸= 0 for equation (1.1) is equivalent to c̃(0, 0) ̸= 0. proof formula (4.18) shows that the surface f given by the equation x = f(y, p) contains the y-axis (x = p = 0). this means that the substitution x = p = 0 into (4.20) gives a certain identity by y. this yields c(x, y) ∣∣ x=0 ≡ 0, and the statement follows from the hadamard lemma. lemma 4.3: in (4.20), the equality by(0) = 0 holds true. copyright © 2023 assa. adv syst sci appl (2023) 16 y. s. agakhanova, a. s. nikachev, i. n. oblasova proof substitution x = f(y, p) with the function f from (4.18) brought equation (4.20) to a certain identity by y, p. the left hand side of this identity contains the monomial by(0)yp, while the right hand side is identically zero. this proves the statement. lemma 4.4: hessian of the function fp at zero is equal to h[fp](0) = 4φ2(0)∆, where ∆ = 3 2 byy(0)− a2y(0), whence the type of singularity of π is determined by the sign of ∆: lips if ∆ > 0 and beaks if ∆ < 0. proof from (4.18) we have: fp(y, p) = fp(x, y, p) = φp(x, y, p)(p 3 + a(x, y)p2 + b(x, y)p+ c(x, y))+ +φ(x, y, p)(3p2 + 2a(x, y)p+ b(x, y)), where x is replaced with f(y, p). let us calculate the quadratic term of the obtained function on y, p. from lemmas 4.2 and 4.3 it follows that the 2-jet of the function p3 + a(f(y, p), y)p2 + b(f(y, p), y)p+ c(f(y, p), y) is zero and the 1-jet of 3p2 + 2a(f(y, p), y)p+ b(f(y, p), y) (4.21) is zero as well. therefore, the quadratic part of the germ fp(y, p) coincides with the quadratic part of (4.21) multiplied by φ(0). taking into account 4.3, a simple calculation shows that the quadratic part of (4.21) is 3p2 + 2ay(0)py + 1 2 byyy 3. this yields the statement of the lemma. theorem 4.2: the local phase portrait of equation (1.1) at proper singular point whose mapping π has a singularity lips is presented in fig. 4.7. here the discriminant set of the equation is a single point, and integral curves are regular at all points except for this point. theorem 4.3: the local phase portrait of equation (1.1) at proper singular point whose mapping π has a singularity lips is presented in fig. 4.8. here the discriminant set of the equation is a pair of two curves tangent at a single point, which separate the (x, y)-plane into four open domains with different stricture of integral curves: 1-webs or 3-webs. proof to prove theorems 4.2 and 4.3, consider the left hand side of equation (4.20) as a polynomial on the variable p with coefficients smoothly depending on x, y: m(p) = p3 + a(x, y)p2 + b(x, y)p+ c(x, y). (4.22) in the space with cartesian coordinates a, b, c consider the discriminant surface of the polynomial m(p) = p3 + ap2 + bp+ c, that is, a surface that consists of points (a, b, c) that the polynomial m(p) has multiple roots. this surface w defined by the equations m(p) =m ′(p) = 0 is called the cuspidal edge, it is presented in fig. 4.9 in orange. it copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 17 fig. 4.7. left: local phase portrait of equation (1.1) whose mapping π has lips. right: the corresponding phase portrait of the lifted field. here the criminant and the discriminant set are isolated points. fig. 4.8. left: local phase portrait of equation (1.1) whose mapping π has beaks. right: the corresponding phase portrait of the lifted field. the criminant and the discriminant set of the equation are depicted with dotted lines. separates the whole space into two open domains u1, u3, where m(p) has 1, 3 real roots, respectively. the domain u3 is more narrow. in the (a, b, c)-space equation (4.15) defines a surface s defined by (4.17). the proof of theorems 4.2, 4.3 is based on the intersection of the surfaces s and w . using lemmas 4.2 – 4.4, it is not hard to show that the mutual position of s and w is exactly as it is presented in fig. 4.9 for lips and beaks. here the surfaces s and w are depicted in blue and orange, respectively. in the case ∆ < 0 (lips) the surface s belongs to the open domain u1 except for a point, which corresponds to the unique singular point of the mapping π. in the case ∆ > 0 (beaks) the surface s is cutted by w into four parts, two of which belong to u1 and two belong to u3. this proves the theorems. fig. 4.9. the mutual position of surfaces w and s for singularities lips (left) and beaks (right). copyright © 2023 assa. adv syst sci appl (2023) 18 y. s. agakhanova, a. s. nikachev, i. n. oblasova 5. bifurcations on one-parameter families 5.1. bifurcations of the swallowtail consider a generic one-parameter perturbation of equation (4.11): p4 + a(x, y, µ)p2 + b(x, y, µ)p+ c(x, y, µ) = 0, p = dy/dx, (5.23) where µ is a small parameter. in the space with cartesian coordinates a, b, c consider the family of surfaces sµ defined by the coefficients a(x, y, µ), b(x, y, µ), c(x, y, µ) of equation (5.23). in the same space consider the discriminant surface w of the polynomial m(p) = p4 + ap2 + bp+ c, which is the swallowtail. the bifurcation of the phase portrait of equation (5.23) is determined by the intersection of the family sµ with w . assume that ∂a ∂µ (0, 0, 0) ̸= 0. (5.24) then the function a(0, 0, µ) changes its sign when the parameters µ passes through zero. without loss of generality we shall further assume that the sign of the function a(0, 0, µ) coincides with the sign of µ. then the mutual position of sµ and w is presented in fig. 5.10, and the corresponding bifurcation of the phase portrait of equation (5.23) is presented in fig. 5.11. it is worth observing that for all µ < 0 the exists a narrow domain on the (x, y)plane where integral curves constitute a 4-web. as µ→ −0 this domain tends to a point, and for µ ≥ 0 integral curves have no quadruple intersections. fig. 5.10. the surfaces sµ and w for a generic family (5.23) depicted in blue and red respectively. from the left to right: µ < 0, µ = 0, µ > 0. fig. 5.11. bifurcation of a generic family (5.23): the phase portrait on the (x, y)-plane, the dotted line is the discriminant curve. from the left to right: µ < 0, µ = 0, µ > 0. copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 19 5.2. bifurcations of lips and beaks consider a generic one-parameter perturbation of equation (4.18) with singularity of the type lips or beaks: f (x, y, p) = p3 + a(x, y)p2 + b(x, y)p+ c(x, y), p = dy/dx, (5.25) where µ is a small parameter. in the space with cartesian coordinates a, b, c consider the family of surfaces sµ defined by the coefficients a(x, y, µ), b(x, y, µ), c(x, y, µ) of equation (5.25). in the same space consider the discriminant surface w of the polynomial m(p) = p3 + ap2 + bp+ c, which is the cuspidal edge. the bifurcation of the phase portrait of equation (5.25) is determined by the intersection of the family sµ with w . assume that ∂b ∂µ (0, 0, 0) ̸= 0, then the function b(0, 0, µ) changes its sign when the parameters µ passes through zero. without loss of generality we shall further assume that the sign of the function b(0, 0, µ) coincides with the sign of µ. the mutual position of sµ and w for a generic family (5.25) with singularity of the type lips is presented in fig. 5.12. the corresponding bifurcation of the phase portrait is presented in fig. 5.13. for all µ > 0 there exists a narrow bounded domain on the (x, y)-play (which looks like lips), where integral curves constitute a 3-web. as µ→ +0 this domain tends to a point, and for µ ≥ 0 integral curves have no triple intersections. fig. 5.12. the surfaces sµ and w for a generic family (5.25) with lips depicted in blue and orange respectively. from the left to right: µ > 0, µ = 0, µ < 0. fig. 5.13. bifurcation of a generic family (5.25) with lips. from the left to right: µ > 0, µ = 0, µ < 0. the closed bold curve in the left image is the discriminant set, which degenerates into a point when µ = 0. the mutual position of sµ and w for a generic family (5.25) with singularity of the type beaks is presented in fig. 5.14. the corresponding bifurcation of the phase portrait is presented in fig. 5.15. copyright © 2023 assa. adv syst sci appl (2023) 20 y. s. agakhanova, a. s. nikachev, i. n. oblasova fig. 5.14. the surfaces sµ and w for a generic family (5.25) with beaks depicted in blue and orange respectively. from the left to right: µ > 0, µ = 0, µ < 0. fig. 5.15. bifurcation of a generic family (5.25) with beaks. from the left to right: µ > 0, µ = 0, µ < 0. the dotted curve is the discriminant set. 6. conclusion in this paper, we studied local phase portraits of first-order implicit differential equations, that is, equations in the form f (x, y, y′) = 0, in a neighborhood of their proper singular points of codimension 3. this study, together with previous works, completes qualitative investigation of singularities of such equations up to codimension— 3. moreover, we investigated generic bifurcations in one-parameter families of such equations. it is worth observing that an interest to such equations is motivated by their applications in various mathematical and technical problems; see, for example, [6] and the lists of references in [11, 14]. references 1. arnol’d, v. i. (1988) geometrical methods in the theory of ordinary differential equations, berlin, germany: springer. 2. arnol’d, v. i. & ilyashenko, yu. s. (1988). ordinary differential equations, dynamical systems i. encycl. math. sci. 1, 1–148. 3. bruce, j. w. (1984). a note on first-order differential equations of degree greater than one and wavefront evolution. bull. london math. soc., 16, 139–144. 4. bruce, j. w., fletcher, g. j., & tari, f. (2000). bifurcations of binary differential equations, proc. royal society of edinburg, 130 a, 485–506. 5. bruce, j. w., fletcher, g. j., & tari, f. (2004). zero curves of families of curve congruences, contemporary math., 354, 1–18. 6. barlukova, a. m. & chupakhin, a. p. (2012). partially invariant solutions in gas dynamics and implicit equations, j. appl. mech. tech. phys., 53(6), 812–824. 7. chertovskih, r. a. & remizov, a. o. (2014). on pleated singular points of first-order implicit differential equations, j. dyn. control syst., 20(2), 197–206. copyright © 2023 assa. adv syst sci appl (2023) singular points of implicit differential equations 21 8. davydov, a. a., ishikawa, g., izumiya, s., & sun, w.-z. (2008). generic singularities of implicit systems of first order differential equations on the plane, jpn. j. math., 3(1), 93–119. 9. davydov, a. a. (1985). the normal form of a differential equation, that is not solved with respect to the derivative, in the neighborhood of its singular point, funktsional. anal. i prilozhen., 19(2), 1–10. 10. davydov, a. a. (1994). qualitative theory of control systems, math. monographs, 141, ams, providence, rhode island. 11. kotyukov, a. m., nikanorov, s. o., & pavlova, n.g. (2020). local normal forms of autonomous quasi-linear constrained differential systems, adv. syst. sci. appl., 20(1), 119–127. 12. pavlova, n. g. & remizov, a. o. (2021). smooth local normal forms of hyperbolic roussarie vector fields, moscow math. j., 21(2), 413–426. 13. pavlova, n. g. & remizov, a. o. (2022). oscillating and proper solutions of singular quasi-linear differential equations, adv. syst. sci. appl., 22(4), 51–64. 14. pazij, n. d. & pavlova, n. g. (2022). local analytic classification for quasi-linear implicit differential systems at transversal singular points. j. dyn. control syst., 28(3), 453–464. 15. remizov, a. o. (2002). implicit differential equations and vector fields with nonisolated singular points, sb. math., 193(11), 1671–1690. 16. remizov, a. o. (2008). multidimensional poincaré construction and singularities of lifted fields for implicit differential equations, j. math. sci., 151(6), 3561–3602. 17. seiler, w. m. & seiss, m. (2021). singular initial value problems for scalar quasi-linear ordinary differential equations, j. differ. equations, 281, 258–288. 18. tari, f. (2007). geometric properties of the integral curves of an implicit differential equation, discrete contin. dyn. syst., 17(2), 349–364. 19. sotomayor, j. & zhitomirskii, m. (2001). impasse singularities of differential systems of the form a(x)x′ = f (x), j. differ. equations, 169(2), 567–587. 20. medrado, j. c. r. & silva, l. a. (2022). codimension-one singularities of constrained systems on r3, bull. sci. math., article id 103142, 22 p. 21. arnol’d, v. i., gusein-zade, s. m., & varchenko, a. n. (1985). singularities of differentiable maps. volume i: the classification of critical points, caustics and wave fronts, monogr. math., 82, boston-basel-stuttgart: birkhauser. 22. broker, th. & lander, l. (1975). differential germs and catastrophes, cambridge, uk: cambridge university press. 23. golubitsky, m. & guillemin, v. (1973). stable mappings and their singularities, grad. texts math., 14, new york-heidelberg-berlin: springer-verlag. 24. rieger, j. h. (1987). families of maps from the plane to the plane, j. london math. soc., ii, 36(1–2), 351–369. 25. saji, k. (2010). criteria for singularities of smooth maps from the plane into the plane and their applications, hiroshima math. j., 40(2), 229–239. copyright © 2023 assa. adv syst sci appl (2023) introduction lifting of equation classification of singular points preliminary results fold and pleat singularities of codimension 3 swallowtail lips and beaks bifurcations on one-parameter families bifurcations of the swallowtail bifurcations of lips and beaks conclusion adv syst sci appl 2019; 03; 131-139 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/718 modeling and estimating the impact of the opec agreement on oil production in russia valery akinfiev * v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: akinf@ipu.ru received march 27, 2019; revised august 28, 2019; published october 1, 2019 abstract: the shale revolution in the united states dramatically changed the situation on the global oil market. the response to the shale boom of 2013-2014 was the opec + agreement, which aims to regulate the market in order to stabilize oil prices. the article analyzes the impact of the opec+ agreement on the global oil market, taking into account the dynamics of us oil shale production. a simulation dynamic game model is proposed, which describes the relationship between the oil demand, the strategy of oil supplying by the players and the dynamics of oil price. the model allows, by setting different strategies for the behavior of players, to simulate the movement of oil prices, which in turn affect the choice of players in their decisions on oil production and its supply to the market. the model includes three players: us shale companies, opec+ countries (including russia), and non-opec countries (including conventional us oil production). the research methodology is based on scenario modeling and analysis of opec+ strategies for managing the supply and demand balance (decrease or increase in production) for targeting oil prices at specified target levels. we tried in this study to answer the question: is the opec agreement favorable for russia? calculations show that for russia, in terms of maintaining market share, the option of targeting oil prices at lower levels is the most profitable, but at the same time, gross oil revenues of russia are noticeably reduced. keywords: opec agreement, oil market models, supply and demand balance. 1. introduction the emergence of fracking shale oil production technology in the united states has changed the outlook for the global petroleum industry in recent years. it was the us shale industry that became the main reason for the sharp increase in oil supply in the market in 2013-2015, which led to a sharp drop in oil prices to 30 usd/bbl in mid-2014. fig. 1. us oil production dynamics. * corresponding author: akinf@ipu.ru mailto:akinf@ipu.ru 132 v. akinfiev copyright ©2019 assa adv. in systems science and appl. (2019) figure 1 plots the dynamics of the us crude oil production, for 1985-2015 (in million barrels per day). source: u.s. energy information administration (eia), 2015. the explosive growth of shale oil production in the united states in 2013-2015 came as a surprise to the market. in 2015, the volume of shale oil produced in the united states almost reached the level of traditionally produced oil more than 4.5 million barrels per day. for reference, it is almost half of the volume of oil production in saudi arabia or russia. if in the period up to 2013, the increase in oil supply on the market was offset by an increase in demand, in 2014 oil consumption grew at a lower rate, and the difference between supply and demand reached 1.5-2.0 million barrels per day. such an imbalance of supply and demand led to a shock reduction in the price of oil. opec could reduce the level of oil production and thus eliminate the excess supply, which would help the price to return to a higher level. however, in 2014, opec abandoned the use of such a strategy. this decision is aimed at maintaining its share by lowering oil prices. shale oil producers reacted to the decline in oil prices by lowering drilling volumes at the least profitable fields: the number of active drilling rigs almost halved by 2016. but by 2016, the world oil market has moved from the phase of competition and the price war of 2014–2015 to market regulation aimed at stabilizing prices at the desired levels. therefore, in november 2016, opec agreed to reduce oil production by 1.2 million barrels per day. in december, another 11 non-opec countries, including russia, joined the agreement. the parties to the agreement agreed to reduce production by 1.75 million barrels per day compared with the level of october 2016. a similar decision was made in december 2018. according to it, the opec + countries, including russia, decided to reduce oil production by 1.2 million barrels per day. the target for the oil price set by the agreement was 70 + usd/bbl. saudi arabia and other opec members in the current situation are counting on cooperation with oil-producing countries led by russia, which will help reduce the negative effects of the american shale revolution on the world oil market. opec + solutions aimed at restoring the balance of supply and demand in the oil market can be quite effective. the reduction in production of opec + leads to an increase in reserve production capacity, which reduces price volatility. nevertheless, thanks to the implementation of the opec + strategy to regulate the market, shale oil companies can develop and increase production at high rates. according to the opec forecast, the annual growth of shale oil production is expected on average by 1.4 million barrels per day in 2019-2020 [10]. then, after 2023, a slowdown in the growth of shale oil production in the united states is predicted. in the period 2027-2028 peak shale oil production is projected at 14.3 million barrels per day (or approximately 13.5% of the market share). then, by 2040, production will drop to 12.1 million barrels per day (approximately 10.5% of the market share). for comparison, in 2018, shale oil production in the united states reached 7.5 million barrels per day, which accounted for about 7.5% of the market share [10]. opec also expects growth in global oil demand in the medium term by 1.2 million barrels per day to 104.5 million barrels per day by 2023. then, a “gradual slowdown in growth rates” of demand is projected, which will reach a maximum in 2040 at the level of 110.0–111.5 million barrels per day. this forecast is consistent with the iea forecast, which also implies a slower growth in global oil demand after 2040. leading global oil companies, such as royal dutch shell and bp, expect a peak in global demand in the years 3035–2040 at 110.3–110.5 million barrels per day. the important question is how profitable is the opec + agreement for russia? the opec forecast was made on the condition that oil prices remain at 70 usd/bbl. in this situation, the share of shale oil in 7-9 years can almost double. at the same time, the share of the remaining producing countries, including russia, will decrease. it is clear that it is desirable for russia to maintain its share in the global market, which in 2018 was 11.8%. but, clearly, this possibility occurs at lower price levels. modeling and estimating the impact of the opec agreement 133 copyright ©2019 assa. adv. in systems science and appl. (2019) it should be noted that the growth rate of shale oil production largely depends on the price of oil on the global market. of practical interest is the study of scenarios in which opec + will support other price ranges, for example, (60-70), (50-60), (40-50) usd/bbl. everywhere else, if it is not specifically stated, the price of oil refers to the brent price. what is oil price russia needs to fulfill budget commitments and economic development, as well as to save its share in the global oil market? the article is devoted to the discussion of these important issues. 2. model of behavior of u. s. shale producers the us shale revolution has changed the parameters of the global oil market model: the supply of oil by producers has become more elastic. us shale companies increase oil production in response to price increases, and, conversely, reduce production if oil becomes cheaper. a quick response from the us shale industry to the changing situation in the oil market is capable of keeping prices within a narrow range, which also depends on the oil production strategy of other of other producers. before the us shale revolution, oil price shocks were common due to the inelastic supply of conventional oil. an important question arises how can we describe the behavior of us shale companies depending on the dynamics of world oil prices? currently, the shale oil companies are drilling and producing at five shale plays, which differ significantly among themselves, including production profiles, well productivity, well flow rate growth dynamics, transport costs and selling price. the main indicator that affects the profitability of investments in shale projects and their ability to generate positive cash flow is the breakeven price. figure 2 shows the distribution of wti breakeven prices for all 2q 2018 horizontal oil completions for the major u.s. liquids plays (source: rystad energy shale well cube, august 2018). the breakeven price calculation is based on the traditional dcf model, which includes all investments, operating and fiscal costs (excluding land value) plus a 5% yield. fig. 2. breakeven oil prices distribution for main shale plays it should be borne in mind that due to the exploitation of shale basins, companies will have to move from areas with a lower break-even price to areas with a higher break-even price. for this reason, in the future, the average break-even price for the main us shale basins will rise, despite improvements in production technology and lower costs for some of the cost items. another important factor that can influence the rate of u.s. shale oil production is also the size of the proven reserves. according to the latest data provided by the sec, reserves in the permian basin (the most promising in terms of development) can reach only 3.8 billion barrels. the eagle ford and bakken pools can hold 5 billion barrels each. a number of publications are devoted to the problem of forecasting us shale oil production [6, 8, 9]. the most popular is the approach associated with the use of various modifications of the autoregressive models. in [6], to estimate future supplies of shale oil (using the example of the bakken basin), it is proposed to use a non-linear and linear forecasting model based on the auto regressive integrated moving average method 134 v. akinfiev copyright ©2019 assa adv. in systems science and appl. (2019) (arima). in [8], a new hybrid model nmgm-arima was proposed to forecast u.s. shale oil production, in which the arima method was supplemented with a non-linear model to correct the residual term of the sequence. empirical results show that the nmgm-arima method can improve forecast efficiency [8]. these methods are designed to forecast the horizon for several quarters. another approach is to use models in which the relationships between a large numbers of factors are modeled and expert information is widely used. these methods allow you to make forecasts with a horizon of several years. note that in these methods, the market price of oil is also exogenous and is set in the form of several scenarios. in [9], the results of studies obtained using the medium-term model of the us shale industry are presented. the model generates forecasts for 6 key indicators for each of the five shale basins (bakken, eagle ford, anadarko, niobrara and perm). the model allows building forecasts of production dynamics depending on the wti oil prices, possible growth of drilling equipment productivity and other technological and resource constraints. the simulation results show that, if the wti oil price is at the level of usd 70 per barrel and higher, then annual production growth in the next two to three years may reach 2.0-2.5 million barrels per day [9]. if the price of wti oil is in the range of 60-70 us dollars per barrel, then production growth is projected at the level of global demand growth, that is, at the level of 1.2-1.4 million barrels per day. when the price of wti crude oil is about $ 50 per barrel, production growth will be negligible or even zero. with prices below $ 50 a barrel, shale projects will no longer be attractive to investors. it should be borne in mind that the price of wti crude oil is on average 10% lower than the price of brent crude oil. we will use these results and the results of other analysts when building a model for the behavior of us shale producers. to implement the proposed approach, it is necessary to build the dependence of u.s. shale oil production on two parameters: the market price of oil and the forecast period. the solution to this problem depends on a large number of factors, many of which can only be evaluated expertly. when constructing these dependencies in fig. 3, the author used the results obtained in [9], as well as research results and forecasts from other sources. let tgshale the increase in shale oil production in period t. tgshale depends on the dynamics of oil prices and the time elapsed from the beginning of the forecast period. then ),(g 1t tpt shale  , where tp is the price of oil on the market. this dependence can be defined as a series of graphs depending on the period of time and the dynamics of production growth in previous periods. the averaged dependences of the increase in u.s. shale oil production on the wti oil price for the periods (2019-2020), (2023-2024) and (2026-2027) are shown in figure 3. fig. 3. the dependence of u.s. shale oil production growth on the wti oil price -4,00 -3,00 -2,00 -1,00 0,00 1,00 2,00 3,00 30 40 50 60 70 80 90 м ill io n b ar re ls p e r d ay wti oil price, usd per barrel the annual production growth 2019 2020 2023 2024 2026 2027 modeling and estimating the impact of the opec agreement 135 copyright ©2019 assa. adv. in systems science and appl. (2019) then the production of shale oil is given by the following recurrent formula: ),(ss 11-tt tpt shaleshale  (1) please note that the price of oil )( tp is not known in advance and is determined based on the market pricing mechanism, depending on the ratio of supply and demand. we will discuss this relationship in detail in the next section. 3. the model of the oil market based on the supply and demand balance in [1, 4, 5, 7], the problem of modeling the oil market and forecasting oil prices was considered, taking into account the dynamics of u.s. shale oil production. in [4], a game approach was proposed for analyzing the strategies of opec +, taking into account the dynamics of u.s. shale oil production. various cournot equilibrium models were used, in which the pricing model in the market is defined as a function of linear demand. it is assumed that the oil demand on the part of consumers depends linearly on its market price. this approach does not take into account the low elasticity of oil demand from its price. based on calculations and comparison of the obtained results with historical data for the period 2014–2016, the authors concluded that the proposed approach does not explain the real data and needs to be improved. in [5], using the structural vector auto regression (var) model, factors affecting the appearance of oil market shocks is analyzed. the study made it possible to draw a number of interesting conclusions regarding the importance of the demand and supply balance factor in explaining fluctuations in the oil market, as well as the effect of various assumptions regarding the elasticity indicators of the oil market. the relationship between oil prices and macroeconomic indicators is analyzed in [7] based on the general equilibrium model, which takes into account technological shocks depending on the development of new technologies, oil supply shocks and forecasts for future oil supplies. in [1] the problem of choosing the investment strategies of oil companies with production of conventional and shale oil is considered. a mathematical model is proposed that describes the relationship between the investment strategies of companies and the price of oil, which depends on the ratio of supply and demand on the world oil market. the solution is reduced to the analysis of the bimatrix game, in which the payoff matrix is formed as a result of numerical simulation. we present a simulation dynamic game model, which is shown in figure 4. the model describes the relationship between oil demand, the strategy of oil supply by producers and the dynamics of changes in the price of oil. the model allows, by setting various strategies of producers 'behavior, to calculate the movement of oil prices, which, in turn, affects the choice of producers' decisions regarding the oil supply. the model includes three producers (let's call them players): us shale companies, opec+ countries (including russia) and nonopec countries (including conventional us oil producers) the research methodology is based on scenario modeling and analysis of the opec + management strategies to maintain the supply and demand balance (decrease or increase in production) for targeting oil prices at specified target levels. when modeling each scenario, two criteria are assessed: gross oil revenues for a period and the dynamics of changes in market share. 136 v. akinfiev copyright ©2019 assa adv. in systems science and appl. (2019) fig. 4. integrated model let t opecx control by opec + in the form of a decrease or increase oil production in period t. the value t opecx is determined depending on the ratio of market oil prices in the previous period t-1 and the boundaries of a given interval. the choice mechanism is denoted by , ),,(x вн tt opec ppp , where нp and вp are the boundaries of the target range. for example, this mechanism may be as follows: tttt opec ksd    )(x 11 , where the coefficient tk , which depends on the degree of deviation 1tp from the target level )10(  tk . t kd the scenarios of global oil demand. it is assumed that the demand t kd is unknown in advance to market participants. the pricing model on the market is given by the recurrence formula [2]:   ) s ()(1( 2 2 t 2*       t t t k t t t t t k t d d ttpp  (2) here ts is the total oil supply on the market, t nonopec t opecshale ss   tt ss . in accordance with [2], the coefficient of price elasticity in the model   )(t , if 0st t kd and   )(t otherwise. the medium value  is equal to 9.75 and, accordingly,  is equal to 28.7. the model uses the information on the dynamics of supply and demand for the two previous periods to calculate the forecast price. the value  tp* is calculated as follows:    0** ptp  , where  0*p is the oil price in the initial period. if at some time "t    tstd  , then  tp* "tp for all "tt  . the model takes into account the "hysteresis" property when calculating the price. the new equilibrium value of the price can be determined at a level different from the initial one. in [2], it was shown that the coefficient of elasticity in the zone of oil deficiency is less than the coefficient of elasticity in the zone of surplus. this is due to the fact that in a period of high oil prices even with a shortage of oil (demand exceeds supply), it is more profitable for oil companies to increase production and not raise prices too high, since this may lead to an irreversible decrease in demand due to market adaptation to new conditions. during the period of oil surplus on the market (supply exceeds demand), it is more profitable for oil companies to reduce the oil price and thereby stimulate oil demand. moreover, companies with a low break-even point have the advantage of reducing the oil model behavior us shale companies market model demand scenarios opec +: production and supply of oil scenarios of production of non-opec countries modeling and estimating the impact of the opec agreement 137 copyright ©2019 assa. adv. in systems science and appl. (2019) price without losing a positive profitability. as a result, some players leave the market and, accordingly, the oil supplies are reduced [3]. figures 5 and 6 show the results of testing the proposed model on historical data on demand, supply and oil price in the period 2q 2013 4q 2018. source: international energy agency (https://www.iea.org/oilmarketreport/omrpublic). fig. 5. oil demand and supply balance fig. 6. price and balance on the oil market the results show that the oil price obtained using the model is quite consistent with historical data. the deviation of the forecast price from the historical one in some periods can be explained by the increasing role of factors unrelated to the balance of supply and demand in the market. for example, in the period 2q 2013 4q 2018, the oil price was also influenced by foreign factors, mainly related to the events in venezuela and the us sanctions against iran. the model described in this section makes it possible, using data on demand options, production scenarios of nonopec countries, the opec quota formation mechanism, to calculate the oil price dynamics based on model (2). the problem is to choose t opecx in each period t the value at which the oil price will remain in a given target range. the software simulation model (fig. 4) and the optimization and parameter selection methods are implemented using ms excel. 138 v. akinfiev copyright ©2019 assa adv. in systems science and appl. (2019) 4. analysis of the results the oil demand is not elastic, that is, it does not depend on changes in the market price and is set exogenously in the model in the form of a set of scenarios. the main scenario for modeling is the opec forecast of global demand in the period 2019 2028 [10]. table 1 presents the results of calculations of the dynamics of oil production and market share for major producers. it also presents the results of the calculation of the decrease / increase in oil production by opec countries to maintain the price at the target level of usd 65 per barrel. table 1. forecast of the dynamics of oil production and market share in the period 2019 2028 1 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2 101,2 102,2 103,2 104,1 105,0 105,9 106,7 107,6 108,4 109,2 3 7,40 9,20 10,64 11,79 12,71 13,45 14,04 14,51 14,89 15,08 4 7,3% 9,0% 10,3% 11,3% 12,1% 12,7% 13,2% 13,5% 13,7% 13,8% 5 48,4 47,6 47,1 46,9 46,9 47,0 47,3 47,7 48,1 48,7 6 47,8% 46,6% 45,7% 45,1% 44,6% 44,4% 44,3% 44,3% 44,4% 44,6% 7 0,0 -0,8 -0,5 -0,2 0,0 0,2 0,3 0,4 0,4 0,6 8 45,4 45,4 45,4 45,4 45,4 45,4 45,4 45,4 45,4 45,4 in the first column of table 1, the following notation is used: 1 periods, 2 demand, million barrels per day (opec forecast), 3 us shale oil production, million barrels per day, 4 the us shale market share, 5 opec + oil production, million barrels per day, 6 opec + market share, 7 the decrease / increase in oil production by opec (year-over-year), 8 non-opec + production, million barrels per day. in addition, two more scenarios were considered: the opec forecast for growth global oil demand plus 15% and the opec forecast for growth global oil demand minus 15%. the results of the analysis of the market share and currency income of russia from the sale of oil, depending on the target level of the price are presented in table. 2 table 2. market share of russia depending on the target oil price level the target oil price level 40 50 50 60 60 70 70+ opec forecast of growth in world oil demand in 2019-2028 market share by 2028, % 12,0% 11,1% 10,5% 10,1% gross oil revenues of russia, trillion usd 13,338 15,280 17,334 19,443 opec forecast of growth in world oil demand in 2019 2028 plus 15% market share by 2028, % 12,1% 11,2% 10,7% 10,3% gross oil revenues of russia, trillion usd 13,596 15,595 17,706 19,873 opec forecast of growth in world oil demand in 2019 2028 minus 15% market share by 2028, % 11,8% 10,9% 10,4% 10,0% modeling and estimating the impact of the opec agreement 139 copyright ©2019 assa. adv. in systems science and appl. (2019) gross oil revenues of russia, trillion usd 13,083 14,968 16,965 19,017 conclusion the calculations (table 1) confirm that opec +, by reducing production in order to maintain the oil price at a high target level, creates a free niche in the market. this niche is occupied by producers outside the agreement, mainly shale oil producers, who can continue to increase production. with lower target oil prices set by opec +, oil production growth by shale companies will be limited and, under certain growth scenarios for global demand, opec + quotas can be set to increase production, which will help maintain their market share. however, with the current strategy of opec + (the target price level of oil in the range of 6070 usd per barrel), the united states will remain the leader in production growth. calculations carried out on the model show that by 2028, us shale oil production compared to 2018 could almost double and reach 15.1 million barrels per day. us shale oil will occupy 13.8% of the market, while the market share of russian oil will drop to 10.5%. calculations show that for russia, in terms of maintaining market share, the option of targeting oil prices at lower levels is the most profitable, but at the same time, gross oil revenues of russia are noticeably reduced. on the whole, it can be concluded that russia's participation in the opec agreement, as well as the participation of other countries in it, under current conditions is a reasonable strategy. references [1] akinfiev, v. (2017) a model of competition between oil companies with conventional and unconventional oil production. large-scale systems control. issue 67. moscow: ics ras. 52-80, [in russian]. [2] akinfiev, v. (2018) an analysis and forecasting volatility of crude oil market. proceedings of the 11th international conference "management of large-scale system development" (mlsd). moscow: ieee, https://doi.org/10.1109/mlsd.2018.8551767 [3] akinfiev, v. (2018). modeling competition in the oil market. herald of computer and information technologies. № 2. 18-27, [in russian], https://doi.org/10.14489/vkit.2018.02. [4] ansari, d. (2017). opec, saudi arabia, and the shale revolution: insights from equilibrium modeling and oil politics. energy policy, volume 111, december, 166-178. [5] caldara, d., cavallo, m., & iacoviello, m. (2019). oil price elasticities and oil price fluctuations. journal of monetary economics, 103, 1-20. [6] smith, jl. (2018) estimating the future supply of shale oil: a bakken case study energy economics, elsevier energy economics 69, 395–403. [7] olovsson, c. (2019). oil prices in a general equilibrium model with precautionary demand for oil. review of economic dynamics, volume 32, april, 1-17. [8] wang, q., song, x. & li, r. (2018) a novel hybridization of nonlinear grey model and linear arima residual correction for forecasting u.s. shale oil production. energy 165, 1320-1331. [9] salikhov, м. & kurilov, v. (2018) new wave of growth of shale oil production in the usa: scenario forecasts. institute of energy and finance, analytics, october. [in russian]. [online]. available https://fief.ru/img/files/slanceva__otrasl__20181031.pdf [10] world oil outlook 2018. (2018). the annual opec forecast. [online]. available https://woo.opec.org/. https://ieeexplore.ieee.org/author/37086244644 https://doi.org/10.1109/mlsd.2018.8551767 https://www.sciencedirect.com/science/journal/03014215 https://www.sciencedirect.com/science/journal/03014215/111/supp/c https://scholar.google.ru/citations?user=yc2mxdaaaaaj&hl=ru&oi=sra https://www.sciencedirect.com/science/article/pii/s0140988317304188 https://www.fief.ru/employee/detail.2.htm https://www.fief.ru/employee/detail.2.htm adv syst sci appl 2024; 02:203–218 published online at https://ijassa.ipu.ru. solving the energy consumption barrier in brackish water reverse osmosis desalination plants: a genetic algorithm and energy recovery approach moumni mohammed*, massour el aoud mohamed, moumni fatima zahra sultan moulay slimane university, khouribga, morocco abstract: reverse osmosis desalination is an effective technology for supplying potable water to regions facing water stress. however, this process consumes a significant amount of energy, limiting its widespread adoption globally. this study aims to analyze the variations in specific energy consumption (sec) in the reverse osmosis desalination process for brackish water at a plant in morocco, considering the feed water parameters. additionally, the study examines energy consumption with and without the implementation of energy recovery devices at this plant, which produces 10 million cubic meters of water annually. a genetic algorithm is utilized to identify the optimal combination of design and operational parameters to achieve the lowest sec. the findings indicate that incorporating energy recovery devices in the future design of the plant could reduce the sec by up to 30%. keywords: desalination, brackish water, reverse osmosis, energy recovery 1. introduction at present, increases in the global population, coupled with drought and desertifcation due to climate change, will undoubtedly aggravate water security [1]. desalination technologies have been emerged to supply water from unconventional water sources, and reverse osmosis systems account for a large share of desalination facilities [2]. reverse osmosis (ro) process is an important filtration process that is used extensively for the desalination of sea and brackish water all over the world [3]. reverse osmosis is generating growing interest due to its energy efficiency and versatility compared to other desalination technologies. the need to supply drinking water to populations in regions suffering from water shortages has made the development of this technology essential. [4]. reducing the specific energy consumption (sec) poses a significant challenge, driving technological innovation and research in the desalination industry. the energy cost of the desalination process (ro) can account for up to half of the total cost of producing one cubic metre of drinking water. [5]. nevertheless, compared with other desalination technologies, such as multi-stage flash (msf) [7], [8], multiple-effect distillation (med) [9], membrane distillation (md) [ [10], [11], [12], [13],and electrodialysis (ed) [15], the ro process consumes relatively little energy. consequently, most large-scale seawater and brackish water desalination plants have been designed to use the ro method. according to several research studies, the swro process consumes between 2 and 5 kwh/m³ depending on the feed characteristics [24]. on the other hand, the total energy consumption of the bwro process, including electrical energy, is between 1.5 and 2.5 kwh/m³ [20]. many researchers and engineers have dedicated significant effort to finding innovative ways ∗corresponding author: moumni.mohammed@gmail.com 204 m. moumni, m. el aoud mohamed, m. fatima zahra to reduce the energy consumption of the reverse osmosis process. they have focused on several key areas and methods to achieve this objective, including the following: • improvement of ro membranes: developing and utilizing high-performance membranes to enhance filtration efficiency and reduce energy requirements. • reduction of membrane fouling: implementing prevention and cleaning techniques to decrease membrane fouling, maintaining optimal efficiency and lowering energy consumption. • efficiency of energy-consuming equipment: enhancing the efficiency of pumps and other energy-intensive equipment to minimize energy use. • energy recovery devices (erds): using energy recovery devices to capture and reuse energy, thereby reducing the overall energy demand. • innovative membrane module designs: introducing new membrane module configurations to reduce pressure drops in the membrane channel, thus improving energy efficiency. • process optimization: applying optimization algorithms to dynamically adjust operational parameters and maximize energy efficiency in real-time. • integration of renewable energy: utilizing renewable energy sources, such as solar or wind power, to supply desalination plants, reducing reliance on traditional energy sources. first, advances in ro membranes have contributed to energy savings. high-performance reverse osmosis membranes, characterized by increased water flow and salt rejection rates, mitigate the excess pressure required above the osmotic pressure of seawater. researchers have successfully developed an efficient ro system for low sec water desalination by adjusting membrane type and size [25] and [26], respectively. furthermore, low fouling propensity is critical for lowering energy consumption in the swro process because membrane fouling reduces water permeability, which requires higher operating pressure. surface modification has been used to develop fouling-resistant ro membranes by improving hydrophilicity, reducing surface roughness, and decreasing concentration polarization at the membrane’s surface [16] and [17] . moreover, to achieve an optimal combination for a reverse osmosis system with the lowest sec, some researchers are focusing on the design of single-stage or two-stage ro systems. they analyze the performance of these systems with varying feed parameters, membrane permeabilities, and recovery rates. ( [1, 27–30]). two-stage reverse osmosis can achieve a recovery rate of over 50% ( [31]).mingheng [32]. it has been demonstrated that, without inter-stage booster pumps, single-stage is more energy efficient than two-stage in bwro due to lower retentate pressure drop. li et al. [33] investigated the validation of a model-based optimization for bwro operations. in addition, energy-intensive equipment such as high-pressure pumps (hpps), booster pumps (bps), and erds have been developed and improved to reduce the overall energy consumption of the swro process. the efficiency of hpp has increased by about 90%, which is a practical limit for centrifugal pump efficiency. large pumps are recommended for improving hpp efficiency, with pump efficiency in a large-scale desalination plant reaching up to 85% [18], [22]. in other areas, the integration of photovoltaic (pv) and wind power in desalination processes has been evaluated, with promising results in regions with high solar and wind potential. for example, yahiaoui et al. optimized a pv-diesel generator-battery hybrid system for the city of djanet, algeria, using the grey wolf optimizer [23]. in the last 20 years, advancements in erd have played a significant role in reducing the energy consumption of the swro process. the development of erd equipment, including francis turbines (fts) and pressure exchangers (pxs), has significantly reduced energy consumption [18], [19]. the efficiency of px can exceed 95% [18]. as a result, advancements in energy-consuming units have nearly copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 205 reached their practical limits. many research investigations have been conducted to compare erds for saline waters. alexander et al. analyzed the choice between two types of erd (turbocharger and isobaric erd) for brackish water in his study [35], taking into account the conversion rate parameter and the life cycle cost analysis. in addition to its efficiency of between 94 and 98 percent for a discharge rate of up to 110 m3 per hour, which is close to the discharge rate of the plant in this case study, (alexander et al. concluded that the isobaric erd is more energy efficient than the turbocharger, particularly at a conversion rate of 75 to 81 percent, which is also the case for the plant in this study. as a result, the isobaric erd will be considered an energy recovery system for the ro process used in this plant, which lacks erd. furthermore, because the organization intends to expand this plant by adding more ro trains, an analysis of the sec with and without erd will be conducted to justify the feasibility of implementing this erd at this plant. this analysis will determine whether implementing an energy recovery device (erd) in the expanded plant is a viable option. to achieve this, a genetic algorithm will be applied to identify the optimal erd configuration, maximizing energy efficiency while considering the specific parameters of the reverse osmosis process. the genetic algorithm, known for its ability to solve complex optimization problems, will help explore a wide range of possible solutions and converge on the best configuration. 2. description of bwro desalination plant the bwro desalination plant consists of three production lines, as depicted in fig 2.1. each production line is structured in two stages. in this setup, the concentrate from the first stage serves as the feed water for the subsequent stage, allowing for additional permeate production. the permeate collected from the first stage is combined with the permeate from the second stage. the hpp increases the pressure of the pre-treated brackish water to a suitable value for fig. 2.1. schematic diagram the membrane. the pressure required depends on the concentration and temperature of the feed water. osmotic pressure increases with concentration, so the operating pressure must be higher than the osmotic pressure corresponding to the concentration of the rejected brine flow rate at the membrane rack outlet, as well as membrane fouling. this plant’s feed water concentration varies from 1g/l to 3g/l throughout the year, as does its temperature, which ranges from copyright © 2024 assa. adv syst sci appl (2024) 206 m. moumni, m. el aoud mohamed, m. fatima zahra table 2.1. equipment characteristics number of production line 3 number of membrane by line 600 membrane sectional area 35m2 nominal flow of hp pump 550m3/h nominal pressure of hp pump 20 bar nominal speed of hp pump 3000 rpm 7c° to 32c°. to manage the changing characteristics of brackish water and keep within the operating range of this plant, the recommended conversion rate for each production line with its two stages is 75%, with a production flow rate of about 390 m3/h. to ensure that the above instructions are followed, each hp feed pump is equipped with a vfd that allows the pump speed to be varied, which changes both the flow rate and the pressure of the membrane feed water. the article [36] used fuzzy logic to effectively control these parameters in the station object of this study. table 2.1 shows the equipment used in this station’s ro process. 3. materiels and methods before evaluating the specific energy consumption (sec) of the (ro) process implemented in this station, it is essential to first model the ro process and develop a simulation model. this initial step was thoroughly addressed in the work presented in article [37].using matlab/simulink software tools, a numerical model was developed and validated against real plant values, as well as against the model used by arun-joseph and vasanthi damodaran in table 3.1. ro process equations definition equation osmotic perssure (bar) π = i.cf .r.t (3.1) permeate flux (m.s−1) jw = aw(∆p −∆π) (3.2) solvent permeability (m.s−1.pa−1) aw = a0exp(6.433− 1885 t ) (3.3) permeat flow (m3/h) qp = s.jw (3.4) drop pressure across membrane (bar) ∆p = ( pf + pc 2 − pp) (3.5) ∆p = ∆π + ( qp s.aw ) (3.6) conversion rate y = qp qf (3.7) their study [38]. the equations used for this modeling are summarized in table 3.1. with this model in place, the current sec of the ro process can be evaluated both with and without an energy recovery device (erd). this comparison will demonstrate the potential benefits of installing an erd in the plant. following the conclusion to install the erd, it is necessary to determine the optimal parameters for the ro process, taking into account the parameters of copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 207 the erd, which will also be modeled using equations. this optimization will be carried out using a genetic algorithm to ensure the most efficient operation of the expanded plant. 4. evaluation of sec at bwro electricity consumption is a major obstacle in the development of the ro process, with highpressure pumps (hpps) being the primary energy consumers. these pumps need substantial energy to overcome the osmotic pressure of saline water. optimizing energy use involves lowering the feed’s osmotic pressure and selecting the most efficient pump. hpps account for about 75% of total specific energy consumption, with the remainder coming from membrane operations [26]. this study, therefore, focuses on the energy consumption of hpps in the ro process for freshwater production. the equation below shows a mathematical method for estimating the sec of the ro membrane without the need for an erd system: es = pf .qf qp.η (4.1) the model generated in matlab/simulink to evaluate the fluctuation of sec in the station under study can be used to get an understanding of the predicted energy consumption and losses at the discharge. three additional scenarios are added to the three cases given in the article [37] to study the energy fluctuation, as shown in table 4.1. furthermore, the energy assumed to be lost during discharge can be calculated by adding the flow rate and pressure of the concentrates. given that the organization intends to expand this station with more trains, and based on the results, which show that the energy lost to create a cubic meter of permeate can average 0.3 kwh, a erd for the new entity must be installed. 5. modeling erd as previously mentioned, an isobaric energy recovery device (erd) will be utilized for the ro process implemented in this plant. rotary isobaric devices use a small rotor to recover hydraulic energy from the concentrate stream. the rotor contains ducts that alternately fill with high-pressure brine and low-pressure feed water. as the rotor turns, these ducts are exposed to highand low-pressure zones alternately, effectively replacing the high-pressure brine with saltwater on a 1-to-1 basis. the timing of the water exchange ensures that the chamber is never completely empty, creating a static water piston that prevents the two streams from mixing. [39]. the equations presented below are used to determine the best design for the ro process with an erd. the installation of an erd in an existing ro process necessitates knowledge of its operating parameters. figure 5.1 shows the erd system’s inputs and outputs. fig. 5.1. the erd system’s inputs and outputs copyright © 2024 assa. adv syst sci appl (2024) 208 m. moumni, m. el aoud mohamed, m. fatima zahra 5.1. flow calculation the erd requires for its adequate choice the preknowledge of the inlet flows(qic, qif ) and outlet flows (qoc,qof ) of this equipment. water flows from the concentrate of ro process and the low pressure pumps are the two types of water flows that enter the erd. qic can be determined as is defined, the concentrate flow rate, qc which is a function of the ro process feed flow qf and its conversion rate y : qic = qf .(1− y ) where : qf = qp y = qc +qp (5.1) each erd is is characterised by brine flow loss or leak (l) and overflush (of) [34] where : qif , qof and qoc can be defined using the following equations: flow balance: qic +qif = qof +qoc (5.2) overflush (of): overflush range is provided by isobaric erds manufactures: of = qif −qof qof (5.3) table 4.1. simulation values and actual values comparaison cnd unit case1 case2 case3 t c° 20 13,5 24 cf g/l 2 1.6 1.26 pa st* mb** er % st mb er % st mb er % pf bar 15,9 15,2 4,4 18,1 17,8 1,65 12,9 14,1 9,3 pc1 bar 13,9 13,2 5,03 15,2 15,7 3,28 10,5 12,1 15,2 pc2 bar 11,9 11,2 5,88 12,9 13,7 6,2 8,95 10,1 12,84 qp1 m3/h 305,4 299,04 2,1 296,9 323 8,7 298,9 287,3 3,88 qp2 m3/h 76,05 74,4 2,16 91,3 80,4 11,9 80 71,5 10,6 qc1 m3/h 230,45 225,6 2,1 215 243 13 226,6 216,6 4,4 qc2 m3/h 154,4 151,15 2,1 140,2 163,1 16,33 146,6 145.1 1 cp1 g/l 0,04 0,02 50 0,01 0,01 0 0,02 0,02 0 cp2 g/l 0,09 0,06 33,3 0,02 0,03 50 0,06 0,04 33,3 y % 71,2 72,7 2,1 73,5 71,1 3,2 72,1 71,2 1,2 es kwh/m3 0,8 0,76 5 0,85 0,89 4,4 0.68 0.71 4,4 ec kwh/m3 0,17 0,15 7,5 0,16 0,2 23,5 0,15 0,16 6 cnd unit case4 case5 case6 t c° 24 10 12 cf g/l 1 2 1.2 pa st* mb** er % st mb er % st mb er % pf bar 13,4 13,95 4,1 19,1 19,55 2,35 17,8 18,09 1,6 pc1 bar 11,6 11,95 3 17,7 17,55 1,1 15,9 16,09 1,1 pc2 bar 10,2 9,95 2,4 15,9 15,5 2,5 14,3 14,09 1,4 qp1 m3/h 291,04 285,62 1,8 345,2 338,1 2 330,2 325,28 1,4 qp2 m3/h 72,05 71,11 1,3 96,3 84,17 12,5 82,1 80,98 1,3 qc1 m3/h 220,45 215,47 2,2 257,2 255,07 0,8 243,6 245,39 0,7 qc2 m3/h 149,4 144,36 3,3 173,2 170,9 1,3 167,6 164,41 1,9 cp1 g/l 0,04 0,02 50 0,05 0,04 20 0,02 0,02 0 cp2 g/l 0,07 0,05 28,5 0,11 0,09 18,1 0,07 0,06 14,2 y % 71,19 72,1 1,2 71,1 71,1 0 72,1 71,2 1,2 es kwh/m3 0,66 0,69 4,5 0,92 0,98 6,5 0,88 0,91 3,4 ec kwh/m3 0.15 0,13 13,3 0,25 0,24 4 0,22 0,21 3 copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 209 brine flow loss or leak (l): is provided by isobaric erds manufactures as a function of temperature and brine input flow per isobaric erds unit [40]: l = qic −qof qic (5.4) therefore the output flow rate from the side of feed membranes qof can be expressed as : qof = qic.(1− l) (5.5) and qif : qif = (of + 1).qof (5.6) finally by using flow balance: qic +qif = qoc +qof (5.7) the output flow rate from erd to the rejection can be expressed as : qoc = qic +qif −qof (5.8) 5.2. salinity calculation after defining the inlet and outlet flows of the erd system, the salinity of the flows in the erd sides can be determined. cicis the salinity of the concentrate of the ro process cc and cif is the salinity of feed water with low-pressure flow which is equal to the salinity cf . it remains to determine the concentration of outflows from the erd system coc and cof which are respectively the salinity of reject flow and the salinity of pressured flow by the erd. each isobaric erd has its own mixing mx ratio [41] which is the ratio of the volume of brine that transfers into a volume feed water and can be calculated with the following equation independent of the pressure exchanger high and low pressure flow balance: mx = cof − cif cic − cif (5.9) using the equation (5.9) cof can be expressed as a function of the previously known parameters: cof = mx(cic − cif ) + cif (5.10) according to the salinity balance : qic.cic +qif .cif = qoc.coc +qof .cof (5.11) coc can be expressed as: cic = (qic.cic +qif .cif )−qof qoc (5.12) 5.3. pressure calculation the erd’s input pressures pic and pif are respectively, the one of the from the concentrate flow rate of the ro process pc, and the second is the pressure of low pressure flow rate coming from the low pressure pump, which is generally the circulation water pressure taken at approximately 2 bar. moreover, the erd outlet pressure may be estimated as follows: the output pressure recovered by erd pof : pof = pic −∆p1 (5.13) the pressure of the discharge from the erd poc: poc = pif −∆p2 (5.14) where : ∆p1 and ∆p2 are the losses across the isobaric erd. copyright © 2024 assa. adv syst sci appl (2024) 210 m. moumni, m. el aoud mohamed, m. fatima zahra 6. booster pump a multi-stage brackish system without an interstage boost can be designed in a similar manner to a single-stage system. in this scenario, the previous stage’s concentrate is used to pressurize a feedwater stream for the first stage. the circulation pump compensates for pressure losses in the membrane stages, pipework, and erd. because the brackish and seawater ro processes are not the same, the implementation of isobaric erds in brackish water ro systems must be handled differently. the small amount of pressure loss caused by membranes, friction in the erd, and the piping circuit necessitates the use of a booster pump in isobaric erd. this pump is used as the erd’s output in single-stage seawater systems. the erd booster pump, on the other hand, can play two important roles in a two-stage brackish water system by being installed between stages one and two. the erd booster pump acts as an inter-stage booster pump in this configuration, lowering the required pressure from the main high-pressure feed pump and balancing flux between stages 1 and 2 [35]. in this study, the booster pump will function as a circulation pump to recover pressure lost in the circuit (membranes, erd, and pipes). 7. the sec of bwro with erd in order to evaluate the sec by a bwro equiped with a erd, firstly it is necessary to express the energy recovered in the concentrate part of the process, which is function of the ro brine pressure pc , flow rate of the concentrate qc, and the erd efficiency ηerd as shown : eerd = pc.qc.ηerd (7.1) by introducing a booster pump (bp) with efficiency ηbp , the energy consumed by this pump must be considered in order to evaluate the energy saving envisaged in the new design. in fact, this pump takes the water flow qof coming from the erd with the pressure pof in order to achieve the required pressure at the ro process inletpf , therfore, the energy consumed by(bop) can be expressed as follows: ebp = (pf − pof )qof ηbp (7.2) the pressure provided by the booster pump pbp : pbp = pf − pof (7.3) the sec by the bwro with a erd will be the combined energy consumed by the hp pump and the booster pump [?]: sec = (qf −qof ).pf ηhp .qp + pbp .qof ηbp .qp (7.4) 8. results and discussions the matlab/simulink model developed in the article [37] and validated by the real values of the studied station makes it possible to evaluate the sec by the process ro and especially the hp pumps. furthermore, the values considered in the above study are taken at the start of exploitation of the said plant. however, after two years of operation, it was discovered that feed pressure values had increased for the same feedwater properties (salinity and temperature). this is due to membrane fouling and its life span. following the development copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 211 of the operating parameters of the ro process in this plant, it was observed that feed pressure and concentrate pressure rise by 2 to 3 bars. as a result, the sec increases, proving the significance of the current study. using the erd modeling equations and the ro model developed in matlab/simulink, the erd has been implemented into this model while taking into account the inclusion of a booster pump, which is essential for an ro process with an erd. the proposed design is illustrated in figure 8.3. then, the specific energy consumption (sec) of the three reverse osmosis (ro) trains implemented in the plant under study is evaluated by considering various states of the saltwater feed, specifically by varying its salinity and temperature. these two parameters have a direct impact on the sec of an ro process, as indicated in several studies ( [6,20,21]), regardless of the system’s design and structure. the graphs below illustrate the evolution of sec for the ro process in this plant with and without energy recovery device (erd). when the feed water salinity is set at 1.85 g/l and its temperature is varied within the recorded range at the station, the sec of the ro process with erd decreases from 2.1 kwh/m³ to 1.45 kwh/m³ as the feed water temperature increases (figure 8.1). in contrast, the sec for the current design without erd decreases from 3.05 kwh/m³ to 2.15 kwh/m³ under the same conditions (figure 8.2). in the same way, the feed water temperature was set to 22°c, and the salinity was adjusted within the station’s operating range. the sec with an erd ranges from 1.55 kwh/m3 to 1.72 kwh/m3 (figure 8.5), while the sec without an erd ranges from 2.27 kwh/m3 to 2.53 kwh/m3. implementing the erd at this station can reduce the sec by an average of 30%, even with the installation of a booster pump, according to the study’s findings. in addition, the sec formula (7.4) takes into account the energy used by this pump. to evaluate the sec fig. 8.1. sec with t(with erd) fig. 8.2. sec with t(without erd) table 8.1. algorithm genetic parameters parameter number of variables population size mutation migration crossover iteration value 6 50 for five 0,71 0,41 0,7 155 definition of variables xi x1 x2 x3 x4 x5 x6 variable y l qif of qf qp lower value lb 0.7 0.003 136 0 490 370 upper value ub 0.75 0.013 162 0.05 550 400 copyright © 2024 assa. adv syst sci appl (2024) 212 m. moumni, m. el aoud mohamed, m. fatima zahra fig. 8.3. the proposed design provided by the ro process in the three production lines, several salt water scenarios that closely resemble reality were captured and then simulated on the matlab platform. it is clear that the inclusion of an erd has an important effect on energy saving, and according to the 3d multivariate representation (figure 8.6 and figure 8.7), it can be observed that the temperature variation has a significant effect on the sec more than the salinity variation fig. 8.4. sec with cf (without erd) fig. 8.5. sec with cf (with erd) table 8.2. results of genetic algorithm pressure (bar) 10 14 17 20 22 25 sec(kwh/m3) 1,07 1,45 1,73 2,01 2,29 2,48 y 0,7 0,7 0,7 0,7 0,7 0,7 l 0,03 0,03 0,03 0,03 0,03 0,03 qif (m3/h) 136,64 136,92 136,16 136,66 136,99 136,65 of (m3/h) 0,04 0,04 0,04 0,04 0,04 0,04 qf (m3/h) 549,54 549,88 549,85 549,45 549,99 549,58 qp(m3/h) 412,99 412,95 412,68 412,78 412,99 412,93 copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 213 fig. 8.6. sec without erd fig. 8.7. sec with erd fig. 8.8. sec compared to pf fig. 8.9. energy saving and qp because its range of variation in this station is limited. as concluded in various studies, including the article [37], changes in saltwater parameters (temperature and salinity) cause variations in the feed pressure. the relationship between sec and feed pressure is illustrated in figure 8.8, which shows that sec increases with rising feed pressure. figure 8.9 depicts the energy savings, defined as the difference in sec between the ro process with erd and the process without erd. the curve indicates that energy savings increase with the produced flow (permeate flow). the reduction rate of sec with the erd remains approximately 30%, regardless of changes in permeate flow. therefore, deploying the erd is beneficial for ro stations with high energy consumption, especially for saline water within the same variable range of salinity and temperature. this advantage can be further supported by a techno-economic analysis. by incorporating the genetic algorithm-based optimization approach, significant improvements in energy efficiency can be achieved for the ro process in water desalination plants. the observed 30% reduction in the sec by adding an erd in the ro process highlights the potential benefits of using this approach. additionally, the genetic algorithmbased optimization can identify the most suitable erd configuration based on the specific parameters of the ro process, leading to further improvement in energy efficiency. the copyright © 2024 assa. adv syst sci appl (2024) 214 m. moumni, m. el aoud mohamed, m. fatima zahra fig. 8.10. best individual combination of erd deployment and genetic algorithm-based optimization could provide a more sustainable and cost-effective approach to water desalination. in order to optimize the specific energy consumption (sec) of the ro process using a genetic algorithm, the formula (7.4)for sec it’s developed as a function of the parameters that contribute to the choice of the erd(overflush: of and leak: l). the equation can be expressed as follows: sec = (1− (1− y )(1− l)).pf .qf ηhp .qp + pbp .qif (of + 1).ηbp .qp (8.1) the optimal values of the sec were obtained by using the genetic algorithm on the given formula, considering the feed pressure values ranging from 7 to 25 bar. it is important to note that the feed pressure values were calculated based on the salinity and temperature of the water. the parameters of the applied genetic algorithm are detailed in table 8.1 and the obtained optimization results are presented in table 8.2. upon analyzing the values obtained through the genetic algorithm, it can be concluded that the specific energy consumption (sec) of the implemented ro process in the studied station ranges from 0,73 to 2,3. this means that, at the maximum feed pressure values considered in this station, energy consumption can be reduced by over 40%. table 8.2 summarizes the optimal values of the variables xi required for the operation of the ro process in this station, taking into account the parameters that characterize the erd. for the process to function optimally, an appropriate erd should have an overflush value of 0,04 and a leak value of 0,03. the figure 8.10 shows that variables x3,x5 and x6 are the parameters that have the most influence on energy optimization in this ro process. 9. conclusions the analysis of specific energy consumption (sec) in the moroccan plant’s reverse osmosis desalination process for brackish water reveals significant potential for energy savings. the study shows that by using a genetic algorithm to optimize design and operational parameters, the incorporation of energy recovery devices can reduce sec by up to 30%. these findings highlight the significance of energy optimization in increasing the feasibility and sustainability of reverse osmosis desalination. implementing these technologies can make the process more cost-effective and environmentally friendly, encouraging widespread copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 215 adoption in water-scarce regions. future research will focus on implementing intelligent control systems for the operating parameters to ensure consistent production under optimal conditions. this approach aims to maintain the desired output quality while adapting to varying feed water characteristics, thus ensuring the process remains efficient and reliable. 10. abbreviations [erd] energy recoverey device [ro] reverse osmosis [hp] high pressure [bw] brackish water [bwro] brackish water reverse osmosis [swro] sea water reverse osmosis [hpp] high pressure pump [bp] booster pump [lpp] low pressure pump [vfd] variable frequency drive [sec] specific energy consumed (kwh/m3) [pf ] feed pressure (bar) [pc] concentrate pressure (bar) [pp] permeat pressure (bar) [qf ] feed flow rate (m3/h) [qc] concentrate flow rate (m3/h) [qp] permeat flow rate (m3/h) [cf ] feed water salinity (g/l) [cc] concentrate salinity (g/l) [cp] permeat salinity (g/l) [t ] temperature (c°) [es] specific energy consumed without erd (kwh/m3) [ec] energy lost on the rejection (kwh) [eerd] energy recovered by erd (kwh) [ebp ] energy consumed by bp (kwh) [ηhp ] hpp efficiency [ηbp ] bp efficiency [ηerd] erd efficiency [pif ] erd input pressure from lpp (bar) [pic ] erd input pressure from the concentrate (bar) [pof ] erd output pressure from erd to bp (bar) [poc] erd output pressure from erd to the rejection (bar) [qif ] erd input flow rate from lpp (m3/h) [qic] erd input flow rate from the concentrate (m3/h) [qof ] erd output flow rate from erd to bp (m3/h) [qoc] erd output flow rate from erd to the rejection (m3/h) [cif ] erd input salinity of flow rate from lpp (g/l) [cic] erd input salinity of flow rate from the concentrate (g/l) [cof ] erd output salinity of flow rate from erd to bp(g/l) [coc] erd output salinity of flow rate from erd to the rejection (g/l) [l] erd flow loss [of ] erd over flush range copyright © 2024 assa. adv syst sci appl (2024) 216 m. moumni, m. el aoud mohamed, m. fatima zahra [mx] mixing ratio [∗st ] station values [∗ ∗mb] matlab/simulink values [er] margin of error between st and mb [i] number of ions dissociated in the case of an electrolyte [r] ideal gas constant r = 8.314 (j .mol−1.k−1) [s] surface of the membrane (m2) references 1. shah abedi, m., hashemi, s. h. & fazeli, m. (2022). feasibility of increasing water recovery of inland reverse osmosis systems and the use of reject brine, arab j sci eng, 47, 6525–6534. doi: 10.1007/s13369-021-06451-4 2. eltamaly, a. m., ali, e., bumazza, m., et al. (2021). optimal design of hybrid renewable energy system for a reverse osmosis desalination system in arar, saudi arabia, arab j sci eng, 46, 9879–9897. doi: 10.1007/s13369-021-05645-0 3. emad, a., ajbar, a. & almutaz, i. (2011). periodic control of a reverse osmosis desalination process, journal of process control, 22, 218–227. doi: 10.1016/j.jprocont.2011.09.001 4. bartman, a. r., zhu, a., christofides, p. d. & cohen, y. (2011). minimizing energy consumption in reverse osmosis membrane desalination using optimization-based control, journal of process control, 20, 1261–1269. doi: 10.1016/j.jprocont.2010.09.004 5. adda, a., naceur, w.m. & abbas, m. (2016). modelisation et optimisation de la consommation d’energie d’une station de dessalement par procede d’osmose inverse en algerie, revue des energies renouvelables, 19(2), 157–64. 6. al-obaidi, m. a., alsarayreh, a. a., al-hroub, a. m., alsadaie, s. & mujtaba, i. m. (2018). performance analysis of a medium-sized industrial reverse osmosis brackish water desalination plant, desalination, 443, 272–284. doi: 10.1016/j.desal.2018.06.010. 7. al-karaghouli, a. & kazmerski, l. l. (2013). energy consumption and water production cost of conventional and renewable-energy-powered desalination processes, renew. sust. energ. rev., 24, 343–356. 8. semiat, r. (2008). energy issues in desalination processes, environ. sci. technol., 42, 8193–8201. 9. miller, s., shemer, h. & semiat, r. (2015). energy and environmental issues in desalination, desalination, 366, 2–8. 10. al-obaidani, s., curcio, e., macedonio, f., di profio, g., al-hinai, h., et al. (2008). potential of membrane distillation in seawater desalination: thermal efficiency, sensitivity study and cost estimation, j. membr. sci., 323, 85–98. 11. jeong, k., park, m., ki, s. j. & kim, j. h. (2017). a systematic optimization of internally staged design (isd) for a full-scale reverse osmosis process, j. membr. sci., 540, 285–296. 12. khayet, m. (2013). solar desalination by membrane distillation: dispersion in energy consumption analysis and water production costs (a review), desalination, 308, 89–101. 13. kim, j., kim, j. & hong, s. (2018). recovery of water and minerals from shale gas produced water by membrane distillation crystallization, water res., 129, 447–459. 14. park, k., kim, d. y. & yang, d. r. (2017). theoretical analysis of pressure retarded membrane distillation (prmd) process for simultaneous production of water and electricity, ind. eng. chem. res., 56, 14888–14901. 15. al-karaghouli, a. & kazmerski, l. l. (2013). energy consumption and water production cost of conventional and renewable-energy-powered desalination processes, renew. sust. energ. rev., 24, 343–356. copyright © 2024 assa. adv syst sci appl (2024) solving the energy consumption barrier... 217 16. miller, d. j., dreyer, d. r., bielawski, c. w., paul, d. r. & freeman, b. d (2017). surface modification of water purification membranes, angew. chem. int. ed., 56, 4662– 4711. 17. guha, r., xiong, b., geitner, m., moore, t., wood t.k., et al (2017). reactive micromixing eliminates fouling and concentration polarization in reverse osmosis membranes, j. membr. sci., 542, 8–17. 18. voutchkov, n. (2018). energy use for membrane seawater desalination–current status and trends, desalination, 431, 2–14. 19. qureshi, b. a. & zubair, s. m. (2016). energy-exergy analysis of seawater reverse osmosis plants, desalination, 385, 138–147. 20. alsarayreh, a. a., al-obaidi, m. a., al-hroub, a. m., patel, r. & mujtaba, i. m. (2019). evaluation and minimisation of energy consumption in a medium-scale reverse osmosis brackish water desalination plant, journal of cleaner production, 248, 119220. doi: 10.1016/j.jclepro.2019.119220 21. anqi, a. e., alkhamis, n. & oztekin, a. (2015). numerical simulation of brackish water desalination by a reverse osmosis membrane, desalination, 369, 156–164. doi: 10.1016/j.desal.2015.05.007 22. ruiz-garcia, a., nuez, i., carrascosa-chisvert, m. d. & santana, j. j. (2020).simulations of bwro systems under different feedwater characteristics. analysis of operation windows and optimal operating points, desalination, 491, 114582. doi: 10.1016/j.desal.2020.114582 23. eltamaly, a. m., ali, e., bumazza, m., et al. (2021). optimal design of hybrid renewable energy system for a reverse osmosis desalination system in arar, saudi arabia, arab j sci eng, 46, 9879–9897. doi: 10.1007/s13369-021-05645-0 24. xevgenos, d., moustakas, k., malamis, d. & loizidou, m. (2016). an overview on desalination and sustainability: renewable energy-driven desalination and brine management, desalination and water treatment, 57(5), 2304–2314. doi: 10.1080/19443994.2014.984927 25. fethi, k. (2003). optimization of energy consumption in the 3300 m3/d ro kerkennah plant, desalination, 157(1–3), 145–149. doi: 10.1016/s0011-9164(03)00394-1. 26. al-obaidi, m. a., kara-zaitri, c. & mujtaba, i. m. (2018). significant energy savings by optimising membrane design in the multi-stage reverse osmosis wastewater treatment process, environmental science: water research and technology, 4(3), 449–460. doi: 10.1039/c7ew00455a 27. alanood, a. alsarayreh, a., mudhar, a., al-obaidi, b., shekhah k., et al. (2021). performance evaluation of a medium-scale industrial reverse osmosis brackish water desalination plant with different brands of membranes. a simulation study, desalination, 503, 114927. doi: 10.1016/j.desal.2020.114927. 28. qiu, t. & davies, p.a. (2012).comparison of configurations for high-recovery inland desalination systems, water, 4(3), 690–706. doi: 10.3390/w4030690 29. wei, q. j., mcgovern, r. k. & lienhard, j. h. v. (2017). saving energy with an optimized two-stage reverse osmosis system, environmental science: water research & technology, 4. doi: 10.1039/c7ew00069c 30. al-huwaidi, j. s., al-obaidi, m., jarullah, a. t., kara-zaı̈tri, c. & mujtaba, i. m. (2021). modeling and simulation of a hybrid system of trickle bed reactor and multistage reverse osmosis process for the removal of phenol from wastewater, computers and chemical engineering, 153, 107452. doi: 10.1016/j.compchemeng.2021.107452 31. kim, j., park, k. (2020). optimization of two-stage seawater reverse osmosis membrane processes with practical design aspects for improving energy efficiency, journal of membrane science, 601, 117889. doi: 10.1016/j.memsci.2020.117889 32. li., m. (2020). optimization and plant validation of bwro operation, in analysis and design of membrane processes: a systems approach. melville, ny: aip publishing. doi: 10.1063/9780735421790-005 copyright © 2024 assa. adv syst sci appl (2024) 218 m. moumni, m. el aoud mohamed, m. fatima zahra 33. li, m. & noh, b. (2012). validation of model-based optimization of brackish water reverse osmosis (bwro) plant operation, desalination, 304, 20–24. doi: 10.1016/j.desal.2012.07.029 34. sanchez, j. m. s., castillo, n. s. & castillo, r. s. (2007). mathematical model for isobaric energy recovery devices, proc. of ida world congress-maspalomas (gran canaria, spain), 21–26. 35. drak, a., adato, m. (2014). energy recovery consideration in brackish water desalination, desalination, 339, 34–39. doi: 10.1016/j.desal.2014.02.008 36. moumni, m. & massour, m. (2022). fuzzy logic control of a brackish water reverse osmosis desalination process, computers and chemical engineering, 167, 108026. doi: 10.1016/j.compchemeng.2022.108026. 37. moumni, m. & massour, m. (2021). modeling of reverse osmosis process at a brackish water desalination station, proc. of 7th international conference on optimization and applications (icoa) (wolfenbüttel, germany). doi: 10.1109/icoa51614.2021.9442632. 38. arun, j., vasanthi, d. (2019).dynamic simulation of the reverse osmosis process for seawater using labview and an analysis of the process performance, computers and chemical engineering, 121, 294–305. doi: 10.1016/j.compchemeng.2018.11.001 39. lance, r., pinto, l. & pinto, j. m. (2015). energy recovery in desalination: returning alternative water supplies to consideration, florida water resources journal. 40. heinz, l. (2022). reverse osmosis seawater desalination volume 2: planning, process design and engineering. a manual for study and practice. berlin, germany: springer nature. 41. huang, b., pu, k., wu, p., wu, d. & leng, j. (2020). design, selection and application of energy recovery device in seawater desalination: a review, energies, 13, 4150. doi:10.3390/en13164150 copyright © 2024 assa. adv syst sci appl (2024) introduction description of bwro desalination plant materiels and methods evaluation of sec at bwro modeling erd flow calculation salinity calculation pressure calculation booster pump the sec of bwro with erd results and discussions conclusions abbreviations microsoft word 1077 article text, copyedited.docx adv syst sci appl 2021; 02; 117-132 published online at https://ijassa.ipu.ru. mathematical modeling of color printing devices: color gamut visualization and colorimetric measurements regularization dmitry v. tunitsky* v.a. trapeznikov institute of control sciences, russian academy of sciences, moscow, russia e-mail: dtunitsky@yahoo.com abstract: the article concerns a piecewise linear modeling of a vast range of color printing devices: various printers of different kind and nature, presses, etc. the central purpose is not a presentation of ready to use computational algorithms, or, moreover, accomplished software solutions. the point is to reveal some aspects of mathematical modeling, which could serve as both a guideline for creation of practical and robust engineer solutions and an example of nontrivial application of piecewise linear topology. keywords: colorant space, reference space, piecewise linear model, singular faces, proper printing device, non-degenerate printing device, regular printing device 1. preliminaries consider a stimuli space ℝ! of m colorants and a reference space ℝ" of n colors. in color management practice, the dimension of reference space n is always equal to three, n = 3, since any color can be described exactly by three real numbers who are the coordinates of this color in some color space, e.g., xyz, lab, etc. the dimension m of stimuli space can vary in accordance with the number of colorants the considered printing device uses. this dimension is called a dimension of a printing device. in most practical cases the dimension of printing device is either three, m = 3, or four, m = 4. in the first case we deal with a three-dimensional printing device and in the second case we deal with a four-dimensional printing device. not all the combinations (x1,…,xm) of real numbers represent combination of colorants, because the amount xk, 1£k£m, of the k-th colorant is measured in percents and can vary from zero percents (do not put the colorant at all) to one hundred percents (put as much colorant as possible). a printing device renders a combination (x1,…,xm) of colorants into the corresponding color (y1, y2, y3). it means that a printing device can be described by a map f: wm ® ℝ#, where wm = {( x1, … , xm), 0£ x1, … , xm£100} is the m-dimensional colorants cube in the stimuli space ℝ! and f(x1, … , xm) = (y1, y2, y3) is the point, whose coordinates correspond to the color rendered by mixture of the colorants combination. we will call such a map a printing device. to find the value of this function we should with the printing device under consideration reproduce on a sheet of paper the patch formed by the colorants combination (x1,…,xm) and measure the color coordinates (y1, y2, y3) of this patch by a colorimetric measurement tool. the number of colorants combinations (x1,…,xm) in the colorants cube is infinite and of cause it’s impossible to print and measure all of them, since actually we can print and measure only the finite subset of such combinations. consider a finite set {wi} ì wm, i.e., a * corresponding author: dtunitsky@yahoo.com mw 118 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) mesh of fixed points wi, i=1,…,m, inside the colorants cube and a set {pi} î ℝ# = {( y1, y2, y3), ¥ < y1, y2, y3 <+ ¥} of the corresponding values pi, f(wi) = pi, i=1,…,m, in three-dimensional color space. call this mesh function a measurement data for the printing device under consideration. hence, a measurement data is a discrete map f: {wi} ® {pi} such that f(wi) = pi = f(wi) for i=1,…,m. for simplicity, we will restrict ourselves to the case of a regular mesh only. recall its definition. definition. let wm = [a1, b1] ´ … ´ [am, bm] be an m-dimensional rectangular parallelepiped with a1 < b1, …, am < bm. for k = 1,…,m consider finite sets zk = {xk0 , …, xkm(k)}, ak = xk0 < … < xkm(k) = bk, of m(k)+1 real numbers. the product mesh {wi} = z1 ´ … ´ zm ì wm of m points, m = ( m(1)+1)…(m(m)+1), is called regular inside the m-dimensional parallelepiped . there are two principal problems of mathematical modeling of a printing device. direct problem. to find a continuous map f: wm ® ℝ# being a satisfactory approximation to a given discrete function f, i.e., to the measurement data of the printing device under consideration. we will call the solution of this problem a model of the printing device or for conciseness just a printing device described by the corresponding measurement data (see section 1). inverse problem. to find a continuous map g: f(wm) ® wm, being an inverse map to f, i.e., the composition of the maps g and f should be the identical map of the set f(wm), . the problem, naturally, involves description of the gamut, i.e., the image f(wm) to the map f. moreover, in case m>3 an inverse map should satisfy to some additional conditions, for example, the maximal (minimal) possible amount of the m-th colorant component xm. we will restrict our examination of these problems to the class of piecewise linear maps. remind the necessary notions. a k-dimensional simplex is a convex hull of k+1 affinely independent points of an m-dimensional space ℝ! , where 0£k£m. the boundary of a kdimensional simplex consists of faces with different dimensions: k-1, k-2, …, 0. onedimensional faces are called edges, and zero-dimensional faces are called vertices. definition. suppose the colorants cube wm is decomposed into a union of n, n>0, sets dj, . this decomposition is called simplex if for j=1,…, n the set dj, is an m-dimensional simplex and the intersection of any two simplexes dj and dk is either empty, dj ç dk = æ, or is a union of whole (m–1)-, (m–2)-, …, and 0-dimentional faces to these simplexes. definition. a continuous map f: wm ® ℝ#, is called piecewise linear if there exists a simplex decomposition of the mdimensional colorant cube such that all the restrictions f½dj : dj ® ℝ# of the map f to tetrahedrons dj are linear maps. in other words, f½dj (x) = cj + bjx, where x = ( x1, … , xm)t mw mw )( mwfidgf =! ! nj j mw ,...,1= d= ! nj j mw ,...,1= d= mw mathematical modeling of color printing devices: color gamut visualization… 119 copyright ©2021 assa. adv. in systems science and appl. (2021) is an m-dimensional vector of colorants combination, bj is a 3´m matrix, and cj is a threedimensional vector, cj î ℝ#, for j=1,…, n (cf. [1, sections 1.4 and 2.3]. remark. generally speaking, for real applications the formulated above notions of simplex decomposition of the colorants cube wm and its piecewise linear map is not sufficient. indeed, rather often a colorants limitations, for example, , where , should be applied to the colorants cube . such limitations have practical sense. in particular, they allow to avoid putting too much of colorants to a sheet of paper. moreover, sometimes several color limitations should be applied to the colorants cube. as a result in general case we will get not a cube decomposed into simplexes but a convex polyhedron decomposed into intact and truncated simplexes. nevertheless all the constructions below actually are through in this general case. therefore the reader, especially the one who is mostly interested in ideas rather than in technical details, can easily content himself with the case of cubes and simplexes only because in general case all the same ideas and methods are practiced. 2. direct problem for three-dimensional printing devices firstly examine a problem of approximation for a three-dimensional printing device. let a finite set {wi} ì w3 of points wi, i=1,…, m, inside the colorants cube w3 be a regular mesh. consider a discrete map f: {wi} ® {pi}, of measurement data, where pi = f(wi) = f(wi) for i=1,…, n. to approximate the given discrete map f by a continuous one f: w3 ® ℝ#, use a piecewise linear or, what is the same in three-dimensional case, a tetrahedral interpolation. we will mostly follow the text-book [2] in our constructions. by definition of a regular mesh, for k=1,2,3 there exist the one-dimensional meshes zk = {xk0, … , xkm(k)}, ak = xk0 < … < xkm(k) = bk, of m(k)+1 real numbers such that {wi} = z1 ´ z2 ´ z3 ì w3 and m = (m(1)+1)(m(2)+1)(m(3)+1). it means that the three-dimensional colorants cube w3 can be decomposed into the union of the mesh parallelepiped cells pi,j,k = [x1i-1, x1i] ´ [x2j-1, x2j] ´ [x3k-1, x3k], i=1,…,m(1), j=1,…,m(2), k=1,…,m(3). inside each of these parallelepiped cells the continuous approximation f of the measurement discrete map f is constructed by the following way. consider an arbitrary three-dimensional rectangular parallelepiped p = [a1, b1] ´ [a2, b2] ´ [a3, b3] = {(x1, x2, x3), a1£ x1£ b1, a2£ x2£ b2, a3£ x3£ b3}, where a1, b1, a2, b2, a3, and b3 are fixed real numbers. there is an obvious one-to-one correspondence of the 8 vertices to the rectangular parallelepiped p and the 8 vertices (0,0,0), (0,0,1), (0,1,0), (0,1,1), (1,0,0), (1,0,1), (1,1,0), (1,1,1) to the unit three-dimensional cube p1 = {(x1, x2, x3), 0£ x1£ 1, 0£ x2£1, 0£ x3£ 1}. cx m i i £å =1 mc 100< mw ! )3(,...,1 )2(,...,1 )1(,...,1 ,, 3 mk mj mi kjiw = = = p= 120 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) numerate all the 8 vertices of the rectangular parallelepiped p by means of the corresponding vertices of the unit cube p1, x000, x001, x010, x011, x100, x101, x110, x111. apply the same numeration to the values of the discrete map f, i.e., put pijl = f(xijl) for i,j,l=0,1. define the map f inside the rectangular parallelepiped p, yn = fn(x1, x2, x3) = pn000 + rn1dx1 + rn2dx2 + rn3dx3, where n, n=1,2,3, is the number of component of the map f in three-dimensional color space ℝ# and dxi = (xi xi0)/(xi1 xi0) for i=1,2,3. the coefficients rni, i=1,2,3, are determined in correspondence with the following table (cf. [2, p. 70–72]). table 1. coefficients rni, i=1,2,3 the interpolation under consideration has pure geometrical sense. we decompose a threedimensional rectangular parallelepiped into six tetrahedrons. these tetrahedrons are defined by the conditions in the second column of the table above. inside each tetrahedron the map f is constructed by linear interpolation of the values pijl, i,j,l=0,1, of the discrete map f at the vertices to the tetrahedrons. thus we have constructed the piecewise linear map f, being a solution to the direct problem of mathematical modeling for a three-dimensional printing device. in other words, we have constructed a piecewise linear model of a three-dimensional printing device. 3. direct problem for four-dimensional printing devices now consider a four-dimensional printing device, which is modelled in a way similar to the three-dimensional case. let a finite set {wi} ì w4 of points wi, i=1,…,m, of the colorants cube w4 be a regular mesh. the measurement data pi = f(wi) = f(wi), i=1,…, m, determines a discrete map f: {wi} ® {pi}, to approximate the discrete map f by a continuous map f: w4 ® ℝ#, use a piecewise linear or, which in four-dimensional case is the same, pentahedral interpolation. by definition of a regular mesh, for k=1,2,3,4 there exist the one-dimensional meshes zk = {xk0 ,…, xkm(k)}, ak = xk0 < … < xkm(k) = bk, of m(k)+1 real numbers such that {wi} = z1 ´ z2 ´ z3 ´ z4 ì w4 and m = (m(1)+1)(m(2)+1)(m(3)+1)(m(4)+1). it means that the fourdimensional colorants cube w4 can be decomposed into the union of the mesh of parallelepiped cells: , pi,j,k = [x1i-1, x1i] ´ [x2j-1, x2j] ´ [x3k-1, x3k] ´ [x4k-1, x4k]. inside each parallelepiped cell pi,j,k the continuous approximation f of the discrete map f is constructed by the following way. no conditions rn1 rn2 rn3 1 dx1 ³ dx2 ³ dx3 pn100pn000 pn110pn100 pn111pn110 2 dx1 ³ d x3 ³ dx2 pn100pn000 pn111pn101 pn101pn100 3 dx3 ³ dx1 ³ dx2 pn101pn001 pn111pn101 pn001pn000 4 dx2 ³ dx1 ³ dx3 pn110pn010 pn010pn000 pn111pn110 5 dx2 ³ dx3 ³ dx1 pn111pn011 pn010pn000 pn011pn010 6 dx3 ³ dx2 ³ dx1 pn111pn011 pn011pn001 pn001pn000 ! )4(,...,1 )3(,...,1 )2(,...,1 )1(,...,1 ,,, 3 ml mk mj mi lkjiw = = = = p= mathematical modeling of color printing devices: color gamut visualization… 121 copyright ©2021 assa. adv. in systems science and appl. (2021) consider an arbitrary four-dimensional rectangular parallelepiped p = {(x1, x2, x3, x4), a1£ x1£ b1, a2£ x2£ b2, a3£ x3£ b3, a4£ x4£ b4}. there is an obvious one-to-one correspondence between the vertices of the parallelepiped p and the vertices (0,0,0,0), (0,0,0,1), (0,0,1,0), (0,0,1,1), (0,1,0,0), (0,1,0,1), (0,1,1,0), (0,1,1,1), (1,0,0,0), (1,0,0,1), (1,0,1,0), (1,0,1,1), (1,1,0,0), (1,1,0,1), (1,1,1,0), (1,1,1,1) of the unit four-dimensional cube p1 = {(x1, x2, x3, x4), 0£ x1£1, 0£ x2£1, 0£ x3£1, 0£ x4£1}. numerate the vertices of the rectangular parallelepiped p by the corresponding vertices of the cube p1: x0000, x0001, x0010, x0011, x0100, x0101, x0110, x0111,x1000, x1001, x1010, x1011, x1100, x1101, x1110, x1111. apply this numeration also to the values of the map f, i.e., put pijkl = f(xijkl) for i,j,k,l=0,1, and define the components of the map f on the parallelepiped p: yn = fn(x1, x2, x3, x4) = pn000 + rn1d x1 + rn2d x2 + rn3d x3 + rn4d x4, where n=1,2,3, dxi = (xi xi0)/( xi1 xi0) for i=1,2,3,4, and rni are determined by the table: table 2. coefficients rni, i=1,2,3,4 no conditions rn1 rn2 rn3 rn4 1 dx1 ³ dx2 ³ dx3 ³ dx4 pn1000-pn0000 pn1100-pn1000 pn1110-pn1100 pn1111-pn1110 2 dx1 ³ dx2 ³ dx4 ³ dx3 pn1000-pn0000 pn1100-pn1000 pn1111-pn1101 pn1101-pn1100 3 dx1 ³ dx4 ³ dx2 ³ dx3 pn1000-pn0000 pn1101-pn1001 pn1111-pn1101 pn1001-pn1000 4 dx4 ³ dx1 ³ dx2 ³ dx3 pn1001-pn0001 pn1101-pn1001 pn1111-pn1101 pn0001-pn0000 5 dx1 ³ dx3 ³ dx2 ³ dx4 pn1000-pn0000 pn1110-pn1010 pn1010-pn1000 pn1111-pn1110 6 dx1 ³ dx3 ³ dx4 ³ dx2 pn1000-pn0000 pn1111-pn1011 pn1010-pn1000 pn1011-pn1010 7 dx1 ³ dx4 ³ dx3 ³ dx2 pn1000-pn0000 pn1111-pn1011 pn1011-pn1001 pn1001-pn1000 8 dx4 ³ dx1 ³ dx3 ³ dx2 pn1001-pn0001 pn1111-pn1011 pn1011-pn1001 pn0001-pn0000 9 dx3 ³ dx1 ³ dx2 ³ dx4 pn1010-pn0010 pn1110-pn1010 pn0010-pn0000 pn1111-pn1110 10 dx3 ³ dx1 ³ dx4 ³ dx2 pn1010-pn0010 pn1111-pn1011 pn0010-pn0000 pn1011-pn1010 11 dx3 ³ dx4 ³ dx1 ³ dx2 pn1011-pn0011 pn1111-pn1011 pn0010-pn0000 pn0011-pn0010 12 dx4 ³ dx3 ³ dx1 ³ dx2 pn1011-pn0011 pn1111-pn1011 pn0011-pn0001 pn0001-pn0000 13 dx2 ³ dx1 ³ dx3 ³ dx4 pn1100-pn0100 pn0100-pn0000 pn1110-pn1100 pn1111-pn1110 14 dx2 ³ dx1 ³dx4 ³ dx3 pn1100-pn0100 pn0100-pn0000 pn1111-pn1101 pn1101-pn1100 15 dx2 ³ dx4 ³ dx1 ³ dx3 pn1101-pn0101 pn0100-pn0000 pn1111-pn1101 pn0101-pn0100 16 dx4 ³ dx2 ³ dx1 ³ dx3 pn1101-pn0101 pn0101-pn0001 pn1111-pn1101 pn0001-pn0000 17 dx2 ³ dx3 ³ dx1 ³ dx4 pn1110-pn0110 pn0100-pn0000 pn0110-pn0100 pn1111-pn1110 18 dx2 ³ dx3 ³ dx4 ³ dx1 pn1111-pn0111 pn0100-pn0000 pn0110-pn0100 pn0111-pn0110 19 dx2 ³ dx4 ³ dx3 ³ dx1 pn1111-pn0111 pn0100-pn0000 pn0111-pn0101 pn0101-pn0100 20 dx4 ³ dx2 ³ dx3 ³ dx1 pn1111-pn0111 pn0101-pn0001 pn0111-pn0101 pn0001-pn0000 21 dx3 ³ dx2 ³ dx1 ³ dx4 pn1110-pn0110 pn0110-pn0010 pn0010-pn0000 pn1111-pn1110 22 dx3 ³ dx2 ³ dx4 ³ dx1 pn1111-pn0111 pn0110-pn0010 pn0010-pn0000 pn0111-pn0110 23 dx3 ³ dx4 ³ dx2 ³ dx1 pn1111-pn0111 pn0111-pn0011 pn0010-pn0000 pn0011-pn0010 24 dx4 ³ dx3 ³ dx2 ³ dx1 pn1111-pn0111 pn0111-pn0011 pn0011-pn0001 pn0001-pn0000 the interpolation under consideration has pure geometrical sense. we decompose a fourdimensional rectangular parallelepiped into 24 pentahedrons. these pentahedrons are defined by the conditions in the second column of the table above. inside each tetrahedron the map f is constructed by linear interpolation of the values pikjl, i,j,k,l=0,1, of the discrete map f at the vertices to the pentahedrons. 122 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) thus we have constructed the piecewise linear map f, being a solution to the direct problem of mathematical modeling for a four-dimensional printing device. in other words, we have constructed a piecewise linear model of a four-dimensional printing device. 4. gamut description for three-dimensional printing devices consider a piecewise linear model f: w3 ® ℝ# of a given three-dimensional printing device. for practical applications it is important to find in color space the gamut of the printing device, i.e., the image f(w3) of the piecewise linear map f. in particular, we need it to go on with the inverse problem. by definition of a piecewise linear map f, we have the simplex decomposition of the three-dimensional colorants cube w3 into the union of n, n>0, tetrahedrons dj: . each tetrahedron has four two-dimensional faces. these faces are triangles and each triangle belongs either to one or several tetrahedrons of the set {dj}. in a standard way the faces of these tetrahedrons can be divided into two classes. definition. fix a tetrahedron dl, l=1,…,n, and consider it’s two-dimensional face d, which is a triangle. the face d is called boundary if it doesn’t belong to any other tetrahedron of the set {dj}. in other words, d ì dl, and d ë dk for k=1,…,l-1,l+1,…,n. the face d is called internal if there is a tetrahedron dk from the set {dj} such that d belongs to both dl and dk, i.e., d í dl ç dk. denote the set of all the boundary faces of the colorants cube w3 by q. remark. the set q of all the boundary faces doesn’t depend on the choice of the threedimensional printing device, i.e., on the choice of the corresponding piecewise linear map f. indeed, the union of all these faces always coincides with the boundary ¶w3 of the threedimensional colorants cube: . suppose the printing device under consideration is non-degenerate, i.e., the corresponding piecewise linear map f is non-degenerate. by definition, it means that all the restrictions f½dj : dj ® ℝ#, of the map f to tetrahedrons dj are non-degenerate linear maps f½dj (x) = cj + bjx (see section 1). in other words, the determinant of the corresponding matrix bj is either positive, det bj > 0, or negative, det bj < 0. definition. fix a number l, l=1,…,n, and consider two-dimensional internal face d of the tetrahedron dl, which is a triangle. the internal face d is called singular if there exists a tetrahedron dk from the set {dj} such that d belongs to both dl and dk, d í dl ç dk, and the determinants of the corresponding matrixes bl and bk have different signs, i.e., det bl × det bk < 0. denote the set of all the singular faces of the given three-dimensional printing device by s. remark. on the contrary to the set q of all the boundary faces, the set s of all the singular faces essentially depends on the choice of the three-dimensional printing device, i.e., on the choice of the corresponding piecewise linear map f. for example, for some printing devices this set is empty and for some it is not (see section 6). it is possible to describe the gamut boundary of a non-degenerate printing device in terms of boundary and singular faces. the following theorem is through. theorem. for any non-degenerate three-dimensional printing device the boundary of the gamut is a subset of the images of all the boundary and singular faces, i.e., ! nj jw ,...,1 3 = d= 3w¶= qî ! d d mathematical modeling of color printing devices: color gamut visualization… 123 copyright ©2021 assa. adv. in systems science and appl. (2021) ¶f(w3) í f(q) è f(s). this theorem can seem to be a pure abstract mathematical proposition. and, of course, it really is. but nevertheless, it does significantly more because it gives a strict mathematical ground for creation of a wide range of computational algorithms for practical approximation of color gamut boundaries of various three-dimensional printing devices. in particular, the author designed a variant of such an algorithm, which constructs a three-dimensional gamut boundary for any real color printing device with three colorants. as an input it takes a file with colorimetric measurements data of the printing device under consideration (see section 1), and as an output it constructs the color gamut boundary, i.e., the corresponding boundary of the three-dimensional body in a reference space of three colors. as it was said in the preface, the same details of computational algorithms are not presented in this article. therefore, we will just give an example of what the designed algorithm does in practice. the picture below visualizes a piecewise linear approximation of the gamut boundary of a rather standard printing device with three colorants, which means that the stimuli space is three-dimensional in this case. in reference space the lab coordinate system is chosen, and one colorants limitation is applied. fig. 1. a piecewise linear approximation of the gamut boundary for a three-dimensional printing device the next picture gives another example of the output of the same algorithm, which is a visualized color gamut boundary. here we handle exactly the same printing device as in the previous case but the colorants limitation is changed to the tighter one, namely to . 250 3 1 £å =i ix 150 3 1 £å =i ix 124 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) the way of visualization is also exactly the same, including angle of view, the scale, and the coordinate system. fig. 2. results of the algorithm for the same printing device with tighter colorants limitation 5. gamut description for four-dimensional printing devices consider a piecewise linear model f: w4 ® ℝ#, of a four-dimensional printing device. definition. if the image of the boundary ¶w4 of the four-dimensional colorants cube coincides with the image of the whole cube w4, i.e., f(w4) = f(¶w4), then the fourdimensional printing device is called proper. remark. from technological point of view, the assumption of a four-dimensional printing device to be proper is reasonable for most real four-dimensional printing devices. by definition of the piecewise linear map f, all the restrictions f½dj : dj ® ℝ#, of the map f to pentahedrons dj are linear maps, i.e., f½dj (x) = cj + bjx, where bj is a 3´4 matrix for j=1,…,n. let bji be the 3´3 matrix obtained by throwing away the i-th column from the 3´4 matrix bj and put cj = ( det bj1, det bj2, det bj3, det bj4 ) for j=1,…, n. definition. consider the four-dimensional printing device corresponding to the piecewise linear map f. a vector field c on the colorants cube w4 is called the characteristic vector field of the printing device under consideration if c½dj = cj for j=1,…, n . the four-dimensional printing device is non-degenerate if the corresponding characteristic vector field c is nondegenerate, i.e., cj ¹ 0 for all j=1,…,n. mathematical modeling of color printing devices: color gamut visualization… 125 copyright ©2021 assa. adv. in systems science and appl. (2021) remark. by definition, the characteristic vector field of any four-dimensional printing device is a four-dimensional piecewise constant vector field on the four-dimensional colorants cube w4. in this section we will describe the gamut of a proper non-degenerate four-dimensional printing device, i.e., the image f(w4) of the corresponding piecewise linear map f in color space. by definition of a piecewise linear map, we have the simplex decomposition of the fourdimensional colorants cube w4 into the union of n, n>0, pentahedrons dj, . each pentahedron has five three-dimensional faces. these faces are tetrahedrons and each tetrahedron either belongs to one or several pentahedrons of the set {dj}. definition. fix a pentahedron dl, l=1,…,n, and consider it’s three-dimensional face, which is a tetrahedron d. the face d is called boundary if it doesn’t belong to any other pentahedron of the set {dj}. in other words, d ì dl, and d ë dk for k=1,…,l-1,l+1,…,n. denote the set of all the boundary faces of the colorants cube w4 by q. remark. the set q of all the boundary faces doesn’t depend on the choice of the fourdimensional printing device, i.e., on the choice of the corresponding piecewise linear map f. the union of all these faces always coincides with the boundary ¶w4 of the four-dimensional colorants cube, . on the boundary ¶w4 of the four-dimensional colorants cube w4 there exists the normal vector field n to this cube. let dj, j=1,…,n, be a boundary face of the four-dimensional colorants cube w4 belonging to the pentahedron d j. denote by nj the restriction of the normal vector field n to this face: nj = n½dj. let dk and dl be boundary faces of the four-dimensional colorants cube w4 such that dk ì dk, dl ì dl for some pentahedrons dk and dl, k,l=1,…,n. by definition, these boundary faces are tetrahedrons. suppose they have a two-dimensional face, which is a triangle d, in common, d = dk ç dl. definition. the triangle d is called a singular face of a non-degenerate four-dimensional printing device that is defined by the piecewise linear map f if the inner products (nk, ck) and (nl, cl) of the normal vector field n and the characteristic vector field c have different signs: (nk, ck) × (nl, cl) < 0. denote the set of all the singular faces of the given four-dimensional printing device by s. remark. the set q of all the boundary faces doesn’t depend on the choice of a printing device. on the contrary, the set s of all the singular faces essentially depends on the choice of a four-dimensional printing device, i.e., on the choice of the corresponding piecewise linear map f. moreover, a boundary face is a three-dimensional simplex, i.e., a tetrahedron, whereas a singular face is a two-dimensional simplex, i.e., a triangle. there is also a serious difference between threeand four-dimensional cases of printing devices. indeed, in three-dimensional case the boundary faces are two-dimensional and in four-dimensional case – four-dimensional. an other important difference is the following: for most three-dimensional printing devices the set s of all the singular faces is empty, while for any four-dimensional printing devices the set s of all the singular faces is not empty. it is possible to describe the gamut boundary of a proper non-degenerate printing device in terms of singular faces only. the following theorem is through. theorem. for any proper non-degenerate four-dimensional printing device the boundary of the gamut is a subset of the images of all the singular faces, i.e., ¶f(w4) í f(s). all the remarks that were made in the previous section after the theorem about the gamut boundaries of the three-dimensional printing devices are entirely correct in the case described by the theorem under consideration. in particular, a wide range of algorithms that approximate ! nj jw ,...,1 4 = d= 4w¶= qî ! d d 126 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) the color gamut boundary of any real color printing device with four colorants can be designed grounded on it. as in the case of three-dimensional stimuli space an input for this algorithm is a file with colorimetric measurements data of the printing device under consideration (see the previous section). and as an output this algorithm produces piecewise linear approximation of the color gamut boundary, i.e., the corresponding three-dimensional body boundary in a reference space of three colors. the following picture gives an example of such approximation. it shows the color gamut boundary in reference space of a rather standard printing device with four colorants (cyan, magenta, yellow, and black), which means that the stimuli space is four-dimensional. to be definite, call the measurement file of this device cmyk.dat. in reference space the lab coordinate system chosen, and one colorants limitation is applied. fig. 3. the color gamut boundary for a four-dimensional printing device the next picture gives another example of the algorithm output. here we have exactly the same printing device as in the previous picture, which is described by measurement data from cmyk.dat file, but the colorants limitation is changed to a tighter one, namely to . the way of rendering is exactly the same, including angle of view, the scale, and the coordinate system. 250 4 1 £å =i ix 150 4 1 £å =i ix mathematical modeling of color printing devices: color gamut visualization… 127 copyright ©2021 assa. adv. in systems science and appl. (2021) fig. 4. results of the algorithm for the same printing device with tighter colorants limitation 6. three-dimensional regular printing devices consider a piecewise linear model f: w3 ® ℝ# of a three-dimensional printing device. definition. the three-dimensional printing device is called regular if the piecewise linear map f is an injection. lemma. let a topological space w be compact and a map f, f: w ® f(w), be a continuous injection. then there exists the unique continuous inverse map g = f-1: f(w) ® w. in other words, then the map f is a homeomorphism. proof. see [3, ch. 2, § 8]. remark. since the three-dimensional cube w3 is a compact topological space the lemma under consideration gives a satisfactory approach to solution of the inverse problem of modeling of three-dimensional regular printing devices (see section 1). important to note, that most of the three-dimensional printing devices are regular though sometimes singular printers are met. 128 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) by definition of a piecewise linear map, we have the simplex decomposition of the threedimensional colorants cube w3 into the set of n, n>0, tetrahedrons dj, , such that all the restrictions f½dj : dj ® ℝ#, of the map f to tetrahedrons dj are linear maps, i.e., f½dj (x) = cj + bjx, where bj is a 3´3 matrix, and x, cj are three-dimensional vectors for j=1,…, n. definition. a three-dimensional printing device is called strictly non-degenerate if all the determinates of the matrixes bj are of the same sign, i.e., det bj × det bj > 0 for all the indexes i,j=1,…,n. remark. by definition of a singular face (see section 4), a three-dimensional printing device is strongly non-degenerate if and only if the set s of all its singular faces is empty, s = æ. any three-dimensional strongly non-degenerate printing device is non-degenerate. the inverse statement is false because there exist three-dimensional non-degenerate printing devices that are not strongly non-degenerate. there is an effective criterion of a three-dimensional printing device to be regular. theorem. let f: w3 ® ℝ# be a piecewise linear model of a three-dimensional printing device. this printing device is regular if and only if it is strongly non-degenerate and the restriction f½¶w3 : ¶w3 ® ℝ# of the map f to the boundary ¶w3 of the three-dimensional colorants cube w3 is an injection. by this theorem, the necessary condition of a three-dimensional printing device to be regular is its strong non-degeneracy. by definition, it means that all the determinates of the matrixes bj of the piecewise linear map f have the same sign. describe the scheme of an algorithmic approach to forcing a three-dimensional printing device to become strictly nondegenerate. at the first step count the number n+ of positive determinants and the number nof negative determinants. for clarity, assume that n+ > n-. at the second step define a positive threshold e, e > 0, which is usually a small real number, and construct the error functional r, , where rj = rj(p1, … , pm) = 0 if det bj ≥ e, and rj = rj(p1, … , pm) = (e det bj)2 if det bj < e, j=1,…, n. here p1, … , pm are the three-dimensional vectors in color space, forming the measurement data of the three-dimensional printing device under consideration (see section 1). by construction of the direct problem solution, all the determinants det bj of the piecewise linear map f are third order polynomials with respect to measurement data p1,…, pm for j=1,…,n (see section 2). hence, all the functions rj are smooth for j=1,…,n and the error functional r = r(p1, … , pm) is smooth with respect to measurement data p1,…,pm too. at the third step minimize the error functional r with respect to measurement data p1,…,pm, i.e., r(p1, …, pm) ® min, by some minimization method. the resulting argument (p10, … , pm0) of the minimal value is the measurement data for regularized three-dimensional printing device. remark. there are m three-dimensional vectors in measurement data. therefore the total dimension of the space is 3m. thus we have a 3m-dimensional non-convex minimization problem. by construction, the error functional r is not convex and can have more than one minimal point. ! nj jw ,...,1 3 = d= å = == n j mjm pprpprr 1 11 ),...,(),...,( mathematical modeling of color printing devices: color gamut visualization… 129 copyright ©2021 assa. adv. in systems science and appl. (2021) 7. four-dimensional regular printing devices consider a piecewise linear model f: w4 ® ℝ# of a four-dimensional printing device. definition. the four-dimensional printing device is called regular if the following three conditions hold for the piecewise linear map f. (1) the color gamut f(w4) is homeomorphic to closed three-dimensional disk d3. (2) for any internal point p of the color gamut f(w4), p î int f(w4), the pre-image f-1(p) is homeomorphic to a segment [a, b], a nifor i=1,2,3 and n4+ < n4. at the second step define a positive threshold e, e > 0, which is usually a small real number, and construct the error functional r, . here rji = rji(p1, …, pm) = 0 if (-1)i+1det bji ≥ e and rj = rj(p1, … , pm) = (e det bj)2 if (-1)i+1det bji < e for i=1,2,3. for i=4 rj4 = rj4(p1, …, pm) = 0 if det bji ≥ e and rj = rj(p1,…, pm) = (e det bj)2 if det bji < e for j=1,…,n. in both cases p1, …, pm are the three-dimensional vectors in color space that form the measurement data of the four-dimensional printing device under consideration (see section 1). by construction of the piecewise linear map f, all the determinants det bji are the third order polynomials with respect to measurement data vectors åå = = == 4 1 1 11 ),...,(),...,( i n j m i jm pprpprr 130 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) p1,…,pm for j=1,…,n and i=1,2,3,4 (see section 3). thence all the functions rji are smooth for j=1,…,n, i=1,2,3,4, and the error functional r = r(p1, …, pm) is smooth with respect to measurement data vectors p1,…,pm too. at the third step minimize the error functional r with respect to measurement data p1,…,pm, i.e., r(p1, …, pm) ® min, by some minimization method. the resulting argument (p10, … , pm0) of the minimal value is the measurement data for regularized four-dimensional printing device. remark. the dimension of the space of measurement data is 3m. thus we obtain a 3mdimensional non-convex minimization problem. by construction, the error functional r is not convex and can have more than one minimal point. consider the measurement file cmyk.dat that we have used in the example of section 5. direct calculations show that this printing device is not strictly non-degenerate. indeed, the table below indicates that some coordinates of the characteristic vector field are positive whereas some are negative: fig. 5. distribution for coordinates of the characteristic vector field from geometrical point of view, it means that we have too many singular faces that lead to formation of redundant faces or singularities inside the color gamut. visualization of these singularities can be seen in the picture below: fig. 6. singularities inside the color gamut mathematical modeling of color printing devices: color gamut visualization… 131 copyright ©2021 assa. adv. in systems science and appl. (2021) in particular, the singularities are the polyhedrons in white, magenta, yellow, and red. therefore, it makes sense to implement the described above algorithmic scheme and apply the resulting tool to the measurement data from cmyk.dat file. the author have implemented such a tool and successfully applied it to force the data to become strictly non-degenerate. the table below shows that after the work of this tool the data became really strictly non-degenerate: fig. 7. distribution for coordinates after applying the algorithm from geometrical point of view, it means that all the redundant singular faces have been removed and, therefore, no singularities exist anymore. visualization of color gamut of the printing device after forcing it to become strictly non-degenerated is given in the picture below: fig. 8. the color gamut without singularities after applying the algorithm indeed, no singularities are present anymore. 132 d.v. tunitsky copyright ©2021 assa. adv. in systems science and appl. (2021) acknowledgements this research was carried out with the support of the russian foundation for basic research under grant no. 20-01-00610 a and the russian scientific foundation under grant no. 19-1100223. references 1. rurk c.p., sanderson b.j. (1982) an introduction to piecewise linear topology. springer–verlag, berlin–new york. 2. kang h. (1997) color technology for electronic imaging devices. spie: bellingham, washington. 3. seifert h., trelfall w. (1980) a textbook in topology. algebraic press, new york. microsoft word 1080-source texts-5245-1-11-20220930.docx adv syst sci appl 2022; 03; 18-35 published online at https://ijassa.ipu.ru. models of industrial risk control systems mikhail geraskin1, elena rostova2* 1) samara national research university, samara, russia e-mail: innovation@ssau.ru 2) samara national research university, samara, russia e-mail: el_rostova@mail.ru abstract: we investigate a problem of searching for pareto equilibrium sets of an insurance rate and an industrial damage utilization price. we consider a system, which, in the case of industrial accidents, arises around an industrial firm. an industrial firm, a waste utilization firm, and an insurance company are considered as the system’s agents. we develop profit functions for the agents, and we determine compromise prices on waste utilization and insurance, which provide the system’s stability. we analyze the set of industrial risk control systems with a various number of the agents and the agent’s relations. a problem of determining an optimal solution is solved on the basis of maximizing agents’ profit functions. the sets of an equilibrium industrial damage utilization price and an equilibrium insurance rate are defined as pareto equilibrium. a problem of determining the set of an insurance rate is solved taking into account constraints according to requirements of an industrial firm and an insurance company. a problem of determining the set of an industrial damage utilization price is solved taking into account constraints according to requirements of an industrial firm and a waste utilization firm. we consider the following models of industrial risk control systems: agents have a strong relation and a weak relation, additionally, one agent of each type and of many agents of the same type. keywords: industrial risk, insurance, waste utilization, optimization, risk control 1. introduction the industrial risk control is an important problem for every firm because the influence of different external market factors. the risk control problems cover a financial risk, human errors, a non-fulfillment of contracts, an industrial risk, an environmental risk, etc. these problems were solved by means of the following methods: the scenario method [17], the multi-agent systems [1, 7, 19], the multi-criteria models [6, 27, 31]. additionally, this problem was analyzed on various levels: the world market risk [11], the regional economic system risk [20, 24, 26], the firm’s risk [2, 28], the technology operation risk [3, 8]. the risk control problems are related to various aspects, in particular, an assessment of the risk factors, a choice of the risk management method, a prediction of the damage, etc. rasmussen and svedung emphasized that «risk management can no longer be based on responses to past accidents and incidents, but must be increasingly proactive» [21]. therefore, an importance of developing measures to prevent risks prevails over minimizing the damage from accidents. wu, olson, and choi indicated that «optimization and risk minimization inherently run counter to each other» [30]. consequently, a choice of the risk management method should be based on an assessment of the preventive measures economic efficiency. these problems were solved on the basis of the multicriteria decision making (mcdm) methods [6, 27, 31], and the biconvex models and algorithms [25]. * corresponding author: first@ras.ru models of industrial risk control systems 19 copyright ©2022 assa. adv. in systems science and appl. (2022) for mcdm, heller [27] proposed using a pair wise comparison of the risk competing objectives. the following criteria were analyzed: a power outage, a fire, a flood, an earthquake, a hurricane, a destruction of buildings, network failures, etc. as a result, a matrix of the risk criteria assessments was formed, which is used in a qualitative risk analysis for three buildings that differ in qualitative features. abla et al. [6] used mcdm to derive an aggregated risk score on the basis of the fuzzy logic. for assessing development scenarios of the risk situation, the decision-making process was considered under the following criteria: a technical and functional efficiency, an economic sustainability, a social sustainability, an institutional environmental sustainability, an overall sustainability. the risk management strategy was selected based on a weighted sum of these criteria. in the case of the flood risk, this decision making technique was applied to select the most sustainable strategy under uncertainty. yazdani et al. [31] proposed mcdm type for assessing risks of the cultivated areas’ flooding; they ranked various agricultural projects, which can mitigate the flood risks. dudin m. n. et al. [5] investigated the risks of an industrial enterprise and calculated external factors influence weights and a likelihood of unforeseen events for political, macroeconomic, social, and technological factors. on the basis of these weights, the external risk average level of russian industrial enterprises was calculated. internal risks of the enterprise were assessed according to the following criteria: a fulfillment of the production plan, an economic security of current obligations, an economic security of supply contracts, a human factor, and a likelihood of success in innovations commercialization. the authors examined hedging and insurance methods for the risk management of an industrial enterprise, however, the insurers and other related organizations were not considered as separate agents, and their interactions were not investigated. krokhina j.a. et al. [16] explored the environmental risks of industrial enterprises using a tree method and assessed the integral risk of an industrial facility as a result of its negative impact on an environment, a human health, and an enterprise’s economy. the authors applied the tree method to assessing the risk of an accident on the main pipeline and assessed an economic efficiency of the measures to reduce the environmental risk of industrial enterprises, but did not take into account an interaction of the enterprise with other economic agents. in contrast to the aforementioned literature, multi-agent systems (mas) models expanded a range of the risk management techniques. mas models were studied in production design and development systems, a production planning and management, and a supply chain management (scm) [19]. in a set of agents, mas models described the participants in the production process supply chain. in scm, an influence of small and medium-sized enterprises in the interaction with large firms [7] and buyers with suppliers [1] on the risk level was considered. in the case of different-scale enterprises in scm, finch [7] emphasized varying degrees of the supply chain disruption risk, because large and small enterprises are subject to different degrees of the risks and their sensitivity to the risk factors is different. ahn and park [1] studied the information exchange processes between participants in the supply chain as mas agents and assessed an impact of an agent’s awareness on scm. in general, mas model was applied for the risk management in systems, which consist of an industrial enterprise and its suppliers, i.e., the participants in the supply chain. finally, it should be emphasized that the aforementioned studies were carried out for specific industries. for example, the problems of risk in the chemical and oil industries were solved [3, 8] and technical, human and organizational factors of the industrial facility risk were taken into account; on the basis of the fuzzy logic, the risk of the industrial facility in the oil and gas industry was analyzed using the criteria of frequency, detectability and damage value [30]. therefore, the results of these studies cannot be applied to all industries. thus, we demonstrate the following research gap in the problem framework of the industrial risk management. on the one hand, the industrial researchers pointed to a need for 20 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) the insurance and the technical measures to prevent or eliminate consequences of industrial accidents. on the other hand, they did not investigate the interaction mechanisms of an enterprise with insurers and waste utilization firms as the specific agents. additionally, they did not generalize the results for the universal industrial enterprise; they were limited to the specific industry. at the same time, mas researchers did not study the principles of applying mas in the process of industrial risk management. hence, we can formulate the following research question of the industrial risk management problem: to describe the industrial risk management process for the universal industrial enterprise within the system of interconnected economic agents and calculate the equilibrium prices for services that circulate within this system. our study aims at calculating the price equilibrium in the system, which appears as a result of the preventive measures to minimize the consequences of technical accidents in the industry. 2. problem framework in this paper, the risk is considered at the firm’s level, and it includes an internal damage and an external damage. the internal damage causes a reduction in the firm’s assets. the external damage is the property wastes of other firms, individuals and the environment. additionally, the fiscal penalties (the ecology payment, the penalty for a damage to health, and a property of other firms and individuals, the compensation caused by the non-fulfillment contracts) depend on a value of the external damage. the industrial firm can reduce the internal/external damage by means of additional expenses on the risk reduction. these costs expenses are named the voluntary risk costs (vrc). we analyze the problem of the industrial risk control for a system with three agents: the industrial firm, the waste utilization firm, and the insurance company. these agents are in various relationships in the system. we consider the following problem: to search for a compromise price of waste utilization and a compromise insurance rate, which are compliant with all agents of the system. we assume that each participant in the system is intended to increase its profits, and he chooses the optimal price. if these optimal prices are different, then the participants may not agree to conclude a contract, then the industrial risk management will not be implemented. therefore, in this case, we determine the set of possible values of the insurance rate and the price of waste utilization, at which the participants in the system will agree. we introduce the following assumptions, which determine the applicability limits of the model. assumption 1. the product price is an exogenous constant, that is, the firm does not affect the price ,0 dq dp (1) where p is the price of the production, q is the production volume. the waste utilization firms and the insurance companies are in the monopolistic competition market, that is, the following conditions are fulfilled: ,0,0  u y u y dx dp dy dp (2) ,0,0       ss y t x t (3) models of industrial risk control systems 21 copyright ©2022 assa. adv. in systems science and appl. (2022) where py is the price of the utilization of a conventional waste unit, t is the insurance rate, xu and yu are the internal and external utilized damage, xs and ys are the internal and external insured damage. assumption 2. the production growth leads to a decreasing return: ,0qqc (4) where c is a value of the firm’s costs. assumption 3. an increase in the production assets leads to an increasing in the possible damage; the internal damage and the external damage are reduced with an increase in vrc; the internal damage is limited from above due to technology features and the production volume 0],,0(,0,0 maxmax       xxx f x q x . (5) where xmax is the maximum possible internal damage, x is the internal damage, f is vrc. assumption 4. the external damage y is proportional to the internal damage x: 0   x y . (6) assumption 5. the voluntary combination insurance is considered, the wear is not included. the insurance indemnity w is proportional to the insured damage xs and ys, the indemnity does not exceed the damage: .,0,0 ss ss yxw y w x w       (7) assumption 6. the cost of the utilization of a conventional waste unit cy is a constant. cy = const. (8) assumption 7. the firm’s external damage y=ys+yu+yres consists of the insured external damage ys= s y, the utilized external damage yu= u y, and the residual external damage yres= res y. the firm’s internal damage x=xs+xu+xres consists of the insured external damage xs= s x, the utilized internal damage xu= u x, and the residual internal damage xres= res x. 0,,,1  resusresus  , (9) 0,,,1  resusresus  . (10) the production costs function, according to assumption 2, has the following form [4], [29]. bqqсq )( , 0],2,1(],,1( maxmax  b , (11) where b and β are the parameters of the production costs function, βmax is the maximum possible parameter value. the internal damage function satisfies assumption 3, and it has the following form: feqfqx   )(),( , ],1,0(],,0( maxmax   0)(  q . (12) this function х(q) expresses an exponential distribution of the damage, which corresponds to man-made accidents, φ(q) is the dependence of the damage on the production volume q,  is the parameters of the internal damage function, max is the maximum possible parameter value. 22 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) the external damage function satisfies assumption 4: 0,)(  xxy . (13) the coefficient of the accident consequences expansion μ expresses the ratio of the external damage and the internal damage, taking into account the specifics of the industrial complex in a region, geographical features, etc. the insurance indemnity satisfies assumption 5: ,10),(),(   ssss yxyxw (14) where α is the coefficient of the insurance indemnity. the penalty function has the following form: ,0,  aхaayh  (15) where a is the parameter of a relationship between the penalty and the external damage. we consider the systems, which include the agents of three types: the 1st agent is the industrial firm, the 2nd agent is the waste utilization firm, and the 3rd agent is the insurance company. the industrial firm (the 1st agent) produces the production volume q, and it sells the product at the price p. we introduce the following notation: cq is the production costs function, x is the internal damage, y is the external damage, f is vrc, h(y) is the penalty function, f(х,y) is the value of waste utilization costs, v(x, y) is the insurance premium, w(x, y) is the insurance indemnity, qmax is the maximum possible production volume, fmax is the maximum possible vrc. the revenue function of the 1st agent is wqpr  . (16) the total costs function of the 1st agent is fhvxfcс res q   . (17) the profit function of the 1st agent is пi=r – c∑. (18) we formulate the problem of the firm’s choice as follows: to search for the production volume and vrc function, which maximize the profit of the 1st agent, that is: .maxarg*}*,{ , i aqaf пqf qf   (19) }0,:{ maxmax   qqqrqaq (20) )},0(,)(:)({ maxmax if rfffrfa   (21)                     ).( ),( , ),( , , ,)( ss ss res uu y q f yxtv yxw ayh xypf bqc xy eqx         (22) the waste utilization firm (the 2nd agent) has the following parameters. the utilized damage uu xy   does not exceed the level y , the price py does not exceed the level yp , where yp is the maximum possible price, y is the maximum possible waste utilization. consequently, the inverse demand function of the 2nd agent, according to assumption 1, has the form: )( uuy yy xy y p pp   . if 0yp , then yxyyx uu  :, models of industrial risk control systems 23 copyright ©2022 assa. adv. in systems science and appl. (2022) the profit function of the 2nd agent is ))(( uu yyii xycpп   . (23) we formulate the problem of the choice of py : to search for the price of the utilization of a conventional waste unit, which maximizes the profit of the 2nd agent, that is: ii rp y пp y   maxarg* (24) ).( uuy yy xy y p pp   (25) the insurance company (the 3rd agent) has the following parameters. the insurance premium depends on the insurance rate t and the insured damage ss xy   ; t is the maximum possible insurance rate, x is the maximum possible insured damage. consequently, the inverse demand function of the 3rd agent, according to assumption 1, has the form: x t xytt ss )(   . if 0t , then xxyyx ss  :, . the profit function of the 3rd agent is wvпiii  . (26) we formulate the problem of the choice of the insurance rate: to search for the insurance rate, which maximizes the profit of the 3rd agent, that is: iii tst пт )1,0( maxarg*   (27)             . ),( ,)( ),( xxy xyw x t xytt xytv ss ss ss ss     (28) we consider the following problem of the agent’s optimal control: to search for the pair (q*, f*), which is optimal according to criterion (19), to search for the price py*, which is optimal according to criterion (24), and to search for the rate t*, which is optimal according to criterion (27). this system consists of three agents, and, if we vary the parameters q, f, py, t, then three agents achieve maximums of their profits. we consider the system of the industrial damage control of the following types. the agents have a strong relation, if they have the collective criterion function, and their costs are not separable. the agents have a weak relation, if the costs are separable, and each agent has the individual criterion function. we investigate the types of the system in the following models. model 1: the 1st agent is the customer of the waste utilization and the insurance, the 2nd agent is the contractor of the waste utilization, the 3rd agent is the insurer of the internal damage and the external damage of the 1st agent. this system has the weak relation type. (fig. 1) 24 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) fig. 1. the agent’s interaction schema in model 1. the 1st agent pays the sum )( uu y xypf   to the 2nd agent, and the agents achieve the contract, if the price py complies with everyone. we formulate the problem with two criteria to search for the compromise price com yp , under which the system is stable. by using an analogy with the previous case, the 1st agent pays the sum )( ss xytv   to the 3rd agent, and the agents achieve the contract, if the insurance rate t complies with everyone. we formulate the problem with two criteria to search for the compromise rate comt , under which the system is stable. thus, the system of three agents is stable, if the compromise price com yp and compromise rate comt are indicated in the contracts. consequently, we formulate the following problems: to search for the compromise price com yp and the compromise rate comt , which satisfy the following conditions: ii gp i gp пп yy   maxmax , (29) iii t i t пп  maxmax , (30) }0)(0)(|{  yiiyiy pпpпpg , (31) }0)(0)()1,0(|{  tпtпtt iiii . (32) in formulas (29), (30), the symbol of the conjunction “ “ means that the maximums are determined according to both criteria, taking into account the pareto optimal principle. model 2: the 1st agent and the 2nd agent have the strong relation; the 3rd agent is the insurer of the internal damage and the external damage of the 1st and the 2nd agents; the 3rd agent has the weak relation to other agents. (fig. 2) models of industrial risk control systems 25 copyright ©2022 assa. adv. in systems science and appl. (2022) fig. 2. the agent’s interaction schema in model 2. the aggregate profit function of the 1st and the 2nd agents in the system is resuu yqiiiiii xxychvfcwqpппп   )(, . (33) the problem of searching for the optimal production volume and vrc for the 1st agent, and, additionally, the optimal rate of the 3rd agent, has the following form: iii aqaf пqf qf , , maxarg*}*,{   , (34)                     ).( ),( , , , , ,)( ss ss res uu q f yxtv yxw ayh yxy bqc xy eqx         (35) these variables influence on the waste utilization, the insured damage, the penalties, and the insurance rate. we formulate the problem of searching for the compromise rate accounting to the following conditions: iii t iii t пп 11 maxmax ,   , (36) }0)(0)()1,0(|{ ,1  tпtпtt iiiiii . (37) model 3: two agents of the 1st type are customers of the waste utilization and the insurance; the 2nd agent is the contractor of the waste utilization; the 3rd agent is the insurer of the internal damage and the external damage of the 1st agent. this system has the weak relation. (fig 3) 26 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) fig. 3. the agent’s interaction schema in model 3. the problem of searching for vrc and the optimal production volume of the 1st type agents is iii res iiiiqiiiii fhvxfcwpqп   , 2,1i , (38) i aqaf ii qifi qf maxarg , *}*,{   , (39)                     ).( ),( , ),( , , ,)( s ii s iii s ii s iiii res iii u ii u iiyi iiiq ii f iii yxtv yxw ayh xypf qbc xy eqx i i         2,1i (40) where }2,1,{  iп iii is the vector of criteria in model 3. the profit function (23) of the 2nd agent is the sum of the revenues, which are provided by two agents of the 1st type, and it has the following form:    2 1 )()( i u ii u iiyyii xycpп  . (41) the problem of searching for the price py* according to maximization of the 2nd agent’s profit function is ii rp y пp y   maxarg* , (42) .2,1),(  ixy y p pp u ii u ii y yy  (43) by using an analogy with the previous case, the profit function (26) of the 3rd agent is    2 1 )( i iiiii wvп . (44) the problem of searching for the insurance rate t according to the maximization of the 3rd agent’s profit function is iii tst пт )1,0( maxarg*   (45) models of industrial risk control systems 27 copyright ©2022 assa. adv. in systems science and appl. (2022)                .)( ),( ,)( ),( 2 1 xxy xyw x t xytt xytv i s ii s ii s ii s iiii s ii s ii s ii s iii     (46) we formulate the problem of searching for the compromise price com yp and the compromise rate comt , accounting to the following conditions: ii gp i gp i gp ппп yyy 111 maxmaxmax 21   . (47) }0)(0)(0)(|{ 211  yiiyiyiy pпpпpпpg . (48) iii t i t i t ппп 222 maxmaxmax 21   (49) }.0)(0)(0)()1,0(|{ 212  tпtпtпtt iiiii (50) 3. results assertion 1. the function |)(|ln 1 * kqf    , u y u y resssresss attk   and the value q*, which is calculated from the equation 0 )( )(1     q q qbp    , are the solution of the problem (19 – 22) for the continuously differentiable functions )( and under conditions 0))(())()1(()( 222   keqkeqqbkeq fff   and k > 0, 0)()()( 2  qqq  . the maximal profit of the 1st agent is   1 ****  fbqpqпi . proof. the profit function of the 1st agent (18) is   resresfss i ayeqfbqyxtqpп   )())(( )( uu y yxp   . we search for the partial derivatives of this function, which are equal to zero: 0)]())([()(1     uu y resresssfi pateq f п   0)]())([()(1     uu y resresssfi pateqqbp q п   we solve these equations as follows: 1)]())([()(  uu y resresssf pateq   28 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) we introduce the denotation )())(( uu y resresss patk   , then we can write the optimal function vri as follows: ))(ln( 1 * kqf    . we write the equation 0)(1   keqqbp f  , and, for the aforementioned symbols k and f*, we get 0 )( )(1     q q qbp    . we solve this equation, and we search for q*. we consider a function )( )( )( 1 q q qbpqh       . this function is continuous function qaq . due to the features of continuous functions, the equation h(q)=0 has a solution, because 0)( qh for )( )( | 1 q q qbpaq q       , 0)( qh for )( )( | 1 q q qbpaq q       , and 0 )( )()()( )1()( 2 2 2     q qqq qbqh    if 0)()()( 2  qqq  . we check a fulfillment of the maximum sufficient condition for the function пi(q, f) at f=f* and q=q*. for this purpose, we define the sign of 22 2 2 2 2                 qf п q п f п iii . keq f п fi     )(2 2 2 , keqqb q п fi       )()1( 2 2 2 , keq qf п fi     )( 2 . 0))(())()1(()( 222   keqkeqqbkeq fff   . if k>0, then 2 2 f п i   <0 and 2 2 q пi   <0, consequently, the pair (f*,q*) is the maximum point according to sylvester’s criterion. █ assertion 2. the value 2 * yy y pc p   is the solution to the problem (24 – 25) for the continuously differentiable function y(py). the maximal profit of the 2nd agent is пii*=пii( * yp )= ))(( 4 1 2 yy y cpy p  . proof: we write the profit function of the 2nd agent (23), which is subjected to the condition (25): models of industrial risk control systems 29 copyright ©2022 assa. adv. in systems science and appl. (2022) .))(( y yyyyii p y ppcpп  we transform this expression as follows: .2 yy y y y yii cyc p y yp p y pп              this expression is the second-order power function, and it has the parabola graph with the maximum at . 2 * yy y cp p   . we search for the maximum of the profit function *)(* yiiii pпп  of the 2nd agent:                            yy y yy y yy ii cyc p y y cp p ycp п 22 * 2 2)( 4 yy y cp p y  .█ assertion 3. the insurance rate 2 *   t t is the solution to the problem (27 – 28) for the continuously differentiable functions x(t) and y(t). the maximal profit of the 3rd agent is 2)( 4 1 *)(*  tx t tпп iiiiii . proof: the profit function of the 3rd agent is t tt xtxytwvп ss iii   )())((  . we transform this function as follows: x t x xt t x tп iii               2 . this function has the parabola graph with the maximum point at 2 *   t t . we determine the maximum of the profit function of the 3rd agent as follows: 2 2 )( 422 *)(*                             t t x x t x x t t xt tпп iiiiii .█ assertion 4. if 2 y y p c  , then     2 ; y y com y p cp is the solution of the problem (29), (31) for the continuously differentiable functions )( , else com yp ø. proof: the profit function of the 1st agent is )()( uu y res qyi yxphvxfcwqppп   the utilized damage uu yx   corresponds to the demand function: . y yuu p p yyyx   then the profit function of the 1st agent is 1 2 1 1)( kyp p y p p p ypkpп y y y y y yyi        , 30 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) where 1k is hvxfcwqp res q   . the maximum point of this function is 2 y y p p  , and the maximum of the profit function of the 1st agent )( yi pп is 142 k ypp п yy i       . the profit function пii(py) is analyzed in the proof of assertion 2. fig. 4. compromise set of the 1st and the 2nd agents the value 0 yp is yc . the figure 4 shows that the 1st agent prefers price     2 ,0 y y p p , the 2nd agent prefers price . 2 ,0       yy yy pc pp if 2 0 y y p p  then     2 ,0 y y com y p pp else         2 ,0 2 ,0 yyy y ppc p ø █ the value com yp is the solution of the problem (29), (31), and it enables us to establish the price yp , which complies with the 1st and the 2nd agents. if one of the agents changes this price, then the profit of other agent decreases, therefore, this is pareto optimal equilibrium set for 1st and the 2nd agents’ prices according to profit function (18) as the criterion of the 1st agent and profit function (23) as the criterion of the 2nd agent. assertion 5. if 2 t  , then        2 ; t t com  is the solution of the problem (30), (32) for the continuously differentiable functions )( , else comt ø. proof: the profit function 1st agent is )()( ssres qi yxtfhxfcwqptп   . models of industrial risk control systems 31 copyright ©2022 assa. adv. in systems science and appl. (2022) the insured damage ss yx   corresponds to the demand function in the insurance market: . t t xxyx ss   we transform the profit function of 1st agent as follows: .1)( 2 2 2 kxt t x t t t xtktпi        the value 2k is hfxfcwqp res q   . the minimum point of profit function )(tпi is 2 t t  and the minimum of the function is 242 k xtt пi       . the profit function пiii(t) is analyzed in the proof of assertion 3. the graphs of the 1st and the 3rd agents’ profits for problem (30), (32) are demonstrated in fig. 5. fig. 5. compromise set of the 1st and the 3rd agents the acceptable set comt is interval       2 ; t . these values enable us to transact of the insurance contract, because the value comt complies with the 1st and the 3rd agents. the deviation of the insurance rate relative to comt leads to a decrease in the profit of one of the agents. thus, we identify the set of the possible values of the insurance rate comt and the price of waste utilization com yp , at which the system agents interact, and the process of the industrial risk management is implemented. similar to the conclusion from assertion 4, we establish the pareto optimal equilibrium set for the 1st and the 3rd agents’ prices according to profit function (18) as the criterion of the 1st agent and profit function (26) as the criterion of the 3rd agent. 32 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) 4. discussion the currently accepted risk management model is based on the standards [13] [15] and, in fact, implements the intra-firm management [3], [4], [6], [8], [16], [27], [31]. in comparison with the aforementioned literature, we investigate the process of the industrial risk management in the system of several organizations interconnected within the framework of this process. our system consists of three agents (the industrial firm, the insurer, and the waste utilization firm) and provides a comprehensive examination of the risk management process. in contrast to use of mas for scm, we consider the mas model, which includes the agents that are not linked in a supply chain. we add the insurers and the waste utilization firm to this model; therefore, our results extend the studies [18], [23]. in addition, our results are industry invariant; consequently, it can be used for any industrial production. this generality of results distinguishes our study from authors who considered risks for a particular industry, for example [3], [8]. we derive an analytical form of the preventive risk costs function for an industrial enterprise and prove that these costs depend logarithmically on the production volume of an enterprise. this pattern means that the preventive risk costs increase more slowly in comparison with a growth of the production. this feature encourages enterprises to take the preventive measures and develops the concept of the proactive risk management [21]. therefore, we prove that the proactive risk management strategy is optimal, i.e., it corresponds to the minimum costs. we calculate the optimal values of the waste utilization price and the insurance rate, which maximize the objective functions of the waste utilization firm and the insurer. in comparison with the approach [4], [16], which is based on the exogenously specified insurance rate, in our model, the insurance rate is determined endogenously. in other words, our model is based on such values of the waste utilization price and the insurance rate that induce these firms to participate in contracts with the industrial enterprise. consequently, in these conditions, the decentralized decision-making system is configured according to the mechanism of the centralized system. in addition, we find the conditions of the compromise domain, i.e., the ranges of the waste utilization price and the insurance rate, within which the system participants are interested in concluding contracts. therefore, we expand the mas approach [1, 7, 19] related to the revenue sharing, and reformulated it into the mechanism of the price sharing contract. next, we prove that the contract prices within the specified ranges are pareto efficient. therefore, when the price varies within the compromise domain, the profit of one participant grows, and the profit of the counterparty decreases, i.e., the price sharing contract corresponds to the profit sharing contract. this is an important advantage of our approach: our model not only allows us to estimate the damage from an industrial accident, but also to calculate the economic effects of all participants in the risk management process and choose the preventive costs sum that corresponds to the optimal solutions for all participants. the results of this article can be used by industrial enterprises, insurance companies, and waste utilization firms to determine the insurance rate and the utilization price. the resulting compromise values of the rate and the price demonstrate the set of acceptable values at which a contract will be concluded. finally, we briefly outline the directions for further research within our version of the mas model. our results are obtained under certain restrictions on the type of a market (assumption 1), the production function (assumption 2), and the damage function (assumption 3). in the future, we plan to expand the study to consider other types of markets and other production and damage functions. models of industrial risk control systems 33 copyright ©2022 assa. adv. in systems science and appl. (2022) 5. conclusion the optimization problems of searching for the firms’ risk prevention costs were investigated in our previous articles [9], [22]. in particular, the optimal vrc function in the case of the internal damage prevention according to the condition of the firm’s profit maximization was derived. the problem of the external damage control was analyzed, and the optimal vrc function with regard to fiscal penalties for the environmental damage and the civil penalties for individuals’ property damages was proved. the problem of searching for the optimal risk costs, taking into account the reinvestment of the firm’s profit, was considered, and the optimal vrc function for a firm’s activity in the consequent periods was obtained. in this paper, the problem of the industrial risk control in different system’s structures is investigated. we consider a risk control system of an industrial firm, which includes the insurance company and the waste utilization firm. this system enables us to determine the conditions of the insurance contract and the waste utilization contract. the problem of the industrial risk control is a problem of the agents’ interests congruence. the calculated values of the insurance rate and the waste utilization price are determined as the pareto equilibrium set for price and insurance rate. we obtain the following results. the problem of the firm’s choice of the product volume and vrc function in the system, which includes an industrial firm, a waste utilization firm, and an insurance company, is solved in assertion 1. the problem of searching for the price of the utilization, which maximizes the profit of a waste utilization firm, is solved in assertion 2. the problem of searching for the insurance rate, which maximizes the profit of an insurance company, is solved in assertion 3. the problem of searching for the compromise utilization price and the compromise insurance rate, taking into account pareto optimal principle for this variable, is solved in assertion 5. the final result provides the set of acceptable values at which a contract will be concluded. references 1. ahn, h.j., park, s.j. (2003) modeling of a multi-agent system for coordination of supply chains with complexity and uncertainty. in: lee j., barley m. (eds) intelligent agents and multi-agent systems. prima 2003. lecture notes in computer science, vol. 2891. springer, berlin, heidelberg. https://doi.org/10.1007/978-3-540-39896-7_2 2. arena, m., arnaboldi, m. & azzone, g. (2011) is enterprise risk management real?, journal of risk research, 14, 779 – 797. 3. bouloiz, h., garbolino, e. (2019) system dynamics applied to the human, technical and organizational factors of industrial safety. safety dynamics (p. 93 – 106). cham: springer. 4. choi, t. m., chan, h. k., yue, x. (2016) recent development in big data analytics for business operations and risk management // ieee transactions on cybernetics, vol. 47, no. 1, 81-92. 5. dudin, m.n., frolova, е.е., lubenets, n.a., sekerin, v.d., bank, s.v., gorohova, a.e. (2016) methodology of analysis and assessment of risks of the operation and development of industrial enterprises // calitatea, vol. 17, no. 153, 53. 6. edjossan-sossou, a.m., galvez, d., deck, o., heib, m.a., verdel, t., dupont, l., chery, o., camargo, m., morel, l. (2020) sustainable risk management strategy selection using a fuzzy multi-criteria decision approach // international journal of disaster risk reduction. vol. 45, may 2020, 101474 https://doi.org/10.1016/j.ijdrr.2020.101474 34 m. geraskin, e. rostova copyright ©0000 assa adv. in systems science and appl. (0000) 7. finch, p. (2004) supply chain risk management // supply chain management, vol. 9 no. 2, 183-196. https://doi.org/10.1108/13598540410527079 8. gallab, m., bouloiz, h., youssef, l.a., tkiouat, m. (2019) risk assessment of maintenance activities using fuzzy logic // procedia computer science, vol. 148, 226-235 9. geraskin, m., rostova, e. (2018) costs function optimization for prevention costs function optimization for prevention of firm’s industrial risks with regard to reinvestment of profit, advances in systems science and applications, vol. 18, no. 4, 52-63. 10. gorecki, s. et al.(2019) risk management and distributed simulation in papyrus tool for decision making in industrial context //computers & industrial engineering, vol. 137, 106039. 11. hanson, d. & white, r (2004) regimes of risk management in corporate annual reports: a case study of one globalizing australian company, journal of risk research, 7, 445 – 460. 12. hay, d., morris, d. (1991) the theory of industrial organization, oxford university press: revised edition. 13. iso 31000:2009 «risk management – principles and guidelines », 2009 14. iso guide 73:2009 «risk management – vocabulary», 2009 15. iso/iec 31010:2009 «risk management – risk assessment techniques» 16. krokhina, j. a. et al. (2018) environmental risk management system projecting of industrial enterprises, ekoloji, vol. 27, no. 106, 735-744. 17. kulba, v., schelkov, a., chernov, i., zaikin, o. (2016) scenario analysis in the management of regional security and social stability, intelligent systems reference library, 98, 249 – 268. 18. lee, j., lee, d. k. (2018) application of industrial risk management practices to control natural hazards, facilitating risk communication, isprs international journal of geo-information, vol 7, № 9, 377 19. lee, j.& kim, с. (2008) multi-agent systems applications in manufacturing systems and supply chain management: a review paper, international journal of production research, vol. 46, issue 1, 233-265. https://doi.org/10.1080/00207540701441921 20. pazdnikova, n.p., shipitsyna, s. y. (2014) stress analysis in managing the region’s budget risks, risk factors for the regional economic growth, economy of region, 3(39), 208 – 217. 21. rasmussen j., svedung, i. proactive risk management in a dynamic society. – swedish rescue services agency, 2000. 22. rostova, e.p., geraskin m.i. (2018) optimization of costs function for prevention of firms' industrial risks with penalties. the proceedings of the third workshop on computer modeling in decision making (cmdm 2018). acsradvances in computer science research. vol. 85, 26-30. 23. samanlioglu, f. (2013) a multi-objective mathematical model for the industrial hazardous waste location-routing problem, european journal of operational research, vol. 226, no. 2, 332-340 24. sapiro, e.s., miroljubova, t.v. (2008) risk factors for the regional economic growth, economy of region, 1, 39 – 49. 25. sherali, h.d., alameddine, a., glickman, t.s. (1994) biconvex models and algorithms for risk management problems, american journal of mathematical and management sciences, vol. 14, no. 3-4, 197-228. 26. shorikov, a.f. (2012) dynamic model of minimax control over economic security state of the region in the presence of risks, economy of region, 2(30), 258 – 266. models of industrial risk control systems 35 copyright ©2022 assa. adv. in systems science and appl. (2022) 27. heller, s. (2006) managing industrial risk—having a tested and proven system to prevent and assess risk // journal of hazardous materials. vol. 130, issues 1– 2, 17 march, 58-63 https://doi.org/10.1016/j.jhazmat.2005.07.067 28. thun, j.-h., drüke, m. & hoenig, d. (2011) managing uncertainty – an empirical analysis of supply chain risk management in small and medium-sized enterprises, international journal of production research, 49, 5511 – 5525. 29. walters, a.a. (1963). production and cost functions: an econometric survey. econometrica. the econometric society, econometrica. 30. wu, d.d., olson, d.l., choi t.m. (2017) guest editorial special issue on risk analytics in industrial systems, ieee systems journal. vol 11, №. 3, 1476-1478. 31. yazdani, m., gonzalez, e.d.r.s. and chatterjee, p. (2019), a multi-criteria decision-making framework for agriculture supply chain risk management under a circular economy context, management decision, vol. ahead-of-print no. ahead-of-print. https://doi.org/10.1108/md-10-2018-1088 microsoft word 868_done adv syst sci appl 2023; 02; 1-11 published online at https://ijassa.ipu.ru. pretreatment performance evaluation of the seawater desalination plant of beni saf bwc mourad berrabah1*, hassiba bouabdesselam1 , noreddine ghaffor2 1) lte research laboratory, national polytechnic school of oran, algeria 2) king abdullah university of sciences and technology, thuwal, saudi arabia abstract: the beni saf seawater desalination plant is one of the largest projects undertaken by the algerian state for supplying drinking water, with a capacity of 200,000 m3 per day. the plant uses the reverse osmosis technique as a desalination process. the performance of such systems requires the production of pretreated, good-quality feed water. moreover, a number of large-scale experiments have shown that pretreatment of seawater before reverse osmosis (ro) desalination is key to retard fouling in osmosis membranes. this work aims to study the performance of the pretreatment selected by the beni saf desalination plant. to achieve this, we determined the physico-chemical and bacteriological parameters of the raw seawater used by the plant. variation of sdi, which is a key parameter in controlling fouling potential, was monitored during the year 2019. the pretreatment performance was investigated by monitoring the efficiency of each stage of the pretreatment process. the results obtained show the efficiency of the pretreatment adopted by this desalination plant. keywords: desalination; fouling; pretreatment; retarding. 1. introduction to address water shortage, algeria has reinforced its water resources by waste water treatment and seawater desalination plants. the desalination plant of beni saf produces 200,000 m3/d of water to meet the needs of the region of ain temouchent and oran. located on the mediterranean coast of algeria in the wilaya of ain temouchent, the plant covers an area of 65700 m2. the desalination process adopted for this plant is the reverse osmosis (ro) technique. a major challenge experienced by the plant is the poor quality of the raw seawater and therefore the rapid fouling of the reverse osmosis membranes. feed water is characterised by turbidity, bacterial content and tds (total dissolved solids), which causes fouling and results in the degradation of the membrane performance. when a membrane is affected by fouling, there is deterioration of its basic functions; this includes salt passage, decreased permeate flow and pressure drop across the membrane. on the other hand, inorganic scaling caused by exceeding the solubility limit of soluble salts is considered less problematic since it can be controlled by ph and addition of antiscalants [7]. deposits on the membrane or on the pores leads to a decrease in the performance of the membrane [5], and to prevent rapid formation of these deposits, the seawater desalination plant of beni saf adopted a conventional physico-chemical treatment. the objectives of the pretreatment are as follows: * improving the quality of water, enhancing the performance of the membranes (reducing membrane fouling). pretreatment has proved to be efficient in improving the performance of the membrane and reducing the cost of its replacement since it eliminates turbidity, bacteria and tds (total dissolved solids). without pretreatment, ro * corresponding author: berrabahmourad534@gmail.com membranes are prone to rapid fouling which results in frequent replacement with the consequential increase in operational costs. to address this issue, a pretreatment method 2 m.berrabah, h.bouabdesselam copyright ©2023 assa adv. in systems science and appl. (2023) must be integrated before feed water enters the ro system to help reduce membrane fouling and improve thereby its performance [1]. 2. materials and methods 2.1. physico-chemical and biological characteristics of the raw seawater from the desalination plant of beni saf the raw seawater of the beni saf plant was characterized by taking a sample of raw seawater and analyzing the physico-chemical and bacteriological parameters. water sampling is a sensitive operation. it allows defining the quality of water at a given time or over a more or less long specified period. therefore, it should be handled with the greatest care [3]. table 1 shows the results of the physico-chemical and bacteriological analyses performed on a sample taken on 27/11/2019. table 1. results of the physico-chemical and bacteriological analyses of the raw seawater from beni saf plant physico-chemical analysis of raw water turbidity (ntu) 1.76 temperature (°c) 19 ph 8.12 conductivity ms.cm-1 54.85 tds (mg.l-) 39.55 dbo5 5.00 dco 10.00 alkalinity (mg.l-de caco3) 142.7 hardness (mg.lde caco3) 7158.4 ionic balance ca++ (mg.l-) 417.20 mg++ (mg.l-) 1487.93 na++ (mg.l-) 12165 k+ (mg.l-) 420.71 sr++ (mg.l-) 11.25 fe++ (mg.l-) 0.03 ag+ (mg.l-) 0.11 so4 2(mg.l-) 2,950.00 cl(mg.l-) 21,950.12 hco3 (mg.l-) 158.92 f(mg.l-) 1.25 co3 2(mg.l-) 8.62 bacteriological analyses total califorms 3/100ml fecal califorms 0 total streptococci 0 fecal streptococci 0 aerobic germs22°c 2/100ml salmonella abs pretreatment performance evaluation of the seawater dessalination plant… 3 copyright ©2023 assa adv. in systems science and appl. (2023) compared to other parameters such as turbidity or suspended solids, the silt density index (sdi) is key in controlling the fouling potential of seawater, and therefore, requires continuous monitoring [2]. to this end, the evolution of the sdi of the seawater from this desalination plant was monitored over a certain period. sdi tests were performed as per astm d 4189-95 standard by filtrating a water sample through a membrane of 0.45 μm (microfiltration) and 47mm of diameter at a constant transmembrane pressure (2.1 bars). the sdi is determined by comparing filtration times, t1 and t2 required to achieve a fixed filtration volume at time 0 and after time t respectively. sdi15 (t = 15 minutes) is defined by the astm as the time required for accurate and standardized testing. sdi15 (t = 15 minutes). figure 1 represents the variation of the raw seawater sdi3 of the desalination plant of beni saf as a function of time for 2019. fig.1. variation of the raw seawater sdi of the desalination plant of beni saf as a function of time for 2019 from the sdi3 versus time plot, one can see that the average sdi3 in autumn and winter is low with little evolution but then increases in spring and summer. 2.2. selection of pretreatment for the desalination plant of beni saf knowledge of the physico-chemical characteristics of the raw seawater requires the selection of the most adapted pretreatment technology to produce good quality supply water, and delay fouling of ro membranes. the pretreatment selected by the desalination plant of beni saf is a conventional physical and chemical treatment. figure 2 illustrates the different pretreatment stages used by the plant. 4 m.berrabah, h.bouabdesselam copyright ©2023 assa adv. in systems science and appl. (2023) fig .2. the pretreatment process used by the desalination plant of beni saf 2.2 performance evaluation of the pretreatment adopted by beni saf plant to monitor the performance of the pretreatment used in this plant, the pretreatment process was subdivided into three systems. then samples were taken downstream of each system (sampling points represented in the diagram fig.2) at regular time intervals, and the parameters most influenced at this stage were measured (sdi15, turbidity, temperature, ph). samples were taken during the daily the production over a one-year period (2019). system 1: description of the system injection of sodium hypochlorite into seawater (pre-chlorination) for pre-chlorination, the plant uses commercial sodium hypochlorite as a biocide. hypochlorite is dosed automatically and continuously into the seawater collection tank to remove organic matter. this operation is handled by three dosing pumps with a rate of 25%, i.e., 187.51 l/h under a pressure of 2 bars. sulfuric acid dosing concentrated sulfuric acid (98%) is dosed upstream of the sand filters, to set the ph and prevent precipitation of carbonates and bicarbonates on the ro membranes. this operation is handled by three dosing pumps with 25% rate, i.e., 83.75l/h under a pressure of 6 bars. ferric chloride dosing (coagulation) coagulation is the most popular treatment process used for the elimination of potential impurities such as colloidal and particulate matter. the purpose of coagulation is to combine small particles into larger aggregates/flocs (i.e. large groups of loosely bound suspended particles) by neutralizing the charges of the particles [6] . performed upstream of the sand filters, the coagulation serves to destabilize colloids and suspended solids so that these substances are retained during filtration. beni saf desalination plant uses 40% commercial iron chloride as coagulant. ferric chloride is automatically dosed into the discharge manifold of the seawater intake pump upstream of the sand filters at a rate that depends on the seawater flow rate. five dosing pumps with 25% rate, i.e., 83.75 l/h are dedicated to this operation. pretreatment performance evaluation of the seawater dessalination plant… 5 copyright ©2023 assa adv. in systems science and appl. (2023) sand filtration a filtration bed constituted of sand and silica with two different sizes (dual layer) is employed in this plant. this dual layer filtration is sufficient to achieve an sdi < 4 and effectively eliminate the pigments of algae if the installation is fed with good quality raw seawater [4]. the layer of coarse sand (base layer) at the bottom of the filter is topped by a finer sand layer (filtration layer). the 1000 mm thick filtration layer (upper layer) with its low granulometry serves to trap particles. the filtration sand has the following characteristics: table 2. characteristics of the filtration sand (upper layer) sio2 content >99,5 effective size (mm) 0.9 uniformity coefficient 1.4 absolute density (kg/m3) 2650 bulk density (kg/m3) 1600 grain size (micron) 139 the thickness of the coarse sand bed (base layer) is ~300 mm. the sand has the following characteristics: table 3. characteristics of the filtration sand (base layer) the turbidity analyzer console of each train includes a sampling device (sampling point 1) for the measurement of sdi and turbidity together with other parameters. system 2: description of the system anthracite filtration was added to this system. the anthracite filtration system consists of two filter trains comprising seven modules or filter pairs, making up a total of twenty-eight filters. each module comprises two identical filters operating as a single filtration unit. in normal operation, the flow rate required to feed the anthracite filter system is 18,000 m3/hour. the desalination plant of beni saf employs anthracite filters with a filtration bed of 1000 mm. the anthracite used has the following characteristics: table 4. the characteristics of the anthracite used in the desalination plant of beni saf rate of charcoal % >90 effective size (mm) 1.2 uniformity coefficient 1.4 absolute density 1400 bulk density (kg/m3) 1050 grain size (micron) 186 the turbidity analyzer console of each train comprises a sampling device, intended for measuring the rate of sdi matter, turbidity together with other parameters (sampling point 2). sio2 content >99,5 effective size (mm) 2 uniformitycoefficient 1.4 absolute density 2650 bulk density (kg/m3) 1600 grain size (micron) 309 6 m.berrabah, h.bouabdesselam copyright ©2023 assa adv. in systems science and appl. (2023) system 3: description of the system physico-chemical treatments were added to this system antisclants dosing this chemical treatment is carried out upstream of the cartridge filters. a commercial dispersant with a concentration of 95% diluted to 10% is dosed using 5 pumps with a flow rate of 25%, i.e., 56.75l/h under a pressure of 6 bars. sodium bisulphite dosing chemical treatment is performed downstream of the cartridge filters using commercial sodium bisulphate 61% diluted to 10%. the operation is handled by five pumps with a flow rate of 25%, i.e., 133.51 l/h under 6 bars of pressure for dosing sodium bisulphite cartridge filtration this is the first pretreatment stage aimed at retaining particulate <5 µg and ensuring the protection of the reverse osmosis installation. the filter cartridges used are made of polypropylene because of its low water absorption. each filter contains 380 cartridges type pp-5, with an outer diameter of 61 mm and a length of 1270 mm with a rating of 5 microns. at the end of this process, a connection for taking samples is placed on the filtered water outlet line to measure the sdi fouling index, turbidity and other parameters (sampling point 3). 3. results and discussion 3.1. study of the pretreatment temperature, ph, turbidity and sdi measurements were taken after each pre-treatment system for the year 2019 by computing the monthly average. the results are consolidated in the following tables. table.5. results of analyses after sand filtration at bwc plant for 2019 physical characteristics ph temperature (°c) sdi15 (%/min) turbidity standard 7-8 16-25 5-6 0.7-1 january 7.45 18.77 5.2 0.76 february 7.61 17.3 5.25 0.76 march 7.71 15.73 5.6 0.77 april 7.51 17.28 5.7 0.8 may 7.45 18.85 5.73 0.86 june 7.40 21.03 5.63 0.87 july 7.67 24.5 5.95 0.85 august 7.52 24.4 5.75 0.88 september 7.63 22.3 5.5 0.78 october 7.76 23.2 5.4 0.77 november 7.61 19.44 5.3 0.75 december 7.64 16.9 5.1 0.71 pretreatment performance evaluation of the seawater dessalination plant… 7 copyright ©2023 assa adv. in systems science and appl. (2023) figure 3 shows the evolution of these parameters as a function of time for the first system fig. 3. evolution of temperature, ph, sdi15, turbidity after the 1st filtration system of the plant for 2019. by comparing the results found and those of the raw seawater, a significant reduction in turbidity and sdi15 was observed with a difference of 10-15% and 0.7-1 ntu for the sdi15 and turbidity respectively, which shows the effectiveness of the system (coagulation+ sand filtration). the sand-based filtration aims at retaining suspended solids present in the seawater and flocs that formed at the preliminary stage of coagulation. the filter medium used by the plant is of two sizes (dual layer); a layer of coarse sand (base layer) located at the bottom of the filter is topped by a layer of finer sand (filtration layer) that serves to retain finer particles. an increase in sdi15 and turbidity was observed. this increase is caused by the fouling of the filtration bed by trapped particles. therefore, the filters must be washed when the head loss reaches a limit value, which results in a drop in filters output. table 6. analyses results after anthracite filter at bwc plant physical characteristics ph temperature (°c) sdi15(%/min) turbidity standard 7-8 16-25 44,5 0.70,5 january 7.45 18.77 44 0.65 february 7.61 17.3 4.45 0.6 march 7.71 15.73 4.3 0.53 april 7.51 17.28 4.25 0.55 may 7.45 18.85 4.15 0.50 june 7.40 21.03 4.13 0.52 july 7.67 24.5 4.1 0.51 august 7.52 24.4 4 0.5 september 7.63 22.3 4.2 0.58 october 7.76 23.2 4.39 0.6 november 7.61 19.44 4.5 0.68 december 7.64 16.9 4.45 0.7 8 m.berrabah, h.bouabdesselam copyright ©2023 assa adv. in systems science and appl. (2023) figure 4 shows the evolution of these parameters as a function of time for the 2nd filtration system. fig. 4. evolution of temperature, ph, sdi15, turbidity after the 2nd filtration system of bwc plant for 2019 figure 4 shows the filtration performance of the 2nd system in the elimination of sdi15 and turbidity. the sdi15 recorded after filtration by anthracite filters fares better than that recorded after a single filter (sand filter only). a discrepancy of 0.5-1.5 was observed between the two configurations (system 1 & 2).the second system shows that anthracite filtration operates as an adsorbent. the anthracite fulfils a number of functions among which the elimination of residual concentrations of oxidant agents such as chlorine and ozone, together with other carcinogenic derivatives resulting from the treatments. anthracite adsorbs these substances or reduces them by catalysis to benign forms it also traps organic matter, algae, pesticides and generally any compounds that may alter the odour or taste of the drinking water. therefore, the anthracite filters are necessary after the sand filters in order to produce water with good quality. table7. results of the analyses after the cartridge filters (pretreated water) at bwc plan for 2019 physical characteristics ph temperature (°c) sdi15(%/min) turbidity standard 78 1625 <4 <0.5 january 7.45 18.77 3.6 0.35 february 7.61 17.3 3.4 0.38 march 7.71 15.73 3.45 0.4 april 7.51 17.28 3.2 0.42 may 7.45 18.85 3.1 0.35 june 7.40 21.03 3 0.3 july 7.67 24.5 3.12 0.35 august 7.52 24.4 3.15 0.30 september 7.63 22.3 3.25 0.38 october 7.76 23.2 3.3 0.35 november 7.61 19.44 3.4 0.4 december 7.64 16.9 3.48 0.38 pretreatment performance evaluation of the seawater dessalination plant… 9 copyright ©2023 assa adv. in systems science and appl. (2023) figure 5 below shows the evolution of these parameters as a function of time for the 3rd system. fig. 5. evolution of temperature, ph, sdi15, turbidity after the 3rd filtration system of bwc plant for 2019. the recommended limit values for sdi and turbidity, upstream of the ro membranes, are < 4 (95 % of the time), and<0.5 ntu respectively [5]. figure 5 shows variation in turbidity and sdi15 as a function of time. as can be seen, the recorded sdi15 and turbidity values are below the recommended limits upstream of the reverse osmosis system, confirming thereby the efficiency of the pretreatment used by the desalination plant. filtration by cartridge filters is the final pretreatment stage; it serves to protect the highpressure unit as well as the ro membranes, and retain smaller particles that may cause damage to equipment. it also serves to reduce subsequently the sdi15 and turbidity to produce water ready for passage through the ro membranes. upstream of the cartridge filters, antiscalant dosing allows for avoiding salts precipitation on the ro membranes. down the cartridge filters, the dosing of sodium bisulphite allows for the elimination of residual chlorine resulting from the sodium hypochlorite employed for the disinfection of seawater. 3.2. study of the influence of the pretreatment used by the desalination plant of beni saf on the performance of reverse osmosis the desalination process is carried out by passing pretreated seawater through semipermeable membrane modules (hydranautics swc5). these modules require the following conditions (table 8). table 8. operating conditions conditions recommended values maximum pressure applied 82,73 bars maximum chlorine concentration < 0.1 ppm maximum operating temperature 113 °f (45 °c) ph range, continuous (cleaning) 2-11 (1-13)* maximum feed water turbidity 1.0 ntu maximum feed water sdi (15 min) 5.0 maximum feed flow 17.0 m3 /h maximum pressure drop for each element 1,03 bar 10 m.berrabah, h.bouabdesselam copyright ©2023 assa adv. in systems science and appl. (2023) these requirements must be fulfilled by the plant to ensure stable membrane performance. to investigate the influence of pretreatment on the performance of ro, variation of the pressure drop of both stages of line no. 1 together with its permeate flow rate were monitored. the following figures show the variation of permeate flow and head losses as a function of time (line 1). les measurements were taken during the daily production. fig .6. evolution of permeate flow as a function of time fig.7.evolution of head losses as a function of time as can be seen from figure 3, which illustrates permeate flow variation (line 1), there is a quasi-stability of the permeate flow during for january-march. for april, one can observe a decline in flow rate (~10%) attributable to the fouling phenomenon, which necessitates a chemical cleaning of the membranes. as shown in figure 7, the variation curve of head losses illustrates the stability of both the first and second stages for jan-march, which can be explained by the absence of the fouling phenomenon. by contrast, april sees an increase in head loss due to fouling. the results found show that albeit the seawater of the desalination plant of beni saf is properly pretreated, there is yet a persistent fouling. currently, a very noteworthy trend includes the use of membrane pretreatment (microfiltrationultrafiltration), to improve the quality of the pretreated water and consequently the performance of the ro installation. pretreatment performance evaluation of the seawater dessalination plant… 11 copyright ©2023 assa adv. in systems science and appl. (2023) 4. conclusion the main challenges faced by the desalination plant of beni saf is the rapid fouling of ro membranes associated to particulate/colloids (deposit), organic compounds (scaling) and to biological growth (biofilm). results of physico-chemical and bacteriological analyses on the raw seawater of the desalination plant show that this seawater necessitates a pretreatment to eliminate the materials responsible for fouling.poor feed water leads inevitalby to shorter life cycle of the ro membrane. for this reason, the desalination plant of beni saf employs a conventional pretreatment (physico-chemical treatment) to fight the formation of deposits on the ro membrane. our study shows that each stage of pretreatment used by the plant leads to significant improvement in the quality of the feed water. the first system shows that pre-chlorination, coagulation, and sand filtration reduces sdi15 by 10-15% /min whereas the second pretreatment system shows the effectiveness of the anthracite filtration, and a discrepancy in sdi15 between the two systems ( 1 & 2) of 0.5-1.5. whereas the last system shows the effectiveness of the filtration based on cartridges (0.5 µm) and addition of antiscalants.sdi15 and turbidity values were found to be in the range < 4 (95 % of the time), and <0.5 ntu, respectively. these results demonstrate the effectiveness of the pretreatment used by the desalination plant of beni saf. in spite of the effectiveness of the pretreatment employed by the plant, the fouling phenomenon persists, which requires an additional pretreatment process such as membrane pretreatment (desalination plant of maktaa) to improve the quality of pretreated seawater and increase the life cycle of ro membranes. references 1. hashlamon, a., ahmad, a. & choonhong, l. (2015). pre-treatment methods for seawater desalination and industrial wastewater treatment: a brief review. int. j. sci. res. sci. eng. technol., 1, 2394–4099. 2. kader, g. (2011). a large review of the pre treatment, expanding issues in desalination. [online] available http://www.intechopen.com/books/expandingissuesin-desalination/a-large-review-of-the-pre-treatment 3. legrand, c., solerieu, m., & goglio, e. (2006). traitement des sites et sols pollués [treatment of polluted sites and soils]. voiron, france: territorial ed., 85, [in french]. 4. leparc, j., rapenne, s., courties, c., lebaron p., croué, j. p., et al. (2007). water quality and performance evaluation at seawater reverse osmosis plants through the use of advanced analytical tools. desalination, 203, 243–255. 5. mosqueda, j., daniella, b. & huck, p. m., (2009). effect of biofiltration as pretreatment on the fouling of nanofiltration membranes. desalination, 245(1–3), 60– 72. 6. sinha, s., yoon, y., amy, g. & yoon, j. (2004). determining the effectiveness of conventional and alternative coagulants through effective characterization. schemes. chemosphere, 57, 1115–1122. 7. yoon, j., yoon, y., amy, g. & her, n. (2005). determination of perchlorate rejection and associated inorganic fouling (scaling) for reverse osmosis and nanofiltration membranes under various operating conditions. j. environ. eng., 131, 726–733. adv syst sci appl 2017; 3:1–8 published online at http://ijassa.ipu.ru/ojs/ijassa/article/view/502 auv thrust allocation with variable constraints v.v. kostenko1, a.yu. tolstonogov1,2∗ 1institute for marine technology problems feb ras, sukhanova st. 5a, vladivostok, 690091, russia 2 far eastern federal university, sukhanova st. 8, vladivostok, 690091, russia abstract: modern multipurpose autonomous underwater vehicles (auvs) represent the next generation of robotic systems with new technological tasks faced by researchers. one method of extending the functionality of the vehicle is installing tunnel thrusters in addition to the stern propulsion system. thereby the vehicle becomes able to undertake both survey-style missions and low speed interactions with the environment. but the efficiency of the tunnel thruster depends strongly on the vehicle’s velocity, due to hydrodynamic aspects. the usual solution to this problem uses different control models and thrust allocation methods for these types of mission. a unified approach to thrust allocation is presented in this paper. the approach is based on solving the allocation problem by quadratic programming with variable constraints, depending on the velocity of the auv and the thrusters’ angle of attack. the well-known implementation of active set method was used to model the proposed allocation method. keywords: auv, propulsion system, thruster hydrodynamics, control allocation problem, quadratic programming, interior point method, active set method 1. introduction the design of a control algorithm for an underwater vehicle is often divided into several levels (fig. 1.1) [1]. first, a high-level motion control algorithm is designed to compute a vector of virtual unbounded inputs to the vehicle fc (eq. (1.1)) from the target and current vehicle states and the control type. fig. 1.1. control system structure. fc is the vector of commanded inputs, f is the virtual vector of inputs, τ is the vector of allocated thrusts, u is the vector of low-level actuator inputs. ∗corresponding author: tolstonogov.anton@gmail.com http://ijassa.ipu.ru/ojs/ijassa/article/view/502 2 v.v. kostenko, a.yu. tolstonogov ẋ(t) = a(t)x +b(t)f y(t) = c(t)x (1.1) here, x(t) ∈ rn is the system state vector, t is time, y(t) is the vector of outputs, f ∈ rm is the vector of virtual inputs, which should be equal to the command vector fc of the high level motion algorithm, anda(t), b(t), c(t) are the coefficient matrices of the mechanical system. the virtual inputs are usually chosen to have a number of forces and torques that equals the number of degrees of freedom, m, that the motion control system wants to control. second, the control allocation algorithm is designed in order to map the vector of virtual input forces and torques fc to individual thruster forces τ such that the total forces and torques generated by all thrusters f amount to the commanded virtual input fc. f = b(x, t)τ (1.2) here, b(x, t) is the thruster configuration matrix. it contains the location and orientation of all thrusters, in the vehicle body-fixed frame. τ ∈ a ⊂ rp is the vector of actuator thrusts generated by the vehicle propulsion system, a are constraints from the saturation of the thrusters or other physical constraints, and p is the number of actuators. third, there is a separate high-frequency low-level controller for each actuator, that controls the desired thrust τi by a low-level control input ui. for each actuator, τi = hi(x, t, ui) (1.3) where h is a function, τi is the thrust of actuator i, and ui ∈ u ⊂ r is the low-level control input of actuator i. usually, the effector model is linear in u: τi = hi(x, t, ui) = g(x, t)ui (1.4) but the relation between the actuator thrust and the low-level control input is not considered in this article. we assume that each low-level controller maintains the desired thrust τi with sufficient quality. this modularity allows the high-level motion control algorithm to be designed without detailed knowledge about the vehicle propulsion system. in addition to coordinating the effects of the different thrusters in the system, issues such as thruster/fault tolerance, redundancy, and control constraints are typically handled within the control allocation module. in the case of an over-actuated propulsion system, when the number of thrusters is greater than the number of degrees of freedom (dof) controlled by the vehicle (p > m), the control allocation module solves an optimization problem to achieve a minimal power consumption of the propulsion system. there are different approaches to the high-level motion control of a vehicle. the pid (proportional-integral-derivative) is a widespread control method, but new methods such as linear quadratic gaussian (lqg) [2], the h∞ control method [3], or fuzzy logic control, are being rapidly developed. an interesting method is model predictive control [4], which solves a suboptimal motion control problem at each iteration, with the ability to combine the first and second levels of the motion, with the consequent increase in the computational complexity. in this paper, the problem of high-level control is assumed to be solved and the only the problem of optimal thrust allocation in the case of an over-actuated vehicle is considered. this problem is well studied. there is a survey devoted to this problem [1]. different approaches to this problem, including linear iterative approach as well as the quadratic programming approach satisfying optimality criteria, are treated in that survey. there are also papers published more recently [5–7]. these papers devoted to control allocation are focused on reducing the computational complexity of the optimal thrust allocation, but there is no mention of the fact that the thruster constraints can drift dynamically due to thruster hydrodynamics. a new approach, taking into account the vehicle speed and the thrusters’ angle of attack, is proposed in the present article. copyright c© 2017 assa. adv syst sci appl (2017) auv thrust allocation with variable constraints 3 this approach allows controlling both survey-style missions and low speed interactions with the environment for a vehicle equipped with tunnel thrusters. 2. problem formulation let the vector of virtual inputs computed by the high-level motion control or vehicle operator be denoted by the generalized force vector f = (fx, fy, fz,mx,my,mz) t , here fi is the projection of the force onto the axis i (i equals to the x, y and z axis) and mi is the projection of the torque onto the axis i in the body-fixed coordinate frame. the body-fixed reference frame [8] is used in this work. the x axis is directed along the longitudinal vehicle axis from the vehicle’s stern to the fore, the y axis is directed along the latitudinal vehicle to the starboard, and the z axis completes the frame to a right-handed coordinate system. assume that the system is equipped with p thrusters with control thrust τi(i = 1, . . . , p). this leads to the following relation between τ = (τ1, τ2, . . . , τp) and the virtual inputs f : bτ = f (2.5) according to [9], the optimization problem can be written as min τ,s 1 2 (τtqτ + strs) bτ = f + s (2.6) τmin ≤ τ ≤ τmax where s is the vector of slack variables used to penalize |bτ− f |, q and r are the weight diagonal matrices for the thrusters and dofs, respectively, τmin and τmax are the vector of thrust constraints according to bollard pull tests. the thrust allocation of an overactuated auv with tunnel vertical thrusters is a velocitydependent problem. for example, the pitch motion of the vehicle on zero velocity is better created by tunnel thrusters due to their large thrust arm. but the effectiveness of a tunnel thruster tends to zero at high velocities (fig. 2.2). we propose a variable constraints method for the thrust allocation problem. this method allows reallocating the thrust depending on the vehicle velocity. the method considers thrust constraints that depend on the vehicle velocity, as well as the thrusters’ angle of attack: τmin(v, θ) ≤ τ ≤ τmax(v, θ) where τmin(v, θ) and τmax(v, θ) are the variable thrust constraints depending on the vehicle velocity and the thrusters’ angle of attack. 3. the model of thruster constraints the thrust and torque generated by a thruster can be described by the following formulas: τ = kτ(j0)ρω |ω|d4 m = km(j0)ρω |ω|d5 where ρ is the density of water, ω is the rotational speed of the thruster propeller, andd is the diameter of the propeller.kτ(j0), km(j0) are the coefficients of thrust and torque determined copyright c© 2017 assa. adv syst sci appl (2017) 4 v.v. kostenko, a.yu. tolstonogov (a) v < 0.7ms−1 (b) v > 0.7ms−1 fig. 2.2. thrust allocation for pitch motion depending on vehicle velocity. by the form of the propeller and the characteristics of the motor. j0 = v/ωd is the advance ratio (where v is a velocity of thruster input flow, this velocity equals to velocity of the vehicle in case of the absence of currents). the function kτ(j0) can be fitted as a linear function of j0 at j0 > 0 and v > 0 [10]: k+ τ (j0) = k0 τ (0)− a1j0 (3.7) where k+ τ (j0) is the thrust coefficient kτ(j0) at j0 > 0 and equi-directional state of ambient flow and axial thruster flow. k0 τ is the thrust coefficient at j0 = 0 and a1 is a coefficient determined by the shape of the propeller. the equation may be rewritten as a function of velocity by fixing the value of ω: kτ(v) = k0 τ − b1v (3.8) where b1 = a1/dω. the thrusters’ efficiency also depends on the thrusters’ angle of attack θ. at small angle of attack (θ < θ∗1, where θ∗1 is the first critical incoming angle [10]), kτ(v, θ) can be written as kθ τ (v, θ) = k0 τ + ( k0 τ −k+ τ (v) ) [ sin ( θ θ∗1 π 2 ) − 1 ] (3.9) where kθ τ (v, θ) is the thrust coefficient at vehicle velocity equals to v and the thrusters’ angle of attack equals to θ,k+ τ (v) is determined by equation 3.8, and θ∗1 is the first critical incoming angle: θ∗1 = π/2− a1v. during bollard pull tests (at v = 0), the thrust of the actuators can be obtained as a function of the input code or the current. the maximum thrust of the thruster can be written as a function of kτ: copyright c© 2017 assa. adv syst sci appl (2017) auv thrust allocation with variable constraints 5 τmax = k0 τρωmax ∣∣ωmax ∣∣d4 (3.10) where τmax is the maximum thrust generated by the thruster during the bollard pull test and ωmax is the number of revolution per second of the propeller at maximum thrust. hence, the maximum thrust τmax for a stern thruster at vehicle velocity v and the thrusters’ angle of attack θ can be obtained from equations (3.7), (3.9) and (3.10): τmax(v, θ) = τmax − cv [ 1− sin ( θ θ∗1 π 2 )] (3.11) where τmax(v, θ) is the maximum thrust generated by the thruster at vehicle velocity v and the thrusters’ angle of attack θ, c = b1/ρωmax ∣∣ωmax ∣∣d4. more complicated cases, when j < 0 or v < 0, can be obtained from the appropriate equations [10]. the steady state performance of a tunnel thruster can be fitted as an exponential function of the velocity ratio v/vjet [11]: τ(v) = τ(0)exp [ −c ( v vjet )2 ] where τ(v) is the the resulting force of the tunnel thruster for a vehicle velocity of v, τ(0) is the thruster’s performance in the bollard pull test, and vjet is the axial velocity of flow generated by the thruster: vjet = √ τ(0) ρs . (3.12) here, ρ is the density of water and s is the cross sectional area of the tunnel thruster. hence, the tunnel thruster constraints can be written as τ(v)max = τ(0)maxexp [ −c ( v vjet,lim )2 ] (3.13) where τ(v)max is the maximum thrust when the vehicle’s motion has velocity v and τ(0)max is the thruster constraint provided by the bollard pull tests of the thruster. 4. propulsion system setup the model of the auv “mt-2012” [12] propulsion system was used for algorithm simulation. the propulsion system consists of five thrusters: four thrusters located at the stern of the vehicle at an angle of α to the longitudinal axis (x axis), and a vertical tunnel thruster located at the forward part of the vehicle (fig. 4.3). the propulsion system of the vehicle can be described by the thrusters’ configuration matrix b [13]: b = [bb, bl, bu, br, bf ] where bd, bl, bu, br are column vectors for (down, left, up, right) the stern thrusters and bf is the column vector for the forward tunnel thruster. the column vectors in four dofs (surge, heave, pitch and yaw) take the form described below. copyright c© 2017 assa. adv syst sci appl (2017) 6 v.v. kostenko, a.yu. tolstonogov fig. 4.3. propulsion system of “mt-2012”. for the up and down stern thrusters: bu =  cosα − sinα 0 luz cosα− lux sinα  , bd =  cosα sinα 0 ldz cosα− ldx sinα  . for the left and right stern thrusters: bl =  cosα 0 lly cosα− llx sinα 0  , br =  cosα 0 lry cosα− lrx sinα 0  . for the vertical tunnel thruster: bf =  0 1 −ltunnelx 0  where ri = [lix, l i y, l i z] is the arm vector of thruster i (where i = u, d, l, r, f ). for this propulsion system, for reasons of symmetry, luz = lry = −ldz = −lly = lsternzy and lux = ldx = llx = lrx = lsternx . the parameters of the propulsion system are shown in table 4.1. table 4.1. parameters of the auv “mt-2012” propulsion system lsternx , m 1.88 α, grad 22.5◦ lsternzy , m 0.23 f sternlim , n 120.1/-68.2 ltunnelx , m 1.20 f tunnellim , n 122.1/-122.0 5. simulation setup the proposed allocation control method with variable constraints was tested on depth maneuvering at different velocities, 0.3, . . . , 1.8 ms−1 (corresponding to forward forces 3.6, . . . , 129.6 n, and 50 nm of pitch torque. copyright c© 2017 assa. adv syst sci appl (2017) auv thrust allocation with variable constraints 7 the optimal thrust allocation problem for different velocities was solved by matlab optimization toolbox. since q > 0 and r > 0 (equation (2.6)) this is a convex quadratic problem in τ, so the interior point method or active set method for convex optimization problem can be used. the results of the thrust allocation for different vehicle velocities are shown in fig. 5.4. fig. 5.4. thrust allocation for different velocities. there are some implementations of quadratic programming solvers in the c/c++ languages. these implementations are faster than the matlab version, and are easy to integrate in vehicle software. there are different types of quadratic programming solvers: qpoases implements the active set method [14], which allows using a hot start to reduce the calculation time for each new iteration using the data of the previous one, and there are several libraries with implementations of the interior point method: [15] and [16]. 6. conclusion to create a multi-purpose auv capable of undertaking both survey-style missions and low speed interactions requires an over-actuated propulsion system and a sophisticated thrust allocation algorithm. one method of extending the functionality of the vehicle is installing tunnel thrusters to the vehicle propulsion system. a new approach for thrust allocation algorithm is proposed in the present article. the method based on the quadratic cost optimal problem with constraints. the main feature of the method is that thrust constraints depend on the vehicle velocity, as well as the thrusters’ angle of attack. this method takes into account the decrease in the thrust of tunnel and stern thrusters at high velocities, and reallocates the thrust accordingly. the proposed algorithm was tested with a simulation model of the real auv “mt-2012” propulsion system. the propulsion system consists of five thrusters: four thrusters located at the stern of the vehicle and the vertical tunnel thruster located at the forward part of the vehicle. the result shows the adaptive thrust reallocation depending on vehicle velocity in case of depth maneuvering. copyright c© 2017 assa. adv syst sci appl (2017) 8 v.v. kostenko, a.yu. tolstonogov references 1. fossen, t. i. & johansen, t. a. (2013) control allocation—a survey. automatica, 1, 1087–1103. 2. naeem, w., sutton, r., & chudley, j. (2006) soft computing design of a linear quadratic gaussian controller for an unmanned surface vehicle. in 14th mediterranean conference on control and automation, palermo, italy, 1–6. 3. moreira, l. & guedes soares, c. (2008) h2 and hinf designs for diving and course control of an autonomous underwater vehicle in presence of waves. ieee journal of oceanic engineering, 33, 69–88, https://dx.doi.org/10.1109/icsmc.2012.6378129. 4. yan, z., chung, s. f., & wang, j. (2012) model predictive control of autonomous underwater vehicles based on the simplified dual neural network. in 2012 ieee international conference on systems, man, and cybernetics (smc), seoul, korea, 2551–2556, https://dx.doi.org/10.1109/icsmc.2012.6378129. 5. pham, c. d., spiten, c., & from, p. j. (2015) a control allocation approach to haptic control of underwater robots. in 2015 ieee international workshop on advanced robotics and its social impacts (arso), lyon, france, 1–6, https://dx.doi.org/10.1109/arso.2015.7428209. 6. saback, r., cesar, d., arnold, s., lepikson, h., santos, t., & albiez, j. (2016) fault-tolerant control allocation technique based on explicit optimization applied to an autonomous underwater vehicle. in oceans 2016 mts/ieee monterey, monterey, ca, 1–8, https://dx.doi.org/10.1109/oceans.2016.7761251. 7. huang, h., zhang, g., yang, y., xu, j., li, j., & wan, l. (2016) thrust optimal allocation for broad types of underwater vehicles. in advances in swarm intelligence: 7th international conference, icsi 2016, bali, indonesia, 491–502, https://dx.doi.org/10.1007/978-3-319-41009-8 53. 8. ben, m. & lee, t. h. (2011) unmanned rotorcraft systems. berlin, germany: springerverlag. 9. johansen, t., fossen, t., & tøndel, p. (2005) efficient optimal constrained control allocation via multiparametric programming. journal of guidance, control, and dynamics, 28, 506–515. 10. kim, j. (2009) thruster modeling and controller design for unmanned underwater vehicles (uuvs). in inzartsev, a. v. (ed.), underwater vehicles, (pp. 235–250), intechg, https://dx.doi.org/10.5772/6705. 11. palmer, a., hearn, g., & stevenson, p. (2008) modelling tunnel thrusters for autonomous underwater vehicles. ifac proceedings volumes, 41, 91–96. 12. matvienko, y. v., boreyko, a. v., lvov, o. y., & vaulin, y. v. (2015) kompleks robototehnicheskih sredstv dlja vypolnenija poiskovyh rabot i obsledovanija podvodnoj infrastruktury na shel’fe [complex robotic tools to perform searches and surveys of underwater infrastructure on the shelf]. podvodnye issledovaniya i robototekhnika, 1, 4–15, [in russian]. 13. fossen, t. (2011) handbook of marine craft hydrodynamics and motion control. chichester, uk: john wiley & sons, ltd. 14. ferreau, h., kirches, c., potschka, a., bock, h., & diehl, m. (2014) qpoases: a parametric active-set algorithm for quadratic programming. mathematical programming computation, 6, 327–363. 15. gertz, e. m. & wright, s. j. (2003) object-oriented software for quadratic programming. acm transactions on mathematical software, 1, 58–81. 16. yan, y. (2014), c++ implementation of the interior point methods (cpipm). [online]. available https://github.com/yimingyan/cppipm. copyright c© 2017 assa. adv syst sci appl (2017) https://dx.doi.org/10.1109/icsmc.2012.6378129 https://dx.doi.org/10.1109/icsmc.2012.6378129 https://dx.doi.org/10.1109/arso.2015.7428209 https://dx.doi.org/10.1109/oceans.2016.7761251 https://dx.doi.org/10.1007/978-3-319-41009-8_53 https://dx.doi.org/10.5772/6705 https://github.com/yimingyan/cppipm introduction problem formulation the model of thruster constraints propulsion system setup simulation setup conclusion microsoft word 729 concept of a universal smart machine with protection against spark discharge adv syst sci appl 2020; 01; 1-12 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/729 concept of a universal smart circuit breaker with protection against spark discharge andrey i. vlasov1*, sergei s. filin2, alexey i. krivoshein1,2 1) bauman moscow state technical university, moscow, russia 2) llc “konnekt”, moscow, russia abstract: the concept of a universal device that can provide diagnostics and monitoring of individual and collective power supply systems, the construction of life support systems, the state of household appliances, the internet of things devices and housing and communal services devices is considered. this paper proposes new design solutions able to improve the existing technologies for assessment of the quality of the energy flow by two orders of magnitude in terms of overall characteristics relative to the existing analogues. the developed smart microcontroller is designed to operate as an intelligent unit, to analyze the parameters of the ac voltage network and control the electromechanical elements of the smart circuit breaker, as well as to send data through a standard interface. the implementation of protection against spark discharge is proposed. the proposed solution is based on the application of cloud and boundary computing technology. the possibilities for improvement of the efficiency of design technologies and construction of energy-efficient buildings, provision of the industry and households with a new type of information management services (monitoring energy infrastructure, etc.) are presented. keywords: smart circuit breaker, monitoring, control, smart networks, power supply systems, internet of things. 1. introduction the creation of a technological basis for the development of innovative housing infrastructure that provides diagnostics and monitoring of the energy status of the infrastructures, individual and collective energy supply systems, building life support systems, monitoring of the state of household appliances, internet of things (iot) devices and housing and utilities infrastructure is a very relevant task and corresponds to the goals of the energy.net roadmap [12, 18]. this process is accompanied by the development of technological solutions in the field of energy status monitoring and predictive maintenance and repair [2, 3, 5, 14, 27-29]. one of the directions for the implementation of the roadmap action plan is the transition of the electrical support systems to the iot concept [20-22, 25, 26, 28]. at the moment, there are many solutions in the market aimed at protecting the household network against various malfunctions [1, 10, 12, 18, 23]. however, in general, they implement only certain types of protection. there are practically no solutions to implement comprehensive protection, including the following functions: overload and short circuit protection; current leakage protection; spark protection; overvoltage or overload protection; determining the cause of the shutdown after the automatic triggering; sending data to an external device; ability to shut off the circuit breaker at the command of the remote user. some companies have the solutions, partially performing these functions, in the form of several individual devices. for example, the solution of the company abb (https://new.abb.com) has the form of 5 separate devices, with a total width of 11 standard * corresponding author: vlasov@iu4.ru 2 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) modules. the solution based on schneider electric devices (https://www.se.com) represents 5 separate devices, with a total width of 7 standard modules, and legrand (https://legrand.ru/) has 4 separate devices in 7 standard modules. the installation of a system of several devices imposes additional requirements on the qualifications of electricians. it is obvious that modern requirements for the qualitative characteristics of energy consumption require wider use of intelligent controls and predictive services, which can be distinguished by the complexity of their functions and would be implemented in a publicprivate partnership [6, 8]. to solve this problem, the goal was set to develop a universal smart circuit breaker with protection against spark discharge, the concept of which is presented in this study. 2. power protection requirements power quality takes into account all aspects of electromagnetic compatibility, but characterizes only the electrical network. the permissible levels of electromagnetic compatibility established for it are called power quality indicators (pqi). the regulatory values of the pqi and the list of them are set by the standards (in the european union – en 50160, iec 61000-4-30: 2008; in russia – gost r 54149, gost 13109-97, etc.), which are the guidance for the developers of the electrical equipment connected to the network, in terms of their noise immunity, on the one hand, and the level of interference, on the other. if the noise immunity level of these technical devices is higher than the maximum permissible pqi values for the network, electromagnetic compatibility (emc) will be ensured [10]. the actual pqi values should be monitored using specialized measuring instruments under operating conditions, and the corresponding characteristics of emc – by the necessary tests during their development and production. indicators of power quality can be divided into three groups. the first group includes frequency deviations and voltage deviations, which are associated with the peculiarities of the process of production and transmission of electricity. the quality of frequency and voltage deviations regulation determines their level in the electric power system. the second group can be attributed to the pqis, characterizing the non-sinusoidal shape of the voltage curve, the asymmetry and voltage fluctuations. the sources of these distortions (emitters) are mainly electrical receivers. in order to take into account the impact of electromotive forces generated by such electrical receivers, it is necessary to take technical measures both at the stage of development and production and during their operation. the third group includes the pqis characterizing random electromagnetic phenomena and electrical processes, which are inseparably connected with the technological process of production, transmission, and consumption of electricity. these include voltage dips, overvoltage, and voltage pulses. the pqis of the first two groups are standardized, and two permissible levels are set for them: normal and limit. the pqis of the third group are not standardized, but the statistical information about them is of great importance for the normal operation of the electric power system. the frequency f is a system-wide parameter of the sos mode and is determined by the active power balance. in the event of a shortage of generated power in the system, the frequency is reduced to a value at which a new balance of generated and consumed power is established. on the contrary, with an excess of generated power, the frequency increases. the voltage deviation from the nominal value in the mode of the maximum (δu2 max) and minimum (δu2 min) load may differ from the allowable values. in the electrical installation rules, it is recommended to maintain the voltage in the cpu at a level not lower than 105% of the nominal under the highest load and not higher than 100% – at the lowest load. this requirement meets the principle of counter voltage regulation. the means of voltage regulation are used for the implementation. if voltage deviations are created under the influence of relatively slow load changes determined by its schedule, then rapid load changes create voltage fluctuations. voltage fluctuations are determined by the envelope of the concept of a universal smart circuit breaker with protection against spark discharge 3 copyright ©2020 assa adv. in systems science and appl. (2020) effective or amplitude voltage values and are characterized by the span δut and the frequency of repetition of voltage changes or the intervals between voltage changes. a significant share of the load in the electrical network is represented by the eds with nonlinear current-voltage characteristic. such eds consume current, the form of which differs significantly from the sinusoidal one. 3. development features of the smart circuit breaker 3.1. concept of a universal smart circuit breaker due to the lack of an integrated solution that implements the listed functions, a conceptual solution of a universal smart circuit breaker is proposed (fig. 1). when designing the circuit breaker, modern requirements for the protection of consumer and industrial electronics are taken into account. the functions of the circuit breaker include:  thermal overload protection;  electromagnetic protection against overcurrent and short circuit;  protection against leakage (triggered when a difference in phase and neutral currents occurs);  protection against the occurrence of an arc on the line (triggered when an arc of more than 0.15 seconds in duration appears);  shutdown by external user command (via rs-485 interfaces);  measurement of current and power consumed by the load;  measurement of voltage on the line, as well as the fundamental frequency and frequency spectrum of the network;  saving information about the reasons for shutting down the circuit breaker and sending data to external devices (via the rs-485 interface, since the device is designed to be used in conjunction with automation controllers, which typically have an rs-485 interface; if necessary, it is possible to transfer data from a smart circuit breaker via the wi-fi interface through the automation and monitoring controller). fig. 1. 3d model of a smart circuit breaker fig. 2. 3d model of the internal system of the smart circuit breaker the internal system of the smart circuit breaker, according to the principle of operation, is conditionally divided into two functional units (fig. 2): a) an electromechanical unit, including electromagnetic and thermal protection mechanisms, as well as transformers acting as current sensors; and b) an electrical unit including an analysis board with an intelligent controller installed on it (intelligent control unit). 4 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) in the electromechanical unit during development it is advisable to implement the functions of thermal overload protection, electromagnetic short-circuit protection, current leakage protection function, an electrical circuit disconnection system, and a current sensor for measuring the phase current. the electrical unit must perform the functions of a smart circuit breaker that are impossible or impractical to implement in the form of electromechanical automation. the intelligent electronics unit must perform data processing functions from the phase current sensor and transformer to measure the differential current, measure the voltage at the input and output of the smart circuit breaker, determine the network parameters (such as power consumption, frequency, sine wave distortion level), switch off the circuit breaker in the case of overvoltage or arc striking on the line, send data to an external device and control the display. a block diagram of a universal smart circuit breaker is shown in fig. 3. fig. 3. block diagram of the smart circuit breaker: 1 – 220 v mains; 2 – electromechanical unit; 3 – analysis module; 4 – electromechanical module; 5 – 200 v input; 6 – thermal protection unit; 7 – short circuit protection unit; 8 – differential current sensor; 9 – current in phase sensor; 10 – relay with arc suppression; 11 – output 220v; 12 – load; 13 – voltage divisor with offset; 14 – adc; 15 – power supply; 16 – uart; 17 – pwm outputs; 18 – inputs; 19 – control output; 20 – uart-rs485 converter; 21 – rgb-indication; 22 – control buttons; 23 – amplifier; 24 – rs-485 interface. considering that the data processing and decision-making functions in the electronics unit are implemented by the microchip, the functions of the analysis board are reduced to ensuring the functioning of the chip, converting signals from the electromechanical unit to acceptable levels for the microcontroller, and converting the control signal of the microcircuit to shutting off the electromechanical unit. in addition, on the analysis board, the indication elements of the state of the circuit breaker should be installed, and buttons for manual control of the smart circuit breaker. the necessary functions of the analysis board include: 1. conversion of an alternating analog signal from the ca input (peak voltage is 400v) to a proportional analog signal in the range of 0-uadc (with an offset of 0.5uadc), where uadc is the maximum voltage of the microcontroller adc; 2. the conversion of the ac analog signal from the ca output (peak voltage is 400v) to the analog signal in the range of 0-uadc (with an offset of 0.5uadc); concept of a universal smart circuit breaker with protection against spark discharge 5 copyright ©2020 assa adv. in systems science and appl. (2020) 3. conversion of an alternating analog signal from a phase current sensor (peak voltage is 7v) to an analog signal in the range of 0-uadc (with an offset of 0.5uadc); 4. conversion of an alternating analog signal from a differential current sensor (peak voltage is 7v) to an analog signal in the range of 0-uadc (with an offset of 0.5uadc); 5. amplification of the control signal of the ic to a power sufficient to trigger the electromagnetic pusher and turn off the smart circuit breaker; 6. the presence of an indication that allows displaying the status of the circuit breaker and the reason for shutdown; 7. the presence of a button for manual control of the mode of operation of the smart circuit breaker, and a button for resetting the electrical unit of the smart circuit breaker; 8. the presence of a converter interface data output of the microcontroller to the rs-485 interface. such a set of functions imposes increased requirements on the element base, which makes it expedient to create a specialized “smart controller” as a special ic. 3.2. approach to the implementation of protection against spark an important component of the electrical protection of the housing infrastructure is protection against the occurrence of a spark or arc, both longitudinal and transverse. both types of arc result in heating of the conductor, melting of the insulation, and may cause a fire. at the same time, none of the standard means of automatic protection, such as a circuit breaker, or a differential circuit breaker, can detect such a fault. there are specialized means of protection against the occurrence of an arc in a line, the so-called afdds (arc fault detection devices). unlike standard circuit breakers based exclusively on electro-mechanics, almost all afdd devices are designed as a combination of an electrical board analyzing network parameters and a network disconnecting switch triggered by command of an electrical board [9, 17]. recently, relatively cheap afdd devices have begun to appear, but their sensitivity remains controversial. based on literature review [15, 16, 17, 26, 27, etc.], one can state that currently, most developers use simplified mechanisms to generate sparks: interrupted circuits based on mechanical contact opening, twisted wires, etc. in some cases, a spark generation unit with the graphite-copper contact is used to demonstrate the operation of the spark detection device. however, the sparking of graphite-copper contact is very different from the sparking of the copper-copper contact under real-life conditions. this paper proposes a new approach based on a special device that provides automatic spark generation at the copper-copper contact; an experimental test bench was built with a spark generating device between conductors (fig. 4). 6 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) fig. 4. device generating a spark in the composition of the experimental stand the proposed solution provides for fixing two wires and ensuring their contact using a rotary mechanism. then, voltage is applied and the adaptive spark control algorithm is initiated, which controls the interruption and restoration of the contact based on the current sensor readings in order to ensure maximum possible sparking. the experiments with a spark under various types of load have shown that when the arc occurs, a sinusoid of alternating current is distorted and the noise-like signals are generated. spark detection is based on analyzing the network spectrum of a signal by current and voltage. the occurrence of a spark generates a noise-like distortion of the network, often of small amplitude, especially when a longitudinal spark occurs. due to the high electrical capacitance and inductance of household wiring, the amplitude of the signal generated by the spark weakens with increasing distance from the point where the spark originated. also, the signal weakens when the active or impulse load in the network is connected between the location of the spark and the location of the afdd device [4, 15, 16]. the determination of the presence of the arc on the line is carried out through the analysis of the spectrum of the network, as well as the analysis of the shape of the sinusoid ac current. when the arc occurs, the breakdown occurs at a certain phase of the sinusoid of the mains voltage, then the load increases sharply and the sinusoid is distorted. this causes the occurrence of high-frequency harmonics, which can be detected by continuously analyzing the spectrum of the network. in this case, another problem arises due to the absence of a difference between signal distortions caused by the appearance of an arc on the line before and after the circuit breaker. the solution to this problem is the introduction of a filtering element into the circuit of an automaton that changes the amplitude of high harmonics. by the difference of the signal before and after such an element, it will be possible to determine which side of the circuit is the source of high harmonics. if the arc originated from the voltage source, the level of high harmonics from the load side will be lower. in the case of an arc on the load side, the level of harmonics will be lower on the side of the voltage source. in the simplest case, such an element could be an inductor. the measurement of the amplitude of high harmonics before and after the inductance, taking into account the unpredictability of their composition, is made by measuring the amplitude of the signal passed through the hf filter and the half-wave rectifier. the calculation of the inductance of the input coil, necessary for the possibility of determining the difference of the amplitudes of the hf signal before and after the filter element, is carried out as follows. since the voltage difference at the outputs of the rectifiers concept of a universal smart circuit breaker with protection against spark discharge 7 copyright ©2020 assa adv. in systems science and appl. (2020) is actually measured, it is first necessary to determine the minimum step of the adc of the microcontroller. the resolution of the microcontroller adc is 12 bits, the range of the measured voltage is 0-1v. calculate the minimum step by the formula (1): n 12 u 1 0.244 2 2 adc stp v u mv   (1) where stpu is the minimum measurement step for the adc, uadc is the upper limit of the adc measurement range, n is the precision of the adc. the noise level at the output of the rectifier, determined empirically, does not exceed 0.1 mv, the error level of the adc is 0.5 minimum step. i.e. the useful step of the adc is approximately 0.5 mv, or 0.05% of the range measured by the microcontroller voltage. the arc can occur both in the presence of a load and in its absence. in this case, the arc itself will be the load, the value of which will depend on the distance between the conductors on which the arc originated. the drop in the amplitude of the rf signal at the coil will depend on the value of the coil impedance and the load, i.e. on the load level and inductance value. the greater the impedance of the load, the smaller the drop in rf on the coil, and hence the difference in the signal level. therefore, one can calculate the required inductance value at the lowest possible load. the occurrence of the arc, according to the source, is possible at a voltage not lower than 10 v and a current of at least 0.1 a. based on these parameters, one can calculate the required resistance of the filter element (2-4): 0.05%induct induct r x x x   , (2) 0.0005 0.9995induct xr x  , (3) 0.0005induct rx x  (4) with a minimum current of 0.1 a and a voltage of 250 v, the total resistance of the arc is (formula 5): 2500 max r min u x om i   (5) then 1.25 inductx om the higher the average frequency of the arc spectrum, the greater the inductance resistance and the greater the voltage difference after rectification. let the minimum average frequency of harmonics f = 1 khz. the resistance of inductance inductx is determined by the formula 6: 2inductx fl (6) then the inductance is calculated by the formula 7: 200 2 ix l µh f    (7) thus, in order to determine the position of the source of the arc, it is necessary to include in the circuit of the automaton an inductance coil at least l = 200 µh. 8 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) 4. results modern comprehensive power grid protection should include: cable overheating protection, short circuit protection, sparking protection, protection against significant voltage deviation, and current leakage protection. the monitoring function is designed to control parameters such as voltage, current strength, power, and electricity consumption. in addition, monitoring in smart security systems should make it possible to determine the state of the system and the cause of the shutdown. the concept of implementing an intelligent smart circuit breaker for the intellectualization of the energy sector infrastructure includes two main functional blocks: an integrated power grid protection unit and a unit for monitoring the power grid parameters. fig. 5 shows the interaction of components of the integrated power system of the russian federation [31]. individual retail companies interact with end-users; they ensure the accounting and management of end consumers’ operation modes. the aim of the work is to develop a concept of implementing universal, intelligent smart circuit breakers, interacting with specific consumers, which provide additional functions to monitor the state of the power grid and provide recommendations for predictive repair. fig. 5. functional scheme of integration of smart circuit breakers into intelligent energy infrastructure by the example of the russian federation the proposed smart circuit breaker can be used anywhere, both in private households and in the manufacturing sector. this device provides comprehensive power grid protection and occupies a small part of the electrical control panel. in addition to the basic protection functions, the smart circuit breaker makes it possible to facilitate keeping records of electricity consumption and the quality of power supply, as well as to reduce the number of emergencies due to deeper diagnosis. the advanced functionality of the proposed smart circuit breaker is ensured by using a new specialized microcontroller (smc). concept of a universal smart circuit breaker with protection against spark discharge 9 copyright ©2020 assa adv. in systems science and appl. (2020) the use of the proposed smc ensures the functioning of an intelligent unit for analyzing the parameters of the ac voltage network and controlling the electromechanical elements of the smart automat, as well as sending data through a standard interface. the smc as part of the smart circuit breaker allows measurement and analysis of the following parameters:  ac input voltage;  current in phase;  difference of phase and zero currents;  power consumed by the load connected via a smart circuit breaker;  frequency spectrum of alternating voltage (analysis of occurrence of a spark). switching off the smart circuit breaker on a command from the smc occurs through a signal to its control outputs connected to the electromechanical controls of the smart circuit breaker. the logic of switching off the automaton is given by the algorithm of its operation. the use of a special microcontroller allows reducing the dimensions of the analysis board. also, the use of an smc makes it possible to reduce the prime cost of a smart circuit breaker in the case of mass production. 5. conclusion the proposed solution is based on the application of cloud and boundary computing technology, and fully allows for remote diagnostics and monitoring of the energy status of the light and digital infrastructures, individual and collective energy systems, life support of buildings, the state of household appliances, iot devices and housing and communal services. this underlines the ability of the smart circuit breaker to be widely used [7, 11, 13, 19, 24, 27, 30]. as part of the technological solution, this device provides comprehensive protection of the electrical network, while occupying a small part of the electrical panel, and surpasses the existing analogs. in addition to the basic protection functions, the smart circuit breaker will make it easier to record data on electricity consumption and the quality of power supply, and reduce the number of emergency situations by simplifying their diagnosis. it can be stated that innovative smart circuit breakers, the operation of which is based on the principles of the internet of things, become an important aspect of social infrastructure. the creation of specialized controllers should ensure at the minimum expense the implementation of monitoring data analysis functions, the transfer of smart circuit breaker data through the service interface to external automation controllers. this paper proposes an implementation concept and recommendations on the configuration of a specialized controller for a smart circuit breaker, designed to analyze the parameters of an ac network. according to the results of the implementation of the technological basis for the development of innovative social infrastructure based on smart circuit breakers using a specialized microcontroller, it can be argued that this approach simplifies development by reducing energy and resource consumption, and the selected controller configuration is the most efficient and fully complies with the requirements imposed on innovative smart circuit breakers. directions of further research will be focused on improving the computational efficiency of monitoring algorithms and forecasting the state of the power grid with the possibility of implementing the predictive repair functions. acknowledgments the research was conducted with the support of the ministry of science and education of russia within the framework of the project under the agreement no. 14.579.21.0158, id rfmefi57918x0158. 10 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) references 1. akberdina, v.v., tretyakova, o.v., & vlasov, a.i. (2017). a methodological approach to forecasting spatial distribution of workplaces in an industrial metropolis, problems and perspectives in management, 15(4(87), 50-61. https://doi.org/10.21511/ppm.15(4).2017.05 2. amruthnath, n., & gupta, t. (2018). a research study on unsupervised machine learning algorithms for early fault detection in predictive maintenance. 2018 5th international conference on industrial engineering and applications, ieee, 355-361. http://dx.doi.org/10.13140/rg.2.2.28822.24648 3. andreev, k.a., vlasov, a.i., & shakhnov, v.a. (2016). silicon pressure transmitters with overload protection, automation and remote control, 77(7), 1281-1285. https://doi.org/10.1134/s0005117916070146 4. bain, j., sankar, l., potsdam, m., & brentner, k.s. (2010). aeromechanic and aeroacoustic predictions of the boeing smart rotor using coupled cfd/csd analysis. annual forum proceedings – ahs international 66th forum of the american helicopter society: “rising to new heights in vertical lift technology”, ahs forum 66, phoenix, az, 2013-2030. 5. baptista, m., sankararaman, s., de medeiros, i.p., nascimento, c., prendinger, h. et al. (2018). forecasting fault events for predictive maintenance using data-driven techniques and arma modeling, computers & industrial engineering, 115, 41-53, http://dx.doi.org/10.1016/j.cie.2017.10.033 6. berdyugina, o.n., vlasov, a.i., & kuzmin, e.a.. (2017). investment capacity of the economy during the implementation of projects of public-private partnership, investment management and financial innovations, 14(3), 189-198. https://doi.org/10.21511/imfi.14(3-1).2017.03 7. chen, m.-t., & lin, c.-m. (2016). development of a smart home energy saving system combining multiple smart devices. 2016 ieee international conference on consumer electronics-taiwan, icce-tw 2016 3, iot in consumer electronic: from vision to reality, 7521072. https://doi.org/10.1109/icce-tw.2016.7521072 8. fund “center of strategical developments ‘north-west’”. (2012). “intelligent” environments, “intelligent” systems, “intelligent” productions: series of reports (green books) within the framework of the project “industrial and technological foresight of the russian federation” (issue 4). saint petersburg. 9. jain, v.k., & chapman, g.h. (2011). enhanced defect tolerance through matrixed deployment of intelligent sensors for the smart power grid. proceedings – 2011 ieee international symposium on defect and fault tolerance in vlsi and nanotechnology systems, dft 2011, 235-242. https://doi.org/10.1109/dft.2011.40 10. kolarzh, v.v. (2017). analysis of housing and utilities branch as a perspective field of entrepreneurship, scientific journal of sru itmo. series “economy and economical management”, 1, 10-14. 11. kumar, n., vasilakos, a.v., & rodrigues j.j.p.c. (2017). a multi-tenant cloud-based dc nano grid for self-sustained smart buildings in smart cities, ieee communications magazine, 55(3), 14-21. https://doi.org/10.1109/mcom.2017.1600228cm 12. limba, t., stankevičius, a., & andrulevičius, a. (2018). industry 4.0 and national security: the phenomenon of disruptive technology, entrepreneurship and sustainability issues, 6(3), 1328-1335. https://doi.org/10.9770/jssi.2019.6.3(33) 13. lund, h., østergaard, p.a., connolly, d., & mathiesen, b.v. (2017). smart energy and smart energy systems, energy, 137, 556-565. https://doi.org/10.1016/j.energy.2017.05.123 14. mobley, r.k. (2002). an introduction to predictive maintenance. elsevier science. https://doi.org/10.1016/b978-075067531-4/50006-3 concept of a universal smart circuit breaker with protection against spark discharge 11 copyright ©2020 assa adv. in systems science and appl. (2020) 15. monakov, v.k., kudryavtsev, d.yu., & peshkun, v.a. (2016). protection against arching fault. time-current characteristics of the device, energy security and energy saving, 6, 5-8. https://doi.org/10.18635/2071-2219-2016-6-5-8 16. monakov, v.k., kudryavtsev, d.yu., & smirnov, v.v. (2016). calculation of timecurrent specification of the protection device against arcing fault/ break on the basis of a serial arc circuit with a cable conductor opening, fire and explosion safety, 25(11), 45-50. https://doi.org/10.18322/pvb.2016.25.11.45-50 17. nagay, i., nagay, v., kireev, p., & sarry, s. (2017). relay protection designing as a solution to the problem of recognizing the regimes of the electrical grid. proceedings of the 9th international scientific symposium on electrical power engineering, elektroenergetika, 9, 409-413. 18. plan of measures (road map) to improve the legislation and eliminate the administrative barriers in order to ensure the implementation of the national technology initiative in the field of energynet. (2018, april 28). approved by the order of the government of the russian federation no. 830-r. 19. praca, d., & barral, c. (2001). from smart cards to smart objects: the road to new smart technologies, computer networks, 36(4), 381. https://doi.org/10.1016/s13891286(01)00161-x 20. prause, g., & atari, s. (2017). on sustainable production networks for industry 4.0, entrepreneurship and sustainability issues, 4(4), 421-431. https://doi.org/10.9770/jesi.2017.4.4(2) 21. ragulina, y.v., semenova, e.i., zueva, i.a., kletskova, e.v., & belkina, e.n. (2018). perspectives of solving the problems of regional development with the help of new internet technologies, entrepreneurship and sustainability issues, 5(4), 890-898. https://doi.org/10.9770/jesi.2018.5.4(13) 22. rogalev, a., komarov, i., kindra, v., & zlyvk, o. (2018). entrepreneurial assessment of sustainable development technologies for power energy sector, enterpreneurship and sustainability issues, 6(1), 429-445. http://doi.org/10.9770/jesi.2018.6.1(26) 23. rudenko, l.g. (2015). analysis of the state of housing and utilities infrastructure of russia under current conditions, bulletin of vitte moscow university. series 1: economy and management, 2, 67-78. 24. smart buildings pay. smart buildings sell. (1993). buildings, 87(10), 58. 25. tadviser. (2019, january). internet of things, iot, m2m global market. [online]. available www.tadviser.ru 26. vlasov, a.i., berdyugina, o.n., & krivoshein, a.i. (2018). technological platform for innovative social infrastructure development on basis of smart machines and principles of internet of things. 2018 global smart industry conference (glosic), 1-7. https://doi.org/10.1109/glosic.2018.8570062 27. vlasov, a.i., grigoriev, p.v., krivoshein, a.i, shakhnov, v.a., filin, s.s. et al. (2018). smart management of technologies: predictive maintenance of industrial equipment using wireless sensor networks, entrepreneurship and sustainability issues, 6(2), 489502. http://doi.org/10.9770/jesi.2018.6.2(2) 28. vlasov, a.i., yudin, a.v., salmina, m.a., shakhnov, v.a., & usov, k.a. (2017). design methods of teaching the development of internet of things components with considering predictive maintenance on the basis of mechatronic devices, international journal of applied engineering research, 12(20), 9390-9396. 29. whitaker, d.a., egan, d., o’brien, e., & kinnear, d. (2018). application of multivariate data analysis to machine power measurements as a means of tool life predictive maintenance for reducing product waste. [online]. available https://arxiv.org/abs/1802.08338 30. yip, s.-c., hew, w.-p., gan, m.-t., phan, r.c.-w., tan, s.-w. et al. (2017). detection of energy theft and defective smart meters in smart grids using linear regression, 12 a.i. vlasov, s.s. filin, a.i. krivoshein copyright ©2020 assa adv. in systems science and appl. (2020) international journal of electrical power & energy systems, 91, 230-240. https://doi.org/10.1016/j.ijepes.2017.04.005 31. vlasov, a.i., shakhnov, v.a., filin, s.s., & krivoshein, a.i. (2019). sustainable energy systems in the digital economy: concept of smart machines. entrepreneurship and sustainabilility issues, 6(4): 1975-1986. adv syst sci appl 2018; 2:11–25 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/604 surrogates for the matrix `0-quasinorm in sparse feedback design: numerical study of the efficiency alexey bykov1, pavel shcherbakov1,2∗ 1institute of control sciences, russian academy of sciences, moscow, russia 2institute for systems analysis, federal research center “computer science and control”, russian academy of sciences, moscow, russia abstract: some formulations of the optimal control problem require the resulting controller to be sparse; i.e., to contain zero elements in the gain matrix. on one hand, sparse feedback leads to the drop of performance as compared to the optimal control; on the other hand, it confers useful properties to the system. for instance, sparse controllers allow to design distributed systems with decentralized feedback. some sparse formulations require the gain matrix of the controller to have a special sparse structure which is characterized by the presence of zero rows in the matrix. in this paper, various approximations to the number of nonzero rows of a matrix are considered and applied to sparse feedback design in optimal control problems for linear systems. along with a popular approach based on using the matrix `1-norm, more complex nonconvex surrogates are proposed and discussed, those surrogates being minimized via special numerical procedures. the efficiency of the approximations is compared via numerical experiments. keywords: sparse control, `1-optimization, linear systems, optimal control, linear matrix inequalities 1. introduction optimal control problems in linear theory may have a requirement imposed on the structure of the controller along with the necessity of optimizing the objective function. one example of such requirement would be sparsity of the controller, which is understood as having many zero elements in its gain matrix; e.g., see [12–14,22]. elementwise and block sparsity allows to design distributed systems with decentralized control. this additional sparsity requirement naturally leads to a drop in performance, which is explained by the fact that the set of sparse controllers is contained in the feasible set for the corresponding classical formulation. nevertheless, in some formulations of control problems, a reasonable drop in performance is allowed for the sake of having a sparse control. for instance, the objective function in the standard lqr problem has the form of a quadratic functional, which is related to energy consumption in the system; e.g. fuel consumption of an aircraft. on one hand, a sparse control leads to worse (higher) values of the objective function, i.e., fuel consumption is increased; on the other hand, it augments the system with various useful properties. for example, decentralized control makes the system more reliable and fail-safe; decreasing the number of active actuators allows to slow down an equipment deterioration, thereby increasing the lifetime of the system. in [15], a special kind of sparsity was introduced, which is different from a typical elementwise sparsity. the authors considered a controller to be sparse if its gain matrix has ∗corresponding author: alexey.bykov.mipt@gmail.com, cavour118@mail.ru http://ijassa.ipu.ru/ 12 a. bykov, p. shcherbakov many zero rows or columns. such controllers facilitate ease of hardware implementation of the control systems, such as a reduction of the number of actuators or sensors and the amount of information transmitted via control channels. a direct minimization of the number of nonzero rows or columns of a matrix is known to be np-hard, since it involves combinatorial search. as the dimensions of the problem increase, the direct approach becomes inapplicable due to its exponential time complexity. in [15, 16], an approach was proposed which is based on solving convex surrogate problem instead of the original nonconvex one. this substitution is implemented using special matrix norms which are convex approximations (also called surrogates) to the number of nonzero rows/columns of a matrix (nonconvex matrix `0-quasinorms). this heuristic, however, does not guarantee an occurrence of zero rows/columns in the gain matrix, and by all appearances strict results cannot be obtained due to nonconvexity of the original problem. nevertheless, the exploited heuristic is shown to be efficient and the approach seems to work for numerous examples. such non-strict approaches require computer simulations for proving their efficiency, and in the current paper we put emphasis on the numerical study. apart from convex surrogates, some papers [2–4, 13] suggest using nonconvex approximations for more efficient detection of sparse solutions in case of vector and matrix variables. the main goal of this paper consists in comparing the efficiency of different approximations to the matrix `0-quasinorm, which can be used for sparse feedback design in optimal control problems. also, a general scheme for gaining trade-off between optimality and sparsity of the solution is proposed which allows to use various approximations in a similar way. the key aspects of this paper are the formulation of several surrogates for the matrix `0-quasinorm, which can be used for promoting matrix row-sparsity, designing the scheme of numerical study, and the analysis of the results of experiments. both models of simple systems and linearized models of real systems were chosen as test cases for conducting the numerical experiment. also, various software for numerical computing [5, 18, 19], and algorithms for nonconvex optimization [3, 21] were used. numerical study is the central part of this paper due to the following reason. although the approach based on convex surrogates heuristic is said to be efficient, one cannot be sure about getting zero rows in the gain matrix. thus, different surrogates demonstrate different efficiency for different problems; this paper aims to study the efficiency of various approximations. our paper essentially exploits linear matrix inequality (lmi, [1]) technique, which allows to formulate many classical problems from control theory in the form of semidefinite programs (sdp, [20]). 2. sparse feedback design in optimal control problems consider the following control system: ẋ = ax+bu, x(0) = x0, (2.1) where x(t) ∈ rn is the state vector, u(t) ∈ rp is the control input,a ∈ rn×n,b ∈ rn×p, and the pair (a,b) is assumed to be controllable. if system (2.1) is not stable, a stabilizing control is to be found, which also has to minimize an objective function along the system’s trajectories. various formulations differ in the form of the control input; in this paper we exploit a static state linear feedback form: u = kx, k ∈ rp×n. (2.2) from this point, we consider the lqr (linear control regulation) problem, the approach and the numerical experiment being applicable to other optimal control problems. in the lqr copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 13 problem, the objective function has the form of the following quadratic functional: j = ∞∫ 0 ( x>rx+ u>su ) dt, (2.3) where r ∈ rn×n and s ∈ rp×p are positive–definite matrices. it is worth emphasizing that minimizing functional (2.3) automatically leads to stabilizing system (2.1). indeed, if the system was unstable, integral (2.3) would be divergent; however, since the pair (a,b) is controllable, there exists a controller, which makes the functional j finitely valued. the standard way of solving the lqr problem consists in finding a solution to the associated algebraic riccati equation; however, in this paper we adhere to another approach (e.g., see [7]) based on the linear matrix inequality technique. in [7], the classical lqr problem is reduced to the following sdp problem: problem 1: let popt, yopt be the solution of the sdp trp −→ max s.t.( ap + pa> +by + y >b> p y > ∗ −r−1 0 ∗ ∗ −s−1 ) 4 0, (2.4) then the controller (2.2) with gain matrix kopt = yoptp −1 opt (2.5) stabilizes system (2.1), the objective functional (2.3) being minimized, and the optimal value of the functional is equal to jopt = x>0 p −1 optx0. (2.6) this approach involving lmis can be very useful, for instance, for robust formulations where the matrices of the system contain uncertainty. but what’s more important for the current paper, it is the lmi-based approach that allows to solve the problem in sparse formulations. as it was mentioned, the sparse formulation that we consider, imposes an additional requirement on the structure of the gain matrix; namely the intention is to have as many zero rows in matrix k as possible. clearly, due to the structural constraint on the matrix k, the feasible set of controllers becomes smaller, so that a drop in performance of the sparse controller is unavoidable, and we do not want it to be significant. the problem of comparison of the two controllers with respect to a quadratic performance index is not as straightforward as it may seem; e.g., see [8]. if one compares the values of the functional j(p ) = x>p−1x corresponding to the two different stabilizing controllers k1 and k2 with the corresponding matrices p1 and p2, the sign of the inequality j(p1) ≶ j(p2) may change depending on the initial condition x0. in this paper we compare sparse controllers with the optimal ones, which, by construction, yield the optimal value of the functional j for any initial condition x0. in other words, the sparse controller with matrix psp will always lose the optimal controller with matrix popt in terms of the value of the quadratic functional j . the magnitude of the loss (performance drop) for a given initial condition x0 is jsp jopt = x>0 p −1 sp x0 x>0 p −1 optx0 ; copyright c© 2018 assa. adv syst sci appl (2018) 14 a. bykov, p. shcherbakov this value, obviously, can change significantly for various x0. there is a well-known approach [11] based on averaging the values of j over initial conditions uniformly distributed on the unit n-dimensional sphere. it is not hard to show that the mathematical expectation of j is equal to e(x>0 p −1x0) = 1 n trp−1; thus the magnitude of the sparse-to-optimal control loss “on the average” equals jsp jopt = trp−1sp trp−1opt . (2.7) in this paper we adhere to the described approach based on “averaging”; i.e., we use relation (2.7). in [16], a procedure is proposed to design sparse controllers in optimal control problems, in particular, in the lqr problem. the key stage of the procedure, which is responsible for the detection of the sparse structure of the controller, consists in minimizing the 1∞-norm of the gain matrix. this norm plays the role of a convex surrogate for the number of nonzero rows of a matrix and is defined as follows: ‖x‖1∞ = n∑ i=1 max 16j6p |xij|, x ∈ rn×p. (2.8) the three-step procedure from [16], generalized to the case of arbitrary approximation to the matrix `0-quasinorm, was used as a basis for the numerical study in this paper. the key aspects of the procedure are given below. step 1 (optimal control design). the classical lqr problem is solved (problem 1) and the value trp−1opt is fixed. this value corresponds to the optimal value jopt of the quadratic functional j . step 2 (detection of sparse structure). let α > 1 denote the maximum admissible loss in the quadratic performance index that appears due to sparsity; i.e., the following condition is imposed: trp−1sp trp−1opt 6 α, α > 1. so, α is the coefficient of the allowed performance drop “on average” and it acts as the parameter of the following sdp problem: problem 2: ‖y ‖(sp) −→ min s.t. (2.4),( x i i p ) < 0, and trx 6 α trp−1opt . (2.9) note that, similarly to problem 1, the gain matrix k does not explicitly appear in the constraints or in the objective function of the sdp. thus, the variables are positive–definite matrices p and x and an auxiliary matrix variable y = kp ; the latter was introduced to transform the constraints to the form of linear matrix inequalities. the minimized function in problem 2 is denoted by ‖·‖(sp) and it plays the role of an approximation to the matrix `0-quasinorm. the 1∞-norm is one of such approximations and copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 15 it allows to solve convex problem instead of the nonconvex one. there are more complex and efficient surrogates, some of which are described in the next section. let the matrix y0 be the solution of problem 2 and assume it contains rows with all elements being equal or close to zero. these rows are memorized and kept as strict zeros, which automatically means that the gain matrix k0 = y0p −1 will also contain the same zero rows. the latter fact follows from the rules of the matrix multiplication. note that the value of the objective function does not exceed αjopt “on average”. step 3 (optimization over sparse-structured controllers). the lqr problem is again solved but with sparse structure of the controller being fixed as found at step 2. in other words, problem 1 is solved with additional constraints on the matrix variable y . to sum up, the procedure described above gives a way to design a controller which may have a sparse structure at the expense of a reasonable drop in performance; thus, a trade-off between optimality and sparsity of the control can be achieved. 3. approximations to the matrix `0-quasinorm so, the key step in the sparse feedback design procedure is the sparsity detection step. the detection is performed via minimizing a surrogate for the number of nonzero rows of a matrix. below we consider four various approximations to the matrix `0-quasinorm, efficiency of which is compared via numerical experiments. all approximations are described as functions of ri, where ri qualifies the magnitude of the values of the ith row of the matrix y . if a row of the matrix is treated as a vector, then ri can be, e.g., one of the standard vector norms: l1-, l2-, or l∞-norm. a matrix row is called a zero row if and only if all of its elements are zeros. also, if any of the mentioned norm approaches zero, then all of the row elements approach zero. numerical experiments show that the choice of the particular norm is not essential, so, to be specific, we chose l∞-norm, the maximum of the absolute values of the components. thus, we define ri as follows: ri = max 16j6n |yij|, y = ‖yij‖ ∈ rp×n. (3.10) 3.1. 1∞-norm the 1∞-norm (2.8) was introduced in [16]; in terms of ri (3.10) it can be written as: ‖y ‖1∞ = p∑ i=1 ri. (3.11) this function is convex, therefore its minimization subject to constraints (2.4), (2.9) is an sdp problem, which can be efficiently solved by means of convex optimization tools [20]. 3.2. weighted 1∞-norm in [3], the authors developed the ideas of `1-optimization for promoting sparsity in case of vector variables. they introduced a new iterative procedure which consists in solving a sequence of weighted `1-optimization problems. at each iteration, the weight coefficients are updated due to the solution received at the previous step. the weighted 1∞-norm is defined as follows: ‖y ‖w1∞ = p∑ i=1 wiri. (3.12) the iterative procedure from [3] can be schematically written down as follows: algorithm 1: copyright c© 2018 assa. adv syst sci appl (2018) 16 a. bykov, p. shcherbakov 1. initial estimate: l = 0, w (0) i = 1, i = 1, . . . , p. 2. the convex problem with currently accepted weights is solved: r(l) = argmin r [ p∑ i=1 w (l) i ri ] s.t. (2.4), (2.9). 3. the weights are updated: w (l+1) i = 1 |r(l)i |+ ε , ε > 0. 4. quit the loop if the procedure converged or the maximum number of iterations is exceeded. if both conditions are false, go back to step 2. the authors of the original paper [3] confirmed via numerous examples that the new approach involving the minimization of the weighted `1-norm performs much better than the “standard” non-weighted approach. this can be explained by the fact that the sequence of convex problems can better approximate the original nonconvex problem of minimizing the matrix `0-quasinorm. 3.3. nonconvex sparsity detector, nsd in [2], a nonconvex sparsity-promoting surrogate was introduced: ‖y ‖nsd = p∑ i=1 ri ∏ j 6=i rj rj + 1 . (3.13) the authors provide some motivation for this particular approximation and demonstrate its efficiency as compared to the 1∞-norm. the function (3.13) is shown to be a dc-function (difference of convex, see [6]) and for its minimization, the concave-convex procedure (abbr. cccp, e.g., see [21]) is proposed. let the function (3.13) be represented as a dc-function: nsd(r) = u(r)− v (r). then the procedure for its minimization is given below: algorithm 2: 1. initial estimate: l = 0, r (0) i = max 16j6n |yij|, i = 1, . . . , p, where yopt = ‖yij‖ is the solution of problem 1. 2. the following convex problem is solved: r(l+1) = argmin r [ u(r)− r>∇v ( r(l) )] s.t. (2.4), (2.9). 3. quit the loop if the procedure converged or the maximum number of iterations is exceeded. if both conditions are false, go back to step 2. in [9], the results on the convergence of the cccp procedure are presented; in the general case, global convergence is not guaranteed, since the problem is not convex. however, for many applications this procedure yields an appropriate solution. copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 17 3.4. nonconvex approximation log-sum in [3, 4, 13], the function log-sum (see definition below) is shown to present a more precise approximation to the `0-norm of a vector than convex approximations. indeed, let us consider the following scalar functions (see fig. 3.1) which can be used as penalty functions for nonzero values: f0(x) = 1[x 6=0], f1(x) = |x|, flog,ε(x) = cε log ( 1 + |x| ε ) ; the factor cε is chosen in a such way that the equality flog,ε(1) = 1 = f0(1) = f1(1) holds. -1 0 1 x 0 1 f 0 (x) f 1 (x) f log, ǫ (x) fig. 3.1. approximations to the `0-quasinorm the derivative of the function flog,ε(x) in the positive vicinity of the origin grows approximately like 1 ε as ε −→ 0. this leads to a relatively big amount of penalty for small but not exactly zero values. moreover, note that flog,ε(x) −→ f0(x) as ε −→ 0. unfortunately, as ε tends to zero, the minimization of the function log-sum becomes difficult; for heuristic of choosing ε see [3]. the function log-sum adapted for promoting matrix row-sparsity can be written as ‖y ‖log-sum = p∑ i=1 log (1 + ri) . (3.14) since the function (3.14) is concave, it cannot be minimized via convex optimization framework. however, it can be treated as a special case of dc-function, so, it can be minimized via concave-convex procedure (cccp) just like the nsd function. 4. results of numerical experiments in the previous section, several well-known and relatively new approximations to the number of nonzero rows of a matrix were considered. among these surrogates are the 1∞-norm copyright c© 2018 assa. adv syst sci appl (2018) 18 a. bykov, p. shcherbakov exploited in [16], and the nsd function introduced in [2]. both approximations, however, have not yet received enough approval from numerical experiments; therefore, they were chosen as contestants for the experiment in this paper. also, two other approximations (the weighted 1∞-norm and the function log-sum), which were earlier applied in the case of vector variables, were adapted for promoting matrix row-sparsity. all these surrogates were used at step 2 of the three-step procedure to design sparse controllers in the numerical experiment. all the examples described below were solved with matlab using the cvx framework [5] for solving convex optimization problems and the sdpt3 solver [19]. it was checked that using other solvers such as sedumi [18] did not give any specific benefits or drawbacks. the main goal of the experiment was to compare the efficiency of approximations (3.11), (3.12), (3.13), (3.14) used for getting zero rows in the matrix y0 at step 2 of the threestep procedure, which allows to design sparse controllers in the lqr problem. during the experiments, both models of simple systems and linearized models of real systems were considered. the latter were taken from the compleib, a popular collection of test problems from control system design and related fields. the general scheme of the experiment is as follows: algorithm 3 (scheme of experiment): 1. the classical lqr problem is solved (problem 1) and the value trp−1opt is fixed. this value corresponds to the optimal value jopt of the quadratic functional j . 2. the initial value of the maximum admissible loss in the quadratic performance index is chosen: αmin = 1 + ε, ε > 0. 3. the magnitude of α is increased inside the loop until the most sparse control can be acquired. for every value of α the following steps are executed: (a) the detection of zero rows is performed (problem 2) for a given α̂. the zero structure obtained from the solution is fixed; the number of zero rows denoted by nnz is memorized. (b) the lqr problem (problem 1) is solved with the sparse structure of the controller being fixed as found at the previous step. the performance drop is determined by the relation αsp = trp−1 sp trp−1 opt . hence, every iteration of the loop gives us two points at the plane with coordinates the number of nonzero rows and the performance drop. these points are (nnz, α̂) and (nnz, αsp), where α̂ is the value of the maximum admissible loss which leads to the detection of zero rows, and αsp is the actual loss in performance, which corresponds to the detected sparse structure of the controller. it is worth noting that αsp might be drastically less than α̂. 4. the brute-force enumeration of all possible combinations of zero rows in the gain matrix is performed and for every combination and the corresponding sparse structure the lqr problem (problem 1) is solved. note that such a brute-force search is performed exceptionally within the experimental setup for illustration purposes only; it is not needed in the respective theorems. for the combination with index k ∈ [1; 2p − 2] the performance drop is determined by the relation αbf,k = trp−1 bf,k trp−1 opt . note that we consider 2p − 2 combinations of zero rows instead of 2p, since two cases are excluded. the first one corresponds to the absence of zero rows, i.e., to the classical formulation without the sparsity requirement. the second obvious case corresponds to the situation where all rows of the gain matrix are zero rows. the points (nnz, αsp) obtained as an output of the algorithm 3 are related to the pareto optimality, a well-known concept from the theory of multiobjective optimization. namely, a copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 19 point is said to be pareto-optimal if none of the criteria can be improved without worsening the others. also, note that the values of αsp might not be optimal for a given number of nonzero rows nnz, since we solve the surrogate for the original nonconvex problem and the global convergence is not guaranteed. nevertheless, varying the importance of one or another criterion we can gain a necessary trade-off between optimality and sparsity. 4.1. examples from compleib the experiment was conducted in accordance with the described scheme for several problems from the compleib collection: “ac1”, “ac9”, “ac12”, “he3”, “he4” (linearized models of aircrafts and helicopters). these examples are of interest in terms of the dimension of the control input. the behaviour of various approximations to the matrix l0-quasinorm for all five models is quite similar, so it is reasonable to present the results for just one example – model “he4”. this eight-order model represents a twin-engine, multi-purpose military helicopter. let us recall that the behaviour of the system is described by equations (2.1). the values of the entries of the matrices a and b are not presented for space considerations; they can be found in the compleib documentation. the weight matrices r and s in the quadratic functional (2.3) were set to identity. the results of the experiment conducted in accordance with algorithm 3 are presented in figs. 4.2, 4.3, and 4.4. 2 3 4 the number of nonzero rows 1 1.002 1.004 1.006 1.008 1.01 1.012 1.014 1.016 1.018 r e la x a ti o n c o e ff ic ie n t α results for problem he4 1∞-norm w1∞-norm nsd sum-log fig. 4.2. “he4”. the detection of zero rows for various α̂ copyright c© 2018 assa. adv syst sci appl (2018) 20 a. bykov, p. shcherbakov figure 4.2 shows the points (nnz, α̂) obtained for the functions (3.11), (3.12), (3.13), (3.14) in accordance with algorithm 3. this plot demonstrates how big one should set the admissible loss in performance in order to get zero rows in the gain matrix of the controller. one can see that using approximations (3.12), (3.13), (3.14) leads to the detection of sparse structure for smaller values of α̂. it turned out that for any surrogate except for the 1∞-norm, in order to detect two zero rows out of the total four, we needed to set α̂ slightly less than 1.005, which corresponds to the performance drop of just 0.5%. the transition from one zero row to two zero rows happens at close values of α̂, therefore using an insufficiently dense grid of α̂ values can mislead into believing that both zero rows appear simultaneously. such simultaneous appearances of two zero rows were not noticed during the experiment, and they obviously should not happen due to the following reason. every next zero row of the gain matrix makes the feasible set of controllers smaller, so the optimal value of the quadratic functional for the controller with n+ 1 zero rows is a priori worse than the corresponding value for the controller with n zero rows. 2 3 4 the number of nonzero rows 1 1.002 1.004 1.006 1.008 1.01 1.012 1.014 1.016 1.018 a c tu a l lo s s results for problem he4 1∞-norm w1∞-norm nsd sum-log fig. 4.3. “he4”. the actual loss in performance index figure 4.3 shows the points (nnz, αsp) obtained from the algorithm 3, which correspond to the actual sparse-to-optimal loss in performance. for the 1∞-norm we had α̂� αsp, while using the other surrogates yields a sparse structure for α̂ close to αsp, which is surely an advantage of these approximations. indeed, if one had to set α̂� αsp in order to get a sparse control, the interpretation of α̂ as a maximum admissible performance loss would be incorrect. for instance, using the 1∞-norm in example “he4”, in order to detect sparse control which yields 0.5% of the actual performance drop, we had to allow for a maximum admissible performance loss of 15− 20% during the detection step. moreover, using the copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 21 1∞-norm in some examples can lead to a large difference between α̂ and αsp. using other approximations usually allows to detect a sparse structure for the values α̂ which are close to the corresponding values of αsp; i.e., α̂ is a more accurate estimate of αsp. 2 3 4 the number of nonzero rows 10 0 10 1 10 2 a c tu a l lo s s results for problem he4 1∞-norm w1∞-norm nsd sum-log fig. 4.4. “he4”. the actual loss in performance index. all possible combinations of zero rows figure 4.3 depicts the points marked as black crosses which were obtained at step 4 of the algorithm 3 during the brute-force enumeration. it turned out that for our example the threestep procedure yielded the optimal sparse structures for every number of zero rows. figure 4.3 is just a part of the whole picture presented in fig. 4.4, where all possible combinations of zero rows of the gain matrix can be seen. as we can see, a poor choice of the sparse structure can lead to a dramatic loss in the performance index. it is important to understand that the occurrence of large values of performance drop in fig. 4.4 is explained by the specific properties of the system itself rather than by sparsity detection methods; some systems just do not tolerate zeroing out specific control inputs. this fact also leads to an interesting “sideeffect” which consists in the ability to filter out sparse controllers with “poor” structure a priori by choosing appropriate values for α̂. similar results were obtained for other examples from compleib such as “ac1”, “ac9”, “ac12”, “he3”. therefore, all the observations and conclusions made for model “he4” are also valid for all listed examples. 4.2. mass spring system consider the system which consists of n masses mis connected via springs with stiffness coefficients kis; the masses can slide frictionless along the line (see fig. 4.5). let us denote copyright c© 2018 assa. adv syst sci appl (2018) 22 a. bykov, p. shcherbakov the displacement of the ith mass from its reference position as pi and let the state variables be x1 = [p1 . . . pn ] > and x2 = ẋ1. fig. 4.5. model “mass spring system” similarly to the previous section, the behaviour of the system is described by the equation (2.1). in this example we used identity weight matrices r and s in the quadratic functional (2.3). 2 3 4 5 6 7 8 9 10 the number of nonzero rows 1 1.5 2 2.5 3 3.5 4 r e la x a ti o n c o e ff ic ie n t α results for problem ms1 10 1∞-norm w1∞-norm nsd sum-log fig. 4.6. “ms”. the detection of zero rows for various α̂ for the sake of simplicity we considered the system with the following parameters: m1 = · · · = mn = 1, k1 = · · · = kn = 1, copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 23 so that the matrices of system (2.1) has the form a = [ o i t o ] , b = [ o i ] , where t ∈ rn×n is a tridiagonal toeplitz matrix with −2 on its main diagonal and 1 on its first superand sub-diagonal, i ∈ rn×n is the identity matrix, and o ∈ rn×n is the zero matrix. this model is useful, since it allows to choose arbitrary values for the number of masses n , i.e., to vary the dimension of the problem. the experiment was conducted for several values of n , the observed results being quite similar. we present the results for n = 10, since even for n = 10 the number of all possible combinations of zero rows is equal to 210 − 2 = 1022, thus, the brute-force approach is still possible yet challenging. 2 3 4 5 6 7 8 9 10 the number of nonzero rows 1 1.5 2 2.5 3 3.5 4 a c tu a l lo s s results for problem ms1 10 1∞-norm w1∞-norm nsd sum-log fig. 4.7. “ms”. the actual loss in performance index the results of the experiment for this example are presented in figs. 4.6 and 4.7. the figures show that use of the 1∞-norm in this problem did not lead to the successful detection of a sparse structure even for large values of α̂. this reminds us of the example “he4” from the previous section, where we noticed that using the 1∞-norm makes α̂ a poor estimate of αsp. this observation is valid for all values of n tested during the experiment. it is hard to explain such an inefficiency of the 1∞-norm. in this particular example one of the possible reasons might be a homogeneous internal structure of the system, which probably makes the choice of the control inputs to be zeroed out a challenging task. copyright c© 2018 assa. adv syst sci appl (2018) 24 a. bykov, p. shcherbakov although the weighted 1∞-norm succeeded in yielding sparse controllers, it is less efficient than the nonconvex approximations (3.13) and (3.14). nevertheless, algorithm 1 for minimizing the weighted 1∞-norm is quite straightforward in implementation and allows to achieve better results than the standard minimization of the 1∞-norm. there is an interesting detail about the process of the detection of zero rows. this process is sequential, i.e., the number of zero rows grows step-wise or one by one. if the values in the row approach zero, say become less than 10−10, then, according to the algorithm, the weight for this row becomes very large, since it is inversely proportional to the maximum among the absolute values of the elements of this row. for a row with the weight 1010 it is hardly possible to gain nonzero values. thus, if at any iteration some row becomes zero, it will stay zero till the end of the algorithm. implementation of the algorithm to minimize approximations (3.13), (3.14) (cccp) requires calculating a gradient of the corresponding functions, its iterative nature being similar to algorithm 1. in most cases, the nonconvex surrogate nsd performs as good as the log-sum approximations, and sometimes nsd happens to be better. figure 4.7 shows that using functions (3.13), (3.14) yielded sparse controllers which are close or equal to optimal sparse controllers. in terms of optimality, among sparse controllers with the fixed number of zero rows the weighted 1∞-norm (3.12) is less efficient than nonconvex surrogates. this detail is important due to the fact that global convergence is not guaranteed. on the ground of the results of the numerical experiments described above we arrived at the following conclusions: • using the 1∞-norm can lead to situations where the actual sparse-to-optimal performance drop αsp is significantly smaller than the maximum admissible loss in performance α̂. other surrogates usually do not suffer from this drawback. • using approximations to the matrix `0-quasinorm usually yields a sparse controller which is optimal among all controllers with the same number of zero rows. • in some problems (e.g., see the example from section 4.2) the 1∞-norm does not succeed in the detection of sparse structure even for large values of α̂. • the nsd function appears to be effective similarly to the nonconvex surrogate log-sum, and in some cases the nsd happens to perform better. 5. conclusion in this paper we considered various approximations to the matrix `0-quasinorm, which can be applied to sparse feedback design in optimal control problems. the results of the numerical experiments show that to some extent all the surrogates analyzed can be applied to the detection of zero rows of the gain matrix, but nonconvex surrogates perform better. future efforts to be made in this area include the analysis of examples where one or another approximation to the matrix `0-quasinorm happens to be inefficient, and studying the alternative numerical procedures which can be applied to sparse feedback design, e.g., admm, alternating direction method of multipliers. acknowledgements this work is supported by the russian science foundation through grant no. 16-1110015. references copyright c© 2018 assa. adv syst sci appl (2018) surrogates for the matrix `0-quasinorm 25 1. boyd, s., ghaoui, l.e., feron, e. & balakrishnan, v. (1994) linear matrix inequalities in system and control theory. philadelphia, pa: siam studies in applied mathematics. 2. bykov, a., shcherbakov, p. & ding, m. (2016) a tractable nonconvex surrogate for the matrix `0-quasinorm: applications to sparse feedback design, ifac-papersonline, 49(13), 53–58. 3. candès, e.j., wakin, m.b. & boyd, s.p. (2008) enhancing sparsity by reweighted `1 minimization, journal of fourier analysis and applications, 14(5), 877–905. 4. fazel, m., hindi, h. & boyd, s. (2003) log-det heuristic for matrix rank minimization with applications to hankel and euclidean distance matrices, proc. of american control conference, 2156–2162. 5. grant, m. & boyd, s. (2014) cvx: matlab software for disciplined convex programming, version 2.1, url: http://cvxr.com/cvx. 6. hartman, p. (1959) on functions representable as a difference of convex functions, pacific journal of mathematics, 9(3), 707–713. 7. khlebnikov, m.v., shcherbakov, p.s. & chestnov v.n. (2015) linear-quadratic regulator. i. a new solution, autom. remote control, 76(12), 2143–2155. 8. khlebnikov, m.v. (2014) an ellipsoid approach to comparison of quadratic performance criteria, stochastic optimization in computer science, 10(1), 145–156. in russian. 9. lanckriet, g.r. & rangarajan, b.k. (2009) on the convergence of the concave-convex procedure, advances in neural information processing systems, 22, 1759–1767. 10. leibfritz, f. & lipinski, w. (2003) description of the benchmark examples in compleib 1.0, url: http://www.complib.de. 11. levine, w.s. & athans, m. (1970) on the determination of the optimal constant output feedback gains for linear multivariable systems, ieee trans. on automatic control, 15(1), 44–48. 12. lin, f., fardad, m. & jovanović, m.r. (2011) augmented lagrangian approach to design of structured optimal state feedback gains, ieee trans. on automatic control, 56(12), 2923–2929. 13. lin, f., fardad, m. & jovanović, m.r. (2013) design of optimal sparse feedback gains via the alternating direction method of multipliers, ieee trans. on automatic control, 58(9), 2426–2431. 14. münz, u., pfister, m. & wolfrum p. (2014) sensor and actuator placement for linear systems based onh2 andh∞ optimization, ieee trans. on automatic control, 59(11), 2984–2989. 15. polyak, b., khlebnikov, m. & shcherbakov, p. (2013) an lmi approach to structured sparse feedback design in linear control systems, proc. of european control conference, 833–838. 16. polyak, b.t., khlebnikov, m.v. & shcherbakov, p.s. (2014) sparse feedback in linear control systems, autom. remote control, 75(12), 2099–2111. 17. polyak, b.t., khlebnikov, m.v. & shcherbakov, p.s. (2014) control of linear systems subjected to exogenous disturbances: the linear matrix inequality technique. moscow, russia: urss. in russian. 18. sturm, j.f. (1999) using sedumi 1.02, a matlab toolbox for optimization over symmetric cones, optimization methods and software, 11–12, 625–653. 19. toh, k.c., todd, m.j. & tutuncu, r.h. (1999) sdpt3 – a matlab software package for semidefinite programming, optimization methods and software, 11, 545–581. 20. vandenberghe, l. & boyd, s. (1996) semidefinite programming, siam review, 38(1), 49–95. 21. yuille, a.l. & rangarajan, a. (2003) the concave-convex procedure, neural computation, 15, 915–936. 22. zelazo, d., schuler, s. & allgöwer f. (2013) performance and design of cycles in consensus networks, systems and control letters, 62(1), 85–96. copyright c© 2018 assa. adv syst sci appl (2018) introduction sparse feedback design in optimal control problems approximations to the matrix 0-quasinorm 1-norm weighted 1-norm nonconvex sparsity detector, nsd nonconvex approximation log-sum results of numerical experiments examples from compleib mass spring system conclusion adv syst sci appl 2019; 02; 8-22 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/727 models of uncertain-random programming georgy veresnikov1*, ludmila pankova1, valeriya pronina1 1) v.a. trapeznikov institute of control sciences of russian academy of scienses moscow, russian e-mail: veresnikov@mail.ru received february 12, 2018; revised december 9, 2018; published december 31, 2018 abstract: the article proposes the models of optimization with constraints under conditions of parametric mixed uncertainty ‒ aleatory and epistemic. we model parameters with aleatory uncertainty by random values with probability distribution functions obtained from statistical data. we model parameters with epistemic uncertainty by uncertain values introduced in the uncertainty theory of liu b. experts define the uncertainty distribution functions. we model a function of random and uncertain parameters by uncertain-random value, interpreted as epistemic value parameterized by random values. optimization criteria (deterministic duplicates of objective functions) are combination of different characteristics of random and uncertain values, which allows both to average objective functions and to take into account risks or reliability arising from the variability of random and uncertain values. using the proposed models of uncertain-random programming, we formalized as a two-criterion optimization problem with constraints and solved the task of preliminary aerodynamic design in the conditions of parametric mixed uncertainty ‒ calculation of aircraft weight parameters. the uncertainty theory makes possible under certain conditions (for sufficiently wide class of functions) to obtain analytical expressions for characteristics of uncertain functions, that significantly reduces computational costs. to calculate weight parameters of aircraft, we use multicriteria genetic algorithm and statistical modeling. we investigate the dependence of the optimization result on the given probability levels for random values and the expert belief degree for epistemic values reflecting the reliability of the obtained solution. keywords: aleatory uncertainty, epistemic uncertainty, uncertain-random quantity, deterministic duplicate, uncertain-random programming, mixed uncertainty. 1. introduction aleatory (objective) and epistemic (subjective) uncertainties are two types of uncertainty, reflecting nondeterminism. in the context of modeling technical objects and decision making aleatory uncertainty occurs when information about a stochastic parameter is accumulated in statistical data and parameters are modeled by random variables with certain distributions. epistemic uncertainty arises when information about a parameter is obtained from experts, while the parameter may be either stochastic, but there are no or insufficient statistical data, or deterministic, but its value is unknown to date. the parameters with epistemic uncertainty are modeled by fuzzy, possibilistic, uncertain [1], and others values. decision making in the design of technical objects, as a rule, occurs under conditions of mixed uncertainty, when there are parameters both aleatory and epistemic. the existing methods and design tools do not take into account the presence of parameters with epistemic uncertainty. however if we consider the epistemic parameters as random variables, it can lead to errors, which is associated with the nonadditivity of expert judgments (the measure of uncertainty is nonadditive in contrast to the probabilistic measure). * corresponding author: veresnikov@mail.ru mailto:veresnikov@mail.ru models of uncertain-random programming 9 copyright ©2019 assa. adv. in systems science and appl. (2019) it is important to note that in construction of optimization models in conditions of uncertainty, transition from nondeterministic objective functions and restrictions to their deterministic duplicates is necessary. deterministic duplicates of objective functions are their numerical characteristics (mean, variance, quantile, etc.). let ),(f  be a function of epistemic values and aleatory values , i.e. a value with mixed uncertainty. there are two known approaches to modeling a value with mixed uncertainty and its characteristics (deterministic duplicates) based on different interpretations of a value with mixed uncertainty [1-4]. the first approach considers ),(f  as random value parameterized by epistemic values. we determine the characteristic of function ),(f  in two stages. first, we calculate the numerical characteristic of the function as a random quantity for each implementation of epistemic quantities. this characteristic does not depend on random quantities and, as a function of epistemic quantities, is an epistemic quantity. then we calculate the characteristic of this epistemic quantity, which is the characteristic of the function with mixed uncertainty. the second approach considers ),(f  as epistemic value parameterized by random values. we determine the characteristic of function ),(f  in two stages. first, we calculate the characteristic of the function as an epistemic quantity for each implementation of random quantities. this characteristic does not depend on epistemic quantities and, as a function of random quantities, is a random quantity. then we calculate the numerical characteristic of this random quantity, which is the characteristic of the function with mixed uncertainty. calculating characteristics of mixed uncertainty values is very expensive, especially in absence of explicit formulas for epistemic and/or random characteristics [2-4]. in this regard, it is relevant to highlight conditions under which there are analytical expressions of characteristics. thus, in [4] in framework of the first approach, the authors investigate values with mixed uncertainty, where possibilistic values model epistemic values, while possibilistic-random values have a shift-scale representation: f (ξ, ω) =α(ω)+σ(ω)∙ β(ξ), where α(ω) and σ(ω) are random values, β(ξ) is a possibilistic value. the authors obtain formulas for calculating the characteristics of the weighted sum of the possibilistic-random values of the shift-scale form with a certain type distribution for the possibilistic and random values. uncertainty theory makes possible for wider class of functions to obtain analytical expressions for characteristics of uncertain functions. that significantly reduces computational costs. in this paper, we model epistemic values by uncertain values introduced in theory of uncertainty [1]. further, we will call uncertain only such values. chance theory [1] introduces uncertain-random value, where uncertain values model epistemic values. the chance is a measure of mixed uncertainty. uncertain-random value is a real function on uncertain-random space with chance measure and have distribution function of chance measure (definitions 10-12, appendix). characteristics of uncertain-random values, defined as the characteristics of the chance measure distribution function (for example, definition 13, appendix), will always be averaged epistemic characteristics, since chance measure is the mathematical expectation of uncertain measure (definition 1, appendix). characteristics are actually determined as in the second approach, where at the first stage the epistemic characteristic is calculated, at the second stage the mathematical expectation is always calculated [1]. thus, use of chance theory in solving optimization problems leads to models with averaging criteria, i.e. the solution will be effective only “on average”, while risk of unwanted solutions is not considered. when modeling constraints using a chance measure, fulfillment of constraints is required only “on average” and risk of not fulfilling the constraints is not considered [1]. the paper proposes the models of uncertain-random programming, that is, the optimization models with constraints under conditions of parametric mixed uncertainty within the 10 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) framework of the second approach, which makes it possible to use analytical expressions for sufficiently wide class of functions of uncertain values. the choice of duplicates of objective functions and constraints allows to take into account risk and reliability requirements. section 2 describes the models of uncertain-random programming. in section 3 the task of calculating the weight parameters of an aircraft under conditions of parametric mixed uncertainty is formalized as a multicriteria optimization problem with constraints. we use two proposed models: with averages and with quantiles. the method for solving the task of calculating weight parameters and the results of calculations are given. the appendix contains definitions and statements from the theory of uncertainty and chance theory. 2. models of uncertain-random programming let ),(  ,xf be the objective function, x is solution vector, ξi, i = 1, …, n, are uncertain values, ωj, j = 1, ..., m, are random values that is, ),(  ,xf is uncertain-random value for each x . we interpret ),x(f  , as the uncertain value parameterized by the random values in the framework of second approach. we determine the characteristic of the function ),x(f  , in two stages. first, we calculate the characteristic of uncertain function ),x(f  , , where random values are parameters. then, we calculate the characteristic of this random value, which is characteristic of the function ),x(f  , with mixed uncertainty. mathematical expectation/expected value, quantile (definition 9, appendix), variance, probability/ belief degree of non-exceedance of a given value by the function can be chosen as characteristics of random and uncertain values. using combination of characteristics, you can build different optimization models. the choice of the optimization criterion depends on the specific task, reliability requirements and is the prerogative of the decision maker. the optimization model with mathematical expectation/expected value criteria, averaging the objective function, gives an effective solution “on average”, while risk or reliability is not taken into account. in robust optimization, when it is necessary to ensure the least variability of the objective function, the variance is used. in tasks of optimal control and design of aircraft under conditions of aleatory uncertainty, quantile and the probability of not exceeding (while minimizing) the objective function of a given threshold value have become widespread, since they are aimed at making optimal decisions based on risk or reliability requirements. we consider some models of uncertain-random programming and reduce them to mathematical or stochastic programming models. let )x(f  ,, be the objective function, x be the solution vector, ξ1, ξ2, …, ξn be independent uncertain values with uncertainty distribution functions φ1, φ2, …, φn , having inverse distribution functions, ω1, …, ωm be independent random values with probability distribution functions ψ1, ψ2, …, ψm. let )x(f  ,, be a continuous strictly increasing function with respect to ξ1, …, ξk and a strictly decreasing function with respect to ξk+1, …, ξn. for the model 1 (ee), we will select expected value em as characteristic at first stage, and mathematical expectation as characteristic at second stage. then the model 1 is: ))).(((min  ,,xfee mp x expected value of function )x(f  ,, where random values  are parameters is (theorem 2, appendix):  dxf,,xfe nkk m )),1(),...,1(),(),...,(,())(( 1 0 11 1 11 1     . (1) models of uncertain-random programming 11 copyright ©2019 assa. adv. in systems science and appl. (2019) mathematical expectation of random value ))x(f(e m  ,, is: ....)),1(),...,1(),(),...,(,( )))((( 1 1 0 11 1 11 1 m r nkk mp dddxf ,,xfee m          the model 1 has the form: ....)),1(),...,1(),(),...,(,(min 1 1 0 11 1 11 1 m r nkk x dddxf m       thus, we have reduced the model 1 to a deterministic model of mathematical programming. for the model 2 (qe), we use expected value as characteristic at first stage, and quantile of the random variable as characteristic at second stage. then the model 2 has the form: ,10.5 ,))),,((( where ,min   rxfep r m x or (according to the formula (1)): .10,)))),1(),...,1(),(),...,(,((( where ,min 1 0 11 1 1 k 1 1      rdxfp r nk x thus, we reduce the model 2 to stochastic model with quantile criterion. for the model 3 (eq), we use quantile as characteristic at first stage, and mathematical expectation as characteristic at second stage. then the model 3 has the form: ))),(((infmin  ,xfe p x , where in accordance with theorem 2 (d) of the appendix: ).),1(),...,1(),(),...,(,()),(( 11 1 11 1      nkk xf,xfinf (2) thus, we have reduced the model 3 to a deterministic model of mathematical programming. for the model 4 (qq), we use quantile of uncertain value as characteristic at first stage, and quantile of random value as characteristic at second stage. then the model 4 has the form: ,min r x where p (infα( ),,( xf ) < r} > β, 1 0 . thus, we reduce the model 4 to stochastic model with quantile criterion (formula (2)). consider representation of duplicate for constraint 0,, )x(g  , where g is an uncertainrandom value, and x is the solution vector, ξ1, ξ2, …, ξn are independent uncertain values with uncertainty distribution functions φ1, φ2, …, φn, having inverse functions, ω1, …, ωm are independent random values with probability distribution functions ψ1, ψ2, …, ψm. 12 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) in the case of hard constraint, the constraint must be satisfied for any implementation of random and uncertain values, and duplicate value has form: 0)(max ,   ,,xg . in the case of soft constraint, the constraint performs with guaranteed belief degree and probability. we will build duplicate of soft constraint. we interpret )x(g  ,, as uncertain value parameterized by random values. then in the first stage, where random variables are parameters, the duplicate of constraint 0,, )x(g  has the form:   ))x(g(m 0,, , where m is the degree of expert confidence (see appendix), α is given by the designer. if )x(g  ,, is continuous strictly increasing function with respect to ξ1, …, ξk and a strictly decreasing function with respect to ξk+1,…, ξn, then the inequality   ))x(g(m 0,, is equivalent to the inequality (theorem 2, appendix): 0)),1(),...,1(),(),...,(,( 11 1 11 1      nkkxg . at the second stage, we replace this inequality containing random values by the probability that this inequality is satisfied: ,)0)),1(),...,1(),(),...,(,(( 11 1 11 1      nkkxgp where β is the confidence level of probability specified by the designer. thus, we reduced constraint with uncertain-random value to probabilistic constraint. reducing the model of uncertain-random programming to a deterministic model of mathematical programming greatly simplifies the solution. reducing model of uncertainrandom programming to model of stochastic programming simplifies solution, but in general case they remain quite complicated. this is due to complexity of finding the analytical form of probabilistic criterion and constraint, and, in absence of one, complexity of numerical solution methods. in particular cases, deterministic equivalents of probabilistic criteria and constraints can be found [5, 6]. 3. calculation of aircraft weight parameters 3.1. models we formalize the one of the tasks of calculating weight parameters of passenger aircraft as optimization problem with constraints under conditions of mixed uncertainty based on the method of calculating the weight report of passenger aircraft [7]. in the original form, the functions for calculating the basic weight parameters of aircraft at the preliminary design stage are: pepayloadraeecppchairframefuel mmmmmmmmm  0 pepayloadpcaa mmvvm  )(0  wetsmairframe aqkm )1(  aaqsqss vvkkq lg)lg( 10  eqairframech kmm 3.0 )( 10 wet a engengengpp a v kktm   models of uncertain-random programming 13 copyright ©2019 assa. adv. in systems science and appl. (2019) where 0m is takeoff mass; airframem is airframe mass; chm is equipment mass for control and hydraulics; ppm is power plant mass; fuelm is fuel mass; ecm is weight of crew and equipment; raem is mass of board radioelectronic equipment; payloadm is commercial load; pem is passenger equipment; a is the average density of aircraft; av is aircraft volume; pcv is volume of the passenger compartment; mk is composites utilization coefficient; sq is specific weight per square meter of airframe; weta is area of washed surface of aircraft; 0qsk is statistical coefficient; 1qsk is statistical coefficient; eqk is coefficient taking into account production technology; eng is technological coefficient, which shows the ratio of weight of engine to maximum thrust at h = 0, v = 0 (h, v are height and speed, respectively); t is required thrust at h = 0, v = 0; 0engk and 1engk are statistical coefficients. to calculate the optimal design parameters we select optimization criteria that affect the fulfillment of specified technical requirements: operating costs and flight range. we minimize takeoff weight to reduce fuel consumption, as the main part of operating costs consists of fuel costs. on the other hand, we maximize the fuel supply, which increases the range of the aircraft. in this regard, the task of weight calculation is formalized and presented in form of following optimization model:        landingfuel0 a fuel0 mmm m m 8.0 ,500400 ,max,min  (1) based on the study of technical requirements, the designer choices the design parameters. the design parameters are weta , av , payloadm , pem , raem . the available information about the parameters determines the separation of the parameters into random and uncertain. statistics provide information on random parameters. experts provide information on uncertain parameters. the parameters with epistemic uncertainty are mk , eqk , pcv , ecm , a . the parameters with aleatory uncertainty are 0qsk , 1qsk , 0engk , 1engk , eng , t. the expression landingfuel0 mmm  8.0 determines the constraint on landing weight of aircraft, associated with length of run during landing. let 1 a  , 1 pcv , 1 ecm , 1 eqk , 1 mk be inverse uncertainty distribution functions of uncertainty for parameters a , pcv , ecm , eqk , mk ; 0qsk , 1qsk , 0engk , , 1engk eng , t be probability distribution functions of uncertainty for parameters 0qsk , 1qsk , 0engk , 1engk , eng , t. to solve this problem, as well as to compare the results and study the models, we construct two two-criterion models based on the single-criterion models proposed in section 2. in model a, for each criterion, select the "averaging" model (ee). in model b, for each criterion we choose the "quantile" model (qq). model a: 14 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019)              ],1,0[,))8.0(( ,)400( ,)500( ],[max],[min landingmlandingmlandingmlandingfuel0 aa aa fuel0 ppmmmmp m m meme      where m is measure of uncertainty (expert belief degree),  a  ,  a  , landingm  are levels of belief degrees that the inequalities in question are satisfied, landingm p is level of probability that belief degree in the implementation of inequality landingfuel0 mmm  8.0 will be greater than landingm  . to solve this task, first the designer sets  a  ,  a  , landingm  , . landingm p then the multicriteria optimization algorithm is applied. for each combination of design parameters varied in process of optimization, we calculate expected values of objective functions according to the formulas:    1 0 11 0 ,)))1()(((][  dmmvme pepayloadpcvaa ,)) )1(((][ 1010 1 0 1 0 tengengkengkqskqskpepayloadrae mr ecmppchairframe ' fuel dddddddmmm mmmmme        ;))1()(( 11 0 pepayloadpcvaa ' mmvm    );1(3.0 1    eqkairframech mm .))1(1( 1 wetsmkairframe aqm    model b:                  ],1,0[,))8.0(( ,)400( ,)500( ],1,0[,)][(sup ,max],[min landingmlandingmlandingmlandingfuel0 aa aa fuelmfuelmfuel fuelm 0 0m ppmmmmp m m pprmp rminf        where fuelm p is level of probability that rm fuel fuelm ][sup , fuelm is level of expert belief degree for quantile criterion. as the objective functions are used: pepayloadmvam mmvminf 0pc0a0m   ))1()((][ 11 0   , r, subject to models of uncertain-random programming 15 copyright ©2019 assa. adv. in systems science and appl. (2019) ;) )(( 1' 0 fuelmpepayloadrae fuelmecmppchairframe prmmm mmmmp     ;))()(1( 11' 0 pepayloadfuelmpcvafuelma mmvm    );(3.0 1 fuelmeqkairframech mm   wetsfuelmmkairframe aqm ))(1( 1   . note that the first criterion does not depend on random variables and is deterministic. we account for constraints for both models as follows. the constraint  aam  )500( is equivalent to 500)(1  aa   . the constraint  aam  )400( is equivalent to .400)1(1   aa   to check constraint on landing weight of aircraft, we calculate probability that belief degree in constraint performing will be greater than the specified level: landinglanding mmlandingfuel0 pmmmmp  ))8.0((  . this inequality is equivalent to inequality (theorem 2d, appendix): landingmlandingfuel0 pmmp(m  )08.0 , ,))1()(( 11 0 pepayloadlandingmpcvalandingma mmvm    ,)(1 0 pepayloadraelandingmecmppchairframefuel mmmmmmmm    ),(3.0 1 landingmeqkairframech mm   .))(1( 1 wetslandingmmkairframe aqm   to calculate weight parameters of aircraft, we use multicriteria genetic algorithm and statistical modeling. 3.2. methods and results we apply a genetic algorithm of multicriteria optimization with values  a  ,  a  , sbm  , sbm p given by the designer. we calculate the probabilistic characteristic of uncertain-random values based on the monte carlo statistical modeling method. as an example, we present the calculation algorithm for model a presented in section 3.1. for each combination of variable design parameters, we calculate the values of objective functions using following formulas:    1 0 11 0 ,)))1()(((][  dmmvme pepayloadva pca ,) )((][   dmmm 1mmmm n 1 me pepayloadrae n 1i 1 0 1 ecmppchairframe ' 0fuel      ,))()(( pepayload 1 pcva 1 a ' 0 mm1vm    ),1(3.0 1    eqkairframech mm ,))(( wets 1 mkairframe aq11m    16 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) ,lg)lg}({}{ aai1qsi0qss vvkkq  ).}{}({}{}{ 10 wet a iengiengiiengpp a v kktm   where 1 a  , 1 pcv , 1 ecm , 1 eqk , 1 mk are inverse uncertainty distributions for the parameters a , pcv , ecm , eqk , mk ; n is number of value combinations for parameter 0qsk , 1qsk , 0engk , 1engk , eng , t while we form the values of the parameters in accordance with the specified probability distributions; {●}i is the i-th element of value set for parameter {●} formed by statistical modeling. the constraint  aam  )500( is equivalent to 500)(1  aa   . the constraint  aam  )400( is equivalent to 400)1(1   aa   . to check the restrictions on the landing weight, an approximate probability value is calculated: ,].[)).((    n 1i fuel0landingmlandingfuel0 m80mi n 1 mm80mmp  where       иначе0 mm80m1 m80mi landingfuel0 fuel0 , ., ].[ , ,))()(( pepayloadlandingm 1 pcvalandingm 1 a0 mm1vm    ,)( pepayloadraelandingm 1 ecmppchairframe0fuel mmmmmmmm    ),(. landingm 1 eqkairframech m30m   ,))(( wetslandingm 1 mkairframe aq1m   the constraint is satisfied subject to landingmlandingmlandingfuel0 pmm50mmp  )).((  . data on uncertain and random parameters obtained from experts and statistics collection are approximated by uncertainty and probability distribution functions. normal distributions are most often used in the design of technical objects for approximation. we will conduct research for these types of distribution of uncertainty and probability. to generate the values of each random parameter, we used the standard normal distribution.                          2 2 ( ( ( 2 )( 1 2 1 )( ) ) ) x erfx     ,  x tdtexerf 2 0 1 )(  . for each epistemic parameter, we set the normal uncertainty distribution function: 1 ( ( ( 3 )( 1)(                           )e ) ) xe expx    ,           1 ln 3 )( ( ( 1 )e )e . models of uncertain-random programming 17 copyright ©2019 assa. adv. in systems science and appl. (2019) we used the nominal (mean) values of random and epistemic parameters as the )( and )e ( respectively. we set the standard deviations )(  and )e (  as a percentage of )( and )e ( . we investigate the influence of model parameters that reflect the characteristics of uncertainty and preferences of design maker. we present in fig. 1 and 2 the results of applying the optimization models a and b for different variants of the initial data: standard deviations of parameters )(  , )e (  , and values related to reliability landingm  , landingm p , 0m , fuelm , fuelm p ,  a  ,  a  . the upper line in fig. 1-3 is the result of applying the optimization model with deterministic values of uncertain and random parameters equal to their average values. the pareto-front, which is the result of the application of model a, practically coincides with the upper line, but narrows with increasing reliability. the values landingm  , landingm p , 0m , fuelm , fuelm p ,  a  ,  a  are equal and vary from 0.6 to 0.9 with a step of 0.1 (bottom four lines in fig. 1 and 2). set for all random parameters ))   (( 03,0   , for all epistemic parameters ))e e   (( 03,0 . then we obtained the pareto fronts shown in fig.1. fig. 1. pareto fronts for model b at ) )   (03,0 (   and )e e )   (03,0 (  (the parameters are random and epistemic) set for all random parameters )(05,0 )(     , for all epistemic parameters .05,0 (( ))e e   the parameters landingm p , fuelm p , landingm  , 0m , fuelm ,  a  ,  a  are equal and vary from 0.6 to 0.9 with a step of 0.1. then we obtained the pareto fronts shown in fig.2. 18 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) fig. 2. pareto fronts for model b at ) )   (05,0 (   and )e e )   (05,0 (  (the parameters are random and epistemic) the figures show that the pareto-fronts obtained as a result of performing optimization calculations using the model b with increasing landingm p , fuelm p , landingm  , 0m , fuelm ,  a  ,  a  are shifted to the region of the worst values of the objective functions. while ensuring high reliability of the solution (p and α close to 1) at fixed take-off weight the difference in fuel mass can reach more than 4000 kg. for comparison, consider the case where the parameters mk , eqk , pcv , ecm , a , 0qsk , 1qsk , 0engk , 1engk , eng , t are random. we use the following optimization model:                     .]1,0[,)8.0( ],1,0[,)400( ],1,0[,)500( ],1,0[,)( ],1,0[,)( ,max,min 2 0010 21 landingmlandingmlandingfuel0 aaa aaa fuelmfuelmfuel mm ppmmmp ppp ppp pprmp pprmp rr     let for all random parameters ))   (( 03,0   . the parameters 0m p , landingm p , fuelm p ,  a p ,  a p are equal and vary from 0.6 to 0.9 with a step of 0.1 (bottom four lines). then we obtained the pareto fronts shown in fig.3. the upper line in figures 3 is the result of applying the optimization model with deterministic values of uncertain and random parameters equal to their average values. models of uncertain-random programming 19 copyright ©2019 assa. adv. in systems science and appl. (2019) fig. 3. pareto fronts at ) )   (03,0 (   (all parameters are random) pareto fronts also move to the area of the worst values of the objective functions, but less than in fig. 1 and 2. the explorations show that the greater the degree of uncertainty and the greater the level of reliability, the more conservative obtained solutions. the obtained results do not contradict the experience and intuition of the designer, which confirms the adequacy of the models. 5. conclusion the models of the optimization problem with constraints under conditions of parametric mixed uncertainty are proposed. uncertain values introduced in the theory of uncertainty models parameters with epistemic uncertainty. under certain conditions (for rather wide class of functions), optimization models with mixed uncertainty are reduced to deterministic models or to stochastic programming models with probabilistic criteria. we formalize and solve the task of preliminary aerodynamic design in conditions of parametric mixed uncertainty using the proposed models. acknowledgements this work was supported by the russian foundation for basic research (grant 18-08-00822a) references [1] liu b. (2010). uncertainty theory: a branch of mathematics for modeling human uncertainty. springer-verlag, berlin. [2] eldred s., swiler t., tang g. (2011). mixed aleatory-epistemic uncertainty quantification with stochastic expansions and optimization-based interval estimation. reliability engineering and system safety. 96 (9), 1092-1113. 20 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) [3] yazenin a.v. (2016). osnovnye ponyatiya teorii vozmozhnostej [basic concepts of the theory of possibilities]. m., fizmatlit. [in russian]. [4] lockwood b., anitescu m., mavripilis d. (2012) mixed aleatory/epistemic uncertainty quantification for hypersonic flows via gradient-based optimization and surrogate models. 50th aiaa aerospace sciences meeting, aiaa-2012-1254. [5] vishnaykov v. b., kibzun a. i. (2006). deterministic equivalents for the problems of stochastic programming with probabilistic criteria. automation and remote control, 67 (6), 945–961. [6] geletu, a., kl¨oppel m. zhag h., li p. (2012). advances and applications of chanceconstrained approaches to systems optimization under uncertainty. international journal of systems science, 44 (7), 1209-1232. [7] kolokolova l.g. (1995). metod obobshchennykh modelej svojstv samoleta dlya etapa rannego proektirovaniya [method of generalized models of the properties of the aircraft for the early design stage]. tekhika vozdushnogo flota, № 5-6. [in russian]. appendix [1] definition 1: let γ be a nonempty set, let λ be a σ-algebra over γ, and let m be an uncertain measure. then the triplet (γ, l, m) is called an uncertainty space. a number m{λ} will be assigned to each event λ to indicate the belief degree with which we believe λ will happen. the measure of uncertainty m is dual, subadditive, the measure of the product of events is equal to the minimum of the measures of these events. definition 2: an uncertain variable is a function ξ from γ to the set of real numbers such that {γ |ξ(γ)∈ b}is an event for any borel set b of real numbers. definition 3: the uncertainty distribution of an uncertain variable ξ is defined by φ(x) = m {ξ ≤ x} for any real number x. definition 4: an uncertainty distribution φ(x) is said to be regular if it is a continuous and strictly increasing function with respect to x at which 0 < φ(x) < 1, and definition 5: let ξ be an uncertain variable with regular uncertainty distribution φ(x). then the inverse function φ−1 (α) is called the inverse uncertainty distribution of ξ. definition 6: the uncertain variables ξ1, ξ2, …, ξn are said to be independent if )()}({ 11 ii n k ii n i bmbm    .1)(lim,0)(lim   xx xx models of uncertain-random programming 21 copyright ©2019 assa. adv. in systems science and appl. (2019) for any borel sets b1, b2, …, bn of real numbers. theorem 1: let ξ1, ξ2, …, ξn be uncertain variables, and let f be a real-valued measurable function. then )...,,,( 21 nf  is an uncertain variable. definition 7: let ξ be an uncertain variable. then the expected value of ξ is defined by .}{}{][ 0 0 drrmdrrme      definition 8: let ξ be an uncertain variable with finite expected value e[ξ]. then the variance of ξ is v[ξ] = e[(ξ – e[ξ])]2. theorem 2: let f(ξ1, ξ2, …, ξn) be continuous, strictly increasing with respect to ξ1, ξ2, …,ξm continuous, strictly decreasing with respect to ξm+1, ξm+2,…, ξn. if ξ1, ξ2, …, ξn are independent uncertain variables with regular uncertainty distributions φ1, φ2, …, φn, respectively, then: a) ξ = f(ξ1, ξ2, …, ξn) is uncertain variable with inverse uncertainty distribution: )),1(...,),1(),(...,),(),(()(1 11 1 11 2 1 1      nmmf b) ,))1(),...,1(),(),...,(),((][ 11 1 11 2 1 1  dfe nmm     c) ,])[))1(...,),1(),(),...,(),(((][ 1 0 211 1 11 2 1 1  defv nmm     d) for any   [0, 1] m{f(ξ1, ξ2, …, ξn) < 0} > α is equivalent .0))1(),...,1(),(),...,(),(( 11 1 11 2 1 1      nmmf definition 9: a) the quantiles of the random variable ξ: supα[ξ ] = sup{r | pr { ξ > r} > α} – α-optimistic value, infα[ξ ] = inf{r | pr { ξ < r} > α} – α-pessimistic value,   [0, 1], b) the quantiles of the uncertain variable ξ: supα[ξ ] = sup{r | m{ ξ > r} > α} – α-optimistic value, infα[ξ ] = inf{r | m{ ξ < r} > α} – α-pessimistic value,   [0, 1]. definition 10: let (ω , a, pr) be a probability space and (γ, λ, m) be an uncertainty space. then the product (γ × ω, λ × a, m × pr) is called a chance space, where γ × ω = {(γ, ω)| γ ∈ γ, ω ∈ ω}. then the chance measure of an event θ of λ × a is: 22 g.s. veresnikov, l.a. pankova, v.a. pronina copyright ©2019 assa adv. in systems science and appl. (2019) x}dxθ}γ|(γ,ω)ω|m{γ{ω}сh{θ   1 0 pr . m{γ ∈ γ |(γ, ω) ∈ θ} is just the uncertain measure of cross section of θ at ω. since m{γ ∈ γ |(γ, ω) ∈ θ} can be regarded as a function from the probability space (ω , a, pr) to [0, 1], it is a random variable. thus the chance measure ch{θ} is just the expected value (i.e., average value) of this random variable. the chance-measure is monotonous, dual and subadditive. definition 11: an uncertain random variable is a function ξ from γ ×ω to the set of real numbers such that {(γ, ω)| γ ∈ γ, ω ∈ ω |ξ (γ, ω) ∈ b} is an event in l × a for any borel set b of real numbers. definition 12: let ξ be an uncertain-random variable. then its chance distribution is defined by φ(x) = ch{ξ ≤ x} for any real x. definition 13: the expected value of uncertain-random variable ξ is:     0 0 ][ x}dxch{dxx}ch{e  provided that at least one of the two integrals is finite. theorem 3: let η1, η2, … , ηm be independent random variables with probability distributions ψ1, ψ2, … , ψm, and let τ1, τ2, … , τn be independent uncertain variables with uncertainty distributions y1, y2, …, yn, respectively. if f(η1, η2, … , ηm, τ1, τ2, … , τn) is a measurable function, then has an expected value e[ξ] is: )(y)...dψ(ydψyyg),...,τ, τ, ...,ηf(ηe mmm mr nm 11111 ),...,(][  , where )],...,,,...,([) 111 nmm yyfe, ...,yg(y  − is the expected value of the uncertain variable ),...,,,...,( 11 nmyyf  for any real numbers ,,...,1 myy and is determined by y1, y2, …, yn. adv syst sci appl 2020; 01:104–113 published online at http://ijassa.ipu.ru/index.php/ijassa/article/view/860 comparison of some molecular structures by minimizing the comparison function natalia chernikova1*, eugeniy laneev1 1peoples’ friendship university of russia, moscow, russia abstract: the geometric difference in the relative arrangement of atoms in molecules with the same chemical formula significantly affects the properties of the crystal. for quantitative comparison of the spatial geometric structure of two molecules, we used the method of optimal superposition of molecules through the minimization of a certain comparison function by rotating the molecules after superposing their centers of mass. the search for the minimum of the comparison function over euler rotation angles is performed by the rosenbrock method. the assumed criterion for comparing molecules allows us to quantify the proximity in the spatial structure of molecules. the implementation of the method is shown by comparing the structure of eight molecules in four crystalline substances. keywords: conformation, euler angles, rotation matrix, optimization methods, rosenbrock method 1. introduction the paper deals with the mathematical problem that arises in structural chemistry, namely: the numerical study of the spatial structure of molecules with the same chemical formula, which may differ in the geometry of the mutual arrangement of atoms. the study of specific crystals was carried out in the framework of a mathematical model for comparing two molecules, studied in [1, 2]. the specified model is reduced to a comparison of two objects consisting of n ordered points with rigid geometry that behave like rigid bodies in r3. the comparison principle in this model is based on optimizing the alignment of these two objects by shifting and rotating. the optimal superposition of these objects involves the minimization over shifts and rotation angles of a certain comparison function of the geometry of the objects, which is the sum of the squares of the distances between the points of two objects with the same indices. it is proved that the minimum over shifts is achieved by superposing some characteristic points, conventionally called centers of mass of molecules. to minimize the [3–5] comparison function over rotation angles, the rosenbrock zero-order method [6] is used. the results obtained are used to study the geometry of molecules in real crystal structures. four organic crystalline substances (i, ii, iii, iv) are investigated. substances ii-iv consist of molecules with the same chemical formula and bond structure (same graphs), substance i is only a part of this graph. in addition, substances ii and iii contain symmetrically independent molecules (iia, iib, iic, iid, and iiia, iiib) that may have the same graph, but may differ in spatial structure, i.e., in the geometry of the arrangement of atoms. the problem arises of objectively comparing the geometry in the space of all these eight molecules and ∗corresponding author: cherni@list.ru comparison of some molecular structures... 105 quantifying the differences that exist in them, since the standard characteristics – bond lengths (interatomic distances) and valence angles (angles between valence bonds) do not always show differences in the geometry of molecules. in addition, it may be necessary to compare fragments of different molecules. for this purpose, [1, 2] provides a method for pairwise quantitative comparison of the geometry of molecules based on minimizing the comparison function by shifting and rotating the molecules. here is a brief description of the comparison method. 2. assumed mathematical model the study of molecules with the same chemical formula by comparing their geometries uses the concept of structure. here is the definition of the structure adopted in [1, 2] definition 2.1: a geometric structure, or in short – structure, will mean a rigid geometric construction of n ordered points in r3 with coordinates (xi, yi, zi), i = 1, ..., n , whose movement in r3 is performed like the motion of a rigid body. we assume that each i−th point of the structure has a weight coefficient wi > 0, and∑n i=1wi = w > 0. let ik, k = 1, ..., k – all indices corresponding to non-zero weight coefficients wi. then w = n∑ i=1 wi = k∑ k=1 wik > 0. (2.1) in accordance with [1, 2], we will associate the optimal criterion for superimposing two geometric structures with the minimum over shifts and rotation angles (euler angles [7]) of the comparison function of the form u(r1,0, r2,0, ϕ, θ, ψ) = n∑ i=1 wi|r1,i − r1,0 −q(r2,i − r2,0)|2, (2.2) where the vectors r1,i and r2,i determine the position of points in the first and second structure, the vectors r1,0 and r2,0 determine the offsets of the corresponding points of the first and second structure from the coordinate center, q = q(ϕ, θ, ψ) — the rotation matrix for euler angles (ψ — precession angle, θ — nutation angle, ϕ — self rotation angle): q = ( cosψ cosϕ− sinψ sinϕ cos θ − cosψ sinϕ− sinψ cosϕ cos θ sinψ sin θ sinψ cosϕ+ cosψ sinϕ cos θ − sinψ sinϕ+ cosψ cosϕ cos θ − cosψ sin θ sinϕ sin θ cosϕ sin θ cos θ ) (2.3) thus, the u function of the form (2.2) is the sum of the squares of the distances between “eponymous” points of two geometric structures with weights wi after matching points defined by vectors r1,0 and r2,0, with the coordinate center and rotation of the second structure relative to the first. by matching some characteristic points of two geometric structures, called “centers of mass” you can reduce the task of minimizing the comparison function over a full set of variables to finding the minimum of the u function over the rotation angles ϕ, θ, ψ. the following theorem is proved in [1]. theorem 2.1: the minimum of the u function of the form (2.2) is achieved at the point corresponding to copyright © 2020 assa. adv syst sci appl (2020) 106 n. chernikova, e. laneev “center of mass” of two geometric structures and is defined by vectors rj,0 = 1 w n∑ i=1 wirj,i, j = 1, 2. (2.4) so, the minimum of the function (2.2) corresponds to the shift of “centers of mass” of geometric structures defined by the formula (2.4) to the origin. the function (2.2) can now be considered as a function of rotation angles u(ϕ, θ, ψ) = n∑ i=1 wi|r1,i − r1,0 −q(ϕ, θ, ψ)(r2,i − r2,0)|2, (2.5) where r1,i, r2,i – given coordinates of geometric structures, and “centers of mass” r1,0 and r2,0 are calculated using the formula (2.4). the minimum of the function (2.5) is determined by the rotation angles ϕ, θ, ψ. let the minimum of the function u = u(ϕ, θ, ψ) of the form (2.5) be reached at the point (ϕ0, θ0, ψ0). in accordance with [1, 2], we will call the value s = (u(ϕ0, θ0, ψ0) w )1/2 . (2.6) the measure of proximity of two geometric structures. this value can be considered as a quantitative characteristic of the proximity of geometric structures, since it represents the average distance between points with the same indices in two structures after “superposing”. we will call two geometric structures equal if they can be superposed by points with the same indices by various displacements of these structures as solids. as shown in [1], this superposition is achieved by minimizing the comparison function (2.5) and corresponds to s = 0. we will call two structures approximately equal if s = ( 1 w min ϕ,θ,ψ u(ϕ, θ, ψ) )1/2 = (u(ϕ0, θ0, ψ0) w )1/2 6 s0, (2.7) where s0 is the specified value (in applications, it is determined by the specific practical situation). the inequality s 6 s0 will be called the criterion of proximity of structures. practical experience in studying the conformation of molecules, based on the results of comparing a significant number of structures [8, 9], resulted in the following conditional classification: s 6 s0 = 0, 1å – molecules are practically equal, 0, 1å < s 6 0, 2å – molecules are close, s > 0, 2å – molecules are different. note that if the minimum point is not unique and there is a point (ϕ1, θ1, ψ1) such that u(ϕ1, θ1, ψ1) = u(ϕ0, θ0, ψ0), then the value of s does not change and the non-uniqueness of the minimum point does not affect the proximity criterion of structures. the test for comparing the geometry of two geometric structures consists of three stages: 1) shifting the center of mass of each geometric structure to the coordinate center; 2) minimizing the function (2.5) over angles; 3) calculating the value of s using the formula (2.7) and inferring the proximity of the structures. additional information is provided by residual values ∆ri = |r1,i − r2,i|. when comparing two structures in accordance with the algorithm 1)-3), determination of the minimum of the comparison function u(ϕ, θ, ψ) over rotation angles after superposing the centers of mass of structures is performed numerically. copyright © 2020 assa. adv syst sci appl (2020) comparison of some molecular structures... 107 fig. 3.1. molecule c14h13no2 4,6-dimethyl-3-(phenylcarbonyl) pyridin-2 (1h) -one (i). 3. numerical study of crystal structures this section presents the results of the numerical study comparing chemically identical molecules in a crystal by minimizing the comparison function (2.5). in this case, we will consider molecules with an ordered structure ofn points-atoms as a geometric structure. the coordinates of the molecules in the crystal are obtained from x-ray analysis. in this study, the combined atoms in the molecules are assigned a weight of wi = 1. the remaining atoms are assigned a weight of wi = 0. the coordinates of the hydrogen atoms were not included in the calculations, since the accuracy of their determination is lower than that of other atoms and their presence is insignificant in the present task of superposing molecules, so they are assigned a weight of wi = 0. when comparing molecules, we will be interested in differences in the spatial structure of molecules with a single structural formula (graph). since the comparison function (2.5) is not convex, the rosenbrock zero-order method [6] is used to numerically minimize it. the program that implements the comparison algorithm 1)-3) uses the optimization library module [10] to minimize the comparison function (2.5) using the rosenbrock method. in this paper, we set the task of comparing the spatial structure of 4,6-dimethyl-3(phenylcarbonyl) pyridin-2 (1h) -one (i) molecules with the chemical formula c14h13no2 where n = 30 and in 3-benzoyl-4,6crystals-dimethyl-1-[2-(4-methylphenyl)-2-oxoethyl] pyridin-2 (1h)-one (ii, iii, iv) with the chemical formula c23h21no3, where n = 48. the composition, bonds and geometry of these molecules are schematically depicted in fig. 3.1 and fig. 3.2. the atomic coordinates were obtained by the method of x-ray analysis [11]. the characteristics of these substances in crystallography are described as follows: i c14h13no2 space group p21/c, z = 4 (molecule occupies 1 position); ii c23h21no3 space group pi, z = 8 (four independent molecules a, b, c, d); iii c23h21no3 · 12h2o (two independent molecules a, b); iv c23h21no3 · 12h2o space group c 2/c, z = 8 (molecule occupies one position). space group [12] describes all possible symmetries of an infinite number of periodically located points in three-dimensional space. the number z shows the number of molecules in the unit cell. in crystals i and iv, the molecules occupy one system of positions, that is, they are transformed into each other using symmetry operations of the space group. in crystals ii and iii, the molecules occupy independent positions, respectively, four (iia,iib,iic,iid) and two (iiia,iiib), that is, the molecules are not transformed into each other by symmetry operations of the space group of the crystal and therefore can be different in structure. copyright © 2020 assa. adv syst sci appl (2020) 108 n. chernikova, e. laneev fig. 3.2. molecule c23h21no3,3-benzoyl-4,6crystals-dimethyl-1-[2-(4-methylphenyl)-2-oxoethyl] pyridin-2 (1h)-one (ii, iii, iv) table 3.1. comparison of four independent molecules a, b, c, d in crystals ii the value of “residuals” ∆ri, characteristic s in crystals ii atom a− b a− c a− d b− c b− c b− d c− d n 1 0.20 0.06 0.03 0.23 0.05 0.18 0.08 c 2 0.19 0.16 0.04 0.34 0.08 0.15 0.20 c 3 0.06 0.15 0.05 0.20 0.02 0.03 0.19 c 4 0.05 0.10 0.06 0.09 0.03 0.10 0.16 c 5 0.02 0.08 0.05 0.08 0.08 0.07 0.11 c 7 0.05 0.06 0.01 0.04 0.04 0.04 0.04 c 8 0.19 0.03 0.03 0.17 0.13 0.16 0.02 c 9 0.03 0.10 0.03 0.08 0.06 0.08 0.08 o 10 0.27 0.22 0.04 0.47 0.07 0.29 0.23 c 11 0.07 0.24 0.04 0.30 0.14 0.04 0.26 o 12 0.22 0.35 0.08 0.56 0.30* 0.14 0.42 c 13 0.09 0.15 0.06 0.24 0.44* 0.13 0.12 c 14 0.30 0.35 0.06 0.65 0.69* 0.24 0.41 c 16 0.42 0.46 0.16 0.88 0.92* 0.26 0.63 c 18 0.37 0.41 0.15 0.78 0.96* 0.22 0.56 c 20 0.23 0.24 0.06 0.45 0.77* 0.18 0.30 c 22 0.12 0.09 0.07 0.19 0.53* 0.09 0.10 c 24 0.18 0.15 0.11 0.20 0.03 0.24 0.26 c 33 0.16 0.12 0.08 0.20 0.34* 0.15 0.12 o 34 0.31 0.28 0.15 0.39 0.56* 0.22 0.21 c 35 0.10 0.10 0.06 0.18 0.50* 0.14 0.14 c 36 0.04 0.29 0.04 0.27 0.40* 0.06 0.32 c 38 0.13 0.33 0.04 0.38 0.65* 0.12 0.30 c 40 0.24 0.14 0.06 0.37 0.91* 0.21 0.17 c 41 0.28 0.25 0.06 0.43 1.06* 0.24 0.26 c 43 0.19 0.24 0.08 0.35 0.88* 0.18 0.21 c 45 0.32 0.27 0.11 0.57 1.06* 0.35 0.22 s 0.21 0.23 0.08 0.40 0.08 0.18 0.27 copyright © 2020 assa. adv syst sci appl (2020) comparison of some molecular structures... 109 note that a comparison of bond lengths and bond angles does not reveal differences in the spatial structure of molecules. in eponymic bonds, the differences in length are no more than 0,02å, and in bond angles no more than 5◦. some differences are observed in individual torsion angles up to 19◦; however, how substantially this affects the general geometry of the molecules is not always clear. thus, it is not known in advance whether these molecules are similar in structure or whether their differences are significant. in this situation, there is a need for an objective, quantitative, comparison of the structure of molecules. for this purpose, the technique described in the previous section was applied to the study of these molecules. comparison of molecules was carried out by minimizing the comparison function (2.5). as already mentioned, the coordinates r1,i, r2,i, i = 1, ..., 48 (for substances ii iv), i = 1, ..., 30 (for substance i) for each pair of compared molecules were obtained on based on x-ray analysis data. all crystals (i iv) are centrosymmetric (both left and right molecules are present), therefore, to compare some pairs, one of the molecules was taken with inversion. the centers of mass were determined by the formula (2.4). the comparison function (2.5) was minimized numerically using the rosenbrock method over three euler angles after superposing the centers of mass of the two molecules being compared. the essential thing in our method of quantitative comparison of the spatial structure of molecules is the variation in the values of the weight factors wi. the ability to superpose molecules with separate fragments is achieved by setting the values of wi = 1 for compatible atoms and wi = 0 for other atoms. this allows you to change the center of gravity of the molecules, combine the molecules along the closest or most significant parts and thereby more clearly visualise certain differences in the structure of the molecules and their parts. as already mentioned, in our calculations we set the values wi = 0 for hydrogen atoms h. in table 3.1 we present the results of applying the comparison algorithm 1) 3) for matching 4 symmetrically independent molecules (a, b, c, d) that coexist in one crystalline substance ii. here, the residuals ∆ri = |r1,i − r2,i|, i = 1, ..., 48, i.e., the distances between the eponymic atoms in the two molecules under comparison are calculated after matching (at the point of minimum of the comparison function u ), as well as magnitude s – magnitude measure of proximity. results are shown for non-hydrogen atoms. the (*) indicates atoms for which wi = 0 was taken in this calculation. let us analyze the data obtained. in practice, only the pair of a-d molecules (s = 0.08å) are practically equal in their spatial structure; another pair of b− dmolecules (s = 0.18å) are close. other pairs of molecules have greater or lesser differences. the largest difference in the structure of molecules, is in the b− c pair (s = 0.40å). for this pair, another calculation was performed in which only for atoms of the central pyridine ring and atoms directly connected with it with indices i = 1, ..., 11, 24 and weight wi = 1. in this comparison, s = 0, 08å, which means the equality of these fragments and, in addition, this comparison highlighted the differences in the rotation of the phenyl rings. the result of this comparison is presented in fig. 3.3. as the calculations showed, two symmetrically independent molecules a and b in crystal iii (s = 0.04å), as well as pairs of molecules iv and iiia (s = 0.03å), iv and iiib (s = 0.04å) are almost identical. calculations comparing molecules from iv with four independent molecules from crystal ii indicated their significant differences in geometry. to identify differences for the same pairs of molecules, additional calculations were performed with varying weight factors. the ∆ri values for atoms with wi = 0 are also marked with (*). the calculation results are given in table 3.2. the results of the comparison point to parts of the molecules where there are differences and make us look for geometric characteristics to visually describe these differences. these characteristics were found as the differences in the angles between the planes (. 3.3) plane 1: through the atoms nc2c3c4c5c6, plane 2: through the atoms c13c14c16c18c20c22, plane 3: through the atoms (c35c36c38c40c41c43). copyright © 2020 assa. adv syst sci appl (2020) 110 n. chernikova, e. laneev fig. 3.3. the best combination of molecules iib and iic with wi = 1 for i = 1, ..., 11, 24. table 3.2. comparison of a molecule from crystal iv with a, b, c, d molecules in crystal ii. the value of “residuals” ∆ri, characteristic s atom iv &iia iv &iia iv &iib iv &iib iv &iic iv &iic iv &iid iv &iid n 1 0.43 0.04 0.38 0.03 0.41 0.02 0.41 0.02 c 2 0.31 0.01 0.31 0.03 0.38 0.05 0.29 0.04 c 3 0.22 0.02 0.16 0.04 0.33 0.02 0.17 0.05 c 4 0.35 0.06 0,40 0.04 0.43 0.04 0.29 0.02 c 5 0.44 0.06 0.47 0.06 0.47 0.08 0.40 0.05 c 7 0.45 0.01 0.41 0.04 0.40 0.06 0.44 0.04 c 8 0.49 0.06 0.46 0.13 0.50 0.03 0.51 0.06 c 9 0.57 0.10 0.57 0.09 0.55 0.13 0.58 0.07 o 10 0.43 0.04 0.50 0.03 0.40 0.06 0.44 0.05 c 11 0.13 0.08 0.17 0.06 0.22 0.17 0.13 0.13 o 12 0.08 0.29* 0.21 0.21* 0.36 0.50* 0.08 0.29* c 13 0.43 0.34* 0.34 0.29* 0.56 0.72* 0.45 0.34* c 14 0.27 0.17* 0.40 0.10* 0.48 0.69* 0.27 0.20* c 16 0.42 0.40* 0.27 0.27* 0.84 1.14* 0.32 0.26* c 18 0.99 0.92* 0.62 0.85* 1.40 1.66* 0.84 0.72* c 20 1.36 1.15* 1.22 1.21* 1.60 1.73* 1.32 1.04* c 22 1.06 0.85* 1.05 0.90* 1.15 1.25* 1.11 0.84* c 24 0.42 0.07 0.61 0.07 0.51 0.05 0.38 0.04 c 33 0.30 0.45* 0.31 0.54* 0.39 0.29* 0.36 0.37* o 34 0.35 0.46* 0.35 0.54* 0.60 0.09* 0.40 0.30* c 35 0.26 0.84* 0.20 1.05* 0.36 0.68* 0.27 0.80* c 36 0.18 0.83* 0.20 2.38* 0.35 0.93* 0.16 0.84* c 38 0.39 1.28* 0.35 2.54* 0.69 1.37* 0.41 1.37* c 40 0.78 1.80* 0.64 2.07* 0.88 1.53* 0.79 1.87* c 41 0.87 1.82* 0.67 3.37* 0.81 1.30* 0.83 1.81* c 43 0.58 1.36* 0.44 3.02* 0.55 0.88* 0.51 1.29* c 45 1.14 2.34* 0.96 2.65* 1.38 2.18* 1.22 2.50* s 0,60 0,05 0,54 0,06 0,72 0,08 0,59 0,05 table 3.3. angles (in degrees) between planes in molecules iia, iib, iic, iid and molecule in crystal iv planes iia iib iic iid iv planes 1 2 84.4 85.9 78.2 89.4 86.7 planes 1 3 86.5 87.8 71.7 85.7 71.2 planes 2 3 55.4 51.8 61.8 52.5 84.1 in addition, we present in table 3.4 the distances between the far points of the c18 and c45 molecules. the maximum distance is observed in molecules iic and iid, the minimum in iv. this illustrates the largest difference s = 0, 72å for the iic iv pair. fig. 3.4 shows the copyright © 2020 assa. adv syst sci appl (2020) comparison of some molecular structures... 111 table 3.4. distances (in angstroms) between the “far points” c18 and c45 of the iia, iib, iic, iid molecules in crystals ii and in the crystal iv molecule. iia iib iic iid iv distance c18c45 12.8 12.3 13.4 12.8 11.3 fig. 3.4. the best combination of molecules iib and iv with wi = 1 for i = 1, ..., 11, 24. result of combining iic iv with wi = 1 for the pyridine nucleus and the nearest atoms (nc2c3c4c5c7c9c11c24o2o10). the application of the comparison algorithm 1) 3) also works well for comparing molecules from i with the corresponding fragments of molecules iia, iib, iic, iid, iiia, iiib, iv. the results are presented in table 3.5. the table contains the calculated residuals ∆ri = |r1,i − r2,i|, i = 1, ..., 30. the results are shown for non-hydrogen atoms (except for h8 from crystal i). comparison of molecule i with fragments of molecules iv, iiia, iiib gave the same results, s = 0, 27å (there are some differences), which is explainable, since molecules iv, iiia, iiib are practically identical to each other in their geometry. the molecule in crystal i has the biggest differences with molecule iic, s = 0, 54å. to identify these differences, we use the same technique variation of weighting factors. the calculation by combining the planes of the pyridine rings and the directly bonded c and o atoms gave s = 0, 17å (parts of the fragments are close). this result indicates differences in the relative positions of the pyridine and phenyl rings. in this regard, the angles between plane 1 s = 0, 17å and plane 2 (c13c14c16c18c20c22) are calculated. for the molecule of their crystal i, the angle is 81, 1◦; the angle in iic is the most different from it, which is 78, 2◦. in addition, differences in the orientation of c = o bonds were found: differences in the torsion angles of c4c3c11o12, o12c11c13c22 and c3c11c13c22 are 19, 7◦, 19.9◦ and 12.5◦, respectively. 4. conclusion the calculations presented here show the effectiveness of the technique used to study the geometry of molecules in crystals using the comparison function. interesting facts were copyright © 2020 assa. adv syst sci appl (2020) 112 n. chernikova, e. laneev table 3.5. comparison of molecule i with corresponding fragments of molecules iia, iib, iic, iid, iiia, iiib, iv ((*) marked atoms for which wi = 0 is calculated) the value of “residuals” ∆ri, characteristic s atom i&iia i&iib i&iic i&iic i&iid i&iiic i&iiib i&iv n 1 0.28 0.27 0.46 0.05 0.24 0.19 0.20 0.20 c 2 0.37 0.39 0.45 0.07 0.36 0.30 0.31 0.31 c 3 0.20 0.23 0.26 0.04 0.19 0.15 0.13 0.13 c 4 0.33 0.33 0.45 0.04 0.29 0.15 0.14 0.15 c 5 0.37 0.37 0.45 0.08 0.34 0.24 0.24 0.24 c 7 0.17 0.24 0.21 0.10 0.18 0.13 0.15 0.14 () 8* 0.78* 0.68* 1.03* 0.52* 0.77* 0.77* 0.83* 0.69* c 9 0.45 0.53 0.44 0.10 0.44 0.44 0.44 0.45 o 10 0.70 0.71 0.83 0.10 0.61 0.56 0.56 0.57 c 11 0.38 0.37 0.55 0.06 0.37 0.17 0.18 0.19 o 12 0.78 0.75 0.96 0.73* 0.75 0.37 0.37 0.36 c 13 0.10 0.13 0.03 0.97* 0.05 0.05 0.05 0.04 c 14 0.31 0.41 0.15 1.19* 0.33 0.14 0.13 0.12 c 16 0.16 0.23 0.17 1.96* 0.19 0.23 0.23 0.23 c 18 0.55 0.55 0.63 2.57* 0.46 0.31 0.31 0.31 c 20 0.73 0.83 0.78 2.39* 0.73 0.25 0.26 0.25 c 22 0.40 0.47 0.41 1,55* 0.45 0,08 0.09 0.09 c 24 0.62 0.56 0.88 0.08* 0.49 0.28 0.27 0.27 s 0.45 0.47 0.54 0.17 0.42 0.27 0.27 0.27 obtained 3-benzoyl-4,6-dimethyl-1[2(4-methylphenyl) -2-oxoethyl] pyridin-2 (1h) one molecules have the same conformation in different crystals of iii and iv, and different conformation in one crystal ii. acknowledgements the above calculations were carried out by the authors using the program comparison in the language c], written by n.v. bakhtadze, to whom the authors express their sincere gratitude. the work was supported by the russian foundation for basic research (grant 17-5112064, grant 18-01-00106, grant 18-01-00590). references 1. laneev, e.b. & chernikova, n.yu. (2019). comparing the spatial structure of molecules by minimizing a comparison function, computational mathematics and mathematical physics, 59(1), 128-135. 2. chernikova, n.yu.& laneev, e.b. (2020). on method of nonlinear optimization in the study of molecules spatial structure yugoslav journal of operations research, 30(1), 91-99. 3. izmailov, a.f., & solodov, m.v. (2003) chislenniye metodi optimizatsii [numerical methods of optimization]. moscow, russia: fizmatlit, [in russian]. 4. arutyunov, a.v.& tynyanskii, n.t. (1985). the maximum principle in a problem with phase constraints. soviet journal of computer and system sciences, 23, 28-35. 5. arutyunov, a a.& mishchenko, a.s. (2019). a smooth version of johnson’s problem on derivations of group algebras, sb. math., 210(6) 756-782 6. himmelblau, d.m. (1972) applied nonlinear programming. new york, usa: mcgraw-hill book company. 7. lurie, a.i., (2002) analytical mechanics. berlin, germany: springer. copyright © 2020 assa. adv syst sci appl (2020) comparison of some molecular structures... 113 8. chernikova, n.yu., lavut, e.e. & zorkii, p.m. (1979) comparing the structure of symmetrically independent molecules in crystals of coordination compounds, russian journal of coordination chemistry, 5(8), 1265–1270. 9. lavut, e.e., zorkii, p.m., & chernikova, n.yu. (1982) contact conformers in crystals of coordination compounds, 2020 journal of structural chemistry 22(5), 715-718. 10. kniaz k. optimization.net, url: http://www.kniaz.net/software/rosnm.aspx. 11. feklicheva, e. m., rybakov, v. b., babaev, e. v., chernikova, n. yu. & ofitserov, e. n. (2020) the structure of products of phenacylation and subsequent (re)cyclizations of 3benzoyl-4,6-dimethylpyridin-2(1h)-one according to x-ray structural analysis, arabian journal of chemistry (in press). 12. international tables for crystallography. volume a: space-group symmetry / edited by m. i. aroyo. international union of crystallography, 2016. isbn 978-0-470-97423-0. copyright © 2020 assa. adv syst sci appl (2020) introduction assumed mathematical model numerical study of crystal structures conclusion adv syst sci appl 2020; 03:105–112 published online at https://ijassa.ipu.ru. estimating the response time of a cloud computing system with the help of neural networks anastasia v. gorbunova1*, vladimir m. vishnevsky1 1v.a. trapeznikov institute of control sciences of russian academy of sciences, moscow, russia abstract: the article presents a new approach to assessing the average response time of a cloud computing system and its dispersion. a fork-join system or a system with request splitting was chosen as a queuing model, and artificial neural networks were used as a method for estimating a variable of interest. the analysis showed that the estimates obtained were more accurate than those previously known. besides, the proposed approach allows expanding the analysis of the cloud system to the case of a model with a non-poisson input stream and non-exponential service time, as well as obtaining estimates for a larger number of performance indicators of the cloud system, which was not previously possible. keywords: cloud computing, parallel computing, queuing system, parallel service of requests, average response time, multilayer perceptron, artificial neural networks, machine learning methods 1. introduction cloud computing is a technology that allows a remote user to access various services that, as a rule, have significant computing resources upon request via the internet [2]. at the same time, even despite the presence of powerful computing infrastructure, the task of ensuring the required level of quality of the services provided and, at the same time, a logical interest in reducing the overall load on the equipment continue to be relevant for cloud providers [1, 9, 10, 23]. this is due to the growing popularity of the services they provide for many objective reasons, as well as the understandable need to reduce costs due to, for example, increased energy efficiency. one of the possible solutions to this problem is parallel data processing, which could be accomplished due to the use of virtualization technology, which results in a decrease in response time, and, accordingly, in maintaining or even improving the quality of the services provided in the cloud. thus, a cloud center consisting of several physical machines that can be used together to process complex user requests served in the order they arrive in the system is considered [5]. one of the most important indicators of the functioning of the cloud system is the response time, the evaluation of which has a large number of publications, for example, works [1,5,6,9,10,23] devoted to. nevertheless, the study of this parameter is still in demand due to the relevance of the practical task, and as a consequence of the need to improve the quality of existing estimates of this characteristic. the article is organized as follows: section 2 describes the mathematical model of the cloud center, section 3 discusses the possibility of applying a new approach to the analysis ∗corresponding author: avgorbunova@list.ru 106 a.v. gorbunova, v.m. vishnevsky of average response time, namely, artificial neural networks (anns), provides a numerical example, and finally, section 4 highlights the prospects for further research. 2. the mathematical model of a cloud computing system a queuing system (qs) of the fork-join type is considered as a model of a cloud system. systems of this type are also called “qs with parallel service of requests” [7, 8]. fork-join systems are models for many real physical systems in which the parallelization of the tasks is performed. for example, using the method of datagrams in computer networks, when a message arrives at the source node, it is divided into packets that are transmitted by different routes to the address node, in the buffer of which the message is assembled. other examples are the processes of assembling orders at the warehouse or the processes of functioning of high-performance applications in the production, medical, financial, scientific or any other areas, based on the concept of distributed or parallel computing [11, 15]. features of the functioning of the qs with the splitting of requests are as follows (fig. 2.1): 1) at the time the request enters the system, it is instantly split into k (k > 2) smaller components, i.e., sub-requests, each of which, in turn, becomes serviced by the device (if it is busy) or immediately begins to be serviced if the corresponding device is free. it is considered that each sub-request has its type, which must correspond to the number of the device on which it will be serviced; 2) after the service is ended the sub-request falls into the so-called synchronization buffer and remains there until all related sub-requests, that is, the sub-requests originally belonging to one request, finish their service. then the whole request is instantly assembled and only after this the request is considered served and may leave the system. since modeling a cloud center in the context of solving complex computational problems is discussed in this case, virtual machines (vms) will act as devices in the fork-join models. at the time of receipt, a complex user request is divided into k smaller components, each of which is processed on the corresponding vm. then it enters the synchronization buffer, where it remains until the last of the subtasks is processed, after which the user request is considered served. fig. 2.1. the model of the cloud center system of the fork-join type note that other sub-request servicing schemes are known. for example, when devices are blocked until all parts of the request are serviced or the sub-requests are indistinguishable in the sense that any k sub-requests are required to assemble the request, etc. a more detailed overview can be found, for example, in [7, 20]. one of the main performance indicators of any queuing system is its response time. the fork-join system in this sense is no exception. therefore, this concept is specified in the context of splitting requests. since the request is considered to be served only at the moment when the last sub-request is serviced, to calculate the residence time of the request in the copyright © 2020 assa. adv syst sci appl (2020) estimating the response time of a cloud computing system... 107 system, taking into account the fact that the moments of the appearance of all its sub-requests in the system coincide, it is sufficient to determine the maximum of all the times of the residence of its sub-requests. however, the task of calculating the exact value of the average response time, even despite the poisson nature of the incoming stream and the exponential service time on all devices in the case of splitting into more than 2 sub-requests, remains unresolved. the expression in closed form exists only for k = 2 [14], and for k > 2 only approximations of various degrees of accuracy were obtained [14, 21, 22]. this is explained by the complexity of the analysis of the residence times of the parts of the request in the system because of the existing relationship between them due to the general moments of receipt. the residence times of subrequests in the system are positively associated random variables. therefore, their maximum is stochastically no greater than the maximum of independent random variables with the same distribution. below are formulas that estimate the average response time in the system with splitting requests in the case of the same service intensities on all devices µk = µ, k = 1, k, and the incoming stream intensity equal to λ [14, 21, 22]. e[wk ] ≈ [ hk h2 + 4 11 ( 1− hk h2 ) ρ ] 12− ρ 8 1 µ− λ , (2.1) e[wk ] ≈ [ hk + ( k∑ i=1 ( k i ) (−1)i−1 i∑ m=1 ( i m ) (m− 1)! im+1 −hk ) λ µ ] 1 µ− λ , (2.2) e[wk ] ≈ 1 µ [ hk + ρ 2(1− ρ) ( k∑ k=1 1 k − ρ + (1− 2ρ) k∑ k=1 1 k(k − ρ) )] , (2.3) e[wk ] ≈ 1 µ− λ hk , (2.4) e[wk ] ≈ 1 µ− λ ( 1 + k − 1√ (2k − 1) ) , (2.5) where hk = ∑k i=1 1/i is the partial sum of the harmonic series, and ρ = λ/µ < 1 is the load factor. the first two of the above formulas are obtained due to the combination of the analytical approach with the empirical one, which consists in observing the behavior of the random variable under investigation through simulation and the subsequent use of the results of this observation in the preparation of the corresponding estimates. the third formula is derived by interpolation of high and weak input loads. approximation (2.4) is the result of analyzing the average response time for k independent m |m |1 qss operating in parallel, which is not only natural but also objectively justified assumption, since the expressions for the marginal probabilities of the number of k-type sub-requests (k = 1, k) in the original fork-join system coincide with the expressions for stationary probabilities in the m |m |1 system [5]. finally, the last expression (2.5) from the above is the upper limit of the average value of the maximum of order statistics for independent exponentially distributed parameterized random variables with (µ− λ). in addition to estimating the average response time, the approximations for variance, and higher moments of the studied value are worth noticing. these numerical characteristics make it possible to formulate a complete picture of the response time behavior. however, there is no longer such a variety of analytical expressions, and it is desirable to improve the existing estimates. so, in [6], an expression is presented for the variance of the average response time copyright © 2020 assa. adv syst sci appl (2020) 108 a.v. gorbunova, v.m. vishnevsky for both identical µ and different µk parameters of exponential service times, k = 1, k d[wk ] ≈ k∑ l=1 2 (µl − λ)2 − ∑ 16l